Compare commits

...
Author SHA1 Message Date
arkonandClaude Opus 4.7 e8a809ea80 chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 03:36:56 +02:00
Ark0NandClaude Opus 4.7 8006cc5db3 fix: allowlist opusContext1mEnabled in SettingsUpdateSchema (#78)
Same root cause as the thinkingEffort fix in #73: the schema is
.strict(), so unknown keys in PUT /api/settings are rejected with
INVALID_INPUT and the toggle never persists. Verified live:
pre-fix returned {"errorCode":"INVALID_INPUT"}, post-fix accepts.

The frontend has been reading and writing this key for a while
(settings-ui.js:336, :1137; session-ui.js:331), so saves were
silently failing — users never noticed because the load path
falls back to false on missing keys.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 03:29:48 +02:00
arkonandClaude Opus 4.7 0ded279b55 chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 03:25:53 +02:00
Tenggan ZhangandTeigen aa5724c390 feat: improve Resume Conversation UX on mobile (#77)
Default layout was a single nowrap row with path + date + size, which on
narrow screens truncated both the first prompt and the directory suffix
(where project names actually live). The /Users/ home shorthand was also
never applied on macOS.

Changes:
- Title now uses 2-line clamp so more of the first prompt is visible.
- Subtitle resolves workingDir against known cases: exact match shows
  "#caseName", subpath shows "#caseName/sub", otherwise falls back to
  basename. Case labels are styled distinctly.
- Normalize both /home/<user>/ and /Users/<user>/ prefixes to "~/".
- Each history item gets a "..." toggle that expands an in-place detail
  panel with the full prompt, full path, timestamp, size, and short
  session id. Collapses back on second click; clicking the card body
  still triggers resume.

Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
2026-04-28 03:14:56 +02:00
Tenggan ZhangandTeigen f21df2a9fb fix: eye icon follows /clear to the new Claude conversation (#76)
Interactive Claude CLI never emits session_id on stdout, so the
Session's _claudeSessionId stayed pinned to the pre-/clear jsonl and
the last-response viewer kept showing the old conversation.

Two complementary update paths:

- Session.adoptClaudeSessionId() — public setter mirroring the existing
  no-op-if-same guard. Called from POST /api/hook-event when Claude Code
  hooks carry data.session_id (works once hooks are configured).

- /api/sessions/:id/last-response now resolves the active id from
  ~/.claude/history.jsonl before reading the transcript. This is the
  only source-of-truth that does not require hooks, and we intentionally
  don't write hooks into arbitrary user repos.

History scan filters out sessionIds held by other Codeman sessions in
the same cwd, and validates via jsonl mtime to avoid inheriting a dead
prior session's id.

Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
2026-04-28 03:03:11 +02:00
e549e15cb8 feat(response-viewer): ASCII diagram wrap toggle, mobile code blocks, chrome-stripping fallback (#75)
* fix: restore clear message separation + proper table layout in response viewer

* fix: capture Claude CLI's real session ID + robust ANSI/CLI-chrome stripping in response viewer fallback

Session constructor seeded _claudeSessionId with Codeman's session.id as a
placeholder, and the message-driven update was gated on !_claudeSessionId —
meaning Claude CLI's actual session UUID was never adopted. This broke
/api/sessions/:id/last-response JSONL lookups, silently falling through to
the terminal-buffer path whose ANSI regex missed \x1b[>c / \x1b[>q queries.

- session.ts: update _claudeSessionId whenever a message's session_id differs
  from current (covers placeholder and stale-resume cases)
- app.js: extract _cleanTerminalBuffer with proper CSI regex (param bytes
  0x30-0x3F now covers > ? < =) plus a chrome filter for status bar,
  progress bar, spinner, shell prompt, and hint lines

* fix: wrap regular code blocks on mobile, keep ASCII diagrams rigid with scroll hint

* feat: add per-block wrap toggle on ASCII-diagram code blocks

* fix: wrap by default, pin toggle button outside scroll container

* fix: narrow diagram detection to box-drawing + block elements only

* feat: show last-response viewer eye icon on desktop too

The response viewer button was mobile-only via a display:none default with a
mobile.css override. Flip the default to inline-flex and drop the override so
the eye icon appears in the header on every form factor — desktop users get
the same quick "Last Response" pane as mobile.

* fix(response-viewer): restore HTML sanitizer + fix undefined `src` in _renderMarkdown

- `_renderMarkdown` referenced an undefined `src` (should be `text`),
  causing a ReferenceError on every markdown render. The try/catch
  swallowed it, so the new table-wrap and ASCII-diagram features
  never actually ran — output silently fell through to plain-text.
  app.js is excluded from ESLint, so this wasn't caught at lint time.
- `_sanitizeHtml` was removed when refactoring the response viewer,
  leaving `marked.parse()` output going straight into `innerHTML`
  without sanitization (XSS regression vs. master). Restored the
  helper and re-applied it before any post-processing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
Co-authored-by: arkon <arkon.85@hotmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 02:58:42 +02:00
Ark0N d07b59db4e Merge pull request #74 from TeigenZhang/refactor/envoverrides-tmux-export
refactor: pass envOverrides via tmux export instead of disk write
2026-04-28 02:45:11 +02:00
arkonandClaude Opus 4.7 a5a7e0c94c Merge master into refactor/envoverrides-tmux-export
Resolved conflict in src/web/public/session-ui.js by keeping this
PR's buildEnvOverrides() helper — it already covers both
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS (this PR) and
CLAUDE_CODE_EFFORT_LEVEL (added in #73), so the master-side inline
block is fully replaced.

Also fixed test/session-manager.test.ts MockSession to add a
getEnvOverridesForPersist() stub — without it,
SessionManager.updateSessionState's new call breaks 19 tests with
"TypeError: session.getEnvOverridesForPersist is not a function".

Verified: typecheck, lint, format:check, build, and
test/{session-manager,session-state,tmux-manager,tmux-restart-recovery}.test.ts
all pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 02:43:31 +02:00
Ark0N 79d7117e6d Merge pull request #73 from TeigenZhang/feat/thinking-effort
feat: thinking effort setting for new sessions (with xhigh/max)
2026-04-28 02:21:57 +02:00
arkonandClaude Opus 4.7 3cf486730b fix: allowlist thinkingEffort in SettingsUpdateSchema
Without this, PUT /api/settings rejects the new field with
INVALID_INPUT (schema is .strict()), so the dropdown's value
never persists.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 02:20:11 +02:00
Tenggan Zhang 996b096849 fix: prevent vertical scroll on mobile keyboard accessory bar (#72)
Thank you @TeigenZhang for the clean mobile fix!
2026-04-28 02:15:00 +02:00
ffa7fcf839 fix(tmux-manager): use '|' separator in reconcileSessions (#71)
* fix(tmux-manager): use '|' separator in reconcileSessions

Under non-tty execution contexts (launchd on macOS, systemd without TTY),
tmux emits '\t' in FORMAT strings as the literal two characters `\` + `t`
rather than as a tab. The parser's `line.indexOf('\t')` (a real tab char)
therefore never matches, `activeSessions` stays empty, `reconcileSessions`
returns `alive: []` / `discovered: []`, and `cleanupStaleSessions()` wipes
every entry in `state.json` — even though the underlying tmux sessions are
still alive. On the next startup the user sees an empty session list.

The bug reproduces reliably when codeman is launched via a user LaunchAgent
or a systemd unit without `TTYPath`. Interactive `npm run dev` hides it
because tmux's format parser does interpret `\t` when stdout is a TTY.

Fix: use `|` as the separator. tmux passes it through verbatim in every
environment, and `|` is not a valid tmux session-name character so it
cannot collide with the codeman-<uuid> / claudeman-<uuid> naming scheme.

* test(tmux-manager): cover parsePaneList separator contract

Extract the inline pane-list parser from `reconcileSessions` into an
exported `parsePaneList()` helper plus `PANE_LIST_SEP` / `PANE_LIST_FORMAT`
constants, so the '|' separator contract can be unit-tested directly.

The new tests lock in:
- Well-formed parsing into name -> pid Map
- Empty / blank-line / missing-separator handling
- Non-numeric pid and empty-name rejection
- A literal `\t` (backslash + t) in the input is NOT treated as a
  delimiter — guards against the launchd/systemd regression that
  motivated PR #71.
- Splitting on the first separator only.

No behavior change in `reconcileSessions`; the body now delegates to the
helper.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
Co-authored-by: arkon <arkon.85@hotmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 02:11:11 +02:00
Teigen a1c69f7405 refactor: pass envOverrides via tmux export instead of disk write
CLAUDE_CODE_EFFORT_LEVEL (and any CLAUDE_CODE_* / OPENCODE_* key) now flows:
  UI dropdown → POST /api/sessions { envOverrides }
             → new Session({ envOverrides })
             → this._envOverrides
             → tmux-manager.buildEnvExports appends `export KEY=<shellescape(VALUE)>`

Previously the API wrote envOverrides to <case>/.claude/settings.local.json, which
created stale state (UI dropdown disagreeing with disk) and polluted user project
directories. Now envOverrides are ephemeral spawn-time state, preserved across
respawnPane cycles via this._envOverrides and across server restart via
SessionState.envOverrides in state.json.

Also removes the now-unused updateCaseEnvVars import from session-routes.ts.
2026-04-24 09:49:52 +08:00
Teigen 534899bc2b feat: add xhigh effort option and /effort max mobile shortcut
Add XHigh option to Thinking Effort dropdown (between High and Max),
and add a Max quick button to the mobile keyboard accessory bar that
sends /effort max as a slash command.
2026-04-24 09:48:53 +08:00
Teigen 03d91ffddd feat: add thinking effort setting for new sessions
Allow configuring CLAUDE_CODE_EFFORT_LEVEL (low/medium/high/max) from
Settings → Claude Permissions. Applied as envOverride on session creation.
2026-04-24 09:48:18 +08:00
arkonandClaude Opus 4.7 6280998bd8 chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 20:22:30 +02:00
arkonandClaude Opus 4.7 02e2f3e8b5 refactor: remove dead code and narrow internal exports (knip sweep)
Knip-driven cleanup. All changes verified with tsc --noEmit, lint, and
build.

Removed (zero consumers):
- VERIFICATION_PROMPT constant + its barrel re-export
- createInitialOrchestratorPersistState factory
- transcriptWatcher singleton export
- createAnsiPatternFull / createAnsiPatternSimple factories
- TimerInfo interface + unused AiCheckResult/AiPlanCheckResult imports
  in respawn-controller.ts
- 35 unused Zod z.infer \`*Input\` types in schemas.ts
- Dead re-exports: SessionMode from session.ts, AuthSessionRecord from
  web/ports/index.ts, EnhancedPlanTask/CheckpointReview from
  ralph-tracker.ts, 7 unused entries in utils/index.ts
- 14 event/config interfaces that lived only as JSDoc hints (no TS type
  position usage): Session/Respawn/RalphLoop/RalphTracker/
  SessionManager/SessionAutoOps/Subagent/TaskQueue/TaskTracker/
  TranscriptWatcher/Image/OrchestratorLoop Events + RespawnPreset +
  SessionOutput

Narrowed to module scope (kept but no longer exported):
- buildPermissionArgs in session-cli-builder.ts
- 28 type/interface declarations used only within their own file:
  Ai{Idle,Plan}Check{Config,State}, BashToolParser{Events,Config},
  FileStream/CreateStream{Options,Result}, PlanSubagentEvent,
  SubagentCallback, RalphLoopConfig, RalphLoop{Events,Options},
  ActiveTimerInfo, DetectionStatus, ActionLogEntry, AutoOpsCallbacks,
  TunnelStatus, Timer/LRUMap/StaleExpirationMap Options, AuthState,
  SessionListenerDeps, SseStreamManagerDeps, and 8 more

Docs: CLAUDE.md advice for global-regex `lastIndex` now points to the
remaining `execPattern()` helper instead of the deleted factories.

Knip delta: unused files 42→0, unused exports 161→16, unused types 92→0.
The 16 remaining exports are a mobile-test helper toolkit intentionally
kept for upcoming tests.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:57:00 +02:00
arkonandClaude Opus 4.7 41300f0a34 chore: add knip config for dead-code detection
knip.json declares the real entry points (scripts, tests, Remotion roots)
and the devDeps invoked only as external CLIs (esbuild for build,
agent-browser/remotion via npx) so future scans surface only true
findings.

Add \`npm run knip\` as the canonical invocation.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:56:31 +02:00
arkonandClaude Opus 4.7 adbc083426 chore: remove dead test and script files
Dead-code sweep via knip. All files below have zero importers and were
leftovers from the local-echo overlay exploration or duplicated by files
under test/mocks/.

Deleted:
- test/respawn-test-utils.ts (728-line duplicate of test/mocks/*)
- test/input-echo-test.mjs
- test/local-echo-*.mjs (7 files)
- test/manual/*.mjs (10 files; dir removed)
- scripts/remotion/components/TerminalScreen.tsx (unused Remotion demo)

Also cleaned stale JSDoc references to the removed
respawn-test-utils.ts in test/mocks/mock-session.ts and
test/mocks/test-helpers.ts.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:56:22 +02:00
arkonandClaude Opus 4.7 3754bcd1aa docs: tighten CLAUDE.md — update counts and CI line
- CI line: include `check:lockfile` step now run in workflow
- Frontend: app.js is ~2.9K lines (was ~2.8K)
- Types: document `src/types/index.ts` as the barrel (14 domain files) plus `src/types.ts` root re-export
- API routes: updated handler counts (128 total; sessions 27, cases 9)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:55:56 +02:00
arkonandClaude Opus 4.7 f2f909ca9c docs: tighten CLAUDE.md — fix counts, remove footer redundancy
Fix stale counts (types 14 to 15, SSE events ~118 to ~120). Remove
redundant footer sections (References list duplicated inline citations;
Common Workflows bullets were self-evident or already stated; Tunnel and
Memory Leak Prevention folded into neighboring sections). 251 to 234 lines.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:14:16 +02:00
arkonandClaude Opus 4.7 1c3f2f6571 docs: tighten CLAUDE.md and archive 22 completed plan docs
CLAUDE.md: fix stale counts (types 14 to 15, SSE events ~118 to ~120),
remove redundant footer sections (References list duplicated inline citations;
Common Workflows bullets were self-evident or already stated; Tunnel/Memory
Leak Prevention folded into neighboring sections). 251 to 234 lines.

Move 22 completed implementation/phase/audit plans to docs/archive/ via
git mv so history is preserved. Living reference docs remain in docs/.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:14:07 +02:00
arkonandClaude Opus 4.7 8da1bdf690 chore: prevent package-lock.json version drift permanently
Makes the drift that PR #70 caught impossible to repeat:

- `version-packages` script now runs `changeset version && npm install
  --package-lock-only && check-lockfile-sync`, so the lockfile is always
  regenerated and verified as part of consuming a changeset
- New `scripts/check-lockfile-sync.mjs` compares package.json#.version against
  package-lock.json's root and packages[""] version fields (npm ci does not
  enforce these, which is why the prior drift slipped through CI)
- CI now runs `npm run check:lockfile` on every push/PR — any future drift
  fails the build before merge
- COM workflow in CLAUDE.md collapsed back to a single release-bump step now
  that lockfile sync is automatic

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:01:51 +02:00
arkonandClaude Opus 4.7 a93325b312 fix: sync package-lock.json to 0.6.0 and document lockfile step in COM workflow
package.json has been at 0.6.0 since release, but package-lock.json stayed at
0.3.11 because `npm run version-packages` (changesets) does not regenerate the
lockfile. This left `npm ci` broken against the committed state.

Also adds step 4 (`npm install --package-lock-only`) to the COM workflow in
CLAUDE.md so future releases keep the lockfile in sync automatically.

Credit to @Matt2012 (#70) for catching the lockfile drift.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:00:10 +02:00
arkonandClaude Opus 4.7 2d03e4efc9 docs: update CLAUDE.md and README for clipboard API and tab shortcuts
Reflect changes from PRs #65–#68: bumped route/SSE counts, added
Ctrl+Shift+{/} (tab reorder), Alt+1-9 (tab switch), Ctrl+Shift+V
(voice), and POST /api/clipboard to the keyboard and API references.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-18 04:16:36 +02:00
arkonandClaude Opus 4.7 ab7c502c2a chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-18 04:13:26 +02:00
Ark0N 546bbcbe7c Merge pull request #68 from aakhter/feat/clipboard-api
feat: add clipboard API for remote browser clipboard access
2026-04-18 04:10:31 +02:00
Ark0N 774d5ff321 Merge pull request #67 from aakhter/feat/tab-badges
feat: improve active tab visibility and add Alt+N number badges
2026-04-18 04:08:09 +02:00
Ark0N 98ceb5da1d Merge pull request #66 from aakhter/feat/tab-reorder-shortcuts
feat: add Ctrl+Shift+{/} to reorder session tabs
2026-04-18 04:07:01 +02:00
Ark0N 34fb5e49f8 Merge pull request #65 from aakhter/fix/android-shift-double-char
fix: prevent double character input on Android Shift+key
2026-04-18 04:05:21 +02:00
arkonandClaude Opus 4.7 002cf81b1e chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-18 04:02:02 +02:00
Aamer Akhter 9b4aab2502 feat: add clipboard API for remote browser clipboard access
POST /api/clipboard with { text } broadcasts to all connected browsers
via SSE. Browser attempts navigator.clipboard.writeText, falling back
to a modal with manual copy button if blocked.

Enables remote clipboard workflows: CLI tools can push text to the
user's browser clipboard across the network.
2026-04-12 10:07:15 -04:00
Aamer Akhter 829c797726 feat: improve active tab visibility and add Alt+N number badges
- Active tab: bright green border with color-matched glow per session color
- Tab number badges (1-9) showing Alt+N shortcut hints
- Badges update on tab reorder and re-render
2026-04-12 10:04:34 -04:00
Aamer Akhter 85da3bb898 feat: add Ctrl+Shift+{/} keyboard shortcuts to reorder session tabs
Matches WezTerm/terminal emulator tab reorder conventions. Uses the
existing sessionOrder array and saveSessionOrder() for persistence.
2026-04-12 10:02:12 -04:00
Aamer Akhter 29b2653801 fix: prevent double character input on Android Shift+key
On Android tablets, pressing Shift+A produces "AA" because the input
event listener re-sends characters that xterm already processed via
its keydown handler. Track keydown timestamps and skip input events
that fire within 50ms of a handled keydown.

Only affects touch devices (listener gated by isTouchDevice()).
2026-04-12 10:01:29 -04:00
arkonandClaude Opus 4.6 7b8b175133 chore: version packages
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 21:01:32 +02:00
Ark0N c3027b21e1 Merge pull request #63 from Typhon0/fix/quick-start-linked-cases
Good catch — quick-start was ignoring linked-cases.json. Thanks @Typhon0!
2026-04-11 20:58:57 +02:00
Loïc Sculier 0b231edd43 Fix quick-start to resolve linked cases before codeman-cases fallback
/api/quick-start was always resolving caseName against CASES_DIR,
ignoring any entries in ~/.codeman/linked-cases.json. This caused
sessions to start in ~/codeman-cases/<name> even when the case was
linked to an external project directory.

Fix: read linked-cases.json first and prefer that path, falling back
to validatePathWithinBase only when no link is found.
2026-04-11 17:45:41 +02:00
arkon ea1c2ee4ec chore: version packages 2026-04-11 07:21:18 +02:00
arkonandClaude Opus 4.6 b4a808adcf fix: security hardening and cleanup from community PR cherry-picks
- Add HTML sanitizer for markdown rendering (XSS prevention)
- Switch service worker to network-first caching (deploys take effect immediately)
- Sanitize Content-Disposition filenames (header injection prevention)
- Expose session.muxName getter, replace unsafe `as any` cast
- Static import for execFile, update CLAUDE.md keyboard shortcuts

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 07:20:09 +02:00
arkonandClaude Opus 4.6 f3cbe9bca6 feat: cherry-pick keyboard UX and file download from community PRs
Cherry-picked from PR #60 (keyboard UX) and PR #61 (file download):

- Alt+1-9 session switching
- Disable Ctrl+K (too easy to trigger accidentally)
- Session rename with prefix preservation (w1-case: description)
- Shift+Enter / Ctrl+Enter multiline input via tmux send-keys -H
- Android virtual keyboard fix for non-composition input
- File download button in browser file explorer (?download=true)

Dropped from PR #60: stale package-lock.json, upload popup (missing upload.html)
Dropped from PR #61: standalone /api/download endpoint (arbitrary fs access)
Fixed from PR #60: execFileSync replaced with async execFile

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 06:59:50 +02:00
Ark0N a11bcb0029 Merge pull request #59 from aakhter/feat/pwa-support
feat: add PWA support for Android/iOS home screen install
2026-04-11 06:54:06 +02:00
Ark0N 47fd9a922f Merge pull request #58 from ToRvaLDz/master
feat: add named Cloudflare tunnel support
2026-04-11 06:53:56 +02:00
Ark0N 14f7d8298d Merge pull request #62 from TeigenZhang/feat/mobile-response-viewer
feat: mobile response viewer with markdown rendering
2026-04-11 06:53:47 +02:00
Teigen d32f4debb2 feat: markdown rendering for response viewer
Add marked.js (39KB) for rich text display in the response viewer.
Renders headings, code blocks, lists, tables, blockquotes, and
inline formatting with dark theme styling.

Falls back to escaped plain text if marked.js fails to load.
2026-04-10 19:12:06 +08:00
Teigen 3cb7b510f8 feat: mobile response viewer — read full Claude responses via native scroll
Claude Code's Ink framework uses alternate screen buffer + VPA cursor
positioning, resulting in near-zero xterm.js scrollback on mobile.
Instead of fighting terminal scrollback, this adds a native scrollable
overlay that reads structured responses from Claude's JSONL transcripts.

- New API: GET /api/sessions/:id/last-response reads transcript JSONL
  - ?context=full returns full conversation thread (user + assistant)
  - Fallback to terminal buffer with ANSI stripping if no transcript
- Response viewer panel: bottom sheet with native iOS/Android scroll
- "More" button loads full conversation context as threaded view
- Eye icon in header bar (mobile only), no toolbar space impact
2026-04-10 19:07:02 +08:00
Aamer AkhterandClaude Opus 4.6 6a12a72c9c local: add PWA support for Android home screen install
Add app icons, update manifest with icon entries, and add app-shell
caching to service worker for offline/instant startup.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 22:15:24 -04:00
Marco Migozzi f1a126efeb feat: add named Cloudflare tunnel support
Add named tunnel mode alongside existing quick tunnel, with systemd
service and setup helper. All tunnel parameters are configurable via
environment variables:

  CLOUDFLARED_TUNNEL_NAME   — tunnel name (default: codeman)
  CLOUDFLARED_TUNNEL_ID     — tunnel UUID (from: cloudflared tunnel list)
  CODEMAN_TUNNEL_HOSTNAME   — public hostname

Backward compatible: ./tunnel.sh [start|stop|status|url] still works.
2026-04-04 13:41:28 +02:00
Marco Migozzi 12fd780af8 feat: add named Cloudflare tunnel support
Add named tunnel mode alongside existing quick tunnel, with systemd
service and setup helper. Tunnel ID and hostname are configurable via
CLOUDFLARED_TUNNEL_ID and CODEMAN_TUNNEL_HOSTNAME env vars.
2026-04-04 13:37:20 +02:00
Teigen fd74a42933 Merge remote-tracking branch 'origin/master' into dev 2026-04-04 08:07:15 +08:00
arkonandClaude Opus 4.6 7101e64800 refactor: restructure repo for cleaner GitHub landing page
Reduce visible top-level items from 21 to 14:
- Untrack test-results/, tmp/, public symlink (added to .gitignore)
- Move agent-teams/ → docs/agent-teams/
- Move mobile-test/ → test/mobile/
- Move tools/remotion/ → scripts/remotion/
- Move eslint.config.js, vitest.config.ts → config/

All path references updated across CLAUDE.md, package.json,
.prettierignore, vitest configs, and capture scripts.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 16:23:54 +02:00
Teigen 1b10d9b733 feat: mobile logo, expandable history, fix session resume
- Show Codeman logo on mobile as compact home button (was hidden)
- Add "Show More" button for history sessions (initial 4, expand all)
- Deduplicate by projectKey instead of workingDir (lossy decode fix)
- Fix project key decoding: handle '_' encoded as '-' with look-ahead
- Pre-validate resumeSessionId before passing to Claude CLI
- Apply content validation to all session files regardless of size
2026-04-03 11:02:43 +08:00
arkon 5078f5251d chore: version packages 2026-04-03 04:28:36 +02:00
arkonandClaude Opus 4.6 196af8fba7 fix: allow bracket chars in model flag for opus[1m] context window
The model validation regex rejected brackets, silently dropping models
like opus[1m]. Also quote the model flag to prevent bash glob expansion.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 04:27:32 +02:00
arkonandClaude Opus 4.6 a9b22b86a4 docs: use launchctl bootstrap instead of deprecated load
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 04:20:42 +02:00
arkonandClaude Opus 4.6 28cace5858 docs: clean up README install and service sections
Remove fork/branch install instructions and env vars table for cleaner
first impression. Reformat systemd and launchd service blocks as
readable multi-line heredocs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 04:19:45 +02:00
arkon 0a594b61bd chore: version packages 2026-04-03 04:17:33 +02:00
arkon 89d787a949 chore: version packages 2026-04-03 04:01:08 +02:00
arkonandClaude Opus 4.6 bd9797b68c fix: sanitize case names from filesystem to prevent XSS in inline handlers
Filter readdir and linked-case names through /^[a-zA-Z0-9_-]+$/ before
returning them from GET /api/cases. Prevents XSS via maliciously-named
directories reaching frontend inline onclick handlers where escapeHtml
is insufficient (HTML-decoded back to quotes before JS execution).

Also fix misleading "Drag or use arrows" hint (no drag-and-drop exists).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 03:58:29 +02:00
Ark0N 0e6cd94312 Merge pull request #56 from TeigenZhang/feat/case-manage-reorder-delete
feat: add case reorder and delete in Manage tab
2026-04-03 03:53:25 +02:00
arkonandClaude Opus 4.6 24a6f1cac8 chore: remove accidentally committed build artifact and dev-specific script
Remove dist/state-store.js (compiled build artifact that should not be tracked)
and scripts/claudeman-launchd-wrapper.sh (developer-specific launchd wrapper
with hardcoded paths) that were included in #55.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 03:51:00 +02:00
Ark0N 8e679a280b Merge pull request #55 from TeigenZhang/fix/auto-attach-on-restart
fix: auto-attach PTY on server restart
2026-04-03 03:50:31 +02:00
Teigen c642689bbd feat: add case reorder and delete in Manage tab
Add a "Manage" tab to the create-case modal with up/down reorder
buttons and delete for each case. Linked cases are unlinked (folder
preserved); CASES_DIR cases are permanently deleted.

Backend:
- DELETE /api/cases/:name — unlink or delete
- PUT /api/cases/order — persist ordering to settings.json
- GET /api/cases now respects saved caseOrder

Frontend:
- Third "Manage" tab in createCaseModal with case list
- Delete button in mobile case picker bottom sheet
- SSE events: case:deleted, case:order-changed
2026-04-03 09:45:19 +08:00
Teigen 28a6247c27 fix: auto-attach PTY to surviving tmux sessions on server restart
Previously, restoreMuxSessions() only created Session objects without
attaching PTY processes. Sessions stayed at pid=null until the client
manually selected them, causing terminals to appear "closed" after deploy.

Now the server calls startInteractive() for each recovered session during
startup, so all sessions resume capturing output immediately. The frontend
auto-attach condition is also relaxed from (pid===null && status==='idle')
to (pid===null && !_ended) as a safety net for edge cases.
2026-04-02 21:35:00 +08:00
Teigen 0ceb455c4b feat: add Ctrl+O button to mobile keyboard accessory bar 2026-04-02 21:13:17 +08:00
Teigen e51117dfa9 feat: add left/right arrow buttons to mobile keyboard accessory bar
Support cursor left/right movement on mobile, using the same blue
accessory-btn-arrow style as the existing up/down arrows.
2026-04-02 15:19:15 +08:00
Teigen 13d41cf7c7 feat: add Tab, Esc, ⌥Enter buttons to mobile keyboard accessory bar
- Add Tab (forward), Esc, and Option+Enter (newline) buttons
- Reorder buttons: ↑ ↓ 📋 Tab ⇧Tab ⌥Enter Esc /init /clear /compact dismiss
- Unify dismiss button style with arrow buttons (was oversized with custom class)
- Remove unused .accessory-btn-dismiss CSS rules
2026-04-02 14:32:54 +08:00
Teigen 2c7557d002 fix state store temp file collisions 2026-04-02 14:24:10 +08:00
arkonandClaude Opus 4.6 53b473708f fix: macOS support — HTML cache, launchd service, trust dialog
Three fixes for macOS deployments:

1. HTML cache bug: @fastify/static with preCompressed serves .html.br/.html.gz
   files, so path.endsWith('.html') missed them — HTML got 1-year immutable
   cache headers instead of no-cache, causing stale pages after deploys.

2. Installer launchd support: macOS now gets proper LaunchAgent setup (like
   systemd on Linux). Removes competing LaunchDaemons to prevent duplicate
   services fighting over the port. Update/uninstall also handle launchd.

3. Trust dialog auto-accept: Claude CLI 2.x shows a workspace trust prompt
   on first launch per directory. Sessions detect "trust this folder" in PTY
   output and auto-send Enter, preventing sessions from hanging on startup.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-01 08:51:37 +02:00
arkonandClaude Opus 4.6 2cba393ae5 fix: installer fails on macOS when piped via curl | bash
When running `curl | bash`, stdin is the pipe, not the terminal.
Homebrew and sudo need TTY access to prompt for the password.
Redirect /dev/tty as stdin for these subprocesses.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-31 19:34:39 +02:00
arkon 64b8ea30b2 chore: version packages 2026-03-31 03:10:22 +02:00
Ark0N 5743af3339 Merge pull request #52 from TeigenZhang/feat/default-model-support
Thanks for the contribution @TeigenZhang! Clean, well-scoped change — applied consistently across all session creation paths. 🎉
2026-03-30 19:19:04 +02:00
Teigen 2011bd8d89 fix: terminal flicker regression — move viewport clear inside dimension guard
Three fixes from the WIP flicker branch that were lost during master merges:

1. Move viewport clear (\x1b[3J\x1b[H\x1b[2J) inside the dimension-change
   guard so it only fires when cols/rows actually change. Previously every
   resize event cleared the screen even at identical dimensions, causing
   visible flicker with no subsequent Ink redraw to repaint.

2. Sync _lastResizeDims in sendResize() so restoreTerminalSize() doesn't
   trigger a redundant viewport clear on the next throttledResize tick.

3. Add didScroll tracking to touch events — tap (no scroll) now refocuses
   xterm's hidden textarea, fixing mobile keyboard input routing after
   tapping the terminal area.
2026-03-30 16:45:20 +08:00
Teigen b76724690d Merge branch 'feat/mobile-shift-tab' into dev 2026-03-30 16:45:16 +08:00
Teigen f277f9664c feat: add Shift+Tab button to mobile keyboard accessory bar
Mobile users cannot press Shift+Tab on virtual keyboards. Add a ⇧Tab
button that sends the escape sequence (\x1b[Z) to the PTY, enabling
mode switching on mobile devices.

Also fix accessory bar overflow on narrow screens by making it
horizontally scrollable with hidden scrollbar.
2026-03-30 16:44:43 +08:00
Teigen cd49171bbc feat: support "Default (CLI default)" option for model selection
Allow users to leave the default model unset, so sessions use whatever
the Claude CLI defaults to rather than forcing a specific model.

- Add empty-value "Default (CLI default)" option to the model dropdown
- Treat empty string as undefined when passing model to Session
- Apply consistently across session creation, quick-start, and Ralph
2026-03-30 16:44:00 +08:00
arkon 0f57342b10 chore: version packages 2026-03-29 05:10:16 +02:00
arkonandClaude Opus 4.6 e1f0ac993a fix: default new sessions to opus[1m] (1M context) instead of opus (200k)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-29 05:09:24 +02:00
arkon a84ef52992 chore: version packages 2026-03-28 16:49:55 +01:00
arkonandClaude Opus 4.6 692c894760 fix: correct process tree detection and prevent timer starvation
1. Rewrote getActiveChildProcesses() to use a single `ps --ppid` call
   instead of two-level pgrep. The pane PID is typically claude itself
   (bash exec'd into it), not a bash wrapper — so direct children of
   pane_pid ARE the tool processes.

2. Added timer restart in tryStartAiCheck() when skipping due to child
   processes. Without this, the pre-filter and no-output timers (both
   one-shot) would never fire again, permanently stalling idle detection
   for sessions with silent long-running processes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-28 05:14:39 +01:00
arkonandClaude Opus 4.6 ad0acb6d58 feat: detect active child processes to prevent false idle during running tools
When Claude Code spawns bash tools (test suites, builds, servers), the
respawn controller could falsely detect idle if terminal output paused.
Now checks the process tree for active children of the Claude process
before triggering AI idle checks or confirming idle state.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-28 04:34:53 +01:00
arkonandClaude Opus 4.6 d866c8f30e chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-27 01:43:59 +01:00
arkonandClaude Opus 4.6 28537de39d refactor: pass 3 — extract helpers, split long functions, deduplicate patterns
Backend:
- subagent-watcher: split 176L processEntry() into 5 focused methods; extract
  _resolveDescription() deduplicating 3 call sites for description acquisition
- bash-tool-parser: split 152L processCleanLine() into 4 handlers; extract
  _createActiveTool() factory and _scheduleAutoRemove() helper
- session: extract _setupOrAttachMuxSession() deduplicating ~80L between
  startInteractive/startShell; extract _handleTerminalOutput()
- respawn-controller: split 180L handleTerminalData() into 3 detection layers;
  data-driven validation loop replacing 9 individual calls
- plan-orchestrator: extract _extractJsonFromResponse(), _emitAgentFailure(),
  _formatResearchSection() helpers
- orchestrator-loop: extract _finalizeTask() unifying task completion/failure;
  _clearTimer() utility for correct clearInterval/clearTimeout dispatch
- ralph-status-parser: config-driven FIELD_PARSERS[] replacing 8 near-identical
  field-matching blocks; split updateCircuitBreaker() into focused handlers
- state-store: extract _mergeWithInitialState() and _resetCircuitBreaker()

Frontend:
- app.js: add _notifySession() helper used by 18 call sites across 5 modules
- panels-ui.js: extract _addActivityEntry() replacing 4 duplicate blocks
- settings-ui.js: extract _updateTunnelUrlRow() deduplicating 2 blocks
- ralph-panel.js, respawn-ui.js: convert to _notifySession()

Routes:
- route-helpers: add toggleService() helper
- system-routes: use toggleService() for watcher toggles; extract collectActiveTokens()
- orchestrator-routes: data-driven EVENT_MAP replacing 10 identical listeners

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-26 22:50:25 +01:00
arkonandClaude Opus 4.6 ba09184efa fix: wizard "No JSON found" — Claude CLI stream-json returns empty result field
Claude CLI's --output-format stream-json now returns "result": "" in the result
message. The actual response text lives in assistant message text blocks, which
_textOutput correctly accumulates. runPrompt() was returning the empty
resultMsg.result without falling back to _textOutput.value.

Also improved plan-orchestrator JSON extraction to try code-block-wrapped JSON
first (```json {...} ```) before the greedy regex, plus debug logging.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-25 23:53:57 +01:00
arkonandClaude Opus 4.6 93719b41cd chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-25 23:34:14 +01:00
arkonandClaude Opus 4.6 a448983be3 refactor: pass 2 — extract shared helpers and simplify patterns
app.js:
- Add _clearTimer() helper replacing 11 inline clearTimeout patterns
- Add _isStaleSelect() helper for generation check + cleanup
- Replace 11 keyboard shortcut if-blocks with data-driven lookup table
- Extract _cleanupPreviousSession() from selectSession() (~75 lines)
- Extract _resetAllAppState() from handleInit() (~75 lines)

tmux-manager:
- Extract buildEnvExports() eliminating duplication in createSession/respawnPane
- Extract buildPathExport() for CLI path resolution
- Extract _configureOpenCode() for OpenCode setup

routes:
- Add readJsonConfig() to route-helpers, replacing 5 inline JSON-read patterns
- Add validateSessionFilePath() to route-helpers, replacing 2 identical path
  traversal validation blocks in file-routes

session-auto-ops:
- Convert executeWhenIdle() from 8 positional params to options object
- Extract validateThreshold() for shared compact/clear validation

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-25 23:32:28 +01:00
arkonandClaude Opus 4.6 3145eac6d9 refactor: extract helper methods to reduce duplication and improve readability
DRY up repeated patterns across 7 core files:
- state-store: extract serializeState() and split assembleStateJson() into 3 focused methods
- session: extract _resetBuffers(), _clearAllTimers(), _handleJsonMessage()
- ralph-tracker: extract completeAllTodos() (was 4x duplicated), emitValidationWarning(), similarity constants
- subagent-watcher: extract markSubagentAsCompleted(), extractFirstTextContent(), emitToolResult(), findOldestInactiveAgent()
- respawn-controller: extract recoveryResetToWatching(), canAutoAccept(), formatRemainingSeconds(), validatePositiveTimeout()
- tmux-manager: replace 15 path.includes() checks with single UNSAFE_PATH_CHARS regex
- session-auto-ops: extract executeWhenIdle() shared retry helper for checkAutoCompact/checkAutoClear

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-25 23:21:38 +01:00
arkonandClaude Opus 4.6 e3c609f5f0 test: add coverage for lastUsedCase partial update and strict schema rejection
Tests that partial PUT /api/settings with just lastUsedCase works correctly
and that including modelConfig triggers strict Zod schema rejection (the bug
fixed in #49).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-25 13:39:47 +01:00
Tenggan Zhang 52e774f83c fix: case selection not persisting across page refresh (#49)
Thank you for the clean fix! The root cause analysis in the PR description was excellent — the strict Zod schema rejecting modelConfig during the GET-then-PUT pattern was a subtle bug.
2026-03-25 13:39:16 +01:00
arkon 82d08df53f chore: version packages 2026-03-25 00:18:14 +01:00
arkonandClaude Opus 4.6 2709b2fe49 feat: make buffer size limits configurable via environment variables
Allow overriding MAX_TERMINAL_BUFFER_SIZE, TRIM_TERMINAL_TO, MAX_TEXT_OUTPUT_SIZE,
TRIM_TEXT_TO, and MAX_MESSAGES via CODEMAN_* env vars, falling back to existing
defaults. Enables users with fewer sessions or more RAM to tune buffer sizes
without patching source.

Closes #48

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-24 18:40:40 +01:00
Ark0N b1d3b27e5b Merge pull request #47 from TeigenZhang/fix/mobile-cjk-input-and-layout
fix: mobile CJK input, terminal flicker, and layout overflow
2026-03-24 18:22:48 +01:00
Teigen 47963b54fa fix: mobile CJK input, terminal flicker, and layout overflow
Terminal flicker:
- Skip buffer-recovered/clear-terminal events during active buffer load
  to prevent competing clear+rewrite cycles (app.js)
- Move viewport+scrollback clear inside dimension-change guard so resize
  without actual SIGWINCH doesn't blank the terminal (terminal-ui.js)
- Sync _lastResizeDims on explicit resize to prevent redundant clears

CJK input rewrite (input-cjk.js):
- Use InputEvent.inputType to distinguish insertText (final) from
  insertCompositionText (tentative) — fixes Chinese punctuation and
  English text being swallowed during Android IME composition
- Remove isComposing guard on Enter so it always sends
- Phantom character (U+200B) keeps textarea non-empty so Android
  long-press backspace generates continuous deleteContentBackward
  events at the keyboard's native repeat rate

CJK input settings:
- Add "CJK Input" toggle in Settings > Input (index.html, settings-ui.js)
- Store as device-specific setting (cjkInputEnabled), not synced to server
- Replace INPUT_CJK_FORM env var dependency with user-controlled setting
  (env var still works as server override)

Mobile layout:
- Fix welcome screen overflow on phones by constraining .welcome-content
  to calc(100vw - 1.5rem) (mobile.css)
- Move xterm helper textarea on-screen for touch devices to fix iOS
  keyboard input (styles.css)
- Focus terminal synchronously in user-gesture context for iOS Safari
  keyboard activation (session-ui.js, app.js)
- Refocus terminal on tap (not scroll) in touch handler (terminal-ui.js)
2026-03-24 09:22:06 +08:00
arkonandClaude Opus 4.6 b7c3c30c8c fix: send Ctrl+L after tab switch to clear stale Ink CUP frames
Tailed terminal buffers contain multiple CUP-positioned Ink frames from
different time points. When replayed in xterm, old frames at viewport
positions not covered by the latest frame persist as ghost content
(e.g. duplicate "bypass permissions" bars). After buffer load, send
Ctrl+L via the session input API to trigger a full Ink redraw, which
overwrites all stale frame content with the correct current state.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-23 16:59:48 +01:00
arkonandClaude Opus 4.6 0d80524f10 fix: prevent duplicate terminal output on tab switch to busy sessions
Two fixes for the tab-switching corruption bug:

1. _finishBufferLoad() now discards queued SSE events instead of flushing
   them. The loaded API buffer is the source of truth — queued events
   overlap with it, and flushing them writes duplicate Ink cursor-up
   redraws that corrupt the terminal display (garbled text, wrong cursor
   positions).

2. Skip stale cache write for busy sessions. When a session is actively
   working, the cache is always outdated — writing it first and then
   rewriting with the fresh API buffer caused a jarring double-render
   flash. Now busy sessions get a single clean clear+write transition.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-23 12:32:33 +01:00
arkonandClaude Opus 4.6 a9d83ec4e3 chore: version packages
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 23:55:39 +01:00
arkonandClaude Opus 4.6 84137cdba4 chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:59:27 +01:00
arkonandClaude Opus 4.6 6eb3969816 fix: avoid no-control-regex lint error for ANSI strip pattern
Use RegExp constructor with String.raw to express \x1b without
a literal control character in the source, matching the pattern
used elsewhere in the codebase.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:57:36 +01:00
arkonandClaude Opus 4.6 de49437a6f docs: add browser-testing-guide to CLAUDE.md references, clarify route count
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:38:14 +01:00
arkonandClaude Opus 4.6 6a27639083 fix: increase Ink frame search window from 4KB to 64KB to prevent partial frames
Single Ink frames with response content can be 10-20KB, so the 4KB tail
was too small and caused blank gaps. Now searches the last 64KB for VPA
row drops to find the last complete frame boundary.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:38:09 +01:00
arkonandClaude Opus 4.6 eb1b38c718 fix: align case select group height — stretch buttons to match dropdown
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:06:15 +01:00
arkonandClaude Opus 4.6 ea7b103b47 fix: prevent stale terminal data on tab switch — add chunkedTerminalWrite cancellation
chunkedTerminalWrite used requestAnimationFrame to write buffer chunks across
frames but had no cancellation. When switching tabs, old session's remaining
chunks continued writing stale data into the new session's terminal, causing
visual artifacts and garbled content.

- Add _chunkedWriteGen generation counter to abort in-flight chunked writes
- Bump gen early in selectSession() and SSE reconnect to immediately cancel
- Guard finish() so aborted writes don't flush SSE queue for wrong session
- Add fitAddon.fit() before buffer writes to sync terminal dimensions
- Add fitAddon.fit() in sendResize() to ensure local/server dim parity

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:00:58 +01:00
arkonandClaude Opus 4.6 bec8e2f9ee fix: improve history prompt extraction — filter expanded commands, add tail scan fallback
Skip /init expansions, slash commands, orchestrator prompts, ANSI codes, secrets,
and short/vague messages. When head scan finds no usable prompt (e.g. /init sessions),
read last 32KB of transcript to find a recent meaningful user message.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 02:42:50 +01:00
arkonandClaude Opus 4.6 2203f3a347 feat: visual redesign — glass morphism, refined colors, polished UI
Modernize the entire UI with a cohesive "refined dark glass" aesthetic
while preserving all existing functionality.

- Header/toolbar: backdrop-filter blur(16px), semi-transparent backgrounds
- Buttons: 6px radius, multi-stop gradients, inner glow, cubic-bezier transitions
- Welcome screen: gradient text title, radial bg glow, pill-shaped buttons with hover lift
- Panels/modals: glass backgrounds, 12px radius, layered shadows
- Color palette: cooler blue-tinted darks replacing flat blacks
- Forms: refined inputs with focus rings, glass toggle switches
- New CSS vars: --glass-bg, --glass-border, --btn-radius, --transition-smooth

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 02:24:27 +01:00
arkonandClaude Opus 4.6 867a10d78a refactor: optimize history endpoint — reuse buffer, extract readFileHead, use line iterator
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 02:01:56 +01:00
arkon e54d7badc4 chore: version packages 2026-03-22 01:57:49 +01:00
arkon 40dfac3534 feat: improve session history with first prompt + clickable monitor rows (closes #45) 2026-03-22 01:57:17 +01:00
arkonandClaude Opus 4.6 e899a43a18 chore: hide orchestrator button until feature is fully tested
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 01:48:11 +01:00
arkonandClaude Opus 4.6 0cab8a7ece fix: stop subagent monitor windows from auto-opening on discovery
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 01:44:28 +01:00
arkonandClaude Opus 4.6 7b7cf958c0 feat: add live progress during orchestrator plan generation
New SSE event orchestrator:planProgress streams phase/detail updates
from the planner to the frontend in real-time. The panel now shows
a scrollable log of planning steps instead of just "Generating plan..."

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 20:03:24 +01:00
arkonandClaude Opus 4.6 6e64ddd853 feat: add Orchestrator button to toolbar
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 19:59:06 +01:00
arkonandClaude Opus 4.6 afea91b92b fix: patch 3 production bugs found during deep audit
1. Post-phase verify timer leak — setTimeout for verifyCurrentPhase was
   never stored, so pause() couldn't cancel it. Timer now tracked in
   postPhaseTimer field and cleared in clearPhasePoll().

2. Event forwarding flag survives loop replacement — boolean
   eventForwardingAttached stayed true when a new loop was created,
   so the new loop never got SSE forwarding. Now tracks the loop
   instance reference instead of a boolean.

3. Replan stuck when no sessions — replanPhase() returned without
   setting up task handlers or polling when no idle sessions were
   available. Now starts polling so the queued task gets picked up
   when a session becomes idle.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 19:23:37 +01:00
arkonandClaude Opus 4.6 9449a8f157 test: expand OrchestratorLoop coverage to 60 tests — 17 new deep paths
New coverage:
- Task failure & retry (handleTaskFailed retry when retries < 2)
- Phase error auto-retry (handlePhaseError when attempts < maxAttempts)
- Verification with actual criteria (verifier call, pass/fail flow)
- Verification failure → replan → retry cycle
- Max verification attempts → phase failure
- Multi-phase sequential advancement
- Compact between phases (writeViaMux('/compact'))
- Crash recovery from verifying/replanning/paused states
- Single-task vs multi-task prompt generation
- Team phase sendInput error handling
- Verification session fallback (no sessions → skip)
- taskAssigned, phaseCompleted, phaseFailed event emissions

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 19:15:42 +01:00
arkonandClaude Opus 4.6 d322f17f73 test: add 43 deep integration tests for OrchestratorLoop state machine
Covers full lifecycle: start → plan → approve → execute → verify → complete.
Tests state transitions, event emissions, persistence/recovery, pause/resume,
skip/retry, team phase execution, error handling, and edge cases.

Also fixes bugs found during review:
- Route context snapshot: use getter for orchestratorLoop (was null forever)
- Event listener stacking: guard setupEventForwarding with boolean flag
- Replan completion: create tracked TaskQueue task instead of raw sendInput
- Pause cleanup: call cleanupTaskHandlers() on pause
- Phase timeout: add phaseTimeoutTimer enforcement

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 15:33:02 +01:00
arkonandClaude Opus 4.6 61b5ec095c feat: add Orchestrator Loop — phased plan execution with team agents
Adds a new autonomous loop that accepts high-level goals, generates
phased execution plans via AI, and executes them step-by-step with
verification gates between phases.

Core components:
- OrchestratorLoop: state machine (idle→planning→approval→executing→verifying→completed)
- OrchestratorPlanner: plan generation via PlanOrchestrator, Kahn's algorithm phase grouping
- OrchestratorVerifier: phase verification (strict/moderate/lenient modes)
- Prompt templates for phase execution, team delegation, verification, replanning

API (10 endpoints):
- POST start/approve/reject/pause/resume/stop
- GET status/plan
- POST phase/:id/skip, phase/:id/retry

Frontend: orchestrator-panel.js with SSE-driven state, phase progress, task tracking

Tests: 22 tests (18 route + 4 unit), all passing. Typecheck/lint/format clean.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 07:20:18 +01:00
arkonandClaude Opus 4.6 497ca4891a fix: restore mobile terminal scrollback — use JS scrollLines() instead of broken native scroll
xterm.js DOM renderer doesn't populate .xterm-viewport's scroll area (the div
is empty, scrollHeight === clientHeight), so native CSS scrolling via
touch-action:pan-y and overflow-y:scroll had nothing to scroll. Desktop worked
only because the wheel handler called terminal.scrollLines() directly.

- Replace split mobile/desktop touch handlers with unified JS-driven handler
  that converts touch deltas to terminal.scrollLines() calls (with pixel
  accumulation for slow swipes and momentum scrolling)
- Change touch-action from pan-y to none on terminal elements so browser
  doesn't fight the JS handler
- Remove now-unnecessary xterm-viewport position/overflow/z-index overrides
  and iOS -webkit-overflow-scrolling rules
- Fix _shrinkPaddingToFit() arithmetic (was adding gap instead of subtracting)
- Minor: add route-helpers.ts to CLAUDE.md, fix sse-events.ts comment count

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 09:38:18 +01:00
arkon 580b7a3f90 chore: version packages 2026-03-19 12:35:39 +01:00
arkonandClaude Opus 4.6 34c3d8f5ff fix: tighten mobile keyboard layout — eliminate dead space and toolbar overlap
- Remove redundant 50px CSS padding on terminal-container when keyboard visible
- Reduce JS paddingBottom constant from +94 to +84 (exact toolbar + accessory)
- Add _shrinkPaddingToFit() to eliminate terminal row quantization gap
- Add CSS padding-bottom on .main for fixed toolbar clearance (keyboard hidden)
- Match iOS Safari toolbar offset (100vh - --app-height) in .main padding

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-19 12:35:08 +01:00
arkonandClaude Opus 4.6 2491471ba5 fix: prevent mobile page scroll when typing with keyboard open
iOS Safari scrolls the document to bring xterm's hidden textarea into
view when the user types, pushing the entire UI off-screen. Fix with:
- CSS position:fixed on .app when keyboard is visible
- window.scroll listener to reset scroll position as safety net
- scroll reset in onKeyboardShow before and after fit/resize

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-19 11:47:18 +01:00
arkonandClaude Opus 4.6 0c4aac8029 chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-19 10:05:16 +01:00
arkonandClaude Opus 4.6 d436c6375f fix: strip Ink spinner bloat from terminal buffer before tailing
During long thinking phases, Ink's TUI rewrites the spinner/status bar
thousands of times via absolute cursor positioning (VPA/CUP). These
500KB+ of redraw frames pushed real content out of the 128KB tail
window, making the terminal appear empty when switching tabs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-16 10:55:24 +01:00
arkonandClaude Opus 4.6 e96baf9f66 chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-16 01:42:16 +01:00
arkon 7b8aa529f2 chore: version packages 2026-03-15 03:52:30 +01:00
arkonandClaude Opus 4.6 0ad4e0ea24 fix: correct resolveCasePath priority order and suppress JSON parse warnings
- resolveCasePath now checks linked cases first (matching original behavior
  of /api/cases/:name and /api/cases/:name/fix-plan handlers)
- readLinkedCases only warns on real I/O errors, not JSON parse errors
  (SyntaxError has no .code property, so check for .code existence first)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 03:50:45 +01:00
arkonandClaude Opus 4.6 6bc403d88d refactor: clean up case routes DRY violations, remove dead export, standardize reply API
- Extract readLinkedCases() helper and resolveCasePath() to eliminate 6x duplicated
  linked-cases.json path construction and 5x duplicated file read/parse logic
- Replace O(n) .some() duplicate check with O(1) Set.has() in case listing
- Un-export isError() in types/api.ts (only used internally by getErrorMessage)
- Standardize reply.status() → reply.code() in system-routes (Fastify canonical API)
- Update CLAUDE.md: accurate frontend module listing, SSE event count (~106)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 03:47:18 +01:00
arkon 1ad05a5a42 chore: version packages 2026-03-14 22:22:52 +01:00
arkonandClaude Opus 4.6 192690911f refactor: extract app.js into 6 domain modules with deferred init
Split the monolithic app.js (~12.5K lines) into 6 focused mixin modules
that extend CodemanApp.prototype via Object.assign:

- terminal-ui.js — terminal setup, rendering pipeline, controls
- respawn-ui.js — respawn banner, countdown, presets, run summary
- ralph-panel.js — Ralph state panel, fix_plan, plan versioning
- settings-ui.js — app settings, visibility, web push, tunnel/QR, help
- panels-ui.js — subagent panel, teams, insights, file browser, log viewer
- session-ui.js — quick start, session options, case settings

Fix deferred script init ordering: wrap CodemanApp instantiation in
DOMContentLoaded so all defer'd mixin modules execute their
Object.assign before the constructor runs. Without this, init() calls
methods like applyHeaderVisibilitySettings() that don't exist yet.

Guard missing cleanupWizardDragging() call in subagent-windows.js.
Update build.mjs to minify/hash all new modules. Update CLAUDE.md
with new frontend architecture and load order.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 21:24:22 +01:00
arkon 551461cb31 chore: version packages 2026-03-14 19:26:09 +01:00
arkonandClaude Opus 4.6 c4bae75c59 fix: add onerror handler for lazy-loaded WebGL addon script
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 19:23:15 +01:00
arkonandClaude Opus 4.6 88c415fc37 perf: V8 compile cache, lazy-load WebGL, preload hints, batch tmux reconciliation
- Enable NODE_COMPILE_CACHE in systemd service and npm start for 10-20% faster cold starts
- Lazy-load xterm-addon-webgl.min.js (244KB) only on desktop — mobile never downloads it
- Add <link rel="preload"> hints for critical scripts (xterm, constants, app) in <head>
- Replace per-session tmux subprocess calls with single batch `list-panes -a` call
  (N*2+1+M execSync calls → 1 for reconcileSessions)
- Fix CLAUDE.md frontend module count (10 → 11)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 19:20:44 +01:00
arkonandClaude Opus 4.6 7175e4b350 docs: update CLAUDE.md and README.md to reflect current codebase
Correct stale counts and add missing entries: route modules 12→13
(ws-routes.ts), frontend modules 9→10 (input-cjk.js), handler count
~111→~114, utilities section expanded, TypeScript badge 5.5→5.9,
frontend extracted modules 8→9.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:48:25 +01:00
arkonandClaude Opus 4.6 08a417997f ci: upgrade actions/checkout and actions/setup-node to v6 (Node 24)
Replace v4 (Node 20) with v6 (Node 24 native) to eliminate the
deprecation warning. Remove the FORCE_JAVASCRIPT_ACTIONS_TO_NODE24
workaround since v6 doesn't need it.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:42:22 +01:00
arkonandClaude Opus 4.6 e6cb89b0cd ci: use Node.js 24 runtime for actions and bump release node to 22
Opt into Node.js 24 for GitHub Actions runners (actions/checkout@v4,
actions/setup-node@v4) to silence deprecation warnings. Also bump
release.yml from node 20 to 22 to match ci.yml.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:40:44 +01:00
arkon d072e773d8 chore: version packages 2026-03-14 18:38:28 +01:00
arkonandClaude Opus 4.6 a649c91b68 fix: WS session lifecycle, reconnection, and CJK session-switch cleanup
- Close WebSocket when session exits (exit event listener) to prevent
  orphaned listeners and stale writes to dead PTY
- Add readyState guard in onTerminal to stop buffering after socket closes
- Simplify heartbeat: remove redundant alive flag, use pongTimeout only
- Add exponential backoff reconnection on unexpected WS close (skip for
  server rejections 4004/4008/4009)
- Clear CJK textarea on session switch to prevent wrong-session input

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:37:10 +01:00
arkonandClaude Opus 4.6 3383c23099 fix: address code review findings across WS, CJK input, install.sh, and README
WebSocket route: add socket error handler to prevent process crashes, enforce
per-session connection limit (max 5), track/decrement counts on close.

CJK input: add destroy() method with proper listener cleanup, guard against
double-init, add maxlength/aria-label to textarea, use language-neutral
placeholder, explicitly clear cjkActive on hide.

install.sh: fix update() to use $BRANCH and $REPO_URL instead of hardcoded
origin/master — fork users were silently switched back to master on update.

README: fix broken markdown table (paragraph concatenated into last cell),
add CODEMAN_NODE_VERSION to env var table.

Tests: add 8 new test cases for batch coalescing, flush threshold, unknown
message types, connection limit, heartbeat, readyState guards. Import
MAX_INPUT_LENGTH from config, add connectWs timeout, replace setTimeout
with vi.waitFor in cleanup test.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:27:58 +01:00
arkonandClaude Opus 4.6 c3e1e731ef fix: use generic placeholders in fork install README example
Replace hardcoded contributor fork URL with <user>/<branch> placeholders
so the documentation is useful for any contributor.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:11:43 +01:00
Ark0N 405b711c3a Merge pull request #41 from douchekr/feat/input-cjk-form
feat: add CJK IME input textarea and fork/branch install support
2026-03-14 18:10:09 +01:00
arkon 4295faefc9 chore: version packages 2026-03-14 18:04:51 +01:00
arkon 93e1ba5110 Merge remote-tracking branch 'origin/feat/ws-terminal-io-upstream' 2026-03-14 18:03:36 +01:00
Ark0N abbbf9e90a Merge pull request #43 from Ark0N/feat/ws-tests
test: add WebSocket terminal I/O route tests
2026-03-14 18:03:04 +01:00
Ark0N 3a41de7b57 Merge pull request #42 from Ark0N/feat/ws-heartbeat
feat: add ping/pong heartbeat to WebSocket connections
2026-03-14 18:03:02 +01:00
Ark0N 8267edc6fe Merge pull request #40 from Spirotot/feat/ws-terminal-io-upstream
feat: WebSocket terminal I/O with server-side DEC 2026 sync
2026-03-14 18:02:55 +01:00
arkonandClaude Opus 4.6 78c568e5f7 test: add automated tests for WebSocket terminal I/O route
16 tests covering session-not-found close code, terminal output with
DEC 2026 sync markers, client input forwarding, resize bounds
validation, malformed message handling, and connection cleanup of
session event listeners.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:00:55 +01:00
arkonandClaude Opus 4.6 cc624d2575 feat: add ping/pong heartbeat to WebSocket connections
Detect stale connections that TCP keepalive won't catch for minutes,
especially through tunnels and proxies. Pings every 30s with a 10s
pong timeout — if the client doesn't respond, the socket is terminated
and all timers cleaned up.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 17:58:50 +01:00
arkonandClaude Opus 4.6 5844720525 fix: validate WS resize dimensions to match HTTP route bounds
The HTTP resize route validates via ResizeSchema (cols: 1-500, rows:
1-200, integers only). The WS handler only checked typeof === 'number',
allowing floats, negatives, and extreme values through to ptyProcess.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 17:57:20 +01:00
jayparkandClaude Opus 4.6 393a2d9c28 fix: use BRANCH variable in install.sh no-changes update path
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:52:28 +09:00
jayparkandClaude Opus 4.6 809bf6a614 fix: use BRANCH variable in install.sh update path
The update path was hardcoded to origin/master. Now uses the
CODEMAN_BRANCH variable and updates the remote URL on upgrade.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:51:20 +09:00
jayparkandClaude Opus 4.6 da71d8d01c feat: support custom repo URL and branch in install.sh
Add CODEMAN_REPO_URL and CODEMAN_BRANCH env vars to install.sh
for installing from forks or feature branches. Update README with
fork installation instructions and env var reference table.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:50:06 +09:00
jayparkandClaude Opus 4.6 e5aca6aa4c feat: add CJK IME input textarea with env toggle
Add a dedicated textarea below the terminal for CJK (Korean/Japanese/Chinese)
IME input. xterm.js intercepts IME composition events, preventing composed
characters from displaying correctly. This textarea bypasses xterm entirely
by using native browser IME handling — text accumulates until Enter, then
sends to PTY in one shot.

- Always-visible textarea below terminal (inside .terminal-wrap flex column)
- focus/blur sets window.cjkActive flag to block xterm onData
- Enter sends textarea.value + \r to PTY, Escape clears
- Arrow keys, Ctrl+C/D/L/Z, Tab, Backspace pass through to PTY when empty
- attachCustomKeyEventHandler suppresses xterm key handling during composition
- INPUT_CJK_FORM=ON|OFF env var toggle (default: off, passed via SSE init)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:23:25 +09:00
Aaron FieldsandClaude Opus 4.6 ceaf4624a1 feat: add WebSocket terminal I/O with server-side DEC 2026 sync
Replace per-keystroke HTTP POST + SSE terminal output with a single
bidirectional WebSocket connection for dramatically lower input latency.
The existing SSE+POST paths remain fully functional as fallback.

Server-side: ws-routes.ts provides /ws/sessions/:id/terminal with 8ms
micro-batching and 16KB flush threshold. Each batch is wrapped in
DEC 2026 synchronized update markers so xterm.js renders atomically —
Ink's DA capability negotiation fails through the PTY→server→WS proxy
chain, so without server-injected markers, cursor-up redraws flicker.

Frontend: _connectWs/_disconnectWs manage per-session WS lifecycle.
Input and resize use WS fast path with HTTP POST fallback. SSE terminal
events are suppressed when WS is active to prevent double rendering.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 21:02:28 -04:00
arkonandClaude Opus 4.6 a6597e4a9a fix: patch 5 dependency vulnerabilities (basic-ftp, fastify, minimatch, serialize-javascript)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 00:26:56 +01:00
arkon f869e823af chore: version packages 2026-03-12 23:59:16 +01:00
arkonandClaude Opus 4.6 8d0b179f94 fix: repair 15 pre-existing subagent-watcher test failures
Root causes:
- Mock readline (EventEmitter) lacked .close() method, causing TypeError
  that blocked extractDescriptionFromFile's Promise from ever resolving
- Mock stream lacked .destroy() method (same issue after .close() fix)
- Entry-processing tests shared one readline mock between description
  extraction and tailing — events emitted before tailFile started were lost
- Liveness checker marked agents as 'completed' instead of 'idle' because
  fixed stat timestamps became stale after fake timer advancement

Fixes:
- Add createMockRl() helper with .close() method
- Use { destroy: vi.fn() } for stream mocks
- Use mockReturnValueOnce() for two-readline pattern in 7 entry tests
- Use mockImplementation() for dynamic stat timestamps

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 16:08:39 +01:00
arkonandClaude Opus 4.6 98fa55b7b2 chore: codebase cleanup — remove dead code, consolidate imports, extract constants
- Remove 3 unused exported constants (TRIM_MESSAGES_TO, MAX_TERMINAL_COLS, MAX_TERMINAL_ROWS)
- Consolidate 8 direct util imports into barrel imports (./utils/index.js)
- Extract magic number 8191 to FILE_PEEK_BYTES constant in buffer-limits.ts
- Add explanatory comments to 9 undocumented .catch(() => {}) handlers

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:50:40 +01:00
arkonandClaude Opus 4.6 c46ac30631 fix: hide subagent monitor panel by default
Change showSubagents default from true to false so the subagent
panel doesn't auto-show on page load. Users can still enable it
via Settings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:35:27 +01:00
arkonandClaude Opus 4.6 dfcc14bfd2 fix: one-liner restart command that works for background processes
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:34:42 +01:00
arkonandClaude Opus 4.6 a068008409 fix: clarify restart instructions — stop first, then start
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:33:34 +01:00
arkonandClaude Opus 4.6 0aa31f100e fix: show restart command when codeman-web is not a systemd service
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:31:19 +01:00
arkonandClaude Opus 4.6 314a160458 feat: auto-restart codeman-web service after update if running
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:23:55 +01:00
arkonandClaude Opus 4.6 e7ee5595c5 feat: auto-detect existing install and run update instead of fresh install
Re-running the install script now detects ~/.codeman/app/.git and
automatically updates instead of re-installing. Removes the separate
`bash -s update` instructions from README since it's no longer needed.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:09:37 +01:00
arkon 625d4976d3 chore: version packages 2026-03-11 19:37:17 +01:00
arkonandClaude Opus 4.6 d02cddece6 fix: correct claudeSessionId for resumed sessions and clean up DEC sync dead code
Use resumeSessionId for Claude conversation ID when resuming sessions,
increase default font size to 14, extract shared history fetch logic,
and remove unused DEC 2026 sync constants/functions (xterm.js 6.0 handles natively).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-11 19:36:58 +01:00
Ark0N abbc4b13fd Merge pull request #39 from sunnyzhouy/master
feat: session resume, xterm.js 6.0 upgrade, and resize fix
2026-03-11 19:20:28 +01:00
arkonandClaude Opus 4.6 a14e47e19c chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-11 19:13:49 +01:00
sunnyzhouy 754a966b53 Merge branch 'Ark0N:master' into master 2026-03-12 01:06:14 +08:00
zhouyuan 28dfc279d4 fix: resolve terminal resize scrollback ghost renders
- Switch resize handler to 300ms trailing-edge debounce for single reflow
- Add \x1b[3J (Erase Saved Lines) to clear scrollback reflow debris
- Remove client-side cursor-up flicker filter and DEC 2026 marker
  stripping — xterm.js 6.0 handles synchronized output natively
- Remove server-side DEC 2026 wrapping to prevent premature sync exit
  from non-reference-counted nested markers
2026-03-12 01:04:21 +08:00
arkonandClaude Opus 4.6 da85e9738b fix: iPad tablet toolbar styling and PR #34 refinements
- Scope toolbar bottom-offset to phone breakpoint only (position:fixed);
  prevents double-correction on iPad where toolbar is position:relative
- Extract keyboard accessory bar styles to top-level mobile.css so
  /init, /clear, /compact buttons render correctly on iPad
- Use desktop-style toolbar sizing on tablet (430-768px): smaller font,
  no forced min-height, proper gap between buttons
- Show voice/mic button on tablet (was hidden at <1023px with no
  mobile replacement above 430px)
- Bump CSS cache-bust version to 0.1633

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-11 12:59:47 +01:00
Ark0N 8b8907c4ec Merge pull request #34 from arnlaugsson/fix/ipad-safari-toolbar-viewport
fix: toolbar off-screen on iPad Safari with tabs
2026-03-11 01:26:55 +01:00
zhouyuan 2329dab240 feat: upgrade xterm.js 5.3 to 6.0 for native DEC 2026 synchronized output
xterm.js 6.0.0 natively handles DEC mode 2026 (synchronized output),
which renders Ink's cursor-up redraws atomically at the parser level.
This eliminates split-frame rendering that caused table header loss
and garbled overlapping text in Claude CLI sessions.

- Migrate from xterm/xterm-addon-* to @xterm/* scoped packages
- Update build.mjs and postinstall.js vendor paths
- Remove old xterm 5.x dependencies
2026-03-10 14:59:26 +08:00
zhouyuan 31ce7405a6 perf: increase terminal scrollback from 5000 to 20000 lines
Long-running Claude sessions can exceed 5000 lines easily, causing
earlier content to be lost. 20000 lines retains ~4x more history.
2026-03-10 14:36:10 +08:00
zhouyuan 06f7d40c42 feat: reduce default font size and persist tabs across refresh
- Default terminal font 14px → 12px, min 10px → 8px
- Save tab metadata to localStorage on every render
- Restore ended sessions as dimmed tabs after page refresh
- Ended tabs show "Session ended" message on click
2026-03-10 14:30:21 +08:00
zhouyuan 05eba70598 feat: improve session resume reliability and persist user settings
- Filter empty sessions from history API (check for conversation content)
- Add --resume fallback to new session if resume fails (prevents dead panes)
- Pass resumeSessionId through respawnPane for dead pane recovery
- Persist respawn presets and runMode to server settings (cross-device sync)
- Fix mobile touch handling for Recent Sessions dropdown (DOM API + touch CSS)
2026-03-10 14:03:42 +08:00
zhouyuan d27974ff6e chore: update package-lock.json 2026-03-10 02:23:15 +08:00
zhouyuan 3cca5380ba fix: route shell sessions to correct endpoint on tab click
selectSession() was always calling /interactive for restored idle
sessions regardless of mode. Shell sessions now correctly call /shell.
Also add loadHistorySessions and resumeHistorySession to frontend.
2026-03-10 02:23:10 +08:00
zhouyuan 6d7efc13e6 feat: add history session resume UI and API
Add GET /api/history/sessions endpoint that scans Claude conversation
files for resume. Add welcome overlay UI with clickable history items.
Path decoding validates existence via fs.access with HOME fallback.
2026-03-10 02:23:02 +08:00
zhouyuan 63f86807ad feat: add resumeSessionId support for conversation resume after reboot
Add resumeSessionId field throughout the session creation pipeline,
allowing sessions to resume previous Claude conversations via --resume
flag instead of --session-id.
2026-03-10 02:22:53 +08:00
Skúli Arnlaugsson 2e4e646c06 fix: toolbar pushed off-screen on iPad Safari with tabs
On iPad Safari with the tab bar visible, `100vh` extends behind the
browser chrome, pushing the fixed-position toolbar out of view.

- Add `viewport-fit=cover` to viewport meta tag
- Use `100dvh` with `100vh` fallback for body/.app height
- Set `--app-height` CSS variable from `visualViewport.height` via JS
- Offset fixed toolbar on iOS Safari using the layout/visual viewport delta
2026-03-08 23:54:15 +00:00
arkon 507423b776 chore: version packages 2026-03-08 16:06:51 +01:00
arkonandClaude Opus 4.6 67d0b0b538 feat: add tunnel status indicator with control panel in header
Green pulsing dot in the desktop header shows when Cloudflare tunnel is active.
Clicking opens a dropdown panel with tunnel URL, remote client count, auth
sessions, and start/stop/QR/revoke controls. Detects tunnel clients via
Cf-Connecting-Ip header to exclude local connections from the count.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-08 16:06:15 +01:00
Ark0N 26cfd8b7ef Merge pull request #33 from arnlaugsson/fix/macos-install-platform-deps
fix: move Linux-only native deps to optionalDependencies
2026-03-08 15:55:16 +01:00
Skúli ArnlaugssonandClaude Opus 4.6 208e6bc175 fix: move Linux-only native deps to optionalDependencies
`@remotion/compositor-linux-x64-gnu` and `@rspack/binding-linux-x64-gnu`
are Linux x64 binaries that cause npm install to fail on macOS (arm64)
with EBADPLATFORM. Moving them to optionalDependencies allows npm to
skip them gracefully on unsupported platforms.

Fixes #32

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 20:59:57 +00:00
arkonandClaude Opus 4.6 6d52b16edc docs: add zerolag demo video to README
Side-by-side comparison of local echo (0ms) vs server echo (600ms-2.7s)
rendered from Remotion ZerolagDemo composition.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 07:43:52 +01:00
arkonandClaude Opus 4.6 4988e85901 docs: add Operation Lightspeed to v0.3.7 changelog
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 07:37:22 +01:00
arkonandClaude Opus 4.6 e799c83b39 chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 07:34:24 +01:00
arkonandClaude Opus 4.6 0717cfbfec refactor: codebase cleanup — dead code, regex helper, centralized constants, tests
- Remove unused validateTokenCounts/validateTokensAndCost exports and PlanPhase type alias
- Add execPattern() helper to eliminate 8 repetitive .lastIndex=0 + exec() loops
- Centralize 11 magic number constants into config/ai-defaults.ts and config/server-timing.ts
- Remove stale src/tui from tsconfig.json exclude
- Fix CLAUDE.md inaccuracies (session helpers, app.js line count, module count)
- Add 316 new tests: LRUMap (38), StaleExpirationMap (42), BufferAccumulator (33),
  respawn helpers (142), system-routes expansion (11→61)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 07:33:44 +01:00
arkonandClaude Opus 4.6 b7b2555dc0 test: add Operation Lightspeed tests — tab switching, local echo, SSE filters
14 new tests covering tab switch SSE reconnect, terminal buffer edge cases
(tail=0, negative, non-numeric, huge values), extractSessionId filtering,
session lifecycle churn, and concurrent SSE client limits.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 07:04:27 +01:00
arkonandClaude Opus 4.6 a609c435fa fix: use TERMINAL_TAIL_SIZE constant and add client-drop recovery
Two hardcoded `256 * 1024` tail sizes in app.js bypassed the
TERMINAL_TAIL_SIZE constant (128KB), causing stale cached browsers
to fetch 256KB buffers even after the constant was reduced to prevent
WebGL GPU stalls. Also adds a self-recovery timer that reloads the
terminal buffer after client-side data drops, preventing permanent
display corruption.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-07 07:00:18 +01:00
arkonandClaude Opus 4.6 3268a12e5e fix: Operation Lightspeed review fixes — padding, dead code, tablet WebGL
- Add SSE padding to backpressure drain write for tunnel clients
- Remove dead SessionTerminal from broadcast padding check
- Trim whitespace in SSE session filter query params
- Remove unused _bufferLazyTerminalData scaffolding code
- Skip WebGL on tablets too, not just phones (canvas fallback)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 06:40:09 +01:00
arkonandClaude Opus 4.6 1f30ed445c perf: Operation Lightspeed — 5 parallel performance optimizations
1. Session-scoped SSE subscriptions: server filters events by session ID,
   clients can subscribe via ?sessions=id1,id2 (backwards-compatible)
2. Lazy xterm.js for subagent windows: terminals created on restore,
   disposed on minimize — saves ~3.75MB DOM at 50 agents
3. Targeted badge updates: badge count changes update the <span> directly
   instead of rebuilding the entire session tab sidebar (O(1) vs O(n))
4. Conditional SSE padding: 8KB Cloudflare padding only on session:terminal
   and session:needsRefresh, not every event (~70% bandwidth reduction)
5. Canvas renderer on mobile: skip WebGL addon on mobile devices to reduce
   GPU pressure and prevent context loss on weaker mobile GPUs

All 5 implemented in parallel via isolated git worktrees, merged conflict-free.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 06:14:16 +01:00
arkonandClaude Opus 4.6 415f02e680 fix: multi-layer backpressure to prevent terminal write freezes
Add three layers of protection against oversized terminal.write() calls
that freeze Chrome's main thread:

1. SSE entry cap: _onSessionTerminal drops data when total queued bytes
   (pendingWrites + flickerFilterBuffer) exceeds 128KB. Server sends
   session:needsRefresh to recover dropped content.

2. Flush cap: flushPendingWrites splits at DEC 2026 sync segment
   boundaries with 64KB per-frame budget. Excess segments deferred to
   next requestAnimationFrame. Segment-level splitting preserves Ink
   redraw atomicity (no flicker).

3. Reduced tail size: initial buffer fetch reduced to 128KB (from 256KB)
   to limit data volume during tab switch.

Re-enable WebGL renderer — root cause was unbounded terminal.write()
volume, not the GPU renderer itself.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-06 17:10:25 +01:00
255 changed files with 33000 additions and 20501 deletions
+5 -2
View File
@@ -11,10 +11,10 @@ jobs:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@v6
- name: Setup Node.js
uses: actions/setup-node@v4
uses: actions/setup-node@v6
with:
node-version: 22
cache: 'npm'
@@ -22,6 +22,9 @@ jobs:
- name: Install dependencies
run: npm ci
- name: Check package-lock.json version sync
run: npm run check:lockfile
- name: Type check
run: npm run typecheck
+3 -3
View File
@@ -16,12 +16,12 @@ jobs:
pull-requests: write
steps:
- name: Checkout repo
uses: actions/checkout@v4
uses: actions/checkout@v6
- name: Setup Node.js
uses: actions/setup-node@v4
uses: actions/setup-node@v6
with:
node-version: 20
node-version: 22
cache: npm
registry-url: https://registry.npmjs.org
+6 -1
View File
@@ -48,7 +48,12 @@ Thumbs.db
# Generated output
out/
screenshots-echo-diag/
tools/remotion/out/
scripts/remotion/out/
# Artifacts that should not be tracked
test-results/
tmp/
public
# Claude Code plan tracking
plan.json
+1 -1
View File
@@ -6,4 +6,4 @@ src/web/public/app.js
src/web/public/styles.css
src/web/public/mobile.css
src/web/public/index.html
tools/
scripts/remotion/
+336
View File
@@ -1,5 +1,341 @@
# aicodeman
## 0.6.3
### Patch Changes
- **Fix**
- Allowlist `opusContext1mEnabled` in `SettingsUpdateSchema`. Without this entry, the strict schema rejected `PUT /api/settings {"opusContext1mEnabled":...}` with `INVALID_INPUT`, so the toggle's value never persisted across reloads. The frontend was already reading and writing this key (`settings-ui.js:336/1137`, `session-ui.js:340`), so saves were silently failing — users never noticed because the load path falls back to `false` on missing keys, hiding the bug. (#78)
## 0.6.2
### Patch Changes
- **Mobile UX**
- Resume Conversation list (welcome page) reworked for narrow screens: 2-line title clamp so more of the first prompt is visible; case-aware subtitle that renders `#caseName` (or `#caseName/sub`) when `workingDir` matches a known case, otherwise falls back to the directory basename; inline `⋯` toggle that expands a detail panel with full prompt, full path, timestamp, size, and short session id; `/Users/<user>/` now collapses to `~/` alongside `/home/<user>/`. (#77)
- Response viewer: ASCII diagram wrap toggle, dedicated mobile code-block layout, and chrome-stripping fallback when the model wraps its reply in extra markup. (#75)
- Mobile keyboard accessory bar no longer triggers vertical scroll. (#72)
**Sessions & settings**
- New `thinkingEffort` setting on session creation, with `xhigh` option and `/effort max` mobile shortcut. (#73)
- `thinkingEffort` is now allowlisted in `SettingsUpdateSchema` so it round-trips through PATCH /api/settings.
- `envOverrides` (`CLAUDE_CODE_*` / `OPENCODE_*`) are now passed to Claude via tmux env exports at spawn time instead of being written to `<case>/.claude/settings.local.json`. Eliminates UI/disk drift; the value lives on `Session._envOverrides`, is exported by `tmux-manager.buildEnvExports()`, and is persisted in `SessionState.envOverrides`. (#74)
**Fixes**
- Eye icon (active-session indicator) now follows `/clear` to the new Claude conversation instead of getting stuck on the previous transcript. (#76)
- `tmux-manager.reconcileSessions` now uses `|` as the field separator, fixing parsing when session names contain other delimiters. (#71)
**Docs**
- CLAUDE.md: added `npm run knip` to the dead-code sweep table and a `Common Gotchas` entry documenting the `envOverrides` → tmux export flow.
## 0.6.1
### Patch Changes
- Internal cleanup and release hygiene:
- **Dead-code sweep via knip**: added `knip.json` for dead-code detection and ran a full sweep — removed unused test files, unused scripts, and narrowed internal module exports to the minimum surface area actually consumed.
- **Lockfile drift prevention**: `version-packages` now runs `npm install --package-lock-only` and verifies the lockfile is in sync via `scripts/check-lockfile-sync.mjs`; CI runs the same check on every push/PR so version drift fails the build instead of reaching production. Resolves the `package-lock.json` / `package.json` version mismatch that shipped in 0.6.0.
- **Docs tightening**: archived 22 completed plan docs from `docs/`, corrected file/handler counts in `CLAUDE.md`, documented the lockfile step in the COM workflow, and removed footer redundancy.
## 0.6.0
### Minor Changes
- Community contributions from @aakhter:
- **feat (#66): Tab reorder shortcuts** — `Ctrl+Shift+{` and `Ctrl+Shift+}` move the active session tab left/right, matching WezTerm convention. Order persists across reloads via `saveSessionOrder()`.
- **feat (#67): Active tab visibility + Alt+N badges** — active tab now has a bright green border with color-matched glow, and the first 9 tabs display number badges hinting at the `Alt+N` switch shortcut. Badges update on reorder/rerender.
- **feat (#68): Clipboard API** — new `POST /api/clipboard` accepting `{text}` broadcasts a `clipboard:write` SSE event; connected browsers attempt `navigator.clipboard.writeText()` with a manual-copy modal fallback when the page isn't focused. Auth-protected via the standard middleware. Useful for pushing snippets from remote sessions to the user's local clipboard.
- **fix (#65): Android Shift+key double character** — pressing `Shift+A` on attached Android keyboards no longer produces "AA". Tracks xterm-handled keydown timestamps and skips the orphaned-input listener for 50ms after a real keydown, while still catching Gboard symbol-keyboard inputs (keyCode 229).
## 0.5.13
### Patch Changes
- Fix "Case path not found" error in Quick Start when `~/codeman-cases/` does not exist (issue #64). Two bugs in `session-ui.js`:
- `runClaude()` auto-create read `createCaseData.case`, but `POST /api/cases` returns `{ success, data: { case } }` — corrected to `createCaseData.data.case`.
- `runShell()` had no auto-create logic and would immediately throw on a missing case directory — now mirrors `runClaude()`'s create-on-demand flow.
## 0.5.12
### Patch Changes
- Fix quick-start to resolve linked cases before codeman-cases fallback. `/api/quick-start` was always resolving `caseName` against `CASES_DIR`, ignoring entries in `~/.codeman/linked-cases.json`. Sessions started via quick-start now correctly honour linked external project directories, consistent with regular case routes.
## 0.5.11
### Patch Changes
- Community contributions and security hardening:
- Mobile response viewer: native-scroll panel for reading full Claude responses with markdown rendering via marked.js (PR #62)
- PWA support: service worker caching, web app manifest, and Android home screen install (PR #59)
- Named Cloudflare tunnel support (PR #58)
- Markdown rendering for response viewer with HTML sanitization (XSS prevention) — strips dangerous elements, event handlers, and javascript: URIs
- Service worker switched from stale-while-revalidate to network-first caching so deploys take effect immediately
- Content-Disposition filename sanitization to prevent header injection in file downloads
- Expose session.muxName public getter, replace unsafe `as any` cast in session-routes
- Static import for execFile in session-routes
- Keyboard shortcut updates: Alt+1-9 tab switching, Shift+Enter newline
- Repo restructure for cleaner GitHub landing page
- Mobile logo, expandable history, session resume fixes
## 0.5.10
### Patch Changes
- fix: allow bracket characters in model validation regex so models like opus[1m] (1M context window) are accepted instead of silently dropped. Quote the model flag value in tmux spawn commands to prevent bash glob expansion of bracket patterns.
docs: update macOS launchd instructions to use `launchctl bootstrap` instead of deprecated `load`. Clean up README install and service sections.
## 0.5.9
### Patch Changes
- Mobile keyboard accessory bar: add configurable "Extended Keyboard Bar" setting (Settings > Display > Input) that toggles between simple mode (up/down arrows, /init, /clear, /compact, paste, dismiss) and extended mode (adds left/right arrows, Tab, Shift+Tab, Ctrl+O, Alt+Enter, Esc). Default is simple mode. Setting is device-specific (not synced to server).
Restyle dismiss button: muted steel-blue tone, fills remaining bar space via flex, larger tap target. Arrow buttons now blue.
Fix paste overlay visibility on mobile: dialog repositioned to top of screen (15vh from top) so the virtual keyboard doesn't cover it. Textarea enlarged for better usability.
(Also includes all v0.5.8 changes: case reorder/delete, XSS sanitization, auto-attach PTY on restart, mobile keyboard buttons, macOS installer fixes, terminal flicker fix, state store collision fix.)
## 0.5.8
### Patch Changes
- Case management: add Manage tab with reorder (up/down arrows) and delete for cases; linked cases are unlinked (folder preserved), CASES_DIR cases are permanently deleted. New endpoints: DELETE /api/cases/:name, PUT /api/cases/order. SSE events: case:deleted, case:order-changed.
Security: sanitize case names from filesystem with /^[a-zA-Z0-9_-]+$/ regex before returning from GET /api/cases to prevent XSS via maliciously-named directories reaching frontend inline onclick handlers.
Auto-attach PTY: server now calls startInteractive() for recovered tmux sessions during startup so all sessions resume capturing output immediately after deploy, instead of waiting for client selection. Frontend auto-attach condition relaxed from (pid===null && status==='idle') to (pid===null && !\_ended).
Mobile keyboard accessory: add Shift+Tab, Tab, Esc, Alt+Enter, Left/Right arrow, and Ctrl+O buttons.
Terminal: fix flicker regression by moving viewport clear inside dimension guard.
State store: fix temp file collisions on concurrent writes.
macOS: fix installer failures when piped via curl | bash, add HTML cache support, launchd service template, and trust dialog handling.
Housekeeping: remove accidentally committed dist/state-store.js build artifact.
## 0.5.7
### Patch Changes
- feat: support "Default (CLI default)" option for model selection. Adds a new empty-value option to the model dropdown that defers to the CLI's own default model instead of forcing a specific model. Ensures empty defaultModel values are treated as undefined when passed to session creation and Ralph loop start, preventing empty strings from being sent as model flags.
## 0.5.6
### Patch Changes
- fix: default new sessions to opus[1m] (1M context window) instead of plain opus (200k context)
## 0.5.5
### Patch Changes
- Add 1M Opus context quick setting — per-case and global toggle that writes `model: "opus[1m]"` to `.claude/settings.local.json` when creating new sessions. Fix mobile layout: banners (respawn, timer, orchestrator) between header and main content now visible by switching from margin-top on `.main` to padding-top on `.app`. Add tablet-optimized respawn banner styles and mobile phone banner refinements.
## 0.5.4
### Patch Changes
- Fix terminal flicker regression — re-add server-side DEC 2026 synchronized output wrapping around batched terminal data. Ink spinner frames (cursor-up + redraw cycles) do not emit their own DEC 2026 markers, so without the server wrapper each partial cursor update rendered individually causing visible flicker. Also: extract SSE stream management, session listener wiring, and respawn event wiring from server.ts into dedicated modules; deduplicate error message extraction across 7 files with shared getErrorMessage() helper; update SSE event count in CLAUDE.md (106 → 117).
## 0.5.3
### Patch Changes
- Readability refactor across 12 core files, extracting ~35 helper methods to reduce duplication:
- state-store: extract serializeState(), split assembleStateJson() into focused sub-methods
- session: extract \_resetBuffers() (3x dedup), \_clearAllTimers() (10 timer cleanups), \_handleJsonMessage()
- ralph-tracker: extract completeAllTodos() (4x dedup), emitValidationWarning(), named similarity constants
- subagent-watcher: extract markSubagentAsCompleted(), extractFirstTextContent(), emitToolResult(), findOldestInactiveAgent()
- respawn-controller: extract recoveryResetToWatching(), canAutoAccept(), formatRemainingSeconds(), validatePositiveTimeout()
- tmux-manager: replace 15 path.includes() with UNSAFE_PATH_CHARS regex, extract buildEnvExports/buildPathExport/\_configureOpenCode helpers
- session-auto-ops: extract executeWhenIdle() shared retry helper, convert to options object, add validateThreshold()
- app.js: add \_clearTimer() (11 call sites), \_isStaleSelect(), keyboard shortcut lookup table, \_cleanupPreviousSession(), \_resetAllAppState()
- route-helpers: add readJsonConfig() (5 inline patterns replaced), validateSessionFilePath() (2 duplicated blocks replaced)
## 0.5.2
### Patch Changes
- Make buffer size limits configurable via CODEMAN\_\* environment variables (MAX_TERMINAL_BUFFER, TRIM_TERMINAL_TO, MAX_TEXT_OUTPUT, TRIM_TEXT_TO, MAX_MESSAGES), falling back to existing defaults. Allows users with fewer sessions or more RAM to tune buffer sizes without patching source.
Fix duplicate terminal output on tab switch to busy sessions by clearing the terminal before writing the new buffer.
Fix stale Ink CUP frames after tab switch by sending Ctrl+L to force a clean redraw.
Fix mobile CJK input handling: resolve textarea positioning, terminal flicker during composition, and layout overflow on small screens. Improve CJK composition lifecycle with better event handling and fallback flush timers.
## 0.5.1
### Patch Changes
- refactor: codebase cleanup — extract route helpers, eliminate boilerplate, optimize hot paths
- Add `parseBody()` helper to route-helpers.ts: validates request body against Zod schema with structured 400 error on failure, replacing 37 identical safeParse + error-check blocks across 10 route files
- Add `persistAndBroadcastSession()` helper: combines persist + SessionUpdated broadcast into one call, replacing 5 repeated 2-line pairs
- Migrate session-routes.ts to use `findSessionOrFail()` consistently (17 inline session lookups replaced) and `parseBody()` (12 patterns)
- Migrate ralph-routes.ts to use `findSessionOrFail()` (9 lookups) and `parseBody()` (4 patterns)
- Migrate 8 remaining route files to use `parseBody()` (21 patterns total)
- Fix O(n log n) eviction in bash-tool-parser.ts: replace `Array.from().sort()[0]` with O(n) min-scan for oldest active tool
- Extract `_debouncedCall()` utility in frontend: replaces 4 manual debounce patterns (7 lines each → 1 line) in app.js, panels-ui.js, ralph-panel.js
- Net reduction: 208 lines removed across 16 files
## 0.5.0
### Minor Changes
- Visual redesign with glass morphism, refined colors, and polished UI. Optimize history endpoint with buffer reuse and line iterator. Fix Ink frame search window (4KB→64KB) to prevent partial frames. Fix stale terminal data on tab switch via chunkedTerminalWrite cancellation. Improve history prompt extraction with expanded command filtering and tail scan fallback. Align case select group height to match dropdown. Fix no-control-regex lint error for ANSI strip pattern. Add browser-testing-guide to CLAUDE.md references.
## 0.4.7
### Patch Changes
- feat: improve session navigability in history and monitor panel (closes #45)
- History items now show the first user prompt as the title with the project path as a subtitle, making it much easier to distinguish sessions from the same project
- The `/api/history/sessions` endpoint extracts the first user message from each transcript JSONL, stripping system-injected XML tags and command artifacts, truncating to 120 chars
- Monitor panel session rows are now clickable — clicking navigates directly to that session's tab via `selectSession()`; Kill button retains independent behavior via `stopPropagation()`
- Updated CLAUDE.md architecture tables to reflect Orchestrator Loop additions (14 route modules, 15 type files, orchestrator domain files, orchestrator-panel.js frontend module)
- fix: stop subagent monitor windows from auto-opening on discovery
- feat: add Orchestrator Loop with phased plan execution, live progress during plan generation, and toolbar button (hidden until fully tested)
- fix: patch 3 production bugs found during deep audit
- fix: restore mobile terminal scrollback using JS scrollLines() instead of broken native scroll
## 0.4.6
### Patch Changes
- Fix mobile keyboard scroll and layout issues:
- Prevent iOS Safari from scrolling the page when typing with the keyboard open (position:fixed on .app + window.scroll reset)
- Eliminate dead space between terminal and keyboard accessory bar by removing redundant CSS padding, tightening JS padding constant, and adding row quantization gap compensation
- Fix toolbar overlapping terminal content when keyboard is hidden by adding proper padding-bottom to .main, including iOS Safari bottom bar offset
- Strip Ink spinner bloat from terminal buffer before tailing
- Fix resolveCasePath priority order and suppress JSON parse warnings
## 0.4.5
### Patch Changes
- Fix mobile keyboard toolbar positioning on iOS Safari: toolbar (Run/Stop/Run Shell) was hidden behind the accessory bar when virtual keyboard was active due to overlapping CSS positions. Remove the aggressive safety check in `updateLayoutForKeyboard()` that incorrectly dismissed keyboard state when iOS scrolled the visual viewport during typing. Add Safari-bar CSS offset to accessory bar so it properly stacks above the toolbar. Remove the double-counted Safari-bar offset when keyboard is visible since the JS transform already covers the full distance.
## 0.4.4
### Patch Changes
- fix: mobile keyboard hides terminal content on iPhone
Fixed a bug where opening the virtual keyboard on iPhone left zero visible terminal space. Two independent mechanisms were both accounting for the keyboard height: `MobileDetection.updateAppHeight()` shrunk `--app-height` to the visual viewport height, while `KeyboardHandler.updateLayoutForKeyboard()` added a large `paddingBottom`. These double-counted, leaving negative space for the terminal (user saw accessory bar + toolbar but no terminal content).
Fix: `updateAppHeight()` now skips when the keyboard is visible, and `handleViewportResize()` restores `--app-height` to the pre-keyboard value on first detection (since MobileDetection's listener fires before KeyboardHandler's). On keyboard close, `--app-height` is re-synced to the current visual viewport.
## 0.4.3
### Patch Changes
- Refactor case routes: extract readLinkedCases() and resolveCasePath() helpers to eliminate 6x duplicated linked-cases.json path construction and 5x duplicated file read/parse logic. Replace O(n) .some() duplicate check with O(1) Set.has() in case listing. Un-export unused isError() type guard. Standardize reply.status() to reply.code() in system routes. Update CLAUDE.md frontend module listing and SSE event count.
## 0.4.2
### Patch Changes
- Extract monolithic app.js (~12.5K lines) into 6 focused domain modules that extend CodemanApp.prototype via Object.assign: terminal-ui.js (terminal setup, rendering pipeline, controls), respawn-ui.js (respawn banner, countdown, presets, run summary), ralph-panel.js (Ralph state panel, fix_plan, plan versioning), settings-ui.js (app settings, visibility, web push, tunnel/QR, help), panels-ui.js (subagent panel, teams, insights, file browser, log viewer), session-ui.js (quick start, session options, case settings). Fix critical deferred script init ordering bug: wrap CodemanApp instantiation in DOMContentLoaded so all defer'd mixin modules execute their Object.assign before the constructor runs. Guard missing cleanupWizardDragging() call in subagent-windows.js. Update build.mjs to minify/hash all new modules.
## 0.4.1
### Patch Changes
- Performance optimizations: V8 compile cache for 10-20% faster cold starts, lazy-load WebGL addon (244KB saved on mobile), preload hints for critical scripts, batch tmux reconciliation (N subprocess calls → 1). Also: WebSocket session lifecycle fixes, CJK IME input support, CI upgrade to Node 24/actions v6, install.sh fork support, and CLAUDE.md/README documentation refresh.
## 0.4.0
### Minor Changes
- Add CJK IME input textarea for xterm.js terminal (env toggle INPUT_CJK_FORM=ON). Always-visible textarea below terminal handles native browser IME composition, forwarding completed text to PTY on Enter. Supports arrow keys, Ctrl combos, backspace passthrough, and Escape to clear.
Add fork installation support to install.sh with CODEMAN_REPO_URL and CODEMAN_BRANCH env vars, allowing custom repository and branch for git clone/update operations. README updated with fork installation instructions.
Fix WebSocket session lifecycle: close WS connections when session exits (prevents orphaned listeners and stale writes to dead PTY), add readyState guard in onTerminal to stop buffering after socket closes, simplify heartbeat by removing redundant alive flag.
Add WebSocket reconnection with exponential backoff (1s-10s) on unexpected close, skipping server rejection codes (4004/4008/4009). Falls back gracefully to SSE+POST during reconnection.
Clear CJK textarea on session switch to prevent sending stale text to wrong session.
## 0.3.12
### Patch Changes
- Add WebSocket terminal I/O with server-side DEC 2026 synchronized update markers. Replaces per-keystroke HTTP POST + SSE terminal output with a single bidirectional WebSocket connection for dramatically lower input latency. Server-side 8ms micro-batching with 16KB flush threshold groups rapid PTY events into single WS frames wrapped in DEC 2026 markers for flicker-free atomic rendering. Includes 30s ping/pong heartbeat with 10s timeout for stale connection detection through tunnels. Existing SSE + HTTP POST paths remain fully functional as transparent fallback. Resize messages validated to match HTTP route bounds (cols 1-500, rows 1-200, integers only). 16 automated route tests added for WS endpoint. Also patches 5 dependency vulnerabilities (basic-ftp, fastify, minimatch, serialize-javascript).
## 0.3.11
### Patch Changes
- ### Session Resume & History
- Add `resumeSessionId` support for conversation resume after reboot
- Add history session resume UI and API with route shell sessions routing fix
- Improve session resume reliability and persist user settings across refresh
- Correct `claudeSessionId` for resumed sessions
### Terminal & Frontend
- Upgrade xterm.js 5.3 → 6.0 with native DEC 2026 synchronized output
- Increase terminal scrollback from 5,000 to 20,000 lines
- Reduce default font size and persist tab state across refresh
- Resolve terminal resize scrollback ghost renders
- Hide subagent monitor panel by default
### Installer
- Auto-detect existing install and run update instead of fresh install
- Auto-restart codeman-web service after update if running
- Show restart command when codeman-web is not a systemd service
- Fix one-liner restart command for background processes
### Codebase Quality
- Remove dead code, consolidate imports, extract constants
- Repair 15 pre-existing subagent-watcher test failures
- Clean up DEC sync dead code
## 0.3.10
### Patch Changes
- - feat: upgrade xterm.js from 5.3 to 6.0 with native DEC 2026 synchronized output support
- feat: add history session resume UI and API — resume Claude conversations after reboot
- feat: add resumeSessionId support for conversation resume across session restarts
- feat: persist active tabs across page refresh
- feat: improve session resume reliability and persist user settings
- perf: increase terminal scrollback from 5,000 to 20,000 lines
- fix: resolve terminal resize scrollback ghost renders
- fix: route shell sessions to correct endpoint on tab click
- fix: correct claudeSessionId for resumed sessions (use original Claude conversation ID)
- fix: increase default desktop font size from 12 to 14
- refactor: extract shared \_fetchHistorySessions() method to eliminate duplication
- refactor: remove dead DEC 2026 sync code (extractSyncSegments, DEC_SYNC_START/END constants)
## 0.3.9
### Patch Changes
- Add content-hash cache busting for static assets — build step now renames JS/CSS files with MD5 content hashes (e.g. app.js → app.94b71235.js) and rewrites index.html references. HTML served with Cache-Control: no-cache so browsers always revalidate and pick up new hashed filenames after deploys. Hashed assets keep immutable 1-year cache. Eliminates the need for manual hard refresh (Ctrl+Shift+R) after deployments.
Refactor path traversal validation into shared validatePathWithinBase() helper in route-helpers.ts, replacing 6 duplicate inline checks across case-routes, plan-routes, and session-routes.
Deduplicate stripAnsi in bash-tool-parser.ts — use shared utility from utils/index.ts instead of private method.
## 0.3.8
### Patch Changes
- Add tunnel status indicator with control panel — green pulsing dot in header when Cloudflare tunnel is active, dropdown with URL, remote clients, auth sessions, and start/stop/QR/revoke controls
## 0.3.7
### Patch Changes
- Operation Lightspeed: 5 parallel performance optimizations — multi-layer backpressure to prevent terminal write freezes, TERMINAL_TAIL_SIZE constant with client-drop recovery, tab switching SSE gating, and local echo improvements
- Codebase cleanup: remove dead code (unused token validation exports, PlanPhase alias), add execPattern() regex helper to eliminate repetitive .lastIndex resets, centralize 11 magic number constants into config files, fix CLAUDE.md inaccuracies, and add 316 new tests for utilities, respawn helpers, and system-routes
## 0.3.6
### Patch Changes
+37 -47
View File
@@ -6,11 +6,12 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
| Task | Command |
|------|---------|
| Dev server | `npx tsx src/index.ts web` |
| Dev server | `npm run dev` (or `npx tsx src/index.ts web`) |
| Type check | `tsc --noEmit` |
| Lint | `npm run lint` (fix: `npm run lint:fix`) |
| Format | `npm run format` (check: `npm run format:check`) |
| Single test | `npx vitest run test/<file>.test.ts` |
| Single test | `npm test -- test/<file>.test.ts` (or `npx vitest run --config config/vitest.config.ts test/<file>.test.ts`) |
| Build | `npm run build` (esbuild via `scripts/build.mjs`, NOT tsc — `tsc --noEmit` is type-check only) |
| Production | `npm run build && systemctl --user restart codeman-web` |
## CRITICAL: Session Safety
@@ -44,15 +45,17 @@ When user says "COM":
"aicodeman": patch
---
Description of changes
Detailed description of ALL changes since last release (not just the most recent commit — review full git log since last version tag)
CHANGESET
```
Replace `patch` with `minor` or `major` as needed. Include `"xterm-zerolag-input": patch` on a separate line if that package changed too.
3. **Consume the changeset**: `npm run version-packages` (bumps versions in `package.json` files and updates `CHANGELOG.md`)
3. **Consume the changeset**: `npm run version-packages` (auto-bumps `package.json` files, updates `CHANGELOG.md`, runs `npm install --package-lock-only`, and verifies lockfile sync via `scripts/check-lockfile-sync.mjs` — all in one command; never hand-edit `CHANGELOG.md` or `package-lock.json` versions)
4. **Sync CLAUDE.md version**: Update the `**Version**` line below to match the new version from `package.json`
5. **Commit and deploy**: `git add -A && git commit -m "chore: version packages" && git push && npm run build && systemctl --user restart codeman-web`
**Version**: 0.3.6 (must match `package.json`)
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
**Version**: 0.6.3 (must match `package.json`)
## Project Overview
@@ -75,19 +78,21 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| Dev with TLS | `npx tsx src/index.ts web --https` |
| Continuous typecheck | `tsc --noEmit --watch` |
| Test coverage | `npm run test:coverage` |
| Dead-code sweep | `npm run knip` (config in `knip.json`) |
| Production start | `npm run start` |
| Production logs | `journalctl --user -u codeman-web -f` |
**CI**: `.github/workflows/ci.yml` runs `typecheck`, `lint`, `format:check` on push to master (Node 22). Tests excluded (they spawn tmux).
**CI**: `.github/workflows/ci.yml` runs `check:lockfile`, `typecheck`, `lint`, `format:check` on push to master/main and on PRs (Node 22). Tests excluded (they spawn tmux).
**Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`). ESLint allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Does not lint `app.js` or `scripts/**/*.mjs`.
**Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`). ESLint flat config (`config/eslint.config.js`) allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Ignores: `app.js`, `scripts/**/*.mjs`, `src/web/public/vendor/**`, `scripts/remotion/**`.
## Common Gotchas
- **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink
- **ESM only** — Never `require()`, use `await import()`. `tsx` masks CJS/ESM issues in dev but production breaks
- **Package ≠ product name** — npm: `aicodeman`, product: **Codeman**. Release renames tags accordingly
- **Global regex `lastIndex`** — Use `createAnsiPatternFull/Simple()` factories, not shared `g`-flag patterns in loops
- **Global regex `lastIndex`** — Shared `g`-flag patterns in loops must reset `lastIndex = 0` first, or use the `execPattern()` helper in `utils/regex-patterns.ts` (resets automatically)
- **`envOverrides` flow `CLAUDE_CODE_*` / `OPENCODE_*` env vars** — Set via `POST /api/sessions { envOverrides }`, stored on `Session._envOverrides`, exported by `tmux-manager.buildEnvExports()` at spawn time, persisted in `SessionState.envOverrides`. **Do NOT** write these to `<case>/.claude/settings.local.json` — that's the old path and creates UI/disk drift
**Import conventions**: Utils from `./utils`, types from `./types` (barrel), config from specific `./config/*` files.
@@ -98,27 +103,28 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| Domain | Key files | Notes |
|--------|-----------|-------|
| **Entry** | `src/index.ts`, `src/cli.ts` | |
| **Session** | `src/session.ts` ★, `src/session-manager.ts`, `src/session-auto-ops.ts`, `src/session-cli-builder.ts` | |
| **Session** | `src/session.ts` ★, `src/session-manager.ts`, `src/session-auto-ops.ts`, `src/session-cli-builder.ts`, `src/session-lifecycle-log.ts`, `src/session-task-cache.ts` | |
| **Mux** | `src/mux-interface.ts`, `src/mux-factory.ts`, `src/tmux-manager.ts` | |
| **Respawn** | `src/respawn-controller.ts` ★ + 4 helpers (`-adaptive-timing`, `-health`, `-metrics`, `-patterns`) | Read `docs/respawn-state-machine.md` first |
| **Ralph** | `src/ralph-tracker.ts` ★, `src/ralph-loop.ts` + 5 helpers (`-config`, `-fix-plan-watcher`, `-plan-tracker`, `-stall-detector`, `-status-parser`) | Read `docs/ralph-wiggum-guide.md` first |
| **Orchestrator** | `src/orchestrator-loop.ts`, `src/orchestrator-planner.ts`, `src/orchestrator-verifier.ts` | Read `docs/orchestrator-loop-architecture.md` first |
| **Agents** | `src/subagent-watcher.ts` ★, `src/team-watcher.ts`, `src/bash-tool-parser.ts`, `src/transcript-watcher.ts` | |
| **AI** | `src/ai-checker-base.ts`, `src/ai-idle-checker.ts`, `src/ai-plan-checker.ts` | |
| **Tasks** | `src/task.ts`, `src/task-queue.ts`, `src/task-tracker.ts` | |
| **State** | `src/state-store.ts`, `src/run-summary.ts`, `src/session-lifecycle-log.ts` | |
| **Infra** | `src/hooks-config.ts`, `src/push-store.ts`, `src/tunnel-manager.ts`, `src/image-watcher.ts`, `src/file-stream-manager.ts` | |
| **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/claude-md.ts` | |
| **Web** | `src/web/server.ts`, `src/web/sse-events.ts`, `src/web/routes/*.ts` (13 modules), `src/web/ports/*.ts`, `src/web/middleware/auth.ts`, `src/web/schemas.ts` | |
| **Frontend** | `src/web/public/app.js` ★ (~11.8K lines) + 10 JS modules (incl. `sw.js` service worker) | |
| **Types** | `src/types/index.ts` → 14 domain files | See `@fileoverview` in index.ts |
| **Web** | `src/web/server.ts`, `src/web/sse-events.ts`, `src/web/routes/*.ts` (15 route modules + barrel), `src/web/route-helpers.ts`, `src/web/ports/*.ts`, `src/web/middleware/auth.ts`, `src/web/schemas.ts` | |
| **Frontend** | `src/web/public/app.js` (~2.9K lines, core) + 5 infra modules (`constants.js`, `mobile-handlers.js`, `voice-input.js`, `notification-manager.js`, `keyboard-accessory.js`) + 7 domain modules (`terminal-ui.js`, `respawn-ui.js`, `ralph-panel.js`, `orchestrator-panel.js`, `settings-ui.js`, `panels-ui.js`, `session-ui.js`) + 4 feature modules (`ralph-wizard.js`, `api-client.js`, `subagent-windows.js`, `input-cjk.js`) + `sw.js` | |
| **Types** | `src/types/index.ts` (barrel) → 14 domain files; also `src/types.ts` root re-export | See `@fileoverview` in index.ts |
★ = Large file (>50KB). All files have `@fileoverview` JSDoc — read that before diving in.
★ = Large file (>50KB). All files have `@fileoverview` JSDoc — read that before diving in. Discovery aid: `grep -l '@fileoverview' src/web/routes/*.ts` lists all route modules; same grep works for `src/types/`, `src/web/public/*.js`.
**Local package**: `packages/xterm-zerolag-input/` — local echo overlay for xterm.js; copy embedded in `app.js`.
**Config**: `src/config/` — 9 files. Import from specific files, not barrel.
**Utilities**: `src/utils/` — re-exported via index. Key: `CleanupManager`, `LRUMap`, `StaleExpirationMap`, `BufferAccumulator`, `stripAnsi`, `Debouncer`.
**Utilities**: `src/utils/` — re-exported via index. Key: `CleanupManager`, `LRUMap`, `StaleExpirationMap`, `BufferAccumulator`, `stripAnsi`, `Debouncer`, `KeyedDebouncer`. Also: `claude-cli-resolver`/`opencode-cli-resolver` (CLI path resolution), `string-similarity` (fuzzy matching), `regex-patterns` (ANSI/token/spinner patterns), `assertNever` (exhaustive checks), `token-validation` (auth tokens), `nice-wrapper` (process priority).
### Data Flow
@@ -135,7 +141,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Hook events**: Claude Code hooks trigger via `/api/hook-event`. Key events: `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`. See `src/hooks-config.ts`.
**Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `agent-teams/`.
**Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `docs/agent-teams/`.
**Circuit breaker**: Prevents respawn thrashing. States: `CLOSED` → `HALF_OPEN` → `OPEN`. Reset: `/api/sessions/:id/ralph-circuit-breaker/reset`.
@@ -143,13 +149,13 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
### Frontend
Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `app.js`(6) → `ralph-wizard.js`(7) → `api-client.js`(8) → `subagent-windows.js`(9).
Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `input-cjk.js`(5.5) → `app.js`(6) → `terminal-ui.js`(7) → `respawn-ui.js`(8) → `ralph-panel.js`(9) → `orchestrator-panel.js`(9.5) → `settings-ui.js`(10) → `panels-ui.js`(11) → `session-ui.js`(12) → `ralph-wizard.js`(13) → `api-client.js`(14) → `subagent-windows.js`(15). `input-cjk.js` handles CJK IME composition via an always-visible textarea below the terminal (`window.cjkActive` blocks xterm's onData).
**Z-index layers**: subagent windows (1000), plan agents (1100), log viewers (2000), image popups (3000), local echo overlay (7).
**Respawn presets**: `solo-work` (3s/60min), `subagent-workflow` (45s/240min), `team-lead` (90s/480min), `ralph-todo` (8s/480min), `overnight-autonomous` (10s/480min).
**Keyboard shortcuts**: Escape (close), Ctrl+? (help), Ctrl+Enter (quick start), Ctrl+W (kill), Ctrl+Tab (next), Ctrl+K (kill all), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl/Cmd +/- (font).
**Keyboard shortcuts**: Escape (close), Ctrl+? (help), Ctrl+W (kill), Ctrl+Tab (next), Alt+1-9 (switch tab), Ctrl+Shift+{/} (move tab left/right), Shift+Enter (newline), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl+Shift+V (voice input), Ctrl/Cmd +/- (font).
### Security
@@ -166,11 +172,11 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
### SSE Event Registry
~100 event types in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). Both must be kept in sync.
~120 event types in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). Both must be kept in sync.
### API Routes
~111 handlers across 13 route files in `src/web/routes/`: system (35), sessions (24), ralph (9), plan (8), respawn (7), cases (7), files (5), mux (5), scheduled (4), push (4), teams (2), hooks (1). Each file has `@fileoverview` with endpoint details.
~128 handlers across 15 route files in `src/web/routes/`: system (36), sessions (27), orchestrator (10), cases (9), ralph (9), plan (8), respawn (7), files (5), mux (5), push (4), scheduled (4), teams (2), hooks (1), clipboard (1), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
## Adding Features
@@ -179,7 +185,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
- **Session setting**: Add to `SessionState`, include in `session.toState()`, call `persistSessionState()`
- **Hook event**: Add to `HookEventType`, add hook in `hooks-config.ts:generateHooksConfig()`, update `HookEventSchema`
- **Mobile feature**: Add to relevant singleton, guard with `MobileDetection.isMobile()`
- **New test**: Pick unique port (search `const PORT =`). Integration: ports 3099-3211. Route tests: `app.inject()` — see `test/routes/_route-test-utils.ts`.
- **New test**: Pick unique port (search `const PORT =`). Route tests use `app.inject()` (no port needed) — see `test/routes/_route-test-utils.ts`.
**Validation**: Zod v4 (different API from v3). Define schemas in `schemas.ts`, use `.parse()`/`.safeParse()`.
@@ -192,22 +198,20 @@ All in `~/.codeman/`: `state.json` (sessions, settings, respawn), `mux-sessions.
**CRITICAL: You are running inside a Codeman-managed tmux session.** Never run `npx vitest run` (full suite) — it spawns/kills tmux sessions and will crash your own session. Only run individual files:
```bash
npx vitest run test/<specific-file>.test.ts # Single file (SAFE)
npx vitest run -t "pattern" # By name (SAFE)
# npx vitest run # DANGEROUS — DON'T DO THIS
npm test -- test/<specific-file>.test.ts # Single file (SAFE, uses config/vitest.config.ts)
npm test -- -t "pattern" # By name (SAFE)
# npm test # DANGEROUS — runs full suite, DON'T DO THIS
```
Raw `npx vitest` skips `config/vitest.config.ts`; always use `npm test --` or pass `--config config/vitest.config.ts`.
**Config**: Vitest with `globals: true`, `fileParallelism: false`. Timeout 30s, teardown 60s.
**Safety**: `test/setup.ts` snapshots pre-existing tmux sessions and never kills them. Only `registerTestTmuxSession()` sessions get cleaned up.
**Ports**: Pick unique ports manually. Search `const PORT =` before adding new tests.
**Respawn tests**: Use `MockSession` from `test/respawn-test-utils.ts`. **Route tests**: `app.inject()` in `test/routes/`. **Mobile tests**: Playwright suite in `mobile-test/` (135 device profiles).
## Screenshots
Mobile screenshots in `~/.codeman/screenshots/`. API: `GET /api/screenshots`, `POST /api/screenshots`.
**Respawn tests**: Use `MockSession` from `test/respawn-test-utils.ts`. **Route tests**: `app.inject()` in `test/routes/`. **Mobile tests**: Playwright suite in `test/mobile/` (135 device profiles).
## Debugging
@@ -219,28 +223,14 @@ curl localhost:3000/api/subagents | jq # Background agents
cat ~/.codeman/state.json | jq # Persisted state
```
Mobile screenshots: `~/.codeman/screenshots/`, accessed via `GET/POST /api/screenshots`.
## Performance & Limits
Target: 20 sessions, 50 agent windows at 60fps. Limits in `src/config/`: terminal 2MB, text 1MB, messages 1000, max agents 500, max sessions 50, max SSE clients 100. Use `LRUMap` for bounded caches, `StaleExpirationMap` for TTL cleanup. Anti-flicker pipeline: `docs/terminal-anti-flicker.md`.
## References
**Memory leaks (24+ hour sessions)**: use `CleanupManager`, clear Maps in `stop()`, guard async with `if (this.cleanup.isStopped) return`. Frontend: store handler refs, clean in `close*()`. Verify: `npm test -- test/memory-leak-prevention.test.ts`.
Deep-dive docs in `docs/`: `respawn-state-machine.md`, `ralph-wiggum-guide.md`, `claude-code-hooks-reference.md`, `terminal-anti-flicker.md`, `opencode-integration.md`, `qr-auth-plan.md`. Agent Teams: `agent-teams/README.md`. SSE events: `src/web/sse-events.ts` + `constants.js`.
## Scripts & Tunnel
## Scripts
Key: `scripts/tmux-manager.sh` (safe tmux mgmt), `scripts/tunnel.sh` (tunnel start/stop/url). Production: `scripts/codeman-web.service`, `scripts/codeman-tunnel.service`.
## Memory Leak Prevention
24+ hour sessions: use `CleanupManager`, clear Maps in `stop()`, guard async with `if (this.cleanup.isStopped) return`. Frontend: store handler refs, clean in `close*()`. Verify: `npx vitest run test/memory-leak-prevention.test.ts`.
## Common Workflows
**Bug investigation**: Dev server → reproduce in browser → check terminal + `~/.codeman/state.json`.
**API endpoint**: Types in `src/types/*.ts` → route in `src/web/routes/*-routes.ts` → SSE event if needed → handle in `app.js`.
**Respawn changes**: Read `docs/respawn-state-machine.md` first. Use `MockSession` from `test/respawn-test-utils.ts`.
## Tunnel
`./scripts/tunnel.sh start|stop|url`. **Always set `CODEMAN_PASSWORD`** before exposing via tunnel.
Key scripts: `scripts/tmux-manager.sh` (safe tmux mgmt), `scripts/tunnel.sh start|stop|url` (tunnel). Production services: `scripts/codeman-web.service`, `scripts/codeman-tunnel.service`. **Always set `CODEMAN_PASSWORD`** before exposing via tunnel.
+57 -11
View File
@@ -11,7 +11,7 @@
<p align="center">
<a href="https://opensource.org/licenses/MIT"><img src="https://img.shields.io/badge/License-MIT-1e3a5f?style=flat-square" alt="License: MIT"></a>
<a href="https://nodejs.org/"><img src="https://img.shields.io/badge/Node.js-18%2B-22c55e?style=flat-square&logo=node.js&logoColor=white" alt="Node.js 18+"></a>
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/TypeScript-5.5-3b82f6?style=flat-square&logo=typescript&logoColor=white" alt="TypeScript 5.5"></a>
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/TypeScript-5.9-3b82f6?style=flat-square&logo=typescript&logoColor=white" alt="TypeScript 5.9"></a>
<a href="https://fastify.dev/"><img src="https://img.shields.io/badge/Fastify-5.x-1e3a5f?style=flat-square&logo=fastify&logoColor=white" alt="Fastify"></a>
<img src="https://img.shields.io/badge/Tests-1435%20total-22c55e?style=flat-square" alt="Tests">
</p>
@@ -28,29 +28,67 @@
curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash
```
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it. You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code) or [OpenCode](https://opencode.ai) (or both). After install:
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code) or [OpenCode](https://opencode.ai) (or both). After install:
```bash
codeman web
# Open http://localhost:3000 — press Ctrl+Enter to start your first session
```
**Update to latest version:**
```bash
curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash -s update
```
<details>
<summary><strong>Run as a background service</strong></summary>
**Linux (systemd):**
```bash
mkdir -p ~/.config/systemd/user && printf '[Unit]\nDescription=Codeman Web Server\nAfter=network.target\n\n[Service]\nType=simple\nExecStart=%s %s/dist/index.js web\nRestart=always\nRestartSec=10\n\n[Install]\nWantedBy=default.target\n' "$(which node)" "$HOME/.codeman/app" > ~/.config/systemd/user/codeman-web.service && systemctl --user daemon-reload && systemctl --user enable --now codeman-web && loginctl enable-linger $USER
mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/codeman-web.service << EOF
[Unit]
Description=Codeman Web Server
After=network.target
[Service]
Type=simple
ExecStart=$(which node) $HOME/.codeman/app/dist/index.js web
Restart=always
RestartSec=10
[Install]
WantedBy=default.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now codeman-web
loginctl enable-linger $USER
```
**macOS (launchd):**
```bash
mkdir -p ~/Library/LaunchAgents && printf '<?xml version="1.0" encoding="UTF-8"?>\n<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">\n<plist version="1.0"><dict><key>Label</key><string>com.codeman.web</string><key>ProgramArguments</key><array><string>%s</string><string>%s/dist/index.js</string><string>web</string></array><key>RunAtLoad</key><true/><key>KeepAlive</key><true/><key>StandardOutPath</key><string>/tmp/codeman.log</string><key>StandardErrorPath</key><string>/tmp/codeman.log</string></dict></plist>\n' "$(which node)" "$HOME/.codeman/app" > ~/Library/LaunchAgents/com.codeman.web.plist && launchctl load ~/Library/LaunchAgents/com.codeman.web.plist
mkdir -p ~/Library/LaunchAgents
cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
"http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.codeman.web</string>
<key>ProgramArguments</key>
<array>
<string>$(which node)</string>
<string>$HOME/.codeman/app/dist/index.js</string>
<string>web</string>
</array>
<key>RunAtLoad</key><true/>
<key>KeepAlive</key><true/>
<key>StandardOutPath</key>
<string>/tmp/codeman.log</string>
<key>StandardErrorPath</key>
<string>/tmp/codeman.log</string>
</dict>
</plist>
EOF
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
```
</details>
@@ -143,6 +181,10 @@ Watch background agents work in real-time. Codeman monitors agent activity and d
## Zero-Lag Input Overlay
<p align="center">
<img src="docs/images/zerolag-demo.gif" alt="Zerolag Demo — local echo vs server echo side-by-side" width="900">
</p>
When accessing your coding agent remotely (VPN, Tailscale, SSH tunnel), every keystroke normally takes 200-300ms to round-trip. Codeman implements a **Mosh-inspired local echo system** that makes typing feel instant regardless of latency.
A pixel-perfect DOM overlay inside xterm.js renders keystrokes at 0ms. Background forwarding silently sends every character to the PTY in 50ms debounced batches, so Tab completion, `Ctrl+R` history search, and all shell features work normally. When the server echo arrives 200-300ms later, the overlay seamlessly disappears and the real terminal text takes over — the transition is invisible.
@@ -359,9 +401,12 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
| `Ctrl+Enter` | Quick-start session |
| `Ctrl+W` | Close session |
| `Ctrl+Tab` | Next session |
| `Alt+1`–`Alt+9` | Switch to tab N |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move active tab left / right |
| `Ctrl+K` | Kill all sessions |
| `Ctrl+L` | Clear terminal |
| `Ctrl+Shift+R` | Restore terminal size |
| `Ctrl+Shift+V` | Toggle voice input |
| `Ctrl/Cmd +/-` | Font size |
| `Escape` | Close panels |
@@ -404,6 +449,7 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
| `GET` | `/api/events` | SSE stream |
| `GET` | `/api/status` | Full app state |
| `POST` | `/api/hook-event` | Hook callbacks |
| `POST` | `/api/clipboard` | Push text to all connected browsers (`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | Timeline + stats |
---
@@ -480,9 +526,9 @@ The codebase went through a comprehensive 7-phase refactoring that eliminated go
| Phase | What changed | Impact |
|-------|-------------|--------|
| **Performance** | Cached endpoints, SSE adaptive batching, buffer chunking | Sub-16ms terminal latency |
| **Route extraction** | `server.ts` split into 12 domain route modules + auth middleware + port interfaces | **−60%** server.ts LOC (6,736 → 2,697) |
| **Route extraction** | `server.ts` split into 13 domain route modules + auth middleware + port interfaces | **−60%** server.ts LOC (6,736 → 2,697) |
| **Domain splitting** | `types.ts` → 14 domain files, `ralph-tracker` → 7 files, `respawn-controller` → 5 files, `session` → 6 files | No more god files |
| **Frontend modules** | `app.js` → 8 extracted modules (constants, mobile, voice, notifications, keyboard, API, subagent windows) | **−24%** app.js LOC (15.2K → 11.5K) |
| **Frontend modules** | `app.js` → 9 extracted modules (constants, mobile, voice, notifications, keyboard, CJK input, API, Ralph wizard, subagent windows) | **−24%** app.js LOC (15.2K → 11.5K) |
| **Config consolidation** | ~70 scattered magic numbers → 9 domain-focused config files | Zero cross-file duplicates |
| **Test infrastructure** | Shared mock library, 12 route test files, consolidated MockSession | Testable route handlers via `app.inject()` |
+1 -2
View File
@@ -22,8 +22,7 @@ export default tseslint.config(
'src/web/public/vendor/**',
'src/web/public/app.js',
'scripts/**/*.mjs',
'tools/**',
'remotion/**',
'scripts/remotion/**',
],
}
);
@@ -1,7 +1,11 @@
import { resolve } from 'node:path';
import { defineConfig } from 'vitest/config';
const root = resolve(import.meta.dirname, '..');
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: ['test/**/*.test.ts'],
+423
View File
@@ -0,0 +1,423 @@
# Performance Analysis & Optimization Opportunities
**Date**: 2026-03-07
**Scope**: Full-stack performance analysis — backend PTY handling, SSE broadcasting, frontend terminal rendering, local echo overlay, DOM updates, config/scaling limits.
**Constraint**: All recommendations preserve existing functionality including local echo, backpressure, anti-flicker pipeline, and mobile support.
---
## Executive Summary
The codebase is already well-optimized in critical paths. The multi-layer backpressure system, adaptive terminal batching, DEC 2026 sync markers, and incremental state serialization are strong. The main opportunities are in **reducing unnecessary work** (SSE filtering, DOM rebuilds, lazy terminal init) rather than algorithmic changes.
**Top 5 high-impact opportunities:**
| # | Optimization | Impact | Risk | Effort |
|---|-------------|--------|------|--------|
| 1 | Session-scoped SSE subscriptions | Bandwidth -60-80%, CPU -40% | Medium | Medium |
| 2 | Lazy xterm.js for minimized subagent windows | Memory -3.5MB at 50 agents | Low | Low |
| 3 | Targeted badge update (skip full tab rebuild) | Eliminates O(n) reflow on badge change | Low | Low |
| 4 | Conditional SSE padding (tunnel-only, terminal-only) | Bandwidth -70% when tunneled | Low | Low |
| 5 | Canvas renderer on mobile | GPU pressure reduction, battery savings | Low | Low |
---
## 1. SSE Broadcasting
### Current State
- **92 event types** broadcast to all connected clients (max 100)
- Single `JSON.stringify()` per event, shared across all clients (efficient)
- **No per-client filtering** — every client receives every event regardless of which session they're viewing
- 8KB padding appended to **every** event when tunnel is active (forces Cloudflare proxy flush)
- Backpressure: clients marked as backpressured if `reply.raw.write()` returns false; recovery via `session:needsRefresh`
### Bottlenecks
**B1: No session-scoped SSE subscriptions** (`server.ts:1986`)
- Client viewing session A still receives all events for sessions B through T
- With 20 active sessions, ~95% of terminal events are irrelevant to any given client
- Cost: wasted bandwidth, CPU for JSON parsing, and event handler dispatch on client
**B2: Unconditional 8KB padding** (`server.ts:1977`)
- Every event gets 8KB comment padding when tunnel is active
- A `task:updated` event (~200 bytes payload) becomes ~8.2KB
- High-frequency events like `session:terminal` need the padding; low-frequency events like `session:created` don't
### Recommendations
**R1: Session-scoped SSE subscriptions** (High impact)
- Add `?sessions=id1,id2` query param to `/api/events` SSE endpoint
- Server filters events by session ID before broadcasting
- Client subscribes to active session + "global" events (session lifecycle, system)
- Re-subscribes on tab switch (or subscribe to all with client-side filter as fallback)
- **Savings**: ~80% bandwidth reduction for single-session viewers; ~60% for multi-session dashboards
**R2: Tiered SSE padding** (Medium impact)
- Only pad `session:terminal` events and SSE heartbeats (the two that need proxy flush)
- Skip padding for low-frequency structural events (`session:created`, `task:updated`, etc.)
- **Savings**: ~70% padding overhead reduction; terminal events already large enough to flush
---
## 2. Terminal Rendering
### Current State (Well-Optimized)
- **6-layer anti-flicker pipeline**: Server batching (adaptive 16-50ms) → DEC 2026 sync wrap → single JSON serialize → client rAF batching → sync segment parser → chunked buffer loading (32KB/frame)
- **64KB/frame write budget** with DEC 2026 sync-segment awareness (prevents 141KB single-frame freezes)
- **3-layer backpressure**: SSE cap (128KB queued → drop + refresh), frame budget (64KB/frame), chunked restore (32KB/frame)
- WebGL renderer enabled by default with canvas fallback on context loss
- Typical latency: 16-32ms; worst case: ~115ms (50ms server batch + 50ms sync wait + 16ms rAF)
### Bottlenecks
**B3: WebGL on mobile** (`app.js:627-637`)
- Mobile GPUs are weaker; WebGL context loss more likely on low-end devices
- Canvas renderer is sufficient for mobile (typically 1 session, smaller viewport)
**B4: Static scrollback for all sessions** (`app.js:572`)
- Default 5000 lines scrollback for all sessions regardless of activity level
- Heavy output sessions (build logs, test runners) accumulate large scroll buffers
**B5: No addon lazy loading**
- FitAddon, Unicode11Addon, and WebGLAddon all loaded at terminal init
- Unicode11Addon only needed for CJK content; WebGLAddon is large
### Recommendations
**R3: Force canvas renderer on mobile** (Low risk)
- Detect `MobileDetection.isMobile()` and skip WebGL addon loading
- Reduces GPU memory pressure, prevents context loss crashes
- Mobile typically has 1-2 sessions — canvas performance is more than adequate
**R4: Dynamic scrollback based on session activity** (Low risk)
- Active sessions (working state): 5000 lines (current default)
- Inactive/idle sessions: reduce to 2000 lines
- Restore on session select (fetch from server buffer)
- **Savings**: ~60% scrollback memory for idle sessions
**R5: Lazy-load Unicode11Addon** (Low risk)
- Only load when CJK content is detected in terminal output
- Detection: check for characters in CJK Unicode ranges during ANSI stripping (already iterating)
- Most sessions never need it
---
## 3. DOM & Session Tab Rendering
### Current State
- Session tabs use **intelligent incremental updates** with debounced 100ms rendering
- Incremental path: only updates changed properties (classes, textContent, badges) when session list is stable
- Full rebuild path: triggered when sessions added/removed **or badge count changes**
- Subagent windows: per-window xterm.js instances, even when minimized
### Bottlenecks
**B6: Badge count change triggers full tab rebuild** (`app.js:3207-3209`)
- A single subagent badge increment on one tab triggers `_fullRenderSessionTabs()` — rebuilds entire sidebar HTML via `innerHTML =`
- With 20 sessions, this is an O(n) reflow for a single badge number change
- Badge changes are frequent during active subagent work
**B7: Minimized subagent windows retain xterm.js instances** (`subagent-windows.js`)
- 50 subagent windows × ~75KB per xterm.js instance = ~3.75MB DOM memory
- Minimized windows are invisible but their terminals remain in DOM
- xterm.js instances continue processing resize events even when hidden
**B8: `backdrop-filter: blur()` on overlays** (`styles.css:2246-2247, 3098`)
- Forces new stacking context, disables browser compositing optimizations
- 50-100ms layout thrashing on modal open/close
- Only 2 uses, but they're on frequently toggled overlays
### Recommendations
**R6: Targeted badge update without full rebuild** (Low risk)
- When badge count changes but session list is stable, update only the badge `<span>` textContent
- Keep incremental path for badge changes; only use full rebuild for structural changes (add/remove sessions)
- **Savings**: Eliminates O(n) reflow per badge change; reduces to O(1) targeted update
**R7: Lazy xterm.js initialization for subagent windows** (Medium impact)
- Only create xterm.js Terminal instance when window is restored/maximized
- On minimize: serialize terminal buffer, dispose Terminal instance, keep buffer in memory
- On restore: create new Terminal, write buffer back
- **Savings**: ~3.5MB DOM reduction at 50 minimized agents; eliminates hidden resize processing
- **Trade-off**: ~200-500ms restore delay (buffer write), mitigated by chunked loading
**R8: Replace `backdrop-filter: blur()` with `background: rgba()`** (Low risk)
- Use semi-transparent background instead of blur effect
- Or use `will-change: transform` hint if blur is kept
- **Savings**: Eliminates forced recomposition layer; 50-100ms faster overlay open
---
## 4. Backend PTY & State Management
### Current State (Excellent)
- **BufferAccumulator**: Array-based chunking with lazy join on read — avoids O(n) string concatenation
- **ANSI stripping**: Throttled at 150ms intervals with lazy evaluation (not per-chunk)
- **State persistence**: 500ms debounce + incremental JSON caching per session (only dirty sessions re-serialized)
- **Expensive parsers**: Throttled to 150ms window, accumulated data capped at 64KB
- **Memory**: All buffers have hard limits (2MB terminal, 1MB text, 1000 messages, 64KB line buffer)
### Bottlenecks
**B9: Pending clean data cap at 64KB** (`session.ts:1097-1133`)
- Between 150ms processing windows, raw PTY data accumulates in `_pendingCleanData`
- Capped at 64KB — excess data rolls off (old data discarded)
- During heavy output (large build logs), this means parsers may miss content
- Acceptable trade-off for performance, but worth documenting
**B10: `LRUMap.delete()` is O(n) worst case** (`utils/lru-map.ts:137-138`)
- When deleting the newest entry, iterates all keys to find new newest
- Rare in practice (delete is uncommon; set/get are hot paths)
- Could matter during mass cleanup of 500 agents
### Recommendations
**R9: Consider adaptive pending data cap** (Low priority)
- During idle detection (critical to get right), increase cap to 128KB
- During active working state, keep at 64KB (parsers less critical)
- **Benefit**: More accurate idle detection during heavy output
**R10: Track second-newest in LRUMap** (Low priority)
- Maintain a `_secondNewestKey` alongside `_newestKey`
- On delete of newest, promote second-newest without iteration
- Only matters at scale (500+ agents with frequent eviction)
---
## 5. Local Echo & Input Path
### Current State (Well-Designed)
- **DOM overlay approach** — `<span>` elements in `.xterm-screen` at z-index 7, completely independent of `terminal.write()`
- **Render caching**: `_lastRenderKey` includes text, position, column offsets — skips redundant re-renders
- **Input flow**: Char accumulation → Enter triggers flush → 80ms delay before `\r` (ensures text reaches PTY first)
- **Tab completion**: Baseline snapshot → detect buffer change → 300ms fallback timer
- **CJK support**: Per-character width detection with `terminal.unicode.getStringCellWidth()` preferred, manual fallback
- **Prompt detection**: Bottom-up line scan, O(rows) — cached position, column-lock prevents jitter
### Bottlenecks
**B11: tmux send-keys latency** (~50-100ms per input)
- Each `writeViaMux()` spawns a child process (`tmux send-keys`)
- Text and Enter sent separately with 50ms delay between
- For rapid typing: characters batch before Enter, so overhead is per-command not per-keystroke
- **Acceptable trade-off** for session persistence (tmux survives server restarts)
**B12: 80ms delay between text flush and Enter** (`app.js:872-875`)
- Intentional: ensures text reaches PTY before Enter, preventing Ink from processing empty input
- Adds 80ms to perceived Enter-to-response latency
- Could potentially be reduced with acknowledgment-based approach
**B13: Scroll listener on terminal viewport** (`zerolag-input-addon.ts:139`)
- 50ms debounced re-render on scroll — acceptable but fires frequently during heavy output
- Overlay hidden when scrolled up (correct behavior), shown when at bottom
### Recommendations
**R11: Reduce Enter delay from 80ms to 50ms** (Low risk, test carefully)
- The tmux `send-keys` already has 50ms internal delay
- Combined with network latency, 80ms client-side may be excessive
- Test with Ink-heavy sessions (Claude Code's status bar) — if text arrives before Enter at 50ms, reduce
- **Savings**: 30ms perceived latency reduction per command
**R12: Batch tmux send-keys via stdin pipe** (Medium effort, high impact for rapid input)
- Instead of spawning `tmux send-keys` per input, maintain a persistent connection
- Use `tmux -C` (control mode) for programmatic interaction without child process spawning
- **Savings**: Eliminate ~50-100ms process spawn overhead per input
- **Risk**: Control mode has different semantics; needs careful testing with session persistence
**R13: Skip overlay re-render during heavy output scroll** (Low risk)
- When terminal is receiving >10KB/s output, hide overlay entirely (user isn't typing during heavy output)
- Re-show overlay after 500ms of output silence
- **Savings**: Eliminates unnecessary DOM overlay re-renders during build logs / test output
---
## 6. Polling & File Watchers
### Current State
- **SubagentWatcher**: 1s base poll, full scan throttled to every 5s, fs.watch() on known directories
- **TranscriptWatcher**: 1 per session, fs.watch() primary with 1s poll fallback
- **ImageWatcher**: chokidar per session with 100ms stability poll, burst limit 20/10s
- **TeamWatcher**: chokidar primary with 30s poll fallback, LRU caches (50 teams, 200 tasks)
- **RalphTracker**: Todo cleanup every 5 minutes
### Scaling Profile (20 sessions)
| Component | Instances | Frequency | Total ops/sec |
|-----------|-----------|-----------|---------------|
| SubagentWatcher | 1 (global) | Full scan every 5s | 0.2/s |
| TranscriptWatcher | 20 | 1s poll (fallback) | 20/s max |
| ImageWatcher | 20 | 100ms poll (during writes only) | 200/s burst |
| TeamWatcher | 1 (global) | 30s poll (fallback) | 0.03/s |
| SSE heartbeat | 1 (global) | 15s | 0.07/s |
| SSE dead client check | 1 (global) | 30s | 0.03/s |
| Mux stats collection | 1 (global) | 2s | 0.5/s |
| **Total steady-state** | | | **~21/s** |
### Recommendations
**R14: Increase TranscriptWatcher poll interval to 2s** (Low risk)
- Transcript changes are infrequent (new messages every few seconds at most)
- fs.watch() is the primary mechanism; polling is fallback
- **Savings**: Halves fallback filesystem checks (20/s → 10/s for 20 sessions)
**R15: Share chokidar instances for co-located session directories** (Medium effort)
- Sessions in the same parent directory could share a single chokidar watcher with depth:3
- Common case: multiple sessions in `~/projects/foo/` — one watcher covers all
- **Savings**: Reduce chokidar instances from 20 to ~5-10 for typical workloads
---
## 7. Frontend Asset Delivery
### Current State
- **app.js**: 12,027 lines (source) → esbuild minified → gzip/brotli compressed (~30-40KB gzipped)
- **Static caching**: `maxAge: '1y'` via `@fastify/static`
- **Service worker**: Push notification handler only — no asset caching
- **No code splitting**: Single monolithic app.js bundle
### Bottlenecks
**B14: No cache-busting mechanism**
- `maxAge: '1y'` means browsers cache aggressively
- After deployment, users need `Ctrl+Shift+R` to see updates
- No content hash in filenames or ETags for automatic invalidation
**B15: Monolithic app.js**
- All 12K lines loaded on initial page load regardless of which features are used
- Ralph wizard, plan orchestrator UI, team management — all loaded upfront
- Mobile loads the same bundle as desktop
### Recommendations
**R16: Add content hash to asset filenames** (Medium impact)
- Build step: rename `app.js` → `app.[hash].js`
- Generate a manifest or inject hash into HTML template
- Keep `maxAge: '1y'` — cache invalidation happens via filename change
- **Savings**: Eliminates stale cache issues after deployment; removes need for manual hard refresh
**R17: Code-split app.js into core + feature modules** (High effort, medium impact)
- Core (~4K lines): terminal, SSE, session management, tabs, input handling
- Deferred (~8K lines): Ralph wizard, plan UI, team management, subagent windows, image viewer
- Load deferred modules on first use via dynamic `import()` or lazy `<script>` injection
- **Savings**: ~60% reduction in initial load size; faster time-to-interactive
- **Risk**: Complexity increase; need to handle loading states for deferred features
- **Note**: May not be worth the effort given the app is already gzipped to ~30-40KB
---
## 8. CSS Performance
### Current State
- **styles.css**: 7,153 lines with ~45 box-shadow uses, 2 backdrop-filter uses
- Animations: GPU-accelerated keyframes for pulsing alerts, loading spinners
- Z-index layering: well-organized (subagent 1000, plan 1100, log 2000, image 3000, overlay 7)
### Recommendations
**R18: Replace backdrop-filter with opaque overlay** (Low risk, covered in R8)
**R19: Use `contain: content` on subagent windows** (Low risk)
- Add CSS containment to subagent window containers
- Prevents layout changes inside windows from triggering reflow on parent
- Especially valuable with 50 windows: changes in one window won't invalidate others
- ```css
.subagent-window { contain: content; }
```
- **Savings**: Reduces layout recalculation scope from global to per-window
**R20: Use `content-visibility: auto` on off-screen subagent windows** (Low risk)
- Browser skips rendering of off-screen windows entirely
- Combined with `contain-intrinsic-size` to prevent layout shift
- ```css
.subagent-window.minimized { content-visibility: hidden; }
```
- **Savings**: Browser skips paint/layout for minimized windows; complements R7
---
## 9. Memory & Scaling Limits
### Current Budget (20 sessions)
| Component | Per Session | Total | Status |
|-----------|-----------|-------|--------|
| Terminal buffer | 2MB | 40MB | Hard-limited, auto-trim |
| Text output | 1MB | 20MB | Hard-limited, auto-trim |
| Messages | ~1MB | 20MB | Capped at 1000, trims to 800 |
| Respawn buffer | 1MB | 20MB | Hard-limited |
| **Buffers total** | | **100MB** | Acceptable |
| TranscriptWatcher | ~100KB | 2MB | |
| ImageWatcher | ~50KB | 1MB | |
| SubagentWatcher | ~500KB | 500KB | Global |
| Frontend terminal cache | ~256KB | 5MB | LRU, max 20 entries |
| **Total estimated** | | **~110MB** | Comfortable |
### At Max Scale (50 sessions)
- Buffers: ~250MB
- Watchers: ~5MB
- **Total: ~255MB** + Node.js overhead — acceptable on modern hardware
### Potential Leak Vectors (All Mitigated)
- `_shortIdCache` in server — unbounded Map, but entries are tiny (string→string); grows at O(sessions created), not O(events)
- All CleanupManager-registered resources tracked and disposed on session stop
- `isStopped` guard prevents new timers after session cleanup
---
## 10. Implementation Priority Matrix
### Phase 1 — Quick Wins (1-2 hours each, low risk)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R6 | Targeted badge update | `app.js` (3207-3209) |
| R3 | Canvas renderer on mobile | `app.js` (627-637) |
| R8 | Replace backdrop-filter blur | `styles.css` (2246, 3098) |
| R19 | CSS containment on subagent windows | `styles.css` |
| R20 | `content-visibility: hidden` on minimized windows | `styles.css` |
### Phase 2 — Medium Effort (half-day each)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R2 | Tiered SSE padding | `server.ts` (broadcast function) |
| R7 | Lazy xterm.js for minimized subagents | `subagent-windows.js` |
| R11 | Reduce Enter delay to 50ms | `app.js` (872-875), test with Ink |
| R14 | TranscriptWatcher 2s poll | `transcript-watcher.ts` |
| R16 | Content-hash asset filenames | `build.mjs`, `server.ts` |
### Phase 3 — Larger Initiatives (1-2 days each)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R1 | Session-scoped SSE subscriptions | `server.ts`, `app.js` (SSE connect) |
| R5 | Lazy Unicode11Addon loading | `app.js`, build pipeline |
| R12 | Persistent tmux control mode | `tmux-manager.ts` |
| R17 | Code-split app.js | `app.js`, `build.mjs`, HTML template |
### Not Recommended (Low ROI or High Risk)
| # | Why Not |
|---|---------|
| R4 | Dynamic scrollback adds complexity; memory savings marginal vs total budget |
| R9 | Adaptive pending data cap adds state; current 64KB cap rarely matters |
| R10 | LRUMap.delete() O(n) is theoretical; never triggered at current scale |
| R15 | Shared chokidar instances add directory-matching complexity for minimal gain |
---
## Appendix: Key File Locations
| Area | File | Key Lines |
|------|------|-----------|
| SSE broadcast | `src/web/server.ts` | 1961-1989 (broadcast), 1934-1959 (backpressure) |
| Terminal batching | `src/web/server.ts` | 1994-2048 (per-session adaptive batching) |
| Frame budget | `src/web/public/app.js` | 1370-1478 (flushPendingWrites, 64KB cap) |
| Flicker filter | `src/web/public/app.js` | 1176-1255 (50ms sync wait, 256KB safety) |
| Tab rendering | `src/web/public/app.js` | 3108-3357 (incremental + full rebuild) |
| Tab switching | `src/web/public/app.js` | 3560-3760 (cache + chunked load + deferred UI) |
| Local echo | `packages/xterm-zerolag-input/src/` | All files (overlay, prompt, CJK) |
| Local echo integration | `src/web/public/app.js` | 640, 815-988 (input flow) |
| Subagent windows | `src/web/public/subagent-windows.js` | Full file (window mgmt, drag, minimize) |
| State persistence | `src/state-store.ts` | 161-250 (debounced save, incremental JSON) |
| Buffer accumulator | `src/utils/buffer-accumulator.ts` | Full file (array chunks, lazy join) |
| PTY handling | `src/session.ts` | 1046-1133 (data flow), 1173-1230 (parsing) |
| Config limits | `src/config/` | 9 files (buffer, map, timing, auth, etc.) |
| Anti-flicker docs | `docs/terminal-anti-flicker.md` | Architecture reference |
| CSS | `src/web/public/styles.css` | 2246 (backdrop-filter), full file |
| Build pipeline | `scripts/build.mjs` | 59-68 (minify + compress) |
+74
View File
@@ -0,0 +1,74 @@
# Codeman Performance Optimization Plan
## Current State
The backend is **already production-grade** — SSE broadcasting, state persistence, terminal batching, buffer management, and memory patterns are all well-optimized. The biggest gains are on the **frontend delivery** side.
## Implemented Optimizations
### 1. V8 Compile Cache (10-20% faster cold start)
**Files:** `scripts/codeman-web.service`, `package.json`
Node.js re-parses and compiles all JS on every cold start. `NODE_COMPILE_CACHE` caches V8 compiled bytecode to disk, reusing it on subsequent starts.
- Added `Environment=NODE_COMPILE_CACHE=/home/arkon/.codeman/compile-cache` to systemd service
- Added to `npm start` script for non-systemd usage
- Zero code changes, immediate win on every restart
### 2. WebGL Addon Lazy-Loading (244KB saved on mobile, non-blocking on desktop)
**Files:** `src/web/public/index.html`, `src/web/public/app.js`
`xterm-addon-webgl.min.js` (244KB) was loaded eagerly for all users via `<script defer>`, but only used on desktop with WebGL2 support.
- Removed `<script defer>` from `index.html`
- Added dynamic script loading in `app.js` — only downloads on desktop when WebGL is needed
- Mobile users never download the file at all (244KB saved)
- Desktop: loads in parallel with page rendering, addon initializes when ready
- Graceful fallback: canvas renderer used if WebGL unavailable or script fails
### 3. Preload Hints (~50-100ms faster perceived load)
**Files:** `src/web/public/index.html`
Browser discovers `<script defer>` tags only when the parser reaches them at the bottom of `<body>`. By then, the HTML parse has blocked for hundreds of lines.
- Added `<link rel="preload" as="script">` in `<head>` for `vendor/xterm.min.js`, `constants.js`, `app.js`
- Browser starts fetching critical scripts immediately during HTML parse (before reaching `<body>`)
- Zero runtime overhead — just hints for the browser's preload scanner
### 4. Batch Tmux Reconciliation (N subprocess calls → 1)
**Files:** `src/tmux-manager.ts`
`reconcileSessions()` previously called `tmux has-session` + `tmux display-message` per known session, plus `tmux list-sessions` for discovery, plus `tmux display-message` per discovered session. With 20 sessions: 41+ subprocess calls.
- Replaced with single `tmux list-panes -a -F '#{session_name}\t#{pane_pid}'` call
- Builds a Map from the result, then does O(1) lookups for both known and discovered sessions
- Also replaced inner O(n) `isKnown` scan with a Set lookup
- 20 sessions: 41 subprocess calls → 1, with faster lookups
### 5. Asset Hashing / Cache Busting (already implemented)
**Files:** `scripts/build.mjs` (pre-existing)
Content-hash cache busting was already implemented in the build script:
- All app JS/CSS files get content hashes (`app.abc123.js`)
- `index.html` rewritten to reference hashed filenames
- Pre-compressed with gzip + Brotli
- 1-year immutable cache works correctly — new deploys get new filenames
## Already Optimized (No Action Needed)
| Area | Why It's Fine |
|------|---------------|
| **SSE Broadcasting** | Single serialization per broadcast, preformatted frames, backpressure handling, session subscription filtering |
| **State Persistence** | 500ms debounce, incremental per-session JSON caching, async atomic writes, circuit breaker on failures |
| **Terminal Batching** | Adaptive intervals (16-50ms), per-session queues, immediate flush at 32KB, array-based accumulation |
| **Buffer Management** | BufferAccumulator (array-push, lazy join), auto-trim at 2MB/1MB, no string concatenation in hot paths |
| **ANSI Stripping** | Pre-compiled regex via factory functions, single-pass processing |
| **Static File Serving** | @fastify/static with 1-year cache, pre-compressed Brotli/gzip, no-cache for HTML |
| **Memory Management** | CleanupManager, LRUMap, StaleExpirationMap, bounded buffers, explicit listener cleanup |
| **Import Patterns** | Pure ESM, lazy web server import, no circular deps, no dynamic imports in hot paths |
| **Config Loading** | Small constant files, no I/O at import time, specific imports (no barrel) |
Binary file not shown.
Binary file not shown.

After

Width:  |  Height:  |  Size: 806 KiB

+367
View File
@@ -0,0 +1,367 @@
# Orchestrator Loop — Architecture & Data Flow
> Technical architecture document. Not for GitHub.
## System Overview
```
┌─────────────────────────────────────────────────────────────────────┐
│ CODEMAN WEB UI │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ Orchestrator Dashboard │ │
│ │ [Goal Input] [Plan View] [Phase Progress] [Agent Activity] │ │
│ └───────────────────────────┬──────────────────────────────────┘ │
│ │ SSE Events │
│ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ Orchestrator API Routes (/api/orchestrator/*) │ │
│ └───────────────────────────┬──────────────────────────────────┘ │
└───────────────────────────────┼─────────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────────────┐
│ ORCHESTRATOR LOOP │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────────────┐ │
│ │ Orchestrator │ │ Orchestrator │ │ Orchestrator │ │
│ │ Planner │ │ Loop (state │ │ Verifier │ │
│ │ │ │ machine) │ │ │ │
│ │ • Research │◄──►│ • Phase mgmt │◄──►│ • Test runner │ │
│ │ • Plan gen │ │ • Task queue │ │ • AI review │ │
│ │ • Phasing │ │ • Event loop │ │ • Output checks │ │
│ └──────┬───────┘ └──────┬───────┘ └──────────┬───────────┘ │
│ │ │ │ │
│ ▼ ▼ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ EXISTING CODEMAN INFRASTRUCTURE │ │
│ │ │ │
│ │ SessionManager ←→ Sessions ←→ PTY (Claude CLI) │ │
│ │ ↑ ↑ ↑ │ │
│ │ │ │ │ │ │
│ │ TaskQueue RalphTracker RespawnController │ │
│ │ StateStore HooksConfig TeamWatcher │ │
│ │ Auto-Ops SubagentWatcher SSE Broadcast │ │
│ └──────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
```
## Data Flow: Complete Lifecycle
### 1. User Submits Goal
```
User → POST /api/orchestrator/start { goal: "Build a REST API...", config: {...} }
→ OrchestratorLoop.start(goal)
→ state = PLANNING
→ emit('stateChanged', 'planning')
→ SSE: orchestrator:stateChanged
```
### 2. Planning Phase
```
OrchestratorPlanner.generatePlan(goal)
→ PlanOrchestrator.generateDetailedPlan(goal)
→ [Research Agent] → enriched task description
→ [Planner Agent] → PlanItem[]
→ groupIntoPhases(planItems)
→ topological sort by dependencies
→ group into layers
→ assign team strategies
→ OrchestratorPlan { phases: [...] }
→ state = APPROVAL
→ emit('planReady', plan)
→ SSE: orchestrator:planReady
```
### 3. User Approves Plan
```
User → POST /api/orchestrator/approve
→ OrchestratorLoop.approvePlan()
→ state = EXECUTING
→ executePhase(phases[0])
```
### 4. Phase Execution
```
executePhase(phase)
→ For each task in phase:
→ Convert to CreateTaskOptions
→ Add to TaskQueue with completion phrase "PHASE_{N}_TASK_{M}_DONE"
→ If phase.teamStrategy.type === 'team':
→ Start session with AGENT_TEAMS enabled
→ Send team orchestration prompt to lead
→ Else:
→ Assign tasks to available sessions (same as RalphLoop)
→ Listen for task completion events:
→ TaskQueue emits taskCompleted
→ Check: all phase tasks done?
→ Yes → state = VERIFYING → verifyPhase(phase)
→ No → wait for more completions
```
### 5. Verification
```
verifyPhase(phase)
→ OrchestratorVerifier.verify(phase, session)
→ Run test commands via session
→ Check file existence
→ AI review (optional)
→ If passed:
→ phase.status = 'passed'
→ emit('phaseCompleted', phase)
→ If more phases: executePhase(nextPhase)
→ If last phase: state = COMPLETED
→ If failed:
→ phase.attempts++
→ If attempts < maxAttempts:
→ state = REPLANNING
→ Generate recovery tasks
→ state = EXECUTING (retry)
→ Else:
→ state = FAILED
→ emit('phaseFailed', phase, reason)
```
### 6. Context Management Between Phases
```
After phase completion:
→ If config.compactBetweenPhases:
→ session.sendInput('/compact')
→ Wait for compact to complete
→ If config.respawnBetweenMilestones && phase is a milestone:
→ Save orchestrator state to StateStore
→ Respawn session (kill + recreate)
→ Send resume prompt with phase context
```
## File Layout
```
src/
├── orchestrator-loop.ts # Main state machine (~400 lines)
├── orchestrator-planner.ts # Plan generation + phase grouping (~300 lines)
├── orchestrator-verifier.ts # Phase verification (~200 lines)
├── types/
│ └── orchestrator.ts # All orchestrator types (~150 lines)
├── prompts/
│ └── orchestrator.ts # Prompt templates (~200 lines)
├── web/
│ ├── routes/
│ │ └── orchestrator-routes.ts # API endpoints (~250 lines)
│ └── public/
│ └── orchestrator-ui.js # Frontend panel (~500 lines)
```
## Integration Points with Existing Code
### StateStore (`src/state-store.ts`)
```typescript
// Add to AppState interface
orchestrator?: OrchestratorPersistState;
// Add methods
getOrchestratorState(): OrchestratorPersistState;
setOrchestratorState(state: Partial<OrchestratorPersistState>): void;
```
### SSE Events (`src/web/sse-events.ts`)
```typescript
// Add ~8 new events
export const SseEvent = {
// ... existing
ORCHESTRATOR_STATE_CHANGED: 'orchestrator:stateChanged',
ORCHESTRATOR_PLAN_READY: 'orchestrator:planReady',
ORCHESTRATOR_PHASE_STARTED: 'orchestrator:phaseStarted',
ORCHESTRATOR_PHASE_COMPLETED: 'orchestrator:phaseCompleted',
ORCHESTRATOR_PHASE_FAILED: 'orchestrator:phaseFailed',
ORCHESTRATOR_VERIFICATION: 'orchestrator:verificationResult',
ORCHESTRATOR_COMPLETED: 'orchestrator:completed',
ORCHESTRATOR_ERROR: 'orchestrator:error',
} as const;
```
### Frontend Constants (`src/web/public/constants.js`)
```javascript
// Mirror SSE events
SSE_EVENTS.ORCHESTRATOR_STATE_CHANGED = 'orchestrator:stateChanged';
// ... etc
```
### Route Registration (`src/web/routes/index.ts`)
```typescript
import { registerOrchestratorRoutes } from './orchestrator-routes.js';
// Add to barrel export
```
### Server (`src/web/server.ts`)
```typescript
// Initialize OrchestratorLoop alongside RalphLoop
const orchestratorLoop = new OrchestratorLoop(config);
// Register routes
registerOrchestratorRoutes(app, { ...ctx, orchestrator: orchestratorLoop });
```
### Port Interface (`src/web/ports/`)
```typescript
// New port
export interface OrchestratorPort {
orchestrator: OrchestratorLoop;
}
```
## Prompt Flow Through System
The key insight is how prompts flow from Orchestrator → Session → Claude:
```
OrchestratorLoop decides to execute Phase 3, Task 2
│
▼
Converts OrchestratorTask to CreateTaskOptions:
{
prompt: "Implement the rate limiter middleware. Read src/middleware/auth.ts
for the pattern. Add to src/middleware/rate-limiter.ts. Must export
a Fastify plugin. When done: <promise>PHASE_3_TASK_2_DONE</promise>",
priority: 100,
dependencies: ["phase-3-task-1"], // Must finish auth middleware first
completionPhrase: "PHASE_3_TASK_2_DONE",
timeoutMs: 600000 // 10 minutes
}
│
▼
TaskQueue.addTask(options)
│
▼
RalphLoop.tick() → assignTasks() // OR OrchestratorLoop does its own assignment
│
▼
session.sendInput(task.prompt)
│
▼
writeViaMux() → tmux send-keys -l "prompt..." + Enter
│
▼
Claude CLI receives prompt, executes, outputs results
│
▼
RalphTracker.processData() → detects "PHASE_3_TASK_2_DONE"
│
▼
emit('completionDetected') → OrchestratorLoop.handleTaskCompleted()
│
▼
Check: all tasks in Phase 3 done? → If yes → verifyPhase(phase3)
```
## Team Agent Flow (When Enabled)
```
Phase has teamStrategy.type === 'team'
│
▼
OrchestratorLoop creates/reuses a session with:
env: { CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS: '1' }
│
▼
Sends team orchestration prompt:
"You're the team lead for Phase 3: Core Implementation.
Your team should work on these tasks in parallel:
1. Rate limiter middleware (teammate 1)
2. Error handling middleware (teammate 2)
3. Validation layer (teammate 3)
Context files to read first: [...]
Each teammate should output their task's completion phrase when done.
When ALL tasks are complete, output: <promise>PHASE_3_COMPLETE</promise>"
│
▼
Claude Code team-lead spawns teammates
│
▼
TeamWatcher detects new team in ~/.claude/teams/
→ Matches to session via leadSessionId
→ Tracks teammate activity
│
▼
Teammates work in parallel (in-process threads)
│
▼
hook: teammate_idle → POST /api/hook-event
→ OrchestratorLoop notes teammate finished
│
▼
hook: task_completed → POST /api/hook-event
→ Or: RalphTracker detects PHASE_3_COMPLETE
→ OrchestratorLoop → phase complete → verify
```
## Error Recovery Strategy
```
Task fails (timeout, error, session crash)
│
├─ Task-level retry (up to 2 retries per task)
│ → Reset task to pending
│ → Re-queue with modified prompt: "Previous attempt failed: {error}. Try again..."
│
├─ Phase-level retry (up to 3 retries per phase)
│ → Respawn session (fresh context)
│ → Re-execute entire phase with learnings from failure
│ → Modified prompt includes what went wrong
│
└─ Orchestration-level failure
→ All retries exhausted
→ state = FAILED
→ Notify user with detailed failure report
→ User can: modify plan → retry, skip phase → continue, or stop
```
## Interaction with Ralph Loop
Ralph Loop and Orchestrator Loop are **mutually exclusive** on the same sessions:
```
if (orchestratorLoop.isRunning()) {
// Orchestrator controls task assignment
// Ralph Loop should not interfere
// Respawn Controller uses 'orchestrator' preset
}
if (ralphLoop.isRunning()) {
// Ralph controls task assignment
// Orchestrator should not start
}
```
The Orchestrator can optionally USE the Ralph Loop internally for phase execution (delegate phase tasks to Ralph's queue), or manage task assignment directly. Decision: **manage directly** — gives more control over phase boundaries and verification timing.
## Summary of What Touches What
| Existing File | Change |
|---|---|
| `src/types/index.ts` | Export orchestrator types |
| `src/state-store.ts` | Add orchestrator state persistence |
| `src/web/sse-events.ts` | Add ~8 orchestrator events |
| `src/web/routes/index.ts` | Register orchestrator routes |
| `src/web/server.ts` | Initialize OrchestratorLoop |
| `src/web/public/constants.js` | Mirror SSE events |
| `src/web/public/app.js` | Add orchestrator event listeners, panel toggle |
| `src/web/route-helpers.ts` | Add 'orchestrator' respawn preset |
| New File | Purpose |
|---|---|
| `src/orchestrator-loop.ts` | Core state machine |
| `src/orchestrator-planner.ts` | Plan generation + phasing |
| `src/orchestrator-verifier.ts` | Phase verification |
| `src/types/orchestrator.ts` | Type definitions |
| `src/prompts/orchestrator.ts` | Prompt templates |
| `src/web/routes/orchestrator-routes.ts` | API endpoints |
| `src/web/public/orchestrator-ui.js` | Frontend panel |
| `src/web/ports/orchestrator-port.ts` | Port interface |
+633
View File
@@ -0,0 +1,633 @@
# Orchestrator Loop — Detailed Implementation Plan (v2)
> Internal research/planning document. Not for GitHub.
## Vision
The **Orchestrator Loop** is a new autonomous execution mode that transforms high-level user goals into phased, verified, team-coordinated implementations. Unlike Ralph Loop (flat task queue → idle sessions), the Orchestrator manages the full lifecycle: **plan → approve → execute → verify → adapt → complete**.
```
USER: "Add OAuth2 login with Google/GitHub, role-based access control, and API key management"
ORCHESTRATOR:
Phase 1: Research & Setup ✅ (3m) — scaffold, deps, config
Phase 2: Auth Core ✅ (8m) — OAuth2 flow, session mgmt
Phase 3: Provider Integration 🔄 (12m) — Google + GitHub (parallel via team agents)
Phase 4: RBAC ⏳ — roles, permissions, middleware
Phase 5: API Keys ⏳ — generation, validation, rate limits
Phase 6: Testing & Review ⏳ — integration tests, security review
Progress: ━━━━━━━━━━━━━━━━━━━━ 40% | Agents: 3 active | Time: 23m
```
## Architecture
```
┌─────────────────────────────────────────────────────────────────┐
│ OrchestratorLoop │
│ │
│ ┌────────────────┐ ┌────────────────┐ ┌──────────────────┐ │
│ │ Orchestrator │ │ Orchestrator │ │ Orchestrator │ │
│ │ Planner │ │ Executor │ │ Verifier │ │
│ │ │ │ │ │ │ │
│ │ PlanOrchestrator│ │ TaskQueue │ │ AI review │ │
│ │ + phase grouper│ │ SessionManager │ │ Test commands │ │
│ │ + team strategy│ │ Team prompts │ │ File checks │ │
│ └───────┬────────┘ └───────┬────────┘ └─────────┬────────┘ │
│ │ │ │ │
│ └───────────────────┼──────────────────────┘ │
│ │ │
│ ┌─────────▼─────────┐ │
│ │ Existing Codeman │ │
│ │ Infrastructure │ │
│ │ │ │
│ │ SessionManager │ │
│ │ TaskQueue │ │
│ │ RespawnController │ │
│ │ TeamWatcher │ │
│ │ PlanOrchestrator │ │
│ │ StateStore │ │
│ │ Hooks + SSE │ │
│ └────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
```
## State Machine
```
┌─────────┐
│ IDLE │
└────┬────┘
│ start(goal)
▼
┌─────────┐
┌────────│PLANNING │────────┐
│ fail └────┬────┘ │
▼ │ plan ready │ user cancels
┌────────┐ ▼ ▼
│ FAILED │ ┌─────────┐ ┌────────┐
└────────┘ │APPROVAL │ │ IDLE │
▲ └────┬────┘ └────────┘
│ │ approve
│ ▼
│ ┌──────────┐
│ ┌───►│EXECUTING │◄────────────────────┐
│ │ └────┬─────┘ │
│ │ │ all tasks in phase done │
│ │ ▼ │
│ │ ┌──────────┐ │
│ │ │VERIFYING │ │
│ │ └────┬─────┘ │
│ │ pass │ │ fail │
│ │ ▼ ▼ │
│ │ more ┌──────────┐ │
│ │ phases?│REPLANNING│── retry ────────┘
│ │ │ └────┬─────┘
│ │ │ │ max retries
│ │ │ ▼
│ │ │ ┌────────┐
│ └────┘ │ FAILED │
│ next └────────┘
│ phase
│ │
│ ▼
│ ┌───────────┐
└─│ COMPLETED │
└───────────┘
```
**States:** `idle` | `planning` | `approval` | `executing` | `verifying` | `replanning` | `completed` | `failed` | `paused`
Transitions are event-driven. The state machine is the single source of truth — all methods check `this.state` before acting.
## Type Definitions
### `src/types/orchestrator.ts`
```typescript
// ═══════════════════════════════════════════════════════════════
// State Machine
// ═══════════════════════════════════════════════════════════════
export type OrchestratorState =
| 'idle'
| 'planning'
| 'approval'
| 'executing'
| 'verifying'
| 'replanning'
| 'completed'
| 'failed'
| 'paused';
// ═══════════════════════════════════════════════════════════════
// Plan Structure
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorPlan {
id: string;
goal: string;
createdAt: number;
phases: OrchestratorPhase[];
metadata: {
totalTasks: number;
estimatedComplexity: 'low' | 'medium' | 'high';
modelUsed: string;
planDurationMs: number;
};
}
export interface OrchestratorPhase {
id: string; // "phase-1", "phase-2"
name: string; // Human-readable name
description: string;
order: number;
status: PhaseStatus;
tasks: OrchestratorTask[];
verificationCriteria: string[];
testCommands: string[];
maxAttempts: number; // Default: 3
attempts: number; // Current attempt count
startedAt: number | null;
completedAt: number | null;
durationMs: number | null;
teamStrategy: TeamStrategy;
}
export type PhaseStatus =
| 'pending'
| 'executing'
| 'verifying'
| 'passed'
| 'failed'
| 'skipped';
export interface OrchestratorTask {
id: string; // "phase-1-task-1"
phaseId: string;
prompt: string; // Single-line prompt for Claude
status: 'pending' | 'running' | 'completed' | 'failed';
assignedSessionId: string | null;
queueTaskId: string | null; // Links to TaskQueue task
parallel: boolean; // Can run in parallel with sibling tasks
completionPhrase: string; // Unique phrase for completion detection
timeoutMs: number;
startedAt: number | null;
completedAt: number | null;
error: string | null;
retries: number;
}
// ═══════════════════════════════════════════════════════════════
// Team Strategy
// ═══════════════════════════════════════════════════════════════
export type TeamStrategy =
| { type: 'single' } // One session handles all
| { type: 'parallel'; maxSessions: number } // Multiple sessions
| { type: 'team'; config: TeamSetup } // Agent teams
export interface TeamSetup {
leadPrompt: string;
suggestedTeammates: string[]; // Role descriptions
maxTeammates: number;
}
// ═══════════════════════════════════════════════════════════════
// Verification
// ═══════════════════════════════════════════════════════════════
export interface VerificationResult {
passed: boolean;
checks: VerificationCheck[];
summary: string;
suggestions: string[]; // Recovery hints for replanning
}
export interface VerificationCheck {
type: 'test_command' | 'ai_review' | 'file_check';
description: string;
passed: boolean;
output?: string;
}
// ═══════════════════════════════════════════════════════════════
// Configuration
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorConfig {
plannerModel: string; // Default: 'opus'
researchEnabled: boolean; // Default: true
autoApprove: boolean; // Default: false
maxPhaseRetries: number; // Default: 3
phaseTimeoutMs: number; // Default: 1800000 (30min)
enableTeamAgents: boolean; // Default: true
maxParallelSessions: number; // Default: 3
verificationMode: 'strict' | 'moderate' | 'lenient';
compactBetweenPhases: boolean; // Default: true
}
// ═══════════════════════════════════════════════════════════════
// Persistence (saved to ~/.codeman/state.json)
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorPersistState {
state: OrchestratorState;
plan: OrchestratorPlan | null;
currentPhaseIndex: number;
startedAt: number | null;
completedAt: number | null;
config: OrchestratorConfig;
stats: OrchestratorStats;
}
export interface OrchestratorStats {
phasesCompleted: number;
phasesFailed: number;
totalTasksCompleted: number;
totalTasksFailed: number;
totalDurationMs: number;
replanCount: number;
}
```
## New Files (Implementation Order)
### Step 1: `src/types/orchestrator.ts` — Type definitions
All interfaces above. No dependencies. ~120 lines.
### Step 2: `src/orchestrator-planner.ts` — Plan generation + phase grouping
~300 lines. Wraps existing PlanOrchestrator.
```typescript
/**
* @fileoverview Orchestrator plan generation — converts goals into phased plans.
*
* Uses PlanOrchestrator for AI plan generation, then groups PlanItems into
* sequential phases with team strategies and verification criteria.
*
* @module orchestrator-planner
*/
export class OrchestratorPlanner {
constructor(mux: TerminalMultiplexer, workingDir: string, config: OrchestratorConfig);
/** Generate plan from goal. Uses PlanOrchestrator internally. */
async generatePlan(goal: string, onProgress?: ProgressCallback): Promise<OrchestratorPlan>;
/** Cancel in-progress plan generation. */
async cancel(): Promise<void>;
// Internal
private groupIntoPhases(items: PlanItem[], goal: string): OrchestratorPhase[];
private assignTeamStrategies(phases: OrchestratorPhase[]): void;
private generateCompletionPhrases(plan: OrchestratorPlan): void;
}
```
**Phase grouping algorithm:**
1. Topological sort by `PlanItem.dependencies`
2. Group into dependency layers (Kahn's algorithm)
3. Within each layer, sub-group by `tddPhase` (setup → test → impl → verify → review)
4. Merge adjacent small phases (< 2 tasks) if they share the same tddPhase
5. Assign team strategies:
- 1-2 tasks → `{ type: 'single' }`
- 3+ independent tasks → `{ type: 'parallel', maxSessions: Math.min(taskCount, config.maxParallelSessions) }`
- 4+ tasks with high complexity → `{ type: 'team', config: { ... } }`
6. Generate unique completion phrases per task: `ORCH_P{phaseOrder}_T{taskIndex}`
### Step 3: `src/orchestrator-verifier.ts` — Phase verification
~200 lines.
```typescript
/**
* @fileoverview Orchestrator phase verification.
*
* Runs verification checks after each phase completes:
* test commands, AI review, and file existence checks.
*
* @module orchestrator-verifier
*/
export class OrchestratorVerifier {
constructor(config: OrchestratorConfig);
/** Run all verification checks for a completed phase. */
async verifyPhase(
phase: OrchestratorPhase,
session: Session,
mode: 'strict' | 'moderate' | 'lenient'
): Promise<VerificationResult>;
// Verification strategies
private async runTestCommands(commands: string[], session: Session): Promise<VerificationCheck[]>;
private async aiReview(phase: OrchestratorPhase, session: Session): Promise<VerificationCheck>;
}
```
**Verification modes:**
- `strict`: ALL test commands must pass AND AI review must approve
- `moderate`: Test commands must pass, AI review is advisory
- `lenient`: At least one test command passes, AI review skipped
**AI review prompt (sent as a task to the session):**
```
Review Phase "{phase.name}" completion. Check:
1. Expected functionality works
2. No obvious regressions
3. Code quality is acceptable
Criteria: {phase.verificationCriteria.join('\n')}
If ALL criteria are met, respond: ORCH_VERIFY_PASS
If ANY criteria fail, respond: ORCH_VERIFY_FAIL and explain what failed.
```
### Step 4: `src/orchestrator-loop.ts` — Core state machine
~500 lines. Main orchestrator engine.
```typescript
/**
* @fileoverview Orchestrator Loop — phased plan execution with team agents.
*
* State machine that generates plans from user goals, executes them
* phase-by-phase with verification gates, and adapts on failure.
*
* @module orchestrator-loop
*/
export interface OrchestratorLoopEvents {
stateChanged: (state: OrchestratorState, prevState: OrchestratorState) => void;
planReady: (plan: OrchestratorPlan) => void;
phaseStarted: (phase: OrchestratorPhase) => void;
phaseCompleted: (phase: OrchestratorPhase) => void;
phaseFailed: (phase: OrchestratorPhase, reason: string) => void;
taskAssigned: (task: OrchestratorTask, sessionId: string) => void;
taskCompleted: (task: OrchestratorTask) => void;
taskFailed: (task: OrchestratorTask, error: string) => void;
verificationResult: (phase: OrchestratorPhase, result: VerificationResult) => void;
completed: (stats: OrchestratorStats) => void;
error: (error: Error) => void;
}
export class OrchestratorLoop extends EventEmitter {
private state: OrchestratorState = 'idle';
private plan: OrchestratorPlan | null = null;
private currentPhaseIndex = 0;
private config: OrchestratorConfig;
private planner: OrchestratorPlanner;
private verifier: OrchestratorVerifier;
private sessionManager: SessionManager;
private taskQueue: TaskQueue;
private store: StateStore;
private stats: OrchestratorStats;
private cleanup: CleanupManager;
private pausedState: OrchestratorState | null = null; // State before pause
// ── Lifecycle ──────────────────────────────────────────────
constructor(mux: TerminalMultiplexer, workingDir: string, config?: Partial<OrchestratorConfig>);
/** Start orchestration with a goal. Transitions: idle → planning */
async start(goal: string): Promise<void>;
/** Approve the generated plan. Transitions: approval → executing */
async approve(): Promise<void>;
/** Reject plan with feedback. Transitions: approval → planning (regenerate) */
async reject(feedback: string): Promise<void>;
/** Pause execution. Saves current state. */
pause(): void;
/** Resume from pause. */
resume(): void;
/** Stop everything and clean up. → idle */
async stop(): Promise<void>;
/** Skip current phase. → executing (next phase) or completed */
async skipPhase(phaseId: string): Promise<void>;
/** Retry a failed phase. → executing */
async retryPhase(phaseId: string): Promise<void>;
// ── Getters ────────────────────────────────────────────────
getState(): OrchestratorState;
getPlan(): OrchestratorPlan | null;
getCurrentPhase(): OrchestratorPhase | null;
getStats(): OrchestratorStats;
getStatus(): OrchestratorPersistState;
// ── Internal: Phase Execution ──────────────────────────────
private async executeCurrentPhase(): Promise<void>;
private async executePhase(phase: OrchestratorPhase): Promise<void>;
private async assignPhaseTasks(phase: OrchestratorPhase): Promise<void>;
private handleTaskCompleted(taskId: string): void;
private handleTaskFailed(taskId: string, error: string): void;
private async onPhaseTasksComplete(phase: OrchestratorPhase): Promise<void>;
// ── Internal: Verification ─────────────────────────────────
private async verifyCurrentPhase(): Promise<void>;
private async handleVerificationResult(phase: OrchestratorPhase, result: VerificationResult): Promise<void>;
// ── Internal: Replanning ───────────────────────────────────
private async replanPhase(phase: OrchestratorPhase, failures: string[]): Promise<void>;
// ── Internal: State Machine ────────────────────────────────
private setState(newState: OrchestratorState): void;
private advanceToNextPhase(): Promise<void>;
private persist(): void;
private restore(): void;
}
```
**Key execution flow in `executePhase()`:**
1. Mark phase as `executing`, emit `phaseStarted`
2. For each task in phase:
- Create a `CreateTaskOptions` from `OrchestratorTask`
- Add to `TaskQueue` with proper dependencies + completion phrase
- Store the TaskQueue task ID in `OrchestratorTask.queueTaskId`
3. Poll task completion (listen to TaskQueue events)
4. When all tasks complete → call `onPhaseTasksComplete()`
5. `onPhaseTasksComplete()` triggers verification
**How tasks get assigned to sessions:**
The OrchestratorLoop does NOT manage session assignment directly. It adds tasks to the existing TaskQueue and starts a mini poll loop that assigns pending tasks to idle sessions — the same pattern as RalphLoop's `assignTasks()`. This reuses existing session management.
**Team agent flow:**
For phases with `teamStrategy.type === 'team'`:
- Start a single session with `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
- Instead of adding individual tasks to TaskQueue, send ONE comprehensive prompt to the lead
- The prompt instructs the lead to create teammates and delegate
- Monitor via TeamWatcher for team task completion + hook events
- Phase completion is detected via the lead's completion phrase
### Step 5: `src/web/routes/orchestrator-routes.ts` — API endpoints
~300 lines.
```
POST /api/orchestrator/start — { goal, config? } → start planning
POST /api/orchestrator/approve — approve generated plan
POST /api/orchestrator/reject — { feedback } → reject + replan
POST /api/orchestrator/pause — pause execution
POST /api/orchestrator/resume — resume execution
POST /api/orchestrator/stop — stop orchestration
GET /api/orchestrator/status — full state + plan + stats
GET /api/orchestrator/plan — plan details only
POST /api/orchestrator/phase/:id/skip — skip a phase
POST /api/orchestrator/phase/:id/retry — retry a failed phase
```
Port dependency: `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort`
The route module receives the OrchestratorLoop instance via the InfraPort (added to `createRouteContext()`).
### Step 6: SSE Events — `src/web/sse-events.ts` additions
```typescript
// ─── Orchestrator ────────────────────────────────────────────────────────────
/** Orchestrator state machine transitioned. */
export const OrchestratorStateChanged = 'orchestrator:stateChanged' as const;
/** Orchestrator plan generated and ready for approval. */
export const OrchestratorPlanReady = 'orchestrator:planReady' as const;
/** Orchestrator phase started executing. */
export const OrchestratorPhaseStarted = 'orchestrator:phaseStarted' as const;
/** Orchestrator phase completed successfully. */
export const OrchestratorPhaseCompleted = 'orchestrator:phaseCompleted' as const;
/** Orchestrator phase failed. */
export const OrchestratorPhaseFailed = 'orchestrator:phaseFailed' as const;
/** Orchestrator verification result for a phase. */
export const OrchestratorVerification = 'orchestrator:verification' as const;
/** Orchestrator task assigned to session. */
export const OrchestratorTaskAssigned = 'orchestrator:taskAssigned' as const;
/** Orchestrator task completed. */
export const OrchestratorTaskCompleted = 'orchestrator:taskCompleted' as const;
/** Orchestrator task failed. */
export const OrchestratorTaskFailed = 'orchestrator:taskFailed' as const;
/** All phases completed successfully. */
export const OrchestratorCompleted = 'orchestrator:completed' as const;
/** Orchestrator error. */
export const OrchestratorError = 'orchestrator:error' as const;
```
11 new events. Add to `SseEvent` namespace object + mirror in `constants.js`.
### Step 7: State persistence — `src/state-store.ts` additions
Add to `AppState`:
```typescript
orchestrator?: OrchestratorPersistState;
```
Add methods:
```typescript
getOrchestratorState(): OrchestratorPersistState | null;
setOrchestratorState(state: Partial<OrchestratorPersistState>): void;
clearOrchestratorState(): void;
```
### Step 8: Server integration — `src/web/server.ts` modifications
1. Import `OrchestratorLoop` and `registerOrchestratorRoutes`
2. Add `private orchestratorLoop: OrchestratorLoop` field
3. Initialize in constructor (lazy — created on first start, not at boot)
4. Add to `createRouteContext()` InfraPort: `orchestratorLoop: this.orchestratorLoop`
5. Wire up OrchestratorLoop events → SSE broadcasts
6. Register routes: `registerOrchestratorRoutes(this.app, ctx)`
7. Clean up in `stop()`
### Step 9: `src/web/public/orchestrator-ui.js` — Frontend panel
~500 lines. New frontend module.
**Load order**: After `panels-ui.js` (11), before `ralph-wizard.js` (13). So load order = 11.5.
**UI elements:**
- Goal input form (text area + config toggles)
- Plan approval view (phase list, task details, approve/reject buttons)
- Execution dashboard (progress bar, phase cards, task status indicators)
- Agent activity panel (session count, team status)
- Controls (pause, resume, stop, skip phase, retry phase)
**SSE listeners:**
- All 11 orchestrator events → update UI state
- Reuses existing session/respawn/team event handlers for agent monitoring
### Step 10: `src/prompts/orchestrator.ts` — Prompt templates
~200 lines.
Templates for:
- Phase execution prompt (tells Claude what to do in this phase)
- Team lead delegation prompt (instructs lead to create and coordinate teammates)
- Verification prompt (asks Claude to verify phase output)
- Replan prompt (gives failure context, asks for recovery steps)
### Step 11: Constants, schemas, route barrel updates
- `src/web/public/constants.js` — Add 11 SSE event mirrors
- `src/web/schemas.ts` — Add Zod schemas for orchestrator API input validation
- `src/web/routes/index.ts` — Export `registerOrchestratorRoutes`
- `src/web/ports/infra-port.ts` — Add `orchestratorLoop` to InfraPort
- `src/types/index.ts` — Export orchestrator types
## Existing File Modifications Summary
| File | Change | Lines |
|------|--------|-------|
| `src/types/index.ts` | Add orchestrator barrel export | +1 |
| `src/web/sse-events.ts` | Add 11 orchestrator events + SseEvent entries | +30 |
| `src/web/public/constants.js` | Mirror 11 SSE events | +15 |
| `src/web/routes/index.ts` | Export registerOrchestratorRoutes | +1 |
| `src/web/ports/infra-port.ts` | Add orchestratorLoop to InfraPort | +3 |
| `src/web/server.ts` | Initialize OrchestratorLoop, wire events, register routes | +40 |
| `src/web/schemas.ts` | Add orchestrator Zod schemas | +20 |
| `src/state-store.ts` | Add orchestrator state persistence | +20 |
| `src/web/public/app.js` | Add orchestrator SSE listeners + panel toggle | +30 |
| `src/web/public/index.html` | Add orchestrator-ui.js script tag | +1 |
**Total new code**: ~2,300 lines across 6 new files
**Total modifications**: ~160 lines across 10 existing files
## Implementation Execution Order
This is the actual build order — each step is a commit checkpoint:
1. **Types** — `src/types/orchestrator.ts` + barrel export. Zero risk, pure types.
2. **SSE events** — Add all 11 events to both `sse-events.ts` and `constants.js`. Wire in SseEvent namespace.
3. **State persistence** — Add orchestrator state to StateStore. Small, isolated change.
4. **Schemas** — Add Zod validation schemas for API input.
5. **Planner** — `src/orchestrator-planner.ts`. Can test in isolation.
6. **Verifier** — `src/orchestrator-verifier.ts`. Can test in isolation.
7. **Core loop** — `src/orchestrator-loop.ts`. The big one. Depends on planner + verifier.
8. **Prompts** — `src/prompts/orchestrator.ts`. Templates used by core loop.
9. **Port + routes** — `src/web/ports/infra-port.ts` update + `src/web/routes/orchestrator-routes.ts`.
10. **Server integration** — Wire OrchestratorLoop into WebServer. Routes become live.
11. **Frontend** — `src/web/public/orchestrator-ui.js` + app.js listeners + index.html script tag.
12. **Tests** — `test/orchestrator-*.test.ts`.
13. **Typecheck + lint** — Fix all issues, ensure CI passes.
## Edge Cases & Error Handling
- **Session limit reached**: Queue tasks and wait for sessions to free up (existing SessionManager handles this)
- **All sessions crash during phase**: Mark phase as failed, attempt replan
- **Verification flaky**: `moderate` mode allows test retries; `lenient` skips AI review
- **Plan too large**: Cap at 10 phases, 50 total tasks. Warn user.
- **Context overflow**: Auto-compact between phases. Respawn if needed (orchestrator state is external).
- **User pauses mid-phase**: Pause task assignment, don't cancel running tasks. Resume picks up where it left off.
- **Network/API errors during planning**: Retry plan generation up to 2 times, then fail with clear message.
- **Orchestrator vs Ralph conflict**: Mutually exclusive. Starting orchestrator stops Ralph if running. Starting Ralph stops orchestrator.
## Testing Strategy
- **Unit tests**: `test/orchestrator-planner.test.ts` — phase grouping algorithm, team strategy assignment
- **Unit tests**: `test/orchestrator-verifier.test.ts` — verification logic with mocked sessions
- **Integration tests**: `test/orchestrator-loop.test.ts` — state machine transitions, task lifecycle
- **Route tests**: `test/routes/orchestrator-routes.test.ts` — API validation, status responses
All tests use `MockSession` pattern from existing test infrastructure. No real tmux needed.
+157
View File
@@ -0,0 +1,157 @@
# Orchestrator Loop — Research Findings
> Research doc for the new "Orchestrator Loop" feature. Not for GitHub.
## What We're Building
A new autonomous loop variant — **Orchestrator Loop** — that takes high-level user tasks, decomposes them into a detailed plan using team agents, and executes the plan step-by-step with quality gates. Unlike Ralph Loop (which executes a flat task queue), the Orchestrator coordinates **planning, delegation, and verification** as a continuous cycle.
**Core idea**: User inputs a goal → Orchestrator creates a detailed plan → spins up team agents for parallel execution → validates each step → adapts the plan based on results → delivers polished output.
## Existing Infrastructure Analysis
### What We Can Reuse
#### 1. Ralph Loop (`src/ralph-loop.ts`)
- **Pattern**: Poll loop with `start() → tick() → stop()` lifecycle
- **Reusable**: Event-driven task assignment, session completion handling, timeout management
- **Limitation**: Flat task queue — no concept of phases, dependencies between task groups, or adaptive replanning
- **Key insight**: `assignTaskToSession()` uses `session.sendInput(task.prompt)` — simple prompt injection into PTY
#### 2. Task Queue (`src/task-queue.ts`) + Task (`src/task.ts`)
- **Already has**: Priority ordering, dependency tracking between tasks, completion phrase detection
- **Limitation**: No task *groups* or *phases*. Dependencies are task-to-task, not phase-to-phase
- **Key insight**: Tasks support `completionPhrase` — a string the task watches for in output. This is how Ralph knows a task is done
#### 3. Plan Orchestrator (`src/plan-orchestrator.ts`)
- **Already has**: 2-agent plan generation (Research Agent → Planner Agent), TDD-aware plan items with P0/P1/P2 priorities
- **Output**: `PlanItem[]` with dependencies, verification criteria, TDD phases, complexity ratings
- **Limitation**: Plan generation only — no execution. Plans are generated then sit in state/UI for human review
- **Key insight**: Uses `Session` directly to run Claude subagent instances for research and planning. Returns structured JSON
#### 4. Team Agents (`src/team-watcher.ts`, `~/.claude/teams/`)
- **Already has**: Team creation, member tracking, filesystem inbox messaging, task management via `~/.claude/tasks/{team-name}/`
- **Limitation**: Codeman can only *observe* teams (TeamWatcher is read-only polling), not *create* or *orchestrate* them
- **Key insight**: Teams are a Claude Code feature. Codeman monitors them but doesn't control them. We can't programmatically create teammates — Claude Code does that when you use `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
#### 5. Respawn Controller (`src/respawn-controller.ts`)
- **Already has**: Preset-based automation (ralph-todo, overnight-autonomous), circuit breaker, health scoring
- **Key insight**: The `ralph-todo` preset (8s idle, 480min max) is designed for autonomous task execution. We'd need a new preset or make Orchestrator Loop set its own timing
#### 6. Session Auto-Ops (`src/session-auto-ops.ts`)
- **Already has**: Auto-compact at token thresholds, auto-clear for context management
- **Key insight**: Critical for long Orchestrator runs — prevents context overflow during multi-step execution
#### 7. Hooks (`src/hooks-config.ts`)
- **Already has**: `idle_prompt`, `stop`, `teammate_idle`, `task_completed` hook events
- **Key insight**: Hooks fire POST to `/api/hook-event` — this is how Codeman knows when Claude is idle, stopped, or completed a task. The Orchestrator Loop can listen to these same events
### What We Need to Build New
1. **Plan → Task decomposition**: Convert PlanOrchestrator output (PlanItem[]) into executable task groups with phase ordering
2. **Multi-phase execution engine**: Execute plan phases sequentially, tasks within phases in parallel
3. **Verification gates**: After each phase, run verification (test commands, AI review) before proceeding
4. **Adaptive replanning**: When a task fails or verification fails, generate a recovery plan
5. **Team agent orchestration**: Leverage Claude Code's agent teams for parallel execution within phases
6. **Progress tracking & UI**: Real-time dashboard showing plan progress, phase status, agent activity
## How Teams Actually Work (Important Constraint)
After deep research, here's the reality of agent teams:
```
User starts session with CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
→ Claude Code creates a team-lead
→ Team-lead spawns teammates (in-process threads)
→ Teammates appear as subagents (detected by SubagentWatcher)
→ Communication via ~/.claude/teams/{name}/inboxes/{member}.json
→ Tasks tracked in ~/.claude/tasks/{team-name}/{N}.json
```
**Codeman cannot programmatically create team members.** This is a Claude Code internal feature. However, Codeman CAN:
- Start a session that has teams enabled
- Send a prompt to the lead that instructs it to use agent teams
- Monitor team activity via TeamWatcher
- React to teammate_idle and task_completed hook events
- Read team task status from the filesystem
**This means**: The Orchestrator Loop orchestrates at the *session prompt* level, not the *team member* level. We tell the lead what to do, and the lead decides how to use its team.
## Architecture Decision: Prompt-Level Orchestration
Given the team constraint, the Orchestrator Loop works by:
1. **Planning phase**: Use PlanOrchestrator to generate a detailed plan from user input
2. **Execution phase**: Feed plan steps as prompts to sessions, one phase at a time
3. **Verification phase**: After each phase, run verification prompts and check results
4. **Adaptation phase**: If verification fails, generate recovery prompts
The "team agents" aspect works by:
- Starting sessions with `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
- Crafting prompts that *instruct the lead to delegate* to teammates
- Monitoring team activity to track parallel progress
- The lead agent is smart enough to decompose work across its team
## Key Technical Findings
### Session Input Mechanics
```typescript
// From session.ts - how we send prompts
await session.sendInput(task.prompt); // Uses writeViaMux() internally
// writeViaMux() does: tmux send-keys -l "prompt text" + tmux send-keys Enter
// CRITICAL: Single-line only! Multi-line breaks Ink rendering
```
### Completion Detection Chain
```
PTY output → RalphTracker.processData() → completion phrase fuzzy match
→ CompletionConfidence scoring (multi-signal: promise tag + todos + exit signal)
→ If confident → emit 'completionDetected'
→ RalphLoop listens → marks task complete → assigns next
```
### How Plan Items Map to Tasks
```typescript
// PlanItem has:
interface PlanItem {
id: string; // "P0-001"
content: string; // "Implement error handling for API endpoints"
priority: 'P0' | 'P1' | 'P2';
dependencies: string[]; // ["P0-000"] — other PlanItem IDs
verificationCriteria: string;
testCommand: string;
tddPhase: 'setup' | 'test' | 'impl' | 'verify' | 'review';
complexity: 'low' | 'medium' | 'high';
}
// Task has:
interface CreateTaskOptions {
prompt: string;
priority: number;
dependencies: string[]; // Task IDs
completionPhrase: string;
timeoutMs: number;
}
// Natural mapping: PlanItem.content → Task.prompt
// PlanItem.dependencies → Task.dependencies
// PlanItem.priority → Task.priority (P0=100, P1=50, P2=10)
// PlanItem.verificationCriteria → verification task prompt
```
### Context Management for Long Runs
- Auto-compact at ~110k tokens (configurable)
- Auto-clear at ~140k tokens (configurable)
- Respawn cycling: kill + restart session to reset context entirely
- For Orchestrator: we want compact between phases, respawn between major milestones
## Risk Assessment
| Risk | Severity | Mitigation |
|------|----------|------------|
| Context overflow during complex phases | High | Auto-compact between tasks, respawn between phases |
| Team agents not predictable | Medium | Orchestrate at session level, let Claude decide team delegation |
| Plan too ambitious → infinite loop | High | Phase budgets (max attempts per phase), circuit breaker |
| Verification too strict → blocks progress | Medium | Configurable strictness, human override via UI |
| Single-line prompt limit | Medium | Use CLAUDE.md file for complex instructions, prompt references file |
| Long planning phase delays execution | Low | Show plan for approval before execution |
+167 -21
View File
@@ -7,8 +7,10 @@
# Environment variables:
# CODEMAN_NONINTERACTIVE=1 - Skip all prompts (for CI/automation)
# CODEMAN_INSTALL_DIR - Custom install directory (default: ~/.codeman/app)
# CODEMAN_SKIP_SYSTEMD=1 - Skip systemd service setup prompt
# CODEMAN_SKIP_SYSTEMD=1 - Skip systemd/launchd service setup prompt
# CODEMAN_NODE_VERSION - Node.js major version to install (default: 22)
# CODEMAN_REPO_URL - Custom git repository URL (default: upstream Codeman)
# CODEMAN_BRANCH - Git branch to install (default: master)
set -euo pipefail
@@ -17,7 +19,8 @@ set -euo pipefail
# ============================================================================
INSTALL_DIR="${CODEMAN_INSTALL_DIR:-$HOME/.codeman/app}"
REPO_URL="https://github.com/Ark0N/Codeman.git"
REPO_URL="${CODEMAN_REPO_URL:-https://github.com/Ark0N/Codeman.git}"
BRANCH="${CODEMAN_BRANCH:-master}"
MIN_NODE_VERSION=18
TARGET_NODE_VERSION="${CODEMAN_NODE_VERSION:-22}"
NONINTERACTIVE="${CODEMAN_NONINTERACTIVE:-0}"
@@ -350,9 +353,16 @@ ensure_sudo() {
die "sudo is required but not installed. Please install packages manually or run as root."
fi
# Validate sudo access
if ! sudo -v 2>/dev/null; then
# When piped (curl | bash), stdin is the pipe — redirect from /dev/tty so sudo can prompt
if [[ -e /dev/tty ]]; then
if ! sudo -v 2>/dev/null < /dev/tty; then
die "Failed to obtain sudo privileges."
fi
else
if ! sudo -v 2>/dev/null; then
die "Failed to obtain sudo privileges. Try running the script directly instead of piping."
fi
fi
}
run_as_root() {
@@ -369,7 +379,12 @@ ensure_homebrew() {
fi
info "Installing Homebrew first..."
/bin/bash -c "$(download_to_stdout https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"
# When piped (curl | bash), stdin is the pipe — Homebrew needs TTY for sudo password prompt
if [[ -e /dev/tty ]]; then
/bin/bash -c "$(download_to_stdout https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)" < /dev/tty
else
NONINTERACTIVE=1 /bin/bash -c "$(download_to_stdout https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"
fi
# Add Homebrew to PATH for Apple Silicon
if [[ -f /opt/homebrew/bin/brew ]]; then
@@ -784,9 +799,84 @@ setup_sc_alias() {
}
# ============================================================================
# Systemd Service Setup (Linux only)
# Service Setup (Linux systemd / macOS launchd)
# ============================================================================
setup_launchd_service() {
local plist_label="com.codeman.web"
local agent_dir="$HOME/Library/LaunchAgents"
local agent_plist="$agent_dir/$plist_label.plist"
local daemon_plist="/Library/LaunchDaemons/$plist_label.plist"
info "Setting up macOS LaunchAgent..."
# Remove any existing LaunchDaemon (system-level) to prevent duplicates.
# We standardize on LaunchAgent (user-level) — it doesn't require sudo,
# inherits the user's environment, and is the correct choice for user apps.
if [[ -f "$daemon_plist" ]]; then
warn "Found system-level LaunchDaemon at $daemon_plist — removing to prevent duplicate"
sudo launchctl unload "$daemon_plist" 2>/dev/null || true
sudo rm -f "$daemon_plist"
success "Removed duplicate LaunchDaemon"
fi
# Unload existing agent before overwriting
if [[ -f "$agent_plist" ]]; then
launchctl unload "$agent_plist" 2>/dev/null || true
fi
mkdir -p "$agent_dir"
# Build PATH: ensure /opt/homebrew/bin (Apple Silicon) and ~/.local/bin are included
local svc_path="/opt/homebrew/bin:/usr/local/bin:$HOME/.local/bin:/usr/bin:/bin:/usr/sbin:/sbin"
# Find node binary path
local node_path
node_path=$(command -v node)
cat > "$agent_plist" << EOF
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>$plist_label</string>
<key>ProgramArguments</key>
<array>
<string>$node_path</string>
<string>$INSTALL_DIR/dist/index.js</string>
<string>web</string>
</array>
<key>EnvironmentVariables</key>
<dict>
<key>PATH</key>
<string>$svc_path</string>
<key>HOME</key>
<string>$HOME</string>
<key>LANG</key>
<string>en_US.UTF-8</string>
</dict>
<key>WorkingDirectory</key>
<string>$HOME</string>
<key>RunAtLoad</key>
<true/>
<key>KeepAlive</key>
<true/>
<key>ThrottleInterval</key>
<integer>10</integer>
<key>StandardOutPath</key>
<string>/tmp/codeman.log</string>
<key>StandardErrorPath</key>
<string>/tmp/codeman.log</string>
</dict>
</plist>
EOF
launchctl load "$agent_plist" 2>/dev/null || true
success "LaunchAgent installed and started"
}
setup_systemd_service() {
local service_dir="$HOME/.config/systemd/user"
local service_file="$service_dir/codeman-web.service"
@@ -1062,26 +1152,27 @@ main() {
if [[ -d "$INSTALL_DIR/.git" ]]; then
info "Existing installation found, updating..."
cd "$INSTALL_DIR"
git remote set-url origin "$REPO_URL" 2>/dev/null || true
# Check for local changes
if ! git diff --quiet 2>/dev/null || ! git diff --staged --quiet 2>/dev/null; then
warn "Local changes detected in $INSTALL_DIR"
if prompt_yes_no "Discard local changes and update?" "n"; then
git fetch --quiet origin
git reset --hard origin/master --quiet
git reset --hard "origin/$BRANCH" --quiet
else
info "Keeping existing installation, skipping update"
fi
else
git fetch --quiet origin
git reset --hard origin/master --quiet
git reset --hard "origin/$BRANCH" --quiet
fi
else
# Create parent directory
mkdir -p "$(dirname "$INSTALL_DIR")"
# Clone repository (shallow for speed)
git clone --quiet --depth 1 "$REPO_URL" "$INSTALL_DIR"
git clone --quiet --depth 1 --branch "$BRANCH" "$REPO_URL" "$INSTALL_DIR"
cd "$INSTALL_DIR"
fi
@@ -1135,17 +1226,25 @@ main() {
echo ""
local launch_choice=""
local has_systemd=false
local has_service=false
local service_type=""
if [[ "$os" == "linux" ]] && [[ "$SKIP_SYSTEMD" != "1" ]] && command -v systemctl &>/dev/null; then
has_systemd=true
has_service=true
service_type="systemd"
elif [[ "$os" == "macos" ]] && [[ "$SKIP_SYSTEMD" != "1" ]]; then
has_service=true
service_type="launchd"
fi
if [[ "$has_systemd" == "true" ]]; then
if [[ "$has_service" == "true" ]]; then
local service_label="systemd service"
[[ "$service_type" == "launchd" ]] && service_label="LaunchAgent"
echo -e " ${BOLD}How would you like to run Codeman?${NC}"
echo ""
echo -e " ${CYAN}1)${NC} Run now in this terminal"
echo -e " ${CYAN}2)${NC} Install as systemd service (auto-start on boot)"
echo -e " ${CYAN}2)${NC} Install as $service_label (auto-start on boot)"
echo -e " ${CYAN}3)${NC} Don't start — I'll run it later"
echo ""
@@ -1162,7 +1261,7 @@ main() {
done
fi
else
# macOS or no systemd — only offer run now or skip
# No service manager available — only offer run now or skip
echo -e " ${BOLD}Would you like to start Codeman now?${NC}"
echo ""
echo -e " ${CYAN}1)${NC} Run now in this terminal"
@@ -1188,12 +1287,16 @@ main() {
echo ""
# Handle systemd setup
# Handle service setup
if [[ "$launch_choice" == "2" ]]; then
if [[ "$service_type" == "launchd" ]]; then
setup_launchd_service
else
setup_systemd_service
fi
# Offer tunnel service if cloudflared is available
if check_cloudflared && [[ -f "$INSTALL_DIR/scripts/codeman-tunnel.service" ]]; then
# Offer tunnel service if cloudflared is available (Linux only — systemd tunnel service)
if [[ "$service_type" == "systemd" ]] && check_cloudflared && [[ -f "$INSTALL_DIR/scripts/codeman-tunnel.service" ]]; then
echo ""
if prompt_yes_no "Also set up Cloudflare tunnel service? (requires CODEMAN_PASSWORD)" "n"; then
setup_tunnel_service
@@ -1208,10 +1311,16 @@ main() {
echo ""
echo -e " ${BOLD}Manage the service:${NC}"
echo ""
if [[ "$service_type" == "launchd" ]]; then
echo -e " ${CYAN}launchctl unload ~/Library/LaunchAgents/com.codeman.web.plist${NC} # Stop"
echo -e " ${CYAN}launchctl load ~/Library/LaunchAgents/com.codeman.web.plist${NC} # Start"
echo -e " ${CYAN}tail -f /tmp/codeman.log${NC} # View logs"
else
echo -e " ${CYAN}systemctl --user stop codeman-web${NC} # Stop"
echo -e " ${CYAN}systemctl --user restart codeman-web${NC} # Restart"
echo -e " ${CYAN}systemctl --user status codeman-web${NC} # Check status"
echo -e " ${CYAN}journalctl --user -u codeman-web -f${NC} # View logs"
fi
echo ""
fi
@@ -1277,13 +1386,29 @@ update() {
info "Updating Codeman..."
cd "$INSTALL_DIR"
git remote set-url origin "$REPO_URL" 2>/dev/null || true
git fetch --quiet origin
git reset --hard origin/master --quiet
git reset --hard "origin/$BRANCH" --quiet
npm install --quiet --no-fund --no-audit 2>/dev/null || npm install --no-fund --no-audit
npm run build --quiet 2>/dev/null || npm run build
success "Updated to $(node -e "console.log(require('./package.json').version)")"
echo ""
echo -e " ${DIM}Restart codeman web to use the new version.${NC}"
# Auto-restart service if running, otherwise tell the user
local agent_plist="$HOME/Library/LaunchAgents/com.codeman.web.plist"
if systemctl --user is-active codeman-web.service &>/dev/null 2>&1; then
info "Restarting codeman-web service..."
systemctl --user restart codeman-web.service
success "codeman-web service restarted"
elif [[ -f "$agent_plist" ]]; then
info "Restarting LaunchAgent..."
launchctl unload "$agent_plist" 2>/dev/null || true
launchctl load "$agent_plist" 2>/dev/null || true
success "LaunchAgent restarted"
else
echo -e " ${DIM}Restart codeman web to use the new version:${NC}"
echo -e " ${CYAN}pkill -f 'codeman.*web'; codeman web &${NC}"
fi
echo ""
}
@@ -1292,9 +1417,9 @@ uninstall() {
info "Uninstalling Codeman..."
echo ""
# Stop and remove systemd services
# Stop and remove systemd services (Linux)
for svc in codeman-web codeman-tunnel; do
if systemctl --user is-active "${svc}.service" &>/dev/null; then
if systemctl --user is-active "${svc}.service" &>/dev/null 2>&1; then
info "Stopping ${svc} service..."
systemctl --user stop "${svc}.service"
fi
@@ -1310,6 +1435,20 @@ uninstall() {
done
systemctl --user daemon-reload 2>/dev/null || true
# Stop and remove launchd services (macOS)
local agent_plist="$HOME/Library/LaunchAgents/com.codeman.web.plist"
local daemon_plist="/Library/LaunchDaemons/com.codeman.web.plist"
if [[ -f "$agent_plist" ]]; then
launchctl unload "$agent_plist" 2>/dev/null || true
rm -f "$agent_plist"
success "Removed LaunchAgent"
fi
if [[ -f "$daemon_plist" ]]; then
sudo launchctl unload "$daemon_plist" 2>/dev/null || true
sudo rm -f "$daemon_plist"
success "Removed LaunchDaemon"
fi
# Remove symlinks
local symlink_dir="$HOME/.local/bin"
if [[ -L "$symlink_dir/codeman" ]]; then
@@ -1355,5 +1494,12 @@ uninstall() {
case "${1:-}" in
update) update ;;
uninstall) uninstall ;;
*) main "$@" ;;
*)
if [[ -z "${1:-}" && -d "$INSTALL_DIR/.git" ]]; then
print_banner
update
else
main "$@"
fi
;;
esac
+16
View File
@@ -0,0 +1,16 @@
{
"$schema": "https://unpkg.com/knip@5/schema.json",
"entry": [
"scripts/*.mjs",
"scripts/*.js",
"scripts/watch-subagents.ts",
"scripts/remotion/Root.tsx",
"scripts/remotion/index.ts",
"test/**/*.test.ts",
"test/mobile/vitest.config.ts",
"test/**/*.mjs"
],
"project": ["src/**/*.{ts,tsx}", "scripts/**/*.{ts,tsx,mjs,js}", "test/**/*.{ts,mjs}"],
"ignoreExportsUsedInFile": true,
"ignoreDependencies": ["@remotion/cli", "@remotion/transitions", "esbuild", "agent-browser"]
}
+140 -56
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "0.3.1",
"version": "0.6.3",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "0.3.1",
"version": "0.6.3",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
@@ -17,8 +17,11 @@
"@fastify/compress": "^8.3.1",
"@fastify/cookie": "^11.0.2",
"@fastify/static": "^8.0.0",
"@remotion/compositor-linux-x64-gnu": "^4.0.432",
"@rspack/binding-linux-x64-gnu": "^1.7.7",
"@fastify/websocket": "^11.2.0",
"@xterm/addon-fit": "^0.11.0",
"@xterm/addon-unicode11": "^0.9.0",
"@xterm/addon-webgl": "^0.19.0",
"@xterm/xterm": "^6.0.0",
"chalk": "^5.3.0",
"chokidar": "^3.6.0",
"commander": "^12.1.0",
@@ -27,10 +30,6 @@
"qrcode": "^1.5.4",
"uuid": "^10.0.0",
"web-push": "^3.6.7",
"xterm": "^5.3.0",
"xterm-addon-fit": "^0.8.0",
"xterm-addon-unicode11": "^0.6.0",
"xterm-addon-webgl": "^0.16.0",
"zod": "^4.3.6"
},
"bin": {
@@ -47,6 +46,7 @@
"@types/react": "^19.2.14",
"@types/uuid": "^10.0.0",
"@types/web-push": "^3.6.4",
"@types/ws": "^8.18.1",
"@vitest/coverage-v8": "^4.0.18",
"agent-browser": "^0.6.0",
"esbuild": "^0.27.3",
@@ -64,6 +64,10 @@
},
"engines": {
"node": ">=18.0.0"
},
"optionalDependencies": {
"@remotion/compositor-linux-x64-gnu": "^4.0.432",
"@rspack/binding-linux-x64-gnu": "^1.7.7"
}
},
"node_modules/@asamuzakjp/css-color": {
@@ -943,6 +947,53 @@
"glob": "^11.0.0"
}
},
"node_modules/@fastify/websocket": {
"version": "11.2.0",
"resolved": "https://registry.npmjs.org/@fastify/websocket/-/websocket-11.2.0.tgz",
"integrity": "sha512-3HrDPbAG1CzUCqnslgJxppvzaAZffieOVbLp1DAy1huCSynUWPifSvfdEDUR8HlJLp3sp1A36uOM2tJogADS8w==",
"funding": [
{
"type": "github",
"url": "https://github.com/sponsors/fastify"
},
{
"type": "opencollective",
"url": "https://opencollective.com/fastify"
}
],
"license": "MIT",
"dependencies": {
"duplexify": "^4.1.3",
"fastify-plugin": "^5.0.0",
"ws": "^8.16.0"
}
},
"node_modules/@fastify/websocket/node_modules/duplexify": {
"version": "4.1.3",
"resolved": "https://registry.npmjs.org/duplexify/-/duplexify-4.1.3.tgz",
"integrity": "sha512-M3BmBhwJRZsSx38lZyhE53Csddgzl5R7xGJNk7CVddZD6CcmwMCH8J+7AprIrQKH7TonKxaCjcv27Qmf+sQ+oA==",
"license": "MIT",
"dependencies": {
"end-of-stream": "^1.4.1",
"inherits": "^2.0.3",
"readable-stream": "^3.1.1",
"stream-shift": "^1.0.2"
}
},
"node_modules/@fastify/websocket/node_modules/readable-stream": {
"version": "3.6.2",
"resolved": "https://registry.npmjs.org/readable-stream/-/readable-stream-3.6.2.tgz",
"integrity": "sha512-9u/sniCrY3D5WdsERHzHE4G2YCXqoG5FTHUiCC4SIbr6XcLZBY05ya9EKjYek9O5xOAwjGq+1JdGBAS7Q9ScoA==",
"license": "MIT",
"dependencies": {
"inherits": "^2.0.3",
"string_decoder": "^1.1.1",
"util-deprecate": "^1.0.1"
},
"engines": {
"node": ">= 6"
}
},
"node_modules/@humanfs/core": {
"version": "0.19.1",
"dev": true,
@@ -1432,6 +1483,7 @@
"cpu": [
"x64"
],
"optional": true,
"os": [
"linux"
]
@@ -1861,6 +1913,7 @@
"x64"
],
"license": "MIT",
"optional": true,
"os": [
"linux"
]
@@ -2114,6 +2167,16 @@
"@types/node": "*"
}
},
"node_modules/@types/ws": {
"version": "8.18.1",
"resolved": "https://registry.npmjs.org/@types/ws/-/ws-8.18.1.tgz",
"integrity": "sha512-ThVF6DCVhA8kUGy+aazFQ4kXQ7E1Ty7A3ypFOe0IcJV8O/M511G99AW24irKrW56Wt44yG9+ij8FaqoBGkuBXg==",
"dev": true,
"license": "MIT",
"dependencies": {
"@types/node": "*"
}
},
"node_modules/@types/yauzl": {
"version": "2.10.3",
"dev": true,
@@ -2599,6 +2662,33 @@
"@xtuc/long": "4.2.2"
}
},
"node_modules/@xterm/addon-fit": {
"version": "0.11.0",
"resolved": "https://registry.npmjs.org/@xterm/addon-fit/-/addon-fit-0.11.0.tgz",
"integrity": "sha512-jYcgT6xtVYhnhgxh3QgYDnnNMYTcf8ElbxxFzX0IZo+vabQqSPAjC3c1wJrKB5E19VwQei89QCiZZP86DCPF7g==",
"license": "MIT"
},
"node_modules/@xterm/addon-unicode11": {
"version": "0.9.0",
"resolved": "https://registry.npmjs.org/@xterm/addon-unicode11/-/addon-unicode11-0.9.0.tgz",
"integrity": "sha512-FxDnYcyuXhNl+XSqGZL/t0U9eiNb/q3EWT5rYkQT/zuig8Gz/VagnQANKHdDWFM2lTMk9ly0EFQxxxtZUoRetw==",
"license": "MIT"
},
"node_modules/@xterm/addon-webgl": {
"version": "0.19.0",
"resolved": "https://registry.npmjs.org/@xterm/addon-webgl/-/addon-webgl-0.19.0.tgz",
"integrity": "sha512-b3fMOsyLVuCeNJWxolACEUED0vm7qC0cy4wRvf3oURSzDTYVQiGPhTnhWZwIHdvC48Y+oLhvYXnY4XDXPoJo6A==",
"license": "MIT"
},
"node_modules/@xterm/xterm": {
"version": "6.0.0",
"resolved": "https://registry.npmjs.org/@xterm/xterm/-/xterm-6.0.0.tgz",
"integrity": "sha512-TQwDdQGtwwDt+2cgKDLn0IRaSxYu1tSUjgKarSDkUM0ZNiSRXFpjxEsvc/Zgc5kq5omJ+V0a8/kIM2WD3sMOYg==",
"license": "MIT",
"workspaces": [
"addons/*"
]
},
"node_modules/@xtuc/ieee754": {
"version": "1.2.0",
"dev": true,
@@ -2988,7 +3078,9 @@
}
},
"node_modules/basic-ftp": {
"version": "5.1.0",
"version": "5.2.0",
"resolved": "https://registry.npmjs.org/basic-ftp/-/basic-ftp-5.2.0.tgz",
"integrity": "sha512-VoMINM2rqJwJgfdHq6RiUudKt2BV+FY5ZFezP/ypmwayk68+NzzAQy4XXLlqsGD4MCzq3DrmNFD/uUmBJuGoXw==",
"dev": true,
"license": "MIT",
"engines": {
@@ -4385,7 +4477,9 @@
"license": "BSD-3-Clause"
},
"node_modules/fastify": {
"version": "5.7.4",
"version": "5.8.2",
"resolved": "https://registry.npmjs.org/fastify/-/fastify-5.8.2.tgz",
"integrity": "sha512-lZmt3navvZG915IE+f7/TIVamxIwmBd+OMB+O9WBzcpIwOo6F0LTh0sluoMFk5VkrKTvvrwIaoJPkir4Z+jtAg==",
"funding": [
{
"type": "github",
@@ -4407,7 +4501,7 @@
"fast-json-stringify": "^6.0.0",
"find-my-way": "^9.0.0",
"light-my-request": "^6.0.0",
"pino": "^10.1.0",
"pino": "^9.14.0 || ^10.1.0",
"process-warning": "^5.0.0",
"rfdc": "^1.3.1",
"secure-json-parse": "^4.0.0",
@@ -4570,6 +4664,20 @@
"dev": true,
"license": "Unlicense"
},
"node_modules/fsevents": {
"version": "2.3.3",
"resolved": "https://registry.npmjs.org/fsevents/-/fsevents-2.3.3.tgz",
"integrity": "sha512-5xoDfX+fL7faATnagmWPpbFtwh/R77WmMMqqHGS65C3vvB0YHrgF+B1YmZ3441tMj5n63k0212XNoJwzlhffQw==",
"hasInstallScript": true,
"license": "MIT",
"optional": true,
"os": [
"darwin"
],
"engines": {
"node": "^8.16.0 || ^10.6.0 || >=11.0.0"
}
},
"node_modules/function-bind": {
"version": "1.1.2",
"dev": true,
@@ -5610,7 +5718,9 @@
"license": "ISC"
},
"node_modules/minimatch": {
"version": "10.2.2",
"version": "10.2.4",
"resolved": "https://registry.npmjs.org/minimatch/-/minimatch-10.2.4.tgz",
"integrity": "sha512-oRjTw/97aTBN0RHbYCdtF1MQfvusSIBQM0IZEgzl6426+8jSC0nF1a/GmnVLpfB9yyr6g6FTqWqiZVbxrtaCIg==",
"license": "BlueOak-1.0.0",
"dependencies": {
"brace-expansion": "^5.0.2"
@@ -6132,6 +6242,21 @@
"node": ">=18"
}
},
"node_modules/playwright/node_modules/fsevents": {
"version": "2.3.2",
"resolved": "https://registry.npmjs.org/fsevents/-/fsevents-2.3.2.tgz",
"integrity": "sha512-xiqMQR4xAeHTuB9uWm+fFRcIOgKBMiOBP+eXiyT7jsgVCq1bkVygt00oASowB7EdtpOHaaPgKt812P9ab+DDKA==",
"dev": true,
"hasInstallScript": true,
"license": "MIT",
"optional": true,
"os": [
"darwin"
],
"engines": {
"node": "^8.16.0 || ^10.6.0 || >=11.0.0"
}
},
"node_modules/pngjs": {
"version": "7.0.0",
"dev": true,
@@ -6594,14 +6719,6 @@
"version": "4.0.4",
"license": "MIT"
},
"node_modules/randombytes": {
"version": "2.1.0",
"dev": true,
"license": "MIT",
"dependencies": {
"safe-buffer": "^5.1.0"
}
},
"node_modules/react": {
"version": "19.2.4",
"dev": true,
@@ -6992,14 +7109,6 @@
"node": ">=10"
}
},
"node_modules/serialize-javascript": {
"version": "6.0.2",
"dev": true,
"license": "BSD-3-Clause",
"dependencies": {
"randombytes": "^2.1.0"
}
},
"node_modules/set-blocking": {
"version": "2.0.0",
"license": "ISC"
@@ -7370,14 +7479,15 @@
}
},
"node_modules/terser-webpack-plugin": {
"version": "5.3.16",
"version": "5.4.0",
"resolved": "https://registry.npmjs.org/terser-webpack-plugin/-/terser-webpack-plugin-5.4.0.tgz",
"integrity": "sha512-Bn5vxm48flOIfkdl5CaD2+1CiUVbonWQ3KQPyP7/EuIl9Gbzq/gQFOzaMFUEgVjB1396tcK0SG8XcNJ/2kDH8g==",
"dev": true,
"license": "MIT",
"dependencies": {
"@jridgewell/trace-mapping": "^0.3.25",
"jest-worker": "^27.4.5",
"schema-utils": "^4.3.0",
"serialize-javascript": "^6.0.2",
"terser": "^5.31.1"
},
"engines": {
@@ -8457,7 +8567,6 @@
},
"node_modules/ws": {
"version": "8.19.0",
"dev": true,
"license": "MIT",
"engines": {
"node": ">=10.0.0"
@@ -8495,31 +8604,6 @@
"node": ">=0.4"
}
},
"node_modules/xterm": {
"version": "5.3.0",
"license": "MIT"
},
"node_modules/xterm-addon-fit": {
"version": "0.8.0",
"license": "MIT",
"peerDependencies": {
"xterm": "^5.0.0"
}
},
"node_modules/xterm-addon-unicode11": {
"version": "0.6.0",
"license": "MIT",
"peerDependencies": {
"xterm": "^5.0.0"
}
},
"node_modules/xterm-addon-webgl": {
"version": "0.16.0",
"license": "MIT",
"peerDependencies": {
"xterm": "^5.0.0"
}
},
"node_modules/xterm-zerolag-input": {
"resolved": "packages/xterm-zerolag-input",
"link": true
+20 -14
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "0.3.6",
"version": "0.6.3",
"description": "The missing control plane for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -11,21 +11,23 @@
"scripts": {
"postinstall": "node scripts/postinstall.js",
"build": "node scripts/build.mjs",
"start": "node dist/index.js",
"start": "NODE_COMPILE_CACHE=${HOME}/.codeman/compile-cache node dist/index.js",
"dev": "tsx src/index.ts web",
"web": "node dist/index.js web",
"clean": "rm -rf dist",
"test": "vitest run",
"test:watch": "vitest",
"test:coverage": "vitest run --coverage",
"test": "vitest run --config config/vitest.config.ts",
"test:watch": "vitest --config config/vitest.config.ts",
"test:coverage": "vitest run --config config/vitest.config.ts --coverage",
"typecheck": "tsc --noEmit",
"lint": "eslint 'src/**/*.ts'",
"lint:fix": "eslint 'src/**/*.ts' --fix",
"lint": "eslint --config config/eslint.config.js 'src/**/*.ts'",
"lint:fix": "eslint --config config/eslint.config.js 'src/**/*.ts' --fix",
"format": "prettier --write 'src/**/*.ts'",
"format:check": "prettier --check 'src/**/*.ts'",
"capture:subagents": "node scripts/capture-subagent-screenshots.mjs",
"changeset": "changeset",
"version-packages": "changeset version",
"version-packages": "changeset version && npm install --package-lock-only && node scripts/check-lockfile-sync.mjs",
"check:lockfile": "node scripts/check-lockfile-sync.mjs",
"knip": "npx --yes knip@latest",
"release": "changeset publish"
},
"workspaces": [
@@ -51,8 +53,11 @@
"@fastify/compress": "^8.3.1",
"@fastify/cookie": "^11.0.2",
"@fastify/static": "^8.0.0",
"@remotion/compositor-linux-x64-gnu": "^4.0.432",
"@rspack/binding-linux-x64-gnu": "^1.7.7",
"@fastify/websocket": "^11.2.0",
"@xterm/addon-fit": "^0.11.0",
"@xterm/addon-unicode11": "^0.9.0",
"@xterm/addon-webgl": "^0.19.0",
"@xterm/xterm": "^6.0.0",
"chalk": "^5.3.0",
"chokidar": "^3.6.0",
"commander": "^12.1.0",
@@ -61,10 +66,6 @@
"qrcode": "^1.5.4",
"uuid": "^10.0.0",
"web-push": "^3.6.7",
"xterm": "^5.3.0",
"xterm-addon-fit": "^0.8.0",
"xterm-addon-unicode11": "^0.6.0",
"xterm-addon-webgl": "^0.16.0",
"zod": "^4.3.6"
},
"devDependencies": {
@@ -78,6 +79,7 @@
"@types/react": "^19.2.14",
"@types/uuid": "^10.0.0",
"@types/web-push": "^3.6.4",
"@types/ws": "^8.18.1",
"@vitest/coverage-v8": "^4.0.18",
"agent-browser": "^0.6.0",
"esbuild": "^0.27.3",
@@ -93,6 +95,10 @@
"typescript-eslint": "^8.0.0",
"vitest": "^4.0.18"
},
"optionalDependencies": {
"@remotion/compositor-linux-x64-gnu": "^4.0.432",
"@rspack/binding-linux-x64-gnu": "^1.7.7"
},
"engines": {
"node": ">=18.0.0"
},
-1
View File
@@ -1 +0,0 @@
tools/remotion/public
+70 -11
View File
@@ -8,13 +8,15 @@
* 2. Copy static assets (web/public, templates)
* 3. Build vendor xterm bundles
* 4. Minify frontend assets (app.js, styles.css, mobile.css)
* 5. Compress with gzip + brotli
* 5. Content-hash cache busting (rename assets, rewrite index.html)
* 6. Compress with gzip + brotli
*/
import { execSync } from 'child_process';
import { appendFileSync } from 'fs';
import { appendFileSync, readFileSync, writeFileSync, renameSync } from 'fs';
import { createHash } from 'crypto';
import { fileURLToPath } from 'url';
import { join } from 'path';
import { join, extname, basename, dirname } from 'path';
const ROOT = join(fileURLToPath(import.meta.url), '..', '..');
@@ -27,17 +29,18 @@ function run(label, cmd) {
run('tsc', 'tsc');
run('chmod dist/index.js', 'chmod +x dist/index.js');
// 2. Copy static assets
// 2. Copy static assets (clean first to remove stale hashed files from previous builds)
run('clean public', 'rm -rf dist/web/public');
run('prepare dirs', 'mkdir -p dist/web dist/templates dist/web/public/vendor');
run('copy web assets', 'cp -r src/web/public dist/web/');
run('copy template', 'cp src/templates/case-template.md dist/templates/');
// 3. Vendor xterm bundles
run('xterm css', 'cp node_modules/xterm/css/xterm.css dist/web/public/vendor/');
run('xterm js', 'npx esbuild node_modules/xterm/lib/xterm.js --minify --outfile=dist/web/public/vendor/xterm.min.js');
run('xterm-addon-fit', 'npx esbuild node_modules/xterm-addon-fit/lib/xterm-addon-fit.js --minify --outfile=dist/web/public/vendor/xterm-addon-fit.min.js');
run('xterm-addon-webgl', 'cp node_modules/xterm-addon-webgl/lib/xterm-addon-webgl.js dist/web/public/vendor/xterm-addon-webgl.min.js');
run('xterm-addon-unicode11', 'npx esbuild node_modules/xterm-addon-unicode11/lib/xterm-addon-unicode11.js --minify --outfile=dist/web/public/vendor/xterm-addon-unicode11.min.js');
// 3. Vendor xterm bundles (xterm.js 6.x — @xterm scoped packages)
run('xterm css', 'cp node_modules/@xterm/xterm/css/xterm.css dist/web/public/vendor/');
run('xterm js', 'npx esbuild node_modules/@xterm/xterm/lib/xterm.js --minify --outfile=dist/web/public/vendor/xterm.min.js');
run('xterm-addon-fit', 'npx esbuild node_modules/@xterm/addon-fit/lib/addon-fit.js --minify --outfile=dist/web/public/vendor/xterm-addon-fit.min.js');
run('xterm-addon-webgl', 'cp node_modules/@xterm/addon-webgl/lib/addon-webgl.js dist/web/public/vendor/xterm-addon-webgl.min.js');
run('xterm-addon-unicode11', 'npx esbuild node_modules/@xterm/addon-unicode11/lib/addon-unicode11.js --minify --outfile=dist/web/public/vendor/xterm-addon-unicode11.min.js');
run('xterm-zerolag-input', 'npx esbuild packages/xterm-zerolag-input/src/zerolag-input-addon.ts --bundle --minify --format=iife --global-name=XtermZerolagInput --outfile=dist/web/public/vendor/xterm-zerolag-input.js');
// Append global aliases so app.js can use `new LocalEchoOverlay(terminal)`
@@ -56,11 +59,67 @@ appendFileSync(
);
// 4. Minify frontend assets
run('minify input-cjk.js', 'npx esbuild dist/web/public/input-cjk.js --minify --outfile=dist/web/public/input-cjk.js --allow-overwrite');
run('minify app.js', 'npx esbuild dist/web/public/app.js --minify --outfile=dist/web/public/app.js --allow-overwrite');
run('minify terminal-ui.js', 'npx esbuild dist/web/public/terminal-ui.js --minify --outfile=dist/web/public/terminal-ui.js --allow-overwrite');
run('minify respawn-ui.js', 'npx esbuild dist/web/public/respawn-ui.js --minify --outfile=dist/web/public/respawn-ui.js --allow-overwrite');
run('minify ralph-panel.js', 'npx esbuild dist/web/public/ralph-panel.js --minify --outfile=dist/web/public/ralph-panel.js --allow-overwrite');
run('minify settings-ui.js', 'npx esbuild dist/web/public/settings-ui.js --minify --outfile=dist/web/public/settings-ui.js --allow-overwrite');
run('minify panels-ui.js', 'npx esbuild dist/web/public/panels-ui.js --minify --outfile=dist/web/public/panels-ui.js --allow-overwrite');
run('minify session-ui.js', 'npx esbuild dist/web/public/session-ui.js --minify --outfile=dist/web/public/session-ui.js --allow-overwrite');
run('minify styles.css', 'npx esbuild dist/web/public/styles.css --minify --outfile=dist/web/public/styles.css --allow-overwrite');
run('minify mobile.css', 'npx esbuild dist/web/public/mobile.css --minify --outfile=dist/web/public/mobile.css --allow-overwrite');
// 5. Compress with gzip + brotli
// 5. Content-hash cache busting
console.log('\n[build] content-hash cache busting');
{
const distPublic = join(ROOT, 'dist/web/public');
const HASHABLE = [
'styles.css',
'mobile.css',
'constants.js',
'mobile-handlers.js',
'voice-input.js',
'notification-manager.js',
'keyboard-accessory.js',
'input-cjk.js',
'app.js',
'terminal-ui.js',
'respawn-ui.js',
'ralph-panel.js',
'settings-ui.js',
'panels-ui.js',
'session-ui.js',
'ralph-wizard.js',
'api-client.js',
'subagent-windows.js',
'vendor/xterm-zerolag-input.js',
];
const manifest = {};
for (const file of HASHABLE) {
const filePath = join(distPublic, file);
const content = readFileSync(filePath);
const hash = createHash('md5').update(content).digest('hex').slice(0, 8);
const ext = extname(file);
const base = basename(file, ext);
const dir = dirname(file);
const hashed = dir === '.' ? `${base}.${hash}${ext}` : `${dir}/${base}.${hash}${ext}`;
renameSync(filePath, join(distPublic, hashed));
manifest[file] = hashed;
}
// Rewrite index.html to reference hashed filenames
let html = readFileSync(join(distPublic, 'index.html'), 'utf8');
for (const [original, hashed] of Object.entries(manifest)) {
html = html.replaceAll(`"${original}"`, `"${hashed}"`);
}
writeFileSync(join(distPublic, 'index.html'), html);
console.log(' Hashed files:');
for (const [orig, hashed] of Object.entries(manifest)) {
console.log(` ${orig} -> ${hashed}`);
}
}
// 6. Compress with gzip + brotli
run(
'compress',
`for f in dist/web/public/*.js dist/web/public/*.css dist/web/public/*.html dist/web/public/vendor/*.js dist/web/public/vendor/*.css; do` +
+1 -1
View File
@@ -428,7 +428,7 @@ const SUBAGENT_ACTIVITY = {
'agent-002': [
{ type: 'tool', tool: 'Glob', input: { pattern: 'test/**/*.test.ts' }, timestamp: new Date().toISOString(), agentId: 'agent-002' },
{ type: 'tool', tool: 'Read', input: { file_path: '/home/arkon/codeman/test/respawn-test-utils.ts' }, timestamp: new Date().toISOString(), agentId: 'agent-002' },
{ type: 'tool', tool: 'Read', input: { file_path: '/home/arkon/codeman/vitest.config.ts' }, timestamp: new Date().toISOString(), agentId: 'agent-002' },
{ type: 'tool', tool: 'Read', input: { file_path: '/home/arkon/codeman/config/vitest.config.ts' }, timestamp: new Date().toISOString(), agentId: 'agent-002' },
{ type: 'message', role: 'assistant', text: 'Analyzing test patterns: MockSession, unique ports, fileParallelism: false...', timestamp: new Date().toISOString(), agentId: 'agent-002' },
],
};
+14 -14
View File
@@ -9,7 +9,7 @@
*
* Usage: node scripts/capture-video-screenshots.mjs
* Port: 3198 (static file server)
* Output: remotion/public/ (6 PNGs)
* Output: scripts/scripts/remotion/public/ (6 PNGs)
*/
import { chromium } from 'playwright';
@@ -21,7 +21,7 @@ import { fileURLToPath } from 'url';
const __dirname = fileURLToPath(new URL('.', import.meta.url));
const PROJECT_ROOT = join(__dirname, '..');
const PUBLIC_DIR = join(PROJECT_ROOT, 'src', 'web', 'public');
const OUTPUT_DIR = join(PROJECT_ROOT, 'remotion', 'public');
const OUTPUT_DIR = join(PROJECT_ROOT, 'scripts', 'remotion', 'public');
const PORT = 3198;
const DESKTOP_VIEWPORT = { width: 1920, height: 1080 };
@@ -516,7 +516,7 @@ async function captureDesktopWelcome(browser) {
path: join(OUTPUT_DIR, 'desktop-welcome.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/desktop-welcome.png');
console.log(' Saved: scripts/remotion/public/desktop-welcome.png');
} finally {
await context.close();
}
@@ -545,7 +545,7 @@ async function captureDesktopClaude(browser) {
path: join(OUTPUT_DIR, 'desktop-claude.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/desktop-claude.png');
console.log(' Saved: scripts/remotion/public/desktop-claude.png');
} finally {
await context.close();
}
@@ -575,7 +575,7 @@ async function captureDesktopBothClaude(browser) {
path: join(OUTPUT_DIR, 'desktop-both-claude.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/desktop-both-claude.png');
console.log(' Saved: scripts/remotion/public/desktop-both-claude.png');
} finally {
await context.close();
}
@@ -605,7 +605,7 @@ async function captureDesktopBothOpencode(browser) {
path: join(OUTPUT_DIR, 'desktop-both-opencode.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/desktop-both-opencode.png');
console.log(' Saved: scripts/remotion/public/desktop-both-opencode.png');
} finally {
await context.close();
}
@@ -634,7 +634,7 @@ async function captureMobileClaude(browser) {
path: join(OUTPUT_DIR, 'mobile-claude.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/mobile-claude.png');
console.log(' Saved: scripts/remotion/public/mobile-claude.png');
} finally {
await context.close();
}
@@ -663,7 +663,7 @@ async function captureMobileOpencode(browser) {
path: join(OUTPUT_DIR, 'mobile-opencode.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/mobile-opencode.png');
console.log(' Saved: scripts/remotion/public/mobile-opencode.png');
} finally {
await context.close();
}
@@ -706,12 +706,12 @@ async function main() {
console.log('All 6 screenshots captured!');
console.log('='.repeat(60));
console.log('\nOutput files:');
console.log(' remotion/public/desktop-welcome.png');
console.log(' remotion/public/desktop-claude.png');
console.log(' remotion/public/desktop-both-claude.png');
console.log(' remotion/public/desktop-both-opencode.png');
console.log(' remotion/public/mobile-claude.png');
console.log(' remotion/public/mobile-opencode.png');
console.log(' scripts/remotion/public/desktop-welcome.png');
console.log(' scripts/remotion/public/desktop-claude.png');
console.log(' scripts/remotion/public/desktop-both-claude.png');
console.log(' scripts/remotion/public/desktop-both-opencode.png');
console.log(' scripts/remotion/public/mobile-claude.png');
console.log(' scripts/remotion/public/mobile-opencode.png');
} catch (err) {
console.error('\nFatal error:', err.message);
console.error(err.stack);
+29
View File
@@ -0,0 +1,29 @@
#!/usr/bin/env node
// Fails if package-lock.json's version fields don't match package.json.
// Changesets bumps package.json but NOT the lockfile — this catches that drift
// (the top-level `version` in lockfiles is metadata, so `npm ci` won't flag it).
import { readFileSync } from 'node:fs';
import { resolve, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
const repoRoot = resolve(dirname(fileURLToPath(import.meta.url)), '..');
const pkg = JSON.parse(readFileSync(resolve(repoRoot, 'package.json'), 'utf8'));
const lock = JSON.parse(readFileSync(resolve(repoRoot, 'package-lock.json'), 'utf8'));
const expected = pkg.version;
const rootVersion = lock.version;
const selfVersion = lock.packages?.['']?.version;
const mismatches = [];
if (rootVersion !== expected) mismatches.push(` package-lock.json#.version = ${rootVersion} (expected ${expected})`);
if (selfVersion !== expected) mismatches.push(` package-lock.json#.packages[""].version = ${selfVersion} (expected ${expected})`);
if (mismatches.length > 0) {
console.error(`\nLockfile version drift detected (package.json is ${expected}):`);
console.error(mismatches.join('\n'));
console.error('\nFix: run `npm install --package-lock-only` and commit the updated package-lock.json.\n');
process.exit(1);
}
console.log(`Lockfile in sync with package.json (${expected}).`);
+19
View File
@@ -0,0 +1,19 @@
[Unit]
Description=Codeman Cloudflare Named Tunnel
After=network-online.target codeman-web.service
Wants=network-online.target
[Service]
Type=simple
ExecStart=/usr/bin/cloudflared tunnel --config %h/.cloudflared/codeman.yml run codeman
Restart=always
RestartSec=5
KillMode=process
# Logging
StandardOutput=journal
StandardError=journal
SyslogIdentifier=codeman-tunnel-named
[Install]
WantedBy=default.target
+1
View File
@@ -11,6 +11,7 @@ RestartSec=5
KillMode=process
Environment=NODE_ENV=production
Environment=HOME=/home/arkon
Environment=NODE_COMPILE_CACHE=/home/arkon/.codeman/compile-cache
# Logging
StandardOutput=journal
+9 -9
View File
@@ -250,10 +250,10 @@ if (isGlobalInstall) {
} else {
try {
const require = createRequire(import.meta.url);
const xtermDir = join(require.resolve('xterm'), '..', '..');
const fitDir = join(require.resolve('xterm-addon-fit'), '..', '..');
const webglDir = join(require.resolve('xterm-addon-webgl'), '..', '..');
const unicode11Dir = join(require.resolve('xterm-addon-unicode11'), '..', '..');
const xtermDir = join(require.resolve('@xterm/xterm'), '..', '..');
const fitDir = join(require.resolve('@xterm/addon-fit'), '..', '..');
const webglDir = join(require.resolve('@xterm/addon-webgl'), '..', '..');
const unicode11Dir = join(require.resolve('@xterm/addon-unicode11'), '..', '..');
const vendorDir = join(srcDir, 'web', 'public', 'vendor');
const { mkdirSync, copyFileSync } = await import('fs');
@@ -263,19 +263,19 @@ if (isGlobalInstall) {
// Minify xterm JS for dev vendor dir (npm packages don't ship .min.js)
try {
execSync(`npx esbuild "${join(xtermDir, 'lib', 'xterm.js')}" --minify --outfile="${join(vendorDir, 'xterm.min.js')}"`, { stdio: 'pipe' });
execSync(`npx esbuild "${join(fitDir, 'lib', 'xterm-addon-fit.js')}" --minify --outfile="${join(vendorDir, 'xterm-addon-fit.min.js')}"`, { stdio: 'pipe' });
execSync(`npx esbuild "${join(unicode11Dir, 'lib', 'xterm-addon-unicode11.js')}" --minify --outfile="${join(vendorDir, 'xterm-addon-unicode11.min.js')}"`, { stdio: 'pipe' });
execSync(`npx esbuild "${join(fitDir, 'lib', 'addon-fit.js')}" --minify --outfile="${join(vendorDir, 'xterm-addon-fit.min.js')}"`, { stdio: 'pipe' });
execSync(`npx esbuild "${join(unicode11Dir, 'lib', 'addon-unicode11.js')}" --minify --outfile="${join(vendorDir, 'xterm-addon-unicode11.min.js')}"`, { stdio: 'pipe' });
console.log(colors.green('✓ xterm vendor files copied to src/web/public/vendor/'));
} catch {
// Fallback: copy unminified
copyFileSync(join(xtermDir, 'lib', 'xterm.js'), join(vendorDir, 'xterm.min.js'));
copyFileSync(join(fitDir, 'lib', 'xterm-addon-fit.js'), join(vendorDir, 'xterm-addon-fit.min.js'));
copyFileSync(join(unicode11Dir, 'lib', 'xterm-addon-unicode11.js'), join(vendorDir, 'xterm-addon-unicode11.min.js'));
copyFileSync(join(fitDir, 'lib', 'addon-fit.js'), join(vendorDir, 'xterm-addon-fit.min.js'));
copyFileSync(join(unicode11Dir, 'lib', 'addon-unicode11.js'), join(vendorDir, 'xterm-addon-unicode11.min.js'));
console.log(colors.green('✓ xterm vendor files copied') + colors.dim(' (unminified — esbuild not available)'));
}
// WebGL addon: copy unminified (matches build script behavior)
copyFileSync(join(webglDir, 'lib', 'xterm-addon-webgl.js'), join(vendorDir, 'xterm-addon-webgl.min.js'));
copyFileSync(join(webglDir, 'lib', 'addon-webgl.js'), join(vendorDir, 'xterm-addon-webgl.min.js'));
// xterm-zerolag-input: bundle local package as IIFE for <script> tag loading
try {
@@ -16,69 +16,54 @@ import { IOSKeyboard } from '../components/IOSKeyboard';
// ─── Scene timing (frames @ 30fps) ───
const TITLE_DUR = 60;
const PHONES_DUR = 30;
const TYPING_DUR = 468;
const HOLD_DUR = 60;
const TYPING_DUR = 610;
const HOLD_DUR = 50;
const OUTRO_DUR = 45;
const TITLE_START = 0;
const PHONES_START = TITLE_DUR; // 60
const TYPING_START = PHONES_START + PHONES_DUR; // 90
const HOLD_START = TYPING_START + TYPING_DUR; // 558
const OUTRO_START = HOLD_START + HOLD_DUR; // 618
const HOLD_START = TYPING_START + TYPING_DUR; // 700
const OUTRO_START = HOLD_START + HOLD_DUR; // 750
export const ZEROLAG_TOTAL_FRAMES = OUTRO_START + OUTRO_DUR; // 663
export const ZEROLAG_TOTAL_FRAMES = OUTRO_START + OUTRO_DUR; // 795
// ─── iPhone 17 Pro safe area ───
const SAFE_AREA_TOP = 59; // Below Dynamic Island
const PHONE_SCALE = 1.12; // Scale up to fill more of the frame
// The Codeman screenshot starts content at y=0 (the session tab).
// On a real device it would sit below the safe area, so we offset it.
const SCREENSHOT_Y_OFFSET = SAFE_AREA_TOP;
// Claude Code header: tab bar + session info + prompt context from screenshot
const HEADER_H = 120;
// Terminal typing overlay position (relative to screenshot top)
// Session tab is ~44px, then terminal starts. Adding safe area offset:
const TERMINAL_TOP = SAFE_AREA_TOP + 52;
// Terminal typing overlay: aligned with the ❯ prompt position in the Claude Code screenshot
const TERMINAL_TOP = 185;
const TERMINAL_LEFT = 14;
const TERMINAL_FONT = 22; // Large for video readability
const TERMINAL_FONT = 21; // Slightly smaller to fit toolbar below
// ─── Typing schedule (with typo + backspace correction) ───
const CORRECT_TEXT = 'fix the auth bug in the login flow';
const FRAME_GAP = 12; // ~400ms between keystrokes
const TYPO_INDEX = 28; // After "logi", type "m" instead of "n"
// Codeman toolbar from screenshot (bottom section showing /init, /clear, Run, etc.)
const TOOLBAR_H = 95;
// Remote connection lag: 600ms-1.2s+ per char (18-36+ frames)
// ─── Typing schedule ───
const CORRECT_TEXT =
'zerolag technology brings in a visual dom overlay to make typing instant, even if your codeman server is on the other side of the world';
const FRAME_GAP = 4; // ~133ms per keystroke (~75 WPM)
// Remote connection lag: 600ms–2.7s per char (18–80 frames @ 30fps).
// Periodic spikes simulate packet loss / retransmission bursts.
// TCP head-of-line blocking causes a single spike to freeze all subsequent chars.
const LAGGY_DELAYS = [
24, 30, 36, 32, 26, 22, 34, 28, 38, 20, 30, 24, 32, 26, 36, 22,
30, 24, 32, 28, 34, 26, 30, 22, 28, 36, 24, 30, 32, 26, 34, 28, 24, 30,
26, 32, 28, 34,
24, 30, 26, 32, 72, 28, 22, 34, 26, 30, 20, 28, 36, 24, 30, 22, 26, 34, 28, 20, 68, 30, 24, 32,
26, 22, 28, 34, 30, 26, 32, 24, 80, 22, 30, 26, 28, 34, 24, 30,
];
type KeyAction = { frame: number; action: 'type' | 'backspace'; char: string; lagDelay: number };
type KeyAction = { frame: number; char: string; lagDelay: number };
const buildSchedule = (): KeyAction[] => {
const actions: KeyAction[] = [];
let idx = 0;
const lag = (i: number) => LAGGY_DELAYS[i % LAGGY_DELAYS.length];
// Type correctly up to typo point: "fix the auth bug in the logi"
for (let i = 0; i < TYPO_INDEX; i++) {
actions.push({ frame: idx * FRAME_GAP, action: 'type', char: CORRECT_TEXT[i], lagDelay: lag(idx) });
idx++;
}
// Typo: type "m" instead of "n"
actions.push({ frame: idx * FRAME_GAP, action: 'type', char: 'm', lagDelay: lag(idx) });
idx++;
// Backspace to fix it
actions.push({ frame: idx * FRAME_GAP, action: 'backspace', char: '⌫', lagDelay: lag(idx) });
idx++;
// Type correct remaining: "n flow"
for (let i = TYPO_INDEX; i < CORRECT_TEXT.length; i++) {
actions.push({ frame: idx * FRAME_GAP, action: 'type', char: CORRECT_TEXT[i], lagDelay: lag(idx) });
idx++;
for (let i = 0; i < CORRECT_TEXT.length; i++) {
actions.push({ frame: i * FRAME_GAP, char: CORRECT_TEXT[i], lagDelay: lag(i) });
}
return actions;
@@ -86,14 +71,16 @@ const buildSchedule = (): KeyAction[] => {
const TYPING_SCHEDULE = buildSchedule();
/** Replay actions in order up to current frame, computing the visible text buffer */
/**
* Replay actions in order up to current frame, computing the visible text buffer.
* TCP-ordered: stops at first unresolved echo (head-of-line blocking).
*/
const computeVisibleText = (frame: number, withLag: boolean): string => {
let buffer = '';
for (const a of TYPING_SCHEDULE) {
const threshold = withLag ? a.frame + a.lagDelay : a.frame;
if (frame < threshold) break; // TCP-ordered: stop at first unresolved
if (a.action === 'backspace') buffer = buffer.slice(0, -1);
else buffer += a.char;
if (frame < threshold) break;
buffer += a.char;
}
return buffer;
};
@@ -128,8 +115,20 @@ const IOSStatusBar: React.FC = () => (
</svg>
{/* WiFi */}
<svg width="16" height="12" viewBox="0 0 16 12">
<path d="M4.5 8.5C5.5 7.2 6.7 6.5 8 6.5s2.5.7 3.5 2" stroke="#fff" strokeWidth="1.5" fill="none" strokeLinecap="round" />
<path d="M1.5 5.5C3.5 3 5.7 1.5 8 1.5s4.5 1.5 6.5 4" stroke="#fff" strokeWidth="1.5" fill="none" strokeLinecap="round" />
<path
d="M4.5 8.5C5.5 7.2 6.7 6.5 8 6.5s2.5.7 3.5 2"
stroke="#fff"
strokeWidth="1.5"
fill="none"
strokeLinecap="round"
/>
<path
d="M1.5 5.5C3.5 3 5.7 1.5 8 1.5s4.5 1.5 6.5 4"
stroke="#fff"
strokeWidth="1.5"
fill="none"
strokeLinecap="round"
/>
<circle cx="8" cy="11" r="1.5" fill="#fff" />
</svg>
{/* Battery */}
@@ -145,9 +144,8 @@ const IOSStatusBar: React.FC = () => (
// ─── Typing overlay ───
const TypingOverlay: React.FC<{
typed: string;
overlayChars?: { char: string; confirmed: boolean }[];
cursorVisible: boolean;
}> = ({ typed, overlayChars, cursorVisible }) => {
}> = ({ typed, cursorVisible }) => {
const frame = useCurrentFrame();
const cursorOpacity = cursorVisible ? (Math.floor(frame / 18) % 2 === 0 ? 0.85 : 0.5) : 0;
const lineH = Math.round(TERMINAL_FONT * 1.4);
@@ -167,13 +165,7 @@ const TypingOverlay: React.FC<{
}}
>
<span style={{ color: '#339af0', fontWeight: 700 }}>{'❯ '}</span>
{overlayChars
? overlayChars.map((oc, i) => (
<span key={i} style={{ color: oc.confirmed ? '#e0e0e0' : '#666' }}>
{oc.char}
</span>
))
: <span style={{ color: '#e0e0e0' }}>{typed}</span>}
<span style={{ color: '#e0e0e0' }}>{typed}</span>
<span
style={{
display: 'inline-block',
@@ -188,37 +180,116 @@ const TypingOverlay: React.FC<{
);
};
// ─── Single phone: iPhone 17 Pro + real screenshot + typing + keyboard ───
// ─── Single phone: iPhone 17 Pro + Claude Code header + typing + keyboard ───
const MobileCodeman: React.FC<{
typed: string;
overlayChars?: { char: string; confirmed: boolean }[];
cursorVisible: boolean;
activeKey?: string;
pressAge?: number;
showKeyboard?: boolean;
noAnimation?: boolean;
}> = ({ typed, overlayChars, cursorVisible, activeKey, pressAge, showKeyboard = true, noAnimation }) => (
showPromo?: boolean;
}> = ({ typed, cursorVisible, activeKey, pressAge, showKeyboard = true, noAnimation, showPromo }) => (
<IPhone17ProFrame noAnimation={noAnimation}>
<div style={{ width: SCREEN_W, height: SCREEN_H, position: 'relative', overflow: 'hidden', background: '#000' }}>
{/* Real Codeman mobile screenshot, pushed down by safe area */}
<div
style={{ width: SCREEN_W, height: SCREEN_H, position: 'relative', overflow: 'hidden', background: '#0d0d0d' }}
>
{/* iOS status bar in the safe area */}
<IOSStatusBar />
{/* Claude Code header from screenshot (tabs + session info) */}
<div
style={{
position: 'absolute',
top: SAFE_AREA_TOP,
left: 0,
right: 0,
height: HEADER_H,
overflow: 'hidden',
zIndex: 5,
}}
>
<Img
src={staticFile('mobile-claude.png')}
style={{
width: SCREEN_W,
height: SCREEN_H - SCREENSHOT_Y_OFFSET,
objectFit: 'cover',
objectPosition: 'top',
position: 'absolute',
top: SCREENSHOT_Y_OFFSET,
left: 0,
objectPosition: 'top left',
}}
/>
{/* Fade to terminal background */}
<div
style={{
position: 'absolute',
bottom: 0,
left: 0,
right: 0,
height: 30,
background: 'linear-gradient(transparent, #0d0d0d)',
}}
/>
</div>
{/* iOS status bar in the safe area */}
<IOSStatusBar />
{/* Dark mask: hides screenshot text below Claude Code info (e.g. "Try edit...") */}
<div
style={{
position: 'absolute',
top: SAFE_AREA_TOP + 95,
left: 0,
right: 0,
bottom: 0,
zIndex: 8,
}}
>
{/* Smooth gradient blend from screenshot into dark terminal */}
<div
style={{
height: 18,
background: 'linear-gradient(transparent, #0d0d0d)',
}}
/>
<div style={{ flex: 1, background: '#0d0d0d' }} />
</div>
{/* Typing animation */}
<TypingOverlay typed={typed} overlayChars={overlayChars} cursorVisible={cursorVisible} />
<TypingOverlay typed={typed} cursorVisible={cursorVisible} />
{/* Codeman toolbar from screenshot (/init, /clear, /compact, Run, Run Shell, voice) */}
<div
style={{
position: 'absolute',
bottom: 232,
left: 0,
right: 0,
height: TOOLBAR_H,
overflow: 'hidden',
zIndex: 15,
}}
>
{/* Gradient blend at top */}
<div
style={{
position: 'absolute',
top: 0,
left: 0,
right: 0,
height: 18,
background: 'linear-gradient(#0d0d0d, transparent)',
zIndex: 1,
}}
/>
<Img
src={staticFile('mobile-claude.png')}
style={{
width: SCREEN_W,
position: 'absolute',
bottom: 0,
}}
/>
</div>
{/* Promo cards between text and toolbar (left phone only) */}
{showPromo && <PromoBanners />}
{/* iOS keyboard at bottom */}
{showKeyboard && (
@@ -230,22 +301,22 @@ const MobileCodeman: React.FC<{
</IPhone17ProFrame>
);
// ─── Label beneath phone ───
// ─── Label above phone (prominent header) ───
const PhoneLabel: React.FC<{
title: string;
detail: string;
dotColor: string;
detailColor: string;
}> = ({ title, detail, dotColor, detailColor }) => (
<div style={{ textAlign: 'center', marginTop: 20 }}>
<div style={{ textAlign: 'center', marginBottom: 16 }}>
<div
style={{
display: 'flex',
alignItems: 'center',
justifyContent: 'center',
gap: 10,
fontSize: 24,
fontWeight: 600,
fontSize: 28,
fontWeight: 700,
fontFamily: fonts.ui,
color: '#fff',
}}
@@ -267,6 +338,62 @@ const PhoneLabel: React.FC<{
</div>
);
// ─── Promo banners (inside left phone, between text and keyboard) ───
const PromoBanners: React.FC = () => (
<div
style={{
position: 'absolute',
bottom: 332,
left: 8,
right: 8,
display: 'flex',
flexDirection: 'column',
gap: 8,
zIndex: 15,
}}
>
<div
style={{
display: 'flex',
alignItems: 'center',
gap: 12,
padding: '12px 14px',
background: 'rgba(255,255,255,0.06)',
borderRadius: 12,
border: '1px solid rgba(255,255,255,0.1)',
}}
>
<svg width="30" height="30" viewBox="0 0 16 16" fill="#ccc" style={{ flexShrink: 0 }}>
<path d="M8 0C3.58 0 0 3.58 0 8c0 3.54 2.29 6.53 5.47 7.59.4.07.55-.17.55-.38 0-.19-.01-.82-.01-1.49-2.01.37-2.53-.49-2.69-.94-.09-.23-.48-.94-.82-1.13-.28-.15-.68-.52-.01-.53.63-.01 1.08.58 1.23.82.72 1.21 1.87.87 2.33.66.07-.52.28-.87.51-1.07-1.78-.2-3.64-.89-3.64-3.95 0-.87.31-1.59.82-2.15-.08-.2-.36-1.02.08-2.12 0 0 .67-.21 2.2.82.64-.18 1.32-.27 2-.27.68 0 1.36.09 2 .27 1.53-1.04 2.2-.82 2.2-.82.44 1.1.16 1.92.08 2.12.51.56.82 1.27.82 2.15 0 3.07-1.87 3.75-3.65 3.95.29.25.54.73.54 1.48 0 1.07-.01 1.93-.01 2.2 0 .21.15.46.55.38A8.013 8.013 0 0016 8c0-4.42-3.58-8-8-8z" />
</svg>
<div>
<div style={{ fontSize: 18, fontWeight: 700, fontFamily: fonts.ui, color: '#e0e0e0' }}>Codeman</div>
<div style={{ fontSize: 12, fontFamily: fonts.mono, color: '#888' }}>github.com/Ark0N/Codeman</div>
</div>
</div>
<div
style={{
display: 'flex',
alignItems: 'center',
gap: 12,
padding: '12px 14px',
background: 'rgba(255,255,255,0.06)',
borderRadius: 12,
border: '1px solid rgba(255,255,255,0.1)',
}}
>
<svg width="30" height="30" viewBox="0 0 16 16" style={{ flexShrink: 0 }}>
<rect width="16" height="16" rx="2" fill="#cb3837" />
<path d="M3 3h10v10H8V5H5v8H3z" fill="#fff" />
</svg>
<div>
<div style={{ fontSize: 18, fontWeight: 700, fontFamily: fonts.ui, color: '#e0e0e0' }}>xterm-zerolag-input</div>
<div style={{ fontSize: 12, fontFamily: fonts.mono, color: '#888' }}>npmjs.com/package/xterm-zerolag-input</div>
</div>
</div>
</div>
);
// ─── Title scene ───
const TitleScene: React.FC = () => {
const frame = useCurrentFrame();
@@ -357,12 +484,22 @@ const TypingDemo: React.FC = () => {
}}
>
<div>
<MobileCodeman typed={zerolagTyped} cursorVisible activeKey={activeKey} pressAge={pressAge} noAnimation />
<PhoneLabel title="With Zerolag" detail="0ms delay" dotColor={colors.accent.green} detailColor={colors.accent.green} />
<PhoneLabel
title="With Zerolag"
detail="0ms local echo"
dotColor={colors.accent.green}
detailColor={colors.accent.green}
/>
<MobileCodeman typed={zerolagTyped} cursorVisible activeKey={activeKey} pressAge={pressAge} noAnimation showPromo />
</div>
<div>
<PhoneLabel
title="Without Zerolag"
detail="600ms–2.7s server echo"
dotColor={colors.accent.red}
detailColor={colors.accent.red}
/>
<MobileCodeman typed={laggyTyped} cursorVisible activeKey={activeKey} pressAge={pressAge} noAnimation />
<PhoneLabel title="Without Zerolag" detail="600ms–1.2s delay" dotColor={colors.accent.red} detailColor={colors.accent.red} />
</div>
</div>
</AbsoluteFill>
@@ -389,12 +526,22 @@ const PanelsEntrance: React.FC = () => {
>
<div style={{ display: 'flex', gap: 50, alignItems: 'flex-start' }}>
<div>
<PhoneLabel
title="With Zerolag"
detail="0ms local echo"
dotColor={colors.accent.green}
detailColor={colors.accent.green}
/>
<MobileCodeman typed="" cursorVisible noAnimation />
<PhoneLabel title="With Zerolag" detail="0ms delay" dotColor={colors.accent.green} detailColor={colors.accent.green} />
</div>
<div>
<PhoneLabel
title="Without Zerolag"
detail="600ms–2.7s server echo"
dotColor={colors.accent.red}
detailColor={colors.accent.red}
/>
<MobileCodeman typed="" cursorVisible noAnimation />
<PhoneLabel title="Without Zerolag" detail="600ms–1.2s delay" dotColor={colors.accent.red} detailColor={colors.accent.red} />
</div>
</div>
</AbsoluteFill>

Before

Width:  |  Height:  |  Size: 41 KiB

After

Width:  |  Height:  |  Size: 41 KiB

Before

Width:  |  Height:  |  Size: 37 KiB

After

Width:  |  Height:  |  Size: 37 KiB

Before

Width:  |  Height:  |  Size: 39 KiB

After

Width:  |  Height:  |  Size: 39 KiB

Before

Width:  |  Height:  |  Size: 57 KiB

After

Width:  |  Height:  |  Size: 57 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 390 KiB

Before

Width:  |  Height:  |  Size: 22 KiB

After

Width:  |  Height:  |  Size: 22 KiB

+192 -28
View File
@@ -1,45 +1,209 @@
#!/usr/bin/env bash
# Quick Cloudflare Tunnel for Codeman
# Usage: ./scripts/tunnel.sh [start|stop|status|url]
# Cloudflare Tunnel manager for Codeman
# Usage: ./scripts/tunnel.sh [quick|named] [start|stop|status|url]
#
# Modes:
# quick — Quick tunnel with random trycloudflare.com URL (default)
# named — Named tunnel on a fixed hostname (requires setup, see below)
#
# Environment variables:
# CLOUDFLARED_TUNNEL_NAME — tunnel name (default: codeman)
# CLOUDFLARED_TUNNEL_ID — tunnel UUID (from: cloudflared tunnel list)
# CODEMAN_TUNNEL_HOSTNAME — public hostname (e.g. codeman.example.com)
#
# First-time named tunnel setup:
# cloudflared tunnel login
# cloudflared tunnel create <tunnel-name>
# cloudflared tunnel route dns <tunnel-name> <hostname>
# ./scripts/tunnel.sh named setup # writes ~/.cloudflared/<tunnel-name>.yml
set -euo pipefail
SERVICE="codeman-tunnel"
QUICK_SERVICE="codeman-tunnel"
NAMED_SERVICE="codeman-tunnel-named"
TUNNEL_NAME="${CLOUDFLARED_TUNNEL_NAME:-codeman}"
TUNNEL_HOSTNAME="${CODEMAN_TUNNEL_HOSTNAME:-codeman.example.com}"
CODEMAN_PORT="3000"
LOG_FILE="$HOME/.codeman/tunnel.log"
case "${1:-start}" in
start)
if ! systemctl --user is-active "$SERVICE" &>/dev/null; then
# Install service if not already
if ! systemctl --user cat "$SERVICE" &>/dev/null 2>&1; then
cp "$(dirname "$0")/codeman-tunnel.service" "$HOME/.config/systemd/user/"
systemctl --user daemon-reload
# ── helpers ──────────────────────────────────────────────────────────────────
_require_cloudflared() {
if ! command -v cloudflared &>/dev/null; then
echo "Error: cloudflared not found. Install with: yay -S cloudflared" >&2
exit 1
fi
systemctl --user start "$SERVICE"
echo "Tunnel starting... waiting for URL"
}
_cloudflared_bin() {
command -v cloudflared
}
_install_service() {
local svc_file="$1"
local svc_name="$2"
if ! systemctl --user cat "$svc_name" &>/dev/null 2>&1; then
cp "$(dirname "$0")/$svc_file" "$HOME/.config/systemd/user/"
systemctl --user daemon-reload
echo "Service $svc_name installed."
fi
}
_install_named_service() {
if ! systemctl --user cat "$NAMED_SERVICE" &>/dev/null 2>&1; then
# Generate service file with the configured tunnel name
sed "s/codeman\.yml/$TUNNEL_NAME.yml/g; s/run codeman/run $TUNNEL_NAME/g" \
"$(dirname "$0")/codeman-tunnel-named.service" \
> "$HOME/.config/systemd/user/codeman-tunnel-named.service"
systemctl --user daemon-reload
echo "Service $NAMED_SERVICE installed (tunnel: $TUNNEL_NAME)."
fi
}
# ── named tunnel setup ───────────────────────────────────────────────────────
_named_setup() {
_require_cloudflared
local creds_dir="$HOME/.cloudflared"
local config_file="$creds_dir/$TUNNEL_NAME.yml"
# Replace with your tunnel ID (from: cloudflared tunnel list)
local tunnel_id="${CLOUDFLARED_TUNNEL_ID:-YOUR_TUNNEL_ID_HERE}"
local creds_file="$creds_dir/$tunnel_id.json"
if [ ! -f "$creds_file" ]; then
echo "Credentials not found: $creds_file"
echo "Run: cloudflared tunnel create $TUNNEL_NAME"
exit 1
fi
cat > "$config_file" <<EOF
tunnel: $tunnel_id
credentials-file: $creds_file
ingress:
- hostname: $TUNNEL_HOSTNAME
service: http://localhost:$CODEMAN_PORT
- service: http_status:404
EOF
echo "Config written to $config_file"
echo "Tunnel ID: $tunnel_id"
echo "Hostname: $TUNNEL_HOSTNAME"
echo ""
echo "Next steps:"
echo " 1. Add Cloudflare Access policy for $TUNNEL_HOSTNAME (Zero Trust dashboard)"
echo " 2. ./scripts/tunnel.sh named start"
}
# ── quick mode ───────────────────────────────────────────────────────────────
_quick_start() {
if ! systemctl --user is-active "$QUICK_SERVICE" &>/dev/null; then
_install_service "codeman-tunnel.service" "$QUICK_SERVICE"
systemctl --user start "$QUICK_SERVICE"
echo "Quick tunnel starting... waiting for URL"
sleep 6
fi
# Extract the tunnel URL from journal
URL=$(grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$HOME/.codeman/tunnel.log" 2>/dev/null | tail -1)
if [ -n "$URL" ]; then
echo "$URL"
local url
url=$(grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$LOG_FILE" 2>/dev/null | tail -1)
if [ -n "$url" ]; then
echo "$url"
else
echo "URL not ready yet, try: $0 url"
echo "URL not ready yet, try: $0 quick url"
fi
;;
stop)
systemctl --user stop "$SERVICE"
echo "Tunnel stopped"
;;
status)
systemctl --user status "$SERVICE" --no-pager 2>&1 | head -10
}
_quick_stop() {
systemctl --user stop "$QUICK_SERVICE"
echo "Quick tunnel stopped"
}
_quick_status() {
systemctl --user status "$QUICK_SERVICE" --no-pager 2>&1 | head -10
echo ""
echo "URL:"
grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$HOME/.codeman/tunnel.log" 2>/dev/null | tail -1
grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$LOG_FILE" 2>/dev/null | tail -1
}
_quick_url() {
grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$LOG_FILE" 2>/dev/null | tail -1
}
# ── named mode ───────────────────────────────────────────────────────────────
_named_start() {
_require_cloudflared
if [ ! -f "$HOME/.cloudflared/$TUNNEL_NAME.yml" ]; then
echo "Config not found. Run: $0 named setup"
exit 1
fi
if ! systemctl --user is-active "$NAMED_SERVICE" &>/dev/null; then
_install_named_service
systemctl --user start "$NAMED_SERVICE"
echo "Named tunnel starting..."
sleep 3
fi
echo "https://$TUNNEL_HOSTNAME"
}
_named_stop() {
systemctl --user stop "$NAMED_SERVICE"
echo "Named tunnel stopped"
}
_named_status() {
systemctl --user status "$NAMED_SERVICE" --no-pager 2>&1 | head -10
echo ""
echo "URL: https://$TUNNEL_HOSTNAME"
}
_named_enable() {
_install_named_service
systemctl --user enable "$NAMED_SERVICE"
echo "Named tunnel enabled at boot."
}
_named_disable() {
systemctl --user disable "$NAMED_SERVICE"
echo "Named tunnel disabled."
}
# ── dispatch ─────────────────────────────────────────────────────────────────
MODE="${1:-quick}"
CMD="${2:-start}"
case "$MODE" in
quick)
case "$CMD" in
start) _quick_start ;;
stop) _quick_stop ;;
status) _quick_status ;;
url) _quick_url ;;
*) echo "Usage: $0 quick [start|stop|status|url]"; exit 1 ;;
esac
;;
url)
grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$HOME/.codeman/tunnel.log" 2>/dev/null | tail -1
named)
case "$CMD" in
start) _named_start ;;
stop) _named_stop ;;
status) _named_status ;;
url) echo "https://$TUNNEL_HOSTNAME" ;;
setup) _named_setup ;;
enable) _named_enable ;;
disable) _named_disable ;;
*) echo "Usage: $0 named [start|stop|status|url|setup|enable|disable]"; exit 1 ;;
esac
;;
# backward compat: no mode prefix → quick tunnel
start) _quick_start ;;
stop) _quick_stop ;;
status) _quick_status ;;
url) _quick_url ;;
*)
echo "Usage: $0 [start|stop|status|url]"
echo "Usage: $0 [quick|named] [start|stop|status|url]"
echo " $0 named setup # first-time named tunnel configuration"
echo " $0 named enable # start at boot"
exit 1
;;
esac
+6 -10
View File
@@ -28,8 +28,9 @@ import { existsSync, readFileSync, unlinkSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { EventEmitter } from 'node:events';
import { getAugmentedPath } from './utils/claude-cli-resolver.js';
import { ANSI_ESCAPE_PATTERN_SIMPLE } from './utils/index.js';
import { getAugmentedPath, ANSI_ESCAPE_PATTERN_SIMPLE } from './utils/index.js';
import { AI_CHECK_MAX_BACKOFF_MS } from './config/ai-defaults.js';
import { getErrorMessage } from './types.js';
// ========== Security Validation ==========
@@ -293,7 +294,7 @@ export abstract class AiCheckerBase<
this.emit('checkCompleted', result);
return result;
} catch (err) {
const errorMsg = err instanceof Error ? err.message : String(err);
const errorMsg = getErrorMessage(err);
this.handleError(errorMsg);
const result = this.createErrorResult(errorMsg, Date.now() - this.checkStartTime);
this.emit('checkFailed', errorMsg);
@@ -412,9 +413,7 @@ export abstract class AiCheckerBase<
});
muxProcess.unref();
} catch (err) {
throw new Error(
`Failed to spawn ${this.checkDescription} tmux session: ${err instanceof Error ? err.message : String(err)}`
);
throw new Error(`Failed to spawn ${this.checkDescription} tmux session: ${getErrorMessage(err)}`);
}
// Poll the temp file for completion
@@ -534,10 +533,7 @@ export abstract class AiCheckerBase<
// P1-005: Exponential backoff for errors
// Base cooldown * 2^(consecutiveErrors-1), capped at 5 minutes
const backoffMultiplier = Math.pow(2, this.consecutiveErrors - 1);
const backoffCooldownMs = Math.min(
this.config.errorCooldownMs * backoffMultiplier,
5 * 60 * 1000 // Max 5 minutes
);
const backoffCooldownMs = Math.min(this.config.errorCooldownMs * backoffMultiplier, AI_CHECK_MAX_BACKOFF_MS);
this.log(`Exponential backoff: ${Math.round(backoffCooldownMs / 1000)}s (error #${this.consecutiveErrors})`);
this.startCooldown(backoffCooldownMs);
}
+13 -6
View File
@@ -30,11 +30,18 @@ import {
type AiCheckerResultBase,
type AiCheckerStateBase,
} from './ai-checker-base.js';
import { AI_CHECK_MODEL, AI_IDLE_CHECK_MAX_CONTEXT } from './config/ai-defaults.js';
import {
AI_CHECK_MODEL,
AI_IDLE_CHECK_MAX_CONTEXT,
AI_IDLE_CHECK_TIMEOUT_MS,
AI_IDLE_CHECK_COOLDOWN_MS,
AI_IDLE_CHECK_ERROR_COOLDOWN_MS,
AI_CHECK_MAX_CONSECUTIVE_ERRORS,
} from './config/ai-defaults.js';
// ========== Types ==========
export type AiIdleCheckConfig = AiCheckerConfigBase;
type AiIdleCheckConfig = AiCheckerConfigBase;
export type AiCheckVerdict = 'IDLE' | 'WORKING' | 'ERROR';
@@ -48,10 +55,10 @@ const DEFAULT_AI_CHECK_CONFIG: AiIdleCheckConfig = {
enabled: true,
model: AI_CHECK_MODEL,
maxContextChars: AI_IDLE_CHECK_MAX_CONTEXT,
checkTimeoutMs: 90000,
cooldownMs: 180000,
errorCooldownMs: 60000,
maxConsecutiveErrors: 3,
checkTimeoutMs: AI_IDLE_CHECK_TIMEOUT_MS,
cooldownMs: AI_IDLE_CHECK_COOLDOWN_MS,
errorCooldownMs: AI_IDLE_CHECK_ERROR_COOLDOWN_MS,
maxConsecutiveErrors: AI_CHECK_MAX_CONSECUTIVE_ERRORS,
};
/** Pattern to match IDLE or WORKING as the first word of output */
+14 -7
View File
@@ -29,17 +29,24 @@ import {
type AiCheckerResultBase,
type AiCheckerStateBase,
} from './ai-checker-base.js';
import { AI_CHECK_MODEL, AI_PLAN_CHECK_MAX_CONTEXT } from './config/ai-defaults.js';
import {
AI_CHECK_MODEL,
AI_PLAN_CHECK_MAX_CONTEXT,
AI_PLAN_CHECK_TIMEOUT_MS,
AI_PLAN_CHECK_COOLDOWN_MS,
AI_PLAN_CHECK_ERROR_COOLDOWN_MS,
AI_CHECK_MAX_CONSECUTIVE_ERRORS,
} from './config/ai-defaults.js';
// ========== Types ==========
export type AiPlanCheckConfig = AiCheckerConfigBase;
type AiPlanCheckConfig = AiCheckerConfigBase;
export type AiPlanCheckVerdict = 'PLAN_MODE' | 'NOT_PLAN_MODE' | 'ERROR';
export type AiPlanCheckResult = AiCheckerResultBase<AiPlanCheckVerdict>;
export type AiPlanCheckState = AiCheckerStateBase<AiPlanCheckVerdict>;
type AiPlanCheckState = AiCheckerStateBase<AiPlanCheckVerdict>;
// ========== Constants ==========
@@ -47,10 +54,10 @@ const DEFAULT_PLAN_CHECK_CONFIG: AiPlanCheckConfig = {
enabled: true,
model: AI_CHECK_MODEL,
maxContextChars: AI_PLAN_CHECK_MAX_CONTEXT,
checkTimeoutMs: 60000,
cooldownMs: 30000,
errorCooldownMs: 30000,
maxConsecutiveErrors: 3,
checkTimeoutMs: AI_PLAN_CHECK_TIMEOUT_MS,
cooldownMs: AI_PLAN_CHECK_COOLDOWN_MS,
errorCooldownMs: AI_PLAN_CHECK_ERROR_COOLDOWN_MS,
maxConsecutiveErrors: AI_CHECK_MAX_CONSECUTIVE_ERRORS,
};
/** Pattern to match PLAN_MODE or NOT_PLAN_MODE as the first word(s) of output */
+70 -88
View File
@@ -15,7 +15,7 @@
import { EventEmitter } from 'node:events';
import { v4 as uuidv4 } from 'uuid';
import { ActiveBashTool } from './types.js';
import { CleanupManager, Debouncer } from './utils/index.js';
import { CleanupManager, Debouncer, stripAnsi } from './utils/index.js';
// ========== Configuration Constants ==========
@@ -99,7 +99,7 @@ const LOG_FILE_MENTION_PATTERN = /([/~][^\s'"<>|;&\n]*(?:\.log|\.txt|\.out|\/log
/**
* Events emitted by BashToolParser.
*/
export interface BashToolParserEvents {
interface BashToolParserEvents {
/** New Bash tool with file paths started */
toolStart: [tool: ActiveBashTool];
/** Bash tool completed */
@@ -111,7 +111,7 @@ export interface BashToolParserEvents {
/**
* Configuration options for BashToolParser.
*/
export interface BashToolParserConfig {
interface BashToolParserConfig {
/** Session ID this parser belongs to */
sessionId: string;
/** Whether the parser is enabled (default: true) */
@@ -462,7 +462,7 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
* Process a single line of terminal output (raw — will strip ANSI).
*/
private processLine(line: string): void {
const cleanLine = this.stripAnsi(line);
const cleanLine = stripAnsi(line);
this.processCleanLine(cleanLine);
}
@@ -470,38 +470,42 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
* Process a single pre-stripped line of terminal output.
*/
private processCleanLine(cleanLine: string): void {
// Check for tool start
if (this._handleToolStart(cleanLine)) return;
if (this._handleToolCompletion(cleanLine)) return;
if (this._handleTextCommand(cleanLine)) return;
this._handleLogFileMention(cleanLine);
}
private _handleToolStart(cleanLine: string): boolean {
const startMatch = cleanLine.match(BASH_TOOL_START_PATTERN);
if (startMatch) {
if (!startMatch) return false;
const command = startMatch[1];
const timeout = startMatch[2]?.trim();
// Check if this is a file-viewing command
if (this.isFileViewerCommand(command)) {
if (!this.isFileViewerCommand(command)) return true;
const filePaths = this.extractFilePaths(command);
// Skip if any file path is already tracked (cross-pattern dedup)
if (filePaths.some((fp) => this.isFilePathTracked(fp))) {
return;
}
if (filePaths.some((fp) => this.isFilePathTracked(fp))) return true;
if (filePaths.length > 0) {
const tool: ActiveBashTool = {
id: uuidv4(),
command,
filePaths,
timeout,
startedAt: Date.now(),
status: 'running',
sessionId: this._sessionId,
};
const tool = this._createActiveTool(command, filePaths, 'running', timeout);
// Enforce max tools limit
if (this._activeTools.size >= MAX_ACTIVE_TOOLS) {
// Remove oldest tool
const oldest = Array.from(this._activeTools.entries()).sort((a, b) => a[1].startedAt - b[1].startedAt)[0];
if (oldest) {
this._activeTools.delete(oldest[0]);
// Remove oldest tool (O(n) min-scan instead of O(n log n) sort)
let oldestKey: string | undefined;
let oldestTime = Infinity;
for (const [key, entry] of this._activeTools) {
if (entry.startedAt < oldestTime) {
oldestTime = entry.startedAt;
oldestKey = key;
}
}
if (oldestKey) {
this._activeTools.delete(oldestKey);
}
}
@@ -511,72 +515,46 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
this.emit('toolStart', tool);
this.scheduleUpdate();
}
}
return;
return true;
}
// Check for tool completion
if (TOOL_COMPLETION_PATTERN.test(cleanLine) && this._lastToolId) {
private _handleToolCompletion(cleanLine: string): boolean {
if (!TOOL_COMPLETION_PATTERN.test(cleanLine) || !this._lastToolId) return false;
const tool = this._activeTools.get(this._lastToolId);
if (tool && tool.status === 'running') {
tool.status = 'completed';
this.emit('toolEnd', tool);
this.scheduleUpdate();
// Remove completed tool after a short delay to allow UI to show completion
this.cleanup.setTimeout(
() => {
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this.scheduleUpdate();
},
2000,
{ description: 'auto-remove completed tool' }
);
this._scheduleAutoRemove(tool.id, 2000, 'auto-remove completed tool');
}
this._lastToolId = null;
return;
return true;
}
// Fallback: Check for command suggestions in plain text (e.g., "tail -f /tmp/file.log")
private _handleTextCommand(cleanLine: string): boolean {
const textCmdMatch = cleanLine.match(TEXT_COMMAND_PATTERN);
if (textCmdMatch) {
if (!textCmdMatch) return false;
const filePath = textCmdMatch[2];
// Create a suggestion tool (marked as 'suggestion' status)
const tool: ActiveBashTool = {
id: uuidv4(),
command: cleanLine.trim(),
filePaths: [filePath],
timeout: undefined,
startedAt: Date.now(),
status: 'running', // Shows as clickable
sessionId: this._sessionId,
};
// Don't add if file path already tracked (cross-pattern dedup)
if (this.isFilePathTracked(filePath)) {
return;
}
if (this.isFilePathTracked(filePath)) return true;
const tool = this._createActiveTool(cleanLine.trim(), [filePath], 'running');
this._activeTools.set(tool.id, tool);
this.emit('toolStart', tool);
this.scheduleUpdate();
// Auto-remove suggestions after 30 seconds
this.cleanup.setTimeout(
() => {
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this.scheduleUpdate();
},
30000,
{ description: 'auto-remove suggestion tool' }
);
return;
this._scheduleAutoRemove(tool.id, 30000, 'auto-remove suggestion tool');
return true;
}
// Last fallback: Check for log file paths mentioned anywhere in the line
private _handleLogFileMention(cleanLine: string): void {
LOG_FILE_MENTION_PATTERN.lastIndex = 0;
let logMatch;
while ((logMatch = LOG_FILE_MENTION_PATTERN.exec(cleanLine)) !== null) {
@@ -588,32 +566,45 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
// Skip if file path already tracked (cross-pattern dedup)
if (this.isFilePathTracked(filePath)) continue;
const tool: ActiveBashTool = {
id: uuidv4(),
command: `View: ${filePath}`,
filePaths: [filePath],
timeout: undefined,
startedAt: Date.now(),
status: 'running',
sessionId: this._sessionId,
};
const tool = this._createActiveTool(`View: ${filePath}`, [filePath], 'running');
this._activeTools.set(tool.id, tool);
this.emit('toolStart', tool);
this.scheduleUpdate();
// Auto-remove after 60 seconds
this._scheduleAutoRemove(tool.id, 60000, 'auto-remove log file tool');
}
}
private _createActiveTool(
command: string,
filePaths: string[],
status: ActiveBashTool['status'],
timeout?: string
): ActiveBashTool {
return {
id: uuidv4(),
command,
filePaths,
timeout,
startedAt: Date.now(),
status,
sessionId: this._sessionId,
};
}
private _scheduleAutoRemove(toolId: string, delayMs: number, description: string): void {
this.cleanup.setTimeout(
() => {
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this._activeTools.delete(toolId);
this.scheduleUpdate();
},
60000,
{ description: 'auto-remove log file tool' }
delayMs,
{ description }
);
}
}
/**
* Check if a command is a file-viewing command worth tracking.
@@ -668,15 +659,6 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
return this.deduplicatePaths(rawPaths);
}
/**
* Strip ANSI escape codes from a string.
*/
private stripAnsi(str: string): string {
// Comprehensive ANSI pattern
// eslint-disable-next-line no-control-regex
return str.replace(/\x1b(?:\[[0-9;?]*[A-Za-z]|\][^\x07\x1b]*(?:\x07|\x1b\\)|[=>])/g, '');
}
/**
* Schedule a debounced update emission.
*/
+44 -4
View File
@@ -1,13 +1,17 @@
/**
* @fileoverview Default model and context limits for AI-powered checkers.
* @fileoverview Default model, context limits, and timing for AI-powered checkers.
*
* Centralizes the AI model identifier and context window sizes used by
* the idle checker, plan checker, respawn controller defaults, and
* respawn route fallbacks. Change the model here when upgrading.
* Centralizes the AI model identifier, context window sizes, and timeout/cooldown
* defaults used by the idle checker, plan checker, respawn controller defaults,
* and respawn route fallbacks. Change values here when tuning AI check behavior.
*
* @module config/ai-defaults
*/
// ============================================================================
// Model & Context
// ============================================================================
/** Default model for AI idle and plan checkers */
export const AI_CHECK_MODEL = 'claude-opus-4-5-20251101';
@@ -16,3 +20,39 @@ export const AI_IDLE_CHECK_MAX_CONTEXT = 16000;
/** Max context chars for plan checker (~2k tokens, plan mode UI is compact) */
export const AI_PLAN_CHECK_MAX_CONTEXT = 8000;
// ============================================================================
// AI Idle Checker Timing
// ============================================================================
/** Timeout for AI idle check (90 seconds — thinking can be slow) */
export const AI_IDLE_CHECK_TIMEOUT_MS = 90_000;
/** Cooldown after WORKING verdict (3 minutes) */
export const AI_IDLE_CHECK_COOLDOWN_MS = 180_000;
/** Cooldown after AI idle check error (1 minute) */
export const AI_IDLE_CHECK_ERROR_COOLDOWN_MS = 60_000;
// ============================================================================
// AI Plan Checker Timing
// ============================================================================
/** Timeout for AI plan check (60 seconds — allows time for thinking) */
export const AI_PLAN_CHECK_TIMEOUT_MS = 60_000;
/** Cooldown after NOT_PLAN_MODE verdict (30 seconds) */
export const AI_PLAN_CHECK_COOLDOWN_MS = 30_000;
/** Cooldown after AI plan check error (30 seconds) */
export const AI_PLAN_CHECK_ERROR_COOLDOWN_MS = 30_000;
// ============================================================================
// Shared AI Checker Limits
// ============================================================================
/** Max consecutive errors before disabling an AI checker */
export const AI_CHECK_MAX_CONSECUTIVE_ERRORS = 3;
/** Maximum exponential backoff cap for AI checker errors (5 minutes) */
export const AI_CHECK_MAX_BACKOFF_MS = 5 * 60 * 1000;
+21 -10
View File
@@ -22,14 +22,16 @@
* Maximum terminal buffer size in characters.
* Contains raw terminal output with ANSI escape sequences.
* Reduced from 5MB to 2MB for better render performance.
* Override: CODEMAN_MAX_TERMINAL_BUFFER (bytes)
*/
export const MAX_TERMINAL_BUFFER_SIZE = 2 * 1024 * 1024; // 2MB
export const MAX_TERMINAL_BUFFER_SIZE = parseInt(process.env.CODEMAN_MAX_TERMINAL_BUFFER || '') || 2 * 1024 * 1024;
/**
* Size to trim terminal buffer to when max is exceeded.
* Keeps the most recent portion to preserve context.
* Override: CODEMAN_TRIM_TERMINAL_TO (bytes)
*/
export const TRIM_TERMINAL_TO = 1.5 * 1024 * 1024; // 1.5MB
export const TRIM_TERMINAL_TO = parseInt(process.env.CODEMAN_TRIM_TERMINAL_TO || '') || 1.5 * 1024 * 1024;
// ============================================================================
// Text Output Buffer Limits
@@ -38,13 +40,15 @@ export const TRIM_TERMINAL_TO = 1.5 * 1024 * 1024; // 1.5MB
/**
* Maximum text output buffer size in characters.
* Contains ANSI-stripped text for search and analysis.
* Override: CODEMAN_MAX_TEXT_OUTPUT (bytes)
*/
export const MAX_TEXT_OUTPUT_SIZE = 1 * 1024 * 1024; // 1MB
export const MAX_TEXT_OUTPUT_SIZE = parseInt(process.env.CODEMAN_MAX_TEXT_OUTPUT || '') || 1 * 1024 * 1024;
/**
* Size to trim text output buffer to when max is exceeded.
* Override: CODEMAN_TRIM_TEXT_TO (bytes)
*/
export const TRIM_TEXT_TO = 768 * 1024; // 768KB
export const TRIM_TEXT_TO = parseInt(process.env.CODEMAN_TRIM_TEXT_TO || '') || 768 * 1024;
// ============================================================================
// Message Buffer Limits
@@ -53,13 +57,9 @@ export const TRIM_TEXT_TO = 768 * 1024; // 768KB
/**
* Maximum number of Claude JSON messages to keep in memory per session.
* Older messages are discarded when limit is exceeded.
* Override: CODEMAN_MAX_MESSAGES (count)
*/
export const MAX_MESSAGES = 1000;
/**
* Number of messages to keep when trimming (80% of max).
*/
export const TRIM_MESSAGES_TO = 800;
export const MAX_MESSAGES = parseInt(process.env.CODEMAN_MAX_MESSAGES || '') || 1000;
// ============================================================================
// Line Buffer Limits
@@ -85,3 +85,14 @@ export const MAX_RESPAWN_BUFFER_SIZE = 1 * 1024 * 1024; // 1MB
* Size to trim respawn buffer to when max is exceeded.
*/
export const TRIM_RESPAWN_BUFFER_TO = 512 * 1024; // 512KB
// ============================================================================
// File Peek Limits
// ============================================================================
/**
* Maximum bytes to read when peeking at the beginning of a file.
* Used with `createReadStream({ end })` (inclusive) to read the first 8KB,
* which is enough to extract metadata from the first few JSONL lines.
*/
export const FILE_PEEK_BYTES = 8 * 1024 - 1; // 8KB (inclusive end offset)
+13
View File
@@ -73,3 +73,16 @@ export const MAX_CONSECUTIVE_ERRORS = 5;
/** Error counter reset interval — forgives errors after quiet period (ms) */
export const ERROR_RESET_MS = 60_000;
// ============================================================================
// Common Cleanup Intervals
// ============================================================================
/** Standard 1-minute cleanup/check interval used by multiple subsystems (ms) */
export const CLEANUP_CHECK_INTERVAL_MS = 60_000;
/** Standard 1-hour max age for stale/completed data (ms) */
export const STALE_DATA_MAX_AGE_MS = 60 * 60 * 1000;
/** Standard 5-minute inactivity timeout for streams and caches (ms) */
export const INACTIVITY_TIMEOUT_MS = 5 * 60 * 1000;
-6
View File
@@ -11,11 +11,5 @@
/** Max input length per API request (bytes) */
export const MAX_INPUT_LENGTH = 64 * 1024;
/** Max terminal columns for resize requests */
export const MAX_TERMINAL_COLS = 500;
/** Max terminal rows for resize requests */
export const MAX_TERMINAL_ROWS = 200;
/** Max session name length (chars) */
export const MAX_SESSION_NAME_LENGTH = 128;
+8 -9
View File
@@ -16,6 +16,8 @@ import { existsSync, statSync, realpathSync } from 'node:fs';
import { resolve, relative, isAbsolute } from 'node:path';
import { homedir } from 'node:os';
import { EventEmitter } from 'node:events';
import { getErrorMessage } from './types.js';
import { CLEANUP_CHECK_INTERVAL_MS, INACTIVITY_TIMEOUT_MS } from './config/server-timing.js';
// ========== Configuration Constants ==========
@@ -39,14 +41,14 @@ const MAX_STREAMS_PER_SESSION = 5;
* Inactivity timeout for streams (5 minutes).
* Streams with no data for this long will be auto-closed.
*/
const STREAM_INACTIVITY_TIMEOUT_MS = 5 * 60 * 1000;
const STREAM_INACTIVITY_TIMEOUT_MS = INACTIVITY_TIMEOUT_MS;
// ========== Types ==========
/**
* Represents an active file stream.
*/
export interface FileStream {
interface FileStream {
/** Unique stream identifier */
id: string;
/** Session this stream belongs to */
@@ -72,7 +74,7 @@ export interface FileStream {
/**
* Options for creating a file stream.
*/
export interface CreateStreamOptions {
interface CreateStreamOptions {
/** Session ID requesting the stream */
sessionId: string;
/** Path to the file to stream */
@@ -92,7 +94,7 @@ export interface CreateStreamOptions {
/**
* Result of creating a stream.
*/
export interface CreateStreamResult {
interface CreateStreamResult {
success: boolean;
streamId?: string;
error?: string;
@@ -129,7 +131,7 @@ export class FileStreamManager extends EventEmitter {
constructor() {
super();
// Start cleanup timer for inactive streams
this.cleanupTimer = setInterval(() => this.cleanupInactiveStreams(), 60 * 1000);
this.cleanupTimer = setInterval(() => this.cleanupInactiveStreams(), CLEANUP_CHECK_INTERVAL_MS);
}
// ========== Public Methods ==========
@@ -171,10 +173,7 @@ export class FileStreamManager extends EventEmitter {
}
} catch (err) {
const errorCode = err instanceof Error && 'code' in err ? (err as NodeJS.ErrnoException).code : 'UNKNOWN';
console.warn(
`[FileStreamManager] Failed to stat file "${absolutePath}" (${errorCode}):`,
err instanceof Error ? err.message : String(err)
);
console.warn(`[FileStreamManager] Failed to stat file "${absolutePath}" (${errorCode}):`, getErrorMessage(err));
return { success: false, error: 'File not found or not accessible' };
}
+64
View File
@@ -84,6 +84,42 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
};
}
/**
* Remove a subset of env keys from .claude/settings.local.json.env if present.
* Used during the disk→tmux-setenv migration: when the caller is actively setting
* a fresh value for a Codeman-managed key, any stale disk entry for THAT KEY is
* superseded and should be removed. Keys NOT in `keysToRemove` are left alone
* (they may be user-managed). No-op if the file/keys don't exist.
*/
export async function stripCaseEnvKeys(casePath: string, keysToRemove: readonly string[]): Promise<void> {
if (keysToRemove.length === 0) return;
const settingsPath = join(casePath, '.claude', 'settings.local.json');
if (!existsSync(settingsPath)) return;
let existing: Record<string, unknown>;
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
return; // Malformed — don't rewrite it
}
const env = existing.env as Record<string, string> | undefined;
if (!env) return;
let changed = false;
for (const key of keysToRemove) {
if (key in env) {
delete env[key];
changed = true;
}
}
if (!changed) return;
existing.env = env;
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
}
/**
* Updates env vars in .claude/settings.local.json for the given case path.
* Merges with existing env field; removes vars set to empty string.
@@ -116,6 +152,34 @@ export async function updateCaseEnvVars(casePath: string, envVars: Record<string
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
}
/**
* Updates the `model` field in .claude/settings.local.json for the given case path.
* Pass a non-empty string to set, or empty/null to remove.
*/
export async function updateCaseModel(casePath: string, model: string | null): Promise<void> {
const claudeDir = join(casePath, '.claude');
if (!existsSync(claudeDir)) {
await mkdir(claudeDir, { recursive: true });
}
const settingsPath = join(claudeDir, 'settings.local.json');
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
existing = {};
}
if (model) {
existing.model = model;
} else {
delete existing.model;
}
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
}
/**
* Writes hooks config to .claude/settings.local.json in the given case path.
* Merges with existing file content, only touching the `hooks` key.
-5
View File
@@ -17,11 +17,6 @@ import { KeyedDebouncer } from './utils/index.js';
// ========== Types ==========
export interface ImageWatcherEvents {
'image:detected': (event: ImageDetectedEvent) => void;
'image:error': (error: Error, sessionId?: string) => void;
}
// ========== Constants ==========
/** Supported image file extensions (lowercase) */
+8
View File
@@ -61,6 +61,10 @@ export interface CreateSessionOptions {
claudeMode?: ClaudeMode;
allowedTools?: string;
openCodeConfig?: OpenCodeConfig;
/** When restoring after reboot, resume a previous Claude conversation by its session ID */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (e.g., CLAUDE_CODE_EFFORT_LEVEL). Ephemeral — not written to disk. */
envOverrides?: Record<string, string>;
}
/** Options for respawning a dead pane. */
@@ -73,6 +77,10 @@ export interface RespawnPaneOptions {
claudeMode?: ClaudeMode;
allowedTools?: string;
openCodeConfig?: OpenCodeConfig;
/** Resume a previous Claude conversation when respawning */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (preserved across respawns). */
envOverrides?: Record<string, string>;
}
/**
+978
View File
@@ -0,0 +1,978 @@
/**
* @fileoverview Orchestrator Loop — phased plan execution with team agents.
*
* State machine that generates plans from user goals, executes them
* phase-by-phase with verification gates, and adapts on failure.
*
* States: idle → planning → approval → executing → verifying → (replanning) → completed/failed
*
* Key exports:
* - `OrchestratorLoop` class — main engine, extends EventEmitter
* - `OrchestratorLoopEvents` interface — typed event map
*
* Lifecycle: `start(goal)` → plan → approve → execute phases → verify → complete
*
* @dependencies orchestrator-planner (plan generation), orchestrator-verifier (phase verification),
* session-manager (sessions), task-queue (task execution), state-store (persistence),
* prompts/orchestrator (prompt templates)
* @consumedby web/server (orchestrator routes, SSE)
* @emits stateChanged, planReady, phaseStarted, phaseCompleted, phaseFailed,
* taskAssigned, taskCompleted, taskFailed, verificationResult, completed, error
* @persistence Orchestrator state saved to `~/.codeman/state.json` (orchestrator key)
*
* @module orchestrator-loop
*/
import { EventEmitter } from 'node:events';
import { getSessionManager, SessionManager } from './session-manager.js';
import { getTaskQueue, TaskQueue } from './task-queue.js';
import { getStore, StateStore } from './state-store.js';
import { OrchestratorPlanner } from './orchestrator-planner.js';
import { OrchestratorVerifier } from './orchestrator-verifier.js';
import { PHASE_EXECUTION_PROMPT, REPLAN_PROMPT, SINGLE_TASK_PROMPT, TEAM_LEAD_PROMPT } from './prompts/index.js';
import type { TerminalMultiplexer } from './mux-interface.js';
import type { CreateTaskOptions } from './task.js';
import {
type OrchestratorState,
type OrchestratorPlan,
type OrchestratorPhase,
type OrchestratorTask,
type OrchestratorConfig,
type OrchestratorStats,
type OrchestratorPersistState,
type VerificationResult,
DEFAULT_ORCHESTRATOR_CONFIG,
createInitialOrchestratorStats,
getErrorMessage,
} from './types.js';
// ═══════════════════════════════════════════════════════════════
// Constants
// ═══════════════════════════════════════════════════════════════
/** Poll interval for checking task completion within a phase (2 seconds) */
const PHASE_POLL_INTERVAL_MS = 2000;
/** Delay between phase completion and verification (1 second) */
const POST_PHASE_DELAY_MS = 1000;
// ═══════════════════════════════════════════════════════════════
// Events
// ═══════════════════════════════════════════════════════════════
// ═══════════════════════════════════════════════════════════════
// OrchestratorLoop
// ═══════════════════════════════════════════════════════════════
export class OrchestratorLoop extends EventEmitter {
private _state: OrchestratorState = 'idle';
private plan: OrchestratorPlan | null = null;
private currentPhaseIndex = 0;
private config: OrchestratorConfig;
private stats: OrchestratorStats;
private startedAt: number | null = null;
private completedAt: number | null = null;
private workingDir: string;
private planner: OrchestratorPlanner;
private verifier: OrchestratorVerifier;
private sessionManager: SessionManager;
private taskQueue: TaskQueue;
private store: StateStore;
/** State before pause (to resume to correct state) */
private pausedState: OrchestratorState | null = null;
/** Phase poll timer for checking task completion */
private phasePollTimer: NodeJS.Timeout | null = null;
/** Phase-level timeout timer */
private phaseTimeoutTimer: NodeJS.Timeout | null = null;
/** Post-phase delay timer before verification */
private postPhaseTimer: NodeJS.Timeout | null = null;
/** Session completion listener (bound for cleanup) */
private sessionCompletionListener: ((sessionId: string, phrase: string) => void) | null = null;
/** Active sessions assigned to current phase */
private phaseSessionIds: Set<string> = new Set();
constructor(mux: TerminalMultiplexer, workingDir: string, config?: Partial<OrchestratorConfig>) {
super();
this.workingDir = workingDir;
this.config = { ...DEFAULT_ORCHESTRATOR_CONFIG, ...config };
this.stats = createInitialOrchestratorStats();
this.sessionManager = getSessionManager();
this.taskQueue = getTaskQueue();
this.store = getStore();
this.planner = new OrchestratorPlanner(mux, workingDir, this.config);
this.verifier = new OrchestratorVerifier(this.config);
// Restore state if crashed while running
this.restore();
}
// ═══════════════════════════════════════════════════════════════
// Public API — Lifecycle
// ═══════════════════════════════════════════════════════════════
/** Start orchestration with a goal. Transitions: idle → planning */
async start(goal: string): Promise<void> {
if (this._state !== 'idle' && this._state !== 'failed' && this._state !== 'completed') {
throw new Error(`Cannot start from state "${this._state}"`);
}
this.reset();
this.startedAt = Date.now();
this.setState('planning');
try {
const plan = await this.planner.generatePlan(goal, (phase, detail) => {
this.emit('planProgress', phase, detail);
});
if (this.currentState() !== 'planning') {
// Cancelled during planning
return;
}
this.plan = plan;
this.persist();
if (this.config.autoApprove) {
this.setState('executing');
await this.executeCurrentPhase();
} else {
this.setState('approval');
this.emit('planReady', plan);
}
} catch (err) {
this.handleError(err);
}
}
/** Approve the generated plan. Transitions: approval → executing */
async approve(): Promise<void> {
this.requireState('approval');
if (!this.plan) {
throw new Error('No plan to approve');
}
this.setState('executing');
await this.executeCurrentPhase();
}
/** Reject plan with feedback. Transitions: approval → planning (regenerate) */
async reject(feedback: string): Promise<void> {
this.requireState('approval');
if (!this.plan) {
throw new Error('No plan to reject');
}
const goal = this.plan.goal + '\n\nFeedback on previous plan: ' + feedback;
this.plan = null;
this.setState('planning');
try {
const plan = await this.planner.generatePlan(goal);
if ((this._state as OrchestratorState) !== 'planning') return;
this.plan = plan;
this.persist();
this.setState('approval');
this.emit('planReady', plan);
} catch (err) {
this.handleError(err);
}
}
/** Pause execution. Saves current state. */
pause(): void {
if (this._state === 'idle' || this._state === 'paused' || this._state === 'completed' || this._state === 'failed') {
return;
}
this.pausedState = this._state;
this.clearPhasePoll();
this.cleanupTaskHandlers();
this.setState('paused');
}
/** Resume from pause. */
async resume(): Promise<void> {
if (this._state !== 'paused' || !this.pausedState) {
throw new Error('Not paused');
}
const resumeTo = this.pausedState;
this.pausedState = null;
this.setState(resumeTo);
// Re-enter the appropriate phase of execution
if (resumeTo === 'executing') {
await this.executeCurrentPhase();
} else if (resumeTo === 'verifying') {
await this.verifyCurrentPhase();
}
}
/** Stop everything and clean up. */
async stop(): Promise<void> {
this.clearPhasePoll();
this.cleanupTaskHandlers();
await this.planner.cancel();
this.setState('idle');
this.store.clearOrchestratorState();
}
/** Skip a specific phase. */
async skipPhase(phaseId: string): Promise<void> {
if (!this.plan) return;
const phase = this.plan.phases.find((p) => p.id === phaseId);
if (!phase) throw new Error(`Phase "${phaseId}" not found`);
phase.status = 'skipped';
phase.completedAt = Date.now();
this.persist();
// If this is the current phase, advance
if (this.plan.phases[this.currentPhaseIndex]?.id === phaseId) {
await this.advanceToNextPhase();
}
}
/** Retry a failed phase. */
async retryPhase(phaseId: string): Promise<void> {
if (!this.plan) return;
if (this._state !== 'executing' && this._state !== 'failed') {
throw new Error(`Cannot retry from state "${this._state}"`);
}
const phaseIndex = this.plan.phases.findIndex((p) => p.id === phaseId);
if (phaseIndex === -1) throw new Error(`Phase "${phaseId}" not found`);
const phase = this.plan.phases[phaseIndex];
phase.status = 'pending';
phase.attempts = 0;
for (const task of phase.tasks) {
task.status = 'pending';
task.error = null;
task.assignedSessionId = null;
task.queueTaskId = null;
}
this.currentPhaseIndex = phaseIndex;
this.setState('executing');
await this.executeCurrentPhase();
}
// ═══════════════════════════════════════════════════════════════
// Public API — Getters
// ═══════════════════════════════════════════════════════════════
get state(): OrchestratorState {
return this._state;
}
getPlan(): OrchestratorPlan | null {
return this.plan;
}
getCurrentPhase(): OrchestratorPhase | null {
if (!this.plan) return null;
return this.plan.phases[this.currentPhaseIndex] ?? null;
}
getStats(): OrchestratorStats {
return { ...this.stats };
}
getStatus(): OrchestratorPersistState {
return {
state: this._state,
plan: this.plan,
currentPhaseIndex: this.currentPhaseIndex,
startedAt: this.startedAt,
completedAt: this.completedAt,
config: this.config,
stats: this.stats,
};
}
isRunning(): boolean {
return this._state !== 'idle' && this._state !== 'completed' && this._state !== 'failed';
}
// ═══════════════════════════════════════════════════════════════
// Internal — Phase Execution
// ═══════════════════════════════════════════════════════════════
private async executeCurrentPhase(): Promise<void> {
if (!this.plan || this._state !== 'executing') return;
const phase = this.plan.phases[this.currentPhaseIndex];
if (!phase) {
// All phases done
await this.handleCompletion();
return;
}
// Skip already completed/skipped phases
if (phase.status === 'passed' || phase.status === 'skipped') {
await this.advanceToNextPhase();
return;
}
phase.status = 'executing';
phase.startedAt = Date.now();
phase.attempts++;
this.persist();
this.emit('phaseStarted', phase);
try {
await this.assignPhaseTasks(phase);
this.startPhasePoll(phase);
} catch (err) {
this.handlePhaseError(phase, getErrorMessage(err));
}
}
private async assignPhaseTasks(phase: OrchestratorPhase): Promise<void> {
// For team strategy, send a single comprehensive prompt to a lead session
if (phase.teamStrategy.type === 'team') {
await this.assignTeamPhase(phase);
return;
}
// For single/parallel strategy, add individual tasks to TaskQueue
for (const task of phase.tasks) {
if (task.status !== 'pending') continue;
const prompt = this.buildTaskPrompt(task, phase);
const taskOptions: CreateTaskOptions = {
prompt,
workingDir: this.workingDir,
priority: 100 - phase.order, // Earlier phases get higher priority
completionPhrase: task.completionPhrase,
timeoutMs: Math.min(task.timeoutMs, this.config.phaseTimeoutMs),
};
const queueTask = this.taskQueue.addTask(taskOptions);
task.queueTaskId = queueTask.id;
task.status = 'running';
}
this.persist();
this.setupTaskHandlers();
// Manually assign tasks to idle sessions
await this.assignQueuedTasksToSessions();
}
private async assignTeamPhase(phase: OrchestratorPhase): Promise<void> {
const teamConfig = phase.teamStrategy.type === 'team' ? phase.teamStrategy.config : null;
if (!teamConfig) return;
// Find or use an idle session
const sessions = this.sessionManager.getIdleSessions();
if (sessions.length === 0) {
throw new Error('No idle sessions available for team phase execution');
}
const session = sessions[0];
this.phaseSessionIds.add(session.id);
// Mark all tasks as running under this session
for (const task of phase.tasks) {
task.status = 'running';
task.assignedSessionId = session.id;
}
// Build and send the team lead prompt
const prompt = TEAM_LEAD_PROMPT.replace('{PHASE_NAME}', phase.name)
.replace('{TASK_LIST}', phase.tasks.map((t, i) => `${i + 1}. ${t.prompt}`).join('\n'))
.replace('{TEAMMATE_HINTS}', teamConfig.suggestedTeammates.map((h, i) => `${i + 1}. ${h}`).join('\n'))
.replace('{COMPLETION_PHRASE}', `${phase.id.toUpperCase()}_COMPLETE`);
// Create a TaskQueue task for the entire phase
const queueTask = this.taskQueue.addTask({
prompt,
workingDir: this.workingDir,
priority: 100 - phase.order,
completionPhrase: `${phase.id.toUpperCase()}_COMPLETE`,
timeoutMs: this.config.phaseTimeoutMs,
});
// Link all phase tasks to this single queue task
for (const task of phase.tasks) {
task.queueTaskId = queueTask.id;
}
this.persist();
this.setupTaskHandlers();
// Assign the task to the session
try {
queueTask.assign(session.id);
session.assignTask(queueTask.id);
this.taskQueue.updateTask(queueTask);
await session.sendInput(prompt);
} catch (err) {
queueTask.fail(getErrorMessage(err));
this.taskQueue.updateTask(queueTask);
throw err;
}
}
private async assignQueuedTasksToSessions(): Promise<void> {
const idleSessions = this.sessionManager.getIdleSessions();
const maxSessions =
this.getCurrentPhase()?.teamStrategy.type === 'parallel'
? (this.getCurrentPhase()?.teamStrategy as { type: 'parallel'; maxSessions: number }).maxSessions
: 1;
const sessionsToUse = idleSessions.slice(0, maxSessions);
for (const session of sessionsToUse) {
const task = this.taskQueue.next();
if (!task) break;
try {
task.assign(session.id);
session.assignTask(task.id);
this.taskQueue.updateTask(task);
await session.sendInput(task.prompt);
this.phaseSessionIds.add(session.id);
// Find the orchestrator task linked to this queue task
const orchTask = this.findOrchestratorTaskByQueueId(task.id);
if (orchTask) {
orchTask.assignedSessionId = session.id;
orchTask.startedAt = Date.now();
this.emit('taskAssigned', orchTask, session.id);
}
} catch (err) {
task.fail(getErrorMessage(err));
session.clearTask();
this.taskQueue.updateTask(task);
}
}
}
// ═══════════════════════════════════════════════════════════════
// Internal — Task Completion Tracking
// ═══════════════════════════════════════════════════════════════
private setupTaskHandlers(): void {
this.cleanupTaskHandlers();
this.sessionCompletionListener = (_sessionId: string, _phrase: string) => {
// Session completion — check if it's related to our phase tasks
this.checkPhaseCompletion();
};
this.sessionManager.on('sessionCompletion', this.sessionCompletionListener);
}
private cleanupTaskHandlers(): void {
if (this.sessionCompletionListener) {
this.sessionManager.off('sessionCompletion', this.sessionCompletionListener);
this.sessionCompletionListener = null;
}
}
private _finalizeTask(queueTaskId: string, status: 'completed' | 'failed', error?: string): OrchestratorTask | null {
const orchTask = this.findOrchestratorTaskByQueueId(queueTaskId);
if (!orchTask) return null;
orchTask.status = status;
if (status === 'completed') {
orchTask.completedAt = Date.now();
this.stats.totalTasksCompleted++;
} else {
orchTask.error = error ?? null;
this.stats.totalTasksFailed++;
}
this.persist();
return orchTask;
}
private handleTaskCompleted(queueTaskId: string): void {
const orchTask = this._finalizeTask(queueTaskId, 'completed');
if (!orchTask) return;
this.emit('taskCompleted', orchTask);
this.checkPhaseCompletion();
}
private handleTaskFailed(queueTaskId: string, error: string): void {
const orchTask = this._finalizeTask(queueTaskId, 'failed', error);
if (!orchTask) return;
this.emit('taskFailed', orchTask, error);
// Check if we should retry the task or fail the phase
if (orchTask.retries < 2) {
orchTask.retries++;
orchTask.status = 'pending';
orchTask.error = null;
orchTask.queueTaskId = null;
// Will be re-queued on next poll
} else {
this.checkPhaseCompletion();
}
}
private startPhasePoll(phase: OrchestratorPhase): void {
this.clearPhasePoll();
this.phasePollTimer = setInterval(() => {
if (this._state !== 'executing') {
this.clearPhasePoll();
return;
}
this.pollPhaseStatus(phase);
}, PHASE_POLL_INTERVAL_MS);
// Phase-level timeout — fail the phase if it exceeds the configured timeout
this.phaseTimeoutTimer = setTimeout(() => {
if (this._state === 'executing' && phase.status === 'executing') {
console.warn(`[Orchestrator] Phase "${phase.name}" timed out after ${this.config.phaseTimeoutMs}ms`);
this.handlePhaseError(phase, `Phase timed out after ${Math.round(this.config.phaseTimeoutMs / 60000)} minutes`);
}
}, this.config.phaseTimeoutMs);
}
private _clearTimer(
timerKey: 'phasePollTimer' | 'phaseTimeoutTimer' | 'postPhaseTimer',
clearFn: typeof clearInterval | typeof clearTimeout
): void {
if (this[timerKey]) {
clearFn(this[timerKey]);
this[timerKey] = null;
}
}
private clearPhasePoll(): void {
this._clearTimer('phasePollTimer', clearInterval);
this._clearTimer('phaseTimeoutTimer', clearTimeout);
this._clearTimer('postPhaseTimer', clearTimeout);
}
private pollPhaseStatus(phase: OrchestratorPhase): void {
// Check for queued tasks that need assignment
const pendingTasks = phase.tasks.filter((t) => t.status === 'pending' && !t.queueTaskId);
if (pendingTasks.length > 0) {
// Re-queue pending tasks
for (const task of pendingTasks) {
const prompt = this.buildTaskPrompt(task, phase);
const queueTask = this.taskQueue.addTask({
prompt,
workingDir: this.workingDir,
priority: 100 - phase.order,
completionPhrase: task.completionPhrase,
timeoutMs: Math.min(task.timeoutMs, this.config.phaseTimeoutMs),
});
task.queueTaskId = queueTask.id;
task.status = 'running';
}
this.assignQueuedTasksToSessions().catch(() => {}); // Best effort
}
// Check completion status of queue tasks
for (const task of phase.tasks) {
if (task.status === 'running' && task.queueTaskId) {
const queueTask = this.taskQueue.getTask(task.queueTaskId);
if (queueTask) {
if (queueTask.isCompleted()) {
this.handleTaskCompleted(task.queueTaskId);
} else if (queueTask.isFailed()) {
this.handleTaskFailed(task.queueTaskId, queueTask.error || 'Task failed');
}
}
}
}
this.checkPhaseCompletion();
}
private checkPhaseCompletion(): void {
if (this._state !== 'executing') return;
const phase = this.getCurrentPhase();
if (!phase) return;
const allDone = phase.tasks.every((t) => t.status === 'completed' || t.status === 'failed');
if (!allDone) return;
const anyFailed = phase.tasks.some((t) => t.status === 'failed');
this.clearPhasePoll();
if (anyFailed) {
// Phase has failed tasks
this.handlePhaseError(phase, 'One or more tasks failed');
} else {
// All tasks completed — run verification after brief delay
this.postPhaseTimer = setTimeout(() => {
this.postPhaseTimer = null;
this.verifyCurrentPhase().catch((err) => this.handleError(err));
}, POST_PHASE_DELAY_MS);
}
}
// ═══════════════════════════════════════════════════════════════
// Internal — Verification
// ═══════════════════════════════════════════════════════════════
private async verifyCurrentPhase(): Promise<void> {
if (!this.plan) return;
const phase = this.plan.phases[this.currentPhaseIndex];
if (!phase) return;
// Skip verification if no criteria defined
if (phase.verificationCriteria.length === 0 && phase.testCommands.length === 0) {
phase.status = 'passed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesCompleted++;
this.persist();
this.emit('phaseCompleted', phase);
await this.advanceToNextPhase();
return;
}
this.setState('verifying');
// Get a session for verification — wait briefly for sessions to become idle
let sessions = this.sessionManager.getIdleSessions();
if (sessions.length === 0) {
// Wait up to 10s for a session to become idle
await new Promise((resolve) => setTimeout(resolve, 10_000));
sessions = this.sessionManager.getIdleSessions();
}
if (sessions.length === 0) {
// Still no sessions — log warning and skip verification (don't silently pass)
console.warn('[Orchestrator] No idle sessions for verification — skipping (marking passed with warning)');
phase.status = 'passed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesCompleted++;
this.persist();
this.emit('phaseCompleted', phase);
this.setState('executing');
await this.advanceToNextPhase();
return;
}
try {
const result = await this.verifier.verifyPhase(phase, sessions[0]);
this.emit('verificationResult', phase, result);
if (result.passed) {
phase.status = 'passed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesCompleted++;
this.persist();
this.emit('phaseCompleted', phase);
this.setState('executing');
await this.advanceToNextPhase();
} else {
// Verification failed — attempt replan
await this.handleVerificationFailure(phase, result);
}
} catch (err) {
// Verification error — treat as pass (don't block on verification bugs)
console.warn('[Orchestrator] Verification error, treating as pass:', err);
phase.status = 'passed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesCompleted++;
this.persist();
this.emit('phaseCompleted', phase);
this.setState('executing');
await this.advanceToNextPhase();
}
}
private async handleVerificationFailure(phase: OrchestratorPhase, result: VerificationResult): Promise<void> {
if (phase.attempts >= phase.maxAttempts) {
// Max retries exceeded
phase.status = 'failed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesFailed++;
this.persist();
this.emit('phaseFailed', phase, `Verification failed after ${phase.attempts} attempts: ${result.summary}`);
this.setState('failed');
return;
}
// Replan and retry
this.stats.replanCount++;
this.setState('replanning');
try {
await this.replanPhase(phase, result);
// Reset task states for retry
for (const task of phase.tasks) {
task.status = 'pending';
task.error = null;
task.assignedSessionId = null;
task.queueTaskId = null;
task.completedAt = null;
task.startedAt = null;
}
phase.status = 'pending';
phase.startedAt = null;
this.persist();
this.setState('executing');
await this.executeCurrentPhase();
} catch (err) {
this.handleError(err);
}
}
private async replanPhase(phase: OrchestratorPhase, result: VerificationResult): Promise<void> {
const completionPhrase = phase.tasks[0]?.completionPhrase || `${phase.id.toUpperCase()}_FIXED`;
const prompt = REPLAN_PROMPT.replace('{PHASE_NAME}', phase.name)
.replace('{ATTEMPT_NUMBER}', String(phase.attempts))
.replace('{MAX_ATTEMPTS}', String(phase.maxAttempts))
.replace('{FAILURE_SUMMARY}', result.summary)
.replace('{SUGGESTIONS}', result.suggestions.join('\n'))
.replace('{ORIGINAL_TASKS}', phase.tasks.map((t, i) => `${i + 1}. ${t.prompt}`).join('\n'))
.replace('{COMPLETION_PHRASE}', completionPhrase);
// Create a tracked queue task for the replan (so completion is detected)
const queueTask = this.taskQueue.addTask({
prompt,
workingDir: this.workingDir,
priority: 100,
completionPhrase,
timeoutMs: this.config.phaseTimeoutMs,
});
// Link to first phase task for tracking
if (phase.tasks[0]) {
phase.tasks[0].queueTaskId = queueTask.id;
phase.tasks[0].status = 'running';
}
this.persist();
// Set up handlers so task completion is tracked
this.setupTaskHandlers();
// Assign to a session
const sessions = this.sessionManager.getIdleSessions();
if (sessions.length === 0) {
console.warn('[Orchestrator] No idle sessions for replan — task queued, will pick up on next poll');
// Start polling so the task gets assigned when a session becomes idle
this.startPhasePoll(phase);
return;
}
try {
queueTask.assign(sessions[0].id);
sessions[0].assignTask(queueTask.id);
this.taskQueue.updateTask(queueTask);
await sessions[0].sendInput(prompt);
} catch (err) {
queueTask.fail(getErrorMessage(err));
this.taskQueue.updateTask(queueTask);
}
}
// ═══════════════════════════════════════════════════════════════
// Internal — State Machine
// ═══════════════════════════════════════════════════════════════
/** Read current state (bypasses TypeScript narrowing from guards) */
private currentState(): OrchestratorState {
return this._state;
}
/** Assert state matches expected or throw */
private requireState(...expected: OrchestratorState[]): void {
if (!expected.includes(this._state)) {
throw new Error(`Expected state "${expected.join('|')}", got "${this._state}"`);
}
}
private setState(newState: OrchestratorState): void {
const prev = this._state;
if (prev === newState) return;
this._state = newState;
this.persist();
this.emit('stateChanged', newState, prev);
}
private async advanceToNextPhase(): Promise<void> {
this.currentPhaseIndex++;
this.phaseSessionIds.clear();
this.persist();
if (!this.plan || this.currentPhaseIndex >= this.plan.phases.length) {
await this.handleCompletion();
} else {
// Compact between phases if configured
if (this.config.compactBetweenPhases) {
const sessions = this.sessionManager.getIdleSessions();
for (const session of sessions) {
try {
await session.writeViaMux('/compact');
} catch {
// Best effort
}
}
// Brief delay for compact to take effect
await new Promise((resolve) => setTimeout(resolve, 2000));
}
await this.executeCurrentPhase();
}
}
private async handleCompletion(): Promise<void> {
this.completedAt = Date.now();
this.stats.totalDurationMs = this.startedAt ? this.completedAt - this.startedAt : 0;
this.clearPhasePoll();
this.cleanupTaskHandlers();
this.setState('completed');
this.emit('completed', this.stats);
}
private handlePhaseError(phase: OrchestratorPhase, error: string): void {
if (phase.attempts >= phase.maxAttempts) {
phase.status = 'failed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesFailed++;
this.persist();
this.emit('phaseFailed', phase, error);
this.setState('failed');
} else {
// Retry the phase
for (const task of phase.tasks) {
if (task.status === 'failed') {
task.status = 'pending';
task.error = null;
task.queueTaskId = null;
task.assignedSessionId = null;
}
}
phase.status = 'pending';
this.persist();
this.executeCurrentPhase().catch((err) => this.handleError(err));
}
}
private handleError(err: unknown): void {
const error = err instanceof Error ? err : new Error(getErrorMessage(err));
console.error('[Orchestrator] Error:', error.message);
this.setState('failed');
this.emit('error', error);
}
// ═══════════════════════════════════════════════════════════════
// Internal — Persistence
// ═══════════════════════════════════════════════════════════════
private persist(): void {
this.store.setOrchestratorState(this.getStatus());
}
private restore(): void {
const saved = this.store.getOrchestratorState();
if (!saved) return;
// If we crashed while running, reset to failed
if (saved.state === 'executing' || saved.state === 'verifying' || saved.state === 'replanning') {
this._state = 'failed';
this.plan = saved.plan;
this.currentPhaseIndex = saved.currentPhaseIndex;
this.startedAt = saved.startedAt;
this.config = saved.config;
this.stats = saved.stats;
this.store.setOrchestratorState({ ...saved, state: 'failed' });
} else if (saved.state === 'planning' || saved.state === 'approval') {
// Planning/approval — reset to idle (plan is lost)
this.store.clearOrchestratorState();
} else if (saved.state === 'completed' || saved.state === 'failed') {
// Preserve completed/failed state for UI display
this._state = saved.state;
this.plan = saved.plan;
this.currentPhaseIndex = saved.currentPhaseIndex;
this.startedAt = saved.startedAt;
this.completedAt = saved.completedAt;
this.config = saved.config;
this.stats = saved.stats;
}
}
private reset(): void {
this._state = 'idle';
this.plan = null;
this.currentPhaseIndex = 0;
this.startedAt = null;
this.completedAt = null;
this.stats = createInitialOrchestratorStats();
this.pausedState = null;
this.phaseSessionIds.clear();
this.clearPhasePoll();
this.cleanupTaskHandlers();
}
// ═══════════════════════════════════════════════════════════════
// Internal — Helpers
// ═══════════════════════════════════════════════════════════════
private buildTaskPrompt(task: OrchestratorTask, phase: OrchestratorPhase): string {
if (phase.tasks.length === 1) {
// Single task — use simpler prompt
const completedPhases = this.getCompletedPhasesSummary();
return SINGLE_TASK_PROMPT.replace('{TASK}', task.prompt)
.replace('{GOAL}', this.plan?.goal || '')
.replace('{CONTEXT}', completedPhases ? `Previous phases completed: ${completedPhases}` : '')
.replace('{COMPLETION_PHRASE}', task.completionPhrase);
}
// Multi-task phase — use full prompt
return PHASE_EXECUTION_PROMPT.replace('{PHASE_NAME}', phase.name)
.replace('{GOAL}', this.plan?.goal || '')
.replace('{COMPLETED_PHASES}', this.getCompletedPhasesSummary() || 'None yet')
.replace('{TASK_LIST}', phase.tasks.map((t, i) => `${i + 1}. ${t.prompt}`).join('\n'))
.replace('{VERIFICATION_CRITERIA}', phase.verificationCriteria.join('\n') || 'No specific criteria')
.replace('{COMPLETION_PHRASE}', task.completionPhrase);
}
private getCompletedPhasesSummary(): string {
if (!this.plan) return '';
return this.plan.phases
.filter((p) => p.status === 'passed' || p.status === 'skipped')
.map((p) => `${p.name}: ${p.status}`)
.join(', ');
}
private findOrchestratorTaskByQueueId(queueTaskId: string): OrchestratorTask | null {
if (!this.plan) return null;
for (const phase of this.plan.phases) {
for (const task of phase.tasks) {
if (task.queueTaskId === queueTaskId) return task;
}
}
return null;
}
/** Clean up resources when the loop is being destroyed. */
destroy(): void {
this.clearPhasePoll();
this.cleanupTaskHandlers();
}
}
+412
View File
@@ -0,0 +1,412 @@
/**
* @fileoverview Orchestrator plan generation — converts goals into phased plans.
*
* Wraps PlanOrchestrator for AI-powered plan generation, then groups the
* resulting PlanItems into sequential phases with team strategies and
* verification criteria.
*
* Phase grouping algorithm:
* 1. Topological sort by dependencies (Kahn's algorithm)
* 2. Group into dependency layers
* 3. Sub-group by TDD phase within layers
* 4. Merge small adjacent phases
* 5. Assign team strategies based on parallelism potential
*
* Key exports:
* - `OrchestratorPlanner` class — plan generation + phase grouping
*
* @dependencies plan-orchestrator (AI plan generation), types (OrchestratorPlan, PlanItem)
* @consumedby orchestrator-loop
*
* @module orchestrator-planner
*/
import { v4 as uuidv4 } from 'uuid';
import { PlanOrchestrator, type DetailedPlanResult, type ProgressCallback } from './plan-orchestrator.js';
import type { TerminalMultiplexer } from './mux-interface.js';
import type {
PlanItem,
TddPhase,
OrchestratorPlan,
OrchestratorPhase,
OrchestratorTask,
OrchestratorConfig,
TeamStrategy,
PhaseStatus,
} from './types.js';
// ═══════════════════════════════════════════════════════════════
// Constants
// ═══════════════════════════════════════════════════════════════
/** Maximum number of phases (prevents runaway plans) */
const MAX_PHASES = 10;
/** Maximum total tasks across all phases */
const MAX_TOTAL_TASKS = 50;
/** Default task timeout (10 minutes) */
const DEFAULT_TASK_TIMEOUT_MS = 10 * 60 * 1000;
/** Minimum tasks in a phase before it gets merged with adjacent */
const MIN_PHASE_TASKS = 2;
/** TDD phase ordering for grouping */
const TDD_PHASE_ORDER: Record<TddPhase, number> = {
setup: 0,
test: 1,
impl: 2,
verify: 3,
review: 4,
};
// ═══════════════════════════════════════════════════════════════
// OrchestratorPlanner
// ═══════════════════════════════════════════════════════════════
export class OrchestratorPlanner {
private mux: TerminalMultiplexer;
private workingDir: string;
private config: OrchestratorConfig;
private orchestrator: PlanOrchestrator | null = null;
constructor(mux: TerminalMultiplexer, workingDir: string, config: OrchestratorConfig) {
this.mux = mux;
this.workingDir = workingDir;
this.config = config;
}
/**
* Generate a phased plan from a user goal.
*
* Uses PlanOrchestrator for AI plan generation, then groups results into phases.
*/
async generatePlan(goal: string, onProgress?: ProgressCallback): Promise<OrchestratorPlan> {
const startTime = Date.now();
// Create a PlanOrchestrator for this plan generation
this.orchestrator = new PlanOrchestrator(this.mux, this.workingDir, undefined, {
defaultModel: this.config.plannerModel,
});
try {
onProgress?.('planning', 'Generating detailed plan...');
const result: DetailedPlanResult = await this.orchestrator.generateDetailedPlan(goal, onProgress);
if (!result.success || !result.items || result.items.length === 0) {
throw new Error(result.error || 'Plan generation returned no items');
}
// Cap total tasks
const items = result.items.slice(0, MAX_TOTAL_TASKS);
onProgress?.('grouping', 'Organizing plan into phases...');
// Group items into phases
const phases = this.groupIntoPhases(items, goal);
// Assign team strategies
this.assignTeamStrategies(phases);
// Generate unique completion phrases
this.generateCompletionPhrases(phases);
const plan: OrchestratorPlan = {
id: uuidv4(),
goal,
createdAt: Date.now(),
phases,
metadata: {
totalTasks: phases.reduce((sum, p) => sum + p.tasks.length, 0),
estimatedComplexity: this.estimateComplexity(items),
modelUsed: this.config.plannerModel,
planDurationMs: Date.now() - startTime,
},
};
return plan;
} finally {
this.orchestrator = null;
}
}
/** Cancel in-progress plan generation. */
async cancel(): Promise<void> {
if (this.orchestrator) {
await this.orchestrator.cancel();
this.orchestrator = null;
}
}
// ═══════════════════════════════════════════════════════════════
// Phase Grouping
// ═══════════════════════════════════════════════════════════════
/**
* Group PlanItems into sequential phases.
*
* Algorithm:
* 1. Build dependency graph and assign IDs to items without them
* 2. Topological sort into dependency layers (Kahn's algorithm)
* 3. Sub-group within each layer by TDD phase
* 4. Merge small phases with their neighbors
*/
private groupIntoPhases(items: PlanItem[], _goal: string): OrchestratorPhase[] {
// Ensure all items have IDs
const indexedItems = items.map((item, i) => ({
...item,
id: item.id || `task-${i}`,
}));
// Build adjacency and in-degree for Kahn's algorithm
const idSet = new Set(indexedItems.map((item) => item.id!));
const inDegree = new Map<string, number>();
const dependents = new Map<string, string[]>(); // id → items that depend on it
for (const item of indexedItems) {
inDegree.set(item.id!, 0);
dependents.set(item.id!, []);
}
for (const item of indexedItems) {
const deps = (item.dependencies || []).filter((d) => idSet.has(d));
inDegree.set(item.id!, deps.length);
for (const dep of deps) {
dependents.get(dep)!.push(item.id!);
}
}
// Kahn's algorithm — produce dependency layers
const layers: PlanItem[][] = [];
const remaining = new Set(indexedItems.map((item) => item.id!));
while (remaining.size > 0) {
// Find items with no remaining dependencies (in-degree 0)
const layer: PlanItem[] = [];
for (const id of remaining) {
if (inDegree.get(id)! === 0) {
layer.push(indexedItems.find((item) => item.id === id)!);
}
}
if (layer.length === 0) {
// Circular dependency — add all remaining items as a single layer
for (const id of remaining) {
layer.push(indexedItems.find((item) => item.id === id)!);
}
}
layers.push(layer);
// Remove this layer's items and update in-degrees
for (const item of layer) {
remaining.delete(item.id!);
for (const dep of dependents.get(item.id!) || []) {
if (remaining.has(dep)) {
inDegree.set(dep, Math.max(0, inDegree.get(dep)! - 1));
}
}
}
}
// Sub-group each layer by TDD phase
const rawPhases: PlanItem[][] = [];
for (const layer of layers) {
const byPhase = new Map<string, PlanItem[]>();
for (const item of layer) {
const phase = item.tddPhase || 'impl';
if (!byPhase.has(phase)) byPhase.set(phase, []);
byPhase.get(phase)!.push(item);
}
// Sort sub-groups by TDD phase order
const sorted = [...byPhase.entries()].sort(
([a], [b]) => (TDD_PHASE_ORDER[a as TddPhase] ?? 2) - (TDD_PHASE_ORDER[b as TddPhase] ?? 2)
);
for (const [, items] of sorted) {
rawPhases.push(items);
}
}
// Merge small phases with their previous neighbor
const mergedPhases: PlanItem[][] = [];
for (const phase of rawPhases) {
if (mergedPhases.length > 0 && phase.length < MIN_PHASE_TASKS) {
const prev = mergedPhases[mergedPhases.length - 1];
if (prev.length < MIN_PHASE_TASKS) {
// Merge with previous
prev.push(...phase);
continue;
}
}
mergedPhases.push([...phase]);
}
// Cap at MAX_PHASES by merging tail phases
while (mergedPhases.length > MAX_PHASES) {
const last = mergedPhases.pop()!;
mergedPhases[mergedPhases.length - 1].push(...last);
}
// Convert to OrchestratorPhase objects
return mergedPhases.map((phaseItems, index) => this.createPhase(phaseItems, index));
}
private createPhase(items: PlanItem[], order: number): OrchestratorPhase {
// Derive phase name from TDD phases and priorities
const tddPhases = [...new Set(items.map((i) => i.tddPhase).filter(Boolean))];
const name = this.generatePhaseName(items, tddPhases as TddPhase[], order);
const description = items.map((i) => i.content).join('; ');
const tasks: OrchestratorTask[] = items.map((item, i) => ({
id: `phase-${order + 1}-task-${i + 1}`,
phaseId: `phase-${order + 1}`,
prompt: item.content,
status: 'pending' as const,
assignedSessionId: null,
queueTaskId: null,
parallel: items.length > 1, // Tasks within a phase are parallel by default
completionPhrase: '', // Assigned later
timeoutMs: DEFAULT_TASK_TIMEOUT_MS,
startedAt: null,
completedAt: null,
error: null,
retries: 0,
}));
// Extract verification criteria and test commands from items
const verificationCriteria = items
.map((i) => i.verificationCriteria)
.filter((v): v is string => v != null && v.length > 0);
const testCommands = items.map((i) => i.testCommand).filter((t): t is string => t != null && t.length > 0);
return {
id: `phase-${order + 1}`,
name,
description,
order,
status: 'pending' as PhaseStatus,
tasks,
verificationCriteria,
testCommands,
maxAttempts: this.config.maxPhaseRetries,
attempts: 0,
startedAt: null,
completedAt: null,
durationMs: null,
teamStrategy: { type: 'single' }, // Assigned later
};
}
private generatePhaseName(items: PlanItem[], tddPhases: TddPhase[], order: number): string {
// Try to create a meaningful name based on content
const priorities = [...new Set(items.map((i) => i.priority).filter(Boolean))];
if (tddPhases.length === 1) {
const phaseNames: Record<TddPhase, string> = {
setup: 'Setup & Configuration',
test: 'Test Definition',
impl: 'Implementation',
verify: 'Verification',
review: 'Review & Polish',
};
return `Phase ${order + 1}: ${phaseNames[tddPhases[0]]}`;
}
if (priorities.includes('P0') && priorities.length === 1) {
return `Phase ${order + 1}: Critical Foundation`;
}
return `Phase ${order + 1}: ${items.length > 1 ? 'Parallel Tasks' : items[0].content.slice(0, 50)}`;
}
// ═══════════════════════════════════════════════════════════════
// Team Strategy Assignment
// ═══════════════════════════════════════════════════════════════
private assignTeamStrategies(phases: OrchestratorPhase[]): void {
for (const phase of phases) {
phase.teamStrategy = this.computeTeamStrategy(phase);
}
}
private computeTeamStrategy(phase: OrchestratorPhase): TeamStrategy {
const taskCount = phase.tasks.length;
const parallelTasks = phase.tasks.filter((t) => t.parallel).length;
// Single task or no parallel potential → single session
if (taskCount <= 2 || parallelTasks <= 1) {
return { type: 'single' };
}
// If team agents are disabled, use parallel sessions instead
if (!this.config.enableTeamAgents) {
return {
type: 'parallel',
maxSessions: Math.min(parallelTasks, this.config.maxParallelSessions),
};
}
// 4+ parallel tasks with team agents enabled → team mode
if (parallelTasks >= 4) {
return {
type: 'team',
config: {
leadPrompt: this.buildTeamLeadPrompt(phase),
suggestedTeammates: phase.tasks.slice(0, 4).map((t) => `Specialist for: ${t.prompt.slice(0, 80)}`),
maxTeammates: Math.min(parallelTasks, 4),
},
};
}
// 3 parallel tasks → parallel sessions
return {
type: 'parallel',
maxSessions: Math.min(parallelTasks, this.config.maxParallelSessions),
};
}
private buildTeamLeadPrompt(phase: OrchestratorPhase): string {
const taskList = phase.tasks.map((t, i) => `${i + 1}. ${t.prompt}`).join('\n');
return [
`You are the team lead for "${phase.name}".`,
`Create teammates and delegate the following tasks for parallel execution:`,
'',
taskList,
'',
`Each teammate should focus on one task area.`,
`When all tasks are complete, verify the results and output: <promise>${phase.id.toUpperCase()}_COMPLETE</promise>`,
].join('\n');
}
// ═══════════════════════════════════════════════════════════════
// Completion Phrases
// ═══════════════════════════════════════════════════════════════
private generateCompletionPhrases(phases: OrchestratorPhase[]): void {
for (const phase of phases) {
for (const task of phase.tasks) {
// Generate a unique, deterministic completion phrase per task
task.completionPhrase = `ORCH_P${phase.order + 1}_T${phase.tasks.indexOf(task) + 1}`;
}
}
}
// ═══════════════════════════════════════════════════════════════
// Helpers
// ═══════════════════════════════════════════════════════════════
private estimateComplexity(items: PlanItem[]): 'low' | 'medium' | 'high' {
const total = items.length;
const highComplexity = items.filter((i) => i.complexity === 'high').length;
const p0Count = items.filter((i) => i.priority === 'P0').length;
if (total > 20 || highComplexity > 5 || p0Count > 8) return 'high';
if (total > 10 || highComplexity > 2 || p0Count > 4) return 'medium';
return 'low';
}
}
+298
View File
@@ -0,0 +1,298 @@
/**
* @fileoverview Orchestrator phase verification.
*
* Runs verification checks after each phase completes:
* - Test commands (shell commands via session)
* - AI review (ask Claude to evaluate phase results)
*
* Three verification modes:
* - strict: ALL test commands must pass AND AI review must approve
* - moderate: Test commands must pass, AI review is advisory
* - lenient: At least one test command passes, AI review skipped
*
* Key exports:
* - `OrchestratorVerifier` class — phase verification engine
*
* @dependencies types (OrchestratorPhase, VerificationResult, VerificationCheck, OrchestratorConfig)
* @consumedby orchestrator-loop
*
* @module orchestrator-verifier
*/
import type { Session } from './session.js';
import {
getErrorMessage,
type OrchestratorPhase,
type OrchestratorConfig,
type VerificationResult,
type VerificationCheck,
} from './types.js';
// ═══════════════════════════════════════════════════════════════
// Constants
// ═══════════════════════════════════════════════════════════════
/** Timeout for individual test command execution (2 minutes) */
const TEST_COMMAND_TIMEOUT_MS = 2 * 60 * 1000;
/** Timeout for AI review (3 minutes) */
const AI_REVIEW_TIMEOUT_MS = 3 * 60 * 1000;
/** Completion phrase for AI verification pass */
const VERIFY_PASS_PHRASE = 'ORCH_VERIFY_PASS';
/** Completion phrase for AI verification fail */
const VERIFY_FAIL_PHRASE = 'ORCH_VERIFY_FAIL';
// ═══════════════════════════════════════════════════════════════
// OrchestratorVerifier
// ═══════════════════════════════════════════════════════════════
export class OrchestratorVerifier {
private config: OrchestratorConfig;
constructor(config: OrchestratorConfig) {
this.config = config;
}
/**
* Run all verification checks for a completed phase.
*
* @param phase - The phase to verify
* @param session - Session to use for running commands/reviews
* @returns Verification result with pass/fail and suggestions
*/
async verifyPhase(phase: OrchestratorPhase, session: Session): Promise<VerificationResult> {
const checks: VerificationCheck[] = [];
const mode = this.config.verificationMode;
// Skip verification entirely in lenient mode with no test commands
if (mode === 'lenient' && phase.testCommands.length === 0 && phase.verificationCriteria.length === 0) {
return {
passed: true,
checks: [],
summary: 'Verification skipped (lenient mode, no checks defined)',
suggestions: [],
};
}
// Run test commands if any are defined
if (phase.testCommands.length > 0) {
const testChecks = await this.runTestCommands(phase.testCommands, session);
checks.push(...testChecks);
}
// Run AI review in strict and moderate modes
if (mode !== 'lenient' && phase.verificationCriteria.length > 0) {
const aiCheck = await this.aiReview(phase, session);
checks.push(aiCheck);
}
// Determine pass/fail based on mode
const passed = this.evaluateChecks(checks, mode);
// Generate suggestions for failed checks
const suggestions = this.generateSuggestions(checks, phase);
const passedCount = checks.filter((c) => c.passed).length;
const summary =
checks.length === 0 ? 'No verification checks defined' : `${passedCount}/${checks.length} checks passed`;
return { passed, checks, summary, suggestions };
}
// ═══════════════════════════════════════════════════════════════
// Test Command Execution
// ═══════════════════════════════════════════════════════════════
private async runTestCommands(commands: string[], session: Session): Promise<VerificationCheck[]> {
const checks: VerificationCheck[] = [];
for (const command of commands) {
try {
const check = await this.runSingleTestCommand(command, session);
checks.push(check);
} catch (err) {
checks.push({
type: 'test_command',
description: `Run: ${command}`,
passed: false,
output: getErrorMessage(err),
});
}
}
return checks;
}
private async runSingleTestCommand(command: string, session: Session): Promise<VerificationCheck> {
// Send the test command to the session and wait for completion
// We use a unique marker to detect when the command finishes
const marker = `ORCH_TEST_${Date.now()}`;
const wrappedCommand = `${command} && echo ${marker}_PASS || echo ${marker}_FAIL`;
const result = await this.sendAndWaitForMarker(session, wrappedCommand, marker, TEST_COMMAND_TIMEOUT_MS);
return {
type: 'test_command',
description: `Run: ${command}`,
passed: result.includes(`${marker}_PASS`),
output: result.slice(0, 2000), // Truncate output
};
}
// ═══════════════════════════════════════════════════════════════
// AI Review
// ═══════════════════════════════════════════════════════════════
private async aiReview(phase: OrchestratorPhase, session: Session): Promise<VerificationCheck> {
const prompt = this.buildVerificationPrompt(phase);
try {
const result = await this.sendAndWaitForMarker(
session,
prompt,
VERIFY_PASS_PHRASE,
AI_REVIEW_TIMEOUT_MS,
VERIFY_FAIL_PHRASE
);
const passed = result.includes(VERIFY_PASS_PHRASE);
return {
type: 'ai_review',
description: `AI review of "${phase.name}"`,
passed,
output: result.slice(0, 3000),
};
} catch (err) {
return {
type: 'ai_review',
description: `AI review of "${phase.name}"`,
passed: false,
output: `AI review timed out or failed: ${getErrorMessage(err)}`,
};
}
}
private buildVerificationPrompt(phase: OrchestratorPhase): string {
const criteria = phase.verificationCriteria.map((c, i) => `${i + 1}. ${c}`).join('\n');
return [
`Review the work done in "${phase.name}". Check these criteria:`,
'',
criteria,
'',
`If ALL criteria are met, respond with: ${VERIFY_PASS_PHRASE}`,
`If ANY criteria fail, respond with: ${VERIFY_FAIL_PHRASE} and explain what failed.`,
].join('\n');
}
// ═══════════════════════════════════════════════════════════════
// Evaluation
// ═══════════════════════════════════════════════════════════════
private evaluateChecks(checks: VerificationCheck[], mode: OrchestratorConfig['verificationMode']): boolean {
if (checks.length === 0) return true;
const testChecks = checks.filter((c) => c.type === 'test_command');
const aiChecks = checks.filter((c) => c.type === 'ai_review');
switch (mode) {
case 'strict':
// ALL checks must pass
return checks.every((c) => c.passed);
case 'moderate':
// All test commands must pass; AI review is advisory
return testChecks.length === 0 || testChecks.every((c) => c.passed);
case 'lenient':
// At least one test passes (AI review skipped in lenient mode)
return testChecks.length === 0 || testChecks.some((c) => c.passed);
default:
return aiChecks.every((c) => c.passed) && testChecks.every((c) => c.passed);
}
}
private generateSuggestions(checks: VerificationCheck[], phase: OrchestratorPhase): string[] {
const suggestions: string[] = [];
const failedChecks = checks.filter((c) => !c.passed);
if (failedChecks.length === 0) return suggestions;
for (const check of failedChecks) {
if (check.type === 'test_command') {
suggestions.push(`Fix failing test: ${check.description}`);
} else if (check.type === 'ai_review' && check.output) {
// Extract failure reasons from AI review output
suggestions.push(`Address AI review feedback for "${phase.name}"`);
}
}
return suggestions;
}
// ═══════════════════════════════════════════════════════════════
// Session Communication
// ═══════════════════════════════════════════════════════════════
/**
* Send a prompt to a session and wait for a marker phrase in the output.
*
* @param session - Session to send to
* @param input - Prompt/command to send
* @param marker - Primary marker to watch for
* @param timeoutMs - Maximum wait time
* @param altMarker - Alternative marker (for pass/fail detection)
* @returns Captured output containing the marker
*/
private sendAndWaitForMarker(
session: Session,
input: string,
marker: string,
timeoutMs: number,
altMarker?: string
): Promise<string> {
return new Promise<string>((resolve, reject) => {
let output = '';
let resolved = false;
const timer = setTimeout(() => {
if (!resolved) {
resolved = true;
cleanup();
reject(new Error(`Timeout waiting for marker "${marker}" after ${timeoutMs}ms`));
}
}, timeoutMs);
const handler = (data: string) => {
if (resolved) return;
output += data;
if (output.includes(marker) || (altMarker && output.includes(altMarker))) {
resolved = true;
cleanup();
resolve(output);
}
};
const cleanup = () => {
clearTimeout(timer);
session.off('terminal', handler);
};
session.on('terminal', handler);
// Send the input
session.sendInput(input).catch((err) => {
if (!resolved) {
resolved = true;
cleanup();
reject(err);
}
});
});
}
}
+98 -92
View File
@@ -20,7 +20,7 @@ import type { TerminalMultiplexer } from './mux-interface.js';
import { existsSync, mkdirSync, writeFileSync } from 'node:fs';
import { join } from 'node:path';
import { RESEARCH_AGENT_PROMPT, PLANNER_PROMPT } from './prompts/index.js';
import type { PlanItem } from './types.js';
import { getErrorMessage, type PlanItem } from './types.js';
// Re-export for backward compatibility
export type { PlanItem };
@@ -68,7 +68,7 @@ export interface DetailedPlanResult {
export type ProgressCallback = (phase: string, detail: string) => void;
export interface PlanSubagentEvent {
interface PlanSubagentEvent {
type: 'started' | 'progress' | 'completed' | 'failed';
agentId: string;
agentType: 'research' | 'planner';
@@ -80,7 +80,7 @@ export interface PlanSubagentEvent {
error?: string;
}
export type SubagentCallback = (event: PlanSubagentEvent) => void;
type SubagentCallback = (event: PlanSubagentEvent) => void;
// ============================================================================
// JSON Repair Helper
@@ -231,6 +231,49 @@ export class PlanOrchestrator {
return md;
}
private _extractJsonFromResponse(response: string): string | null {
let jsonMatch = response.match(/```(?:json)?\s*(\{[\s\S]*?\})\s*```/);
if (jsonMatch) {
jsonMatch = [jsonMatch[1]]; // Use captured group (inside code block)
} else {
jsonMatch = response.match(/\{[\s\S]*\}/);
}
return jsonMatch ? jsonMatch[0] : null;
}
private _emitAgentFailure(
onSubagent: SubagentCallback | undefined,
agentId: string,
agentType: 'research' | 'planner',
model: string,
error: string,
durationMs: number
): void {
onSubagent?.({
type: 'failed',
agentId,
agentType,
model,
status: 'failed',
error,
durationMs,
});
}
private _formatResearchSection(
parts: string[],
title: string,
items: unknown[],
formatter: (item: unknown) => string[]
): void {
if (items.length === 0) return;
parts.push(title);
for (const item of items.slice(0, 5)) {
parts.push(...formatter(item));
}
parts.push('');
}
async cancel(): Promise<void> {
this.cancelled = true;
// Stop all running sessions and await cleanup to prevent PTY process leaks
@@ -312,7 +355,7 @@ export class PlanOrchestrator {
} catch (err) {
return {
success: false,
error: err instanceof Error ? err.message : String(err),
error: getErrorMessage(err),
};
}
}
@@ -322,32 +365,23 @@ export class PlanOrchestrator {
const parts: string[] = ['## Research Context\n'];
if (research.findings.externalResources.length > 0) {
parts.push('### External Resources');
for (const r of research.findings.externalResources.slice(0, 5)) {
parts.push(`- ${r.title}${r.url ? ` (${r.url})` : ''}`);
this._formatResearchSection(parts, '### External Resources', research.findings.externalResources, (item) => {
const r = item as ResearchResult['findings']['externalResources'][number];
const lines = [`- ${r.title}${r.url ? ` (${r.url})` : ''}`];
if (r.keyInsights.length > 0) {
parts.push(` Key insights: ${r.keyInsights.slice(0, 3).join(', ')}`);
}
}
parts.push('');
lines.push(` Key insights: ${r.keyInsights.slice(0, 3).join(', ')}`);
}
return lines;
});
if (research.findings.codebasePatterns.length > 0) {
parts.push('### Existing Codebase Patterns');
for (const p of research.findings.codebasePatterns.slice(0, 5)) {
parts.push(`- ${p.pattern} at ${p.location}`);
}
parts.push('');
}
this._formatResearchSection(parts, '### Existing Codebase Patterns', research.findings.codebasePatterns, (item) => {
const p = item as ResearchResult['findings']['codebasePatterns'][number];
return [`- ${p.pattern} at ${p.location}`];
});
if (research.findings.technicalRecommendations.length > 0) {
parts.push('### Recommendations');
for (const r of research.findings.technicalRecommendations.slice(0, 5)) {
parts.push(`- ${r}`);
}
parts.push('');
}
this._formatResearchSection(parts, '### Recommendations', research.findings.technicalRecommendations, (item) => [
`- ${item as string}`,
]);
return parts.join('\n');
}
@@ -414,18 +448,20 @@ export class PlanOrchestrator {
const durationMs = Date.now() - startTime;
// Extract JSON from response
const jsonMatch = response.match(/\{[\s\S]*\}/);
if (!jsonMatch) {
onSubagent?.({
type: 'failed',
agentId,
agentType: 'research',
model: this.researchModel,
status: 'failed',
error: 'No JSON found',
durationMs,
});
console.log(
`[PlanOrchestrator] Research response length: ${response.length}, first 500 chars:`,
response.substring(0, 500)
);
// Extract JSON from response — try multiple strategies
const jsonStr = this._extractJsonFromResponse(response);
if (!jsonStr) {
console.error(
`[PlanOrchestrator] No JSON found in research response. Full response:`,
response.substring(0, 2000)
);
this._emitAgentFailure(onSubagent, agentId, 'research', this.researchModel, 'No JSON found', durationMs);
return {
success: false,
findings: {
@@ -441,17 +477,9 @@ export class PlanOrchestrator {
};
}
const parsed = tryParseJSON(jsonMatch[0]);
const parsed = tryParseJSON(jsonStr);
if (!parsed.success) {
onSubagent?.({
type: 'failed',
agentId,
agentType: 'research',
model: this.researchModel,
status: 'failed',
error: parsed.error,
durationMs,
});
this._emitAgentFailure(onSubagent, agentId, 'research', this.researchModel, parsed.error!, durationMs);
return {
success: false,
findings: {
@@ -495,16 +523,8 @@ export class PlanOrchestrator {
return result;
} catch (err) {
const durationMs = Date.now() - startTime;
const error = err instanceof Error ? err.message : String(err);
onSubagent?.({
type: 'failed',
agentId,
agentType: 'research',
model: this.researchModel,
status: 'failed',
error,
durationMs,
});
const error = getErrorMessage(err);
this._emitAgentFailure(onSubagent, agentId, 'research', this.researchModel, error, durationMs);
return {
success: false,
findings: {
@@ -521,7 +541,7 @@ export class PlanOrchestrator {
} finally {
// Always clean up session and progress interval — centralizing here
// prevents the race where cancel() and catch both try to manage the set
await session.stop().catch(() => {});
await session.stop().catch(() => {}); // Ignore - session cleanup is best-effort in finally block
this.runningSessions.delete(session);
clearInterval(progressInterval);
}
@@ -587,32 +607,26 @@ export class PlanOrchestrator {
const durationMs = Date.now() - startTime;
// Extract JSON from response
const jsonMatch = response.match(/\{[\s\S]*\}/);
if (!jsonMatch) {
onSubagent?.({
type: 'failed',
agentId,
agentType: 'planner',
model: this.plannerModel,
status: 'failed',
error: 'No JSON found',
durationMs,
});
console.log(
`[PlanOrchestrator] Planner response length: ${response.length}, first 500 chars:`,
response.substring(0, 500)
);
// Extract JSON from response — try multiple strategies
const jsonStr = this._extractJsonFromResponse(response);
if (!jsonStr) {
console.error(
`[PlanOrchestrator] No JSON found in planner response. Full response:`,
response.substring(0, 2000)
);
this._emitAgentFailure(onSubagent, agentId, 'planner', this.plannerModel, 'No JSON found', durationMs);
return { success: false, error: 'No JSON in response' };
}
const parsed = tryParseJSON(jsonMatch[0]);
const parsed = tryParseJSON(jsonStr);
if (!parsed.success) {
onSubagent?.({
type: 'failed',
agentId,
agentType: 'planner',
model: this.plannerModel,
status: 'failed',
error: parsed.error,
durationMs,
});
this._emitAgentFailure(onSubagent, agentId, 'planner', this.plannerModel, parsed.error!, durationMs);
return { success: false, error: parsed.error };
}
@@ -637,21 +651,13 @@ export class PlanOrchestrator {
return { success: true, items, gaps, warnings };
} catch (err) {
const durationMs = Date.now() - startTime;
const error = err instanceof Error ? err.message : String(err);
onSubagent?.({
type: 'failed',
agentId,
agentType: 'planner',
model: this.plannerModel,
status: 'failed',
error,
durationMs,
});
const error = getErrorMessage(err);
this._emitAgentFailure(onSubagent, agentId, 'planner', this.plannerModel, error, durationMs);
return { success: false, error };
} finally {
// Always clean up session and progress interval — centralizing here
// prevents the race where cancel() and catch both try to manage the set
await session.stop().catch(() => {});
await session.stop().catch(() => {}); // Ignore - session cleanup is best-effort in finally block
this.runningSessions.delete(session);
clearInterval(progressInterval);
}
+1
View File
@@ -7,3 +7,4 @@
export { RESEARCH_AGENT_PROMPT } from './research-agent.js';
export { PLANNER_PROMPT } from './planner.js';
export { PHASE_EXECUTION_PROMPT, TEAM_LEAD_PROMPT, REPLAN_PROMPT, SINGLE_TASK_PROMPT } from './orchestrator.js';
+100
View File
@@ -0,0 +1,100 @@
/**
* @fileoverview Orchestrator Loop prompt templates.
*
* Templates for phase execution, team delegation, verification, and replanning.
* Placeholders use {VARIABLE} syntax and are replaced at runtime.
*
* @module prompts/orchestrator
*/
/**
* Phase execution prompt — tells Claude what to accomplish in this phase.
*
* Placeholders:
* - {PHASE_NUMBER}: Phase index (1-based)
* - {PHASE_NAME}: Human-readable phase name
* - {GOAL}: Original user goal
* - {COMPLETED_PHASES}: Summary of previously completed phases
* - {TASK_LIST}: Numbered task list for this phase
* - {VERIFICATION_CRITERIA}: What will be checked after this phase
* - {COMPLETION_PHRASE}: The phrase to output when done
*/
export const PHASE_EXECUTION_PROMPT = `You are executing {PHASE_NAME} of a larger project.
OVERALL GOAL: {GOAL}
COMPLETED SO FAR:
{COMPLETED_PHASES}
YOUR TASKS FOR THIS PHASE:
{TASK_LIST}
Complete each task thoroughly. Run tests after each change to catch issues early.
VERIFICATION (will be checked after you finish):
{VERIFICATION_CRITERIA}
When ALL tasks in this phase are complete and verified, output: <promise>{COMPLETION_PHRASE}</promise>`;
/**
* Team lead delegation prompt — instructs a lead to coordinate teammates.
*
* Placeholders:
* - {PHASE_NAME}: Phase name
* - {TASK_LIST}: Numbered task list
* - {TEAMMATE_HINTS}: Suggested teammate specializations
* - {COMPLETION_PHRASE}: Phrase for when all work is done
*/
export const TEAM_LEAD_PROMPT = `You are the team lead for {PHASE_NAME}.
Create teammates and delegate the following tasks for parallel execution:
{TASK_LIST}
Suggested teammate roles:
{TEAMMATE_HINTS}
Each teammate should focus on their assigned task area. Monitor their progress.
When ALL tasks are complete and you've verified the results, output: <promise>{COMPLETION_PHRASE}</promise>`;
/**
* Replan prompt — gives failure context and asks for recovery.
*
* Placeholders:
* - {PHASE_NAME}: Phase name
* - {ATTEMPT_NUMBER}: Current retry attempt
* - {MAX_ATTEMPTS}: Maximum attempts allowed
* - {FAILURE_SUMMARY}: What went wrong
* - {SUGGESTIONS}: Recovery suggestions from verification
* - {ORIGINAL_TASKS}: The original task list
* - {COMPLETION_PHRASE}: Phrase for when recovery is done
*/
export const REPLAN_PROMPT = `Phase "{PHASE_NAME}" verification failed (attempt {ATTEMPT_NUMBER}/{MAX_ATTEMPTS}).
WHAT WENT WRONG:
{FAILURE_SUMMARY}
SUGGESTIONS:
{SUGGESTIONS}
ORIGINAL TASKS:
{ORIGINAL_TASKS}
Fix the issues identified above. Focus on making the verification criteria pass.
When the fixes are complete, output: <promise>{COMPLETION_PHRASE}</promise>`;
/**
* Single-task execution prompt — for phases with a single task.
*
* Placeholders:
* - {TASK}: The task description
* - {GOAL}: Original user goal
* - {CONTEXT}: Any relevant context
* - {COMPLETION_PHRASE}: Phrase for when done
*/
export const SINGLE_TASK_PROMPT = `{TASK}
Context: This is part of a larger project — {GOAL}
{CONTEXT}
When done, output: <promise>{COMPLETION_PHRASE}</promise>`;
+4 -5
View File
@@ -9,6 +9,7 @@
import { existsSync, readFileSync } from 'node:fs';
import { join } from 'node:path';
import { execPattern } from './utils/index.js';
// Pattern to extract completion phrase from CLAUDE.md
// Matches <promise>PHRASE</promise> with optional whitespace
@@ -23,7 +24,7 @@ const YAML_LINE_PATTERN = /^([a-zA-Z_-]+):\s*"?([^"\n]+)"?\s*$/gm;
/**
* Ralph Loop configuration from .claude/ralph-loop.local.md
*/
export interface RalphLoopConfig {
interface RalphLoopConfig {
enabled: boolean;
iteration: number;
maxIterations: number | null;
@@ -83,9 +84,7 @@ export function parseRalphLoopConfigFromContent(content: string): RalphLoopConfi
};
// Parse each YAML line
let match;
YAML_LINE_PATTERN.lastIndex = 0;
while ((match = YAML_LINE_PATTERN.exec(yaml)) !== null) {
execPattern(YAML_LINE_PATTERN, yaml, (match) => {
const key = match[1].toLowerCase();
const value = match[2].trim();
@@ -103,7 +102,7 @@ export function parseRalphLoopConfigFromContent(content: string): RalphLoopConfi
config.completionPromise = value.toUpperCase();
break;
}
}
});
return config;
}
+1 -9
View File
@@ -34,19 +34,11 @@ import { RalphLoopStatus, getErrorMessage } from './types.js';
/**
* Events emitted by RalphLoop
*/
export interface RalphLoopEvents {
started: () => void;
stopped: () => void;
taskAssigned: (taskId: string, sessionId: string) => void;
taskCompleted: (taskId: string) => void;
taskFailed: (taskId: string, error: string) => void;
error: (error: Error) => void;
}
/**
* Configuration options for RalphLoop
*/
export interface RalphLoopOptions {
interface RalphLoopOptions {
/** How often to check for new tasks (default from config) */
pollIntervalMs?: number;
/** Minimum time to run before stopping (null = no minimum) */
+2 -1
View File
@@ -10,6 +10,7 @@
*/
import { EventEmitter } from 'node:events';
import { CLEANUP_CHECK_INTERVAL_MS } from './config/server-timing.js';
/**
* RalphStallDetector - Detects iteration stalls in the Ralph loop.
@@ -57,7 +58,7 @@ export class RalphStallDetector extends EventEmitter {
// Check every minute
this._iterationStallTimer = setInterval(() => {
this.checkIterationStall();
}, 60 * 1000);
}, CLEANUP_CHECK_INTERVAL_MS);
}
/**
+106 -104
View File
@@ -91,6 +91,67 @@ const COMPLETION_INDICATOR_PATTERNS = [
/project\s+(?:is\s+)?(?:completed?|done|finished)/i,
];
interface FieldParser<T> {
pattern: RegExp;
field: keyof RalphStatusBlock;
validate: (value: string) => boolean;
transform: (value: string) => T;
errorMsg: (value: string) => string;
}
const FIELD_PARSERS: FieldParser<RalphStatusValue | RalphTestsStatus | RalphWorkType | number | boolean | string>[] = [
{
pattern: RALPH_STATUS_FIELD_PATTERN,
field: 'status',
validate: (v) => ['IN_PROGRESS', 'COMPLETE', 'BLOCKED'].includes(v.toUpperCase()),
transform: (v) => v.toUpperCase() as RalphStatusValue,
errorMsg: (v) => `Invalid STATUS value: "${v}". Expected: IN_PROGRESS, COMPLETE, or BLOCKED`,
},
{
pattern: RALPH_TASKS_COMPLETED_PATTERN,
field: 'tasksCompletedThisLoop',
validate: (v) => !Number.isNaN(parseInt(v, 10)) && parseInt(v, 10) >= 0,
transform: (v) => parseInt(v, 10),
errorMsg: (v) => `Invalid TASKS_COMPLETED_THIS_LOOP value: "${v}". Expected: non-negative integer`,
},
{
pattern: RALPH_FILES_MODIFIED_PATTERN,
field: 'filesModified',
validate: (v) => !Number.isNaN(parseInt(v, 10)) && parseInt(v, 10) >= 0,
transform: (v) => parseInt(v, 10),
errorMsg: (v) => `Invalid FILES_MODIFIED value: "${v}". Expected: non-negative integer`,
},
{
pattern: RALPH_TESTS_STATUS_PATTERN,
field: 'testsStatus',
validate: (v) => ['PASSING', 'FAILING', 'NOT_RUN'].includes(v.toUpperCase()),
transform: (v) => v.toUpperCase() as RalphTestsStatus,
errorMsg: (v) => `Invalid TESTS_STATUS value: "${v}". Expected: PASSING, FAILING, or NOT_RUN`,
},
{
pattern: RALPH_WORK_TYPE_PATTERN,
field: 'workType',
validate: (v) => ['IMPLEMENTATION', 'TESTING', 'DOCUMENTATION', 'REFACTORING'].includes(v.toUpperCase()),
transform: (v) => v.toUpperCase() as RalphWorkType,
errorMsg: (v) =>
`Invalid WORK_TYPE value: "${v}". Expected: IMPLEMENTATION, TESTING, DOCUMENTATION, or REFACTORING`,
},
{
pattern: RALPH_EXIT_SIGNAL_PATTERN,
field: 'exitSignal',
validate: () => true,
transform: (v) => v.toLowerCase() === 'true',
errorMsg: () => '',
},
{
pattern: RALPH_RECOMMENDATION_PATTERN,
field: 'recommendation',
validate: () => true,
transform: (v) => v.trim(),
errorMsg: () => '',
},
];
/**
* RalphStatusParser - Parses RALPH_STATUS blocks and manages circuit breaker.
*
@@ -303,85 +364,21 @@ export class RalphStatusParser extends EventEmitter {
const trimmedLine = line.trim();
if (!trimmedLine) continue;
// Track whether this line matched any known field
let matched = false;
// STATUS field (required)
const statusMatch = trimmedLine.match(RALPH_STATUS_FIELD_PATTERN);
if (statusMatch) {
const value = statusMatch[1].toUpperCase();
if (['IN_PROGRESS', 'COMPLETE', 'BLOCKED'].includes(value)) {
block.status = value as RalphStatusValue;
for (const parser of FIELD_PARSERS) {
const match = trimmedLine.match(parser.pattern);
if (match) {
const rawValue = match[1];
if (parser.validate(rawValue)) {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
(block as any)[parser.field] = parser.transform(rawValue);
} else {
parseErrors.push(`Invalid STATUS value: "${value}". Expected: IN_PROGRESS, COMPLETE, or BLOCKED`);
parseErrors.push(parser.errorMsg(rawValue));
}
matched = true;
break;
}
// TASKS_COMPLETED_THIS_LOOP field
const tasksMatch = trimmedLine.match(RALPH_TASKS_COMPLETED_PATTERN);
if (tasksMatch) {
const value = parseInt(tasksMatch[1], 10);
if (!Number.isNaN(value) && value >= 0) {
block.tasksCompletedThisLoop = value;
} else {
parseErrors.push(
`Invalid TASKS_COMPLETED_THIS_LOOP value: "${tasksMatch[1]}". Expected: non-negative integer`
);
}
matched = true;
}
// FILES_MODIFIED field
const filesMatch = trimmedLine.match(RALPH_FILES_MODIFIED_PATTERN);
if (filesMatch) {
const value = parseInt(filesMatch[1], 10);
if (!Number.isNaN(value) && value >= 0) {
block.filesModified = value;
} else {
parseErrors.push(`Invalid FILES_MODIFIED value: "${filesMatch[1]}". Expected: non-negative integer`);
}
matched = true;
}
// TESTS_STATUS field
const testsMatch = trimmedLine.match(RALPH_TESTS_STATUS_PATTERN);
if (testsMatch) {
const value = testsMatch[1].toUpperCase();
if (['PASSING', 'FAILING', 'NOT_RUN'].includes(value)) {
block.testsStatus = value as RalphTestsStatus;
} else {
parseErrors.push(`Invalid TESTS_STATUS value: "${value}". Expected: PASSING, FAILING, or NOT_RUN`);
}
matched = true;
}
// WORK_TYPE field
const workMatch = trimmedLine.match(RALPH_WORK_TYPE_PATTERN);
if (workMatch) {
const value = workMatch[1].toUpperCase();
if (['IMPLEMENTATION', 'TESTING', 'DOCUMENTATION', 'REFACTORING'].includes(value)) {
block.workType = value as RalphWorkType;
} else {
parseErrors.push(
`Invalid WORK_TYPE value: "${value}". Expected: IMPLEMENTATION, TESTING, DOCUMENTATION, or REFACTORING`
);
}
matched = true;
}
// EXIT_SIGNAL field
const exitMatch = trimmedLine.match(RALPH_EXIT_SIGNAL_PATTERN);
if (exitMatch) {
block.exitSignal = exitMatch[1].toLowerCase() === 'true';
matched = true;
}
// RECOMMENDATION field
const recMatch = trimmedLine.match(RALPH_RECOMMENDATION_PATTERN);
if (recMatch) {
block.recommendation = recMatch[1].trim();
matched = true;
}
// Track unknown fields for debugging (only if looks like a field)
@@ -475,38 +472,9 @@ export class RalphStatusParser extends EventEmitter {
const prevState = this._circuitBreaker.state;
if (hasProgress) {
// Progress detected - reset counters, possibly close circuit
this._circuitBreaker.consecutiveNoProgress = 0;
this._circuitBreaker.consecutiveSameError = 0;
this._circuitBreaker.lastProgressIteration = this._cycleCount;
if (this._circuitBreaker.state === 'HALF_OPEN') {
this._circuitBreaker.state = 'CLOSED';
this._circuitBreaker.reason = 'Progress detected, circuit closed';
this._circuitBreaker.reasonCode = 'progress_detected';
}
this._handleProgressDetected();
} else {
// No progress
this._circuitBreaker.consecutiveNoProgress++;
// State transitions based on consecutive no-progress
if (this._circuitBreaker.state === 'CLOSED') {
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = `No progress for ${this._circuitBreaker.consecutiveNoProgress} iterations`;
this._circuitBreaker.reasonCode = 'no_progress_open';
} else if (this._circuitBreaker.consecutiveNoProgress >= 2) {
this._circuitBreaker.state = 'HALF_OPEN';
this._circuitBreaker.reason = 'Warning: no progress detected';
this._circuitBreaker.reasonCode = 'no_progress_warning';
}
} else if (this._circuitBreaker.state === 'HALF_OPEN') {
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = `No progress for ${this._circuitBreaker.consecutiveNoProgress} iterations`;
this._circuitBreaker.reasonCode = 'no_progress_open';
}
}
this._handleNoProgress();
}
// Track tests failure
@@ -535,6 +503,40 @@ export class RalphStatusParser extends EventEmitter {
}
}
private _handleProgressDetected(): void {
this._circuitBreaker.consecutiveNoProgress = 0;
this._circuitBreaker.consecutiveSameError = 0;
this._circuitBreaker.lastProgressIteration = this._cycleCount;
if (this._circuitBreaker.state === 'HALF_OPEN') {
this._circuitBreaker.state = 'CLOSED';
this._circuitBreaker.reason = 'Progress detected, circuit closed';
this._circuitBreaker.reasonCode = 'progress_detected';
}
}
private _handleNoProgress(): void {
this._circuitBreaker.consecutiveNoProgress++;
if (this._circuitBreaker.state === 'CLOSED') {
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = `No progress for ${this._circuitBreaker.consecutiveNoProgress} iterations`;
this._circuitBreaker.reasonCode = 'no_progress_open';
} else if (this._circuitBreaker.consecutiveNoProgress >= 2) {
this._circuitBreaker.state = 'HALF_OPEN';
this._circuitBreaker.reason = 'Warning: no progress detected';
this._circuitBreaker.reasonCode = 'no_progress_warning';
}
} else if (this._circuitBreaker.state === 'HALF_OPEN') {
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = `No progress for ${this._circuitBreaker.consecutiveNoProgress} iterations`;
this._circuitBreaker.reasonCode = 'no_progress_open';
}
}
}
/**
* Check line for completion indicators (natural language patterns).
* Used for dual-condition exit gate.
+70 -112
View File
@@ -19,7 +19,6 @@
* Key exports:
* - `RalphTracker` class — main tracker, extends EventEmitter
* - `RalphTrackerEvents` interface — typed event map
* - Re-exports: `EnhancedPlanTask`, `CheckpointReview` from ralph-plan-tracker
*
* Key methods: `processData(data)` — feed terminal output, `getState()`,
* `getTodos()`, `getCompletionHistory()`, `getPlanTasks()`, `reset()`
@@ -55,6 +54,7 @@ import {
stringSimilarity,
Debouncer,
CleanupManager,
execPattern,
} from './utils/index.js';
import { MAX_LINE_BUFFER_SIZE } from './config/buffer-limits.js';
import { MAX_TODOS_PER_SESSION } from './config/map-limits.js';
@@ -63,18 +63,16 @@ import type { EnhancedPlanTask, CheckpointReview } from './ralph-plan-tracker.js
import { RalphFixPlanWatcher, generateFixPlanMarkdown, importFixPlanMarkdown } from './ralph-fix-plan-watcher.js';
import { RalphStallDetector } from './ralph-stall-detector.js';
import { RalphStatusParser } from './ralph-status-parser.js';
// Re-export sub-module types for backward compatibility
export type { EnhancedPlanTask, CheckpointReview } from './ralph-plan-tracker.js';
import { STALE_DATA_MAX_AGE_MS, INACTIVITY_TIMEOUT_MS } from './config/server-timing.js';
// ========== Configuration Constants ==========
// Note: MAX_TODOS_PER_SESSION and MAX_LINE_BUFFER_SIZE are imported from config modules
/**
* Todo items older than this duration (in milliseconds) will be auto-expired.
* Default: 1 hour (60 * 60 * 1000)
* Default: 1 hour
*/
const TODO_EXPIRY_MS = 60 * 60 * 1000;
const TODO_EXPIRY_MS = STALE_DATA_MAX_AGE_MS;
/**
* Minimum interval between on-demand cleanup checks (in milliseconds).
@@ -88,7 +86,7 @@ const CLEANUP_THROTTLE_MS = 30 * 1000;
* Actively purges expired todos even when no terminal data is flowing.
* Default: 5 minutes
*/
const TODO_CLEANUP_INTERVAL_MS = 5 * 60 * 1000;
const TODO_CLEANUP_INTERVAL_MS = INACTIVITY_TIMEOUT_MS;
/**
* Similarity threshold for todo deduplication.
@@ -98,6 +96,18 @@ const TODO_CLEANUP_INTERVAL_MS = 5 * 60 * 1000;
*/
const TODO_SIMILARITY_THRESHOLD = 0.85;
/**
* Similarity threshold for short todo content (<30 chars).
* Higher threshold reduces false positive deduplication of short strings.
*/
const SIMILARITY_THRESHOLD_SHORT = 0.95;
/**
* Similarity threshold for medium-length todo content (30-60 chars).
* Slightly relaxed compared to short strings.
*/
const SIMILARITY_THRESHOLD_MEDIUM = 0.9;
/**
* Debounce interval for event emissions (milliseconds).
* Prevents UI jitter from rapid consecutive updates.
@@ -376,32 +386,6 @@ const P2_PRIORITY_PATTERNS = [
* @event circuitBreakerUpdate - Fired when circuit breaker state changes
* @event exitGateMet - Fired when dual-condition exit gate is met
*/
export interface RalphTrackerEvents {
/** Emitted when loop state changes */
loopUpdate: (state: RalphTrackerState) => void;
/** Emitted when todo list is modified */
todoUpdate: (todos: RalphTodoItem[]) => void;
/** Emitted when completion phrase detected (loop finished) */
completionDetected: (phrase: string) => void;
/** Emitted when tracker auto-enables from disabled state */
enabled: () => void;
/** Emitted when a RALPH_STATUS block is parsed */
statusBlockDetected: (block: RalphStatusBlock) => void;
/** Emitted when circuit breaker state changes */
circuitBreakerUpdate: (status: CircuitBreakerStatus) => void;
/** Emitted when dual-condition exit gate is met (completion indicators >= 2 AND EXIT_SIGNAL: true) */
exitGateMet: (data: { completionIndicators: number; exitSignal: boolean }) => void;
/** Emitted when iteration count hasn't changed for an extended period (stall warning) */
iterationStallWarning: (data: { iteration: number; stallDurationMs: number }) => void;
/** Emitted when iteration count hasn't changed for critical period (stall critical) */
iterationStallCritical: (data: { iteration: number; stallDurationMs: number }) => void;
/** Emitted when a common/risky completion phrase is detected (P1-002) */
phraseValidationWarning: (data: {
phrase: string;
reason: 'common' | 'short' | 'numeric';
suggestedPhrase: string;
}) => void;
}
/**
* RalphTracker - Parses terminal output to detect Ralph Wiggum loops and todos
@@ -1297,6 +1281,24 @@ export class RalphTracker extends EventEmitter {
this.detectTodoItems(trimmed);
}
/**
* Mark all tracked todos as completed and emit todoUpdate if any changed.
* @returns true if any todo was updated
*/
private completeAllTodos(): boolean {
let updated = false;
for (const todo of this._todos.values()) {
if (todo.status !== 'completed') {
todo.status = 'completed';
updated = true;
}
}
if (updated) {
this.emit('todoUpdate', this.todos);
}
return updated;
}
/**
* Detect "all tasks complete" messages.
*/
@@ -1316,16 +1318,7 @@ export class RalphTracker extends EventEmitter {
return;
}
let updated = false;
for (const todo of this._todos.values()) {
if (todo.status !== 'completed') {
todo.status = 'completed';
updated = true;
}
}
if (updated) {
this.emit('todoUpdate', this.todos);
}
this.completeAllTodos();
if (this._loopState.completionPhrase) {
this._loopState.active = false;
@@ -1423,16 +1416,7 @@ export class RalphTracker extends EventEmitter {
if (bareCount > 1) return;
let updated = false;
for (const todo of this._todos.values()) {
if (todo.status !== 'completed') {
todo.status = 'completed';
updated = true;
}
}
if (updated) {
this.emit('todoUpdate', this.todos);
}
this.completeAllTodos();
this._loopState.active = false;
this._loopState.lastActivity = Date.now();
@@ -1478,16 +1462,7 @@ export class RalphTracker extends EventEmitter {
if (canonicalCount >= 2 || this._loopState.active) {
this._loopState.active = false;
this._loopState.lastActivity = Date.now();
let updated = false;
for (const todo of this._todos.values()) {
if (todo.status !== 'completed') {
todo.status = 'completed';
updated = true;
}
}
if (updated) {
this.emit('todoUpdate', this.todos);
}
this.completeAllTodos();
this.emit('completionDetected', matchedPhrase);
this.emit('loopUpdate', this.loopState);
return;
@@ -1495,16 +1470,7 @@ export class RalphTracker extends EventEmitter {
}
if (this._loopState.active || count >= 2) {
let updated = false;
for (const todo of this._todos.values()) {
if (todo.status !== 'completed') {
todo.status = 'completed';
updated = true;
}
}
if (updated) {
this.emit('todoUpdate', this.todos);
}
this.completeAllTodos();
this._loopState.active = false;
this._loopState.lastActivity = Date.now();
@@ -1530,40 +1496,38 @@ export class RalphTracker extends EventEmitter {
const suggestedPhrase = `${phrase}_${uniqueSuffix}`;
if (COMMON_COMPLETION_PHRASES.has(normalized)) {
console.warn(
`[RalphTracker] Warning: Completion phrase "${phrase}" is very common and may cause false positives. Consider using: "${suggestedPhrase}"`
);
this.emit('phraseValidationWarning', {
phrase,
reason: 'common',
suggestedPhrase,
});
this.emitValidationWarning(phrase, 'common', suggestedPhrase);
return;
}
if (normalized.length < MIN_RECOMMENDED_PHRASE_LENGTH) {
console.warn(
`[RalphTracker] Warning: Completion phrase "${phrase}" is too short (${normalized.length} chars). Consider using: "${suggestedPhrase}"`
);
this.emit('phraseValidationWarning', {
phrase,
reason: 'short',
suggestedPhrase,
});
this.emitValidationWarning(phrase, 'short', suggestedPhrase);
return;
}
if (/^\d+$/.test(normalized)) {
this.emitValidationWarning(phrase, 'numeric', suggestedPhrase);
}
}
/**
* Emit a phrase validation warning with a console message and event.
*/
private emitValidationWarning(phrase: string, reason: 'common' | 'short' | 'numeric', suggestedPhrase: string): void {
const descriptions: Record<'common' | 'short' | 'numeric', string> = {
common: 'is very common and may cause false positives',
short: `is too short (${phrase.toUpperCase().replace(/[\s_\-.]+/g, '').length} chars)`,
numeric: 'is numeric-only and may cause false positives',
};
console.warn(
`[RalphTracker] Warning: Completion phrase "${phrase}" is numeric-only and may cause false positives. Consider using: "${suggestedPhrase}"`
`[RalphTracker] Warning: Completion phrase "${phrase}" ${descriptions[reason]}. Consider using: "${suggestedPhrase}"`
);
this.emit('phraseValidationWarning', {
phrase,
reason: 'numeric',
reason,
suggestedPhrase,
});
}
}
/**
* Activate the loop if not already active.
@@ -1666,35 +1630,32 @@ export class RalphTracker extends EventEmitter {
let match: RegExpExecArray | null;
if (hasCheckbox) {
TODO_CHECKBOX_PATTERN.lastIndex = 0;
while ((match = TODO_CHECKBOX_PATTERN.exec(line)) !== null) {
execPattern(TODO_CHECKBOX_PATTERN, line, (match) => {
const checked = match[1].toLowerCase() === 'x';
const content = match[2].trim();
const status: RalphTodoStatus = checked ? 'completed' : 'pending';
this.upsertTodo(content, status);
updated = true;
}
});
}
if (hasTodoIndicator) {
TODO_INDICATOR_PATTERN.lastIndex = 0;
while ((match = TODO_INDICATOR_PATTERN.exec(line)) !== null) {
execPattern(TODO_INDICATOR_PATTERN, line, (match) => {
const icon = match[1];
const content = match[2].trim();
const status = this.iconToStatus(icon);
this.upsertTodo(content, status);
updated = true;
}
});
}
if (hasStatus) {
TODO_STATUS_PATTERN.lastIndex = 0;
while ((match = TODO_STATUS_PATTERN.exec(line)) !== null) {
execPattern(TODO_STATUS_PATTERN, line, (match) => {
const content = match[1].trim();
const status = match[2] as RalphTodoStatus;
this.upsertTodo(content, status);
updated = true;
}
});
}
if (hasNativeCheckbox) {
@@ -1715,8 +1676,7 @@ export class RalphTracker extends EventEmitter {
}
if (hasCheckmark) {
TODO_TASK_CREATED_PATTERN.lastIndex = 0;
while ((match = TODO_TASK_CREATED_PATTERN.exec(line)) !== null) {
execPattern(TODO_TASK_CREATED_PATTERN, line, (match) => {
const taskNum = parseInt(match[1], 10);
const content = match[2].trim();
if (content.length >= 5) {
@@ -1725,10 +1685,9 @@ export class RalphTracker extends EventEmitter {
this.upsertTodo(content, 'pending');
updated = true;
}
}
});
TODO_TASK_SUMMARY_PATTERN.lastIndex = 0;
while ((match = TODO_TASK_SUMMARY_PATTERN.exec(line)) !== null) {
execPattern(TODO_TASK_SUMMARY_PATTERN, line, (match) => {
const taskNum = parseInt(match[1], 10);
const content = match[2].trim();
if (content.length >= 5) {
@@ -1739,10 +1698,9 @@ export class RalphTracker extends EventEmitter {
this.upsertTodo(this._taskNumberToContent.get(taskNum) || content, 'pending');
updated = true;
}
}
});
TODO_TASK_STATUS_PATTERN.lastIndex = 0;
while ((match = TODO_TASK_STATUS_PATTERN.exec(line)) !== null) {
execPattern(TODO_TASK_STATUS_PATTERN, line, (match) => {
const taskNum = parseInt(match[1], 10);
const statusStr = match[2].trim();
const status: RalphTodoStatus =
@@ -1752,7 +1710,7 @@ export class RalphTracker extends EventEmitter {
this.upsertTodo(content, status);
updated = true;
}
}
});
if (!updated) {
TODO_PLAIN_CHECKMARK_PATTERN.lastIndex = 0;
@@ -1981,9 +1939,9 @@ export class RalphTracker extends EventEmitter {
let threshold: number;
if (normalized.length < 30) {
threshold = 0.95;
threshold = SIMILARITY_THRESHOLD_SHORT;
} else if (normalized.length < 60) {
threshold = 0.9;
threshold = SIMILARITY_THRESHOLD_MEDIUM;
} else {
threshold = TODO_SIMILARITY_THRESHOLD;
}
+185 -161
View File
@@ -27,7 +27,7 @@
* - `RespawnController` class — state machine, extends EventEmitter
* - `RespawnConfig` interface — all configuration options
* - `RespawnState` type — union of all state machine states
* - `DetectionStatus`, `ActiveTimerInfo`, `RespawnEvents` — status/event types
* - `DetectionStatus`, `ActiveTimerInfo` — status types
*
* Key methods: `start()`, `stop()`, `getStatus()`, `getConfig()`,
* `getDetectionStatus()`, `getActiveTimers()`, `getAggregateMetrics()`,
@@ -46,11 +46,10 @@
import { EventEmitter } from 'node:events';
import { randomUUID } from 'node:crypto';
import { Session } from './session.js';
import { AiIdleChecker, type AiCheckResult, type AiCheckState } from './ai-idle-checker.js';
import { AiPlanChecker, type AiPlanCheckResult } from './ai-plan-checker.js';
import { AiIdleChecker, type AiCheckState } from './ai-idle-checker.js';
import { AiPlanChecker } from './ai-plan-checker.js';
import type { TeamWatcher } from './team-watcher.js';
import { BufferAccumulator } from './utils/buffer-accumulator.js';
import { ANSI_ESCAPE_PATTERN_SIMPLE, assertNever, CleanupManager } from './utils/index.js';
import { BufferAccumulator, ANSI_ESCAPE_PATTERN_SIMPLE, assertNever, CleanupManager } from './utils/index.js';
import { MAX_RESPAWN_BUFFER_SIZE, TRIM_RESPAWN_BUFFER_TO as RESPAWN_BUFFER_TRIM_SIZE } from './config/buffer-limits.js';
import {
isCompletionMessage,
@@ -62,13 +61,22 @@ import {
import { RespawnAdaptiveTiming } from './respawn-adaptive-timing.js';
import { RespawnCycleMetricsTracker } from './respawn-metrics.js';
import { calculateHealthScore, shouldSkipClear, type HealthInputs } from './respawn-health.js';
import { AI_CHECK_MODEL, AI_IDLE_CHECK_MAX_CONTEXT, AI_PLAN_CHECK_MAX_CONTEXT } from './config/ai-defaults.js';
import type {
RespawnCycleMetrics,
RespawnAggregateMetrics,
RalphLoopHealthScore,
TimingHistory,
CycleOutcome,
import {
AI_CHECK_MODEL,
AI_IDLE_CHECK_MAX_CONTEXT,
AI_PLAN_CHECK_MAX_CONTEXT,
AI_IDLE_CHECK_TIMEOUT_MS,
AI_IDLE_CHECK_COOLDOWN_MS,
AI_PLAN_CHECK_TIMEOUT_MS,
AI_PLAN_CHECK_COOLDOWN_MS,
} from './config/ai-defaults.js';
import {
getErrorMessage,
type RespawnCycleMetrics,
type RespawnAggregateMetrics,
type RalphLoopHealthScore,
type TimingHistory,
type CycleOutcome,
} from './types.js';
// ========== Constants ==========
@@ -91,13 +99,13 @@ const PLAN_MODE_SELECTOR_PATTERN = /[❯>]\s*\d+\./;
* Each layer provides a confidence signal that Claude has finished working.
*/
/** Active timer info for UI display */
export interface ActiveTimerInfo {
interface ActiveTimerInfo {
name: string;
remainingMs: number;
totalMs: number;
}
export interface DetectionStatus {
interface DetectionStatus {
/** Layer 0: Stop hook received (highest priority - definitive signal) */
stopHookReceived: boolean;
/** Timestamp when Stop hook was received */
@@ -480,68 +488,19 @@ export interface RespawnConfig {
* @event error - Fired on errors
* @event log - Fired for debug logging
*/
/** Timer info for countdown display */
export interface TimerInfo {
name: string;
durationMs: number;
endsAt: number;
reason?: string;
}
/** Action log entry for detailed UI feedback */
export interface ActionLogEntry {
interface ActionLogEntry {
type: string;
detail: string;
timestamp: number;
}
export interface RespawnEvents {
/** State machine transition */
stateChanged: (state: RespawnState, prevState: RespawnState) => void;
/** New respawn cycle started */
respawnCycleStarted: (cycleNumber: number) => void;
/** Respawn cycle finished */
respawnCycleCompleted: (cycleNumber: number) => void;
/** Command sent to session */
stepSent: (step: string, input: string) => void;
/** Step completed (ready indicator detected) */
stepCompleted: (step: string) => void;
/** Detection status update for UI display */
detectionUpdate: (status: DetectionStatus) => void;
/** Auto-accept sent for plan mode approval */
autoAcceptSent: () => void;
/** AI idle check started */
aiCheckStarted: () => void;
/** AI idle check completed with verdict */
aiCheckCompleted: (result: AiCheckResult) => void;
/** AI idle check failed */
aiCheckFailed: (error: string) => void;
/** AI idle check cooldown state changed */
aiCheckCooldown: (active: boolean, endsAt: number | null) => void;
/** AI plan check started */
planCheckStarted: () => void;
/** AI plan check completed with verdict */
planCheckCompleted: (result: AiPlanCheckResult) => void;
/** AI plan check failed */
planCheckFailed: (error: string) => void;
/** Timer started for countdown display */
timerStarted: (timer: TimerInfo) => void;
/** Timer cancelled */
timerCancelled: (timerName: string, reason?: string) => void;
/** Timer completed */
timerCompleted: (timerName: string) => void;
/** Verbose action log for detailed UI feedback */
actionLog: (action: ActionLogEntry) => void;
/** Error occurred */
error: (error: Error) => void;
/** Debug log message */
log: (message: string) => void;
/** Stuck state warning emitted */
stuckStateWarning: (state: RespawnState, durationMs: number) => void;
/** Stuck state recovery triggered */
stuckStateRecovery: (state: RespawnState, durationMs: number, attempt: number) => void;
/** Respawn blocked by external signal */
respawnBlocked: (data: { reason: string; details: string }) => void;
/**
* Convert milliseconds to a non-negative whole number of seconds for countdown display.
* Rounds up so that e.g. 1200 ms shows as 2 s (never under-reports remaining time).
*/
function formatRemainingSeconds(ms: number): number {
return Math.max(0, Math.ceil(ms / 1000));
}
/** Default configuration values */
@@ -559,13 +518,13 @@ const DEFAULT_CONFIG: RespawnConfig = {
aiIdleCheckEnabled: true, // use AI to confirm idle state
aiIdleCheckModel: AI_CHECK_MODEL,
aiIdleCheckMaxContext: AI_IDLE_CHECK_MAX_CONTEXT,
aiIdleCheckTimeoutMs: 90000, // 90 seconds (thinking can be slow)
aiIdleCheckCooldownMs: 180000, // 3 minutes after WORKING verdict
aiIdleCheckTimeoutMs: AI_IDLE_CHECK_TIMEOUT_MS,
aiIdleCheckCooldownMs: AI_IDLE_CHECK_COOLDOWN_MS,
aiPlanCheckEnabled: true, // use AI to confirm plan mode before auto-accept
aiPlanCheckModel: AI_CHECK_MODEL,
aiPlanCheckMaxContext: AI_PLAN_CHECK_MAX_CONTEXT,
aiPlanCheckTimeoutMs: 60000, // 60 seconds (thinking can be slow)
aiPlanCheckCooldownMs: 30000, // 30 seconds after NOT_PLAN_MODE
aiPlanCheckTimeoutMs: AI_PLAN_CHECK_TIMEOUT_MS,
aiPlanCheckCooldownMs: AI_PLAN_CHECK_COOLDOWN_MS,
stuckStateDetectionEnabled: true, // detect stuck states
stuckStateWarningMs: 300000, // 5 minutes warning threshold
stuckStateRecoveryMs: 600000, // 10 minutes recovery threshold
@@ -832,27 +791,42 @@ export class RespawnController extends EventEmitter {
private validateConfig(): void {
const c = this.config;
// Ensure timeouts are positive
if (c.idleTimeoutMs <= 0) c.idleTimeoutMs = DEFAULT_CONFIG.idleTimeoutMs;
if (c.completionConfirmMs <= 0) c.completionConfirmMs = DEFAULT_CONFIG.completionConfirmMs;
if (c.noOutputTimeoutMs <= 0) c.noOutputTimeoutMs = DEFAULT_CONFIG.noOutputTimeoutMs;
if (c.autoAcceptDelayMs < 0) c.autoAcceptDelayMs = DEFAULT_CONFIG.autoAcceptDelayMs;
if (c.interStepDelayMs <= 0) c.interStepDelayMs = DEFAULT_CONFIG.interStepDelayMs;
/**
* Validate that a timeout value is positive (or non-negative when allowZero is true).
* Falls back to the DEFAULT_CONFIG value if invalid.
*/
const validatePositiveTimeout = (field: keyof RespawnConfig, allowZero = false): void => {
const value = c[field] as number;
const invalid = allowZero ? value < 0 : value <= 0;
if (invalid) {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
(c as any)[field] = DEFAULT_CONFIG[field];
}
};
const REQUIRED_TIMEOUT_FIELDS = [
'idleTimeoutMs',
'completionConfirmMs',
'noOutputTimeoutMs',
'interStepDelayMs',
'aiIdleCheckTimeoutMs',
'aiIdleCheckMaxContext',
'aiPlanCheckTimeoutMs',
'aiPlanCheckMaxContext',
] as const;
for (const field of REQUIRED_TIMEOUT_FIELDS) {
validatePositiveTimeout(field);
}
const ALLOW_ZERO_FIELDS = ['autoAcceptDelayMs', 'aiIdleCheckCooldownMs', 'aiPlanCheckCooldownMs'] as const;
for (const field of ALLOW_ZERO_FIELDS) {
validatePositiveTimeout(field, true);
}
// Ensure completion confirm doesn't exceed no-output timeout
if (c.completionConfirmMs > c.noOutputTimeoutMs) {
c.completionConfirmMs = c.noOutputTimeoutMs;
}
// Ensure AI check timeouts are positive
if (c.aiIdleCheckTimeoutMs <= 0) c.aiIdleCheckTimeoutMs = DEFAULT_CONFIG.aiIdleCheckTimeoutMs;
if (c.aiIdleCheckCooldownMs < 0) c.aiIdleCheckCooldownMs = DEFAULT_CONFIG.aiIdleCheckCooldownMs;
if (c.aiIdleCheckMaxContext <= 0) c.aiIdleCheckMaxContext = DEFAULT_CONFIG.aiIdleCheckMaxContext;
// Ensure plan check timeouts are positive
if (c.aiPlanCheckTimeoutMs <= 0) c.aiPlanCheckTimeoutMs = DEFAULT_CONFIG.aiPlanCheckTimeoutMs;
if (c.aiPlanCheckCooldownMs < 0) c.aiPlanCheckCooldownMs = DEFAULT_CONFIG.aiPlanCheckCooldownMs;
if (c.aiPlanCheckMaxContext <= 0) c.aiPlanCheckMaxContext = DEFAULT_CONFIG.aiPlanCheckMaxContext;
}
/** Wire up AI checker events to controller events (removes existing listeners first to prevent duplicates) */
@@ -980,11 +954,11 @@ export class RespawnController extends EventEmitter {
waitingFor = 'AI verdict (IDLE or WORKING)';
} else if (this._state === 'confirming_idle') {
statusText = `Confirming idle (${confidence}% confidence)`;
waitingFor = `${Math.max(0, Math.ceil((this.config.completionConfirmMs - msSinceLastOutput) / 1000))}s more silence`;
waitingFor = `${formatRemainingSeconds(this.config.completionConfirmMs - msSinceLastOutput)}s more silence`;
} else if (this._state === 'watching') {
const aiState = this.aiChecker.getState();
if (aiState.status === 'cooldown') {
const remaining = Math.ceil(this.aiChecker.getCooldownRemainingMs() / 1000);
const remaining = formatRemainingSeconds(this.aiChecker.getCooldownRemainingMs());
statusText = `AI Check: WORKING (cooldown ${remaining}s)`;
waitingFor = 'Cooldown to expire';
} else if (completionMessageDetected) {
@@ -1357,10 +1331,40 @@ export class RespawnController extends EventEmitter {
this.lastTokenChangeTime = now;
}
// Detect completion message FIRST (Layer 1) - PRIMARY DETECTION
// Check this before working patterns because completion message indicates
// the work is done, even if working patterns are still in the rolling window
if (isCompletionMessage(data)) {
// Layer 1: Completion message (PRIMARY) — checked before working patterns
if (this._detectCompletionMessage(data, now)) return;
// Layer 4: Working patterns
if (this._detectWorkingPattern(data, now)) return;
// Substantial output during confirming_idle/ai_checking cancels the flow
if (this._state === 'confirming_idle' || this._state === 'ai_checking') {
// Strip ANSI escape codes to check if there's real content
ANSI_ESCAPE_PATTERN_SIMPLE.lastIndex = 0;
const stripped = data.replace(ANSI_ESCAPE_PATTERN_SIMPLE, '').trim();
if (stripped.length > 2) {
if (this._state === 'ai_checking') {
this.log(`Substantial output during AI check ("${stripped.substring(0, 40)}..."), cancelling`);
this.aiChecker.cancel();
this.setState('watching');
} else {
// Real content (not just escape codes or single chars) - cancel confirmation
this.log(
`Substantial output during confirmation ("${stripped.substring(0, 40)}..."), cancelling idle detection`
);
this.cancelCompletionConfirm();
}
return;
}
}
// Legacy fallback: prompt detection
this._detectPrompt(data);
}
private _detectCompletionMessage(data: string, now: number): boolean {
if (!isCompletionMessage(data)) return false;
// Clear the rolling window - completion marks a transition point
this.clearWorkingPatternWindow();
this.workingDetected = false;
@@ -1371,7 +1375,7 @@ export class RespawnController extends EventEmitter {
// In watching state, start completion confirmation timer
if (this._state === 'watching') {
this.startCompletionConfirmTimer();
return;
return true;
}
// In waiting states, also use confirmation timer (same detection logic)
@@ -1404,12 +1408,13 @@ export class RespawnController extends EventEmitter {
default:
assertNever(this._state, `Unhandled RespawnState in completion detection: ${this._state}`);
}
return;
return true;
}
// Detect working patterns (Layer 4)
private _detectWorkingPattern(data: string, now: number): boolean {
const isWorking = this.checkWorkingPattern(data);
if (isWorking) {
if (!isWorking) return false;
this.workingDetected = true;
this.promptDetected = false;
this.elicitationDetected = false; // Clear on new work cycle
@@ -1444,34 +1449,13 @@ export class RespawnController extends EventEmitter {
this.emit('stepCompleted', 'init');
this.completeCycle();
}
return;
return true;
}
// In confirming_idle or ai_checking state, substantial output cancels the flow.
// This prevents false triggers when Claude pauses briefly mid-work.
if (this._state === 'confirming_idle' || this._state === 'ai_checking') {
// Strip ANSI escape codes to check if there's real content
ANSI_ESCAPE_PATTERN_SIMPLE.lastIndex = 0;
const stripped = data.replace(ANSI_ESCAPE_PATTERN_SIMPLE, '').trim();
if (stripped.length > 2) {
if (this._state === 'ai_checking') {
this.log(`Substantial output during AI check ("${stripped.substring(0, 40)}..."), cancelling`);
this.aiChecker.cancel();
this.setState('watching');
} else {
// Real content (not just escape codes or single chars) - cancel confirmation
this.log(
`Substantial output during confirmation ("${stripped.substring(0, 40)}..."), cancelling idle detection`
);
this.cancelCompletionConfirm();
}
return;
}
}
// Legacy fallback: detect prompt characters (still useful for waiting_* states)
private _detectPrompt(data: string): void {
const hasPrompt = PROMPT_PATTERNS.some((pattern) => data.includes(pattern));
if (hasPrompt) {
if (!hasPrompt) return;
this.promptDetected = true;
this.workingDetected = false;
@@ -1507,7 +1491,6 @@ export class RespawnController extends EventEmitter {
assertNever(this._state, `Unhandled RespawnState in prompt detection: ${this._state}`);
}
}
}
/**
* Handle update step completion.
@@ -1791,18 +1774,23 @@ export class RespawnController extends EventEmitter {
case 'sending_init':
case 'sending_kickstart':
// For sending states, retry the send
this.log('Recovery: returning to watching state');
this.setState('watching');
this.startNoOutputTimer();
this.startPreFilterTimer();
if (this.config.autoAcceptPrompts) {
this.startAutoAcceptTimer();
}
this.recoveryResetToWatching('returning to watching state');
break;
default:
// Fallback: reset to watching
this.log('Recovery: fallback to watching state');
this.recoveryResetToWatching('fallback to watching state');
}
}
/**
* Reset the controller to watching state during stuck-state recovery.
* Sets state to watching and restarts all detection timers.
*
* @param reason - Human-readable reason for the reset (logged)
*/
private recoveryResetToWatching(reason: string): void {
this.log(`Recovery: ${reason}`);
this.setState('watching');
this.startNoOutputTimer();
this.startPreFilterTimer();
@@ -1810,7 +1798,6 @@ export class RespawnController extends EventEmitter {
this.startAutoAcceptTimer();
}
}
}
/**
* Get stuck-state metrics for UI display.
@@ -2034,6 +2021,18 @@ export class RespawnController extends EventEmitter {
return;
}
// Check for active child processes (bash tools, test suites, builds, etc.)
// These may produce no terminal output, so restart timers to retry periodically.
const activeProcesses = this.session.getActiveChildProcesses();
if (activeProcesses.length > 0) {
const names = activeProcesses.map((p) => p.command).join(', ');
this.log(`Skipping AI check - ${activeProcesses.length} active child process(es): ${names}`);
this.logAction('detection', `Skipped AI check: child processes running (${names})`);
this.startNoOutputTimer();
this.startPreFilterTimer();
return;
}
// If AI check is disabled or errored out, fall back to direct idle confirmation
if (!this.config.aiIdleCheckEnabled || this.aiChecker.status === 'disabled') {
this.log(`AI check unavailable (${this.aiChecker.status}), confirming idle directly via: ${reason}`);
@@ -2044,7 +2043,7 @@ export class RespawnController extends EventEmitter {
// If on cooldown, don't start check - wait for cooldown to expire
if (this.aiChecker.isOnCooldown()) {
this.log(
`AI check on cooldown (${Math.ceil(this.aiChecker.getCooldownRemainingMs() / 1000)}s remaining), waiting...`
`AI check on cooldown (${formatRemainingSeconds(this.aiChecker.getCooldownRemainingMs())}s remaining), waiting...`
);
return;
}
@@ -2129,7 +2128,7 @@ export class RespawnController extends EventEmitter {
}
if (this._state === 'stopped') return; // Guard against stopped state
if (this._state === 'ai_checking') {
const errorMsg = err instanceof Error ? err.message : String(err);
const errorMsg = getErrorMessage(err);
this.logAction('ai-check', `Failed: ${errorMsg.substring(0, 50)}`);
this.emit('aiCheckFailed', errorMsg);
this.setState('watching');
@@ -2191,36 +2190,15 @@ export class RespawnController extends EventEmitter {
* @fires planCheckStarted
*/
private tryAutoAccept(): void {
// Only auto-accept in watching state (not during a respawn cycle)
if (this._state !== 'watching') return;
if (!this.canAutoAccept()) return;
// Don't auto-accept if a completion message was detected (normal idle handles it)
if (this.completionMessageTime !== null) return;
// Don't auto-accept if disabled
if (!this.config.autoAcceptPrompts) return;
// Don't auto-accept if we haven't received any output yet (prevents spurious Enter on fresh start)
if (!this.hasReceivedOutput) return;
// Don't auto-accept if an elicitation dialog (AskUserQuestion) was detected
if (this.elicitationDetected) {
this.log('Skipping auto-accept: elicitation dialog detected (AskUserQuestion)');
return;
}
// Stage 1: Pre-filter — check if buffer looks like plan mode
const buffer = this.terminalBuffer.value;
if (!this.isPlanModePreFilterMatch(buffer)) {
this.log('Skipping auto-accept: pre-filter did not match plan mode patterns');
return;
}
// Stage 2: AI confirmation (if enabled and available)
if (this.config.aiPlanCheckEnabled && this.planChecker.status !== 'disabled') {
if (this.planChecker.isOnCooldown()) {
this.log(
`Skipping auto-accept: plan checker on cooldown (${Math.ceil(this.planChecker.getCooldownRemainingMs() / 1000)}s remaining)`
`Skipping auto-accept: plan checker on cooldown (${formatRemainingSeconds(this.planChecker.getCooldownRemainingMs())}s remaining)`
);
return;
}
@@ -2237,6 +2215,40 @@ export class RespawnController extends EventEmitter {
this.sendAutoAcceptEnter();
}
/**
* Check whether all preconditions for auto-accept are met.
* Validates state, config, and pre-filter conditions before attempting auto-accept.
*
* @returns True if auto-accept should proceed to the AI confirmation stage
*/
private canAutoAccept(): boolean {
// Only auto-accept in watching state (not during a respawn cycle)
if (this._state !== 'watching') return false;
// Don't auto-accept if a completion message was detected (normal idle handles it)
if (this.completionMessageTime !== null) return false;
// Don't auto-accept if disabled
if (!this.config.autoAcceptPrompts) return false;
// Don't auto-accept if we haven't received any output yet (prevents spurious Enter on fresh start)
if (!this.hasReceivedOutput) return false;
// Don't auto-accept if an elicitation dialog (AskUserQuestion) was detected
if (this.elicitationDetected) {
this.log('Skipping auto-accept: elicitation dialog detected (AskUserQuestion)');
return false;
}
// Stage 1: Pre-filter — check if buffer looks like plan mode
if (!this.isPlanModePreFilterMatch(this.terminalBuffer.value)) {
this.log('Skipping auto-accept: pre-filter did not match plan mode patterns');
return false;
}
return true;
}
/**
* Check if the terminal buffer matches plan mode pre-filter patterns.
* Only checks the last 2000 chars (plan mode UI appears at the bottom).
@@ -2315,7 +2327,7 @@ export class RespawnController extends EventEmitter {
}
})
.catch((err) => {
const errorMsg = err instanceof Error ? err.message : String(err);
const errorMsg = getErrorMessage(err);
this.emit('planCheckFailed', errorMsg);
this.logAction('plan-check', `Failed: ${errorMsg.substring(0, 50)}`);
});
@@ -2625,6 +2637,18 @@ export class RespawnController extends EventEmitter {
return;
}
// Safety check: if child processes are running (bash tools, test suites, builds, etc.)
const activeProcesses = this.session.getActiveChildProcesses();
if (activeProcesses.length > 0) {
const names = activeProcesses.map((p) => p.command).join(', ');
this.log(`Idle confirmation rejected - ${activeProcesses.length} active child process(es): ${names}`);
this.logAction('detection', `Rejected: child processes running (${names})`);
this.setState('watching');
this.startNoOutputTimer();
this.startPreFilterTimer();
return;
}
this.log(`Idle confirmed via: ${reason}`);
const status = this.getDetectionStatus();
this.log(
+2 -1
View File
@@ -22,6 +22,7 @@ import {
RunSummaryStats,
createInitialRunSummaryStats,
} from './types.js';
import { CLEANUP_CHECK_INTERVAL_MS } from './config/server-timing.js';
/** Maximum events to keep per session (FIFO trimming) */
const MAX_EVENTS = 1000;
@@ -36,7 +37,7 @@ const TOKEN_MILESTONE_INTERVAL = 50000;
const STATE_STUCK_WARNING_MS = 10 * 60 * 1000; // 10 minutes
/** State stuck check interval (ms) */
const STATE_STUCK_CHECK_INTERVAL = 60 * 1000; // 1 minute
const STATE_STUCK_CHECK_INTERVAL = CLEANUP_CHECK_INTERVAL_MS;
/**
* Tracks events and statistics for a session's run summary.
+90 -49
View File
@@ -27,6 +27,57 @@ const COMPACT_COOLDOWN_MS = 10000;
/** Cooldown after clear completes before re-enabling (5 seconds) */
const CLEAR_COOLDOWN_MS = 5000;
/**
* Executes an action when the session becomes idle, retrying if currently working.
*
* @param action - The async action to execute once idle
* @param isActive - Returns whether this operation is still active (not cancelled)
* @param isWorking - Returns whether the session is currently working
* @param isStopped - Returns whether the session has been stopped
* @param retryMs - Delay between retry attempts when working
* @param cooldownMs - Delay after action completes before calling onCooldownDone
* @param setTimer - Stores the timer reference for cleanup
* @param onCooldownDone - Called after cooldown to reset state
*/
async function executeWhenIdle(
action: () => Promise<void>,
isActive: () => boolean,
isWorking: () => boolean,
isStopped: () => boolean,
retryMs: number,
cooldownMs: number,
setTimer: (timer: NodeJS.Timeout | null) => void,
onCooldownDone: () => void
): Promise<void> {
if (isStopped()) return;
if (!isActive()) return;
if (!isWorking()) {
if (isStopped()) return;
await action();
if (!isStopped()) {
setTimer(
setTimeout(() => {
if (isStopped()) return;
setTimer(null);
onCooldownDone();
}, cooldownMs)
);
}
} else {
if (!isStopped()) {
setTimer(
setTimeout(
() => executeWhenIdle(action, isActive, isWorking, isStopped, retryMs, cooldownMs, setTimer, onCooldownDone),
retryMs
)
);
}
}
}
/** Minimum valid threshold for auto-clear/compact (1000 tokens) */
const MIN_AUTO_THRESHOLD = 1000;
@@ -42,7 +93,7 @@ const DEFAULT_AUTO_COMPACT_THRESHOLD = 110_000;
/**
* Callbacks required by SessionAutoOps to interact with the parent Session.
*/
export interface AutoOpsCallbacks {
interface AutoOpsCallbacks {
/** Send a command via the terminal multiplexer */
writeCommand: (command: string) => Promise<boolean>;
/** Check if Claude is currently working */
@@ -58,12 +109,6 @@ export interface AutoOpsCallbacks {
/**
* Events emitted by SessionAutoOps.
*/
export interface SessionAutoOpsEvents {
/** Auto-compact was triggered and the /compact command was sent */
autoCompact: (data: { tokens: number; threshold: number; prompt?: string }) => void;
/** Auto-clear was triggered and the /clear command was sent */
autoClear: (data: { tokens: number; threshold: number }) => void;
}
/**
* Manages auto-compact and auto-clear automation for a Session.
@@ -181,13 +226,7 @@ export class SessionAutoOps extends EventEmitter {
`[SessionAutoOps] Auto-compact triggered: ${totalTokens} tokens >= ${this._autoCompactThreshold} threshold`
);
const checkAndCompact = async () => {
if (this.callbacks.isStopped()) return;
if (!this._isCompacting) return;
if (!this.callbacks.isWorking()) {
if (this.callbacks.isStopped()) return;
const action = async () => {
const compactCmd = this._autoCompactPrompt ? `/compact ${this._autoCompactPrompt}\r` : '/compact\r';
await this.callbacks.writeCommand(compactCmd);
this.emit('autoCompact', {
@@ -195,23 +234,27 @@ export class SessionAutoOps extends EventEmitter {
threshold: this._autoCompactThreshold,
prompt: this._autoCompactPrompt || undefined,
});
if (!this.callbacks.isStopped()) {
this._autoCompactTimer = setTimeout(() => {
if (this.callbacks.isStopped()) return;
this._autoCompactTimer = null;
this._isCompacting = false;
}, COMPACT_COOLDOWN_MS);
}
} else {
if (!this.callbacks.isStopped()) {
this._autoCompactTimer = setTimeout(checkAndCompact, AUTO_RETRY_DELAY_MS);
}
}
};
if (!this.callbacks.isStopped()) {
this._autoCompactTimer = setTimeout(checkAndCompact, AUTO_INITIAL_DELAY_MS);
this._autoCompactTimer = setTimeout(
() =>
executeWhenIdle(
action,
() => this._isCompacting,
() => this.callbacks.isWorking(),
() => this.callbacks.isStopped(),
AUTO_RETRY_DELAY_MS,
COMPACT_COOLDOWN_MS,
(timer) => {
this._autoCompactTimer = timer;
},
() => {
this._isCompacting = false;
}
),
AUTO_INITIAL_DELAY_MS
);
}
}
}
@@ -231,32 +274,30 @@ export class SessionAutoOps extends EventEmitter {
`[SessionAutoOps] Auto-clear triggered: ${totalTokens} tokens >= ${this._autoClearThreshold} threshold`
);
const checkAndClear = async () => {
if (this.callbacks.isStopped()) return;
if (!this._isClearing) return;
if (!this.callbacks.isWorking()) {
if (this.callbacks.isStopped()) return;
const action = async () => {
await this.callbacks.writeCommand('/clear\r');
this.emit('autoClear', { tokens: totalTokens, threshold: this._autoClearThreshold });
if (!this.callbacks.isStopped()) {
this._autoClearTimer = setTimeout(() => {
if (this.callbacks.isStopped()) return;
this._autoClearTimer = null;
this._isClearing = false;
}, CLEAR_COOLDOWN_MS);
}
} else {
if (!this.callbacks.isStopped()) {
this._autoClearTimer = setTimeout(checkAndClear, AUTO_RETRY_DELAY_MS);
}
}
};
if (!this.callbacks.isStopped()) {
this._autoClearTimer = setTimeout(checkAndClear, AUTO_INITIAL_DELAY_MS);
this._autoClearTimer = setTimeout(
() =>
executeWhenIdle(
action,
() => this._isClearing,
() => this.callbacks.isWorking(),
() => this.callbacks.isStopped(),
AUTO_RETRY_DELAY_MS,
CLEAR_COOLDOWN_MS,
(timer) => {
this._autoClearTimer = timer;
},
() => {
this._isClearing = false;
}
),
AUTO_INITIAL_DELAY_MS
);
}
}
}
+2 -2
View File
@@ -9,13 +9,13 @@
*/
import type { ClaudeMode } from './types.js';
import { getAugmentedPath } from './utils/claude-cli-resolver.js';
import { getAugmentedPath } from './utils/index.js';
/**
* Build Claude CLI permission flags based on the configured mode.
* Returns an array of args to pass to the CLI.
*/
export function buildPermissionArgs(claudeMode: ClaudeMode, allowedTools?: string): string[] {
function buildPermissionArgs(claudeMode: ClaudeMode, allowedTools?: string): string[] {
switch (claudeMode) {
case 'dangerously-skip-permissions':
return ['--dangerously-skip-permissions'];
+10 -14
View File
@@ -30,18 +30,6 @@ import { SessionState } from './types.js';
/**
* Events emitted by SessionManager
*/
export interface SessionManagerEvents {
/** Fired when a new session starts successfully */
sessionStarted: (session: Session) => void;
/** Fired when a session stops (graceful or forced) */
sessionStopped: (sessionId: string) => void;
/** Fired when a session encounters an error */
sessionError: (sessionId: string, error: string) => void;
/** Fired when a session produces terminal output */
sessionOutput: (sessionId: string, output: string) => void;
/** Fired when a completion phrase is detected */
sessionCompletion: (sessionId: string, phrase: string) => void;
}
/**
* Manages multiple Claude sessions with lifecycle coordination.
@@ -164,7 +152,7 @@ export class SessionManager extends EventEmitter {
await session.start();
this.sessions.set(session.id, session);
this.store.setSession(session.id, session.toState());
this.updateSessionState(session);
this.emit('sessionStarted', session);
return session;
@@ -259,7 +247,15 @@ export class SessionManager extends EventEmitter {
}
private updateSessionState(session: Session): void {
this.store.setSession(session.id, session.toState());
// envOverrides is intentionally NOT on SessionState (API safety). For disk
// persistence we augment the stored object with __envOverrides so reboot
// recovery can restore them without leaking through any API serializer.
// The key uses the reserved `__` prefix so it is visibly "internal" to any
// future reader of state.json.
const state = session.toState();
const envOverrides = session.getEnvOverridesForPersist();
const toStore = envOverrides ? { ...state, __envOverrides: envOverrides } : state;
this.store.setSession(session.id, toStore as SessionState);
}
/** Gets all sessions from persistent storage (including stopped). */
+337 -315
View File
@@ -29,6 +29,7 @@
*/
import { EventEmitter } from 'node:events';
import { execSync } from 'node:child_process';
import { v4 as uuidv4 } from 'uuid';
import * as pty from 'node-pty';
import {
@@ -40,6 +41,7 @@ import {
ActiveBashTool,
NiceConfig,
DEFAULT_NICE_CONFIG,
getErrorMessage,
type ClaudeMode,
type SessionMode,
type OpenCodeConfig,
@@ -48,8 +50,14 @@ import type { TerminalMultiplexer, MuxSession } from './mux-interface.js';
import { TaskTracker, type BackgroundTask } from './task-tracker.js';
import { RalphTracker } from './ralph-tracker.js';
import { BashToolParser } from './bash-tool-parser.js';
import { BufferAccumulator } from './utils/buffer-accumulator.js';
import { ANSI_ESCAPE_PATTERN_FULL, TOKEN_PATTERN, SPINNER_PATTERN, MAX_SESSION_TOKENS } from './utils/index.js';
import {
BufferAccumulator,
ANSI_ESCAPE_PATTERN_FULL,
TOKEN_PATTERN,
SPINNER_PATTERN,
MAX_SESSION_TOKENS,
execPattern,
} from './utils/index.js';
import {
MAX_TERMINAL_BUFFER_SIZE,
TRIM_TERMINAL_TO as TERMINAL_BUFFER_TRIM_SIZE,
@@ -58,6 +66,7 @@ import {
MAX_MESSAGES,
MAX_LINE_BUFFER_SIZE,
} from './config/buffer-limits.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
import {
buildInteractiveArgs,
buildPromptArgs,
@@ -145,63 +154,6 @@ export interface ClaudeMessage {
* Event signatures emitted by the Session class.
* Subscribe using `session.on('eventName', handler)`.
*/
export interface SessionEvents {
/** Processed text output (ANSI stripped) */
output: (data: string) => void;
/** Parsed JSON message from Claude CLI */
message: (msg: ClaudeMessage) => void;
/** Error output from the session */
error: (data: string) => void;
/** Session process exited */
exit: (code: number | null) => void;
/** One-shot prompt completed with result and cost */
completion: (result: string, cost: number) => void;
/** Raw terminal data (includes ANSI codes) */
terminal: (data: string) => void;
/** Signal to clear terminal display (after mux attach) */
clearTerminal: () => void;
/** New background task started */
taskCreated: (task: BackgroundTask) => void;
/** Background task status changed */
taskUpdated: (task: BackgroundTask) => void;
/** Background task finished successfully */
taskCompleted: (task: BackgroundTask) => void;
/** Background task failed with error */
taskFailed: (task: BackgroundTask, error: string) => void;
/** Auto-clear triggered due to token threshold */
autoClear: (data: { tokens: number; threshold: number }) => void;
/** Auto-compact triggered due to token threshold */
autoCompact: (data: { tokens: number; threshold: number; prompt?: string }) => void;
/** Ralph loop state changed */
ralphLoopUpdate: (state: RalphTrackerState) => void;
/** Ralph todo list updated */
ralphTodoUpdate: (todos: RalphTodoItem[]) => void;
/** Ralph completion phrase detected */
ralphCompletionDetected: (phrase: string) => void;
/** RALPH_STATUS block detected */
ralphStatusBlockDetected: (block: import('./types.js').RalphStatusBlock) => void;
/** Circuit breaker state changed */
ralphCircuitBreakerUpdate: (status: import('./types.js').CircuitBreakerStatus) => void;
/** Dual-condition exit gate met */
ralphExitGateMet: (data: { completionIndicators: number; exitSignal: boolean }) => void;
/** Bash tool with file paths started */
bashToolStart: (tool: ActiveBashTool) => void;
/** Bash tool completed */
bashToolEnd: (tool: ActiveBashTool) => void;
/** Active Bash tools list updated */
bashToolsUpdate: (tools: ActiveBashTool[]) => void;
/** CLI info (version, model, account) updated */
cliInfoUpdated: (info: {
version: string | null;
model: string | null;
accountType: string | null;
latestVersion: string | null;
}) => void;
}
// SessionMode is imported from types.ts (single source of truth)
// Re-export for backwards compatibility with any external consumers
export type { SessionMode } from './types.js';
/**
* Core session class that wraps a PTY process running Claude CLI or a shell.
@@ -265,6 +217,7 @@ export class Session extends EventEmitter {
private _lastPromptTime: number = 0;
private activityTimeout: NodeJS.Timeout | null = null;
private _awaitingIdleConfirmation: boolean = false; // Prevents timeout reset during idle detection
private _trustDialogAccepted: boolean = false; // Prevents repeated trust dialog auto-accept
private _taskTracker: TaskTracker;
// Token tracking for auto-clear
@@ -318,6 +271,11 @@ export class Session extends EventEmitter {
// OpenCode configuration (only for mode === 'opencode')
private _openCodeConfig: OpenCodeConfig | undefined;
private _resumeSessionId: string | undefined;
// Ephemeral env overrides (e.g., CLAUDE_CODE_EFFORT_LEVEL). Exported by tmux at spawn,
// preserved across respawns via persisted state. Not written to .claude/settings.local.json.
private _envOverrides: Record<string, string> | undefined;
// Session color for visual differentiation
private _color: import('./types.js').SessionColor = 'default';
@@ -376,6 +334,10 @@ export class Session extends EventEmitter {
allowedTools?: string;
/** OpenCode configuration (only for mode === 'opencode') */
openCodeConfig?: OpenCodeConfig;
/** Resume a previous Claude conversation (used after server reboot) */
resumeSessionId?: string;
/** Extra env vars exported to the CLI at spawn time (no disk persistence) */
envOverrides?: Record<string, string>;
}
) {
super();
@@ -392,12 +354,10 @@ export class Session extends EventEmitter {
this.createdAt = config.createdAt || Date.now();
this.mode = config.mode || 'claude';
this._name = config.name || '';
this._resumeSessionId = config.resumeSessionId;
this._lastActivityAt = this.createdAt;
// Set claudeSessionId immediately — Codeman always passes --session-id ${this.id}
// to Claude CLI, so the Claude session ID always matches the Codeman session ID.
// This ensures subagent matching works even for recovered sessions (where
// startInteractive() hasn't been called yet).
this._claudeSessionId = this.id;
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = config.resumeSessionId || this.id;
this._mux = config.mux || null;
this._useMux = config.useMux ?? (this._mux !== null && this._mux.isAvailable());
this._muxSession = config.muxSession || null;
@@ -425,6 +385,11 @@ export class Session extends EventEmitter {
this._openCodeConfig = config.openCodeConfig;
}
// Apply env overrides (exported at spawn, not persisted to disk)
if (config.envOverrides && Object.keys(config.envOverrides).length > 0) {
this._envOverrides = { ...config.envOverrides };
}
// Initialize task tracker and forward events (store handlers for cleanup)
this._taskTracker = new TaskTracker();
this._taskTrackerHandlers = {
@@ -519,6 +484,20 @@ export class Session extends EventEmitter {
return this._claudeSessionId;
}
// Adopt a Claude conversation ID observed from an external source (e.g. hook
// payload). In interactive PTY mode Claude CLI emits no JSON to stdout, so
// `_handleJsonMessage` never sees `session_id`; hooks are the only signal
// that conveys a post-/clear conversation switch.
adoptClaudeSessionId(newId: string): void {
if (!newId || newId === this._claudeSessionId) return;
this._claudeSessionId = newId;
}
/** The tmux session name, if the session is running inside a mux */
get muxName(): string | null {
return this._muxSession?.muxName ?? null;
}
get totalCost(): number {
return this._totalCost;
}
@@ -531,6 +510,49 @@ export class Session extends EventEmitter {
return this._isWorking;
}
/**
* Check if the session's process tree has active child processes beyond Claude itself.
* Detects running bash tools, test suites, builds, servers, etc. that Claude spawned.
*
* The tmux pane PID is typically "claude" directly (bash exec'd into it). When Claude
* runs a bash tool, it spawns child processes: claude → bash → npm/node/python/etc.
* We check direct children of the pane PID, filtering out "claude" itself (for the rare
* case where bash wraps claude and didn't exec).
*
* Returns an array of {pid, command} for each child process, or empty array if none.
* Returns empty array if no mux session or on error (fail-open to avoid blocking respawn).
*/
getActiveChildProcesses(): { pid: number; command: string }[] {
if (!this._muxSession) return [];
try {
const panePid = this._muxSession.pid;
// Single call: get direct children with their command names
const output = execSync(`ps -o pid=,comm= --ppid ${panePid} 2>/dev/null`, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (!output) return [];
const activeProcesses: { pid: number; command: string }[] = [];
for (const line of output.split('\n')) {
const match = line.trim().match(/^(\d+)\s+(.+)/);
if (!match) continue;
const pid = parseInt(match[1], 10);
const command = match[2].trim();
// Skip the claude process itself (pane_pid may be bash wrapping claude)
if (command === 'claude') continue;
activeProcesses.push({ pid, command });
}
return activeProcesses;
} catch {
// ps returns exit code 1 when no matches — normal (no children)
return [];
}
}
get lastPromptTime(): number {
return this._lastPromptTime;
}
@@ -786,9 +808,30 @@ export class Session extends EventEmitter {
cliAccountType: this._cliAccountType || undefined,
cliLatestVersion: this._cliLatestVersion || undefined,
openCodeConfig: this._openCodeConfig,
resumeSessionId: this._resumeSessionId,
// envOverrides intentionally NOT on the public SessionState type — they must not
// leak into SSE / GET /api/sessions broadcasts (schema allows OPENCODE_*, which
// can carry secrets). For disk persistence, session-manager calls
// getEnvOverridesForPersist() and writes alongside state.
};
}
/**
* Returns a subset of env overrides safe for disk persistence (state.json).
* Only non-sensitive `CLAUDE_CODE_*` keys are included. `OPENCODE_*` keys are
* filtered out because the schema permits them and they can carry secrets
* (e.g., OPENCODE_API_KEY); secrets must not land in `~/.codeman/state.json`.
* Must NOT be included in any API-bound serializer — see toState() comment.
*/
getEnvOverridesForPersist(): Record<string, string> | undefined {
if (!this._envOverrides) return undefined;
const safe: Record<string, string> = {};
for (const [key, value] of Object.entries(this._envOverrides)) {
if (key.startsWith('CLAUDE_CODE_')) safe[key] = value;
}
return Object.keys(safe).length > 0 ? safe : undefined;
}
toDetailedState() {
return {
...this.toLightDetailedState(),
@@ -863,29 +906,15 @@ export class Session extends EventEmitter {
* session.write('help me with this code\r');
* ```
*/
async startInteractive(): Promise<void> {
if (this.ptyProcess) {
throw new Error('Session already has a running process');
}
private async _setupOrAttachMuxSession(options: {
respawnPaneOptions: import('./mux-interface.js').RespawnPaneOptions;
createSessionOptions: import('./mux-interface.js').CreateSessionOptions;
spawnErrLabel: string;
}): Promise<{ isRestored: boolean }> {
const mux = this._mux!;
this._status = 'busy';
this._terminalBuffer.clear();
this._textOutput.clear();
this._errorBuffer = '';
this._messages = [];
this._lineBuffer = '';
this._lastActivityAt = Date.now();
const modeLabel = this.mode === 'opencode' ? 'OpenCode' : 'Claude';
console.log(
`[Session] Starting interactive ${modeLabel} session` + (this._useMux ? ` (with ${this._mux!.backend})` : '')
);
// If mux wrapping is enabled, create or attach to a mux session
if (this._useMux && this._mux) {
try {
// Verify stale mux session — tmux may have been destroyed (e.g., killed externally)
if (this._muxSession && !this._mux.muxSessionExists(this._muxSession.muxName)) {
if (this._muxSession && !mux.muxSessionExists(this._muxSession.muxName)) {
console.log('[Session] Stale mux session detected (tmux gone):', this._muxSession.muxName);
this._muxSession = null;
}
@@ -893,18 +922,9 @@ export class Session extends EventEmitter {
// Check if session exists but pane is dead (remain-on-exit keeps it alive)
// Respawn the pane instead of creating a whole new session — preserves tmux scrollback
let needsNewSession = false;
if (this._muxSession && this._mux.isPaneDead(this._muxSession.muxName)) {
if (this._muxSession && mux.isPaneDead(this._muxSession.muxName)) {
console.log('[Session] Dead pane detected, respawning:', this._muxSession.muxName);
const newPid = await this._mux.respawnPane({
sessionId: this.id,
workingDir: this.workingDir,
mode: this.mode,
niceConfig: this._niceConfig,
model: this._model,
claudeMode: this._claudeMode,
allowedTools: this._allowedTools,
openCodeConfig: this._openCodeConfig,
});
const newPid = await mux.respawnPane(options.respawnPaneOptions);
if (!newPid) {
console.error('[Session] Failed to respawn pane, will create new session');
needsNewSession = true;
@@ -915,12 +935,71 @@ export class Session extends EventEmitter {
}
// Check if we already have a mux session (restored session)
const isRestoredSession = this._muxSession !== null && !needsNewSession;
if (isRestoredSession) {
const isRestored = this._muxSession !== null && !needsNewSession;
if (isRestored) {
console.log('[Session] Attaching to existing mux session:', this._muxSession!.muxName);
} else {
// Create a new mux session
this._muxSession = await this._mux.createSession({
this._muxSession = await mux.createSession(options.createSessionOptions);
console.log('[Session] Created mux session:', this._muxSession.muxName);
// No extra sleep — createSession() already waits for tmux readiness
}
// Attach to the mux session via PTY
try {
this.ptyProcess = pty.spawn(mux.getAttachCommand(), mux.getAttachArgs(this._muxSession!.muxName), {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildMuxAttachEnv(),
});
} catch (spawnErr) {
console.error(`[Session] Failed to spawn PTY for ${options.spawnErrLabel}:`, spawnErr);
this.emit('error', `Failed to attach to mux session: ${spawnErr}`);
throw spawnErr;
}
return { isRestored };
}
private _handleTerminalOutput(data: string): void {
// BufferAccumulator handles auto-trimming when max size exceeded
this._terminalBuffer.append(data);
this._lastActivityAt = Date.now();
this.emit('terminal', data);
this.emit('output', data);
}
async startInteractive(): Promise<void> {
if (this.ptyProcess) {
throw new Error('Session already has a running process');
}
this._resetBuffers();
const modeLabel = this.mode === 'opencode' ? 'OpenCode' : 'Claude';
console.log(
`[Session] Starting interactive ${modeLabel} session` + (this._useMux ? ` (with ${this._mux!.backend})` : '')
);
// If mux wrapping is enabled, create or attach to a mux session
if (this._useMux && this._mux) {
try {
const { isRestored } = await this._setupOrAttachMuxSession({
respawnPaneOptions: {
sessionId: this.id,
workingDir: this.workingDir,
mode: this.mode,
niceConfig: this._niceConfig,
model: this._model,
claudeMode: this._claudeMode,
allowedTools: this._allowedTools,
openCodeConfig: this._openCodeConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
},
createSessionOptions: {
sessionId: this.id,
workingDir: this.workingDir,
mode: this.mode,
@@ -930,37 +1009,18 @@ export class Session extends EventEmitter {
claudeMode: this._claudeMode,
allowedTools: this._allowedTools,
openCodeConfig: this._openCodeConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
},
spawnErrLabel: 'mux attachment',
});
console.log('[Session] Created mux session:', this._muxSession.muxName);
// No extra sleep — createSession() already waits for tmux readiness
}
// Attach to the mux session via PTY
try {
this.ptyProcess = pty.spawn(
this._mux.getAttachCommand(),
this._mux.getAttachArgs(this._muxSession!.muxName),
{
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildMuxAttachEnv(),
}
);
// Set claudeSessionId immediately since we passed --session-id to Claude
// The mux manager passes --session-id ${sessionId} to Claude
this._claudeSessionId = this.id;
} catch (spawnErr) {
console.error('[Session] Failed to spawn PTY for mux attachment:', spawnErr);
this.emit('error', `Failed to attach to mux session: ${spawnErr}`);
throw spawnErr;
}
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = this._resumeSessionId || this.id;
// For NEW mux sessions: wait for readiness then clean buffer
// For RESTORED mux sessions: don't do anything - client will fetch buffer on tab switch
if (!isRestoredSession) {
if (!isRestored) {
if (this.mode === 'opencode') {
// OpenCode uses Bubble Tea TUI — no ❯ prompt to detect.
// Wait for TUI to stabilize (output stops changing), then mark ready.
@@ -1026,7 +1086,8 @@ export class Session extends EventEmitter {
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildClaudeEnv(this.id),
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
});
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY:', spawnErr);
@@ -1036,9 +1097,8 @@ export class Session extends EventEmitter {
}
}
// Set the claudeSessionId immediately since we passed --session-id
// This ensures subagent matching works without waiting for JSON messages
this._claudeSessionId = this.id;
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = this._resumeSessionId || this.id;
this._pid = this.ptyProcess.pid;
console.log('[Session] Interactive PTY spawned with PID:', this._pid);
@@ -1048,12 +1108,17 @@ export class Session extends EventEmitter {
const data = rawData.replace(FOCUS_ESCAPE_FILTER, '').replace(CTRL_L_PATTERN, ''); // Remove Ctrl+L
if (!data) return; // Skip if only filtered sequences
// BufferAccumulator handles auto-trimming when max size exceeded
this._terminalBuffer.append(data);
this._lastActivityAt = Date.now();
this._handleTerminalOutput(data);
this.emit('terminal', data);
this.emit('output', data);
// === Auto-accept workspace trust dialog ===
// Claude CLI 2.x shows "Yes, I trust this folder" prompt on first launch per directory.
// Codeman sessions always use --dangerously-skip-permissions, so auto-accept.
if (!this._trustDialogAccepted && data.includes('trust this folder')) {
this._trustDialogAccepted = true;
console.log(`[Session] Auto-accepting workspace trust dialog for: ${this.id}`);
// Send Enter to accept the default selection ("Yes, I trust this folder")
this.writeViaMux('\r');
}
// === Idle/working detection runs on every chunk (latency-sensitive) ===
// Detect if Claude is working or at prompt
@@ -1249,13 +1314,7 @@ export class Session extends EventEmitter {
throw new Error('Session already has a running process');
}
this._status = 'busy';
this._terminalBuffer.clear();
this._textOutput.clear();
this._errorBuffer = '';
this._messages = [];
this._lineBuffer = '';
this._lastActivityAt = Date.now();
this._resetBuffers();
// Use user's default shell or bash
const shell = process.env.SHELL || '/bin/bash';
@@ -1267,69 +1326,28 @@ export class Session extends EventEmitter {
// If mux wrapping is enabled, create or attach to a mux session
if (this._useMux && this._mux) {
try {
// Verify stale mux session — tmux may have been destroyed externally
if (this._muxSession && !this._mux.muxSessionExists(this._muxSession.muxName)) {
console.log('[Session] Stale mux session detected (tmux gone):', this._muxSession.muxName);
this._muxSession = null;
}
// Check if session exists but pane is dead (remain-on-exit keeps it alive)
let needsNewSession = false;
if (this._muxSession && this._mux.isPaneDead(this._muxSession.muxName)) {
console.log('[Session] Dead pane detected, respawning:', this._muxSession.muxName);
const newPid = await this._mux.respawnPane({
const { isRestored } = await this._setupOrAttachMuxSession({
respawnPaneOptions: {
sessionId: this.id,
workingDir: this.workingDir,
mode: 'shell',
niceConfig: this._niceConfig,
});
if (!newPid) {
console.error('[Session] Failed to respawn pane, will create new session');
needsNewSession = true;
} else {
await new Promise((resolve) => setTimeout(resolve, MUX_STARTUP_DELAY_MS));
}
}
// Check if we already have a mux session (restored session)
const isRestoredSession = this._muxSession !== null && !needsNewSession;
if (isRestoredSession) {
console.log('[Session] Attaching to existing mux session:', this._muxSession!.muxName);
} else {
// Create a new mux session
this._muxSession = await this._mux.createSession({
envOverrides: this._envOverrides,
},
createSessionOptions: {
sessionId: this.id,
workingDir: this.workingDir,
mode: 'shell',
name: this._name,
niceConfig: this._niceConfig,
envOverrides: this._envOverrides,
},
spawnErrLabel: 'shell mux attachment',
});
console.log('[Session] Created mux session:', this._muxSession.muxName);
// No extra sleep — createSession() already waits for tmux readiness
}
// Attach to the mux session via PTY
try {
this.ptyProcess = pty.spawn(
this._mux.getAttachCommand(),
this._mux.getAttachArgs(this._muxSession!.muxName),
{
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildMuxAttachEnv(),
}
);
} catch (spawnErr) {
console.error('[Session] Failed to spawn PTY for shell mux attachment:', spawnErr);
this.emit('error', `Failed to attach to mux session: ${spawnErr}`);
throw spawnErr;
}
// For NEW sessions: clear by sending 'clear' command to the shell
// For RESTORED sessions: don't clear - we want to see the existing output
if (!isRestoredSession) {
if (!isRestored) {
setTimeout(() => {
if (this.ptyProcess) {
this._terminalBuffer.clear();
@@ -1370,12 +1388,7 @@ export class Session extends EventEmitter {
const data = rawData.replace(FOCUS_ESCAPE_FILTER, '');
if (!data) return; // Skip if only focus sequences
// BufferAccumulator handles auto-trimming when max size exceeded
this._terminalBuffer.append(data);
this._lastActivityAt = Date.now();
this.emit('terminal', data);
this.emit('output', data);
this._handleTerminalOutput(data);
});
this.ptyProcess.onExit(({ exitCode }) => {
@@ -1440,13 +1453,7 @@ export class Session extends EventEmitter {
return;
}
this._status = 'busy';
this._terminalBuffer.clear();
this._textOutput.clear();
this._errorBuffer = '';
this._messages = [];
this._lineBuffer = '';
this._lastActivityAt = Date.now();
this._resetBuffers();
this._promptResolved = false; // Reset race condition guard
this.resolvePromise = resolve;
@@ -1469,7 +1476,8 @@ export class Session extends EventEmitter {
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildClaudeEnv(this.id),
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
});
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY for runPrompt:', spawnErr);
@@ -1489,12 +1497,7 @@ export class Session extends EventEmitter {
const data = rawData.replace(FOCUS_ESCAPE_FILTER, '');
if (!data) return; // Skip if only focus sequences
// BufferAccumulator handles auto-trimming when max size exceeded
this._terminalBuffer.append(data);
this._lastActivityAt = Date.now();
this.emit('terminal', data);
this.emit('output', data);
this._handleTerminalOutput(data);
// Also try to parse JSON lines for structured data
this.processOutput(data);
@@ -1526,9 +1529,11 @@ export class Session extends EventEmitter {
this._status = 'idle';
const cost = resultMsg.total_cost_usd || 0;
this._totalCost += cost;
this.emit('completion', resultMsg.result || '', cost);
// Claude CLI stream-json may return empty result field — fall back to accumulated text output
const result = resultMsg.result || this._textOutput.value || '';
this.emit('completion', result, cost);
if (resolve) {
resolve({ result: resultMsg.result || '', cost });
resolve({ result, cost });
}
} else if (exitCode !== 0 || (resultMsg && resultMsg.is_error)) {
this._status = 'error';
@@ -1557,6 +1562,120 @@ export class Session extends EventEmitter {
});
}
private _resetBuffers(): void {
this._status = 'busy';
this._terminalBuffer.clear();
this._textOutput.clear();
this._errorBuffer = '';
this._messages = [];
this._lineBuffer = '';
this._lastActivityAt = Date.now();
}
private _clearAllTimers(): void {
// Clear activity timeout to prevent memory leak
if (this.activityTimeout) {
clearTimeout(this.activityTimeout);
this.activityTimeout = null;
}
// Clear line buffer flush timer
if (this._lineBufferFlushTimer) {
clearTimeout(this._lineBufferFlushTimer);
this._lineBufferFlushTimer = null;
}
// Destroy auto-compact/auto-clear automation (clears its timers)
this._autoOps.destroy();
// Clear prompt check timers
if (this._promptCheckInterval) {
clearInterval(this._promptCheckInterval);
this._promptCheckInterval = null;
}
if (this._promptCheckTimeout) {
clearTimeout(this._promptCheckTimeout);
this._promptCheckTimeout = null;
}
// Clear shell idle timer
if (this._shellIdleTimer) {
clearTimeout(this._shellIdleTimer);
this._shellIdleTimer = null;
}
// Clear expensive processing timer
if (this._expensiveProcessTimer) {
clearTimeout(this._expensiveProcessTimer);
this._expensiveProcessTimer = null;
}
this._pendingCleanData = '';
}
private _handleJsonMessage(cleanLine: string, rawLine: string): void {
try {
const msg = JSON.parse(cleanLine) as ClaudeMessage;
this._messages.push(msg);
this.emit('message', msg);
// Trim messages array for long-running sessions
if (this._messages.length > MAX_MESSAGES) {
this._messages = this._messages.slice(-Math.floor(MAX_MESSAGES * 0.8));
}
// Extract Claude session ID from messages (can be in any message type).
// Support both sessionId (camelCase) and session_id (snake_case).
// The constructor seeds _claudeSessionId with this.id as a placeholder;
// once Claude CLI emits its real session ID, adopt it so JSONL lookups
// (e.g. /api/sessions/:id/last-response) can find the transcript file.
const msgSessionId =
((msg as unknown as Record<string, unknown>).sessionId as string | undefined) ?? msg.session_id;
if (msgSessionId && msgSessionId !== this._claudeSessionId) {
this._claudeSessionId = msgSessionId;
}
// Process message for task tracking
this._taskTracker.processMessage(msg);
if (msg.type === 'assistant' && msg.message?.content) {
for (const block of msg.message.content) {
if (block.type === 'text' && block.text) {
this._textOutput.append(block.text);
}
}
// Track tokens from usage (with validation)
if (msg.message.usage) {
const inputDelta = msg.message.usage.input_tokens || 0;
const outputDelta = msg.message.usage.output_tokens || 0;
// Sanity check: max 100k tokens per message (generous limit)
const MAX_TOKENS_PER_MESSAGE = 100_000;
if (inputDelta > 0 && inputDelta <= MAX_TOKENS_PER_MESSAGE) {
this._totalInputTokens += inputDelta;
}
if (outputDelta > 0 && outputDelta <= MAX_TOKENS_PER_MESSAGE) {
this._totalOutputTokens += outputDelta;
}
// Check if we should auto-compact or auto-clear
this._autoOps.checkAutoCompact();
this._autoOps.checkAutoClear();
}
}
if (msg.type === 'result' && msg.total_cost_usd) {
this._totalCost = msg.total_cost_usd;
}
} catch (parseErr) {
// Not JSON, just regular output - this is expected for non-JSON lines
console.debug(
'[Session] Line not JSON (expected for text output):',
parseErr instanceof Error ? parseErr.message : parseErr
);
this._textOutput.append(rawLine + '\n');
}
}
private processOutput(data: string): void {
// Early return if session is stopped to prevent any processing or timer creation
if (this._isStopped) return;
@@ -1598,64 +1717,7 @@ export class Session extends EventEmitter {
const cleanLine = trimmed.replace(ANSI_ESCAPE_PATTERN_FULL, '');
if (cleanLine.startsWith('{') && cleanLine.endsWith('}')) {
try {
const msg = JSON.parse(cleanLine) as ClaudeMessage;
this._messages.push(msg);
this.emit('message', msg);
// Trim messages array for long-running sessions
if (this._messages.length > MAX_MESSAGES) {
this._messages = this._messages.slice(-Math.floor(MAX_MESSAGES * 0.8));
}
// Extract Claude session ID from messages (can be in any message type)
// Support both sessionId (camelCase) and session_id (snake_case)
const msgSessionId =
((msg as unknown as Record<string, unknown>).sessionId as string | undefined) ?? msg.session_id;
if (msgSessionId && !this._claudeSessionId) {
this._claudeSessionId = msgSessionId;
}
// Process message for task tracking
this._taskTracker.processMessage(msg);
if (msg.type === 'assistant' && msg.message?.content) {
for (const block of msg.message.content) {
if (block.type === 'text' && block.text) {
this._textOutput.append(block.text);
}
}
// Track tokens from usage (with validation)
if (msg.message.usage) {
const inputDelta = msg.message.usage.input_tokens || 0;
const outputDelta = msg.message.usage.output_tokens || 0;
// Sanity check: max 100k tokens per message (generous limit)
const MAX_TOKENS_PER_MESSAGE = 100_000;
if (inputDelta > 0 && inputDelta <= MAX_TOKENS_PER_MESSAGE) {
this._totalInputTokens += inputDelta;
}
if (outputDelta > 0 && outputDelta <= MAX_TOKENS_PER_MESSAGE) {
this._totalOutputTokens += outputDelta;
}
// Check if we should auto-compact or auto-clear
this._autoOps.checkAutoCompact();
this._autoOps.checkAutoClear();
}
}
if (msg.type === 'result' && msg.total_cost_usd) {
this._totalCost = msg.total_cost_usd;
}
} catch (parseErr) {
// Not JSON, just regular output - this is expected for non-JSON lines
console.debug(
'[Session] Line not JSON (expected for text output):',
parseErr instanceof Error ? parseErr.message : parseErr
);
this._textOutput.append(line + '\n');
}
this._handleJsonMessage(cleanLine, line);
} else if (trimmed) {
this._textOutput.append(line + '\n');
}
@@ -1692,16 +1754,12 @@ export class Session extends EventEmitter {
// Quick pre-check: skip expensive regex if no common tool patterns present
if (!cleanLine.includes('(') || !cleanLine.includes(')')) return;
// Reset regex lastIndex for global pattern
TASK_TOOL_PATTERN.lastIndex = 0;
let match;
while ((match = TASK_TOOL_PATTERN.exec(cleanLine)) !== null) {
execPattern(TASK_TOOL_PATTERN, cleanLine, (match) => {
const description = match[2].trim();
if (description && description.length > 0) {
this._taskCache.add(Date.now(), description);
}
}
});
}
/**
@@ -1951,7 +2009,7 @@ export class Session extends EventEmitter {
this._status = 'busy';
this._lastActivityAt = Date.now();
this.runPrompt(input).catch((err) => {
const errorMsg = err instanceof Error ? err.message : String(err);
const errorMsg = getErrorMessage(err);
// Clean up task state so the task queue doesn't get stuck
if (this._currentTaskId) {
const taskId = this._currentTaskId;
@@ -2026,43 +2084,7 @@ export class Session extends EventEmitter {
// Set stopped flag first to prevent new timers from being created
this._isStopped = true;
// Clear activity timeout to prevent memory leak
if (this.activityTimeout) {
clearTimeout(this.activityTimeout);
this.activityTimeout = null;
}
// Clear line buffer flush timer
if (this._lineBufferFlushTimer) {
clearTimeout(this._lineBufferFlushTimer);
this._lineBufferFlushTimer = null;
}
// Destroy auto-compact/auto-clear automation (clears its timers)
this._autoOps.destroy();
// Clear prompt check timers
if (this._promptCheckInterval) {
clearInterval(this._promptCheckInterval);
this._promptCheckInterval = null;
}
if (this._promptCheckTimeout) {
clearTimeout(this._promptCheckTimeout);
this._promptCheckTimeout = null;
}
// Clear shell idle timer
if (this._shellIdleTimer) {
clearTimeout(this._shellIdleTimer);
this._shellIdleTimer = null;
}
// Clear expensive processing timer
if (this._expensiveProcessTimer) {
clearTimeout(this._expensiveProcessTimer);
this._expensiveProcessTimer = null;
}
this._pendingCleanData = '';
this._clearAllTimers();
// Immediately cleanup Promise callbacks to prevent orphaned references
// during the rest of stop() processing (e.g., if mux kill times out)
+98 -80
View File
@@ -116,6 +116,26 @@ export class StateStore {
this.loadRalphStates();
}
private _mergeWithInitialState(parsed: Partial<AppState>): AppState {
const initial = createInitialState();
return {
...initial,
...parsed,
sessions: { ...parsed.sessions },
tasks: { ...parsed.tasks },
ralphLoop: { ...initial.ralphLoop, ...parsed.ralphLoop },
config: { ...initial.config, ...parsed.config },
};
}
private _resetCircuitBreaker(): void {
this.consecutiveSaveFailures = 0;
if (this.circuitBreakerOpen) {
console.log('[StateStore] Circuit breaker CLOSED - save succeeded');
this.circuitBreakerOpen = false;
}
}
private ensureDir(): void {
const dir = dirname(this.filePath);
if (!existsSync(dir)) {
@@ -132,15 +152,7 @@ export class StateStore {
if (existsSync(path)) {
const data = readFileSync(path, 'utf-8');
const parsed = JSON.parse(data) as Partial<AppState>;
const initial = createInitialState();
const result = {
...initial,
...parsed,
sessions: { ...parsed.sessions },
tasks: { ...parsed.tasks },
ralphLoop: { ...initial.ralphLoop, ...parsed.ralphLoop },
config: { ...initial.config, ...parsed.config },
};
const result = this._mergeWithInitialState(parsed);
if (path !== this.filePath) {
console.warn(`[StateStore] Recovered state from backup: ${path}`);
}
@@ -195,16 +207,7 @@ export class StateStore {
* Only dirty sessions are re-serialized; clean sessions use cached JSON fragments.
*/
private assembleStateJson(): string {
// Re-serialize dirty sessions and update cache
for (const id of this.dirtySessions) {
const session = this.state.sessions[id];
if (session) {
this.cachedSessionJsons.set(id, JSON.stringify(session));
} else {
this.cachedSessionJsons.delete(id);
}
}
this.dirtySessions.clear();
this.updateDirtySessionCache();
// Build sessions object from cached fragments
const sessionParts: string[] = [];
@@ -218,6 +221,25 @@ export class StateStore {
sessionParts.push(`${JSON.stringify(id)}:${json}`);
}
this.pruneStaleCacheEntries();
return this.buildPartialJson(sessionParts);
}
private updateDirtySessionCache(): void {
// Re-serialize dirty sessions and update cache
for (const id of this.dirtySessions) {
const session = this.state.sessions[id];
if (session) {
this.cachedSessionJsons.set(id, JSON.stringify(session));
} else {
this.cachedSessionJsons.delete(id);
}
}
this.dirtySessions.clear();
}
private pruneStaleCacheEntries(): void {
// Prune stale cache entries (sessions removed via direct state mutation)
if (this.cachedSessionJsons.size > Object.keys(this.state.sessions).length) {
for (const cachedId of this.cachedSessionJsons.keys()) {
@@ -226,7 +248,9 @@ export class StateStore {
}
}
}
}
private buildPartialJson(sessionParts: string[]): string {
// Build final JSON: sessions from cache, everything else re-serialized (tiny)
const sessionsJson = `{${sessionParts.join(',')}}`;
@@ -249,6 +273,28 @@ export class StateStore {
return `{${parts.join(',')}}`;
}
private serializeState(): string | null {
try {
return this.assembleStateJson();
} catch (assembleErr) {
// Fallback to full serialization if incremental assembly fails
console.warn('[StateStore] assembleStateJson failed, falling back to full serialize:', assembleErr);
this.cachedSessionJsons.clear();
this.dirtySessions.clear();
try {
return JSON.stringify(this.state);
} catch (err) {
console.error('[StateStore] Failed to serialize state (circular reference or invalid data):', err);
this.consecutiveSaveFailures++;
if (this.consecutiveSaveFailures >= MAX_CONSECUTIVE_FAILURES) {
console.error('[StateStore] Circuit breaker OPEN - serialization failing repeatedly');
this.circuitBreakerOpen = true;
}
return null;
}
}
}
private async _doSaveAsync(): Promise<void> {
this.saveDeb.cancel();
if (!this.dirty) {
@@ -263,30 +309,12 @@ export class StateStore {
this.ensureDir();
const tempPath = this.filePath + '.tmp';
const tempPath = `${this.filePath}.${process.pid}.${Date.now()}.${Math.random().toString(36).slice(2)}.tmp`;
const backupPath = this.filePath + '.bak';
let json: string;
// Step 1: Serialize state (validates it's JSON-safe)
try {
json = this.assembleStateJson();
} catch (assembleErr) {
// Fallback to full serialization if incremental assembly fails
console.warn('[StateStore] assembleStateJson failed, falling back to full serialize:', assembleErr);
this.cachedSessionJsons.clear();
this.dirtySessions.clear();
try {
json = JSON.stringify(this.state);
} catch (err) {
console.error('[StateStore] Failed to serialize state (circular reference or invalid data):', err);
this.consecutiveSaveFailures++;
if (this.consecutiveSaveFailures >= MAX_CONSECUTIVE_FAILURES) {
console.error('[StateStore] Circuit breaker OPEN - serialization failing repeatedly');
this.circuitBreakerOpen = true;
}
return;
}
}
const json = this.serializeState();
if (json === null) return;
// Clear dirty flag BEFORE async I/O so mutations during write re-set it.
// The state snapshot is already captured in `json` above.
@@ -305,11 +333,7 @@ export class StateStore {
await writeFile(tempPath, json, 'utf-8');
await rename(tempPath, this.filePath);
this.consecutiveSaveFailures = 0;
if (this.circuitBreakerOpen) {
console.log('[StateStore] Circuit breaker CLOSED - save succeeded');
this.circuitBreakerOpen = false;
}
this._resetCircuitBreaker();
} catch (err) {
console.error('[StateStore] Failed to write state file:', err);
// Re-mark dirty so the data is retried on the next save cycle
@@ -349,29 +373,11 @@ export class StateStore {
this.ensureDir();
const tempPath = this.filePath + '.tmp';
const tempPath = `${this.filePath}.${process.pid}.${Date.now()}.${Math.random().toString(36).slice(2)}.tmp`;
const backupPath = this.filePath + '.bak';
let json: string;
try {
json = this.assembleStateJson();
} catch (assembleErr) {
// Fallback to full serialization if incremental assembly fails
console.warn('[StateStore] assembleStateJson failed, falling back to full serialize:', assembleErr);
this.cachedSessionJsons.clear();
this.dirtySessions.clear();
try {
json = JSON.stringify(this.state);
} catch (err) {
console.error('[StateStore] Failed to serialize state (circular reference or invalid data):', err);
this.consecutiveSaveFailures++;
if (this.consecutiveSaveFailures >= MAX_CONSECUTIVE_FAILURES) {
console.error('[StateStore] Circuit breaker OPEN - serialization failing repeatedly');
this.circuitBreakerOpen = true;
}
return;
}
}
const json = this.serializeState();
if (json === null) return;
// Backup via atomic copy (avoids reading entire file into memory)
try {
@@ -387,11 +393,7 @@ export class StateStore {
renameSync(tempPath, this.filePath);
// Clear dirty flag only AFTER successful write
this.dirty = false;
this.consecutiveSaveFailures = 0;
if (this.circuitBreakerOpen) {
console.log('[StateStore] Circuit breaker CLOSED - save succeeded');
this.circuitBreakerOpen = false;
}
this._resetCircuitBreaker();
} catch (err) {
console.error('[StateStore] Failed to write state file:', err);
this.consecutiveSaveFailures++;
@@ -417,15 +419,7 @@ export class StateStore {
if (existsSync(backupPath)) {
const backupContent = readFileSync(backupPath, 'utf-8');
const parsed = JSON.parse(backupContent) as Partial<AppState>;
const initial = createInitialState();
this.state = {
...initial,
...parsed,
sessions: { ...parsed.sessions },
tasks: { ...parsed.tasks },
ralphLoop: { ...initial.ralphLoop, ...parsed.ralphLoop },
config: { ...initial.config, ...parsed.config },
};
this.state = this._mergeWithInitialState(parsed);
console.log('[StateStore] Successfully recovered state from backup');
// Reset circuit breaker after successful recovery
this.circuitBreakerOpen = false;
@@ -547,6 +541,30 @@ export class StateStore {
this.save();
}
// ========== Orchestrator Loop State Methods ==========
/** Returns the orchestrator loop state, or null if never initialized. */
getOrchestratorState() {
return this.state.orchestrator ?? null;
}
/** Updates orchestrator loop state (partial merge) and triggers a debounced save. */
setOrchestratorState(orchestrator: Partial<NonNullable<AppState['orchestrator']>>) {
if (this.state.orchestrator) {
this.state.orchestrator = { ...this.state.orchestrator, ...orchestrator };
} else {
// First initialization — caller must provide full state
this.state.orchestrator = orchestrator as NonNullable<AppState['orchestrator']>;
}
this.save();
}
/** Clears orchestrator state and triggers a debounced save. */
clearOrchestratorState() {
this.state.orchestrator = undefined;
this.save();
}
/** Returns the application configuration. */
getConfig() {
return this.state.config;
+188 -161
View File
@@ -17,7 +17,7 @@
* Tracks per-agent: status, token counts, model, description, tool call count, liveness (PID).
*
* @dependencies config/map-limits (MAX_TRACKED_AGENTS, PENDING_TOOL_CALL_TTL_MS),
* utils (CleanupManager, KeyedDebouncer)
* config/buffer-limits (FILE_PEEK_BYTES), utils (CleanupManager, KeyedDebouncer)
* @consumedby web/server (SSE broadcast), session (subagent-session correlation)
* @emits subagent:discovered, subagent:updated, subagent:tool_call, subagent:tool_result,
* subagent:progress, subagent:message, subagent:completed
@@ -34,6 +34,8 @@ import { join, basename } from 'node:path';
import { execFile } from 'node:child_process';
import { readFile, readdir, stat as statAsync } from 'node:fs/promises';
import { PENDING_TOOL_CALL_TTL_MS, MAX_PENDING_TOOL_CALLS, MAX_TRACKED_AGENTS } from './config/map-limits.js';
import { STALE_DATA_MAX_AGE_MS } from './config/server-timing.js';
import { FILE_PEEK_BYTES } from './config/buffer-limits.js';
import { CleanupManager, KeyedDebouncer } from './utils/index.js';
// ========== Types ==========
@@ -132,17 +134,6 @@ export interface SubagentToolResult {
isError: boolean; // Whether result is an error
}
export interface SubagentEvents {
'subagent:discovered': (info: SubagentInfo) => void;
'subagent:updated': (info: SubagentInfo) => void;
'subagent:tool_call': (data: SubagentToolCall) => void;
'subagent:tool_result': (data: SubagentToolResult) => void;
'subagent:progress': (data: SubagentProgress) => void;
'subagent:message': (data: SubagentMessage) => void;
'subagent:completed': (info: SubagentInfo) => void;
'subagent:error': (error: Error, agentId?: string) => void;
}
// ========== Constants ==========
const CLAUDE_PROJECTS_DIR = join(homedir(), '.claude/projects');
@@ -151,7 +142,7 @@ const POLL_INTERVAL_MS = 1000; // Base poll interval (lightweight checks)
const FULL_SCAN_EVERY_N_POLLS = 5; // Full directory traversal every 5th poll (5s)
const LIVENESS_CHECK_MS = 10000; // Check if subagent processes are still alive every 10s
const FILE_ALIVE_THRESHOLD_MS = 30000; // File mtime within 30s = agent alive (primary check)
const STALE_COMPLETED_MAX_AGE_MS = 60 * 60 * 1000; // Remove completed agents older than 1 hour
const STALE_COMPLETED_MAX_AGE_MS = STALE_DATA_MAX_AGE_MS; // Remove completed agents older than 1 hour
const STALE_IDLE_MAX_AGE_MS = 4 * 60 * 60 * 1000; // Remove idle agents older than 4 hours
const STARTUP_MAX_FILE_AGE_MS = 4 * 60 * 60 * 1000; // Only load files modified in last 4 hours on startup
@@ -220,6 +211,83 @@ export class SubagentWatcher extends EventEmitter {
return INTERNAL_AGENT_PATTERNS.some((pattern) => pattern.test(description));
}
/**
* Mark a subagent as completed: clear PID, set status, clean up pending tool calls, emit event.
*/
private markSubagentAsCompleted(info: SubagentInfo): void {
info.pid = undefined;
info.status = 'completed';
this.pendingToolCalls.delete(info.agentId);
this.emit('subagent:completed', info);
}
/**
* Extract text from message content, handling both string and array formats.
* For array content, returns the text from the first 'text' block.
*/
private extractFirstTextContent(
content: string | Array<{ type: string; text?: string }> | undefined
): string | undefined {
if (!content) return undefined;
if (typeof content === 'string') {
const trimmed = content.trim();
return trimmed.length > 0 ? trimmed : undefined;
}
if (Array.isArray(content)) {
const firstContent = content[0];
if (firstContent?.type === 'text' && firstContent.text) {
const trimmed = firstContent.text.trim();
return trimmed.length > 0 ? trimmed : undefined;
}
}
return undefined;
}
/**
* Process a tool_result content block: look up pending tool call, emit tool_result event.
*/
private emitToolResult(
content: { tool_use_id: string; content?: string | Array<{ type: string; text?: string }>; is_error?: boolean },
agentId: string,
sessionId: string,
timestamp: string
): void {
const resultContent = this.extractToolResultContent(content.content);
const agentPendingCalls = this.pendingToolCalls.get(agentId);
const pendingCall = agentPendingCalls?.get(content.tool_use_id);
const toolName = pendingCall?.toolName;
// Delete after lookup to prevent memory leak
agentPendingCalls?.delete(content.tool_use_id);
const toolResult: SubagentToolResult = {
agentId,
sessionId,
timestamp,
toolUseId: content.tool_use_id,
tool: toolName,
preview: resultContent.substring(0, MESSAGE_TEXT_LIMIT),
contentLength: resultContent.length,
isError: content.is_error || false,
};
this.emit('subagent:tool_result', toolResult);
}
/**
* Find the oldest inactive (non-active) agent for LRU eviction.
* Returns the agent ID of the oldest inactive agent, or null if all are active.
*/
private findOldestInactiveAgent(): string | null {
let oldestId: string | null = null;
let oldestTime = Infinity;
for (const [id, existing] of this.agentInfo) {
if (existing.status !== 'active' && existing.lastActivityAt < oldestTime) {
oldestTime = existing.lastActivityAt;
oldestId = id;
}
}
return oldestId;
}
/**
* Extract short model identifier from full model name
*/
@@ -305,10 +373,7 @@ export class SubagentWatcher extends EventEmitter {
const alive = this.checkSubagentAliveFromPidMap(info, pidMap);
if (!alive) {
info.pid = undefined;
info.status = 'completed';
this.pendingToolCalls.delete(info.agentId);
this.emit('subagent:completed', info);
this.markSubagentAsCompleted(info);
}
}
}
@@ -675,10 +740,7 @@ export class SubagentWatcher extends EventEmitter {
const pid = await this.findSubagentProcess(info.sessionId);
if (pid) {
process.kill(pid, 'SIGTERM');
info.pid = undefined;
info.status = 'completed';
this.pendingToolCalls.delete(info.agentId);
this.emit('subagent:completed', info);
this.markSubagentAsCompleted(info);
return true;
}
} catch {
@@ -686,10 +748,7 @@ export class SubagentWatcher extends EventEmitter {
}
// Mark as completed even if we couldn't find the process
info.pid = undefined;
info.status = 'completed';
this.pendingToolCalls.delete(info.agentId);
this.emit('subagent:completed', info);
this.markSubagentAsCompleted(info);
return true;
}
@@ -841,20 +900,10 @@ export class SubagentWatcher extends EventEmitter {
}
} else if (entry.type === 'user' && entry.message?.content) {
// Handle both string and array content formats
if (typeof entry.message.content === 'string') {
const text = entry.message.content.trim();
if (text.length < 100 && !text.includes('{')) {
const text = this.extractFirstTextContent(entry.message.content);
if (text && text.length < 100 && !text.includes('{')) {
lines.push(`${this.formatTime(entry.timestamp)} 📥 User: ${text.substring(0, USER_TEXT_PREVIEW_LENGTH)}`);
}
} else {
const firstContent = entry.message.content[0];
if (firstContent?.type === 'text' && firstContent.text) {
const text = firstContent.text.trim();
if (text.length < 100 && !text.includes('{')) {
lines.push(`${this.formatTime(entry.timestamp)} 📥 User: ${text.substring(0, USER_TEXT_PREVIEW_LENGTH)}`);
}
}
}
}
}
@@ -928,6 +977,24 @@ export class SubagentWatcher extends EventEmitter {
return truncated.replace(/[.!?,:\s]+$/, '');
}
private async _resolveDescription(
projectHash: string,
sessionId: string,
agentId: string,
filePath: string,
fallbackText?: string
): Promise<string | undefined> {
// First try parent transcript (most reliable)
const fromParent = await this.extractDescriptionFromParentTranscript(projectHash, sessionId, agentId);
if (fromParent) return fromParent;
// Fallback: inline text (from processEntry) or file extraction
if (fallbackText) {
return this.extractSmartTitle(fallbackText);
}
return this.extractDescriptionFromFile(filePath);
}
/**
* Extract the short description from the parent session's transcript.
* This is the most reliable method because it reads the actual Task tool result
@@ -1008,7 +1075,7 @@ export class SubagentWatcher extends EventEmitter {
private async extractDescriptionFromFile(filePath: string): Promise<string | undefined> {
try {
// Only read the first 8KB — more than enough for 5 JSONL lines
const stream = createReadStream(filePath, { end: 8191 });
const stream = createReadStream(filePath, { end: FILE_PEEK_BYTES });
const rl = createInterface({ input: stream });
return await new Promise<string | undefined>((resolve) => {
@@ -1025,15 +1092,7 @@ export class SubagentWatcher extends EventEmitter {
try {
const entry = JSON.parse(line);
if (entry.type === 'user' && entry.message?.content) {
let text: string | undefined;
if (typeof entry.message.content === 'string') {
text = entry.message.content.trim();
} else if (Array.isArray(entry.message.content)) {
const firstContent = entry.message.content[0];
if (firstContent?.type === 'text' && firstContent.text) {
text = firstContent.text.trim();
}
}
const text = this.extractFirstTextContent(entry.message.content);
if (text) {
resolved = true;
rl.close();
@@ -1140,10 +1199,10 @@ export class SubagentWatcher extends EventEmitter {
if (this.fileAgentContext.has(filePath)) {
// Known file — handle content change
this.handleFileChange(filePath).catch(() => {});
this.handleFileChange(filePath).catch(() => {}); // Ignore - errors logged internally, don't crash watcher callback
} else {
// New file — register it
this.registerAgentFile(filePath, projectHash, sessionId).catch(() => {});
this.registerAgentFile(filePath, projectHash, sessionId).catch(() => {}); // Ignore - errors logged internally, don't crash watcher callback
}
});
});
@@ -1193,18 +1252,13 @@ export class SubagentWatcher extends EventEmitter {
// Retry description extraction if missing (race condition fix)
if (!existingInfo.description) {
// First try parent transcript (most reliable)
let extractedDescription = await this.extractDescriptionFromParentTranscript(
const extractedDescription = await this._resolveDescription(
existingInfo.projectHash,
existingInfo.sessionId,
agentId
agentId,
filePath
);
// Fallback to subagent file
if (!extractedDescription) {
extractedDescription = await this.extractDescriptionFromFile(filePath);
}
if (extractedDescription) {
// Check if this is an internal agent - if so, remove it
if (this.isInternalAgent(extractedDescription)) {
this.removeAgent(agentId);
return;
@@ -1255,13 +1309,7 @@ export class SubagentWatcher extends EventEmitter {
}
// Extract description - prefer reading from parent transcript (most reliable)
// The parent transcript has the exact Task tool call with description parameter
let description = await this.extractDescriptionFromParentTranscript(projectHash, sessionId, agentId);
// Fallback: extract a smart title from the subagent's prompt if parent lookup failed
if (!description) {
description = await this.extractDescriptionFromFile(filePath);
}
const description = await this._resolveDescription(projectHash, sessionId, agentId, filePath);
// Skip internal Claude Code agents (e.g., suggestion mode) - not real subagents
if (this.isInternalAgent(description)) {
@@ -1284,14 +1332,7 @@ export class SubagentWatcher extends EventEmitter {
// Enforce MAX_TRACKED_AGENTS during insertion — evict oldest inactive agent
if (this.agentInfo.size >= MAX_TRACKED_AGENTS) {
let oldestId: string | null = null;
let oldestTime = Infinity;
for (const [id, existing] of this.agentInfo) {
if (existing.status !== 'active' && existing.lastActivityAt < oldestTime) {
oldestTime = existing.lastActivityAt;
oldestId = id;
}
}
const oldestId = this.findOldestInactiveAgent();
if (oldestId) {
this.removeAgent(oldestId);
}
@@ -1361,51 +1402,11 @@ export class SubagentWatcher extends EventEmitter {
private async processEntry(entry: SubagentTranscriptEntry, agentId: string, sessionId: string): Promise<void> {
const info = this.agentInfo.get(agentId);
// Extract model from assistant messages (first one sets the model)
if (info && entry.type === 'assistant' && entry.message?.model && !info.model) {
info.model = entry.message.model;
info.modelShort = this.extractModelShort(entry.message.model);
this.emit('subagent:updated', info);
}
if (info) {
this._processModelInfo(entry, info);
this._processTokenInfo(entry, info);
// Aggregate token usage from messages
if (info && entry.message?.usage) {
if (entry.message.usage.input_tokens) {
info.totalInputTokens = (info.totalInputTokens || 0) + entry.message.usage.input_tokens;
}
if (entry.message.usage.output_tokens) {
info.totalOutputTokens = (info.totalOutputTokens || 0) + entry.message.usage.output_tokens;
}
}
// Check if this is first user message and description is missing
if (info && !info.description && entry.type === 'user' && entry.message?.content) {
// First try parent transcript (most reliable)
let description = await this.extractDescriptionFromParentTranscript(info.projectHash, info.sessionId, agentId);
// Fallback: extract smart title from the prompt content
if (!description) {
let text: string | undefined;
if (typeof entry.message.content === 'string') {
text = entry.message.content.trim();
} else if (Array.isArray(entry.message.content)) {
const firstContent = entry.message.content[0];
if (firstContent?.type === 'text' && firstContent.text) {
text = firstContent.text.trim();
}
}
if (text) {
description = this.extractSmartTitle(text);
}
}
if (description) {
// Check if this is an internal agent - if so, remove it
if (this.isInternalAgent(description)) {
this.removeAgent(agentId);
return;
}
info.description = description;
this.emit('subagent:updated', info);
}
if (await this._processDescription(entry, agentId, info)) return;
}
if (entry.type === 'progress' && entry.data) {
@@ -1416,7 +1417,6 @@ export class SubagentWatcher extends EventEmitter {
progressType: entry.data.type,
query: entry.data.query,
resultCount: entry.data.resultCount,
// Extract hook event info if present
hookEvent: entry.data.hookEvent,
hookName:
entry.data.hookName ||
@@ -1426,9 +1426,60 @@ export class SubagentWatcher extends EventEmitter {
};
this.emit('subagent:progress', progress);
} else if (entry.type === 'assistant' && entry.message?.content) {
// Handle both string and array content formats
if (typeof entry.message.content === 'string') {
const text = entry.message.content.trim();
this._processAssistantContent(entry, agentId, sessionId);
} else if (entry.type === 'user' && entry.message?.content) {
this._processUserContent(entry, agentId, sessionId);
}
}
private _processModelInfo(entry: SubagentTranscriptEntry, agent: SubagentInfo): void {
if (entry.type === 'assistant' && entry.message?.model && !agent.model) {
agent.model = entry.message.model;
agent.modelShort = this.extractModelShort(entry.message.model);
this.emit('subagent:updated', agent);
}
}
private _processTokenInfo(entry: SubagentTranscriptEntry, agent: SubagentInfo): void {
if (!entry.message?.usage) return;
if (entry.message.usage.input_tokens) {
agent.totalInputTokens = (agent.totalInputTokens || 0) + entry.message.usage.input_tokens;
}
if (entry.message.usage.output_tokens) {
agent.totalOutputTokens = (agent.totalOutputTokens || 0) + entry.message.usage.output_tokens;
}
}
private async _processDescription(
entry: SubagentTranscriptEntry,
agentId: string,
agent: SubagentInfo
): Promise<boolean> {
if (agent.description || entry.type !== 'user' || !entry.message?.content) return false;
const fallbackText = this.extractFirstTextContent(entry.message.content);
const description = await this._resolveDescription(
agent.projectHash,
agent.sessionId,
agentId,
agent.filePath,
fallbackText
);
if (description) {
if (this.isInternalAgent(description)) {
this.removeAgent(agentId);
return true;
}
agent.description = description;
this.emit('subagent:updated', agent);
}
return false;
}
private _processAssistantContent(entry: SubagentTranscriptEntry, agentId: string, sessionId: string): void {
const messageContent = entry.message!.content;
if (typeof messageContent === 'string') {
const text = messageContent.trim();
if (text.length > 0) {
const message: SubagentMessage = {
agentId,
@@ -1440,7 +1491,7 @@ export class SubagentWatcher extends EventEmitter {
this.emit('subagent:message', message);
}
} else {
for (const content of entry.message.content) {
for (const content of messageContent) {
if (content.type === 'tool_use' && content.name) {
// Store toolUseId for linking to results, with timestamp for TTL cleanup
if (content.id) {
@@ -1477,25 +1528,12 @@ export class SubagentWatcher extends EventEmitter {
agentInfo.toolCallCount++;
}
} else if (content.type === 'tool_result' && content.tool_use_id) {
// Extract tool result
const resultContent = this.extractToolResultContent(content.content);
const agentPendingCalls = this.pendingToolCalls.get(agentId);
const pendingCall = agentPendingCalls?.get(content.tool_use_id);
const toolName = pendingCall?.toolName;
// Delete after lookup to prevent memory leak
agentPendingCalls?.delete(content.tool_use_id);
const toolResult: SubagentToolResult = {
this.emitToolResult(
{ tool_use_id: content.tool_use_id, content: content.content, is_error: content.is_error },
agentId,
sessionId,
timestamp: entry.timestamp,
toolUseId: content.tool_use_id,
tool: toolName,
preview: resultContent.substring(0, MESSAGE_TEXT_LIMIT),
contentLength: resultContent.length,
isError: content.is_error || false,
};
this.emit('subagent:tool_result', toolResult);
entry.timestamp
);
} else if (content.type === 'text' && content.text) {
const text = content.text.trim();
if (text.length > 0) {
@@ -1504,17 +1542,19 @@ export class SubagentWatcher extends EventEmitter {
sessionId,
timestamp: entry.timestamp,
role: 'assistant',
text: text.substring(0, MESSAGE_TEXT_LIMIT), // Limit text length
text: text.substring(0, MESSAGE_TEXT_LIMIT),
};
this.emit('subagent:message', message);
}
}
}
}
} else if (entry.type === 'user' && entry.message?.content) {
// Handle both string and array content formats - also check for tool_result in user messages
if (typeof entry.message.content === 'string') {
const userText = entry.message.content.trim();
}
private _processUserContent(entry: SubagentTranscriptEntry, agentId: string, sessionId: string): void {
const messageContent = entry.message!.content;
if (typeof messageContent === 'string') {
const userText = messageContent.trim();
if (userText.length > 0 && userText.length < 500) {
const message: SubagentMessage = {
agentId,
@@ -1527,26 +1567,14 @@ export class SubagentWatcher extends EventEmitter {
}
} else {
// Check for tool_result blocks in user messages (common pattern)
for (const content of entry.message.content) {
for (const content of messageContent) {
if (content.type === 'tool_result' && content.tool_use_id) {
const resultContent = this.extractToolResultContent(content.content);
const agentPendingCalls = this.pendingToolCalls.get(agentId);
const pendingCall = agentPendingCalls?.get(content.tool_use_id);
const toolName = pendingCall?.toolName;
// Delete after lookup to prevent memory leak
agentPendingCalls?.delete(content.tool_use_id);
const toolResult: SubagentToolResult = {
this.emitToolResult(
{ tool_use_id: content.tool_use_id, content: content.content, is_error: content.is_error },
agentId,
sessionId,
timestamp: entry.timestamp,
toolUseId: content.tool_use_id,
tool: toolName,
preview: resultContent.substring(0, MESSAGE_TEXT_LIMIT),
contentLength: resultContent.length,
isError: content.is_error || false,
};
this.emit('subagent:tool_result', toolResult);
entry.timestamp
);
} else if (content.type === 'text' && content.text) {
const userText = content.text.trim();
if (userText.length > 0 && userText.length < 500) {
@@ -1563,7 +1591,6 @@ export class SubagentWatcher extends EventEmitter {
}
}
}
}
/**
* Extract text content from tool_result content field
-8
View File
@@ -17,14 +17,6 @@ import { getStore } from './state-store.js';
/**
* Events emitted by TaskQueue
*/
export interface TaskQueueEvents {
/** Fired when a task is added to the queue */
taskAdded: (task: Task) => void;
/** Fired when a task is removed from the queue */
taskRemoved: (taskId: string) => void;
/** Fired when a task's state changes */
taskUpdated: (task: Task) => void;
}
/**
* Priority queue for managing tasks with dependency support.

Some files were not shown because too many files have changed in this diff Show More