Compare commits

...
Author SHA1 Message Date
arkon 00721069e1 chore: version packages 2026-05-12 10:25:20 +02:00
arkonandClaude Opus 4.7 453a5383d2 test: cover hostname title (#82) and tmux size-query (#80)
Backfill the two regression gaps flagged on master after the recent
hostname-title and tmux-flicker fixes shipped without server-side
assertions.

* test/server-index-title.test.ts (8 tests) — exercises WebServer's
  index.html templating path: default os.hostname(), --title-hostname
  override, HTML-escape against `<script>`-style breakout, ampersand
  non-double-encoding, exact-once substitution, and byte-identical
  template-tail invariance.

* test/tmux-window-size-query.test.ts (15 tests) — mocks
  child_process.execFileSync and walks the helper through the
  browser-resize-between-attaches happy path, query-then-die race,
  zero/negative/empty/non-numeric output, plus argv-form/timeout
  assertions to lock down the no-shell-interpolation guarantee.

* src/session.ts — extracts the inline 14-line tmux size query into
  a named `queryTmuxWindowSize()` export so the test surface is a
  pure function. Behavior unchanged.

* src/web/public/notification-manager.js — Browser Notification API
  (layer 3) now uses `${this.originalTitle}: ${title}` so OS-level
  desktop pop-ups carry the same `codeman:<host>` prefix that the
  tab title and Web Push payloads already do, finishing the
  hostname plumb-through started in #82.

* CLAUDE.md, README.md — document the dual-CLI env-prefix discipline
  (CLAUDE_CODE_* vs OPENCODE_*), expand the xterm-zerolag-input
  duplication gotcha to mention the published-package side-effect,
  and note that the hostname prefix now applies uniformly to tab
  title, tab-flash, and OS notifications.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 10:23:44 +02:00
arkonandClaude Opus 4.7 e7b95ae579 test(routes): regression coverage for stripInkRedrawBloat
The clustering rewrite of stripInkRedrawBloat() shipped silently inside
the v0.6.7 "chore: version packages" commit (dcc814f). The previous
implementation discarded everything after the first VPA escape — silently
dropping 100KB+ of legitimate streamed response text on every long
Claude turn. The fix landed without any test coverage, so a regression
back to the old shape would be invisible until users noticed missing
conversation history.

Export the function (it's a pure (string)=>string helper) and add 12
tests covering:
  - The early-out paths (empty buffer, no VPAs, fewer than 10 VPAs)
  - Small clusters preserved (< MIN_BLOAT_SIZE = 32KB span)
  - Big clusters collapsed to a single trailing VPA
  - The silent-data-loss bug: response text BETWEEN two big clusters
    is preserved (input >280KB so any "keep just the tail" approach
    would push the response text out of its window — verified locally
    that a simulated old impl fails the assertion)
  - FRAME_GAP boundary on both sides (>8KB splits clusters; <=8KB merges)
  - Mixed small + big in the same buffer
  - Big cluster at end-of-buffer keeps the last frame
  - Idempotency: a second pass is a no-op
  - Realistic 200KB+ input shrinks by an order of magnitude

Total runtime ~12ms.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 10:11:47 +02:00
arkonandClaude Opus 4.7 56c2c29009 feat(push): plumb hostname-aware prefix into Web Push notifications
Closes the Web Push gap left by #82: in-page Notification API and tab
title flash both showed `codeman:<host>` after that PR, but OS-level
notifications dispatched via the service worker — the surface that
matters most when the tab is closed and the user is reading their
system notification center across multiple Codeman instances —
still hardcoded the literal "Codeman" prefix.

Service workers run in an isolated context with no access to
document.title or any in-page state, so the hostname has to ride
along in the push payload itself.

Server (server.ts:sendPushNotifications): emit `hostTitle: this.windowTitle`
in the JSON payload alongside the existing `title` (event-specific text
like "Permission Required"). The two stay separate so the SW can compose
them — the server knows the host, the SW knows the OS context.

Service worker (sw.js): compose `${hostTitle}: ${title}` when both
present, mirroring the in-page Notification format from
notification-manager.js. Fall back to `title || hostTitle || 'Codeman'`
so older servers (which omit hostTitle) keep working — the field is
purely additive on the wire.

Tests (test/push-payload-host-title.test.ts): mock the `web-push` module
via vi.hoisted(), instantiate WebServer without binding a port, stub
the push store with one fake subscription, and verify the JSON payload
shipped to webpush.sendNotification carries the right hostTitle for
both --title-hostname overrides and the os.hostname() default. Also
mirrors the SW's title-composition logic in a small helper so any
future change to the format breaks the test instead of being caught
only by users running multiple Codeman instances.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 10:03:59 +02:00
arkonandClaude Opus 4.7 7beec7194a fix(client): harden inline rename against CJK, mid-rename deletion, and double-fire
Three follow-up fixes to the inline rename input introduced in #81:

1. IME composition guard. Pressing Enter to confirm a Chinese pinyin
   candidate (or any IME composition) was committing the half-composed
   text as the session name. Skip the keydown handler when isComposing
   is true or when keyCode is the legacy 229 sentinel that older
   Safari/Edge versions report on the Enter that triggers compositionend.

2. Ghost tab on mid-rename deletion. If a session was deleted via SSE
   while its tab was being renamed, the render-skip flag suppressed
   _renderSessionTabs() and the orphaned <input> stayed on screen until
   blur — at which point the rename PUT 404'd against the dead session.
   Replace the boolean _inlineRenameActive with a _activeRename
   {sessionId, cancel} object so _cleanupSessionData can abort an
   in-flight rename targeting the deleted session, and finishRename
   skips the API call when the session is gone.

3. Stuck-flag risk. Move the settle-once guard into a closure-local
   `settled` boolean so blur / Enter / Escape / external cancel all
   converge to a single idempotent path. Register _activeRename only
   after the input is fully wired so a throw earlier in setup can't
   strand state.

Adds test/inline-rename.test.ts with 7 Playwright tests that drive
startInlineRename via page.evaluate() against a stubbed session and
synthetic .tab-name node — no real PTY/tmux needed, runs in ~1.3s.

Also fixes test/mobile/helpers/server.ts which imported the WebServer
via a path one directory short of the repo root, breaking the entire
mobile test suite under the main vitest config.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 09:57:08 +02:00
arkonandClaude Opus 4.7 dcc814f40c chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 09:18:14 +02:00
aakhterandClaude Opus 4.6 b7e94e7068 feat: hostname-aware window title (#82)
Set the browser tab title to codeman:${hostname} instead of the bare
"Codeman" literal. Useful for users running multiple Codeman instances
across hosts (laptop, dev box, NAS) — the OS hostname disambiguates
which tab points at which backend.

Implementation:

- src/cli.ts: new --title-hostname <hostname> flag overrides the
  detected hostname (handy for cosmetic naming or when os.hostname()
  returns something noisy).
- src/web/server.ts: WebServer now accepts an optional titleHostname
  constructor arg (defaults to os.hostname()), composes
  windowTitle = codeman:${titleHostname}, and serves / and
  /index.html by templating that title into the cached index.html
  template (with HTML escaping of the title text).
- src/web/public/notification-manager.js: title-flash logic now uses
  this.originalTitle instead of the hardcoded "Codeman" literal, so
  the tab flash respects the per-host title.
- scripts/browser-comparison.mjs + test/file-link-click.test.ts:
  expectations updated from === "Codeman" to a startsWith("codeman:")
  predicate so they pass regardless of host.

The new index.html templating is intentionally narrow — it only
substitutes the <title> tag and continues to serve everything else
from the static template. No JS-side title injection, so it works
without JavaScript and shows the correct title from the very first
paint.

Note: test/file-link-click.test.ts shows ~49 prettier-reformat lines
that are not part of the feature — they are pre-existing prettier
debt that the pre-commit hook required me to clear. The single
behavioral change is the browserAvailable line.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-12 09:11:33 +02:00
aakhterandClaude Opus 4.6 eade261763 fix(client): preserve inline rename input across tab re-renders (#81)
When the inline session-rename input is open, any incoming SSE event
that triggers renderSessionTabs() (a sibling session updating, a hook
firing, a status change) destroys the input element mid-keystroke and
the user loses what they were typing.

Add a _inlineRenameActive flag that:
- guards the two render paths (renderSessionTabs and
  _fullRenderSessionTabs) so they bail out early while a rename is
  in progress;
- is set true when the inline input mounts (session-ui.js);
- is cleared in finishRename, which then explicitly calls
  renderSessionTabs to restore the normal tab structure.

Also add a re-entrance guard at the top of finishRename so the blur
event and the Enter keydown do not both fire it (was a latent
double-call).

Drive-by: replace tabName.innerHTML = "" with explicit child removal.
The preceding textContent = "" already clears the element; this avoids
an innerHTML write on a node that takes user-supplied content on the
next line.

Follow-up to the inline-rename feature cherry-picked from #60.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-12 09:10:34 +02:00
arkonandClaude Opus 4.7 41a82fcf02 chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 22:41:11 +02:00
aakhterandClaude Opus 4.6 eecf74c001 fix: prevent tmux flicker on restart by matching existing window size (#80)
When a PTY client re-attaches to an existing tmux session, it currently
hardcodes the PTY size to 120x40 and tmux resizes the window to match.
The xterm.js client then resizes back to its actual viewport on the
next render tick, so every restart causes a visible flicker and loses
one repaint of buffer content.

Also remove the hardcoded `-x 120 -y 40` from `tmux new-session` so
initial size adapts to the first client.

Changes:
- session.ts: query existing window size via `tmux display -p
  #{window_width} #{window_height}` before pty.spawn, fall back to
  120x40 only if tmux is unreachable.
- tmux-manager.ts: drop -x/-y from new-session args.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-11 22:32:17 +02:00
arkonandClaude Opus 4.7 23b4dfcd82 chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 03:23:34 +02:00
arkonandClaude Opus 4.7 e017b275fe chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 13:16:05 +02:00
arkonandClaude Opus 4.7 e8a809ea80 chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 03:36:56 +02:00
Ark0NandClaude Opus 4.7 8006cc5db3 fix: allowlist opusContext1mEnabled in SettingsUpdateSchema (#78)
Same root cause as the thinkingEffort fix in #73: the schema is
.strict(), so unknown keys in PUT /api/settings are rejected with
INVALID_INPUT and the toggle never persists. Verified live:
pre-fix returned {"errorCode":"INVALID_INPUT"}, post-fix accepts.

The frontend has been reading and writing this key for a while
(settings-ui.js:336, :1137; session-ui.js:331), so saves were
silently failing — users never noticed because the load path
falls back to false on missing keys.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 03:29:48 +02:00
arkonandClaude Opus 4.7 0ded279b55 chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 03:25:53 +02:00
Tenggan ZhangandTeigen aa5724c390 feat: improve Resume Conversation UX on mobile (#77)
Default layout was a single nowrap row with path + date + size, which on
narrow screens truncated both the first prompt and the directory suffix
(where project names actually live). The /Users/ home shorthand was also
never applied on macOS.

Changes:
- Title now uses 2-line clamp so more of the first prompt is visible.
- Subtitle resolves workingDir against known cases: exact match shows
  "#caseName", subpath shows "#caseName/sub", otherwise falls back to
  basename. Case labels are styled distinctly.
- Normalize both /home/<user>/ and /Users/<user>/ prefixes to "~/".
- Each history item gets a "..." toggle that expands an in-place detail
  panel with the full prompt, full path, timestamp, size, and short
  session id. Collapses back on second click; clicking the card body
  still triggers resume.

Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
2026-04-28 03:14:56 +02:00
Tenggan ZhangandTeigen f21df2a9fb fix: eye icon follows /clear to the new Claude conversation (#76)
Interactive Claude CLI never emits session_id on stdout, so the
Session's _claudeSessionId stayed pinned to the pre-/clear jsonl and
the last-response viewer kept showing the old conversation.

Two complementary update paths:

- Session.adoptClaudeSessionId() — public setter mirroring the existing
  no-op-if-same guard. Called from POST /api/hook-event when Claude Code
  hooks carry data.session_id (works once hooks are configured).

- /api/sessions/:id/last-response now resolves the active id from
  ~/.claude/history.jsonl before reading the transcript. This is the
  only source-of-truth that does not require hooks, and we intentionally
  don't write hooks into arbitrary user repos.

History scan filters out sessionIds held by other Codeman sessions in
the same cwd, and validates via jsonl mtime to avoid inheriting a dead
prior session's id.

Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
2026-04-28 03:03:11 +02:00
e549e15cb8 feat(response-viewer): ASCII diagram wrap toggle, mobile code blocks, chrome-stripping fallback (#75)
* fix: restore clear message separation + proper table layout in response viewer

* fix: capture Claude CLI's real session ID + robust ANSI/CLI-chrome stripping in response viewer fallback

Session constructor seeded _claudeSessionId with Codeman's session.id as a
placeholder, and the message-driven update was gated on !_claudeSessionId —
meaning Claude CLI's actual session UUID was never adopted. This broke
/api/sessions/:id/last-response JSONL lookups, silently falling through to
the terminal-buffer path whose ANSI regex missed \x1b[>c / \x1b[>q queries.

- session.ts: update _claudeSessionId whenever a message's session_id differs
  from current (covers placeholder and stale-resume cases)
- app.js: extract _cleanTerminalBuffer with proper CSI regex (param bytes
  0x30-0x3F now covers > ? < =) plus a chrome filter for status bar,
  progress bar, spinner, shell prompt, and hint lines

* fix: wrap regular code blocks on mobile, keep ASCII diagrams rigid with scroll hint

* feat: add per-block wrap toggle on ASCII-diagram code blocks

* fix: wrap by default, pin toggle button outside scroll container

* fix: narrow diagram detection to box-drawing + block elements only

* feat: show last-response viewer eye icon on desktop too

The response viewer button was mobile-only via a display:none default with a
mobile.css override. Flip the default to inline-flex and drop the override so
the eye icon appears in the header on every form factor — desktop users get
the same quick "Last Response" pane as mobile.

* fix(response-viewer): restore HTML sanitizer + fix undefined `src` in _renderMarkdown

- `_renderMarkdown` referenced an undefined `src` (should be `text`),
  causing a ReferenceError on every markdown render. The try/catch
  swallowed it, so the new table-wrap and ASCII-diagram features
  never actually ran — output silently fell through to plain-text.
  app.js is excluded from ESLint, so this wasn't caught at lint time.
- `_sanitizeHtml` was removed when refactoring the response viewer,
  leaving `marked.parse()` output going straight into `innerHTML`
  without sanitization (XSS regression vs. master). Restored the
  helper and re-applied it before any post-processing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
Co-authored-by: arkon <arkon.85@hotmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 02:58:42 +02:00
Ark0N d07b59db4e Merge pull request #74 from TeigenZhang/refactor/envoverrides-tmux-export
refactor: pass envOverrides via tmux export instead of disk write
2026-04-28 02:45:11 +02:00
arkonandClaude Opus 4.7 a5a7e0c94c Merge master into refactor/envoverrides-tmux-export
Resolved conflict in src/web/public/session-ui.js by keeping this
PR's buildEnvOverrides() helper — it already covers both
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS (this PR) and
CLAUDE_CODE_EFFORT_LEVEL (added in #73), so the master-side inline
block is fully replaced.

Also fixed test/session-manager.test.ts MockSession to add a
getEnvOverridesForPersist() stub — without it,
SessionManager.updateSessionState's new call breaks 19 tests with
"TypeError: session.getEnvOverridesForPersist is not a function".

Verified: typecheck, lint, format:check, build, and
test/{session-manager,session-state,tmux-manager,tmux-restart-recovery}.test.ts
all pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 02:43:31 +02:00
Ark0N 79d7117e6d Merge pull request #73 from TeigenZhang/feat/thinking-effort
feat: thinking effort setting for new sessions (with xhigh/max)
2026-04-28 02:21:57 +02:00
arkonandClaude Opus 4.7 3cf486730b fix: allowlist thinkingEffort in SettingsUpdateSchema
Without this, PUT /api/settings rejects the new field with
INVALID_INPUT (schema is .strict()), so the dropdown's value
never persists.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 02:20:11 +02:00
Tenggan Zhang 996b096849 fix: prevent vertical scroll on mobile keyboard accessory bar (#72)
Thank you @TeigenZhang for the clean mobile fix!
2026-04-28 02:15:00 +02:00
ffa7fcf839 fix(tmux-manager): use '|' separator in reconcileSessions (#71)
* fix(tmux-manager): use '|' separator in reconcileSessions

Under non-tty execution contexts (launchd on macOS, systemd without TTY),
tmux emits '\t' in FORMAT strings as the literal two characters `\` + `t`
rather than as a tab. The parser's `line.indexOf('\t')` (a real tab char)
therefore never matches, `activeSessions` stays empty, `reconcileSessions`
returns `alive: []` / `discovered: []`, and `cleanupStaleSessions()` wipes
every entry in `state.json` — even though the underlying tmux sessions are
still alive. On the next startup the user sees an empty session list.

The bug reproduces reliably when codeman is launched via a user LaunchAgent
or a systemd unit without `TTYPath`. Interactive `npm run dev` hides it
because tmux's format parser does interpret `\t` when stdout is a TTY.

Fix: use `|` as the separator. tmux passes it through verbatim in every
environment, and `|` is not a valid tmux session-name character so it
cannot collide with the codeman-<uuid> / claudeman-<uuid> naming scheme.

* test(tmux-manager): cover parsePaneList separator contract

Extract the inline pane-list parser from `reconcileSessions` into an
exported `parsePaneList()` helper plus `PANE_LIST_SEP` / `PANE_LIST_FORMAT`
constants, so the '|' separator contract can be unit-tested directly.

The new tests lock in:
- Well-formed parsing into name -> pid Map
- Empty / blank-line / missing-separator handling
- Non-numeric pid and empty-name rejection
- A literal `\t` (backslash + t) in the input is NOT treated as a
  delimiter — guards against the launchd/systemd regression that
  motivated PR #71.
- Splitting on the first separator only.

No behavior change in `reconcileSessions`; the body now delegates to the
helper.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
Co-authored-by: arkon <arkon.85@hotmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-28 02:11:11 +02:00
Teigen a1c69f7405 refactor: pass envOverrides via tmux export instead of disk write
CLAUDE_CODE_EFFORT_LEVEL (and any CLAUDE_CODE_* / OPENCODE_* key) now flows:
  UI dropdown → POST /api/sessions { envOverrides }
             → new Session({ envOverrides })
             → this._envOverrides
             → tmux-manager.buildEnvExports appends `export KEY=<shellescape(VALUE)>`

Previously the API wrote envOverrides to <case>/.claude/settings.local.json, which
created stale state (UI dropdown disagreeing with disk) and polluted user project
directories. Now envOverrides are ephemeral spawn-time state, preserved across
respawnPane cycles via this._envOverrides and across server restart via
SessionState.envOverrides in state.json.

Also removes the now-unused updateCaseEnvVars import from session-routes.ts.
2026-04-24 09:49:52 +08:00
Teigen 534899bc2b feat: add xhigh effort option and /effort max mobile shortcut
Add XHigh option to Thinking Effort dropdown (between High and Max),
and add a Max quick button to the mobile keyboard accessory bar that
sends /effort max as a slash command.
2026-04-24 09:48:53 +08:00
Teigen 03d91ffddd feat: add thinking effort setting for new sessions
Allow configuring CLAUDE_CODE_EFFORT_LEVEL (low/medium/high/max) from
Settings → Claude Permissions. Applied as envOverride on session creation.
2026-04-24 09:48:18 +08:00
arkonandClaude Opus 4.7 6280998bd8 chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 20:22:30 +02:00
arkonandClaude Opus 4.7 02e2f3e8b5 refactor: remove dead code and narrow internal exports (knip sweep)
Knip-driven cleanup. All changes verified with tsc --noEmit, lint, and
build.

Removed (zero consumers):
- VERIFICATION_PROMPT constant + its barrel re-export
- createInitialOrchestratorPersistState factory
- transcriptWatcher singleton export
- createAnsiPatternFull / createAnsiPatternSimple factories
- TimerInfo interface + unused AiCheckResult/AiPlanCheckResult imports
  in respawn-controller.ts
- 35 unused Zod z.infer \`*Input\` types in schemas.ts
- Dead re-exports: SessionMode from session.ts, AuthSessionRecord from
  web/ports/index.ts, EnhancedPlanTask/CheckpointReview from
  ralph-tracker.ts, 7 unused entries in utils/index.ts
- 14 event/config interfaces that lived only as JSDoc hints (no TS type
  position usage): Session/Respawn/RalphLoop/RalphTracker/
  SessionManager/SessionAutoOps/Subagent/TaskQueue/TaskTracker/
  TranscriptWatcher/Image/OrchestratorLoop Events + RespawnPreset +
  SessionOutput

Narrowed to module scope (kept but no longer exported):
- buildPermissionArgs in session-cli-builder.ts
- 28 type/interface declarations used only within their own file:
  Ai{Idle,Plan}Check{Config,State}, BashToolParser{Events,Config},
  FileStream/CreateStream{Options,Result}, PlanSubagentEvent,
  SubagentCallback, RalphLoopConfig, RalphLoop{Events,Options},
  ActiveTimerInfo, DetectionStatus, ActionLogEntry, AutoOpsCallbacks,
  TunnelStatus, Timer/LRUMap/StaleExpirationMap Options, AuthState,
  SessionListenerDeps, SseStreamManagerDeps, and 8 more

Docs: CLAUDE.md advice for global-regex `lastIndex` now points to the
remaining `execPattern()` helper instead of the deleted factories.

Knip delta: unused files 42→0, unused exports 161→16, unused types 92→0.
The 16 remaining exports are a mobile-test helper toolkit intentionally
kept for upcoming tests.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:57:00 +02:00
arkonandClaude Opus 4.7 41300f0a34 chore: add knip config for dead-code detection
knip.json declares the real entry points (scripts, tests, Remotion roots)
and the devDeps invoked only as external CLIs (esbuild for build,
agent-browser/remotion via npx) so future scans surface only true
findings.

Add \`npm run knip\` as the canonical invocation.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:56:31 +02:00
arkonandClaude Opus 4.7 adbc083426 chore: remove dead test and script files
Dead-code sweep via knip. All files below have zero importers and were
leftovers from the local-echo overlay exploration or duplicated by files
under test/mocks/.

Deleted:
- test/respawn-test-utils.ts (728-line duplicate of test/mocks/*)
- test/input-echo-test.mjs
- test/local-echo-*.mjs (7 files)
- test/manual/*.mjs (10 files; dir removed)
- scripts/remotion/components/TerminalScreen.tsx (unused Remotion demo)

Also cleaned stale JSDoc references to the removed
respawn-test-utils.ts in test/mocks/mock-session.ts and
test/mocks/test-helpers.ts.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:56:22 +02:00
arkonandClaude Opus 4.7 3754bcd1aa docs: tighten CLAUDE.md — update counts and CI line
- CI line: include `check:lockfile` step now run in workflow
- Frontend: app.js is ~2.9K lines (was ~2.8K)
- Types: document `src/types/index.ts` as the barrel (14 domain files) plus `src/types.ts` root re-export
- API routes: updated handler counts (128 total; sessions 27, cases 9)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:55:56 +02:00
arkonandClaude Opus 4.7 f2f909ca9c docs: tighten CLAUDE.md — fix counts, remove footer redundancy
Fix stale counts (types 14 to 15, SSE events ~118 to ~120). Remove
redundant footer sections (References list duplicated inline citations;
Common Workflows bullets were self-evident or already stated; Tunnel and
Memory Leak Prevention folded into neighboring sections). 251 to 234 lines.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:14:16 +02:00
arkonandClaude Opus 4.7 1c3f2f6571 docs: tighten CLAUDE.md and archive 22 completed plan docs
CLAUDE.md: fix stale counts (types 14 to 15, SSE events ~118 to ~120),
remove redundant footer sections (References list duplicated inline citations;
Common Workflows bullets were self-evident or already stated; Tunnel/Memory
Leak Prevention folded into neighboring sections). 251 to 234 lines.

Move 22 completed implementation/phase/audit plans to docs/archive/ via
git mv so history is preserved. Living reference docs remain in docs/.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:14:07 +02:00
arkonandClaude Opus 4.7 8da1bdf690 chore: prevent package-lock.json version drift permanently
Makes the drift that PR #70 caught impossible to repeat:

- `version-packages` script now runs `changeset version && npm install
  --package-lock-only && check-lockfile-sync`, so the lockfile is always
  regenerated and verified as part of consuming a changeset
- New `scripts/check-lockfile-sync.mjs` compares package.json#.version against
  package-lock.json's root and packages[""] version fields (npm ci does not
  enforce these, which is why the prior drift slipped through CI)
- CI now runs `npm run check:lockfile` on every push/PR — any future drift
  fails the build before merge
- COM workflow in CLAUDE.md collapsed back to a single release-bump step now
  that lockfile sync is automatic

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:01:51 +02:00
arkonandClaude Opus 4.7 a93325b312 fix: sync package-lock.json to 0.6.0 and document lockfile step in COM workflow
package.json has been at 0.6.0 since release, but package-lock.json stayed at
0.3.11 because `npm run version-packages` (changesets) does not regenerate the
lockfile. This left `npm ci` broken against the committed state.

Also adds step 4 (`npm install --package-lock-only`) to the COM workflow in
CLAUDE.md so future releases keep the lockfile in sync automatically.

Credit to @Matt2012 (#70) for catching the lockfile drift.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:00:10 +02:00
arkonandClaude Opus 4.7 2d03e4efc9 docs: update CLAUDE.md and README for clipboard API and tab shortcuts
Reflect changes from PRs #65–#68: bumped route/SSE counts, added
Ctrl+Shift+{/} (tab reorder), Alt+1-9 (tab switch), Ctrl+Shift+V
(voice), and POST /api/clipboard to the keyboard and API references.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-18 04:16:36 +02:00
arkonandClaude Opus 4.7 ab7c502c2a chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-18 04:13:26 +02:00
Ark0N 546bbcbe7c Merge pull request #68 from aakhter/feat/clipboard-api
feat: add clipboard API for remote browser clipboard access
2026-04-18 04:10:31 +02:00
Ark0N 774d5ff321 Merge pull request #67 from aakhter/feat/tab-badges
feat: improve active tab visibility and add Alt+N number badges
2026-04-18 04:08:09 +02:00
Ark0N 98ceb5da1d Merge pull request #66 from aakhter/feat/tab-reorder-shortcuts
feat: add Ctrl+Shift+{/} to reorder session tabs
2026-04-18 04:07:01 +02:00
Ark0N 34fb5e49f8 Merge pull request #65 from aakhter/fix/android-shift-double-char
fix: prevent double character input on Android Shift+key
2026-04-18 04:05:21 +02:00
arkonandClaude Opus 4.7 002cf81b1e chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-18 04:02:02 +02:00
Aamer Akhter 9b4aab2502 feat: add clipboard API for remote browser clipboard access
POST /api/clipboard with { text } broadcasts to all connected browsers
via SSE. Browser attempts navigator.clipboard.writeText, falling back
to a modal with manual copy button if blocked.

Enables remote clipboard workflows: CLI tools can push text to the
user's browser clipboard across the network.
2026-04-12 10:07:15 -04:00
Aamer Akhter 829c797726 feat: improve active tab visibility and add Alt+N number badges
- Active tab: bright green border with color-matched glow per session color
- Tab number badges (1-9) showing Alt+N shortcut hints
- Badges update on tab reorder and re-render
2026-04-12 10:04:34 -04:00
Aamer Akhter 85da3bb898 feat: add Ctrl+Shift+{/} keyboard shortcuts to reorder session tabs
Matches WezTerm/terminal emulator tab reorder conventions. Uses the
existing sessionOrder array and saveSessionOrder() for persistence.
2026-04-12 10:02:12 -04:00
Aamer Akhter 29b2653801 fix: prevent double character input on Android Shift+key
On Android tablets, pressing Shift+A produces "AA" because the input
event listener re-sends characters that xterm already processed via
its keydown handler. Track keydown timestamps and skip input events
that fire within 50ms of a handled keydown.

Only affects touch devices (listener gated by isTouchDevice()).
2026-04-12 10:01:29 -04:00
arkonandClaude Opus 4.6 7b8b175133 chore: version packages
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 21:01:32 +02:00
Ark0N c3027b21e1 Merge pull request #63 from Typhon0/fix/quick-start-linked-cases
Good catch — quick-start was ignoring linked-cases.json. Thanks @Typhon0!
2026-04-11 20:58:57 +02:00
Loïc Sculier 0b231edd43 Fix quick-start to resolve linked cases before codeman-cases fallback
/api/quick-start was always resolving caseName against CASES_DIR,
ignoring any entries in ~/.codeman/linked-cases.json. This caused
sessions to start in ~/codeman-cases/<name> even when the case was
linked to an external project directory.

Fix: read linked-cases.json first and prefer that path, falling back
to validatePathWithinBase only when no link is found.
2026-04-11 17:45:41 +02:00
arkon ea1c2ee4ec chore: version packages 2026-04-11 07:21:18 +02:00
arkonandClaude Opus 4.6 b4a808adcf fix: security hardening and cleanup from community PR cherry-picks
- Add HTML sanitizer for markdown rendering (XSS prevention)
- Switch service worker to network-first caching (deploys take effect immediately)
- Sanitize Content-Disposition filenames (header injection prevention)
- Expose session.muxName getter, replace unsafe `as any` cast
- Static import for execFile, update CLAUDE.md keyboard shortcuts

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 07:20:09 +02:00
arkonandClaude Opus 4.6 f3cbe9bca6 feat: cherry-pick keyboard UX and file download from community PRs
Cherry-picked from PR #60 (keyboard UX) and PR #61 (file download):

- Alt+1-9 session switching
- Disable Ctrl+K (too easy to trigger accidentally)
- Session rename with prefix preservation (w1-case: description)
- Shift+Enter / Ctrl+Enter multiline input via tmux send-keys -H
- Android virtual keyboard fix for non-composition input
- File download button in browser file explorer (?download=true)

Dropped from PR #60: stale package-lock.json, upload popup (missing upload.html)
Dropped from PR #61: standalone /api/download endpoint (arbitrary fs access)
Fixed from PR #60: execFileSync replaced with async execFile

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-11 06:59:50 +02:00
Ark0N a11bcb0029 Merge pull request #59 from aakhter/feat/pwa-support
feat: add PWA support for Android/iOS home screen install
2026-04-11 06:54:06 +02:00
Ark0N 47fd9a922f Merge pull request #58 from ToRvaLDz/master
feat: add named Cloudflare tunnel support
2026-04-11 06:53:56 +02:00
Ark0N 14f7d8298d Merge pull request #62 from TeigenZhang/feat/mobile-response-viewer
feat: mobile response viewer with markdown rendering
2026-04-11 06:53:47 +02:00
Teigen d32f4debb2 feat: markdown rendering for response viewer
Add marked.js (39KB) for rich text display in the response viewer.
Renders headings, code blocks, lists, tables, blockquotes, and
inline formatting with dark theme styling.

Falls back to escaped plain text if marked.js fails to load.
2026-04-10 19:12:06 +08:00
Teigen 3cb7b510f8 feat: mobile response viewer — read full Claude responses via native scroll
Claude Code's Ink framework uses alternate screen buffer + VPA cursor
positioning, resulting in near-zero xterm.js scrollback on mobile.
Instead of fighting terminal scrollback, this adds a native scrollable
overlay that reads structured responses from Claude's JSONL transcripts.

- New API: GET /api/sessions/:id/last-response reads transcript JSONL
  - ?context=full returns full conversation thread (user + assistant)
  - Fallback to terminal buffer with ANSI stripping if no transcript
- Response viewer panel: bottom sheet with native iOS/Android scroll
- "More" button loads full conversation context as threaded view
- Eye icon in header bar (mobile only), no toolbar space impact
2026-04-10 19:07:02 +08:00
Aamer AkhterandClaude Opus 4.6 6a12a72c9c local: add PWA support for Android home screen install
Add app icons, update manifest with icon entries, and add app-shell
caching to service worker for offline/instant startup.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-04 22:15:24 -04:00
Marco Migozzi f1a126efeb feat: add named Cloudflare tunnel support
Add named tunnel mode alongside existing quick tunnel, with systemd
service and setup helper. All tunnel parameters are configurable via
environment variables:

  CLOUDFLARED_TUNNEL_NAME   — tunnel name (default: codeman)
  CLOUDFLARED_TUNNEL_ID     — tunnel UUID (from: cloudflared tunnel list)
  CODEMAN_TUNNEL_HOSTNAME   — public hostname

Backward compatible: ./tunnel.sh [start|stop|status|url] still works.
2026-04-04 13:41:28 +02:00
Marco Migozzi 12fd780af8 feat: add named Cloudflare tunnel support
Add named tunnel mode alongside existing quick tunnel, with systemd
service and setup helper. Tunnel ID and hostname are configurable via
CLOUDFLARED_TUNNEL_ID and CODEMAN_TUNNEL_HOSTNAME env vars.
2026-04-04 13:37:20 +02:00
Teigen fd74a42933 Merge remote-tracking branch 'origin/master' into dev 2026-04-04 08:07:15 +08:00
arkonandClaude Opus 4.6 7101e64800 refactor: restructure repo for cleaner GitHub landing page
Reduce visible top-level items from 21 to 14:
- Untrack test-results/, tmp/, public symlink (added to .gitignore)
- Move agent-teams/ → docs/agent-teams/
- Move mobile-test/ → test/mobile/
- Move tools/remotion/ → scripts/remotion/
- Move eslint.config.js, vitest.config.ts → config/

All path references updated across CLAUDE.md, package.json,
.prettierignore, vitest configs, and capture scripts.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-03 16:23:54 +02:00
Teigen 1b10d9b733 feat: mobile logo, expandable history, fix session resume
- Show Codeman logo on mobile as compact home button (was hidden)
- Add "Show More" button for history sessions (initial 4, expand all)
- Deduplicate by projectKey instead of workingDir (lossy decode fix)
- Fix project key decoding: handle '_' encoded as '-' with look-ahead
- Pre-validate resumeSessionId before passing to Claude CLI
- Apply content validation to all session files regardless of size
2026-04-03 11:02:43 +08:00
arkon 5078f5251d chore: version packages 2026-04-03 04:28:36 +02:00
arkonandClaude Opus 4.6 196af8fba7 fix: allow bracket chars in model flag for opus[1m] context window
The model validation regex rejected brackets, silently dropping models
like opus[1m]. Also quote the model flag to prevent bash glob expansion.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 04:27:32 +02:00
arkonandClaude Opus 4.6 a9b22b86a4 docs: use launchctl bootstrap instead of deprecated load
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 04:20:42 +02:00
arkonandClaude Opus 4.6 28cace5858 docs: clean up README install and service sections
Remove fork/branch install instructions and env vars table for cleaner
first impression. Reformat systemd and launchd service blocks as
readable multi-line heredocs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 04:19:45 +02:00
arkon 0a594b61bd chore: version packages 2026-04-03 04:17:33 +02:00
arkon 89d787a949 chore: version packages 2026-04-03 04:01:08 +02:00
arkonandClaude Opus 4.6 bd9797b68c fix: sanitize case names from filesystem to prevent XSS in inline handlers
Filter readdir and linked-case names through /^[a-zA-Z0-9_-]+$/ before
returning them from GET /api/cases. Prevents XSS via maliciously-named
directories reaching frontend inline onclick handlers where escapeHtml
is insufficient (HTML-decoded back to quotes before JS execution).

Also fix misleading "Drag or use arrows" hint (no drag-and-drop exists).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 03:58:29 +02:00
Ark0N 0e6cd94312 Merge pull request #56 from TeigenZhang/feat/case-manage-reorder-delete
feat: add case reorder and delete in Manage tab
2026-04-03 03:53:25 +02:00
arkonandClaude Opus 4.6 24a6f1cac8 chore: remove accidentally committed build artifact and dev-specific script
Remove dist/state-store.js (compiled build artifact that should not be tracked)
and scripts/claudeman-launchd-wrapper.sh (developer-specific launchd wrapper
with hardcoded paths) that were included in #55.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-03 03:51:00 +02:00
Ark0N 8e679a280b Merge pull request #55 from TeigenZhang/fix/auto-attach-on-restart
fix: auto-attach PTY on server restart
2026-04-03 03:50:31 +02:00
Teigen c642689bbd feat: add case reorder and delete in Manage tab
Add a "Manage" tab to the create-case modal with up/down reorder
buttons and delete for each case. Linked cases are unlinked (folder
preserved); CASES_DIR cases are permanently deleted.

Backend:
- DELETE /api/cases/:name — unlink or delete
- PUT /api/cases/order — persist ordering to settings.json
- GET /api/cases now respects saved caseOrder

Frontend:
- Third "Manage" tab in createCaseModal with case list
- Delete button in mobile case picker bottom sheet
- SSE events: case:deleted, case:order-changed
2026-04-03 09:45:19 +08:00
Teigen 28a6247c27 fix: auto-attach PTY to surviving tmux sessions on server restart
Previously, restoreMuxSessions() only created Session objects without
attaching PTY processes. Sessions stayed at pid=null until the client
manually selected them, causing terminals to appear "closed" after deploy.

Now the server calls startInteractive() for each recovered session during
startup, so all sessions resume capturing output immediately. The frontend
auto-attach condition is also relaxed from (pid===null && status==='idle')
to (pid===null && !_ended) as a safety net for edge cases.
2026-04-02 21:35:00 +08:00
Teigen 0ceb455c4b feat: add Ctrl+O button to mobile keyboard accessory bar 2026-04-02 21:13:17 +08:00
Teigen e51117dfa9 feat: add left/right arrow buttons to mobile keyboard accessory bar
Support cursor left/right movement on mobile, using the same blue
accessory-btn-arrow style as the existing up/down arrows.
2026-04-02 15:19:15 +08:00
Teigen 13d41cf7c7 feat: add Tab, Esc, ⌥Enter buttons to mobile keyboard accessory bar
- Add Tab (forward), Esc, and Option+Enter (newline) buttons
- Reorder buttons: ↑ ↓ 📋 Tab ⇧Tab ⌥Enter Esc /init /clear /compact dismiss
- Unify dismiss button style with arrow buttons (was oversized with custom class)
- Remove unused .accessory-btn-dismiss CSS rules
2026-04-02 14:32:54 +08:00
Teigen 2c7557d002 fix state store temp file collisions 2026-04-02 14:24:10 +08:00
arkonandClaude Opus 4.6 53b473708f fix: macOS support — HTML cache, launchd service, trust dialog
Three fixes for macOS deployments:

1. HTML cache bug: @fastify/static with preCompressed serves .html.br/.html.gz
   files, so path.endsWith('.html') missed them — HTML got 1-year immutable
   cache headers instead of no-cache, causing stale pages after deploys.

2. Installer launchd support: macOS now gets proper LaunchAgent setup (like
   systemd on Linux). Removes competing LaunchDaemons to prevent duplicate
   services fighting over the port. Update/uninstall also handle launchd.

3. Trust dialog auto-accept: Claude CLI 2.x shows a workspace trust prompt
   on first launch per directory. Sessions detect "trust this folder" in PTY
   output and auto-send Enter, preventing sessions from hanging on startup.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-01 08:51:37 +02:00
arkonandClaude Opus 4.6 2cba393ae5 fix: installer fails on macOS when piped via curl | bash
When running `curl | bash`, stdin is the pipe, not the terminal.
Homebrew and sudo need TTY access to prompt for the password.
Redirect /dev/tty as stdin for these subprocesses.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-31 19:34:39 +02:00
arkon 64b8ea30b2 chore: version packages 2026-03-31 03:10:22 +02:00
Ark0N 5743af3339 Merge pull request #52 from TeigenZhang/feat/default-model-support
Thanks for the contribution @TeigenZhang! Clean, well-scoped change — applied consistently across all session creation paths. 🎉
2026-03-30 19:19:04 +02:00
Teigen 2011bd8d89 fix: terminal flicker regression — move viewport clear inside dimension guard
Three fixes from the WIP flicker branch that were lost during master merges:

1. Move viewport clear (\x1b[3J\x1b[H\x1b[2J) inside the dimension-change
   guard so it only fires when cols/rows actually change. Previously every
   resize event cleared the screen even at identical dimensions, causing
   visible flicker with no subsequent Ink redraw to repaint.

2. Sync _lastResizeDims in sendResize() so restoreTerminalSize() doesn't
   trigger a redundant viewport clear on the next throttledResize tick.

3. Add didScroll tracking to touch events — tap (no scroll) now refocuses
   xterm's hidden textarea, fixing mobile keyboard input routing after
   tapping the terminal area.
2026-03-30 16:45:20 +08:00
Teigen b76724690d Merge branch 'feat/mobile-shift-tab' into dev 2026-03-30 16:45:16 +08:00
Teigen f277f9664c feat: add Shift+Tab button to mobile keyboard accessory bar
Mobile users cannot press Shift+Tab on virtual keyboards. Add a ⇧Tab
button that sends the escape sequence (\x1b[Z) to the PTY, enabling
mode switching on mobile devices.

Also fix accessory bar overflow on narrow screens by making it
horizontally scrollable with hidden scrollbar.
2026-03-30 16:44:43 +08:00
Teigen cd49171bbc feat: support "Default (CLI default)" option for model selection
Allow users to leave the default model unset, so sessions use whatever
the Claude CLI defaults to rather than forcing a specific model.

- Add empty-value "Default (CLI default)" option to the model dropdown
- Treat empty string as undefined when passing model to Session
- Apply consistently across session creation, quick-start, and Ralph
2026-03-30 16:44:00 +08:00
arkon 0f57342b10 chore: version packages 2026-03-29 05:10:16 +02:00
arkonandClaude Opus 4.6 e1f0ac993a fix: default new sessions to opus[1m] (1M context) instead of opus (200k)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-29 05:09:24 +02:00
arkon a84ef52992 chore: version packages 2026-03-28 16:49:55 +01:00
arkonandClaude Opus 4.6 692c894760 fix: correct process tree detection and prevent timer starvation
1. Rewrote getActiveChildProcesses() to use a single `ps --ppid` call
   instead of two-level pgrep. The pane PID is typically claude itself
   (bash exec'd into it), not a bash wrapper — so direct children of
   pane_pid ARE the tool processes.

2. Added timer restart in tryStartAiCheck() when skipping due to child
   processes. Without this, the pre-filter and no-output timers (both
   one-shot) would never fire again, permanently stalling idle detection
   for sessions with silent long-running processes.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-28 05:14:39 +01:00
arkonandClaude Opus 4.6 ad0acb6d58 feat: detect active child processes to prevent false idle during running tools
When Claude Code spawns bash tools (test suites, builds, servers), the
respawn controller could falsely detect idle if terminal output paused.
Now checks the process tree for active children of the Claude process
before triggering AI idle checks or confirming idle state.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-28 04:34:53 +01:00
arkonandClaude Opus 4.6 d866c8f30e chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-27 01:43:59 +01:00
arkonandClaude Opus 4.6 28537de39d refactor: pass 3 — extract helpers, split long functions, deduplicate patterns
Backend:
- subagent-watcher: split 176L processEntry() into 5 focused methods; extract
  _resolveDescription() deduplicating 3 call sites for description acquisition
- bash-tool-parser: split 152L processCleanLine() into 4 handlers; extract
  _createActiveTool() factory and _scheduleAutoRemove() helper
- session: extract _setupOrAttachMuxSession() deduplicating ~80L between
  startInteractive/startShell; extract _handleTerminalOutput()
- respawn-controller: split 180L handleTerminalData() into 3 detection layers;
  data-driven validation loop replacing 9 individual calls
- plan-orchestrator: extract _extractJsonFromResponse(), _emitAgentFailure(),
  _formatResearchSection() helpers
- orchestrator-loop: extract _finalizeTask() unifying task completion/failure;
  _clearTimer() utility for correct clearInterval/clearTimeout dispatch
- ralph-status-parser: config-driven FIELD_PARSERS[] replacing 8 near-identical
  field-matching blocks; split updateCircuitBreaker() into focused handlers
- state-store: extract _mergeWithInitialState() and _resetCircuitBreaker()

Frontend:
- app.js: add _notifySession() helper used by 18 call sites across 5 modules
- panels-ui.js: extract _addActivityEntry() replacing 4 duplicate blocks
- settings-ui.js: extract _updateTunnelUrlRow() deduplicating 2 blocks
- ralph-panel.js, respawn-ui.js: convert to _notifySession()

Routes:
- route-helpers: add toggleService() helper
- system-routes: use toggleService() for watcher toggles; extract collectActiveTokens()
- orchestrator-routes: data-driven EVENT_MAP replacing 10 identical listeners

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-26 22:50:25 +01:00
arkonandClaude Opus 4.6 ba09184efa fix: wizard "No JSON found" — Claude CLI stream-json returns empty result field
Claude CLI's --output-format stream-json now returns "result": "" in the result
message. The actual response text lives in assistant message text blocks, which
_textOutput correctly accumulates. runPrompt() was returning the empty
resultMsg.result without falling back to _textOutput.value.

Also improved plan-orchestrator JSON extraction to try code-block-wrapped JSON
first (```json {...} ```) before the greedy regex, plus debug logging.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-25 23:53:57 +01:00
arkonandClaude Opus 4.6 93719b41cd chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-25 23:34:14 +01:00
arkonandClaude Opus 4.6 a448983be3 refactor: pass 2 — extract shared helpers and simplify patterns
app.js:
- Add _clearTimer() helper replacing 11 inline clearTimeout patterns
- Add _isStaleSelect() helper for generation check + cleanup
- Replace 11 keyboard shortcut if-blocks with data-driven lookup table
- Extract _cleanupPreviousSession() from selectSession() (~75 lines)
- Extract _resetAllAppState() from handleInit() (~75 lines)

tmux-manager:
- Extract buildEnvExports() eliminating duplication in createSession/respawnPane
- Extract buildPathExport() for CLI path resolution
- Extract _configureOpenCode() for OpenCode setup

routes:
- Add readJsonConfig() to route-helpers, replacing 5 inline JSON-read patterns
- Add validateSessionFilePath() to route-helpers, replacing 2 identical path
  traversal validation blocks in file-routes

session-auto-ops:
- Convert executeWhenIdle() from 8 positional params to options object
- Extract validateThreshold() for shared compact/clear validation

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-25 23:32:28 +01:00
arkonandClaude Opus 4.6 3145eac6d9 refactor: extract helper methods to reduce duplication and improve readability
DRY up repeated patterns across 7 core files:
- state-store: extract serializeState() and split assembleStateJson() into 3 focused methods
- session: extract _resetBuffers(), _clearAllTimers(), _handleJsonMessage()
- ralph-tracker: extract completeAllTodos() (was 4x duplicated), emitValidationWarning(), similarity constants
- subagent-watcher: extract markSubagentAsCompleted(), extractFirstTextContent(), emitToolResult(), findOldestInactiveAgent()
- respawn-controller: extract recoveryResetToWatching(), canAutoAccept(), formatRemainingSeconds(), validatePositiveTimeout()
- tmux-manager: replace 15 path.includes() checks with single UNSAFE_PATH_CHARS regex
- session-auto-ops: extract executeWhenIdle() shared retry helper for checkAutoCompact/checkAutoClear

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-25 23:21:38 +01:00
arkonandClaude Opus 4.6 e3c609f5f0 test: add coverage for lastUsedCase partial update and strict schema rejection
Tests that partial PUT /api/settings with just lastUsedCase works correctly
and that including modelConfig triggers strict Zod schema rejection (the bug
fixed in #49).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-25 13:39:47 +01:00
Tenggan Zhang 52e774f83c fix: case selection not persisting across page refresh (#49)
Thank you for the clean fix! The root cause analysis in the PR description was excellent — the strict Zod schema rejecting modelConfig during the GET-then-PUT pattern was a subtle bug.
2026-03-25 13:39:16 +01:00
arkon 82d08df53f chore: version packages 2026-03-25 00:18:14 +01:00
arkonandClaude Opus 4.6 2709b2fe49 feat: make buffer size limits configurable via environment variables
Allow overriding MAX_TERMINAL_BUFFER_SIZE, TRIM_TERMINAL_TO, MAX_TEXT_OUTPUT_SIZE,
TRIM_TEXT_TO, and MAX_MESSAGES via CODEMAN_* env vars, falling back to existing
defaults. Enables users with fewer sessions or more RAM to tune buffer sizes
without patching source.

Closes #48

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-24 18:40:40 +01:00
Ark0N b1d3b27e5b Merge pull request #47 from TeigenZhang/fix/mobile-cjk-input-and-layout
fix: mobile CJK input, terminal flicker, and layout overflow
2026-03-24 18:22:48 +01:00
Teigen 47963b54fa fix: mobile CJK input, terminal flicker, and layout overflow
Terminal flicker:
- Skip buffer-recovered/clear-terminal events during active buffer load
  to prevent competing clear+rewrite cycles (app.js)
- Move viewport+scrollback clear inside dimension-change guard so resize
  without actual SIGWINCH doesn't blank the terminal (terminal-ui.js)
- Sync _lastResizeDims on explicit resize to prevent redundant clears

CJK input rewrite (input-cjk.js):
- Use InputEvent.inputType to distinguish insertText (final) from
  insertCompositionText (tentative) — fixes Chinese punctuation and
  English text being swallowed during Android IME composition
- Remove isComposing guard on Enter so it always sends
- Phantom character (U+200B) keeps textarea non-empty so Android
  long-press backspace generates continuous deleteContentBackward
  events at the keyboard's native repeat rate

CJK input settings:
- Add "CJK Input" toggle in Settings > Input (index.html, settings-ui.js)
- Store as device-specific setting (cjkInputEnabled), not synced to server
- Replace INPUT_CJK_FORM env var dependency with user-controlled setting
  (env var still works as server override)

Mobile layout:
- Fix welcome screen overflow on phones by constraining .welcome-content
  to calc(100vw - 1.5rem) (mobile.css)
- Move xterm helper textarea on-screen for touch devices to fix iOS
  keyboard input (styles.css)
- Focus terminal synchronously in user-gesture context for iOS Safari
  keyboard activation (session-ui.js, app.js)
- Refocus terminal on tap (not scroll) in touch handler (terminal-ui.js)
2026-03-24 09:22:06 +08:00
arkonandClaude Opus 4.6 b7c3c30c8c fix: send Ctrl+L after tab switch to clear stale Ink CUP frames
Tailed terminal buffers contain multiple CUP-positioned Ink frames from
different time points. When replayed in xterm, old frames at viewport
positions not covered by the latest frame persist as ghost content
(e.g. duplicate "bypass permissions" bars). After buffer load, send
Ctrl+L via the session input API to trigger a full Ink redraw, which
overwrites all stale frame content with the correct current state.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-23 16:59:48 +01:00
arkonandClaude Opus 4.6 0d80524f10 fix: prevent duplicate terminal output on tab switch to busy sessions
Two fixes for the tab-switching corruption bug:

1. _finishBufferLoad() now discards queued SSE events instead of flushing
   them. The loaded API buffer is the source of truth — queued events
   overlap with it, and flushing them writes duplicate Ink cursor-up
   redraws that corrupt the terminal display (garbled text, wrong cursor
   positions).

2. Skip stale cache write for busy sessions. When a session is actively
   working, the cache is always outdated — writing it first and then
   rewriting with the fresh API buffer caused a jarring double-render
   flash. Now busy sessions get a single clean clear+write transition.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-23 12:32:33 +01:00
arkonandClaude Opus 4.6 a9d83ec4e3 chore: version packages
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 23:55:39 +01:00
arkonandClaude Opus 4.6 84137cdba4 chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:59:27 +01:00
arkonandClaude Opus 4.6 6eb3969816 fix: avoid no-control-regex lint error for ANSI strip pattern
Use RegExp constructor with String.raw to express \x1b without
a literal control character in the source, matching the pattern
used elsewhere in the codebase.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:57:36 +01:00
arkonandClaude Opus 4.6 de49437a6f docs: add browser-testing-guide to CLAUDE.md references, clarify route count
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:38:14 +01:00
arkonandClaude Opus 4.6 6a27639083 fix: increase Ink frame search window from 4KB to 64KB to prevent partial frames
Single Ink frames with response content can be 10-20KB, so the 4KB tail
was too small and caused blank gaps. Now searches the last 64KB for VPA
row drops to find the last complete frame boundary.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:38:09 +01:00
arkonandClaude Opus 4.6 eb1b38c718 fix: align case select group height — stretch buttons to match dropdown
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:06:15 +01:00
arkonandClaude Opus 4.6 ea7b103b47 fix: prevent stale terminal data on tab switch — add chunkedTerminalWrite cancellation
chunkedTerminalWrite used requestAnimationFrame to write buffer chunks across
frames but had no cancellation. When switching tabs, old session's remaining
chunks continued writing stale data into the new session's terminal, causing
visual artifacts and garbled content.

- Add _chunkedWriteGen generation counter to abort in-flight chunked writes
- Bump gen early in selectSession() and SSE reconnect to immediately cancel
- Guard finish() so aborted writes don't flush SSE queue for wrong session
- Add fitAddon.fit() before buffer writes to sync terminal dimensions
- Add fitAddon.fit() in sendResize() to ensure local/server dim parity

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 03:00:58 +01:00
arkonandClaude Opus 4.6 bec8e2f9ee fix: improve history prompt extraction — filter expanded commands, add tail scan fallback
Skip /init expansions, slash commands, orchestrator prompts, ANSI codes, secrets,
and short/vague messages. When head scan finds no usable prompt (e.g. /init sessions),
read last 32KB of transcript to find a recent meaningful user message.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 02:42:50 +01:00
arkonandClaude Opus 4.6 2203f3a347 feat: visual redesign — glass morphism, refined colors, polished UI
Modernize the entire UI with a cohesive "refined dark glass" aesthetic
while preserving all existing functionality.

- Header/toolbar: backdrop-filter blur(16px), semi-transparent backgrounds
- Buttons: 6px radius, multi-stop gradients, inner glow, cubic-bezier transitions
- Welcome screen: gradient text title, radial bg glow, pill-shaped buttons with hover lift
- Panels/modals: glass backgrounds, 12px radius, layered shadows
- Color palette: cooler blue-tinted darks replacing flat blacks
- Forms: refined inputs with focus rings, glass toggle switches
- New CSS vars: --glass-bg, --glass-border, --btn-radius, --transition-smooth

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 02:24:27 +01:00
arkonandClaude Opus 4.6 867a10d78a refactor: optimize history endpoint — reuse buffer, extract readFileHead, use line iterator
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 02:01:56 +01:00
arkon e54d7badc4 chore: version packages 2026-03-22 01:57:49 +01:00
arkon 40dfac3534 feat: improve session history with first prompt + clickable monitor rows (closes #45) 2026-03-22 01:57:17 +01:00
arkonandClaude Opus 4.6 e899a43a18 chore: hide orchestrator button until feature is fully tested
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 01:48:11 +01:00
arkonandClaude Opus 4.6 0cab8a7ece fix: stop subagent monitor windows from auto-opening on discovery
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-22 01:44:28 +01:00
arkonandClaude Opus 4.6 7b7cf958c0 feat: add live progress during orchestrator plan generation
New SSE event orchestrator:planProgress streams phase/detail updates
from the planner to the frontend in real-time. The panel now shows
a scrollable log of planning steps instead of just "Generating plan..."

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 20:03:24 +01:00
arkonandClaude Opus 4.6 6e64ddd853 feat: add Orchestrator button to toolbar
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 19:59:06 +01:00
arkonandClaude Opus 4.6 afea91b92b fix: patch 3 production bugs found during deep audit
1. Post-phase verify timer leak — setTimeout for verifyCurrentPhase was
   never stored, so pause() couldn't cancel it. Timer now tracked in
   postPhaseTimer field and cleared in clearPhasePoll().

2. Event forwarding flag survives loop replacement — boolean
   eventForwardingAttached stayed true when a new loop was created,
   so the new loop never got SSE forwarding. Now tracks the loop
   instance reference instead of a boolean.

3. Replan stuck when no sessions — replanPhase() returned without
   setting up task handlers or polling when no idle sessions were
   available. Now starts polling so the queued task gets picked up
   when a session becomes idle.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 19:23:37 +01:00
arkonandClaude Opus 4.6 9449a8f157 test: expand OrchestratorLoop coverage to 60 tests — 17 new deep paths
New coverage:
- Task failure & retry (handleTaskFailed retry when retries < 2)
- Phase error auto-retry (handlePhaseError when attempts < maxAttempts)
- Verification with actual criteria (verifier call, pass/fail flow)
- Verification failure → replan → retry cycle
- Max verification attempts → phase failure
- Multi-phase sequential advancement
- Compact between phases (writeViaMux('/compact'))
- Crash recovery from verifying/replanning/paused states
- Single-task vs multi-task prompt generation
- Team phase sendInput error handling
- Verification session fallback (no sessions → skip)
- taskAssigned, phaseCompleted, phaseFailed event emissions

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 19:15:42 +01:00
arkonandClaude Opus 4.6 d322f17f73 test: add 43 deep integration tests for OrchestratorLoop state machine
Covers full lifecycle: start → plan → approve → execute → verify → complete.
Tests state transitions, event emissions, persistence/recovery, pause/resume,
skip/retry, team phase execution, error handling, and edge cases.

Also fixes bugs found during review:
- Route context snapshot: use getter for orchestratorLoop (was null forever)
- Event listener stacking: guard setupEventForwarding with boolean flag
- Replan completion: create tracked TaskQueue task instead of raw sendInput
- Pause cleanup: call cleanupTaskHandlers() on pause
- Phase timeout: add phaseTimeoutTimer enforcement

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 15:33:02 +01:00
arkonandClaude Opus 4.6 61b5ec095c feat: add Orchestrator Loop — phased plan execution with team agents
Adds a new autonomous loop that accepts high-level goals, generates
phased execution plans via AI, and executes them step-by-step with
verification gates between phases.

Core components:
- OrchestratorLoop: state machine (idle→planning→approval→executing→verifying→completed)
- OrchestratorPlanner: plan generation via PlanOrchestrator, Kahn's algorithm phase grouping
- OrchestratorVerifier: phase verification (strict/moderate/lenient modes)
- Prompt templates for phase execution, team delegation, verification, replanning

API (10 endpoints):
- POST start/approve/reject/pause/resume/stop
- GET status/plan
- POST phase/:id/skip, phase/:id/retry

Frontend: orchestrator-panel.js with SSE-driven state, phase progress, task tracking

Tests: 22 tests (18 route + 4 unit), all passing. Typecheck/lint/format clean.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-21 07:20:18 +01:00
arkonandClaude Opus 4.6 497ca4891a fix: restore mobile terminal scrollback — use JS scrollLines() instead of broken native scroll
xterm.js DOM renderer doesn't populate .xterm-viewport's scroll area (the div
is empty, scrollHeight === clientHeight), so native CSS scrolling via
touch-action:pan-y and overflow-y:scroll had nothing to scroll. Desktop worked
only because the wheel handler called terminal.scrollLines() directly.

- Replace split mobile/desktop touch handlers with unified JS-driven handler
  that converts touch deltas to terminal.scrollLines() calls (with pixel
  accumulation for slow swipes and momentum scrolling)
- Change touch-action from pan-y to none on terminal elements so browser
  doesn't fight the JS handler
- Remove now-unnecessary xterm-viewport position/overflow/z-index overrides
  and iOS -webkit-overflow-scrolling rules
- Fix _shrinkPaddingToFit() arithmetic (was adding gap instead of subtracting)
- Minor: add route-helpers.ts to CLAUDE.md, fix sse-events.ts comment count

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-20 09:38:18 +01:00
arkon 580b7a3f90 chore: version packages 2026-03-19 12:35:39 +01:00
arkonandClaude Opus 4.6 34c3d8f5ff fix: tighten mobile keyboard layout — eliminate dead space and toolbar overlap
- Remove redundant 50px CSS padding on terminal-container when keyboard visible
- Reduce JS paddingBottom constant from +94 to +84 (exact toolbar + accessory)
- Add _shrinkPaddingToFit() to eliminate terminal row quantization gap
- Add CSS padding-bottom on .main for fixed toolbar clearance (keyboard hidden)
- Match iOS Safari toolbar offset (100vh - --app-height) in .main padding

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-19 12:35:08 +01:00
arkonandClaude Opus 4.6 2491471ba5 fix: prevent mobile page scroll when typing with keyboard open
iOS Safari scrolls the document to bring xterm's hidden textarea into
view when the user types, pushing the entire UI off-screen. Fix with:
- CSS position:fixed on .app when keyboard is visible
- window.scroll listener to reset scroll position as safety net
- scroll reset in onKeyboardShow before and after fit/resize

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-19 11:47:18 +01:00
arkonandClaude Opus 4.6 0c4aac8029 chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-19 10:05:16 +01:00
arkonandClaude Opus 4.6 d436c6375f fix: strip Ink spinner bloat from terminal buffer before tailing
During long thinking phases, Ink's TUI rewrites the spinner/status bar
thousands of times via absolute cursor positioning (VPA/CUP). These
500KB+ of redraw frames pushed real content out of the 128KB tail
window, making the terminal appear empty when switching tabs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-16 10:55:24 +01:00
arkonandClaude Opus 4.6 e96baf9f66 chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-16 01:42:16 +01:00
arkon 7b8aa529f2 chore: version packages 2026-03-15 03:52:30 +01:00
arkonandClaude Opus 4.6 0ad4e0ea24 fix: correct resolveCasePath priority order and suppress JSON parse warnings
- resolveCasePath now checks linked cases first (matching original behavior
  of /api/cases/:name and /api/cases/:name/fix-plan handlers)
- readLinkedCases only warns on real I/O errors, not JSON parse errors
  (SyntaxError has no .code property, so check for .code existence first)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 03:50:45 +01:00
arkonandClaude Opus 4.6 6bc403d88d refactor: clean up case routes DRY violations, remove dead export, standardize reply API
- Extract readLinkedCases() helper and resolveCasePath() to eliminate 6x duplicated
  linked-cases.json path construction and 5x duplicated file read/parse logic
- Replace O(n) .some() duplicate check with O(1) Set.has() in case listing
- Un-export isError() in types/api.ts (only used internally by getErrorMessage)
- Standardize reply.status() → reply.code() in system-routes (Fastify canonical API)
- Update CLAUDE.md: accurate frontend module listing, SSE event count (~106)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-15 03:47:18 +01:00
arkon 1ad05a5a42 chore: version packages 2026-03-14 22:22:52 +01:00
arkonandClaude Opus 4.6 192690911f refactor: extract app.js into 6 domain modules with deferred init
Split the monolithic app.js (~12.5K lines) into 6 focused mixin modules
that extend CodemanApp.prototype via Object.assign:

- terminal-ui.js — terminal setup, rendering pipeline, controls
- respawn-ui.js — respawn banner, countdown, presets, run summary
- ralph-panel.js — Ralph state panel, fix_plan, plan versioning
- settings-ui.js — app settings, visibility, web push, tunnel/QR, help
- panels-ui.js — subagent panel, teams, insights, file browser, log viewer
- session-ui.js — quick start, session options, case settings

Fix deferred script init ordering: wrap CodemanApp instantiation in
DOMContentLoaded so all defer'd mixin modules execute their
Object.assign before the constructor runs. Without this, init() calls
methods like applyHeaderVisibilitySettings() that don't exist yet.

Guard missing cleanupWizardDragging() call in subagent-windows.js.
Update build.mjs to minify/hash all new modules. Update CLAUDE.md
with new frontend architecture and load order.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 21:24:22 +01:00
arkon 551461cb31 chore: version packages 2026-03-14 19:26:09 +01:00
arkonandClaude Opus 4.6 c4bae75c59 fix: add onerror handler for lazy-loaded WebGL addon script
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 19:23:15 +01:00
arkonandClaude Opus 4.6 88c415fc37 perf: V8 compile cache, lazy-load WebGL, preload hints, batch tmux reconciliation
- Enable NODE_COMPILE_CACHE in systemd service and npm start for 10-20% faster cold starts
- Lazy-load xterm-addon-webgl.min.js (244KB) only on desktop — mobile never downloads it
- Add <link rel="preload"> hints for critical scripts (xterm, constants, app) in <head>
- Replace per-session tmux subprocess calls with single batch `list-panes -a` call
  (N*2+1+M execSync calls → 1 for reconcileSessions)
- Fix CLAUDE.md frontend module count (10 → 11)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 19:20:44 +01:00
arkonandClaude Opus 4.6 7175e4b350 docs: update CLAUDE.md and README.md to reflect current codebase
Correct stale counts and add missing entries: route modules 12→13
(ws-routes.ts), frontend modules 9→10 (input-cjk.js), handler count
~111→~114, utilities section expanded, TypeScript badge 5.5→5.9,
frontend extracted modules 8→9.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:48:25 +01:00
arkonandClaude Opus 4.6 08a417997f ci: upgrade actions/checkout and actions/setup-node to v6 (Node 24)
Replace v4 (Node 20) with v6 (Node 24 native) to eliminate the
deprecation warning. Remove the FORCE_JAVASCRIPT_ACTIONS_TO_NODE24
workaround since v6 doesn't need it.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:42:22 +01:00
arkonandClaude Opus 4.6 e6cb89b0cd ci: use Node.js 24 runtime for actions and bump release node to 22
Opt into Node.js 24 for GitHub Actions runners (actions/checkout@v4,
actions/setup-node@v4) to silence deprecation warnings. Also bump
release.yml from node 20 to 22 to match ci.yml.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:40:44 +01:00
arkon d072e773d8 chore: version packages 2026-03-14 18:38:28 +01:00
arkonandClaude Opus 4.6 a649c91b68 fix: WS session lifecycle, reconnection, and CJK session-switch cleanup
- Close WebSocket when session exits (exit event listener) to prevent
  orphaned listeners and stale writes to dead PTY
- Add readyState guard in onTerminal to stop buffering after socket closes
- Simplify heartbeat: remove redundant alive flag, use pongTimeout only
- Add exponential backoff reconnection on unexpected WS close (skip for
  server rejections 4004/4008/4009)
- Clear CJK textarea on session switch to prevent wrong-session input

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:37:10 +01:00
arkonandClaude Opus 4.6 3383c23099 fix: address code review findings across WS, CJK input, install.sh, and README
WebSocket route: add socket error handler to prevent process crashes, enforce
per-session connection limit (max 5), track/decrement counts on close.

CJK input: add destroy() method with proper listener cleanup, guard against
double-init, add maxlength/aria-label to textarea, use language-neutral
placeholder, explicitly clear cjkActive on hide.

install.sh: fix update() to use $BRANCH and $REPO_URL instead of hardcoded
origin/master — fork users were silently switched back to master on update.

README: fix broken markdown table (paragraph concatenated into last cell),
add CODEMAN_NODE_VERSION to env var table.

Tests: add 8 new test cases for batch coalescing, flush threshold, unknown
message types, connection limit, heartbeat, readyState guards. Import
MAX_INPUT_LENGTH from config, add connectWs timeout, replace setTimeout
with vi.waitFor in cleanup test.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:27:58 +01:00
arkonandClaude Opus 4.6 c3e1e731ef fix: use generic placeholders in fork install README example
Replace hardcoded contributor fork URL with <user>/<branch> placeholders
so the documentation is useful for any contributor.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:11:43 +01:00
Ark0N 405b711c3a Merge pull request #41 from douchekr/feat/input-cjk-form
feat: add CJK IME input textarea and fork/branch install support
2026-03-14 18:10:09 +01:00
arkon 4295faefc9 chore: version packages 2026-03-14 18:04:51 +01:00
arkon 93e1ba5110 Merge remote-tracking branch 'origin/feat/ws-terminal-io-upstream' 2026-03-14 18:03:36 +01:00
Ark0N abbbf9e90a Merge pull request #43 from Ark0N/feat/ws-tests
test: add WebSocket terminal I/O route tests
2026-03-14 18:03:04 +01:00
Ark0N 3a41de7b57 Merge pull request #42 from Ark0N/feat/ws-heartbeat
feat: add ping/pong heartbeat to WebSocket connections
2026-03-14 18:03:02 +01:00
Ark0N 8267edc6fe Merge pull request #40 from Spirotot/feat/ws-terminal-io-upstream
feat: WebSocket terminal I/O with server-side DEC 2026 sync
2026-03-14 18:02:55 +01:00
arkonandClaude Opus 4.6 78c568e5f7 test: add automated tests for WebSocket terminal I/O route
16 tests covering session-not-found close code, terminal output with
DEC 2026 sync markers, client input forwarding, resize bounds
validation, malformed message handling, and connection cleanup of
session event listeners.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:00:55 +01:00
arkonandClaude Opus 4.6 cc624d2575 feat: add ping/pong heartbeat to WebSocket connections
Detect stale connections that TCP keepalive won't catch for minutes,
especially through tunnels and proxies. Pings every 30s with a 10s
pong timeout — if the client doesn't respond, the socket is terminated
and all timers cleaned up.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 17:58:50 +01:00
arkonandClaude Opus 4.6 5844720525 fix: validate WS resize dimensions to match HTTP route bounds
The HTTP resize route validates via ResizeSchema (cols: 1-500, rows:
1-200, integers only). The WS handler only checked typeof === 'number',
allowing floats, negatives, and extreme values through to ptyProcess.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 17:57:20 +01:00
jayparkandClaude Opus 4.6 393a2d9c28 fix: use BRANCH variable in install.sh no-changes update path
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:52:28 +09:00
jayparkandClaude Opus 4.6 809bf6a614 fix: use BRANCH variable in install.sh update path
The update path was hardcoded to origin/master. Now uses the
CODEMAN_BRANCH variable and updates the remote URL on upgrade.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:51:20 +09:00
jayparkandClaude Opus 4.6 da71d8d01c feat: support custom repo URL and branch in install.sh
Add CODEMAN_REPO_URL and CODEMAN_BRANCH env vars to install.sh
for installing from forks or feature branches. Update README with
fork installation instructions and env var reference table.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:50:06 +09:00
jayparkandClaude Opus 4.6 e5aca6aa4c feat: add CJK IME input textarea with env toggle
Add a dedicated textarea below the terminal for CJK (Korean/Japanese/Chinese)
IME input. xterm.js intercepts IME composition events, preventing composed
characters from displaying correctly. This textarea bypasses xterm entirely
by using native browser IME handling — text accumulates until Enter, then
sends to PTY in one shot.

- Always-visible textarea below terminal (inside .terminal-wrap flex column)
- focus/blur sets window.cjkActive flag to block xterm onData
- Enter sends textarea.value + \r to PTY, Escape clears
- Arrow keys, Ctrl+C/D/L/Z, Tab, Backspace pass through to PTY when empty
- attachCustomKeyEventHandler suppresses xterm key handling during composition
- INPUT_CJK_FORM=ON|OFF env var toggle (default: off, passed via SSE init)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:23:25 +09:00
Aaron FieldsandClaude Opus 4.6 ceaf4624a1 feat: add WebSocket terminal I/O with server-side DEC 2026 sync
Replace per-keystroke HTTP POST + SSE terminal output with a single
bidirectional WebSocket connection for dramatically lower input latency.
The existing SSE+POST paths remain fully functional as fallback.

Server-side: ws-routes.ts provides /ws/sessions/:id/terminal with 8ms
micro-batching and 16KB flush threshold. Each batch is wrapped in
DEC 2026 synchronized update markers so xterm.js renders atomically —
Ink's DA capability negotiation fails through the PTY→server→WS proxy
chain, so without server-injected markers, cursor-up redraws flicker.

Frontend: _connectWs/_disconnectWs manage per-session WS lifecycle.
Input and resize use WS fast path with HTTP POST fallback. SSE terminal
events are suppressed when WS is active to prevent double rendering.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 21:02:28 -04:00
arkonandClaude Opus 4.6 a6597e4a9a fix: patch 5 dependency vulnerabilities (basic-ftp, fastify, minimatch, serialize-javascript)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 00:26:56 +01:00
arkon f869e823af chore: version packages 2026-03-12 23:59:16 +01:00
arkonandClaude Opus 4.6 8d0b179f94 fix: repair 15 pre-existing subagent-watcher test failures
Root causes:
- Mock readline (EventEmitter) lacked .close() method, causing TypeError
  that blocked extractDescriptionFromFile's Promise from ever resolving
- Mock stream lacked .destroy() method (same issue after .close() fix)
- Entry-processing tests shared one readline mock between description
  extraction and tailing — events emitted before tailFile started were lost
- Liveness checker marked agents as 'completed' instead of 'idle' because
  fixed stat timestamps became stale after fake timer advancement

Fixes:
- Add createMockRl() helper with .close() method
- Use { destroy: vi.fn() } for stream mocks
- Use mockReturnValueOnce() for two-readline pattern in 7 entry tests
- Use mockImplementation() for dynamic stat timestamps

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 16:08:39 +01:00
arkonandClaude Opus 4.6 98fa55b7b2 chore: codebase cleanup — remove dead code, consolidate imports, extract constants
- Remove 3 unused exported constants (TRIM_MESSAGES_TO, MAX_TERMINAL_COLS, MAX_TERMINAL_ROWS)
- Consolidate 8 direct util imports into barrel imports (./utils/index.js)
- Extract magic number 8191 to FILE_PEEK_BYTES constant in buffer-limits.ts
- Add explanatory comments to 9 undocumented .catch(() => {}) handlers

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:50:40 +01:00
arkonandClaude Opus 4.6 c46ac30631 fix: hide subagent monitor panel by default
Change showSubagents default from true to false so the subagent
panel doesn't auto-show on page load. Users can still enable it
via Settings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:35:27 +01:00
arkonandClaude Opus 4.6 dfcc14bfd2 fix: one-liner restart command that works for background processes
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:34:42 +01:00
arkonandClaude Opus 4.6 a068008409 fix: clarify restart instructions — stop first, then start
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:33:34 +01:00
arkonandClaude Opus 4.6 0aa31f100e fix: show restart command when codeman-web is not a systemd service
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:31:19 +01:00
arkonandClaude Opus 4.6 314a160458 feat: auto-restart codeman-web service after update if running
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:23:55 +01:00
arkonandClaude Opus 4.6 e7ee5595c5 feat: auto-detect existing install and run update instead of fresh install
Re-running the install script now detects ~/.codeman/app/.git and
automatically updates instead of re-installing. Removes the separate
`bash -s update` instructions from README since it's no longer needed.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:09:37 +01:00
251 changed files with 27225 additions and 20557 deletions
+5 -2
View File
@@ -11,10 +11,10 @@ jobs:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/checkout@v6
- name: Setup Node.js
uses: actions/setup-node@v4
uses: actions/setup-node@v6
with:
node-version: 22
cache: 'npm'
@@ -22,6 +22,9 @@ jobs:
- name: Install dependencies
run: npm ci
- name: Check package-lock.json version sync
run: npm run check:lockfile
- name: Type check
run: npm run typecheck
+3 -3
View File
@@ -16,12 +16,12 @@ jobs:
pull-requests: write
steps:
- name: Checkout repo
uses: actions/checkout@v4
uses: actions/checkout@v6
- name: Setup Node.js
uses: actions/setup-node@v4
uses: actions/setup-node@v6
with:
node-version: 20
node-version: 22
cache: npm
registry-url: https://registry.npmjs.org
+6 -1
View File
@@ -48,7 +48,12 @@ Thumbs.db
# Generated output
out/
screenshots-echo-diag/
tools/remotion/out/
scripts/remotion/out/
# Artifacts that should not be tracked
test-results/
tmp/
public
# Claude Code plan tracking
plan.json
+1 -1
View File
@@ -6,4 +6,4 @@ src/web/public/app.js
src/web/public/styles.css
src/web/public/mobile.css
src/web/public/index.html
tools/
scripts/remotion/
+349
View File
@@ -1,5 +1,354 @@
# aicodeman
## 0.6.8
### Patch Changes
- Finish the hostname-aware notification plumbing started in 0.6.7 and lock down the recent UI/runtime fixes with regression tests.
- Browser Notification API (OS-level desktop pop-ups, layer 3 of the 5-layer notification system) now uses `${originalTitle}: ${title}` instead of the hardcoded `Codeman:` literal — so multi-host users running Codeman on laptop / dev box / NAS see `codeman:<host>: <event>` consistently across tab title, tab-flash, Web Push, and OS notifications.
- Inline session rename hardened against three corner cases: IME composition commits (Chinese pinyin Enter no longer ships half-composed text as the session name), mid-rename SSE deletion (orphaned `<input>` no longer 404s on blur), and double-fire on stuck settle-once flag (closure-local `settled` boolean replaces the boolean instance flag).
- Test coverage backfilled for two prior shipped fixes:
- `<title>codeman:<host></title>` server-side templating (#82): 8 tests covering default `os.hostname()`, `--title-hostname` override, HTML-escape against `<script>`-style breakout, ampersand non-double-encoding, and template-tail byte-identical invariance.
- tmux size-query helper (#80): 15 tests covering the browser-resize-between-attaches happy path, the query-then-die race, zero/negative/empty/non-numeric output fallbacks, and argv-form/timeout assertions that lock down the no-shell-interpolation guarantee. Inline 14-line query block extracted into a named `queryTmuxWindowSize()` export in `session.ts` so the test surface is a pure function.
- Regression coverage added for `stripInkRedrawBloat` route helper.
- CLAUDE.md and README.md updated to document dual-CLI env-prefix discipline (`CLAUDE_CODE_*` vs `OPENCODE_*`), the `xterm-zerolag-input` published-package side-effect of overlay edits, and the unified hostname prefix across tab title / tab-flash / OS notifications.
## 0.6.7
### Patch Changes
- - **fix(client): preserve inline rename input across tab re-renders** (#81) — Right-click → rename on a session tab no longer loses keystrokes when SSE traffic from sibling sessions triggers a tab re-render. Adds an `_inlineRenameActive` guard at the top of `renderSessionTabs()` and `_fullRenameSessionTabs()` so the in-progress input isn't destroyed mid-typing. Also fixes a latent double-fire of `finishRename` (blur + Enter could both invoke it). Drive-by: safer DOM child clearing in place of `innerHTML = ''`.
- **feat: hostname-aware window title** (#82) — The browser tab title is now `codeman:<hostname>` instead of the bare `Codeman` literal, so users running Codeman on multiple hosts (laptop, dev box, NAS) can tell at a glance which tab points at which backend. New `--title-hostname <name>` CLI flag overrides the detected `os.hostname()` when it's noisy or you want a cosmetic name. The title is templated into the served HTML on first byte (with narrow HTML escaping), so it's correct from the first paint and works without JavaScript. Title-flash logic now respects the per-host title.
- **perf: larger terminal tail on tab switch** — `TERMINAL_TAIL_SIZE` raised from 128KB to 1MB. When switching back to a busy session tab you now get ~8× more scrollback restored immediately.
- **fix: preserve response text in Ink redraw stripping** — `stripInkRedrawBloat()` rewritten from a first-VPA approach to cluster-based detection. The previous algorithm assumed all VPA escapes after the first one belonged to a single redraw region and discarded everything in between, which silently lost 100KB+ of legitimate Claude response text once a render had occurred. The new approach groups VPAs into clusters separated by ≥8KB gaps and only collapses clusters spanning ≥32KB, so streamed response content between redraw bursts is preserved.
- **docs**: `CLAUDE.md` Additional Commands gains the `--title-hostname` row; `README.md` gets a "Hostname-Aware Window Title" subsection under Multi-Session Dashboard.
## 0.6.6
### Patch Changes
- **Terminal scrollback significantly increased** — both the xterm.js viewport and the tmux backing buffer were bottlenecking how far back you could scroll. Three changes:
- `DEFAULT_SCROLLBACK` raised from 20000 → 50000 lines (xterm.js, main terminal). The previous bump from 5000 only helped users with empty localStorage; existing users were stuck on whatever value they first picked up. The loader now treats `DEFAULT_SCROLLBACK` as a floor — if your stored value is below the new minimum, you're raised to it automatically.
- Subagent / teammate terminals (`panels-ui.js`) were stuck at 5000; now use the same `DEFAULT_SCROLLBACK` constant (50000).
- New tmux sessions now run with `history-limit 50000` (tmux defaults to 2000). This matters for hard-reload / re-attach — without it, only the last ~2000 lines survive the round-trip back into a fresh xterm.
**Tmux flicker on session re-attach fixed (PR #80 by @aakhter)**: the PTY now queries the existing tmux window size via `tmux display -p` before spawning, instead of hardcoding 120x40. Previously, every re-attach forced tmux to resize down to 120x40, causing a visible flicker and one frame of scrollback loss. The `-x 120 -y 40` flag was also dropped from `tmux new-session` so the initial size matches the first attaching client. Uses `execFileSync` (not shell) for safety and falls back to 120x40 on any error.
**Docs**: CLAUDE.md now documents two recurring foot-guns — the `xterm-zerolag-input` overlay code is duplicated between `packages/xterm-zerolag-input/src/` and inline inside `src/web/public/app.js`, so any overlay change must touch both; and the COM workflow explicitly includes a post-push `gh run watch` step to confirm CI before considering the release done.
## 0.6.5
### Patch Changes
- **Mobile fix**
- Android virtual keyboard: space character was silently dropped on touch devices using GBoard / SwiftKey / similar IMEs. Root cause: the input-event handler in `terminal-ui.js` treated any whitespace-only textarea value as proof that xterm had already processed the input. A lone space (`' '.trim() === ''`) tripped this guard, so the space was consumed but never forwarded. Now skips only when the textarea is truly empty (or whitespace from a non-space key). Reported and diagnosed by @coolk8 in #79.
**Docs**
- `CLAUDE.md`: added Zod `.optional()`-vs-`null` gotcha (recurring trap from 0.6.3 / 0.6.4 incidents) and a more visible warning against running bare `npm test` (kills the host tmux session).
- `docs/local-echo-overlay-plan.md`: marked SHIPPED, corrected xterm version reference (v5.3.0 → `@xterm/xterm` ^6.0.0).
## 0.6.4
### Patch Changes
- Fix "Failed to enable respawn: Invalid request body" error when selecting infinity duration (∞) in the respawn modal. Frontend was sending `durationMinutes: null`, which Zod's `.optional()` schema rejected (it accepts `undefined` only). The body now omits the field when no duration is selected.
## 0.6.3
### Patch Changes
- **Fix**
- Allowlist `opusContext1mEnabled` in `SettingsUpdateSchema`. Without this entry, the strict schema rejected `PUT /api/settings {"opusContext1mEnabled":...}` with `INVALID_INPUT`, so the toggle's value never persisted across reloads. The frontend was already reading and writing this key (`settings-ui.js:336/1137`, `session-ui.js:340`), so saves were silently failing — users never noticed because the load path falls back to `false` on missing keys, hiding the bug. (#78)
## 0.6.2
### Patch Changes
- **Mobile UX**
- Resume Conversation list (welcome page) reworked for narrow screens: 2-line title clamp so more of the first prompt is visible; case-aware subtitle that renders `#caseName` (or `#caseName/sub`) when `workingDir` matches a known case, otherwise falls back to the directory basename; inline `⋯` toggle that expands a detail panel with full prompt, full path, timestamp, size, and short session id; `/Users/<user>/` now collapses to `~/` alongside `/home/<user>/`. (#77)
- Response viewer: ASCII diagram wrap toggle, dedicated mobile code-block layout, and chrome-stripping fallback when the model wraps its reply in extra markup. (#75)
- Mobile keyboard accessory bar no longer triggers vertical scroll. (#72)
**Sessions & settings**
- New `thinkingEffort` setting on session creation, with `xhigh` option and `/effort max` mobile shortcut. (#73)
- `thinkingEffort` is now allowlisted in `SettingsUpdateSchema` so it round-trips through PATCH /api/settings.
- `envOverrides` (`CLAUDE_CODE_*` / `OPENCODE_*`) are now passed to Claude via tmux env exports at spawn time instead of being written to `<case>/.claude/settings.local.json`. Eliminates UI/disk drift; the value lives on `Session._envOverrides`, is exported by `tmux-manager.buildEnvExports()`, and is persisted in `SessionState.envOverrides`. (#74)
**Fixes**
- Eye icon (active-session indicator) now follows `/clear` to the new Claude conversation instead of getting stuck on the previous transcript. (#76)
- `tmux-manager.reconcileSessions` now uses `|` as the field separator, fixing parsing when session names contain other delimiters. (#71)
**Docs**
- CLAUDE.md: added `npm run knip` to the dead-code sweep table and a `Common Gotchas` entry documenting the `envOverrides` → tmux export flow.
## 0.6.1
### Patch Changes
- Internal cleanup and release hygiene:
- **Dead-code sweep via knip**: added `knip.json` for dead-code detection and ran a full sweep — removed unused test files, unused scripts, and narrowed internal module exports to the minimum surface area actually consumed.
- **Lockfile drift prevention**: `version-packages` now runs `npm install --package-lock-only` and verifies the lockfile is in sync via `scripts/check-lockfile-sync.mjs`; CI runs the same check on every push/PR so version drift fails the build instead of reaching production. Resolves the `package-lock.json` / `package.json` version mismatch that shipped in 0.6.0.
- **Docs tightening**: archived 22 completed plan docs from `docs/`, corrected file/handler counts in `CLAUDE.md`, documented the lockfile step in the COM workflow, and removed footer redundancy.
## 0.6.0
### Minor Changes
- Community contributions from @aakhter:
- **feat (#66): Tab reorder shortcuts** — `Ctrl+Shift+{` and `Ctrl+Shift+}` move the active session tab left/right, matching WezTerm convention. Order persists across reloads via `saveSessionOrder()`.
- **feat (#67): Active tab visibility + Alt+N badges** — active tab now has a bright green border with color-matched glow, and the first 9 tabs display number badges hinting at the `Alt+N` switch shortcut. Badges update on reorder/rerender.
- **feat (#68): Clipboard API** — new `POST /api/clipboard` accepting `{text}` broadcasts a `clipboard:write` SSE event; connected browsers attempt `navigator.clipboard.writeText()` with a manual-copy modal fallback when the page isn't focused. Auth-protected via the standard middleware. Useful for pushing snippets from remote sessions to the user's local clipboard.
- **fix (#65): Android Shift+key double character** — pressing `Shift+A` on attached Android keyboards no longer produces "AA". Tracks xterm-handled keydown timestamps and skips the orphaned-input listener for 50ms after a real keydown, while still catching Gboard symbol-keyboard inputs (keyCode 229).
## 0.5.13
### Patch Changes
- Fix "Case path not found" error in Quick Start when `~/codeman-cases/` does not exist (issue #64). Two bugs in `session-ui.js`:
- `runClaude()` auto-create read `createCaseData.case`, but `POST /api/cases` returns `{ success, data: { case } }` — corrected to `createCaseData.data.case`.
- `runShell()` had no auto-create logic and would immediately throw on a missing case directory — now mirrors `runClaude()`'s create-on-demand flow.
## 0.5.12
### Patch Changes
- Fix quick-start to resolve linked cases before codeman-cases fallback. `/api/quick-start` was always resolving `caseName` against `CASES_DIR`, ignoring entries in `~/.codeman/linked-cases.json`. Sessions started via quick-start now correctly honour linked external project directories, consistent with regular case routes.
## 0.5.11
### Patch Changes
- Community contributions and security hardening:
- Mobile response viewer: native-scroll panel for reading full Claude responses with markdown rendering via marked.js (PR #62)
- PWA support: service worker caching, web app manifest, and Android home screen install (PR #59)
- Named Cloudflare tunnel support (PR #58)
- Markdown rendering for response viewer with HTML sanitization (XSS prevention) — strips dangerous elements, event handlers, and javascript: URIs
- Service worker switched from stale-while-revalidate to network-first caching so deploys take effect immediately
- Content-Disposition filename sanitization to prevent header injection in file downloads
- Expose session.muxName public getter, replace unsafe `as any` cast in session-routes
- Static import for execFile in session-routes
- Keyboard shortcut updates: Alt+1-9 tab switching, Shift+Enter newline
- Repo restructure for cleaner GitHub landing page
- Mobile logo, expandable history, session resume fixes
## 0.5.10
### Patch Changes
- fix: allow bracket characters in model validation regex so models like opus[1m] (1M context window) are accepted instead of silently dropped. Quote the model flag value in tmux spawn commands to prevent bash glob expansion of bracket patterns.
docs: update macOS launchd instructions to use `launchctl bootstrap` instead of deprecated `load`. Clean up README install and service sections.
## 0.5.9
### Patch Changes
- Mobile keyboard accessory bar: add configurable "Extended Keyboard Bar" setting (Settings > Display > Input) that toggles between simple mode (up/down arrows, /init, /clear, /compact, paste, dismiss) and extended mode (adds left/right arrows, Tab, Shift+Tab, Ctrl+O, Alt+Enter, Esc). Default is simple mode. Setting is device-specific (not synced to server).
Restyle dismiss button: muted steel-blue tone, fills remaining bar space via flex, larger tap target. Arrow buttons now blue.
Fix paste overlay visibility on mobile: dialog repositioned to top of screen (15vh from top) so the virtual keyboard doesn't cover it. Textarea enlarged for better usability.
(Also includes all v0.5.8 changes: case reorder/delete, XSS sanitization, auto-attach PTY on restart, mobile keyboard buttons, macOS installer fixes, terminal flicker fix, state store collision fix.)
## 0.5.8
### Patch Changes
- Case management: add Manage tab with reorder (up/down arrows) and delete for cases; linked cases are unlinked (folder preserved), CASES_DIR cases are permanently deleted. New endpoints: DELETE /api/cases/:name, PUT /api/cases/order. SSE events: case:deleted, case:order-changed.
Security: sanitize case names from filesystem with /^[a-zA-Z0-9_-]+$/ regex before returning from GET /api/cases to prevent XSS via maliciously-named directories reaching frontend inline onclick handlers.
Auto-attach PTY: server now calls startInteractive() for recovered tmux sessions during startup so all sessions resume capturing output immediately after deploy, instead of waiting for client selection. Frontend auto-attach condition relaxed from (pid===null && status==='idle') to (pid===null && !\_ended).
Mobile keyboard accessory: add Shift+Tab, Tab, Esc, Alt+Enter, Left/Right arrow, and Ctrl+O buttons.
Terminal: fix flicker regression by moving viewport clear inside dimension guard.
State store: fix temp file collisions on concurrent writes.
macOS: fix installer failures when piped via curl | bash, add HTML cache support, launchd service template, and trust dialog handling.
Housekeeping: remove accidentally committed dist/state-store.js build artifact.
## 0.5.7
### Patch Changes
- feat: support "Default (CLI default)" option for model selection. Adds a new empty-value option to the model dropdown that defers to the CLI's own default model instead of forcing a specific model. Ensures empty defaultModel values are treated as undefined when passed to session creation and Ralph loop start, preventing empty strings from being sent as model flags.
## 0.5.6
### Patch Changes
- fix: default new sessions to opus[1m] (1M context window) instead of plain opus (200k context)
## 0.5.5
### Patch Changes
- Add 1M Opus context quick setting — per-case and global toggle that writes `model: "opus[1m]"` to `.claude/settings.local.json` when creating new sessions. Fix mobile layout: banners (respawn, timer, orchestrator) between header and main content now visible by switching from margin-top on `.main` to padding-top on `.app`. Add tablet-optimized respawn banner styles and mobile phone banner refinements.
## 0.5.4
### Patch Changes
- Fix terminal flicker regression — re-add server-side DEC 2026 synchronized output wrapping around batched terminal data. Ink spinner frames (cursor-up + redraw cycles) do not emit their own DEC 2026 markers, so without the server wrapper each partial cursor update rendered individually causing visible flicker. Also: extract SSE stream management, session listener wiring, and respawn event wiring from server.ts into dedicated modules; deduplicate error message extraction across 7 files with shared getErrorMessage() helper; update SSE event count in CLAUDE.md (106 → 117).
## 0.5.3
### Patch Changes
- Readability refactor across 12 core files, extracting ~35 helper methods to reduce duplication:
- state-store: extract serializeState(), split assembleStateJson() into focused sub-methods
- session: extract \_resetBuffers() (3x dedup), \_clearAllTimers() (10 timer cleanups), \_handleJsonMessage()
- ralph-tracker: extract completeAllTodos() (4x dedup), emitValidationWarning(), named similarity constants
- subagent-watcher: extract markSubagentAsCompleted(), extractFirstTextContent(), emitToolResult(), findOldestInactiveAgent()
- respawn-controller: extract recoveryResetToWatching(), canAutoAccept(), formatRemainingSeconds(), validatePositiveTimeout()
- tmux-manager: replace 15 path.includes() with UNSAFE_PATH_CHARS regex, extract buildEnvExports/buildPathExport/\_configureOpenCode helpers
- session-auto-ops: extract executeWhenIdle() shared retry helper, convert to options object, add validateThreshold()
- app.js: add \_clearTimer() (11 call sites), \_isStaleSelect(), keyboard shortcut lookup table, \_cleanupPreviousSession(), \_resetAllAppState()
- route-helpers: add readJsonConfig() (5 inline patterns replaced), validateSessionFilePath() (2 duplicated blocks replaced)
## 0.5.2
### Patch Changes
- Make buffer size limits configurable via CODEMAN\_\* environment variables (MAX_TERMINAL_BUFFER, TRIM_TERMINAL_TO, MAX_TEXT_OUTPUT, TRIM_TEXT_TO, MAX_MESSAGES), falling back to existing defaults. Allows users with fewer sessions or more RAM to tune buffer sizes without patching source.
Fix duplicate terminal output on tab switch to busy sessions by clearing the terminal before writing the new buffer.
Fix stale Ink CUP frames after tab switch by sending Ctrl+L to force a clean redraw.
Fix mobile CJK input handling: resolve textarea positioning, terminal flicker during composition, and layout overflow on small screens. Improve CJK composition lifecycle with better event handling and fallback flush timers.
## 0.5.1
### Patch Changes
- refactor: codebase cleanup — extract route helpers, eliminate boilerplate, optimize hot paths
- Add `parseBody()` helper to route-helpers.ts: validates request body against Zod schema with structured 400 error on failure, replacing 37 identical safeParse + error-check blocks across 10 route files
- Add `persistAndBroadcastSession()` helper: combines persist + SessionUpdated broadcast into one call, replacing 5 repeated 2-line pairs
- Migrate session-routes.ts to use `findSessionOrFail()` consistently (17 inline session lookups replaced) and `parseBody()` (12 patterns)
- Migrate ralph-routes.ts to use `findSessionOrFail()` (9 lookups) and `parseBody()` (4 patterns)
- Migrate 8 remaining route files to use `parseBody()` (21 patterns total)
- Fix O(n log n) eviction in bash-tool-parser.ts: replace `Array.from().sort()[0]` with O(n) min-scan for oldest active tool
- Extract `_debouncedCall()` utility in frontend: replaces 4 manual debounce patterns (7 lines each → 1 line) in app.js, panels-ui.js, ralph-panel.js
- Net reduction: 208 lines removed across 16 files
## 0.5.0
### Minor Changes
- Visual redesign with glass morphism, refined colors, and polished UI. Optimize history endpoint with buffer reuse and line iterator. Fix Ink frame search window (4KB→64KB) to prevent partial frames. Fix stale terminal data on tab switch via chunkedTerminalWrite cancellation. Improve history prompt extraction with expanded command filtering and tail scan fallback. Align case select group height to match dropdown. Fix no-control-regex lint error for ANSI strip pattern. Add browser-testing-guide to CLAUDE.md references.
## 0.4.7
### Patch Changes
- feat: improve session navigability in history and monitor panel (closes #45)
- History items now show the first user prompt as the title with the project path as a subtitle, making it much easier to distinguish sessions from the same project
- The `/api/history/sessions` endpoint extracts the first user message from each transcript JSONL, stripping system-injected XML tags and command artifacts, truncating to 120 chars
- Monitor panel session rows are now clickable — clicking navigates directly to that session's tab via `selectSession()`; Kill button retains independent behavior via `stopPropagation()`
- Updated CLAUDE.md architecture tables to reflect Orchestrator Loop additions (14 route modules, 15 type files, orchestrator domain files, orchestrator-panel.js frontend module)
- fix: stop subagent monitor windows from auto-opening on discovery
- feat: add Orchestrator Loop with phased plan execution, live progress during plan generation, and toolbar button (hidden until fully tested)
- fix: patch 3 production bugs found during deep audit
- fix: restore mobile terminal scrollback using JS scrollLines() instead of broken native scroll
## 0.4.6
### Patch Changes
- Fix mobile keyboard scroll and layout issues:
- Prevent iOS Safari from scrolling the page when typing with the keyboard open (position:fixed on .app + window.scroll reset)
- Eliminate dead space between terminal and keyboard accessory bar by removing redundant CSS padding, tightening JS padding constant, and adding row quantization gap compensation
- Fix toolbar overlapping terminal content when keyboard is hidden by adding proper padding-bottom to .main, including iOS Safari bottom bar offset
- Strip Ink spinner bloat from terminal buffer before tailing
- Fix resolveCasePath priority order and suppress JSON parse warnings
## 0.4.5
### Patch Changes
- Fix mobile keyboard toolbar positioning on iOS Safari: toolbar (Run/Stop/Run Shell) was hidden behind the accessory bar when virtual keyboard was active due to overlapping CSS positions. Remove the aggressive safety check in `updateLayoutForKeyboard()` that incorrectly dismissed keyboard state when iOS scrolled the visual viewport during typing. Add Safari-bar CSS offset to accessory bar so it properly stacks above the toolbar. Remove the double-counted Safari-bar offset when keyboard is visible since the JS transform already covers the full distance.
## 0.4.4
### Patch Changes
- fix: mobile keyboard hides terminal content on iPhone
Fixed a bug where opening the virtual keyboard on iPhone left zero visible terminal space. Two independent mechanisms were both accounting for the keyboard height: `MobileDetection.updateAppHeight()` shrunk `--app-height` to the visual viewport height, while `KeyboardHandler.updateLayoutForKeyboard()` added a large `paddingBottom`. These double-counted, leaving negative space for the terminal (user saw accessory bar + toolbar but no terminal content).
Fix: `updateAppHeight()` now skips when the keyboard is visible, and `handleViewportResize()` restores `--app-height` to the pre-keyboard value on first detection (since MobileDetection's listener fires before KeyboardHandler's). On keyboard close, `--app-height` is re-synced to the current visual viewport.
## 0.4.3
### Patch Changes
- Refactor case routes: extract readLinkedCases() and resolveCasePath() helpers to eliminate 6x duplicated linked-cases.json path construction and 5x duplicated file read/parse logic. Replace O(n) .some() duplicate check with O(1) Set.has() in case listing. Un-export unused isError() type guard. Standardize reply.status() to reply.code() in system routes. Update CLAUDE.md frontend module listing and SSE event count.
## 0.4.2
### Patch Changes
- Extract monolithic app.js (~12.5K lines) into 6 focused domain modules that extend CodemanApp.prototype via Object.assign: terminal-ui.js (terminal setup, rendering pipeline, controls), respawn-ui.js (respawn banner, countdown, presets, run summary), ralph-panel.js (Ralph state panel, fix_plan, plan versioning), settings-ui.js (app settings, visibility, web push, tunnel/QR, help), panels-ui.js (subagent panel, teams, insights, file browser, log viewer), session-ui.js (quick start, session options, case settings). Fix critical deferred script init ordering bug: wrap CodemanApp instantiation in DOMContentLoaded so all defer'd mixin modules execute their Object.assign before the constructor runs. Guard missing cleanupWizardDragging() call in subagent-windows.js. Update build.mjs to minify/hash all new modules.
## 0.4.1
### Patch Changes
- Performance optimizations: V8 compile cache for 10-20% faster cold starts, lazy-load WebGL addon (244KB saved on mobile), preload hints for critical scripts, batch tmux reconciliation (N subprocess calls → 1). Also: WebSocket session lifecycle fixes, CJK IME input support, CI upgrade to Node 24/actions v6, install.sh fork support, and CLAUDE.md/README documentation refresh.
## 0.4.0
### Minor Changes
- Add CJK IME input textarea for xterm.js terminal (env toggle INPUT_CJK_FORM=ON). Always-visible textarea below terminal handles native browser IME composition, forwarding completed text to PTY on Enter. Supports arrow keys, Ctrl combos, backspace passthrough, and Escape to clear.
Add fork installation support to install.sh with CODEMAN_REPO_URL and CODEMAN_BRANCH env vars, allowing custom repository and branch for git clone/update operations. README updated with fork installation instructions.
Fix WebSocket session lifecycle: close WS connections when session exits (prevents orphaned listeners and stale writes to dead PTY), add readyState guard in onTerminal to stop buffering after socket closes, simplify heartbeat by removing redundant alive flag.
Add WebSocket reconnection with exponential backoff (1s-10s) on unexpected close, skipping server rejection codes (4004/4008/4009). Falls back gracefully to SSE+POST during reconnection.
Clear CJK textarea on session switch to prevent sending stale text to wrong session.
## 0.3.12
### Patch Changes
- Add WebSocket terminal I/O with server-side DEC 2026 synchronized update markers. Replaces per-keystroke HTTP POST + SSE terminal output with a single bidirectional WebSocket connection for dramatically lower input latency. Server-side 8ms micro-batching with 16KB flush threshold groups rapid PTY events into single WS frames wrapped in DEC 2026 markers for flicker-free atomic rendering. Includes 30s ping/pong heartbeat with 10s timeout for stale connection detection through tunnels. Existing SSE + HTTP POST paths remain fully functional as transparent fallback. Resize messages validated to match HTTP route bounds (cols 1-500, rows 1-200, integers only). 16 automated route tests added for WS endpoint. Also patches 5 dependency vulnerabilities (basic-ftp, fastify, minimatch, serialize-javascript).
## 0.3.11
### Patch Changes
- ### Session Resume & History
- Add `resumeSessionId` support for conversation resume after reboot
- Add history session resume UI and API with route shell sessions routing fix
- Improve session resume reliability and persist user settings across refresh
- Correct `claudeSessionId` for resumed sessions
### Terminal & Frontend
- Upgrade xterm.js 5.3 → 6.0 with native DEC 2026 synchronized output
- Increase terminal scrollback from 5,000 to 20,000 lines
- Reduce default font size and persist tab state across refresh
- Resolve terminal resize scrollback ghost renders
- Hide subagent monitor panel by default
### Installer
- Auto-detect existing install and run update instead of fresh install
- Auto-restart codeman-web service after update if running
- Show restart command when codeman-web is not a systemd service
- Fix one-liner restart command for background processes
### Codebase Quality
- Remove dead code, consolidate imports, extract constants
- Repair 15 pre-existing subagent-watcher test failures
- Clean up DEC sync dead code
## 0.3.10
### Patch Changes
+39 -43
View File
@@ -6,11 +6,12 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
| Task | Command |
|------|---------|
| Dev server | `npx tsx src/index.ts web` |
| Dev server | `npm run dev` (or `npx tsx src/index.ts web`) |
| Type check | `tsc --noEmit` |
| Lint | `npm run lint` (fix: `npm run lint:fix`) |
| Format | `npm run format` (check: `npm run format:check`) |
| Single test | `npx vitest run test/<file>.test.ts` |
| Single test | `npm test -- test/<file>.test.ts` (or `npx vitest run --config config/vitest.config.ts test/<file>.test.ts`) — ⚠ **never** run bare `npm test`, see Testing section |
| Build | `npm run build` (esbuild via `scripts/build.mjs`, NOT tsc — `tsc --noEmit` is type-check only) |
| Production | `npm run build && systemctl --user restart codeman-web` |
## CRITICAL: Session Safety
@@ -48,11 +49,14 @@ When user says "COM":
CHANGESET
```
Replace `patch` with `minor` or `major` as needed. Include `"xterm-zerolag-input": patch` on a separate line if that package changed too.
3. **Consume the changeset**: `npm run version-packages` (bumps versions in `package.json` files and updates `CHANGELOG.md`)
3. **Consume the changeset**: `npm run version-packages` (auto-bumps `package.json` files, updates `CHANGELOG.md`, runs `npm install --package-lock-only`, and verifies lockfile sync via `scripts/check-lockfile-sync.mjs` — all in one command; never hand-edit `CHANGELOG.md` or `package-lock.json` versions)
4. **Sync CLAUDE.md version**: Update the `**Version**` line below to match the new version from `package.json`
5. **Commit and deploy**: `git add -A && git commit -m "chore: version packages" && git push && npm run build && systemctl --user restart codeman-web`
6. **Wait for CI**: after `git push`, find the run with `gh run list -L 1 --json databaseId,headBranch -q '.[0].databaseId'` and watch it with `gh run watch <id> --exit-status`. Confirm all checks pass before considering the release done.
**Version**: 0.3.10 (must match `package.json`)
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
**Version**: 0.6.8 (must match `package.json`)
## Project Overview
@@ -73,21 +77,27 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| Task | Command |
|------|---------|
| Dev with TLS | `npx tsx src/index.ts web --https` |
| Override window title hostname | `npx tsx src/index.ts web --title-hostname <name>` (default: `os.hostname()` — `codeman:<name>` is used for tab title, title-flash, and OS desktop notification prefix) |
| Continuous typecheck | `tsc --noEmit --watch` |
| Test coverage | `npm run test:coverage` |
| Dead-code sweep | `npm run knip` (config in `knip.json`) |
| Production start | `npm run start` |
| Production logs | `journalctl --user -u codeman-web -f` |
**CI**: `.github/workflows/ci.yml` runs `typecheck`, `lint`, `format:check` on push to master (Node 22). Tests excluded (they spawn tmux).
**CI**: `.github/workflows/ci.yml` runs `check:lockfile`, `typecheck`, `lint`, `format:check` on push to master/main and on PRs (Node 22). Tests excluded (they spawn tmux).
**Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`). ESLint flat config (`eslint.config.js`) allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Ignores: `app.js`, `scripts/**/*.mjs`, `src/web/public/vendor/**`, `tools/**`, `remotion/**`.
**Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`). ESLint flat config (`config/eslint.config.js`) allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Ignores: `app.js`, `scripts/**/*.mjs`, `src/web/public/vendor/**`, `scripts/remotion/**`.
## Common Gotchas
- **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink
- **ESM only** — Never `require()`, use `await import()`. `tsx` masks CJS/ESM issues in dev but production breaks
- **Package ≠ product name** — npm: `aicodeman`, product: **Codeman**. Release renames tags accordingly
- **Global regex `lastIndex`** — Use `createAnsiPatternFull/Simple()` factories, not shared `g`-flag patterns in loops
- **Global regex `lastIndex`** — Shared `g`-flag patterns in loops must reset `lastIndex = 0` first, or use the `execPattern()` helper in `utils/regex-patterns.ts` (resets automatically)
- **`envOverrides` flow `CLAUDE_CODE_*` / `OPENCODE_*` env vars** — Set via `POST /api/sessions { envOverrides }`, stored on `Session._envOverrides`, exported by `tmux-manager.buildEnvExports()` at spawn time, persisted in `SessionState.envOverrides`. **Do NOT** write these to `<case>/.claude/settings.local.json` — that's the old path and creates UI/disk drift
- **Dual-CLI prefix discipline** — Codeman supports both Claude Code and OpenCode (`claude-cli-resolver.ts` / `opencode-cli-resolver.ts`); env-var prefix is CLI-specific (`CLAUDE_CODE_*` vs `OPENCODE_*`) and the allowlist in `schemas.ts` enforces this. When adding settings, decide which CLI(s) it applies to and gate the env export accordingly — don't blindly forward both prefixes
- **Zod `.optional()` rejects `null`** — accepts `undefined` only. When the frontend builds a request body with `JSON.stringify`, an explicit `null` field is preserved on the wire and fails validation with `INVALID_INPUT`. Convert `null` → `undefined` before stringifying (e.g. `field: value ?? undefined`), or declare the schema `.nullish()`. Real bugs caused: 0.6.4 (`durationMinutes` for ∞ respawn), and the same shape pattern hit `opusContext1mEnabled` in 0.6.3
- **`xterm-zerolag-input` is duplicated** — the local-echo overlay lives in BOTH `packages/xterm-zerolag-input/src/` (published to npm as a standalone library for external consumers — see README "Published Packages") AND inline inside `src/web/public/app.js` (runtime copy the web UI actually loads, since the page ships as plain JS without a bundler). Any change to overlay behavior MUST be applied to both, or dev and prod diverge — and a public API break in the package warrants a separate version bump for `xterm-zerolag-input` in the changeset. Always test on mobile after touching it.
**Import conventions**: Utils from `./utils`, types from `./types` (barrel), config from specific `./config/*` files.
@@ -102,23 +112,24 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| **Mux** | `src/mux-interface.ts`, `src/mux-factory.ts`, `src/tmux-manager.ts` | |
| **Respawn** | `src/respawn-controller.ts` ★ + 4 helpers (`-adaptive-timing`, `-health`, `-metrics`, `-patterns`) | Read `docs/respawn-state-machine.md` first |
| **Ralph** | `src/ralph-tracker.ts` ★, `src/ralph-loop.ts` + 5 helpers (`-config`, `-fix-plan-watcher`, `-plan-tracker`, `-stall-detector`, `-status-parser`) | Read `docs/ralph-wiggum-guide.md` first |
| **Orchestrator** | `src/orchestrator-loop.ts`, `src/orchestrator-planner.ts`, `src/orchestrator-verifier.ts` | Read `docs/orchestrator-loop-architecture.md` first |
| **Agents** | `src/subagent-watcher.ts` ★, `src/team-watcher.ts`, `src/bash-tool-parser.ts`, `src/transcript-watcher.ts` | |
| **AI** | `src/ai-checker-base.ts`, `src/ai-idle-checker.ts`, `src/ai-plan-checker.ts` | |
| **Tasks** | `src/task.ts`, `src/task-queue.ts`, `src/task-tracker.ts` | |
| **State** | `src/state-store.ts`, `src/run-summary.ts`, `src/session-lifecycle-log.ts` | |
| **Infra** | `src/hooks-config.ts`, `src/push-store.ts`, `src/tunnel-manager.ts`, `src/image-watcher.ts`, `src/file-stream-manager.ts` | |
| **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/claude-md.ts` | |
| **Web** | `src/web/server.ts`, `src/web/sse-events.ts`, `src/web/routes/*.ts` (12 route modules + barrel), `src/web/ports/*.ts`, `src/web/middleware/auth.ts`, `src/web/schemas.ts` | |
| **Frontend** | `src/web/public/app.js` ★ (~12.1K lines) + 9 JS modules (incl. `sw.js` service worker) | |
| **Types** | `src/types/index.ts` → 14 domain files | See `@fileoverview` in index.ts |
| **Web** | `src/web/server.ts`, `src/web/sse-events.ts`, `src/web/routes/*.ts` (15 route modules + barrel), `src/web/route-helpers.ts`, `src/web/ports/*.ts`, `src/web/middleware/auth.ts`, `src/web/schemas.ts` | |
| **Frontend** | `src/web/public/app.js` (~2.9K lines, core) + 5 infra modules (`constants.js`, `mobile-handlers.js`, `voice-input.js`, `notification-manager.js`, `keyboard-accessory.js`) + 7 domain modules (`terminal-ui.js`, `respawn-ui.js`, `ralph-panel.js`, `orchestrator-panel.js`, `settings-ui.js`, `panels-ui.js`, `session-ui.js`) + 4 feature modules (`ralph-wizard.js`, `api-client.js`, `subagent-windows.js`, `input-cjk.js`) + `sw.js` | |
| **Types** | `src/types/index.ts` (barrel) → 14 domain files; also `src/types.ts` root re-export | See `@fileoverview` in index.ts |
★ = Large file (>50KB). All files have `@fileoverview` JSDoc — read that before diving in.
★ = Large file (>50KB). All files have `@fileoverview` JSDoc — read that before diving in. Discovery aid: `grep -l '@fileoverview' src/web/routes/*.ts` lists all route modules; same grep works for `src/types/`, `src/web/public/*.js`.
**Local package**: `packages/xterm-zerolag-input/` — local echo overlay for xterm.js; copy embedded in `app.js`.
**Config**: `src/config/` — 9 files. Import from specific files, not barrel.
**Utilities**: `src/utils/` — re-exported via index. Key: `CleanupManager`, `LRUMap`, `StaleExpirationMap`, `BufferAccumulator`, `stripAnsi`, `Debouncer`.
**Utilities**: `src/utils/` — re-exported via index. Key: `CleanupManager`, `LRUMap`, `StaleExpirationMap`, `BufferAccumulator`, `stripAnsi`, `Debouncer`, `KeyedDebouncer`. Also: `claude-cli-resolver`/`opencode-cli-resolver` (CLI path resolution), `string-similarity` (fuzzy matching), `regex-patterns` (ANSI/token/spinner patterns), `assertNever` (exhaustive checks), `token-validation` (auth tokens), `nice-wrapper` (process priority).
### Data Flow
@@ -135,7 +146,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Hook events**: Claude Code hooks trigger via `/api/hook-event`. Key events: `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`. See `src/hooks-config.ts`.
**Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `agent-teams/`.
**Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `docs/agent-teams/`.
**Circuit breaker**: Prevents respawn thrashing. States: `CLOSED` → `HALF_OPEN` → `OPEN`. Reset: `/api/sessions/:id/ralph-circuit-breaker/reset`.
@@ -143,13 +154,13 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
### Frontend
Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `app.js`(6) → `ralph-wizard.js`(7) → `api-client.js`(8) → `subagent-windows.js`(9).
Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `input-cjk.js`(5.5) → `app.js`(6) → `terminal-ui.js`(7) → `respawn-ui.js`(8) → `ralph-panel.js`(9) → `orchestrator-panel.js`(9.5) → `settings-ui.js`(10) → `panels-ui.js`(11) → `session-ui.js`(12) → `ralph-wizard.js`(13) → `api-client.js`(14) → `subagent-windows.js`(15). `input-cjk.js` handles CJK IME composition via an always-visible textarea below the terminal (`window.cjkActive` blocks xterm's onData).
**Z-index layers**: subagent windows (1000), plan agents (1100), log viewers (2000), image popups (3000), local echo overlay (7).
**Respawn presets**: `solo-work` (3s/60min), `subagent-workflow` (45s/240min), `team-lead` (90s/480min), `ralph-todo` (8s/480min), `overnight-autonomous` (10s/480min).
**Keyboard shortcuts**: Escape (close), Ctrl+? (help), Ctrl+Enter (quick start), Ctrl+W (kill), Ctrl+Tab (next), Ctrl+K (kill all), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl/Cmd +/- (font).
**Keyboard shortcuts**: Escape (close), Ctrl+? (help), Ctrl+W (kill), Ctrl+Tab (next), Alt+1-9 (switch tab), Ctrl+Shift+{/} (move tab left/right), Shift+Enter (newline), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl+Shift+V (voice input), Ctrl/Cmd +/- (font).
### Security
@@ -166,11 +177,11 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
### SSE Event Registry
~100 event types in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). Both must be kept in sync.
~120 event types in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). Both must be kept in sync.
### API Routes
~111 handlers across 12 route files in `src/web/routes/`: system (35), sessions (24), ralph (9), plan (8), respawn (7), cases (7), files (5), mux (5), scheduled (4), push (4), teams (2), hooks (1). Each file has `@fileoverview` with endpoint details.
~128 handlers across 15 route files in `src/web/routes/`: system (36), sessions (27), orchestrator (10), cases (9), ralph (9), plan (8), respawn (7), files (5), mux (5), push (4), scheduled (4), teams (2), hooks (1), clipboard (1), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
## Adding Features
@@ -192,22 +203,20 @@ All in `~/.codeman/`: `state.json` (sessions, settings, respawn), `mux-sessions.
**CRITICAL: You are running inside a Codeman-managed tmux session.** Never run `npx vitest run` (full suite) — it spawns/kills tmux sessions and will crash your own session. Only run individual files:
```bash
npx vitest run test/<specific-file>.test.ts # Single file (SAFE)
npx vitest run -t "pattern" # By name (SAFE)
# npx vitest run # DANGEROUS — DON'T DO THIS
npm test -- test/<specific-file>.test.ts # Single file (SAFE, uses config/vitest.config.ts)
npm test -- -t "pattern" # By name (SAFE)
# npm test # DANGEROUS — runs full suite, DON'T DO THIS
```
Raw `npx vitest` skips `config/vitest.config.ts`; always use `npm test --` or pass `--config config/vitest.config.ts`.
**Config**: Vitest with `globals: true`, `fileParallelism: false`. Timeout 30s, teardown 60s.
**Safety**: `test/setup.ts` snapshots pre-existing tmux sessions and never kills them. Only `registerTestTmuxSession()` sessions get cleaned up.
**Ports**: Pick unique ports manually. Search `const PORT =` before adding new tests.
**Respawn tests**: Use `MockSession` from `test/respawn-test-utils.ts`. **Route tests**: `app.inject()` in `test/routes/`. **Mobile tests**: Playwright suite in `mobile-test/` (135 device profiles).
## Screenshots
Mobile screenshots in `~/.codeman/screenshots/`. API: `GET /api/screenshots`, `POST /api/screenshots`.
**Respawn tests**: Use `MockSession` from `test/respawn-test-utils.ts`. **Route tests**: `app.inject()` in `test/routes/`. **Mobile tests**: Playwright suite in `test/mobile/` (135 device profiles).
## Debugging
@@ -219,27 +228,14 @@ curl localhost:3000/api/subagents | jq # Background agents
cat ~/.codeman/state.json | jq # Persisted state
```
Mobile screenshots: `~/.codeman/screenshots/`, accessed via `GET/POST /api/screenshots`.
## Performance & Limits
Target: 20 sessions, 50 agent windows at 60fps. Limits in `src/config/`: terminal 2MB, text 1MB, messages 1000, max agents 500, max sessions 50, max SSE clients 100. Use `LRUMap` for bounded caches, `StaleExpirationMap` for TTL cleanup. Anti-flicker pipeline: `docs/terminal-anti-flicker.md`.
## References
**Memory leaks (24+ hour sessions)**: use `CleanupManager`, clear Maps in `stop()`, guard async with `if (this.cleanup.isStopped) return`. Frontend: store handler refs, clean in `close*()`. Verify: `npm test -- test/memory-leak-prevention.test.ts`.
Deep-dive docs in `docs/`: `respawn-state-machine.md`, `ralph-wiggum-guide.md`, `claude-code-hooks-reference.md`, `terminal-anti-flicker.md`, `opencode-integration.md`, `qr-auth-plan.md`. Agent Teams: `agent-teams/README.md`. SSE events: `src/web/sse-events.ts` + `constants.js`.
## Scripts & Tunnel
## Scripts
Key: `scripts/tmux-manager.sh` (safe tmux mgmt), `scripts/tunnel.sh` (tunnel start/stop/url). Production: `scripts/codeman-web.service`, `scripts/codeman-tunnel.service`.
## Memory Leak Prevention
24+ hour sessions: use `CleanupManager`, clear Maps in `stop()`, guard async with `if (this.cleanup.isStopped) return`. Frontend: store handler refs, clean in `close*()`. Verify: `npx vitest run test/memory-leak-prevention.test.ts`.
## Common Workflows
**Bug investigation**: Dev server → reproduce in browser → check terminal + `~/.codeman/state.json`.
**Respawn changes**: Read `docs/respawn-state-machine.md` first. Use `MockSession` from `test/respawn-test-utils.ts`.
## Tunnel
`./scripts/tunnel.sh start|stop|url`. **Always set `CODEMAN_PASSWORD`** before exposing via tunnel.
Key scripts: `scripts/tmux-manager.sh` (safe tmux mgmt), `scripts/tunnel.sh start|stop|url` (tunnel). Production services: `scripts/codeman-web.service`, `scripts/codeman-tunnel.service`. **Always set `CODEMAN_PASSWORD`** before exposing via tunnel.
+64 -11
View File
@@ -11,7 +11,7 @@
<p align="center">
<a href="https://opensource.org/licenses/MIT"><img src="https://img.shields.io/badge/License-MIT-1e3a5f?style=flat-square" alt="License: MIT"></a>
<a href="https://nodejs.org/"><img src="https://img.shields.io/badge/Node.js-18%2B-22c55e?style=flat-square&logo=node.js&logoColor=white" alt="Node.js 18+"></a>
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/TypeScript-5.5-3b82f6?style=flat-square&logo=typescript&logoColor=white" alt="TypeScript 5.5"></a>
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/TypeScript-5.9-3b82f6?style=flat-square&logo=typescript&logoColor=white" alt="TypeScript 5.9"></a>
<a href="https://fastify.dev/"><img src="https://img.shields.io/badge/Fastify-5.x-1e3a5f?style=flat-square&logo=fastify&logoColor=white" alt="Fastify"></a>
<img src="https://img.shields.io/badge/Tests-1435%20total-22c55e?style=flat-square" alt="Tests">
</p>
@@ -28,29 +28,67 @@
curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash
```
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it. You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code) or [OpenCode](https://opencode.ai) (or both). After install:
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code) or [OpenCode](https://opencode.ai) (or both). After install:
```bash
codeman web
# Open http://localhost:3000 — press Ctrl+Enter to start your first session
```
**Update to latest version:**
```bash
curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash -s update
```
<details>
<summary><strong>Run as a background service</strong></summary>
**Linux (systemd):**
```bash
mkdir -p ~/.config/systemd/user && printf '[Unit]\nDescription=Codeman Web Server\nAfter=network.target\n\n[Service]\nType=simple\nExecStart=%s %s/dist/index.js web\nRestart=always\nRestartSec=10\n\n[Install]\nWantedBy=default.target\n' "$(which node)" "$HOME/.codeman/app" > ~/.config/systemd/user/codeman-web.service && systemctl --user daemon-reload && systemctl --user enable --now codeman-web && loginctl enable-linger $USER
mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/codeman-web.service << EOF
[Unit]
Description=Codeman Web Server
After=network.target
[Service]
Type=simple
ExecStart=$(which node) $HOME/.codeman/app/dist/index.js web
Restart=always
RestartSec=10
[Install]
WantedBy=default.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now codeman-web
loginctl enable-linger $USER
```
**macOS (launchd):**
```bash
mkdir -p ~/Library/LaunchAgents && printf '<?xml version="1.0" encoding="UTF-8"?>\n<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">\n<plist version="1.0"><dict><key>Label</key><string>com.codeman.web</string><key>ProgramArguments</key><array><string>%s</string><string>%s/dist/index.js</string><string>web</string></array><key>RunAtLoad</key><true/><key>KeepAlive</key><true/><key>StandardOutPath</key><string>/tmp/codeman.log</string><key>StandardErrorPath</key><string>/tmp/codeman.log</string></dict></plist>\n' "$(which node)" "$HOME/.codeman/app" > ~/Library/LaunchAgents/com.codeman.web.plist && launchctl load ~/Library/LaunchAgents/com.codeman.web.plist
mkdir -p ~/Library/LaunchAgents
cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
"http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.codeman.web</string>
<key>ProgramArguments</key>
<array>
<string>$(which node)</string>
<string>$HOME/.codeman/app/dist/index.js</string>
<string>web</string>
</array>
<key>RunAtLoad</key><true/>
<key>KeepAlive</key><true/>
<key>StandardOutPath</key>
<string>/tmp/codeman.log</string>
<key>StandardErrorPath</key>
<string>/tmp/codeman.log</string>
</dict>
</plist>
EOF
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
```
</details>
@@ -188,6 +226,17 @@ Run **20 parallel sessions** with full visibility — real-time xterm.js termina
Every session runs inside **tmux** — sessions survive server restarts, network drops, and machine sleep. Auto-recovery on startup with dual redundancy. Ghost session discovery finds orphaned tmux sessions. Managed sessions are environment-tagged so the agent won't kill its own session.
### Hostname-Aware Window Title
Running Codeman on multiple hosts (laptop, dev box, NAS)? The browser tab title is `codeman:<hostname>` so you can tell which backend each tab points at without clicking in:
```bash
codeman web # codeman:<os.hostname()>
codeman web --title-hostname dev-box # codeman:dev-box (manual override for noisy hostnames)
```
The title is templated into the served HTML on first byte, so it's correct from the very first paint and works without JavaScript. The same hostname prefix is applied to the tab-flash format (`⚠️ (N) codeman:<host>`) and to OS-level desktop notifications (`codeman:<host>: <event>`), so cross-host alerts in the system notification center are also unambiguous.
### Smart Token Management
| Threshold | Action | Result |
@@ -363,9 +412,12 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
| `Ctrl+Enter` | Quick-start session |
| `Ctrl+W` | Close session |
| `Ctrl+Tab` | Next session |
| `Alt+1`–`Alt+9` | Switch to tab N |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move active tab left / right |
| `Ctrl+K` | Kill all sessions |
| `Ctrl+L` | Clear terminal |
| `Ctrl+Shift+R` | Restore terminal size |
| `Ctrl+Shift+V` | Toggle voice input |
| `Ctrl/Cmd +/-` | Font size |
| `Escape` | Close panels |
@@ -408,6 +460,7 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
| `GET` | `/api/events` | SSE stream |
| `GET` | `/api/status` | Full app state |
| `POST` | `/api/hook-event` | Hook callbacks |
| `POST` | `/api/clipboard` | Push text to all connected browsers (`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | Timeline + stats |
---
@@ -484,9 +537,9 @@ The codebase went through a comprehensive 7-phase refactoring that eliminated go
| Phase | What changed | Impact |
|-------|-------------|--------|
| **Performance** | Cached endpoints, SSE adaptive batching, buffer chunking | Sub-16ms terminal latency |
| **Route extraction** | `server.ts` split into 12 domain route modules + auth middleware + port interfaces | **−60%** server.ts LOC (6,736 → 2,697) |
| **Route extraction** | `server.ts` split into 13 domain route modules + auth middleware + port interfaces | **−60%** server.ts LOC (6,736 → 2,697) |
| **Domain splitting** | `types.ts` → 14 domain files, `ralph-tracker` → 7 files, `respawn-controller` → 5 files, `session` → 6 files | No more god files |
| **Frontend modules** | `app.js` → 8 extracted modules (constants, mobile, voice, notifications, keyboard, API, subagent windows) | **−24%** app.js LOC (15.2K → 11.5K) |
| **Frontend modules** | `app.js` → 9 extracted modules (constants, mobile, voice, notifications, keyboard, CJK input, API, Ralph wizard, subagent windows) | **−24%** app.js LOC (15.2K → 11.5K) |
| **Config consolidation** | ~70 scattered magic numbers → 9 domain-focused config files | Zero cross-file duplicates |
| **Test infrastructure** | Shared mock library, 12 route test files, consolidated MockSession | Testable route handlers via `app.inject()` |
+1 -2
View File
@@ -22,8 +22,7 @@ export default tseslint.config(
'src/web/public/vendor/**',
'src/web/public/app.js',
'scripts/**/*.mjs',
'tools/**',
'remotion/**',
'scripts/remotion/**',
],
}
);
@@ -1,7 +1,11 @@
import { resolve } from 'node:path';
import { defineConfig } from 'vitest/config';
const root = resolve(import.meta.dirname, '..');
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: ['test/**/*.test.ts'],
+74
View File
@@ -0,0 +1,74 @@
# Codeman Performance Optimization Plan
## Current State
The backend is **already production-grade** — SSE broadcasting, state persistence, terminal batching, buffer management, and memory patterns are all well-optimized. The biggest gains are on the **frontend delivery** side.
## Implemented Optimizations
### 1. V8 Compile Cache (10-20% faster cold start)
**Files:** `scripts/codeman-web.service`, `package.json`
Node.js re-parses and compiles all JS on every cold start. `NODE_COMPILE_CACHE` caches V8 compiled bytecode to disk, reusing it on subsequent starts.
- Added `Environment=NODE_COMPILE_CACHE=/home/arkon/.codeman/compile-cache` to systemd service
- Added to `npm start` script for non-systemd usage
- Zero code changes, immediate win on every restart
### 2. WebGL Addon Lazy-Loading (244KB saved on mobile, non-blocking on desktop)
**Files:** `src/web/public/index.html`, `src/web/public/app.js`
`xterm-addon-webgl.min.js` (244KB) was loaded eagerly for all users via `<script defer>`, but only used on desktop with WebGL2 support.
- Removed `<script defer>` from `index.html`
- Added dynamic script loading in `app.js` — only downloads on desktop when WebGL is needed
- Mobile users never download the file at all (244KB saved)
- Desktop: loads in parallel with page rendering, addon initializes when ready
- Graceful fallback: canvas renderer used if WebGL unavailable or script fails
### 3. Preload Hints (~50-100ms faster perceived load)
**Files:** `src/web/public/index.html`
Browser discovers `<script defer>` tags only when the parser reaches them at the bottom of `<body>`. By then, the HTML parse has blocked for hundreds of lines.
- Added `<link rel="preload" as="script">` in `<head>` for `vendor/xterm.min.js`, `constants.js`, `app.js`
- Browser starts fetching critical scripts immediately during HTML parse (before reaching `<body>`)
- Zero runtime overhead — just hints for the browser's preload scanner
### 4. Batch Tmux Reconciliation (N subprocess calls → 1)
**Files:** `src/tmux-manager.ts`
`reconcileSessions()` previously called `tmux has-session` + `tmux display-message` per known session, plus `tmux list-sessions` for discovery, plus `tmux display-message` per discovered session. With 20 sessions: 41+ subprocess calls.
- Replaced with single `tmux list-panes -a -F '#{session_name}\t#{pane_pid}'` call
- Builds a Map from the result, then does O(1) lookups for both known and discovered sessions
- Also replaced inner O(n) `isKnown` scan with a Set lookup
- 20 sessions: 41 subprocess calls → 1, with faster lookups
### 5. Asset Hashing / Cache Busting (already implemented)
**Files:** `scripts/build.mjs` (pre-existing)
Content-hash cache busting was already implemented in the build script:
- All app JS/CSS files get content hashes (`app.abc123.js`)
- `index.html` rewritten to reference hashed filenames
- Pre-compressed with gzip + Brotli
- 1-year immutable cache works correctly — new deploys get new filenames
## Already Optimized (No Action Needed)
| Area | Why It's Fine |
|------|---------------|
| **SSE Broadcasting** | Single serialization per broadcast, preformatted frames, backpressure handling, session subscription filtering |
| **State Persistence** | 500ms debounce, incremental per-session JSON caching, async atomic writes, circuit breaker on failures |
| **Terminal Batching** | Adaptive intervals (16-50ms), per-session queues, immediate flush at 32KB, array-based accumulation |
| **Buffer Management** | BufferAccumulator (array-push, lazy join), auto-trim at 2MB/1MB, no string concatenation in hot paths |
| **ANSI Stripping** | Pre-compiled regex via factory functions, single-pass processing |
| **Static File Serving** | @fastify/static with 1-year cache, pre-compressed Brotli/gzip, no-cache for HTML |
| **Memory Management** | CleanupManager, LRUMap, StaleExpirationMap, bounded buffers, explicit listener cleanup |
| **Import Patterns** | Pure ESM, lazy web server import, no circular deps, no dynamic imports in hot paths |
| **Config Loading** | Small constant files, no I/O at import time, specific imports (no barrel) |
Binary file not shown.
+6 -4
View File
@@ -1,5 +1,7 @@
# Local Echo Overlay — Implementation Plan
> **Status: SHIPPED.** Implementation lives in `packages/xterm-zerolag-input/src/` (overlay-renderer.ts, prompt-finder.ts, cell-dimensions.ts, zerolag-input-addon.ts) with the embedded copy in `src/web/public/app.js`. This document is retained as historical design context.
## Context
User accesses Codeman remotely from Thailand to Switzerland over Tailscale (~200-300ms RTT).
@@ -18,9 +20,9 @@ redraws. A DOM overlay sits in a separate rendering layer (z-index 7) and doesn'
with Ink's cursor management or screen redraws at all. When Ink redraws (server output arrives),
we simply hide the overlay.
**Why it will look indistinguishable:** We use the DOM renderer (not canvas/WebGL) in our
xterm.js v5.3.0, so both terminal text and overlay text are rendered by the same browser
font engine with identical sub-pixel rendering.
**Why it will look indistinguishable:** We use the DOM renderer (not canvas/WebGL), so both
terminal text and overlay text are rendered by the same browser font engine with identical
sub-pixel rendering. (Originally designed against xterm.js v5.3.0; project now on `@xterm/xterm` ^6.0.0 — the internal `_core._renderService.dimensions` access path still works in v6.)
## Key Technical Details (from research)
@@ -36,7 +38,7 @@ const top = cursorY * dims.css.cell.height; // CSS pixels, relative to .xterm-
- `cursorY` = `terminal.buffer.active.cursorY` (0 to terminal.rows-1, ALREADY viewport-relative)
- No scroll offset math needed
### Cell Dimensions (v5.3.0 — no public API, use internal)
### Cell Dimensions (no public API in v5/v6 — use internal; public in v7+)
```js
const dims = terminal._core._renderService.dimensions;
dims.css.cell.width // e.g., 8.4px
+367
View File
@@ -0,0 +1,367 @@
# Orchestrator Loop — Architecture & Data Flow
> Technical architecture document. Not for GitHub.
## System Overview
```
┌─────────────────────────────────────────────────────────────────────┐
│ CODEMAN WEB UI │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ Orchestrator Dashboard │ │
│ │ [Goal Input] [Plan View] [Phase Progress] [Agent Activity] │ │
│ └───────────────────────────┬──────────────────────────────────┘ │
│ │ SSE Events │
│ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ Orchestrator API Routes (/api/orchestrator/*) │ │
│ └───────────────────────────┬──────────────────────────────────┘ │
└───────────────────────────────┼─────────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────────────┐
│ ORCHESTRATOR LOOP │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────────────┐ │
│ │ Orchestrator │ │ Orchestrator │ │ Orchestrator │ │
│ │ Planner │ │ Loop (state │ │ Verifier │ │
│ │ │ │ machine) │ │ │ │
│ │ • Research │◄──►│ • Phase mgmt │◄──►│ • Test runner │ │
│ │ • Plan gen │ │ • Task queue │ │ • AI review │ │
│ │ • Phasing │ │ • Event loop │ │ • Output checks │ │
│ └──────┬───────┘ └──────┬───────┘ └──────────┬───────────┘ │
│ │ │ │ │
│ ▼ ▼ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ EXISTING CODEMAN INFRASTRUCTURE │ │
│ │ │ │
│ │ SessionManager ←→ Sessions ←→ PTY (Claude CLI) │ │
│ │ ↑ ↑ ↑ │ │
│ │ │ │ │ │ │
│ │ TaskQueue RalphTracker RespawnController │ │
│ │ StateStore HooksConfig TeamWatcher │ │
│ │ Auto-Ops SubagentWatcher SSE Broadcast │ │
│ └──────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
```
## Data Flow: Complete Lifecycle
### 1. User Submits Goal
```
User → POST /api/orchestrator/start { goal: "Build a REST API...", config: {...} }
→ OrchestratorLoop.start(goal)
→ state = PLANNING
→ emit('stateChanged', 'planning')
→ SSE: orchestrator:stateChanged
```
### 2. Planning Phase
```
OrchestratorPlanner.generatePlan(goal)
→ PlanOrchestrator.generateDetailedPlan(goal)
→ [Research Agent] → enriched task description
→ [Planner Agent] → PlanItem[]
→ groupIntoPhases(planItems)
→ topological sort by dependencies
→ group into layers
→ assign team strategies
→ OrchestratorPlan { phases: [...] }
→ state = APPROVAL
→ emit('planReady', plan)
→ SSE: orchestrator:planReady
```
### 3. User Approves Plan
```
User → POST /api/orchestrator/approve
→ OrchestratorLoop.approvePlan()
→ state = EXECUTING
→ executePhase(phases[0])
```
### 4. Phase Execution
```
executePhase(phase)
→ For each task in phase:
→ Convert to CreateTaskOptions
→ Add to TaskQueue with completion phrase "PHASE_{N}_TASK_{M}_DONE"
→ If phase.teamStrategy.type === 'team':
→ Start session with AGENT_TEAMS enabled
→ Send team orchestration prompt to lead
→ Else:
→ Assign tasks to available sessions (same as RalphLoop)
→ Listen for task completion events:
→ TaskQueue emits taskCompleted
→ Check: all phase tasks done?
→ Yes → state = VERIFYING → verifyPhase(phase)
→ No → wait for more completions
```
### 5. Verification
```
verifyPhase(phase)
→ OrchestratorVerifier.verify(phase, session)
→ Run test commands via session
→ Check file existence
→ AI review (optional)
→ If passed:
→ phase.status = 'passed'
→ emit('phaseCompleted', phase)
→ If more phases: executePhase(nextPhase)
→ If last phase: state = COMPLETED
→ If failed:
→ phase.attempts++
→ If attempts < maxAttempts:
→ state = REPLANNING
→ Generate recovery tasks
→ state = EXECUTING (retry)
→ Else:
→ state = FAILED
→ emit('phaseFailed', phase, reason)
```
### 6. Context Management Between Phases
```
After phase completion:
→ If config.compactBetweenPhases:
→ session.sendInput('/compact')
→ Wait for compact to complete
→ If config.respawnBetweenMilestones && phase is a milestone:
→ Save orchestrator state to StateStore
→ Respawn session (kill + recreate)
→ Send resume prompt with phase context
```
## File Layout
```
src/
├── orchestrator-loop.ts # Main state machine (~400 lines)
├── orchestrator-planner.ts # Plan generation + phase grouping (~300 lines)
├── orchestrator-verifier.ts # Phase verification (~200 lines)
├── types/
│ └── orchestrator.ts # All orchestrator types (~150 lines)
├── prompts/
│ └── orchestrator.ts # Prompt templates (~200 lines)
├── web/
│ ├── routes/
│ │ └── orchestrator-routes.ts # API endpoints (~250 lines)
│ └── public/
│ └── orchestrator-ui.js # Frontend panel (~500 lines)
```
## Integration Points with Existing Code
### StateStore (`src/state-store.ts`)
```typescript
// Add to AppState interface
orchestrator?: OrchestratorPersistState;
// Add methods
getOrchestratorState(): OrchestratorPersistState;
setOrchestratorState(state: Partial<OrchestratorPersistState>): void;
```
### SSE Events (`src/web/sse-events.ts`)
```typescript
// Add ~8 new events
export const SseEvent = {
// ... existing
ORCHESTRATOR_STATE_CHANGED: 'orchestrator:stateChanged',
ORCHESTRATOR_PLAN_READY: 'orchestrator:planReady',
ORCHESTRATOR_PHASE_STARTED: 'orchestrator:phaseStarted',
ORCHESTRATOR_PHASE_COMPLETED: 'orchestrator:phaseCompleted',
ORCHESTRATOR_PHASE_FAILED: 'orchestrator:phaseFailed',
ORCHESTRATOR_VERIFICATION: 'orchestrator:verificationResult',
ORCHESTRATOR_COMPLETED: 'orchestrator:completed',
ORCHESTRATOR_ERROR: 'orchestrator:error',
} as const;
```
### Frontend Constants (`src/web/public/constants.js`)
```javascript
// Mirror SSE events
SSE_EVENTS.ORCHESTRATOR_STATE_CHANGED = 'orchestrator:stateChanged';
// ... etc
```
### Route Registration (`src/web/routes/index.ts`)
```typescript
import { registerOrchestratorRoutes } from './orchestrator-routes.js';
// Add to barrel export
```
### Server (`src/web/server.ts`)
```typescript
// Initialize OrchestratorLoop alongside RalphLoop
const orchestratorLoop = new OrchestratorLoop(config);
// Register routes
registerOrchestratorRoutes(app, { ...ctx, orchestrator: orchestratorLoop });
```
### Port Interface (`src/web/ports/`)
```typescript
// New port
export interface OrchestratorPort {
orchestrator: OrchestratorLoop;
}
```
## Prompt Flow Through System
The key insight is how prompts flow from Orchestrator → Session → Claude:
```
OrchestratorLoop decides to execute Phase 3, Task 2
│
▼
Converts OrchestratorTask to CreateTaskOptions:
{
prompt: "Implement the rate limiter middleware. Read src/middleware/auth.ts
for the pattern. Add to src/middleware/rate-limiter.ts. Must export
a Fastify plugin. When done: <promise>PHASE_3_TASK_2_DONE</promise>",
priority: 100,
dependencies: ["phase-3-task-1"], // Must finish auth middleware first
completionPhrase: "PHASE_3_TASK_2_DONE",
timeoutMs: 600000 // 10 minutes
}
│
▼
TaskQueue.addTask(options)
│
▼
RalphLoop.tick() → assignTasks() // OR OrchestratorLoop does its own assignment
│
▼
session.sendInput(task.prompt)
│
▼
writeViaMux() → tmux send-keys -l "prompt..." + Enter
│
▼
Claude CLI receives prompt, executes, outputs results
│
▼
RalphTracker.processData() → detects "PHASE_3_TASK_2_DONE"
│
▼
emit('completionDetected') → OrchestratorLoop.handleTaskCompleted()
│
▼
Check: all tasks in Phase 3 done? → If yes → verifyPhase(phase3)
```
## Team Agent Flow (When Enabled)
```
Phase has teamStrategy.type === 'team'
│
▼
OrchestratorLoop creates/reuses a session with:
env: { CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS: '1' }
│
▼
Sends team orchestration prompt:
"You're the team lead for Phase 3: Core Implementation.
Your team should work on these tasks in parallel:
1. Rate limiter middleware (teammate 1)
2. Error handling middleware (teammate 2)
3. Validation layer (teammate 3)
Context files to read first: [...]
Each teammate should output their task's completion phrase when done.
When ALL tasks are complete, output: <promise>PHASE_3_COMPLETE</promise>"
│
▼
Claude Code team-lead spawns teammates
│
▼
TeamWatcher detects new team in ~/.claude/teams/
→ Matches to session via leadSessionId
→ Tracks teammate activity
│
▼
Teammates work in parallel (in-process threads)
│
▼
hook: teammate_idle → POST /api/hook-event
→ OrchestratorLoop notes teammate finished
│
▼
hook: task_completed → POST /api/hook-event
→ Or: RalphTracker detects PHASE_3_COMPLETE
→ OrchestratorLoop → phase complete → verify
```
## Error Recovery Strategy
```
Task fails (timeout, error, session crash)
│
├─ Task-level retry (up to 2 retries per task)
│ → Reset task to pending
│ → Re-queue with modified prompt: "Previous attempt failed: {error}. Try again..."
│
├─ Phase-level retry (up to 3 retries per phase)
│ → Respawn session (fresh context)
│ → Re-execute entire phase with learnings from failure
│ → Modified prompt includes what went wrong
│
└─ Orchestration-level failure
→ All retries exhausted
→ state = FAILED
→ Notify user with detailed failure report
→ User can: modify plan → retry, skip phase → continue, or stop
```
## Interaction with Ralph Loop
Ralph Loop and Orchestrator Loop are **mutually exclusive** on the same sessions:
```
if (orchestratorLoop.isRunning()) {
// Orchestrator controls task assignment
// Ralph Loop should not interfere
// Respawn Controller uses 'orchestrator' preset
}
if (ralphLoop.isRunning()) {
// Ralph controls task assignment
// Orchestrator should not start
}
```
The Orchestrator can optionally USE the Ralph Loop internally for phase execution (delegate phase tasks to Ralph's queue), or manage task assignment directly. Decision: **manage directly** — gives more control over phase boundaries and verification timing.
## Summary of What Touches What
| Existing File | Change |
|---|---|
| `src/types/index.ts` | Export orchestrator types |
| `src/state-store.ts` | Add orchestrator state persistence |
| `src/web/sse-events.ts` | Add ~8 orchestrator events |
| `src/web/routes/index.ts` | Register orchestrator routes |
| `src/web/server.ts` | Initialize OrchestratorLoop |
| `src/web/public/constants.js` | Mirror SSE events |
| `src/web/public/app.js` | Add orchestrator event listeners, panel toggle |
| `src/web/route-helpers.ts` | Add 'orchestrator' respawn preset |
| New File | Purpose |
|---|---|
| `src/orchestrator-loop.ts` | Core state machine |
| `src/orchestrator-planner.ts` | Plan generation + phasing |
| `src/orchestrator-verifier.ts` | Phase verification |
| `src/types/orchestrator.ts` | Type definitions |
| `src/prompts/orchestrator.ts` | Prompt templates |
| `src/web/routes/orchestrator-routes.ts` | API endpoints |
| `src/web/public/orchestrator-ui.js` | Frontend panel |
| `src/web/ports/orchestrator-port.ts` | Port interface |
+633
View File
@@ -0,0 +1,633 @@
# Orchestrator Loop — Detailed Implementation Plan (v2)
> Internal research/planning document. Not for GitHub.
## Vision
The **Orchestrator Loop** is a new autonomous execution mode that transforms high-level user goals into phased, verified, team-coordinated implementations. Unlike Ralph Loop (flat task queue → idle sessions), the Orchestrator manages the full lifecycle: **plan → approve → execute → verify → adapt → complete**.
```
USER: "Add OAuth2 login with Google/GitHub, role-based access control, and API key management"
ORCHESTRATOR:
Phase 1: Research & Setup ✅ (3m) — scaffold, deps, config
Phase 2: Auth Core ✅ (8m) — OAuth2 flow, session mgmt
Phase 3: Provider Integration 🔄 (12m) — Google + GitHub (parallel via team agents)
Phase 4: RBAC ⏳ — roles, permissions, middleware
Phase 5: API Keys ⏳ — generation, validation, rate limits
Phase 6: Testing & Review ⏳ — integration tests, security review
Progress: ━━━━━━━━━━━━━━━━━━━━ 40% | Agents: 3 active | Time: 23m
```
## Architecture
```
┌─────────────────────────────────────────────────────────────────┐
│ OrchestratorLoop │
│ │
│ ┌────────────────┐ ┌────────────────┐ ┌──────────────────┐ │
│ │ Orchestrator │ │ Orchestrator │ │ Orchestrator │ │
│ │ Planner │ │ Executor │ │ Verifier │ │
│ │ │ │ │ │ │ │
│ │ PlanOrchestrator│ │ TaskQueue │ │ AI review │ │
│ │ + phase grouper│ │ SessionManager │ │ Test commands │ │
│ │ + team strategy│ │ Team prompts │ │ File checks │ │
│ └───────┬────────┘ └───────┬────────┘ └─────────┬────────┘ │
│ │ │ │ │
│ └───────────────────┼──────────────────────┘ │
│ │ │
│ ┌─────────▼─────────┐ │
│ │ Existing Codeman │ │
│ │ Infrastructure │ │
│ │ │ │
│ │ SessionManager │ │
│ │ TaskQueue │ │
│ │ RespawnController │ │
│ │ TeamWatcher │ │
│ │ PlanOrchestrator │ │
│ │ StateStore │ │
│ │ Hooks + SSE │ │
│ └────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
```
## State Machine
```
┌─────────┐
│ IDLE │
└────┬────┘
│ start(goal)
▼
┌─────────┐
┌────────│PLANNING │────────┐
│ fail └────┬────┘ │
▼ │ plan ready │ user cancels
┌────────┐ ▼ ▼
│ FAILED │ ┌─────────┐ ┌────────┐
└────────┘ │APPROVAL │ │ IDLE │
▲ └────┬────┘ └────────┘
│ │ approve
│ ▼
│ ┌──────────┐
│ ┌───►│EXECUTING │◄────────────────────┐
│ │ └────┬─────┘ │
│ │ │ all tasks in phase done │
│ │ ▼ │
│ │ ┌──────────┐ │
│ │ │VERIFYING │ │
│ │ └────┬─────┘ │
│ │ pass │ │ fail │
│ │ ▼ ▼ │
│ │ more ┌──────────┐ │
│ │ phases?│REPLANNING│── retry ────────┘
│ │ │ └────┬─────┘
│ │ │ │ max retries
│ │ │ ▼
│ │ │ ┌────────┐
│ └────┘ │ FAILED │
│ next └────────┘
│ phase
│ │
│ ▼
│ ┌───────────┐
└─│ COMPLETED │
└───────────┘
```
**States:** `idle` | `planning` | `approval` | `executing` | `verifying` | `replanning` | `completed` | `failed` | `paused`
Transitions are event-driven. The state machine is the single source of truth — all methods check `this.state` before acting.
## Type Definitions
### `src/types/orchestrator.ts`
```typescript
// ═══════════════════════════════════════════════════════════════
// State Machine
// ═══════════════════════════════════════════════════════════════
export type OrchestratorState =
| 'idle'
| 'planning'
| 'approval'
| 'executing'
| 'verifying'
| 'replanning'
| 'completed'
| 'failed'
| 'paused';
// ═══════════════════════════════════════════════════════════════
// Plan Structure
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorPlan {
id: string;
goal: string;
createdAt: number;
phases: OrchestratorPhase[];
metadata: {
totalTasks: number;
estimatedComplexity: 'low' | 'medium' | 'high';
modelUsed: string;
planDurationMs: number;
};
}
export interface OrchestratorPhase {
id: string; // "phase-1", "phase-2"
name: string; // Human-readable name
description: string;
order: number;
status: PhaseStatus;
tasks: OrchestratorTask[];
verificationCriteria: string[];
testCommands: string[];
maxAttempts: number; // Default: 3
attempts: number; // Current attempt count
startedAt: number | null;
completedAt: number | null;
durationMs: number | null;
teamStrategy: TeamStrategy;
}
export type PhaseStatus =
| 'pending'
| 'executing'
| 'verifying'
| 'passed'
| 'failed'
| 'skipped';
export interface OrchestratorTask {
id: string; // "phase-1-task-1"
phaseId: string;
prompt: string; // Single-line prompt for Claude
status: 'pending' | 'running' | 'completed' | 'failed';
assignedSessionId: string | null;
queueTaskId: string | null; // Links to TaskQueue task
parallel: boolean; // Can run in parallel with sibling tasks
completionPhrase: string; // Unique phrase for completion detection
timeoutMs: number;
startedAt: number | null;
completedAt: number | null;
error: string | null;
retries: number;
}
// ═══════════════════════════════════════════════════════════════
// Team Strategy
// ═══════════════════════════════════════════════════════════════
export type TeamStrategy =
| { type: 'single' } // One session handles all
| { type: 'parallel'; maxSessions: number } // Multiple sessions
| { type: 'team'; config: TeamSetup } // Agent teams
export interface TeamSetup {
leadPrompt: string;
suggestedTeammates: string[]; // Role descriptions
maxTeammates: number;
}
// ═══════════════════════════════════════════════════════════════
// Verification
// ═══════════════════════════════════════════════════════════════
export interface VerificationResult {
passed: boolean;
checks: VerificationCheck[];
summary: string;
suggestions: string[]; // Recovery hints for replanning
}
export interface VerificationCheck {
type: 'test_command' | 'ai_review' | 'file_check';
description: string;
passed: boolean;
output?: string;
}
// ═══════════════════════════════════════════════════════════════
// Configuration
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorConfig {
plannerModel: string; // Default: 'opus'
researchEnabled: boolean; // Default: true
autoApprove: boolean; // Default: false
maxPhaseRetries: number; // Default: 3
phaseTimeoutMs: number; // Default: 1800000 (30min)
enableTeamAgents: boolean; // Default: true
maxParallelSessions: number; // Default: 3
verificationMode: 'strict' | 'moderate' | 'lenient';
compactBetweenPhases: boolean; // Default: true
}
// ═══════════════════════════════════════════════════════════════
// Persistence (saved to ~/.codeman/state.json)
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorPersistState {
state: OrchestratorState;
plan: OrchestratorPlan | null;
currentPhaseIndex: number;
startedAt: number | null;
completedAt: number | null;
config: OrchestratorConfig;
stats: OrchestratorStats;
}
export interface OrchestratorStats {
phasesCompleted: number;
phasesFailed: number;
totalTasksCompleted: number;
totalTasksFailed: number;
totalDurationMs: number;
replanCount: number;
}
```
## New Files (Implementation Order)
### Step 1: `src/types/orchestrator.ts` — Type definitions
All interfaces above. No dependencies. ~120 lines.
### Step 2: `src/orchestrator-planner.ts` — Plan generation + phase grouping
~300 lines. Wraps existing PlanOrchestrator.
```typescript
/**
* @fileoverview Orchestrator plan generation — converts goals into phased plans.
*
* Uses PlanOrchestrator for AI plan generation, then groups PlanItems into
* sequential phases with team strategies and verification criteria.
*
* @module orchestrator-planner
*/
export class OrchestratorPlanner {
constructor(mux: TerminalMultiplexer, workingDir: string, config: OrchestratorConfig);
/** Generate plan from goal. Uses PlanOrchestrator internally. */
async generatePlan(goal: string, onProgress?: ProgressCallback): Promise<OrchestratorPlan>;
/** Cancel in-progress plan generation. */
async cancel(): Promise<void>;
// Internal
private groupIntoPhases(items: PlanItem[], goal: string): OrchestratorPhase[];
private assignTeamStrategies(phases: OrchestratorPhase[]): void;
private generateCompletionPhrases(plan: OrchestratorPlan): void;
}
```
**Phase grouping algorithm:**
1. Topological sort by `PlanItem.dependencies`
2. Group into dependency layers (Kahn's algorithm)
3. Within each layer, sub-group by `tddPhase` (setup → test → impl → verify → review)
4. Merge adjacent small phases (< 2 tasks) if they share the same tddPhase
5. Assign team strategies:
- 1-2 tasks → `{ type: 'single' }`
- 3+ independent tasks → `{ type: 'parallel', maxSessions: Math.min(taskCount, config.maxParallelSessions) }`
- 4+ tasks with high complexity → `{ type: 'team', config: { ... } }`
6. Generate unique completion phrases per task: `ORCH_P{phaseOrder}_T{taskIndex}`
### Step 3: `src/orchestrator-verifier.ts` — Phase verification
~200 lines.
```typescript
/**
* @fileoverview Orchestrator phase verification.
*
* Runs verification checks after each phase completes:
* test commands, AI review, and file existence checks.
*
* @module orchestrator-verifier
*/
export class OrchestratorVerifier {
constructor(config: OrchestratorConfig);
/** Run all verification checks for a completed phase. */
async verifyPhase(
phase: OrchestratorPhase,
session: Session,
mode: 'strict' | 'moderate' | 'lenient'
): Promise<VerificationResult>;
// Verification strategies
private async runTestCommands(commands: string[], session: Session): Promise<VerificationCheck[]>;
private async aiReview(phase: OrchestratorPhase, session: Session): Promise<VerificationCheck>;
}
```
**Verification modes:**
- `strict`: ALL test commands must pass AND AI review must approve
- `moderate`: Test commands must pass, AI review is advisory
- `lenient`: At least one test command passes, AI review skipped
**AI review prompt (sent as a task to the session):**
```
Review Phase "{phase.name}" completion. Check:
1. Expected functionality works
2. No obvious regressions
3. Code quality is acceptable
Criteria: {phase.verificationCriteria.join('\n')}
If ALL criteria are met, respond: ORCH_VERIFY_PASS
If ANY criteria fail, respond: ORCH_VERIFY_FAIL and explain what failed.
```
### Step 4: `src/orchestrator-loop.ts` — Core state machine
~500 lines. Main orchestrator engine.
```typescript
/**
* @fileoverview Orchestrator Loop — phased plan execution with team agents.
*
* State machine that generates plans from user goals, executes them
* phase-by-phase with verification gates, and adapts on failure.
*
* @module orchestrator-loop
*/
export interface OrchestratorLoopEvents {
stateChanged: (state: OrchestratorState, prevState: OrchestratorState) => void;
planReady: (plan: OrchestratorPlan) => void;
phaseStarted: (phase: OrchestratorPhase) => void;
phaseCompleted: (phase: OrchestratorPhase) => void;
phaseFailed: (phase: OrchestratorPhase, reason: string) => void;
taskAssigned: (task: OrchestratorTask, sessionId: string) => void;
taskCompleted: (task: OrchestratorTask) => void;
taskFailed: (task: OrchestratorTask, error: string) => void;
verificationResult: (phase: OrchestratorPhase, result: VerificationResult) => void;
completed: (stats: OrchestratorStats) => void;
error: (error: Error) => void;
}
export class OrchestratorLoop extends EventEmitter {
private state: OrchestratorState = 'idle';
private plan: OrchestratorPlan | null = null;
private currentPhaseIndex = 0;
private config: OrchestratorConfig;
private planner: OrchestratorPlanner;
private verifier: OrchestratorVerifier;
private sessionManager: SessionManager;
private taskQueue: TaskQueue;
private store: StateStore;
private stats: OrchestratorStats;
private cleanup: CleanupManager;
private pausedState: OrchestratorState | null = null; // State before pause
// ── Lifecycle ──────────────────────────────────────────────
constructor(mux: TerminalMultiplexer, workingDir: string, config?: Partial<OrchestratorConfig>);
/** Start orchestration with a goal. Transitions: idle → planning */
async start(goal: string): Promise<void>;
/** Approve the generated plan. Transitions: approval → executing */
async approve(): Promise<void>;
/** Reject plan with feedback. Transitions: approval → planning (regenerate) */
async reject(feedback: string): Promise<void>;
/** Pause execution. Saves current state. */
pause(): void;
/** Resume from pause. */
resume(): void;
/** Stop everything and clean up. → idle */
async stop(): Promise<void>;
/** Skip current phase. → executing (next phase) or completed */
async skipPhase(phaseId: string): Promise<void>;
/** Retry a failed phase. → executing */
async retryPhase(phaseId: string): Promise<void>;
// ── Getters ────────────────────────────────────────────────
getState(): OrchestratorState;
getPlan(): OrchestratorPlan | null;
getCurrentPhase(): OrchestratorPhase | null;
getStats(): OrchestratorStats;
getStatus(): OrchestratorPersistState;
// ── Internal: Phase Execution ──────────────────────────────
private async executeCurrentPhase(): Promise<void>;
private async executePhase(phase: OrchestratorPhase): Promise<void>;
private async assignPhaseTasks(phase: OrchestratorPhase): Promise<void>;
private handleTaskCompleted(taskId: string): void;
private handleTaskFailed(taskId: string, error: string): void;
private async onPhaseTasksComplete(phase: OrchestratorPhase): Promise<void>;
// ── Internal: Verification ─────────────────────────────────
private async verifyCurrentPhase(): Promise<void>;
private async handleVerificationResult(phase: OrchestratorPhase, result: VerificationResult): Promise<void>;
// ── Internal: Replanning ───────────────────────────────────
private async replanPhase(phase: OrchestratorPhase, failures: string[]): Promise<void>;
// ── Internal: State Machine ────────────────────────────────
private setState(newState: OrchestratorState): void;
private advanceToNextPhase(): Promise<void>;
private persist(): void;
private restore(): void;
}
```
**Key execution flow in `executePhase()`:**
1. Mark phase as `executing`, emit `phaseStarted`
2. For each task in phase:
- Create a `CreateTaskOptions` from `OrchestratorTask`
- Add to `TaskQueue` with proper dependencies + completion phrase
- Store the TaskQueue task ID in `OrchestratorTask.queueTaskId`
3. Poll task completion (listen to TaskQueue events)
4. When all tasks complete → call `onPhaseTasksComplete()`
5. `onPhaseTasksComplete()` triggers verification
**How tasks get assigned to sessions:**
The OrchestratorLoop does NOT manage session assignment directly. It adds tasks to the existing TaskQueue and starts a mini poll loop that assigns pending tasks to idle sessions — the same pattern as RalphLoop's `assignTasks()`. This reuses existing session management.
**Team agent flow:**
For phases with `teamStrategy.type === 'team'`:
- Start a single session with `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
- Instead of adding individual tasks to TaskQueue, send ONE comprehensive prompt to the lead
- The prompt instructs the lead to create teammates and delegate
- Monitor via TeamWatcher for team task completion + hook events
- Phase completion is detected via the lead's completion phrase
### Step 5: `src/web/routes/orchestrator-routes.ts` — API endpoints
~300 lines.
```
POST /api/orchestrator/start — { goal, config? } → start planning
POST /api/orchestrator/approve — approve generated plan
POST /api/orchestrator/reject — { feedback } → reject + replan
POST /api/orchestrator/pause — pause execution
POST /api/orchestrator/resume — resume execution
POST /api/orchestrator/stop — stop orchestration
GET /api/orchestrator/status — full state + plan + stats
GET /api/orchestrator/plan — plan details only
POST /api/orchestrator/phase/:id/skip — skip a phase
POST /api/orchestrator/phase/:id/retry — retry a failed phase
```
Port dependency: `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort`
The route module receives the OrchestratorLoop instance via the InfraPort (added to `createRouteContext()`).
### Step 6: SSE Events — `src/web/sse-events.ts` additions
```typescript
// ─── Orchestrator ────────────────────────────────────────────────────────────
/** Orchestrator state machine transitioned. */
export const OrchestratorStateChanged = 'orchestrator:stateChanged' as const;
/** Orchestrator plan generated and ready for approval. */
export const OrchestratorPlanReady = 'orchestrator:planReady' as const;
/** Orchestrator phase started executing. */
export const OrchestratorPhaseStarted = 'orchestrator:phaseStarted' as const;
/** Orchestrator phase completed successfully. */
export const OrchestratorPhaseCompleted = 'orchestrator:phaseCompleted' as const;
/** Orchestrator phase failed. */
export const OrchestratorPhaseFailed = 'orchestrator:phaseFailed' as const;
/** Orchestrator verification result for a phase. */
export const OrchestratorVerification = 'orchestrator:verification' as const;
/** Orchestrator task assigned to session. */
export const OrchestratorTaskAssigned = 'orchestrator:taskAssigned' as const;
/** Orchestrator task completed. */
export const OrchestratorTaskCompleted = 'orchestrator:taskCompleted' as const;
/** Orchestrator task failed. */
export const OrchestratorTaskFailed = 'orchestrator:taskFailed' as const;
/** All phases completed successfully. */
export const OrchestratorCompleted = 'orchestrator:completed' as const;
/** Orchestrator error. */
export const OrchestratorError = 'orchestrator:error' as const;
```
11 new events. Add to `SseEvent` namespace object + mirror in `constants.js`.
### Step 7: State persistence — `src/state-store.ts` additions
Add to `AppState`:
```typescript
orchestrator?: OrchestratorPersistState;
```
Add methods:
```typescript
getOrchestratorState(): OrchestratorPersistState | null;
setOrchestratorState(state: Partial<OrchestratorPersistState>): void;
clearOrchestratorState(): void;
```
### Step 8: Server integration — `src/web/server.ts` modifications
1. Import `OrchestratorLoop` and `registerOrchestratorRoutes`
2. Add `private orchestratorLoop: OrchestratorLoop` field
3. Initialize in constructor (lazy — created on first start, not at boot)
4. Add to `createRouteContext()` InfraPort: `orchestratorLoop: this.orchestratorLoop`
5. Wire up OrchestratorLoop events → SSE broadcasts
6. Register routes: `registerOrchestratorRoutes(this.app, ctx)`
7. Clean up in `stop()`
### Step 9: `src/web/public/orchestrator-ui.js` — Frontend panel
~500 lines. New frontend module.
**Load order**: After `panels-ui.js` (11), before `ralph-wizard.js` (13). So load order = 11.5.
**UI elements:**
- Goal input form (text area + config toggles)
- Plan approval view (phase list, task details, approve/reject buttons)
- Execution dashboard (progress bar, phase cards, task status indicators)
- Agent activity panel (session count, team status)
- Controls (pause, resume, stop, skip phase, retry phase)
**SSE listeners:**
- All 11 orchestrator events → update UI state
- Reuses existing session/respawn/team event handlers for agent monitoring
### Step 10: `src/prompts/orchestrator.ts` — Prompt templates
~200 lines.
Templates for:
- Phase execution prompt (tells Claude what to do in this phase)
- Team lead delegation prompt (instructs lead to create and coordinate teammates)
- Verification prompt (asks Claude to verify phase output)
- Replan prompt (gives failure context, asks for recovery steps)
### Step 11: Constants, schemas, route barrel updates
- `src/web/public/constants.js` — Add 11 SSE event mirrors
- `src/web/schemas.ts` — Add Zod schemas for orchestrator API input validation
- `src/web/routes/index.ts` — Export `registerOrchestratorRoutes`
- `src/web/ports/infra-port.ts` — Add `orchestratorLoop` to InfraPort
- `src/types/index.ts` — Export orchestrator types
## Existing File Modifications Summary
| File | Change | Lines |
|------|--------|-------|
| `src/types/index.ts` | Add orchestrator barrel export | +1 |
| `src/web/sse-events.ts` | Add 11 orchestrator events + SseEvent entries | +30 |
| `src/web/public/constants.js` | Mirror 11 SSE events | +15 |
| `src/web/routes/index.ts` | Export registerOrchestratorRoutes | +1 |
| `src/web/ports/infra-port.ts` | Add orchestratorLoop to InfraPort | +3 |
| `src/web/server.ts` | Initialize OrchestratorLoop, wire events, register routes | +40 |
| `src/web/schemas.ts` | Add orchestrator Zod schemas | +20 |
| `src/state-store.ts` | Add orchestrator state persistence | +20 |
| `src/web/public/app.js` | Add orchestrator SSE listeners + panel toggle | +30 |
| `src/web/public/index.html` | Add orchestrator-ui.js script tag | +1 |
**Total new code**: ~2,300 lines across 6 new files
**Total modifications**: ~160 lines across 10 existing files
## Implementation Execution Order
This is the actual build order — each step is a commit checkpoint:
1. **Types** — `src/types/orchestrator.ts` + barrel export. Zero risk, pure types.
2. **SSE events** — Add all 11 events to both `sse-events.ts` and `constants.js`. Wire in SseEvent namespace.
3. **State persistence** — Add orchestrator state to StateStore. Small, isolated change.
4. **Schemas** — Add Zod validation schemas for API input.
5. **Planner** — `src/orchestrator-planner.ts`. Can test in isolation.
6. **Verifier** — `src/orchestrator-verifier.ts`. Can test in isolation.
7. **Core loop** — `src/orchestrator-loop.ts`. The big one. Depends on planner + verifier.
8. **Prompts** — `src/prompts/orchestrator.ts`. Templates used by core loop.
9. **Port + routes** — `src/web/ports/infra-port.ts` update + `src/web/routes/orchestrator-routes.ts`.
10. **Server integration** — Wire OrchestratorLoop into WebServer. Routes become live.
11. **Frontend** — `src/web/public/orchestrator-ui.js` + app.js listeners + index.html script tag.
12. **Tests** — `test/orchestrator-*.test.ts`.
13. **Typecheck + lint** — Fix all issues, ensure CI passes.
## Edge Cases & Error Handling
- **Session limit reached**: Queue tasks and wait for sessions to free up (existing SessionManager handles this)
- **All sessions crash during phase**: Mark phase as failed, attempt replan
- **Verification flaky**: `moderate` mode allows test retries; `lenient` skips AI review
- **Plan too large**: Cap at 10 phases, 50 total tasks. Warn user.
- **Context overflow**: Auto-compact between phases. Respawn if needed (orchestrator state is external).
- **User pauses mid-phase**: Pause task assignment, don't cancel running tasks. Resume picks up where it left off.
- **Network/API errors during planning**: Retry plan generation up to 2 times, then fail with clear message.
- **Orchestrator vs Ralph conflict**: Mutually exclusive. Starting orchestrator stops Ralph if running. Starting Ralph stops orchestrator.
## Testing Strategy
- **Unit tests**: `test/orchestrator-planner.test.ts` — phase grouping algorithm, team strategy assignment
- **Unit tests**: `test/orchestrator-verifier.test.ts` — verification logic with mocked sessions
- **Integration tests**: `test/orchestrator-loop.test.ts` — state machine transitions, task lifecycle
- **Route tests**: `test/routes/orchestrator-routes.test.ts` — API validation, status responses
All tests use `MockSession` pattern from existing test infrastructure. No real tmux needed.
+157
View File
@@ -0,0 +1,157 @@
# Orchestrator Loop — Research Findings
> Research doc for the new "Orchestrator Loop" feature. Not for GitHub.
## What We're Building
A new autonomous loop variant — **Orchestrator Loop** — that takes high-level user tasks, decomposes them into a detailed plan using team agents, and executes the plan step-by-step with quality gates. Unlike Ralph Loop (which executes a flat task queue), the Orchestrator coordinates **planning, delegation, and verification** as a continuous cycle.
**Core idea**: User inputs a goal → Orchestrator creates a detailed plan → spins up team agents for parallel execution → validates each step → adapts the plan based on results → delivers polished output.
## Existing Infrastructure Analysis
### What We Can Reuse
#### 1. Ralph Loop (`src/ralph-loop.ts`)
- **Pattern**: Poll loop with `start() → tick() → stop()` lifecycle
- **Reusable**: Event-driven task assignment, session completion handling, timeout management
- **Limitation**: Flat task queue — no concept of phases, dependencies between task groups, or adaptive replanning
- **Key insight**: `assignTaskToSession()` uses `session.sendInput(task.prompt)` — simple prompt injection into PTY
#### 2. Task Queue (`src/task-queue.ts`) + Task (`src/task.ts`)
- **Already has**: Priority ordering, dependency tracking between tasks, completion phrase detection
- **Limitation**: No task *groups* or *phases*. Dependencies are task-to-task, not phase-to-phase
- **Key insight**: Tasks support `completionPhrase` — a string the task watches for in output. This is how Ralph knows a task is done
#### 3. Plan Orchestrator (`src/plan-orchestrator.ts`)
- **Already has**: 2-agent plan generation (Research Agent → Planner Agent), TDD-aware plan items with P0/P1/P2 priorities
- **Output**: `PlanItem[]` with dependencies, verification criteria, TDD phases, complexity ratings
- **Limitation**: Plan generation only — no execution. Plans are generated then sit in state/UI for human review
- **Key insight**: Uses `Session` directly to run Claude subagent instances for research and planning. Returns structured JSON
#### 4. Team Agents (`src/team-watcher.ts`, `~/.claude/teams/`)
- **Already has**: Team creation, member tracking, filesystem inbox messaging, task management via `~/.claude/tasks/{team-name}/`
- **Limitation**: Codeman can only *observe* teams (TeamWatcher is read-only polling), not *create* or *orchestrate* them
- **Key insight**: Teams are a Claude Code feature. Codeman monitors them but doesn't control them. We can't programmatically create teammates — Claude Code does that when you use `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
#### 5. Respawn Controller (`src/respawn-controller.ts`)
- **Already has**: Preset-based automation (ralph-todo, overnight-autonomous), circuit breaker, health scoring
- **Key insight**: The `ralph-todo` preset (8s idle, 480min max) is designed for autonomous task execution. We'd need a new preset or make Orchestrator Loop set its own timing
#### 6. Session Auto-Ops (`src/session-auto-ops.ts`)
- **Already has**: Auto-compact at token thresholds, auto-clear for context management
- **Key insight**: Critical for long Orchestrator runs — prevents context overflow during multi-step execution
#### 7. Hooks (`src/hooks-config.ts`)
- **Already has**: `idle_prompt`, `stop`, `teammate_idle`, `task_completed` hook events
- **Key insight**: Hooks fire POST to `/api/hook-event` — this is how Codeman knows when Claude is idle, stopped, or completed a task. The Orchestrator Loop can listen to these same events
### What We Need to Build New
1. **Plan → Task decomposition**: Convert PlanOrchestrator output (PlanItem[]) into executable task groups with phase ordering
2. **Multi-phase execution engine**: Execute plan phases sequentially, tasks within phases in parallel
3. **Verification gates**: After each phase, run verification (test commands, AI review) before proceeding
4. **Adaptive replanning**: When a task fails or verification fails, generate a recovery plan
5. **Team agent orchestration**: Leverage Claude Code's agent teams for parallel execution within phases
6. **Progress tracking & UI**: Real-time dashboard showing plan progress, phase status, agent activity
## How Teams Actually Work (Important Constraint)
After deep research, here's the reality of agent teams:
```
User starts session with CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
→ Claude Code creates a team-lead
→ Team-lead spawns teammates (in-process threads)
→ Teammates appear as subagents (detected by SubagentWatcher)
→ Communication via ~/.claude/teams/{name}/inboxes/{member}.json
→ Tasks tracked in ~/.claude/tasks/{team-name}/{N}.json
```
**Codeman cannot programmatically create team members.** This is a Claude Code internal feature. However, Codeman CAN:
- Start a session that has teams enabled
- Send a prompt to the lead that instructs it to use agent teams
- Monitor team activity via TeamWatcher
- React to teammate_idle and task_completed hook events
- Read team task status from the filesystem
**This means**: The Orchestrator Loop orchestrates at the *session prompt* level, not the *team member* level. We tell the lead what to do, and the lead decides how to use its team.
## Architecture Decision: Prompt-Level Orchestration
Given the team constraint, the Orchestrator Loop works by:
1. **Planning phase**: Use PlanOrchestrator to generate a detailed plan from user input
2. **Execution phase**: Feed plan steps as prompts to sessions, one phase at a time
3. **Verification phase**: After each phase, run verification prompts and check results
4. **Adaptation phase**: If verification fails, generate recovery prompts
The "team agents" aspect works by:
- Starting sessions with `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
- Crafting prompts that *instruct the lead to delegate* to teammates
- Monitoring team activity to track parallel progress
- The lead agent is smart enough to decompose work across its team
## Key Technical Findings
### Session Input Mechanics
```typescript
// From session.ts - how we send prompts
await session.sendInput(task.prompt); // Uses writeViaMux() internally
// writeViaMux() does: tmux send-keys -l "prompt text" + tmux send-keys Enter
// CRITICAL: Single-line only! Multi-line breaks Ink rendering
```
### Completion Detection Chain
```
PTY output → RalphTracker.processData() → completion phrase fuzzy match
→ CompletionConfidence scoring (multi-signal: promise tag + todos + exit signal)
→ If confident → emit 'completionDetected'
→ RalphLoop listens → marks task complete → assigns next
```
### How Plan Items Map to Tasks
```typescript
// PlanItem has:
interface PlanItem {
id: string; // "P0-001"
content: string; // "Implement error handling for API endpoints"
priority: 'P0' | 'P1' | 'P2';
dependencies: string[]; // ["P0-000"] — other PlanItem IDs
verificationCriteria: string;
testCommand: string;
tddPhase: 'setup' | 'test' | 'impl' | 'verify' | 'review';
complexity: 'low' | 'medium' | 'high';
}
// Task has:
interface CreateTaskOptions {
prompt: string;
priority: number;
dependencies: string[]; // Task IDs
completionPhrase: string;
timeoutMs: number;
}
// Natural mapping: PlanItem.content → Task.prompt
// PlanItem.dependencies → Task.dependencies
// PlanItem.priority → Task.priority (P0=100, P1=50, P2=10)
// PlanItem.verificationCriteria → verification task prompt
```
### Context Management for Long Runs
- Auto-compact at ~110k tokens (configurable)
- Auto-clear at ~140k tokens (configurable)
- Respawn cycling: kill + restart session to reset context entirely
- For Orchestrator: we want compact between phases, respawn between major milestones
## Risk Assessment
| Risk | Severity | Mitigation |
|------|----------|------------|
| Context overflow during complex phases | High | Auto-compact between tasks, respawn between phases |
| Team agents not predictable | Medium | Orchestrate at session level, let Claude decide team delegation |
| Plan too ambitious → infinite loop | High | Phase budgets (max attempts per phase), circuit breaker |
| Verification too strict → blocks progress | Medium | Configurable strictness, human override via UI |
| Single-line prompt limit | Medium | Use CLAUDE.md file for complex instructions, prompt references file |
| Long planning phase delays execution | Low | Show plan for approval before execution |
+173 -27
View File
@@ -7,8 +7,10 @@
# Environment variables:
# CODEMAN_NONINTERACTIVE=1 - Skip all prompts (for CI/automation)
# CODEMAN_INSTALL_DIR - Custom install directory (default: ~/.codeman/app)
# CODEMAN_SKIP_SYSTEMD=1 - Skip systemd service setup prompt
# CODEMAN_SKIP_SYSTEMD=1 - Skip systemd/launchd service setup prompt
# CODEMAN_NODE_VERSION - Node.js major version to install (default: 22)
# CODEMAN_REPO_URL - Custom git repository URL (default: upstream Codeman)
# CODEMAN_BRANCH - Git branch to install (default: master)
set -euo pipefail
@@ -17,7 +19,8 @@ set -euo pipefail
# ============================================================================
INSTALL_DIR="${CODEMAN_INSTALL_DIR:-$HOME/.codeman/app}"
REPO_URL="https://github.com/Ark0N/Codeman.git"
REPO_URL="${CODEMAN_REPO_URL:-https://github.com/Ark0N/Codeman.git}"
BRANCH="${CODEMAN_BRANCH:-master}"
MIN_NODE_VERSION=18
TARGET_NODE_VERSION="${CODEMAN_NODE_VERSION:-22}"
NONINTERACTIVE="${CODEMAN_NONINTERACTIVE:-0}"
@@ -350,8 +353,15 @@ ensure_sudo() {
die "sudo is required but not installed. Please install packages manually or run as root."
fi
# Validate sudo access
if ! sudo -v 2>/dev/null; then
die "Failed to obtain sudo privileges."
# When piped (curl | bash), stdin is the pipe — redirect from /dev/tty so sudo can prompt
if [[ -e /dev/tty ]]; then
if ! sudo -v 2>/dev/null < /dev/tty; then
die "Failed to obtain sudo privileges."
fi
else
if ! sudo -v 2>/dev/null; then
die "Failed to obtain sudo privileges. Try running the script directly instead of piping."
fi
fi
}
@@ -369,7 +379,12 @@ ensure_homebrew() {
fi
info "Installing Homebrew first..."
/bin/bash -c "$(download_to_stdout https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"
# When piped (curl | bash), stdin is the pipe — Homebrew needs TTY for sudo password prompt
if [[ -e /dev/tty ]]; then
/bin/bash -c "$(download_to_stdout https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)" < /dev/tty
else
NONINTERACTIVE=1 /bin/bash -c "$(download_to_stdout https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"
fi
# Add Homebrew to PATH for Apple Silicon
if [[ -f /opt/homebrew/bin/brew ]]; then
@@ -784,9 +799,84 @@ setup_sc_alias() {
}
# ============================================================================
# Systemd Service Setup (Linux only)
# Service Setup (Linux systemd / macOS launchd)
# ============================================================================
setup_launchd_service() {
local plist_label="com.codeman.web"
local agent_dir="$HOME/Library/LaunchAgents"
local agent_plist="$agent_dir/$plist_label.plist"
local daemon_plist="/Library/LaunchDaemons/$plist_label.plist"
info "Setting up macOS LaunchAgent..."
# Remove any existing LaunchDaemon (system-level) to prevent duplicates.
# We standardize on LaunchAgent (user-level) — it doesn't require sudo,
# inherits the user's environment, and is the correct choice for user apps.
if [[ -f "$daemon_plist" ]]; then
warn "Found system-level LaunchDaemon at $daemon_plist — removing to prevent duplicate"
sudo launchctl unload "$daemon_plist" 2>/dev/null || true
sudo rm -f "$daemon_plist"
success "Removed duplicate LaunchDaemon"
fi
# Unload existing agent before overwriting
if [[ -f "$agent_plist" ]]; then
launchctl unload "$agent_plist" 2>/dev/null || true
fi
mkdir -p "$agent_dir"
# Build PATH: ensure /opt/homebrew/bin (Apple Silicon) and ~/.local/bin are included
local svc_path="/opt/homebrew/bin:/usr/local/bin:$HOME/.local/bin:/usr/bin:/bin:/usr/sbin:/sbin"
# Find node binary path
local node_path
node_path=$(command -v node)
cat > "$agent_plist" << EOF
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>$plist_label</string>
<key>ProgramArguments</key>
<array>
<string>$node_path</string>
<string>$INSTALL_DIR/dist/index.js</string>
<string>web</string>
</array>
<key>EnvironmentVariables</key>
<dict>
<key>PATH</key>
<string>$svc_path</string>
<key>HOME</key>
<string>$HOME</string>
<key>LANG</key>
<string>en_US.UTF-8</string>
</dict>
<key>WorkingDirectory</key>
<string>$HOME</string>
<key>RunAtLoad</key>
<true/>
<key>KeepAlive</key>
<true/>
<key>ThrottleInterval</key>
<integer>10</integer>
<key>StandardOutPath</key>
<string>/tmp/codeman.log</string>
<key>StandardErrorPath</key>
<string>/tmp/codeman.log</string>
</dict>
</plist>
EOF
launchctl load "$agent_plist" 2>/dev/null || true
success "LaunchAgent installed and started"
}
setup_systemd_service() {
local service_dir="$HOME/.config/systemd/user"
local service_file="$service_dir/codeman-web.service"
@@ -1062,26 +1152,27 @@ main() {
if [[ -d "$INSTALL_DIR/.git" ]]; then
info "Existing installation found, updating..."
cd "$INSTALL_DIR"
git remote set-url origin "$REPO_URL" 2>/dev/null || true
# Check for local changes
if ! git diff --quiet 2>/dev/null || ! git diff --staged --quiet 2>/dev/null; then
warn "Local changes detected in $INSTALL_DIR"
if prompt_yes_no "Discard local changes and update?" "n"; then
git fetch --quiet origin
git reset --hard origin/master --quiet
git reset --hard "origin/$BRANCH" --quiet
else
info "Keeping existing installation, skipping update"
fi
else
git fetch --quiet origin
git reset --hard origin/master --quiet
git reset --hard "origin/$BRANCH" --quiet
fi
else
# Create parent directory
mkdir -p "$(dirname "$INSTALL_DIR")"
# Clone repository (shallow for speed)
git clone --quiet --depth 1 "$REPO_URL" "$INSTALL_DIR"
git clone --quiet --depth 1 --branch "$BRANCH" "$REPO_URL" "$INSTALL_DIR"
cd "$INSTALL_DIR"
fi
@@ -1135,17 +1226,25 @@ main() {
echo ""
local launch_choice=""
local has_systemd=false
local has_service=false
local service_type=""
if [[ "$os" == "linux" ]] && [[ "$SKIP_SYSTEMD" != "1" ]] && command -v systemctl &>/dev/null; then
has_systemd=true
has_service=true
service_type="systemd"
elif [[ "$os" == "macos" ]] && [[ "$SKIP_SYSTEMD" != "1" ]]; then
has_service=true
service_type="launchd"
fi
if [[ "$has_systemd" == "true" ]]; then
if [[ "$has_service" == "true" ]]; then
local service_label="systemd service"
[[ "$service_type" == "launchd" ]] && service_label="LaunchAgent"
echo -e " ${BOLD}How would you like to run Codeman?${NC}"
echo ""
echo -e " ${CYAN}1)${NC} Run now in this terminal"
echo -e " ${CYAN}2)${NC} Install as systemd service (auto-start on boot)"
echo -e " ${CYAN}2)${NC} Install as $service_label (auto-start on boot)"
echo -e " ${CYAN}3)${NC} Don't start — I'll run it later"
echo ""
@@ -1162,7 +1261,7 @@ main() {
done
fi
else
# macOS or no systemd — only offer run now or skip
# No service manager available — only offer run now or skip
echo -e " ${BOLD}Would you like to start Codeman now?${NC}"
echo ""
echo -e " ${CYAN}1)${NC} Run now in this terminal"
@@ -1188,12 +1287,16 @@ main() {
echo ""
# Handle systemd setup
# Handle service setup
if [[ "$launch_choice" == "2" ]]; then
setup_systemd_service
if [[ "$service_type" == "launchd" ]]; then
setup_launchd_service
else
setup_systemd_service
fi
# Offer tunnel service if cloudflared is available
if check_cloudflared && [[ -f "$INSTALL_DIR/scripts/codeman-tunnel.service" ]]; then
# Offer tunnel service if cloudflared is available (Linux only — systemd tunnel service)
if [[ "$service_type" == "systemd" ]] && check_cloudflared && [[ -f "$INSTALL_DIR/scripts/codeman-tunnel.service" ]]; then
echo ""
if prompt_yes_no "Also set up Cloudflare tunnel service? (requires CODEMAN_PASSWORD)" "n"; then
setup_tunnel_service
@@ -1208,10 +1311,16 @@ main() {
echo ""
echo -e " ${BOLD}Manage the service:${NC}"
echo ""
echo -e " ${CYAN}systemctl --user stop codeman-web${NC} # Stop"
echo -e " ${CYAN}systemctl --user restart codeman-web${NC} # Restart"
echo -e " ${CYAN}systemctl --user status codeman-web${NC} # Check status"
echo -e " ${CYAN}journalctl --user -u codeman-web -f${NC} # View logs"
if [[ "$service_type" == "launchd" ]]; then
echo -e " ${CYAN}launchctl unload ~/Library/LaunchAgents/com.codeman.web.plist${NC} # Stop"
echo -e " ${CYAN}launchctl load ~/Library/LaunchAgents/com.codeman.web.plist${NC} # Start"
echo -e " ${CYAN}tail -f /tmp/codeman.log${NC} # View logs"
else
echo -e " ${CYAN}systemctl --user stop codeman-web${NC} # Stop"
echo -e " ${CYAN}systemctl --user restart codeman-web${NC} # Restart"
echo -e " ${CYAN}systemctl --user status codeman-web${NC} # Check status"
echo -e " ${CYAN}journalctl --user -u codeman-web -f${NC} # View logs"
fi
echo ""
fi
@@ -1277,13 +1386,29 @@ update() {
info "Updating Codeman..."
cd "$INSTALL_DIR"
git remote set-url origin "$REPO_URL" 2>/dev/null || true
git fetch --quiet origin
git reset --hard origin/master --quiet
git reset --hard "origin/$BRANCH" --quiet
npm install --quiet --no-fund --no-audit 2>/dev/null || npm install --no-fund --no-audit
npm run build --quiet 2>/dev/null || npm run build
success "Updated to $(node -e "console.log(require('./package.json').version)")"
echo ""
echo -e " ${DIM}Restart codeman web to use the new version.${NC}"
# Auto-restart service if running, otherwise tell the user
local agent_plist="$HOME/Library/LaunchAgents/com.codeman.web.plist"
if systemctl --user is-active codeman-web.service &>/dev/null 2>&1; then
info "Restarting codeman-web service..."
systemctl --user restart codeman-web.service
success "codeman-web service restarted"
elif [[ -f "$agent_plist" ]]; then
info "Restarting LaunchAgent..."
launchctl unload "$agent_plist" 2>/dev/null || true
launchctl load "$agent_plist" 2>/dev/null || true
success "LaunchAgent restarted"
else
echo -e " ${DIM}Restart codeman web to use the new version:${NC}"
echo -e " ${CYAN}pkill -f 'codeman.*web'; codeman web &${NC}"
fi
echo ""
}
@@ -1292,9 +1417,9 @@ uninstall() {
info "Uninstalling Codeman..."
echo ""
# Stop and remove systemd services
# Stop and remove systemd services (Linux)
for svc in codeman-web codeman-tunnel; do
if systemctl --user is-active "${svc}.service" &>/dev/null; then
if systemctl --user is-active "${svc}.service" &>/dev/null 2>&1; then
info "Stopping ${svc} service..."
systemctl --user stop "${svc}.service"
fi
@@ -1310,6 +1435,20 @@ uninstall() {
done
systemctl --user daemon-reload 2>/dev/null || true
# Stop and remove launchd services (macOS)
local agent_plist="$HOME/Library/LaunchAgents/com.codeman.web.plist"
local daemon_plist="/Library/LaunchDaemons/com.codeman.web.plist"
if [[ -f "$agent_plist" ]]; then
launchctl unload "$agent_plist" 2>/dev/null || true
rm -f "$agent_plist"
success "Removed LaunchAgent"
fi
if [[ -f "$daemon_plist" ]]; then
sudo launchctl unload "$daemon_plist" 2>/dev/null || true
sudo rm -f "$daemon_plist"
success "Removed LaunchDaemon"
fi
# Remove symlinks
local symlink_dir="$HOME/.local/bin"
if [[ -L "$symlink_dir/codeman" ]]; then
@@ -1355,5 +1494,12 @@ uninstall() {
case "${1:-}" in
update) update ;;
uninstall) uninstall ;;
*) main "$@" ;;
*)
if [[ -z "${1:-}" && -d "$INSTALL_DIR/.git" ]]; then
print_banner
update
else
main "$@"
fi
;;
esac
+16
View File
@@ -0,0 +1,16 @@
{
"$schema": "https://unpkg.com/knip@5/schema.json",
"entry": [
"scripts/*.mjs",
"scripts/*.js",
"scripts/watch-subagents.ts",
"scripts/remotion/Root.tsx",
"scripts/remotion/index.ts",
"test/**/*.test.ts",
"test/mobile/vitest.config.ts",
"test/**/*.mjs"
],
"project": ["src/**/*.{ts,tsx}", "scripts/**/*.{ts,tsx,mjs,js}", "test/**/*.{ts,mjs}"],
"ignoreExportsUsedInFile": true,
"ignoreDependencies": ["@remotion/cli", "@remotion/transitions", "esbuild", "agent-browser"]
}
+103 -25
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "0.3.8",
"version": "0.6.8",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "0.3.8",
"version": "0.6.8",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
@@ -17,6 +17,7 @@
"@fastify/compress": "^8.3.1",
"@fastify/cookie": "^11.0.2",
"@fastify/static": "^8.0.0",
"@fastify/websocket": "^11.2.0",
"@xterm/addon-fit": "^0.11.0",
"@xterm/addon-unicode11": "^0.9.0",
"@xterm/addon-webgl": "^0.19.0",
@@ -45,6 +46,7 @@
"@types/react": "^19.2.14",
"@types/uuid": "^10.0.0",
"@types/web-push": "^3.6.4",
"@types/ws": "^8.18.1",
"@vitest/coverage-v8": "^4.0.18",
"agent-browser": "^0.6.0",
"esbuild": "^0.27.3",
@@ -945,6 +947,53 @@
"glob": "^11.0.0"
}
},
"node_modules/@fastify/websocket": {
"version": "11.2.0",
"resolved": "https://registry.npmjs.org/@fastify/websocket/-/websocket-11.2.0.tgz",
"integrity": "sha512-3HrDPbAG1CzUCqnslgJxppvzaAZffieOVbLp1DAy1huCSynUWPifSvfdEDUR8HlJLp3sp1A36uOM2tJogADS8w==",
"funding": [
{
"type": "github",
"url": "https://github.com/sponsors/fastify"
},
{
"type": "opencollective",
"url": "https://opencollective.com/fastify"
}
],
"license": "MIT",
"dependencies": {
"duplexify": "^4.1.3",
"fastify-plugin": "^5.0.0",
"ws": "^8.16.0"
}
},
"node_modules/@fastify/websocket/node_modules/duplexify": {
"version": "4.1.3",
"resolved": "https://registry.npmjs.org/duplexify/-/duplexify-4.1.3.tgz",
"integrity": "sha512-M3BmBhwJRZsSx38lZyhE53Csddgzl5R7xGJNk7CVddZD6CcmwMCH8J+7AprIrQKH7TonKxaCjcv27Qmf+sQ+oA==",
"license": "MIT",
"dependencies": {
"end-of-stream": "^1.4.1",
"inherits": "^2.0.3",
"readable-stream": "^3.1.1",
"stream-shift": "^1.0.2"
}
},
"node_modules/@fastify/websocket/node_modules/readable-stream": {
"version": "3.6.2",
"resolved": "https://registry.npmjs.org/readable-stream/-/readable-stream-3.6.2.tgz",
"integrity": "sha512-9u/sniCrY3D5WdsERHzHE4G2YCXqoG5FTHUiCC4SIbr6XcLZBY05ya9EKjYek9O5xOAwjGq+1JdGBAS7Q9ScoA==",
"license": "MIT",
"dependencies": {
"inherits": "^2.0.3",
"string_decoder": "^1.1.1",
"util-deprecate": "^1.0.1"
},
"engines": {
"node": ">= 6"
}
},
"node_modules/@humanfs/core": {
"version": "0.19.1",
"dev": true,
@@ -2118,6 +2167,16 @@
"@types/node": "*"
}
},
"node_modules/@types/ws": {
"version": "8.18.1",
"resolved": "https://registry.npmjs.org/@types/ws/-/ws-8.18.1.tgz",
"integrity": "sha512-ThVF6DCVhA8kUGy+aazFQ4kXQ7E1Ty7A3ypFOe0IcJV8O/M511G99AW24irKrW56Wt44yG9+ij8FaqoBGkuBXg==",
"dev": true,
"license": "MIT",
"dependencies": {
"@types/node": "*"
}
},
"node_modules/@types/yauzl": {
"version": "2.10.3",
"dev": true,
@@ -3019,7 +3078,9 @@
}
},
"node_modules/basic-ftp": {
"version": "5.1.0",
"version": "5.2.0",
"resolved": "https://registry.npmjs.org/basic-ftp/-/basic-ftp-5.2.0.tgz",
"integrity": "sha512-VoMINM2rqJwJgfdHq6RiUudKt2BV+FY5ZFezP/ypmwayk68+NzzAQy4XXLlqsGD4MCzq3DrmNFD/uUmBJuGoXw==",
"dev": true,
"license": "MIT",
"engines": {
@@ -4416,7 +4477,9 @@
"license": "BSD-3-Clause"
},
"node_modules/fastify": {
"version": "5.7.4",
"version": "5.8.2",
"resolved": "https://registry.npmjs.org/fastify/-/fastify-5.8.2.tgz",
"integrity": "sha512-lZmt3navvZG915IE+f7/TIVamxIwmBd+OMB+O9WBzcpIwOo6F0LTh0sluoMFk5VkrKTvvrwIaoJPkir4Z+jtAg==",
"funding": [
{
"type": "github",
@@ -4438,7 +4501,7 @@
"fast-json-stringify": "^6.0.0",
"find-my-way": "^9.0.0",
"light-my-request": "^6.0.0",
"pino": "^10.1.0",
"pino": "^9.14.0 || ^10.1.0",
"process-warning": "^5.0.0",
"rfdc": "^1.3.1",
"secure-json-parse": "^4.0.0",
@@ -4601,6 +4664,20 @@
"dev": true,
"license": "Unlicense"
},
"node_modules/fsevents": {
"version": "2.3.3",
"resolved": "https://registry.npmjs.org/fsevents/-/fsevents-2.3.3.tgz",
"integrity": "sha512-5xoDfX+fL7faATnagmWPpbFtwh/R77WmMMqqHGS65C3vvB0YHrgF+B1YmZ3441tMj5n63k0212XNoJwzlhffQw==",
"hasInstallScript": true,
"license": "MIT",
"optional": true,
"os": [
"darwin"
],
"engines": {
"node": "^8.16.0 || ^10.6.0 || >=11.0.0"
}
},
"node_modules/function-bind": {
"version": "1.1.2",
"dev": true,
@@ -5641,7 +5718,9 @@
"license": "ISC"
},
"node_modules/minimatch": {
"version": "10.2.2",
"version": "10.2.4",
"resolved": "https://registry.npmjs.org/minimatch/-/minimatch-10.2.4.tgz",
"integrity": "sha512-oRjTw/97aTBN0RHbYCdtF1MQfvusSIBQM0IZEgzl6426+8jSC0nF1a/GmnVLpfB9yyr6g6FTqWqiZVbxrtaCIg==",
"license": "BlueOak-1.0.0",
"dependencies": {
"brace-expansion": "^5.0.2"
@@ -6163,6 +6242,21 @@
"node": ">=18"
}
},
"node_modules/playwright/node_modules/fsevents": {
"version": "2.3.2",
"resolved": "https://registry.npmjs.org/fsevents/-/fsevents-2.3.2.tgz",
"integrity": "sha512-xiqMQR4xAeHTuB9uWm+fFRcIOgKBMiOBP+eXiyT7jsgVCq1bkVygt00oASowB7EdtpOHaaPgKt812P9ab+DDKA==",
"dev": true,
"hasInstallScript": true,
"license": "MIT",
"optional": true,
"os": [
"darwin"
],
"engines": {
"node": "^8.16.0 || ^10.6.0 || >=11.0.0"
}
},
"node_modules/pngjs": {
"version": "7.0.0",
"dev": true,
@@ -6625,14 +6719,6 @@
"version": "4.0.4",
"license": "MIT"
},
"node_modules/randombytes": {
"version": "2.1.0",
"dev": true,
"license": "MIT",
"dependencies": {
"safe-buffer": "^5.1.0"
}
},
"node_modules/react": {
"version": "19.2.4",
"dev": true,
@@ -7023,14 +7109,6 @@
"node": ">=10"
}
},
"node_modules/serialize-javascript": {
"version": "6.0.2",
"dev": true,
"license": "BSD-3-Clause",
"dependencies": {
"randombytes": "^2.1.0"
}
},
"node_modules/set-blocking": {
"version": "2.0.0",
"license": "ISC"
@@ -7401,14 +7479,15 @@
}
},
"node_modules/terser-webpack-plugin": {
"version": "5.3.16",
"version": "5.4.0",
"resolved": "https://registry.npmjs.org/terser-webpack-plugin/-/terser-webpack-plugin-5.4.0.tgz",
"integrity": "sha512-Bn5vxm48flOIfkdl5CaD2+1CiUVbonWQ3KQPyP7/EuIl9Gbzq/gQFOzaMFUEgVjB1396tcK0SG8XcNJ/2kDH8g==",
"dev": true,
"license": "MIT",
"dependencies": {
"@jridgewell/trace-mapping": "^0.3.25",
"jest-worker": "^27.4.5",
"schema-utils": "^4.3.0",
"serialize-javascript": "^6.0.2",
"terser": "^5.31.1"
},
"engines": {
@@ -8488,7 +8567,6 @@
},
"node_modules/ws": {
"version": "8.19.0",
"dev": true,
"license": "MIT",
"engines": {
"node": ">=10.0.0"
+12 -8
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "0.3.10",
"version": "0.6.8",
"description": "The missing control plane for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -11,21 +11,23 @@
"scripts": {
"postinstall": "node scripts/postinstall.js",
"build": "node scripts/build.mjs",
"start": "node dist/index.js",
"start": "NODE_COMPILE_CACHE=${HOME}/.codeman/compile-cache node dist/index.js",
"dev": "tsx src/index.ts web",
"web": "node dist/index.js web",
"clean": "rm -rf dist",
"test": "vitest run",
"test:watch": "vitest",
"test:coverage": "vitest run --coverage",
"test": "vitest run --config config/vitest.config.ts",
"test:watch": "vitest --config config/vitest.config.ts",
"test:coverage": "vitest run --config config/vitest.config.ts --coverage",
"typecheck": "tsc --noEmit",
"lint": "eslint 'src/**/*.ts'",
"lint:fix": "eslint 'src/**/*.ts' --fix",
"lint": "eslint --config config/eslint.config.js 'src/**/*.ts'",
"lint:fix": "eslint --config config/eslint.config.js 'src/**/*.ts' --fix",
"format": "prettier --write 'src/**/*.ts'",
"format:check": "prettier --check 'src/**/*.ts'",
"capture:subagents": "node scripts/capture-subagent-screenshots.mjs",
"changeset": "changeset",
"version-packages": "changeset version",
"version-packages": "changeset version && npm install --package-lock-only && node scripts/check-lockfile-sync.mjs",
"check:lockfile": "node scripts/check-lockfile-sync.mjs",
"knip": "npx --yes knip@latest",
"release": "changeset publish"
},
"workspaces": [
@@ -51,6 +53,7 @@
"@fastify/compress": "^8.3.1",
"@fastify/cookie": "^11.0.2",
"@fastify/static": "^8.0.0",
"@fastify/websocket": "^11.2.0",
"@xterm/addon-fit": "^0.11.0",
"@xterm/addon-unicode11": "^0.9.0",
"@xterm/addon-webgl": "^0.19.0",
@@ -76,6 +79,7 @@
"@types/react": "^19.2.14",
"@types/uuid": "^10.0.0",
"@types/web-push": "^3.6.4",
"@types/ws": "^8.18.1",
"@vitest/coverage-v8": "^4.0.18",
"agent-browser": "^0.6.0",
"esbuild": "^0.27.3",
-1
View File
@@ -1 +0,0 @@
tools/remotion/public
+8 -4
View File
@@ -19,6 +19,10 @@ const PORTS = {
const results = [];
function isCodemanTitle(title) {
return typeof title === 'string' && title.startsWith('codeman:');
}
function logSection(title) {
console.log('\n' + '='.repeat(60));
console.log(` ${title}`);
@@ -88,7 +92,7 @@ async function main() {
const page = await playwrightBrowser.newPage();
await page.goto(`http://localhost:${PORTS.playwright}`);
const title = await page.title();
if (title !== 'Codeman') throw new Error(`Expected Codeman, got ${title}`);
if (!isCodemanTitle(title)) throw new Error(`Expected codeman:<hostname>, got ${title}`);
await page.close();
});
@@ -149,7 +153,7 @@ async function main() {
const page = await puppeteerBrowser.newPage();
await page.goto(`http://localhost:${PORTS.puppeteer}`);
const title = await page.title();
if (title !== 'Codeman') throw new Error(`Expected Codeman, got ${title}`);
if (!isCodemanTitle(title)) throw new Error(`Expected codeman:<hostname>, got ${title}`);
await page.close();
});
@@ -202,7 +206,7 @@ async function main() {
agentBrowser(`open http://localhost:${PORTS.agentBrowser}`);
await new Promise(r => setTimeout(r, 2000));
const title = agentBrowserJson('get title');
agentBrowserAvailable = title.title === 'Codeman';
agentBrowserAvailable = isCodemanTitle(title.title);
console.log(' Browser launched');
// Test 1: Page load
@@ -210,7 +214,7 @@ async function main() {
agentBrowser(`open http://localhost:${PORTS.agentBrowser}`);
await new Promise(r => setTimeout(r, 1000));
const title = agentBrowserJson('get title');
if (title.title !== 'Codeman') throw new Error(`Expected Codeman, got ${title.title}`);
if (!isCodemanTitle(title.title)) throw new Error(`Expected codeman:<hostname>, got ${title.title}`);
});
// Test 2: Element selection
+14
View File
@@ -59,7 +59,14 @@ appendFileSync(
);
// 4. Minify frontend assets
run('minify input-cjk.js', 'npx esbuild dist/web/public/input-cjk.js --minify --outfile=dist/web/public/input-cjk.js --allow-overwrite');
run('minify app.js', 'npx esbuild dist/web/public/app.js --minify --outfile=dist/web/public/app.js --allow-overwrite');
run('minify terminal-ui.js', 'npx esbuild dist/web/public/terminal-ui.js --minify --outfile=dist/web/public/terminal-ui.js --allow-overwrite');
run('minify respawn-ui.js', 'npx esbuild dist/web/public/respawn-ui.js --minify --outfile=dist/web/public/respawn-ui.js --allow-overwrite');
run('minify ralph-panel.js', 'npx esbuild dist/web/public/ralph-panel.js --minify --outfile=dist/web/public/ralph-panel.js --allow-overwrite');
run('minify settings-ui.js', 'npx esbuild dist/web/public/settings-ui.js --minify --outfile=dist/web/public/settings-ui.js --allow-overwrite');
run('minify panels-ui.js', 'npx esbuild dist/web/public/panels-ui.js --minify --outfile=dist/web/public/panels-ui.js --allow-overwrite');
run('minify session-ui.js', 'npx esbuild dist/web/public/session-ui.js --minify --outfile=dist/web/public/session-ui.js --allow-overwrite');
run('minify styles.css', 'npx esbuild dist/web/public/styles.css --minify --outfile=dist/web/public/styles.css --allow-overwrite');
run('minify mobile.css', 'npx esbuild dist/web/public/mobile.css --minify --outfile=dist/web/public/mobile.css --allow-overwrite');
@@ -75,7 +82,14 @@ console.log('\n[build] content-hash cache busting');
'voice-input.js',
'notification-manager.js',
'keyboard-accessory.js',
'input-cjk.js',
'app.js',
'terminal-ui.js',
'respawn-ui.js',
'ralph-panel.js',
'settings-ui.js',
'panels-ui.js',
'session-ui.js',
'ralph-wizard.js',
'api-client.js',
'subagent-windows.js',
+1 -1
View File
@@ -428,7 +428,7 @@ const SUBAGENT_ACTIVITY = {
'agent-002': [
{ type: 'tool', tool: 'Glob', input: { pattern: 'test/**/*.test.ts' }, timestamp: new Date().toISOString(), agentId: 'agent-002' },
{ type: 'tool', tool: 'Read', input: { file_path: '/home/arkon/codeman/test/respawn-test-utils.ts' }, timestamp: new Date().toISOString(), agentId: 'agent-002' },
{ type: 'tool', tool: 'Read', input: { file_path: '/home/arkon/codeman/vitest.config.ts' }, timestamp: new Date().toISOString(), agentId: 'agent-002' },
{ type: 'tool', tool: 'Read', input: { file_path: '/home/arkon/codeman/config/vitest.config.ts' }, timestamp: new Date().toISOString(), agentId: 'agent-002' },
{ type: 'message', role: 'assistant', text: 'Analyzing test patterns: MockSession, unique ports, fileParallelism: false...', timestamp: new Date().toISOString(), agentId: 'agent-002' },
],
};
+14 -14
View File
@@ -9,7 +9,7 @@
*
* Usage: node scripts/capture-video-screenshots.mjs
* Port: 3198 (static file server)
* Output: remotion/public/ (6 PNGs)
* Output: scripts/scripts/remotion/public/ (6 PNGs)
*/
import { chromium } from 'playwright';
@@ -21,7 +21,7 @@ import { fileURLToPath } from 'url';
const __dirname = fileURLToPath(new URL('.', import.meta.url));
const PROJECT_ROOT = join(__dirname, '..');
const PUBLIC_DIR = join(PROJECT_ROOT, 'src', 'web', 'public');
const OUTPUT_DIR = join(PROJECT_ROOT, 'remotion', 'public');
const OUTPUT_DIR = join(PROJECT_ROOT, 'scripts', 'remotion', 'public');
const PORT = 3198;
const DESKTOP_VIEWPORT = { width: 1920, height: 1080 };
@@ -516,7 +516,7 @@ async function captureDesktopWelcome(browser) {
path: join(OUTPUT_DIR, 'desktop-welcome.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/desktop-welcome.png');
console.log(' Saved: scripts/remotion/public/desktop-welcome.png');
} finally {
await context.close();
}
@@ -545,7 +545,7 @@ async function captureDesktopClaude(browser) {
path: join(OUTPUT_DIR, 'desktop-claude.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/desktop-claude.png');
console.log(' Saved: scripts/remotion/public/desktop-claude.png');
} finally {
await context.close();
}
@@ -575,7 +575,7 @@ async function captureDesktopBothClaude(browser) {
path: join(OUTPUT_DIR, 'desktop-both-claude.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/desktop-both-claude.png');
console.log(' Saved: scripts/remotion/public/desktop-both-claude.png');
} finally {
await context.close();
}
@@ -605,7 +605,7 @@ async function captureDesktopBothOpencode(browser) {
path: join(OUTPUT_DIR, 'desktop-both-opencode.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/desktop-both-opencode.png');
console.log(' Saved: scripts/remotion/public/desktop-both-opencode.png');
} finally {
await context.close();
}
@@ -634,7 +634,7 @@ async function captureMobileClaude(browser) {
path: join(OUTPUT_DIR, 'mobile-claude.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/mobile-claude.png');
console.log(' Saved: scripts/remotion/public/mobile-claude.png');
} finally {
await context.close();
}
@@ -663,7 +663,7 @@ async function captureMobileOpencode(browser) {
path: join(OUTPUT_DIR, 'mobile-opencode.png'),
fullPage: false,
});
console.log(' Saved: remotion/public/mobile-opencode.png');
console.log(' Saved: scripts/remotion/public/mobile-opencode.png');
} finally {
await context.close();
}
@@ -706,12 +706,12 @@ async function main() {
console.log('All 6 screenshots captured!');
console.log('='.repeat(60));
console.log('\nOutput files:');
console.log(' remotion/public/desktop-welcome.png');
console.log(' remotion/public/desktop-claude.png');
console.log(' remotion/public/desktop-both-claude.png');
console.log(' remotion/public/desktop-both-opencode.png');
console.log(' remotion/public/mobile-claude.png');
console.log(' remotion/public/mobile-opencode.png');
console.log(' scripts/remotion/public/desktop-welcome.png');
console.log(' scripts/remotion/public/desktop-claude.png');
console.log(' scripts/remotion/public/desktop-both-claude.png');
console.log(' scripts/remotion/public/desktop-both-opencode.png');
console.log(' scripts/remotion/public/mobile-claude.png');
console.log(' scripts/remotion/public/mobile-opencode.png');
} catch (err) {
console.error('\nFatal error:', err.message);
console.error(err.stack);
+29
View File
@@ -0,0 +1,29 @@
#!/usr/bin/env node
// Fails if package-lock.json's version fields don't match package.json.
// Changesets bumps package.json but NOT the lockfile — this catches that drift
// (the top-level `version` in lockfiles is metadata, so `npm ci` won't flag it).
import { readFileSync } from 'node:fs';
import { resolve, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
const repoRoot = resolve(dirname(fileURLToPath(import.meta.url)), '..');
const pkg = JSON.parse(readFileSync(resolve(repoRoot, 'package.json'), 'utf8'));
const lock = JSON.parse(readFileSync(resolve(repoRoot, 'package-lock.json'), 'utf8'));
const expected = pkg.version;
const rootVersion = lock.version;
const selfVersion = lock.packages?.['']?.version;
const mismatches = [];
if (rootVersion !== expected) mismatches.push(` package-lock.json#.version = ${rootVersion} (expected ${expected})`);
if (selfVersion !== expected) mismatches.push(` package-lock.json#.packages[""].version = ${selfVersion} (expected ${expected})`);
if (mismatches.length > 0) {
console.error(`\nLockfile version drift detected (package.json is ${expected}):`);
console.error(mismatches.join('\n'));
console.error('\nFix: run `npm install --package-lock-only` and commit the updated package-lock.json.\n');
process.exit(1);
}
console.log(`Lockfile in sync with package.json (${expected}).`);
+19
View File
@@ -0,0 +1,19 @@
[Unit]
Description=Codeman Cloudflare Named Tunnel
After=network-online.target codeman-web.service
Wants=network-online.target
[Service]
Type=simple
ExecStart=/usr/bin/cloudflared tunnel --config %h/.cloudflared/codeman.yml run codeman
Restart=always
RestartSec=5
KillMode=process
# Logging
StandardOutput=journal
StandardError=journal
SyslogIdentifier=codeman-tunnel-named
[Install]
WantedBy=default.target
+1
View File
@@ -11,6 +11,7 @@ RestartSec=5
KillMode=process
Environment=NODE_ENV=production
Environment=HOME=/home/arkon
Environment=NODE_COMPILE_CACHE=/home/arkon/.codeman/compile-cache
# Logging
StandardOutput=journal

Before

Width:  |  Height:  |  Size: 41 KiB

After

Width:  |  Height:  |  Size: 41 KiB

Before

Width:  |  Height:  |  Size: 37 KiB

After

Width:  |  Height:  |  Size: 37 KiB

Before

Width:  |  Height:  |  Size: 39 KiB

After

Width:  |  Height:  |  Size: 39 KiB

Before

Width:  |  Height:  |  Size: 57 KiB

After

Width:  |  Height:  |  Size: 57 KiB

Before

Width:  |  Height:  |  Size: 390 KiB

After

Width:  |  Height:  |  Size: 390 KiB

Before

Width:  |  Height:  |  Size: 22 KiB

After

Width:  |  Height:  |  Size: 22 KiB

+199 -35
View File
@@ -1,45 +1,209 @@
#!/usr/bin/env bash
# Quick Cloudflare Tunnel for Codeman
# Usage: ./scripts/tunnel.sh [start|stop|status|url]
# Cloudflare Tunnel manager for Codeman
# Usage: ./scripts/tunnel.sh [quick|named] [start|stop|status|url]
#
# Modes:
# quick — Quick tunnel with random trycloudflare.com URL (default)
# named — Named tunnel on a fixed hostname (requires setup, see below)
#
# Environment variables:
# CLOUDFLARED_TUNNEL_NAME — tunnel name (default: codeman)
# CLOUDFLARED_TUNNEL_ID — tunnel UUID (from: cloudflared tunnel list)
# CODEMAN_TUNNEL_HOSTNAME — public hostname (e.g. codeman.example.com)
#
# First-time named tunnel setup:
# cloudflared tunnel login
# cloudflared tunnel create <tunnel-name>
# cloudflared tunnel route dns <tunnel-name> <hostname>
# ./scripts/tunnel.sh named setup # writes ~/.cloudflared/<tunnel-name>.yml
set -euo pipefail
SERVICE="codeman-tunnel"
QUICK_SERVICE="codeman-tunnel"
NAMED_SERVICE="codeman-tunnel-named"
TUNNEL_NAME="${CLOUDFLARED_TUNNEL_NAME:-codeman}"
TUNNEL_HOSTNAME="${CODEMAN_TUNNEL_HOSTNAME:-codeman.example.com}"
CODEMAN_PORT="3000"
LOG_FILE="$HOME/.codeman/tunnel.log"
case "${1:-start}" in
start)
if ! systemctl --user is-active "$SERVICE" &>/dev/null; then
# Install service if not already
if ! systemctl --user cat "$SERVICE" &>/dev/null 2>&1; then
cp "$(dirname "$0")/codeman-tunnel.service" "$HOME/.config/systemd/user/"
systemctl --user daemon-reload
fi
systemctl --user start "$SERVICE"
echo "Tunnel starting... waiting for URL"
sleep 6
fi
# Extract the tunnel URL from journal
URL=$(grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$HOME/.codeman/tunnel.log" 2>/dev/null | tail -1)
if [ -n "$URL" ]; then
echo "$URL"
else
echo "URL not ready yet, try: $0 url"
fi
# ── helpers ──────────────────────────────────────────────────────────────────
_require_cloudflared() {
if ! command -v cloudflared &>/dev/null; then
echo "Error: cloudflared not found. Install with: yay -S cloudflared" >&2
exit 1
fi
}
_cloudflared_bin() {
command -v cloudflared
}
_install_service() {
local svc_file="$1"
local svc_name="$2"
if ! systemctl --user cat "$svc_name" &>/dev/null 2>&1; then
cp "$(dirname "$0")/$svc_file" "$HOME/.config/systemd/user/"
systemctl --user daemon-reload
echo "Service $svc_name installed."
fi
}
_install_named_service() {
if ! systemctl --user cat "$NAMED_SERVICE" &>/dev/null 2>&1; then
# Generate service file with the configured tunnel name
sed "s/codeman\.yml/$TUNNEL_NAME.yml/g; s/run codeman/run $TUNNEL_NAME/g" \
"$(dirname "$0")/codeman-tunnel-named.service" \
> "$HOME/.config/systemd/user/codeman-tunnel-named.service"
systemctl --user daemon-reload
echo "Service $NAMED_SERVICE installed (tunnel: $TUNNEL_NAME)."
fi
}
# ── named tunnel setup ───────────────────────────────────────────────────────
_named_setup() {
_require_cloudflared
local creds_dir="$HOME/.cloudflared"
local config_file="$creds_dir/$TUNNEL_NAME.yml"
# Replace with your tunnel ID (from: cloudflared tunnel list)
local tunnel_id="${CLOUDFLARED_TUNNEL_ID:-YOUR_TUNNEL_ID_HERE}"
local creds_file="$creds_dir/$tunnel_id.json"
if [ ! -f "$creds_file" ]; then
echo "Credentials not found: $creds_file"
echo "Run: cloudflared tunnel create $TUNNEL_NAME"
exit 1
fi
cat > "$config_file" <<EOF
tunnel: $tunnel_id
credentials-file: $creds_file
ingress:
- hostname: $TUNNEL_HOSTNAME
service: http://localhost:$CODEMAN_PORT
- service: http_status:404
EOF
echo "Config written to $config_file"
echo "Tunnel ID: $tunnel_id"
echo "Hostname: $TUNNEL_HOSTNAME"
echo ""
echo "Next steps:"
echo " 1. Add Cloudflare Access policy for $TUNNEL_HOSTNAME (Zero Trust dashboard)"
echo " 2. ./scripts/tunnel.sh named start"
}
# ── quick mode ───────────────────────────────────────────────────────────────
_quick_start() {
if ! systemctl --user is-active "$QUICK_SERVICE" &>/dev/null; then
_install_service "codeman-tunnel.service" "$QUICK_SERVICE"
systemctl --user start "$QUICK_SERVICE"
echo "Quick tunnel starting... waiting for URL"
sleep 6
fi
local url
url=$(grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$LOG_FILE" 2>/dev/null | tail -1)
if [ -n "$url" ]; then
echo "$url"
else
echo "URL not ready yet, try: $0 quick url"
fi
}
_quick_stop() {
systemctl --user stop "$QUICK_SERVICE"
echo "Quick tunnel stopped"
}
_quick_status() {
systemctl --user status "$QUICK_SERVICE" --no-pager 2>&1 | head -10
echo ""
echo "URL:"
grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$LOG_FILE" 2>/dev/null | tail -1
}
_quick_url() {
grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$LOG_FILE" 2>/dev/null | tail -1
}
# ── named mode ───────────────────────────────────────────────────────────────
_named_start() {
_require_cloudflared
if [ ! -f "$HOME/.cloudflared/$TUNNEL_NAME.yml" ]; then
echo "Config not found. Run: $0 named setup"
exit 1
fi
if ! systemctl --user is-active "$NAMED_SERVICE" &>/dev/null; then
_install_named_service
systemctl --user start "$NAMED_SERVICE"
echo "Named tunnel starting..."
sleep 3
fi
echo "https://$TUNNEL_HOSTNAME"
}
_named_stop() {
systemctl --user stop "$NAMED_SERVICE"
echo "Named tunnel stopped"
}
_named_status() {
systemctl --user status "$NAMED_SERVICE" --no-pager 2>&1 | head -10
echo ""
echo "URL: https://$TUNNEL_HOSTNAME"
}
_named_enable() {
_install_named_service
systemctl --user enable "$NAMED_SERVICE"
echo "Named tunnel enabled at boot."
}
_named_disable() {
systemctl --user disable "$NAMED_SERVICE"
echo "Named tunnel disabled."
}
# ── dispatch ─────────────────────────────────────────────────────────────────
MODE="${1:-quick}"
CMD="${2:-start}"
case "$MODE" in
quick)
case "$CMD" in
start) _quick_start ;;
stop) _quick_stop ;;
status) _quick_status ;;
url) _quick_url ;;
*) echo "Usage: $0 quick [start|stop|status|url]"; exit 1 ;;
esac
;;
stop)
systemctl --user stop "$SERVICE"
echo "Tunnel stopped"
;;
status)
systemctl --user status "$SERVICE" --no-pager 2>&1 | head -10
echo ""
echo "URL:"
grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$HOME/.codeman/tunnel.log" 2>/dev/null | tail -1
;;
url)
grep -oP 'https://[a-z0-9-]+\.trycloudflare\.com' "$HOME/.codeman/tunnel.log" 2>/dev/null | tail -1
named)
case "$CMD" in
start) _named_start ;;
stop) _named_stop ;;
status) _named_status ;;
url) echo "https://$TUNNEL_HOSTNAME" ;;
setup) _named_setup ;;
enable) _named_enable ;;
disable) _named_disable ;;
*) echo "Usage: $0 named [start|stop|status|url|setup|enable|disable]"; exit 1 ;;
esac
;;
# backward compat: no mode prefix → quick tunnel
start) _quick_start ;;
stop) _quick_stop ;;
status) _quick_status ;;
url) _quick_url ;;
*)
echo "Usage: $0 [start|stop|status|url]"
echo "Usage: $0 [quick|named] [start|stop|status|url]"
echo " $0 named setup # first-time named tunnel configuration"
echo " $0 named enable # start at boot"
exit 1
;;
esac
+4 -6
View File
@@ -28,9 +28,9 @@ import { existsSync, readFileSync, unlinkSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { EventEmitter } from 'node:events';
import { getAugmentedPath } from './utils/claude-cli-resolver.js';
import { ANSI_ESCAPE_PATTERN_SIMPLE } from './utils/index.js';
import { getAugmentedPath, ANSI_ESCAPE_PATTERN_SIMPLE } from './utils/index.js';
import { AI_CHECK_MAX_BACKOFF_MS } from './config/ai-defaults.js';
import { getErrorMessage } from './types.js';
// ========== Security Validation ==========
@@ -294,7 +294,7 @@ export abstract class AiCheckerBase<
this.emit('checkCompleted', result);
return result;
} catch (err) {
const errorMsg = err instanceof Error ? err.message : String(err);
const errorMsg = getErrorMessage(err);
this.handleError(errorMsg);
const result = this.createErrorResult(errorMsg, Date.now() - this.checkStartTime);
this.emit('checkFailed', errorMsg);
@@ -413,9 +413,7 @@ export abstract class AiCheckerBase<
});
muxProcess.unref();
} catch (err) {
throw new Error(
`Failed to spawn ${this.checkDescription} tmux session: ${err instanceof Error ? err.message : String(err)}`
);
throw new Error(`Failed to spawn ${this.checkDescription} tmux session: ${getErrorMessage(err)}`);
}
// Poll the temp file for completion
+1 -1
View File
@@ -41,7 +41,7 @@ import {
// ========== Types ==========
export type AiIdleCheckConfig = AiCheckerConfigBase;
type AiIdleCheckConfig = AiCheckerConfigBase;
export type AiCheckVerdict = 'IDLE' | 'WORKING' | 'ERROR';
+2 -2
View File
@@ -40,13 +40,13 @@ import {
// ========== Types ==========
export type AiPlanCheckConfig = AiCheckerConfigBase;
type AiPlanCheckConfig = AiCheckerConfigBase;
export type AiPlanCheckVerdict = 'PLAN_MODE' | 'NOT_PLAN_MODE' | 'ERROR';
export type AiPlanCheckResult = AiCheckerResultBase<AiPlanCheckVerdict>;
export type AiPlanCheckState = AiCheckerStateBase<AiPlanCheckVerdict>;
type AiPlanCheckState = AiCheckerStateBase<AiPlanCheckVerdict>;
// ========== Constants ==========
+104 -113
View File
@@ -99,7 +99,7 @@ const LOG_FILE_MENTION_PATTERN = /([/~][^\s'"<>|;&\n]*(?:\.log|\.txt|\.out|\/log
/**
* Events emitted by BashToolParser.
*/
export interface BashToolParserEvents {
interface BashToolParserEvents {
/** New Bash tool with file paths started */
toolStart: [tool: ActiveBashTool];
/** Bash tool completed */
@@ -111,7 +111,7 @@ export interface BashToolParserEvents {
/**
* Configuration options for BashToolParser.
*/
export interface BashToolParserConfig {
interface BashToolParserConfig {
/** Session ID this parser belongs to */
sessionId: string;
/** Whether the parser is enabled (default: true) */
@@ -470,113 +470,91 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
* Process a single pre-stripped line of terminal output.
*/
private processCleanLine(cleanLine: string): void {
// Check for tool start
if (this._handleToolStart(cleanLine)) return;
if (this._handleToolCompletion(cleanLine)) return;
if (this._handleTextCommand(cleanLine)) return;
this._handleLogFileMention(cleanLine);
}
private _handleToolStart(cleanLine: string): boolean {
const startMatch = cleanLine.match(BASH_TOOL_START_PATTERN);
if (startMatch) {
const command = startMatch[1];
const timeout = startMatch[2]?.trim();
if (!startMatch) return false;
// Check if this is a file-viewing command
if (this.isFileViewerCommand(command)) {
const filePaths = this.extractFilePaths(command);
const command = startMatch[1];
const timeout = startMatch[2]?.trim();
// Skip if any file path is already tracked (cross-pattern dedup)
if (filePaths.some((fp) => this.isFilePathTracked(fp))) {
return;
}
if (!this.isFileViewerCommand(command)) return true;
if (filePaths.length > 0) {
const tool: ActiveBashTool = {
id: uuidv4(),
command,
filePaths,
timeout,
startedAt: Date.now(),
status: 'running',
sessionId: this._sessionId,
};
const filePaths = this.extractFilePaths(command);
// Enforce max tools limit
if (this._activeTools.size >= MAX_ACTIVE_TOOLS) {
// Remove oldest tool
const oldest = Array.from(this._activeTools.entries()).sort((a, b) => a[1].startedAt - b[1].startedAt)[0];
if (oldest) {
this._activeTools.delete(oldest[0]);
}
// Skip if any file path is already tracked (cross-pattern dedup)
if (filePaths.some((fp) => this.isFilePathTracked(fp))) return true;
if (filePaths.length > 0) {
const tool = this._createActiveTool(command, filePaths, 'running', timeout);
// Enforce max tools limit
if (this._activeTools.size >= MAX_ACTIVE_TOOLS) {
// Remove oldest tool (O(n) min-scan instead of O(n log n) sort)
let oldestKey: string | undefined;
let oldestTime = Infinity;
for (const [key, entry] of this._activeTools) {
if (entry.startedAt < oldestTime) {
oldestTime = entry.startedAt;
oldestKey = key;
}
this._activeTools.set(tool.id, tool);
this._lastToolId = tool.id;
this.emit('toolStart', tool);
this.scheduleUpdate();
}
}
return;
}
// Check for tool completion
if (TOOL_COMPLETION_PATTERN.test(cleanLine) && this._lastToolId) {
const tool = this._activeTools.get(this._lastToolId);
if (tool && tool.status === 'running') {
tool.status = 'completed';
this.emit('toolEnd', tool);
this.scheduleUpdate();
// Remove completed tool after a short delay to allow UI to show completion
this.cleanup.setTimeout(
() => {
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this.scheduleUpdate();
},
2000,
{ description: 'auto-remove completed tool' }
);
}
this._lastToolId = null;
return;
}
// Fallback: Check for command suggestions in plain text (e.g., "tail -f /tmp/file.log")
const textCmdMatch = cleanLine.match(TEXT_COMMAND_PATTERN);
if (textCmdMatch) {
const filePath = textCmdMatch[2];
// Create a suggestion tool (marked as 'suggestion' status)
const tool: ActiveBashTool = {
id: uuidv4(),
command: cleanLine.trim(),
filePaths: [filePath],
timeout: undefined,
startedAt: Date.now(),
status: 'running', // Shows as clickable
sessionId: this._sessionId,
};
// Don't add if file path already tracked (cross-pattern dedup)
if (this.isFilePathTracked(filePath)) {
return;
if (oldestKey) {
this._activeTools.delete(oldestKey);
}
}
this._activeTools.set(tool.id, tool);
this._lastToolId = tool.id;
this.emit('toolStart', tool);
this.scheduleUpdate();
// Auto-remove suggestions after 30 seconds
this.cleanup.setTimeout(
() => {
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this.scheduleUpdate();
},
30000,
{ description: 'auto-remove suggestion tool' }
);
return;
}
// Last fallback: Check for log file paths mentioned anywhere in the line
return true;
}
private _handleToolCompletion(cleanLine: string): boolean {
if (!TOOL_COMPLETION_PATTERN.test(cleanLine) || !this._lastToolId) return false;
const tool = this._activeTools.get(this._lastToolId);
if (tool && tool.status === 'running') {
tool.status = 'completed';
this.emit('toolEnd', tool);
this.scheduleUpdate();
this._scheduleAutoRemove(tool.id, 2000, 'auto-remove completed tool');
}
this._lastToolId = null;
return true;
}
private _handleTextCommand(cleanLine: string): boolean {
const textCmdMatch = cleanLine.match(TEXT_COMMAND_PATTERN);
if (!textCmdMatch) return false;
const filePath = textCmdMatch[2];
// Don't add if file path already tracked (cross-pattern dedup)
if (this.isFilePathTracked(filePath)) return true;
const tool = this._createActiveTool(cleanLine.trim(), [filePath], 'running');
this._activeTools.set(tool.id, tool);
this.emit('toolStart', tool);
this.scheduleUpdate();
// Auto-remove suggestions after 30 seconds
this._scheduleAutoRemove(tool.id, 30000, 'auto-remove suggestion tool');
return true;
}
private _handleLogFileMention(cleanLine: string): void {
LOG_FILE_MENTION_PATTERN.lastIndex = 0;
let logMatch;
while ((logMatch = LOG_FILE_MENTION_PATTERN.exec(cleanLine)) !== null) {
@@ -588,33 +566,46 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
// Skip if file path already tracked (cross-pattern dedup)
if (this.isFilePathTracked(filePath)) continue;
const tool: ActiveBashTool = {
id: uuidv4(),
command: `View: ${filePath}`,
filePaths: [filePath],
timeout: undefined,
startedAt: Date.now(),
status: 'running',
sessionId: this._sessionId,
};
const tool = this._createActiveTool(`View: ${filePath}`, [filePath], 'running');
this._activeTools.set(tool.id, tool);
this.emit('toolStart', tool);
this.scheduleUpdate();
// Auto-remove after 60 seconds
this.cleanup.setTimeout(
() => {
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this.scheduleUpdate();
},
60000,
{ description: 'auto-remove log file tool' }
);
this._scheduleAutoRemove(tool.id, 60000, 'auto-remove log file tool');
}
}
private _createActiveTool(
command: string,
filePaths: string[],
status: ActiveBashTool['status'],
timeout?: string
): ActiveBashTool {
return {
id: uuidv4(),
command,
filePaths,
timeout,
startedAt: Date.now(),
status,
sessionId: this._sessionId,
};
}
private _scheduleAutoRemove(toolId: string, delayMs: number, description: string): void {
this.cleanup.setTimeout(
() => {
if (this._destroyed) return;
this._activeTools.delete(toolId);
this.scheduleUpdate();
},
delayMs,
{ description }
);
}
/**
* Check if a command is a file-viewing command worth tracking.
*/
+3 -1
View File
@@ -485,16 +485,18 @@ program
.description('Start the web interface')
.option('-p, --port <port>', 'Port to listen on', '3000')
.option('--https', 'Enable HTTPS with self-signed certificate (only needed for remote access, not localhost)')
.option('--title-hostname <hostname>', 'Override the hostname shown in the browser title')
.action(async (options) => {
const { startWebServer } = await import('./web/server.js');
const port = parseInt(options.port, 10);
const https = !!options.https;
const titleHostname = options.titleHostname;
const protocol = https ? 'https' : 'http';
console.log(chalk.cyan(`Starting Codeman web interface on port ${port}${https ? ' (HTTPS)' : ''}...`));
try {
const server = await startWebServer(port, https);
const server = await startWebServer(port, https, false, titleHostname);
console.log(chalk.green(`\n✓ Web interface running at ${protocol}://localhost:${port}`));
if (https) {
console.log(chalk.yellow(' Note: Accept the self-signed certificate in your browser on first visit'));
+21 -10
View File
@@ -22,14 +22,16 @@
* Maximum terminal buffer size in characters.
* Contains raw terminal output with ANSI escape sequences.
* Reduced from 5MB to 2MB for better render performance.
* Override: CODEMAN_MAX_TERMINAL_BUFFER (bytes)
*/
export const MAX_TERMINAL_BUFFER_SIZE = 2 * 1024 * 1024; // 2MB
export const MAX_TERMINAL_BUFFER_SIZE = parseInt(process.env.CODEMAN_MAX_TERMINAL_BUFFER || '') || 2 * 1024 * 1024;
/**
* Size to trim terminal buffer to when max is exceeded.
* Keeps the most recent portion to preserve context.
* Override: CODEMAN_TRIM_TERMINAL_TO (bytes)
*/
export const TRIM_TERMINAL_TO = 1.5 * 1024 * 1024; // 1.5MB
export const TRIM_TERMINAL_TO = parseInt(process.env.CODEMAN_TRIM_TERMINAL_TO || '') || 1.5 * 1024 * 1024;
// ============================================================================
// Text Output Buffer Limits
@@ -38,13 +40,15 @@ export const TRIM_TERMINAL_TO = 1.5 * 1024 * 1024; // 1.5MB
/**
* Maximum text output buffer size in characters.
* Contains ANSI-stripped text for search and analysis.
* Override: CODEMAN_MAX_TEXT_OUTPUT (bytes)
*/
export const MAX_TEXT_OUTPUT_SIZE = 1 * 1024 * 1024; // 1MB
export const MAX_TEXT_OUTPUT_SIZE = parseInt(process.env.CODEMAN_MAX_TEXT_OUTPUT || '') || 1 * 1024 * 1024;
/**
* Size to trim text output buffer to when max is exceeded.
* Override: CODEMAN_TRIM_TEXT_TO (bytes)
*/
export const TRIM_TEXT_TO = 768 * 1024; // 768KB
export const TRIM_TEXT_TO = parseInt(process.env.CODEMAN_TRIM_TEXT_TO || '') || 768 * 1024;
// ============================================================================
// Message Buffer Limits
@@ -53,13 +57,9 @@ export const TRIM_TEXT_TO = 768 * 1024; // 768KB
/**
* Maximum number of Claude JSON messages to keep in memory per session.
* Older messages are discarded when limit is exceeded.
* Override: CODEMAN_MAX_MESSAGES (count)
*/
export const MAX_MESSAGES = 1000;
/**
* Number of messages to keep when trimming (80% of max).
*/
export const TRIM_MESSAGES_TO = 800;
export const MAX_MESSAGES = parseInt(process.env.CODEMAN_MAX_MESSAGES || '') || 1000;
// ============================================================================
// Line Buffer Limits
@@ -85,3 +85,14 @@ export const MAX_RESPAWN_BUFFER_SIZE = 1 * 1024 * 1024; // 1MB
* Size to trim respawn buffer to when max is exceeded.
*/
export const TRIM_RESPAWN_BUFFER_TO = 512 * 1024; // 512KB
// ============================================================================
// File Peek Limits
// ============================================================================
/**
* Maximum bytes to read when peeking at the beginning of a file.
* Used with `createReadStream({ end })` (inclusive) to read the first 8KB,
* which is enough to extract metadata from the first few JSONL lines.
*/
export const FILE_PEEK_BYTES = 8 * 1024 - 1; // 8KB (inclusive end offset)
-6
View File
@@ -11,11 +11,5 @@
/** Max input length per API request (bytes) */
export const MAX_INPUT_LENGTH = 64 * 1024;
/** Max terminal columns for resize requests */
export const MAX_TERMINAL_COLS = 500;
/** Max terminal rows for resize requests */
export const MAX_TERMINAL_ROWS = 200;
/** Max session name length (chars) */
export const MAX_SESSION_NAME_LENGTH = 128;
+5 -7
View File
@@ -16,6 +16,7 @@ import { existsSync, statSync, realpathSync } from 'node:fs';
import { resolve, relative, isAbsolute } from 'node:path';
import { homedir } from 'node:os';
import { EventEmitter } from 'node:events';
import { getErrorMessage } from './types.js';
import { CLEANUP_CHECK_INTERVAL_MS, INACTIVITY_TIMEOUT_MS } from './config/server-timing.js';
// ========== Configuration Constants ==========
@@ -47,7 +48,7 @@ const STREAM_INACTIVITY_TIMEOUT_MS = INACTIVITY_TIMEOUT_MS;
/**
* Represents an active file stream.
*/
export interface FileStream {
interface FileStream {
/** Unique stream identifier */
id: string;
/** Session this stream belongs to */
@@ -73,7 +74,7 @@ export interface FileStream {
/**
* Options for creating a file stream.
*/
export interface CreateStreamOptions {
interface CreateStreamOptions {
/** Session ID requesting the stream */
sessionId: string;
/** Path to the file to stream */
@@ -93,7 +94,7 @@ export interface CreateStreamOptions {
/**
* Result of creating a stream.
*/
export interface CreateStreamResult {
interface CreateStreamResult {
success: boolean;
streamId?: string;
error?: string;
@@ -172,10 +173,7 @@ export class FileStreamManager extends EventEmitter {
}
} catch (err) {
const errorCode = err instanceof Error && 'code' in err ? (err as NodeJS.ErrnoException).code : 'UNKNOWN';
console.warn(
`[FileStreamManager] Failed to stat file "${absolutePath}" (${errorCode}):`,
err instanceof Error ? err.message : String(err)
);
console.warn(`[FileStreamManager] Failed to stat file "${absolutePath}" (${errorCode}):`, getErrorMessage(err));
return { success: false, error: 'File not found or not accessible' };
}
+64
View File
@@ -84,6 +84,42 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
};
}
/**
* Remove a subset of env keys from .claude/settings.local.json.env if present.
* Used during the disk→tmux-setenv migration: when the caller is actively setting
* a fresh value for a Codeman-managed key, any stale disk entry for THAT KEY is
* superseded and should be removed. Keys NOT in `keysToRemove` are left alone
* (they may be user-managed). No-op if the file/keys don't exist.
*/
export async function stripCaseEnvKeys(casePath: string, keysToRemove: readonly string[]): Promise<void> {
if (keysToRemove.length === 0) return;
const settingsPath = join(casePath, '.claude', 'settings.local.json');
if (!existsSync(settingsPath)) return;
let existing: Record<string, unknown>;
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
return; // Malformed — don't rewrite it
}
const env = existing.env as Record<string, string> | undefined;
if (!env) return;
let changed = false;
for (const key of keysToRemove) {
if (key in env) {
delete env[key];
changed = true;
}
}
if (!changed) return;
existing.env = env;
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
}
/**
* Updates env vars in .claude/settings.local.json for the given case path.
* Merges with existing env field; removes vars set to empty string.
@@ -116,6 +152,34 @@ export async function updateCaseEnvVars(casePath: string, envVars: Record<string
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
}
/**
* Updates the `model` field in .claude/settings.local.json for the given case path.
* Pass a non-empty string to set, or empty/null to remove.
*/
export async function updateCaseModel(casePath: string, model: string | null): Promise<void> {
const claudeDir = join(casePath, '.claude');
if (!existsSync(claudeDir)) {
await mkdir(claudeDir, { recursive: true });
}
const settingsPath = join(claudeDir, 'settings.local.json');
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
existing = {};
}
if (model) {
existing.model = model;
} else {
delete existing.model;
}
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
}
/**
* Writes hooks config to .claude/settings.local.json in the given case path.
* Merges with existing file content, only touching the `hooks` key.
-5
View File
@@ -17,11 +17,6 @@ import { KeyedDebouncer } from './utils/index.js';
// ========== Types ==========
export interface ImageWatcherEvents {
'image:detected': (event: ImageDetectedEvent) => void;
'image:error': (error: Error, sessionId?: string) => void;
}
// ========== Constants ==========
/** Supported image file extensions (lowercase) */
+4
View File
@@ -63,6 +63,8 @@ export interface CreateSessionOptions {
openCodeConfig?: OpenCodeConfig;
/** When restoring after reboot, resume a previous Claude conversation by its session ID */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (e.g., CLAUDE_CODE_EFFORT_LEVEL). Ephemeral — not written to disk. */
envOverrides?: Record<string, string>;
}
/** Options for respawning a dead pane. */
@@ -77,6 +79,8 @@ export interface RespawnPaneOptions {
openCodeConfig?: OpenCodeConfig;
/** Resume a previous Claude conversation when respawning */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (preserved across respawns). */
envOverrides?: Record<string, string>;
}
/**
+978
View File
@@ -0,0 +1,978 @@
/**
* @fileoverview Orchestrator Loop — phased plan execution with team agents.
*
* State machine that generates plans from user goals, executes them
* phase-by-phase with verification gates, and adapts on failure.
*
* States: idle → planning → approval → executing → verifying → (replanning) → completed/failed
*
* Key exports:
* - `OrchestratorLoop` class — main engine, extends EventEmitter
* - `OrchestratorLoopEvents` interface — typed event map
*
* Lifecycle: `start(goal)` → plan → approve → execute phases → verify → complete
*
* @dependencies orchestrator-planner (plan generation), orchestrator-verifier (phase verification),
* session-manager (sessions), task-queue (task execution), state-store (persistence),
* prompts/orchestrator (prompt templates)
* @consumedby web/server (orchestrator routes, SSE)
* @emits stateChanged, planReady, phaseStarted, phaseCompleted, phaseFailed,
* taskAssigned, taskCompleted, taskFailed, verificationResult, completed, error
* @persistence Orchestrator state saved to `~/.codeman/state.json` (orchestrator key)
*
* @module orchestrator-loop
*/
import { EventEmitter } from 'node:events';
import { getSessionManager, SessionManager } from './session-manager.js';
import { getTaskQueue, TaskQueue } from './task-queue.js';
import { getStore, StateStore } from './state-store.js';
import { OrchestratorPlanner } from './orchestrator-planner.js';
import { OrchestratorVerifier } from './orchestrator-verifier.js';
import { PHASE_EXECUTION_PROMPT, REPLAN_PROMPT, SINGLE_TASK_PROMPT, TEAM_LEAD_PROMPT } from './prompts/index.js';
import type { TerminalMultiplexer } from './mux-interface.js';
import type { CreateTaskOptions } from './task.js';
import {
type OrchestratorState,
type OrchestratorPlan,
type OrchestratorPhase,
type OrchestratorTask,
type OrchestratorConfig,
type OrchestratorStats,
type OrchestratorPersistState,
type VerificationResult,
DEFAULT_ORCHESTRATOR_CONFIG,
createInitialOrchestratorStats,
getErrorMessage,
} from './types.js';
// ═══════════════════════════════════════════════════════════════
// Constants
// ═══════════════════════════════════════════════════════════════
/** Poll interval for checking task completion within a phase (2 seconds) */
const PHASE_POLL_INTERVAL_MS = 2000;
/** Delay between phase completion and verification (1 second) */
const POST_PHASE_DELAY_MS = 1000;
// ═══════════════════════════════════════════════════════════════
// Events
// ═══════════════════════════════════════════════════════════════
// ═══════════════════════════════════════════════════════════════
// OrchestratorLoop
// ═══════════════════════════════════════════════════════════════
export class OrchestratorLoop extends EventEmitter {
private _state: OrchestratorState = 'idle';
private plan: OrchestratorPlan | null = null;
private currentPhaseIndex = 0;
private config: OrchestratorConfig;
private stats: OrchestratorStats;
private startedAt: number | null = null;
private completedAt: number | null = null;
private workingDir: string;
private planner: OrchestratorPlanner;
private verifier: OrchestratorVerifier;
private sessionManager: SessionManager;
private taskQueue: TaskQueue;
private store: StateStore;
/** State before pause (to resume to correct state) */
private pausedState: OrchestratorState | null = null;
/** Phase poll timer for checking task completion */
private phasePollTimer: NodeJS.Timeout | null = null;
/** Phase-level timeout timer */
private phaseTimeoutTimer: NodeJS.Timeout | null = null;
/** Post-phase delay timer before verification */
private postPhaseTimer: NodeJS.Timeout | null = null;
/** Session completion listener (bound for cleanup) */
private sessionCompletionListener: ((sessionId: string, phrase: string) => void) | null = null;
/** Active sessions assigned to current phase */
private phaseSessionIds: Set<string> = new Set();
constructor(mux: TerminalMultiplexer, workingDir: string, config?: Partial<OrchestratorConfig>) {
super();
this.workingDir = workingDir;
this.config = { ...DEFAULT_ORCHESTRATOR_CONFIG, ...config };
this.stats = createInitialOrchestratorStats();
this.sessionManager = getSessionManager();
this.taskQueue = getTaskQueue();
this.store = getStore();
this.planner = new OrchestratorPlanner(mux, workingDir, this.config);
this.verifier = new OrchestratorVerifier(this.config);
// Restore state if crashed while running
this.restore();
}
// ═══════════════════════════════════════════════════════════════
// Public API — Lifecycle
// ═══════════════════════════════════════════════════════════════
/** Start orchestration with a goal. Transitions: idle → planning */
async start(goal: string): Promise<void> {
if (this._state !== 'idle' && this._state !== 'failed' && this._state !== 'completed') {
throw new Error(`Cannot start from state "${this._state}"`);
}
this.reset();
this.startedAt = Date.now();
this.setState('planning');
try {
const plan = await this.planner.generatePlan(goal, (phase, detail) => {
this.emit('planProgress', phase, detail);
});
if (this.currentState() !== 'planning') {
// Cancelled during planning
return;
}
this.plan = plan;
this.persist();
if (this.config.autoApprove) {
this.setState('executing');
await this.executeCurrentPhase();
} else {
this.setState('approval');
this.emit('planReady', plan);
}
} catch (err) {
this.handleError(err);
}
}
/** Approve the generated plan. Transitions: approval → executing */
async approve(): Promise<void> {
this.requireState('approval');
if (!this.plan) {
throw new Error('No plan to approve');
}
this.setState('executing');
await this.executeCurrentPhase();
}
/** Reject plan with feedback. Transitions: approval → planning (regenerate) */
async reject(feedback: string): Promise<void> {
this.requireState('approval');
if (!this.plan) {
throw new Error('No plan to reject');
}
const goal = this.plan.goal + '\n\nFeedback on previous plan: ' + feedback;
this.plan = null;
this.setState('planning');
try {
const plan = await this.planner.generatePlan(goal);
if ((this._state as OrchestratorState) !== 'planning') return;
this.plan = plan;
this.persist();
this.setState('approval');
this.emit('planReady', plan);
} catch (err) {
this.handleError(err);
}
}
/** Pause execution. Saves current state. */
pause(): void {
if (this._state === 'idle' || this._state === 'paused' || this._state === 'completed' || this._state === 'failed') {
return;
}
this.pausedState = this._state;
this.clearPhasePoll();
this.cleanupTaskHandlers();
this.setState('paused');
}
/** Resume from pause. */
async resume(): Promise<void> {
if (this._state !== 'paused' || !this.pausedState) {
throw new Error('Not paused');
}
const resumeTo = this.pausedState;
this.pausedState = null;
this.setState(resumeTo);
// Re-enter the appropriate phase of execution
if (resumeTo === 'executing') {
await this.executeCurrentPhase();
} else if (resumeTo === 'verifying') {
await this.verifyCurrentPhase();
}
}
/** Stop everything and clean up. */
async stop(): Promise<void> {
this.clearPhasePoll();
this.cleanupTaskHandlers();
await this.planner.cancel();
this.setState('idle');
this.store.clearOrchestratorState();
}
/** Skip a specific phase. */
async skipPhase(phaseId: string): Promise<void> {
if (!this.plan) return;
const phase = this.plan.phases.find((p) => p.id === phaseId);
if (!phase) throw new Error(`Phase "${phaseId}" not found`);
phase.status = 'skipped';
phase.completedAt = Date.now();
this.persist();
// If this is the current phase, advance
if (this.plan.phases[this.currentPhaseIndex]?.id === phaseId) {
await this.advanceToNextPhase();
}
}
/** Retry a failed phase. */
async retryPhase(phaseId: string): Promise<void> {
if (!this.plan) return;
if (this._state !== 'executing' && this._state !== 'failed') {
throw new Error(`Cannot retry from state "${this._state}"`);
}
const phaseIndex = this.plan.phases.findIndex((p) => p.id === phaseId);
if (phaseIndex === -1) throw new Error(`Phase "${phaseId}" not found`);
const phase = this.plan.phases[phaseIndex];
phase.status = 'pending';
phase.attempts = 0;
for (const task of phase.tasks) {
task.status = 'pending';
task.error = null;
task.assignedSessionId = null;
task.queueTaskId = null;
}
this.currentPhaseIndex = phaseIndex;
this.setState('executing');
await this.executeCurrentPhase();
}
// ═══════════════════════════════════════════════════════════════
// Public API — Getters
// ═══════════════════════════════════════════════════════════════
get state(): OrchestratorState {
return this._state;
}
getPlan(): OrchestratorPlan | null {
return this.plan;
}
getCurrentPhase(): OrchestratorPhase | null {
if (!this.plan) return null;
return this.plan.phases[this.currentPhaseIndex] ?? null;
}
getStats(): OrchestratorStats {
return { ...this.stats };
}
getStatus(): OrchestratorPersistState {
return {
state: this._state,
plan: this.plan,
currentPhaseIndex: this.currentPhaseIndex,
startedAt: this.startedAt,
completedAt: this.completedAt,
config: this.config,
stats: this.stats,
};
}
isRunning(): boolean {
return this._state !== 'idle' && this._state !== 'completed' && this._state !== 'failed';
}
// ═══════════════════════════════════════════════════════════════
// Internal — Phase Execution
// ═══════════════════════════════════════════════════════════════
private async executeCurrentPhase(): Promise<void> {
if (!this.plan || this._state !== 'executing') return;
const phase = this.plan.phases[this.currentPhaseIndex];
if (!phase) {
// All phases done
await this.handleCompletion();
return;
}
// Skip already completed/skipped phases
if (phase.status === 'passed' || phase.status === 'skipped') {
await this.advanceToNextPhase();
return;
}
phase.status = 'executing';
phase.startedAt = Date.now();
phase.attempts++;
this.persist();
this.emit('phaseStarted', phase);
try {
await this.assignPhaseTasks(phase);
this.startPhasePoll(phase);
} catch (err) {
this.handlePhaseError(phase, getErrorMessage(err));
}
}
private async assignPhaseTasks(phase: OrchestratorPhase): Promise<void> {
// For team strategy, send a single comprehensive prompt to a lead session
if (phase.teamStrategy.type === 'team') {
await this.assignTeamPhase(phase);
return;
}
// For single/parallel strategy, add individual tasks to TaskQueue
for (const task of phase.tasks) {
if (task.status !== 'pending') continue;
const prompt = this.buildTaskPrompt(task, phase);
const taskOptions: CreateTaskOptions = {
prompt,
workingDir: this.workingDir,
priority: 100 - phase.order, // Earlier phases get higher priority
completionPhrase: task.completionPhrase,
timeoutMs: Math.min(task.timeoutMs, this.config.phaseTimeoutMs),
};
const queueTask = this.taskQueue.addTask(taskOptions);
task.queueTaskId = queueTask.id;
task.status = 'running';
}
this.persist();
this.setupTaskHandlers();
// Manually assign tasks to idle sessions
await this.assignQueuedTasksToSessions();
}
private async assignTeamPhase(phase: OrchestratorPhase): Promise<void> {
const teamConfig = phase.teamStrategy.type === 'team' ? phase.teamStrategy.config : null;
if (!teamConfig) return;
// Find or use an idle session
const sessions = this.sessionManager.getIdleSessions();
if (sessions.length === 0) {
throw new Error('No idle sessions available for team phase execution');
}
const session = sessions[0];
this.phaseSessionIds.add(session.id);
// Mark all tasks as running under this session
for (const task of phase.tasks) {
task.status = 'running';
task.assignedSessionId = session.id;
}
// Build and send the team lead prompt
const prompt = TEAM_LEAD_PROMPT.replace('{PHASE_NAME}', phase.name)
.replace('{TASK_LIST}', phase.tasks.map((t, i) => `${i + 1}. ${t.prompt}`).join('\n'))
.replace('{TEAMMATE_HINTS}', teamConfig.suggestedTeammates.map((h, i) => `${i + 1}. ${h}`).join('\n'))
.replace('{COMPLETION_PHRASE}', `${phase.id.toUpperCase()}_COMPLETE`);
// Create a TaskQueue task for the entire phase
const queueTask = this.taskQueue.addTask({
prompt,
workingDir: this.workingDir,
priority: 100 - phase.order,
completionPhrase: `${phase.id.toUpperCase()}_COMPLETE`,
timeoutMs: this.config.phaseTimeoutMs,
});
// Link all phase tasks to this single queue task
for (const task of phase.tasks) {
task.queueTaskId = queueTask.id;
}
this.persist();
this.setupTaskHandlers();
// Assign the task to the session
try {
queueTask.assign(session.id);
session.assignTask(queueTask.id);
this.taskQueue.updateTask(queueTask);
await session.sendInput(prompt);
} catch (err) {
queueTask.fail(getErrorMessage(err));
this.taskQueue.updateTask(queueTask);
throw err;
}
}
private async assignQueuedTasksToSessions(): Promise<void> {
const idleSessions = this.sessionManager.getIdleSessions();
const maxSessions =
this.getCurrentPhase()?.teamStrategy.type === 'parallel'
? (this.getCurrentPhase()?.teamStrategy as { type: 'parallel'; maxSessions: number }).maxSessions
: 1;
const sessionsToUse = idleSessions.slice(0, maxSessions);
for (const session of sessionsToUse) {
const task = this.taskQueue.next();
if (!task) break;
try {
task.assign(session.id);
session.assignTask(task.id);
this.taskQueue.updateTask(task);
await session.sendInput(task.prompt);
this.phaseSessionIds.add(session.id);
// Find the orchestrator task linked to this queue task
const orchTask = this.findOrchestratorTaskByQueueId(task.id);
if (orchTask) {
orchTask.assignedSessionId = session.id;
orchTask.startedAt = Date.now();
this.emit('taskAssigned', orchTask, session.id);
}
} catch (err) {
task.fail(getErrorMessage(err));
session.clearTask();
this.taskQueue.updateTask(task);
}
}
}
// ═══════════════════════════════════════════════════════════════
// Internal — Task Completion Tracking
// ═══════════════════════════════════════════════════════════════
private setupTaskHandlers(): void {
this.cleanupTaskHandlers();
this.sessionCompletionListener = (_sessionId: string, _phrase: string) => {
// Session completion — check if it's related to our phase tasks
this.checkPhaseCompletion();
};
this.sessionManager.on('sessionCompletion', this.sessionCompletionListener);
}
private cleanupTaskHandlers(): void {
if (this.sessionCompletionListener) {
this.sessionManager.off('sessionCompletion', this.sessionCompletionListener);
this.sessionCompletionListener = null;
}
}
private _finalizeTask(queueTaskId: string, status: 'completed' | 'failed', error?: string): OrchestratorTask | null {
const orchTask = this.findOrchestratorTaskByQueueId(queueTaskId);
if (!orchTask) return null;
orchTask.status = status;
if (status === 'completed') {
orchTask.completedAt = Date.now();
this.stats.totalTasksCompleted++;
} else {
orchTask.error = error ?? null;
this.stats.totalTasksFailed++;
}
this.persist();
return orchTask;
}
private handleTaskCompleted(queueTaskId: string): void {
const orchTask = this._finalizeTask(queueTaskId, 'completed');
if (!orchTask) return;
this.emit('taskCompleted', orchTask);
this.checkPhaseCompletion();
}
private handleTaskFailed(queueTaskId: string, error: string): void {
const orchTask = this._finalizeTask(queueTaskId, 'failed', error);
if (!orchTask) return;
this.emit('taskFailed', orchTask, error);
// Check if we should retry the task or fail the phase
if (orchTask.retries < 2) {
orchTask.retries++;
orchTask.status = 'pending';
orchTask.error = null;
orchTask.queueTaskId = null;
// Will be re-queued on next poll
} else {
this.checkPhaseCompletion();
}
}
private startPhasePoll(phase: OrchestratorPhase): void {
this.clearPhasePoll();
this.phasePollTimer = setInterval(() => {
if (this._state !== 'executing') {
this.clearPhasePoll();
return;
}
this.pollPhaseStatus(phase);
}, PHASE_POLL_INTERVAL_MS);
// Phase-level timeout — fail the phase if it exceeds the configured timeout
this.phaseTimeoutTimer = setTimeout(() => {
if (this._state === 'executing' && phase.status === 'executing') {
console.warn(`[Orchestrator] Phase "${phase.name}" timed out after ${this.config.phaseTimeoutMs}ms`);
this.handlePhaseError(phase, `Phase timed out after ${Math.round(this.config.phaseTimeoutMs / 60000)} minutes`);
}
}, this.config.phaseTimeoutMs);
}
private _clearTimer(
timerKey: 'phasePollTimer' | 'phaseTimeoutTimer' | 'postPhaseTimer',
clearFn: typeof clearInterval | typeof clearTimeout
): void {
if (this[timerKey]) {
clearFn(this[timerKey]);
this[timerKey] = null;
}
}
private clearPhasePoll(): void {
this._clearTimer('phasePollTimer', clearInterval);
this._clearTimer('phaseTimeoutTimer', clearTimeout);
this._clearTimer('postPhaseTimer', clearTimeout);
}
private pollPhaseStatus(phase: OrchestratorPhase): void {
// Check for queued tasks that need assignment
const pendingTasks = phase.tasks.filter((t) => t.status === 'pending' && !t.queueTaskId);
if (pendingTasks.length > 0) {
// Re-queue pending tasks
for (const task of pendingTasks) {
const prompt = this.buildTaskPrompt(task, phase);
const queueTask = this.taskQueue.addTask({
prompt,
workingDir: this.workingDir,
priority: 100 - phase.order,
completionPhrase: task.completionPhrase,
timeoutMs: Math.min(task.timeoutMs, this.config.phaseTimeoutMs),
});
task.queueTaskId = queueTask.id;
task.status = 'running';
}
this.assignQueuedTasksToSessions().catch(() => {}); // Best effort
}
// Check completion status of queue tasks
for (const task of phase.tasks) {
if (task.status === 'running' && task.queueTaskId) {
const queueTask = this.taskQueue.getTask(task.queueTaskId);
if (queueTask) {
if (queueTask.isCompleted()) {
this.handleTaskCompleted(task.queueTaskId);
} else if (queueTask.isFailed()) {
this.handleTaskFailed(task.queueTaskId, queueTask.error || 'Task failed');
}
}
}
}
this.checkPhaseCompletion();
}
private checkPhaseCompletion(): void {
if (this._state !== 'executing') return;
const phase = this.getCurrentPhase();
if (!phase) return;
const allDone = phase.tasks.every((t) => t.status === 'completed' || t.status === 'failed');
if (!allDone) return;
const anyFailed = phase.tasks.some((t) => t.status === 'failed');
this.clearPhasePoll();
if (anyFailed) {
// Phase has failed tasks
this.handlePhaseError(phase, 'One or more tasks failed');
} else {
// All tasks completed — run verification after brief delay
this.postPhaseTimer = setTimeout(() => {
this.postPhaseTimer = null;
this.verifyCurrentPhase().catch((err) => this.handleError(err));
}, POST_PHASE_DELAY_MS);
}
}
// ═══════════════════════════════════════════════════════════════
// Internal — Verification
// ═══════════════════════════════════════════════════════════════
private async verifyCurrentPhase(): Promise<void> {
if (!this.plan) return;
const phase = this.plan.phases[this.currentPhaseIndex];
if (!phase) return;
// Skip verification if no criteria defined
if (phase.verificationCriteria.length === 0 && phase.testCommands.length === 0) {
phase.status = 'passed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesCompleted++;
this.persist();
this.emit('phaseCompleted', phase);
await this.advanceToNextPhase();
return;
}
this.setState('verifying');
// Get a session for verification — wait briefly for sessions to become idle
let sessions = this.sessionManager.getIdleSessions();
if (sessions.length === 0) {
// Wait up to 10s for a session to become idle
await new Promise((resolve) => setTimeout(resolve, 10_000));
sessions = this.sessionManager.getIdleSessions();
}
if (sessions.length === 0) {
// Still no sessions — log warning and skip verification (don't silently pass)
console.warn('[Orchestrator] No idle sessions for verification — skipping (marking passed with warning)');
phase.status = 'passed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesCompleted++;
this.persist();
this.emit('phaseCompleted', phase);
this.setState('executing');
await this.advanceToNextPhase();
return;
}
try {
const result = await this.verifier.verifyPhase(phase, sessions[0]);
this.emit('verificationResult', phase, result);
if (result.passed) {
phase.status = 'passed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesCompleted++;
this.persist();
this.emit('phaseCompleted', phase);
this.setState('executing');
await this.advanceToNextPhase();
} else {
// Verification failed — attempt replan
await this.handleVerificationFailure(phase, result);
}
} catch (err) {
// Verification error — treat as pass (don't block on verification bugs)
console.warn('[Orchestrator] Verification error, treating as pass:', err);
phase.status = 'passed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesCompleted++;
this.persist();
this.emit('phaseCompleted', phase);
this.setState('executing');
await this.advanceToNextPhase();
}
}
private async handleVerificationFailure(phase: OrchestratorPhase, result: VerificationResult): Promise<void> {
if (phase.attempts >= phase.maxAttempts) {
// Max retries exceeded
phase.status = 'failed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesFailed++;
this.persist();
this.emit('phaseFailed', phase, `Verification failed after ${phase.attempts} attempts: ${result.summary}`);
this.setState('failed');
return;
}
// Replan and retry
this.stats.replanCount++;
this.setState('replanning');
try {
await this.replanPhase(phase, result);
// Reset task states for retry
for (const task of phase.tasks) {
task.status = 'pending';
task.error = null;
task.assignedSessionId = null;
task.queueTaskId = null;
task.completedAt = null;
task.startedAt = null;
}
phase.status = 'pending';
phase.startedAt = null;
this.persist();
this.setState('executing');
await this.executeCurrentPhase();
} catch (err) {
this.handleError(err);
}
}
private async replanPhase(phase: OrchestratorPhase, result: VerificationResult): Promise<void> {
const completionPhrase = phase.tasks[0]?.completionPhrase || `${phase.id.toUpperCase()}_FIXED`;
const prompt = REPLAN_PROMPT.replace('{PHASE_NAME}', phase.name)
.replace('{ATTEMPT_NUMBER}', String(phase.attempts))
.replace('{MAX_ATTEMPTS}', String(phase.maxAttempts))
.replace('{FAILURE_SUMMARY}', result.summary)
.replace('{SUGGESTIONS}', result.suggestions.join('\n'))
.replace('{ORIGINAL_TASKS}', phase.tasks.map((t, i) => `${i + 1}. ${t.prompt}`).join('\n'))
.replace('{COMPLETION_PHRASE}', completionPhrase);
// Create a tracked queue task for the replan (so completion is detected)
const queueTask = this.taskQueue.addTask({
prompt,
workingDir: this.workingDir,
priority: 100,
completionPhrase,
timeoutMs: this.config.phaseTimeoutMs,
});
// Link to first phase task for tracking
if (phase.tasks[0]) {
phase.tasks[0].queueTaskId = queueTask.id;
phase.tasks[0].status = 'running';
}
this.persist();
// Set up handlers so task completion is tracked
this.setupTaskHandlers();
// Assign to a session
const sessions = this.sessionManager.getIdleSessions();
if (sessions.length === 0) {
console.warn('[Orchestrator] No idle sessions for replan — task queued, will pick up on next poll');
// Start polling so the task gets assigned when a session becomes idle
this.startPhasePoll(phase);
return;
}
try {
queueTask.assign(sessions[0].id);
sessions[0].assignTask(queueTask.id);
this.taskQueue.updateTask(queueTask);
await sessions[0].sendInput(prompt);
} catch (err) {
queueTask.fail(getErrorMessage(err));
this.taskQueue.updateTask(queueTask);
}
}
// ═══════════════════════════════════════════════════════════════
// Internal — State Machine
// ═══════════════════════════════════════════════════════════════
/** Read current state (bypasses TypeScript narrowing from guards) */
private currentState(): OrchestratorState {
return this._state;
}
/** Assert state matches expected or throw */
private requireState(...expected: OrchestratorState[]): void {
if (!expected.includes(this._state)) {
throw new Error(`Expected state "${expected.join('|')}", got "${this._state}"`);
}
}
private setState(newState: OrchestratorState): void {
const prev = this._state;
if (prev === newState) return;
this._state = newState;
this.persist();
this.emit('stateChanged', newState, prev);
}
private async advanceToNextPhase(): Promise<void> {
this.currentPhaseIndex++;
this.phaseSessionIds.clear();
this.persist();
if (!this.plan || this.currentPhaseIndex >= this.plan.phases.length) {
await this.handleCompletion();
} else {
// Compact between phases if configured
if (this.config.compactBetweenPhases) {
const sessions = this.sessionManager.getIdleSessions();
for (const session of sessions) {
try {
await session.writeViaMux('/compact');
} catch {
// Best effort
}
}
// Brief delay for compact to take effect
await new Promise((resolve) => setTimeout(resolve, 2000));
}
await this.executeCurrentPhase();
}
}
private async handleCompletion(): Promise<void> {
this.completedAt = Date.now();
this.stats.totalDurationMs = this.startedAt ? this.completedAt - this.startedAt : 0;
this.clearPhasePoll();
this.cleanupTaskHandlers();
this.setState('completed');
this.emit('completed', this.stats);
}
private handlePhaseError(phase: OrchestratorPhase, error: string): void {
if (phase.attempts >= phase.maxAttempts) {
phase.status = 'failed';
phase.completedAt = Date.now();
phase.durationMs = phase.startedAt ? Date.now() - phase.startedAt : null;
this.stats.phasesFailed++;
this.persist();
this.emit('phaseFailed', phase, error);
this.setState('failed');
} else {
// Retry the phase
for (const task of phase.tasks) {
if (task.status === 'failed') {
task.status = 'pending';
task.error = null;
task.queueTaskId = null;
task.assignedSessionId = null;
}
}
phase.status = 'pending';
this.persist();
this.executeCurrentPhase().catch((err) => this.handleError(err));
}
}
private handleError(err: unknown): void {
const error = err instanceof Error ? err : new Error(getErrorMessage(err));
console.error('[Orchestrator] Error:', error.message);
this.setState('failed');
this.emit('error', error);
}
// ═══════════════════════════════════════════════════════════════
// Internal — Persistence
// ═══════════════════════════════════════════════════════════════
private persist(): void {
this.store.setOrchestratorState(this.getStatus());
}
private restore(): void {
const saved = this.store.getOrchestratorState();
if (!saved) return;
// If we crashed while running, reset to failed
if (saved.state === 'executing' || saved.state === 'verifying' || saved.state === 'replanning') {
this._state = 'failed';
this.plan = saved.plan;
this.currentPhaseIndex = saved.currentPhaseIndex;
this.startedAt = saved.startedAt;
this.config = saved.config;
this.stats = saved.stats;
this.store.setOrchestratorState({ ...saved, state: 'failed' });
} else if (saved.state === 'planning' || saved.state === 'approval') {
// Planning/approval — reset to idle (plan is lost)
this.store.clearOrchestratorState();
} else if (saved.state === 'completed' || saved.state === 'failed') {
// Preserve completed/failed state for UI display
this._state = saved.state;
this.plan = saved.plan;
this.currentPhaseIndex = saved.currentPhaseIndex;
this.startedAt = saved.startedAt;
this.completedAt = saved.completedAt;
this.config = saved.config;
this.stats = saved.stats;
}
}
private reset(): void {
this._state = 'idle';
this.plan = null;
this.currentPhaseIndex = 0;
this.startedAt = null;
this.completedAt = null;
this.stats = createInitialOrchestratorStats();
this.pausedState = null;
this.phaseSessionIds.clear();
this.clearPhasePoll();
this.cleanupTaskHandlers();
}
// ═══════════════════════════════════════════════════════════════
// Internal — Helpers
// ═══════════════════════════════════════════════════════════════
private buildTaskPrompt(task: OrchestratorTask, phase: OrchestratorPhase): string {
if (phase.tasks.length === 1) {
// Single task — use simpler prompt
const completedPhases = this.getCompletedPhasesSummary();
return SINGLE_TASK_PROMPT.replace('{TASK}', task.prompt)
.replace('{GOAL}', this.plan?.goal || '')
.replace('{CONTEXT}', completedPhases ? `Previous phases completed: ${completedPhases}` : '')
.replace('{COMPLETION_PHRASE}', task.completionPhrase);
}
// Multi-task phase — use full prompt
return PHASE_EXECUTION_PROMPT.replace('{PHASE_NAME}', phase.name)
.replace('{GOAL}', this.plan?.goal || '')
.replace('{COMPLETED_PHASES}', this.getCompletedPhasesSummary() || 'None yet')
.replace('{TASK_LIST}', phase.tasks.map((t, i) => `${i + 1}. ${t.prompt}`).join('\n'))
.replace('{VERIFICATION_CRITERIA}', phase.verificationCriteria.join('\n') || 'No specific criteria')
.replace('{COMPLETION_PHRASE}', task.completionPhrase);
}
private getCompletedPhasesSummary(): string {
if (!this.plan) return '';
return this.plan.phases
.filter((p) => p.status === 'passed' || p.status === 'skipped')
.map((p) => `${p.name}: ${p.status}`)
.join(', ');
}
private findOrchestratorTaskByQueueId(queueTaskId: string): OrchestratorTask | null {
if (!this.plan) return null;
for (const phase of this.plan.phases) {
for (const task of phase.tasks) {
if (task.queueTaskId === queueTaskId) return task;
}
}
return null;
}
/** Clean up resources when the loop is being destroyed. */
destroy(): void {
this.clearPhasePoll();
this.cleanupTaskHandlers();
}
}
+412
View File
@@ -0,0 +1,412 @@
/**
* @fileoverview Orchestrator plan generation — converts goals into phased plans.
*
* Wraps PlanOrchestrator for AI-powered plan generation, then groups the
* resulting PlanItems into sequential phases with team strategies and
* verification criteria.
*
* Phase grouping algorithm:
* 1. Topological sort by dependencies (Kahn's algorithm)
* 2. Group into dependency layers
* 3. Sub-group by TDD phase within layers
* 4. Merge small adjacent phases
* 5. Assign team strategies based on parallelism potential
*
* Key exports:
* - `OrchestratorPlanner` class — plan generation + phase grouping
*
* @dependencies plan-orchestrator (AI plan generation), types (OrchestratorPlan, PlanItem)
* @consumedby orchestrator-loop
*
* @module orchestrator-planner
*/
import { v4 as uuidv4 } from 'uuid';
import { PlanOrchestrator, type DetailedPlanResult, type ProgressCallback } from './plan-orchestrator.js';
import type { TerminalMultiplexer } from './mux-interface.js';
import type {
PlanItem,
TddPhase,
OrchestratorPlan,
OrchestratorPhase,
OrchestratorTask,
OrchestratorConfig,
TeamStrategy,
PhaseStatus,
} from './types.js';
// ═══════════════════════════════════════════════════════════════
// Constants
// ═══════════════════════════════════════════════════════════════
/** Maximum number of phases (prevents runaway plans) */
const MAX_PHASES = 10;
/** Maximum total tasks across all phases */
const MAX_TOTAL_TASKS = 50;
/** Default task timeout (10 minutes) */
const DEFAULT_TASK_TIMEOUT_MS = 10 * 60 * 1000;
/** Minimum tasks in a phase before it gets merged with adjacent */
const MIN_PHASE_TASKS = 2;
/** TDD phase ordering for grouping */
const TDD_PHASE_ORDER: Record<TddPhase, number> = {
setup: 0,
test: 1,
impl: 2,
verify: 3,
review: 4,
};
// ═══════════════════════════════════════════════════════════════
// OrchestratorPlanner
// ═══════════════════════════════════════════════════════════════
export class OrchestratorPlanner {
private mux: TerminalMultiplexer;
private workingDir: string;
private config: OrchestratorConfig;
private orchestrator: PlanOrchestrator | null = null;
constructor(mux: TerminalMultiplexer, workingDir: string, config: OrchestratorConfig) {
this.mux = mux;
this.workingDir = workingDir;
this.config = config;
}
/**
* Generate a phased plan from a user goal.
*
* Uses PlanOrchestrator for AI plan generation, then groups results into phases.
*/
async generatePlan(goal: string, onProgress?: ProgressCallback): Promise<OrchestratorPlan> {
const startTime = Date.now();
// Create a PlanOrchestrator for this plan generation
this.orchestrator = new PlanOrchestrator(this.mux, this.workingDir, undefined, {
defaultModel: this.config.plannerModel,
});
try {
onProgress?.('planning', 'Generating detailed plan...');
const result: DetailedPlanResult = await this.orchestrator.generateDetailedPlan(goal, onProgress);
if (!result.success || !result.items || result.items.length === 0) {
throw new Error(result.error || 'Plan generation returned no items');
}
// Cap total tasks
const items = result.items.slice(0, MAX_TOTAL_TASKS);
onProgress?.('grouping', 'Organizing plan into phases...');
// Group items into phases
const phases = this.groupIntoPhases(items, goal);
// Assign team strategies
this.assignTeamStrategies(phases);
// Generate unique completion phrases
this.generateCompletionPhrases(phases);
const plan: OrchestratorPlan = {
id: uuidv4(),
goal,
createdAt: Date.now(),
phases,
metadata: {
totalTasks: phases.reduce((sum, p) => sum + p.tasks.length, 0),
estimatedComplexity: this.estimateComplexity(items),
modelUsed: this.config.plannerModel,
planDurationMs: Date.now() - startTime,
},
};
return plan;
} finally {
this.orchestrator = null;
}
}
/** Cancel in-progress plan generation. */
async cancel(): Promise<void> {
if (this.orchestrator) {
await this.orchestrator.cancel();
this.orchestrator = null;
}
}
// ═══════════════════════════════════════════════════════════════
// Phase Grouping
// ═══════════════════════════════════════════════════════════════
/**
* Group PlanItems into sequential phases.
*
* Algorithm:
* 1. Build dependency graph and assign IDs to items without them
* 2. Topological sort into dependency layers (Kahn's algorithm)
* 3. Sub-group within each layer by TDD phase
* 4. Merge small phases with their neighbors
*/
private groupIntoPhases(items: PlanItem[], _goal: string): OrchestratorPhase[] {
// Ensure all items have IDs
const indexedItems = items.map((item, i) => ({
...item,
id: item.id || `task-${i}`,
}));
// Build adjacency and in-degree for Kahn's algorithm
const idSet = new Set(indexedItems.map((item) => item.id!));
const inDegree = new Map<string, number>();
const dependents = new Map<string, string[]>(); // id → items that depend on it
for (const item of indexedItems) {
inDegree.set(item.id!, 0);
dependents.set(item.id!, []);
}
for (const item of indexedItems) {
const deps = (item.dependencies || []).filter((d) => idSet.has(d));
inDegree.set(item.id!, deps.length);
for (const dep of deps) {
dependents.get(dep)!.push(item.id!);
}
}
// Kahn's algorithm — produce dependency layers
const layers: PlanItem[][] = [];
const remaining = new Set(indexedItems.map((item) => item.id!));
while (remaining.size > 0) {
// Find items with no remaining dependencies (in-degree 0)
const layer: PlanItem[] = [];
for (const id of remaining) {
if (inDegree.get(id)! === 0) {
layer.push(indexedItems.find((item) => item.id === id)!);
}
}
if (layer.length === 0) {
// Circular dependency — add all remaining items as a single layer
for (const id of remaining) {
layer.push(indexedItems.find((item) => item.id === id)!);
}
}
layers.push(layer);
// Remove this layer's items and update in-degrees
for (const item of layer) {
remaining.delete(item.id!);
for (const dep of dependents.get(item.id!) || []) {
if (remaining.has(dep)) {
inDegree.set(dep, Math.max(0, inDegree.get(dep)! - 1));
}
}
}
}
// Sub-group each layer by TDD phase
const rawPhases: PlanItem[][] = [];
for (const layer of layers) {
const byPhase = new Map<string, PlanItem[]>();
for (const item of layer) {
const phase = item.tddPhase || 'impl';
if (!byPhase.has(phase)) byPhase.set(phase, []);
byPhase.get(phase)!.push(item);
}
// Sort sub-groups by TDD phase order
const sorted = [...byPhase.entries()].sort(
([a], [b]) => (TDD_PHASE_ORDER[a as TddPhase] ?? 2) - (TDD_PHASE_ORDER[b as TddPhase] ?? 2)
);
for (const [, items] of sorted) {
rawPhases.push(items);
}
}
// Merge small phases with their previous neighbor
const mergedPhases: PlanItem[][] = [];
for (const phase of rawPhases) {
if (mergedPhases.length > 0 && phase.length < MIN_PHASE_TASKS) {
const prev = mergedPhases[mergedPhases.length - 1];
if (prev.length < MIN_PHASE_TASKS) {
// Merge with previous
prev.push(...phase);
continue;
}
}
mergedPhases.push([...phase]);
}
// Cap at MAX_PHASES by merging tail phases
while (mergedPhases.length > MAX_PHASES) {
const last = mergedPhases.pop()!;
mergedPhases[mergedPhases.length - 1].push(...last);
}
// Convert to OrchestratorPhase objects
return mergedPhases.map((phaseItems, index) => this.createPhase(phaseItems, index));
}
private createPhase(items: PlanItem[], order: number): OrchestratorPhase {
// Derive phase name from TDD phases and priorities
const tddPhases = [...new Set(items.map((i) => i.tddPhase).filter(Boolean))];
const name = this.generatePhaseName(items, tddPhases as TddPhase[], order);
const description = items.map((i) => i.content).join('; ');
const tasks: OrchestratorTask[] = items.map((item, i) => ({
id: `phase-${order + 1}-task-${i + 1}`,
phaseId: `phase-${order + 1}`,
prompt: item.content,
status: 'pending' as const,
assignedSessionId: null,
queueTaskId: null,
parallel: items.length > 1, // Tasks within a phase are parallel by default
completionPhrase: '', // Assigned later
timeoutMs: DEFAULT_TASK_TIMEOUT_MS,
startedAt: null,
completedAt: null,
error: null,
retries: 0,
}));
// Extract verification criteria and test commands from items
const verificationCriteria = items
.map((i) => i.verificationCriteria)
.filter((v): v is string => v != null && v.length > 0);
const testCommands = items.map((i) => i.testCommand).filter((t): t is string => t != null && t.length > 0);
return {
id: `phase-${order + 1}`,
name,
description,
order,
status: 'pending' as PhaseStatus,
tasks,
verificationCriteria,
testCommands,
maxAttempts: this.config.maxPhaseRetries,
attempts: 0,
startedAt: null,
completedAt: null,
durationMs: null,
teamStrategy: { type: 'single' }, // Assigned later
};
}
private generatePhaseName(items: PlanItem[], tddPhases: TddPhase[], order: number): string {
// Try to create a meaningful name based on content
const priorities = [...new Set(items.map((i) => i.priority).filter(Boolean))];
if (tddPhases.length === 1) {
const phaseNames: Record<TddPhase, string> = {
setup: 'Setup & Configuration',
test: 'Test Definition',
impl: 'Implementation',
verify: 'Verification',
review: 'Review & Polish',
};
return `Phase ${order + 1}: ${phaseNames[tddPhases[0]]}`;
}
if (priorities.includes('P0') && priorities.length === 1) {
return `Phase ${order + 1}: Critical Foundation`;
}
return `Phase ${order + 1}: ${items.length > 1 ? 'Parallel Tasks' : items[0].content.slice(0, 50)}`;
}
// ═══════════════════════════════════════════════════════════════
// Team Strategy Assignment
// ═══════════════════════════════════════════════════════════════
private assignTeamStrategies(phases: OrchestratorPhase[]): void {
for (const phase of phases) {
phase.teamStrategy = this.computeTeamStrategy(phase);
}
}
private computeTeamStrategy(phase: OrchestratorPhase): TeamStrategy {
const taskCount = phase.tasks.length;
const parallelTasks = phase.tasks.filter((t) => t.parallel).length;
// Single task or no parallel potential → single session
if (taskCount <= 2 || parallelTasks <= 1) {
return { type: 'single' };
}
// If team agents are disabled, use parallel sessions instead
if (!this.config.enableTeamAgents) {
return {
type: 'parallel',
maxSessions: Math.min(parallelTasks, this.config.maxParallelSessions),
};
}
// 4+ parallel tasks with team agents enabled → team mode
if (parallelTasks >= 4) {
return {
type: 'team',
config: {
leadPrompt: this.buildTeamLeadPrompt(phase),
suggestedTeammates: phase.tasks.slice(0, 4).map((t) => `Specialist for: ${t.prompt.slice(0, 80)}`),
maxTeammates: Math.min(parallelTasks, 4),
},
};
}
// 3 parallel tasks → parallel sessions
return {
type: 'parallel',
maxSessions: Math.min(parallelTasks, this.config.maxParallelSessions),
};
}
private buildTeamLeadPrompt(phase: OrchestratorPhase): string {
const taskList = phase.tasks.map((t, i) => `${i + 1}. ${t.prompt}`).join('\n');
return [
`You are the team lead for "${phase.name}".`,
`Create teammates and delegate the following tasks for parallel execution:`,
'',
taskList,
'',
`Each teammate should focus on one task area.`,
`When all tasks are complete, verify the results and output: <promise>${phase.id.toUpperCase()}_COMPLETE</promise>`,
].join('\n');
}
// ═══════════════════════════════════════════════════════════════
// Completion Phrases
// ═══════════════════════════════════════════════════════════════
private generateCompletionPhrases(phases: OrchestratorPhase[]): void {
for (const phase of phases) {
for (const task of phase.tasks) {
// Generate a unique, deterministic completion phrase per task
task.completionPhrase = `ORCH_P${phase.order + 1}_T${phase.tasks.indexOf(task) + 1}`;
}
}
}
// ═══════════════════════════════════════════════════════════════
// Helpers
// ═══════════════════════════════════════════════════════════════
private estimateComplexity(items: PlanItem[]): 'low' | 'medium' | 'high' {
const total = items.length;
const highComplexity = items.filter((i) => i.complexity === 'high').length;
const p0Count = items.filter((i) => i.priority === 'P0').length;
if (total > 20 || highComplexity > 5 || p0Count > 8) return 'high';
if (total > 10 || highComplexity > 2 || p0Count > 4) return 'medium';
return 'low';
}
}
+298
View File
@@ -0,0 +1,298 @@
/**
* @fileoverview Orchestrator phase verification.
*
* Runs verification checks after each phase completes:
* - Test commands (shell commands via session)
* - AI review (ask Claude to evaluate phase results)
*
* Three verification modes:
* - strict: ALL test commands must pass AND AI review must approve
* - moderate: Test commands must pass, AI review is advisory
* - lenient: At least one test command passes, AI review skipped
*
* Key exports:
* - `OrchestratorVerifier` class — phase verification engine
*
* @dependencies types (OrchestratorPhase, VerificationResult, VerificationCheck, OrchestratorConfig)
* @consumedby orchestrator-loop
*
* @module orchestrator-verifier
*/
import type { Session } from './session.js';
import {
getErrorMessage,
type OrchestratorPhase,
type OrchestratorConfig,
type VerificationResult,
type VerificationCheck,
} from './types.js';
// ═══════════════════════════════════════════════════════════════
// Constants
// ═══════════════════════════════════════════════════════════════
/** Timeout for individual test command execution (2 minutes) */
const TEST_COMMAND_TIMEOUT_MS = 2 * 60 * 1000;
/** Timeout for AI review (3 minutes) */
const AI_REVIEW_TIMEOUT_MS = 3 * 60 * 1000;
/** Completion phrase for AI verification pass */
const VERIFY_PASS_PHRASE = 'ORCH_VERIFY_PASS';
/** Completion phrase for AI verification fail */
const VERIFY_FAIL_PHRASE = 'ORCH_VERIFY_FAIL';
// ═══════════════════════════════════════════════════════════════
// OrchestratorVerifier
// ═══════════════════════════════════════════════════════════════
export class OrchestratorVerifier {
private config: OrchestratorConfig;
constructor(config: OrchestratorConfig) {
this.config = config;
}
/**
* Run all verification checks for a completed phase.
*
* @param phase - The phase to verify
* @param session - Session to use for running commands/reviews
* @returns Verification result with pass/fail and suggestions
*/
async verifyPhase(phase: OrchestratorPhase, session: Session): Promise<VerificationResult> {
const checks: VerificationCheck[] = [];
const mode = this.config.verificationMode;
// Skip verification entirely in lenient mode with no test commands
if (mode === 'lenient' && phase.testCommands.length === 0 && phase.verificationCriteria.length === 0) {
return {
passed: true,
checks: [],
summary: 'Verification skipped (lenient mode, no checks defined)',
suggestions: [],
};
}
// Run test commands if any are defined
if (phase.testCommands.length > 0) {
const testChecks = await this.runTestCommands(phase.testCommands, session);
checks.push(...testChecks);
}
// Run AI review in strict and moderate modes
if (mode !== 'lenient' && phase.verificationCriteria.length > 0) {
const aiCheck = await this.aiReview(phase, session);
checks.push(aiCheck);
}
// Determine pass/fail based on mode
const passed = this.evaluateChecks(checks, mode);
// Generate suggestions for failed checks
const suggestions = this.generateSuggestions(checks, phase);
const passedCount = checks.filter((c) => c.passed).length;
const summary =
checks.length === 0 ? 'No verification checks defined' : `${passedCount}/${checks.length} checks passed`;
return { passed, checks, summary, suggestions };
}
// ═══════════════════════════════════════════════════════════════
// Test Command Execution
// ═══════════════════════════════════════════════════════════════
private async runTestCommands(commands: string[], session: Session): Promise<VerificationCheck[]> {
const checks: VerificationCheck[] = [];
for (const command of commands) {
try {
const check = await this.runSingleTestCommand(command, session);
checks.push(check);
} catch (err) {
checks.push({
type: 'test_command',
description: `Run: ${command}`,
passed: false,
output: getErrorMessage(err),
});
}
}
return checks;
}
private async runSingleTestCommand(command: string, session: Session): Promise<VerificationCheck> {
// Send the test command to the session and wait for completion
// We use a unique marker to detect when the command finishes
const marker = `ORCH_TEST_${Date.now()}`;
const wrappedCommand = `${command} && echo ${marker}_PASS || echo ${marker}_FAIL`;
const result = await this.sendAndWaitForMarker(session, wrappedCommand, marker, TEST_COMMAND_TIMEOUT_MS);
return {
type: 'test_command',
description: `Run: ${command}`,
passed: result.includes(`${marker}_PASS`),
output: result.slice(0, 2000), // Truncate output
};
}
// ═══════════════════════════════════════════════════════════════
// AI Review
// ═══════════════════════════════════════════════════════════════
private async aiReview(phase: OrchestratorPhase, session: Session): Promise<VerificationCheck> {
const prompt = this.buildVerificationPrompt(phase);
try {
const result = await this.sendAndWaitForMarker(
session,
prompt,
VERIFY_PASS_PHRASE,
AI_REVIEW_TIMEOUT_MS,
VERIFY_FAIL_PHRASE
);
const passed = result.includes(VERIFY_PASS_PHRASE);
return {
type: 'ai_review',
description: `AI review of "${phase.name}"`,
passed,
output: result.slice(0, 3000),
};
} catch (err) {
return {
type: 'ai_review',
description: `AI review of "${phase.name}"`,
passed: false,
output: `AI review timed out or failed: ${getErrorMessage(err)}`,
};
}
}
private buildVerificationPrompt(phase: OrchestratorPhase): string {
const criteria = phase.verificationCriteria.map((c, i) => `${i + 1}. ${c}`).join('\n');
return [
`Review the work done in "${phase.name}". Check these criteria:`,
'',
criteria,
'',
`If ALL criteria are met, respond with: ${VERIFY_PASS_PHRASE}`,
`If ANY criteria fail, respond with: ${VERIFY_FAIL_PHRASE} and explain what failed.`,
].join('\n');
}
// ═══════════════════════════════════════════════════════════════
// Evaluation
// ═══════════════════════════════════════════════════════════════
private evaluateChecks(checks: VerificationCheck[], mode: OrchestratorConfig['verificationMode']): boolean {
if (checks.length === 0) return true;
const testChecks = checks.filter((c) => c.type === 'test_command');
const aiChecks = checks.filter((c) => c.type === 'ai_review');
switch (mode) {
case 'strict':
// ALL checks must pass
return checks.every((c) => c.passed);
case 'moderate':
// All test commands must pass; AI review is advisory
return testChecks.length === 0 || testChecks.every((c) => c.passed);
case 'lenient':
// At least one test passes (AI review skipped in lenient mode)
return testChecks.length === 0 || testChecks.some((c) => c.passed);
default:
return aiChecks.every((c) => c.passed) && testChecks.every((c) => c.passed);
}
}
private generateSuggestions(checks: VerificationCheck[], phase: OrchestratorPhase): string[] {
const suggestions: string[] = [];
const failedChecks = checks.filter((c) => !c.passed);
if (failedChecks.length === 0) return suggestions;
for (const check of failedChecks) {
if (check.type === 'test_command') {
suggestions.push(`Fix failing test: ${check.description}`);
} else if (check.type === 'ai_review' && check.output) {
// Extract failure reasons from AI review output
suggestions.push(`Address AI review feedback for "${phase.name}"`);
}
}
return suggestions;
}
// ═══════════════════════════════════════════════════════════════
// Session Communication
// ═══════════════════════════════════════════════════════════════
/**
* Send a prompt to a session and wait for a marker phrase in the output.
*
* @param session - Session to send to
* @param input - Prompt/command to send
* @param marker - Primary marker to watch for
* @param timeoutMs - Maximum wait time
* @param altMarker - Alternative marker (for pass/fail detection)
* @returns Captured output containing the marker
*/
private sendAndWaitForMarker(
session: Session,
input: string,
marker: string,
timeoutMs: number,
altMarker?: string
): Promise<string> {
return new Promise<string>((resolve, reject) => {
let output = '';
let resolved = false;
const timer = setTimeout(() => {
if (!resolved) {
resolved = true;
cleanup();
reject(new Error(`Timeout waiting for marker "${marker}" after ${timeoutMs}ms`));
}
}, timeoutMs);
const handler = (data: string) => {
if (resolved) return;
output += data;
if (output.includes(marker) || (altMarker && output.includes(altMarker))) {
resolved = true;
cleanup();
resolve(output);
}
};
const cleanup = () => {
clearTimeout(timer);
session.off('terminal', handler);
};
session.on('terminal', handler);
// Send the input
session.sendInput(input).catch((err) => {
if (!resolved) {
resolved = true;
cleanup();
reject(err);
}
});
});
}
}
+99 -93
View File
@@ -20,7 +20,7 @@ import type { TerminalMultiplexer } from './mux-interface.js';
import { existsSync, mkdirSync, writeFileSync } from 'node:fs';
import { join } from 'node:path';
import { RESEARCH_AGENT_PROMPT, PLANNER_PROMPT } from './prompts/index.js';
import type { PlanItem } from './types.js';
import { getErrorMessage, type PlanItem } from './types.js';
// Re-export for backward compatibility
export type { PlanItem };
@@ -68,7 +68,7 @@ export interface DetailedPlanResult {
export type ProgressCallback = (phase: string, detail: string) => void;
export interface PlanSubagentEvent {
interface PlanSubagentEvent {
type: 'started' | 'progress' | 'completed' | 'failed';
agentId: string;
agentType: 'research' | 'planner';
@@ -80,7 +80,7 @@ export interface PlanSubagentEvent {
error?: string;
}
export type SubagentCallback = (event: PlanSubagentEvent) => void;
type SubagentCallback = (event: PlanSubagentEvent) => void;
// ============================================================================
// JSON Repair Helper
@@ -231,6 +231,49 @@ export class PlanOrchestrator {
return md;
}
private _extractJsonFromResponse(response: string): string | null {
let jsonMatch = response.match(/```(?:json)?\s*(\{[\s\S]*?\})\s*```/);
if (jsonMatch) {
jsonMatch = [jsonMatch[1]]; // Use captured group (inside code block)
} else {
jsonMatch = response.match(/\{[\s\S]*\}/);
}
return jsonMatch ? jsonMatch[0] : null;
}
private _emitAgentFailure(
onSubagent: SubagentCallback | undefined,
agentId: string,
agentType: 'research' | 'planner',
model: string,
error: string,
durationMs: number
): void {
onSubagent?.({
type: 'failed',
agentId,
agentType,
model,
status: 'failed',
error,
durationMs,
});
}
private _formatResearchSection(
parts: string[],
title: string,
items: unknown[],
formatter: (item: unknown) => string[]
): void {
if (items.length === 0) return;
parts.push(title);
for (const item of items.slice(0, 5)) {
parts.push(...formatter(item));
}
parts.push('');
}
async cancel(): Promise<void> {
this.cancelled = true;
// Stop all running sessions and await cleanup to prevent PTY process leaks
@@ -312,7 +355,7 @@ export class PlanOrchestrator {
} catch (err) {
return {
success: false,
error: err instanceof Error ? err.message : String(err),
error: getErrorMessage(err),
};
}
}
@@ -322,32 +365,23 @@ export class PlanOrchestrator {
const parts: string[] = ['## Research Context\n'];
if (research.findings.externalResources.length > 0) {
parts.push('### External Resources');
for (const r of research.findings.externalResources.slice(0, 5)) {
parts.push(`- ${r.title}${r.url ? ` (${r.url})` : ''}`);
if (r.keyInsights.length > 0) {
parts.push(` Key insights: ${r.keyInsights.slice(0, 3).join(', ')}`);
}
this._formatResearchSection(parts, '### External Resources', research.findings.externalResources, (item) => {
const r = item as ResearchResult['findings']['externalResources'][number];
const lines = [`- ${r.title}${r.url ? ` (${r.url})` : ''}`];
if (r.keyInsights.length > 0) {
lines.push(` Key insights: ${r.keyInsights.slice(0, 3).join(', ')}`);
}
parts.push('');
}
return lines;
});
if (research.findings.codebasePatterns.length > 0) {
parts.push('### Existing Codebase Patterns');
for (const p of research.findings.codebasePatterns.slice(0, 5)) {
parts.push(`- ${p.pattern} at ${p.location}`);
}
parts.push('');
}
this._formatResearchSection(parts, '### Existing Codebase Patterns', research.findings.codebasePatterns, (item) => {
const p = item as ResearchResult['findings']['codebasePatterns'][number];
return [`- ${p.pattern} at ${p.location}`];
});
if (research.findings.technicalRecommendations.length > 0) {
parts.push('### Recommendations');
for (const r of research.findings.technicalRecommendations.slice(0, 5)) {
parts.push(`- ${r}`);
}
parts.push('');
}
this._formatResearchSection(parts, '### Recommendations', research.findings.technicalRecommendations, (item) => [
`- ${item as string}`,
]);
return parts.join('\n');
}
@@ -414,18 +448,20 @@ export class PlanOrchestrator {
const durationMs = Date.now() - startTime;
// Extract JSON from response
const jsonMatch = response.match(/\{[\s\S]*\}/);
if (!jsonMatch) {
onSubagent?.({
type: 'failed',
agentId,
agentType: 'research',
model: this.researchModel,
status: 'failed',
error: 'No JSON found',
durationMs,
});
console.log(
`[PlanOrchestrator] Research response length: ${response.length}, first 500 chars:`,
response.substring(0, 500)
);
// Extract JSON from response — try multiple strategies
const jsonStr = this._extractJsonFromResponse(response);
if (!jsonStr) {
console.error(
`[PlanOrchestrator] No JSON found in research response. Full response:`,
response.substring(0, 2000)
);
this._emitAgentFailure(onSubagent, agentId, 'research', this.researchModel, 'No JSON found', durationMs);
return {
success: false,
findings: {
@@ -441,17 +477,9 @@ export class PlanOrchestrator {
};
}
const parsed = tryParseJSON(jsonMatch[0]);
const parsed = tryParseJSON(jsonStr);
if (!parsed.success) {
onSubagent?.({
type: 'failed',
agentId,
agentType: 'research',
model: this.researchModel,
status: 'failed',
error: parsed.error,
durationMs,
});
this._emitAgentFailure(onSubagent, agentId, 'research', this.researchModel, parsed.error!, durationMs);
return {
success: false,
findings: {
@@ -495,16 +523,8 @@ export class PlanOrchestrator {
return result;
} catch (err) {
const durationMs = Date.now() - startTime;
const error = err instanceof Error ? err.message : String(err);
onSubagent?.({
type: 'failed',
agentId,
agentType: 'research',
model: this.researchModel,
status: 'failed',
error,
durationMs,
});
const error = getErrorMessage(err);
this._emitAgentFailure(onSubagent, agentId, 'research', this.researchModel, error, durationMs);
return {
success: false,
findings: {
@@ -521,7 +541,7 @@ export class PlanOrchestrator {
} finally {
// Always clean up session and progress interval — centralizing here
// prevents the race where cancel() and catch both try to manage the set
await session.stop().catch(() => {});
await session.stop().catch(() => {}); // Ignore - session cleanup is best-effort in finally block
this.runningSessions.delete(session);
clearInterval(progressInterval);
}
@@ -587,32 +607,26 @@ export class PlanOrchestrator {
const durationMs = Date.now() - startTime;
// Extract JSON from response
const jsonMatch = response.match(/\{[\s\S]*\}/);
if (!jsonMatch) {
onSubagent?.({
type: 'failed',
agentId,
agentType: 'planner',
model: this.plannerModel,
status: 'failed',
error: 'No JSON found',
durationMs,
});
console.log(
`[PlanOrchestrator] Planner response length: ${response.length}, first 500 chars:`,
response.substring(0, 500)
);
// Extract JSON from response — try multiple strategies
const jsonStr = this._extractJsonFromResponse(response);
if (!jsonStr) {
console.error(
`[PlanOrchestrator] No JSON found in planner response. Full response:`,
response.substring(0, 2000)
);
this._emitAgentFailure(onSubagent, agentId, 'planner', this.plannerModel, 'No JSON found', durationMs);
return { success: false, error: 'No JSON in response' };
}
const parsed = tryParseJSON(jsonMatch[0]);
const parsed = tryParseJSON(jsonStr);
if (!parsed.success) {
onSubagent?.({
type: 'failed',
agentId,
agentType: 'planner',
model: this.plannerModel,
status: 'failed',
error: parsed.error,
durationMs,
});
this._emitAgentFailure(onSubagent, agentId, 'planner', this.plannerModel, parsed.error!, durationMs);
return { success: false, error: parsed.error };
}
@@ -637,21 +651,13 @@ export class PlanOrchestrator {
return { success: true, items, gaps, warnings };
} catch (err) {
const durationMs = Date.now() - startTime;
const error = err instanceof Error ? err.message : String(err);
onSubagent?.({
type: 'failed',
agentId,
agentType: 'planner',
model: this.plannerModel,
status: 'failed',
error,
durationMs,
});
const error = getErrorMessage(err);
this._emitAgentFailure(onSubagent, agentId, 'planner', this.plannerModel, error, durationMs);
return { success: false, error };
} finally {
// Always clean up session and progress interval — centralizing here
// prevents the race where cancel() and catch both try to manage the set
await session.stop().catch(() => {});
await session.stop().catch(() => {}); // Ignore - session cleanup is best-effort in finally block
this.runningSessions.delete(session);
clearInterval(progressInterval);
}
+1
View File
@@ -7,3 +7,4 @@
export { RESEARCH_AGENT_PROMPT } from './research-agent.js';
export { PLANNER_PROMPT } from './planner.js';
export { PHASE_EXECUTION_PROMPT, TEAM_LEAD_PROMPT, REPLAN_PROMPT, SINGLE_TASK_PROMPT } from './orchestrator.js';
+100
View File
@@ -0,0 +1,100 @@
/**
* @fileoverview Orchestrator Loop prompt templates.
*
* Templates for phase execution, team delegation, verification, and replanning.
* Placeholders use {VARIABLE} syntax and are replaced at runtime.
*
* @module prompts/orchestrator
*/
/**
* Phase execution prompt — tells Claude what to accomplish in this phase.
*
* Placeholders:
* - {PHASE_NUMBER}: Phase index (1-based)
* - {PHASE_NAME}: Human-readable phase name
* - {GOAL}: Original user goal
* - {COMPLETED_PHASES}: Summary of previously completed phases
* - {TASK_LIST}: Numbered task list for this phase
* - {VERIFICATION_CRITERIA}: What will be checked after this phase
* - {COMPLETION_PHRASE}: The phrase to output when done
*/
export const PHASE_EXECUTION_PROMPT = `You are executing {PHASE_NAME} of a larger project.
OVERALL GOAL: {GOAL}
COMPLETED SO FAR:
{COMPLETED_PHASES}
YOUR TASKS FOR THIS PHASE:
{TASK_LIST}
Complete each task thoroughly. Run tests after each change to catch issues early.
VERIFICATION (will be checked after you finish):
{VERIFICATION_CRITERIA}
When ALL tasks in this phase are complete and verified, output: <promise>{COMPLETION_PHRASE}</promise>`;
/**
* Team lead delegation prompt — instructs a lead to coordinate teammates.
*
* Placeholders:
* - {PHASE_NAME}: Phase name
* - {TASK_LIST}: Numbered task list
* - {TEAMMATE_HINTS}: Suggested teammate specializations
* - {COMPLETION_PHRASE}: Phrase for when all work is done
*/
export const TEAM_LEAD_PROMPT = `You are the team lead for {PHASE_NAME}.
Create teammates and delegate the following tasks for parallel execution:
{TASK_LIST}
Suggested teammate roles:
{TEAMMATE_HINTS}
Each teammate should focus on their assigned task area. Monitor their progress.
When ALL tasks are complete and you've verified the results, output: <promise>{COMPLETION_PHRASE}</promise>`;
/**
* Replan prompt — gives failure context and asks for recovery.
*
* Placeholders:
* - {PHASE_NAME}: Phase name
* - {ATTEMPT_NUMBER}: Current retry attempt
* - {MAX_ATTEMPTS}: Maximum attempts allowed
* - {FAILURE_SUMMARY}: What went wrong
* - {SUGGESTIONS}: Recovery suggestions from verification
* - {ORIGINAL_TASKS}: The original task list
* - {COMPLETION_PHRASE}: Phrase for when recovery is done
*/
export const REPLAN_PROMPT = `Phase "{PHASE_NAME}" verification failed (attempt {ATTEMPT_NUMBER}/{MAX_ATTEMPTS}).
WHAT WENT WRONG:
{FAILURE_SUMMARY}
SUGGESTIONS:
{SUGGESTIONS}
ORIGINAL TASKS:
{ORIGINAL_TASKS}
Fix the issues identified above. Focus on making the verification criteria pass.
When the fixes are complete, output: <promise>{COMPLETION_PHRASE}</promise>`;
/**
* Single-task execution prompt — for phases with a single task.
*
* Placeholders:
* - {TASK}: The task description
* - {GOAL}: Original user goal
* - {CONTEXT}: Any relevant context
* - {COMPLETION_PHRASE}: Phrase for when done
*/
export const SINGLE_TASK_PROMPT = `{TASK}
Context: This is part of a larger project — {GOAL}
{CONTEXT}
When done, output: <promise>{COMPLETION_PHRASE}</promise>`;
+1 -1
View File
@@ -24,7 +24,7 @@ const YAML_LINE_PATTERN = /^([a-zA-Z_-]+):\s*"?([^"\n]+)"?\s*$/gm;
/**
* Ralph Loop configuration from .claude/ralph-loop.local.md
*/
export interface RalphLoopConfig {
interface RalphLoopConfig {
enabled: boolean;
iteration: number;
maxIterations: number | null;
+1 -9
View File
@@ -34,19 +34,11 @@ import { RalphLoopStatus, getErrorMessage } from './types.js';
/**
* Events emitted by RalphLoop
*/
export interface RalphLoopEvents {
started: () => void;
stopped: () => void;
taskAssigned: (taskId: string, sessionId: string) => void;
taskCompleted: (taskId: string) => void;
taskFailed: (taskId: string, error: string) => void;
error: (error: Error) => void;
}
/**
* Configuration options for RalphLoop
*/
export interface RalphLoopOptions {
interface RalphLoopOptions {
/** How often to check for new tasks (default from config) */
pollIntervalMs?: number;
/** Minimum time to run before stopping (null = no minimum) */
+109 -107
View File
@@ -91,6 +91,67 @@ const COMPLETION_INDICATOR_PATTERNS = [
/project\s+(?:is\s+)?(?:completed?|done|finished)/i,
];
interface FieldParser<T> {
pattern: RegExp;
field: keyof RalphStatusBlock;
validate: (value: string) => boolean;
transform: (value: string) => T;
errorMsg: (value: string) => string;
}
const FIELD_PARSERS: FieldParser<RalphStatusValue | RalphTestsStatus | RalphWorkType | number | boolean | string>[] = [
{
pattern: RALPH_STATUS_FIELD_PATTERN,
field: 'status',
validate: (v) => ['IN_PROGRESS', 'COMPLETE', 'BLOCKED'].includes(v.toUpperCase()),
transform: (v) => v.toUpperCase() as RalphStatusValue,
errorMsg: (v) => `Invalid STATUS value: "${v}". Expected: IN_PROGRESS, COMPLETE, or BLOCKED`,
},
{
pattern: RALPH_TASKS_COMPLETED_PATTERN,
field: 'tasksCompletedThisLoop',
validate: (v) => !Number.isNaN(parseInt(v, 10)) && parseInt(v, 10) >= 0,
transform: (v) => parseInt(v, 10),
errorMsg: (v) => `Invalid TASKS_COMPLETED_THIS_LOOP value: "${v}". Expected: non-negative integer`,
},
{
pattern: RALPH_FILES_MODIFIED_PATTERN,
field: 'filesModified',
validate: (v) => !Number.isNaN(parseInt(v, 10)) && parseInt(v, 10) >= 0,
transform: (v) => parseInt(v, 10),
errorMsg: (v) => `Invalid FILES_MODIFIED value: "${v}". Expected: non-negative integer`,
},
{
pattern: RALPH_TESTS_STATUS_PATTERN,
field: 'testsStatus',
validate: (v) => ['PASSING', 'FAILING', 'NOT_RUN'].includes(v.toUpperCase()),
transform: (v) => v.toUpperCase() as RalphTestsStatus,
errorMsg: (v) => `Invalid TESTS_STATUS value: "${v}". Expected: PASSING, FAILING, or NOT_RUN`,
},
{
pattern: RALPH_WORK_TYPE_PATTERN,
field: 'workType',
validate: (v) => ['IMPLEMENTATION', 'TESTING', 'DOCUMENTATION', 'REFACTORING'].includes(v.toUpperCase()),
transform: (v) => v.toUpperCase() as RalphWorkType,
errorMsg: (v) =>
`Invalid WORK_TYPE value: "${v}". Expected: IMPLEMENTATION, TESTING, DOCUMENTATION, or REFACTORING`,
},
{
pattern: RALPH_EXIT_SIGNAL_PATTERN,
field: 'exitSignal',
validate: () => true,
transform: (v) => v.toLowerCase() === 'true',
errorMsg: () => '',
},
{
pattern: RALPH_RECOMMENDATION_PATTERN,
field: 'recommendation',
validate: () => true,
transform: (v) => v.trim(),
errorMsg: () => '',
},
];
/**
* RalphStatusParser - Parses RALPH_STATUS blocks and manages circuit breaker.
*
@@ -303,85 +364,21 @@ export class RalphStatusParser extends EventEmitter {
const trimmedLine = line.trim();
if (!trimmedLine) continue;
// Track whether this line matched any known field
let matched = false;
// STATUS field (required)
const statusMatch = trimmedLine.match(RALPH_STATUS_FIELD_PATTERN);
if (statusMatch) {
const value = statusMatch[1].toUpperCase();
if (['IN_PROGRESS', 'COMPLETE', 'BLOCKED'].includes(value)) {
block.status = value as RalphStatusValue;
} else {
parseErrors.push(`Invalid STATUS value: "${value}". Expected: IN_PROGRESS, COMPLETE, or BLOCKED`);
for (const parser of FIELD_PARSERS) {
const match = trimmedLine.match(parser.pattern);
if (match) {
const rawValue = match[1];
if (parser.validate(rawValue)) {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
(block as any)[parser.field] = parser.transform(rawValue);
} else {
parseErrors.push(parser.errorMsg(rawValue));
}
matched = true;
break;
}
matched = true;
}
// TASKS_COMPLETED_THIS_LOOP field
const tasksMatch = trimmedLine.match(RALPH_TASKS_COMPLETED_PATTERN);
if (tasksMatch) {
const value = parseInt(tasksMatch[1], 10);
if (!Number.isNaN(value) && value >= 0) {
block.tasksCompletedThisLoop = value;
} else {
parseErrors.push(
`Invalid TASKS_COMPLETED_THIS_LOOP value: "${tasksMatch[1]}". Expected: non-negative integer`
);
}
matched = true;
}
// FILES_MODIFIED field
const filesMatch = trimmedLine.match(RALPH_FILES_MODIFIED_PATTERN);
if (filesMatch) {
const value = parseInt(filesMatch[1], 10);
if (!Number.isNaN(value) && value >= 0) {
block.filesModified = value;
} else {
parseErrors.push(`Invalid FILES_MODIFIED value: "${filesMatch[1]}". Expected: non-negative integer`);
}
matched = true;
}
// TESTS_STATUS field
const testsMatch = trimmedLine.match(RALPH_TESTS_STATUS_PATTERN);
if (testsMatch) {
const value = testsMatch[1].toUpperCase();
if (['PASSING', 'FAILING', 'NOT_RUN'].includes(value)) {
block.testsStatus = value as RalphTestsStatus;
} else {
parseErrors.push(`Invalid TESTS_STATUS value: "${value}". Expected: PASSING, FAILING, or NOT_RUN`);
}
matched = true;
}
// WORK_TYPE field
const workMatch = trimmedLine.match(RALPH_WORK_TYPE_PATTERN);
if (workMatch) {
const value = workMatch[1].toUpperCase();
if (['IMPLEMENTATION', 'TESTING', 'DOCUMENTATION', 'REFACTORING'].includes(value)) {
block.workType = value as RalphWorkType;
} else {
parseErrors.push(
`Invalid WORK_TYPE value: "${value}". Expected: IMPLEMENTATION, TESTING, DOCUMENTATION, or REFACTORING`
);
}
matched = true;
}
// EXIT_SIGNAL field
const exitMatch = trimmedLine.match(RALPH_EXIT_SIGNAL_PATTERN);
if (exitMatch) {
block.exitSignal = exitMatch[1].toLowerCase() === 'true';
matched = true;
}
// RECOMMENDATION field
const recMatch = trimmedLine.match(RALPH_RECOMMENDATION_PATTERN);
if (recMatch) {
block.recommendation = recMatch[1].trim();
matched = true;
}
// Track unknown fields for debugging (only if looks like a field)
@@ -475,38 +472,9 @@ export class RalphStatusParser extends EventEmitter {
const prevState = this._circuitBreaker.state;
if (hasProgress) {
// Progress detected - reset counters, possibly close circuit
this._circuitBreaker.consecutiveNoProgress = 0;
this._circuitBreaker.consecutiveSameError = 0;
this._circuitBreaker.lastProgressIteration = this._cycleCount;
if (this._circuitBreaker.state === 'HALF_OPEN') {
this._circuitBreaker.state = 'CLOSED';
this._circuitBreaker.reason = 'Progress detected, circuit closed';
this._circuitBreaker.reasonCode = 'progress_detected';
}
this._handleProgressDetected();
} else {
// No progress
this._circuitBreaker.consecutiveNoProgress++;
// State transitions based on consecutive no-progress
if (this._circuitBreaker.state === 'CLOSED') {
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = `No progress for ${this._circuitBreaker.consecutiveNoProgress} iterations`;
this._circuitBreaker.reasonCode = 'no_progress_open';
} else if (this._circuitBreaker.consecutiveNoProgress >= 2) {
this._circuitBreaker.state = 'HALF_OPEN';
this._circuitBreaker.reason = 'Warning: no progress detected';
this._circuitBreaker.reasonCode = 'no_progress_warning';
}
} else if (this._circuitBreaker.state === 'HALF_OPEN') {
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = `No progress for ${this._circuitBreaker.consecutiveNoProgress} iterations`;
this._circuitBreaker.reasonCode = 'no_progress_open';
}
}
this._handleNoProgress();
}
// Track tests failure
@@ -535,6 +503,40 @@ export class RalphStatusParser extends EventEmitter {
}
}
private _handleProgressDetected(): void {
this._circuitBreaker.consecutiveNoProgress = 0;
this._circuitBreaker.consecutiveSameError = 0;
this._circuitBreaker.lastProgressIteration = this._cycleCount;
if (this._circuitBreaker.state === 'HALF_OPEN') {
this._circuitBreaker.state = 'CLOSED';
this._circuitBreaker.reason = 'Progress detected, circuit closed';
this._circuitBreaker.reasonCode = 'progress_detected';
}
}
private _handleNoProgress(): void {
this._circuitBreaker.consecutiveNoProgress++;
if (this._circuitBreaker.state === 'CLOSED') {
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = `No progress for ${this._circuitBreaker.consecutiveNoProgress} iterations`;
this._circuitBreaker.reasonCode = 'no_progress_open';
} else if (this._circuitBreaker.consecutiveNoProgress >= 2) {
this._circuitBreaker.state = 'HALF_OPEN';
this._circuitBreaker.reason = 'Warning: no progress detected';
this._circuitBreaker.reasonCode = 'no_progress_warning';
}
} else if (this._circuitBreaker.state === 'HALF_OPEN') {
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = `No progress for ${this._circuitBreaker.consecutiveNoProgress} iterations`;
this._circuitBreaker.reasonCode = 'no_progress_open';
}
}
}
/**
* Check line for completion indicators (natural language patterns).
* Used for dual-condition exit gate.
+58 -96
View File
@@ -19,7 +19,6 @@
* Key exports:
* - `RalphTracker` class — main tracker, extends EventEmitter
* - `RalphTrackerEvents` interface — typed event map
* - Re-exports: `EnhancedPlanTask`, `CheckpointReview` from ralph-plan-tracker
*
* Key methods: `processData(data)` — feed terminal output, `getState()`,
* `getTodos()`, `getCompletionHistory()`, `getPlanTasks()`, `reset()`
@@ -66,9 +65,6 @@ import { RalphStallDetector } from './ralph-stall-detector.js';
import { RalphStatusParser } from './ralph-status-parser.js';
import { STALE_DATA_MAX_AGE_MS, INACTIVITY_TIMEOUT_MS } from './config/server-timing.js';
// Re-export sub-module types for backward compatibility
export type { EnhancedPlanTask, CheckpointReview } from './ralph-plan-tracker.js';
// ========== Configuration Constants ==========
// Note: MAX_TODOS_PER_SESSION and MAX_LINE_BUFFER_SIZE are imported from config modules
@@ -100,6 +96,18 @@ const TODO_CLEANUP_INTERVAL_MS = INACTIVITY_TIMEOUT_MS;
*/
const TODO_SIMILARITY_THRESHOLD = 0.85;
/**
* Similarity threshold for short todo content (<30 chars).
* Higher threshold reduces false positive deduplication of short strings.
*/
const SIMILARITY_THRESHOLD_SHORT = 0.95;
/**
* Similarity threshold for medium-length todo content (30-60 chars).
* Slightly relaxed compared to short strings.
*/
const SIMILARITY_THRESHOLD_MEDIUM = 0.9;
/**
* Debounce interval for event emissions (milliseconds).
* Prevents UI jitter from rapid consecutive updates.
@@ -378,32 +386,6 @@ const P2_PRIORITY_PATTERNS = [
* @event circuitBreakerUpdate - Fired when circuit breaker state changes
* @event exitGateMet - Fired when dual-condition exit gate is met
*/
export interface RalphTrackerEvents {
/** Emitted when loop state changes */
loopUpdate: (state: RalphTrackerState) => void;
/** Emitted when todo list is modified */
todoUpdate: (todos: RalphTodoItem[]) => void;
/** Emitted when completion phrase detected (loop finished) */
completionDetected: (phrase: string) => void;
/** Emitted when tracker auto-enables from disabled state */
enabled: () => void;
/** Emitted when a RALPH_STATUS block is parsed */
statusBlockDetected: (block: RalphStatusBlock) => void;
/** Emitted when circuit breaker state changes */
circuitBreakerUpdate: (status: CircuitBreakerStatus) => void;
/** Emitted when dual-condition exit gate is met (completion indicators >= 2 AND EXIT_SIGNAL: true) */
exitGateMet: (data: { completionIndicators: number; exitSignal: boolean }) => void;
/** Emitted when iteration count hasn't changed for an extended period (stall warning) */
iterationStallWarning: (data: { iteration: number; stallDurationMs: number }) => void;
/** Emitted when iteration count hasn't changed for critical period (stall critical) */
iterationStallCritical: (data: { iteration: number; stallDurationMs: number }) => void;
/** Emitted when a common/risky completion phrase is detected (P1-002) */
phraseValidationWarning: (data: {
phrase: string;
reason: 'common' | 'short' | 'numeric';
suggestedPhrase: string;
}) => void;
}
/**
* RalphTracker - Parses terminal output to detect Ralph Wiggum loops and todos
@@ -1299,6 +1281,24 @@ export class RalphTracker extends EventEmitter {
this.detectTodoItems(trimmed);
}
/**
* Mark all tracked todos as completed and emit todoUpdate if any changed.
* @returns true if any todo was updated
*/
private completeAllTodos(): boolean {
let updated = false;
for (const todo of this._todos.values()) {
if (todo.status !== 'completed') {
todo.status = 'completed';
updated = true;
}
}
if (updated) {
this.emit('todoUpdate', this.todos);
}
return updated;
}
/**
* Detect "all tasks complete" messages.
*/
@@ -1318,16 +1318,7 @@ export class RalphTracker extends EventEmitter {
return;
}
let updated = false;
for (const todo of this._todos.values()) {
if (todo.status !== 'completed') {
todo.status = 'completed';
updated = true;
}
}
if (updated) {
this.emit('todoUpdate', this.todos);
}
this.completeAllTodos();
if (this._loopState.completionPhrase) {
this._loopState.active = false;
@@ -1425,16 +1416,7 @@ export class RalphTracker extends EventEmitter {
if (bareCount > 1) return;
let updated = false;
for (const todo of this._todos.values()) {
if (todo.status !== 'completed') {
todo.status = 'completed';
updated = true;
}
}
if (updated) {
this.emit('todoUpdate', this.todos);
}
this.completeAllTodos();
this._loopState.active = false;
this._loopState.lastActivity = Date.now();
@@ -1480,16 +1462,7 @@ export class RalphTracker extends EventEmitter {
if (canonicalCount >= 2 || this._loopState.active) {
this._loopState.active = false;
this._loopState.lastActivity = Date.now();
let updated = false;
for (const todo of this._todos.values()) {
if (todo.status !== 'completed') {
todo.status = 'completed';
updated = true;
}
}
if (updated) {
this.emit('todoUpdate', this.todos);
}
this.completeAllTodos();
this.emit('completionDetected', matchedPhrase);
this.emit('loopUpdate', this.loopState);
return;
@@ -1497,16 +1470,7 @@ export class RalphTracker extends EventEmitter {
}
if (this._loopState.active || count >= 2) {
let updated = false;
for (const todo of this._todos.values()) {
if (todo.status !== 'completed') {
todo.status = 'completed';
updated = true;
}
}
if (updated) {
this.emit('todoUpdate', this.todos);
}
this.completeAllTodos();
this._loopState.active = false;
this._loopState.lastActivity = Date.now();
@@ -1532,41 +1496,39 @@ export class RalphTracker extends EventEmitter {
const suggestedPhrase = `${phrase}_${uniqueSuffix}`;
if (COMMON_COMPLETION_PHRASES.has(normalized)) {
console.warn(
`[RalphTracker] Warning: Completion phrase "${phrase}" is very common and may cause false positives. Consider using: "${suggestedPhrase}"`
);
this.emit('phraseValidationWarning', {
phrase,
reason: 'common',
suggestedPhrase,
});
this.emitValidationWarning(phrase, 'common', suggestedPhrase);
return;
}
if (normalized.length < MIN_RECOMMENDED_PHRASE_LENGTH) {
console.warn(
`[RalphTracker] Warning: Completion phrase "${phrase}" is too short (${normalized.length} chars). Consider using: "${suggestedPhrase}"`
);
this.emit('phraseValidationWarning', {
phrase,
reason: 'short',
suggestedPhrase,
});
this.emitValidationWarning(phrase, 'short', suggestedPhrase);
return;
}
if (/^\d+$/.test(normalized)) {
console.warn(
`[RalphTracker] Warning: Completion phrase "${phrase}" is numeric-only and may cause false positives. Consider using: "${suggestedPhrase}"`
);
this.emit('phraseValidationWarning', {
phrase,
reason: 'numeric',
suggestedPhrase,
});
this.emitValidationWarning(phrase, 'numeric', suggestedPhrase);
}
}
/**
* Emit a phrase validation warning with a console message and event.
*/
private emitValidationWarning(phrase: string, reason: 'common' | 'short' | 'numeric', suggestedPhrase: string): void {
const descriptions: Record<'common' | 'short' | 'numeric', string> = {
common: 'is very common and may cause false positives',
short: `is too short (${phrase.toUpperCase().replace(/[\s_\-.]+/g, '').length} chars)`,
numeric: 'is numeric-only and may cause false positives',
};
console.warn(
`[RalphTracker] Warning: Completion phrase "${phrase}" ${descriptions[reason]}. Consider using: "${suggestedPhrase}"`
);
this.emit('phraseValidationWarning', {
phrase,
reason,
suggestedPhrase,
});
}
/**
* Activate the loop if not already active.
*/
@@ -1977,9 +1939,9 @@ export class RalphTracker extends EventEmitter {
let threshold: number;
if (normalized.length < 30) {
threshold = 0.95;
threshold = SIMILARITY_THRESHOLD_SHORT;
} else if (normalized.length < 60) {
threshold = 0.9;
threshold = SIMILARITY_THRESHOLD_MEDIUM;
} else {
threshold = TODO_SIMILARITY_THRESHOLD;
}
+269 -253
View File
@@ -27,7 +27,7 @@
* - `RespawnController` class — state machine, extends EventEmitter
* - `RespawnConfig` interface — all configuration options
* - `RespawnState` type — union of all state machine states
* - `DetectionStatus`, `ActiveTimerInfo`, `RespawnEvents` — status/event types
* - `DetectionStatus`, `ActiveTimerInfo` — status types
*
* Key methods: `start()`, `stop()`, `getStatus()`, `getConfig()`,
* `getDetectionStatus()`, `getActiveTimers()`, `getAggregateMetrics()`,
@@ -46,11 +46,10 @@
import { EventEmitter } from 'node:events';
import { randomUUID } from 'node:crypto';
import { Session } from './session.js';
import { AiIdleChecker, type AiCheckResult, type AiCheckState } from './ai-idle-checker.js';
import { AiPlanChecker, type AiPlanCheckResult } from './ai-plan-checker.js';
import { AiIdleChecker, type AiCheckState } from './ai-idle-checker.js';
import { AiPlanChecker } from './ai-plan-checker.js';
import type { TeamWatcher } from './team-watcher.js';
import { BufferAccumulator } from './utils/buffer-accumulator.js';
import { ANSI_ESCAPE_PATTERN_SIMPLE, assertNever, CleanupManager } from './utils/index.js';
import { BufferAccumulator, ANSI_ESCAPE_PATTERN_SIMPLE, assertNever, CleanupManager } from './utils/index.js';
import { MAX_RESPAWN_BUFFER_SIZE, TRIM_RESPAWN_BUFFER_TO as RESPAWN_BUFFER_TRIM_SIZE } from './config/buffer-limits.js';
import {
isCompletionMessage,
@@ -71,12 +70,13 @@ import {
AI_PLAN_CHECK_TIMEOUT_MS,
AI_PLAN_CHECK_COOLDOWN_MS,
} from './config/ai-defaults.js';
import type {
RespawnCycleMetrics,
RespawnAggregateMetrics,
RalphLoopHealthScore,
TimingHistory,
CycleOutcome,
import {
getErrorMessage,
type RespawnCycleMetrics,
type RespawnAggregateMetrics,
type RalphLoopHealthScore,
type TimingHistory,
type CycleOutcome,
} from './types.js';
// ========== Constants ==========
@@ -99,13 +99,13 @@ const PLAN_MODE_SELECTOR_PATTERN = /[❯>]\s*\d+\./;
* Each layer provides a confidence signal that Claude has finished working.
*/
/** Active timer info for UI display */
export interface ActiveTimerInfo {
interface ActiveTimerInfo {
name: string;
remainingMs: number;
totalMs: number;
}
export interface DetectionStatus {
interface DetectionStatus {
/** Layer 0: Stop hook received (highest priority - definitive signal) */
stopHookReceived: boolean;
/** Timestamp when Stop hook was received */
@@ -488,68 +488,19 @@ export interface RespawnConfig {
* @event error - Fired on errors
* @event log - Fired for debug logging
*/
/** Timer info for countdown display */
export interface TimerInfo {
name: string;
durationMs: number;
endsAt: number;
reason?: string;
}
/** Action log entry for detailed UI feedback */
export interface ActionLogEntry {
interface ActionLogEntry {
type: string;
detail: string;
timestamp: number;
}
export interface RespawnEvents {
/** State machine transition */
stateChanged: (state: RespawnState, prevState: RespawnState) => void;
/** New respawn cycle started */
respawnCycleStarted: (cycleNumber: number) => void;
/** Respawn cycle finished */
respawnCycleCompleted: (cycleNumber: number) => void;
/** Command sent to session */
stepSent: (step: string, input: string) => void;
/** Step completed (ready indicator detected) */
stepCompleted: (step: string) => void;
/** Detection status update for UI display */
detectionUpdate: (status: DetectionStatus) => void;
/** Auto-accept sent for plan mode approval */
autoAcceptSent: () => void;
/** AI idle check started */
aiCheckStarted: () => void;
/** AI idle check completed with verdict */
aiCheckCompleted: (result: AiCheckResult) => void;
/** AI idle check failed */
aiCheckFailed: (error: string) => void;
/** AI idle check cooldown state changed */
aiCheckCooldown: (active: boolean, endsAt: number | null) => void;
/** AI plan check started */
planCheckStarted: () => void;
/** AI plan check completed with verdict */
planCheckCompleted: (result: AiPlanCheckResult) => void;
/** AI plan check failed */
planCheckFailed: (error: string) => void;
/** Timer started for countdown display */
timerStarted: (timer: TimerInfo) => void;
/** Timer cancelled */
timerCancelled: (timerName: string, reason?: string) => void;
/** Timer completed */
timerCompleted: (timerName: string) => void;
/** Verbose action log for detailed UI feedback */
actionLog: (action: ActionLogEntry) => void;
/** Error occurred */
error: (error: Error) => void;
/** Debug log message */
log: (message: string) => void;
/** Stuck state warning emitted */
stuckStateWarning: (state: RespawnState, durationMs: number) => void;
/** Stuck state recovery triggered */
stuckStateRecovery: (state: RespawnState, durationMs: number, attempt: number) => void;
/** Respawn blocked by external signal */
respawnBlocked: (data: { reason: string; details: string }) => void;
/**
* Convert milliseconds to a non-negative whole number of seconds for countdown display.
* Rounds up so that e.g. 1200 ms shows as 2 s (never under-reports remaining time).
*/
function formatRemainingSeconds(ms: number): number {
return Math.max(0, Math.ceil(ms / 1000));
}
/** Default configuration values */
@@ -840,27 +791,42 @@ export class RespawnController extends EventEmitter {
private validateConfig(): void {
const c = this.config;
// Ensure timeouts are positive
if (c.idleTimeoutMs <= 0) c.idleTimeoutMs = DEFAULT_CONFIG.idleTimeoutMs;
if (c.completionConfirmMs <= 0) c.completionConfirmMs = DEFAULT_CONFIG.completionConfirmMs;
if (c.noOutputTimeoutMs <= 0) c.noOutputTimeoutMs = DEFAULT_CONFIG.noOutputTimeoutMs;
if (c.autoAcceptDelayMs < 0) c.autoAcceptDelayMs = DEFAULT_CONFIG.autoAcceptDelayMs;
if (c.interStepDelayMs <= 0) c.interStepDelayMs = DEFAULT_CONFIG.interStepDelayMs;
/**
* Validate that a timeout value is positive (or non-negative when allowZero is true).
* Falls back to the DEFAULT_CONFIG value if invalid.
*/
const validatePositiveTimeout = (field: keyof RespawnConfig, allowZero = false): void => {
const value = c[field] as number;
const invalid = allowZero ? value < 0 : value <= 0;
if (invalid) {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
(c as any)[field] = DEFAULT_CONFIG[field];
}
};
const REQUIRED_TIMEOUT_FIELDS = [
'idleTimeoutMs',
'completionConfirmMs',
'noOutputTimeoutMs',
'interStepDelayMs',
'aiIdleCheckTimeoutMs',
'aiIdleCheckMaxContext',
'aiPlanCheckTimeoutMs',
'aiPlanCheckMaxContext',
] as const;
for (const field of REQUIRED_TIMEOUT_FIELDS) {
validatePositiveTimeout(field);
}
const ALLOW_ZERO_FIELDS = ['autoAcceptDelayMs', 'aiIdleCheckCooldownMs', 'aiPlanCheckCooldownMs'] as const;
for (const field of ALLOW_ZERO_FIELDS) {
validatePositiveTimeout(field, true);
}
// Ensure completion confirm doesn't exceed no-output timeout
if (c.completionConfirmMs > c.noOutputTimeoutMs) {
c.completionConfirmMs = c.noOutputTimeoutMs;
}
// Ensure AI check timeouts are positive
if (c.aiIdleCheckTimeoutMs <= 0) c.aiIdleCheckTimeoutMs = DEFAULT_CONFIG.aiIdleCheckTimeoutMs;
if (c.aiIdleCheckCooldownMs < 0) c.aiIdleCheckCooldownMs = DEFAULT_CONFIG.aiIdleCheckCooldownMs;
if (c.aiIdleCheckMaxContext <= 0) c.aiIdleCheckMaxContext = DEFAULT_CONFIG.aiIdleCheckMaxContext;
// Ensure plan check timeouts are positive
if (c.aiPlanCheckTimeoutMs <= 0) c.aiPlanCheckTimeoutMs = DEFAULT_CONFIG.aiPlanCheckTimeoutMs;
if (c.aiPlanCheckCooldownMs < 0) c.aiPlanCheckCooldownMs = DEFAULT_CONFIG.aiPlanCheckCooldownMs;
if (c.aiPlanCheckMaxContext <= 0) c.aiPlanCheckMaxContext = DEFAULT_CONFIG.aiPlanCheckMaxContext;
}
/** Wire up AI checker events to controller events (removes existing listeners first to prevent duplicates) */
@@ -988,11 +954,11 @@ export class RespawnController extends EventEmitter {
waitingFor = 'AI verdict (IDLE or WORKING)';
} else if (this._state === 'confirming_idle') {
statusText = `Confirming idle (${confidence}% confidence)`;
waitingFor = `${Math.max(0, Math.ceil((this.config.completionConfirmMs - msSinceLastOutput) / 1000))}s more silence`;
waitingFor = `${formatRemainingSeconds(this.config.completionConfirmMs - msSinceLastOutput)}s more silence`;
} else if (this._state === 'watching') {
const aiState = this.aiChecker.getState();
if (aiState.status === 'cooldown') {
const remaining = Math.ceil(this.aiChecker.getCooldownRemainingMs() / 1000);
const remaining = formatRemainingSeconds(this.aiChecker.getCooldownRemainingMs());
statusText = `AI Check: WORKING (cooldown ${remaining}s)`;
waitingFor = 'Cooldown to expire';
} else if (completionMessageDetected) {
@@ -1365,98 +1331,13 @@ export class RespawnController extends EventEmitter {
this.lastTokenChangeTime = now;
}
// Detect completion message FIRST (Layer 1) - PRIMARY DETECTION
// Check this before working patterns because completion message indicates
// the work is done, even if working patterns are still in the rolling window
if (isCompletionMessage(data)) {
// Clear the rolling window - completion marks a transition point
this.clearWorkingPatternWindow();
this.workingDetected = false;
this.completionMessageTime = now;
this.cancelAutoAcceptTimer(); // Normal idle flow handles this
this.log(`Completion message detected: "${data.trim().substring(0, 50)}..."`);
// Layer 1: Completion message (PRIMARY) — checked before working patterns
if (this._detectCompletionMessage(data, now)) return;
// In watching state, start completion confirmation timer
if (this._state === 'watching') {
this.startCompletionConfirmTimer();
return;
}
// Layer 4: Working patterns
if (this._detectWorkingPattern(data, now)) return;
// In waiting states, also use confirmation timer (same detection logic)
// This ensures we wait for Claude to finish before proceeding
// Note: 'watching' is already handled above and returns early
switch (this._state) {
case 'waiting_update':
this.startStepConfirmTimer('update');
break;
case 'waiting_clear':
this.checkClearComplete(); // /clear is quick, no need to wait
break;
case 'waiting_init':
this.startStepConfirmTimer('init');
break;
case 'waiting_kickstart':
this.startStepConfirmTimer('kickstart');
break;
// Non-waiting states: completion message is ignored
case 'confirming_idle':
case 'ai_checking':
case 'sending_update':
case 'sending_clear':
case 'sending_init':
case 'monitoring_init':
case 'sending_kickstart':
case 'stopped':
// Completion message during these states is ignored
break;
default:
assertNever(this._state, `Unhandled RespawnState in completion detection: ${this._state}`);
}
return;
}
// Detect working patterns (Layer 4)
const isWorking = this.checkWorkingPattern(data);
if (isWorking) {
this.workingDetected = true;
this.promptDetected = false;
this.elicitationDetected = false; // Clear on new work cycle
this.resetHookState(); // Clear hook signals on new work
this.lastWorkingPatternTime = now;
// Cancel hook confirmation timer if running
this.cancelTrackedTimer('hook-confirm', 'working patterns detected');
// Cancel any pending completion confirmation
this.cancelCompletionConfirm();
// Cancel any pending step confirmation (Claude is still working)
this.cancelStepConfirm();
// If AI check is running, cancel it (Claude is working)
if (this._state === 'ai_checking') {
this.log('Working patterns detected during AI check, cancelling');
this.aiChecker.cancel();
this.setState('watching');
}
// Cancel plan check if running (Claude started working)
if (this.planChecker.status === 'checking') {
this.log('Working patterns detected during plan check, cancelling');
this.planChecker.cancel();
}
// If we're monitoring init and work started, go to watching (no kickstart needed)
if (this._state === 'monitoring_init') {
this.log('/init triggered work, skipping kickstart');
this.emit('stepCompleted', 'init');
this.completeCycle();
}
return;
}
// In confirming_idle or ai_checking state, substantial output cancels the flow.
// This prevents false triggers when Claude pauses briefly mid-work.
// Substantial output during confirming_idle/ai_checking cancels the flow
if (this._state === 'confirming_idle' || this._state === 'ai_checking') {
// Strip ANSI escape codes to check if there's real content
ANSI_ESCAPE_PATTERN_SIMPLE.lastIndex = 0;
@@ -1477,43 +1358,137 @@ export class RespawnController extends EventEmitter {
}
}
// Legacy fallback: detect prompt characters (still useful for waiting_* states)
const hasPrompt = PROMPT_PATTERNS.some((pattern) => data.includes(pattern));
if (hasPrompt) {
this.promptDetected = true;
this.workingDetected = false;
// Legacy fallback: prompt detection
this._detectPrompt(data);
}
// Handle legacy detection in waiting states - also use confirmation timers
switch (this._state) {
case 'waiting_update':
this.startStepConfirmTimer('update');
break;
case 'waiting_clear':
this.checkClearComplete(); // /clear is quick, no need to wait
break;
case 'waiting_init':
this.startStepConfirmTimer('init');
break;
case 'monitoring_init':
this.checkMonitoringInitIdle();
break;
case 'waiting_kickstart':
this.startStepConfirmTimer('kickstart');
break;
// Non-waiting states: prompt detection is informational only
case 'watching':
case 'confirming_idle':
case 'ai_checking':
case 'sending_update':
case 'sending_clear':
case 'sending_init':
case 'sending_kickstart':
case 'stopped':
// Prompt detection during these states doesn't trigger action
break;
default:
assertNever(this._state, `Unhandled RespawnState in prompt detection: ${this._state}`);
}
private _detectCompletionMessage(data: string, now: number): boolean {
if (!isCompletionMessage(data)) return false;
// Clear the rolling window - completion marks a transition point
this.clearWorkingPatternWindow();
this.workingDetected = false;
this.completionMessageTime = now;
this.cancelAutoAcceptTimer(); // Normal idle flow handles this
this.log(`Completion message detected: "${data.trim().substring(0, 50)}..."`);
// In watching state, start completion confirmation timer
if (this._state === 'watching') {
this.startCompletionConfirmTimer();
return true;
}
// In waiting states, also use confirmation timer (same detection logic)
// This ensures we wait for Claude to finish before proceeding
// Note: 'watching' is already handled above and returns early
switch (this._state) {
case 'waiting_update':
this.startStepConfirmTimer('update');
break;
case 'waiting_clear':
this.checkClearComplete(); // /clear is quick, no need to wait
break;
case 'waiting_init':
this.startStepConfirmTimer('init');
break;
case 'waiting_kickstart':
this.startStepConfirmTimer('kickstart');
break;
// Non-waiting states: completion message is ignored
case 'confirming_idle':
case 'ai_checking':
case 'sending_update':
case 'sending_clear':
case 'sending_init':
case 'monitoring_init':
case 'sending_kickstart':
case 'stopped':
// Completion message during these states is ignored
break;
default:
assertNever(this._state, `Unhandled RespawnState in completion detection: ${this._state}`);
}
return true;
}
private _detectWorkingPattern(data: string, now: number): boolean {
const isWorking = this.checkWorkingPattern(data);
if (!isWorking) return false;
this.workingDetected = true;
this.promptDetected = false;
this.elicitationDetected = false; // Clear on new work cycle
this.resetHookState(); // Clear hook signals on new work
this.lastWorkingPatternTime = now;
// Cancel hook confirmation timer if running
this.cancelTrackedTimer('hook-confirm', 'working patterns detected');
// Cancel any pending completion confirmation
this.cancelCompletionConfirm();
// Cancel any pending step confirmation (Claude is still working)
this.cancelStepConfirm();
// If AI check is running, cancel it (Claude is working)
if (this._state === 'ai_checking') {
this.log('Working patterns detected during AI check, cancelling');
this.aiChecker.cancel();
this.setState('watching');
}
// Cancel plan check if running (Claude started working)
if (this.planChecker.status === 'checking') {
this.log('Working patterns detected during plan check, cancelling');
this.planChecker.cancel();
}
// If we're monitoring init and work started, go to watching (no kickstart needed)
if (this._state === 'monitoring_init') {
this.log('/init triggered work, skipping kickstart');
this.emit('stepCompleted', 'init');
this.completeCycle();
}
return true;
}
private _detectPrompt(data: string): void {
const hasPrompt = PROMPT_PATTERNS.some((pattern) => data.includes(pattern));
if (!hasPrompt) return;
this.promptDetected = true;
this.workingDetected = false;
// Handle legacy detection in waiting states - also use confirmation timers
switch (this._state) {
case 'waiting_update':
this.startStepConfirmTimer('update');
break;
case 'waiting_clear':
this.checkClearComplete(); // /clear is quick, no need to wait
break;
case 'waiting_init':
this.startStepConfirmTimer('init');
break;
case 'monitoring_init':
this.checkMonitoringInitIdle();
break;
case 'waiting_kickstart':
this.startStepConfirmTimer('kickstart');
break;
// Non-waiting states: prompt detection is informational only
case 'watching':
case 'confirming_idle':
case 'ai_checking':
case 'sending_update':
case 'sending_clear':
case 'sending_init':
case 'sending_kickstart':
case 'stopped':
// Prompt detection during these states doesn't trigger action
break;
default:
assertNever(this._state, `Unhandled RespawnState in prompt detection: ${this._state}`);
}
}
@@ -1799,24 +1774,28 @@ export class RespawnController extends EventEmitter {
case 'sending_init':
case 'sending_kickstart':
// For sending states, retry the send
this.log('Recovery: returning to watching state');
this.setState('watching');
this.startNoOutputTimer();
this.startPreFilterTimer();
if (this.config.autoAcceptPrompts) {
this.startAutoAcceptTimer();
}
this.recoveryResetToWatching('returning to watching state');
break;
default:
// Fallback: reset to watching
this.log('Recovery: fallback to watching state');
this.setState('watching');
this.startNoOutputTimer();
this.startPreFilterTimer();
if (this.config.autoAcceptPrompts) {
this.startAutoAcceptTimer();
}
this.recoveryResetToWatching('fallback to watching state');
}
}
/**
* Reset the controller to watching state during stuck-state recovery.
* Sets state to watching and restarts all detection timers.
*
* @param reason - Human-readable reason for the reset (logged)
*/
private recoveryResetToWatching(reason: string): void {
this.log(`Recovery: ${reason}`);
this.setState('watching');
this.startNoOutputTimer();
this.startPreFilterTimer();
if (this.config.autoAcceptPrompts) {
this.startAutoAcceptTimer();
}
}
@@ -2042,6 +2021,18 @@ export class RespawnController extends EventEmitter {
return;
}
// Check for active child processes (bash tools, test suites, builds, etc.)
// These may produce no terminal output, so restart timers to retry periodically.
const activeProcesses = this.session.getActiveChildProcesses();
if (activeProcesses.length > 0) {
const names = activeProcesses.map((p) => p.command).join(', ');
this.log(`Skipping AI check - ${activeProcesses.length} active child process(es): ${names}`);
this.logAction('detection', `Skipped AI check: child processes running (${names})`);
this.startNoOutputTimer();
this.startPreFilterTimer();
return;
}
// If AI check is disabled or errored out, fall back to direct idle confirmation
if (!this.config.aiIdleCheckEnabled || this.aiChecker.status === 'disabled') {
this.log(`AI check unavailable (${this.aiChecker.status}), confirming idle directly via: ${reason}`);
@@ -2052,7 +2043,7 @@ export class RespawnController extends EventEmitter {
// If on cooldown, don't start check - wait for cooldown to expire
if (this.aiChecker.isOnCooldown()) {
this.log(
`AI check on cooldown (${Math.ceil(this.aiChecker.getCooldownRemainingMs() / 1000)}s remaining), waiting...`
`AI check on cooldown (${formatRemainingSeconds(this.aiChecker.getCooldownRemainingMs())}s remaining), waiting...`
);
return;
}
@@ -2137,7 +2128,7 @@ export class RespawnController extends EventEmitter {
}
if (this._state === 'stopped') return; // Guard against stopped state
if (this._state === 'ai_checking') {
const errorMsg = err instanceof Error ? err.message : String(err);
const errorMsg = getErrorMessage(err);
this.logAction('ai-check', `Failed: ${errorMsg.substring(0, 50)}`);
this.emit('aiCheckFailed', errorMsg);
this.setState('watching');
@@ -2199,36 +2190,15 @@ export class RespawnController extends EventEmitter {
* @fires planCheckStarted
*/
private tryAutoAccept(): void {
// Only auto-accept in watching state (not during a respawn cycle)
if (this._state !== 'watching') return;
if (!this.canAutoAccept()) return;
// Don't auto-accept if a completion message was detected (normal idle handles it)
if (this.completionMessageTime !== null) return;
// Don't auto-accept if disabled
if (!this.config.autoAcceptPrompts) return;
// Don't auto-accept if we haven't received any output yet (prevents spurious Enter on fresh start)
if (!this.hasReceivedOutput) return;
// Don't auto-accept if an elicitation dialog (AskUserQuestion) was detected
if (this.elicitationDetected) {
this.log('Skipping auto-accept: elicitation dialog detected (AskUserQuestion)');
return;
}
// Stage 1: Pre-filter — check if buffer looks like plan mode
const buffer = this.terminalBuffer.value;
if (!this.isPlanModePreFilterMatch(buffer)) {
this.log('Skipping auto-accept: pre-filter did not match plan mode patterns');
return;
}
// Stage 2: AI confirmation (if enabled and available)
if (this.config.aiPlanCheckEnabled && this.planChecker.status !== 'disabled') {
if (this.planChecker.isOnCooldown()) {
this.log(
`Skipping auto-accept: plan checker on cooldown (${Math.ceil(this.planChecker.getCooldownRemainingMs() / 1000)}s remaining)`
`Skipping auto-accept: plan checker on cooldown (${formatRemainingSeconds(this.planChecker.getCooldownRemainingMs())}s remaining)`
);
return;
}
@@ -2245,6 +2215,40 @@ export class RespawnController extends EventEmitter {
this.sendAutoAcceptEnter();
}
/**
* Check whether all preconditions for auto-accept are met.
* Validates state, config, and pre-filter conditions before attempting auto-accept.
*
* @returns True if auto-accept should proceed to the AI confirmation stage
*/
private canAutoAccept(): boolean {
// Only auto-accept in watching state (not during a respawn cycle)
if (this._state !== 'watching') return false;
// Don't auto-accept if a completion message was detected (normal idle handles it)
if (this.completionMessageTime !== null) return false;
// Don't auto-accept if disabled
if (!this.config.autoAcceptPrompts) return false;
// Don't auto-accept if we haven't received any output yet (prevents spurious Enter on fresh start)
if (!this.hasReceivedOutput) return false;
// Don't auto-accept if an elicitation dialog (AskUserQuestion) was detected
if (this.elicitationDetected) {
this.log('Skipping auto-accept: elicitation dialog detected (AskUserQuestion)');
return false;
}
// Stage 1: Pre-filter — check if buffer looks like plan mode
if (!this.isPlanModePreFilterMatch(this.terminalBuffer.value)) {
this.log('Skipping auto-accept: pre-filter did not match plan mode patterns');
return false;
}
return true;
}
/**
* Check if the terminal buffer matches plan mode pre-filter patterns.
* Only checks the last 2000 chars (plan mode UI appears at the bottom).
@@ -2323,7 +2327,7 @@ export class RespawnController extends EventEmitter {
}
})
.catch((err) => {
const errorMsg = err instanceof Error ? err.message : String(err);
const errorMsg = getErrorMessage(err);
this.emit('planCheckFailed', errorMsg);
this.logAction('plan-check', `Failed: ${errorMsg.substring(0, 50)}`);
});
@@ -2633,6 +2637,18 @@ export class RespawnController extends EventEmitter {
return;
}
// Safety check: if child processes are running (bash tools, test suites, builds, etc.)
const activeProcesses = this.session.getActiveChildProcesses();
if (activeProcesses.length > 0) {
const names = activeProcesses.map((p) => p.command).join(', ');
this.log(`Idle confirmation rejected - ${activeProcesses.length} active child process(es): ${names}`);
this.logAction('detection', `Rejected: child processes running (${names})`);
this.setState('watching');
this.startNoOutputTimer();
this.startPreFilterTimer();
return;
}
this.log(`Idle confirmed via: ${reason}`);
const status = this.getDetectionStatus();
this.log(
+99 -58
View File
@@ -27,6 +27,57 @@ const COMPACT_COOLDOWN_MS = 10000;
/** Cooldown after clear completes before re-enabling (5 seconds) */
const CLEAR_COOLDOWN_MS = 5000;
/**
* Executes an action when the session becomes idle, retrying if currently working.
*
* @param action - The async action to execute once idle
* @param isActive - Returns whether this operation is still active (not cancelled)
* @param isWorking - Returns whether the session is currently working
* @param isStopped - Returns whether the session has been stopped
* @param retryMs - Delay between retry attempts when working
* @param cooldownMs - Delay after action completes before calling onCooldownDone
* @param setTimer - Stores the timer reference for cleanup
* @param onCooldownDone - Called after cooldown to reset state
*/
async function executeWhenIdle(
action: () => Promise<void>,
isActive: () => boolean,
isWorking: () => boolean,
isStopped: () => boolean,
retryMs: number,
cooldownMs: number,
setTimer: (timer: NodeJS.Timeout | null) => void,
onCooldownDone: () => void
): Promise<void> {
if (isStopped()) return;
if (!isActive()) return;
if (!isWorking()) {
if (isStopped()) return;
await action();
if (!isStopped()) {
setTimer(
setTimeout(() => {
if (isStopped()) return;
setTimer(null);
onCooldownDone();
}, cooldownMs)
);
}
} else {
if (!isStopped()) {
setTimer(
setTimeout(
() => executeWhenIdle(action, isActive, isWorking, isStopped, retryMs, cooldownMs, setTimer, onCooldownDone),
retryMs
)
);
}
}
}
/** Minimum valid threshold for auto-clear/compact (1000 tokens) */
const MIN_AUTO_THRESHOLD = 1000;
@@ -42,7 +93,7 @@ const DEFAULT_AUTO_COMPACT_THRESHOLD = 110_000;
/**
* Callbacks required by SessionAutoOps to interact with the parent Session.
*/
export interface AutoOpsCallbacks {
interface AutoOpsCallbacks {
/** Send a command via the terminal multiplexer */
writeCommand: (command: string) => Promise<boolean>;
/** Check if Claude is currently working */
@@ -58,12 +109,6 @@ export interface AutoOpsCallbacks {
/**
* Events emitted by SessionAutoOps.
*/
export interface SessionAutoOpsEvents {
/** Auto-compact was triggered and the /compact command was sent */
autoCompact: (data: { tokens: number; threshold: number; prompt?: string }) => void;
/** Auto-clear was triggered and the /clear command was sent */
autoClear: (data: { tokens: number; threshold: number }) => void;
}
/**
* Manages auto-compact and auto-clear automation for a Session.
@@ -181,37 +226,35 @@ export class SessionAutoOps extends EventEmitter {
`[SessionAutoOps] Auto-compact triggered: ${totalTokens} tokens >= ${this._autoCompactThreshold} threshold`
);
const checkAndCompact = async () => {
if (this.callbacks.isStopped()) return;
if (!this._isCompacting) return;
if (!this.callbacks.isWorking()) {
if (this.callbacks.isStopped()) return;
const compactCmd = this._autoCompactPrompt ? `/compact ${this._autoCompactPrompt}\r` : '/compact\r';
await this.callbacks.writeCommand(compactCmd);
this.emit('autoCompact', {
tokens: totalTokens,
threshold: this._autoCompactThreshold,
prompt: this._autoCompactPrompt || undefined,
});
if (!this.callbacks.isStopped()) {
this._autoCompactTimer = setTimeout(() => {
if (this.callbacks.isStopped()) return;
this._autoCompactTimer = null;
this._isCompacting = false;
}, COMPACT_COOLDOWN_MS);
}
} else {
if (!this.callbacks.isStopped()) {
this._autoCompactTimer = setTimeout(checkAndCompact, AUTO_RETRY_DELAY_MS);
}
}
const action = async () => {
const compactCmd = this._autoCompactPrompt ? `/compact ${this._autoCompactPrompt}\r` : '/compact\r';
await this.callbacks.writeCommand(compactCmd);
this.emit('autoCompact', {
tokens: totalTokens,
threshold: this._autoCompactThreshold,
prompt: this._autoCompactPrompt || undefined,
});
};
if (!this.callbacks.isStopped()) {
this._autoCompactTimer = setTimeout(checkAndCompact, AUTO_INITIAL_DELAY_MS);
this._autoCompactTimer = setTimeout(
() =>
executeWhenIdle(
action,
() => this._isCompacting,
() => this.callbacks.isWorking(),
() => this.callbacks.isStopped(),
AUTO_RETRY_DELAY_MS,
COMPACT_COOLDOWN_MS,
(timer) => {
this._autoCompactTimer = timer;
},
() => {
this._isCompacting = false;
}
),
AUTO_INITIAL_DELAY_MS
);
}
}
}
@@ -231,32 +274,30 @@ export class SessionAutoOps extends EventEmitter {
`[SessionAutoOps] Auto-clear triggered: ${totalTokens} tokens >= ${this._autoClearThreshold} threshold`
);
const checkAndClear = async () => {
if (this.callbacks.isStopped()) return;
if (!this._isClearing) return;
if (!this.callbacks.isWorking()) {
if (this.callbacks.isStopped()) return;
await this.callbacks.writeCommand('/clear\r');
this.emit('autoClear', { tokens: totalTokens, threshold: this._autoClearThreshold });
if (!this.callbacks.isStopped()) {
this._autoClearTimer = setTimeout(() => {
if (this.callbacks.isStopped()) return;
this._autoClearTimer = null;
this._isClearing = false;
}, CLEAR_COOLDOWN_MS);
}
} else {
if (!this.callbacks.isStopped()) {
this._autoClearTimer = setTimeout(checkAndClear, AUTO_RETRY_DELAY_MS);
}
}
const action = async () => {
await this.callbacks.writeCommand('/clear\r');
this.emit('autoClear', { tokens: totalTokens, threshold: this._autoClearThreshold });
};
if (!this.callbacks.isStopped()) {
this._autoClearTimer = setTimeout(checkAndClear, AUTO_INITIAL_DELAY_MS);
this._autoClearTimer = setTimeout(
() =>
executeWhenIdle(
action,
() => this._isClearing,
() => this.callbacks.isWorking(),
() => this.callbacks.isStopped(),
AUTO_RETRY_DELAY_MS,
CLEAR_COOLDOWN_MS,
(timer) => {
this._autoClearTimer = timer;
},
() => {
this._isClearing = false;
}
),
AUTO_INITIAL_DELAY_MS
);
}
}
}
+2 -2
View File
@@ -9,13 +9,13 @@
*/
import type { ClaudeMode } from './types.js';
import { getAugmentedPath } from './utils/claude-cli-resolver.js';
import { getAugmentedPath } from './utils/index.js';
/**
* Build Claude CLI permission flags based on the configured mode.
* Returns an array of args to pass to the CLI.
*/
export function buildPermissionArgs(claudeMode: ClaudeMode, allowedTools?: string): string[] {
function buildPermissionArgs(claudeMode: ClaudeMode, allowedTools?: string): string[] {
switch (claudeMode) {
case 'dangerously-skip-permissions':
return ['--dangerously-skip-permissions'];
+10 -14
View File
@@ -30,18 +30,6 @@ import { SessionState } from './types.js';
/**
* Events emitted by SessionManager
*/
export interface SessionManagerEvents {
/** Fired when a new session starts successfully */
sessionStarted: (session: Session) => void;
/** Fired when a session stops (graceful or forced) */
sessionStopped: (sessionId: string) => void;
/** Fired when a session encounters an error */
sessionError: (sessionId: string, error: string) => void;
/** Fired when a session produces terminal output */
sessionOutput: (sessionId: string, output: string) => void;
/** Fired when a completion phrase is detected */
sessionCompletion: (sessionId: string, phrase: string) => void;
}
/**
* Manages multiple Claude sessions with lifecycle coordination.
@@ -164,7 +152,7 @@ export class SessionManager extends EventEmitter {
await session.start();
this.sessions.set(session.id, session);
this.store.setSession(session.id, session.toState());
this.updateSessionState(session);
this.emit('sessionStarted', session);
return session;
@@ -259,7 +247,15 @@ export class SessionManager extends EventEmitter {
}
private updateSessionState(session: Session): void {
this.store.setSession(session.id, session.toState());
// envOverrides is intentionally NOT on SessionState (API safety). For disk
// persistence we augment the stored object with __envOverrides so reboot
// recovery can restore them without leaking through any API serializer.
// The key uses the reserved `__` prefix so it is visibly "internal" to any
// future reader of state.json.
const state = session.toState();
const envOverrides = session.getEnvOverridesForPersist();
const toStore = envOverrides ? { ...state, __envOverrides: envOverrides } : state;
this.store.setSession(session.id, toStore as SessionState);
}
/** Gets all sessions from persistent storage (including stopped). */
+353 -302
View File
@@ -29,6 +29,7 @@
*/
import { EventEmitter } from 'node:events';
import { execSync, execFileSync } from 'node:child_process';
import { v4 as uuidv4 } from 'uuid';
import * as pty from 'node-pty';
import {
@@ -40,6 +41,7 @@ import {
ActiveBashTool,
NiceConfig,
DEFAULT_NICE_CONFIG,
getErrorMessage,
type ClaudeMode,
type SessionMode,
type OpenCodeConfig,
@@ -48,8 +50,8 @@ import type { TerminalMultiplexer, MuxSession } from './mux-interface.js';
import { TaskTracker, type BackgroundTask } from './task-tracker.js';
import { RalphTracker } from './ralph-tracker.js';
import { BashToolParser } from './bash-tool-parser.js';
import { BufferAccumulator } from './utils/buffer-accumulator.js';
import {
BufferAccumulator,
ANSI_ESCAPE_PATTERN_FULL,
TOKEN_PATTERN,
SPINNER_PATTERN,
@@ -64,6 +66,7 @@ import {
MAX_MESSAGES,
MAX_LINE_BUFFER_SIZE,
} from './config/buffer-limits.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
import {
buildInteractiveArgs,
buildPromptArgs,
@@ -118,6 +121,37 @@ const NEWLINE_SPLIT_PATTERN = /\r?\n/;
// Note: Claude CLI PATH resolution moved to session-cli-builder.ts (buildClaudeEnv)
/** PTY fallback geometry when tmux can't be queried (matches pre-#80 hardcoded values). */
const DEFAULT_PTY_COLS = 120;
const DEFAULT_PTY_ROWS = 40;
const TMUX_DISPLAY_TIMEOUT_MS = 2000;
/**
* Ask tmux for the current window geometry of `muxName` so a re-attaching PTY
* client can spawn at the same size and avoid the resize-flicker / scrollback
* loss documented in #80. Returns `{ cols: 120, rows: 40 }` on any failure
* (tmux dead, muxName unknown, malformed output) — caller never has to
* differentiate "tmux unreachable" from "size 120x40".
*
* Argv form (execFileSync, not execSync) keeps `muxName` out of any shell so
* a hostile session name can't inject options.
*/
export function queryTmuxWindowSize(muxName: string): { cols: number; rows: number } {
try {
const sizeStr = execFileSync('tmux', ['display', '-t', muxName, '-p', '#{window_width} #{window_height}'], {
timeout: TMUX_DISPLAY_TIMEOUT_MS,
encoding: 'utf8',
}).trim();
const [w, h] = sizeStr.split(' ').map(Number);
if (w > 0 && h > 0) {
return { cols: w, rows: h };
}
} catch {
/* fall back below */
}
return { cols: DEFAULT_PTY_COLS, rows: DEFAULT_PTY_ROWS };
}
/**
* Represents a JSON message from Claude CLI's stream-json output format.
* Messages are newline-delimited JSON objects parsed from PTY output.
@@ -151,63 +185,6 @@ export interface ClaudeMessage {
* Event signatures emitted by the Session class.
* Subscribe using `session.on('eventName', handler)`.
*/
export interface SessionEvents {
/** Processed text output (ANSI stripped) */
output: (data: string) => void;
/** Parsed JSON message from Claude CLI */
message: (msg: ClaudeMessage) => void;
/** Error output from the session */
error: (data: string) => void;
/** Session process exited */
exit: (code: number | null) => void;
/** One-shot prompt completed with result and cost */
completion: (result: string, cost: number) => void;
/** Raw terminal data (includes ANSI codes) */
terminal: (data: string) => void;
/** Signal to clear terminal display (after mux attach) */
clearTerminal: () => void;
/** New background task started */
taskCreated: (task: BackgroundTask) => void;
/** Background task status changed */
taskUpdated: (task: BackgroundTask) => void;
/** Background task finished successfully */
taskCompleted: (task: BackgroundTask) => void;
/** Background task failed with error */
taskFailed: (task: BackgroundTask, error: string) => void;
/** Auto-clear triggered due to token threshold */
autoClear: (data: { tokens: number; threshold: number }) => void;
/** Auto-compact triggered due to token threshold */
autoCompact: (data: { tokens: number; threshold: number; prompt?: string }) => void;
/** Ralph loop state changed */
ralphLoopUpdate: (state: RalphTrackerState) => void;
/** Ralph todo list updated */
ralphTodoUpdate: (todos: RalphTodoItem[]) => void;
/** Ralph completion phrase detected */
ralphCompletionDetected: (phrase: string) => void;
/** RALPH_STATUS block detected */
ralphStatusBlockDetected: (block: import('./types.js').RalphStatusBlock) => void;
/** Circuit breaker state changed */
ralphCircuitBreakerUpdate: (status: import('./types.js').CircuitBreakerStatus) => void;
/** Dual-condition exit gate met */
ralphExitGateMet: (data: { completionIndicators: number; exitSignal: boolean }) => void;
/** Bash tool with file paths started */
bashToolStart: (tool: ActiveBashTool) => void;
/** Bash tool completed */
bashToolEnd: (tool: ActiveBashTool) => void;
/** Active Bash tools list updated */
bashToolsUpdate: (tools: ActiveBashTool[]) => void;
/** CLI info (version, model, account) updated */
cliInfoUpdated: (info: {
version: string | null;
model: string | null;
accountType: string | null;
latestVersion: string | null;
}) => void;
}
// SessionMode is imported from types.ts (single source of truth)
// Re-export for backwards compatibility with any external consumers
export type { SessionMode } from './types.js';
/**
* Core session class that wraps a PTY process running Claude CLI or a shell.
@@ -271,6 +248,7 @@ export class Session extends EventEmitter {
private _lastPromptTime: number = 0;
private activityTimeout: NodeJS.Timeout | null = null;
private _awaitingIdleConfirmation: boolean = false; // Prevents timeout reset during idle detection
private _trustDialogAccepted: boolean = false; // Prevents repeated trust dialog auto-accept
private _taskTracker: TaskTracker;
// Token tracking for auto-clear
@@ -326,6 +304,10 @@ export class Session extends EventEmitter {
private _openCodeConfig: OpenCodeConfig | undefined;
private _resumeSessionId: string | undefined;
// Ephemeral env overrides (e.g., CLAUDE_CODE_EFFORT_LEVEL). Exported by tmux at spawn,
// preserved across respawns via persisted state. Not written to .claude/settings.local.json.
private _envOverrides: Record<string, string> | undefined;
// Session color for visual differentiation
private _color: import('./types.js').SessionColor = 'default';
@@ -385,6 +367,8 @@ export class Session extends EventEmitter {
openCodeConfig?: OpenCodeConfig;
/** Resume a previous Claude conversation (used after server reboot) */
resumeSessionId?: string;
/** Extra env vars exported to the CLI at spawn time (no disk persistence) */
envOverrides?: Record<string, string>;
}
) {
super();
@@ -432,6 +416,11 @@ export class Session extends EventEmitter {
this._openCodeConfig = config.openCodeConfig;
}
// Apply env overrides (exported at spawn, not persisted to disk)
if (config.envOverrides && Object.keys(config.envOverrides).length > 0) {
this._envOverrides = { ...config.envOverrides };
}
// Initialize task tracker and forward events (store handlers for cleanup)
this._taskTracker = new TaskTracker();
this._taskTrackerHandlers = {
@@ -526,6 +515,20 @@ export class Session extends EventEmitter {
return this._claudeSessionId;
}
// Adopt a Claude conversation ID observed from an external source (e.g. hook
// payload). In interactive PTY mode Claude CLI emits no JSON to stdout, so
// `_handleJsonMessage` never sees `session_id`; hooks are the only signal
// that conveys a post-/clear conversation switch.
adoptClaudeSessionId(newId: string): void {
if (!newId || newId === this._claudeSessionId) return;
this._claudeSessionId = newId;
}
/** The tmux session name, if the session is running inside a mux */
get muxName(): string | null {
return this._muxSession?.muxName ?? null;
}
get totalCost(): number {
return this._totalCost;
}
@@ -538,6 +541,49 @@ export class Session extends EventEmitter {
return this._isWorking;
}
/**
* Check if the session's process tree has active child processes beyond Claude itself.
* Detects running bash tools, test suites, builds, servers, etc. that Claude spawned.
*
* The tmux pane PID is typically "claude" directly (bash exec'd into it). When Claude
* runs a bash tool, it spawns child processes: claude → bash → npm/node/python/etc.
* We check direct children of the pane PID, filtering out "claude" itself (for the rare
* case where bash wraps claude and didn't exec).
*
* Returns an array of {pid, command} for each child process, or empty array if none.
* Returns empty array if no mux session or on error (fail-open to avoid blocking respawn).
*/
getActiveChildProcesses(): { pid: number; command: string }[] {
if (!this._muxSession) return [];
try {
const panePid = this._muxSession.pid;
// Single call: get direct children with their command names
const output = execSync(`ps -o pid=,comm= --ppid ${panePid} 2>/dev/null`, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (!output) return [];
const activeProcesses: { pid: number; command: string }[] = [];
for (const line of output.split('\n')) {
const match = line.trim().match(/^(\d+)\s+(.+)/);
if (!match) continue;
const pid = parseInt(match[1], 10);
const command = match[2].trim();
// Skip the claude process itself (pane_pid may be bash wrapping claude)
if (command === 'claude') continue;
activeProcesses.push({ pid, command });
}
return activeProcesses;
} catch {
// ps returns exit code 1 when no matches — normal (no children)
return [];
}
}
get lastPromptTime(): number {
return this._lastPromptTime;
}
@@ -794,9 +840,29 @@ export class Session extends EventEmitter {
cliLatestVersion: this._cliLatestVersion || undefined,
openCodeConfig: this._openCodeConfig,
resumeSessionId: this._resumeSessionId,
// envOverrides intentionally NOT on the public SessionState type — they must not
// leak into SSE / GET /api/sessions broadcasts (schema allows OPENCODE_*, which
// can carry secrets). For disk persistence, session-manager calls
// getEnvOverridesForPersist() and writes alongside state.
};
}
/**
* Returns a subset of env overrides safe for disk persistence (state.json).
* Only non-sensitive `CLAUDE_CODE_*` keys are included. `OPENCODE_*` keys are
* filtered out because the schema permits them and they can carry secrets
* (e.g., OPENCODE_API_KEY); secrets must not land in `~/.codeman/state.json`.
* Must NOT be included in any API-bound serializer — see toState() comment.
*/
getEnvOverridesForPersist(): Record<string, string> | undefined {
if (!this._envOverrides) return undefined;
const safe: Record<string, string> = {};
for (const [key, value] of Object.entries(this._envOverrides)) {
if (key.startsWith('CLAUDE_CODE_')) safe[key] = value;
}
return Object.keys(safe).length > 0 ? safe : undefined;
}
toDetailedState() {
return {
...this.toLightDetailedState(),
@@ -871,18 +937,79 @@ export class Session extends EventEmitter {
* session.write('help me with this code\r');
* ```
*/
private async _setupOrAttachMuxSession(options: {
respawnPaneOptions: import('./mux-interface.js').RespawnPaneOptions;
createSessionOptions: import('./mux-interface.js').CreateSessionOptions;
spawnErrLabel: string;
}): Promise<{ isRestored: boolean }> {
const mux = this._mux!;
// Verify stale mux session — tmux may have been destroyed (e.g., killed externally)
if (this._muxSession && !mux.muxSessionExists(this._muxSession.muxName)) {
console.log('[Session] Stale mux session detected (tmux gone):', this._muxSession.muxName);
this._muxSession = null;
}
// Check if session exists but pane is dead (remain-on-exit keeps it alive)
// Respawn the pane instead of creating a whole new session — preserves tmux scrollback
let needsNewSession = false;
if (this._muxSession && mux.isPaneDead(this._muxSession.muxName)) {
console.log('[Session] Dead pane detected, respawning:', this._muxSession.muxName);
const newPid = await mux.respawnPane(options.respawnPaneOptions);
if (!newPid) {
console.error('[Session] Failed to respawn pane, will create new session');
needsNewSession = true;
} else {
// Wait a moment for the respawned process to fully start
await new Promise((resolve) => setTimeout(resolve, MUX_STARTUP_DELAY_MS));
}
}
// Check if we already have a mux session (restored session)
const isRestored = this._muxSession !== null && !needsNewSession;
if (isRestored) {
console.log('[Session] Attaching to existing mux session:', this._muxSession!.muxName);
} else {
// Create a new mux session
this._muxSession = await mux.createSession(options.createSessionOptions);
console.log('[Session] Created mux session:', this._muxSession.muxName);
// No extra sleep — createSession() already waits for tmux readiness
}
// Attach to the mux session via PTY
// Query existing tmux window size so re-attach matches (avoids flicker from 120x40 default)
const { cols: ptyCols, rows: ptyRows } = queryTmuxWindowSize(this._muxSession!.muxName);
try {
this.ptyProcess = pty.spawn(mux.getAttachCommand(), mux.getAttachArgs(this._muxSession!.muxName), {
name: 'xterm-256color',
cols: ptyCols,
rows: ptyRows,
cwd: this.workingDir,
env: buildMuxAttachEnv(),
});
} catch (spawnErr) {
console.error(`[Session] Failed to spawn PTY for ${options.spawnErrLabel}:`, spawnErr);
this.emit('error', `Failed to attach to mux session: ${spawnErr}`);
throw spawnErr;
}
return { isRestored };
}
private _handleTerminalOutput(data: string): void {
// BufferAccumulator handles auto-trimming when max size exceeded
this._terminalBuffer.append(data);
this._lastActivityAt = Date.now();
this.emit('terminal', data);
this.emit('output', data);
}
async startInteractive(): Promise<void> {
if (this.ptyProcess) {
throw new Error('Session already has a running process');
}
this._status = 'busy';
this._terminalBuffer.clear();
this._textOutput.clear();
this._errorBuffer = '';
this._messages = [];
this._lineBuffer = '';
this._lastActivityAt = Date.now();
this._resetBuffers();
const modeLabel = this.mode === 'opencode' ? 'OpenCode' : 'Claude';
console.log(
@@ -892,18 +1019,8 @@ export class Session extends EventEmitter {
// If mux wrapping is enabled, create or attach to a mux session
if (this._useMux && this._mux) {
try {
// Verify stale mux session — tmux may have been destroyed (e.g., killed externally)
if (this._muxSession && !this._mux.muxSessionExists(this._muxSession.muxName)) {
console.log('[Session] Stale mux session detected (tmux gone):', this._muxSession.muxName);
this._muxSession = null;
}
// Check if session exists but pane is dead (remain-on-exit keeps it alive)
// Respawn the pane instead of creating a whole new session — preserves tmux scrollback
let needsNewSession = false;
if (this._muxSession && this._mux.isPaneDead(this._muxSession.muxName)) {
console.log('[Session] Dead pane detected, respawning:', this._muxSession.muxName);
const newPid = await this._mux.respawnPane({
const { isRestored } = await this._setupOrAttachMuxSession({
respawnPaneOptions: {
sessionId: this.id,
workingDir: this.workingDir,
mode: this.mode,
@@ -913,23 +1030,9 @@ export class Session extends EventEmitter {
allowedTools: this._allowedTools,
openCodeConfig: this._openCodeConfig,
resumeSessionId: this._resumeSessionId,
});
if (!newPid) {
console.error('[Session] Failed to respawn pane, will create new session');
needsNewSession = true;
} else {
// Wait a moment for the respawned process to fully start
await new Promise((resolve) => setTimeout(resolve, MUX_STARTUP_DELAY_MS));
}
}
// Check if we already have a mux session (restored session)
const isRestoredSession = this._muxSession !== null && !needsNewSession;
if (isRestoredSession) {
console.log('[Session] Attaching to existing mux session:', this._muxSession!.muxName);
} else {
// Create a new mux session
this._muxSession = await this._mux.createSession({
envOverrides: this._envOverrides,
},
createSessionOptions: {
sessionId: this.id,
workingDir: this.workingDir,
mode: this.mode,
@@ -940,36 +1043,17 @@ export class Session extends EventEmitter {
allowedTools: this._allowedTools,
openCodeConfig: this._openCodeConfig,
resumeSessionId: this._resumeSessionId,
});
console.log('[Session] Created mux session:', this._muxSession.muxName);
// No extra sleep — createSession() already waits for tmux readiness
}
envOverrides: this._envOverrides,
},
spawnErrLabel: 'mux attachment',
});
// Attach to the mux session via PTY
try {
this.ptyProcess = pty.spawn(
this._mux.getAttachCommand(),
this._mux.getAttachArgs(this._muxSession!.muxName),
{
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildMuxAttachEnv(),
}
);
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = this._resumeSessionId || this.id;
} catch (spawnErr) {
console.error('[Session] Failed to spawn PTY for mux attachment:', spawnErr);
this.emit('error', `Failed to attach to mux session: ${spawnErr}`);
throw spawnErr;
}
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = this._resumeSessionId || this.id;
// For NEW mux sessions: wait for readiness then clean buffer
// For RESTORED mux sessions: don't do anything - client will fetch buffer on tab switch
if (!isRestoredSession) {
if (!isRestored) {
if (this.mode === 'opencode') {
// OpenCode uses Bubble Tea TUI — no ❯ prompt to detect.
// Wait for TUI to stabilize (output stops changing), then mark ready.
@@ -1035,7 +1119,8 @@ export class Session extends EventEmitter {
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildClaudeEnv(this.id),
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
});
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY:', spawnErr);
@@ -1056,12 +1141,17 @@ export class Session extends EventEmitter {
const data = rawData.replace(FOCUS_ESCAPE_FILTER, '').replace(CTRL_L_PATTERN, ''); // Remove Ctrl+L
if (!data) return; // Skip if only filtered sequences
// BufferAccumulator handles auto-trimming when max size exceeded
this._terminalBuffer.append(data);
this._lastActivityAt = Date.now();
this._handleTerminalOutput(data);
this.emit('terminal', data);
this.emit('output', data);
// === Auto-accept workspace trust dialog ===
// Claude CLI 2.x shows "Yes, I trust this folder" prompt on first launch per directory.
// Codeman sessions always use --dangerously-skip-permissions, so auto-accept.
if (!this._trustDialogAccepted && data.includes('trust this folder')) {
this._trustDialogAccepted = true;
console.log(`[Session] Auto-accepting workspace trust dialog for: ${this.id}`);
// Send Enter to accept the default selection ("Yes, I trust this folder")
this.writeViaMux('\r');
}
// === Idle/working detection runs on every chunk (latency-sensitive) ===
// Detect if Claude is working or at prompt
@@ -1257,13 +1347,7 @@ export class Session extends EventEmitter {
throw new Error('Session already has a running process');
}
this._status = 'busy';
this._terminalBuffer.clear();
this._textOutput.clear();
this._errorBuffer = '';
this._messages = [];
this._lineBuffer = '';
this._lastActivityAt = Date.now();
this._resetBuffers();
// Use user's default shell or bash
const shell = process.env.SHELL || '/bin/bash';
@@ -1275,69 +1359,28 @@ export class Session extends EventEmitter {
// If mux wrapping is enabled, create or attach to a mux session
if (this._useMux && this._mux) {
try {
// Verify stale mux session — tmux may have been destroyed externally
if (this._muxSession && !this._mux.muxSessionExists(this._muxSession.muxName)) {
console.log('[Session] Stale mux session detected (tmux gone):', this._muxSession.muxName);
this._muxSession = null;
}
// Check if session exists but pane is dead (remain-on-exit keeps it alive)
let needsNewSession = false;
if (this._muxSession && this._mux.isPaneDead(this._muxSession.muxName)) {
console.log('[Session] Dead pane detected, respawning:', this._muxSession.muxName);
const newPid = await this._mux.respawnPane({
const { isRestored } = await this._setupOrAttachMuxSession({
respawnPaneOptions: {
sessionId: this.id,
workingDir: this.workingDir,
mode: 'shell',
niceConfig: this._niceConfig,
});
if (!newPid) {
console.error('[Session] Failed to respawn pane, will create new session');
needsNewSession = true;
} else {
await new Promise((resolve) => setTimeout(resolve, MUX_STARTUP_DELAY_MS));
}
}
// Check if we already have a mux session (restored session)
const isRestoredSession = this._muxSession !== null && !needsNewSession;
if (isRestoredSession) {
console.log('[Session] Attaching to existing mux session:', this._muxSession!.muxName);
} else {
// Create a new mux session
this._muxSession = await this._mux.createSession({
envOverrides: this._envOverrides,
},
createSessionOptions: {
sessionId: this.id,
workingDir: this.workingDir,
mode: 'shell',
name: this._name,
niceConfig: this._niceConfig,
});
console.log('[Session] Created mux session:', this._muxSession.muxName);
// No extra sleep — createSession() already waits for tmux readiness
}
// Attach to the mux session via PTY
try {
this.ptyProcess = pty.spawn(
this._mux.getAttachCommand(),
this._mux.getAttachArgs(this._muxSession!.muxName),
{
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildMuxAttachEnv(),
}
);
} catch (spawnErr) {
console.error('[Session] Failed to spawn PTY for shell mux attachment:', spawnErr);
this.emit('error', `Failed to attach to mux session: ${spawnErr}`);
throw spawnErr;
}
envOverrides: this._envOverrides,
},
spawnErrLabel: 'shell mux attachment',
});
// For NEW sessions: clear by sending 'clear' command to the shell
// For RESTORED sessions: don't clear - we want to see the existing output
if (!isRestoredSession) {
if (!isRestored) {
setTimeout(() => {
if (this.ptyProcess) {
this._terminalBuffer.clear();
@@ -1378,12 +1421,7 @@ export class Session extends EventEmitter {
const data = rawData.replace(FOCUS_ESCAPE_FILTER, '');
if (!data) return; // Skip if only focus sequences
// BufferAccumulator handles auto-trimming when max size exceeded
this._terminalBuffer.append(data);
this._lastActivityAt = Date.now();
this.emit('terminal', data);
this.emit('output', data);
this._handleTerminalOutput(data);
});
this.ptyProcess.onExit(({ exitCode }) => {
@@ -1448,13 +1486,7 @@ export class Session extends EventEmitter {
return;
}
this._status = 'busy';
this._terminalBuffer.clear();
this._textOutput.clear();
this._errorBuffer = '';
this._messages = [];
this._lineBuffer = '';
this._lastActivityAt = Date.now();
this._resetBuffers();
this._promptResolved = false; // Reset race condition guard
this.resolvePromise = resolve;
@@ -1477,7 +1509,8 @@ export class Session extends EventEmitter {
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildClaudeEnv(this.id),
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
});
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY for runPrompt:', spawnErr);
@@ -1497,12 +1530,7 @@ export class Session extends EventEmitter {
const data = rawData.replace(FOCUS_ESCAPE_FILTER, '');
if (!data) return; // Skip if only focus sequences
// BufferAccumulator handles auto-trimming when max size exceeded
this._terminalBuffer.append(data);
this._lastActivityAt = Date.now();
this.emit('terminal', data);
this.emit('output', data);
this._handleTerminalOutput(data);
// Also try to parse JSON lines for structured data
this.processOutput(data);
@@ -1534,9 +1562,11 @@ export class Session extends EventEmitter {
this._status = 'idle';
const cost = resultMsg.total_cost_usd || 0;
this._totalCost += cost;
this.emit('completion', resultMsg.result || '', cost);
// Claude CLI stream-json may return empty result field — fall back to accumulated text output
const result = resultMsg.result || this._textOutput.value || '';
this.emit('completion', result, cost);
if (resolve) {
resolve({ result: resultMsg.result || '', cost });
resolve({ result, cost });
}
} else if (exitCode !== 0 || (resultMsg && resultMsg.is_error)) {
this._status = 'error';
@@ -1565,6 +1595,120 @@ export class Session extends EventEmitter {
});
}
private _resetBuffers(): void {
this._status = 'busy';
this._terminalBuffer.clear();
this._textOutput.clear();
this._errorBuffer = '';
this._messages = [];
this._lineBuffer = '';
this._lastActivityAt = Date.now();
}
private _clearAllTimers(): void {
// Clear activity timeout to prevent memory leak
if (this.activityTimeout) {
clearTimeout(this.activityTimeout);
this.activityTimeout = null;
}
// Clear line buffer flush timer
if (this._lineBufferFlushTimer) {
clearTimeout(this._lineBufferFlushTimer);
this._lineBufferFlushTimer = null;
}
// Destroy auto-compact/auto-clear automation (clears its timers)
this._autoOps.destroy();
// Clear prompt check timers
if (this._promptCheckInterval) {
clearInterval(this._promptCheckInterval);
this._promptCheckInterval = null;
}
if (this._promptCheckTimeout) {
clearTimeout(this._promptCheckTimeout);
this._promptCheckTimeout = null;
}
// Clear shell idle timer
if (this._shellIdleTimer) {
clearTimeout(this._shellIdleTimer);
this._shellIdleTimer = null;
}
// Clear expensive processing timer
if (this._expensiveProcessTimer) {
clearTimeout(this._expensiveProcessTimer);
this._expensiveProcessTimer = null;
}
this._pendingCleanData = '';
}
private _handleJsonMessage(cleanLine: string, rawLine: string): void {
try {
const msg = JSON.parse(cleanLine) as ClaudeMessage;
this._messages.push(msg);
this.emit('message', msg);
// Trim messages array for long-running sessions
if (this._messages.length > MAX_MESSAGES) {
this._messages = this._messages.slice(-Math.floor(MAX_MESSAGES * 0.8));
}
// Extract Claude session ID from messages (can be in any message type).
// Support both sessionId (camelCase) and session_id (snake_case).
// The constructor seeds _claudeSessionId with this.id as a placeholder;
// once Claude CLI emits its real session ID, adopt it so JSONL lookups
// (e.g. /api/sessions/:id/last-response) can find the transcript file.
const msgSessionId =
((msg as unknown as Record<string, unknown>).sessionId as string | undefined) ?? msg.session_id;
if (msgSessionId && msgSessionId !== this._claudeSessionId) {
this._claudeSessionId = msgSessionId;
}
// Process message for task tracking
this._taskTracker.processMessage(msg);
if (msg.type === 'assistant' && msg.message?.content) {
for (const block of msg.message.content) {
if (block.type === 'text' && block.text) {
this._textOutput.append(block.text);
}
}
// Track tokens from usage (with validation)
if (msg.message.usage) {
const inputDelta = msg.message.usage.input_tokens || 0;
const outputDelta = msg.message.usage.output_tokens || 0;
// Sanity check: max 100k tokens per message (generous limit)
const MAX_TOKENS_PER_MESSAGE = 100_000;
if (inputDelta > 0 && inputDelta <= MAX_TOKENS_PER_MESSAGE) {
this._totalInputTokens += inputDelta;
}
if (outputDelta > 0 && outputDelta <= MAX_TOKENS_PER_MESSAGE) {
this._totalOutputTokens += outputDelta;
}
// Check if we should auto-compact or auto-clear
this._autoOps.checkAutoCompact();
this._autoOps.checkAutoClear();
}
}
if (msg.type === 'result' && msg.total_cost_usd) {
this._totalCost = msg.total_cost_usd;
}
} catch (parseErr) {
// Not JSON, just regular output - this is expected for non-JSON lines
console.debug(
'[Session] Line not JSON (expected for text output):',
parseErr instanceof Error ? parseErr.message : parseErr
);
this._textOutput.append(rawLine + '\n');
}
}
private processOutput(data: string): void {
// Early return if session is stopped to prevent any processing or timer creation
if (this._isStopped) return;
@@ -1606,64 +1750,7 @@ export class Session extends EventEmitter {
const cleanLine = trimmed.replace(ANSI_ESCAPE_PATTERN_FULL, '');
if (cleanLine.startsWith('{') && cleanLine.endsWith('}')) {
try {
const msg = JSON.parse(cleanLine) as ClaudeMessage;
this._messages.push(msg);
this.emit('message', msg);
// Trim messages array for long-running sessions
if (this._messages.length > MAX_MESSAGES) {
this._messages = this._messages.slice(-Math.floor(MAX_MESSAGES * 0.8));
}
// Extract Claude session ID from messages (can be in any message type)
// Support both sessionId (camelCase) and session_id (snake_case)
const msgSessionId =
((msg as unknown as Record<string, unknown>).sessionId as string | undefined) ?? msg.session_id;
if (msgSessionId && !this._claudeSessionId) {
this._claudeSessionId = msgSessionId;
}
// Process message for task tracking
this._taskTracker.processMessage(msg);
if (msg.type === 'assistant' && msg.message?.content) {
for (const block of msg.message.content) {
if (block.type === 'text' && block.text) {
this._textOutput.append(block.text);
}
}
// Track tokens from usage (with validation)
if (msg.message.usage) {
const inputDelta = msg.message.usage.input_tokens || 0;
const outputDelta = msg.message.usage.output_tokens || 0;
// Sanity check: max 100k tokens per message (generous limit)
const MAX_TOKENS_PER_MESSAGE = 100_000;
if (inputDelta > 0 && inputDelta <= MAX_TOKENS_PER_MESSAGE) {
this._totalInputTokens += inputDelta;
}
if (outputDelta > 0 && outputDelta <= MAX_TOKENS_PER_MESSAGE) {
this._totalOutputTokens += outputDelta;
}
// Check if we should auto-compact or auto-clear
this._autoOps.checkAutoCompact();
this._autoOps.checkAutoClear();
}
}
if (msg.type === 'result' && msg.total_cost_usd) {
this._totalCost = msg.total_cost_usd;
}
} catch (parseErr) {
// Not JSON, just regular output - this is expected for non-JSON lines
console.debug(
'[Session] Line not JSON (expected for text output):',
parseErr instanceof Error ? parseErr.message : parseErr
);
this._textOutput.append(line + '\n');
}
this._handleJsonMessage(cleanLine, line);
} else if (trimmed) {
this._textOutput.append(line + '\n');
}
@@ -1955,7 +2042,7 @@ export class Session extends EventEmitter {
this._status = 'busy';
this._lastActivityAt = Date.now();
this.runPrompt(input).catch((err) => {
const errorMsg = err instanceof Error ? err.message : String(err);
const errorMsg = getErrorMessage(err);
// Clean up task state so the task queue doesn't get stuck
if (this._currentTaskId) {
const taskId = this._currentTaskId;
@@ -2030,43 +2117,7 @@ export class Session extends EventEmitter {
// Set stopped flag first to prevent new timers from being created
this._isStopped = true;
// Clear activity timeout to prevent memory leak
if (this.activityTimeout) {
clearTimeout(this.activityTimeout);
this.activityTimeout = null;
}
// Clear line buffer flush timer
if (this._lineBufferFlushTimer) {
clearTimeout(this._lineBufferFlushTimer);
this._lineBufferFlushTimer = null;
}
// Destroy auto-compact/auto-clear automation (clears its timers)
this._autoOps.destroy();
// Clear prompt check timers
if (this._promptCheckInterval) {
clearInterval(this._promptCheckInterval);
this._promptCheckInterval = null;
}
if (this._promptCheckTimeout) {
clearTimeout(this._promptCheckTimeout);
this._promptCheckTimeout = null;
}
// Clear shell idle timer
if (this._shellIdleTimer) {
clearTimeout(this._shellIdleTimer);
this._shellIdleTimer = null;
}
// Clear expensive processing timer
if (this._expensiveProcessTimer) {
clearTimeout(this._expensiveProcessTimer);
this._expensiveProcessTimer = null;
}
this._pendingCleanData = '';
this._clearAllTimers();
// Immediately cleanup Promise callbacks to prevent orphaned references
// during the rest of stop() processing (e.g., if mux kill times out)
+98 -80
View File
@@ -116,6 +116,26 @@ export class StateStore {
this.loadRalphStates();
}
private _mergeWithInitialState(parsed: Partial<AppState>): AppState {
const initial = createInitialState();
return {
...initial,
...parsed,
sessions: { ...parsed.sessions },
tasks: { ...parsed.tasks },
ralphLoop: { ...initial.ralphLoop, ...parsed.ralphLoop },
config: { ...initial.config, ...parsed.config },
};
}
private _resetCircuitBreaker(): void {
this.consecutiveSaveFailures = 0;
if (this.circuitBreakerOpen) {
console.log('[StateStore] Circuit breaker CLOSED - save succeeded');
this.circuitBreakerOpen = false;
}
}
private ensureDir(): void {
const dir = dirname(this.filePath);
if (!existsSync(dir)) {
@@ -132,15 +152,7 @@ export class StateStore {
if (existsSync(path)) {
const data = readFileSync(path, 'utf-8');
const parsed = JSON.parse(data) as Partial<AppState>;
const initial = createInitialState();
const result = {
...initial,
...parsed,
sessions: { ...parsed.sessions },
tasks: { ...parsed.tasks },
ralphLoop: { ...initial.ralphLoop, ...parsed.ralphLoop },
config: { ...initial.config, ...parsed.config },
};
const result = this._mergeWithInitialState(parsed);
if (path !== this.filePath) {
console.warn(`[StateStore] Recovered state from backup: ${path}`);
}
@@ -195,16 +207,7 @@ export class StateStore {
* Only dirty sessions are re-serialized; clean sessions use cached JSON fragments.
*/
private assembleStateJson(): string {
// Re-serialize dirty sessions and update cache
for (const id of this.dirtySessions) {
const session = this.state.sessions[id];
if (session) {
this.cachedSessionJsons.set(id, JSON.stringify(session));
} else {
this.cachedSessionJsons.delete(id);
}
}
this.dirtySessions.clear();
this.updateDirtySessionCache();
// Build sessions object from cached fragments
const sessionParts: string[] = [];
@@ -218,6 +221,25 @@ export class StateStore {
sessionParts.push(`${JSON.stringify(id)}:${json}`);
}
this.pruneStaleCacheEntries();
return this.buildPartialJson(sessionParts);
}
private updateDirtySessionCache(): void {
// Re-serialize dirty sessions and update cache
for (const id of this.dirtySessions) {
const session = this.state.sessions[id];
if (session) {
this.cachedSessionJsons.set(id, JSON.stringify(session));
} else {
this.cachedSessionJsons.delete(id);
}
}
this.dirtySessions.clear();
}
private pruneStaleCacheEntries(): void {
// Prune stale cache entries (sessions removed via direct state mutation)
if (this.cachedSessionJsons.size > Object.keys(this.state.sessions).length) {
for (const cachedId of this.cachedSessionJsons.keys()) {
@@ -226,7 +248,9 @@ export class StateStore {
}
}
}
}
private buildPartialJson(sessionParts: string[]): string {
// Build final JSON: sessions from cache, everything else re-serialized (tiny)
const sessionsJson = `{${sessionParts.join(',')}}`;
@@ -249,6 +273,28 @@ export class StateStore {
return `{${parts.join(',')}}`;
}
private serializeState(): string | null {
try {
return this.assembleStateJson();
} catch (assembleErr) {
// Fallback to full serialization if incremental assembly fails
console.warn('[StateStore] assembleStateJson failed, falling back to full serialize:', assembleErr);
this.cachedSessionJsons.clear();
this.dirtySessions.clear();
try {
return JSON.stringify(this.state);
} catch (err) {
console.error('[StateStore] Failed to serialize state (circular reference or invalid data):', err);
this.consecutiveSaveFailures++;
if (this.consecutiveSaveFailures >= MAX_CONSECUTIVE_FAILURES) {
console.error('[StateStore] Circuit breaker OPEN - serialization failing repeatedly');
this.circuitBreakerOpen = true;
}
return null;
}
}
}
private async _doSaveAsync(): Promise<void> {
this.saveDeb.cancel();
if (!this.dirty) {
@@ -263,30 +309,12 @@ export class StateStore {
this.ensureDir();
const tempPath = this.filePath + '.tmp';
const tempPath = `${this.filePath}.${process.pid}.${Date.now()}.${Math.random().toString(36).slice(2)}.tmp`;
const backupPath = this.filePath + '.bak';
let json: string;
// Step 1: Serialize state (validates it's JSON-safe)
try {
json = this.assembleStateJson();
} catch (assembleErr) {
// Fallback to full serialization if incremental assembly fails
console.warn('[StateStore] assembleStateJson failed, falling back to full serialize:', assembleErr);
this.cachedSessionJsons.clear();
this.dirtySessions.clear();
try {
json = JSON.stringify(this.state);
} catch (err) {
console.error('[StateStore] Failed to serialize state (circular reference or invalid data):', err);
this.consecutiveSaveFailures++;
if (this.consecutiveSaveFailures >= MAX_CONSECUTIVE_FAILURES) {
console.error('[StateStore] Circuit breaker OPEN - serialization failing repeatedly');
this.circuitBreakerOpen = true;
}
return;
}
}
const json = this.serializeState();
if (json === null) return;
// Clear dirty flag BEFORE async I/O so mutations during write re-set it.
// The state snapshot is already captured in `json` above.
@@ -305,11 +333,7 @@ export class StateStore {
await writeFile(tempPath, json, 'utf-8');
await rename(tempPath, this.filePath);
this.consecutiveSaveFailures = 0;
if (this.circuitBreakerOpen) {
console.log('[StateStore] Circuit breaker CLOSED - save succeeded');
this.circuitBreakerOpen = false;
}
this._resetCircuitBreaker();
} catch (err) {
console.error('[StateStore] Failed to write state file:', err);
// Re-mark dirty so the data is retried on the next save cycle
@@ -349,29 +373,11 @@ export class StateStore {
this.ensureDir();
const tempPath = this.filePath + '.tmp';
const tempPath = `${this.filePath}.${process.pid}.${Date.now()}.${Math.random().toString(36).slice(2)}.tmp`;
const backupPath = this.filePath + '.bak';
let json: string;
try {
json = this.assembleStateJson();
} catch (assembleErr) {
// Fallback to full serialization if incremental assembly fails
console.warn('[StateStore] assembleStateJson failed, falling back to full serialize:', assembleErr);
this.cachedSessionJsons.clear();
this.dirtySessions.clear();
try {
json = JSON.stringify(this.state);
} catch (err) {
console.error('[StateStore] Failed to serialize state (circular reference or invalid data):', err);
this.consecutiveSaveFailures++;
if (this.consecutiveSaveFailures >= MAX_CONSECUTIVE_FAILURES) {
console.error('[StateStore] Circuit breaker OPEN - serialization failing repeatedly');
this.circuitBreakerOpen = true;
}
return;
}
}
const json = this.serializeState();
if (json === null) return;
// Backup via atomic copy (avoids reading entire file into memory)
try {
@@ -387,11 +393,7 @@ export class StateStore {
renameSync(tempPath, this.filePath);
// Clear dirty flag only AFTER successful write
this.dirty = false;
this.consecutiveSaveFailures = 0;
if (this.circuitBreakerOpen) {
console.log('[StateStore] Circuit breaker CLOSED - save succeeded');
this.circuitBreakerOpen = false;
}
this._resetCircuitBreaker();
} catch (err) {
console.error('[StateStore] Failed to write state file:', err);
this.consecutiveSaveFailures++;
@@ -417,15 +419,7 @@ export class StateStore {
if (existsSync(backupPath)) {
const backupContent = readFileSync(backupPath, 'utf-8');
const parsed = JSON.parse(backupContent) as Partial<AppState>;
const initial = createInitialState();
this.state = {
...initial,
...parsed,
sessions: { ...parsed.sessions },
tasks: { ...parsed.tasks },
ralphLoop: { ...initial.ralphLoop, ...parsed.ralphLoop },
config: { ...initial.config, ...parsed.config },
};
this.state = this._mergeWithInitialState(parsed);
console.log('[StateStore] Successfully recovered state from backup');
// Reset circuit breaker after successful recovery
this.circuitBreakerOpen = false;
@@ -547,6 +541,30 @@ export class StateStore {
this.save();
}
// ========== Orchestrator Loop State Methods ==========
/** Returns the orchestrator loop state, or null if never initialized. */
getOrchestratorState() {
return this.state.orchestrator ?? null;
}
/** Updates orchestrator loop state (partial merge) and triggers a debounced save. */
setOrchestratorState(orchestrator: Partial<NonNullable<AppState['orchestrator']>>) {
if (this.state.orchestrator) {
this.state.orchestrator = { ...this.state.orchestrator, ...orchestrator };
} else {
// First initialization — caller must provide full state
this.state.orchestrator = orchestrator as NonNullable<AppState['orchestrator']>;
}
this.save();
}
/** Clears orchestrator state and triggers a debounced save. */
clearOrchestratorState() {
this.state.orchestrator = undefined;
this.save();
}
/** Returns the application configuration. */
getConfig() {
return this.state.config;
+258 -232
View File
@@ -17,7 +17,7 @@
* Tracks per-agent: status, token counts, model, description, tool call count, liveness (PID).
*
* @dependencies config/map-limits (MAX_TRACKED_AGENTS, PENDING_TOOL_CALL_TTL_MS),
* utils (CleanupManager, KeyedDebouncer)
* config/buffer-limits (FILE_PEEK_BYTES), utils (CleanupManager, KeyedDebouncer)
* @consumedby web/server (SSE broadcast), session (subagent-session correlation)
* @emits subagent:discovered, subagent:updated, subagent:tool_call, subagent:tool_result,
* subagent:progress, subagent:message, subagent:completed
@@ -35,6 +35,7 @@ import { execFile } from 'node:child_process';
import { readFile, readdir, stat as statAsync } from 'node:fs/promises';
import { PENDING_TOOL_CALL_TTL_MS, MAX_PENDING_TOOL_CALLS, MAX_TRACKED_AGENTS } from './config/map-limits.js';
import { STALE_DATA_MAX_AGE_MS } from './config/server-timing.js';
import { FILE_PEEK_BYTES } from './config/buffer-limits.js';
import { CleanupManager, KeyedDebouncer } from './utils/index.js';
// ========== Types ==========
@@ -133,17 +134,6 @@ export interface SubagentToolResult {
isError: boolean; // Whether result is an error
}
export interface SubagentEvents {
'subagent:discovered': (info: SubagentInfo) => void;
'subagent:updated': (info: SubagentInfo) => void;
'subagent:tool_call': (data: SubagentToolCall) => void;
'subagent:tool_result': (data: SubagentToolResult) => void;
'subagent:progress': (data: SubagentProgress) => void;
'subagent:message': (data: SubagentMessage) => void;
'subagent:completed': (info: SubagentInfo) => void;
'subagent:error': (error: Error, agentId?: string) => void;
}
// ========== Constants ==========
const CLAUDE_PROJECTS_DIR = join(homedir(), '.claude/projects');
@@ -221,6 +211,83 @@ export class SubagentWatcher extends EventEmitter {
return INTERNAL_AGENT_PATTERNS.some((pattern) => pattern.test(description));
}
/**
* Mark a subagent as completed: clear PID, set status, clean up pending tool calls, emit event.
*/
private markSubagentAsCompleted(info: SubagentInfo): void {
info.pid = undefined;
info.status = 'completed';
this.pendingToolCalls.delete(info.agentId);
this.emit('subagent:completed', info);
}
/**
* Extract text from message content, handling both string and array formats.
* For array content, returns the text from the first 'text' block.
*/
private extractFirstTextContent(
content: string | Array<{ type: string; text?: string }> | undefined
): string | undefined {
if (!content) return undefined;
if (typeof content === 'string') {
const trimmed = content.trim();
return trimmed.length > 0 ? trimmed : undefined;
}
if (Array.isArray(content)) {
const firstContent = content[0];
if (firstContent?.type === 'text' && firstContent.text) {
const trimmed = firstContent.text.trim();
return trimmed.length > 0 ? trimmed : undefined;
}
}
return undefined;
}
/**
* Process a tool_result content block: look up pending tool call, emit tool_result event.
*/
private emitToolResult(
content: { tool_use_id: string; content?: string | Array<{ type: string; text?: string }>; is_error?: boolean },
agentId: string,
sessionId: string,
timestamp: string
): void {
const resultContent = this.extractToolResultContent(content.content);
const agentPendingCalls = this.pendingToolCalls.get(agentId);
const pendingCall = agentPendingCalls?.get(content.tool_use_id);
const toolName = pendingCall?.toolName;
// Delete after lookup to prevent memory leak
agentPendingCalls?.delete(content.tool_use_id);
const toolResult: SubagentToolResult = {
agentId,
sessionId,
timestamp,
toolUseId: content.tool_use_id,
tool: toolName,
preview: resultContent.substring(0, MESSAGE_TEXT_LIMIT),
contentLength: resultContent.length,
isError: content.is_error || false,
};
this.emit('subagent:tool_result', toolResult);
}
/**
* Find the oldest inactive (non-active) agent for LRU eviction.
* Returns the agent ID of the oldest inactive agent, or null if all are active.
*/
private findOldestInactiveAgent(): string | null {
let oldestId: string | null = null;
let oldestTime = Infinity;
for (const [id, existing] of this.agentInfo) {
if (existing.status !== 'active' && existing.lastActivityAt < oldestTime) {
oldestTime = existing.lastActivityAt;
oldestId = id;
}
}
return oldestId;
}
/**
* Extract short model identifier from full model name
*/
@@ -306,10 +373,7 @@ export class SubagentWatcher extends EventEmitter {
const alive = this.checkSubagentAliveFromPidMap(info, pidMap);
if (!alive) {
info.pid = undefined;
info.status = 'completed';
this.pendingToolCalls.delete(info.agentId);
this.emit('subagent:completed', info);
this.markSubagentAsCompleted(info);
}
}
}
@@ -676,10 +740,7 @@ export class SubagentWatcher extends EventEmitter {
const pid = await this.findSubagentProcess(info.sessionId);
if (pid) {
process.kill(pid, 'SIGTERM');
info.pid = undefined;
info.status = 'completed';
this.pendingToolCalls.delete(info.agentId);
this.emit('subagent:completed', info);
this.markSubagentAsCompleted(info);
return true;
}
} catch {
@@ -687,10 +748,7 @@ export class SubagentWatcher extends EventEmitter {
}
// Mark as completed even if we couldn't find the process
info.pid = undefined;
info.status = 'completed';
this.pendingToolCalls.delete(info.agentId);
this.emit('subagent:completed', info);
this.markSubagentAsCompleted(info);
return true;
}
@@ -842,19 +900,9 @@ export class SubagentWatcher extends EventEmitter {
}
} else if (entry.type === 'user' && entry.message?.content) {
// Handle both string and array content formats
if (typeof entry.message.content === 'string') {
const text = entry.message.content.trim();
if (text.length < 100 && !text.includes('{')) {
lines.push(`${this.formatTime(entry.timestamp)} 📥 User: ${text.substring(0, USER_TEXT_PREVIEW_LENGTH)}`);
}
} else {
const firstContent = entry.message.content[0];
if (firstContent?.type === 'text' && firstContent.text) {
const text = firstContent.text.trim();
if (text.length < 100 && !text.includes('{')) {
lines.push(`${this.formatTime(entry.timestamp)} 📥 User: ${text.substring(0, USER_TEXT_PREVIEW_LENGTH)}`);
}
}
const text = this.extractFirstTextContent(entry.message.content);
if (text && text.length < 100 && !text.includes('{')) {
lines.push(`${this.formatTime(entry.timestamp)} 📥 User: ${text.substring(0, USER_TEXT_PREVIEW_LENGTH)}`);
}
}
}
@@ -929,6 +977,24 @@ export class SubagentWatcher extends EventEmitter {
return truncated.replace(/[.!?,:\s]+$/, '');
}
private async _resolveDescription(
projectHash: string,
sessionId: string,
agentId: string,
filePath: string,
fallbackText?: string
): Promise<string | undefined> {
// First try parent transcript (most reliable)
const fromParent = await this.extractDescriptionFromParentTranscript(projectHash, sessionId, agentId);
if (fromParent) return fromParent;
// Fallback: inline text (from processEntry) or file extraction
if (fallbackText) {
return this.extractSmartTitle(fallbackText);
}
return this.extractDescriptionFromFile(filePath);
}
/**
* Extract the short description from the parent session's transcript.
* This is the most reliable method because it reads the actual Task tool result
@@ -1009,7 +1075,7 @@ export class SubagentWatcher extends EventEmitter {
private async extractDescriptionFromFile(filePath: string): Promise<string | undefined> {
try {
// Only read the first 8KB — more than enough for 5 JSONL lines
const stream = createReadStream(filePath, { end: 8191 });
const stream = createReadStream(filePath, { end: FILE_PEEK_BYTES });
const rl = createInterface({ input: stream });
return await new Promise<string | undefined>((resolve) => {
@@ -1026,15 +1092,7 @@ export class SubagentWatcher extends EventEmitter {
try {
const entry = JSON.parse(line);
if (entry.type === 'user' && entry.message?.content) {
let text: string | undefined;
if (typeof entry.message.content === 'string') {
text = entry.message.content.trim();
} else if (Array.isArray(entry.message.content)) {
const firstContent = entry.message.content[0];
if (firstContent?.type === 'text' && firstContent.text) {
text = firstContent.text.trim();
}
}
const text = this.extractFirstTextContent(entry.message.content);
if (text) {
resolved = true;
rl.close();
@@ -1141,10 +1199,10 @@ export class SubagentWatcher extends EventEmitter {
if (this.fileAgentContext.has(filePath)) {
// Known file — handle content change
this.handleFileChange(filePath).catch(() => {});
this.handleFileChange(filePath).catch(() => {}); // Ignore - errors logged internally, don't crash watcher callback
} else {
// New file — register it
this.registerAgentFile(filePath, projectHash, sessionId).catch(() => {});
this.registerAgentFile(filePath, projectHash, sessionId).catch(() => {}); // Ignore - errors logged internally, don't crash watcher callback
}
});
});
@@ -1194,18 +1252,13 @@ export class SubagentWatcher extends EventEmitter {
// Retry description extraction if missing (race condition fix)
if (!existingInfo.description) {
// First try parent transcript (most reliable)
let extractedDescription = await this.extractDescriptionFromParentTranscript(
const extractedDescription = await this._resolveDescription(
existingInfo.projectHash,
existingInfo.sessionId,
agentId
agentId,
filePath
);
// Fallback to subagent file
if (!extractedDescription) {
extractedDescription = await this.extractDescriptionFromFile(filePath);
}
if (extractedDescription) {
// Check if this is an internal agent - if so, remove it
if (this.isInternalAgent(extractedDescription)) {
this.removeAgent(agentId);
return;
@@ -1256,13 +1309,7 @@ export class SubagentWatcher extends EventEmitter {
}
// Extract description - prefer reading from parent transcript (most reliable)
// The parent transcript has the exact Task tool call with description parameter
let description = await this.extractDescriptionFromParentTranscript(projectHash, sessionId, agentId);
// Fallback: extract a smart title from the subagent's prompt if parent lookup failed
if (!description) {
description = await this.extractDescriptionFromFile(filePath);
}
const description = await this._resolveDescription(projectHash, sessionId, agentId, filePath);
// Skip internal Claude Code agents (e.g., suggestion mode) - not real subagents
if (this.isInternalAgent(description)) {
@@ -1285,14 +1332,7 @@ export class SubagentWatcher extends EventEmitter {
// Enforce MAX_TRACKED_AGENTS during insertion — evict oldest inactive agent
if (this.agentInfo.size >= MAX_TRACKED_AGENTS) {
let oldestId: string | null = null;
let oldestTime = Infinity;
for (const [id, existing] of this.agentInfo) {
if (existing.status !== 'active' && existing.lastActivityAt < oldestTime) {
oldestTime = existing.lastActivityAt;
oldestId = id;
}
}
const oldestId = this.findOldestInactiveAgent();
if (oldestId) {
this.removeAgent(oldestId);
}
@@ -1362,51 +1402,11 @@ export class SubagentWatcher extends EventEmitter {
private async processEntry(entry: SubagentTranscriptEntry, agentId: string, sessionId: string): Promise<void> {
const info = this.agentInfo.get(agentId);
// Extract model from assistant messages (first one sets the model)
if (info && entry.type === 'assistant' && entry.message?.model && !info.model) {
info.model = entry.message.model;
info.modelShort = this.extractModelShort(entry.message.model);
this.emit('subagent:updated', info);
}
if (info) {
this._processModelInfo(entry, info);
this._processTokenInfo(entry, info);
// Aggregate token usage from messages
if (info && entry.message?.usage) {
if (entry.message.usage.input_tokens) {
info.totalInputTokens = (info.totalInputTokens || 0) + entry.message.usage.input_tokens;
}
if (entry.message.usage.output_tokens) {
info.totalOutputTokens = (info.totalOutputTokens || 0) + entry.message.usage.output_tokens;
}
}
// Check if this is first user message and description is missing
if (info && !info.description && entry.type === 'user' && entry.message?.content) {
// First try parent transcript (most reliable)
let description = await this.extractDescriptionFromParentTranscript(info.projectHash, info.sessionId, agentId);
// Fallback: extract smart title from the prompt content
if (!description) {
let text: string | undefined;
if (typeof entry.message.content === 'string') {
text = entry.message.content.trim();
} else if (Array.isArray(entry.message.content)) {
const firstContent = entry.message.content[0];
if (firstContent?.type === 'text' && firstContent.text) {
text = firstContent.text.trim();
}
}
if (text) {
description = this.extractSmartTitle(text);
}
}
if (description) {
// Check if this is an internal agent - if so, remove it
if (this.isInternalAgent(description)) {
this.removeAgent(agentId);
return;
}
info.description = description;
this.emit('subagent:updated', info);
}
if (await this._processDescription(entry, agentId, info)) return;
}
if (entry.type === 'progress' && entry.data) {
@@ -1417,7 +1417,6 @@ export class SubagentWatcher extends EventEmitter {
progressType: entry.data.type,
query: entry.data.query,
resultCount: entry.data.resultCount,
// Extract hook event info if present
hookEvent: entry.data.hookEvent,
hookName:
entry.data.hookName ||
@@ -1427,139 +1426,166 @@ export class SubagentWatcher extends EventEmitter {
};
this.emit('subagent:progress', progress);
} else if (entry.type === 'assistant' && entry.message?.content) {
// Handle both string and array content formats
if (typeof entry.message.content === 'string') {
const text = entry.message.content.trim();
if (text.length > 0) {
const message: SubagentMessage = {
this._processAssistantContent(entry, agentId, sessionId);
} else if (entry.type === 'user' && entry.message?.content) {
this._processUserContent(entry, agentId, sessionId);
}
}
private _processModelInfo(entry: SubagentTranscriptEntry, agent: SubagentInfo): void {
if (entry.type === 'assistant' && entry.message?.model && !agent.model) {
agent.model = entry.message.model;
agent.modelShort = this.extractModelShort(entry.message.model);
this.emit('subagent:updated', agent);
}
}
private _processTokenInfo(entry: SubagentTranscriptEntry, agent: SubagentInfo): void {
if (!entry.message?.usage) return;
if (entry.message.usage.input_tokens) {
agent.totalInputTokens = (agent.totalInputTokens || 0) + entry.message.usage.input_tokens;
}
if (entry.message.usage.output_tokens) {
agent.totalOutputTokens = (agent.totalOutputTokens || 0) + entry.message.usage.output_tokens;
}
}
private async _processDescription(
entry: SubagentTranscriptEntry,
agentId: string,
agent: SubagentInfo
): Promise<boolean> {
if (agent.description || entry.type !== 'user' || !entry.message?.content) return false;
const fallbackText = this.extractFirstTextContent(entry.message.content);
const description = await this._resolveDescription(
agent.projectHash,
agent.sessionId,
agentId,
agent.filePath,
fallbackText
);
if (description) {
if (this.isInternalAgent(description)) {
this.removeAgent(agentId);
return true;
}
agent.description = description;
this.emit('subagent:updated', agent);
}
return false;
}
private _processAssistantContent(entry: SubagentTranscriptEntry, agentId: string, sessionId: string): void {
const messageContent = entry.message!.content;
if (typeof messageContent === 'string') {
const text = messageContent.trim();
if (text.length > 0) {
const message: SubagentMessage = {
agentId,
sessionId,
timestamp: entry.timestamp,
role: 'assistant',
text: text.substring(0, MESSAGE_TEXT_LIMIT),
};
this.emit('subagent:message', message);
}
} else {
for (const content of messageContent) {
if (content.type === 'tool_use' && content.name) {
// Store toolUseId for linking to results, with timestamp for TTL cleanup
if (content.id) {
if (!this.pendingToolCalls.has(agentId)) {
this.pendingToolCalls.set(agentId, new Map());
}
const agentCalls = this.pendingToolCalls.get(agentId)!;
// Enforce size limit to prevent memory leak from rapid tool calls
if (agentCalls.size >= MAX_PENDING_TOOL_CALLS) {
// FIFO eviction: delete first (oldest) entry using Map insertion order
const firstKey = agentCalls.keys().next().value;
if (firstKey !== undefined) agentCalls.delete(firstKey);
}
agentCalls.set(content.id, {
toolName: content.name,
timestamp: Date.now(),
});
}
const toolCall: SubagentToolCall = {
agentId,
sessionId,
timestamp: entry.timestamp,
role: 'assistant',
text: text.substring(0, MESSAGE_TEXT_LIMIT),
tool: content.name,
input: this.getTruncatedInput(content.name, content.input || {}),
toolUseId: content.id,
fullInput: content.input || {},
};
this.emit('subagent:message', message);
}
} else {
for (const content of entry.message.content) {
if (content.type === 'tool_use' && content.name) {
// Store toolUseId for linking to results, with timestamp for TTL cleanup
if (content.id) {
if (!this.pendingToolCalls.has(agentId)) {
this.pendingToolCalls.set(agentId, new Map());
}
const agentCalls = this.pendingToolCalls.get(agentId)!;
// Enforce size limit to prevent memory leak from rapid tool calls
if (agentCalls.size >= MAX_PENDING_TOOL_CALLS) {
// FIFO eviction: delete first (oldest) entry using Map insertion order
const firstKey = agentCalls.keys().next().value;
if (firstKey !== undefined) agentCalls.delete(firstKey);
}
agentCalls.set(content.id, {
toolName: content.name,
timestamp: Date.now(),
});
}
this.emit('subagent:tool_call', toolCall);
const toolCall: SubagentToolCall = {
// Update tool call count
const agentInfo = this.agentInfo.get(agentId);
if (agentInfo) {
agentInfo.toolCallCount++;
}
} else if (content.type === 'tool_result' && content.tool_use_id) {
this.emitToolResult(
{ tool_use_id: content.tool_use_id, content: content.content, is_error: content.is_error },
agentId,
sessionId,
entry.timestamp
);
} else if (content.type === 'text' && content.text) {
const text = content.text.trim();
if (text.length > 0) {
const message: SubagentMessage = {
agentId,
sessionId,
timestamp: entry.timestamp,
tool: content.name,
input: this.getTruncatedInput(content.name, content.input || {}),
toolUseId: content.id,
fullInput: content.input || {},
role: 'assistant',
text: text.substring(0, MESSAGE_TEXT_LIMIT),
};
this.emit('subagent:tool_call', toolCall);
// Update tool call count
const agentInfo = this.agentInfo.get(agentId);
if (agentInfo) {
agentInfo.toolCallCount++;
}
} else if (content.type === 'tool_result' && content.tool_use_id) {
// Extract tool result
const resultContent = this.extractToolResultContent(content.content);
const agentPendingCalls = this.pendingToolCalls.get(agentId);
const pendingCall = agentPendingCalls?.get(content.tool_use_id);
const toolName = pendingCall?.toolName;
// Delete after lookup to prevent memory leak
agentPendingCalls?.delete(content.tool_use_id);
const toolResult: SubagentToolResult = {
agentId,
sessionId,
timestamp: entry.timestamp,
toolUseId: content.tool_use_id,
tool: toolName,
preview: resultContent.substring(0, MESSAGE_TEXT_LIMIT),
contentLength: resultContent.length,
isError: content.is_error || false,
};
this.emit('subagent:tool_result', toolResult);
} else if (content.type === 'text' && content.text) {
const text = content.text.trim();
if (text.length > 0) {
const message: SubagentMessage = {
agentId,
sessionId,
timestamp: entry.timestamp,
role: 'assistant',
text: text.substring(0, MESSAGE_TEXT_LIMIT), // Limit text length
};
this.emit('subagent:message', message);
}
this.emit('subagent:message', message);
}
}
}
} else if (entry.type === 'user' && entry.message?.content) {
// Handle both string and array content formats - also check for tool_result in user messages
if (typeof entry.message.content === 'string') {
const userText = entry.message.content.trim();
if (userText.length > 0 && userText.length < 500) {
const message: SubagentMessage = {
}
}
private _processUserContent(entry: SubagentTranscriptEntry, agentId: string, sessionId: string): void {
const messageContent = entry.message!.content;
if (typeof messageContent === 'string') {
const userText = messageContent.trim();
if (userText.length > 0 && userText.length < 500) {
const message: SubagentMessage = {
agentId,
sessionId,
timestamp: entry.timestamp,
role: 'user',
text: userText,
};
this.emit('subagent:message', message);
}
} else {
// Check for tool_result blocks in user messages (common pattern)
for (const content of messageContent) {
if (content.type === 'tool_result' && content.tool_use_id) {
this.emitToolResult(
{ tool_use_id: content.tool_use_id, content: content.content, is_error: content.is_error },
agentId,
sessionId,
timestamp: entry.timestamp,
role: 'user',
text: userText,
};
this.emit('subagent:message', message);
}
} else {
// Check for tool_result blocks in user messages (common pattern)
for (const content of entry.message.content) {
if (content.type === 'tool_result' && content.tool_use_id) {
const resultContent = this.extractToolResultContent(content.content);
const agentPendingCalls = this.pendingToolCalls.get(agentId);
const pendingCall = agentPendingCalls?.get(content.tool_use_id);
const toolName = pendingCall?.toolName;
// Delete after lookup to prevent memory leak
agentPendingCalls?.delete(content.tool_use_id);
const toolResult: SubagentToolResult = {
entry.timestamp
);
} else if (content.type === 'text' && content.text) {
const userText = content.text.trim();
if (userText.length > 0 && userText.length < 500) {
const message: SubagentMessage = {
agentId,
sessionId,
timestamp: entry.timestamp,
toolUseId: content.tool_use_id,
tool: toolName,
preview: resultContent.substring(0, MESSAGE_TEXT_LIMIT),
contentLength: resultContent.length,
isError: content.is_error || false,
role: 'user',
text: userText,
};
this.emit('subagent:tool_result', toolResult);
} else if (content.type === 'text' && content.text) {
const userText = content.text.trim();
if (userText.length > 0 && userText.length < 500) {
const message: SubagentMessage = {
agentId,
sessionId,
timestamp: entry.timestamp,
role: 'user',
text: userText,
};
this.emit('subagent:message', message);
}
this.emit('subagent:message', message);
}
}
}
-8
View File
@@ -17,14 +17,6 @@ import { getStore } from './state-store.js';
/**
* Events emitted by TaskQueue
*/
export interface TaskQueueEvents {
/** Fired when a task is added to the queue */
taskAdded: (task: Task) => void;
/** Fired when a task is removed from the queue */
taskRemoved: (taskId: string) => void;
/** Fired when a task's state changes */
taskUpdated: (task: Task) => void;
}
/**
* Priority queue for managing tasks with dependency support.
-10
View File
@@ -153,16 +153,6 @@ export interface BackgroundTask {
* @event taskCompleted - Task finished successfully
* @event taskFailed - Task finished with error
*/
export interface TaskTrackerEvents {
/** New task created */
taskCreated: (task: BackgroundTask) => void;
/** Task state updated */
taskUpdated: (task: BackgroundTask) => void;
/** Task completed successfully */
taskCompleted: (task: BackgroundTask) => void;
/** Task failed with error */
taskFailed: (task: BackgroundTask, error: string) => void;
}
/**
* TaskTracker - Detects and tracks background tasks in Claude Code sessions.
+5 -5
View File
@@ -61,7 +61,7 @@ export class TeamWatcher extends EventEmitter {
persistent: false,
});
const teamsHandler = () => this.pollAsync().catch(() => {});
const teamsHandler = () => this.pollAsync().catch(() => {}); // Ignore - poll errors are non-fatal, next poll will retry
this.teamsWatcher.on('add', teamsHandler);
this.teamsWatcher.on('change', teamsHandler);
this.teamsWatcher.on('unlink', teamsHandler);
@@ -82,8 +82,8 @@ export class TeamWatcher extends EventEmitter {
persistent: false,
});
this.tasksWatcher.on('add', () => this.pollTasks().catch(() => {}));
this.tasksWatcher.on('change', () => this.pollTasks().catch(() => {}));
this.tasksWatcher.on('add', () => this.pollTasks().catch(() => {})); // Ignore - poll errors are non-fatal, next poll will retry
this.tasksWatcher.on('change', () => this.pollTasks().catch(() => {})); // Ignore - poll errors are non-fatal, next poll will retry
this.tasksWatcher.on('error', (err) => {
console.warn('[TeamWatcher] chokidar tasks watcher error:', err);
});
@@ -95,11 +95,11 @@ export class TeamWatcher extends EventEmitter {
stop(): void {
// Close chokidar watchers
if (this.teamsWatcher) {
this.teamsWatcher.close().catch(() => {});
this.teamsWatcher.close().catch(() => {}); // Ignore - watcher cleanup is best-effort during shutdown
this.teamsWatcher = null;
}
if (this.tasksWatcher) {
this.tasksWatcher.close().catch(() => {});
this.tasksWatcher.close().catch(() => {}); // Ignore - watcher cleanup is best-effort during shutdown
this.tasksWatcher = null;
}
if (this.pollTimer) {
+191 -123
View File
@@ -40,8 +40,7 @@ import {
type SessionMode,
type OpenCodeConfig,
} from './types.js';
import { wrapWithNice } from './utils/nice-wrapper.js';
import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';
import { wrapWithNice, SAFE_PATH_PATTERN, findClaudeDir, resolveOpenCodeDir } from './utils/index.js';
import type {
TerminalMultiplexer,
MuxSession,
@@ -50,11 +49,6 @@ import type {
RespawnPaneOptions,
} from './mux-interface.js';
// Claude CLI PATH resolution — shared utility
import { findClaudeDir } from './utils/claude-cli-resolver.js';
// OpenCode CLI PATH resolution
import { resolveOpenCodeDir } from './utils/opencode-cli-resolver.js';
// ============================================================================
// Timing Constants
// ============================================================================
@@ -104,6 +98,47 @@ const LEGACY_MUX_NAME_PATTERN = /^claudeman-[a-f0-9-]+$/;
/** Regex to validate tmux pane targets (e.g., "%0", "%1", "0", "1") */
const SAFE_PANE_TARGET_PATTERN = /^(%\d+|\d+)$/;
/**
* Separator used in `tmux list-panes -F` output between session name and pid.
*
* Must NOT be a backslash-escape (e.g. `\t`, `\n`): under non-tty execution
* contexts (launchd on macOS, systemd without TTYPath) tmux can emit such
* escapes as the literal two characters `\` + letter rather than the control
* byte, breaking the parser and causing every tracked session to be classified
* as dead — which wipes state.json on restart. '|' is passed through verbatim
* in every environment and is rejected by tmux's own session-name validation,
* so it cannot appear inside `#{session_name}` and cause a false split.
*/
const PANE_LIST_SEP = '|';
/** Format string for `tmux list-panes -F`. Keep in sync with {@link parsePaneList}. */
const PANE_LIST_FORMAT = `#{session_name}${PANE_LIST_SEP}#{pane_pid}`;
/**
* Parse the output of `tmux list-panes -a -F '#{session_name}|#{pane_pid}'`
* into a Map of session-name → pane pid. Exported for unit testing.
*
* - Skips empty lines and lines without the separator.
* - Skips entries with a non-numeric pid or empty name.
*/
export function parsePaneList(output: string): Map<string, number> {
const result = new Map<string, number>();
for (const line of output.split('\n')) {
if (!line) continue;
const sep = line.indexOf(PANE_LIST_SEP);
if (sep === -1) continue;
const name = line.slice(0, sep);
const pid = parseInt(line.slice(sep + 1), 10);
if (name && !Number.isNaN(pid)) {
result.set(name, pid);
}
}
return result;
}
/** Characters unsafe in paths — shell metacharacters, quotes, and control chars */
const UNSAFE_PATH_CHARS = /[;&|$`(){}<>'"\n\r]/;
/**
* Validates that a session name contains only safe characters.
* Prevents command injection via malformed session IDs.
@@ -117,23 +152,7 @@ function isValidMuxName(name: string): boolean {
* Prevents command injection via malformed paths.
*/
function isValidPath(path: string): boolean {
if (
path.includes(';') ||
path.includes('&') ||
path.includes('|') ||
path.includes('$') ||
path.includes('`') ||
path.includes('(') ||
path.includes(')') ||
path.includes('{') ||
path.includes('}') ||
path.includes('<') ||
path.includes('>') ||
path.includes("'") ||
path.includes('"') ||
path.includes('\n') ||
path.includes('\r')
) {
if (UNSAFE_PATH_CHARS.test(path)) {
return false;
}
if (path.includes('..')) {
@@ -204,8 +223,8 @@ function buildSpawnCommand(options: {
}): string {
if (options.mode === 'claude') {
// Validate model to prevent command injection
const safeModel = options.model && /^[a-zA-Z0-9._-]+$/.test(options.model) ? options.model : undefined;
const modelFlag = safeModel ? ` --model ${safeModel}` : '';
const safeModel = options.model && /^[a-zA-Z0-9._\-[\]]+$/.test(options.model) ? options.model : undefined;
const modelFlag = safeModel ? ` --model "${safeModel}"` : '';
// Use --resume to restore a previous conversation, otherwise --session-id for new sessions.
// Wrap --resume in a fallback: if it exits non-zero (session not found, corrupt, etc.),
// fall back to a new session with --session-id so the pane doesn't die.
@@ -377,6 +396,85 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
}
/**
* Build the array of environment export commands shared by createSession() and respawnPane().
* Includes locale, mux markers, session identity, and API URL.
*
* User-supplied envOverrides are NOT inlined here — they go through applyEnvOverrides()
* via `tmux setenv` so secret values (e.g., OPENCODE_API_KEY) never appear in the bash
* command line (visible in `ps`). This also sidesteps shell-metachar injection via keys.
*/
private buildEnvExports(sessionId: string, muxName: string, mode: SessionMode): string[] {
const exports = [
'export LANG=en_US.UTF-8',
'export LC_ALL=en_US.UTF-8',
'unset COLORTERM',
'export CODEMAN_MUX=1',
`export CODEMAN_SESSION_ID=${sessionId}`,
`export CODEMAN_MUX_NAME=${muxName}`,
`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL || 'http://localhost:3000'}`,
];
// Only unset CLAUDECODE for Claude sessions
if (mode === 'claude') exports.splice(2, 0, 'unset CLAUDECODE');
return exports;
}
/**
* Apply user-supplied env overrides to a tmux session via `tmux setenv`.
* Values stay off the bash command line (not visible in `ps`), and are inherited
* by new panes — including `respawn-pane`. Persists at tmux-session level, so
* Codeman server restarts don't lose the setting as long as the tmux session lives.
*
* Key validation is strict (`/^[A-Z_][A-Z0-9_]*$/`) as defense-in-depth against
* shell-metachar injection even if upstream schema check is bypassed.
*/
private applyEnvOverrides(muxName: string, envOverrides?: Record<string, string>): void {
if (!envOverrides) return;
const VALID_KEY = /^[A-Z_][A-Z0-9_]*$/;
for (const [key, value] of Object.entries(envOverrides)) {
if (!value) continue; // Skip empty — nothing to set
if (!VALID_KEY.test(key)) {
console.warn(`[TmuxManager] Skipping invalid env override key: ${JSON.stringify(key)}`);
continue;
}
try {
execSync(`tmux setenv -t ${shellescape(muxName)} ${key} ${shellescape(value)}`, {
timeout: EXEC_TIMEOUT_MS,
stdio: ['pipe', 'pipe', 'pipe'],
});
} catch (err) {
console.warn(`[TmuxManager] Failed to set env override ${key}:`, err);
}
}
}
/**
* Resolve the CLI binary directory and return the PATH export prefix string.
* Returns '' if no override is needed (shell mode) or the binary dir is not found.
* In createSession(), a missing binary dir throws — the caller handles that separately.
*/
private buildPathExport(mode: SessionMode): { pathExport: string; dir: string | null } {
if (mode === 'claude') {
const dir = findClaudeDir();
return { pathExport: dir ? `export PATH="${dir}:$PATH" && ` : '', dir };
}
if (mode === 'opencode') {
const dir = resolveOpenCodeDir();
return { pathExport: dir ? `export PATH="${dir}:$PATH" && ` : '', dir };
}
return { pathExport: '', dir: null };
}
/**
* Configure OpenCode-specific environment on a tmux session.
* Sets sensitive API keys and config content via tmux setenv
* (not visible in ps output or tmux history, inherited by panes).
*/
private _configureOpenCode(muxName: string, openCodeConfig?: OpenCodeConfig): void {
setOpenCodeEnvVars(muxName);
setOpenCodeConfigContent(muxName, openCodeConfig);
}
/**
* Creates a new tmux session wrapping Claude CLI or a shell.
* In test mode: creates an in-memory session only (no real tmux session).
@@ -393,6 +491,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
allowedTools,
openCodeConfig,
resumeSessionId,
envOverrides,
} = options;
const muxName = `codeman-${sessionId.slice(0, 8)}`;
@@ -421,33 +520,15 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
// Resolve CLI binary directory based on mode
let pathExport = '';
if (mode === 'claude') {
const claudeDir = findClaudeDir();
if (!claudeDir) {
throw new Error('Claude CLI not found. Install it with: curl -fsSL https://claude.ai/install.sh | bash');
}
pathExport = `export PATH="${claudeDir}:$PATH" && `;
} else if (mode === 'opencode') {
const openCodeDir = resolveOpenCodeDir();
if (!openCodeDir) {
throw new Error('OpenCode CLI not found. Install with: curl -fsSL https://opencode.ai/install | bash');
}
pathExport = `export PATH="${openCodeDir}:$PATH" && `;
const { pathExport, dir: cliDir } = this.buildPathExport(mode);
if (mode === 'claude' && !cliDir) {
throw new Error('Claude CLI not found. Install it with: curl -fsSL https://claude.ai/install.sh | bash');
}
if (mode === 'opencode' && !cliDir) {
throw new Error('OpenCode CLI not found. Install with: curl -fsSL https://opencode.ai/install | bash');
}
const envExports = [
'export LANG=en_US.UTF-8',
'export LC_ALL=en_US.UTF-8',
'unset COLORTERM',
'export CODEMAN_MUX=1',
`export CODEMAN_SESSION_ID=${sessionId}`,
`export CODEMAN_MUX_NAME=${muxName}`,
`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL || 'http://localhost:3000'}`,
];
// Only unset CLAUDECODE for Claude sessions
if (mode === 'claude') envExports.splice(2, 0, 'unset CLAUDECODE');
const envExportsStr = envExports.join(' && ');
const envExportsStr = this.buildEnvExports(sessionId, muxName, mode).join(' && ');
const baseCmd = buildSpawnCommand({
mode,
@@ -475,7 +556,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
// (Production uses systemd which has a clean env, but dev/test may be nested.)
const cleanEnv = { ...process.env };
delete cleanEnv.TMUX;
execSync(`tmux new-session -ds "${muxName}" -c "${workingDir}" -x 120 -y 40`, {
execSync(`tmux new-session -ds "${muxName}" -c "${workingDir}"`, {
cwd: workingDir,
timeout: EXEC_TIMEOUT_MS,
stdio: 'ignore',
@@ -495,10 +576,13 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
// For OpenCode: set sensitive env vars and config via tmux setenv
// (not visible in ps output or tmux history, inherited by panes)
if (mode === 'opencode') {
setOpenCodeEnvVars(muxName);
setOpenCodeConfigContent(muxName, openCodeConfig);
this._configureOpenCode(muxName, openCodeConfig);
}
// Apply user-supplied env overrides (e.g., CLAUDE_CODE_EFFORT_LEVEL) via tmux setenv
// so secret values stay off the bash command line. Must run before respawn-pane.
this.applyEnvOverrides(muxName, envOverrides);
// Replace the shell with the actual command (no echo in terminal)
execSync(`tmux respawn-pane -k -t "${muxName}" bash -c ${JSON.stringify(fullCmd)}`, {
timeout: EXEC_TIMEOUT_MS,
@@ -525,6 +609,13 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
.catch(() => {
/* Already set globally as fallback */
}),
// Raise tmux scrollback from its 2000-line default so re-attach preserves
// more context. Matches the xterm-side default in constants.js.
execAsync(`tmux set-option -t "${muxName}" history-limit 50000`, { timeout: EXEC_TIMEOUT_MS })
.then(() => {})
.catch(() => {
/* Non-critical — falls back to tmux default */
}),
];
// Enable 24-bit true color passthrough — server-wide, set once per lifetime
@@ -639,6 +730,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
allowedTools,
openCodeConfig,
resumeSessionId,
envOverrides,
} = options;
const session = this.sessions.get(sessionId);
if (!session) return null;
@@ -647,26 +739,9 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
if (!isValidMuxName(muxName) || !isValidPath(workingDir)) return null;
// Resolve CLI binary directory based on mode
let pathExport = '';
if (mode === 'claude') {
const claudeDir = findClaudeDir();
pathExport = claudeDir ? `export PATH="${claudeDir}:$PATH" && ` : '';
} else if (mode === 'opencode') {
const openCodeDir = resolveOpenCodeDir();
pathExport = openCodeDir ? `export PATH="${openCodeDir}:$PATH" && ` : '';
}
const { pathExport } = this.buildPathExport(mode);
const envExports = [
'export LANG=en_US.UTF-8',
'export LC_ALL=en_US.UTF-8',
'unset COLORTERM',
'export CODEMAN_MUX=1',
`export CODEMAN_SESSION_ID=${sessionId}`,
`export CODEMAN_MUX_NAME=${muxName}`,
`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL || 'http://localhost:3000'}`,
];
if (mode === 'claude') envExports.splice(2, 0, 'unset CLAUDECODE');
const envExportsStr = envExports.join(' && ');
const envExportsStr = this.buildEnvExports(sessionId, muxName, mode).join(' && ');
const baseCmd = buildSpawnCommand({
mode,
@@ -684,10 +759,12 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
try {
// For OpenCode: set sensitive env vars via tmux setenv before respawn
if (mode === 'opencode') {
setOpenCodeEnvVars(muxName);
setOpenCodeConfigContent(muxName, openCodeConfig);
this._configureOpenCode(muxName, openCodeConfig);
}
// Re-apply user env overrides before respawn so the new shell inherits them.
this.applyEnvOverrides(muxName, envOverrides);
await execAsync(`tmux respawn-pane -k -t "${muxName}" bash -c ${JSON.stringify(fullCmd)}`, {
timeout: EXEC_TIMEOUT_MS,
});
@@ -911,13 +988,24 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const dead: string[] = [];
const discovered: string[] = [];
// Check known sessions
// Batch: single tmux call to get all session names + pane PIDs (replaces N per-session subprocess calls)
let activeSessions = new Map<string, number>();
try {
const output = execSync(`tmux list-panes -a -F '${PANE_LIST_FORMAT}' 2>/dev/null || true`, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
activeSessions = parsePaneList(output);
} catch (err) {
console.error('[TmuxManager] Failed to list tmux panes:', err);
}
// Check known sessions against the batch result (O(1) map lookup instead of subprocess per session)
for (const [sessionId, session] of this.sessions) {
if (this.sessionExists(session.muxName)) {
const pid = activeSessions.get(session.muxName);
if (pid !== undefined) {
alive.push(sessionId);
// Update PID if it changed
const pid = this.getPanePid(session.muxName);
if (pid && pid !== session.pid) {
if (pid !== session.pid) {
session.pid = pid;
}
} else {
@@ -927,51 +1015,31 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
}
// Discover unknown codeman sessions
try {
const output = execSync("tmux list-sessions -F '#{session_name}' 2>/dev/null || true", {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
// Discover unknown codeman/claudeman sessions from the same batch result
const knownMuxNames = new Set<string>();
for (const session of this.sessions.values()) {
knownMuxNames.add(session.muxName);
}
for (const line of output.split('\n')) {
const sessionName = line.trim();
if (!sessionName || (!sessionName.startsWith('codeman-') && !sessionName.startsWith('claudeman-'))) continue;
for (const [sessionName, pid] of activeSessions) {
if (!sessionName.startsWith('codeman-') && !sessionName.startsWith('claudeman-')) continue;
if (knownMuxNames.has(sessionName)) continue;
// Check if this session is already known
let isKnown = false;
for (const session of this.sessions.values()) {
if (session.muxName === sessionName) {
isKnown = true;
break;
}
}
if (!isKnown) {
// Extract session ID fragment from name
const fragment = sessionName.replace(/^(?:codeman|claudeman)-/, '');
const sessionId = `restored-${fragment}`;
const pid = this.getPanePid(sessionName);
if (pid) {
const session: MuxSession = {
sessionId,
muxName: sessionName,
pid,
createdAt: Date.now(),
workingDir: process.cwd(),
mode: 'claude',
attached: false,
name: `Restored: ${sessionName}`,
};
this.sessions.set(sessionId, session);
discovered.push(sessionId);
console.log(`[TmuxManager] Discovered unknown tmux session: ${sessionName} (PID ${pid})`);
}
}
}
} catch (err) {
console.error('[TmuxManager] Failed to discover sessions:', err);
const fragment = sessionName.replace(/^(?:codeman|claudeman)-/, '');
const sessionId = `restored-${fragment}`;
const session: MuxSession = {
sessionId,
muxName: sessionName,
pid,
createdAt: Date.now(),
workingDir: process.cwd(),
mode: 'claude',
attached: false,
name: `Restored: ${sessionName}`,
};
this.sessions.set(sessionId, session);
discovered.push(sessionId);
console.log(`[TmuxManager] Discovered unknown tmux session: ${sessionName} (PID ${pid})`);
}
if (dead.length > 0 || discovered.length > 0) {

Some files were not shown because too many files have changed in this diff Show More