Compare commits

...
Author SHA1 Message Date
arkon d072e773d8 chore: version packages 2026-03-14 18:38:28 +01:00
arkonandClaude Opus 4.6 a649c91b68 fix: WS session lifecycle, reconnection, and CJK session-switch cleanup
- Close WebSocket when session exits (exit event listener) to prevent
  orphaned listeners and stale writes to dead PTY
- Add readyState guard in onTerminal to stop buffering after socket closes
- Simplify heartbeat: remove redundant alive flag, use pongTimeout only
- Add exponential backoff reconnection on unexpected WS close (skip for
  server rejections 4004/4008/4009)
- Clear CJK textarea on session switch to prevent wrong-session input

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:37:10 +01:00
arkonandClaude Opus 4.6 3383c23099 fix: address code review findings across WS, CJK input, install.sh, and README
WebSocket route: add socket error handler to prevent process crashes, enforce
per-session connection limit (max 5), track/decrement counts on close.

CJK input: add destroy() method with proper listener cleanup, guard against
double-init, add maxlength/aria-label to textarea, use language-neutral
placeholder, explicitly clear cjkActive on hide.

install.sh: fix update() to use $BRANCH and $REPO_URL instead of hardcoded
origin/master — fork users were silently switched back to master on update.

README: fix broken markdown table (paragraph concatenated into last cell),
add CODEMAN_NODE_VERSION to env var table.

Tests: add 8 new test cases for batch coalescing, flush threshold, unknown
message types, connection limit, heartbeat, readyState guards. Import
MAX_INPUT_LENGTH from config, add connectWs timeout, replace setTimeout
with vi.waitFor in cleanup test.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:27:58 +01:00
arkonandClaude Opus 4.6 c3e1e731ef fix: use generic placeholders in fork install README example
Replace hardcoded contributor fork URL with <user>/<branch> placeholders
so the documentation is useful for any contributor.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:11:43 +01:00
Ark0N 405b711c3a Merge pull request #41 from douchekr/feat/input-cjk-form
feat: add CJK IME input textarea and fork/branch install support
2026-03-14 18:10:09 +01:00
arkon 4295faefc9 chore: version packages 2026-03-14 18:04:51 +01:00
arkon 93e1ba5110 Merge remote-tracking branch 'origin/feat/ws-terminal-io-upstream' 2026-03-14 18:03:36 +01:00
Ark0N abbbf9e90a Merge pull request #43 from Ark0N/feat/ws-tests
test: add WebSocket terminal I/O route tests
2026-03-14 18:03:04 +01:00
Ark0N 3a41de7b57 Merge pull request #42 from Ark0N/feat/ws-heartbeat
feat: add ping/pong heartbeat to WebSocket connections
2026-03-14 18:03:02 +01:00
Ark0N 8267edc6fe Merge pull request #40 from Spirotot/feat/ws-terminal-io-upstream
feat: WebSocket terminal I/O with server-side DEC 2026 sync
2026-03-14 18:02:55 +01:00
arkonandClaude Opus 4.6 78c568e5f7 test: add automated tests for WebSocket terminal I/O route
16 tests covering session-not-found close code, terminal output with
DEC 2026 sync markers, client input forwarding, resize bounds
validation, malformed message handling, and connection cleanup of
session event listeners.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 18:00:55 +01:00
arkonandClaude Opus 4.6 cc624d2575 feat: add ping/pong heartbeat to WebSocket connections
Detect stale connections that TCP keepalive won't catch for minutes,
especially through tunnels and proxies. Pings every 30s with a 10s
pong timeout — if the client doesn't respond, the socket is terminated
and all timers cleaned up.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 17:58:50 +01:00
arkonandClaude Opus 4.6 5844720525 fix: validate WS resize dimensions to match HTTP route bounds
The HTTP resize route validates via ResizeSchema (cols: 1-500, rows:
1-200, integers only). The WS handler only checked typeof === 'number',
allowing floats, negatives, and extreme values through to ptyProcess.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-14 17:57:20 +01:00
jayparkandClaude Opus 4.6 393a2d9c28 fix: use BRANCH variable in install.sh no-changes update path
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:52:28 +09:00
jayparkandClaude Opus 4.6 809bf6a614 fix: use BRANCH variable in install.sh update path
The update path was hardcoded to origin/master. Now uses the
CODEMAN_BRANCH variable and updates the remote URL on upgrade.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:51:20 +09:00
jayparkandClaude Opus 4.6 da71d8d01c feat: support custom repo URL and branch in install.sh
Add CODEMAN_REPO_URL and CODEMAN_BRANCH env vars to install.sh
for installing from forks or feature branches. Update README with
fork installation instructions and env var reference table.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:50:06 +09:00
jayparkandClaude Opus 4.6 e5aca6aa4c feat: add CJK IME input textarea with env toggle
Add a dedicated textarea below the terminal for CJK (Korean/Japanese/Chinese)
IME input. xterm.js intercepts IME composition events, preventing composed
characters from displaying correctly. This textarea bypasses xterm entirely
by using native browser IME handling — text accumulates until Enter, then
sends to PTY in one shot.

- Always-visible textarea below terminal (inside .terminal-wrap flex column)
- focus/blur sets window.cjkActive flag to block xterm onData
- Enter sends textarea.value + \r to PTY, Escape clears
- Arrow keys, Ctrl+C/D/L/Z, Tab, Backspace pass through to PTY when empty
- attachCustomKeyEventHandler suppresses xterm key handling during composition
- INPUT_CJK_FORM=ON|OFF env var toggle (default: off, passed via SSE init)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 16:23:25 +09:00
Aaron FieldsandClaude Opus 4.6 ceaf4624a1 feat: add WebSocket terminal I/O with server-side DEC 2026 sync
Replace per-keystroke HTTP POST + SSE terminal output with a single
bidirectional WebSocket connection for dramatically lower input latency.
The existing SSE+POST paths remain fully functional as fallback.

Server-side: ws-routes.ts provides /ws/sessions/:id/terminal with 8ms
micro-batching and 16KB flush threshold. Each batch is wrapped in
DEC 2026 synchronized update markers so xterm.js renders atomically —
Ink's DA capability negotiation fails through the PTY→server→WS proxy
chain, so without server-injected markers, cursor-up redraws flicker.

Frontend: _connectWs/_disconnectWs manage per-session WS lifecycle.
Input and resize use WS fast path with HTTP POST fallback. SSE terminal
events are suppressed when WS is active to prevent double rendering.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 21:02:28 -04:00
arkonandClaude Opus 4.6 a6597e4a9a fix: patch 5 dependency vulnerabilities (basic-ftp, fastify, minimatch, serialize-javascript)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-13 00:26:56 +01:00
arkon f869e823af chore: version packages 2026-03-12 23:59:16 +01:00
arkonandClaude Opus 4.6 8d0b179f94 fix: repair 15 pre-existing subagent-watcher test failures
Root causes:
- Mock readline (EventEmitter) lacked .close() method, causing TypeError
  that blocked extractDescriptionFromFile's Promise from ever resolving
- Mock stream lacked .destroy() method (same issue after .close() fix)
- Entry-processing tests shared one readline mock between description
  extraction and tailing — events emitted before tailFile started were lost
- Liveness checker marked agents as 'completed' instead of 'idle' because
  fixed stat timestamps became stale after fake timer advancement

Fixes:
- Add createMockRl() helper with .close() method
- Use { destroy: vi.fn() } for stream mocks
- Use mockReturnValueOnce() for two-readline pattern in 7 entry tests
- Use mockImplementation() for dynamic stat timestamps

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 16:08:39 +01:00
arkonandClaude Opus 4.6 98fa55b7b2 chore: codebase cleanup — remove dead code, consolidate imports, extract constants
- Remove 3 unused exported constants (TRIM_MESSAGES_TO, MAX_TERMINAL_COLS, MAX_TERMINAL_ROWS)
- Consolidate 8 direct util imports into barrel imports (./utils/index.js)
- Extract magic number 8191 to FILE_PEEK_BYTES constant in buffer-limits.ts
- Add explanatory comments to 9 undocumented .catch(() => {}) handlers

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:50:40 +01:00
arkonandClaude Opus 4.6 c46ac30631 fix: hide subagent monitor panel by default
Change showSubagents default from true to false so the subagent
panel doesn't auto-show on page load. Users can still enable it
via Settings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:35:27 +01:00
arkonandClaude Opus 4.6 dfcc14bfd2 fix: one-liner restart command that works for background processes
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:34:42 +01:00
arkonandClaude Opus 4.6 a068008409 fix: clarify restart instructions — stop first, then start
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:33:34 +01:00
arkonandClaude Opus 4.6 0aa31f100e fix: show restart command when codeman-web is not a systemd service
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:31:19 +01:00
arkonandClaude Opus 4.6 314a160458 feat: auto-restart codeman-web service after update if running
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:23:55 +01:00
arkonandClaude Opus 4.6 e7ee5595c5 feat: auto-detect existing install and run update instead of fresh install
Re-running the install script now detects ~/.codeman/app/.git and
automatically updates instead of re-installing. Removes the separate
`bash -s update` instructions from README since it's no longer needed.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-12 15:09:37 +01:00
arkon 625d4976d3 chore: version packages 2026-03-11 19:37:17 +01:00
arkonandClaude Opus 4.6 d02cddece6 fix: correct claudeSessionId for resumed sessions and clean up DEC sync dead code
Use resumeSessionId for Claude conversation ID when resuming sessions,
increase default font size to 14, extract shared history fetch logic,
and remove unused DEC 2026 sync constants/functions (xterm.js 6.0 handles natively).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-11 19:36:58 +01:00
Ark0N abbc4b13fd Merge pull request #39 from sunnyzhouy/master
feat: session resume, xterm.js 6.0 upgrade, and resize fix
2026-03-11 19:20:28 +01:00
arkonandClaude Opus 4.6 a14e47e19c chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-11 19:13:49 +01:00
sunnyzhouy 754a966b53 Merge branch 'Ark0N:master' into master 2026-03-12 01:06:14 +08:00
zhouyuan 28dfc279d4 fix: resolve terminal resize scrollback ghost renders
- Switch resize handler to 300ms trailing-edge debounce for single reflow
- Add \x1b[3J (Erase Saved Lines) to clear scrollback reflow debris
- Remove client-side cursor-up flicker filter and DEC 2026 marker
  stripping — xterm.js 6.0 handles synchronized output natively
- Remove server-side DEC 2026 wrapping to prevent premature sync exit
  from non-reference-counted nested markers
2026-03-12 01:04:21 +08:00
arkonandClaude Opus 4.6 da85e9738b fix: iPad tablet toolbar styling and PR #34 refinements
- Scope toolbar bottom-offset to phone breakpoint only (position:fixed);
  prevents double-correction on iPad where toolbar is position:relative
- Extract keyboard accessory bar styles to top-level mobile.css so
  /init, /clear, /compact buttons render correctly on iPad
- Use desktop-style toolbar sizing on tablet (430-768px): smaller font,
  no forced min-height, proper gap between buttons
- Show voice/mic button on tablet (was hidden at <1023px with no
  mobile replacement above 430px)
- Bump CSS cache-bust version to 0.1633

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-11 12:59:47 +01:00
Ark0N 8b8907c4ec Merge pull request #34 from arnlaugsson/fix/ipad-safari-toolbar-viewport
fix: toolbar off-screen on iPad Safari with tabs
2026-03-11 01:26:55 +01:00
zhouyuan 2329dab240 feat: upgrade xterm.js 5.3 to 6.0 for native DEC 2026 synchronized output
xterm.js 6.0.0 natively handles DEC mode 2026 (synchronized output),
which renders Ink's cursor-up redraws atomically at the parser level.
This eliminates split-frame rendering that caused table header loss
and garbled overlapping text in Claude CLI sessions.

- Migrate from xterm/xterm-addon-* to @xterm/* scoped packages
- Update build.mjs and postinstall.js vendor paths
- Remove old xterm 5.x dependencies
2026-03-10 14:59:26 +08:00
zhouyuan 31ce7405a6 perf: increase terminal scrollback from 5000 to 20000 lines
Long-running Claude sessions can exceed 5000 lines easily, causing
earlier content to be lost. 20000 lines retains ~4x more history.
2026-03-10 14:36:10 +08:00
zhouyuan 06f7d40c42 feat: reduce default font size and persist tabs across refresh
- Default terminal font 14px → 12px, min 10px → 8px
- Save tab metadata to localStorage on every render
- Restore ended sessions as dimmed tabs after page refresh
- Ended tabs show "Session ended" message on click
2026-03-10 14:30:21 +08:00
zhouyuan 05eba70598 feat: improve session resume reliability and persist user settings
- Filter empty sessions from history API (check for conversation content)
- Add --resume fallback to new session if resume fails (prevents dead panes)
- Pass resumeSessionId through respawnPane for dead pane recovery
- Persist respawn presets and runMode to server settings (cross-device sync)
- Fix mobile touch handling for Recent Sessions dropdown (DOM API + touch CSS)
2026-03-10 14:03:42 +08:00
zhouyuan d27974ff6e chore: update package-lock.json 2026-03-10 02:23:15 +08:00
zhouyuan 3cca5380ba fix: route shell sessions to correct endpoint on tab click
selectSession() was always calling /interactive for restored idle
sessions regardless of mode. Shell sessions now correctly call /shell.
Also add loadHistorySessions and resumeHistorySession to frontend.
2026-03-10 02:23:10 +08:00
zhouyuan 6d7efc13e6 feat: add history session resume UI and API
Add GET /api/history/sessions endpoint that scans Claude conversation
files for resume. Add welcome overlay UI with clickable history items.
Path decoding validates existence via fs.access with HOME fallback.
2026-03-10 02:23:02 +08:00
zhouyuan 63f86807ad feat: add resumeSessionId support for conversation resume after reboot
Add resumeSessionId field throughout the session creation pipeline,
allowing sessions to resume previous Claude conversations via --resume
flag instead of --session-id.
2026-03-10 02:22:53 +08:00
Skúli Arnlaugsson 2e4e646c06 fix: toolbar pushed off-screen on iPad Safari with tabs
On iPad Safari with the tab bar visible, `100vh` extends behind the
browser chrome, pushing the fixed-position toolbar out of view.

- Add `viewport-fit=cover` to viewport meta tag
- Use `100dvh` with `100vh` fallback for body/.app height
- Set `--app-height` CSS variable from `visualViewport.height` via JS
- Offset fixed toolbar on iOS Safari using the layout/visual viewport delta
2026-03-08 23:54:15 +00:00
arkon 507423b776 chore: version packages 2026-03-08 16:06:51 +01:00
arkonandClaude Opus 4.6 67d0b0b538 feat: add tunnel status indicator with control panel in header
Green pulsing dot in the desktop header shows when Cloudflare tunnel is active.
Clicking opens a dropdown panel with tunnel URL, remote client count, auth
sessions, and start/stop/QR/revoke controls. Detects tunnel clients via
Cf-Connecting-Ip header to exclude local connections from the count.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-08 16:06:15 +01:00
Ark0N 26cfd8b7ef Merge pull request #33 from arnlaugsson/fix/macos-install-platform-deps
fix: move Linux-only native deps to optionalDependencies
2026-03-08 15:55:16 +01:00
Skúli ArnlaugssonandClaude Opus 4.6 208e6bc175 fix: move Linux-only native deps to optionalDependencies
`@remotion/compositor-linux-x64-gnu` and `@rspack/binding-linux-x64-gnu`
are Linux x64 binaries that cause npm install to fail on macOS (arm64)
with EBADPLATFORM. Moving them to optionalDependencies allows npm to
skip them gracefully on unsupported platforms.

Fixes #32

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 20:59:57 +00:00
arkonandClaude Opus 4.6 6d52b16edc docs: add zerolag demo video to README
Side-by-side comparison of local echo (0ms) vs server echo (600ms-2.7s)
rendered from Remotion ZerolagDemo composition.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 07:43:52 +01:00
arkonandClaude Opus 4.6 4988e85901 docs: add Operation Lightspeed to v0.3.7 changelog
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 07:37:22 +01:00
arkonandClaude Opus 4.6 e799c83b39 chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 07:34:24 +01:00
arkonandClaude Opus 4.6 0717cfbfec refactor: codebase cleanup — dead code, regex helper, centralized constants, tests
- Remove unused validateTokenCounts/validateTokensAndCost exports and PlanPhase type alias
- Add execPattern() helper to eliminate 8 repetitive .lastIndex=0 + exec() loops
- Centralize 11 magic number constants into config/ai-defaults.ts and config/server-timing.ts
- Remove stale src/tui from tsconfig.json exclude
- Fix CLAUDE.md inaccuracies (session helpers, app.js line count, module count)
- Add 316 new tests: LRUMap (38), StaleExpirationMap (42), BufferAccumulator (33),
  respawn helpers (142), system-routes expansion (11→61)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 07:33:44 +01:00
arkonandClaude Opus 4.6 b7b2555dc0 test: add Operation Lightspeed tests — tab switching, local echo, SSE filters
14 new tests covering tab switch SSE reconnect, terminal buffer edge cases
(tail=0, negative, non-numeric, huge values), extractSessionId filtering,
session lifecycle churn, and concurrent SSE client limits.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 07:04:27 +01:00
arkonandClaude Opus 4.6 a609c435fa fix: use TERMINAL_TAIL_SIZE constant and add client-drop recovery
Two hardcoded `256 * 1024` tail sizes in app.js bypassed the
TERMINAL_TAIL_SIZE constant (128KB), causing stale cached browsers
to fetch 256KB buffers even after the constant was reduced to prevent
WebGL GPU stalls. Also adds a self-recovery timer that reloads the
terminal buffer after client-side data drops, preventing permanent
display corruption.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-07 07:00:18 +01:00
arkonandClaude Opus 4.6 3268a12e5e fix: Operation Lightspeed review fixes — padding, dead code, tablet WebGL
- Add SSE padding to backpressure drain write for tunnel clients
- Remove dead SessionTerminal from broadcast padding check
- Trim whitespace in SSE session filter query params
- Remove unused _bufferLazyTerminalData scaffolding code
- Skip WebGL on tablets too, not just phones (canvas fallback)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 06:40:09 +01:00
arkonandClaude Opus 4.6 1f30ed445c perf: Operation Lightspeed — 5 parallel performance optimizations
1. Session-scoped SSE subscriptions: server filters events by session ID,
   clients can subscribe via ?sessions=id1,id2 (backwards-compatible)
2. Lazy xterm.js for subagent windows: terminals created on restore,
   disposed on minimize — saves ~3.75MB DOM at 50 agents
3. Targeted badge updates: badge count changes update the <span> directly
   instead of rebuilding the entire session tab sidebar (O(1) vs O(n))
4. Conditional SSE padding: 8KB Cloudflare padding only on session:terminal
   and session:needsRefresh, not every event (~70% bandwidth reduction)
5. Canvas renderer on mobile: skip WebGL addon on mobile devices to reduce
   GPU pressure and prevent context loss on weaker mobile GPUs

All 5 implemented in parallel via isolated git worktrees, merged conflict-free.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 06:14:16 +01:00
arkonandClaude Opus 4.6 415f02e680 fix: multi-layer backpressure to prevent terminal write freezes
Add three layers of protection against oversized terminal.write() calls
that freeze Chrome's main thread:

1. SSE entry cap: _onSessionTerminal drops data when total queued bytes
   (pendingWrites + flickerFilterBuffer) exceeds 128KB. Server sends
   session:needsRefresh to recover dropped content.

2. Flush cap: flushPendingWrites splits at DEC 2026 sync segment
   boundaries with 64KB per-frame budget. Excess segments deferred to
   next requestAnimationFrame. Segment-level splitting preserves Ink
   redraw atomicity (no flicker).

3. Reduced tail size: initial buffer fetch reduced to 128KB (from 256KB)
   to limit data volume during tab switch.

Re-enable WebGL renderer — root cause was unbounded terminal.write()
volume, not the GPU renderer itself.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-06 17:10:25 +01:00
arkon d5814947d6 chore: version packages 2026-03-05 23:09:24 +01:00
arkonandClaude Opus 4.6 b620511d0e fix: cap terminal writes at 48KB/frame to prevent page unresponsive freezes
Root cause was NOT WebGL — breadcrumbs showed 141KB single-frame
terminal.write() calls freezing Chrome for 2+ minutes even with the
canvas renderer. During heavy Ink output, multiple SSE terminal events
accumulate between animation frames and flush all at once.

Fix: split flushPendingWrites at DEC 2026 sync segment boundaries with
a 48KB per-frame budget. Each segment is a complete Ink redraw, so
splitting between them preserves atomicity (no flicker). Excess
segments are deferred to the next requestAnimationFrame.

Also re-enable WebGL since it was not the cause — the flush cap
protects both renderers equally.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 18:36:27 +01:00
arkon 525f02f502 chore: version packages 2026-03-05 17:59:33 +01:00
arkonandClaude Opus 4.6 e07c59477d fix: disable WebGL renderer to prevent Chrome page unresponsive crashes
Root cause: xterm.js WebGL addon performs synchronous GPU ReadPixels
calls during terminal.write(). When heavy terminal output floods in
(~1MB/4s from active Claude sessions), single-frame writes of 70-105KB
block Chrome's main thread for 10+ seconds, triggering "page
unresponsive" dialogs. This happens both during tab switches (buffer
load + live SSE data competing) and during normal use (Ink redraw
bursts).

Fix: disable WebGL by default, use canvas renderer instead. Canvas
handles the same workloads without GPU stalls. Re-enable with ?webgl
URL param for testing.

Also: gate live SSE terminal writes during the entire selectSession()
buffer load sequence (not just during chunkedTerminalWrite), preventing
live data from competing with historical buffer restoration.

Crash investigation data (from server-side breadcrumb collection):
- 27 flushes totaling 937KB in 4 seconds preceded crash
- Two 105KB and one 104KB single-frame flushes observed
- Crash occurred when user switched tabs during output flood
- Tab froze for 2m54s before recovering
- Backend always stable (0 crashes); pure frontend GPU issue

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 17:52:28 +01:00
arkonandClaude Opus 4.6 2b8a522cbd fix: add server-side crash breadcrumbs and remove --drop:console from build
The --drop:console esbuild flag was silently stripping all diagnostic
console.* calls from production builds, making crash investigation
impossible. Remove it temporarily for debugging.

Add server-side crash breadcrumb collection:
- Frontend writes breadcrumbs to localStorage AND POSTs to /api/crash-diag
- Server stores latest breadcrumbs in memory, readable via GET /api/crash-diag
- Granular breadcrumbs in selectSession: CACHE_WRITE, FETCH_START,
  FETCH_DONE, REWRITE, FOCUS, SELECT_DONE
- 2s heartbeat beacon so breadcrumbs survive tab freezes
- text/plain content-type parser for navigator.sendBeacon compatibility

Initial findings from breadcrumbs:
- Crash happens during selectSession() for sessions with large buffers
- Pattern: cached buffer exists from prior visit + live SSE data arriving
- Backend is always stable (0 crashes); this is a pure frontend issue

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 16:44:17 +01:00
arkonandClaude Opus 4.6 4abe055182 feat: add frontend crash diagnostics for page unresponsive investigation
Add global error handlers, long task detection, WebGL context tracking,
and performance timing to identify root cause of intermittent Chrome
"page unresponsive" freezes during session switching and typing.

Diagnostics:
- window error/unhandledrejection handlers ([CRASH-DIAG] prefix)
- PerformanceObserver for long tasks (>200ms main thread blocks)
- WebGL context loss/restore tracking on all canvases
- ?nowebgl URL param to disable WebGL renderer for testing
- Timing on flushPendingWrites, chunkedTerminalWrite, selectSession

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-05 12:42:32 +01:00
arkon 15a3b3996b chore: version packages 2026-03-05 00:28:13 +01:00
arkonandSigurður Guðbrandsson 1b76e6e2e2 fix: prevent Chrome freeze and shell feedback delay from flicker filter
Bug 1: Every incoming SSE terminal event reset the 50ms flush timer, not
just cursor-up events. During active Claude runs the timer never fired,
accumulating MBs in flickerFilterBuffer that froze Chrome on flush.
Fix: only reset timer on cursor-up events; add 256KB safety valve.

Bug 2: Shell sessions emit cursor-up on every keystroke for readline
prompt redraws, triggering the flicker filter and delaying feedback.
Fix: skip cursor-up filter for shell mode; disable local echo overlay.

Based on PR #31 by @SGudbrandsson.

Co-Authored-By: Sigurður Guðbrandsson <SGudbrandsson@users.noreply.github.com>
2026-03-05 00:27:07 +01:00
arkonandClaude Opus 4.6 f8b81b8478 fix: eliminate WebGL re-render flicker during tab switch
Stop toggling WebGL renderer off/on around large buffer writes in
chunkedTerminalWrite(). The dispose+loadAddon cycle caused visible
re-render flashes and (before the deferred fix) synchronous GPU stalls
from ReadPixels blocking the main thread. Instead, keep WebGL active
and rely on 32KB chunked writes to keep per-frame render work under
~5ms.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-04 16:01:16 +01:00
arkonandClaude Opus 4.6 d79e25d5d8 test: comprehensive CJK wide character test plan for PR #30
Covers all 7 test plan items: Chinese overlap prevention, double-width
spacing, Japanese/Korean input, cursor positioning, line wrapping at
column boundaries, ASCII regression, and teammate terminal panels.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 13:29:23 +01:00
arkonandClaude Opus 4.6 b1df5319d5 fix: prevent page unresponsive crashes from WebGL GPU stalls during session switch
Disable WebGL renderer during large buffer loads (>32KB) and fall back to
canvas, which handles bulk ANSI writes without synchronous GPU ReadPixels
calls. Re-enable WebGL after the buffer load completes so live terminal
streaming still benefits from GPU acceleration.

Also reduce chunk size from 128KB to 32KB and use chunked writes for
cached buffer restores instead of synchronous terminal.write().

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-04 13:18:02 +01:00
arkonandClaude Opus 4.6 fc92d8a8a4 fix: restore CJK wide character support in ZerolagInput overlay (#30)
Add charCellWidth/stringCellWidth helpers for Unicode-aware width detection,
fix makeLine to use for...of iteration with visual column positioning, and
fix line splitting in _render to use visual column widths instead of string
length. CJK/fullwidth characters now correctly occupy 2 cell widths.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 13:12:30 +01:00
arkonandClaude Opus 4.6 c7cd4f9e17 fix: prevent page unresponsive crashes from WebGL GPU stalls during session switch
Reduce terminal chunk size from 128KB to 32KB and use double-RAF between
chunks to give the WebGL renderer time to flush GPU operations. Also switch
cached buffer restore from direct terminal.write() to chunked writes —
the synchronous 256KB write was blocking the main thread for 5+ seconds
via synchronous ReadPixels calls.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-04 13:09:10 +01:00
arkonandClaude Opus 4.6 9433de75b8 chore: version packages
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-04 12:17:38 +01:00
arkon 49e9b9e8c9 chore: version packages 2026-03-03 22:56:40 +01:00
Ark0NandClaude Opus 4.6 2ee9ad72e8 refactor: SSE event handlers, LLM context optimization, @fileoverview docs (#29)
* refactor: extract SSE event handlers into named class methods

Replace ~80 inline addListener closures in connectSSE() with a
declarative _SSE_HANDLER_MAP array that drives registration in a
single loop. Each handler is now a named _on* method on CodemanApp,
making them individually addressable for LLM navigation.

Add SSE_EVENTS constant object in constants.js to eliminate magic
event-type strings scattered across the frontend.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: fix inaccuracies in CLAUDE.md

- Fix types barrel path: src/types.ts → src/types/index.ts
- Update app.js line count: ~12K → ~11.5K
- Correct route handler counts (113 → 111, per-group fixes)
- Add code style, ESM gotcha, env vars, route test, lifecycle log docs
- Add Node 22 CI note, test teardown timeout, port range

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: add mobile screenshots and QR auth security writeup to README

Add 3 mobile screenshots (landing, idle, active) and expand the
mobile section with QR auth security design details and a
touch-optimized interface subsection.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: bundle xterm-zerolag-input as vendor IIFE and add pre-commit hook

Build and postinstall now bundle the local xterm-zerolag-input package
as an IIFE at vendor/xterm-zerolag-input.js with global LocalEchoOverlay
shim. Add git pre-commit hook that runs prettier --check on staged .ts
files to catch format issues before CI.

Also bump constants.js and app.js cache-bust versions to 0.3.0 and add
tunnel upload URL display row in settings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add cloudflared install support and interactive launch menu

- Add optional cloudflared dependency detection and installation
  across 6 distro families (macOS, Debian, Fedora, Arch, Alpine, SUSE)
- Add tunnel systemd service setup helper
- Replace post-install instructions with interactive launch menu
  (run now / systemd service / skip)
- Uninstall now cleans up both codeman-web and codeman-tunnel services

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore: gitignore readme-preview.mjs

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: WIP — SSE event constants, @fileoverview docs, CLAUDE.md compression

- Migrate broadcast() string literals → SseEvent.* typed constants
- Add @fileoverview with cross-domain references to all 13 type domain files
- Add @fileoverview to frontend JS modules (constants, mobile, voice, etc.)
- Add section dividers to route files for LLM scanability
- Compress CLAUDE.md: flat file list → domain table, fix counts

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: optimize codebase for LLM context window efficiency

CLAUDE.md: 456 → 309 lines (32% reduction)
- Merge Commands into compact table, remove redundant bash block
- Convert Security section to dense table format
- Merge Performance + Resource Limits, Debugging + Troubleshooting
- Compress Tunnel, Memory Leak, Scripts, Screenshots sections
- Remove Key Patterns that duplicate @fileoverview in source files

Backend @fileoverview enhancements (10 priority files):
- session.ts: key methods, events, cross-domain refs
- respawn-controller.ts: state machine, idle detection layers
- ralph-tracker.ts: exports, circuit breaker, events
- ralph-loop.ts: lifecycle, persistence, events
- subagent-watcher.ts: watched patterns, teammate detection
- server.ts: coordination list, port interfaces
- state-store.ts: dual-file persistence, migration
- session-manager.ts: lifecycle methods, mutex guard
- hooks-config.ts: hook events list, categories
- sse-events.ts: category breakdown (~90 events, 17 categories)

Frontend app.js: add 6 section dividers, update @fileoverview line refs

Fix: escape glob `*/` in JSDoc that broke ESLint parser

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address PR #29 review bugs

- server.ts: replace hardcoded 'session:needsRefresh' with SseEvent constant
- install.sh: fix Alpine cloudflared install for non-root (download to tmpfile first)
- install.sh: replace Arch pacman (AUR-only) with direct binary download
- index.html: bump all 8 remaining cache-bust versions from v0.2.9 to v0.3.0
- mobile-handlers.js: fix @dependency annotation (keyboard-accessory.js, not constants.js)
- types/push.ts: fix layer number (4, not 5)
- subagent-watcher.ts: fix watched pattern path to include {session} segment
- constants.js: fix SSE_EVENTS count in @fileoverview (~73, not ~65)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 22:40:18 +01:00
arkonandClaude Opus 4.6 14462f7bfe fix: remove hard-coded rollup-linux-x64-gnu dependency
The @rollup/rollup-linux-x64-gnu native binding was a hard dependency,
causing npm install to fail on arm64 and macOS platforms. Nothing in the
codebase uses rollup directly (build uses esbuild); rollup is only
pulled in transitively by @remotion/cli and manages its own
platform-specific bindings via its own optionalDependencies.

Closes #28

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-03 14:07:18 +01:00
arkonandClaude Opus 4.6 efa6487361 fix: eliminate SSE padding overhead and debounce subagent window renders
Two performance fixes for browser hanging during active agent work:

1. SSE padding (8KB per event) now only applied when Cloudflare tunnel is
   active — direct/Tailscale connections skip the padding entirely. Previously
   every broadcast event got 8KB of comment padding even on local connections,
   causing 40-160KB/s of wasted bandwidth during active subagent work.

2. Subagent window content renders (tool_call, progress, message, tool_result)
   now debounced at 100ms per agent via scheduleSubagentWindowRender(). Previously
   each SSE event triggered an immediate DOM rewrite, causing 10-30+ rewrites/sec
   that starved the terminal rendering pipeline.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-02 21:48:31 +01:00
arkonandClaude Opus 4.6 f7552c7cb5 style: fix prettier formatting in server.ts
Expand one-liner try/catch to multi-line to satisfy prettier check.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 23:48:14 +01:00
arkonandClaude Opus 4.6 2b61a8db1b fix: add SSE padding to flush Cloudflare tunnel buffers for real-time events
Cloudflare quick tunnels buffer small SSE responses, causing tab creation
and other UI events to arrive late on mobile. Adds ~8KB SSE comment padding
(ignored by EventSource) to force the proxy to flush immediately.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 18:24:14 +01:00
arkonandClaude Opus 4.6 8764a684ac fix: replace require('qrcode') with await import() for ESM compatibility
require() is not defined in ESM modules. The production build (tsc)
outputs ESM, causing 'require is not defined' → 500 → 'QR unavailable'.
Vitest/tsx shimmed require() so tests passed but production was broken.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 18:04:23 +01:00
arkonandClaude Opus 4.6 211abe60ca test: add 8 integration tests for GET /api/tunnel/qr SVG endpoint
The QR SVG endpoint had zero test coverage for success paths — only
the 404 (tunnel not running) case was tested. This adds tests for
auth/no-auth SVG generation, the 500 error when token rotation isn't
started, SVG caching consistency, and cache invalidation on regeneration.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 17:57:46 +01:00
arkonandClaude Opus 4.6 888a2d8e9a chore: move manual test scripts to test/manual/
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 17:38:45 +01:00
arkon 87abc0d301 fix: release tag rename — use databaseId and rename before tag deletion 2026-03-01 17:36:54 +01:00
arkon b79ad497f4 chore: version packages 2026-03-01 17:34:09 +01:00
arkonandClaude Opus 4.6 18217268bf fix: extract QR auth magic numbers into named constants, add 16 security tests
Replace hardcoded per-IP rate limit (10) and cookie maxAge (86400) in
system-routes.ts with QR_AUTH_FAILURE_MAX and AUTH_SESSION_TTL_MS/1000
so both auth paths stay in sync if constants change.

Add 16 new tests: grace period boundary precision, base62 charset
validation, current+previous token during grace, stopTokenRotation
state cleanup, rate limit reset, consumed token eviction, full
end-to-end QR flow, per-IP 429, cookie attributes, concurrent race,
regenerate invalidation, URL encoding, path traversal, /q without
param, and session record method:qr.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 17:32:02 +01:00
arkonandClaude Opus 4.6 d094d9ec50 fix: code cleanup — path traversal, test leaks, dead code, consistency
- Add path traversal protection to GET /api/cases/:name and fix-plan
- Use safePathSchema for LinkCaseSchema.path
- Fix QR auth test timer leak (afterAll → afterEach) and env var try/finally
- Remove dead terminal size check after Zod validation in resize route
- Remove no-op sampleCount guard in adaptive timing
- Replace hardcoded values with constants in notification-manager and subagent-windows
- Add Zod validation to POST /api/auth/revoke
- Use _apiPut instead of raw fetch in subagent-windows
- Add SwipeHandler.cleanup() for consistency with other mobile handlers
- Move NiceConfig/ProcessStats from types/plan.ts to types/common.ts

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 17:26:31 +01:00
arkonandClaude Opus 4.6 f94ab08cbe fix: stale types test and unreachable ralph full-reset API path
- Remove tests for createSuccessResponse and ErrorMessages which were
  removed/made private during the type system refactoring (15 failures)
- Fix RalphConfigSchema to accept 'full' string for reset field,
  matching the route handler's fullReset() code path

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 17:17:05 +01:00
arkonandClaude Opus 4.6 bdaa43bc5c fix: formatting and async test bug in hooks-config
- Fix Prettier formatting in ralph-tracker.ts and respawn-controller.ts
  (whitespace drift from Phase 2/4 refactoring)
- Add missing `await` to writeHooksConfig() calls in hooks-config.test.ts
  (async function was called without await, causing ENOENT race condition)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 17:08:41 +01:00
arkonandClaude Opus 4.6 8e5f95d6ea test: add unit tests for Debouncer, CleanupManager, and timer migration refactors
New test files for utilities that previously had no dedicated coverage,
plus migration tests validating the ralph-tracker and respawn-controller
timer refactorings work correctly.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 16:50:23 +01:00
arkonandClaude Opus 4.6 71276c5d6e refactor: complete Phase 2 — migrate ralph-tracker and respawn-controller to managed timers
ralph-tracker.ts: Replace 4 manual timer/flag fields (_todoUpdateTimer, _loopUpdateTimer,
_todoUpdatePending, _loopUpdatePending) with 2 Debouncer instances.

respawn-controller.ts: Replace 10 manual timer fields (stepTimer, completionConfirmTimer,
noOutputTimer, detectionUpdateTimer, autoAcceptTimer, preFilterTimer, hookConfirmTimer,
clearFallbackTimer, stepConfirmTimer, stuckStateTimer) with CleanupManager + timerIds Map.
startTrackedTimer/cancelTrackedTimer preserved as wrappers for UI countdown display.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 16:43:45 +01:00
arkonandClaude Opus 4.6 a8a6c1a648 test: add comprehensive overlay tests for visual fixes and setPrompt
New cell-dimensions.test.ts (8 tests): DPR conversion, charTop/charHeight,
null cases. Overlay renderer tests (+12): cellH+1 seam fix, span centering,
no-transform, ligature disabling, multi-line cursor. Addon tests (+9):
setPrompt() strategy switching, tab-switch cycle, ghost artifact prevention.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 16:40:26 +01:00
arkonandClaude Opus 4.6 e8fd358923 fix: overlay rendering — vertical alignment, line artifact, and duplicate class members
Overlay renderer (xterm-zerolag-input):
- Add charTop/charHeight to CellDimensions and RenderParams for precise
  vertical text positioning matching xterm's canvas renderer
- Convert device.char dimensions to CSS pixels via devicePixelRatio
- Extend line div background 1px past cell boundary to cover compositing
  seam between overlay layer (z-index:7) and canvas layer below
- Remove -webkit-font-smoothing/text-rendering overrides that made overlay
  text thinner than canvas text
- Add per-span height/lineHeight for natural CSS vertical centering
- Add setPrompt() method for runtime prompt strategy switching (fixes tab
  switching crash with "setPrompt is not a function")

app.js duplicate class members:
- Remove dead formatTokens duplicate (line ~5590 shadowed precise version)
- Remove fire-and-forget resetCircuitBreaker duplicate (shadowed notification version)
- Rename mux-panel killAllSessions to killAllMuxSessions (was shadowing
  Codeman session killer, breaking Ctrl+K)

Other:
- Update index.html onclick to use killAllMuxSessions
- Add getTeamTasks mock to test route context

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 16:27:40 +01:00
arkonandClaude Opus 4.6 3915f4ea9d docs: add detailed QR code authentication section to README
Document the ephemeral single-use QR token system — how it works,
security design informed by USENIX Security 2025 research (6 flaws
addressed), timing-safe lookup, dual-layer rate limiting, QR version
optimization, desktop experience, threat coverage, and comparison
with Discord/WhatsApp/Signal QR auth models.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 16:27:05 +01:00
arkonandClaude Opus 4.6 090fc71fd0 docs: fix inaccurate Phase 2 and Phase 7 status in code-structure-findings.md
Phase 2: respawn-controller.ts and ralph-tracker.ts were never migrated
to CleanupManager/Debouncer despite being marked complete. Corrected
status to reflect actual codebase state (6 of 8 files migrated).

Phase 7: All 12 route test files now exist, updated from "9 remaining".

Scorecard adjusted: Resource Cleanup 9→8/10, Test Coverage 7→7.5/10.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 15:36:15 +01:00
arkonandClaude Opus 4.6 570daf2ee3 fix: show clear error when microphone used without HTTPS
navigator.mediaDevices is undefined in insecure contexts, causing
"undefined is not an object" error. Now shows actionable message.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 14:57:05 +01:00
arkonandClaude Opus 4.6 df551b5356 docs: mark all 7 phases complete in code-structure-findings.md
Update scorecard with before/after scores, mark phases 3-7 as
complete with verified line counts and file inventories. Fix 3
CLAUDE.md inaccuracies: static cache 1h→1y, route count ~160→~113,
remove stale src/tui reference.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 14:50:17 +01:00
arkonandClaude Opus 4.6 442cc19c3c test: add shared mock infrastructure and route tests (phase 7)
Consolidate duplicated MockSession/MockStateStore into test/mocks/,
migrate respawn tests to shared mocks, and add 58 route tests for
session, system, and respawn endpoints using Fastify app.inject().

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 14:33:18 +01:00
arkonandClaude Opus 4.6 97ca13916a refactor: consolidate ~70 scattered constants into 6 new config files (phase 6)
Create 6 domain-focused config files in src/config/ to centralize operational
tuning knobs that were scattered across 15+ source files:

- server-timing.ts: SSE batching, state debounce, scheduled runs, error recovery
- auth-config.ts: session TTL, rate limits, hook timeout
- tunnel-config.ts: QR token rotation, tunnel process lifecycle
- terminal-limits.ts: input length, terminal cols/rows, session name
- ai-defaults.ts: AI model string, idle/plan context limits
- team-config.ts: poll interval, cache sizes

Eliminates all cross-file duplicates:
- AI model string: 5 occurrences → 1 (config only)
- STATS_COLLECTION_INTERVAL_MS: 2 → 1
- timeout: 10000 in hooks-config: 6 → 0 (now HOOK_TIMEOUT_MS)
- MAX_TRACKED_AGENTS shadow: removed, imports from map-limits

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 13:19:16 +01:00
arkonandClaude Opus 4.6 ed398294c6 docs: fix inaccurate counts in CLAUDE.md
Route modules (13→12), type domain files (14→13), SSE events (~80→~100),
total route handlers (~110→~160), and all per-group API route counts.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 13:09:02 +01:00
arkonandClaude Opus 4.6 746004c461 feat: implement QR code authentication for tunnel access
Adds ephemeral single-use QR tokens for passwordless tunnel login.
Scanning the QR auto-authenticates; bare tunnel URL requires Basic Auth.

Backend:
- TunnelManager: 60s token rotation, 90s grace, rejection-sampled 6-char
  base62 short codes, Map-based O(1) lookup, SVG caching, global rate limit
- Auth middleware: /q/ bypass, separate qrAuthFailures counter, enhanced
  AuthSessionRecord with device context (ip, ua, createdAt, method)
- Routes: GET /q/:code (consume + cookie + redirect), POST /api/tunnel/qr/
  regenerate, POST /api/auth/revoke, updated GET /api/tunnel/qr with cache
- SSE: tunnel:qrRotated, tunnel:qrRegenerated, tunnel:qrAuthUsed events
- Audit: qr_auth lifecycle log entries

Frontend:
- Auto-refresh QR via inline SVG in SSE (fallback fetch if absent)
- 60s countdown indicator on QR badge
- Regenerate QR button
- QRLjacking detection toast with [Revoke All] action button (10s duration)
- showToast enhanced with optional duration and action button support

Fixes:
- /api/logout now invalidates server-side session token (was only clearing
  browser cookie, leaving token valid for replay)

Tests: 20 new tests in test/qr-auth.test.ts covering token lifecycle,
bias check, rate limiting, SVG caching, and full server integration.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 06:05:26 +01:00
arkonandClaude Opus 4.6 295190cc72 docs: update CLAUDE.md with new files from recent refactors
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 05:04:14 +01:00
arkonandClaude Opus 4.6 f44f0a912f refactor: split app.js into focused frontend modules (phase 5)
Extract 8 modules from the 15K-line app.js monolith:
- constants.js: shared constants, timing, escapeHtml(), extractSyncSegments()
- mobile-handlers.js: MobileDetection, KeyboardHandler, SwipeHandler
- voice-input.js: DeepgramProvider, VoiceInput
- notification-manager.js: NotificationManager class
- keyboard-accessory.js: KeyboardAccessoryBar, FocusTrap
- api-client.js: _api(), _apiJson(), _apiPost(), _apiDelete() helpers
- subagent-windows.js: 13 subagent window methods (open, close, drag, lines)
- vendor/xterm-zerolag-input.js: IIFE build replacing inlined copy

app.js reduced from ~15,200 to ~11,500 lines. All modules use global
scope with <script defer> ordering. Prototype extensions use
Object.assign(CodemanApp.prototype, {...}) pattern.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 04:57:37 +01:00
arkonandClaude Opus 4.6 e7a9cbe442 refactor: split god files into focused modules (phase 4, steps 2-4)
Split 3 large files into 11 focused sub-modules via composition:

ralph-tracker.ts (3,868 → ~2,400 LOC):
- ralph-plan-tracker.ts: plan task tracking, checkpoints, history
- ralph-fix-plan-watcher.ts: @fix_plan.md file watching
- ralph-stall-detector.ts: iteration stall detection
- ralph-status-parser.ts: RALPH_STATUS block parsing, circuit breaker

respawn-controller.ts (3,611 → ~3,200 LOC):
- respawn-patterns.ts: pure pattern detection functions
- respawn-adaptive-timing.ts: adaptive timing with percentile calc
- respawn-metrics.ts: cycle metrics tracking + aggregation
- respawn-health.ts: pure health scoring functions

session.ts (2,418 → ~1,800 LOC):
- session-cli-builder.ts: CLI argument construction
- session-auto-ops.ts: auto-compact/clear automation
- session-task-cache.ts: task description LRU cache

All external APIs preserved via delegation. Events forwarded
through parent classes. Zero behavioral changes — all 436 tests pass.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 04:04:35 +01:00
arkonandClaude Opus 4.6 8d2d51e8f0 refactor: split types.ts into domain modules (phase 4, step 1)
Split 1,443-line types.ts into 14 focused domain files under src/types/:
common.ts, session.ts, task.ts, app-state.ts, respawn.ts, ralph.ts,
api.ts, lifecycle.ts, run-summary.ts, tools.ts, teams.ts, push.ts,
plan.ts, and index.ts barrel.

Moved PlanItem interface from plan-orchestrator.ts into types/plan.ts
to break circular dependency. Original types.ts replaced with barrel
re-export — zero changes to 36 import sites.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 04:01:58 +01:00
arkonandClaude Opus 4.6 a27be18a5a fix: address review findings from route extraction refactor
- Add addSession() to SessionPort, replacing unsafe ReadonlyMap casts
- Restore getLightSessionsState() 1s TTL cache for GET /api/sessions
- Deduplicate AUTH_COOKIE_NAME, CASES_DIR, SETTINGS_PATH constants
- Fix eslint no-require-imports error in system-routes.ts
- Remove unused imports (homedir) from cleaned-up route modules

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 03:26:08 +01:00
arkonandClaude Opus 4.6 e05d507254 refactor: extract server.ts routes into domain modules (phase 3)
Split 6,710-line server.ts into focused route modules using port
interfaces for dependency injection. 107/109 routes extracted into
12 domain files with auth middleware, 5 port interfaces, and shared
helpers. Server.ts retains orchestration (SSE, lifecycle, state).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 19:12:45 +01:00
arkonandClaude Opus 4.6 e0a2774d37 perf: implement phase 1-3 performance optimizations
Add implementation plans and code structure analysis for a 3-phase
performance optimization effort. Refactor core modules to reduce
timer overhead, consolidate regex usage, extract exec timeout config,
add debouncer utility, and streamline server/schema validation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 17:39:44 +01:00
arkonandClaude Opus 4.6 562b14ab61 fix: rename GitHub release tags from aicodeman to codeman
npm package stays as aicodeman (name taken), but GitHub releases
now show as codeman@x.y.z via a post-publish retag step.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 15:05:12 +01:00
arkonandClaude Opus 4.6 db550a0e28 fix: format server.ts to pass CI prettier check
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 14:50:59 +01:00
213 changed files with 46086 additions and 17311 deletions
+22
View File
@@ -42,3 +42,25 @@ jobs:
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Rename release tag to codeman
if: steps.changesets.outputs.published == 'true'
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
VERSION=$(node -p "require('./package.json').version")
OLD_TAG="aicodeman@${VERSION}"
NEW_TAG="codeman@${VERSION}"
# Update the GitHub release BEFORE deleting the old tag
RELEASE_ID=$(gh release view "$OLD_TAG" --json databaseId -q .databaseId 2>/dev/null || true)
if [ -n "$RELEASE_ID" ]; then
gh api -X PATCH "repos/${{ github.repository }}/releases/${RELEASE_ID}" \
-f tag_name="$NEW_TAG" \
-f name="$NEW_TAG"
fi
# Retag
git tag "$NEW_TAG" "$OLD_TAG" 2>/dev/null || true
git tag -d "$OLD_TAG" 2>/dev/null || true
git push origin "$NEW_TAG" ":refs/tags/$OLD_TAG" 2>/dev/null || true
+1
View File
@@ -60,3 +60,4 @@ media-assets/
commands
todo.md
@fix_plan.md
readme-preview.mjs
+130
View File
@@ -1,5 +1,135 @@
# aicodeman
## 0.4.0
### Minor Changes
- Add CJK IME input textarea for xterm.js terminal (env toggle INPUT_CJK_FORM=ON). Always-visible textarea below terminal handles native browser IME composition, forwarding completed text to PTY on Enter. Supports arrow keys, Ctrl combos, backspace passthrough, and Escape to clear.
Add fork installation support to install.sh with CODEMAN_REPO_URL and CODEMAN_BRANCH env vars, allowing custom repository and branch for git clone/update operations. README updated with fork installation instructions.
Fix WebSocket session lifecycle: close WS connections when session exits (prevents orphaned listeners and stale writes to dead PTY), add readyState guard in onTerminal to stop buffering after socket closes, simplify heartbeat by removing redundant alive flag.
Add WebSocket reconnection with exponential backoff (1s-10s) on unexpected close, skipping server rejection codes (4004/4008/4009). Falls back gracefully to SSE+POST during reconnection.
Clear CJK textarea on session switch to prevent sending stale text to wrong session.
## 0.3.12
### Patch Changes
- Add WebSocket terminal I/O with server-side DEC 2026 synchronized update markers. Replaces per-keystroke HTTP POST + SSE terminal output with a single bidirectional WebSocket connection for dramatically lower input latency. Server-side 8ms micro-batching with 16KB flush threshold groups rapid PTY events into single WS frames wrapped in DEC 2026 markers for flicker-free atomic rendering. Includes 30s ping/pong heartbeat with 10s timeout for stale connection detection through tunnels. Existing SSE + HTTP POST paths remain fully functional as transparent fallback. Resize messages validated to match HTTP route bounds (cols 1-500, rows 1-200, integers only). 16 automated route tests added for WS endpoint. Also patches 5 dependency vulnerabilities (basic-ftp, fastify, minimatch, serialize-javascript).
## 0.3.11
### Patch Changes
- ### Session Resume & History
- Add `resumeSessionId` support for conversation resume after reboot
- Add history session resume UI and API with route shell sessions routing fix
- Improve session resume reliability and persist user settings across refresh
- Correct `claudeSessionId` for resumed sessions
### Terminal & Frontend
- Upgrade xterm.js 5.3 → 6.0 with native DEC 2026 synchronized output
- Increase terminal scrollback from 5,000 to 20,000 lines
- Reduce default font size and persist tab state across refresh
- Resolve terminal resize scrollback ghost renders
- Hide subagent monitor panel by default
### Installer
- Auto-detect existing install and run update instead of fresh install
- Auto-restart codeman-web service after update if running
- Show restart command when codeman-web is not a systemd service
- Fix one-liner restart command for background processes
### Codebase Quality
- Remove dead code, consolidate imports, extract constants
- Repair 15 pre-existing subagent-watcher test failures
- Clean up DEC sync dead code
## 0.3.10
### Patch Changes
- - feat: upgrade xterm.js from 5.3 to 6.0 with native DEC 2026 synchronized output support
- feat: add history session resume UI and API — resume Claude conversations after reboot
- feat: add resumeSessionId support for conversation resume across session restarts
- feat: persist active tabs across page refresh
- feat: improve session resume reliability and persist user settings
- perf: increase terminal scrollback from 5,000 to 20,000 lines
- fix: resolve terminal resize scrollback ghost renders
- fix: route shell sessions to correct endpoint on tab click
- fix: correct claudeSessionId for resumed sessions (use original Claude conversation ID)
- fix: increase default desktop font size from 12 to 14
- refactor: extract shared \_fetchHistorySessions() method to eliminate duplication
- refactor: remove dead DEC 2026 sync code (extractSyncSegments, DEC_SYNC_START/END constants)
## 0.3.9
### Patch Changes
- Add content-hash cache busting for static assets — build step now renames JS/CSS files with MD5 content hashes (e.g. app.js → app.94b71235.js) and rewrites index.html references. HTML served with Cache-Control: no-cache so browsers always revalidate and pick up new hashed filenames after deploys. Hashed assets keep immutable 1-year cache. Eliminates the need for manual hard refresh (Ctrl+Shift+R) after deployments.
Refactor path traversal validation into shared validatePathWithinBase() helper in route-helpers.ts, replacing 6 duplicate inline checks across case-routes, plan-routes, and session-routes.
Deduplicate stripAnsi in bash-tool-parser.ts — use shared utility from utils/index.ts instead of private method.
## 0.3.8
### Patch Changes
- Add tunnel status indicator with control panel — green pulsing dot in header when Cloudflare tunnel is active, dropdown with URL, remote clients, auth sessions, and start/stop/QR/revoke controls
## 0.3.7
### Patch Changes
- Operation Lightspeed: 5 parallel performance optimizations — multi-layer backpressure to prevent terminal write freezes, TERMINAL_TAIL_SIZE constant with client-drop recovery, tab switching SSE gating, and local echo improvements
- Codebase cleanup: remove dead code (unused token validation exports, PlanPhase alias), add execPattern() regex helper to eliminate repetitive .lastIndex resets, centralize 11 magic number constants into config files, fix CLAUDE.md inaccuracies, and add 316 new tests for utilities, respawn helpers, and system-routes
## 0.3.6
### Patch Changes
- Re-enable WebGL renderer with 48KB/frame flush cap protection against GPU stalls
## 0.3.5
### Patch Changes
- Fix Chrome "page unresponsive" crashes caused by xterm.js WebGL renderer GPU stalls during heavy terminal output. Disable WebGL by default (canvas renderer used instead), gate SSE terminal writes during tab switches, and add crash diagnostics with server-side breadcrumb collection.
## 0.3.4
### Patch Changes
- Fix Chrome tab freeze from flicker filter buffer accumulation during active sessions, and fix shell mode feedback delay by excluding shell sessions from cursor-up filter
## 0.3.3
### Patch Changes
- fix: eliminate WebGL re-render flicker during tab switch by keeping renderer active instead of toggling it off/on around large buffer writes
## 0.3.2
### Patch Changes
- Make file browser panel draggable by its header
## 0.3.1
### Patch Changes
- LLM context optimization and performance improvements: compress CLAUDE.md 21%, MEMORY.md 61%; SSE broadcast early return, cached tunnel state, cache invalidation fix, ralph todo cleanup timer; frontend SSE listener leak fix, short ID caching, subagent window handle cleanup; 100% @fileoverview coverage
## 0.3.0
### Minor Changes
- QR code authentication for tunnel access, 7-phase codebase refactor (route extraction, type domain modules, frontend module split, config consolidation, managed timers, test infrastructure), overlay rendering fixes, and security hardening
## 0.2.9
### Patch Changes
+98 -407
View File
@@ -29,7 +29,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
2. **Frontend changes**: Use Playwright to load the page and assert the UI renders correctly. Use `waitUntil: 'domcontentloaded'` (not `networkidle` — SSE keeps the connection open). Wait 3-4s for polling/async data to populate, then check element visibility, text content, and CSS values
3. **Only after verification passes**, proceed with COM
The production server caches static files for 1 hour (`maxAge: '1h'` in `server.ts`). After deploying frontend changes, users may need a hard refresh (Ctrl+Shift+R) to see updates.
The production server caches static files for 1 year (`maxAge: '1y'` in `server.ts`). After deploying frontend changes, users may need a hard refresh (Ctrl+Shift+R) to see updates.
## COM Shorthand (Deployment)
@@ -44,7 +44,7 @@ When user says "COM":
"aicodeman": patch
---
Description of changes
Detailed description of ALL changes since last release (not just the most recent commit — review full git log since last version tag)
CHANGESET
```
Replace `patch` with `minor` or `major` as needed. Include `"xterm-zerolag-input": patch` on a separate line if that package changed too.
@@ -52,7 +52,7 @@ When user says "COM":
4. **Sync CLAUDE.md version**: Update the `**Version**` line below to match the new version from `package.json`
5. **Commit and deploy**: `git add -A && git commit -m "chore: version packages" && git push && npm run build && systemctl --user restart codeman-web`
**Version**: 0.2.9 (must match `package.json`)
**Version**: 0.4.0 (must match `package.json`)
## Project Overview
@@ -60,135 +60,65 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Tech Stack**: TypeScript (ES2022/NodeNext, strict mode), Node.js, Fastify, node-pty, xterm.js. Supports both Claude Code and OpenCode AI CLIs via pluggable CLI resolvers.
**TypeScript Strictness** (see `tsconfig.json`): `noUnusedLocals`, `noUnusedParameters`, `noImplicitReturns`, `noImplicitOverride`, `noFallthroughCasesInSwitch`, `allowUnreachableCode: false`, `allowUnusedLabels: false`. Note: `src/tui` is excluded from compilation (legacy/deprecated code path).
**TypeScript Strictness** (see `tsconfig.json`): `noUnusedLocals`, `noUnusedParameters`, `noImplicitReturns`, `noImplicitOverride`, `noFallthroughCasesInSwitch`, `allowUnreachableCode: false`, `allowUnusedLabels: false`.
**Requirements**: Node.js 18+, Claude CLI, tmux
## Commands
**Git**: Main branch is `master`. SSH session chooser: `sc` (interactive), `sc 2` (quick attach), `sc -l` (list).
**Note**: `npm run dev` starts the web server (equivalent to `npx tsx src/index.ts web`).
## Additional Commands
**Default port**: `3000` (web UI at `http://localhost:3000`)
`npm run dev` = dev server. Default port: `3000`. Commands not in Quick Reference:
```bash
# Setup
npm install # Install dependencies
| Task | Command |
|------|---------|
| Dev with TLS | `npx tsx src/index.ts web --https` |
| Continuous typecheck | `tsc --noEmit --watch` |
| Test coverage | `npm run test:coverage` |
| Production start | `npm run start` |
| Production logs | `journalctl --user -u codeman-web -f` |
# Development
npx tsx src/index.ts web # Dev server (RECOMMENDED)
npx tsx src/index.ts web --https # With TLS (only needed for remote access)
npm run typecheck # Type check
tsc --noEmit --watch # Continuous type checking
npm run lint # ESLint
npm run lint:fix # ESLint with auto-fix
npm run format # Prettier format
npm run format:check # Prettier check only
**CI**: `.github/workflows/ci.yml` runs `typecheck`, `lint`, `format:check` on push to master (Node 22). Tests excluded (they spawn tmux).
# Testing (see "Testing" section for CRITICAL safety warnings)
npx vitest run test/<file>.test.ts # Single file (SAFE)
npx vitest run -t "pattern" # Tests matching name
npm run test:coverage # With coverage report
# Production
npm run build # esbuild via scripts/build.mjs (not tsc)
npm run start # node dist/index.js (production)
systemctl --user restart codeman-web
journalctl --user -u codeman-web -f
```
**CI**: `.github/workflows/ci.yml` runs `typecheck`, `lint`, and `format:check` on push to master. Tests are intentionally excluded from CI (they spawn tmux).
**Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`). ESLint flat config (`eslint.config.js`) allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Ignores: `app.js`, `scripts/**/*.mjs`, `src/web/public/vendor/**`, `tools/**`, `remotion/**`.
## Common Gotchas
- **Single-line prompts only** — `writeViaMux()` sends text and Enter separately; multi-line breaks Ink
- **Don't kill tmux sessions blindly** — Check `$CODEMAN_MUX` first; you might be inside one
- **Global regex `lastIndex` sharing** — `ANSI_ESCAPE_PATTERN_FULL/SIMPLE` have `g` flag; use `createAnsiPatternFull/Simple()` factory functions for fresh instances in loops
- **DEC 2026 sync blocks** — Never discard incomplete sync blocks (START without END); buffer up to 50ms then flush. See `app.js:extractSyncSegments()`
- **Terminal writes during buffer load** — Live SSE writes are queued while `_isLoadingBuffer` is true to prevent interleaving with historical data
- **Local echo prompt scanning** — Does NOT use `buffer.cursorY` (Ink moves it); scans buffer bottom-up for visible `>` prompt marker
- **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink
- **ESM only** — Never `require()`, use `await import()`. `tsx` masks CJS/ESM issues in dev but production breaks
- **Package ≠ product name** — npm: `aicodeman`, product: **Codeman**. Release renames tags accordingly
- **Global regex `lastIndex`** — Use `createAnsiPatternFull/Simple()` factories, not shared `g`-flag patterns in loops
## Import Conventions
- **Utilities**: Import from `./utils` (re-exports all): `import { LRUMap, stripAnsi } from './utils'`
- **Types**: Use type imports: `import type { SessionState } from './types'`
- **Config**: Import from specific files: `import { MAX_TERMINAL_BUFFER_SIZE } from './config/buffer-limits'`
**Import conventions**: Utils from `./utils`, types from `./types` (barrel), config from specific `./config/*` files.
## Architecture
### Core Files
### Core Files (by domain)
| File | Purpose |
|------|---------|
| `src/index.ts` | CLI entry point: global error recovery, uncaught exception guard, `MAX_CONSECUTIVE_ERRORS` auto-restart |
| `src/session.ts` | PTY wrapper: `runPrompt()`, `startInteractive()`, `startShell()` |
| `src/mux-interface.ts` | `TerminalMultiplexer` interface + `MuxSession` type |
| `src/mux-factory.ts` | Create tmux multiplexer instance |
| `src/tmux-manager.ts` | tmux session management |
| `src/session-manager.ts` | Session lifecycle, cleanup |
| `src/state-store.ts` | State persistence to `~/.codeman/state.json` |
| `src/respawn-controller.ts` | State machine for autonomous cycling |
| `src/ralph-tracker.ts` | Detects `<promise>PHRASE</promise>`, todos |
| `src/ralph-loop.ts` | Autonomous task execution loop (polls queue, assigns tasks) |
| `src/ralph-config.ts` | Parses `.claude/ralph-loop.local.md` plugin config |
| `src/task.ts` | Task model for prompt execution |
| `src/task-queue.ts` | Priority queue for tasks with dependencies |
| `src/task-tracker.ts` | Background task tracker for subagent detection |
| `src/subagent-watcher.ts` | Monitors Claude Code's Task tool (background agents) |
| `src/team-watcher.ts` | Polls `~/.claude/teams/` for agent team activity; matches teams to sessions via `leadSessionId` |
| `src/run-summary.ts` | Timeline events for "what happened while away" |
| `src/ai-checker-base.ts` | Base class for AI-powered checkers (shared by idle + plan checkers) |
| `src/ai-idle-checker.ts` | AI-powered idle detection |
| `src/ai-plan-checker.ts` | AI-powered plan completion checker |
| `src/bash-tool-parser.ts` | Parses Claude's bash tool invocations from output |
| `src/transcript-watcher.ts` | Watches Claude's transcript files for changes |
| `src/hooks-config.ts` | Manages `.claude/settings.local.json` hook configuration |
| `src/push-store.ts` | VAPID key auto-gen + push subscription CRUD for Web Push |
| `src/session-lifecycle-log.ts` | Append-only JSONL audit log at `~/.codeman/session-lifecycle.jsonl` |
| `src/image-watcher.ts` | Watches for image file creation (screenshots, etc.) |
| `src/file-stream-manager.ts` | Manages `tail -f` processes for live log viewing |
| `src/plan-orchestrator.ts` | Multi-agent plan generation with research and planning phases |
| `src/prompts/index.ts` | Barrel export for all agent prompts |
| `src/prompts/*.ts` | Agent prompts (research-agent, planner) |
| `src/templates/claude-md.ts` | CLAUDE.md generation for new cases |
| `src/tunnel-manager.ts` | Manages cloudflared child process for Cloudflare tunnel remote access |
| `src/cli.ts` | Command-line interface handlers |
| `src/web/server.ts` | Fastify REST API + SSE at `/api/events` (~280 routes) |
| `src/web/schemas.ts` | Zod v4 validation schemas with path/env security allowlists |
| `src/web/public/app.js` | Frontend: xterm.js, tab management, subagent windows, mobile support (~15K lines) |
| `src/types.ts` | All TypeScript interfaces (~70 type/interface/enum defs, ~1450 lines) |
| Domain | Key files | Notes |
|--------|-----------|-------|
| **Entry** | `src/index.ts`, `src/cli.ts` | |
| **Session** | `src/session.ts` ★, `src/session-manager.ts`, `src/session-auto-ops.ts`, `src/session-cli-builder.ts`, `src/session-lifecycle-log.ts`, `src/session-task-cache.ts` | |
| **Mux** | `src/mux-interface.ts`, `src/mux-factory.ts`, `src/tmux-manager.ts` | |
| **Respawn** | `src/respawn-controller.ts` ★ + 4 helpers (`-adaptive-timing`, `-health`, `-metrics`, `-patterns`) | Read `docs/respawn-state-machine.md` first |
| **Ralph** | `src/ralph-tracker.ts` ★, `src/ralph-loop.ts` + 5 helpers (`-config`, `-fix-plan-watcher`, `-plan-tracker`, `-stall-detector`, `-status-parser`) | Read `docs/ralph-wiggum-guide.md` first |
| **Agents** | `src/subagent-watcher.ts` ★, `src/team-watcher.ts`, `src/bash-tool-parser.ts`, `src/transcript-watcher.ts` | |
| **AI** | `src/ai-checker-base.ts`, `src/ai-idle-checker.ts`, `src/ai-plan-checker.ts` | |
| **Tasks** | `src/task.ts`, `src/task-queue.ts`, `src/task-tracker.ts` | |
| **State** | `src/state-store.ts`, `src/run-summary.ts`, `src/session-lifecycle-log.ts` | |
| **Infra** | `src/hooks-config.ts`, `src/push-store.ts`, `src/tunnel-manager.ts`, `src/image-watcher.ts`, `src/file-stream-manager.ts` | |
| **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/claude-md.ts` | |
| **Web** | `src/web/server.ts`, `src/web/sse-events.ts`, `src/web/routes/*.ts` (12 route modules + barrel), `src/web/ports/*.ts`, `src/web/middleware/auth.ts`, `src/web/schemas.ts` | |
| **Frontend** | `src/web/public/app.js` ★ (~12.5K lines) + 9 JS modules (incl. `sw.js` service worker) | |
| **Types** | `src/types/index.ts` → 13 domain files | See `@fileoverview` in index.ts |
**Large files** (>50KB): `app.js`, `ralph-tracker.ts`, `respawn-controller.ts`, `session.ts`, `subagent-watcher.ts` — these contain complex state machines; read `docs/respawn-state-machine.md` before modifying.
★ = Large file (>50KB). All files have `@fileoverview` JSDoc — read that before diving in.
### Local Packages
**Local package**: `packages/xterm-zerolag-input/` — local echo overlay for xterm.js; copy embedded in `app.js`.
| Package | Purpose |
|---------|---------|
| `packages/xterm-zerolag-input/` | Instant keystroke feedback overlay for xterm.js — eliminates perceived input latency over high-RTT connections. Source of truth for `LocalEchoOverlay`; a copy is embedded in `app.js`. Build: `npm run build` (tsup). |
**Config**: `src/config/` — 9 files. Import from specific files, not barrel.
### Config Files (`src/config/`)
| File | Purpose |
|------|---------|
| `buffer-limits.ts` | Terminal/text buffer size limits |
| `map-limits.ts` | Global limits for Maps, sessions, watchers |
### Utilities (`src/utils/`)
Re-exported via `src/utils/index.ts`. Key exports:
| File | Exports |
|------|---------|
| `cleanup-manager.ts` | `CleanupManager` — centralized disposal for timers, intervals, watchers, listeners, streams |
| `lru-map.ts` | `LRUMap` — bounded cache with eviction |
| `stale-expiration-map.ts` | `StaleExpirationMap` — TTL-based map with automatic cleanup |
| `regex-patterns.ts` | `ANSI_ESCAPE_PATTERN_FULL/SIMPLE`, `createAnsiPatternFull/Simple()`, `stripAnsi`, `TOKEN_PATTERN`, `SPINNER_PATTERN` |
| `buffer-accumulator.ts` | `BufferAccumulator` — batches rapid writes into single flushes |
| `claude-cli-resolver.ts` | `findClaudeDir`, `getAugmentedPath` — resolves Claude CLI paths |
| `opencode-cli-resolver.ts` | `resolveOpenCodeDir`, `isOpenCodeAvailable` — OpenCode CLI support |
| `string-similarity.ts` | `stringSimilarity`, `fuzzyPhraseMatch`, `todoContentHash` |
| `token-validation.ts` | `validateTokenCounts`, `validateTokensAndCost` |
| `nice-wrapper.ts` | `wrapWithNice` — wraps commands with `nice`/`ionice` for lower priority |
| `type-safety.ts` | `assertNever` — exhaustive switch/case guard |
**Utilities**: `src/utils/` — re-exported via index. Key: `CleanupManager`, `LRUMap`, `StaleExpirationMap`, `BufferAccumulator`, `stripAnsi`, `Debouncer`.
### Data Flow
@@ -199,356 +129,117 @@ Re-exported via `src/utils/index.ts`. Key exports:
### Key Patterns
**Input to sessions**: Use `session.writeViaMux()` for programmatic input (respawn, auto-compact). Uses tmux `send-keys -l` (literal text) + `send-keys Enter`. All prompts must be single-line.
**Terminal multiplexer**: `TerminalMultiplexer` interface (`src/mux-interface.ts`) abstracts the backend. `createMultiplexer()` from `src/mux-factory.ts` creates the tmux backend.
**Input**: `session.writeViaMux()` for programmatic input — tmux `send-keys -l` (literal) + `send-keys Enter`. Single-line only.
**Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`.
**Token tracking**: Interactive mode parses status line ("123.4k tokens"), estimates 60/40 input/output split.
**Hook events**: Claude Code hooks trigger via `/api/hook-event`. Key events: `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`. See `src/hooks-config.ts`.
**Hook events**: Claude Code hooks trigger notifications via `/api/hook-event`. Key events: `permission_prompt` (tool approval needed), `elicitation_dialog` (Claude asking question), `idle_prompt` (waiting for input), `stop` (response complete), `teammate_idle` (Agent Teams), `task_completed` (Agent Teams). See `src/hooks-config.ts`.
**Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `agent-teams/`.
**Web Push**: Layer 5 of the notification system. Service worker (`sw.js`) receives push events and shows OS-level notifications even when the browser tab is closed. VAPID keys auto-generated on first use and persisted to `~/.codeman/push-keys.json`. Per-subscription per-event preferences stored in `~/.codeman/push-subscriptions.json`. Expired subscriptions (410/404) auto-cleaned. Requires HTTPS or localhost. iOS requires PWA installed to home screen. See `src/push-store.ts`.
**Circuit breaker**: Prevents respawn thrashing. States: `CLOSED` → `HALF_OPEN` → `OPEN`. Reset: `/api/sessions/:id/ralph-circuit-breaker/reset`.
**Agent Teams (experimental)**: `TeamWatcher` polls `~/.claude/teams/` for team configs and matches teams to sessions via `leadSessionId`. Teammates are in-process threads (not separate OS processes) and appear as standard subagents. RespawnController checks `TeamWatcher.hasActiveTeammates()` before triggering respawn. Enable via `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1` env var in `settings.local.json`. See `agent-teams/` for full docs.
**Port interfaces**: Routes declare dependencies via port interfaces (`src/web/ports/`). Routes use intersection types (e.g., `SessionPort & EventPort`).
**Circuit breaker**: Prevents respawn thrashing when Claude is stuck. States: `CLOSED` (normal) → `HALF_OPEN` (testing) → `OPEN` (blocked). Tracks consecutive no-progress, same-error-repeated, and tests-failing-too-long. Reset via API at `/api/sessions/:id/ralph-circuit-breaker/reset`.
### Frontend
**Respawn cycle metrics & health scoring**: `RespawnCycleMetrics` tracks per-cycle outcomes (success, stuck_recovery, blocked, error). `RalphLoopHealthScore` computes 0-100 health with component scores (cycleSuccess, circuitBreaker, iterationProgress, aiChecker, stuckRecovery). Available via respawn status API.
**Subagent-session correlation**: Session parses Task tool output via `BashToolParser` → `SubagentWatcher` discovers new agent → calls `session.findTaskDescriptionNear()` to match description for window title.
### Frontend Files
| File | Purpose |
|------|---------|
| `src/web/public/index.html` | HTML entry point with inline critical CSS and async vendor loading |
| `src/web/public/app.js` | Core UI: xterm.js, tab management, subagent windows, mobile support (~15K lines) |
| `src/web/public/ralph-wizard.js` | Ralph Loop wizard UI extracted from app.js (~1K lines) |
| `src/web/public/styles.css` | Main styling (dark theme, layout, components) |
| `src/web/public/mobile.css` | Responsive overrides for screens <1024px (loaded conditionally via `media` attribute) |
| `src/web/public/upload.html` | Screenshot upload page served at `/upload.html` |
| `src/web/public/sw.js` | Service worker for Web Push notifications |
| `src/web/public/manifest.json` | Minimal PWA manifest (required for push on Android) |
| `src/web/public/vendor/` | Self-hosted xterm.js + addons (eliminates CDN latency) |
### Frontend Architecture (`app.js`)
The frontend is a single ~15K-line vanilla JS file with these key systems:
| System | Key Classes/Functions | Purpose |
|--------|----------------------|---------|
| **Terminal rendering** | `batchTerminalWrite()`, `flushPendingWrites()`, `chunkedTerminalWrite()` | 60fps batched writes with DEC 2026 sync |
| **Local echo overlay** | `LocalEchoOverlay` class | DOM overlay for instant mobile keystroke feedback |
| **Mobile support** | `MobileDetection`, `KeyboardHandler`, `SwipeHandler`, `KeyboardAccessoryBar` | Touch input, viewport adaptation, swipe navigation |
| **Subagent windows** | `openSubagentWindow()`, `closeSubagentWindow()`, `updateConnectionLines()` | Floating terminal windows with parent connection lines |
| **Notifications** | `NotificationManager` class | 5-layer: in-app drawer, tab flash, browser API, web push, audio beep |
| **SSE connection** | `connectSSE()`, `addListener()` | EventSource with exponential backoff (1-30s), offline queue (64KB) |
| **Settings** | `openAppSettings()`, `apply*Visibility()` | Server-backed + localStorage persistence |
| **Focus management** | `FocusTrap` class | Modal keyboard navigation with focus restore |
Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `app.js`(6) → `ralph-wizard.js`(7) → `api-client.js`(8) → `subagent-windows.js`(9).
**Z-index layers**: subagent windows (1000), plan agents (1100), log viewers (2000), image popups (3000), local echo overlay (7).
**Built-in respawn presets**: `solo-work` (3s idle, 60min), `subagent-workflow` (45s idle, 240min), `team-lead` (90s idle, 480min), `ralph-todo` (8s idle, 480min, works through @fix_plan.md tasks), `overnight-autonomous` (10s idle, 480min, full reset).
**Respawn presets**: `solo-work` (3s/60min), `subagent-workflow` (45s/240min), `team-lead` (90s/480min), `ralph-todo` (8s/480min), `overnight-autonomous` (10s/480min).
**Keyboard shortcuts**: Escape (close panels), Ctrl+? (help), Ctrl+Enter (quick start), Ctrl+W (kill session), Ctrl+Tab (next session), Ctrl+K (kill all), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl/Cmd +/- (font size).
**Keyboard shortcuts**: Escape (close), Ctrl+? (help), Ctrl+Enter (quick start), Ctrl+W (kill), Ctrl+Tab (next), Ctrl+K (kill all), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl/Cmd +/- (font).
### Security
- **HTTP Basic Auth**: Optional via `CODEMAN_USERNAME`/`CODEMAN_PASSWORD` env vars
- **Session cookies**: After Basic Auth, a 24h session cookie (`codeman_session`) is issued so credentials aren't re-sent on every request. Active sessions auto-extend. SSE works via same-origin cookie (`EventSource` can't send custom headers).
- **Rate limiting**: 10 failed auth attempts per IP triggers 429 rejection (15-minute decay window). Manual `StaleExpirationMap` counter — no `@fastify/rate-limit` needed.
- **Hook bypass**: `/api/hook-event` POST is exempt from auth — Claude Code hooks curl this from localhost and can't present credentials. Safe: validated by `HookEventSchema`, only triggers broadcasts.
- **CORS**: Restricted to localhost only
- **Security headers**: X-Content-Type-Options, X-Frame-Options, CSP; HSTS if HTTPS
- **Path validation** (`schemas.ts`): Strict allowlist regex, no shell metacharacters, no traversal, must be absolute
- **Env var allowlist**: Only `CLAUDE_CODE_*` prefixes allowed; blocks `PATH`, `LD_PRELOAD`, `NODE_OPTIONS`, `CODEMAN_*` keys
- **File streaming TOCTOU protection**: `FileStreamManager` calls `realpathSync()` twice (at validation and before spawn) to catch symlink swaps
| Layer | Details |
|-------|---------|
| **Auth** | Optional HTTP Basic via `CODEMAN_USERNAME`/`CODEMAN_PASSWORD` env vars |
| **QR Auth** | Single-use 6-char tokens (60s TTL) for tunnel login. See `docs/qr-auth-plan.md` |
| **Sessions** | 24h cookie (`codeman_session`), auto-extend, device context audit |
| **Rate limit** | 10 failed auth/IP → 429 (15min decay). QR has separate limiter |
| **Hook bypass** | `/api/hook-event` exempt from auth (localhost-only, schema-validated) |
| **Env vars** | `CODEMAN_MUX` (managed session), `CODEMAN_API_URL` (auto-set for hooks) |
| **Validation** | Zod schemas, path allowlist regex, `CLAUDE_CODE_*` env prefix allowlist |
| **Headers** | CORS localhost-only, CSP, X-Frame-Options, HSTS if HTTPS |
### SSE Event Categories
### SSE Event Registry
~80+ event types broadcast via `broadcast()`. Key categories:
~100 event types in `src/web/sse-events.ts` (backend) and `SSE_EVENTS` in `constants.js` (frontend). Both must be kept in sync.
| Category | Events | Purpose |
|----------|--------|---------|
| Session | `session:created/updated/deleted/working/idle/exit/error/completion` | Lifecycle |
| Terminal | `session:terminal`, `session:clearTerminal`, `session:needsRefresh` | Output streaming |
| Respawn | `respawn:stateChanged/cycleStarted/blocked/aiCheck*/planCheck*/timer*` | Respawn state machine |
| Subagent | `subagent:discovered/updated/completed/tool_call/progress` | Background agents |
| Ralph | `session:ralphLoopUpdate/ralphTodoUpdate/ralphCompletionDetected` | Ralph tracking |
| Hooks | `hook:{eventName}` (dynamic) | Claude Code hook events |
| Plan | `plan:started/progress/completed/cancelled/subagent` | Plan orchestration |
| Mux | `mux:created/killed/died/statsUpdated` | tmux process monitor |
| Image | `image:detected` | Screenshot detection |
### API Routes
### API Route Categories
~280 route handlers in `server.ts:buildServer()`. Key groups:
| Group | Prefix | Count | Key endpoints |
|-------|--------|-------|---------------|
| Sessions | `/api/sessions` | ~20 | CRUD, input, resize, interactive, shell |
| Respawn | `/api/sessions/:id/respawn` | 5 | start, stop, enable, config |
| Ralph | `/api/sessions/:id/ralph-*` | 6 | state, status, config, circuit-breaker |
| Plan | `/api/sessions/:id/plan/*` | 5 | task CRUD, checkpoint, history, rollback |
| Subagents | `/api/subagents` | 7 | list, transcript, kill, cleanup |
| Cases | `/api/cases` | 5 | CRUD, link, fix-plan |
| Scheduled | `/api/scheduled` | 4 | CRUD for scheduled runs |
| Push | `/api/push` | 4 | VAPID key, subscribe, update prefs, unsubscribe |
| System | `/api/status`, `/api/stats`, `/api/config`, `/api/settings` | 8 | App state, config |
| Files | `/api/sessions/:id/file*`, `tail-file` | 5 | Browser, preview, raw, tail stream |
| Mux | `/api/mux-sessions` | 4 | tmux management, stats |
~111 handlers across 12 route files in `src/web/routes/`: system (35), sessions (24), ralph (9), plan (8), respawn (7), cases (7), files (5), mux (5), scheduled (4), push (4), teams (2), hooks (1). Each file has `@fileoverview` with endpoint details.
## Adding Features
- **API endpoint**: Types in `types.ts`, route in `server.ts:buildServer()`, use `createErrorResponse()`. Validate request bodies with Zod schemas in `schemas.ts`.
- **SSE event**: Emit via `broadcast()`, handle in `app.js` SSE listener section (search `addListener(`)
- **Session setting**: Add to `SessionState` in `types.ts`, include in `session.toState()`, call `persistSessionState()`
- **Hook event**: Add to `HookEventType` in `types.ts`, add hook command in `hooks-config.ts:generateHooksConfig()`, update `HookEventSchema` in `schemas.ts`
- **Mobile feature**: Add to relevant mobile singleton (`KeyboardHandler`, `KeyboardAccessoryBar`, etc.), test with `MobileDetection.isMobile()` guard
- **New test**: Pick unique port (search `const PORT =`), add port comment to test file header. Tests use ports 3150+.
- **API endpoint**: Types in `src/types/` domain file, route in `src/web/routes/*-routes.ts`, use `createErrorResponse()`. Validate with Zod schemas in `schemas.ts`.
- **SSE event**: Add to `src/web/sse-events.ts` + `SSE_EVENTS` in `constants.js`, emit via `broadcast()`, handle in `app.js` (`addListener(`)
- **Session setting**: Add to `SessionState`, include in `session.toState()`, call `persistSessionState()`
- **Hook event**: Add to `HookEventType`, add hook in `hooks-config.ts:generateHooksConfig()`, update `HookEventSchema`
- **Mobile feature**: Add to relevant singleton, guard with `MobileDetection.isMobile()`
- **New test**: Pick unique port (search `const PORT =`). Route tests use `app.inject()` (no port needed) — see `test/routes/_route-test-utils.ts`.
**Validation**: Uses Zod v4 for request validation. Define schemas in `schemas.ts` and use `.parse()` or `.safeParse()`. Note: Zod v4 has different API from v3 (e.g., `z.object()` options changed, error formatting differs).
**Validation**: Zod v4 (different API from v3). Define schemas in `schemas.ts`, use `.parse()`/`.safeParse()`.
## State Files
| File | Purpose |
|------|---------|
| `~/.codeman/state.json` | Sessions, settings, tokens, respawn config |
| `~/.codeman/mux-sessions.json` | Tmux session metadata for recovery |
| `~/.codeman/settings.json` | User preferences |
| `~/.codeman/push-keys.json` | VAPID key pair for Web Push (auto-generated) |
| `~/.codeman/push-subscriptions.json` | Registered push notification subscriptions |
## Default Settings
UI defaults are set in `src/web/public/app.js` using `??` fallbacks. To change defaults, edit `openAppSettings()` and `apply*Visibility()` functions.
**Key defaults:** Most panels hidden (monitor, subagents shown), notifications enabled (audio disabled), subagent tracking on, Ralph tracking off.
All in `~/.codeman/`: `state.json` (sessions, settings, respawn), `mux-sessions.json` (tmux recovery), `settings.json` (user prefs), `push-keys.json` (VAPID), `push-subscriptions.json`, `session-lifecycle.jsonl` (audit log).
## Testing
**CRITICAL: You are running inside a Codeman-managed tmux session.** Never run `npx vitest run` (full suite) — it spawns/kills tmux sessions and will crash your own session. Instead:
**CRITICAL: You are running inside a Codeman-managed tmux session.** Never run `npx vitest run` (full suite) — it spawns/kills tmux sessions and will crash your own session. Only run individual files:
```bash
# Safe: run individual test files
npx vitest run test/<specific-file>.test.ts
# Safe: run tests matching a pattern
npx vitest run -t "pattern"
# DANGEROUS from inside Codeman — will kill your tmux session:
# npx vitest run ← DON'T DO THIS
npx vitest run test/<specific-file>.test.ts # Single file (SAFE)
npx vitest run -t "pattern" # By name (SAFE)
# npx vitest run # DANGEROUS — DON'T DO THIS
```
**Ports**: Unit tests pick unique ports manually. Search `const PORT =` before adding new tests.
**Config**: Vitest with `globals: true`, `fileParallelism: false`. Timeout 30s, teardown 60s.
**Config**: Vitest with `globals: true`, `fileParallelism: false`. Unit timeout 30s.
**Safety**: `test/setup.ts` snapshots pre-existing tmux sessions and never kills them. Only `registerTestTmuxSession()` sessions get cleaned up.
**Safety**: `test/setup.ts` snapshots pre-existing tmux sessions at load time and never kills them. Only sessions registered via `registerTestTmuxSession()` get cleaned up.
**Ports**: Pick unique ports manually. Search `const PORT =` before adding new tests.
**Respawn tests**: Use MockSession from `test/respawn-test-utils.ts` to avoid spawning real Claude processes.
**Respawn tests**: Use `MockSession` from `test/respawn-test-utils.ts`. **Route tests**: `app.inject()` in `test/routes/`. **Mobile tests**: Playwright suite in `mobile-test/` (135 device profiles).
**Mobile tests**: Separate Playwright-based suite in `mobile-test/` with 135 device profiles. Run via `npx vitest run --config mobile-test/vitest.config.ts`. See `mobile-test/README.md`.
## Screenshots
## Screenshots ("sc")
When the user says "check the sc", "screenshot", or "sc", they mean uploaded screenshots from their mobile device. Screenshots are saved to `~/.codeman/screenshots/` and uploaded via `/upload.html` on the Codeman web UI. To view them, use the Read tool on the image files:
```bash
ls ~/.codeman/screenshots/ # List uploaded screenshots
# Then use Read tool on individual files — Claude Code can view images natively
```
API: `GET /api/screenshots` (list), `GET /api/screenshots/:name` (serve), `POST /api/screenshots` (upload multipart/form-data). Source: `src/web/public/upload.html`.
Mobile screenshots in `~/.codeman/screenshots/`. API: `GET /api/screenshots`, `POST /api/screenshots`.
## Debugging
```bash
tmux list-sessions # List tmux sessions
tmux attach-session -t <name> # Attach (Ctrl+B D to detach)
curl localhost:3000/api/sessions # Check sessions
curl localhost:3000/api/status | jq # Full app state
cat ~/.codeman/state.json | jq # View persisted state
curl localhost:3000/api/subagents # List background agents
curl localhost:3000/api/sessions/:id/run-summary | jq # Session timeline
tmux list-sessions # List tmux sessions
curl localhost:3000/api/sessions | jq # Check sessions
curl localhost:3000/api/status | jq # Full app state
curl localhost:3000/api/subagents | jq # Background agents
cat ~/.codeman/state.json | jq # Persisted state
```
## Troubleshooting
## Performance & Limits
| Problem | Check | Fix |
|---------|-------|-----|
| Session won't start | `tmux list-sessions` for orphans | Kill orphaned sessions, check Claude CLI installed |
| Port 3000 in use | `lsof -i :3000` | Kill conflicting process or use `--port` flag |
| SSE not connecting | Browser console for errors | Check CORS, ensure server running |
| Respawn not triggering | Session settings → Respawn enabled? | Enable respawn, check idle timeout config |
| Terminal blank on tab switch | Network tab for `/api/sessions/:id/buffer` | Check session exists, restart server |
| Tests failing on session limits | `tmux list-sessions \| wc -l` | Clean up: `tmux list-sessions \| grep test \| awk -F: '{print $1}' \| xargs -I{} tmux kill-session -t {}` |
| State not persisting | `cat ~/.codeman/state.json` | Check file permissions, disk space |
Target: 20 sessions, 50 agent windows at 60fps. Limits in `src/config/`: terminal 2MB, text 1MB, messages 1000, max agents 500, max sessions 50, max SSE clients 100. Use `LRUMap` for bounded caches, `StaleExpirationMap` for TTL cleanup. Anti-flicker pipeline: `docs/terminal-anti-flicker.md`.
## Performance Constraints
## References
The app must stay fast with 20 sessions and 50 agent windows:
- 60fps terminal (16ms batching + `requestAnimationFrame`)
- Auto-trimming buffers (2MB terminal max)
- Debounced state persistence (500ms)
- SSE adaptive batching: 16ms (normal), 32ms (moderate), 50ms (rapid); immediate flush at 32KB
- SSE backpressure handling: skip writes to backpressured clients, recover via `session:needsRefresh` on drain
- Cached endpoints: `/api/sessions` and `/api/status` use 1s TTL caches to avoid expensive serialization
- Frontend buffer loads: 128KB chunks via `requestAnimationFrame` to prevent UI jank
## Terminal Anti-Flicker System
Claude Code uses Ink (React for terminals), which redraws the screen on every state change. Codeman implements a 6-layer anti-flicker pipeline for smooth 60fps output:
```
PTY Output → Server Batching (16-50ms) → DEC 2026 Wrap → SSE → Client rAF → xterm.js
```
**Key functions:** `server.ts:batchTerminalData()`, `server.ts:flushTerminalBatches()`, `app.js:batchTerminalWrite()`, `app.js:extractSyncSegments()`
**Typical latency:** 16-32ms. Optional per-session flicker filter adds ~50ms for problematic terminals.
See `docs/terminal-anti-flicker.md` for full implementation details (adaptive batching, DEC 2026 markers, edge cases).
## Resource Limits
Limits are centralized in `src/config/buffer-limits.ts` and `src/config/map-limits.ts`.
**Buffer limits** (per session):
| Buffer | Max | Trim To |
|--------|-----|---------|
| Terminal | 2MB | 1.5MB |
| Text output | 1MB | 768KB |
| Messages | 1000 | 800 |
**Map limits** (global):
| Resource | Max |
|----------|-----|
| Tracked agents | 500 |
| Concurrent sessions | 50 |
| SSE clients total | 100 |
| File watchers | 500 |
Use `LRUMap` for bounded caches with eviction, `StaleExpirationMap` for TTL-based cleanup.
## Where to Find More Information
| Topic | Location |
|-------|----------|
| **Respawn state machine** | `docs/respawn-state-machine.md` |
| **Ralph Loop guide** | `docs/ralph-wiggum-guide.md` |
| **Claude Code hooks** | `docs/claude-code-hooks-reference.md` |
| **Terminal anti-flicker** | `docs/terminal-anti-flicker.md` |
| **Agent Teams (experimental)** | `agent-teams/README.md`, `agent-teams/design.md` |
| **API routes** | `src/web/server.ts:buildServer()` or README.md |
| **SSE events** | Search `broadcast(` in `server.ts` |
| **Session statuses** | `SessionStatus` type in `src/types.ts` |
| **Error codes** | `createErrorResponse()` in `src/types.ts` |
| **Test utilities** | `test/respawn-test-utils.ts` |
| **Mobile test suite** | `mobile-test/README.md` |
| **OpenCode integration** | `docs/opencode-integration.md` |
| **Local echo overlay** | `docs/local-echo-overlay-plan.md` |
| **Performance investigation** | `docs/performance-investigation-report.md` |
| **First-load optimization** | `docs/first-load-optimization-plan.md`, `docs/perf-audit-first-load.md` |
| **Dead code audit** | `docs/cleanup-findings.md` |
| **TypeScript improvements** | `docs/typescript-improvement-suggestions.md` |
| **Browser testing** | `docs/browser-testing-guide.md` |
| **Mobile testing report** | `docs/mobile-testing-report.md` |
| **Voice input** | `docs/voice-input-plan.md` |
| **Improvement roadmaps** | `docs/respawn-improvement-plan.md`, `docs/ralph-improvement-plan.md`, `docs/plan-improvement-roadmap.md` |
| **Background keystroke forwarding** | `docs/background-keystroke-forwarding-merged-plan.md` |
| **Run summary** | `docs/run-summary-plan.md` |
Additional design docs and investigation reports are in the `docs/` directory.
Deep-dive docs in `docs/`: `respawn-state-machine.md`, `ralph-wiggum-guide.md`, `claude-code-hooks-reference.md`, `terminal-anti-flicker.md`, `opencode-integration.md`, `qr-auth-plan.md`. Agent Teams: `agent-teams/README.md`. SSE events: `src/web/sse-events.ts` + `constants.js`.
## Scripts
| Script | Purpose |
|--------|---------|
| `scripts/tmux-manager.sh` | Safe tmux session management (use instead of direct kill commands) |
| `scripts/monitor-respawn.sh` | Monitor respawn state machine in real-time |
| `scripts/watch-subagents.ts` | Real-time subagent transcript watcher (list, follow by session/agent ID) |
| `scripts/codeman-web.service` | systemd service file for production deployment |
| `scripts/codeman-tunnel.service` | systemd service file for persistent Cloudflare tunnel |
| `scripts/tunnel.sh` | Start/stop/check Cloudflare quick tunnel (`./scripts/tunnel.sh start\|stop\|url`) |
| `scripts/build.mjs` | esbuild-based production build (called by `npm run build`) |
| `scripts/postinstall.js` | npm postinstall hook for setup |
Additional scripts in `scripts/` for screenshots, demos, Ralph wizards, and browser testing.
Key: `scripts/tmux-manager.sh` (safe tmux mgmt), `scripts/tunnel.sh` (tunnel start/stop/url). Production: `scripts/codeman-web.service`, `scripts/codeman-tunnel.service`.
## Memory Leak Prevention
Frontend runs long (24+ hour sessions); all Maps/timers must be cleaned up.
### Cleanup Patterns
When adding new event listeners or timers:
1. Store handler references for later removal
2. Add cleanup to appropriate `stop()` or `cleanup*()` method
3. For singleton watchers, store refs in class properties and remove in server `stop()`
**Backend**: Clear Maps in `stop()`, null promise callbacks on error, remove watcher listeners on shutdown. Use `CleanupManager` for centralized disposal — supports timers, intervals, watchers, listeners, streams. Guard async callbacks with `if (this.cleanup.isStopped) return`.
**Frontend**: Store drag/resize handlers on elements, clean up in `close*()` functions. SSE reconnect calls `handleInit()` which resets state. SSE listeners are tracked in an array and removed on reconnect to prevent accumulation.
Run `npx vitest run test/memory-leak-prevention.test.ts` to verify patterns.
24+ hour sessions: use `CleanupManager`, clear Maps in `stop()`, guard async with `if (this.cleanup.isStopped) return`. Frontend: store handler refs, clean in `close*()`. Verify: `npx vitest run test/memory-leak-prevention.test.ts`.
## Common Workflows
**Investigating a bug**: Start dev server (`npx tsx src/index.ts web`), reproduce in browser, check terminal output and `~/.codeman/state.json` for clues.
**Bug investigation**: Dev server → reproduce in browser → check terminal + `~/.codeman/state.json`.
**Respawn changes**: Read `docs/respawn-state-machine.md` first. Use `MockSession` from `test/respawn-test-utils.ts`.
**Adding a new API endpoint**: Define types in `types.ts`, add route in `server.ts:buildServer()`, broadcast SSE events if needed, handle in `app.js:handleSSEEvent()`.
## Tunnel
**Modifying respawn behavior**: Study `docs/respawn-state-machine.md` first. The state machine is in `respawn-controller.ts`. Use MockSession from `test/respawn-test-utils.ts` for testing.
**Modifying mobile behavior**: Mobile singletons (`MobileDetection`, `KeyboardHandler`, `SwipeHandler`, `KeyboardAccessoryBar`) all have `init()`/`cleanup()` lifecycle. KeyboardHandler uses `visualViewport` API for iOS keyboard detection (100px threshold for address bar drift). All mobile handlers are re-initialized after SSE reconnect to prevent stale closures.
**Adding a file watcher**: Use `ImageWatcher` as a template pattern — chokidar with `awaitWriteFinish`, burst throttling (max 20/10s), debouncing (200ms), and auto-ignore of `node_modules/.git/dist/`.
## Tunnel Setup (Remote Access)
Access Codeman from mobile/remote devices via Cloudflare quick tunnel.
```
Browser → Cloudflare Edge (HTTPS) → cloudflared → localhost:3000
```
**Prerequisites**: `cloudflared` installed (`cloudflared --version`), `CODEMAN_PASSWORD` set in environment.
### Quick Start
```bash
# Via CLI
./scripts/tunnel.sh start # Start tunnel, prints public URL
./scripts/tunnel.sh url # Show current URL
./scripts/tunnel.sh stop # Stop tunnel
# Via web UI: Settings → Tunnel → Toggle On
```
### systemd Service (Persistent)
```bash
# Install and enable
cp scripts/codeman-tunnel.service ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now codeman-tunnel
# Check logs
journalctl --user -u codeman-tunnel -f
```
### Auth Flow
1. First request → browser shows Basic Auth prompt (username: `admin` or `CODEMAN_USERNAME`)
2. On success → server issues `codeman_session` HttpOnly cookie (24h TTL, auto-extends on activity)
3. Subsequent requests → cookie authenticates silently (no more prompts)
4. SSE works automatically — `EventSource` sends same-origin cookies
5. 10 failed attempts per IP → 429 rate limit (15-minute decay)
### Security Requirements
- **Always set `CODEMAN_PASSWORD`** before exposing via tunnel — without it, anyone with the URL has full access
- Session cookies are `Secure` when using `--https` flag; through Cloudflare tunnel without `--https`, cookies are non-Secure but traffic is still encrypted end-to-end via Cloudflare
- `/api/hook-event` bypasses auth (localhost-only Claude Code hooks need unauthenticated access)
`./scripts/tunnel.sh start|stop|url`. **Always set `CODEMAN_PASSWORD`** before exposing via tunnel.
+143 -10
View File
@@ -28,18 +28,33 @@
curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash
```
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it. You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code) or [OpenCode](https://opencode.ai) (or both). After install:
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it.
**Install from a fork or specific branch:**
```bash
curl -fsSL https://raw.githubusercontent.com/<user>/Codeman/<branch>/install.sh | \
CODEMAN_REPO_URL=https://github.com/<user>/Codeman.git \
CODEMAN_BRANCH=<branch> bash
```
The installer supports these environment variables:
| Variable | Default | Description |
|----------|---------|-------------|
| `CODEMAN_REPO_URL` | upstream Codeman | Custom git repository URL |
| `CODEMAN_BRANCH` | `master` | Git branch to install |
| `CODEMAN_INSTALL_DIR` | `~/.codeman/app` | Custom install directory |
| `CODEMAN_SKIP_SYSTEMD` | `0` | Skip systemd service setup prompt |
| `CODEMAN_NODE_VERSION` | `22` | Node.js major version to install |
| `CODEMAN_NONINTERACTIVE` | `0` | Skip all prompts (for CI/automation) |
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code) or [OpenCode](https://opencode.ai) (or both). After install:
```bash
codeman web
# Open http://localhost:3000 — press Ctrl+Enter to start your first session
```
**Update to latest version:**
```bash
curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash -s update
```
<details>
<summary><strong>Run as a background service</strong></summary>
@@ -68,11 +83,23 @@ Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/e
## Mobile-Optimized Web UI
The most responsive AI coding agent experience on any phone. Full xterm.js terminal with local echo, swipe navigation, and a touch-optimized interface designed for real remote work.
The most responsive AI coding agent experience on any phone. Full xterm.js terminal with local echo, swipe navigation, and a touch-optimized interface designed for real remote work — not a desktop UI crammed onto a small screen.
<table>
<tr>
<td align="center" width="33%"><img src="docs/screenshots/mobile-landing-qr.png" alt="Mobile — landing page with QR auth" width="260"></td>
<td align="center" width="33%"><img src="docs/screenshots/mobile-session-idle.png" alt="Mobile — idle session with keyboard accessory" width="260"></td>
<td align="center" width="33%"><img src="docs/screenshots/mobile-session-active.png" alt="Mobile — active agent session" width="260"></td>
</tr>
<tr>
<td align="center"><em>Landing page with QR auth</em></td>
<td align="center"><em>Keyboard accessory bar</em></td>
<td align="center"><em>Agent working in real-time</em></td>
</tr>
</table>
<table>
<tr>
<td rowspan="8" width="320"><img src="docs/screenshots/mobile-keyboard-open.png" alt="Mobile — keyboard open" width="300"></td>
<th>Terminal Apps</th>
<th>Codeman Mobile</th>
</tr>
@@ -82,11 +109,22 @@ The most responsive AI coding agent experience on any phone. Full xterm.js termi
<tr><td>No notifications</td><td>Push alerts for approvals and idle</td></tr>
<tr><td>Manual reconnect</td><td>tmux persistence</td></tr>
<tr><td>No agent visibility</td><td>Background agents in real-time</td></tr>
<tr><td>Copy-paste slash commands</td><td>One-tap <code>/init</code></tr>
<tr><td>Copy-paste slash commands</td><td>One-tap <code>/init</code>, <code>/clear</code>, <code>/compact</code></td></tr>
<tr><td>Password typing on phone</td><td><b>QR code scan — instant auth</b></td></tr>
</table>
- **Swipe navigation** — left/right on the terminal to switch sessions (80px threshold, 300ms)
### Secure QR Code Authentication
Typing passwords on a phone keyboard is miserable. Codeman replaces it with **cryptographically secure single-use QR tokens** — scan the code displayed on your desktop and your phone is authenticated instantly.
Each QR encodes a URL containing a 6-character short code that maps to a 256-bit secret (`crypto.randomBytes(32)`) on the server. Tokens auto-rotate every **60 seconds**, are **atomically consumed on first scan** (replays always fail), and use **hash-based `Map.get()` lookup** that leaks nothing through response timing. The short code is an opaque pointer — the real secret never appears in browser history, `Referer` headers, or Cloudflare edge logs.
The security design addresses all 6 critical QR auth flaws identified in ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) (USENIX Security 2025, which found 47 of the top-100 websites vulnerable): single-use enforcement, short TTL, cryptographic randomness, server-side generation, real-time desktop notification on scan (QRLjacking detection), and IP + User-Agent session binding with manual revocation. Dual-layer rate limiting (per-IP + global) makes brute force infeasible across 62^6 = 56.8 billion possible codes. Full security analysis: [`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
### Touch-Optimized Interface
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard. Destructive commands (`/clear`, `/compact`) require a double-press to confirm — first tap arms the button, second tap executes — so you never fire one by accident on a bumpy commute
- **Swipe navigation** — left/right on the terminal to switch sessions (80px threshold, 300ms)
- **Smart keyboard handling** — toolbar and terminal shift up when keyboard opens (uses `visualViewport` API with 100px threshold for iOS address bar drift)
- **Safe area support** — respects iPhone notch and home indicator via `env(safe-area-inset-*)`
- **44px touch targets** — all buttons meet iOS Human Interface Guidelines minimum sizes
@@ -120,6 +158,10 @@ Watch background agents work in real-time. Codeman monitors agent activity and d
## Zero-Lag Input Overlay
<p align="center">
<img src="docs/images/zerolag-demo.gif" alt="Zerolag Demo — local echo vs server echo side-by-side" width="900">
</p>
When accessing your coding agent remotely (VPN, Tailscale, SSH tunnel), every keystroke normally takes 200-300ms to round-trip. Codeman implements a **Mosh-inspired local echo system** that makes typing feel instant regardless of latency.
A pixel-perfect DOM overlay inside xterm.js renders keystrokes at 0ms. Background forwarding silently sends every character to the PTY in 50ms debounced batches, so Tab completion, `Ctrl+R` history search, and all shell features work normally. When the server echo arrives 200-300ms later, the overlay seamlessly disappears and the real terminal text takes over — the transition is invisible.
@@ -239,6 +281,80 @@ loginctl enable-linger $USER
</details>
### QR Code Authentication
Typing a password on a phone keyboard is terrible. Codeman solves this with **ephemeral single-use QR tokens** — scan the code on your desktop, and your phone is instantly authenticated. No password prompt, no typing, no clipboard.
```
Desktop displays QR → Phone scans → GET /q/Xk9mQ3 → Server validates
→ Token atomically consumed (single-use) → Session cookie issued → 302 to /
→ Desktop notified: "Device authenticated via QR" → New QR auto-generated
```
Someone who only has the bare tunnel URL (without the QR) still hits the standard password prompt. The QR is the fast path; the password is the fallback.
#### How It Works
The server maintains a rotating pool of short-lived, single-use tokens. Each token consists of a 256-bit secret (`crypto.randomBytes(32)`) paired with a 6-character base62 short code used as an opaque lookup key in the URL path. The QR code encodes a URL like `https://abc-xyz.trycloudflare.com/q/Xk9mQ3` — the short code is a pointer, not the secret itself, so it never leaks through browser history, `Referer` headers, or Cloudflare edge logs.
Every **60 seconds**, the server automatically rotates to a fresh token. The previous token remains valid for a **90-second grace period** to handle the race where you scan right as rotation happens — after that, it's dead. Each token is **single-use**: the moment a phone successfully scans it, the token is atomically consumed and a new one is immediately generated for the desktop display.
#### Security Design
The design is informed by ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) (USENIX Security 2025), which found 47 of the top-100 websites vulnerable to QR auth attacks due to 6 critical design flaws across 42 CVEs. Codeman addresses all six:
| USENIX Flaw | Mitigation |
|-------------|------------|
| **Flaw-1**: Missing single-use enforcement | Token atomically consumed on first scan — replays always fail |
| **Flaw-2**: Long-lived tokens | 60s TTL with 90s grace, auto-rotation via timer |
| **Flaw-3**: Predictable token generation | `crypto.randomBytes(32)` — 256-bit entropy. Short codes use rejection sampling to eliminate modulo bias |
| **Flaw-4**: Client-side token generation | Server-side only — tokens never leave the server until embedded in the QR |
| **Flaw-5**: Missing status notification | Desktop toast: *"Device [IP] authenticated via QR (Safari). Not you? [Revoke]"* — real-time QRLjacking detection |
| **Flaw-6**: Inadequate session binding | IP + User-Agent stored for audit. Manual session revocation via API. HttpOnly + Secure + SameSite=lax cookies |
#### Timing-Safe Lookup
Short codes are stored in a `Map<shortCode, TokenRecord>`. Validation uses `Map.get()` — a hash-based O(1) lookup that reveals nothing about the target string through response timing. There is no character-by-character string comparison anywhere in the hot path, eliminating timing side-channel attacks entirely.
#### Rate Limiting (Dual Layer)
QR auth has its own rate limiting, completely independent from password auth:
- **Per-IP**: 10 failed QR attempts per IP trigger a 429 block (15-minute decay window) — separate counter from Basic Auth failures, so a fat-fingered password doesn't burn your QR budget
- **Global**: 30 QR attempts per minute across all IPs combined — defends against distributed brute force. With 62^6 = 56.8 billion possible short codes and only ~2 valid at any time, brute force is computationally infeasible regardless
#### QR Code Size Optimization
The URL is kept deliberately short (`/q/` path + 6-char code = ~53-56 total characters) to target **QR Version 4** (33x33 modules) instead of Version 5 (37x37). Smaller QR codes scan faster on budget phones — modern devices read Version 4 in 100-300ms. The `/q/` prefix saves 7 bytes compared to `/qr-auth/`, which alone is the difference between QR versions.
#### Desktop Experience
The QR display auto-refreshes every 60 seconds via SSE with the SVG embedded directly in the event payload (~2-5KB) — no extra HTTP fetch, sub-50ms refresh. A countdown timer shows time remaining. A "Regenerate" button instantly invalidates all existing tokens and creates a fresh one (useful if you suspect the QR was photographed).
When someone authenticates via QR, the desktop shows a notification toast with the device's IP and browser — if it wasn't you, one click revokes all sessions.
#### Threat Coverage
| Threat | Why it doesn't work |
|--------|-------------------|
| **QR screenshot shared** | Single-use: consumed on first scan. 60s TTL: expired before the attacker can act. Desktop notification alerts you immediately. |
| **Replay attack** | Atomic single-use consumption + 60s TTL. Old URLs always return 401. |
| **Cloudflare edge logs** | Short code is an opaque 6-char lookup key, not the real 256-bit token. Single-use means replaying from logs always fails. |
| **Brute force** | 56.8 billion combinations, ~2 valid at any time, dual-layer rate limiting blocks well before statistical feasibility. |
| **QRLjacking** | 60s rotation forces real-time relay. Desktop toast provides instant detection. Self-hosted single-user context makes phishing implausible. |
| **Timing attack** | Hash-based Map lookup — no string comparison timing leak. |
| **Session cookie theft** | HttpOnly + Secure + SameSite=lax + 24h TTL. Manual revocation at `POST /api/auth/revoke`. |
#### How It Compares
| Platform | Model | Comparison |
|----------|-------|------------|
| **Discord** | Long-lived token, no confirmation, [repeatedly exploited](https://owasp.org/www-community/attacks/Qrljacking) | Codeman: single-use + TTL + notification |
| **WhatsApp Web** | Phone confirms "Link device?", ~60s rotation | Comparable rotation; WhatsApp adds explicit confirmation (acceptable tradeoff for single-user) |
| **Signal** | Ephemeral public key, E2E encrypted channel | Stronger crypto, but [exploited by Russian state actors in 2025](https://cloud.google.com/blog/topics/threat-intelligence/russia-targeting-signal-messenger) via social engineering despite it |
> Full design rationale, security analysis, and implementation details: [`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
---
## SSH Alternative (`sc`)
@@ -376,6 +492,23 @@ See [CLAUDE.md](./CLAUDE.md) for full documentation.
---
## Codebase Quality
The codebase went through a comprehensive 7-phase refactoring that eliminated god objects, centralized configuration, and established modular architecture:
| Phase | What changed | Impact |
|-------|-------------|--------|
| **Performance** | Cached endpoints, SSE adaptive batching, buffer chunking | Sub-16ms terminal latency |
| **Route extraction** | `server.ts` split into 12 domain route modules + auth middleware + port interfaces | **−60%** server.ts LOC (6,736 → 2,697) |
| **Domain splitting** | `types.ts` → 14 domain files, `ralph-tracker` → 7 files, `respawn-controller` → 5 files, `session` → 6 files | No more god files |
| **Frontend modules** | `app.js` → 8 extracted modules (constants, mobile, voice, notifications, keyboard, API, subagent windows) | **−24%** app.js LOC (15.2K → 11.5K) |
| **Config consolidation** | ~70 scattered magic numbers → 9 domain-focused config files | Zero cross-file duplicates |
| **Test infrastructure** | Shared mock library, 12 route test files, consolidated MockSession | Testable route handlers via `app.inject()` |
Full details: [`docs/code-structure-findings.md`](docs/code-structure-findings.md)
---
## Published Packages
### [`xterm-zerolag-input`](https://www.npmjs.com/package/xterm-zerolag-input)
+977
View File
@@ -0,0 +1,977 @@
# Code Structure & Quality Findings
**Date**: 2026-02-28
**Scope**: Full codebase analysis across 5 dimensions: frontend, backend, TypeScript, testing, and utilities/config.
This document contains detailed findings for agent teams to write implementation plans and execute improvements. Each section includes severity, specific locations, and recommended fixes.
---
## Table of Contents
1. [Critical: server.ts God Object (6,736 LOC)](#1-critical-serverts-god-object)
2. [Critical: app.js Monolith (15,196 LOC)](#2-critical-appjs-monolith)
3. [Critical: CleanupManager Unused Despite Existing](#3-critical-cleanupmanager-unused)
4. [High: Duplicated Debounce/Timer Patterns](#4-high-duplicated-debouncetimer-patterns)
5. [High: Large Domain Files Need Splitting](#5-high-large-domain-files-need-splitting)
6. [High: types.ts God File (1,443 LOC)](#6-high-typests-god-file)
7. [High: Zod Schemas Duplicate TypeScript Types](#7-high-zod-schemas-duplicate-typescript-types)
8. [High: Test Coverage Gaps](#8-high-test-coverage-gaps)
9. [High: Duplicated Test Mocks](#9-high-duplicated-test-mocks)
10. [Medium: Hardcoded Magic Values](#10-medium-hardcoded-magic-values)
11. [Medium: Frontend Global State Monolith](#11-medium-frontend-global-state-monolith)
12. [Medium: Frontend Code Duplication](#12-medium-frontend-code-duplication)
13. [Medium: Inconsistent Logging](#13-medium-inconsistent-logging)
14. [Medium: Utils Barrel Export Gaps](#14-medium-utils-barrel-export-gaps)
15. [Medium: Non-Null Assertion Risks](#15-medium-non-null-assertion-risks)
16. [Low: Dead Utility Functions](#16-low-dead-utility-functions)
17. [Low: No Dependency Injection for File I/O](#17-low-no-dependency-injection-for-file-io)
18. [Scorecard & Prioritized Roadmap](#18-scorecard--prioritized-roadmap)
---
## 1. Critical: server.ts God Object
**File**: `src/web/server.ts` (6,736 lines)
**Severity**: CRITICAL
**Impact**: Hardest file to maintain, test, and extend. Imports 38 modules.
### Problem
The `WebServer` class handles everything: HTTP routing (~110 routes), authentication, SSE broadcasting, terminal data batching, state persistence, session lifecycle, respawn orchestration, file serving, tunnel management, plan orchestration, and subagent coordination.
**Key metrics**:
- 40+ private properties (Maps, timers, caches)
- 70+ methods
- `setupRoutes()` is 2,000+ LOC of inline route handlers
- Zero test coverage
### Current Structure (Bad)
```
WebServer class (6,736 LOC)
├── Auth session management (lines 469, 668-698)
├── SSE client management (lines 407-408, 5843-5880)
├── Terminal data batching (lines 414-416, 5909-5966)
├── Task update batching (line 426, 5995-6028)
├── State persistence batching (lines 429-430, 6028-6061)
├── Respawn lifecycle (lines 445-451, 5425-5534)
├── Session cleanup (lines 4769-4961)
├── Listener setup (lines 544-643)
└── setupRoutes() (lines 645+, 2000+ LOC)
├── /api/sessions/* (30+ routes inline)
├── /api/respawn/* (7 routes inline)
├── /api/subagents/* (7 routes inline)
├── /api/plan/* (5 routes inline)
├── /api/push/* (4 routes inline)
└── ... 60+ more inline
```
### Recommended Structure
```
src/web/
├── server.ts (~500 LOC - HTTP setup, route registration only)
├── routes/
│ ├── session-routes.ts (session CRUD, input, resize)
│ ├── respawn-routes.ts (respawn control endpoints)
│ ├── subagent-routes.ts (background agent tracking)
│ ├── plan-routes.ts (plan generation & management)
│ ├── push-routes.ts (web push subscriptions)
│ ├── mux-routes.ts (tmux management)
│ ├── case-routes.ts (case management)
│ ├── file-routes.ts (file browsing/serving)
│ └── system-routes.ts (status, stats, config, settings)
├── middleware/
│ ├── auth.ts (Basic Auth + session cookies)
│ └── error-handler.ts (centralized error responses)
└── services/
├── sse-manager.ts (SSE client + broadcast)
├── terminal-batcher.ts (60fps terminal batching)
└── session-lifecycle.ts (listener setup/teardown)
```
### Duplication in server.ts
**Error response pattern** repeated 189 times:
```typescript
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Session not found');
```
**Fix**: Extract `findSessionOrFail()` middleware:
```typescript
const findSessionOrFail = (sessionId: string) => {
const session = this.sessions.get(sessionId);
if (!session) throw new NotFoundError('Session not found');
return session;
};
```
**Event listener setup** copy-pasted for subagent watcher, image watcher, and team watcher (lines 544-643). Same attach/detach pattern duplicated 3 times.
---
## 2. Critical: app.js Monolith
**File**: `src/web/public/app.js` (15,196 lines)
**Severity**: CRITICAL
**Impact**: Untestable, hard to navigate, tightly coupled systems.
### Extractable Modules (by priority)
| Module | Lines | Current Location | Impact |
|--------|-------|------------------|--------|
| Mobile handlers (MobileDetection, KeyboardHandler, SwipeHandler) | ~300 | lines 168-620 | High |
| Voice input (DeepgramProvider, VoiceInput) | ~830 | lines 631-1471 | High |
| NotificationManager | ~450 | lines 2218-2663 | High |
| xterm-zerolag-input (inlined copy from packages/) | ~400 | lines 1756-2153 | High |
| KeyboardAccessoryBar | ~195 | lines 1480-1680 | Medium |
| FocusTrap | ~60 | lines 1690-1748 | Medium |
### CodemanApp Class (12,000+ LOC)
The main `CodemanApp` class starting at line 2665 has:
- **60+ Maps/Sets** in the constructor (lines 2667-2805)
- **18 Map instances** with complex cross-references (subagents, parents, teams, windows)
- **10+ monolithic methods** exceeding 100 lines each
**Largest methods**:
| Method | Lines | Size |
|--------|-------|------|
| `renderAppSettings()` | 14400-14700 | ~300 LOC |
| `selectSession()` | 6028-6250 | ~220 LOC |
| `batchTerminalWrite()` | 7482-7700 | ~200 LOC |
| `renderSessionTabs()` | 5814-6000 | ~180 LOC |
| `openSubagentWindow()` | 11927-12100 | ~170 LOC |
| `handleInit()` | 5183-5350 | ~170 LOC |
### Recommended Split
```
src/web/public/
├── app.js (~4000 LOC - core app, session mgmt, SSE)
├── mobile.js (~300 LOC - MobileDetection, KeyboardHandler, SwipeHandler)
├── voice.js (~830 LOC - DeepgramProvider, VoiceInput)
├── notifications.js (~450 LOC - NotificationManager)
├── keyboard-accessory.js (~200 LOC - KeyboardAccessoryBar)
├── api-client.js (~100 LOC - fetch wrapper with error handling)
└── config.js (~50 LOC - magic numbers, z-index layers)
```
---
## 3. Critical: CleanupManager Unused
**File**: `src/utils/cleanup-manager.ts` (320 lines)
**Severity**: CRITICAL
**Impact**: Memory leak risk. Well-designed utility exists but is never used. Every file manages cleanup manually.
### Current State
`CleanupManager` is exported from the utils barrel but has **0 instantiations** in production code. Instead, every file implements manual cleanup:
**respawn-controller.ts** (worst offender):
```typescript
// 11 timer properties, manually cleared in stop()
private stepTimer: NodeJS.Timeout | null = null;
private completionConfirmTimer: NodeJS.Timeout | null = null;
private noOutputTimer: NodeJS.Timeout | null = null;
// ... 8 more
stop() {
if (this.stepTimer) clearTimeout(this.stepTimer);
if (this.completionConfirmTimer) clearTimeout(this.completionConfirmTimer);
// ... 9 more clearTimeout/clearInterval calls
}
```
**Files that should use CleanupManager**:
| File | Timer/Listener Count | Current Cleanup |
|------|---------------------|-----------------|
| `respawn-controller.ts` | 11 timers + intervals | 11 manual clearTimeout/clearInterval |
| `web/server.ts` | 6+ timers, debounce map | Manual in stop(), some may leak |
| `state-store.ts` | 2 debounce timers | Manual clearTimeout |
| `push-store.ts` | 1 save timer | Manual clearTimeout |
| `subagent-watcher.ts` | debounce map + watchers | Manual clear + close |
| `ralph-tracker.ts` | 3 debounce timers | Manual clear |
| `bash-tool-parser.ts` | 1 debounce timer | Manual clear |
| `image-watcher.ts` | 1 debounce map | Manual clear |
### Fix
Migrate all timer management to use `CleanupManager`. Example for respawn-controller.ts:
```typescript
// Before: 11 fields + 11 clearTimeout calls
private stepTimer: NodeJS.Timeout | null = null;
// ...
// After: 1 field, auto-cleanup
private cleanup = new CleanupManager();
startStep() {
this.cleanup.setTimeout(() => { ... }, 5000, 'step');
}
stop() {
this.cleanup.dispose(); // Clears everything
}
```
---
## 4. High: Duplicated Debounce/Timer Patterns
**Severity**: HIGH
**Impact**: 8+ files implement debounce independently. Bug fixes need to be applied everywhere.
### Pattern Inventory
```typescript
// Pattern 1: Manual timer ref (used in 6 files)
private saveTimer: NodeJS.Timeout | null = null;
debouncedSave() {
if (this.saveTimer) clearTimeout(this.saveTimer);
this.saveTimer = setTimeout(() => this.save(), 500);
}
// Pattern 2: Timer Map (used in 3 files)
private fileDebouncers = new Map<string, NodeJS.Timeout>();
debounce(key: string) {
const existing = this.fileDebouncers.get(key);
if (existing) clearTimeout(existing);
this.fileDebouncers.set(key, setTimeout(() => { ... }, 100));
}
// Pattern 3: State flag (used in 2 files)
private isSaving = false;
```
### Locations
| File | Debounce Vars | Delay (ms) |
|------|---------------|------------|
| `state-store.ts` | `saveTimeout`, `ralphStateSaveTimeout` | 500 |
| `push-store.ts` | `saveTimer` | 500 |
| `web/server.ts` | `persistDebounceTimers` (Map) | 500 |
| `subagent-watcher.ts` | `fileDebouncers` (Map) | 100 |
| `ralph-tracker.ts` | 3 debounce timers | 50, 30000 |
| `bash-tool-parser.ts` | `EVENT_DEBOUNCE_MS` | 50 |
| `image-watcher.ts` | debounce map | 200 |
| `respawn-controller.ts` | 11 timer fields | various |
### Fix
Create a `Debouncer` utility:
```typescript
// src/utils/debouncer.ts
export class Debouncer {
private timer: NodeJS.Timeout | null = null;
constructor(private readonly delayMs: number) {}
run(fn: () => void): void {
if (this.timer) clearTimeout(this.timer);
this.timer = setTimeout(fn, this.delayMs);
}
cancel(): void {
if (this.timer) clearTimeout(this.timer);
this.timer = null;
}
}
// Usage:
private saveDeb = new Debouncer(500);
this.saveDeb.run(() => this.save());
// cleanup: this.saveDeb.cancel();
```
---
## 5. High: Large Domain Files Need Splitting
**Severity**: HIGH
**Impact**: Complex state machines spanning 3,000+ lines are hard to understand and test.
### ralph-tracker.ts (3,905 LOC)
**5 responsibilities mixed**:
1. Output Parsing (~900 LOC) - Line-by-line parsing, state extraction
2. Todo Management (~700 LOC) - Parsing, dedup, expiry
3. Plan Tracking (~800 LOC) - Enhanced plan tasks, checkpoints
4. Circuit Breaker (~400 LOC) - State machine for stuck detection
5. File Watching (~300 LOC) - Monitor external state files
**Recommended split**:
```
ralph-tracker.ts (core output parsing, ~1200 LOC)
ralph-todo-manager.ts (todo parsing + management, ~700 LOC)
ralph-plan-tracker.ts (plan tasks + checkpoints, ~800 LOC)
ralph-circuit-breaker.ts (circuit breaker logic, ~400 LOC)
```
### respawn-controller.ts (3,611 LOC)
**6 responsibilities mixed**:
1. State Machine (~1,000 LOC) - 6+ states, transitions
2. Idle Detection (~800 LOC) - 5 layers + multi-signal combining
3. AI Checkers (~600 LOC) - Idle + plan checkers integration
4. Health Scoring (~500 LOC) - Metrics, circuit breaker, scoring
5. Action Logging (~300 LOC) - Timeline, detection status
6. Stuck-State Detection (~250 LOC) - Timeout tracking
**Recommended split**:
```
respawn-controller.ts (state machine core, ~1000 LOC)
respawn-idle-detection.ts (all 5 idle detection layers, ~800 LOC)
respawn-health-scorer.ts (metrics & health scoring, ~500 LOC)
```
### session.ts (2,418 LOC)
**8 responsibilities mixed**:
1. PTY Management (~600 LOC)
2. Terminal I/O (~400 LOC)
3. Token Tracking (~200 LOC)
4. Task Tracking (~250 LOC)
5. Ralph Integration (~200 LOC)
6. Auto-Clear/Compact (~300 LOC)
7. Image Watching (~100 LOC)
8. CLI Detection (~150 LOC)
**Recommended split**:
```
session.ts (PTY + terminal I/O core, ~1000 LOC)
session-tracking.ts (token + task + Ralph, ~500 LOC)
session-auto-ops.ts (auto-clear/compact + image, ~300 LOC)
```
---
## 6. High: types.ts God File
**File**: `src/types.ts` (1,443 lines, 72 exported definitions)
**Severity**: HIGH
**Impact**: Every file imports from types.ts. Hard to find relevant types.
### Current Contents
- 46 interfaces
- 25 types
- 1 enum (ApiErrorCode)
- 9 factory functions (createInitialState, etc.)
### Recommended Split
```
src/types/
├── index.ts (barrel export - transparent migration)
├── session.ts (SessionState, SessionConfig, SessionMode, SessionColor)
├── task.ts (TaskState, TaskDefinition, TaskStatus)
├── respawn.ts (RespawnConfig, RespawnState, CircuitBreakerStatus)
├── ralph.ts (RalphLoopState, RalphTrackerState, RalphTodoItem)
├── api.ts (ApiResponse, ApiErrorCode, HookEventType, all route types)
├── lifecycle.ts (LifecycleEventType, LifecycleEntry)
└── common.ts (Disposable, BufferConfig, CleanupResourceType)
```
The barrel export makes this a transparent refactor - existing `import from './types'` continues to work.
---
## 7. High: Zod Schemas Duplicate TypeScript Types
**File**: `src/web/schemas.ts` (508 lines)
**Severity**: HIGH
**Impact**: When a type changes, the Zod schema must be manually updated too. Source of bugs.
### Problem
Zod schemas manually duplicate TypeScript interfaces. **Zero `z.infer` usage found.**
```typescript
// types.ts (manual interface)
export interface CreateSessionRequest {
workingDir?: string;
mode?: SessionMode;
name?: string;
}
// schemas.ts (manual Zod schema - duplicated!)
export const CreateSessionSchema = z.object({
workingDir: safePathSchema.optional(),
mode: z.enum(['claude', 'shell', 'opencode']).optional(),
name: z.string().max(100).optional(),
});
```
### Fix
Use `z.infer` to derive TypeScript types from Zod schemas (single source of truth):
```typescript
// schemas.ts
export const CreateSessionSchema = z.object({
workingDir: safePathSchema.optional(),
mode: z.enum(['claude', 'shell', 'opencode']).optional(),
name: z.string().max(100).optional(),
});
// types.ts (auto-derived)
export type CreateSessionRequest = z.infer<typeof CreateSessionSchema>;
```
**Affected schemas** (~10):
- CreateSessionSchema
- RunPromptSchema
- ResizeSchema
- CreateCaseSchema
- QuickStartSchema
- HookEventSchema
- RespawnConfigSchema
- ConfigUpdateSchema
- SettingsUpdateSchema
---
## 8. High: Test Coverage Gaps
**Severity**: HIGH
**Impact**: Critical code paths untested. Regressions go unnoticed.
### Untested Source Files
| File | Lines | Risk |
|------|-------|------|
| `src/web/server.ts` | 6,736 | CRITICAL - Core REST API, 280+ routes |
| `src/plan-orchestrator.ts` | ~500 | HIGH - Multi-agent plan generation |
| `src/tunnel-manager.ts` | ~200 | MEDIUM - Cloudflare tunnel |
| `src/session-lifecycle-log.ts` | ~150 | MEDIUM - JSONL audit log |
| `src/ai-plan-checker.ts` | ~300 | MEDIUM - Plan completion detection |
| `src/templates/claude-md.ts` | ~200 | LOW - CLAUDE.md generation |
| `src/utils/claude-cli-resolver.ts` | ~100 | LOW - CLI path resolution |
| `src/utils/opencode-cli-resolver.ts` | ~100 | LOW - OpenCode CLI support |
| `src/utils/regex-patterns.ts` | ~100 | LOW - Used everywhere! |
| `src/utils/token-validation.ts` | ~50 | LOW - Token counting |
### Test Quality Issues
**10 "not.toThrow()" tests without behavior verification**:
```typescript
// BAD: Only checks it doesn't crash
expect(() => tracker.processMessage(null)).not.toThrow();
// GOOD: Also verify defensive behavior
expect(() => tracker.processMessage(null)).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
```
Locations:
- `task-tracker.test.ts` - 5 instances
- `image-watcher.test.ts` - 1 instance
- `task-queue.test.ts` - 1 instance
- Others scattered
---
## 9. High: Duplicated Test Mocks
**Severity**: HIGH
**Impact**: Mock changes need updating in 4 places. Inconsistent mock behavior.
### MockSession Defined 4 Times
| File | Usage |
|------|-------|
| `test/respawn-controller.test.ts` | Full mock with event emitter |
| `test/session-manager.test.ts` | Simpler mock |
| `test/respawn-team-awareness.test.ts` | Copy of respawn-controller mock |
| `test/respawn-test-utils.ts` | **Comprehensive mock - UNUSED!** |
### MockStateStore Defined 2 Times
| File | Usage |
|------|-------|
| `test/session-manager.test.ts` | Basic mock |
| `test/ralph-loop.test.ts` | Separate implementation |
### Unused Test Utilities
`test/respawn-test-utils.ts` exports these utilities that **no test file imports**:
- `createTimeController()` - Abstraction over vitest fake timers
- `MockAiIdleChecker` - Fully mocked AI idle checker
- `MockAiPlanChecker` - Fully mocked plan checker
- Factory functions for pre-configured controllers
### Fix
Create `test/mocks/` directory:
```
test/
├── mocks/
│ ├── mock-session.ts (single MockSession, used everywhere)
│ ├── mock-state-store.ts (single MockStateStore)
│ └── index.ts (barrel export)
├── utils/
│ └── time-controller.ts (from respawn-test-utils.ts)
└── ... test files
```
---
## 10. Medium: Hardcoded Magic Values
**Severity**: MEDIUM
**Impact**: Hard to tune, inconsistent when same value appears in multiple places.
### Already Centralized (Good)
- `src/config/buffer-limits.ts` - All buffer sizes
- `src/config/map-limits.ts` - All collection limits
### NOT Centralized (40+ values scattered)
**In server.ts** (lines 145-194):
```typescript
const TASK_UPDATE_BATCH_INTERVAL = 100;
const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
const SESSIONS_LIST_CACHE_TTL = 1000;
const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
const SSE_HEALTH_CHECK_INTERVAL = 30 * 1000;
const MAX_TERMINAL_COLS = 500;
const MAX_TERMINAL_ROWS = 200;
const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
const MAX_AUTH_SESSIONS = 100;
const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
const STATS_COLLECTION_INTERVAL_MS = 2000;
const MAX_INPUT_LENGTH = 64 * 1024;
```
**In hooks-config.ts**: `timeout: 10000` hardcoded 6 times.
**In respawn-controller.ts** (lines 538-565): 10 timing constants.
**In utils**: `EXEC_TIMEOUT_MS = 5000` duplicated in both `claude-cli-resolver.ts` and `opencode-cli-resolver.ts`.
**In app.js**:
```javascript
// line 27: 600000 - stuck detection threshold
// line 24: 5000 - default scrollback
// lines 34-35: 128*1024, 256*1024 - chunk sizes
// lines 152-155: 150, 100 - keyboard detection thresholds
// lines 573-575: 80, 300, 100 - swipe detection params
```
### Fix
Create additional config files:
```
src/config/
├── buffer-limits.ts (existing)
├── map-limits.ts (existing)
├── server-config.ts (NEW - web server intervals, auth, caching)
├── timing-config.ts (NEW - debounce delays, check intervals)
└── terminal-config.ts (NEW - max cols/rows, batch intervals)
```
---
## 11. Medium: Frontend Global State Monolith
**Severity**: MEDIUM
**Impact**: All state in single CodemanApp class. Tight coupling between unrelated systems.
### 60+ State Variables in CodemanApp Constructor (lines 2667-2805)
```javascript
this.sessions = new Map(); // Session data
this.subagents = new Map(); // Agent tracking
this.subagentActivity = new Map(); // Tool call tracking
this.subagentToolResults = new Map(); // Result caching
this.subagentParentMap = new Map(); // Agent-to-session mapping
this.teams = new Map(); // Team tracking
this.teamTasks = new Map(); // Team task state
this.planSubagents = new Map(); // Plan agent tracking
this.pendingWrites = []; // Terminal write queue
this.terminalBufferCache = new Map(); // Buffer caching (unbounded!)
this.projectInsights = new Map(); // Bash tool insights
// ... 40+ more
```
### Problems
1. **18 Map instances** with complex cross-references (no garbage collection strategy)
2. **No domain separation**: Session, subagent, notification, UI, and network state mixed
3. **Implicit dependencies**: `selectSession()` requires 5+ Maps to be in consistent state
4. **`terminalBufferCache`** has no max size - can grow unbounded with many sessions
### Recommended Domain Split
```javascript
// Instead of 60+ flat properties:
class SessionState {
sessions = new Map();
sessionOrder = [];
terminalBuffers = new Map();
tabAlerts = new Map();
}
class SubagentState {
subagents = new Map();
activity = new Map();
parentMap = new Map();
windows = new Map();
minimized = new Map();
}
class TeamState {
teams = new Map();
tasks = new Map();
teammates = new Map();
}
class UIState {
activeSessionId = null;
draggedTabId = null;
isLoadingBuffer = false;
}
```
---
## 12. Medium: Frontend Code Duplication
**Severity**: MEDIUM
**Impact**: Repeated patterns increase maintenance burden and inconsistency risk.
### Duplicated Patterns
**API fetch calls** (~50 instances):
```javascript
// Repeated everywhere:
fetch(`/api/sessions/${sessionId}/...`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({...})
}).catch(() => {})
```
**Fix**: Extract `ApiClient` class.
**`innerHTML` usage** (104 instances):
- Mix of template strings, createElement chains, and direct innerHTML
- Some with manual XSS escaping (`text.replace(/</g, '&lt;')`), some without
- No consistent DOM creation pattern
**`typeof app !== 'undefined'` checks** (20+ instances):
- Lines 458, 467, 481, 614, 617, 1549, etc.
- **Fix**: Ensure `app` is always defined as global singleton.
**Element visibility toggling** (212+ occurrences):
```javascript
element.classList.add('active')
element.classList.remove('active')
```
**Fix**: Create `toggleClass(el, className, condition)` utility.
### Event Listener Issues
- **152 `addEventListener` calls** with fragile cleanup
- **Mix of inline (`onclick="app.method()"`) and addEventListener** - hard to track
- **Element cache (`_elemCache`) never invalidated** if DOM elements are recreated (line 2808)
- **Tab drag-and-drop listeners** may not clean up if user switches tabs mid-drag
---
## 13. Medium: Inconsistent Logging
**Severity**: MEDIUM
**Impact**: Hard to debug in production. Can't filter by severity or component.
### Current State
- **345 console calls** across source files
- **No structured logging** - all `console.log/error` directly
- **No log levels** (DEBUG, INFO, WARN, ERROR)
### Inconsistent Prefixes
```typescript
// Some files use brackets:
console.log('[Session] Starting interactive...');
console.log('[RalphLoop] Task assigned...');
console.log('[TunnelManager] Tunnel started');
// Others use no prefix:
console.error('Failed to spawn PTY:', err);
console.log('Server listening on port', port);
```
### Positive: CleanupManager Has Debug Mode
`src/utils/cleanup-manager.ts` has a `debugMode` flag for conditional debug logging - good pattern not replicated elsewhere.
### Fix
Either:
1. Enforce consistent `[ComponentName]` prefixes via lint rule
2. Create lightweight logger abstraction (not a heavy framework)
---
## 14. Medium: Utils Barrel Export Gaps
**File**: `src/utils/index.ts`
**Severity**: MEDIUM
**Impact**: Forces deep imports, unclear public API.
### Missing Exports
These functions are defined but NOT exported from the barrel:
- `createAnsiPatternFull()` and `createAnsiPatternSimple()` (factory functions from `regex-patterns.ts`)
- `SAFE_PATH_PATTERN` (from `regex-patterns.ts`)
- `validateTokenCounts()` and `validateTokensAndCost()` (from `token-validation.ts`)
- `isSimilar()`, `isSimilarByDistance()`, `levenshteinDistance()`, `normalizePhrase()` (from `string-similarity.ts` - though some are dead code, see finding #16)
### Deep Import Anti-Pattern (16 instances)
Some files bypass the barrel unnecessarily:
```typescript
// Could use barrel:
import { BufferAccumulator } from './utils/buffer-accumulator.js';
import { LRUMap } from './utils/lru-map.js';
// Must deep import (not in barrel):
import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';
```
### Fix
Add missing exports to `src/utils/index.ts` and update import sites.
---
## 15. Medium: Non-Null Assertion Risks
**Severity**: MEDIUM
**Impact**: Runtime crashes if assumptions violated. 37 instances found.
### Distribution
| File | Count | Risk Level |
|------|-------|------------|
| `src/web/server.ts` | 10 | Low (auth flow verified) |
| `src/session.ts` | 6 | **High** (mux/terminal refs) |
| `src/respawn-controller.ts` | 4 | Low (config validated) |
| `src/lru-map.ts` | 3 | Low (checked lookups) |
| `src/subagent-watcher.ts` | 2 | Low (pending tool calls) |
| Others | 12 | Low |
### High-Risk Examples (session.ts)
```typescript
// Line 915 - _mux could be null if startInteractive called during cleanup
`[Session] Starting interactive (with ${this._mux!.backend})`
// Line 954 - _muxSession could be null in race condition
this._muxSession!.muxName
```
### Fix
Add null guards before assertions, or document invariants:
```typescript
// Before:
this._mux!.backend
// After:
if (!this._mux) throw new Error('Invariant: _mux must be initialized before startInteractive');
this._mux.backend
```
### Positive Notes
- **0 instances of `as any`**
- **0 instances of `@ts-ignore` or `@ts-expect-error`**
- TypeScript overall score: 8.5/10
---
## 16. Low: Dead Utility Functions
**File**: `src/utils/string-similarity.ts`
**Severity**: LOW
**Impact**: Code clutter, confusion about what's actually used.
### Unused Functions
These are defined and exported but **never imported anywhere**:
- `isSimilar(a, b, threshold)` - similarity check with threshold
- `isSimilarByDistance(a, b, maxDistance)` - Levenshtein-based check
- `levenshteinDistance(a, b)` - raw edit distance
- `normalizePhrase(phrase)` - phrase normalization
### Actually Used
Only these are imported from the barrel:
- `stringSimilarity()` - used in ralph-tracker.ts
- `fuzzyPhraseMatch()` - used in ralph-tracker.ts
- `todoContentHash()` - used in ralph-tracker.ts
### Fix
Delete unused functions or mark as `@internal` if kept for future use.
---
## 17. Low: No Dependency Injection for File I/O
**Severity**: LOW (practical impact limited at current scale)
**Impact**: Can't mock filesystem for unit tests. 68+ hard-coded filesystem calls.
### Examples
```typescript
// state-store.ts - directly imports and uses fs
import { readFileSync, writeFileSync, existsSync, mkdirSync } from 'node:fs';
// push-store.ts - hard-coded paths
const KEYS_FILE = join(DATA_DIR, 'push-keys.json');
const SUBS_FILE = join(DATA_DIR, 'push-subscriptions.json');
// ai-checker-base.ts - direct execSync
execSync(`tmux kill-session -t "${this.checkMuxName}"`, { timeout: 3000 });
```
### Why This Is Lower Priority
- The codebase uses integration tests (spawning real processes/tmux sessions) rather than unit tests
- Most filesystem operations are in infrastructure code, not business logic
- Adding DI would be a large refactor with limited near-term benefit
---
## 18. Scorecard & Prioritized Roadmap
### Overall Scores (Post-Implementation)
| Category | Before | After | Notes |
|----------|--------|-------|-------|
| TypeScript Safety | 8.5/10 | 9/10 | 0 `any`, 0 `@ts-ignore`, Zod `z.infer` eliminates type drift |
| Error Handling | 8/10 | 8/10 | Unchanged — already strong |
| Async/Promise Safety | 9.5/10 | 9.5/10 | Unchanged — already strong |
| Resource Cleanup | 7/10 | 8/10 | CleanupManager adopted in server.ts, subagent-watcher, bash-tool-parser; Debouncer in 6 files. **Gaps**: respawn-controller (10+ manual timers) and ralph-tracker (2 manual timers) not migrated |
| Module Organization | 5/10 | 8/10 | Routes extracted (12 modules), types split (14 domain files), domain files split (ralph: 7, respawn: 5, session: 6) |
| Test Coverage | 6/10 | 7.5/10 | Shared mock infrastructure, 12 route test files, MockSession/MockStateStore consolidated |
| Config Centralization | 6/10 | 9/10 | 9 config files, ~65 constants centralized, 0 cross-file duplicates |
| Frontend Architecture | 4/10 | 7/10 | 8 extracted modules (3,453 LOC), app.js reduced 24% (15.2K → 11.5K), xterm-zerolag-input vendor build |
| Code Duplication | 5/10 | 8/10 | Debouncer utility, shared test mocks, barrel exports, config consolidation |
### Implementation Phases
**Phase 1 - Quick Wins (1-2 days)** ✅ COMPLETE
1. ✅ Export missing functions from utils barrel (~30 min) — `createAnsiPatternFull`, `createAnsiPatternSimple`, `SAFE_PATH_PATTERN`, `validateTokenCounts`, `validateTokensAndCost` all now exported from `src/utils/index.ts`
2. ✅ Delete dead utility functions (~15 min) — `isSimilar()` removed from `string-similarity.ts`; `levenshteinDistance()`, `isSimilarByDistance()`, `normalizePhrase()` made private (used internally by `fuzzyPhraseMatch`/`stringSimilarity`)
3. ✅ Consolidate duplicated `EXEC_TIMEOUT_MS` constant (~15 min) — Created `src/config/exec-timeout.ts` as single source of truth; `claude-cli-resolver.ts`, `opencode-cli-resolver.ts`, and `tmux-manager.ts` all import from it
4. ✅ Add `z.infer` to Zod schemas (~2 hours) — `src/web/schemas.ts` now has 36 `z.infer` type exports (lines 512-547) covering all schemas
5. ✅ Fix 10 weak "not.toThrow()" tests (~1 hour) — All `not.toThrow()` calls now have behavior assertions: `task-tracker.test.ts` (6 instances all followed by state checks), `image-watcher.test.ts` (1 instance followed by length check), `session-manager.test.ts` (1 instance followed by count check)
**Phase 2 - CleanupManager & Debounce (2-3 days)** ✅ COMPLETE
1. ✅ Create `Debouncer` utility class (~1 hour) — Created `src/utils/debouncer.ts` with `Debouncer` and `KeyedDebouncer` classes; exported from `src/utils/index.ts`
2. ✅ Migrate all 8 files from manual debounce to Debouncer — `state-store.ts` (2 Debouncers), `push-store.ts` (1 Debouncer), `bash-tool-parser.ts` (1 Debouncer), `image-watcher.ts` (1 KeyedDebouncer), `subagent-watcher.ts` (2 KeyedDebouncers), `server.ts` (1 KeyedDebouncer for persist timers), `ralph-tracker.ts` (2 Debouncers replacing 4 manual fields: `_todoUpdateTimer`, `_loopUpdateTimer`, `_todoUpdatePending`, `_loopUpdatePending`)
3. ✅ Migrate respawn-controller to CleanupManager — 10 manual timer fields replaced with single `CleanupManager` instance + `timerIds` Map. `startTrackedTimer()`/`cancelTrackedTimer()` preserved as wrappers for UI countdown display and timer events. `clearTimers()` uses dispose-and-recreate pattern for state transitions.
4. ✅ Migrate server.ts timer cleanup to CleanupManager (~2 hours) — `private cleanup = new CleanupManager()` present; terminal batch timers and pending respawn starts left as manual Maps (complex lifecycle)
5. ✅ Migrate remaining files — `bash-tool-parser.ts` (CleanupManager ✅), `subagent-watcher.ts` (CleanupManager ✅), `ralph-tracker.ts` (Debouncer ✅)
**Phase 3 - server.ts Route Extraction (3-4 days)** ✅ COMPLETE
1. ✅ Created `src/web/routes/` with 12 domain route modules + index barrel (4,090 LOC total): session (909), system (768), ralph (533), plan (459), respawn (315), case, file, hook-event, mux, push, scheduled, team
2. ✅ Created `src/web/middleware/auth.ts` (193 LOC) — Basic Auth, session cookies, rate limiting, security headers, CORS
3. ✅ Created `src/web/ports/` with 7 typed port interfaces (142 LOC) — SessionPort, EventPort, RespawnPort, ConfigPort, InfraPort, AuthPort; routes declare dependencies via intersection types
4. ✅ Created `src/web/route-helpers.ts` (154 LOC) — `findSessionOrFail()`, `formatUptime()`, `sanitizeHookData()`, `autoConfigureRalph()`
5. ✅ Reduced `server.ts` from 6,736 → 2,697 LOC (60% reduction). Remaining LOC is justified infrastructure: session lifecycle, SSE broadcast engine, terminal batching, respawn integration, resource cleanup
**Phase 4 - Domain File Splitting (2-3 days)** ✅ COMPLETE
1. ✅ Split `types.ts` into `src/types/` directory — 14 domain files (1,469 LOC total): common, session, task, app-state, respawn, ralph, api, lifecycle, run-summary, tools, teams, push, plan + index barrel. Original `types.ts` is now a 1-line re-export
2. ✅ Split `ralph-tracker.ts` into 7 files (exceeded plan of 4) — ralph-tracker (2,391), ralph-plan-tracker (477), ralph-status-parser (552), ralph-fix-plan-watcher (366), ralph-stall-detector (166), ralph-config (153), ralph-loop (522)
3. ✅ Split `respawn-controller.ts` into 5 files (exceeded plan of 3) — respawn-controller (3,228), respawn-health (229), respawn-metrics (229), respawn-patterns (131), respawn-adaptive-timing (134)
4. ✅ Split `session.ts` into 6 files (exceeded plan of 3) — session (2,168), session-manager (298), session-auto-ops (284), session-cli-builder (132), session-task-cache (101), session-lifecycle-log (114)
**Phase 5 - Frontend Modularization (3-4 days)** ✅ COMPLETE
1. ✅ Extracted `constants.js` (238 LOC) — shared constants, timing values, Z-index layers, `escapeHtml()`, `extractSyncSegments()`
2. ✅ Extracted `mobile-handlers.js` (449 LOC) — `MobileDetection`, `KeyboardHandler`, `SwipeHandler`
3. ✅ Extracted `voice-input.js` (853 LOC) — `DeepgramProvider`, `VoiceInput`
4. ✅ Extracted `notification-manager.js` (445 LOC) — `NotificationManager` class (5-layer system)
5. ✅ Extracted `keyboard-accessory.js` (279 LOC) — `KeyboardAccessoryBar`, `FocusTrap`
6. ✅ Extracted `api-client.js` (70 LOC) — `_api()`, `_apiJson()`, `_apiPost()`, `_apiPut()`
7. ✅ Extracted `subagent-windows.js` (1,119 LOC) — 13 subagent window methods
8. ✅ Removed inlined xterm-zerolag-input copy → built to `vendor/xterm-zerolag-input.js` from `packages/xterm-zerolag-input/`
9. ✅ Reduced `app.js` from ~15,200 → 11,473 LOC (24% reduction). All scripts loaded in correct dependency order in `index.html`
**Phase 6 - Config Consolidation (1 day)** ✅ COMPLETE
1. ✅ Created 6 new domain-focused config files (better than plan's 2 generic files): `server-timing.ts` (13 constants), `auth-config.ts` (5 constants), `tunnel-config.ts` (8 constants), `terminal-limits.ts` (4 constants), `ai-defaults.ts` (3 constants), `team-config.ts` (3 constants)
2. ✅ Total: 9 config files in `src/config/`, ~65 constants centralized
3. ✅ Eliminated all cross-file duplicates: `STATS_COLLECTION_INTERVAL_MS` (was in 2 files), `timeout: 10000` (was 6× inline in hooks-config.ts → `HOOK_TIMEOUT_MS`), AI model string (was in 5 files → `AI_CHECK_MODEL`), `MAX_TRACKED_AGENTS` (was shadowed in subagent-watcher.ts)
4. ✅ CLAUDE.md updated with config files table, import conventions, resource limits references
**Phase 7 - Test Infrastructure (2-3 days)** ✅ COMPLETE
1. ✅ Created `test/mocks/` directory with 5 files (541 LOC): `mock-session.ts` (312), `mock-state-store.ts` (60), `mock-route-context.ts` (121), `test-helpers.ts` (37), `index.ts` (11 — barrel export)
2. ✅ Consolidated MockSession into single shared definition — no duplicate class definitions remain (2 `vi.mock()`-based copies intentionally left in session-manager.test.ts and ralph-loop.test.ts)
3. ✅ `respawn-test-utils.ts` converted to backward-compatibility shim — re-exports from `test/mocks/`, retains respawn-specific utilities (MockAiIdleChecker, TimeController, etc.)
4. ✅ Created initial 3 route test files with 58 total tests: `session-routes.test.ts` (34 tests), `respawn-routes.test.ts` (13 tests), `system-routes.test.ts` (11 tests). Route test harness uses `app.inject()` — no real ports needed
5. ✅ All 12 route modules now have dedicated test files in `test/routes/`: session, respawn, system, ralph, plan, push, team, mux, file, scheduled, hook-event, case
---
## Appendix: File Size Inventory (Post-Implementation)
### Before vs After
| File | Before | After | Change |
|------|--------|-------|--------|
| `src/web/server.ts` | 6,736 | 2,697 | **−60%** (routes, auth, ports extracted) |
| `src/web/public/app.js` | 15,196 | 11,473 | **−24%** (8 modules extracted) |
| `src/ralph-tracker.ts` | 3,905 | 2,391 | **−39%** (6 companion files extracted) |
| `src/respawn-controller.ts` | 3,611 | 3,228 | **−11%** (4 companion files extracted) |
| `src/session.ts` | 2,418 | 2,168 | **−10%** (5 companion files extracted) |
| `src/types.ts` | 1,443 | 1 | **−99%** (14 domain files in `src/types/`) |
### New Infrastructure Created
| Directory | Files | Total LOC | Purpose |
|-----------|-------|-----------|---------|
| `src/web/routes/` | 13 | 4,090 | Domain route modules |
| `src/web/ports/` | 7 | 142 | Port interfaces for DI |
| `src/web/middleware/` | 1 | 193 | Auth middleware |
| `src/types/` | 14 | 1,469 | Domain type files |
| `src/config/` | 9 | ~450 | Centralized config |
| `test/mocks/` | 5 | 541 | Shared test mocks |
| `test/routes/` | 4 | ~500 | Route handler tests |
### Extracted Frontend Modules
| Module | Lines | Purpose |
|--------|-------|---------|
| `subagent-windows.js` | 1,119 | Subagent window management |
| `voice-input.js` | 853 | DeepgramProvider, VoiceInput |
| `mobile-handlers.js` | 449 | MobileDetection, KeyboardHandler, SwipeHandler |
| `notification-manager.js` | 445 | 5-layer notification system |
| `keyboard-accessory.js` | 279 | KeyboardAccessoryBar, FocusTrap |
| `constants.js` | 238 | Shared constants, timing, Z-index |
| `api-client.js` | 70 | API fetch wrapper |
### What's Working Well
These patterns should be **preserved, not refactored**:
- Clean one-way dependency graph (no circular deps)
- EventEmitter-based decoupling between domain models
- Proper `import type` usage (19 files, consistent)
- Utility type adoption (101 instances of Record, Partial, Omit, etc.)
- `assertNever()` for exhaustive switch checking
- `StaleExpirationMap` and `LRUMap` for bounded collections
- State persistence circuit breaker pattern
- TypeScript strict mode with all safety flags enabled
- `CleanupManager` for centralized timer/watcher disposal
- `Debouncer`/`KeyedDebouncer` for consistent debounce patterns
- Port interfaces for route module dependency injection
- `Object.assign(CodemanApp.prototype, ...)` for frontend module composition
Binary file not shown.
Binary file not shown.

After

Width:  |  Height:  |  Size: 806 KiB

+423
View File
@@ -0,0 +1,423 @@
# Performance Analysis & Optimization Opportunities
**Date**: 2026-03-07
**Scope**: Full-stack performance analysis — backend PTY handling, SSE broadcasting, frontend terminal rendering, local echo overlay, DOM updates, config/scaling limits.
**Constraint**: All recommendations preserve existing functionality including local echo, backpressure, anti-flicker pipeline, and mobile support.
---
## Executive Summary
The codebase is already well-optimized in critical paths. The multi-layer backpressure system, adaptive terminal batching, DEC 2026 sync markers, and incremental state serialization are strong. The main opportunities are in **reducing unnecessary work** (SSE filtering, DOM rebuilds, lazy terminal init) rather than algorithmic changes.
**Top 5 high-impact opportunities:**
| # | Optimization | Impact | Risk | Effort |
|---|-------------|--------|------|--------|
| 1 | Session-scoped SSE subscriptions | Bandwidth -60-80%, CPU -40% | Medium | Medium |
| 2 | Lazy xterm.js for minimized subagent windows | Memory -3.5MB at 50 agents | Low | Low |
| 3 | Targeted badge update (skip full tab rebuild) | Eliminates O(n) reflow on badge change | Low | Low |
| 4 | Conditional SSE padding (tunnel-only, terminal-only) | Bandwidth -70% when tunneled | Low | Low |
| 5 | Canvas renderer on mobile | GPU pressure reduction, battery savings | Low | Low |
---
## 1. SSE Broadcasting
### Current State
- **92 event types** broadcast to all connected clients (max 100)
- Single `JSON.stringify()` per event, shared across all clients (efficient)
- **No per-client filtering** — every client receives every event regardless of which session they're viewing
- 8KB padding appended to **every** event when tunnel is active (forces Cloudflare proxy flush)
- Backpressure: clients marked as backpressured if `reply.raw.write()` returns false; recovery via `session:needsRefresh`
### Bottlenecks
**B1: No session-scoped SSE subscriptions** (`server.ts:1986`)
- Client viewing session A still receives all events for sessions B through T
- With 20 active sessions, ~95% of terminal events are irrelevant to any given client
- Cost: wasted bandwidth, CPU for JSON parsing, and event handler dispatch on client
**B2: Unconditional 8KB padding** (`server.ts:1977`)
- Every event gets 8KB comment padding when tunnel is active
- A `task:updated` event (~200 bytes payload) becomes ~8.2KB
- High-frequency events like `session:terminal` need the padding; low-frequency events like `session:created` don't
### Recommendations
**R1: Session-scoped SSE subscriptions** (High impact)
- Add `?sessions=id1,id2` query param to `/api/events` SSE endpoint
- Server filters events by session ID before broadcasting
- Client subscribes to active session + "global" events (session lifecycle, system)
- Re-subscribes on tab switch (or subscribe to all with client-side filter as fallback)
- **Savings**: ~80% bandwidth reduction for single-session viewers; ~60% for multi-session dashboards
**R2: Tiered SSE padding** (Medium impact)
- Only pad `session:terminal` events and SSE heartbeats (the two that need proxy flush)
- Skip padding for low-frequency structural events (`session:created`, `task:updated`, etc.)
- **Savings**: ~70% padding overhead reduction; terminal events already large enough to flush
---
## 2. Terminal Rendering
### Current State (Well-Optimized)
- **6-layer anti-flicker pipeline**: Server batching (adaptive 16-50ms) → DEC 2026 sync wrap → single JSON serialize → client rAF batching → sync segment parser → chunked buffer loading (32KB/frame)
- **64KB/frame write budget** with DEC 2026 sync-segment awareness (prevents 141KB single-frame freezes)
- **3-layer backpressure**: SSE cap (128KB queued → drop + refresh), frame budget (64KB/frame), chunked restore (32KB/frame)
- WebGL renderer enabled by default with canvas fallback on context loss
- Typical latency: 16-32ms; worst case: ~115ms (50ms server batch + 50ms sync wait + 16ms rAF)
### Bottlenecks
**B3: WebGL on mobile** (`app.js:627-637`)
- Mobile GPUs are weaker; WebGL context loss more likely on low-end devices
- Canvas renderer is sufficient for mobile (typically 1 session, smaller viewport)
**B4: Static scrollback for all sessions** (`app.js:572`)
- Default 5000 lines scrollback for all sessions regardless of activity level
- Heavy output sessions (build logs, test runners) accumulate large scroll buffers
**B5: No addon lazy loading**
- FitAddon, Unicode11Addon, and WebGLAddon all loaded at terminal init
- Unicode11Addon only needed for CJK content; WebGLAddon is large
### Recommendations
**R3: Force canvas renderer on mobile** (Low risk)
- Detect `MobileDetection.isMobile()` and skip WebGL addon loading
- Reduces GPU memory pressure, prevents context loss crashes
- Mobile typically has 1-2 sessions — canvas performance is more than adequate
**R4: Dynamic scrollback based on session activity** (Low risk)
- Active sessions (working state): 5000 lines (current default)
- Inactive/idle sessions: reduce to 2000 lines
- Restore on session select (fetch from server buffer)
- **Savings**: ~60% scrollback memory for idle sessions
**R5: Lazy-load Unicode11Addon** (Low risk)
- Only load when CJK content is detected in terminal output
- Detection: check for characters in CJK Unicode ranges during ANSI stripping (already iterating)
- Most sessions never need it
---
## 3. DOM & Session Tab Rendering
### Current State
- Session tabs use **intelligent incremental updates** with debounced 100ms rendering
- Incremental path: only updates changed properties (classes, textContent, badges) when session list is stable
- Full rebuild path: triggered when sessions added/removed **or badge count changes**
- Subagent windows: per-window xterm.js instances, even when minimized
### Bottlenecks
**B6: Badge count change triggers full tab rebuild** (`app.js:3207-3209`)
- A single subagent badge increment on one tab triggers `_fullRenderSessionTabs()` — rebuilds entire sidebar HTML via `innerHTML =`
- With 20 sessions, this is an O(n) reflow for a single badge number change
- Badge changes are frequent during active subagent work
**B7: Minimized subagent windows retain xterm.js instances** (`subagent-windows.js`)
- 50 subagent windows × ~75KB per xterm.js instance = ~3.75MB DOM memory
- Minimized windows are invisible but their terminals remain in DOM
- xterm.js instances continue processing resize events even when hidden
**B8: `backdrop-filter: blur()` on overlays** (`styles.css:2246-2247, 3098`)
- Forces new stacking context, disables browser compositing optimizations
- 50-100ms layout thrashing on modal open/close
- Only 2 uses, but they're on frequently toggled overlays
### Recommendations
**R6: Targeted badge update without full rebuild** (Low risk)
- When badge count changes but session list is stable, update only the badge `<span>` textContent
- Keep incremental path for badge changes; only use full rebuild for structural changes (add/remove sessions)
- **Savings**: Eliminates O(n) reflow per badge change; reduces to O(1) targeted update
**R7: Lazy xterm.js initialization for subagent windows** (Medium impact)
- Only create xterm.js Terminal instance when window is restored/maximized
- On minimize: serialize terminal buffer, dispose Terminal instance, keep buffer in memory
- On restore: create new Terminal, write buffer back
- **Savings**: ~3.5MB DOM reduction at 50 minimized agents; eliminates hidden resize processing
- **Trade-off**: ~200-500ms restore delay (buffer write), mitigated by chunked loading
**R8: Replace `backdrop-filter: blur()` with `background: rgba()`** (Low risk)
- Use semi-transparent background instead of blur effect
- Or use `will-change: transform` hint if blur is kept
- **Savings**: Eliminates forced recomposition layer; 50-100ms faster overlay open
---
## 4. Backend PTY & State Management
### Current State (Excellent)
- **BufferAccumulator**: Array-based chunking with lazy join on read — avoids O(n) string concatenation
- **ANSI stripping**: Throttled at 150ms intervals with lazy evaluation (not per-chunk)
- **State persistence**: 500ms debounce + incremental JSON caching per session (only dirty sessions re-serialized)
- **Expensive parsers**: Throttled to 150ms window, accumulated data capped at 64KB
- **Memory**: All buffers have hard limits (2MB terminal, 1MB text, 1000 messages, 64KB line buffer)
### Bottlenecks
**B9: Pending clean data cap at 64KB** (`session.ts:1097-1133`)
- Between 150ms processing windows, raw PTY data accumulates in `_pendingCleanData`
- Capped at 64KB — excess data rolls off (old data discarded)
- During heavy output (large build logs), this means parsers may miss content
- Acceptable trade-off for performance, but worth documenting
**B10: `LRUMap.delete()` is O(n) worst case** (`utils/lru-map.ts:137-138`)
- When deleting the newest entry, iterates all keys to find new newest
- Rare in practice (delete is uncommon; set/get are hot paths)
- Could matter during mass cleanup of 500 agents
### Recommendations
**R9: Consider adaptive pending data cap** (Low priority)
- During idle detection (critical to get right), increase cap to 128KB
- During active working state, keep at 64KB (parsers less critical)
- **Benefit**: More accurate idle detection during heavy output
**R10: Track second-newest in LRUMap** (Low priority)
- Maintain a `_secondNewestKey` alongside `_newestKey`
- On delete of newest, promote second-newest without iteration
- Only matters at scale (500+ agents with frequent eviction)
---
## 5. Local Echo & Input Path
### Current State (Well-Designed)
- **DOM overlay approach** — `<span>` elements in `.xterm-screen` at z-index 7, completely independent of `terminal.write()`
- **Render caching**: `_lastRenderKey` includes text, position, column offsets — skips redundant re-renders
- **Input flow**: Char accumulation → Enter triggers flush → 80ms delay before `\r` (ensures text reaches PTY first)
- **Tab completion**: Baseline snapshot → detect buffer change → 300ms fallback timer
- **CJK support**: Per-character width detection with `terminal.unicode.getStringCellWidth()` preferred, manual fallback
- **Prompt detection**: Bottom-up line scan, O(rows) — cached position, column-lock prevents jitter
### Bottlenecks
**B11: tmux send-keys latency** (~50-100ms per input)
- Each `writeViaMux()` spawns a child process (`tmux send-keys`)
- Text and Enter sent separately with 50ms delay between
- For rapid typing: characters batch before Enter, so overhead is per-command not per-keystroke
- **Acceptable trade-off** for session persistence (tmux survives server restarts)
**B12: 80ms delay between text flush and Enter** (`app.js:872-875`)
- Intentional: ensures text reaches PTY before Enter, preventing Ink from processing empty input
- Adds 80ms to perceived Enter-to-response latency
- Could potentially be reduced with acknowledgment-based approach
**B13: Scroll listener on terminal viewport** (`zerolag-input-addon.ts:139`)
- 50ms debounced re-render on scroll — acceptable but fires frequently during heavy output
- Overlay hidden when scrolled up (correct behavior), shown when at bottom
### Recommendations
**R11: Reduce Enter delay from 80ms to 50ms** (Low risk, test carefully)
- The tmux `send-keys` already has 50ms internal delay
- Combined with network latency, 80ms client-side may be excessive
- Test with Ink-heavy sessions (Claude Code's status bar) — if text arrives before Enter at 50ms, reduce
- **Savings**: 30ms perceived latency reduction per command
**R12: Batch tmux send-keys via stdin pipe** (Medium effort, high impact for rapid input)
- Instead of spawning `tmux send-keys` per input, maintain a persistent connection
- Use `tmux -C` (control mode) for programmatic interaction without child process spawning
- **Savings**: Eliminate ~50-100ms process spawn overhead per input
- **Risk**: Control mode has different semantics; needs careful testing with session persistence
**R13: Skip overlay re-render during heavy output scroll** (Low risk)
- When terminal is receiving >10KB/s output, hide overlay entirely (user isn't typing during heavy output)
- Re-show overlay after 500ms of output silence
- **Savings**: Eliminates unnecessary DOM overlay re-renders during build logs / test output
---
## 6. Polling & File Watchers
### Current State
- **SubagentWatcher**: 1s base poll, full scan throttled to every 5s, fs.watch() on known directories
- **TranscriptWatcher**: 1 per session, fs.watch() primary with 1s poll fallback
- **ImageWatcher**: chokidar per session with 100ms stability poll, burst limit 20/10s
- **TeamWatcher**: chokidar primary with 30s poll fallback, LRU caches (50 teams, 200 tasks)
- **RalphTracker**: Todo cleanup every 5 minutes
### Scaling Profile (20 sessions)
| Component | Instances | Frequency | Total ops/sec |
|-----------|-----------|-----------|---------------|
| SubagentWatcher | 1 (global) | Full scan every 5s | 0.2/s |
| TranscriptWatcher | 20 | 1s poll (fallback) | 20/s max |
| ImageWatcher | 20 | 100ms poll (during writes only) | 200/s burst |
| TeamWatcher | 1 (global) | 30s poll (fallback) | 0.03/s |
| SSE heartbeat | 1 (global) | 15s | 0.07/s |
| SSE dead client check | 1 (global) | 30s | 0.03/s |
| Mux stats collection | 1 (global) | 2s | 0.5/s |
| **Total steady-state** | | | **~21/s** |
### Recommendations
**R14: Increase TranscriptWatcher poll interval to 2s** (Low risk)
- Transcript changes are infrequent (new messages every few seconds at most)
- fs.watch() is the primary mechanism; polling is fallback
- **Savings**: Halves fallback filesystem checks (20/s → 10/s for 20 sessions)
**R15: Share chokidar instances for co-located session directories** (Medium effort)
- Sessions in the same parent directory could share a single chokidar watcher with depth:3
- Common case: multiple sessions in `~/projects/foo/` — one watcher covers all
- **Savings**: Reduce chokidar instances from 20 to ~5-10 for typical workloads
---
## 7. Frontend Asset Delivery
### Current State
- **app.js**: 12,027 lines (source) → esbuild minified → gzip/brotli compressed (~30-40KB gzipped)
- **Static caching**: `maxAge: '1y'` via `@fastify/static`
- **Service worker**: Push notification handler only — no asset caching
- **No code splitting**: Single monolithic app.js bundle
### Bottlenecks
**B14: No cache-busting mechanism**
- `maxAge: '1y'` means browsers cache aggressively
- After deployment, users need `Ctrl+Shift+R` to see updates
- No content hash in filenames or ETags for automatic invalidation
**B15: Monolithic app.js**
- All 12K lines loaded on initial page load regardless of which features are used
- Ralph wizard, plan orchestrator UI, team management — all loaded upfront
- Mobile loads the same bundle as desktop
### Recommendations
**R16: Add content hash to asset filenames** (Medium impact)
- Build step: rename `app.js` → `app.[hash].js`
- Generate a manifest or inject hash into HTML template
- Keep `maxAge: '1y'` — cache invalidation happens via filename change
- **Savings**: Eliminates stale cache issues after deployment; removes need for manual hard refresh
**R17: Code-split app.js into core + feature modules** (High effort, medium impact)
- Core (~4K lines): terminal, SSE, session management, tabs, input handling
- Deferred (~8K lines): Ralph wizard, plan UI, team management, subagent windows, image viewer
- Load deferred modules on first use via dynamic `import()` or lazy `<script>` injection
- **Savings**: ~60% reduction in initial load size; faster time-to-interactive
- **Risk**: Complexity increase; need to handle loading states for deferred features
- **Note**: May not be worth the effort given the app is already gzipped to ~30-40KB
---
## 8. CSS Performance
### Current State
- **styles.css**: 7,153 lines with ~45 box-shadow uses, 2 backdrop-filter uses
- Animations: GPU-accelerated keyframes for pulsing alerts, loading spinners
- Z-index layering: well-organized (subagent 1000, plan 1100, log 2000, image 3000, overlay 7)
### Recommendations
**R18: Replace backdrop-filter with opaque overlay** (Low risk, covered in R8)
**R19: Use `contain: content` on subagent windows** (Low risk)
- Add CSS containment to subagent window containers
- Prevents layout changes inside windows from triggering reflow on parent
- Especially valuable with 50 windows: changes in one window won't invalidate others
- ```css
.subagent-window { contain: content; }
```
- **Savings**: Reduces layout recalculation scope from global to per-window
**R20: Use `content-visibility: auto` on off-screen subagent windows** (Low risk)
- Browser skips rendering of off-screen windows entirely
- Combined with `contain-intrinsic-size` to prevent layout shift
- ```css
.subagent-window.minimized { content-visibility: hidden; }
```
- **Savings**: Browser skips paint/layout for minimized windows; complements R7
---
## 9. Memory & Scaling Limits
### Current Budget (20 sessions)
| Component | Per Session | Total | Status |
|-----------|-----------|-------|--------|
| Terminal buffer | 2MB | 40MB | Hard-limited, auto-trim |
| Text output | 1MB | 20MB | Hard-limited, auto-trim |
| Messages | ~1MB | 20MB | Capped at 1000, trims to 800 |
| Respawn buffer | 1MB | 20MB | Hard-limited |
| **Buffers total** | | **100MB** | Acceptable |
| TranscriptWatcher | ~100KB | 2MB | |
| ImageWatcher | ~50KB | 1MB | |
| SubagentWatcher | ~500KB | 500KB | Global |
| Frontend terminal cache | ~256KB | 5MB | LRU, max 20 entries |
| **Total estimated** | | **~110MB** | Comfortable |
### At Max Scale (50 sessions)
- Buffers: ~250MB
- Watchers: ~5MB
- **Total: ~255MB** + Node.js overhead — acceptable on modern hardware
### Potential Leak Vectors (All Mitigated)
- `_shortIdCache` in server — unbounded Map, but entries are tiny (string→string); grows at O(sessions created), not O(events)
- All CleanupManager-registered resources tracked and disposed on session stop
- `isStopped` guard prevents new timers after session cleanup
---
## 10. Implementation Priority Matrix
### Phase 1 — Quick Wins (1-2 hours each, low risk)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R6 | Targeted badge update | `app.js` (3207-3209) |
| R3 | Canvas renderer on mobile | `app.js` (627-637) |
| R8 | Replace backdrop-filter blur | `styles.css` (2246, 3098) |
| R19 | CSS containment on subagent windows | `styles.css` |
| R20 | `content-visibility: hidden` on minimized windows | `styles.css` |
### Phase 2 — Medium Effort (half-day each)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R2 | Tiered SSE padding | `server.ts` (broadcast function) |
| R7 | Lazy xterm.js for minimized subagents | `subagent-windows.js` |
| R11 | Reduce Enter delay to 50ms | `app.js` (872-875), test with Ink |
| R14 | TranscriptWatcher 2s poll | `transcript-watcher.ts` |
| R16 | Content-hash asset filenames | `build.mjs`, `server.ts` |
### Phase 3 — Larger Initiatives (1-2 days each)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R1 | Session-scoped SSE subscriptions | `server.ts`, `app.js` (SSE connect) |
| R5 | Lazy Unicode11Addon loading | `app.js`, build pipeline |
| R12 | Persistent tmux control mode | `tmux-manager.ts` |
| R17 | Code-split app.js | `app.js`, `build.mjs`, HTML template |
### Not Recommended (Low ROI or High Risk)
| # | Why Not |
|---|---------|
| R4 | Dynamic scrollback adds complexity; memory savings marginal vs total budget |
| R9 | Adaptive pending data cap adds state; current 64KB cap rarely matters |
| R10 | LRUMap.delete() O(n) is theoretical; never triggered at current scale |
| R15 | Shared chokidar instances add directory-matching complexity for minimal gain |
---
## Appendix: Key File Locations
| Area | File | Key Lines |
|------|------|-----------|
| SSE broadcast | `src/web/server.ts` | 1961-1989 (broadcast), 1934-1959 (backpressure) |
| Terminal batching | `src/web/server.ts` | 1994-2048 (per-session adaptive batching) |
| Frame budget | `src/web/public/app.js` | 1370-1478 (flushPendingWrites, 64KB cap) |
| Flicker filter | `src/web/public/app.js` | 1176-1255 (50ms sync wait, 256KB safety) |
| Tab rendering | `src/web/public/app.js` | 3108-3357 (incremental + full rebuild) |
| Tab switching | `src/web/public/app.js` | 3560-3760 (cache + chunked load + deferred UI) |
| Local echo | `packages/xterm-zerolag-input/src/` | All files (overlay, prompt, CJK) |
| Local echo integration | `src/web/public/app.js` | 640, 815-988 (input flow) |
| Subagent windows | `src/web/public/subagent-windows.js` | Full file (window mgmt, drag, minimize) |
| State persistence | `src/state-store.ts` | 161-250 (debounced save, incremental JSON) |
| Buffer accumulator | `src/utils/buffer-accumulator.ts` | Full file (array chunks, lazy join) |
| PTY handling | `src/session.ts` | 1046-1133 (data flow), 1173-1230 (parsing) |
| Config limits | `src/config/` | 9 files (buffer, map, timing, auth, etc.) |
| Anti-flicker docs | `docs/terminal-anti-flicker.md` | Architecture reference |
| CSS | `src/web/public/styles.css` | 2246 (backdrop-filter), full file |
| Build pipeline | `scripts/build.mjs` | 59-68 (minify + compress) |
+38 -55
View File
@@ -1,7 +1,7 @@
# Performance & Responsiveness Optimization Plan
**Date**: 2026-02-28
**Status**: In Progress
**Status**: Phases 1–4 Complete. Phase 5 optional/deferred.
---
@@ -13,7 +13,7 @@ Three independent research passes analyzed the Codeman codebase for performance
---
## Phase 1: Quick Wins — ALREADY IMPLEMENTED
## Phase 1: Quick Wins — COMPLETE
All Phase 1 items were found to already exist in the codebase during verification:
@@ -27,7 +27,7 @@ All Phase 1 items were found to already exist in the codebase during verificatio
---
## Phase 2: Frontend Responsiveness — MOSTLY ALREADY IMPLEMENTED
## Phase 2: Frontend Responsiveness — COMPLETE
### 2.1 Batch `getBoundingClientRect()` in connection lines — DONE
- **Files**: `src/web/public/app.js` (`_updateConnectionLinesImmediate()`)
@@ -51,7 +51,7 @@ All Phase 1 items were found to already exist in the codebase during verificatio
---
## Phase 3: Backend Hot Paths
## Phase 3: Backend Hot Paths — COMPLETE
### 3.1 State diff broadcasts — ALREADY OPTIMIZED
- `broadcastSessionStateDebounced()` already batches at 500ms intervals
@@ -84,35 +84,37 @@ All Phase 1 items were found to already exist in the codebase during verificatio
---
## Phase 4: System-Level Improvements
## Phase 4: System-Level Improvements — COMPLETE
### 4.1 Incremental state persistence
- **Files**: `src/state-store.ts` (~lines 145-160)
- **Problem**: Every 500ms debounce writes the entire `AppState` (all sessions, tasks, config) via `JSON.stringify()`. With 50 sessions, state can be tens of MB. Serialization alone costs 50-100ms.
- **Fix**: Track dirty sessions. On persist, only re-serialize dirty sessions; cache serialized JSON for clean sessions. Assemble final output from cached fragments.
- **Impact**: Reduces serialization cost from O(all sessions) to O(dirty sessions). Typical steady-state: 1-2 dirty sessions instead of 50.
### 4.1 Incremental state persistence — DONE
- **Files**: `src/state-store.ts` (`assembleStateJson()`, `setSession()`)
- **Change**: Added `dirtySessions` Set and `cachedSessionJsons` Map. On persist, only dirty sessions are re-serialized; clean sessions reuse cached JSON fragments. `setSession()` marks sessions dirty; `assembleStateJson()` rebuilds only changed fragments.
- **Impact**: Serialization cost reduced from O(all sessions) to O(dirty sessions). Typical steady-state: 1-2 dirty sessions instead of 50.
### 4.2 Replace polling with fs watchers for team watcher
- **Files**: `src/team-watcher.ts` (~lines 148-180)
- **Problem**: Polls `~/.claude/teams/` every 5s via `readdir()` + `stat()`. Blocks event loop for 100-200ms on large directories.
- **Fix**: Use `chokidar` (already a dependency) or `fs.watch()` to react to changes. Keep a 30s fallback poll for reliability.
- **Impact**: Eliminates 5s polling overhead; near-instant team detection.
### 4.2 Replace polling with fs watchers for team watcher — DONE
- **Files**: `src/team-watcher.ts` (`setupFsWatchers()`)
- **Change**: Added chokidar watchers on both `~/.claude/teams/` and `~/.claude/tasks/` directories for instant event-driven detection. Lock files ignored via chokidar config. Mtime-based dedup skips unchanged files. Polling interval relaxed from 5s to 30s as a fallback.
- **Impact**: Near-instant team detection; polling overhead eliminated for normal operation.
### 4.3 Consolidate subagent file watchers
- **Files**: `src/subagent-watcher.ts` (~line 229+)
- **Problem**: One chokidar watcher per agent directory. With 500 agents, that's 500 inotify watchers consuming kernel resources.
- **Fix**: Watch at the session level (one watcher per session's subagent directory), not per-agent. Parse events to route to correct agent.
- **Impact**: Reduces inotify watchers from 500 to ~50 (one per session).
### 4.3 Consolidate subagent file watchers — DONE
- **Files**: `src/subagent-watcher.ts` (`setupDirectoryWatcher()`)
- **Change**: Replaced per-agent chokidar watchers with one `fs.watch()` per session subagent directory. Events are routed to the correct agent via filename. Per-file debouncing (100ms) prevents hammering on bulk discovery.
- **Impact**: Inotify watchers reduced from potentially 500 (one per agent) to ~50 (one per session directory).
### 4.4 Stream transcript files instead of full reads
- **Files**: `src/subagent-watcher.ts` (~lines 959-964)
- **Problem**: `loadTranscript()` reads entire transcript file (can be >100KB). With 500 agents discovered at once, that's 50MB of file reads.
- **Fix**: Only read last 10KB for display (tail). Full file on-demand only (e.g., when user opens transcript viewer).
- **Impact**: Reduces file I/O from 50MB to 5MB for bulk agent discovery.
### 4.4 Stream transcript files instead of full reads — DONE
- **Files**: `src/subagent-watcher.ts` (`tailFile()`, `findDescriptionInAgentFile()`, parent transcript lookup)
- **Change**: Multiple streaming strategies implemented:
- **Live monitoring**: Position-based `tailFile()` with `createReadStream({ start: fromPosition })` — only reads new content
- **Parent transcript lookup**: Streams only last 16KB (`createReadStream({ start: offset })`)
- **Description extraction**: Streams only first 8KB, exits early after 5 lines
- **Full read**: Only for on-demand transcript review panel (with optional `limit` parameter)
- **Impact**: File I/O for bulk agent discovery reduced from ~50MB to ~5MB.
---
## Phase 5: Long-Term Architectural (Optional)
## Phase 5: Long-Term Architectural (Optional) — NOT STARTED
These items are deferred until scaling demands justify the complexity.
### 5.1 Worker thread for PTY processing
- **Files**: `src/session.ts`
@@ -134,42 +136,23 @@ All Phase 1 items were found to already exist in the codebase during verificatio
---
## Priority Matrix (Remaining Work)
## Completion Summary
| # | Item | Impact | Risk | Effort |
|---|------|--------|------|--------|
| 3.1 | State diff broadcasts | **Very High** | Medium | 3-4h |
| 3.2 | Fix session cache invalidation | **High** | Low | 1h |
| 3.3 | Skip PTY processing for hidden sessions | **High** | Medium | 2-3h |
| 3.5 | Throttle detection broadcasts | **Medium** | Low | 1h |
| 3.4 | Batch liveness checks | **Medium** | Low | 1-2h |
| 4.1 | Incremental state persistence | **Medium** | Medium | 3-4h |
| 4.2 | Team watcher fs events | **Low-Med** | Medium | 2h |
| 4.3 | Consolidate file watchers | **Low-Med** | Medium | 2h |
| 4.4 | Stream transcripts | **Low-Med** | Low | 1h |
| 5.1 | Worker thread PTY | **Med** (at scale) | High | 8h |
| 5.2 | Per-session SSE subs | **Med** (at scale) | High | 4h |
| 5.3 | O(1) LRUMap | **Very Low** | Medium | 2h |
| Phase | Scope | Status | Items |
|-------|-------|--------|-------|
| 1 | Quick Wins | **Complete** | 5/5 (all pre-existing) |
| 2 | Frontend Responsiveness | **Complete** | 3/3 actionable done, 2 skipped |
| 3 | Backend Hot Paths | **Complete** | 4/4 actionable done, 1 deferred |
| 4 | System-Level | **Complete** | 4/4 done |
| 5 | Long-Term Architectural | **Not started** | 0/3 — deferred until needed |
---
## Recommended Execution Order
**Sprint 1** (Phase 3 — Backend Hot Paths): Items 3.1, 3.2, 3.3, 3.5
- Backend serialization and broadcast efficiency
- Highest remaining impact; requires careful testing with multiple active sessions
**Sprint 2** (Phase 4 — System Level): Items 4.1, 3.4, 4.3, 4.4
- State persistence, liveness checks, watcher consolidation
- Medium-complexity refactors
**Sprint 3** (Phase 5 — Architectural): Items 5.1, 5.2 — only if scaling demands it
**Overall**: 16/16 actionable items complete. 3 optional items deferred.
---
## Measurement
Before starting implementation, establish baselines:
Before starting Phase 5, establish baselines:
1. **Frontend**: Record Chrome DevTools Performance trace with 10 sessions open. Measure:
- Frame rate during rapid terminal output
+788
View File
@@ -0,0 +1,788 @@
# Phase 4: Domain File Splitting — Implementation Plan
**Date**: 2026-03-01
**Prerequisites**: Phase 1-3 complete (utils cleanup, CleanupManager/Debouncer migration, route extraction)
**Goal**: Split 4 god files into focused modules with barrel exports for transparent migration.
---
## Table of Contents
1. [Split types.ts into types/ directory](#1-split-typests-into-types-directory)
2. [Split ralph-tracker.ts into focused modules](#2-split-ralph-trackerts-into-focused-modules)
3. [Split respawn-controller.ts into focused modules](#3-split-respawn-controllerts-into-focused-modules)
4. [Split session.ts into focused modules](#4-split-sessionts-into-focused-modules)
5. [Execution Order & Dependencies](#5-execution-order--dependencies)
6. [Validation Checklist](#6-validation-checklist)
---
## 1. Split types.ts into types/ directory
**Current**: 1,443 lines, 71 exports, imported by 36 files.
**Risk**: LOW — pure type refactor, no runtime behavior change.
### Target Structure
```
src/types/
├── index.ts (barrel re-export — transparent migration)
├── common.ts (Disposable, BufferConfig, CleanupResourceType, CleanupRegistration)
├── session.ts (SessionStatus, SessionMode, ClaudeMode, SessionConfig, SessionColor,
│ SessionState, OpenCodeConfig, SessionOutput)
├── task.ts (TaskStatus, TaskDefinition, TaskState)
├── app-state.ts (AppState, AppConfig, GlobalStats, TokenUsageEntry, TokenStats,
│ DEFAULT_CONFIG, createInitialState, createInitialGlobalStats)
├── respawn.ts (RespawnConfig, PersistedRespawnConfig, CycleOutcome,
│ RespawnCycleMetrics, RespawnAggregateMetrics, HealthStatus,
│ RalphLoopHealthScore, TimingHistory, RespawnPreset)
├── ralph.ts (RalphLoopStatus, RalphLoopState, RalphTodoStatus, RalphTodoPriority,
│ RalphTodoItem, RalphTodoProgress, RalphSessionState,
│ RalphStatusValue, RalphTestsStatus, RalphWorkType, RalphStatusBlock,
│ CompletionConfidence, RalphTrackerState,
│ CircuitBreakerState, CircuitBreakerReason, CircuitBreakerStatus,
│ createInitialCircuitBreakerStatus, createInitialRalphTrackerState,
│ createInitialRalphSessionState)
├── api.ts (ApiErrorCode, ApiResponse, HookEventType, QuickStartResponse,
│ CaseInfo, createErrorResponse, isError, getErrorMessage)
├── lifecycle.ts (LifecycleEventType, LifecycleEntry)
├── run-summary.ts (RunSummaryEventType, RunSummaryEventSeverity, RunSummaryEvent,
│ RunSummaryStats, RunSummary, createInitialRunSummaryStats)
├── tools.ts (ActiveBashToolStatus, ActiveBashTool, ImageDetectedEvent)
├── teams.ts (TeamConfig, TeamMember, TeamTask, InboxMessage, PaneInfo)
├── push.ts (PushSubscriptionRecord, VapidKeys)
└── plan.ts (PlanTaskStatus, TddPhase, PlanItem re-export, NiceConfig,
DEFAULT_NICE_CONFIG, ProcessStats)
```
### Steps
1. **Create `src/types/` directory** and each domain file above.
2. **Move types** from `src/types.ts` into their domain files. Preserve all JSDoc comments. Each file should import from siblings as needed (e.g., `ralph.ts` imports `CircuitBreakerState` within itself — no cross-file deps needed since they're in the same file).
3. **Create barrel `src/types/index.ts`** that re-exports everything:
```typescript
export * from './common.js';
export * from './session.js';
export * from './task.js';
export * from './app-state.js';
export * from './respawn.js';
export * from './ralph.js';
export * from './api.js';
export * from './lifecycle.js';
export * from './run-summary.js';
export * from './tools.js';
export * from './teams.js';
export * from './push.js';
export * from './plan.js';
```
4. **Delete old `src/types.ts`** and replace with a single-line re-export barrel:
```typescript
export * from './types/index.js';
```
This ensures `import { ... } from './types.js'` continues to work everywhere — zero changes to 36 import sites.
5. **Verify**: `tsc --noEmit` and `npm run lint` must pass. No runtime changes.
### Internal Dependencies Between Domain Files
Some types reference others across domains. Handle with imports:
| File | Imports From |
|------|-------------|
| `app-state.ts` | `session.ts` (SessionState), `task.ts` (TaskState), `ralph.ts` (RalphLoopState, RalphSessionState) |
| `respawn.ts` | None (self-contained) |
| `ralph.ts` | None (self-contained) |
| `run-summary.ts` | None (self-contained) |
| `api.ts` | None (self-contained) |
| `session.ts` | `respawn.ts` (RespawnConfig), `ralph.ts` (RalphTrackerState, RalphTodoItem, CircuitBreakerStatus, RalphSessionState, RunSummaryEvent) |
Wait — `SessionState` references `RespawnConfig`, `RalphTrackerState`, `CircuitBreakerStatus`, and `RunSummaryEvent`. This creates imports from `session.ts` → `respawn.ts`, `ralph.ts`, `run-summary.ts`. This is fine (one-way deps, no cycles).
---
## 2. Split ralph-tracker.ts into focused modules
**Current**: 3,868 lines, single `RalphTracker` class with 5 responsibilities.
**Risk**: MEDIUM — class has shared mutable state, but extractable modules are well-isolated.
### Coupling Analysis Summary
| Module | Coupling | Extractability |
|--------|----------|----------------|
| Plan task tracking | LOW | HIGH — only reads `cycleCount` |
| Fix-plan file watching | LOW | HIGH — callback-based todo replacement |
| Iteration stall detection | LOW | HIGH — notification-based |
| RALPH_STATUS block parsing + circuit breaker | MEDIUM | MEDIUM — callback for circuit breaker updates |
| Todo parsing, loop detection, completion | HIGH | LOW — deeply entangled shared state |
### Target Structure
```
src/
├── ralph-tracker.ts (~1,800 LOC — core: output parsing, loop state,
│ todo management, completion detection)
├── ralph-plan-tracker.ts (~600 LOC — plan tasks, checkpoints, history, rollback)
├── ralph-status-parser.ts (~300 LOC — RALPH_STATUS block parsing, circuit breaker)
├── ralph-fix-plan-watcher.ts (~150 LOC — @fix_plan.md file watching)
└── ralph-stall-detector.ts (~80 LOC — iteration stall detection)
```
### Step 2a: Extract `RalphPlanTracker` (~600 LOC)
**Why first**: Lowest coupling. Only dependency is `cycleCount` for checkpoint detection.
**Extract these from `RalphTracker`**:
Types to export:
- `EnhancedPlanTask` (interface, currently lines 56-87)
- `CheckpointReview` (interface, currently lines 90-139)
Properties to move:
- `_planVersion: number`
- `_planHistory: Array<{version, timestamp, tasks, summary}>`
- `_planTasks: Map<string, EnhancedPlanTask>`
- `_checkpointIterations: number[]`
- `_lastCheckpointIteration: number`
Methods to move:
- `initializePlanTasks(items)`
- `updatePlanTask(taskId, update)`
- `addPlanTask(params)`
- `getPlanTasks()`
- `generateCheckpointReview()`
- `getPlanHistory()`
- `rollbackToVersion(version)`
- `isCheckpointDue()`
- `planVersion` getter
- `_savePlanToHistory()` (private)
- `_unblockDependentTasks()` (private)
- `_checkForCheckpoint()` (private)
Events emitted (define in new class):
- `planInitialized`
- `planTaskUpdate`
- `taskBlocked`
- `taskUnblocked`
- `planCheckpoint`
**Interface with parent**:
```typescript
export class RalphPlanTracker extends EventEmitter {
constructor() { ... }
// Parent calls this when iteration changes (for checkpoint detection)
notifyCycleCount(cycleCount: number): void { ... }
// Full public API moves here unchanged
initializePlanTasks(items: PlanItem[]): void { ... }
updatePlanTask(taskId: string, update: { ... }): { ... } | null { ... }
// ...etc
}
```
**In `RalphTracker`**: Replace plan methods with delegation:
```typescript
readonly planTracker = new RalphPlanTracker();
// Forward plan events
this.planTracker.on('planInitialized', (...args) => this.emit('planInitialized', ...args));
// ...etc
// In detectLoopStatus(), when cycleCount changes:
this.planTracker.notifyCycleCount(this._loopState.cycleCount);
```
### Step 2b: Extract `RalphFixPlanWatcher` (~150 LOC)
**Extract these**:
Properties:
- `_workingDir: string | null`
- `_fixPlanPath: string | null`
- `_fixPlanWatcher: FSWatcher | null`
- `_fixPlanWatcherErrorHandler`
- `_fixPlanReloadDeb`
Methods:
- `setWorkingDir(workingDir)`
- `loadFixPlanFromDisk()`
- `startWatchingFixPlan()`
- `stopWatchingFixPlan()`
- `handleFixPlanChange()`
- `isFileAuthoritative` getter
**Interface with parent**:
```typescript
export class RalphFixPlanWatcher extends EventEmitter {
get isFileAuthoritative(): boolean { ... }
setWorkingDir(workingDir: string): void { ... }
stop(): void { ... }
}
// Events:
// 'todosLoaded' → (todos: Array<{id, content, status, priority}>) — parent replaces _todos
```
**In `RalphTracker`**:
```typescript
readonly fixPlanWatcher = new RalphFixPlanWatcher();
constructor() {
this.fixPlanWatcher.on('todosLoaded', (items) => {
// Replace _todos with file-based items
this._todos.clear();
for (const item of items) {
this.addOrUpdateTodo(item.id, item.content, item.status, item.priority);
}
});
}
// Delegate isFileAuthoritative
get isFileAuthoritative(): boolean {
return this.fixPlanWatcher.isFileAuthoritative;
}
```
### Step 2c: Extract `RalphStallDetector` (~80 LOC)
**Extract these**:
Properties:
- `_lastIterationChangeTime`
- `_lastObservedIteration`
- `_iterationStallTimerId`
- `_iterationStallWarningMs`
- `_iterationStallCriticalMs`
- `_iterationStallWarned`
Methods:
- `startIterationStallDetection()`
- `stopIterationStallDetection()`
- `checkIterationStall()`
- `getIterationStallMetrics()`
- `configureIterationStallThresholds(warningMs, criticalMs)`
**Interface with parent**:
```typescript
export class RalphStallDetector extends EventEmitter {
constructor(private cleanup: CleanupManager) { ... }
start(): void { ... }
stop(): void { ... }
// Parent calls when iteration changes
notifyIterationChanged(iteration: number): void {
this._lastIterationChangeTime = Date.now();
this._lastObservedIteration = iteration;
this._iterationStallWarned = false;
}
// Parent calls to check if loop is active
setLoopActive(active: boolean): void { ... }
getIterationStallMetrics(): { ... } { ... }
}
// Events: 'iterationStallWarning', 'iterationStallCritical'
```
### Step 2d: Extract `RalphStatusParser` (~300 LOC)
**Extract these**:
Properties:
- `_circuitBreaker: CircuitBreakerStatus`
- `_statusBlockBuffer: string[]`
- `_inStatusBlock: boolean`
- `_lastStatusBlock: RalphStatusBlock | null`
- `_completionIndicators: number`
- `_exitGateMet: boolean`
- `_totalFilesModified: number`
- `_totalTasksCompleted: number`
Methods:
- `processStatusBlockLine(line)`
- `parseStatusBlock(lines)`
- `detectCompletionIndicators(line)`
- `updateCircuitBreaker(hasProgress, testsStatus, status)`
- `resetCircuitBreaker()`
- `circuitBreakerStatus` getter
- `lastStatusBlock` getter
- `cumulativeStats` getter
- `exitGateMet` getter
Regex patterns to move:
- `RALPH_STATUS_START_PATTERN` through `RALPH_RECOMMENDATION_PATTERN`
- `COMPLETION_INDICATOR_PATTERNS`
**Interface with parent**:
```typescript
export class RalphStatusParser extends EventEmitter {
processLine(line: string): void { ... } // calls processStatusBlockLine + detectCompletionIndicators
get circuitBreakerStatus(): CircuitBreakerStatus { ... }
get lastStatusBlock(): RalphStatusBlock | null { ... }
get exitGateMet(): boolean { ... }
get cumulativeStats(): { ... } { ... }
resetCircuitBreaker(): void { ... }
reset(): void { ... }
}
// Events: 'statusBlockDetected', 'circuitBreakerUpdate', 'exitGateMet'
```
**In `RalphTracker.processLine()`**:
```typescript
// Replace inline status block handling with delegation
this.statusParser.processLine(line);
```
### Step 2e: Keep in `ralph-tracker.ts` (~1,800 LOC)
The core remains tightly coupled and stays together:
- Output parsing pipeline (`processTerminalData`, `processCleanData`, `processLine`)
- Loop state management (`_loopState`, `detectLoopStatus`, `enable/disable/startLoop/stopLoop`)
- Todo management (`_todos`, `detectTodoItems`, `addOrUpdateTodo`, `updateTodoStatus`, `getTodoStats`)
- Completion detection (`detectCompletionPhrase`, `handleCompletionPhrase`, `calculateCompletionConfidence`)
- All-tasks-complete detection (`detectAllTasksComplete`)
- Auto-enable logic (`shouldAutoEnable`)
- Lifecycle (`reset`, `fullReset`, `clear`, `restoreState`, `destroy`)
- Event debouncing and buffering
The class coordinates the extracted modules via composition:
```typescript
export class RalphTracker extends EventEmitter {
readonly planTracker = new RalphPlanTracker();
readonly fixPlanWatcher = new RalphFixPlanWatcher();
readonly stallDetector: RalphStallDetector;
readonly statusParser = new RalphStatusParser();
constructor() {
super();
this.stallDetector = new RalphStallDetector(this.cleanup);
this._wireSubModuleEvents();
}
private _wireSubModuleEvents(): void {
// Forward all sub-module events through RalphTracker
// so external consumers don't need to know about the split
for (const event of ['planInitialized', 'planTaskUpdate', ...]) {
this.planTracker.on(event, (...args) => this.emit(event, ...args));
}
// ...same for statusParser, stallDetector, fixPlanWatcher
}
}
```
### Migration Safety
- All events continue to be emitted from `RalphTracker` (forwarded from sub-modules)
- All public methods stay on `RalphTracker` (delegated to sub-modules)
- External consumers (`session.ts`, `case-routes.ts`) see zero API changes
- New sub-modules are exposed as `readonly` properties for direct access where needed
---
## 3. Split respawn-controller.ts into focused modules
**Current**: 3,611 lines, single `RespawnController` class with 6 responsibilities.
**Risk**: MEDIUM — health scoring and metrics are cleanly decoupled; detection is tightly coupled.
### Coupling Analysis Summary
| Module | Coupling | Extractability |
|--------|----------|----------------|
| Health scoring | NONE | HIGH — pure calculations from metrics |
| Cycle metrics | LOW | HIGH — standalone tracking |
| Adaptive timing | LOW | HIGH — standalone timing adjustments |
| Stuck-state detection | LOW | MEDIUM — needs state + config refs |
| Pattern detection utilities | NONE | HIGH — pure functions |
| State machine + idle detection + AI checkers | HIGH | LOW — deeply entangled |
### Target Structure
```
src/
├── respawn-controller.ts (~2,200 LOC — state machine, idle detection,
│ AI checkers, terminal handling, hook signals,
│ auto-accept, step execution)
├── respawn-health.ts (~250 LOC — health scoring + recommendations)
├── respawn-metrics.ts (~200 LOC — cycle metrics + aggregate stats)
├── respawn-adaptive-timing.ts (~100 LOC — adaptive timing with percentile calc)
└── respawn-patterns.ts (~50 LOC — terminal pattern detection utilities)
```
### Step 3a: Extract `RespawnPatterns` (~50 LOC)
**Pure utility functions, zero coupling**.
Move:
- `isCompletionMessage(data): boolean`
- `hasWorkingPattern(data, window): boolean`
- `extractTokenCount(data): number | null`
- `PROMPT_PATTERNS` array
- `WORKING_PATTERNS` array
```typescript
// src/respawn-patterns.ts
import { TOKEN_PATTERN, SPINNER_PATTERN } from './utils/index.js';
export const PROMPT_PATTERNS = ['❯', '>', '$', '%', '#'];
export const WORKING_PATTERNS = [/* 70+ patterns */];
export function isCompletionMessage(data: string): boolean { ... }
export function hasWorkingPattern(data: string, window: string): boolean { ... }
export function extractTokenCount(data: string): number | null { ... }
```
**In `RespawnController`**: Import and call:
```typescript
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
```
### Step 3b: Extract `RespawnAdaptiveTiming` (~100 LOC)
**Self-contained timing controller**.
Move properties:
- `timingHistory: TimingHistory`
Move methods:
- `recordTimingData(idleDetectionMs, cycleDurationMs)`
- `updateAdaptiveTiming()`
- `getTimingHistory()`
- `getAdaptiveCompletionConfirmMs()`
```typescript
export class RespawnAdaptiveTiming {
private timingHistory: TimingHistory;
constructor(private config: { adaptiveMinConfirmMs: number; adaptiveMaxConfirmMs: number }) {
this.timingHistory = { recentIdleDetectionMs: [], recentCycleDurationMs: [], ... };
}
recordTimingData(idleDetectionMs: number, cycleDurationMs: number): void { ... }
getAdaptiveCompletionConfirmMs(): number { ... }
getTimingHistory(): TimingHistory { ... }
reset(): void { ... }
}
```
### Step 3c: Extract `RespawnCycleMetrics` (~200 LOC)
**Standalone metrics tracker**.
Move properties:
- `currentCycleMetrics`
- `recentCycleMetrics[]`
- `aggregateMetrics`
- `MAX_CYCLE_METRICS_IN_MEMORY`
Move methods:
- `startCycleMetrics(idleReason)`
- `recordCycleStep(step)`
- `completeCycleMetrics(outcome, errorMessage?)`
- `updateAggregateMetrics(metrics)`
- `getAggregateMetrics()`
- `getRecentCycleMetrics(limit?)`
```typescript
export class RespawnCycleMetricsTracker {
private currentCycleMetrics: Partial<RespawnCycleMetrics> | null = null;
private recentCycleMetrics: RespawnCycleMetrics[] = [];
private aggregateMetrics: RespawnAggregateMetrics;
startCycle(sessionId: string, cycleNumber: number, idleReason: string): void { ... }
recordStep(step: string): void { ... }
completeCycle(outcome: CycleOutcome, errorMessage?: string): RespawnCycleMetrics | null { ... }
getAggregate(): RespawnAggregateMetrics { ... }
getRecent(limit?: number): RespawnCycleMetrics[] { ... }
reset(): void { ... }
}
```
**Callback**: `completeCycle()` returns the completed metrics so the controller can pass them to `adaptiveTiming.recordTimingData()`.
### Step 3d: Extract `RespawnHealthCalculator` (~250 LOC)
**Pure calculation — no state of its own**.
Move methods:
- `calculateHealthScore()`
- `calculateCycleSuccessScore()`
- `calculateCircuitBreakerScore()`
- `calculateIterationProgressScore()`
- `calculateAiCheckerScore()`
- `calculateStuckRecoveryScore()`
- `generateHealthRecommendations(components)`
- `generateHealthSummary(score, status, components)`
- `shouldSkipClear()` (belongs here since it's a pure calculation on token/config)
```typescript
export interface HealthInputs {
aggregateMetrics: RespawnAggregateMetrics;
circuitBreakerStatus: CircuitBreakerStatus;
iterationStallMetrics: { stallDurationMs: number; warningMs: number; criticalMs: number } | null;
aiCheckerState: { disabled: boolean; inCooldown: boolean; hasErrors: boolean };
stuckRecoveryCount: number;
maxStuckRecoveries: number;
}
export function calculateHealthScore(inputs: HealthInputs): RalphLoopHealthScore { ... }
export function shouldSkipClear(
lastTokenCount: number,
skipClearThresholdPercent: number,
maxContextTokens: number
): boolean { ... }
```
**Made as pure functions** (not a class) since they hold no state.
### Step 3e: Keep in `respawn-controller.ts` (~2,200 LOC)
The core state machine, idle detection, and AI checker integration stays:
- State machine transitions (`setState`, `start`, `stop`, `pause`, `resume`)
- Terminal data handling (`handleTerminalData`)
- All 5 idle detection layers + hook signals
- AI checker integration (`tryStartAiCheck`, `startAiCheck`, `startPlanCheck`)
- Auto-accept logic
- Step execution (`sendUpdateDocs`, `sendClear`, `sendInit`, `sendKickstart`)
- Timer management (`startTrackedTimer`, `cancelTrackedTimer`)
- Stuck-state detection and recovery
- Action logging
The class composes extracted modules:
```typescript
import { RespawnAdaptiveTiming } from './respawn-adaptive-timing.js';
import { RespawnCycleMetricsTracker } from './respawn-metrics.js';
import { calculateHealthScore, shouldSkipClear } from './respawn-health.js';
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
export class RespawnController extends EventEmitter {
private adaptiveTiming: RespawnAdaptiveTiming;
private cycleMetrics: RespawnCycleMetricsTracker;
calculateHealthScore(): RalphLoopHealthScore {
return calculateHealthScore({
aggregateMetrics: this.cycleMetrics.getAggregate(),
circuitBreakerStatus: this.session.ralphTracker.circuitBreakerStatus,
iterationStallMetrics: this.session.ralphTracker.getIterationStallMetrics(),
aiCheckerState: { ... },
stuckRecoveryCount: this.stuckRecoveryCount,
maxStuckRecoveries: this.config.maxStuckRecoveries ?? 3,
});
}
}
```
---
## 4. Split session.ts into focused modules
**Current**: 2,418 lines, single `Session` class.
**Risk**: LOW-MEDIUM — extractable pieces are utility-like with clear boundaries.
### Coupling Analysis Summary
| Module | Coupling | Extractability |
|--------|----------|----------------|
| CLI arg builder | NONE | HIGH — pure functions used at spawn time |
| Auto-compact/clear | LOW | HIGH — self-contained automation with config |
| Token tracking | LOW | MEDIUM — reads PTY output, writes state |
| Task description cache | LOW | HIGH — separate LRU cache |
| PTY + mux lifecycle | HIGH | KEEP — core of the class |
| Tracker integration | HIGH | KEEP — event forwarding plumbing |
### Target Structure
```
src/
├── session.ts (~1,600 LOC — PTY lifecycle, terminal I/O,
│ tracker integration, output processing,
│ token tracking, state management)
├── session-cli-builder.ts (~250 LOC — Claude/OpenCode CLI arg construction)
├── session-auto-ops.ts (~300 LOC — auto-compact, auto-clear automation)
└── session-task-cache.ts (~100 LOC — task description LRU cache)
```
### Step 4a: Extract `SessionCliBuilder` (~250 LOC)
**Pure functions — zero coupling to Session instance**.
Move:
- `buildClaudeArgs()` logic (currently inlined in `startInteractive` and `runPrompt`)
- `buildOpenCodeArgs()` logic
- Model mapping constants
- Claude mode to flag mapping
- Environment variable construction
```typescript
// src/session-cli-builder.ts
export interface CliBuilderConfig {
claudeMode: ClaudeMode;
model?: string;
workingDir: string;
sessionId: string;
niceConfig?: NiceConfig;
isOpenCode?: boolean;
openCodeConfig?: OpenCodeConfig;
}
export function buildInteractiveArgs(config: CliBuilderConfig): string[] { ... }
export function buildPromptArgs(config: CliBuilderConfig, prompt: string): string[] { ... }
export function buildShellArgs(shell?: string): string[] { ... }
export function buildClaudeEnv(config: CliBuilderConfig): Record<string, string> { ... }
```
### Step 4b: Extract `SessionAutoOps` (~300 LOC)
**Self-contained automation with config-based thresholds**.
Move properties:
- `_autoCompactThreshold`
- `_autoClearThreshold`
- `_isAutoCompacting`
- `_isAutoClearing`
- `_autoCompactCount`
- `_autoClearCount`
- `_lastAutoCompactTime`
- `_lastAutoClearTime`
Move methods:
- `checkAutoCompact(tokenCount)`
- `performAutoCompact()`
- `checkAutoClear(tokenCount)`
- `performAutoClear()`
- Auto-compact/clear threshold configuration
```typescript
export class SessionAutoOps extends EventEmitter {
constructor(
private writeCommand: (command: string) => Promise<void>,
private getTokenCount: () => number,
config: { compactThreshold: number; clearThreshold: number }
) { ... }
/** Called after token count updates. Checks thresholds and triggers if needed. */
checkThresholds(tokenCount: number): void { ... }
updateConfig(config: { compactThreshold?: number; clearThreshold?: number }): void { ... }
getStats(): { autoCompactCount: number; autoClearCount: number; ... } { ... }
}
// Events: 'autoCompact', 'autoClear'
```
**In `Session`**: Compose and wire:
```typescript
private autoOps = new SessionAutoOps(
(cmd) => this.writeViaMux(cmd),
() => this._state.tokenCount,
{ compactThreshold: 110_000, clearThreshold: 140_000 }
);
```
### Step 4c: Extract `SessionTaskCache` (~100 LOC)
**Isolated LRU cache for task descriptions**.
Move:
- `_taskDescriptionCache: LRUMap<number, { description: string; timestamp: number }>`
- `_taskDescriptionMaxAge`
- `findTaskDescriptionNear(lineNumber)`
- `cacheTaskDescription(lineNumber, description)`
```typescript
export class SessionTaskCache {
private cache: LRUMap<number, { description: string; timestamp: number }>;
private maxAgeMs: number;
constructor(maxSize: number = 50, maxAgeMs: number = 30_000) { ... }
find(lineNumber: number, searchRadius: number = 50): string | null { ... }
add(lineNumber: number, description: string): void { ... }
clear(): void { ... }
}
```
### Step 4d: Keep in `session.ts` (~1,600 LOC)
The core stays together:
- PTY process management (`spawn`, `kill`, `resize`, `writeViaMux`)
- Data streaming pipeline (PTY → buffer → ANSI strip → JSON parse → events)
- Tracker initialization and event forwarding (RalphTracker, BashToolParser, TaskTracker)
- Output processing (message extraction, completion detection)
- Token tracking (status line parsing)
- State management (`toState()`, `updateState()`)
- Session lifecycle (`startInteractive`, `startShell`, `runPrompt`)
- CLI info detection (version, model, account)
---
## 5. Execution Order & Dependencies
Execute in this order to minimize risk. Each step is independently deployable.
```
Step 1: types.ts split
↓ (no runtime change, just file reorganization)
Step 2a: RalphPlanTracker extraction
↓ (independent of types split)
Step 2b: RalphFixPlanWatcher extraction
Step 2c: RalphStallDetector extraction
Step 2d: RalphStatusParser extraction
↓ (ralph-tracker.ts now ~1,800 LOC)
Step 3a: RespawnPatterns extraction
Step 3b: RespawnAdaptiveTiming extraction
Step 3c: RespawnCycleMetrics extraction
Step 3d: RespawnHealthCalculator extraction
↓ (respawn-controller.ts now ~2,200 LOC)
Step 4a: SessionCliBuilder extraction
Step 4b: SessionAutoOps extraction
Step 4c: SessionTaskCache extraction
↓ (session.ts now ~1,600 LOC)
```
**Parallelization**: Steps 1, 2a-2d, 3a-3d, and 4a-4c can be done by separate agents in parallel since they touch different files. However, within each group, sequential execution is safer.
### Risk Mitigation
- **Barrel exports**: Every split uses delegation + barrel re-export so external consumers see zero API changes
- **Event forwarding**: Sub-modules emit events, parent class forwards them — no event contract changes
- **Incremental**: Each step can be verified independently with `tsc --noEmit` + `npm run lint`
- **No test changes needed**: External API stays identical; existing tests continue to pass
---
## 6. Validation Checklist
After each step, verify:
- [ ] `tsc --noEmit` passes (no type errors)
- [ ] `npm run lint` passes (no unused imports, etc.)
- [ ] `npm run format:check` passes
- [ ] `npx vitest run test/respawn-controller.test.ts` passes (for respawn splits)
- [ ] `npx vitest run test/ralph-tracker.test.ts` passes (for ralph splits)
- [ ] `npx vitest run test/session-manager.test.ts` passes (for session splits)
- [ ] Dev server starts: `npx tsx src/index.ts web`
- [ ] Existing sessions work (create, interact, delete)
- [ ] Respawn cycle works (enable respawn, verify idle detection fires)
- [ ] No new circular dependencies: `npx madge --circular src/`
### Size Targets
| File | Before | After |
|------|--------|-------|
| `src/types.ts` | 1,443 LOC | 1 LOC (re-export barrel) |
| `src/ralph-tracker.ts` | 3,868 LOC | ~1,800 LOC |
| `src/respawn-controller.ts` | 3,611 LOC | ~2,200 LOC |
| `src/session.ts` | 2,418 LOC | ~1,600 LOC |
| **Total new files** | — | 12 files |
| **Net LOC change** | — | ~0 (refactor only) |
+738
View File
@@ -0,0 +1,738 @@
# Phase 1 Implementation Plan: Quick Wins
**Source**: `docs/code-structure-findings.md` (Phase 1 - Quick Wins section)
**Estimated effort**: 1-2 days
**Tasks**: 5 independent tasks (can be done in parallel unless noted)
---
## Safety Constraints
Before starting ANY work, read and follow these rules:
1. **Never run `npx vitest run`** (full suite) -- it kills tmux sessions. You are running inside a Codeman-managed tmux session.
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
3. **Never test on port 3000** -- the live dev server runs there. Tests use ports 3150+.
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
6. **Never kill tmux sessions** -- check `echo $CODEMAN_MUX` first.
---
## Task Dependencies
All 5 tasks are independent and can be done in parallel. However:
- Task 1 (barrel exports) is a prerequisite if you want to update import sites to use the barrel after Task 3 (consolidate EXEC_TIMEOUT_MS). The EXEC_TIMEOUT_MS consolidation creates a new export that should be added to the barrel.
- Task 2 (delete dead functions) removes functions that Task 1 would otherwise need to add to the barrel. Do Task 2 first or simultaneously with Task 1 to avoid adding exports for dead code.
**Recommended order**: Task 2 -> Task 1 -> Task 3 -> Task 4 -> Task 5
---
## Task 1: Export Missing Functions from Utils Barrel
**File**: `src/utils/index.ts`
**Time**: ~30 minutes
### Problem
The barrel file (`src/utils/index.ts`) is missing exports for several functions that are defined in util modules, forcing consumers to use deep imports or preventing usage entirely.
### Missing Exports
From `src/utils/regex-patterns.ts`:
- `createAnsiPatternFull()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
- `createAnsiPatternSimple()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
- `stripAnsi()` -- ANSI stripping utility
- `SAFE_PATH_PATTERN` -- regex for safe file paths (currently deep-imported by `schemas.ts` and `tmux-manager.ts`)
From `src/utils/token-validation.ts`:
- `validateTokenCounts()` -- token count validation (documented in CLAUDE.md)
- `validateTokensAndCost()` -- token + cost validation (documented in CLAUDE.md)
**Note**: Do NOT export `isSimilar`, `isSimilarByDistance`, `levenshteinDistance`, or `normalizePhrase` from `string-similarity.ts` -- these are dead code (see Task 2).
### Edit 1: Add missing regex-patterns exports
**File**: `src/utils/index.ts`
**Old code** (lines 13-18):
```typescript
export {
ANSI_ESCAPE_PATTERN_FULL,
ANSI_ESCAPE_PATTERN_SIMPLE,
TOKEN_PATTERN,
SPINNER_PATTERN,
} from './regex-patterns.js';
```
**New code**:
```typescript
export {
ANSI_ESCAPE_PATTERN_FULL,
ANSI_ESCAPE_PATTERN_SIMPLE,
TOKEN_PATTERN,
SPINNER_PATTERN,
createAnsiPatternFull,
createAnsiPatternSimple,
stripAnsi,
SAFE_PATH_PATTERN,
} from './regex-patterns.js';
```
### Edit 2: Add missing token-validation exports
**File**: `src/utils/index.ts`
**Old code** (line 19):
```typescript
export { MAX_SESSION_TOKENS } from './token-validation.js';
```
**New code**:
```typescript
export { MAX_SESSION_TOKENS, validateTokenCounts, validateTokensAndCost } from './token-validation.js';
```
### Optional follow-up: Update deep imports to use barrel
These files currently deep-import `SAFE_PATH_PATTERN` and could be updated to use the barrel instead:
- `src/web/schemas.ts` line 11: `import { SAFE_PATH_PATTERN } from '../utils/regex-patterns.js';` could become `import { SAFE_PATH_PATTERN } from '../utils/index.js';`
- `src/tmux-manager.ts` line 44: `import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';` could become part of existing barrel import
This is a low-priority cosmetic change. The barrel export itself is the important fix.
### Verification
```bash
tsc --noEmit
npm run lint
```
---
## Task 2: Delete Dead Utility Functions
**File**: `src/utils/string-similarity.ts`
**Time**: ~15 minutes
### Problem
Four exported functions in `string-similarity.ts` are never imported anywhere in the codebase:
- `levenshteinDistance()` (lines 27-69)
- `isSimilar()` (lines 106-108)
- `isSimilarByDistance()` (lines 123-125)
- `normalizePhrase()` (lines 139-144)
Only three functions are actually used (all by `ralph-tracker.ts` via the barrel):
- `stringSimilarity()` -- uses `levenshteinDistance()` internally
- `fuzzyPhraseMatch()` -- uses `normalizePhrase()` and `isSimilarByDistance()` internally
- `todoContentHash()`
### Strategy
`levenshteinDistance()` is called by `stringSimilarity()`, and `normalizePhrase()` and `isSimilarByDistance()` are called by `fuzzyPhraseMatch()`. So they cannot be deleted -- they just need to be un-exported (made private to the module).
`isSimilar()` is truly dead -- not called by anything. Delete it entirely.
### Edit 1: Remove `export` from `levenshteinDistance`
**File**: `src/utils/string-similarity.ts`
**Old code** (line 27):
```typescript
export function levenshteinDistance(a: string, b: string): number {
```
**New code**:
```typescript
function levenshteinDistance(a: string, b: string): number {
```
### Edit 2: Delete `isSimilar` function entirely
**File**: `src/utils/string-similarity.ts`
**Old code** (lines 94-108):
```typescript
/**
* Check if two strings are similar within a given threshold.
*
* @param a - First string
* @param b - Second string
* @param threshold - Minimum similarity ratio (default: 0.85 = 85% similar)
* @returns True if similarity >= threshold
*
* @example
* isSimilar('COMPLETE', 'COMPLET', 0.85) // true (87.5% similar)
* isSimilar('COMPLETE', 'DONE', 0.85) // false (0% similar)
*/
export function isSimilar(a: string, b: string, threshold = 0.85): boolean {
return stringSimilarity(a, b) >= threshold;
}
```
**New code**: (delete entirely -- replace with empty string)
### Edit 3: Remove `export` from `isSimilarByDistance`
**File**: `src/utils/string-similarity.ts`
**Old code** (line 123):
```typescript
export function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
```
**New code**:
```typescript
function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
```
### Edit 4: Remove `export` from `normalizePhrase`
**File**: `src/utils/string-similarity.ts`
**Old code** (line 139):
```typescript
export function normalizePhrase(phrase: string): string {
```
**New code**:
```typescript
function normalizePhrase(phrase: string): string {
```
### Verification
```bash
tsc --noEmit
npx vitest run test/string-utilities.test.ts
npm run lint
```
Note: If `test/string-utilities.test.ts` imports any of the now-unexported functions, those test imports will fail. Check the test file and remove tests for `isSimilar` (deleted) and update any direct tests for `levenshteinDistance`, `isSimilarByDistance`, `normalizePhrase` to test them indirectly through the public API (`stringSimilarity`, `fuzzyPhraseMatch`), or remove those tests.
---
## Task 3: Consolidate Duplicated `EXEC_TIMEOUT_MS` Constant
**Files**:
- `src/utils/claude-cli-resolver.ts` (line 17)
- `src/utils/opencode-cli-resolver.ts` (line 16)
- `src/tmux-manager.ts` (line 63) -- also has its own copy
**Time**: ~15 minutes
### Problem
`EXEC_TIMEOUT_MS = 5000` is defined identically in three files. Changes need to happen in all three places.
### Strategy
Create a shared constant and export it. The natural home is a new config file since the existing config files (`buffer-limits.ts`, `map-limits.ts`) follow this pattern. However, to keep it minimal, we can add it to an existing config file or create a small one.
**Recommended approach**: Add to `src/config/timing-config.ts` (new file) as a single constant. This file can grow later in Phase 6 to hold other timing constants.
Alternatively, the simplest approach: export from one of the existing utils and import in the others. Since both CLI resolvers are in `src/utils/`, the cleanest approach is to put it in a shared location.
### Option A: Add to existing config (simpler)
Create `src/config/exec-timeout.ts`:
**New file**: `src/config/exec-timeout.ts`
```typescript
/**
* Timeout for child process exec commands (e.g., `which claude`, `which opencode`, tmux commands).
* Used across CLI resolvers and tmux manager.
*/
export const EXEC_TIMEOUT_MS = 5000;
```
### Edit 1: Update `claude-cli-resolver.ts`
**File**: `src/utils/claude-cli-resolver.ts`
**Old code** (lines 11-17):
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { delimiter, dirname, join } from 'node:path';
import { homedir } from 'node:os';
/** Timeout for exec commands (5 seconds) */
const EXEC_TIMEOUT_MS = 5000;
```
**New code**:
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { delimiter, dirname, join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
```
### Edit 2: Update `opencode-cli-resolver.ts`
**File**: `src/utils/opencode-cli-resolver.ts`
**Old code** (lines 10-16):
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { homedir } from 'node:os';
/** Timeout for exec commands (5 seconds) */
const EXEC_TIMEOUT_MS = 5000;
```
**New code**:
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
```
### Edit 3: Update `tmux-manager.ts`
**File**: `src/tmux-manager.ts`
**Old code** (line 63):
```typescript
const EXEC_TIMEOUT_MS = 5000;
```
**New code**:
```typescript
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
```
Note: `tmux-manager.ts` already has many imports at the top of the file. Add this import near the other local imports (around lines 43-56). The `const EXEC_TIMEOUT_MS = 5000;` on line 63 should be deleted entirely (replaced with the import).
### Verification
```bash
tsc --noEmit
npm run lint
```
---
## Task 4: Add `z.infer` to Zod Schemas
**Files**:
- `src/web/schemas.ts` (add type exports)
- `src/types.ts` (replace manual interfaces with `z.infer` re-exports where applicable)
**Time**: ~2 hours
### Problem
All 30+ Zod schemas in `schemas.ts` define validation rules, but zero use `z.infer` to derive TypeScript types. Instead, `types.ts` manually duplicates interfaces that match the schemas. When a schema changes, the type must be manually updated too.
### Strategy
Add `z.infer` type exports to `schemas.ts` for each exported schema. This creates derived types as the single source of truth. For schemas that have corresponding manual interfaces in `types.ts`, the manual interface can be replaced with a re-export of the inferred type.
**Important**: Not all schemas have matching interfaces in `types.ts`. The `RespawnConfig` interface in `types.ts` (line 395) has all required fields, while `RespawnConfigSchema` has all optional fields (it's for partial updates). These are NOT the same type and should NOT be unified.
### Edit 1: Add inferred type exports to `schemas.ts`
**File**: `src/web/schemas.ts`
After each schema definition, add a corresponding type export. Add the following lines at the **end of the file** (after line 509):
**Old code** (end of file, lines 506-509):
```typescript
.optional(),
});
```
Wait -- the end of file is actually at line 509 after the `RalphLoopStartSchema`. Add the type exports after the last schema:
**Append to end of file** `src/web/schemas.ts`:
```typescript
// ========== Inferred Types ==========
// Derive TypeScript types from Zod schemas (single source of truth)
export type CreateSessionInput = z.infer<typeof CreateSessionSchema>;
export type RunPromptInput = z.infer<typeof RunPromptSchema>;
export type ResizeInput = z.infer<typeof ResizeSchema>;
export type CreateCaseInput = z.infer<typeof CreateCaseSchema>;
export type QuickStartInput = z.infer<typeof QuickStartSchema>;
export type HookEventInput = z.infer<typeof HookEventSchema>;
export type RespawnConfigInput = z.infer<typeof RespawnConfigSchema>;
export type ConfigUpdateInput = z.infer<typeof ConfigUpdateSchema>;
export type SettingsUpdateInput = z.infer<typeof SettingsUpdateSchema>;
export type SessionInputWithLimitInput = z.infer<typeof SessionInputWithLimitSchema>;
export type SessionNameInput = z.infer<typeof SessionNameSchema>;
export type SessionColorInput = z.infer<typeof SessionColorSchema>;
export type RalphConfigInput = z.infer<typeof RalphConfigSchema>;
export type FixPlanImportInput = z.infer<typeof FixPlanImportSchema>;
export type RalphPromptWriteInput = z.infer<typeof RalphPromptWriteSchema>;
export type AutoClearInput = z.infer<typeof AutoClearSchema>;
export type AutoCompactInput = z.infer<typeof AutoCompactSchema>;
export type ImageWatcherInput = z.infer<typeof ImageWatcherSchema>;
export type FlickerFilterInput = z.infer<typeof FlickerFilterSchema>;
export type QuickRunInput = z.infer<typeof QuickRunSchema>;
export type ScheduledRunInput = z.infer<typeof ScheduledRunSchema>;
export type LinkCaseInput = z.infer<typeof LinkCaseSchema>;
export type GeneratePlanInput = z.infer<typeof GeneratePlanSchema>;
export type GeneratePlanDetailedInput = z.infer<typeof GeneratePlanDetailedSchema>;
export type CancelPlanInput = z.infer<typeof CancelPlanSchema>;
export type PlanTaskUpdateInput = z.infer<typeof PlanTaskUpdateSchema>;
export type PlanTaskAddInput = z.infer<typeof PlanTaskAddSchema>;
export type CpuLimitInput = z.infer<typeof CpuLimitSchema>;
export type SubagentWindowStatesInput = z.infer<typeof SubagentWindowStatesSchema>;
export type SubagentParentMapInput = z.infer<typeof SubagentParentMapSchema>;
export type InteractiveRespawnInput = z.infer<typeof InteractiveRespawnSchema>;
export type RespawnEnableInput = z.infer<typeof RespawnEnableSchema>;
export type PushSubscribeInput = z.infer<typeof PushSubscribeSchema>;
export type PushPreferencesUpdateInput = z.infer<typeof PushPreferencesUpdateSchema>;
export type RalphLoopStartInput = z.infer<typeof RalphLoopStartSchema>;
```
### What NOT to do
Do NOT replace the `RespawnConfig` interface in `types.ts` with `z.infer<typeof RespawnConfigSchema>`. The schema has all optional fields (for partial config updates), but the interface has required fields (for the full config object). These are intentionally different shapes.
Similarly, do NOT try to unify every interface in `types.ts` with a schema -- most interfaces in `types.ts` represent internal domain objects (SessionState, TaskState, etc.) that have no corresponding Zod schema. The schemas only exist for API request validation.
### Future opportunity
In a future phase, route handlers in `server.ts` can use these inferred types for request body typing:
```typescript
const body = CreateSessionSchema.parse(request.body) as CreateSessionInput;
```
This task only adds the type exports. Migrating route handlers to use them is out of scope.
### Verification
```bash
tsc --noEmit
npm run lint
npm run format:check
```
---
## Task 5: Fix Weak `not.toThrow()` Tests with Behavioral Assertions
**Files**:
- `test/task-tracker.test.ts` -- 6 instances
- `test/image-watcher.test.ts` -- 1 instance
- `test/task-queue.test.ts` -- 1 instance
- `test/hooks-config.test.ts` -- 1 instance
- `test/session-manager.test.ts` -- 1 instance
**Time**: ~1 hour
### Problem
10 tests only assert `not.toThrow()` without verifying the actual defensive behavior. These tests prove the code doesn't crash but don't verify it does the right thing.
### Fix Strategy
After each `not.toThrow()`, add a behavioral assertion that verifies the state is correct (e.g., no tasks were created, no side effects occurred).
### Edit 1: `task-tracker.test.ts` -- null message (line 566)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle null message', () => {
expect(() => tracker.processMessage(null)).not.toThrow();
});
```
**New code**:
```typescript
it('should handle null message', () => {
expect(() => tracker.processMessage(null)).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
expect(tracker.getRunningCount()).toBe(0);
});
```
### Edit 2: `task-tracker.test.ts` -- message without content (line 569-571)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle message without content', () => {
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
});
```
**New code**:
```typescript
it('should handle message without content', () => {
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
});
```
### Edit 3: `task-tracker.test.ts` -- empty content array (line 573-575)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle empty content array', () => {
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
});
```
**New code**:
```typescript
it('should handle empty content array', () => {
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
});
```
### Edit 4: `task-tracker.test.ts` -- tool_result for unknown task (lines 577-590)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle tool_result for unknown task', () => {
expect(() => {
tracker.processMessage({
message: {
content: [{
type: 'tool_result',
tool_use_id: 'unknown-task',
is_error: false,
content: 'Done',
}],
},
});
}).not.toThrow();
});
```
**New code**:
```typescript
it('should handle tool_result for unknown task', () => {
expect(() => {
tracker.processMessage({
message: {
content: [{
type: 'tool_result',
tool_use_id: 'unknown-task',
is_error: false,
content: 'Done',
}],
},
});
}).not.toThrow();
expect(tracker.getTask('unknown-task')).toBeUndefined();
expect(tracker.getAllTasks().size).toBe(0);
});
```
### Edit 5: `task-tracker.test.ts` -- empty terminal output (lines 592-595)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle empty terminal output', () => {
expect(() => tracker.processTerminalOutput('')).not.toThrow();
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
});
```
**New code**:
```typescript
it('should handle empty terminal output', () => {
expect(() => tracker.processTerminalOutput('')).not.toThrow();
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
expect(tracker.getRunningCount()).toBe(0);
});
```
### Edit 6: `image-watcher.test.ts` -- unwatchSession for non-watched session (line 123)
**File**: `test/image-watcher.test.ts`
**Old code**:
```typescript
it('should be safe to call for non-watched session', () => {
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
});
```
**New code**:
```typescript
it('should be safe to call for non-watched session', () => {
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
expect(watcher.getWatchedSessions()).toHaveLength(0);
});
```
### Edit 7: `task-queue.test.ts` -- dependencies on non-existent tasks (lines 538-542)
**File**: `test/task-queue.test.ts`
**Old code**:
```typescript
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
// Dependencies on non-existent tasks are valid - they just won't be satisfied
expect(() => {
queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
}).not.toThrow();
});
```
**New code**:
```typescript
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
// Dependencies on non-existent tasks are valid - they just won't be satisfied
let task: ReturnType<typeof queue.addTask> | undefined;
expect(() => {
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
}).not.toThrow();
expect(task).toBeDefined();
expect(task!.dependencies).toEqual(['non-existent-id']);
// Task should be pending but blocked (dependency unsatisfied)
expect(queue.next()?.prompt).toBeUndefined();
});
```
Wait -- `queue.next()` returns `null` when no next task is available (all blocked). Let me adjust:
**New code** (corrected):
```typescript
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
// Dependencies on non-existent tasks are valid - they just won't be satisfied
let task: ReturnType<typeof queue.addTask> | undefined;
expect(() => {
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
}).not.toThrow();
expect(task).toBeDefined();
expect(task!.dependencies).toEqual(['non-existent-id']);
// Task exists but is blocked (dependency unsatisfied), so next() skips it
expect(queue.getAllTasks()).toHaveLength(1);
expect(queue.next()).toBeNull();
});
```
### Edit 8: `hooks-config.test.ts` -- valid JSON check (line 129)
**File**: `test/hooks-config.test.ts`
**Old code**:
```typescript
it('should write valid JSON', () => {
writeHooksConfig(testDir);
const settingsPath = join(testDir, '.claude', 'settings.local.json');
const content = readFileSync(settingsPath, 'utf-8');
expect(() => JSON.parse(content)).not.toThrow();
});
```
**New code**:
```typescript
it('should write valid JSON', () => {
writeHooksConfig(testDir);
const settingsPath = join(testDir, '.claude', 'settings.local.json');
const content = readFileSync(settingsPath, 'utf-8');
const parsed = JSON.parse(content);
expect(parsed).toBeDefined();
expect(typeof parsed).toBe('object');
expect(parsed.hooks).toBeDefined();
});
```
### Edit 9: `session-manager.test.ts` -- stopSession for non-existent (line 216)
**File**: `test/session-manager.test.ts`
**Old code**:
```typescript
it('should handle non-existent session gracefully', async () => {
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
});
```
**New code**:
```typescript
it('should handle non-existent session gracefully', async () => {
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
expect(manager.getSessionCount()).toBe(0);
});
```
### Verification
Run each test file individually:
```bash
npx vitest run test/task-tracker.test.ts
npx vitest run test/image-watcher.test.ts
npx vitest run test/task-queue.test.ts
npx vitest run test/hooks-config.test.ts
npx vitest run test/session-manager.test.ts
```
**Important**: `hooks-config.test.ts` and `session-manager.test.ts` spawn real servers on ports 3130-3131. Only run them if you are NOT running other tests that use those ports.
---
## Final Verification Checklist
After all 5 tasks are complete, run the following in order:
```bash
# 1. TypeScript type checking
tsc --noEmit
# 2. Linting
npm run lint
# 3. Formatting
npm run format:check
# 4. Run affected test files individually (NOT the full suite)
npx vitest run test/string-utilities.test.ts
npx vitest run test/task-tracker.test.ts
npx vitest run test/image-watcher.test.ts
npx vitest run test/task-queue.test.ts
npx vitest run test/session-manager.test.ts
npx vitest run test/hooks-config.test.ts
```
If any formatting issues arise, fix with:
```bash
npm run format
```
If any lint issues arise, fix with:
```bash
npm run lint:fix
```
### Summary of Changes
| Task | Files Modified | Files Created |
|------|---------------|---------------|
| 1. Barrel exports | `src/utils/index.ts` | -- |
| 2. Dead functions | `src/utils/string-similarity.ts` | -- |
| 3. EXEC_TIMEOUT_MS | `src/utils/claude-cli-resolver.ts`, `src/utils/opencode-cli-resolver.ts`, `src/tmux-manager.ts` | `src/config/exec-timeout.ts` |
| 4. z.infer types | `src/web/schemas.ts` | -- |
| 5. Weak tests | `test/task-tracker.test.ts`, `test/image-watcher.test.ts`, `test/task-queue.test.ts`, `test/hooks-config.test.ts`, `test/session-manager.test.ts` | -- |
**Total files modified**: 10
**Total files created**: 1
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+689
View File
@@ -0,0 +1,689 @@
# Phase 6 Implementation Plan: Config Consolidation
**Source**: `docs/code-structure-findings.md` (Phase 6 — Config Consolidation)
**Estimated effort**: 1 day
**Tasks**: 8 tasks with dependencies (see dependency graph below)
---
## Safety Constraints
Before starting ANY work, read and follow these rules:
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
7. **Verify the dev server starts**: After each task, run `npx tsx src/index.ts web --port 3099 &` on a non-production port, confirm `curl -s http://localhost:3099/api/status | jq .status` returns `"ok"`, then kill the background process.
---
## Goal
Consolidate ~70 scattered numeric constants from 15+ source files into 6 new domain-focused config files, eliminating cross-file duplicates (including a 5x-duplicated AI model string) and making all tuning knobs discoverable in `src/config/`.
**Non-goal**: Moving every constant. Module-internal implementation details (like regex patterns, algorithm-specific magic numbers, or constants only used once in deeply coupled logic) stay where they are. The goal is discoverability of operational tuning knobs, not mechanical relocation.
---
## Design Decisions
### What gets centralized (and why)
Constants are candidates for centralization when they meet **any** of these criteria:
1. **Duplicated across files** — DRY violation (e.g., `STATS_COLLECTION_INTERVAL_MS` in `server.ts` and `mux-routes.ts`, AI model string in 5 files)
2. **Operational tuning knobs** — values an operator might want to adjust for performance, security, or behavior without understanding the implementation (e.g., SSE health check interval, auth session TTL, rate limits)
3. **Cross-cutting concerns** — values that establish system-wide contracts (e.g., max terminal dimensions used by both server routes and frontend)
### What stays in place (and why)
Constants that are **internal implementation details** of a single module stay where they are:
- **Algorithm parameters** — `TODO_SIMILARITY_THRESHOLD`, `adaptiveCompletionConfirmMs`, confidence weights. These are meaningless without understanding the algorithm.
- **Display/UI formatting** — `TEXT_PREVIEW_LENGTH`, `SMART_TITLE_MAX_LENGTH`, `COMMAND_DISPLAY_LENGTH` in `subagent-watcher.ts`. Only used locally, tightly coupled to rendering logic.
- **Module-internal timing** — `LINE_BUFFER_FLUSH_INTERVAL` in `session.ts`, `AI_CHECK_POLL_INTERVAL` in `ai-checker-base.ts`. Internal implementation of specific features.
- **Frontend constants** — `constants.js` already centralizes frontend values well. Don't mix frontend and backend config.
- **Respawn `DEFAULT_CONFIG`** — these are user-configurable defaults for the respawn config interface, not system constants. They live properly in `respawn-controller.ts`. The AI model/context defaults within it are replaced with imports from the new `ai-defaults.ts` (Task 5).
- **Session auto-ops thresholds** — `AUTO_RETRY_DELAY_MS`, `COMPACT_COOLDOWN_MS`, etc. in `session-auto-ops.ts` are internal to that module's retry logic and already well-documented in place.
### File organization: domain-based, not category-based
A single `timing-config.ts` with 70 unrelated timing values would be worse than the current state — developers would need to grep it just like they grep the whole codebase now. Instead, constants are grouped by **the system they configure**:
| New File | Domain | Developer Question It Answers |
|----------|--------|-------------------------------|
| `server-timing.ts` | Web server performance | "How do I tune SSE batching / terminal throughput?" |
| `auth-config.ts` | Authentication & security | "What are the rate limits and session TTLs?" |
| `tunnel-config.ts` | QR auth & Cloudflare tunnel | "What are the QR token rotation parameters?" |
| `terminal-limits.ts` | Terminal dimensions & input | "What are the max cols/rows/input size?" |
| `ai-defaults.ts` | AI checker model & context | "What model do the AI checkers use? What's the context limit?" |
| `team-config.ts` | Agent Teams polling & caching | "How often does team polling run? What are the cache limits?" |
---
## Task Dependencies
```
Task 1 (server-timing.ts)
Task 2 (auth-config.ts)
Task 3 (tunnel-config.ts)
Task 4 (terminal-limits.ts)
Task 5 (ai-defaults.ts)
Task 6 (team-config.ts)
└──> Task 7 (Fix remaining duplicates)
└──> Task 8 (Update CLAUDE.md + final verification)
```
**Tasks 1–6** are independent and can run in parallel.
**Task 7** depends on Tasks 1–6 (needs the new config files to exist).
**Task 8** depends on Task 7.
---
## Task 1: Create `src/config/server-timing.ts`
**Estimated effort**: 30 minutes
**Files created**: `src/config/server-timing.ts`
**Files modified**: `src/web/server.ts`, `src/web/routes/mux-routes.ts`
### Constants to extract from `src/web/server.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `TERMINAL_BATCH_INTERVAL` | `16` | Terminal data batching interval (60fps) |
| `TASK_UPDATE_BATCH_INTERVAL` | `100` | Task event batching interval (ms) |
| `STATE_UPDATE_DEBOUNCE_INTERVAL` | `500` | State persistence debounce (ms) |
| `SESSIONS_LIST_CACHE_TTL` | `1000` | Sessions list cache TTL (ms) |
| `SCHEDULED_CLEANUP_INTERVAL` | `300000` | Scheduled runs cleanup check (5 min) |
| `SCHEDULED_RUN_MAX_AGE` | `3600000` | Completed scheduled run max age (1 hour) |
| `SSE_HEALTH_CHECK_INTERVAL` | `30000` | SSE client health check (30s) |
| `SESSION_LIMIT_WAIT_MS` | `5000` | Session limit retry wait (5s) |
| `ITERATION_PAUSE_MS` | `2000` | Scheduled run iteration pause (2s) |
| `BATCH_FLUSH_THRESHOLD` | `32768` | Terminal batch immediate flush threshold (32KB) |
| `STATS_COLLECTION_INTERVAL_MS` | `2000` | Mux stats collection interval (2s) |
### Implementation
1. Create `src/config/server-timing.ts` with all 11 constants, preserving existing JSDoc comments.
2. In `src/web/server.ts`: Remove the 11 local constant declarations (lines ~92–121). Add `import { TERMINAL_BATCH_INTERVAL, ... } from '../config/server-timing.js'`.
3. In `src/web/routes/mux-routes.ts`: Remove the duplicate `STATS_COLLECTION_INTERVAL_MS` (line 10) and its comment. Add `import { STATS_COLLECTION_INTERVAL_MS } from '../../config/server-timing.js'`. This fixes a **duplicate constant** (finding #10).
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Web server performance and scheduling constants.
*
* Controls terminal batching throughput, SSE health checking,
* state persistence debouncing, and scheduled run timing.
*
* @module config/server-timing
*/
// ============================================================================
// Terminal & SSE Performance
// ============================================================================
/** Terminal data batching interval — targets 60fps (ms) */
export const TERMINAL_BATCH_INTERVAL = 16;
/** Immediate flush threshold for terminal batches (bytes).
* Set high (32KB) to allow effective batching; avg Ink events are ~14KB. */
export const BATCH_FLUSH_THRESHOLD = 32 * 1024;
/** Task event batching interval (ms) */
export const TASK_UPDATE_BATCH_INTERVAL = 100;
/** SSE client health check interval (ms) */
export const SSE_HEALTH_CHECK_INTERVAL = 30 * 1000;
// ============================================================================
// State Persistence
// ============================================================================
/** State update debounce — batches expensive toDetailedState() calls (ms) */
export const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
/** Sessions list cache TTL — avoids re-serializing on every SSE init (ms) */
export const SESSIONS_LIST_CACHE_TTL = 1000;
// ============================================================================
// Scheduled Runs
// ============================================================================
/** Scheduled runs cleanup check interval (ms) */
export const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
/** Completed scheduled run max age before cleanup (ms) */
export const SCHEDULED_RUN_MAX_AGE = 60 * 60 * 1000;
/** Session limit retry wait before retrying (ms) */
export const SESSION_LIMIT_WAIT_MS = 5000;
/** Pause between scheduled run iterations (ms) */
export const ITERATION_PAUSE_MS = 2000;
// ============================================================================
// Mux Stats
// ============================================================================
/** Mux stats collection interval (ms) */
export const STATS_COLLECTION_INTERVAL_MS = 2000;
```
### Verification
```bash
tsc --noEmit
npx tsx src/index.ts web --port 3099 &
curl -s http://localhost:3099/api/status | jq .status # "ok"
kill %1
```
---
## Task 2: Create `src/config/auth-config.ts`
**Estimated effort**: 20 minutes
**Files created**: `src/config/auth-config.ts`
**Files modified**: `src/web/middleware/auth.ts`, `src/hooks-config.ts`
### Constants to extract from `src/web/middleware/auth.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `AUTH_SESSION_TTL_MS` | `86400000` | Auth session cookie TTL (24h) |
| `MAX_AUTH_SESSIONS` | `100` | Max concurrent auth sessions |
| `AUTH_FAILURE_MAX` | `10` | Max failed auth attempts per IP |
| `AUTH_FAILURE_WINDOW_MS` | `900000` | Failed auth tracking window (15 min) |
### Constants to extract from `src/hooks-config.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `HOOK_TIMEOUT_MS` | `10000` | Timeout for Claude Code hook commands |
The `timeout: 10000` value is hardcoded 6 times in `hooks-config.ts` as inline literals. Extract to a single named constant.
### Implementation
1. Create `src/config/auth-config.ts` with the 5 constants.
2. In `src/web/middleware/auth.ts`: Remove the 4 local constant declarations (lines 17–25). Add import from `../../config/auth-config.js`. Keep `AUTH_COOKIE_NAME` in place — it's a string identifier, not a tunable numeric constant.
3. In `src/hooks-config.ts`: Replace all 6 inline `timeout: 10000` occurrences with `timeout: HOOK_TIMEOUT_MS`. Add import from `./config/auth-config.js`.
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Authentication, rate limiting, and hook security constants.
*
* Controls auth session lifecycle, brute-force protection,
* and Claude Code hook timeouts.
*
* @module config/auth-config
*/
// ============================================================================
// Session Cookies
// ============================================================================
/** Auth session cookie TTL — matches autonomous run length (ms) */
export const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
/** Max concurrent auth sessions per server */
export const MAX_AUTH_SESSIONS = 100;
// ============================================================================
// Rate Limiting
// ============================================================================
/** Max failed auth attempts per IP before 429 rejection */
export const AUTH_FAILURE_MAX = 10;
/** Failed auth attempt tracking window (ms) */
export const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
// ============================================================================
// Hooks
// ============================================================================
/** Timeout for Claude Code hook curl commands (ms) */
export const HOOK_TIMEOUT_MS = 10000;
```
### Verification
```bash
tsc --noEmit
npm run lint
```
---
## Task 3: Create `src/config/tunnel-config.ts`
**Estimated effort**: 20 minutes
**Files created**: `src/config/tunnel-config.ts`
**Files modified**: `src/tunnel-manager.ts`
### Constants to extract from `src/tunnel-manager.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `QR_TOKEN_TTL_MS` | `60000` | QR token auto-rotation interval (60s) |
| `QR_TOKEN_GRACE_MS` | `90000` | Grace period for previous token (90s) |
| `SHORT_CODE_LENGTH` | `6` | Length of QR short code |
| `QR_RATE_LIMIT_MAX` | `30` | Global QR attempt rate limit |
| `QR_RATE_LIMIT_WINDOW_MS` | `60000` | QR rate limit reset window (60s) |
| `URL_TIMEOUT_MS` | `30000` | Cloudflared URL fetch timeout (30s) |
| `RESTART_DELAY_MS` | `5000` | Tunnel restart delay after crash (5s) |
| `FORCE_KILL_MS` | `5000` | SIGTERM → SIGKILL escalation timeout (5s) |
### Implementation
1. Create `src/config/tunnel-config.ts` with all 8 constants.
2. In `src/tunnel-manager.ts`: Remove the 8 local constant declarations (lines ~39–75). Add `import { QR_TOKEN_TTL_MS, ... } from './config/tunnel-config.js'`.
3. Keep the `TUNNEL_URL_REGEX` in `tunnel-manager.ts` — it's a parsing detail, not a tuning knob.
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Cloudflare tunnel and QR authentication constants.
*
* Controls QR token rotation timing, rate limiting,
* and tunnel process lifecycle.
*
* @module config/tunnel-config
*/
// ============================================================================
// QR Token Rotation
// ============================================================================
/** QR token auto-rotation interval (ms) */
export const QR_TOKEN_TTL_MS = 60_000;
/** Grace period — previous token still valid during rotation (ms) */
export const QR_TOKEN_GRACE_MS = 90_000;
/** Length of the short code in QR URL path (chars) */
export const SHORT_CODE_LENGTH = 6;
// ============================================================================
// QR Rate Limiting
// ============================================================================
/** Global rate limit for QR auth attempts across all IPs */
export const QR_RATE_LIMIT_MAX = 30;
/** QR rate limit reset window (ms) */
export const QR_RATE_LIMIT_WINDOW_MS = 60_000;
// ============================================================================
// Tunnel Process Lifecycle
// ============================================================================
/** Max time to wait for cloudflared URL before timeout (ms) */
export const URL_TIMEOUT_MS = 30_000;
/** Restart delay after unexpected tunnel exit (ms) */
export const RESTART_DELAY_MS = 5_000;
/** SIGTERM → SIGKILL escalation timeout (ms) */
export const FORCE_KILL_MS = 5_000;
```
### Verification
```bash
tsc --noEmit
```
---
## Task 4: Create `src/config/terminal-limits.ts`
**Estimated effort**: 20 minutes
**Files created**: `src/config/terminal-limits.ts`
**Files modified**: `src/web/routes/session-routes.ts`
### Constants to extract from `src/web/routes/session-routes.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `MAX_INPUT_LENGTH` | `65536` | Max input length per request (64KB) |
| `MAX_TERMINAL_COLS` | `500` | Max terminal columns |
| `MAX_TERMINAL_ROWS` | `200` | Max terminal rows |
| `MAX_SESSION_NAME_LENGTH` | `128` | Max session name length (chars) |
### Why a separate file instead of adding to `buffer-limits.ts`
`buffer-limits.ts` covers memory buffer sizes (2MB terminal, 1MB text). These constants are **validation limits** for API inputs — different concern. A terminal resize request must not exceed `MAX_TERMINAL_COLS`; this has nothing to do with buffer trimming.
### Implementation
1. Create `src/config/terminal-limits.ts` with all 4 constants.
2. In `src/web/routes/session-routes.ts`: Remove the 4 local constant declarations (lines 45–48). Add `import { MAX_INPUT_LENGTH, MAX_TERMINAL_COLS, MAX_TERMINAL_ROWS, MAX_SESSION_NAME_LENGTH } from '../../config/terminal-limits.js'`.
3. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Terminal dimension and input validation limits.
*
* Used by API routes to validate resize, input, and session
* creation requests. Separate from buffer-limits.ts which
* controls memory buffer sizes.
*
* @module config/terminal-limits
*/
/** Max input length per API request (bytes) */
export const MAX_INPUT_LENGTH = 64 * 1024;
/** Max terminal columns for resize requests */
export const MAX_TERMINAL_COLS = 500;
/** Max terminal rows for resize requests */
export const MAX_TERMINAL_ROWS = 200;
/** Max session name length (chars) */
export const MAX_SESSION_NAME_LENGTH = 128;
```
### Verification
```bash
tsc --noEmit
```
---
## Task 5: Create `src/config/ai-defaults.ts`
**Estimated effort**: 30 minutes
**Files created**: `src/config/ai-defaults.ts`
**Files modified**: `src/respawn-controller.ts`, `src/ai-idle-checker.ts`, `src/ai-plan-checker.ts`, `src/web/routes/respawn-routes.ts`
### Problem: AI model string duplicated 5 times
The model identifier `'claude-opus-4-5-20251101'` appears in 5 places across 4 files. When the model changes, all 5 must be updated — a guaranteed source of bugs. The context limits (`16000`, `8000`) are similarly scattered across 3 files each.
| Constant | Current Value | Duplicated In |
|----------|---------------|---------------|
| `AI_CHECK_MODEL` | `'claude-opus-4-5-20251101'` | `respawn-controller.ts` (×2: idle + plan), `ai-idle-checker.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` (×2: idle + plan) |
| `AI_IDLE_CHECK_MAX_CONTEXT` | `16000` | `respawn-controller.ts`, `ai-idle-checker.ts`, `respawn-routes.ts` |
| `AI_PLAN_CHECK_MAX_CONTEXT` | `8000` | `respawn-controller.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` |
### Implementation
1. Create `src/config/ai-defaults.ts` with the 3 constants.
2. In `src/respawn-controller.ts` `DEFAULT_CONFIG` (line 538): Replace `aiIdleCheckModel: 'claude-opus-4-5-20251101'` with `aiIdleCheckModel: AI_CHECK_MODEL`, `aiIdleCheckMaxContext: 16000` with `aiIdleCheckMaxContext: AI_IDLE_CHECK_MAX_CONTEXT`, `aiPlanCheckModel: 'claude-opus-4-5-20251101'` with `aiPlanCheckModel: AI_CHECK_MODEL`, `aiPlanCheckMaxContext: 8000` with `aiPlanCheckMaxContext: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
3. In `src/ai-idle-checker.ts` `DEFAULT_AI_CHECK_CONFIG` (line 46): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 16000` with `maxContextChars: AI_IDLE_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
4. In `src/ai-plan-checker.ts` `DEFAULT_PLAN_CHECK_CONFIG` (line 45): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 8000` with `maxContextChars: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
5. In `src/web/routes/respawn-routes.ts` config merge block (lines 173–179): Replace all 4 inline fallback values with imports from `../../config/ai-defaults.js`.
6. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Default model and context limits for AI-powered checkers.
*
* Centralizes the AI model identifier and context window sizes used by
* the idle checker, plan checker, respawn controller defaults, and
* respawn route fallbacks. Change the model here when upgrading.
*
* @module config/ai-defaults
*/
/** Default model for AI idle and plan checkers */
export const AI_CHECK_MODEL = 'claude-opus-4-5-20251101';
/** Max context chars for idle checker (~4k tokens) */
export const AI_IDLE_CHECK_MAX_CONTEXT = 16000;
/** Max context chars for plan checker (~2k tokens, plan mode UI is compact) */
export const AI_PLAN_CHECK_MAX_CONTEXT = 8000;
```
### Verification
```bash
tsc --noEmit
# Verify no remaining hardcoded model strings
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
```
---
## Task 6: Create `src/config/team-config.ts`
**Estimated effort**: 15 minutes
**Files created**: `src/config/team-config.ts`
**Files modified**: `src/team-watcher.ts`
### Constants to extract from `src/team-watcher.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `TEAM_POLL_INTERVAL_MS` | `30000` | Team directory poll interval (30s) |
| `MAX_CACHED_TEAMS` | `50` | LRU cache size for team configs |
| `MAX_CACHED_TASKS` | `200` | LRU cache size for team tasks + inboxes |
### Why centralize these
Team polling frequency and cache sizes are operational knobs that affect both performance (polling too often wastes CPU) and responsiveness (polling too rarely means stale team state in the UI). They're also the kind of values a developer tuning for a large team deployment would want to find quickly. `MAX_CACHED_TASKS` is used for both the task cache and inbox cache — worth documenting.
### Implementation
1. Create `src/config/team-config.ts` with the 3 constants.
2. In `src/team-watcher.ts`: Remove the 3 local constants (lines 23–25). Add `import { TEAM_POLL_INTERVAL_MS, MAX_CACHED_TEAMS, MAX_CACHED_TASKS } from './config/team-config.js'`. Note: rename `POLL_INTERVAL_MS` → `TEAM_POLL_INTERVAL_MS` to avoid ambiguity with the identically-named constant in `subagent-watcher.ts`.
3. Update the usage site: `setInterval(... POLL_INTERVAL_MS)` → `setInterval(... TEAM_POLL_INTERVAL_MS)`.
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Agent Teams polling and cache configuration.
*
* Controls how frequently TeamWatcher polls ~/.claude/teams/
* and how many teams/tasks are cached in memory.
*
* @module config/team-config
*/
/** Team directory poll interval (ms) */
export const TEAM_POLL_INTERVAL_MS = 30_000;
/** Max cached team configs (LRU eviction) */
export const MAX_CACHED_TEAMS = 50;
/** Max cached team tasks and inbox messages (LRU eviction).
* Used for both teamTasks and inboxCache maps. */
export const MAX_CACHED_TASKS = 200;
```
### Verification
```bash
tsc --noEmit
```
---
## Task 7: Fix remaining cross-file duplicates
**Estimated effort**: 30 minutes
**Files modified**: `src/index.ts`, `src/subagent-watcher.ts`
### Duplicate 1: `STATS_COLLECTION_INTERVAL_MS`
Already fixed in Task 1 — both `server.ts` and `mux-routes.ts` now import from `server-timing.ts`.
### Duplicate 2: AI model string
Already fixed in Task 5 — all 5 occurrences now import from `ai-defaults.ts`.
### Duplicate 3: `MAX_SCREENSHOT_SIZE` / `MAX_TEXT_FILE_SIZE` / `MAX_RAW_FILE_SIZE`
These file size limits in `file-routes.ts` and `system-routes.ts` are **API-specific validation limits**. They're only used in their respective route files and aren't duplicated. **Leave in place** — they're local to their route module and well-commented.
### Action A: Move `MAX_CONSECUTIVE_ERRORS` and `ERROR_RESET_MS` to config
`src/index.ts` has two process-level constants that are operational tuning knobs:
| Constant | Value | Purpose |
|----------|-------|---------|
| `MAX_CONSECUTIVE_ERRORS` | `5` | Max consecutive unhandled errors before process exit |
| `ERROR_RESET_MS` | `60000` | Error counter reset interval (1 min) |
These belong in a config file since they control server reliability behavior. Add them to `src/config/server-timing.ts` (they're server operational constants).
1. Add to `src/config/server-timing.ts`:
```typescript
// ============================================================================
// Process Error Recovery
// ============================================================================
/** Max consecutive unhandled errors before auto-restart */
export const MAX_CONSECUTIVE_ERRORS = 5;
/** Error counter reset interval — forgives errors after quiet period (ms) */
export const ERROR_RESET_MS = 60_000;
```
2. In `src/index.ts`: Remove lines 19–20, add import from `./config/server-timing.js`.
3. Run `tsc --noEmit`.
### Action B: Fix `MAX_TRACKED_AGENTS` shadow in `subagent-watcher.ts`
`subagent-watcher.ts` defines its own `MAX_TRACKED_AGENTS = 500` locally instead of importing the identical value from `config/map-limits.ts`. This is a latent bug — if someone changes the config value, the subagent watcher's copy stays stale.
1. In `src/subagent-watcher.ts`: Remove the local `MAX_TRACKED_AGENTS` constant. Add `import { MAX_TRACKED_AGENTS } from './config/map-limits.js'` (the value there is `MAX_TODOS_PER_SESSION = 500` — **verify** the map-limits constant is actually named `MAX_TRACKED_AGENTS` or if it needs to be added). If the constant doesn't exist in `map-limits.ts` under that name, add it.
2. Run `tsc --noEmit`.
### Verification
```bash
tsc --noEmit
npm run lint
npm run format:check
```
---
## Task 8: Update CLAUDE.md and final verification
**Estimated effort**: 20 minutes
**Files modified**: `CLAUDE.md`
### Updates to CLAUDE.md
1. **Config Files table** (`src/config/`): Add the 6 new files:
| File | Purpose |
|------|---------|
| `buffer-limits.ts` | Terminal/text buffer size limits |
| `map-limits.ts` | Global limits for Maps, sessions, watchers |
| `exec-timeout.ts` | Execution timeout configuration |
| `server-timing.ts` | Web server batching, SSE, scheduled run timing |
| `auth-config.ts` | Auth session TTL, rate limits, hook timeout |
| `tunnel-config.ts` | QR token rotation, tunnel process lifecycle |
| `terminal-limits.ts` | Terminal dimension and input validation limits |
| `ai-defaults.ts` | AI checker model and context limits |
| `team-config.ts` | Agent Teams polling and cache sizes |
2. **Import Conventions** section: Add:
```
- **Config**: Import from specific files: `import { MAX_TERMINAL_COLS } from './config/terminal-limits'`
```
3. **Phase 6 status** in `docs/code-structure-findings.md`: Mark as COMPLETE with summary of what was done.
### Final verification checklist
```bash
# Type checking
tsc --noEmit
# Linting
npm run lint
# Formatting
npm run format:check
# Dev server starts
npx tsx src/index.ts web --port 3099 &
curl -s http://localhost:3099/api/status | jq .status # "ok"
kill %1
# Verify no remaining duplicates
grep -rn 'STATS_COLLECTION_INTERVAL_MS' src/ # Should only appear in config + import sites
grep -rn 'timeout: 10000' src/hooks-config.ts # Should be 0 — all replaced with HOOK_TIMEOUT_MS
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
```
---
## What is NOT in scope (and why)
These constants were considered but deliberately left in their current files:
### Respawn controller defaults (`src/respawn-controller.ts`)
The `DEFAULT_CONFIG` object (lines 538–578) contains ~30 default values for the `RespawnConfig` interface. These are **user-facing configuration defaults**, not system constants — they're the starting values for a config object that users can modify via the API and UI. Centralizing them would break the locality between the config interface definition and its defaults. They already have excellent JSDoc with `@default` tags. The only values extracted are the AI model/context constants (Task 5) which are duplicated in other files.
### Subagent watcher timing (`src/subagent-watcher.ts`)
The 18 constants at lines 129–158 are all internal to the subagent watcher's polling/lifecycle algorithm. Moving them to a config file would force developers to context-switch between two files to understand the polling logic. They're already grouped with clear comments. Exception: `MAX_TRACKED_AGENTS` is consolidated with `map-limits.ts` (Task 7B) since it duplicates a global limit.
### Session auto-ops timing (`src/session-auto-ops.ts`)
The 8 constants at lines 19–40 are internal to the auto-compact/clear retry state machine. They form a coherent group that's meaningless without the surrounding implementation context.
### Run summary constants (`src/run-summary.ts`)
`MAX_EVENTS`, `TRIM_TO_EVENTS`, `TOKEN_MILESTONE_INTERVAL`, `STATE_STUCK_WARNING_MS`, `STATE_STUCK_CHECK_INTERVAL` — all module-internal. The buffer-style limits (`MAX_EVENTS`/`TRIM_TO_EVENTS`) follow the same pattern as `buffer-limits.ts` but are only used in this one file.
### Frontend (`src/web/public/constants.js`)
Already well-centralized. Frontend and backend run in different environments — mixing them in TypeScript config files would create import problems. If frontend constants need expansion, do it in `constants.js`. Note: `app.js` has 2 inline uses of `256 * 1024` that should use the existing `TERMINAL_TAIL_SIZE` from `constants.js` — a minor cleanup that can be done opportunistically but is not worth a task here.
### Tmux manager timing (`src/tmux-manager.ts`)
The 6 constants (lines 65–78) are internal to tmux process lifecycle management. They're low-level retry/wait values that are meaningless without understanding the tmux spawn sequence.
### Process-internal constants
`image-watcher.ts`, `bash-tool-parser.ts`, `transcript-watcher.ts`, `ralph-tracker.ts`, `task-tracker.ts`, `file-stream-manager.ts`, `session-lifecycle-log.ts`, `session-task-cache.ts`, `respawn-metrics.ts`, `respawn-adaptive-timing.ts`, `ai-checker-base.ts` — all have module-local constants that are internal implementation details.
### `localhost:3000` default URL
The string `'http://localhost:3000'` or port `3000` appears as a fallback default in ~5 files (`session-cli-builder.ts`, `tmux-manager.ts`, `tunnel-manager.ts`, `server.ts`, CLI). While technically duplicated, extracting it provides little value — each usage has a different fallback chain (env var → config → hardcoded) and the port is also baked into systemd service files and documentation. The risk of a missed update is low since port 3000 is deeply conventional.
### `SAVE_DEBOUNCE_MS = 500` in `state-store.ts` / `push-store.ts`
Same value (500ms), but they debounce different persistence targets (state.json vs push-subscriptions.json). If one needed faster/slower debouncing, they'd diverge. Coupling them would be misleading.
---
## Summary
| Metric | Before | After |
|--------|--------|-------|
| Config files in `src/config/` | 3 | 9 |
| Constants centralized | ~25 | ~65 |
| Cross-file duplicates | 9+ (`STATS_COLLECTION_INTERVAL_MS`, `timeout: 10000` ×6, AI model ×5, context limits ×3 each, `MAX_TRACKED_AGENTS`) | 0 |
| Files with `timeout: 10000` inline | 1 (6 occurrences) | 0 |
| Files with hardcoded AI model string | 4 (5 occurrences) | 1 (config only) |
| Files modified | — | 11 |
| Files created | — | 6 |
+953
View File
@@ -0,0 +1,953 @@
# Phase 7 Implementation Plan: Test Infrastructure
**Source**: `docs/code-structure-findings.md` (Phase 7 — Test Infrastructure)
**Estimated effort**: 2–3 days
**Tasks**: 11 tasks with dependencies (see dependency graph below)
---
## Safety Constraints
Before starting ANY work, read and follow these rules:
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
7. **Port assignments for this phase**: New tests use ports 3220–3229 (see individual tasks for assignments).
---
## Goal
Eliminate duplicated test mocks, activate the unused `respawn-test-utils.ts` utilities, and add route-level test coverage for the server's 12 route modules — the single largest untested area in the codebase (162 route handlers, 0 dedicated tests).
**Non-goals**:
- Full end-to-end integration tests (those require real Claude CLI / tmux sessions)
- 100% route coverage in this phase — focus on the highest-value route modules first
- Refactoring test patterns in existing passing tests that don't use shared mocks
- Migrating `vi.mock()`-based module replacement mocks (different pattern, see Task 6/7)
---
## Current State
### Mock Duplication (Finding #9)
`MockSession` is defined **4 times** across test files with varying levels of completeness:
| File | Properties | Methods | EventEmitter | Notes |
|------|-----------|---------|-------------|-------|
| `test/respawn-test-utils.ts` | 6 | 20+ | Yes | **Most complete**. Includes terminal simulation, token count, ANSI output, plan mode prompts. **Never imported by any test.** |
| `test/respawn-controller.test.ts` | 6 | 9 | Yes | Subset of respawn-test-utils. Missing token simulation, ANSI helpers. |
| `test/respawn-team-awareness.test.ts` | ~6 | ~9 | Yes | Near-copy of respawn-controller.test.ts version. |
| `test/session-manager.test.ts` | 4 | 8 | Yes | **Inside `vi.mock()` factory** — replaces `../src/session.js` module. Different shape: `start()`/`stop()`/`toState()`/`sendInput()` for lifecycle testing. |
`MockStateStore` is defined **2 times** (both inside `vi.mock()` factories):
| File | Shape | Methods | Mock Pattern |
|------|-------|---------|-------------|
| `test/session-manager.test.ts` | `{ sessions, config }` | `getConfig`, `getSessions`, `getSession`, `setSession`, `removeSession` | `vi.mock('../src/state-store.js')` |
| `test/ralph-loop.test.ts` | `{ ralphLoop, tasks, config }` | `getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask` | `vi.mock('../src/state-store.js')` |
### Important: Two distinct mocking patterns
The codebase uses two different mocking patterns that require different migration strategies:
1. **Direct instantiation** (respawn-controller, respawn-team-awareness): `MockSession` is defined at file scope and instantiated directly in tests. These can be migrated to shared mocks via simple import replacement.
2. **Module replacement** (session-manager, ralph-loop): Mocks are defined inside `vi.mock()` factories that replace entire modules (`../src/session.js`, `../src/state-store.js`). These factories run in an isolated scope and return `{ Session: MockClass }` or `{ getStore: vi.fn(() => instance) }`. Migrating these requires either `vi.hoisted()` or restructuring the test's module mocking — higher risk for limited benefit.
### Unused Test Utilities
`test/respawn-test-utils.ts` exports these utilities that **no test file imports**:
- `TimeController` / `createTimeController()` — abstraction over vitest fake timers
- `MockAiIdleChecker` / `MockAiPlanChecker` — fully mocked AI checkers with result queueing
- `createStateTracker()` / `createEventRecorder()` — state transition and event recording
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — pre-configured RespawnConfig objects
- `waitForState()` / `waitForEvent()` / `createDeferred()` — async test helpers
- `terminalOutputs` — factory object for common terminal output patterns
### Route Test Coverage
Currently **zero** dedicated tests for the 12 route modules in `src/web/routes/`. The existing test files that touch API endpoints:
| Test File | What It Tests | Approach |
|-----------|--------------|----------|
| `test/api-responses.test.ts` | Response structure validation | Imports types, no HTTP calls |
| `test/api-generate-plan.test.ts` | Plan generation API | Mocks validation logic, Port 3191 declared |
| `test/auth-security.test.ts` | Auth middleware | Integration tests with WebServer, Ports 3160/3161 |
| `test/qr-auth.test.ts` | QR authentication | Integration + unit tests, Port 3162 |
None of these test the route handlers themselves with real HTTP requests against a running Fastify instance.
---
## Design Decisions
### Shared mocks: Superset strategy
Rather than creating a lowest-common-denominator mock, `MockSession` in `test/mocks/` will be the **superset** from `respawn-test-utils.ts` (the most complete version). Test files that need a simpler mock can just ignore the extra methods — having unused methods costs nothing, but missing methods forces local re-definition.
### vi.mock() tests: Don't migrate
The `session-manager.test.ts` and `ralph-loop.test.ts` tests define mocks inside `vi.mock()` factories. These use **module-level replacement** (replacing `../src/session.js` and `../src/state-store.js` entirely), which is fundamentally different from the direct-instantiation pattern. Migrating them would require `vi.hoisted()` or factory restructuring — high complexity for limited benefit since these mocks are already working. We leave these as-is and create the shared mocks for **new** tests and for the two direct-instantiation tests (Tasks 4–5).
### MockStateStore: Union of both shapes
The shared `MockStateStore` in `test/mocks/` will include methods from both existing definitions (session management + Ralph loop), so any **new** test can use it. Methods default to no-ops via `vi.fn()`. Existing `vi.mock()`-based tests are not migrated.
### Route testing strategy: Lightweight Fastify instances
Each route test file will:
1. Create a minimal `Fastify` instance
2. Register **only** the route module under test
3. Provide a mock context object satisfying the port interfaces
4. Use `app.inject()` (Fastify's built-in test helper) — no real HTTP, no port needed
This avoids port conflicts entirely and runs fast. Only tests that need SSE or WebSocket behavior will use a real listening server with assigned ports.
### Port assignments (for tests needing real servers)
| Port | Test File | Purpose |
|------|-----------|---------|
| 3220 | `test/routes/session-routes.test.ts` | SSE integration (if needed) |
| 3221 | `test/routes/system-routes.test.ts` | Status/stats endpoints |
| 3222 | `test/routes/respawn-routes.test.ts` | Respawn API |
| 3223 | `test/routes/ralph-routes.test.ts` | Ralph API |
| 3224–3229 | Reserved | Future route tests |
Most tests should NOT need real ports — `app.inject()` is preferred. Verified: ports 3220–3229 are completely unused by existing tests (highest used port is 3211 in `opencode-resize.test.ts`).
---
## Task Dependencies
```
Task 1 (Consolidate MockSession)
Task 2 (Consolidate MockStateStore)
└──> Task 3 (Create test/mocks/ barrel)
├──> Task 4 (Migrate respawn-controller.test.ts)
├──> Task 5 (Migrate respawn-team-awareness.test.ts)
└──> Task 6 (Route test scaffold + helpers)
├──> Task 7 (Session routes tests)
└──> Task 8 (System + respawn routes tests)
Task 9 (Slim down respawn-test-utils.ts) — depends on Tasks 4, 5
```
**Tasks 1–2** are independent and can run in parallel.
**Task 3** depends on Tasks 1–2.
**Tasks 4–6** depend on Task 3 and can run in parallel.
**Tasks 7–8** depend on Task 6 and can run in parallel.
**Task 9** depends on Tasks 4, 5 (must verify migrations work before removing duplicates from source).
---
## Task 1: Consolidate MockSession into `test/mocks/mock-session.ts`
**Estimated effort**: 2 hours
**Files created**: `test/mocks/mock-session.ts`
**Files modified**: None yet (consumers migrate in Tasks 4–5)
### Source
The canonical MockSession comes from `test/respawn-test-utils.ts` (lines 89–241). It is the most complete version with:
- All properties needed by `RespawnController`: `id`, `workingDir`, `status`, `writeBuffer`, `terminalBuffer`, `muxName`
- `write()` / `writeViaMux()` for input simulation
- Buffer inspection: `lastWrite`, `hasWritten(pattern)`, `clearWriteBuffer()`
- Terminal simulation: `simulateTerminalOutput()`, `simulatePrompt()`, `simulateReady()`, `simulateCompletionMessage()`, `simulateWorking()`, `simulateClearComplete()`, `simulateInitComplete()`, `simulatePlanModePrompt()`, `simulateElicitationDialog()`, `simulateTokenCount()`, `simulateAnsiOutput()`
- Lifecycle: `close()`
### Implementation
1. Create `test/mocks/` directory.
2. Create `test/mocks/mock-session.ts`:
- Copy the `MockSession` class **exactly** from `test/respawn-test-utils.ts` (lines 89–241)
- Copy `terminalOutputs` helper object (tightly coupled to mock)
- Copy `createMockSession()` factory function
- Export all three: `export { MockSession, createMockSession, terminalOutputs }`
- Ensure all `vi` imports come from `vitest`
**CRITICAL**: Copy the source verbatim — do NOT rewrite the simulation methods. The respawn controller's detection logic matches specific output patterns (e.g., `'\u276f '` for prompt, `'\u273b Worked for'` for completion). Using different patterns would cause test failures.
### Template
```typescript
/**
* Shared MockSession for tests that need terminal simulation.
*
* Copied from test/respawn-test-utils.ts (the canonical, most complete version).
* Used by respawn, route, and subagent tests.
*/
import { EventEmitter } from 'node:events';
// Copy MockSession class exactly from test/respawn-test-utils.ts lines 89–241
export class MockSession extends EventEmitter {
// ... (copy verbatim from respawn-test-utils.ts)
}
/**
* Factory for common terminal output strings.
* Must match the patterns used in MockSession's simulate* methods.
*/
export const terminalOutputs = {
// ... (copy verbatim from respawn-test-utils.ts)
};
/**
* Convenience factory.
*/
export function createMockSession(id?: string): MockSession {
return new MockSession(id);
}
```
### Verification
```bash
tsc --noEmit # Ensure file compiles
```
---
## Task 2: Consolidate MockStateStore into `test/mocks/mock-state-store.ts`
**Estimated effort**: 1 hour
**Files created**: `test/mocks/mock-state-store.ts`
**Files modified**: None (existing vi.mock()-based tests are NOT migrated; this is for new route tests)
### Source
Union of both existing definitions:
- From `test/session-manager.test.ts`: session CRUD methods (`getConfig`, `getSession`, `setSession`, `removeSession`, `getSessions`)
- From `test/ralph-loop.test.ts`: Ralph state methods (`getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask`)
### Template
```typescript
/**
* Shared MockStateStore for tests.
*
* Includes methods for both session management and Ralph loop testing.
* All methods are vi.fn() spies — tests can override return values as needed.
*
* NOTE: This is for direct instantiation in new tests. Existing tests that
* use vi.mock('../src/state-store.js') keep their inline definitions.
*/
import { vi } from 'vitest';
export class MockStateStore {
state: Record<string, unknown> = {
sessions: {} as Record<string, unknown>,
config: { maxConcurrentSessions: 5 },
ralphLoop: { status: 'stopped' },
tasks: {} as Record<string, unknown>,
};
// Session methods
getConfig = vi.fn(() => this.state.config);
getSessions = vi.fn(() => this.state.sessions as Record<string, unknown>);
getSession = vi.fn((id: string) => (this.state.sessions as Record<string, unknown>)[id]);
setSession = vi.fn((id: string, state: unknown) => {
(this.state.sessions as Record<string, unknown>)[id] = state;
});
removeSession = vi.fn((id: string) => {
delete (this.state.sessions as Record<string, unknown>)[id];
});
// Ralph state methods
getRalphLoopState = vi.fn(() => this.state.ralphLoop);
setRalphLoopState = vi.fn((update: Record<string, unknown>) => {
this.state.ralphLoop = { ...(this.state.ralphLoop as Record<string, unknown>), ...update };
});
// Task methods
getTasks = vi.fn(() => this.state.tasks);
setTask = vi.fn();
removeTask = vi.fn();
// Settings methods
getSettings = vi.fn(() => ({}));
setSettings = vi.fn();
// Generic persistence
save = vi.fn();
load = vi.fn();
/** Reset all state and mocks for clean test isolation */
reset(): void {
this.state = {
sessions: {},
config: { maxConcurrentSessions: 5 },
ralphLoop: { status: 'stopped' },
tasks: {},
};
vi.clearAllMocks();
}
}
```
### Verification
```bash
tsc --noEmit
```
---
## Task 3: Create `test/mocks/index.ts` barrel export
**Estimated effort**: 30 minutes
**Depends on**: Tasks 1, 2
**Files created**: `test/mocks/index.ts`, `test/mocks/test-helpers.ts`
**Files modified**: None
### Implementation
1. Create `test/mocks/test-helpers.ts` with the async utilities from `respawn-test-utils.ts`:
```typescript
/**
* Reusable async test helpers.
* Extracted from respawn-test-utils.ts.
*/
/** Wait for an EventEmitter to emit a specific event, with timeout */
export function waitForEvent(
emitter: { once: (event: string, listener: (...args: unknown[]) => void) => void },
event: string,
timeoutMs = 5000,
): Promise<unknown> {
return new Promise((resolve, reject) => {
const timer = setTimeout(
() => reject(new Error(`Timed out waiting for event "${event}" after ${timeoutMs}ms`)),
timeoutMs,
);
emitter.once(event, (...args: unknown[]) => {
clearTimeout(timer);
resolve(args.length === 1 ? args[0] : args);
});
});
}
/** Create a deferred promise with external resolve/reject */
export function createDeferred<T = void>(): {
promise: Promise<T>;
resolve: (value: T) => void;
reject: (reason?: unknown) => void;
} {
let resolve!: (value: T) => void;
let reject!: (reason?: unknown) => void;
const promise = new Promise<T>((res, rej) => {
resolve = res;
reject = rej;
});
return { promise, resolve, reject };
}
```
2. Create `test/mocks/index.ts` barrel:
```typescript
/**
* Shared test mocks — import from here instead of defining inline.
*
* @example
* import { MockSession, MockStateStore, terminalOutputs } from './mocks/index.js';
*/
export { MockSession, createMockSession, terminalOutputs } from './mock-session.js';
export { MockStateStore } from './mock-state-store.js';
export { waitForEvent, createDeferred } from './test-helpers.js';
```
### Verification
```bash
tsc --noEmit
```
---
## Task 4: Migrate `respawn-controller.test.ts` to shared mocks
**Estimated effort**: 30 minutes
**Depends on**: Task 3
**Files modified**: `test/respawn-controller.test.ts`
### Steps
1. Remove the local `MockSession` class definition (approx. 50 lines).
2. Add: `import { MockSession } from './mocks/index.js';`
3. Verify all test methods still exist on the shared mock. The shared mock is a superset, so all existing usage should work.
4. If the local mock had any test-specific customizations (e.g., extra properties added in `beforeEach`), keep those in the test file as inline assignments on the shared instance.
5. Run the test to confirm it passes.
### Potential issues
- The local mock's `simulateCompletionMessage()` may have a slightly different output format than the shared mock's (from respawn-test-utils.ts). Verify the respawn controller's completion detection regex matches the shared mock's output pattern (`'\u273b Worked for ...'`).
- If the local mock adds `pid` or `isWorking` properties that the shared mock doesn't have, add inline assignments in `beforeEach`.
### Verification
```bash
npx vitest run test/respawn-controller.test.ts
```
---
## Task 5: Migrate `respawn-team-awareness.test.ts` to shared mocks
**Estimated effort**: 30 minutes
**Depends on**: Task 3
**Files modified**: `test/respawn-team-awareness.test.ts`
### Steps
1. Remove the local `MockSession` class definition.
2. Add: `import { MockSession } from './mocks/index.js';`
3. Keep `MockTeamWatcher` in this file — it's test-specific and extends the real `TeamWatcher`, not a general-purpose mock.
4. Run the test to confirm it passes.
### Verification
```bash
npx vitest run test/respawn-team-awareness.test.ts
```
---
## Task 6: Create route test scaffold and helpers
**Estimated effort**: 2 hours
**Depends on**: Task 3
**Files created**: `test/mocks/mock-route-context.ts`, `test/routes/` directory, `test/routes/_route-test-utils.ts`
### Problem
The 12 route modules in `src/web/routes/` have zero dedicated test coverage. Each route module takes `(app: FastifyInstance, ctx: PortIntersection)` — we need a reusable way to create mock context objects that satisfy the port interfaces.
### Design
Create a `MockRouteContext` factory that builds a mock object satisfying all port interfaces. Each port's methods are `vi.fn()` stubs. Tests can override specific methods as needed.
### Route registration signatures (verified)
Each route module requires a specific port intersection. The mock must satisfy all of them:
| Route Module | Required Ports |
|-------------|----------------|
| `registerSessionRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
| `registerSystemRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
| `registerRespawnRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
| `registerRalphRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
| `registerPlanRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort` |
| `registerCaseRoutes` | `EventPort & ConfigPort` |
| `registerScheduledRoutes` | `SessionPort & EventPort & InfraPort` |
| `registerFileRoutes` | `SessionPort` |
| `registerMuxRoutes` | `InfraPort` |
| `registerPushRoutes` | `InfraPort` |
| `registerTeamRoutes` | `InfraPort` |
| `registerHookEventRoutes` | `EventPort & AuthPort` |
### Implementation
1. Create `test/mocks/mock-route-context.ts`:
```typescript
/**
* Mock context for route handler testing.
*
* Satisfies ALL port interfaces (SessionPort, EventPort, RespawnPort,
* ConfigPort, InfraPort, AuthPort) so any route module can be tested.
* Override specific methods in individual tests as needed.
*
* Verified against actual port interfaces in src/web/ports/:
* - SessionPort: 6 methods (sessions, addSession, cleanupSession,
* setupSessionListeners, persistSessionState, persistSessionStateNow,
* getSessionStateWithRespawn)
* - EventPort: 5 methods (broadcast, sendPushNotifications, batchTerminalData,
* broadcastSessionStateDebounced, batchTaskUpdate)
* - RespawnPort: 2 maps + 4 methods
* - ConfigPort: 5 readonly + 7 methods (incl getDefaultClaudeMdPath,
* getLightState, getLightSessionsState, stopTranscriptWatcher)
* - InfraPort: 7 readonly + 2 methods (startScheduledRun, stopScheduledRun)
* - AuthPort: 3 readonly (authSessions, qrAuthFailures, https)
*/
import { vi } from 'vitest';
import { MockSession, createMockSession } from './mock-session.js';
/**
* Creates a mock context that satisfies all port interfaces.
* Pre-populated with one session for convenience.
*/
export function createMockRouteContext(options?: { sessionId?: string }) {
const sessionId = options?.sessionId ?? 'test-session-1';
const session = createMockSession(sessionId);
const sessions = new Map<string, MockSession>();
sessions.set(sessionId, session);
return {
// -- SessionPort --
sessions,
addSession: vi.fn(),
cleanupSession: vi.fn(),
setupSessionListeners: vi.fn(),
persistSessionState: vi.fn(),
persistSessionStateNow: vi.fn(),
getSessionStateWithRespawn: vi.fn((s: unknown) => s),
// -- EventPort --
broadcast: vi.fn(),
sendPushNotifications: vi.fn(),
batchTerminalData: vi.fn(),
broadcastSessionStateDebounced: vi.fn(),
batchTaskUpdate: vi.fn(),
// -- RespawnPort --
respawnControllers: new Map(),
respawnTimers: new Map(),
setupRespawnListeners: vi.fn(),
setupTimedRespawn: vi.fn(),
restoreRespawnController: vi.fn(),
saveRespawnConfig: vi.fn(),
// -- ConfigPort --
store: {
getConfig: vi.fn(() => ({})),
getSessions: vi.fn(() => ({})),
getSession: vi.fn(),
setSession: vi.fn(),
removeSession: vi.fn(),
getSettings: vi.fn(() => ({})),
setSettings: vi.fn(),
getRalphLoopState: vi.fn(() => ({})),
setRalphLoopState: vi.fn(),
getTasks: vi.fn(() => ({})),
save: vi.fn(),
load: vi.fn(),
},
port: 3000,
https: false,
testMode: true,
serverStartTime: Date.now(),
getGlobalNiceConfig: vi.fn(async () => undefined),
getModelConfig: vi.fn(async () => null),
getClaudeModeConfig: vi.fn(async () => ({})),
getDefaultClaudeMdPath: vi.fn(async () => undefined),
getLightState: vi.fn(() => ({ sessions: [], status: 'ok' })),
getLightSessionsState: vi.fn(() => []),
startTranscriptWatcher: vi.fn(),
stopTranscriptWatcher: vi.fn(),
// -- InfraPort --
mux: {
createSession: vi.fn(),
killSession: vi.fn(),
listSessions: vi.fn(() => []),
getStats: vi.fn(() => ({})),
},
runSummaryTrackers: new Map(),
activePlanOrchestrators: new Map(),
scheduledRuns: new Map(),
teamWatcher: { getTeams: vi.fn(() => []), hasActiveTeammates: vi.fn(() => false) },
tunnelManager: null,
pushStore: null,
startScheduledRun: vi.fn(),
stopScheduledRun: vi.fn(),
// -- AuthPort --
authSessions: null,
qrAuthFailures: null,
// https already declared above in ConfigPort (shared property)
// Convenience accessors (not part of any port interface)
_session: session,
_sessionId: sessionId,
};
}
export type MockRouteContext = ReturnType<typeof createMockRouteContext>;
```
2. Add to `test/mocks/index.ts` barrel:
```typescript
export { createMockRouteContext, type MockRouteContext } from './mock-route-context.js';
```
3. Create `test/routes/` directory for route test files.
4. Create `test/routes/_route-test-utils.ts` with Fastify test helpers:
```typescript
/**
* Shared utilities for route testing.
*
* Creates minimal Fastify instances with just the route module under test
* and a mock context. Uses app.inject() for HTTP testing without real ports.
*/
import Fastify, { type FastifyInstance } from 'fastify';
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
export interface RouteTestHarness {
app: FastifyInstance;
ctx: MockRouteContext;
}
/**
* Creates a Fastify instance with a route module registered against a mock context.
*
* @param registerFn - The route registration function (e.g., registerSessionRoutes).
* Uses `any` for ctx parameter because route functions expect typed port intersections
* that MockRouteContext satisfies structurally but not nominally.
* @param ctxOptions - Optional overrides for the mock context
*/
export async function createRouteTestHarness(
// eslint-disable-next-line @typescript-eslint/no-explicit-any
registerFn: (app: FastifyInstance, ctx: any) => void,
ctxOptions?: { sessionId?: string },
): Promise<RouteTestHarness> {
const app = Fastify({ logger: false });
const ctx = createMockRouteContext(ctxOptions);
registerFn(app, ctx);
await app.ready();
return { app, ctx };
}
```
### Why `ctx: any` in the harness
Route registration functions like `registerSessionRoutes(app, ctx: SessionPort & EventPort & ConfigPort & InfraPort & AuthPort)` expect specific port intersection types. TypeScript won't accept `unknown` here because it's not assignable to the port types. The `MockRouteContext` satisfies the interfaces structurally (it has all the required properties and methods), but since it's not declared as implementing them, we need `any` at the call site. This is the standard pattern for test mocks in TypeScript.
### Verification
```bash
tsc --noEmit
```
---
## Task 7: Add session routes tests
**Estimated effort**: 4 hours
**Depends on**: Task 6
**Files created**: `test/routes/session-routes.test.ts`
**Port**: 3220 (only if SSE tests needed; prefer `app.inject()`)
### Coverage targets
`src/web/routes/session-routes.ts` is the largest route module (43 handlers). Focus on the most critical endpoints first:
#### Priority 1: Session CRUD (must test)
| Method | Path | What to test |
|--------|------|-------------|
| `GET` | `/api/sessions` | Returns session list; empty when no sessions |
| `GET` | `/api/sessions/:id` | Returns session state; 404 for unknown ID |
| `POST` | `/api/sessions` | Creates session; validates workingDir; rejects invalid paths |
| `DELETE` | `/api/sessions/:id` | Calls cleanupSession; 404 for unknown ID |
#### Priority 2: Session I/O
| Method | Path | What to test |
|--------|------|-------------|
| `POST` | `/api/sessions/:id/input` | Sends input to session; validates input length; 404 for unknown |
| `POST` | `/api/sessions/:id/resize` | Validates cols/rows bounds; 404 for unknown |
| `GET` | `/api/sessions/:id/buffer` | Returns terminal buffer; 404 for unknown |
#### Priority 3: Session actions
| Method | Path | What to test |
|--------|------|-------------|
| `POST` | `/api/sessions/:id/run` | Runs prompt on session |
| `POST` | `/api/sessions/:id/clear` | Clears session |
| `POST` | `/api/sessions/:id/compact` | Compacts session |
| `POST` | `/api/sessions/:id/interactive` | Starts interactive mode |
| `POST` | `/api/sessions/:id/quick-start` | Quick start flow |
### Test pattern
```typescript
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { createRouteTestHarness, type RouteTestHarness } from './_route-test-utils.js';
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
describe('session-routes', () => {
let harness: RouteTestHarness;
beforeEach(async () => {
harness = await createRouteTestHarness(registerSessionRoutes);
});
afterEach(async () => {
await harness.app.close();
});
describe('GET /api/sessions', () => {
it('returns empty array when no sessions', async () => {
harness.ctx.sessions.clear();
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
expect(res.statusCode).toBe(200);
expect(JSON.parse(res.body)).toEqual([]);
});
it('returns session list with one session', async () => {
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
expect(res.statusCode).toBe(200);
const sessions = JSON.parse(res.body);
expect(sessions).toHaveLength(1);
});
});
describe('GET /api/sessions/:id', () => {
it('returns 404 for unknown session', async () => {
const res = await harness.app.inject({
method: 'GET',
url: '/api/sessions/nonexistent',
});
expect(res.statusCode).toBe(404);
});
});
describe('POST /api/sessions/:id/input', () => {
it('rejects input exceeding max length', async () => {
const res = await harness.app.inject({
method: 'POST',
url: `/api/sessions/${harness.ctx._sessionId}/input`,
payload: { input: 'x'.repeat(65537) },
});
expect(res.statusCode).toBe(400);
});
});
describe('POST /api/sessions/:id/resize', () => {
it('rejects cols exceeding max', async () => {
const res = await harness.app.inject({
method: 'POST',
url: `/api/sessions/${harness.ctx._sessionId}/resize`,
payload: { cols: 501, rows: 24 },
});
expect(res.statusCode).toBe(400);
});
});
});
```
### Key assertions to include
- **404 for unknown sessions**: Every `:id` endpoint must return 404 for nonexistent IDs
- **Input validation**: Bad paths, oversized inputs, invalid resize dimensions
- **Side effects**: Verify `ctx.broadcast()` was called with correct event type after mutations
- **Response shape**: Verify response bodies match expected API types
### Verification
```bash
npx vitest run test/routes/session-routes.test.ts
```
---
## Task 8: Add system + respawn routes tests
**Estimated effort**: 4 hours
**Depends on**: Task 6
**Files created**: `test/routes/system-routes.test.ts`, `test/routes/respawn-routes.test.ts`
### System routes (`src/web/routes/system-routes.ts`)
Focus on status and configuration endpoints:
| Method | Path | What to test |
|--------|------|-------------|
| `GET` | `/api/status` | Returns server status with uptime, session count |
| `GET` | `/api/stats` | Returns mux stats |
| `GET` | `/api/config` | Returns current config |
| `PUT` | `/api/config` | Updates config; validates input |
| `GET` | `/api/settings` | Returns user settings |
| `PUT` | `/api/settings` | Updates settings; validates input |
| `GET` | `/api/subagents` | Returns subagent list |
| `GET` | `/api/screenshots` | Returns screenshot list |
### Respawn routes (`src/web/routes/respawn-routes.ts`)
| Method | Path | What to test |
|--------|------|-------------|
| `GET` | `/api/sessions/:id/respawn` | Returns respawn status; null when not configured |
| `POST` | `/api/sessions/:id/respawn/start` | Starts respawn; 404 for unknown session |
| `POST` | `/api/sessions/:id/respawn/stop` | Stops respawn; 404 for unknown session |
| `PUT` | `/api/sessions/:id/respawn/config` | Updates respawn config; validates |
| `POST` | `/api/sessions/:id/respawn/enable` | Enables respawn loop |
| `POST` | `/api/sessions/:id/respawn/disable` | Disables respawn loop |
### Test patterns
Same pattern as Task 7 — `createRouteTestHarness` with `registerSystemRoutes` / `registerRespawnRoutes`.
For respawn tests, pre-populate `ctx.respawnControllers` with a mock controller in `beforeEach`:
```typescript
beforeEach(async () => {
harness = await createRouteTestHarness(registerRespawnRoutes);
// Add a mock respawn controller for the default session
harness.ctx.respawnControllers.set(harness.ctx._sessionId, {
getState: vi.fn(() => 'idle'),
getConfig: vi.fn(() => ({})),
getStatus: vi.fn(() => ({ state: 'idle', health: 100 })),
start: vi.fn(),
stop: vi.fn(),
updateConfig: vi.fn(),
enable: vi.fn(),
disable: vi.fn(),
});
});
```
### Verification
```bash
npx vitest run test/routes/system-routes.test.ts
npx vitest run test/routes/respawn-routes.test.ts
```
---
## Task 9: Slim down `respawn-test-utils.ts`
**Estimated effort**: 30 minutes
**Depends on**: Tasks 4, 5
**Files modified**: `test/respawn-test-utils.ts`
After Tasks 4–5 are verified passing with shared mocks, slim down `respawn-test-utils.ts` to remove duplicates.
### Steps
1. **Remove** from `respawn-test-utils.ts` what has been moved to shared mocks:
- `MockSession` class → now in `test/mocks/mock-session.ts`
- `createMockSession()` → now in `test/mocks/mock-session.ts`
- `terminalOutputs` → now in `test/mocks/mock-session.ts`
- `waitForEvent()` / `createDeferred()` → now in `test/mocks/test-helpers.ts`
2. **Keep** respawn-specific utilities that don't belong in the general mocks:
- `TimeController` / `createTimeController()` — respawn-specific timer control
- `MockAiIdleChecker` / `MockAiPlanChecker` — respawn-specific AI mocks
- `createStateTracker()` / `createEventRecorder()` — respawn state tracking
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — respawn config presets
- `waitForState()` — respawn state machine waiter
3. **Update imports** in `respawn-test-utils.ts` to re-use shared mocks:
```typescript
import { MockSession, createMockSession, terminalOutputs } from './mocks/index.js';
import { waitForEvent, createDeferred } from './mocks/index.js';
export { MockSession, createMockSession, terminalOutputs, waitForEvent, createDeferred };
```
This preserves backward compatibility for any future tests that import from `respawn-test-utils.ts` directly while eliminating the duplication.
### Verification
```bash
tsc --noEmit
npx vitest run test/respawn-controller.test.ts
npx vitest run test/respawn-team-awareness.test.ts
```
---
## What is NOT in scope (and why)
### Migrating `session-manager.test.ts` and `ralph-loop.test.ts` mocks
Both files define mocks inside `vi.mock()` factories that replace entire modules:
```typescript
// session-manager.test.ts — mock replaces ../src/session.js
vi.mock('../src/session.js', () => {
class MockSession extends EventEmitter { ... }
return { Session: MockSession };
});
// ralph-loop.test.ts — mock replaces ../src/state-store.js
vi.mock('../src/state-store.js', () => {
class MockStateStore { ... }
return { getStore: vi.fn(() => instance), StateStore: MockStateStore };
});
```
These are fundamentally different from the direct-instantiation pattern:
- The `vi.mock()` factory runs in an isolated scope — outer imports are not available
- The mock class must be returned with the exact export names (`Session`, `getStore`, `StateStore`)
- The `session-manager.test.ts` MockSession auto-registers into a shared `mockState.sessions` Map (tight coupling with test setup)
Migrating would require `vi.hoisted()` to share the class between factory and test scope, plus restructuring the test's module-mocking setup. This is high-complexity, high-risk refactoring with limited benefit since these tests already work. The shared `MockStateStore` in `test/mocks/` is available for **new** tests (like route tests) that use direct instantiation instead.
### Full integration tests with real Fastify server
Route tests use `app.inject()` which simulates HTTP without opening ports. Full integration tests that spin up `WebServer`, create real sessions, and stream SSE would be valuable but are a separate effort requiring:
- A test WebServer factory
- Session lifecycle management in tests
- SSE client test utilities
- Significantly more setup/teardown complexity
### Testing auth middleware in route tests
Route tests bypass authentication (no auth middleware registered on the test Fastify instance). Auth middleware has its own dedicated tests in `auth-security.test.ts` and `qr-auth.test.ts`. Testing auth + routes together is a future integration test concern.
### Testing SSE event streaming
SSE integration requires a running server with `EventSource` client. This is significantly more complex than `app.inject()` tests and is deferred. The existing `sse-events.test.ts` covers SSE patterns.
### Complete route coverage for all 12 modules
This phase covers the 3 highest-value route modules (session, system, respawn — 98 of 162 handlers). The remaining 9 modules (ralph, plan, push, team, mux, file, scheduled, hook-event, case) should be added incrementally in follow-up work.
---
## Summary
| Metric | Before | After |
|--------|--------|-------|
| MockSession definitions | 4 (across 4 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
| MockStateStore definitions | 2 (across 2 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
| Files importing from `respawn-test-utils.ts` | 0 | Utilities split into `test/mocks/` |
| Route test files | 0 | 3 (session, system, respawn) |
| Route handlers with dedicated tests | 0 | ~30 (highest-priority endpoints) |
| Shared mock directory | None | `test/mocks/` with 5 files + barrel |
### Final verification checklist
```bash
# Type checking
tsc --noEmit
# Linting
npm run lint
# Formatting
npm run format:check
# Run all affected tests individually
npx vitest run test/respawn-controller.test.ts
npx vitest run test/respawn-team-awareness.test.ts
npx vitest run test/routes/session-routes.test.ts
npx vitest run test/routes/system-routes.test.ts
npx vitest run test/routes/respawn-routes.test.ts
# Verify unchanged tests still pass
npx vitest run test/session-manager.test.ts
npx vitest run test/ralph-loop.test.ts
# Dev server still starts
npx tsx src/index.ts web --port 3099 &
curl -s http://localhost:3099/api/status | jq .status # "ok"
kill %1
```
+723
View File
@@ -0,0 +1,723 @@
# QR Code Authentication Plan
> Ephemeral, single-use auth tokens embedded in the tunnel QR code — scan to auto-authenticate, while the bare tunnel URL stays password-protected.
## Problem
When the Cloudflare tunnel is active, anyone who discovers the `*.trycloudflare.com` URL can access Codeman (they just need the Basic Auth password, or if no password is set, full open access). The QR code currently encodes the raw tunnel URL — it provides no additional security. We want:
1. **Scanning the QR code** → seamless, instant access (no password prompt)
2. **Having only the URL** → blocked by Basic Auth (no access without credentials)
## Design
### Core Concept: Ephemeral Single-Use QR Tokens
The server maintains a rotating pool of short-lived, single-use tokens. The QR code encodes a short URL containing a lookup code that maps to the real token server-side. When scanned, the server validates the token, atomically consumes it, issues a session cookie, and redirects to `/`. The token is **not** the password — it's a separate, independent, ephemeral authentication pathway.
```
Desktop → displays QR (auto-refreshes every 60s via SSE)
QR Code → https://abc-xyz.trycloudflare.com/q/Xk9mQ3
Phone → scans, GET /q/Xk9mQ3
Server → looks up short code via Map (hash-based, timing-safe)
→ finds token record → validates TTL
→ atomically consumes token (single-use)
→ issues codeman_session cookie
→ 302 redirect to /
→ SSE push: new QR with embedded SVG for desktop display
→ desktop toast: "Device [IP] authenticated via QR"
→ audit log entry to session-lifecycle.jsonl
User → lands on app, fully authenticated
```
Someone who only has `https://abc-xyz.trycloudflare.com/` gets the standard Basic Auth prompt.
### Token Properties
| Property | Value |
|----------|-------|
| Length | 32 bytes (256 bits entropy) |
| Generation | `crypto.randomBytes(32).toString('hex')` |
| Short code | 6 chars base62, rejection-sampled (no modulo bias) |
| Short code derivation | Independent random generation (not derived from token) |
| Storage | In-memory `Map<shortCode, QrTokenRecord>` (no disk persistence) |
| TTL | 60 seconds (auto-rotation via timer), 90s grace for previous token |
| Effective window | Up to 90 seconds for the previous token (documented, not hidden) |
| Usage | **Single-use** — atomically consumed on first valid scan |
| URL format | Short code in path (`/q/Xk9mQ3`), not query params |
| URL length | ~53-56 chars total — targets QR Version 4 (33x33) for fast scanning |
| Scope | Only valid when `CODEMAN_PASSWORD` is set (no point without auth) |
| Lookup | `Map.get()` — hash-based O(1), no timing side-channel |
### Why This Design?
**Why not embed the password directly?**
- Password would appear in browser history, Cloudflare edge logs, and URL bars
- Password can't be rotated independently from QR access
**Why not a long-lived multi-use token? (original design)**
- A static token is functionally a second password — if the QR image leaks (screenshot shared, shoulder surfing, Cloudflare logs), the attacker has permanent access
- The USENIX Security 2025 paper ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) found 47 of the top-100 websites vulnerable due to exactly this pattern — missing single-use enforcement and long-lived tokens were 2 of the 6 critical design flaws identified
**Why short codes in the URL path instead of query params?**
- Query params (`?t=TOKEN`) leak into browser history, address bar, `Referer` headers, and Cloudflare edge logs
- Path-based short codes (`/q/Xk9mQ3`) are opaque references — the real token never appears in URLs
- Short codes are 6-char base62 (62^6 = 56.8 billion combinations), sufficient for lookup since they're backed by the full 256-bit token for validation and rate-limited to 10 attempts/IP
- The short `/q/` path (vs `/qr-auth/`) saves 7 bytes, helping keep the QR at Version 4 (33x33 modules) instead of Version 5 (37x37) — faster scanning on budget phones
## Auth Flow Diagram
```
┌─────────────┐ scan QR ┌──────────────────────────────────────┐
│ Mobile │ ────────────→ │ GET /q/Xk9mQ3 │
│ Device │ │ │
└─────────────┘ │ 1. Auth middleware sees /q/ │
│ → skips Basic Auth check │
│ 2. Route handler: Map.get(shortCode) │
│ → hash-based lookup (timing-safe) │
│ 3. Checks TTL (90s grace for prev) │
│ → token not expired? │
│ 4. Checks consumed flag │
│ → not already used? │
│ 5. Atomically marks token consumed │
│ 6. Issues codeman_session cookie │
│ 7. 302 redirect to / │
│ 8. Audit log → session-lifecycle.jsonl│
│ 9. SSE push: tunnel:qrRegenerated │
│ → desktop refreshes QR (SVG inline)│
│ 10. Desktop toast: "Device auth'd" │
└──────────────────────────────────────┘
┌─────────────┐ replay URL ┌──────────────────────────────────────┐
│ Attacker │ ────────────→ │ GET /q/Xk9mQ3 │
│ (stale code) │ │ │
└─────────────┘ │ 1. Map.get(shortCode) → not found │
│ OR token consumed OR expired │
│ 2. Increment QR rate limit counter │
│ (separate from Basic Auth counter) │
│ 3. 401 Unauthorized │
└──────────────────────────────────────┘
┌─────────────┐ URL only ┌──────────────────────────────────────┐
│ Attacker │ ────────────→ │ GET / │
│ (no token) │ │ │
└─────────────┘ │ 1. Auth middleware checks cookie │
│ → no cookie │
│ 2. Checks Basic Auth header │
│ → no header │
│ 3. Returns 401 + WWW-Authenticate │
│ → Browser shows password popup │
└──────────────────────────────────────┘
```
## Implementation
### 1. Token Manager — `src/tunnel-manager.ts`
Add a `QrTokenRecord` type and token rotation logic to `TunnelManager`. The token rotates every 60 seconds. A consumed token is immediately replaced. Up to 2 tokens can be valid simultaneously (current + previous, to handle the race where someone scans right as rotation happens). The previous token has a 90s grace period (not a full extra 60s — only enough to cover the scan-during-rotation race).
**Design decisions from security review:**
- **Map-based lookup** (not array scan) — `Map.get()` uses hash-based O(1) lookup, eliminating timing side-channels from string comparison
- **Rejection sampling** for short codes — avoids modulo bias (`256 % 62 != 0` gives 25% overrepresentation for first 6 charset chars)
- **SVG cache** — stores generated QR SVG per rotation cycle to avoid regenerating on every `/api/tunnel/qr` poll
- **Separate rate limit counter** — QR auth failures tracked independently from Basic Auth failures
```typescript
import { randomBytes } from 'node:crypto';
interface QrTokenRecord {
token: string; // 64 hex chars (256 bits)
shortCode: string; // 6 chars base62 (for URL path)
createdAt: number; // Date.now()
consumed: boolean; // single-use flag
}
const QR_TOKEN_TTL_MS = 60_000; // 60 seconds
const QR_TOKEN_GRACE_MS = 90_000; // 90s grace for previous token (scan-during-rotation)
const SHORT_CODE_LENGTH = 6;
const QR_RATE_LIMIT_MAX = 30; // global rate limit across all IPs
const QR_RATE_LIMIT_WINDOW_MS = 60_000; // 1 minute window
/** Rejection-sampled short code generation — no modulo bias */
function generateShortCode(): string {
const chars = 'ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789';
const maxUnbiased = 248; // largest multiple of 62 that fits in a byte (248 = 62 * 4)
const result: string[] = [];
while (result.length < SHORT_CODE_LENGTH) {
const [byte] = randomBytes(1);
if (byte < maxUnbiased) result.push(chars[byte % 62]);
// else: discard and re-draw (rejection sampling)
}
return result.join('');
}
export class TunnelManager extends EventEmitter {
// Map-based lookup: shortCode → QrTokenRecord (timing-safe, no string comparison)
private qrTokensByCode = new Map<string, QrTokenRecord>();
private currentShortCode: string | null = null;
private rotationTimer: ReturnType<typeof setInterval> | null = null;
// SVG cache — regenerated only on token rotation, not per request
private cachedQrSvg: { shortCode: string; svg: string } | null = null;
// Global rate limit counter (separate from Basic Auth rate limiting)
private qrAttemptCount = 0;
private qrRateLimitResetTimer: ReturnType<typeof setInterval> | null = null;
constructor() {
super();
this.rotateToken();
this.rotationTimer = setInterval(() => this.rotateToken(), QR_TOKEN_TTL_MS);
this.qrRateLimitResetTimer = setInterval(() => { this.qrAttemptCount = 0; }, QR_RATE_LIMIT_WINDOW_MS);
}
private rotateToken(): void {
const record: QrTokenRecord = {
token: randomBytes(32).toString('hex'),
shortCode: generateShortCode(),
createdAt: Date.now(),
consumed: false,
};
// Evict expired tokens from the Map
const now = Date.now();
for (const [code, rec] of this.qrTokensByCode) {
if (now - rec.createdAt > QR_TOKEN_GRACE_MS || rec.consumed) {
this.qrTokensByCode.delete(code);
}
}
this.qrTokensByCode.set(record.shortCode, record);
this.currentShortCode = record.shortCode;
this.cachedQrSvg = null; // invalidate SVG cache
this.emit('qrTokenRotated');
}
/** Get the current (newest) token's short code for QR URL */
getCurrentShortCode(): string | undefined {
return this.currentShortCode ?? undefined;
}
/** Get cached QR SVG, regenerating only if the short code changed */
async getQrSvg(tunnelUrl: string): Promise<string> {
const code = this.currentShortCode;
if (!code) throw new Error('No QR token available');
if (this.cachedQrSvg?.shortCode === code) return this.cachedQrSvg.svg;
const QRCode = require('qrcode');
const svg = await QRCode.toString(`${tunnelUrl}/q/${code}`, { type: 'svg', margin: 2, width: 256 });
this.cachedQrSvg = { shortCode: code, svg };
return svg;
}
/**
* Validate and atomically consume a token by short code.
* Returns { success, ip?, ua? } for audit logging on success.
* Map.get() is hash-based — no timing side-channel from string comparison.
*/
consumeToken(shortCode: string): boolean {
// Global rate limit (across all IPs)
if (this.qrAttemptCount >= QR_RATE_LIMIT_MAX) return false;
this.qrAttemptCount++;
const record = this.qrTokensByCode.get(shortCode);
if (!record) return false;
if (record.consumed) return false;
const now = Date.now();
if (now - record.createdAt > QR_TOKEN_GRACE_MS) return false;
// Atomic consume (single-threaded JS = no race)
record.consumed = true;
// Immediately rotate so desktop gets a fresh QR
this.rotateToken();
this.emit('qrTokenRegenerated');
return true;
}
/** Force-regenerate (manual revocation via API) */
regenerateQrToken(): void {
// Invalidate all existing tokens
this.qrTokensByCode.clear();
this.currentShortCode = null;
this.rotateToken();
this.emit('qrTokenRegenerated');
}
stopRotation(): void {
if (this.rotationTimer) {
clearInterval(this.rotationTimer);
this.rotationTimer = null;
}
if (this.qrRateLimitResetTimer) {
clearInterval(this.qrRateLimitResetTimer);
this.qrRateLimitResetTimer = null;
}
}
}
```
### 2. Auth Middleware Bypass — `src/web/middleware/auth.ts`
Add `/q/` to the bypass list (same pattern as `/api/hook-event`). The route handler itself handles token validation and rate limiting.
```typescript
// In the onRequest hook, add before Basic Auth check:
if (req.url.startsWith('/q/')) {
done(); // Let the route handler deal with token validation
return;
}
```
**Important**: Unlike `/api/hook-event` (localhost-only), `/q/` must be reachable from any IP (remote devices scan the QR). Rate limiting is handled by two independent mechanisms:
1. **Per-IP rate limit** — reuses the `authFailures` StaleExpirationMap (10 attempts/IP/15min), but tracked via a **separate counter** from Basic Auth failures (so a user who fat-fingers their password doesn't burn their QR attempts)
2. **Global path rate limit** — `TunnelManager.qrAttemptCount` caps total QR attempts to 30/minute across all IPs, defending against distributed brute force
### 3. Auto-Auth Route — `src/web/routes/system-routes.ts`
Add `GET /q/:code` as a top-level route (not under `/api/`):
```typescript
app.get('/q/:code', async (req, reply) => {
const shortCode = (req.params as { code: string }).code;
const authPassword = process.env.CODEMAN_PASSWORD;
// No point if auth isn't enabled
if (!authPassword) {
return reply.redirect('/');
}
// Per-IP rate limit (separate counter from Basic Auth failures)
const clientIp = req.ip;
const qrFailures = ctx.authState.qrAuthFailures?.get(clientIp) ?? 0;
if (qrFailures >= 10) {
return reply.code(429).send('Too Many Requests');
}
// Validate and atomically consume the token
// consumeToken() also checks the global rate limit (30/min across all IPs)
if (!shortCode || !ctx.tunnelManager.consumeToken(shortCode)) {
ctx.authState.qrAuthFailures?.set(clientIp, qrFailures + 1);
return reply.code(401).send('Invalid or expired QR code');
}
// Issue session cookie (same as Basic Auth success path)
const sessionToken = randomBytes(32).toString('hex');
const clientUA = req.headers['user-agent'] ?? '';
ctx.authState.authSessions?.set(sessionToken, {
ip: clientIp,
ua: clientUA,
createdAt: Date.now(),
});
ctx.authState.qrAuthFailures?.delete(clientIp);
// Audit log — write to session-lifecycle.jsonl for forensic analysis
ctx.lifecycleLog?.append({
event: 'qr_auth',
ip: clientIp,
ua: clientUA,
timestamp: Date.now(),
shortCodePrefix: shortCode.slice(0, 3) + '***', // partial for privacy
});
reply.setCookie(AUTH_COOKIE_NAME, sessionToken, {
httpOnly: true,
secure: ctx.https,
sameSite: 'lax',
maxAge: 86400, // 24h
path: '/',
});
// Broadcast auth notification — desktop sees who authenticated (QRLjacking detection)
broadcast('tunnel:qrAuthUsed', {
ip: clientIp,
ua: clientUA,
timestamp: Date.now(),
});
return reply.redirect('/');
});
```
### 4. Update QR Code URL — `src/web/routes/system-routes.ts`
Modify `/api/tunnel/qr` to encode the short-code URL. Uses the `TunnelManager.getQrSvg()` cache — SVG is regenerated only when the token rotates, not on every request.
```typescript
app.get('/api/tunnel/qr', async (_req, reply) => {
const url = ctx.tunnelManager.getUrl();
if (!url) {
return reply.code(404).send(createErrorResponse(ApiErrorCode.NOT_FOUND, 'Tunnel not running'));
}
const authPassword = process.env.CODEMAN_PASSWORD;
// If auth is enabled, use the cached SVG with embedded short code
if (authPassword) {
const svg = await ctx.tunnelManager.getQrSvg(url);
return { svg, authEnabled: true };
}
// No auth — just encode the raw tunnel URL
const QRCode = require('qrcode');
const svg = await QRCode.toString(url, { type: 'svg', margin: 2, width: 256 });
return { svg, authEnabled: false };
});
```
### 5. Token Regeneration Endpoint — `src/web/routes/system-routes.ts`
Manual revocation — invalidates ALL existing tokens and creates a fresh one:
```typescript
app.post('/api/tunnel/qr/regenerate', async () => {
ctx.tunnelManager.regenerateQrToken();
return { success: true };
});
```
### 6. Frontend Updates — `src/web/public/app.js`
#### QR Overlay Changes
- **Auto-refresh via inline SVG**: Listen for `tunnel:qrRotated` SSE events which now include the SVG directly in the payload — no extra HTTP fetch needed, sub-50ms refresh on desktop.
- **Countdown indicator**: Small "expires in Xs" text under the QR that counts down from 60. Reassures the user the QR is live and not stale.
- **Regenerate button**: "Regenerate QR" button. Calls `POST /api/tunnel/qr/regenerate` — SSE event delivers the new SVG.
- **Auth badge**: Lock icon or "Single-use auth" label when auth is active.
- **URL display**: Show the raw tunnel URL (not the auth URL) for manual copy — users who copy the URL authenticate via Basic Auth. The QR is the fast path.
- **Auth notification toast**: When `tunnel:qrAuthUsed` fires, show a 10-second toast: "Device [IP] authenticated via QR (Safari). Not you? [Revoke]". This is the primary QRLjacking detection mechanism (USENIX Flaw-5).
```javascript
// Auto-refresh QR on rotation — SVG is inline in the event payload
addListener('tunnel:qrRotated', (data) => {
if (data.svg) {
updateQrDisplay(data.svg); // direct DOM update, no fetch
} else {
refreshTunnelQR(); // fallback: fetch from API
}
});
// Also refresh on manual regeneration
addListener('tunnel:qrRegenerated', (data) => {
if (data.svg) {
updateQrDisplay(data.svg);
} else {
refreshTunnelQR();
}
});
// QRLjacking detection — notify desktop user when QR is consumed
addListener('tunnel:qrAuthUsed', (data) => {
showNotificationToast(
`Device authenticated via QR (${parseUAFamily(data.ua)}, ${data.ip}). Not you?`,
{
duration: 10000,
action: { label: 'Revoke', onClick: () => revokeAllSessions() },
}
);
});
// In showTunnelQR(), after fetching /api/tunnel/qr:
if (data.authEnabled) {
const badge = document.createElement('div');
badge.textContent = 'Single-use auth \u00b7 refreshes every 60s';
badge.style.cssText = 'margin-top:8px;font-size:11px;color:var(--text-secondary)';
container.parentElement.appendChild(badge);
}
```
#### Welcome Screen QR
Same auto-refresh behavior applies to `_updateWelcomeTunnelBtn()` — the QR is fetched from `/api/tunnel/qr` so token embedding happens automatically.
### 7. SSE Events
Three events for the frontend. QR rotation events embed the SVG directly in the payload to eliminate an extra HTTP fetch — the desktop gets the new QR in a single SSE push (~2-5KB SVG, well within SSE limits).
```typescript
// In server.ts, listen for tunnelManager events:
// Auto-rotation every 60s — desktop refreshes QR silently (SVG inline)
tunnelManager.on('qrTokenRotated', async () => {
const url = tunnelManager.getUrl();
if (url && process.env.CODEMAN_PASSWORD) {
const svg = await tunnelManager.getQrSvg(url);
broadcast('tunnel:qrRotated', { svg });
} else {
broadcast('tunnel:qrRotated', {});
}
});
// Manual regeneration or post-consumption — desktop refreshes QR (SVG inline)
tunnelManager.on('qrTokenRegenerated', async () => {
const url = tunnelManager.getUrl();
if (url && process.env.CODEMAN_PASSWORD) {
const svg = await tunnelManager.getQrSvg(url);
broadcast('tunnel:qrRegenerated', { svg });
} else {
broadcast('tunnel:qrRegenerated', {});
}
});
// QR auth consumed — desktop shows notification toast (QRLjacking detection)
// Note: this is broadcast from the route handler, not tunnelManager
// Event: tunnel:qrAuthUsed { ip, ua, timestamp }
```
### 8. Session Cookie Binding & Revocation
Enhance session records to include device context for audit purposes. The UA is stored for **logging only** — not for blocking.
**Why no UA-family blocking (`majorUAChanged`)?** Security review found this is security theater:
- UA strings are trivially spoofable by any attacker who can steal a cookie
- Chrome UA reduction (2022+) makes family detection unreliable
- Mobile WebView → browser switches trigger false positives on the same device
- HttpOnly + Secure + SameSite=lax + 24h TTL already protect against cookie theft
- The attacker who can exfiltrate a cookie can also replay the exact UA
Instead, provide **manual session revocation** as the active defense:
```typescript
// Session record stores device context for audit logging (not blocking):
ctx.authState.authSessions?.set(sessionToken, {
ip: clientIp,
ua: req.headers['user-agent'] ?? '',
createdAt: Date.now(),
method: 'qr', // 'qr' | 'basic' — tracks how session was created
});
// Manual revocation endpoint — kill specific session or all sessions
app.post('/api/auth/revoke', async (req, reply) => {
const { sessionToken: target } = req.body as { sessionToken?: string };
if (target) {
ctx.authState.authSessions?.delete(target);
} else {
// Revoke all sessions (nuclear option)
ctx.authState.authSessions?.clear();
}
return { success: true };
});
```
**Note**: This is a breaking type change. The `AuthState` interface must be updated from `StaleExpirationMap<string, string>` (token → clientIp) to `StaleExpirationMap<string, { ip, ua, createdAt, method }>`. All session validation code in `auth.ts` must be updated simultaneously.
### 9. Cleanup — `src/tunnel-manager.ts`
Stop the rotation timer in the `stop()` method:
```typescript
stop(): void {
this.stopRotation();
// ... existing cleanup
}
```
## Security Analysis
### Threat Model
| Threat | Attack Vector | Mitigation | Residual Risk |
|--------|--------------|------------|---------------|
| **QR screenshot shared** | Attacker gets image of QR code | Single-use: token consumed on first scan. 60s TTL: expired by the time attacker tries. Desktop toast notification alerts user if someone else scans. | If attacker scans faster than legitimate user (~seconds), they win the race. Low risk: requires physical proximity + speed. User sees notification and can revoke. |
| **Cloudflare edge logs** | Cloudflare logs the full URL path | Short code is opaque (6-char lookup key), not the real token. Single-use: replaying from logs always fails. 60s TTL (90s grace): expired before log review. `trycloudflare.com` quick tunnels have no customer-accessible logging controls — the privacy implications are inherent to using free quick tunnels. | Cloudflare has TLS termination access regardless. Ephemeral short codes are far less valuable than a permanent token. |
| **Brute force short code** | Attacker guesses `/q/XXXXXX` | Per-IP rate limiting (10/IP/15min) + global path rate limit (30/min across all IPs). 62^6 = 56.8 billion combinations. Only ~2 valid codes at any time. | Infeasible: expected guesses to hit = ~2.8×10^10, rate limits block well before. |
| **Replay attack** | Reuse a previously valid URL | Single-use consumption + 60s TTL (90s grace). Old codes always 401. | None — replay is impossible by design. |
| **QRLjacking** | Attacker displays your QR on phishing site | No companion app = limited mitigation. However: 60s rotation means attacker must relay in real-time. Desktop toast notification ("Device [IP] authenticated via QR. Not you? [Revoke]") provides real-time detection. Self-hosted single-user context makes phishing implausible. | Theoretical risk for multi-user deployments. Mitigated by notification toast for single-user. Note: Signal's linked-device QR flow was exploited by Russian state actors (UNC5792/Sandworm) via quishing in 2025 — but that targeted a multi-user messaging platform, not a self-hosted dev tool. |
| **Session cookie theft** | XSS or network sniffing steals cookie | HttpOnly + Secure flags. SameSite=lax prevents CSRF. 24h TTL limits exposure window. Manual revocation via `/api/auth/revoke`. | Standard web cookie risks apply. Mitigated by security headers (CSP, etc.). |
| **Token in server logs** | Access log captures URL path | Log `/q/*` with short code masked or omitted. Configure Fastify logger to redact `/q/` paths. | Path still appears in server access logs (mitigated by masking). |
| **Timing attack** | Measure response time to leak short code | Map-based lookup (`Map.get()`) — hash-based O(1), no character-by-character timing leak. No string comparison in the hot path. | None — timing side channel eliminated by design. |
| **Token not in query params** | N/A (this is a mitigation) | Short code in URL path avoids browser history, Referer headers, and address bar exposure. | Path still appears in server access logs (mitigated by masking). |
| **Distributed brute force** | Multiple IPs guess codes simultaneously | Global rate limit (30/min total across all IPs) in addition to per-IP limit. | Infeasible given keyspace. Global limit prevents botnet-scale attempts. |
| **CSRF on regenerate** | Cross-origin POST to `/api/tunnel/qr/regenerate` | SameSite=lax cookies are NOT sent with cross-origin POST requests, providing CSRF protection. Endpoint requires authenticated session. | Verify SameSite=lax behavior through cloudflared tunnel. |
### USENIX Security 2025 Flaw Coverage
The [Zhang et al. paper](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) (USENIX Security 2025, 47 of top-100 websites vulnerable, 42 CVEs) identified 6 critical design flaws. Coverage:
| USENIX Flaw | Status | Implementation |
|-------------|--------|----------------|
| Flaw-1: Missing single-use enforcement | **Fixed** | Atomic `consumed` flag, Map-based lookup |
| Flaw-2: Long-lived tokens | **Fixed** | 60s TTL, 90s grace, auto-rotation |
| Flaw-3: Predictable QrId generation | **Fixed** | `crypto.randomBytes(32)` — 256-bit entropy, rejection-sampled short codes |
| Flaw-4: Client-side QrId generation | **Fixed** | Server-side generation only |
| Flaw-5: Missing status notification | **Fixed** | Desktop toast notification via `tunnel:qrAuthUsed` SSE event. Shows device IP/UA with [Revoke] button. |
| Flaw-6: Inadequate session binding | **Partial** | IP + UA stored for audit. No cryptographic channel binding (requires companion app / FIDO2 — overkill for single-user). Manual revocation as active defense. |
### Industry Comparison
| Platform | Model | How This Plan Compares |
|----------|-------|----------------------|
| **Discord** | Long-lived session token, no confirmation, repeatedly exploited via QRLjacking | **Better** — single-use + TTL + notification toast |
| **WhatsApp Web** | Pre-authenticated phone confirms "Link device?", ~60s rotation | **Comparable** rotation model; missing WhatsApp's explicit confirmation prompt (acceptable: single-user, no account selection) |
| **Signal** | Ephemeral public key in QR, E2E encrypted channel via Signal protocol | **Below** — no cryptographic channel binding. Note: Signal's QR flow was exploited by state actors in 2025 despite stronger crypto, showing that protocol strength alone doesn't prevent social engineering. |
| **1Password** | Noise framework E2E channel, post-quantum pre-shared keys, confirmation codes | **Below** — but 1Password is a credential manager with different threat model. Overkill for a dev tool. |
| **FIDO2 CTAP 2.2** | BLE proximity + cryptographic binding + biometric verification | **Below** — but requires BLE stack, FIDO server, and companion authenticator. Completely inappropriate here. |
### Comparison to Prior Design
| Property | Original Plan | Current Plan |
|----------|--------------|--------------|
| Token TTL | Infinite (until restart) | 60 seconds (90s grace for previous token) |
| Reuse | Multi-use (same QR works forever) | Single-use (consumed atomically on first scan) |
| Secret in URL | Query param (`?t=64-char-hex`) | Opaque short code in path (`/q/Xk9mQ3`) |
| Leak impact | Permanent access until manual revoke | Worthless after first use or 90s, whichever comes first |
| Desktop QR refresh | Manual only | Auto-refresh every 60s via SSE with inline SVG |
| Session binding | IP only | IP + UA stored for audit (not blocking). Manual revocation endpoint. |
| Auth notification | None | Desktop toast: "Device [IP] authenticated via QR. Not you? [Revoke]" |
| Audit logging | None | `session-lifecycle.jsonl` entry on every QR auth event |
| Rate limiting | Per-IP only, shared with Basic Auth | Per-IP (separate counter) + global path limit (30/min) |
| Short code generation | Modulo-biased | Rejection-sampled (no bias) |
| Short code lookup | Array scan (timing leak) | Map-based O(1) (timing-safe) |
| Connect latency | ~50ms (localhost only) | ~150-300ms through Cloudflare tunnel (honest estimate) |
### What This Does NOT Protect Against
- **FIDO2/passkey-level phishing resistance**: Would require BLE proximity verification and cryptographic channel binding. Overkill for a self-hosted single-user dev tool. The FIDO2 CTAP 2.2 hybrid transport is the gold standard but requires BLE hardware and a companion authenticator.
- **Compromised phone**: If the attacker has physical access to the phone that scans, no QR scheme helps.
- **Compromised Cloudflare tunnel**: Cloudflare terminates TLS and can inspect all traffic. This is inherent to using `trycloudflare.com` quick tunnels — use `--https` for end-to-end encryption if this matters.
- **State-sponsored quishing**: Sophisticated attackers could create convincing phishing pages that relay the QR in real-time. The 60s rotation and desktop notification toast mitigate this for the single-user case, but a dedicated attacker with social engineering could theoretically succeed within the TTL window.
### Standards Compliance Note
This design is **inspired by but does not conform to** [OASIS SQRAP v1.0](https://docs.oasis-open.org/esat/sqrap/v1.0/cs01/sqrap-v1.0-cs01.html). SQRAP's architecture requires a companion mobile app with stored identity keys, public key channel binding, back-channel authentication, and user presence verification (biometric/PIN). These are fundamentally incompatible with a browser-scan-to-authenticate flow. SQRAP is referenced for awareness of formal QR auth standards, not as a compliance target.
## Performance
The design prioritizes speed on connect. Latency depends on whether the request goes through a Cloudflare tunnel or is localhost:
### Localhost (no tunnel)
| Step | Latency |
|------|---------|
| QR scan (physical) | ~1-2s (user action) |
| `GET /q/:code` → Map.get() lookup + consume | <1ms |
| Cookie set + 302 redirect | <1ms |
| Browser follows redirect to `/` | <5ms |
| **Total (after scan)** | **<10ms** |
### Through Cloudflare Tunnel (typical mobile use case)
Each request traverses: phone → Cloudflare edge (TLS termination) → cloudflared → localhost. The 302 redirect means **two full round trips** through the tunnel.
| Step | Latency |
|------|---------|
| QR scan (physical) | ~1-2s (user action) |
| DNS resolution for `*.trycloudflare.com` | 20-80ms (first request, cached after) |
| TLS handshake to Cloudflare edge | 50-100ms (first request, 0 with TLS resumption) |
| `GET /q/:code` through tunnel (request + response) | 30-90ms |
| Browser follows 302 redirect: `GET /` through tunnel | 30-90ms |
| **Total first connection (cold)** | **~200-400ms** |
| **Total subsequent (TLS/DNS cached)** | **~100-200ms** |
This is still fast — **imperceptible after the 1-2s physical QR scan action**. For comparison, VS Code Remote Tunnels (through Azure) adds 20-100ms per hop.
### Why Not Eliminate the Redirect?
The 302 means two round trips. Alternatives considered:
- **200 + serve `index.html` directly**: URL bar shows `/q/Xk9mQ3`, relative paths break, couples auth to static serving. Not worth the complexity.
- **200 + `<meta http-equiv="refresh">`**: Still two requests, plus HTML parse delay. Actually slower.
- **200 + JavaScript redirect**: Same problem, plus fails if JS disabled.
The 302 is clean, universally supported, and the extra 30-90ms is invisible to users.
### QR Code Size Optimization
The URL `https://xxx-yyy.trycloudflare.com/q/Xk9mQ3` is ~53-56 characters. At QR Error Correction Level M:
| QR Version | Grid Size | Byte Capacity | Fits? |
|------------|-----------|---------------|-------|
| Version 3 | 29x29 | 42 bytes | No |
| Version 4 | 33x33 | 62 bytes | Yes (comfortably) |
| Version 5 | 37x37 | 84 bytes | Yes |
The shortened `/q/` path (vs `/qr-auth/`) and 6-char code (vs 8-char) save 9 bytes, targeting Version 4 (33x33) for faster scanning on budget Android phones. Modern phones scan Version 4 QR codes in 100-300ms — the user action of pointing the camera dominates.
### Desktop QR Refresh
Token rotation SSE events now embed the SVG directly in the payload (~2-5KB). The desktop gets the new QR in a single SSE push — no extra HTTP fetch needed. Refresh latency: **sub-50ms** (SSE adaptive batching at 16-50ms).
### SVG Caching
QR SVG is cached per rotation cycle on `TunnelManager.cachedQrSvg`. The SVG is regenerated only when the token rotates (every 60s), not on every `/api/tunnel/qr` request. SVG format is optimal: resolution-independent (retina-safe), inline-able (no extra HTTP request), ~2-5KB, renders in <1ms.
## Edge Cases
1. **Scan during rotation**: The server keeps 2 tokens (current + previous). If the user scans right as rotation happens, the previous token is still valid for up to 60s more. Seamless.
2. **Server restart**: All tokens cleared (in-memory). New token generated immediately. Tunnel URL also changes (trycloudflare gives a new subdomain), so old QR codes are doubly dead.
3. **Multiple devices**: Each scan consumes the token and triggers a fresh one. To auth a second device, wait for the QR to refresh (≤60s) or hit "Regenerate QR" on the desktop, then scan the new code.
4. **Token without tunnel**: `/qr-auth/:code` works even on localhost. If you have the code and it's valid, you get authenticated regardless of access method.
5. **Tunnel restart (same server)**: Tokens survive tunnel restarts (stored on `TunnelManager` instance). But new tunnel URL = new QR code generated. Short code stays valid until consumed or expired.
6. **Desktop browser closed during scan**: Token is consumed server-side. The scanning phone gets authenticated. When the desktop reopens, SSE reconnects and shows a fresh QR. No state corruption.
7. **Race condition: two phones scan same QR**: First scanner wins (atomic `consumed = true`). Second scanner gets 401. This is correct behavior — single-use by design.
## Files to Modify
| File | Changes |
|------|---------|
| `src/tunnel-manager.ts` | `QrTokenRecord` type, `Map<shortCode, record>` token pool, rejection-sampled `generateShortCode()`, rotation timer, `consumeToken()`, `getCurrentShortCode()`, `getQrSvg()` (cached), `regenerateQrToken()`, global rate limit counter, cleanup in `stop()` |
| `src/web/middleware/auth.ts` | Add `/q/` bypass in `onRequest` hook. Enhance session record type from `string` to `{ ip, ua, createdAt, method }` (**breaking type change** — all consumers must update). Add `qrAuthFailures` StaleExpirationMap (separate from Basic Auth `authFailures`). |
| `src/web/routes/system-routes.ts` | Modify `/api/tunnel/qr` to use `getQrSvg()` cache. Add `GET /q/:code` with atomic consume, audit log, and `tunnel:qrAuthUsed` broadcast. Add `POST /api/tunnel/qr/regenerate`. Add `POST /api/auth/revoke`. |
| `src/web/server.ts` | Pass `authState` + `lifecycleLog` to route context. Listen for `qrTokenRotated` and `qrTokenRegenerated` events → broadcast SSE with inline SVG. |
| `src/web/public/app.js` | Auto-refresh QR from inline SSE SVG payload (no extra fetch). Countdown timer. Regenerate button. Auth badge. Auth notification toast on `tunnel:qrAuthUsed` with [Revoke] action. |
| `src/session-lifecycle-log.ts` | Add `qr_auth` event type to lifecycle log schema |
| `src/types/api.ts` | Update `AuthState` interface: `authSessions` value type, add `qrAuthFailures` map |
## Complexity Estimate
Medium change. Core logic (Map-based token pool, rejection-sampled short codes, SVG cache, atomic consumption, cookie issuance, audit logging) is ~120 lines. Rate limiting (separate QR counter + global path limit) adds ~20 lines. SSE plumbing with inline SVG adds ~30 lines. Frontend (inline SVG refresh, auth notification toast with revoke, countdown) is ~40 lines. Auth type migration (session record type change) touches ~10 lines across middleware. No new dependencies — `crypto` and `qrcode` are already available.
## Testing
### Automated
```bash
# Unit test for token manager
npx vitest run test/qr-auth.test.ts
```
Test cases:
- Token rotation generates unique short codes (6-char, base62)
- Short codes have uniform character distribution (no modulo bias — verify with chi-squared test over 10K samples)
- `consumeToken()` returns true on first use, false on second
- Expired tokens (>90s old) return false
- Previous token still works during 90s grace period
- Token at exactly 60s still valid (within grace), token at 91s rejected
- `regenerateQrToken()` invalidates all existing tokens (Map cleared)
- Short code lookup is case-sensitive
- Per-IP rate limiting increments on invalid codes (separate from Basic Auth counter)
- Global rate limit (30/min) blocks attempts across all IPs
- SVG cache returns same string for same short code, regenerates on rotation
- Audit log entry written on successful QR auth
- `tunnel:qrAuthUsed` SSE event broadcast on successful QR auth
- `tunnel:qrRotated` SSE event includes inline SVG payload
- Map-based lookup does not leak timing information (no string comparison in hot path)
### Manual
1. Start server with `CODEMAN_PASSWORD=test`
2. Enable tunnel
3. Verify `/api/tunnel/qr` returns QR encoding `https://...trycloudflare.com/q/Xk9mQ3`
4. Open the QR URL in incognito → should auto-redirect to `/` with session cookie
5. Verify desktop shows notification toast: "Device [IP] authenticated via QR"
6. Open the **same** URL again → should get 401 (single-use consumed)
7. Wait 60s → verify QR display auto-updated (new short code, inline SVG via SSE)
8. Open just the tunnel URL → should get Basic Auth prompt
9. Call `POST /api/tunnel/qr/regenerate` → old QR URL returns 401, new QR appears
10. Verify per-IP rate limiting: 10+ failed `/q/badcode` → 429
11. Verify Basic Auth failures don't consume QR rate limit budget (and vice versa)
12. Check `~/.codeman/session-lifecycle.jsonl` for `qr_auth` entries after successful scan
13. Click [Revoke] on the notification toast → verify session is invalidated
## References
- [USENIX Security 2025: "Demystifying the (In)Security of QR Code-based Login in Real-world Deployments"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) — 6 design flaws, 5 attack types, 42 CVEs across 47 of top-100 websites. Primary design reference for this plan.
- [OWASP QRLJacking](https://owasp.org/www-community/attacks/Qrljacking) — canonical QR session hijacking reference
- [OASIS SQRAP v1.0 Standard](https://docs.oasis-open.org/esat/sqrap/v1.0/cs01/sqrap-v1.0-cs01.html) — formal standard for secure QR authentication. **Not a compliance target** for this plan (requires companion app + PKI). Referenced for awareness only.
- [FIDO2 CTAP 2.2 Hybrid Transport](https://fidoalliance.org/specs/fido-v2.2-rd-20230321/fido-client-to-authenticator-protocol-v2.2-rd-20230321.html) — gold standard for cross-device auth (overkill for this use case)
- [Google GTIG: Signal QR quishing by Russian state actors (2025)](https://cloud.google.com/blog/topics/threat-intelligence/russia-targeting-signal-messenger) — UNC5792/Sandworm exploited Signal's linked-device QR flow via phishing. Demonstrates that even cryptographically strong QR auth can be defeated by social engineering.
- [CVE-2026-2144: Magic Login QR Code Plugin race condition](https://www.cvedetails.com/cve/CVE-2026-2144/) — QR token stored as predictable static file, race window between creation and deletion. Validates this plan's in-memory-only approach.
Binary file not shown.

After

Width:  |  Height:  |  Size: 894 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 576 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 390 KiB

+273 -53
View File
@@ -9,6 +9,8 @@
# CODEMAN_INSTALL_DIR - Custom install directory (default: ~/.codeman/app)
# CODEMAN_SKIP_SYSTEMD=1 - Skip systemd service setup prompt
# CODEMAN_NODE_VERSION - Node.js major version to install (default: 22)
# CODEMAN_REPO_URL - Custom git repository URL (default: upstream Codeman)
# CODEMAN_BRANCH - Git branch to install (default: master)
set -euo pipefail
@@ -17,7 +19,8 @@ set -euo pipefail
# ============================================================================
INSTALL_DIR="${CODEMAN_INSTALL_DIR:-$HOME/.codeman/app}"
REPO_URL="https://github.com/Ark0N/Codeman.git"
REPO_URL="${CODEMAN_REPO_URL:-https://github.com/Ark0N/Codeman.git}"
BRANCH="${CODEMAN_BRANCH:-master}"
MIN_NODE_VERSION=18
TARGET_NODE_VERSION="${CODEMAN_NODE_VERSION:-22}"
NONINTERACTIVE="${CODEMAN_NONINTERACTIVE:-0}"
@@ -312,6 +315,32 @@ get_opencode_path() {
done
}
check_cloudflared() {
# Check ~/.local/bin first (matches tunnel-manager.ts resolution order)
if [[ -x "$HOME/.local/bin/cloudflared" ]]; then
return 0
fi
if [[ -x "/usr/local/bin/cloudflared" ]]; then
return 0
fi
if command -v cloudflared &>/dev/null; then
return 0
fi
return 1
}
get_cloudflared_path() {
if [[ -x "$HOME/.local/bin/cloudflared" ]]; then
echo "$HOME/.local/bin/cloudflared"
return
fi
if [[ -x "/usr/local/bin/cloudflared" ]]; then
echo "/usr/local/bin/cloudflared"
return
fi
command -v cloudflared 2>/dev/null
}
# ============================================================================
# Dependency Installation
# ============================================================================
@@ -541,6 +570,82 @@ install_git_suse() {
run_as_root zypper install -y git
}
install_cloudflared_macos() {
info "Installing cloudflared via Homebrew..."
ensure_homebrew
brew install cloudflared
}
install_cloudflared_debian() {
info "Installing cloudflared..."
ensure_sudo
local arch
arch="$(dpkg --print-architecture 2>/dev/null || echo "amd64")"
local tmp
tmp="$(mktemp)"
download "https://github.com/cloudflare/cloudflared/releases/latest/download/cloudflared-linux-$arch.deb" "$tmp"
run_as_root dpkg -i "$tmp"
rm -f "$tmp"
}
install_cloudflared_fedora() {
info "Installing cloudflared..."
ensure_sudo
local arch
arch="$(uname -m)"
local rpm_arch="$arch"
[[ "$arch" == "x86_64" ]] && rpm_arch="x86_64"
[[ "$arch" == "aarch64" ]] && rpm_arch="aarch64"
local tmp
tmp="$(mktemp)"
download "https://github.com/cloudflare/cloudflared/releases/latest/download/cloudflared-linux-$rpm_arch.rpm" "$tmp"
run_as_root rpm -i "$tmp" || run_as_root rpm -U "$tmp"
rm -f "$tmp"
}
install_cloudflared_arch() {
info "Installing cloudflared binary..."
local arch
arch="$(uname -m)"
local cf_arch="amd64"
[[ "$arch" == "aarch64" ]] && cf_arch="arm64"
[[ "$arch" == "armv7l" ]] && cf_arch="arm"
ensure_sudo
local tmp
tmp="$(mktemp)"
download "https://github.com/cloudflare/cloudflared/releases/latest/download/cloudflared-linux-$cf_arch" "$tmp"
run_as_root mv "$tmp" /usr/local/bin/cloudflared
run_as_root chmod +x /usr/local/bin/cloudflared
}
install_cloudflared_alpine() {
info "Installing cloudflared binary..."
local arch
arch="$(uname -m)"
local cf_arch="amd64"
[[ "$arch" == "aarch64" ]] && cf_arch="arm64"
[[ "$arch" == "armv7l" ]] && cf_arch="arm"
ensure_sudo
local tmp
tmp="$(mktemp)"
download "https://github.com/cloudflare/cloudflared/releases/latest/download/cloudflared-linux-$cf_arch" "$tmp"
run_as_root mv "$tmp" /usr/local/bin/cloudflared
run_as_root chmod +x /usr/local/bin/cloudflared
}
install_cloudflared_suse() {
info "Installing cloudflared..."
ensure_sudo
local arch
arch="$(uname -m)"
local rpm_arch="$arch"
local tmp
tmp="$(mktemp)"
download "https://github.com/cloudflare/cloudflared/releases/latest/download/cloudflared-linux-$rpm_arch.rpm" "$tmp"
run_as_root rpm -i "$tmp" || run_as_root rpm -U "$tmp"
rm -f "$tmp"
}
# ============================================================================
# Interactive Prompts
# ============================================================================
@@ -733,6 +838,22 @@ EOF
success "Systemd service installed and started"
}
setup_tunnel_service() {
local service_dir="$HOME/.config/systemd/user"
local service_file="$service_dir/codeman-tunnel.service"
info "Setting up Cloudflare tunnel systemd service..."
mkdir -p "$service_dir"
cp "$INSTALL_DIR/scripts/codeman-tunnel.service" "$service_file"
systemctl --user daemon-reload
systemctl --user enable codeman-tunnel.service 2>/dev/null || true
success "Tunnel service installed (start with: systemctl --user start codeman-tunnel)"
echo -e " ${DIM}Note: Set CODEMAN_PASSWORD env var before starting the tunnel for security.${NC}"
}
# ============================================================================
# Installation Helpers
# ============================================================================
@@ -915,6 +1036,24 @@ main() {
fi
fi
# cloudflared (optional — for remote/mobile access via Cloudflare Tunnel)
info "Checking cloudflared (optional, for remote access)..."
if check_cloudflared; then
success "cloudflared found at $(get_cloudflared_path)"
else
if prompt_yes_no "Install cloudflared? (enables remote/mobile access via Cloudflare Tunnel)" "n"; then
install_dependency "cloudflared" "$os" "$distro"
hash -r 2>/dev/null || true
if check_cloudflared; then
success "cloudflared installed at $(get_cloudflared_path)"
else
warn "cloudflared installation failed. You can install it manually later."
fi
else
info "Skipped (you can install cloudflared later for remote access)"
fi
fi
echo ""
# ========================================================================
@@ -926,26 +1065,27 @@ main() {
if [[ -d "$INSTALL_DIR/.git" ]]; then
info "Existing installation found, updating..."
cd "$INSTALL_DIR"
git remote set-url origin "$REPO_URL" 2>/dev/null || true
# Check for local changes
if ! git diff --quiet 2>/dev/null || ! git diff --staged --quiet 2>/dev/null; then
warn "Local changes detected in $INSTALL_DIR"
if prompt_yes_no "Discard local changes and update?" "n"; then
git fetch --quiet origin
git reset --hard origin/master --quiet
git reset --hard "origin/$BRANCH" --quiet
else
info "Keeping existing installation, skipping update"
fi
else
git fetch --quiet origin
git reset --hard origin/master --quiet
git reset --hard "origin/$BRANCH" --quiet
fi
else
# Create parent directory
mkdir -p "$(dirname "$INSTALL_DIR")"
# Clone repository (shallow for speed)
git clone --quiet --depth 1 "$REPO_URL" "$INSTALL_DIR"
git clone --quiet --depth 1 --branch "$BRANCH" "$REPO_URL" "$INSTALL_DIR"
cd "$INSTALL_DIR"
fi
@@ -989,18 +1129,7 @@ main() {
fi
# ========================================================================
# Systemd Service (Linux only)
# ========================================================================
if [[ "$os" == "linux" ]] && [[ "$SKIP_SYSTEMD" != "1" ]] && command -v systemctl &>/dev/null; then
echo ""
if prompt_yes_no "Set up systemd service for auto-start?" "n"; then
setup_systemd_service
fi
fi
# ========================================================================
# Success!
# Launch Options
# ========================================================================
echo ""
@@ -1009,13 +1138,73 @@ main() {
echo -e "${GREEN}${BOLD}============================================================${NC}"
echo ""
# Check if systemd service is running (we just started it above)
local service_running=false
if systemctl --user is-active codeman-web.service &>/dev/null; then
service_running=true
local launch_choice=""
local has_systemd=false
if [[ "$os" == "linux" ]] && [[ "$SKIP_SYSTEMD" != "1" ]] && command -v systemctl &>/dev/null; then
has_systemd=true
fi
if [[ "$service_running" == "true" ]]; then
if [[ "$has_systemd" == "true" ]]; then
echo -e " ${BOLD}How would you like to run Codeman?${NC}"
echo ""
echo -e " ${CYAN}1)${NC} Run now in this terminal"
echo -e " ${CYAN}2)${NC} Install as systemd service (auto-start on boot)"
echo -e " ${CYAN}3)${NC} Don't start — I'll run it later"
echo ""
if [[ "$NONINTERACTIVE" == "1" ]] || [[ ! -t 0 ]]; then
launch_choice="3"
else
while true; do
echo -en "${CYAN}Choose [1/2/3]:${NC} " >&2
read -r launch_choice
case "$launch_choice" in
1|2|3) break ;;
*) echo "Please enter 1, 2, or 3." >&2 ;;
esac
done
fi
else
# macOS or no systemd — only offer run now or skip
echo -e " ${BOLD}Would you like to start Codeman now?${NC}"
echo ""
echo -e " ${CYAN}1)${NC} Run now in this terminal"
echo -e " ${CYAN}2)${NC} Don't start — I'll run it later"
echo ""
if [[ "$NONINTERACTIVE" == "1" ]] || [[ ! -t 0 ]]; then
launch_choice="2"
else
while true; do
echo -en "${CYAN}Choose [1/2]:${NC} " >&2
read -r launch_choice
case "$launch_choice" in
1) break ;;
2) break ;;
*) echo "Please enter 1 or 2." >&2 ;;
esac
done
fi
# Remap: no-systemd choice "2" (skip) → internal "3"
[[ "$launch_choice" == "2" ]] && launch_choice="3"
fi
echo ""
# Handle systemd setup
if [[ "$launch_choice" == "2" ]]; then
setup_systemd_service
# Offer tunnel service if cloudflared is available
if check_cloudflared && [[ -f "$INSTALL_DIR/scripts/codeman-tunnel.service" ]]; then
echo ""
if prompt_yes_no "Also set up Cloudflare tunnel service? (requires CODEMAN_PASSWORD)" "n"; then
setup_tunnel_service
fi
fi
echo ""
echo -e " ${GREEN}${BOLD}Codeman is running now!${NC}"
echo ""
echo -e " ${CYAN}# Open in browser${NC}"
@@ -1028,20 +1217,29 @@ main() {
echo -e " ${CYAN}systemctl --user status codeman-web${NC} # Check status"
echo -e " ${CYAN}journalctl --user -u codeman-web -f${NC} # View logs"
echo ""
else
fi
# Show quick-start help for non-service paths
if [[ "$launch_choice" != "2" ]]; then
echo -e " ${BOLD}Quick Start:${NC}"
echo ""
echo -e " ${CYAN}# Start the web server${NC}"
echo -e " codeman web"
echo ""
echo -e " ${CYAN}# Start with HTTPS (only needed for remote access)${NC}"
echo -e " codeman web --https"
echo -e " ${CYAN}codeman web${NC} # Start the web server"
echo -e " ${CYAN}codeman web --https${NC} # With HTTPS (for remote access)"
echo ""
echo -e " ${CYAN}# Open in browser${NC}"
echo -e " http://localhost:3000"
echo ""
fi
if check_cloudflared; then
echo -e " ${BOLD}Remote Access (Cloudflare Tunnel):${NC}"
echo ""
echo -e " ${CYAN}./scripts/tunnel.sh start${NC} # Start tunnel"
echo -e " ${CYAN}./scripts/tunnel.sh url${NC} # Show tunnel URL"
echo -e " ${CYAN}./scripts/tunnel.sh stop${NC} # Stop tunnel"
echo ""
fi
echo -e " ${BOLD}Mobile Access (Termius/SSH):${NC}"
echo ""
echo -e " ${CYAN}sc${NC} # Interactive tmux session chooser"
@@ -1060,16 +1258,19 @@ main() {
echo ""
fi
# Check if PATH needs reload in user's shell (only relevant if service not running)
if [[ "$service_running" != "true" ]]; then
# Run now in foreground (must be last — exec replaces the shell)
if [[ "$launch_choice" == "1" ]]; then
local profile
profile=$(detect_shell_profile)
if ! command -v codeman &>/dev/null 2>&1; then
echo -e " ${YELLOW}Run this to start using codeman now:${NC}"
echo ""
echo -e " ${CYAN}source $profile && codeman web${NC}"
echo ""
fi
echo -e " ${GREEN}${BOLD}Starting Codeman...${NC}"
echo -e " ${DIM}Press Ctrl+C to stop${NC}"
echo ""
# Source profile to pick up PATH changes, then exec codeman
# shellcheck disable=SC1090
source "$profile" 2>/dev/null || true
exec node "$INSTALL_DIR/dist/index.js" web
fi
}
@@ -1080,13 +1281,23 @@ update() {
info "Updating Codeman..."
cd "$INSTALL_DIR"
git remote set-url origin "$REPO_URL" 2>/dev/null || true
git fetch --quiet origin
git reset --hard origin/master --quiet
git reset --hard "origin/$BRANCH" --quiet
npm install --quiet --no-fund --no-audit 2>/dev/null || npm install --no-fund --no-audit
npm run build --quiet 2>/dev/null || npm run build
success "Updated to $(node -e "console.log(require('./package.json').version)")"
echo ""
echo -e " ${DIM}Restart codeman web to use the new version.${NC}"
# Auto-restart systemd service if it's running, otherwise tell the user
if systemctl --user is-active codeman-web.service &>/dev/null; then
info "Restarting codeman-web service..."
systemctl --user restart codeman-web.service
success "codeman-web service restarted"
else
echo -e " ${DIM}Restart codeman web to use the new version:${NC}"
echo -e " ${CYAN}pkill -f 'codeman.*web'; codeman web &${NC}"
fi
echo ""
}
@@ -1095,21 +1306,23 @@ uninstall() {
info "Uninstalling Codeman..."
echo ""
# Stop and remove systemd service
if systemctl --user is-active codeman-web.service &>/dev/null; then
info "Stopping codeman-web service..."
systemctl --user stop codeman-web.service
fi
if systemctl --user is-enabled codeman-web.service &>/dev/null 2>&1; then
info "Disabling codeman-web service..."
systemctl --user disable codeman-web.service 2>/dev/null || true
fi
local service_file="$HOME/.config/systemd/user/codeman-web.service"
if [[ -f "$service_file" ]]; then
rm -f "$service_file"
systemctl --user daemon-reload 2>/dev/null || true
success "Systemd service removed"
fi
# Stop and remove systemd services
for svc in codeman-web codeman-tunnel; do
if systemctl --user is-active "${svc}.service" &>/dev/null; then
info "Stopping ${svc} service..."
systemctl --user stop "${svc}.service"
fi
if systemctl --user is-enabled "${svc}.service" &>/dev/null 2>&1; then
info "Disabling ${svc} service..."
systemctl --user disable "${svc}.service" 2>/dev/null || true
fi
local svc_file="$HOME/.config/systemd/user/${svc}.service"
if [[ -f "$svc_file" ]]; then
rm -f "$svc_file"
success "Removed ${svc} service"
fi
done
systemctl --user daemon-reload 2>/dev/null || true
# Remove symlinks
local symlink_dir="$HOME/.local/bin"
@@ -1156,5 +1369,12 @@ uninstall() {
case "${1:-}" in
update) update ;;
uninstall) uninstall ;;
*) main "$@" ;;
*)
if [[ -z "${1:-}" && -d "$INSTALL_DIR/.git" ]]; then
print_banner
update
else
main "$@"
fi
;;
esac
+415 -258
View File
File diff suppressed because it is too large Load Diff
+12 -5
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "0.2.9",
"version": "0.4.0",
"description": "The missing control plane for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -51,6 +51,11 @@
"@fastify/compress": "^8.3.1",
"@fastify/cookie": "^11.0.2",
"@fastify/static": "^8.0.0",
"@fastify/websocket": "^11.2.0",
"@xterm/addon-fit": "^0.11.0",
"@xterm/addon-unicode11": "^0.9.0",
"@xterm/addon-webgl": "^0.19.0",
"@xterm/xterm": "^6.0.0",
"chalk": "^5.3.0",
"chokidar": "^3.6.0",
"commander": "^12.1.0",
@@ -59,10 +64,6 @@
"qrcode": "^1.5.4",
"uuid": "^10.0.0",
"web-push": "^3.6.7",
"xterm": "^5.3.0",
"xterm-addon-fit": "^0.8.0",
"xterm-addon-unicode11": "^0.6.0",
"xterm-addon-webgl": "^0.16.0",
"zod": "^4.3.6"
},
"devDependencies": {
@@ -72,9 +73,11 @@
"@remotion/transitions": "4.0.429",
"@types/node": "^20.19.33",
"@types/pngjs": "^6.0.5",
"@types/qrcode": "^1.5.6",
"@types/react": "^19.2.14",
"@types/uuid": "^10.0.0",
"@types/web-push": "^3.6.4",
"@types/ws": "^8.18.1",
"@vitest/coverage-v8": "^4.0.18",
"agent-browser": "^0.6.0",
"esbuild": "^0.27.3",
@@ -90,6 +93,10 @@
"typescript-eslint": "^8.0.0",
"vitest": "^4.0.18"
},
"optionalDependencies": {
"@remotion/compositor-linux-x64-gnu": "^4.0.432",
"@rspack/binding-linux-x64-gnu": "^1.7.7"
},
"engines": {
"node": ">=18.0.0"
},
+1 -1
View File
@@ -17,7 +17,7 @@
"dist/"
],
"scripts": {
"build": "tsup src/index.ts --format cjs,esm --dts --clean",
"build": "tsup",
"test": "vitest run",
"typecheck": "tsc --noEmit",
"prepublishOnly": "npm run build"
@@ -4,18 +4,26 @@ import type { XtermTerminal, CellDimensions } from './types.js';
* Get cell dimensions from the terminal, handling xterm.js v5 (private API)
* and v7+ (public API).
*
* Returns CSS-pixel values. xterm's `device.char` is in device pixels, so
* we divide by `devicePixelRatio` to stay consistent with `css.cell`.
*
* Returns `null` if the terminal is not yet rendered or dimensions are
* unavailable.
*/
export function getCellDimensions(terminal: XtermTerminal): CellDimensions | null {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const t = terminal as any;
const dpr = typeof devicePixelRatio === 'number' && devicePixelRatio > 0
? devicePixelRatio : 1;
// Try v7+ public API first
if (t.dimensions?.css?.cell) {
const cellH = t.dimensions.css.cell.height;
return {
width: t.dimensions.css.cell.width,
height: t.dimensions.css.cell.height,
height: cellH,
charTop: (t.dimensions?.device?.char?.top ?? 0) / dpr,
charHeight: (t.dimensions?.device?.char?.height ?? (cellH * dpr)) / dpr,
};
}
@@ -23,9 +31,12 @@ export function getCellDimensions(terminal: XtermTerminal): CellDimensions | nul
try {
const dims = t._core?._renderService?.dimensions;
if (dims?.css?.cell) {
const cellH = dims.css.cell.height;
return {
width: dims.css.cell.width,
height: dims.css.cell.height,
height: cellH,
charTop: (dims.device?.char?.top ?? 0) / dpr,
charHeight: (dims.device?.char?.height ?? (cellH * dpr)) / dpr,
};
}
} catch {
@@ -1,4 +1,50 @@
import type { RenderParams, FontStyle } from './types.js';
import type { RenderParams, FontStyle, XtermTerminal } from './types.js';
// ─── CJK / fullwidth character width detection ───────────────────────
/**
* Get visual cell width of a single character.
* CJK wide characters occupy 2 cells, others occupy 1.
* Prefers the terminal's Unicode addon when available.
*/
export function charCellWidth(terminal: XtermTerminal | null | undefined, ch: string): number {
if (terminal?.unicode?.getStringCellWidth) {
return terminal.unicode.getStringCellWidth(ch);
}
// Fallback: detect CJK wide characters by Unicode range
const code = ch.codePointAt(0);
if (
code !== undefined &&
code >= 0x1100 &&
(code <= 0x115f || // Hangul Jamo
(code >= 0x2e80 && code <= 0x303e) || // CJK Radicals, Kangxi, Ideographic
(code >= 0x3040 && code <= 0x33bf) || // Hiragana, Katakana, Bopomofo, CJK Compat
(code >= 0x3400 && code <= 0x4dbf) || // CJK Unified Ext A
(code >= 0x4e00 && code <= 0xa4cf) || // CJK Unified, Yi
(code >= 0xa960 && code <= 0xa97c) || // Hangul Jamo Extended-A
(code >= 0xac00 && code <= 0xd7a3) || // Hangul Syllables
(code >= 0xf900 && code <= 0xfaff) || // CJK Compat Ideographs
(code >= 0xfe30 && code <= 0xfe6f) || // CJK Compat Forms
(code >= 0xff01 && code <= 0xff60) || // Fullwidth Forms
(code >= 0xffe0 && code <= 0xffe6) || // Fullwidth Signs
(code >= 0x1f000 && code <= 0x1fbff) || // Mahjong, Domino, Emoji
(code >= 0x20000 && code <= 0x2ffff) || // CJK Unified Ext B-F
(code >= 0x30000 && code <= 0x3ffff)) // CJK Unified Ext G+
)
return 2;
return 1;
}
/**
* Get visual cell width of a string (sum of all character widths).
*/
export function stringCellWidth(terminal: XtermTerminal | null | undefined, str: string): number {
let w = 0;
for (const ch of str) w += charCellWidth(terminal, ch);
return w;
}
// ─── Overlay rendering ────────────────────────────────────────────────
/**
* Render the overlay content into the container element.
@@ -6,91 +52,114 @@ import type { RenderParams, FontStyle } from './types.js';
* Creates per-character `<span>` elements positioned on an exact grid
* matching xterm.js's canvas renderer. This avoids sub-pixel drift that
* occurs with normal DOM text flow.
*
* CJK wide characters are rendered with double-width spans.
*/
export function renderOverlay(container: HTMLDivElement, params: RenderParams): void {
const { lines, startCol, totalCols, cellW, cellH, promptRow, font, showCursor, cursorColor } = params;
const {
lines,
startCol,
totalCols,
cellW,
cellH,
charTop,
charHeight,
promptRow,
font,
showCursor,
cursorColor,
terminal,
} = params;
// Position container at prompt row
container.style.left = '0px';
container.style.top = (promptRow * cellH) + 'px';
// Position container at prompt row.
container.style.left = '0px';
container.style.top = promptRow * cellH + 'px';
// Clear and rebuild (typically 1-3 line divs, negligible cost)
container.innerHTML = '';
const fullWidthPx = totalCols * cellW;
// Clear and rebuild (typically 1-3 line divs, negligible cost)
container.innerHTML = '';
const fullWidthPx = totalCols * cellW;
for (let i = 0; i < lines.length; i++) {
const leftPx = i === 0 ? startCol * cellW : 0;
const widthPx = i === 0 ? (fullWidthPx - leftPx) : fullWidthPx;
const topPx = i * cellH;
const lineEl = makeLine(lines[i], leftPx, topPx, widthPx, cellH, cellW, font);
container.appendChild(lineEl);
for (let i = 0; i < lines.length; i++) {
const leftPx = i === 0 ? startCol * cellW : 0;
const widthPx = i === 0 ? fullWidthPx - leftPx : fullWidthPx;
const topPx = i * cellH;
const lineEl = makeLine(lines[i], leftPx, topPx, widthPx, cellH, cellW, charTop, charHeight, font, terminal);
container.appendChild(lineEl);
}
// Block cursor at end of last line (use visual width for CJK support)
if (showCursor) {
const lastLine = lines[lines.length - 1];
const lastLineLeft = lines.length === 1 ? startCol : 0;
const cursorCol = lastLineLeft + stringCellWidth(terminal, lastLine);
if (cursorCol < totalCols) {
const cursor = document.createElement('span');
cursor.style.cssText = 'position:absolute;display:inline-block';
cursor.style.left = cursorCol * cellW + 'px';
cursor.style.top = (lines.length - 1) * cellH + 'px';
cursor.style.width = cellW + 'px';
cursor.style.height = cellH + 'px';
cursor.style.backgroundColor = cursorColor;
container.appendChild(cursor);
}
}
// Block cursor at end of last line
if (showCursor) {
const lastLine = lines[lines.length - 1];
const lastLineLeft = lines.length === 1 ? startCol : 0;
const cursorCol = lastLineLeft + lastLine.length;
if (cursorCol < totalCols) {
const cursor = document.createElement('span');
cursor.style.cssText = 'position:absolute;display:inline-block';
cursor.style.left = (cursorCol * cellW) + 'px';
cursor.style.top = ((lines.length - 1) * cellH) + 'px';
cursor.style.width = cellW + 'px';
cursor.style.height = cellH + 'px';
cursor.style.backgroundColor = cursorColor;
container.appendChild(cursor);
}
}
container.style.display = '';
container.style.display = '';
}
/**
* Create a styled line `<div>` with per-character grid positioning.
*
* Each character gets its own `<span>` placed at `i * cellW` pixels.
* This matches xterm's canvas renderer where each glyph occupies exactly
* one cell width, regardless of the actual glyph metrics.
* Each character gets its own `<span>` positioned by visual column offset.
* CJK wide characters occupy 2 cell widths.
*/
function makeLine(
text: string,
leftPx: number,
topPx: number,
widthPx: number,
cellH: number,
cellW: number,
font: FontStyle,
text: string,
leftPx: number,
topPx: number,
widthPx: number,
cellH: number,
cellW: number,
_charTop: number,
_charHeight: number,
font: FontStyle,
terminal?: XtermTerminal | null
): HTMLDivElement {
const el = document.createElement('div');
el.style.cssText = 'position:absolute;pointer-events:none';
el.style.backgroundColor = font.backgroundColor;
el.style.left = leftPx + 'px';
el.style.top = topPx + 'px';
el.style.width = widthPx + 'px';
el.style.height = (cellH + 1) + 'px';
el.style.lineHeight = cellH + 'px';
const el = document.createElement('div');
el.style.cssText = 'position:absolute;pointer-events:none';
el.style.backgroundColor = font.backgroundColor;
el.style.left = leftPx + 'px';
el.style.top = topPx + 'px';
el.style.width = widthPx + 'px';
// Extend background 1px past cell boundary to cover the compositing
// seam between the overlay layer (z-index:7) and the canvas layer below.
// The extra 1px lands in the next row's charTop gap (empty area before
// text rendering starts), so no canvas content is obscured.
el.style.height = cellH + 1 + 'px';
for (let i = 0; i < text.length; i++) {
const span = document.createElement('span');
// Match xterm.js canvas text rendering:
// - antialiased smoothing (canvas uses grayscale, not LCD subpixel)
// - geometricPrecision for consistent glyph sizing
// - no ligatures (canvas renders each glyph independently)
span.style.cssText =
'position:absolute;display:inline-block;text-align:center;pointer-events:none;' +
'-webkit-font-smoothing:antialiased;-moz-osx-font-smoothing:grayscale;' +
"text-rendering:geometricPrecision;font-feature-settings:'liga' 0,'calt' 0";
span.style.left = (i * cellW) + 'px';
span.style.width = cellW + 'px';
span.style.fontFamily = font.fontFamily;
span.style.fontSize = font.fontSize;
span.style.fontWeight = font.fontWeight;
span.style.color = font.color;
if (font.letterSpacing) span.style.letterSpacing = font.letterSpacing;
span.textContent = text[i];
el.appendChild(span);
}
// CJK wide chars occupy 2 cells — position by visual column offset
let colOffset = 0;
for (const ch of text) {
const cw = charCellWidth(terminal, ch);
const span = document.createElement('span');
// No ligatures — canvas renders each glyph independently.
span.style.cssText =
'position:absolute;display:inline-block;text-align:center;pointer-events:none;' +
"font-feature-settings:'liga' 0,'calt' 0";
span.style.left = colOffset * cellW + 'px';
span.style.top = '0px';
span.style.width = cw * cellW + 'px';
span.style.height = cellH + 'px';
span.style.lineHeight = cellH + 'px';
span.style.fontFamily = font.fontFamily;
span.style.fontSize = font.fontSize;
span.style.fontWeight = font.fontWeight;
span.style.color = font.color;
if (font.letterSpacing) span.style.letterSpacing = font.letterSpacing;
span.textContent = ch;
el.appendChild(span);
colOffset += cw;
}
return el;
return el;
}
+114 -97
View File
@@ -5,28 +5,35 @@
* Consumers pass their real Terminal instance — we only use these properties.
*/
export interface XtermTerminal {
readonly element: HTMLElement | undefined;
readonly cols: number;
readonly rows: number;
readonly options: {
fontFamily?: string;
fontSize?: number;
fontWeight?: string | number;
theme?: {
background?: string;
foreground?: string;
cursor?: string;
};
readonly element: HTMLElement | undefined;
readonly cols: number;
readonly rows: number;
readonly options: {
fontFamily?: string;
fontSize?: number;
fontWeight?: string | number;
theme?: {
background?: string;
foreground?: string;
cursor?: string;
};
readonly buffer: {
readonly active: {
readonly viewportY: number;
readonly baseY: number;
getLine(y: number): {
translateToString(trimRight?: boolean): string;
} | undefined;
};
};
readonly buffer: {
readonly active: {
readonly viewportY: number;
readonly baseY: number;
getLine(y: number):
| {
translateToString(trimRight?: boolean): string;
}
| undefined;
};
};
/** Unicode addon (e.g. Unicode11Addon) for CJK wide character width */
readonly unicode?: {
getStringCellWidth(str: string): number;
activeVersion?: string;
};
}
/**
@@ -35,18 +42,18 @@ export interface XtermTerminal {
* The consumer calls `terminal.loadAddon(addon)` which invokes `activate()`.
*/
export interface XtermAddon {
activate(terminal: XtermTerminal): void;
dispose(): void;
activate(terminal: XtermTerminal): void;
dispose(): void;
}
/**
* Position of the prompt in the terminal viewport.
*/
export interface PromptPosition {
/** Viewport-relative row (0 = top of viewport) */
row: number;
/** Column of the prompt marker character */
col: number;
/** Viewport-relative row (0 = top of viewport) */
row: number;
/** Column of the prompt marker character */
col: number;
}
/**
@@ -60,104 +67,114 @@ export interface PromptPosition {
* - `custom`: Full escape hatch — provide your own finder function
*/
export type PromptFinder =
| { type: 'character'; char: string; offset?: number }
| { type: 'regex'; pattern: RegExp; offset?: number }
| { type: 'custom'; find: (terminal: XtermTerminal) => PromptPosition | null; offset?: number };
| { type: 'character'; char: string; offset?: number }
| { type: 'regex'; pattern: RegExp; offset?: number }
| { type: 'custom'; find: (terminal: XtermTerminal) => PromptPosition | null; offset?: number };
/**
* Configuration options for ZerolagInputAddon.
*/
export interface ZerolagInputOptions {
/**
* How to find the prompt in the terminal buffer.
*
* The `offset` controls how many characters after the prompt marker
* the user input begins (e.g., `"> "` = offset 2).
*
* @default { type: 'character', char: '>', offset: 2 }
*/
prompt?: PromptFinder;
/**
* How to find the prompt in the terminal buffer.
*
* The `offset` controls how many characters after the prompt marker
* the user input begins (e.g., `"> "` = offset 2).
*
* @default { type: 'character', char: '>', offset: 2 }
*/
prompt?: PromptFinder;
/**
* Z-index for the overlay element.
* @default 7
*/
zIndex?: number;
/**
* Z-index for the overlay element.
* @default 7
*/
zIndex?: number;
/**
* Background color for the overlay.
* Set to `'transparent'` to disable the opaque background.
* @default Read from terminal.options.theme.background
*/
backgroundColor?: string;
/**
* Background color for the overlay.
* Set to `'transparent'` to disable the opaque background.
* @default Read from terminal.options.theme.background
*/
backgroundColor?: string;
/**
* Foreground color for overlay text.
* @default Read from terminal.options.theme.foreground
*/
foregroundColor?: string;
/**
* Foreground color for overlay text.
* @default Read from terminal.options.theme.foreground
*/
foregroundColor?: string;
/**
* Whether to show a block cursor at the end of the overlay text.
* @default true
*/
showCursor?: boolean;
/**
* Whether to show a block cursor at the end of the overlay text.
* @default true
*/
showCursor?: boolean;
/**
* Cursor color (block cursor at end of text).
* @default Read from terminal.options.theme.cursor
*/
cursorColor?: string;
/**
* Cursor color (block cursor at end of text).
* @default Read from terminal.options.theme.cursor
*/
cursorColor?: string;
/**
* Scroll debounce time in ms for re-rendering when user scrolls
* back to the bottom of the terminal.
* @default 50
*/
scrollDebounceMs?: number;
/**
* Scroll debounce time in ms for re-rendering when user scrolls
* back to the bottom of the terminal.
* @default 50
*/
scrollDebounceMs?: number;
}
/**
* Read-only state snapshot of the overlay.
*/
export interface ZerolagInputState {
/** Characters typed but not yet acknowledged by the server */
pendingText: string;
/** Number of characters flushed to PTY but echo not yet received */
flushedLength: number;
/** Text content of the flushed portion */
flushedText: string;
/** Whether the overlay is currently visible */
visible: boolean;
/** Last detected prompt position, if any */
promptPosition: PromptPosition | null;
/** Characters typed but not yet acknowledged by the server */
pendingText: string;
/** Number of characters flushed to PTY but echo not yet received */
flushedLength: number;
/** Text content of the flushed portion */
flushedText: string;
/** Whether the overlay is currently visible */
visible: boolean;
/** Last detected prompt position, if any */
promptPosition: PromptPosition | null;
}
/** Cell dimensions in CSS pixels. */
export interface CellDimensions {
width: number;
height: number;
width: number;
height: number;
/** Vertical offset (px) from cell top to where characters render. */
charTop: number;
/** Height of the character rendering area (px). */
charHeight: number;
}
/** Parameters for the overlay renderer. */
export interface RenderParams {
lines: string[];
startCol: number;
totalCols: number;
cellW: number;
cellH: number;
promptRow: number;
font: FontStyle;
showCursor: boolean;
cursorColor: string;
lines: string[];
startCol: number;
totalCols: number;
cellW: number;
cellH: number;
/** Vertical offset (px) from cell top to character rendering area. */
charTop: number;
/** Height of the character rendering area (px). */
charHeight: number;
promptRow: number;
font: FontStyle;
showCursor: boolean;
cursorColor: string;
/** Terminal instance for CJK wide character width detection */
terminal?: XtermTerminal | null;
}
/** Cached font style properties for overlay rendering. */
export interface FontStyle {
fontFamily: string;
fontSize: string;
fontWeight: string;
color: string;
backgroundColor: string;
letterSpacing: string;
fontFamily: string;
fontSize: string;
fontWeight: string;
color: string;
backgroundColor: string;
letterSpacing: string;
}
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,127 @@
import { describe, it, expect, afterEach, beforeEach } from 'vitest';
import { getCellDimensions } from '../src/cell-dimensions.js';
import { createMockTerminal } from './helpers.js';
import type { XtermTerminal } from '../src/types.js';
let cleanups: (() => void)[] = [];
afterEach(() => {
for (const fn of cleanups) fn();
cleanups = [];
});
describe('getCellDimensions', () => {
describe('v5 private API (mock _core._renderService)', () => {
it('returns cell width and height from css.cell', () => {
const mock = createMockTerminal({ cellWidth: 8.4, cellHeight: 19 });
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
expect(dims!.width).toBe(8.4);
expect(dims!.height).toBe(19);
});
it('returns charTop from device.char.top divided by DPR', () => {
const mock = createMockTerminal({
cellWidth: 8, cellHeight: 19,
deviceCharTop: 2,
});
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
// DPR=1 in jsdom, so charTop = 2 / 1 = 2
expect(dims!.charTop).toBe(2);
});
it('returns charHeight from device.char.height divided by DPR', () => {
const mock = createMockTerminal({
cellWidth: 8, cellHeight: 19,
deviceCharHeight: 16,
});
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
// DPR=1, so charHeight = 16 / 1 = 16
expect(dims!.charHeight).toBe(16);
});
it('defaults charTop to 0 when device.char not present', () => {
// Default mock has deviceCharTop=0
const mock = createMockTerminal({ cellWidth: 8, cellHeight: 19 });
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims!.charTop).toBe(0);
});
it('defaults charHeight to cellH when device.char.height not set', () => {
// Default mock has deviceCharHeight=cellH
const mock = createMockTerminal({ cellWidth: 8, cellHeight: 19 });
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims!.charHeight).toBe(19);
});
});
describe('DPR simulation', () => {
const originalDPR = globalThis.devicePixelRatio;
beforeEach(() => {
// Set DPR=2 to test division
Object.defineProperty(globalThis, 'devicePixelRatio', {
value: 2,
writable: true,
configurable: true,
});
});
afterEach(() => {
Object.defineProperty(globalThis, 'devicePixelRatio', {
value: originalDPR,
writable: true,
configurable: true,
});
});
it('divides device.char.top by DPR', () => {
const mock = createMockTerminal({
cellWidth: 16, cellHeight: 38,
deviceCharTop: 4,
deviceCharHeight: 32,
});
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
// charTop = 4 / 2 = 2
expect(dims!.charTop).toBe(2);
// charHeight = 32 / 2 = 16
expect(dims!.charHeight).toBe(16);
});
});
describe('null cases', () => {
it('returns null for terminal without _core', () => {
const terminal = {
element: document.createElement('div'),
cols: 80,
rows: 24,
options: {},
buffer: { active: { viewportY: 0, baseY: 0, getLine: () => undefined } },
} as unknown as XtermTerminal;
const dims = getCellDimensions(terminal);
expect(dims).toBeNull();
});
it('returns null for terminal with no dimensions', () => {
const terminal = {
element: document.createElement('div'),
cols: 80,
rows: 24,
options: {},
buffer: { active: { viewportY: 0, baseY: 0, getLine: () => undefined } },
_core: { _renderService: {} },
} as unknown as XtermTerminal;
const dims = getCellDimensions(terminal);
expect(dims).toBeNull();
});
});
});
@@ -0,0 +1,843 @@
/**
* Comprehensive CJK wide character test plan for PR #30.
*
* Tests all 7 items from the test plan:
* 1. Chinese text input — no character overlap
* 2. CJK text renders with correct double-width spacing
* 3. Japanese (こんにちは) and Korean (안녕하세요) input
* 4. Cursor positions correctly after CJK characters
* 5. Long CJK input wraps at correct column boundary
* 6. Existing ASCII input is unaffected
* 7. Teammate terminal panels render CJK correctly (app.js embedded copy)
*/
import { describe, it, expect, afterEach } from 'vitest';
import { renderOverlay, charCellWidth, stringCellWidth } from '../src/overlay-renderer.js';
import type { RenderParams, FontStyle } from '../src/types.js';
import { createMockTerminal } from './helpers.js';
import { ZerolagInputAddon } from '../src/zerolag-input-addon.js';
// ─── Shared fixtures ─────────────────────────────────────────────────
const FONT: FontStyle = {
fontFamily: 'monospace',
fontSize: '14px',
fontWeight: 'normal',
color: '#eeeeee',
backgroundColor: '#0d0d0d',
letterSpacing: '',
};
function makeParams(overrides: Partial<RenderParams> = {}): RenderParams {
return {
lines: ['hello'],
startCol: 2,
totalCols: 80,
cellW: 10,
cellH: 17,
charTop: 2,
charHeight: 14,
promptRow: 0,
font: FONT,
showCursor: true,
cursorColor: '#e0e0e0',
...overrides,
};
}
/** Extract span data from a rendered line div */
function getSpans(lineDiv: HTMLDivElement) {
const spans: { text: string; left: number; width: number }[] = [];
for (let i = 0; i < lineDiv.children.length; i++) {
const span = lineDiv.children[i] as HTMLSpanElement;
spans.push({
text: span.textContent || '',
left: parseFloat(span.style.left),
width: parseFloat(span.style.width),
});
}
return spans;
}
// ─── Addon setup helpers ─────────────────────────────────────────────
let cleanups: (() => void)[] = [];
afterEach(() => {
for (const fn of cleanups) fn();
cleanups = [];
});
function tracked(lines: string[] = ['$ '], promptChar = '$') {
const mock = createMockTerminal({ buffer: { lines }, cols: 80 });
const addon = new ZerolagInputAddon({
prompt: { type: 'character', char: promptChar, offset: 2 },
});
mock.terminal.loadAddon(addon);
cleanups.push(() => {
addon.dispose();
mock.cleanup();
});
return { addon, mock };
}
// ═══════════════════════════════════════════════════════════════════════
// TEST 1: Chinese text input — no character overlap
// ═══════════════════════════════════════════════════════════════════════
describe('Test 1: Chinese text — no character overlap', () => {
it('consecutive Chinese characters have contiguous non-overlapping spans', () => {
const container = document.createElement('div');
const text = '你好世界';
renderOverlay(container, makeParams({ lines: [text], cellW: 10 }));
const lineDiv = container.children[0] as HTMLDivElement;
const spans = getSpans(lineDiv);
expect(spans.length).toBe(4);
// Each span: left = previous span's (left + width), width = 20px (2 cells)
for (let i = 0; i < spans.length; i++) {
expect(spans[i].width).toBe(20); // 2 cells * 10px
if (i > 0) {
const expectedLeft = spans[i - 1].left + spans[i - 1].width;
expect(spans[i].left).toBe(expectedLeft);
}
}
});
it('Chinese sentence (simulated pinyin output) renders without gaps or overlaps', () => {
const container = document.createElement('div');
const text = '我是一个测试';
renderOverlay(container, makeParams({ lines: [text], cellW: 8 }));
const lineDiv = container.children[0] as HTMLDivElement;
const spans = getSpans(lineDiv);
expect(spans.length).toBe(6);
// Verify contiguous positioning
let expectedLeft = 0;
for (const span of spans) {
expect(span.left).toBe(expectedLeft);
expect(span.width).toBe(16); // 2 * 8px
expectedLeft += span.width;
}
// Total visual width should be 6 chars * 2 cells * 8px = 96px
expect(expectedLeft).toBe(96);
});
it('mixed Chinese + ASCII has no gaps between spans', () => {
const container = document.createElement('div');
const text = 'hello你好world';
renderOverlay(container, makeParams({ lines: [text], cellW: 10 }));
const lineDiv = container.children[0] as HTMLDivElement;
const spans = getSpans(lineDiv);
// h(10) e(10) l(10) l(10) o(10) 你(20) 好(20) w(10) o(10) r(10) l(10) d(10)
expect(spans.length).toBe(12);
// Verify contiguity: no gaps between any adjacent spans
for (let i = 1; i < spans.length; i++) {
const prevEnd = spans[i - 1].left + spans[i - 1].width;
expect(spans[i].left).toBe(prevEnd);
}
});
it('addChar with Chinese characters accumulates correctly in addon', () => {
const { addon } = tracked();
for (const ch of '你好世界') {
addon.addChar(ch);
}
expect(addon.pendingText).toBe('你好世界');
expect(addon.hasPending).toBe(true);
});
});
// ═══════════════════════════════════════════════════════════════════════
// TEST 2: CJK text renders with correct double-width spacing
// ═══════════════════════════════════════════════════════════════════════
describe('Test 2: CJK double-width spacing', () => {
it('charCellWidth returns 2 for CJK Unified Ideographs (0x4E00-0x9FFF)', () => {
// Common Chinese characters
const chars = '中文测试你好世界天地人';
for (const ch of chars) {
expect(charCellWidth(null, ch)).toBe(2);
}
});
it('charCellWidth returns 2 for CJK Extension A (0x3400-0x4DBF)', () => {
expect(charCellWidth(null, '\u3400')).toBe(2); // First Extension A char
expect(charCellWidth(null, '\u4DB5')).toBe(2); // One of the last Extension A chars
});
it('charCellWidth returns 2 for CJK Radicals (0x2E80-0x2EFF)', () => {
expect(charCellWidth(null, '\u2E80')).toBe(2); // CJK Radical Repeat
});
it('charCellWidth returns 2 for fullwidth ASCII forms (0xFF01-0xFF5E)', () => {
expect(charCellWidth(null, '\uFF01')).toBe(2); // !
expect(charCellWidth(null, '\uFF21')).toBe(2); // A
expect(charCellWidth(null, '\uFF41')).toBe(2); // a
expect(charCellWidth(null, '\uFF10')).toBe(2); // 0
});
it('charCellWidth returns 2 for CJK Compatibility Ideographs (0xF900-0xFAFF)', () => {
expect(charCellWidth(null, '\uF900')).toBe(2);
});
it('stringCellWidth calculates correct visual width for CJK strings', () => {
expect(stringCellWidth(null, '你好')).toBe(4); // 2 + 2
expect(stringCellWidth(null, '你好世界')).toBe(8); // 4 * 2
expect(stringCellWidth(null, 'hi你好')).toBe(6); // 1 + 1 + 2 + 2
expect(stringCellWidth(null, '你a好b')).toBe(6); // 2 + 1 + 2 + 1
expect(stringCellWidth(null, 'abc你好def')).toBe(10); // 3 + 4 + 3
});
it('CJK span widths are exactly 2 * cellW pixels', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['你'], cellW: 8.4 }));
const lineDiv = container.children[0] as HTMLDivElement;
const span = lineDiv.children[0] as HTMLSpanElement;
expect(span.style.width).toBe('16.8px'); // 2 * 8.4
});
it('ASCII span widths remain 1 * cellW pixels', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['a'], cellW: 8.4 }));
const lineDiv = container.children[0] as HTMLDivElement;
const span = lineDiv.children[0] as HTMLSpanElement;
expect(span.style.width).toBe('8.4px'); // 1 * 8.4
});
it('terminal unicode addon is preferred over fallback when available', () => {
const mockTerminal = {
unicode: {
getStringCellWidth: (s: string) => {
// Custom width: treat 'W' as wide
return s === 'W' ? 2 : 1;
},
},
} as any;
expect(charCellWidth(mockTerminal, 'W')).toBe(2);
expect(charCellWidth(mockTerminal, 'a')).toBe(1);
// Fallback ignores terminal addon result for actual CJK
expect(charCellWidth(null, '你')).toBe(2);
});
});
// ═══════════════════════════════════════════════════════════════════════
// TEST 3: Japanese (こんにちは) and Korean (안녕하세요) input
// ═══════════════════════════════════════════════════════════════════════
describe('Test 3: Japanese and Korean input', () => {
describe('Japanese', () => {
it('charCellWidth returns 2 for Hiragana (0x3040-0x309F)', () => {
const hiragana = 'あいうえおかきくけこさしすせそ';
for (const ch of hiragana) {
expect(charCellWidth(null, ch)).toBe(2);
}
});
it('charCellWidth returns 2 for Katakana (0x30A0-0x30FF)', () => {
const katakana = 'アイウエオカキクケコサシスセソ';
for (const ch of katakana) {
expect(charCellWidth(null, ch)).toBe(2);
}
});
it('Japanese greeting renders with correct span positions', () => {
const container = document.createElement('div');
const text = 'こんにちは';
renderOverlay(container, makeParams({ lines: [text], cellW: 10 }));
const lineDiv = container.children[0] as HTMLDivElement;
const spans = getSpans(lineDiv);
expect(spans.length).toBe(5);
// こ(0,20) ん(20,20) に(40,20) ち(60,20) は(80,20)
expect(spans[0]).toEqual({ text: 'こ', left: 0, width: 20 });
expect(spans[1]).toEqual({ text: 'ん', left: 20, width: 20 });
expect(spans[2]).toEqual({ text: 'に', left: 40, width: 20 });
expect(spans[3]).toEqual({ text: 'ち', left: 60, width: 20 });
expect(spans[4]).toEqual({ text: 'は', left: 80, width: 20 });
});
it('mixed Japanese + ASCII positions correctly', () => {
const container = document.createElement('div');
const text = 'hello こんにちは';
renderOverlay(container, makeParams({ lines: [text], cellW: 10 }));
const lineDiv = container.children[0] as HTMLDivElement;
const spans = getSpans(lineDiv);
// h(0) e(10) l(20) l(30) o(40) space(50) こ(60) ん(80) に(100) ち(120) は(140)
expect(spans.length).toBe(11);
expect(spans[5]).toEqual({ text: ' ', left: 50, width: 10 }); // space
expect(spans[6]).toEqual({ text: 'こ', left: 60, width: 20 }); // first CJK after ASCII
});
it('addon handles Japanese input via addChar', () => {
const { addon } = tracked();
for (const ch of 'こんにちは') {
addon.addChar(ch);
}
expect(addon.pendingText).toBe('こんにちは');
});
it('addon handles Japanese input via appendText (paste)', () => {
const { addon } = tracked();
addon.appendText('こんにちは世界');
expect(addon.pendingText).toBe('こんにちは世界');
expect(addon.hasPending).toBe(true);
});
it('stringCellWidth correct for Japanese greeting', () => {
expect(stringCellWidth(null, 'こんにちは')).toBe(10); // 5 * 2
});
});
describe('Korean', () => {
it('charCellWidth returns 2 for Hangul Syllables (0xAC00-0xD7A3)', () => {
const hangul = '가나다라마바사아자차카타파하';
for (const ch of hangul) {
expect(charCellWidth(null, ch)).toBe(2);
}
});
it('charCellWidth returns 2 for Hangul Jamo (0x1100-0x115F)', () => {
expect(charCellWidth(null, '\u1100')).toBe(2); // ᄀ
expect(charCellWidth(null, '\u1112')).toBe(2); // ᄒ
});
it('Korean greeting renders with correct span positions', () => {
const container = document.createElement('div');
const text = '안녕하세요';
renderOverlay(container, makeParams({ lines: [text], cellW: 10 }));
const lineDiv = container.children[0] as HTMLDivElement;
const spans = getSpans(lineDiv);
expect(spans.length).toBe(5);
expect(spans[0]).toEqual({ text: '안', left: 0, width: 20 });
expect(spans[1]).toEqual({ text: '녕', left: 20, width: 20 });
expect(spans[2]).toEqual({ text: '하', left: 40, width: 20 });
expect(spans[3]).toEqual({ text: '세', left: 60, width: 20 });
expect(spans[4]).toEqual({ text: '요', left: 80, width: 20 });
});
it('addon handles Korean input', () => {
const { addon } = tracked();
addon.appendText('안녕하세요');
expect(addon.pendingText).toBe('안녕하세요');
expect(addon.hasPending).toBe(true);
});
it('removeChar removes Korean characters one at a time', () => {
const { addon } = tracked();
addon.appendText('안녕');
addon.removeChar();
expect(addon.pendingText).toBe('안');
addon.removeChar();
expect(addon.pendingText).toBe('');
});
it('stringCellWidth correct for Korean greeting', () => {
expect(stringCellWidth(null, '안녕하세요')).toBe(10); // 5 * 2
});
});
describe('mixed CJK scripts', () => {
it('Chinese + Japanese + Korean in one string', () => {
const text = '你好こんにちは안녕';
expect(stringCellWidth(null, text)).toBe(18); // 9 chars * 2 each
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: [text], cellW: 10 }));
const lineDiv = container.children[0] as HTMLDivElement;
const spans = getSpans(lineDiv);
// All 9 CJK chars: contiguous double-width spans
expect(spans.length).toBe(9);
let expectedLeft = 0;
for (const span of spans) {
expect(span.left).toBe(expectedLeft);
expect(span.width).toBe(20);
expectedLeft += 20;
}
});
});
});
// ═══════════════════════════════════════════════════════════════════════
// TEST 4: Cursor positions correctly after CJK characters
// ═══════════════════════════════════════════════════════════════════════
describe('Test 4: Cursor positioning after CJK', () => {
it('cursor after Chinese text on first line', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['你好世界'],
startCol: 2,
cellW: 10,
showCursor: true,
})
);
// cursorCol = startCol(2) + stringCellWidth('你好世界')(8) = 10
const cursor = container.children[container.children.length - 1] as HTMLSpanElement;
expect(cursor.style.left).toBe('100px'); // 10 * 10px
});
it('cursor after Japanese text on first line', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['こんにちは'],
startCol: 2,
cellW: 10,
showCursor: true,
})
);
// cursorCol = 2 + 10 = 12
const cursor = container.children[container.children.length - 1] as HTMLSpanElement;
expect(cursor.style.left).toBe('120px');
});
it('cursor after Korean text on first line', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['안녕하세요'],
startCol: 2,
cellW: 10,
showCursor: true,
})
);
// cursorCol = 2 + 10 = 12
const cursor = container.children[container.children.length - 1] as HTMLSpanElement;
expect(cursor.style.left).toBe('120px');
});
it('cursor after mixed ASCII + CJK', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['hi你好'],
startCol: 3,
cellW: 10,
showCursor: true,
})
);
// cursorCol = 3 + stringCellWidth('hi你好')(6) = 9
const cursor = container.children[container.children.length - 1] as HTMLSpanElement;
expect(cursor.style.left).toBe('90px');
});
it('cursor on wrapped line after CJK text', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['你好世界', '再见'],
startCol: 2,
cellW: 10,
cellH: 20,
showCursor: true,
})
);
// Last line is '再见', starts at col 0 (wrapped line), width = 4
const cursor = container.children[container.children.length - 1] as HTMLSpanElement;
expect(cursor.style.left).toBe('40px'); // 0 + 4 = 4, * 10 = 40
expect(cursor.style.top).toBe('20px'); // row 1 * cellH
});
it('cursor hidden when it would exceed totalCols', () => {
const container = document.createElement('div');
// 4 CJK chars = 8 visual cols, startCol=73, cursorCol = 73 + 8 = 81 > 80
renderOverlay(
container,
makeParams({
lines: ['你好世界'],
startCol: 73,
totalCols: 80,
cellW: 10,
showCursor: true,
})
);
// Should be 1 line div, no cursor span (cursor at col 81 >= totalCols 80)
// Actually cursor check is cursorCol < totalCols, so at 81 it's hidden
const children = container.children;
// If cursor is rendered, last child would be a cursor span
// With cursorCol 81 >= 80, cursor should NOT be rendered
expect(children.length).toBe(1); // only line div
});
});
// ═══════════════════════════════════════════════════════════════════════
// TEST 5: Long CJK input wraps at correct column boundary
// ═══════════════════════════════════════════════════════════════════════
describe('Test 5: CJK line wrapping at column boundaries', () => {
it('CJK chars wrap when they would overflow first line', () => {
const container = document.createElement('div');
// totalCols=10, startCol=2 → firstLineCols=8
// Each CJK char = 2 cols → first line fits 4 chars (8 cols)
// 5th char wraps to second line
const text = '你好世界啊'; // 5 chars = 10 visual cols
renderOverlay(
container,
makeParams({
lines: ['你好世界', '啊'],
startCol: 2,
totalCols: 10,
cellW: 10,
cellH: 20,
})
);
expect(container.children.length).toBe(3); // 2 line divs + cursor
const line1 = container.children[0] as HTMLDivElement;
expect(line1.children.length).toBe(4); // 你好世界
expect(line1.style.left).toBe('20px'); // startCol * cellW
const line2 = container.children[1] as HTMLDivElement;
expect(line2.children.length).toBe(1); // 啊
expect(line2.style.left).toBe('0px'); // wrapped line starts at col 0
expect(line2.style.top).toBe('20px'); // second row
});
it('CJK char that would partially overflow stays on next line', () => {
// totalCols=9, startCol=2 → firstLineCols=7
// CJK chars are 2-wide. 3 chars = 6 cols (fits). 4th char = 8 cols > 7. Wraps.
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['你好世', '界'],
startCol: 2,
totalCols: 9,
cellW: 10,
})
);
const line1 = container.children[0] as HTMLDivElement;
expect(line1.children.length).toBe(3); // 3 CJK chars fit (6 cols <= 7)
const line2 = container.children[1] as HTMLDivElement;
expect(line2.children.length).toBe(1); // 界 wraps
});
it('addon _render splits CJK text into visual lines correctly', () => {
// Narrow terminal: 10 cols, startCol=2 → firstLineCols=8 → 4 CJK chars
const mock = createMockTerminal({
buffer: { lines: ['$ '] },
cols: 10,
});
const addon = new ZerolagInputAddon({
prompt: { type: 'character', char: '$', offset: 2 },
});
mock.terminal.loadAddon(addon);
cleanups.push(() => {
addon.dispose();
mock.cleanup();
});
// Type 6 CJK chars (12 visual cols)
for (const ch of '你好世界再见') {
addon.addChar(ch);
}
expect(addon.pendingText).toBe('你好世界再见');
expect(addon.hasPending).toBe(true);
});
it('mixed ASCII + CJK wraps correctly', () => {
const container = document.createElement('div');
// totalCols=10, startCol=2 → firstLineCols=8
// 'ab' = 2 cols, '你好' = 4 cols, 'cd' = 2 cols → total 8 cols (fits line 1)
// '世' = 2 cols → wraps to line 2
renderOverlay(
container,
makeParams({
lines: ['ab你好cd', '世'],
startCol: 2,
totalCols: 10,
cellW: 10,
})
);
const line1 = container.children[0] as HTMLDivElement;
expect(line1.children.length).toBe(6); // a,b,你,好,c,d
const line2 = container.children[1] as HTMLDivElement;
expect(line2.children.length).toBe(1); // 世
});
it('odd column count: CJK char does not split across boundary', () => {
// totalCols=11, startCol=2 → firstLineCols=9
// 4 CJK chars = 8 cols (fits). 5th CJK = 10 cols > 9 (wraps).
// 1 empty column remains on first line (CJK can't fit in 1 col).
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['你好世界', '啊'],
startCol: 2,
totalCols: 11,
cellW: 10,
})
);
const line1 = container.children[0] as HTMLDivElement;
expect(line1.children.length).toBe(4); // 4 chars = 8 cols, 5th would need 10 > 9
const line2 = container.children[1] as HTMLDivElement;
expect(line2.children.length).toBe(1);
});
it('CJK fills entire wrapped line', () => {
const container = document.createElement('div');
// totalCols=6, startCol=0 → firstLineCols=6
// 3 CJK chars = 6 cols (fills line 1), next 3 fill line 2
renderOverlay(
container,
makeParams({
lines: ['你好世', '界再见'],
startCol: 0,
totalCols: 6,
cellW: 10,
})
);
const line1 = container.children[0] as HTMLDivElement;
expect(line1.children.length).toBe(3);
const line2 = container.children[1] as HTMLDivElement;
expect(line2.children.length).toBe(3);
});
});
// ═══════════════════════════════════════════════════════════════════════
// TEST 6: Existing ASCII input is unaffected
// ═══════════════════════════════════════════════════════════════════════
describe('Test 6: ASCII input — no regression', () => {
it('ASCII characters still get 1-cell-wide spans', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['abcdef'], cellW: 10 }));
const lineDiv = container.children[0] as HTMLDivElement;
const spans = getSpans(lineDiv);
for (const span of spans) {
expect(span.width).toBe(10); // 1 * cellW
}
});
it('ASCII span positions are sequential at 1-cell intervals', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['xyz'], cellW: 8.4 }));
const lineDiv = container.children[0] as HTMLDivElement;
const spans = getSpans(lineDiv);
expect(spans[0].left).toBe(0);
expect(spans[1].left).toBeCloseTo(8.4, 5);
expect(spans[2].left).toBeCloseTo(16.8, 5);
});
it('ASCII cursor positions at correct column', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['hello'],
startCol: 3,
cellW: 10,
showCursor: true,
})
);
// cursor at startCol(3) + 5 = 8
const cursor = container.children[container.children.length - 1] as HTMLSpanElement;
expect(cursor.style.left).toBe('80px');
});
it('charCellWidth returns 1 for all printable ASCII', () => {
for (let code = 32; code < 127; code++) {
const ch = String.fromCharCode(code);
expect(charCellWidth(null, ch)).toBe(1);
}
});
it('stringCellWidth equals length for pure ASCII', () => {
expect(stringCellWidth(null, 'hello world')).toBe(11);
expect(stringCellWidth(null, 'test123!@#')).toBe(10);
expect(stringCellWidth(null, '')).toBe(0);
});
it('ASCII multi-line wrapping is unaffected', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['abcde', 'fgh'],
startCol: 5,
totalCols: 10,
cellW: 10,
cellH: 20,
})
);
expect(container.children.length).toBe(3); // 2 lines + cursor
const line1 = container.children[0] as HTMLDivElement;
const line2 = container.children[1] as HTMLDivElement;
expect(line1.children.length).toBe(5);
expect(line2.children.length).toBe(3);
expect(line1.style.left).toBe('50px'); // startCol * cellW
expect(line2.style.left).toBe('0px');
});
it('addon addChar/removeChar/clear work for ASCII', () => {
const { addon } = tracked();
addon.addChar('h');
addon.addChar('e');
addon.addChar('l');
addon.addChar('l');
addon.addChar('o');
expect(addon.pendingText).toBe('hello');
addon.removeChar();
expect(addon.pendingText).toBe('hell');
addon.clear();
expect(addon.pendingText).toBe('');
expect(addon.hasPending).toBe(false);
});
});
// ═══════════════════════════════════════════════════════════════════════
// TEST 7: Teammate terminal panels render CJK correctly
//
// Teammate panels share the global xterm-zerolag-input.js vendor bundle
// (loaded via <script> in index.html). This is the IIFE build of the
// same package source tested in tests 1-6. The build pipeline is:
// src/overlay-renderer.ts → tsup → dist/index.global.js → vendor copy
//
// Since all terminals (main + teammate) use the same LocalEchoOverlay
// class from the global scope, CJK correctness is guaranteed by:
// (a) The package source handles CJK correctly (tests 1-6 above)
// (b) The IIFE build bundles the exact same charCellWidth / makeLine code
//
// These tests verify the exported module includes CJK-aware functions
// and that teammate-style terminal instances work identically.
// ═══════════════════════════════════════════════════════════════════════
describe('Test 7: Teammate terminal panels render CJK correctly', () => {
it('package exports charCellWidth and stringCellWidth', () => {
// These are the CJK-aware functions that the IIFE build exposes
expect(typeof charCellWidth).toBe('function');
expect(typeof stringCellWidth).toBe('function');
});
it('ZerolagInputAddon (used by LocalEchoOverlay) handles CJK in teammate terminals', () => {
// Simulate a teammate terminal panel: separate terminal instance,
// same addon class, different prompt character
const mock = createMockTerminal({
buffer: { lines: ['❯ '] },
cols: 40,
});
const addon = new ZerolagInputAddon({
prompt: { type: 'character', char: '❯', offset: 2 },
});
mock.terminal.loadAddon(addon);
cleanups.push(() => {
addon.dispose();
mock.cleanup();
});
// Type CJK in teammate terminal
for (const ch of '你好世界') {
addon.addChar(ch);
}
expect(addon.pendingText).toBe('你好世界');
expect(addon.hasPending).toBe(true);
});
it('teammate terminal with narrow width wraps CJK correctly', () => {
// Teammate panels are often narrower (sidebar, split view)
const mock = createMockTerminal({
buffer: { lines: ['❯ '] },
cols: 12, // narrow panel
});
const addon = new ZerolagInputAddon({
prompt: { type: 'character', char: '❯', offset: 2 },
});
mock.terminal.loadAddon(addon);
cleanups.push(() => {
addon.dispose();
mock.cleanup();
});
// Type 6 CJK chars (12 visual cols) with startCol=2 → only 10 available
// First line: 5 CJK = 10 cols. 6th wraps.
for (const ch of '你好世界再见') {
addon.addChar(ch);
}
expect(addon.pendingText).toBe('你好世界再见');
});
it('mixed CJK scripts render identically in teammate and main terminals', () => {
// Verify that the same input produces the same overlay output
// regardless of which terminal instance it's on
const text = '你好こんにちは안녕';
// "Main" terminal
const container1 = document.createElement('div');
renderOverlay(container1, makeParams({ lines: [text], cellW: 10 }));
const line1 = container1.children[0] as HTMLDivElement;
const spans1 = getSpans(line1);
// "Teammate" terminal (same params — same bundle)
const container2 = document.createElement('div');
renderOverlay(container2, makeParams({ lines: [text], cellW: 10 }));
const line2 = container2.children[0] as HTMLDivElement;
const spans2 = getSpans(line2);
// Identical rendering
expect(spans1.length).toBe(spans2.length);
for (let i = 0; i < spans1.length; i++) {
expect(spans1[i]).toEqual(spans2[i]);
}
});
it('CJK rendering correct in overlay with teammate-typical prompt (❯)', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['こんにちは世界'],
startCol: 2, // after ❯ prompt
cellW: 10,
showCursor: true,
})
);
const lineDiv = container.children[0] as HTMLDivElement;
const spans = getSpans(lineDiv);
expect(spans.length).toBe(7);
// All CJK, all double-width, contiguous
let expectedLeft = 0;
for (const span of spans) {
expect(span.left).toBe(expectedLeft);
expect(span.width).toBe(20);
expectedLeft += 20;
}
// Cursor position: startCol(2) + 14 visual cols = 16
const cursor = container.children[container.children.length - 1] as HTMLSpanElement;
expect(cursor.style.left).toBe('160px');
});
});
@@ -31,6 +31,10 @@ interface MockTerminalOptions {
};
cellWidth?: number;
cellHeight?: number;
/** Device-pixel char top offset (for charTop calculation). Default: 0 */
deviceCharTop?: number;
/** Device-pixel char height (for charHeight calculation). Default: cellHeight * dpr */
deviceCharHeight?: number;
}
export function createMockTerminal(opts: MockTerminalOptions = {}) {
@@ -95,6 +99,12 @@ export function createMockTerminal(opts: MockTerminalOptions = {}) {
css: {
cell: { width: cellW, height: cellH },
},
device: {
char: {
top: opts.deviceCharTop ?? 0,
height: opts.deviceCharHeight ?? cellH,
},
},
},
},
},
@@ -1,169 +1,420 @@
import { describe, it, expect } from 'vitest';
import { renderOverlay } from '../src/overlay-renderer.js';
import { renderOverlay, charCellWidth, stringCellWidth } from '../src/overlay-renderer.js';
import type { RenderParams, FontStyle } from '../src/types.js';
const FONT: FontStyle = {
fontFamily: 'monospace',
fontSize: '14px',
fontWeight: 'normal',
color: '#eeeeee',
backgroundColor: '#0d0d0d',
letterSpacing: '',
fontFamily: 'monospace',
fontSize: '14px',
fontWeight: 'normal',
color: '#eeeeee',
backgroundColor: '#0d0d0d',
letterSpacing: '',
};
function makeParams(overrides: Partial<RenderParams> = {}): RenderParams {
return {
lines: ['hello'],
startCol: 2,
totalCols: 80,
cellW: 8.4,
cellH: 17,
promptRow: 10,
font: FONT,
showCursor: true,
cursorColor: '#e0e0e0',
...overrides,
};
return {
lines: ['hello'],
startCol: 2,
totalCols: 80,
cellW: 8.4,
cellH: 17,
charTop: 2,
charHeight: 14,
promptRow: 10,
font: FONT,
showCursor: true,
cursorColor: '#e0e0e0',
...overrides,
};
}
describe('renderOverlay', () => {
it('positions container at prompt row', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ promptRow: 5 }));
expect(container.style.top).toBe((5 * 17) + 'px');
expect(container.style.left).toBe('0px');
});
it('positions container at prompt row', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ promptRow: 5 }));
expect(container.style.top).toBe(5 * 17 + 'px');
expect(container.style.left).toBe('0px');
});
it('creates per-character spans in a line div', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['abc'] }));
it('creates per-character spans in a line div', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['abc'] }));
// Line div + cursor span
expect(container.children.length).toBe(2);
// Line div + cursor span
expect(container.children.length).toBe(2);
const lineDiv = container.children[0] as HTMLDivElement;
expect(lineDiv.children.length).toBe(3); // a, b, c
const lineDiv = container.children[0] as HTMLDivElement;
expect(lineDiv.children.length).toBe(3); // a, b, c
const spanA = lineDiv.children[0] as HTMLSpanElement;
expect(spanA.textContent).toBe('a');
expect(spanA.style.left).toBe('0px');
const spanA = lineDiv.children[0] as HTMLSpanElement;
expect(spanA.textContent).toBe('a');
expect(spanA.style.left).toBe('0px');
const spanB = lineDiv.children[1] as HTMLSpanElement;
expect(spanB.textContent).toBe('b');
expect(spanB.style.left).toBe('8.4px');
const spanB = lineDiv.children[1] as HTMLSpanElement;
expect(spanB.textContent).toBe('b');
expect(spanB.style.left).toBe('8.4px');
const spanC = lineDiv.children[2] as HTMLSpanElement;
expect(spanC.textContent).toBe('c');
expect(spanC.style.left).toBe('16.8px');
});
const spanC = lineDiv.children[2] as HTMLSpanElement;
expect(spanC.textContent).toBe('c');
expect(spanC.style.left).toBe('16.8px');
});
it('sets span width to cellW', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['x'], cellW: 9.5 }));
const lineDiv = container.children[0] as HTMLDivElement;
const span = lineDiv.children[0] as HTMLSpanElement;
expect(span.style.width).toBe('9.5px');
});
it('sets span width to cellW', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['x'], cellW: 9.5 }));
const lineDiv = container.children[0] as HTMLDivElement;
const span = lineDiv.children[0] as HTMLSpanElement;
expect(span.style.width).toBe('9.5px');
});
it('applies font styles to spans', () => {
const font: FontStyle = {
fontFamily: 'Fira Code',
fontSize: '16px',
fontWeight: 'bold',
color: '#ff0000',
backgroundColor: '#000000',
letterSpacing: '0.5px',
};
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['A'], font }));
it('applies font styles to spans', () => {
const font: FontStyle = {
fontFamily: 'Fira Code',
fontSize: '16px',
fontWeight: 'bold',
color: '#ff0000',
backgroundColor: '#000000',
letterSpacing: '0.5px',
};
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['A'], font }));
const lineDiv = container.children[0] as HTMLDivElement;
// jsdom normalizes hex to rgb()
expect(lineDiv.style.backgroundColor).toBe('rgb(0, 0, 0)');
const lineDiv = container.children[0] as HTMLDivElement;
// jsdom normalizes hex to rgb()
expect(lineDiv.style.backgroundColor).toBe('rgb(0, 0, 0)');
const span = lineDiv.children[0] as HTMLSpanElement;
expect(span.style.fontFamily).toBe('Fira Code');
expect(span.style.fontSize).toBe('16px');
expect(span.style.fontWeight).toBe('bold');
expect(span.style.color).toBe('rgb(255, 0, 0)');
expect(span.style.letterSpacing).toBe('0.5px');
});
const span = lineDiv.children[0] as HTMLSpanElement;
expect(span.style.fontFamily).toBe('Fira Code');
expect(span.style.fontSize).toBe('16px');
expect(span.style.fontWeight).toBe('bold');
expect(span.style.color).toBe('rgb(255, 0, 0)');
expect(span.style.letterSpacing).toBe('0.5px');
});
it('offsets first line by startCol', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['hi'], startCol: 5, cellW: 10 }));
const lineDiv = container.children[0] as HTMLDivElement;
// First line left = startCol * cellW
expect(lineDiv.style.left).toBe('50px');
});
it('offsets first line by startCol', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['hi'], startCol: 5, cellW: 10 }));
const lineDiv = container.children[0] as HTMLDivElement;
// First line left = startCol * cellW
expect(lineDiv.style.left).toBe('50px');
});
it('renders cursor at end of text', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({
lines: ['ab'],
startCol: 3,
cellW: 10,
cellH: 20,
showCursor: true,
cursorColor: '#ff00ff',
}));
it('renders cursor at end of text', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['ab'],
startCol: 3,
cellW: 10,
cellH: 20,
showCursor: true,
cursorColor: '#ff00ff',
})
);
// Last child is cursor (after line div)
const cursor = container.children[container.children.length - 1] as HTMLSpanElement;
// cursorCol = startCol(3) + text.length(2) = 5
expect(cursor.style.left).toBe('50px');
expect(cursor.style.width).toBe('10px');
expect(cursor.style.height).toBe('20px');
// jsdom normalizes hex to rgb()
expect(cursor.style.backgroundColor).toBe('rgb(255, 0, 255)');
});
// Last child is cursor (after line div)
const cursor = container.children[container.children.length - 1] as HTMLSpanElement;
// cursorCol = startCol(3) + text.length(2) = 5
expect(cursor.style.left).toBe('50px');
expect(cursor.style.width).toBe('10px');
expect(cursor.style.height).toBe('20px');
// jsdom normalizes hex to rgb()
expect(cursor.style.backgroundColor).toBe('rgb(255, 0, 255)');
});
it('does not render cursor when showCursor is false', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['ab'], showCursor: false }));
// Only line div, no cursor
expect(container.children.length).toBe(1);
});
it('does not render cursor when showCursor is false', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['ab'], showCursor: false }));
// Only line div, no cursor
expect(container.children.length).toBe(1);
});
it('renders multi-line text', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({
lines: ['first', 'second'],
startCol: 5,
cellW: 10,
cellH: 20,
}));
it('renders multi-line text', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['first', 'second'],
startCol: 5,
cellW: 10,
cellH: 20,
})
);
// 2 line divs + cursor
expect(container.children.length).toBe(3);
// 2 line divs + cursor
expect(container.children.length).toBe(3);
const line1 = container.children[0] as HTMLDivElement;
expect(line1.style.left).toBe('50px'); // startCol * cellW
expect(line1.style.top).toBe('0px');
expect(line1.children.length).toBe(5); // 'first'
const line1 = container.children[0] as HTMLDivElement;
expect(line1.style.left).toBe('50px'); // startCol * cellW
expect(line1.style.top).toBe('0px');
expect(line1.children.length).toBe(5); // 'first'
const line2 = container.children[1] as HTMLDivElement;
expect(line2.style.left).toBe('0px'); // wrapped lines start at col 0
expect(line2.style.top).toBe('20px'); // second row
expect(line2.children.length).toBe(6); // 'second'
});
const line2 = container.children[1] as HTMLDivElement;
expect(line2.style.left).toBe('0px'); // wrapped lines start at col 0
expect(line2.style.top).toBe('20px'); // second row
expect(line2.children.length).toBe(6); // 'second'
});
it('clears previous content on re-render', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['abc'] }));
expect(container.children.length).toBe(2); // line + cursor
it('clears previous content on re-render', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['abc'] }));
expect(container.children.length).toBe(2); // line + cursor
renderOverlay(container, makeParams({ lines: ['xy'] }));
expect(container.children.length).toBe(2); // line + cursor (rebuilt)
renderOverlay(container, makeParams({ lines: ['xy'] }));
expect(container.children.length).toBe(2); // line + cursor (rebuilt)
const lineDiv = container.children[0] as HTMLDivElement;
expect(lineDiv.children.length).toBe(2); // x, y
});
const lineDiv = container.children[0] as HTMLDivElement;
expect(lineDiv.children.length).toBe(2); // x, y
});
it('shows container (display not none)', () => {
const container = document.createElement('div');
container.style.display = 'none';
renderOverlay(container, makeParams());
expect(container.style.display).toBe('');
});
it('shows container (display not none)', () => {
const container = document.createElement('div');
container.style.display = 'none';
renderOverlay(container, makeParams());
expect(container.style.display).toBe('');
});
// ─── Anti-flicker / compositing seam tests ────────────────────
it('line div height extends 1px past cellH to cover compositing seam', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['abc'], cellH: 19 }));
const lineDiv = container.children[0] as HTMLDivElement;
// cellH + 1 = 20px — the extra 1px covers the compositing seam
expect(lineDiv.style.height).toBe('20px');
});
it('line div height is cellH+1 for various cell heights', () => {
for (const cellH of [15, 17, 19, 22]) {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['x'], cellH }));
const lineDiv = container.children[0] as HTMLDivElement;
expect(lineDiv.style.height).toBe(cellH + 1 + 'px');
}
});
it('multi-line overlay has cellH+1 height on each line div', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['first', 'second'],
cellH: 19,
})
);
const line1 = container.children[0] as HTMLDivElement;
const line2 = container.children[1] as HTMLDivElement;
expect(line1.style.height).toBe('20px');
expect(line2.style.height).toBe('20px');
});
// ─── Span vertical centering tests ────────────────────────────
it('span uses full cellH for height and lineHeight (CSS centering)', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['a'], cellH: 19 }));
const lineDiv = container.children[0] as HTMLDivElement;
const span = lineDiv.children[0] as HTMLSpanElement;
expect(span.style.height).toBe('19px');
expect(span.style.lineHeight).toBe('19px');
});
it('span top is 0px (no vertical offset / no transform)', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['a'], cellH: 19 }));
const lineDiv = container.children[0] as HTMLDivElement;
const span = lineDiv.children[0] as HTMLSpanElement;
expect(span.style.top).toBe('0px');
// No translateY transform — sub-pixel overhang causes artifacts
expect(span.style.transform).toBe('');
});
// ─── Font rendering tests ─────────────────────────────────────
it('span disables ligatures via font-feature-settings', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['fi'] }));
const lineDiv = container.children[0] as HTMLDivElement;
const span = lineDiv.children[0] as HTMLSpanElement;
// Check cssText includes the ligature-disabling settings
// jsdom may normalize whitespace; check that both liga and calt are disabled
expect(span.style.cssText).toContain('font-feature-settings:');
expect(span.style.cssText).toContain("'liga' 0");
expect(span.style.cssText).toContain("'calt' 0");
});
it('span has text-align: center for glyph centering', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['m'] }));
const lineDiv = container.children[0] as HTMLDivElement;
const span = lineDiv.children[0] as HTMLSpanElement;
expect(span.style.textAlign).toBe('center');
});
it('span has pointer-events: none', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['a'] }));
const lineDiv = container.children[0] as HTMLDivElement;
const span = lineDiv.children[0] as HTMLSpanElement;
expect(span.style.pointerEvents).toBe('none');
});
// ─── Multi-line cursor positioning ────────────────────────────
it('cursor on wrapped line uses col 0 as base', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['first', 'ab'],
startCol: 5,
cellW: 10,
cellH: 20,
showCursor: true,
})
);
// Cursor at end of second line: col = 0 + 2 = 2
const cursor = container.children[container.children.length - 1] as HTMLSpanElement;
expect(cursor.style.left).toBe('20px'); // 2 * 10
expect(cursor.style.top).toBe('20px'); // row 1 * cellH
});
// ─── charTop/charHeight passed through ────────────────────────
it('accepts charTop and charHeight params without error', () => {
const container = document.createElement('div');
expect(() =>
renderOverlay(
container,
makeParams({
lines: ['test'],
charTop: 2,
charHeight: 14,
})
)
).not.toThrow();
expect(container.children.length).toBeGreaterThan(0);
});
// ─── Line div positioning regression ──────────────────────────
it('line div background color matches font.backgroundColor', () => {
const container = document.createElement('div');
const font: FontStyle = { ...FONT, backgroundColor: '#1a1a1a' };
renderOverlay(container, makeParams({ lines: ['x'], font }));
const lineDiv = container.children[0] as HTMLDivElement;
// jsdom normalizes hex to rgb()
expect(lineDiv.style.backgroundColor).toBe('rgb(26, 26, 26)');
});
it('empty line produces line div with no spans', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: [''] }));
const lineDiv = container.children[0] as HTMLDivElement;
expect(lineDiv.children.length).toBe(0);
});
// ─── CJK wide character support ───────────────────────────────
it('CJK characters get double-width spans', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['a你b'], cellW: 10 }));
const lineDiv = container.children[0] as HTMLDivElement;
expect(lineDiv.children.length).toBe(3);
const spanA = lineDiv.children[0] as HTMLSpanElement;
expect(spanA.textContent).toBe('a');
expect(spanA.style.left).toBe('0px');
expect(spanA.style.width).toBe('10px'); // 1 cell
const spanCJK = lineDiv.children[1] as HTMLSpanElement;
expect(spanCJK.textContent).toBe('你');
expect(spanCJK.style.left).toBe('10px'); // col 1
expect(spanCJK.style.width).toBe('20px'); // 2 cells
const spanB = lineDiv.children[2] as HTMLSpanElement;
expect(spanB.textContent).toBe('b');
expect(spanB.style.left).toBe('30px'); // col 3
expect(spanB.style.width).toBe('10px'); // 1 cell
});
it('cursor position accounts for CJK width', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({
lines: ['你好'],
startCol: 2,
cellW: 10,
showCursor: true,
})
);
// 你(2) + 好(2) = 4 visual cols, cursor at startCol(2) + 4 = 6
const cursor = container.children[container.children.length - 1] as HTMLSpanElement;
expect(cursor.style.left).toBe('60px');
});
it('mixed ASCII and CJK characters position correctly', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: ['hi你'], cellW: 8 }));
const lineDiv = container.children[0] as HTMLDivElement;
// h(col 0), i(col 1), 你(col 2, width 2)
const spanH = lineDiv.children[0] as HTMLSpanElement;
expect(spanH.style.left).toBe('0px');
const spanI = lineDiv.children[1] as HTMLSpanElement;
expect(spanI.style.left).toBe('8px');
const spanCJK = lineDiv.children[2] as HTMLSpanElement;
expect(spanCJK.style.left).toBe('16px');
expect(spanCJK.style.width).toBe('16px');
});
});
describe('charCellWidth', () => {
it('returns 1 for ASCII characters', () => {
expect(charCellWidth(null, 'a')).toBe(1);
expect(charCellWidth(null, '!')).toBe(1);
expect(charCellWidth(null, ' ')).toBe(1);
});
it('returns 2 for CJK ideographs', () => {
expect(charCellWidth(null, '你')).toBe(2);
expect(charCellWidth(null, '好')).toBe(2);
expect(charCellWidth(null, '中')).toBe(2);
});
it('returns 2 for Japanese hiragana', () => {
expect(charCellWidth(null, 'こ')).toBe(2);
expect(charCellWidth(null, 'ん')).toBe(2);
});
it('returns 2 for Korean syllables', () => {
expect(charCellWidth(null, '안')).toBe(2);
expect(charCellWidth(null, '녕')).toBe(2);
});
it('returns 2 for fullwidth forms', () => {
expect(charCellWidth(null, '\uff01')).toBe(2); // !
expect(charCellWidth(null, '\uff21')).toBe(2); // A
});
it('uses terminal unicode addon when available', () => {
const mockTerminal = {
unicode: { getStringCellWidth: (s: string) => (s === 'W' ? 2 : 1) },
} as any;
expect(charCellWidth(mockTerminal, 'W')).toBe(2);
expect(charCellWidth(mockTerminal, 'n')).toBe(1);
});
});
describe('stringCellWidth', () => {
it('sums individual character widths', () => {
expect(stringCellWidth(null, 'abc')).toBe(3);
expect(stringCellWidth(null, '你好')).toBe(4);
expect(stringCellWidth(null, 'a你b')).toBe(4);
});
it('returns 0 for empty string', () => {
expect(stringCellWidth(null, '')).toBe(0);
});
});
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,33 @@
import { defineConfig } from 'tsup';
export default defineConfig([
// Standard builds (CJS + ESM + DTS)
{
entry: ['src/index.ts'],
format: ['cjs', 'esm'],
dts: true,
clean: true,
},
// IIFE build for browser <script> tag usage
{
entry: ['src/index.ts'],
format: ['iife'],
globalName: 'XtermZerolagInput',
outDir: 'dist',
// Append global aliases so app.js can access classes directly
footer: {
js: [
'// Global aliases for browser usage',
'if(typeof window!=="undefined"){',
' window.ZerolagInputAddon=XtermZerolagInput.ZerolagInputAddon;',
' window.LocalEchoOverlay=class extends XtermZerolagInput.ZerolagInputAddon{',
' constructor(terminal){',
' super({prompt:{type:"character",char:"\\u276f",offset:2}});',
' this.activate(terminal);',
' }',
' };',
'}',
].join('\n'),
},
},
]);
Symlink
+1
View File
@@ -0,0 +1 @@
tools/remotion/public
+75 -11
View File
@@ -8,12 +8,15 @@
* 2. Copy static assets (web/public, templates)
* 3. Build vendor xterm bundles
* 4. Minify frontend assets (app.js, styles.css, mobile.css)
* 5. Compress with gzip + brotli
* 5. Content-hash cache busting (rename assets, rewrite index.html)
* 6. Compress with gzip + brotli
*/
import { execSync } from 'child_process';
import { appendFileSync, readFileSync, writeFileSync, renameSync } from 'fs';
import { createHash } from 'crypto';
import { fileURLToPath } from 'url';
import { join } from 'path';
import { join, extname, basename, dirname } from 'path';
const ROOT = join(fileURLToPath(import.meta.url), '..', '..');
@@ -26,24 +29,85 @@ function run(label, cmd) {
run('tsc', 'tsc');
run('chmod dist/index.js', 'chmod +x dist/index.js');
// 2. Copy static assets
// 2. Copy static assets (clean first to remove stale hashed files from previous builds)
run('clean public', 'rm -rf dist/web/public');
run('prepare dirs', 'mkdir -p dist/web dist/templates dist/web/public/vendor');
run('copy web assets', 'cp -r src/web/public dist/web/');
run('copy template', 'cp src/templates/case-template.md dist/templates/');
// 3. Vendor xterm bundles
run('xterm css', 'cp node_modules/xterm/css/xterm.css dist/web/public/vendor/');
run('xterm js', 'npx esbuild node_modules/xterm/lib/xterm.js --minify --outfile=dist/web/public/vendor/xterm.min.js');
run('xterm-addon-fit', 'npx esbuild node_modules/xterm-addon-fit/lib/xterm-addon-fit.js --minify --outfile=dist/web/public/vendor/xterm-addon-fit.min.js');
run('xterm-addon-webgl', 'cp node_modules/xterm-addon-webgl/lib/xterm-addon-webgl.js dist/web/public/vendor/xterm-addon-webgl.min.js');
run('xterm-addon-unicode11', 'npx esbuild node_modules/xterm-addon-unicode11/lib/xterm-addon-unicode11.js --minify --outfile=dist/web/public/vendor/xterm-addon-unicode11.min.js');
// 3. Vendor xterm bundles (xterm.js 6.x — @xterm scoped packages)
run('xterm css', 'cp node_modules/@xterm/xterm/css/xterm.css dist/web/public/vendor/');
run('xterm js', 'npx esbuild node_modules/@xterm/xterm/lib/xterm.js --minify --outfile=dist/web/public/vendor/xterm.min.js');
run('xterm-addon-fit', 'npx esbuild node_modules/@xterm/addon-fit/lib/addon-fit.js --minify --outfile=dist/web/public/vendor/xterm-addon-fit.min.js');
run('xterm-addon-webgl', 'cp node_modules/@xterm/addon-webgl/lib/addon-webgl.js dist/web/public/vendor/xterm-addon-webgl.min.js');
run('xterm-addon-unicode11', 'npx esbuild node_modules/@xterm/addon-unicode11/lib/addon-unicode11.js --minify --outfile=dist/web/public/vendor/xterm-addon-unicode11.min.js');
run('xterm-zerolag-input', 'npx esbuild packages/xterm-zerolag-input/src/zerolag-input-addon.ts --bundle --minify --format=iife --global-name=XtermZerolagInput --outfile=dist/web/public/vendor/xterm-zerolag-input.js');
// Append global aliases so app.js can use `new LocalEchoOverlay(terminal)`
appendFileSync(
join(ROOT, 'dist/web/public/vendor/xterm-zerolag-input.js'),
'\n// Global aliases for browser usage\n' +
'if(typeof window!=="undefined"){' +
'window.ZerolagInputAddon=XtermZerolagInput.ZerolagInputAddon;' +
'window.LocalEchoOverlay=class extends XtermZerolagInput.ZerolagInputAddon{' +
'constructor(terminal){' +
'super({prompt:{type:"character",char:"\\u276f",offset:2}});' +
'this.activate(terminal);' +
'}' +
'};' +
'}\n'
);
// 4. Minify frontend assets
run('minify app.js', 'npx esbuild dist/web/public/app.js --minify --drop:console --outfile=dist/web/public/app.js --allow-overwrite');
run('minify input-cjk.js', 'npx esbuild dist/web/public/input-cjk.js --minify --outfile=dist/web/public/input-cjk.js --allow-overwrite');
run('minify app.js', 'npx esbuild dist/web/public/app.js --minify --outfile=dist/web/public/app.js --allow-overwrite');
run('minify styles.css', 'npx esbuild dist/web/public/styles.css --minify --outfile=dist/web/public/styles.css --allow-overwrite');
run('minify mobile.css', 'npx esbuild dist/web/public/mobile.css --minify --outfile=dist/web/public/mobile.css --allow-overwrite');
// 5. Compress with gzip + brotli
// 5. Content-hash cache busting
console.log('\n[build] content-hash cache busting');
{
const distPublic = join(ROOT, 'dist/web/public');
const HASHABLE = [
'styles.css',
'mobile.css',
'constants.js',
'mobile-handlers.js',
'voice-input.js',
'notification-manager.js',
'keyboard-accessory.js',
'input-cjk.js',
'app.js',
'ralph-wizard.js',
'api-client.js',
'subagent-windows.js',
'vendor/xterm-zerolag-input.js',
];
const manifest = {};
for (const file of HASHABLE) {
const filePath = join(distPublic, file);
const content = readFileSync(filePath);
const hash = createHash('md5').update(content).digest('hex').slice(0, 8);
const ext = extname(file);
const base = basename(file, ext);
const dir = dirname(file);
const hashed = dir === '.' ? `${base}.${hash}${ext}` : `${dir}/${base}.${hash}${ext}`;
renameSync(filePath, join(distPublic, hashed));
manifest[file] = hashed;
}
// Rewrite index.html to reference hashed filenames
let html = readFileSync(join(distPublic, 'index.html'), 'utf8');
for (const [original, hashed] of Object.entries(manifest)) {
html = html.replaceAll(`"${original}"`, `"${hashed}"`);
}
writeFileSync(join(distPublic, 'index.html'), html);
console.log(' Hashed files:');
for (const [orig, hashed] of Object.entries(manifest)) {
console.log(` ${orig} -> ${hashed}`);
}
}
// 6. Compress with gzip + brotli
run(
'compress',
`for f in dist/web/public/*.js dist/web/public/*.css dist/web/public/*.html dist/web/public/vendor/*.js dist/web/public/vendor/*.css; do` +
+67 -9
View File
@@ -250,10 +250,10 @@ if (isGlobalInstall) {
} else {
try {
const require = createRequire(import.meta.url);
const xtermDir = join(require.resolve('xterm'), '..', '..');
const fitDir = join(require.resolve('xterm-addon-fit'), '..', '..');
const webglDir = join(require.resolve('xterm-addon-webgl'), '..', '..');
const unicode11Dir = join(require.resolve('xterm-addon-unicode11'), '..', '..');
const xtermDir = join(require.resolve('@xterm/xterm'), '..', '..');
const fitDir = join(require.resolve('@xterm/addon-fit'), '..', '..');
const webglDir = join(require.resolve('@xterm/addon-webgl'), '..', '..');
const unicode11Dir = join(require.resolve('@xterm/addon-unicode11'), '..', '..');
const vendorDir = join(srcDir, 'web', 'public', 'vendor');
const { mkdirSync, copyFileSync } = await import('fs');
@@ -263,19 +263,47 @@ if (isGlobalInstall) {
// Minify xterm JS for dev vendor dir (npm packages don't ship .min.js)
try {
execSync(`npx esbuild "${join(xtermDir, 'lib', 'xterm.js')}" --minify --outfile="${join(vendorDir, 'xterm.min.js')}"`, { stdio: 'pipe' });
execSync(`npx esbuild "${join(fitDir, 'lib', 'xterm-addon-fit.js')}" --minify --outfile="${join(vendorDir, 'xterm-addon-fit.min.js')}"`, { stdio: 'pipe' });
execSync(`npx esbuild "${join(unicode11Dir, 'lib', 'xterm-addon-unicode11.js')}" --minify --outfile="${join(vendorDir, 'xterm-addon-unicode11.min.js')}"`, { stdio: 'pipe' });
execSync(`npx esbuild "${join(fitDir, 'lib', 'addon-fit.js')}" --minify --outfile="${join(vendorDir, 'xterm-addon-fit.min.js')}"`, { stdio: 'pipe' });
execSync(`npx esbuild "${join(unicode11Dir, 'lib', 'addon-unicode11.js')}" --minify --outfile="${join(vendorDir, 'xterm-addon-unicode11.min.js')}"`, { stdio: 'pipe' });
console.log(colors.green('✓ xterm vendor files copied to src/web/public/vendor/'));
} catch {
// Fallback: copy unminified
copyFileSync(join(xtermDir, 'lib', 'xterm.js'), join(vendorDir, 'xterm.min.js'));
copyFileSync(join(fitDir, 'lib', 'xterm-addon-fit.js'), join(vendorDir, 'xterm-addon-fit.min.js'));
copyFileSync(join(unicode11Dir, 'lib', 'xterm-addon-unicode11.js'), join(vendorDir, 'xterm-addon-unicode11.min.js'));
copyFileSync(join(fitDir, 'lib', 'addon-fit.js'), join(vendorDir, 'xterm-addon-fit.min.js'));
copyFileSync(join(unicode11Dir, 'lib', 'addon-unicode11.js'), join(vendorDir, 'xterm-addon-unicode11.min.js'));
console.log(colors.green('✓ xterm vendor files copied') + colors.dim(' (unminified — esbuild not available)'));
}
// WebGL addon: copy unminified (matches build script behavior)
copyFileSync(join(webglDir, 'lib', 'xterm-addon-webgl.js'), join(vendorDir, 'xterm-addon-webgl.min.js'));
copyFileSync(join(webglDir, 'lib', 'addon-webgl.js'), join(vendorDir, 'xterm-addon-webgl.min.js'));
// xterm-zerolag-input: bundle local package as IIFE for <script> tag loading
try {
const zerolagSrc = join(import.meta.dirname, '..', 'packages', 'xterm-zerolag-input', 'src', 'zerolag-input-addon.ts');
const zerolagOut = join(vendorDir, 'xterm-zerolag-input.js');
execSync(
`npx esbuild "${zerolagSrc}" --bundle --format=iife --global-name=XtermZerolagInput --outfile="${zerolagOut}"`,
{ stdio: 'pipe' }
);
// Append global aliases so app.js can use `new LocalEchoOverlay(terminal)`
const { appendFileSync } = await import('fs');
appendFileSync(
zerolagOut,
'\n// Global aliases for browser usage\n' +
'if(typeof window!=="undefined"){' +
'window.ZerolagInputAddon=XtermZerolagInput.ZerolagInputAddon;' +
'window.LocalEchoOverlay=class extends XtermZerolagInput.ZerolagInputAddon{' +
'constructor(terminal){' +
'super({prompt:{type:"character",char:"\\u276f",offset:2}});' +
'this.activate(terminal);' +
'}' +
'};' +
'}\n'
);
console.log(colors.green('✓ xterm-zerolag-input bundled to vendor/'));
} catch {
console.log(colors.yellow('⚠ Failed to bundle xterm-zerolag-input — overlay may not work in dev mode'));
}
} catch (err) {
hasWarnings = true;
console.log(colors.yellow('⚠ Failed to copy xterm vendor files'));
@@ -284,6 +312,36 @@ if (isGlobalInstall) {
}
}
// ----------------------------------------------------------------------------
// 5. Install git pre-commit hook (format check)
// ----------------------------------------------------------------------------
if (!isGlobalInstall) {
try {
const { writeFileSync, mkdirSync } = await import('fs');
const gitHooksDir = join(import.meta.dirname, '..', '.git', 'hooks');
if (existsSync(join(import.meta.dirname, '..', '.git'))) {
mkdirSync(gitHooksDir, { recursive: true });
const hook = `#!/bin/bash
# Auto-installed by postinstall — prevents CI format failures
staged_ts=$(git diff --cached --name-only --diff-filter=ACM -- '*.ts')
[ -z "$staged_ts" ] && exit 0
echo "$staged_ts" | xargs npx prettier --check 2>&1
if [ $? -ne 0 ]; then
echo ""
echo "Pre-commit: Prettier check failed. Run 'npm run format' to fix."
exit 1
fi
`;
const hookPath = join(gitHooksDir, 'pre-commit');
writeFileSync(hookPath, hook, { mode: 0o755 });
console.log(colors.green('✓ Git pre-commit hook installed (prettier check)'));
}
} catch {
// Non-critical — git hook is a convenience
}
}
// ----------------------------------------------------------------------------
// Summary
// ----------------------------------------------------------------------------
+3 -6
View File
@@ -28,8 +28,8 @@ import { existsSync, readFileSync, unlinkSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { EventEmitter } from 'node:events';
import { getAugmentedPath } from './utils/claude-cli-resolver.js';
import { ANSI_ESCAPE_PATTERN_SIMPLE } from './utils/index.js';
import { getAugmentedPath, ANSI_ESCAPE_PATTERN_SIMPLE } from './utils/index.js';
import { AI_CHECK_MAX_BACKOFF_MS } from './config/ai-defaults.js';
// ========== Security Validation ==========
@@ -534,10 +534,7 @@ export abstract class AiCheckerBase<
// P1-005: Exponential backoff for errors
// Base cooldown * 2^(consecutiveErrors-1), capped at 5 minutes
const backoffMultiplier = Math.pow(2, this.consecutiveErrors - 1);
const backoffCooldownMs = Math.min(
this.config.errorCooldownMs * backoffMultiplier,
5 * 60 * 1000 // Max 5 minutes
);
const backoffCooldownMs = Math.min(this.config.errorCooldownMs * backoffMultiplier, AI_CHECK_MAX_BACKOFF_MS);
this.log(`Exponential backoff: ${Math.round(backoffCooldownMs / 1000)}s (error #${this.consecutiveErrors})`);
this.startCooldown(backoffCooldownMs);
}
+14 -6
View File
@@ -30,6 +30,14 @@ import {
type AiCheckerResultBase,
type AiCheckerStateBase,
} from './ai-checker-base.js';
import {
AI_CHECK_MODEL,
AI_IDLE_CHECK_MAX_CONTEXT,
AI_IDLE_CHECK_TIMEOUT_MS,
AI_IDLE_CHECK_COOLDOWN_MS,
AI_IDLE_CHECK_ERROR_COOLDOWN_MS,
AI_CHECK_MAX_CONSECUTIVE_ERRORS,
} from './config/ai-defaults.js';
// ========== Types ==========
@@ -45,12 +53,12 @@ export type AiCheckState = AiCheckerStateBase<AiCheckVerdict>;
const DEFAULT_AI_CHECK_CONFIG: AiIdleCheckConfig = {
enabled: true,
model: 'claude-opus-4-5-20251101',
maxContextChars: 16000,
checkTimeoutMs: 90000,
cooldownMs: 180000,
errorCooldownMs: 60000,
maxConsecutiveErrors: 3,
model: AI_CHECK_MODEL,
maxContextChars: AI_IDLE_CHECK_MAX_CONTEXT,
checkTimeoutMs: AI_IDLE_CHECK_TIMEOUT_MS,
cooldownMs: AI_IDLE_CHECK_COOLDOWN_MS,
errorCooldownMs: AI_IDLE_CHECK_ERROR_COOLDOWN_MS,
maxConsecutiveErrors: AI_CHECK_MAX_CONSECUTIVE_ERRORS,
};
/** Pattern to match IDLE or WORKING as the first word of output */
+14 -6
View File
@@ -29,6 +29,14 @@ import {
type AiCheckerResultBase,
type AiCheckerStateBase,
} from './ai-checker-base.js';
import {
AI_CHECK_MODEL,
AI_PLAN_CHECK_MAX_CONTEXT,
AI_PLAN_CHECK_TIMEOUT_MS,
AI_PLAN_CHECK_COOLDOWN_MS,
AI_PLAN_CHECK_ERROR_COOLDOWN_MS,
AI_CHECK_MAX_CONSECUTIVE_ERRORS,
} from './config/ai-defaults.js';
// ========== Types ==========
@@ -44,12 +52,12 @@ export type AiPlanCheckState = AiCheckerStateBase<AiPlanCheckVerdict>;
const DEFAULT_PLAN_CHECK_CONFIG: AiPlanCheckConfig = {
enabled: true,
model: 'claude-opus-4-5-20251101',
maxContextChars: 8000,
checkTimeoutMs: 60000,
cooldownMs: 30000,
errorCooldownMs: 30000,
maxConsecutiveErrors: 3,
model: AI_CHECK_MODEL,
maxContextChars: AI_PLAN_CHECK_MAX_CONTEXT,
checkTimeoutMs: AI_PLAN_CHECK_TIMEOUT_MS,
cooldownMs: AI_PLAN_CHECK_COOLDOWN_MS,
errorCooldownMs: AI_PLAN_CHECK_ERROR_COOLDOWN_MS,
maxConsecutiveErrors: AI_CHECK_MAX_CONSECUTIVE_ERRORS,
};
/** Pattern to match PLAN_MODE or NOT_PLAN_MODE as the first word(s) of output */
+36 -50
View File
@@ -15,6 +15,7 @@
import { EventEmitter } from 'node:events';
import { v4 as uuidv4 } from 'uuid';
import { ActiveBashTool } from './types.js';
import { CleanupManager, Debouncer, stripAnsi } from './utils/index.js';
// ========== Configuration Constants ==========
@@ -145,15 +146,14 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
private _workingDir: string;
private _homeDir: string;
// Track auto-remove timers for cleanup
private _autoRemoveTimers: Set<ReturnType<typeof setTimeout>> = new Set();
// Centralized resource cleanup for auto-remove timers
private cleanup = new CleanupManager();
// Flag to prevent operations after destroy
private _destroyed: boolean = false;
// Debouncing
private _pendingUpdate: boolean = false;
private _updateTimer: ReturnType<typeof setTimeout> | null = null;
private _updateDeb = new Debouncer(EVENT_DEBOUNCE_MS);
constructor(config: BashToolParserConfig) {
super();
@@ -462,7 +462,7 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
* Process a single line of terminal output (raw — will strip ANSI).
*/
private processLine(line: string): void {
const cleanLine = this.stripAnsi(line);
const cleanLine = stripAnsi(line);
this.processCleanLine(cleanLine);
}
@@ -524,13 +524,15 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
this.scheduleUpdate();
// Remove completed tool after a short delay to allow UI to show completion
const timer = setTimeout(() => {
this._autoRemoveTimers.delete(timer);
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this.scheduleUpdate();
}, 2000);
this._autoRemoveTimers.add(timer);
this.cleanup.setTimeout(
() => {
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this.scheduleUpdate();
},
2000,
{ description: 'auto-remove completed tool' }
);
}
this._lastToolId = null;
return;
@@ -562,13 +564,15 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
this.scheduleUpdate();
// Auto-remove suggestions after 30 seconds
const timer = setTimeout(() => {
this._autoRemoveTimers.delete(timer);
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this.scheduleUpdate();
}, 30000);
this._autoRemoveTimers.add(timer);
this.cleanup.setTimeout(
() => {
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this.scheduleUpdate();
},
30000,
{ description: 'auto-remove suggestion tool' }
);
return;
}
@@ -599,13 +603,15 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
this.scheduleUpdate();
// Auto-remove after 60 seconds
const timer = setTimeout(() => {
this._autoRemoveTimers.delete(timer);
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this.scheduleUpdate();
}, 60000);
this._autoRemoveTimers.add(timer);
this.cleanup.setTimeout(
() => {
if (this._destroyed) return;
this._activeTools.delete(tool.id);
this.scheduleUpdate();
},
60000,
{ description: 'auto-remove log file tool' }
);
}
}
@@ -662,26 +668,13 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
return this.deduplicatePaths(rawPaths);
}
/**
* Strip ANSI escape codes from a string.
*/
private stripAnsi(str: string): string {
// Comprehensive ANSI pattern
// eslint-disable-next-line no-control-regex
return str.replace(/\x1b(?:\[[0-9;?]*[A-Za-z]|\][^\x07\x1b]*(?:\x07|\x1b\\)|[=>])/g, '');
}
/**
* Schedule a debounced update emission.
*/
private scheduleUpdate(): void {
if (this._pendingUpdate) return;
this._pendingUpdate = true;
this._updateTimer = setTimeout(() => {
this._pendingUpdate = false;
this._updateDeb.schedule(() => {
this.emitUpdate();
}, EVENT_DEBOUNCE_MS);
});
}
/**
@@ -696,15 +689,8 @@ export class BashToolParser extends EventEmitter<BashToolParserEvents> {
*/
destroy(): void {
this._destroyed = true;
if (this._updateTimer) {
clearTimeout(this._updateTimer);
this._updateTimer = null;
}
// Clear all auto-remove timers to prevent orphaned callbacks
for (const timer of this._autoRemoveTimers) {
clearTimeout(timer);
}
this._autoRemoveTimers.clear();
this._updateDeb.dispose();
this.cleanup.dispose();
this._activeTools.clear();
this.removeAllListeners();
}
+58
View File
@@ -0,0 +1,58 @@
/**
* @fileoverview Default model, context limits, and timing for AI-powered checkers.
*
* Centralizes the AI model identifier, context window sizes, and timeout/cooldown
* defaults used by the idle checker, plan checker, respawn controller defaults,
* and respawn route fallbacks. Change values here when tuning AI check behavior.
*
* @module config/ai-defaults
*/
// ============================================================================
// Model & Context
// ============================================================================
/** Default model for AI idle and plan checkers */
export const AI_CHECK_MODEL = 'claude-opus-4-5-20251101';
/** Max context chars for idle checker (~4k tokens) */
export const AI_IDLE_CHECK_MAX_CONTEXT = 16000;
/** Max context chars for plan checker (~2k tokens, plan mode UI is compact) */
export const AI_PLAN_CHECK_MAX_CONTEXT = 8000;
// ============================================================================
// AI Idle Checker Timing
// ============================================================================
/** Timeout for AI idle check (90 seconds — thinking can be slow) */
export const AI_IDLE_CHECK_TIMEOUT_MS = 90_000;
/** Cooldown after WORKING verdict (3 minutes) */
export const AI_IDLE_CHECK_COOLDOWN_MS = 180_000;
/** Cooldown after AI idle check error (1 minute) */
export const AI_IDLE_CHECK_ERROR_COOLDOWN_MS = 60_000;
// ============================================================================
// AI Plan Checker Timing
// ============================================================================
/** Timeout for AI plan check (60 seconds — allows time for thinking) */
export const AI_PLAN_CHECK_TIMEOUT_MS = 60_000;
/** Cooldown after NOT_PLAN_MODE verdict (30 seconds) */
export const AI_PLAN_CHECK_COOLDOWN_MS = 30_000;
/** Cooldown after AI plan check error (30 seconds) */
export const AI_PLAN_CHECK_ERROR_COOLDOWN_MS = 30_000;
// ============================================================================
// Shared AI Checker Limits
// ============================================================================
/** Max consecutive errors before disabling an AI checker */
export const AI_CHECK_MAX_CONSECUTIVE_ERRORS = 3;
/** Maximum exponential backoff cap for AI checker errors (5 minutes) */
export const AI_CHECK_MAX_BACKOFF_MS = 5 * 60 * 1000;
+35
View File
@@ -0,0 +1,35 @@
/**
* @fileoverview Authentication, rate limiting, and hook security constants.
*
* Controls auth session lifecycle, brute-force protection,
* and Claude Code hook timeouts.
*
* @module config/auth-config
*/
// ============================================================================
// Session Cookies
// ============================================================================
/** Auth session cookie TTL — matches autonomous run length (ms) */
export const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
/** Max concurrent auth sessions per server */
export const MAX_AUTH_SESSIONS = 100;
// ============================================================================
// Rate Limiting
// ============================================================================
/** Max failed auth attempts per IP before 429 rejection */
export const AUTH_FAILURE_MAX = 10;
/** Failed auth attempt tracking window (ms) */
export const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
// ============================================================================
// Hooks
// ============================================================================
/** Timeout for Claude Code hook curl commands (ms) */
export const HOOK_TIMEOUT_MS = 10000;
+11 -5
View File
@@ -56,11 +56,6 @@ export const TRIM_TEXT_TO = 768 * 1024; // 768KB
*/
export const MAX_MESSAGES = 1000;
/**
* Number of messages to keep when trimming (80% of max).
*/
export const TRIM_MESSAGES_TO = 800;
// ============================================================================
// Line Buffer Limits
// ============================================================================
@@ -85,3 +80,14 @@ export const MAX_RESPAWN_BUFFER_SIZE = 1 * 1024 * 1024; // 1MB
* Size to trim respawn buffer to when max is exceeded.
*/
export const TRIM_RESPAWN_BUFFER_TO = 512 * 1024; // 512KB
// ============================================================================
// File Peek Limits
// ============================================================================
/**
* Maximum bytes to read when peeking at the beginning of a file.
* Used with `createReadStream({ end })` (inclusive) to read the first 8KB,
* which is enough to extract metadata from the first few JSONL lines.
*/
export const FILE_PEEK_BYTES = 8 * 1024 - 1; // 8KB (inclusive end offset)
+10
View File
@@ -0,0 +1,10 @@
/**
* @fileoverview Shared exec timeout constant.
*
* Used by CLI resolvers and tmux-manager for execSync/exec calls.
*
* @module config/exec-timeout
*/
/** Timeout for exec commands (5 seconds) */
export const EXEC_TIMEOUT_MS = 5000;
+9
View File
@@ -47,6 +47,15 @@ export const MAX_TODOS_PER_SESSION = 500;
// Pending Tool Calls Limits
// ============================================================================
// ============================================================================
// Agent Tracking Limits
// ============================================================================
/**
* Maximum agents to track across all sessions (LRU eviction when exceeded).
*/
export const MAX_TRACKED_AGENTS = 500;
/**
* Maximum pending tool calls to track per subagent.
* Entries should be cleaned up on tool_result, but this prevents leaks.
+88
View File
@@ -0,0 +1,88 @@
/**
* @fileoverview Web server performance and scheduling constants.
*
* Controls terminal batching throughput, SSE health checking,
* state persistence debouncing, and scheduled run timing.
*
* @module config/server-timing
*/
// ============================================================================
// Terminal & SSE Performance
// ============================================================================
/** Terminal data batching interval — targets 60fps (ms) */
export const TERMINAL_BATCH_INTERVAL = 16;
/** Immediate flush threshold for terminal batches (bytes).
* Set high (32KB) to allow effective batching; avg Ink events are ~14KB. */
export const BATCH_FLUSH_THRESHOLD = 32 * 1024;
/** Task event batching interval (ms) */
export const TASK_UPDATE_BATCH_INTERVAL = 100;
/** SSE heartbeat interval — sends padded keepalive to flush proxy buffers (ms).
* 15s is fast enough to keep Cloudflare tunnel buffers flushed while avoiding
* excessive bandwidth. Also serves as dead-client detection. */
export const SSE_HEARTBEAT_INTERVAL = 15 * 1000;
/** SSE padding size (bytes). Cloudflare quick tunnels buffer small SSE events;
* appending ~8KB of SSE comment padding forces the proxy to flush immediately.
* SSE comments (lines starting with ':') are silently ignored by EventSource. */
export const SSE_PADDING_SIZE = 8 * 1024;
// ============================================================================
// State Persistence
// ============================================================================
/** State update debounce — batches expensive toDetailedState() calls (ms) */
export const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
/** Sessions list cache TTL — avoids re-serializing on every SSE init (ms) */
export const SESSIONS_LIST_CACHE_TTL = 1000;
// ============================================================================
// Scheduled Runs
// ============================================================================
/** Scheduled runs cleanup check interval (ms) */
export const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
/** Completed scheduled run max age before cleanup (ms) */
export const SCHEDULED_RUN_MAX_AGE = 60 * 60 * 1000;
/** Session limit retry wait before retrying (ms) */
export const SESSION_LIMIT_WAIT_MS = 5000;
/** Pause between scheduled run iterations (ms) */
export const ITERATION_PAUSE_MS = 2000;
// ============================================================================
// Mux Stats
// ============================================================================
/** Mux stats collection interval (ms) */
export const STATS_COLLECTION_INTERVAL_MS = 2000;
// ============================================================================
// Process Error Recovery
// ============================================================================
/** Max consecutive unhandled errors before auto-restart */
export const MAX_CONSECUTIVE_ERRORS = 5;
/** Error counter reset interval — forgives errors after quiet period (ms) */
export const ERROR_RESET_MS = 60_000;
// ============================================================================
// Common Cleanup Intervals
// ============================================================================
/** Standard 1-minute cleanup/check interval used by multiple subsystems (ms) */
export const CLEANUP_CHECK_INTERVAL_MS = 60_000;
/** Standard 1-hour max age for stale/completed data (ms) */
export const STALE_DATA_MAX_AGE_MS = 60 * 60 * 1000;
/** Standard 5-minute inactivity timeout for streams and caches (ms) */
export const INACTIVITY_TIMEOUT_MS = 5 * 60 * 1000;
+18
View File
@@ -0,0 +1,18 @@
/**
* @fileoverview Agent Teams polling and cache configuration.
*
* Controls how frequently TeamWatcher polls ~/.claude/teams/
* and how many teams/tasks are cached in memory.
*
* @module config/team-config
*/
/** Team directory poll interval (ms) */
export const TEAM_POLL_INTERVAL_MS = 30_000;
/** Max cached team configs (LRU eviction) */
export const MAX_CACHED_TEAMS = 50;
/** Max cached team tasks and inbox messages (LRU eviction).
* Used for both teamTasks and inboxCache maps. */
export const MAX_CACHED_TASKS = 200;
+15
View File
@@ -0,0 +1,15 @@
/**
* @fileoverview Terminal dimension and input validation limits.
*
* Used by API routes to validate resize, input, and session
* creation requests. Separate from buffer-limits.ts which
* controls memory buffer sizes.
*
* @module config/terminal-limits
*/
/** Max input length per API request (bytes) */
export const MAX_INPUT_LENGTH = 64 * 1024;
/** Max session name length (chars) */
export const MAX_SESSION_NAME_LENGTH = 128;
+47
View File
@@ -0,0 +1,47 @@
/**
* @fileoverview Cloudflare tunnel and QR authentication constants.
*
* Controls QR token rotation timing, rate limiting,
* and tunnel process lifecycle.
*
* @module config/tunnel-config
*/
// ============================================================================
// QR Token Rotation
// ============================================================================
/** QR token auto-rotation interval (ms) */
export const QR_TOKEN_TTL_MS = 60_000;
/** Grace period — previous token still valid during rotation (ms) */
export const QR_TOKEN_GRACE_MS = 90_000;
/** Length of the short code in QR URL path (chars) */
export const SHORT_CODE_LENGTH = 6;
// ============================================================================
// QR Rate Limiting
// ============================================================================
/** Global rate limit for QR auth attempts across all IPs */
export const QR_RATE_LIMIT_MAX = 30;
/** QR rate limit reset window (ms) */
export const QR_RATE_LIMIT_WINDOW_MS = 60_000;
/** Per-IP rate limit for QR auth failures (separate from Basic Auth AUTH_FAILURE_MAX) */
export const QR_AUTH_FAILURE_MAX = 10;
// ============================================================================
// Tunnel Process Lifecycle
// ============================================================================
/** Max time to wait for cloudflared URL before timeout (ms) */
export const URL_TIMEOUT_MS = 30_000;
/** Restart delay after unexpected tunnel exit (ms) */
export const RESTART_DELAY_MS = 5_000;
/** SIGTERM → SIGKILL escalation timeout (ms) */
export const FORCE_KILL_MS = 5_000;
+3 -2
View File
@@ -16,6 +16,7 @@ import { existsSync, statSync, realpathSync } from 'node:fs';
import { resolve, relative, isAbsolute } from 'node:path';
import { homedir } from 'node:os';
import { EventEmitter } from 'node:events';
import { CLEANUP_CHECK_INTERVAL_MS, INACTIVITY_TIMEOUT_MS } from './config/server-timing.js';
// ========== Configuration Constants ==========
@@ -39,7 +40,7 @@ const MAX_STREAMS_PER_SESSION = 5;
* Inactivity timeout for streams (5 minutes).
* Streams with no data for this long will be auto-closed.
*/
const STREAM_INACTIVITY_TIMEOUT_MS = 5 * 60 * 1000;
const STREAM_INACTIVITY_TIMEOUT_MS = INACTIVITY_TIMEOUT_MS;
// ========== Types ==========
@@ -129,7 +130,7 @@ export class FileStreamManager extends EventEmitter {
constructor() {
super();
// Start cleanup timer for inactive streams
this.cleanupTimer = setInterval(() => this.cleanupInactiveStreams(), 60 * 1000);
this.cleanupTimer = setInterval(() => this.cleanupInactiveStreams(), CLEANUP_CHECK_INTERVAL_MS);
}
// ========== Public Methods ==========
+28 -12
View File
@@ -1,11 +1,26 @@
/**
* @fileoverview Claude Code hooks configuration generator
* @fileoverview Claude Code hooks configuration generator.
*
* Generates .claude/settings.local.json with hook definitions that POST
* to Codeman's /api/hook-event endpoint when Claude Code fires
* notification or stop hooks. Uses $CODEMAN_API_URL and
* $CODEMAN_SESSION_ID env vars (set on every managed session) so the
* config is static per case directory.
* Generates `.claude/settings.local.json` with hook definitions that POST
* to Codeman's `/api/hook-event` endpoint when Claude Code fires hooks.
* Uses `$CODEMAN_API_URL` and `$CODEMAN_SESSION_ID` env vars (set on every
* managed session) so the config is static per case directory.
*
* Key exports:
* - `generateHooksConfig()` — returns hooks object for settings.local.json
* - `writeHooksConfig(casePath)` — writes hooks + env config to disk
* - `updateCaseEnvVars(casePath, envVars)` — merges env vars into settings
*
* Hook events generated: `idle_prompt`, `permission_prompt`, `elicitation_dialog`,
* `stop`, `teammate_idle`, `task_completed`
*
* Hook categories: `Notification` (3 matchers), `Stop` (1), `TeammateIdle` (1),
* `TaskCompleted` (1)
*
* @dependencies types (HookEventType), config/auth-config (HOOK_TIMEOUT_MS)
* @consumedby web/server (session creation), session-cli-builder (env setup)
*
* @module hooks-config
*/
import { existsSync } from 'node:fs';
@@ -13,6 +28,7 @@ import { readFile, writeFile, mkdir } from 'node:fs/promises';
import { join } from 'node:path';
import type { HookEventType } from './types.js';
import { HOOK_TIMEOUT_MS } from './config/auth-config.js';
/**
* Generates the hooks section for .claude/settings.local.json
@@ -38,30 +54,30 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
Notification: [
{
matcher: 'idle_prompt',
hooks: [{ type: 'command', command: curlCmd('idle_prompt'), timeout: 10000 }],
hooks: [{ type: 'command', command: curlCmd('idle_prompt'), timeout: HOOK_TIMEOUT_MS }],
},
{
matcher: 'permission_prompt',
hooks: [{ type: 'command', command: curlCmd('permission_prompt'), timeout: 10000 }],
hooks: [{ type: 'command', command: curlCmd('permission_prompt'), timeout: HOOK_TIMEOUT_MS }],
},
{
matcher: 'elicitation_dialog',
hooks: [{ type: 'command', command: curlCmd('elicitation_dialog'), timeout: 10000 }],
hooks: [{ type: 'command', command: curlCmd('elicitation_dialog'), timeout: HOOK_TIMEOUT_MS }],
},
],
Stop: [
{
hooks: [{ type: 'command', command: curlCmd('stop'), timeout: 10000 }],
hooks: [{ type: 'command', command: curlCmd('stop'), timeout: HOOK_TIMEOUT_MS }],
},
],
TeammateIdle: [
{
hooks: [{ type: 'command', command: curlCmd('teammate_idle'), timeout: 10000 }],
hooks: [{ type: 'command', command: curlCmd('teammate_idle'), timeout: HOOK_TIMEOUT_MS }],
},
],
TaskCompleted: [
{
hooks: [{ type: 'command', command: curlCmd('task_completed'), timeout: 10000 }],
hooks: [{ type: 'command', command: curlCmd('task_completed'), timeout: HOOK_TIMEOUT_MS }],
},
],
},
+17 -31
View File
@@ -13,6 +13,7 @@ import { watch, type FSWatcher } from 'chokidar';
import { basename, extname, relative } from 'node:path';
import { statSync } from 'node:fs';
import type { ImageDetectedEvent } from './types.js';
import { KeyedDebouncer } from './utils/index.js';
// ========== Types ==========
@@ -65,11 +66,11 @@ export class ImageWatcher extends EventEmitter {
/** Map of sessionId -> working directory path */
private sessionDirs = new Map<string, string>();
/** Debounce timers for rapid image creation (keyed by filePath) */
private debounceTimers = new Map<string, NodeJS.Timeout>();
/** Per-file debouncer for rapid image creation */
private fileDeb = new KeyedDebouncer(DEBOUNCE_DELAY_MS);
/** Track which session owns each debounce timer (for cleanup) */
private timerToSession = new Map<string, string>();
/** Track which session owns each debounced file (for cleanup) */
private fileToSession = new Map<string, string>();
/** Per-session burst tracking: sessionId -> { count, windowStart } */
private burstTrackers = new Map<string, { count: number; windowStart: number }>();
@@ -118,11 +119,8 @@ export class ImageWatcher extends EventEmitter {
this.sessionDirs.clear();
// Clear all debounce timers
for (const timer of this.debounceTimers.values()) {
clearTimeout(timer);
}
this.debounceTimers.clear();
this.timerToSession.clear();
this.fileDeb.dispose();
this.fileToSession.clear();
this.burstTrackers.clear();
}
@@ -212,20 +210,15 @@ export class ImageWatcher extends EventEmitter {
this.sessionDirs.delete(sessionId);
// Clear any pending debounce timers for this session
// Collect keys first to avoid iterator invalidation during deletion
const toDelete: string[] = [];
for (const [filePath, ownerId] of this.timerToSession) {
const toCancel: string[] = [];
for (const [filePath, ownerId] of this.fileToSession) {
if (ownerId === sessionId) {
toDelete.push(filePath);
toCancel.push(filePath);
}
}
for (const filePath of toDelete) {
const timer = this.debounceTimers.get(filePath);
if (timer) {
clearTimeout(timer);
this.debounceTimers.delete(filePath);
}
this.timerToSession.delete(filePath);
for (const filePath of toCancel) {
this.fileDeb.cancelKey(filePath);
this.fileToSession.delete(filePath);
}
this.burstTrackers.delete(sessionId);
}
@@ -269,22 +262,15 @@ export class ImageWatcher extends EventEmitter {
}
// Debounce rapid file creation (e.g., multiple screenshots quickly)
const existingTimer = this.debounceTimers.get(filePath);
if (existingTimer) {
clearTimeout(existingTimer);
}
const timer = setTimeout(() => {
this.debounceTimers.delete(filePath);
this.timerToSession.delete(filePath);
this.fileDeb.schedule(filePath, () => {
this.fileToSession.delete(filePath);
this.emitImageDetected(sessionId, filePath);
// Increment burst count on actual emission (not on detection)
const b = this.burstTrackers.get(sessionId);
if (b) b.count++;
}, DEBOUNCE_DELAY_MS);
});
this.debounceTimers.set(filePath, timer);
this.timerToSession.set(filePath, sessionId);
this.fileToSession.set(filePath, sessionId);
}
/**
+2 -2
View File
@@ -14,10 +14,10 @@ import { program } from './cli.js';
// In web mode, we should NOT exit on transient errors — log and continue
const isWebMode = process.argv.includes('web');
import { MAX_CONSECUTIVE_ERRORS, ERROR_RESET_MS } from './config/server-timing.js';
// Track consecutive unhandled errors in web mode — restart after too many
let consecutiveErrors = 0;
const MAX_CONSECUTIVE_ERRORS = 5;
const ERROR_RESET_MS = 60000; // Reset counter after 1 minute of no errors
let errorResetTimer: ReturnType<typeof setTimeout> | null = null;
function trackError(): void {
+4
View File
@@ -61,6 +61,8 @@ export interface CreateSessionOptions {
claudeMode?: ClaudeMode;
allowedTools?: string;
openCodeConfig?: OpenCodeConfig;
/** When restoring after reboot, resume a previous Claude conversation by its session ID */
resumeSessionId?: string;
}
/** Options for respawning a dead pane. */
@@ -73,6 +75,8 @@ export interface RespawnPaneOptions {
claudeMode?: ClaudeMode;
allowedTools?: string;
openCodeConfig?: OpenCodeConfig;
/** Resume a previous Claude conversation when respawning */
resumeSessionId?: string;
}
/**
+5 -31
View File
@@ -20,36 +20,10 @@ import type { TerminalMultiplexer } from './mux-interface.js';
import { existsSync, mkdirSync, writeFileSync } from 'node:fs';
import { join } from 'node:path';
import { RESEARCH_AGENT_PROMPT, PLANNER_PROMPT } from './prompts/index.js';
import { PlanTaskStatus, TddPhase } from './types.js';
import type { PlanItem } from './types.js';
// ============================================================================
// Types
// ============================================================================
/** Development phase in TDD cycle (alias for TddPhase) */
export type PlanPhase = TddPhase;
/**
* Plan item with TDD structure.
*/
export interface PlanItem {
id?: string;
content: string;
priority: 'P0' | 'P1' | 'P2' | null;
source?: string;
rationale?: string;
verificationCriteria?: string;
testCommand?: string;
dependencies?: string[];
status?: PlanTaskStatus;
attempts?: number;
lastError?: string;
completedAt?: number;
complexity?: 'low' | 'medium' | 'high';
tddPhase?: PlanPhase;
pairedWith?: string;
reviewChecklist?: string[];
}
// Re-export for backward compatibility
export type { PlanItem };
export interface ResearchResult {
success: boolean;
@@ -547,7 +521,7 @@ export class PlanOrchestrator {
} finally {
// Always clean up session and progress interval — centralizing here
// prevents the race where cancel() and catch both try to manage the set
await session.stop().catch(() => {});
await session.stop().catch(() => {}); // Ignore - session cleanup is best-effort in finally block
this.runningSessions.delete(session);
clearInterval(progressInterval);
}
@@ -677,7 +651,7 @@ export class PlanOrchestrator {
} finally {
// Always clean up session and progress interval — centralizing here
// prevents the race where cancel() and catch both try to manage the set
await session.stop().catch(() => {});
await session.stop().catch(() => {}); // Ignore - session cleanup is best-effort in finally block
this.runningSessions.delete(session);
clearInterval(progressInterval);
}
+7 -8
View File
@@ -1,14 +1,13 @@
/**
* Planner Prompt - Single agent for TDD plan generation
*
* Combines what was previously 5 separate agents:
* - Requirements Analyst (redundant)
* - Architecture Planner (redundant)
* - Testing Specialist (kept - TDD focus)
* - Risk Analyst (redundant)
* - Verification Expert (kept - structure)
* @fileoverview Planner Prompt — single TDD plan generator combining
* requirements analysis, architecture, testing, risk, and verification
* into one agent (previously 5 separate agents).
*
* Placeholders: {TASK}, {RESEARCH_CONTEXT}
*
* @dependencies none (pure template)
* @consumedby prompts/index (re-export), plan-orchestrator
* @module prompts/planner
*/
export const PLANNER_PROMPT = `You are a TDD Plan Generator. Create a complete implementation plan with test-first approach.
+6 -4
View File
@@ -1,10 +1,12 @@
/**
* Research Agent Prompt
*
* Gathers external resources, codebase patterns, and technical context
* before other agents analyze the task.
* @fileoverview Research Agent Prompt — gathers codebase patterns, external
* resources, and technical context before the planner analyzes the task.
*
* Placeholders: {TASK}, {WORKING_DIR}
*
* @dependencies none (pure template)
* @consumedby prompts/index (re-export), plan-orchestrator
* @module prompts/research-agent
*/
export const RESEARCH_AGENT_PROMPT = `You are a Research Specialist preparing context for an implementation task. Your job is to gather all relevant information that will help the development team succeed.
+4 -11
View File
@@ -11,6 +11,7 @@ import { join } from 'node:path';
import { homedir } from 'node:os';
import webpush from 'web-push';
import type { VapidKeys, PushSubscriptionRecord } from './types.js';
import { Debouncer } from './utils/index.js';
const DATA_DIR = join(homedir(), '.codeman');
const KEYS_FILE = join(DATA_DIR, 'push-keys.json');
@@ -20,7 +21,7 @@ const SAVE_DEBOUNCE_MS = 500;
export class PushSubscriptionStore {
private vapidKeys: VapidKeys | null = null;
private subscriptions: Map<string, PushSubscriptionRecord> = new Map();
private saveTimer: NodeJS.Timeout | null = null;
private saveDeb = new Debouncer(SAVE_DEBOUNCE_MS);
private _disposed = false;
constructor() {
@@ -149,10 +150,7 @@ export class PushSubscriptionStore {
/** Schedule a debounced save */
private scheduleSave(): void {
if (this._disposed) return;
if (this.saveTimer) clearTimeout(this.saveTimer);
this.saveTimer = setTimeout(() => {
this.flushSave();
}, SAVE_DEBOUNCE_MS);
this.saveDeb.schedule(() => this.flushSave());
}
/** Immediately persist subscriptions to disk */
@@ -169,11 +167,6 @@ export class PushSubscriptionStore {
dispose(): void {
if (this._disposed) return;
this._disposed = true;
if (this.saveTimer) {
clearTimeout(this.saveTimer);
this.saveTimer = null;
}
// Final flush
this.flushSave();
this.saveDeb.flush(() => this.flushSave());
}
}
+3 -4
View File
@@ -9,6 +9,7 @@
import { existsSync, readFileSync } from 'node:fs';
import { join } from 'node:path';
import { execPattern } from './utils/index.js';
// Pattern to extract completion phrase from CLAUDE.md
// Matches <promise>PHRASE</promise> with optional whitespace
@@ -83,9 +84,7 @@ export function parseRalphLoopConfigFromContent(content: string): RalphLoopConfi
};
// Parse each YAML line
let match;
YAML_LINE_PATTERN.lastIndex = 0;
while ((match = YAML_LINE_PATTERN.exec(yaml)) !== null) {
execPattern(YAML_LINE_PATTERN, yaml, (match) => {
const key = match[1].toLowerCase();
const value = match[2].trim();
@@ -103,7 +102,7 @@ export function parseRalphLoopConfigFromContent(content: string): RalphLoopConfi
config.completionPromise = value.toUpperCase();
break;
}
}
});
return config;
}
+366
View File
@@ -0,0 +1,366 @@
/**
* @fileoverview RalphFixPlanWatcher - Watches @fix_plan.md for changes
*
* Monitors the @fix_plan.md file in the session's working directory
* for changes, parsing todo items from the markdown format.
*
* Extracted from ralph-tracker.ts as part of domain splitting.
*
* @module ralph-fix-plan-watcher
*/
import { EventEmitter } from 'node:events';
import { readFile } from 'node:fs/promises';
import { existsSync, FSWatcher, watch as fsWatch } from 'node:fs';
import { join } from 'node:path';
import type { RalphTodoStatus, RalphTodoPriority, RalphTodoItem } from './types.js';
// ========== @fix_plan.md Generation & Import Utility Functions ==========
/**
* Generate @fix_plan.md content from todo items.
* Groups todos by priority and status.
*
* @param todos - Array of todo items
* @returns Markdown content for @fix_plan.md
*/
export function generateFixPlanMarkdown(todos: RalphTodoItem[]): string {
const lines: string[] = ['# Fix Plan', ''];
// Group by priority
const p0: RalphTodoItem[] = [];
const p1: RalphTodoItem[] = [];
const p2: RalphTodoItem[] = [];
const noPriority: RalphTodoItem[] = [];
const completed: RalphTodoItem[] = [];
for (const todo of todos) {
if (todo.status === 'completed') {
completed.push(todo);
} else if (todo.priority === 'P0') {
p0.push(todo);
} else if (todo.priority === 'P1') {
p1.push(todo);
} else if (todo.priority === 'P2') {
p2.push(todo);
} else {
noPriority.push(todo);
}
}
// High Priority (P0)
if (p0.length > 0) {
lines.push('## High Priority (P0)');
for (const todo of p0) {
const checkbox = todo.status === 'in_progress' ? '[-]' : '[ ]';
lines.push(`- ${checkbox} ${todo.content}`);
}
lines.push('');
}
// Standard (P1)
if (p1.length > 0) {
lines.push('## Standard (P1)');
for (const todo of p1) {
const checkbox = todo.status === 'in_progress' ? '[-]' : '[ ]';
lines.push(`- ${checkbox} ${todo.content}`);
}
lines.push('');
}
// Nice to Have (P2)
if (p2.length > 0) {
lines.push('## Nice to Have (P2)');
for (const todo of p2) {
const checkbox = todo.status === 'in_progress' ? '[-]' : '[ ]';
lines.push(`- ${checkbox} ${todo.content}`);
}
lines.push('');
}
// Tasks (no priority)
if (noPriority.length > 0) {
lines.push('## Tasks');
for (const todo of noPriority) {
const checkbox = todo.status === 'in_progress' ? '[-]' : '[ ]';
lines.push(`- ${checkbox} ${todo.content}`);
}
lines.push('');
}
// Completed
if (completed.length > 0) {
lines.push('## Completed');
for (const todo of completed) {
lines.push(`- [x] ${todo.content}`);
}
lines.push('');
}
return lines.join('\n');
}
/**
* Parse @fix_plan.md content and return parsed todo items.
*
* @param content - Markdown content from @fix_plan.md
* @param parsePriority - Function to parse priority from content text
* @param generateTodoId - Function to generate stable todo ID from content
* @returns Array of parsed todo items
*/
export function importFixPlanMarkdown(
content: string,
parsePriority: (content: string) => RalphTodoPriority,
generateTodoId: (content: string) => string
): RalphTodoItem[] {
const lines = content.split('\n');
const newTodos: RalphTodoItem[] = [];
let currentPriority: RalphTodoPriority = null;
// Patterns for section headers
const p0HeaderPattern = /^##\s*(High Priority|Critical|P0)/i;
const p1HeaderPattern = /^##\s*(Standard|P1|Medium Priority)/i;
const p2HeaderPattern = /^##\s*(Nice to Have|P2|Low Priority)/i;
const completedHeaderPattern = /^##\s*Completed/i;
const tasksHeaderPattern = /^##\s*Tasks/i;
// Pattern for todo items
const todoPattern = /^-\s*\[([ x-])\]\s*(.+)$/;
let inCompletedSection = false;
for (const line of lines) {
const trimmed = line.trim();
// Check for section headers
if (p0HeaderPattern.test(trimmed)) {
currentPriority = 'P0';
inCompletedSection = false;
continue;
}
if (p1HeaderPattern.test(trimmed)) {
currentPriority = 'P1';
inCompletedSection = false;
continue;
}
if (p2HeaderPattern.test(trimmed)) {
currentPriority = 'P2';
inCompletedSection = false;
continue;
}
if (completedHeaderPattern.test(trimmed)) {
inCompletedSection = true;
continue;
}
if (tasksHeaderPattern.test(trimmed)) {
currentPriority = null;
inCompletedSection = false;
continue;
}
// Parse todo item
const match = trimmed.match(todoPattern);
if (match) {
const [, checkboxState, todoContent] = match;
let status: RalphTodoStatus;
if (inCompletedSection || checkboxState === 'x' || checkboxState === 'X') {
status = 'completed';
} else if (checkboxState === '-') {
status = 'in_progress';
} else {
status = 'pending';
}
// Parse priority from content if not in a priority section
const parsedPriority = inCompletedSection ? null : currentPriority || parsePriority(todoContent);
const id = generateTodoId(todoContent);
newTodos.push({
id,
content: todoContent.trim(),
status,
detectedAt: Date.now(),
priority: parsedPriority,
});
}
}
return newTodos;
}
/**
* RalphFixPlanWatcher - Watches @fix_plan.md for changes.
*
* Events emitted:
* - `todosLoaded` - Emits parsed todo items when @fix_plan.md is loaded/changed
* - `enabled` - Emits when tracker should be auto-enabled (todos loaded from file)
*/
export class RalphFixPlanWatcher extends EventEmitter {
/** Working directory for @fix_plan.md watching */
private _workingDir: string | null = null;
/** Path to the @fix_plan.md file being watched */
private _fixPlanPath: string | null = null;
/** File watcher for @fix_plan.md */
private _fixPlanWatcher: FSWatcher | null = null;
/** Error handler for FSWatcher (stored for cleanup to prevent memory leak) */
private _fixPlanWatcherErrorHandler: ((err: Error) => void) | null = null;
/** Debounce timer for file change events */
private _fixPlanReloadTimer: NodeJS.Timeout | null = null;
/** Priority parser injected from parent (for importFixPlanMarkdown) */
private _parsePriority: (content: string) => RalphTodoPriority;
/** Todo ID generator injected from parent */
private _generateTodoId: (content: string) => string;
constructor(parsePriority: (content: string) => RalphTodoPriority, generateTodoId: (content: string) => string) {
super();
this._parsePriority = parsePriority;
this._generateTodoId = generateTodoId;
}
/**
* When @fix_plan.md is active, treat it as the source of truth for todo status.
* This prevents output-based detection from overriding file-based status.
*/
get isFileAuthoritative(): boolean {
return this._fixPlanPath !== null;
}
/**
* Set the working directory and start watching @fix_plan.md.
* Automatically loads existing @fix_plan.md if present.
* @param workingDir - The session's working directory
*/
setWorkingDir(workingDir: string): void {
this._workingDir = workingDir;
this._fixPlanPath = join(workingDir, '@fix_plan.md');
// Try to load existing @fix_plan.md
this.loadFixPlanFromDisk();
// Start watching for changes
this.startWatchingFixPlan();
}
/**
* Load @fix_plan.md from disk if it exists.
* Called on initialization and when file changes are detected.
*/
async loadFixPlanFromDisk(): Promise<number> {
if (!this._fixPlanPath) return 0;
try {
if (!existsSync(this._fixPlanPath)) {
return 0;
}
const content = await readFile(this._fixPlanPath, 'utf-8');
const todos = importFixPlanMarkdown(content, this._parsePriority, this._generateTodoId);
if (todos.length > 0) {
this.emit('todosLoaded', todos);
console.log(`[RalphFixPlanWatcher] Loaded ${todos.length} todos from @fix_plan.md`);
}
return todos.length;
} catch (err) {
// File doesn't exist or can't be read - that's OK
console.log(`[RalphFixPlanWatcher] Could not load @fix_plan.md: ${err}`);
return 0;
}
}
/**
* Start watching @fix_plan.md for changes.
* Reloads todos when the file is modified.
*/
private startWatchingFixPlan(): void {
if (!this._fixPlanPath || !this._workingDir) return;
// Stop existing watcher if any
this.stopWatchingFixPlan();
try {
// Only watch if the file exists
if (!existsSync(this._fixPlanPath)) {
// Watch the directory instead for file creation
this._fixPlanWatcher = fsWatch(this._workingDir, (_eventType, filename) => {
if (filename === '@fix_plan.md') {
this.handleFixPlanChange();
}
});
} else {
// Watch the file directly
this._fixPlanWatcher = fsWatch(this._fixPlanPath, () => {
this.handleFixPlanChange();
});
}
// Add error handler to prevent unhandled errors and clean up on failure
// Store handler reference for proper cleanup in stopWatchingFixPlan()
if (this._fixPlanWatcher) {
this._fixPlanWatcherErrorHandler = (err: Error) => {
console.log(`[RalphFixPlanWatcher] FSWatcher error for @fix_plan.md: ${err.message}`);
this.stopWatchingFixPlan();
};
this._fixPlanWatcher.on('error', this._fixPlanWatcherErrorHandler);
}
} catch (err) {
console.log(`[RalphFixPlanWatcher] Could not watch @fix_plan.md: ${err}`);
}
}
/**
* Handle @fix_plan.md file change with debouncing.
*/
private handleFixPlanChange(): void {
// Debounce rapid changes (e.g., multiple writes)
if (this._fixPlanReloadTimer) {
clearTimeout(this._fixPlanReloadTimer);
}
this._fixPlanReloadTimer = setTimeout(() => {
this._fixPlanReloadTimer = null;
this.loadFixPlanFromDisk();
}, 500); // 500ms debounce
}
/**
* Stop watching @fix_plan.md.
*/
stopWatchingFixPlan(): void {
if (this._fixPlanWatcher) {
// Remove error handler before closing to prevent memory leak
if (this._fixPlanWatcherErrorHandler) {
this._fixPlanWatcher.off('error', this._fixPlanWatcherErrorHandler);
this._fixPlanWatcherErrorHandler = null;
}
this._fixPlanWatcher.close();
this._fixPlanWatcher = null;
}
if (this._fixPlanReloadTimer) {
clearTimeout(this._fixPlanReloadTimer);
this._fixPlanReloadTimer = null;
}
}
/**
* Stop watching and clean up all resources.
*/
stop(): void {
this.stopWatchingFixPlan();
}
/**
* Clean up all resources.
*/
destroy(): void {
this.stop();
this.removeAllListeners();
}
}
+14 -4
View File
@@ -1,14 +1,24 @@
/**
* @fileoverview Ralph Loop - Autonomous task execution engine
* @fileoverview Ralph Loop - Autonomous task execution engine.
*
* The Ralph Loop orchestrates autonomous Claude sessions by:
* Orchestrates autonomous Claude sessions by:
* - Polling for available tasks from the task queue
* - Assigning tasks to idle sessions
* - Monitoring completion and handling failures
* - Auto-generating follow-up tasks when min duration not reached
*
* Named after Ralph Wiggum's persistence ("I'm in danger!"),
* this loop keeps Claude working until all tasks are done.
* Key exports:
* - `RalphLoop` class — the loop engine, extends EventEmitter
* - `RalphLoopEvents` interface — typed event map
* - `RalphLoopOptions` interface — configuration options
*
* Lifecycle: `start()` → poll loop → `stop()` (when all tasks done + min duration met)
*
* @dependencies session-manager (session lifecycle), task-queue (task FIFO),
* state-store (persistence), session (PTY execution), task (task model)
* @consumedby web/server (ralph routes, SSE)
* @emits started, stopped, taskAssigned, taskCompleted, taskFailed, error
* @persistence Ralph loop state saved to `~/.codeman/state.json` (ralphLoop key)
*
* @module ralph-loop
*/
+477
View File
@@ -0,0 +1,477 @@
/**
* @fileoverview RalphPlanTracker - Enhanced plan task management
*
* Manages plan tasks with verification criteria, dependencies,
* execution tracking, TDD workflow support, and plan versioning.
*
* Extracted from ralph-tracker.ts as part of domain splitting.
*
* @module ralph-plan-tracker
*/
import { EventEmitter } from 'node:events';
import type { PlanTaskStatus, TddPhase } from './types.js';
// ========== Enhanced Plan Task Interface ==========
/**
* Enhanced plan task with verification criteria, dependencies, and execution tracking.
* Supports TDD workflow, failure tracking, and plan versioning.
*/
export interface EnhancedPlanTask {
/** Unique identifier (e.g., "P0-001") */
id: string;
/** Task description */
content: string;
/** Criticality level */
priority: 'P0' | 'P1' | 'P2' | null;
/** How to verify completion */
verificationCriteria?: string;
/** Command to run for verification */
testCommand?: string;
/** IDs of tasks that must complete first */
dependencies: string[];
/** Current execution status */
status: PlanTaskStatus;
/** How many times attempted */
attempts: number;
/** Most recent failure reason */
lastError?: string;
/** Timestamp of completion */
completedAt?: number;
/** Plan version this belongs to */
version: number;
/** TDD phase category */
tddPhase?: TddPhase;
/** ID of paired test/impl task */
pairedWith?: string;
/** Estimated complexity */
complexity?: 'low' | 'medium' | 'high';
/** Checklist items for review tasks (tddPhase: 'review') */
reviewChecklist?: string[];
}
/** Checkpoint review data */
export interface CheckpointReview {
iteration: number;
timestamp: number;
summary: {
total: number;
completed: number;
failed: number;
blocked: number;
pending: number;
inProgress: number;
};
stuckTasks: Array<{
id: string;
content: string;
attempts: number;
lastError?: string;
}>;
recommendations: string[];
}
const MAX_PLAN_HISTORY = 10;
/**
* RalphPlanTracker - Manages enhanced plan tasks with versioning and checkpoints.
*
* Events emitted:
* - `planInitialized` - When a new plan is initialized
* - `planTaskUpdate` - When a plan task is updated
* - `taskBlocked` - When a task becomes blocked after too many failures
* - `taskUnblocked` - When a task's dependencies are all met
* - `planCheckpoint` - When a checkpoint review is triggered
* - `planTaskAdded` - When a new task is added to the plan
* - `planRollback` - When the plan is rolled back to a previous version
*/
export class RalphPlanTracker extends EventEmitter {
/** Current version of the plan (incremented on changes) */
private _planVersion: number = 1;
/** History of plan versions for rollback support */
private _planHistory: Array<{
version: number;
timestamp: number;
tasks: Map<string, EnhancedPlanTask>;
summary: string;
}> = [];
/** Enhanced plan tasks with execution tracking */
private _planTasks: Map<string, EnhancedPlanTask> = new Map();
/** Checkpoint intervals (iterations at which to trigger review) */
private _checkpointIterations: number[] = [5, 10, 20, 30, 50, 75, 100];
/** Last checkpoint iteration */
private _lastCheckpointIteration: number = 0;
/** Current cycle count (fed by parent via notifyCycleCount) */
private _cycleCount: number = 0;
constructor() {
super();
}
/**
* Notify the plan tracker of the current cycle count.
* Called by parent when iteration changes (for checkpoint detection).
*/
notifyCycleCount(cycleCount: number): void {
this._cycleCount = cycleCount;
}
/**
* Initialize plan tasks from generated plan items.
* Called when wizard generates a new plan.
*/
initializePlanTasks(
items: Array<{
id?: string;
content: string;
priority?: 'P0' | 'P1' | 'P2' | null;
verificationCriteria?: string;
testCommand?: string;
dependencies?: string[];
tddPhase?: TddPhase;
pairedWith?: string;
complexity?: 'low' | 'medium' | 'high';
}>
): void {
// Save current plan to history before replacing
if (this._planTasks.size > 0) {
this._savePlanToHistory('Plan replaced with new generation');
}
// Clear and rebuild
this._planTasks.clear();
this._planVersion++;
items.forEach((item, idx) => {
const id = item.id || `task-${idx}`;
const task: EnhancedPlanTask = {
id,
content: item.content,
priority: item.priority || null,
verificationCriteria: item.verificationCriteria,
testCommand: item.testCommand,
dependencies: item.dependencies || [],
status: 'pending',
attempts: 0,
version: this._planVersion,
tddPhase: item.tddPhase,
pairedWith: item.pairedWith,
complexity: item.complexity,
};
this._planTasks.set(id, task);
});
this.emit('planInitialized', { version: this._planVersion, taskCount: this._planTasks.size });
}
/**
* Update a specific plan task's status, attempts, or error.
*/
updatePlanTask(
taskId: string,
update: {
status?: PlanTaskStatus;
error?: string;
incrementAttempts?: boolean;
}
): { success: boolean; task?: EnhancedPlanTask; error?: string } {
const task = this._planTasks.get(taskId);
if (!task) {
return { success: false, error: 'Task not found' };
}
if (update.status) {
task.status = update.status;
if (update.status === 'completed') {
task.completedAt = Date.now();
}
}
if (update.error) {
task.lastError = update.error;
}
if (update.incrementAttempts) {
task.attempts++;
// After 3 failed attempts, mark as blocked and emit warning
if (task.attempts >= 3 && task.status === 'failed') {
task.status = 'blocked';
this.emit('taskBlocked', {
taskId,
content: task.content,
attempts: task.attempts,
lastError: task.lastError,
});
}
}
// Update blocked tasks when a dependency completes
if (update.status === 'completed') {
this._unblockDependentTasks(taskId);
}
// Check for checkpoint
this._checkForCheckpoint();
this.emit('planTaskUpdate', { taskId, task });
return { success: true, task };
}
/**
* Add a new task to the plan (for runtime adaptation).
*/
addPlanTask(task: {
content: string;
priority?: 'P0' | 'P1' | 'P2';
verificationCriteria?: string;
dependencies?: string[];
insertAfter?: string;
}): { task: EnhancedPlanTask } {
// Generate unique ID
const existingIds = Array.from(this._planTasks.keys());
const prefix = task.priority || 'P1';
let counter = existingIds.filter((id) => id.startsWith(prefix)).length + 1;
let id = `${prefix}-${String(counter).padStart(3, '0')}`;
while (this._planTasks.has(id)) {
counter++;
id = `${prefix}-${String(counter).padStart(3, '0')}`;
}
const newTask: EnhancedPlanTask = {
id,
content: task.content,
priority: task.priority || null,
verificationCriteria: task.verificationCriteria || 'Task completed successfully',
dependencies: task.dependencies || [],
status: 'pending',
attempts: 0,
version: this._planVersion,
};
this._planTasks.set(id, newTask);
this.emit('planTaskAdded', { task: newTask });
return { task: newTask };
}
/**
* Get all plan tasks.
*/
getPlanTasks(): EnhancedPlanTask[] {
return Array.from(this._planTasks.values());
}
/**
* Generate a checkpoint review summarizing plan progress and stuck tasks.
*/
generateCheckpointReview(): CheckpointReview {
const tasks = Array.from(this._planTasks.values());
const summary = {
total: tasks.length,
completed: tasks.filter((t) => t.status === 'completed').length,
failed: tasks.filter((t) => t.status === 'failed').length,
blocked: tasks.filter((t) => t.status === 'blocked').length,
pending: tasks.filter((t) => t.status === 'pending').length,
inProgress: tasks.filter((t) => t.status === 'in_progress').length,
};
// Find stuck tasks (3+ attempts or blocked)
const stuckTasks = tasks
.filter((t) => t.attempts >= 3 || t.status === 'blocked')
.map((t) => ({
id: t.id,
content: t.content,
attempts: t.attempts,
lastError: t.lastError,
}));
// Generate recommendations
const recommendations: string[] = [];
if (stuckTasks.length > 0) {
recommendations.push(`${stuckTasks.length} task(s) are stuck. Consider breaking them into smaller steps.`);
}
if (summary.failed > summary.completed && summary.total > 5) {
recommendations.push('More tasks have failed than completed. Review approach and consider plan adjustment.');
}
const progressPercent = summary.total > 0 ? Math.round((summary.completed / summary.total) * 100) : 0;
if (progressPercent < 20 && this._cycleCount > 10) {
recommendations.push('Progress is slow. Consider simplifying tasks or reviewing dependencies.');
}
if (summary.total > 0 && summary.blocked > summary.total / 3) {
recommendations.push('Many tasks are blocked. Review dependency chain for bottlenecks.');
}
return {
iteration: this._cycleCount,
timestamp: Date.now(),
summary,
stuckTasks,
recommendations,
};
}
/**
* Get plan version history.
*/
getPlanHistory(): Array<{
version: number;
timestamp: number;
summary: string;
stats: { total: number; completed: number; failed: number };
}> {
return this._planHistory.map((h) => {
const tasks = Array.from(h.tasks.values());
return {
version: h.version,
timestamp: h.timestamp,
summary: h.summary,
stats: {
total: tasks.length,
completed: tasks.filter((t) => t.status === 'completed').length,
failed: tasks.filter((t) => t.status === 'failed').length,
},
};
});
}
/**
* Rollback to a previous plan version.
*/
rollbackToVersion(version: number): {
success: boolean;
plan?: EnhancedPlanTask[];
error?: string;
} {
const historyEntry = this._planHistory.find((h) => h.version === version);
if (!historyEntry) {
return { success: false, error: `Version ${version} not found in history` };
}
// Save current state first
this._savePlanToHistory(`Rolled back from v${this._planVersion} to v${version}`);
// Restore the historical version
this._planTasks.clear();
for (const [id, task] of historyEntry.tasks) {
// Reset execution state for retry
this._planTasks.set(id, {
...task,
status: task.status === 'completed' ? 'completed' : 'pending',
attempts: task.status === 'completed' ? task.attempts : 0,
lastError: undefined,
});
}
this._planVersion++;
this.emit('planRollback', { version, newVersion: this._planVersion });
return { success: true, plan: Array.from(this._planTasks.values()) };
}
/**
* Check if checkpoint review is due for current iteration.
*/
isCheckpointDue(): boolean {
return this._checkpointIterations.includes(this._cycleCount) && this._cycleCount > this._lastCheckpointIteration;
}
/**
* Get current plan version.
*/
get planVersion(): number {
return this._planVersion;
}
/**
* Reset plan state (soft reset - keeps version history).
*/
reset(): void {
// Don't clear history or version on soft reset
}
/**
* Full reset - clears all plan state.
*/
fullReset(): void {
this._planTasks.clear();
this._planHistory.length = 0;
this._planVersion = 1;
this._lastCheckpointIteration = 0;
this._cycleCount = 0;
}
/**
* Clean up all resources.
*/
destroy(): void {
this._planTasks.clear();
this._planHistory.length = 0;
this.removeAllListeners();
}
/**
* Unblock tasks that were waiting on a completed dependency.
*/
private _unblockDependentTasks(completedTaskId: string): void {
for (const [_, task] of this._planTasks) {
if (task.dependencies.includes(completedTaskId)) {
// Check if all dependencies are now complete
const allDepsComplete = task.dependencies.every((depId) => {
const dep = this._planTasks.get(depId);
return dep && dep.status === 'completed';
});
if (allDepsComplete && task.status === 'blocked') {
task.status = 'pending';
this.emit('taskUnblocked', { taskId: task.id });
}
}
}
}
/**
* Check if current iteration is a checkpoint and emit review if so.
*/
private _checkForCheckpoint(): void {
if (this._checkpointIterations.includes(this._cycleCount) && this._cycleCount > this._lastCheckpointIteration) {
this._lastCheckpointIteration = this._cycleCount;
const checkpoint = this.generateCheckpointReview();
this.emit('planCheckpoint', checkpoint);
}
}
/**
* Save current plan state to history.
*/
private _savePlanToHistory(summary: string): void {
// Clone current tasks
const tasksCopy = new Map<string, EnhancedPlanTask>();
for (const [id, task] of this._planTasks) {
tasksCopy.set(id, { ...task });
}
this._planHistory.push({
version: this._planVersion,
timestamp: Date.now(),
tasks: tasksCopy,
summary,
});
// Limit history size
if (this._planHistory.length > MAX_PLAN_HISTORY) {
this._planHistory.shift();
}
}
}
+167
View File
@@ -0,0 +1,167 @@
/**
* @fileoverview RalphStallDetector - Iteration stall detection
*
* Monitors iteration progress and emits warnings when the loop
* appears to be stalled (no iteration changes for extended periods).
*
* Extracted from ralph-tracker.ts as part of domain splitting.
*
* @module ralph-stall-detector
*/
import { EventEmitter } from 'node:events';
import { CLEANUP_CHECK_INTERVAL_MS } from './config/server-timing.js';
/**
* RalphStallDetector - Detects iteration stalls in the Ralph loop.
*
* Events emitted:
* - `iterationStallWarning` - When iteration hasn't changed for warning threshold
* - `iterationStallCritical` - When iteration hasn't changed for critical threshold
*/
export class RalphStallDetector extends EventEmitter {
/** Timestamp when iteration count last changed */
private _lastIterationChangeTime: number = 0;
/** Last observed iteration count for stall detection */
private _lastObservedIteration: number = 0;
/** Timer for iteration stall detection */
private _iterationStallTimer: NodeJS.Timeout | null = null;
/** Iteration stall warning threshold (ms) - default 10 minutes */
private _iterationStallWarningMs: number = 10 * 60 * 1000;
/** Iteration stall critical threshold (ms) - default 20 minutes */
private _iterationStallCriticalMs: number = 20 * 60 * 1000;
/** Whether stall warning has been emitted */
private _iterationStallWarned: boolean = false;
/** Whether the loop is currently active */
private _loopActive: boolean = false;
constructor() {
super();
this._lastIterationChangeTime = Date.now();
}
/**
* Start iteration stall detection timer.
* Should be called when the loop becomes active.
*/
startIterationStallDetection(): void {
this.stopIterationStallDetection();
this._lastIterationChangeTime = Date.now();
this._iterationStallWarned = false;
// Check every minute
this._iterationStallTimer = setInterval(() => {
this.checkIterationStall();
}, CLEANUP_CHECK_INTERVAL_MS);
}
/**
* Stop iteration stall detection timer.
*/
stopIterationStallDetection(): void {
if (this._iterationStallTimer) {
clearInterval(this._iterationStallTimer);
this._iterationStallTimer = null;
}
}
/**
* Notify the detector that the iteration has changed.
* Resets stall tracking state.
*/
notifyIterationChanged(iteration: number): void {
this._lastIterationChangeTime = Date.now();
this._lastObservedIteration = iteration;
this._iterationStallWarned = false;
}
/**
* Set whether the loop is currently active.
* Stall detection only fires when loop is active.
*/
setLoopActive(active: boolean): void {
this._loopActive = active;
}
/**
* Check for iteration stall and emit appropriate events.
*/
private checkIterationStall(): void {
if (!this._loopActive) return;
const stallDurationMs = Date.now() - this._lastIterationChangeTime;
// Critical stall (longer duration)
if (stallDurationMs >= this._iterationStallCriticalMs) {
this.emit('iterationStallCritical', {
iteration: this._lastObservedIteration,
stallDurationMs,
});
return;
}
// Warning stall
if (stallDurationMs >= this._iterationStallWarningMs && !this._iterationStallWarned) {
this._iterationStallWarned = true;
this.emit('iterationStallWarning', {
iteration: this._lastObservedIteration,
stallDurationMs,
});
}
}
/**
* Get iteration stall metrics for monitoring.
*/
getIterationStallMetrics(): {
lastIterationChangeTime: number;
stallDurationMs: number;
warningThresholdMs: number;
criticalThresholdMs: number;
isWarned: boolean;
currentIteration: number;
} {
return {
lastIterationChangeTime: this._lastIterationChangeTime,
stallDurationMs: Date.now() - this._lastIterationChangeTime,
warningThresholdMs: this._iterationStallWarningMs,
criticalThresholdMs: this._iterationStallCriticalMs,
isWarned: this._iterationStallWarned,
currentIteration: this._lastObservedIteration,
};
}
/**
* Configure iteration stall thresholds.
* @param warningMs - Warning threshold in milliseconds
* @param criticalMs - Critical threshold in milliseconds
*/
configureIterationStallThresholds(warningMs: number, criticalMs: number): void {
this._iterationStallWarningMs = warningMs;
this._iterationStallCriticalMs = criticalMs;
}
/**
* Reset stall detector state.
*/
reset(): void {
this._lastIterationChangeTime = Date.now();
this._lastObservedIteration = 0;
this._iterationStallWarned = false;
this._loopActive = false;
}
/**
* Clean up all resources.
*/
destroy(): void {
this.stopIterationStallDetection();
this.removeAllListeners();
}
}
+552
View File
@@ -0,0 +1,552 @@
/**
* @fileoverview RalphStatusParser - RALPH_STATUS block parsing and circuit breaker
*
* Parses structured RALPH_STATUS blocks from Claude Code output
* and manages the circuit breaker state machine.
*
* Extracted from ralph-tracker.ts as part of domain splitting.
*
* @module ralph-status-parser
*/
import { EventEmitter } from 'node:events';
import type {
RalphStatusBlock,
RalphStatusValue,
RalphTestsStatus,
RalphWorkType,
CircuitBreakerStatus,
} from './types.js';
import { createInitialCircuitBreakerStatus } from './types.js';
// ---------- RALPH_STATUS Block Patterns ----------
// Based on Ralph Claude Code structured status reporting
/**
* Matches the start of a RALPH_STATUS block
* Pattern: ---RALPH_STATUS---
*/
const RALPH_STATUS_START_PATTERN = /^---RALPH_STATUS---\s*$/;
/**
* Matches the end of a RALPH_STATUS block
* Pattern: ---END_RALPH_STATUS---
*/
const RALPH_STATUS_END_PATTERN = /^---END_RALPH_STATUS---\s*$/;
/**
* Matches STATUS field in RALPH_STATUS block
* Captures: IN_PROGRESS | COMPLETE | BLOCKED
*/
const RALPH_STATUS_FIELD_PATTERN = /^STATUS:\s*(IN_PROGRESS|COMPLETE|BLOCKED)\s*$/i;
/**
* Matches TASKS_COMPLETED_THIS_LOOP field
* Captures: number
*/
const RALPH_TASKS_COMPLETED_PATTERN = /^TASKS_COMPLETED_THIS_LOOP:\s*(\d+)\s*$/i;
/**
* Matches FILES_MODIFIED field
* Captures: number
*/
const RALPH_FILES_MODIFIED_PATTERN = /^FILES_MODIFIED:\s*(\d+)\s*$/i;
/**
* Matches TESTS_STATUS field
* Captures: PASSING | FAILING | NOT_RUN
*/
const RALPH_TESTS_STATUS_PATTERN = /^TESTS_STATUS:\s*(PASSING|FAILING|NOT_RUN)\s*$/i;
/**
* Matches WORK_TYPE field
* Captures: IMPLEMENTATION | TESTING | DOCUMENTATION | REFACTORING
*/
const RALPH_WORK_TYPE_PATTERN = /^WORK_TYPE:\s*(IMPLEMENTATION|TESTING|DOCUMENTATION|REFACTORING)\s*$/i;
/**
* Matches EXIT_SIGNAL field
* Captures: true | false
*/
const RALPH_EXIT_SIGNAL_PATTERN = /^EXIT_SIGNAL:\s*(true|false)\s*$/i;
/**
* Matches RECOMMENDATION field
* Captures: any text
*/
const RALPH_RECOMMENDATION_PATTERN = /^RECOMMENDATION:\s*(.+)$/i;
// ---------- Completion Indicator Patterns (for dual-condition exit) ----------
/**
* Patterns that indicate potential completion (natural language)
* Count >= 2 along with EXIT_SIGNAL: true triggers exit
*/
const COMPLETION_INDICATOR_PATTERNS = [
/all\s+(?:tasks?|items?|work)\s+(?:are\s+)?(?:completed?|done|finished)/i,
/(?:completed?|finished)\s+all\s+(?:tasks?|items?|work)/i,
/nothing\s+(?:left|remaining)\s+to\s+do/i,
/no\s+more\s+(?:tasks?|items?|work)/i,
/everything\s+(?:is\s+)?(?:completed?|done)/i,
/project\s+(?:is\s+)?(?:completed?|done|finished)/i,
];
/**
* RalphStatusParser - Parses RALPH_STATUS blocks and manages circuit breaker.
*
* Events emitted:
* - `statusBlockDetected` - When a complete RALPH_STATUS block is parsed
* - `circuitBreakerUpdate` - When circuit breaker state changes
* - `exitGateMet` - When dual-condition exit gate is met
*/
export class RalphStatusParser extends EventEmitter {
/** Circuit breaker state tracking */
private _circuitBreaker: CircuitBreakerStatus;
/** Buffer for RALPH_STATUS block lines */
private _statusBlockBuffer: string[] = [];
/** Flag indicating we're inside a RALPH_STATUS block */
private _inStatusBlock: boolean = false;
/** Last parsed RALPH_STATUS block */
private _lastStatusBlock: RalphStatusBlock | null = null;
/** Count of completion indicators detected (for dual-condition exit) */
private _completionIndicators: number = 0;
/** Whether dual-condition exit gate has been met */
private _exitGateMet: boolean = false;
/** Cumulative files modified across all iterations */
private _totalFilesModified: number = 0;
/** Cumulative tasks completed across all iterations */
private _totalTasksCompleted: number = 0;
/** Current cycle count (fed by parent) */
private _cycleCount: number = 0;
constructor() {
super();
this._circuitBreaker = createInitialCircuitBreakerStatus();
}
/**
* Process a line for status block detection and completion indicators.
* Main entry point - call this for each trimmed line.
*/
processLine(line: string): void {
this.processStatusBlockLine(line);
this.detectCompletionIndicators(line);
}
/**
* Set the current cycle count (fed by parent for circuit breaker tracking).
*/
setCycleCount(cycleCount: number): void {
this._cycleCount = cycleCount;
}
/**
* Notify of iteration progress (for circuit breaker reset on progress).
* Called by parent when iteration count changes.
*/
notifyIterationProgress(currentIteration: number): void {
if (
this._circuitBreaker.state === 'HALF_OPEN' ||
this._circuitBreaker.consecutiveNoProgress > 0 ||
this._circuitBreaker.consecutiveSameError > 0 ||
this._circuitBreaker.consecutiveTestsFailure > 0
) {
this._circuitBreaker.consecutiveNoProgress = 0;
this._circuitBreaker.consecutiveSameError = 0;
this._circuitBreaker.lastProgressIteration = currentIteration;
if (this._circuitBreaker.state === 'HALF_OPEN') {
this._circuitBreaker.state = 'CLOSED';
this._circuitBreaker.reason = 'Iteration progress detected';
this._circuitBreaker.reasonCode = 'progress_detected';
this.emit('circuitBreakerUpdate', { ...this._circuitBreaker });
}
}
}
/**
* Get current circuit breaker status.
*/
get circuitBreakerStatus(): CircuitBreakerStatus {
return { ...this._circuitBreaker };
}
/**
* Get last parsed RALPH_STATUS block.
*/
get lastStatusBlock(): RalphStatusBlock | null {
return this._lastStatusBlock ? { ...this._lastStatusBlock } : null;
}
/**
* Get cumulative stats from status blocks.
*/
get cumulativeStats(): {
filesModified: number;
tasksCompleted: number;
completionIndicators: number;
} {
return {
filesModified: this._totalFilesModified,
tasksCompleted: this._totalTasksCompleted,
completionIndicators: this._completionIndicators,
};
}
/**
* Whether dual-condition exit gate has been met.
*/
get exitGateMet(): boolean {
return this._exitGateMet;
}
/**
* Manually reset circuit breaker to CLOSED state.
* Use when user acknowledges the issue is resolved.
*
* @fires circuitBreakerUpdate
*/
resetCircuitBreaker(): void {
this._circuitBreaker = createInitialCircuitBreakerStatus();
this._circuitBreaker.reason = 'Manual reset';
this._circuitBreaker.reasonCode = 'manual_reset';
this.emit('circuitBreakerUpdate', { ...this._circuitBreaker });
}
/**
* Reset status parser state (soft reset).
* Clears status block buffer and completion indicators.
* Keeps circuit breaker state (it tracks across iterations).
*/
reset(): void {
this._statusBlockBuffer = [];
this._inStatusBlock = false;
this._lastStatusBlock = null;
this._completionIndicators = 0;
this._exitGateMet = false;
this._totalFilesModified = 0;
this._totalTasksCompleted = 0;
// Keep circuit breaker state on soft reset (it tracks across iterations)
}
/**
* Full reset - clears all state including circuit breaker.
*/
fullReset(): void {
this.reset();
this._circuitBreaker = createInitialCircuitBreakerStatus();
}
/**
* Clean up all resources.
*/
destroy(): void {
this._statusBlockBuffer.length = 0;
this.removeAllListeners();
}
// ========== Private Methods ==========
/**
* Process a line for RALPH_STATUS block detection.
* Buffers lines between ---RALPH_STATUS--- and ---END_RALPH_STATUS---
* then parses the complete block.
*
* @param line - Single line to process (already trimmed)
* @fires statusBlockDetected - When a complete block is parsed
*/
private processStatusBlockLine(line: string): void {
// Check for block start
if (RALPH_STATUS_START_PATTERN.test(line)) {
this._inStatusBlock = true;
this._statusBlockBuffer = [];
return;
}
// Check for block end
if (this._inStatusBlock && RALPH_STATUS_END_PATTERN.test(line)) {
this._inStatusBlock = false;
this.parseStatusBlock(this._statusBlockBuffer);
this._statusBlockBuffer = [];
return;
}
// Buffer lines while in block
if (this._inStatusBlock) {
this._statusBlockBuffer.push(line);
}
}
/**
* Parse buffered RALPH_STATUS block lines into structured data.
*
* P1-004: Enhanced with schema validation and error recovery
*
* @param lines - Array of lines between block markers
* @fires statusBlockDetected - When parsing succeeds
*/
private parseStatusBlock(lines: string[]): void {
const block: Partial<RalphStatusBlock> = {
parsedAt: Date.now(),
};
const parseErrors: string[] = [];
const unknownFields: string[] = [];
for (const line of lines) {
const trimmedLine = line.trim();
if (!trimmedLine) continue;
// Track whether this line matched any known field
let matched = false;
// STATUS field (required)
const statusMatch = trimmedLine.match(RALPH_STATUS_FIELD_PATTERN);
if (statusMatch) {
const value = statusMatch[1].toUpperCase();
if (['IN_PROGRESS', 'COMPLETE', 'BLOCKED'].includes(value)) {
block.status = value as RalphStatusValue;
} else {
parseErrors.push(`Invalid STATUS value: "${value}". Expected: IN_PROGRESS, COMPLETE, or BLOCKED`);
}
matched = true;
}
// TASKS_COMPLETED_THIS_LOOP field
const tasksMatch = trimmedLine.match(RALPH_TASKS_COMPLETED_PATTERN);
if (tasksMatch) {
const value = parseInt(tasksMatch[1], 10);
if (!Number.isNaN(value) && value >= 0) {
block.tasksCompletedThisLoop = value;
} else {
parseErrors.push(
`Invalid TASKS_COMPLETED_THIS_LOOP value: "${tasksMatch[1]}". Expected: non-negative integer`
);
}
matched = true;
}
// FILES_MODIFIED field
const filesMatch = trimmedLine.match(RALPH_FILES_MODIFIED_PATTERN);
if (filesMatch) {
const value = parseInt(filesMatch[1], 10);
if (!Number.isNaN(value) && value >= 0) {
block.filesModified = value;
} else {
parseErrors.push(`Invalid FILES_MODIFIED value: "${filesMatch[1]}". Expected: non-negative integer`);
}
matched = true;
}
// TESTS_STATUS field
const testsMatch = trimmedLine.match(RALPH_TESTS_STATUS_PATTERN);
if (testsMatch) {
const value = testsMatch[1].toUpperCase();
if (['PASSING', 'FAILING', 'NOT_RUN'].includes(value)) {
block.testsStatus = value as RalphTestsStatus;
} else {
parseErrors.push(`Invalid TESTS_STATUS value: "${value}". Expected: PASSING, FAILING, or NOT_RUN`);
}
matched = true;
}
// WORK_TYPE field
const workMatch = trimmedLine.match(RALPH_WORK_TYPE_PATTERN);
if (workMatch) {
const value = workMatch[1].toUpperCase();
if (['IMPLEMENTATION', 'TESTING', 'DOCUMENTATION', 'REFACTORING'].includes(value)) {
block.workType = value as RalphWorkType;
} else {
parseErrors.push(
`Invalid WORK_TYPE value: "${value}". Expected: IMPLEMENTATION, TESTING, DOCUMENTATION, or REFACTORING`
);
}
matched = true;
}
// EXIT_SIGNAL field
const exitMatch = trimmedLine.match(RALPH_EXIT_SIGNAL_PATTERN);
if (exitMatch) {
block.exitSignal = exitMatch[1].toLowerCase() === 'true';
matched = true;
}
// RECOMMENDATION field
const recMatch = trimmedLine.match(RALPH_RECOMMENDATION_PATTERN);
if (recMatch) {
block.recommendation = recMatch[1].trim();
matched = true;
}
// Track unknown fields for debugging (only if looks like a field)
if (!matched && trimmedLine.includes(':')) {
const fieldName = trimmedLine.split(':')[0].trim().toUpperCase();
if (fieldName && !['#', '//'].some((c) => fieldName.startsWith(c))) {
unknownFields.push(fieldName);
}
}
}
// Log parse errors if any
if (parseErrors.length > 0) {
console.warn(`[RalphStatusParser] RALPH_STATUS parse errors:\n - ${parseErrors.join('\n - ')}`);
}
// Log unknown fields if any
if (unknownFields.length > 0) {
console.warn(`[RalphStatusParser] RALPH_STATUS unknown fields: ${unknownFields.join(', ')}`);
}
// Validate required field: STATUS
if (block.status === undefined) {
console.warn('[RalphStatusParser] RALPH_STATUS block missing required STATUS field, skipping');
return;
}
// Fill in defaults for missing optional fields
const fullBlock: RalphStatusBlock = {
status: block.status,
tasksCompletedThisLoop: block.tasksCompletedThisLoop ?? 0,
filesModified: block.filesModified ?? 0,
testsStatus: block.testsStatus ?? 'NOT_RUN',
workType: block.workType ?? 'IMPLEMENTATION',
exitSignal: block.exitSignal ?? false,
recommendation: block.recommendation ?? '',
parsedAt: block.parsedAt!,
};
this._lastStatusBlock = fullBlock;
this.handleStatusBlock(fullBlock);
}
/**
* Handle a parsed RALPH_STATUS block.
* Updates circuit breaker, checks exit conditions.
*
* @param block - Parsed status block
* @fires statusBlockDetected - With the block data
* @fires circuitBreakerUpdate - If state changes
* @fires exitGateMet - If dual-condition exit triggered
*/
private handleStatusBlock(block: RalphStatusBlock): void {
// Update cumulative counts
this._totalFilesModified += block.filesModified;
this._totalTasksCompleted += block.tasksCompletedThisLoop;
// Check for progress (for circuit breaker)
const hasProgress = block.filesModified > 0 || block.tasksCompletedThisLoop > 0;
// Update circuit breaker
this.updateCircuitBreaker(hasProgress, block.testsStatus, block.status);
// Check completion indicators
if (block.status === 'COMPLETE') {
this._completionIndicators++;
}
// Check dual-condition exit gate
if (block.exitSignal && this._completionIndicators >= 2 && !this._exitGateMet) {
this._exitGateMet = true;
this.emit('exitGateMet', {
completionIndicators: this._completionIndicators,
exitSignal: true,
});
}
// Emit the status block
this.emit('statusBlockDetected', block);
}
/**
* Update circuit breaker state based on iteration results.
*
* @param hasProgress - Whether this iteration made progress
* @param testsStatus - Current test status
* @param status - Overall status from RALPH_STATUS
* @fires circuitBreakerUpdate - If state changes
*/
private updateCircuitBreaker(hasProgress: boolean, testsStatus: RalphTestsStatus, status: RalphStatusValue): void {
const prevState = this._circuitBreaker.state;
if (hasProgress) {
// Progress detected - reset counters, possibly close circuit
this._circuitBreaker.consecutiveNoProgress = 0;
this._circuitBreaker.consecutiveSameError = 0;
this._circuitBreaker.lastProgressIteration = this._cycleCount;
if (this._circuitBreaker.state === 'HALF_OPEN') {
this._circuitBreaker.state = 'CLOSED';
this._circuitBreaker.reason = 'Progress detected, circuit closed';
this._circuitBreaker.reasonCode = 'progress_detected';
}
} else {
// No progress
this._circuitBreaker.consecutiveNoProgress++;
// State transitions based on consecutive no-progress
if (this._circuitBreaker.state === 'CLOSED') {
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = `No progress for ${this._circuitBreaker.consecutiveNoProgress} iterations`;
this._circuitBreaker.reasonCode = 'no_progress_open';
} else if (this._circuitBreaker.consecutiveNoProgress >= 2) {
this._circuitBreaker.state = 'HALF_OPEN';
this._circuitBreaker.reason = 'Warning: no progress detected';
this._circuitBreaker.reasonCode = 'no_progress_warning';
}
} else if (this._circuitBreaker.state === 'HALF_OPEN') {
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = `No progress for ${this._circuitBreaker.consecutiveNoProgress} iterations`;
this._circuitBreaker.reasonCode = 'no_progress_open';
}
}
}
// Track tests failure
if (testsStatus === 'FAILING') {
this._circuitBreaker.consecutiveTestsFailure++;
if (this._circuitBreaker.consecutiveTestsFailure >= 5 && this._circuitBreaker.state !== 'OPEN') {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = `Tests failing for ${this._circuitBreaker.consecutiveTestsFailure} iterations`;
this._circuitBreaker.reasonCode = 'tests_failing_too_long';
}
} else {
this._circuitBreaker.consecutiveTestsFailure = 0;
}
// Track blocked status
if (status === 'BLOCKED' && this._circuitBreaker.state !== 'OPEN') {
this._circuitBreaker.state = 'OPEN';
this._circuitBreaker.reason = 'Claude reported BLOCKED status';
this._circuitBreaker.reasonCode = 'same_error_repeated';
}
// Emit if state changed
if (prevState !== this._circuitBreaker.state) {
this._circuitBreaker.lastTransitionAt = Date.now();
this.emit('circuitBreakerUpdate', { ...this._circuitBreaker });
}
}
/**
* Check line for completion indicators (natural language patterns).
* Used for dual-condition exit gate.
*
* @param line - Line to check
*/
private detectCompletionIndicators(line: string): void {
for (const pattern of COMPLETION_INDICATOR_PATTERNS) {
if (pattern.test(line)) {
this._completionIndicators++;
break; // Only count once per line
}
}
}
}
+520 -2043
View File
File diff suppressed because it is too large Load Diff
+129
View File
@@ -0,0 +1,129 @@
/**
* @fileoverview Adaptive timing controller for respawn idle detection.
*
* Extracted from respawn-controller.ts for modularity. Tracks historical timing
* data and adjusts the completion confirm timeout dynamically based on the 75th
* percentile of recent idle detection durations.
*
* @module respawn-adaptive-timing
*/
import type { TimingHistory } from './types.js';
/**
* Configuration for adaptive timing bounds.
*/
export interface AdaptiveTimingConfig {
/** Minimum adaptive completion confirm timeout (ms) */
adaptiveMinConfirmMs: number;
/** Maximum adaptive completion confirm timeout (ms) */
adaptiveMaxConfirmMs: number;
}
/**
* Manages adaptive timing for respawn idle detection.
*
* Uses historical idle detection durations to calculate an optimal completion
* confirm timeout. The timeout is based on the 75th percentile of recent
* durations with a 20% safety buffer, clamped to configured bounds.
*/
export class RespawnAdaptiveTiming {
private timingHistory: TimingHistory;
constructor(private config: AdaptiveTimingConfig) {
this.timingHistory = {
recentIdleDetectionMs: [],
recentCycleDurationMs: [],
adaptiveCompletionConfirmMs: 10000, // Start with default
sampleCount: 0,
maxSamples: 20, // Keep last 20 samples for rolling average
lastUpdatedAt: Date.now(),
};
}
/**
* Record timing data from a completed cycle for adaptive adjustments.
*
* @param idleDetectionMs - Time spent detecting idle
* @param cycleDurationMs - Total cycle duration
*/
recordTimingData(idleDetectionMs: number, cycleDurationMs: number): void {
const history = this.timingHistory;
// Add to rolling windows
history.recentIdleDetectionMs.push(idleDetectionMs);
history.recentCycleDurationMs.push(cycleDurationMs);
// Trim to max samples
if (history.recentIdleDetectionMs.length > history.maxSamples) {
history.recentIdleDetectionMs.shift();
}
if (history.recentCycleDurationMs.length > history.maxSamples) {
history.recentCycleDurationMs.shift();
}
history.sampleCount = history.recentIdleDetectionMs.length;
history.lastUpdatedAt = Date.now();
// Recalculate adaptive timing
this.updateAdaptiveTiming();
}
/**
* Get the current adaptive completion confirm timeout.
* Returns the calculated value, or the default if not enough samples.
*
* @returns Completion confirm timeout in milliseconds
*/
getAdaptiveCompletionConfirmMs(): number {
return this.timingHistory.adaptiveCompletionConfirmMs;
}
/**
* Get the current timing history for monitoring.
* @returns Copy of timing history
*/
getTimingHistory(): TimingHistory {
return { ...this.timingHistory };
}
/**
* Reset all timing history.
*/
reset(): void {
this.timingHistory = {
recentIdleDetectionMs: [],
recentCycleDurationMs: [],
adaptiveCompletionConfirmMs: 10000,
sampleCount: 0,
maxSamples: 20,
lastUpdatedAt: Date.now(),
};
}
/**
* Recalculate the adaptive completion confirm timeout based on historical data.
* Uses the 75th percentile of recent idle detection times as the new timeout,
* with a 20% buffer for safety.
*/
private updateAdaptiveTiming(): void {
const history = this.timingHistory;
const minMs = this.config.adaptiveMinConfirmMs;
const maxMs = this.config.adaptiveMaxConfirmMs;
if (history.recentIdleDetectionMs.length < 5) return;
// Sort for percentile calculation
const sorted = [...history.recentIdleDetectionMs].sort((a, b) => a - b);
// Use 75th percentile with 20% buffer
const p75Index = Math.floor(sorted.length * 0.75);
const p75Value = sorted[p75Index];
const withBuffer = Math.round(p75Value * 1.2);
// Clamp to configured bounds
const clamped = Math.max(minMs, Math.min(maxMs, withBuffer));
history.adaptiveCompletionConfirmMs = clamped;
}
}
+235 -680
View File
File diff suppressed because it is too large Load Diff
+229
View File
@@ -0,0 +1,229 @@
/**
* @fileoverview Pure health scoring functions for respawn controller.
*
* Extracted from respawn-controller.ts for modularity. All functions are pure
* (no side effects, no state) and take a HealthInputs interface that decouples
* them from direct access to Session, RalphTracker, or AiChecker instances.
*
* @module respawn-health
*/
import type { RespawnAggregateMetrics, RalphLoopHealthScore, HealthStatus, CircuitBreakerStatus } from './types.js';
/**
* Input data for health score calculation.
* Decouples the health calculation from direct access to controller internals.
*/
export interface HealthInputs {
/** Aggregate cycle metrics */
aggregateMetrics: RespawnAggregateMetrics;
/** Current circuit breaker status */
circuitBreakerStatus: CircuitBreakerStatus | null;
/** Iteration stall metrics, or null if tracker unavailable */
iterationStallMetrics: {
stallDurationMs: number;
warningThresholdMs: number;
criticalThresholdMs: number;
} | null;
/** AI checker state summary */
aiCheckerState: {
status: string;
consecutiveErrors: number;
};
/** Number of stuck-state recovery attempts */
stuckRecoveryCount: number;
/** Maximum allowed stuck recoveries */
maxStuckRecoveries: number;
}
/**
* Calculate a comprehensive health score for the Ralph Loop system.
* Aggregates multiple health signals into a single score (0-100).
*
* @param inputs - Health calculation inputs
* @returns Health score with component breakdown
*/
export function calculateHealthScore(inputs: HealthInputs): RalphLoopHealthScore {
const now = Date.now();
const components = {
cycleSuccess: calculateCycleSuccessScore(inputs.aggregateMetrics),
circuitBreaker: calculateCircuitBreakerScore(inputs.circuitBreakerStatus),
iterationProgress: calculateIterationProgressScore(inputs.iterationStallMetrics),
aiChecker: calculateAiCheckerScore(inputs.aiCheckerState),
stuckRecovery: calculateStuckRecoveryScore(inputs.stuckRecoveryCount, inputs.maxStuckRecoveries),
};
// Weighted average (cycle success is most important)
const weights = {
cycleSuccess: 0.35,
circuitBreaker: 0.2,
iterationProgress: 0.2,
aiChecker: 0.15,
stuckRecovery: 0.1,
};
const score = Math.round(
components.cycleSuccess * weights.cycleSuccess +
components.circuitBreaker * weights.circuitBreaker +
components.iterationProgress * weights.iterationProgress +
components.aiChecker * weights.aiChecker +
components.stuckRecovery * weights.stuckRecovery
);
// Determine status
let status: HealthStatus;
if (score >= 90) status = 'excellent';
else if (score >= 70) status = 'good';
else if (score >= 50) status = 'degraded';
else status = 'critical';
// Generate recommendations
const recommendations = generateHealthRecommendations(components);
// Generate summary
const summary = generateHealthSummary(score, status, components);
return {
score,
status,
components,
summary,
recommendations,
calculatedAt: now,
};
}
/**
* Determine whether to skip the /clear step based on current context usage.
* Skips if token count is below the configured threshold percentage.
*
* @param lastTokenCount - Current token count from the session
* @param skipClearThresholdPercent - Threshold percentage below which to skip /clear
* @param maxContextTokens - Approximate max context window size
* @returns True if /clear should be skipped
*/
export function shouldSkipClear(
lastTokenCount: number,
skipClearThresholdPercent: number,
maxContextTokens: number
): boolean {
if (lastTokenCount === 0) return false; // Can't determine, don't skip
const usagePercent = (lastTokenCount / maxContextTokens) * 100;
return usagePercent < skipClearThresholdPercent;
}
/**
* Calculate score based on recent cycle success rate.
*/
function calculateCycleSuccessScore(aggregateMetrics: RespawnAggregateMetrics): number {
if (aggregateMetrics.totalCycles === 0) return 100; // No data = assume healthy
return aggregateMetrics.successRate;
}
/**
* Calculate score based on circuit breaker state.
*/
function calculateCircuitBreakerScore(circuitBreakerStatus: CircuitBreakerStatus | null): number {
if (!circuitBreakerStatus) return 100;
switch (circuitBreakerStatus.state) {
case 'CLOSED':
return 100;
case 'HALF_OPEN':
return 50;
case 'OPEN':
return 0;
default:
return 100;
}
}
/**
* Calculate score based on iteration progress.
*/
function calculateIterationProgressScore(
stallMetrics: { stallDurationMs: number; warningThresholdMs: number; criticalThresholdMs: number } | null
): number {
if (!stallMetrics) return 100;
const { stallDurationMs, warningThresholdMs, criticalThresholdMs } = stallMetrics;
if (stallDurationMs >= criticalThresholdMs) return 0;
if (stallDurationMs >= warningThresholdMs) return 30;
if (stallDurationMs >= warningThresholdMs / 2) return 70;
return 100;
}
/**
* Calculate score based on AI checker health.
*/
function calculateAiCheckerScore(aiCheckerState: { status: string; consecutiveErrors: number }): number {
if (aiCheckerState.status === 'disabled') return 30;
if (aiCheckerState.status === 'cooldown') return 70;
if (aiCheckerState.consecutiveErrors > 0) return 50;
return 100;
}
/**
* Calculate score based on stuck-state recovery count.
*/
function calculateStuckRecoveryScore(stuckRecoveryCount: number, maxStuckRecoveries: number): number {
if (stuckRecoveryCount === 0) return 100;
if (stuckRecoveryCount >= maxStuckRecoveries) return 0;
return Math.round(100 - (stuckRecoveryCount / maxStuckRecoveries) * 100);
}
/**
* Generate health recommendations based on component scores.
*/
function generateHealthRecommendations(components: RalphLoopHealthScore['components']): string[] {
const recommendations: string[] = [];
if (components.cycleSuccess < 70) {
recommendations.push('Cycle success rate is low. Check for recurring errors or stuck states.');
}
if (components.circuitBreaker < 50) {
recommendations.push('Circuit breaker is open or half-open. Review recent errors and consider manual reset.');
}
if (components.iterationProgress < 50) {
recommendations.push('Iteration progress has stalled. Check if Claude is stuck on a task.');
}
if (components.aiChecker < 50) {
recommendations.push('AI idle checker has errors. May need to check Claude CLI availability.');
}
if (components.stuckRecovery < 50) {
recommendations.push('Multiple stuck-state recoveries occurred. Consider increasing timeouts.');
}
if (recommendations.length === 0) {
recommendations.push('System is healthy. No action needed.');
}
return recommendations;
}
/**
* Generate a human-readable health summary.
*/
function generateHealthSummary(
score: number,
status: HealthStatus,
components: RalphLoopHealthScore['components']
): string {
const lowest = Object.entries(components).reduce((min, [key, val]) => (val < min.val ? { key, val } : min), {
key: '',
val: 100,
});
if (status === 'excellent') {
return `Ralph Loop is operating excellently (${score}/100). All systems healthy.`;
}
if (status === 'good') {
return `Ralph Loop is operating well (${score}/100). Minor issues in ${lowest.key}.`;
}
if (status === 'degraded') {
return `Ralph Loop is degraded (${score}/100). Primary issue: ${lowest.key} (${lowest.val}/100).`;
}
return `Ralph Loop is in critical state (${score}/100). Immediate attention needed: ${lowest.key}.`;
}
+229
View File
@@ -0,0 +1,229 @@
/**
* @fileoverview Cycle metrics tracker for respawn controller.
*
* Extracted from respawn-controller.ts for modularity. Tracks per-cycle metrics
* and maintains aggregate statistics across all tracked cycles.
*
* @module respawn-metrics
*/
import { assertNever } from './utils/index.js';
import type { RespawnCycleMetrics, RespawnAggregateMetrics, CycleOutcome } from './types.js';
/**
* Maximum number of cycle metrics to keep in memory.
*/
const MAX_CYCLE_METRICS_IN_MEMORY = 100;
/**
* Tracks respawn cycle metrics and maintains aggregate statistics.
*
* Each respawn cycle is tracked from start to completion, recording timing,
* steps completed, and outcome. Aggregate metrics provide a rolling view
* of system health across recent cycles.
*/
export class RespawnCycleMetricsTracker {
/** Current cycle being tracked */
private currentCycleMetrics: Partial<RespawnCycleMetrics> | null = null;
/** Recent cycle metrics (rolling window for aggregate calculation) */
private recentCycleMetrics: RespawnCycleMetrics[] = [];
/** Aggregate metrics across all tracked cycles */
private aggregateMetrics: RespawnAggregateMetrics = {
totalCycles: 0,
successfulCycles: 0,
stuckRecoveryCycles: 0,
blockedCycles: 0,
errorCycles: 0,
avgCycleDurationMs: 0,
avgIdleDetectionMs: 0,
p90CycleDurationMs: 0,
successRate: 100,
lastUpdatedAt: Date.now(),
};
/**
* Start tracking metrics for a new cycle.
* Called when a respawn cycle begins.
*
* @param sessionId - The session this cycle belongs to
* @param cycleNumber - The cycle number within the session
* @param idleReason - What triggered idle detection
* @param idleDetectionStartTime - Timestamp when idle detection started
* @param lastTokenCount - Token count at start of cycle
* @param adaptiveCompletionConfirmMs - Completion confirm timeout used
*/
startCycle(
sessionId: string,
cycleNumber: number,
idleReason: string,
idleDetectionStartTime: number,
lastTokenCount: number,
adaptiveCompletionConfirmMs: number
): void {
const now = Date.now();
this.currentCycleMetrics = {
cycleId: `${sessionId}:${cycleNumber}`,
sessionId,
cycleNumber,
startedAt: now,
idleReason,
idleDetectionMs: now - idleDetectionStartTime,
stepsCompleted: [],
clearSkipped: false,
tokenCountAtStart: lastTokenCount,
completionConfirmMsUsed: adaptiveCompletionConfirmMs,
};
}
/**
* Record a completed step in the current cycle.
* @param step - Name of the step (e.g., 'update', 'clear', 'init')
*/
recordStep(step: string): void {
if (!this.currentCycleMetrics) return;
this.currentCycleMetrics.stepsCompleted?.push(step);
}
/**
* Mark that /clear was skipped in the current cycle.
*/
markClearSkipped(): void {
if (this.currentCycleMetrics) {
this.currentCycleMetrics.clearSkipped = true;
}
}
/**
* Get the current in-progress cycle metrics (for external inspection).
* @returns The current cycle metrics, or null if no cycle is in progress
*/
getCurrentCycle(): Partial<RespawnCycleMetrics> | null {
return this.currentCycleMetrics;
}
/**
* Complete the current cycle metrics with outcome.
* Adds to recent metrics and updates aggregates.
*
* @param outcome - Outcome of the cycle
* @param lastTokenCount - Token count at end of cycle
* @param errorMessage - Optional error message if outcome is 'error'
* @returns The completed cycle metrics, or null if no cycle was in progress
*/
completeCycle(outcome: CycleOutcome, lastTokenCount: number, errorMessage?: string): RespawnCycleMetrics | null {
if (!this.currentCycleMetrics) return null;
const now = Date.now();
const metrics: RespawnCycleMetrics = {
...(this.currentCycleMetrics as RespawnCycleMetrics),
completedAt: now,
durationMs: now - (this.currentCycleMetrics.startedAt ?? now),
outcome,
errorMessage,
tokenCountAtEnd: lastTokenCount,
};
// Add to recent metrics
this.recentCycleMetrics.push(metrics);
if (this.recentCycleMetrics.length > MAX_CYCLE_METRICS_IN_MEMORY) {
this.recentCycleMetrics.shift();
}
// Update aggregate metrics
this.updateAggregateMetrics(metrics);
// Clear current cycle
this.currentCycleMetrics = null;
return metrics;
}
/**
* Get aggregate metrics for monitoring.
* @returns Copy of aggregate metrics
*/
getAggregate(): RespawnAggregateMetrics {
return { ...this.aggregateMetrics };
}
/**
* Get recent cycle metrics for analysis.
* @param limit - Maximum number of metrics to return (default: 20)
* @returns Recent cycle metrics, newest first
*/
getRecent(limit: number = 20): RespawnCycleMetrics[] {
return this.recentCycleMetrics.slice(-limit).reverse();
}
/**
* Reset all metrics state.
*/
reset(): void {
this.currentCycleMetrics = null;
this.recentCycleMetrics = [];
this.aggregateMetrics = {
totalCycles: 0,
successfulCycles: 0,
stuckRecoveryCycles: 0,
blockedCycles: 0,
errorCycles: 0,
avgCycleDurationMs: 0,
avgIdleDetectionMs: 0,
p90CycleDurationMs: 0,
successRate: 100,
lastUpdatedAt: Date.now(),
};
}
/**
* Update aggregate metrics with a new cycle's data.
* @param metrics - The completed cycle metrics
*/
private updateAggregateMetrics(metrics: RespawnCycleMetrics): void {
const agg = this.aggregateMetrics;
agg.totalCycles++;
switch (metrics.outcome) {
case 'success':
agg.successfulCycles++;
break;
case 'stuck_recovery':
agg.stuckRecoveryCycles++;
break;
case 'blocked':
agg.blockedCycles++;
break;
case 'error':
agg.errorCycles++;
break;
case 'cancelled':
// Cancelled cycles don't count towards any specific category
// but are still counted in totalCycles
break;
default:
assertNever(metrics.outcome, `Unhandled CycleOutcome: ${metrics.outcome}`);
}
// Recalculate averages using all recent metrics
const durations = this.recentCycleMetrics.map((m) => m.durationMs);
const idleTimes = this.recentCycleMetrics.map((m) => m.idleDetectionMs);
if (durations.length > 0) {
agg.avgCycleDurationMs = Math.round(durations.reduce((a, b) => a + b, 0) / durations.length);
agg.avgIdleDetectionMs = Math.round(idleTimes.reduce((a, b) => a + b, 0) / idleTimes.length);
// Calculate P90
const sortedDurations = [...durations].sort((a, b) => a - b);
const p90Index = Math.floor(sortedDurations.length * 0.9);
agg.p90CycleDurationMs = sortedDurations[p90Index];
}
// Calculate success rate
agg.successRate = agg.totalCycles > 0 ? Math.round((agg.successfulCycles / agg.totalCycles) * 100) : 100;
agg.lastUpdatedAt = Date.now();
}
}
+131
View File
@@ -0,0 +1,131 @@
/**
* @fileoverview Pure utility functions for terminal pattern detection in respawn controller.
*
* Extracted from respawn-controller.ts for modularity. These are stateless functions
* and constants used to detect completion messages, working patterns, and token counts
* in terminal output.
*
* @module respawn-patterns
*/
import { TOKEN_PATTERN } from './utils/index.js';
// ========== Constants ==========
/**
* Pattern to detect completion messages from Claude Code.
* Requires "Worked for" prefix to avoid false positives from bare time durations
* in regular text (e.g., "wait for 5s", "run for 2m").
*
* Matches: "Worked for 2m 46s", "Worked for 46s", "Worked for 1h 2m 3s"
* Does NOT match: "wait for 5s", "run for 2m", "for 3s the system..."
*/
const COMPLETION_TIME_PATTERN = /\bWorked\s+for\s+\d+[hms](\s*\d+[hms])*/i;
/**
* Patterns indicating Claude is ready for input (legacy fallback).
* Used as secondary signals, not primary detection.
*/
export const PROMPT_PATTERNS = [
'❯', // Standard prompt
'\u276f', // Unicode variant
'⏵', // Claude Code prompt variant
];
/**
* Patterns indicating Claude is actively working.
* When detected, resets all idle detection timers.
* Note: ✻ and ✽ removed - they appear in completion messages too.
*/
export const WORKING_PATTERNS = [
'Thinking',
'Writing',
'Reading',
'Running',
'Searching',
'Editing',
'Creating',
'Deleting',
'Analyzing',
'Executing',
'Synthesizing',
'Brewing', // Claude's processing indicators
'Compiling',
'Building',
'Installing',
'Fetching',
'Downloading',
'Processing',
'Generating',
'Loading',
'Starting',
'Updating',
'Checking',
'Validating',
'Testing',
'Formatting',
'Linting',
'⠋',
'⠙',
'⠹',
'⠸',
'⠼',
'⠴',
'⠦',
'⠧',
'⠇',
'⠏', // Spinner chars
'◐',
'◓',
'◑',
'◒', // Alternative spinners
'⣾',
'⣽',
'⣻',
'⢿',
'⡿',
'⣟',
'⣯',
'⣷', // Braille spinners
];
/**
* Check if data contains a completion message pattern.
* Matches "Worked for Xh Xm Xs" time duration patterns.
*
* @param data - Raw terminal output data
* @returns True if completion message pattern is found
*/
export function isCompletionMessage(data: string): boolean {
return COMPLETION_TIME_PATTERN.test(data);
}
/**
* Check if a rolling window of terminal output contains working patterns.
* The rolling window catches patterns split across chunks (e.g., "Thin" + "king").
*
* @param window - Rolling window of recent terminal output (already includes current data)
* @returns True if any working pattern is found in the window
*/
export function hasWorkingPattern(window: string): boolean {
return WORKING_PATTERNS.some((pattern) => window.includes(pattern));
}
/**
* Extract token count from data if present.
* Parses patterns like "123.4k tokens" or "1.5M tokens".
*
* @param data - Raw terminal output data
* @returns Parsed token count, or null if no token pattern found
*/
export function extractTokenCount(data: string): number | null {
const match = data.match(TOKEN_PATTERN);
if (!match) return null;
let count = parseFloat(match[1]);
const suffix = match[2]?.toLowerCase();
if (suffix === 'k') count *= 1000;
else if (suffix === 'm') count *= 1000000;
return Math.round(count);
}
+2 -1
View File
@@ -22,6 +22,7 @@ import {
RunSummaryStats,
createInitialRunSummaryStats,
} from './types.js';
import { CLEANUP_CHECK_INTERVAL_MS } from './config/server-timing.js';
/** Maximum events to keep per session (FIFO trimming) */
const MAX_EVENTS = 1000;
@@ -36,7 +37,7 @@ const TOKEN_MILESTONE_INTERVAL = 50000;
const STATE_STUCK_WARNING_MS = 10 * 60 * 1000; // 10 minutes
/** State stuck check interval (ms) */
const STATE_STUCK_CHECK_INTERVAL = 60 * 1000; // 1 minute
const STATE_STUCK_CHECK_INTERVAL = CLEANUP_CHECK_INTERVAL_MS;
/**
* Tracks events and statistics for a session's run summary.
+284
View File
@@ -0,0 +1,284 @@
/**
* @fileoverview Auto-compact and auto-clear automation for Session.
*
* Monitors token counts and triggers /compact or /clear commands when
* configurable thresholds are reached. Waits for Claude to be idle
* before sending commands, with retry logic and mutual exclusion
* (compact and clear never run simultaneously).
*
* @module session-auto-ops
*/
import { EventEmitter } from 'node:events';
// ============================================================================
// Timing Constants
// ============================================================================
/** Delay for auto-compact/clear retry attempts (2 seconds) */
const AUTO_RETRY_DELAY_MS = 2000;
/** Delay for auto-compact/clear initial check (1 second) */
const AUTO_INITIAL_DELAY_MS = 1000;
/** Cooldown after compact completes before re-enabling (10 seconds) */
const COMPACT_COOLDOWN_MS = 10000;
/** Cooldown after clear completes before re-enabling (5 seconds) */
const CLEAR_COOLDOWN_MS = 5000;
/** Minimum valid threshold for auto-clear/compact (1000 tokens) */
const MIN_AUTO_THRESHOLD = 1000;
/** Maximum valid threshold for auto-clear/compact (500k tokens) */
const MAX_AUTO_THRESHOLD = 500_000;
/** Default auto-clear threshold when invalid value provided */
const DEFAULT_AUTO_CLEAR_THRESHOLD = 140_000;
/** Default auto-compact threshold when invalid value provided */
const DEFAULT_AUTO_COMPACT_THRESHOLD = 110_000;
/**
* Callbacks required by SessionAutoOps to interact with the parent Session.
*/
export interface AutoOpsCallbacks {
/** Send a command via the terminal multiplexer */
writeCommand: (command: string) => Promise<boolean>;
/** Check if Claude is currently working */
isWorking: () => boolean;
/** Check if the session has been stopped */
isStopped: () => boolean;
/** Get current total token count (input + output) */
getTotalTokens: () => number;
/** Get session ID for logging */
getSessionId: () => string;
}
/**
* Events emitted by SessionAutoOps.
*/
export interface SessionAutoOpsEvents {
/** Auto-compact was triggered and the /compact command was sent */
autoCompact: (data: { tokens: number; threshold: number; prompt?: string }) => void;
/** Auto-clear was triggered and the /clear command was sent */
autoClear: (data: { tokens: number; threshold: number }) => void;
}
/**
* Manages auto-compact and auto-clear automation for a Session.
*
* When enabled, monitors token counts after each update and triggers
* /compact or /clear commands when thresholds are exceeded. Ensures
* mutual exclusion between compact and clear operations.
*/
export class SessionAutoOps extends EventEmitter {
// Auto-compact state
private _autoCompactThreshold: number;
private _autoCompactEnabled: boolean = false;
private _autoCompactPrompt: string = '';
private _isCompacting: boolean = false;
private _autoCompactTimer: NodeJS.Timeout | null = null;
// Auto-clear state
private _autoClearThreshold: number;
private _autoClearEnabled: boolean = false;
private _isClearing: boolean = false;
private _autoClearTimer: NodeJS.Timeout | null = null;
private readonly callbacks: AutoOpsCallbacks;
constructor(callbacks: AutoOpsCallbacks, config?: { compactThreshold?: number; clearThreshold?: number }) {
super();
this.callbacks = callbacks;
this._autoCompactThreshold = config?.compactThreshold ?? DEFAULT_AUTO_COMPACT_THRESHOLD;
this._autoClearThreshold = config?.clearThreshold ?? DEFAULT_AUTO_CLEAR_THRESHOLD;
}
// ============================================================================
// Auto-compact getters/setters
// ============================================================================
get autoCompactThreshold(): number {
return this._autoCompactThreshold;
}
get autoCompactEnabled(): boolean {
return this._autoCompactEnabled;
}
get autoCompactPrompt(): string {
return this._autoCompactPrompt;
}
get isCompacting(): boolean {
return this._isCompacting;
}
setAutoCompact(enabled: boolean, threshold?: number, prompt?: string): void {
this._autoCompactEnabled = enabled;
if (threshold !== undefined) {
if (threshold < MIN_AUTO_THRESHOLD || threshold > MAX_AUTO_THRESHOLD) {
console.warn(
`[SessionAutoOps ${this.callbacks.getSessionId()}] Invalid autoCompact threshold ${threshold}, must be between ${MIN_AUTO_THRESHOLD} and ${MAX_AUTO_THRESHOLD}. Using default ${DEFAULT_AUTO_COMPACT_THRESHOLD}.`
);
this._autoCompactThreshold = DEFAULT_AUTO_COMPACT_THRESHOLD;
} else {
this._autoCompactThreshold = threshold;
}
}
if (prompt !== undefined) {
this._autoCompactPrompt = prompt;
}
}
// ============================================================================
// Auto-clear getters/setters
// ============================================================================
get autoClearThreshold(): number {
return this._autoClearThreshold;
}
get autoClearEnabled(): boolean {
return this._autoClearEnabled;
}
get isClearing(): boolean {
return this._isClearing;
}
setAutoClear(enabled: boolean, threshold?: number): void {
this._autoClearEnabled = enabled;
if (threshold !== undefined) {
if (threshold < MIN_AUTO_THRESHOLD || threshold > MAX_AUTO_THRESHOLD) {
console.warn(
`[SessionAutoOps ${this.callbacks.getSessionId()}] Invalid autoClear threshold ${threshold}, must be between ${MIN_AUTO_THRESHOLD} and ${MAX_AUTO_THRESHOLD}. Using default ${DEFAULT_AUTO_CLEAR_THRESHOLD}.`
);
this._autoClearThreshold = DEFAULT_AUTO_CLEAR_THRESHOLD;
} else {
this._autoClearThreshold = threshold;
}
}
}
// ============================================================================
// Threshold checks
// ============================================================================
/**
* Check if auto-compact should be triggered based on current token count.
* Called after token count updates.
*/
checkAutoCompact(): void {
if (this.callbacks.isStopped()) return;
if (!this._autoCompactEnabled || this._isCompacting || this._isClearing) return;
const totalTokens = this.callbacks.getTotalTokens();
if (totalTokens >= this._autoCompactThreshold) {
this._isCompacting = true;
console.log(
`[SessionAutoOps] Auto-compact triggered: ${totalTokens} tokens >= ${this._autoCompactThreshold} threshold`
);
const checkAndCompact = async () => {
if (this.callbacks.isStopped()) return;
if (!this._isCompacting) return;
if (!this.callbacks.isWorking()) {
if (this.callbacks.isStopped()) return;
const compactCmd = this._autoCompactPrompt ? `/compact ${this._autoCompactPrompt}\r` : '/compact\r';
await this.callbacks.writeCommand(compactCmd);
this.emit('autoCompact', {
tokens: totalTokens,
threshold: this._autoCompactThreshold,
prompt: this._autoCompactPrompt || undefined,
});
if (!this.callbacks.isStopped()) {
this._autoCompactTimer = setTimeout(() => {
if (this.callbacks.isStopped()) return;
this._autoCompactTimer = null;
this._isCompacting = false;
}, COMPACT_COOLDOWN_MS);
}
} else {
if (!this.callbacks.isStopped()) {
this._autoCompactTimer = setTimeout(checkAndCompact, AUTO_RETRY_DELAY_MS);
}
}
};
if (!this.callbacks.isStopped()) {
this._autoCompactTimer = setTimeout(checkAndCompact, AUTO_INITIAL_DELAY_MS);
}
}
}
/**
* Check if auto-clear should be triggered based on current token count.
* Called after token count updates.
*/
checkAutoClear(): void {
if (this.callbacks.isStopped()) return;
if (!this._autoClearEnabled || this._isClearing || this._isCompacting) return;
const totalTokens = this.callbacks.getTotalTokens();
if (totalTokens >= this._autoClearThreshold) {
this._isClearing = true;
console.log(
`[SessionAutoOps] Auto-clear triggered: ${totalTokens} tokens >= ${this._autoClearThreshold} threshold`
);
const checkAndClear = async () => {
if (this.callbacks.isStopped()) return;
if (!this._isClearing) return;
if (!this.callbacks.isWorking()) {
if (this.callbacks.isStopped()) return;
await this.callbacks.writeCommand('/clear\r');
this.emit('autoClear', { tokens: totalTokens, threshold: this._autoClearThreshold });
if (!this.callbacks.isStopped()) {
this._autoClearTimer = setTimeout(() => {
if (this.callbacks.isStopped()) return;
this._autoClearTimer = null;
this._isClearing = false;
}, CLEAR_COOLDOWN_MS);
}
} else {
if (!this.callbacks.isStopped()) {
this._autoClearTimer = setTimeout(checkAndClear, AUTO_RETRY_DELAY_MS);
}
}
};
if (!this.callbacks.isStopped()) {
this._autoClearTimer = setTimeout(checkAndClear, AUTO_INITIAL_DELAY_MS);
}
}
}
// ============================================================================
// Cleanup
// ============================================================================
/**
* Clear all timers and reset state. Called when the session stops.
*/
destroy(): void {
if (this._autoCompactTimer) {
clearTimeout(this._autoCompactTimer);
this._autoCompactTimer = null;
}
this._isCompacting = false;
if (this._autoClearTimer) {
clearTimeout(this._autoClearTimer);
this._autoClearTimer = null;
}
this._isClearing = false;
}
}
+132
View File
@@ -0,0 +1,132 @@
/**
* @fileoverview Pure functions for building CLI arguments and environment variables
* for Claude and OpenCode CLI spawning.
*
* Extracted from Session to keep argument construction logic testable and
* separate from PTY lifecycle management.
*
* @module session-cli-builder
*/
import type { ClaudeMode } from './types.js';
import { getAugmentedPath } from './utils/index.js';
/**
* Build Claude CLI permission flags based on the configured mode.
* Returns an array of args to pass to the CLI.
*/
export function buildPermissionArgs(claudeMode: ClaudeMode, allowedTools?: string): string[] {
switch (claudeMode) {
case 'dangerously-skip-permissions':
return ['--dangerously-skip-permissions'];
case 'allowedTools':
if (allowedTools) {
return ['--allowedTools', allowedTools];
}
// Fall back to normal mode if no tools specified
return [];
case 'normal':
default:
return [];
}
}
/**
* Build args for an interactive Claude CLI session (direct PTY, non-mux fallback).
*
* @param sessionId - The Codeman session ID (passed as --session-id to Claude)
* @param claudeMode - Permission mode for the CLI
* @param model - Optional model override (e.g., 'opus', 'sonnet')
* @param allowedTools - Optional comma-separated allowed tools list
* @returns Array of CLI arguments
*/
export function buildInteractiveArgs(
sessionId: string,
claudeMode: ClaudeMode,
model?: string,
allowedTools?: string
): string[] {
const args = [...buildPermissionArgs(claudeMode, allowedTools), '--session-id', sessionId];
if (model) args.push('--model', model);
return args;
}
/**
* Build args for a one-shot Claude CLI prompt (runPrompt mode).
*
* @param prompt - The prompt text to send
* @param model - Optional model override
* @returns Array of CLI arguments
*/
export function buildPromptArgs(prompt: string, model?: string): string[] {
const args = ['-p', '--verbose', '--dangerously-skip-permissions', '--output-format', 'stream-json'];
if (model) {
args.push('--model', model);
}
args.push(prompt);
return args;
}
/**
* Build environment variables for Claude CLI processes (direct PTY, non-mux).
*
* Augments process.env with:
* - UTF-8 locale settings
* - Augmented PATH (includes Claude CLI directory)
* - xterm-256color terminal type
* - Codeman session identification vars
*
* @param sessionId - The Codeman session ID
* @returns Environment variables object for pty.spawn
*/
export function buildClaudeEnv(sessionId: string): Record<string, string | undefined> {
return {
...process.env,
LANG: 'en_US.UTF-8',
LC_ALL: 'en_US.UTF-8',
PATH: getAugmentedPath(),
TERM: 'xterm-256color',
COLORTERM: undefined,
CLAUDECODE: undefined,
// Inform Claude it's running within Codeman (helps prevent self-termination)
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: sessionId,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
};
}
/**
* Build environment variables for mux-attached PTY sessions (tmux attach).
* Lighter than buildClaudeEnv — no PATH augmentation or Codeman vars needed
* since the mux session already has those set.
*
* @returns Environment variables object for pty.spawn
*/
export function buildMuxAttachEnv(): Record<string, string | undefined> {
return {
...process.env,
LANG: 'en_US.UTF-8',
LC_ALL: 'en_US.UTF-8',
TERM: 'xterm-256color',
COLORTERM: undefined,
CLAUDECODE: undefined,
};
}
/**
* Build environment variables for a direct shell session (non-mux fallback).
*
* @param sessionId - The Codeman session ID
* @returns Environment variables object for pty.spawn
*/
export function buildShellEnv(sessionId: string): Record<string, string | undefined> {
return {
...process.env,
LANG: 'en_US.UTF-8',
LC_ALL: 'en_US.UTF-8',
TERM: 'xterm-256color',
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: sessionId,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
};
}
+17 -5
View File
@@ -1,11 +1,23 @@
/**
* @fileoverview Session Manager for coordinating multiple Claude sessions
* @fileoverview Session Manager for coordinating multiple Claude sessions.
*
* Provides lifecycle management for Claude CLI sessions:
* - Session creation with working directory configuration
* - Event forwarding from individual sessions
* Lifecycle management for Claude CLI sessions:
* - Session creation with working directory, concurrent session limits (mutex-guarded)
* - Event forwarding from individual sessions to subscribers
* - State persistence via StateStore
* - Concurrent session limits
* - Graceful shutdown of all sessions
*
* Key exports:
* - `SessionManager` class — coordinator, extends EventEmitter
* - `SessionManagerEvents` interface — typed event map
* - `getSessionManager()` — singleton accessor
*
* Key methods: `createSession(workingDir)`, `getSession(id)`, `getAllSessions()`,
* `removeSession(id)`, `stopAll()`
*
* @dependencies session (Session class), state-store (persistence), types (SessionState)
* @consumedby web/server, ralph-loop, respawn-controller
* @emits sessionStarted, sessionStopped, sessionError, sessionOutput, sessionCompletion
*
* @module session-manager
*/
+101
View File
@@ -0,0 +1,101 @@
/**
* @fileoverview LRU cache for task descriptions parsed from terminal output.
*
* Stores descriptions extracted from Claude Code's Task tool invocations
* (e.g., "Explore(Check files)") keyed by timestamp. Used to correlate
* with SubagentWatcher discoveries for better window titles.
*
* @module session-task-cache
*/
import { LRUMap } from './utils/lru-map.js';
/** Default maximum number of task descriptions to keep */
const DEFAULT_MAX_SIZE = 100;
/** Default maximum age for task descriptions (30 seconds) */
const DEFAULT_MAX_AGE_MS = 30_000;
/**
* LRU cache for task descriptions parsed from terminal output.
*
* Descriptions are keyed by the timestamp when they were parsed.
* Old entries are automatically cleaned up based on maxAgeMs.
* Size is bounded by the underlying LRUMap.
*/
export class SessionTaskCache {
private readonly cache: LRUMap<number, string>;
private readonly maxAgeMs: number;
constructor(maxSize: number = DEFAULT_MAX_SIZE, maxAgeMs: number = DEFAULT_MAX_AGE_MS) {
this.cache = new LRUMap<number, string>({ maxSize });
this.maxAgeMs = maxAgeMs;
}
/**
* Add a task description at the given timestamp.
*/
add(timestamp: number, description: string): void {
this.cache.set(timestamp, description);
}
/**
* Remove task descriptions older than maxAgeMs.
* LRUMap maintains insertion order, so we can break early
* once we find a non-expired entry.
*/
private cleanupOld(): void {
const cutoff = Date.now() - this.maxAgeMs;
for (const timestamp of this.cache.keysInOrder()) {
if (timestamp < cutoff) {
this.cache.delete(timestamp);
} else {
break;
}
}
}
/**
* Get all recent task descriptions sorted by timestamp (most recent first).
*/
getAll(): Array<{ timestamp: number; description: string }> {
this.cleanupOld();
const results: Array<{ timestamp: number; description: string }> = [];
for (const [timestamp, description] of this.cache) {
results.push({ timestamp, description });
}
return results.sort((a, b) => b.timestamp - a.timestamp);
}
/**
* Find a task description that was parsed close to a given timestamp.
* Used to correlate with SubagentWatcher discoveries.
*
* @param subagentStartTime - The timestamp when the subagent was discovered
* @param maxAgeMs - Maximum age difference to consider (default 10 seconds)
* @returns The matching description or undefined
*/
findNear(subagentStartTime: number, maxAgeMs: number = 10000): string | undefined {
this.cleanupOld();
let bestMatch: { timestamp: number; description: string } | undefined;
let bestDiff = Infinity;
for (const [timestamp, description] of this.cache) {
const diff = Math.abs(subagentStartTime - timestamp);
if (diff < maxAgeMs && diff < bestDiff) {
bestMatch = { timestamp, description };
bestDiff = diff;
}
}
return bestMatch?.description;
}
/**
* Clear all cached descriptions.
*/
clear(): void {
this.cache.clear();
}
}
+110 -343
View File
@@ -1,16 +1,29 @@
/**
* @fileoverview Core PTY session wrapper for Claude CLI interactions.
*
* This module provides the Session class which manages a PTY (pseudo-terminal)
* process running the Claude CLI. It supports three operation modes:
* Manages a PTY (pseudo-terminal) process running Claude CLI or OpenCode CLI.
* Three operation modes:
* 1. **One-shot** (`runPrompt`): Single prompt → JSON response
* 2. **Interactive** (`startInteractive`): Persistent interactive session
* 3. **Shell** (`startShell`): Plain bash shell for debugging
*
* 1. **One-shot mode** (`runPrompt`): Execute a single prompt and get JSON response
* 2. **Interactive mode** (`startInteractive`): Start an interactive Claude session
* 3. **Shell mode**: Run a plain bash shell for debugging/testing
* Optionally wraps in a tmux session for persistence across disconnects.
* Tracks tokens, costs, background tasks, and auto-compact/clear.
*
* The session can optionally run inside a tmux session for persistence across disconnects.
* It tracks tokens, costs, background tasks, and supports
* auto-clear/auto-compact functionality when token limits are approached.
* Key exports:
* - `Session` class — main entity, extends EventEmitter
* - `ClaudeMessage` interface — parsed JSON messages from Claude output
* - `SessionEvents` interface — typed event map
*
* Key methods: `runPrompt()`, `startInteractive()`, `startShell()`,
* `writeViaMux()`, `toState()`, `stop()`, `resize()`, `isIdle()`,
* `setAutoCompact()`, `findTaskDescriptionNear()`, `getTerminalBuffer()`
*
* @dependencies session-cli-builder (args/env), session-auto-ops (auto-compact/clear),
* ralph-tracker (todo/completion parsing), bash-tool-parser (tool invocation tracking),
* task-tracker (background tasks), mux-interface (tmux abstraction)
* @consumedby session-manager, web/server, respawn-controller
* @emits session:terminal, session:idle, session:working, session:completion, session:exit
*
* @module session
*/
@@ -35,9 +48,14 @@ import type { TerminalMultiplexer, MuxSession } from './mux-interface.js';
import { TaskTracker, type BackgroundTask } from './task-tracker.js';
import { RalphTracker } from './ralph-tracker.js';
import { BashToolParser } from './bash-tool-parser.js';
import { BufferAccumulator } from './utils/buffer-accumulator.js';
import { LRUMap } from './utils/lru-map.js';
import { ANSI_ESCAPE_PATTERN_FULL, TOKEN_PATTERN, SPINNER_PATTERN, MAX_SESSION_TOKENS } from './utils/index.js';
import {
BufferAccumulator,
ANSI_ESCAPE_PATTERN_FULL,
TOKEN_PATTERN,
SPINNER_PATTERN,
MAX_SESSION_TOKENS,
execPattern,
} from './utils/index.js';
import {
MAX_TERMINAL_BUFFER_SIZE,
TRIM_TERMINAL_TO as TERMINAL_BUFFER_TRIM_SIZE,
@@ -46,6 +64,15 @@ import {
MAX_MESSAGES,
MAX_LINE_BUFFER_SIZE,
} from './config/buffer-limits.js';
import {
buildInteractiveArgs,
buildPromptArgs,
buildClaudeEnv,
buildMuxAttachEnv,
buildShellEnv,
} from './session-cli-builder.js';
import { SessionAutoOps } from './session-auto-ops.js';
import { SessionTaskCache } from './session-task-cache.js';
export type { BackgroundTask } from './task-tracker.js';
export type { RalphTrackerState, RalphTodoItem, ActiveBashTool } from './types.js';
@@ -63,11 +90,7 @@ const MUX_STARTUP_DELAY_MS = 300;
/** Delay before declaring session idle after last output (2 seconds) */
const IDLE_DETECTION_DELAY_MS = 2000;
/** Delay for auto-compact/clear retry attempts (2 seconds) */
const AUTO_RETRY_DELAY_MS = 2000;
/** Delay for auto-compact/clear initial check (1 second) */
const AUTO_INITIAL_DELAY_MS = 1000;
// Note: Auto-compact/clear timing constants moved to session-auto-ops.ts
/** Graceful shutdown delay when stopping session (100ms) */
const GRACEFUL_SHUTDOWN_DELAY_MS = 100;
@@ -93,8 +116,7 @@ const CTRL_L_PATTERN = /\x0c/g;
/** Pattern to split by newlines (CR or LF) */
const NEWLINE_SPLIT_PATTERN = /\r?\n/;
// Claude CLI PATH resolution — shared utility
import { getAugmentedPath } from './utils/claude-cli-resolver.js';
// Note: Claude CLI PATH resolution moved to session-cli-builder.ts (buildClaudeEnv)
/**
* Represents a JSON message from Claude CLI's stream-json output format.
@@ -224,9 +246,8 @@ export class Session extends EventEmitter {
readonly createdAt: number;
readonly mode: SessionMode;
/** Maximum number of task descriptions to keep (LRUMap handles size limit automatically) */
private static readonly MAX_TASK_DESCRIPTIONS = 100;
private static readonly TASK_DESCRIPTION_MAX_AGE_MS = 30000; // Keep descriptions for 30 seconds
// Task description cache (extracted to SessionTaskCache)
private _taskCache = new SessionTaskCache();
private _name: string;
private ptyProcess: pty.IPty | null = null;
@@ -255,15 +276,9 @@ export class Session extends EventEmitter {
// Token tracking for auto-clear
private _totalInputTokens: number = 0;
private _totalOutputTokens: number = 0;
private _autoClearThreshold: number = 140000; // Default 140k tokens
private _autoClearEnabled: boolean = false;
private _isClearing: boolean = false; // Prevent recursive clearing
// Auto-compact settings
private _autoCompactThreshold: number = 110000; // Default 110k tokens (lower than clear)
private _autoCompactEnabled: boolean = false;
private _autoCompactPrompt: string = ''; // Optional prompt for compact
private _isCompacting: boolean = false; // Prevent recursive compacting
// Auto-compact/auto-clear automation (extracted to SessionAutoOps)
private _autoOps!: SessionAutoOps;
// Image watcher setting (per-session toggle)
private _imageWatcherEnabled: boolean = false;
@@ -279,8 +294,6 @@ export class Session extends EventEmitter {
private _cliInfoParsed: boolean = false; // Only parse once per session
// Timer tracking for cleanup (prevents memory leaks)
private _autoCompactTimer: NodeJS.Timeout | null = null;
private _autoClearTimer: NodeJS.Timeout | null = null;
private _promptCheckInterval: NodeJS.Timeout | null = null;
private _promptCheckTimeout: NodeJS.Timeout | null = null;
private _shellIdleTimer: NodeJS.Timeout | null = null;
@@ -311,6 +324,7 @@ export class Session extends EventEmitter {
// OpenCode configuration (only for mode === 'opencode')
private _openCodeConfig: OpenCodeConfig | undefined;
private _resumeSessionId: string | undefined;
// Session color for visual differentiation
private _color: import('./types.js').SessionColor = 'default';
@@ -340,12 +354,7 @@ export class Session extends EventEmitter {
toolsUpdate: (tools: ActiveBashTool[]) => void;
} | null = null;
// Task descriptions parsed from terminal output (e.g., "Explore(Description)")
// Used to correlate with SubagentWatcher discoveries for better window titles
// Uses LRUMap for automatic eviction at MAX_TASK_DESCRIPTIONS limit
private _recentTaskDescriptions: LRUMap<number, string> = new LRUMap({
maxSize: Session.MAX_TASK_DESCRIPTIONS,
});
// Task descriptions parsed from terminal output — delegated to SessionTaskCache
// Throttle expensive PTY processing (Ralph, bash parser, task descriptions)
// Accumulates clean data between processing windows to avoid running regex on every chunk
@@ -374,6 +383,8 @@ export class Session extends EventEmitter {
allowedTools?: string;
/** OpenCode configuration (only for mode === 'opencode') */
openCodeConfig?: OpenCodeConfig;
/** Resume a previous Claude conversation (used after server reboot) */
resumeSessionId?: string;
}
) {
super();
@@ -390,12 +401,10 @@ export class Session extends EventEmitter {
this.createdAt = config.createdAt || Date.now();
this.mode = config.mode || 'claude';
this._name = config.name || '';
this._resumeSessionId = config.resumeSessionId;
this._lastActivityAt = this.createdAt;
// Set claudeSessionId immediately — Codeman always passes --session-id ${this.id}
// to Claude CLI, so the Claude session ID always matches the Codeman session ID.
// This ensures subagent matching works even for recovered sessions (where
// startInteractive() hasn't been called yet).
this._claudeSessionId = this.id;
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = config.resumeSessionId || this.id;
this._mux = config.mux || null;
this._useMux = config.useMux ?? (this._mux !== null && this._mux.isAvailable());
this._muxSession = config.muxSession || null;
@@ -463,6 +472,22 @@ export class Session extends EventEmitter {
this._bashToolParser.on('toolStart', this._bashToolHandlers.toolStart);
this._bashToolParser.on('toolEnd', this._bashToolHandlers.toolEnd);
this._bashToolParser.on('toolsUpdate', this._bashToolHandlers.toolsUpdate);
// Initialize auto-compact/auto-clear automation and forward events
this._autoOps = new SessionAutoOps({
writeCommand: (cmd) => this.writeViaMux(cmd),
isWorking: () => this._isWorking,
isStopped: () => this._isStopped,
getTotalTokens: () => this._totalInputTokens + this._totalOutputTokens,
getSessionId: () => this.id,
});
this._autoOps.on('autoCompact', (data) => this.emit('autoCompact', data));
this._autoOps.on('autoClear', (data) => {
// Reset token counts on clear
this._totalInputTokens = 0;
this._totalOutputTokens = 0;
this.emit('autoClear', data);
});
}
get status(): SessionStatus {
@@ -597,25 +622,7 @@ export class Session extends EventEmitter {
return this._allowedTools;
}
/**
* Build Claude CLI permission flags based on the configured mode.
* Returns an array of args to pass to the CLI.
*/
private _buildPermissionArgs(): string[] {
switch (this._claudeMode) {
case 'dangerously-skip-permissions':
return ['--dangerously-skip-permissions'];
case 'allowedTools':
if (this._allowedTools) {
return ['--allowedTools', this._allowedTools];
}
// Fall back to normal mode if no tools specified
return [];
case 'normal':
default:
return [];
}
}
// Note: _buildPermissionArgs removed — now using buildInteractiveArgs from session-cli-builder.ts
/**
* Set CPU priority configuration.
@@ -689,11 +696,11 @@ export class Session extends EventEmitter {
}
get autoClearThreshold(): number {
return this._autoClearThreshold;
return this._autoOps.autoClearThreshold;
}
get autoClearEnabled(): boolean {
return this._autoClearEnabled;
return this._autoOps.autoClearEnabled;
}
get name(): string {
@@ -704,58 +711,24 @@ export class Session extends EventEmitter {
this._name = value;
}
/** Minimum valid threshold for auto-clear/compact (1000 tokens) */
private static readonly MIN_AUTO_THRESHOLD = 1000;
/** Maximum valid threshold for auto-clear/compact (500k tokens) */
private static readonly MAX_AUTO_THRESHOLD = 500_000;
/** Default auto-clear threshold when invalid value provided */
private static readonly DEFAULT_AUTO_CLEAR_THRESHOLD = 140_000;
/** Default auto-compact threshold when invalid value provided */
private static readonly DEFAULT_AUTO_COMPACT_THRESHOLD = 110_000;
setAutoClear(enabled: boolean, threshold?: number): void {
this._autoClearEnabled = enabled;
if (threshold !== undefined) {
// Validate threshold bounds
if (threshold < Session.MIN_AUTO_THRESHOLD || threshold > Session.MAX_AUTO_THRESHOLD) {
console.warn(
`[Session ${this.id}] Invalid autoClear threshold ${threshold}, must be between ${Session.MIN_AUTO_THRESHOLD} and ${Session.MAX_AUTO_THRESHOLD}. Using default ${Session.DEFAULT_AUTO_CLEAR_THRESHOLD}.`
);
this._autoClearThreshold = Session.DEFAULT_AUTO_CLEAR_THRESHOLD;
} else {
this._autoClearThreshold = threshold;
}
}
this._autoOps.setAutoClear(enabled, threshold);
}
get autoCompactThreshold(): number {
return this._autoCompactThreshold;
return this._autoOps.autoCompactThreshold;
}
get autoCompactEnabled(): boolean {
return this._autoCompactEnabled;
return this._autoOps.autoCompactEnabled;
}
get autoCompactPrompt(): string {
return this._autoCompactPrompt;
return this._autoOps.autoCompactPrompt;
}
setAutoCompact(enabled: boolean, threshold?: number, prompt?: string): void {
this._autoCompactEnabled = enabled;
if (threshold !== undefined) {
// Validate threshold bounds
if (threshold < Session.MIN_AUTO_THRESHOLD || threshold > Session.MAX_AUTO_THRESHOLD) {
console.warn(
`[Session ${this.id}] Invalid autoCompact threshold ${threshold}, must be between ${Session.MIN_AUTO_THRESHOLD} and ${Session.MAX_AUTO_THRESHOLD}. Using default ${Session.DEFAULT_AUTO_COMPACT_THRESHOLD}.`
);
this._autoCompactThreshold = Session.DEFAULT_AUTO_COMPACT_THRESHOLD;
} else {
this._autoCompactThreshold = threshold;
}
}
if (prompt !== undefined) {
this._autoCompactPrompt = prompt;
}
this._autoOps.setAutoCompact(enabled, threshold, prompt);
}
get imageWatcherEnabled(): boolean {
@@ -797,11 +770,11 @@ export class Session extends EventEmitter {
lastActivityAt: this._lastActivityAt,
name: this._name,
mode: this.mode,
autoClearEnabled: this._autoClearEnabled,
autoClearThreshold: this._autoClearThreshold,
autoCompactEnabled: this._autoCompactEnabled,
autoCompactThreshold: this._autoCompactThreshold,
autoCompactPrompt: this._autoCompactPrompt,
autoClearEnabled: this._autoOps.autoClearEnabled,
autoClearThreshold: this._autoOps.autoClearThreshold,
autoCompactEnabled: this._autoOps.autoCompactEnabled,
autoCompactThreshold: this._autoOps.autoCompactThreshold,
autoCompactPrompt: this._autoOps.autoCompactPrompt,
imageWatcherEnabled: this._imageWatcherEnabled,
totalCost: this._totalCost,
inputTokens: this._totalInputTokens,
@@ -820,6 +793,7 @@ export class Session extends EventEmitter {
cliAccountType: this._cliAccountType || undefined,
cliLatestVersion: this._cliLatestVersion || undefined,
openCodeConfig: this._openCodeConfig,
resumeSessionId: this._resumeSessionId,
};
}
@@ -865,8 +839,8 @@ export class Session extends EventEmitter {
total: this._totalInputTokens + this._totalOutputTokens,
},
autoClear: {
enabled: this._autoClearEnabled,
threshold: this._autoClearThreshold,
enabled: this._autoOps.autoClearEnabled,
threshold: this._autoOps.autoClearThreshold,
},
// CPU priority configuration
nice: {
@@ -938,6 +912,7 @@ export class Session extends EventEmitter {
claudeMode: this._claudeMode,
allowedTools: this._allowedTools,
openCodeConfig: this._openCodeConfig,
resumeSessionId: this._resumeSessionId,
});
if (!newPid) {
console.error('[Session] Failed to respawn pane, will create new session');
@@ -964,6 +939,7 @@ export class Session extends EventEmitter {
claudeMode: this._claudeMode,
allowedTools: this._allowedTools,
openCodeConfig: this._openCodeConfig,
resumeSessionId: this._resumeSessionId,
});
console.log('[Session] Created mux session:', this._muxSession.muxName);
// No extra sleep — createSession() already waits for tmux readiness
@@ -979,20 +955,12 @@ export class Session extends EventEmitter {
cols: 120,
rows: 40,
cwd: this.workingDir,
env: {
...process.env,
LANG: 'en_US.UTF-8',
LC_ALL: 'en_US.UTF-8',
TERM: 'xterm-256color',
COLORTERM: undefined,
CLAUDECODE: undefined,
},
env: buildMuxAttachEnv(),
}
);
// Set claudeSessionId immediately since we passed --session-id to Claude
// The mux manager passes --session-id ${sessionId} to Claude
this._claudeSessionId = this.id;
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = this._resumeSessionId || this.id;
} catch (spawnErr) {
console.error('[Session] Failed to spawn PTY for mux attachment:', spawnErr);
this.emit('error', `Failed to attach to mux session: ${spawnErr}`);
@@ -1061,26 +1029,13 @@ export class Session extends EventEmitter {
try {
// Pass --session-id to use the SAME ID as the Codeman session
// This ensures subagents can be directly matched to the correct tab
const args = [...this._buildPermissionArgs(), '--session-id', this.id];
if (this._model) args.push('--model', this._model);
const args = buildInteractiveArgs(this.id, this._claudeMode, this._model, this._allowedTools);
this.ptyProcess = pty.spawn('claude', args, {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
env: {
...process.env,
LANG: 'en_US.UTF-8',
LC_ALL: 'en_US.UTF-8',
PATH: getAugmentedPath(),
TERM: 'xterm-256color',
COLORTERM: undefined,
CLAUDECODE: undefined,
// Inform Claude it's running within Codeman (helps prevent self-termination)
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: this.id,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
},
env: buildClaudeEnv(this.id),
});
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY:', spawnErr);
@@ -1090,9 +1045,8 @@ export class Session extends EventEmitter {
}
}
// Set the claudeSessionId immediately since we passed --session-id
// This ensures subagent matching works without waiting for JSON messages
this._claudeSessionId = this.id;
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = this._resumeSessionId || this.id;
this._pid = this.ptyProcess.pid;
console.log('[Session] Interactive PTY spawned with PID:', this._pid);
@@ -1372,14 +1326,7 @@ export class Session extends EventEmitter {
cols: 120,
rows: 40,
cwd: this.workingDir,
env: {
...process.env,
LANG: 'en_US.UTF-8',
LC_ALL: 'en_US.UTF-8',
TERM: 'xterm-256color',
COLORTERM: undefined,
CLAUDECODE: undefined,
},
env: buildMuxAttachEnv(),
}
);
} catch (spawnErr) {
@@ -1413,15 +1360,7 @@ export class Session extends EventEmitter {
cols: 120,
rows: 40,
cwd: this.workingDir,
env: {
...process.env,
LANG: 'en_US.UTF-8',
LC_ALL: 'en_US.UTF-8',
TERM: 'xterm-256color',
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: this.id,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
},
env: buildShellEnv(this.id),
});
} catch (spawnErr) {
console.error('[Session] Failed to spawn shell PTY:', spawnErr);
@@ -1530,11 +1469,7 @@ export class Session extends EventEmitter {
model ? `(model: ${model})` : ''
);
const args = ['-p', '--verbose', '--dangerously-skip-permissions', '--output-format', 'stream-json'];
if (model) {
args.push('--model', model);
}
args.push(prompt);
const args = buildPromptArgs(prompt, model);
try {
this.ptyProcess = pty.spawn('claude', args, {
@@ -1542,19 +1477,7 @@ export class Session extends EventEmitter {
cols: 120,
rows: 40,
cwd: this.workingDir,
env: {
...process.env,
LANG: 'en_US.UTF-8',
LC_ALL: 'en_US.UTF-8',
PATH: getAugmentedPath(),
TERM: 'xterm-256color',
COLORTERM: undefined,
CLAUDECODE: undefined,
// Inform Claude it's running within Codeman
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: this.id,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
},
env: buildClaudeEnv(this.id),
});
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY for runPrompt:', spawnErr);
@@ -1725,8 +1648,8 @@ export class Session extends EventEmitter {
}
// Check if we should auto-compact or auto-clear
this.checkAutoCompact();
this.checkAutoClear();
this._autoOps.checkAutoCompact();
this._autoOps.checkAutoClear();
}
}
@@ -1777,39 +1700,12 @@ export class Session extends EventEmitter {
// Quick pre-check: skip expensive regex if no common tool patterns present
if (!cleanLine.includes('(') || !cleanLine.includes(')')) return;
// Reset regex lastIndex for global pattern
TASK_TOOL_PATTERN.lastIndex = 0;
let match;
while ((match = TASK_TOOL_PATTERN.exec(cleanLine)) !== null) {
execPattern(TASK_TOOL_PATTERN, cleanLine, (match) => {
const description = match[2].trim();
if (description && description.length > 0) {
const now = Date.now();
this._recentTaskDescriptions.set(now, description);
// Cleanup old entries
this.cleanupOldTaskDescriptions();
this._taskCache.add(Date.now(), description);
}
}
}
/**
* Remove task descriptions older than TASK_DESCRIPTION_MAX_AGE_MS.
* Size limit is handled automatically by LRUMap eviction on set().
*/
private cleanupOldTaskDescriptions(): void {
const cutoff = Date.now() - Session.TASK_DESCRIPTION_MAX_AGE_MS;
// Keys are timestamps - iterate and delete expired entries
// LRUMap maintains insertion order, so we can break early once we find a non-expired entry
for (const timestamp of this._recentTaskDescriptions.keysInOrder()) {
if (timestamp < cutoff) {
this._recentTaskDescriptions.delete(timestamp);
} else {
// Keys are ordered by insertion time (which is the timestamp)
// Once we find a non-expired one, all subsequent are also non-expired
break;
}
}
});
}
/**
@@ -1817,12 +1713,7 @@ export class Session extends EventEmitter {
* Returns descriptions sorted by timestamp (most recent first).
*/
getRecentTaskDescriptions(): Array<{ timestamp: number; description: string }> {
this.cleanupOldTaskDescriptions();
const results: Array<{ timestamp: number; description: string }> = [];
for (const [timestamp, description] of this._recentTaskDescriptions) {
results.push({ timestamp, description });
}
return results.sort((a, b) => b.timestamp - a.timestamp);
return this._taskCache.getAll();
}
/**
@@ -1834,21 +1725,7 @@ export class Session extends EventEmitter {
* @returns The matching description or undefined
*/
findTaskDescriptionNear(subagentStartTime: number, maxAgeMs: number = 10000): string | undefined {
this.cleanupOldTaskDescriptions();
// Find the most recent description that was parsed before or around the subagent start time
let bestMatch: { timestamp: number; description: string } | undefined;
let bestDiff = Infinity;
for (const [timestamp, description] of this._recentTaskDescriptions) {
const diff = Math.abs(subagentStartTime - timestamp);
if (diff < maxAgeMs && diff < bestDiff) {
bestMatch = { timestamp, description };
bestDiff = diff;
}
}
return bestMatch?.description;
return this._taskCache.findNear(subagentStartTime, maxAgeMs);
}
// Parse token count from Claude's status line in interactive mode
@@ -1913,8 +1790,8 @@ export class Session extends EventEmitter {
this._totalOutputTokens += Math.round(delta * 0.4);
// Check if we should auto-compact or auto-clear
this.checkAutoCompact();
this.checkAutoClear();
this._autoOps.checkAutoCompact();
this._autoOps.checkAutoClear();
}
}
}
@@ -1997,107 +1874,7 @@ export class Session extends EventEmitter {
}
}
// Check if we should auto-compact based on token threshold
private checkAutoCompact(): void {
if (this._isStopped) return; // Early exit check
if (!this._autoCompactEnabled || this._isCompacting || this._isClearing) return;
const totalTokens = this._totalInputTokens + this._totalOutputTokens;
if (totalTokens >= this._autoCompactThreshold) {
this._isCompacting = true;
console.log(`[Session] Auto-compact triggered: ${totalTokens} tokens >= ${this._autoCompactThreshold} threshold`);
// Wait for Claude to be idle before compacting
const checkAndCompact = async () => {
// Check if session is still valid (not stopped) - must be first check
if (this._isStopped) return;
if (!this._isCompacting) return;
if (!this._isWorking) {
// Re-check stopped state after async operation might have completed
if (this._isStopped) return;
// Send /compact command with optional prompt
const compactCmd = this._autoCompactPrompt ? `/compact ${this._autoCompactPrompt}\r` : '/compact\r';
await this.writeViaMux(compactCmd);
this.emit('autoCompact', {
tokens: totalTokens,
threshold: this._autoCompactThreshold,
prompt: this._autoCompactPrompt || undefined,
});
// Wait a moment then re-enable (longer than clear since compact takes time)
if (!this._isStopped) {
this._autoCompactTimer = setTimeout(() => {
if (this._isStopped) return; // Check at callback start
this._autoCompactTimer = null;
this._isCompacting = false;
}, 10000);
}
} else {
// Check again after delay
if (!this._isStopped) {
this._autoCompactTimer = setTimeout(checkAndCompact, AUTO_RETRY_DELAY_MS);
}
}
};
// Start checking after a short delay
if (!this._isStopped) {
this._autoCompactTimer = setTimeout(checkAndCompact, AUTO_INITIAL_DELAY_MS);
}
}
}
// Check if we should auto-clear based on token threshold
private checkAutoClear(): void {
if (this._isStopped) return; // Early exit check
if (!this._autoClearEnabled || this._isClearing || this._isCompacting) return;
const totalTokens = this._totalInputTokens + this._totalOutputTokens;
if (totalTokens >= this._autoClearThreshold) {
this._isClearing = true;
console.log(`[Session] Auto-clear triggered: ${totalTokens} tokens >= ${this._autoClearThreshold} threshold`);
// Wait for Claude to be idle before clearing
const checkAndClear = async () => {
// Check if session is still valid (not stopped) - must be first check
if (this._isStopped) return;
if (!this._isClearing) return;
if (!this._isWorking) {
// Re-check stopped state after async operation might have completed
if (this._isStopped) return;
// Send /clear command
await this.writeViaMux('/clear\r');
// Reset token counts
this._totalInputTokens = 0;
this._totalOutputTokens = 0;
this.emit('autoClear', { tokens: totalTokens, threshold: this._autoClearThreshold });
// Wait a moment then re-enable
if (!this._isStopped) {
this._autoClearTimer = setTimeout(() => {
if (this._isStopped) return; // Check at callback start
this._autoClearTimer = null;
this._isClearing = false;
}, 5000);
}
} else {
// Check again after delay
if (!this._isStopped) {
this._autoClearTimer = setTimeout(checkAndClear, AUTO_RETRY_DELAY_MS);
}
}
};
// Start checking after a short delay
if (!this._isStopped) {
this._autoClearTimer = setTimeout(checkAndClear, AUTO_INITIAL_DELAY_MS);
}
}
}
// Note: checkAutoCompact/checkAutoClear moved to SessionAutoOps (this._autoOps)
/**
* Sends input directly to the PTY process.
@@ -2265,18 +2042,8 @@ export class Session extends EventEmitter {
this._lineBufferFlushTimer = null;
}
// Clear auto-compact/auto-clear timers to prevent memory leaks
if (this._autoCompactTimer) {
clearTimeout(this._autoCompactTimer);
this._autoCompactTimer = null;
}
this._isCompacting = false;
if (this._autoClearTimer) {
clearTimeout(this._autoClearTimer);
this._autoClearTimer = null;
}
this._isClearing = false;
// Destroy auto-compact/auto-clear automation (clears its timers)
this._autoOps.destroy();
// Clear prompt check timers
if (this._promptCheckInterval) {
@@ -2357,7 +2124,7 @@ export class Session extends EventEmitter {
this._currentTaskId = null;
// Clear task description cache and agent tree to prevent memory leak
this._recentTaskDescriptions.clear();
this._taskCache.clear();
this._childAgentIds = [];
// Kill the associated mux session if requested
@@ -2413,6 +2180,6 @@ export class Session extends EventEmitter {
this._messages = [];
this._taskTracker.clear();
this._ralphTracker.clear();
this._recentTaskDescriptions.clear();
this._taskCache.clear();
}
}
+29 -32
View File
@@ -1,15 +1,25 @@
/**
* @fileoverview Persistent JSON state storage for Codeman.
*
* This module provides the StateStore class which persists application state
* to `~/.codeman/state.json` with debounced writes to prevent excessive disk I/O.
*
* Persists application state with debounced writes (500ms) to prevent excessive disk I/O.
* State is split into two files:
* - `state.json`: Main app state (sessions, tasks, config)
* - `state-inner.json`: Inner loop state (todos, Ralph loop state per session)
* - `~/.codeman/state.json` — main app state (sessions, tasks, config, global stats)
* - `~/.codeman/state-inner.json` — Ralph loop state per session (changes rapidly)
*
* The separation reduces write frequency since Ralph state changes rapidly
* during Ralph Wiggum loops.
* Key exports:
* - `StateStore` class — singleton store with circuit breaker for save failures
* - `getStore(filePath?)` — factory/singleton accessor
*
* Key methods: `getState()`, `getSessions()`, `setSession()`, `getConfig()`,
* `setConfig()`, `getGlobalStats()`, `getAggregateStats()`, `getTokenStats()`,
* `getDailyStats()`, `getRalphState()`, `setRalphState()`, `save()`, `saveNow()`
*
* Auto-migrates legacy `~/.claudeman/` → `~/.codeman/` on first load.
*
* @dependencies types (AppState, RalphSessionState, GlobalStats, TokenStats),
* utils (Debouncer, MAX_SESSION_TOKENS)
* @consumedby session-manager, ralph-loop, web/server, respawn-controller,
* hooks-config, and most subsystems
*
* @module state-store
*/
@@ -28,7 +38,7 @@ import {
TokenStats,
TokenUsageEntry,
} from './types.js';
import { MAX_SESSION_TOKENS } from './utils/index.js';
import { Debouncer, MAX_SESSION_TOKENS } from './utils/index.js';
/** Debounce delay for batching state writes (ms) */
const SAVE_DEBOUNCE_MS = 500;
@@ -60,7 +70,7 @@ const MAX_CONSECUTIVE_FAILURES = 3;
export class StateStore {
private state: AppState;
private filePath: string;
private saveTimeout: NodeJS.Timeout | null = null;
private saveDeb = new Debouncer(SAVE_DEBOUNCE_MS);
private dirty: boolean = false;
private dirtySessions = new Set<string>();
private cachedSessionJsons = new Map<string, string>();
@@ -68,7 +78,7 @@ export class StateStore {
// Inner state storage (separate from main state to reduce write frequency)
private ralphStates: Map<string, RalphSessionState> = new Map();
private ralphStatePath: string;
private ralphStateSaveTimeout: NodeJS.Timeout | null = null;
private ralphStateSaveDeb = new Debouncer(SAVE_DEBOUNCE_MS);
private ralphStateDirty: boolean = false;
// Circuit breaker for save failures (prevents hammering disk on persistent errors)
@@ -150,14 +160,12 @@ export class StateStore {
*/
save(): void {
this.dirty = true;
if (this.saveTimeout) {
return; // Already scheduled
}
this.saveTimeout = setTimeout(() => {
if (this.saveDeb.isPending) return; // Already scheduled
this.saveDeb.schedule(() => {
this.saveNowAsync().catch((err) => {
console.error('[StateStore] Async save failed:', err);
});
}, SAVE_DEBOUNCE_MS);
});
}
/**
@@ -242,10 +250,7 @@ export class StateStore {
}
private async _doSaveAsync(): Promise<void> {
if (this.saveTimeout) {
clearTimeout(this.saveTimeout);
this.saveTimeout = null;
}
this.saveDeb.cancel();
if (!this.dirty) {
return;
}
@@ -332,10 +337,7 @@ export class StateStore {
* Prefer saveNowAsync() for normal operation.
*/
saveNow(): void {
if (this.saveTimeout) {
clearTimeout(this.saveTimeout);
this.saveTimeout = null;
}
this.saveDeb.cancel();
if (!this.dirty) {
return;
}
@@ -775,12 +777,10 @@ export class StateStore {
// Debounced save for inner states
private saveRalphStates(): void {
this.ralphStateDirty = true;
if (this.ralphStateSaveTimeout) {
return; // Already scheduled
}
this.ralphStateSaveTimeout = setTimeout(() => {
if (this.ralphStateSaveDeb.isPending) return; // Already scheduled
this.ralphStateSaveDeb.schedule(() => {
this.saveRalphStatesNow();
}, SAVE_DEBOUNCE_MS);
});
}
/**
@@ -788,10 +788,7 @@ export class StateStore {
* Writes to temp file first, then renames to prevent corruption on crash.
*/
private saveRalphStatesNow(): void {
if (this.ralphStateSaveTimeout) {
clearTimeout(this.ralphStateSaveTimeout);
this.ralphStateSaveTimeout = null;
}
this.ralphStateSaveDeb.cancel();
if (!this.ralphStateDirty) {
return;
}
+102 -114
View File
@@ -1,8 +1,28 @@
/**
* @fileoverview Subagent Watcher - Real-time monitoring of Claude Code background agents
* @fileoverview Subagent Watcher - Real-time monitoring of Claude Code background agents.
*
* Watches ~/.claude/projects/{project}/{session}/subagents/agent-{id}.jsonl files
* and emits structured events for tool calls, progress, and messages.
* Watches `~/.claude/projects/{project}/{session}/subagents/agent-{id}.jsonl` files
* and emits structured events for tool calls, progress, messages, and tool results.
* Also detects Agent Teams teammates (distinguished by `<teammate-message>` in description).
*
* Key exports:
* - `SubagentWatcher` class — singleton watcher, extends EventEmitter
* - `subagentWatcher` — pre-instantiated singleton instance
* - `SubagentInfo`, `SubagentToolCall`, `SubagentProgress`, `SubagentMessage`,
* `SubagentToolResult`, `SubagentTranscriptEntry` — data interfaces
* - `SubagentEvents` — typed event map
*
* Watched patterns: `~/.claude/projects/{project}/{session}/subagents/agent-{id}.jsonl`
* Parses JSONL entries: user/assistant messages, tool_use/tool_result blocks, progress events.
* Tracks per-agent: status, token counts, model, description, tool call count, liveness (PID).
*
* @dependencies config/map-limits (MAX_TRACKED_AGENTS, PENDING_TOOL_CALL_TTL_MS),
* config/buffer-limits (FILE_PEEK_BYTES), utils (CleanupManager, KeyedDebouncer)
* @consumedby web/server (SSE broadcast), session (subagent-session correlation)
* @emits subagent:discovered, subagent:updated, subagent:tool_call, subagent:tool_result,
* subagent:progress, subagent:message, subagent:completed
*
* @module subagent-watcher
*/
import { EventEmitter } from 'node:events';
@@ -13,7 +33,10 @@ import { homedir } from 'node:os';
import { join, basename } from 'node:path';
import { execFile } from 'node:child_process';
import { readFile, readdir, stat as statAsync } from 'node:fs/promises';
import { PENDING_TOOL_CALL_TTL_MS, MAX_PENDING_TOOL_CALLS } from './config/map-limits.js';
import { PENDING_TOOL_CALL_TTL_MS, MAX_PENDING_TOOL_CALLS, MAX_TRACKED_AGENTS } from './config/map-limits.js';
import { STALE_DATA_MAX_AGE_MS } from './config/server-timing.js';
import { FILE_PEEK_BYTES } from './config/buffer-limits.js';
import { CleanupManager, KeyedDebouncer } from './utils/index.js';
// ========== Types ==========
@@ -130,10 +153,9 @@ const POLL_INTERVAL_MS = 1000; // Base poll interval (lightweight checks)
const FULL_SCAN_EVERY_N_POLLS = 5; // Full directory traversal every 5th poll (5s)
const LIVENESS_CHECK_MS = 10000; // Check if subagent processes are still alive every 10s
const FILE_ALIVE_THRESHOLD_MS = 30000; // File mtime within 30s = agent alive (primary check)
const STALE_COMPLETED_MAX_AGE_MS = 60 * 60 * 1000; // Remove completed agents older than 1 hour
const STALE_COMPLETED_MAX_AGE_MS = STALE_DATA_MAX_AGE_MS; // Remove completed agents older than 1 hour
const STALE_IDLE_MAX_AGE_MS = 4 * 60 * 60 * 1000; // Remove idle agents older than 4 hours
const STARTUP_MAX_FILE_AGE_MS = 4 * 60 * 60 * 1000; // Only load files modified in last 4 hours on startup
const MAX_TRACKED_AGENTS = 500; // Maximum agents to track (LRU eviction when exceeded)
// Internal Claude Code agent patterns to filter out (not real user-initiated subagents)
const INTERNAL_AGENT_PATTERNS = [
@@ -161,12 +183,11 @@ const FILE_CONTENT_DEBOUNCE_MS = 100; // Debounce delay for file content updates
export class SubagentWatcher extends EventEmitter {
private filePositions = new Map<string, number>();
private dirWatchers = new Map<string, FSWatcher>();
// Per-file debounce timers for directory watcher (replaces per-file FSWatchers)
private fileDebouncers = new Map<string, NodeJS.Timeout>();
// Per-file debouncer for directory watcher (replaces per-file FSWatchers)
private fileDeb = new KeyedDebouncer(FILE_CONTENT_DEBOUNCE_MS);
private agentInfo = new Map<string, SubagentInfo>();
private idleTimers = new Map<string, NodeJS.Timeout>();
private pollInterval: NodeJS.Timeout | null = null;
private livenessInterval: NodeJS.Timeout | null = null;
private idleDeb = new KeyedDebouncer(IDLE_TIMEOUT_MS);
private cleanup = new CleanupManager();
private _isRunning = false;
private knownSubagentDirs = new Set<string>();
// Map of agentId -> Map of toolUseId -> { toolName, timestamp } (for linking tool_result to tool_call)
@@ -228,12 +249,16 @@ export class SubagentWatcher extends EventEmitter {
// Periodic scan for new subagent directories
// Full directory traversal only every FULL_SCAN_EVERY_N_POLLS polls (~5s)
// FSWatchers handle known directories between full scans
this.pollInterval = setInterval(() => {
this._pollCount++;
if (this._pollCount % FULL_SCAN_EVERY_N_POLLS === 0) {
this.scanForSubagents().catch((err) => this.emit('subagent:error', err as Error));
}
}, POLL_INTERVAL_MS);
this.cleanup.setInterval(
() => {
this._pollCount++;
if (this._pollCount % FULL_SCAN_EVERY_N_POLLS === 0) {
this.scanForSubagents().catch((err) => this.emit('subagent:error', err as Error));
}
},
POLL_INTERVAL_MS,
{ description: 'subagent directory poll' }
);
// Periodic liveness check for active subagents
this.startLivenessChecker();
@@ -249,54 +274,56 @@ export class SubagentWatcher extends EventEmitter {
* 3. Full pgrep scan (expensive, ~500ms) — only for agents that fail tiers 1+2
*/
private startLivenessChecker(): void {
if (this.livenessInterval) return;
this.cleanup.setInterval(
async () => {
// Guard: prevent concurrent liveness checks (avoids duplicate completed events)
if (this._isCheckingLiveness) return;
this._isCheckingLiveness = true;
this.livenessInterval = setInterval(async () => {
// Guard: prevent concurrent liveness checks (avoids duplicate completed events)
if (this._isCheckingLiveness) return;
this._isCheckingLiveness = true;
try {
// Collect agents that need the expensive pgrep scan
const needsFullScan: SubagentInfo[] = [];
try {
// Collect agents that need the expensive pgrep scan
const needsFullScan: SubagentInfo[] = [];
for (const [_agentId, info] of this.agentInfo) {
if (info.status !== 'active' && info.status !== 'idle') continue;
// Tier 1: File mtime check (~0.3ms per agent)
if (await this.checkSubagentFileAlive(info)) continue;
// Tier 2: Cached PID check (~0.1ms per agent)
if (info.pid && (await this.checkPidAlive(info.pid))) continue;
// Tiers 1+2 failed — need expensive scan for this agent
needsFullScan.push(info);
}
// Tier 3: Full pgrep scan — only if any agents failed cheap checks
if (needsFullScan.length > 0) {
const pidMap = await this.getClaudePids();
for (const info of needsFullScan) {
// Re-check status in case another check completed this agent
for (const [_agentId, info] of this.agentInfo) {
if (info.status !== 'active' && info.status !== 'idle') continue;
const alive = this.checkSubagentAliveFromPidMap(info, pidMap);
if (!alive) {
info.pid = undefined;
info.status = 'completed';
this.pendingToolCalls.delete(info.agentId);
this.emit('subagent:completed', info);
// Tier 1: File mtime check (~0.3ms per agent)
if (await this.checkSubagentFileAlive(info)) continue;
// Tier 2: Cached PID check (~0.1ms per agent)
if (info.pid && (await this.checkPidAlive(info.pid))) continue;
// Tiers 1+2 failed — need expensive scan for this agent
needsFullScan.push(info);
}
// Tier 3: Full pgrep scan — only if any agents failed cheap checks
if (needsFullScan.length > 0) {
const pidMap = await this.getClaudePids();
for (const info of needsFullScan) {
// Re-check status in case another check completed this agent
if (info.status !== 'active' && info.status !== 'idle') continue;
const alive = this.checkSubagentAliveFromPidMap(info, pidMap);
if (!alive) {
info.pid = undefined;
info.status = 'completed';
this.pendingToolCalls.delete(info.agentId);
this.emit('subagent:completed', info);
}
}
}
}
// Periodically clean up stale completed agents (older than 24 hours)
this.cleanupStaleAgents();
} finally {
this._isCheckingLiveness = false;
}
}, LIVENESS_CHECK_MS);
// Periodically clean up stale completed agents (older than 24 hours)
this.cleanupStaleAgents();
} finally {
this._isCheckingLiveness = false;
}
},
LIVENESS_CHECK_MS,
{ description: 'subagent liveness check' }
);
}
/**
@@ -418,21 +445,12 @@ export class SubagentWatcher extends EventEmitter {
stop(): void {
this._isRunning = false;
if (this.pollInterval) {
clearInterval(this.pollInterval);
this.pollInterval = null;
}
if (this.livenessInterval) {
clearInterval(this.livenessInterval);
this.livenessInterval = null;
}
// Dispose poll and liveness intervals, then re-create for potential restart
this.cleanup.dispose();
this.cleanup = new CleanupManager();
// Clear file debouncers
for (const timer of this.fileDebouncers.values()) {
clearTimeout(timer);
}
this.fileDebouncers.clear();
this.fileDeb.dispose();
this.fileAgentContext.clear();
// Remove error handlers before closing watchers to prevent memory leak
@@ -447,10 +465,7 @@ export class SubagentWatcher extends EventEmitter {
}
this.dirWatchers.clear();
for (const timer of this.idleTimers.values()) {
clearTimeout(timer);
}
this.idleTimers.clear();
this.idleDeb.dispose();
// Clear all state for clean restart
this.filePositions.clear();
@@ -532,16 +547,8 @@ export class SubagentWatcher extends EventEmitter {
this.pendingToolCalls.delete(agentId);
this.filePositions.delete(info.filePath);
this.fileAgentContext.delete(info.filePath);
const debounceTimer = this.fileDebouncers.get(info.filePath);
if (debounceTimer) {
clearTimeout(debounceTimer);
this.fileDebouncers.delete(info.filePath);
}
const timer = this.idleTimers.get(agentId);
if (timer) {
clearTimeout(timer);
this.idleTimers.delete(agentId);
}
this.fileDeb.cancelKey(info.filePath);
this.idleDeb.cancelKey(agentId);
}
}
@@ -634,9 +641,9 @@ export class SubagentWatcher extends EventEmitter {
return {
agentCount: this.agentInfo.size,
fileDebouncerCount: this.fileDebouncers.size,
fileDebouncerCount: this.fileDeb.size,
dirWatcherCount: this.dirWatchers.size,
idleTimerCount: this.idleTimers.size,
idleTimerCount: this.idleDeb.size,
pendingToolCallsCount,
knownDirsCount: this.knownSubagentDirs.size,
filePositionsCount: this.filePositions.size,
@@ -1003,7 +1010,7 @@ export class SubagentWatcher extends EventEmitter {
private async extractDescriptionFromFile(filePath: string): Promise<string | undefined> {
try {
// Only read the first 8KB — more than enough for 5 JSONL lines
const stream = createReadStream(filePath, { end: 8191 });
const stream = createReadStream(filePath, { end: FILE_PEEK_BYTES });
const rl = createInterface({ input: stream });
return await new Promise<string | undefined>((resolve) => {
@@ -1129,25 +1136,18 @@ export class SubagentWatcher extends EventEmitter {
if (!filename?.endsWith('.jsonl')) return;
const filePath = join(dir, filename);
// Clear existing debounce for this file
const existing = this.fileDebouncers.get(filePath);
if (existing) clearTimeout(existing);
// Debounce 100ms to batch rapid writes
const timer = setTimeout(() => {
this.fileDebouncers.delete(filePath);
this.fileDeb.schedule(filePath, () => {
if (!existsSync(filePath)) return;
if (this.fileAgentContext.has(filePath)) {
// Known file — handle content change
this.handleFileChange(filePath).catch(() => {});
this.handleFileChange(filePath).catch(() => {}); // Ignore - errors logged internally, don't crash watcher callback
} else {
// New file — register it
this.registerAgentFile(filePath, projectHash, sessionId).catch(() => {});
this.registerAgentFile(filePath, projectHash, sessionId).catch(() => {}); // Ignore - errors logged internally, don't crash watcher callback
}
}, FILE_CONTENT_DEBOUNCE_MS);
this.fileDebouncers.set(filePath, timer);
});
});
// Handle watcher errors to prevent unhandled exceptions
@@ -1602,25 +1602,13 @@ export class SubagentWatcher extends EventEmitter {
* Reset idle timer for an agent
*/
private resetIdleTimer(agentId: string): void {
const existing = this.idleTimers.get(agentId);
if (existing) {
clearTimeout(existing);
}
const timer = setTimeout(() => {
// Guard against race condition: agent may have been deleted before timer fires
this.idleDeb.schedule(agentId, () => {
const info = this.agentInfo.get(agentId);
if (!info) {
// Agent was deleted - clean up timer reference
this.idleTimers.delete(agentId);
return;
}
if (!info) return;
if (info.status === 'active') {
info.status = 'idle';
}
}, IDLE_TIMEOUT_MS);
this.idleTimers.set(agentId, timer);
});
}
/**
+2 -1
View File
@@ -22,6 +22,7 @@
import { EventEmitter } from 'node:events';
import { assertNever } from './utils/index.js';
import { STALE_DATA_MAX_AGE_MS } from './config/server-timing.js';
// ========== Configuration Constants ==========
@@ -36,7 +37,7 @@ const MAX_COMPLETED_TASKS = 100;
* Entries older than this are cleaned up to prevent unbounded growth.
* Default: 1 hour
*/
const PENDING_TOOL_USE_MAX_AGE_MS = 60 * 60 * 1000;
const PENDING_TOOL_USE_MAX_AGE_MS = STALE_DATA_MAX_AGE_MS;
/**
* Maximum number of pending tool uses to allow.
+7 -12
View File
@@ -17,12 +17,7 @@ import { watch as chokidarWatch, type FSWatcher as ChokidarWatcher } from 'choki
import type { TeamConfig, TeamMember, TeamTask, InboxMessage } from './types.js';
import { LRUMap } from './utils/lru-map.js';
// ========== Constants ==========
const POLL_INTERVAL_MS = 30000;
const MAX_CACHED_TEAMS = 50;
const MAX_CACHED_TASKS = 200;
import { TEAM_POLL_INTERVAL_MS, MAX_CACHED_TEAMS, MAX_CACHED_TASKS } from './config/team-config.js';
// ========== TeamWatcher Class ==========
@@ -52,7 +47,7 @@ export class TeamWatcher extends EventEmitter {
start(): void {
if (this.pollTimer) return;
this.poll();
this.pollTimer = setInterval(() => this.poll(), POLL_INTERVAL_MS);
this.pollTimer = setInterval(() => this.poll(), TEAM_POLL_INTERVAL_MS);
this.setupFsWatchers();
}
@@ -66,7 +61,7 @@ export class TeamWatcher extends EventEmitter {
persistent: false,
});
const teamsHandler = () => this.pollAsync().catch(() => {});
const teamsHandler = () => this.pollAsync().catch(() => {}); // Ignore - poll errors are non-fatal, next poll will retry
this.teamsWatcher.on('add', teamsHandler);
this.teamsWatcher.on('change', teamsHandler);
this.teamsWatcher.on('unlink', teamsHandler);
@@ -87,8 +82,8 @@ export class TeamWatcher extends EventEmitter {
persistent: false,
});
this.tasksWatcher.on('add', () => this.pollTasks().catch(() => {}));
this.tasksWatcher.on('change', () => this.pollTasks().catch(() => {}));
this.tasksWatcher.on('add', () => this.pollTasks().catch(() => {})); // Ignore - poll errors are non-fatal, next poll will retry
this.tasksWatcher.on('change', () => this.pollTasks().catch(() => {})); // Ignore - poll errors are non-fatal, next poll will retry
this.tasksWatcher.on('error', (err) => {
console.warn('[TeamWatcher] chokidar tasks watcher error:', err);
});
@@ -100,11 +95,11 @@ export class TeamWatcher extends EventEmitter {
stop(): void {
// Close chokidar watchers
if (this.teamsWatcher) {
this.teamsWatcher.close().catch(() => {});
this.teamsWatcher.close().catch(() => {}); // Ignore - watcher cleanup is best-effort during shutdown
this.teamsWatcher = null;
}
if (this.tasksWatcher) {
this.tasksWatcher.close().catch(() => {});
this.tasksWatcher.close().catch(() => {}); // Ignore - watcher cleanup is best-effort during shutdown
this.tasksWatcher = null;
}
if (this.pollTimer) {
+40 -12
View File
@@ -40,8 +40,7 @@ import {
type SessionMode,
type OpenCodeConfig,
} from './types.js';
import { wrapWithNice } from './utils/nice-wrapper.js';
import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';
import { wrapWithNice, SAFE_PATH_PATTERN, findClaudeDir, resolveOpenCodeDir } from './utils/index.js';
import type {
TerminalMultiplexer,
MuxSession,
@@ -50,17 +49,11 @@ import type {
RespawnPaneOptions,
} from './mux-interface.js';
// Claude CLI PATH resolution — shared utility
import { findClaudeDir } from './utils/claude-cli-resolver.js';
// OpenCode CLI PATH resolution
import { resolveOpenCodeDir } from './utils/opencode-cli-resolver.js';
// ============================================================================
// Timing Constants
// ============================================================================
/** Timeout for exec commands (5 seconds) */
const EXEC_TIMEOUT_MS = 5000;
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
/** Delay after tmux session creation — enough for detached tmux to be queryable */
const TMUX_CREATION_WAIT_MS = 100;
@@ -201,12 +194,24 @@ function buildSpawnCommand(options: {
claudeMode?: ClaudeMode;
allowedTools?: string;
openCodeConfig?: OpenCodeConfig;
resumeSessionId?: string;
}): string {
if (options.mode === 'claude') {
// Validate model to prevent command injection
const safeModel = options.model && /^[a-zA-Z0-9._-]+$/.test(options.model) ? options.model : undefined;
const modelFlag = safeModel ? ` --model ${safeModel}` : '';
return `claude${buildClaudePermissionFlags(options.claudeMode, options.allowedTools)} --session-id "${options.sessionId}"${modelFlag}`;
// Use --resume to restore a previous conversation, otherwise --session-id for new sessions.
// Wrap --resume in a fallback: if it exits non-zero (session not found, corrupt, etc.),
// fall back to a new session with --session-id so the pane doesn't die.
const safeResumeId =
options.resumeSessionId && /^[a-f0-9-]+$/.test(options.resumeSessionId) ? options.resumeSessionId : undefined;
const permFlags = buildClaudePermissionFlags(options.claudeMode, options.allowedTools);
if (safeResumeId) {
const resumeCmd = `claude${permFlags} --resume "${safeResumeId}"${modelFlag}`;
const fallbackCmd = `claude${permFlags} --session-id "${options.sessionId}"${modelFlag}`;
return `${resumeCmd} || ${fallbackCmd}`;
}
return `claude${permFlags} --session-id "${options.sessionId}"${modelFlag}`;
}
if (options.mode === 'opencode') {
return buildOpenCodeCommand(options.openCodeConfig);
@@ -371,7 +376,18 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
* In test mode: creates an in-memory session only (no real tmux session).
*/
async createSession(options: CreateSessionOptions): Promise<MuxSession> {
const { sessionId, workingDir, mode, name, niceConfig, model, claudeMode, allowedTools, openCodeConfig } = options;
const {
sessionId,
workingDir,
mode,
name,
niceConfig,
model,
claudeMode,
allowedTools,
openCodeConfig,
resumeSessionId,
} = options;
const muxName = `codeman-${sessionId.slice(0, 8)}`;
if (!isValidMuxName(muxName)) {
@@ -434,6 +450,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
claudeMode,
allowedTools,
openCodeConfig,
resumeSessionId,
});
const config = niceConfig || DEFAULT_NICE_CONFIG;
@@ -606,7 +623,17 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
* preserving the session and its scrollback buffer.
*/
async respawnPane(options: RespawnPaneOptions): Promise<number | null> {
const { sessionId, workingDir, mode, niceConfig, model, claudeMode, allowedTools, openCodeConfig } = options;
const {
sessionId,
workingDir,
mode,
niceConfig,
model,
claudeMode,
allowedTools,
openCodeConfig,
resumeSessionId,
} = options;
const session = this.sessions.get(sessionId);
if (!session) return null;
const muxName = session.muxName;
@@ -642,6 +669,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
claudeMode,
allowedTools,
openCodeConfig,
resumeSessionId,
});
const config = niceConfig || DEFAULT_NICE_CONFIG;
const cmd = wrapWithNice(baseCmd, config);
+150 -10
View File
@@ -18,6 +18,17 @@ import { spawn, type ChildProcess } from 'node:child_process';
import { existsSync } from 'node:fs';
import { join } from 'node:path';
import { homedir } from 'node:os';
import { randomBytes } from 'node:crypto';
import {
QR_TOKEN_TTL_MS,
QR_TOKEN_GRACE_MS,
SHORT_CODE_LENGTH,
QR_RATE_LIMIT_MAX,
QR_RATE_LIMIT_WINDOW_MS,
URL_TIMEOUT_MS,
RESTART_DELAY_MS,
FORCE_KILL_MS,
} from './config/tunnel-config.js';
// ========== Types ==========
@@ -26,20 +37,29 @@ export interface TunnelStatus {
url: string | null;
}
// ========== Constants ==========
interface QrTokenRecord {
token: string; // 64 hex chars (256 bits)
shortCode: string; // 6 chars base62 (for URL path)
createdAt: number; // Date.now()
consumed: boolean; // single-use flag
}
/** Rejection-sampled base62 short code — no modulo bias */
function generateShortCode(): string {
const chars = 'ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789';
const maxUnbiased = 248; // largest multiple of 62 that fits in a byte (248 = 62 * 4)
const result: string[] = [];
while (result.length < SHORT_CODE_LENGTH) {
const [byte] = randomBytes(1);
if (byte < maxUnbiased) result.push(chars[byte % 62]);
// else: discard and re-draw (rejection sampling)
}
return result.join('');
}
/** Regex to extract the trycloudflare.com URL from cloudflared output */
const TUNNEL_URL_REGEX = /https:\/\/[a-z0-9-]+\.trycloudflare\.com/;
/** Max time to wait for URL before considering it a timeout (ms) */
const URL_TIMEOUT_MS = 30_000;
/** Restart delay after unexpected exit (ms) */
const RESTART_DELAY_MS = 5_000;
/** Force-kill timeout after SIGTERM (ms) */
const FORCE_KILL_MS = 5_000;
// ========== TunnelManager Class ==========
export class TunnelManager extends EventEmitter {
@@ -54,6 +74,17 @@ export class TunnelManager extends EventEmitter {
private localPort = 3000;
private useHttps = false;
// ========== QR Token State ==========
/** Map-based lookup: shortCode → QrTokenRecord (hash-based, timing-safe) */
private qrTokensByCode = new Map<string, QrTokenRecord>();
private currentShortCode: string | null = null;
private rotationTimer: ReturnType<typeof setInterval> | null = null;
/** SVG cache — regenerated only on token rotation, not per request */
private cachedQrSvg: { shortCode: string; svg: string } | null = null;
/** Global rate limit counter (separate from Basic Auth rate limiting) */
private qrAttemptCount = 0;
private qrRateLimitResetTimer: ReturnType<typeof setInterval> | null = null;
/**
* Resolve cloudflared binary path.
* Checks ~/.local/bin first, then falls back to PATH.
@@ -170,6 +201,10 @@ export class TunnelManager extends EventEmitter {
// Detach listeners — no need to parse further output
this.process?.stdout?.off('data', handleOutput);
this.process?.stderr?.off('data', handleOutput);
// Start QR token rotation when tunnel URL is acquired (only if auth enabled)
if (process.env.CODEMAN_PASSWORD) {
this.startTokenRotation();
}
this.emit('started', { url: this.url });
}
};
@@ -243,6 +278,7 @@ export class TunnelManager extends EventEmitter {
stop(): void {
this.stopped = true;
this.clearTimers();
this.stopTokenRotation();
if (this.process) {
const pid = this.process.pid;
@@ -264,6 +300,110 @@ export class TunnelManager extends EventEmitter {
}
}
// ========== QR Token Management ==========
/** Start token rotation — called after tunnel URL is acquired */
startTokenRotation(): void {
this.stopTokenRotation();
this.rotateToken();
this.rotationTimer = setInterval(() => this.rotateToken(), QR_TOKEN_TTL_MS);
this.qrRateLimitResetTimer = setInterval(() => {
this.qrAttemptCount = 0;
}, QR_RATE_LIMIT_WINDOW_MS);
}
/** Stop token rotation and clear all tokens */
stopTokenRotation(): void {
if (this.rotationTimer) {
clearInterval(this.rotationTimer);
this.rotationTimer = null;
}
if (this.qrRateLimitResetTimer) {
clearInterval(this.qrRateLimitResetTimer);
this.qrRateLimitResetTimer = null;
}
this.qrTokensByCode.clear();
this.currentShortCode = null;
this.cachedQrSvg = null;
this.qrAttemptCount = 0;
}
/** Create a new token, evict expired/consumed ones, emit rotation event */
private rotateToken(): void {
const record: QrTokenRecord = {
token: randomBytes(32).toString('hex'),
shortCode: generateShortCode(),
createdAt: Date.now(),
consumed: false,
};
// Evict expired or consumed tokens
const now = Date.now();
for (const [code, rec] of this.qrTokensByCode) {
if (now - rec.createdAt > QR_TOKEN_GRACE_MS || rec.consumed) {
this.qrTokensByCode.delete(code);
}
}
this.qrTokensByCode.set(record.shortCode, record);
this.currentShortCode = record.shortCode;
this.cachedQrSvg = null; // invalidate SVG cache
this.emit('qrTokenRotated');
}
/** Get the current (newest) token's short code for QR URL */
getCurrentShortCode(): string | undefined {
return this.currentShortCode ?? undefined;
}
/** Get cached QR SVG, regenerating only if the short code changed */
async getQrSvg(tunnelUrl: string): Promise<string> {
const code = this.currentShortCode;
if (!code) throw new Error('No QR token available');
if (this.cachedQrSvg?.shortCode === code) return this.cachedQrSvg.svg;
const QRCode = await import('qrcode');
const svg: string = await QRCode.toString(`${tunnelUrl}/q/${code}`, {
type: 'svg',
margin: 2,
width: 256,
});
this.cachedQrSvg = { shortCode: code, svg };
return svg;
}
/**
* Validate and atomically consume a token by short code.
* Map.get() is hash-based — no timing side-channel from string comparison.
*/
consumeToken(shortCode: string): boolean {
// Global rate limit (across all IPs)
if (this.qrAttemptCount >= QR_RATE_LIMIT_MAX) return false;
this.qrAttemptCount++;
const record = this.qrTokensByCode.get(shortCode);
if (!record) return false;
if (record.consumed) return false;
const now = Date.now();
if (now - record.createdAt > QR_TOKEN_GRACE_MS) return false;
// Atomic consume (single-threaded JS = no race)
record.consumed = true;
// Immediately rotate so desktop gets a fresh QR
this.rotateToken();
this.emit('qrTokenRegenerated');
return true;
}
/** Force-regenerate (manual revocation via API) */
regenerateQrToken(): void {
this.qrTokensByCode.clear();
this.currentShortCode = null;
this.rotateToken();
this.emit('qrTokenRegenerated');
}
isRunning(): boolean {
return this.process !== null || this.restartTimer !== null;
}
+8 -1439
View File
File diff suppressed because it is too large Load Diff
+151
View File
@@ -0,0 +1,151 @@
/**
* @fileoverview API types and error handling.
*
* Defines the standardized API response envelope (ApiResponse), error codes,
* hook event types from Claude Code's hooks system, and utility functions
* for error message extraction. Used by all route modules in `src/web/routes/`.
*
* Key exports:
* - ApiResponse<T> — discriminated union envelope (success with data or error with code)
* - ApiErrorCode — enum of standard error codes with user-friendly messages
* - HookEventType — union of Claude Code hook event names (idle_prompt, stop, etc.)
* - createErrorResponse() — factory for consistent error responses
* - getErrorMessage() — safe extraction from unknown catch values
* - CaseInfo, QuickStartResponse — case folder metadata types
*
* No dependencies on other domain modules. Consumed by all route modules
* and validated via Zod schemas in `src/web/schemas.ts`.
*/
/**
* Standard error codes for API responses
*/
export enum ApiErrorCode {
/** Resource not found */
NOT_FOUND = 'NOT_FOUND',
/** Invalid input provided */
INVALID_INPUT = 'INVALID_INPUT',
/** Session is currently busy */
SESSION_BUSY = 'SESSION_BUSY',
/** Operation failed */
OPERATION_FAILED = 'OPERATION_FAILED',
/** Resource already exists */
ALREADY_EXISTS = 'ALREADY_EXISTS',
/** Internal server error */
INTERNAL_ERROR = 'INTERNAL_ERROR',
}
/**
* User-friendly error messages for each error code
*/
const ErrorMessages: Record<ApiErrorCode, string> = {
[ApiErrorCode.NOT_FOUND]: 'The requested resource was not found',
[ApiErrorCode.INVALID_INPUT]: 'Invalid input provided',
[ApiErrorCode.SESSION_BUSY]: 'Session is currently busy',
[ApiErrorCode.OPERATION_FAILED]: 'The operation failed',
[ApiErrorCode.ALREADY_EXISTS]: 'Resource already exists',
[ApiErrorCode.INTERNAL_ERROR]: 'An internal error occurred',
};
/**
* Hook event types triggered by Claude Code's hooks system
*/
export type HookEventType =
| 'idle_prompt'
| 'permission_prompt'
| 'elicitation_dialog'
| 'stop'
| 'teammate_idle'
| 'task_completed';
// ========== API Response Types ==========
/**
* Standard API response wrapper (discriminated union for type safety)
* @template T Type of the data payload
*/
export type ApiResponse<T = unknown> =
| { success: true; data?: T }
| { success: false; error: string; errorCode: ApiErrorCode };
/**
* Creates a standardized error response
* @param code Error code
* @param details Optional detailed error message
* @returns Formatted error response
*/
export function createErrorResponse(code: ApiErrorCode, details?: string): ApiResponse<never> {
return {
success: false,
error: details || ErrorMessages[code],
errorCode: code,
};
}
/**
* Response for quick start operation
*/
export interface QuickStartResponse {
/** Whether the request succeeded */
success: boolean;
/** Created session ID */
sessionId?: string;
/** Path to case folder */
casePath?: string;
/** Case name */
caseName?: string;
/** Error message if failed */
error?: string;
}
/**
* Information about a case folder
*/
export interface CaseInfo {
/** Case name */
name: string;
/** Full path to case folder */
path: string;
/** Whether CLAUDE.md exists */
hasClaudeMd?: boolean;
}
// ========== Error Handling Utilities ==========
/**
* Type guard to check if a value is an Error instance
* @param value The value to check
* @returns True if the value is an Error instance
*/
export function isError(value: unknown): value is Error {
return value instanceof Error;
}
/**
* Safely extracts an error message from an unknown caught value.
* Handles the TypeScript 4.4+ unknown error type in catch blocks.
*
* @param error The caught error (type unknown in strict mode)
* @returns A string error message
*
* @example
* ```typescript
* try {
* await riskyOperation();
* } catch (err) {
* console.error('Failed:', getErrorMessage(err));
* }
* ```
*/
export function getErrorMessage(error: unknown): string {
if (isError(error)) {
return error.message;
}
if (typeof error === 'string') {
return error;
}
if (error && typeof error === 'object' && 'message' in error) {
return String((error as { message: unknown }).message);
}
return 'An unknown error occurred';
}
+171
View File
@@ -0,0 +1,171 @@
/**
* @fileoverview Application state type definitions.
*
* Defines the top-level persisted state structure (AppState) which composes
* types from multiple domains: SessionState (session), TaskState (task),
* RalphLoopState (ralph), and RespawnConfig (respawn).
*
* Key exports:
* - AppState — root state object (sessions, tasks, ralphLoop, config, globalStats, tokenStats)
* - AppConfig — app configuration including default RespawnConfig
* - GlobalStats — cumulative usage stats across all sessions (lifetime)
* - TokenStats / TokenUsageEntry — daily token usage history
* - DEFAULT_CONFIG — default AppConfig values
* - createInitialState() — factory for fresh AppState
*
* Persisted to `~/.codeman/state.json` via StateStore (debounced 500ms writes).
* Served at `GET /api/status` (full state) and `GET /api/config` (config subset).
*
* Cross-domain imports: SessionState, TaskState, RalphLoopState, RespawnConfig.
*/
import type { SessionState } from './session.js';
import type { TaskState } from './task.js';
import type { RalphLoopState } from './ralph.js';
import type { RespawnConfig } from './respawn.js';
// ========== Global Stats Types ==========
/**
* Global statistics across all sessions (including deleted ones).
* Persisted to track cumulative usage over time.
*/
export interface GlobalStats {
/** Total input tokens used across all sessions */
totalInputTokens: number;
/** Total output tokens used across all sessions */
totalOutputTokens: number;
/** Total cost in USD across all sessions */
totalCost: number;
/** Total number of sessions created (lifetime) */
totalSessionsCreated: number;
/** Timestamp when stats were first recorded */
firstRecordedAt: number;
/** Timestamp of last update */
lastUpdatedAt: number;
}
// ========== Token Usage History Types ==========
/**
* Daily token usage entry for historical tracking.
*/
export interface TokenUsageEntry {
/** Date in YYYY-MM-DD format */
date: string;
/** Input tokens used on this day */
inputTokens: number;
/** Output tokens used on this day */
outputTokens: number;
/** Estimated cost in USD */
estimatedCost: number;
/** Number of sessions that contributed to this day's usage */
sessions: number;
}
/**
* Token usage statistics with daily tracking.
*/
export interface TokenStats {
/** Daily usage entries (most recent first) */
daily: TokenUsageEntry[];
/** Timestamp of last update */
lastUpdated: number;
}
/**
* Application configuration
*/
export interface AppConfig {
/** Interval for polling session status (ms) */
pollIntervalMs: number;
/** Default timeout for tasks (ms) */
defaultTimeoutMs: number;
/** Maximum concurrent sessions allowed */
maxConcurrentSessions: number;
/** Path to state file */
stateFilePath: string;
/** Respawn controller configuration */
respawn: RespawnConfig;
/** Last used case name (for default selection) */
lastUsedCase: string | null;
/** Whether Ralph/Todo tracker is globally enabled for all new sessions */
ralphEnabled: boolean;
}
/**
* Complete application state
*/
export interface AppState {
/** Map of session ID to session state */
sessions: Record<string, SessionState>;
/** Map of task ID to task state */
tasks: Record<string, TaskState>;
/** Ralph Loop controller state */
ralphLoop: RalphLoopState;
/** Application configuration */
config: AppConfig;
/** Global statistics (cumulative across all sessions) */
globalStats?: GlobalStats;
/** Daily token usage statistics */
tokenStats?: TokenStats;
}
// ========== Default Configuration ==========
/**
* Default application configuration values
*/
export const DEFAULT_CONFIG: AppConfig = {
pollIntervalMs: 1000,
defaultTimeoutMs: 300000, // 5 minutes
maxConcurrentSessions: 5,
stateFilePath: '',
respawn: {
idleTimeoutMs: 5000, // 5 seconds of no activity after prompt
updatePrompt: 'update all the docs and CLAUDE.md',
interStepDelayMs: 1000, // 1 second between steps
enabled: false, // disabled by default
sendClear: true, // send /clear after update prompt
sendInit: true, // send /init after /clear
},
lastUsedCase: null,
ralphEnabled: false,
};
/**
* Creates initial application state
* @returns Fresh application state with defaults
*/
export function createInitialState(): AppState {
return {
sessions: {},
tasks: {},
ralphLoop: {
status: 'stopped',
startedAt: null,
minDurationMs: null,
tasksCompleted: 0,
tasksGenerated: 0,
lastCheckAt: null,
},
config: { ...DEFAULT_CONFIG },
globalStats: createInitialGlobalStats(),
};
}
/**
* Creates initial global stats object
* @returns Fresh global stats with zero values
*/
export function createInitialGlobalStats(): GlobalStats {
const now = Date.now();
return {
totalInputTokens: 0,
totalOutputTokens: 0,
totalCost: 0,
totalSessionsCreated: 0,
firstRecordedAt: now,
lastUpdatedAt: now,
};
}
+88
View File
@@ -0,0 +1,88 @@
/**
* @fileoverview Common/shared type definitions.
*
* Base types used across multiple domains. No dependencies on other domain modules.
*
* Key exports:
* - Disposable — interface for objects requiring explicit cleanup (timers, watchers)
* - BufferConfig — size-limited storage config (terminal: 2MB, text: 1MB)
* - CleanupRegistration / CleanupResourceType — entries for the centralized CleanupManager
* - NiceConfig / DEFAULT_NICE_CONFIG — process priority settings for `nice`/`ionice`
* - ProcessStats — memory/CPU/child-count snapshot for resource monitoring
*/
/**
* Interface for objects that hold resources requiring explicit cleanup.
* Implementing classes should release timers, watchers, and other resources in dispose().
*/
export interface Disposable {
/** Release all held resources. Safe to call multiple times. */
dispose(): void;
/** Whether this object has been disposed */
readonly isDisposed: boolean;
}
/**
* Configuration for buffer accumulator instances.
* Used for terminal buffers, text output, and other size-limited string storage.
*/
export interface BufferConfig {
/** Maximum buffer size in bytes before trimming */
maxSize: number;
/** Size to trim to when maxSize is exceeded */
trimSize: number;
/** Optional callback invoked when buffer is trimmed */
onTrim?: (trimmedBytes: number) => void;
}
/**
* Resource types that can be registered for cleanup.
*/
/**
* Configuration for process priority using `nice`.
* Lower priority reduces CPU contention with other processes.
*/
export interface NiceConfig {
/** Whether nice priority is enabled */
enabled: boolean;
/** Nice value (-20 to 19, default: 10 = lower priority) */
niceValue: number;
}
export const DEFAULT_NICE_CONFIG: NiceConfig = {
enabled: false,
niceValue: 10,
};
/**
* Process resource statistics
*/
export interface ProcessStats {
/** Memory usage in megabytes */
memoryMB: number;
/** CPU usage percentage */
cpuPercent: number;
/** Number of child processes */
childCount: number;
/** Timestamp of stats collection */
updatedAt: number;
}
export type CleanupResourceType = 'timer' | 'interval' | 'watcher' | 'listener' | 'stream';
/**
* Registration entry for a cleanup resource.
* Used by CleanupManager to track and dispose resources.
*/
export interface CleanupRegistration {
/** Unique identifier for this registration */
id: string;
/** Type of resource */
type: CleanupResourceType;
/** Human-readable description for debugging */
description: string;
/** Cleanup function to call on dispose */
cleanup: () => void;
/** Timestamp when registered */
registeredAt: number;
}
+66
View File
@@ -0,0 +1,66 @@
/**
* @fileoverview Barrel re-export for all Codeman type definitions.
*
* The type system is split into 13 domain modules for maintainability.
* Import from `'./types'` (or `'./types/index.js'`) to access any type:
*
* ```ts
* import type { SessionState, AppState, RespawnConfig } from './types';
* import { createErrorResponse, ApiErrorCode } from './types';
* ```
*
* ## Domain modules
*
* | Module | Key exports | Persistence / API |
* |--------------|-----------------------------------------------------------------------|-------------------------------------------------|
* | common | Disposable, BufferConfig, CleanupRegistration, NiceConfig, ProcessStats | In-memory only |
* | session | SessionState, SessionConfig, SessionStatus, SessionMode, ClaudeMode, OpenCodeConfig | `~/.codeman/state.json` → `GET /api/sessions` |
* | task | TaskDefinition, TaskState, TaskStatus | `~/.codeman/state.json` → `GET /api/tasks` |
* | app-state | AppState, AppConfig, GlobalStats, TokenStats, DEFAULT_CONFIG | `~/.codeman/state.json` → `GET /api/status` |
* | respawn | RespawnConfig, RespawnPreset, RespawnCycleMetrics, RalphLoopHealthScore, TimingHistory | Per-session in state.json → `GET /api/sessions/:id/respawn` |
* | ralph | RalphTrackerState, RalphTodoItem, CircuitBreakerStatus, RalphStatusBlock, RalphSessionState | Per-session → `GET /api/sessions/:id/ralph-state` |
* | api | ApiResponse, ApiErrorCode, HookEventType, CaseInfo, createErrorResponse, getErrorMessage | Used by all route handlers |
* | lifecycle | LifecycleEntry, LifecycleEventType | `~/.codeman/session-lifecycle.jsonl` (append-only) |
* | run-summary | RunSummary, RunSummaryEvent, RunSummaryStats | In-memory → `GET /api/sessions/:id/run-summary` |
* | tools | ActiveBashTool, ImageDetectedEvent | In-memory, broadcast via SSE |
* | teams | TeamConfig, TeamMember, TeamTask, InboxMessage, PaneInfo | `~/.claude/teams/`, `~/.claude/tasks/` → `GET /api/teams` |
* | push | PushSubscriptionRecord, VapidKeys | `~/.codeman/push-keys.json`, `~/.codeman/push-subscriptions.json` |
* | plan | PlanItem, PlanTaskStatus, TddPhase | In-memory → `GET /api/sessions/:id/plan/tasks` |
*
* ## Cross-domain relationship map
*
* ```
* AppState (app-state)
* ├── sessions: Record<id, SessionState> ← session domain
* │ ├── respawnConfig?: RespawnConfig ← respawn domain (per-session settings)
* │ ├── ralphEnabled?: boolean ← toggles ralph tracking
* │ └── id ← referenced by:
* │ ├── RalphSessionState.sessionId ← ralph domain
* │ ├── RunSummary.sessionId ← run-summary domain
* │ ├── ActiveBashTool.sessionId ← tools domain
* │ ├── RespawnCycleMetrics.sessionId ← respawn domain
* │ └── TeamConfig.leadSessionId ← teams domain
* ├── tasks: Record<id, TaskState> ← task domain
* │ └── assignedSessionId → SessionState.id
* ├── ralphLoop: RalphLoopState ← ralph domain (global loop state)
* └── config: AppConfig
* └── respawn: RespawnConfig ← respawn domain (global defaults)
*
* RalphLoopHealthScore (respawn)
* └── components.circuitBreaker ← derived from CircuitBreakerStatus (ralph)
* ```
*/
export * from './common.js';
export * from './session.js';
export * from './task.js';
export * from './app-state.js';
export * from './respawn.js';
export * from './ralph.js';
export * from './api.js';
export * from './lifecycle.js';
export * from './run-summary.js';
export * from './tools.js';
export * from './teams.js';
export * from './push.js';
export * from './plan.js';
+40
View File
@@ -0,0 +1,40 @@
/**
* @fileoverview Session lifecycle audit types.
*
* Types for the append-only JSONL audit log at `~/.codeman/session-lifecycle.jsonl`.
* Records session creation, PTY launch/exit, server start/stop, tmux recovery,
* and QR auth events. Written by `SessionLifecycleLog`, read for debugging.
*
* Key exports:
* - LifecycleEventType — union of 11 event types (created, started, exit, qr_auth, etc.)
* - LifecycleEntry — a single timestamped log entry with event, sessionId, and optional metadata
*
* No dependencies on other domain modules. LifecycleEntry.sessionId links
* back to SessionState.id for correlation.
*/
/** Types of session lifecycle events recorded to the audit log */
export type LifecycleEventType =
| 'created' // Session object created
| 'started' // PTY process launched (interactive/shell/prompt)
| 'exit' // PTY process exited (with exit code)
| 'deleted' // cleanupSession() called — session removed
| 'detached' // Server shutdown — PTY left alive in tmux for recovery
| 'recovered' // Session restored from tmux on server restart
| 'stale_cleaned' // Removed from state.json by cleanupStaleSessions()
| 'mux_died' // tmux session died (detected by reconciliation)
| 'server_started' // Server started (marker for restart detection)
| 'server_stopped' // Server shutting down
| 'qr_auth'; // Device authenticated via QR code scan
/** A single entry in the session lifecycle audit log */
export interface LifecycleEntry {
ts: number;
event: LifecycleEventType;
sessionId: string;
name?: string;
mode?: string;
reason?: string;
exitCode?: number | null;
extra?: Record<string, unknown>;
}
+45
View File
@@ -0,0 +1,45 @@
/**
* @fileoverview Plan orchestrator type definitions.
*
* Types for the 2-agent plan generation system (optional research agent → planner agent).
*
* Key exports:
* - PlanItem — a single task with priority (P0/P1/P2), TDD phase, dependencies, verification criteria
* - PlanTaskStatus — 'pending' | 'in_progress' | 'completed' | 'failed' | 'blocked'
* - TddPhase — 'setup' | 'test' | 'impl' | 'verify' | 'review'
*
* Used by PlanOrchestrator (`src/plan-orchestrator.ts`) and the plan API routes
* (`src/web/routes/plan-routes.ts`). Served at `GET /api/sessions/:id/plan/tasks`.
*
* PlanItem was moved here from plan-orchestrator.ts to break a circular dependency.
* No dependencies on other domain modules.
*/
/** Task execution status for plan tracking */
export type PlanTaskStatus = 'pending' | 'in_progress' | 'completed' | 'failed' | 'blocked';
/** TDD phase categories */
export type TddPhase = 'setup' | 'test' | 'impl' | 'verify' | 'review';
/**
* A single plan item for plan orchestration.
* Moved here from plan-orchestrator.ts to break circular dependency.
*/
export interface PlanItem {
id?: string;
content: string;
priority: 'P0' | 'P1' | 'P2' | null;
source?: string;
rationale?: string;
verificationCriteria?: string;
testCommand?: string;
dependencies?: string[];
status?: PlanTaskStatus;
attempts?: number;
lastError?: string;
completedAt?: number;
complexity?: 'low' | 'medium' | 'high';
tddPhase?: TddPhase;
pairedWith?: string;
reviewChecklist?: string[];
}
+34
View File
@@ -0,0 +1,34 @@
/**
* @fileoverview Web Push notification type definitions.
*
* Types for the Web Push notification layer (layer 4 of the 5-layer notification system).
*
* Key exports:
* - PushSubscriptionRecord — a registered push endpoint with per-event preferences
* - VapidKeys — VAPID key pair (public + private) for Web Push authentication
*
* Persistence:
* - VAPID keys: `~/.codeman/push-keys.json` (auto-generated on first use)
* - Subscriptions: `~/.codeman/push-subscriptions.json` (expired auto-cleaned on 410/404)
*
* Managed by PushStore (`src/push-store.ts`). Served at `GET /api/push/vapid-key`,
* `POST /api/push/subscribe`. No dependencies on other domain modules.
*/
/** A registered push subscription */
export interface PushSubscriptionRecord {
id: string;
endpoint: string;
keys: { p256dh: string; auth: string };
userAgent: string;
createdAt: number;
lastUsedAt: number;
pushPreferences: Record<string, boolean>;
}
/** VAPID key pair for Web Push */
export interface VapidKeys {
publicKey: string;
privateKey: string;
generatedAt: number;
}
+321
View File
@@ -0,0 +1,321 @@
/**
* @fileoverview Ralph Loop / todo tracking type definitions.
*
* Covers the autonomous task execution system: loop state, todo items,
* completion confidence scoring, RALPH_STATUS block parsing, and the
* circuit breaker for stuck-loop detection.
*
* Key exports:
* - RalphLoopState / RalphLoopStatus — global loop controller state (embedded in AppState)
* - RalphTrackerState — per-session loop tracking (cycle count, completion phrase, plan version)
* - RalphTodoItem / RalphTodoProgress — detected todo items with priority and progress estimation
* - RalphSessionState — composite per-session state (loop + todos), linked via sessionId
* - CompletionConfidence — multi-signal scoring for completion detection (0-100)
* - RalphStatusBlock — parsed RALPH_STATUS block from Claude output (status, tests, exit signal)
* - CircuitBreakerStatus / CircuitBreakerState — stuck-loop detection state machine (CLOSED → HALF_OPEN → OPEN)
* - Factory functions: createInitialCircuitBreakerStatus(), createInitialRalphTrackerState(), createInitialRalphSessionState()
*
* Cross-domain relationships:
* - RalphLoopState is embedded in AppState.ralphLoop (app-state domain)
* - RalphSessionState.sessionId links to SessionState.id (session domain)
* - CircuitBreakerStatus feeds into RalphLoopHealthScore.components.circuitBreaker (respawn domain)
*
* Served at `GET /api/sessions/:id/ralph-state` and `GET /api/sessions/:id/ralph-status`.
*/
/** Status of the Ralph Loop controller */
export type RalphLoopStatus = 'stopped' | 'running' | 'paused';
/**
* State of the Ralph Loop controller
*/
export interface RalphLoopState {
/** Current loop status */
status: RalphLoopStatus;
/** Timestamp when loop started */
startedAt: number | null;
/** Minimum duration to run in milliseconds */
minDurationMs: number | null;
/** Number of tasks completed in this run */
tasksCompleted: number;
/** Number of tasks auto-generated */
tasksGenerated: number;
/** Timestamp of last status check */
lastCheckAt: number | null;
}
/** Status of a detected todo item */
export type RalphTodoStatus = 'pending' | 'in_progress' | 'completed';
/**
* Confidence scoring for completion detection.
* Helps distinguish genuine completion signals from false positives.
*/
export interface CompletionConfidence {
/** Overall confidence level (0-100) */
score: number;
/** Whether score is above threshold for triggering completion */
isConfident: boolean;
/** Individual signal contributions */
signals: {
/** Promise tag detected with proper formatting */
hasPromiseTag: boolean;
/** Phrase matches expected completion phrase */
matchesExpected: boolean;
/** All todos are marked complete */
allTodosComplete: boolean;
/** EXIT_SIGNAL: true in RALPH_STATUS block */
hasExitSignal: boolean;
/** Multiple completion indicators present */
multipleIndicators: boolean;
/** Output context suggests completion (not in prompt/explanation) */
contextAppropriate: boolean;
};
/** Timestamp of last confidence calculation */
calculatedAt: number;
}
export interface RalphTrackerState {
/** Whether the tracker is actively monitoring (disabled by default) */
enabled: boolean;
/** Whether a loop is currently active */
active: boolean;
/** Detected completion phrase (primary) */
completionPhrase: string | null;
/** Additional valid completion phrases (P1-003: multi-phrase support) */
alternateCompletionPhrases?: string[];
/** Timestamp when loop started */
startedAt: number | null;
/** Number of cycles/iterations detected */
cycleCount: number;
/** Maximum iterations if detected */
maxIterations: number | null;
/** Timestamp of last activity */
lastActivity: number;
/** Elapsed hours if detected */
elapsedHours: number | null;
/** Current plan version (for versioning UI) */
planVersion?: number;
/** Number of versions in history (for versioning UI) */
planHistoryLength?: number;
/** Last completion confidence assessment */
completionConfidence?: CompletionConfidence;
}
/**
* Priority levels for todo items.
* Matches @fix_plan.md format (P0=critical, P1=high, P2=normal).
*/
export type RalphTodoPriority = 'P0' | 'P1' | 'P2' | null;
/**
* A detected todo item from Claude Code output
*/
export interface RalphTodoItem {
/** Unique identifier based on content hash */
id: string;
/** Todo item text content */
content: string;
/** Current status */
status: RalphTodoStatus;
/** Timestamp when detected */
detectedAt: number;
/** Priority level (P0=critical, P1=high, P2=normal) */
priority: RalphTodoPriority;
/** P1-009: Estimated time to complete (ms), based on historical patterns */
estimatedDurationMs?: number;
/** P1-009: Complexity category for progress estimation */
estimatedComplexity?: 'trivial' | 'simple' | 'moderate' | 'complex';
}
/**
* Progress estimation for the todo list
*/
export interface RalphTodoProgress {
/** Total number of todos */
total: number;
/** Number completed */
completed: number;
/** Number in progress */
inProgress: number;
/** Number pending */
pending: number;
/** Completion percentage (0-100) */
percentComplete: number;
/** Estimated remaining time (ms), based on historical completion rate */
estimatedRemainingMs: number | null;
/** Average time per todo completion (ms) */
avgCompletionTimeMs: number | null;
/** Projected completion timestamp (epoch ms) */
projectedCompletionAt: number | null;
}
/**
* Complete Ralph/todo state for a session
*/
export interface RalphSessionState {
/** Session this state belongs to */
sessionId: string;
/** Loop tracking state */
loop: RalphTrackerState;
/** Detected todo items */
todos: RalphTodoItem[];
/** Timestamp of last update */
lastUpdated: number;
}
// ========== RALPH_STATUS Block Types ==========
/**
* Status values from RALPH_STATUS block.
* - IN_PROGRESS: Work is ongoing
* - COMPLETE: All tasks finished
* - BLOCKED: Needs human intervention
*/
export type RalphStatusValue = 'IN_PROGRESS' | 'COMPLETE' | 'BLOCKED';
/**
* Test status from RALPH_STATUS block.
*/
export type RalphTestsStatus = 'PASSING' | 'FAILING' | 'NOT_RUN';
/**
* Work type classification for current iteration.
*/
export type RalphWorkType = 'IMPLEMENTATION' | 'TESTING' | 'DOCUMENTATION' | 'REFACTORING';
/**
* Parsed RALPH_STATUS block from Claude output.
*
* Claude outputs this at the end of every response:
* ```
* ---RALPH_STATUS---
* STATUS: IN_PROGRESS
* TASKS_COMPLETED_THIS_LOOP: 3
* FILES_MODIFIED: 5
* TESTS_STATUS: PASSING
* WORK_TYPE: IMPLEMENTATION
* EXIT_SIGNAL: false
* RECOMMENDATION: Continue with database migration
* ---END_RALPH_STATUS---
* ```
*/
export interface RalphStatusBlock {
/** Overall loop status */
status: RalphStatusValue;
/** Number of tasks completed in current iteration */
tasksCompletedThisLoop: number;
/** Number of files modified in current iteration */
filesModified: number;
/** Current state of tests */
testsStatus: RalphTestsStatus;
/** Type of work being performed */
workType: RalphWorkType;
/** Whether Claude is signaling completion */
exitSignal: boolean;
/** Claude's recommendation for next steps */
recommendation: string;
/** Timestamp when this block was parsed */
parsedAt: number;
}
// ========== Circuit Breaker Types ==========
/**
* Circuit breaker states for detecting stuck loops.
* - CLOSED: Normal operation, all checks passing
* - HALF_OPEN: Warning state, some checks failing
* - OPEN: Loop is stuck, requires intervention
*/
export type CircuitBreakerState = 'CLOSED' | 'HALF_OPEN' | 'OPEN';
/**
* Reason codes for circuit breaker state transitions.
*/
export type CircuitBreakerReason =
| 'normal_operation'
| 'no_progress_warning'
| 'no_progress_open'
| 'same_error_repeated'
| 'tests_failing_too_long'
| 'progress_detected'
| 'manual_reset';
/**
* Circuit breaker status for tracking loop health.
*
* Transitions:
* - CLOSED -> HALF_OPEN: consecutive_no_progress >= 2
* - CLOSED -> OPEN: consecutive_no_progress >= 3 OR consecutive_same_error >= 5
* - HALF_OPEN -> CLOSED: progress detected
* - HALF_OPEN -> OPEN: consecutive_no_progress >= 3
* - OPEN -> CLOSED: manual reset only
*/
export interface CircuitBreakerStatus {
/** Current state of the circuit breaker */
state: CircuitBreakerState;
/** Number of consecutive iterations with no progress */
consecutiveNoProgress: number;
/** Number of consecutive iterations with the same error */
consecutiveSameError: number;
/** Number of consecutive iterations with failing tests */
consecutiveTestsFailure: number;
/** Last iteration number that showed progress */
lastProgressIteration: number;
/** Human-readable reason for current state */
reason: string;
/** Reason code for programmatic handling */
reasonCode: CircuitBreakerReason;
/** Timestamp of last state transition */
lastTransitionAt: number;
/** Last error message seen (for same-error tracking) */
lastErrorMessage: string | null;
}
/**
* Creates initial circuit breaker status.
*/
export function createInitialCircuitBreakerStatus(): CircuitBreakerStatus {
return {
state: 'CLOSED',
consecutiveNoProgress: 0,
consecutiveSameError: 0,
consecutiveTestsFailure: 0,
lastProgressIteration: 0,
reason: 'Initial state',
reasonCode: 'normal_operation',
lastTransitionAt: Date.now(),
lastErrorMessage: null,
};
}
/**
* Creates initial Ralph tracker state
* @returns Fresh Ralph tracker state with defaults
*/
export function createInitialRalphTrackerState(): RalphTrackerState {
return {
enabled: false, // Disabled by default, auto-enables when Ralph patterns detected
active: false,
completionPhrase: null,
startedAt: null,
cycleCount: 0,
maxIterations: null,
lastActivity: Date.now(),
elapsedHours: null,
};
}
/**
* Creates initial Ralph session state
* @param sessionId Session ID this state belongs to
* @returns Fresh Ralph session state
*/
export function createInitialRalphSessionState(sessionId: string): RalphSessionState {
return {
sessionId,
loop: createInitialRalphTrackerState(),
todos: [],
lastUpdated: Date.now(),
};
}
+299
View File
@@ -0,0 +1,299 @@
/**
* @fileoverview Respawn controller type definitions.
*
* Covers the autonomous session cycling system: configuration, presets,
* per-cycle metrics, aggregate health scoring, and adaptive timing.
*
* Key exports:
* - RespawnConfig — full respawn settings (idle timeout, AI checks, adaptive timing, skip-clear)
* - PersistedRespawnConfig — subset saved to disk for mux session recovery
* - RespawnPreset — named preset for quick setup (solo-work, team-lead, overnight-autonomous, etc.)
* - RespawnCycleMetrics — per-cycle outcome tracking (duration, idle reason, steps, tokens)
* - RespawnAggregateMetrics — aggregate stats across cycles (success rate, p90 duration)
* - RalphLoopHealthScore — composite 0-100 health score with 5 component scores
* - HealthStatus — 'excellent' | 'good' | 'degraded' | 'critical'
* - TimingHistory — rolling window of timing data for adaptive adjustments
* - CycleOutcome — 'success' | 'stuck_recovery' | 'blocked' | 'error' | 'cancelled'
*
* Cross-domain relationships:
* - RespawnConfig is embedded in AppConfig.respawn (app-state) and SessionState.respawnConfig (session)
* - RalphLoopHealthScore.components.circuitBreaker derives from CircuitBreakerStatus (ralph domain)
* - RespawnCycleMetrics.sessionId links to SessionState.id (session domain)
*
* Served at `GET /api/sessions/:id/respawn` (config + state).
*/
/**
* Configuration for the Respawn Controller
*
* The respawn controller keeps interactive sessions productive by
* automatically cycling through update prompts when Claude goes idle.
*/
export interface RespawnConfig {
/** How long to wait after seeing prompt before considering truly idle (ms) */
idleTimeoutMs: number;
/** The prompt to send for updating docs */
updatePrompt: string;
/** Delay between sending steps (ms) */
interStepDelayMs: number;
/** Whether to enable respawn loop */
enabled: boolean;
/** Whether to send /clear after update prompt */
sendClear: boolean;
/** Whether to send /init after /clear */
sendInit: boolean;
/** Optional prompt to send if /init doesn't trigger work */
kickstartPrompt?: string;
/** Time to wait after completion message before confirming idle (ms) */
completionConfirmMs?: number;
/** Fallback timeout when no output received at all (ms) */
noOutputTimeoutMs?: number;
/** Whether to auto-accept plan mode prompts by pressing Enter (not questions) */
autoAcceptPrompts?: boolean;
/** Delay before auto-accepting plan mode prompts when no output and no completion message (ms) */
autoAcceptDelayMs?: number;
/** Whether AI idle check is enabled */
aiIdleCheckEnabled?: boolean;
/** Model to use for AI idle check */
aiIdleCheckModel?: string;
/** Maximum characters of terminal buffer for AI check */
aiIdleCheckMaxContext?: number;
/** Timeout for AI check in ms */
aiIdleCheckTimeoutMs?: number;
/** Cooldown after WORKING verdict in ms */
aiIdleCheckCooldownMs?: number;
/** Whether AI plan mode check is enabled for auto-accept */
aiPlanCheckEnabled?: boolean;
/** Model to use for AI plan mode check */
aiPlanCheckModel?: string;
/** Maximum characters of terminal buffer for plan check */
aiPlanCheckMaxContext?: number;
/** Timeout for AI plan check in ms */
aiPlanCheckTimeoutMs?: number;
/** Cooldown after NOT_PLAN_MODE verdict in ms */
aiPlanCheckCooldownMs?: number;
// ========== P2-001: Adaptive Timing ==========
/** Whether to use adaptive timing based on historical patterns */
adaptiveTimingEnabled?: boolean;
/** Minimum value for adaptive completion confirm (ms) */
adaptiveMinConfirmMs?: number;
/** Maximum value for adaptive completion confirm (ms) */
adaptiveMaxConfirmMs?: number;
// ========== P2-002: Skip-Clear Optimization ==========
/** Whether to skip /clear when context is below threshold */
skipClearWhenLowContext?: boolean;
/** Token percentage threshold below which /clear is skipped (0-100) */
skipClearThresholdPercent?: number;
// ========== P2-004: Cycle Metrics ==========
/** Whether to track and persist cycle metrics */
trackCycleMetrics?: boolean;
}
// ========== P2-004: Respawn Cycle Metrics ==========
/**
* Outcome of a respawn cycle
*/
export type CycleOutcome =
| 'success' // Cycle completed normally
| 'stuck_recovery' // Stuck-state recovery triggered
| 'blocked' // Blocked by circuit breaker or exit signal
| 'error' // Error during cycle
| 'cancelled'; // Cancelled (e.g., controller stopped)
/**
* Metrics for a single respawn cycle.
* Persisted for post-mortem analysis of long-running loops.
*/
export interface RespawnCycleMetrics {
/** Unique cycle ID (session-id:cycle-number) */
cycleId: string;
/** Session ID this cycle belongs to */
sessionId: string;
/** Cycle number within the session */
cycleNumber: number;
/** Timestamp when cycle started */
startedAt: number;
/** Timestamp when cycle completed */
completedAt: number;
/** Total duration of cycle (ms) */
durationMs: number;
/** What triggered idle detection */
idleReason: string;
/** Time spent detecting idle (from start of watching to idle confirmed) */
idleDetectionMs: number;
/** Steps completed in this cycle */
stepsCompleted: string[];
/** Whether /clear was skipped (P2-002) */
clearSkipped: boolean;
/** Outcome of the cycle */
outcome: CycleOutcome;
/** Error message if outcome is 'error' */
errorMessage?: string;
/** Token count at start of cycle */
tokenCountAtStart?: number;
/** Token count at end of cycle */
tokenCountAtEnd?: number;
/** Completion confirm time used (may be adaptive) */
completionConfirmMsUsed: number;
}
/**
* Aggregate metrics across multiple cycles for health scoring.
*/
export interface RespawnAggregateMetrics {
/** Total cycles tracked */
totalCycles: number;
/** Successful cycles */
successfulCycles: number;
/** Cycles that required stuck-state recovery */
stuckRecoveryCycles: number;
/** Blocked cycles */
blockedCycles: number;
/** Error cycles */
errorCycles: number;
/** Average cycle duration (ms) */
avgCycleDurationMs: number;
/** Average idle detection time (ms) */
avgIdleDetectionMs: number;
/** 90th percentile cycle duration (ms) */
p90CycleDurationMs: number;
/** Success rate (0-100) */
successRate: number;
/** Last updated timestamp */
lastUpdatedAt: number;
}
// ========== P2-005: Ralph Loop Health Score ==========
/**
* Health status levels for the Ralph Loop system.
*/
export type HealthStatus = 'excellent' | 'good' | 'degraded' | 'critical';
/**
* Comprehensive health score for a Ralph Loop session.
* Aggregates multiple health signals into a single score.
*/
export interface RalphLoopHealthScore {
/** Overall health score (0-100) */
score: number;
/** Health status based on score thresholds */
status: HealthStatus;
/** Individual component scores (0-100 each) */
components: {
/** Based on recent cycle success rate */
cycleSuccess: number;
/** Based on circuit breaker state */
circuitBreaker: number;
/** Based on iteration stall metrics */
iterationProgress: number;
/** Based on AI checker error rate */
aiChecker: number;
/** Based on stuck-state recovery count */
stuckRecovery: number;
};
/** Human-readable summary of health */
summary: string;
/** Recommendations for improvement */
recommendations: string[];
/** Timestamp when score was calculated */
calculatedAt: number;
}
// ========== Timing History for Adaptive Timing ==========
/**
* Historical timing data for adaptive adjustments.
*/
export interface TimingHistory {
/** Rolling window of recent idle detection durations (ms) */
recentIdleDetectionMs: number[];
/** Rolling window of recent cycle durations (ms) */
recentCycleDurationMs: number[];
/** Calculated adaptive completion confirm value (ms) */
adaptiveCompletionConfirmMs: number;
/** Number of samples in rolling windows */
sampleCount: number;
/** Maximum samples to keep */
maxSamples: number;
/** Last updated timestamp */
lastUpdatedAt: number;
}
/**
* Named respawn configuration preset for quick setup
*/
export interface RespawnPreset {
/** Unique preset identifier */
id: string;
/** User-friendly preset name */
name: string;
/** Description of when to use this preset */
description?: string;
/** The respawn configuration (without enabled flag) */
config: Omit<RespawnConfig, 'enabled'>;
/** Duration in minutes (optional default) */
durationMinutes?: number;
/** Whether this is a built-in preset */
builtIn?: boolean;
/** Timestamp when created */
createdAt: number;
}
/**
* Persisted respawn configuration for mux sessions.
* Subset of RespawnConfig that gets saved to disk.
*/
export interface PersistedRespawnConfig {
/** Whether respawn was enabled */
enabled: boolean;
/** How long to wait after seeing prompt before considering truly idle (ms) */
idleTimeoutMs: number;
/** The prompt to send for updating docs */
updatePrompt: string;
/** Delay between sending steps (ms) */
interStepDelayMs: number;
/** Whether to send /clear after update prompt */
sendClear: boolean;
/** Whether to send /init after /clear */
sendInit: boolean;
/** Optional prompt to send if /init doesn't trigger work */
kickstartPrompt?: string;
/** Whether to auto-accept plan mode prompts by pressing Enter (not questions) */
autoAcceptPrompts?: boolean;
/** Delay before auto-accepting prompts (ms) */
autoAcceptDelayMs?: number;
/** Time to wait after completion message before confirming idle (ms) */
completionConfirmMs?: number;
/** Fallback timeout when no output received at all (ms) */
noOutputTimeoutMs?: number;
/** Whether AI idle check is enabled */
aiIdleCheckEnabled?: boolean;
/** Model to use for AI idle check */
aiIdleCheckModel?: string;
/** Maximum characters of terminal buffer for AI check */
aiIdleCheckMaxContext?: number;
/** Timeout for AI check in ms */
aiIdleCheckTimeoutMs?: number;
/** Cooldown after WORKING verdict in ms */
aiIdleCheckCooldownMs?: number;
/** Whether AI plan mode check is enabled for auto-accept */
aiPlanCheckEnabled?: boolean;
/** Model to use for AI plan mode check */
aiPlanCheckModel?: string;
/** Maximum characters of terminal buffer for plan check */
aiPlanCheckMaxContext?: number;
/** Timeout for AI plan check in ms */
aiPlanCheckTimeoutMs?: number;
/** Cooldown after NOT_PLAN_MODE verdict in ms */
aiPlanCheckCooldownMs?: number;
/** Duration in minutes if timed respawn was set */
durationMinutes?: number;
}
+131
View File
@@ -0,0 +1,131 @@
/**
* @fileoverview Run summary type definitions.
*
* Types for the "what happened while away" session timeline. RunSummary
* aggregates events (respawn cycles, errors, token milestones, AI checks)
* into a per-session historical view.
*
* Key exports:
* - RunSummary — complete per-session summary (events timeline + aggregated stats)
* - RunSummaryEvent — a single timestamped event (16 event types, 4 severity levels)
* - RunSummaryStats — aggregated metrics (cycles, tokens, active/idle time, error count)
* - RunSummaryEventType — union of event types (session_started, respawn_cycle_*, error, etc.)
* - RunSummaryEventSeverity — 'info' | 'warning' | 'error' | 'success'
* - createInitialRunSummaryStats() — factory for fresh stats
*
* Cross-domain: RunSummary.sessionId links to SessionState.id (session domain).
* In-memory only (not persisted to disk). Served at `GET /api/sessions/:id/run-summary`.
*/
/**
* Types of events tracked in the run summary.
* These provide a historical view of what happened during a session.
*/
export type RunSummaryEventType =
| 'session_started'
| 'session_stopped'
| 'respawn_cycle_started'
| 'respawn_cycle_completed'
| 'respawn_state_change'
| 'error'
| 'warning'
| 'token_milestone'
| 'auto_compact'
| 'auto_clear'
| 'idle_detected'
| 'working_detected'
| 'ralph_completion'
| 'ai_check_result'
| 'hook_event'
| 'state_stuck';
/**
* Severity levels for run summary events.
*/
export type RunSummaryEventSeverity = 'info' | 'warning' | 'error' | 'success';
/**
* A single event in the run summary timeline.
*/
export interface RunSummaryEvent {
/** Unique event identifier */
id: string;
/** Timestamp when event occurred */
timestamp: number;
/** Type of event */
type: RunSummaryEventType;
/** Severity level for display */
severity: RunSummaryEventSeverity;
/** Short title for the event */
title: string;
/** Optional detailed description */
details?: string;
/** Optional additional metadata */
metadata?: Record<string, unknown>;
}
/**
* Statistics aggregated from run summary events.
*/
export interface RunSummaryStats {
/** Number of respawn cycles completed */
totalRespawnCycles: number;
/** Total tokens used during this run */
totalTokensUsed: number;
/** Peak token count observed */
peakTokens: number;
/** Total time Claude was actively working (ms) */
totalTimeActiveMs: number;
/** Total time Claude was idle (ms) */
totalTimeIdleMs: number;
/** Number of errors encountered */
errorCount: number;
/** Number of warnings encountered */
warningCount: number;
/** Number of AI idle checks performed */
aiCheckCount: number;
/** Timestamp when last became idle */
lastIdleAt: number | null;
/** Timestamp when last started working */
lastWorkingAt: number | null;
/** Total number of state transitions */
stateTransitions: number;
}
/**
* Complete run summary for a session.
* Provides a historical view of session activity for users returning after absence.
*/
export interface RunSummary {
/** Session ID this summary belongs to */
sessionId: string;
/** Session display name */
sessionName: string;
/** Timestamp when tracking started */
startedAt: number;
/** Timestamp of last update */
lastUpdatedAt: number;
/** Timeline of events (most recent last) */
events: RunSummaryEvent[];
/** Aggregated statistics */
stats: RunSummaryStats;
}
/**
* Creates initial run summary stats.
*/
export function createInitialRunSummaryStats(): RunSummaryStats {
return {
totalRespawnCycles: 0,
totalTokensUsed: 0,
peakTokens: 0,
totalTimeActiveMs: 0,
totalTimeIdleMs: 0,
errorCount: 0,
warningCount: 0,
aiCheckCount: 0,
lastIdleAt: null,
lastWorkingAt: null,
stateTransitions: 0,
};
}
+160
View File
@@ -0,0 +1,160 @@
/**
* @fileoverview Session type definitions.
*
* Core domain type — SessionState is the primary entity in the system.
*
* Key exports:
* - SessionState — full session state (status, tokens, respawn, ralph, CLI metadata)
* - SessionConfig — creation-time config (id, workingDir, createdAt)
* - SessionOutput — captured stdout/stderr/exitCode
* - SessionStatus — 'idle' | 'busy' | 'stopped' | 'error'
* - SessionMode — 'claude' | 'shell' | 'opencode' (which CLI backend)
* - ClaudeMode — CLI permission mode ('dangerously-skip-permissions' | 'normal' | 'allowedTools')
* - SessionColor — visual differentiation color
* - OpenCodeConfig — OpenCode-specific settings (model, autoAllowTools, continueSession)
*
* Cross-domain relationships:
* - SessionState.respawnConfig embeds RespawnConfig (respawn domain)
* - SessionState.id is referenced by: RalphSessionState.sessionId (ralph),
* RunSummary.sessionId (run-summary), ActiveBashTool.sessionId (tools),
* TeamConfig.leadSessionId (teams), RespawnCycleMetrics.sessionId (respawn),
* TaskState.assignedSessionId (task)
*
* Persisted to `~/.codeman/state.json`. Served at `GET /api/sessions` and
* `GET /api/sessions/:id`.
*/
import type { RespawnConfig } from './respawn.js';
/** Status of a Claude session */
export type SessionStatus = 'idle' | 'busy' | 'stopped' | 'error';
/**
* Claude CLI startup permission mode.
* - `'dangerously-skip-permissions'`: Bypass all permission prompts (default)
* - `'normal'`: Standard mode with permission prompts
* - `'allowedTools'`: Only allow specific tools (requires allowedTools list)
*/
export type ClaudeMode = 'dangerously-skip-permissions' | 'normal' | 'allowedTools';
/** Session mode: which CLI backend a session runs */
export type SessionMode = 'claude' | 'shell' | 'opencode';
/** OpenCode session configuration */
export interface OpenCodeConfig {
/** Model identifier (e.g., "anthropic/claude-sonnet-4-5", "openai/gpt-5.2", "ollama/codellama") */
model?: string;
/** Whether to auto-allow all tool executions (sets permission.* = allow) */
autoAllowTools?: boolean;
/** Session ID to continue from */
continueSession?: string;
/** Whether to fork when continuing (branch the conversation) */
forkSession?: boolean;
/** Custom inline config JSON (passed via OPENCODE_CONFIG_CONTENT) */
configContent?: string;
}
/**
* Configuration for creating a new session
*/
export interface SessionConfig {
/** Unique session identifier */
id: string;
/** Working directory for the session */
workingDir: string;
/** Timestamp when session was created */
createdAt: number;
}
/**
* Available session colors for visual differentiation
*/
export type SessionColor = 'default' | 'red' | 'orange' | 'yellow' | 'green' | 'blue' | 'purple' | 'pink';
/**
* Current state of a session
*/
export interface SessionState {
/** Unique session identifier */
id: string;
/** Process ID of the PTY process, null if not running */
pid: number | null;
/** Current session status */
status: SessionStatus;
/** Working directory path */
workingDir: string;
/** ID of currently assigned task, null if none */
currentTaskId: string | null;
/** Timestamp when session was created */
createdAt: number;
/** Timestamp of last activity */
lastActivityAt: number;
/** Session display name */
name?: string;
/** Session mode */
mode?: SessionMode;
/** Auto-clear enabled */
autoClearEnabled?: boolean;
/** Auto-clear token threshold */
autoClearThreshold?: number;
/** Auto-compact enabled */
autoCompactEnabled?: boolean;
/** Auto-compact token threshold */
autoCompactThreshold?: number;
/** Auto-compact prompt */
autoCompactPrompt?: string;
/** Image watcher enabled for this session */
imageWatcherEnabled?: boolean;
/** Total cost in USD */
totalCost?: number;
/** Input tokens used */
inputTokens?: number;
/** Output tokens used */
outputTokens?: number;
/** Whether respawn controller is currently enabled/running */
respawnEnabled?: boolean;
/** Respawn controller config (if enabled) */
respawnConfig?: RespawnConfig & { durationMinutes?: number };
/** Ralph / Todo tracker enabled */
ralphEnabled?: boolean;
/** Ralph auto-enable disabled (user explicitly turned off Ralph) */
ralphAutoEnableDisabled?: boolean;
/** Ralph completion phrase (if set) */
ralphCompletionPhrase?: string;
/** Parent agent ID if this session is a spawned agent */
parentAgentId?: string;
/** Child agent IDs spawned by this session */
childAgentIds?: string[];
/** Nice priority enabled */
niceEnabled?: boolean;
/** Nice value (-20 to 19) */
niceValue?: number;
/** User-assigned color for visual differentiation */
color?: SessionColor;
/** Flicker filter enabled (buffers output after screen clears) */
flickerFilterEnabled?: boolean;
/** Claude Code CLI version (parsed from terminal, e.g., "2.1.27") */
cliVersion?: string;
/** Claude model in use (parsed from terminal, e.g., "Opus 4.5") */
cliModel?: string;
/** Account type (parsed from terminal, e.g., "Claude Max", "API") */
cliAccountType?: string;
/** Latest CLI version available (parsed from version check) */
cliLatestVersion?: string;
/** OpenCode-specific configuration (only for mode === 'opencode') */
openCodeConfig?: OpenCodeConfig;
/** Claude conversation session ID to resume after reboot (set by restore script) */
resumeSessionId?: string;
}
/**
* Output captured from a session
*/
export interface SessionOutput {
/** Standard output content */
stdout: string;
/** Standard error content */
stderr: string;
/** Exit code of the process, null if still running */
exitCode: number | null;
}
+74
View File
@@ -0,0 +1,74 @@
/**
* @fileoverview Task queue type definitions.
*
* Types for the prompt execution queue: TaskDefinition (input) and
* TaskState (persisted state with execution details).
*
* Key exports:
* - TaskDefinition — input type for creating a task (prompt, workingDir, priority, dependencies)
* - TaskState — persisted state with execution details (status, assignedSessionId, output, error)
* - TaskStatus — 'pending' | 'running' | 'completed' | 'failed'
*
* Cross-domain relationships:
* - TaskState.assignedSessionId links to SessionState.id (session domain)
* - TaskState is stored in AppState.tasks (app-state domain)
*
* Persisted to `~/.codeman/state.json`. No dependencies on other domain modules.
*/
/** Status of a task in the queue */
export type TaskStatus = 'pending' | 'running' | 'completed' | 'failed';
/**
* Definition of a task to be executed
*/
export interface TaskDefinition {
/** Unique task identifier */
id: string;
/** Prompt to send to Claude */
prompt: string;
/** Working directory for task execution */
workingDir: string;
/** Priority level (higher = processed first) */
priority: number;
/** IDs of tasks that must complete first */
dependencies: string[];
/** Custom phrase to detect task completion */
completionPhrase?: string;
/** Timeout in milliseconds */
timeoutMs?: number;
}
/**
* Full state of a task including execution details
*/
export interface TaskState {
/** Unique task identifier */
id: string;
/** Prompt sent to Claude */
prompt: string;
/** Working directory for task execution */
workingDir: string;
/** Priority level (higher = processed first) */
priority: number;
/** IDs of tasks that must complete first */
dependencies: string[];
/** Custom phrase to detect task completion */
completionPhrase?: string;
/** Timeout in milliseconds */
timeoutMs?: number;
/** Current task status */
status: TaskStatus;
/** ID of session running this task, null if not assigned */
assignedSessionId: string | null;
/** Timestamp when task was created */
createdAt: number;
/** Timestamp when task started executing */
startedAt: number | null;
/** Timestamp when task completed */
completedAt: number | null;
/** Captured output from Claude */
output: string;
/** Error message if task failed */
error: string | null;
}
+77
View File
@@ -0,0 +1,77 @@
/**
* @fileoverview Agent Teams type definitions (experimental).
*
* Types for Claude Code's Agent Teams feature: team configuration,
* member metadata, task tracking, and inbox messaging.
*
* Key exports:
* - TeamConfig — team from `~/.claude/teams/{name}/config.json` (name, leadSessionId, members[])
* - TeamMember — member entry (agentId, name, agentType, color)
* - TeamTask — task from `~/.claude/tasks/{name}/{N}.json` (subject, status, blocks/blockedBy, owner)
* - InboxMessage — message from `~/.claude/teams/{name}/inboxes/{member}.json`
* - PaneInfo — tmux pane metadata for teammate pane management
*
* Cross-domain relationships:
* - TeamConfig.leadSessionId links to SessionState.id (session domain)
* - Teammates appear as standard subagents (detected by SubagentWatcher)
*
* Served at `GET /api/teams` (list) and `GET /api/teams/:name/tasks`.
* Requires `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1` env var.
* No dependencies on other domain modules.
*/
/** Team configuration from ~/.claude/teams/{name}/config.json */
export interface TeamConfig {
name: string;
leadSessionId: string;
members: TeamMember[];
}
/** A single team member (lead or teammate) */
export interface TeamMember {
agentId: string;
name: string;
agentType: 'team-lead' | 'general-purpose' | string;
color?: string;
backendType?: string;
prompt?: string;
tmuxPaneId?: string;
}
/** A task from ~/.claude/tasks/{team-name}/{N}.json */
export interface TeamTask {
id: string;
subject: string;
description?: string;
activeForm?: string;
status: 'pending' | 'in_progress' | 'completed' | string;
blocks: string[];
blockedBy: string[];
owner?: string;
metadata?: Record<string, unknown>;
}
/** An inbox message from ~/.claude/teams/{name}/inboxes/{member}.json */
export interface InboxMessage {
from: string;
text: string;
timestamp: string;
read?: boolean;
}
/**
* Information about a tmux pane within a session.
* Used for agent team teammate pane management.
*/
export interface PaneInfo {
/** Pane ID (e.g., "%0", "%1") — immutable within a tmux session */
paneId: string;
/** Pane index within the window (0, 1, 2...) */
paneIndex: number;
/** PID of the process running in the pane */
panePid: number;
/** Pane width in columns */
width: number;
/** Pane height in rows */
height: number;
}
+63
View File
@@ -0,0 +1,63 @@
/**
* @fileoverview Tool-related type definitions.
*
* Types for tracking Claude's tool invocations in real-time.
*
* Key exports:
* - ActiveBashTool — a live bash command with extracted file paths and status
* - ActiveBashToolStatus — 'running' | 'completed'
* - ImageDetectedEvent — screenshot/image file detection trigger for UI popup
*
* Cross-domain relationships:
* - ActiveBashTool.sessionId links to SessionState.id (session domain)
* - ImageDetectedEvent.sessionId links to SessionState.id (session domain)
*
* Both types are in-memory only (not persisted). Broadcast via SSE events
* `subagent:tool_call` and `image:detected`. Parsed by BashToolParser
* (`src/bash-tool-parser.ts`).
*/
/**
* Status of an active Bash tool command.
*/
export type ActiveBashToolStatus = 'running' | 'completed';
/**
* Represents an active Bash tool command detected in Claude's output.
* Used to display clickable file paths for file-viewing commands.
*/
export interface ActiveBashTool {
/** Unique identifier for this tool invocation */
id: string;
/** The full command being executed */
command: string;
/** Extracted file paths from the command (clickable) */
filePaths: string[];
/** Timeout string if specified (e.g., "16m 0s") */
timeout?: string;
/** Timestamp when the tool started */
startedAt: number;
/** Current status */
status: ActiveBashToolStatus;
/** Session ID this tool belongs to */
sessionId: string;
}
/**
* Event emitted when a new image file is detected in a session's working directory.
* Used to trigger automatic image popup display in the web UI.
*/
export interface ImageDetectedEvent {
/** Codeman session ID where the image was detected */
sessionId: string;
/** Full path to the detected image file */
filePath: string;
/** Path relative to the session's working directory (for file-raw endpoint) */
relativePath: string;
/** Image file name (basename) */
fileName: string;
/** Timestamp when the image was detected */
timestamp: number;
/** File size in bytes */
size: number;
}
+1 -3
View File
@@ -12,9 +12,7 @@ import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { delimiter, dirname, join } from 'node:path';
import { homedir } from 'node:os';
/** Timeout for exec commands (5 seconds) */
const EXEC_TIMEOUT_MS = 5000;
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
/** Common directories where the Claude CLI binary may be installed */
const CLAUDE_SEARCH_DIRS = [

Some files were not shown because too many files have changed in this diff Show More