mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-05 15:09:42 +02:00
docs: tighten CLAUDE.md and archive 22 completed plan docs
CLAUDE.md: fix stale counts (types 14 to 15, SSE events ~118 to ~120), remove redundant footer sections (References list duplicated inline citations; Common Workflows bullets were self-evident or already stated; Tunnel/Memory Leak Prevention folded into neighboring sections). 251 to 234 lines. Move 22 completed implementation/phase/audit plans to docs/archive/ via git mv so history is preserved. Living reference docs remain in docs/. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,331 @@
|
||||
# Plan: Background Keystroke Forwarding (Local Echo Mode)
|
||||
|
||||
> **Supersedes**: This document merges two previous plan drafts into a single authoritative reference:
|
||||
> - `docs/background-keystroke-forwarding-plan.md` (detailed design doc)
|
||||
> - `.claude/plans/jazzy-bubbling-salamander.md` (Claude-generated implementation plan)
|
||||
>
|
||||
> The docs plan was used as the base. The Claude plan was a correct but simplified subset; its Context paragraph is incorporated below as a lead-in.
|
||||
|
||||
## Context
|
||||
|
||||
When local echo is enabled, keystrokes accumulate in the `LocalEchoOverlay.pendingText` and are only sent to the server when Enter is pressed. This means switching tabs loses the input from the actual Claude Code PTY (the overlay caches text client-side, but the PTY has nothing). If the session respawns or resets, accumulated input is lost entirely.
|
||||
|
||||
## Problem
|
||||
|
||||
When local echo is enabled, keystrokes accumulate **only** in `LocalEchoOverlay.pendingText` (a client-side string). Nothing reaches the server PTY until Enter is pressed. This creates three failure modes:
|
||||
|
||||
1. **Tab switch loses PTY state** — switching sessions saves overlay text to `localEchoTextCache` (a Map), but the actual Claude Code Ink process has no knowledge of what was typed. If respawn or `/clear` fires on that session, the cached text is meaningless.
|
||||
2. **Session death loses input** — if the session crashes or respawns while text is pending in the overlay, that input is gone (localStorage backup `codeman_local_echo_pending` only survives page reloads, not session resets).
|
||||
3. **Tab completion impossible** — pressing Tab with pending overlay text sends the raw Tab character to a PTY that has no knowledge of the typed text, so completion fails.
|
||||
|
||||
## Goal
|
||||
|
||||
Send every keystroke to the server in the background (debounced), while the overlay continues providing instant visual feedback. The overlay sits at z-index 7 with an opaque background over `.xterm-screen`, masking Ink's echo of the background-sent characters. Input persists in the actual Claude Code readline buffer across tab switches and respawns.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
User keystroke
|
||||
|
|
||||
v
|
||||
xterm.js onData(data)
|
||||
|
|
||||
+---> LocalEchoOverlay.addChar(data) [instant visual feedback]
|
||||
|
|
||||
+---> _localEchoBgBuffer += data [queue for background send]
|
||||
| clearTimeout + setTimeout(50ms)
|
||||
| |
|
||||
| v (50ms debounce fires)
|
||||
| _flushBgInput()
|
||||
| |
|
||||
| v
|
||||
| _sendInputAsync(sessionId, buffer) [promise chain preserves order]
|
||||
| |
|
||||
| v
|
||||
| POST /api/sessions/:id/input [{ input: "hel" }]
|
||||
| |
|
||||
| v
|
||||
| session.write(inputStr) [direct PTY write, synchronous]
|
||||
| |
|
||||
| v
|
||||
| Ink readline echoes "hel" [hidden behind overlay's opaque bg]
|
||||
|
|
||||
+--- On Enter:
|
||||
1. clearTimeout(_localEchoBgTimer)
|
||||
2. flush _localEchoBgBuffer via _sendInputAsync (remaining chars)
|
||||
3. clear overlay
|
||||
4. 120ms later: send \r via _sendInputAsync (Ink text/Enter split)
|
||||
5. Ink processes "hello\r" → overlay gone, terminal visible with output
|
||||
```
|
||||
|
||||
### Two Input Paths (important context)
|
||||
|
||||
The codebase has **two separate input paths** to the server:
|
||||
|
||||
| Path | Used by | Promise chain? | `useMux`? |
|
||||
|------|---------|---------------|-----------|
|
||||
| `_sendInputAsync()` (line 3626) | `onData` handler, `flushInput()` | Yes (`_inputSendChain`) | No (direct PTY write) |
|
||||
| `sendInput()` (line 8755) | Mobile accessory bar, programmatic commands | **No** (raw `fetch`) | Yes (tmux `send-keys`) |
|
||||
|
||||
Background keystroke forwarding uses **only** the `_sendInputAsync` path, which guarantees ordering via the promise chain. The `sendInput()` path is unaffected and unmodified.
|
||||
|
||||
### Server-Side Input Flow
|
||||
|
||||
```
|
||||
POST /api/sessions/:id/input { input: "hel" }
|
||||
|
|
||||
+-- useMux? No (default)
|
||||
| session.write("hel") → ptyProcess.write("hel") [sync]
|
||||
|
|
||||
+-- useMux? Yes
|
||||
session.writeViaMux("hel") → tmux send-keys -l "hel" [async]
|
||||
```
|
||||
|
||||
Background sends use the default path (no `useMux`), which is a synchronous direct PTY write — faster than spawning a tmux subprocess for each character batch.
|
||||
|
||||
## Implementation
|
||||
|
||||
All changes in **one file**: `src/web/public/app.js`
|
||||
|
||||
### Step 1: Add background send state (in terminal setup, after line ~1999)
|
||||
|
||||
```js
|
||||
this._localEchoBgBuffer = ''; // Characters queued for background send
|
||||
this._localEchoBgTimer = null; // 50ms debounce timer ID
|
||||
```
|
||||
|
||||
Add an atomic drain helper alongside existing `flushInput` (after line ~2008):
|
||||
|
||||
```js
|
||||
// Atomically drain background buffer — returns contents and cancels pending timer.
|
||||
// Single point of extraction prevents double-flush race conditions.
|
||||
const drainBgBuffer = () => {
|
||||
if (this._localEchoBgTimer) {
|
||||
clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = null;
|
||||
}
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
return buf;
|
||||
};
|
||||
|
||||
const scheduleBgFlush = () => {
|
||||
if (this._localEchoBgTimer) clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = setTimeout(() => {
|
||||
this._localEchoBgTimer = null;
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
if (buf && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, buf);
|
||||
}
|
||||
}, 50);
|
||||
};
|
||||
```
|
||||
|
||||
**Why `drainBgBuffer` exists**: Every exit path (Enter, Ctrl+C, tab switch, echo disable) needs to flush the buffer AND cancel the timer atomically. Without a single extraction point, it's easy to forget one of the two operations, leading to double-sends when the timer fires after a manual flush.
|
||||
|
||||
### Step 2: Modify `onData` handler — local echo path (lines 2023–2067)
|
||||
|
||||
**Printable characters** (lines 2063–2067 → replace):
|
||||
```js
|
||||
if (data.length === 1 && data.charCodeAt(0) >= 32) {
|
||||
this._localEchoOverlay?.addChar(data);
|
||||
// Background: queue char for server send (50ms debounce batches rapid typing)
|
||||
this._localEchoBgBuffer += data;
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Backspace** (lines 2024–2028 → replace):
|
||||
```js
|
||||
if (data === '\x7f') {
|
||||
this._localEchoOverlay?.removeChar();
|
||||
// Background: queue DEL for server (Ink's readline handles backspace via \x7f)
|
||||
this._localEchoBgBuffer += '\x7f';
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Enter** (lines 2029–2050 → replace):
|
||||
```js
|
||||
if (/^[\r\n]+$/.test(data)) {
|
||||
this._localEchoOverlay?.clear();
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
// Flush any remaining background chars (e.g., last 50ms batch not yet sent)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send \r after 120ms — Ink needs text and Enter as separate events.
|
||||
// The promise chain in _sendInputAsync guarantees the remaining chars
|
||||
// are dispatched before \r, regardless of timing.
|
||||
setTimeout(() => {
|
||||
this._pendingInput += '\r';
|
||||
flushInput();
|
||||
}, 120);
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Key change from original plan**: The Enter handler no longer checks `if (text)` and branches on whether the overlay had content. With background sends, the PTY already has most/all of the text. We just flush any remainder and unconditionally send `\r` after 120ms. This simplifies the flow and handles edge cases like "user typed nothing but pressed Enter" (remainder is empty, just `\r` is sent).
|
||||
|
||||
**Control characters and paste** (lines 2052–2061 → replace):
|
||||
```js
|
||||
if (data.charCodeAt(0) < 32 || data.length > 1) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush background buffer so PTY has full text state before control char
|
||||
// (critical for Tab completion — PTY needs typed text to complete against)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send control char / paste text via normal path
|
||||
this._pendingInput += data;
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
flushInput();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Note on paste**: Desktop paste arrives via `onData` as a single multi-character string (`data.length > 1`). This falls into the control char path above, which:
|
||||
1. Clears the overlay (existing behavior)
|
||||
2. Flushes background buffer (new — ensures PTY has prefix text)
|
||||
3. Sends paste text immediately (existing behavior)
|
||||
|
||||
Mobile paste via `KeyboardAccessoryBar.pasteFromClipboard()` uses `app.sendInput()` which bypasses `onData` entirely — no change needed.
|
||||
|
||||
### Step 3: Flush on tab switch (`selectSession()`, line ~4533)
|
||||
|
||||
Insert before the existing overlay save/clear block (before line 4534):
|
||||
|
||||
```js
|
||||
// Flush background send buffer for outgoing session
|
||||
if (this.activeSessionId) {
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This ensures the PTY receives all typed characters before the tab switch. When the user switches back, the terminal buffer will show the text (echoed by Ink) and the overlay will restore its cached copy on top.
|
||||
|
||||
### Step 4: Cleanup on local echo disable (`_updateLocalEchoState()`, lines 2362–2371)
|
||||
|
||||
Expand the disable transition (line 2367–2368):
|
||||
```js
|
||||
if (this._localEchoEnabled && !shouldEnable) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush any pending background chars before disabling
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Step 5: Cleanup on session delete (`deleteSession()`)
|
||||
|
||||
When a session is deleted, cancel any pending background timer for that session:
|
||||
```js
|
||||
// In deleteSession(), after removing the session from this.sessions:
|
||||
drainBgBuffer(); // Discard — session is gone, nowhere to send
|
||||
this.localEchoTextCache.delete(sessionId);
|
||||
```
|
||||
|
||||
## Visual Timeline
|
||||
|
||||
```
|
||||
t=0ms User types "h" → overlay: "h" bgBuffer: "h" timer: 50ms
|
||||
t=30ms User types "e" → overlay: "he" bgBuffer: "he" timer: reset 50ms
|
||||
t=60ms User types "l" → overlay: "hel" bgBuffer: "hel" timer: reset 50ms
|
||||
t=110ms Debounce fires → overlay: "hel" bgBuffer: "" POST "hel" → PTY
|
||||
t=115ms Ink echoes "hel" → terminal: "❯ hel" (hidden behind overlay)
|
||||
t=140ms User types "l" → overlay: "hell" bgBuffer: "l" timer: 50ms
|
||||
t=170ms User types "o" → overlay: "hello" bgBuffer: "lo" timer: reset 50ms
|
||||
t=220ms Debounce fires → overlay: "hello" bgBuffer: "" POST "lo" → PTY
|
||||
t=250ms User hits Enter → drainBgBuffer()="" overlay: cleared
|
||||
t=370ms \r sent via chain → Ink processes "hello\r" → output appears
|
||||
```
|
||||
|
||||
**Tab switch scenario:**
|
||||
```
|
||||
t=0ms User types "wor" → overlay: "wor" bgBuffer: "wor" timer: 50ms
|
||||
t=25ms User switches tab → drainBgBuffer() sends "wor" to old session PTY
|
||||
overlay text "wor" saved to localEchoTextCache
|
||||
overlay cleared, new session loaded
|
||||
...later...
|
||||
t=5000ms User switches back → terminal shows "❯ wor" (Ink echo from background send)
|
||||
overlay restores "wor" from cache, masks terminal
|
||||
user continues typing seamlessly
|
||||
```
|
||||
|
||||
## Edge Cases & Mitigations
|
||||
|
||||
### Confirmed Safe (JS single-threaded guarantee)
|
||||
|
||||
| Scenario | Why it's safe |
|
||||
|----------|--------------|
|
||||
| **Debounce fires during Enter handler** | Impossible. JS event loop is single-threaded — the Enter handler runs atomically. `drainBgBuffer()` cancels the timer before it can fire. |
|
||||
| **Debounce fires during tab switch** | Same reason. `selectSession()` calls `drainBgBuffer()` synchronously, canceling the timer. |
|
||||
| **Double-send of background buffer** | `drainBgBuffer()` atomically clears both buffer and timer. Once drained, subsequent drain returns empty string. |
|
||||
| **`_pendingInput` conflict** | In local echo mode, `_pendingInput` is only used for Enter (`\r`) and control chars. Background chars use a separate `_localEchoBgBuffer`. No overlap. |
|
||||
|
||||
### Handled by Design
|
||||
|
||||
| Scenario | Handling |
|
||||
|----------|---------|
|
||||
| **Rapid typing / paste** | 50ms debounce batches rapid chars. At 100 WPM (~50ms/char), sends ~1 char per batch. For paste (multi-char string, `data.length > 1`), the control char path bypasses the buffer entirely and sends immediately. |
|
||||
| **Network failure** | `_sendInputAsync` catches fetch failures and calls `_enqueueInput()` for retry. `_drainInputQueues()` replays on reconnect. Background chars use the same retry path. |
|
||||
| **Offline mode** | `_sendInputAsync` checks `this.isOnline` and immediately enqueues if offline. Same behavior for background sends. 64KB queue cap prevents memory growth. |
|
||||
| **Tab completion** | Ctrl+Tab path flushes background buffer BEFORE sending Tab char. PTY has full text for readline completion. |
|
||||
| **Session respawn** | PTY already has typed text (sent in background). On respawn, Claude exits and restarts — Ink's readline buffer is lost, but the text was already processed or is no longer relevant. The overlay clears on session status change via `_updateLocalEchoState()`. |
|
||||
| **SSE reconnect** | `handleInit()` saves overlay text before `selectSession()` clears it, then restores after reload (line ~3945–3966). Background buffer is cleared on reconnect since state is reset. |
|
||||
|
||||
### Network Ordering
|
||||
|
||||
**Question**: Can background sends arrive at the server out of order?
|
||||
|
||||
**Answer**: No, for practical purposes.
|
||||
|
||||
1. `_sendInputAsync` uses a **promise chain** (`_inputSendChain`) — each fetch is dispatched only after the previous one has been dispatched. This means requests are sent in order.
|
||||
2. Localhost connections (HTTP/1.1) are inherently sequential on a single TCP connection.
|
||||
3. Even with HTTP/2 multiplexing, Fastify (Node.js) is single-threaded — request handlers execute via the event loop in arrival order.
|
||||
4. The server's `session.write()` is synchronous — it writes to the PTY immediately within the request handler.
|
||||
|
||||
### Known Limitations (Not Addressed)
|
||||
|
||||
| Limitation | Impact | Notes |
|
||||
|-----------|--------|-------|
|
||||
| **IME composition** | CJK input via IME would send partial composition sequences to PTY | No IME handling exists in the codebase today (line count: 0 references to `compositionstart/end/update`). Fixing this is a separate feature. |
|
||||
| **`sendInput()` ordering** | Mobile accessory bar commands (`/init`, `/clear`, paste) use `sendInput()` which bypasses `_inputSendChain` — no ordering guarantee relative to background sends | Unlikely to conflict in practice: accessory bar clears the overlay first, and the commands are typically sent when no typing is in progress. |
|
||||
| **localStorage stale text** | After background sends, localStorage still has overlay text. On hard reload, overlay restores text that the PTY already has → visual duplicate behind overlay | Harmless — overlay masks the terminal. On Enter, overlay clears and terminal shows correct state. Could be fixed by clearing localStorage after successful background flush, but adds complexity for minimal benefit. |
|
||||
|
||||
## Verification Checklist
|
||||
|
||||
1. **Basic typing**: Enable local echo → type "hello" → overlay shows instantly → check Network tab for batched POST requests (~50ms intervals) → press Enter → command executes
|
||||
2. **Tab switch persistence**: Type "test" → switch to another tab → switch back → text visible in both overlay AND terminal prompt
|
||||
3. **Backspace**: Type "helloo" → press backspace → overlay shows "hello" → check PTY received \x7f
|
||||
4. **Paste**: Type "hel" → paste "lo world" → overlay clears → "lo world" sent immediately → PTY has "hello world"
|
||||
5. **Tab completion**: Type "src/w" → press Tab → PTY completes to "src/web/" (background send gave PTY the prefix)
|
||||
6. **Ctrl+C**: Type "hello" → press Ctrl+C → overlay clears → PTY receives pending chars + \x03
|
||||
7. **Network tab**: Verify POST /api/sessions/:id/input requests appear as you type (batched ~50ms)
|
||||
8. **Offline resilience**: Disconnect network → type "hello" → reconnect → verify chars are replayed via drain queue
|
||||
9. **Session delete**: Type text → delete session → no console errors from orphaned timer
|
||||
10. **Mobile keyboard**: Test on mobile device — typing goes through same onData path, same behavior expected
|
||||
|
||||
## Files Modified
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/web/public/app.js` | ~40 lines changed across 5 locations (Steps 1–5) |
|
||||
|
||||
No server-side changes. No new files. No new dependencies.
|
||||
@@ -0,0 +1,321 @@
|
||||
# Plan: Background Keystroke Forwarding (Local Echo Mode)
|
||||
|
||||
## Problem
|
||||
|
||||
When local echo is enabled, keystrokes accumulate **only** in `LocalEchoOverlay.pendingText` (a client-side string). Nothing reaches the server PTY until Enter is pressed. This creates three failure modes:
|
||||
|
||||
1. **Tab switch loses PTY state** — switching sessions saves overlay text to `localEchoTextCache` (a Map), but the actual Claude Code Ink process has no knowledge of what was typed. If respawn or `/clear` fires on that session, the cached text is meaningless.
|
||||
2. **Session death loses input** — if the session crashes or respawns while text is pending in the overlay, that input is gone (localStorage backup `codeman_local_echo_pending` only survives page reloads, not session resets).
|
||||
3. **Tab completion impossible** — pressing Tab with pending overlay text sends the raw Tab character to a PTY that has no knowledge of the typed text, so completion fails.
|
||||
|
||||
## Goal
|
||||
|
||||
Send every keystroke to the server in the background (debounced), while the overlay continues providing instant visual feedback. The overlay sits at z-index 7 with an opaque background over `.xterm-screen`, masking Ink's echo of the background-sent characters. Input persists in the actual Claude Code readline buffer across tab switches and respawns.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
User keystroke
|
||||
|
|
||||
v
|
||||
xterm.js onData(data)
|
||||
|
|
||||
+---> LocalEchoOverlay.addChar(data) [instant visual feedback]
|
||||
|
|
||||
+---> _localEchoBgBuffer += data [queue for background send]
|
||||
| clearTimeout + setTimeout(50ms)
|
||||
| |
|
||||
| v (50ms debounce fires)
|
||||
| _flushBgInput()
|
||||
| |
|
||||
| v
|
||||
| _sendInputAsync(sessionId, buffer) [promise chain preserves order]
|
||||
| |
|
||||
| v
|
||||
| POST /api/sessions/:id/input [{ input: "hel" }]
|
||||
| |
|
||||
| v
|
||||
| session.write(inputStr) [direct PTY write, synchronous]
|
||||
| |
|
||||
| v
|
||||
| Ink readline echoes "hel" [hidden behind overlay's opaque bg]
|
||||
|
|
||||
+--- On Enter:
|
||||
1. clearTimeout(_localEchoBgTimer)
|
||||
2. flush _localEchoBgBuffer via _sendInputAsync (remaining chars)
|
||||
3. clear overlay
|
||||
4. 120ms later: send \r via _sendInputAsync (Ink text/Enter split)
|
||||
5. Ink processes "hello\r" → overlay gone, terminal visible with output
|
||||
```
|
||||
|
||||
### Two Input Paths (important context)
|
||||
|
||||
The codebase has **two separate input paths** to the server:
|
||||
|
||||
| Path | Used by | Promise chain? | `useMux`? |
|
||||
|------|---------|---------------|-----------|
|
||||
| `_sendInputAsync()` (line 3626) | `onData` handler, `flushInput()` | Yes (`_inputSendChain`) | No (direct PTY write) |
|
||||
| `sendInput()` (line 8755) | Mobile accessory bar, programmatic commands | **No** (raw `fetch`) | Yes (tmux `send-keys`) |
|
||||
|
||||
Background keystroke forwarding uses **only** the `_sendInputAsync` path, which guarantees ordering via the promise chain. The `sendInput()` path is unaffected and unmodified.
|
||||
|
||||
### Server-Side Input Flow
|
||||
|
||||
```
|
||||
POST /api/sessions/:id/input { input: "hel" }
|
||||
|
|
||||
+-- useMux? No (default)
|
||||
| session.write("hel") → ptyProcess.write("hel") [sync]
|
||||
|
|
||||
+-- useMux? Yes
|
||||
session.writeViaMux("hel") → tmux send-keys -l "hel" [async]
|
||||
```
|
||||
|
||||
Background sends use the default path (no `useMux`), which is a synchronous direct PTY write — faster than spawning a tmux subprocess for each character batch.
|
||||
|
||||
## Implementation
|
||||
|
||||
All changes in **one file**: `src/web/public/app.js`
|
||||
|
||||
### Step 1: Add background send state (in terminal setup, after line ~1999)
|
||||
|
||||
```js
|
||||
this._localEchoBgBuffer = ''; // Characters queued for background send
|
||||
this._localEchoBgTimer = null; // 50ms debounce timer ID
|
||||
```
|
||||
|
||||
Add an atomic drain helper alongside existing `flushInput` (after line ~2008):
|
||||
|
||||
```js
|
||||
// Atomically drain background buffer — returns contents and cancels pending timer.
|
||||
// Single point of extraction prevents double-flush race conditions.
|
||||
const drainBgBuffer = () => {
|
||||
if (this._localEchoBgTimer) {
|
||||
clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = null;
|
||||
}
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
return buf;
|
||||
};
|
||||
|
||||
const scheduleBgFlush = () => {
|
||||
if (this._localEchoBgTimer) clearTimeout(this._localEchoBgTimer);
|
||||
this._localEchoBgTimer = setTimeout(() => {
|
||||
this._localEchoBgTimer = null;
|
||||
const buf = this._localEchoBgBuffer;
|
||||
this._localEchoBgBuffer = '';
|
||||
if (buf && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, buf);
|
||||
}
|
||||
}, 50);
|
||||
};
|
||||
```
|
||||
|
||||
**Why `drainBgBuffer` exists**: Every exit path (Enter, Ctrl+C, tab switch, echo disable) needs to flush the buffer AND cancel the timer atomically. Without a single extraction point, it's easy to forget one of the two operations, leading to double-sends when the timer fires after a manual flush.
|
||||
|
||||
### Step 2: Modify `onData` handler — local echo path (lines 2023–2067)
|
||||
|
||||
**Printable characters** (lines 2063–2067 → replace):
|
||||
```js
|
||||
if (data.length === 1 && data.charCodeAt(0) >= 32) {
|
||||
this._localEchoOverlay?.addChar(data);
|
||||
// Background: queue char for server send (50ms debounce batches rapid typing)
|
||||
this._localEchoBgBuffer += data;
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Backspace** (lines 2024–2028 → replace):
|
||||
```js
|
||||
if (data === '\x7f') {
|
||||
this._localEchoOverlay?.removeChar();
|
||||
// Background: queue DEL for server (Ink's readline handles backspace via \x7f)
|
||||
this._localEchoBgBuffer += '\x7f';
|
||||
scheduleBgFlush();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Enter** (lines 2029–2050 → replace):
|
||||
```js
|
||||
if (/^[\r\n]+$/.test(data)) {
|
||||
this._localEchoOverlay?.clear();
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
// Flush any remaining background chars (e.g., last 50ms batch not yet sent)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send \r after 120ms — Ink needs text and Enter as separate events.
|
||||
// The promise chain in _sendInputAsync guarantees the remaining chars
|
||||
// are dispatched before \r, regardless of timing.
|
||||
setTimeout(() => {
|
||||
this._pendingInput += '\r';
|
||||
flushInput();
|
||||
}, 120);
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Key change from original plan**: The Enter handler no longer checks `if (text)` and branches on whether the overlay had content. With background sends, the PTY already has most/all of the text. We just flush any remainder and unconditionally send `\r` after 120ms. This simplifies the flow and handles edge cases like "user typed nothing but pressed Enter" (remainder is empty, just `\r` is sent).
|
||||
|
||||
**Control characters and paste** (lines 2052–2061 → replace):
|
||||
```js
|
||||
if (data.charCodeAt(0) < 32 || data.length > 1) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush background buffer so PTY has full text state before control char
|
||||
// (critical for Tab completion — PTY needs typed text to complete against)
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
// Send control char / paste text via normal path
|
||||
this._pendingInput += data;
|
||||
if (this._inputFlushTimeout) {
|
||||
clearTimeout(this._inputFlushTimeout);
|
||||
this._inputFlushTimeout = null;
|
||||
}
|
||||
flushInput();
|
||||
return;
|
||||
}
|
||||
```
|
||||
|
||||
**Note on paste**: Desktop paste arrives via `onData` as a single multi-character string (`data.length > 1`). This falls into the control char path above, which:
|
||||
1. Clears the overlay (existing behavior)
|
||||
2. Flushes background buffer (new — ensures PTY has prefix text)
|
||||
3. Sends paste text immediately (existing behavior)
|
||||
|
||||
Mobile paste via `KeyboardAccessoryBar.pasteFromClipboard()` uses `app.sendInput()` which bypasses `onData` entirely — no change needed.
|
||||
|
||||
### Step 3: Flush on tab switch (`selectSession()`, line ~4533)
|
||||
|
||||
Insert before the existing overlay save/clear block (before line 4534):
|
||||
|
||||
```js
|
||||
// Flush background send buffer for outgoing session
|
||||
if (this.activeSessionId) {
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This ensures the PTY receives all typed characters before the tab switch. When the user switches back, the terminal buffer will show the text (echoed by Ink) and the overlay will restore its cached copy on top.
|
||||
|
||||
### Step 4: Cleanup on local echo disable (`_updateLocalEchoState()`, lines 2362–2371)
|
||||
|
||||
Expand the disable transition (line 2367–2368):
|
||||
```js
|
||||
if (this._localEchoEnabled && !shouldEnable) {
|
||||
this._localEchoOverlay?.clear();
|
||||
// Flush any pending background chars before disabling
|
||||
const remaining = drainBgBuffer();
|
||||
if (remaining && this.activeSessionId) {
|
||||
this._sendInputAsync(this.activeSessionId, remaining);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Step 5: Cleanup on session delete (`deleteSession()`)
|
||||
|
||||
When a session is deleted, cancel any pending background timer for that session:
|
||||
```js
|
||||
// In deleteSession(), after removing the session from this.sessions:
|
||||
drainBgBuffer(); // Discard — session is gone, nowhere to send
|
||||
this.localEchoTextCache.delete(sessionId);
|
||||
```
|
||||
|
||||
## Visual Timeline
|
||||
|
||||
```
|
||||
t=0ms User types "h" → overlay: "h" bgBuffer: "h" timer: 50ms
|
||||
t=30ms User types "e" → overlay: "he" bgBuffer: "he" timer: reset 50ms
|
||||
t=60ms User types "l" → overlay: "hel" bgBuffer: "hel" timer: reset 50ms
|
||||
t=110ms Debounce fires → overlay: "hel" bgBuffer: "" POST "hel" → PTY
|
||||
t=115ms Ink echoes "hel" → terminal: "❯ hel" (hidden behind overlay)
|
||||
t=140ms User types "l" → overlay: "hell" bgBuffer: "l" timer: 50ms
|
||||
t=170ms User types "o" → overlay: "hello" bgBuffer: "lo" timer: reset 50ms
|
||||
t=220ms Debounce fires → overlay: "hello" bgBuffer: "" POST "lo" → PTY
|
||||
t=250ms User hits Enter → drainBgBuffer()="" overlay: cleared
|
||||
t=370ms \r sent via chain → Ink processes "hello\r" → output appears
|
||||
```
|
||||
|
||||
**Tab switch scenario:**
|
||||
```
|
||||
t=0ms User types "wor" → overlay: "wor" bgBuffer: "wor" timer: 50ms
|
||||
t=25ms User switches tab → drainBgBuffer() sends "wor" to old session PTY
|
||||
overlay text "wor" saved to localEchoTextCache
|
||||
overlay cleared, new session loaded
|
||||
...later...
|
||||
t=5000ms User switches back → terminal shows "❯ wor" (Ink echo from background send)
|
||||
overlay restores "wor" from cache, masks terminal
|
||||
user continues typing seamlessly
|
||||
```
|
||||
|
||||
## Edge Cases & Mitigations
|
||||
|
||||
### Confirmed Safe (JS single-threaded guarantee)
|
||||
|
||||
| Scenario | Why it's safe |
|
||||
|----------|--------------|
|
||||
| **Debounce fires during Enter handler** | Impossible. JS event loop is single-threaded — the Enter handler runs atomically. `drainBgBuffer()` cancels the timer before it can fire. |
|
||||
| **Debounce fires during tab switch** | Same reason. `selectSession()` calls `drainBgBuffer()` synchronously, canceling the timer. |
|
||||
| **Double-send of background buffer** | `drainBgBuffer()` atomically clears both buffer and timer. Once drained, subsequent drain returns empty string. |
|
||||
| **`_pendingInput` conflict** | In local echo mode, `_pendingInput` is only used for Enter (`\r`) and control chars. Background chars use a separate `_localEchoBgBuffer`. No overlap. |
|
||||
|
||||
### Handled by Design
|
||||
|
||||
| Scenario | Handling |
|
||||
|----------|---------|
|
||||
| **Rapid typing / paste** | 50ms debounce batches rapid chars. At 100 WPM (~50ms/char), sends ~1 char per batch. For paste (multi-char string, `data.length > 1`), the control char path bypasses the buffer entirely and sends immediately. |
|
||||
| **Network failure** | `_sendInputAsync` catches fetch failures and calls `_enqueueInput()` for retry. `_drainInputQueues()` replays on reconnect. Background chars use the same retry path. |
|
||||
| **Offline mode** | `_sendInputAsync` checks `this.isOnline` and immediately enqueues if offline. Same behavior for background sends. 64KB queue cap prevents memory growth. |
|
||||
| **Tab completion** | Ctrl+Tab path flushes background buffer BEFORE sending Tab char. PTY has full text for readline completion. |
|
||||
| **Session respawn** | PTY already has typed text (sent in background). On respawn, Claude exits and restarts — Ink's readline buffer is lost, but the text was already processed or is no longer relevant. The overlay clears on session status change via `_updateLocalEchoState()`. |
|
||||
| **SSE reconnect** | `handleInit()` saves overlay text before `selectSession()` clears it, then restores after reload (line ~3945–3966). Background buffer is cleared on reconnect since state is reset. |
|
||||
|
||||
### Network Ordering
|
||||
|
||||
**Question**: Can background sends arrive at the server out of order?
|
||||
|
||||
**Answer**: No, for practical purposes.
|
||||
|
||||
1. `_sendInputAsync` uses a **promise chain** (`_inputSendChain`) — each fetch is dispatched only after the previous one has been dispatched. This means requests are sent in order.
|
||||
2. Localhost connections (HTTP/1.1) are inherently sequential on a single TCP connection.
|
||||
3. Even with HTTP/2 multiplexing, Fastify (Node.js) is single-threaded — request handlers execute via the event loop in arrival order.
|
||||
4. The server's `session.write()` is synchronous — it writes to the PTY immediately within the request handler.
|
||||
|
||||
### Known Limitations (Not Addressed)
|
||||
|
||||
| Limitation | Impact | Notes |
|
||||
|-----------|--------|-------|
|
||||
| **IME composition** | CJK input via IME would send partial composition sequences to PTY | No IME handling exists in the codebase today (line count: 0 references to `compositionstart/end/update`). Fixing this is a separate feature. |
|
||||
| **`sendInput()` ordering** | Mobile accessory bar commands (`/init`, `/clear`, paste) use `sendInput()` which bypasses `_inputSendChain` — no ordering guarantee relative to background sends | Unlikely to conflict in practice: accessory bar clears the overlay first, and the commands are typically sent when no typing is in progress. |
|
||||
| **localStorage stale text** | After background sends, localStorage still has overlay text. On hard reload, overlay restores text that the PTY already has → visual duplicate behind overlay | Harmless — overlay masks the terminal. On Enter, overlay clears and terminal shows correct state. Could be fixed by clearing localStorage after successful background flush, but adds complexity for minimal benefit. |
|
||||
|
||||
## Verification Checklist
|
||||
|
||||
1. **Basic typing**: Enable local echo → type "hello" → overlay shows instantly → check Network tab for batched POST requests (~50ms intervals) → press Enter → command executes
|
||||
2. **Tab switch persistence**: Type "test" → switch to another tab → switch back → text visible in both overlay AND terminal prompt
|
||||
3. **Backspace**: Type "helloo" → press backspace → overlay shows "hello" → check PTY received \x7f
|
||||
4. **Paste**: Type "hel" → paste "lo world" → overlay clears → "lo world" sent immediately → PTY has "hello world"
|
||||
5. **Tab completion**: Type "src/w" → press Tab → PTY completes to "src/web/" (background send gave PTY the prefix)
|
||||
6. **Ctrl+C**: Type "hello" → press Ctrl+C → overlay clears → PTY receives pending chars + \x03
|
||||
7. **Network tab**: Verify POST /api/sessions/:id/input requests appear as you type (batched ~50ms)
|
||||
8. **Offline resilience**: Disconnect network → type "hello" → reconnect → verify chars are replayed via drain queue
|
||||
9. **Session delete**: Type text → delete session → no console errors from orphaned timer
|
||||
10. **Mobile keyboard**: Test on mobile device — typing goes through same onData path, same behavior expected
|
||||
|
||||
## Files Modified
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/web/public/app.js` | ~40 lines changed across 5 locations (Steps 1–5) |
|
||||
|
||||
No server-side changes. No new files. No new dependencies.
|
||||
@@ -0,0 +1,409 @@
|
||||
# First-Load Performance Optimization Plan
|
||||
|
||||
**Date**: 2026-02-18
|
||||
**Audit by**: 4-agent team (css-analyst, js-analyst, server-analyst, deps-analyst)
|
||||
**Scope**: First browser load of Codeman web UI at `/`
|
||||
|
||||
---
|
||||
|
||||
## Current State (Baseline)
|
||||
|
||||
### Payload
|
||||
|
||||
| Asset | Raw | Compressed | Render-Blocking? |
|
||||
|-------|-----|-----------|-----------------|
|
||||
| `index.html` | 82 KB | ~15 KB | N/A (document) |
|
||||
| `styles.css` | 154 KB | ~25 KB | **YES** |
|
||||
| `mobile.css` | 34 KB | ~7 KB | **YES** (missing media query) |
|
||||
| `xterm.css` (CDN) | 2 KB | ~2 KB | No (preload pattern) |
|
||||
| `xterm.min.js` (CDN) | 67 KB | ~65 KB | No (defer) |
|
||||
| `xterm-addon-fit` (CDN) | 1 KB | ~1 KB | No (defer) |
|
||||
| `app.js` | 563 KB | ~126 KB | No (defer) |
|
||||
| **Total** | **903 KB** | **~241 KB** | |
|
||||
|
||||
### Request Waterfall (13 requests on first load)
|
||||
|
||||
```
|
||||
T=0 GET / (82KB doc)
|
||||
T+20ms ├── styles.css?v=0.1536 (154KB — BLOCKS RENDER)
|
||||
├── mobile.css?v=0.1536 (34KB — BLOCKS RENDER on all viewports!)
|
||||
├── xterm.css (CDN, preloaded) (2KB — non-blocking, already async)
|
||||
├── xterm.min.js (CDN, defer) (67KB)
|
||||
├── xterm-addon-fit.min.js (CDN) (1KB)
|
||||
└── app.js?v=0.1536 (defer) (563KB)
|
||||
|
||||
[FIRST PAINT blocked by: styles.css + mobile.css]
|
||||
|
||||
T+200ms JS execution starts
|
||||
├── new Terminal() + terminal.open() ← HEAVY sync (canvas creation)
|
||||
├── connectSSE() → /api/events ← SSE stream
|
||||
├── loadState() → /api/status ← DUPLICATE of SSE init!
|
||||
├── loadQuickStartCases()
|
||||
│ ├── /api/settings ← fetched TWICE
|
||||
│ └── /api/cases?_t=<timestamp> ← cache-busted unnecessarily
|
||||
├── startSystemStatsPolling()
|
||||
│ └── /api/system/stats ← starts immediately, every 2s
|
||||
└── loadAppSettingsFromServer()
|
||||
└── /api/settings ← DUPLICATE #2
|
||||
|
||||
T+500ms First Meaningful Paint (terminal + header visible)
|
||||
```
|
||||
|
||||
### Problems
|
||||
|
||||
1. **2 render-blocking CSS files** — mobile.css blocks desktop for no reason
|
||||
2. **Double handleInit()** — SSE init + /api/status both call full state reset
|
||||
3. **Duplicate /api/settings** — fetched twice in init chain
|
||||
4. **563KB unminified JS monolith** — no build minification at all
|
||||
5. **154KB unminified CSS** — 70% is for modals/wizards (below-the-fold)
|
||||
6. **Sync terminal.open()** — heaviest single call, blocks before first paint
|
||||
7. **12 modals pre-rendered** — ~600+ hidden DOM nodes, ~60KB HTML
|
||||
8. **Stats polling starts immediately** — even with 0 sessions
|
||||
9. **No loading skeleton** — blank black screen until all CSS+JS loads
|
||||
10. **CDN dependency** — 3 xterm files from jsdelivr (DNS+TLS latency)
|
||||
11. **1h cache for versioned assets** — could be 1yr+immutable with ?v= busting
|
||||
12. **No HTTP/2** — 6-connection limit queues some requests
|
||||
13. **On-the-fly compression** — no pre-compressed .gz/.br files
|
||||
|
||||
### What's Already Good (don't touch)
|
||||
|
||||
- Single shared Terminal instance (buffer swapping)
|
||||
- Teammate terminals created lazily on window open
|
||||
- Subagent windows use HTML logs, not Terminal instances
|
||||
- `getLightState()` has 1s TTL cache
|
||||
- SSE init sends lightweight state (no terminal buffers)
|
||||
- Buffer hydration uses chunked writes (128KB via rAF)
|
||||
- `selectSession()` defers secondary panels via requestIdleCallback
|
||||
- System fonts only — zero web font loading
|
||||
- xterm.css already uses async preload pattern
|
||||
- Proper SSE reconnection with exponential backoff
|
||||
- CSS `contain` on header/tabs for layout isolation
|
||||
|
||||
---
|
||||
|
||||
## Implementation Plan (15 steps, ordered by impact/effort)
|
||||
|
||||
### Phase 1: Quick Wins (1-line to 15-min changes)
|
||||
|
||||
#### Step 1: Add media attribute to mobile.css
|
||||
**Impact**: HIGH — 34KB stops blocking render on desktop
|
||||
**File**: `src/web/public/index.html:14`
|
||||
|
||||
```html
|
||||
<!-- BEFORE -->
|
||||
<link rel="stylesheet" href="mobile.css?v=0.1536">
|
||||
|
||||
<!-- AFTER -->
|
||||
<link rel="stylesheet" href="mobile.css?v=0.1536" media="(max-width: 1023px)">
|
||||
```
|
||||
|
||||
Browser still downloads it (for potential resize) but won't block rendering on desktop. The mobile.css file header says this was intended but never implemented.
|
||||
|
||||
---
|
||||
|
||||
#### Step 2: Remove duplicate /api/status + double handleInit()
|
||||
**Impact**: HIGH — eliminates redundant API call + double state reset (clears 15+ Maps, 7+ timers, runs cleanupAllFloatingWindows(), double renderSessionTabs())
|
||||
**Files**: `src/web/public/app.js`
|
||||
|
||||
The SSE `init` event (server.ts:618) sends `getLightState()`. The `loadState()` in `init()` at `app.js:1554` fetches identical data from `/api/status`. Both call `handleInit()` which wipes state. The `_initGeneration` guard only protects session-restore, NOT the expensive cleanup (lines 3389-3503).
|
||||
|
||||
```js
|
||||
// In init() — REMOVE this.loadState(), add SSE fallback:
|
||||
this.connectSSE();
|
||||
// Remove: this.loadState();
|
||||
this._initFallbackTimer = setTimeout(() => {
|
||||
if (this._initGeneration === 0) this.loadState();
|
||||
}, 3000);
|
||||
|
||||
// In handleInit() — clear fallback timer:
|
||||
handleInit(data) {
|
||||
if (this._initFallbackTimer) {
|
||||
clearTimeout(this._initFallbackTimer);
|
||||
this._initFallbackTimer = null;
|
||||
}
|
||||
// ... rest of handleInit
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
#### Step 3: Deduplicate /api/settings fetch
|
||||
**Impact**: MEDIUM — removes 1 redundant API call
|
||||
**Files**: `src/web/public/app.js:7341` (loadQuickStartCases), `app.js:9964` (loadAppSettingsFromServer)
|
||||
|
||||
```js
|
||||
// In init() — fetch settings once, share the promise:
|
||||
const settingsPromise = fetch('/api/settings').then(r => r.json());
|
||||
this.loadQuickStartCases(null, settingsPromise);
|
||||
this.loadAppSettingsFromServer(settingsPromise);
|
||||
```
|
||||
|
||||
Both functions need to accept an optional pre-fetched settings promise parameter.
|
||||
|
||||
---
|
||||
|
||||
#### Step 4: Remove cache-busting from /api/cases
|
||||
**Impact**: LOW — allows browser caching
|
||||
**File**: `src/web/public/app.js:7351`
|
||||
|
||||
```js
|
||||
// BEFORE
|
||||
const res = await fetch('/api/cases?_t=' + Date.now());
|
||||
// AFTER
|
||||
const res = await fetch('/api/cases');
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
#### Step 5: Defer system stats polling
|
||||
**Impact**: MEDIUM — removes API call every 2s when idle
|
||||
**Files**: `src/web/public/app.js:1567`, `app.js:15261`
|
||||
|
||||
Move `startSystemStatsPolling()` out of `init()`. Start it in `handleInit()` only when `data.sessions.length > 0`.
|
||||
|
||||
---
|
||||
|
||||
### Phase 2: Build Pipeline (30-min changes, highest payload impact)
|
||||
|
||||
#### Step 6: Self-host xterm.js assets
|
||||
**Impact**: MEDIUM-HIGH — eliminates CDN DNS/TLS latency (~100ms even with preconnect)
|
||||
**Files**: `src/web/public/index.html`, `package.json` build script
|
||||
|
||||
xterm is NOT in package.json — add it:
|
||||
```bash
|
||||
npm install xterm@5.3.0 @xterm/addon-fit@0.8.0 --save
|
||||
```
|
||||
|
||||
Build script addition:
|
||||
```bash
|
||||
mkdir -p dist/web/public/vendor
|
||||
cp node_modules/xterm/css/xterm.css dist/web/public/vendor/
|
||||
cp node_modules/xterm/lib/xterm.min.js dist/web/public/vendor/
|
||||
cp node_modules/@xterm/addon-fit/lib/xterm-addon-fit.min.js dist/web/public/vendor/
|
||||
```
|
||||
|
||||
Update index.html CDN URLs to `/vendor/xterm.min.js` etc. Remove preconnect/dns-prefetch for jsdelivr.
|
||||
|
||||
---
|
||||
|
||||
#### Step 7: Add esbuild minification to build
|
||||
**Impact**: HIGH — biggest single optimization for payload size
|
||||
**File**: `package.json` build script
|
||||
|
||||
Current build just does `cp -r src/web/public dist/web/`. No minification.
|
||||
|
||||
app.js stats: 1,525 comment lines (10%), 89 console.* statements, 23% whitespace.
|
||||
|
||||
```bash
|
||||
# Add to build script after cp:
|
||||
npx esbuild dist/web/public/app.js --minify --drop:console --outfile=dist/web/public/app.js --allow-overwrite
|
||||
npx esbuild dist/web/public/styles.css --minify --outfile=dist/web/public/styles.css --allow-overwrite
|
||||
npx esbuild dist/web/public/mobile.css --minify --outfile=dist/web/public/mobile.css --allow-overwrite
|
||||
```
|
||||
|
||||
Expected savings:
|
||||
| File | Before (gzip) | After (gzip) | Saved |
|
||||
|------|---------------|-------------|-------|
|
||||
| app.js | ~126 KB | ~85 KB | ~41 KB (33%) |
|
||||
| styles.css | ~25 KB | ~18 KB | ~7 KB (28%) |
|
||||
| mobile.css | ~7 KB | ~5 KB | ~2 KB (29%) |
|
||||
| **Total** | **~158 KB** | **~108 KB** | **~50 KB** |
|
||||
|
||||
---
|
||||
|
||||
#### Step 8: Pre-compress static assets at build time
|
||||
**Impact**: MEDIUM — eliminates per-request CPU compression
|
||||
**Files**: `package.json` build script, potentially `src/web/server.ts`
|
||||
|
||||
```bash
|
||||
# Add to build script after minification:
|
||||
for f in dist/web/public/*.{js,css,html}; do
|
||||
gzip -9 -k "$f"
|
||||
brotli -9 -k "$f"
|
||||
done
|
||||
```
|
||||
|
||||
Check if `@fastify/static` supports `preCompressed: true` option. If not, serve pre-compressed files via custom Accept-Encoding check.
|
||||
|
||||
---
|
||||
|
||||
#### Step 9: Extend cache duration for versioned assets
|
||||
**Impact**: LOW (first load) / HIGH (repeat visits)
|
||||
**File**: `src/web/server.ts:601`
|
||||
|
||||
```js
|
||||
// BEFORE
|
||||
maxAge: '1h'
|
||||
|
||||
// AFTER
|
||||
maxAge: '1y',
|
||||
immutable: true
|
||||
```
|
||||
|
||||
Safe because all assets use `?v=0.1536` cache-busting. First-load unaffected, but all repeat visits serve from disk cache instantly.
|
||||
|
||||
---
|
||||
|
||||
### Phase 3: Perceived Performance (30-60min, user experience)
|
||||
|
||||
#### Step 10: Add loading skeleton
|
||||
**Impact**: MEDIUM-HIGH — instant visual structure instead of black screen
|
||||
**File**: `src/web/public/index.html`
|
||||
|
||||
Add minimal inline `<style>` + skeleton HTML in `<body>`:
|
||||
|
||||
```html
|
||||
<style>
|
||||
.skeleton { display: flex; flex-direction: column; height: 100vh; background: #0a0a0a; }
|
||||
.skeleton-header { height: 40px; background: #111; border-bottom: 1px solid #222; }
|
||||
.skeleton-terminal { flex: 1; background: #0d0d0d; }
|
||||
.app-loaded .skeleton { display: none; }
|
||||
</style>
|
||||
<div class="skeleton">
|
||||
<div class="skeleton-header"></div>
|
||||
<div class="skeleton-terminal"></div>
|
||||
</div>
|
||||
```
|
||||
|
||||
In `app.js` init() end: `document.body.classList.add('app-loaded');`
|
||||
|
||||
---
|
||||
|
||||
#### Step 11: Defer terminal creation to after first paint
|
||||
**Impact**: MEDIUM-HIGH — terminal.open() is heaviest sync call
|
||||
**File**: `src/web/public/app.js:1545`
|
||||
|
||||
```js
|
||||
init() {
|
||||
// ... mobile detection, visibility settings ...
|
||||
document.documentElement.classList.remove('mobile-init');
|
||||
|
||||
// Show skeleton immediately, defer heavy terminal init
|
||||
requestAnimationFrame(() => {
|
||||
this.initTerminal();
|
||||
this.connectSSE();
|
||||
// ... rest of init
|
||||
});
|
||||
}
|
||||
```
|
||||
|
||||
Lets browser paint header/tabs/skeleton before canvas creation.
|
||||
|
||||
---
|
||||
|
||||
### Phase 4: DOM + Payload Reduction (1-3 hours)
|
||||
|
||||
#### Step 12: Lazy-create modals on first open
|
||||
**Impact**: HIGH — removes ~600+ DOM nodes, ~60KB hidden HTML
|
||||
**Files**: `src/web/public/index.html`, `src/web/public/app.js`
|
||||
|
||||
12 modals pre-rendered in index.html:
|
||||
- `helpModal` (lines 227-447)
|
||||
- `sessionOptionsModal` (lines 448-714) — 266 lines
|
||||
- `appSettingsModal` (lines 715-900+)
|
||||
- `createCaseModal`, `mobileCasePickerModal`, `ralphWizardModal`, `killAllModal`, `closeConfirmModal`, `savePresetModal`, `tokenStatsModal`, `filePreviewModal`, notification drawer
|
||||
|
||||
Replace each modal's HTML with `<div id="helpModal" class="modal"></div>`. On first open, inject full HTML via template function. Cache after creation.
|
||||
|
||||
---
|
||||
|
||||
#### Step 13: Batch initial API calls into one endpoint
|
||||
**Impact**: MEDIUM — reduces 4+ API calls to 1
|
||||
**Files**: `src/web/server.ts`, `src/web/public/app.js`
|
||||
|
||||
Create `GET /api/init-bundle`:
|
||||
```json
|
||||
{
|
||||
"status": { /* getLightState() */ },
|
||||
"cases": [ /* case list */ ],
|
||||
"settings": { /* user settings */ }
|
||||
}
|
||||
```
|
||||
|
||||
Use as SSE init fallback (step 2's timeout). Saves HTTP round trips.
|
||||
|
||||
---
|
||||
|
||||
#### Step 14: Trim SSE init payload
|
||||
**Impact**: LOW-MEDIUM — reduces init payload by removing data not needed for first paint
|
||||
**File**: `src/web/server.ts`
|
||||
|
||||
Remove from SSE init event: `taskTree`, `ralphTodos`, `ralphTodoStats` per session. These can be fetched on-demand when user opens a session's details panel.
|
||||
|
||||
---
|
||||
|
||||
#### Step 15: Enable HTTP/2
|
||||
**Impact**: MEDIUM — multiplexed loading over single connection
|
||||
**File**: `src/web/server.ts`
|
||||
|
||||
```js
|
||||
// BEFORE
|
||||
const server = Fastify({ logger: false });
|
||||
|
||||
// AFTER (when HTTPS is enabled)
|
||||
const server = Fastify({
|
||||
logger: false,
|
||||
http2: true // Only works with HTTPS
|
||||
});
|
||||
```
|
||||
|
||||
Only applicable for `--https` mode. HTTP/2 multiplexing eliminates the 6-connection limit queuing.
|
||||
|
||||
---
|
||||
|
||||
## Expected Combined Impact
|
||||
|
||||
| Metric | Before | After | Improvement |
|
||||
|--------|--------|-------|-------------|
|
||||
| First Paint | ~300ms | ~100ms | **-200ms** (skeleton visible instantly) |
|
||||
| First Contentful Paint | ~400ms | ~200ms | **-200ms** (no mobile.css blocking desktop) |
|
||||
| Time to Interactive | ~600ms | ~350ms | **-250ms** (fewer API calls, deferred terminal) |
|
||||
| Total compressed payload | ~241 KB | ~191 KB | **-50 KB (21%)** via minification |
|
||||
| Init API calls | 6-7 (2 dupes) | 2-3 | **-60%** fewer requests |
|
||||
| Initial DOM nodes | ~1800+ | ~1200 | **-600** (lazy modals) |
|
||||
|
||||
---
|
||||
|
||||
## Verification
|
||||
|
||||
After each step, verify with Playwright:
|
||||
|
||||
```js
|
||||
const { chromium } = require('playwright');
|
||||
const browser = await chromium.launch();
|
||||
const page = await browser.newPage();
|
||||
|
||||
// Measure first paint
|
||||
await page.goto('http://localhost:3000', { waitUntil: 'domcontentloaded' });
|
||||
await page.waitForTimeout(4000); // Wait for async data
|
||||
|
||||
// Check UI renders correctly
|
||||
const header = await page.locator('.header').isVisible();
|
||||
const tabs = await page.locator('.session-tabs').isVisible();
|
||||
const terminal = await page.locator('.terminal-container').isVisible();
|
||||
console.log({ header, tabs, terminal });
|
||||
|
||||
await browser.close();
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Files Changed Per Step (for implementation agent)
|
||||
|
||||
| Step | Files Modified |
|
||||
|------|---------------|
|
||||
| 1 | `index.html` |
|
||||
| 2 | `app.js` |
|
||||
| 3 | `app.js` |
|
||||
| 4 | `app.js` |
|
||||
| 5 | `app.js` |
|
||||
| 6 | `index.html`, `package.json` |
|
||||
| 7 | `package.json` |
|
||||
| 8 | `package.json`, optionally `server.ts` |
|
||||
| 9 | `server.ts` |
|
||||
| 10 | `index.html`, `app.js` |
|
||||
| 11 | `app.js` |
|
||||
| 12 | `index.html`, `app.js` |
|
||||
| 13 | `server.ts`, `app.js` |
|
||||
| 14 | `server.ts` |
|
||||
| 15 | `server.ts` |
|
||||
@@ -0,0 +1,388 @@
|
||||
# Performance Audit: First Page Load
|
||||
|
||||
**Date**: 2026-02-18
|
||||
**Scope**: Browser first-load of Codeman web UI (`/`)
|
||||
**Method**: Static analysis by 4 parallel audit agents (server, frontend, SSE/xterm, asset pipeline)
|
||||
|
||||
---
|
||||
|
||||
## Current State Summary
|
||||
|
||||
### Payload Sizes (measured from live server, port 3000)
|
||||
|
||||
| Asset | Raw Size | Gzip | Brotli | Lines | Render-Blocking? |
|
||||
|-------|----------|------|--------|-------|-----------------|
|
||||
| `index.html` | 82 KB | 15 KB | 15 KB | 1,479 | N/A (document) |
|
||||
| `app.js` | 562 KB | 126 KB | 125 KB | 15,354 | No (`defer`) |
|
||||
| `styles.css` | 154 KB | 25 KB | 27 KB | 8,199 | **YES** |
|
||||
| `mobile.css` | 34 KB | 7 KB | 7 KB | 1,493 | **YES** (no media query!) |
|
||||
| `xterm.css` (CDN) | 2 KB | 2 KB | — | — | **YES** (external CDN) |
|
||||
| `xterm.min.js` (CDN) | 67 KB | 65 KB | — | — | No (`defer`) |
|
||||
| `xterm-addon-fit` (CDN) | 1 KB | 1 KB | — | — | No (`defer`) |
|
||||
| **Total local** | **832 KB** | **173 KB** | **174 KB** | | |
|
||||
| **Total w/ CDN** | **~902 KB** | **~241 KB** | | | |
|
||||
|
||||
**Server compression**: Brotli preferred (`Content-Encoding: br`), via `@fastify/compress` with threshold 1024. Compression is **on-the-fly per request** — no pre-compressed files exist.
|
||||
|
||||
**HTTP headers verified**: `Cache-Control: public, max-age=3600`, weak ETags auto-generated by `@fastify/static`, `Vary: accept-encoding`, CSP + security headers present.
|
||||
|
||||
### Request Waterfall on First Load (6-7 API calls!)
|
||||
|
||||
```
|
||||
Browser hits /
|
||||
├── index.html ............................ (82 KB document)
|
||||
├── styles.css?v=0.1533 .................. (render-blocking CSS, 154 KB)
|
||||
├── mobile.css?v=0.1533 .................. (render-blocking CSS, 34 KB — wasted on desktop!)
|
||||
├── xterm.css (CDN) ...................... (render-blocking CSS — external!)
|
||||
├── xterm.min.js (CDN, defer) ........... (67 KB, parallel download)
|
||||
├── xterm-addon-fit.min.js (CDN, defer) .. (1 KB, parallel download)
|
||||
├── app.js?v=0.1533 (defer) ............. (562 KB, parallel download)
|
||||
│
|
||||
│ [FIRST PAINT blocked until ALL CSS downloaded + parsed]
|
||||
│
|
||||
├── JS executes: new CodemanApp().init()
|
||||
│ ├── initTerminal() ................... (SYNC: new Terminal() + terminal.open() → canvas creation)
|
||||
│ ├── connectSSE() → /api/events ....... (SSE → fires 'init' with getLightState())
|
||||
│ ├── loadState() → /api/status ........ (DUPLICATE #1: same data as SSE init!)
|
||||
│ ├── loadQuickStartCases()
|
||||
│ │ ├── /api/settings ................ (settings fetch #1)
|
||||
│ │ └── /api/cases?_t=<timestamp> ... (case list, cache-busted!)
|
||||
│ ├── startSystemStatsPolling() → /api/system/stats (every 2s, starts immediately)
|
||||
│ └── loadAppSettingsFromServer() → /api/settings (DUPLICATE #2: settings fetched again!)
|
||||
```
|
||||
|
||||
**Total init API calls**: 6-7 requests, with **2 duplicates** (`/api/status` = SSE init, `/api/settings` fetched twice).
|
||||
|
||||
### Critical Path Bottlenecks
|
||||
|
||||
1. **3 render-blocking CSS files** (one from CDN, one wasted on desktop)
|
||||
2. **Synchronous `terminal.open()`** blocks main thread during init (canvas creation)
|
||||
3. **Double `handleInit()` execution** — SSE init + `/api/status` both call it, causing full state reset + cleanup twice within ~100ms
|
||||
4. **`/api/settings` fetched twice** — once in `loadQuickStartCases()`, once in `loadAppSettingsFromServer()`
|
||||
5. **No loading skeleton** — blank `#0a0a0a` screen until CSS+JS fully loaded
|
||||
6. **12 modals pre-rendered** in HTML — ~600+ DOM elements, ~60KB of invisible HTML
|
||||
7. **562KB monolith `app.js`** unminified — 1,525 comment lines (10%), 89 `console.*` statements, 23% whitespace
|
||||
8. **No minification in build** — `cp -r` copies raw source to dist
|
||||
9. **Stats polling starts immediately** — 2s interval even with no sessions
|
||||
10. **Version query strings stale** — HTML has `?v=0.1533`, package.json is `0.1534`
|
||||
|
||||
### What's Already Good
|
||||
|
||||
- Only **1 xterm Terminal instance** shared across all sessions (buffer swapping on tab switch)
|
||||
- Teammate terminals created **lazily** on window open (with `requestAnimationFrame` defer)
|
||||
- Subagent windows use **HTML activity logs**, not additional Terminal instances
|
||||
- `getLightState()` has a **1-second TTL cache** — no duplicate server-side computation
|
||||
- SSE init sends **lightweight state** (no terminal buffers) — buffers fetched on-demand per tab
|
||||
- Buffer hydration uses **chunked writes** (128KB chunks via `requestAnimationFrame`) — no UI jank
|
||||
- `selectSession()` defers secondary panels via **`requestIdleCallback`**
|
||||
- Buffer fetch is **tail-mode** (last 256KB only, not full 2MB)
|
||||
- **System fonts only** — no web font downloads blocking paint
|
||||
- All JS scripts use **`defer`**
|
||||
- SSE reconnection has **proper exponential backoff** with timeout cleanup
|
||||
|
||||
---
|
||||
|
||||
## Optimization Plan
|
||||
|
||||
### Phase 1: Quick Wins (High Impact, Low Effort)
|
||||
|
||||
#### 1.1 Add `media` attribute to mobile.css
|
||||
**Impact**: HIGH — 34KB CSS stops blocking render on desktop
|
||||
**Effort**: 1 line change
|
||||
**File**: `src/web/public/index.html:14`
|
||||
|
||||
```html
|
||||
<!-- Before -->
|
||||
<link rel="stylesheet" href="mobile.css?v=...">
|
||||
|
||||
<!-- After -->
|
||||
<link rel="stylesheet" href="mobile.css?v=..." media="(max-width: 1023px)">
|
||||
```
|
||||
|
||||
The browser still downloads it (for potential resize) but won't block rendering on desktop. The `mobile.css` comment on line 4 says this was *intended* but never implemented.
|
||||
|
||||
#### 1.2 Eliminate duplicate `/api/status` fetch + double `handleInit()`
|
||||
**Impact**: HIGH — removes 1 redundant API call + eliminates double state reset (clearing 15+ Maps, 7+ timers, `cleanupAllFloatingWindows()`, double `renderSessionTabs()`, double async subagent restore chain)
|
||||
**Effort**: Small
|
||||
**Files**: `src/web/public/app.js:1554`, `app.js:3566-3574`
|
||||
|
||||
The SSE `init` event (`server.ts:618`) already sends `getLightState()`. The `loadState()` at `app.js:1554` fetches identical data from `/api/status`. Both call `handleInit()` which does a full state reset — whichever arrives second **wipes all state from the first** and rebuilds from scratch.
|
||||
|
||||
The `_initGeneration` guard (line 3373/3549) only protects the session-restore at the end, NOT the expensive full cleanup (lines 3389-3503).
|
||||
|
||||
**Approach**: Remove `this.loadState()` from `init()`. Add a fallback timeout:
|
||||
|
||||
```js
|
||||
// In init():
|
||||
this.connectSSE();
|
||||
// Remove: this.loadState();
|
||||
this._initFallbackTimer = setTimeout(() => {
|
||||
if (this._initGeneration === 0) this.loadState();
|
||||
}, 3000);
|
||||
```
|
||||
|
||||
Clear the timer in `handleInit()`:
|
||||
```js
|
||||
handleInit(data) {
|
||||
if (this._initFallbackTimer) {
|
||||
clearTimeout(this._initFallbackTimer);
|
||||
this._initFallbackTimer = null;
|
||||
}
|
||||
// ... rest of handleInit
|
||||
}
|
||||
```
|
||||
|
||||
#### 1.3 Deduplicate `/api/settings` fetch
|
||||
**Impact**: MEDIUM — removes 1 redundant API call
|
||||
**Effort**: Small
|
||||
**Files**: `src/web/public/app.js:7341` (in `loadQuickStartCases`), `app.js:9964` (in `loadAppSettingsFromServer`)
|
||||
|
||||
Both fetch `/api/settings`. Fetch it once, pass the result to both consumers:
|
||||
|
||||
```js
|
||||
// In init():
|
||||
const settingsPromise = fetch('/api/settings').then(r => r.json());
|
||||
this.loadQuickStartCases(null, settingsPromise);
|
||||
this.loadAppSettingsFromServer(settingsPromise);
|
||||
```
|
||||
|
||||
#### 1.4 Defer system stats polling
|
||||
**Impact**: MEDIUM — removes 1 API call every 2s when idle
|
||||
**Effort**: Small
|
||||
**Files**: `src/web/public/app.js:1567`, `app.js:15261-15271`
|
||||
|
||||
`fetchSystemStats()` already has a visibility guard (line 15282: skips if `#headerSystemStats` is `display: none`), but the interval still ticks. Move `startSystemStatsPolling()` out of `init()` — start it in `handleInit()` only when `data.sessions.length > 0`.
|
||||
|
||||
#### 1.5 Preload xterm.css to unblock render
|
||||
**Impact**: MEDIUM — external CDN CSS currently blocks first paint
|
||||
**Effort**: 2 line change
|
||||
**File**: `src/web/public/index.html:15`
|
||||
|
||||
```html
|
||||
<!-- Before -->
|
||||
<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css">
|
||||
|
||||
<!-- After -->
|
||||
<link rel="preload" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css" as="style" onload="this.onload=null;this.rel='stylesheet'">
|
||||
<noscript><link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/xterm@5.3.0/css/xterm.css"></noscript>
|
||||
```
|
||||
|
||||
Terminal won't display until xterm.js executes anyway, so the CSS doesn't need to block initial paint.
|
||||
|
||||
#### 1.6 Fix stale version query strings
|
||||
**Impact**: LOW — prevents serving cached stale assets after deploy
|
||||
**Effort**: Small
|
||||
**File**: COM script in CLAUDE.md
|
||||
|
||||
The HTML references `?v=0.1533` while package.json is already at `0.1534`. The COM workflow should auto-update HTML version strings. Add to the COM script:
|
||||
|
||||
```bash
|
||||
# After incrementing version in package.json + CLAUDE.md:
|
||||
sed -i "s/?v=[0-9.]*/?v=$NEW_VERSION/g" src/web/public/index.html
|
||||
```
|
||||
|
||||
#### 1.7 Remove cache-busting from `/api/cases`
|
||||
**Impact**: LOW — allows HTTP caching of case list
|
||||
**Effort**: 1 line change
|
||||
**File**: `src/web/public/app.js:7351`
|
||||
|
||||
```js
|
||||
// Before:
|
||||
const res = await fetch('/api/cases?_t=' + Date.now());
|
||||
// After:
|
||||
const res = await fetch('/api/cases');
|
||||
```
|
||||
|
||||
The case list rarely changes during a session. Let the browser cache it.
|
||||
|
||||
---
|
||||
|
||||
### Phase 2: Medium Effort (High Impact)
|
||||
|
||||
#### 2.1 Add loading skeleton
|
||||
**Impact**: MEDIUM-HIGH — perceived performance improvement (instant visual structure)
|
||||
**Effort**: Small-Medium
|
||||
**File**: `src/web/public/index.html`
|
||||
|
||||
Add minimal inline `<style>` + skeleton HTML in `<body>` showing a dark header bar + terminal placeholder. Hidden by `app.js` once init completes:
|
||||
|
||||
```html
|
||||
<style>
|
||||
.skeleton { display: flex; flex-direction: column; height: 100vh; }
|
||||
.skeleton-header { height: 40px; background: #111; border-bottom: 1px solid #222; }
|
||||
.skeleton-terminal { flex: 1; background: #0d0d0d; }
|
||||
.app-loaded .skeleton { display: none; }
|
||||
</style>
|
||||
<div class="skeleton">
|
||||
<div class="skeleton-header"></div>
|
||||
<div class="skeleton-terminal"></div>
|
||||
</div>
|
||||
```
|
||||
|
||||
In `app.js` init(), add `document.body.classList.add('app-loaded')` at the end.
|
||||
|
||||
#### 2.2 Defer xterm.js terminal creation to after first paint
|
||||
**Impact**: MEDIUM-HIGH — `terminal.open()` is the heaviest synchronous call in init
|
||||
**Effort**: Medium
|
||||
**Files**: `src/web/public/app.js:1545`, `app.js:1578-1639`
|
||||
|
||||
```js
|
||||
init() {
|
||||
// ... mobile detection, visibility settings ...
|
||||
document.documentElement.classList.remove('mobile-init');
|
||||
|
||||
// Show skeleton/header immediately, defer heavy terminal init
|
||||
requestAnimationFrame(() => {
|
||||
this.initTerminal();
|
||||
this.connectSSE();
|
||||
// ... rest of init
|
||||
});
|
||||
}
|
||||
```
|
||||
|
||||
Lets the browser paint the header/tabs before the terminal canvas is created.
|
||||
|
||||
#### 2.3 Batch initial API calls into one endpoint
|
||||
**Impact**: MEDIUM — reduces 4+ API calls to 1
|
||||
**Effort**: Medium
|
||||
**Files**: `src/web/server.ts`, `src/web/public/app.js`
|
||||
|
||||
Create `/api/init-bundle`:
|
||||
```json
|
||||
{
|
||||
"status": { /* getLightState() */ },
|
||||
"cases": [ /* case list */ ],
|
||||
"settings": { /* user settings */ }
|
||||
}
|
||||
```
|
||||
|
||||
Replaces `/api/status` (fallback), `/api/cases`, `/api/settings`. Saves HTTP round trips and server-side work.
|
||||
|
||||
#### 2.4 Lazy-create modals on first open
|
||||
**Impact**: HIGH — removes ~600+ DOM elements from initial parse (~60KB of HTML)
|
||||
**Effort**: Medium-High
|
||||
**Files**: `src/web/public/index.html`, `src/web/public/app.js`
|
||||
|
||||
12 modals pre-rendered in `index.html`:
|
||||
- `helpModal` (lines 227-447)
|
||||
- `sessionOptionsModal` (lines 448-714) — **266 lines alone**
|
||||
- `appSettingsModal` (lines 715-900+)
|
||||
- `createCaseModal`, `mobileCasePickerModal`, `ralphWizardModal`, `killAllModal`, `closeConfirmModal`, `savePresetModal`, `tokenStatsModal`, `filePreviewModal`, notification drawer
|
||||
|
||||
**Approach**: Replace each modal's HTML with `<div id="helpModal" class="modal"></div>`. On first open, inject full HTML via `createModalContent()`. Cache after creation.
|
||||
|
||||
---
|
||||
|
||||
### Phase 3: Build Pipeline (Highest Impact)
|
||||
|
||||
#### 3.1 Self-host xterm.js assets
|
||||
**Impact**: MEDIUM — eliminates CDN dependency + latency, enables local caching
|
||||
**Effort**: Low-Medium
|
||||
**Files**: `src/web/public/index.html`, `package.json` build script
|
||||
|
||||
```bash
|
||||
# Build script addition:
|
||||
mkdir -p dist/web/public/vendor
|
||||
cp node_modules/xterm/css/xterm.css dist/web/public/vendor/
|
||||
cp node_modules/xterm/lib/xterm.min.js dist/web/public/vendor/
|
||||
cp node_modules/@xterm/addon-fit/lib/xterm-addon-fit.min.js dist/web/public/vendor/
|
||||
```
|
||||
|
||||
Update HTML to reference `/vendor/xterm.min.js` etc. Removes render-blocking CDN CSS entirely.
|
||||
|
||||
#### 3.2 Add esbuild minification to build
|
||||
**Impact**: HIGH — ~38 KB compressed savings (16% of local payload)
|
||||
**Effort**: Medium
|
||||
**Files**: `package.json` (build script)
|
||||
|
||||
Current build just does `cp -r src/web/public dist/web/`. No minification at all.
|
||||
|
||||
**app.js specifics**: 1,525 comment lines (10%), 89 `console.*` statements, 23% whitespace.
|
||||
|
||||
```bash
|
||||
# Add to build script:
|
||||
npx esbuild dist/web/public/app.js --minify --drop:console --outfile=dist/web/public/app.js --allow-overwrite
|
||||
npx esbuild dist/web/public/styles.css --minify --outfile=dist/web/public/styles.css --allow-overwrite
|
||||
npx esbuild dist/web/public/mobile.css --minify --outfile=dist/web/public/mobile.css --allow-overwrite
|
||||
```
|
||||
|
||||
Expected: `app.js` 562KB → ~350KB minified → ~90KB gzip (from 126KB). `--drop:console` removes all 89 debug statements.
|
||||
|
||||
Note: `app.js` is vanilla JS (not modules), so esbuild works directly as a minifier.
|
||||
|
||||
#### 3.3 Pre-compress static assets at build time
|
||||
**Impact**: MEDIUM — eliminates per-request CPU compression work
|
||||
**Effort**: Low
|
||||
**Files**: `package.json` build script, `src/web/server.ts`
|
||||
|
||||
Currently `@fastify/compress` compresses on-the-fly for every request. Pre-compress at build time:
|
||||
|
||||
```bash
|
||||
# Build script:
|
||||
for f in dist/web/public/*.{js,css,html}; do
|
||||
gzip -9 -k "$f"
|
||||
brotli -9 -k "$f"
|
||||
done
|
||||
```
|
||||
|
||||
Then configure `@fastify/static` with `preCompressed: true` (if supported) or serve pre-compressed files via custom logic.
|
||||
|
||||
#### 3.4 Extract critical CSS inline
|
||||
**Impact**: MEDIUM — eliminates render-blocking `styles.css` for first paint
|
||||
**Effort**: Medium-High
|
||||
**Files**: `src/web/public/styles.css`, `src/web/public/index.html`
|
||||
|
||||
Identify ~2-3KB of CSS needed for first paint (body, header, tab bar, terminal container) and inline it in `<head>`. Load full `styles.css` asynchronously:
|
||||
|
||||
```html
|
||||
<style>/* ~50 lines of critical CSS */</style>
|
||||
<link rel="preload" href="styles.css?v=..." as="style" onload="this.onload=null;this.rel='stylesheet'">
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Impact Estimates
|
||||
|
||||
| # | Optimization | First Paint | TTI | Effort |
|
||||
|---|-------------|-------------|-----|--------|
|
||||
| 1.1 | mobile.css media query | -50ms | — | 1 min |
|
||||
| 1.2 | Remove duplicate fetch + double handleInit | — | -100-200ms | 15 min |
|
||||
| 1.3 | Deduplicate settings fetch | — | -50ms | 10 min |
|
||||
| 1.4 | Defer stats polling | — | -20ms | 10 min |
|
||||
| 1.5 | Preload xterm.css | -100-300ms | — | 5 min |
|
||||
| 1.6 | Fix stale version strings | cache correctness | — | 5 min |
|
||||
| 1.7 | Remove cases cache-bust | — | -10ms | 1 min |
|
||||
| 2.1 | Loading skeleton | perceived -500ms | — | 30 min |
|
||||
| 2.2 | Defer terminal init | -50-100ms | -50ms | 30 min |
|
||||
| 2.3 | Batch API endpoint | — | -100-200ms | 1 hr |
|
||||
| 2.4 | Lazy modals | -30-50ms parse | -50ms | 2-3 hrs |
|
||||
| 3.1 | Self-host xterm | -100-300ms | — | 20 min |
|
||||
| 3.2 | Minify JS/CSS | -50-100ms parse | — | 30 min |
|
||||
| 3.3 | Pre-compress assets | -10-30ms TTFB | — | 20 min |
|
||||
| 3.4 | Critical CSS inline | -200-400ms | — | 2 hrs |
|
||||
|
||||
**Combined estimate**: First paint **300-800ms faster**, TTI **200-500ms faster**.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Order (for implementation agent)
|
||||
|
||||
Do these in order — each step is independently testable:
|
||||
|
||||
1. **1.1** — mobile.css media query (1 line, instant win)
|
||||
2. **1.5** — Preload xterm.css (2 lines, big render-blocking fix)
|
||||
3. **1.2** — Remove duplicate `/api/status` + double handleInit
|
||||
4. **1.3** — Deduplicate `/api/settings` fetch
|
||||
5. **1.7** — Remove cache-busting from `/api/cases`
|
||||
6. **3.1** — Self-host xterm.js (removes CDN dependency entirely)
|
||||
7. **3.2** — Add esbuild minification to build
|
||||
8. **1.4** — Defer stats polling
|
||||
9. **1.6** — Fix stale version strings in COM workflow
|
||||
10. **2.1** — Loading skeleton
|
||||
11. **2.2** — Defer terminal init after first paint
|
||||
12. **2.3** — Batch init API endpoint
|
||||
13. **2.4** — Lazy modals (biggest refactor, do last)
|
||||
14. **3.3** — Pre-compress assets (nice-to-have)
|
||||
15. **3.4** — Critical CSS extraction (only if still needed after above)
|
||||
|
||||
**Verification after each step**: Use Playwright to load the page with `waitUntil: 'domcontentloaded'`, measure first paint timing, check that the UI renders correctly with 3-4s wait for async data.
|
||||
@@ -0,0 +1,423 @@
|
||||
# Performance Analysis & Optimization Opportunities
|
||||
|
||||
**Date**: 2026-03-07
|
||||
**Scope**: Full-stack performance analysis — backend PTY handling, SSE broadcasting, frontend terminal rendering, local echo overlay, DOM updates, config/scaling limits.
|
||||
**Constraint**: All recommendations preserve existing functionality including local echo, backpressure, anti-flicker pipeline, and mobile support.
|
||||
|
||||
---
|
||||
|
||||
## Executive Summary
|
||||
|
||||
The codebase is already well-optimized in critical paths. The multi-layer backpressure system, adaptive terminal batching, DEC 2026 sync markers, and incremental state serialization are strong. The main opportunities are in **reducing unnecessary work** (SSE filtering, DOM rebuilds, lazy terminal init) rather than algorithmic changes.
|
||||
|
||||
**Top 5 high-impact opportunities:**
|
||||
|
||||
| # | Optimization | Impact | Risk | Effort |
|
||||
|---|-------------|--------|------|--------|
|
||||
| 1 | Session-scoped SSE subscriptions | Bandwidth -60-80%, CPU -40% | Medium | Medium |
|
||||
| 2 | Lazy xterm.js for minimized subagent windows | Memory -3.5MB at 50 agents | Low | Low |
|
||||
| 3 | Targeted badge update (skip full tab rebuild) | Eliminates O(n) reflow on badge change | Low | Low |
|
||||
| 4 | Conditional SSE padding (tunnel-only, terminal-only) | Bandwidth -70% when tunneled | Low | Low |
|
||||
| 5 | Canvas renderer on mobile | GPU pressure reduction, battery savings | Low | Low |
|
||||
|
||||
---
|
||||
|
||||
## 1. SSE Broadcasting
|
||||
|
||||
### Current State
|
||||
- **92 event types** broadcast to all connected clients (max 100)
|
||||
- Single `JSON.stringify()` per event, shared across all clients (efficient)
|
||||
- **No per-client filtering** — every client receives every event regardless of which session they're viewing
|
||||
- 8KB padding appended to **every** event when tunnel is active (forces Cloudflare proxy flush)
|
||||
- Backpressure: clients marked as backpressured if `reply.raw.write()` returns false; recovery via `session:needsRefresh`
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B1: No session-scoped SSE subscriptions** (`server.ts:1986`)
|
||||
- Client viewing session A still receives all events for sessions B through T
|
||||
- With 20 active sessions, ~95% of terminal events are irrelevant to any given client
|
||||
- Cost: wasted bandwidth, CPU for JSON parsing, and event handler dispatch on client
|
||||
|
||||
**B2: Unconditional 8KB padding** (`server.ts:1977`)
|
||||
- Every event gets 8KB comment padding when tunnel is active
|
||||
- A `task:updated` event (~200 bytes payload) becomes ~8.2KB
|
||||
- High-frequency events like `session:terminal` need the padding; low-frequency events like `session:created` don't
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R1: Session-scoped SSE subscriptions** (High impact)
|
||||
- Add `?sessions=id1,id2` query param to `/api/events` SSE endpoint
|
||||
- Server filters events by session ID before broadcasting
|
||||
- Client subscribes to active session + "global" events (session lifecycle, system)
|
||||
- Re-subscribes on tab switch (or subscribe to all with client-side filter as fallback)
|
||||
- **Savings**: ~80% bandwidth reduction for single-session viewers; ~60% for multi-session dashboards
|
||||
|
||||
**R2: Tiered SSE padding** (Medium impact)
|
||||
- Only pad `session:terminal` events and SSE heartbeats (the two that need proxy flush)
|
||||
- Skip padding for low-frequency structural events (`session:created`, `task:updated`, etc.)
|
||||
- **Savings**: ~70% padding overhead reduction; terminal events already large enough to flush
|
||||
|
||||
---
|
||||
|
||||
## 2. Terminal Rendering
|
||||
|
||||
### Current State (Well-Optimized)
|
||||
- **6-layer anti-flicker pipeline**: Server batching (adaptive 16-50ms) → DEC 2026 sync wrap → single JSON serialize → client rAF batching → sync segment parser → chunked buffer loading (32KB/frame)
|
||||
- **64KB/frame write budget** with DEC 2026 sync-segment awareness (prevents 141KB single-frame freezes)
|
||||
- **3-layer backpressure**: SSE cap (128KB queued → drop + refresh), frame budget (64KB/frame), chunked restore (32KB/frame)
|
||||
- WebGL renderer enabled by default with canvas fallback on context loss
|
||||
- Typical latency: 16-32ms; worst case: ~115ms (50ms server batch + 50ms sync wait + 16ms rAF)
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B3: WebGL on mobile** (`app.js:627-637`)
|
||||
- Mobile GPUs are weaker; WebGL context loss more likely on low-end devices
|
||||
- Canvas renderer is sufficient for mobile (typically 1 session, smaller viewport)
|
||||
|
||||
**B4: Static scrollback for all sessions** (`app.js:572`)
|
||||
- Default 5000 lines scrollback for all sessions regardless of activity level
|
||||
- Heavy output sessions (build logs, test runners) accumulate large scroll buffers
|
||||
|
||||
**B5: No addon lazy loading**
|
||||
- FitAddon, Unicode11Addon, and WebGLAddon all loaded at terminal init
|
||||
- Unicode11Addon only needed for CJK content; WebGLAddon is large
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R3: Force canvas renderer on mobile** (Low risk)
|
||||
- Detect `MobileDetection.isMobile()` and skip WebGL addon loading
|
||||
- Reduces GPU memory pressure, prevents context loss crashes
|
||||
- Mobile typically has 1-2 sessions — canvas performance is more than adequate
|
||||
|
||||
**R4: Dynamic scrollback based on session activity** (Low risk)
|
||||
- Active sessions (working state): 5000 lines (current default)
|
||||
- Inactive/idle sessions: reduce to 2000 lines
|
||||
- Restore on session select (fetch from server buffer)
|
||||
- **Savings**: ~60% scrollback memory for idle sessions
|
||||
|
||||
**R5: Lazy-load Unicode11Addon** (Low risk)
|
||||
- Only load when CJK content is detected in terminal output
|
||||
- Detection: check for characters in CJK Unicode ranges during ANSI stripping (already iterating)
|
||||
- Most sessions never need it
|
||||
|
||||
---
|
||||
|
||||
## 3. DOM & Session Tab Rendering
|
||||
|
||||
### Current State
|
||||
- Session tabs use **intelligent incremental updates** with debounced 100ms rendering
|
||||
- Incremental path: only updates changed properties (classes, textContent, badges) when session list is stable
|
||||
- Full rebuild path: triggered when sessions added/removed **or badge count changes**
|
||||
- Subagent windows: per-window xterm.js instances, even when minimized
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B6: Badge count change triggers full tab rebuild** (`app.js:3207-3209`)
|
||||
- A single subagent badge increment on one tab triggers `_fullRenderSessionTabs()` — rebuilds entire sidebar HTML via `innerHTML =`
|
||||
- With 20 sessions, this is an O(n) reflow for a single badge number change
|
||||
- Badge changes are frequent during active subagent work
|
||||
|
||||
**B7: Minimized subagent windows retain xterm.js instances** (`subagent-windows.js`)
|
||||
- 50 subagent windows × ~75KB per xterm.js instance = ~3.75MB DOM memory
|
||||
- Minimized windows are invisible but their terminals remain in DOM
|
||||
- xterm.js instances continue processing resize events even when hidden
|
||||
|
||||
**B8: `backdrop-filter: blur()` on overlays** (`styles.css:2246-2247, 3098`)
|
||||
- Forces new stacking context, disables browser compositing optimizations
|
||||
- 50-100ms layout thrashing on modal open/close
|
||||
- Only 2 uses, but they're on frequently toggled overlays
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R6: Targeted badge update without full rebuild** (Low risk)
|
||||
- When badge count changes but session list is stable, update only the badge `<span>` textContent
|
||||
- Keep incremental path for badge changes; only use full rebuild for structural changes (add/remove sessions)
|
||||
- **Savings**: Eliminates O(n) reflow per badge change; reduces to O(1) targeted update
|
||||
|
||||
**R7: Lazy xterm.js initialization for subagent windows** (Medium impact)
|
||||
- Only create xterm.js Terminal instance when window is restored/maximized
|
||||
- On minimize: serialize terminal buffer, dispose Terminal instance, keep buffer in memory
|
||||
- On restore: create new Terminal, write buffer back
|
||||
- **Savings**: ~3.5MB DOM reduction at 50 minimized agents; eliminates hidden resize processing
|
||||
- **Trade-off**: ~200-500ms restore delay (buffer write), mitigated by chunked loading
|
||||
|
||||
**R8: Replace `backdrop-filter: blur()` with `background: rgba()`** (Low risk)
|
||||
- Use semi-transparent background instead of blur effect
|
||||
- Or use `will-change: transform` hint if blur is kept
|
||||
- **Savings**: Eliminates forced recomposition layer; 50-100ms faster overlay open
|
||||
|
||||
---
|
||||
|
||||
## 4. Backend PTY & State Management
|
||||
|
||||
### Current State (Excellent)
|
||||
- **BufferAccumulator**: Array-based chunking with lazy join on read — avoids O(n) string concatenation
|
||||
- **ANSI stripping**: Throttled at 150ms intervals with lazy evaluation (not per-chunk)
|
||||
- **State persistence**: 500ms debounce + incremental JSON caching per session (only dirty sessions re-serialized)
|
||||
- **Expensive parsers**: Throttled to 150ms window, accumulated data capped at 64KB
|
||||
- **Memory**: All buffers have hard limits (2MB terminal, 1MB text, 1000 messages, 64KB line buffer)
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B9: Pending clean data cap at 64KB** (`session.ts:1097-1133`)
|
||||
- Between 150ms processing windows, raw PTY data accumulates in `_pendingCleanData`
|
||||
- Capped at 64KB — excess data rolls off (old data discarded)
|
||||
- During heavy output (large build logs), this means parsers may miss content
|
||||
- Acceptable trade-off for performance, but worth documenting
|
||||
|
||||
**B10: `LRUMap.delete()` is O(n) worst case** (`utils/lru-map.ts:137-138`)
|
||||
- When deleting the newest entry, iterates all keys to find new newest
|
||||
- Rare in practice (delete is uncommon; set/get are hot paths)
|
||||
- Could matter during mass cleanup of 500 agents
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R9: Consider adaptive pending data cap** (Low priority)
|
||||
- During idle detection (critical to get right), increase cap to 128KB
|
||||
- During active working state, keep at 64KB (parsers less critical)
|
||||
- **Benefit**: More accurate idle detection during heavy output
|
||||
|
||||
**R10: Track second-newest in LRUMap** (Low priority)
|
||||
- Maintain a `_secondNewestKey` alongside `_newestKey`
|
||||
- On delete of newest, promote second-newest without iteration
|
||||
- Only matters at scale (500+ agents with frequent eviction)
|
||||
|
||||
---
|
||||
|
||||
## 5. Local Echo & Input Path
|
||||
|
||||
### Current State (Well-Designed)
|
||||
- **DOM overlay approach** — `<span>` elements in `.xterm-screen` at z-index 7, completely independent of `terminal.write()`
|
||||
- **Render caching**: `_lastRenderKey` includes text, position, column offsets — skips redundant re-renders
|
||||
- **Input flow**: Char accumulation → Enter triggers flush → 80ms delay before `\r` (ensures text reaches PTY first)
|
||||
- **Tab completion**: Baseline snapshot → detect buffer change → 300ms fallback timer
|
||||
- **CJK support**: Per-character width detection with `terminal.unicode.getStringCellWidth()` preferred, manual fallback
|
||||
- **Prompt detection**: Bottom-up line scan, O(rows) — cached position, column-lock prevents jitter
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B11: tmux send-keys latency** (~50-100ms per input)
|
||||
- Each `writeViaMux()` spawns a child process (`tmux send-keys`)
|
||||
- Text and Enter sent separately with 50ms delay between
|
||||
- For rapid typing: characters batch before Enter, so overhead is per-command not per-keystroke
|
||||
- **Acceptable trade-off** for session persistence (tmux survives server restarts)
|
||||
|
||||
**B12: 80ms delay between text flush and Enter** (`app.js:872-875`)
|
||||
- Intentional: ensures text reaches PTY before Enter, preventing Ink from processing empty input
|
||||
- Adds 80ms to perceived Enter-to-response latency
|
||||
- Could potentially be reduced with acknowledgment-based approach
|
||||
|
||||
**B13: Scroll listener on terminal viewport** (`zerolag-input-addon.ts:139`)
|
||||
- 50ms debounced re-render on scroll — acceptable but fires frequently during heavy output
|
||||
- Overlay hidden when scrolled up (correct behavior), shown when at bottom
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R11: Reduce Enter delay from 80ms to 50ms** (Low risk, test carefully)
|
||||
- The tmux `send-keys` already has 50ms internal delay
|
||||
- Combined with network latency, 80ms client-side may be excessive
|
||||
- Test with Ink-heavy sessions (Claude Code's status bar) — if text arrives before Enter at 50ms, reduce
|
||||
- **Savings**: 30ms perceived latency reduction per command
|
||||
|
||||
**R12: Batch tmux send-keys via stdin pipe** (Medium effort, high impact for rapid input)
|
||||
- Instead of spawning `tmux send-keys` per input, maintain a persistent connection
|
||||
- Use `tmux -C` (control mode) for programmatic interaction without child process spawning
|
||||
- **Savings**: Eliminate ~50-100ms process spawn overhead per input
|
||||
- **Risk**: Control mode has different semantics; needs careful testing with session persistence
|
||||
|
||||
**R13: Skip overlay re-render during heavy output scroll** (Low risk)
|
||||
- When terminal is receiving >10KB/s output, hide overlay entirely (user isn't typing during heavy output)
|
||||
- Re-show overlay after 500ms of output silence
|
||||
- **Savings**: Eliminates unnecessary DOM overlay re-renders during build logs / test output
|
||||
|
||||
---
|
||||
|
||||
## 6. Polling & File Watchers
|
||||
|
||||
### Current State
|
||||
- **SubagentWatcher**: 1s base poll, full scan throttled to every 5s, fs.watch() on known directories
|
||||
- **TranscriptWatcher**: 1 per session, fs.watch() primary with 1s poll fallback
|
||||
- **ImageWatcher**: chokidar per session with 100ms stability poll, burst limit 20/10s
|
||||
- **TeamWatcher**: chokidar primary with 30s poll fallback, LRU caches (50 teams, 200 tasks)
|
||||
- **RalphTracker**: Todo cleanup every 5 minutes
|
||||
|
||||
### Scaling Profile (20 sessions)
|
||||
| Component | Instances | Frequency | Total ops/sec |
|
||||
|-----------|-----------|-----------|---------------|
|
||||
| SubagentWatcher | 1 (global) | Full scan every 5s | 0.2/s |
|
||||
| TranscriptWatcher | 20 | 1s poll (fallback) | 20/s max |
|
||||
| ImageWatcher | 20 | 100ms poll (during writes only) | 200/s burst |
|
||||
| TeamWatcher | 1 (global) | 30s poll (fallback) | 0.03/s |
|
||||
| SSE heartbeat | 1 (global) | 15s | 0.07/s |
|
||||
| SSE dead client check | 1 (global) | 30s | 0.03/s |
|
||||
| Mux stats collection | 1 (global) | 2s | 0.5/s |
|
||||
| **Total steady-state** | | | **~21/s** |
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R14: Increase TranscriptWatcher poll interval to 2s** (Low risk)
|
||||
- Transcript changes are infrequent (new messages every few seconds at most)
|
||||
- fs.watch() is the primary mechanism; polling is fallback
|
||||
- **Savings**: Halves fallback filesystem checks (20/s → 10/s for 20 sessions)
|
||||
|
||||
**R15: Share chokidar instances for co-located session directories** (Medium effort)
|
||||
- Sessions in the same parent directory could share a single chokidar watcher with depth:3
|
||||
- Common case: multiple sessions in `~/projects/foo/` — one watcher covers all
|
||||
- **Savings**: Reduce chokidar instances from 20 to ~5-10 for typical workloads
|
||||
|
||||
---
|
||||
|
||||
## 7. Frontend Asset Delivery
|
||||
|
||||
### Current State
|
||||
- **app.js**: 12,027 lines (source) → esbuild minified → gzip/brotli compressed (~30-40KB gzipped)
|
||||
- **Static caching**: `maxAge: '1y'` via `@fastify/static`
|
||||
- **Service worker**: Push notification handler only — no asset caching
|
||||
- **No code splitting**: Single monolithic app.js bundle
|
||||
|
||||
### Bottlenecks
|
||||
|
||||
**B14: No cache-busting mechanism**
|
||||
- `maxAge: '1y'` means browsers cache aggressively
|
||||
- After deployment, users need `Ctrl+Shift+R` to see updates
|
||||
- No content hash in filenames or ETags for automatic invalidation
|
||||
|
||||
**B15: Monolithic app.js**
|
||||
- All 12K lines loaded on initial page load regardless of which features are used
|
||||
- Ralph wizard, plan orchestrator UI, team management — all loaded upfront
|
||||
- Mobile loads the same bundle as desktop
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R16: Add content hash to asset filenames** (Medium impact)
|
||||
- Build step: rename `app.js` → `app.[hash].js`
|
||||
- Generate a manifest or inject hash into HTML template
|
||||
- Keep `maxAge: '1y'` — cache invalidation happens via filename change
|
||||
- **Savings**: Eliminates stale cache issues after deployment; removes need for manual hard refresh
|
||||
|
||||
**R17: Code-split app.js into core + feature modules** (High effort, medium impact)
|
||||
- Core (~4K lines): terminal, SSE, session management, tabs, input handling
|
||||
- Deferred (~8K lines): Ralph wizard, plan UI, team management, subagent windows, image viewer
|
||||
- Load deferred modules on first use via dynamic `import()` or lazy `<script>` injection
|
||||
- **Savings**: ~60% reduction in initial load size; faster time-to-interactive
|
||||
- **Risk**: Complexity increase; need to handle loading states for deferred features
|
||||
- **Note**: May not be worth the effort given the app is already gzipped to ~30-40KB
|
||||
|
||||
---
|
||||
|
||||
## 8. CSS Performance
|
||||
|
||||
### Current State
|
||||
- **styles.css**: 7,153 lines with ~45 box-shadow uses, 2 backdrop-filter uses
|
||||
- Animations: GPU-accelerated keyframes for pulsing alerts, loading spinners
|
||||
- Z-index layering: well-organized (subagent 1000, plan 1100, log 2000, image 3000, overlay 7)
|
||||
|
||||
### Recommendations
|
||||
|
||||
**R18: Replace backdrop-filter with opaque overlay** (Low risk, covered in R8)
|
||||
|
||||
**R19: Use `contain: content` on subagent windows** (Low risk)
|
||||
- Add CSS containment to subagent window containers
|
||||
- Prevents layout changes inside windows from triggering reflow on parent
|
||||
- Especially valuable with 50 windows: changes in one window won't invalidate others
|
||||
- ```css
|
||||
.subagent-window { contain: content; }
|
||||
```
|
||||
- **Savings**: Reduces layout recalculation scope from global to per-window
|
||||
|
||||
**R20: Use `content-visibility: auto` on off-screen subagent windows** (Low risk)
|
||||
- Browser skips rendering of off-screen windows entirely
|
||||
- Combined with `contain-intrinsic-size` to prevent layout shift
|
||||
- ```css
|
||||
.subagent-window.minimized { content-visibility: hidden; }
|
||||
```
|
||||
- **Savings**: Browser skips paint/layout for minimized windows; complements R7
|
||||
|
||||
---
|
||||
|
||||
## 9. Memory & Scaling Limits
|
||||
|
||||
### Current Budget (20 sessions)
|
||||
| Component | Per Session | Total | Status |
|
||||
|-----------|-----------|-------|--------|
|
||||
| Terminal buffer | 2MB | 40MB | Hard-limited, auto-trim |
|
||||
| Text output | 1MB | 20MB | Hard-limited, auto-trim |
|
||||
| Messages | ~1MB | 20MB | Capped at 1000, trims to 800 |
|
||||
| Respawn buffer | 1MB | 20MB | Hard-limited |
|
||||
| **Buffers total** | | **100MB** | Acceptable |
|
||||
| TranscriptWatcher | ~100KB | 2MB | |
|
||||
| ImageWatcher | ~50KB | 1MB | |
|
||||
| SubagentWatcher | ~500KB | 500KB | Global |
|
||||
| Frontend terminal cache | ~256KB | 5MB | LRU, max 20 entries |
|
||||
| **Total estimated** | | **~110MB** | Comfortable |
|
||||
|
||||
### At Max Scale (50 sessions)
|
||||
- Buffers: ~250MB
|
||||
- Watchers: ~5MB
|
||||
- **Total: ~255MB** + Node.js overhead — acceptable on modern hardware
|
||||
|
||||
### Potential Leak Vectors (All Mitigated)
|
||||
- `_shortIdCache` in server — unbounded Map, but entries are tiny (string→string); grows at O(sessions created), not O(events)
|
||||
- All CleanupManager-registered resources tracked and disposed on session stop
|
||||
- `isStopped` guard prevents new timers after session cleanup
|
||||
|
||||
---
|
||||
|
||||
## 10. Implementation Priority Matrix
|
||||
|
||||
### Phase 1 — Quick Wins (1-2 hours each, low risk)
|
||||
| # | Optimization | Files to Change |
|
||||
|---|-------------|-----------------|
|
||||
| R6 | Targeted badge update | `app.js` (3207-3209) |
|
||||
| R3 | Canvas renderer on mobile | `app.js` (627-637) |
|
||||
| R8 | Replace backdrop-filter blur | `styles.css` (2246, 3098) |
|
||||
| R19 | CSS containment on subagent windows | `styles.css` |
|
||||
| R20 | `content-visibility: hidden` on minimized windows | `styles.css` |
|
||||
|
||||
### Phase 2 — Medium Effort (half-day each)
|
||||
| # | Optimization | Files to Change |
|
||||
|---|-------------|-----------------|
|
||||
| R2 | Tiered SSE padding | `server.ts` (broadcast function) |
|
||||
| R7 | Lazy xterm.js for minimized subagents | `subagent-windows.js` |
|
||||
| R11 | Reduce Enter delay to 50ms | `app.js` (872-875), test with Ink |
|
||||
| R14 | TranscriptWatcher 2s poll | `transcript-watcher.ts` |
|
||||
| R16 | Content-hash asset filenames | `build.mjs`, `server.ts` |
|
||||
|
||||
### Phase 3 — Larger Initiatives (1-2 days each)
|
||||
| # | Optimization | Files to Change |
|
||||
|---|-------------|-----------------|
|
||||
| R1 | Session-scoped SSE subscriptions | `server.ts`, `app.js` (SSE connect) |
|
||||
| R5 | Lazy Unicode11Addon loading | `app.js`, build pipeline |
|
||||
| R12 | Persistent tmux control mode | `tmux-manager.ts` |
|
||||
| R17 | Code-split app.js | `app.js`, `build.mjs`, HTML template |
|
||||
|
||||
### Not Recommended (Low ROI or High Risk)
|
||||
| # | Why Not |
|
||||
|---|---------|
|
||||
| R4 | Dynamic scrollback adds complexity; memory savings marginal vs total budget |
|
||||
| R9 | Adaptive pending data cap adds state; current 64KB cap rarely matters |
|
||||
| R10 | LRUMap.delete() O(n) is theoretical; never triggered at current scale |
|
||||
| R15 | Shared chokidar instances add directory-matching complexity for minimal gain |
|
||||
|
||||
---
|
||||
|
||||
## Appendix: Key File Locations
|
||||
|
||||
| Area | File | Key Lines |
|
||||
|------|------|-----------|
|
||||
| SSE broadcast | `src/web/server.ts` | 1961-1989 (broadcast), 1934-1959 (backpressure) |
|
||||
| Terminal batching | `src/web/server.ts` | 1994-2048 (per-session adaptive batching) |
|
||||
| Frame budget | `src/web/public/app.js` | 1370-1478 (flushPendingWrites, 64KB cap) |
|
||||
| Flicker filter | `src/web/public/app.js` | 1176-1255 (50ms sync wait, 256KB safety) |
|
||||
| Tab rendering | `src/web/public/app.js` | 3108-3357 (incremental + full rebuild) |
|
||||
| Tab switching | `src/web/public/app.js` | 3560-3760 (cache + chunked load + deferred UI) |
|
||||
| Local echo | `packages/xterm-zerolag-input/src/` | All files (overlay, prompt, CJK) |
|
||||
| Local echo integration | `src/web/public/app.js` | 640, 815-988 (input flow) |
|
||||
| Subagent windows | `src/web/public/subagent-windows.js` | Full file (window mgmt, drag, minimize) |
|
||||
| State persistence | `src/state-store.ts` | 161-250 (debounced save, incremental JSON) |
|
||||
| Buffer accumulator | `src/utils/buffer-accumulator.ts` | Full file (array chunks, lazy join) |
|
||||
| PTY handling | `src/session.ts` | 1046-1133 (data flow), 1173-1230 (parsing) |
|
||||
| Config limits | `src/config/` | 9 files (buffer, map, timing, auth, etc.) |
|
||||
| Anti-flicker docs | `docs/terminal-anti-flicker.md` | Architecture reference |
|
||||
| CSS | `src/web/public/styles.css` | 2246 (backdrop-filter), full file |
|
||||
| Build pipeline | `scripts/build.mjs` | 59-68 (minify + compress) |
|
||||
@@ -0,0 +1,266 @@
|
||||
# Codeman Performance Investigation Report
|
||||
|
||||
**Date**: 2026-02-20
|
||||
**Scope**: Why Codeman feels sluggish when multiple Claude tabs are very busy
|
||||
**Method**: 4-agent parallel analysis of server, PTY pipeline, frontend, and background systems
|
||||
|
||||
---
|
||||
|
||||
## Executive Summary
|
||||
|
||||
When multiple Claude sessions are actively producing heavy terminal output (e.g., building, writing files, running tests), Codeman's UI becomes sluggish. This investigation identified **14 bottlenecks** across 4 layers of the stack. The root cause is **cumulative event loop blocking** — no single operation is catastrophically slow, but dozens of small synchronous operations run on every PTY data chunk, and with N busy sessions producing chunks every few milliseconds, the event loop gets saturated.
|
||||
|
||||
The most impactful findings are ranked by severity below.
|
||||
|
||||
---
|
||||
|
||||
## Critical Findings (Event Loop Blockers)
|
||||
|
||||
### 1. PTY Data Handler Chain — O(output_volume) per session, synchronous
|
||||
**File**: `src/session.ts:986-1086`
|
||||
**Severity**: CRITICAL
|
||||
|
||||
Every chunk of PTY output from a busy Claude session runs through this synchronous chain on the Node.js event loop:
|
||||
|
||||
```
|
||||
PTY onData → ANSI strip regex → ralph-tracker → bash-tool-parser →
|
||||
token parser → CLI info parser → task description parser →
|
||||
idle/working detection → emit('terminal') → emit('output')
|
||||
```
|
||||
|
||||
**Key costs per chunk:**
|
||||
- `ANSI_ESCAPE_PATTERN_FULL` regex (line 999): Complex regex with alternation, runs on every chunk where any consumer needs clean data
|
||||
- `ralphTracker.processCleanData()` (line 1014): Splits into lines, runs regex per line, checks multi-line patterns
|
||||
- `bashToolParser.processCleanData()` (line 1020): Similar line-by-line regex processing
|
||||
- `parseTaskDescriptionsFromTerminalData()` (line 1038): Regex scan for parenthesized descriptions
|
||||
- Working/idle detection (lines 1043-1085): Multiple `includes()` checks plus `getCleanData()` calls
|
||||
|
||||
**The lazy `getCleanData()` pattern (line 997-1002)** was a good optimization — it avoids ANSI stripping when no consumer needs it. But when Ralph tracking is enabled (common during active work), `getCleanData()` is called on every chunk, negating the optimization.
|
||||
|
||||
**With 5 busy sessions** producing 50+ chunks/second each, this means 250+ synchronous processing chains per second on the event loop. Each chain involves string allocation, regex matching, and line splitting.
|
||||
|
||||
### 2. Broadcast Serialization — JSON.stringify on every flush
|
||||
**File**: `src/web/server.ts:4941-4967`
|
||||
**Severity**: CRITICAL
|
||||
|
||||
The `broadcast()` method calls `JSON.stringify(data)` synchronously for every event. Terminal data is the highest-frequency event. During `flushTerminalBatches()` (line 5030), broadcast is called once per session with pending data. With 10 busy sessions flushing every 16-50ms, that's 200-625 `JSON.stringify` calls per second on terminal data alone.
|
||||
|
||||
The terminal data payload is a string that gets double-encoded: the raw terminal string is embedded inside a JSON object `{id, data}`, then that object is JSON.stringify'd. For large chunks (up to 32KB per the `BATCH_FLUSH_THRESHOLD`), this creates significant garbage collection pressure.
|
||||
|
||||
**Additionally**, the `session:updated` broadcast includes `toLightDetailedState()` which serializes `taskTree`, `tokens`, `bufferStats`, and `respawnConfig` — this is called on many state changes, not just terminal data.
|
||||
|
||||
### 3. Single-Timer Batching — All sessions share one setTimeout
|
||||
**File**: `src/web/server.ts:5017-5027`
|
||||
**Severity**: HIGH
|
||||
|
||||
The `batchTerminalData()` method uses a **single shared timer** (`this.terminalBatchTimer`) for all sessions. When the timer fires, `flushTerminalBatches()` iterates ALL pending sessions and broadcasts each one. This means:
|
||||
|
||||
- One extremely busy session's rapid data can force the timer to fire at the minimum interval (16ms), flushing ALL sessions at that rate
|
||||
- The flush itself iterates all pending sessions synchronously
|
||||
- The `_minBatchInterval` optimization (line 5003) means the fastest session dictates the timer for everyone
|
||||
|
||||
This creates a **thundering herd** effect: all session flushes happen in a single synchronous burst rather than being staggered.
|
||||
|
||||
### 4. State Persistence Storms
|
||||
**File**: `src/web/server.ts:3879-3917`
|
||||
**Severity**: HIGH
|
||||
|
||||
`persistSessionState()` is called from **28+ locations** in server.ts. Each call sets a 100ms debounce timer per session. During heavy activity, this means:
|
||||
|
||||
- Frequent timer creation/cancellation (GC pressure)
|
||||
- The actual persist (`_persistSessionStateNow`) calls `session.toState()` which creates a new object, then `store.setSession()` which triggers `JSON.stringify` of the entire state store and `writeFileSync` to disk
|
||||
|
||||
The `StateStore` (via `state-store.ts`) debounces its own write, but the overhead is in the per-session `toState()` serialization and object creation, not just the disk write.
|
||||
|
||||
---
|
||||
|
||||
## High-Severity Findings
|
||||
|
||||
### 5. Ralph Tracker Line Processing — O(lines) per chunk
|
||||
**File**: `src/ralph-tracker.ts:1337-1375`
|
||||
**Severity**: HIGH (when Ralph tracking is enabled)
|
||||
|
||||
When enabled, `processCleanData()`:
|
||||
1. Appends to a line buffer (string concatenation)
|
||||
2. Splits on `\n` (creates array)
|
||||
3. Calls `processLine()` on each line (regex matching per line)
|
||||
4. Calls `checkMultiLinePatterns()` (additional regex on full chunk)
|
||||
5. Calls `maybeCleanupExpiredTodos()` (iterates todos Map)
|
||||
|
||||
For a busy session producing 100+ lines/second, this is significant. The line buffer can grow up to `MAX_LINE_BUFFER_SIZE` before being truncated, and the split/iterate pattern creates garbage on every chunk.
|
||||
|
||||
### 6. Subagent Watcher Polling — O(agents) every 1-10 seconds
|
||||
**File**: `src/subagent-watcher.ts:225-274`
|
||||
**Severity**: MEDIUM-HIGH
|
||||
|
||||
Three periodic operations:
|
||||
- **Poll interval** (1s): Lightweight check, but full directory scan every 5th poll (5s)
|
||||
- **Liveness check** (10s): Runs `pgrep` (child process spawn), then iterates ALL tracked agents to check if alive. With 50+ subagents (common with agent teams), this is a non-trivial burst.
|
||||
- **File watchers**: One `chokidar` watcher per tracked agent directory, plus transcript file watchers. With many agents, this means many active file watchers consuming kernel inotify resources.
|
||||
|
||||
The `getClaudePids()` call spawns a child process (`pgrep`) every 10 seconds. Under heavy load, child process spawning competes with the event loop.
|
||||
|
||||
### 7. SSE Client Iteration — O(clients) per broadcast
|
||||
**File**: `src/web/server.ts:4964-4966`
|
||||
**Severity**: MEDIUM
|
||||
|
||||
Every `broadcast()` iterates all SSE clients to send the pre-formatted message. With multiple browser tabs or mobile clients, each flush sends data to every client. The `reply.raw.write()` call goes through Node's HTTP stream, which is generally non-blocking but can cause backpressure cascades.
|
||||
|
||||
The backpressure handling (line 4916-4938) correctly skips backpressured clients, but the `once('drain')` handler sends a `session:needsRefresh` event, which the client responds to by fetching the full buffer — potentially a 2MB request — amplifying the problem.
|
||||
|
||||
### 8. Event Emitter Fan-Out in Session
|
||||
**File**: `src/session.ts:1008-1009`
|
||||
**Severity**: MEDIUM
|
||||
|
||||
Every PTY data chunk emits TWO events: `terminal` and `output`. The `terminal` event triggers `batchTerminalData()` in server.ts. The `output` event may trigger additional handlers. EventEmitter dispatch is synchronous — all listeners run before the next operation in the PTY handler continues.
|
||||
|
||||
With busy sessions, this means every chunk blocks the event loop for: PTY processing + all terminal listeners + all output listeners.
|
||||
|
||||
---
|
||||
|
||||
## Medium-Severity Findings
|
||||
|
||||
### 9. Respawn Controller Timer Accumulation
|
||||
**File**: `src/respawn-controller.ts` (various)
|
||||
**Severity**: MEDIUM
|
||||
|
||||
Each session with respawn enabled runs multiple timers:
|
||||
- Idle detection timeout
|
||||
- AI checker interval (when active)
|
||||
- Output silence detection interval
|
||||
- Token stability interval
|
||||
- Circuit breaker state timeouts
|
||||
|
||||
With 10 sessions with respawn, that's 50+ active timers. While individual timers are cheap, the cumulative effect on the event loop's timer queue is non-trivial — the libuv timer heap has O(log n) insertion but all callbacks run synchronously.
|
||||
|
||||
### 10. Team Watcher Polling
|
||||
**File**: `src/team-watcher.ts`
|
||||
**Severity**: MEDIUM (when agent teams are active)
|
||||
|
||||
Polls `~/.claude/teams/` directory every few seconds. Each poll reads config.json files and task files. With active teams, this adds filesystem reads to the event loop's I/O budget.
|
||||
|
||||
### 11. Frontend Terminal Write Batching
|
||||
**File**: `src/web/public/app.js` (batchTerminalWrite/flushPendingWrites)
|
||||
**Severity**: MEDIUM
|
||||
|
||||
The frontend batches terminal writes at `requestAnimationFrame` rate (16ms). When receiving SSE events from multiple busy sessions:
|
||||
- `batchTerminalWrite()` is called for EVERY session's data, even sessions not currently displayed
|
||||
- Terminal instances exist for all sessions (not just the active tab)
|
||||
- Each `flushPendingWrites()` calls `terminal.write()` which triggers xterm.js rendering
|
||||
|
||||
Hidden tabs still process terminal writes, consuming CPU for rendering that's never displayed.
|
||||
|
||||
### 12. Frontend Connection Line Rendering
|
||||
**File**: `src/web/public/app.js` (updateConnectionLines)
|
||||
**Severity**: LOW-MEDIUM
|
||||
|
||||
Connection lines between parent/child agent windows are recalculated on window moves, resizes, and potentially on terminal writes. With many subagent windows open, this involves DOM reads (getBoundingClientRect) that force layout recalculation.
|
||||
|
||||
### 13. Image Watcher File System Events
|
||||
**File**: `src/image-watcher.ts`
|
||||
**Severity**: LOW
|
||||
|
||||
Uses chokidar to watch for image files in session working directories. With many sessions in the same or overlapping directories, watchers may generate redundant events. The `awaitWriteFinish` and burst throttling mitigate this, but the kernel inotify resources add up.
|
||||
|
||||
### 14. ANSI Escape Regex Complexity
|
||||
**File**: `src/session.ts:999`
|
||||
**Severity**: LOW (but cumulative)
|
||||
|
||||
`ANSI_ESCAPE_PATTERN_FULL` is a complex regex with multiple alternation branches. While V8's regex engine handles this well for typical terminal data, adversarial input (deeply nested escape sequences) could cause superlinear matching time. The `FOCUS_ESCAPE_FILTER` regex runs first on every chunk.
|
||||
|
||||
---
|
||||
|
||||
## Scaling Analysis
|
||||
|
||||
| Resource | Per Session | 10 Sessions | 20 Sessions |
|
||||
|----------|-------------|-------------|-------------|
|
||||
| PTY data handlers | 1 synchronous chain | 10 chains competing for event loop | 20 chains — event loop saturation likely |
|
||||
| Broadcast calls (terminal only) | 20-60/sec | 200-600/sec | 400-1200/sec |
|
||||
| JSON.stringify (terminal) | 20-60/sec | 200-600/sec | 400-1200/sec |
|
||||
| Active timers | ~5 | ~50 | ~100 |
|
||||
| File watchers (subagents) | 2-5 | 20-50 | 40-100 |
|
||||
| SSE writes per flush | N clients | N clients x 10 sessions | N clients x 20 sessions |
|
||||
| Ralph line processing | O(lines/sec) | O(10 x lines/sec) | O(20 x lines/sec) |
|
||||
|
||||
**The critical threshold appears to be 5-8 simultaneously busy sessions**, where the cumulative PTY processing + broadcast serialization + timer callbacks start to exceed the event loop's capacity for responsive handling.
|
||||
|
||||
---
|
||||
|
||||
## Root Cause Architecture Diagram
|
||||
|
||||
```
|
||||
Busy Claude Session 1 ─┐
|
||||
Busy Claude Session 2 ─┤ ┌──────────────────────┐
|
||||
Busy Claude Session 3 ─┼───→│ Node.js Event Loop │
|
||||
Busy Claude Session 4 ─┤ │ (SINGLE THREAD) │
|
||||
Busy Claude Session 5 ─┘ │ │
|
||||
│ PTY handlers (sync) │◄── BOTTLENECK 1
|
||||
│ ANSI strip regex │
|
||||
│ Ralph tracker │
|
||||
│ Bash tool parser │
|
||||
│ Idle detection │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ EventEmitter.emit() │◄── BOTTLENECK 2
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ batchTerminalData() │
|
||||
│ (shared timer) │◄── BOTTLENECK 3
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ flushTerminalBatches() │
|
||||
│ broadcast() per session│
|
||||
│ JSON.stringify() each │◄── BOTTLENECK 4
|
||||
│ write() to N clients │
|
||||
│ │
|
||||
│ + persistSessionState │◄── BOTTLENECK 5
|
||||
│ + respawn timers │
|
||||
│ + subagent polling │
|
||||
│ + team watcher │
|
||||
└────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Recommendations (Not Implemented — For Discussion)
|
||||
|
||||
### Tier 1: Highest Impact, Lowest Risk
|
||||
1. **Disable processing for non-visible sessions**: Skip Ralph tracking, bash tool parsing, and task description parsing for sessions that no active SSE client is viewing. Only buffer terminal data.
|
||||
2. **Per-session flush staggering**: Instead of one shared timer flushing all sessions, use individual timers offset by `index * (interval/N)` to spread flushes across the batch window.
|
||||
3. **Skip hidden tab terminal writes on frontend**: Don't call `terminal.write()` for terminals not in the active tab. Lazy-load on tab switch.
|
||||
|
||||
### Tier 2: Medium Impact
|
||||
4. **Worker thread for ANSI stripping and parsing**: Move the regex-heavy ANSI strip + Ralph parsing to a worker thread pool. PTY data → worker → clean data back to main thread.
|
||||
5. **Pre-formatted SSE messages for terminal data**: Since terminal events are just `{id, data}`, build the SSE message string directly without `JSON.stringify`.
|
||||
6. **Adaptive processing based on load**: When event loop lag exceeds a threshold (measured via `setTimeout(0)` drift), reduce processing — skip Ralph, increase batch intervals, reduce subagent poll frequency.
|
||||
|
||||
### Tier 3: Longer-Term Architectural
|
||||
7. **Process-per-session or cluster mode**: Move each session's PTY handling to a separate Node.js worker or process, communicating to the main server via IPC.
|
||||
8. **Binary protocol for terminal data**: Replace JSON-encoded SSE terminal events with binary frames (e.g., MessagePack or raw binary WebSocket frames) to eliminate double-encoding.
|
||||
9. **Selective SSE subscriptions**: Clients subscribe to specific sessions instead of receiving all events. The server only broadcasts to interested clients.
|
||||
|
||||
---
|
||||
|
||||
## How to Validate
|
||||
|
||||
To confirm these findings, instrument with:
|
||||
```typescript
|
||||
// Add to event loop — measures how long synchronous work takes
|
||||
let lastCheck = Date.now();
|
||||
setInterval(() => {
|
||||
const now = Date.now();
|
||||
const lag = now - lastCheck - 100; // 100ms interval
|
||||
if (lag > 10) console.log(`[PERF] Event loop lag: ${lag}ms`);
|
||||
lastCheck = now;
|
||||
}, 100);
|
||||
```
|
||||
|
||||
And in `flushTerminalBatches()`:
|
||||
```typescript
|
||||
const start = performance.now();
|
||||
// ... existing flush logic ...
|
||||
const elapsed = performance.now() - start;
|
||||
if (elapsed > 5) console.log(`[PERF] Flush took ${elapsed.toFixed(1)}ms for ${this.terminalBatches.size} sessions`);
|
||||
```
|
||||
|
||||
This will show exactly when and how much the event loop is being blocked during heavy session activity.
|
||||
@@ -0,0 +1,168 @@
|
||||
# Performance & Responsiveness Optimization Plan
|
||||
|
||||
**Date**: 2026-02-28
|
||||
**Status**: Phases 1–4 Complete. Phase 5 optional/deferred.
|
||||
|
||||
---
|
||||
|
||||
## Executive Summary
|
||||
|
||||
Three independent research passes analyzed the Codeman codebase for performance bottlenecks across frontend rendering, backend hot paths, and system-level resource usage. The codebase already has strong foundational optimizations (per-session adaptive batching, rAF terminal writes, DEC 2026 sync markers, backpressure handling). This plan targets the remaining high-impact opportunities.
|
||||
|
||||
**Key finding**: The biggest wins come from **skipping unnecessary work** — serializing unchanged state, processing output nobody is watching, and reducing broadcast volume.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Quick Wins — COMPLETE
|
||||
|
||||
All Phase 1 items were found to already exist in the codebase during verification:
|
||||
|
||||
| # | Item | Status | Evidence |
|
||||
|---|------|--------|----------|
|
||||
| 1.1 | Skip terminal writes for hidden tabs | Done | SSE handler filters by `activeSessionId` (app.js:4076) |
|
||||
| 1.2 | mobile.css media query | Done | `media="(max-width: 1023px)"` on link tag (index.html:13) |
|
||||
| 1.3 | Deduplicate init API calls | Done | `_initGeneration` dedup + 3s fallback timer (app.js:2901-2904) |
|
||||
| 1.4 | Remove cache-busting timestamps | Done | No `?_t=` patterns found anywhere |
|
||||
| 1.5 | JS/CSS minification + compression | Done | esbuild minify + gzip + brotli in build.mjs (lines 42-51) |
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Frontend Responsiveness — COMPLETE
|
||||
|
||||
### 2.1 Batch `getBoundingClientRect()` in connection lines — DONE
|
||||
- **Files**: `src/web/public/app.js` (`_updateConnectionLinesImmediate()`)
|
||||
- **Change**: Refactored to batch all layout reads into Phase 1 (collect all rects into a Map), then perform all SVG writes in Phase 2 using cached values. Classic read-then-write pattern prevents interleaved forced reflows.
|
||||
|
||||
### 2.2 Clean up ResizeObservers — Already implemented
|
||||
- `forceCloseSubagentWindow()` disconnects observers (app.js:12618-12620)
|
||||
- `cleanupAllFloatingWindows()` disconnects all on reconnect (app.js:12649-12653)
|
||||
- Observer refs stored on `windowData.resizeObserver` (app.js:12492)
|
||||
|
||||
### 2.3 Drag handler cleanup — Already implemented
|
||||
- `makeWindowDraggable()` returns listener refs, stored in `windowData.dragListeners`
|
||||
- `forceCloseSubagentWindow()` removes all document-level drag listeners (app.js:12622-12630)
|
||||
- Panel drags add listeners on mousedown, remove on mouseup (app.js:10253-10284)
|
||||
|
||||
### 2.4 Mobile window position cache — Skipped
|
||||
- O(n) loop over max ~20 windows; complexity of cached counter not justified
|
||||
|
||||
### 2.5 Lazy modal DOM — Skipped
|
||||
- Large effort, marginal benefit for a vanilla JS app with fast DOM construction
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Backend Hot Paths — COMPLETE
|
||||
|
||||
### 3.1 State diff broadcasts — ALREADY OPTIMIZED
|
||||
- `broadcastSessionStateDebounced()` already batches at 500ms intervals
|
||||
- `toLightDetailedState()` excludes heavy buffers (textOutput, terminalBuffer)
|
||||
- Per-session serialization is <1ms; with debouncing, only 1-3 sessions serialize per flush
|
||||
- JSON.stringify happens once per broadcast (not per client) — serialization cost is negligible
|
||||
- Full state diffs would add significant frontend complexity for marginal gain
|
||||
|
||||
### 3.2 Improve session list cache hit rate — DONE
|
||||
- **Files**: `src/web/server.ts` (`broadcast()` method)
|
||||
- **Change**: Cache now only invalidated on truly structural events (`session:created`, `session:deleted`, `session:updated`) instead of on every `session:*` and `respawn:*` event. High-frequency events like `session:working`, `session:idle`, `session:completion`, `respawn:stateChanged` no longer defeat the 1s TTL cache.
|
||||
- **Impact**: Cache hit ratio from ~0% to ~80%+ during active sessions. The debounced `session:updated` still refreshes the cache within 500ms of any state change.
|
||||
|
||||
### 3.3 Skip PTY processing — ALREADY OPTIMIZED
|
||||
- `_processExpensiveParsers()` is already throttled to every 150ms (not per-chunk)
|
||||
- Lazy ANSI stripping via `getCleanData()` closure — only computed when a consumer needs it
|
||||
- Quick pre-checks skip parsers when content is irrelevant (e.g., token parser only runs if data contains "token")
|
||||
- OpenCode sessions skip all Claude-specific parsers entirely
|
||||
- Further optimization would require visibility-aware processing, adding complexity for marginal gain
|
||||
|
||||
### 3.4 Batch subagent liveness checks — Deferred
|
||||
- `/proc/{pid}` stat calls are ~0.1ms each; even with 500 agents, total is 50ms every 10s
|
||||
- Current approach is simple and reliable; batching adds race condition risk
|
||||
- Consider only if profiling shows this as a bottleneck
|
||||
|
||||
### 3.5 Deduplicate detection update emissions — DONE
|
||||
- **Files**: `src/respawn-controller.ts` (`startDetectionUpdates()`)
|
||||
- **Change**: Detection status now only emitted when key fields (confidenceLevel, statusText, controller state) actually change. Previously emitted every 2s regardless, broadcasting identical status to all SSE clients.
|
||||
- **Impact**: For stable/idle sessions, eliminates ~100% of redundant detection broadcasts. For active sessions, reduces broadcasts to only meaningful state transitions.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4: System-Level Improvements — COMPLETE
|
||||
|
||||
### 4.1 Incremental state persistence — DONE
|
||||
- **Files**: `src/state-store.ts` (`assembleStateJson()`, `setSession()`)
|
||||
- **Change**: Added `dirtySessions` Set and `cachedSessionJsons` Map. On persist, only dirty sessions are re-serialized; clean sessions reuse cached JSON fragments. `setSession()` marks sessions dirty; `assembleStateJson()` rebuilds only changed fragments.
|
||||
- **Impact**: Serialization cost reduced from O(all sessions) to O(dirty sessions). Typical steady-state: 1-2 dirty sessions instead of 50.
|
||||
|
||||
### 4.2 Replace polling with fs watchers for team watcher — DONE
|
||||
- **Files**: `src/team-watcher.ts` (`setupFsWatchers()`)
|
||||
- **Change**: Added chokidar watchers on both `~/.claude/teams/` and `~/.claude/tasks/` directories for instant event-driven detection. Lock files ignored via chokidar config. Mtime-based dedup skips unchanged files. Polling interval relaxed from 5s to 30s as a fallback.
|
||||
- **Impact**: Near-instant team detection; polling overhead eliminated for normal operation.
|
||||
|
||||
### 4.3 Consolidate subagent file watchers — DONE
|
||||
- **Files**: `src/subagent-watcher.ts` (`setupDirectoryWatcher()`)
|
||||
- **Change**: Replaced per-agent chokidar watchers with one `fs.watch()` per session subagent directory. Events are routed to the correct agent via filename. Per-file debouncing (100ms) prevents hammering on bulk discovery.
|
||||
- **Impact**: Inotify watchers reduced from potentially 500 (one per agent) to ~50 (one per session directory).
|
||||
|
||||
### 4.4 Stream transcript files instead of full reads — DONE
|
||||
- **Files**: `src/subagent-watcher.ts` (`tailFile()`, `findDescriptionInAgentFile()`, parent transcript lookup)
|
||||
- **Change**: Multiple streaming strategies implemented:
|
||||
- **Live monitoring**: Position-based `tailFile()` with `createReadStream({ start: fromPosition })` — only reads new content
|
||||
- **Parent transcript lookup**: Streams only last 16KB (`createReadStream({ start: offset })`)
|
||||
- **Description extraction**: Streams only first 8KB, exits early after 5 lines
|
||||
- **Full read**: Only for on-demand transcript review panel (with optional `limit` parameter)
|
||||
- **Impact**: File I/O for bulk agent discovery reduced from ~50MB to ~5MB.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5: Long-Term Architectural (Optional) — NOT STARTED
|
||||
|
||||
These items are deferred until scaling demands justify the complexity.
|
||||
|
||||
### 5.1 Worker thread for PTY processing
|
||||
- **Files**: `src/session.ts`
|
||||
- **Problem**: ANSI stripping, Ralph tracking, and bash tool parsing all run on the main event loop. At scale (50 busy sessions), this consumes 300-500ms CPU/sec.
|
||||
- **Fix**: Offload ANSI strip + line processing to a worker thread pool. Main thread receives clean text + parsed events.
|
||||
- **Impact**: Frees event loop for I/O operations. Most impactful at 10+ concurrent busy sessions.
|
||||
|
||||
### 5.2 Per-session SSE subscriptions
|
||||
- **Files**: `src/web/server.ts`
|
||||
- **Problem**: Every SSE event is broadcast to all connected clients. A client watching session A still receives events for sessions B through Z.
|
||||
- **Fix**: Clients subscribe to specific session IDs. Server only sends events to interested clients.
|
||||
- **Impact**: Reduces SSE broadcast fan-out from N clients to ~1-2 per event. Major improvement at 100 SSE clients.
|
||||
|
||||
### 5.3 O(1) LRUMap via doubly-linked list
|
||||
- **Files**: `src/utils/lru-map.ts` (~lines 98-110)
|
||||
- **Problem**: `get()` uses delete + re-insert to refresh position — O(n) on Map iteration for delete.
|
||||
- **Fix**: Implement classic LRU with doubly-linked list + Map for O(1) get/put/evict.
|
||||
- **Impact**: Low — current sizes (max 500) make this barely measurable. Only worthwhile if LRUMap is used on hot paths.
|
||||
|
||||
---
|
||||
|
||||
## Completion Summary
|
||||
|
||||
| Phase | Scope | Status | Items |
|
||||
|-------|-------|--------|-------|
|
||||
| 1 | Quick Wins | **Complete** | 5/5 (all pre-existing) |
|
||||
| 2 | Frontend Responsiveness | **Complete** | 3/3 actionable done, 2 skipped |
|
||||
| 3 | Backend Hot Paths | **Complete** | 4/4 actionable done, 1 deferred |
|
||||
| 4 | System-Level | **Complete** | 4/4 done |
|
||||
| 5 | Long-Term Architectural | **Not started** | 0/3 — deferred until needed |
|
||||
|
||||
**Overall**: 16/16 actionable items complete. 3 optional items deferred.
|
||||
|
||||
---
|
||||
|
||||
## Measurement
|
||||
|
||||
Before starting Phase 5, establish baselines:
|
||||
|
||||
1. **Frontend**: Record Chrome DevTools Performance trace with 10 sessions open. Measure:
|
||||
- Frame rate during rapid terminal output
|
||||
- Long tasks (>50ms) count per 30s
|
||||
- Heap size after 1h session
|
||||
|
||||
2. **Backend**: Add `performance.now()` instrumentation around:
|
||||
- `flushSessionTerminalBatch()` — time per flush
|
||||
- `broadcastSessionStateDebounced()` — serialization time
|
||||
- `StateStore.save()` — persist time
|
||||
- Event loop lag via `monitorEventLoopDelay()`
|
||||
|
||||
3. **First load**: Lighthouse score on desktop and mobile (simulated 3G)
|
||||
@@ -0,0 +1,74 @@
|
||||
# Codeman Performance Optimization Plan
|
||||
|
||||
## Current State
|
||||
|
||||
The backend is **already production-grade** — SSE broadcasting, state persistence, terminal batching, buffer management, and memory patterns are all well-optimized. The biggest gains are on the **frontend delivery** side.
|
||||
|
||||
## Implemented Optimizations
|
||||
|
||||
### 1. V8 Compile Cache (10-20% faster cold start)
|
||||
|
||||
**Files:** `scripts/codeman-web.service`, `package.json`
|
||||
|
||||
Node.js re-parses and compiles all JS on every cold start. `NODE_COMPILE_CACHE` caches V8 compiled bytecode to disk, reusing it on subsequent starts.
|
||||
|
||||
- Added `Environment=NODE_COMPILE_CACHE=/home/arkon/.codeman/compile-cache` to systemd service
|
||||
- Added to `npm start` script for non-systemd usage
|
||||
- Zero code changes, immediate win on every restart
|
||||
|
||||
### 2. WebGL Addon Lazy-Loading (244KB saved on mobile, non-blocking on desktop)
|
||||
|
||||
**Files:** `src/web/public/index.html`, `src/web/public/app.js`
|
||||
|
||||
`xterm-addon-webgl.min.js` (244KB) was loaded eagerly for all users via `<script defer>`, but only used on desktop with WebGL2 support.
|
||||
|
||||
- Removed `<script defer>` from `index.html`
|
||||
- Added dynamic script loading in `app.js` — only downloads on desktop when WebGL is needed
|
||||
- Mobile users never download the file at all (244KB saved)
|
||||
- Desktop: loads in parallel with page rendering, addon initializes when ready
|
||||
- Graceful fallback: canvas renderer used if WebGL unavailable or script fails
|
||||
|
||||
### 3. Preload Hints (~50-100ms faster perceived load)
|
||||
|
||||
**Files:** `src/web/public/index.html`
|
||||
|
||||
Browser discovers `<script defer>` tags only when the parser reaches them at the bottom of `<body>`. By then, the HTML parse has blocked for hundreds of lines.
|
||||
|
||||
- Added `<link rel="preload" as="script">` in `<head>` for `vendor/xterm.min.js`, `constants.js`, `app.js`
|
||||
- Browser starts fetching critical scripts immediately during HTML parse (before reaching `<body>`)
|
||||
- Zero runtime overhead — just hints for the browser's preload scanner
|
||||
|
||||
### 4. Batch Tmux Reconciliation (N subprocess calls → 1)
|
||||
|
||||
**Files:** `src/tmux-manager.ts`
|
||||
|
||||
`reconcileSessions()` previously called `tmux has-session` + `tmux display-message` per known session, plus `tmux list-sessions` for discovery, plus `tmux display-message` per discovered session. With 20 sessions: 41+ subprocess calls.
|
||||
|
||||
- Replaced with single `tmux list-panes -a -F '#{session_name}\t#{pane_pid}'` call
|
||||
- Builds a Map from the result, then does O(1) lookups for both known and discovered sessions
|
||||
- Also replaced inner O(n) `isKnown` scan with a Set lookup
|
||||
- 20 sessions: 41 subprocess calls → 1, with faster lookups
|
||||
|
||||
### 5. Asset Hashing / Cache Busting (already implemented)
|
||||
|
||||
**Files:** `scripts/build.mjs` (pre-existing)
|
||||
|
||||
Content-hash cache busting was already implemented in the build script:
|
||||
- All app JS/CSS files get content hashes (`app.abc123.js`)
|
||||
- `index.html` rewritten to reference hashed filenames
|
||||
- Pre-compressed with gzip + Brotli
|
||||
- 1-year immutable cache works correctly — new deploys get new filenames
|
||||
|
||||
## Already Optimized (No Action Needed)
|
||||
|
||||
| Area | Why It's Fine |
|
||||
|------|---------------|
|
||||
| **SSE Broadcasting** | Single serialization per broadcast, preformatted frames, backpressure handling, session subscription filtering |
|
||||
| **State Persistence** | 500ms debounce, incremental per-session JSON caching, async atomic writes, circuit breaker on failures |
|
||||
| **Terminal Batching** | Adaptive intervals (16-50ms), per-session queues, immediate flush at 32KB, array-based accumulation |
|
||||
| **Buffer Management** | BufferAccumulator (array-push, lazy join), auto-trim at 2MB/1MB, no string concatenation in hot paths |
|
||||
| **ANSI Stripping** | Pre-compiled regex via factory functions, single-pass processing |
|
||||
| **Static File Serving** | @fastify/static with 1-year cache, pre-compressed Brotli/gzip, no-cache for HTML |
|
||||
| **Memory Management** | CleanupManager, LRUMap, StaleExpirationMap, bounded buffers, explicit listener cleanup |
|
||||
| **Import Patterns** | Pure ESM, lazy web server import, no circular deps, no dynamic imports in hot paths |
|
||||
| **Config Loading** | Small constant files, no I/O at import time, specific imports (no barrel) |
|
||||
@@ -0,0 +1,788 @@
|
||||
# Phase 4: Domain File Splitting — Implementation Plan
|
||||
|
||||
**Date**: 2026-03-01
|
||||
**Prerequisites**: Phase 1-3 complete (utils cleanup, CleanupManager/Debouncer migration, route extraction)
|
||||
**Goal**: Split 4 god files into focused modules with barrel exports for transparent migration.
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Split types.ts into types/ directory](#1-split-typests-into-types-directory)
|
||||
2. [Split ralph-tracker.ts into focused modules](#2-split-ralph-trackerts-into-focused-modules)
|
||||
3. [Split respawn-controller.ts into focused modules](#3-split-respawn-controllerts-into-focused-modules)
|
||||
4. [Split session.ts into focused modules](#4-split-sessionts-into-focused-modules)
|
||||
5. [Execution Order & Dependencies](#5-execution-order--dependencies)
|
||||
6. [Validation Checklist](#6-validation-checklist)
|
||||
|
||||
---
|
||||
|
||||
## 1. Split types.ts into types/ directory
|
||||
|
||||
**Current**: 1,443 lines, 71 exports, imported by 36 files.
|
||||
**Risk**: LOW — pure type refactor, no runtime behavior change.
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/types/
|
||||
├── index.ts (barrel re-export — transparent migration)
|
||||
├── common.ts (Disposable, BufferConfig, CleanupResourceType, CleanupRegistration)
|
||||
├── session.ts (SessionStatus, SessionMode, ClaudeMode, SessionConfig, SessionColor,
|
||||
│ SessionState, OpenCodeConfig, SessionOutput)
|
||||
├── task.ts (TaskStatus, TaskDefinition, TaskState)
|
||||
├── app-state.ts (AppState, AppConfig, GlobalStats, TokenUsageEntry, TokenStats,
|
||||
│ DEFAULT_CONFIG, createInitialState, createInitialGlobalStats)
|
||||
├── respawn.ts (RespawnConfig, PersistedRespawnConfig, CycleOutcome,
|
||||
│ RespawnCycleMetrics, RespawnAggregateMetrics, HealthStatus,
|
||||
│ RalphLoopHealthScore, TimingHistory, RespawnPreset)
|
||||
├── ralph.ts (RalphLoopStatus, RalphLoopState, RalphTodoStatus, RalphTodoPriority,
|
||||
│ RalphTodoItem, RalphTodoProgress, RalphSessionState,
|
||||
│ RalphStatusValue, RalphTestsStatus, RalphWorkType, RalphStatusBlock,
|
||||
│ CompletionConfidence, RalphTrackerState,
|
||||
│ CircuitBreakerState, CircuitBreakerReason, CircuitBreakerStatus,
|
||||
│ createInitialCircuitBreakerStatus, createInitialRalphTrackerState,
|
||||
│ createInitialRalphSessionState)
|
||||
├── api.ts (ApiErrorCode, ApiResponse, HookEventType, QuickStartResponse,
|
||||
│ CaseInfo, createErrorResponse, isError, getErrorMessage)
|
||||
├── lifecycle.ts (LifecycleEventType, LifecycleEntry)
|
||||
├── run-summary.ts (RunSummaryEventType, RunSummaryEventSeverity, RunSummaryEvent,
|
||||
│ RunSummaryStats, RunSummary, createInitialRunSummaryStats)
|
||||
├── tools.ts (ActiveBashToolStatus, ActiveBashTool, ImageDetectedEvent)
|
||||
├── teams.ts (TeamConfig, TeamMember, TeamTask, InboxMessage, PaneInfo)
|
||||
├── push.ts (PushSubscriptionRecord, VapidKeys)
|
||||
└── plan.ts (PlanTaskStatus, TddPhase, PlanItem re-export, NiceConfig,
|
||||
DEFAULT_NICE_CONFIG, ProcessStats)
|
||||
```
|
||||
|
||||
### Steps
|
||||
|
||||
1. **Create `src/types/` directory** and each domain file above.
|
||||
|
||||
2. **Move types** from `src/types.ts` into their domain files. Preserve all JSDoc comments. Each file should import from siblings as needed (e.g., `ralph.ts` imports `CircuitBreakerState` within itself — no cross-file deps needed since they're in the same file).
|
||||
|
||||
3. **Create barrel `src/types/index.ts`** that re-exports everything:
|
||||
```typescript
|
||||
export * from './common.js';
|
||||
export * from './session.js';
|
||||
export * from './task.js';
|
||||
export * from './app-state.js';
|
||||
export * from './respawn.js';
|
||||
export * from './ralph.js';
|
||||
export * from './api.js';
|
||||
export * from './lifecycle.js';
|
||||
export * from './run-summary.js';
|
||||
export * from './tools.js';
|
||||
export * from './teams.js';
|
||||
export * from './push.js';
|
||||
export * from './plan.js';
|
||||
```
|
||||
|
||||
4. **Delete old `src/types.ts`** and replace with a single-line re-export barrel:
|
||||
```typescript
|
||||
export * from './types/index.js';
|
||||
```
|
||||
This ensures `import { ... } from './types.js'` continues to work everywhere — zero changes to 36 import sites.
|
||||
|
||||
5. **Verify**: `tsc --noEmit` and `npm run lint` must pass. No runtime changes.
|
||||
|
||||
### Internal Dependencies Between Domain Files
|
||||
|
||||
Some types reference others across domains. Handle with imports:
|
||||
|
||||
| File | Imports From |
|
||||
|------|-------------|
|
||||
| `app-state.ts` | `session.ts` (SessionState), `task.ts` (TaskState), `ralph.ts` (RalphLoopState, RalphSessionState) |
|
||||
| `respawn.ts` | None (self-contained) |
|
||||
| `ralph.ts` | None (self-contained) |
|
||||
| `run-summary.ts` | None (self-contained) |
|
||||
| `api.ts` | None (self-contained) |
|
||||
| `session.ts` | `respawn.ts` (RespawnConfig), `ralph.ts` (RalphTrackerState, RalphTodoItem, CircuitBreakerStatus, RalphSessionState, RunSummaryEvent) |
|
||||
|
||||
Wait — `SessionState` references `RespawnConfig`, `RalphTrackerState`, `CircuitBreakerStatus`, and `RunSummaryEvent`. This creates imports from `session.ts` → `respawn.ts`, `ralph.ts`, `run-summary.ts`. This is fine (one-way deps, no cycles).
|
||||
|
||||
---
|
||||
|
||||
## 2. Split ralph-tracker.ts into focused modules
|
||||
|
||||
**Current**: 3,868 lines, single `RalphTracker` class with 5 responsibilities.
|
||||
**Risk**: MEDIUM — class has shared mutable state, but extractable modules are well-isolated.
|
||||
|
||||
### Coupling Analysis Summary
|
||||
|
||||
| Module | Coupling | Extractability |
|
||||
|--------|----------|----------------|
|
||||
| Plan task tracking | LOW | HIGH — only reads `cycleCount` |
|
||||
| Fix-plan file watching | LOW | HIGH — callback-based todo replacement |
|
||||
| Iteration stall detection | LOW | HIGH — notification-based |
|
||||
| RALPH_STATUS block parsing + circuit breaker | MEDIUM | MEDIUM — callback for circuit breaker updates |
|
||||
| Todo parsing, loop detection, completion | HIGH | LOW — deeply entangled shared state |
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/
|
||||
├── ralph-tracker.ts (~1,800 LOC — core: output parsing, loop state,
|
||||
│ todo management, completion detection)
|
||||
├── ralph-plan-tracker.ts (~600 LOC — plan tasks, checkpoints, history, rollback)
|
||||
├── ralph-status-parser.ts (~300 LOC — RALPH_STATUS block parsing, circuit breaker)
|
||||
├── ralph-fix-plan-watcher.ts (~150 LOC — @fix_plan.md file watching)
|
||||
└── ralph-stall-detector.ts (~80 LOC — iteration stall detection)
|
||||
```
|
||||
|
||||
### Step 2a: Extract `RalphPlanTracker` (~600 LOC)
|
||||
|
||||
**Why first**: Lowest coupling. Only dependency is `cycleCount` for checkpoint detection.
|
||||
|
||||
**Extract these from `RalphTracker`**:
|
||||
|
||||
Types to export:
|
||||
- `EnhancedPlanTask` (interface, currently lines 56-87)
|
||||
- `CheckpointReview` (interface, currently lines 90-139)
|
||||
|
||||
Properties to move:
|
||||
- `_planVersion: number`
|
||||
- `_planHistory: Array<{version, timestamp, tasks, summary}>`
|
||||
- `_planTasks: Map<string, EnhancedPlanTask>`
|
||||
- `_checkpointIterations: number[]`
|
||||
- `_lastCheckpointIteration: number`
|
||||
|
||||
Methods to move:
|
||||
- `initializePlanTasks(items)`
|
||||
- `updatePlanTask(taskId, update)`
|
||||
- `addPlanTask(params)`
|
||||
- `getPlanTasks()`
|
||||
- `generateCheckpointReview()`
|
||||
- `getPlanHistory()`
|
||||
- `rollbackToVersion(version)`
|
||||
- `isCheckpointDue()`
|
||||
- `planVersion` getter
|
||||
- `_savePlanToHistory()` (private)
|
||||
- `_unblockDependentTasks()` (private)
|
||||
- `_checkForCheckpoint()` (private)
|
||||
|
||||
Events emitted (define in new class):
|
||||
- `planInitialized`
|
||||
- `planTaskUpdate`
|
||||
- `taskBlocked`
|
||||
- `taskUnblocked`
|
||||
- `planCheckpoint`
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphPlanTracker extends EventEmitter {
|
||||
constructor() { ... }
|
||||
|
||||
// Parent calls this when iteration changes (for checkpoint detection)
|
||||
notifyCycleCount(cycleCount: number): void { ... }
|
||||
|
||||
// Full public API moves here unchanged
|
||||
initializePlanTasks(items: PlanItem[]): void { ... }
|
||||
updatePlanTask(taskId: string, update: { ... }): { ... } | null { ... }
|
||||
// ...etc
|
||||
}
|
||||
```
|
||||
|
||||
**In `RalphTracker`**: Replace plan methods with delegation:
|
||||
```typescript
|
||||
readonly planTracker = new RalphPlanTracker();
|
||||
|
||||
// Forward plan events
|
||||
this.planTracker.on('planInitialized', (...args) => this.emit('planInitialized', ...args));
|
||||
// ...etc
|
||||
|
||||
// In detectLoopStatus(), when cycleCount changes:
|
||||
this.planTracker.notifyCycleCount(this._loopState.cycleCount);
|
||||
```
|
||||
|
||||
### Step 2b: Extract `RalphFixPlanWatcher` (~150 LOC)
|
||||
|
||||
**Extract these**:
|
||||
|
||||
Properties:
|
||||
- `_workingDir: string | null`
|
||||
- `_fixPlanPath: string | null`
|
||||
- `_fixPlanWatcher: FSWatcher | null`
|
||||
- `_fixPlanWatcherErrorHandler`
|
||||
- `_fixPlanReloadDeb`
|
||||
|
||||
Methods:
|
||||
- `setWorkingDir(workingDir)`
|
||||
- `loadFixPlanFromDisk()`
|
||||
- `startWatchingFixPlan()`
|
||||
- `stopWatchingFixPlan()`
|
||||
- `handleFixPlanChange()`
|
||||
- `isFileAuthoritative` getter
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphFixPlanWatcher extends EventEmitter {
|
||||
get isFileAuthoritative(): boolean { ... }
|
||||
|
||||
setWorkingDir(workingDir: string): void { ... }
|
||||
stop(): void { ... }
|
||||
}
|
||||
|
||||
// Events:
|
||||
// 'todosLoaded' → (todos: Array<{id, content, status, priority}>) — parent replaces _todos
|
||||
```
|
||||
|
||||
**In `RalphTracker`**:
|
||||
```typescript
|
||||
readonly fixPlanWatcher = new RalphFixPlanWatcher();
|
||||
|
||||
constructor() {
|
||||
this.fixPlanWatcher.on('todosLoaded', (items) => {
|
||||
// Replace _todos with file-based items
|
||||
this._todos.clear();
|
||||
for (const item of items) {
|
||||
this.addOrUpdateTodo(item.id, item.content, item.status, item.priority);
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
// Delegate isFileAuthoritative
|
||||
get isFileAuthoritative(): boolean {
|
||||
return this.fixPlanWatcher.isFileAuthoritative;
|
||||
}
|
||||
```
|
||||
|
||||
### Step 2c: Extract `RalphStallDetector` (~80 LOC)
|
||||
|
||||
**Extract these**:
|
||||
|
||||
Properties:
|
||||
- `_lastIterationChangeTime`
|
||||
- `_lastObservedIteration`
|
||||
- `_iterationStallTimerId`
|
||||
- `_iterationStallWarningMs`
|
||||
- `_iterationStallCriticalMs`
|
||||
- `_iterationStallWarned`
|
||||
|
||||
Methods:
|
||||
- `startIterationStallDetection()`
|
||||
- `stopIterationStallDetection()`
|
||||
- `checkIterationStall()`
|
||||
- `getIterationStallMetrics()`
|
||||
- `configureIterationStallThresholds(warningMs, criticalMs)`
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphStallDetector extends EventEmitter {
|
||||
constructor(private cleanup: CleanupManager) { ... }
|
||||
|
||||
start(): void { ... }
|
||||
stop(): void { ... }
|
||||
|
||||
// Parent calls when iteration changes
|
||||
notifyIterationChanged(iteration: number): void {
|
||||
this._lastIterationChangeTime = Date.now();
|
||||
this._lastObservedIteration = iteration;
|
||||
this._iterationStallWarned = false;
|
||||
}
|
||||
|
||||
// Parent calls to check if loop is active
|
||||
setLoopActive(active: boolean): void { ... }
|
||||
|
||||
getIterationStallMetrics(): { ... } { ... }
|
||||
}
|
||||
|
||||
// Events: 'iterationStallWarning', 'iterationStallCritical'
|
||||
```
|
||||
|
||||
### Step 2d: Extract `RalphStatusParser` (~300 LOC)
|
||||
|
||||
**Extract these**:
|
||||
|
||||
Properties:
|
||||
- `_circuitBreaker: CircuitBreakerStatus`
|
||||
- `_statusBlockBuffer: string[]`
|
||||
- `_inStatusBlock: boolean`
|
||||
- `_lastStatusBlock: RalphStatusBlock | null`
|
||||
- `_completionIndicators: number`
|
||||
- `_exitGateMet: boolean`
|
||||
- `_totalFilesModified: number`
|
||||
- `_totalTasksCompleted: number`
|
||||
|
||||
Methods:
|
||||
- `processStatusBlockLine(line)`
|
||||
- `parseStatusBlock(lines)`
|
||||
- `detectCompletionIndicators(line)`
|
||||
- `updateCircuitBreaker(hasProgress, testsStatus, status)`
|
||||
- `resetCircuitBreaker()`
|
||||
- `circuitBreakerStatus` getter
|
||||
- `lastStatusBlock` getter
|
||||
- `cumulativeStats` getter
|
||||
- `exitGateMet` getter
|
||||
|
||||
Regex patterns to move:
|
||||
- `RALPH_STATUS_START_PATTERN` through `RALPH_RECOMMENDATION_PATTERN`
|
||||
- `COMPLETION_INDICATOR_PATTERNS`
|
||||
|
||||
**Interface with parent**:
|
||||
```typescript
|
||||
export class RalphStatusParser extends EventEmitter {
|
||||
processLine(line: string): void { ... } // calls processStatusBlockLine + detectCompletionIndicators
|
||||
|
||||
get circuitBreakerStatus(): CircuitBreakerStatus { ... }
|
||||
get lastStatusBlock(): RalphStatusBlock | null { ... }
|
||||
get exitGateMet(): boolean { ... }
|
||||
get cumulativeStats(): { ... } { ... }
|
||||
|
||||
resetCircuitBreaker(): void { ... }
|
||||
reset(): void { ... }
|
||||
}
|
||||
|
||||
// Events: 'statusBlockDetected', 'circuitBreakerUpdate', 'exitGateMet'
|
||||
```
|
||||
|
||||
**In `RalphTracker.processLine()`**:
|
||||
```typescript
|
||||
// Replace inline status block handling with delegation
|
||||
this.statusParser.processLine(line);
|
||||
```
|
||||
|
||||
### Step 2e: Keep in `ralph-tracker.ts` (~1,800 LOC)
|
||||
|
||||
The core remains tightly coupled and stays together:
|
||||
- Output parsing pipeline (`processTerminalData`, `processCleanData`, `processLine`)
|
||||
- Loop state management (`_loopState`, `detectLoopStatus`, `enable/disable/startLoop/stopLoop`)
|
||||
- Todo management (`_todos`, `detectTodoItems`, `addOrUpdateTodo`, `updateTodoStatus`, `getTodoStats`)
|
||||
- Completion detection (`detectCompletionPhrase`, `handleCompletionPhrase`, `calculateCompletionConfidence`)
|
||||
- All-tasks-complete detection (`detectAllTasksComplete`)
|
||||
- Auto-enable logic (`shouldAutoEnable`)
|
||||
- Lifecycle (`reset`, `fullReset`, `clear`, `restoreState`, `destroy`)
|
||||
- Event debouncing and buffering
|
||||
|
||||
The class coordinates the extracted modules via composition:
|
||||
```typescript
|
||||
export class RalphTracker extends EventEmitter {
|
||||
readonly planTracker = new RalphPlanTracker();
|
||||
readonly fixPlanWatcher = new RalphFixPlanWatcher();
|
||||
readonly stallDetector: RalphStallDetector;
|
||||
readonly statusParser = new RalphStatusParser();
|
||||
|
||||
constructor() {
|
||||
super();
|
||||
this.stallDetector = new RalphStallDetector(this.cleanup);
|
||||
this._wireSubModuleEvents();
|
||||
}
|
||||
|
||||
private _wireSubModuleEvents(): void {
|
||||
// Forward all sub-module events through RalphTracker
|
||||
// so external consumers don't need to know about the split
|
||||
for (const event of ['planInitialized', 'planTaskUpdate', ...]) {
|
||||
this.planTracker.on(event, (...args) => this.emit(event, ...args));
|
||||
}
|
||||
// ...same for statusParser, stallDetector, fixPlanWatcher
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Migration Safety
|
||||
|
||||
- All events continue to be emitted from `RalphTracker` (forwarded from sub-modules)
|
||||
- All public methods stay on `RalphTracker` (delegated to sub-modules)
|
||||
- External consumers (`session.ts`, `case-routes.ts`) see zero API changes
|
||||
- New sub-modules are exposed as `readonly` properties for direct access where needed
|
||||
|
||||
---
|
||||
|
||||
## 3. Split respawn-controller.ts into focused modules
|
||||
|
||||
**Current**: 3,611 lines, single `RespawnController` class with 6 responsibilities.
|
||||
**Risk**: MEDIUM — health scoring and metrics are cleanly decoupled; detection is tightly coupled.
|
||||
|
||||
### Coupling Analysis Summary
|
||||
|
||||
| Module | Coupling | Extractability |
|
||||
|--------|----------|----------------|
|
||||
| Health scoring | NONE | HIGH — pure calculations from metrics |
|
||||
| Cycle metrics | LOW | HIGH — standalone tracking |
|
||||
| Adaptive timing | LOW | HIGH — standalone timing adjustments |
|
||||
| Stuck-state detection | LOW | MEDIUM — needs state + config refs |
|
||||
| Pattern detection utilities | NONE | HIGH — pure functions |
|
||||
| State machine + idle detection + AI checkers | HIGH | LOW — deeply entangled |
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/
|
||||
├── respawn-controller.ts (~2,200 LOC — state machine, idle detection,
|
||||
│ AI checkers, terminal handling, hook signals,
|
||||
│ auto-accept, step execution)
|
||||
├── respawn-health.ts (~250 LOC — health scoring + recommendations)
|
||||
├── respawn-metrics.ts (~200 LOC — cycle metrics + aggregate stats)
|
||||
├── respawn-adaptive-timing.ts (~100 LOC — adaptive timing with percentile calc)
|
||||
└── respawn-patterns.ts (~50 LOC — terminal pattern detection utilities)
|
||||
```
|
||||
|
||||
### Step 3a: Extract `RespawnPatterns` (~50 LOC)
|
||||
|
||||
**Pure utility functions, zero coupling**.
|
||||
|
||||
Move:
|
||||
- `isCompletionMessage(data): boolean`
|
||||
- `hasWorkingPattern(data, window): boolean`
|
||||
- `extractTokenCount(data): number | null`
|
||||
- `PROMPT_PATTERNS` array
|
||||
- `WORKING_PATTERNS` array
|
||||
|
||||
```typescript
|
||||
// src/respawn-patterns.ts
|
||||
import { TOKEN_PATTERN, SPINNER_PATTERN } from './utils/index.js';
|
||||
|
||||
export const PROMPT_PATTERNS = ['❯', '>', '$', '%', '#'];
|
||||
|
||||
export const WORKING_PATTERNS = [/* 70+ patterns */];
|
||||
|
||||
export function isCompletionMessage(data: string): boolean { ... }
|
||||
export function hasWorkingPattern(data: string, window: string): boolean { ... }
|
||||
export function extractTokenCount(data: string): number | null { ... }
|
||||
```
|
||||
|
||||
**In `RespawnController`**: Import and call:
|
||||
```typescript
|
||||
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
|
||||
```
|
||||
|
||||
### Step 3b: Extract `RespawnAdaptiveTiming` (~100 LOC)
|
||||
|
||||
**Self-contained timing controller**.
|
||||
|
||||
Move properties:
|
||||
- `timingHistory: TimingHistory`
|
||||
|
||||
Move methods:
|
||||
- `recordTimingData(idleDetectionMs, cycleDurationMs)`
|
||||
- `updateAdaptiveTiming()`
|
||||
- `getTimingHistory()`
|
||||
- `getAdaptiveCompletionConfirmMs()`
|
||||
|
||||
```typescript
|
||||
export class RespawnAdaptiveTiming {
|
||||
private timingHistory: TimingHistory;
|
||||
|
||||
constructor(private config: { adaptiveMinConfirmMs: number; adaptiveMaxConfirmMs: number }) {
|
||||
this.timingHistory = { recentIdleDetectionMs: [], recentCycleDurationMs: [], ... };
|
||||
}
|
||||
|
||||
recordTimingData(idleDetectionMs: number, cycleDurationMs: number): void { ... }
|
||||
getAdaptiveCompletionConfirmMs(): number { ... }
|
||||
getTimingHistory(): TimingHistory { ... }
|
||||
reset(): void { ... }
|
||||
}
|
||||
```
|
||||
|
||||
### Step 3c: Extract `RespawnCycleMetrics` (~200 LOC)
|
||||
|
||||
**Standalone metrics tracker**.
|
||||
|
||||
Move properties:
|
||||
- `currentCycleMetrics`
|
||||
- `recentCycleMetrics[]`
|
||||
- `aggregateMetrics`
|
||||
- `MAX_CYCLE_METRICS_IN_MEMORY`
|
||||
|
||||
Move methods:
|
||||
- `startCycleMetrics(idleReason)`
|
||||
- `recordCycleStep(step)`
|
||||
- `completeCycleMetrics(outcome, errorMessage?)`
|
||||
- `updateAggregateMetrics(metrics)`
|
||||
- `getAggregateMetrics()`
|
||||
- `getRecentCycleMetrics(limit?)`
|
||||
|
||||
```typescript
|
||||
export class RespawnCycleMetricsTracker {
|
||||
private currentCycleMetrics: Partial<RespawnCycleMetrics> | null = null;
|
||||
private recentCycleMetrics: RespawnCycleMetrics[] = [];
|
||||
private aggregateMetrics: RespawnAggregateMetrics;
|
||||
|
||||
startCycle(sessionId: string, cycleNumber: number, idleReason: string): void { ... }
|
||||
recordStep(step: string): void { ... }
|
||||
completeCycle(outcome: CycleOutcome, errorMessage?: string): RespawnCycleMetrics | null { ... }
|
||||
getAggregate(): RespawnAggregateMetrics { ... }
|
||||
getRecent(limit?: number): RespawnCycleMetrics[] { ... }
|
||||
reset(): void { ... }
|
||||
}
|
||||
```
|
||||
|
||||
**Callback**: `completeCycle()` returns the completed metrics so the controller can pass them to `adaptiveTiming.recordTimingData()`.
|
||||
|
||||
### Step 3d: Extract `RespawnHealthCalculator` (~250 LOC)
|
||||
|
||||
**Pure calculation — no state of its own**.
|
||||
|
||||
Move methods:
|
||||
- `calculateHealthScore()`
|
||||
- `calculateCycleSuccessScore()`
|
||||
- `calculateCircuitBreakerScore()`
|
||||
- `calculateIterationProgressScore()`
|
||||
- `calculateAiCheckerScore()`
|
||||
- `calculateStuckRecoveryScore()`
|
||||
- `generateHealthRecommendations(components)`
|
||||
- `generateHealthSummary(score, status, components)`
|
||||
- `shouldSkipClear()` (belongs here since it's a pure calculation on token/config)
|
||||
|
||||
```typescript
|
||||
export interface HealthInputs {
|
||||
aggregateMetrics: RespawnAggregateMetrics;
|
||||
circuitBreakerStatus: CircuitBreakerStatus;
|
||||
iterationStallMetrics: { stallDurationMs: number; warningMs: number; criticalMs: number } | null;
|
||||
aiCheckerState: { disabled: boolean; inCooldown: boolean; hasErrors: boolean };
|
||||
stuckRecoveryCount: number;
|
||||
maxStuckRecoveries: number;
|
||||
}
|
||||
|
||||
export function calculateHealthScore(inputs: HealthInputs): RalphLoopHealthScore { ... }
|
||||
|
||||
export function shouldSkipClear(
|
||||
lastTokenCount: number,
|
||||
skipClearThresholdPercent: number,
|
||||
maxContextTokens: number
|
||||
): boolean { ... }
|
||||
```
|
||||
|
||||
**Made as pure functions** (not a class) since they hold no state.
|
||||
|
||||
### Step 3e: Keep in `respawn-controller.ts` (~2,200 LOC)
|
||||
|
||||
The core state machine, idle detection, and AI checker integration stays:
|
||||
- State machine transitions (`setState`, `start`, `stop`, `pause`, `resume`)
|
||||
- Terminal data handling (`handleTerminalData`)
|
||||
- All 5 idle detection layers + hook signals
|
||||
- AI checker integration (`tryStartAiCheck`, `startAiCheck`, `startPlanCheck`)
|
||||
- Auto-accept logic
|
||||
- Step execution (`sendUpdateDocs`, `sendClear`, `sendInit`, `sendKickstart`)
|
||||
- Timer management (`startTrackedTimer`, `cancelTrackedTimer`)
|
||||
- Stuck-state detection and recovery
|
||||
- Action logging
|
||||
|
||||
The class composes extracted modules:
|
||||
```typescript
|
||||
import { RespawnAdaptiveTiming } from './respawn-adaptive-timing.js';
|
||||
import { RespawnCycleMetricsTracker } from './respawn-metrics.js';
|
||||
import { calculateHealthScore, shouldSkipClear } from './respawn-health.js';
|
||||
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
|
||||
|
||||
export class RespawnController extends EventEmitter {
|
||||
private adaptiveTiming: RespawnAdaptiveTiming;
|
||||
private cycleMetrics: RespawnCycleMetricsTracker;
|
||||
|
||||
calculateHealthScore(): RalphLoopHealthScore {
|
||||
return calculateHealthScore({
|
||||
aggregateMetrics: this.cycleMetrics.getAggregate(),
|
||||
circuitBreakerStatus: this.session.ralphTracker.circuitBreakerStatus,
|
||||
iterationStallMetrics: this.session.ralphTracker.getIterationStallMetrics(),
|
||||
aiCheckerState: { ... },
|
||||
stuckRecoveryCount: this.stuckRecoveryCount,
|
||||
maxStuckRecoveries: this.config.maxStuckRecoveries ?? 3,
|
||||
});
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Split session.ts into focused modules
|
||||
|
||||
**Current**: 2,418 lines, single `Session` class.
|
||||
**Risk**: LOW-MEDIUM — extractable pieces are utility-like with clear boundaries.
|
||||
|
||||
### Coupling Analysis Summary
|
||||
|
||||
| Module | Coupling | Extractability |
|
||||
|--------|----------|----------------|
|
||||
| CLI arg builder | NONE | HIGH — pure functions used at spawn time |
|
||||
| Auto-compact/clear | LOW | HIGH — self-contained automation with config |
|
||||
| Token tracking | LOW | MEDIUM — reads PTY output, writes state |
|
||||
| Task description cache | LOW | HIGH — separate LRU cache |
|
||||
| PTY + mux lifecycle | HIGH | KEEP — core of the class |
|
||||
| Tracker integration | HIGH | KEEP — event forwarding plumbing |
|
||||
|
||||
### Target Structure
|
||||
|
||||
```
|
||||
src/
|
||||
├── session.ts (~1,600 LOC — PTY lifecycle, terminal I/O,
|
||||
│ tracker integration, output processing,
|
||||
│ token tracking, state management)
|
||||
├── session-cli-builder.ts (~250 LOC — Claude/OpenCode CLI arg construction)
|
||||
├── session-auto-ops.ts (~300 LOC — auto-compact, auto-clear automation)
|
||||
└── session-task-cache.ts (~100 LOC — task description LRU cache)
|
||||
```
|
||||
|
||||
### Step 4a: Extract `SessionCliBuilder` (~250 LOC)
|
||||
|
||||
**Pure functions — zero coupling to Session instance**.
|
||||
|
||||
Move:
|
||||
- `buildClaudeArgs()` logic (currently inlined in `startInteractive` and `runPrompt`)
|
||||
- `buildOpenCodeArgs()` logic
|
||||
- Model mapping constants
|
||||
- Claude mode to flag mapping
|
||||
- Environment variable construction
|
||||
|
||||
```typescript
|
||||
// src/session-cli-builder.ts
|
||||
export interface CliBuilderConfig {
|
||||
claudeMode: ClaudeMode;
|
||||
model?: string;
|
||||
workingDir: string;
|
||||
sessionId: string;
|
||||
niceConfig?: NiceConfig;
|
||||
isOpenCode?: boolean;
|
||||
openCodeConfig?: OpenCodeConfig;
|
||||
}
|
||||
|
||||
export function buildInteractiveArgs(config: CliBuilderConfig): string[] { ... }
|
||||
export function buildPromptArgs(config: CliBuilderConfig, prompt: string): string[] { ... }
|
||||
export function buildShellArgs(shell?: string): string[] { ... }
|
||||
export function buildClaudeEnv(config: CliBuilderConfig): Record<string, string> { ... }
|
||||
```
|
||||
|
||||
### Step 4b: Extract `SessionAutoOps` (~300 LOC)
|
||||
|
||||
**Self-contained automation with config-based thresholds**.
|
||||
|
||||
Move properties:
|
||||
- `_autoCompactThreshold`
|
||||
- `_autoClearThreshold`
|
||||
- `_isAutoCompacting`
|
||||
- `_isAutoClearing`
|
||||
- `_autoCompactCount`
|
||||
- `_autoClearCount`
|
||||
- `_lastAutoCompactTime`
|
||||
- `_lastAutoClearTime`
|
||||
|
||||
Move methods:
|
||||
- `checkAutoCompact(tokenCount)`
|
||||
- `performAutoCompact()`
|
||||
- `checkAutoClear(tokenCount)`
|
||||
- `performAutoClear()`
|
||||
- Auto-compact/clear threshold configuration
|
||||
|
||||
```typescript
|
||||
export class SessionAutoOps extends EventEmitter {
|
||||
constructor(
|
||||
private writeCommand: (command: string) => Promise<void>,
|
||||
private getTokenCount: () => number,
|
||||
config: { compactThreshold: number; clearThreshold: number }
|
||||
) { ... }
|
||||
|
||||
/** Called after token count updates. Checks thresholds and triggers if needed. */
|
||||
checkThresholds(tokenCount: number): void { ... }
|
||||
|
||||
updateConfig(config: { compactThreshold?: number; clearThreshold?: number }): void { ... }
|
||||
getStats(): { autoCompactCount: number; autoClearCount: number; ... } { ... }
|
||||
}
|
||||
|
||||
// Events: 'autoCompact', 'autoClear'
|
||||
```
|
||||
|
||||
**In `Session`**: Compose and wire:
|
||||
```typescript
|
||||
private autoOps = new SessionAutoOps(
|
||||
(cmd) => this.writeViaMux(cmd),
|
||||
() => this._state.tokenCount,
|
||||
{ compactThreshold: 110_000, clearThreshold: 140_000 }
|
||||
);
|
||||
```
|
||||
|
||||
### Step 4c: Extract `SessionTaskCache` (~100 LOC)
|
||||
|
||||
**Isolated LRU cache for task descriptions**.
|
||||
|
||||
Move:
|
||||
- `_taskDescriptionCache: LRUMap<number, { description: string; timestamp: number }>`
|
||||
- `_taskDescriptionMaxAge`
|
||||
- `findTaskDescriptionNear(lineNumber)`
|
||||
- `cacheTaskDescription(lineNumber, description)`
|
||||
|
||||
```typescript
|
||||
export class SessionTaskCache {
|
||||
private cache: LRUMap<number, { description: string; timestamp: number }>;
|
||||
private maxAgeMs: number;
|
||||
|
||||
constructor(maxSize: number = 50, maxAgeMs: number = 30_000) { ... }
|
||||
|
||||
find(lineNumber: number, searchRadius: number = 50): string | null { ... }
|
||||
add(lineNumber: number, description: string): void { ... }
|
||||
clear(): void { ... }
|
||||
}
|
||||
```
|
||||
|
||||
### Step 4d: Keep in `session.ts` (~1,600 LOC)
|
||||
|
||||
The core stays together:
|
||||
- PTY process management (`spawn`, `kill`, `resize`, `writeViaMux`)
|
||||
- Data streaming pipeline (PTY → buffer → ANSI strip → JSON parse → events)
|
||||
- Tracker initialization and event forwarding (RalphTracker, BashToolParser, TaskTracker)
|
||||
- Output processing (message extraction, completion detection)
|
||||
- Token tracking (status line parsing)
|
||||
- State management (`toState()`, `updateState()`)
|
||||
- Session lifecycle (`startInteractive`, `startShell`, `runPrompt`)
|
||||
- CLI info detection (version, model, account)
|
||||
|
||||
---
|
||||
|
||||
## 5. Execution Order & Dependencies
|
||||
|
||||
Execute in this order to minimize risk. Each step is independently deployable.
|
||||
|
||||
```
|
||||
Step 1: types.ts split
|
||||
↓ (no runtime change, just file reorganization)
|
||||
Step 2a: RalphPlanTracker extraction
|
||||
↓ (independent of types split)
|
||||
Step 2b: RalphFixPlanWatcher extraction
|
||||
Step 2c: RalphStallDetector extraction
|
||||
Step 2d: RalphStatusParser extraction
|
||||
↓ (ralph-tracker.ts now ~1,800 LOC)
|
||||
Step 3a: RespawnPatterns extraction
|
||||
Step 3b: RespawnAdaptiveTiming extraction
|
||||
Step 3c: RespawnCycleMetrics extraction
|
||||
Step 3d: RespawnHealthCalculator extraction
|
||||
↓ (respawn-controller.ts now ~2,200 LOC)
|
||||
Step 4a: SessionCliBuilder extraction
|
||||
Step 4b: SessionAutoOps extraction
|
||||
Step 4c: SessionTaskCache extraction
|
||||
↓ (session.ts now ~1,600 LOC)
|
||||
```
|
||||
|
||||
**Parallelization**: Steps 1, 2a-2d, 3a-3d, and 4a-4c can be done by separate agents in parallel since they touch different files. However, within each group, sequential execution is safer.
|
||||
|
||||
### Risk Mitigation
|
||||
|
||||
- **Barrel exports**: Every split uses delegation + barrel re-export so external consumers see zero API changes
|
||||
- **Event forwarding**: Sub-modules emit events, parent class forwards them — no event contract changes
|
||||
- **Incremental**: Each step can be verified independently with `tsc --noEmit` + `npm run lint`
|
||||
- **No test changes needed**: External API stays identical; existing tests continue to pass
|
||||
|
||||
---
|
||||
|
||||
## 6. Validation Checklist
|
||||
|
||||
After each step, verify:
|
||||
|
||||
- [ ] `tsc --noEmit` passes (no type errors)
|
||||
- [ ] `npm run lint` passes (no unused imports, etc.)
|
||||
- [ ] `npm run format:check` passes
|
||||
- [ ] `npx vitest run test/respawn-controller.test.ts` passes (for respawn splits)
|
||||
- [ ] `npx vitest run test/ralph-tracker.test.ts` passes (for ralph splits)
|
||||
- [ ] `npx vitest run test/session-manager.test.ts` passes (for session splits)
|
||||
- [ ] Dev server starts: `npx tsx src/index.ts web`
|
||||
- [ ] Existing sessions work (create, interact, delete)
|
||||
- [ ] Respawn cycle works (enable respawn, verify idle detection fires)
|
||||
- [ ] No new circular dependencies: `npx madge --circular src/`
|
||||
|
||||
### Size Targets
|
||||
|
||||
| File | Before | After |
|
||||
|------|--------|-------|
|
||||
| `src/types.ts` | 1,443 LOC | 1 LOC (re-export barrel) |
|
||||
| `src/ralph-tracker.ts` | 3,868 LOC | ~1,800 LOC |
|
||||
| `src/respawn-controller.ts` | 3,611 LOC | ~2,200 LOC |
|
||||
| `src/session.ts` | 2,418 LOC | ~1,600 LOC |
|
||||
| **Total new files** | — | 12 files |
|
||||
| **Net LOC change** | — | ~0 (refactor only) |
|
||||
@@ -0,0 +1,738 @@
|
||||
# Phase 1 Implementation Plan: Quick Wins
|
||||
|
||||
**Source**: `docs/code-structure-findings.md` (Phase 1 - Quick Wins section)
|
||||
**Estimated effort**: 1-2 days
|
||||
**Tasks**: 5 independent tasks (can be done in parallel unless noted)
|
||||
|
||||
---
|
||||
|
||||
## Safety Constraints
|
||||
|
||||
Before starting ANY work, read and follow these rules:
|
||||
|
||||
1. **Never run `npx vitest run`** (full suite) -- it kills tmux sessions. You are running inside a Codeman-managed tmux session.
|
||||
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
|
||||
3. **Never test on port 3000** -- the live dev server runs there. Tests use ports 3150+.
|
||||
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
|
||||
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
|
||||
6. **Never kill tmux sessions** -- check `echo $CODEMAN_MUX` first.
|
||||
|
||||
---
|
||||
|
||||
## Task Dependencies
|
||||
|
||||
All 5 tasks are independent and can be done in parallel. However:
|
||||
- Task 1 (barrel exports) is a prerequisite if you want to update import sites to use the barrel after Task 3 (consolidate EXEC_TIMEOUT_MS). The EXEC_TIMEOUT_MS consolidation creates a new export that should be added to the barrel.
|
||||
- Task 2 (delete dead functions) removes functions that Task 1 would otherwise need to add to the barrel. Do Task 2 first or simultaneously with Task 1 to avoid adding exports for dead code.
|
||||
|
||||
**Recommended order**: Task 2 -> Task 1 -> Task 3 -> Task 4 -> Task 5
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Export Missing Functions from Utils Barrel
|
||||
|
||||
**File**: `src/utils/index.ts`
|
||||
**Time**: ~30 minutes
|
||||
|
||||
### Problem
|
||||
|
||||
The barrel file (`src/utils/index.ts`) is missing exports for several functions that are defined in util modules, forcing consumers to use deep imports or preventing usage entirely.
|
||||
|
||||
### Missing Exports
|
||||
|
||||
From `src/utils/regex-patterns.ts`:
|
||||
- `createAnsiPatternFull()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
|
||||
- `createAnsiPatternSimple()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
|
||||
- `stripAnsi()` -- ANSI stripping utility
|
||||
- `SAFE_PATH_PATTERN` -- regex for safe file paths (currently deep-imported by `schemas.ts` and `tmux-manager.ts`)
|
||||
|
||||
From `src/utils/token-validation.ts`:
|
||||
- `validateTokenCounts()` -- token count validation (documented in CLAUDE.md)
|
||||
- `validateTokensAndCost()` -- token + cost validation (documented in CLAUDE.md)
|
||||
|
||||
**Note**: Do NOT export `isSimilar`, `isSimilarByDistance`, `levenshteinDistance`, or `normalizePhrase` from `string-similarity.ts` -- these are dead code (see Task 2).
|
||||
|
||||
### Edit 1: Add missing regex-patterns exports
|
||||
|
||||
**File**: `src/utils/index.ts`
|
||||
|
||||
**Old code** (lines 13-18):
|
||||
```typescript
|
||||
export {
|
||||
ANSI_ESCAPE_PATTERN_FULL,
|
||||
ANSI_ESCAPE_PATTERN_SIMPLE,
|
||||
TOKEN_PATTERN,
|
||||
SPINNER_PATTERN,
|
||||
} from './regex-patterns.js';
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
export {
|
||||
ANSI_ESCAPE_PATTERN_FULL,
|
||||
ANSI_ESCAPE_PATTERN_SIMPLE,
|
||||
TOKEN_PATTERN,
|
||||
SPINNER_PATTERN,
|
||||
createAnsiPatternFull,
|
||||
createAnsiPatternSimple,
|
||||
stripAnsi,
|
||||
SAFE_PATH_PATTERN,
|
||||
} from './regex-patterns.js';
|
||||
```
|
||||
|
||||
### Edit 2: Add missing token-validation exports
|
||||
|
||||
**File**: `src/utils/index.ts`
|
||||
|
||||
**Old code** (line 19):
|
||||
```typescript
|
||||
export { MAX_SESSION_TOKENS } from './token-validation.js';
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
export { MAX_SESSION_TOKENS, validateTokenCounts, validateTokensAndCost } from './token-validation.js';
|
||||
```
|
||||
|
||||
### Optional follow-up: Update deep imports to use barrel
|
||||
|
||||
These files currently deep-import `SAFE_PATH_PATTERN` and could be updated to use the barrel instead:
|
||||
|
||||
- `src/web/schemas.ts` line 11: `import { SAFE_PATH_PATTERN } from '../utils/regex-patterns.js';` could become `import { SAFE_PATH_PATTERN } from '../utils/index.js';`
|
||||
- `src/tmux-manager.ts` line 44: `import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';` could become part of existing barrel import
|
||||
|
||||
This is a low-priority cosmetic change. The barrel export itself is the important fix.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Delete Dead Utility Functions
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
**Time**: ~15 minutes
|
||||
|
||||
### Problem
|
||||
|
||||
Four exported functions in `string-similarity.ts` are never imported anywhere in the codebase:
|
||||
- `levenshteinDistance()` (lines 27-69)
|
||||
- `isSimilar()` (lines 106-108)
|
||||
- `isSimilarByDistance()` (lines 123-125)
|
||||
- `normalizePhrase()` (lines 139-144)
|
||||
|
||||
Only three functions are actually used (all by `ralph-tracker.ts` via the barrel):
|
||||
- `stringSimilarity()` -- uses `levenshteinDistance()` internally
|
||||
- `fuzzyPhraseMatch()` -- uses `normalizePhrase()` and `isSimilarByDistance()` internally
|
||||
- `todoContentHash()`
|
||||
|
||||
### Strategy
|
||||
|
||||
`levenshteinDistance()` is called by `stringSimilarity()`, and `normalizePhrase()` and `isSimilarByDistance()` are called by `fuzzyPhraseMatch()`. So they cannot be deleted -- they just need to be un-exported (made private to the module).
|
||||
|
||||
`isSimilar()` is truly dead -- not called by anything. Delete it entirely.
|
||||
|
||||
### Edit 1: Remove `export` from `levenshteinDistance`
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (line 27):
|
||||
```typescript
|
||||
export function levenshteinDistance(a: string, b: string): number {
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
function levenshteinDistance(a: string, b: string): number {
|
||||
```
|
||||
|
||||
### Edit 2: Delete `isSimilar` function entirely
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (lines 94-108):
|
||||
```typescript
|
||||
/**
|
||||
* Check if two strings are similar within a given threshold.
|
||||
*
|
||||
* @param a - First string
|
||||
* @param b - Second string
|
||||
* @param threshold - Minimum similarity ratio (default: 0.85 = 85% similar)
|
||||
* @returns True if similarity >= threshold
|
||||
*
|
||||
* @example
|
||||
* isSimilar('COMPLETE', 'COMPLET', 0.85) // true (87.5% similar)
|
||||
* isSimilar('COMPLETE', 'DONE', 0.85) // false (0% similar)
|
||||
*/
|
||||
export function isSimilar(a: string, b: string, threshold = 0.85): boolean {
|
||||
return stringSimilarity(a, b) >= threshold;
|
||||
}
|
||||
```
|
||||
|
||||
**New code**: (delete entirely -- replace with empty string)
|
||||
|
||||
### Edit 3: Remove `export` from `isSimilarByDistance`
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (line 123):
|
||||
```typescript
|
||||
export function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
|
||||
```
|
||||
|
||||
### Edit 4: Remove `export` from `normalizePhrase`
|
||||
|
||||
**File**: `src/utils/string-similarity.ts`
|
||||
|
||||
**Old code** (line 139):
|
||||
```typescript
|
||||
export function normalizePhrase(phrase: string): string {
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
function normalizePhrase(phrase: string): string {
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npx vitest run test/string-utilities.test.ts
|
||||
npm run lint
|
||||
```
|
||||
|
||||
Note: If `test/string-utilities.test.ts` imports any of the now-unexported functions, those test imports will fail. Check the test file and remove tests for `isSimilar` (deleted) and update any direct tests for `levenshteinDistance`, `isSimilarByDistance`, `normalizePhrase` to test them indirectly through the public API (`stringSimilarity`, `fuzzyPhraseMatch`), or remove those tests.
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Consolidate Duplicated `EXEC_TIMEOUT_MS` Constant
|
||||
|
||||
**Files**:
|
||||
- `src/utils/claude-cli-resolver.ts` (line 17)
|
||||
- `src/utils/opencode-cli-resolver.ts` (line 16)
|
||||
- `src/tmux-manager.ts` (line 63) -- also has its own copy
|
||||
|
||||
**Time**: ~15 minutes
|
||||
|
||||
### Problem
|
||||
|
||||
`EXEC_TIMEOUT_MS = 5000` is defined identically in three files. Changes need to happen in all three places.
|
||||
|
||||
### Strategy
|
||||
|
||||
Create a shared constant and export it. The natural home is a new config file since the existing config files (`buffer-limits.ts`, `map-limits.ts`) follow this pattern. However, to keep it minimal, we can add it to an existing config file or create a small one.
|
||||
|
||||
**Recommended approach**: Add to `src/config/timing-config.ts` (new file) as a single constant. This file can grow later in Phase 6 to hold other timing constants.
|
||||
|
||||
Alternatively, the simplest approach: export from one of the existing utils and import in the others. Since both CLI resolvers are in `src/utils/`, the cleanest approach is to put it in a shared location.
|
||||
|
||||
### Option A: Add to existing config (simpler)
|
||||
|
||||
Create `src/config/exec-timeout.ts`:
|
||||
|
||||
**New file**: `src/config/exec-timeout.ts`
|
||||
```typescript
|
||||
/**
|
||||
* Timeout for child process exec commands (e.g., `which claude`, `which opencode`, tmux commands).
|
||||
* Used across CLI resolvers and tmux manager.
|
||||
*/
|
||||
export const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
### Edit 1: Update `claude-cli-resolver.ts`
|
||||
|
||||
**File**: `src/utils/claude-cli-resolver.ts`
|
||||
|
||||
**Old code** (lines 11-17):
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { delimiter, dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
|
||||
/** Timeout for exec commands (5 seconds) */
|
||||
const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { delimiter, dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
|
||||
```
|
||||
|
||||
### Edit 2: Update `opencode-cli-resolver.ts`
|
||||
|
||||
**File**: `src/utils/opencode-cli-resolver.ts`
|
||||
|
||||
**Old code** (lines 10-16):
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
|
||||
/** Timeout for exec commands (5 seconds) */
|
||||
const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
import { execSync } from 'node:child_process';
|
||||
import { existsSync } from 'node:fs';
|
||||
import { dirname, join } from 'node:path';
|
||||
import { homedir } from 'node:os';
|
||||
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
|
||||
```
|
||||
|
||||
### Edit 3: Update `tmux-manager.ts`
|
||||
|
||||
**File**: `src/tmux-manager.ts`
|
||||
|
||||
**Old code** (line 63):
|
||||
```typescript
|
||||
const EXEC_TIMEOUT_MS = 5000;
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
|
||||
```
|
||||
|
||||
Note: `tmux-manager.ts` already has many imports at the top of the file. Add this import near the other local imports (around lines 43-56). The `const EXEC_TIMEOUT_MS = 5000;` on line 63 should be deleted entirely (replaced with the import).
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Add `z.infer` to Zod Schemas
|
||||
|
||||
**Files**:
|
||||
- `src/web/schemas.ts` (add type exports)
|
||||
- `src/types.ts` (replace manual interfaces with `z.infer` re-exports where applicable)
|
||||
|
||||
**Time**: ~2 hours
|
||||
|
||||
### Problem
|
||||
|
||||
All 30+ Zod schemas in `schemas.ts` define validation rules, but zero use `z.infer` to derive TypeScript types. Instead, `types.ts` manually duplicates interfaces that match the schemas. When a schema changes, the type must be manually updated too.
|
||||
|
||||
### Strategy
|
||||
|
||||
Add `z.infer` type exports to `schemas.ts` for each exported schema. This creates derived types as the single source of truth. For schemas that have corresponding manual interfaces in `types.ts`, the manual interface can be replaced with a re-export of the inferred type.
|
||||
|
||||
**Important**: Not all schemas have matching interfaces in `types.ts`. The `RespawnConfig` interface in `types.ts` (line 395) has all required fields, while `RespawnConfigSchema` has all optional fields (it's for partial updates). These are NOT the same type and should NOT be unified.
|
||||
|
||||
### Edit 1: Add inferred type exports to `schemas.ts`
|
||||
|
||||
**File**: `src/web/schemas.ts`
|
||||
|
||||
After each schema definition, add a corresponding type export. Add the following lines at the **end of the file** (after line 509):
|
||||
|
||||
**Old code** (end of file, lines 506-509):
|
||||
```typescript
|
||||
.optional(),
|
||||
});
|
||||
```
|
||||
|
||||
Wait -- the end of file is actually at line 509 after the `RalphLoopStartSchema`. Add the type exports after the last schema:
|
||||
|
||||
**Append to end of file** `src/web/schemas.ts`:
|
||||
|
||||
```typescript
|
||||
|
||||
// ========== Inferred Types ==========
|
||||
// Derive TypeScript types from Zod schemas (single source of truth)
|
||||
|
||||
export type CreateSessionInput = z.infer<typeof CreateSessionSchema>;
|
||||
export type RunPromptInput = z.infer<typeof RunPromptSchema>;
|
||||
export type ResizeInput = z.infer<typeof ResizeSchema>;
|
||||
export type CreateCaseInput = z.infer<typeof CreateCaseSchema>;
|
||||
export type QuickStartInput = z.infer<typeof QuickStartSchema>;
|
||||
export type HookEventInput = z.infer<typeof HookEventSchema>;
|
||||
export type RespawnConfigInput = z.infer<typeof RespawnConfigSchema>;
|
||||
export type ConfigUpdateInput = z.infer<typeof ConfigUpdateSchema>;
|
||||
export type SettingsUpdateInput = z.infer<typeof SettingsUpdateSchema>;
|
||||
export type SessionInputWithLimitInput = z.infer<typeof SessionInputWithLimitSchema>;
|
||||
export type SessionNameInput = z.infer<typeof SessionNameSchema>;
|
||||
export type SessionColorInput = z.infer<typeof SessionColorSchema>;
|
||||
export type RalphConfigInput = z.infer<typeof RalphConfigSchema>;
|
||||
export type FixPlanImportInput = z.infer<typeof FixPlanImportSchema>;
|
||||
export type RalphPromptWriteInput = z.infer<typeof RalphPromptWriteSchema>;
|
||||
export type AutoClearInput = z.infer<typeof AutoClearSchema>;
|
||||
export type AutoCompactInput = z.infer<typeof AutoCompactSchema>;
|
||||
export type ImageWatcherInput = z.infer<typeof ImageWatcherSchema>;
|
||||
export type FlickerFilterInput = z.infer<typeof FlickerFilterSchema>;
|
||||
export type QuickRunInput = z.infer<typeof QuickRunSchema>;
|
||||
export type ScheduledRunInput = z.infer<typeof ScheduledRunSchema>;
|
||||
export type LinkCaseInput = z.infer<typeof LinkCaseSchema>;
|
||||
export type GeneratePlanInput = z.infer<typeof GeneratePlanSchema>;
|
||||
export type GeneratePlanDetailedInput = z.infer<typeof GeneratePlanDetailedSchema>;
|
||||
export type CancelPlanInput = z.infer<typeof CancelPlanSchema>;
|
||||
export type PlanTaskUpdateInput = z.infer<typeof PlanTaskUpdateSchema>;
|
||||
export type PlanTaskAddInput = z.infer<typeof PlanTaskAddSchema>;
|
||||
export type CpuLimitInput = z.infer<typeof CpuLimitSchema>;
|
||||
export type SubagentWindowStatesInput = z.infer<typeof SubagentWindowStatesSchema>;
|
||||
export type SubagentParentMapInput = z.infer<typeof SubagentParentMapSchema>;
|
||||
export type InteractiveRespawnInput = z.infer<typeof InteractiveRespawnSchema>;
|
||||
export type RespawnEnableInput = z.infer<typeof RespawnEnableSchema>;
|
||||
export type PushSubscribeInput = z.infer<typeof PushSubscribeSchema>;
|
||||
export type PushPreferencesUpdateInput = z.infer<typeof PushPreferencesUpdateSchema>;
|
||||
export type RalphLoopStartInput = z.infer<typeof RalphLoopStartSchema>;
|
||||
```
|
||||
|
||||
### What NOT to do
|
||||
|
||||
Do NOT replace the `RespawnConfig` interface in `types.ts` with `z.infer<typeof RespawnConfigSchema>`. The schema has all optional fields (for partial config updates), but the interface has required fields (for the full config object). These are intentionally different shapes.
|
||||
|
||||
Similarly, do NOT try to unify every interface in `types.ts` with a schema -- most interfaces in `types.ts` represent internal domain objects (SessionState, TaskState, etc.) that have no corresponding Zod schema. The schemas only exist for API request validation.
|
||||
|
||||
### Future opportunity
|
||||
|
||||
In a future phase, route handlers in `server.ts` can use these inferred types for request body typing:
|
||||
```typescript
|
||||
const body = CreateSessionSchema.parse(request.body) as CreateSessionInput;
|
||||
```
|
||||
This task only adds the type exports. Migrating route handlers to use them is out of scope.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
npm run format:check
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Fix Weak `not.toThrow()` Tests with Behavioral Assertions
|
||||
|
||||
**Files**:
|
||||
- `test/task-tracker.test.ts` -- 6 instances
|
||||
- `test/image-watcher.test.ts` -- 1 instance
|
||||
- `test/task-queue.test.ts` -- 1 instance
|
||||
- `test/hooks-config.test.ts` -- 1 instance
|
||||
- `test/session-manager.test.ts` -- 1 instance
|
||||
|
||||
**Time**: ~1 hour
|
||||
|
||||
### Problem
|
||||
|
||||
10 tests only assert `not.toThrow()` without verifying the actual defensive behavior. These tests prove the code doesn't crash but don't verify it does the right thing.
|
||||
|
||||
### Fix Strategy
|
||||
|
||||
After each `not.toThrow()`, add a behavioral assertion that verifies the state is correct (e.g., no tasks were created, no side effects occurred).
|
||||
|
||||
### Edit 1: `task-tracker.test.ts` -- null message (line 566)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle null message', () => {
|
||||
expect(() => tracker.processMessage(null)).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle null message', () => {
|
||||
expect(() => tracker.processMessage(null)).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
expect(tracker.getRunningCount()).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 2: `task-tracker.test.ts` -- message without content (line 569-571)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle message without content', () => {
|
||||
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle message without content', () => {
|
||||
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 3: `task-tracker.test.ts` -- empty content array (line 573-575)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle empty content array', () => {
|
||||
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle empty content array', () => {
|
||||
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 4: `task-tracker.test.ts` -- tool_result for unknown task (lines 577-590)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle tool_result for unknown task', () => {
|
||||
expect(() => {
|
||||
tracker.processMessage({
|
||||
message: {
|
||||
content: [{
|
||||
type: 'tool_result',
|
||||
tool_use_id: 'unknown-task',
|
||||
is_error: false,
|
||||
content: 'Done',
|
||||
}],
|
||||
},
|
||||
});
|
||||
}).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle tool_result for unknown task', () => {
|
||||
expect(() => {
|
||||
tracker.processMessage({
|
||||
message: {
|
||||
content: [{
|
||||
type: 'tool_result',
|
||||
tool_use_id: 'unknown-task',
|
||||
is_error: false,
|
||||
content: 'Done',
|
||||
}],
|
||||
},
|
||||
});
|
||||
}).not.toThrow();
|
||||
expect(tracker.getTask('unknown-task')).toBeUndefined();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 5: `task-tracker.test.ts` -- empty terminal output (lines 592-595)
|
||||
|
||||
**File**: `test/task-tracker.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle empty terminal output', () => {
|
||||
expect(() => tracker.processTerminalOutput('')).not.toThrow();
|
||||
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle empty terminal output', () => {
|
||||
expect(() => tracker.processTerminalOutput('')).not.toThrow();
|
||||
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
|
||||
expect(tracker.getAllTasks().size).toBe(0);
|
||||
expect(tracker.getRunningCount()).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 6: `image-watcher.test.ts` -- unwatchSession for non-watched session (line 123)
|
||||
|
||||
**File**: `test/image-watcher.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should be safe to call for non-watched session', () => {
|
||||
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should be safe to call for non-watched session', () => {
|
||||
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
|
||||
expect(watcher.getWatchedSessions()).toHaveLength(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 7: `task-queue.test.ts` -- dependencies on non-existent tasks (lines 538-542)
|
||||
|
||||
**File**: `test/task-queue.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
|
||||
// Dependencies on non-existent tasks are valid - they just won't be satisfied
|
||||
expect(() => {
|
||||
queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
|
||||
}).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
|
||||
// Dependencies on non-existent tasks are valid - they just won't be satisfied
|
||||
let task: ReturnType<typeof queue.addTask> | undefined;
|
||||
expect(() => {
|
||||
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
|
||||
}).not.toThrow();
|
||||
expect(task).toBeDefined();
|
||||
expect(task!.dependencies).toEqual(['non-existent-id']);
|
||||
// Task should be pending but blocked (dependency unsatisfied)
|
||||
expect(queue.next()?.prompt).toBeUndefined();
|
||||
});
|
||||
```
|
||||
|
||||
Wait -- `queue.next()` returns `null` when no next task is available (all blocked). Let me adjust:
|
||||
|
||||
**New code** (corrected):
|
||||
```typescript
|
||||
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
|
||||
// Dependencies on non-existent tasks are valid - they just won't be satisfied
|
||||
let task: ReturnType<typeof queue.addTask> | undefined;
|
||||
expect(() => {
|
||||
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
|
||||
}).not.toThrow();
|
||||
expect(task).toBeDefined();
|
||||
expect(task!.dependencies).toEqual(['non-existent-id']);
|
||||
// Task exists but is blocked (dependency unsatisfied), so next() skips it
|
||||
expect(queue.getAllTasks()).toHaveLength(1);
|
||||
expect(queue.next()).toBeNull();
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 8: `hooks-config.test.ts` -- valid JSON check (line 129)
|
||||
|
||||
**File**: `test/hooks-config.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should write valid JSON', () => {
|
||||
writeHooksConfig(testDir);
|
||||
const settingsPath = join(testDir, '.claude', 'settings.local.json');
|
||||
const content = readFileSync(settingsPath, 'utf-8');
|
||||
expect(() => JSON.parse(content)).not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should write valid JSON', () => {
|
||||
writeHooksConfig(testDir);
|
||||
const settingsPath = join(testDir, '.claude', 'settings.local.json');
|
||||
const content = readFileSync(settingsPath, 'utf-8');
|
||||
const parsed = JSON.parse(content);
|
||||
expect(parsed).toBeDefined();
|
||||
expect(typeof parsed).toBe('object');
|
||||
expect(parsed.hooks).toBeDefined();
|
||||
});
|
||||
```
|
||||
|
||||
### Edit 9: `session-manager.test.ts` -- stopSession for non-existent (line 216)
|
||||
|
||||
**File**: `test/session-manager.test.ts`
|
||||
|
||||
**Old code**:
|
||||
```typescript
|
||||
it('should handle non-existent session gracefully', async () => {
|
||||
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
|
||||
});
|
||||
```
|
||||
|
||||
**New code**:
|
||||
```typescript
|
||||
it('should handle non-existent session gracefully', async () => {
|
||||
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
|
||||
expect(manager.getSessionCount()).toBe(0);
|
||||
});
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
Run each test file individually:
|
||||
|
||||
```bash
|
||||
npx vitest run test/task-tracker.test.ts
|
||||
npx vitest run test/image-watcher.test.ts
|
||||
npx vitest run test/task-queue.test.ts
|
||||
npx vitest run test/hooks-config.test.ts
|
||||
npx vitest run test/session-manager.test.ts
|
||||
```
|
||||
|
||||
**Important**: `hooks-config.test.ts` and `session-manager.test.ts` spawn real servers on ports 3130-3131. Only run them if you are NOT running other tests that use those ports.
|
||||
|
||||
---
|
||||
|
||||
## Final Verification Checklist
|
||||
|
||||
After all 5 tasks are complete, run the following in order:
|
||||
|
||||
```bash
|
||||
# 1. TypeScript type checking
|
||||
tsc --noEmit
|
||||
|
||||
# 2. Linting
|
||||
npm run lint
|
||||
|
||||
# 3. Formatting
|
||||
npm run format:check
|
||||
|
||||
# 4. Run affected test files individually (NOT the full suite)
|
||||
npx vitest run test/string-utilities.test.ts
|
||||
npx vitest run test/task-tracker.test.ts
|
||||
npx vitest run test/image-watcher.test.ts
|
||||
npx vitest run test/task-queue.test.ts
|
||||
npx vitest run test/session-manager.test.ts
|
||||
npx vitest run test/hooks-config.test.ts
|
||||
```
|
||||
|
||||
If any formatting issues arise, fix with:
|
||||
```bash
|
||||
npm run format
|
||||
```
|
||||
|
||||
If any lint issues arise, fix with:
|
||||
```bash
|
||||
npm run lint:fix
|
||||
```
|
||||
|
||||
### Summary of Changes
|
||||
|
||||
| Task | Files Modified | Files Created |
|
||||
|------|---------------|---------------|
|
||||
| 1. Barrel exports | `src/utils/index.ts` | -- |
|
||||
| 2. Dead functions | `src/utils/string-similarity.ts` | -- |
|
||||
| 3. EXEC_TIMEOUT_MS | `src/utils/claude-cli-resolver.ts`, `src/utils/opencode-cli-resolver.ts`, `src/tmux-manager.ts` | `src/config/exec-timeout.ts` |
|
||||
| 4. z.infer types | `src/web/schemas.ts` | -- |
|
||||
| 5. Weak tests | `test/task-tracker.test.ts`, `test/image-watcher.test.ts`, `test/task-queue.test.ts`, `test/hooks-config.test.ts`, `test/session-manager.test.ts` | -- |
|
||||
|
||||
**Total files modified**: 10
|
||||
**Total files created**: 1
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,689 @@
|
||||
# Phase 6 Implementation Plan: Config Consolidation
|
||||
|
||||
**Source**: `docs/code-structure-findings.md` (Phase 6 — Config Consolidation)
|
||||
**Estimated effort**: 1 day
|
||||
**Tasks**: 8 tasks with dependencies (see dependency graph below)
|
||||
|
||||
---
|
||||
|
||||
## Safety Constraints
|
||||
|
||||
Before starting ANY work, read and follow these rules:
|
||||
|
||||
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
|
||||
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
|
||||
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
|
||||
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
|
||||
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
|
||||
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
|
||||
7. **Verify the dev server starts**: After each task, run `npx tsx src/index.ts web --port 3099 &` on a non-production port, confirm `curl -s http://localhost:3099/api/status | jq .status` returns `"ok"`, then kill the background process.
|
||||
|
||||
---
|
||||
|
||||
## Goal
|
||||
|
||||
Consolidate ~70 scattered numeric constants from 15+ source files into 6 new domain-focused config files, eliminating cross-file duplicates (including a 5x-duplicated AI model string) and making all tuning knobs discoverable in `src/config/`.
|
||||
|
||||
**Non-goal**: Moving every constant. Module-internal implementation details (like regex patterns, algorithm-specific magic numbers, or constants only used once in deeply coupled logic) stay where they are. The goal is discoverability of operational tuning knobs, not mechanical relocation.
|
||||
|
||||
---
|
||||
|
||||
## Design Decisions
|
||||
|
||||
### What gets centralized (and why)
|
||||
|
||||
Constants are candidates for centralization when they meet **any** of these criteria:
|
||||
|
||||
1. **Duplicated across files** — DRY violation (e.g., `STATS_COLLECTION_INTERVAL_MS` in `server.ts` and `mux-routes.ts`, AI model string in 5 files)
|
||||
2. **Operational tuning knobs** — values an operator might want to adjust for performance, security, or behavior without understanding the implementation (e.g., SSE health check interval, auth session TTL, rate limits)
|
||||
3. **Cross-cutting concerns** — values that establish system-wide contracts (e.g., max terminal dimensions used by both server routes and frontend)
|
||||
|
||||
### What stays in place (and why)
|
||||
|
||||
Constants that are **internal implementation details** of a single module stay where they are:
|
||||
|
||||
- **Algorithm parameters** — `TODO_SIMILARITY_THRESHOLD`, `adaptiveCompletionConfirmMs`, confidence weights. These are meaningless without understanding the algorithm.
|
||||
- **Display/UI formatting** — `TEXT_PREVIEW_LENGTH`, `SMART_TITLE_MAX_LENGTH`, `COMMAND_DISPLAY_LENGTH` in `subagent-watcher.ts`. Only used locally, tightly coupled to rendering logic.
|
||||
- **Module-internal timing** — `LINE_BUFFER_FLUSH_INTERVAL` in `session.ts`, `AI_CHECK_POLL_INTERVAL` in `ai-checker-base.ts`. Internal implementation of specific features.
|
||||
- **Frontend constants** — `constants.js` already centralizes frontend values well. Don't mix frontend and backend config.
|
||||
- **Respawn `DEFAULT_CONFIG`** — these are user-configurable defaults for the respawn config interface, not system constants. They live properly in `respawn-controller.ts`. The AI model/context defaults within it are replaced with imports from the new `ai-defaults.ts` (Task 5).
|
||||
- **Session auto-ops thresholds** — `AUTO_RETRY_DELAY_MS`, `COMPACT_COOLDOWN_MS`, etc. in `session-auto-ops.ts` are internal to that module's retry logic and already well-documented in place.
|
||||
|
||||
### File organization: domain-based, not category-based
|
||||
|
||||
A single `timing-config.ts` with 70 unrelated timing values would be worse than the current state — developers would need to grep it just like they grep the whole codebase now. Instead, constants are grouped by **the system they configure**:
|
||||
|
||||
| New File | Domain | Developer Question It Answers |
|
||||
|----------|--------|-------------------------------|
|
||||
| `server-timing.ts` | Web server performance | "How do I tune SSE batching / terminal throughput?" |
|
||||
| `auth-config.ts` | Authentication & security | "What are the rate limits and session TTLs?" |
|
||||
| `tunnel-config.ts` | QR auth & Cloudflare tunnel | "What are the QR token rotation parameters?" |
|
||||
| `terminal-limits.ts` | Terminal dimensions & input | "What are the max cols/rows/input size?" |
|
||||
| `ai-defaults.ts` | AI checker model & context | "What model do the AI checkers use? What's the context limit?" |
|
||||
| `team-config.ts` | Agent Teams polling & caching | "How often does team polling run? What are the cache limits?" |
|
||||
|
||||
---
|
||||
|
||||
## Task Dependencies
|
||||
|
||||
```
|
||||
Task 1 (server-timing.ts)
|
||||
Task 2 (auth-config.ts)
|
||||
Task 3 (tunnel-config.ts)
|
||||
Task 4 (terminal-limits.ts)
|
||||
Task 5 (ai-defaults.ts)
|
||||
Task 6 (team-config.ts)
|
||||
└──> Task 7 (Fix remaining duplicates)
|
||||
└──> Task 8 (Update CLAUDE.md + final verification)
|
||||
```
|
||||
|
||||
**Tasks 1–6** are independent and can run in parallel.
|
||||
**Task 7** depends on Tasks 1–6 (needs the new config files to exist).
|
||||
**Task 8** depends on Task 7.
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Create `src/config/server-timing.ts`
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Files created**: `src/config/server-timing.ts`
|
||||
**Files modified**: `src/web/server.ts`, `src/web/routes/mux-routes.ts`
|
||||
|
||||
### Constants to extract from `src/web/server.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `TERMINAL_BATCH_INTERVAL` | `16` | Terminal data batching interval (60fps) |
|
||||
| `TASK_UPDATE_BATCH_INTERVAL` | `100` | Task event batching interval (ms) |
|
||||
| `STATE_UPDATE_DEBOUNCE_INTERVAL` | `500` | State persistence debounce (ms) |
|
||||
| `SESSIONS_LIST_CACHE_TTL` | `1000` | Sessions list cache TTL (ms) |
|
||||
| `SCHEDULED_CLEANUP_INTERVAL` | `300000` | Scheduled runs cleanup check (5 min) |
|
||||
| `SCHEDULED_RUN_MAX_AGE` | `3600000` | Completed scheduled run max age (1 hour) |
|
||||
| `SSE_HEALTH_CHECK_INTERVAL` | `30000` | SSE client health check (30s) |
|
||||
| `SESSION_LIMIT_WAIT_MS` | `5000` | Session limit retry wait (5s) |
|
||||
| `ITERATION_PAUSE_MS` | `2000` | Scheduled run iteration pause (2s) |
|
||||
| `BATCH_FLUSH_THRESHOLD` | `32768` | Terminal batch immediate flush threshold (32KB) |
|
||||
| `STATS_COLLECTION_INTERVAL_MS` | `2000` | Mux stats collection interval (2s) |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/server-timing.ts` with all 11 constants, preserving existing JSDoc comments.
|
||||
2. In `src/web/server.ts`: Remove the 11 local constant declarations (lines ~92–121). Add `import { TERMINAL_BATCH_INTERVAL, ... } from '../config/server-timing.js'`.
|
||||
3. In `src/web/routes/mux-routes.ts`: Remove the duplicate `STATS_COLLECTION_INTERVAL_MS` (line 10) and its comment. Add `import { STATS_COLLECTION_INTERVAL_MS } from '../../config/server-timing.js'`. This fixes a **duplicate constant** (finding #10).
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Web server performance and scheduling constants.
|
||||
*
|
||||
* Controls terminal batching throughput, SSE health checking,
|
||||
* state persistence debouncing, and scheduled run timing.
|
||||
*
|
||||
* @module config/server-timing
|
||||
*/
|
||||
|
||||
// ============================================================================
|
||||
// Terminal & SSE Performance
|
||||
// ============================================================================
|
||||
|
||||
/** Terminal data batching interval — targets 60fps (ms) */
|
||||
export const TERMINAL_BATCH_INTERVAL = 16;
|
||||
|
||||
/** Immediate flush threshold for terminal batches (bytes).
|
||||
* Set high (32KB) to allow effective batching; avg Ink events are ~14KB. */
|
||||
export const BATCH_FLUSH_THRESHOLD = 32 * 1024;
|
||||
|
||||
/** Task event batching interval (ms) */
|
||||
export const TASK_UPDATE_BATCH_INTERVAL = 100;
|
||||
|
||||
/** SSE client health check interval (ms) */
|
||||
export const SSE_HEALTH_CHECK_INTERVAL = 30 * 1000;
|
||||
|
||||
// ============================================================================
|
||||
// State Persistence
|
||||
// ============================================================================
|
||||
|
||||
/** State update debounce — batches expensive toDetailedState() calls (ms) */
|
||||
export const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
|
||||
|
||||
/** Sessions list cache TTL — avoids re-serializing on every SSE init (ms) */
|
||||
export const SESSIONS_LIST_CACHE_TTL = 1000;
|
||||
|
||||
// ============================================================================
|
||||
// Scheduled Runs
|
||||
// ============================================================================
|
||||
|
||||
/** Scheduled runs cleanup check interval (ms) */
|
||||
export const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
|
||||
|
||||
/** Completed scheduled run max age before cleanup (ms) */
|
||||
export const SCHEDULED_RUN_MAX_AGE = 60 * 60 * 1000;
|
||||
|
||||
/** Session limit retry wait before retrying (ms) */
|
||||
export const SESSION_LIMIT_WAIT_MS = 5000;
|
||||
|
||||
/** Pause between scheduled run iterations (ms) */
|
||||
export const ITERATION_PAUSE_MS = 2000;
|
||||
|
||||
// ============================================================================
|
||||
// Mux Stats
|
||||
// ============================================================================
|
||||
|
||||
/** Mux stats collection interval (ms) */
|
||||
export const STATS_COLLECTION_INTERVAL_MS = 2000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npx tsx src/index.ts web --port 3099 &
|
||||
curl -s http://localhost:3099/api/status | jq .status # "ok"
|
||||
kill %1
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Create `src/config/auth-config.ts`
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files created**: `src/config/auth-config.ts`
|
||||
**Files modified**: `src/web/middleware/auth.ts`, `src/hooks-config.ts`
|
||||
|
||||
### Constants to extract from `src/web/middleware/auth.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `AUTH_SESSION_TTL_MS` | `86400000` | Auth session cookie TTL (24h) |
|
||||
| `MAX_AUTH_SESSIONS` | `100` | Max concurrent auth sessions |
|
||||
| `AUTH_FAILURE_MAX` | `10` | Max failed auth attempts per IP |
|
||||
| `AUTH_FAILURE_WINDOW_MS` | `900000` | Failed auth tracking window (15 min) |
|
||||
|
||||
### Constants to extract from `src/hooks-config.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `HOOK_TIMEOUT_MS` | `10000` | Timeout for Claude Code hook commands |
|
||||
|
||||
The `timeout: 10000` value is hardcoded 6 times in `hooks-config.ts` as inline literals. Extract to a single named constant.
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/auth-config.ts` with the 5 constants.
|
||||
2. In `src/web/middleware/auth.ts`: Remove the 4 local constant declarations (lines 17–25). Add import from `../../config/auth-config.js`. Keep `AUTH_COOKIE_NAME` in place — it's a string identifier, not a tunable numeric constant.
|
||||
3. In `src/hooks-config.ts`: Replace all 6 inline `timeout: 10000` occurrences with `timeout: HOOK_TIMEOUT_MS`. Add import from `./config/auth-config.js`.
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Authentication, rate limiting, and hook security constants.
|
||||
*
|
||||
* Controls auth session lifecycle, brute-force protection,
|
||||
* and Claude Code hook timeouts.
|
||||
*
|
||||
* @module config/auth-config
|
||||
*/
|
||||
|
||||
// ============================================================================
|
||||
// Session Cookies
|
||||
// ============================================================================
|
||||
|
||||
/** Auth session cookie TTL — matches autonomous run length (ms) */
|
||||
export const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
|
||||
|
||||
/** Max concurrent auth sessions per server */
|
||||
export const MAX_AUTH_SESSIONS = 100;
|
||||
|
||||
// ============================================================================
|
||||
// Rate Limiting
|
||||
// ============================================================================
|
||||
|
||||
/** Max failed auth attempts per IP before 429 rejection */
|
||||
export const AUTH_FAILURE_MAX = 10;
|
||||
|
||||
/** Failed auth attempt tracking window (ms) */
|
||||
export const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
|
||||
|
||||
// ============================================================================
|
||||
// Hooks
|
||||
// ============================================================================
|
||||
|
||||
/** Timeout for Claude Code hook curl commands (ms) */
|
||||
export const HOOK_TIMEOUT_MS = 10000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Create `src/config/tunnel-config.ts`
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files created**: `src/config/tunnel-config.ts`
|
||||
**Files modified**: `src/tunnel-manager.ts`
|
||||
|
||||
### Constants to extract from `src/tunnel-manager.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `QR_TOKEN_TTL_MS` | `60000` | QR token auto-rotation interval (60s) |
|
||||
| `QR_TOKEN_GRACE_MS` | `90000` | Grace period for previous token (90s) |
|
||||
| `SHORT_CODE_LENGTH` | `6` | Length of QR short code |
|
||||
| `QR_RATE_LIMIT_MAX` | `30` | Global QR attempt rate limit |
|
||||
| `QR_RATE_LIMIT_WINDOW_MS` | `60000` | QR rate limit reset window (60s) |
|
||||
| `URL_TIMEOUT_MS` | `30000` | Cloudflared URL fetch timeout (30s) |
|
||||
| `RESTART_DELAY_MS` | `5000` | Tunnel restart delay after crash (5s) |
|
||||
| `FORCE_KILL_MS` | `5000` | SIGTERM → SIGKILL escalation timeout (5s) |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/tunnel-config.ts` with all 8 constants.
|
||||
2. In `src/tunnel-manager.ts`: Remove the 8 local constant declarations (lines ~39–75). Add `import { QR_TOKEN_TTL_MS, ... } from './config/tunnel-config.js'`.
|
||||
3. Keep the `TUNNEL_URL_REGEX` in `tunnel-manager.ts` — it's a parsing detail, not a tuning knob.
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Cloudflare tunnel and QR authentication constants.
|
||||
*
|
||||
* Controls QR token rotation timing, rate limiting,
|
||||
* and tunnel process lifecycle.
|
||||
*
|
||||
* @module config/tunnel-config
|
||||
*/
|
||||
|
||||
// ============================================================================
|
||||
// QR Token Rotation
|
||||
// ============================================================================
|
||||
|
||||
/** QR token auto-rotation interval (ms) */
|
||||
export const QR_TOKEN_TTL_MS = 60_000;
|
||||
|
||||
/** Grace period — previous token still valid during rotation (ms) */
|
||||
export const QR_TOKEN_GRACE_MS = 90_000;
|
||||
|
||||
/** Length of the short code in QR URL path (chars) */
|
||||
export const SHORT_CODE_LENGTH = 6;
|
||||
|
||||
// ============================================================================
|
||||
// QR Rate Limiting
|
||||
// ============================================================================
|
||||
|
||||
/** Global rate limit for QR auth attempts across all IPs */
|
||||
export const QR_RATE_LIMIT_MAX = 30;
|
||||
|
||||
/** QR rate limit reset window (ms) */
|
||||
export const QR_RATE_LIMIT_WINDOW_MS = 60_000;
|
||||
|
||||
// ============================================================================
|
||||
// Tunnel Process Lifecycle
|
||||
// ============================================================================
|
||||
|
||||
/** Max time to wait for cloudflared URL before timeout (ms) */
|
||||
export const URL_TIMEOUT_MS = 30_000;
|
||||
|
||||
/** Restart delay after unexpected tunnel exit (ms) */
|
||||
export const RESTART_DELAY_MS = 5_000;
|
||||
|
||||
/** SIGTERM → SIGKILL escalation timeout (ms) */
|
||||
export const FORCE_KILL_MS = 5_000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Create `src/config/terminal-limits.ts`
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files created**: `src/config/terminal-limits.ts`
|
||||
**Files modified**: `src/web/routes/session-routes.ts`
|
||||
|
||||
### Constants to extract from `src/web/routes/session-routes.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `MAX_INPUT_LENGTH` | `65536` | Max input length per request (64KB) |
|
||||
| `MAX_TERMINAL_COLS` | `500` | Max terminal columns |
|
||||
| `MAX_TERMINAL_ROWS` | `200` | Max terminal rows |
|
||||
| `MAX_SESSION_NAME_LENGTH` | `128` | Max session name length (chars) |
|
||||
|
||||
### Why a separate file instead of adding to `buffer-limits.ts`
|
||||
|
||||
`buffer-limits.ts` covers memory buffer sizes (2MB terminal, 1MB text). These constants are **validation limits** for API inputs — different concern. A terminal resize request must not exceed `MAX_TERMINAL_COLS`; this has nothing to do with buffer trimming.
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/terminal-limits.ts` with all 4 constants.
|
||||
2. In `src/web/routes/session-routes.ts`: Remove the 4 local constant declarations (lines 45–48). Add `import { MAX_INPUT_LENGTH, MAX_TERMINAL_COLS, MAX_TERMINAL_ROWS, MAX_SESSION_NAME_LENGTH } from '../../config/terminal-limits.js'`.
|
||||
3. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Terminal dimension and input validation limits.
|
||||
*
|
||||
* Used by API routes to validate resize, input, and session
|
||||
* creation requests. Separate from buffer-limits.ts which
|
||||
* controls memory buffer sizes.
|
||||
*
|
||||
* @module config/terminal-limits
|
||||
*/
|
||||
|
||||
/** Max input length per API request (bytes) */
|
||||
export const MAX_INPUT_LENGTH = 64 * 1024;
|
||||
|
||||
/** Max terminal columns for resize requests */
|
||||
export const MAX_TERMINAL_COLS = 500;
|
||||
|
||||
/** Max terminal rows for resize requests */
|
||||
export const MAX_TERMINAL_ROWS = 200;
|
||||
|
||||
/** Max session name length (chars) */
|
||||
export const MAX_SESSION_NAME_LENGTH = 128;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Create `src/config/ai-defaults.ts`
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Files created**: `src/config/ai-defaults.ts`
|
||||
**Files modified**: `src/respawn-controller.ts`, `src/ai-idle-checker.ts`, `src/ai-plan-checker.ts`, `src/web/routes/respawn-routes.ts`
|
||||
|
||||
### Problem: AI model string duplicated 5 times
|
||||
|
||||
The model identifier `'claude-opus-4-5-20251101'` appears in 5 places across 4 files. When the model changes, all 5 must be updated — a guaranteed source of bugs. The context limits (`16000`, `8000`) are similarly scattered across 3 files each.
|
||||
|
||||
| Constant | Current Value | Duplicated In |
|
||||
|----------|---------------|---------------|
|
||||
| `AI_CHECK_MODEL` | `'claude-opus-4-5-20251101'` | `respawn-controller.ts` (×2: idle + plan), `ai-idle-checker.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` (×2: idle + plan) |
|
||||
| `AI_IDLE_CHECK_MAX_CONTEXT` | `16000` | `respawn-controller.ts`, `ai-idle-checker.ts`, `respawn-routes.ts` |
|
||||
| `AI_PLAN_CHECK_MAX_CONTEXT` | `8000` | `respawn-controller.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/ai-defaults.ts` with the 3 constants.
|
||||
2. In `src/respawn-controller.ts` `DEFAULT_CONFIG` (line 538): Replace `aiIdleCheckModel: 'claude-opus-4-5-20251101'` with `aiIdleCheckModel: AI_CHECK_MODEL`, `aiIdleCheckMaxContext: 16000` with `aiIdleCheckMaxContext: AI_IDLE_CHECK_MAX_CONTEXT`, `aiPlanCheckModel: 'claude-opus-4-5-20251101'` with `aiPlanCheckModel: AI_CHECK_MODEL`, `aiPlanCheckMaxContext: 8000` with `aiPlanCheckMaxContext: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
|
||||
3. In `src/ai-idle-checker.ts` `DEFAULT_AI_CHECK_CONFIG` (line 46): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 16000` with `maxContextChars: AI_IDLE_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
|
||||
4. In `src/ai-plan-checker.ts` `DEFAULT_PLAN_CHECK_CONFIG` (line 45): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 8000` with `maxContextChars: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
|
||||
5. In `src/web/routes/respawn-routes.ts` config merge block (lines 173–179): Replace all 4 inline fallback values with imports from `../../config/ai-defaults.js`.
|
||||
6. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Default model and context limits for AI-powered checkers.
|
||||
*
|
||||
* Centralizes the AI model identifier and context window sizes used by
|
||||
* the idle checker, plan checker, respawn controller defaults, and
|
||||
* respawn route fallbacks. Change the model here when upgrading.
|
||||
*
|
||||
* @module config/ai-defaults
|
||||
*/
|
||||
|
||||
/** Default model for AI idle and plan checkers */
|
||||
export const AI_CHECK_MODEL = 'claude-opus-4-5-20251101';
|
||||
|
||||
/** Max context chars for idle checker (~4k tokens) */
|
||||
export const AI_IDLE_CHECK_MAX_CONTEXT = 16000;
|
||||
|
||||
/** Max context chars for plan checker (~2k tokens, plan mode UI is compact) */
|
||||
export const AI_PLAN_CHECK_MAX_CONTEXT = 8000;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
# Verify no remaining hardcoded model strings
|
||||
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 6: Create `src/config/team-config.ts`
|
||||
|
||||
**Estimated effort**: 15 minutes
|
||||
**Files created**: `src/config/team-config.ts`
|
||||
**Files modified**: `src/team-watcher.ts`
|
||||
|
||||
### Constants to extract from `src/team-watcher.ts`
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `TEAM_POLL_INTERVAL_MS` | `30000` | Team directory poll interval (30s) |
|
||||
| `MAX_CACHED_TEAMS` | `50` | LRU cache size for team configs |
|
||||
| `MAX_CACHED_TASKS` | `200` | LRU cache size for team tasks + inboxes |
|
||||
|
||||
### Why centralize these
|
||||
|
||||
Team polling frequency and cache sizes are operational knobs that affect both performance (polling too often wastes CPU) and responsiveness (polling too rarely means stale team state in the UI). They're also the kind of values a developer tuning for a large team deployment would want to find quickly. `MAX_CACHED_TASKS` is used for both the task cache and inbox cache — worth documenting.
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `src/config/team-config.ts` with the 3 constants.
|
||||
2. In `src/team-watcher.ts`: Remove the 3 local constants (lines 23–25). Add `import { TEAM_POLL_INTERVAL_MS, MAX_CACHED_TEAMS, MAX_CACHED_TASKS } from './config/team-config.js'`. Note: rename `POLL_INTERVAL_MS` → `TEAM_POLL_INTERVAL_MS` to avoid ambiguity with the identically-named constant in `subagent-watcher.ts`.
|
||||
3. Update the usage site: `setInterval(... POLL_INTERVAL_MS)` → `setInterval(... TEAM_POLL_INTERVAL_MS)`.
|
||||
4. Run `tsc --noEmit`.
|
||||
|
||||
### New file template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* @fileoverview Agent Teams polling and cache configuration.
|
||||
*
|
||||
* Controls how frequently TeamWatcher polls ~/.claude/teams/
|
||||
* and how many teams/tasks are cached in memory.
|
||||
*
|
||||
* @module config/team-config
|
||||
*/
|
||||
|
||||
/** Team directory poll interval (ms) */
|
||||
export const TEAM_POLL_INTERVAL_MS = 30_000;
|
||||
|
||||
/** Max cached team configs (LRU eviction) */
|
||||
export const MAX_CACHED_TEAMS = 50;
|
||||
|
||||
/** Max cached team tasks and inbox messages (LRU eviction).
|
||||
* Used for both teamTasks and inboxCache maps. */
|
||||
export const MAX_CACHED_TASKS = 200;
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 7: Fix remaining cross-file duplicates
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Files modified**: `src/index.ts`, `src/subagent-watcher.ts`
|
||||
|
||||
### Duplicate 1: `STATS_COLLECTION_INTERVAL_MS`
|
||||
|
||||
Already fixed in Task 1 — both `server.ts` and `mux-routes.ts` now import from `server-timing.ts`.
|
||||
|
||||
### Duplicate 2: AI model string
|
||||
|
||||
Already fixed in Task 5 — all 5 occurrences now import from `ai-defaults.ts`.
|
||||
|
||||
### Duplicate 3: `MAX_SCREENSHOT_SIZE` / `MAX_TEXT_FILE_SIZE` / `MAX_RAW_FILE_SIZE`
|
||||
|
||||
These file size limits in `file-routes.ts` and `system-routes.ts` are **API-specific validation limits**. They're only used in their respective route files and aren't duplicated. **Leave in place** — they're local to their route module and well-commented.
|
||||
|
||||
### Action A: Move `MAX_CONSECUTIVE_ERRORS` and `ERROR_RESET_MS` to config
|
||||
|
||||
`src/index.ts` has two process-level constants that are operational tuning knobs:
|
||||
|
||||
| Constant | Value | Purpose |
|
||||
|----------|-------|---------|
|
||||
| `MAX_CONSECUTIVE_ERRORS` | `5` | Max consecutive unhandled errors before process exit |
|
||||
| `ERROR_RESET_MS` | `60000` | Error counter reset interval (1 min) |
|
||||
|
||||
These belong in a config file since they control server reliability behavior. Add them to `src/config/server-timing.ts` (they're server operational constants).
|
||||
|
||||
1. Add to `src/config/server-timing.ts`:
|
||||
```typescript
|
||||
// ============================================================================
|
||||
// Process Error Recovery
|
||||
// ============================================================================
|
||||
|
||||
/** Max consecutive unhandled errors before auto-restart */
|
||||
export const MAX_CONSECUTIVE_ERRORS = 5;
|
||||
|
||||
/** Error counter reset interval — forgives errors after quiet period (ms) */
|
||||
export const ERROR_RESET_MS = 60_000;
|
||||
```
|
||||
2. In `src/index.ts`: Remove lines 19–20, add import from `./config/server-timing.js`.
|
||||
3. Run `tsc --noEmit`.
|
||||
|
||||
### Action B: Fix `MAX_TRACKED_AGENTS` shadow in `subagent-watcher.ts`
|
||||
|
||||
`subagent-watcher.ts` defines its own `MAX_TRACKED_AGENTS = 500` locally instead of importing the identical value from `config/map-limits.ts`. This is a latent bug — if someone changes the config value, the subagent watcher's copy stays stale.
|
||||
|
||||
1. In `src/subagent-watcher.ts`: Remove the local `MAX_TRACKED_AGENTS` constant. Add `import { MAX_TRACKED_AGENTS } from './config/map-limits.js'` (the value there is `MAX_TODOS_PER_SESSION = 500` — **verify** the map-limits constant is actually named `MAX_TRACKED_AGENTS` or if it needs to be added). If the constant doesn't exist in `map-limits.ts` under that name, add it.
|
||||
2. Run `tsc --noEmit`.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npm run lint
|
||||
npm run format:check
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 8: Update CLAUDE.md and final verification
|
||||
|
||||
**Estimated effort**: 20 minutes
|
||||
**Files modified**: `CLAUDE.md`
|
||||
|
||||
### Updates to CLAUDE.md
|
||||
|
||||
1. **Config Files table** (`src/config/`): Add the 6 new files:
|
||||
|
||||
| File | Purpose |
|
||||
|------|---------|
|
||||
| `buffer-limits.ts` | Terminal/text buffer size limits |
|
||||
| `map-limits.ts` | Global limits for Maps, sessions, watchers |
|
||||
| `exec-timeout.ts` | Execution timeout configuration |
|
||||
| `server-timing.ts` | Web server batching, SSE, scheduled run timing |
|
||||
| `auth-config.ts` | Auth session TTL, rate limits, hook timeout |
|
||||
| `tunnel-config.ts` | QR token rotation, tunnel process lifecycle |
|
||||
| `terminal-limits.ts` | Terminal dimension and input validation limits |
|
||||
| `ai-defaults.ts` | AI checker model and context limits |
|
||||
| `team-config.ts` | Agent Teams polling and cache sizes |
|
||||
|
||||
2. **Import Conventions** section: Add:
|
||||
```
|
||||
- **Config**: Import from specific files: `import { MAX_TERMINAL_COLS } from './config/terminal-limits'`
|
||||
```
|
||||
|
||||
3. **Phase 6 status** in `docs/code-structure-findings.md`: Mark as COMPLETE with summary of what was done.
|
||||
|
||||
### Final verification checklist
|
||||
|
||||
```bash
|
||||
# Type checking
|
||||
tsc --noEmit
|
||||
|
||||
# Linting
|
||||
npm run lint
|
||||
|
||||
# Formatting
|
||||
npm run format:check
|
||||
|
||||
# Dev server starts
|
||||
npx tsx src/index.ts web --port 3099 &
|
||||
curl -s http://localhost:3099/api/status | jq .status # "ok"
|
||||
kill %1
|
||||
|
||||
# Verify no remaining duplicates
|
||||
grep -rn 'STATS_COLLECTION_INTERVAL_MS' src/ # Should only appear in config + import sites
|
||||
grep -rn 'timeout: 10000' src/hooks-config.ts # Should be 0 — all replaced with HOOK_TIMEOUT_MS
|
||||
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## What is NOT in scope (and why)
|
||||
|
||||
These constants were considered but deliberately left in their current files:
|
||||
|
||||
### Respawn controller defaults (`src/respawn-controller.ts`)
|
||||
|
||||
The `DEFAULT_CONFIG` object (lines 538–578) contains ~30 default values for the `RespawnConfig` interface. These are **user-facing configuration defaults**, not system constants — they're the starting values for a config object that users can modify via the API and UI. Centralizing them would break the locality between the config interface definition and its defaults. They already have excellent JSDoc with `@default` tags. The only values extracted are the AI model/context constants (Task 5) which are duplicated in other files.
|
||||
|
||||
### Subagent watcher timing (`src/subagent-watcher.ts`)
|
||||
|
||||
The 18 constants at lines 129–158 are all internal to the subagent watcher's polling/lifecycle algorithm. Moving them to a config file would force developers to context-switch between two files to understand the polling logic. They're already grouped with clear comments. Exception: `MAX_TRACKED_AGENTS` is consolidated with `map-limits.ts` (Task 7B) since it duplicates a global limit.
|
||||
|
||||
### Session auto-ops timing (`src/session-auto-ops.ts`)
|
||||
|
||||
The 8 constants at lines 19–40 are internal to the auto-compact/clear retry state machine. They form a coherent group that's meaningless without the surrounding implementation context.
|
||||
|
||||
### Run summary constants (`src/run-summary.ts`)
|
||||
|
||||
`MAX_EVENTS`, `TRIM_TO_EVENTS`, `TOKEN_MILESTONE_INTERVAL`, `STATE_STUCK_WARNING_MS`, `STATE_STUCK_CHECK_INTERVAL` — all module-internal. The buffer-style limits (`MAX_EVENTS`/`TRIM_TO_EVENTS`) follow the same pattern as `buffer-limits.ts` but are only used in this one file.
|
||||
|
||||
### Frontend (`src/web/public/constants.js`)
|
||||
|
||||
Already well-centralized. Frontend and backend run in different environments — mixing them in TypeScript config files would create import problems. If frontend constants need expansion, do it in `constants.js`. Note: `app.js` has 2 inline uses of `256 * 1024` that should use the existing `TERMINAL_TAIL_SIZE` from `constants.js` — a minor cleanup that can be done opportunistically but is not worth a task here.
|
||||
|
||||
### Tmux manager timing (`src/tmux-manager.ts`)
|
||||
|
||||
The 6 constants (lines 65–78) are internal to tmux process lifecycle management. They're low-level retry/wait values that are meaningless without understanding the tmux spawn sequence.
|
||||
|
||||
### Process-internal constants
|
||||
|
||||
`image-watcher.ts`, `bash-tool-parser.ts`, `transcript-watcher.ts`, `ralph-tracker.ts`, `task-tracker.ts`, `file-stream-manager.ts`, `session-lifecycle-log.ts`, `session-task-cache.ts`, `respawn-metrics.ts`, `respawn-adaptive-timing.ts`, `ai-checker-base.ts` — all have module-local constants that are internal implementation details.
|
||||
|
||||
### `localhost:3000` default URL
|
||||
|
||||
The string `'http://localhost:3000'` or port `3000` appears as a fallback default in ~5 files (`session-cli-builder.ts`, `tmux-manager.ts`, `tunnel-manager.ts`, `server.ts`, CLI). While technically duplicated, extracting it provides little value — each usage has a different fallback chain (env var → config → hardcoded) and the port is also baked into systemd service files and documentation. The risk of a missed update is low since port 3000 is deeply conventional.
|
||||
|
||||
### `SAVE_DEBOUNCE_MS = 500` in `state-store.ts` / `push-store.ts`
|
||||
|
||||
Same value (500ms), but they debounce different persistence targets (state.json vs push-subscriptions.json). If one needed faster/slower debouncing, they'd diverge. Coupling them would be misleading.
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
| Metric | Before | After |
|
||||
|--------|--------|-------|
|
||||
| Config files in `src/config/` | 3 | 9 |
|
||||
| Constants centralized | ~25 | ~65 |
|
||||
| Cross-file duplicates | 9+ (`STATS_COLLECTION_INTERVAL_MS`, `timeout: 10000` ×6, AI model ×5, context limits ×3 each, `MAX_TRACKED_AGENTS`) | 0 |
|
||||
| Files with `timeout: 10000` inline | 1 (6 occurrences) | 0 |
|
||||
| Files with hardcoded AI model string | 4 (5 occurrences) | 1 (config only) |
|
||||
| Files modified | — | 11 |
|
||||
| Files created | — | 6 |
|
||||
@@ -0,0 +1,953 @@
|
||||
# Phase 7 Implementation Plan: Test Infrastructure
|
||||
|
||||
**Source**: `docs/code-structure-findings.md` (Phase 7 — Test Infrastructure)
|
||||
**Estimated effort**: 2–3 days
|
||||
**Tasks**: 11 tasks with dependencies (see dependency graph below)
|
||||
|
||||
---
|
||||
|
||||
## Safety Constraints
|
||||
|
||||
Before starting ANY work, read and follow these rules:
|
||||
|
||||
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
|
||||
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
|
||||
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
|
||||
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
|
||||
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
|
||||
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
|
||||
7. **Port assignments for this phase**: New tests use ports 3220–3229 (see individual tasks for assignments).
|
||||
|
||||
---
|
||||
|
||||
## Goal
|
||||
|
||||
Eliminate duplicated test mocks, activate the unused `respawn-test-utils.ts` utilities, and add route-level test coverage for the server's 12 route modules — the single largest untested area in the codebase (162 route handlers, 0 dedicated tests).
|
||||
|
||||
**Non-goals**:
|
||||
- Full end-to-end integration tests (those require real Claude CLI / tmux sessions)
|
||||
- 100% route coverage in this phase — focus on the highest-value route modules first
|
||||
- Refactoring test patterns in existing passing tests that don't use shared mocks
|
||||
- Migrating `vi.mock()`-based module replacement mocks (different pattern, see Task 6/7)
|
||||
|
||||
---
|
||||
|
||||
## Current State
|
||||
|
||||
### Mock Duplication (Finding #9)
|
||||
|
||||
`MockSession` is defined **4 times** across test files with varying levels of completeness:
|
||||
|
||||
| File | Properties | Methods | EventEmitter | Notes |
|
||||
|------|-----------|---------|-------------|-------|
|
||||
| `test/respawn-test-utils.ts` | 6 | 20+ | Yes | **Most complete**. Includes terminal simulation, token count, ANSI output, plan mode prompts. **Never imported by any test.** |
|
||||
| `test/respawn-controller.test.ts` | 6 | 9 | Yes | Subset of respawn-test-utils. Missing token simulation, ANSI helpers. |
|
||||
| `test/respawn-team-awareness.test.ts` | ~6 | ~9 | Yes | Near-copy of respawn-controller.test.ts version. |
|
||||
| `test/session-manager.test.ts` | 4 | 8 | Yes | **Inside `vi.mock()` factory** — replaces `../src/session.js` module. Different shape: `start()`/`stop()`/`toState()`/`sendInput()` for lifecycle testing. |
|
||||
|
||||
`MockStateStore` is defined **2 times** (both inside `vi.mock()` factories):
|
||||
|
||||
| File | Shape | Methods | Mock Pattern |
|
||||
|------|-------|---------|-------------|
|
||||
| `test/session-manager.test.ts` | `{ sessions, config }` | `getConfig`, `getSessions`, `getSession`, `setSession`, `removeSession` | `vi.mock('../src/state-store.js')` |
|
||||
| `test/ralph-loop.test.ts` | `{ ralphLoop, tasks, config }` | `getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask` | `vi.mock('../src/state-store.js')` |
|
||||
|
||||
### Important: Two distinct mocking patterns
|
||||
|
||||
The codebase uses two different mocking patterns that require different migration strategies:
|
||||
|
||||
1. **Direct instantiation** (respawn-controller, respawn-team-awareness): `MockSession` is defined at file scope and instantiated directly in tests. These can be migrated to shared mocks via simple import replacement.
|
||||
|
||||
2. **Module replacement** (session-manager, ralph-loop): Mocks are defined inside `vi.mock()` factories that replace entire modules (`../src/session.js`, `../src/state-store.js`). These factories run in an isolated scope and return `{ Session: MockClass }` or `{ getStore: vi.fn(() => instance) }`. Migrating these requires either `vi.hoisted()` or restructuring the test's module mocking — higher risk for limited benefit.
|
||||
|
||||
### Unused Test Utilities
|
||||
|
||||
`test/respawn-test-utils.ts` exports these utilities that **no test file imports**:
|
||||
|
||||
- `TimeController` / `createTimeController()` — abstraction over vitest fake timers
|
||||
- `MockAiIdleChecker` / `MockAiPlanChecker` — fully mocked AI checkers with result queueing
|
||||
- `createStateTracker()` / `createEventRecorder()` — state transition and event recording
|
||||
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — pre-configured RespawnConfig objects
|
||||
- `waitForState()` / `waitForEvent()` / `createDeferred()` — async test helpers
|
||||
- `terminalOutputs` — factory object for common terminal output patterns
|
||||
|
||||
### Route Test Coverage
|
||||
|
||||
Currently **zero** dedicated tests for the 12 route modules in `src/web/routes/`. The existing test files that touch API endpoints:
|
||||
|
||||
| Test File | What It Tests | Approach |
|
||||
|-----------|--------------|----------|
|
||||
| `test/api-responses.test.ts` | Response structure validation | Imports types, no HTTP calls |
|
||||
| `test/api-generate-plan.test.ts` | Plan generation API | Mocks validation logic, Port 3191 declared |
|
||||
| `test/auth-security.test.ts` | Auth middleware | Integration tests with WebServer, Ports 3160/3161 |
|
||||
| `test/qr-auth.test.ts` | QR authentication | Integration + unit tests, Port 3162 |
|
||||
|
||||
None of these test the route handlers themselves with real HTTP requests against a running Fastify instance.
|
||||
|
||||
---
|
||||
|
||||
## Design Decisions
|
||||
|
||||
### Shared mocks: Superset strategy
|
||||
|
||||
Rather than creating a lowest-common-denominator mock, `MockSession` in `test/mocks/` will be the **superset** from `respawn-test-utils.ts` (the most complete version). Test files that need a simpler mock can just ignore the extra methods — having unused methods costs nothing, but missing methods forces local re-definition.
|
||||
|
||||
### vi.mock() tests: Don't migrate
|
||||
|
||||
The `session-manager.test.ts` and `ralph-loop.test.ts` tests define mocks inside `vi.mock()` factories. These use **module-level replacement** (replacing `../src/session.js` and `../src/state-store.js` entirely), which is fundamentally different from the direct-instantiation pattern. Migrating them would require `vi.hoisted()` or factory restructuring — high complexity for limited benefit since these mocks are already working. We leave these as-is and create the shared mocks for **new** tests and for the two direct-instantiation tests (Tasks 4–5).
|
||||
|
||||
### MockStateStore: Union of both shapes
|
||||
|
||||
The shared `MockStateStore` in `test/mocks/` will include methods from both existing definitions (session management + Ralph loop), so any **new** test can use it. Methods default to no-ops via `vi.fn()`. Existing `vi.mock()`-based tests are not migrated.
|
||||
|
||||
### Route testing strategy: Lightweight Fastify instances
|
||||
|
||||
Each route test file will:
|
||||
1. Create a minimal `Fastify` instance
|
||||
2. Register **only** the route module under test
|
||||
3. Provide a mock context object satisfying the port interfaces
|
||||
4. Use `app.inject()` (Fastify's built-in test helper) — no real HTTP, no port needed
|
||||
|
||||
This avoids port conflicts entirely and runs fast. Only tests that need SSE or WebSocket behavior will use a real listening server with assigned ports.
|
||||
|
||||
### Port assignments (for tests needing real servers)
|
||||
|
||||
| Port | Test File | Purpose |
|
||||
|------|-----------|---------|
|
||||
| 3220 | `test/routes/session-routes.test.ts` | SSE integration (if needed) |
|
||||
| 3221 | `test/routes/system-routes.test.ts` | Status/stats endpoints |
|
||||
| 3222 | `test/routes/respawn-routes.test.ts` | Respawn API |
|
||||
| 3223 | `test/routes/ralph-routes.test.ts` | Ralph API |
|
||||
| 3224–3229 | Reserved | Future route tests |
|
||||
|
||||
Most tests should NOT need real ports — `app.inject()` is preferred. Verified: ports 3220–3229 are completely unused by existing tests (highest used port is 3211 in `opencode-resize.test.ts`).
|
||||
|
||||
---
|
||||
|
||||
## Task Dependencies
|
||||
|
||||
```
|
||||
Task 1 (Consolidate MockSession)
|
||||
Task 2 (Consolidate MockStateStore)
|
||||
└──> Task 3 (Create test/mocks/ barrel)
|
||||
├──> Task 4 (Migrate respawn-controller.test.ts)
|
||||
├──> Task 5 (Migrate respawn-team-awareness.test.ts)
|
||||
└──> Task 6 (Route test scaffold + helpers)
|
||||
├──> Task 7 (Session routes tests)
|
||||
└──> Task 8 (System + respawn routes tests)
|
||||
|
||||
Task 9 (Slim down respawn-test-utils.ts) — depends on Tasks 4, 5
|
||||
```
|
||||
|
||||
**Tasks 1–2** are independent and can run in parallel.
|
||||
**Task 3** depends on Tasks 1–2.
|
||||
**Tasks 4–6** depend on Task 3 and can run in parallel.
|
||||
**Tasks 7–8** depend on Task 6 and can run in parallel.
|
||||
**Task 9** depends on Tasks 4, 5 (must verify migrations work before removing duplicates from source).
|
||||
|
||||
---
|
||||
|
||||
## Task 1: Consolidate MockSession into `test/mocks/mock-session.ts`
|
||||
|
||||
**Estimated effort**: 2 hours
|
||||
**Files created**: `test/mocks/mock-session.ts`
|
||||
**Files modified**: None yet (consumers migrate in Tasks 4–5)
|
||||
|
||||
### Source
|
||||
|
||||
The canonical MockSession comes from `test/respawn-test-utils.ts` (lines 89–241). It is the most complete version with:
|
||||
|
||||
- All properties needed by `RespawnController`: `id`, `workingDir`, `status`, `writeBuffer`, `terminalBuffer`, `muxName`
|
||||
- `write()` / `writeViaMux()` for input simulation
|
||||
- Buffer inspection: `lastWrite`, `hasWritten(pattern)`, `clearWriteBuffer()`
|
||||
- Terminal simulation: `simulateTerminalOutput()`, `simulatePrompt()`, `simulateReady()`, `simulateCompletionMessage()`, `simulateWorking()`, `simulateClearComplete()`, `simulateInitComplete()`, `simulatePlanModePrompt()`, `simulateElicitationDialog()`, `simulateTokenCount()`, `simulateAnsiOutput()`
|
||||
- Lifecycle: `close()`
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `test/mocks/` directory.
|
||||
2. Create `test/mocks/mock-session.ts`:
|
||||
- Copy the `MockSession` class **exactly** from `test/respawn-test-utils.ts` (lines 89–241)
|
||||
- Copy `terminalOutputs` helper object (tightly coupled to mock)
|
||||
- Copy `createMockSession()` factory function
|
||||
- Export all three: `export { MockSession, createMockSession, terminalOutputs }`
|
||||
- Ensure all `vi` imports come from `vitest`
|
||||
|
||||
**CRITICAL**: Copy the source verbatim — do NOT rewrite the simulation methods. The respawn controller's detection logic matches specific output patterns (e.g., `'\u276f '` for prompt, `'\u273b Worked for'` for completion). Using different patterns would cause test failures.
|
||||
|
||||
### Template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared MockSession for tests that need terminal simulation.
|
||||
*
|
||||
* Copied from test/respawn-test-utils.ts (the canonical, most complete version).
|
||||
* Used by respawn, route, and subagent tests.
|
||||
*/
|
||||
import { EventEmitter } from 'node:events';
|
||||
|
||||
// Copy MockSession class exactly from test/respawn-test-utils.ts lines 89–241
|
||||
export class MockSession extends EventEmitter {
|
||||
// ... (copy verbatim from respawn-test-utils.ts)
|
||||
}
|
||||
|
||||
/**
|
||||
* Factory for common terminal output strings.
|
||||
* Must match the patterns used in MockSession's simulate* methods.
|
||||
*/
|
||||
export const terminalOutputs = {
|
||||
// ... (copy verbatim from respawn-test-utils.ts)
|
||||
};
|
||||
|
||||
/**
|
||||
* Convenience factory.
|
||||
*/
|
||||
export function createMockSession(id?: string): MockSession {
|
||||
return new MockSession(id);
|
||||
}
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit # Ensure file compiles
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 2: Consolidate MockStateStore into `test/mocks/mock-state-store.ts`
|
||||
|
||||
**Estimated effort**: 1 hour
|
||||
**Files created**: `test/mocks/mock-state-store.ts`
|
||||
**Files modified**: None (existing vi.mock()-based tests are NOT migrated; this is for new route tests)
|
||||
|
||||
### Source
|
||||
|
||||
Union of both existing definitions:
|
||||
|
||||
- From `test/session-manager.test.ts`: session CRUD methods (`getConfig`, `getSession`, `setSession`, `removeSession`, `getSessions`)
|
||||
- From `test/ralph-loop.test.ts`: Ralph state methods (`getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask`)
|
||||
|
||||
### Template
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared MockStateStore for tests.
|
||||
*
|
||||
* Includes methods for both session management and Ralph loop testing.
|
||||
* All methods are vi.fn() spies — tests can override return values as needed.
|
||||
*
|
||||
* NOTE: This is for direct instantiation in new tests. Existing tests that
|
||||
* use vi.mock('../src/state-store.js') keep their inline definitions.
|
||||
*/
|
||||
import { vi } from 'vitest';
|
||||
|
||||
export class MockStateStore {
|
||||
state: Record<string, unknown> = {
|
||||
sessions: {} as Record<string, unknown>,
|
||||
config: { maxConcurrentSessions: 5 },
|
||||
ralphLoop: { status: 'stopped' },
|
||||
tasks: {} as Record<string, unknown>,
|
||||
};
|
||||
|
||||
// Session methods
|
||||
getConfig = vi.fn(() => this.state.config);
|
||||
getSessions = vi.fn(() => this.state.sessions as Record<string, unknown>);
|
||||
getSession = vi.fn((id: string) => (this.state.sessions as Record<string, unknown>)[id]);
|
||||
setSession = vi.fn((id: string, state: unknown) => {
|
||||
(this.state.sessions as Record<string, unknown>)[id] = state;
|
||||
});
|
||||
removeSession = vi.fn((id: string) => {
|
||||
delete (this.state.sessions as Record<string, unknown>)[id];
|
||||
});
|
||||
|
||||
// Ralph state methods
|
||||
getRalphLoopState = vi.fn(() => this.state.ralphLoop);
|
||||
setRalphLoopState = vi.fn((update: Record<string, unknown>) => {
|
||||
this.state.ralphLoop = { ...(this.state.ralphLoop as Record<string, unknown>), ...update };
|
||||
});
|
||||
|
||||
// Task methods
|
||||
getTasks = vi.fn(() => this.state.tasks);
|
||||
setTask = vi.fn();
|
||||
removeTask = vi.fn();
|
||||
|
||||
// Settings methods
|
||||
getSettings = vi.fn(() => ({}));
|
||||
setSettings = vi.fn();
|
||||
|
||||
// Generic persistence
|
||||
save = vi.fn();
|
||||
load = vi.fn();
|
||||
|
||||
/** Reset all state and mocks for clean test isolation */
|
||||
reset(): void {
|
||||
this.state = {
|
||||
sessions: {},
|
||||
config: { maxConcurrentSessions: 5 },
|
||||
ralphLoop: { status: 'stopped' },
|
||||
tasks: {},
|
||||
};
|
||||
vi.clearAllMocks();
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 3: Create `test/mocks/index.ts` barrel export
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Tasks 1, 2
|
||||
**Files created**: `test/mocks/index.ts`, `test/mocks/test-helpers.ts`
|
||||
**Files modified**: None
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `test/mocks/test-helpers.ts` with the async utilities from `respawn-test-utils.ts`:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Reusable async test helpers.
|
||||
* Extracted from respawn-test-utils.ts.
|
||||
*/
|
||||
|
||||
/** Wait for an EventEmitter to emit a specific event, with timeout */
|
||||
export function waitForEvent(
|
||||
emitter: { once: (event: string, listener: (...args: unknown[]) => void) => void },
|
||||
event: string,
|
||||
timeoutMs = 5000,
|
||||
): Promise<unknown> {
|
||||
return new Promise((resolve, reject) => {
|
||||
const timer = setTimeout(
|
||||
() => reject(new Error(`Timed out waiting for event "${event}" after ${timeoutMs}ms`)),
|
||||
timeoutMs,
|
||||
);
|
||||
emitter.once(event, (...args: unknown[]) => {
|
||||
clearTimeout(timer);
|
||||
resolve(args.length === 1 ? args[0] : args);
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
/** Create a deferred promise with external resolve/reject */
|
||||
export function createDeferred<T = void>(): {
|
||||
promise: Promise<T>;
|
||||
resolve: (value: T) => void;
|
||||
reject: (reason?: unknown) => void;
|
||||
} {
|
||||
let resolve!: (value: T) => void;
|
||||
let reject!: (reason?: unknown) => void;
|
||||
const promise = new Promise<T>((res, rej) => {
|
||||
resolve = res;
|
||||
reject = rej;
|
||||
});
|
||||
return { promise, resolve, reject };
|
||||
}
|
||||
```
|
||||
|
||||
2. Create `test/mocks/index.ts` barrel:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared test mocks — import from here instead of defining inline.
|
||||
*
|
||||
* @example
|
||||
* import { MockSession, MockStateStore, terminalOutputs } from './mocks/index.js';
|
||||
*/
|
||||
|
||||
export { MockSession, createMockSession, terminalOutputs } from './mock-session.js';
|
||||
export { MockStateStore } from './mock-state-store.js';
|
||||
export { waitForEvent, createDeferred } from './test-helpers.js';
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 4: Migrate `respawn-controller.test.ts` to shared mocks
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Task 3
|
||||
**Files modified**: `test/respawn-controller.test.ts`
|
||||
|
||||
### Steps
|
||||
|
||||
1. Remove the local `MockSession` class definition (approx. 50 lines).
|
||||
2. Add: `import { MockSession } from './mocks/index.js';`
|
||||
3. Verify all test methods still exist on the shared mock. The shared mock is a superset, so all existing usage should work.
|
||||
4. If the local mock had any test-specific customizations (e.g., extra properties added in `beforeEach`), keep those in the test file as inline assignments on the shared instance.
|
||||
5. Run the test to confirm it passes.
|
||||
|
||||
### Potential issues
|
||||
|
||||
- The local mock's `simulateCompletionMessage()` may have a slightly different output format than the shared mock's (from respawn-test-utils.ts). Verify the respawn controller's completion detection regex matches the shared mock's output pattern (`'\u273b Worked for ...'`).
|
||||
- If the local mock adds `pid` or `isWorking` properties that the shared mock doesn't have, add inline assignments in `beforeEach`.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/respawn-controller.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 5: Migrate `respawn-team-awareness.test.ts` to shared mocks
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Task 3
|
||||
**Files modified**: `test/respawn-team-awareness.test.ts`
|
||||
|
||||
### Steps
|
||||
|
||||
1. Remove the local `MockSession` class definition.
|
||||
2. Add: `import { MockSession } from './mocks/index.js';`
|
||||
3. Keep `MockTeamWatcher` in this file — it's test-specific and extends the real `TeamWatcher`, not a general-purpose mock.
|
||||
4. Run the test to confirm it passes.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/respawn-team-awareness.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 6: Create route test scaffold and helpers
|
||||
|
||||
**Estimated effort**: 2 hours
|
||||
**Depends on**: Task 3
|
||||
**Files created**: `test/mocks/mock-route-context.ts`, `test/routes/` directory, `test/routes/_route-test-utils.ts`
|
||||
|
||||
### Problem
|
||||
|
||||
The 12 route modules in `src/web/routes/` have zero dedicated test coverage. Each route module takes `(app: FastifyInstance, ctx: PortIntersection)` — we need a reusable way to create mock context objects that satisfy the port interfaces.
|
||||
|
||||
### Design
|
||||
|
||||
Create a `MockRouteContext` factory that builds a mock object satisfying all port interfaces. Each port's methods are `vi.fn()` stubs. Tests can override specific methods as needed.
|
||||
|
||||
### Route registration signatures (verified)
|
||||
|
||||
Each route module requires a specific port intersection. The mock must satisfy all of them:
|
||||
|
||||
| Route Module | Required Ports |
|
||||
|-------------|----------------|
|
||||
| `registerSessionRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
|
||||
| `registerSystemRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
|
||||
| `registerRespawnRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
|
||||
| `registerRalphRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
|
||||
| `registerPlanRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort` |
|
||||
| `registerCaseRoutes` | `EventPort & ConfigPort` |
|
||||
| `registerScheduledRoutes` | `SessionPort & EventPort & InfraPort` |
|
||||
| `registerFileRoutes` | `SessionPort` |
|
||||
| `registerMuxRoutes` | `InfraPort` |
|
||||
| `registerPushRoutes` | `InfraPort` |
|
||||
| `registerTeamRoutes` | `InfraPort` |
|
||||
| `registerHookEventRoutes` | `EventPort & AuthPort` |
|
||||
|
||||
### Implementation
|
||||
|
||||
1. Create `test/mocks/mock-route-context.ts`:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Mock context for route handler testing.
|
||||
*
|
||||
* Satisfies ALL port interfaces (SessionPort, EventPort, RespawnPort,
|
||||
* ConfigPort, InfraPort, AuthPort) so any route module can be tested.
|
||||
* Override specific methods in individual tests as needed.
|
||||
*
|
||||
* Verified against actual port interfaces in src/web/ports/:
|
||||
* - SessionPort: 6 methods (sessions, addSession, cleanupSession,
|
||||
* setupSessionListeners, persistSessionState, persistSessionStateNow,
|
||||
* getSessionStateWithRespawn)
|
||||
* - EventPort: 5 methods (broadcast, sendPushNotifications, batchTerminalData,
|
||||
* broadcastSessionStateDebounced, batchTaskUpdate)
|
||||
* - RespawnPort: 2 maps + 4 methods
|
||||
* - ConfigPort: 5 readonly + 7 methods (incl getDefaultClaudeMdPath,
|
||||
* getLightState, getLightSessionsState, stopTranscriptWatcher)
|
||||
* - InfraPort: 7 readonly + 2 methods (startScheduledRun, stopScheduledRun)
|
||||
* - AuthPort: 3 readonly (authSessions, qrAuthFailures, https)
|
||||
*/
|
||||
import { vi } from 'vitest';
|
||||
import { MockSession, createMockSession } from './mock-session.js';
|
||||
|
||||
/**
|
||||
* Creates a mock context that satisfies all port interfaces.
|
||||
* Pre-populated with one session for convenience.
|
||||
*/
|
||||
export function createMockRouteContext(options?: { sessionId?: string }) {
|
||||
const sessionId = options?.sessionId ?? 'test-session-1';
|
||||
const session = createMockSession(sessionId);
|
||||
const sessions = new Map<string, MockSession>();
|
||||
sessions.set(sessionId, session);
|
||||
|
||||
return {
|
||||
// -- SessionPort --
|
||||
sessions,
|
||||
addSession: vi.fn(),
|
||||
cleanupSession: vi.fn(),
|
||||
setupSessionListeners: vi.fn(),
|
||||
persistSessionState: vi.fn(),
|
||||
persistSessionStateNow: vi.fn(),
|
||||
getSessionStateWithRespawn: vi.fn((s: unknown) => s),
|
||||
|
||||
// -- EventPort --
|
||||
broadcast: vi.fn(),
|
||||
sendPushNotifications: vi.fn(),
|
||||
batchTerminalData: vi.fn(),
|
||||
broadcastSessionStateDebounced: vi.fn(),
|
||||
batchTaskUpdate: vi.fn(),
|
||||
|
||||
// -- RespawnPort --
|
||||
respawnControllers: new Map(),
|
||||
respawnTimers: new Map(),
|
||||
setupRespawnListeners: vi.fn(),
|
||||
setupTimedRespawn: vi.fn(),
|
||||
restoreRespawnController: vi.fn(),
|
||||
saveRespawnConfig: vi.fn(),
|
||||
|
||||
// -- ConfigPort --
|
||||
store: {
|
||||
getConfig: vi.fn(() => ({})),
|
||||
getSessions: vi.fn(() => ({})),
|
||||
getSession: vi.fn(),
|
||||
setSession: vi.fn(),
|
||||
removeSession: vi.fn(),
|
||||
getSettings: vi.fn(() => ({})),
|
||||
setSettings: vi.fn(),
|
||||
getRalphLoopState: vi.fn(() => ({})),
|
||||
setRalphLoopState: vi.fn(),
|
||||
getTasks: vi.fn(() => ({})),
|
||||
save: vi.fn(),
|
||||
load: vi.fn(),
|
||||
},
|
||||
port: 3000,
|
||||
https: false,
|
||||
testMode: true,
|
||||
serverStartTime: Date.now(),
|
||||
getGlobalNiceConfig: vi.fn(async () => undefined),
|
||||
getModelConfig: vi.fn(async () => null),
|
||||
getClaudeModeConfig: vi.fn(async () => ({})),
|
||||
getDefaultClaudeMdPath: vi.fn(async () => undefined),
|
||||
getLightState: vi.fn(() => ({ sessions: [], status: 'ok' })),
|
||||
getLightSessionsState: vi.fn(() => []),
|
||||
startTranscriptWatcher: vi.fn(),
|
||||
stopTranscriptWatcher: vi.fn(),
|
||||
|
||||
// -- InfraPort --
|
||||
mux: {
|
||||
createSession: vi.fn(),
|
||||
killSession: vi.fn(),
|
||||
listSessions: vi.fn(() => []),
|
||||
getStats: vi.fn(() => ({})),
|
||||
},
|
||||
runSummaryTrackers: new Map(),
|
||||
activePlanOrchestrators: new Map(),
|
||||
scheduledRuns: new Map(),
|
||||
teamWatcher: { getTeams: vi.fn(() => []), hasActiveTeammates: vi.fn(() => false) },
|
||||
tunnelManager: null,
|
||||
pushStore: null,
|
||||
startScheduledRun: vi.fn(),
|
||||
stopScheduledRun: vi.fn(),
|
||||
|
||||
// -- AuthPort --
|
||||
authSessions: null,
|
||||
qrAuthFailures: null,
|
||||
// https already declared above in ConfigPort (shared property)
|
||||
|
||||
// Convenience accessors (not part of any port interface)
|
||||
_session: session,
|
||||
_sessionId: sessionId,
|
||||
};
|
||||
}
|
||||
|
||||
export type MockRouteContext = ReturnType<typeof createMockRouteContext>;
|
||||
```
|
||||
|
||||
2. Add to `test/mocks/index.ts` barrel:
|
||||
|
||||
```typescript
|
||||
export { createMockRouteContext, type MockRouteContext } from './mock-route-context.js';
|
||||
```
|
||||
|
||||
3. Create `test/routes/` directory for route test files.
|
||||
|
||||
4. Create `test/routes/_route-test-utils.ts` with Fastify test helpers:
|
||||
|
||||
```typescript
|
||||
/**
|
||||
* Shared utilities for route testing.
|
||||
*
|
||||
* Creates minimal Fastify instances with just the route module under test
|
||||
* and a mock context. Uses app.inject() for HTTP testing without real ports.
|
||||
*/
|
||||
import Fastify, { type FastifyInstance } from 'fastify';
|
||||
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
|
||||
|
||||
export interface RouteTestHarness {
|
||||
app: FastifyInstance;
|
||||
ctx: MockRouteContext;
|
||||
}
|
||||
|
||||
/**
|
||||
* Creates a Fastify instance with a route module registered against a mock context.
|
||||
*
|
||||
* @param registerFn - The route registration function (e.g., registerSessionRoutes).
|
||||
* Uses `any` for ctx parameter because route functions expect typed port intersections
|
||||
* that MockRouteContext satisfies structurally but not nominally.
|
||||
* @param ctxOptions - Optional overrides for the mock context
|
||||
*/
|
||||
export async function createRouteTestHarness(
|
||||
// eslint-disable-next-line @typescript-eslint/no-explicit-any
|
||||
registerFn: (app: FastifyInstance, ctx: any) => void,
|
||||
ctxOptions?: { sessionId?: string },
|
||||
): Promise<RouteTestHarness> {
|
||||
const app = Fastify({ logger: false });
|
||||
const ctx = createMockRouteContext(ctxOptions);
|
||||
|
||||
registerFn(app, ctx);
|
||||
await app.ready();
|
||||
|
||||
return { app, ctx };
|
||||
}
|
||||
```
|
||||
|
||||
### Why `ctx: any` in the harness
|
||||
|
||||
Route registration functions like `registerSessionRoutes(app, ctx: SessionPort & EventPort & ConfigPort & InfraPort & AuthPort)` expect specific port intersection types. TypeScript won't accept `unknown` here because it's not assignable to the port types. The `MockRouteContext` satisfies the interfaces structurally (it has all the required properties and methods), but since it's not declared as implementing them, we need `any` at the call site. This is the standard pattern for test mocks in TypeScript.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 7: Add session routes tests
|
||||
|
||||
**Estimated effort**: 4 hours
|
||||
**Depends on**: Task 6
|
||||
**Files created**: `test/routes/session-routes.test.ts`
|
||||
**Port**: 3220 (only if SSE tests needed; prefer `app.inject()`)
|
||||
|
||||
### Coverage targets
|
||||
|
||||
`src/web/routes/session-routes.ts` is the largest route module (43 handlers). Focus on the most critical endpoints first:
|
||||
|
||||
#### Priority 1: Session CRUD (must test)
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `GET` | `/api/sessions` | Returns session list; empty when no sessions |
|
||||
| `GET` | `/api/sessions/:id` | Returns session state; 404 for unknown ID |
|
||||
| `POST` | `/api/sessions` | Creates session; validates workingDir; rejects invalid paths |
|
||||
| `DELETE` | `/api/sessions/:id` | Calls cleanupSession; 404 for unknown ID |
|
||||
|
||||
#### Priority 2: Session I/O
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `POST` | `/api/sessions/:id/input` | Sends input to session; validates input length; 404 for unknown |
|
||||
| `POST` | `/api/sessions/:id/resize` | Validates cols/rows bounds; 404 for unknown |
|
||||
| `GET` | `/api/sessions/:id/buffer` | Returns terminal buffer; 404 for unknown |
|
||||
|
||||
#### Priority 3: Session actions
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `POST` | `/api/sessions/:id/run` | Runs prompt on session |
|
||||
| `POST` | `/api/sessions/:id/clear` | Clears session |
|
||||
| `POST` | `/api/sessions/:id/compact` | Compacts session |
|
||||
| `POST` | `/api/sessions/:id/interactive` | Starts interactive mode |
|
||||
| `POST` | `/api/sessions/:id/quick-start` | Quick start flow |
|
||||
|
||||
### Test pattern
|
||||
|
||||
```typescript
|
||||
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
|
||||
import { createRouteTestHarness, type RouteTestHarness } from './_route-test-utils.js';
|
||||
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
|
||||
|
||||
describe('session-routes', () => {
|
||||
let harness: RouteTestHarness;
|
||||
|
||||
beforeEach(async () => {
|
||||
harness = await createRouteTestHarness(registerSessionRoutes);
|
||||
});
|
||||
|
||||
afterEach(async () => {
|
||||
await harness.app.close();
|
||||
});
|
||||
|
||||
describe('GET /api/sessions', () => {
|
||||
it('returns empty array when no sessions', async () => {
|
||||
harness.ctx.sessions.clear();
|
||||
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
|
||||
expect(res.statusCode).toBe(200);
|
||||
expect(JSON.parse(res.body)).toEqual([]);
|
||||
});
|
||||
|
||||
it('returns session list with one session', async () => {
|
||||
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
|
||||
expect(res.statusCode).toBe(200);
|
||||
const sessions = JSON.parse(res.body);
|
||||
expect(sessions).toHaveLength(1);
|
||||
});
|
||||
});
|
||||
|
||||
describe('GET /api/sessions/:id', () => {
|
||||
it('returns 404 for unknown session', async () => {
|
||||
const res = await harness.app.inject({
|
||||
method: 'GET',
|
||||
url: '/api/sessions/nonexistent',
|
||||
});
|
||||
expect(res.statusCode).toBe(404);
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/sessions/:id/input', () => {
|
||||
it('rejects input exceeding max length', async () => {
|
||||
const res = await harness.app.inject({
|
||||
method: 'POST',
|
||||
url: `/api/sessions/${harness.ctx._sessionId}/input`,
|
||||
payload: { input: 'x'.repeat(65537) },
|
||||
});
|
||||
expect(res.statusCode).toBe(400);
|
||||
});
|
||||
});
|
||||
|
||||
describe('POST /api/sessions/:id/resize', () => {
|
||||
it('rejects cols exceeding max', async () => {
|
||||
const res = await harness.app.inject({
|
||||
method: 'POST',
|
||||
url: `/api/sessions/${harness.ctx._sessionId}/resize`,
|
||||
payload: { cols: 501, rows: 24 },
|
||||
});
|
||||
expect(res.statusCode).toBe(400);
|
||||
});
|
||||
});
|
||||
});
|
||||
```
|
||||
|
||||
### Key assertions to include
|
||||
|
||||
- **404 for unknown sessions**: Every `:id` endpoint must return 404 for nonexistent IDs
|
||||
- **Input validation**: Bad paths, oversized inputs, invalid resize dimensions
|
||||
- **Side effects**: Verify `ctx.broadcast()` was called with correct event type after mutations
|
||||
- **Response shape**: Verify response bodies match expected API types
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/routes/session-routes.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 8: Add system + respawn routes tests
|
||||
|
||||
**Estimated effort**: 4 hours
|
||||
**Depends on**: Task 6
|
||||
**Files created**: `test/routes/system-routes.test.ts`, `test/routes/respawn-routes.test.ts`
|
||||
|
||||
### System routes (`src/web/routes/system-routes.ts`)
|
||||
|
||||
Focus on status and configuration endpoints:
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `GET` | `/api/status` | Returns server status with uptime, session count |
|
||||
| `GET` | `/api/stats` | Returns mux stats |
|
||||
| `GET` | `/api/config` | Returns current config |
|
||||
| `PUT` | `/api/config` | Updates config; validates input |
|
||||
| `GET` | `/api/settings` | Returns user settings |
|
||||
| `PUT` | `/api/settings` | Updates settings; validates input |
|
||||
| `GET` | `/api/subagents` | Returns subagent list |
|
||||
| `GET` | `/api/screenshots` | Returns screenshot list |
|
||||
|
||||
### Respawn routes (`src/web/routes/respawn-routes.ts`)
|
||||
|
||||
| Method | Path | What to test |
|
||||
|--------|------|-------------|
|
||||
| `GET` | `/api/sessions/:id/respawn` | Returns respawn status; null when not configured |
|
||||
| `POST` | `/api/sessions/:id/respawn/start` | Starts respawn; 404 for unknown session |
|
||||
| `POST` | `/api/sessions/:id/respawn/stop` | Stops respawn; 404 for unknown session |
|
||||
| `PUT` | `/api/sessions/:id/respawn/config` | Updates respawn config; validates |
|
||||
| `POST` | `/api/sessions/:id/respawn/enable` | Enables respawn loop |
|
||||
| `POST` | `/api/sessions/:id/respawn/disable` | Disables respawn loop |
|
||||
|
||||
### Test patterns
|
||||
|
||||
Same pattern as Task 7 — `createRouteTestHarness` with `registerSystemRoutes` / `registerRespawnRoutes`.
|
||||
|
||||
For respawn tests, pre-populate `ctx.respawnControllers` with a mock controller in `beforeEach`:
|
||||
|
||||
```typescript
|
||||
beforeEach(async () => {
|
||||
harness = await createRouteTestHarness(registerRespawnRoutes);
|
||||
// Add a mock respawn controller for the default session
|
||||
harness.ctx.respawnControllers.set(harness.ctx._sessionId, {
|
||||
getState: vi.fn(() => 'idle'),
|
||||
getConfig: vi.fn(() => ({})),
|
||||
getStatus: vi.fn(() => ({ state: 'idle', health: 100 })),
|
||||
start: vi.fn(),
|
||||
stop: vi.fn(),
|
||||
updateConfig: vi.fn(),
|
||||
enable: vi.fn(),
|
||||
disable: vi.fn(),
|
||||
});
|
||||
});
|
||||
```
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
npx vitest run test/routes/system-routes.test.ts
|
||||
npx vitest run test/routes/respawn-routes.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Task 9: Slim down `respawn-test-utils.ts`
|
||||
|
||||
**Estimated effort**: 30 minutes
|
||||
**Depends on**: Tasks 4, 5
|
||||
**Files modified**: `test/respawn-test-utils.ts`
|
||||
|
||||
After Tasks 4–5 are verified passing with shared mocks, slim down `respawn-test-utils.ts` to remove duplicates.
|
||||
|
||||
### Steps
|
||||
|
||||
1. **Remove** from `respawn-test-utils.ts` what has been moved to shared mocks:
|
||||
- `MockSession` class → now in `test/mocks/mock-session.ts`
|
||||
- `createMockSession()` → now in `test/mocks/mock-session.ts`
|
||||
- `terminalOutputs` → now in `test/mocks/mock-session.ts`
|
||||
- `waitForEvent()` / `createDeferred()` → now in `test/mocks/test-helpers.ts`
|
||||
|
||||
2. **Keep** respawn-specific utilities that don't belong in the general mocks:
|
||||
- `TimeController` / `createTimeController()` — respawn-specific timer control
|
||||
- `MockAiIdleChecker` / `MockAiPlanChecker` — respawn-specific AI mocks
|
||||
- `createStateTracker()` / `createEventRecorder()` — respawn state tracking
|
||||
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — respawn config presets
|
||||
- `waitForState()` — respawn state machine waiter
|
||||
|
||||
3. **Update imports** in `respawn-test-utils.ts` to re-use shared mocks:
|
||||
```typescript
|
||||
import { MockSession, createMockSession, terminalOutputs } from './mocks/index.js';
|
||||
import { waitForEvent, createDeferred } from './mocks/index.js';
|
||||
export { MockSession, createMockSession, terminalOutputs, waitForEvent, createDeferred };
|
||||
```
|
||||
|
||||
This preserves backward compatibility for any future tests that import from `respawn-test-utils.ts` directly while eliminating the duplication.
|
||||
|
||||
### Verification
|
||||
|
||||
```bash
|
||||
tsc --noEmit
|
||||
npx vitest run test/respawn-controller.test.ts
|
||||
npx vitest run test/respawn-team-awareness.test.ts
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## What is NOT in scope (and why)
|
||||
|
||||
### Migrating `session-manager.test.ts` and `ralph-loop.test.ts` mocks
|
||||
|
||||
Both files define mocks inside `vi.mock()` factories that replace entire modules:
|
||||
|
||||
```typescript
|
||||
// session-manager.test.ts — mock replaces ../src/session.js
|
||||
vi.mock('../src/session.js', () => {
|
||||
class MockSession extends EventEmitter { ... }
|
||||
return { Session: MockSession };
|
||||
});
|
||||
|
||||
// ralph-loop.test.ts — mock replaces ../src/state-store.js
|
||||
vi.mock('../src/state-store.js', () => {
|
||||
class MockStateStore { ... }
|
||||
return { getStore: vi.fn(() => instance), StateStore: MockStateStore };
|
||||
});
|
||||
```
|
||||
|
||||
These are fundamentally different from the direct-instantiation pattern:
|
||||
- The `vi.mock()` factory runs in an isolated scope — outer imports are not available
|
||||
- The mock class must be returned with the exact export names (`Session`, `getStore`, `StateStore`)
|
||||
- The `session-manager.test.ts` MockSession auto-registers into a shared `mockState.sessions` Map (tight coupling with test setup)
|
||||
|
||||
Migrating would require `vi.hoisted()` to share the class between factory and test scope, plus restructuring the test's module-mocking setup. This is high-complexity, high-risk refactoring with limited benefit since these tests already work. The shared `MockStateStore` in `test/mocks/` is available for **new** tests (like route tests) that use direct instantiation instead.
|
||||
|
||||
### Full integration tests with real Fastify server
|
||||
|
||||
Route tests use `app.inject()` which simulates HTTP without opening ports. Full integration tests that spin up `WebServer`, create real sessions, and stream SSE would be valuable but are a separate effort requiring:
|
||||
- A test WebServer factory
|
||||
- Session lifecycle management in tests
|
||||
- SSE client test utilities
|
||||
- Significantly more setup/teardown complexity
|
||||
|
||||
### Testing auth middleware in route tests
|
||||
|
||||
Route tests bypass authentication (no auth middleware registered on the test Fastify instance). Auth middleware has its own dedicated tests in `auth-security.test.ts` and `qr-auth.test.ts`. Testing auth + routes together is a future integration test concern.
|
||||
|
||||
### Testing SSE event streaming
|
||||
|
||||
SSE integration requires a running server with `EventSource` client. This is significantly more complex than `app.inject()` tests and is deferred. The existing `sse-events.test.ts` covers SSE patterns.
|
||||
|
||||
### Complete route coverage for all 12 modules
|
||||
|
||||
This phase covers the 3 highest-value route modules (session, system, respawn — 98 of 162 handlers). The remaining 9 modules (ralph, plan, push, team, mux, file, scheduled, hook-event, case) should be added incrementally in follow-up work.
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
| Metric | Before | After |
|
||||
|--------|--------|-------|
|
||||
| MockSession definitions | 4 (across 4 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
|
||||
| MockStateStore definitions | 2 (across 2 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
|
||||
| Files importing from `respawn-test-utils.ts` | 0 | Utilities split into `test/mocks/` |
|
||||
| Route test files | 0 | 3 (session, system, respawn) |
|
||||
| Route handlers with dedicated tests | 0 | ~30 (highest-priority endpoints) |
|
||||
| Shared mock directory | None | `test/mocks/` with 5 files + barrel |
|
||||
|
||||
### Final verification checklist
|
||||
|
||||
```bash
|
||||
# Type checking
|
||||
tsc --noEmit
|
||||
|
||||
# Linting
|
||||
npm run lint
|
||||
|
||||
# Formatting
|
||||
npm run format:check
|
||||
|
||||
# Run all affected tests individually
|
||||
npx vitest run test/respawn-controller.test.ts
|
||||
npx vitest run test/respawn-team-awareness.test.ts
|
||||
npx vitest run test/routes/session-routes.test.ts
|
||||
npx vitest run test/routes/system-routes.test.ts
|
||||
npx vitest run test/routes/respawn-routes.test.ts
|
||||
|
||||
# Verify unchanged tests still pass
|
||||
npx vitest run test/session-manager.test.ts
|
||||
npx vitest run test/ralph-loop.test.ts
|
||||
|
||||
# Dev server still starts
|
||||
npx tsx src/index.ts web --port 3099 &
|
||||
curl -s http://localhost:3099/api/status | jq .status # "ok"
|
||||
kill %1
|
||||
```
|
||||
@@ -0,0 +1,247 @@
|
||||
# Ralph Loop Plan Improvement Roadmap
|
||||
|
||||
> Research-backed improvements for rock-solid AI planning with auto-improvement capabilities.
|
||||
|
||||
**Created**: 2026-01-27
|
||||
**Status**: Implementation in Progress
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
1. [Research Summary](#research-summary)
|
||||
2. [Current State Analysis](#current-state-analysis)
|
||||
3. [Proposed Improvements](#proposed-improvements)
|
||||
4. [Implementation Plan](#implementation-plan)
|
||||
5. [Sources](#sources)
|
||||
|
||||
---
|
||||
|
||||
## Research Summary
|
||||
|
||||
### Key Insights from Industry Best Practices
|
||||
|
||||
#### 1. Self-Verification is Critical
|
||||
> "Claude performs dramatically better when it can verify its own work, like run tests, compare screenshots, and validate outputs. Without clear success criteria, it might produce something that looks right but actually doesn't work."
|
||||
> — [Anthropic Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
|
||||
|
||||
#### 2. Iterative Refinement Patterns (AWS)
|
||||
> "A generator agent produces output, an evaluator agent reviews using evaluation rubric, and based on feedback, an optimizer agent revises the output. Loop repeats until criteria met."
|
||||
> — [AWS Agentic AI Patterns](https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-patterns/evaluator-reflect-refine-loop-patterns.html)
|
||||
|
||||
#### 3. Dynamic Task Decomposition (TDAG Framework)
|
||||
> "Dynamically decomposes complex tasks into smaller subtasks and assigns each to a specifically generated subagent, enhancing adaptability in diverse and unpredictable real-world tasks."
|
||||
> — [TDAG Framework - arXiv](https://arxiv.org/abs/2402.10178)
|
||||
|
||||
#### 4. Multi-Stage Verification Workflow
|
||||
> "o3: Generate plan → Sonnet: Verify and create task list → Sonnet: Execute → Sonnet: Verify against plan → o3: Final verification → Issues bake back into plan"
|
||||
> — [Claude Code Best Practices Community](https://rosmur.github.io/claudecode-best-practices/)
|
||||
|
||||
#### 5. Self-Improving Agents
|
||||
> "Through an iterative refinement process (analyze outcome → adjust approach → try again), the agent becomes more adept at handling tasks over time. It effectively builds a growing knowledge base of what strategies work best."
|
||||
> — [Self-Improving Data Agents](https://powerdrill.ai/blog/self-improving-data-agents)
|
||||
|
||||
#### 6. Memory Architecture for Planning
|
||||
> "Agents use three memory layers: working memory for short-lived calculations, episodic memory for step-by-step histories, and semantic memory for long-term knowledge."
|
||||
> — [LLM Agent Research](https://www.promptingguide.ai/research/llm-agents)
|
||||
|
||||
---
|
||||
|
||||
## Current State Analysis
|
||||
|
||||
### What We Have
|
||||
|
||||
The current plan generation system (`/api/generate-plan` and `/api/generate-plan-detailed`):
|
||||
|
||||
1. **Standard Mode**: Single Opus 4.5 call with TDD-focused prompt
|
||||
2. **Enhanced Mode**: 4 parallel subagents (Requirements, Architecture, Testing, Risks) + Verification
|
||||
|
||||
### Current Plan Item Structure
|
||||
|
||||
```json
|
||||
{
|
||||
"content": "Implement login endpoint",
|
||||
"priority": "P0"
|
||||
}
|
||||
```
|
||||
|
||||
### Limitations
|
||||
|
||||
| Issue | Impact |
|
||||
|-------|--------|
|
||||
| No verification criteria | Can't automatically validate completion |
|
||||
| No test pairing | TDD not enforced structurally |
|
||||
| Static plans | No adaptation during execution |
|
||||
| No dependencies | Can't track blocking relationships |
|
||||
| No failure tracking | Same errors repeat |
|
||||
| No checkpoints | Plans run until completion or failure |
|
||||
|
||||
---
|
||||
|
||||
## Proposed Improvements
|
||||
|
||||
### Enhanced Plan Item Structure
|
||||
|
||||
```typescript
|
||||
interface EnhancedPlanItem {
|
||||
id: string; // Unique identifier (e.g., "P0-001")
|
||||
content: string; // Task description
|
||||
priority: 'P0' | 'P1' | 'P2'; // Criticality
|
||||
phase: 'setup' | 'test' | 'impl' | 'verify'; // Development phase
|
||||
|
||||
// NEW: Verification
|
||||
verificationCriteria: string; // How to know it's done
|
||||
testCommand?: string; // Command to run for verification
|
||||
|
||||
// NEW: Dependencies
|
||||
dependencies: string[]; // IDs of tasks that must complete first
|
||||
blockedBy?: string[]; // Runtime: tasks blocking this one
|
||||
|
||||
// NEW: Execution tracking
|
||||
status: 'pending' | 'in_progress' | 'completed' | 'failed' | 'blocked';
|
||||
attempts: number; // How many times attempted
|
||||
lastError?: string; // Most recent failure reason
|
||||
completedAt?: number; // Timestamp of completion
|
||||
|
||||
// NEW: Metadata
|
||||
estimatedComplexity: 'low' | 'medium' | 'high';
|
||||
rollbackStrategy?: string; // How to undo if needed
|
||||
version: number; // Plan version this belongs to
|
||||
}
|
||||
```
|
||||
|
||||
### Runtime Plan Adaptation Flow
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────┐
|
||||
│ RUNTIME PLAN LOOP │
|
||||
├─────────────────────────────────────────────────────────────────┤
|
||||
│ │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │ Execute │──▶│ Verify │──▶│ Success? │──▶│ Mark │ │
|
||||
│ │ Task │ │ Output │ │ │ │ Complete │ │
|
||||
│ └──────────┘ └──────────┘ └────┬─────┘ └──────────┘ │
|
||||
│ │ No │
|
||||
│ ▼ │
|
||||
│ ┌──────────┐ │
|
||||
│ │ Analyze │ │
|
||||
│ │ Failure │ │
|
||||
│ └────┬─────┘ │
|
||||
│ │ │
|
||||
│ ┌──────────────┼──────────────┐ │
|
||||
│ ▼ ▼ ▼ │
|
||||
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
|
||||
│ │ Retry │ │ Add Fix │ │ Escalate │ │
|
||||
│ │ (< 3x) │ │ Sub-Task │ │ BLOCKED │ │
|
||||
│ └──────────┘ └──────────┘ └──────────┘ │
|
||||
│ │
|
||||
└───────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### Checkpoint Review System
|
||||
|
||||
At iterations 5, 10, 20, 30, 50:
|
||||
1. Pause execution
|
||||
2. Summarize progress (completed/failed/pending)
|
||||
3. Identify stuck items (3+ failures)
|
||||
4. Generate alternative approaches for stuck items
|
||||
5. Update plan with new strategies
|
||||
6. Continue with refined plan
|
||||
|
||||
---
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
### Phase 1: Quick Wins (Implementing Now)
|
||||
|
||||
#### 1.1 Add Verification Criteria to Plan Items
|
||||
- Modify plan generation prompts to require `verificationCriteria`
|
||||
- Update `PlanItem` interface in `types.ts`
|
||||
- Update plan orchestrator prompts
|
||||
|
||||
#### 1.2 Pair Test/Implementation Steps
|
||||
- Ensure every implementation step has a corresponding test step
|
||||
- Group items: test → implement → verify
|
||||
- Add phase field to track TDD cycle
|
||||
|
||||
#### 1.3 Checkpoint Review Prompts
|
||||
- Add checkpoint logic to Ralph tracker
|
||||
- At iterations 5, 10, 20: inject review prompt
|
||||
- Generate progress summary and stuck item analysis
|
||||
|
||||
### Phase 2: Medium Effort (Implementing Now)
|
||||
|
||||
#### 2.1 Failure Tracking
|
||||
- Track `attempts` and `lastError` per task
|
||||
- After 3 failures, auto-generate debug sub-task
|
||||
- Record failure patterns in plan history
|
||||
|
||||
#### 2.2 Plan Versioning
|
||||
- Add `version` field to plans
|
||||
- Keep history in `@fix_plan.md` with version markers
|
||||
- Allow rollback to previous versions
|
||||
- Track which version each task belongs to
|
||||
|
||||
#### 2.3 Dependency Tracking
|
||||
- Add `dependencies` field to plan items
|
||||
- Validate dependency graph (no cycles)
|
||||
- Block tasks until dependencies complete
|
||||
- Show dependency status in UI
|
||||
|
||||
### Phase 3: Future Enhancements
|
||||
|
||||
#### 3.1 Full Runtime Adaptation
|
||||
- TDAG-style dynamic decomposition
|
||||
- Auto-generate sub-tasks for complex items
|
||||
- Learning from failure patterns
|
||||
|
||||
#### 3.2 Multi-Model Verification
|
||||
- Haiku: Fast initial generation
|
||||
- Sonnet: Verification and refinement
|
||||
- Opus: Final quality check
|
||||
|
||||
#### 3.3 Plan Memory System
|
||||
- Episodic memory: What worked/failed in this session
|
||||
- Semantic memory: Patterns across projects
|
||||
- Use for future plan generation
|
||||
|
||||
---
|
||||
|
||||
## File Changes Required
|
||||
|
||||
### New/Modified Files
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/types.ts` | Add `EnhancedPlanItem` interface |
|
||||
| `src/plan-orchestrator.ts` | Update prompts, add versioning |
|
||||
| `src/ralph-tracker.ts` | Add checkpoint logic, failure tracking |
|
||||
| `src/web/server.ts` | New endpoints for plan updates |
|
||||
| `src/web/public/app.js` | UI for enhanced plan display |
|
||||
|
||||
### New Endpoints
|
||||
|
||||
| Method | Endpoint | Purpose |
|
||||
|--------|----------|---------|
|
||||
| PATCH | `/api/sessions/:id/plan/task/:taskId` | Update task status |
|
||||
| POST | `/api/sessions/:id/plan/checkpoint` | Trigger checkpoint review |
|
||||
| GET | `/api/sessions/:id/plan/history` | Get plan version history |
|
||||
| POST | `/api/sessions/:id/plan/rollback/:version` | Rollback to version |
|
||||
|
||||
---
|
||||
|
||||
## Sources
|
||||
|
||||
- [Anthropic Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
|
||||
- [AWS Agentic AI Patterns](https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-patterns/evaluator-reflect-refine-loop-patterns.html)
|
||||
- [TDAG: Multi-Agent Task Decomposition Framework](https://arxiv.org/abs/2402.10178)
|
||||
- [Self-Improving Data Agents](https://powerdrill.ai/blog/self-improving-data-agents)
|
||||
- [OpenAI Self-Evolving Agents Cookbook](https://cookbook.openai.com/examples/partners/self_evolving_agents/autonomous_agent_retraining)
|
||||
- [Task Decomposition for Coding Agents](https://mgx.dev/insights/task-decomposition-for-coding-agents-architectures-advancements-and-future-directions/)
|
||||
- [Claude Code Best Practices Community Guide](https://rosmur.github.io/claudecode-best-practices/)
|
||||
- [LLM Agents Prompt Engineering Guide](https://www.promptingguide.ai/research/llm-agents)
|
||||
- [Agentic AI Implementation Guide](https://www.sketchdev.io/blog/agentic-ai-implementation-guide)
|
||||
|
||||
---
|
||||
|
||||
*This document is part of the Codeman project. See [CLAUDE.md](../CLAUDE.md) for main documentation.*
|
||||
@@ -0,0 +1,251 @@
|
||||
# Ralph Loop Improvements Plan
|
||||
|
||||
## Overview
|
||||
|
||||
This plan details improvements to Codeman's Ralph Loop system based on best practices from the Ralph Claude Code repository (https://github.com/frankbria/ralph-claude-code).
|
||||
|
||||
## Key Concepts to Implement
|
||||
|
||||
### RALPH_STATUS Block Format
|
||||
|
||||
Claude outputs this structured block at the end of every response for better tracking:
|
||||
|
||||
```
|
||||
---RALPH_STATUS---
|
||||
STATUS: IN_PROGRESS | COMPLETE | BLOCKED
|
||||
TASKS_COMPLETED_THIS_LOOP: <number>
|
||||
FILES_MODIFIED: <number>
|
||||
TESTS_STATUS: PASSING | FAILING | NOT_RUN
|
||||
WORK_TYPE: IMPLEMENTATION | TESTING | DOCUMENTATION | REFACTORING
|
||||
EXIT_SIGNAL: false | true
|
||||
RECOMMENDATION: <one line summary of what to do next>
|
||||
---END_RALPH_STATUS---
|
||||
```
|
||||
|
||||
### Dual-Condition Exit Gate
|
||||
|
||||
Exit requires BOTH conditions:
|
||||
1. `completion_indicators >= 2` (heuristic detection from natural language patterns)
|
||||
2. Claude's explicit `EXIT_SIGNAL: true` in the RALPH_STATUS block
|
||||
|
||||
### Circuit Breaker Pattern
|
||||
|
||||
Three states: CLOSED → HALF_OPEN → OPEN
|
||||
|
||||
| From State | Condition | To State |
|
||||
|------------|-----------|----------|
|
||||
| CLOSED | consecutive_no_progress >= 2 | HALF_OPEN |
|
||||
| CLOSED | consecutive_no_progress >= 3 | OPEN |
|
||||
| CLOSED | consecutive_same_error >= 5 | OPEN |
|
||||
| HALF_OPEN | progress detected | CLOSED |
|
||||
| HALF_OPEN | consecutive_no_progress >= 3 | OPEN |
|
||||
| OPEN | Manual reset | CLOSED |
|
||||
|
||||
### @fix_plan.md Structure
|
||||
|
||||
```markdown
|
||||
# Fix Plan
|
||||
|
||||
## High Priority (P0)
|
||||
- [ ] Critical: Fix authentication bug
|
||||
- [ ] Blocker: Database connection timeout
|
||||
|
||||
## Standard (P1)
|
||||
- [ ] Feature: Add user profile page
|
||||
|
||||
## Nice to Have (P2)
|
||||
- [ ] Improvement: Add dark mode
|
||||
|
||||
## Completed
|
||||
- [x] Setup: Initialize project structure
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase 1: Quick Wins (1-2 days)
|
||||
|
||||
### 1.1 RALPH_STATUS Block Parsing
|
||||
|
||||
**What**: Add parsing support for the structured RALPH_STATUS block format in RalphTracker.
|
||||
|
||||
**Implementation**:
|
||||
- Add regex pattern to detect `---RALPH_STATUS---` blocks
|
||||
- Parse fields: STATUS, TASKS_COMPLETED_THIS_LOOP, FILES_MODIFIED, TESTS_STATUS, WORK_TYPE, EXIT_SIGNAL, RECOMMENDATION
|
||||
- Store in extended `RalphTrackerState` type
|
||||
- Emit new events: `ralphStatusUpdate`
|
||||
|
||||
**Files**: `ralph-tracker.ts`, `types.ts`
|
||||
|
||||
### 1.2 Enhanced Status Display in UI
|
||||
|
||||
**What**: Display RALPH_STATUS fields in the Ralph State Panel.
|
||||
|
||||
**Implementation**:
|
||||
- Add UI elements: WORK_TYPE indicator, TESTS_STATUS badge, FILES_MODIFIED count
|
||||
- Show RECOMMENDATION text in expanded view
|
||||
- Color-code status (IN_PROGRESS=blue, COMPLETE=green, BLOCKED=red)
|
||||
|
||||
**Files**: `app.js`, `styles.css`, `index.html`
|
||||
|
||||
### 1.3 Prompt Template Improvements
|
||||
|
||||
**What**: Add specification-by-example exit scenarios to prompts.
|
||||
|
||||
**Implementation**:
|
||||
- Add "Exit Scenarios" section to case-template.md
|
||||
- Document when to continue vs. when to output completion
|
||||
- Include testing limits guidance (max 20% effort on tests)
|
||||
- Add RALPH_STATUS block instructions
|
||||
|
||||
**Files**: `case-template.md`
|
||||
|
||||
### 1.4 Better Wizard Validation
|
||||
|
||||
**What**: Add client-side validation and helpful warnings.
|
||||
|
||||
**Implementation**:
|
||||
- Warn if task description < 50 chars
|
||||
- Warn if no success criteria mentioned
|
||||
- Suggest adding test requirements if none detected
|
||||
- Validate completion phrase is uppercase alphanumeric
|
||||
|
||||
**Files**: `app.js`
|
||||
|
||||
---
|
||||
|
||||
## Phase 2: Core Improvements (3-5 days)
|
||||
|
||||
### 2.1 Circuit Breaker Pattern
|
||||
|
||||
**What**: Implement three-state circuit breaker to detect stuck loops.
|
||||
|
||||
**Implementation**:
|
||||
- Create `CircuitBreaker` class with CLOSED, HALF_OPEN, OPEN states
|
||||
- Track: files_modified, tasks_completed, error_patterns per iteration
|
||||
- Triggers: N consecutive no-progress, same error M times, tests failing K iterations
|
||||
- Emit events: `circuitBreakerStateChange`
|
||||
|
||||
**Files**: New `circuit-breaker.ts`, integrate into `ralph-tracker.ts`
|
||||
|
||||
### 2.2 Circuit Breaker UI
|
||||
|
||||
**What**: Visual indicator in Ralph panel.
|
||||
|
||||
**Implementation**:
|
||||
- Badge: green (CLOSED), yellow (HALF_OPEN), red (OPEN)
|
||||
- Warning before tripping
|
||||
- Notification when circuit opens
|
||||
- Manual reset button
|
||||
|
||||
**Files**: `app.js`, `styles.css`, `index.html`
|
||||
|
||||
### 2.3 @fix_plan.md Integration
|
||||
|
||||
**What**: Generate and track structured task plan file.
|
||||
|
||||
**Implementation**:
|
||||
- Generate `@fix_plan.md` in working directory when loop starts
|
||||
- Watch file for changes and sync with RalphTracker todos
|
||||
- Parse priority levels (P0, P1, P2)
|
||||
- Show priority in UI
|
||||
|
||||
**Files**: New `fix-plan.ts`, `ralph-tracker.ts`, `server.ts`
|
||||
|
||||
### 2.4 Wizard Plan Generation Step
|
||||
|
||||
**What**: Add third wizard step for AI-assisted plan generation.
|
||||
|
||||
**Implementation**:
|
||||
- Step 2: "Plan Generation" between Task Setup and Launch
|
||||
- Use Claude to break down task into fix plan items
|
||||
- Allow edit/reorder before launch
|
||||
- Generate @fix_plan.md with selected items
|
||||
|
||||
**Files**: `app.js`, `index.html`, `server.ts`
|
||||
|
||||
### 2.5 Smart Respawn Integration
|
||||
|
||||
**What**: Use RALPH_STATUS for respawn decisions.
|
||||
|
||||
**Implementation**:
|
||||
- Use EXIT_SIGNAL field for respawn decisions
|
||||
- If STATUS=BLOCKED, trigger circuit breaker instead of respawn
|
||||
- Pass RECOMMENDATION to respawn update prompt
|
||||
|
||||
**Files**: `respawn-controller.ts`, `ralph-tracker.ts`
|
||||
|
||||
---
|
||||
|
||||
## Phase 3: Advanced Features (5+ days)
|
||||
|
||||
### 3.1 Template Library
|
||||
- Bug Fix, Feature, Refactoring, Test Coverage, Documentation templates
|
||||
- Template selector in wizard
|
||||
- Custom templates in `~/.codeman/templates/`
|
||||
|
||||
### 3.2 Tool Permissions
|
||||
- Configure allowed Claude tools per loop
|
||||
- Generate hook configuration
|
||||
- Store in session config
|
||||
|
||||
### 3.3 Per-Iteration Timeout
|
||||
- Max time per iteration (5-60 min)
|
||||
- Auto-continue on timeout
|
||||
- Log timeout events
|
||||
|
||||
### 3.4 Rate Limiting
|
||||
- Max tokens per iteration
|
||||
- Max API calls per minute
|
||||
- Cooldown between iterations
|
||||
|
||||
### 3.5 Metrics Dashboard
|
||||
- Time-series charts (files modified, tasks completed, tokens)
|
||||
- Aggregate statistics
|
||||
- Export to JSON/CSV
|
||||
|
||||
---
|
||||
|
||||
## Priority Matrix
|
||||
|
||||
| Item | Effort | Impact | Priority |
|
||||
|------|--------|--------|----------|
|
||||
| 1.1 RALPH_STATUS Parsing | Low | High | P0 |
|
||||
| 1.2 Status Display UI | Low | Medium | P0 |
|
||||
| 1.3 Prompt Templates | Low | High | P0 |
|
||||
| 1.4 Wizard Validation | Low | Medium | P1 |
|
||||
| 2.1 Circuit Breaker | Medium | High | P1 |
|
||||
| 2.2 Circuit Breaker UI | Medium | Medium | P1 |
|
||||
| 2.3 Fix Plan Integration | Medium | High | P1 |
|
||||
| 2.4 Plan Generation Step | Medium | Medium | P2 |
|
||||
| 2.5 Respawn Integration | Medium | High | P1 |
|
||||
| 3.1 Template Selection | High | Medium | P2 |
|
||||
| 3.2 Tool Permissions | High | Medium | P3 |
|
||||
| 3.3 Per-Iteration Timeout | High | Medium | P2 |
|
||||
| 3.4 Rate Limiting | High | Low | P3 |
|
||||
| 3.5 Metrics Dashboard | High | Medium | P3 |
|
||||
|
||||
---
|
||||
|
||||
## Reference: Ralph Claude Code Best Practices
|
||||
|
||||
### Testing Guidelines
|
||||
- LIMIT testing to ~20% of total effort per loop
|
||||
- PRIORITIZE: Implementation > Documentation > Tests
|
||||
- Only write tests for NEW functionality
|
||||
- Do NOT refactor existing tests unless broken
|
||||
|
||||
### What NOT to Do
|
||||
- Do NOT continue with busy work when EXIT_SIGNAL should be true
|
||||
- Do NOT run tests repeatedly without implementing new features
|
||||
- Do NOT refactor code that is already working
|
||||
- Do NOT add features not in specifications
|
||||
- Do NOT forget the status block
|
||||
|
||||
### Exit Scenarios (Specification by Example)
|
||||
|
||||
1. **Successful Completion**: All tasks done → EXIT_SIGNAL=true
|
||||
2. **Test-Only Loop**: No implementation, only testing → continue but warn
|
||||
3. **Stuck on Error**: Same error 5 times → circuit breaker opens
|
||||
4. **No Work Remaining**: All specs done → EXIT_SIGNAL=true
|
||||
5. **Making Progress**: Normal flow → continue
|
||||
6. **Blocked**: Needs human intervention → STATUS=BLOCKED
|
||||
@@ -0,0 +1,434 @@
|
||||
# Ralph Tracker Phase 1 Implementation Plan
|
||||
|
||||
## Overview
|
||||
|
||||
This plan details how to enhance the existing RalphTracker with RALPH_STATUS block parsing, circuit breaker pattern, and dual-condition exit gate.
|
||||
|
||||
---
|
||||
|
||||
## 1. Current State Analysis
|
||||
|
||||
### What RalphTracker Already Does Well
|
||||
|
||||
- **Todo Detection**: Supports 5 formats (checkboxes, indicators, status in parentheses, native TodoWrite, checkmark-based)
|
||||
- **Completion Phrases**: Detects `<promise>PHRASE</promise>` with occurrence-based logic (1st = store, 2nd = complete)
|
||||
- **Loop State Tracking**: Tracks active/inactive, iteration counts, max iterations, elapsed hours, cycle counts
|
||||
- **Auto-Enable**: Disabled by default, auto-enables when Ralph patterns detected
|
||||
- **Event System**: Emits `loopUpdate`, `todoUpdate`, `completionDetected`, `enabled` events
|
||||
- **SSE Integration**: Events forwarded via `session:ralphLoopUpdate`, `session:ralphTodoUpdate`, `session:ralphCompletionDetected`
|
||||
- **Debouncing**: EVENT_DEBOUNCE_MS (50ms) for rapid updates to prevent UI jitter
|
||||
- **Cleanup**: MAX_TODO_ITEMS (50), TODO_EXPIRY_MS (1 hour), throttled cleanup
|
||||
|
||||
### Current Limitations
|
||||
|
||||
| Feature | Status |
|
||||
|---------|--------|
|
||||
| RALPH_STATUS block parsing | Missing |
|
||||
| Circuit breaker pattern | Missing |
|
||||
| Priority-based todos (P0/P1/P2) | Missing |
|
||||
| Dual-condition exit gate | Missing |
|
||||
| Files modified tracking | Missing |
|
||||
| Tests status tracking | Missing |
|
||||
| Work type classification | Missing |
|
||||
|
||||
---
|
||||
|
||||
## 2. New Type Definitions (types.ts)
|
||||
|
||||
```typescript
|
||||
// ========== RALPH_STATUS Block Types ==========
|
||||
|
||||
export type RalphStatusValue = 'IN_PROGRESS' | 'COMPLETE' | 'BLOCKED';
|
||||
export type RalphTestsStatus = 'PASSING' | 'FAILING' | 'NOT_RUN';
|
||||
export type RalphWorkType = 'IMPLEMENTATION' | 'TESTING' | 'DOCUMENTATION' | 'REFACTORING';
|
||||
|
||||
/**
|
||||
* Parsed RALPH_STATUS block from Claude output.
|
||||
*/
|
||||
export interface RalphStatusBlock {
|
||||
status: RalphStatusValue;
|
||||
tasksCompletedThisLoop: number;
|
||||
filesModified: number;
|
||||
testsStatus: RalphTestsStatus;
|
||||
workType: RalphWorkType;
|
||||
exitSignal: boolean;
|
||||
recommendation: string;
|
||||
parsedAt: number;
|
||||
}
|
||||
|
||||
// ========== Circuit Breaker Types ==========
|
||||
|
||||
export type CircuitBreakerState = 'CLOSED' | 'HALF_OPEN' | 'OPEN';
|
||||
|
||||
export type CircuitBreakerReason =
|
||||
| 'normal_operation'
|
||||
| 'no_progress_warning'
|
||||
| 'no_progress_open'
|
||||
| 'same_error_repeated'
|
||||
| 'tests_failing_too_long'
|
||||
| 'progress_detected'
|
||||
| 'manual_reset';
|
||||
|
||||
export interface CircuitBreakerStatus {
|
||||
state: CircuitBreakerState;
|
||||
consecutiveNoProgress: number;
|
||||
consecutiveSameError: number;
|
||||
consecutiveTestsFailure: number;
|
||||
lastProgressIteration: number;
|
||||
reason: string;
|
||||
reasonCode: CircuitBreakerReason;
|
||||
lastTransitionAt: number;
|
||||
lastErrorMessage: string | null;
|
||||
}
|
||||
|
||||
// ========== Priority Todo Types ==========
|
||||
|
||||
export type RalphTodoPriority = 'P0' | 'P1' | 'P2' | null;
|
||||
|
||||
// ========== Helper Functions ==========
|
||||
|
||||
export function createInitialCircuitBreakerStatus(): CircuitBreakerStatus {
|
||||
return {
|
||||
state: 'CLOSED',
|
||||
consecutiveNoProgress: 0,
|
||||
consecutiveSameError: 0,
|
||||
consecutiveTestsFailure: 0,
|
||||
lastProgressIteration: 0,
|
||||
reason: 'Initial state',
|
||||
reasonCode: 'normal_operation',
|
||||
lastTransitionAt: Date.now(),
|
||||
lastErrorMessage: null,
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. New Regex Patterns (ralph-tracker.ts)
|
||||
|
||||
```typescript
|
||||
// ---------- RALPH_STATUS Block Patterns ----------
|
||||
|
||||
const RALPH_STATUS_START_PATTERN = /^---RALPH_STATUS---\s*$/;
|
||||
const RALPH_STATUS_END_PATTERN = /^---END_RALPH_STATUS---\s*$/;
|
||||
const RALPH_STATUS_FIELD_PATTERN = /^STATUS:\s*(IN_PROGRESS|COMPLETE|BLOCKED)\s*$/i;
|
||||
const RALPH_TASKS_COMPLETED_PATTERN = /^TASKS_COMPLETED_THIS_LOOP:\s*(\d+)\s*$/i;
|
||||
const RALPH_FILES_MODIFIED_PATTERN = /^FILES_MODIFIED:\s*(\d+)\s*$/i;
|
||||
const RALPH_TESTS_STATUS_PATTERN = /^TESTS_STATUS:\s*(PASSING|FAILING|NOT_RUN)\s*$/i;
|
||||
const RALPH_WORK_TYPE_PATTERN = /^WORK_TYPE:\s*(IMPLEMENTATION|TESTING|DOCUMENTATION|REFACTORING)\s*$/i;
|
||||
const RALPH_EXIT_SIGNAL_PATTERN = /^EXIT_SIGNAL:\s*(true|false)\s*$/i;
|
||||
const RALPH_RECOMMENDATION_PATTERN = /^RECOMMENDATION:\s*(.+)$/i;
|
||||
|
||||
// ---------- Completion Indicator Patterns ----------
|
||||
|
||||
const COMPLETION_INDICATOR_PATTERNS = [
|
||||
/all\s+(?:tasks?|items?|work)\s+(?:are\s+)?(?:completed?|done|finished)/i,
|
||||
/(?:completed?|finished)\s+all\s+(?:tasks?|items?|work)/i,
|
||||
/nothing\s+(?:left|remaining)\s+to\s+do/i,
|
||||
/no\s+more\s+(?:tasks?|items?|work)/i,
|
||||
/everything\s+(?:is\s+)?(?:completed?|done)/i,
|
||||
];
|
||||
|
||||
// ---------- Priority Pattern ----------
|
||||
|
||||
const TODO_PRIORITY_PATTERN = /^\s*(?:\[.\])?\s*(?:Critical:|Blocker:|Feature:|Improvement:)?\s*\(?(P[012])\)?:?\s*/i;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. New State Properties (ralph-tracker.ts)
|
||||
|
||||
```typescript
|
||||
// Add to RalphTracker class
|
||||
|
||||
// Circuit breaker state tracking
|
||||
private _circuitBreaker: CircuitBreakerStatus;
|
||||
|
||||
// RALPH_STATUS block parsing state
|
||||
private _statusBlockBuffer: string[] = [];
|
||||
private _inStatusBlock: boolean = false;
|
||||
private _lastStatusBlock: RalphStatusBlock | null = null;
|
||||
|
||||
// Dual-condition exit tracking
|
||||
private _completionIndicators: number = 0;
|
||||
private _exitGateMet: boolean = false;
|
||||
|
||||
// Cumulative tracking
|
||||
private _totalFilesModified: number = 0;
|
||||
private _totalTasksCompleted: number = 0;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. New Methods to Implement
|
||||
|
||||
### 5.1 RALPH_STATUS Block Parsing
|
||||
|
||||
```typescript
|
||||
private processStatusBlockLine(line: string): void {
|
||||
const trimmed = line.trim();
|
||||
|
||||
if (RALPH_STATUS_START_PATTERN.test(trimmed)) {
|
||||
this._inStatusBlock = true;
|
||||
this._statusBlockBuffer = [];
|
||||
return;
|
||||
}
|
||||
|
||||
if (this._inStatusBlock && RALPH_STATUS_END_PATTERN.test(trimmed)) {
|
||||
this._inStatusBlock = false;
|
||||
this.parseStatusBlock(this._statusBlockBuffer);
|
||||
this._statusBlockBuffer = [];
|
||||
return;
|
||||
}
|
||||
|
||||
if (this._inStatusBlock) {
|
||||
this._statusBlockBuffer.push(trimmed);
|
||||
}
|
||||
}
|
||||
|
||||
private parseStatusBlock(lines: string[]): void {
|
||||
const block: Partial<RalphStatusBlock> = { parsedAt: Date.now() };
|
||||
|
||||
for (const line of lines) {
|
||||
// Parse each field...
|
||||
}
|
||||
|
||||
if (block.status !== undefined) {
|
||||
this._lastStatusBlock = fullBlock;
|
||||
this.handleStatusBlock(fullBlock);
|
||||
}
|
||||
}
|
||||
|
||||
private handleStatusBlock(block: RalphStatusBlock): void {
|
||||
this._totalFilesModified += block.filesModified;
|
||||
this._totalTasksCompleted += block.tasksCompletedThisLoop;
|
||||
|
||||
const hasProgress = block.filesModified > 0 || block.tasksCompletedThisLoop > 0;
|
||||
this.updateCircuitBreaker(hasProgress, block.testsStatus, block.status);
|
||||
|
||||
if (block.status === 'COMPLETE') {
|
||||
this._completionIndicators++;
|
||||
}
|
||||
|
||||
if (block.exitSignal && this._completionIndicators >= 2) {
|
||||
this._exitGateMet = true;
|
||||
this.emit('exitGateMet', { completionIndicators: this._completionIndicators, exitSignal: true });
|
||||
}
|
||||
|
||||
this.emit('statusBlockDetected', block);
|
||||
}
|
||||
```
|
||||
|
||||
### 5.2 Circuit Breaker Logic
|
||||
|
||||
```typescript
|
||||
private updateCircuitBreaker(
|
||||
hasProgress: boolean,
|
||||
testsStatus: RalphTestsStatus,
|
||||
status: RalphStatusValue
|
||||
): void {
|
||||
const prevState = this._circuitBreaker.state;
|
||||
|
||||
if (hasProgress) {
|
||||
this._circuitBreaker.consecutiveNoProgress = 0;
|
||||
this._circuitBreaker.lastProgressIteration = this._loopState.cycleCount;
|
||||
|
||||
if (this._circuitBreaker.state === 'HALF_OPEN') {
|
||||
this._circuitBreaker.state = 'CLOSED';
|
||||
this._circuitBreaker.reasonCode = 'progress_detected';
|
||||
}
|
||||
} else {
|
||||
this._circuitBreaker.consecutiveNoProgress++;
|
||||
|
||||
if (this._circuitBreaker.state === 'CLOSED') {
|
||||
if (this._circuitBreaker.consecutiveNoProgress >= 3) {
|
||||
this._circuitBreaker.state = 'OPEN';
|
||||
this._circuitBreaker.reasonCode = 'no_progress_open';
|
||||
} else if (this._circuitBreaker.consecutiveNoProgress >= 2) {
|
||||
this._circuitBreaker.state = 'HALF_OPEN';
|
||||
this._circuitBreaker.reasonCode = 'no_progress_warning';
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (prevState !== this._circuitBreaker.state) {
|
||||
this._circuitBreaker.lastTransitionAt = Date.now();
|
||||
this.emit('circuitBreakerUpdate', { ...this._circuitBreaker });
|
||||
}
|
||||
}
|
||||
|
||||
resetCircuitBreaker(): void {
|
||||
this._circuitBreaker = createInitialCircuitBreakerStatus();
|
||||
this._circuitBreaker.reasonCode = 'manual_reset';
|
||||
this.emit('circuitBreakerUpdate', { ...this._circuitBreaker });
|
||||
}
|
||||
```
|
||||
|
||||
### 5.3 Update processLine Method
|
||||
|
||||
```typescript
|
||||
private processLine(line: string): void {
|
||||
const trimmed = line.trim();
|
||||
if (!trimmed) return;
|
||||
|
||||
// NEW: Check for RALPH_STATUS block
|
||||
this.processStatusBlockLine(trimmed);
|
||||
|
||||
// NEW: Check for completion indicators
|
||||
this.detectCompletionIndicators(trimmed);
|
||||
|
||||
// EXISTING: Rest of the detection methods...
|
||||
this.detectCompletionPhrase(trimmed);
|
||||
this.detectAllTasksComplete(trimmed);
|
||||
this.detectTaskCompletion(trimmed);
|
||||
this.detectLoopStatus(trimmed);
|
||||
this.detectTodoItems(trimmed);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. New Events to Add
|
||||
|
||||
```typescript
|
||||
export interface RalphTrackerEvents {
|
||||
// Existing events
|
||||
loopUpdate: (state: RalphTrackerState) => void;
|
||||
todoUpdate: (todos: RalphTodoItem[]) => void;
|
||||
completionDetected: (phrase: string) => void;
|
||||
enabled: () => void;
|
||||
|
||||
// New events
|
||||
statusBlockDetected: (block: RalphStatusBlock) => void;
|
||||
circuitBreakerUpdate: (status: CircuitBreakerStatus) => void;
|
||||
exitGateMet: (data: { completionIndicators: number; exitSignal: boolean }) => void;
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. Server Integration (server.ts)
|
||||
|
||||
```typescript
|
||||
// Add new SSE event handlers in setupSessionListeners()
|
||||
|
||||
session.on('ralphStatusBlockDetected', (block: RalphStatusBlock) => {
|
||||
this.broadcast('session:ralphStatusUpdate', { sessionId: session.id, block });
|
||||
});
|
||||
|
||||
session.on('ralphCircuitBreakerUpdate', (status: CircuitBreakerStatus) => {
|
||||
this.broadcast('session:circuitBreakerUpdate', { sessionId: session.id, status });
|
||||
});
|
||||
|
||||
session.on('ralphExitGateMet', (data) => {
|
||||
this.broadcast('session:exitGateMet', { sessionId: session.id, ...data });
|
||||
});
|
||||
|
||||
// Add API endpoint for circuit breaker reset
|
||||
this.app.post('/api/sessions/:id/ralph-circuit-breaker/reset', async (req) => {
|
||||
const session = this.sessions.get(req.params.id);
|
||||
if (!session) return { success: false, error: 'Session not found' };
|
||||
|
||||
session.ralphTracker?.resetCircuitBreaker();
|
||||
return { success: true };
|
||||
});
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 8. Frontend Changes (app.js)
|
||||
|
||||
### New SSE Event Listeners
|
||||
|
||||
```javascript
|
||||
this.eventSource.addEventListener('session:ralphStatusUpdate', (e) => {
|
||||
const data = JSON.parse(e.data);
|
||||
this.updateRalphStatusBlock(data.sessionId, data.block);
|
||||
});
|
||||
|
||||
this.eventSource.addEventListener('session:circuitBreakerUpdate', (e) => {
|
||||
const data = JSON.parse(e.data);
|
||||
this.updateCircuitBreaker(data.sessionId, data.status);
|
||||
});
|
||||
```
|
||||
|
||||
### New Rendering Methods
|
||||
|
||||
```javascript
|
||||
updateRalphStatusBlock(sessionId, block) {
|
||||
// Store and render status block
|
||||
}
|
||||
|
||||
renderRalphStatusBlock(block) {
|
||||
// Render STATUS, WORK_TYPE, TESTS_STATUS, RECOMMENDATION
|
||||
}
|
||||
|
||||
updateCircuitBreaker(sessionId, status) {
|
||||
// Store and render circuit breaker state
|
||||
}
|
||||
|
||||
renderCircuitBreaker(status) {
|
||||
// Render badge: green (CLOSED), yellow (HALF_OPEN), red (OPEN)
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 9. Implementation Order
|
||||
|
||||
| Step | Task | Time |
|
||||
|------|------|------|
|
||||
| 1 | Add type definitions to `types.ts` | 30 min |
|
||||
| 2 | Add regex patterns to `ralph-tracker.ts` | 30 min |
|
||||
| 3 | Add state properties to RalphTracker class | 15 min |
|
||||
| 4 | Implement RALPH_STATUS parsing methods | 1.5 hr |
|
||||
| 5 | Implement circuit breaker logic | 1 hr |
|
||||
| 6 | Implement completion indicators | 30 min |
|
||||
| 7 | Update events interface | 15 min |
|
||||
| 8 | Add server SSE handlers and API endpoint | 45 min |
|
||||
| 9 | Add frontend event listeners and rendering | 1 hr |
|
||||
| 10 | Add CSS styles | 30 min |
|
||||
| 11 | Update HTML structure | 15 min |
|
||||
| 12 | Write unit tests | 1.5 hr |
|
||||
|
||||
**Total: ~8 hours**
|
||||
|
||||
---
|
||||
|
||||
## 10. Files to Modify
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/types.ts` | Add RalphStatusBlock, CircuitBreakerStatus, helper functions |
|
||||
| `src/ralph-tracker.ts` | Add patterns, state, parsing methods, circuit breaker |
|
||||
| `src/web/server.ts` | Add SSE handlers, circuit breaker reset endpoint |
|
||||
| `src/web/public/app.js` | Add event listeners, rendering methods |
|
||||
| `src/web/public/styles.css` | Add status block and circuit breaker styles |
|
||||
| `src/web/public/index.html` | Add UI elements to Ralph panel |
|
||||
| `test/ralph-tracker.test.ts` | Add tests for new functionality |
|
||||
|
||||
---
|
||||
|
||||
## 11. Test Cases to Add
|
||||
|
||||
1. **RALPH_STATUS Parsing**
|
||||
- Parse valid status block with all fields
|
||||
- Parse block with missing optional fields
|
||||
- Ignore malformed blocks
|
||||
- Handle multiple blocks in sequence
|
||||
|
||||
2. **Circuit Breaker State Transitions**
|
||||
- CLOSED → HALF_OPEN on 2 no-progress
|
||||
- HALF_OPEN → OPEN on 3 no-progress
|
||||
- HALF_OPEN → CLOSED on progress
|
||||
- Manual reset from OPEN
|
||||
|
||||
3. **Dual-Condition Exit Gate**
|
||||
- Exit when indicators >= 2 AND exitSignal = true
|
||||
- No exit when indicators >= 2 but exitSignal = false
|
||||
- No exit when exitSignal = true but indicators < 2
|
||||
|
||||
4. **Integration Tests**
|
||||
- SSE events broadcast correctly
|
||||
- UI updates on status block detection
|
||||
- Circuit breaker badge updates
|
||||
@@ -0,0 +1,385 @@
|
||||
# Respawn Controller Idle Detection Improvement Plan
|
||||
|
||||
## Executive Summary
|
||||
|
||||
The current respawn controller relies primarily on **parsing terminal output** to detect idle states. This approach is fragile and leads to false positives/negatives (e.g., the w3-reddit-analyse session).
|
||||
|
||||
**Key insight**: Claude Code provides **direct, authoritative signals** via hooks and files that definitively indicate session state. We're receiving some of these signals but not using them for idle detection!
|
||||
|
||||
---
|
||||
|
||||
## Current Detection Layers (What We Have)
|
||||
|
||||
| Layer | Signal | Source | Reliability |
|
||||
|-------|--------|--------|-------------|
|
||||
| 1 | Completion message ("Worked for Xm Xs") | Terminal parsing | Medium - can miss edge cases |
|
||||
| 2 | Output silence (configurable duration) | Terminal activity | Low - Claude can be processing silently |
|
||||
| 3 | Token stability | Terminal parsing | Low - tokens don't change during I/O waits |
|
||||
| 4 | Working pattern absence | Terminal parsing | Medium - patterns can be missed |
|
||||
| 5 | AI idle check | Spawned Claude CLI | High but slow (90s timeout) |
|
||||
|
||||
**Problem**: All layers depend on **parsing terminal output**, which is inherently unreliable.
|
||||
|
||||
---
|
||||
|
||||
## Available Claude Code Signals (Not Fully Utilized)
|
||||
|
||||
### 1. `Stop` Hook ⭐ CRITICAL - DEFINITIVE SIGNAL
|
||||
|
||||
**What it is**: Fires when the main Claude Code agent **finishes responding**.
|
||||
|
||||
**From docs**: "Runs when the main Claude Code agent has finished responding. Does not run if the stoppage occurred due to a user interrupt."
|
||||
|
||||
**Current status**: We receive it via `/api/hook-event` but **don't use it for idle detection**!
|
||||
|
||||
**Input received**:
|
||||
```json
|
||||
{
|
||||
"session_id": "abc123",
|
||||
"transcript_path": "~/.claude/projects/.../00893aaf.jsonl",
|
||||
"hook_event_name": "Stop",
|
||||
"stop_hook_active": true // Important for preventing loops
|
||||
}
|
||||
```
|
||||
|
||||
**Action needed**: The `Stop` hook should be the **PRIMARY** idle detection signal. When Claude fires Stop, the agent has definitively finished its response cycle.
|
||||
|
||||
### 2. `idle_prompt` Notification ⭐ HIGH VALUE
|
||||
|
||||
**What it is**: Fires after **60+ seconds of idle time** when Claude is waiting for user input.
|
||||
|
||||
**From docs**: "When Claude is waiting for user input (after 60+ seconds of idle time)"
|
||||
|
||||
**Current status**: We receive it but only forward it to the UI for notification display.
|
||||
|
||||
**Action needed**: Use `idle_prompt` as a **definitive confirmation** that Claude is idle. If we receive this, there's no need for AI idle checks or output silence timers.
|
||||
|
||||
### 3. Transcript JSONL File ⭐ HIGH VALUE
|
||||
|
||||
**What it is**: Complete conversation history at `~/.claude/projects/{project-hash}/{session-id}.jsonl`
|
||||
|
||||
**Current status**: We already watch subagent transcripts but **not the main session transcript**.
|
||||
|
||||
**Data available**:
|
||||
- Every message (user, assistant, system)
|
||||
- Every tool call with inputs/outputs
|
||||
- Progress events
|
||||
- Structured, parseable JSON
|
||||
|
||||
**Action needed**:
|
||||
- Monitor the main transcript file (path provided in every hook input)
|
||||
- Parse the last few entries to detect:
|
||||
- Tool completion
|
||||
- Assistant message completion
|
||||
- Error states
|
||||
- Plan mode prompts
|
||||
|
||||
### 4. `PostToolUse` Hook - Tool Completion Tracking
|
||||
|
||||
**What it is**: Fires immediately after any tool completes successfully.
|
||||
|
||||
**Use case**: Track exactly when tools finish to understand execution flow.
|
||||
|
||||
**Current status**: Not implemented.
|
||||
|
||||
**Action needed**: Add PostToolUse hooks to track tool completion events.
|
||||
|
||||
### 5. `SubagentStop` Hook - Background Agent Completion
|
||||
|
||||
**What it is**: Fires when a subagent (Task tool) finishes responding.
|
||||
|
||||
**Current status**: Not implemented in hooks config (we watch JSONL files separately).
|
||||
|
||||
**Action needed**: Add to hooks config for redundant detection.
|
||||
|
||||
### 6. `permission_prompt` and `elicitation_dialog` - Blocking State Detection
|
||||
|
||||
**What it is**: Fires when Claude needs user input (permission or question).
|
||||
|
||||
**Current status**: We receive and use for auto-accept blocking.
|
||||
|
||||
**Enhancement**: Use as definitive "Claude is NOT idle - it's waiting for user action".
|
||||
|
||||
---
|
||||
|
||||
## Proposed Architecture: Multi-Signal Idle Detection
|
||||
|
||||
### New Detection Hierarchy
|
||||
|
||||
```
|
||||
Priority 1 (Definitive):
|
||||
└── Stop hook received → CONFIRMED IDLE
|
||||
└── idle_prompt received → CONFIRMED IDLE (60s+ idle)
|
||||
|
||||
Priority 2 (Blocking):
|
||||
└── permission_prompt received → NOT IDLE (waiting for permission)
|
||||
└── elicitation_dialog received → NOT IDLE (waiting for answer)
|
||||
└── Working patterns in terminal → NOT IDLE
|
||||
|
||||
Priority 3 (Supporting):
|
||||
└── Transcript analysis → Check last entries for completion
|
||||
└── Output silence + token stability → Weak idle signal
|
||||
|
||||
Priority 4 (Fallback):
|
||||
└── AI idle check → Only if no definitive signals after timeout
|
||||
```
|
||||
|
||||
### State Machine Changes
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────┐
|
||||
│ │
|
||||
▼ │
|
||||
┌─────────────────────┐ │
|
||||
│ WATCHING │◄──────────────────────────────┤
|
||||
└─────────────────────┘ │
|
||||
│ │ │
|
||||
│ │ Stop hook or idle_prompt │
|
||||
│ └────────────────────────┐ │
|
||||
│ ▼ │
|
||||
│ Output silence ┌────────────┐ │
|
||||
│ (no definitive signals) │ HOOK_IDLE │───────┤
|
||||
│ └────────────┘ │
|
||||
▼ (skip AI check) │
|
||||
┌────────────────────┐ │
|
||||
│ CONFIRMING_IDLE │ │
|
||||
└────────────────────┘ │
|
||||
│ │
|
||||
│ Silence confirmed │
|
||||
▼ │
|
||||
┌────────────────────┐ │
|
||||
│ AI_CHECKING │──── IDLE verdict ─────────────┤
|
||||
└────────────────────┘ │
|
||||
│ │
|
||||
│ WORKING verdict │
|
||||
└─────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### New State: `hook_idle`
|
||||
|
||||
When a definitive hook signal is received:
|
||||
1. Skip AI idle check entirely (saves time and API calls)
|
||||
2. Short confirmation period (2-3s) to handle race conditions
|
||||
3. Proceed directly to respawn sequence
|
||||
|
||||
---
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
### Phase 1: Use Stop Hook for Idle Detection ✅ COMPLETED
|
||||
|
||||
**Files modified**:
|
||||
- `src/respawn-controller.ts`
|
||||
- `src/web/server.ts`
|
||||
- `test/respawn-controller.test.ts`
|
||||
|
||||
**Changes implemented**:
|
||||
1. Added `stopHookReceived`, `stopHookTime`, `idlePromptReceived`, `idlePromptTime` fields to `DetectionStatus`
|
||||
2. Added `hookConfirmTimer` for short confirmation after hook signal (3s)
|
||||
3. Added `signalStopHook()` method:
|
||||
- Sets `stopHookReceived = true` and timestamp
|
||||
- Cancels any running AI check (hook is definitive)
|
||||
- Starts 3s confirmation timer
|
||||
- If no new output during confirmation → triggers respawn cycle
|
||||
4. Added `signalIdlePrompt()` method:
|
||||
- Sets `idlePromptReceived = true` and timestamp
|
||||
- Immediately confirms idle (skips confirmation timer - 60s+ already proven)
|
||||
5. Added `resetHookState()` to clear hook flags on:
|
||||
- Controller start
|
||||
- Working patterns detected
|
||||
- Cycle completion
|
||||
6. Updated server.ts `/api/hook-event` endpoint to call:
|
||||
- `controller.signalStopHook()` for `stop` events
|
||||
- `controller.signalIdlePrompt()` for `idle_prompt` events
|
||||
7. Updated `getDetectionStatus()`:
|
||||
- Returns hook signal states
|
||||
- Sets confidence to 100% when hook received
|
||||
- Updates statusText to show hook status
|
||||
|
||||
**Tests added** (9 new tests in `RespawnController Hook-Based Idle Detection` describe block):
|
||||
- `should expose signalStopHook method`
|
||||
- `should expose signalIdlePrompt method`
|
||||
- `should set stopHookReceived in detection status when Stop hook signaled`
|
||||
- `should include hook status in statusText when Stop hook received`
|
||||
- `should trigger respawn cycle after Stop hook confirmation`
|
||||
- `should immediately confirm idle when idle_prompt signaled (skip confirmation)`
|
||||
- `should cancel Stop hook confirmation if working patterns detected`
|
||||
- `should ignore Stop hook when not in watching state`
|
||||
- `should have 100% confidence when hook signal is received`
|
||||
|
||||
**Detection status update** (implemented):
|
||||
```typescript
|
||||
interface DetectionStatus {
|
||||
/** Layer 0: Stop hook received (highest priority - definitive signal) */
|
||||
stopHookReceived: boolean;
|
||||
stopHookTime: number | null;
|
||||
/** Layer 0: idle_prompt notification received (definitive signal) */
|
||||
idlePromptReceived: boolean;
|
||||
idlePromptTime: number | null;
|
||||
|
||||
// Existing fields...
|
||||
}
|
||||
```
|
||||
|
||||
### Phase 2: Use idle_prompt for Definitive Idle ✅ COMPLETED (in Phase 1)
|
||||
|
||||
**Already implemented in Phase 1**:
|
||||
1. `signalIdlePrompt()` method sets `idlePromptReceived = true`
|
||||
2. Immediately calls `onIdleConfirmed()` - skips all other detection
|
||||
3. Server.ts calls `controller.signalIdlePrompt()` when `idle_prompt` event received
|
||||
4. 60s+ of Claude waiting = definitive idle signal
|
||||
|
||||
### Phase 3: Transcript File Monitoring ✅ COMPLETED
|
||||
|
||||
**New file**: `src/transcript-watcher.ts`
|
||||
|
||||
**Functionality implemented**:
|
||||
1. Watch the session's transcript JSONL file using `fs.watch()`
|
||||
2. Parse new entries as they're appended (incremental reading from last position)
|
||||
3. Detect:
|
||||
- `result` entry → `transcript:complete` event (isComplete = true)
|
||||
- `tool_use` content block → `transcript:tool_start` event
|
||||
- `tool_result` content block → `transcript:tool_end` event
|
||||
- `AskUserQuestion` or `ExitPlanMode` tools → `transcript:plan_mode` event
|
||||
- Error conditions in result entries
|
||||
4. Emit structured events consumed by respawn controller
|
||||
|
||||
**Integration implemented**:
|
||||
- `transcript_path` added to allowed hook data fields in `sanitizeHookData()`
|
||||
- `transcriptWatchers` Map added to WebServer for per-session watchers
|
||||
- `startTranscriptWatcher()` creates watcher and wires up events:
|
||||
- `transcript:complete` → `controller.signalTranscriptComplete()`
|
||||
- `transcript:plan_mode` → `controller.signalTranscriptPlanMode()`
|
||||
- `stopTranscriptWatcher()` cleans up on session cleanup
|
||||
- Hook events with `transcript_path` automatically start watching
|
||||
|
||||
**RespawnController methods added**:
|
||||
- `signalTranscriptComplete()` - Supporting signal that can accelerate idle detection
|
||||
- `signalTranscriptPlanMode()` - Cancels auto-accept timer (like elicitation)
|
||||
|
||||
**Tests added** (13 tests in `test/transcript-watcher.test.ts`):
|
||||
- Initialization tests
|
||||
- File watching tests (existing file, non-existent file, stop, updatePath)
|
||||
- Entry processing tests (user entry, result entry, tool execution, plan mode, errors)
|
||||
- State management tests
|
||||
|
||||
### Phase 4: Enhanced Hook Configuration
|
||||
|
||||
**Update `src/hooks-config.ts`**:
|
||||
|
||||
```typescript
|
||||
export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
|
||||
return {
|
||||
hooks: {
|
||||
Notification: [
|
||||
{ matcher: 'idle_prompt', hooks: [...] },
|
||||
{ matcher: 'permission_prompt', hooks: [...] },
|
||||
{ matcher: 'elicitation_dialog', hooks: [...] },
|
||||
],
|
||||
Stop: [{ hooks: [...] }],
|
||||
// NEW: Add these
|
||||
PostToolUse: [
|
||||
{ matcher: '*', hooks: [...] } // Track all tool completions
|
||||
],
|
||||
SubagentStop: [{ hooks: [...] }],
|
||||
PreCompact: [
|
||||
{ matcher: '*', hooks: [...] } // Track compaction
|
||||
],
|
||||
},
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
### Phase 5: Confidence Scoring Overhaul
|
||||
|
||||
Replace current confidence calculation with weighted signals:
|
||||
|
||||
```typescript
|
||||
function calculateConfidence(): number {
|
||||
let confidence = 0;
|
||||
|
||||
// Definitive signals (100% confidence)
|
||||
if (this.stopHookReceived) confidence = 100;
|
||||
if (this.idlePromptReceived) confidence = 100;
|
||||
|
||||
// Blocking signals (0% confidence)
|
||||
if (this.permissionPromptReceived) return 0;
|
||||
if (this.elicitationReceived) return 0;
|
||||
if (this.workingPatternRecent) return 0;
|
||||
|
||||
// Supporting signals (build up to ~80%)
|
||||
if (confidence < 100) {
|
||||
if (this.outputSilent) confidence += 30;
|
||||
if (this.tokensStable) confidence += 20;
|
||||
if (this.transcriptShowsCompletion) confidence += 30;
|
||||
}
|
||||
|
||||
return Math.min(100, confidence);
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Expected Benefits
|
||||
|
||||
| Metric | Current | After Implementation |
|
||||
|--------|---------|---------------------|
|
||||
| False positive rate | ~15-20% | <5% |
|
||||
| Detection latency | 10-90s (AI check) | 3-5s (hook-based) |
|
||||
| API calls for AI check | Every idle detection | Only when hooks unavailable |
|
||||
| Reliability | Medium | High (definitive signals) |
|
||||
|
||||
---
|
||||
|
||||
## Testing Strategy
|
||||
|
||||
### Unit Tests
|
||||
1. `Stop` hook triggers immediate idle confirmation
|
||||
2. `idle_prompt` skips all other detection
|
||||
3. `permission_prompt` blocks idle detection
|
||||
4. Transcript parsing correctly identifies completion
|
||||
5. Fallback to AI check when no hooks received
|
||||
|
||||
### Integration Tests
|
||||
1. End-to-end with real Claude session
|
||||
2. Hook event delivery and handling
|
||||
3. Transcript file monitoring
|
||||
4. Race condition handling
|
||||
|
||||
### Scenarios to Test
|
||||
1. Normal completion → Stop hook → respawn
|
||||
2. Long-running task → idle_prompt → respawn
|
||||
3. Permission needed → wait for user action
|
||||
4. AskUserQuestion → wait for user answer
|
||||
5. Plan mode → auto-accept → continue
|
||||
6. Hooks disabled/unavailable → fallback to AI check
|
||||
|
||||
---
|
||||
|
||||
## Migration Path
|
||||
|
||||
1. **Implement Phase 1** - Stop hook detection (low risk, high value)
|
||||
2. **Deploy and monitor** - Verify Stop hooks are reliable
|
||||
3. **Implement Phase 2** - idle_prompt (simple addition)
|
||||
4. **Implement Phase 3** - Transcript monitoring (more complex)
|
||||
5. **Implement Phase 4** - Enhanced hooks (optional, for completeness)
|
||||
6. **Implement Phase 5** - Refactor confidence scoring
|
||||
|
||||
---
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. **Stop hook reliability**: Does it fire 100% of the time? Edge cases?
|
||||
2. **Transcript file location**: Always at the path in hook input?
|
||||
3. **Hook delivery latency**: How quickly do hooks fire after state change?
|
||||
4. **Race conditions**: What if Stop hook and new work happen simultaneously?
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
- [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)
|
||||
- [Agent SDK Documentation](https://platform.claude.com/docs/en/agent-sdk/overview)
|
||||
- Current implementation: `src/respawn-controller.ts`
|
||||
- Hooks config: `src/hooks-config.ts`
|
||||
- Subagent watcher: `src/subagent-watcher.ts`
|
||||
@@ -0,0 +1,172 @@
|
||||
# Run Summary Feature - Implementation Plan
|
||||
|
||||
## Overview
|
||||
|
||||
The Run Summary feature provides users with a consolidated view of what happened in their session while they were away. It tracks significant events, issues, and statistics, presenting them in an easy-to-digest format.
|
||||
|
||||
## Data Structures
|
||||
|
||||
### RunSummaryEventType
|
||||
```typescript
|
||||
type RunSummaryEventType =
|
||||
| 'session_started'
|
||||
| 'session_stopped'
|
||||
| 'respawn_cycle_started'
|
||||
| 'respawn_cycle_completed'
|
||||
| 'respawn_state_change'
|
||||
| 'error'
|
||||
| 'warning'
|
||||
| 'token_milestone'
|
||||
| 'auto_compact'
|
||||
| 'auto_clear'
|
||||
| 'idle_detected'
|
||||
| 'working_detected'
|
||||
| 'ralph_completion'
|
||||
| 'ai_check_result'
|
||||
| 'hook_event'
|
||||
| 'state_stuck';
|
||||
```
|
||||
|
||||
### RunSummaryEvent
|
||||
```typescript
|
||||
interface RunSummaryEvent {
|
||||
id: string;
|
||||
timestamp: number;
|
||||
type: RunSummaryEventType;
|
||||
severity: 'info' | 'warning' | 'error' | 'success';
|
||||
title: string;
|
||||
details?: string;
|
||||
metadata?: Record<string, unknown>;
|
||||
}
|
||||
```
|
||||
|
||||
### RunSummary
|
||||
```typescript
|
||||
interface RunSummary {
|
||||
sessionId: string;
|
||||
sessionName: string;
|
||||
startedAt: number;
|
||||
lastUpdatedAt: number;
|
||||
events: RunSummaryEvent[];
|
||||
stats: {
|
||||
totalRespawnCycles: number;
|
||||
totalTokensUsed: number;
|
||||
peakTokens: number;
|
||||
totalTimeActiveMs: number;
|
||||
totalTimeIdleMs: number;
|
||||
errorCount: number;
|
||||
warningCount: number;
|
||||
aiCheckCount: number;
|
||||
lastIdleAt: number | null;
|
||||
lastWorkingAt: number | null;
|
||||
stateTransitions: number;
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
## Files to Create/Modify
|
||||
|
||||
### 1. `src/run-summary.ts` (NEW)
|
||||
- `RunSummaryTracker` class
|
||||
- Event tracking and aggregation
|
||||
- Statistics calculation
|
||||
- Max 1000 events per session (FIFO trimming)
|
||||
|
||||
### 2. `src/types.ts` (MODIFY)
|
||||
- Add `RunSummaryEvent`, `RunSummaryEventType`, `RunSummary` interfaces
|
||||
- Add `RunSummaryEventSeverity` type
|
||||
|
||||
### 3. `src/web/server.ts` (MODIFY)
|
||||
- Create `RunSummaryTracker` per session
|
||||
- Subscribe to session events and forward to tracker
|
||||
- Subscribe to respawn controller events
|
||||
- Add API endpoint: `GET /api/sessions/:id/run-summary`
|
||||
- Broadcast `session:runSummaryUpdate` SSE event
|
||||
|
||||
### 4. `src/web/public/app.js` (MODIFY)
|
||||
- Add "Run Summary" button to session header
|
||||
- Create modal to display summary
|
||||
- Handle `session:runSummaryUpdate` SSE event
|
||||
- Timeline view for events
|
||||
- Stats cards at top
|
||||
|
||||
### 5. `src/web/public/index.html` (MODIFY)
|
||||
- Add modal HTML structure for run summary
|
||||
|
||||
### 6. `src/web/public/styles.css` (MODIFY)
|
||||
- Styles for run summary modal and timeline
|
||||
|
||||
## Event Sources
|
||||
|
||||
| Event Type | Source | Trigger |
|
||||
|------------|--------|---------|
|
||||
| session_started | Session | `startInteractive()` / `startShell()` |
|
||||
| session_stopped | Session | `stop()` |
|
||||
| respawn_cycle_started | RespawnController | State → `sending_update` |
|
||||
| respawn_cycle_completed | RespawnController | State → `watching` (after cycle) |
|
||||
| respawn_state_change | RespawnController | Any state transition |
|
||||
| error | Various | Errors caught in try/catch |
|
||||
| warning | RunSummaryTracker | State stuck > 5min, high tokens |
|
||||
| token_milestone | Session | Every 50k tokens |
|
||||
| auto_compact | Session | `autoCompact` event |
|
||||
| auto_clear | Session | `autoClear` event |
|
||||
| idle_detected | Session | `idle` event |
|
||||
| working_detected | Session | `working` event |
|
||||
| ralph_completion | RalphTracker | `completionDetected` event |
|
||||
| ai_check_result | RespawnController | AI check completes |
|
||||
| hook_event | Server | `/api/hook-event` endpoint |
|
||||
| state_stuck | RunSummaryTracker | Same state > 10min |
|
||||
|
||||
## API Endpoint
|
||||
|
||||
### GET /api/sessions/:id/run-summary
|
||||
Returns the full run summary for a session.
|
||||
|
||||
Response:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"summary": {
|
||||
"sessionId": "...",
|
||||
"sessionName": "...",
|
||||
"startedAt": 1234567890,
|
||||
"lastUpdatedAt": 1234567890,
|
||||
"events": [...],
|
||||
"stats": {...}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## UI Design
|
||||
|
||||
### Summary Modal
|
||||
- Header: Session name, duration, status indicator
|
||||
- Stats Cards Row:
|
||||
- Respawn Cycles: count
|
||||
- Tokens Used: peak / current
|
||||
- Active Time: formatted duration
|
||||
- Issues: errors + warnings count
|
||||
- Timeline:
|
||||
- Vertical timeline of events
|
||||
- Color-coded by severity (green=success, blue=info, yellow=warning, red=error)
|
||||
- Expandable details
|
||||
- Filter by event type
|
||||
- Footer: "Close" button
|
||||
|
||||
## Implementation Steps
|
||||
|
||||
1. Add types to `types.ts`
|
||||
2. Create `run-summary.ts` with RunSummaryTracker class
|
||||
3. Integrate tracker with server.ts (create per session, wire events)
|
||||
4. Add API endpoint
|
||||
5. Add frontend modal and button
|
||||
6. Test with live session
|
||||
|
||||
## Storage
|
||||
|
||||
Run summaries are kept in memory only (not persisted to disk) since:
|
||||
- They're session-specific and regenerated on session start
|
||||
- Persisting thousands of events would bloat state.json
|
||||
- Server restart = fresh session anyway
|
||||
|
||||
If persistence is needed later, could add to `state-inner.json` with per-session limits.
|
||||
@@ -0,0 +1,387 @@
|
||||
# Codeman TypeScript Improvement Suggestions
|
||||
|
||||
**Generated**: February 2026
|
||||
**Based on**: Research into TypeScript best practices (2024-2025) and codebase analysis
|
||||
|
||||
---
|
||||
|
||||
## 🔴 High Priority (Low effort, high impact)
|
||||
|
||||
### 1. Use the Already-Installed Zod for API Validation
|
||||
|
||||
Zod v4.3.6 is in `package.json` but **never imported**. API routes use unsafe type assertions:
|
||||
|
||||
```typescript
|
||||
// Current (unsafe)
|
||||
const body = req.body as CreateSessionRequest;
|
||||
|
||||
// Recommended
|
||||
const result = CreateSessionSchema.safeParse(req.body);
|
||||
if (!result.success) return createErrorResponse(ApiErrorCode.INVALID_INPUT, ...);
|
||||
```
|
||||
|
||||
**Impact**: Prevents runtime errors from malformed client requests.
|
||||
|
||||
**Files to update**: `src/web/server.ts` (all POST/PUT routes)
|
||||
|
||||
---
|
||||
|
||||
### 2. Add `assertNever` for Exhaustive Switch Checking
|
||||
|
||||
Switch statements on union types (e.g., `respawn-controller.ts:1072`, `ralph-tracker.ts:2088`) lack exhaustive checking. Adding new union members won't cause compile errors.
|
||||
|
||||
```typescript
|
||||
// Add to src/utils/type-safety.ts
|
||||
export function assertNever(x: never, message?: string): never {
|
||||
throw new Error(message ?? `Unexpected value: ${JSON.stringify(x)}`);
|
||||
}
|
||||
|
||||
// Usage in switch statements
|
||||
switch (status) {
|
||||
case 'idle': return handleIdle();
|
||||
case 'busy': return handleBusy();
|
||||
case 'stopped': return handleStopped();
|
||||
case 'error': return handleError();
|
||||
default: return assertNever(status);
|
||||
}
|
||||
```
|
||||
|
||||
**Impact**: Compile-time guarantee all cases are handled.
|
||||
|
||||
**Files affected**: `respawn-controller.ts`, `ralph-tracker.ts`, any file with switch on union types
|
||||
|
||||
---
|
||||
|
||||
### 3. Standardize `createErrorResponse` Usage
|
||||
|
||||
Currently only used in 2 files despite being a good pattern. Many routes still use ad-hoc error responses.
|
||||
|
||||
**Impact**: Consistent API error format across all endpoints.
|
||||
|
||||
---
|
||||
|
||||
## 🟡 Medium Priority (Medium effort, significant benefit)
|
||||
|
||||
### 4. Convert `ApiResponse<T>` to Discriminated Union
|
||||
|
||||
Current interface has optional properties; discriminated union enables better narrowing:
|
||||
|
||||
```typescript
|
||||
// Current (types.ts)
|
||||
interface ApiResponse<T> { success: boolean; error?: string; data?: T; }
|
||||
|
||||
// Better
|
||||
type ApiResponse<T> =
|
||||
| { success: true; data: T }
|
||||
| { success: false; error: string; errorCode: ApiErrorCode };
|
||||
|
||||
// Usage with exhaustive checking
|
||||
function handleResponse<T>(response: ApiResponse<T>): T {
|
||||
if (response.success) {
|
||||
return response.data; // TypeScript knows data exists
|
||||
} else {
|
||||
throw new Error(response.error); // TypeScript knows error exists
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 5. Add Branded Types for Token Counts
|
||||
|
||||
Prevents mixing input/output tokens in calculations:
|
||||
|
||||
```typescript
|
||||
// src/types/branded.ts
|
||||
type Brand<K, T extends string> = K & { readonly __brand: T };
|
||||
|
||||
export type InputTokens = Brand<number, 'InputTokens'>;
|
||||
export type OutputTokens = Brand<number, 'OutputTokens'>;
|
||||
export type TokenCount = Brand<number, 'TokenCount'>;
|
||||
export type Milliseconds = Brand<number, 'Milliseconds'>;
|
||||
|
||||
// Constructor functions
|
||||
export function inputTokens(value: number): InputTokens {
|
||||
if (value < 0) throw new Error('Token count cannot be negative');
|
||||
return value as InputTokens;
|
||||
}
|
||||
```
|
||||
|
||||
**Use cases**:
|
||||
- Token counts (`_totalInputTokens`, `_totalOutputTokens`)
|
||||
- Timeout values (`idleTimeoutMs`, `completionConfirmMs`, `noOutputTimeoutMs`)
|
||||
- IDs (`SessionId`, `TaskId`, `CycleId`)
|
||||
|
||||
---
|
||||
|
||||
### 6. Dependency Injection for Core Services
|
||||
|
||||
Replace hidden singleton dependencies with constructor injection for better testability:
|
||||
|
||||
```typescript
|
||||
// Current: Hidden dependencies
|
||||
export class RalphLoop extends EventEmitter {
|
||||
constructor() {
|
||||
this.sessionManager = getSessionManager();
|
||||
this.store = getStore();
|
||||
}
|
||||
}
|
||||
|
||||
// Better: Explicit dependencies
|
||||
export interface RalphLoopDeps {
|
||||
sessionManager: SessionManager;
|
||||
taskQueue: TaskQueue;
|
||||
store: StateStore;
|
||||
}
|
||||
|
||||
export class RalphLoop extends EventEmitter {
|
||||
constructor(deps: RalphLoopDeps, options?: RalphLoopOptions) {
|
||||
this.sessionManager = deps.sessionManager;
|
||||
// ...
|
||||
}
|
||||
}
|
||||
|
||||
// Production factory
|
||||
export function createRalphLoop(options?: RalphLoopOptions): RalphLoop {
|
||||
return new RalphLoop({
|
||||
sessionManager: getSessionManager(),
|
||||
taskQueue: getTaskQueue(),
|
||||
store: getStore(),
|
||||
}, options);
|
||||
}
|
||||
```
|
||||
|
||||
**Start with**: `RalphLoop` (has the most dependencies)
|
||||
|
||||
**Benefits**: Easier testing, explicit dependencies, SOLID compliance
|
||||
|
||||
---
|
||||
|
||||
### 7. Enforce Consistent `import type` Usage
|
||||
|
||||
Mixed usage across codebase. Add ESLint rule:
|
||||
|
||||
```json
|
||||
{
|
||||
"rules": {
|
||||
"@typescript-eslint/consistent-type-imports": ["error", {
|
||||
"prefer": "type-imports",
|
||||
"fixStyle": "separate-type-imports"
|
||||
}]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Benefits**: Reduced bundle size, better tree-shaking, cleaner separation
|
||||
|
||||
---
|
||||
|
||||
### 8. Add Circular Dependency Detection
|
||||
|
||||
```bash
|
||||
npm install -D dpdm
|
||||
```
|
||||
|
||||
Add to `package.json`:
|
||||
```json
|
||||
{
|
||||
"scripts": {
|
||||
"check:circular": "dpdm --no-warning --no-tree src/index.ts"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Potential risk areas identified**:
|
||||
- `ralph-loop.ts` → `session-manager.ts` → `session.ts`
|
||||
- `respawn-controller.ts` → `session.ts` → `ai-idle-checker.ts`
|
||||
|
||||
---
|
||||
|
||||
## 🟢 Lower Priority (Higher effort, situational benefit)
|
||||
|
||||
### 9. Apply `as const satisfies` to Default Configs
|
||||
|
||||
Preserves literal types while validating structure:
|
||||
|
||||
```typescript
|
||||
// Current
|
||||
export const DEFAULT_NICE_CONFIG: NiceConfig = {
|
||||
enabled: false,
|
||||
niceValue: 10,
|
||||
};
|
||||
// niceValue is type: number
|
||||
|
||||
// Better
|
||||
export const DEFAULT_NICE_CONFIG = {
|
||||
enabled: false,
|
||||
niceValue: 10,
|
||||
} as const satisfies NiceConfig;
|
||||
// niceValue is type: 10 (literal)
|
||||
```
|
||||
|
||||
**Files**: `types.ts`, `respawn-controller.ts` (DEFAULT_CONFIG)
|
||||
|
||||
---
|
||||
|
||||
### 10. Create Custom Error Class Hierarchy
|
||||
|
||||
Replace string-based errors with typed errors:
|
||||
|
||||
```typescript
|
||||
// src/errors.ts
|
||||
export class CodemanError extends Error {
|
||||
constructor(
|
||||
message: string,
|
||||
public code: string,
|
||||
public context?: Record<string, unknown>
|
||||
) {
|
||||
super(message);
|
||||
Object.setPrototypeOf(this, CodemanError.prototype);
|
||||
this.name = 'CodemanError';
|
||||
}
|
||||
}
|
||||
|
||||
export class SessionError extends CodemanError {
|
||||
constructor(message: string, code: string, public sessionId: string) {
|
||||
super(message, code, { sessionId });
|
||||
this.name = 'SessionError';
|
||||
}
|
||||
}
|
||||
|
||||
export class ValidationError extends CodemanError {
|
||||
constructor(message: string, public field: string, public value: unknown) {
|
||||
super(message, 'VALIDATION_ERROR', { field, value });
|
||||
this.name = 'ValidationError';
|
||||
}
|
||||
}
|
||||
|
||||
export class ScreenError extends CodemanError {
|
||||
constructor(message: string, public screenName: string, public operation: string) {
|
||||
super(message, 'SCREEN_ERROR', { screenName, operation });
|
||||
this.name = 'ScreenError';
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 11. Split Large Files
|
||||
|
||||
**`types.ts` (~1500 lines)**:
|
||||
```
|
||||
src/types/
|
||||
index.ts # Re-exports all
|
||||
session.types.ts # Session-related types
|
||||
task.types.ts # Task-related types
|
||||
ralph.types.ts # Ralph loop types
|
||||
api.types.ts # API request/response types
|
||||
config.types.ts # Configuration types
|
||||
factories.ts # createInitialState(), etc.
|
||||
```
|
||||
|
||||
**`server.ts`**:
|
||||
```
|
||||
src/web/
|
||||
server.ts # Main Fastify setup
|
||||
routes/
|
||||
sessions.ts # Session management routes
|
||||
respawn.ts # Respawn control routes
|
||||
scheduled.ts # Scheduled run routes
|
||||
system.ts # System status routes
|
||||
sse/
|
||||
manager.ts # SSE client management
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 12. Formalize Result Pattern
|
||||
|
||||
Existing `validateTokenCounts` returns `{ isValid, reason }` which is essentially a Result.
|
||||
|
||||
**Option A: Simple Result type (no dependency)**:
|
||||
```typescript
|
||||
// src/utils/result.ts
|
||||
export type Result<T, E = Error> =
|
||||
| { success: true; data: T }
|
||||
| { success: false; error: E };
|
||||
|
||||
export const ok = <T>(data: T): Result<T, never> => ({ success: true, data });
|
||||
export const err = <E>(error: E): Result<never, E> => ({ success: false, error });
|
||||
```
|
||||
|
||||
**Option B: Install neverthrow**:
|
||||
```bash
|
||||
npm install neverthrow
|
||||
```
|
||||
|
||||
Provides chaining (`map`, `andThen`, `match`) and `ResultAsync` for async operations.
|
||||
|
||||
---
|
||||
|
||||
### 13. Template Literal Types for IDs
|
||||
|
||||
Enforce ID formats at compile time:
|
||||
|
||||
```typescript
|
||||
type CycleIdFormat = `${string}:cycle-${number}`;
|
||||
type ScreenSessionName = `codeman-${string}`;
|
||||
|
||||
interface RespawnCycleMetrics {
|
||||
cycleId: CycleIdFormat; // Enforces format at compile time
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Summary Table
|
||||
|
||||
| # | Suggestion | Category | Effort | Impact |
|
||||
|---|------------|----------|--------|--------|
|
||||
| 1 | Use Zod for API validation | Error Handling | Low | High |
|
||||
| 2 | Add `assertNever` utility | Type Safety | Low | High |
|
||||
| 3 | Standardize `createErrorResponse` | Error Handling | Low | Medium |
|
||||
| 4 | Discriminated union for `ApiResponse` | Type Safety | Medium | High |
|
||||
| 5 | Branded types for tokens | Type Safety | Medium | Medium |
|
||||
| 6 | Dependency injection for services | Architecture | Medium | High |
|
||||
| 7 | Enforce `import type` | Architecture | Low | Medium |
|
||||
| 8 | Circular dependency detection | Architecture | Low | Medium |
|
||||
| 9 | `as const satisfies` for configs | Type Safety | Low | Low |
|
||||
| 10 | Custom error classes | Error Handling | Medium | Medium |
|
||||
| 11 | Split large files | Architecture | High | Medium |
|
||||
| 12 | Formalize Result pattern | Error Handling | Medium | Medium |
|
||||
| 13 | Template literal types for IDs | Type Safety | Low | Low |
|
||||
|
||||
---
|
||||
|
||||
## Notable Strengths to Keep
|
||||
|
||||
These patterns are already well-implemented and should be preserved:
|
||||
|
||||
- **Circuit breaker pattern** in `state-store.ts` and `ai-checker-base.ts` (excellent resilience)
|
||||
- **`getErrorMessage()` utility** (solid, used in 8 files)
|
||||
- **Barrel files for `utils/` and `prompts/`** (appropriate size, good organization)
|
||||
- **Strict TypeScript config** (comprehensive strictness settings)
|
||||
- **Well-documented configuration** in `src/config/`
|
||||
- **Extensive union types** for status tracking (18+ well-defined types)
|
||||
- **Type guards** like `isError()` for runtime narrowing
|
||||
|
||||
---
|
||||
|
||||
## References
|
||||
|
||||
### Type Safety
|
||||
- [TypeScript Handbook: Narrowing](https://www.typescriptlang.org/docs/handbook/2/narrowing.html)
|
||||
- [Fullstory: Discriminated Unions](https://www.fullstory.com/blog/discriminated-unions-and-exhaustiveness-checking-in-typescript/)
|
||||
- [Learning TypeScript: Branded Types](https://www.learningtypescript.com/articles/branded-types)
|
||||
- [Total TypeScript: satisfies Operator](https://www.totaltypescript.com/how-to-use-satisfies-operator)
|
||||
|
||||
### Error Handling
|
||||
- [neverthrow GitHub](https://github.com/supermacro/neverthrow)
|
||||
- [Zod Documentation](https://zod.dev/)
|
||||
- [Custom Errors in TypeScript](https://medium.com/@Nelsonalfonso/understanding-custom-errors-in-typescript-a-complete-guide-f47a1df9354c)
|
||||
|
||||
### Architecture
|
||||
- [Please Stop Using Barrel Files - TkDodo](https://tkdodo.eu/blog/please-stop-using-barrel-files)
|
||||
- [TypeScript Dependency Injection](https://softwarepatternslexicon.com/js/typescript-and-javascript-design-patterns/dependency-injection-with-typescript/)
|
||||
- [dpdm - Circular Dependency Detector](https://github.com/acrazing/dpdm)
|
||||
- [Consistent Type Imports - typescript-eslint](https://typescript-eslint.io/blog/consistent-type-imports-and-exports-why-and-how/)
|
||||
@@ -0,0 +1,155 @@
|
||||
# Voice Input V2 — Implementation Plan
|
||||
|
||||
## Executive Summary
|
||||
|
||||
Fix and improve the existing VoiceInput implementation. The core class is solid but has **critical integration bugs** that prevent it from working on mobile, plus several UX improvements needed to make it feel fast and polished.
|
||||
|
||||
---
|
||||
|
||||
## Current State: What Exists
|
||||
|
||||
The `VoiceInput` singleton (app.js:602-830) is already committed and uses the Web Speech API with:
|
||||
- Toggle mode (tap start/stop), 5s silence auto-stop
|
||||
- `interimResults: true` for streaming transcription preview
|
||||
- iOS Safari `isFinal` workaround (750ms stability timer)
|
||||
- Desktop button in `toolbar-right`, mobile button in `KeyboardAccessoryBar`
|
||||
- `voice-pulse` CSS animation, `.voice-preview` overlay
|
||||
- Cleanup on SSE reconnect, haptic feedback on mobile
|
||||
|
||||
## Critical Bugs Found (Must Fix)
|
||||
|
||||
### Bug 1: Mobile button NEVER shows (CRITICAL)
|
||||
`KeyboardAccessoryBar.init()` runs at line 2239, BEFORE `VoiceInput.init()` at line 2240. The accessory bar template checks `VoiceInput.supported` at render time — but `init()` hasn't run yet, so `supported` is still `false`. The inline `style="${VoiceInput.supported ? '' : 'display:none'}"` always resolves to `display:none`.
|
||||
|
||||
**Fix:** Move `VoiceInput.init()` BEFORE `KeyboardAccessoryBar.init()`, OR remove the inline style check and have `VoiceInput.init()` show/hide the mobile button after the fact (like it does for desktop).
|
||||
|
||||
### Bug 2: `_showButtons()` ignores mobile button
|
||||
`_showButtons()` only targets `#voiceInputBtn` (desktop). It never removes `display:none` from the mobile `[data-action="voice"]` button.
|
||||
|
||||
**Fix:** Add mobile button selector to `_showButtons()`.
|
||||
|
||||
### Bug 3: Recognition instance leak on cleanup
|
||||
`cleanup()` stops recording and removes the preview element, but doesn't null out `this.recognition`. After `cleanup()` + `init()` on SSE reconnect, the old `SpeechRecognition` instance with its handlers is orphaned.
|
||||
|
||||
**Fix:** Add `this.recognition = null` in `cleanup()`.
|
||||
|
||||
## UX Improvements (Should Fix)
|
||||
|
||||
### Improvement 1: Consider auto-sending after voice
|
||||
Currently, voice text is inserted but the user must press Enter. This is safe but adds friction. Two options:
|
||||
- **Option A (safe, current):** Insert text, user presses Enter — good for a terminal where wrong commands matter
|
||||
- **Option B (fast):** Insert text + auto-send `\r` after a brief 500ms delay — feels more "voice assistant"-like
|
||||
- **Recommendation:** Keep Option A as default, but add an optional setting for auto-send
|
||||
|
||||
### Improvement 2: Shorter silence timeout for commands
|
||||
5 seconds of silence before auto-stop feels slow for short terminal commands. Consider:
|
||||
- 3 seconds for auto-stop (still generous for natural pauses)
|
||||
- Or make it configurable via settings
|
||||
|
||||
### Improvement 3: Better visual state on mobile
|
||||
The blue-tinted voice button in the accessory bar is distinctive but subtle. When recording:
|
||||
- The `.recording` class turns it red with pulse — good
|
||||
- But the button is small among other buttons — easy to miss the state change
|
||||
- Consider: also show a small red dot indicator in the header or terminal area during recording
|
||||
|
||||
## Architecture Decision: Keep Web Speech API
|
||||
|
||||
Confirmed by research: Web Speech API is the right choice.
|
||||
- **Free, fast (150-300ms interim), trivial complexity**
|
||||
- Chrome + Safari = ~70% of users, ~95% of Codeman's target audience (devs on Chrome)
|
||||
- Works on localhost without HTTPS
|
||||
- Accuracy is adequate for English command dictation
|
||||
- Deepgram streaming (Phase 2 optional) only if accuracy complaints arise
|
||||
- Skip Whisper batch entirely (too slow for interactive voice input)
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
### Phase 1: Fix Critical Bugs (Priority)
|
||||
|
||||
**File: `src/web/public/app.js`**
|
||||
|
||||
1. **Fix init order** — Move `VoiceInput.init()` BEFORE `KeyboardAccessoryBar.init()`:
|
||||
```
|
||||
// Current (broken):
|
||||
KeyboardAccessoryBar.init();
|
||||
VoiceInput.init();
|
||||
|
||||
// Fixed:
|
||||
VoiceInput.init();
|
||||
KeyboardAccessoryBar.init();
|
||||
```
|
||||
|
||||
2. **Fix `_showButtons()` to handle mobile** — Add mobile button selector:
|
||||
```javascript
|
||||
_showButtons() {
|
||||
const desktopBtn = document.getElementById('voiceInputBtn');
|
||||
if (desktopBtn) desktopBtn.style.display = '';
|
||||
// Also show mobile button (may not exist yet if KeyboardAccessoryBar hasn't init'd)
|
||||
const mobileBtn = document.querySelector('[data-action="voice"]');
|
||||
if (mobileBtn) mobileBtn.style.display = '';
|
||||
}
|
||||
```
|
||||
|
||||
3. **Fix cleanup leak** — Null out recognition instance:
|
||||
```javascript
|
||||
cleanup() {
|
||||
if (this.isRecording) this.stop();
|
||||
if (this.previewEl) {
|
||||
this.previewEl.remove();
|
||||
this.previewEl = null;
|
||||
}
|
||||
this.recognition = null; // <-- add this
|
||||
clearTimeout(this.silenceTimeout);
|
||||
clearTimeout(this._stabilityTimer);
|
||||
// ... rest
|
||||
}
|
||||
```
|
||||
|
||||
4. **Remove inline style from mobile button template** — Since `_showButtons()` will handle visibility, the template should always render the button visible and let `init()` hide it if unsupported:
|
||||
```
|
||||
// Current (broken):
|
||||
style="${VoiceInput.supported ? '' : 'display:none'}"
|
||||
|
||||
// Fixed: remove the style attr entirely, let _showButtons/_hideButtons manage it
|
||||
```
|
||||
Actually better: **always show the button** if we init VoiceInput before KeyboardAccessoryBar. The `VoiceInput.supported` will be set correctly by then.
|
||||
|
||||
### Phase 2: UX Polish
|
||||
|
||||
5. **Reduce silence timeout** from 5s to 3s for snappier feel
|
||||
|
||||
6. **Add recording indicator** — When recording, add a subtle pulsing red dot to the session header or status area so the recording state is visible even if the button is off-screen
|
||||
|
||||
7. **Voice input setting** — Add a toggle in App Settings to enable/disable voice input (some users may not want the button). Default: enabled on supported browsers.
|
||||
|
||||
### Phase 3: Future Enhancements (Not in this PR)
|
||||
|
||||
- Language selector (currently hardcoded `en-US`)
|
||||
- Auto-send option (insert text + `\r` automatically)
|
||||
- Deepgram WebSocket fallback for Firefox/Edge
|
||||
- Waveform visualization during recording
|
||||
- Voice command recognition ("clear", "compact", "new session")
|
||||
|
||||
## Files to Modify
|
||||
|
||||
| File | Changes |
|
||||
|------|---------|
|
||||
| `src/web/public/app.js` | Fix init order, fix `_showButtons()`, fix `cleanup()`, remove inline style, reduce silence timeout |
|
||||
| `src/web/public/mobile.css` | (optional) Adjust voice preview positioning if needed |
|
||||
|
||||
## Testing Plan
|
||||
|
||||
1. **Desktop Chrome:** Verify mic button visible in toolbar-right, click toggles recording state, interim text shows in preview, final text inserted at prompt
|
||||
2. **Mobile Chrome (emulated):** Verify mic button visible in accessory bar, tap toggles recording, pulse animation plays
|
||||
3. **Firefox:** Verify mic button is hidden (no SpeechRecognition support)
|
||||
4. **SSE reconnect:** Verify cleanup stops recording and re-init works
|
||||
5. **No active session:** Verify toast "No active session" shows when tapping mic with no session
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Risk | Impact | Mitigation |
|
||||
|------|--------|------------|
|
||||
| iOS Safari isFinal bug | Medium | Already handled by 750ms stability timer |
|
||||
| Chrome auto-stops after 60s | Low | Prompts are short; 3s silence timeout covers this |
|
||||
| Mic permission denied | Low | Error toast with clear message |
|
||||
| Init order regression | High | Integration test to verify button visibility |
|
||||
Reference in New Issue
Block a user