chore: bump version to 0.1402

This commit is contained in:
arkon
2026-01-27 23:15:25 +01:00
parent 9fc3d2eda3
commit 46434ee6d0
11 changed files with 2253 additions and 870 deletions
+94 -815
View File
@@ -2,877 +2,156 @@
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## ⚠️ CRITICAL: Screen Session Safety ## CRITICAL: Screen Session Safety
**You may be running inside a Claudeman-managed screen session.** Before killing ANY screen or Claude process: **You may be running inside a Claudeman-managed screen session.** Before killing ANY screen or Claude process:
1. **Check environment**: `echo $CLAUDEMAN_SCREEN` - if it returns `1`, you're in a managed session 1. Check: `echo $CLAUDEMAN_SCREEN` - if `1`, you're in a managed session
2. **NEVER run** `screen -X quit`, `pkill screen`, or `pkill claude` without first confirming you're not killing yourself 2. **NEVER** run `screen -X quit`, `pkill screen`, or `pkill claude` without confirming
3. **Safe debugging**: Use `screen -ls` to LIST sessions, but don't kill them blindly 3. Use the web UI or `./scripts/screen-manager.sh` instead of direct kill commands
4. **If you need to kill screens**: Use the web UI or `./scripts/screen-manager.sh` instead of direct commands
**Why this matters**: Killing your own screen terminates your session mid-work, losing context and potentially corrupting state. ## COM Shorthand (Deployment)
## ⚡ COM Shorthand (Deployment)
When user says "COM": When user says "COM":
1. Increment version in BOTH `package.json` AND `CLAUDE.md` (keep them in sync) 1. Increment version in BOTH `package.json` AND `CLAUDE.md`
2. Run: `git add -A && git commit -m "chore: bump version to X.XXXX" && git push && npm run build && systemctl --user restart claudeman-web` 2. Run: `git add -A && git commit -m "chore: bump version to X.XXXX" && git push && npm run build && systemctl --user restart claudeman-web`
Always bump version on every COM, even for small changes. **Version**: 0.1402 (must match `package.json`)
## Project Overview ## Project Overview
Claudeman is a Claude Code session manager with a web interface and autonomous Ralph Loop. It spawns Claude CLI processes via PTY, streams output in real-time via SSE, and supports scheduled/timed runs. Claudeman is a Claude Code session manager with web interface and autonomous Ralph Loop. Spawns Claude CLI via PTY, streams via SSE, supports respawn cycling for 24+ hour autonomous runs.
### Mission: Rock-Solid Performance **Tech Stack**: TypeScript (ES2022/NodeNext, strict mode with `noUnusedLocals`, `noUnusedParameters`, `noImplicitReturns`), Node.js, Fastify, node-pty, xterm.js
**The app MUST remain fast, responsive, and never hang** — even with many tabs open and many subagent windows active. This is a core design principle, not a nice-to-have. Every feature must be built with this constraint in mind: **Requirements**: Node.js 18+, Claude CLI, GNU Screen
- **60fps terminal rendering** with batched writes and `requestAnimationFrame`
- **Auto-trimming buffers** prevent memory bloat in long-running sessions
- **Debounced state persistence** (500ms) avoids disk thrashing
- **SSE batching** (16ms) reduces client-side event storms
- **Lazy loading** of agent transcripts and historical data
- **No blocking operations** in the main event loop
When adding new features, always ask: "Will this maintain responsiveness with 20 sessions and 50 agent windows?"
**Version**: 0.1400 (must match `package.json`)
**Tech Stack**: TypeScript (ES2022/NodeNext, strict mode), Node.js, Fastify, Server-Sent Events, node-pty
**Key Dependencies**: fastify (REST API), node-pty (PTY spawning), ink/react (TUI), xterm.js (web terminal), @modelcontextprotocol/sdk (MCP server for spawn protocol)
**Requirements**: Node.js 18+, Claude CLI (`claude`) installed, GNU Screen (`apt install screen` / `brew install screen`)
> **Note**: `claude` does not need to be in the server process's PATH. Claudeman auto-discovers the binary from common install locations (`~/.local/bin`, `~/.claude/local`, `/usr/local/bin`, etc.) and augments PATH for spawned sessions.
> **Runtime**: The web server runs as a systemd user service (`claudeman-web.service`) on HTTPS port 3000 with a self-signed certificate. It auto-restarts and survives logout.
## First-Time Setup
```bash
npm install
```
## Commands ## Commands
**CRITICAL**: `npm run dev` runs CLI help, NOT the web server. Use `npx tsx src/index.ts web` for development. **CRITICAL**: `npm run dev` shows CLI help, NOT the web server.
**Quick reference**:
- Dev server: `npx tsx src/index.ts web` (or `web --https` for notifications)
- Type check: `npx tsc --noEmit`
- Single test: `npx vitest run test/<file>.test.ts`
- Restart prod: `systemctl --user restart claudeman-web`
### Build & Clean
```bash ```bash
npm run build # Compile TS + copy static files + templates + make bins executable # Development
npm run clean # Remove dist/ npx tsx src/index.ts web # Dev server (RECOMMENDED)
npm run typecheck # Type check without building (or: npx tsc --noEmit) npx tsx src/index.ts web --https # With TLS for notifications
``` npx tsc --noEmit # Type check
### Web Server # Testing
npx vitest run # All tests
npx vitest run test/<file>.test.ts # Single file
npm run test:e2e # Browser E2E tests
```bash # Production
npx tsx src/index.ts web # Dev mode - no build needed (RECOMMENDED) npm run build
npx tsx src/index.ts web -p 8080 # Dev mode with custom port systemctl --user restart claudeman-web
npx tsx src/index.ts web --https # Dev mode with self-signed TLS (enables browser notifications) journalctl --user -u claudeman-web -f
npm run web # After npm run build (shorthand)
node dist/index.js web # After npm run build
claudeman web # After npm link
```
### TUI (Terminal User Interface)
```bash
npx tsx src/index.ts tui # Dev mode - prompts to start web if not running
claudeman tui # After npm link
claudeman tui --with-web # Auto-start web server if not running (no prompt)
claudeman tui --no-web # Skip web server check entirely
claudeman tui -p 8080 # Specify web server port
```
### Testing
```bash
npm run test # Run all tests once
npm run test:watch # Watch mode
npm run test:coverage # With coverage report
npx vitest run test/session.test.ts # Single file
npx vitest run -t "should create session" # By pattern
```
**Test Configuration** (vitest.config.ts):
- `globals: true` - no imports needed for `describe`/`it`/`expect`
- `testTimeout: 30000` - 30s for integration tests
- `teardownTimeout: 60000` - 60s ensures cleanup runs even on failures
- `fileParallelism: false` - sequential file execution to respect screen session limits
- Coverage excludes entry points: `src/index.ts`, `src/cli.ts`
**Test Port Allocation** (integration tests spawn servers):
| Port | Test File |
|------|-----------|
| 3099 | quick-start.test.ts |
| 3102 | session.test.ts |
| 3105 | scheduled-runs.test.ts |
| 3107 | sse-events.test.ts |
| 3110 | edge-cases.test.ts |
| 3115 | integration-flows.test.ts |
| 3120 | session-cleanup.test.ts |
| 3125 | ralph-integration.test.ts |
| 3127 | respawn-integration.test.ts (reserved) |
| 3130 | hooks-config.test.ts (Hook Event API) |
| 3131 | hooks-config.test.ts (Hook Data Sanitization) |
| 3150 | browser-e2e.test.ts (main browser tests) |
| 3151 | browser-e2e.test.ts (SSE events tests) |
| 3152 | browser-e2e.test.ts (hook events tests) |
| 3153 | browser-e2e.test.ts (Ralph panel tests) |
| 3154 | file-link-click.test.ts |
| 3155 | browser-playwright.test.ts |
| 3156 | browser-puppeteer.test.ts |
| 3157 | browser-agent.test.ts |
| 3158-3160 | browser-comparison.test.ts |
| 3180-3182 | scripts/browser-comparison.mjs (benchmark) |
| 3183 | test/e2e/workflows/quick-start.e2e.ts |
| 3184 | test/e2e/workflows/session-input.e2e.ts |
| 3185 | test/e2e/workflows/session-delete.e2e.ts |
| 3186 | test/e2e/workflows/multi-session.e2e.ts |
| 3187 | test/e2e/workflows/agent-interactions.e2e.ts |
| 3188 | test/e2e/workflows/input-interactions.e2e.ts |
| 3189 | test/e2e/workflows/respawn-flow.e2e.ts |
**Next available port**: 3190
**Browser Testing**: Three frameworks available (Playwright, Puppeteer, Agent-Browser). See `docs/browser-testing-guide.md` for full comparison and patterns.
```bash
# Run browser benchmark (standalone - recommended)
npx tsx scripts/browser-comparison.mjs
# Run existing browser E2E tests
npm test -- test/browser-e2e.test.ts
```
**Browser Testing Key Points**:
- **Vitest hook issue**: Browser tests using `beforeAll`/`afterAll` timeout even when tests pass. Use standalone scripts or run browser code directly in tests.
- **Recommended framework**: Playwright for most cases (auto-waiting, debugging). Puppeteer for Chrome-specific/CDP features.
- **Required browser args**: `--no-sandbox`, `--disable-setuid-sandbox`, `--disable-dev-shm-usage`
- **Install browsers**: `npx playwright install chromium` after npm install
### E2E Test Suite
Real browser-based end-to-end tests using Playwright that validate actual user workflows. These tests catch issues that unit tests miss (like the cpulimit bug that broke screen creation).
**Setup**:
```bash
npm install # Install dependencies (pixelmatch, pngjs)
npx playwright install chromium # Install browser
```
**Running E2E Tests**:
```bash
npm run test:e2e # Run all E2E tests
npm run test:e2e:quick # Run quick-start test only (critical path)
npx vitest run test/e2e/workflows/quick-start.e2e.ts # Run single test file
npx vitest run test/e2e/ -t "should create session" # Run by pattern
```
**Test Fixtures** (`test/e2e/fixtures/`):
| Fixture | Purpose |
|---------|---------|
| `server.fixture.ts` | `createServerFixture(port)` / `destroyServerFixture()` - Server lifecycle |
| `browser.fixture.ts` | `createBrowserFixture()` / `destroyBrowserFixture()` - Playwright browser with required args |
| `cleanup.fixture.ts` | `CleanupTracker` - Tracks sessions, cases, screens for cleanup |
| `screenshot.fixture.ts` | `captureAndCompare()` - Visual regression testing with pixelmatch |
**Workflow Tests** (`test/e2e/workflows/`):
| Test | Port | What it validates |
|------|------|-------------------|
| `quick-start.e2e.ts` | 3183 | **Critical path**: click button → session created → screen created → terminal visible |
| `session-input.e2e.ts` | 3184 | Terminal input, Ctrl+C cancel, multi-line input |
| `session-delete.e2e.ts` | 3185 | Delete button → screen killed → UI updated |
| `multi-session.e2e.ts` | 3186 | Multiple sessions, tab switching, Ctrl+Tab shortcut |
| `agent-interactions.e2e.ts` | 3187 | Subagent windows, parent attachment, visibility |
| `input-interactions.e2e.ts` | 3188 | Modals, checkboxes, Ctrl+Enter/W shortcuts |
| `respawn-flow.e2e.ts` | 3189 | Respawn enable/start/stop via API and UI |
| `ralph-loop.e2e.ts` | 3190 | Ralph Loop wizard: open, configure, start, verify tracker enabled |
**Screenshot Validation**:
- Baselines stored in `test/e2e/screenshots/baselines/`
- Current screenshots in `test/e2e/screenshots/current/`
- Diff images (on failure) in `test/e2e/screenshots/diffs/`
- First run auto-creates baselines; subsequent runs compare
- Default threshold: 5% pixel difference allowed
**Cleanup Behavior**:
- All test cases use `e2e-test-*` prefix for easy identification
- `CleanupTracker.forceCleanupAll()` removes ALL `e2e-test-*` resources
- try/finally patterns ensure cleanup even on test failures
- `afterAll` hooks call cleanup as safety net
**E2E Test Pattern** (avoids Vitest hook timeout issue):
```typescript
it('should create session', async () => {
let browser: BrowserFixture | null = null;
const caseName = generateCaseName('test');
try {
serverFixture = await createServerFixture(PORT);
cleanup = new CleanupTracker(serverFixture.baseUrl);
cleanup.trackCase(caseName);
browser = await createBrowserFixture();
// ... test code ...
} finally {
if (browser) await destroyBrowserFixture(browser);
}
}, 90000);
```
**E2E Installation Summary**:
| Component | Details |
|-----------|---------|
| **Test Files** | 14 TypeScript files in `test/e2e/` |
| **Fixtures** | server, browser, cleanup, screenshot, index, pixelmatch types |
| **Workflows** | quick-start, session-input, session-delete, multi-session, agent-interactions, input-interactions, respawn-flow |
| **Dependencies** | `playwright`, `pixelmatch`, `pngjs`, `@types/pngjs` |
| **Browser** | Chromium via `npx playwright install chromium` |
| **NPM Scripts** | `npm run test:e2e` (all), `npm run test:e2e:quick` (critical path) |
| **Ports** | 3183-3189 (see workflow tests table above) |
| **Gitignore** | `test/e2e/screenshots/current/` and `diffs/` ignored; `baselines/` tracked |
**Unit tests** (no server needed): `respawn-controller`, `ralph-tracker`, `pty-interactive`, `task-queue`, `task`, `ralph-loop`, `session-manager`, `state-store`, `types`, `templates`, `ralph-config`, `spawn-detector`, `spawn-types`, `spawn-orchestrator`, `ai-idle-checker`, `ai-plan-checker`
**Test Utilities**: `test/respawn-test-utils.ts` provides MockSession, MockAiIdleChecker, MockAiPlanChecker, time controller, state tracker, and event recorder for respawn controller testing. See `test/respawn-test-plan.md` for architecture and `test/respawn-scenarios.md` for comprehensive test scenarios.
**Test Safety**: `test/setup.ts` enforces max 10 concurrent screens, performs orphan cleanup, and protects its own process tree. You can safely run tests from within a Claudeman-managed session - the cleanup will not kill your own Claude instance. The respawn-controller tests use MockSession (not real screens).
**Test Cleanup Patterns**: Integration tests track resources in `createdSessions` and `createdCases` arrays, cleaned up by `afterAll`/`afterEach` hooks. However, some tests perform cleanup in the test body itself (e.g., `edge-cases.test.ts:273-302` creates 5 sessions and cleans them in a loop). If assertions fail before cleanup code runs, resources leak.
**Known Cleanup Issues** (technical debt):
- `pty-interactive.test.ts`: Uses `await session.stop()` at end of each test, not in `afterEach`. Test failures leave sessions running.
- `edge-cases.test.ts`: Multiple sessions created in test body with cleanup at end; failures leak sessions.
- Test cases (`~/claudeman-cases/`): Cases named `flow-test-*`, `ralph-track-loop-*`, `session-detail-*` may persist after test failures.
**Manual Cleanup**:
```bash
# Remove orphaned test cases
rm -rf ~/claudeman-cases/flow-test-* ~/claudeman-cases/ralph-track-loop-* ~/claudeman-cases/session-detail-*
# Kill orphaned test screens (only detached claudeman screens)
screen -ls | grep -E 'Detached.*claudeman' | cut -d. -f1 | xargs -I{} screen -S {} -X quit
```
### MCP Server
```bash
npx tsx src/mcp-server.ts # Dev mode (stdio transport)
```
Configure in Claude Code's MCP settings:
```json
{ "command": "node", "args": ["dist/mcp-server.js"], "env": { "CLAUDEMAN_API_URL": "http://localhost:3000", "CLAUDEMAN_SESSION_ID": "<id>" } }
```
### Debugging
```bash
screen -ls # List GNU screen sessions
screen -r <name> # Attach to screen session (Ctrl+A D to detach)
curl localhost:3000/api/sessions # Check active sessions
curl localhost:3000/api/status | jq . # Full app state including respawn
cat ~/.claudeman/state.json | jq . # View main state
cat ~/.claudeman/state-inner.json | jq . # View Ralph loop state
```
### Systemd Service
```bash
systemctl --user status claudeman-web # Check status
systemctl --user restart claudeman-web # Restart
systemctl --user stop claudeman-web # Stop
journalctl --user -u claudeman-web -f # Stream logs
```
Install: `ln -sf scripts/claudeman-web.service ~/.config/systemd/user/`
Enable: `systemctl --user enable claudeman-web && loginctl enable-linger $USER`
### Kill Stuck Screens
```bash
screen -X -S <name> quit # Graceful quit
pkill -f "SCREEN.*claudeman" # Force kill all claudeman screens
```
## CLI Commands
```bash
claudeman session [s] # Manage Claude sessions
start # Start new session
stop <id> # Stop session
list [ls] # List all
logs <id> # View output
claudeman task [t] # Manage tasks
add <prompt> # Add task
list [ls] # List tasks
status <id> # Task details
remove [rm] <id> # Remove task
clear # Clear completed/failed
claudeman ralph [r] # Control Ralph loop
start # Start loop
stop # Stop loop
status # Show status
claudeman web # Start web interface
claudeman tui # Start TUI
claudeman status # Overall status
claudeman reset # Reset all state
``` ```
## Architecture ## Architecture
### Key Files ### Core Files
**Core Session Management:**
| File | Purpose |
|------|---------|
| `src/session.ts` | Core PTY wrapper for Claude CLI. Modes: `runPrompt()`, `startInteractive()`, `startShell()` |
| `src/screen-manager.ts` | GNU screen persistence, ghost discovery, 4-strategy kill |
| `src/session-manager.ts` | Session lifecycle, task assignment, cleanup |
| `src/state-store.ts` | JSON persistence to `~/.claudeman/` with debounced writes |
| `src/types.ts` | All TypeScript interfaces |
**Autonomous Features:**
| File | Purpose |
|------|---------|
| `src/respawn-controller.ts` | State machine for autonomous session cycling |
| `src/ai-idle-checker.ts` | Spawns Claude to analyze terminal output for IDLE/WORKING verdict |
| `src/ai-plan-checker.ts` | Spawns Claude to detect plan mode prompts for auto-accept |
| `src/ralph-tracker.ts` | Detects `<promise>PHRASE</promise>`, todos, loop status |
| `src/ralph-config.ts` | Parses `.claude/ralph-loop.local.md` and CLAUDE.md for Ralph config |
| `src/plan-orchestrator.ts` | Multi-agent plan generation with parallel analysis + verification |
| `src/run-summary.ts` | Tracks session events for "what happened while away" summaries |
**Spawn Protocol (Autonomous Agents):**
| File | Purpose |
|------|---------|
| `src/spawn-orchestrator.ts` | Full agent lifecycle: spawn, monitor, budget, queue, cleanup |
| `src/mcp-server.ts` | MCP server binary exposing spawn tools to Claude Code |
| `src/subagent-watcher.ts` | Monitors Claude Code background agents in `~/.claude/projects/*/subagents/*.jsonl` |
**Web & TUI:**
| File | Purpose | | File | Purpose |
|------|---------| |------|---------|
| `src/session.ts` | PTY wrapper: `runPrompt()`, `startInteractive()`, `startShell()` |
| `src/screen-manager.ts` | GNU screen persistence, ghost discovery |
| `src/session-manager.ts` | Session lifecycle, cleanup |
| `src/respawn-controller.ts` | State machine for autonomous cycling |
| `src/ralph-tracker.ts` | Detects `<promise>PHRASE</promise>`, todos |
| `src/plan-orchestrator.ts` | Multi-agent plan generation |
| `src/web/server.ts` | Fastify REST API + SSE at `/api/events` | | `src/web/server.ts` | Fastify REST API + SSE at `/api/events` |
| `src/web/public/app.js` | Frontend: SSE handling, xterm.js, tab management | | `src/web/public/app.js` | Frontend: xterm.js, tab management, subagent windows |
| `src/tui/App.tsx` | TUI main component (Ink/React) | | `src/types.ts` | All TypeScript interfaces |
| `src/tui/DirectAttach.ts` | Full-screen console attach with tab switching |
### Data Flow ### Data Flow
1. **Session** spawns `claude -p --dangerously-skip-permissions` via `node-pty` (PATH augmented to include claude's install directory) 1. Session spawns `claude --dangerously-skip-permissions` via node-pty
2. PTY output is buffered, ANSI stripped, and parsed for JSON messages 2. PTY output buffered, ANSI stripped, parsed for JSON messages
3. **WebServer** broadcasts events to SSE clients at `/api/events` 3. WebServer broadcasts to SSE clients at `/api/events`
4. Full session state (settings, tokens, respawn config, Ralph state) persists to `~/.claudeman/state.json` via **StateStore** 4. State persists to `~/.claudeman/state.json` via StateStore
5. Screen metadata persists separately to `~/.claudeman/screens.json` for session recovery
### Respawn State Machine ### Key Patterns
State machine for autonomous session cycling: `watching` → `confirming_idle` → `ai_checking` → `sending_update` → `waiting_update` → `sending_clear` → `waiting_clear` → `sending_init` → `waiting_init` → `monitoring_init` → (optionally) `sending_kickstart`. Steps can be skipped via config (`sendClear: false`, `sendInit: false`). After each step, waits for `completionConfirmMs` (10s) of output silence before proceeding. AI idle check uses a fresh Claude session to analyze terminal output for IDLE/WORKING verdict. **Input to sessions**: Use `session.writeViaScreen()` for programmatic input (respawn, auto-compact). Text and Enter sent as separate `screen -X stuff` commands due to Ink's requirements. All prompts must be single-line.
See `docs/respawn-state-machine.md` for the full state diagram, idle detection layers, and auto-accept behavior. **Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`.
### Spawn1337 Protocol (Autonomous Agents) **Token tracking**: Interactive mode parses status line ("123.4k tokens"), estimates 60/40 input/output split.
Spawned agents are full-power Claude sessions in their own screen sessions, managed via MCP tools (`spawn_agent`, `list_agents`, `get_agent_status`, `get_agent_result`, `send_agent_message`, `cancel_agent`). Max 5 concurrent, max depth 3, default timeout 30min. Agents communicate via filesystem (`spawn-comms/`) and signal completion via `<promise>PHRASE</promise>`. **Memory leak prevention**: Frontend runs long; clear all Maps/timers on SSE reconnect in `handleInit()`.
See `docs/spawn-protocol.md` for the full protocol flow, directory structure, resource governance, and MCP configuration. ## Adding Features
### Subagent Watcher (Claude Code Background Agents)
Monitors Claude Code's internal background agents (the `Task` tool) in real-time. Watches `~/.claude/projects/{project}/{session}/subagents/agent-{id}.jsonl` files and emits structured events.
**Events**: `subagent:discovered`, `subagent:tool_call`, `subagent:progress`, `subagent:message`, `subagent:completed`
**API**:
- `GET /api/subagents` - List all known subagents (optional `?minutes=60` for recent only)
- `GET /api/subagents/:agentId` - Get subagent info
- `GET /api/subagents/:agentId/transcript` - Get transcript (`?limit=N`, `?format=formatted`)
- `DELETE /api/subagents/:agentId` - Kill subagent process
- `GET /api/sessions/:id/subagents` - Get subagents for session's working directory
**Status lifecycle**: `active` → `idle` (30s no activity) → `completed` (process exited or file stale)
**Settings**: Can be disabled via App Settings → Display → "Enable Subagent Tracking" (default: enabled). Setting is stored in `~/.claudeman/settings.json` as `subagentTrackingEnabled`.
Implementation: `src/subagent-watcher.ts` - singleton `subagentWatcher` started on server boot (if enabled).
### Subagent Window Management (Frontend)
Floating subagent windows in the web UI show real-time activity from Claude Code's background agents. Each window displays tool calls, progress, and messages, with visual connection lines to the parent session tab.
**Parent Session Discovery** (`app.js:findParentSessionForSubagent()`):
1. When a subagent is discovered via SSE event, the frontend must determine which Claudeman session spawned it
2. Matching uses `workingDir` → `projectHash` conversion: `/home/user/project` becomes `-home-user-project`
3. For each session, calls `/api/sessions/{id}/subagents` which returns subagents matching that session's `workingDir`
4. First match wins: checks active session first, then iterates through other sessions
5. Once found, caches `parentSessionId` and `parentSessionName` in the agent object
**Window Visibility** (`app.js:updateSubagentWindowVisibility()`):
- "Show for Active Tab Only" setting controls whether windows are hidden when their parent session isn't active
- Windows with unknown parents (discovery pending) are always shown
- Minimized windows stay minimized regardless of visibility setting
**Tab Badge System**:
- Minimized subagents appear as a badge on their parent session's tab
- Badge dropdown allows restoring or permanently dismissing each agent
- `minimizedSubagents` Map tracks agentIds per sessionId
**Parent Matching Algorithm** (to handle multiple sessions with same workingDir):
1. **Strategy 1 - Sibling matching**: If another subagent with the same Claude `sessionId` already has a parent, use that same parent. This ensures all subagents from the same Claude session go to the same Claudeman session.
2. **Strategy 2 - WorkingDir matching with heuristics**: Find all sessions matching by `workingDir`. If multiple match, prefer the most recently created session (newest session likely spawned the subagent).
**Session Rename Handling**: When a session is renamed, `updateSubagentParentNames()` updates all subagent objects and their window headers via the `session:updated` SSE handler.
**Key Frontend Data Structures** (`app.js`):
```javascript
this.subagents = new Map(); // agentId → SubagentInfo (includes parentSessionId, parentSessionName)
this.subagentWindows = new Map(); // agentId → { element, minimized, hidden, position }
this.minimizedSubagents = new Map(); // sessionId → Set<agentId> (for tab badges)
```
**Implementation Files**:
- Server: `src/subagent-watcher.ts` (discovery, monitoring, events)
- Server: `src/web/server.ts` (SSE broadcast, REST endpoints)
- Frontend: `src/web/public/app.js:6183-6550` (window management, parent discovery, visibility)
### Session Modes
Sessions have a `mode` property (`SessionMode` type):
- **`'claude'`**: Runs Claude CLI for AI interactions (default)
- **`'shell'`**: Runs a plain bash shell for debugging/testing
### Screen-Aware Sessions
All Claude sessions spawned by Claudeman receive environment variables indicating they're running in a managed screen:
| Variable | Value | Purpose |
|----------|-------|---------|
| `CLAUDEMAN_SCREEN` | `1` | Indicates session is managed by Claudeman |
| `CLAUDEMAN_SESSION_ID` | `<uuid>` | Unique session identifier |
| `CLAUDEMAN_SCREEN_NAME` | `claudeman-<name>` | GNU screen session name |
This prevents Claude from accidentally killing its own screen session. The default CLAUDE.md template includes guidance about this.
**Implementation**: Set in `screen-manager.ts:createScreen()` for screen-based sessions and `session.ts:startInteractive()`/`startShell()` for PTY-only sessions. Both paths also augment `PATH` with the claude binary's directory to ensure discovery in restricted environments (systemd, non-login shells).
## Code Patterns
### Memory Leak Prevention
The frontend (`app.js`) runs for extended periods and must avoid memory leaks. Key patterns:
**SSE Reconnection**: When EventSource reconnects, all event listeners are re-registered on the new instance. The old EventSource is closed, but any orphaned reconnect timeouts must be cleared:
```javascript
// Clear pending reconnect timeout before creating new connection
if (this.sseReconnectTimeout) {
clearTimeout(this.sseReconnectTimeout);
this.sseReconnectTimeout = null;
}
```
**Cleanup on Init**: When `handleInit()` is called (SSE reconnect), clear all Maps and timers that could contain stale data:
- `idleTimers` - Clear all timeouts, then clear the Map
- `subagentActivity`, `subagentToolResults` - Clear to remove stale agent data
- `pendingHooks`, `tabAlerts`, `_shownCompletions` - Clear state tracking Sets
**Interval Management**: Always provide a stop method for intervals:
```javascript
startSystemStatsPolling() {
this.stopSystemStatsPolling(); // Clear existing before starting
this.systemStatsInterval = setInterval(...);
}
stopSystemStatsPolling() {
if (this.systemStatsInterval) {
clearInterval(this.systemStatsInterval);
this.systemStatsInterval = null;
}
}
```
**Session Cleanup**: When a session is deleted, clean up ALL associated resources:
- Respawn state (`respawnStatus`, `respawnTimers`, `respawnCountdownTimers`, `respawnActionLogs`)
- Subagent data (windows, activity, tool results)
- Timers (idle timers, pending hooks)
**Subagent Data**: Clean up activity/toolResults for completed agents after 5 minutes to prevent unbounded growth during long sessions.
### Pre-compiled Regex Patterns
For performance, regex patterns that are used frequently should be compiled once at module level:
```typescript
// Good - compile once
const ANSI_ESCAPE_PATTERN = /\x1b\[[0-9;]*m/g;
const TOKEN_PATTERN = /(\d+(?:\.\d+)?)\s*([kKmM])?\s*tokens/;
// Bad - recompiles on each call
function parse(line: string) {
return line.replace(/\x1b\[[0-9;]*m/g, '');
}
```
### Claude Message Parsing
Claude CLI outputs newline-delimited JSON. Strip ANSI codes before parsing:
```typescript
const cleanLine = line.replace(ANSI_ESCAPE_PATTERN, '');
const msg = JSON.parse(cleanLine) as ClaudeMessage;
// msg.type: 'system' | 'assistant' | 'user' | 'result'
// msg.message?.content: Array<{ type: 'text', text: string }>
// msg.total_cost_usd: number (on result messages)
```
### PTY Spawn Modes
All PTY spawns pass `PATH: getAugmentedPath()` in the env to ensure `claude` is discoverable even when the server runs in a restricted environment (e.g., systemd). The `getAugmentedPath()` function (in `session.ts`) resolves the claude binary's directory once at startup and prepends it to PATH. The `screen-manager.ts` equivalent (`findClaudeDir()`) does the same for screen-based spawns via `export PATH="<dir>:$PATH"` in the bash command.
```typescript
// One-shot mode (JSON output for token tracking)
pty.spawn('claude', ['-p', '--dangerously-skip-permissions', '--output-format', 'stream-json', prompt], {
env: { ...process.env, PATH: getAugmentedPath(), ... }
})
// Interactive mode (tokens parsed from status line)
pty.spawn('claude', ['--dangerously-skip-permissions'], {
env: { ...process.env, PATH: getAugmentedPath(), ... }
})
// Shell mode (debugging/testing - no Claude CLI)
pty.spawn('bash', [], { ... })
```
**PATH resolution search order** (both `session.ts` and `screen-manager.ts`):
1. `which claude` (respects current PATH)
2. `~/.local/bin/claude`
3. `~/.claude/local/claude`
4. `/usr/local/bin/claude`
5. `~/.npm-global/bin/claude`
6. `~/bin/claude`
### Sending Input to Sessions
Two methods:
1. **`session.write(data)`** - Direct PTY write (used by `/api/sessions/:id/input` endpoint)
2. **`session.writeViaScreen(data)`** - Via GNU screen (RECOMMENDED for programmatic input). Used by RespawnController, auto-compact, auto-clear.
**How `writeViaScreen` works internally** (in `screen-manager.ts:sendInput`):
1. Strips all `\n` newlines and `\r` carriage returns from text
2. Sends text first: `screen -S name -p 0 -X stuff "text"`
3. Sends Enter separately: `screen -S name -p 0 -X stuff "$(printf '\015')"`
**Why separate commands?** Claude CLI uses Ink (React for terminals) which requires text and Enter as separate `screen -X stuff` commands. Combining them doesn't work. This is a critical implementation detail when debugging input issues.
**IMPORTANT**: All prompts sent via `writeViaScreen` (respawn updatePrompt, kickstartPrompt, auto-compact, etc.) must be **single-line**. Newlines are stripped, so multi-line prompts become one long line.
### Idle Detection
**Session**: emits `idle`/`working` events on prompt detection + 2s activity timeout.
**RespawnController**: Multi-layer detection (completion message → AI idle check → output silence → token stability → working pattern absence). See `docs/respawn-state-machine.md` for full details.
### Token Tracking
- **One-shot mode**: Uses `--output-format stream-json` for detailed token usage from JSON
- **Interactive mode**: Parses tokens from Claude's status line (e.g., "123.4k tokens"), estimates 60/40 input/output split
### Auto-Compact & Auto-Clear
| Feature | Default Threshold | Action |
|---------|------------------|--------|
| Auto-Compact | 110k tokens | `/compact` with optional prompt |
| Auto-Clear | 140k tokens | `/clear` to reset context |
Both wait for idle. Configure via `session.setAutoCompact()` / `session.setAutoClear()`.
### Ralph / Todo Tracking
Detects Ralph loops and todos inside Claude sessions. **Disabled by default** but auto-enables when Ralph-related patterns are detected (promise tags, TodoWrite, iteration patterns, etc.). See `ralph-tracker.ts:shouldAutoEnable()` for the full pattern list.
**Auto-Loading @fix_plan.md**: When a session starts, the Ralph tracker automatically:
1. Loads existing `@fix_plan.md` from the working directory (if present)
2. Imports todos and shows them in the Ralph panel
3. Watches the file for changes and reloads when modified by Claude
4. Auto-enables the tracker if todos are found
This means you can create a session in a folder with an existing `@fix_plan.md` and the tracker will immediately show and track those tasks.
**Auto-Configuration from Ralph Plugin State**: When a session starts, Claudeman reads `.claude/ralph-loop.local.md` to auto-configure:
```yaml
---
enabled: true
iteration: 5
max-iterations: 50
completion-promise: "COMPLETE"
---
```
Priority: 1) `.claude/ralph-loop.local.md` (official Ralph plugin state), 2) `CLAUDE.md` `<promise>` tags (fallback). See `src/ralph-config.ts`.
**Completion Detection** (multi-strategy):
- 1st occurrence of `<promise>PHRASE</promise>`: Stores as expected phrase (likely in prompt)
- 2nd occurrence: Emits `completionDetected` event (actual completion)
- **Bare phrase detection**: Also detects phrase without tags once expected phrase is known
- **All complete detection**: When "All X files/tasks created/completed" detected, marks all todos complete and emits completion
- If loop is already active (via `/ralph-loop:ralph-loop`): Emits immediately on first occurrence
**Session Lifecycle**: Each session has its own independent tracker:
- New session → Fresh tracker (no carryover)
- Close tab → Tracker state cleared, UI panel hides
- `tracker.reset()` → Clears todos/state, keeps enabled status
- `tracker.fullReset()` → Complete reset to initial state
- `tracker.configure({ enabled?, completionPhrase?, maxIterations? })` → Partial config update
**API**:
- `GET /api/sessions/:id/ralph-state` - Get loop state and todos
- `POST /api/sessions/:id/ralph-config` - Configure tracker:
- `{ enabled: boolean }` - Enable/disable
- `{ reset: true }` - Soft reset (keep enabled)
- `{ reset: "full" }` - Full reset
### Terminal Display Fix
Tab switch/new session fix: clear xterm → write buffer → resize PTY → Ctrl+L redraw. Uses `pendingCtrlL` Set, triggered on `session:idle`/`session:working` events.
### SSE Events
All events broadcast to `/api/events` with format: `{ type: string, sessionId?: string, data: any }`.
Event prefixes: `session:`, `task:`, `respawn:`, `spawn:`, `subagent:`, `hook:`, `scheduled:`, `case:`, `screen:`, `init`.
Key events (see `app.js:handleSSEEvent()`):
- `session:idle`, `session:working` - Status indicators
- `session:terminal`, `session:clearTerminal` - Terminal content
- `session:ralphLoopUpdate`, `session:ralphTodoUpdate`, `session:ralphCompletionDetected` - Ralph tracking
- `respawn:detectionUpdate` - Idle detection status
- `spawn:queued`, `spawn:started`, `spawn:completed`, `spawn:failed` - Agent lifecycle
- `subagent:discovered`, `subagent:tool_call`, `subagent:progress`, `subagent:message`, `subagent:completed` - Claude Code background agents
- `hook:idle_prompt`, `hook:permission_prompt`, `hook:elicitation_dialog`, `hook:stop` - Claude Code hooks
### Run Summary
Per-session event tracking for "what happened while you were away" view. Click the chart icon (📊) on any session tab to view.
**Tracked Events**: session start/stop, respawn cycles, state changes, idle/working transitions, token milestones (every 50k), auto-compact/clear, Ralph completions, AI check results, hook events, errors/warnings, state stuck warnings (>10min same state).
**Stats**: respawn cycles, peak tokens, active time, idle time, error/warning counts.
**Storage**: In-memory only (not persisted). Fresh tracker created per session; cleared when session is deleted.
**Implementation**: `RunSummaryTracker` class in `src/run-summary.ts`, integrated via `setupSessionListeners()` and `setupRespawnListeners()` in `server.ts`.
### Frontend (app.js)
Vanilla JS + xterm.js. 60fps rendering: server batches terminal data every 16ms, client uses `requestAnimationFrame` to batch xterm.js writes.
### HTTPS & Browser Notifications
**HTTPS**: The `--https` flag generates/reuses self-signed certificates in `~/.claudeman/certs/`. Required for the Web Notification API.
**Notifications** (`NotificationManager` in `app.js`): In-app drawer, tab title flashing, Web Notification API (rate limited 3s), audio alerts (critical only), tab blinking (red=action, yellow=idle).
**Hook Event Data**: `/api/hook-event` forwards `data` field into SSE broadcast. Hook events: `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `stop`.
### State Store
Writes debounced (500ms) to `~/.claudeman/state.json` via `persistSessionState()` on every meaningful change.
**Per-session fields stored** (`SessionState` in `types.ts`):
- `id`, `pid`, `status`, `name`, `mode` - Core identity
- `workingDir`, `createdAt`, `lastActivityAt` - Location and timestamps
- `autoClearEnabled/Threshold`, `autoCompactEnabled/Threshold/Prompt` - Context management
- `ralphEnabled`, `ralphCompletionPhrase` - Ralph tracker state
- `respawnEnabled`, `respawnConfig` - Respawn controller state
- `totalCost`, `inputTokens`, `outputTokens` - Token tracking
- `parentAgentId`, `childAgentIds` - Spawn agent tree
CLI commands (`claudeman status/session list`) read from `state.json` to display web-managed sessions.
### TypeScript Config
Module resolution: NodeNext. Target: ES2022. Strict mode enabled with all additional strictness flags (`noUnusedLocals`, `noUnusedParameters`, `noImplicitReturns`, `noFallthroughCasesInSwitch`, etc.). No ESLint/Prettier configured - rely on TypeScript strict mode.
TUI uses React JSX (`jsx: react-jsx`, `jsxImportSource: react`) for Ink components.
## Adding New Features
- **API endpoint**: Add types in `types.ts`, route in `server.ts:buildServer()`, use `createErrorResponse()` for errors
- **SSE event**: Emit via `broadcast()` in server.ts, handle in `app.js:handleSSEEvent()` switch
- **Session event**: Add to `SessionEvents` interface in `session.ts`, emit via `this.emit()`, subscribe in server.ts, handle in frontend
- **Session setting**: Add field to `SessionState` in `types.ts`, include in `session.toState()`, call `this.persistSessionState(session)` in server.ts after the change
- **MCP tool**: Add tool definition in `mcp-server.ts` using `server.tool()`, use `apiRequest()` to call Claudeman REST API
- **New test file**: Create `test/<name>.test.ts`, pick unique port (next available: 3190), add to port allocation table above
### API Error Codes
Use `createErrorResponse(code, details?)` from `types.ts`:
| Code | Use Case |
|------|----------|
| `NOT_FOUND` | Session/resource doesn't exist |
| `INVALID_INPUT` | Bad request parameters |
| `SESSION_BUSY` | Session is currently processing |
| `OPERATION_FAILED` | Action couldn't complete |
| `ALREADY_EXISTS` | Duplicate resource |
| `INTERNAL_ERROR` | Unexpected server error |
## Session Lifecycle & Cleanup
- **Limit**: Web server: `MAX_CONCURRENT_SESSIONS = 50` (`server.ts:56`), UI tab limit: 20, CLI default: 5 (`types.ts:DEFAULT_CONFIG`)
- **Kill** (`killScreen()`): child PIDs → process group → screen quit → SIGKILL
- **Ghost discovery**: `reconcileScreens()` finds orphaned screens on startup
- **Cleanup** (`cleanupSession()`): stops respawn, clears buffers/timers, kills screen, removes from `state.json`
- **State sync**: Every session create/delete/update calls `persistSessionState()` which writes full state (including respawn config from controller) to `state.json`
- **Recovery on restart**: Server reads `state.json` first (has all settings), falls back to `screens.json` for any sessions not found. Settings (auto-compact, auto-clear, respawn config, Ralph state) restored to live session objects after screen reattachment.
## TUI (WIP)
Ink/React-based TUI in `src/tui/`. Client to the web server, uses `/api/*` endpoints and attaches to screens via GNU screen.
**Key files**: `App.tsx` (main component), `DirectAttach.ts` (full-screen attach with tab switching), `SessionList.tsx`, `SessionView.tsx`.
**Current state**: Basic session list and direct attach work. Missing: full session management UI, settings panel, notifications.
## Buffer Limits
| Buffer | Max Size | Trim To |
|--------|----------|---------|
| Terminal | 2MB | 1.5MB |
| Text output | 1MB | 768KB |
| Messages | 1000 | 800 |
| Line buffer | 64KB | (flushed every 100ms) |
| Respawn buffer | 1MB | 512KB |
Tab switch uses `tail=256KB` for fast initial load, then chunked 64KB writes via `requestAnimationFrame`.
## API Routes
All routes defined in `server.ts:buildServer()`. Key endpoint groups:
- `/api/events` - SSE stream | `/api/status` - Full app state
- `/api/sessions` - CRUD + `/input`, `/resize`, `/interactive`
- `/api/sessions/:id/respawn/*` - Start/stop/enable/config respawn controller
- `/api/sessions/:id/ralph-*` - Ralph tracker config and state
- `/api/sessions/:id/auto-compact`, `/auto-clear` - Token threshold settings
- `/api/quick-start` - Create case + start session (`{mode?: 'claude'|'shell'}`)
- `/api/cases`, `/api/screens` - Case and screen management
- `/api/spawn/*` - Agent lifecycle (list, status, result, messages, cancel, trigger)
- `/api/subagents` - List/get/kill Claude Code background agents, get transcripts
- `/api/sessions/:id/subagents` - Get subagents for a specific session's working directory
- `/api/sessions/:id/run-summary` - Get run summary (events, stats) for "what happened while away"
- `/api/hook-event` - Claude Code hook callbacks (`{event, sessionId, data?}`)
- **API endpoint**: Types in `types.ts`, route in `server.ts:buildServer()`, use `createErrorResponse()`
- **SSE event**: Emit via `broadcast()`, handle in `app.js:handleSSEEvent()`
- **Session setting**: Add to `SessionState` in `types.ts`, include in `session.toState()`, call `persistSessionState()`
- **New test**: Pick unique port (see below), add port comment to test file header
## State Files ## State Files
| File | Purpose | | File | Purpose |
|------|---------| |------|---------|
| `~/.claudeman/state.json` | Full session state (all settings, tokens, respawn config, Ralph state), tasks, app config | | `~/.claudeman/state.json` | Sessions, settings, tokens, respawn config |
| `~/.claudeman/state-inner.json` | Ralph loop/todo state per session (separate to reduce writes) | | `~/.claudeman/screens.json` | Screen metadata for recovery |
| `~/.claudeman/screens.json` | Screen session metadata (for recovery after restart) | | `~/.claudeman/settings.json` | User preferences |
| `~/.claudeman/settings.json` | User preferences (lastUsedCase, custom template path, subagentTrackingEnabled) |
| `~/.claudeman/certs/` | Self-signed TLS certificates for `--https` mode |
**Recovery**: On restart, sessions restored from `state.json` (primary) with `screens.json` as fallback. All settings re-applied to live sessions. Cases created in `~/claudeman-cases/` by default. ## Testing
### CLAUDE.md Templates **Port allocation**: E2E tests use centralized ports in `test/e2e/e2e.config.ts` (E2E_PORTS: 3183-3190). Unit/integration tests pick unique ports manually (up to 3157). Search `const PORT =` or `TEST_PORT` in test files to see used ports. **Next available: 3191**
New cases get a CLAUDE.md generated from `src/templates/case-template.md` (bundled with the project). Template resolution order: **E2E tests**: Use Playwright. Run `npx playwright install chromium` first. See `test/e2e/fixtures/` for helpers. E2E config provides ports, timeouts, and helpers.
1. Custom path from `~/.claudeman/settings.json` (`defaultClaudeMdPath` field)
2. Bundled `case-template.md` (copied to `dist/templates/` during build)
3. Minimal fallback (if bundled template is missing)
Placeholders replaced: **Test config**: Vitest runs with `globals: true` (no imports needed for `describe`/`it`/`expect`) and `fileParallelism: false` (files run sequentially to respect screen limits).
- `[PROJECT_NAME]` → Case name
- `[PROJECT_DESCRIPTION]` → Description
- `[DATE]` → Current date (YYYY-MM-DD)
## Screen Session Manager (CLI Tool) **Test safety**: `test/setup.ts` provides:
- Screen concurrency limiter (max 10)
- Pre-existing screen protection (never kills screens present before tests)
- Tracked resource cleanup (only kills screens/processes tests register)
- Safe to run from within Claudeman-managed sessions
`./scripts/screen-manager.sh` - Interactive bash tool for managing screen sessions. Commands: `list`, `attach N`, `kill N,M`, `kill-all`, `info N`. Requires `jq` and `screen`. Respawn tests use MockSession to avoid spawning real Claude processes.
## Web UI Keyboard Shortcuts ## Debugging
| Shortcut | Action |
|----------|--------|
| `Ctrl+Enter` | Quick-start session |
| `Ctrl+W` | Close session |
| `Ctrl+Tab` | Next session |
| `Ctrl+K` | Kill all sessions |
| `Ctrl+L` | Clear terminal |
## Documentation
**Reference docs** (read these for deep dives):
- `docs/respawn-state-machine.md` - Respawn controller states, idle detection, auto-accept
- `docs/spawn-protocol.md` - Spawn1337 agent protocol, MCP tools, resource governance
- `docs/ralph-wiggum-guide.md` - Ralph Wiggum loop guide (plugin reference, prompt templates)
- `docs/claude-code-hooks-reference.md` - Claude Code hooks documentation
- `docs/browser-testing-guide.md` - Browser testing frameworks comparison, patterns, and known issues
**Internal planning** (historical context, may be outdated):
- `docs/respawn-improvement-plan.md` - Planned respawn improvements
- `docs/run-summary-plan.md` - Run summary feature design
### Ralph Wiggum Loops
**Core Pattern**: `<promise>PHRASE</promise>` - The completion signal that tells the loop to stop.
**Skill Commands**:
```bash ```bash
/ralph-loop:ralph-loop # Start Ralph Loop in current session screen -ls # List screens
/ralph-loop:cancel-ralph # Cancel active Ralph Loop screen -r <name> # Attach (Ctrl+A D to detach)
/ralph-loop:help # Show help and usage curl localhost:3000/api/sessions # Check sessions
curl localhost:3000/api/status | jq # Full app state
cat ~/.claudeman/state.json | jq # View persisted state
``` ```
The `RalphTracker` class (`src/ralph-tracker.ts`) detects Ralph patterns in Claude output and tracks loop state, todos, and completion phrases. It auto-enables when Ralph-related patterns are detected. ## Performance Constraints
**Ralph Loop Wizard**: The web UI wizard (`app.js:startRalphLoop()`) provides a guided setup for Ralph Loops. It: The app must stay fast with 20 sessions and 50 agent windows:
1. Creates/selects a case and starts a session - 60fps terminal (16ms batching + `requestAnimationFrame`)
2. Configures the Ralph tracker with completion phrase and max iterations - Auto-trimming buffers (2MB terminal max)
3. Optionally generates a task plan (`@fix_plan.md`) - Debounced state persistence (500ms)
4. Sends the initial prompt with iteration protocol - SSE batching (16ms)
**Plan Generation Modes** (Step 2 of wizard): ## Buffer Limits
| Mode | Description | API Endpoint |
|------|-------------|--------------|
| **Brief** | High-level milestones only | `/api/generate-plan` |
| **Standard** | Balanced implementation steps | `/api/generate-plan` |
| **Enhanced** | Multi-agent orchestration with verification | `/api/generate-plan-detailed` |
**Enhanced Plan Generation** (`src/plan-orchestrator.ts`): When "Enhanced" mode is selected, the plan is generated using parallel subagent orchestration: | Buffer | Max | Trim To |
1. **Phase 1 - Parallel Analysis**: Spawns 4 specialist subagents simultaneously: |--------|-----|---------|
- Requirements Analyst → Extracts explicit/implicit requirements | Terminal | 2MB | 1.5MB |
- Architecture Planner → Identifies modules, interfaces, types | Text output | 1MB | 768KB |
- TDD Specialist → Designs test-first approach, edge cases | Messages | 1000 | 800 |
- Risk Analyst → Identifies failure points, dependencies, blockers
2. **Phase 2 - Synthesis**: Merges outputs, deduplicates, orders by dependency
3. **Phase 3 - Verification**: Review subagent validates plan, assigns P0/P1/P2 priorities, identifies gaps
The enhanced mode takes longer (~60-90s) but produces more thorough plans with quality scores. Switching between modes auto-regenerates the plan. ## Where to Find More Information
**Respawn for Ralph Loops**: Disabled by default (checkbox unchecked). When enabled, the respawn controller uses Ralph-specific prompts: | Topic | Location |
- **Update Prompt**: Instructs Claude to document progress to CLAUDE.md, update planning files (`@fix_plan.md`), mark completed tasks, and write a summary before `/clear` |-------|----------|
- **Kickstart Prompt**: After `/init`, tells Claude it's in a Ralph Wiggum loop and to continue work by reading `@fix_plan.md` and CLAUDE.md notes, then resume on uncompleted tasks | **Respawn state machine** | `docs/respawn-state-machine.md` |
| **Spawn agent protocol** | `docs/spawn-protocol.md` |
| **Ralph Loop guide** | `docs/ralph-wiggum-guide.md` |
| **Claude Code hooks** | `docs/claude-code-hooks-reference.md` |
| **Browser/E2E testing** | `docs/browser-testing-guide.md` |
| **API routes** | `src/web/server.ts:buildServer()` or README.md |
| **SSE events** | Search `broadcast(` in `server.ts` |
| **CLI commands** | `claudeman --help` |
| **Frontend patterns** | `src/web/public/app.js` (subagent windows, notifications) |
| **Session modes** | `SessionMode` type in `src/types.ts` |
| **Error codes** | `createErrorResponse()` in `src/types.ts` |
| **Test fixtures** | `test/e2e/fixtures/` |
| **Test utilities** | `test/respawn-test-utils.ts` |
| **Keyboard shortcuts** | README.md or App Settings in web UI |
+247
View File
@@ -0,0 +1,247 @@
# Ralph Loop Plan Improvement Roadmap
> Research-backed improvements for rock-solid AI planning with auto-improvement capabilities.
**Created**: 2026-01-27
**Status**: Implementation in Progress
---
## Table of Contents
1. [Research Summary](#research-summary)
2. [Current State Analysis](#current-state-analysis)
3. [Proposed Improvements](#proposed-improvements)
4. [Implementation Plan](#implementation-plan)
5. [Sources](#sources)
---
## Research Summary
### Key Insights from Industry Best Practices
#### 1. Self-Verification is Critical
> "Claude performs dramatically better when it can verify its own work, like run tests, compare screenshots, and validate outputs. Without clear success criteria, it might produce something that looks right but actually doesn't work."
> — [Anthropic Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
#### 2. Iterative Refinement Patterns (AWS)
> "A generator agent produces output, an evaluator agent reviews using evaluation rubric, and based on feedback, an optimizer agent revises the output. Loop repeats until criteria met."
> — [AWS Agentic AI Patterns](https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-patterns/evaluator-reflect-refine-loop-patterns.html)
#### 3. Dynamic Task Decomposition (TDAG Framework)
> "Dynamically decomposes complex tasks into smaller subtasks and assigns each to a specifically generated subagent, enhancing adaptability in diverse and unpredictable real-world tasks."
> — [TDAG Framework - arXiv](https://arxiv.org/abs/2402.10178)
#### 4. Multi-Stage Verification Workflow
> "o3: Generate plan → Sonnet: Verify and create task list → Sonnet: Execute → Sonnet: Verify against plan → o3: Final verification → Issues bake back into plan"
> — [Claude Code Best Practices Community](https://rosmur.github.io/claudecode-best-practices/)
#### 5. Self-Improving Agents
> "Through an iterative refinement process (analyze outcome → adjust approach → try again), the agent becomes more adept at handling tasks over time. It effectively builds a growing knowledge base of what strategies work best."
> — [Self-Improving Data Agents](https://powerdrill.ai/blog/self-improving-data-agents)
#### 6. Memory Architecture for Planning
> "Agents use three memory layers: working memory for short-lived calculations, episodic memory for step-by-step histories, and semantic memory for long-term knowledge."
> — [LLM Agent Research](https://www.promptingguide.ai/research/llm-agents)
---
## Current State Analysis
### What We Have
The current plan generation system (`/api/generate-plan` and `/api/generate-plan-detailed`):
1. **Standard Mode**: Single Opus 4.5 call with TDD-focused prompt
2. **Enhanced Mode**: 4 parallel subagents (Requirements, Architecture, Testing, Risks) + Verification
### Current Plan Item Structure
```json
{
"content": "Implement login endpoint",
"priority": "P0"
}
```
### Limitations
| Issue | Impact |
|-------|--------|
| No verification criteria | Can't automatically validate completion |
| No test pairing | TDD not enforced structurally |
| Static plans | No adaptation during execution |
| No dependencies | Can't track blocking relationships |
| No failure tracking | Same errors repeat |
| No checkpoints | Plans run until completion or failure |
---
## Proposed Improvements
### Enhanced Plan Item Structure
```typescript
interface EnhancedPlanItem {
id: string; // Unique identifier (e.g., "P0-001")
content: string; // Task description
priority: 'P0' | 'P1' | 'P2'; // Criticality
phase: 'setup' | 'test' | 'impl' | 'verify'; // Development phase
// NEW: Verification
verificationCriteria: string; // How to know it's done
testCommand?: string; // Command to run for verification
// NEW: Dependencies
dependencies: string[]; // IDs of tasks that must complete first
blockedBy?: string[]; // Runtime: tasks blocking this one
// NEW: Execution tracking
status: 'pending' | 'in_progress' | 'completed' | 'failed' | 'blocked';
attempts: number; // How many times attempted
lastError?: string; // Most recent failure reason
completedAt?: number; // Timestamp of completion
// NEW: Metadata
estimatedComplexity: 'low' | 'medium' | 'high';
rollbackStrategy?: string; // How to undo if needed
version: number; // Plan version this belongs to
}
```
### Runtime Plan Adaptation Flow
```
┌─────────────────────────────────────────────────────────────────┐
│ RUNTIME PLAN LOOP │
├─────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ Execute │──▶│ Verify │──▶│ Success? │──▶│ Mark │ │
│ │ Task │ │ Output │ │ │ │ Complete │ │
│ └──────────┘ └──────────┘ └────┬─────┘ └──────────┘ │
│ │ No │
│ ▼ │
│ ┌──────────┐ │
│ │ Analyze │ │
│ │ Failure │ │
│ └────┬─────┘ │
│ │ │
│ ┌──────────────┼──────────────┐ │
│ ▼ ▼ ▼ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ Retry │ │ Add Fix │ │ Escalate │ │
│ │ (< 3x) │ │ Sub-Task │ │ BLOCKED │ │
│ └──────────┘ └──────────┘ └──────────┘ │
│ │
└───────────────────────────────────────────────────────────────┘
```
### Checkpoint Review System
At iterations 5, 10, 20, 30, 50:
1. Pause execution
2. Summarize progress (completed/failed/pending)
3. Identify stuck items (3+ failures)
4. Generate alternative approaches for stuck items
5. Update plan with new strategies
6. Continue with refined plan
---
## Implementation Plan
### Phase 1: Quick Wins (Implementing Now)
#### 1.1 Add Verification Criteria to Plan Items
- Modify plan generation prompts to require `verificationCriteria`
- Update `PlanItem` interface in `types.ts`
- Update plan orchestrator prompts
#### 1.2 Pair Test/Implementation Steps
- Ensure every implementation step has a corresponding test step
- Group items: test → implement → verify
- Add phase field to track TDD cycle
#### 1.3 Checkpoint Review Prompts
- Add checkpoint logic to Ralph tracker
- At iterations 5, 10, 20: inject review prompt
- Generate progress summary and stuck item analysis
### Phase 2: Medium Effort (Implementing Now)
#### 2.1 Failure Tracking
- Track `attempts` and `lastError` per task
- After 3 failures, auto-generate debug sub-task
- Record failure patterns in plan history
#### 2.2 Plan Versioning
- Add `version` field to plans
- Keep history in `@fix_plan.md` with version markers
- Allow rollback to previous versions
- Track which version each task belongs to
#### 2.3 Dependency Tracking
- Add `dependencies` field to plan items
- Validate dependency graph (no cycles)
- Block tasks until dependencies complete
- Show dependency status in UI
### Phase 3: Future Enhancements
#### 3.1 Full Runtime Adaptation
- TDAG-style dynamic decomposition
- Auto-generate sub-tasks for complex items
- Learning from failure patterns
#### 3.2 Multi-Model Verification
- Haiku: Fast initial generation
- Sonnet: Verification and refinement
- Opus: Final quality check
#### 3.3 Plan Memory System
- Episodic memory: What worked/failed in this session
- Semantic memory: Patterns across projects
- Use for future plan generation
---
## File Changes Required
### New/Modified Files
| File | Changes |
|------|---------|
| `src/types.ts` | Add `EnhancedPlanItem` interface |
| `src/plan-orchestrator.ts` | Update prompts, add versioning |
| `src/ralph-tracker.ts` | Add checkpoint logic, failure tracking |
| `src/web/server.ts` | New endpoints for plan updates |
| `src/web/public/app.js` | UI for enhanced plan display |
### New Endpoints
| Method | Endpoint | Purpose |
|--------|----------|---------|
| PATCH | `/api/sessions/:id/plan/task/:taskId` | Update task status |
| POST | `/api/sessions/:id/plan/checkpoint` | Trigger checkpoint review |
| GET | `/api/sessions/:id/plan/history` | Get plan version history |
| POST | `/api/sessions/:id/plan/rollback/:version` | Rollback to version |
---
## Sources
- [Anthropic Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices)
- [AWS Agentic AI Patterns](https://docs.aws.amazon.com/prescriptive-guidance/latest/agentic-ai-patterns/evaluator-reflect-refine-loop-patterns.html)
- [TDAG: Multi-Agent Task Decomposition Framework](https://arxiv.org/abs/2402.10178)
- [Self-Improving Data Agents](https://powerdrill.ai/blog/self-improving-data-agents)
- [OpenAI Self-Evolving Agents Cookbook](https://cookbook.openai.com/examples/partners/self_evolving_agents/autonomous_agent_retraining)
- [Task Decomposition for Coding Agents](https://mgx.dev/insights/task-decomposition-for-coding-agents-architectures-advancements-and-future-directions/)
- [Claude Code Best Practices Community Guide](https://rosmur.github.io/claudecode-best-practices/)
- [LLM Agents Prompt Engineering Guide](https://www.promptingguide.ai/research/llm-agents)
- [Agentic AI Implementation Guide](https://www.sketchdev.io/blog/agentic-ai-implementation-guide)
---
*This document is part of the Claudeman project. See [CLAUDE.md](../CLAUDE.md) for main documentation.*
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -1,6 +1,6 @@
{ {
"name": "claudeman", "name": "claudeman",
"version": "0.1400", "version": "0.1402",
"description": "The missing control plane for Claude Code - run 20 autonomous agents with real-time monitoring and session persistence", "description": "The missing control plane for Claude Code - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module", "type": "module",
"main": "dist/index.js", "main": "dist/index.js",
+195 -14
View File
@@ -124,14 +124,29 @@ export interface DetailedPlanResult {
export type ProgressCallback = (phase: string, detail: string) => void; export type ProgressCallback = (phase: string, detail: string) => void;
/** Event types for plan subagent visibility */
export interface PlanSubagentEvent {
type: 'started' | 'progress' | 'completed' | 'failed';
agentId: string;
agentType: 'requirements' | 'architecture' | 'testing' | 'risks' | 'verification';
model: string;
status: string;
detail?: string;
itemCount?: number;
durationMs?: number;
error?: string;
}
export type SubagentCallback = (event: PlanSubagentEvent) => void;
// ============================================================================ // ============================================================================
// Constants // Constants
// ============================================================================ // ============================================================================
const SUBAGENT_TIMEOUT_MS = 45000; // 45 seconds per subagent const SUBAGENT_TIMEOUT_MS = 300000; // 5 minutes per subagent (Opus needs time for complex analysis)
const VERIFICATION_TIMEOUT_MS = 60000; // 60 seconds for verification const VERIFICATION_TIMEOUT_MS = 480000; // 8 minutes for verification (Opus + large plans)
const MODEL_ANALYSIS = 'haiku'; // Fast model for parallel analysis const MODEL_ANALYSIS = 'opus'; // Best model for thorough analysis
const MODEL_VERIFICATION = 'sonnet'; // Better reasoning for verification const MODEL_VERIFICATION = 'opus'; // Best model for verification
// ============================================================================ // ============================================================================
// Subagent Prompts // Subagent Prompts
@@ -375,12 +390,35 @@ Be critical but constructive. A thorough review catches issues that tests miss.`
export class PlanOrchestrator { export class PlanOrchestrator {
private screenManager: ScreenManager; private screenManager: ScreenManager;
private workingDir: string; private workingDir: string;
private runningSessions: Set<Session> = new Set();
private cancelled = false;
constructor(screenManager: ScreenManager, workingDir: string = process.cwd()) { constructor(screenManager: ScreenManager, workingDir: string = process.cwd()) {
this.screenManager = screenManager; this.screenManager = screenManager;
this.workingDir = workingDir; this.workingDir = workingDir;
} }
/**
* Cancel all running subagent sessions.
* Call this when the client disconnects or user clicks Stop.
*/
async cancel(): Promise<void> {
this.cancelled = true;
console.log(`[PlanOrchestrator] Cancelling ${this.runningSessions.size} running sessions...`);
const stopPromises = Array.from(this.runningSessions).map(async (session) => {
try {
await session.stop();
} catch {
// Ignore cleanup errors
}
});
await Promise.all(stopPromises);
this.runningSessions.clear();
console.log('[PlanOrchestrator] All sessions cancelled');
}
/** /**
* Generate a detailed implementation plan using subagent orchestration. * Generate a detailed implementation plan using subagent orchestration.
* *
@@ -391,7 +429,8 @@ export class PlanOrchestrator {
*/ */
async generateDetailedPlan( async generateDetailedPlan(
taskDescription: string, taskDescription: string,
onProgress?: ProgressCallback onProgress?: ProgressCallback,
onSubagent?: SubagentCallback
): Promise<DetailedPlanResult> { ): Promise<DetailedPlanResult> {
const startTime = Date.now(); const startTime = Date.now();
let totalCost = 0; let totalCost = 0;
@@ -399,7 +438,7 @@ export class PlanOrchestrator {
try { try {
// Phase 1: Parallel Analysis // Phase 1: Parallel Analysis
onProgress?.('parallel-analysis', 'Spawning analysis subagents...'); onProgress?.('parallel-analysis', 'Spawning analysis subagents...');
const subagentResults = await this.runParallelAnalysis(taskDescription, onProgress); const subagentResults = await this.runParallelAnalysis(taskDescription, onProgress, onSubagent);
totalCost += subagentResults.reduce((sum, r) => sum + (r.success ? 0.002 : 0), 0); // Estimate totalCost += subagentResults.reduce((sum, r) => sum + (r.success ? 0.002 : 0), 0); // Estimate
@@ -421,7 +460,8 @@ export class PlanOrchestrator {
const verificationResult = await this.runVerification( const verificationResult = await this.runVerification(
taskDescription, taskDescription,
synthesisResult.items, synthesisResult.items,
onProgress onProgress,
onSubagent
); );
totalCost += 0.01; // Verification cost estimate totalCost += 0.01; // Verification cost estimate
@@ -462,7 +502,8 @@ export class PlanOrchestrator {
*/ */
private async runParallelAnalysis( private async runParallelAnalysis(
taskDescription: string, taskDescription: string,
onProgress?: ProgressCallback onProgress?: ProgressCallback,
onSubagent?: SubagentCallback
): Promise<SubagentResult[]> { ): Promise<SubagentResult[]> {
const subagents: Array<{ const subagents: Array<{
type: SubagentResult['agentType']; type: SubagentResult['agentType'];
@@ -476,7 +517,7 @@ export class PlanOrchestrator {
// Run all subagents in parallel // Run all subagents in parallel
const promises = subagents.map(({ type, prompt }) => const promises = subagents.map(({ type, prompt }) =>
this.runSubagent(type, prompt, onProgress) this.runSubagent(type, prompt, onProgress, onSubagent)
); );
return Promise.all(promises); return Promise.all(promises);
@@ -488,10 +529,34 @@ export class PlanOrchestrator {
private async runSubagent( private async runSubagent(
agentType: SubagentResult['agentType'], agentType: SubagentResult['agentType'],
prompt: string, prompt: string,
onProgress?: ProgressCallback onProgress?: ProgressCallback,
onSubagent?: SubagentCallback
): Promise<SubagentResult> { ): Promise<SubagentResult> {
const agentId = `plan-${agentType}-${Date.now()}`;
// Check if already cancelled
if (this.cancelled) {
return {
agentType,
items: [],
success: false,
error: 'Cancelled',
durationMs: 0,
};
}
const startTime = Date.now(); const startTime = Date.now();
// Emit started event
onSubagent?.({
type: 'started',
agentId,
agentType,
model: MODEL_ANALYSIS,
status: 'running',
detail: `Analyzing ${agentType}...`,
});
const session = new Session({ const session = new Session({
workingDir: this.workingDir, workingDir: this.workingDir,
screenManager: this.screenManager, screenManager: this.screenManager,
@@ -499,6 +564,9 @@ export class PlanOrchestrator {
mode: 'claude', mode: 'claude',
}); });
// Track this session for cancellation
this.runningSessions.add(session);
try { try {
onProgress?.('subagent', `Running ${agentType} analysis...`); onProgress?.('subagent', `Running ${agentType} analysis...`);
@@ -507,9 +575,38 @@ export class PlanOrchestrator {
this.timeout(SUBAGENT_TIMEOUT_MS), this.timeout(SUBAGENT_TIMEOUT_MS),
]); ]);
// Check if cancelled during execution
if (this.cancelled) {
onSubagent?.({
type: 'failed',
agentId,
agentType,
model: MODEL_ANALYSIS,
status: 'cancelled',
error: 'Cancelled',
durationMs: Date.now() - startTime,
});
return {
agentType,
items: [],
success: false,
error: 'Cancelled',
durationMs: Date.now() - startTime,
};
}
// Parse JSON from result // Parse JSON from result
const jsonMatch = result.match(/\[[\s\S]*\]/); const jsonMatch = result.match(/\[[\s\S]*\]/);
if (!jsonMatch) { if (!jsonMatch) {
onSubagent?.({
type: 'failed',
agentId,
agentType,
model: MODEL_ANALYSIS,
status: 'failed',
error: 'No JSON array found in response',
durationMs: Date.now() - startTime,
});
return { return {
agentType, agentType,
items: [], items: [],
@@ -521,6 +618,15 @@ export class PlanOrchestrator {
const parsed = JSON.parse(jsonMatch[0]); const parsed = JSON.parse(jsonMatch[0]);
if (!Array.isArray(parsed)) { if (!Array.isArray(parsed)) {
onSubagent?.({
type: 'failed',
agentId,
agentType,
model: MODEL_ANALYSIS,
status: 'failed',
error: 'Response is not an array',
durationMs: Date.now() - startTime,
});
return { return {
agentType, agentType,
items: [], items: [],
@@ -544,6 +650,17 @@ export class PlanOrchestrator {
onProgress?.('subagent', `${agentType} complete (${items.length} items)`); onProgress?.('subagent', `${agentType} complete (${items.length} items)`);
// Emit completed event
onSubagent?.({
type: 'completed',
agentId,
agentType,
model: MODEL_ANALYSIS,
status: 'completed',
itemCount: items.length,
durationMs: Date.now() - startTime,
});
return { return {
agentType, agentType,
items, items,
@@ -551,14 +668,26 @@ export class PlanOrchestrator {
durationMs: Date.now() - startTime, durationMs: Date.now() - startTime,
}; };
} catch (err) { } catch (err) {
const errorMsg = err instanceof Error ? err.message : String(err);
onSubagent?.({
type: 'failed',
agentId,
agentType,
model: MODEL_ANALYSIS,
status: 'failed',
error: errorMsg,
durationMs: Date.now() - startTime,
});
return { return {
agentType, agentType,
items: [], items: [],
success: false, success: false,
error: err instanceof Error ? err.message : String(err), error: errorMsg,
durationMs: Date.now() - startTime, durationMs: Date.now() - startTime,
}; };
} finally { } finally {
// Remove from tracking and clean up
this.runningSessions.delete(session);
try { try {
await session.stop(); await session.stop();
} catch { } catch {
@@ -702,8 +831,28 @@ export class PlanOrchestrator {
private async runVerification( private async runVerification(
taskDescription: string, taskDescription: string,
synthesizedItems: PlanItem[], synthesizedItems: PlanItem[],
onProgress?: ProgressCallback onProgress?: ProgressCallback,
onSubagent?: SubagentCallback
): Promise<VerificationResult> { ): Promise<VerificationResult> {
const agentId = `plan-verification-${Date.now()}`;
// Check if already cancelled
if (this.cancelled) {
return this.fallbackVerification(synthesizedItems);
}
// Emit started event
onSubagent?.({
type: 'started',
agentId,
agentType: 'verification',
model: MODEL_VERIFICATION,
status: 'running',
detail: 'Validating and prioritizing plan...',
});
const startTime = Date.now();
const session = new Session({ const session = new Session({
workingDir: this.workingDir, workingDir: this.workingDir,
screenManager: this.screenManager, screenManager: this.screenManager,
@@ -711,6 +860,9 @@ export class PlanOrchestrator {
mode: 'claude', mode: 'claude',
}); });
// Track this session for cancellation
this.runningSessions.add(session);
try { try {
// Format plan for verification // Format plan for verification
const planText = synthesizedItems const planText = synthesizedItems
@@ -728,6 +880,11 @@ export class PlanOrchestrator {
this.timeout(VERIFICATION_TIMEOUT_MS), this.timeout(VERIFICATION_TIMEOUT_MS),
]); ]);
// Check if cancelled during execution
if (this.cancelled) {
return this.fallbackVerification(synthesizedItems);
}
// Parse JSON from result // Parse JSON from result
const jsonMatch = result.match(/\{[\s\S]*\}/); const jsonMatch = result.match(/\{[\s\S]*\}/);
if (!jsonMatch) { if (!jsonMatch) {
@@ -787,6 +944,17 @@ export class PlanOrchestrator {
onProgress?.('verification', `Verification complete (quality: ${Math.round((parsed.qualityScore || 0.8) * 100)}%)`); onProgress?.('verification', `Verification complete (quality: ${Math.round((parsed.qualityScore || 0.8) * 100)}%)`);
// Emit completed event
onSubagent?.({
type: 'completed',
agentId,
agentType: 'verification',
model: MODEL_VERIFICATION,
status: 'completed',
itemCount: validatedPlan.length,
durationMs: Date.now() - startTime,
});
return { return {
validatedPlan, validatedPlan,
gaps: Array.isArray(parsed.gaps) ? parsed.gaps.map(String) : [], gaps: Array.isArray(parsed.gaps) ? parsed.gaps.map(String) : [],
@@ -795,8 +963,19 @@ export class PlanOrchestrator {
}; };
} catch (err) { } catch (err) {
console.error('[PlanOrchestrator] Verification failed:', err); console.error('[PlanOrchestrator] Verification failed:', err);
onSubagent?.({
type: 'failed',
agentId,
agentType: 'verification',
model: MODEL_VERIFICATION,
status: 'failed',
error: err instanceof Error ? err.message : String(err),
durationMs: Date.now() - startTime,
});
return this.fallbackVerification(synthesizedItems); return this.fallbackVerification(synthesizedItems);
} finally { } finally {
// Remove from tracking and clean up
this.runningSessions.delete(session);
try { try {
await session.stop(); await session.stop();
} catch { } catch {
@@ -808,8 +987,10 @@ export class PlanOrchestrator {
/** /**
* Fallback verification when the verification subagent fails. * Fallback verification when the verification subagent fails.
* Assigns heuristic priorities and adds basic verification criteria. * Assigns heuristic priorities and adds basic verification criteria.
* Note: This is a silent fallback - no warning shown to user since the result is still useful.
*/ */
private fallbackVerification(items: PlanItem[]): VerificationResult { private fallbackVerification(items: PlanItem[]): VerificationResult {
console.log('[PlanOrchestrator] Using heuristic priorities (verification subagent timed out or failed)');
return { return {
validatedPlan: items.map((item, idx) => ({ validatedPlan: items.map((item, idx) => ({
...item, ...item,
@@ -824,8 +1005,8 @@ export class PlanOrchestrator {
version: 1, version: 1,
})), })),
gaps: [], gaps: [],
warnings: ['Verification subagent failed - using heuristic priorities'], warnings: [], // Silent fallback - heuristics are good enough, no need to alarm user
qualityScore: 0.7, qualityScore: 0.75, // Slightly higher since fallback still produces useful results
}; };
} }
+416 -1
View File
@@ -31,6 +31,72 @@ import {
createInitialCircuitBreakerStatus, createInitialCircuitBreakerStatus,
} from './types.js'; } from './types.js';
// ========== Enhanced Plan Task Interface ==========
/** Task execution status for plan tracking */
export type PlanTaskStatus = 'pending' | 'in_progress' | 'completed' | 'failed' | 'blocked';
/** TDD phase categories */
export type TddPhase = 'setup' | 'test' | 'impl' | 'verify' | 'review';
/**
* Enhanced plan task with verification criteria, dependencies, and execution tracking.
* Supports TDD workflow, failure tracking, and plan versioning.
*/
export interface EnhancedPlanTask {
/** Unique identifier (e.g., "P0-001") */
id: string;
/** Task description */
content: string;
/** Criticality level */
priority: 'P0' | 'P1' | 'P2' | null;
/** How to verify completion */
verificationCriteria?: string;
/** Command to run for verification */
testCommand?: string;
/** IDs of tasks that must complete first */
dependencies: string[];
/** Current execution status */
status: PlanTaskStatus;
/** How many times attempted */
attempts: number;
/** Most recent failure reason */
lastError?: string;
/** Timestamp of completion */
completedAt?: number;
/** Plan version this belongs to */
version: number;
/** TDD phase category */
tddPhase?: TddPhase;
/** ID of paired test/impl task */
pairedWith?: string;
/** Estimated complexity */
complexity?: 'low' | 'medium' | 'high';
/** Checklist items for review tasks (tddPhase: 'review') */
reviewChecklist?: string[];
}
/** Checkpoint review data */
export interface CheckpointReview {
iteration: number;
timestamp: number;
summary: {
total: number;
completed: number;
failed: number;
blocked: number;
pending: number;
inProgress: number;
};
stuckTasks: Array<{
id: string;
content: string;
attempts: number;
lastError?: string;
}>;
recommendations: string[];
}
// ========== Configuration Constants ========== // ========== Configuration Constants ==========
/** /**
@@ -444,6 +510,28 @@ export class RalphTracker extends EventEmitter {
/** Path to the @fix_plan.md file being watched */ /** Path to the @fix_plan.md file being watched */
private _fixPlanPath: string | null = null; private _fixPlanPath: string | null = null;
// ========== Enhanced Plan Management ==========
/** Current version of the plan (incremented on changes) */
private _planVersion: number = 1;
/** History of plan versions for rollback support */
private _planHistory: Array<{
version: number;
timestamp: number;
tasks: Map<string, EnhancedPlanTask>;
summary: string;
}> = [];
/** Enhanced plan tasks with execution tracking */
private _planTasks: Map<string, EnhancedPlanTask> = new Map();
/** Checkpoint intervals (iterations at which to trigger review) */
private _checkpointIterations: number[] = [5, 10, 20, 30, 50, 75, 100];
/** Last checkpoint iteration */
private _lastCheckpointIteration: number = 0;
/** /**
* Creates a new RalphTracker instance. * Creates a new RalphTracker instance.
* Starts in disabled state until Ralph patterns are detected. * Starts in disabled state until Ralph patterns are detected.
@@ -778,7 +866,11 @@ export class RalphTracker extends EventEmitter {
* @returns Shallow copy of loop state (safe to modify) * @returns Shallow copy of loop state (safe to modify)
*/ */
get loopState(): RalphTrackerState { get loopState(): RalphTrackerState {
return { ...this._loopState }; return {
...this._loopState,
planVersion: this._planVersion,
planHistoryLength: this._planHistory.length,
};
} }
/** /**
@@ -2338,4 +2430,327 @@ export class RalphTracker extends EventEmitter {
return newTodos.length; return newTodos.length;
} }
// ========== Enhanced Plan Management Methods ==========
/**
* Initialize plan tasks from generated plan items.
* Called when wizard generates a new plan.
*/
initializePlanTasks(items: Array<{
id?: string;
content: string;
priority?: 'P0' | 'P1' | 'P2' | null;
verificationCriteria?: string;
testCommand?: string;
dependencies?: string[];
tddPhase?: TddPhase;
pairedWith?: string;
complexity?: 'low' | 'medium' | 'high';
}>): void {
// Save current plan to history before replacing
if (this._planTasks.size > 0) {
this._savePlanToHistory('Plan replaced with new generation');
}
// Clear and rebuild
this._planTasks.clear();
this._planVersion++;
items.forEach((item, idx) => {
const id = item.id || `task-${idx}`;
const task: EnhancedPlanTask = {
id,
content: item.content,
priority: item.priority || null,
verificationCriteria: item.verificationCriteria,
testCommand: item.testCommand,
dependencies: item.dependencies || [],
status: 'pending',
attempts: 0,
version: this._planVersion,
tddPhase: item.tddPhase,
pairedWith: item.pairedWith,
complexity: item.complexity,
};
this._planTasks.set(id, task);
});
this.emit('planInitialized', { version: this._planVersion, taskCount: this._planTasks.size });
}
/**
* Update a specific plan task's status, attempts, or error.
*/
updatePlanTask(taskId: string, update: {
status?: PlanTaskStatus;
error?: string;
incrementAttempts?: boolean;
}): { success: boolean; task?: EnhancedPlanTask; error?: string } {
const task = this._planTasks.get(taskId);
if (!task) {
return { success: false, error: 'Task not found' };
}
if (update.status) {
task.status = update.status;
if (update.status === 'completed') {
task.completedAt = Date.now();
}
}
if (update.error) {
task.lastError = update.error;
}
if (update.incrementAttempts) {
task.attempts++;
// After 3 failed attempts, mark as blocked and emit warning
if (task.attempts >= 3 && task.status === 'failed') {
task.status = 'blocked';
this.emit('taskBlocked', {
taskId,
content: task.content,
attempts: task.attempts,
lastError: task.lastError,
});
}
}
// Update blocked tasks when a dependency completes
if (update.status === 'completed') {
this._unblockDependentTasks(taskId);
}
// Check for checkpoint
this._checkForCheckpoint();
this.emit('planTaskUpdate', { taskId, task });
return { success: true, task };
}
/**
* Unblock tasks that were waiting on a completed dependency.
*/
private _unblockDependentTasks(completedTaskId: string): void {
for (const [_, task] of this._planTasks) {
if (task.dependencies.includes(completedTaskId)) {
// Check if all dependencies are now complete
const allDepsComplete = task.dependencies.every(depId => {
const dep = this._planTasks.get(depId);
return dep && dep.status === 'completed';
});
if (allDepsComplete && task.status === 'blocked') {
task.status = 'pending';
this.emit('taskUnblocked', { taskId: task.id });
}
}
}
}
/**
* Check if current iteration is a checkpoint and emit review if so.
*/
private _checkForCheckpoint(): void {
const currentIteration = this._loopState.cycleCount;
if (this._checkpointIterations.includes(currentIteration) &&
currentIteration > this._lastCheckpointIteration) {
this._lastCheckpointIteration = currentIteration;
const checkpoint = this.generateCheckpointReview();
this.emit('planCheckpoint', checkpoint);
}
}
/**
* Generate a checkpoint review summarizing plan progress and stuck tasks.
*/
generateCheckpointReview(): CheckpointReview {
const tasks = Array.from(this._planTasks.values());
const summary = {
total: tasks.length,
completed: tasks.filter(t => t.status === 'completed').length,
failed: tasks.filter(t => t.status === 'failed').length,
blocked: tasks.filter(t => t.status === 'blocked').length,
pending: tasks.filter(t => t.status === 'pending').length,
inProgress: tasks.filter(t => t.status === 'in_progress').length,
};
// Find stuck tasks (3+ attempts or blocked)
const stuckTasks = tasks
.filter(t => t.attempts >= 3 || t.status === 'blocked')
.map(t => ({
id: t.id,
content: t.content,
attempts: t.attempts,
lastError: t.lastError,
}));
// Generate recommendations
const recommendations: string[] = [];
if (stuckTasks.length > 0) {
recommendations.push(`${stuckTasks.length} task(s) are stuck. Consider breaking them into smaller steps.`);
}
if (summary.failed > summary.completed && summary.total > 5) {
recommendations.push('More tasks have failed than completed. Review approach and consider plan adjustment.');
}
const progressPercent = Math.round((summary.completed / summary.total) * 100);
if (progressPercent < 20 && this._loopState.cycleCount > 10) {
recommendations.push('Progress is slow. Consider simplifying tasks or reviewing dependencies.');
}
if (summary.blocked > summary.total / 3) {
recommendations.push('Many tasks are blocked. Review dependency chain for bottlenecks.');
}
return {
iteration: this._loopState.cycleCount,
timestamp: Date.now(),
summary,
stuckTasks,
recommendations,
};
}
/**
* Save current plan state to history.
*/
private _savePlanToHistory(summary: string): void {
// Clone current tasks
const tasksCopy = new Map<string, EnhancedPlanTask>();
for (const [id, task] of this._planTasks) {
tasksCopy.set(id, { ...task });
}
this._planHistory.push({
version: this._planVersion,
timestamp: Date.now(),
tasks: tasksCopy,
summary,
});
// Limit history size
if (this._planHistory.length > 10) {
this._planHistory.shift();
}
}
/**
* Get plan version history.
*/
getPlanHistory(): Array<{
version: number;
timestamp: number;
summary: string;
stats: { total: number; completed: number; failed: number };
}> {
return this._planHistory.map(h => {
const tasks = Array.from(h.tasks.values());
return {
version: h.version,
timestamp: h.timestamp,
summary: h.summary,
stats: {
total: tasks.length,
completed: tasks.filter(t => t.status === 'completed').length,
failed: tasks.filter(t => t.status === 'failed').length,
},
};
});
}
/**
* Rollback to a previous plan version.
*/
rollbackToVersion(version: number): { success: boolean; plan?: EnhancedPlanTask[]; error?: string } {
const historyEntry = this._planHistory.find(h => h.version === version);
if (!historyEntry) {
return { success: false, error: `Version ${version} not found in history` };
}
// Save current state first
this._savePlanToHistory(`Rolled back from v${this._planVersion} to v${version}`);
// Restore the historical version
this._planTasks.clear();
for (const [id, task] of historyEntry.tasks) {
// Reset execution state for retry
this._planTasks.set(id, {
...task,
status: task.status === 'completed' ? 'completed' : 'pending',
attempts: task.status === 'completed' ? task.attempts : 0,
lastError: undefined,
});
}
this._planVersion++;
this.emit('planRollback', { version, newVersion: this._planVersion });
return { success: true, plan: Array.from(this._planTasks.values()) };
}
/**
* Add a new task to the plan (for runtime adaptation).
*/
addPlanTask(task: {
content: string;
priority?: 'P0' | 'P1' | 'P2';
verificationCriteria?: string;
dependencies?: string[];
insertAfter?: string;
}): { task: EnhancedPlanTask } {
// Generate unique ID
const existingIds = Array.from(this._planTasks.keys());
const prefix = task.priority || 'P1';
let counter = existingIds.filter(id => id.startsWith(prefix)).length + 1;
let id = `${prefix}-${String(counter).padStart(3, '0')}`;
while (this._planTasks.has(id)) {
counter++;
id = `${prefix}-${String(counter).padStart(3, '0')}`;
}
const newTask: EnhancedPlanTask = {
id,
content: task.content,
priority: task.priority || null,
verificationCriteria: task.verificationCriteria || 'Task completed successfully',
dependencies: task.dependencies || [],
status: 'pending',
attempts: 0,
version: this._planVersion,
};
this._planTasks.set(id, newTask);
this.emit('planTaskAdded', { task: newTask });
return { task: newTask };
}
/**
* Get all plan tasks.
*/
getPlanTasks(): EnhancedPlanTask[] {
return Array.from(this._planTasks.values());
}
/**
* Get current plan version.
*/
get planVersion(): number {
return this._planVersion;
}
/**
* Check if checkpoint review is due for current iteration.
*/
isCheckpointDue(): boolean {
const currentIteration = this._loopState.cycleCount;
return this._checkpointIterations.includes(currentIteration) &&
currentIteration > this._lastCheckpointIteration;
}
} }
+4
View File
@@ -738,6 +738,10 @@ export interface RalphTrackerState {
lastActivity: number; lastActivity: number;
/** Elapsed hours if detected */ /** Elapsed hours if detected */
elapsedHours: number | null; elapsedHours: number | null;
/** Current plan version (for versioning UI) */
planVersion?: number;
/** Number of versions in history (for versioning UI) */
planHistoryLength?: number;
} }
/** /**
+349 -26
View File
@@ -464,6 +464,11 @@ class ClaudemanApp {
this._subagentHideTimeout = null; // Timeout for hover-based dropdown hide this._subagentHideTimeout = null; // Timeout for hover-based dropdown hide
this.ralphStatePanelCollapsed = true; // Default to collapsed this.ralphStatePanelCollapsed = true; // Default to collapsed
// Plan subagent windows (visible Opus agents during plan generation)
this.planSubagents = new Map(); // Map<agentId, { type, model, status, startTime, element }>
this.planSubagentWindowZIndex = 1100;
this.planGenerationStopped = false; // Flag to ignore SSE events after Stop
// Project Insights tracking (active Bash tools with clickable file paths) // Project Insights tracking (active Bash tools with clickable file paths)
this.projectInsights = new Map(); // Map<sessionId, ActiveBashTool[]> this.projectInsights = new Map(); // Map<sessionId, ActiveBashTool[]>
this.logViewerWindows = new Map(); // Map<windowId, { element, eventSource, filePath }> this.logViewerWindows = new Map(); // Map<windowId, { element, eventSource, filePath }>
@@ -1907,6 +1912,13 @@ class ClaudemanApp {
}, 5 * 60 * 1000); // 5 minutes }, 5 * 60 * 1000); // 5 minutes
}); });
// Plan subagent visibility events (show Opus agents during plan generation)
addListener('plan:subagent', (e) => {
const data = JSON.parse(e.data);
console.log('[Plan Subagent]', data);
this.handlePlanSubagentEvent(data);
});
// Plan generation progress events (for detailed mode with subagents) // Plan generation progress events (for detailed mode with subagents)
addListener('plan:progress', (e) => { addListener('plan:progress', (e) => {
const data = JSON.parse(e.data); const data = JSON.parse(e.data);
@@ -2578,7 +2590,7 @@ class ClaudemanApp {
generatedPlan: null, // [{content, priority, enabled, id}] or null generatedPlan: null, // [{content, priority, enabled, id}] or null
planGenerated: false, planGenerated: false,
skipPlanGeneration: false, skipPlanGeneration: false,
planDetailLevel: 'standard', // 'brief', 'standard', 'detailed' planDetailLevel: 'detailed', // 'brief', 'standard', 'detailed'
}; };
planLoadingTimer = null; planLoadingTimer = null;
planLoadingStartTime = null; planLoadingStartTime = null;
@@ -2596,7 +2608,7 @@ class ClaudemanApp {
generatedPlan: null, generatedPlan: null,
planGenerated: false, planGenerated: false,
skipPlanGeneration: false, skipPlanGeneration: false,
planDetailLevel: 'standard', planDetailLevel: 'detailed',
// Existing plan detection // Existing plan detection
existingPlan: null, // { todos, stats, content } from @fix_plan.md existingPlan: null, // { todos, stats, content } from @fix_plan.md
useExistingPlan: false, // User chose to use existing plan useExistingPlan: false, // User chose to use existing plan
@@ -2620,10 +2632,16 @@ class ClaudemanApp {
this.updateRalphWizardUI(); this.updateRalphWizardUI();
document.getElementById('ralphWizardModal').classList.add('active'); document.getElementById('ralphWizardModal').classList.add('active');
document.getElementById('ralphTaskDescription').focus(); document.getElementById('ralphTaskDescription').focus();
// Update connection lines to show wizard connections
this.updateConnectionLines();
} }
closeRalphWizard() { closeRalphWizard() {
document.getElementById('ralphWizardModal')?.classList.remove('active'); document.getElementById('ralphWizardModal')?.classList.remove('active');
// Update connection lines (wizard closed, revert to normal)
this.updateConnectionLines();
} }
populateRalphCaseSelector() { populateRalphCaseSelector() {
@@ -2925,6 +2943,10 @@ class ClaudemanApp {
const spinnerEl = document.querySelector('.plan-spinner'); const spinnerEl = document.querySelector('.plan-spinner');
if (spinnerEl) spinnerEl.style.display = ''; if (spinnerEl) spinnerEl.style.display = '';
// Reset stopped indicator
const stoppedIndicator = document.getElementById('planStoppedIndicator');
if (stoppedIndicator) stoppedIndicator.style.display = 'none';
// Reset existing plan badge // Reset existing plan badge
const badge = document.getElementById('existingPlanBadge'); const badge = document.getElementById('existingPlanBadge');
if (badge) badge.style.display = 'none'; if (badge) badge.style.display = 'none';
@@ -2937,6 +2959,9 @@ class ClaudemanApp {
// Stop any existing generation first // Stop any existing generation first
this.stopPlanGeneration(); this.stopPlanGeneration();
// Reset stopped flag to allow new SSE events
this.planGenerationStopped = false;
// Create abort controller for this generation // Create abort controller for this generation
this.planGenerationAbortController = new AbortController(); this.planGenerationAbortController = new AbortController();
@@ -3164,6 +3189,10 @@ class ClaudemanApp {
const errorEl = document.getElementById('planGenerationError'); const errorEl = document.getElementById('planGenerationError');
const msgEl = document.getElementById('planErrorMsg'); const msgEl = document.getElementById('planErrorMsg');
const stoppedIndicator = document.getElementById('planStoppedIndicator');
// Hide the stopped indicator for real errors, show error message
if (stoppedIndicator) stoppedIndicator.style.display = 'none';
if (msgEl) msgEl.textContent = message; if (msgEl) msgEl.textContent = message;
errorEl?.classList.remove('hidden'); errorEl?.classList.remove('hidden');
} }
@@ -3176,6 +3205,9 @@ class ClaudemanApp {
document.getElementById('planGenerationError')?.classList.add('hidden'); document.getElementById('planGenerationError')?.classList.add('hidden');
document.getElementById('planEditor')?.classList.remove('hidden'); document.getElementById('planEditor')?.classList.remove('hidden');
// Close plan subagent windows since generation is complete
this.closePlanSubagentWindows();
const list = document.getElementById('planItemsList'); const list = document.getElementById('planItemsList');
if (!list) return; if (!list) return;
@@ -3565,13 +3597,183 @@ class ClaudemanApp {
cancelPlanGeneration() { cancelPlanGeneration() {
this.stopPlanGeneration(); this.stopPlanGeneration();
this.showToast('Plan generation cancelled', 'info'); this.planGenerationStopped = true; // Ignore future SSE events
this.showToast('Plan generation stopped', 'info');
// Allow user to proceed without a plan by clicking Next
this.ralphWizardConfig.skipPlanGeneration = true;
// Show stopped state with clear visual indicator
const errorEl = document.getElementById('planGenerationError'); const errorEl = document.getElementById('planGenerationError');
const msgEl = document.getElementById('planErrorMsg'); const msgEl = document.getElementById('planErrorMsg');
if (msgEl) msgEl.textContent = 'Plan generation was cancelled.'; const stoppedIndicator = document.getElementById('planStoppedIndicator');
errorEl?.classList.remove('hidden');
// Show the stopped indicator, hide error message since this was intentional
if (stoppedIndicator) stoppedIndicator.style.display = 'flex';
if (msgEl) msgEl.textContent = '';
// Hide spinner immediately and show stopped state
document.getElementById('planGenerationLoading')?.classList.add('hidden'); document.getElementById('planGenerationLoading')?.classList.add('hidden');
errorEl?.classList.remove('hidden');
// Close all plan subagent windows
this.closePlanSubagentWindows();
}
// ========== Plan Subagent Windows (Visible Opus Agents) ==========
handlePlanSubagentEvent(event) {
// Ignore events if user has stopped plan generation
if (this.planGenerationStopped) return;
const { type, agentId, agentType, model, status, detail, itemCount, durationMs, error } = event;
if (type === 'started') {
// Create new plan subagent window
this.createPlanSubagentWindow(agentId, agentType, model, detail);
} else if (type === 'completed' || type === 'failed') {
// Update window to show completion/failure
this.updatePlanSubagentWindow(agentId, status, itemCount, durationMs, error);
} else if (type === 'progress') {
// Update progress detail
const windowData = this.planSubagents.get(agentId);
if (windowData?.element) {
const detailEl = windowData.element.querySelector('.plan-subagent-detail');
if (detailEl) detailEl.textContent = detail || '';
}
}
}
createPlanSubagentWindow(agentId, agentType, model, detail) {
if (this.planSubagents.has(agentId)) return;
const wizardModal = document.getElementById('ralphWizardModal');
const wizardContent = wizardModal?.querySelector('.modal-content');
const wizardRect = wizardContent?.getBoundingClientRect();
// Calculate position - alternate left/right of wizard
const windowWidth = 280;
const windowHeight = 100;
const windowCount = this.planSubagents.size;
const viewportWidth = window.innerWidth;
let x, y;
if (wizardRect) {
const centerX = wizardRect.left + wizardRect.width / 2;
const gap = 30;
// Stack windows vertically on each side
const side = windowCount % 2 === 0 ? 'right' : 'left';
const sideIndex = Math.floor(windowCount / 2);
if (side === 'right') {
x = wizardRect.right + gap;
} else {
x = wizardRect.left - windowWidth - gap;
}
y = wizardRect.top + sideIndex * (windowHeight + 15);
// Clamp to viewport
x = Math.max(10, Math.min(x, viewportWidth - windowWidth - 10));
y = Math.max(60, Math.min(y, window.innerHeight - windowHeight - 10));
} else {
// Fallback positioning
x = 50 + windowCount * 30;
y = 120 + windowCount * 30;
}
// Create window element
const win = document.createElement('div');
win.className = 'plan-subagent-window';
win.id = `plan-subagent-${agentId}`;
win.style.left = `${x}px`;
win.style.top = `${y}px`;
win.style.zIndex = ++this.planSubagentWindowZIndex;
const typeLabels = {
requirements: 'Requirements Analyst',
architecture: 'Architecture Planner',
testing: 'TDD Specialist',
risks: 'Risk Analyst',
verification: 'Verification Expert',
};
const typeIcons = {
requirements: '📋',
architecture: '🏗️',
testing: '🧪',
risks: '⚠️',
verification: '✓',
};
win.innerHTML = `
<div class="plan-subagent-header">
<span class="plan-subagent-icon">${typeIcons[agentType] || '🤖'}</span>
<span class="plan-subagent-title">${typeLabels[agentType] || agentType}</span>
<span class="plan-subagent-model">${model}</span>
</div>
<div class="plan-subagent-body">
<div class="plan-subagent-status running">
<span class="plan-subagent-spinner"></span>
<span class="plan-subagent-status-text">Running...</span>
</div>
<div class="plan-subagent-detail">${detail || ''}</div>
</div>
`;
document.body.appendChild(win);
// Store reference
this.planSubagents.set(agentId, {
type: agentType,
model,
status: 'running',
startTime: Date.now(),
element: win,
});
// Update connection lines
this.updateConnectionLines();
}
updatePlanSubagentWindow(agentId, status, itemCount, durationMs, error) {
const windowData = this.planSubagents.get(agentId);
if (!windowData?.element) return;
const win = windowData.element;
const statusEl = win.querySelector('.plan-subagent-status');
const statusTextEl = win.querySelector('.plan-subagent-status-text');
const spinnerEl = win.querySelector('.plan-subagent-spinner');
const detailEl = win.querySelector('.plan-subagent-detail');
windowData.status = status;
if (status === 'completed') {
statusEl?.classList.remove('running');
statusEl?.classList.add('completed');
if (spinnerEl) spinnerEl.style.display = 'none';
if (statusTextEl) statusTextEl.textContent = `Done (${itemCount || 0} items)`;
if (detailEl && durationMs) detailEl.textContent = `${(durationMs / 1000).toFixed(1)}s`;
} else if (status === 'failed' || status === 'cancelled') {
statusEl?.classList.remove('running');
statusEl?.classList.add('failed');
if (spinnerEl) spinnerEl.style.display = 'none';
if (statusTextEl) statusTextEl.textContent = status === 'cancelled' ? 'Cancelled' : 'Failed';
if (detailEl) detailEl.textContent = error || '';
}
// Update connection lines
this.updateConnectionLines();
}
closePlanSubagentWindows() {
for (const [agentId, windowData] of this.planSubagents) {
if (windowData.element) {
windowData.element.remove();
}
}
this.planSubagents.clear();
this.updateConnectionLines();
} }
skipPlanGeneration() { skipPlanGeneration() {
@@ -7768,34 +7970,118 @@ class ClaudemanApp {
svg.innerHTML = ''; svg.innerHTML = '';
// Check if Ralph wizard modal is open
const wizardModal = document.getElementById('ralphWizardModal');
const wizardOpen = wizardModal && !wizardModal.classList.contains('hidden') &&
wizardModal.style.display !== 'none';
const wizardContent = wizardOpen ? wizardModal.querySelector('.modal-content') : null;
for (const [agentId, windowInfo] of this.subagentWindows) { for (const [agentId, windowInfo] of this.subagentWindows) {
if (windowInfo.minimized || windowInfo.hidden) continue; if (windowInfo.minimized || windowInfo.hidden) continue;
const agent = this.subagents.get(agentId); const agent = this.subagents.get(agentId);
if (!agent?.parentSessionId) continue;
const tab = document.querySelector(`.session-tab[data-id="${agent.parentSessionId}"]`);
const win = windowInfo.element; const win = windowInfo.element;
if (!tab || !win) continue; if (!win) continue;
const tabRect = tab.getBoundingClientRect();
const winRect = win.getBoundingClientRect(); const winRect = win.getBoundingClientRect();
// Draw curved line from tab bottom-center to window top-center // If wizard is open, draw lines from wizard to subagent windows
const x1 = tabRect.left + tabRect.width / 2; if (wizardOpen && wizardContent) {
const y1 = tabRect.bottom; const wizardRect = wizardContent.getBoundingClientRect();
const x2 = winRect.left + winRect.width / 2;
const y2 = winRect.top;
// Bezier curve control points for smooth curve // Determine which side of wizard the window is on
const midY = (y1 + y2) / 2; const winCenterX = winRect.left + winRect.width / 2;
const path = `M ${x1} ${y1} C ${x1} ${midY}, ${x2} ${midY}, ${x2} ${y2}`; const wizardCenterX = wizardRect.left + wizardRect.width / 2;
const line = document.createElementNS('http://www.w3.org/2000/svg', 'path'); let x1, y1, x2, y2;
line.setAttribute('d', path);
line.setAttribute('class', 'connection-line'); if (winCenterX < wizardCenterX) {
line.setAttribute('data-agent-id', agentId); // Window is on the left - connect from wizard left edge
svg.appendChild(line); x1 = wizardRect.left;
y1 = wizardRect.top + wizardRect.height / 2;
x2 = winRect.right;
y2 = winRect.top + winRect.height / 2;
} else {
// Window is on the right - connect from wizard right edge
x1 = wizardRect.right;
y1 = wizardRect.top + wizardRect.height / 2;
x2 = winRect.left;
y2 = winRect.top + winRect.height / 2;
}
// Bezier curve for smooth horizontal connection
const midX = (x1 + x2) / 2;
const path = `M ${x1} ${y1} C ${midX} ${y1}, ${midX} ${y2}, ${x2} ${y2}`;
const line = document.createElementNS('http://www.w3.org/2000/svg', 'path');
line.setAttribute('d', path);
line.setAttribute('class', 'connection-line wizard-connection');
line.setAttribute('data-agent-id', agentId);
svg.appendChild(line);
} else if (agent?.parentSessionId) {
// Normal mode - connect to parent session tab
const tab = document.querySelector(`.session-tab[data-id="${agent.parentSessionId}"]`);
if (!tab) continue;
const tabRect = tab.getBoundingClientRect();
// Draw curved line from tab bottom-center to window top-center
const x1 = tabRect.left + tabRect.width / 2;
const y1 = tabRect.bottom;
const x2 = winRect.left + winRect.width / 2;
const y2 = winRect.top;
// Bezier curve control points for smooth curve
const midY = (y1 + y2) / 2;
const path = `M ${x1} ${y1} C ${x1} ${midY}, ${x2} ${midY}, ${x2} ${y2}`;
const line = document.createElementNS('http://www.w3.org/2000/svg', 'path');
line.setAttribute('d', path);
line.setAttribute('class', 'connection-line');
line.setAttribute('data-agent-id', agentId);
svg.appendChild(line);
}
}
// Draw lines from wizard to plan subagent windows (Opus agents during plan generation)
if (wizardOpen && wizardContent && this.planSubagents.size > 0) {
const wizardRect = wizardContent.getBoundingClientRect();
for (const [agentId, windowData] of this.planSubagents) {
const win = windowData.element;
if (!win) continue;
const winRect = win.getBoundingClientRect();
// Determine which side of wizard the window is on
const winCenterX = winRect.left + winRect.width / 2;
const wizardCenterX = wizardRect.left + wizardRect.width / 2;
let x1, y1, x2, y2;
if (winCenterX < wizardCenterX) {
// Window is on the left
x1 = wizardRect.left;
y1 = wizardRect.top + wizardRect.height / 3 + (this.planSubagents.size > 3 ? 0 : 50);
x2 = winRect.right;
y2 = winRect.top + winRect.height / 2;
} else {
// Window is on the right
x1 = wizardRect.right;
y1 = wizardRect.top + wizardRect.height / 3 + (this.planSubagents.size > 3 ? 0 : 50);
x2 = winRect.left;
y2 = winRect.top + winRect.height / 2;
}
const midX = (x1 + x2) / 2;
const path = `M ${x1} ${y1} C ${midX} ${y1}, ${midX} ${y2}, ${x2} ${y2}`;
const line = document.createElementNS('http://www.w3.org/2000/svg', 'path');
line.setAttribute('d', path);
line.setAttribute('class', 'connection-line wizard-connection plan-subagent-line');
line.setAttribute('data-plan-agent-id', agentId);
svg.appendChild(line);
}
} }
} }
@@ -7874,11 +8160,48 @@ class ClaudemanApp {
const windowWidth = 420; const windowWidth = 420;
const windowHeight = 350; const windowHeight = 350;
const gap = 20; const gap = 20;
const startX = 50;
const startY = 120;
const viewportWidth = window.innerWidth; const viewportWidth = window.innerWidth;
const viewportHeight = window.innerHeight; const viewportHeight = window.innerHeight;
const maxCols = Math.floor((viewportWidth - startX - 50) / (windowWidth + gap)) || 1;
// Check if Ralph wizard modal is open - if so, position windows on the sides
const wizardModal = document.getElementById('ralphWizardModal');
const wizardOpen = wizardModal && !wizardModal.classList.contains('hidden') &&
wizardModal.style.display !== 'none';
let startX, startY, maxCols;
if (wizardOpen) {
// Wizard is ~720px wide, centered. Position windows on left/right sides
const wizardWidth = 720;
const centerX = viewportWidth / 2;
const wizardLeft = centerX - wizardWidth / 2;
const wizardRight = centerX + wizardWidth / 2;
// Alternate between left and right sides of the wizard
const leftSideSpace = wizardLeft - 20;
const rightSideSpace = viewportWidth - wizardRight - 20;
if (windowCount % 2 === 0 && rightSideSpace >= windowWidth) {
// Even windows go to the right
startX = wizardRight + 20;
maxCols = Math.floor(rightSideSpace / (windowWidth + gap)) || 1;
} else if (leftSideSpace >= windowWidth) {
// Odd windows go to the left
startX = Math.max(10, wizardLeft - windowWidth - 20);
maxCols = 1; // Usually only room for 1 column on left
} else {
// Not enough side space, use right side
startX = wizardRight + 20;
maxCols = 1;
}
startY = 80; // Start higher when wizard is open
} else {
// Normal positioning
startX = 50;
startY = 120;
maxCols = Math.floor((viewportWidth - startX - 50) / (windowWidth + gap)) || 1;
}
const maxRows = Math.floor((viewportHeight - startY - 50) / (windowHeight + gap)) || 1; const maxRows = Math.floor((viewportHeight - startY - 50) / (windowHeight + gap)) || 1;
const col = windowCount % maxCols; const col = windowCount % maxCols;
const row = Math.floor(windowCount / maxCols) % maxRows; // Wrap rows to stay in viewport const row = Math.floor(windowCount / maxCols) % maxRows; // Wrap rows to stay in viewport
+24 -6
View File
@@ -172,6 +172,13 @@
<span class="ralph-info-label">Elapsed</span> <span class="ralph-info-label">Elapsed</span>
<span id="ralphElapsed">0m</span> <span id="ralphElapsed">0m</span>
</div> </div>
<div class="ralph-info-row" id="ralphVersionRow" style="display: none;">
<span class="ralph-info-label">Plan</span>
<div class="ralph-version-info">
<span class="version-badge" id="ralphPlanVersion">v1</span>
<button class="rollback-btn" id="ralphRollbackBtn" onclick="app.showPlanHistory()" title="View history">history</button>
</div>
</div>
</div> </div>
<div class="ralph-tasks-grid" id="ralphTasksGrid"></div> <div class="ralph-tasks-grid" id="ralphTasksGrid"></div>
</div> </div>
@@ -1013,16 +1020,23 @@
</div> </div>
</div> </div>
<p class="plan-loading-hint" id="planLoadingHint">Initializing deep reasoning model</p> <p class="plan-loading-hint" id="planLoadingHint">Initializing deep reasoning model</p>
<button class="btn-toolbar btn-sm plan-skip-btn" onclick="app.skipPlanGeneration()">Skip</button> <div class="plan-loading-actions">
<button class="btn-toolbar btn-sm plan-cancel-btn" onclick="app.cancelPlanGeneration()">Stop</button>
<button class="btn-toolbar btn-sm" onclick="app.regeneratePlan()">Regenerate</button>
</div>
</div> </div>
<!-- Error state --> <!-- Stopped/Error state -->
<div id="planGenerationError" class="plan-section hidden"> <div id="planGenerationError" class="plan-section hidden">
<div class="plan-stopped-indicator" id="planStoppedIndicator">
<span class="plan-stopped-icon">&#x25A0;</span>
<span class="plan-stopped-text">Stopped</span>
</div>
<p class="plan-error-msg" id="planErrorMsg"></p> <p class="plan-error-msg" id="planErrorMsg"></p>
<div class="plan-actions"> <div class="plan-actions">
<button class="btn-toolbar btn-primary" onclick="app.generatePlan()">Retry</button> <button class="btn-toolbar btn-primary" onclick="app.generatePlan()">Regenerate Plan</button>
<button class="btn-toolbar" onclick="app.skipPlanGeneration()">Skip</button>
</div> </div>
<p class="plan-skip-hint">Use <strong>Back</strong> to edit your task, or <strong>Next</strong> to continue without a plan</p>
</div> </div>
<!-- Editor state --> <!-- Editor state -->
@@ -1034,8 +1048,8 @@
<label>Detail:</label> <label>Detail:</label>
<div class="plan-detail-btns"> <div class="plan-detail-btns">
<button type="button" class="plan-detail-btn" data-detail="brief" onclick="app.setPlanDetail('brief')" title="High-level milestones only">Brief</button> <button type="button" class="plan-detail-btn" data-detail="brief" onclick="app.setPlanDetail('brief')" title="High-level milestones only">Brief</button>
<button type="button" class="plan-detail-btn active" data-detail="standard" onclick="app.setPlanDetail('standard')" title="Single-pass generation with Opus 4.5">Standard</button> <button type="button" class="plan-detail-btn" data-detail="standard" onclick="app.setPlanDetail('standard')" title="Single-pass generation with Opus 4.5">Standard</button>
<button type="button" class="plan-detail-btn" data-detail="detailed" onclick="app.setPlanDetail('detailed')" title="Enhanced: 4 parallel subagents + verification (slower but more thorough)">Enhanced</button> <button type="button" class="plan-detail-btn active" data-detail="detailed" onclick="app.setPlanDetail('detailed')" title="Enhanced: 4 parallel subagents + verification (slower but more thorough)">Enhanced</button>
</div> </div>
</div> </div>
<button class="btn-toolbar btn-sm" onclick="app.regeneratePlan()">Regenerate</button> <button class="btn-toolbar btn-sm" onclick="app.regeneratePlan()">Regenerate</button>
@@ -1045,6 +1059,10 @@
<div id="planItemsList" class="plan-items-list"></div> <div id="planItemsList" class="plan-items-list"></div>
<div class="plan-editor-footer"> <div class="plan-editor-footer">
<button class="btn-toolbar btn-sm" onclick="app.addPlanItem()">+ Add Step</button> <button class="btn-toolbar btn-sm" onclick="app.addPlanItem()">+ Add Step</button>
<div class="plan-view-toggle">
<button type="button" class="plan-view-btn active" data-view="list" onclick="app.setPlanViewMode('list')" title="Flat list">&#x2630;</button>
<button type="button" class="plan-view-btn" data-view="grouped" onclick="app.setPlanViewMode('grouped')" title="Group by TDD phase">&#x229E;</button>
</div>
</div> </div>
</div> </div>
</div> </div>
+738 -7
View File
@@ -4692,6 +4692,19 @@ kbd {
stroke-dashoffset: 1000; stroke-dashoffset: 1000;
} }
/* Wizard connection lines - different color to show plan generation link */
.connection-line.wizard-connection {
stroke: #60a5fa;
filter: drop-shadow(0 0 4px #60a5fa);
stroke-dasharray: 8 4;
animation: wizard-pulse 1.5s ease-in-out infinite;
}
@keyframes wizard-pulse {
0%, 100% { opacity: 0.5; }
50% { opacity: 0.9; }
}
/* ========== Project Insights Panel (Bash File Viewers) ========== */ /* ========== Project Insights Panel (Bash File Viewers) ========== */
.project-insights-panel { .project-insights-panel {
@@ -5366,10 +5379,23 @@ kbd {
gap: 0.5rem; gap: 0.5rem;
} }
/* Wizard Modal - wider layout */ /* Wizard Modal - wider layout with height constraints */
.modal-wizard { .modal-wizard {
width: 720px; width: 720px;
max-width: 95vw; max-width: 95vw;
max-height: 80vh;
display: flex;
flex-direction: column;
}
.modal-wizard .modal-body {
overflow-y: auto;
flex: 1;
min-height: 0;
}
.modal-wizard .wizard-footer {
flex-shrink: 0;
} }
.wizard-page-compact { .wizard-page-compact {
@@ -5641,20 +5667,158 @@ kbd {
display: flex; display: flex;
flex-direction: column; flex-direction: column;
gap: 0.35rem; gap: 0.35rem;
max-height: 300px; max-height: 200px;
overflow-y: auto; overflow-y: auto;
margin-bottom: 0.5rem; margin-bottom: 0.5rem;
} }
.plan-item { .plan-loading-actions {
display: flex;
gap: 0.5rem;
justify-content: center;
}
.plan-cancel-btn {
color: var(--red);
}
/* Plan Stopped State */
.plan-stopped-indicator {
display: flex; display: flex;
align-items: center; align-items: center;
gap: 0.4rem; justify-content: center;
padding: 0.35rem 0.5rem; gap: 0.5rem;
background: var(--bg-input); margin-bottom: 1rem;
border-radius: 4px; padding: 0.75rem;
background: var(--surface);
border-radius: 6px;
border: 1px solid var(--border);
} }
.plan-stopped-icon {
color: var(--red);
font-size: 1rem;
}
.plan-stopped-text {
font-weight: 500;
color: var(--text);
}
.plan-actions-single {
justify-content: center;
}
.plan-skip-hint {
font-size: 0.75rem;
color: var(--text-muted);
text-align: center;
margin-top: 0.75rem;
}
.plan-skip-hint strong {
color: var(--text-dim);
}
/* ========== Plan Subagent Windows (Visible Opus Agents) ========== */
.plan-subagent-window {
position: fixed;
width: 280px;
background: var(--bg-card);
border: 1px solid var(--accent);
border-radius: 8px;
box-shadow: 0 4px 20px rgba(0, 0, 0, 0.4), 0 0 20px rgba(96, 165, 250, 0.2);
overflow: hidden;
animation: planSubagentSpawn 0.3s ease-out;
}
@keyframes planSubagentSpawn {
from {
opacity: 0;
transform: scale(0.8);
}
to {
opacity: 1;
transform: scale(1);
}
}
.plan-subagent-header {
display: flex;
align-items: center;
gap: 0.5rem;
padding: 0.5rem 0.75rem;
background: linear-gradient(135deg, rgba(96, 165, 250, 0.2), rgba(96, 165, 250, 0.1));
border-bottom: 1px solid var(--border);
}
.plan-subagent-icon {
font-size: 1rem;
}
.plan-subagent-title {
flex: 1;
font-weight: 500;
font-size: 0.8rem;
color: var(--text);
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
}
.plan-subagent-model {
font-size: 0.65rem;
padding: 0.15rem 0.4rem;
background: #7c3aed;
color: white;
border-radius: 3px;
text-transform: uppercase;
font-weight: 600;
}
.plan-subagent-body {
padding: 0.75rem;
}
.plan-subagent-status {
display: flex;
align-items: center;
gap: 0.5rem;
margin-bottom: 0.35rem;
}
.plan-subagent-status.running .plan-subagent-spinner {
width: 14px;
height: 14px;
border: 2px solid var(--border);
border-top-color: var(--accent);
border-radius: 50%;
animation: spin 1s linear infinite;
}
.plan-subagent-status.completed .plan-subagent-status-text {
color: var(--green);
}
.plan-subagent-status.failed .plan-subagent-status-text {
color: var(--red);
}
.plan-subagent-status-text {
font-size: 0.8rem;
font-weight: 500;
}
.plan-subagent-detail {
font-size: 0.7rem;
color: var(--text-muted);
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
}
/* Note: main .plan-item styles are in "Enhanced Plan Item Styling" section below */
.plan-item-checkbox { .plan-item-checkbox {
flex-shrink: 0; flex-shrink: 0;
width: 14px; width: 14px;
@@ -5701,6 +5865,8 @@ kbd {
color: var(--red); color: var(--red);
} }
/* Note: .plan-item-phase and .plan-item-verify enhanced styles are in section below */
.plan-item:has(.plan-item-checkbox:not(:checked)) { .plan-item:has(.plan-item-checkbox:not(:checked)) {
opacity: 0.5; opacity: 0.5;
} }
@@ -5761,6 +5927,571 @@ kbd {
color: #fbbf24; color: #fbbf24;
} }
/* Enhanced Plan Item Styling */
.plan-item {
display: flex;
align-items: center;
gap: 0.4rem;
padding: 0.35rem 0.5rem;
background: var(--bg-input);
border-radius: 4px;
border-left: 3px solid transparent;
transition: border-color 0.15s, background-color 0.15s, transform 0.15s;
cursor: default;
}
.plan-item.dragging {
opacity: 0.5;
transform: scale(0.98);
}
.plan-item.drag-over {
background: var(--accent-dim);
border-left-color: var(--accent);
}
/* Priority color borders */
.plan-item.priority-p0 {
border-left-color: #ef4444;
}
.plan-item.priority-p1 {
border-left-color: #f59e0b;
}
.plan-item.priority-p2 {
border-left-color: #3b82f6;
}
/* Drag handle */
.plan-item-drag {
cursor: grab;
color: var(--text-muted);
font-size: 0.9rem;
padding: 0 0.2rem;
opacity: 0.5;
user-select: none;
}
.plan-item-drag:hover {
opacity: 1;
}
.plan-item-drag:active {
cursor: grabbing;
}
/* TDD Phase badge - enhanced */
.plan-item-phase {
flex-shrink: 0;
min-width: 22px;
height: 18px;
display: flex;
align-items: center;
justify-content: center;
font-size: 0.7rem;
border-radius: 3px;
background: var(--bg-card);
}
.plan-item-phase.phase-setup {
background: #6366f1;
color: white;
}
.plan-item-phase.phase-test {
background: #8b5cf6;
color: white;
}
.plan-item-phase.phase-impl {
background: #06b6d4;
color: white;
}
.plan-item-phase.phase-review {
background: #f59e0b;
color: white;
}
.plan-item-phase.phase-verify {
background: #22c55e;
color: white;
}
/* Enhanced verification indicator */
.plan-item-verify {
flex-shrink: 0;
display: flex;
align-items: center;
gap: 0.15rem;
color: var(--green);
font-size: 0.65rem;
cursor: help;
padding: 0.1rem 0.25rem;
background: rgba(34, 197, 94, 0.1);
border-radius: 3px;
opacity: 0.85;
}
.plan-item-verify:hover {
opacity: 1;
background: rgba(34, 197, 94, 0.2);
}
/* Dependencies indicator */
.plan-item-deps {
flex-shrink: 0;
font-size: 0.65rem;
color: var(--text-muted);
padding: 0.1rem 0.25rem;
background: var(--bg-card);
border-radius: 3px;
cursor: help;
}
.plan-item-deps:hover {
color: var(--accent);
}
/* Review checklist indicator */
.plan-item-checklist {
flex-shrink: 0;
display: flex;
align-items: center;
gap: 0.15rem;
color: #f59e0b;
font-size: 0.65rem;
cursor: help;
padding: 0.1rem 0.25rem;
background: rgba(245, 158, 11, 0.1);
border-radius: 3px;
opacity: 0.85;
}
.plan-item-checklist:hover {
opacity: 1;
background: rgba(245, 158, 11, 0.2);
}
/* Checklist modal */
.checklist-modal-container {
position: fixed;
top: 0;
left: 0;
right: 0;
bottom: 0;
background: rgba(0, 0, 0, 0.7);
display: flex;
align-items: center;
justify-content: center;
z-index: 9999;
}
.checklist-modal {
padding: 1rem;
}
.checklist-modal h3 {
margin: 0 0 0.5rem 0;
color: var(--text);
}
.checklist-task {
color: var(--text-muted);
font-size: 0.85rem;
margin-bottom: 1rem;
padding-bottom: 0.5rem;
border-bottom: 1px solid var(--border);
}
.checklist-list {
list-style: none;
padding: 0;
margin: 0 0 1rem 0;
max-height: 300px;
overflow-y: auto;
}
.checklist-item {
display: flex;
align-items: flex-start;
gap: 0.5rem;
padding: 0.35rem 0;
}
.checklist-checkbox {
flex-shrink: 0;
margin-top: 0.15rem;
}
.checklist-item label {
color: var(--text);
font-size: 0.85rem;
cursor: pointer;
}
.checklist-footer {
text-align: right;
}
/* Group headers for TDD phases */
.plan-group-header {
display: flex;
align-items: center;
gap: 0.5rem;
padding: 0.3rem 0.5rem;
font-size: 0.7rem;
font-weight: 600;
text-transform: uppercase;
color: var(--text-muted);
border-bottom: 1px solid var(--border);
margin-top: 0.5rem;
cursor: pointer;
user-select: none;
}
.plan-group-header:first-child {
margin-top: 0;
}
.plan-group-header:hover {
color: var(--text);
}
.plan-group-header .group-icon {
font-size: 0.8rem;
}
.plan-group-header .group-count {
margin-left: auto;
color: var(--text-dim);
font-weight: normal;
}
.plan-group-header .group-chevron {
transition: transform 0.15s;
}
.plan-group-header.collapsed .group-chevron {
transform: rotate(-90deg);
}
.plan-group-items {
display: flex;
flex-direction: column;
gap: 0.25rem;
padding: 0.25rem 0;
}
.plan-group-items.collapsed {
display: none;
}
/* View mode toggle */
.plan-view-toggle {
display: flex;
gap: 0.25rem;
}
.plan-view-btn {
padding: 0.2rem 0.4rem;
font-size: 0.7rem;
background: var(--bg-input);
border: 1px solid var(--border);
border-radius: 3px;
color: var(--text-muted);
cursor: pointer;
transition: all 0.15s;
}
.plan-view-btn:hover {
color: var(--text);
border-color: var(--accent);
}
.plan-view-btn.active {
background: var(--accent);
color: white;
border-color: var(--accent);
}
/* Verification criteria tooltip expansion */
.plan-item-verify-expanded {
position: absolute;
left: 0;
right: 0;
bottom: calc(100% + 4px);
background: var(--bg-card);
border: 1px solid var(--border);
border-radius: 4px;
padding: 0.5rem;
font-size: 0.75rem;
color: var(--text);
z-index: 10;
box-shadow: 0 -2px 8px rgba(0,0,0,0.2);
max-width: 300px;
}
.plan-item-verify-expanded::after {
content: '';
position: absolute;
bottom: -6px;
left: 50%;
transform: translateX(-50%);
border: 5px solid transparent;
border-top-color: var(--border);
}
/* Plan toolbar improvements */
.plan-editor-toolbar {
display: flex;
align-items: center;
justify-content: space-between;
margin-bottom: 0.5rem;
padding-bottom: 0.5rem;
border-bottom: 1px solid var(--border);
}
.plan-toolbar-left {
display: flex;
gap: 0.5rem;
align-items: center;
}
.plan-toolbar-right {
display: flex;
gap: 0.5rem;
align-items: center;
}
/* Plan item inline actions */
.plan-item-actions {
display: flex;
gap: 0.1rem;
opacity: 0;
transition: opacity 0.15s;
}
.plan-item:hover .plan-item-actions {
opacity: 1;
}
.plan-item-action {
background: none;
border: none;
color: var(--text-muted);
cursor: pointer;
font-size: 0.75rem;
padding: 0.1rem 0.2rem;
border-radius: 2px;
}
.plan-item-action:hover {
color: var(--text);
background: var(--bg-card);
}
/* Runtime task card enhancements */
.ralph-task-card {
position: relative;
}
.ralph-task-attempts {
position: absolute;
top: 2px;
right: 2px;
font-size: 0.6rem;
color: var(--text-muted);
background: var(--bg-card);
padding: 0 3px;
border-radius: 2px;
}
.ralph-task-attempts.has-errors {
color: var(--red);
background: rgba(239, 68, 68, 0.1);
}
.ralph-task-verify-badge {
position: absolute;
bottom: 2px;
right: 2px;
font-size: 0.6rem;
color: var(--green);
opacity: 0.7;
}
.ralph-task-deps-indicator {
position: absolute;
top: 2px;
left: 2px;
font-size: 0.55rem;
color: var(--accent);
}
/* Quick actions on task hover */
.ralph-task-actions {
position: absolute;
top: 50%;
left: 50%;
transform: translate(-50%, -50%);
display: none;
gap: 0.25rem;
background: rgba(0,0,0,0.85);
padding: 0.25rem;
border-radius: 4px;
}
.ralph-task-card:hover .ralph-task-actions {
display: flex;
}
.ralph-task-action-btn {
padding: 0.2rem 0.35rem;
font-size: 0.65rem;
background: var(--bg-input);
border: 1px solid var(--border);
border-radius: 3px;
color: var(--text);
cursor: pointer;
}
.ralph-task-action-btn:hover {
background: var(--accent);
color: white;
border-color: var(--accent);
}
/* Plan version indicator */
.plan-version-indicator {
display: flex;
align-items: center;
gap: 0.5rem;
font-size: 0.75rem;
color: var(--text-muted);
padding: 0.25rem 0.5rem;
background: var(--bg-input);
border-radius: 4px;
}
.plan-version-badge {
background: var(--accent);
color: white;
padding: 0.1rem 0.3rem;
border-radius: 3px;
font-size: 0.65rem;
font-weight: 500;
}
.plan-version-history {
cursor: pointer;
color: var(--accent);
}
.plan-version-history:hover {
text-decoration: underline;
}
/* Plan history dropdown */
.plan-history-dropdown {
position: absolute;
top: 100%;
right: 0;
background: var(--bg-card);
border: 1px solid var(--border);
border-radius: 4px;
box-shadow: 0 4px 12px rgba(0,0,0,0.3);
z-index: 100;
min-width: 200px;
max-height: 200px;
overflow-y: auto;
}
.plan-history-list {
max-height: 300px;
overflow-y: auto;
border: 1px solid var(--border);
border-radius: 4px;
}
.plan-history-item {
display: flex;
justify-content: space-between;
align-items: center;
padding: 0.5rem 0.75rem;
border-bottom: 1px solid var(--border);
cursor: pointer;
font-size: 0.8rem;
transition: background-color 0.15s;
}
.plan-history-item:last-child {
border-bottom: none;
}
.plan-history-item:hover {
background: var(--bg-input);
}
.plan-history-item.current {
background: var(--accent-dim);
cursor: default;
}
.plan-history-item.current::after {
content: '(current)';
font-size: 0.65rem;
color: var(--accent);
margin-left: 0.5rem;
}
.plan-history-version {
font-weight: 500;
color: var(--text);
}
.plan-history-tasks {
font-size: 0.7rem;
color: var(--text-muted);
margin-left: 0.5rem;
}
.plan-history-time {
color: var(--text-muted);
font-size: 0.7rem;
}
/* Ralph panel version info */
.ralph-version-info {
display: flex;
align-items: center;
gap: 0.35rem;
font-size: 0.7rem;
color: var(--text-muted);
margin-left: auto;
}
.ralph-version-info .version-badge {
background: var(--bg-card);
padding: 0.1rem 0.25rem;
border-radius: 2px;
font-family: 'SF Mono', Monaco, monospace;
}
.ralph-version-info .rollback-btn {
background: none;
border: none;
color: var(--accent);
cursor: pointer;
font-size: 0.65rem;
padding: 0.1rem 0.25rem;
}
.ralph-version-info .rollback-btn:hover {
text-decoration: underline;
}
/* Textarea with AI Assist button */ /* Textarea with AI Assist button */
.textarea-with-assist { .textarea-with-assist {
position: relative; position: relative;
+184
View File
@@ -0,0 +1,184 @@
/**
* @fileoverview Tests for /api/generate-plan endpoint validation
*
* Tests request validation for the plan generation API.
* Port: 3191
*/
import { describe, it, expect } from 'vitest';
describe('Generate Plan API Validation', () => {
// Validation logic extracted from server.ts for unit testing
const validateGeneratePlanRequest = (req: {
taskDescription?: unknown;
detailLevel?: unknown;
}): { valid: boolean; error?: string } => {
const { taskDescription, detailLevel } = req;
if (!taskDescription || typeof taskDescription !== 'string') {
return { valid: false, error: 'Task description is required' };
}
if (taskDescription.length === 0) {
return { valid: false, error: 'Task description is required' };
}
if (taskDescription.length > 10000) {
return { valid: false, error: 'Task description too long (max 10000 chars)' };
}
// Detail level validation (optional, defaults to 'standard')
if (detailLevel !== undefined) {
const validLevels = ['brief', 'standard', 'detailed'];
if (!validLevels.includes(detailLevel as string)) {
return { valid: false, error: 'Invalid detail level. Must be: brief, standard, or detailed' };
}
}
return { valid: true };
};
describe('taskDescription validation', () => {
it('should reject missing taskDescription', () => {
const result = validateGeneratePlanRequest({});
expect(result.valid).toBe(false);
expect(result.error).toContain('required');
});
it('should reject null taskDescription', () => {
const result = validateGeneratePlanRequest({ taskDescription: null });
expect(result.valid).toBe(false);
expect(result.error).toContain('required');
});
it('should reject empty string taskDescription', () => {
const result = validateGeneratePlanRequest({ taskDescription: '' });
expect(result.valid).toBe(false);
expect(result.error).toContain('required');
});
it('should reject non-string taskDescription', () => {
const result = validateGeneratePlanRequest({ taskDescription: 123 });
expect(result.valid).toBe(false);
expect(result.error).toContain('required');
});
it('should reject taskDescription over 10000 chars', () => {
const longDescription = 'a'.repeat(10001);
const result = validateGeneratePlanRequest({ taskDescription: longDescription });
expect(result.valid).toBe(false);
expect(result.error).toContain('too long');
});
it('should accept taskDescription at exactly 10000 chars', () => {
const maxDescription = 'a'.repeat(10000);
const result = validateGeneratePlanRequest({ taskDescription: maxDescription });
expect(result.valid).toBe(true);
});
it('should accept valid taskDescription', () => {
const result = validateGeneratePlanRequest({
taskDescription: 'Implement user authentication with JWT tokens',
});
expect(result.valid).toBe(true);
});
it('should accept multi-line taskDescription', () => {
const result = validateGeneratePlanRequest({
taskDescription: `Fix the Ralph Loop wizard issues:
1. Two Skip buttons confusing
2. Modal height overflow
3. Dead code references`,
});
expect(result.valid).toBe(true);
});
});
describe('detailLevel validation', () => {
it('should accept brief detail level', () => {
const result = validateGeneratePlanRequest({
taskDescription: 'Build a feature',
detailLevel: 'brief',
});
expect(result.valid).toBe(true);
});
it('should accept standard detail level', () => {
const result = validateGeneratePlanRequest({
taskDescription: 'Build a feature',
detailLevel: 'standard',
});
expect(result.valid).toBe(true);
});
it('should accept detailed detail level', () => {
const result = validateGeneratePlanRequest({
taskDescription: 'Build a feature',
detailLevel: 'detailed',
});
expect(result.valid).toBe(true);
});
it('should accept missing detailLevel (defaults to standard)', () => {
const result = validateGeneratePlanRequest({
taskDescription: 'Build a feature',
});
expect(result.valid).toBe(true);
});
it('should reject invalid detailLevel', () => {
const result = validateGeneratePlanRequest({
taskDescription: 'Build a feature',
detailLevel: 'verbose',
});
expect(result.valid).toBe(false);
expect(result.error).toContain('Invalid detail level');
});
it('should reject numeric detailLevel', () => {
const result = validateGeneratePlanRequest({
taskDescription: 'Build a feature',
detailLevel: 3,
});
expect(result.valid).toBe(false);
expect(result.error).toContain('Invalid detail level');
});
});
describe('edge cases', () => {
it('should handle taskDescription with only whitespace', () => {
// Whitespace-only is technically a valid non-empty string
// The API accepts it (server-side trim could be added if needed)
const result = validateGeneratePlanRequest({ taskDescription: ' ' });
expect(result.valid).toBe(true);
});
it('should handle taskDescription with unicode characters', () => {
const result = validateGeneratePlanRequest({
taskDescription: 'Implement 日本語 support with emoji 🎉',
});
expect(result.valid).toBe(true);
});
it('should handle taskDescription with special characters', () => {
const result = validateGeneratePlanRequest({
taskDescription: 'Fix bug: user@example.com fails with <script>alert(1)</script>',
});
expect(result.valid).toBe(true);
});
it('should handle array as taskDescription', () => {
const result = validateGeneratePlanRequest({
taskDescription: ['item1', 'item2'],
});
expect(result.valid).toBe(false);
});
it('should handle object as taskDescription', () => {
const result = validateGeneratePlanRequest({
taskDescription: { text: 'description' },
});
expect(result.valid).toBe(false);
});
});
});