From 482394e280278a96d6253d79559e301b254ce46c Mon Sep 17 00:00:00 2001 From: arkon Date: Tue, 20 Jan 2026 17:43:36 +0100 Subject: [PATCH] docs: update completion detection documentation - Update "All complete" detection behavior in CLAUDE.md and ralph-wiggum-guide.md - Add UI Behaviors section documenting auto-focus, scroll preservation, consistent icons Co-Authored-By: Claude Opus 4.5 --- CLAUDE.md | 116 +++++++++++++++++++++++++++++-------- docs/ralph-wiggum-guide.md | 116 ++++++++++++++++++++++++++----------- 2 files changed, 176 insertions(+), 56 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 38f9eb07..c0001ec7 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -38,29 +38,35 @@ npm run test:coverage # With coverage report npx vitest run test/session.test.ts # Single file npx vitest run -t "should create session" # By pattern -# Test port allocation (add new tests in next available range): -# 3099-3101: quick-start.test.ts +# Test port allocation (integration tests spawn servers): +# 3099: quick-start.test.ts # 3102: session.test.ts -# 3105-3106: scheduled-runs.test.ts -# 3107-3108: sse-events.test.ts -# 3110-3112: edge-cases.test.ts -# 3115-3116: integration-flows.test.ts -# 3120-3121: session-cleanup.test.ts -# (no port): respawn-controller.test.ts, inner-loop-tracker.test.ts, pty-interactive.test.ts (unit tests) +# 3105: scheduled-runs.test.ts +# 3107: sse-events.test.ts +# 3110: edge-cases.test.ts +# 3115: integration-flows.test.ts +# 3120: session-cleanup.test.ts +# Unit tests (no port needed): respawn-controller, inner-loop-tracker, pty-interactive # Next available: 3122+ # Tests mock PTY - no real Claude CLI spawned # Test timeout: 30s (configured in vitest.config.ts) -# TypeScript checking (no linter configured) +# TypeScript checking npx tsc --noEmit # Type check without building +# Note: No ESLint/Prettier configured - rely on TypeScript strict mode # Debugging screen -ls # List GNU screen sessions -screen -r # Attach to screen session +screen -r # Attach to screen session (Ctrl+A D to detach) curl localhost:3000/api/sessions # Check active sessions +curl localhost:3000/api/status | jq . # Full app state including respawn cat ~/.claudeman/state.json | jq . # View main state cat ~/.claudeman/state-inner.json | jq . # View inner loop state + +# Kill stuck screen sessions +screen -X -S quit # Graceful quit +pkill -f "SCREEN.*claudeman" # Force kill all claudeman screens ``` ## Architecture @@ -74,7 +80,8 @@ cat ~/.claudeman/state-inner.json | jq . # View inner loop state | `src/screen-manager.ts` | GNU screen persistence, ghost discovery, 4-strategy kill | | `src/inner-loop-tracker.ts` | Detects `PHRASE`, todos, loop status in output | | `src/task-tracker.ts` | Parses background task output (agent IDs, status) from Claude CLI | -| `src/state-store.ts` | JSON persistence to `~/.claudeman/` with debounced (100ms) writes | +| `src/session-manager.ts` | Manages session lifecycle, task assignment, and cleanup | +| `src/state-store.ts` | JSON persistence to `~/.claudeman/` with debounced writes | | `src/web/server.ts` | Fastify REST API + SSE at `/api/events` | | `src/web/public/app.js` | Frontend: SSE handling, xterm.js, tab management | | `src/types.ts` | All TypeScript interfaces | @@ -89,12 +96,21 @@ cat ~/.claudeman/state-inner.json | jq . # View inner loop state ### Respawn State Machine ``` -WATCHING → SENDING_UPDATE → WAITING_UPDATE → SENDING_CLEAR → WAITING_CLEAR → SENDING_INIT → WAITING_INIT → WATCHING - ↑ | - └──────────────────────────────────────────────────────────────────────────────────────────────────────────┘ +┌─────────────────────────────────────────────────────────────────────────────────────────────────┐ +│ │ +▼ │ +WATCHING → SENDING_UPDATE → WAITING_UPDATE → SENDING_CLEAR → WAITING_CLEAR │ + │ │ + ▼ │ + SENDING_INIT → WAITING_INIT → MONITORING_INIT ──┬──────────────────┘ + │ + ▼ (if no work triggered) + SENDING_KICKSTART → WAITING_KICKSTART ``` -Steps can be skipped via config (`sendClear: false`, `sendInit: false`). Idle detection triggers state transitions. +**States**: `watching`, `sending_update`, `waiting_update`, `sending_clear`, `waiting_clear`, `sending_init`, `waiting_init`, `monitoring_init`, `sending_kickstart`, `waiting_kickstart`, `stopped` + +Steps can be skipped via config (`sendClear: false`, `sendInit: false`). Optional `kickstartPrompt` triggers if `/init` doesn't start work. Idle detection triggers state transitions. ### Session Modes @@ -169,7 +185,7 @@ session.writeViaScreen('/init\r'); 2. Sends text first: `screen -S name -p 0 -X stuff "text"` 3. Sends Enter separately: `screen -S name -p 0 -X stuff "$(printf '\015')"` -**Why separate commands?** Claude CLI uses Ink (React for terminals) which requires text and Enter as separate `screen -X stuff` commands. Combining them doesn't work. +**Why separate commands?** Claude CLI uses [Ink](https://github.com/vadimdemedes/ink) (React for terminals) which requires text and Enter as separate `screen -X stuff` commands. Combining them doesn't work. This is a critical implementation detail when debugging input issues. #### API Usage ```bash @@ -208,15 +224,38 @@ Both wait for idle. Configure via `session.setAutoCompact()` / `session.setAutoC ### Inner Loop Tracking Detects Ralph loops and todos inside Claude sessions. **Disabled by default** but auto-enables when any of these patterns are detected in terminal output: -- `/ralph-loop` command +- `/ralph-loop:ralph-loop` command - `PHRASE` completion phrases - `TodoWrite` tool usage - Iteration patterns (`Iteration 5/50`, `[5/50]`) - Todo checkboxes (`- [ ]`/`- [x]`) or indicator icons (`☐`/`◐`/`✓`) +- "All tasks complete" messages +- Individual task completion signals (`Task 8 is done`) See `inner-loop-tracker.ts:shouldAutoEnable()` for detection logic. -API: `GET /api/sessions/:id/inner-state`. UI: collapsible panel below tabs. Use `tracker.enable()` / `tracker.disable()` for programmatic control, or `POST /api/sessions/:id/inner-config` with `{ enabled: boolean }` via API. +**Completion Detection**: Uses multi-strategy detection: +- 1st occurrence of `PHRASE`: Stores as expected phrase (likely in prompt) +- 2nd occurrence: Emits `completionDetected` event (actual completion) +- **Bare phrase detection**: Also detects phrase without tags once expected phrase is known +- **All complete detection**: When "All X files/tasks created/completed" detected, marks all todos complete and emits completion +- If loop is already active (via `/ralph-loop:ralph-loop`): Emits immediately on first occurrence + +**Session Lifecycle**: Each session has its own independent tracker: +- New session → Fresh tracker (no carryover) +- Close tab → Tracker state cleared, UI panel hides +- Switch tabs → Panel shows tracker for active session +- `tracker.reset()` → Clears todos/state, keeps enabled status +- `tracker.fullReset()` → Complete reset to initial state + +**API**: +- `GET /api/sessions/:id/inner-state` - Get loop state and todos +- `POST /api/sessions/:id/inner-config` - Configure tracker: + - `{ enabled: boolean }` - Enable/disable + - `{ reset: true }` - Soft reset (keep enabled) + - `{ reset: "full" }` - Full reset + +UI: Collapsible panel below tabs, shows progress ring and todo list. ### Terminal Display Fix @@ -248,7 +287,17 @@ Vanilla JS + xterm.js. Key functions: ### State Store -Writes debounced (100ms) to `~/.claudeman/state.json`. Batches rapid changes. +Writes debounced to `~/.claudeman/state.json`. Batches rapid changes. + +### Timing Constants + +| Constant | Value | Location | +|----------|-------|----------| +| State save debounce | 500ms | `state-store.ts` | +| Line buffer flush | 100ms | `session.ts` | +| Terminal batch interval | 16ms | `server.ts` (60fps) | +| Idle activity timeout | 2s | `session.ts` | +| Respawn idle timeout | 5s default | `RespawnConfig.idleTimeoutMs` | ### TypeScript Config @@ -259,10 +308,24 @@ Module resolution: NodeNext. Target: ES2022. Strict mode enabled. See `tsconfig. - **API endpoint**: Add types in `types.ts`, route in `server.ts:buildServer()`, use `createErrorResponse()` for errors - **SSE event**: Emit via `broadcast()` in server.ts, handle in `app.js:handleSSEEvent()` switch - **Session event**: Add to `SessionEvents` interface in `session.ts`, emit via `this.emit()`, subscribe in server.ts, handle in frontend +- **New test file**: Create `test/.test.ts`, pick unique port (next available: 3122+), add to port allocation comment above + +### API Error Codes + +Use `createErrorResponse(code, details?)` from `types.ts`: + +| Code | Use Case | +|------|----------| +| `NOT_FOUND` | Session/resource doesn't exist | +| `INVALID_INPUT` | Bad request parameters | +| `SESSION_BUSY` | Session is currently processing | +| `OPERATION_FAILED` | Action couldn't complete | +| `ALREADY_EXISTS` | Duplicate resource | +| `INTERNAL_ERROR` | Unexpected server error | ## Session Lifecycle & Cleanup -- **Limit**: `MAX_CONCURRENT_SESSIONS = 50` +- **Limit**: Web server: `MAX_CONCURRENT_SESSIONS = 50` (`server.ts:56`), CLI default: 5 (`types.ts:DEFAULT_CONFIG`) - **Kill** (`killScreen()`): child PIDs → process group → screen quit → SIGKILL - **Ghost discovery**: `reconcileScreens()` finds orphaned screens on startup - **Cleanup** (`cleanupSession()`): stops respawn, clears buffers/timers, kills screen @@ -339,6 +402,12 @@ claudeman status # Overall status | `Ctrl+?` | Show keyboard shortcuts help | | `Escape` | Close panels and modals | +## UI Behaviors + +- **Auto-focus**: When only one session exists and none is active, it auto-selects +- **Scroll preservation**: Expanding/collapsing Ralph panel preserves terminal scroll position +- **Consistent icons**: Ralph and Monitor panels use same detach/attach icons (⧉/⊞) + ## State Files | File | Purpose | @@ -364,10 +433,11 @@ Ralph Wiggum is an autonomous loop technique that lets Claude work iteratively u **Core Pattern**: `PHRASE` - The completion signal that tells the loop to stop. -**Official Plugin Commands**: +**Skill Commands**: ```bash -/ralph-loop "" --max-iterations 50 --completion-promise "COMPLETE" -/cancel-ralph +/ralph-loop:ralph-loop # Start Ralph Loop in current session +/ralph-loop:cancel-ralph # Cancel active Ralph Loop +/ralph-loop:help # Show help and usage ``` **Best Practices** (see full guide for details): diff --git a/docs/ralph-wiggum-guide.md b/docs/ralph-wiggum-guide.md index af033d19..6d33683c 100644 --- a/docs/ralph-wiggum-guide.md +++ b/docs/ralph-wiggum-guide.md @@ -73,33 +73,33 @@ while :; do cat PROMPT.md | claude ; done ### Commands -#### `/ralph-loop` +#### `/ralph-loop:ralph-loop` -Start an autonomous loop. +Start an autonomous loop in the current session. ```bash -/ralph-loop "" --max-iterations --completion-promise "" +/ralph-loop:ralph-loop ``` -**Parameters**: +When invoked, this skill prompts you to configure: +- **Task prompt**: The work to be done (persists across iterations) +- **Max iterations**: Safety limit on iterations (recommended: always set this) +- **Completion promise**: The phrase that signals completion (e.g., `COMPLETE`) -| Parameter | Type | Default | Description | -|-----------|------|---------|-------------| -| `` | string | required | Task description (persists across all iterations) | -| `--max-iterations` | integer | unlimited | Safety limit on iterations | -| `--completion-promise` | string | none | Exact string that signals completion | - -**Example**: -```bash -/ralph-loop "Build a REST API for todos. Requirements: CRUD operations, input validation, tests. Output COMPLETE when done." --completion-promise "COMPLETE" --max-iterations 50 -``` - -#### `/cancel-ralph` +#### `/ralph-loop:cancel-ralph` Cancel the active Ralph loop. ```bash -/cancel-ralph +/ralph-loop:cancel-ralph +``` + +#### `/ralph-loop:help` + +Show help and usage information. + +```bash +/ralph-loop:help ``` ### State File @@ -153,24 +153,38 @@ The official implementation (and Claudeman) prevents false positives when comple - Documentation or examples - Comments -**Solution**: Only mark as complete if the loop was already `active` when the phrase is detected. +**Solution**: Claudeman uses **occurrence-based detection** to distinguish prompts from actual completions: +- **1st occurrence**: Store as expected phrase (likely in the prompt) +- **2nd occurrence**: Emit `completionDetected` (actual completion) +- **If loop already active**: Emit immediately (explicit loop start via `/ralph-loop:ralph-loop`) ```typescript // From claudeman/src/inner-loop-tracker.ts -if (!this._loopState.active) { - // Just record the expected completion phrase without marking as complete +private handleCompletionPhrase(phrase: string): void { + const count = (this._completionPhraseCount.get(phrase) || 0) + 1; + this._completionPhraseCount.set(phrase, count); + + // Store phrase on first occurrence if (!this._loopState.completionPhrase) { this._loopState.completionPhrase = phrase; + this._loopState.lastActivity = Date.now(); + this.emit('loopUpdate', this.loopState); } - return; -} -// Loop was active, this is a real completion -this._loopState.completionPhrase = phrase; -this._loopState.active = false; -this.emit('completionDetected', phrase); + // Emit completion if loop is active OR this is 2nd+ occurrence + if (this._loopState.active || count >= 2) { + this._loopState.active = false; + this._loopState.lastActivity = Date.now(); + this.emit('completionDetected', phrase); + this.emit('loopUpdate', this.loopState); + } +} ``` +This approach handles both scenarios: +1. **Explicit loop start**: User runs `/ralph-loop:ralph-loop`, loop is active, first completion phrase triggers +2. **Implicit completion**: Phrase appears in prompt (1st), then Claude outputs it on completion (2nd) + --- ## TodoWrite Tool Integration @@ -319,7 +333,8 @@ For more sophisticated evaluation, use LLM-based hooks: > This cannot be overstated: always set `--max-iterations`. Autonomous loops consume tokens rapidly. A typical 50-iteration loop on a medium-sized codebase can cost $50-100+ in API usage. ```bash -/ralph-loop "..." --max-iterations 30 --completion-promise "DONE" +/ralph-loop:ralph-loop +# Then configure: max-iterations=30, completion-promise="DONE" ``` ### 2. Define Clear, Measurable Success Criteria @@ -416,10 +431,12 @@ This creates recovery points and shows progress in git history. ```bash # Test with 1 iteration first -/ralph-loop "..." --max-iterations 1 +/ralph-loop:ralph-loop +# Configure: max-iterations=1 # Then run full loop -/ralph-loop "..." --max-iterations 50 +/ralph-loop:ralph-loop +# Configure: max-iterations=50 ``` ### 8. Use Git for Safety @@ -607,12 +624,35 @@ The tracker automatically enables when detecting: | Pattern | Example | Regex | |---------|---------|-------| -| Ralph command | `/ralph-loop` | `/\/ralph-loop\|starting ralph/i` | +| Ralph command | `/ralph-loop:ralph-loop` | `/\/ralph-loop\|starting ralph/i` | | Promise tag | `COMPLETE` | `/([^<]+)<\/promise>/` | | TodoWrite | `Todos have been modified` | `/TodoWrite\|todos?\s*(?:updated\|written)/i` | | Iteration | `Iteration 5/50` or `[5/50]` | `/(?:iteration)\s*#?(\d+)(?:\s*[\/of]\s*(\d+))?/i` | | Todo checkbox | `- [ ] Task` | `/^[-*]\s*\[([xX ])\]\s+(.+)$/gm` | | Todo indicator | `Todo: ☐ Task` | `/Todo:\s*(☐\|◐\|✓)/g` | +| All complete | `All tasks completed` | `/all\s+tasks?\s+completed?\|all\s+done/i` | +| Task done | `Task 8 is done` | `/task\s*#?\d+\s*(?:is\s+)?done/i` | + +### Completion Detection + +Multi-strategy detection to catch various completion signals: + +1. **Tagged phrase**: `PHRASE` - First occurrence stores phrase, second triggers completion +2. **Bare phrase**: Detects phrase without tags once expected phrase is known (e.g., Claude outputs `COMPLETE` instead of `COMPLETE`) +3. **All complete signals**: Detects "All X files/tasks created/completed" messages, marks all todos complete and emits completion +4. **Explicit task completion**: Matches "Task N is done" patterns + +### Session Lifecycle + +Each session has its **own independent tracker**: + +| Action | Result | +|--------|--------| +| New session opened | Fresh tracker, no carryover | +| Tab closed | Tracker state cleared, UI panel hides | +| Switch tabs | Panel shows tracker for active session | +| `tracker.reset()` | Clears todos/state, keeps enabled status | +| `tracker.fullReset()` | Complete reset to initial state | ### State Structure @@ -643,6 +683,16 @@ interface InnerTodoItem { | GET | `/api/sessions/:id/inner-state` | Get loop state and todos | | POST | `/api/sessions/:id/inner-config` | Configure tracker settings | +**POST `/inner-config` Options**: +```json +{ + "enabled": true, // Enable/disable tracker + "reset": true, // Soft reset (clears state, keeps enabled) + "reset": "full", // Full reset (clears everything) + "completionPhrase": "DONE" // Set expected completion phrase +} +``` + **GET Response**: ```json { @@ -676,9 +726,9 @@ interface InnerTodoItem { ### Skill Commands ```bash -/ralph-loop # Start Ralph loop in current session -/cancel-ralph # Cancel active Ralph loop -/ralph-loop:help # Show help for Ralph loop +/ralph-loop:ralph-loop # Start Ralph Loop in current session +/ralph-loop:cancel-ralph # Cancel active Ralph Loop +/ralph-loop:help # Show help and usage ``` ---