diff --git a/CLAUDE.md b/CLAUDE.md index 58c9d103..38f9eb07 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -10,6 +10,12 @@ Claudeman is a Claude Code session manager with a web interface and autonomous R **Requirements**: Node.js 18+, Claude CLI (`claude`) installed and available in PATH +## First-Time Setup + +```bash +npm install +``` + ## Commands **CRITICAL**: `npm run dev` runs CLI help, NOT the web server. Use `npx tsx src/index.ts web` for development. @@ -24,15 +30,27 @@ npx tsx src/index.ts web -p 8080 # Dev mode with custom port node dist/index.js web # After npm run build claudeman web # After npm link -# Testing (vitest with globals: true - no imports needed for describe/it/expect) +# Testing (vitest) +# Note: globals: true configured - no imports needed for describe/it/expect npm run test # Run all tests once npm run test:watch # Watch mode npm run test:coverage # With coverage report npx vitest run test/session.test.ts # Single file npx vitest run -t "should create session" # By pattern -# Tests use ports 3099-3121 to avoid conflicts with dev server (3000) -# Test timeout: 30s (configured in vitest.config.ts for integration tests) + +# Test port allocation (add new tests in next available range): +# 3099-3101: quick-start.test.ts +# 3102: session.test.ts +# 3105-3106: scheduled-runs.test.ts +# 3107-3108: sse-events.test.ts +# 3110-3112: edge-cases.test.ts +# 3115-3116: integration-flows.test.ts +# 3120-3121: session-cleanup.test.ts +# (no port): respawn-controller.test.ts, inner-loop-tracker.test.ts, pty-interactive.test.ts (unit tests) +# Next available: 3122+ + # Tests mock PTY - no real Claude CLI spawned +# Test timeout: 30s (configured in vitest.config.ts) # TypeScript checking (no linter configured) npx tsc --noEmit # Type check without building @@ -41,48 +59,25 @@ npx tsc --noEmit # Type check without building screen -ls # List GNU screen sessions screen -r # Attach to screen session curl localhost:3000/api/sessions # Check active sessions +cat ~/.claudeman/state.json | jq . # View main state +cat ~/.claudeman/state-inner.json | jq . # View inner loop state ``` ## Architecture -``` -src/ -├── index.ts # CLI entry (commander) -├── cli.ts # CLI commands -├── session.ts # Core: PTY wrapper for Claude CLI + token tracking -├── session-manager.ts # Manages multiple sessions -├── screen-manager.ts # GNU screen persistence + process stats -├── respawn-controller.ts # Auto-respawn state machine -├── ralph-loop.ts # Autonomous task assignment -├── task.ts # Task class implementation -├── task-queue.ts # Priority queue with dependencies -├── task-tracker.ts # Background task detection from terminal output -├── inner-loop-tracker.ts # Detect Ralph loops and todos inside Claude sessions -├── state-store.ts # Persistence to ~/.claudeman/state.json -├── types.ts # All TypeScript interfaces -├── web/ -│ ├── server.ts # Fastify REST API + SSE + session restoration -│ └── public/ # Vanilla JS frontend (xterm.js, no bundler) -│ ├── app.js # Main app logic, SSE handling, tab management -│ ├── styles.css # All styles including responsive/mobile -│ └── index.html # Single page with modal templates -└── templates/ - └── claude-md.ts # CLAUDE.md generator for new cases +### Key Files -test/ # All tests use vitest -├── session.test.ts # Core session creation, lifecycle, PTY behavior (port 3102) -├── pty-interactive.test.ts # Interactive mode, terminal input/output (unit test, no server) -├── respawn-controller.test.ts # Respawn state machine, idle detection (unit test, no server) -├── inner-loop-tracker.test.ts # Ralph loop and todo detection parsing (unit test, no server) -├── quick-start.test.ts # Quick-start API endpoint (ports 3099-3101) -├── scheduled-runs.test.ts # Timed/scheduled session runs (ports 3105-3106) -├── sse-events.test.ts # Server-Sent Events broadcasting (ports 3107-3108) -├── integration-flows.test.ts # Multi-step workflow tests (ports 3115-3116) -├── session-cleanup.test.ts # Resource cleanup, buffer trimming (ports 3120-3121) -└── edge-cases.test.ts # Error handling, boundary conditions (ports 3110-3112) -``` - -**Test ports**: Integration tests use unique port ranges (3099-3121) to allow parallel execution. Unit tests don't need a server. Dev server uses port 3000. +| File | Purpose | +|------|---------| +| `src/session.ts` | Core PTY wrapper for Claude CLI. Modes: `runPrompt()`, `startInteractive()`, `startShell()` | +| `src/respawn-controller.ts` | State machine for autonomous session cycling | +| `src/screen-manager.ts` | GNU screen persistence, ghost discovery, 4-strategy kill | +| `src/inner-loop-tracker.ts` | Detects `PHRASE`, todos, loop status in output | +| `src/task-tracker.ts` | Parses background task output (agent IDs, status) from Claude CLI | +| `src/state-store.ts` | JSON persistence to `~/.claudeman/` with debounced (100ms) writes | +| `src/web/server.ts` | Fastify REST API + SSE at `/api/events` | +| `src/web/public/app.js` | Frontend: SSE handling, xterm.js, tab management | +| `src/types.ts` | All TypeScript interfaces | ### Data Flow @@ -91,16 +86,6 @@ test/ # All tests use vitest 3. **WebServer** broadcasts events to SSE clients at `/api/events` 4. State persists to `~/.claudeman/state.json` via **StateStore** -### Key Components - -| Component | File | Purpose | -|-----------|------|---------| -| Session | `session.ts` | PTY wrapper for Claude CLI. Modes: `runPrompt()`, `startInteractive()`, `startShell()` | -| RespawnController | `respawn-controller.ts` | State machine for autonomous session cycling (see diagram below) | -| ScreenManager | `screen-manager.ts` | GNU screen persistence, ghost discovery, 4-strategy kill | -| WebServer | `web/server.ts` | Fastify REST + SSE at `/api/events` | -| InnerLoopTracker | `inner-loop-tracker.ts` | Detects `PHRASE`, todos, loop status in output | - ### Respawn State Machine ``` @@ -222,14 +207,16 @@ Both wait for idle. Configure via `session.setAutoCompact()` / `session.setAutoC ### Inner Loop Tracking -Detects Ralph loops and todos inside Claude sessions. **Disabled by default** - auto-enables when Ralph-related patterns are detected: +Detects Ralph loops and todos inside Claude sessions. **Disabled by default** but auto-enables when any of these patterns are detected in terminal output: - `/ralph-loop` command - `PHRASE` completion phrases - `TodoWrite` tool usage - Iteration patterns (`Iteration 5/50`, `[5/50]`) - Todo checkboxes (`- [ ]`/`- [x]`) or indicator icons (`☐`/`◐`/`✓`) -API: `GET /api/sessions/:id/inner-state`. UI: collapsible panel below tabs with enable/disable toggle. Use `tracker.enable()` / `tracker.disable()` for programmatic control, or `POST /api/sessions/:id/inner-config` with `{ enabled: boolean }` via API. +See `inner-loop-tracker.ts:shouldAutoEnable()` for detection logic. + +API: `GET /api/sessions/:id/inner-state`. UI: collapsible panel below tabs. Use `tracker.enable()` / `tracker.disable()` for programmatic control, or `POST /api/sessions/:id/inner-config` with `{ enabled: boolean }` via API. ### Terminal Display Fix @@ -239,32 +226,39 @@ Tab switch/new session fix: clear xterm → write buffer → resize PTY → Ctrl All events broadcast to `/api/events` with format: `{ type: string, sessionId?: string, data: any }`. -Event prefixes: `session:`, `task:`, `respawn:`, `scheduled:`, `case:`, `screen:`, `init`. Key events: `session:idle`, `session:working`, `session:terminal`, `session:clearTerminal`, `session:completion`, `session:autoClear`, `session:autoCompact`, `session:innerLoopUpdate`, `session:innerTodoUpdate`, `session:innerCompletionDetected`. +Event prefixes: `session:`, `task:`, `respawn:`, `scheduled:`, `case:`, `screen:`, `init`. + +Key events for frontend handling (see `app.js:handleSSEEvent()`): +- `session:idle`, `session:working` - Status indicator updates +- `session:terminal`, `session:clearTerminal` - Terminal content +- `session:completion`, `session:autoClear`, `session:autoCompact` - Lifecycle events +- `session:innerLoopUpdate`, `session:innerTodoUpdate`, `session:innerCompletionDetected` - Ralph tracking ### Frontend (app.js) -Vanilla JS + xterm.js. `handleSSEEvent()` dispatches events, `switchToSession()` manages tabs. 60fps: server batches 16ms, client uses `requestAnimationFrame`. +Vanilla JS + xterm.js. Key functions: +- `handleSSEEvent()` - Dispatches events to appropriate handlers +- `switchToSession()` - Tab management and terminal focus +- `createSessionTab()` - Tab creation and xterm setup + +**60fps Rendering Pipeline**: +- Server batches terminal data every 16ms before broadcasting via SSE +- Client uses `requestAnimationFrame` to batch xterm.js writes +- Prevents UI jank during high-throughput Claude output ### State Store Writes debounced (100ms) to `~/.claudeman/state.json`. Batches rapid changes. +### TypeScript Config + +Module resolution: NodeNext. Target: ES2022. Strict mode enabled. See `tsconfig.json` for full settings. + ## Adding New Features -### New API Endpoint -1. Add types to `src/types.ts` -2. Add route in `src/web/server.ts` within `buildServer()` -3. Use `createErrorResponse()` for errors - -### New SSE Event -1. Emit from component via `broadcast()` in server.ts -2. Handle in `src/web/public/app.js` `handleSSEEvent()` switch - -### New Session Event -1. Add to `SessionEvents` interface in `src/session.ts` -2. Emit via `this.emit()` -3. Subscribe in `src/web/server.ts` when wiring session to SSE -4. Handle in frontend SSE listener +- **API endpoint**: Add types in `types.ts`, route in `server.ts:buildServer()`, use `createErrorResponse()` for errors +- **SSE event**: Emit via `broadcast()` in server.ts, handle in `app.js:handleSSEEvent()` switch +- **Session event**: Add to `SessionEvents` interface in `session.ts`, emit via `this.emit()`, subscribe in server.ts, handle in frontend ## Session Lifecycle & Cleanup @@ -290,7 +284,11 @@ Long-running sessions are supported with automatic trimming: Uses `agent-browser` for web UI automation. Full test plan: `.claude/skills/e2e-test.md` ```bash -npx agent-browser open http://localhost:3000 && npx agent-browser snapshot +npx agent-browser open http://localhost:3000 +npx agent-browser wait --load networkidle +npx agent-browser snapshot +npx agent-browser find text "Run Claude" click +npx agent-browser close ``` ## API Routes Quick Reference @@ -328,8 +326,75 @@ claudeman ralph start [--min-hours N] # Start autonomous loop claudeman status # Overall status ``` -## Notes +## Keyboard Shortcuts -- State persists to `~/.claudeman/state.json`, `~/.claudeman/state-inner.json`, and `~/.claudeman/screens.json` -- Inner loop/todo state persists separately in `state-inner.json` to reduce write frequency -- Cases created in `~/claudeman-cases/` by default +| Shortcut | Action | +|----------|--------| +| `Ctrl+Enter` | Run Claude (create case + interactive session) | +| `Ctrl+W` | Close current session | +| `Ctrl+Tab` | Switch to next session | +| `Ctrl+K` | Kill all sessions | +| `Ctrl+L` | Clear terminal | +| `Ctrl++/-` | Increase/decrease font size | +| `Ctrl+?` | Show keyboard shortcuts help | +| `Escape` | Close panels and modals | + +## State Files + +| File | Purpose | +|------|---------| +| `~/.claudeman/state.json` | Sessions, tasks, config | +| `~/.claudeman/state-inner.json` | Inner loop/todo state (separate to reduce writes) | +| `~/.claudeman/screens.json` | Screen session metadata | + +Cases created in `~/claudeman-cases/` by default. + +## Documentation + +Extended documentation is available in the `docs/` directory: + +| Document | Description | +|----------|-------------| +| [`docs/ralph-wiggum-guide.md`](docs/ralph-wiggum-guide.md) | Complete Ralph Wiggum loop guide: official plugin reference, best practices, prompt templates, troubleshooting | +| [`docs/claude-code-hooks-reference.md`](docs/claude-code-hooks-reference.md) | Official Claude Code hooks documentation: all events, configuration, examples | + +### Quick Reference: Ralph Wiggum Loops + +Ralph Wiggum is an autonomous loop technique that lets Claude work iteratively until completion criteria are met. + +**Core Pattern**: `PHRASE` - The completion signal that tells the loop to stop. + +**Official Plugin Commands**: +```bash +/ralph-loop "" --max-iterations 50 --completion-promise "COMPLETE" +/cancel-ralph +``` + +**Best Practices** (see full guide for details): +1. **Always set `--max-iterations`** - Safety limit to prevent runaway costs +2. **Define clear success criteria** - Tests pass, lint clean, specific outputs +3. **Use test-driven verification** - Built-in feedback loop +4. **Include escape hatches** - "If stuck after N iterations, document and stop" +5. **Commit frequently** - Recovery points in git history + +**Claudeman Implementation**: The `InnerLoopTracker` class (`src/inner-loop-tracker.ts`) detects Ralph patterns in Claude output and tracks loop state, todos, and completion phrases. It auto-enables when Ralph-related patterns are detected. + +**API**: +- `GET /api/sessions/:id/inner-state` - Loop state and todos +- `POST /api/sessions/:id/inner-config` - Configure tracker + +**SSE Events**: +- `session:innerLoopUpdate` - Loop state changes +- `session:innerTodoUpdate` - Todo list updates +- `session:innerCompletionDetected` - Completion phrase detected + +### External References + +**Official Anthropic Documentation**: +- [Claude Code Hooks](https://code.claude.com/docs/en/hooks) +- [Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices) +- [Ralph Wiggum Plugin](https://github.com/anthropics/claude-code/tree/main/plugins/ralph-wiggum) + +**Community Resources**: +- [Awesome Claude - Ralph Wiggum](https://awesomeclaude.ai/ralph-wiggum) +- [Claude Fast - Autonomous Loops](https://claudefa.st/blog/guide/mechanics/autonomous-agent-loops) diff --git a/docs/claude-code-hooks-reference.md b/docs/claude-code-hooks-reference.md new file mode 100644 index 00000000..3e69282c --- /dev/null +++ b/docs/claude-code-hooks-reference.md @@ -0,0 +1,596 @@ +# Claude Code Hooks Reference + +> Official documentation for Claude Code hooks system, extracted from [code.claude.com](https://code.claude.com/docs/en/hooks). + +**Last Updated**: 2026-01-20 +**Source**: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks) + +--- + +## Overview + +Hooks are automated scripts that execute at specific events during your Claude Code session. They allow you to: +- Validate, modify, or block tool usage +- Add context to prompts +- Implement custom workflows +- Control agent behavior + +--- + +## Configuration + +Hooks are configured in settings files: + +| File | Scope | +|------|-------| +| `~/.claude/settings.json` | User (global) | +| `.claude/settings.json` | Project | +| `.claude/settings.local.json` | Local project (gitignored) | +| Plugin hook files | Plugin-specific | + +### Basic Structure + +```json +{ + "hooks": { + "EventName": [ + { + "matcher": "ToolPattern", + "hooks": [ + { + "type": "command", + "command": "your-command-here" + } + ] + } + ] + } +} +``` + +**Key Fields**: +- `matcher`: Pattern to match tool names (case-sensitive, supports regex like `Edit|Write` or `*` for all) +- `type`: `"command"` for bash or `"prompt"` for LLM-based evaluation +- `command`: Bash command to execute +- `prompt`: LLM prompt for evaluation (prompt-based hooks only) +- `timeout`: Optional timeout in seconds (default: 60) + +--- + +## Hook Events + +### PreToolUse + +**When**: After Claude creates tool parameters, before processing the tool call. + +**Use Cases**: Approval, denial, or modification of tool calls. + +**Common Matchers**: +- `Bash` - Shell commands +- `Write` - File writing +- `Edit` - File editing +- `Read` - File reading +- `Task` - Subagent tasks +- `WebFetch`, `WebSearch` - Web operations +- `mcp____` - MCP tools + +**Output Control**: +```json +{ + "hookSpecificOutput": { + "hookEventName": "PreToolUse", + "permissionDecision": "allow|deny|ask", + "permissionDecisionReason": "string", + "updatedInput": { + "field_to_modify": "new value" + }, + "additionalContext": "Context for Claude" + } +} +``` + +### PermissionRequest + +**When**: When the user is shown a permission dialog. + +**Use Cases**: Auto-approve or deny permissions. + +**Output Control**: +```json +{ + "hookSpecificOutput": { + "hookEventName": "PermissionRequest", + "decision": { + "behavior": "allow|deny", + "updatedInput": { }, + "message": "deny reason", + "interrupt": false + } + } +} +``` + +### PostToolUse + +**When**: Immediately after a tool completes successfully. + +**Use Cases**: Provide feedback, run formatters/linters, log operations. + +**Output Control**: +```json +{ + "decision": "block", + "reason": "Explanation", + "hookSpecificOutput": { + "hookEventName": "PostToolUse", + "additionalContext": "Additional information" + } +} +``` + +### Notification + +**When**: When Claude Code sends notifications. + +**Matchers**: +- `permission_prompt` +- `idle_prompt` +- `auth_success` +- `elicitation_dialog` + +### UserPromptSubmit + +**When**: When the user submits a prompt, before Claude processes it. + +**Use Cases**: Add context, validate, or block prompts. + +**Output Control**: +```json +{ + "decision": "block", + "reason": "Explanation", + "hookSpecificOutput": { + "hookEventName": "UserPromptSubmit", + "additionalContext": "My additional context" + } +} +``` + +### Stop + +**When**: When the main Claude Code agent finishes responding. + +**Important**: Does NOT run on user interrupt. + +**Use Cases**: **Ralph Wiggum loops** - block exit and refeed prompt. + +**Output Control**: +```json +{ + "decision": "block", + "reason": "Must provide when blocking" +} +``` + +Or to allow exit: +```json +{ + "continue": true, + "stopReason": "optional message" +} +``` + +**Note**: For Stop events, `"continue": false` takes precedence over `"decision": "block"`. + +### SubagentStop + +**When**: When a subagent (Task tool call) finishes responding. + +**Use Cases**: Control nested loops, verify subagent output. + +### PreCompact + +**When**: Before a compact operation. + +**Matchers**: +- `manual` - Invoked from `/compact` +- `auto` - Invoked from auto-compact + +### SessionStart + +**When**: When Claude Code starts or resumes a session. + +**Matchers**: +- `startup` - Fresh start +- `resume` - From `--resume`, `--continue`, or `/resume` +- `clear` - From `/clear` +- `compact` - From auto or manual compact + +**Use Cases**: Load development context, set environment variables. + +**Persisting Environment Variables**: +```bash +#!/bin/bash +if [ -n "$CLAUDE_ENV_FILE" ]; then + echo 'export NODE_ENV=production' >> "$CLAUDE_ENV_FILE" + echo 'export API_KEY=your-api-key' >> "$CLAUDE_ENV_FILE" +fi +exit 0 +``` + +**Output Control**: +```json +{ + "hookSpecificOutput": { + "hookEventName": "SessionStart", + "additionalContext": "Context to load" + } +} +``` + +### SessionEnd + +**When**: When a session ends. + +**Reason Values**: +- `clear` +- `logout` +- `prompt_input_exit` +- `other` + +**Use Cases**: Cleanup tasks, logging. + +--- + +## Hook Input + +Hooks receive JSON via stdin with common fields: + +```json +{ + "session_id": "abc123", + "transcript_path": "/path/to/transcript.jsonl", + "cwd": "/current/directory", + "permission_mode": "default", + "hook_event_name": "PreToolUse", + "tool_name": "Bash", + "tool_input": { }, + "tool_use_id": "toolu_01ABC123..." +} +``` + +### Tool-Specific Input + +**Bash**: +```json +{ + "tool_name": "Bash", + "tool_input": { + "command": "psql -c 'SELECT * FROM users'", + "description": "Query the users table", + "timeout": 120000 + } +} +``` + +**Write**: +```json +{ + "tool_name": "Write", + "tool_input": { + "file_path": "/path/to/file.txt", + "content": "file content" + } +} +``` + +**Edit**: +```json +{ + "tool_name": "Edit", + "tool_input": { + "file_path": "/path/to/file.txt", + "old_string": "original text", + "new_string": "replacement text" + } +} +``` + +--- + +## Hook Output + +### Exit Codes + +| Code | Behavior | +|------|----------| +| 0 | Success. `stdout` processed (shown in verbose or added as context) | +| 2 | Blocking error. Only `stderr` used. Blocks tool/prompt based on event | +| Other | Non-blocking error. `stderr` shown in verbose, execution continues | + +### JSON Output (Exit Code 0) + +```json +{ + "continue": true, + "stopReason": "optional message", + "suppressOutput": true, + "systemMessage": "optional warning" +} +``` + +--- + +## Prompt-Based Hooks + +For Stop and SubagentStop events, you can use LLM-based evaluation: + +```json +{ + "hooks": { + "Stop": [ + { + "hooks": [ + { + "type": "prompt", + "prompt": "Should Claude stop? Context: $ARGUMENTS\n\nCheck if all tasks are complete.", + "timeout": 30 + } + ] + } + ] + } +} +``` + +**LLM Response Format**: +```json +{ + "ok": true, + "reason": "Explanation when ok is false" +} +``` + +--- + +## Component-Scoped Hooks + +Hooks can be defined in Skills, Agents, and Slash Commands using frontmatter: + +```markdown +--- +name: secure-operations +hooks: + PreToolUse: + - matcher: "Bash" + hooks: + - type: command + command: "./scripts/security-check.sh" +--- +``` + +These hooks: +- Are scoped to the component's lifecycle +- Only run when that component is active +- Support: PreToolUse, PostToolUse, Stop + +--- + +## MCP Tools + +MCP tools follow the pattern `mcp____`: + +```json +{ + "hooks": { + "PreToolUse": [ + { + "matcher": "mcp__memory__.*", + "hooks": [ + { + "type": "command", + "command": "echo 'Memory operation' >> ~/mcp.log" + } + ] + }, + { + "matcher": "mcp__.*__write.*", + "hooks": [ + { + "type": "command", + "command": "/home/user/scripts/validate-mcp-write.py" + } + ] + } + ] + } +} +``` + +--- + +## Examples + +### Bash Command Validation + +```python +#!/usr/bin/env python3 +import json +import re +import sys + +VALIDATION_RULES = [ + (r"\bgrep\b(?!.*\|)", "Use 'rg' instead of 'grep'"), + (r"\bfind\s+\S+\s+-name\b", "Use 'rg --files' instead of 'find -name'"), +] + +try: + input_data = json.load(sys.stdin) +except json.JSONDecodeError as e: + print(f"Error: {e}", file=sys.stderr) + sys.exit(1) + +tool_name = input_data.get("tool_name", "") +tool_input = input_data.get("tool_input", {}) +command = tool_input.get("command", "") + +if tool_name != "Bash" or not command: + sys.exit(1) + +issues = [] +for pattern, message in VALIDATION_RULES: + if re.search(pattern, command): + issues.append(message) + +if issues: + for message in issues: + print(f"- {message}", file=sys.stderr) + sys.exit(2) +``` + +### Auto-Approve Documentation Reads + +```python +#!/usr/bin/env python3 +import json +import sys + +try: + input_data = json.load(sys.stdin) +except json.JSONDecodeError as e: + print(f"Error: {e}", file=sys.stderr) + sys.exit(1) + +tool_name = input_data.get("tool_name", "") +tool_input = input_data.get("tool_input", {}) + +if tool_name == "Read": + file_path = tool_input.get("file_path", "") + if file_path.endswith((".md", ".mdx", ".txt", ".json")): + output = { + "hookSpecificOutput": { + "hookEventName": "PreToolUse", + "permissionDecision": "allow", + "permissionDecisionReason": "Documentation file auto-approved" + }, + "suppressOutput": True + } + print(json.dumps(output)) + sys.exit(0) + +sys.exit(0) +``` + +### Post-Write Formatter + +```json +{ + "hooks": { + "PostToolUse": [ + { + "matcher": "Edit|Write", + "hooks": [ + { + "type": "command", + "command": "npx prettier --write \"$TOOL_INPUT_FILE_PATH\" 2>/dev/null || true" + } + ] + } + ] + } +} +``` + +### Ralph Wiggum Stop Hook + +```bash +#!/bin/bash +# ralph-stop-hook.sh + +STATE_FILE=".claude/ralph-loop.local.md" + +# Check if state file exists +if [ ! -f "$STATE_FILE" ]; then + exit 0 # No active loop, allow exit +fi + +# Read state from YAML frontmatter +ENABLED=$(grep -m1 "^enabled:" "$STATE_FILE" | cut -d' ' -f2) +ITERATION=$(grep -m1 "^iteration:" "$STATE_FILE" | cut -d' ' -f2) +MAX_ITER=$(grep -m1 "^max-iterations:" "$STATE_FILE" | cut -d' ' -f2) +PROMISE=$(grep -m1 "^completion-promise:" "$STATE_FILE" | cut -d' ' -f2-) + +# Check if disabled +if [ "$ENABLED" = "false" ]; then + exit 0 +fi + +# Check max iterations +if [ -n "$MAX_ITER" ] && [ "$ITERATION" -ge "$MAX_ITER" ]; then + exit 0 +fi + +# Check for completion promise in output +if [ -n "$PROMISE" ]; then + if echo "$CLAUDE_OUTPUT" | grep -q "$PROMISE"; then + exit 0 + fi +fi + +# Block exit, increment iteration +NEW_ITER=$((ITERATION + 1)) +sed -i "s/^iteration:.*/iteration: $NEW_ITER/" "$STATE_FILE" + +# Output block decision +echo '{"decision": "block", "reason": "Completion promise not found. Iteration '"$NEW_ITER"'."}' +exit 0 +``` + +--- + +## Environment Variables + +| Variable | Description | +|----------|-------------| +| `CLAUDE_PROJECT_DIR` | Project root directory | +| `CLAUDE_CODE_REMOTE` | `"true"` for web, empty for CLI | +| `CLAUDE_ENV_FILE` | Path to write persistent env vars (SessionStart) | + +--- + +## Debugging + +Use `claude --debug` to see detailed hook execution: + +``` +[DEBUG] Executing hooks for PostToolUse:Write +[DEBUG] Found 1 hook matchers in settings +[DEBUG] Matched 1 hooks for query "Write" +[DEBUG] Executing hook command: with timeout 60000ms +[DEBUG] Hook command completed with status 0: +``` + +Use `/hooks` command to view registered hooks and make changes. + +--- + +## Execution Details + +- **Timeout**: 60-second default per hook, configurable +- **Parallelization**: All matching hooks run in parallel +- **Deduplication**: Identical commands deduplicated automatically +- **Matchers**: Only apply to tool-based hooks (PreToolUse, PostToolUse, PostToolUseFailure, PermissionRequest) + +--- + +## Security Best Practices + +1. **Validate and sanitize inputs** - Never trust input data blindly +2. **Always quote shell variables** - Use `"$VAR"` not `$VAR` +3. **Block path traversal** - Check for `..` in file paths +4. **Use absolute paths** - Specify full paths for scripts (use `$CLAUDE_PROJECT_DIR`) +5. **Skip sensitive files** - Avoid `.env`, `.git/`, keys, etc. + +--- + +*Source: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)* diff --git a/docs/ralph-wiggum-guide.md b/docs/ralph-wiggum-guide.md new file mode 100644 index 00000000..af033d19 --- /dev/null +++ b/docs/ralph-wiggum-guide.md @@ -0,0 +1,776 @@ +# Ralph Wiggum Loop: Complete Guide + +> This document consolidates official Anthropic documentation, community best practices, and implementation details for autonomous Claude Code loops. + +**Last Updated**: 2026-01-20 +**Sources**: [Official Anthropic Plugin](https://github.com/anthropics/claude-code/tree/main/plugins/ralph-wiggum), [Claude Code Docs](https://code.claude.com/docs/en/hooks), [Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices) + +--- + +## Table of Contents + +1. [Overview](#overview) +2. [Core Concept](#core-concept) +3. [Official Plugin Reference](#official-plugin-reference) +4. [The Promise Tag Contract](#the-promise-tag-contract) +5. [TodoWrite Tool Integration](#todowrite-tool-integration) +6. [Hooks System](#hooks-system) +7. [Best Practices](#best-practices) +8. [Prompt Templates](#prompt-templates) +9. [When to Use (and Not Use)](#when-to-use-and-not-use) +10. [Real-World Examples](#real-world-examples) +11. [Claudeman Implementation](#claudeman-implementation) +12. [Troubleshooting](#troubleshooting) + +--- + +## Overview + +Ralph Wiggum is an autonomous loop technique for Claude Code, named after The Simpsons character. It enables Claude to work iteratively on tasks for hours without human intervention, self-correcting until completion criteria are met. + +**Core Philosophy**: +- **Iteration > Perfection**: Don't aim for perfect on first try; let the loop refine +- **Failures Are Data**: "Deterministically bad" means failures are predictable and informative +- **Operator Skill Matters**: Success depends on writing good prompts, not just having a good model +- **Persistence Wins**: Keep trying until success; the loop handles retry logic + +**Origin**: Created by Geoffrey Huntley, formalized into an official Anthropic plugin by Boris Cherny (Head of Claude Code) in late 2025. + +--- + +## Core Concept + +The simplest form of a Ralph loop: + +```bash +while :; do cat PROMPT.md | claude ; done +``` + +**How It Works**: +1. Claude processes a task prompt +2. Attempts to exit when "done" +3. A **Stop hook** intercepts the exit +4. Checks for **completion promise** (e.g., `COMPLETE`) +5. If not found, re-feeds the same prompt +6. Files from previous iteration persist, so Claude sees its own work +7. Cycle repeats until completion or max iterations reached + +**Key insight**: The prompt never changes between iterations, but Claude's previous work persists in files, allowing autonomous improvement by reading past work. + +--- + +## Official Plugin Reference + +### Installation + +```bash +# Add Anthropic's official plugin marketplace +/plugin marketplace add anthropics/claude-plugins-official + +# Install Ralph Wiggum plugin +/plugin install ralph-wiggum@claude-plugins-official +``` + +### Commands + +#### `/ralph-loop` + +Start an autonomous loop. + +```bash +/ralph-loop "" --max-iterations --completion-promise "" +``` + +**Parameters**: + +| Parameter | Type | Default | Description | +|-----------|------|---------|-------------| +| `` | string | required | Task description (persists across all iterations) | +| `--max-iterations` | integer | unlimited | Safety limit on iterations | +| `--completion-promise` | string | none | Exact string that signals completion | + +**Example**: +```bash +/ralph-loop "Build a REST API for todos. Requirements: CRUD operations, input validation, tests. Output COMPLETE when done." --completion-promise "COMPLETE" --max-iterations 50 +``` + +#### `/cancel-ralph` + +Cancel the active Ralph loop. + +```bash +/cancel-ralph +``` + +### State File + +The plugin persists state to `.claude/ralph-loop.local.md`: + +```yaml +--- +enabled: true +iteration: 5 +max-iterations: 50 +completion-promise: "COMPLETE" +--- +# Original Prompt + +Build a REST API for todos... +``` + +**YAML Fields**: +- `enabled` (boolean): Controls hook activation +- `iteration` (integer): Current iteration count (0-indexed) +- `max-iterations` (integer): Optional maximum +- `completion-promise` (string): Optional completion text + +--- + +## The Promise Tag Contract + +The completion phrase pattern is the core contract between Claude and the loop system: + +``` +PHRASE +``` + +**Examples**: +- `COMPLETE` - Generic completion +- `TESTS_PASS` - Test-specific completion +- `TIME_COMPLETE` - Time-aware loop completion +- `FIXED` - Bug fix completion + +### How Completion Detection Works + +1. **Exact String Matching**: The `--completion-promise` uses case-sensitive exact matching +2. **Output Scanning**: The Stop hook scans Claude's final output for the promise tag +3. **Exit Control**: If found, exit is allowed. If not, loop continues. + +### False Positive Prevention + +The official implementation (and Claudeman) prevents false positives when completion phrases appear in: +- Initial prompts +- Documentation or examples +- Comments + +**Solution**: Only mark as complete if the loop was already `active` when the phrase is detected. + +```typescript +// From claudeman/src/inner-loop-tracker.ts +if (!this._loopState.active) { + // Just record the expected completion phrase without marking as complete + if (!this._loopState.completionPhrase) { + this._loopState.completionPhrase = phrase; + } + return; +} + +// Loop was active, this is a real completion +this._loopState.completionPhrase = phrase; +this._loopState.active = false; +this.emit('completionDetected', phrase); +``` + +--- + +## TodoWrite Tool Integration + +The **TodoWrite tool** is Claude Code's built-in task management system that integrates with Ralph loops. + +### How It Works + +Claude uses TodoWrite to: +1. Break complex tasks into subtasks +2. Track progress through iterations +3. Provide visibility into current state +4. Resume work after context resets + +### Todo Formats Detected + +**Format 1: Markdown Checkboxes** +```markdown +- [ ] Pending task +- [x] Completed task +- [X] Completed task (uppercase) +``` + +**Format 2: Status Indicators** +``` +Todo: ☐ Pending task +Todo: ◐ In progress task +Todo: ✓ Completed task +``` + +**Format 3: Parenthetical Status** +``` +- Task name (pending) +- Task name (in_progress) +- Task name (completed) +``` + +### System Reminder Integration + +From official Claude Code documentation: + +> After commands like an ls -la run via bash tool, system-reminder tags are injected to remind the model to use the TodoWrite tool if it hasn't been using it so far. + +The system prompt includes: +> "IMPORTANT: Always use the TodoWrite tool to plan and track tasks throughout the conversation." + +### Checklists for Complex Workflows + +From [Anthropic Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices): + +> For large tasks with multiple steps or requiring exhaustive solutions—like code migrations, fixing numerous lint errors, or running complex build scripts—improve performance by having Claude use a Markdown file (or even a GitHub issue!) as a checklist and working scratchpad. + +--- + +## Hooks System + +Ralph loops are powered by Claude Code's hooks system. Understanding hooks is essential for customization. + +### Hook Events Reference + +| Event | When | Use Case | +|-------|------|----------| +| `PreToolUse` | Before tool execution | Validate, modify, or block tool calls | +| `PostToolUse` | After tool completes | Provide feedback, run formatters/linters | +| `Stop` | When Claude finishes | **Ralph loop control** - block exit, refeed prompt | +| `SubagentStop` | When subagent finishes | Control nested loops | +| `UserPromptSubmit` | User submits prompt | Add context, validate input | +| `SessionStart` | Session begins | Load environment, context | +| `SessionEnd` | Session ends | Cleanup, logging | +| `PermissionRequest` | Permission dialog shown | Auto-approve/deny | +| `PreCompact` | Before compact | Backup, preprocessing | + +### Stop Hook for Ralph Loops + +The Stop hook is the key mechanism: + +```json +{ + "hooks": { + "Stop": [ + { + "hooks": [ + { + "type": "command", + "command": "./scripts/ralph-stop-hook.sh" + } + ] + } + ] + } +} +``` + +**Stop Hook Logic**: +1. Check if `.claude/ralph-loop.local.md` exists +2. Read `enabled` flag from YAML frontmatter +3. Check for `completion-promise` in output +4. Check if `iteration >= max-iterations` +5. If none match, block exit and refeed prompt + +### Hook Output for Stop Events + +```json +{ + "decision": "block", + "reason": "Completion promise not found. Restarting iteration." +} +``` + +Or to allow exit: +```json +{ + "continue": true, + "stopReason": "Completion promise detected" +} +``` + +### Prompt-Based Hooks + +For more sophisticated evaluation, use LLM-based hooks: + +```json +{ + "hooks": { + "Stop": [ + { + "hooks": [ + { + "type": "prompt", + "prompt": "Check if the task is complete. Context: $ARGUMENTS\n\nRespond with {\"ok\": true} if done, {\"ok\": false, \"reason\": \"...\"} if not.", + "timeout": 30 + } + ] + } + ] + } +} +``` + +--- + +## Best Practices + +### 1. Always Set `--max-iterations` + +> This cannot be overstated: always set `--max-iterations`. Autonomous loops consume tokens rapidly. A typical 50-iteration loop on a medium-sized codebase can cost $50-100+ in API usage. + +```bash +/ralph-loop "..." --max-iterations 30 --completion-promise "DONE" +``` + +### 2. Define Clear, Measurable Success Criteria + +**Bad**: +``` +Build a todo API and make it good. +``` + +**Good**: +``` +Build a REST API for todos. + +Completion criteria: +- All CRUD endpoints working (GET, POST, PUT, DELETE) +- Input validation with error messages +- Tests passing with >80% coverage +- README with API documentation + +Output COMPLETE when ALL criteria are met. +``` + +### 3. Use Test-Driven Verification + +> The most effective Ralph Loop tasks include built-in verification. This creates a natural feedback loop within the loop. + +``` +Implement user authentication using TDD: + +1. Write failing tests for each requirement +2. Implement feature to make tests pass +3. Run tests after each change +4. If any fail, debug and fix +5. Refactor if needed +6. Output TESTS_PASS when all tests green +``` + +### 4. Include Escape Hatches + +``` +Primary task: Implement feature X + +If stuck after 10 iterations: +- Document what's blocking progress +- List approaches that were attempted +- Suggest alternative approaches +- Output BLOCKED +``` + +### 5. Incremental Goals for Large Tasks + +**Bad**: +``` +Create a complete e-commerce platform. +``` + +**Good**: +``` +Build e-commerce platform in phases: + +Phase 1: User authentication +- JWT-based auth +- Tests passing +- Commit: "feat: add user auth" + +Phase 2: Product catalog +- CRUD for products +- Search functionality +- Tests passing +- Commit: "feat: add product catalog" + +Phase 3: Shopping cart +- Add/remove items +- Persist cart state +- Tests passing +- Commit: "feat: add shopping cart" + +Output COMPLETE when all phases done. +``` + +### 6. Commit Frequently + +``` +After each meaningful completion: +1. git add . +2. git commit -m "descriptive message" + +This creates recovery points and shows progress in git history. +``` + +### 7. Test Before Long Runs + +> Pro tip: Test manually with one iteration before running 50-iteration loops. + +```bash +# Test with 1 iteration first +/ralph-loop "..." --max-iterations 1 + +# Then run full loop +/ralph-loop "..." --max-iterations 50 +``` + +### 8. Use Git for Safety + +> Always run Ralph loops in a git-tracked directory. If something goes wrong, you can revert. Each iteration adds to git history, giving you a clear trail of what changed. + +--- + +## Prompt Templates + +### Template 1: Test-Driven Development + +```markdown +# Task: [FEATURE_NAME] + +## Requirements +- [Requirement 1] +- [Requirement 2] +- [Requirement 3] + +## Approach +Follow TDD methodology: +1. Write failing tests for each requirement +2. Implement minimal code to pass tests +3. Run tests: `npm test` +4. If tests fail, read error, fix, repeat +5. When all tests pass, refactor if needed +6. Commit: `git add . && git commit -m "feat: [feature]"` + +## Completion +Output TESTS_PASS when: +- All tests pass +- Code is committed +- No lint errors +``` + +### Template 2: Migration/Refactor + +```markdown +# Task: Migrate from [OLD] to [NEW] + +## Scope +Files to migrate: `src/**/*.ts` + +## Migration Steps +For each file: +1. Update imports +2. Replace deprecated patterns +3. Run type check: `npx tsc --noEmit` +4. If errors, fix them +5. Run tests: `npm test` +6. Commit: `git commit -m "refactor: migrate [file]"` + +## Completion +Output MIGRATION_COMPLETE when: +- All files migrated +- Type check passes +- All tests pass +- All changes committed +``` + +### Template 3: Bug Fix + +```markdown +# Bug: [BUG_DESCRIPTION] + +## Reproduction +[Steps to reproduce] + +## Investigation +1. Find the root cause +2. Document findings + +## Fix +1. Write a failing test that reproduces the bug +2. Implement the fix +3. Verify test passes +4. Check for regressions: `npm test` +5. Commit: `git commit -m "fix: [description]"` + +## Completion +Output FIXED when: +- Bug is fixed +- Test added to prevent regression +- All tests pass +``` + +### Template 4: Time-Aware Loop + +```markdown +# Task: Optimize API performance + +## Primary Goals +1. Profile existing endpoints +2. Identify bottlenecks +3. Implement optimizations +4. Verify improvements + +## Duration +Minimum runtime: 4 hours + +## Self-Generated Tasks +If primary goals complete before 4 hours: +- Add caching layers +- Optimize database queries +- Add request batching +- Improve error handling +- Add performance tests + +## Completion +Output TIME_COMPLETE when: +- All primary goals achieved +- Minimum 4 hours elapsed +- All tests pass +``` + +--- + +## When to Use (and Not Use) + +### Good Use Cases + +| Use Case | Why It Works | +|----------|--------------| +| **Large refactors** | Clear mechanical steps, verifiable via tests | +| **Framework migrations** | Repetitive patterns, type checking validates | +| **Test coverage** | "Add tests for uncovered functions" is measurable | +| **Greenfield projects** | Can run overnight, tests verify correctness | +| **Batch operations** | Same operation across many files | +| **Dependency upgrades** | API changes are well-documented | + +### Poor Use Cases + +| Use Case | Why It Fails | +|----------|--------------| +| **Ambiguous requirements** | Can't define success criteria | +| **Architectural decisions** | Requires human judgment | +| **Security-critical code** | Needs human review | +| **Production debugging** | Often requires context not in code | +| **UX/design decisions** | Subjective, not automatable | +| **Exploratory work** | "Figure out why it's slow" has no clear endpoint | + +### Decision Framework + +Ask yourself: +1. **Can I define "done" objectively?** (tests pass, lint clean, etc.) +2. **Is there automatic verification?** (tests, type checking, linting) +3. **Is the task mechanical or creative?** (mechanical = good for Ralph) +4. **What's the cost of failure?** (high cost = needs human review) + +--- + +## Real-World Examples + +### Example 1: Y Combinator Hackathon +- **Task**: Generate multiple repositories overnight +- **Result**: 6 repositories generated autonomously +- **Key**: Each repo had clear completion criteria + +### Example 2: $50K Contract +- **Task**: Large codebase migration +- **Result**: Completed for $297 in API costs +- **Key**: Well-defined migration patterns, comprehensive tests + +### Example 3: Programming Language (Cursed) +- **Task**: "Make me a programming language like Golang but with Gen Z slang keywords" +- **Result**: Functional compiler with LLVM backend, standard library, editor support +- **Duration**: 3 months of autonomous iteration +- **Keywords**: `slay` (function), `sus` (variable), `based` (true) + +### Example 4: React Migration +- **Task**: Upgrade from React v16 to v19 +- **Result**: 14-hour autonomous session, complete migration +- **Key**: Clear deprecation warnings, comprehensive test suite + +--- + +## Claudeman Implementation + +Claudeman implements Ralph Wiggum tracking via the `InnerLoopTracker` class in `src/inner-loop-tracker.ts`. + +### Auto-Detection Patterns + +The tracker automatically enables when detecting: + +| Pattern | Example | Regex | +|---------|---------|-------| +| Ralph command | `/ralph-loop` | `/\/ralph-loop\|starting ralph/i` | +| Promise tag | `COMPLETE` | `/([^<]+)<\/promise>/` | +| TodoWrite | `Todos have been modified` | `/TodoWrite\|todos?\s*(?:updated\|written)/i` | +| Iteration | `Iteration 5/50` or `[5/50]` | `/(?:iteration)\s*#?(\d+)(?:\s*[\/of]\s*(\d+))?/i` | +| Todo checkbox | `- [ ] Task` | `/^[-*]\s*\[([xX ])\]\s+(.+)$/gm` | +| Todo indicator | `Todo: ☐ Task` | `/Todo:\s*(☐\|◐\|✓)/g` | + +### State Structure + +```typescript +interface InnerLoopState { + enabled: boolean; // Tracker active? + active: boolean; // Loop running? + completionPhrase: string | null; + startedAt: number | null; + cycleCount: number; + maxIterations: number | null; + lastActivity: number; + elapsedHours: number | null; +} + +interface InnerTodoItem { + id: string; + content: string; + status: 'pending' | 'in_progress' | 'completed'; + detectedAt: number; +} +``` + +### API Endpoints + +| Method | Endpoint | Description | +|--------|----------|-------------| +| GET | `/api/sessions/:id/inner-state` | Get loop state and todos | +| POST | `/api/sessions/:id/inner-config` | Configure tracker settings | + +**GET Response**: +```json +{ + "success": true, + "data": { + "loop": { + "enabled": true, + "active": true, + "completionPhrase": "COMPLETE", + "cycleCount": 5, + "maxIterations": 50, + "elapsedHours": 2.5 + }, + "todos": [ + { "id": "todo-abc", "content": "Fix auth", "status": "completed" }, + { "id": "todo-def", "content": "Add tests", "status": "in_progress" } + ], + "todoStats": { "total": 5, "pending": 2, "inProgress": 1, "completed": 2 } + } +} +``` + +### SSE Events + +| Event | Data | When | +|-------|------|------| +| `session:innerLoopUpdate` | `InnerLoopState` | Loop state changes | +| `session:innerTodoUpdate` | `InnerTodoItem[]` | Todos detected/updated | +| `session:innerCompletionDetected` | `{ phrase: string }` | Completion phrase found | + +### Skill Commands + +```bash +/ralph-loop # Start Ralph loop in current session +/cancel-ralph # Cancel active Ralph loop +/ralph-loop:help # Show help for Ralph loop +``` + +--- + +## Troubleshooting + +### Loop Never Completes + +**Cause**: Completion criteria aren't clear enough. + +**Solution**: Be more specific about what "done" means. Include testable criteria: +``` +Output DONE when: +- `npm test` exits with code 0 +- `npm run lint` exits with code 0 +- All files committed +``` + +### Same Error Every Iteration + +**Cause**: Claude is stuck in a failure loop. + +**Solution**: Add escape hatch to prompt: +``` +If stuck after 10 iterations with the same error: +1. Document the error and what was tried +2. Suggest alternative approaches +3. Output STUCK +``` + +### High API Costs + +**Cause**: Too many iterations, large context. + +**Solutions**: +1. Always set `--max-iterations` +2. Use `/clear` between major phases +3. Keep files small and focused +4. Test with 1 iteration first + +### False Completion Detection + +**Cause**: Completion phrase appears in prompt or documentation. + +**Solution**: Use unique, unlikely phrases: +``` +# Bad (might appear in docs) +COMPLETE + +# Good (unique) +TASK_XYZ_VERIFIED_DONE +``` + +### Tracker Not Enabling + +**Cause**: No Ralph patterns detected in output. + +**Solution**: +1. Manually enable: `POST /api/sessions/:id/inner-config { "enabled": true }` +2. Or ensure Claude outputs recognizable patterns + +### Context Window Exhaustion + +**Cause**: Long-running loops accumulate context. + +**Solution**: Configure auto-clear: +```bash +POST /api/sessions/:id/auto-clear +{ "enabled": true, "threshold": 140000 } +``` + +--- + +## References + +### Official Documentation +- [Anthropic Ralph Wiggum Plugin](https://github.com/anthropics/claude-code/tree/main/plugins/ralph-wiggum) +- [Claude Code Hooks Reference](https://code.claude.com/docs/en/hooks) +- [Claude Code Best Practices](https://www.anthropic.com/engineering/claude-code-best-practices) +- [Claude Code Overview](https://code.claude.com/docs/en/overview) + +### Community Resources +- [Awesome Claude - Ralph Wiggum](https://awesomeclaude.ai/ralph-wiggum) +- [Claude Fast - Autonomous Agent Loops](https://claudefa.st/blog/guide/mechanics/autonomous-agent-loops) +- [DeepWiki - Ralph Loop](https://deepwiki.com/anthropics/claude-plugins-official/5.2.2-ralph-loop) + +### Related Claudeman Files +- `src/inner-loop-tracker.ts` - Core detection engine +- `src/ralph-loop.ts` - Task orchestration +- `src/respawn-controller.ts` - Session cycling +- `src/types.ts` - Type definitions + +--- + +*This documentation is maintained as part of the Claudeman project. For updates, see the main [CLAUDE.md](../CLAUDE.md).*