docs: update completion detection documentation

- Update "All complete" detection behavior in CLAUDE.md and ralph-wiggum-guide.md
- Add UI Behaviors section documenting auto-focus, scroll preservation, consistent icons

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
arkon
2026-01-20 17:43:36 +01:00
co-authored by Claude Opus 4.5
parent 2e31db524a
commit 482394e280
2 changed files with 176 additions and 56 deletions
+93 -23
View File
@@ -38,29 +38,35 @@ npm run test:coverage # With coverage report
npx vitest run test/session.test.ts # Single file
npx vitest run -t "should create session" # By pattern
# Test port allocation (add new tests in next available range):
# 3099-3101: quick-start.test.ts
# Test port allocation (integration tests spawn servers):
# 3099: quick-start.test.ts
# 3102: session.test.ts
# 3105-3106: scheduled-runs.test.ts
# 3107-3108: sse-events.test.ts
# 3110-3112: edge-cases.test.ts
# 3115-3116: integration-flows.test.ts
# 3120-3121: session-cleanup.test.ts
# (no port): respawn-controller.test.ts, inner-loop-tracker.test.ts, pty-interactive.test.ts (unit tests)
# 3105: scheduled-runs.test.ts
# 3107: sse-events.test.ts
# 3110: edge-cases.test.ts
# 3115: integration-flows.test.ts
# 3120: session-cleanup.test.ts
# Unit tests (no port needed): respawn-controller, inner-loop-tracker, pty-interactive
# Next available: 3122+
# Tests mock PTY - no real Claude CLI spawned
# Test timeout: 30s (configured in vitest.config.ts)
# TypeScript checking (no linter configured)
# TypeScript checking
npx tsc --noEmit # Type check without building
# Note: No ESLint/Prettier configured - rely on TypeScript strict mode
# Debugging
screen -ls # List GNU screen sessions
screen -r <name> # Attach to screen session
screen -r <name> # Attach to screen session (Ctrl+A D to detach)
curl localhost:3000/api/sessions # Check active sessions
curl localhost:3000/api/status | jq . # Full app state including respawn
cat ~/.claudeman/state.json | jq . # View main state
cat ~/.claudeman/state-inner.json | jq . # View inner loop state
# Kill stuck screen sessions
screen -X -S <name> quit # Graceful quit
pkill -f "SCREEN.*claudeman" # Force kill all claudeman screens
```
## Architecture
@@ -74,7 +80,8 @@ cat ~/.claudeman/state-inner.json | jq . # View inner loop state
| `src/screen-manager.ts` | GNU screen persistence, ghost discovery, 4-strategy kill |
| `src/inner-loop-tracker.ts` | Detects `<promise>PHRASE</promise>`, todos, loop status in output |
| `src/task-tracker.ts` | Parses background task output (agent IDs, status) from Claude CLI |
| `src/state-store.ts` | JSON persistence to `~/.claudeman/` with debounced (100ms) writes |
| `src/session-manager.ts` | Manages session lifecycle, task assignment, and cleanup |
| `src/state-store.ts` | JSON persistence to `~/.claudeman/` with debounced writes |
| `src/web/server.ts` | Fastify REST API + SSE at `/api/events` |
| `src/web/public/app.js` | Frontend: SSE handling, xterm.js, tab management |
| `src/types.ts` | All TypeScript interfaces |
@@ -89,12 +96,21 @@ cat ~/.claudeman/state-inner.json | jq . # View inner loop state
### Respawn State Machine
```
WATCHING → SENDING_UPDATE → WAITING_UPDATE → SENDING_CLEAR → WAITING_CLEAR → SENDING_INIT → WAITING_INIT → WATCHING
↑ |
└──────────────────────────────────────────────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────────────────────────────────────┐
│ │
▼ │
WATCHING → SENDING_UPDATE → WAITING_UPDATE → SENDING_CLEAR → WAITING_CLEAR │
│ │
▼ │
SENDING_INIT → WAITING_INIT → MONITORING_INIT ──┬──────────────────┘
│
▼ (if no work triggered)
SENDING_KICKSTART → WAITING_KICKSTART
```
Steps can be skipped via config (`sendClear: false`, `sendInit: false`). Idle detection triggers state transitions.
**States**: `watching`, `sending_update`, `waiting_update`, `sending_clear`, `waiting_clear`, `sending_init`, `waiting_init`, `monitoring_init`, `sending_kickstart`, `waiting_kickstart`, `stopped`
Steps can be skipped via config (`sendClear: false`, `sendInit: false`). Optional `kickstartPrompt` triggers if `/init` doesn't start work. Idle detection triggers state transitions.
### Session Modes
@@ -169,7 +185,7 @@ session.writeViaScreen('/init\r');
2. Sends text first: `screen -S name -p 0 -X stuff "text"`
3. Sends Enter separately: `screen -S name -p 0 -X stuff "$(printf '\015')"`
**Why separate commands?** Claude CLI uses Ink (React for terminals) which requires text and Enter as separate `screen -X stuff` commands. Combining them doesn't work.
**Why separate commands?** Claude CLI uses [Ink](https://github.com/vadimdemedes/ink) (React for terminals) which requires text and Enter as separate `screen -X stuff` commands. Combining them doesn't work. This is a critical implementation detail when debugging input issues.
#### API Usage
```bash
@@ -208,15 +224,38 @@ Both wait for idle. Configure via `session.setAutoCompact()` / `session.setAutoC
### Inner Loop Tracking
Detects Ralph loops and todos inside Claude sessions. **Disabled by default** but auto-enables when any of these patterns are detected in terminal output:
- `/ralph-loop` command
- `/ralph-loop:ralph-loop` command
- `<promise>PHRASE</promise>` completion phrases
- `TodoWrite` tool usage
- Iteration patterns (`Iteration 5/50`, `[5/50]`)
- Todo checkboxes (`- [ ]`/`- [x]`) or indicator icons (`☐`/`◐`/`✓`)
- "All tasks complete" messages
- Individual task completion signals (`Task 8 is done`)
See `inner-loop-tracker.ts:shouldAutoEnable()` for detection logic.
API: `GET /api/sessions/:id/inner-state`. UI: collapsible panel below tabs. Use `tracker.enable()` / `tracker.disable()` for programmatic control, or `POST /api/sessions/:id/inner-config` with `{ enabled: boolean }` via API.
**Completion Detection**: Uses multi-strategy detection:
- 1st occurrence of `<promise>PHRASE</promise>`: Stores as expected phrase (likely in prompt)
- 2nd occurrence: Emits `completionDetected` event (actual completion)
- **Bare phrase detection**: Also detects phrase without tags once expected phrase is known
- **All complete detection**: When "All X files/tasks created/completed" detected, marks all todos complete and emits completion
- If loop is already active (via `/ralph-loop:ralph-loop`): Emits immediately on first occurrence
**Session Lifecycle**: Each session has its own independent tracker:
- New session → Fresh tracker (no carryover)
- Close tab → Tracker state cleared, UI panel hides
- Switch tabs → Panel shows tracker for active session
- `tracker.reset()` → Clears todos/state, keeps enabled status
- `tracker.fullReset()` → Complete reset to initial state
**API**:
- `GET /api/sessions/:id/inner-state` - Get loop state and todos
- `POST /api/sessions/:id/inner-config` - Configure tracker:
- `{ enabled: boolean }` - Enable/disable
- `{ reset: true }` - Soft reset (keep enabled)
- `{ reset: "full" }` - Full reset
UI: Collapsible panel below tabs, shows progress ring and todo list.
### Terminal Display Fix
@@ -248,7 +287,17 @@ Vanilla JS + xterm.js. Key functions:
### State Store
Writes debounced (100ms) to `~/.claudeman/state.json`. Batches rapid changes.
Writes debounced to `~/.claudeman/state.json`. Batches rapid changes.
### Timing Constants
| Constant | Value | Location |
|----------|-------|----------|
| State save debounce | 500ms | `state-store.ts` |
| Line buffer flush | 100ms | `session.ts` |
| Terminal batch interval | 16ms | `server.ts` (60fps) |
| Idle activity timeout | 2s | `session.ts` |
| Respawn idle timeout | 5s default | `RespawnConfig.idleTimeoutMs` |
### TypeScript Config
@@ -259,10 +308,24 @@ Module resolution: NodeNext. Target: ES2022. Strict mode enabled. See `tsconfig.
- **API endpoint**: Add types in `types.ts`, route in `server.ts:buildServer()`, use `createErrorResponse()` for errors
- **SSE event**: Emit via `broadcast()` in server.ts, handle in `app.js:handleSSEEvent()` switch
- **Session event**: Add to `SessionEvents` interface in `session.ts`, emit via `this.emit()`, subscribe in server.ts, handle in frontend
- **New test file**: Create `test/<name>.test.ts`, pick unique port (next available: 3122+), add to port allocation comment above
### API Error Codes
Use `createErrorResponse(code, details?)` from `types.ts`:
| Code | Use Case |
|------|----------|
| `NOT_FOUND` | Session/resource doesn't exist |
| `INVALID_INPUT` | Bad request parameters |
| `SESSION_BUSY` | Session is currently processing |
| `OPERATION_FAILED` | Action couldn't complete |
| `ALREADY_EXISTS` | Duplicate resource |
| `INTERNAL_ERROR` | Unexpected server error |
## Session Lifecycle & Cleanup
- **Limit**: `MAX_CONCURRENT_SESSIONS = 50`
- **Limit**: Web server: `MAX_CONCURRENT_SESSIONS = 50` (`server.ts:56`), CLI default: 5 (`types.ts:DEFAULT_CONFIG`)
- **Kill** (`killScreen()`): child PIDs → process group → screen quit → SIGKILL
- **Ghost discovery**: `reconcileScreens()` finds orphaned screens on startup
- **Cleanup** (`cleanupSession()`): stops respawn, clears buffers/timers, kills screen
@@ -339,6 +402,12 @@ claudeman status # Overall status
| `Ctrl+?` | Show keyboard shortcuts help |
| `Escape` | Close panels and modals |
## UI Behaviors
- **Auto-focus**: When only one session exists and none is active, it auto-selects
- **Scroll preservation**: Expanding/collapsing Ralph panel preserves terminal scroll position
- **Consistent icons**: Ralph and Monitor panels use same detach/attach icons (⧉/⊞)
## State Files
| File | Purpose |
@@ -364,10 +433,11 @@ Ralph Wiggum is an autonomous loop technique that lets Claude work iteratively u
**Core Pattern**: `<promise>PHRASE</promise>` - The completion signal that tells the loop to stop.
**Official Plugin Commands**:
**Skill Commands**:
```bash
/ralph-loop "<prompt>" --max-iterations 50 --completion-promise "COMPLETE"
/cancel-ralph
/ralph-loop:ralph-loop # Start Ralph Loop in current session
/ralph-loop:cancel-ralph # Cancel active Ralph Loop
/ralph-loop:help # Show help and usage
```
**Best Practices** (see full guide for details):
+83 -33
View File
@@ -73,33 +73,33 @@ while :; do cat PROMPT.md | claude ; done
### Commands
#### `/ralph-loop`
#### `/ralph-loop:ralph-loop`
Start an autonomous loop.
Start an autonomous loop in the current session.
```bash
/ralph-loop "<prompt>" --max-iterations <n> --completion-promise "<text>"
/ralph-loop:ralph-loop
```
**Parameters**:
When invoked, this skill prompts you to configure:
- **Task prompt**: The work to be done (persists across iterations)
- **Max iterations**: Safety limit on iterations (recommended: always set this)
- **Completion promise**: The phrase that signals completion (e.g., `COMPLETE`)
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `<prompt>` | string | required | Task description (persists across all iterations) |
| `--max-iterations` | integer | unlimited | Safety limit on iterations |
| `--completion-promise` | string | none | Exact string that signals completion |
**Example**:
```bash
/ralph-loop "Build a REST API for todos. Requirements: CRUD operations, input validation, tests. Output <promise>COMPLETE</promise> when done." --completion-promise "COMPLETE" --max-iterations 50
```
#### `/cancel-ralph`
#### `/ralph-loop:cancel-ralph`
Cancel the active Ralph loop.
```bash
/cancel-ralph
/ralph-loop:cancel-ralph
```
#### `/ralph-loop:help`
Show help and usage information.
```bash
/ralph-loop:help
```
### State File
@@ -153,24 +153,38 @@ The official implementation (and Claudeman) prevents false positives when comple
- Documentation or examples
- Comments
**Solution**: Only mark as complete if the loop was already `active` when the phrase is detected.
**Solution**: Claudeman uses **occurrence-based detection** to distinguish prompts from actual completions:
- **1st occurrence**: Store as expected phrase (likely in the prompt)
- **2nd occurrence**: Emit `completionDetected` (actual completion)
- **If loop already active**: Emit immediately (explicit loop start via `/ralph-loop:ralph-loop`)
```typescript
// From claudeman/src/inner-loop-tracker.ts
if (!this._loopState.active) {
// Just record the expected completion phrase without marking as complete
private handleCompletionPhrase(phrase: string): void {
const count = (this._completionPhraseCount.get(phrase) || 0) + 1;
this._completionPhraseCount.set(phrase, count);
// Store phrase on first occurrence
if (!this._loopState.completionPhrase) {
this._loopState.completionPhrase = phrase;
this._loopState.lastActivity = Date.now();
this.emit('loopUpdate', this.loopState);
}
return;
}
// Loop was active, this is a real completion
this._loopState.completionPhrase = phrase;
this._loopState.active = false;
this.emit('completionDetected', phrase);
// Emit completion if loop is active OR this is 2nd+ occurrence
if (this._loopState.active || count >= 2) {
this._loopState.active = false;
this._loopState.lastActivity = Date.now();
this.emit('completionDetected', phrase);
this.emit('loopUpdate', this.loopState);
}
}
```
This approach handles both scenarios:
1. **Explicit loop start**: User runs `/ralph-loop:ralph-loop`, loop is active, first completion phrase triggers
2. **Implicit completion**: Phrase appears in prompt (1st), then Claude outputs it on completion (2nd)
---
## TodoWrite Tool Integration
@@ -319,7 +333,8 @@ For more sophisticated evaluation, use LLM-based hooks:
> This cannot be overstated: always set `--max-iterations`. Autonomous loops consume tokens rapidly. A typical 50-iteration loop on a medium-sized codebase can cost $50-100+ in API usage.
```bash
/ralph-loop "..." --max-iterations 30 --completion-promise "DONE"
/ralph-loop:ralph-loop
# Then configure: max-iterations=30, completion-promise="DONE"
```
### 2. Define Clear, Measurable Success Criteria
@@ -416,10 +431,12 @@ This creates recovery points and shows progress in git history.
```bash
# Test with 1 iteration first
/ralph-loop "..." --max-iterations 1
/ralph-loop:ralph-loop
# Configure: max-iterations=1
# Then run full loop
/ralph-loop "..." --max-iterations 50
/ralph-loop:ralph-loop
# Configure: max-iterations=50
```
### 8. Use Git for Safety
@@ -607,12 +624,35 @@ The tracker automatically enables when detecting:
| Pattern | Example | Regex |
|---------|---------|-------|
| Ralph command | `/ralph-loop` | `/\/ralph-loop\|starting ralph/i` |
| Ralph command | `/ralph-loop:ralph-loop` | `/\/ralph-loop\|starting ralph/i` |
| Promise tag | `<promise>COMPLETE</promise>` | `/<promise>([^<]+)<\/promise>/` |
| TodoWrite | `Todos have been modified` | `/TodoWrite\|todos?\s*(?:updated\|written)/i` |
| Iteration | `Iteration 5/50` or `[5/50]` | `/(?:iteration)\s*#?(\d+)(?:\s*[\/of]\s*(\d+))?/i` |
| Todo checkbox | `- [ ] Task` | `/^[-*]\s*\[([xX ])\]\s+(.+)$/gm` |
| Todo indicator | `Todo: ☐ Task` | `/Todo:\s*(☐\|◐\|✓)/g` |
| All complete | `All tasks completed` | `/all\s+tasks?\s+completed?\|all\s+done/i` |
| Task done | `Task 8 is done` | `/task\s*#?\d+\s*(?:is\s+)?done/i` |
### Completion Detection
Multi-strategy detection to catch various completion signals:
1. **Tagged phrase**: `<promise>PHRASE</promise>` - First occurrence stores phrase, second triggers completion
2. **Bare phrase**: Detects phrase without tags once expected phrase is known (e.g., Claude outputs `COMPLETE` instead of `<promise>COMPLETE</promise>`)
3. **All complete signals**: Detects "All X files/tasks created/completed" messages, marks all todos complete and emits completion
4. **Explicit task completion**: Matches "Task N is done" patterns
### Session Lifecycle
Each session has its **own independent tracker**:
| Action | Result |
|--------|--------|
| New session opened | Fresh tracker, no carryover |
| Tab closed | Tracker state cleared, UI panel hides |
| Switch tabs | Panel shows tracker for active session |
| `tracker.reset()` | Clears todos/state, keeps enabled status |
| `tracker.fullReset()` | Complete reset to initial state |
### State Structure
@@ -643,6 +683,16 @@ interface InnerTodoItem {
| GET | `/api/sessions/:id/inner-state` | Get loop state and todos |
| POST | `/api/sessions/:id/inner-config` | Configure tracker settings |
**POST `/inner-config` Options**:
```json
{
"enabled": true, // Enable/disable tracker
"reset": true, // Soft reset (clears state, keeps enabled)
"reset": "full", // Full reset (clears everything)
"completionPhrase": "DONE" // Set expected completion phrase
}
```
**GET Response**:
```json
{
@@ -676,9 +726,9 @@ interface InnerTodoItem {
### Skill Commands
```bash
/ralph-loop # Start Ralph loop in current session
/cancel-ralph # Cancel active Ralph loop
/ralph-loop:help # Show help for Ralph loop
/ralph-loop:ralph-loop # Start Ralph Loop in current session
/ralph-loop:cancel-ralph # Cancel active Ralph Loop
/ralph-loop:help # Show help and usage
```
---