mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
docs: update completion detection documentation
- Update "All complete" detection behavior in CLAUDE.md and ralph-wiggum-guide.md - Add UI Behaviors section documenting auto-focus, scroll preservation, consistent icons Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
@@ -38,29 +38,35 @@ npm run test:coverage # With coverage report
|
||||
npx vitest run test/session.test.ts # Single file
|
||||
npx vitest run -t "should create session" # By pattern
|
||||
|
||||
# Test port allocation (add new tests in next available range):
|
||||
# 3099-3101: quick-start.test.ts
|
||||
# Test port allocation (integration tests spawn servers):
|
||||
# 3099: quick-start.test.ts
|
||||
# 3102: session.test.ts
|
||||
# 3105-3106: scheduled-runs.test.ts
|
||||
# 3107-3108: sse-events.test.ts
|
||||
# 3110-3112: edge-cases.test.ts
|
||||
# 3115-3116: integration-flows.test.ts
|
||||
# 3120-3121: session-cleanup.test.ts
|
||||
# (no port): respawn-controller.test.ts, inner-loop-tracker.test.ts, pty-interactive.test.ts (unit tests)
|
||||
# 3105: scheduled-runs.test.ts
|
||||
# 3107: sse-events.test.ts
|
||||
# 3110: edge-cases.test.ts
|
||||
# 3115: integration-flows.test.ts
|
||||
# 3120: session-cleanup.test.ts
|
||||
# Unit tests (no port needed): respawn-controller, inner-loop-tracker, pty-interactive
|
||||
# Next available: 3122+
|
||||
|
||||
# Tests mock PTY - no real Claude CLI spawned
|
||||
# Test timeout: 30s (configured in vitest.config.ts)
|
||||
|
||||
# TypeScript checking (no linter configured)
|
||||
# TypeScript checking
|
||||
npx tsc --noEmit # Type check without building
|
||||
# Note: No ESLint/Prettier configured - rely on TypeScript strict mode
|
||||
|
||||
# Debugging
|
||||
screen -ls # List GNU screen sessions
|
||||
screen -r <name> # Attach to screen session
|
||||
screen -r <name> # Attach to screen session (Ctrl+A D to detach)
|
||||
curl localhost:3000/api/sessions # Check active sessions
|
||||
curl localhost:3000/api/status | jq . # Full app state including respawn
|
||||
cat ~/.claudeman/state.json | jq . # View main state
|
||||
cat ~/.claudeman/state-inner.json | jq . # View inner loop state
|
||||
|
||||
# Kill stuck screen sessions
|
||||
screen -X -S <name> quit # Graceful quit
|
||||
pkill -f "SCREEN.*claudeman" # Force kill all claudeman screens
|
||||
```
|
||||
|
||||
## Architecture
|
||||
@@ -74,7 +80,8 @@ cat ~/.claudeman/state-inner.json | jq . # View inner loop state
|
||||
| `src/screen-manager.ts` | GNU screen persistence, ghost discovery, 4-strategy kill |
|
||||
| `src/inner-loop-tracker.ts` | Detects `<promise>PHRASE</promise>`, todos, loop status in output |
|
||||
| `src/task-tracker.ts` | Parses background task output (agent IDs, status) from Claude CLI |
|
||||
| `src/state-store.ts` | JSON persistence to `~/.claudeman/` with debounced (100ms) writes |
|
||||
| `src/session-manager.ts` | Manages session lifecycle, task assignment, and cleanup |
|
||||
| `src/state-store.ts` | JSON persistence to `~/.claudeman/` with debounced writes |
|
||||
| `src/web/server.ts` | Fastify REST API + SSE at `/api/events` |
|
||||
| `src/web/public/app.js` | Frontend: SSE handling, xterm.js, tab management |
|
||||
| `src/types.ts` | All TypeScript interfaces |
|
||||
@@ -89,12 +96,21 @@ cat ~/.claudeman/state-inner.json | jq . # View inner loop state
|
||||
### Respawn State Machine
|
||||
|
||||
```
|
||||
WATCHING → SENDING_UPDATE → WAITING_UPDATE → SENDING_CLEAR → WAITING_CLEAR → SENDING_INIT → WAITING_INIT → WATCHING
|
||||
↑ |
|
||||
└──────────────────────────────────────────────────────────────────────────────────────────────────────────┘
|
||||
┌─────────────────────────────────────────────────────────────────────────────────────────────────┐
|
||||
│ │
|
||||
▼ │
|
||||
WATCHING → SENDING_UPDATE → WAITING_UPDATE → SENDING_CLEAR → WAITING_CLEAR │
|
||||
│ │
|
||||
▼ │
|
||||
SENDING_INIT → WAITING_INIT → MONITORING_INIT ──┬──────────────────┘
|
||||
│
|
||||
▼ (if no work triggered)
|
||||
SENDING_KICKSTART → WAITING_KICKSTART
|
||||
```
|
||||
|
||||
Steps can be skipped via config (`sendClear: false`, `sendInit: false`). Idle detection triggers state transitions.
|
||||
**States**: `watching`, `sending_update`, `waiting_update`, `sending_clear`, `waiting_clear`, `sending_init`, `waiting_init`, `monitoring_init`, `sending_kickstart`, `waiting_kickstart`, `stopped`
|
||||
|
||||
Steps can be skipped via config (`sendClear: false`, `sendInit: false`). Optional `kickstartPrompt` triggers if `/init` doesn't start work. Idle detection triggers state transitions.
|
||||
|
||||
### Session Modes
|
||||
|
||||
@@ -169,7 +185,7 @@ session.writeViaScreen('/init\r');
|
||||
2. Sends text first: `screen -S name -p 0 -X stuff "text"`
|
||||
3. Sends Enter separately: `screen -S name -p 0 -X stuff "$(printf '\015')"`
|
||||
|
||||
**Why separate commands?** Claude CLI uses Ink (React for terminals) which requires text and Enter as separate `screen -X stuff` commands. Combining them doesn't work.
|
||||
**Why separate commands?** Claude CLI uses [Ink](https://github.com/vadimdemedes/ink) (React for terminals) which requires text and Enter as separate `screen -X stuff` commands. Combining them doesn't work. This is a critical implementation detail when debugging input issues.
|
||||
|
||||
#### API Usage
|
||||
```bash
|
||||
@@ -208,15 +224,38 @@ Both wait for idle. Configure via `session.setAutoCompact()` / `session.setAutoC
|
||||
### Inner Loop Tracking
|
||||
|
||||
Detects Ralph loops and todos inside Claude sessions. **Disabled by default** but auto-enables when any of these patterns are detected in terminal output:
|
||||
- `/ralph-loop` command
|
||||
- `/ralph-loop:ralph-loop` command
|
||||
- `<promise>PHRASE</promise>` completion phrases
|
||||
- `TodoWrite` tool usage
|
||||
- Iteration patterns (`Iteration 5/50`, `[5/50]`)
|
||||
- Todo checkboxes (`- [ ]`/`- [x]`) or indicator icons (`☐`/`◐`/`✓`)
|
||||
- "All tasks complete" messages
|
||||
- Individual task completion signals (`Task 8 is done`)
|
||||
|
||||
See `inner-loop-tracker.ts:shouldAutoEnable()` for detection logic.
|
||||
|
||||
API: `GET /api/sessions/:id/inner-state`. UI: collapsible panel below tabs. Use `tracker.enable()` / `tracker.disable()` for programmatic control, or `POST /api/sessions/:id/inner-config` with `{ enabled: boolean }` via API.
|
||||
**Completion Detection**: Uses multi-strategy detection:
|
||||
- 1st occurrence of `<promise>PHRASE</promise>`: Stores as expected phrase (likely in prompt)
|
||||
- 2nd occurrence: Emits `completionDetected` event (actual completion)
|
||||
- **Bare phrase detection**: Also detects phrase without tags once expected phrase is known
|
||||
- **All complete detection**: When "All X files/tasks created/completed" detected, marks all todos complete and emits completion
|
||||
- If loop is already active (via `/ralph-loop:ralph-loop`): Emits immediately on first occurrence
|
||||
|
||||
**Session Lifecycle**: Each session has its own independent tracker:
|
||||
- New session → Fresh tracker (no carryover)
|
||||
- Close tab → Tracker state cleared, UI panel hides
|
||||
- Switch tabs → Panel shows tracker for active session
|
||||
- `tracker.reset()` → Clears todos/state, keeps enabled status
|
||||
- `tracker.fullReset()` → Complete reset to initial state
|
||||
|
||||
**API**:
|
||||
- `GET /api/sessions/:id/inner-state` - Get loop state and todos
|
||||
- `POST /api/sessions/:id/inner-config` - Configure tracker:
|
||||
- `{ enabled: boolean }` - Enable/disable
|
||||
- `{ reset: true }` - Soft reset (keep enabled)
|
||||
- `{ reset: "full" }` - Full reset
|
||||
|
||||
UI: Collapsible panel below tabs, shows progress ring and todo list.
|
||||
|
||||
### Terminal Display Fix
|
||||
|
||||
@@ -248,7 +287,17 @@ Vanilla JS + xterm.js. Key functions:
|
||||
|
||||
### State Store
|
||||
|
||||
Writes debounced (100ms) to `~/.claudeman/state.json`. Batches rapid changes.
|
||||
Writes debounced to `~/.claudeman/state.json`. Batches rapid changes.
|
||||
|
||||
### Timing Constants
|
||||
|
||||
| Constant | Value | Location |
|
||||
|----------|-------|----------|
|
||||
| State save debounce | 500ms | `state-store.ts` |
|
||||
| Line buffer flush | 100ms | `session.ts` |
|
||||
| Terminal batch interval | 16ms | `server.ts` (60fps) |
|
||||
| Idle activity timeout | 2s | `session.ts` |
|
||||
| Respawn idle timeout | 5s default | `RespawnConfig.idleTimeoutMs` |
|
||||
|
||||
### TypeScript Config
|
||||
|
||||
@@ -259,10 +308,24 @@ Module resolution: NodeNext. Target: ES2022. Strict mode enabled. See `tsconfig.
|
||||
- **API endpoint**: Add types in `types.ts`, route in `server.ts:buildServer()`, use `createErrorResponse()` for errors
|
||||
- **SSE event**: Emit via `broadcast()` in server.ts, handle in `app.js:handleSSEEvent()` switch
|
||||
- **Session event**: Add to `SessionEvents` interface in `session.ts`, emit via `this.emit()`, subscribe in server.ts, handle in frontend
|
||||
- **New test file**: Create `test/<name>.test.ts`, pick unique port (next available: 3122+), add to port allocation comment above
|
||||
|
||||
### API Error Codes
|
||||
|
||||
Use `createErrorResponse(code, details?)` from `types.ts`:
|
||||
|
||||
| Code | Use Case |
|
||||
|------|----------|
|
||||
| `NOT_FOUND` | Session/resource doesn't exist |
|
||||
| `INVALID_INPUT` | Bad request parameters |
|
||||
| `SESSION_BUSY` | Session is currently processing |
|
||||
| `OPERATION_FAILED` | Action couldn't complete |
|
||||
| `ALREADY_EXISTS` | Duplicate resource |
|
||||
| `INTERNAL_ERROR` | Unexpected server error |
|
||||
|
||||
## Session Lifecycle & Cleanup
|
||||
|
||||
- **Limit**: `MAX_CONCURRENT_SESSIONS = 50`
|
||||
- **Limit**: Web server: `MAX_CONCURRENT_SESSIONS = 50` (`server.ts:56`), CLI default: 5 (`types.ts:DEFAULT_CONFIG`)
|
||||
- **Kill** (`killScreen()`): child PIDs → process group → screen quit → SIGKILL
|
||||
- **Ghost discovery**: `reconcileScreens()` finds orphaned screens on startup
|
||||
- **Cleanup** (`cleanupSession()`): stops respawn, clears buffers/timers, kills screen
|
||||
@@ -339,6 +402,12 @@ claudeman status # Overall status
|
||||
| `Ctrl+?` | Show keyboard shortcuts help |
|
||||
| `Escape` | Close panels and modals |
|
||||
|
||||
## UI Behaviors
|
||||
|
||||
- **Auto-focus**: When only one session exists and none is active, it auto-selects
|
||||
- **Scroll preservation**: Expanding/collapsing Ralph panel preserves terminal scroll position
|
||||
- **Consistent icons**: Ralph and Monitor panels use same detach/attach icons (⧉/⊞)
|
||||
|
||||
## State Files
|
||||
|
||||
| File | Purpose |
|
||||
@@ -364,10 +433,11 @@ Ralph Wiggum is an autonomous loop technique that lets Claude work iteratively u
|
||||
|
||||
**Core Pattern**: `<promise>PHRASE</promise>` - The completion signal that tells the loop to stop.
|
||||
|
||||
**Official Plugin Commands**:
|
||||
**Skill Commands**:
|
||||
```bash
|
||||
/ralph-loop "<prompt>" --max-iterations 50 --completion-promise "COMPLETE"
|
||||
/cancel-ralph
|
||||
/ralph-loop:ralph-loop # Start Ralph Loop in current session
|
||||
/ralph-loop:cancel-ralph # Cancel active Ralph Loop
|
||||
/ralph-loop:help # Show help and usage
|
||||
```
|
||||
|
||||
**Best Practices** (see full guide for details):
|
||||
|
||||
+83
-33
@@ -73,33 +73,33 @@ while :; do cat PROMPT.md | claude ; done
|
||||
|
||||
### Commands
|
||||
|
||||
#### `/ralph-loop`
|
||||
#### `/ralph-loop:ralph-loop`
|
||||
|
||||
Start an autonomous loop.
|
||||
Start an autonomous loop in the current session.
|
||||
|
||||
```bash
|
||||
/ralph-loop "<prompt>" --max-iterations <n> --completion-promise "<text>"
|
||||
/ralph-loop:ralph-loop
|
||||
```
|
||||
|
||||
**Parameters**:
|
||||
When invoked, this skill prompts you to configure:
|
||||
- **Task prompt**: The work to be done (persists across iterations)
|
||||
- **Max iterations**: Safety limit on iterations (recommended: always set this)
|
||||
- **Completion promise**: The phrase that signals completion (e.g., `COMPLETE`)
|
||||
|
||||
| Parameter | Type | Default | Description |
|
||||
|-----------|------|---------|-------------|
|
||||
| `<prompt>` | string | required | Task description (persists across all iterations) |
|
||||
| `--max-iterations` | integer | unlimited | Safety limit on iterations |
|
||||
| `--completion-promise` | string | none | Exact string that signals completion |
|
||||
|
||||
**Example**:
|
||||
```bash
|
||||
/ralph-loop "Build a REST API for todos. Requirements: CRUD operations, input validation, tests. Output <promise>COMPLETE</promise> when done." --completion-promise "COMPLETE" --max-iterations 50
|
||||
```
|
||||
|
||||
#### `/cancel-ralph`
|
||||
#### `/ralph-loop:cancel-ralph`
|
||||
|
||||
Cancel the active Ralph loop.
|
||||
|
||||
```bash
|
||||
/cancel-ralph
|
||||
/ralph-loop:cancel-ralph
|
||||
```
|
||||
|
||||
#### `/ralph-loop:help`
|
||||
|
||||
Show help and usage information.
|
||||
|
||||
```bash
|
||||
/ralph-loop:help
|
||||
```
|
||||
|
||||
### State File
|
||||
@@ -153,24 +153,38 @@ The official implementation (and Claudeman) prevents false positives when comple
|
||||
- Documentation or examples
|
||||
- Comments
|
||||
|
||||
**Solution**: Only mark as complete if the loop was already `active` when the phrase is detected.
|
||||
**Solution**: Claudeman uses **occurrence-based detection** to distinguish prompts from actual completions:
|
||||
- **1st occurrence**: Store as expected phrase (likely in the prompt)
|
||||
- **2nd occurrence**: Emit `completionDetected` (actual completion)
|
||||
- **If loop already active**: Emit immediately (explicit loop start via `/ralph-loop:ralph-loop`)
|
||||
|
||||
```typescript
|
||||
// From claudeman/src/inner-loop-tracker.ts
|
||||
if (!this._loopState.active) {
|
||||
// Just record the expected completion phrase without marking as complete
|
||||
private handleCompletionPhrase(phrase: string): void {
|
||||
const count = (this._completionPhraseCount.get(phrase) || 0) + 1;
|
||||
this._completionPhraseCount.set(phrase, count);
|
||||
|
||||
// Store phrase on first occurrence
|
||||
if (!this._loopState.completionPhrase) {
|
||||
this._loopState.completionPhrase = phrase;
|
||||
this._loopState.lastActivity = Date.now();
|
||||
this.emit('loopUpdate', this.loopState);
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
// Loop was active, this is a real completion
|
||||
this._loopState.completionPhrase = phrase;
|
||||
this._loopState.active = false;
|
||||
this.emit('completionDetected', phrase);
|
||||
// Emit completion if loop is active OR this is 2nd+ occurrence
|
||||
if (this._loopState.active || count >= 2) {
|
||||
this._loopState.active = false;
|
||||
this._loopState.lastActivity = Date.now();
|
||||
this.emit('completionDetected', phrase);
|
||||
this.emit('loopUpdate', this.loopState);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This approach handles both scenarios:
|
||||
1. **Explicit loop start**: User runs `/ralph-loop:ralph-loop`, loop is active, first completion phrase triggers
|
||||
2. **Implicit completion**: Phrase appears in prompt (1st), then Claude outputs it on completion (2nd)
|
||||
|
||||
---
|
||||
|
||||
## TodoWrite Tool Integration
|
||||
@@ -319,7 +333,8 @@ For more sophisticated evaluation, use LLM-based hooks:
|
||||
> This cannot be overstated: always set `--max-iterations`. Autonomous loops consume tokens rapidly. A typical 50-iteration loop on a medium-sized codebase can cost $50-100+ in API usage.
|
||||
|
||||
```bash
|
||||
/ralph-loop "..." --max-iterations 30 --completion-promise "DONE"
|
||||
/ralph-loop:ralph-loop
|
||||
# Then configure: max-iterations=30, completion-promise="DONE"
|
||||
```
|
||||
|
||||
### 2. Define Clear, Measurable Success Criteria
|
||||
@@ -416,10 +431,12 @@ This creates recovery points and shows progress in git history.
|
||||
|
||||
```bash
|
||||
# Test with 1 iteration first
|
||||
/ralph-loop "..." --max-iterations 1
|
||||
/ralph-loop:ralph-loop
|
||||
# Configure: max-iterations=1
|
||||
|
||||
# Then run full loop
|
||||
/ralph-loop "..." --max-iterations 50
|
||||
/ralph-loop:ralph-loop
|
||||
# Configure: max-iterations=50
|
||||
```
|
||||
|
||||
### 8. Use Git for Safety
|
||||
@@ -607,12 +624,35 @@ The tracker automatically enables when detecting:
|
||||
|
||||
| Pattern | Example | Regex |
|
||||
|---------|---------|-------|
|
||||
| Ralph command | `/ralph-loop` | `/\/ralph-loop\|starting ralph/i` |
|
||||
| Ralph command | `/ralph-loop:ralph-loop` | `/\/ralph-loop\|starting ralph/i` |
|
||||
| Promise tag | `<promise>COMPLETE</promise>` | `/<promise>([^<]+)<\/promise>/` |
|
||||
| TodoWrite | `Todos have been modified` | `/TodoWrite\|todos?\s*(?:updated\|written)/i` |
|
||||
| Iteration | `Iteration 5/50` or `[5/50]` | `/(?:iteration)\s*#?(\d+)(?:\s*[\/of]\s*(\d+))?/i` |
|
||||
| Todo checkbox | `- [ ] Task` | `/^[-*]\s*\[([xX ])\]\s+(.+)$/gm` |
|
||||
| Todo indicator | `Todo: ☐ Task` | `/Todo:\s*(☐\|◐\|✓)/g` |
|
||||
| All complete | `All tasks completed` | `/all\s+tasks?\s+completed?\|all\s+done/i` |
|
||||
| Task done | `Task 8 is done` | `/task\s*#?\d+\s*(?:is\s+)?done/i` |
|
||||
|
||||
### Completion Detection
|
||||
|
||||
Multi-strategy detection to catch various completion signals:
|
||||
|
||||
1. **Tagged phrase**: `<promise>PHRASE</promise>` - First occurrence stores phrase, second triggers completion
|
||||
2. **Bare phrase**: Detects phrase without tags once expected phrase is known (e.g., Claude outputs `COMPLETE` instead of `<promise>COMPLETE</promise>`)
|
||||
3. **All complete signals**: Detects "All X files/tasks created/completed" messages, marks all todos complete and emits completion
|
||||
4. **Explicit task completion**: Matches "Task N is done" patterns
|
||||
|
||||
### Session Lifecycle
|
||||
|
||||
Each session has its **own independent tracker**:
|
||||
|
||||
| Action | Result |
|
||||
|--------|--------|
|
||||
| New session opened | Fresh tracker, no carryover |
|
||||
| Tab closed | Tracker state cleared, UI panel hides |
|
||||
| Switch tabs | Panel shows tracker for active session |
|
||||
| `tracker.reset()` | Clears todos/state, keeps enabled status |
|
||||
| `tracker.fullReset()` | Complete reset to initial state |
|
||||
|
||||
### State Structure
|
||||
|
||||
@@ -643,6 +683,16 @@ interface InnerTodoItem {
|
||||
| GET | `/api/sessions/:id/inner-state` | Get loop state and todos |
|
||||
| POST | `/api/sessions/:id/inner-config` | Configure tracker settings |
|
||||
|
||||
**POST `/inner-config` Options**:
|
||||
```json
|
||||
{
|
||||
"enabled": true, // Enable/disable tracker
|
||||
"reset": true, // Soft reset (clears state, keeps enabled)
|
||||
"reset": "full", // Full reset (clears everything)
|
||||
"completionPhrase": "DONE" // Set expected completion phrase
|
||||
}
|
||||
```
|
||||
|
||||
**GET Response**:
|
||||
```json
|
||||
{
|
||||
@@ -676,9 +726,9 @@ interface InnerTodoItem {
|
||||
### Skill Commands
|
||||
|
||||
```bash
|
||||
/ralph-loop # Start Ralph loop in current session
|
||||
/cancel-ralph # Cancel active Ralph loop
|
||||
/ralph-loop:help # Show help for Ralph loop
|
||||
/ralph-loop:ralph-loop # Start Ralph Loop in current session
|
||||
/ralph-loop:cancel-ralph # Cancel active Ralph Loop
|
||||
/ralph-loop:help # Show help and usage
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
Reference in New Issue
Block a user