Add comprehensive tests for the cleanup patterns used to prevent
memory leaks in long-running sessions:
- Task description cache with TTL expiration
- Promise callback null-after-rejection pattern
- Event listener tracking and removal
- DOM handler storage for frontend cleanup
- Timer/interval management
- Map cleanup patterns and safe iteration
- WeakRef/WeakMap usage patterns
- Cleanup order verification (LIFO)
Update CLAUDE.md with reference to new test file.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Previously killOrphanedTestScreens() would kill ANY detached claudeman
screen that wasn't in preExistingScreens. This was dangerous because:
- User screens can become temporarily detached (web server reconnect)
- Tests might start before user creates sessions
- Race conditions between screen status and cleanup timing
Now we ONLY kill screens that tests explicitly register via
registerTestScreen(). Orphaned screens are warned about but not killed.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Remove TUI from git (moved to .gitignore for local development)
- Remove React/Ink dependencies (unused in published version)
- Remove tui command from CLI
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Implements the Execution Optimizer Rework plan that enables the system
to actually use the optimizer metadata (parallelGroups, agentType,
recommendedModel, requiresFreshContext, estimatedTokens) that was
previously generated but ignored.
New components:
- ExecutionBridge: Central coordinator that loads optimized plans,
manages parallel execution within groups, and coordinates with
SpawnOrchestrator for session-based execution
- ModelSelector: Routes tasks to appropriate models (opus/sonnet/haiku)
based on user defaults and agent type overrides. Optimizer
recommendations are advisory only - user preferences always win
- GroupScheduler: Builds topologically ordered execution groups,
manages dependencies, determines execution mode (session vs task-tool)
- ContextManager: Handles fresh context requirements via /clear+/init
or new session spawning
Features:
- Parallel task execution within groups (configurable limit)
- Group-level dependency tracking (lower groups complete first)
- Partial failure handling (continue with non-dependent tasks)
- Model configuration in App Settings > Models tab
- Agent type overrides (explore, implement, test, review)
- Execution control API endpoints (start, pause, resume, cancel)
- SSE events for real-time execution progress visibility
- Execution history tracking
API endpoints:
- GET/POST/PUT /api/execution/* for execution control
- GET/PUT /api/execution/model-config for model settings
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
setTimeout isn't exact, so the remaining TTL after 100ms wait may be
slightly more than expected (901ms instead of ≤900ms). Adding 10ms
tolerance to prevent flaky test failures.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Automatically removes entries not accessed within TTL
- Periodic cleanup with configurable interval
- Optional onExpire callback for cleanup notifications
- Refresh TTL on get (configurable)
- Touch, peek, getAge, getRemainingTtl methods
- Full iteration support
- Implements Disposable interface
Useful for caching ephemeral data like pending tool calls, subagent activity.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add Disposable, BufferConfig, MemoryMetrics, CleanupRegistration types
- Create src/config/buffer-limits.ts with consolidated buffer size constants
- Create src/config/map-limits.ts with Map size limits to prevent unbounded growth
- Implement BufferAccumulator utility with configurable trim and onTrim callback
- Implement LRUMap with automatic eviction and O(1) operations
- Implement CleanupManager for unified resource cleanup with isStopped guard
- Add comprehensive tests for all new utilities
This lays the foundation for memory leak prevention and performance improvements.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
The plan generation API (/api/generate-plan-detailed) was incorrectly
detecting client disconnection because it listened to req.raw.on('close')
which fires when the HTTP request body finishes parsing, not when the
actual TCP connection closes.
This caused the error "Failed to parse plan - no JSON array found" because
the server would cancel all subagent sessions almost immediately after
starting them.
Fix:
- Changed from req.raw.on('close') to socket.on('close')
- Added responseSent flag to only cancel if response hasn't been sent
- Added E2E test to verify the fix
Tested with:
- Simple plan generation: 61 items, 138.5s, quality 0.82
- Smartphone app plan: 53 items, 93.4s, quality 0.75
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add /api/sessions/:id/ralph-prompt/write endpoint to write prompt to @ralph_prompt.md
- Wizard now writes full prompt to file, then sends simple read command to Claude
- Fix session readiness check to wait for prompt character instead of notWorking flag
- Add E2E test infrastructure for ralph-loop workflow (port 3190)
- Add ralph-wizard-prod.mjs script for production testing
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- restoreSubagentWindowStates now discovers parent sessions before restoring
- closeSubagentWindow discovers parent before minimizing to prevent wrong tab
- Fixed async handling in subagent:completed event handler
chore: bump version to 0.1389
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add 300-char rolling window to catch working patterns split across PTY chunks
- Check completion message BEFORE working patterns (priority fix)
- Clear rolling window on completion message (transition point)
- Increase working pattern absence threshold from 3s to 8s
- Add Session.isWorking safety check before confirming idle
- Add 20+ more working patterns (Compiling, Building, Processing, etc.)
- Make AI idle checker prompt more conservative (err toward WORKING)
- Update documentation
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Implement multi-phase improvement to respawn controller's idle detection:
Phase 1 - Stop Hook Detection:
- Add signalStopHook() method - definitive signal when Claude finishes
- 3s confirmation timer to handle race conditions
- Skip AI check when hook received (100% confidence)
Phase 2 - idle_prompt Detection:
- Add signalIdlePrompt() method - fires after 60s+ of idle
- Immediately confirms idle (no confirmation timer needed)
Phase 3 - Transcript File Monitoring:
- New TranscriptWatcher class watches session JSONL files
- Detects: completion, tool execution, plan mode, errors
- Signals respawn controller for supporting detection
Web UI Updates:
- Hook indicator with purple styling and pulse animation
- Shows "Stop hook received" or "idle_prompt hook received"
- 100% confidence displayed with hook-confirmed style
This significantly improves idle detection reliability by using
definitive signals from Claude Code rather than parsing terminal output.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Stack timers row vertically (timers on top, action log below)
- Reduce action log height to 60px since it's now full width
- Filter action log to show only important entries:
- Commands sent to console
- Plan-check with action taken
- Step completions
- Skip timer starts/cancels, detection updates, ai-check status
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add 100ms debounce for new file detection to allow content to be written
- Add subagent:updated event for retroactive description updates
- Extract description in processEntry when first user message is processed
- Add extractDescriptionFromFile helper with retry in file change handler
- Update SubagentTranscriptEntry.content type to support string | array
- Add test coverage for subagent:updated event
Fixes 43% failure rate where subagents displayed raw IDs instead of descriptions
due to race condition when files were discovered before first line was written.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
The ai-plan-checker.ts was passing the prompt directly as a shell argument,
which can cause E2BIG errors when the terminal buffer is large (8KB+).
This fix applies the same temp file approach already used in ai-idle-checker.ts:
- Write prompt to a temp file instead of passing as shell argument
- Pipe the file to claude via stdin: `cat prompt.txt | claude -p ...`
- Clean up prompt file after check completes
Also fixes a race condition in the cancel() method of both AI checkers where
the poll timer could fire between setting checkCancelled and clearing timers.
Now timers are cleared before resolving the promise to prevent this race.
Includes test utilities and analysis documents for the respawn controller
created by other agents:
- test/respawn-test-utils.ts - MockSession, MockAiIdleChecker utilities
- test/respawn-analysis.md - Code analysis and issue identification
- test/respawn-scenarios.md - Test scenario documentation
- test/respawn-test-plan.md - Testing architecture documentation
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add restoreTokens() method to Session class for recovery after restart
- Add GlobalStats type for cumulative usage tracking across all sessions
- Accumulate tokens from deleted sessions into global stats
- Track lifetime session count with incrementSessionsCreated()
- Add /api/stats endpoint for global stats
- Include globalStats in /api/status response
- Frontend shows aggregate tokens + cost in header
- Fix respawn controller config to filter undefined values
- Add 7 new tests for global stats functionality (v0.1343)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Adds a two-stage gate before auto-accepting plan mode prompts:
1. Strict regex pre-filter - checks for numbered options + selector
2. AI confirmation - spawns Opus to classify as PLAN_MODE or NOT_PLAN_MODE
This prevents spurious Enter key presses when Claude is paused mid-thought
or experiencing network lag, rather than showing a plan approval prompt.
New files:
- src/ai-plan-checker.ts: AI checker following ai-idle-checker.ts pattern
Config fields added to RespawnConfig:
- aiPlanCheckEnabled (default: true)
- aiPlanCheckModel (default: claude-opus-4-5-20251101)
- aiPlanCheckMaxContext (default: 8000)
- aiPlanCheckTimeoutMs (default: 60000)
- aiPlanCheckCooldownMs (default: 30000)
New events: planCheckStarted, planCheckCompleted, planCheckFailed
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Replace the "Worked for Xm Xs" pattern as the sole primary idle
detection signal with a final AI-powered check. When pre-filter
conditions are met (output silence, no working patterns, tokens
stable), a fresh Claude CLI session is spawned in a screen to
analyze terminal output and provide a definitive IDLE/WORKING
verdict before proceeding with the respawn cycle.
New state in state machine: `ai_checking` (between pre-filter
confirmation and `sending_update`). WORKING verdict triggers a
3-minute cooldown. Errors auto-disable after 3 consecutive
failures, falling back to the existing noOutputTimeoutMs safety net.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- ralph-tracker: add ✓ to native todo pattern + pre-check, add configure() method
- session: remove stale message expectation from interactive endpoint test
- buffer-management: fix trim count expectations (1202 items triggers second trim)
- cli-commands: replace non-existent toStartWith with toMatch regex
- edge-cases: update error expectation for session-not-found on respawn config PUT
- ralph-integration: respawn config PUT without controller now saves as pre-config
- session-state: fix debouncer shouldFlush(0) - pass timestamp >= delayMs
- timing-utilities: attach catch handlers before advancing fake timers
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
The hook commands now read stdin JSON from Claude Code (contains tool_name,
tool_input, etc.) and forward it as the data field to the API. This enables
richer notifications showing actual context (e.g., "Bash: docker push prod").
Previously the curl commands only sent event type and session ID, losing
all hook context data.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
The auto-accept feature was too aggressive - it would press Enter for any
silence without a completion message, including AskUserQuestion prompts.
Now uses the elicitation_dialog notification hook to detect when Claude is
asking a question, and blocks auto-accept in that case. Only plan mode
approvals (silence with no completion message AND no elicitation signal)
trigger auto-accept.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Wire Claude Code's official hooks system (Notification, Stop) to POST
back to Claudeman's new /api/hook-event endpoint, which broadcasts SSE
events consumed by the existing NotificationManager for desktop alerts.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Integration tests create screen sessions via the web server, but
server.stop() intentionally preserves them (for reattachment in
production). This left 30+ detached screens after each test run.
Fix by recording pre-existing screens in beforeAll, then killing any
new detached claudeman-* screens in afterAll that weren't there at
test start. Never kills attached sessions (user's active work).
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Instead of parsing terminal output for <spawn1337> tags, spawn capabilities
are now exposed as native MCP tools that Claude Code can call directly.
The MCP server (stdio transport) proxies requests to the existing REST API.
- Add src/mcp-server.ts with 6 tools: spawn_agent, list_agents,
get_agent_status, get_agent_result, send_agent_message, cancel_agent
- Remove src/spawn-detector.ts and all references in session.ts/server.ts
- Add CLAUDEMAN_API_URL env var propagation to sessions and screens
- Write .mcp.json to case directories during creation
- Remove spawn1337 tag documentation from case-template.md
- Add claudeman-mcp bin entry to package.json
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Ralph tracker no longer auto-enables on pattern detection. It must be
explicitly enabled per-session (via API) or globally via the new
`ralphEnabled` AppConfig setting. Adds GET/PUT /api/config endpoints
for runtime configuration. Spawn agent ralph enable is unchanged.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
When Claude enters plan mode or asks a question (AskUserQuestion), output
stops without a completion message. The new autoAcceptPrompts feature
detects this state and sends Enter after a configurable delay (default 8s)
to accept the plan or select the default option, keeping Claude working
autonomously.
Enabled by default. Adds UI checkbox in the respawn config panel.
Safety: only fires once per silence period, requires prior output,
and won't fire during active respawn cycles.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Moves case directory cleanup from afterAll to afterEach in all test
files that create cases. Previously, if the test suite was interrupted
or afterAll timed out, all case directories were left behind. Now each
test cleans up immediately after itself.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Update idle detection from legacy '↵ send' to completion message pattern
("for Xm Xs" time patterns like "Worked for 2m 46s")
- Add confirming_idle state for false positive prevention
- Add completionConfirmMs (5s) and noOutputTimeoutMs (30s) config options
- Add multi-layer detection with confidence scoring (0-100%)
- Fix null pointer error in extractTokenCount with guard clause
- Fix respawn controller not stopping on session cleanup (broadcast respawn:stopped)
- Update tests to use new completion message patterns
- Add detection status UI display (confidence level, waiting state)
- Update CLAUDE.md with new detection documentation
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add test/setup.ts with max 10 concurrent screen sessions limiter
- Add orphaned Claude/screen process cleanup before/after tests
- Add semaphore-based screen slot acquisition for concurrency control
- Update vitest.config.ts with setupFiles and fileParallelism: false
- Add 16 new test files for comprehensive coverage
- Update README badge to show 1337 total tests
- Update CLAUDE.md with test setup documentation
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Rename inner-loop-tracker.ts → ralph-tracker.ts throughout codebase
- Add ralph-config.ts for parsing .claude/ralph-loop.local.md config
- Standardize API error responses using createErrorResponse()
- Add input validation for auto-compact/auto-clear thresholds
- Update UI labels to "Ralph / Todo Tracker" consistently
- Add 46 integration tests for Ralph tracking functionality
- Update test badge to 438 total tests
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
TUI now checks if web server is running on startup and offers to start
it in the background. Added new CLI options: --with-web (auto-start),
--no-web (skip check), -p (port).
TUI feature parity with web interface:
- Shell mode: press 'h' in cases view to start bash instead of Claude
- Multi-start: press 'm' to start 1-20 sessions at once
- Respawn toggle: Ctrl+R to enable/disable respawn on Claude sessions
- Session rename: API support via useSessionManager hook
Security fixes from previous analysis:
- Command injection prevention in screen-manager.ts
- Path traversal protection in server.ts
- Input validation for shell-interpolated values
Also fixes memory leak in session.ts (timer tracking) and flaky test
timeout in session-cleanup.test.ts.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Adds 16 tests covering:
- Default template generation with case name and description
- Date placeholder replacement
- Claudeman environment section inclusion
- Work principles and TodoWrite guidance
- Ralph Wiggum Loop section
- Custom template loading and fallback behavior
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add 10 tests for types.ts helper functions
- Test createErrorResponse with all error codes
- Test createSuccessResponse with and without data
- Test createInitialInnerLoopState defaults
- Test createInitialInnerSessionState structure
- Test createInitialState with Ralph Loop and config
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add 27 tests covering state persistence and retrieval
- Test debounced save behavior with fake timers
- Test session, task, and config CRUD operations
- Test Ralph Loop state management
- Test inner state operations for session tracking
- Test persistence across store instances
- Uses real file system in temp directory for integration tests
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add 29 tests covering session lifecycle management
- Test session creation, stopping, and cleanup
- Test event forwarding (output, error, completion, exit)
- Test max concurrent sessions limit
- Test session state persistence and retrieval
- Mock Session class to isolate unit tests
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add 24 tests covering Ralph Loop lifecycle
- Test start, stop, pause, resume state transitions
- Test elapsed time tracking and min duration
- Test automatic stopping when conditions met
- Test stats reporting
- Use vi.hoisted() for mock state sharing
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add 25 tests covering Task model functionality
- Test constructor options and ID generation
- Test lifecycle methods (assign, complete, fail, reset)
- Test completion detection with promise tags
- Test timeout handling with fake timers
- Test serialization (toDefinition, toState, fromState)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Fix TypeScript error: assignTask() was still using string assignment
- Add extended timeout to session.test.ts afterAll hook to prevent
flaky test failures during cleanup
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
- Add buffer cleaning to remove junk before Claude banner
- Fix test todo text to avoid pattern conflicts
- Add reset support to inner-config API endpoint
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>