Commit Graph
20 Commits
Author SHA1 Message Date
arkonandClaude Opus 4.7 257695ff8e chore: version packages
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-31 05:17:07 +02:00
Tenggan ZhangandTeigen 1b652ceb87 test: repair route harness error rendering + stop AI-checker spawning real processes in tests (#97)
* fix(test): share route error handler with test harness + fix stale assertions

The route test harness built a bare Fastify instance without the production
global error handler (server.ts), so structured errors thrown by route helpers
(findSessionOrFail → 404, parseBody → 400) fell through to Fastify's default
handler — yielding a `{statusCode,error,message}` body instead of the
`{success:false,...}` shape, and the tests asserted the old implicit-200
behavior. 51 route tests across 7 files were red.

- Extract the handler into src/web/route-error-handler.ts; server.ts and the
  test harness now install the identical handler (single source of truth).
- Correct stale assertions across route test files: throw-based error paths
  now assert 404 (unknown session) / 400 (invalid body); genuine in-handler
  `return createErrorResponse(...)` paths (200 + success:false) left untouched.
- Reformat a few test files prettier flagged (pre-existing non-compliance).

Route suite: 307/307 passing (was 256/307). No production behavior change.

* test(respawn): mock child_process so AI checker never spawns real processes

respawn-controller.test.ts drives the AI idle checker (ai-checker-base), whose
runCheck() spawns a real `tmux new-session` running `claude -p`. The AI-enabled
tests only assert the ai_checking state transition (then cancel/stop), so the
spawn produced stray real tmux sessions and claude processes on every run — the
reason `npm test` (full suite) was unsafe to run inside a managed session.

Mock node:child_process here (mirroring ai-idle-checker.test.ts), spreading the
real module so `exec` stays intact for transitively-imported modules
(tmux-manager calls promisify(exec) at load). With this, the full non-mobile
suite runs without spawning any real tmux/claude.

---------

Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
2026-05-25 23:55:57 +02:00
arkonandClaude Opus 4.6 8e5f95d6ea test: add unit tests for Debouncer, CleanupManager, and timer migration refactors
New test files for utilities that previously had no dedicated coverage,
plus migration tests validating the ralph-tracker and respawn-controller
timer refactorings work correctly.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 16:50:23 +01:00
arkonandClaude Opus 4.6 442cc19c3c test: add shared mock infrastructure and route tests (phase 7)
Consolidate duplicated MockSession/MockStateStore into test/mocks/,
migrate respawn tests to shared mocks, and add 58 route tests for
session, system, and respawn endpoints using Fastify app.inject().

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-01 14:33:18 +01:00
arkonandClaude Opus 4.6 fe1f665531 fix: ralph tracker — 3 bugs: first-occurrence completion, double-counting, TodoWrite pattern
Bug 1 (CRITICAL): canonicalCount >= 1 always fired on first <promise> tag
(prompt echo). Changed to >= 2 so only 2nd+ occurrence triggers completion.

Bug 2: checkMultiLinePatterns() re-detected complete tags already handled
by processLine(), double-counting. Now only tries completion when partial
buffer is non-empty (cross-chunk scenario).

Bug 3: TodoWrite ✔ patterns required "Task #N" but real Claude Code output
is plain "✔ content". Added TODO_PLAIN_CHECKMARK_PATTERN fallback.

Includes 71 new deep tests + real-life verification.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-19 14:00:10 +01:00
arkon 88422cab57 chore: bump version to 0.1538 2026-02-18 12:51:35 +01:00
arkonandClaude Opus 4.6 7d2519cacb fix: subagent windows not opening — two cascading bugs
1. handleInit crash: commit cdc822d removed Map initializations for
   teammateTerminals, teammatePanesByName, teams, teamTasks, teammateMap
   from the constructor but left cleanup code that iterates them.
   cleanupAllFloatingWindows() crashed on "not iterable", preventing
   ALL frontend data (sessions, subagents) from loading.

2. claudeSessionId null on recovered sessions: only set inside
   startInteractive(), never in constructor or persisted. After server
   restart, recovered sessions had null claudeSessionId, so the
   hasMatchingTab check always failed → no subagent windows.

Fixes:
- Re-add all 5 missing Map initializations in app.js constructor
- Set _claudeSessionId = this.id in Session constructor (Claudeman
  always passes --session-id to Claude, so they always match)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-13 05:07:51 +01:00
arkon 5b6dd078f1 chore: bump version to 0.1438 2026-01-30 17:20:53 +01:00
arkonandClaude Opus 4.5 1ff1e300f6 feat: add hook-based idle detection for respawn controller
Implement multi-phase improvement to respawn controller's idle detection:

Phase 1 - Stop Hook Detection:
- Add signalStopHook() method - definitive signal when Claude finishes
- 3s confirmation timer to handle race conditions
- Skip AI check when hook received (100% confidence)

Phase 2 - idle_prompt Detection:
- Add signalIdlePrompt() method - fires after 60s+ of idle
- Immediately confirms idle (no confirmation timer needed)

Phase 3 - Transcript File Monitoring:
- New TranscriptWatcher class watches session JSONL files
- Detects: completion, tool execution, plan mode, errors
- Signals respawn controller for supporting detection

Web UI Updates:
- Hook indicator with purple styling and pulse animation
- Shows "Stop hook received" or "idle_prompt hook received"
- 100% confidence displayed with hook-confirmed style

This significantly improves idle detection reliability by using
definitive signals from Claude Code rather than parsing terminal output.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-25 21:29:35 +01:00
arkonandClaude Opus 4.5 40df0748c5 ui: improve respawn controller layout and filter action log noise
- Stack timers row vertically (timers on top, action log below)
- Reduce action log height to 60px since it's now full width
- Filter action log to show only important entries:
  - Commands sent to console
  - Plan-check with action taken
  - Step completions
- Skip timer starts/cancels, detection updates, ai-check status

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-25 19:59:42 +01:00
arkonandClaude Opus 4.5 e56c8626e4 feat: persistent token tracking with global stats aggregation
- Add restoreTokens() method to Session class for recovery after restart
- Add GlobalStats type for cumulative usage tracking across all sessions
- Accumulate tokens from deleted sessions into global stats
- Track lifetime session count with incrementSessionsCreated()
- Add /api/stats endpoint for global stats
- Include globalStats in /api/status response
- Frontend shows aggregate tokens + cost in header
- Fix respawn controller config to filter undefined values
- Add 7 new tests for global stats functionality (v0.1343)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-25 07:26:49 +01:00
arkonandClaude Opus 4.5 ae4c445601 feat: add AI-powered plan mode detection for auto-accept
Adds a two-stage gate before auto-accepting plan mode prompts:
1. Strict regex pre-filter - checks for numbered options + selector
2. AI confirmation - spawns Opus to classify as PLAN_MODE or NOT_PLAN_MODE

This prevents spurious Enter key presses when Claude is paused mid-thought
or experiencing network lag, rather than showing a plan approval prompt.

New files:
- src/ai-plan-checker.ts: AI checker following ai-idle-checker.ts pattern

Config fields added to RespawnConfig:
- aiPlanCheckEnabled (default: true)
- aiPlanCheckModel (default: claude-opus-4-5-20251101)
- aiPlanCheckMaxContext (default: 8000)
- aiPlanCheckTimeoutMs (default: 60000)
- aiPlanCheckCooldownMs (default: 30000)

New events: planCheckStarted, planCheckCompleted, planCheckFailed

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-25 02:55:43 +01:00
arkonandClaude Opus 4.5 0e56e590f3 feat: add AI-powered idle check for respawn controller
Replace the "Worked for Xm Xs" pattern as the sole primary idle
detection signal with a final AI-powered check. When pre-filter
conditions are met (output silence, no working patterns, tokens
stable), a fresh Claude CLI session is spawned in a screen to
analyze terminal output and provide a definitive IDLE/WORKING
verdict before proceeding with the respawn cycle.

New state in state machine: `ai_checking` (between pre-filter
confirmation and `sending_update`). WORKING verdict triggers a
3-minute cooldown. Errors auto-disable after 3 consecutive
failures, falling back to the existing noOutputTimeoutMs safety net.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-24 18:48:37 +01:00
arkonandClaude Opus 4.5 10222075cb fix: restrict auto-accept to plan mode only, block AskUserQuestion prompts
The auto-accept feature was too aggressive - it would press Enter for any
silence without a completion message, including AskUserQuestion prompts.
Now uses the elicitation_dialog notification hook to detect when Claude is
asking a question, and blocks auto-accept in that case. Only plan mode
approvals (silence with no completion message AND no elicitation signal)
trigger auto-accept.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-24 03:10:02 +01:00
arkonandClaude Opus 4.5 0e8c8b8c60 feat: auto-accept plan mode and question prompts in respawn controller
When Claude enters plan mode or asks a question (AskUserQuestion), output
stops without a completion message. The new autoAcceptPrompts feature
detects this state and sends Enter after a configurable delay (default 8s)
to accept the plan or select the default option, keeping Claude working
autonomously.

Enabled by default. Adds UI checkbox in the respawn config panel.
Safety: only fires once per silence period, requires prior output,
and won't fire during active respawn cycles.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-23 01:25:09 +01:00
arkonandClaude Opus 4.5 ec92e5b1cf fix: respawn controller multi-layer detection and cleanup
- Update idle detection from legacy '↵ send' to completion message pattern
  ("for Xm Xs" time patterns like "Worked for 2m 46s")
- Add confirming_idle state for false positive prevention
- Add completionConfirmMs (5s) and noOutputTimeoutMs (30s) config options
- Add multi-layer detection with confidence scoring (0-100%)
- Fix null pointer error in extractTokenCount with guard clause
- Fix respawn controller not stopping on session cleanup (broadcast respawn:stopped)
- Update tests to use new completion message patterns
- Add detection status UI display (confidence level, waiting state)
- Update CLAUDE.md with new detection documentation

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-22 19:24:36 +01:00
arkonandClaude Opus 4.5 34fb33fbc6 test: add global test setup with screen limits and reach 1337 tests
- Add test/setup.ts with max 10 concurrent screen sessions limiter
- Add orphaned Claude/screen process cleanup before/after tests
- Add semaphore-based screen slot acquisition for concurrency control
- Update vitest.config.ts with setupFiles and fileParallelism: false
- Add 16 new test files for comprehensive coverage
- Update README badge to show 1337 total tests
- Update CLAUDE.md with test setup documentation

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-22 11:30:10 +01:00
arkonandClaude Opus 4.5 d741b82966 fix: improve terminal buffer cleaning and test reliability
- Add buffer cleaning to remove junk before Claude banner
- Fix test todo text to avoid pattern conflicts
- Add reset support to inner-config API endpoint

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-21 04:34:45 +01:00
arkonandClaude Opus 4.5 a5b8628e53 refactor: hybrid indicator+timeout detection in respawn controller
- Primary: Use '↵ send' indicator for immediate response when ready
- Fallback: Use timeout-based detection for prompt patterns
- Add working indicators: Synthesizing, Brewing, ✻, ✽
- Cleaner step completion handlers (checkUpdateComplete, etc.)

The hybrid approach responds immediately when Claude shows the
suggestion indicator, while still having timeout fallback for
reliability.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-19 12:32:36 +01:00
arkonandClaude Opus 4.5 c81d014ec1 test: add comprehensive test suite with 129 tests
- Add PTY interactive session tests
- Add SSE event streaming tests
- Add scheduled runs API tests
- Add edge case and error handling tests
- Add integration flow tests for user workflows
- Add respawn controller tests with edge cases
- Add session cleanup and resource management tests

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-18 21:35:34 +01:00