feat: add Orchestrator Loop — phased plan execution with team agents

Adds a new autonomous loop that accepts high-level goals, generates
phased execution plans via AI, and executes them step-by-step with
verification gates between phases.

Core components:
- OrchestratorLoop: state machine (idle→planning→approval→executing→verifying→completed)
- OrchestratorPlanner: plan generation via PlanOrchestrator, Kahn's algorithm phase grouping
- OrchestratorVerifier: phase verification (strict/moderate/lenient modes)
- Prompt templates for phase execution, team delegation, verification, replanning

API (10 endpoints):
- POST start/approve/reject/pause/resume/stop
- GET status/plan
- POST phase/:id/skip, phase/:id/retry

Frontend: orchestrator-panel.js with SSE-driven state, phase progress, task tracking

Tests: 22 tests (18 route + 4 unit), all passing. Typecheck/lint/format clean.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
arkon
2026-03-21 07:20:18 +01:00
co-authored by Claude Opus 4.6
parent 497ca4891a
commit 61b5ec095c
27 changed files with 4718 additions and 1 deletions
+367
View File
@@ -0,0 +1,367 @@
# Orchestrator Loop — Architecture & Data Flow
> Technical architecture document. Not for GitHub.
## System Overview
```
┌─────────────────────────────────────────────────────────────────────┐
│ CODEMAN WEB UI │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ Orchestrator Dashboard │ │
│ │ [Goal Input] [Plan View] [Phase Progress] [Agent Activity] │ │
│ └───────────────────────────┬──────────────────────────────────┘ │
│ │ SSE Events │
│ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ Orchestrator API Routes (/api/orchestrator/*) │ │
│ └───────────────────────────┬──────────────────────────────────┘ │
└───────────────────────────────┼─────────────────────────────────────┘
▼
┌─────────────────────────────────────────────────────────────────────┐
│ ORCHESTRATOR LOOP │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────────────┐ │
│ │ Orchestrator │ │ Orchestrator │ │ Orchestrator │ │
│ │ Planner │ │ Loop (state │ │ Verifier │ │
│ │ │ │ machine) │ │ │ │
│ │ • Research │◄──►│ • Phase mgmt │◄──►│ • Test runner │ │
│ │ • Plan gen │ │ • Task queue │ │ • AI review │ │
│ │ • Phasing │ │ • Event loop │ │ • Output checks │ │
│ └──────┬───────┘ └──────┬───────┘ └──────────┬───────────┘ │
│ │ │ │ │
│ ▼ ▼ ▼ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ EXISTING CODEMAN INFRASTRUCTURE │ │
│ │ │ │
│ │ SessionManager ←→ Sessions ←→ PTY (Claude CLI) │ │
│ │ ↑ ↑ ↑ │ │
│ │ │ │ │ │ │
│ │ TaskQueue RalphTracker RespawnController │ │
│ │ StateStore HooksConfig TeamWatcher │ │
│ │ Auto-Ops SubagentWatcher SSE Broadcast │ │
│ └──────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
```
## Data Flow: Complete Lifecycle
### 1. User Submits Goal
```
User → POST /api/orchestrator/start { goal: "Build a REST API...", config: {...} }
→ OrchestratorLoop.start(goal)
→ state = PLANNING
→ emit('stateChanged', 'planning')
→ SSE: orchestrator:stateChanged
```
### 2. Planning Phase
```
OrchestratorPlanner.generatePlan(goal)
→ PlanOrchestrator.generateDetailedPlan(goal)
→ [Research Agent] → enriched task description
→ [Planner Agent] → PlanItem[]
→ groupIntoPhases(planItems)
→ topological sort by dependencies
→ group into layers
→ assign team strategies
→ OrchestratorPlan { phases: [...] }
→ state = APPROVAL
→ emit('planReady', plan)
→ SSE: orchestrator:planReady
```
### 3. User Approves Plan
```
User → POST /api/orchestrator/approve
→ OrchestratorLoop.approvePlan()
→ state = EXECUTING
→ executePhase(phases[0])
```
### 4. Phase Execution
```
executePhase(phase)
→ For each task in phase:
→ Convert to CreateTaskOptions
→ Add to TaskQueue with completion phrase "PHASE_{N}_TASK_{M}_DONE"
→ If phase.teamStrategy.type === 'team':
→ Start session with AGENT_TEAMS enabled
→ Send team orchestration prompt to lead
→ Else:
→ Assign tasks to available sessions (same as RalphLoop)
→ Listen for task completion events:
→ TaskQueue emits taskCompleted
→ Check: all phase tasks done?
→ Yes → state = VERIFYING → verifyPhase(phase)
→ No → wait for more completions
```
### 5. Verification
```
verifyPhase(phase)
→ OrchestratorVerifier.verify(phase, session)
→ Run test commands via session
→ Check file existence
→ AI review (optional)
→ If passed:
→ phase.status = 'passed'
→ emit('phaseCompleted', phase)
→ If more phases: executePhase(nextPhase)
→ If last phase: state = COMPLETED
→ If failed:
→ phase.attempts++
→ If attempts < maxAttempts:
→ state = REPLANNING
→ Generate recovery tasks
→ state = EXECUTING (retry)
→ Else:
→ state = FAILED
→ emit('phaseFailed', phase, reason)
```
### 6. Context Management Between Phases
```
After phase completion:
→ If config.compactBetweenPhases:
→ session.sendInput('/compact')
→ Wait for compact to complete
→ If config.respawnBetweenMilestones && phase is a milestone:
→ Save orchestrator state to StateStore
→ Respawn session (kill + recreate)
→ Send resume prompt with phase context
```
## File Layout
```
src/
├── orchestrator-loop.ts # Main state machine (~400 lines)
├── orchestrator-planner.ts # Plan generation + phase grouping (~300 lines)
├── orchestrator-verifier.ts # Phase verification (~200 lines)
├── types/
│ └── orchestrator.ts # All orchestrator types (~150 lines)
├── prompts/
│ └── orchestrator.ts # Prompt templates (~200 lines)
├── web/
│ ├── routes/
│ │ └── orchestrator-routes.ts # API endpoints (~250 lines)
│ └── public/
│ └── orchestrator-ui.js # Frontend panel (~500 lines)
```
## Integration Points with Existing Code
### StateStore (`src/state-store.ts`)
```typescript
// Add to AppState interface
orchestrator?: OrchestratorPersistState;
// Add methods
getOrchestratorState(): OrchestratorPersistState;
setOrchestratorState(state: Partial<OrchestratorPersistState>): void;
```
### SSE Events (`src/web/sse-events.ts`)
```typescript
// Add ~8 new events
export const SseEvent = {
// ... existing
ORCHESTRATOR_STATE_CHANGED: 'orchestrator:stateChanged',
ORCHESTRATOR_PLAN_READY: 'orchestrator:planReady',
ORCHESTRATOR_PHASE_STARTED: 'orchestrator:phaseStarted',
ORCHESTRATOR_PHASE_COMPLETED: 'orchestrator:phaseCompleted',
ORCHESTRATOR_PHASE_FAILED: 'orchestrator:phaseFailed',
ORCHESTRATOR_VERIFICATION: 'orchestrator:verificationResult',
ORCHESTRATOR_COMPLETED: 'orchestrator:completed',
ORCHESTRATOR_ERROR: 'orchestrator:error',
} as const;
```
### Frontend Constants (`src/web/public/constants.js`)
```javascript
// Mirror SSE events
SSE_EVENTS.ORCHESTRATOR_STATE_CHANGED = 'orchestrator:stateChanged';
// ... etc
```
### Route Registration (`src/web/routes/index.ts`)
```typescript
import { registerOrchestratorRoutes } from './orchestrator-routes.js';
// Add to barrel export
```
### Server (`src/web/server.ts`)
```typescript
// Initialize OrchestratorLoop alongside RalphLoop
const orchestratorLoop = new OrchestratorLoop(config);
// Register routes
registerOrchestratorRoutes(app, { ...ctx, orchestrator: orchestratorLoop });
```
### Port Interface (`src/web/ports/`)
```typescript
// New port
export interface OrchestratorPort {
orchestrator: OrchestratorLoop;
}
```
## Prompt Flow Through System
The key insight is how prompts flow from Orchestrator → Session → Claude:
```
OrchestratorLoop decides to execute Phase 3, Task 2
│
▼
Converts OrchestratorTask to CreateTaskOptions:
{
prompt: "Implement the rate limiter middleware. Read src/middleware/auth.ts
for the pattern. Add to src/middleware/rate-limiter.ts. Must export
a Fastify plugin. When done: <promise>PHASE_3_TASK_2_DONE</promise>",
priority: 100,
dependencies: ["phase-3-task-1"], // Must finish auth middleware first
completionPhrase: "PHASE_3_TASK_2_DONE",
timeoutMs: 600000 // 10 minutes
}
│
▼
TaskQueue.addTask(options)
│
▼
RalphLoop.tick() → assignTasks() // OR OrchestratorLoop does its own assignment
│
▼
session.sendInput(task.prompt)
│
▼
writeViaMux() → tmux send-keys -l "prompt..." + Enter
│
▼
Claude CLI receives prompt, executes, outputs results
│
▼
RalphTracker.processData() → detects "PHASE_3_TASK_2_DONE"
│
▼
emit('completionDetected') → OrchestratorLoop.handleTaskCompleted()
│
▼
Check: all tasks in Phase 3 done? → If yes → verifyPhase(phase3)
```
## Team Agent Flow (When Enabled)
```
Phase has teamStrategy.type === 'team'
│
▼
OrchestratorLoop creates/reuses a session with:
env: { CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS: '1' }
│
▼
Sends team orchestration prompt:
"You're the team lead for Phase 3: Core Implementation.
Your team should work on these tasks in parallel:
1. Rate limiter middleware (teammate 1)
2. Error handling middleware (teammate 2)
3. Validation layer (teammate 3)
Context files to read first: [...]
Each teammate should output their task's completion phrase when done.
When ALL tasks are complete, output: <promise>PHASE_3_COMPLETE</promise>"
│
▼
Claude Code team-lead spawns teammates
│
▼
TeamWatcher detects new team in ~/.claude/teams/
→ Matches to session via leadSessionId
→ Tracks teammate activity
│
▼
Teammates work in parallel (in-process threads)
│
▼
hook: teammate_idle → POST /api/hook-event
→ OrchestratorLoop notes teammate finished
│
▼
hook: task_completed → POST /api/hook-event
→ Or: RalphTracker detects PHASE_3_COMPLETE
→ OrchestratorLoop → phase complete → verify
```
## Error Recovery Strategy
```
Task fails (timeout, error, session crash)
│
├─ Task-level retry (up to 2 retries per task)
│ → Reset task to pending
│ → Re-queue with modified prompt: "Previous attempt failed: {error}. Try again..."
│
├─ Phase-level retry (up to 3 retries per phase)
│ → Respawn session (fresh context)
│ → Re-execute entire phase with learnings from failure
│ → Modified prompt includes what went wrong
│
└─ Orchestration-level failure
→ All retries exhausted
→ state = FAILED
→ Notify user with detailed failure report
→ User can: modify plan → retry, skip phase → continue, or stop
```
## Interaction with Ralph Loop
Ralph Loop and Orchestrator Loop are **mutually exclusive** on the same sessions:
```
if (orchestratorLoop.isRunning()) {
// Orchestrator controls task assignment
// Ralph Loop should not interfere
// Respawn Controller uses 'orchestrator' preset
}
if (ralphLoop.isRunning()) {
// Ralph controls task assignment
// Orchestrator should not start
}
```
The Orchestrator can optionally USE the Ralph Loop internally for phase execution (delegate phase tasks to Ralph's queue), or manage task assignment directly. Decision: **manage directly** — gives more control over phase boundaries and verification timing.
## Summary of What Touches What
| Existing File | Change |
|---|---|
| `src/types/index.ts` | Export orchestrator types |
| `src/state-store.ts` | Add orchestrator state persistence |
| `src/web/sse-events.ts` | Add ~8 orchestrator events |
| `src/web/routes/index.ts` | Register orchestrator routes |
| `src/web/server.ts` | Initialize OrchestratorLoop |
| `src/web/public/constants.js` | Mirror SSE events |
| `src/web/public/app.js` | Add orchestrator event listeners, panel toggle |
| `src/web/route-helpers.ts` | Add 'orchestrator' respawn preset |
| New File | Purpose |
|---|---|
| `src/orchestrator-loop.ts` | Core state machine |
| `src/orchestrator-planner.ts` | Plan generation + phasing |
| `src/orchestrator-verifier.ts` | Phase verification |
| `src/types/orchestrator.ts` | Type definitions |
| `src/prompts/orchestrator.ts` | Prompt templates |
| `src/web/routes/orchestrator-routes.ts` | API endpoints |
| `src/web/public/orchestrator-ui.js` | Frontend panel |
| `src/web/ports/orchestrator-port.ts` | Port interface |
+633
View File
@@ -0,0 +1,633 @@
# Orchestrator Loop — Detailed Implementation Plan (v2)
> Internal research/planning document. Not for GitHub.
## Vision
The **Orchestrator Loop** is a new autonomous execution mode that transforms high-level user goals into phased, verified, team-coordinated implementations. Unlike Ralph Loop (flat task queue → idle sessions), the Orchestrator manages the full lifecycle: **plan → approve → execute → verify → adapt → complete**.
```
USER: "Add OAuth2 login with Google/GitHub, role-based access control, and API key management"
ORCHESTRATOR:
Phase 1: Research & Setup ✅ (3m) — scaffold, deps, config
Phase 2: Auth Core ✅ (8m) — OAuth2 flow, session mgmt
Phase 3: Provider Integration 🔄 (12m) — Google + GitHub (parallel via team agents)
Phase 4: RBAC ⏳ — roles, permissions, middleware
Phase 5: API Keys ⏳ — generation, validation, rate limits
Phase 6: Testing & Review ⏳ — integration tests, security review
Progress: ━━━━━━━━━━━━━━━━━━━━ 40% | Agents: 3 active | Time: 23m
```
## Architecture
```
┌─────────────────────────────────────────────────────────────────┐
│ OrchestratorLoop │
│ │
│ ┌────────────────┐ ┌────────────────┐ ┌──────────────────┐ │
│ │ Orchestrator │ │ Orchestrator │ │ Orchestrator │ │
│ │ Planner │ │ Executor │ │ Verifier │ │
│ │ │ │ │ │ │ │
│ │ PlanOrchestrator│ │ TaskQueue │ │ AI review │ │
│ │ + phase grouper│ │ SessionManager │ │ Test commands │ │
│ │ + team strategy│ │ Team prompts │ │ File checks │ │
│ └───────┬────────┘ └───────┬────────┘ └─────────┬────────┘ │
│ │ │ │ │
│ └───────────────────┼──────────────────────┘ │
│ │ │
│ ┌─────────▼─────────┐ │
│ │ Existing Codeman │ │
│ │ Infrastructure │ │
│ │ │ │
│ │ SessionManager │ │
│ │ TaskQueue │ │
│ │ RespawnController │ │
│ │ TeamWatcher │ │
│ │ PlanOrchestrator │ │
│ │ StateStore │ │
│ │ Hooks + SSE │ │
│ └────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
```
## State Machine
```
┌─────────┐
│ IDLE │
└────┬────┘
│ start(goal)
▼
┌─────────┐
┌────────│PLANNING │────────┐
│ fail └────┬────┘ │
▼ │ plan ready │ user cancels
┌────────┐ ▼ ▼
│ FAILED │ ┌─────────┐ ┌────────┐
└────────┘ │APPROVAL │ │ IDLE │
▲ └────┬────┘ └────────┘
│ │ approve
│ ▼
│ ┌──────────┐
│ ┌───►│EXECUTING │◄────────────────────┐
│ │ └────┬─────┘ │
│ │ │ all tasks in phase done │
│ │ ▼ │
│ │ ┌──────────┐ │
│ │ │VERIFYING │ │
│ │ └────┬─────┘ │
│ │ pass │ │ fail │
│ │ ▼ ▼ │
│ │ more ┌──────────┐ │
│ │ phases?│REPLANNING│── retry ────────┘
│ │ │ └────┬─────┘
│ │ │ │ max retries
│ │ │ ▼
│ │ │ ┌────────┐
│ └────┘ │ FAILED │
│ next └────────┘
│ phase
│ │
│ ▼
│ ┌───────────┐
└─│ COMPLETED │
└───────────┘
```
**States:** `idle` | `planning` | `approval` | `executing` | `verifying` | `replanning` | `completed` | `failed` | `paused`
Transitions are event-driven. The state machine is the single source of truth — all methods check `this.state` before acting.
## Type Definitions
### `src/types/orchestrator.ts`
```typescript
// ═══════════════════════════════════════════════════════════════
// State Machine
// ═══════════════════════════════════════════════════════════════
export type OrchestratorState =
| 'idle'
| 'planning'
| 'approval'
| 'executing'
| 'verifying'
| 'replanning'
| 'completed'
| 'failed'
| 'paused';
// ═══════════════════════════════════════════════════════════════
// Plan Structure
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorPlan {
id: string;
goal: string;
createdAt: number;
phases: OrchestratorPhase[];
metadata: {
totalTasks: number;
estimatedComplexity: 'low' | 'medium' | 'high';
modelUsed: string;
planDurationMs: number;
};
}
export interface OrchestratorPhase {
id: string; // "phase-1", "phase-2"
name: string; // Human-readable name
description: string;
order: number;
status: PhaseStatus;
tasks: OrchestratorTask[];
verificationCriteria: string[];
testCommands: string[];
maxAttempts: number; // Default: 3
attempts: number; // Current attempt count
startedAt: number | null;
completedAt: number | null;
durationMs: number | null;
teamStrategy: TeamStrategy;
}
export type PhaseStatus =
| 'pending'
| 'executing'
| 'verifying'
| 'passed'
| 'failed'
| 'skipped';
export interface OrchestratorTask {
id: string; // "phase-1-task-1"
phaseId: string;
prompt: string; // Single-line prompt for Claude
status: 'pending' | 'running' | 'completed' | 'failed';
assignedSessionId: string | null;
queueTaskId: string | null; // Links to TaskQueue task
parallel: boolean; // Can run in parallel with sibling tasks
completionPhrase: string; // Unique phrase for completion detection
timeoutMs: number;
startedAt: number | null;
completedAt: number | null;
error: string | null;
retries: number;
}
// ═══════════════════════════════════════════════════════════════
// Team Strategy
// ═══════════════════════════════════════════════════════════════
export type TeamStrategy =
| { type: 'single' } // One session handles all
| { type: 'parallel'; maxSessions: number } // Multiple sessions
| { type: 'team'; config: TeamSetup } // Agent teams
export interface TeamSetup {
leadPrompt: string;
suggestedTeammates: string[]; // Role descriptions
maxTeammates: number;
}
// ═══════════════════════════════════════════════════════════════
// Verification
// ═══════════════════════════════════════════════════════════════
export interface VerificationResult {
passed: boolean;
checks: VerificationCheck[];
summary: string;
suggestions: string[]; // Recovery hints for replanning
}
export interface VerificationCheck {
type: 'test_command' | 'ai_review' | 'file_check';
description: string;
passed: boolean;
output?: string;
}
// ═══════════════════════════════════════════════════════════════
// Configuration
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorConfig {
plannerModel: string; // Default: 'opus'
researchEnabled: boolean; // Default: true
autoApprove: boolean; // Default: false
maxPhaseRetries: number; // Default: 3
phaseTimeoutMs: number; // Default: 1800000 (30min)
enableTeamAgents: boolean; // Default: true
maxParallelSessions: number; // Default: 3
verificationMode: 'strict' | 'moderate' | 'lenient';
compactBetweenPhases: boolean; // Default: true
}
// ═══════════════════════════════════════════════════════════════
// Persistence (saved to ~/.codeman/state.json)
// ═══════════════════════════════════════════════════════════════
export interface OrchestratorPersistState {
state: OrchestratorState;
plan: OrchestratorPlan | null;
currentPhaseIndex: number;
startedAt: number | null;
completedAt: number | null;
config: OrchestratorConfig;
stats: OrchestratorStats;
}
export interface OrchestratorStats {
phasesCompleted: number;
phasesFailed: number;
totalTasksCompleted: number;
totalTasksFailed: number;
totalDurationMs: number;
replanCount: number;
}
```
## New Files (Implementation Order)
### Step 1: `src/types/orchestrator.ts` — Type definitions
All interfaces above. No dependencies. ~120 lines.
### Step 2: `src/orchestrator-planner.ts` — Plan generation + phase grouping
~300 lines. Wraps existing PlanOrchestrator.
```typescript
/**
* @fileoverview Orchestrator plan generation — converts goals into phased plans.
*
* Uses PlanOrchestrator for AI plan generation, then groups PlanItems into
* sequential phases with team strategies and verification criteria.
*
* @module orchestrator-planner
*/
export class OrchestratorPlanner {
constructor(mux: TerminalMultiplexer, workingDir: string, config: OrchestratorConfig);
/** Generate plan from goal. Uses PlanOrchestrator internally. */
async generatePlan(goal: string, onProgress?: ProgressCallback): Promise<OrchestratorPlan>;
/** Cancel in-progress plan generation. */
async cancel(): Promise<void>;
// Internal
private groupIntoPhases(items: PlanItem[], goal: string): OrchestratorPhase[];
private assignTeamStrategies(phases: OrchestratorPhase[]): void;
private generateCompletionPhrases(plan: OrchestratorPlan): void;
}
```
**Phase grouping algorithm:**
1. Topological sort by `PlanItem.dependencies`
2. Group into dependency layers (Kahn's algorithm)
3. Within each layer, sub-group by `tddPhase` (setup → test → impl → verify → review)
4. Merge adjacent small phases (< 2 tasks) if they share the same tddPhase
5. Assign team strategies:
- 1-2 tasks → `{ type: 'single' }`
- 3+ independent tasks → `{ type: 'parallel', maxSessions: Math.min(taskCount, config.maxParallelSessions) }`
- 4+ tasks with high complexity → `{ type: 'team', config: { ... } }`
6. Generate unique completion phrases per task: `ORCH_P{phaseOrder}_T{taskIndex}`
### Step 3: `src/orchestrator-verifier.ts` — Phase verification
~200 lines.
```typescript
/**
* @fileoverview Orchestrator phase verification.
*
* Runs verification checks after each phase completes:
* test commands, AI review, and file existence checks.
*
* @module orchestrator-verifier
*/
export class OrchestratorVerifier {
constructor(config: OrchestratorConfig);
/** Run all verification checks for a completed phase. */
async verifyPhase(
phase: OrchestratorPhase,
session: Session,
mode: 'strict' | 'moderate' | 'lenient'
): Promise<VerificationResult>;
// Verification strategies
private async runTestCommands(commands: string[], session: Session): Promise<VerificationCheck[]>;
private async aiReview(phase: OrchestratorPhase, session: Session): Promise<VerificationCheck>;
}
```
**Verification modes:**
- `strict`: ALL test commands must pass AND AI review must approve
- `moderate`: Test commands must pass, AI review is advisory
- `lenient`: At least one test command passes, AI review skipped
**AI review prompt (sent as a task to the session):**
```
Review Phase "{phase.name}" completion. Check:
1. Expected functionality works
2. No obvious regressions
3. Code quality is acceptable
Criteria: {phase.verificationCriteria.join('\n')}
If ALL criteria are met, respond: ORCH_VERIFY_PASS
If ANY criteria fail, respond: ORCH_VERIFY_FAIL and explain what failed.
```
### Step 4: `src/orchestrator-loop.ts` — Core state machine
~500 lines. Main orchestrator engine.
```typescript
/**
* @fileoverview Orchestrator Loop — phased plan execution with team agents.
*
* State machine that generates plans from user goals, executes them
* phase-by-phase with verification gates, and adapts on failure.
*
* @module orchestrator-loop
*/
export interface OrchestratorLoopEvents {
stateChanged: (state: OrchestratorState, prevState: OrchestratorState) => void;
planReady: (plan: OrchestratorPlan) => void;
phaseStarted: (phase: OrchestratorPhase) => void;
phaseCompleted: (phase: OrchestratorPhase) => void;
phaseFailed: (phase: OrchestratorPhase, reason: string) => void;
taskAssigned: (task: OrchestratorTask, sessionId: string) => void;
taskCompleted: (task: OrchestratorTask) => void;
taskFailed: (task: OrchestratorTask, error: string) => void;
verificationResult: (phase: OrchestratorPhase, result: VerificationResult) => void;
completed: (stats: OrchestratorStats) => void;
error: (error: Error) => void;
}
export class OrchestratorLoop extends EventEmitter {
private state: OrchestratorState = 'idle';
private plan: OrchestratorPlan | null = null;
private currentPhaseIndex = 0;
private config: OrchestratorConfig;
private planner: OrchestratorPlanner;
private verifier: OrchestratorVerifier;
private sessionManager: SessionManager;
private taskQueue: TaskQueue;
private store: StateStore;
private stats: OrchestratorStats;
private cleanup: CleanupManager;
private pausedState: OrchestratorState | null = null; // State before pause
// ── Lifecycle ──────────────────────────────────────────────
constructor(mux: TerminalMultiplexer, workingDir: string, config?: Partial<OrchestratorConfig>);
/** Start orchestration with a goal. Transitions: idle → planning */
async start(goal: string): Promise<void>;
/** Approve the generated plan. Transitions: approval → executing */
async approve(): Promise<void>;
/** Reject plan with feedback. Transitions: approval → planning (regenerate) */
async reject(feedback: string): Promise<void>;
/** Pause execution. Saves current state. */
pause(): void;
/** Resume from pause. */
resume(): void;
/** Stop everything and clean up. → idle */
async stop(): Promise<void>;
/** Skip current phase. → executing (next phase) or completed */
async skipPhase(phaseId: string): Promise<void>;
/** Retry a failed phase. → executing */
async retryPhase(phaseId: string): Promise<void>;
// ── Getters ────────────────────────────────────────────────
getState(): OrchestratorState;
getPlan(): OrchestratorPlan | null;
getCurrentPhase(): OrchestratorPhase | null;
getStats(): OrchestratorStats;
getStatus(): OrchestratorPersistState;
// ── Internal: Phase Execution ──────────────────────────────
private async executeCurrentPhase(): Promise<void>;
private async executePhase(phase: OrchestratorPhase): Promise<void>;
private async assignPhaseTasks(phase: OrchestratorPhase): Promise<void>;
private handleTaskCompleted(taskId: string): void;
private handleTaskFailed(taskId: string, error: string): void;
private async onPhaseTasksComplete(phase: OrchestratorPhase): Promise<void>;
// ── Internal: Verification ─────────────────────────────────
private async verifyCurrentPhase(): Promise<void>;
private async handleVerificationResult(phase: OrchestratorPhase, result: VerificationResult): Promise<void>;
// ── Internal: Replanning ───────────────────────────────────
private async replanPhase(phase: OrchestratorPhase, failures: string[]): Promise<void>;
// ── Internal: State Machine ────────────────────────────────
private setState(newState: OrchestratorState): void;
private advanceToNextPhase(): Promise<void>;
private persist(): void;
private restore(): void;
}
```
**Key execution flow in `executePhase()`:**
1. Mark phase as `executing`, emit `phaseStarted`
2. For each task in phase:
- Create a `CreateTaskOptions` from `OrchestratorTask`
- Add to `TaskQueue` with proper dependencies + completion phrase
- Store the TaskQueue task ID in `OrchestratorTask.queueTaskId`
3. Poll task completion (listen to TaskQueue events)
4. When all tasks complete → call `onPhaseTasksComplete()`
5. `onPhaseTasksComplete()` triggers verification
**How tasks get assigned to sessions:**
The OrchestratorLoop does NOT manage session assignment directly. It adds tasks to the existing TaskQueue and starts a mini poll loop that assigns pending tasks to idle sessions — the same pattern as RalphLoop's `assignTasks()`. This reuses existing session management.
**Team agent flow:**
For phases with `teamStrategy.type === 'team'`:
- Start a single session with `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
- Instead of adding individual tasks to TaskQueue, send ONE comprehensive prompt to the lead
- The prompt instructs the lead to create teammates and delegate
- Monitor via TeamWatcher for team task completion + hook events
- Phase completion is detected via the lead's completion phrase
### Step 5: `src/web/routes/orchestrator-routes.ts` — API endpoints
~300 lines.
```
POST /api/orchestrator/start — { goal, config? } → start planning
POST /api/orchestrator/approve — approve generated plan
POST /api/orchestrator/reject — { feedback } → reject + replan
POST /api/orchestrator/pause — pause execution
POST /api/orchestrator/resume — resume execution
POST /api/orchestrator/stop — stop orchestration
GET /api/orchestrator/status — full state + plan + stats
GET /api/orchestrator/plan — plan details only
POST /api/orchestrator/phase/:id/skip — skip a phase
POST /api/orchestrator/phase/:id/retry — retry a failed phase
```
Port dependency: `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort`
The route module receives the OrchestratorLoop instance via the InfraPort (added to `createRouteContext()`).
### Step 6: SSE Events — `src/web/sse-events.ts` additions
```typescript
// ─── Orchestrator ────────────────────────────────────────────────────────────
/** Orchestrator state machine transitioned. */
export const OrchestratorStateChanged = 'orchestrator:stateChanged' as const;
/** Orchestrator plan generated and ready for approval. */
export const OrchestratorPlanReady = 'orchestrator:planReady' as const;
/** Orchestrator phase started executing. */
export const OrchestratorPhaseStarted = 'orchestrator:phaseStarted' as const;
/** Orchestrator phase completed successfully. */
export const OrchestratorPhaseCompleted = 'orchestrator:phaseCompleted' as const;
/** Orchestrator phase failed. */
export const OrchestratorPhaseFailed = 'orchestrator:phaseFailed' as const;
/** Orchestrator verification result for a phase. */
export const OrchestratorVerification = 'orchestrator:verification' as const;
/** Orchestrator task assigned to session. */
export const OrchestratorTaskAssigned = 'orchestrator:taskAssigned' as const;
/** Orchestrator task completed. */
export const OrchestratorTaskCompleted = 'orchestrator:taskCompleted' as const;
/** Orchestrator task failed. */
export const OrchestratorTaskFailed = 'orchestrator:taskFailed' as const;
/** All phases completed successfully. */
export const OrchestratorCompleted = 'orchestrator:completed' as const;
/** Orchestrator error. */
export const OrchestratorError = 'orchestrator:error' as const;
```
11 new events. Add to `SseEvent` namespace object + mirror in `constants.js`.
### Step 7: State persistence — `src/state-store.ts` additions
Add to `AppState`:
```typescript
orchestrator?: OrchestratorPersistState;
```
Add methods:
```typescript
getOrchestratorState(): OrchestratorPersistState | null;
setOrchestratorState(state: Partial<OrchestratorPersistState>): void;
clearOrchestratorState(): void;
```
### Step 8: Server integration — `src/web/server.ts` modifications
1. Import `OrchestratorLoop` and `registerOrchestratorRoutes`
2. Add `private orchestratorLoop: OrchestratorLoop` field
3. Initialize in constructor (lazy — created on first start, not at boot)
4. Add to `createRouteContext()` InfraPort: `orchestratorLoop: this.orchestratorLoop`
5. Wire up OrchestratorLoop events → SSE broadcasts
6. Register routes: `registerOrchestratorRoutes(this.app, ctx)`
7. Clean up in `stop()`
### Step 9: `src/web/public/orchestrator-ui.js` — Frontend panel
~500 lines. New frontend module.
**Load order**: After `panels-ui.js` (11), before `ralph-wizard.js` (13). So load order = 11.5.
**UI elements:**
- Goal input form (text area + config toggles)
- Plan approval view (phase list, task details, approve/reject buttons)
- Execution dashboard (progress bar, phase cards, task status indicators)
- Agent activity panel (session count, team status)
- Controls (pause, resume, stop, skip phase, retry phase)
**SSE listeners:**
- All 11 orchestrator events → update UI state
- Reuses existing session/respawn/team event handlers for agent monitoring
### Step 10: `src/prompts/orchestrator.ts` — Prompt templates
~200 lines.
Templates for:
- Phase execution prompt (tells Claude what to do in this phase)
- Team lead delegation prompt (instructs lead to create and coordinate teammates)
- Verification prompt (asks Claude to verify phase output)
- Replan prompt (gives failure context, asks for recovery steps)
### Step 11: Constants, schemas, route barrel updates
- `src/web/public/constants.js` — Add 11 SSE event mirrors
- `src/web/schemas.ts` — Add Zod schemas for orchestrator API input validation
- `src/web/routes/index.ts` — Export `registerOrchestratorRoutes`
- `src/web/ports/infra-port.ts` — Add `orchestratorLoop` to InfraPort
- `src/types/index.ts` — Export orchestrator types
## Existing File Modifications Summary
| File | Change | Lines |
|------|--------|-------|
| `src/types/index.ts` | Add orchestrator barrel export | +1 |
| `src/web/sse-events.ts` | Add 11 orchestrator events + SseEvent entries | +30 |
| `src/web/public/constants.js` | Mirror 11 SSE events | +15 |
| `src/web/routes/index.ts` | Export registerOrchestratorRoutes | +1 |
| `src/web/ports/infra-port.ts` | Add orchestratorLoop to InfraPort | +3 |
| `src/web/server.ts` | Initialize OrchestratorLoop, wire events, register routes | +40 |
| `src/web/schemas.ts` | Add orchestrator Zod schemas | +20 |
| `src/state-store.ts` | Add orchestrator state persistence | +20 |
| `src/web/public/app.js` | Add orchestrator SSE listeners + panel toggle | +30 |
| `src/web/public/index.html` | Add orchestrator-ui.js script tag | +1 |
**Total new code**: ~2,300 lines across 6 new files
**Total modifications**: ~160 lines across 10 existing files
## Implementation Execution Order
This is the actual build order — each step is a commit checkpoint:
1. **Types** — `src/types/orchestrator.ts` + barrel export. Zero risk, pure types.
2. **SSE events** — Add all 11 events to both `sse-events.ts` and `constants.js`. Wire in SseEvent namespace.
3. **State persistence** — Add orchestrator state to StateStore. Small, isolated change.
4. **Schemas** — Add Zod validation schemas for API input.
5. **Planner** — `src/orchestrator-planner.ts`. Can test in isolation.
6. **Verifier** — `src/orchestrator-verifier.ts`. Can test in isolation.
7. **Core loop** — `src/orchestrator-loop.ts`. The big one. Depends on planner + verifier.
8. **Prompts** — `src/prompts/orchestrator.ts`. Templates used by core loop.
9. **Port + routes** — `src/web/ports/infra-port.ts` update + `src/web/routes/orchestrator-routes.ts`.
10. **Server integration** — Wire OrchestratorLoop into WebServer. Routes become live.
11. **Frontend** — `src/web/public/orchestrator-ui.js` + app.js listeners + index.html script tag.
12. **Tests** — `test/orchestrator-*.test.ts`.
13. **Typecheck + lint** — Fix all issues, ensure CI passes.
## Edge Cases & Error Handling
- **Session limit reached**: Queue tasks and wait for sessions to free up (existing SessionManager handles this)
- **All sessions crash during phase**: Mark phase as failed, attempt replan
- **Verification flaky**: `moderate` mode allows test retries; `lenient` skips AI review
- **Plan too large**: Cap at 10 phases, 50 total tasks. Warn user.
- **Context overflow**: Auto-compact between phases. Respawn if needed (orchestrator state is external).
- **User pauses mid-phase**: Pause task assignment, don't cancel running tasks. Resume picks up where it left off.
- **Network/API errors during planning**: Retry plan generation up to 2 times, then fail with clear message.
- **Orchestrator vs Ralph conflict**: Mutually exclusive. Starting orchestrator stops Ralph if running. Starting Ralph stops orchestrator.
## Testing Strategy
- **Unit tests**: `test/orchestrator-planner.test.ts` — phase grouping algorithm, team strategy assignment
- **Unit tests**: `test/orchestrator-verifier.test.ts` — verification logic with mocked sessions
- **Integration tests**: `test/orchestrator-loop.test.ts` — state machine transitions, task lifecycle
- **Route tests**: `test/routes/orchestrator-routes.test.ts` — API validation, status responses
All tests use `MockSession` pattern from existing test infrastructure. No real tmux needed.
+157
View File
@@ -0,0 +1,157 @@
# Orchestrator Loop — Research Findings
> Research doc for the new "Orchestrator Loop" feature. Not for GitHub.
## What We're Building
A new autonomous loop variant — **Orchestrator Loop** — that takes high-level user tasks, decomposes them into a detailed plan using team agents, and executes the plan step-by-step with quality gates. Unlike Ralph Loop (which executes a flat task queue), the Orchestrator coordinates **planning, delegation, and verification** as a continuous cycle.
**Core idea**: User inputs a goal → Orchestrator creates a detailed plan → spins up team agents for parallel execution → validates each step → adapts the plan based on results → delivers polished output.
## Existing Infrastructure Analysis
### What We Can Reuse
#### 1. Ralph Loop (`src/ralph-loop.ts`)
- **Pattern**: Poll loop with `start() → tick() → stop()` lifecycle
- **Reusable**: Event-driven task assignment, session completion handling, timeout management
- **Limitation**: Flat task queue — no concept of phases, dependencies between task groups, or adaptive replanning
- **Key insight**: `assignTaskToSession()` uses `session.sendInput(task.prompt)` — simple prompt injection into PTY
#### 2. Task Queue (`src/task-queue.ts`) + Task (`src/task.ts`)
- **Already has**: Priority ordering, dependency tracking between tasks, completion phrase detection
- **Limitation**: No task *groups* or *phases*. Dependencies are task-to-task, not phase-to-phase
- **Key insight**: Tasks support `completionPhrase` — a string the task watches for in output. This is how Ralph knows a task is done
#### 3. Plan Orchestrator (`src/plan-orchestrator.ts`)
- **Already has**: 2-agent plan generation (Research Agent → Planner Agent), TDD-aware plan items with P0/P1/P2 priorities
- **Output**: `PlanItem[]` with dependencies, verification criteria, TDD phases, complexity ratings
- **Limitation**: Plan generation only — no execution. Plans are generated then sit in state/UI for human review
- **Key insight**: Uses `Session` directly to run Claude subagent instances for research and planning. Returns structured JSON
#### 4. Team Agents (`src/team-watcher.ts`, `~/.claude/teams/`)
- **Already has**: Team creation, member tracking, filesystem inbox messaging, task management via `~/.claude/tasks/{team-name}/`
- **Limitation**: Codeman can only *observe* teams (TeamWatcher is read-only polling), not *create* or *orchestrate* them
- **Key insight**: Teams are a Claude Code feature. Codeman monitors them but doesn't control them. We can't programmatically create teammates — Claude Code does that when you use `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
#### 5. Respawn Controller (`src/respawn-controller.ts`)
- **Already has**: Preset-based automation (ralph-todo, overnight-autonomous), circuit breaker, health scoring
- **Key insight**: The `ralph-todo` preset (8s idle, 480min max) is designed for autonomous task execution. We'd need a new preset or make Orchestrator Loop set its own timing
#### 6. Session Auto-Ops (`src/session-auto-ops.ts`)
- **Already has**: Auto-compact at token thresholds, auto-clear for context management
- **Key insight**: Critical for long Orchestrator runs — prevents context overflow during multi-step execution
#### 7. Hooks (`src/hooks-config.ts`)
- **Already has**: `idle_prompt`, `stop`, `teammate_idle`, `task_completed` hook events
- **Key insight**: Hooks fire POST to `/api/hook-event` — this is how Codeman knows when Claude is idle, stopped, or completed a task. The Orchestrator Loop can listen to these same events
### What We Need to Build New
1. **Plan → Task decomposition**: Convert PlanOrchestrator output (PlanItem[]) into executable task groups with phase ordering
2. **Multi-phase execution engine**: Execute plan phases sequentially, tasks within phases in parallel
3. **Verification gates**: After each phase, run verification (test commands, AI review) before proceeding
4. **Adaptive replanning**: When a task fails or verification fails, generate a recovery plan
5. **Team agent orchestration**: Leverage Claude Code's agent teams for parallel execution within phases
6. **Progress tracking & UI**: Real-time dashboard showing plan progress, phase status, agent activity
## How Teams Actually Work (Important Constraint)
After deep research, here's the reality of agent teams:
```
User starts session with CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
→ Claude Code creates a team-lead
→ Team-lead spawns teammates (in-process threads)
→ Teammates appear as subagents (detected by SubagentWatcher)
→ Communication via ~/.claude/teams/{name}/inboxes/{member}.json
→ Tasks tracked in ~/.claude/tasks/{team-name}/{N}.json
```
**Codeman cannot programmatically create team members.** This is a Claude Code internal feature. However, Codeman CAN:
- Start a session that has teams enabled
- Send a prompt to the lead that instructs it to use agent teams
- Monitor team activity via TeamWatcher
- React to teammate_idle and task_completed hook events
- Read team task status from the filesystem
**This means**: The Orchestrator Loop orchestrates at the *session prompt* level, not the *team member* level. We tell the lead what to do, and the lead decides how to use its team.
## Architecture Decision: Prompt-Level Orchestration
Given the team constraint, the Orchestrator Loop works by:
1. **Planning phase**: Use PlanOrchestrator to generate a detailed plan from user input
2. **Execution phase**: Feed plan steps as prompts to sessions, one phase at a time
3. **Verification phase**: After each phase, run verification prompts and check results
4. **Adaptation phase**: If verification fails, generate recovery prompts
The "team agents" aspect works by:
- Starting sessions with `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`
- Crafting prompts that *instruct the lead to delegate* to teammates
- Monitoring team activity to track parallel progress
- The lead agent is smart enough to decompose work across its team
## Key Technical Findings
### Session Input Mechanics
```typescript
// From session.ts - how we send prompts
await session.sendInput(task.prompt); // Uses writeViaMux() internally
// writeViaMux() does: tmux send-keys -l "prompt text" + tmux send-keys Enter
// CRITICAL: Single-line only! Multi-line breaks Ink rendering
```
### Completion Detection Chain
```
PTY output → RalphTracker.processData() → completion phrase fuzzy match
→ CompletionConfidence scoring (multi-signal: promise tag + todos + exit signal)
→ If confident → emit 'completionDetected'
→ RalphLoop listens → marks task complete → assigns next
```
### How Plan Items Map to Tasks
```typescript
// PlanItem has:
interface PlanItem {
id: string; // "P0-001"
content: string; // "Implement error handling for API endpoints"
priority: 'P0' | 'P1' | 'P2';
dependencies: string[]; // ["P0-000"] — other PlanItem IDs
verificationCriteria: string;
testCommand: string;
tddPhase: 'setup' | 'test' | 'impl' | 'verify' | 'review';
complexity: 'low' | 'medium' | 'high';
}
// Task has:
interface CreateTaskOptions {
prompt: string;
priority: number;
dependencies: string[]; // Task IDs
completionPhrase: string;
timeoutMs: number;
}
// Natural mapping: PlanItem.content → Task.prompt
// PlanItem.dependencies → Task.dependencies
// PlanItem.priority → Task.priority (P0=100, P1=50, P2=10)
// PlanItem.verificationCriteria → verification task prompt
```
### Context Management for Long Runs
- Auto-compact at ~110k tokens (configurable)
- Auto-clear at ~140k tokens (configurable)
- Respawn cycling: kill + restart session to reset context entirely
- For Orchestrator: we want compact between phases, respawn between major milestones
## Risk Assessment
| Risk | Severity | Mitigation |
|------|----------|------------|
| Context overflow during complex phases | High | Auto-compact between tasks, respawn between phases |
| Team agents not predictable | Medium | Orchestrate at session level, let Claude decide team delegation |
| Plan too ambitious → infinite loop | High | Phase budgets (max attempts per phase), circuit breaker |
| Verification too strict → blocks progress | Medium | Configurable strictness, human override via UI |
| Single-line prompt limit | Medium | Use CLAUDE.md file for complex instructions, prompt references file |
| Long planning phase delays execution | Low | Show plan for approval before execution |