refactor: simplify ralph wizard from 9 agents to 2

Remove execution layer that never actually controlled execution:
- execution-bridge.ts (model param was ignored)
- group-scheduler.ts (task-tool mode never used)
- model-selector.ts (recommendations were display-only)
- context-manager.ts (never invoked)
- execution-limits.ts

Remove redundant agent prompts (overlapping outputs):
- requirements-analyst, architecture-planner, risk-analyst
- testing-specialist, verification (merged into planner.ts)
- execution-optimizer (output was ignored)
- final-review (scores were cosmetic)

Simplify plan-orchestrator.ts from 2400 LOC to 520 LOC:
- Before: 9 agents, 6 phases, ~40-60 minutes
- After: 2 agents (research + planner), ~18 minutes

Remove /api/execution/* endpoints and related server code.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
This commit is contained in:
arkon
2026-01-30 12:15:09 +01:00
co-authored by Claude Opus 4.5
parent 0020ae5a3e
commit a92826b66d
18 changed files with 329 additions and 5468 deletions
-32
View File
@@ -1,32 +0,0 @@
/**
* Architecture Planner Prompt
*
* Designs software component architecture for the task.
*
* Placeholders: {TASK}, {RESEARCH_CONTEXT}
*/
export const ARCHITECTURE_PLANNER_PROMPT = `You are an Architecture Planner specializing in software component design.
## YOUR TASK
Design the architecture for implementing this task:
## TASK DESCRIPTION
{TASK}
{RESEARCH_CONTEXT}
## INSTRUCTIONS
1. Identify all modules/components needed
2. Define interfaces between components
3. Specify data structures and types
4. Note configuration and setup requirements
5. Consider separation of concerns
## OUTPUT FORMAT
Return ONLY a JSON array:
[
{"category": "module|interface|type|config|infrastructure", "content": "component description", "rationale": "why needed"}
]
Generate 10-20 items. Think about the complete system architecture.`;
-134
View File
@@ -1,134 +0,0 @@
/**
* Execution Optimizer Prompt
*
* Optimizes the plan for Claude Code execution with parallel groups,
* agent types, model recommendations, and token estimates.
*
* Placeholders: {TASK}, {PLAN}
*/
export const EXECUTION_OPTIMIZER_PROMPT = `You are a Claude Code Execution Optimizer. Your job is to analyze an implementation plan and optimize it for efficient execution using Claude Code's agent system.
## ORIGINAL TASK
{TASK}
## CURRENT PLAN
{PLAN}
## YOUR MISSION
Analyze and enhance this plan for optimal Claude Code execution:
### 1. PARALLEL EXECUTION GROUPS
Identify tasks that can run simultaneously in separate agents:
- Tasks with NO dependencies between them
- Tasks that modify DIFFERENT files
- Tasks that read-only operations (exploration, analysis)
- Assign a parallelGroup ID (e.g., "parallel-1", "parallel-2") to related tasks
### 2. AGENT TYPE RECOMMENDATIONS
For each task, recommend the optimal Claude Code agent type:
- "explore": For codebase exploration, finding files, understanding patterns
- "implement": For writing new code, features, modifications
- "test": For writing and running tests
- "review": For code review, security analysis, best practices
- "general": For mixed or unclear tasks
### 3. FRESH CONTEXT RECOMMENDATIONS
Mark tasks that benefit from a fresh context (new conversation):
- After large file modifications (>500 lines changed)
- When switching between unrelated features
- After test failures that need fresh analysis
- When accumulated context might cause confusion
### 4. MODEL RECOMMENDATIONS
Suggest the optimal model for each task:
- "opus": Complex architecture, critical decisions, security review
- "sonnet": Standard implementation, most coding tasks
- "haiku": Quick exploration, simple searches, routine checks
### 5. FILE SCOPE ANALYSIS
For each task, identify:
- inputFiles: Files the task will need to READ
- outputFiles: Files the task will CREATE or MODIFY
### 6. TOKEN ESTIMATION
Estimate token usage for each task:
- Small (exploration, simple changes): 5000-15000
- Medium (feature implementation): 15000-50000
- Large (complex features, refactoring): 50000-100000
## OUTPUT FORMAT
Return ONLY a JSON object:
{
"optimizedPlan": [
{
"id": "P0-001",
"content": "Explore existing auth patterns in codebase",
"priority": "P0",
"tddPhase": "setup",
"verificationCriteria": "Documented auth patterns with file locations",
"dependencies": [],
"parallelGroup": "parallel-1",
"agentType": "explore",
"recommendedModel": "haiku",
"requiresFreshContext": false,
"estimatedTokens": 8000,
"inputFiles": ["src/auth/**/*.ts", "src/middleware/*.ts"],
"outputFiles": [],
"executionNotes": "Quick exploration, can run alongside P0-002"
},
{
"id": "P0-002",
"content": "Explore test patterns and fixtures",
"priority": "P0",
"tddPhase": "setup",
"verificationCriteria": "Understood test setup and conventions",
"dependencies": [],
"parallelGroup": "parallel-1",
"agentType": "explore",
"recommendedModel": "haiku",
"requiresFreshContext": false,
"estimatedTokens": 6000,
"inputFiles": ["test/**/*.test.ts", "test/fixtures/**/*"],
"outputFiles": [],
"executionNotes": "Parallel with P0-001, different file scope"
}
],
"parallelGroups": [
{
"id": "parallel-1",
"tasks": ["P0-001", "P0-002"],
"rationale": "Independent exploration tasks with no file overlap",
"estimatedDuration": "2-3 minutes",
"totalTokens": 14000
}
],
"executionStrategy": {
"totalParallelGroups": 3,
"sequentialBlockers": ["P0-005 blocks all P1 tasks"],
"freshContextPoints": ["After P0-005 (large refactor)", "After P1-003 (test failures)"],
"estimatedTotalTokens": 150000,
"estimatedAgentSpawns": 8,
"criticalPath": ["P0-001", "P0-003", "P0-005", "P1-001"],
"optimizationNotes": [
"Group 1 saves ~3 min by parallelizing exploration",
"Use haiku for 4 exploration tasks to reduce cost",
"Fresh context after auth refactor prevents confusion"
]
}
}
CRITICAL REQUIREMENTS:
1. Every task MUST have parallelGroup, agentType, recommendedModel
2. Parallel groups MUST NOT have overlapping outputFiles
3. Tasks in same parallelGroup MUST NOT depend on each other
4. Preserve all existing task fields (id, content, priority, etc.)
5. Add executionNotes explaining WHY this optimization
PARALLELIZATION GUIDELINES (BE CONSERVATIVE):
- Only parallelize tasks when you are CERTAIN they have no file conflicts
- Prefer sequential execution for complex or risky tasks
- Limit parallel groups to 2-3 tasks maximum per group
- When in doubt, keep tasks sequential - correctness over speed
- Focus parallelization on exploration/read-only tasks, not implementations
- Never parallelize tasks that might share state or side effects`;
-100
View File
@@ -1,100 +0,0 @@
/**
* Final Review Expert Prompt
*
* Provides holistic analysis of the complete implementation plan
* with scoring and improvement suggestions.
*
* Placeholders: {TASK}, {PLAN}
*/
export const FINAL_REVIEW_PROMPT = `You are a Final Review Expert providing a holistic analysis of an implementation plan.
## ORIGINAL TASK
{TASK}
## COMPLETE PLAN
{PLAN}
## YOUR MISSION
Review the ENTIRE plan from a high-level perspective. You have the bird's eye view.
### 1. LOGICAL FLOW ANALYSIS
Check if the plan makes logical sense:
- Does the order of tasks make sense?
- Are there circular dependencies or impossible orderings?
- Is there a clear progression from setup → implementation → testing → review?
- Are foundation tasks (types, configs, setup) done before dependent tasks?
### 2. COMPLETENESS CHECK
Verify nothing is missing:
- Every implementation has a corresponding test?
- Every test has clear verification criteria?
- Error handling and edge cases are covered?
- Setup and teardown steps are included?
- Documentation tasks if needed?
### 3. COHERENCE VALIDATION
Ensure the plan is internally consistent:
- Do task descriptions match their dependencies?
- Are file references consistent across tasks?
- Do parallel groups actually make sense together?
- Are priority levels justified?
### 4. FEASIBILITY ASSESSMENT
Is this plan actually achievable?
- Are any tasks too vague to execute?
- Are there unrealistic expectations?
- Are there hidden complexities not addressed?
- Is the scope creep under control?
### 5. SUGGESTED IMPROVEMENTS
Provide actionable fixes:
- Tasks to add if missing
- Tasks to split if too large
- Tasks to merge if redundant
- Order changes if needed
- Clarifications needed
## OUTPUT FORMAT
Return ONLY a JSON object:
{
"overallAssessment": "ready|needs-revision|major-issues",
"logicScore": 0.85,
"completenessScore": 0.90,
"coherenceScore": 0.88,
"feasibilityScore": 0.82,
"overallScore": 0.86,
"summary": "Brief 2-3 sentence summary of the plan quality",
"logicIssues": [
{"severity": "warning|error", "issue": "Description", "affectedTasks": ["P0-001"], "suggestion": "How to fix"}
],
"missingTasks": [
{"content": "Add database migration script", "reason": "Schema changes require migration", "insertAfter": "P0-002", "priority": "P0"}
],
"tasksToSplit": [
{"taskId": "P1-005", "reason": "Too complex", "splitInto": ["Implement auth logic", "Add session management"]}
],
"tasksToMerge": [
{"taskIds": ["P2-001", "P2-002"], "reason": "Redundant", "mergedContent": "Combined task description"}
],
"orderChanges": [
{"taskId": "P0-003", "currentPosition": 3, "suggestedPosition": 1, "reason": "Should run earlier"}
],
"clarificationsNeeded": [
{"taskId": "P1-002", "issue": "Unclear which API endpoint", "question": "Is this REST or GraphQL?"}
],
"finalRecommendations": [
"Start with P0 tasks in sequence for stable foundation",
"Consider adding integration tests after P1-004",
"Review security implications of auth changes"
]
}
SCORING GUIDELINES:
- 0.9+: Excellent, ready to execute
- 0.8-0.9: Good, minor tweaks recommended
- 0.7-0.8: Acceptable, some issues to address
- 0.6-0.7: Needs revision before execution
- <0.6: Major issues, significant rework needed
Be thorough but constructive. The goal is to catch issues before execution, not to criticize.`;
+1 -7
View File
@@ -6,11 +6,5 @@
*/
export { RESEARCH_AGENT_PROMPT } from './research-agent.js';
export { REQUIREMENTS_ANALYST_PROMPT } from './requirements-analyst.js';
export { ARCHITECTURE_PLANNER_PROMPT } from './architecture-planner.js';
export { TESTING_SPECIALIST_PROMPT } from './testing-specialist.js';
export { RISK_ANALYST_PROMPT } from './risk-analyst.js';
export { PLANNER_PROMPT } from './planner.js';
export { CODE_REVIEWER_PROMPT } from './code-reviewer.js';
export { VERIFICATION_PROMPT } from './verification.js';
export { EXECUTION_OPTIMIZER_PROMPT } from './execution-optimizer.js';
export { FINAL_REVIEW_PROMPT } from './final-review.js';
+83
View File
@@ -0,0 +1,83 @@
/**
* Planner Prompt - Single agent for TDD plan generation
*
* Combines what was previously 5 separate agents:
* - Requirements Analyst (redundant)
* - Architecture Planner (redundant)
* - Testing Specialist (kept - TDD focus)
* - Risk Analyst (redundant)
* - Verification Expert (kept - structure)
*
* Placeholders: {TASK}, {RESEARCH_CONTEXT}
*/
export const PLANNER_PROMPT = `You are a TDD Plan Generator. Create a complete implementation plan with test-first approach.
## TASK DESCRIPTION
{TASK}
{RESEARCH_CONTEXT}
## YOUR MISSION
Generate a complete TDD implementation plan with:
1. Tests BEFORE implementations (red-green-refactor)
2. Review tasks AFTER implementations
3. Clear priorities (P0=blocking, P1=required, P2=polish)
4. Dependencies between tasks
## TDD CYCLE
For each feature:
1. Write failing test first
2. Implement to make test pass
3. Review implementation
## PRIORITY GUIDELINES
- P0: Foundation, types, project setup, blocking dependencies
- P1: Core features, main implementation, error handling
- P2: Polish, optimization, documentation
## OUTPUT FORMAT
Return ONLY a JSON object:
{
"items": [
{
"id": "P0-001",
"content": "Write failing test for user authentication endpoint",
"priority": "P0",
"tddPhase": "test",
"verificationCriteria": "Test file exists, test fails with 'not implemented'",
"testCommand": "npm test -- --grep='auth'",
"dependencies": []
},
{
"id": "P0-002",
"content": "Implement user authentication handler",
"priority": "P0",
"tddPhase": "impl",
"verificationCriteria": "npm test -- --grep='auth' passes",
"pairedWith": "P0-001",
"dependencies": ["P0-001"]
},
{
"id": "P0-003",
"content": "Review auth implementation for security",
"priority": "P0",
"tddPhase": "review",
"verificationCriteria": "No security issues, follows best practices",
"reviewChecklist": ["Input validation", "XSS prevention", "Error handling"],
"pairedWith": "P0-002",
"dependencies": ["P0-002"]
}
],
"gaps": ["any missing requirements noted"],
"warnings": ["any concerns or risks identified"]
}
CRITICAL REQUIREMENTS:
1. Every implementation MUST have a paired test task that comes BEFORE it
2. Every implementation MUST have a review task that comes AFTER it
3. Use sequential IDs: P0-001, P0-002, P1-001, etc.
4. verificationCriteria must be SPECIFIC and observable
5. Dependencies must form a valid DAG (no cycles)
Generate 15-40 items covering the complete implementation.`;
-31
View File
@@ -1,31 +0,0 @@
/**
* Requirements Analyst Prompt
*
* Extracts explicit and implicit requirements from task descriptions.
*
* Placeholders: {TASK}, {RESEARCH_CONTEXT}
*/
export const REQUIREMENTS_ANALYST_PROMPT = `You are a Requirements Analyst specializing in extracting all requirements from task descriptions.
## YOUR TASK
Analyze the following task and extract ALL requirements (explicit and implicit):
## TASK DESCRIPTION
{TASK}
{RESEARCH_CONTEXT}
## INSTRUCTIONS
1. Identify explicit requirements (directly stated)
2. Infer implicit requirements (unstated but necessary)
3. Note any assumptions that should be validated
4. Consider non-functional requirements (performance, security, usability)
## OUTPUT FORMAT
Return ONLY a JSON array:
[
{"category": "functional|non-functional|constraint|assumption", "content": "requirement description", "rationale": "why this is needed"}
]
Generate 8-15 items. Be thorough - missing requirements cause project failures.`;
-32
View File
@@ -1,32 +0,0 @@
/**
* Risk Analyst Prompt
*
* Identifies potential issues, edge cases, and blockers.
*
* Placeholders: {TASK}, {RESEARCH_CONTEXT}
*/
export const RISK_ANALYST_PROMPT = `You are a Risk Analyst identifying potential issues and blockers.
## YOUR TASK
Identify risks and edge cases for this task:
## TASK DESCRIPTION
{TASK}
{RESEARCH_CONTEXT}
## INSTRUCTIONS
1. Identify potential failure points
2. Note edge cases that could cause bugs
3. Consider security vulnerabilities
4. Flag performance concerns
5. Identify dependencies that could block progress
## OUTPUT FORMAT
Return ONLY a JSON array:
[
{"category": "failure|edge-case|security|performance|dependency", "content": "risk description", "rationale": "mitigation approach"}
]
Generate 8-15 items. Being proactive about risks prevents surprises.`;
-76
View File
@@ -1,76 +0,0 @@
/**
* Testing Specialist Prompt
*
* Designs comprehensive, realistic test coverage with TDD approach.
*
* Placeholders: {TASK}, {RESEARCH_CONTEXT}
*/
export const TESTING_SPECIALIST_PROMPT = `You are a TDD Specialist designing a comprehensive, REALISTIC test strategy.
## YOUR TASK
Design detailed, executable test coverage for this task:
## TASK DESCRIPTION
{TASK}
{RESEARCH_CONTEXT}
## INSTRUCTIONS
Create REALISTIC tests that would actually run in a real codebase:
### 1. Unit Tests (test individual functions/methods in isolation)
- Mock external dependencies (databases, APIs, file system)
- Test pure logic with specific input/output examples
- Include exact assertion values, not placeholders
### 2. Integration Tests (test component interactions)
- Test API endpoints with realistic request/response bodies
- Test database operations with actual schema
- Test service-to-service communication
### 3. Edge Cases & Boundary Tests
- Empty inputs, null values, undefined
- Maximum/minimum values, overflow conditions
- Unicode, special characters, injection attempts
- Concurrent access, race conditions
### 4. Error Scenario Tests
- Network failures, timeouts, connection refused
- Invalid input validation with specific error messages
- Authorization failures, permission denied
- Resource not found, conflict states
### 5. Performance & Load Tests (where applicable)
- Response time thresholds
- Memory usage limits
- Concurrent user handling
## REALISTIC TEST EXAMPLE
BAD: "Test user login" (too vague)
GOOD: "Test POST /api/auth/login with valid email 'test@example.com' and password 'ValidPass123!' returns 200 with JWT token containing userId and exp claims, sets httpOnly cookie 'session'"
## OUTPUT FORMAT
Return ONLY a JSON array:
[
{
"category": "unit|integration|edge-case|error|e2e|performance",
"content": "Test POST /api/users with email 'new@test.com' creates user and returns 201 with {id, email, createdAt}",
"rationale": "Validates user creation happy path with all required response fields",
"verificationCriteria": "Response status 201, body contains id (uuid), email matches input, createdAt is valid ISO timestamp",
"testCommand": "npm test -- --grep='POST /api/users creates user'",
"testSetup": "Clear users table, seed with test data",
"testTeardown": "Delete created test user",
"pairedImpl": "Implement POST /api/users endpoint with validation and database insert",
"mockDependencies": ["database connection", "email service"],
"assertionDetails": ["status === 201", "body.id matches UUID regex", "body.email === 'new@test.com'"]
}
]
CRITICAL REQUIREMENTS:
- verificationCriteria: SPECIFIC observable outcomes with exact values
- testCommand: Actual runnable command (npm test, pytest, vitest, etc.)
- pairedImpl: The exact implementation step this test validates
- assertionDetails: List of specific assertions to make
Generate 15-30 detailed test items. Tests MUST be specific enough to implement directly.`;
-96
View File
@@ -1,96 +0,0 @@
/**
* Verification Expert Prompt
*
* Reviews and enhances the synthesized plan with priorities,
* verification criteria, and TDD pairing.
*
* Placeholders: {TASK}, {PLAN}
*/
export const VERIFICATION_PROMPT = `You are a Plan Verification Expert reviewing an implementation plan for completeness and quality.
## ORIGINAL TASK
{TASK}
## SYNTHESIZED PLAN (from multiple analysis subagents)
{PLAN}
## YOUR MISSION
Review and enhance this plan:
1. Assign priorities (P0=critical/blocking, P1=required, P2=enhancement)
2. Add verification criteria to EVERY task (how to know it's done)
3. Pair test tasks with implementation tasks (TDD cycle)
4. Add dependencies where one task blocks another
5. Identify gaps and calculate quality score
## PRIORITY GUIDELINES
- P0: Foundation tasks, type definitions, project setup, blocking dependencies
- P1: Core implementation, tests, main features, error handling
- P2: Polish, optimization, documentation, nice-to-have features
## TDD + REVIEW CYCLE RULES
The complete cycle is: test → impl → review
- Every implementation task should have a corresponding test task AND review task
- Test task comes BEFORE its paired implementation task
- Review task comes AFTER the implementation it reviews
- Use "pairedWith" to link test ↔ implementation ↔ review
- Verification criteria should reference test results where applicable
## REVIEW TASK REQUIREMENTS
After EVERY implementation task, add a review task that checks:
- Best practices for the language/framework
- Security vulnerabilities (OWASP top 10)
- Performance concerns
- Error handling completeness
- Code quality (DRY, SOLID, readability)
## OUTPUT FORMAT
Return ONLY a JSON object:
{
"validatedPlan": [
{
"id": "P0-001",
"content": "Write failing test for user authentication",
"priority": "P0",
"tddPhase": "test",
"verificationCriteria": "Test file exists, test fails with 'not implemented'",
"testCommand": "npm test -- --grep='auth'",
"pairedWith": "P0-002",
"dependencies": [],
"complexity": "low"
},
{
"id": "P0-002",
"content": "Implement user authentication handler",
"priority": "P0",
"tddPhase": "impl",
"verificationCriteria": "npm test -- --grep='auth' passes",
"pairedWith": "P0-001",
"dependencies": ["P0-001"],
"complexity": "medium"
},
{
"id": "P0-003",
"content": "Review auth implementation for security and best practices",
"priority": "P0",
"tddPhase": "review",
"verificationCriteria": "No security issues found, follows TypeScript best practices",
"reviewChecklist": ["Input validation", "XSS prevention", "Session security", "Error handling"],
"pairedWith": "P0-002",
"dependencies": ["P0-002"],
"complexity": "low"
}
],
"gaps": ["missing requirement 1", "missing test coverage for X"],
"warnings": ["consider Y before Z", "potential issue with..."],
"qualityScore": 0.85
}
CRITICAL REQUIREMENTS:
1. EVERY task MUST have verificationCriteria (how to verify completion)
2. Implementation tasks MUST have a paired test task AND a review task
3. Review tasks MUST have a reviewChecklist with specific items to check
4. Dependencies must form a valid DAG (no cycles)
5. Use sequential IDs: P0-001, P0-002, P0-003, P1-001, etc.
Be critical but constructive. A thorough review catches issues that tests miss.`;