Compare commits

...
Author SHA1 Message Date
arkonandClaude Fable 5 beeec63f72 fix(terminal): linear-time link-provider regex; always allow blob workers in CSP
cmdPattern's empty-matchable unbounded arg group backtracked exponentially
on wrapped heredoc/table lines — hovering one froze the tab for minutes.
Non-empty tokens + bounded reps make it O(n); regression test extracts the
shipped patterns and pins timing on the real killer shapes.

worker-src 'self' blob: is now unconditional so terminal-ui's _safeYield
tick worker (throttling escape) isn't CSP-blocked on non-gesture installs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 18:06:45 +02:00
arkonandClaude Fable 5 fad32eeaab chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 17:16:13 +02:00
arkonandClaude Fable 5 c1458d8ab8 feat(self-update): launchd-daemon supervisor — rootless restart on headless Macs
A KeepAlive system-level LaunchDaemon (the right setup for headless Macs,
where no GUI login means LaunchAgents never start) is now detected as
supervisor 'launchd-daemon': the updater kills the server PID (passed via
--server-pid) and launchd respawns it on the new dist/ — no root needed.
Detection requires the daemon plist to be bootstrapped AND KeepAlive=true.

Also: on boot, a 'completed-needs-manual-restart' status auto-completes
when the running version matches the staged target, so the stale
'restart Codeman to apply' instruction no longer lingers in the UI.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 17:08:36 +02:00
22 changed files with 466 additions and 541 deletions
+20
View File
@@ -1,5 +1,25 @@
# aicodeman
## 0.9.11
### Patch Changes
- Fix a terminal freeze on hover (catastrophic regex backtracking) and a CSP violation that disabled the terminal's anti-throttling worker.
**Tab-freezing hover bug**: the terminal link provider's `cmdPattern` (which turns `tail -f /path`-style text into clickable links) used an empty-matchable, unbounded arg group — `(?:[^\s\/]*\s+)*` — that backtracks exponentially on real Claude output, e.g. wrapped `git commit -m "$(cat <<'EOF'` heredoc lines or aligned table rows. Hovering the mouse over such a line hung the page's main thread for minutes ("page unresponsive"). The pattern now uses non-empty tokens with bounded repetition (linear time); all intended command+path link forms still match. New `test/link-provider-regex.test.ts` extracts the shipped patterns from source and pins linear-time behavior on the killer line shapes.
**Blob worker CSP fix**: `worker-src 'self' blob:` is now always present in the CSP (previously only with `CODEMAN_GESTURE=1`). The terminal's `_safeYield` anti-throttling tick worker is created from a Blob URL and was silently blocked on every install, logging a CSP violation on each page load and disabling the worker leg of the render-yield fallback chain.
## 0.9.10
### Patch Changes
- Self-update now restarts automatically on headless Macs supervised by a system LaunchDaemon.
New `launchd-daemon` supervisor kind: when Codeman runs under a bootstrapped, KeepAlive system-level LaunchDaemon (`/Library/LaunchDaemons/com.codeman.web.plist` — the right setup for headless Macs, where LaunchAgents never start because there is no GUI login), the updater no longer ends with "Update staged — restart Codeman to apply". It restarts rootlessly: the update script kills the server PID (passed via `--server-pid`) and launchd respawns it on the freshly built `dist/`. Detection is conservative — the daemon must be bootstrapped in the system domain AND have `KeepAlive` enabled.
Also fixed: a lingering "restart Codeman to apply" status. After a manual restart of a staged update, boot reconciliation now flips `completed-needs-manual-restart` to `completed` once the running version matches the staged target, so the Updates tab stops showing the stale instruction.
## 0.9.9
### Patch Changes
+6 -6
View File
@@ -56,7 +56,7 @@ When user says "COM":
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
**Version**: 0.9.9 (must match `package.json`)
**Version**: 0.9.11 (must match `package.json`)
## Project Overview
@@ -127,7 +127,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| **Tasks** | `src/task.ts`, `src/task-queue.ts`, `src/task-tracker.ts` | |
| **State** | `src/state-store.ts`, `src/run-summary.ts`, `src/session-lifecycle-log.ts` | |
| **Infra** | `src/hooks-config.ts`, `src/push-store.ts`, `src/tunnel-manager.ts`, `src/image-watcher.ts`, `src/file-stream-manager.ts` | |
| **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/claude-md.ts` | |
| **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/` (`claude-md.ts` + `case-template.md`, the CLAUDE.md scaffold generated into new cases) | |
| **Web** | `src/web/server.ts` ★, `src/web/sse-events.ts`, `src/web/routes/*.ts` (15 route modules + barrel; `session-routes.ts` ★), `src/web/route-helpers.ts`, `src/web/ports/*.ts`, `src/web/middleware/auth.ts`, `src/web/schemas.ts`, `src/web/self-update.ts` | |
| **Frontend** | `src/web/public/app.js` (~3.6K lines, core) + 5 infra modules (`constants.js`, `mobile-handlers.js`, `voice-input.js`, `notification-manager.js`, `keyboard-accessory.js`) + 7 domain modules (`terminal-ui.js`, `respawn-ui.js`, `ralph-panel.js`, `orchestrator-panel.js`, `settings-ui.js`, `panels-ui.js`, `session-ui.js`) + 5 feature modules (`ralph-wizard.js`, `api-client.js`, `subagent-windows.js`, `input-cjk.js`, `image-input.js`) + `sw.js` | |
| **Types** | `src/types/index.ts` (barrel) → 15 domain files; also `src/types.ts` root re-export | See `@fileoverview` in index.ts |
@@ -157,13 +157,13 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**External CLI modes (OpenCode, Codex)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior — Ralph tracker, BashToolParser, token/CLI-info parsing, and ❯-prompt readiness detection are all skipped (these CLIs render their own TUIs; readiness = output stabilization instead). Both modes **require tmux — no direct PTY fallback** — because secrets are injected via `tmux setenv`, never on the spawn command line: OpenCode gets `OPENCODE_CONFIG_CONTENT` etc., Codex gets `OPENAI_API_KEY`/`CODEX_API_KEY`/`CODEX_HOME` (`setCodexEnvVars` in `tmux-manager.ts`). Codex specifics: command built by `buildCodexCommand()` (`--model`, `resume <id>`, `--dangerously-bypass-approvals-and-sandbox` from the `codexConfig` payload / `codexDangerouslyBypassApprovals` app setting; `renderMode` is schema-coerced to `'hybrid'`, the only supported mode); tmux exports `COLORTERM=truecolor` + unsets `NO_COLOR` (other modes unset `COLORTERM`); availability via `GET /api/codex/status` — session/quick-start routes fail with `OPERATION_FAILED` and an install hint (`npm install -g @openai/codex`) when the binary is missing. Frontend: run-mode dropdown → `runCodex()` in `session-ui.js` ("Run CX" label), App Settings → Codex CLI tab; Respawn/Ralph options are Claude-only, so session options open on the Summary tab for external CLI sessions. Tests: `test/run-mode-ui.test.ts` (vm-sandbox harness, no real DOM).
**Hook events**: Claude Code hooks trigger via `/api/hook-event`. Key events: `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`. See `src/hooks-config.ts`.
**Hook events**: Claude Code hooks trigger via `/api/hook-event`. Key events: `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `stop`, `teammate_idle`, `task_completed`. See `src/hooks-config.ts`; upstream hook semantics mirrored in `docs/claude-code-hooks-reference.md`.
**Agent Teams**: `TeamWatcher` polls `~/.claude/teams/`, matches to sessions via `leadSessionId`. Teammates are in-process threads appearing as subagents. Enable: `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`. See `docs/agent-teams/`.
**Circuit breaker**: Prevents respawn thrashing. States: `CLOSED` → `HALF_OPEN` → `OPEN`. Reset: `/api/sessions/:id/ralph-circuit-breaker/reset`.
**Self-update** (App Settings → Updates): in-app updater for **git-clone installs** supervised by systemd/launchd. The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` (`git checkout <release tag> && npm install && npm run build && restart`) that outlives the restart; it writes progress to `dataPath('update-status.json')`, which the browser polls across the connection drop. Channel = latest `codeman@X.Y.Z` release tag; dirty trees are auto-stashed. `src/web/self-update.ts` splits PURE helpers (semver/tag parsing, reconcile decision — unit-tested) from IO wrappers (`getInstallInfo`/`checkForUpdate`/`startUpdate`/`reconcileUpdateOnBoot`). Routes: `GET /api/system/update/check`, `POST /api/system/update`, `GET /api/system/update/status`. Types: `src/types/update.ts`. npm installs report as non-updatable.
**Self-update** (App Settings → Updates): in-app updater for **git-clone installs** supervised by systemd/launchd. Supervisors: `systemd` (user unit), `launchd` (GUI LaunchAgent, gui-domain kickstart), `launchd-daemon` (KeepAlive system LaunchDaemon on headless Macs — restarts rootlessly by killing the server PID and letting launchd respawn it; detected only when the daemon is bootstrapped AND KeepAlive), else `none` → "restart manually" message; on next boot a manual-restart status auto-completes when the running version matches the target. The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` (`git checkout <release tag> && npm install && npm run build && restart`) that outlives the restart; it writes progress to `dataPath('update-status.json')`, which the browser polls across the connection drop. Channel = latest `codeman@X.Y.Z` release tag; dirty trees are auto-stashed. `src/web/self-update.ts` splits PURE helpers (semver/tag parsing, reconcile decision — unit-tested) from IO wrappers (`getInstallInfo`/`checkForUpdate`/`startUpdate`/`reconcileUpdateOnBoot`). Routes: `GET /api/system/update/check`, `POST /api/system/update`, `GET /api/system/update/status`. Types: `src/types/update.ts`. npm installs report as non-updatable.
**Port interfaces**: Routes declare dependencies via port interfaces (`src/web/ports/`). Routes use intersection types (e.g., `SessionPort & EventPort`).
@@ -211,7 +211,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
~135 handlers across 15 route files in `src/web/routes/`: system (41, incl. self-update `check`/`status`/`POST /api/system/update`, `POST /api/system/span-displays` → spawns `scripts/span-codeman.sh`, and `GET /api/codex/status`), sessions (28), orchestrator (10), cases (9), ralph (9), plan (8), respawn (7), files (6), mux (5), push (4), scheduled (4), teams (2), hooks (1), clipboard (1), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
**HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`): responses use the `ApiResponse<T>` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`).
**HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse<T>` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`).
## Adding Features
@@ -249,7 +249,7 @@ Raw `npx vitest` skips `config/vitest.config.ts`; always use `npm test --` or pa
**Ports**: Pick unique ports manually. Search `const PORT =` before adding new tests.
**Respawn tests**: Use `MockSession` from `test/respawn-test-utils.ts`. **Route tests**: `app.inject({ method, url, payload })` in `test/routes/` — no live port needed. **Mobile tests**: Playwright suite in `test/mobile/` (135 device profiles).
**Respawn tests**: Use `MockSession` from `test/respawn-test-utils.ts`. **Route tests**: `app.inject({ method, url, payload })` in `test/routes/` — no live port needed. **Mobile tests**: Playwright suite in `test/mobile/` (135 device profiles). Browser-testing infra and practices: `docs/browser-testing-guide.md`.
## Debugging
+2 -2
View File
@@ -2,10 +2,10 @@
<img src="docs/images/codeman-title.svg" alt="Codeman" height="60">
</p>
<h2 align="center">The missing control plane for AI coding agents</h2>
<h2 align="center">Mission control for AI coding agents</h2>
<p align="center">
<em>Agent Visualization &bull; Zero-Lag Input &bull; Mobile-First UI &bull; Hardened Security</em>
<em>Claude Code &bull; OpenCode &bull; Codex &mdash; One Dashboard &bull; Zero-Lag Mobile Input &bull; Any Device</em>
</p>
<p align="center">
+2 -2
View File
@@ -2,10 +2,10 @@
<img src="docs/images/codeman-title.svg" alt="Codeman" height="60">
</p>
<h2 align="center">为 AI 编程智能体而生的「控制平面」</h2>
<h2 align="center">AI 编程智能体的任务控制中心</h2>
<p align="center">
<em>智能体可视化 &bull; 零延迟输入 &bull; 自主编排器 &bull; 重生控制器 &bull; 移动优先 UI &bull; 安全加固</em>
<em>Claude Code &bull; OpenCode &bull; Codex —— 统一仪表盘 &bull; 零延迟移动输入 &bull; 任意设备</em>
</p>
<p align="center">
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "0.9.9",
"version": "0.9.11",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "0.9.9",
"version": "0.9.11",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
+2 -2
View File
@@ -1,7 +1,7 @@
{
"name": "aicodeman",
"version": "0.9.9",
"description": "The missing control plane for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"version": "0.9.11",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
"types": "dist/index.d.ts",
+1 -1
View File
@@ -39,7 +39,7 @@ Server echoes 'h' ←───────────────────
## Origin
This library was extracted from [Codeman](https://github.com/Ark0N/Codeman), the missing control plane for AI coding agents — multi-session management, real-time agent visualization, autonomous respawn loops, and a mobile-first web UI for Claude Code and OpenCode. The local echo system was built to make mobile and remote access feel instant, then battle-tested across thousands of hours of real usage. After 3 deep code audits, it was extracted into this standalone library with 78 tests covering every state transition.
This library was extracted from [Codeman](https://github.com/Ark0N/Codeman), mission control for AI coding agents — multi-session management, real-time agent visualization, autonomous respawn loops, and a mobile-first web UI for Claude Code, OpenCode, and Codex. The local echo system was built to make mobile and remote access feel instant, then battle-tested across thousands of hours of real usage. After 3 deep code audits, it was extracted into this standalone library with 78 tests covering every state transition.
## Install
+15
View File
@@ -31,6 +31,7 @@ export PUPPETEER_SKIP_DOWNLOAD="${PUPPETEER_SKIP_DOWNLOAD:-1}"
REPO=""
TAG=""
SUPERVISOR="none"
SERVER_PID=""
STATUS_FILE=""
UPDATE_ID=""
FROM_VERSION=""
@@ -50,6 +51,7 @@ while [[ $# -gt 0 ]]; do
--node) NODE="$2"; shift 2 ;;
--log) LOG="$2"; shift 2 ;;
--prev-sha) PREV_SHA="$2"; shift 2 ;;
--server-pid) SERVER_PID="$2"; shift 2 ;;
--stash) DO_STASH=1; shift ;;
*) shift ;;
esac
@@ -198,6 +200,19 @@ case "$SUPERVISOR" in
|| fail "Build succeeded but launchd restart failed" "launchctl"
}
;;
launchd-daemon)
# System-level KeepAlive LaunchDaemon (headless Mac): kickstarting the system
# domain needs root, but we don't need it — kill the server and launchd
# respawns it on the new dist/ within ThrottleInterval seconds.
if [[ -n "$SERVER_PID" ]] && kill "$SERVER_PID" 2>/dev/null; then
: # respawn is launchd's job from here
else
MANUAL_CMD="sudo launchctl kickstart -k system/com.codeman.web"
write_status "completed-needs-manual-restart" "Update staged — restart Codeman to apply v$TO_VERSION."
echo "[self-update] launchd-daemon: could not signal server pid '$SERVER_PID' — manual restart required"
exit 0
fi
;;
*)
MANUAL_CMD="pkill -f 'codeman.*web'; codeman web &"
write_status "completed-needs-manual-restart" "Update staged — restart Codeman to apply v$TO_VERSION."
+1
View File
@@ -8,3 +8,4 @@
export { RESEARCH_AGENT_PROMPT } from './research-agent.js';
export { PLANNER_PROMPT } from './planner.js';
export { PHASE_EXECUTION_PROMPT, TEAM_LEAD_PROMPT, REPLAN_PROMPT, SINGLE_TASK_PROMPT } from './orchestrator.js';
export { RALPH_STATUS_CONTRACT, buildRalphLoopPrompt, type RalphLoopPromptOptions } from './ralph.js';
+85
View File
@@ -0,0 +1,85 @@
/**
* @fileoverview Ralph Loop prompt construction
*
* Builds the full `@ralph_prompt.md` content written for a new Ralph loop
* session, including the RALPH_STATUS block contract. The contract travels
* with the loop prompt (not the generated CLAUDE.md) so every Ralph session
* emits parseable status blocks regardless of the project's CLAUDE.md.
*
* @module prompts/ralph
*/
/**
* Structured status-reporting contract appended to every Ralph loop prompt.
*
* `RalphStatusParser` (src/ralph-status-parser.ts) parses this block from
* session output — keep the field names and enum values in sync with its
* patterns.
*/
export const RALPH_STATUS_CONTRACT = `## Status Reporting
End EVERY response with exactly this block — Codeman parses it to track the loop:
\`\`\`
---RALPH_STATUS---
STATUS: IN_PROGRESS | COMPLETE | BLOCKED
TASKS_COMPLETED_THIS_LOOP: <number>
FILES_MODIFIED: <number>
TESTS_STATUS: PASSING | FAILING | NOT_RUN
WORK_TYPE: IMPLEMENTATION | TESTING | DOCUMENTATION | REFACTORING
EXIT_SIGNAL: false | true
RECOMMENDATION: <one line: what to do next>
---END_RALPH_STATUS---
\`\`\`
Rules:
- \`EXIT_SIGNAL: true\` only when ALL tasks are verifiably done — then also output the completion phrase
- \`STATUS: BLOCKED\` when you need human input; describe the blocker in RECOMMENDATION
- Never set \`EXIT_SIGNAL: true\` while tests are failing
`;
export interface RalphLoopPromptOptions {
/** The user's task description (becomes the prompt header) */
taskDescription: string;
/** Completion phrase the session must emit inside <promise></promise> */
completionPhrase: string;
/** Whether a @fix_plan.md task plan was generated for this loop */
hasPlan: boolean;
}
/**
* Builds the full Ralph loop prompt written to `@ralph_prompt.md`.
*/
export function buildRalphLoopPrompt({ taskDescription, completionPhrase, hasPlan }: RalphLoopPromptOptions): string {
let fullPrompt = taskDescription + '\n\n---\n\n';
if (hasPlan) {
fullPrompt += '## Task Plan\n\n';
fullPrompt += 'A task plan has been written to `@fix_plan.md`. Use this to track progress:\n';
fullPrompt += '- Reference the plan at the start of each iteration\n';
fullPrompt += '- Update task checkboxes as you complete items\n';
fullPrompt += '- Work through items in priority order (P0 > P1 > P2)\n\n';
}
fullPrompt += '## Iteration Protocol\n\n';
fullPrompt += 'This is an autonomous loop. Files from previous iterations persist. On each iteration:\n';
fullPrompt += '1. Check what work has already been done\n';
fullPrompt += '2. Make incremental progress toward completion\n';
fullPrompt += '3. Commit meaningful changes with descriptive messages\n\n';
fullPrompt += '## Verification\n\n';
fullPrompt += 'After each significant change:\n';
fullPrompt += '- Run tests to verify (npm test, pytest, etc.)\n';
fullPrompt += '- Check for type/lint errors if applicable\n';
fullPrompt += '- If tests fail, read the error, fix it, and retry\n\n';
fullPrompt += '## Completion Criteria\n\n';
fullPrompt += `Output \`<promise>${completionPhrase}</promise>\` when ALL of the following are true:\n`;
fullPrompt += '- All requirements from the task description are implemented\n';
fullPrompt += '- All tests pass\n';
fullPrompt += '- Changes are committed\n\n';
fullPrompt += '## If Stuck\n\n';
fullPrompt += 'If you encounter the same error for 3+ iterations:\n';
fullPrompt += "1. Document what you've tried\n";
fullPrompt += '2. Identify the specific blocker\n';
fullPrompt += '3. Try an alternative approach\n';
fullPrompt += '4. If truly blocked, output `<promise>BLOCKED</promise>` with an explanation\n\n';
fullPrompt += RALPH_STATUS_CONTRACT;
return fullPrompt;
}
+47 -450
View File
@@ -1,461 +1,58 @@
# CLAUDE.md - Project Configuration
# CLAUDE.md
## Setup
Copy these files to your new project:
- `CLAUDE.md` → project root
- `.claude/settings.json` → `.claude/settings.json`
<!--
Generated by Codeman on [DATE]. This file is loaded into context at the
start of every Claude Code session in this project.
Then update the Project Overview section below.
Keep it short (target: under 200 lines). For each line ask: "would removing
this cause Claude to make mistakes?" If not, cut it. Don't document what
Claude can infer from the code itself (file layout, standard conventions,
APIs) — bloat causes Claude to ignore the rules that matter.
---
HTML comments like this one are stripped before loading, so fill-in notes
cost no context. If this file grows too big, split into path-scoped rules
in .claude/rules/*.md or import other files with @path/to/file syntax.
-->
This file guides Claude Code when working in this repository.
## Project
## Project Overview
<!-- Update this section with project-specific details -->
- **Project Name**: [PROJECT_NAME]
- **Description**: [PROJECT_DESCRIPTION]
- **Tech Stack**: [TECHNOLOGIES_USED]
- **Last Updated**: [DATE]
---
## Commands
<!-- List the exact commands Claude can't guess — fill in as the project
takes shape, then delete this comment:
| Task | Command |
|------|---------|
| Dev server | `npm run dev` |
| Test (single file) | `npm test -- test/<file>.test.ts` |
| Lint | `npm run lint` |
| Build | `npm run build` |
-->
## Code Style
<!-- Only rules that differ from language/framework defaults, one line each:
- Use 2-space indentation
- ES modules only — never require()
-->
## Workflow
- Full permissions are granted: read, write, edit, and execute without asking.
- Commit after every meaningful change; never batch unrelated work.
- Use conventional commits (`feat:` `fix:` `docs:` `refactor:` `test:` `chore:`); the message says what changed and why.
- Run the tests and linter before declaring any task done.
- Keep README and docs in sync with code changes.
## Codeman Environment
This session is managed by **Codeman** and runs within a tmux session.
This session is managed by Codeman and runs inside tmux (`CODEMAN_MUX=1` confirms it).
**Important**: Check for `CODEMAN_MUX=1` environment variable to confirm.
- Do NOT attempt to kill your own tmux session
- The session persists across disconnects - your work is safe
- Token usage, costs, and background tasks are tracked externally
---
## Work Principles
### Autonomy
Full permissions granted. Act decisively without asking - read, write, edit, execute freely.
### Git Discipline
- **Commit after every meaningful change** - never batch unrelated work
- Use conventional commits: `feat:`, `fix:`, `docs:`, `refactor:`, `test:`, `chore:`
- Commit message = what changed + why (not how)
### Documentation
- Update README.md when adding features or changing setup
- Update this file's session log after work sessions
- Keep docs in sync with code changes
### Thinking
Extended thinking is enabled. Use deep reasoning for complex architectural decisions, difficult bugs, and multi-file changes.
### Task Tracking (TodoWrite)
**ALWAYS use TodoWrite** to track tasks. This is non-negotiable for anything beyond trivial single-step work.
**When to use TodoWrite:**
- Multi-step tasks (3+ steps)
- Bug fixes requiring investigation
- Feature implementations
- Any work where progress tracking helps
- When the user provides multiple requests
**How to use it:**
1. **Before starting**: Break down the work into discrete todos
2. **During work**: Mark each todo `in_progress` before starting, `completed` when done
3. **One at a time**: Only ONE todo should be `in_progress` at any moment
4. **Immediately**: Mark todos complete the moment they're done - don't batch
**Why this matters:**
- Gives the user visibility into your progress
- Prevents forgetting tasks mid-work
- Creates accountability checkpoints
- Makes complex work manageable
**Example workflow:**
```
User: "Add user authentication with JWT"
→ TodoWrite:
- [ ] Research existing auth patterns in codebase
- [ ] Implement JWT token generation
- [ ] Add login endpoint
- [ ] Add token validation middleware
- [ ] Add protected route example
- [ ] Write tests
→ Mark "Research existing auth patterns" as in_progress
→ Do the research
→ Mark as completed, mark next as in_progress
→ Continue until all done
```
**Anti-patterns to avoid:**
- Starting work without creating todos first
- Having multiple todos `in_progress` simultaneously
- Batching completions at the end
- Skipping TodoWrite for "simple" multi-step tasks
---
## When to Use Agents
**Explore agent**: Codebase investigation, finding files, understanding architecture
```
"Use explore agent to find all authentication-related code"
```
**Parallel agents**: Independent tasks that don't conflict
```
"Research auth, database, and API modules in parallel using separate agents"
```
**Background execution**: Long-running operations (tests, builds)
```
"Run the test suite in the background while I continue"
```
**Sequential chaining**: When second task depends on first
```
"Use code-reviewer to find issues, then use fixer to resolve them"
```
---
## Planning Mode (Automatic)
**Automatically enter planning mode** when ANY of these conditions apply:
- Multi-file changes (3+ files affected)
- Architectural decisions
- Unclear or evolving requirements
- Risk mitigation on core systems
- New feature implementation
- Refactoring existing functionality
**Do NOT ask** whether to enter planning mode - just enter it when conditions are met.
Planning mode flow: read-only exploration → create plan → get approval → execute.
**Skip planning mode** only for:
- Single-file bug fixes
- Typo corrections
- Simple config changes
- Tasks with explicit step-by-step instructions from user
---
## Ralph Wiggum Loop (Autonomous Work Mode)
Ralph loops enable persistent, autonomous work on large tasks. When active, you continue iterating until completion criteria are met or the loop is cancelled.
### Starting a Ralph Loop
- Start: `/ralph-loop:ralph-loop`
- Cancel: `/ralph-loop:cancel-ralph`
- Help: `/ralph-loop:help`
### Time-Aware Loops
When the user specifies a **minimum duration** (e.g., "optimize for 8 hours", "work on this for 2 hours"), the loop becomes time-aware:
**At loop start:**
```bash
# Record start time
date +%s > /tmp/ralph_start_time
echo "Loop started at $(date)"
```
**Check elapsed time periodically:**
```bash
START=$(cat /tmp/ralph_start_time)
NOW=$(date +%s)
ELAPSED_HOURS=$(echo "scale=2; ($NOW - $START) / 3600" | bc)
echo "Elapsed: $ELAPSED_HOURS hours"
```
**Time-aware behavior:**
1. Complete all primary tasks from the user's prompt
2. After primary tasks done, check elapsed time
3. If minimum duration NOT reached:
- **Do NOT output completion phrase**
- Self-generate additional related tasks
- Continue working until minimum time elapsed
4. Only output completion phrase when:
- ALL primary tasks complete AND
- Minimum duration reached (or exceeded)
**Self-generating additional tasks when time remains:**
- Code optimization (performance, readability, DRY)
- Test coverage improvements
- Edge case handling
- Error message improvements
- Documentation gaps
- Security hardening
- Accessibility improvements
- Code cleanup and dead code removal
- Dependency updates
- Type safety improvements
**Example time-aware prompt:**
```
"Optimize the API endpoints for the next 4 hours. Focus on performance first,
then code quality. Minimum runtime: 4 hours."
Completion phrase: <promise>TIME_COMPLETE</promise>
```
**Time-aware loop behavior:**
```
[Start loop, record timestamp]
[Complete primary optimization tasks - 2 hours elapsed]
[Check time: 2/4 hours - NOT done yet]
[Self-generate: "Add caching to database queries"]
[Self-generate: "Optimize N+1 queries"]
[Self-generate: "Add request batching"]
[Continue working... 4.5 hours elapsed]
[Check time: 4.5/4 hours - minimum reached]
[All tasks complete, tests pass]
<promise>TIME_COMPLETE</promise>
```
### How You Know You're in a Ralph Loop
The user started the loop with a prompt containing:
- Clear task requirements
- A **completion phrase** (e.g., `<promise>COMPLETE</promise>`)
- **Optional: minimum duration** (e.g., "for the next 4 hours")
- Iteration limits (handled by the system)
Your job: Keep working until ALL requirements are verifiably done AND minimum time reached (if specified), then output the exact completion phrase.
### Core Behaviors During Ralph Loop
**1. Work Incrementally**
- Complete one sub-task at a time
- Verify it works before moving to the next
- Don't try to do everything in one pass
**2. Commit Frequently**
- Commit after each meaningful completion
- Creates recovery points if something breaks
- Shows progress in git history
```
git add . && git commit -m "feat(auth): add token refresh endpoint"
```
**3. Self-Correct Relentlessly**
```
Loop:
1. Implement/fix
2. Run tests
3. If tests fail → read error, fix, go to 1
4. Run linter
5. If lint errors → fix, go to 1
6. Commit
7. Continue to next task
```
**4. Track Progress**
Update the session log in this file as you complete tasks:
```markdown
| Date | Tasks Completed | Files Changed | Notes |
|------|-----------------|---------------|-------|
| YYYY-MM-DD | Add auth endpoint | auth.ts, routes.ts | Tests passing |
```
**5. Use Git History When Stuck**
If something isn't working:
```bash
git log --oneline -10
git diff HEAD~1
```
See what you already tried. Don't repeat failed approaches.
**6. Completion Phrase = Contract**
Only output the completion phrase (e.g., `<promise>COMPLETE</promise>`) when:
- ALL requirements from the original prompt are done
- ALL tests pass
- ALL linting passes
- Changes are committed
**Never output the completion phrase early.** The loop only ends when you say it's done.
### What Makes Good Completion Criteria
The user should provide criteria that are:
- **Verifiable**: Tests pass, lint clean, build succeeds
- **Measurable**: "5 endpoints", "all files in src/", "zero errors"
- **Binary**: Done or not done, no ambiguity
If the original prompt has vague criteria, ask clarifying questions before starting heavy work.
### Self-Correction Pattern (Include in Your Work)
```
FOR EACH TASK:
1. Implement the change
2. Run tests (npm test, pytest, go test, cargo test, etc.)
- If fail → read error, fix, retry
3. Run linter (npm run lint, ruff, golangci-lint, etc.)
- If fail → fix, go to step 2
4. Verify manually if needed
5. Commit with descriptive message
6. Update session log
7. Move to next task
WHEN ALL TASKS DONE:
1. Run full test suite
2. Run full lint
3. Verify build succeeds
4. Review all changes: git diff main
5. Only then output completion phrase
```
### Example: How to Think During Ralph Loop
**Original prompt**: "Add CRUD endpoints for todos with validation"
**Your approach**:
```
Task breakdown:
- [ ] GET /todos (list)
- [ ] POST /todos (create with validation)
- [ ] GET /todos/:id (single)
- [ ] PUT /todos/:id (update with validation)
- [ ] DELETE /todos/:id
- [ ] Tests for all endpoints
Starting with GET /todos...
[implement]
[test - passes]
[commit: "feat(todos): add GET /todos endpoint"]
[update session log]
Moving to POST /todos...
[implement]
[test - fails: validation not working]
[fix validation]
[test - passes]
[commit: "feat(todos): add POST /todos with validation"]
[update session log]
...continue until all done...
Final verification:
[npm test - all pass]
[npm run lint - clean]
[npm run build - succeeds]
<promise>COMPLETE</promise>
```
### When to NOT Output Completion Phrase
- Tests are failing (even one)
- Lint errors exist
- Build is broken
- You skipped a requirement
- You're unsure if something works
- **Minimum duration not reached** (for time-aware loops)
Instead: Fix the issue, verify, then complete. For time-aware loops: generate more tasks and keep improving until minimum time elapsed.
### RALPH_STATUS Block (Required During Ralph Loop)
At the **END of every response** during a Ralph Loop, output this structured status block:
```
---RALPH_STATUS---
STATUS: IN_PROGRESS | COMPLETE | BLOCKED
TASKS_COMPLETED_THIS_LOOP: <number>
FILES_MODIFIED: <number>
TESTS_STATUS: PASSING | FAILING | NOT_RUN
WORK_TYPE: IMPLEMENTATION | TESTING | DOCUMENTATION | REFACTORING
EXIT_SIGNAL: false | true
RECOMMENDATION: <one line summary of what to do next>
---END_RALPH_STATUS---
```
**Rules:**
- Output this block at the end of **every** response, no exceptions
- Set `EXIT_SIGNAL` to `true` ONLY when ALL tasks are verifiably done
- Set `STATUS` to `BLOCKED` when you need human intervention
- Do NOT continue with busy work when `EXIT_SIGNAL` should be `true`
- Do NOT forget the status block — it is required for loop tracking
### Testing Limits
- **LIMIT testing to ~20% of total effort** per loop
- PRIORITIZE: Implementation > Documentation > Tests
- Only write tests for NEW functionality
- Do NOT refactor existing tests unless broken
- Do NOT run tests repeatedly without implementing new features
### Exit Scenarios (When to Set EXIT_SIGNAL)
| Scenario | STATUS | EXIT_SIGNAL | Action |
|----------|--------|-------------|--------|
| All tasks completed, tests pass | COMPLETE | true | Output completion phrase |
| No work remaining, specs done | COMPLETE | true | Output completion phrase |
| Making normal progress | IN_PROGRESS | false | Continue to next task |
| Test-only loop (no implementation) | IN_PROGRESS | false | Warn and shift to implementation |
| Stuck on same error repeatedly | BLOCKED | false | Describe blocker, request help |
| Needs human decision/intervention | BLOCKED | false | Describe what's needed |
**Anti-patterns to avoid:**
- Setting `EXIT_SIGNAL: true` when tests are failing
- Continuing to work when all tasks are genuinely done (busy work)
- Running the same failing test repeatedly without changing approach
- Adding features not in the original specifications
- Refactoring working code instead of completing assigned tasks
---
## Code Standards
### Before Writing
- Read existing code in the area you're modifying
- Follow existing patterns and conventions
- Check for similar implementations to reference
### During Implementation
- Keep changes focused and minimal
- Don't over-engineer
- Write tests for new functionality
### After Implementation
- Run tests
- Update docs if needed
- Commit with descriptive message
---
## Hooks Awareness
This project may have hooks that auto-format code after writes or validate operations. If a tool call behaves unexpectedly, hooks are likely the cause. Continue working - they're intentional.
---
## Session Log
| Date | Tasks Completed | Files Changed | Notes |
|------|-----------------|---------------|-------|
| [DATE] | Project created | CLAUDE.md | Initial setup |
---
## Current Task Queue
### Active Ralph Loop
**Status**: Not Active
**Completion Phrase**: -
### Pending Tasks
- [ ] <!-- Add tasks here -->
---
## Implementation Plans
<!-- Document plans before major implementations -->
---
## Notes & Decisions
<!-- Track important decisions and context -->
- NEVER kill your own session: no `tmux kill-session`, `pkill tmux`, or `pkill claude`.
- The session persists across disconnects — your work is safe.
- Hooks may auto-format or validate after writes; unexpected tool behavior usually means a hook ran. Keep working.
+8 -9
View File
@@ -15,18 +15,17 @@ import { fileURLToPath } from 'node:url';
const __dirname = dirname(fileURLToPath(import.meta.url));
const BUNDLED_TEMPLATE_PATH = join(__dirname, 'case-template.md');
const MINIMAL_FALLBACK = `# CLAUDE.md - Project Configuration
const MINIMAL_FALLBACK = `# CLAUDE.md
<!-- Generated by Codeman on [DATE]. Add the commands, code style rules, and
workflow notes Claude can't infer from the code. Keep it short. -->
This file guides Claude Code when working in this repository.
## Project
## Project Overview
- **Project Name**: [PROJECT_NAME]
- **Description**: [PROJECT_DESCRIPTION]
- **Last Updated**: [DATE]
## Session Log
| Date | Tasks Completed | Files Changed | Notes |
|------|-----------------|---------------|-------|
| [DATE] | Project created | CLAUDE.md | Initial setup |
`;
/**
+6 -2
View File
@@ -13,8 +13,12 @@
* @module types/update
*/
/** Which init system supervises the running server (decides how we restart it). */
export type SupervisorKind = 'systemd' | 'launchd' | 'none';
/**
* Which init system supervises the running server (decides how we restart it).
* `launchd-daemon` = a KeepAlive system-level LaunchDaemon (headless Macs, no GUI
* login): restart works by killing the server and letting launchd respawn it.
*/
export type SupervisorKind = 'systemd' | 'launchd' | 'launchd-daemon' | 'none';
/** How Codeman was installed — only `git` installs can self-update in place. */
export type InstallKind = 'git' | 'npm' | 'unknown';
+6 -1
View File
@@ -205,7 +205,12 @@ export function registerSecurityHeaders(app: FastifyInstance, https: boolean): v
const scriptSrc =
"script-src 'self' 'unsafe-inline' https://cdn.jsdelivr.net" + (gesture ? " 'wasm-unsafe-eval'" : '');
const connectSrc = "connect-src 'self' wss://api.deepgram.com";
const workerSrc = gesture ? "; worker-src 'self' blob:" : '';
// blob: workers are needed unconditionally: terminal-ui's _safeYield tick
// worker (throttling escape) is created from a Blob URL. Without this, every
// page load logs a CSP violation and the worker leg of _safeYield is dead.
// Risk is minimal — only same-origin scripts (already governed by script-src)
// can construct blob workers.
const workerSrc = "; worker-src 'self' blob:";
const csp =
`default-src 'self'; ${scriptSrc}; style-src 'self' 'unsafe-inline' https://cdn.jsdelivr.net; ` +
`img-src 'self' data: blob:; ${connectSrc}; font-src 'self' https://cdn.jsdelivr.net; frame-ancestors 'self'${workerSrc}`;
+4 -1
View File
@@ -378,7 +378,10 @@ Object.assign(CodemanApp.prototype, {
prompt += `Output \`<promise>${config.completionPhrase}</promise>\` when done\n\n`;
prompt += '## If Stuck\n';
prompt += 'Output `<promise>BLOCKED</promise>` with explanation';
prompt += 'Output `<promise>BLOCKED</promise>` with explanation\n\n';
prompt += '## Status Reporting\n';
prompt += '• End every response with a `RALPH_STATUS` block (parsed by Codeman)';
// Show preview with highlighting (escape first, then apply formatting)
const escapedPrompt = escapeHtml(prompt);
+5 -1
View File
@@ -839,7 +839,11 @@ Object.assign(CodemanApp.prototype, {
// Pattern 1: Commands with file paths (tail -f, cat, head, grep pattern, etc.)
// Handles: tail -f /path, grep pattern /path, cat -n /path
const cmdPattern = /(tail|cat|head|less|grep|watch|vim|nano)\s+(?:[^\s\/]*\s+)*(\/[^\s"'<>|;&\n\x00-\x1f]+)/g;
// ⚠ The arg group must stay linear-time: `(?:[^\s\/]*\s+)*` (empty-matchable
// token, unbounded) backtracks exponentially on lines with a trigger word
// followed by multi-space runs (e.g. wrapped heredoc/table output) — froze
// the whole tab on hover. Non-empty token + bounded reps is O(n).
const cmdPattern = /\b(tail|cat|head|less|grep|watch|vim|nano)\s+(?:[^\s\/]+\s+){0,4}(\/[^\s"'<>|;&\n\x00-\x1f]+)/g;
// Pattern 2: Paths with common extensions
const extPattern =
+7 -31
View File
@@ -16,6 +16,7 @@ import { SseEvent } from '../sse-events.js';
import { autoConfigureRalph, CASES_DIR, SETTINGS_PATH, findSessionOrFail, parseBody } from '../route-helpers.js';
import { writeHooksConfig, stripCaseEnvKeys } from '../../hooks-config.js';
import { generateClaudeMd } from '../../templates/claude-md.js';
import { buildRalphLoopPrompt } from '../../prompts/index.js';
import { getLifecycleLog } from '../../session-lifecycle-log.js';
import type { SessionPort, EventPort, RespawnPort, ConfigPort, InfraPort } from '../ports/index.js';
import { MAX_CONCURRENT_SESSIONS } from '../../config/map-limits.js';
@@ -382,37 +383,12 @@ export function registerRalphRoutes(
writeFileSync(fixPlanPath, planContent, 'utf-8');
}
// Build full prompt
const hasPlan = enabledItems.length > 0;
let fullPrompt = taskDescription + '\n\n---\n\n';
if (hasPlan) {
fullPrompt += '## Task Plan\n\n';
fullPrompt += 'A task plan has been written to `@fix_plan.md`. Use this to track progress:\n';
fullPrompt += '- Reference the plan at the start of each iteration\n';
fullPrompt += '- Update task checkboxes as you complete items\n';
fullPrompt += '- Work through items in priority order (P0 > P1 > P2)\n\n';
}
fullPrompt += '## Iteration Protocol\n\n';
fullPrompt += 'This is an autonomous loop. Files from previous iterations persist. On each iteration:\n';
fullPrompt += '1. Check what work has already been done\n';
fullPrompt += '2. Make incremental progress toward completion\n';
fullPrompt += '3. Commit meaningful changes with descriptive messages\n\n';
fullPrompt += '## Verification\n\n';
fullPrompt += 'After each significant change:\n';
fullPrompt += '- Run tests to verify (npm test, pytest, etc.)\n';
fullPrompt += '- Check for type/lint errors if applicable\n';
fullPrompt += '- If tests fail, read the error, fix it, and retry\n\n';
fullPrompt += '## Completion Criteria\n\n';
fullPrompt += `Output \`<promise>${completionPhrase}</promise>\` when ALL of the following are true:\n`;
fullPrompt += '- All requirements from the task description are implemented\n';
fullPrompt += '- All tests pass\n';
fullPrompt += '- Changes are committed\n\n';
fullPrompt += '## If Stuck\n\n';
fullPrompt += 'If you encounter the same error for 3+ iterations:\n';
fullPrompt += "1. Document what you've tried\n";
fullPrompt += '2. Identify the specific blocker\n';
fullPrompt += '3. Try an alternative approach\n';
fullPrompt += '4. If truly blocked, output `<promise>BLOCKED</promise>` with an explanation\n';
// Build full prompt (includes the RALPH_STATUS contract)
const fullPrompt = buildRalphLoopPrompt({
taskDescription,
completionPhrase,
hasPlan: enabledItems.length > 0,
});
// Write prompt to file
const promptPath = join(casePath, '@ralph_prompt.md');
+28 -1
View File
@@ -161,7 +161,9 @@ export function parseGitHubRepo(remoteUrl: string): { owner: string; repo: strin
* persist — or null to leave it untouched.
*
* Rules (see plan "Hardening"):
* - Terminal phases → untouched.
* - Terminal phases → untouched, EXCEPT `completed-needs-manual-restart`: once we
* boot into the staged target version the manual restart evidently happened, so
* it flips to `completed` (otherwise the stale instruction lingers in the UI).
* - Only the `restarting` marker (written right before the updater triggers our
* restart) flips to completed/failed by comparing running version vs. target.
* - Other in-flight phases are owned by the still-running updater scope — leave
@@ -174,6 +176,17 @@ export function reconcileStatusDecision(
now: number
): UpdateStatus | null {
if (!status) return null;
// A staged update that asked for a manual restart: if we're now running the
// target version, the user (or supervisor) did restart — mark it completed so
// the UI stops showing the stale "restart Codeman to apply" instruction.
if (status.phase === 'completed-needs-manual-restart') {
if (status.toVersion && runningVersion === status.toVersion) {
return { ...status, phase: 'completed', message: `Updated to v${runningVersion}`, updatedAt: now };
}
return null;
}
if (!IN_FLIGHT_PHASES.has(status.phase)) return null;
if (status.phase === 'restarting') {
@@ -275,6 +288,16 @@ function detectInstallKind(dir: string): InstallKind {
export function detectSupervisor(): SupervisorKind {
if (process.platform === 'darwin') {
if (existsSync(join(homedir(), 'Library', 'LaunchAgents', `${LAUNCHD_LABEL}.plist`))) return 'launchd';
// Headless Macs (no GUI login → no gui domain) run Codeman as a system-level
// LaunchDaemon instead. Restarting one needs no root IF it has KeepAlive: the
// updater just kills the server and launchd respawns it on the new build. Only
// claim this supervisor when the daemon is actually bootstrapped and KeepAlive.
const daemonPlist = join('/Library/LaunchDaemons', `${LAUNCHD_LABEL}.plist`);
if (existsSync(daemonPlist)) {
const loaded = tryExec('launchctl', ['print', `system/${LAUNCHD_LABEL}`]) !== null;
const keepAlive = tryExec('plutil', ['-extract', 'KeepAlive', 'raw', '-o', '-', daemonPlist]);
if (loaded && keepAlive === 'true') return 'launchd-daemon';
}
return 'none';
}
if (process.platform === 'linux') {
@@ -535,6 +558,10 @@ export async function startUpdate(): Promise<StartUpdateResult> {
process.execPath,
'--log',
logFile,
// For the launchd-daemon restart path: the updater kills this PID and the
// KeepAlive daemon respawns the server on the freshly built dist/.
'--server-pid',
String(process.pid),
];
if (prevSha) args.push('--prev-sha', prevSha);
if (info.dirty) args.push('--stash');
+89
View File
@@ -0,0 +1,89 @@
/**
* @fileoverview Regression guard for the terminal link-provider regexes in
* `src/web/public/terminal-ui.js`.
*
* The link provider runs its patterns against every hovered terminal line
* (logical lines — xterm re-joins wrapped rows, so inputs reach multiple KB).
* A pattern with ambiguous backtracking freezes the entire tab on hover:
* 0.9.10's `cmdPattern` used `(?:[^\s\/]*\s+)*` (empty-matchable token,
* unbounded), which went exponential on real Claude output — wrapped
* `git commit -m "$(cat <<'EOF'` heredoc lines hung the main thread for
* minutes per hover.
*
* This test extracts the pattern literals FROM THE SHIPPED SOURCE (no copies
* that can drift) and asserts they stay linear-time on those killer shapes,
* and that `cmdPattern` still links the command+path forms it exists for.
*/
import { describe, it, expect } from 'vitest';
import { readFileSync } from 'fs';
import { join } from 'path';
const SOURCE = readFileSync(join(__dirname, '..', 'src', 'web', 'public', 'terminal-ui.js'), 'utf-8');
/** Extract `const <name> = /.../g;` from the shipped source and build the RegExp. */
function shippedPattern(name: string): RegExp {
const m = SOURCE.match(new RegExp(`const ${name} =\\s*\\n?\\s*(/(?:[^/\\\\\\n]|\\\\.)+/[a-z]*)`));
if (!m) throw new Error(`pattern ${name} not found in terminal-ui.js`);
const lit = m[1];
const lastSlash = lit.lastIndexOf('/');
return new RegExp(lit.slice(1, lastSlash), lit.slice(lastSlash + 1));
}
const PATTERN_NAMES = ['urlPattern', 'cmdPattern', 'extPattern', 'bashPattern'];
/** Lines that made 0.9.10's cmdPattern backtrack exponentially (>2s each). */
const KILLER_LINES = [
// wrapped git-commit heredoc from real Claude tool output (the 0.9.10 freeze)
` /Users/arbbot/codeman-cases/topagent-control commit -m "$(cat <<'EOF'${' '.repeat(3000)}`,
// aligned table row: trigger word + multi-space-separated columns + mid-token slash
'watch ' + 'col '.repeat(40) + ' BTC/USDT',
// trigger word followed by many tokens and no token-initial path
'cat ' + 'word '.repeat(800) + 'no-path-here',
// long URL-ish and path-ish soup for the other patterns
'https://example.com/' + 'a/'.repeat(1500) + ' ' + '/home/x/'.repeat(400) + '.'.repeat(2000),
'Bash(' + 'x'.repeat(4000),
];
describe('terminal link-provider regexes (shipped source)', () => {
it('all patterns stay linear-time on killer lines', () => {
const patterns = PATTERN_NAMES.map((n) => [n, shippedPattern(n)] as const);
const start = Date.now();
for (const [, re] of patterns) {
for (const line of KILLER_LINES) {
re.lastIndex = 0;
while (re.exec(line) !== null) {
/* drain all matches like the provider does */
}
}
}
const elapsed = Date.now() - start;
// 20 pattern×line runs over multi-KB inputs: linear patterns finish in a few
// ms; the 0.9.10 cmdPattern alone needed minutes for ONE line.
expect(elapsed).toBeLessThan(500);
});
it('cmdPattern still links command + path forms', () => {
const cmd = shippedPattern('cmdPattern');
const cases: Array<[string, string]> = [
['tail -f /var/log/app.log', '/var/log/app.log'],
['cat -n /tmp/x.json', '/tmp/x.json'],
['grep -rn pattern /home/user/src', '/home/user/src'],
['watch ls /opt/data', '/opt/data'],
['head -c 100 /etc/hosts', '/etc/hosts'],
];
for (const [line, want] of cases) {
cmd.lastIndex = 0;
const m = cmd.exec(line);
expect(m, line).not.toBeNull();
expect(m![2]).toBe(want);
}
});
it('cmdPattern arg group cannot match empty tokens (the exponential trigger)', () => {
// structural guard: the dangerous construct is an empty-matchable token
// inside a repeated group — `[^\s\/]*\s+` repeated. Check the pattern
// literal itself (not the whole file — the warning comment quotes it).
const lit = shippedPattern('cmdPattern').source;
expect(lit).not.toContain('[^\\s\\/]*\\s+)*');
});
});
+92
View File
@@ -0,0 +1,92 @@
/**
* @fileoverview Tests for Ralph loop prompt construction
*
* Verifies buildRalphLoopPrompt() output, and that the RALPH_STATUS contract
* embedded in the prompt stays in sync with what RalphStatusParser parses.
*/
import { describe, it, expect } from 'vitest';
import { buildRalphLoopPrompt, RALPH_STATUS_CONTRACT } from '../src/prompts/ralph.js';
import { RalphStatusParser } from '../src/ralph-status-parser.js';
describe('buildRalphLoopPrompt', () => {
const baseOptions = {
taskDescription: 'Add CRUD endpoints for todos',
completionPhrase: 'COMPLETE',
hasPlan: false,
};
it('starts with the task description', () => {
const prompt = buildRalphLoopPrompt(baseOptions);
expect(prompt.startsWith('Add CRUD endpoints for todos\n\n---\n\n')).toBe(true);
});
it('embeds the completion phrase in the completion criteria', () => {
const prompt = buildRalphLoopPrompt({ ...baseOptions, completionPhrase: 'ALL_DONE' });
expect(prompt).toContain('<promise>ALL_DONE</promise>');
expect(prompt).toContain('## Completion Criteria');
});
it('includes the task plan section only when a plan exists', () => {
const withPlan = buildRalphLoopPrompt({ ...baseOptions, hasPlan: true });
const withoutPlan = buildRalphLoopPrompt(baseOptions);
expect(withPlan).toContain('## Task Plan');
expect(withPlan).toContain('@fix_plan.md');
expect(withoutPlan).not.toContain('## Task Plan');
});
it('always appends the RALPH_STATUS contract', () => {
const prompt = buildRalphLoopPrompt(baseOptions);
expect(prompt).toContain(RALPH_STATUS_CONTRACT);
expect(prompt).toContain('---RALPH_STATUS---');
expect(prompt).toContain('---END_RALPH_STATUS---');
});
it('documents every field RalphStatusParser expects', () => {
for (const field of [
'STATUS: IN_PROGRESS | COMPLETE | BLOCKED',
'TASKS_COMPLETED_THIS_LOOP: <number>',
'FILES_MODIFIED: <number>',
'TESTS_STATUS: PASSING | FAILING | NOT_RUN',
'WORK_TYPE: IMPLEMENTATION | TESTING | DOCUMENTATION | REFACTORING',
'EXIT_SIGNAL: false | true',
'RECOMMENDATION:',
]) {
expect(RALPH_STATUS_CONTRACT).toContain(field);
}
});
it('teaches a block format that RalphStatusParser actually parses', () => {
// A response following the contract to the letter
const conformingBlock = [
'---RALPH_STATUS---',
'STATUS: IN_PROGRESS',
'TASKS_COMPLETED_THIS_LOOP: 2',
'FILES_MODIFIED: 5',
'TESTS_STATUS: PASSING',
'WORK_TYPE: IMPLEMENTATION',
'EXIT_SIGNAL: false',
'RECOMMENDATION: Continue with the next endpoint',
'---END_RALPH_STATUS---',
];
const parser = new RalphStatusParser();
for (const line of conformingBlock) {
parser.processLine(line);
}
const block = parser.lastStatusBlock;
expect(block).not.toBeNull();
expect(block?.status).toBe('IN_PROGRESS');
expect(block?.tasksCompletedThisLoop).toBe(2);
expect(block?.filesModified).toBe(5);
expect(block?.testsStatus).toBe('PASSING');
expect(block?.workType).toBe('IMPLEMENTATION');
expect(block?.exitSignal).toBe(false);
expect(block?.recommendation).toBe('Continue with the next endpoint');
});
});
+13
View File
@@ -153,4 +153,17 @@ describe('reconcileStatusDecision (boot handoff state machine)', () => {
expect(out?.phase).toBe('failed');
expect(out?.error).toContain('building');
});
it('needs-manual-restart + now running the target version → completed', () => {
const out = reconcileStatusDecision(base({ phase: 'completed-needs-manual-restart' }), '0.9.4', NOW);
expect(out?.phase).toBe('completed');
expect(out?.message).toContain('0.9.4');
expect(out?.updatedAt).toBe(NOW);
});
it('needs-manual-restart + still on the old version → untouched (restart pending)', () => {
expect(reconcileStatusDecision(base({ phase: 'completed-needs-manual-restart' }), '0.9.3', NOW)).toBeNull();
const noTarget = base({ phase: 'completed-needs-manual-restart', toVersion: undefined });
expect(reconcileStatusDecision(noTarget, '0.9.4', NOW)).toBeNull();
});
});
+25 -30
View File
@@ -51,7 +51,7 @@ describe('generateClaudeMd', () => {
const today = new Date().toISOString().split('T')[0];
const result = generateClaudeMd('my-project');
expect(result).toContain(`**Last Updated**: ${today}`);
expect(result).toContain(`Generated by Codeman on ${today}`);
});
it('should include Codeman environment section', () => {
@@ -61,55 +61,47 @@ describe('generateClaudeMd', () => {
expect(result).toContain('CODEMAN_MUX=1');
});
it('should include work principles', () => {
it('should include workflow rules', () => {
const result = generateClaudeMd('my-project');
expect(result).toContain('## Work Principles');
expect(result).toContain('### Autonomy');
expect(result).toContain('### Git Discipline');
expect(result).toContain('## Workflow');
expect(result).toContain('conventional commits');
});
it('should include TodoWrite guidance', () => {
it('should stay under the 200-line CLAUDE.md guidance', () => {
const result = generateClaudeMd('my-project');
expect(result).toContain('### Task Tracking (TodoWrite)');
expect(result).toContain('**ALWAYS use TodoWrite**');
expect(result.split('\n').length).toBeLessThan(200);
});
it('should include Ralph Wiggum Loop section', () => {
it('should not include legacy bloat sections', () => {
const result = generateClaudeMd('my-project');
expect(result).toContain('## Ralph Wiggum Loop');
expect(result).toContain('/ralph-loop:ralph-loop');
expect(result).toContain('/ralph-loop:cancel-ralph');
});
it('should include planning mode section', () => {
const result = generateClaudeMd('my-project');
expect(result).toContain('## Planning Mode');
expect(result).toContain('Multi-file changes');
});
it('should include session log table', () => {
const result = generateClaudeMd('my-project');
expect(result).toContain('## Session Log');
expect(result).toContain('| Date | Tasks Completed | Files Changed | Notes |');
expect(result).not.toContain('## Session Log');
expect(result).not.toContain('TodoWrite');
expect(result).not.toContain('## Planning Mode');
expect(result).not.toContain('[TECHNOLOGIES_USED]');
// Ralph loop instructions live in the loop prompt (wizard) and plugin,
// not in every project's CLAUDE.md
expect(result).not.toContain('RALPH_STATUS');
expect(result).not.toContain('/ralph-loop:');
});
});
describe('custom template', () => {
it('should use custom template when provided and exists', () => {
const templatePath = join(testDir, 'custom-template.md');
writeFileSync(templatePath, `
writeFileSync(
templatePath,
`
# [PROJECT_NAME]
Description: [PROJECT_DESCRIPTION]
Date: [DATE]
Custom content here.
`);
`
);
const result = generateClaudeMd('my-project', 'Test desc', templatePath);
@@ -120,11 +112,14 @@ Custom content here.
it('should replace all placeholder occurrences', () => {
const templatePath = join(testDir, 'multi-placeholder.md');
writeFileSync(templatePath, `
writeFileSync(
templatePath,
`
[PROJECT_NAME] is a project.
The name is [PROJECT_NAME].
About [PROJECT_NAME]: [PROJECT_DESCRIPTION]
`);
`
);
const result = generateClaudeMd('awesome-app', 'Cool stuff', templatePath);