Compare commits

..
Author SHA1 Message Date
arkonandClaude Opus 4.6 7c3688467e fix: align CI node version to .nvmrc (22), remove redundant gotcha
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 13:15:54 +01:00
Noah Waldner 406c18cefe fix layout 2026-02-27 11:46:40 +01:00
Noah Waldner 25293a0349 format 2026-02-27 11:39:03 +01:00
Noah Waldner 1424373895 fix linting 2026-02-27 11:37:40 +01:00
Noah Waldner 8be7c9eb59 adjust formatting 2026-02-27 11:29:24 +01:00
Noah Waldner 8b072c6841 use node 20 2026-02-27 11:26:01 +01:00
Noah WaldnerandClaude Sonnet 4.6 7de939eb16 chore: simplify CI to single Node.js version (22)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-27 11:25:32 +01:00
Noah Waldner 7f40589ece update claude md 2026-02-27 11:23:24 +01:00
Noah Waldner 6da081f02b remove contributing.md 2026-02-27 11:18:05 +01:00
Noah WaldnerandClaude Sonnet 4.6 0af9053aff chore: merge master into cleanup/add-project-setup
Merged changeset scripts/devDeps from master with ESLint/Prettier
tooling added on this branch.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-27 11:01:56 +01:00
Noah Waldner 34c444f7b2 add tools to eslint 2026-02-27 10:54:19 +01:00
Noah Waldner fc6a74ab73 add project setup 2026-02-27 10:54:10 +01:00
1110 changed files with 47711 additions and 312057 deletions
@@ -0,0 +1,61 @@
---
name: remotion-best-practices
description: Best practices for Remotion - Video creation in React
metadata:
tags: remotion, video, react, animation, composition
---
## When to use
Use this skills whenever you are dealing with Remotion code to obtain the domain-specific knowledge.
## Captions
When dealing with captions or subtitles, load the [./rules/subtitles.md](./rules/subtitles.md) file for more information.
## Using FFmpeg
For some video operations, such as trimming videos or detecting silence, FFmpeg should be used. Load the [./rules/ffmpeg.md](./rules/ffmpeg.md) file for more information.
## Audio visualization
When needing to visualize audio (spectrum bars, waveforms, bass-reactive effects), load the [./rules/audio-visualization.md](./rules/audio-visualization.md) file for more information.
## Sound effects
When needing to use sound effects, load the [./rules/sound-effects.md](./rules/sound-effects.md) file for more information.
## How to use
Read individual rule files for detailed explanations and code examples:
- [rules/3d.md](rules/3d.md) - 3D content in Remotion using Three.js and React Three Fiber
- [rules/animations.md](rules/animations.md) - Fundamental animation skills for Remotion
- [rules/assets.md](rules/assets.md) - Importing images, videos, audio, and fonts into Remotion
- [rules/audio.md](rules/audio.md) - Using audio and sound in Remotion - importing, trimming, volume, speed, pitch
- [rules/calculate-metadata.md](rules/calculate-metadata.md) - Dynamically set composition duration, dimensions, and props
- [rules/can-decode.md](rules/can-decode.md) - Check if a video can be decoded by the browser using Mediabunny
- [rules/charts.md](rules/charts.md) - Chart and data visualization patterns for Remotion (bar, pie, line, stock charts)
- [rules/compositions.md](rules/compositions.md) - Defining compositions, stills, folders, default props and dynamic metadata
- [rules/extract-frames.md](rules/extract-frames.md) - Extract frames from videos at specific timestamps using Mediabunny
- [rules/fonts.md](rules/fonts.md) - Loading Google Fonts and local fonts in Remotion
- [rules/get-audio-duration.md](rules/get-audio-duration.md) - Getting the duration of an audio file in seconds with Mediabunny
- [rules/get-video-dimensions.md](rules/get-video-dimensions.md) - Getting the width and height of a video file with Mediabunny
- [rules/get-video-duration.md](rules/get-video-duration.md) - Getting the duration of a video file in seconds with Mediabunny
- [rules/gifs.md](rules/gifs.md) - Displaying GIFs synchronized with Remotion's timeline
- [rules/images.md](rules/images.md) - Embedding images in Remotion using the Img component
- [rules/light-leaks.md](rules/light-leaks.md) - Light leak overlay effects using @remotion/light-leaks
- [rules/lottie.md](rules/lottie.md) - Embedding Lottie animations in Remotion
- [rules/measuring-dom-nodes.md](rules/measuring-dom-nodes.md) - Measuring DOM element dimensions in Remotion
- [rules/measuring-text.md](rules/measuring-text.md) - Measuring text dimensions, fitting text to containers, and checking overflow
- [rules/sequencing.md](rules/sequencing.md) - Sequencing patterns for Remotion - delay, trim, limit duration of items
- [rules/tailwind.md](rules/tailwind.md) - Using TailwindCSS in Remotion
- [rules/text-animations.md](rules/text-animations.md) - Typography and text animation patterns for Remotion
- [rules/timing.md](rules/timing.md) - Interpolation curves in Remotion - linear, easing, spring animations
- [rules/transitions.md](rules/transitions.md) - Scene transition patterns for Remotion
- [rules/transparent-videos.md](rules/transparent-videos.md) - Rendering out a video with transparency
- [rules/trimming.md](rules/trimming.md) - Trimming patterns for Remotion - cut the beginning or end of animations
- [rules/videos.md](rules/videos.md) - Embedding videos in Remotion - trimming, volume, speed, looping, pitch
- [rules/parameters.md](rules/parameters.md) - Make a video parametrizable by adding a Zod schema
- [rules/maps.md](rules/maps.md) - Add a map using Mapbox and animate it
- [rules/voiceover.md](rules/voiceover.md) - Adding AI-generated voiceover to Remotion compositions using ElevenLabs TTS
@@ -0,0 +1,86 @@
---
name: 3d
description: 3D content in Remotion using Three.js and React Three Fiber.
metadata:
tags: 3d, three, threejs
---
# Using Three.js and React Three Fiber in Remotion
Follow React Three Fiber and Three.js best practices.
Only the following Remotion-specific rules need to be followed:
## Prerequisites
First, the `@remotion/three` package needs to be installed.
If it is not, use the following command:
```bash
npx remotion add @remotion/three # If project uses npm
bunx remotion add @remotion/three # If project uses bun
yarn remotion add @remotion/three # If project uses yarn
pnpm exec remotion add @remotion/three # If project uses pnpm
```
## Using ThreeCanvas
You MUST wrap 3D content in `<ThreeCanvas>` and include proper lighting.
`<ThreeCanvas>` MUST have a `width` and `height` prop.
```tsx
import { ThreeCanvas } from "@remotion/three";
import { useVideoConfig } from "remotion";
const { width, height } = useVideoConfig();
<ThreeCanvas width={width} height={height}>
<ambientLight intensity={0.4} />
<directionalLight position={[5, 5, 5]} intensity={0.8} />
<mesh>
<sphereGeometry args={[1, 32, 32]} />
<meshStandardMaterial color="red" />
</mesh>
</ThreeCanvas>;
```
## No animations not driven by `useCurrentFrame()`
Shaders, models etc MUST NOT animate by themselves.
No animations are allowed unless they are driven by `useCurrentFrame()`.
Otherwise, it will cause flickering during rendering.
Using `useFrame()` from `@react-three/fiber` is forbidden.
## Animate using `useCurrentFrame()`
Use `useCurrentFrame()` to perform animations.
```tsx
const frame = useCurrentFrame();
const rotationY = frame * 0.02;
<mesh rotation={[0, rotationY, 0]}>
<boxGeometry args={[2, 2, 2]} />
<meshStandardMaterial color="#4a9eff" />
</mesh>;
```
## Using `<Sequence>` inside `<ThreeCanvas>`
The `layout` prop of any `<Sequence>` inside a `<ThreeCanvas>` must be set to `none`.
```tsx
import { Sequence } from "remotion";
import { ThreeCanvas } from "@remotion/three";
const { width, height } = useVideoConfig();
<ThreeCanvas width={width} height={height}>
<Sequence layout="none">
<mesh>
<boxGeometry args={[2, 2, 2]} />
<meshStandardMaterial color="#4a9eff" />
</mesh>
</Sequence>
</ThreeCanvas>;
```
@@ -0,0 +1,27 @@
---
name: animations
description: Fundamental animation skills for Remotion
metadata:
tags: animations, transitions, frames, useCurrentFrame
---
All animations MUST be driven by the `useCurrentFrame()` hook.
Write animations in seconds and multiply them by the `fps` value from `useVideoConfig()`.
```tsx
import { useCurrentFrame } from "remotion";
export const FadeIn = () => {
const frame = useCurrentFrame();
const { fps } = useVideoConfig();
const opacity = interpolate(frame, [0, 2 * fps], [0, 1], {
extrapolateRight: "clamp",
});
return <div style={{ opacity }}>Hello World!</div>;
};
```
CSS transitions or animations are FORBIDDEN - they will not render correctly.
Tailwind animation class names are FORBIDDEN - they will not render correctly.
@@ -0,0 +1,78 @@
---
name: assets
description: Importing images, videos, audio, and fonts into Remotion
metadata:
tags: assets, staticFile, images, fonts, public
---
# Importing assets in Remotion
## The public folder
Place assets in the `public/` folder at your project root.
## Using staticFile()
You MUST use `staticFile()` to reference files from the `public/` folder:
```tsx
import { Img, staticFile } from "remotion";
export const MyComposition = () => {
return <Img src={staticFile("logo.png")} />;
};
```
The function returns an encoded URL that works correctly when deploying to subdirectories.
## Using with components
**Images:**
```tsx
import { Img, staticFile } from "remotion";
<Img src={staticFile("photo.png")} />;
```
**Videos:**
```tsx
import { Video } from "@remotion/media";
import { staticFile } from "remotion";
<Video src={staticFile("clip.mp4")} />;
```
**Audio:**
```tsx
import { Audio } from "@remotion/media";
import { staticFile } from "remotion";
<Audio src={staticFile("music.mp3")} />;
```
**Fonts:**
```tsx
import { staticFile } from "remotion";
const fontFamily = new FontFace("MyFont", `url(${staticFile("font.woff2")})`);
await fontFamily.load();
document.fonts.add(fontFamily);
```
## Remote URLs
Remote URLs can be used directly without `staticFile()`:
```tsx
<Img src="https://example.com/image.png" />
<Video src="https://remotion.media/video.mp4" />
```
## Important notes
- Remotion components (`<Img>`, `<Video>`, `<Audio>`) ensure assets are fully loaded before rendering
- Special characters in filenames (`#`, `?`, `&`) are automatically encoded
@@ -0,0 +1,173 @@
import {loadFont} from '@remotion/google-fonts/Inter';
import {AbsoluteFill, spring, useCurrentFrame, useVideoConfig} from 'remotion';
const {fontFamily} = loadFont();
const COLOR_BAR = '#D4AF37';
const COLOR_TEXT = '#ffffff';
const COLOR_MUTED = '#888888';
const COLOR_BG = '#0a0a0a';
const COLOR_AXIS = '#333333';
// Ideal composition size: 1280x720
const Title: React.FC<{children: React.ReactNode}> = ({children}) => (
<div style={{textAlign: 'center', marginBottom: 40}}>
<div style={{color: COLOR_TEXT, fontSize: 48, fontWeight: 600}}>
{children}
</div>
</div>
);
const YAxis: React.FC<{steps: number[]; height: number}> = ({
steps,
height,
}) => (
<div
style={{
display: 'flex',
flexDirection: 'column',
justifyContent: 'space-between',
height,
paddingRight: 16,
}}
>
{steps
.slice()
.reverse()
.map((step) => (
<div
key={step}
style={{
color: COLOR_MUTED,
fontSize: 20,
textAlign: 'right',
}}
>
{step.toLocaleString()}
</div>
))}
</div>
);
const Bar: React.FC<{
height: number;
progress: number;
}> = ({height, progress}) => (
<div
style={{
flex: 1,
display: 'flex',
flexDirection: 'column',
justifyContent: 'flex-end',
}}
>
<div
style={{
width: '100%',
height,
backgroundColor: COLOR_BAR,
borderRadius: '8px 8px 0 0',
opacity: progress,
}}
/>
</div>
);
const XAxis: React.FC<{
children: React.ReactNode;
labels: string[];
height: number;
}> = ({children, labels, height}) => (
<div style={{flex: 1, display: 'flex', flexDirection: 'column'}}>
<div
style={{
display: 'flex',
alignItems: 'flex-end',
gap: 16,
height,
borderLeft: `2px solid ${COLOR_AXIS}`,
borderBottom: `2px solid ${COLOR_AXIS}`,
paddingLeft: 16,
}}
>
{children}
</div>
<div
style={{
display: 'flex',
gap: 16,
paddingLeft: 16,
marginTop: 12,
}}
>
{labels.map((label) => (
<div
key={label}
style={{
flex: 1,
textAlign: 'center',
color: COLOR_MUTED,
fontSize: 20,
}}
>
{label}
</div>
))}
</div>
</div>
);
export const MyAnimation = () => {
const frame = useCurrentFrame();
const {fps, height} = useVideoConfig();
const data = [
{month: 'Jan', price: 2039},
{month: 'Mar', price: 2160},
{month: 'May', price: 2327},
{month: 'Jul', price: 2426},
{month: 'Sep', price: 2634},
{month: 'Nov', price: 2672},
];
const minPrice = 2000;
const maxPrice = 2800;
const priceRange = maxPrice - minPrice;
const chartHeight = height - 280;
const yAxisSteps = [2000, 2400, 2800];
return (
<AbsoluteFill
style={{
backgroundColor: COLOR_BG,
padding: 60,
display: 'flex',
flexDirection: 'column',
fontFamily,
}}
>
<Title>Gold Price 2024</Title>
<div style={{display: 'flex', flex: 1}}>
<YAxis steps={yAxisSteps} height={chartHeight} />
<XAxis height={chartHeight} labels={data.map((d) => d.month)}>
{data.map((item, i) => {
const progress = spring({
frame: frame - i * 5 - 10,
fps,
config: {damping: 18, stiffness: 80},
});
const barHeight =
((item.price - minPrice) / priceRange) * chartHeight * progress;
return (
<Bar key={item.month} height={barHeight} progress={progress} />
);
})}
</XAxis>
</div>
</AbsoluteFill>
);
};
@@ -0,0 +1,100 @@
import {
AbsoluteFill,
interpolate,
useCurrentFrame,
useVideoConfig,
} from 'remotion';
const COLOR_BG = '#ffffff';
const COLOR_TEXT = '#000000';
const FULL_TEXT = 'From prompt to motion graphics. This is Remotion.';
const PAUSE_AFTER = 'From prompt to motion graphics.';
const FONT_SIZE = 72;
const FONT_WEIGHT = 700;
const CHAR_FRAMES = 2;
const CURSOR_BLINK_FRAMES = 16;
const PAUSE_SECONDS = 1;
// Ideal composition size: 1280x720
const getTypedText = ({
frame,
fullText,
pauseAfter,
charFrames,
pauseFrames,
}: {
frame: number;
fullText: string;
pauseAfter: string;
charFrames: number;
pauseFrames: number;
}): string => {
const pauseIndex = fullText.indexOf(pauseAfter);
const preLen =
pauseIndex >= 0 ? pauseIndex + pauseAfter.length : fullText.length;
let typedChars = 0;
if (frame < preLen * charFrames) {
typedChars = Math.floor(frame / charFrames);
} else if (frame < preLen * charFrames + pauseFrames) {
typedChars = preLen;
} else {
const postPhase = frame - preLen * charFrames - pauseFrames;
typedChars = Math.min(
fullText.length,
preLen + Math.floor(postPhase / charFrames),
);
}
return fullText.slice(0, typedChars);
};
const Cursor: React.FC<{
frame: number;
blinkFrames: number;
symbol?: string;
}> = ({frame, blinkFrames, symbol = '\u258C'}) => {
const opacity = interpolate(
frame % blinkFrames,
[0, blinkFrames / 2, blinkFrames],
[1, 0, 1],
{extrapolateLeft: 'clamp', extrapolateRight: 'clamp'},
);
return <span style={{opacity}}>{symbol}</span>;
};
export const MyAnimation = () => {
const frame = useCurrentFrame();
const {fps} = useVideoConfig();
const pauseFrames = Math.round(fps * PAUSE_SECONDS);
const typedText = getTypedText({
frame,
fullText: FULL_TEXT,
pauseAfter: PAUSE_AFTER,
charFrames: CHAR_FRAMES,
pauseFrames,
});
return (
<AbsoluteFill
style={{
backgroundColor: COLOR_BG,
}}
>
<div
style={{
color: COLOR_TEXT,
fontSize: FONT_SIZE,
fontWeight: FONT_WEIGHT,
fontFamily: 'sans-serif',
}}
>
<span>{typedText}</span>
<Cursor frame={frame} blinkFrames={CURSOR_BLINK_FRAMES} />
</div>
</AbsoluteFill>
);
};
@@ -0,0 +1,103 @@
import {loadFont} from '@remotion/google-fonts/Inter';
import React from 'react';
import {AbsoluteFill, spring, useCurrentFrame, useVideoConfig} from 'remotion';
/*
* Highlight a word in a sentence with a spring-animated wipe effect.
*/
// Ideal composition size: 1280x720
const COLOR_BG = '#ffffff';
const COLOR_TEXT = '#000000';
const COLOR_HIGHLIGHT = '#A7C7E7';
const FULL_TEXT = 'This is Remotion.';
const HIGHLIGHT_WORD = 'Remotion';
const FONT_SIZE = 72;
const FONT_WEIGHT = 700;
const HIGHLIGHT_START_FRAME = 30;
const HIGHLIGHT_WIPE_DURATION = 18;
const {fontFamily} = loadFont();
const Highlight: React.FC<{
word: string;
color: string;
delay: number;
durationInFrames: number;
}> = ({word, color, delay, durationInFrames}) => {
const frame = useCurrentFrame();
const {fps} = useVideoConfig();
const highlightProgress = spring({
fps,
frame,
config: {damping: 200},
delay,
durationInFrames,
});
const scaleX = Math.max(0, Math.min(1, highlightProgress));
return (
<span style={{position: 'relative', display: 'inline-block'}}>
<span
style={{
position: 'absolute',
left: 0,
right: 0,
top: '50%',
height: '1.05em',
transform: `translateY(-50%) scaleX(${scaleX})`,
transformOrigin: 'left center',
backgroundColor: color,
borderRadius: '0.18em',
zIndex: 0,
}}
/>
<span style={{position: 'relative', zIndex: 1}}>{word}</span>
</span>
);
};
export const MyAnimation = () => {
const highlightIndex = FULL_TEXT.indexOf(HIGHLIGHT_WORD);
const hasHighlight = highlightIndex >= 0;
const preText = hasHighlight ? FULL_TEXT.slice(0, highlightIndex) : FULL_TEXT;
const postText = hasHighlight
? FULL_TEXT.slice(highlightIndex + HIGHLIGHT_WORD.length)
: '';
return (
<AbsoluteFill
style={{
backgroundColor: COLOR_BG,
alignItems: 'center',
justifyContent: 'center',
fontFamily,
}}
>
<div
style={{
color: COLOR_TEXT,
fontSize: FONT_SIZE,
fontWeight: FONT_WEIGHT,
}}
>
{hasHighlight ? (
<>
<span>{preText}</span>
<Highlight
word={HIGHLIGHT_WORD}
color={COLOR_HIGHLIGHT}
delay={HIGHLIGHT_START_FRAME}
durationInFrames={HIGHLIGHT_WIPE_DURATION}
/>
<span>{postText}</span>
</>
) : (
<span>{FULL_TEXT}</span>
)}
</div>
</AbsoluteFill>
);
};
@@ -0,0 +1,198 @@
---
name: audio-visualization
description: Audio visualization patterns - spectrum bars, waveforms, bass-reactive effects
metadata:
tags: audio, visualization, spectrum, waveform, bass, music, audiogram, frequency
---
# Audio Visualization in Remotion
## Prerequisites
```bash
npx remotion add @remotion/media-utils
```
## Loading Audio Data
Use `useWindowedAudioData()` (https://www.remotion.dev/docs/use-windowed-audio-data) to load audio data:
```tsx
import { useWindowedAudioData } from "@remotion/media-utils";
import { staticFile, useCurrentFrame, useVideoConfig } from "remotion";
const frame = useCurrentFrame();
const { fps } = useVideoConfig();
const { audioData, dataOffsetInSeconds } = useWindowedAudioData({
src: staticFile("podcast.wav"),
frame,
fps,
windowInSeconds: 30,
});
```
## Spectrum Bar Visualization
Use `visualizeAudio()` (https://www.remotion.dev/docs/visualize-audio) to get frequency data for bar charts:
```tsx
import { useWindowedAudioData, visualizeAudio } from "@remotion/media-utils";
import { staticFile, useCurrentFrame, useVideoConfig } from "remotion";
const frame = useCurrentFrame();
const { fps } = useVideoConfig();
const { audioData, dataOffsetInSeconds } = useWindowedAudioData({
src: staticFile("music.mp3"),
frame,
fps,
windowInSeconds: 30,
});
if (!audioData) {
return null;
}
const frequencies = visualizeAudio({
fps,
frame,
audioData,
numberOfSamples: 256,
optimizeFor: "speed",
dataOffsetInSeconds,
});
return (
<div style={{ display: "flex", alignItems: "flex-end", height: 200 }}>
{frequencies.map((v, i) => (
<div
key={i}
style={{
flex: 1,
height: `${v * 100}%`,
backgroundColor: "#0b84f3",
margin: "0 1px",
}}
/>
))}
</div>
);
```
- `numberOfSamples` must be power of 2 (32, 64, 128, 256, 512, 1024)
- Values range 0-1; left of array = bass, right = highs
- Use `optimizeFor: "speed"` for Lambda or high sample counts
**Important:** When passing `audioData` to child components, also pass the `frame` from the parent. Do not call `useCurrentFrame()` in each child - this causes discontinuous visualization when children are inside `<Sequence>` with offsets.
## Waveform Visualization
Use `visualizeAudioWaveform()` (https://www.remotion.dev/docs/media-utils/visualize-audio-waveform) with `createSmoothSvgPath()` (https://www.remotion.dev/docs/media-utils/create-smooth-svg-path) for oscilloscope-style displays:
```tsx
import {
createSmoothSvgPath,
useWindowedAudioData,
visualizeAudioWaveform,
} from "@remotion/media-utils";
import { staticFile, useCurrentFrame, useVideoConfig } from "remotion";
const frame = useCurrentFrame();
const { width, fps } = useVideoConfig();
const HEIGHT = 200;
const { audioData, dataOffsetInSeconds } = useWindowedAudioData({
src: staticFile("voice.wav"),
frame,
fps,
windowInSeconds: 30,
});
if (!audioData) {
return null;
}
const waveform = visualizeAudioWaveform({
fps,
frame,
audioData,
numberOfSamples: 256,
windowInSeconds: 0.5,
dataOffsetInSeconds,
});
const path = createSmoothSvgPath({
points: waveform.map((y, i) => ({
x: (i / (waveform.length - 1)) * width,
y: HEIGHT / 2 + (y * HEIGHT) / 2,
})),
});
return (
<svg width={width} height={HEIGHT}>
<path d={path} fill="none" stroke="#0b84f3" strokeWidth={2} />
</svg>
);
```
## Bass-Reactive Effects
Extract low frequencies for beat-reactive animations:
```tsx
const frequencies = visualizeAudio({
fps,
frame,
audioData,
numberOfSamples: 128,
optimizeFor: "speed",
dataOffsetInSeconds,
});
const lowFrequencies = frequencies.slice(0, 32);
const bassIntensity =
lowFrequencies.reduce((sum, v) => sum + v, 0) / lowFrequencies.length;
const scale = 1 + bassIntensity * 0.5;
const opacity = Math.min(0.6, bassIntensity * 0.8);
```
## Volume-Based Waveform
Use `getWaveformPortion()` (https://www.remotion.dev/docs/get-waveform-portion) when you need simplified volume data instead of frequency spectrum:
```tsx
import { getWaveformPortion } from "@remotion/media-utils";
import { useCurrentFrame, useVideoConfig } from "remotion";
const frame = useCurrentFrame();
const { fps } = useVideoConfig();
const currentTimeInSeconds = frame / fps;
const waveform = getWaveformPortion({
audioData,
startTimeInSeconds: currentTimeInSeconds,
durationInSeconds: 5,
numberOfSamples: 50,
});
// Returns array of { index, amplitude } objects (amplitude: 0-1)
waveform.map((bar) => (
<div key={bar.index} style={{ height: bar.amplitude * 100 }} />
));
```
## Postprocessing
Low frequencies naturally dominate. Apply logarithmic scaling for visual balance:
```tsx
const minDb = -100;
const maxDb = -30;
const scaled = frequencies.map((value) => {
const db = 20 * Math.log10(value);
return (db - minDb) / (maxDb - minDb);
});
```
@@ -0,0 +1,169 @@
---
name: audio
description: Using audio and sound in Remotion - importing, trimming, volume, speed, pitch
metadata:
tags: audio, media, trim, volume, speed, loop, pitch, mute, sound, sfx
---
# Using audio in Remotion
## Prerequisites
First, the @remotion/media package needs to be installed.
If it is not installed, use the following command:
```bash
npx remotion add @remotion/media
```
## Importing Audio
Use `<Audio>` from `@remotion/media` to add audio to your composition.
```tsx
import { Audio } from "@remotion/media";
import { staticFile } from "remotion";
export const MyComposition = () => {
return <Audio src={staticFile("audio.mp3")} />;
};
```
Remote URLs are also supported:
```tsx
<Audio src="https://remotion.media/audio.mp3" />
```
By default, audio plays from the start, at full volume and full length.
Multiple audio tracks can be layered by adding multiple `<Audio>` components.
## Trimming
Use `trimBefore` and `trimAfter` to remove portions of the audio. Values are in frames.
```tsx
const { fps } = useVideoConfig();
return (
<Audio
src={staticFile("audio.mp3")}
trimBefore={2 * fps} // Skip the first 2 seconds
trimAfter={10 * fps} // End at the 10 second mark
/>
);
```
The audio still starts playing at the beginning of the composition - only the specified portion is played.
## Delaying
Wrap the audio in a `<Sequence>` to delay when it starts:
```tsx
import { Sequence, staticFile } from "remotion";
import { Audio } from "@remotion/media";
const { fps } = useVideoConfig();
return (
<Sequence from={1 * fps}>
<Audio src={staticFile("audio.mp3")} />
</Sequence>
);
```
The audio will start playing after 1 second.
## Volume
Set a static volume (0 to 1):
```tsx
<Audio src={staticFile("audio.mp3")} volume={0.5} />
```
Or use a callback for dynamic volume based on the current frame:
```tsx
import { interpolate } from "remotion";
const { fps } = useVideoConfig();
return (
<Audio
src={staticFile("audio.mp3")}
volume={(f) =>
interpolate(f, [0, 1 * fps], [0, 1], { extrapolateRight: "clamp" })
}
/>
);
```
The value of `f` starts at 0 when the audio begins to play, not the composition frame.
## Muting
Use `muted` to silence the audio. It can be set dynamically:
```tsx
const frame = useCurrentFrame();
const { fps } = useVideoConfig();
return (
<Audio
src={staticFile("audio.mp3")}
muted={frame >= 2 * fps && frame <= 4 * fps} // Mute between 2s and 4s
/>
);
```
## Speed
Use `playbackRate` to change the playback speed:
```tsx
<Audio src={staticFile("audio.mp3")} playbackRate={2} /> {/* 2x speed */}
<Audio src={staticFile("audio.mp3")} playbackRate={0.5} /> {/* Half speed */}
```
Reverse playback is not supported.
## Looping
Use `loop` to loop the audio indefinitely:
```tsx
<Audio src={staticFile("audio.mp3")} loop />
```
Use `loopVolumeCurveBehavior` to control how the frame count behaves when looping:
- `"repeat"`: Frame count resets to 0 each loop (default)
- `"extend"`: Frame count continues incrementing
```tsx
<Audio
src={staticFile("audio.mp3")}
loop
loopVolumeCurveBehavior="extend"
volume={(f) => interpolate(f, [0, 300], [1, 0])} // Fade out over multiple loops
/>
```
## Pitch
Use `toneFrequency` to adjust the pitch without affecting speed. Values range from 0.01 to 2:
```tsx
<Audio
src={staticFile("audio.mp3")}
toneFrequency={1.5} // Higher pitch
/>
<Audio
src={staticFile("audio.mp3")}
toneFrequency={0.8} // Lower pitch
/>
```
Pitch shifting only works during server-side rendering, not in the Remotion Studio preview or in the `<Player />`.
@@ -0,0 +1,134 @@
---
name: calculate-metadata
description: Dynamically set composition duration, dimensions, and props
metadata:
tags: calculateMetadata, duration, dimensions, props, dynamic
---
# Using calculateMetadata
Use `calculateMetadata` on a `<Composition>` to dynamically set duration, dimensions, and transform props before rendering.
```tsx
<Composition
id="MyComp"
component={MyComponent}
durationInFrames={300}
fps={30}
width={1920}
height={1080}
defaultProps={{ videoSrc: "https://remotion.media/video.mp4" }}
calculateMetadata={calculateMetadata}
/>
```
## Setting duration based on a video
Use the [`getVideoDuration`](./get-video-duration.md) and [`getVideoDimensions`](./get-video-dimensions.md) skills to get the video duration and dimensions:
```tsx
import { CalculateMetadataFunction } from "remotion";
import { getVideoDuration } from "./get-video-duration";
const calculateMetadata: CalculateMetadataFunction<Props> = async ({
props,
}) => {
const durationInSeconds = await getVideoDuration(props.videoSrc);
return {
durationInFrames: Math.ceil(durationInSeconds * 30),
};
};
```
## Matching dimensions of a video
Use the [`getVideoDimensions`](./get-video-dimensions.md) skill to get the video dimensions:
```tsx
import { CalculateMetadataFunction } from "remotion";
import { getVideoDuration } from "./get-video-duration";
import { getVideoDimensions } from "./get-video-dimensions";
const calculateMetadata: CalculateMetadataFunction<Props> = async ({
props,
}) => {
const dimensions = await getVideoDimensions(props.videoSrc);
return {
width: dimensions.width,
height: dimensions.height,
};
};
```
## Setting duration based on multiple videos
```tsx
const calculateMetadata: CalculateMetadataFunction<Props> = async ({
props,
}) => {
const metadataPromises = props.videos.map((video) =>
getVideoDuration(video.src),
);
const allMetadata = await Promise.all(metadataPromises);
const totalDuration = allMetadata.reduce(
(sum, durationInSeconds) => sum + durationInSeconds,
0,
);
return {
durationInFrames: Math.ceil(totalDuration * 30),
};
};
```
## Setting a default outName
Set the default output filename based on props:
```tsx
const calculateMetadata: CalculateMetadataFunction<Props> = async ({
props,
}) => {
return {
defaultOutName: `video-${props.id}.mp4`,
};
};
```
## Transforming props
Fetch data or transform props before rendering:
```tsx
const calculateMetadata: CalculateMetadataFunction<Props> = async ({
props,
abortSignal,
}) => {
const response = await fetch(props.dataUrl, { signal: abortSignal });
const data = await response.json();
return {
props: {
...props,
fetchedData: data,
},
};
};
```
The `abortSignal` cancels stale requests when props change in the Studio.
## Return value
All fields are optional. Returned values override the `<Composition>` props:
- `durationInFrames`: Number of frames
- `width`: Composition width in pixels
- `height`: Composition height in pixels
- `fps`: Frames per second
- `props`: Transformed props passed to the component
- `defaultOutName`: Default output filename
- `defaultCodec`: Default codec for rendering
@@ -0,0 +1,75 @@
---
name: can-decode
description: Check if a video can be decoded by the browser using Mediabunny
metadata:
tags: decode, validation, video, audio, compatibility, browser
---
# Checking if a video can be decoded
Use Mediabunny to check if a video can be decoded by the browser before attempting to play it.
## The `canDecode()` function
This function can be copy-pasted into any project.
```tsx
import { Input, ALL_FORMATS, UrlSource } from "mediabunny";
export const canDecode = async (src: string) => {
const input = new Input({
formats: ALL_FORMATS,
source: new UrlSource(src, {
getRetryDelay: () => null,
}),
});
try {
await input.getFormat();
} catch {
return false;
}
const videoTrack = await input.getPrimaryVideoTrack();
if (videoTrack && !(await videoTrack.canDecode())) {
return false;
}
const audioTrack = await input.getPrimaryAudioTrack();
if (audioTrack && !(await audioTrack.canDecode())) {
return false;
}
return true;
};
```
## Usage
```tsx
const src = "https://remotion.media/video.mp4";
const isDecodable = await canDecode(src);
if (isDecodable) {
console.log("Video can be decoded");
} else {
console.log("Video cannot be decoded by this browser");
}
```
## Using with Blob
For file uploads or drag-and-drop, use `BlobSource`:
```tsx
import { Input, ALL_FORMATS, BlobSource } from "mediabunny";
export const canDecodeBlob = async (blob: Blob) => {
const input = new Input({
formats: ALL_FORMATS,
source: new BlobSource(blob),
});
// Same validation logic as above
};
```
@@ -0,0 +1,120 @@
---
name: charts
description: Chart and data visualization patterns for Remotion. Use when creating bar charts, pie charts, line charts, stock graphs, or any data-driven animations.
metadata:
tags: charts, data, visualization, bar-chart, pie-chart, line-chart, stock-chart, svg-paths, graphs
---
# Charts in Remotion
Create charts using React code - HTML, SVG, and D3.js are all supported.
Disable all animations from third party libraries - they cause flickering.
Drive all animations from `useCurrentFrame()`.
## Bar Chart
```tsx
const STAGGER_DELAY = 5;
const frame = useCurrentFrame();
const { fps } = useVideoConfig();
const bars = data.map((item, i) => {
const height = spring({
frame,
fps,
delay: i * STAGGER_DELAY,
config: { damping: 200 },
});
return <div style={{ height: height * item.value }} />;
});
```
## Pie Chart
Animate segments using stroke-dashoffset, starting from 12 o'clock:
```tsx
const progress = interpolate(frame, [0, 100], [0, 1]);
const circumference = 2 * Math.PI * radius;
const segmentLength = (value / total) * circumference;
const offset = interpolate(progress, [0, 1], [segmentLength, 0]);
<circle
r={radius}
cx={center}
cy={center}
fill="none"
stroke={color}
strokeWidth={strokeWidth}
strokeDasharray={`${segmentLength} ${circumference}`}
strokeDashoffset={offset}
transform={`rotate(-90 ${center} ${center})`}
/>;
```
## Line Chart / Path Animation
Use `@remotion/paths` for animating SVG paths (line charts, stock graphs, signatures).
Install: `npx remotion add @remotion/paths`
Docs: https://remotion.dev/docs/paths.md
### Convert data points to SVG path
```tsx
type Point = { x: number; y: number };
const generateLinePath = (points: Point[]): string => {
if (points.length < 2) return "";
return points.map((p, i) => `${i === 0 ? "M" : "L"} ${p.x} ${p.y}`).join(" ");
};
```
### Draw path with animation
```tsx
import { evolvePath } from "@remotion/paths";
const path = "M 100 200 L 200 150 L 300 180 L 400 100";
const progress = interpolate(frame, [0, 2 * fps], [0, 1], {
extrapolateLeft: "clamp",
extrapolateRight: "clamp",
easing: Easing.out(Easing.quad),
});
const { strokeDasharray, strokeDashoffset } = evolvePath(progress, path);
<path
d={path}
fill="none"
stroke="#FF3232"
strokeWidth={4}
strokeDasharray={strokeDasharray}
strokeDashoffset={strokeDashoffset}
/>;
```
### Follow path with marker/arrow
```tsx
import {
getLength,
getPointAtLength,
getTangentAtLength,
} from "@remotion/paths";
const pathLength = getLength(path);
const point = getPointAtLength(path, progress * pathLength);
const tangent = getTangentAtLength(path, progress * pathLength);
const angle = Math.atan2(tangent.y, tangent.x);
<g
style={{
transform: `translate(${point.x}px, ${point.y}px) rotate(${angle}rad)`,
transformOrigin: "0 0",
}}
>
<polygon points="0,0 -20,-10 -20,10" fill="#FF3232" />
</g>;
```
@@ -0,0 +1,154 @@
---
name: compositions
description: Defining compositions, stills, folders, default props and dynamic metadata
metadata:
tags: composition, still, folder, props, metadata
---
A `<Composition>` defines the component, width, height, fps and duration of a renderable video.
It normally is placed in the `src/Root.tsx` file.
```tsx
import { Composition } from "remotion";
import { MyComposition } from "./MyComposition";
export const RemotionRoot = () => {
return (
<Composition
id="MyComposition"
component={MyComposition}
durationInFrames={100}
fps={30}
width={1080}
height={1080}
/>
);
};
```
## Default Props
Pass `defaultProps` to provide initial values for your component.
Values must be JSON-serializable (`Date`, `Map`, `Set`, and `staticFile()` are supported).
```tsx
import { Composition } from "remotion";
import { MyComposition, MyCompositionProps } from "./MyComposition";
export const RemotionRoot = () => {
return (
<Composition
id="MyComposition"
component={MyComposition}
durationInFrames={100}
fps={30}
width={1080}
height={1080}
defaultProps={
{
title: "Hello World",
color: "#ff0000",
} satisfies MyCompositionProps
}
/>
);
};
```
Use `type` declarations for props rather than `interface` to ensure `defaultProps` type safety.
## Folders
Use `<Folder>` to organize compositions in the sidebar.
Folder names can only contain letters, numbers, and hyphens.
```tsx
import { Composition, Folder } from "remotion";
export const RemotionRoot = () => {
return (
<>
<Folder name="Marketing">
<Composition id="Promo" /* ... */ />
<Composition id="Ad" /* ... */ />
</Folder>
<Folder name="Social">
<Folder name="Instagram">
<Composition id="Story" /* ... */ />
<Composition id="Reel" /* ... */ />
</Folder>
</Folder>
</>
);
};
```
## Stills
Use `<Still>` for single-frame images. It does not require `durationInFrames` or `fps`.
```tsx
import { Still } from "remotion";
import { Thumbnail } from "./Thumbnail";
export const RemotionRoot = () => {
return (
<Still id="Thumbnail" component={Thumbnail} width={1280} height={720} />
);
};
```
## Calculate Metadata
Use `calculateMetadata` to make dimensions, duration, or props dynamic based on data.
```tsx
import { Composition, CalculateMetadataFunction } from "remotion";
import { MyComposition, MyCompositionProps } from "./MyComposition";
const calculateMetadata: CalculateMetadataFunction<
MyCompositionProps
> = async ({ props, abortSignal }) => {
const data = await fetch(`https://api.example.com/video/${props.videoId}`, {
signal: abortSignal,
}).then((res) => res.json());
return {
durationInFrames: Math.ceil(data.duration * 30),
props: {
...props,
videoUrl: data.url,
},
};
};
export const RemotionRoot = () => {
return (
<Composition
id="MyComposition"
component={MyComposition}
durationInFrames={100} // Placeholder, will be overridden
fps={30}
width={1080}
height={1080}
defaultProps={{ videoId: "abc123" }}
calculateMetadata={calculateMetadata}
/>
);
};
```
The function can return `props`, `durationInFrames`, `width`, `height`, `fps`, and codec-related defaults. It runs once before rendering begins.
## Nesting compositions within another
To add a composition within another composition, you can use the `<Sequence>` component with a `width` and `height` prop to specify the size of the composition.
```tsx
<AbsoluteFill>
<Sequence width={COMPOSITION_WIDTH} height={COMPOSITION_HEIGHT}>
<CompositionComponent />
</Sequence>
</AbsoluteFill>
```
@@ -0,0 +1,184 @@
---
name: display-captions
description: Displaying captions in Remotion with TikTok-style pages and word highlighting
metadata:
tags: captions, subtitles, display, tiktok, highlight
---
# Displaying captions in Remotion
This guide explains how to display captions in Remotion, assuming you already have captions in the [`Caption`](https://www.remotion.dev/docs/captions/caption) format.
## Prerequisites
Read [Transcribing audio](transcribe-captions.md) for how to generate captions.
First, the [`@remotion/captions`](https://www.remotion.dev/docs/captions) package needs to be installed.
If it is not installed, use the following command:
```bash
npx remotion add @remotion/captions
```
## Fetching captions
First, fetch your captions JSON file. Use [`useDelayRender()`](https://www.remotion.dev/docs/use-delay-render) to hold the render until the captions are loaded:
```tsx
import { useState, useEffect, useCallback } from "react";
import { AbsoluteFill, staticFile, useDelayRender } from "remotion";
import type { Caption } from "@remotion/captions";
export const MyComponent: React.FC = () => {
const [captions, setCaptions] = useState<Caption[] | null>(null);
const { delayRender, continueRender, cancelRender } = useDelayRender();
const [handle] = useState(() => delayRender());
const fetchCaptions = useCallback(async () => {
try {
// Assuming captions.json is in the public/ folder.
const response = await fetch(staticFile("captions123.json"));
const data = await response.json();
setCaptions(data);
continueRender(handle);
} catch (e) {
cancelRender(e);
}
}, [continueRender, cancelRender, handle]);
useEffect(() => {
fetchCaptions();
}, [fetchCaptions]);
if (!captions) {
return null;
}
return <AbsoluteFill>{/* Render captions here */}</AbsoluteFill>;
};
```
## Creating pages
Use `createTikTokStyleCaptions()` to group captions into pages. The `combineTokensWithinMilliseconds` option controls how many words appear at once:
```tsx
import { useMemo } from "react";
import { createTikTokStyleCaptions } from "@remotion/captions";
import type { Caption } from "@remotion/captions";
// How often captions should switch (in milliseconds)
// Higher values = more words per page
// Lower values = fewer words (more word-by-word)
const SWITCH_CAPTIONS_EVERY_MS = 1200;
const { pages } = useMemo(() => {
return createTikTokStyleCaptions({
captions,
combineTokensWithinMilliseconds: SWITCH_CAPTIONS_EVERY_MS,
});
}, [captions]);
```
## Rendering with Sequences
Map over the pages and render each one in a `<Sequence>`. Calculate the start frame and duration from the page timing:
```tsx
import { Sequence, useVideoConfig, AbsoluteFill } from "remotion";
import type { TikTokPage } from "@remotion/captions";
const CaptionedContent: React.FC = () => {
const { fps } = useVideoConfig();
return (
<AbsoluteFill>
{pages.map((page, index) => {
const nextPage = pages[index + 1] ?? null;
const startFrame = (page.startMs / 1000) * fps;
const endFrame = Math.min(
nextPage ? (nextPage.startMs / 1000) * fps : Infinity,
startFrame + (SWITCH_CAPTIONS_EVERY_MS / 1000) * fps,
);
const durationInFrames = endFrame - startFrame;
if (durationInFrames <= 0) {
return null;
}
return (
<Sequence
key={index}
from={startFrame}
durationInFrames={durationInFrames}
>
<CaptionPage page={page} />
</Sequence>
);
})}
</AbsoluteFill>
);
};
```
## White-space preservation
The captions are whitespace sensitive. You should include spaces in the `text` field before each word. Use `whiteSpace: "pre"` to preserve the whitespace in the captions.
## Separate component for captions
Put captioning logic in a separate component.
Make a new file for it.
## Word highlighting
A caption page contains `tokens` which you can use to highlight the currently spoken word:
```tsx
import { AbsoluteFill, useCurrentFrame, useVideoConfig } from "remotion";
import type { TikTokPage } from "@remotion/captions";
const HIGHLIGHT_COLOR = "#39E508";
const CaptionPage: React.FC<{ page: TikTokPage }> = ({ page }) => {
const frame = useCurrentFrame();
const { fps } = useVideoConfig();
// Current time relative to the start of the sequence
const currentTimeMs = (frame / fps) * 1000;
// Convert to absolute time by adding the page start
const absoluteTimeMs = page.startMs + currentTimeMs;
return (
<AbsoluteFill style={{ justifyContent: "center", alignItems: "center" }}>
<div style={{ fontSize: 80, fontWeight: "bold", whiteSpace: "pre" }}>
{page.tokens.map((token) => {
const isActive =
token.fromMs <= absoluteTimeMs && token.toMs > absoluteTimeMs;
return (
<span
key={token.fromMs}
style={{ color: isActive ? HIGHLIGHT_COLOR : "white" }}
>
{token.text}
</span>
);
})}
</div>
</AbsoluteFill>
);
};
```
## Display captions alongside video content
By default, put the captions alongside the video content, so the captions are in sync.
For each video, make a new captions JSON file.
```tsx
<AbsoluteFill>
<Video src={staticFile("video.mp4")} />
<CaptionPage page={page} />
</AbsoluteFill>
```
@@ -0,0 +1,229 @@
---
name: extract-frames
description: Extract frames from videos at specific timestamps using Mediabunny
metadata:
tags: frames, extract, video, thumbnail, filmstrip, canvas
---
# Extracting frames from videos
Use Mediabunny to extract frames from videos at specific timestamps. This is useful for generating thumbnails, filmstrips, or processing individual frames.
## The `extractFrames()` function
This function can be copy-pasted into any project.
```tsx
import {
ALL_FORMATS,
Input,
UrlSource,
VideoSample,
VideoSampleSink,
} from "mediabunny";
type Options = {
track: { width: number; height: number };
container: string;
durationInSeconds: number | null;
};
export type ExtractFramesTimestampsInSecondsFn = (
options: Options,
) => Promise<number[]> | number[];
export type ExtractFramesProps = {
src: string;
timestampsInSeconds: number[] | ExtractFramesTimestampsInSecondsFn;
onVideoSample: (sample: VideoSample) => void;
signal?: AbortSignal;
};
export async function extractFrames({
src,
timestampsInSeconds,
onVideoSample,
signal,
}: ExtractFramesProps): Promise<void> {
using input = new Input({
formats: ALL_FORMATS,
source: new UrlSource(src),
});
const [durationInSeconds, format, videoTrack] = await Promise.all([
input.computeDuration(),
input.getFormat(),
input.getPrimaryVideoTrack(),
]);
if (!videoTrack) {
throw new Error("No video track found in the input");
}
if (signal?.aborted) {
throw new Error("Aborted");
}
const timestamps =
typeof timestampsInSeconds === "function"
? await timestampsInSeconds({
track: {
width: videoTrack.displayWidth,
height: videoTrack.displayHeight,
},
container: format.name,
durationInSeconds,
})
: timestampsInSeconds;
if (timestamps.length === 0) {
return;
}
if (signal?.aborted) {
throw new Error("Aborted");
}
const sink = new VideoSampleSink(videoTrack);
for await (using videoSample of sink.samplesAtTimestamps(timestamps)) {
if (signal?.aborted) {
break;
}
if (!videoSample) {
continue;
}
onVideoSample(videoSample);
}
}
```
## Basic usage
Extract frames at specific timestamps:
```tsx
await extractFrames({
src: "https://remotion.media/video.mp4",
timestampsInSeconds: [0, 1, 2, 3, 4],
onVideoSample: (sample) => {
const canvas = document.createElement("canvas");
canvas.width = sample.displayWidth;
canvas.height = sample.displayHeight;
const ctx = canvas.getContext("2d");
sample.draw(ctx!, 0, 0);
},
});
```
## Creating a filmstrip
Use a callback function to dynamically calculate timestamps based on video metadata:
```tsx
const canvasWidth = 500;
const canvasHeight = 80;
const fromSeconds = 0;
const toSeconds = 10;
await extractFrames({
src: "https://remotion.media/video.mp4",
timestampsInSeconds: async ({ track, durationInSeconds }) => {
const aspectRatio = track.width / track.height;
const amountOfFramesFit = Math.ceil(
canvasWidth / (canvasHeight * aspectRatio),
);
const segmentDuration = toSeconds - fromSeconds;
const timestamps: number[] = [];
for (let i = 0; i < amountOfFramesFit; i++) {
timestamps.push(
fromSeconds + (segmentDuration / amountOfFramesFit) * (i + 0.5),
);
}
return timestamps;
},
onVideoSample: (sample) => {
console.log(`Frame at ${sample.timestamp}s`);
const canvas = document.createElement("canvas");
canvas.width = sample.displayWidth;
canvas.height = sample.displayHeight;
const ctx = canvas.getContext("2d");
sample.draw(ctx!, 0, 0);
},
});
```
## Cancellation with AbortSignal
Cancel frame extraction after a timeout:
```tsx
const controller = new AbortController();
setTimeout(() => controller.abort(), 5000);
try {
await extractFrames({
src: "https://remotion.media/video.mp4",
timestampsInSeconds: [0, 1, 2, 3, 4],
onVideoSample: (sample) => {
using frame = sample;
const canvas = document.createElement("canvas");
canvas.width = frame.displayWidth;
canvas.height = frame.displayHeight;
const ctx = canvas.getContext("2d");
frame.draw(ctx!, 0, 0);
},
signal: controller.signal,
});
console.log("Frame extraction complete!");
} catch (error) {
console.error("Frame extraction was aborted or failed:", error);
}
```
## Timeout with Promise.race
```tsx
const controller = new AbortController();
const timeoutPromise = new Promise<never>((_, reject) => {
const timeoutId = setTimeout(() => {
controller.abort();
reject(new Error("Frame extraction timed out after 10 seconds"));
}, 10000);
controller.signal.addEventListener("abort", () => clearTimeout(timeoutId), {
once: true,
});
});
try {
await Promise.race([
extractFrames({
src: "https://remotion.media/video.mp4",
timestampsInSeconds: [0, 1, 2, 3, 4],
onVideoSample: (sample) => {
using frame = sample;
const canvas = document.createElement("canvas");
canvas.width = frame.displayWidth;
canvas.height = frame.displayHeight;
const ctx = canvas.getContext("2d");
frame.draw(ctx!, 0, 0);
},
signal: controller.signal,
}),
timeoutPromise,
]);
console.log("Frame extraction complete!");
} catch (error) {
console.error("Frame extraction was aborted or failed:", error);
}
```
@@ -0,0 +1,38 @@
---
name: ffmpeg
description: Using FFmpeg and FFprobe in Remotion
metadata:
tags: ffmpeg, ffprobe, video, trimming
---
## FFmpeg in Remotion
`ffmpeg` and `ffprobe` do not need to be installed. They are available via the `bunx remotion ffmpeg` and `bunx remotion ffprobe`:
```bash
bunx remotion ffmpeg -i input.mp4 output.mp3
bunx remotion ffprobe input.mp4
```
### Trimming videos
You have 2 options for trimming videos:
1. Use the FFMpeg command line. You MUST re-encode the video to avoid frozen frames at the start of the video.
```bash
# Re-encodes from the exact frame
bunx remotion ffmpeg -ss 00:00:05 -i public/input.mp4 -to 00:00:10 -c:v libx264 -c:a aac public/output.mp4
```
2. Use the `trimBefore` and `trimAfter` props of the `<Video>` component. The benefit is that this is non-destructive and you can change the trim at any time.
```tsx
import { Video } from "@remotion/media";
<Video
src={staticFile("video.mp4")}
trimBefore={5 * fps}
trimAfter={10 * fps}
/>;
```
@@ -0,0 +1,152 @@
---
name: fonts
description: Loading Google Fonts and local fonts in Remotion
metadata:
tags: fonts, google-fonts, typography, text
---
# Using fonts in Remotion
## Google Fonts with @remotion/google-fonts
The recommended way to use Google Fonts. It's type-safe and automatically blocks rendering until the font is ready.
### Prerequisites
First, the @remotion/google-fonts package needs to be installed.
If it is not installed, use the following command:
```bash
npx remotion add @remotion/google-fonts # If project uses npm
bunx remotion add @remotion/google-fonts # If project uses bun
yarn remotion add @remotion/google-fonts # If project uses yarn
pnpm exec remotion add @remotion/google-fonts # If project uses pnpm
```
```tsx
import { loadFont } from "@remotion/google-fonts/Lobster";
const { fontFamily } = loadFont();
export const MyComposition = () => {
return <div style={{ fontFamily }}>Hello World</div>;
};
```
Preferrably, specify only needed weights and subsets to reduce file size:
```tsx
import { loadFont } from "@remotion/google-fonts/Roboto";
const { fontFamily } = loadFont("normal", {
weights: ["400", "700"],
subsets: ["latin"],
});
```
### Waiting for font to load
Use `waitUntilDone()` if you need to know when the font is ready:
```tsx
import { loadFont } from "@remotion/google-fonts/Lobster";
const { fontFamily, waitUntilDone } = loadFont();
await waitUntilDone();
```
## Local fonts with @remotion/fonts
For local font files, use the `@remotion/fonts` package.
### Prerequisites
First, install @remotion/fonts:
```bash
npx remotion add @remotion/fonts # If project uses npm
bunx remotion add @remotion/fonts # If project uses bun
yarn remotion add @remotion/fonts # If project uses yarn
pnpm exec remotion add @remotion/fonts # If project uses pnpm
```
### Loading a local font
Place your font file in the `public/` folder and use `loadFont()`:
```tsx
import { loadFont } from "@remotion/fonts";
import { staticFile } from "remotion";
await loadFont({
family: "MyFont",
url: staticFile("MyFont-Regular.woff2"),
});
export const MyComposition = () => {
return <div style={{ fontFamily: "MyFont" }}>Hello World</div>;
};
```
### Loading multiple weights
Load each weight separately with the same family name:
```tsx
import { loadFont } from "@remotion/fonts";
import { staticFile } from "remotion";
await Promise.all([
loadFont({
family: "Inter",
url: staticFile("Inter-Regular.woff2"),
weight: "400",
}),
loadFont({
family: "Inter",
url: staticFile("Inter-Bold.woff2"),
weight: "700",
}),
]);
```
### Available options
```tsx
loadFont({
family: "MyFont", // Required: name to use in CSS
url: staticFile("font.woff2"), // Required: font file URL
format: "woff2", // Optional: auto-detected from extension
weight: "400", // Optional: font weight
style: "normal", // Optional: normal or italic
display: "block", // Optional: font-display behavior
});
```
## Using in components
Call `loadFont()` at the top level of your component or in a separate file that's imported early:
```tsx
import { loadFont } from "@remotion/google-fonts/Montserrat";
const { fontFamily } = loadFont("normal", {
weights: ["400", "700"],
subsets: ["latin"],
});
export const Title: React.FC<{ text: string }> = ({ text }) => {
return (
<h1
style={{
fontFamily,
fontSize: 80,
fontWeight: "bold",
}}
>
{text}
</h1>
);
};
```
@@ -0,0 +1,58 @@
---
name: get-audio-duration
description: Getting the duration of an audio file in seconds with Mediabunny
metadata:
tags: duration, audio, length, time, seconds, mp3, wav
---
# Getting audio duration with Mediabunny
Mediabunny can extract the duration of an audio file. It works in browser, Node.js, and Bun environments.
## Getting audio duration
```tsx title="get-audio-duration.ts"
import { Input, ALL_FORMATS, UrlSource } from "mediabunny";
export const getAudioDuration = async (src: string) => {
const input = new Input({
formats: ALL_FORMATS,
source: new UrlSource(src, {
getRetryDelay: () => null,
}),
});
const durationInSeconds = await input.computeDuration();
return durationInSeconds;
};
```
## Usage
```tsx
const duration = await getAudioDuration("https://remotion.media/audio.mp3");
console.log(duration); // e.g. 180.5 (seconds)
```
## Using with staticFile in Remotion
Make sure to wrap the file path in `staticFile()`:
```tsx
import { staticFile } from "remotion";
const duration = await getAudioDuration(staticFile("audio.mp3"));
```
## In Node.js and Bun
Use `FileSource` instead of `UrlSource`:
```tsx
import { Input, ALL_FORMATS, FileSource } from "mediabunny";
const input = new Input({
formats: ALL_FORMATS,
source: new FileSource(file), // File object from input or drag-drop
});
```
@@ -0,0 +1,68 @@
---
name: get-video-dimensions
description: Getting the width and height of a video file with Mediabunny
metadata:
tags: dimensions, width, height, resolution, size, video
---
# Getting video dimensions with Mediabunny
Mediabunny can extract the width and height of a video file. It works in browser, Node.js, and Bun environments.
## Getting video dimensions
```tsx
import { Input, ALL_FORMATS, UrlSource } from "mediabunny";
export const getVideoDimensions = async (src: string) => {
const input = new Input({
formats: ALL_FORMATS,
source: new UrlSource(src, {
getRetryDelay: () => null,
}),
});
const videoTrack = await input.getPrimaryVideoTrack();
if (!videoTrack) {
throw new Error("No video track found");
}
return {
width: videoTrack.displayWidth,
height: videoTrack.displayHeight,
};
};
```
## Usage
```tsx
const dimensions = await getVideoDimensions("https://remotion.media/video.mp4");
console.log(dimensions.width); // e.g. 1920
console.log(dimensions.height); // e.g. 1080
```
## Using with local files
For local files, use `FileSource` instead of `UrlSource`:
```tsx
import { Input, ALL_FORMATS, FileSource } from "mediabunny";
const input = new Input({
formats: ALL_FORMATS,
source: new FileSource(file), // File object from input or drag-drop
});
const videoTrack = await input.getPrimaryVideoTrack();
const width = videoTrack.displayWidth;
const height = videoTrack.displayHeight;
```
## Using with staticFile in Remotion
```tsx
import { staticFile } from "remotion";
const dimensions = await getVideoDimensions(staticFile("video.mp4"));
```
@@ -0,0 +1,60 @@
---
name: get-video-duration
description: Getting the duration of a video file in seconds with Mediabunny
metadata:
tags: duration, video, length, time, seconds
---
# Getting video duration with Mediabunny
Mediabunny can extract the duration of a video file. It works in browser, Node.js, and Bun environments.
## Getting video duration
```tsx
import { Input, ALL_FORMATS, UrlSource } from "mediabunny";
export const getVideoDuration = async (src: string) => {
const input = new Input({
formats: ALL_FORMATS,
source: new UrlSource(src, {
getRetryDelay: () => null,
}),
});
const durationInSeconds = await input.computeDuration();
return durationInSeconds;
};
```
## Usage
```tsx
const duration = await getVideoDuration("https://remotion.media/video.mp4");
console.log(duration); // e.g. 10.5 (seconds)
```
## Video files from the public/ directory
Make sure to wrap the file path in `staticFile()`:
```tsx
import { staticFile } from "remotion";
const duration = await getVideoDuration(staticFile("video.mp4"));
```
## In Node.js and Bun
Use `FileSource` instead of `UrlSource`:
```tsx
import { Input, ALL_FORMATS, FileSource } from "mediabunny";
const input = new Input({
formats: ALL_FORMATS,
source: new FileSource(file), // File object from input or drag-drop
});
const durationInSeconds = await input.computeDuration();
```
@@ -0,0 +1,141 @@
---
name: gif
description: Displaying GIFs, APNG, AVIF and WebP in Remotion
metadata:
tags: gif, animation, images, animated, apng, avif, webp
---
# Using Animated images in Remotion
## Basic usage
Use `<AnimatedImage>` to display a GIF, APNG, AVIF or WebP image synchronized with Remotion's timeline:
```tsx
import { AnimatedImage, staticFile } from "remotion";
export const MyComposition = () => {
return (
<AnimatedImage src={staticFile("animation.gif")} width={500} height={500} />
);
};
```
Remote URLs are also supported (must have CORS enabled):
```tsx
<AnimatedImage
src="https://example.com/animation.gif"
width={500}
height={500}
/>
```
## Sizing and fit
Control how the image fills its container with the `fit` prop:
```tsx
// Stretch to fill (default)
<AnimatedImage src={staticFile("animation.gif")} width={500} height={300} fit="fill" />
// Maintain aspect ratio, fit inside container
<AnimatedImage src={staticFile("animation.gif")} width={500} height={300} fit="contain" />
// Fill container, crop if needed
<AnimatedImage src={staticFile("animation.gif")} width={500} height={300} fit="cover" />
```
## Playback speed
Use `playbackRate` to control the animation speed:
```tsx
<AnimatedImage src={staticFile("animation.gif")} width={500} height={500} playbackRate={2} /> {/* 2x speed */}
<AnimatedImage src={staticFile("animation.gif")} width={500} height={500} playbackRate={0.5} /> {/* Half speed */}
```
## Looping behavior
Control what happens when the animation finishes:
```tsx
// Loop indefinitely (default)
<AnimatedImage src={staticFile("animation.gif")} width={500} height={500} loopBehavior="loop" />
// Play once, show final frame
<AnimatedImage src={staticFile("animation.gif")} width={500} height={500} loopBehavior="pause-after-finish" />
// Play once, then clear canvas
<AnimatedImage src={staticFile("animation.gif")} width={500} height={500} loopBehavior="clear-after-finish" />
```
## Styling
Use the `style` prop for additional CSS (use `width` and `height` props for sizing):
```tsx
<AnimatedImage
src={staticFile("animation.gif")}
width={500}
height={500}
style={{
borderRadius: 20,
position: "absolute",
top: 100,
left: 50,
}}
/>
```
## Getting GIF duration
Use `getGifDurationInSeconds()` from `@remotion/gif` to get the duration of a GIF.
```bash
npx remotion add @remotion/gif
```
```tsx
import { getGifDurationInSeconds } from "@remotion/gif";
import { staticFile } from "remotion";
const duration = await getGifDurationInSeconds(staticFile("animation.gif"));
console.log(duration); // e.g. 2.5
```
This is useful for setting the composition duration to match the GIF:
```tsx
import { getGifDurationInSeconds } from "@remotion/gif";
import { staticFile, CalculateMetadataFunction } from "remotion";
const calculateMetadata: CalculateMetadataFunction = async () => {
const duration = await getGifDurationInSeconds(staticFile("animation.gif"));
return {
durationInFrames: Math.ceil(duration * 30),
};
};
```
## Alternative
If `<AnimatedImage>` does not work (only supported in Chrome and Firefox), you can use `<Gif>` from `@remotion/gif` instead.
```bash
npx remotion add @remotion/gif # If project uses npm
bunx remotion add @remotion/gif # If project uses bun
yarn remotion add @remotion/gif # If project uses yarn
pnpm exec remotion add @remotion/gif # If project uses pnpm
```
```tsx
import { Gif } from "@remotion/gif";
import { staticFile } from "remotion";
export const MyComposition = () => {
return <Gif src={staticFile("animation.gif")} width={500} height={500} />;
};
```
The `<Gif>` component has the same props as `<AnimatedImage>` but only supports GIF files.
@@ -0,0 +1,134 @@
---
name: images
description: Embedding images in Remotion using the <Img> component
metadata:
tags: images, img, staticFile, png, jpg, svg, webp
---
# Using images in Remotion
## The `<Img>` component
Always use the `<Img>` component from `remotion` to display images:
```tsx
import { Img, staticFile } from "remotion";
export const MyComposition = () => {
return <Img src={staticFile("photo.png")} />;
};
```
## Important restrictions
**You MUST use the `<Img>` component from `remotion`.** Do not use:
- Native HTML `<img>` elements
- Next.js `<Image>` component
- CSS `background-image`
The `<Img>` component ensures images are fully loaded before rendering, preventing flickering and blank frames during video export.
## Local images with staticFile()
Place images in the `public/` folder and use `staticFile()` to reference them:
```
my-video/
├─ public/
│ ├─ logo.png
│ ├─ avatar.jpg
│ └─ icon.svg
├─ src/
├─ package.json
```
```tsx
import { Img, staticFile } from "remotion";
<Img src={staticFile("logo.png")} />;
```
## Remote images
Remote URLs can be used directly without `staticFile()`:
```tsx
<Img src="https://example.com/image.png" />
```
Ensure remote images have CORS enabled.
For animated GIFs, use the `<Gif>` component from `@remotion/gif` instead.
## Sizing and positioning
Use the `style` prop to control size and position:
```tsx
<Img
src={staticFile("photo.png")}
style={{
width: 500,
height: 300,
position: "absolute",
top: 100,
left: 50,
objectFit: "cover",
}}
/>
```
## Dynamic image paths
Use template literals for dynamic file references:
```tsx
import { Img, staticFile, useCurrentFrame } from "remotion";
const frame = useCurrentFrame();
// Image sequence
<Img src={staticFile(`frames/frame${frame}.png`)} />
// Selecting based on props
<Img src={staticFile(`avatars/${props.userId}.png`)} />
// Conditional images
<Img src={staticFile(`icons/${isActive ? "active" : "inactive"}.svg`)} />
```
This pattern is useful for:
- Image sequences (frame-by-frame animations)
- User-specific avatars or profile images
- Theme-based icons
- State-dependent graphics
## Getting image dimensions
Use `getImageDimensions()` to get the dimensions of an image:
```tsx
import { getImageDimensions, staticFile } from "remotion";
const { width, height } = await getImageDimensions(staticFile("photo.png"));
```
This is useful for calculating aspect ratios or sizing compositions:
```tsx
import {
getImageDimensions,
staticFile,
CalculateMetadataFunction,
} from "remotion";
const calculateMetadata: CalculateMetadataFunction = async () => {
const { width, height } = await getImageDimensions(staticFile("photo.png"));
return {
width,
height,
};
};
```
@@ -0,0 +1,69 @@
---
name: import-srt-captions
description: Importing .srt subtitle files into Remotion using @remotion/captions
metadata:
tags: captions, subtitles, srt, import, parse
---
# Importing .srt subtitles into Remotion
If you have an existing `.srt` subtitle file, you can import it into Remotion using `parseSrt()` from `@remotion/captions`.
If you don't have a .srt file, read [Transcribing audio](transcribe-captions.md) for how to generate captions instead.
## Prerequisites
First, the @remotion/captions package needs to be installed.
If it is not installed, use the following command:
```bash
npx remotion add @remotion/captions # If project uses npm
bunx remotion add @remotion/captions # If project uses bun
yarn remotion add @remotion/captions # If project uses yarn
pnpm exec remotion add @remotion/captions # If project uses pnpm
```
## Reading an .srt file
Use `staticFile()` to reference an `.srt` file in your `public` folder, then fetch and parse it:
```tsx
import { useState, useEffect, useCallback } from "react";
import { AbsoluteFill, staticFile, useDelayRender } from "remotion";
import { parseSrt } from "@remotion/captions";
import type { Caption } from "@remotion/captions";
export const MyComponent: React.FC = () => {
const [captions, setCaptions] = useState<Caption[] | null>(null);
const { delayRender, continueRender, cancelRender } = useDelayRender();
const [handle] = useState(() => delayRender());
const fetchCaptions = useCallback(async () => {
try {
const response = await fetch(staticFile("subtitles.srt"));
const text = await response.text();
const { captions: parsed } = parseSrt({ input: text });
setCaptions(parsed);
continueRender(handle);
} catch (e) {
cancelRender(e);
}
}, [continueRender, cancelRender, handle]);
useEffect(() => {
fetchCaptions();
}, [fetchCaptions]);
if (!captions) {
return null;
}
return <AbsoluteFill>{/* Use captions here */}</AbsoluteFill>;
};
```
Remote URLs are also supported - you can `fetch()` a remote file via URL instead of using `staticFile()`.
## Using imported captions
Once parsed, the captions are in the `Caption` format and can be used with all `@remotion/captions` utilities.
@@ -0,0 +1,73 @@
---
name: light-leaks
description: Light leak overlay effects for Remotion using @remotion/light-leaks.
metadata:
tags: light-leaks, overlays, effects, transitions
---
## Light Leaks
This only works from Remotion 4.0.415 and up. Use `npx remotion versions` to check your Remotion version and `npx remotion upgrade` to upgrade your Remotion version.
`<LightLeak>` from `@remotion/light-leaks` renders a WebGL-based light leak effect. It reveals during the first half of its duration and retracts during the second half.
Typically used inside a `<TransitionSeries.Overlay>` to play over the cut point between two scenes. See the **transitions** rule for `<TransitionSeries>` and overlay usage.
## Prerequisites
```bash
npx remotion add @remotion/light-leaks
```
## Basic usage with TransitionSeries
```tsx
import { TransitionSeries } from "@remotion/transitions";
import { LightLeak } from "@remotion/light-leaks";
<TransitionSeries>
<TransitionSeries.Sequence durationInFrames={60}>
<SceneA />
</TransitionSeries.Sequence>
<TransitionSeries.Overlay durationInFrames={30}>
<LightLeak />
</TransitionSeries.Overlay>
<TransitionSeries.Sequence durationInFrames={60}>
<SceneB />
</TransitionSeries.Sequence>
</TransitionSeries>;
```
## Props
- `durationInFrames?` — defaults to the parent sequence/composition duration. The effect reveals during the first half and retracts during the second half.
- `seed?` — determines the shape of the light leak pattern. Different seeds produce different patterns. Default: `0`.
- `hueShift?` — rotates the hue in degrees (`0`–`360`). Default: `0` (yellow-to-orange). `120` = green, `240` = blue.
## Customizing the look
```tsx
import { LightLeak } from "@remotion/light-leaks";
// Blue-tinted light leak with a different pattern
<LightLeak seed={5} hueShift={240} />;
// Green-tinted light leak
<LightLeak seed={2} hueShift={120} />;
```
## Standalone usage
`<LightLeak>` can also be used outside of `<TransitionSeries>`, for example as a decorative overlay in any composition:
```tsx
import { AbsoluteFill } from "remotion";
import { LightLeak } from "@remotion/light-leaks";
const MyComp: React.FC = () => (
<AbsoluteFill>
<MyContent />
<LightLeak durationInFrames={60} seed={3} />
</AbsoluteFill>
);
```
@@ -0,0 +1,70 @@
---
name: lottie
description: Embedding Lottie animations in Remotion.
metadata:
category: Animation
---
# Using Lottie Animations in Remotion
## Prerequisites
First, the @remotion/lottie package needs to be installed.
If it is not, use the following command:
```bash
npx remotion add @remotion/lottie # If project uses npm
bunx remotion add @remotion/lottie # If project uses bun
yarn remotion add @remotion/lottie # If project uses yarn
pnpm exec remotion add @remotion/lottie # If project uses pnpm
```
## Displaying a Lottie file
To import a Lottie animation:
- Fetch the Lottie asset
- Wrap the loading process in `delayRender()` and `continueRender()`
- Save the animation data in a state
- Render the Lottie animation using the `Lottie` component from the `@remotion/lottie` package
```tsx
import { Lottie, LottieAnimationData } from "@remotion/lottie";
import { useEffect, useState } from "react";
import { cancelRender, continueRender, delayRender } from "remotion";
export const MyAnimation = () => {
const [handle] = useState(() => delayRender("Loading Lottie animation"));
const [animationData, setAnimationData] =
useState<LottieAnimationData | null>(null);
useEffect(() => {
fetch("https://assets4.lottiefiles.com/packages/lf20_zyquagfl.json")
.then((data) => data.json())
.then((json) => {
setAnimationData(json);
continueRender(handle);
})
.catch((err) => {
cancelRender(err);
});
}, [handle]);
if (!animationData) {
return null;
}
return <Lottie animationData={animationData} />;
};
```
## Styling and animating
Lottie supports the `style` prop to allow styles and animations:
```tsx
return (
<Lottie animationData={animationData} style={{ width: 400, height: 400 }} />
);
```
@@ -0,0 +1,412 @@
---
name: maps
description: Make map animations with Mapbox
metadata:
tags: map, map animation, mapbox
---
Maps can be added to a Remotion video with Mapbox.
The [Mapbox documentation](https://docs.mapbox.com/mapbox-gl-js/api/) has the API reference.
## Prerequisites
Mapbox and `@turf/turf` need to be installed.
Search the project for lockfiles and run the correct command depending on the package manager:
If `package-lock.json` is found, use the following command:
```bash
npm i mapbox-gl @turf/turf @types/mapbox-gl
```
If `bun.lock` is found, use the following command:
```bash
bun i mapbox-gl @turf/turf @types/mapbox-gl
```
If `yarn.lock` is found, use the following command:
```bash
yarn add mapbox-gl @turf/turf @types/mapbox-gl
```
If `pnpm-lock.yaml` is found, use the following command:
```bash
pnpm i mapbox-gl @turf/turf @types/mapbox-gl
```
The user needs to create a free Mapbox account and create an access token by visiting https://console.mapbox.com/account/access-tokens/.
The mapbox token needs to be added to the `.env` file:
```txt title=".env"
REMOTION_MAPBOX_TOKEN==pk.your-mapbox-access-token
```
## Adding a map
Here is a basic example of a map in Remotion.
```tsx
import { useEffect, useMemo, useRef, useState } from "react";
import { AbsoluteFill, useDelayRender, useVideoConfig } from "remotion";
import mapboxgl, { Map } from "mapbox-gl";
export const lineCoordinates = [
[6.56158447265625, 46.059891147620725],
[6.5691375732421875, 46.05679376154153],
[6.5842437744140625, 46.05059898938315],
[6.594886779785156, 46.04702502069337],
[6.601066589355469, 46.0460718554722],
[6.6089630126953125, 46.0365370783104],
[6.6185760498046875, 46.018420689207964],
];
mapboxgl.accessToken = process.env.REMOTION_MAPBOX_TOKEN as string;
export const MyComposition = () => {
const ref = useRef<HTMLDivElement>(null);
const { delayRender, continueRender } = useDelayRender();
const { width, height } = useVideoConfig();
const [handle] = useState(() => delayRender("Loading map..."));
const [map, setMap] = useState<Map | null>(null);
useEffect(() => {
const _map = new Map({
container: ref.current!,
zoom: 11.53,
center: [6.5615, 46.0598],
pitch: 65,
bearing: 0,
style: "⁠mapbox://styles/mapbox/standard",
interactive: false,
fadeDuration: 0,
});
_map.on("style.load", () => {
// Hide all features from the Mapbox Standard style
const hideFeatures = [
"showRoadsAndTransit",
"showRoads",
"showTransit",
"showPedestrianRoads",
"showRoadLabels",
"showTransitLabels",
"showPlaceLabels",
"showPointOfInterestLabels",
"showPointsOfInterest",
"showAdminBoundaries",
"showLandmarkIcons",
"showLandmarkIconLabels",
"show3dObjects",
"show3dBuildings",
"show3dTrees",
"show3dLandmarks",
"show3dFacades",
];
for (const feature of hideFeatures) {
_map.setConfigProperty("basemap", feature, false);
}
_map.setConfigProperty("basemap", "colorTrunks", "rgba(0, 0, 0, 0)");
_map.addSource("trace", {
type: "geojson",
data: {
type: "Feature",
properties: {},
geometry: {
type: "LineString",
coordinates: lineCoordinates,
},
},
});
_map.addLayer({
type: "line",
source: "trace",
id: "line",
paint: {
"line-color": "black",
"line-width": 5,
},
layout: {
"line-cap": "round",
"line-join": "round",
},
});
});
_map.on("load", () => {
continueRender(handle);
setMap(_map);
});
}, [handle, lineCoordinates]);
const style: React.CSSProperties = useMemo(
() => ({ width, height, position: "absolute" }),
[width, height],
);
return <AbsoluteFill ref={ref} style={style} />;
};
```
The following is important in Remotion:
- Animations must be driven by `useCurrentFrame()` and animations that Mapbox brings itself should be disabled. For example, the `fadeDuration` prop should be set to `0`, `interactive` should be set to `false`, etc.
- Loading the map should be delayed using `useDelayRender()` and the map should be set to `null` until it is loaded.
- The element containing the ref MUST have an explicit width and height and `position: "absolute"`.
- Do not add a `_map.remove();` cleanup function.
## Drawing lines
Unless I request it, do not add a glow effect to the lines.
Unless I request it, do not add additional points to the lines.
## Map style
By default, use the `mapbox://styles/mapbox/standard` style.
Hide the labels from the base map style.
Unless I request otherwise, remove all features from the Mapbox Standard style.
```tsx
// Hide all features from the Mapbox Standard style
const hideFeatures = [
"showRoadsAndTransit",
"showRoads",
"showTransit",
"showPedestrianRoads",
"showRoadLabels",
"showTransitLabels",
"showPlaceLabels",
"showPointOfInterestLabels",
"showPointsOfInterest",
"showAdminBoundaries",
"showLandmarkIcons",
"showLandmarkIconLabels",
"show3dObjects",
"show3dBuildings",
"show3dTrees",
"show3dLandmarks",
"show3dFacades",
];
for (const feature of hideFeatures) {
_map.setConfigProperty("basemap", feature, false);
}
_map.setConfigProperty("basemap", "colorMotorways", "transparent");
_map.setConfigProperty("basemap", "colorRoads", "transparent");
_map.setConfigProperty("basemap", "colorTrunks", "transparent");
```
## Animating the camera
You can animate the camera along the line by adding a `useEffect` hook that updates the camera position based on the current frame.
Unless I ask for it, do not jump between camera angles.
```tsx
import * as turf from "@turf/turf";
import { interpolate } from "remotion";
import { Easing } from "remotion";
import { useCurrentFrame, useVideoConfig, useDelayRender } from "remotion";
const animationDuration = 20;
const cameraAltitude = 4000;
```
```tsx
const frame = useCurrentFrame();
const { fps } = useVideoConfig();
const { delayRender, continueRender } = useDelayRender();
useEffect(() => {
if (!map) {
return;
}
const handle = delayRender("Moving point...");
const routeDistance = turf.length(turf.lineString(lineCoordinates));
const progress = interpolate(
frame / fps,
[0.00001, animationDuration],
[0, 1],
{
easing: Easing.inOut(Easing.sin),
extrapolateLeft: "clamp",
extrapolateRight: "clamp",
},
);
const camera = map.getFreeCameraOptions();
const alongRoute = turf.along(
turf.lineString(lineCoordinates),
routeDistance * progress,
).geometry.coordinates;
camera.lookAtPoint({
lng: alongRoute[0],
lat: alongRoute[1],
});
map.setFreeCameraOptions(camera);
map.once("idle", () => continueRender(handle));
}, [lineCoordinates, fps, frame, handle, map]);
```
Notes:
IMPORTANT: Keep the camera by default so north is up.
IMPORTANT: For multi-step animations, set all properties at all stages (zoom, position, line progress) to prevent jumps. Override initial values.
- The progress is clamped to a minimum value to avoid the line being empty, which can lead to turf errors
- See [Timing](./timing.md) for more options for timing.
- Consider the dimensions of the composition and make the lines thick enough and the label font size large enough to be legible for when the composition is scaled down.
## Animating lines
### Straight lines (linear interpolation)
To animate a line that appears straight on the map, use linear interpolation between coordinates. Do NOT use turf's `lineSliceAlong` or `along` functions, as they use geodesic (great circle) calculations which appear curved on a Mercator projection.
```tsx
const frame = useCurrentFrame();
const { durationInFrames } = useVideoConfig();
useEffect(() => {
if (!map) return;
const animationHandle = delayRender("Animating line...");
const progress = interpolate(frame, [0, durationInFrames - 1], [0, 1], {
extrapolateLeft: "clamp",
extrapolateRight: "clamp",
easing: Easing.inOut(Easing.cubic),
});
// Linear interpolation for a straight line on the map
const start = lineCoordinates[0];
const end = lineCoordinates[1];
const currentLng = start[0] + (end[0] - start[0]) * progress;
const currentLat = start[1] + (end[1] - start[1]) * progress;
const lineData: GeoJSON.Feature<GeoJSON.LineString> = {
type: "Feature",
properties: {},
geometry: {
type: "LineString",
coordinates: [start, [currentLng, currentLat]],
},
};
const source = map.getSource("trace") as mapboxgl.GeoJSONSource;
if (source) {
source.setData(lineData);
}
map.once("idle", () => continueRender(animationHandle));
}, [frame, map, durationInFrames]);
```
### Curved lines (geodesic/great circle)
To animate a line that follows the geodesic (great circle) path between two points, use turf's `lineSliceAlong`. This is useful for showing flight paths or the actual shortest distance on Earth.
```tsx
import * as turf from "@turf/turf";
const routeLine = turf.lineString(lineCoordinates);
const routeDistance = turf.length(routeLine);
const currentDistance = Math.max(0.001, routeDistance * progress);
const slicedLine = turf.lineSliceAlong(routeLine, 0, currentDistance);
const source = map.getSource("route") as mapboxgl.GeoJSONSource;
if (source) {
source.setData(slicedLine);
}
```
## Markers
Add labels, and markers where appropriate.
```tsx
_map.addSource("markers", {
type: "geojson",
data: {
type: "FeatureCollection",
features: [
{
type: "Feature",
properties: { name: "Point 1" },
geometry: { type: "Point", coordinates: [-118.2437, 34.0522] },
},
],
},
});
_map.addLayer({
id: "city-markers",
type: "circle",
source: "markers",
paint: {
"circle-radius": 40,
"circle-color": "#FF4444",
"circle-stroke-width": 4,
"circle-stroke-color": "#FFFFFF",
},
});
_map.addLayer({
id: "labels",
type: "symbol",
source: "markers",
layout: {
"text-field": ["get", "name"],
"text-font": ["DIN Pro Bold", "Arial Unicode MS Bold"],
"text-size": 50,
"text-offset": [0, 0.5],
"text-anchor": "top",
},
paint: {
"text-color": "#FFFFFF",
"text-halo-color": "#000000",
"text-halo-width": 2,
},
});
```
Make sure they are big enough. Check the composition dimensions and scale the labels accordingly.
For a composition size of 1920x1080, the label font size should be at least 40px.
IMPORTANT: Keep the `text-offset` small enough so it is close to the marker. Consider the marker circle radius. For a circle radius of 40, this is a good offset:
```tsx
"text-offset": [0, 0.5],
```
## 3D buildings
To enable 3D buildings, use the following code:
```tsx
_map.setConfigProperty("basemap", "show3dObjects", true);
_map.setConfigProperty("basemap", "show3dLandmarks", true);
_map.setConfigProperty("basemap", "show3dBuildings", true);
```
## Rendering
When rendering a map animation, make sure to render with the following flags:
```
npx remotion render --gl=angle --concurrency=1
```
@@ -0,0 +1,34 @@
---
name: measuring-dom-nodes
description: Measuring DOM element dimensions in Remotion
metadata:
tags: measure, layout, dimensions, getBoundingClientRect, scale
---
# Measuring DOM nodes in Remotion
Remotion applies a `scale()` transform to the video container, which affects values from `getBoundingClientRect()`. Use `useCurrentScale()` to get correct measurements.
## Measuring element dimensions
```tsx
import { useCurrentScale } from "remotion";
import { useRef, useEffect, useState } from "react";
export const MyComponent = () => {
const ref = useRef<HTMLDivElement>(null);
const scale = useCurrentScale();
const [dimensions, setDimensions] = useState({ width: 0, height: 0 });
useEffect(() => {
if (!ref.current) return;
const rect = ref.current.getBoundingClientRect();
setDimensions({
width: rect.width / scale,
height: rect.height / scale,
});
}, [scale]);
return <div ref={ref}>Content to measure</div>;
};
```
@@ -0,0 +1,140 @@
---
name: measuring-text
description: Measuring text dimensions, fitting text to containers, and checking overflow
metadata:
tags: measure, text, layout, dimensions, fitText, fillTextBox
---
# Measuring text in Remotion
## Prerequisites
Install @remotion/layout-utils if it is not already installed:
```bash
npx remotion add @remotion/layout-utils
```
## Measuring text dimensions
Use `measureText()` to calculate the width and height of text:
```tsx
import { measureText } from "@remotion/layout-utils";
const { width, height } = measureText({
text: "Hello World",
fontFamily: "Arial",
fontSize: 32,
fontWeight: "bold",
});
```
Results are cached - duplicate calls return the cached result.
## Fitting text to a width
Use `fitText()` to find the optimal font size for a container:
```tsx
import { fitText } from "@remotion/layout-utils";
const { fontSize } = fitText({
text: "Hello World",
withinWidth: 600,
fontFamily: "Inter",
fontWeight: "bold",
});
return (
<div
style={{
fontSize: Math.min(fontSize, 80), // Cap at 80px
fontFamily: "Inter",
fontWeight: "bold",
}}
>
Hello World
</div>
);
```
## Checking text overflow
Use `fillTextBox()` to check if text exceeds a box:
```tsx
import { fillTextBox } from "@remotion/layout-utils";
const box = fillTextBox({ maxBoxWidth: 400, maxLines: 3 });
const words = ["Hello", "World", "This", "is", "a", "test"];
for (const word of words) {
const { exceedsBox } = box.add({
text: word + " ",
fontFamily: "Arial",
fontSize: 24,
});
if (exceedsBox) {
// Text would overflow, handle accordingly
break;
}
}
```
## Best practices
**Load fonts first:** Only call measurement functions after fonts are loaded.
```tsx
import { loadFont } from "@remotion/google-fonts/Inter";
const { fontFamily, waitUntilDone } = loadFont("normal", {
weights: ["400"],
subsets: ["latin"],
});
waitUntilDone().then(() => {
// Now safe to measure
const { width } = measureText({
text: "Hello",
fontFamily,
fontSize: 32,
});
});
```
**Use validateFontIsLoaded:** Catch font loading issues early:
```tsx
measureText({
text: "Hello",
fontFamily: "MyCustomFont",
fontSize: 32,
validateFontIsLoaded: true, // Throws if font not loaded
});
```
**Match font properties:** Use the same properties for measurement and rendering:
```tsx
const fontStyle = {
fontFamily: "Inter",
fontSize: 32,
fontWeight: "bold" as const,
letterSpacing: "0.5px",
};
const { width } = measureText({
text: "Hello",
...fontStyle,
});
return <div style={fontStyle}>Hello</div>;
```
**Avoid padding and border:** Use `outline` instead of `border` to prevent layout differences:
```tsx
<div style={{ outline: "2px solid red" }}>Text</div>
```
@@ -0,0 +1,109 @@
---
name: parameters
description: Make a video parametrizable by adding a Zod schema
metadata:
tags: parameters, zod, schema
---
To make a video parametrizable, a Zod schema can be added to a composition.
First, `zod` must be installed .
Search the project for lockfiles and run the correct command depending on the package manager:
If `package-lock.json` is found, use the following command:
```bash
npm i zod
```
If `bun.lockb` is found, use the following command:
```bash
bun i zod
```
If `yarn.lock` is found, use the following command:
```bash
yarn add zod
```
If `pnpm-lock.yaml` is found, use the following command:
```bash
pnpm i zod
```
Then, a Zod schema can be defined alongside the component:
```tsx title="src/MyComposition.tsx"
import { z } from "zod";
export const MyCompositionSchema = z.object({
title: z.string(),
});
const MyComponent: React.FC<z.infer<typeof MyCompositionSchema>> = () => {
return (
<div>
<h1>{props.title}</h1>
</div>
);
};
```
In the root file, the schema can be passed to the composition:
```tsx title="src/Root.tsx"
import { Composition } from "remotion";
import { MycComponent, MyCompositionSchema } from "./MyComposition";
export const RemotionRoot = () => {
return (
<Composition
id="MyComposition"
component={MyComponent}
durationInFrames={100}
fps={30}
width={1080}
height={1080}
defaultProps={{ title: "Hello World" }}
schema={MyCompositionSchema}
/>
);
};
```
Now, the user can edit the parameter visually in the sidebar.
All schemas that are supported by Zod are supported by Remotion.
Remotion requires that the top-level type is a z.object(), because the collection of props of a React component is always an object.
## Color picker
For adding a color picker, use `zColor()` from `@remotion/zod-types`.
If it is not installed, use the following command:
```bash
npx remotion add @remotion/zod-types # If project uses npm
bunx remotion add @remotion/zod-types # If project uses bun
yarn remotion add @remotion/zod-types # If project uses yarn
pnpm exec remotion add @remotion/zod-types # If project uses pnpm
```
Then import `zColor` from `@remotion/zod-types`:
```tsx
import { zColor } from "@remotion/zod-types";
```
Then use it in the schema:
```tsx
export const MyCompositionSchema = z.object({
color: zColor(),
});
```
@@ -0,0 +1,118 @@
---
name: sequencing
description: Sequencing patterns for Remotion - delay, trim, limit duration of items
metadata:
tags: sequence, series, timing, delay, trim
---
Use `<Sequence>` to delay when an element appears in the timeline.
```tsx
import { Sequence } from "remotion";
const {fps} = useVideoConfig();
<Sequence from={1 * fps} durationInFrames={2 * fps} premountFor={1 * fps}>
<Title />
</Sequence>
<Sequence from={2 * fps} durationInFrames={2 * fps} premountFor={1 * fps}>
<Subtitle />
</Sequence>
```
This will by default wrap the component in an absolute fill element.
If the items should not be wrapped, use the `layout` prop:
```tsx
<Sequence layout="none">
<Title />
</Sequence>
```
## Premounting
This loads the component in the timeline before it is actually played.
Always premount any `<Sequence>`!
```tsx
<Sequence premountFor={1 * fps}>
<Title />
</Sequence>
```
## Series
Use `<Series>` when elements should play one after another without overlap.
```tsx
import { Series } from "remotion";
<Series>
<Series.Sequence durationInFrames={45}>
<Intro />
</Series.Sequence>
<Series.Sequence durationInFrames={60}>
<MainContent />
</Series.Sequence>
<Series.Sequence durationInFrames={30}>
<Outro />
</Series.Sequence>
</Series>;
```
Same as with `<Sequence>`, the items will be wrapped in an absolute fill element by default when using `<Series.Sequence>`, unless the `layout` prop is set to `none`.
### Series with overlaps
Use negative offset for overlapping sequences:
```tsx
<Series>
<Series.Sequence durationInFrames={60}>
<SceneA />
</Series.Sequence>
<Series.Sequence offset={-15} durationInFrames={60}>
{/* Starts 15 frames before SceneA ends */}
<SceneB />
</Series.Sequence>
</Series>
```
## Frame References Inside Sequences
Inside a Sequence, `useCurrentFrame()` returns the local frame (starting from 0):
```tsx
<Sequence from={60} durationInFrames={30}>
<MyComponent />
{/* Inside MyComponent, useCurrentFrame() returns 0-29, not 60-89 */}
</Sequence>
```
## Nested Sequences
Sequences can be nested for complex timing:
```tsx
<Sequence from={0} durationInFrames={120}>
<Background />
<Sequence from={15} durationInFrames={90} layout="none">
<Title />
</Sequence>
<Sequence from={45} durationInFrames={60} layout="none">
<Subtitle />
</Sequence>
</Sequence>
```
## Nesting compositions within another
To add a composition within another composition, you can use the `<Sequence>` component with a `width` and `height` prop to specify the size of the composition.
```tsx
<AbsoluteFill>
<Sequence width={COMPOSITION_WIDTH} height={COMPOSITION_HEIGHT}>
<CompositionComponent />
</Sequence>
</AbsoluteFill>
```
@@ -0,0 +1,26 @@
---
name: sfx
description: Including sound effects
metadata:
tags: sfx, sound, effect, audio
---
To include a sound effect, use the `<Audio>` tag:
```tsx
import { Audio } from "@remotion/sfx";
<Audio src={"https://remotion.media/whoosh.wav"} />;
```
The following sound effects are available:
- `https://remotion.media/whoosh.wav`
- `https://remotion.media/whip.wav`
- `https://remotion.media/page-turn.wav`
- `https://remotion.media/switch.wav`
- `https://remotion.media/mouse-click.wav`
- `https://remotion.media/shutter-modern.wav`
- `https://remotion.media/shutter-old.wav`
For more sound effects, search the internet. A good resource is https://github.com/kapishdima/soundcn/tree/main/assets.
@@ -0,0 +1,36 @@
---
name: subtitles
description: subtitles and caption rules
metadata:
tags: subtitles, captions, remotion, json
---
All captions must be processed in JSON. The captions must use the `Caption` type which is the following:
```ts
import type { Caption } from "@remotion/captions";
```
This is the definition:
```ts
type Caption = {
text: string;
startMs: number;
endMs: number;
timestampMs: number | null;
confidence: number | null;
};
```
## Generating captions
To transcribe video and audio files to generate captions, load the [./transcribe-captions.md](./transcribe-captions.md) file for more instructions.
## Displaying captions
To display captions in your video, load the [./display-captions.md](./display-captions.md) file for more instructions.
## Importing captions
To import captions from a .srt file, load the [./import-srt-captions.md](./import-srt-captions.md) file for more instructions.
@@ -0,0 +1,11 @@
---
name: tailwind
description: Using TailwindCSS in Remotion.
metadata:
---
You can and should use TailwindCSS in Remotion, if TailwindCSS is installed in the project.
Don't use `transition-*` or `animate-*` classes - always animate using the `useCurrentFrame()` hook.
Tailwind must be installed and enabled first in a Remotion project - fetch https://www.remotion.dev/docs/tailwind using WebFetch for instructions.
@@ -0,0 +1,20 @@
---
name: text-animations
description: Typography and text animation patterns for Remotion.
metadata:
tags: typography, text, typewriter, highlighter ken
---
## Text animations
Based on `useCurrentFrame()`, reduce the string character by character to create a typewriter effect.
## Typewriter Effect
See [Typewriter](assets/text-animations-typewriter.tsx) for an advanced example with a blinking cursor and a pause after the first sentence.
Always use string slicing for typewriter effects. Never use per-character opacity.
## Word Highlighting
See [Word Highlight](assets/text-animations-word-highlight.tsx) for an example for how a word highlight is animated, like with a highlighter pen.
@@ -0,0 +1,179 @@
---
name: timing
description: Interpolation curves in Remotion - linear, easing, spring animations
metadata:
tags: spring, bounce, easing, interpolation
---
A simple linear interpolation is done using the `interpolate` function.
```ts title="Going from 0 to 1 over 100 frames"
import { interpolate } from "remotion";
const opacity = interpolate(frame, [0, 100], [0, 1]);
```
By default, the values are not clamped, so the value can go outside the range [0, 1].
Here is how they can be clamped:
```ts title="Going from 0 to 1 over 100 frames with extrapolation"
const opacity = interpolate(frame, [0, 100], [0, 1], {
extrapolateRight: "clamp",
extrapolateLeft: "clamp",
});
```
## Spring animations
Spring animations have a more natural motion.
They go from 0 to 1 over time.
```ts title="Spring animation from 0 to 1 over 100 frames"
import { spring, useCurrentFrame, useVideoConfig } from "remotion";
const frame = useCurrentFrame();
const { fps } = useVideoConfig();
const scale = spring({
frame,
fps,
});
```
### Physical properties
The default configuration is: `mass: 1, damping: 10, stiffness: 100`.
This leads to the animation having a bit of bounce before it settles.
The config can be overwritten like this:
```ts
const scale = spring({
frame,
fps,
config: { damping: 200 },
});
```
The recommended configuration for a natural motion without a bounce is: `{ damping: 200 }`.
Here are some common configurations:
```tsx
const smooth = { damping: 200 }; // Smooth, no bounce (subtle reveals)
const snappy = { damping: 20, stiffness: 200 }; // Snappy, minimal bounce (UI elements)
const bouncy = { damping: 8 }; // Bouncy entrance (playful animations)
const heavy = { damping: 15, stiffness: 80, mass: 2 }; // Heavy, slow, small bounce
```
### Delay
The animation starts immediately by default.
Use the `delay` parameter to delay the animation by a number of frames.
```tsx
const entrance = spring({
frame: frame - ENTRANCE_DELAY,
fps,
delay: 20,
});
```
### Duration
A `spring()` has a natural duration based on the physical properties.
To stretch the animation to a specific duration, use the `durationInFrames` parameter.
```tsx
const spring = spring({
frame,
fps,
durationInFrames: 40,
});
```
### Combining spring() with interpolate()
Map spring output (0-1) to custom ranges:
```tsx
const springProgress = spring({
frame,
fps,
});
// Map to rotation
const rotation = interpolate(springProgress, [0, 1], [0, 360]);
<div style={{ rotate: rotation + "deg" }} />;
```
### Adding springs
Springs return just numbers, so math can be performed:
```tsx
const frame = useCurrentFrame();
const { fps, durationInFrames } = useVideoConfig();
const inAnimation = spring({
frame,
fps,
});
const outAnimation = spring({
frame,
fps,
durationInFrames: 1 * fps,
delay: durationInFrames - 1 * fps,
});
const scale = inAnimation - outAnimation;
```
## Easing
Easing can be added to the `interpolate` function:
```ts
import { interpolate, Easing } from "remotion";
const value1 = interpolate(frame, [0, 100], [0, 1], {
easing: Easing.inOut(Easing.quad),
extrapolateLeft: "clamp",
extrapolateRight: "clamp",
});
```
The default easing is `Easing.linear`.
There are various other convexities:
- `Easing.in` for starting slow and accelerating
- `Easing.out` for starting fast and slowing down
- `Easing.inOut`
and curves (sorted from most linear to most curved):
- `Easing.quad`
- `Easing.sin`
- `Easing.exp`
- `Easing.circle`
Convexities and curves need be combined for an easing function:
```ts
const value1 = interpolate(frame, [0, 100], [0, 1], {
easing: Easing.inOut(Easing.quad),
extrapolateLeft: "clamp",
extrapolateRight: "clamp",
});
```
Cubic bezier curves are also supported:
```ts
const value1 = interpolate(frame, [0, 100], [0, 1], {
easing: Easing.bezier(0.8, 0.22, 0.96, 0.65),
extrapolateLeft: "clamp",
extrapolateRight: "clamp",
});
```
@@ -0,0 +1,70 @@
---
name: transcribe-captions
description: Transcribing audio to generate captions in Remotion
metadata:
tags: captions, transcribe, whisper, audio, speech-to-text
---
# Transcribing audio
To transcribe audio to generate captions in Remotion, you can use the [`transcribe()`](https://www.remotion.dev/docs/install-whisper-cpp/transcribe) function from the [`@remotion/install-whisper-cpp`](https://www.remotion.dev/docs/install-whisper-cpp) package.
## Prerequisites
First, the @remotion/install-whisper-cpp package needs to be installed.
If it is not installed, use the following command:
```bash
npx remotion add @remotion/install-whisper-cpp
```
## Transcribing
Make a Node.js script to download Whisper.cpp and a model, and transcribe the audio.
```ts
import path from "path";
import {
downloadWhisperModel,
installWhisperCpp,
transcribe,
toCaptions,
} from "@remotion/install-whisper-cpp";
import fs from "fs";
const to = path.join(process.cwd(), "whisper.cpp");
await installWhisperCpp({
to,
version: "1.5.5",
});
await downloadWhisperModel({
model: "medium.en",
folder: to,
});
// Convert the audio to a 16KHz wav file first if needed:
// import {execSync} from 'child_process';
// execSync('ffmpeg -i /path/to/audio.mp4 -ar 16000 /path/to/audio.wav -y');
const whisperCppOutput = await transcribe({
model: "medium.en",
whisperPath: to,
whisperCppVersion: "1.5.5",
inputPath: "/path/to/audio123.wav",
tokenLevelTimestamps: true,
});
// Optional: Apply our recommended postprocessing
const { captions } = toCaptions({
whisperCppOutput,
});
// Write it to the public/ folder so it can be fetched from Remotion
fs.writeFileSync("captions123.json", JSON.stringify(captions, null, 2));
```
Transcribe each clip individually and create multiple JSON files.
See [Displaying captions](display-captions.md) for how to display the captions in Remotion.
@@ -0,0 +1,197 @@
---
name: transitions
description: Scene transitions and overlays for Remotion using TransitionSeries.
metadata:
tags: transitions, overlays, fade, slide, wipe, scenes
---
## TransitionSeries
`<TransitionSeries>` arranges scenes and supports two ways to enhance the cut point between them:
- **Transitions** (`<TransitionSeries.Transition>`) — crossfade, slide, wipe, etc. between two scenes. Shortens the timeline because both scenes play simultaneously during the transition.
- **Overlays** (`<TransitionSeries.Overlay>`) — render an effect (e.g. a light leak) on top of the cut point without shortening the timeline.
Children are absolutely positioned.
## Prerequisites
```bash
npx remotion add @remotion/transitions
```
## Transition example
```tsx
import { TransitionSeries, linearTiming } from "@remotion/transitions";
import { fade } from "@remotion/transitions/fade";
<TransitionSeries>
<TransitionSeries.Sequence durationInFrames={60}>
<SceneA />
</TransitionSeries.Sequence>
<TransitionSeries.Transition
presentation={fade()}
timing={linearTiming({ durationInFrames: 15 })}
/>
<TransitionSeries.Sequence durationInFrames={60}>
<SceneB />
</TransitionSeries.Sequence>
</TransitionSeries>;
```
## Overlay example
Any React component can be used as an overlay. For a ready-made effect, see the **light-leaks** rule.
```tsx
import { TransitionSeries } from "@remotion/transitions";
import { LightLeak } from "@remotion/light-leaks";
<TransitionSeries>
<TransitionSeries.Sequence durationInFrames={60}>
<SceneA />
</TransitionSeries.Sequence>
<TransitionSeries.Overlay durationInFrames={20}>
<LightLeak />
</TransitionSeries.Overlay>
<TransitionSeries.Sequence durationInFrames={60}>
<SceneB />
</TransitionSeries.Sequence>
</TransitionSeries>;
```
## Mixing transitions and overlays
Transitions and overlays can coexist in the same `<TransitionSeries>`, but an overlay cannot be adjacent to a transition or another overlay.
```tsx
import { TransitionSeries, linearTiming } from "@remotion/transitions";
import { fade } from "@remotion/transitions/fade";
import { LightLeak } from "@remotion/light-leaks";
<TransitionSeries>
<TransitionSeries.Sequence durationInFrames={60}>
<SceneA />
</TransitionSeries.Sequence>
<TransitionSeries.Overlay durationInFrames={30}>
<LightLeak />
</TransitionSeries.Overlay>
<TransitionSeries.Sequence durationInFrames={60}>
<SceneB />
</TransitionSeries.Sequence>
<TransitionSeries.Transition
presentation={fade()}
timing={linearTiming({ durationInFrames: 15 })}
/>
<TransitionSeries.Sequence durationInFrames={60}>
<SceneC />
</TransitionSeries.Sequence>
</TransitionSeries>;
```
## Transition props
`<TransitionSeries.Transition>` requires:
- `presentation` — the visual effect (e.g. `fade()`, `slide()`, `wipe()`).
- `timing` — controls speed and easing (e.g. `linearTiming()`, `springTiming()`).
## Overlay props
`<TransitionSeries.Overlay>` accepts:
- `durationInFrames` — how long the overlay is visible (positive integer).
- `offset?` — shifts the overlay relative to the cut point center. Positive = later, negative = earlier. Default: `0`.
## Available transition types
Import transitions from their respective modules:
```tsx
import { fade } from "@remotion/transitions/fade";
import { slide } from "@remotion/transitions/slide";
import { wipe } from "@remotion/transitions/wipe";
import { flip } from "@remotion/transitions/flip";
import { clockWipe } from "@remotion/transitions/clock-wipe";
```
## Slide transition with direction
```tsx
import { slide } from "@remotion/transitions/slide";
<TransitionSeries.Transition
presentation={slide({ direction: "from-left" })}
timing={linearTiming({ durationInFrames: 20 })}
/>;
```
Directions: `"from-left"`, `"from-right"`, `"from-top"`, `"from-bottom"`
## Timing options
```tsx
import { linearTiming, springTiming } from "@remotion/transitions";
// Linear timing - constant speed
linearTiming({ durationInFrames: 20 });
// Spring timing - organic motion
springTiming({ config: { damping: 200 }, durationInFrames: 25 });
```
## Duration calculation
Transitions overlap adjacent scenes, so the total composition length is **shorter** than the sum of all sequence durations. Overlays do **not** affect the total duration.
For example, with two 60-frame sequences and a 15-frame transition:
- Without transitions: `60 + 60 = 120` frames
- With transition: `60 + 60 - 15 = 105` frames
Adding an overlay between two other sequences does not change the total.
### Getting the duration of a transition
Use the `getDurationInFrames()` method on the timing object:
```tsx
import { linearTiming, springTiming } from "@remotion/transitions";
const linearDuration = linearTiming({
durationInFrames: 20,
}).getDurationInFrames({ fps: 30 });
// Returns 20
const springDuration = springTiming({
config: { damping: 200 },
}).getDurationInFrames({ fps: 30 });
// Returns calculated duration based on spring physics
```
For `springTiming` without an explicit `durationInFrames`, the duration depends on `fps` because it calculates when the spring animation settles.
### Calculating total composition duration
```tsx
import { linearTiming } from "@remotion/transitions";
const scene1Duration = 60;
const scene2Duration = 60;
const scene3Duration = 60;
const timing1 = linearTiming({ durationInFrames: 15 });
const timing2 = linearTiming({ durationInFrames: 20 });
const transition1Duration = timing1.getDurationInFrames({ fps: 30 });
const transition2Duration = timing2.getDurationInFrames({ fps: 30 });
const totalDuration =
scene1Duration +
scene2Duration +
scene3Duration -
transition1Duration -
transition2Duration;
// 60 + 60 + 60 - 15 - 20 = 145 frames
```
@@ -0,0 +1,106 @@
---
name: transparent-videos
description: Rendering transparent videos in Remotion
metadata:
tags: transparent, alpha, codec, vp9, prores, webm
---
# Rendering Transparent Videos
Remotion can render transparent videos in two ways: as a ProRes video or as a WebM video.
## Transparent ProRes
Ideal for when importing into video editing software.
**CLI:**
```bash
npx remotion render --image-format=png --pixel-format=yuva444p10le --codec=prores --prores-profile=4444 MyComp out.mov
```
**Default in Studio** (restart Studio after changing):
```ts
// remotion.config.ts
import { Config } from "@remotion/cli/config";
Config.setVideoImageFormat("png");
Config.setPixelFormat("yuva444p10le");
Config.setCodec("prores");
Config.setProResProfile("4444");
```
**Setting it as the default export settings for a composition** (using `calculateMetadata`):
```tsx
import { CalculateMetadataFunction } from "remotion";
const calculateMetadata: CalculateMetadataFunction<Props> = async ({
props,
}) => {
return {
defaultCodec: "prores",
defaultVideoImageFormat: "png",
defaultPixelFormat: "yuva444p10le",
defaultProResProfile: "4444",
};
};
<Composition
id="my-video"
component={MyVideo}
durationInFrames={150}
fps={30}
width={1920}
height={1080}
calculateMetadata={calculateMetadata}
/>;
```
## Transparent WebM (VP9)
Ideal for when playing in a browser.
**CLI:**
```bash
npx remotion render --image-format=png --pixel-format=yuva420p --codec=vp9 MyComp out.webm
```
**Default in Studio** (restart Studio after changing):
```ts
// remotion.config.ts
import { Config } from "@remotion/cli/config";
Config.setVideoImageFormat("png");
Config.setPixelFormat("yuva420p");
Config.setCodec("vp9");
```
**Setting it as the default export settings for a composition** (using `calculateMetadata`):
```tsx
import { CalculateMetadataFunction } from "remotion";
const calculateMetadata: CalculateMetadataFunction<Props> = async ({
props,
}) => {
return {
defaultCodec: "vp8",
defaultVideoImageFormat: "png",
defaultPixelFormat: "yuva420p",
};
};
<Composition
id="my-video"
component={MyVideo}
durationInFrames={150}
fps={30}
width={1920}
height={1080}
calculateMetadata={calculateMetadata}
/>;
```
@@ -0,0 +1,51 @@
---
name: trimming
description: Trimming patterns for Remotion - cut the beginning or end of animations
metadata:
tags: sequence, trim, clip, cut, offset
---
Use `<Sequence>` with a negative `from` value to trim the start of an animation.
## Trim the Beginning
A negative `from` value shifts time backwards, making the animation start partway through:
```tsx
import { Sequence, useVideoConfig } from "remotion";
const fps = useVideoConfig();
<Sequence from={-0.5 * fps}>
<MyAnimation />
</Sequence>;
```
The animation appears 15 frames into its progress - the first 15 frames are trimmed off.
Inside `<MyAnimation>`, `useCurrentFrame()` starts at 15 instead of 0.
## Trim the End
Use `durationInFrames` to unmount content after a specified duration:
```tsx
<Sequence durationInFrames={1.5 * fps}>
<MyAnimation />
</Sequence>
```
The animation plays for 45 frames, then the component unmounts.
## Trim and Delay
Nest sequences to both trim the beginning and delay when it appears:
```tsx
<Sequence from={30}>
<Sequence from={-15}>
<MyAnimation />
</Sequence>
</Sequence>
```
The inner sequence trims 15 frames from the start, and the outer sequence delays the result by 30 frames.
@@ -0,0 +1,171 @@
---
name: videos
description: Embedding videos in Remotion - trimming, volume, speed, looping, pitch
metadata:
tags: video, media, trim, volume, speed, loop, pitch
---
# Using videos in Remotion
## Prerequisites
First, the @remotion/media package needs to be installed.
If it is not, use the following command:
```bash
npx remotion add @remotion/media # If project uses npm
bunx remotion add @remotion/media # If project uses bun
yarn remotion add @remotion/media # If project uses yarn
pnpm exec remotion add @remotion/media # If project uses pnpm
```
Use `<Video>` from `@remotion/media` to embed videos into your composition.
```tsx
import { Video } from "@remotion/media";
import { staticFile } from "remotion";
export const MyComposition = () => {
return <Video src={staticFile("video.mp4")} />;
};
```
Remote URLs are also supported:
```tsx
<Video src="https://remotion.media/video.mp4" />
```
## Trimming
Use `trimBefore` and `trimAfter` to remove portions of the video. Values are in seconds.
```tsx
const { fps } = useVideoConfig();
return (
<Video
src={staticFile("video.mp4")}
trimBefore={2 * fps} // Skip the first 2 seconds
trimAfter={10 * fps} // End at the 10 second mark
/>
);
```
## Delaying
Wrap the video in a `<Sequence>` to delay when it appears:
```tsx
import { Sequence, staticFile } from "remotion";
import { Video } from "@remotion/media";
const { fps } = useVideoConfig();
return (
<Sequence from={1 * fps}>
<Video src={staticFile("video.mp4")} />
</Sequence>
);
```
The video will appear after 1 second.
## Sizing and Position
Use the `style` prop to control size and position:
```tsx
<Video
src={staticFile("video.mp4")}
style={{
width: 500,
height: 300,
position: "absolute",
top: 100,
left: 50,
objectFit: "cover",
}}
/>
```
## Volume
Set a static volume (0 to 1):
```tsx
<Video src={staticFile("video.mp4")} volume={0.5} />
```
Or use a callback for dynamic volume based on the current frame:
```tsx
import { interpolate } from "remotion";
const { fps } = useVideoConfig();
return (
<Video
src={staticFile("video.mp4")}
volume={(f) =>
interpolate(f, [0, 1 * fps], [0, 1], { extrapolateRight: "clamp" })
}
/>
);
```
Use `muted` to silence the video entirely:
```tsx
<Video src={staticFile("video.mp4")} muted />
```
## Speed
Use `playbackRate` to change the playback speed:
```tsx
<Video src={staticFile("video.mp4")} playbackRate={2} /> {/* 2x speed */}
<Video src={staticFile("video.mp4")} playbackRate={0.5} /> {/* Half speed */}
```
Reverse playback is not supported.
## Looping
Use `loop` to loop the video indefinitely:
```tsx
<Video src={staticFile("video.mp4")} loop />
```
Use `loopVolumeCurveBehavior` to control how the frame count behaves when looping:
- `"repeat"`: Frame count resets to 0 each loop (for `volume` callback)
- `"extend"`: Frame count continues incrementing
```tsx
<Video
src={staticFile("video.mp4")}
loop
loopVolumeCurveBehavior="extend"
volume={(f) => interpolate(f, [0, 300], [1, 0])} // Fade out over multiple loops
/>
```
## Pitch
Use `toneFrequency` to adjust the pitch without affecting speed. Values range from 0.01 to 2:
```tsx
<Video
src={staticFile("video.mp4")}
toneFrequency={1.5} // Higher pitch
/>
<Video
src={staticFile("video.mp4")}
toneFrequency={0.8} // Lower pitch
/>
```
Pitch shifting only works during server-side rendering, not in the Remotion Studio preview or in the `<Player />`.
@@ -0,0 +1,99 @@
---
name: voiceover
description: Adding AI-generated voiceover to Remotion compositions using ElevenLabs TTS
metadata:
tags: voiceover, audio, elevenlabs, tts, speech, calculateMetadata, dynamic duration
---
# Adding AI voiceover to a Remotion composition
Use ElevenLabs TTS to generate speech audio per scene, then use [`calculateMetadata`](./calculate-metadata) to dynamically size the composition to match the audio.
## Prerequisites
An **ElevenLabs API key** is required (`ELEVENLABS_API_KEY` environment variable).
**MUST** ask the user for their ElevenLabs API key if `ELEVENLABS_API_KEY` is not set. **MUST NOT** fall back to other TTS tools.
Ensure the environment variable is available when running the generation script:
```bash
node --strip-types generate-voiceover.ts
```
## Generating audio with ElevenLabs
Create a script that reads the config, calls the ElevenLabs API for each scene, and writes MP3 files to the `public/` directory so Remotion can access them via `staticFile()`.
The core API call for a single scene:
```ts title="generate-voiceover.ts"
const response = await fetch(
`https://api.elevenlabs.io/v1/text-to-speech/${voiceId}`,
{
method: "POST",
headers: {
"xi-api-key": process.env.ELEVENLABS_API_KEY!,
"Content-Type": "application/json",
Accept: "audio/mpeg",
},
body: JSON.stringify({
text: "Welcome to the show.",
model_id: "eleven_multilingual_v2",
voice_settings: {
stability: 0.5,
similarity_boost: 0.75,
style: 0.3,
},
}),
},
);
const audioBuffer = Buffer.from(await response.arrayBuffer());
writeFileSync(`public/voiceover/${compositionId}/${scene.id}.mp3`, audioBuffer);
```
## Dynamic composition duration with calculateMetadata
Use [`calculateMetadata`](./calculate-metadata.md) to measure the [audio durations](./get-audio-duration.md) and set the composition length accordingly.
```tsx
import { CalculateMetadataFunction, staticFile } from "remotion";
import { getAudioDuration } from "./get-audio-duration";
const FPS = 30;
const SCENE_AUDIO_FILES = [
"voiceover/my-comp/scene-01-intro.mp3",
"voiceover/my-comp/scene-02-main.mp3",
"voiceover/my-comp/scene-03-outro.mp3",
];
export const calculateMetadata: CalculateMetadataFunction<Props> = async ({
props,
}) => {
const durations = await Promise.all(
SCENE_AUDIO_FILES.map((file) => getAudioDuration(staticFile(file))),
);
const sceneDurations = durations.map((durationInSeconds) => {
return durationInSeconds * FPS;
});
return {
durationInFrames: Math.ceil(sceneDurations.reduce((sum, d) => sum + d, 0)),
};
};
```
The computed `sceneDurations` are passed into the component via a `voiceover` prop so the component knows how long each scene should be.
If the composition uses [`<TransitionSeries>`](./transitions.md), subtract the overlap from total duration: [./transitions.md#calculating-total-composition-duration](./transitions.md#calculating-total-composition-duration)
## Rendering audio in the component
See [audio.md](./audio.md) for more information on how to render audio in the component.
## Delaying audio start
See [audio.md#delaying](./audio.md#delaying) for more information on how to delay the audio start.
@@ -1,5 +0,0 @@
---
"aicodeman": patch
---
fix(terminal): keep the output a pane capture could not contain. Opening a session, a backpressure refresh, a clear-terminal reload and a full-history re-pull all load the screen from a tmux pane capture, and anything the CLI printed between that capture and the end of the load used to be dropped, so its next partial redraw landed on a frame the terminal had never seen: missing or garbled output right after a tab switch or a refresh, plainest in a shell session. Each load now replays exactly the output that arrived after the capture, through one shared rule for all four paths, and a refresh that restores your scroll position no longer snaps back to the bottom afterwards.
-5
View File
@@ -1,5 +0,0 @@
---
"aicodeman": patch
---
fix(input): make sure a prompt sent through the API actually leaves the composer. Claude Code 2.1.277 started ignoring Enter for the first 30 to 50 seconds after the composer paints while still accepting the typed text, so a prompt sent right after a session came up sat unsent in the pane and every waiter (send-and-wait, the agent skill, cron, the maintainer bot) burned its whole timeout on a turn that never started. The server now reads the pane after every programmatic write that carried Enter and presses Enter again, on a 2 to 60 second schedule, only while the composer verifiably still holds the text it sent; an empty composer, other text, or a pane with no composer at all ends it. The agent skill's `sendwait` gets the same loop for servers that predate this, and its preamble version moves to 1.30.1 so an already-seeded agent picks up the fresh copy.
-30
View File
@@ -1,30 +0,0 @@
{
"name": "codeman",
"owner": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
},
"description": "Codeman, self-hosted mission control for AI coding agents. Ships the codeman agent skill: let one Claude Code session spawn, prompt, wait on and read other sessions.",
"plugins": [
{
"name": "codeman",
"source": "./plugins/codeman",
"description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.30.0",
"author": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
},
"homepage": "https://getcodeman.com",
"category": "productivity",
"keywords": [
"codeman",
"orchestration",
"multi-agent",
"session-manager",
"tmux",
"claude-code"
]
}
]
}
-21
View File
@@ -1,21 +0,0 @@
.git
.agents
.claude
.codex
# `**/` matters: a .dockerignore pattern is matched against the WHOLE
# context-relative path, so a bare `.env` excludes ONLY the root file and
# `COPY . .` would bake docker/.env -- CODEMAN_PASSWORD and any provider API
# keys -- into the published image at /opt/codeman/docker/.env (verified).
**/.env
**/.env.*
!**/.env.example
# Same shape: docker/docker-compose.override.yml is the documented home for
# host-specific settings, so it must not ride COPY . . into the image either.
**/docker-compose.override.*
node_modules
dist
coverage
out
test-results
tmp
*.log
-89
View File
@@ -1,89 +0,0 @@
# Contributing to Codeman
Thanks for wanting to help! Codeman is a small project with a fast loop: issues usually get a response within a day, good PRs get reviewed quickly, and every release credits its contributors and bug reporters by name in the release notes. This guide gets you from clone to merged PR without stepping on the traps.
## The short version
1. **Bugs**: open an issue with your OS, install method (installer / npm / git clone), browser, and which CLI + version the session was running.
2. **Questions and ideas**: use [Discussions](https://github.com/Ark0N/Codeman/discussions), not issues.
3. **Small fixes** (docs, typos, a new skin, a translation): just send the PR.
4. **Anything bigger**: open an issue or Discussion first and get a nod before building. Codeman has strong architectural invariants, and a design chat up front is what turns a big idea into a merged PR instead of a stalled one. This flow works: features like Clone Repo (#236) went idea, then design discussion, then review, then shipped.
5. **Security issues**: never a public issue. See [SECURITY.md](SECURITY.md).
## Dev setup
Requirements: Node.js 22+ (see `.nvmrc`), tmux, and at least one supported agent CLI on your PATH (Claude Code is the primary one).
```bash
git clone https://github.com/Ark0N/Codeman.git
cd Codeman
npm install # postinstall builds the vendored xterm addon bundles
npm run dev # dev server on http://localhost:3000
```
The frontend is plain JS served from `src/web/public/` with no bundler in dev: edit a `.js`/`.css` file and reload the page. The one exception is `index.html`, which is read once at server start, so markup changes need a server restart.
## Before you push
CI runs all of these, so save yourself a round trip:
```bash
npm run typecheck # tsc --noEmit, strict mode
npm run lint
npm run format:check
npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules
```
### Tests
```bash
npm test # the gate — exactly what CI runs
npm test -- test/<file>.test.ts # one file
```
`npm test` is the same suite CI runs, so a green run locally means a green run there. It leaves out three suites that cannot pass on an arbitrary machine, each with its own command:
```bash
npm run test:browser # Playwright + chromium (+ a live server; codex-predictive-echo needs a real codex binary)
npm run test:mobile # the above plus environment-specific PNG baselines
npm run test:perf # wall-clock benchmarks — run on an otherwise idle machine
npm run test:all # literally everything, environmental failures included
```
Expect `test:browser`/`test:mobile`/`test:perf` to fail where the machine cannot provide what they need; read that as "not runnable here", not as a regression. `config/test-suites.ts` holds the globs, and both configs derive from it, so the exclusions and those runners cannot drift apart.
If you add a test that binds a port, pick a unique one at 3150 or above (search the repo for `const PORT =` first). Never 3000.
Tests are tmux-safe by design: under vitest, the tmux layer becomes an in-memory mock, so tests cannot touch real sessions.
## Finding your way around
- Every source file starts with a `@fileoverview` JSDoc block. Read it before diving into the file, it is the map.
- [`CLAUDE.md`](../CLAUDE.md) at the repo root is the densest architecture primer in the repo. It is written for AI coding agents, but the invariants and gotchas in it apply to humans exactly the same, and most review feedback on PRs traces back to something already written there.
- Deep mechanisms and the history behind each rule live in [`docs/architecture-invariants.md`](../docs/architecture-invariants.md).
- Third-party extension surfaces are documented in [`docs/extending-codeman.md`](../docs/extending-codeman.md).
## Great first contributions
These are well-fenced areas where a first PR is genuinely easy to get right:
- **A new theme skin.** A skin is four things kept in sync: the `html[data-skin="…"]` token block in `styles.css`, the xterm ANSI palette in `terminal-ui.js`, the pre-paint allowlist and the settings picker (both in `index.html`). `test/skin-themes.test.ts` statically checks the sync, so if the test passes, your skin works.
- **A new language.** `src/web/public/i18n.js` is dependency-free, English is the canonical source, and `zh-CN` is a complete example to copy. Add your language's entries and register it in `SUPPORTED_LANGUAGES`.
- **Docs.** If you got stuck on something and then figured it out, the sentence that would have unstuck you is a PR.
- Anything labeled [`good first issue`](https://github.com/Ark0N/Codeman/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22).
Bigger extension points worth discussing first: new CLI backends (the pluggable resolver pattern has absorbed six CLIs so far; `docs/extending-codeman.md` and `docs/opencode-integration.md` show the shape), and real-device testing reports, especially mobile, which always find things emulation cannot.
## PR expectations
- **One change per PR.** Small and focused reviews fast; a grab-bag stalls.
- Target the `master` branch.
- **Keep your branch mergeable.** A PR with conflicts silently gets no CI runs at all (GitHub quirk), so rebase or merge master when conflicts appear.
- Include or update tests when you change behavior. Route handlers have a lightweight pattern in `test/routes/` using `app.inject()` (no live server needed).
- Formatting is Prettier with a deliberately narrow scope (`npm run format`), several frontend files are hand-formatted on purpose and excluded via `.prettierignore`. Don't "fix" a file by adding it back into Prettier's scope.
- Don't bump versions or touch `CHANGELOG.md`; releases are handled by the maintainer via changesets after merge.
- AI-assisted contributions are welcome (much of Codeman is built that way), with one condition: you must understand what you're submitting and have actually run it. "The model said it works" is not a test.
## Conduct
Be kind, be direct, assume good faith. Report unacceptable behavior privately via the contact in [SECURITY.md](SECURITY.md).
-86
View File
@@ -1,86 +0,0 @@
# Security Policy
Codeman launches AI coding sessions with `--dangerously-skip-permissions`, so the
web UI is **by design a remote-code-execution surface for whoever can reach it**.
The entire security model exists to control *who* that is. Please read this before
exposing an instance beyond `localhost`. The full model lives in
[`docs/security-architecture.md`](../docs/security-architecture.md).
## Supported versions
Security fixes land on the latest published `codeman@X.Y.Z` release and `master`.
Older versions are not patched — upgrade to the latest release (App Settings →
Updates for git-clone installs, or `npm i -g aicodeman@latest`).
| Version | Supported |
| ------- | --------- |
| latest `0.9.x` / `master` | ✅ |
| anything older | ❌ (upgrade) |
## Reporting a vulnerability
**Please do not open a public issue for security problems.**
Report privately via **GitHub's private vulnerability reporting**:
the repository's **Security** tab → **Report a vulnerability**
(<https://github.com/Ark0N/Codeman/security/advisories/new>). This opens a private
advisory thread with the maintainer.
> Maintainer note: enable *Settings → Code security and analysis → Private
> vulnerability reporting* so this channel is live.
When reporting, please include: affected version/commit, the deployment shape
(loopback-only, `CODEMAN_PASSWORD` set, tunnel/`tailscale serve`, custom
reverse proxy), reproduction steps, and impact. We aim to acknowledge within a
few days. Coordinated disclosure is appreciated — we'll agree a disclosure
timeline with you once impact is confirmed.
### In scope
- Authentication / session-cookie bypass when `CODEMAN_PASSWORD` is set
- DNS-rebinding, CSRF/CSWSH, or Origin/Host-guard bypass reaching state-changing routes
- Remote code execution reachable **without** local OS access (e.g. via a browser, a tunnel, or a foreign origin)
- Path traversal / arbitrary file read or write through the HTTP API
- Supply-chain integrity of the in-app self-updater
### Out of scope (by design — see Known limitations)
- Anything requiring an already-trusted **same-machine, same-uid** process. Codeman trusts the local OS user it runs as; a peer process of that user is already inside the boundary.
- Running an authless instance bound to a non-loopback host after dismissing the startup warning (you explicitly acknowledged it).
- The default loopback + no-password posture itself (it is reachable only from the same machine).
## Trust model (summary)
- **Loopback by default.** Binds `127.0.0.1`; the no-password default is safe out of the box. Binding a non-loopback host without `CODEMAN_PASSWORD` *starts but prints a loud warning* with concrete fixes.
- **Always-on Host + Origin guards.** Block DNS-rebinding and cross-site state-changing requests even on the no-auth loopback install (a missing Origin is allowed so CLI/hooks work).
- **Optional auth.** HTTP Basic via `CODEMAN_USERNAME`/`CODEMAN_PASSWORD`; success issues an opaque server-side 256-bit cookie. Per-IP rate limiting on failures.
- **Hardened file serving, tmux launch, transport headers, and multi-instance isolation** — see the full architecture doc.
## Known limitations and accepted risk
A 1.0 release is an implicit statement that the documented model *is* the model, so
these residuals are stated explicitly. Most sit **inside the same-uid OS trust
boundary** or behind the always-on Origin guard; they matter mainly for
shared-host, multi-user, or tunneled deployments.
- **Self-update trusts an unsigned release tag.** The in-app updater does `git checkout <tag> && npm install` (lifecycle scripts run) of a tag matched only by name shape, from whatever `origin` points to — no signature/commit verification. Treat the updater as trusting your `origin` remote and your release pipeline. (Hardening tracked for 1.0.)
- **CSP ships `'unsafe-inline'`.** Inline handlers mean the Content-Security-Policy is defense-in-depth only; all AI-/file-derived sinks are escaped, but a future missed escape would be executable.
- **`workingDir` is unconstrained.** A session may be created with any absolute working directory (e.g. `/`), which becomes the file-route boundary for that session. Scope it to trusted paths on shared hosts.
- **Hook-event auth exemption is loopback-IP-based.** `POST /api/hook-event` is exempt from auth for loopback callers; because tunnels (cloudflared / `tailscale serve`) terminate at `127.0.0.1`, a loopback-terminating tunnel inherits the exemption. Set `CODEMAN_PASSWORD` and prefer a tunnel that preserves the client identity if this matters.
- **Session cookie is not bound to client IP/UA on reuse, and refreshes without an absolute cap.** A stolen cookie replays until its idle TTL elapses.
- **Multi-instance tmux socket is process-wide.** Two Codeman instances on the same `CODEMAN_INSTANCE` share a tmux socket and can attach each other's live sessions — isolate with distinct `CODEMAN_INSTANCE` values.
- **The live log-tail route reads `/var/log` and `~/logs`** in addition to the session working directory (read-only) — a deliberate choice for tailing system/app logs. On a password-protected remote deployment an authenticated user can therefore read those roots outside their session. See `docs/security-architecture.md` §5.
- **The web-tab proxy fetches from the server's network position.** Any authenticated user can save a dashboard URL on loopback or a private range and have Codeman relay to it; that is the feature. Link-local and cloud-metadata addresses are the only refused targets (see below). On a shared host, restrict who holds an account.
Recent hardening (2026-09-04): the web-tab proxy, its "Test" probe and its
WebSocket relay refuse link-local and cloud-metadata targets (`169.254.0.0/16`,
`fe80::/10`, `fd00:ec2::254`, `168.63.129.16`, `100.100.100.200`,
`metadata.google.internal`), judged on the RESOLVED address so a DNS name pointing
there is refused too; proxy capabilities are revoked on logout, admin logout and
user deletion; proxied responses carry `Referrer-Policy: same-origin`. Earlier:
web-push subscription endpoints are restricted to https public hosts (SSRF guard,
rejects internal/metadata IP literals, validated at subscribe and send time), and
tmux session names discovered on the shared socket are validated against the
safe-name pattern before reaching any shell call site.
For the detailed rationale, defenses, and recommended secure setups, see
[`docs/security-architecture.md`](../docs/security-architecture.md).
+5 -147
View File
@@ -11,10 +11,10 @@ jobs:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- uses: actions/checkout@v4
- name: Setup Node.js
uses: actions/setup-node@v6
uses: actions/setup-node@v4
with:
node-version: 22
cache: 'npm'
@@ -22,157 +22,15 @@ jobs:
- name: Install dependencies
run: npm ci
- name: Check package-lock.json version sync
run: npm run check:lockfile
- name: Type check
run: npm run typecheck
- name: Lint
run: npm run lint
- name: Frontend JS syntax check
run: npm run check:frontend-syntax
- name: Format check
run: npm run format:check
# install.sh reaches users through `curl | bash` with nothing between it and
# them, and until now nothing in this repo checked it at all: no shellcheck,
# no bats, and the vitest gate is Node-only.
- name: install.sh syntax
run: bash -n install.sh
# macOS ships bash 3.2 and this runner has bash 5, so the constructs that
# actually break a Mac install are invisible here without a container. This
# step is what catches them — in particular expanding an EMPTY array under
# `set -u`, which bash 3.2 treats as an unbound variable and `bash -n`
# cannot see because it is a runtime error, not a syntax one.
- name: install.sh runs on bash 3.2 (macOS's version)
run: |
set -euo pipefail
docker run --rm -v "$PWD":/w -w /w bash:3.2 bash -n /w/install.sh
docker run --rm -v "$PWD":/w -w /w -e CODEMAN_INSTALL_SH_LIB=1 bash:3.2 bash -c '
set -euo pipefail
. /w/install.sh
detect_all_clis
# `shell` declares no binaries, so its offset/length window is length 0.
# Iterating it is the empty-array case; reaching here means it did not abort.
echo "bash $BASH_VERSION: ${#CLI_IDS[@]} CLIs, $CLI_FOUND_COUNT found"
cli_catalog_names >/dev/null
cli_catalog_print_install_hints >/dev/null
# The install menu with nothing installed and the user answering "s":
# skipping must warn and continue, never trip the "failed to install"
# gate (it did once, aborting the install before the clone).
has_tty() { return 0; }
headless_guard() { return 0; }
read_reply() { eval "$1=s"; }
NONINTERACTIVE=0
k=0; while [[ $k -lt ${#CLI_ALL_BINS[@]} ]]; do CLI_ALL_BINS[$k]="no-such-cli-$k"; k=$((k + 1)); done
k=0; while [[ $k -lt ${#CLI_ALL_PATHS[@]} ]]; do CLI_ALL_PATHS[$k]="/nonexistent/$k"; k=$((k + 1)); done
CLI_DETECT_DONE=""; detect_all_clis
offer_ai_cli_install >/dev/null 2>&1
echo "bash $BASH_VERSION: skipping the AI CLI install menu continues"
'
# Issue #382: the dsh identity probe builds an OPTIONAL `timeout` prefix as an
# array, and on stock macOS there is no `timeout`, so the array is empty and the
# expansion aborts the whole installer under `set -u`. The step above cannot
# reach that branch: this image HAS `timeout`, and with no `dsh` on PATH the
# probe is never called at all. So hide `timeout` and call it directly.
docker run --rm -v "$PWD":/w -w /w -e CODEMAN_INSTALL_SH_LIB=1 bash:3.2 bash -c '
set -euo pipefail
. /w/install.sh
printf "#!/bin/sh\necho \"DeepSeek Harness 0.1\"\n" > /tmp/dsh
printf "#!/bin/sh\necho \"dancer shell (Debian dsh)\"\n" > /tmp/not-dsh
chmod 755 /tmp/dsh /tmp/not-dsh
# A PATH the probe can still work on, minus the binary under test.
mkdir -p /tmp/nobin
for b in grep sh; do ln -sf "$(command -v $b)" "/tmp/nobin/$b"; done
export PATH=/tmp/nobin
if command -v timeout >/dev/null 2>&1; then
echo "timeout is still on PATH, so this is NOT exercising the empty-array branch" >&2
exit 1
fi
dsh_banner_probe /tmp/dsh
if dsh_banner_probe /tmp/not-dsh; then
echo "identity probe accepted a foreign dsh" >&2
exit 1
fi
echo "bash $BASH_VERSION: dsh identity probe survives a missing timeout"
'
- name: CLI catalogue artifacts are in sync with stock.ts
run: npm run generate:cli-catalog -- --check
- name: Server boot smoke test
run: |
set -u
if ! command -v tmux >/dev/null; then
sudo apt-get update -qq
sudo apt-get install -y tmux
fi
npx tsx src/index.ts web --port 3151 > /tmp/boot.log 2>&1 &
SERVER_PID=$!
trap "kill $SERVER_PID 2>/dev/null || true" EXIT
for i in $(seq 1 30); do
if curl -fsS http://localhost:3151/api/status -o /dev/null; then
echo "Server booted in ${i}s"
exit 0
fi
if ! kill -0 $SERVER_PID 2>/dev/null; then
echo "Server exited before becoming ready. Logs:"
cat /tmp/boot.log
exit 1
fi
sleep 1
done
echo "Server did not respond on /api/status within 30s. Logs:"
cat /tmp/boot.log
exit 1
test:
name: Unit & integration tests
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- name: Setup Node.js
uses: actions/setup-node@v6
with:
node-version: 22
cache: 'npm'
- name: Install dependencies
run: npm ci
- name: Install tmux
run: |
if ! command -v tmux >/dev/null; then
sudo apt-get update -qq
sudo apt-get install -y tmux
fi
- name: Run unit & integration tests
# Excludes the suites that need chromium, per-machine PNG baselines or a
# quiet machine — see config/test-suites.ts for the list and the reason
# behind each entry. Identical to what `npm test` runs locally.
# Safe in CI: TmuxManager no-ops all shell commands under VITEST (test/setup.ts).
run: npm run test:ci
- name: Run xterm-zerolag-input package tests
# Layers 1-3 of the predictive-echo suites (unit laws, fixture replay,
# seeded fuzz): deterministic, no browser, no live server. Depends on
# the ROOT `npm ci` above — workspaces hoist the package's vitest into
# the root node_modules; do not add a separate install here.
run: npx vitest run
working-directory: packages/xterm-zerolag-input
# Note: three suites are excluded from CI, each with its own local runner:
# npm run test:browser Playwright + chromium (+ a live server, and a real
# codex binary for codex-predictive-echo)
# npm run test:mobile the above plus environment-specific PNG baselines
# npm run test:perf wall-clock benchmarks; need an otherwise idle machine
# config/test-suites.ts holds the globs; the configs derive from it so the
# exclusions here and those runners cannot drift apart. Everything else runs in
# the `test` job above, which is the same thing `npm test` runs.
# Note: The test suite is intentionally excluded from CI.
# Tests spawn real tmux sessions and require a full system environment.
# Run tests locally with: npx vitest run test/<file>.test.ts
+4 -38
View File
@@ -11,19 +11,15 @@ jobs:
release:
name: Release
runs-on: ubuntu-latest
permissions:
contents: write
pull-requests: write
steps:
- name: Checkout repo
uses: actions/checkout@v6
uses: actions/checkout@v4
- name: Setup Node.js
uses: actions/setup-node@v6
uses: actions/setup-node@v4
with:
node-version: 22
node-version: 20
cache: npm
registry-url: https://registry.npmjs.org
- name: Install dependencies
run: npm ci
@@ -41,34 +37,4 @@ jobs:
commit: "chore: version packages"
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
NODE_AUTH_TOKEN: ${{ secrets.NPM_TOKEN }}
- name: Rename release tag to codeman
if: steps.changesets.outputs.published == 'true'
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
VERSION=$(node -p "require('./package.json').version")
OLD_TAG="aicodeman@${VERSION}"
NEW_TAG="codeman@${VERSION}"
# Update the GitHub release BEFORE deleting the old tag.
# make_latest pins the "Latest" badge to the Codeman release. This repo
# publishes TWO packages (aicodeman + xterm-zerolag-input), changesets
# creates a GitHub release for each, and GitHub awards "Latest" to
# whichever was published LAST. That is a race: 1.9.2 kept the badge,
# 1.9.4 lost it to xterm-zerolag-input@0.1.7 by two seconds. All package
# releases already exist by the time this step runs, so setting it here
# is deterministic.
RELEASE_ID=$(gh release view "$OLD_TAG" --json databaseId -q .databaseId 2>/dev/null || true)
if [ -n "$RELEASE_ID" ]; then
gh api -X PATCH "repos/${{ github.repository }}/releases/${RELEASE_ID}" \
-f tag_name="$NEW_TAG" \
-f name="$NEW_TAG" \
-f make_latest=true
fi
# Retag
git tag "$NEW_TAG" "$OLD_TAG" 2>/dev/null || true
git tag -d "$OLD_TAG" 2>/dev/null || true
git push origin "$NEW_TAG" ":refs/tags/$OLD_TAG" 2>/dev/null || true
NPM_TOKEN: ${{ secrets.NPM_TOKEN }}
-109
View File
@@ -1,109 +0,0 @@
name: Sync Wiki
# Publishes docs/wiki/ to the repository's GitHub wiki.
#
# The wiki is a separate git repo with no CI and no review, so the source of truth
# lives in docs/wiki/ and this workflow mirrors it. Browser edits to the wiki are
# overwritten by the next sync; fix pages with a PR against docs/wiki/ instead.
#
# One-time setup: GitHub only creates <repo>.wiki.git once the first page has been
# saved in the browser. Save a stub page at /wiki/_new before the first run.
#
# Token: GITHUB_TOKEN can push to the wiki on most repos but not all. If a run fails
# with 403, add a fine-grained PAT with wiki write access as the WIKI_TOKEN secret;
# it is preferred automatically when present. Note the 403 usually surfaces on the
# PUSH, not the clone: this repo is public, so a read-only token still clones the
# wiki fine. Both steps carry the hint.
on:
push:
branches: [master]
paths:
- 'docs/wiki/**'
- '.github/workflows/wiki-sync.yml'
workflow_dispatch:
concurrency: ${{ github.workflow }}
jobs:
sync:
name: Push docs/wiki to the wiki
runs-on: ubuntu-latest
permissions:
contents: write
steps:
- name: Checkout repo
uses: actions/checkout@v6
- name: Clone wiki
env:
WIKI_TOKEN: ${{ secrets.WIKI_TOKEN || secrets.GITHUB_TOKEN }}
run: |
set -euo pipefail
if ! git clone "https://x-access-token:${WIKI_TOKEN}@github.com/${GITHUB_REPOSITORY}.wiki.git" wiki 2>"${RUNNER_TEMP}/clone-err.txt"; then
cat "${RUNNER_TEMP}/clone-err.txt"
echo "::error::Could not clone ${GITHUB_REPOSITORY}.wiki.git. If this says 'Repository not found', the wiki has never had a page: save one at https://github.com/${GITHUB_REPOSITORY}/wiki/_new and re-run. If it says 403, add a WIKI_TOKEN secret."
exit 1
fi
- name: Mirror pages
run: |
set -euo pipefail
# The mirror deletes before it copies, so an empty source would wipe
# every published page and the commit step would happily push that. A
# MISSING directory already fails safely (cp aborts under set -e); an
# empty one does not, so check explicitly. This is the one failure mode
# here that destroys something a browser edit cannot get back.
if [ ! -d docs/wiki ]; then
echo "::error::docs/wiki does not exist. Refusing to mirror, which would delete the entire published wiki."
exit 1
fi
pages=$(find docs/wiki -maxdepth 1 -name '*.md' | wc -l)
if [ "$pages" -eq 0 ]; then
echo "::error::docs/wiki contains no .md pages. Refusing to mirror, which would delete the entire published wiki."
exit 1
fi
echo "Mirroring ${pages} pages."
find wiki -mindepth 1 -maxdepth 1 ! -name '.git' -exec rm -rf {} +
cp -R docs/wiki/. wiki/
- name: Stamp the documented version
run: |
set -euo pipefail
# _Footer.md renders on every page and used to carry a hand-written
# version, which went stale on every release because nothing refreshed
# it. It carries {{VERSION}} instead and the series is stamped here.
series="$(node -p "require('./package.json').version.split('.').slice(0,2).join('.') + '.x'")"
# grep exits 1 when it matches nothing, which under `set -o pipefail`
# would fail the step instead of warning, so test before substituting.
if grep -rlq '{{VERSION}}' wiki/; then
grep -rlZ '{{VERSION}}' wiki/ | xargs -0 -r sed -i "s/{{VERSION}}/${series}/g"
else
echo "::warning::No {{VERSION}} placeholder found in docs/wiki. The published version line can no longer be refreshed automatically."
fi
if grep -rq '{{VERSION}}' wiki/; then
echo "::error::A {{VERSION}} placeholder survived substitution and would be published verbatim."
exit 1
fi
echo "Stamped version ${series}."
- name: Commit and push
run: |
set -euo pipefail
cd wiki
git config user.name 'github-actions[bot]'
git config user.email '41898282+github-actions[bot]@users.noreply.github.com'
git add -A
if git diff --quiet --cached; then
echo "Wiki already up to date."
exit 0
fi
git commit -m "docs: sync wiki from docs/wiki @ ${GITHUB_SHA:0:7}"
if ! git push 2>"${RUNNER_TEMP}/push-err.txt"; then
cat "${RUNNER_TEMP}/push-err.txt"
echo "::error::Could not push to ${GITHUB_REPOSITORY}.wiki.git. A 403 here means the token can read the wiki but not write it, which is the usual GITHUB_TOKEN case: add a fine-grained PAT with wiki write access as the WIKI_TOKEN secret."
exit 1
fi
+3 -60
View File
@@ -1,13 +1,3 @@
# Claude Code local files
.agents/
skills-lock.json
# In-session decision scratchpad (context-survival mechanism, not a deliverable)
DECISIONS.md
# Written by install.sh into end-user clones when setup finishes
.install-complete
# Dependencies
node_modules/
@@ -24,10 +14,6 @@ coverage/
test/e2e/screenshots/current/
test/e2e/screenshots/diffs/
# Mobile visual regression failure artifacts
test/mobile/snapshots/*.actual.png
test/mobile/snapshots/*.diff.png
# Logs
*.log
npm-debug.log*
@@ -48,10 +34,6 @@ Thumbs.db
.env.local
.env.*.local
# Local Compose customisation (host-specific, not part of the project)
docker-compose.override.yml
docker-compose.override.yaml
# State files (local to each machine)
.claude/ralph-loop.local.md
@@ -62,54 +44,15 @@ docker-compose.override.yaml
# Generated output
out/
screenshots-echo-diag/
screenshots-readme/
screenshots-readme-real/
screenshots-real/
scripts/remotion/out/
# Local UI/README capture scratch (screenshot runs, design mockups). Not build
# output, but never meant for git — an unqualified `git add -A` during a COM has
# swept dirs like these into a release before.
design-explorations/
# Artifacts that should not be tracked
test-results/
tmp/
# Machine-local working files (never meant for git). ANCHORED so only the root
# dir matches.
/pr/
# Root `public` (a symlink to scripts/remotion/public — local artifact). ANCHORED
# with a leading slash so it does NOT also match src/web/public (a bare `public`
# would swallow the whole web UI source dir and silently un-stage any new asset
# added there). No trailing slash so it still matches the symlink, not just dirs.
/public
# Opt-in gesture overlay runtime assets: large MediaPipe wasm + model (~27 MB)
# fetched at build/install by scripts/fetch-gesture-assets.mjs, kept out of git.
# (The gesture bundle itself, gesture-codeman.js, IS tracked — built from
# packages/gesture-control source by `npm run build:gesture`.)
src/web/public/gesture/wasm/
src/web/public/gesture/*.task
# Gesture-control workspace package build outputs (source is tracked; the
# Codeman bundle is emitted to src/web/public/gesture/gesture-codeman.js instead).
packages/gesture-control/dist/
packages/gesture-control/dist-codeman/
packages/gesture-control/.vite/
tools/remotion/out/
# Claude Code plan tracking
plan.json
# Unfinished TUI (local development only)
src/tui/
.claude/
media-assets/
commands
todo.md
@fix_plan.md
readme-preview.mjs
# Uploaded images land here under each session working dir (runtime artifact)
.claude-images/
# Local-LLM harness smoke-test config (real IPs/keys) — see the .example.json
# alongside it in scripts/, which IS tracked as the template.
scripts/local-llm-test.config.json
+1 -23
View File
@@ -2,30 +2,8 @@ dist/
coverage/
node_modules/
src/web/public/vendor/
src/web/public/gesture/
src/web/public/app.js
src/web/public/styles.css
src/web/public/mobile.css
src/web/public/index.html
# Hand-formatted public JS modules (never prettier-enforced; the new
# check-public-assets.mjs still validates NUL bytes + JS syntax on these).
src/web/public/constants.js
src/web/public/image-input.js
src/web/public/input-cjk.js
src/web/public/keyboard-accessory.js
src/web/public/notification-manager.js
src/web/public/orchestrator-panel.js
src/web/public/panels-ui.js
src/web/public/ralph-panel.js
src/web/public/ralph-wizard.js
src/web/public/respawn-ui.js
src/web/public/session-ui.js
src/web/public/settings-ui.js
src/web/public/sw.js
src/web/public/terminal-ui.js
src/web/public/voice-input.js
src/web/public/upload.html
scripts/remotion/
# Hand-maintained; Prettier escapes underscores in glob paths and corrupts paragraphs.
CLAUDE.md
tools/
+8
View File
@@ -0,0 +1,8 @@
{
"singleQuote": true,
"semi": true,
"tabWidth": 2,
"printWidth": 120,
"trailingComma": "es5",
"endOfLine": "lf"
}
-17
View File
@@ -1,17 +0,0 @@
# Repository Guidelines
Canonical agent/contributor guidance for this repository lives in [CLAUDE.md](CLAUDE.md) —
project structure, build/test/lint commands, code style, testing safety rules
(`npm test` is the CI gate and is safe to run bare; the three excluded suites
have their own runners), security notes, and
the deployment workflow are all maintained there. Please read it before making
changes, and keep it the single source of truth rather than duplicating
sections here.
Quick pointers:
- Type check: `tsc --noEmit` · Lint: `npm run lint` · Format: `npm run format:check`
- Tests: `npm test` (the CI gate, safe to run bare) or `npm test -- test/<file>.test.ts` for one file
- Route tests use `app.inject()`; new tests needing ports must pick a unique `const PORT =`
- Branch off `master` for all work; Conventional Commit-style messages (`fix(mobile): ...`)
- Never commit secrets or local state from `~/.codeman/`
+1 -3217
View File
File diff suppressed because it is too large Load Diff
+425 -335
View File
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -1,6 +1,6 @@
MIT License
Copyright (c) 2024-2026 Codeman Contributors
Copyright (c) 2024 Claudeman Contributors
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
+100 -886
View File
File diff suppressed because it is too large Load Diff
-1135
View File
File diff suppressed because it is too large Load Diff
-281
View File
@@ -1,281 +0,0 @@
[
{
"id": "claude",
"label": "Claude",
"shortBadge": "CC",
"enabled": true,
"order": 0,
"kind": "agent",
"discovery": {
"binaries": [
"claude"
],
"searchDirs": [
"~/.local/bin",
"~/.claude/local",
"/usr/local/bin",
"~/.npm-global/bin",
"~/bin"
],
"install": {
"command": {
"linux": "curl -fsSL https://claude.ai/install.sh | bash",
"darwin": "curl -fsSL https://claude.ai/install.sh | bash",
"wsl": "curl -fsSL https://claude.ai/install.sh | bash"
},
"npmPackage": "@anthropic-ai/claude-code",
"docsUrl": "https://docs.claude.com/claude-code"
}
}
},
{
"id": "shell",
"label": "Shell",
"shortBadge": "SH",
"enabled": true,
"order": 1,
"kind": "shell",
"discovery": {
"binaries": [],
"searchDirs": [],
"install": {
"command": {}
}
}
},
{
"id": "opencode",
"label": "OpenCode",
"shortBadge": "OC",
"enabled": true,
"order": 10,
"kind": "agent",
"discovery": {
"binaries": [
"opencode"
],
"searchDirs": [
"~/.opencode/bin",
"~/.local/bin",
"/usr/local/bin",
"~/go/bin",
"~/.bun/bin",
"~/.npm-global/bin",
"~/bin"
],
"install": {
"command": {
"linux": "curl -fsSL https://opencode.ai/install | bash",
"darwin": "curl -fsSL https://opencode.ai/install | bash"
},
"npmPackage": "opencode-ai",
"docsUrl": "https://opencode.ai/docs"
}
}
},
{
"id": "codex",
"label": "Codex",
"shortBadge": "CX",
"enabled": true,
"order": 20,
"kind": "agent",
"discovery": {
"binaries": [
"codex"
],
"searchDirs": [
"~/.codex/bin",
"~/.local/bin",
"/usr/local/bin",
"~/.bun/bin",
"~/.npm-global/bin",
"~/bin"
],
"install": {
"command": {
"linux": "npm install -g @openai/codex",
"darwin": "npm install -g @openai/codex"
},
"npmPackage": "@openai/codex",
"docsUrl": "https://developers.openai.com/codex/cli"
}
}
},
{
"id": "gemini",
"label": "Gemini",
"shortBadge": "GM",
"enabled": true,
"order": 30,
"kind": "agent",
"discovery": {
"binaries": [
"gemini"
],
"searchDirs": [
"~/.gemini/bin",
"~/.local/bin",
"/usr/local/bin",
"~/.bun/bin",
"~/.npm-global/bin",
"~/bin"
],
"install": {
"command": {
"linux": "npm install -g @google/gemini-cli",
"darwin": "npm install -g @google/gemini-cli"
},
"npmPackage": "@google/gemini-cli",
"docsUrl": "https://github.com/google-gemini/gemini-cli"
}
}
},
{
"id": "antigravity",
"label": "Antigravity",
"shortBadge": "AG",
"enabled": true,
"order": 40,
"kind": "agent",
"discovery": {
"binaries": [
"agy"
],
"searchDirs": [
"~/.local/bin",
"~/.antigravity/bin",
"/usr/local/bin",
"~/bin"
],
"install": {
"command": {
"linux": "curl -fsSL https://antigravity.google/cli/install.sh | bash",
"darwin": "curl -fsSL https://antigravity.google/cli/install.sh | bash"
},
"docsUrl": "https://antigravity.google/cli"
}
}
},
{
"id": "pi",
"label": "Pi",
"shortBadge": "PI",
"enabled": true,
"order": 50,
"kind": "agent",
"discovery": {
"binaries": [
"pi"
],
"searchDirs": [
"~/.local/bin",
"/usr/local/bin",
"~/.bun/bin",
"~/.npm-global/bin",
"~/bin"
],
"install": {
"command": {
"linux": "npm install -g --ignore-scripts @earendil-works/pi-coding-agent",
"darwin": "npm install -g --ignore-scripts @earendil-works/pi-coding-agent"
},
"npmPackage": "@earendil-works/pi-coding-agent",
"docsUrl": "https://pi.dev",
"agentImageLayer": {
"kind": "dedicated",
"reason": "installed with --ignore-scripts in its own layer, so the flag cannot leak to the shared block"
}
}
}
},
{
"id": "grok",
"label": "Grok",
"shortBadge": "GK",
"enabled": true,
"order": 70,
"kind": "agent",
"discovery": {
"binaries": [
"grok"
],
"searchDirs": [
"~/.grok/bin",
"~/.local/bin",
"/usr/local/bin",
"~/bin"
],
"install": {
"command": {
"linux": "curl -fsSL https://x.ai/cli/install.sh | bash",
"darwin": "curl -fsSL https://x.ai/cli/install.sh | bash"
},
"docsUrl": "https://github.com/xai-org/grok-build"
}
}
},
{
"id": "deepseek",
"label": "DeepSeek",
"shortBadge": "DS",
"enabled": true,
"order": 80,
"kind": "agent",
"discovery": {
"binaries": [
"dsh"
],
"searchDirs": [
"~/.local/bin",
"/usr/local/bin",
"~/.npm-global/bin",
"~/bin"
],
"identity": {
"arg": "--help",
"regex": "DeepSeek\\s+Harness"
},
"install": {
"command": {
"linux": "npm install -g @deepseek-ai/dsh",
"darwin": "npm install -g @deepseek-ai/dsh"
},
"npmPackage": "@deepseek-ai/dsh",
"docsUrl": "https://github.com/deepseek-ai/deepseek-harness",
"agentImageLayer": {
"kind": "dedicated",
"reason": "needs pnpm alongside it (dsh plugin, issue #352) and a dsh-tui profile install"
}
}
}
},
{
"id": "omp",
"label": "OMP",
"shortBadge": "OM",
"enabled": true,
"order": 90,
"kind": "agent",
"discovery": {
"binaries": [
"omp"
],
"searchDirs": [
"~/.local/bin",
"~/.omp/bin",
"/usr/local/bin",
"~/.bun/bin",
"~/.npm-global/bin",
"~/bin"
],
"install": {
"command": {
"linux": "curl -fsSL https://omp.sh/install | sh",
"darwin": "brew install can1357/tap/omp"
},
"docsUrl": "https://omp.sh"
}
}
}
]
-16
View File
@@ -1,16 +0,0 @@
{
"$schema": "https://unpkg.com/knip@5/schema.json",
"entry": [
"scripts/*.mjs",
"scripts/*.js",
"scripts/watch-subagents.ts",
"scripts/remotion/Root.tsx",
"scripts/remotion/index.ts",
"test/**/*.test.ts",
"test/mobile/vitest.config.ts",
"test/**/*.mjs"
],
"project": ["src/**/*.{ts,tsx}", "scripts/**/*.{ts,tsx,mjs,js}", "test/**/*.{ts,mjs}"],
"ignoreExportsUsedInFile": true,
"ignoreDependencies": ["@remotion/cli", "@remotion/transitions", "esbuild", "agent-browser"]
}
-50
View File
@@ -1,50 +0,0 @@
/**
* The test suites that `npm test` deliberately does NOT run, in one place.
*
* Why this file exists: the exclusion list used to live only in
* config/vitest.ci.config.ts, as literals. Anything excluded there was
* therefore reachable only by running the everything-config by hand and reading
* past its failures — and a newly excluded file was reachable by nothing at
* all, silently, because nothing pointed at it. Both configs now derive their
* globs from the arrays below, so adding a suite here puts it in exactly one
* runner and takes it out of exactly one gate.
*
* Adding a new test that cannot run in CI: put its glob in the array that
* describes WHY it cannot, not in whichever one is shortest.
*/
/**
* Playwright-driven: needs chromium and, in most cases, a live Codeman server
* on a real port. Deterministic where the environment provides both, which is
* why these are a runnable suite (`npm run test:browser`) rather than skipped.
*/
export const BROWSER_TEST_GLOBS = [
'test/tab-rail-resize.browser.test.ts',
'test/session-sidebar-ux.browser.test.ts',
'test/session-options-responsive.browser.test.ts',
'test/inline-rename.test.ts',
'test/opencode-resize.test.ts',
'test/webgl-fallback.test.ts',
'test/terminal-copy-shortcut.test.ts',
'test/terminal-keycode229-recovery.browser.test.ts',
'test/capture-load-window.browser.test.ts',
'test/codex-predictive-echo.test.ts', // also needs a real codex binary
];
/**
* Wall-clock benchmarks. They assert on durations, so a loaded shared runner
* fails them for reasons that have nothing to do with the diff under test.
*/
export const PERF_TEST_GLOBS = ['test/perf-*.test.ts'];
/**
* Browser + visual regression: chromium AND environment-specific PNG baselines
* that are generated per machine. Has its own config
* (test/mobile/vitest.config.ts) because it needs serial execution, a longer
* timeout and the `pretest:mobile` vendor step — run it with
* `npm run test:mobile`, not through the configs here.
*/
export const MOBILE_TEST_GLOBS = ['test/mobile/**'];
/** Everything `npm test` skips. */
export const NON_CI_TEST_GLOBS = [...MOBILE_TEST_GLOBS, ...PERF_TEST_GLOBS, ...BROWSER_TEST_GLOBS];
-11
View File
@@ -1,11 +0,0 @@
{
"extends": "../tsconfig.json",
"compilerOptions": {
"rootDir": "..",
"noEmit": true,
"declaration": false,
"declarationMap": false,
"sourceMap": false
},
"include": ["../scripts/test-local-llm-harnesses.ts"]
}
-34
View File
@@ -1,34 +0,0 @@
import { resolve } from 'node:path';
import { defineConfig } from 'vitest/config';
import { BROWSER_TEST_GLOBS } from './test-suites';
const root = resolve(import.meta.dirname, '..');
/**
* The Playwright-driven suite `npm test` skips — `npm run test:browser`.
*
* Needs chromium and, for most of these, a live Codeman server on a real port;
* codex-predictive-echo also needs a real codex binary. Expect failures where
* the machine cannot provide those, and read them as "not runnable here", not
* as a regression.
*
* The mobile suite is NOT here: it needs per-machine PNG baselines, serial
* execution and the `pretest:mobile` vendor step, so it keeps its own config
* (test/mobile/vitest.config.ts) behind `npm run test:mobile`.
*
* fileParallelism stays off for the same reason as every other config in this
* directory: these bind real ports and drive real tmux sessions, and two files
* doing that at once fail each other rather than the code.
*/
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: BROWSER_TEST_GLOBS,
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
testTimeout: 60000,
teardownTimeout: 60000,
},
});
-30
View File
@@ -1,30 +0,0 @@
import { resolve } from 'node:path';
import { defineConfig, configDefaults } from 'vitest/config';
import { NON_CI_TEST_GLOBS } from './test-suites';
const root = resolve(import.meta.dirname, '..');
/**
* The default gate — what `npm test` and CI both run.
*
* Same as vitest.config.ts but EXCLUDES the suites that cannot pass on an
* arbitrary machine: browser-driven (Playwright + chromium), visual-regression
* (per-machine PNG baselines) and wall-clock perf. Those are not unmaintained;
* they have their own runners (`test:browser`, `test:mobile`, `test:perf`).
* See config/test-suites.ts for the list and the reason behind each entry.
*
* Keep the rest in sync with config/vitest.config.ts.
*/
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: ['test/**/*.test.ts'],
exclude: [...configDefaults.exclude, ...NON_CI_TEST_GLOBS],
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
testTimeout: 30000,
teardownTimeout: 60000,
},
});
-37
View File
@@ -1,37 +0,0 @@
import { resolve } from 'node:path';
import { defineConfig } from 'vitest/config';
const root = resolve(import.meta.dirname, '..');
/**
* EVERY test in the repo, including the ones that cannot pass on an arbitrary
* machine — `npm run test:all`. Reach for it when you want the complete picture
* and are prepared to read past environmental failures.
*
* This is NOT what `npm test` runs. On a machine without chromium, a free port
* or per-machine PNG baselines this config fails ~87 tests on a clean master,
* which makes it useless as a pass/fail signal: the default gate is
* config/vitest.ci.config.ts, and the suites it leaves out each have their own
* runner (`test:browser`, `test:perf`, `test:mobile`). See config/test-suites.ts.
*/
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: ['test/**/*.test.ts'],
setupFiles: ['./test/setup.ts'],
// Run test files sequentially to respect mux session limits
// Individual tests within files still run in parallel where safe
fileParallelism: false,
coverage: {
provider: 'v8',
reporter: ['text', 'json', 'html'],
include: ['src/**/*.ts'],
exclude: ['src/index.ts', 'src/cli.ts'],
},
testTimeout: 30000, // 30 seconds for integration tests
// Ensure cleanup runs even on test failures
teardownTimeout: 60000,
},
});
-25
View File
@@ -1,25 +0,0 @@
import { resolve } from 'node:path';
import { defineConfig } from 'vitest/config';
import { PERF_TEST_GLOBS } from './test-suites';
const root = resolve(import.meta.dirname, '..');
/**
* The wall-clock benchmarks `npm test` skips — `npm run test:perf`.
*
* These assert on durations, so run them on an otherwise idle machine: a loaded
* runner fails them for reasons that have nothing to do with the diff under
* test, which is exactly why they are not part of the default gate.
*/
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: PERF_TEST_GLOBS,
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
testTimeout: 60000,
teardownTimeout: 60000,
},
});
-82
View File
@@ -1,82 +0,0 @@
# =============================================================================
# Codeman Docker Compose environment template
# Copy this file to .env and set the values for the Docker host.
# =============================================================================
TZ=Australia/Perth
# Optional overrides for direct `docker compose` use. The Bash start script
# detects these values from CODEMAN_APPDATA_PATH automatically. Compose uses
# 1000:1000 when the variables are omitted.
# PUID=1000
# PGID=1000
# Name of the account that runs Codeman and all local CLI sessions. Changing
# this value rebuilds the image with a matching account.
CODEMAN_RUNTIME_USER=codeman
# Required. Persistent Codeman application data, CLI credentials, and session
# state are stored here on the host and mounted at the runtime account's home
# directory in the container.
CODEMAN_APPDATA_PATH=/mnt/user/appdata/codeman
# Optional. Absolute host path of this Codeman checkout, mounted at
# /opt/codeman so App Settings -> Updates can update Codeman in place. The Bash
# start script detects it from the compose file's own location, so it only needs
# setting for direct `docker compose` use or a checkout kept elsewhere. Point it
# at a directory that is not a git checkout and in-app updates are unavailable.
# CODEMAN_REPO_PATH=/mnt/user/appdata/codeman/app
# Required for Docker cases. This must be an absolute path on the Docker host.
# Codeman and each isolated case use this same path, so it cannot be a
# container-only path such as /home/codeman/codeman-cases.
CODEMAN_CASES_PATH=/mnt/user/appdata/codeman/codeman-cases
# Required. Network bind address, host port, and local image tag.
CODEMAN_HOST=0.0.0.0
CODEMAN_PORT=3000
CODEMAN_IMAGE=codeman:local
# Required for any network-accessible Codeman instance. Use a unique, strong
# password. This file is safe to commit; copy it to .env and set the value.
CODEMAN_PASSWORD=changeme
# Required. Username for Codeman HTTP Basic authentication.
CODEMAN_USERNAME=admin
# Optional. Extra Host-header allowlist entries for a reverse-proxied domain
# (comma-separated; a bare `.suffix` matches every subdomain). Without it a
# proxied request is rejected with `403 Forbidden: host not allowed`. See
# README.md, "Reverse-proxy host allowlist".
# CODEMAN_ALLOWED_HOSTS=codeman.example.com,.internal.example.com
# Optional: authenticate Gemini CLI without an interactive login.
GEMINI_API_KEY=
# Linux default. On Docker Desktop, use the socket path supported by your
# Docker installation when it differs from /var/run/docker.sock.
DOCKER_SOCKET=/var/run/docker.sock
# Optional override for direct `docker compose` use. The Bash start script
# detects this from DOCKER_SOCKET automatically. The direct Compose default is
# 999, but the correct value depends on the Docker host.
# DOCKER_SOCKET_GID=999
# Set to 1 only when Docker-case hook callbacks are required.
CODEMAN_DOCKER_BRIDGE_HOOKS=0
# Set to 1 when `docker info` reports `SwapLimit=false`. The case memory limit
# remains active; Codeman omits --memory-swap and filters the daemon's exact
# unsupported-swap warning while preserving all other Docker create errors.
CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=0
# Required only when applying the macvlan example in README.md.
CODEMAN_MACVLAN_NETWORK=br0.11
CODEMAN_IPV4_ADDRESS=10.10.11.236
CODEMAN_MAC_ADDRESS=02:10:11:00:00:EC
# Required only when creating a new managed macvlan network, rather than using
# the external-network macvlan example.
CODEMAN_MACVLAN_PARENT=br0.11
CODEMAN_MACVLAN_SUBNET=10.10.11.0/24
CODEMAN_MACVLAN_GATEWAY=10.10.11.1
-148
View File
@@ -1,148 +0,0 @@
# Codeman Docker deployment
This folder contains the Compose configuration, server image Dockerfile, and environment template for a locally built Codeman server.
## Start
From the repository root, create the runtime environment file and set the required values, especially `CODEMAN_PASSWORD`.
```sh
cp docker/.env.example docker/.env
bash docker/Start-Codeman.sh
```
On PowerShell, use the following commands instead. Running Compose from inside `docker/` with no `-f` lets it discover `docker-compose.override.yml` on its own (see [Local customisation](#local-customisation)); naming the file with `-f docker/docker-compose.yaml` from the repository root silently drops the override unless it is named too.
```powershell
Copy-Item docker/.env.example docker/.env
Set-Location docker
docker compose --env-file .env up --build -d
```
Every required value is defined and explained in `.env.example`. `GEMINI_API_KEY` is intentionally optional and may remain blank.
The container starts as root so `entrypoint.sh` can correct the ownership of a bind source the Docker daemon created (it creates a missing one as `root:root`), then drops to `PUID:PGID` with `setpriv` before the server starts, so Codeman itself never runs privileged. That drop needs `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against the file's `cap_drop: ALL`; a compose file written elsewhere (Unraid's Compose Manager, a hand-written unit) must carry the same additions, and the entrypoint names them when they are missing. A directory owned by neither root nor `PUID:PGID` is never re-owned: it is probed for writability as the runtime account and refused with a clear message if that fails. Setting `user:` in Compose skips the whole step.
On Linux, `Start-Codeman.sh` stops with an error when required paths are missing. It creates the application-data directory when safe, detects its numeric owner as `PUID:PGID`, and detects `DOCKER_SOCKET_GID` from the configured Docker socket. It rejects a root-owned application-data directory because Codeman and its local CLI sessions must remain unprivileged.
Codeman, Claude, OpenCode, and other local sessions run as the unprivileged account named by `CODEMAN_RUNTIME_USER`, which defaults to `codeman`. When Compose is run directly, `PUID` and `PGID` default to `1000:1000`; set them in `.env` when the application-data directory has a different owner. The Bash start script determines them automatically instead.
To retain Docker-case support without root when running Compose directly, set `DOCKER_SOCKET_GID` to the numeric group ID of the host socket. On a standard Linux Docker host, obtain it with `stat -c '%g' /var/run/docker.sock`. The Bash start script detects it automatically.
## Updating
Use **App Settings → Updates** in the web UI. The checkout Compose builds from is
also mounted at `/opt/codeman`, so an update's `git checkout` and rebuild persist
on the host, and the server exiting is what restarts the container onto the new
build.
Releases that change `server.Dockerfile`, `docker-compose.yaml`, or add a key to
`.env.example` cannot be applied that way — the updater detects them, names what
changed, and asks you to run `Start-Codeman.sh` here on the host instead. Details:
[`../docs/docker-self-update.md`](../docs/docker-self-update.md).
## Local customisation
Compose merges `docker-compose.override.yml` on top of `docker-compose.yaml`. Keep host-specific changes there rather than editing `docker-compose.yaml`, so this repository can be updated without losing them. Both `docker-compose.override.yml` and `docker-compose.override.yaml` are ignored by Git.
`Start-Codeman.sh` names the Compose file explicitly, which disables Compose's automatic discovery of the override file, so the script adds it back when one is present and prints the file it used. Running `docker compose` from this folder without any `-f` option finds it automatically. When passing `-f docker/docker-compose.yaml` from the repository root, add `-f docker/docker-compose.override.yml` as well, or the override is silently ignored.
An override file adds to and replaces individual settings. It cannot delete a key from `docker-compose.yaml`, and Compose concatenates rather than replaces `ports`, so removing a published port still requires editing `docker-compose.yaml`. The example below replaces the restart policy and adds a mount, leaving every other setting in place:
```yaml
services:
codeman:
restart: always
volumes:
- /srv/projects:/srv/projects
```
### Reverse-proxy host allowlist
Codeman rejects any request whose `Host` header is not on its own allowlist - a
DNS-rebinding guard, not a Compose or Docker concern. Loopback, any IP literal,
the configured `--host`, and a few tunnel-provider suffixes are allowed by
default; a reverse-proxied domain is not, and is rejected with
`403 Forbidden: host not allowed` before the request reaches any handler.
Add the domain with `CODEMAN_ALLOWED_HOSTS` in `.env`:
```sh
CODEMAN_ALLOWED_HOSTS='codeman.example.com,.internal.example.com'
```
`docker-compose.yaml` forwards it into the container (Compose only passes
through the environment keys it explicitly lists, and this is one of them, with
an empty default so the line is optional in `.env`).
See the application's own `docs/wiki/Remote-Access.md` for the full allowlist
format and the tunnel providers it accepts by default.
## Application data storage
The default configuration uses a host-folder bind mount:
```yaml
volumes:
- type: bind
source: ${CODEMAN_APPDATA_PATH}
target: /home/${CODEMAN_RUNTIME_USER}
```
Set `CODEMAN_APPDATA_PATH` in `.env` to a directory that the Docker daemon can access. The example value is `/mnt/user/appdata/codeman`.
`CODEMAN_CASES_PATH` is the separate host directory for managed case workspaces. It is mounted into Codeman at the same absolute path, allowing the host Docker daemon to bind it into an isolated case container. Set it to a child directory of `CODEMAN_APPDATA_PATH` unless you deliberately store workspaces elsewhere.
Compose also exposes `CODEMAN_APPDATA_PATH` to Codeman as `CODEMAN_DOCKER_HOST_HOME`. This lets Docker case seed files, CLI credentials and the hook secret be mounted using paths that exist in the host daemon's filesystem. Direct host installations do not set this variable and retain their existing behaviour.
Set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` when `docker info` reports `SwapLimit=false`. Codeman continues to apply the configured case memory limit, omits Docker's unsupported `--memory-swap` option, and filters only the daemon's exact swap-capability warning. Every other Docker create error and its exit status remain visible.
For an existing installation created by a root-running image, change ownership of the application-data directory before upgrading so the configured `PUID` and `PGID` can read the saved credentials and state:
```sh
chown -R 99:100 /mnt/user/appdata/codeman
```
Replace `99:100` and the path with the values from your `.env` file.
Do not replace this bind mount with a Docker-managed named volume when Docker cases are enabled. Codeman passes seed, credential, transcript and hook-secret bind sources to the host Docker daemon, so their source files must have stable paths in the daemon's filesystem. A named volume does not provide the required host path mapping.
## Static macvlan networking
The default configuration publishes a host port. It does not use `network_mode: host`. To attach Codeman directly to an existing external macvlan network with a static IP address and MAC address, remove the `ports:` section from `docker-compose.yaml` and add the following to the `codeman` service. The service and network additions can instead be placed in `docker-compose.override.yml`, but the `ports:` removal cannot, as described under [Local customisation](#local-customisation):
```yaml
mac_address: ${CODEMAN_MAC_ADDRESS}
networks:
codeman_lan:
ipv4_address: ${CODEMAN_IPV4_ADDRESS}
```
Then add this top-level network declaration:
```yaml
networks:
codeman_lan:
external: true
name: ${CODEMAN_MACVLAN_NETWORK}
```
Set `CODEMAN_MACVLAN_NETWORK`, `CODEMAN_IPV4_ADDRESS`, and `CODEMAN_MAC_ADDRESS` in `.env`. The values in `.env.example` match the supplied Unraid example network and should be changed for other hosts.
### Create a managed macvlan network
If an external macvlan network does not already exist, use this top-level declaration instead. Do not use it together with the external-network declaration.
```yaml
networks:
codeman_lan:
driver: macvlan
driver_opts:
parent: ${CODEMAN_MACVLAN_PARENT}
ipam:
config:
- subnet: ${CODEMAN_MACVLAN_SUBNET}
gateway: ${CODEMAN_MACVLAN_GATEWAY}
```
Macvlan containers are ordinarily not reachable from their Docker host without additional host-network routing. Confirm the selected address, MAC address, parent interface, and subnet are reserved and valid for the target network before starting the stack.
-309
View File
@@ -1,309 +0,0 @@
#!/usr/bin/env bash
set -euo pipefail
script_dir=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
env_file="$script_dir/.env"
compose_file="$script_dir/docker-compose.yaml"
if [[ ! -f "$env_file" ]]; then
printf 'Error: Docker environment file is missing: %s\n' "$env_file" >&2
printf 'Create it from %s/.env.example before starting Codeman.\n' "$script_dir" >&2
exit 1
fi
# Naming a Compose file explicitly disables Compose's automatic discovery of
# the override file, so it has to be added back by hand. Without this, local
# customisation in docker-compose.override.yml is silently ignored. The
# candidates are checked in Compose's own precedence order - measured on
# Compose v5.5.0 with both present: it uses `.yml` and ignores `.yaml`.
override_yml="$script_dir/docker-compose.override.yml"
override_yaml="$script_dir/docker-compose.override.yaml"
if [[ -f "$override_yml" && -f "$override_yaml" ]]; then
printf 'Warning: both %s and %s exist; Compose uses .yml and ignores .yaml.\n' \
"$override_yml" "$override_yaml" >&2
fi
compose_files=(-f "$compose_file")
for override_file in "$override_yml" "$override_yaml"; do
if [[ -f "$override_file" ]]; then
compose_files+=(-f "$override_file")
printf 'Using Compose override file: %s\n' "$override_file"
break
fi
done
compose_command=(docker compose --env-file "$env_file" "${compose_files[@]}")
appdata_path=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "CODEMAN_APPDATA_PATH" { sub(/^[^=]*=/, ""); print; exit }'
)
cases_path=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "CODEMAN_CASES_PATH" { sub(/^[^=]*=/, ""); print; exit }'
)
docker_socket=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "DOCKER_SOCKET" { sub(/^[^=]*=/, ""); print; exit }'
)
if [[ -z "$appdata_path" ]]; then
printf 'Error: CODEMAN_APPDATA_PATH is not set in %s\n' "$env_file" >&2
exit 1
fi
if [[ ! -d "$appdata_path" ]]; then
if [[ "$EUID" == '0' ]]; then
printf 'Error: Refusing to create CODEMAN_APPDATA_PATH as root: %s\n' "$appdata_path" >&2
printf 'Create it as the unprivileged account that should run Codeman, then retry.\n' >&2
exit 1
fi
mkdir -p -- "$appdata_path"
fi
if [[ -z "$cases_path" ]]; then
printf 'Error: CODEMAN_CASES_PATH is not set in %s\n' "$env_file" >&2
exit 1
fi
# `stat -c` is GNU, `stat -f` is BSD/macOS; the bind sources live on the Docker
# host, so both need to work.
owner_of() {
stat -c '%u:%g' -- "$1" 2>/dev/null || stat -f '%u:%g' "$1" 2>/dev/null
}
if ! owner_ids=$(owner_of "$appdata_path"); then
printf 'Error: Cannot determine the owner of CODEMAN_APPDATA_PATH: %s\n' "$appdata_path" >&2
exit 1
fi
export PUID=${owner_ids%%:*}
export PGID=${owner_ids##*:}
if [[ "$PUID" == '0' ]]; then
printf 'Error: CODEMAN_APPDATA_PATH is owned by root: %s\n' "$appdata_path" >&2
printf 'Change the directory ownership to the unprivileged account that should run Codeman.\n' >&2
exit 1
fi
# Pre-creating this here, exactly like CODEMAN_APPDATA_PATH above, means Compose
# never has to materialise a missing bind source itself - which it does as
# root:root - so the in-container entrypoint's chown never has to run for this
# path at all. It happens AFTER PUID/PGID are known (they come from the appdata
# directory just above) so the new directory can be given that exact owner: a
# plain `mkdir -p` lands as the invoking user's uid and PRIMARY gid, and on a
# host set up the way the README suggests (`chown -R 99:100 <appdata>`) that gid
# is not PGID, which the container would then refuse to run on. Unlike appdata,
# an EXISTING cases directory is left exactly as it is: the README explicitly
# allows pointing this at a normal projects directory the host account already
# owns, and the container checks that it is WRITABLE as PUID:PGID rather than
# who owns it.
if [[ ! -d "$cases_path" ]]; then
mkdir -p -- "$cases_path"
if [[ "$(owner_of "$cases_path")" != "$PUID:$PGID" ]]; then
# As root this always succeeds; as a member of PGID a chgrp does; anyone
# else gets the clear error here, where the fix is obvious, rather than a
# restart loop from the container.
if ! chown -- "$PUID:$PGID" "$cases_path" 2>/dev/null; then
printf 'Error: created CODEMAN_CASES_PATH (%s) but could not make it %s:%s (the owner of CODEMAN_APPDATA_PATH).\n' \
"$cases_path" "$PUID" "$PGID" >&2
printf 'Run `chown %s:%s %s` as root, or create the directory as that account, then retry.\n' \
"$PUID" "$PGID" "$cases_path" >&2
exit 1
fi
fi
fi
if [[ -z "$docker_socket" || ! -S "$docker_socket" ]]; then
printf 'Error: DOCKER_SOCKET is not a Unix socket: %s\n' "${docker_socket:-<unset>}" >&2
exit 1
fi
if socket_ids=$(stat -c '%u:%g' -- "$docker_socket" 2>/dev/null); then
:
elif socket_ids=$(stat -f '%u:%g' "$docker_socket" 2>/dev/null); then
:
else
printf 'Error: Cannot determine the owner of DOCKER_SOCKET: %s\n' "$docker_socket" >&2
exit 1
fi
export DOCKER_SOCKET_GID=${socket_ids##*:}
repo_path=${CODEMAN_REPO_PATH:-$(cd -- "$script_dir/.." && pwd)}
if [[ ! -d "$repo_path" ]]; then
printf 'Error: CODEMAN_REPO_PATH is not a directory: %s\n' "$repo_path" >&2
exit 1
fi
export CODEMAN_REPO_PATH="$repo_path"
# The in-app updater runs `git checkout` and `npm install` against this checkout
# as PUID:PGID. If the directory belongs to someone else, git refuses outright
# ("detected dubious ownership") and the update fails at the first step — so warn
# here, where the fix is obvious, rather than in a failed update hours later.
if repo_owner=$(stat -c '%u' -- "$repo_path" 2>/dev/null || stat -f '%u' "$repo_path" 2>/dev/null); then
if [[ "$repo_owner" != "$PUID" ]]; then
printf 'Warning: %s is owned by UID %s but Codeman runs as UID %s.\n' "$repo_path" "$repo_owner" "$PUID" >&2
printf 'In-app updates will fail until the ownership matches. Codeman itself still starts.\n' >&2
fi
fi
if [[ ! -d "$repo_path/.git" ]]; then
printf 'Note: %s is not a git checkout, so in-app updates are unavailable.\n' "$repo_path" >&2
fi
# Reads HEAD without requiring a `git` binary on the host — this script
# otherwise checks the checkout only by testing for `.git` as a directory, and
# resolving refs by hand keeps that the same "no host git needed" guarantee.
# ⚠️ A worktree checkout has `.git` as a FILE (`gitdir: <path>`), not a
# directory, so this returns nothing there and the volume-refresh check below
# silently no-ops — consistent with the `-d .git` test used everywhere else in
# this script, not a special case, but worth knowing if a worktree checkout
# stops picking up a stale-volume refresh it should have caught.
git_head_commit() {
local git_dir="$1/.git" head_ref ref_path
[[ -d "$git_dir" ]] || return 1
head_ref=$(cat -- "$git_dir/HEAD" 2>/dev/null) || return 1
if [[ "$head_ref" == ref:* ]]; then
ref_path="${head_ref#ref: }"
if [[ -f "$git_dir/$ref_path" ]]; then
cat -- "$git_dir/$ref_path"
else
# Packed after a `git gc`; the loose ref file above is gone.
awk -v ref="$ref_path" '$2 == ref { print $1; exit }' "$git_dir/packed-refs" 2>/dev/null
fi
else
printf '%s' "$head_ref"
fi
}
# Record what the container is about to be built and created FROM. The in-app
# updater compares these against the release it wants to apply: a release that
# changes either file cannot be applied by the container restarting itself (a
# restart reuses the existing image and config), so it is refused and the user
# is sent back here. Written on every start, so the baseline always describes
# the container that is actually running. See docs/docker-self-update.md.
if command -v sha256sum >/dev/null 2>&1; then
sha256_of() { sha256sum -- "$1" | cut -d' ' -f1; }
elif command -v shasum >/dev/null 2>&1; then
sha256_of() { shasum -a 256 -- "$1" | cut -d' ' -f1; }
else
sha256_of() { printf ''; }
fi
dockerfile_sha=$(sha256_of "$script_dir/server.Dockerfile")
compose_sha=$(sha256_of "$compose_file")
if [[ -n "$dockerfile_sha" && -n "$compose_sha" ]]; then
# $CODEMAN_APPDATA_PATH is mounted at the runtime account's home, so this is
# dataPath('docker-env-applied.json') as the server inside the container sees it.
state_dir="$appdata_path/.codeman"
mkdir -p -- "$state_dir"
printf '{\n "dockerfileSha256": "%s",\n "composeSha256": "%s"\n}\n' \
"$dockerfile_sha" "$compose_sha" >"$state_dir/docker-env-applied.json.tmp"
mv -- "$state_dir/docker-env-applied.json.tmp" "$state_dir/docker-env-applied.json"
# A root-run start (common on Unraid) would otherwise leave a root-owned
# `.codeman` on a FIRST start, before the container has created it as PUID,
# and the unprivileged server could then never write its own state there.
if [[ "$EUID" == '0' ]]; then
chown -- "$PUID:$PGID" "$state_dir" "$state_dir/docker-env-applied.json"
fi
else
printf 'Warning: no sha256 tool found; in-app updates will not detect environment changes.\n' >&2
fi
# codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded from
# the image only while EMPTY, so a rebuilt image's fresh output sits unused
# behind old volume content until something clears it. The in-app self-updater
# never hits this — it rebuilds INSIDE the running container, into the very
# volume already in use — but a `docker compose build` triggered from outside
# it (this script, after a `git pull`) does: the container comes back up
# looking unchanged. Detect that here and clear just the affected volume(s) so
# the build below actually takes effect. Best-effort: with no sha256 tool this
# quietly does nothing, same as the environment-gate block above.
volumes_to_refresh=()
if [[ -n "$dockerfile_sha" ]]; then
repo_head=$(git_head_commit "$repo_path" || true)
lockfile_sha=$(sha256_of "$repo_path/package-lock.json" 2>/dev/null || true)
source_state_file="$state_dir/docker-build-source.json"
prev_head=''
prev_lockfile_sha=''
if [[ -f "$source_state_file" ]]; then
prev_head=$(sed -n 's/.*"headCommit": *"\([^"]*\)".*/\1/p' "$source_state_file")
prev_lockfile_sha=$(sed -n 's/.*"lockfileSha256": *"\([^"]*\)".*/\1/p' "$source_state_file")
fi
[[ -n "$repo_head" && "$repo_head" != "$prev_head" ]] && volumes_to_refresh+=('codeman-dist')
[[ -n "$lockfile_sha" && "$lockfile_sha" != "$prev_lockfile_sha" ]] && volumes_to_refresh+=('codeman-node-modules')
fi
if [[ ${#volumes_to_refresh[@]} -eq 0 ]]; then
exec "${compose_command[@]}" up --build -d
fi
# Runs even on this script's very first invocation against an EXISTING
# deployment, deliberately: that deployment's volumes may already be stale
# (there was no earlier version of this check to have caught it), and clearing
# an already-empty or nonexistent volume is a harmless no-op, so there is no
# fresh-install case this needs to avoid.
printf 'Source changed since the last start; refreshing: %s\n' "${volumes_to_refresh[*]}"
# Build BEFORE taking the stack down: the image build is the slow part and needs
# no container stopped, so the deployment is offline only for the recreate.
"${compose_command[@]}" build
# `com.docker.compose.volume` is the volume KEY, not a project-qualified name -
# a second stack on the same host (a beta instance started with a different
# COMPOSE_PROJECT_NAME, say) that also declares a volume keyed `codeman-dist`
# shares that label, and `head -n1` would pick whichever the daemon happens to
# list first. Scope the lookup to THIS stack's own resolved project name so it
# can only ever match this stack's volume. The name is read from the resolved
# config's top-level `name` key, indentation-agnostic (the formatting is not a
# contract), and the FIRST `name` in the output is the project's: nested ones
# (a network's `name:`) come later. `--format json` needs Compose v2.3+.
project_name=$(
"${compose_command[@]}" config --format json 2>/dev/null |
sed -n 's/^[[:space:]]*"name":[[:space:]]*"\([^"]*\)".*$/\1/p' | head -n1
)
"${compose_command[@]}" down
# Track whether the volumes were actually cleared. The marker below is written
# ONLY on success: with an unresolvable project name the label filter would
# match nothing, nothing would be removed, and a marker recording the new HEAD
# would stop this check from ever firing again while the stale volume kept
# serving old code. A failed removal likewise leaves the marker alone, so the
# next start retries, and the stack is brought back up regardless rather than
# left down.
refreshed=1
if [[ -z "$project_name" ]]; then
# The documented reset (docs/docker-self-update.md): both volumes re-seed from
# the image by a plain copy, so clearing the extra one costs a copy, not data.
printf 'Warning: could not resolve the Compose project name; clearing both build-artefact volumes with `down --volumes` instead.\n' >&2
"${compose_command[@]}" down --volumes || refreshed=0
else
for key in "${volumes_to_refresh[@]}"; do
volume_name=$(
docker volume ls -q \
--filter "label=com.docker.compose.volume=$key" \
--filter "label=com.docker.compose.project=$project_name" |
head -n1
)
if [[ -n "$volume_name" ]] && ! docker volume rm -- "$volume_name"; then
printf 'Warning: could not remove volume %s; it will be retried on the next start.\n' "$volume_name" >&2
refreshed=0
fi
done
fi
if [[ "$refreshed" == '1' ]]; then
printf '{\n "headCommit": "%s",\n "lockfileSha256": "%s"\n}\n' \
"$repo_head" "$lockfile_sha" >"$source_state_file.tmp"
mv -- "$source_state_file.tmp" "$source_state_file"
if [[ "$EUID" == '0' ]]; then
chown -- "$PUID:$PGID" "$source_state_file"
fi
else
printf 'Warning: the build-artefact volumes were NOT refreshed; the container may serve stale code until the next successful start.\n' >&2
fi
# Already built above, so no --build here: a second build would only re-check
# the cache.
exec "${compose_command[@]}" up -d
-175
View File
@@ -1,175 +0,0 @@
# Codeman agent base image (built locally by scripts/build-agent-image.mjs).
#
# Contains the agent toolchain (node + the CLIs + git/tmux/ripgrep) but NO
# secrets: credentials are delivered at RUNTIME via bind mounts (~/.claude etc.)
# or name-only `docker exec --env`, never baked in, so `docker save` exports stay
# secret-free. tmux is a HARD prerequisite (the in-container tmux is what makes a
# reconnect durable), so it is installed here and probed before launch.
#
# HOME is made writable by an ARBITRARY host uid via the OpenShift "gid 0,
# group-writable" convention: on Linux we run `--user <hostUid>:0`, so the agent
# uid is the host uid (workspace files stay host-owned) while gid 0 keeps $HOME
# writable even though the uid is not the baked 1000.
FROM node:22-bookworm-slim
# Base toolchain. `curl` is needed for the hook callbacks (`curl -sk $CODEMAN_API_URL`),
# `procps` for `ps`, `tmux` for the durable in-container session.
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
git \
tmux \
ripgrep \
curl \
ca-certificates \
less \
procps \
openssh-client \
&& rm -rf /var/lib/apt/lists/*
# The npm-published agent CLIs, supplied by scripts/build-agent-image.mjs from
# config/clis.stock.json so a new stock CLI needs no edit here. The default is
# today's literal list, so a bare `docker build` still produces the same image.
#
# ⚠️ Expanded UNQUOTED on purpose: word splitting is what turns the list into
# several arguments. Every token is validated against
# ^[@A-Za-z0-9][@A-Za-z0-9/._-]*$ on the producing side
# (scripts/lib/cli-catalog.mjs) precisely because of that.
#
# ⚠️ Filtered on each entry's `enabled` flag, so a CLI that ships disabled is
# never baked into every image.
#
# Pinning is left to the rebuild cadence (see docs/docker-cases-plan.md,
# user-decision 2).
# ⚠️ The default is in REGISTRY order, byte-identical to what the generator emits.
# A different order is a different RUN string, which is a different layer hash and
# so a needless cache miss between a bare `docker build` and a scripted one.
ARG CLI_NPM_PACKAGES="@anthropic-ai/claude-code opencode-ai @openai/codex @google/gemini-cli"
RUN npm install -g ${CLI_NPM_PACKAGES} \
&& npm cache clean --force
# Antigravity (`agy`) is NOT on npm — Google ships a standalone binary through its
# own installer, so it needs its own step. `--dir /usr/local/bin` is load-bearing:
# the installer's default target is `$HOME/.local/bin`, which at build time is
# root's home and would be unreachable by the `agent` user the container runs as.
# ⚠️ This binary is ~190MB on its own; it is the single largest layer in the image.
RUN curl -fsSL https://antigravity.google/cli/install.sh | bash -s -- --dir /usr/local/bin \
&& chmod 755 /usr/local/bin/agy \
&& agy --version
# Pi (pi.dev). Upstream documents --ignore-scripts (pi needs no lifecycle scripts);
# kept out of the shared npm block above so the flag cannot silently change how the
# rest of that block's CLIs install — a fixed count would go stale here since
# CLI_NPM_PACKAGES (above) is now a generated, dynamic list rather than a hand-kept one.
RUN npm install -g --ignore-scripts @earendil-works/pi-coding-agent \
&& npm cache clean --force \
&& pi --version
# Grok Build (`grok`, xAI) is NOT on npm: a standalone ~160MB Rust binary through
# xAI's installer, which targets $HOME/.grok/bin with no --dir override. At build
# time that is root's home and unreachable by the `agent` user, so copy the binary
# into /usr/local/bin and drop root's ~/.grok in the same layer so the image does
# not carry the download twice. The staging cp -T is what makes this survive the
# installer's own behavior EITHER way: newer installers already symlink
# /usr/local/bin/grok -> /root/.grok/bin/grok, and a direct `cp -L` onto that
# symlink fails with "same file" (2026-08-24 rebuild), while removing the link
# first and copying fresh works for both old and new installers.
RUN curl -fsSL https://x.ai/cli/install.sh | bash \
&& cp -L /root/.grok/bin/grok /usr/local/bin/grok.real \
&& rm -f /usr/local/bin/grok \
&& mv /usr/local/bin/grok.real /usr/local/bin/grok \
&& chmod 755 /usr/local/bin/grok \
&& rm -rf /root/.grok /root/.local/bin/grok /root/.local/bin/agent \
&& grok --version
# DeepSeek Harness (`dsh`). A normal npm package, but the ONLY entry here whose
# binary runs nothing on its own: `dsh` is a profile launcher, and DeepSeek ships
# only `web` and `headless`, so without an interactive profile a
# `mode: 'deepseek'` container would start a pane that dies on arrival. The
# profile itself is installed further down, into the `agent` HOME, because
# Codeman deliberately does NOT seed `profiles/` from the host: it is a
# per-profile node_modules tree, host-arch-specific and far too large to copy on
# every container start.
# ⚠️ `pnpm` is a HARD dependency of `dsh plugin`, not optional tooling: the
# subcommand is a thin forwarder that `spawnSync`s a literal `pnpm` with no
# fallback to npm, so on an image without it the profile install below dies
# with `dsh: pnpm not found on PATH` / exit 127 and takes the whole build with
# it (issue #352). It stays on PATH at runtime too, so a container user can run
# `dsh plugin add` themselves.
RUN npm install -g @deepseek-ai/dsh pnpm \
&& npm cache clean --force \
&& dsh --version \
&& pnpm --version
# OMP (Oh My Pi) is NOT on npm: a standalone binary via omp.sh's installer, which
# targets $HOME/.local/bin with no --dir override (verified 2026-08-27 — the
# resolver's OMP_SEARCH_DIRS lists ~/.omp/bin first, which turned out to be the
# WRONG guess for the installer's actual target; build this step for real
# rather than trust that ordering). At build time $HOME is root's home and
# unreachable by the `agent` user, so copy the binary into /usr/local/bin and
# drop root's ~/.local/bin/omp in the same layer so the image does not carry
# the download twice.
RUN curl -fsSL https://omp.sh/install | sh \
&& cp -L /root/.local/bin/omp /usr/local/bin/omp.real \
&& rm -f /usr/local/bin/omp \
&& mv /usr/local/bin/omp.real /usr/local/bin/omp \
&& chmod 755 /usr/local/bin/omp \
&& rm -f /root/.local/bin/omp \
&& omp --version
# `agent` user (gid 0) with an arbitrary-uid-writable HOME. The uid is
# auto-assigned (node:22-slim already occupies uid 1000 with its `node` user); at
# runtime Codeman overrides with `--user <hostUid>:0` on Linux, so the baked uid
# only matters for a hand-run / Docker Desktop container. gid 0 + group-writable
# HOME (OpenShift arbitrary-uid convention) keeps $HOME writable for any uid.
# UTF-8 locale so tmux/Ink render Unicode box-drawing instead of VT100 ACS `q`
# glyphs (C.UTF-8 is built into glibc; no locales package needed). Codeman also
# sets these at run time so containers built before this line still get UTF-8.
ENV LANG=C.UTF-8 LC_ALL=C.UTF-8
ENV HOME=/home/agent
# `.claude` (+ `.claude/projects` mount point) and `.codex` (+ `.codex/sessions`) are
# pre-created gid-0 group-writable so the container owns its OWN credential config
# dirs: tokens/settings/config are seeded in as writable copies and each CLI's runtime
# state (backups, tasks, refreshed tokens) stays container-local, while ONLY the shared
# transcript/rollout dirs (`.claude/projects`, `.codex/sessions`) are bind-mounted from
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir;
# Antigravity nests its state inside `.gemini/antigravity-cli`, so it rides that seed.)
# `.pi/agent` and `.grok` ARE pre-created: both are seeded per-FILE (pi:
# auth/settings/trust/models; grok: auth.json/config.toml/pager.toml), and a
# per-file seed copy, unlike a whole-dir one, does not create its parent directory.
# `.dsh` is pre-created for the same per-file reason (.env/settings.yaml/
# cordis.patch.yml), and the interactive profile is built into it HERE rather than
# after `USER agent`: this layer's closing chgrp/chmod is what makes the whole tree
# writable by the arbitrary uid the container actually runs as, and a profile
# installed after it would miss that fixup. DSH_HOME points the launcher at the
# agent's dir while this still runs as root.
# ⚠️ `dangerouslyAllowAllBuilds` is what keeps that profile install from becoming
# the next #352. pnpm (unlike npm) blocks dependency lifecycle scripts by default
# and FAILS the install over it — `ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on
# pnpm 11.24 — so any package in the tui's tree that ships one stops the build
# dead. An allowlist of the offenders rots: `@deepseek-harness-tui/dsh-tui` is
# resolved by dist-tag, not pinned, and 0.9.3 pulled `@google/genai` (a
# `preinstall: no-op`) where 0.10.0-beta.x does not, so the names to allow move
# under us between rebuilds. Allowing them wholesale is also the SAME exposure
# this image already accepts three layers up: `npm install -g` runs the install
# scripts of every transitive dep of the five CLIs above it, with no gate at all.
# `.omp/agent` is pre-created for the same reason `.codex` is: it is a MIXED
# store (per-file config seeds PLUS a shared `sessions/` RW bind mount for
# Codeman's own host-side history/resume reads), and neither kind of artifact
# creates its own parent directory.
RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
&& mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \
/home/agent/.claude/projects /home/agent/.codex/sessions /home/agent/.pi/agent /home/agent/.grok \
/home/agent/.dsh /home/agent/.omp/agent \
&& DSH_HOME=/home/agent/.dsh HOME=/home/agent \
dsh plugin --profile dsh-tui add --config.dangerouslyAllowAllBuilds=true \
@deepseek-harness-tui/dsh-tui \
&& test -f /home/agent/.dsh/profiles/dsh-tui/package.json \
&& chgrp -R 0 /home/agent \
&& chmod -R g=u /home/agent
USER agent
WORKDIR /home/agent
# Codeman overrides the command with `sleep infinity` at create time; this is the
# fallback so a hand-run container also idles rather than exiting.
CMD ["sleep", "infinity"]
-132
View File
@@ -1,132 +0,0 @@
name: codeman
services:
codeman:
build:
context: ..
dockerfile: docker/server.Dockerfile
args:
CODEMAN_RUNTIME_USER: ${CODEMAN_RUNTIME_USER}
PGID: ${PGID:-1000}
PUID: ${PUID:-1000}
image: ${CODEMAN_IMAGE}
init: true
restart: unless-stopped
ports:
- "${CODEMAN_PORT}:${CODEMAN_PORT}"
environment:
# Tells the self-updater to restart by exiting (the restart policy below
# relaunches it) rather than by looking for an init system that is not
# here. Also set in the image; repeated so a container started without the
# image default still self-identifies.
CODEMAN_IN_CONTAINER: "1"
# This file sets `restart: unless-stopped` below, so the updater may restart
# the server by EXITING. Declared here and only here, never in the image: a
# container started by plain `docker run` has no restart policy unless the
# operator gave it one, and there the updater asks the daemon instead and
# stages the update for a manual restart when it cannot get an answer.
CODEMAN_RESTART_BY_EXIT: "1"
CODEMAN_DOCKER_BRIDGE_HOOKS: ${CODEMAN_DOCKER_BRIDGE_HOOKS}
# Host-side equivalent of the runtime user's HOME. Docker case seed,
# credential and hook mounts are translated into the daemon namespace.
CODEMAN_DOCKER_HOST_HOME: ${CODEMAN_APPDATA_PATH}
CODEMAN_DOCKER_DISABLE_SWAP_LIMIT: ${CODEMAN_DOCKER_DISABLE_SWAP_LIMIT}
CODEMAN_CASES_PATH: ${CODEMAN_CASES_PATH}
# Extra Host-header allowlist entries for a reverse-proxied deployment
# (docker/README.md, "Reverse-proxy host allowlist"). Optional, so it
# defaults to empty rather than requiring a line in every .env.
CODEMAN_ALLOWED_HOSTS: ${CODEMAN_ALLOWED_HOSTS:-}
CODEMAN_HOST: ${CODEMAN_HOST}
CODEMAN_PASSWORD: ${CODEMAN_PASSWORD}
CODEMAN_PORT: ${CODEMAN_PORT}
CODEMAN_USERNAME: ${CODEMAN_USERNAME}
GEMINI_API_KEY: ${GEMINI_API_KEY}
PGID: ${PGID:-1000}
PUID: ${PUID:-1000}
TZ: ${TZ}
group_add:
# Retain access to the host Docker socket without running as root.
- ${DOCKER_SOCKET_GID:-999}
volumes:
# Application data and CLI credentials persist on the configured host
# path, rather than in a Docker-managed volume.
- type: bind
source: ${CODEMAN_APPDATA_PATH}
target: /home/${CODEMAN_RUNTIME_USER}
# Docker cases are sibling containers on the host daemon. Their workspace
# must be visible to Codeman at the same absolute path used by that daemon.
- type: bind
source: ${CODEMAN_CASES_PATH}
target: ${CODEMAN_CASES_PATH}
# Codeman uses the host daemon to create isolated Docker cases. This is
# Docker-outside-of-Docker, not Docker-in-Docker.
- type: bind
source: ${DOCKER_SOCKET}
target: /var/run/docker.sock
# The application source, so App Settings -> Updates can update in place.
# This is the SAME checkout used as the build context above, mounted over
# the image's baked copy: a `git checkout` performed inside the container
# then lands on the host and survives the container being recreated.
# Without it the pull would go to the container's writable layer and be
# silently discarded by the next `up`. See docs/docker-self-update.md.
# Defaults to `..` — the build context above — which Compose resolves
# against the project directory, so plain `docker compose up` works with
# no extra configuration. Set CODEMAN_REPO_PATH only to point elsewhere.
- type: bind
source: ${CODEMAN_REPO_PATH:-..}
target: /opt/codeman
# Build artefacts live in named volumes layered OVER the repo bind mount,
# so `npm install` and `npm run build` inside the container never write
# into the host checkout. That keeps container-compiled native modules
# (node-pty is built from source here) out of a checkout that may also be
# used to run Codeman natively, and keeps `git status` clean. Docker seeds
# an EMPTY named volume from the image, so the first start inherits the
# image's already-built node_modules and dist rather than paying for a
# bootstrap build.
- type: volume
source: codeman-node-modules
target: /opt/codeman/node_modules
- type: volume
source: codeman-dist
target: /opt/codeman/dist
extra_hosts:
- "host.docker.internal:host-gateway"
security_opt:
- no-new-privileges:true
cap_drop:
- ALL
cap_add:
# The entrypoint corrects bind-mount ownership as root before dropping to
# PUID:PGID. Everything not listed here remains dropped by cap_drop above.
# test/docker-entrypoint.test.ts pins this list against what the
# entrypoint and `init: true` actually need, so a capability cannot go
# missing silently again.
- CHOWN
- DAC_OVERRIDE
# `init: true` makes tini PID 1, and tini stays ROOT while the entrypoint
# drops the server to PUID. Signalling a process of a different uid needs
# CAP_KILL; without it tini's SIGTERM forward fails ("Unexpected error
# when forwarding signal: 'Operation not permitted'"), tini dies, and the
# PID namespace teardown SIGKILLs the server instead of letting
# `server.stop()` flush state on every `docker compose down`/`restart`.
- KILL
- SETGID
- SETUID
healthcheck:
test:
- CMD-SHELL
- >-
node -e "fetch('http://127.0.0.1:${CODEMAN_PORT}/api/status').then((response) => process.exit(response.status < 500 ? 0 : 1)).catch(() => process.exit(1))"
interval: 30s
timeout: 5s
retries: 3
start_period: 30s
volumes:
# Container-owned build artefacts. They persist across container recreation,
# so an in-app update's `npm install` output is not thrown away by the next
# `up`, and they are seeded from the image on first use. Removing them (or
# `docker compose down -v`) is the supported reset: the next start rebuilds
# from the image.
codeman-node-modules:
codeman-dist:
-165
View File
@@ -1,165 +0,0 @@
#!/bin/sh
# Corrects ownership - host bind mounts, and the image-baked CLI prefix -
# then drops to PUID:PGID.
#
# Compose binds CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH from the host. When
# either path does not exist yet - a first run, a cleared application-data
# directory, a restored backup - the Docker daemon creates it owned by root,
# and an unprivileged server cannot then create its own state directory. The
# result is a container that restarts forever on:
#
# Failed to start web server: EACCES: permission denied, mkdir '/home/<user>/.codeman'
#
# Running this as root and dropping afterwards removes that failure mode without
# leaving the server privileged. The same root start also lets it re-assert
# /opt/codeman-cli's ownership on every start, not just at image build time -
# see the comment at that chown below for why that matters for anyone who
# runs the compose file directly rather than through Start-Codeman.sh.
#
# Capabilities this script needs against the compose file's `cap_drop: ALL`
# (test/docker-entrypoint.test.ts pins the list against docker-compose.yaml):
# CHOWN + DAC_OVERRIDE the chown of a root-owned bind source below
# SETUID + SETGID the setpriv drop itself
# KILL NOT used here, but required by the container: with
# `init: true` tini is PID 1 and runs as root while the
# server runs as PUID, and signalling a process of a
# different uid needs CAP_KILL. Without it every
# `docker compose down`/`restart` ends in tini dying with
# "Unexpected error when forwarding signal" and the
# server being SIGKILLed instead of stopping cleanly.
set -eu
# Honour an explicit `user:` in Compose: when the container was not started as
# root there is nothing to correct and no privilege to drop.
if [ "$(id -u)" -ne 0 ]; then
exec "$@"
fi
# Everything below runs as root and calls stat, chown, id, setpriv and friends
# by bare name, so the lookup path must not contain a directory the runtime
# account can write to. /opt/codeman-cli/bin is exactly that (it is chowned to
# PUID:PGID so sessions can update the agent CLIs in place), and the image
# appends it to PATH for the server's sake. Resolve root's commands through the
# system directories only, and hand the image's full PATH back to the server at
# the exec below, since Codeman resolves the agent CLIs through it.
runtime_path=$PATH
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
export PATH
: "${PUID:=1000}"
: "${PGID:=1000}"
# The capabilities the compose file must grant, named in the diagnosis below so
# an out-of-tree compose file (Unraid's Compose Manager, a hand-written unit)
# fails with a one-line fix instead of a restart loop.
required_caps='CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID'
# Pre-flight the drop itself before touching anything. A container started with
# `cap_drop: ALL` and none of the additions above fails here, and would otherwise
# die at the final exec with a bare "setpriv: setresuid failed: Operation not
# permitted" after chown had already failed, or worse, misreport a perfectly
# writable directory as unwritable because the probe below could not drop
# privileges to test it.
if ! setpriv --reuid "$PUID" --regid "$PGID" --clear-groups true 2>/dev/null; then
printf 'entrypoint: cannot drop privileges to PUID:PGID (%s:%s).\n' "$PUID" "$PGID" >&2
printf 'entrypoint: this image starts as root and drops with setpriv, which needs\n' >&2
printf 'entrypoint: cap_add: [%s]\n' "$required_caps" >&2
printf 'entrypoint: on top of cap_drop: ALL (see docker/docker-compose.yaml). Add them to the\n' >&2
printf 'entrypoint: compose file that started this container, or set `user:` to skip the drop entirely.\n' >&2
exit 1
fi
# Preserve the supplementary groups Compose granted through group_add - that is
# how the Docker socket stays reachable - while discarding root's own group.
supplementary=$(id -G | tr ' ' '\n' | grep -vx 0 | paste -sd, -)
[ -n "$supplementary" ] || supplementary="$PGID"
# Writable as the account the server is about to become? A real probe, run as
# exactly the identity the final exec below produces (PUID, PGID, the same
# supplementary groups, capabilities dropped), rather than a comparison of
# owners: ownership is not writability. A group-writable tree owned by another
# account, an ACL, or a CIFS/NFS mount that reports some unrelated uid are all
# fine to run on and would all fail an owner check.
writable_as_runtime() {
setpriv --reuid "$PUID" --regid "$PGID" --groups "$supplementary" test -w "$1" 2>/dev/null
}
for target in "${HOME:-}" "${CODEMAN_CASES_PATH:-}"; do
[ -n "$target" ] && [ -d "$target" ] || continue
owner=$(stat -c '%u:%g' "$target")
[ "$owner" = "${PUID}:${PGID}" ] && continue
# Only ever correct a directory the DAEMON created: root-owned, because
# neither PUID nor PGID existed yet when it materialised the missing bind
# source. Anything else - a host tree that legitimately belongs to some
# OTHER account, such as an existing CODEMAN_CASES_PATH the README already
# allows pointing at a normal project directory - is not this container's
# to reassign; recursively chowning it on every mismatch silently rewrote
# a credentials tree or a projects directory to PUID:PGID with one log
# line to explain it. Such a directory is left alone and only PROBED below.
#
# The chown is deliberately not fatal. A bind mount backed by NFS, CIFS or a
# rootless daemon can refuse chown while still being perfectly writable, and
# the probe below is what decides whether the server can run on it.
if [ "${owner%%:*}" = '0' ]; then
if chown -R "${PUID}:${PGID}" "$target" 2>/dev/null; then
printf 'entrypoint: corrected ownership of %s to %s:%s\n' "$target" "$PUID" "$PGID"
else
printf 'entrypoint: warning: cannot change ownership of %s to %s:%s; checking whether it is writable anyway\n' \
"$target" "$PUID" "$PGID" >&2
fi
fi
if writable_as_runtime "$target"; then
if [ "${owner%%:*}" != '0' ]; then
printf 'entrypoint: %s is owned by %s, not %s:%s, but is writable as the runtime account; leaving its ownership alone\n' \
"$target" "$owner" "$PUID" "$PGID"
fi
continue
fi
printf 'entrypoint: %s is not writable as PUID:PGID (%s:%s); it is owned by %s.\n' \
"$target" "$PUID" "$PGID" "$owner" >&2
printf 'entrypoint: refusing to change ownership of a directory this container did not create.\n' >&2
printf 'entrypoint: either chown it on the host, make it writable to %s:%s, or set PUID/PGID to match its owner.\n' \
"$PUID" "$PGID" >&2
exit 1
done
# /opt/codeman-cli (the four agent CLIs) is chowned to PUID:PGID once, at
# image BUILD time, from the PUID/PGID build args - server.Dockerfile's own
# comment on that RUN step explains why it lives in its own prefix rather than
# /usr/local. Unlike HOME/CODEMAN_CASES_PATH above, that bake happens only
# when the image is actually rebuilt (`docker compose up --build`, which
# Start-Codeman.sh always does) - a deployment that instead runs the compose
# file directly (Unraid's Compose Manager, a native Debian systemd unit, any
# `docker compose up`/`restart` with no --build) can change PUID/PGID in .env
# and restart without ever rebuilding, at which point the container runs as
# the NEW uid while the CLI directory is still owned by the OLD one baked into
# the image layer - silently breaking the very "self-update a CLI in place"
# fix this directory exists for. Re-assert it here, every start, unconditionally:
# unlike the host bind mounts above, this is pure image content Codeman itself
# populated, never host data that might legitimately belong to someone else,
# so there is no ownership to be careful about - it is always correct for it
# to be owned by whoever this container is about to run as.
if [ -d /opt/codeman-cli ] && [ "$(stat -c '%u:%g' /opt/codeman-cli)" != "${PUID}:${PGID}" ]; then
chown -R "${PUID}:${PGID}" /opt/codeman-cli
fi
# Discarding group 0 is right for root's own group, but it also discards a
# `group_add: 0` that was there to reach a Docker socket owned by root:root.
# The previous image ran as PUID with that group kept, so say so rather than
# letting Docker-case support vanish silently on such a host.
if [ -S /var/run/docker.sock ] && [ "$(stat -c '%g' /var/run/docker.sock)" = '0' ]; then
printf 'entrypoint: warning: /var/run/docker.sock is owned by group 0, which is dropped along with root;\n' >&2
printf 'entrypoint: warning: Docker cases will not work from this container. Give the socket a dedicated\n' >&2
printf 'entrypoint: warning: group on the host and set DOCKER_SOCKET_GID to it.\n' >&2
fi
# No `--bounding-set -all` here: it is a silent no-op without CAP_SETPCAP, which
# the compose file deliberately does not grant, and `no-new-privileges` already
# makes the bounding set moot. The reuid/regid drop leaves CapPrm/CapEff empty.
# The image's full PATH goes back to the server here; see the top of the file.
exec setpriv --reuid "$PUID" --regid "$PGID" --groups "$supplementary" \
env PATH="$runtime_path" "$@"
-186
View File
@@ -1,186 +0,0 @@
# syntax=docker/dockerfile:1
# Build the application from the checkout supplied as the Docker build context.
# No published Codeman application image is required.
FROM node:22-bookworm-slim AS build
RUN apt-get update \
&& apt-get install -y --no-install-recommends python3 make g++ \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /opt/codeman
COPY . .
# devDependencies are deliberately KEPT (no `npm prune --omit=dev`). The in-app
# updater rebuilds from inside this container, and `npm run build` is tsc +
# esbuild — both devDependencies. Pruning them saves image size and takes the
# self-updater with it. See docs/docker-self-update.md.
RUN npm ci \
&& npm run build \
&& npm cache clean --force
# The Docker CLI talks to the host daemon through the socket mounted by
# docker/docker-compose.yaml. It does not run a Docker daemon in this container.
FROM node:22-bookworm-slim
ARG CODEMAN_RUNTIME_USER=codeman
ARG PUID=1000
ARG PGID=1000
# python3/make/g++ are here for the SELF-UPDATER, not for this build. An update
# runs `npm install` inside the running container, and node-pty ships no Linux
# prebuild, so a release that bumps it compiles from source right here. Without
# a toolchain that install fails and the update rolls back — every time, on the
# releases that need it most. Same reason install.sh installs one on bare hosts.
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
ca-certificates \
curl \
g++ \
git \
make \
openssh-client \
procps \
python3 \
ripgrep \
tmux \
&& rm -rf /var/lib/apt/lists/*
# The Docker CLI, taken from the official image rather than Debian's `docker.io`.
# That package is the full ENGINE: with --no-install-recommends it still pulls 15
# packages including containerd, runc, dmsetup and iptables, none of which a
# client that only talks to a mounted socket can use. Measured on top of this
# base image: `docker.io` costs 266 MB and ships Docker 20.10.24 (2023), while
# these two files cost 108 MB and ship the current CLI (493 MB vs 335 MB total).
#
# The binaries are STATIC Go builds, so they run on this glibc image even though
# the image they come from is Alpine (verified: `docker --version`, `docker ps`
# and `docker build` all work here against a mounted host socket).
#
# buildx is copied on purpose. `scripts/build-agent-image.mjs` shells out to
# `docker build` — Codeman auto-builds the agent image on the first Docker case —
# and without the plugin that silently falls back to the CLASSIC builder, which
# Docker has deprecated and will eventually drop. `docker-compose` is NOT copied:
# Codeman never shells out to it.
COPY --from=docker:29-cli /usr/local/bin/docker /usr/local/bin/docker
COPY --from=docker:29-cli \
/usr/local/libexec/docker/cli-plugins/docker-buildx \
/usr/local/libexec/docker/cli-plugins/docker-buildx
# Keep credentials out of the image. Users authenticate these CLIs at runtime
# through Codeman sessions, and the configured host bind mount retains state.
#
# Installed into a DEDICATED prefix, /opt/codeman-cli, not the base image's
# default /usr/local. A session needs write access to wherever these CLIs live
# so it can self-update one in place (observed via Codex's own
# `npm install -g @openai/codex`, which renames the old package directory
# aside before installing the new one — a rename needs write access to the
# PARENT directory, not just the target, so the runtime account needs that
# access at the directory level). Chowning /usr/local/bin and
# /usr/local/lib/node_modules directly to get it would ALSO hand away
# entrypoint.sh (COPY'd to /usr/local/bin below, root-owned, executed as root
# on every container start with CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node
# binary: owning the DIRECTORY is enough to rename it aside and drop a
# replacement, even though the file itself stays root-owned, which would let a
# compromised session arrange for its own script to run as root at the next
# restart — undoing the "the server itself never runs privileged" guarantee
# the entrypoint exists to provide. /opt/codeman-cli holds nothing else to
# escalate through, so owning it is exactly the CLI-update access it needs and
# no more.
#
# ⚠️ PINNED ON PURPOSE. Unpinned, the agent CLI versions a user ends up with are
# a function of WHEN their image was built, not of any commit — so a Codeman
# release that depends on newer CLI behaviour (the trust-dialog handling is
# pinned to Claude Code 2.1.252's layout; wheel forwarding to >= 2.1.187) breaks
# on an older image with no diff anywhere to explain why. In-app updates make
# rebuilds RARER, which makes that drift worse. Pinning turns "this release needs
# a newer CLI" into a Dockerfile change, which the updater's environment gate
# already detects and refuses (docs/docker-self-update.md).
#
# Bump these deliberately, in a release. `--no-cache` is still needed to rebuild
# this layer when only the pins change upstream.
# The prefix is APPENDED to PATH, never prepended: it is chowned to the runtime
# account below, and entrypoint.sh runs as root calling stat/chown/setpriv by
# bare name. A prefix ahead of /usr/bin would let a session drop a `setpriv`
# there and have it run as root at the next container start (measured with a
# minimal image of this exact shape). The four CLIs live only in this prefix,
# so they still resolve; entrypoint.sh additionally pins its own PATH to the
# system directories for the root part of the start.
ENV NPM_CONFIG_PREFIX=/opt/codeman-cli
ENV PATH=$PATH:/opt/codeman-cli/bin
RUN npm install --global \
@anthropic-ai/claude-code@2.1.258 \
@google/gemini-cli@0.58.0 \
@openai/codex@0.152.1 \
opencode-ai@1.18.26 \
&& npm cache clean --force
# Keep the web server and every local Codeman session unprivileged. PUID and
# PGID match the host-owned application-data directory mounted by Compose. The
# requested GID may not exist in the base image, and a host UID such as 1000 may
# already belong to the baked `node` account, so handle both cases explicitly.
#
# The trailing chown hands the CLI prefix (/opt/codeman-cli, populated above)
# to that same account, so a session can self-update one of the CLIs in place.
# /usr/local stays root-owned throughout — see the comment on the npm install
# above for why that boundary matters.
RUN set -eux; \
case "${PUID}" in ''|*[!0-9]*) echo "PUID must be numeric" >&2; exit 1;; esac; \
case "${PGID}" in ''|*[!0-9]*) echo "PGID must be numeric" >&2; exit 1;; esac; \
if [ "${PUID}" -eq 0 ]; then \
echo "PUID must identify an unprivileged account, not root" >&2; \
exit 1; \
fi; \
if ! getent group "${PGID}" >/dev/null; then \
groupadd --gid "${PGID}" codeman-runtime; \
fi; \
existing_user="$(getent passwd "${PUID}" | cut -d: -f1 || true)"; \
if [ -n "${existing_user}" ]; then \
usermod \
--login "${CODEMAN_RUNTIME_USER}" \
--gid "${PGID}" \
--home "/home/${CODEMAN_RUNTIME_USER}" \
--move-home \
--shell /bin/bash \
"${existing_user}"; \
else \
useradd \
--uid "${PUID}" \
--gid "${PGID}" \
--create-home \
--home-dir "/home/${CODEMAN_RUNTIME_USER}" \
--shell /bin/bash \
"${CODEMAN_RUNTIME_USER}"; \
fi; \
chown -R "${PUID}:${PGID}" /opt/codeman-cli
WORKDIR /opt/codeman
COPY --from=build /opt/codeman /opt/codeman
# CODEMAN_IN_CONTAINER tells the self-updater it must restart by exiting rather
# than by asking an init system that is not here (src/web/self-update.ts).
# NODE_ENV stays `production`; the updater passes `npm install --include=dev`
# explicitly, since that value would otherwise omit the build toolchain.
ENV CODEMAN_IN_CONTAINER=1 \
CODEMAN_PORT=3000 \
HOME=/home/${CODEMAN_RUNTIME_USER} \
NODE_ENV=production
# Runtime defaults for the entrypoint, matching the account created above.
ENV PGID=${PGID} PUID=${PUID}
EXPOSE 3000
# The container starts as root so the entrypoint can correct the ownership of
# the host bind mounts, which the daemon creates as root whenever they do not
# already exist. The entrypoint then drops to PUID:PGID with setpriv, so the
# server itself never runs privileged. Setting `user:` in Compose bypasses both
# steps, leaving the caller in full control.
COPY docker/entrypoint.sh /usr/local/bin/entrypoint.sh
RUN chmod 0755 /usr/local/bin/entrypoint.sh
ENTRYPOINT ["/usr/local/bin/entrypoint.sh"]
CMD ["node", "dist/index.js", "web"]
-104
View File
@@ -1,104 +0,0 @@
# SPEEDRUN.md — Fast-execution protocol for Claude
Read this when the goal is **throughput**: get correct, verified work done with
minimum ceremony. This does **not** relax correctness or the safety rules in
`CLAUDE.md` — those still win. It removes _waste_, not _rigor_.
> Precedence: `CLAUDE.md` > explicit user instructions > this file. If anything
> here conflicts with `CLAUDE.md`, `CLAUDE.md` wins.
---
## The mindset
- **Act, don't announce.** No "I'm going to now…" preamble. Do the thing, report
the result.
- **Cheapest proof that the change works.** Pick the smallest check that actually
demonstrates correctness — not the biggest.
- **Batch aggressively.** Independent reads, greps, and edits go in **one**
message with parallel tool calls. Never serialize work that has no dependency.
- **Momentum over perfection.** Land a correct increment, verify it, move on.
Don't gold-plate untouched code.
---
## Loop (repeat until done)
1. **Orient once** — one parallel burst of reads/greps to load the context you
need. Don't re-read files the harness says are already current.
2. **Change** — make the edit(s). Batch independent edits.
3. **Verify cheaply** — the smallest check that proves _this_ change (see below).
4. **Advance** — next item. Only re-verify what you touched.
5. **Stop** at: list empty, a hard blocker, or a decision that's genuinely the
user's to make.
---
## Verification ladder — climb only as high as the change needs
| Change kind | Cheapest sufficient check |
|-------------|---------------------------|
| Types / signatures / imports | `tsc --noEmit` (or `--watch` already running) |
| One module's logic | `npm test -- test/<file>.test.ts` (the **one** relevant file) |
| A named behavior | `npm test -- -t "pattern"` |
| Route/handler | `app.inject()` route test, or one `curl` against the running dev server |
| Frontend render | Playwright load + assert (`waitUntil: 'domcontentloaded'`, wait 3–4s) |
| Broad / pre-merge | `npm run test:ci` (the CI-equivalent sweep) |
**Hard rules (never skip, even in a rush):**
- ⚠️ **Never run bare `npm test`** — it pulls in browser/visual suites that hang
or fail locally. Always pass a file or `-t`, or use `test:ci`.
- ⚠️ **Never COM without verifying the change actually works** first (curl the
endpoint / Playwright the UI). "Compiles" ≠ "works".
- ⚠️ **Session safety** — check `$CODEMAN_MUX`; never `tmux kill-session` /
`pkill claude` in a managed session.
- ⚠️ **Single-line prompts** for any programmatic session input.
---
## Speed tactics that pay off here
- **Parallel exploration**: dispatch `Explore` subagents (or one parallel grep
burst) instead of serial file-by-file reading when scope is uncertain.
- **`tsc --noEmit --watch`** in the background — instant type feedback, no repeat
cold starts.
- **Target one test file** — `fileParallelism: false` means the suite is serial;
running one file is dramatically faster than the sweep.
- **`curl localhost:3000/api/...`** beats spinning up a browser for backend
checks. Reserve Playwright for actual UI rendering.
- **Trust the harness** — if it says a file you just edited is current, don't
re-Read it to "confirm". The Edit already succeeded or it would have errored.
---
## Anti-patterns (these masquerade as speed, but cost time)
- Running the full test suite to check a one-file change.
- Re-reading files you already have in context.
- Narrating a plan you're about to execute anyway.
- Serial tool calls that have no dependency between them.
- Claiming "done / fixed / passing" **before** running the check that proves it.
- Deploying (COM) on green typecheck alone, without exercising the real flow.
---
## Stop-conditions (don't rush past these)
Stop and surface, don't guess, when you hit:
- A **destructive / hard-to-reverse** action (delete, overwrite, force-push).
- An **outward-facing** action (publishing, sending, deploying) not already
authorized.
- A **genuine product decision** the code can't answer.
- A **failing verification you can't explain** — debug it (see
`superpowers:systematic-debugging`), don't paper over it.
---
## Definition of done
A task is done when **all** hold:
- The change is made.
- The cheapest sufficient check **ran** and **passed** — evidence, not assertion.
- No new type errors / lint errors introduced (`tsc --noEmit`, `npm run lint`).
- You state plainly what was done and what proved it. If a step was skipped or a
test failed, say so — don't hedge, don't overclaim.
-759
View File
@@ -1,759 +0,0 @@
# Agent Control Plan: skill packaging + wait primitives
**Status**: steps 1 to 8 DONE and RELEASED. The wait primitives and the skill itself
(steps 1 to 5) shipped in **1.13.0**; the `codeman skill install` CLI, per-case injection
and `agentSkillEnabled` (step 6) shipped in **1.14.1** and were republished with fixes in
**1.14.2**. Steps 1 to 5 were multi-round verified on 2026-08-08, step 6 on 2026-08-09;
see [§7 Build log](#7-build-log-what-actually-happened) for what shipped, what each
verification round found, and the two items that genuinely remain open (§2.4's footgun
guard and the Part 3 deferrals).
**Date**: 2026-08-08
**Scope**: Part 1 (agent skill) and Part 2 (wait primitives) were specified and built.
Parts 3 to 5 are captured so they are not lost, but remain deliberately deferred.
---
## 0. Where this came from: what herdr does
[herdr](https://github.com/herdrdev/herdr) (Rust, Apache-2.0, ~25.8k stars) is a terminal
multiplexer built around AI coding agents. Relevant findings from the research pass:
| Capability | How herdr does it |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Agent state | Four states (`idle`, `working`, `blocked`, `done`) that roll up pane to tab to workspace in a sidebar |
| Detection | Lifecycle hooks where the agent supports them (it names Pi and MastraCode), otherwise TOML manifests matched against a live bottom-buffer snapshot. Bundled manifests plus remote updates from herdr.dev, local overrides win |
| Control API | Newline-delimited JSON over a Unix socket (`~/.config/herdr/sessions/<name>/herdr.sock`), `{"id":"req_1","method":"pane.split","params":{}}`, dot-notation methods, plus long-lived event subscriptions |
| Discoverability | `herdr api schema` prints a machine-readable schema |
| Agent skill | `npx skills add herdrdev/herdr --skill herdr -g`, a SKILL.md wrapping the CLI, guarded by `test "${HERDR_ENV:-}" = 1` so an agent outside a herdr pane refuses to act |
| Persistence | Background server, detach with `ctrl+b q`, snapshot restore of workspaces/tabs/panes/cwd/layout, experimental screen-history replay, agent resume via native session ids, live PTY handoff across server replacement |
| Plugins | `herdr-plugin.toml` manifest, actions, event hooks, plugin panes, link handlers, GitHub-topic marketplace index |
The commands the skill teaches the agent:
| Group | Commands |
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| workspace | `workspace list`, `workspace create` |
| tab | `tab list --workspace <id>`, `tab create` |
| pane | `pane current`, `pane list`, `pane layout`, `pane split --current --direction right --cwd <path> --no-focus`, `pane run <id> "<cmd>"`, `pane wait-output <id> --match/--regex <p> --timeout <ms>`, `pane read <id> --source visible\|recent\|detection` |
| agent | `agent list`, `agent start <name> --kind <type> --pane <id>`, `agent prompt <name> "<text>" --wait --timeout <ms>`, `agent wait <name> --until <state> --timeout <ms>`, `agent send-keys`, `agent get`, `agent read` |
### The honest comparison
herdr and Codeman are not the same product. herdr is a local, keyboard-first multiplexer with
no server, no web UI, and no autonomy layer. Codeman is a server with a browser and mobile UI,
remote and Docker cases, respawn, Ralph, cron, and the orchestrator, none of which herdr has.
What herdr genuinely does better is being **callable by the agent running inside it**. For
Codeman that is a packaging problem plus one missing primitive, not an architecture problem.
---
## 1. Gap analysis
| herdr capability | Codeman equivalent today | Gap |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------- |
| `pane split` + `agent start` | `POST /api/quick-start`, `POST /api/sessions` | none, already there |
| `agent prompt` | `POST /api/sessions/:id/input` with `clientId`+`seq` exactly-once | no `--wait` |
| `pane read` | `GET /api/sessions/:id/output`, `GET /api/sessions/:id/terminal?full=1` | none |
| `agent list` / `agent get` | `GET /api/sessions`, `GET /api/sessions/unified`, `GET /api/status` | none |
| `agent wait --until <state>` | SSE only (`/api/events`) | **missing**, and SSE is impractical from a shell tool |
| `pane wait-output --match` | nothing | **missing** |
| Skill file | README section "Driving Codeman from an Agent" | **not packaged**, an agent will never find it |
| Env guard `HERDR_ENV=1` | `CODEMAN_MUX=1`, `CODEMAN_API_URL`, `CODEMAN_SESSION_ID` already exported at spawn | none, the guard variables exist |
| `blocked` state | hook events (`permission_prompt`, `elicitation_dialog`) plus CSS classes plus the phone overview NEEDS YOU section | not in the wire contract (`SessionStatus = 'idle' \| 'busy' \| 'stopped' \| 'error'`) |
| `api schema` | hand-written `docs/api-reference.md` | no machine-readable schema |
| Detection manifests | hardcoded in `usage-limit-patterns.ts`, `respawn-*-patterns`, `regex-patterns.ts` | patterns are code, not data |
| Plugin runtime | deliberately refused, see `docs/extending-codeman.md` | not a gap, a decision |
| Session handoff on restart | tmux owns the PTYs, so they already survive a Codeman restart | not a gap, solved by architecture |
**Conclusion**: roughly 90% of the capability surface already exists. Parts 1 and 2 below close
the two real gaps.
The table is the 2026-08-08 snapshot that motivated the work, kept as written. The three rows
marked missing are closed since: `GET .../wait` and `GET .../wait-output` shipped in 1.13.0, and
the skill is packaged at `skills/codeman` (npm tarball included). `blocked` as a wire-contract
state, and the machine-readable schema, are still open (Parts 3 and 4).
---
## 2. Part 1: the Codeman agent skill
### 2.1 Goal
An agent running inside a Codeman session can discover and correctly drive Codeman without the
user pasting API docs into the prompt, and without inventing dangerous calls.
### 2.2 Layout and distribution
The `npx skills` CLI (vercel-labs/skills) clones a GitHub repo and looks for
`skills/<name>/SKILL.md`. Claude Code natively discovers `.claude/skills/<name>/SKILL.md` in a
project and `~/.claude/skills/` globally. Both are satisfied with one source of truth plus a
symlink, which is the pattern this repo already uses for `remotion-best-practices`.
```
skills/
codeman/
SKILL.md <- single source of truth
reference/
endpoints.md <- full endpoint tables, loaded on demand
recipes.md <- worked multi-session orchestration examples
.claude/skills/codeman -> ../../skills/codeman (symlink, dogfooding in this repo)
```
Adding a `skills/` directory to the repo root costs one entry in the GitHub listing. CLAUDE.md
keeps the root short on purpose, so this needs a conscious sign-off; the alternative is
`docs/skills/codeman/` with a `--skill` path argument, which breaks the one-liner install.
**Recommendation**: accept `skills/` at the root, because the install one-liner is the whole
point of shipping a skill.
Install paths, in order of how a user gets it:
1. `npx skills add Ark0N/Codeman --skill codeman -g` (global, any agent, matches the herdr flow).
2. `codeman skill install [--global | --case <name>]`, a new CLI subcommand writing the same
file. This is the path for users who installed via npm and never cloned the repo.
3. **Automatic per-case injection**, modeled exactly on `applyStatusLineConfig(casePath, enabled)`
in `hooks-config.ts`: write `<case>/.claude/skills/codeman/SKILL.md` at case creation,
gated on a new setting. Codeman already writes `<case>/.claude/settings.local.json` hooks
through `writeHooksConfig()`, so this is the same mechanism with the same lifecycle.
Setting name: `agentSkillEnabled`. Synced (not per-device), since it changes on-disk case
content rather than display. Default: **ON after the dogfooding phase, OFF in the first
release**. Rationale for starting OFF: Claude Code loads every skill's name and description
into context on every turn, so an always-on skill has a small permanent token cost, and we
should measure that we are buying something with it first.
### 2.3 SKILL.md content
Frontmatter, per the skills convention (`name` + `description` required):
```yaml
---
name: codeman
description: >-
Control Codeman, the session manager this agent is running inside: list sessions,
start worker sessions, send prompts, read terminal output, and wait for other agents
to finish. Only usable when CODEMAN_MUX=1.
---
```
Body sections, in order:
**1. Guard (first thing, non-negotiable).**
```bash
test "${CODEMAN_MUX:-}" = 1 || { echo "not inside a Codeman session"; exit 1; }
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set, refusing to guess}"
SELF="${CODEMAN_SESSION_ID:-}"
```
If `CODEMAN_MUX` is not `1`, the agent must stop and say it is not running inside a
Codeman-managed session. Same shape as herdr's `HERDR_ENV` guard, and the variables are
already exported by `tmux-manager.buildEnvExports()`. No fallback URL when
`CODEMAN_API_URL` is unset: any guess is the wrong scheme on an HTTPS install (prod is
HTTPS with a self-signed cert, hence `curl -sk` throughout), and a server the agent
cannot identify is not one it should be driving.
**2. Rules of the road.** Lifted and tightened from README lines 666 to 745:
- Single-line input only. Multi-line breaks the agent TUI (Ink).
- Always send `clientId` + a monotonic `seq` on `POST .../input` so a retry cannot double-deliver.
- Envelope is `{success, data}`; a few legacy GETs are bare, so read `body.data ?? body`.
- Add `-u admin:"$CODEMAN_PASSWORD"` when a password is set. Prod is HTTPS, so `curl -sk`.
- Prefer `/api/v1/*`, the stable alias.
**3. Safety rules (the section that does not exist anywhere today).**
- Never act on `$CODEMAN_SESSION_ID`. That is you.
- Only `DELETE` sessions **you created in this conversation**, by exact id. Keep the list.
- Never bulk-delete, never loop a `DELETE` over `/api/sessions`. There is no undo.
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. Use the API.
- Creating a session consumes a slot against the 50-session cap. Clean up what you start.
**4. Recipes**, each one a single copy-pasteable curl:
| Task | Call |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| list sessions | `GET /api/v1/sessions` |
| find yourself | match ids by PREFIX of `$CODEMAN_SESSION_ID` (Docker cases truncate it to 8 chars, so an equality check never fires there) |
| start a worker | `POST /api/v1/quick-start {caseName, mode, effort}` |
| send a prompt | `POST /api/v1/sessions/:id/input {input:"…\r", useMux:true, clientId, seq}` (the trailing `\r` is what sends Enter; without it the text sits on the prompt unsubmitted) |
| send prompt and wait | `POST /api/v1/sessions/:id/input {input:"…\r", wait:"stop", waitTimeout:600000}` (Part 2) |
| wait for a worker | `GET /api/v1/sessions/:id/wait?until=stop,blocked&timeout=300000` (Part 2) |
| wait for a marker | `GET /api/v1/sessions/:id/wait-output?match=DONE_<random>&timeout=120000` (Part 2; unique per call, per §3.3's repaint rule) |
| read output | `GET /api/v1/sessions/:id/output` |
| read full scrollback | `GET /api/v1/sessions/:id/terminal?full=1` |
| watch sub-agents | `GET /api/v1/subagents` |
| schedule work | `POST /api/v1/cron/jobs` |
| clean up | `DELETE /api/v1/sessions/:id` |
**5. Pointer to `reference/endpoints.md`** for anything not in the table, so the always-loaded
part of the skill stays small.
### 2.4 An ergonomics guard worth adding server-side
The skill will tell the agent not to act on itself, but a confused agent can still try. Propose:
the skill sends `X-Codeman-Caller-Session: $CODEMAN_SESSION_ID` on every request, and the server
refuses destructive operations (`DELETE /api/sessions/:id`, kill, respawn stop) when that header
equals the target id, with a clear error.
This is a **footgun guard, not a security control**: any caller can omit the header. Document it
as such so nobody mistakes it for a boundary. It costs about 10 lines in `route-helpers.ts`.
### 2.5 Verification
Per the always-end-to-end-test rule, "the skill exists" is not done. Done is:
1. Symlink it into `.claude/skills/`, start a real throwaway Codeman session, and ask that agent
to "start a worker session that runs the test suite and tell me when it finishes".
2. Confirm from the outside that exactly one new session appeared, got the prompt, and that the
lead agent waited rather than polling in a busy loop.
3. Confirm the guard: run the same prompt in a shell with `CODEMAN_MUX` unset and confirm refusal.
4. Confirm cleanup: the worker session is deleted by exact id and no other session was touched.
Never run this against `w1`/`w2`/`w3`.
### 2.6 Files touched
- `skills/codeman/SKILL.md` (new), `skills/codeman/reference/*.md` (new)
- `.claude/skills/codeman` symlink (new)
- `src/cli.ts` (new `skill install` subcommand)
- `src/hooks-config.ts` (new `applyAgentSkill(casePath, enabled)`, mirroring `applyStatusLineConfig`)
- `src/web/schemas.ts` (`agentSkillEnabled` in `SettingsUpdateSchema`, which is `.strict()`)
- `src/web/routes/system-routes.ts` (settings PUT must resolve the flag from `merged`, never
from the raw body, per the partial-PUT invariant)
- `src/web/public/settings-ui.js` + `index.html` (checkbox)
- `package.json` `files` array, so `skills/` ships to npm
- README pointer, `docs/extending-codeman.md` seam 3 pointer
---
## 3. Part 2: wait primitives
### 3.1 Goal
Make Codeman orchestratable from a shell tool. Today the only "tell me when" channel is SSE,
which a curl-driven agent cannot practically consume: it would have to hold a streaming
connection and parse events inline. herdr solves this with blocking CLI calls. Codeman should
solve it with bounded long-poll endpoints.
All three additions are **additive**, so the versioning policy stays intact (new endpoints and
new optional fields are non-breaking).
### 3.2 The signal model
A waiter resolves on the first of a set of signals. Sources that already exist:
| Signal | Source today |
| --------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
| `idle` | `Session` emits `idle` (session.ts ~1775 for Claude, ~2101 for shell), wired at `session-listener-wiring.ts:402` |
| `working` | `Session` emits `working` (session.ts ~1788), wired at `session-listener-wiring.ts:401` |
| `stop` | `POST /api/hook-event` with `event: 'stop'`, the definitive "Claude finished responding" signal already used by `controller.signalStopHook()` |
| `blocked` | `POST /api/hook-event` with `permission_prompt` or `elicitation_dialog` |
| `exit` | `Session` emits `exit` |
`stop` is the highest-quality signal for "the turn is over" and should be the documented default
for orchestration. `idle` is heuristic: output stabilization plus prompt detection, and it can
flap mid-turn when a spinner pauses. External CLI modes (`isExternalCliMode()`) have no stop
hook at all, so for opencode/codex/gemini/antigravity only `idle`, `working` and `exit` are
available. **The skill and the docs must say which signals exist per mode**, otherwise an agent
waits forever on `stop` in a codex session.
### 3.3 Endpoint specs
#### A. `GET /api/sessions/:id/wait`
| Param | Type | Default | Notes |
| --------- | ---------------------------------------------- | ---------------- | ------------------------------------------------------------ |
| `until` | comma list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on first match |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` (600000) |
| `fresh` | `0`/`1` | `0` | `1` requires a _transition_, ignoring the state at call time |
Response (always 200 unless the session is missing or a cap is hit):
```json
{
"success": true,
"data": {
"signal": "stop",
"timedOut": false,
"immediate": false,
"ended": false,
"waitedMs": 8421,
"status": "idle",
"sessionId": "...",
"until": ["stop", "idle", "exit"],
"limitPaused": false
}
}
```
`until` is echoed back because the server may narrow it: `stop`/`blocked` are dropped
from the DEFAULT set for external CLI modes (asking for them EXPLICITLY is a 400
instead, since omitting `until` must never 400). `limitPaused` tells a caller that a
timeout was expected rather than a stall worth retrying hard.
**A timeout is not an error.** `{"timedOut": true, "signal": null}` with HTTP 200, so a caller
can loop without treating every poll boundary as a failure. Errors are reserved for
`NOT_FOUND` (unknown or not-owned session) and `SESSION_BUSY` (waiter cap exceeded).
`immediate: true` means the session was already in the requested state and `fresh` was not set.
#### B. `GET /api/sessions/:id/wait-output`
| Param | Type | Default | Notes |
| --------- | ------------------------------ | -------- | --------------------------------------------------------- |
| `match` | literal string, 1 to 200 chars | required | substring match against ANSI-stripped output |
| `nocase` | `0`/`1` | `0` | case-insensitive compare |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the existing text buffer first, then waits |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` |
Response: `{ matched: true, timedOut: false, snippet: "...", waitedMs }`.
**No regex in v1, deliberately.** `search-service.ts` already avoids regex specifically so there
is no ReDoS surface, and this endpoint would be even more exposed since the pattern is attacker
supplied and the input is a live stream. herdr can offer `--regex` because Rust's regex crate is
linear-time with no backtracking; JS `RegExp` is not. If regex is wanted later, the honest
options are a length-capped subset compiled once with a match budget, or `re2`. Note it and move on.
Implementation detail that will bite if missed: a match can straddle two PTY chunks. Keep a
carry buffer of `match.length - 1` bytes from the previous chunk and test `carry + chunk`.
⚠️ **`from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, resize, or any TUI redraw, and a repaint arrives as ordinary `terminal`
data. Observed live: a marker echoed a minute earlier matched instantly on a fresh
`from=now` wait. This is inherent to a terminal multiplexer, not fixable in the registry,
so the contract is: **use a marker unique per call** (`echo DONE_$RANDOM`), never a
generic one like `BUILD OK`. The skill's recipes must show that.
The returned snippet is whitespace-collapsed (blank runs to a single newline) for
readability only; matching runs on the raw stripped text. Without it, a real pane's
`\r\n` padding between the prompt and the match fills the whole context window with
nothing, which was the first thing the live test showed.
#### C. `wait` on the existing input endpoint
`POST /api/sessions/:id/input` gains two optional fields:
```json
{ "input": "run the tests\r", "useMux": true, "clientId": "agent-1", "seq": 7, "wait": "stop", "waitTimeout": 600000 }
```
(The trailing `\r` is required on every input body: `sendInput` sends Enter only
when the input contains a carriage return.)
Response gains `"wait": { "signal": "stop", "timedOut": false, "waitedMs": 41230 }`.
This is the important one, because it closes a race the standalone `GET .../wait` cannot: between
"input delivered" and "session flips to working" there is a window where a naive
send-then-wait sees the _pre-existing_ idle state and returns instantly. The combined endpoint
**registers the waiter before writing**, so that window does not exist. This is exactly why herdr
ships `agent prompt --wait` as its own thing.
`wait` accepts `true` (the default signal set) or the same comma grammar as `until`.
Both new fields are `.nullish()`, not `.optional()`: a third-party caller building the
body with `JSON.stringify` keeps an explicit `null` on the wire, and `.optional()`
rejects that with `INVALID_INPUT`. That gotcha has shipped as a real bug twice.
Two behaviors to preserve carefully:
- **`useMux` is fire-and-forget today.** The handler responds without awaiting `writeViaMux`, on
purpose (a tmux child process must not block the HTTP response). With `wait` present the
handler already has to stay open, so it can await delivery, and a `writeViaMux` failure becomes
observable for the first time. The non-wait path must keep its current fire-and-forget shape
byte for byte.
- **Duplicate suppression.** A tagged redelivery (`clientId`+`seq` already applied) returns 200
without writing. With `wait` set it still waits, since the caller's intent is "tell me when
this settles". But it waits with `requireTransition: false`, unlike a fresh delivery: the
original turn may be long over, and requiring a new transition would block a redelivery until
timeout for no reason. Fresh delivery requires a transition, a duplicate answers from the
current state.
- **Capacity rollback.** `shouldApplyInput()` MUTATES (it records the seq), and it runs before
the waiter is registered. If registration then fails on a full pool, the handler must call
`forgetInputSeq` before returning `SESSION_BUSY`, or the caller's retry is rejected as a
duplicate and the input is lost by the very mechanism reliable delivery exists for.
### 3.4 Module design
New file `src/web/session-wait-registry.ts`, with the IO-free core unit-testable in isolation
(same split as `self-update.ts`):
```ts
type WaitSignal = 'idle' | 'working' | 'stop' | 'blocked' | 'exit';
waitForSignal(sessionId, { until: Set<WaitSignal>, timeoutMs, requireTransition }): Promise<WaitResult>
notifySignal(sessionId, signal: WaitSignal): void
waitForOutput(sessionId, { match, nocase, timeoutMs }): Promise<OutputWaitResult>
notifyOutput(sessionId, chunk: string): void
cancelAll(sessionId, reason): void
```
Wiring points, all existing:
- `src/web/session-listener-wiring.ts` around lines 190 and 200 already handles `working` and
`idle` and broadcasts them. Add a `notifySignal()` call next to each broadcast, plus `exit`.
- `src/web/routes/hook-event-routes.ts` already switches on `event` for the respawn controller.
Add `notifySignal(sessionId, 'stop' | 'blocked')` in the same switch.
- Output: `notifyOutput()` rides the ALREADY-attached `terminal` listener in
session-listener-wiring.ts. An earlier draft had the registry hand out attach/detach
callbacks so a listener could be added lazily; that was deleted once it was clear no
second listener is needed at all. The cost is one Map lookup per PTY chunk, which is why
the no-waiter check comes before the ANSI strip.
- Session deletion calls `notifySignal('exit')` then `cancelAll()`, so no promise is left
hanging. Both are required: `_doCleanupSession` detaches the session's listeners BEFORE
`session.stop()`, so on a delete the PTY exit event never reaches the registry, and an
`until=exit` caller would otherwise get a bare `ended` instead of its signal. Found by
live-testing the delete path, not by the unit tests.
Memory-leak discipline, per the 24-hour-session rules: every waiter owns a timer that is cleared
on resolve, the per-session waiter set is deleted when it empties, and the output listener is
removed with it. `test/memory-leak-prevention.test.ts` should grow a case for this.
Caps in a new `src/config/agent-wait.ts` (limits live in `src/config/`, env-overridable):
| Constant | Default | Why |
| ------------------------- | ------- | --------------------------------------- |
| `MAX_WAIT_MS` | 600000 | an unbounded long-poll is a socket leak |
| `DEFAULT_WAIT_MS` | 60000 | short enough to survive most proxies |
| `MAX_WAITERS_PER_SESSION` | 16 | |
| `MAX_WAITERS_TOTAL` | 128 | same reasoning as `MAX_SSE_CLIENTS` |
Exceeding a cap returns `SESSION_BUSY`, not a silent queue.
### 3.5 Transport concerns
Fastify is constructed with defaults in `server.ts:329-331`. `requestTimeout` defaults to 0
(disabled) and `keepAliveTimeout` (72s) applies between requests, not to an in-flight one, so a
10-minute in-process hold is fine. **Verify this on the real instance before relying on it.**
Intermediaries are the actual risk. Prod is reached through `tailscale serve`, and users also run
cloudflared tunnels; both can cut an idle connection. That is why `DEFAULT_WAIT_MS` is 60s and
why the documented pattern is a client-side loop over short waits rather than one 10-minute call.
The skill's recipes must show the loop.
### 3.6 Edge cases to get right
| Case | Behavior |
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Session already idle, `fresh=0` | return immediately, `immediate: true` |
| Session already idle, `fresh=1` | wait for the next transition into a requested state |
| Session dies mid-wait | resolve with `signal: "exit"` if `exit` was requested, otherwise resolve `timedOut:false, signal:null, ended:true`. Never hang |
| Session deleted mid-wait | same, resolve, do not throw. Verified live: `until=exit` gets `signal:"exit"`, a concurrent `until=blocked` gets `ended:true`, both in ~0ms |
| Shutdown with a wait pending | `cancelEverything()` in `stop()`. Verified live: SIGTERM with a 300s wait in flight exits in 1s |
| External CLI mode | `stop` and `blocked` never fire. Reject `until=stop` for those modes with a clear `INVALID_INPUT` rather than hanging until timeout |
| Multi-user | goes through `findSessionOrFail(ctx, id, req)`, which already enforces ownership |
| Remote / Docker cases | signals originate from the same `Session` object, so no special casing. Docker hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for `stop`/`blocked` to arrive at all; without it, only `idle` works. Document it |
| Respawn `/clear` mid-wait | a respawn cycle emits `idle`. Callers waiting on `stop` are unaffected; callers on `idle` may resolve early. Documented, not fixed |
| Limit pause | if the session is paused on a usage limit, nothing will fire until the reset. The wait times out honestly. Consider surfacing `limitPaused: true` in the response so the caller can back off |
### 3.7 Tests
- `test/session-wait-registry.test.ts` (pure): immediate resolve, transition-required, multi-signal
first-wins, timeout, cap exceeded, cancel on session end, no listener leak after resolve,
chunk-straddling output match, case-insensitive match.
- `test/routes/session-wait-routes.test.ts` (`app.inject()`, no port): all three endpoints against
a `MockSession`, including the 200-with-`timedOut` contract and the ownership 404.
- `test/routes/session-input-wait.test.ts`: the send-and-wait race, plus proof that the non-wait
path is unchanged (still returns before `writeViaMux` settles).
- Live verification on a throwaway session before COM, per the always-end-to-end-test rule.
### 3.8 Files touched
- `src/config/agent-wait.ts` (new)
- `src/web/session-wait-registry.ts` (new)
- `src/web/session-listener-wiring.ts` (notify on idle/working/exit)
- `src/web/routes/hook-event-routes.ts` (notify on stop/blocked)
- `src/web/routes/session-routes.ts` (two new routes, `wait` fields on input)
- `src/web/schemas.ts` (`SessionWaitQuerySchema`, `SessionWaitOutputQuerySchema`, extend
`SessionInputWithLimitSchema`. Note: `.optional()` rejects `null`, so the frontend and any
generated client must send `undefined`, never `null`)
- `docs/api-reference.md`, `docs/extending-codeman.md`, README API table
- `skills/codeman/SKILL.md` recipes (Part 1 depends on this)
---
## 4. Deferred: parts 3 to 5
Not in scope now, kept here so they are not lost.
### Part 3: promote `blocked` to a first-class state
`SessionStatus` is `'idle' | 'busy' | 'stopped' | 'error'`. "Needs you" exists three times over:
hook events, the `tab-alert-action` CSS class, and the phone overview NEEDS YOU section, each
re-deriving it. herdr makes `blocked` a real state that rolls up.
Add `blocked` (and possibly `done`) to `SessionStatus`, set it from the same hook events that
Part 2 uses as wait signals, and clear it on the next `working`/`stop`. Then the tab strip, the
mobile overview, the wait endpoints, and any external agent read one field.
Cost: `SessionStatus` is a widely-consumed union, so every exhaustive `switch` (the codebase has
`assertNever` and `noFallthroughCasesInSwitch`) will need a branch. That is a feature, it makes
the compiler find every site. This is a **minor** bump, not a patch: it widens a public type in
the HTTP contract.
### Part 4: `GET /api/schema`
herdr ships `herdr api schema`. Every Codeman route is already Zod-validated, so
`zod-to-json-schema` over `schemas.ts` gives a self-describing API almost free. Value: third-party
tools and the skill stop drifting from hand-written docs. Open question: whether to emit full
OpenAPI (`@fastify/swagger` would need per-route schema registration, which is a much larger
change) or just dump the Zod schemas keyed by name (cheap, 80% of the value).
### Part 5: detection manifests instead of hardcoded patterns
CLI-specific readiness, blocked and usage-limit patterns live in code across
`usage-limit-patterns.ts`, the respawn pattern helpers and `regex-patterns.ts`. Externalizing the
per-CLI ones into data files would make adding a sixth CLI a data change instead of a code change.
**Do not copy the remote-update part.** herdr auto-fetches manifest updates from herdr.dev.
Codeman auto-pulling behavioral rules from a vendor server contradicts its security posture.
Bundled manifests plus local override only, no network.
### Explicit non-goals
- **Plugin runtime and marketplace.** `docs/extending-codeman.md` already argues this: a plugin
runtime means third-party code inside a process that spawns agents with your credentials, on a
server people expose over a tunnel. The reasoning still holds. If the marketplace _pattern_ is
wanted, apply it to data (web tabs, case templates, cron recipes), never to executable code.
- **Live PTY handoff on restart.** herdr needs it because it owns the terminals. Codeman
delegates to tmux, so PTYs already survive a self-update restart.
- **Socket API.** HTTP plus SSE is the existing, documented, stable contract. A second transport
would double the surface for no capability gain.
---
## 5. Sequencing
| Step | Work | Gate |
| ---- | ------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1 ✅ | `src/config/agent-wait.ts` + `session-wait-registry.ts` + unit tests | 48 tests green |
| 2 ✅ | `GET .../wait` + wiring in listener-wiring, hook-event-routes, server teardown | 15 route tests green; live-verified on an isolated `CODEMAN_INSTANCE=waittest` instance (immediate resolve, 400 on a bad signal, 200+`timedOut` on timeout, hook `stop` and `permission_prompt`→`blocked` waking an in-flight wait, delete delivering `exit`, SIGTERM not blocked); full `test:ci` sweep green |
| 3 ✅ | `GET .../wait-output` | 16 route tests green; live-verified on real PTY bytes (`echo MARKER` waking a blocked request in ~1s, `from=buffer` immediate hit, never-seen marker timing out at exactly 2001ms, nocase, `regex` refused with a 400); full `test:ci` sweep green |
| 4 ✅ | `wait` field on `POST .../input`, non-wait path proven unchanged | 16 route tests green; live-verified (no-wait returns in 26ms with the historical bare body; an idle session did NOT satisfy a `wait` request, blocking the full 2001ms, which is the race the endpoint exists to close; the stop hook resolved a send-and-wait at 1510ms and the input was confirmed in the tmux pane; `wait:null` accepted) |
| 5 ✅ | `skills/codeman/SKILL.md` + reference files + `.claude/skills` symlink | live dogfood: a real session orchestrates a worker end to end |
| 6 ✅ | `codeman skill install` CLI + `applyAgentSkill()` + `agentSkillEnabled` setting | 10 unit tests (`test/agent-skill.test.ts`) + real-server case-creation tests (`test/quick-start.test.ts`, incl. the settings PUT accepting the key) green; CLI verified live (install/uninstall, global + `--case`, foreign/symlink refusals) |
| 7 ✅ | Docs: api-reference, extending-codeman, README | plus `architecture-invariants.md` (§agent-wait-primitives), `CLAUDE.md` and the API reference's per-mode signal table |
| 8 ✅ | COM (minor bump: new endpoints, new setting, new optional fields) | released as 1.13.0 (wait primitives + skill); step 6 followed in 1.14.1 and was republished as 1.14.2 after live-testing the packaged skill |
Parts 1 and 2 are independent enough to land separately, but the skill is much less useful
without the wait endpoints, so the wait work goes first.
## 6. Open questions for the owner
1. ✅ `skills/` at the repo root: accepted (built that way; the install one-liner depends on it).
2. ✅ `agentSkillEnabled` default: **OFF** for the first release, per §2.2's rationale (skills
cost context on every turn; measure before defaulting on). Flip later if dogfooding earns it.
3. ✅ Both: global install via `npx skills add` / `codeman skill install`, AND per-case
auto-injection behind the (default-off) setting. Injection is add-only at session create and
marker-guarded, so a user-authored copy is never touched.
4. Is `X-Codeman-Caller-Session` self-protection worth the 10 lines, given it is a footgun guard
and not a security boundary? (Still open, not built with step 6.)
5. ✅ Regex support in `wait-output`: literal-only shipped, and a `regex` query param is
rejected with a 400 rather than ignored, so an agent that assumed otherwise cannot
silently wait on the wrong thing.
---
## 7. Build log: what actually happened
Written at the end of the build so the next person inherits the reasoning, not just the
diff. Process artifacts (per-agent briefs, findings, reports) live in the gitignored
`tmp/agent-wait-review/`; this section is the part worth keeping.
### What shipped
| Piece | Files |
| ------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| Bounds + clamping | `src/config/agent-wait.ts` (new) |
| Blocking-wait registry | `src/web/session-wait-registry.ts` (new, IO-free, unit-tested) |
| `GET .../wait`, `GET .../wait-output`, `wait`/`waitTimeout` on `POST .../input` | `src/web/routes/session-routes.ts` |
| Signal wiring | `session-listener-wiring.ts` (idle/working/exit + output), `hook-event-routes.ts` (stop/blocked), `server.ts` (teardown, shutdown) |
| Agent skill | `skills/codeman/SKILL.md` + `reference/`, `.claude/skills/codeman` symlink, `package.json` `files` |
| Docs | `api-reference.md`, `extending-codeman.md`, `architecture-invariants.md`, `README.md`, `CLAUDE.md` |
| Tests | `test/session-wait-registry.test.ts`, three `test/routes/session-*wait*.test.ts`, `http-contract.test.ts`, `mock-session.ts` |
### Bugs found in ADJACENT code, not in the new feature
These are the highest-value output of the exercise and none were on the plan:
1. **Every Codeman hook was dead on HTTPS installs.** `hooks-config.ts` built the hook
curl as `curl -s` with no `-k` while the statusline exporter 300 lines below used
`curl -sk` and documented why. Proven with the real hook command: `curl exit=60`
without the flag, success with it, and the failure swallowed by the hook's own
`2>/dev/null || true`. This silently killed `stop`, `permission_prompt`,
`elicitation_dialog`, `idle_prompt`, `teammate_idle` and `task_completed`, taking
respawn's definitive idle signals with them. Fixed, **plus** a staleness detector in
`refreshStaleCodemanHooks` that regenerates the on-disk config of already-created
cases (23 of 26 local cases carried the broken form; fixing the generator alone would
have left every one of them broken).
2. **`buildEnvExports()` exported a wrong-scheme `CODEMAN_API_URL`** (`http://` fallback
on an HTTPS install). Now omitted rather than guessed, so in-session guards fail closed.
3. **Programmatic input is only submitted when it contains `\r`.** `sendInput` sends Enter
only if the payload has a carriage return; without it the text sits in the composer
forever. Bit this build repeatedly before it was diagnosed, and had leaked into the
docs' own examples.
### Design decisions worth not re-litigating
- **A timeout is HTTP 200** with `wait.timedOut`, never a 4xx: callers loop over short
waits because tunnels cut idle connections, and every poll boundary would otherwise be
indistinguishable from failure.
- **Send-and-wait must be one endpoint.** A separate POST-then-wait races: between the
write and the flip to `working`, a wait sees the stale `idle` and reports the PREVIOUS
turn as this one. The waiter is registered before the write.
- **`stop`/`blocked` exist for `claude` mode only.** They come from Claude Code hooks;
`shell` installs none either, so keying off `isExternalCliMode()` was wrong.
- **Literal matching only, never regex.** JS `RegExp` backtracks; herdr can offer
`--regex` because Rust's regex crate is linear-time.
- **Client-hangup abort listens on `reply.raw` guarded by `writableFinished`.** On
`req.raw`, `close` fires when the request BODY ends, which on a POST killed every
send-and-wait instantly, and no `app.inject()` test can see it (inject never emits
`close`).
- **Liveness cannot come from `session.pid`.** For a tmux session that is the local
`tmux attach` client, not the worker: a worker exiting inside its pane leaves
`pane_dead=1` with the client alive, so `pid` never goes null. Liveness is probed at
the mux layer, cached (~750 ms) and only on blocking waits, never on the input hot path.
### Verification rounds
Six agents across three rounds, each verifying the previous round's work rather than its
own. Findings that mattered, in order of severity, were: the dead-pane liveness gap; the
`reply.raw` abort regression; abandoned long-polls leaking waiter slots; a crashed session
reporting `idle`; `shell` accepting `until=stop`; and a documented recipe that reported
success without running its task. Two traps recurred often enough to name:
- **Vacuous passes.** `app.inject()` never emits `close`; a latched `cancelEverything()`
in `afterEach` silently killed the registry for every later test in a file; three test
files sharing one session id against the process-wide registry let one file's leftover
waiter fail another's assertion. Any new wait test needs care on all three.
- **HTTP-only test instances.** Every isolated instance used during the build was plain
HTTP, which is exactly why the HTTPS hook bug survived so long. Test the transport the
user actually runs.
### Resolved at wrap-up (2026-08-08, conclusion pass)
- **R2-A**: the fire-and-forget-then-gather-sequentially pattern was **removed from
the skill** rather than patched. Signals are edge-triggered with no history, so a
`stop` that fires before its waiter registers is unobservable afterwards; a
`fresh=0` gather was rejected because the only `until` set that current state can
satisfy answers `idle` for a prompt that never submitted, resurrecting the exact
false-success failure R2-B had just closed. Flow 3b's pattern B now gathers on
latched `wait-output` markers (`from=buffer`), the same mechanism that makes the
shell flows reliable; the limitation is recorded in
`architecture-invariants#agent-wait-primitives` and `endpoints.md`. The durable
fix, a latched last-signal-per-turn on the server, stays with deferred Part 3.
- Docs F7/F8, F4 and the false-`idle` attribution: `api-reference.md`,
`extending-codeman.md` and `architecture-invariants.md` rewritten to the post-fix
matcher (one normalized stream, chunk-straddling found, snippet as a rendering of
the matched window), the real no-PTY answer (`ended:true`, `aborted:false`,
`delivered:false`), and the startup-idle mechanism (a session parked on the trust
dialog emits no further `idle`; the false success is the startup transition).
- Orchestrate #12, #5/R2-B, #6, and R2-C..R2-E: fire-and-forget's empty `data`
documented; every send-and-wait retry loop now treats `duplicate:true` +
`immediate:true` as "no new turn ran" and reads the terminal before believing it;
claude fan-out is pattern A (backgrounded send-and-waits) or the marker gather;
readiness budgets rebalanced (5 s stage 1, 45 s stage 3) with the virgin-case
floor named; the auth fallback now also reads the supervisor definition
(`codeman-web.service` / launchd plist) and accepts `export`-prefixed `.env`
lines; `pid != null` is documented as startup-only, never liveness.
- Both public readiness recipes (extending-codeman.md, README) are bypass-first with
the trust probe as the bounded fallback; the worked recipe carries `-k` and fails
loudly on an empty SID; the hook `-k`/self-heal fix appears in every
"hooks go missing" list; the multi-word-TUI claim is "unreliable", not "never".
### Still open
Both release-checklist items that used to sit here are done: `skills/` is tracked and
ships through `package.json` `files` (published with 1.13.0, republished with 1.14.2),
and the changeset was consumed, committed and deployed. What is left:
- Deferred with Part 3: the latched last-signal-per-turn. Nice-to-haves from the
reviews: N2 (create the death-watcher inside its `try`, still built one line above
it in `GET .../wait`) and converting timeout-shaped test detections into fast
assertions.
- §2.4's `X-Codeman-Caller-Session` footgun guard: still not built (open question 4).
### Step 6 (2026-08-09): install command, per-case injection, the setting
Built to the §2.6 file list, mirroring the statusLine mechanism throughout:
| Piece | Where |
| ----- | ----- |
| `applyAgentSkill(casePath, enabled)` + `installAgentSkillInto` / `removeAgentSkillFrom` | `src/hooks-config.ts` |
| `codeman skill install` / `skill uninstall` (`--global` default, `--case <name>`) | `src/cli.ts` |
| `agentSkillEnabled` (SYNCED, default OFF) | `schemas.ts` (`SettingsUpdateSchema`), `getAgentSkillEnabled()` on `ConfigPort`/`server.ts`, checkbox in `index.html` + `settings-ui.js` |
| Injection call sites (Claude mode only) | `POST /api/sessions` next to `refreshStaleCodemanHooks`; `POST /api/quick-start` after the case-create/self-heal blocks (local + docker cases; remote skipped, its path lives on another host) |
| Tests | `test/agent-skill.test.ts` (10 unit), `test/quick-start.test.ts` (real server: default-off, PUT accepts key, injection on create, shell-mode skipped) |
Decisions worth keeping:
- **Ownership marker, prefix-matched.** The injected SKILL.md ends with
`<!-- codeman-managed-agent-skill: … -->`; install/refresh/remove all refuse a copy
without the marker (a user's own skill) and match on the PREFIX so a wording change
cannot disown older injected copies (the `BACKGROUND_WAKE_MARKER_PREFIX` pattern).
- **Symlink refusal.** This repo's own dogfooding layout
(`.claude/skills/codeman -> ../../skills/codeman`) means the injector must `lstat`
the skill dir AND its `skills/` parent and bail on a symlink, or enabling the
setting in the Codeman repo itself would overwrite the skill source through the link.
- **ADD-ONLY at session create**, same shared-`.claude` rationale as the statusLine:
a create while the setting is off must not yank the skill out from under other live
sessions in the repo. The remove path exists (CLI `skill uninstall`, tests); no
automatic sweep removes on toggle-off.
- **Removal is manifest-based, never `rm -rf`**: only files the packaged source would
have written are deleted, directories are pruned bottom-up only if they emptied, so
a user's extra notes in `reference/` survive an uninstall.
- **Source resolution**: `join(moduleDir, '..', 'skills', 'codeman')` works from
`src/` (tsx), `dist/` (tsc build), and the npm tarball alike, because all three sit
one level below the package root and `files` ships `skills/`.
- **Nothing acts on the setting at PUT time**: injection reads the merged persisted
settings at session create (`readSettings`, ~2s cache), so the partial-PUT invariant
(`toggleService` reading `merged`) is untouched by construction.
### 2026-08-09 addendum: cross-session messaging folded into the skill
Claude Code 2.1.224+ ships cross-session messaging: `ListAgents`/`SendMessage`
tools, a per-session Unix inbox socket, and a registry in
`~/.claude/sessions/<pid>.json`. Codeman's claude workers are ordinary local Claude
Code sessions, so the skill now routes task delivery and result collection over it
when available, while the HTTP primitives keep spawn, readiness, synchronization,
liveness and delete. New `skills/codeman/reference/messaging.md` (ships with zero
installer changes: `readAgentSkillSource()` enumerates `reference/*.md` from disk),
Flow 5 in recipes.md, and §4 in SKILL.md.
Verified live (claude-cli 2.1.226, Linux):
- A message to an idle worker starts a turn and that turn fires the normal `stop`
hook (8.3 s send-to-stop measured), so the HTTP wait primitives compose with
messaging unchanged; delivery to a busy session lands between tool calls.
- First contact needs the `name [ref]` form; the bare name errors with the exact
string to resend. The `uds:` reply address of an inbound message works as a `to`.
- The `tmux codeman-<id8>` column in `ListAgents` (and the registry's `tmux` field)
is the join key to Codeman session ids. The registry's `sessionId` field starts as
the Codeman id (we spawn `claude --session-id <id>`) but drifts after `/clear` or
resume, so it must never be the join key.
- The feature is flag-gated beyond the version: two 2.1.226 sessions on one machine,
one with an inbox socket and one without. Absence is a fallback case, not an error.
- Codeman's default `--dangerously-skip-permissions` spawn puts both ends in the
bypassing class, which delivers; mixed classes hold behind an approval dialog that
expires unattended (upstream default 5 min), which on a headless worker means the
message silently dies. The skill's backstop covers it.
Follow-up, landed in the same PR: local claude spawns now pass
`--name <session name>` so peers carry Codeman session names. The gate is
`buildNameCliArgs()` (session-cli-builder.ts), fail-closed at
`CLAUDE_NAME_FLAG_MIN_VERSION = 2.1.224`: that is the messaging release, the flag's
presence there was verified against the installed 2.1.224 binary, and the version
comes from `getClaudeCliVersion()` (null on probe failure and under vitest), so an
older or unknown CLI gets a command byte-identical to before. That matters because
claude aborts startup on an unknown option, which would kill every session spawn.
The value is allowlist-sanitized (Unicode letters/digits plus ` ._:-`, leading
dashes stripped so it cannot parse as another option, 64-char cap, empty result =
flag omitted) before the double-quoted interpolation in `buildSpawnCommand`, and
only the LOCAL command carries it: the docker/remote builders never see it, since
their CLI is not the binary the probe measured. E2E on an isolated instance
(`CODEMAN_INSTANCE`): process cmdline `claude ... --name w9-msgtest`, registry
`name: "w9-msgtest"`, `ListAgents` lists it under that name, a message round-trip
works, and its replies arrive tagged `from-name="w9-msgtest"` (a derived-name
worker's replies carry no `from-name`). A quick-start without `sessionName` has an
empty Codeman name, so the peer name stays derived: agents should name their
workers. Tests: `test/name-flag-injection.test.ts`.
-627
View File
@@ -1,627 +0,0 @@
# HTTP API Reference
Codeman's HTTP API is a **stable contract** as of 1.0 — see
[`versioning-policy.md`](versioning-policy.md) for the SemVer guarantee. This page
defines the response envelope, status codes, error codes, versioning, and the SSE
event channel.
## Versioning
- The stable, public surface is served under **`/api/v1/...`**. Pin external
clients to this prefix.
- The unversioned **`/api/...`** paths are a permanent alias of the current
version (what the bundled web UI uses). They are kept working, but new external
integrations should use `/api/v1`.
- Breaking changes to the contract ship under a new prefix (`/api/v2`); `/api/v1`
keeps its semantics. Additive changes (new endpoints, new optional fields, new
error codes) are non-breaking and may appear in a minor release.
- The implementation rewrites `/api/v1/*` → `/api/*` at the server level
(`rewriteApiV1Url` in `src/web/server.ts`).
## Response envelope
Every JSON response uses one uniform envelope, applied centrally by a
`preSerialization` hook (`src/web/server.ts`) — handlers return bare data and the
hook wraps it:
**Success** — HTTP `2xx`:
```json
{ "success": true, "data": <payload> }
```
`data` is the endpoint's payload (object, array, or value). Endpoints with no
payload return `{ "success": true, "data": {} }`.
**Error** — HTTP `4xx`/`5xx`:
```json
{ "success": false, "error": "human-readable message", "errorCode": "NOT_FOUND" }
```
`ApiResponse<T>` in `src/types/api.ts` is the canonical type.
> Non-JSON endpoints are exempt from the envelope: `GET /api/sessions/:id/file-raw`,
> `GET /api/sessions/:id/tail-file` (SSE), `GET /api/download`,
> `GET /api/screenshots/:name`, `GET /q/:code` (QR redirect), and the
> `GET /ws/sessions/:id/terminal` WebSocket upgrade.
> The [agent wait endpoints](#long-polling-agent-wait) use the normal envelope but
> are the only JSON endpoints that deliberately **hold the connection open**, for up
> to 600 s. Proxy operators and HTTP clients with a global read timeout need to know
> that before pointing them at Codeman.
⚠️ **A `401` is the one status that is not an envelope.** Authentication is rejected
in a request hook, before any handler runs, and it replies with the bare string
`Unauthorized` (`Unauthorized: hook secret required` on the hook path) plus
`WWW-Authenticate: Basic realm="Codeman"`. There is no `success`, no `error`, and no
`errorCode`, because the wrapping hook only wraps object payloads. So a client that
pipes every response straight into a JSON parser dies with a parse error rather than
reporting an auth failure, which is a confusing way to discover that a password is
set. Branch on the HTTP status **before** parsing.
## Error codes → HTTP status
The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
`src/types/api.ts`. Clients should branch on `errorCode` (stable) and may rely on
the HTTP status.
| `errorCode` | HTTP | Meaning |
|-------------|------|---------|
| `INVALID_INPUT` | 400 | Malformed request / failed validation |
| `UNAUTHORIZED` | 401 | Authentication required or failed |
| `NOT_FOUND` | 404 | Resource does not exist |
| `SESSION_BUSY` | 409 | Session is busy |
| `CONFLICT` | 409 | Conflicts with current state (e.g. already running) |
| `ALREADY_EXISTS` | 409 | Resource already exists |
| `OPERATION_FAILED` | 422 | Well-formed but could not be completed |
| `RATE_LIMITED` | 429 | Too many requests |
| `INTERNAL_ERROR` | 500 | Unexpected server error |
Adding a new error code is non-breaking; removing or renaming one is a major change.
## Long-polling (agent wait)
Three calls block until something happens instead of answering immediately. They
exist because SSE is Codeman's only other "tell me when" channel, and an agent
driving the API from a shell tool cannot practically hold a stream and parse
events inline.
| Call | Blocks until |
|------|--------------|
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
`GET .../wait`. It registers the waiter **before** writing, which closes the window
in which a separate wait sees the session still idle from the previous turn and
answers instantly with the wrong turn's result. Use it whenever you send a prompt
and want to know when that prompt is done.
### Three semantics that break callers who assume otherwise
**1. A timeout is HTTP `200`, not an error.** A wait that ends without its signal
returns `{"success":true, ...,"wait":{"timedOut":true,"signal":null}}`. The
intended pattern is a client-side loop over short waits, because `tailscale serve`
and cloudflared can both cut an idle connection, and turning every poll boundary
into a `4xx` would make that loop indistinguishable from a real failure. `408` is
auto-retried by several clients (silently doubling the polling load), `504` is what
a genuine tunnel failure looks like, and `204` cannot carry `waitedMs` / `status` /
`limitPaused`. Reserve error handling for the four codes in the table below.
**2. `stop` and `blocked` fire only for `claude` sessions.** Both come from Claude
Code hooks, and no other mode installs them: `shell` runs no agent, and the external
CLIs (`opencode`, `codex`, `gemini`, `antigravity`, `pi`) render their own TUIs and post
no hooks. For every non-`claude` mode only `idle`, `working` and `exit` are
accepted, and of those only `exit` is dependable: see the caveats under
[Signals](#signals) before building on `idle`. Requesting `stop` or `blocked`
**explicitly** on such a session is a
`400`; omitting `until` never fails, the server just drops them from the default set
and echoes the narrowed set back as `wait.until`. Three more places hooks can go
missing even in `claude` mode: a **Docker case** needs
`CODEMAN_DOCKER_BRIDGE_HOOKS=1`, since a container cannot reach a loopback-bound
Codeman (without it, only `idle` / `working` / `exit` work); a **remote-SSH
case** runs the agent on another host, whose hooks may never reach this server at
all; and a case whose hook config was written by **Codeman < 1.13.0 against an
`--https` install** carries hook curls without `-k`, which TLS-fail silently (the
hook line ends in `|| true`). Codeman now writes `curl -sk` and repairs a stale
case config the next time a session starts in that case. When in doubt, ask for
`stop,idle,exit` so a session without hooks still resolves on the heuristic
signal.
**3. `from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, on resize, and on any TUI redraw, and a repaint arrives as
ordinary output, so text that was already on screen can satisfy a fresh wait. This
was observed live: a marker echoed a minute earlier matched instantly on a new
`from=now` wait. It is inherent to running the agent under a multiplexer, so the
contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
`echo $MARK`, then wait on `$MARK`), never a generic string like `BUILD OK`.
### Signals
| Signal | Source | Actually fires for |
|--------|--------|--------------------|
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
| `exit` | no process is behind the session | every mode |
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
is omitted is `stop,idle,exit` (`exit` is in there so a worker that crashes resolves
the wait promptly instead of burning the caller's whole timeout on something that
can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit`
once the session is up: the default set's `idle` also resolves on a spinner pause,
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
can land inside your first wait window and report a turn that never ran. Measured:
a session parked on the trust dialog emits no *further* `idle`, so it is the
startup transition, not the dialog, that produces the false success below.
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
server answers from `pid === null` plus a mux-layer pane-death probe, and that
covers a session that exited — including a worker that died *inside* its tmux pane
while the local attach client (and therefore `pid`) lives on — one that was
detached, and one that was **created but never started**. So the first wait
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
milliseconds, and reading that as "the worker died" is wrong: it means start it, or
wait for it to come up. `status` is carried alongside so nothing is hidden. The
alternative (trusting `status`) is worse, because a dead PTY parks the session at
`status: "idle"`, which would answer the default wait with `immediate: true` for a
worker that has crashed. A worker dying while a wait is parked resolves it within
a few seconds (a background death-watcher), not at the timeout.
⚠️ **`blocked` is reachable less often than the table suggests.** It fires on two
hooks, and the default configuration suppresses one of them: Codeman spawns claude
with `--dangerously-skip-permissions`, so permission prompts do not happen unless the
instance is switched to the `auto` Claude mode (App Settings), or the caller is a
multi-user account without the bypass grant, which is forced to `--permission-mode
auto`. What does still fire under the default is `elicitation_dialog`, the agent
asking the user a question. So `until=stop,blocked,exit` is a reasonable belt on a
long turn, but a worker that never comes back is far more likely to be working than
blocked, and polling `blocked` alone will sit at its timeout.
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
session emits its one `idle` at startup and then stays `status: "idle"` forever,
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
requires a transition (and so does `fresh=1`), both can only time out there:
a documented default `wait` on a shell worker running `sleep 4` times out at the
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
instead. The same caution applies to the external CLIs.
### Readiness is not a signal
Nothing here reports "the agent is ready for a prompt", and no combination of
`until`/`fresh` synthesizes one. A freshly created session reads as `exit` (above),
and a `claude` worker in a brand-new case comes up on the CLI's **trust dialog**,
which contains a ❯ prompt of its own. Send-and-wait posted at that moment types the
prompt into the dialog, where the `\r` never gets past it, while the session's
startup `idle` lands inside the wait window: the wait resolves on `idle` in a
couple of seconds with `timedOut: false`, which looks exactly like a completed
turn.
The reliable sequence is: poll `GET /api/v1/sessions/:id` until `.data.pid` is
non-null, then `wait-output` for the composer's own marker (`bypass`, the status
bar of a CLI spawned in bypass mode) with a short timeout, handling the trust
dialog only as the bounded fallback.
⚠️ **The fallback is not a bare `\r`.** Claude Code 2.1.252 unnumbered the dialog's
options, reversed them and highlights `No, exit`, so an Enter sent blind quits the
CLI and the pane dies seconds after the spawn. Read the `❯` marker off the current
frame (`GET /api/v1/sessions/:id/terminal?full=1`), send `ESC [ B` while it is on
`No, exit`, re-read, and confirm only once it is on `Yes, I trust this folder`.
Reading the current frame is also what keeps this correct on later runs: the dialog
text stays in the terminal buffer for the life of the session, so a `trust` probe
with `from=buffer` keeps matching long after the dialog is gone. A worked version is in
[`extending-codeman.md`](extending-codeman.md#seam-3-http-api-and-cli).
### `GET /api/v1/sessions/:id/wait`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
```bash
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
```
Both GET wait routes answer with `Cache-Control: no-store`, because the documented
pattern polls one identical URL in a loop and a cached `{"timedOut":true}` would
turn that loop into a busy spin. `POST .../input` sends no cache header (it is a
POST, which is not heuristically cacheable).
⚠️ **Unknown query parameters are ignored, not rejected**, with one exception
(`regex`, below). In particular `match=` on `/wait` is silently dropped and you get
a plain signal wait, so check the endpoint path before blaming the parameters.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
instead of waiting on the wrong thing. The reasoning is in
[`architecture-invariants.md`](architecture-invariants.md#agent-wait-primitives).
#### What the matcher actually sees
The matcher scans the raw PTY stream, **normalized**: ANSI escape sequences are
stripped — CSI, OSC, and the charset-designation escapes a stock bash prompt emits
on every line (`ESC ( B`), so `match=tnode:` matches a prompt that renders
`…@tnode:` — a partial escape arriving at a chunk boundary is held back until its
tail arrives, and a match may straddle PTY chunks: `printf STRAD; sleep 1; printf
DLEQQ` is matchable as `STRADDLEQQ` (all measured live). Three caveats remain:
⚠️ **It is still the byte stream, not the rendered pane.** `GET .../terminal`
answers from a tmux screen capture (`data.source: "mux-visible"`), the finished
picture; the matcher sees the stream that painted it. For linear output the two
agree once escapes are stripped, but a full-screen TUI composes its picture with
cursor positioning, so what the pane shows and what the stream carries can differ.
Seeing your string in `terminal?tail=` makes a match likely, not guaranteed.
⚠️ **A TUI's text can arrive without its spaces.** Claude Code positions words
with cursor moves rather than printing spaces, so screen text can reach the
matcher as `Quicksafetycheck:Isthisaprojectyoucreated...`. Whether a given phrase
keeps its spaces depends on how the TUI happened to draw it (measured: `I trust
this folder` matched, `Quick safety check` did not), so a multi-word `match`
against a TUI pane is unreliable rather than impossible. Match a **single
space-free token**, ideally one you printed yourself. Plain command output (a
shell worker, an `echo`) keeps its spaces.
⚠️ **The returned `snippet` is a rendering of the matched text, not a quotation of
it.** It is cut from the same normalized stream the match ran against, then
cleaned for display: remaining raw control bytes are removed (an agent pipes the
snippet into its own terminal, so a worker's bytes must not be able to reset that
display) and blank runs are collapsed. A printable needle that matched will appear
in it; a needle containing control bytes or a blank run may not survive verbatim.
```bash
MARK="DONE_$RANDOM"
curl -sG "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'timeout=120000'
```
Build the query with `-G --data-urlencode` rather than by hand: a `+` in a
hand-written query string decodes to a space.
### `POST /api/v1/sessions/:id/input` with `wait`
Two optional fields on the existing endpoint:
| Field | Type | Notes |
|-------|------|-------|
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
"absent" rather than failing validation. That is deliberate: `.optional()` would
reject it, which has shipped as a real bug twice.
The input must end with `\r` (a real carriage return in the JSON string): Enter is
sent only when the input contains one, so text without it is typed onto the
worker's prompt but never submitted, and the wait then runs its full timeout on a
turn that never started. Verified live; this is the most common silent failure on
this endpoint.
```bash
curl -s -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","useMux":true,"clientId":"agent-1","seq":1,
"wait":"stop","waitTimeout":600000}'
```
A **tagged duplicate** (a `clientId` + `seq` pair the server has already applied)
still honors `wait`, because the caller's question is unanswered, but it answers
from the session's current state rather than requiring a new transition: the
original turn may be long over. It comes back as
`"delivered": false, "duplicate": true`.
### Response
All three nest the wait result under `data.wait`, so one client helper works against
any of them:
```json
{ "success": true, "data": {
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
"status": "idle",
"limitPaused": false,
"wait": {
"signal": "stop", "until": ["stop", "idle", "exit"],
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
"waitedMs": 8421, "timeoutMs": 60000
}
}}
```
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
`status` and `limitPaused`. `POST .../input` **without** `wait` is unchanged and
still returns `{"success": true, "data": {}}`.
⚠️ `delivered: false` has **two** meanings, and they must be told apart by
`duplicate`: with `duplicate: true` the input was suppressed as an already-applied
redelivery (harmless, the turn it refers to may be long over), while with
`duplicate: false` the **write failed** (typically no PTY behind the session). A
client that reads `delivered === false` as "duplicate" silently treats a failed send
as a success.
| Field | Type | Meaning |
|-------|------|---------|
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
Read the outcome by discriminator, in this order:
1. `wait.signal !== null` (or `wait.matched === true`): the thing happened.
2. `wait.timedOut`: a poll boundary. Loop again.
3. `wait.ended` or `wait.aborted`: the wait answered nothing, because the session is
gone or was never running. Re-check the session instead of looping.
`wait.immediate` is not a fourth outcome: it rides along with the first one and
means the condition already held at call time, so nothing was actually waited for.
If that is not what you meant, you wanted `fresh=1` or the send-and-wait form. Note
that `{"signal":"exit","immediate":true}` on a session you just created is the
not-started-yet case, not a crash.
**The timeout is clamped, so read it back.** A request for 1800000 ms is silently
reduced to the server's ceiling (600000 ms by default, operator-tunable), and a
request for 1 ms is raised to 1000 ms. `wait.timeoutMs` is the value that was
applied. Without checking it, a caller that asked for 30 minutes and got 10 will
read the timeout as "the worker is wedged" and kill a session that was working fine.
### Errors
| `errorCode` | HTTP | When |
|-------------|------|------|
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
The two capacity codes are deliberately different. A process-wide cap reported as
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
error message names the cap that was hit.
⚠️ A `401` is **not** in this table and is not an envelope at all (see
[Response envelope](#response-envelope)). It matters most here: a polling loop that
pipes each wait straight into `jq` fails with a parse error on every iteration
against a password-protected server, which reads as "the wait endpoints are broken".
Check the status first.
The per-session cap is a **combined** budget: signal waiters and output waiters
count against the same 16, not 16 of each. An abandoned request no longer holds its
slot, because the routes release the waiter when the client disconnects, but a
client that opens many concurrent waits against one session will still hit the cap.
## Session lineage (`parentSessionId`)
A create request may name the session that spawned it, which the web UI draws as a
line between the two tabs. Accepted on `POST /api/v1/sessions` and
`POST /api/v1/quick-start`, either way:
```bash
# as a body field
-d '{"caseName":"worker-1","mode":"claude","parentSessionId":"'"$CODEMAN_SESSION_ID"'"}'
# or as a header, which is what an agent driving many spawns should use: set it once
# on the curl invocation and every spawn call carries it
-H "X-Codeman-Parent-Session: $CODEMAN_SESSION_ID"
```
The body field wins if both are present. The value is resolved against live sessions
(exact id, or a unique prefix of at least 8 characters) and must belong to the same
owner as the session being created.
**It cannot fail your spawn.** An unknown, stale, foreign or malformed value is
silently dropped and the session is created without lineage — never a `400`. It is
also pure decoration: it confers no permission, and a child is unaffected by its
parent exiting. It appears on session state as `parentSessionId` (absent when
unresolved) and survives a server restart.
## Approvals Inbox
Cross-session queue of prompts waiting on a human (permission dialogs,
AskUserQuestion questions, idle prompts). Claude-mode sessions only; items are
in-memory (a server restart drops them; the next prompt re-fires the hook).
Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md).
- `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first,
ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId,
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
toolSummary?, message?, cwd?, context?, options?: {n, label}[],
acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
`options` is present only when the dialog's numbered choices parsed
confidently; `acknowledgedAt` marks an item a human has already looked at
(see `/viewed` below) and tells clients not to re-arm its tab alert. Listing
also runs a staleness sweep over the caller's own items: the pane is
re-captured, and an item whose dialog no longer parses is resolved as
`resolved_in_terminal` instead of being returned (only items whose original
frame parsed `options` can be dropped this way, so an unreadable capture
keeps the item).
- `POST /api/v1/approvals/:id/answer` with `{ action: 'approve' }` (sends the
digit `1`), `{ action: 'deny' }` (sends Esc), `{ action: 'option', option: n }`
(sends the digit; accepted only when `n` is among the item's parsed
`options`), or `{ action: 'text', text }` (idle prompts only; submits the
line as a prompt). `404 NOT_FOUND` when the item is no longer pending,
`409 CONFLICT` when the dialog left the screen or another actor answered
first, `422 OPERATION_FAILED` when the session refused input.
- `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes.
- `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId,
acknowledged: itemId | null }`. Marks the session's pending **idle** item as
seen by a human (the web UI calls it when you open the session's tab): the
item stays pending and answerable, but stops arming the yellow tab alert on
every client, including after a reload. Permission/question items are never
acknowledged this way, since looking at a dialog does not answer it. `404`
for an unknown or inaccessible session; acknowledging twice is a no-op
(`acknowledged: null`).
SSE events: `approval:pending` (full item), `approval:updated` (context/options
re-captured, or the item acknowledged), `approval:resolved` (`{ id, sessionId, kind, resolution }` with
`resolution` one of `answered | resolved_in_terminal | superseded |
session_ended | dismissed | expired`).
## Reboot restore
A host reboot takes the tmux server down with it, so every pane dies and the
board comes up empty. At boot Codeman works out which sessions the reboot
destroyed and holds that plan in memory, and these endpoints let a client offer
it to the user. Nothing creates a pane until the user asks: the boot-time reboot
heuristic decides whether to ASK, never whether to act.
Claude-mode sessions only (others carry their conversation id in their own
config object); remote and docker sessions are never offered, because both need
another host or container to be up. The plan is in-memory, so a server restart
drops it and the offer is gone; the conversations themselves are unaffected,
since they live in the CLI's own transcript store and stay reachable from the
Resume list. A plan nobody spends expires after 24 hours.
- `GET /api/v1/reboot-restore` → `{ sessions: RestorableSession[],
scrollbackRestored: false }`, ownership-scoped in multi-user mode.
`RestorableSession`: `{ id, name?, workingDir, mode, owner? }`. The persisted
record itself is never sent. `scrollbackRestored` is always `false` and exists
so a client states it: a restored session is a NEW pane, so the conversation
continues and the terminal history does not.
- `POST /api/v1/reboot-restore/restore` with `{ sessionIds?: string[] }` (omit
to restore everything the caller can see) → `{ restored: RestorableSession[],
skipped: { sessionId, reason }[] }`. `reason` is one of `workspace-missing`
(the directory is gone), `workspace-forbidden` (in multi-user mode it is
outside the workspace of the user the session belongs to, re-checked against
that owner's current grant rather than the caller's), `already-live` (the conversation is already
open, typically resumed by hand from the Resume list), `capacity-reached`
(the global or per-user session cap), or `rebuild-failed` (the agent would not
start, most often a CLI binary missing from the server's PATH).
`409 CONFLICT` when that caller already has a restore running. Entries are
removed from the plan before any pane is built, so a double-click cannot put
two panes on one conversation; anything that never became a pane goes back on
offer, except `already-live`, which cannot stop being true. A restored session
comes back attached, idle and disarmed: respawn controllers and Ralph loops
are never re-armed automatically.
- `POST /api/v1/reboot-restore/dismiss` → `{ dismissed: n }`. Drops the offer
for everything the caller can see.
Each rebuilt session also emits the ordinary `session:created` SSE event, so
clients other than the one that clicked pick it up without refetching.
## Read My Mind intent profiles
Per-case profiles of what the user is trying to accomplish: user/agent-stated
goals plus the user's recently submitted prompts, captured from the Claude
session transcript while the opt-in `readMyMindEnabled` setting is on (default
OFF). Keyed by owner + workingDir, so the profile survives `/clear`, respawns,
and session churn. Stored in `~/.codeman/intents.json` (mode 0600); never fed
into `/api/v1/search`. Design: [`readmymind-plan.md`](readmymind-plan.md);
user guide: [`readmymind.md`](readmymind.md).
- `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the
session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals,
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
50, each <= 500 chars). A case with nothing recorded answers an empty
profile with `updatedAt: 0`; nothing is persisted by reads.
- `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict
schema) replaces the goals text and answers the updated profile.
`400 INVALID_INPUT` on over-long or unknown fields.
- `DELETE /api/v1/sessions/:id/intent` -> `{ deleted: boolean }` forgets the
case's profile entirely.
- `POST /api/v1/sessions/:id/readmymind` predicts the user's next prompt:
a one-shot model call over the intent profile plus live session signals
(pending approval dialog, transcript tail, git state, run-summary events,
sibling sessions). Body is optional; the rethink flow passes
`{ steer?, rejected? }` (strict schema: `steer` <= 2000 chars, `rejected`
up to 10 strings <= 1000 chars). Answers
`{ suggestions: { prompt, why, kind }[], durationMs }` with 1-3 suggestions
(`kind`: `continue` | `verify` | `redirect`; prompts are single-line).
Claude-mode sessions only (`400 INVALID_INPUT` otherwise); one prediction in
flight per session (`409 CONFLICT`); predictor failures answer
`502 OPERATION_FAILED`. Takes 5-90 s and costs real tokens. Suggestions are
only ever returned, never sent: submitting one is the caller's explicit act.
All four enforce session ownership in multi-user mode; a foreign session id
answers `404 NOT_FOUND` (no existence leak), and profiles of two owners of the
same directory are distinct by construction.
## Voice dictation
Browser dictation transcribed through this server's Claude Code login, i.e. the
same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synced
`claudeVoiceEnabled` setting (default OFF). Design:
[`claude-voice-plan.md`](claude-voice-plan.md).
- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?,
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
signed in to Claude Code on the server), `expired` (the access token elapsed;
running any Claude session refreshes it) or `malformed`. The OAuth token
itself is never returned by this or any other endpoint.
- `GET /ws/voice/stream?language=&keyterms=` (WebSocket, not under `/api`)
relays one dictation. Client sends binary frames of signed 16-bit
little-endian PCM, 16 kHz mono (<= 64 KB per frame), plus JSON control frames
`{"t":"finalize"}` (ask for the final transcript) and `{"t":"stop"}`. Server
sends `{"t":"ready"}`, `{"t":"transcript","text","final"}` (each frame is the
WHOLE running transcript, not a delta), `{"t":"error","message"}` and
`{"t":"closed"}`. Close codes: `4003` disallowed Host/Origin, `4004`
unavailable (reason in the close reason), `4008` too many concurrent streams.
Streams are capped in count and length (`src/config/voice.ts`).
## Authentication
Optional HTTP Basic (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`) → opaque
`codeman_session` cookie. When enabled, unauthenticated requests get
`401 UNAUTHORIZED`; rate-limited requests get `429 RATE_LIMITED`. See
[`security-architecture.md`](security-architecture.md).
## SSE event channel
`GET /api/events` is a Server-Sent Events stream (`text/event-stream`); each
message is `event: <name>` + `data: <json>`. The event-name registry
(`src/web/sse-events.ts`, mirrored in `src/web/public/constants.js`) is part of
the stable contract — event names are not renamed without a major bump. An
optional `?sessions=<id,...>` filter suppresses only the high-volume terminal
stream; lifecycle/metadata events are delivered to all clients regardless.
### `sse:heartbeat` (liveness)
Every 15s the server writes a `sse:heartbeat` frame to every connected client:
```
event: sse:heartbeat
data: {"t":1755100000000}
```
`t` is the server's epoch-ms timestamp at write time. The frame carries no
application state and can be ignored for correctness. It exists so a client can
tell a live stream from a dead one: an `EventSource` whose connection has been
idle-closed by a proxy (or that resumed from sleep on a stale socket) keeps
delivering nothing without ever firing `onerror`. Clients that care should treat
silence longer than about three intervals as a dead stream and reconnect, which
is what the bundled frontend does.
This replaced a `:keepalive` SSE **comment**, which served the same
proxy-flushing purpose but is invisible to `EventSource` by spec and so could
never be observed by a client. Consumers written against the old behavior are
unaffected: `EventSource` dispatches only events that have a registered
listener, so an unknown event name is dropped.
## Consuming from JavaScript
The bundled frontend reads responses through `_apiJson()`
(`src/web/public/api-client.js`), which unwraps `{success:true,data}` → `data` and
returns `null` on a non-2xx / `{success:false}` response. External clients should
do the same: check the HTTP status (or `body.success`), then read `body.data`.
-107
View File
@@ -1,107 +0,0 @@
# Approvals Inbox (design)
One cross-session inbox for every prompt that is waiting on a human: permission dialogs, questions (AskUserQuestion / elicitation), and idle prompts. Cards are answerable in place (option digits, Esc, or a typed prompt) from desktop, phone overview, and push notification action buttons. Inspired by Cloudflare OS's Gatekeeper approval queue (https://github.com/cloudflare/cloudflare-os, asynchronous human-in-the-loop approvals): with a fleet of sessions the human is the bottleneck, and today answering means finding the right tab.
## Problems this fixes (all real today)
1. **No cross-session surface.** Pending prompts exist only as per-tab alert colors (`tab-alert-action`/`tab-alert-idle`) and NEEDS YOU rows on the phone overview. Answering means switching to the session and typing.
2. **Alerts die on reload.** `pendingHooks` lives only in `app.js` memory, fed by transient SSE `hook:*` events. A page reload (or a phone browser evicting the tab) silently loses every pending alert. There is no server-side record.
3. **Push Approve/Deny buttons are dead.** `PUSH_EVENT_MAP` already attaches `approve`/`deny` actions to permission pushes, and `sw.js` forwards `event.action` to the page, but the `notification-click` handler in settings-ui.js ignores it (and when no tab is open, the action is dropped entirely). The buttons render on the lock screen and do nothing.
4. **Card context is missing.** The frontend handlers read `data.question` / `data.message` / `data.tool`, but `sanitizeHookData` never forwards `message`, so notifications show generic fallback text.
## Scope
- Claude mode only (hooks fire only for `claude`; external CLIs keep their output-stabilization heuristics and get no inbox items). This mirrors the wait-primitive `stop`/`blocked` gating.
- Permission prompts occur for sessions running `ClaudeMode` `normal` / `auto` / `allowedTools` (and the trust-folder dialog even under skip-permissions). Question and idle prompts occur in every mode including `dangerously-skip-permissions`.
- In-memory store (plus the frontend seeding from it on load). Server restart drops items; hooks re-fire on the next prompt. No new state file in v1.
## Data model
At most **one active item per session**: the Claude TUI shows one dialog at a time, so a new prompt event supersedes the session's previous item (resolution `superseded`).
```ts
interface ApprovalItem {
id: string; // `${sessionId}:${seq}`
sessionId: string;
sessionName: string;
kind: 'permission' | 'question' | 'idle';
createdAt: number;
toolName?: string; // from sanitized hook data
toolSummary?: string; // command / file_path / description, already bounded
message?: string; // Notification hook `message` (newly allowlisted)
cwd?: string;
context?: string; // ANSI-stripped visible pane frame tail, ≤ 4000 chars
options?: { n: number; label: string }[]; // parsed from context when confident
}
```
Resolutions (server-emitted, item removed from pending): `answered` (via inbox), `resolved_in_terminal` (stop / elicitation_complete / elicitation_response / session went working), `superseded`, `session_ended`, `dismissed`, `expired` (12h TTL sweep).
## Backend
### Store: `src/approval-inbox.ts`
Module-level singleton in the style of `session-wait-registry.ts` (pure, no `Session` import, injected emit callback so there is no import cycle with the server):
- `notePrompt(info)` creates/supersedes the session's item; schedules ONE re-capture ~600ms later (the Notification hook can fire before the dialog finishes painting) which updates `context`/`options` and emits `approval:updated`.
- `resolveForSession(sessionId, reason)`, `dismiss(id)`, `answerable(id)`, `listPending()`, `stop()` (clears timers; tests).
- Option parsing (pure, unit-tested): consecutive `❯? N. label` lines, 2..6 options, labels ≤ 120 chars. Parsed options gate which digits the answer endpoint accepts; when parsing fails the card falls back to Approve(1)/Deny(Esc) only.
- TTL: items expire after 12h (checked on read + a lazy sweep; no standing interval).
### Wiring
- `hook-event-routes.ts`: on `permission_prompt` / `elicitation_dialog` / `idle_prompt`, call `notePrompt` with sanitized data + a pane capture callback (`mux.capturePaneBuffer(muxName)` visible frame, ANSI-stripped via existing utils; fall back to `session.terminalBuffer` tail). On `stop` / `elicitation_complete` / `elicitation_response`, `resolveForSession(id, 'resolved_in_terminal')`.
- `session-listener-wiring.ts`: `working` listener resolves **idle items only** (`working` is heuristic and can flap mid-turn, so it must never clear a pending permission/question dialog); `exit` resolves with `session_ended`. Same singleton-import pattern as `sessionWaits`.
- Session delete route: resolve with `session_ended`.
- **New hook matchers** `elicitation_complete` + `elicitation_response` added to `generateHooksConfig()`, `HookEventType`, `HookEventSchema`, and both SSE registries. `refreshStaleCodemanHooks` gets a staleness probe for them (`hooksJson.includes('elicitation_complete')`) so existing cases heal on next Claude spawn, exactly like the `-k`/secret/marker probes.
- `sanitizeHookData`: allowlist `message` (bounded 500 chars). This also un-deadens the existing notification text paths.
### Routes: `src/web/routes/approval-routes.ts`
Normal authed API (NOT the hook-secret bypass), `ApiResponse` envelope, Zod schemas in `schemas.ts`:
- `GET /api/approvals` → pending items, multi-user filtered by `canAccessOwned` (same policy as session lists). Also sweeps the caller's own items for staleness through `verifyStillAnswerable()`: Claude Code fires no "permission answered" hook, so a dialog answered in the terminal used to sit pending until `stop` and re-arm a red tab alert on the next page load. Only items whose original frame parsed options can be dropped this way, so an unreadable capture keeps the alert.
- `POST /api/approvals/:id/answer` body `{ action: 'approve' | 'deny' | 'option' | 'text', option?, text? }`:
- `approve` → `writeViaMux('1')` (option 1 is always plain Yes; no Enter, menus react to the digit).
- `deny` → `writeViaMux('\x1b')` (Esc is the official No/cancel; precedent: auto-resume sends Esc the same way).
- `option` → digit `String(n)`; accepted only when `n` is within the item's parsed options (prevents blind digit-poking at an unparsed dialog).
- `text` → `idle` items only: single line, embedded newlines stripped, sent as `text\r` (the `\r` discipline from CLAUDE.md).
- Guards: item still pending (404 otherwise), session exists + ownership via `findSessionOrFail`, session mode installs hooks. **Answer-time re-capture**: for items whose frame parsed options, the pane is re-captured before sending; if the dialog no longer parses, the item resolves and the answer is refused with 409 (the keystroke would land in whatever now has focus). Marks `answered` BEFORE the write so a double-tap cannot double-send; rolls back to pending if the write fails.
- `POST /api/approvals/:id/dismiss` → remove without keystrokes.
- `POST /api/approvals/session/:sessionId/viewed` → acknowledge the session's pending **idle** item (`acknowledgedAt`, emitted as `approval:updated`). Added after the owner reported that a yellow tab clicked and checked went yellow again on reload: the view-clears-idle rule lived in one browser's memory, so the seed re-armed it and other devices never saw the clear. Acknowledgement is deliberately **not** resolution (the prompt is still unanswered, so it stays in the inbox and stays available as Read My Mind context), and deliberately **idle-only** (looking at a permission/question dialog does not answer it, so the red alert survives being viewed).
### SSE
`approval:pending`, `approval:updated`, `approval:resolved` in `sse-events.ts` + `SSE_EVENTS` in constants.js (the parity test pins the sync). Broadcasts carry `sessionId`, so multi-user SSE scoping applies unchanged.
### Push
- `sendPushNotifications` payload gains `approvalId` for the three hook events. Both `approvalId` and the Approve/Deny `actions` are **gated on the opt-in setting**: with it off, permission pushes carry no buttons at all (pre-inbox they rendered and did nothing, so stripping them is the honest shape).
- `sw.js` `notificationclick`: when `event.action` is `approve`/`deny`, POST `/api/approvals/:id/answer` directly from the worker (same-origin, cookie credentials) so the buttons work **with no tab open**; on failure fall back to focusing/opening a tab. Non-action clicks keep today's behavior.
- Page-side `notification-click` handler: honor `action` instead of dropping it (also setting-gated, for stale notifications sent before the toggle flipped).
- Question/idle pushes keep no action buttons (options vary per dialog); tapping opens the inbox.
## Frontend
New module `approvals-ui.js` (@loadorder 11.2, after panels-ui.js), prettier-formatted (not added to `.prettierignore`).
- **Seed on connect**: `GET /api/approvals` on init and SSE reconnect; each pending item re-feeds `setPendingHook(...)` so tab alerts and the phone overview survive reload (fixes problem 2 with zero changes to the alert state machine). Items carrying `acknowledgedAt` are skipped, and `markIdleAlertSeen()` (app.js) is what sets it: viewing a session clears its yellow locally and POSTs `.../viewed`, so "I checked it" survives the reload and reaches the user's other devices through `approval:updated`.
- **Desktop**: header bell `btn-approvals` with count badge. Ships default-hidden via marker class `btn-approvals--hidden` (same policy as the attachments button, so `test/mobile-header-buttons-policy.test.ts` excludes it from the default-visible enumeration); JS shows it only while count > 0. Click toggles a drawer of cards: session name + kind, tool/message summary, mono context block, buttons rendered from parsed options (else Approve/Deny), plus Dismiss and Open session. Esc closes; existing z-index layers respected.
- **Phone**: header button stays hidden (`mobile.css`); the phone surface is the overview's NEEDS YOU section, whose rows gain inline ✓/✗ buttons for permission items (tap-through to the session remains the row's main action). Toolbar classes/status language rules from the mobile-overview section of CLAUDE.md apply.
- **i18n**: new strings registered in i18n.js (en + zh-CN); status words carry `data-i18n-skip` where they would collide (mirroring the overview pills).
- **Setting**: `approvalsInboxEnabled`, synced (in `SettingsUpdateSchema`), **default OFF** (owner decision: the entire feature is opt-in, meaning no bell, no drawer, no overview strips, no seeding, and no push action buttons until enabled in App Settings → Panels). Only the store and answer endpoints keep running regardless, so flipping the toggle ON surfaces anything already pending immediately, with no restart.
## Race honesty
The prompt can be answered in the terminal a moment before an inbox answer lands; then the keystroke would hit whatever now has focus (worst case: a digit typed into the composer, not submitted, since no `\r` is ever sent for menu answers). Mitigations, in order: answer-time re-capture (the dialog must still parse on screen or the answer is refused), answered-before-write marking, digit-only/Esc-only writes for menus, and the card's context block showing what the pane looked like when captured. This is the same class of risk `writeViaMux` automation (auto-resume, respawn) already accepts.
## Tests
- `test/approval-inbox.test.ts`: supersede per session, every resolution path, TTL, option parsing fixtures (2-option, 3-option with ❯, unparseable frame), re-capture update.
- `test/routes/approval-routes.test.ts` (`app.inject`, no port): list; hook event creates item; answer approve/deny/option writes the exact bytes (test-PTY echo asserts them); text answers restricted to idle; 404 unknown id; 409 answered twice; option out of range rejected; multi-user scoping.
- Existing suites extended: hook-event schema accepts the two new events; `sanitizeHookData` forwards bounded `message`; SSE parity + mobile-header policy pass as-is by construction.
## Docs
- CLAUDE.md: Key Patterns entry + SSE/route counts + frontend load order.
- `docs/api-reference.md`: the two endpoints + three SSE events (additive, fine under the 0.9.x contract).
File diff suppressed because one or more lines are too long
-983
View File
@@ -1,983 +0,0 @@
> **⚠️ ARCHIVED 2026-05-21 — superseded, kept for history.**
> The "Critical" structural items here are done: `server.ts` 6,736→2,065 LOC,
> `app.js` 15,196→3,083 LOC, `types.ts` 1,443→12 LOC (now a barrel → `src/types/`).
> The phase plans that executed this work are in `docs/archive/phase*-plan.md`.
> Do not treat this as a live TODO; see CLAUDE.md for current architecture.
# Code Structure & Quality Findings
**Date**: 2026-02-28
**Scope**: Full codebase analysis across 5 dimensions: frontend, backend, TypeScript, testing, and utilities/config.
This document contains detailed findings for agent teams to write implementation plans and execute improvements. Each section includes severity, specific locations, and recommended fixes.
---
## Table of Contents
1. [Critical: server.ts God Object (6,736 LOC)](#1-critical-serverts-god-object)
2. [Critical: app.js Monolith (15,196 LOC)](#2-critical-appjs-monolith)
3. [Critical: CleanupManager Unused Despite Existing](#3-critical-cleanupmanager-unused)
4. [High: Duplicated Debounce/Timer Patterns](#4-high-duplicated-debouncetimer-patterns)
5. [High: Large Domain Files Need Splitting](#5-high-large-domain-files-need-splitting)
6. [High: types.ts God File (1,443 LOC)](#6-high-typests-god-file)
7. [High: Zod Schemas Duplicate TypeScript Types](#7-high-zod-schemas-duplicate-typescript-types)
8. [High: Test Coverage Gaps](#8-high-test-coverage-gaps)
9. [High: Duplicated Test Mocks](#9-high-duplicated-test-mocks)
10. [Medium: Hardcoded Magic Values](#10-medium-hardcoded-magic-values)
11. [Medium: Frontend Global State Monolith](#11-medium-frontend-global-state-monolith)
12. [Medium: Frontend Code Duplication](#12-medium-frontend-code-duplication)
13. [Medium: Inconsistent Logging](#13-medium-inconsistent-logging)
14. [Medium: Utils Barrel Export Gaps](#14-medium-utils-barrel-export-gaps)
15. [Medium: Non-Null Assertion Risks](#15-medium-non-null-assertion-risks)
16. [Low: Dead Utility Functions](#16-low-dead-utility-functions)
17. [Low: No Dependency Injection for File I/O](#17-low-no-dependency-injection-for-file-io)
18. [Scorecard & Prioritized Roadmap](#18-scorecard--prioritized-roadmap)
---
## 1. Critical: server.ts God Object
**File**: `src/web/server.ts` (6,736 lines)
**Severity**: CRITICAL
**Impact**: Hardest file to maintain, test, and extend. Imports 38 modules.
### Problem
The `WebServer` class handles everything: HTTP routing (~110 routes), authentication, SSE broadcasting, terminal data batching, state persistence, session lifecycle, respawn orchestration, file serving, tunnel management, plan orchestration, and subagent coordination.
**Key metrics**:
- 40+ private properties (Maps, timers, caches)
- 70+ methods
- `setupRoutes()` is 2,000+ LOC of inline route handlers
- Zero test coverage
### Current Structure (Bad)
```
WebServer class (6,736 LOC)
├── Auth session management (lines 469, 668-698)
├── SSE client management (lines 407-408, 5843-5880)
├── Terminal data batching (lines 414-416, 5909-5966)
├── Task update batching (line 426, 5995-6028)
├── State persistence batching (lines 429-430, 6028-6061)
├── Respawn lifecycle (lines 445-451, 5425-5534)
├── Session cleanup (lines 4769-4961)
├── Listener setup (lines 544-643)
└── setupRoutes() (lines 645+, 2000+ LOC)
├── /api/sessions/* (30+ routes inline)
├── /api/respawn/* (7 routes inline)
├── /api/subagents/* (7 routes inline)
├── /api/plan/* (5 routes inline)
├── /api/push/* (4 routes inline)
└── ... 60+ more inline
```
### Recommended Structure
```
src/web/
├── server.ts (~500 LOC - HTTP setup, route registration only)
├── routes/
│ ├── session-routes.ts (session CRUD, input, resize)
│ ├── respawn-routes.ts (respawn control endpoints)
│ ├── subagent-routes.ts (background agent tracking)
│ ├── plan-routes.ts (plan generation & management)
│ ├── push-routes.ts (web push subscriptions)
│ ├── mux-routes.ts (tmux management)
│ ├── case-routes.ts (case management)
│ ├── file-routes.ts (file browsing/serving)
│ └── system-routes.ts (status, stats, config, settings)
├── middleware/
│ ├── auth.ts (Basic Auth + session cookies)
│ └── error-handler.ts (centralized error responses)
└── services/
├── sse-manager.ts (SSE client + broadcast)
├── terminal-batcher.ts (60fps terminal batching)
└── session-lifecycle.ts (listener setup/teardown)
```
### Duplication in server.ts
**Error response pattern** repeated 189 times:
```typescript
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Session not found');
```
**Fix**: Extract `findSessionOrFail()` middleware:
```typescript
const findSessionOrFail = (sessionId: string) => {
const session = this.sessions.get(sessionId);
if (!session) throw new NotFoundError('Session not found');
return session;
};
```
**Event listener setup** copy-pasted for subagent watcher, image watcher, and team watcher (lines 544-643). Same attach/detach pattern duplicated 3 times.
---
## 2. Critical: app.js Monolith
**File**: `src/web/public/app.js` (15,196 lines)
**Severity**: CRITICAL
**Impact**: Untestable, hard to navigate, tightly coupled systems.
### Extractable Modules (by priority)
| Module | Lines | Current Location | Impact |
|--------|-------|------------------|--------|
| Mobile handlers (MobileDetection, KeyboardHandler, SwipeHandler) | ~300 | lines 168-620 | High |
| Voice input (DeepgramProvider, VoiceInput) | ~830 | lines 631-1471 | High |
| NotificationManager | ~450 | lines 2218-2663 | High |
| xterm-zerolag-input (inlined copy from packages/) | ~400 | lines 1756-2153 | High |
| KeyboardAccessoryBar | ~195 | lines 1480-1680 | Medium |
| FocusTrap | ~60 | lines 1690-1748 | Medium |
### CodemanApp Class (12,000+ LOC)
The main `CodemanApp` class starting at line 2665 has:
- **60+ Maps/Sets** in the constructor (lines 2667-2805)
- **18 Map instances** with complex cross-references (subagents, parents, teams, windows)
- **10+ monolithic methods** exceeding 100 lines each
**Largest methods**:
| Method | Lines | Size |
|--------|-------|------|
| `renderAppSettings()` | 14400-14700 | ~300 LOC |
| `selectSession()` | 6028-6250 | ~220 LOC |
| `batchTerminalWrite()` | 7482-7700 | ~200 LOC |
| `renderSessionTabs()` | 5814-6000 | ~180 LOC |
| `openSubagentWindow()` | 11927-12100 | ~170 LOC |
| `handleInit()` | 5183-5350 | ~170 LOC |
### Recommended Split
```
src/web/public/
├── app.js (~4000 LOC - core app, session mgmt, SSE)
├── mobile.js (~300 LOC - MobileDetection, KeyboardHandler, SwipeHandler)
├── voice.js (~830 LOC - DeepgramProvider, VoiceInput)
├── notifications.js (~450 LOC - NotificationManager)
├── keyboard-accessory.js (~200 LOC - KeyboardAccessoryBar)
├── api-client.js (~100 LOC - fetch wrapper with error handling)
└── config.js (~50 LOC - magic numbers, z-index layers)
```
---
## 3. Critical: CleanupManager Unused
**File**: `src/utils/cleanup-manager.ts` (320 lines)
**Severity**: CRITICAL
**Impact**: Memory leak risk. Well-designed utility exists but is never used. Every file manages cleanup manually.
### Current State
`CleanupManager` is exported from the utils barrel but has **0 instantiations** in production code. Instead, every file implements manual cleanup:
**respawn-controller.ts** (worst offender):
```typescript
// 11 timer properties, manually cleared in stop()
private stepTimer: NodeJS.Timeout | null = null;
private completionConfirmTimer: NodeJS.Timeout | null = null;
private noOutputTimer: NodeJS.Timeout | null = null;
// ... 8 more
stop() {
if (this.stepTimer) clearTimeout(this.stepTimer);
if (this.completionConfirmTimer) clearTimeout(this.completionConfirmTimer);
// ... 9 more clearTimeout/clearInterval calls
}
```
**Files that should use CleanupManager**:
| File | Timer/Listener Count | Current Cleanup |
|------|---------------------|-----------------|
| `respawn-controller.ts` | 11 timers + intervals | 11 manual clearTimeout/clearInterval |
| `web/server.ts` | 6+ timers, debounce map | Manual in stop(), some may leak |
| `state-store.ts` | 2 debounce timers | Manual clearTimeout |
| `push-store.ts` | 1 save timer | Manual clearTimeout |
| `subagent-watcher.ts` | debounce map + watchers | Manual clear + close |
| `ralph-tracker.ts` | 3 debounce timers | Manual clear |
| `bash-tool-parser.ts` | 1 debounce timer | Manual clear |
| `image-watcher.ts` | 1 debounce map | Manual clear |
### Fix
Migrate all timer management to use `CleanupManager`. Example for respawn-controller.ts:
```typescript
// Before: 11 fields + 11 clearTimeout calls
private stepTimer: NodeJS.Timeout | null = null;
// ...
// After: 1 field, auto-cleanup
private cleanup = new CleanupManager();
startStep() {
this.cleanup.setTimeout(() => { ... }, 5000, 'step');
}
stop() {
this.cleanup.dispose(); // Clears everything
}
```
---
## 4. High: Duplicated Debounce/Timer Patterns
**Severity**: HIGH
**Impact**: 8+ files implement debounce independently. Bug fixes need to be applied everywhere.
### Pattern Inventory
```typescript
// Pattern 1: Manual timer ref (used in 6 files)
private saveTimer: NodeJS.Timeout | null = null;
debouncedSave() {
if (this.saveTimer) clearTimeout(this.saveTimer);
this.saveTimer = setTimeout(() => this.save(), 500);
}
// Pattern 2: Timer Map (used in 3 files)
private fileDebouncers = new Map<string, NodeJS.Timeout>();
debounce(key: string) {
const existing = this.fileDebouncers.get(key);
if (existing) clearTimeout(existing);
this.fileDebouncers.set(key, setTimeout(() => { ... }, 100));
}
// Pattern 3: State flag (used in 2 files)
private isSaving = false;
```
### Locations
| File | Debounce Vars | Delay (ms) |
|------|---------------|------------|
| `state-store.ts` | `saveTimeout`, `ralphStateSaveTimeout` | 500 |
| `push-store.ts` | `saveTimer` | 500 |
| `web/server.ts` | `persistDebounceTimers` (Map) | 500 |
| `subagent-watcher.ts` | `fileDebouncers` (Map) | 100 |
| `ralph-tracker.ts` | 3 debounce timers | 50, 30000 |
| `bash-tool-parser.ts` | `EVENT_DEBOUNCE_MS` | 50 |
| `image-watcher.ts` | debounce map | 200 |
| `respawn-controller.ts` | 11 timer fields | various |
### Fix
Create a `Debouncer` utility:
```typescript
// src/utils/debouncer.ts
export class Debouncer {
private timer: NodeJS.Timeout | null = null;
constructor(private readonly delayMs: number) {}
run(fn: () => void): void {
if (this.timer) clearTimeout(this.timer);
this.timer = setTimeout(fn, this.delayMs);
}
cancel(): void {
if (this.timer) clearTimeout(this.timer);
this.timer = null;
}
}
// Usage:
private saveDeb = new Debouncer(500);
this.saveDeb.run(() => this.save());
// cleanup: this.saveDeb.cancel();
```
---
## 5. High: Large Domain Files Need Splitting
**Severity**: HIGH
**Impact**: Complex state machines spanning 3,000+ lines are hard to understand and test.
### ralph-tracker.ts (3,905 LOC)
**5 responsibilities mixed**:
1. Output Parsing (~900 LOC) - Line-by-line parsing, state extraction
2. Todo Management (~700 LOC) - Parsing, dedup, expiry
3. Plan Tracking (~800 LOC) - Enhanced plan tasks, checkpoints
4. Circuit Breaker (~400 LOC) - State machine for stuck detection
5. File Watching (~300 LOC) - Monitor external state files
**Recommended split**:
```
ralph-tracker.ts (core output parsing, ~1200 LOC)
ralph-todo-manager.ts (todo parsing + management, ~700 LOC)
ralph-plan-tracker.ts (plan tasks + checkpoints, ~800 LOC)
ralph-circuit-breaker.ts (circuit breaker logic, ~400 LOC)
```
### respawn-controller.ts (3,611 LOC)
**6 responsibilities mixed**:
1. State Machine (~1,000 LOC) - 6+ states, transitions
2. Idle Detection (~800 LOC) - 5 layers + multi-signal combining
3. AI Checkers (~600 LOC) - Idle + plan checkers integration
4. Health Scoring (~500 LOC) - Metrics, circuit breaker, scoring
5. Action Logging (~300 LOC) - Timeline, detection status
6. Stuck-State Detection (~250 LOC) - Timeout tracking
**Recommended split**:
```
respawn-controller.ts (state machine core, ~1000 LOC)
respawn-idle-detection.ts (all 5 idle detection layers, ~800 LOC)
respawn-health-scorer.ts (metrics & health scoring, ~500 LOC)
```
### session.ts (2,418 LOC)
**8 responsibilities mixed**:
1. PTY Management (~600 LOC)
2. Terminal I/O (~400 LOC)
3. Token Tracking (~200 LOC)
4. Task Tracking (~250 LOC)
5. Ralph Integration (~200 LOC)
6. Auto-Clear/Compact (~300 LOC)
7. Image Watching (~100 LOC)
8. CLI Detection (~150 LOC)
**Recommended split**:
```
session.ts (PTY + terminal I/O core, ~1000 LOC)
session-tracking.ts (token + task + Ralph, ~500 LOC)
session-auto-ops.ts (auto-clear/compact + image, ~300 LOC)
```
---
## 6. High: types.ts God File
**File**: `src/types.ts` (1,443 lines, 72 exported definitions)
**Severity**: HIGH
**Impact**: Every file imports from types.ts. Hard to find relevant types.
### Current Contents
- 46 interfaces
- 25 types
- 1 enum (ApiErrorCode)
- 9 factory functions (createInitialState, etc.)
### Recommended Split
```
src/types/
├── index.ts (barrel export - transparent migration)
├── session.ts (SessionState, SessionConfig, SessionMode, SessionColor)
├── task.ts (TaskState, TaskDefinition, TaskStatus)
├── respawn.ts (RespawnConfig, RespawnState, CircuitBreakerStatus)
├── ralph.ts (RalphLoopState, RalphTrackerState, RalphTodoItem)
├── api.ts (ApiResponse, ApiErrorCode, HookEventType, all route types)
├── lifecycle.ts (LifecycleEventType, LifecycleEntry)
└── common.ts (Disposable, BufferConfig, CleanupResourceType)
```
The barrel export makes this a transparent refactor - existing `import from './types'` continues to work.
---
## 7. High: Zod Schemas Duplicate TypeScript Types
**File**: `src/web/schemas.ts` (508 lines)
**Severity**: HIGH
**Impact**: When a type changes, the Zod schema must be manually updated too. Source of bugs.
### Problem
Zod schemas manually duplicate TypeScript interfaces. **Zero `z.infer` usage found.**
```typescript
// types.ts (manual interface)
export interface CreateSessionRequest {
workingDir?: string;
mode?: SessionMode;
name?: string;
}
// schemas.ts (manual Zod schema - duplicated!)
export const CreateSessionSchema = z.object({
workingDir: safePathSchema.optional(),
mode: z.enum(['claude', 'shell', 'opencode']).optional(),
name: z.string().max(100).optional(),
});
```
### Fix
Use `z.infer` to derive TypeScript types from Zod schemas (single source of truth):
```typescript
// schemas.ts
export const CreateSessionSchema = z.object({
workingDir: safePathSchema.optional(),
mode: z.enum(['claude', 'shell', 'opencode']).optional(),
name: z.string().max(100).optional(),
});
// types.ts (auto-derived)
export type CreateSessionRequest = z.infer<typeof CreateSessionSchema>;
```
**Affected schemas** (~10):
- CreateSessionSchema
- RunPromptSchema
- ResizeSchema
- CreateCaseSchema
- QuickStartSchema
- HookEventSchema
- RespawnConfigSchema
- ConfigUpdateSchema
- SettingsUpdateSchema
---
## 8. High: Test Coverage Gaps
**Severity**: HIGH
**Impact**: Critical code paths untested. Regressions go unnoticed.
### Untested Source Files
| File | Lines | Risk |
|------|-------|------|
| `src/web/server.ts` | 6,736 | CRITICAL - Core REST API, 280+ routes |
| `src/plan-orchestrator.ts` | ~500 | HIGH - Multi-agent plan generation |
| `src/tunnel-manager.ts` | ~200 | MEDIUM - Cloudflare tunnel |
| `src/session-lifecycle-log.ts` | ~150 | MEDIUM - JSONL audit log |
| `src/ai-plan-checker.ts` | ~300 | MEDIUM - Plan completion detection |
| `src/templates/claude-md.ts` | ~200 | LOW - CLAUDE.md generation |
| `src/utils/claude-cli-resolver.ts` | ~100 | LOW - CLI path resolution |
| `src/utils/opencode-cli-resolver.ts` | ~100 | LOW - OpenCode CLI support |
| `src/utils/regex-patterns.ts` | ~100 | LOW - Used everywhere! |
| `src/utils/token-validation.ts` | ~50 | LOW - Token counting |
### Test Quality Issues
**10 "not.toThrow()" tests without behavior verification**:
```typescript
// BAD: Only checks it doesn't crash
expect(() => tracker.processMessage(null)).not.toThrow();
// GOOD: Also verify defensive behavior
expect(() => tracker.processMessage(null)).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
```
Locations:
- `task-tracker.test.ts` - 5 instances
- `image-watcher.test.ts` - 1 instance
- `task-queue.test.ts` - 1 instance
- Others scattered
---
## 9. High: Duplicated Test Mocks
**Severity**: HIGH
**Impact**: Mock changes need updating in 4 places. Inconsistent mock behavior.
### MockSession Defined 4 Times
| File | Usage |
|------|-------|
| `test/respawn-controller.test.ts` | Full mock with event emitter |
| `test/session-manager.test.ts` | Simpler mock |
| `test/respawn-team-awareness.test.ts` | Copy of respawn-controller mock |
| `test/respawn-test-utils.ts` | **Comprehensive mock - UNUSED!** |
### MockStateStore Defined 2 Times
| File | Usage |
|------|-------|
| `test/session-manager.test.ts` | Basic mock |
| `test/ralph-loop.test.ts` | Separate implementation |
### Unused Test Utilities
`test/respawn-test-utils.ts` exports these utilities that **no test file imports**:
- `createTimeController()` - Abstraction over vitest fake timers
- `MockAiIdleChecker` - Fully mocked AI idle checker
- `MockAiPlanChecker` - Fully mocked plan checker
- Factory functions for pre-configured controllers
### Fix
Create `test/mocks/` directory:
```
test/
├── mocks/
│ ├── mock-session.ts (single MockSession, used everywhere)
│ ├── mock-state-store.ts (single MockStateStore)
│ └── index.ts (barrel export)
├── utils/
│ └── time-controller.ts (from respawn-test-utils.ts)
└── ... test files
```
---
## 10. Medium: Hardcoded Magic Values
**Severity**: MEDIUM
**Impact**: Hard to tune, inconsistent when same value appears in multiple places.
### Already Centralized (Good)
- `src/config/buffer-limits.ts` - All buffer sizes
- `src/config/map-limits.ts` - All collection limits
### NOT Centralized (40+ values scattered)
**In server.ts** (lines 145-194):
```typescript
const TASK_UPDATE_BATCH_INTERVAL = 100;
const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
const SESSIONS_LIST_CACHE_TTL = 1000;
const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
const SSE_HEALTH_CHECK_INTERVAL = 30 * 1000;
const MAX_TERMINAL_COLS = 500;
const MAX_TERMINAL_ROWS = 200;
const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
const MAX_AUTH_SESSIONS = 100;
const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
const STATS_COLLECTION_INTERVAL_MS = 2000;
const MAX_INPUT_LENGTH = 64 * 1024;
```
**In hooks-config.ts**: `timeout: 10000` hardcoded 6 times.
**In respawn-controller.ts** (lines 538-565): 10 timing constants.
**In utils**: `EXEC_TIMEOUT_MS = 5000` duplicated in both `claude-cli-resolver.ts` and `opencode-cli-resolver.ts`.
**In app.js**:
```javascript
// line 27: 600000 - stuck detection threshold
// line 24: 5000 - default scrollback
// lines 34-35: 128*1024, 256*1024 - chunk sizes
// lines 152-155: 150, 100 - keyboard detection thresholds
// lines 573-575: 80, 300, 100 - swipe detection params
```
### Fix
Create additional config files:
```
src/config/
├── buffer-limits.ts (existing)
├── map-limits.ts (existing)
├── server-config.ts (NEW - web server intervals, auth, caching)
├── timing-config.ts (NEW - debounce delays, check intervals)
└── terminal-config.ts (NEW - max cols/rows, batch intervals)
```
---
## 11. Medium: Frontend Global State Monolith
**Severity**: MEDIUM
**Impact**: All state in single CodemanApp class. Tight coupling between unrelated systems.
### 60+ State Variables in CodemanApp Constructor (lines 2667-2805)
```javascript
this.sessions = new Map(); // Session data
this.subagents = new Map(); // Agent tracking
this.subagentActivity = new Map(); // Tool call tracking
this.subagentToolResults = new Map(); // Result caching
this.subagentParentMap = new Map(); // Agent-to-session mapping
this.teams = new Map(); // Team tracking
this.teamTasks = new Map(); // Team task state
this.planSubagents = new Map(); // Plan agent tracking
this.pendingWrites = []; // Terminal write queue
this.terminalBufferCache = new Map(); // Buffer caching (unbounded!)
this.projectInsights = new Map(); // Bash tool insights
// ... 40+ more
```
### Problems
1. **18 Map instances** with complex cross-references (no garbage collection strategy)
2. **No domain separation**: Session, subagent, notification, UI, and network state mixed
3. **Implicit dependencies**: `selectSession()` requires 5+ Maps to be in consistent state
4. **`terminalBufferCache`** has no max size - can grow unbounded with many sessions
### Recommended Domain Split
```javascript
// Instead of 60+ flat properties:
class SessionState {
sessions = new Map();
sessionOrder = [];
terminalBuffers = new Map();
tabAlerts = new Map();
}
class SubagentState {
subagents = new Map();
activity = new Map();
parentMap = new Map();
windows = new Map();
minimized = new Map();
}
class TeamState {
teams = new Map();
tasks = new Map();
teammates = new Map();
}
class UIState {
activeSessionId = null;
draggedTabId = null;
isLoadingBuffer = false;
}
```
---
## 12. Medium: Frontend Code Duplication
**Severity**: MEDIUM
**Impact**: Repeated patterns increase maintenance burden and inconsistency risk.
### Duplicated Patterns
**API fetch calls** (~50 instances):
```javascript
// Repeated everywhere:
fetch(`/api/sessions/${sessionId}/...`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({...})
}).catch(() => {})
```
**Fix**: Extract `ApiClient` class.
**`innerHTML` usage** (104 instances):
- Mix of template strings, createElement chains, and direct innerHTML
- Some with manual XSS escaping (`text.replace(/</g, '&lt;')`), some without
- No consistent DOM creation pattern
**`typeof app !== 'undefined'` checks** (20+ instances):
- Lines 458, 467, 481, 614, 617, 1549, etc.
- **Fix**: Ensure `app` is always defined as global singleton.
**Element visibility toggling** (212+ occurrences):
```javascript
element.classList.add('active')
element.classList.remove('active')
```
**Fix**: Create `toggleClass(el, className, condition)` utility.
### Event Listener Issues
- **152 `addEventListener` calls** with fragile cleanup
- **Mix of inline (`onclick="app.method()"`) and addEventListener** - hard to track
- **Element cache (`_elemCache`) never invalidated** if DOM elements are recreated (line 2808)
- **Tab drag-and-drop listeners** may not clean up if user switches tabs mid-drag
---
## 13. Medium: Inconsistent Logging
**Severity**: MEDIUM
**Impact**: Hard to debug in production. Can't filter by severity or component.
### Current State
- **345 console calls** across source files
- **No structured logging** - all `console.log/error` directly
- **No log levels** (DEBUG, INFO, WARN, ERROR)
### Inconsistent Prefixes
```typescript
// Some files use brackets:
console.log('[Session] Starting interactive...');
console.log('[RalphLoop] Task assigned...');
console.log('[TunnelManager] Tunnel started');
// Others use no prefix:
console.error('Failed to spawn PTY:', err);
console.log('Server listening on port', port);
```
### Positive: CleanupManager Has Debug Mode
`src/utils/cleanup-manager.ts` has a `debugMode` flag for conditional debug logging - good pattern not replicated elsewhere.
### Fix
Either:
1. Enforce consistent `[ComponentName]` prefixes via lint rule
2. Create lightweight logger abstraction (not a heavy framework)
---
## 14. Medium: Utils Barrel Export Gaps
**File**: `src/utils/index.ts`
**Severity**: MEDIUM
**Impact**: Forces deep imports, unclear public API.
### Missing Exports
These functions are defined but NOT exported from the barrel:
- `createAnsiPatternFull()` and `createAnsiPatternSimple()` (factory functions from `regex-patterns.ts`)
- `SAFE_PATH_PATTERN` (from `regex-patterns.ts`)
- `validateTokenCounts()` and `validateTokensAndCost()` (from `token-validation.ts`)
- `isSimilar()`, `isSimilarByDistance()`, `levenshteinDistance()`, `normalizePhrase()` (from `string-similarity.ts` - though some are dead code, see finding #16)
### Deep Import Anti-Pattern (16 instances)
Some files bypass the barrel unnecessarily:
```typescript
// Could use barrel:
import { BufferAccumulator } from './utils/buffer-accumulator.js';
import { LRUMap } from './utils/lru-map.js';
// Must deep import (not in barrel):
import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';
```
### Fix
Add missing exports to `src/utils/index.ts` and update import sites.
---
## 15. Medium: Non-Null Assertion Risks
**Severity**: MEDIUM
**Impact**: Runtime crashes if assumptions violated. 37 instances found.
### Distribution
| File | Count | Risk Level |
|------|-------|------------|
| `src/web/server.ts` | 10 | Low (auth flow verified) |
| `src/session.ts` | 6 | **High** (mux/terminal refs) |
| `src/respawn-controller.ts` | 4 | Low (config validated) |
| `src/lru-map.ts` | 3 | Low (checked lookups) |
| `src/subagent-watcher.ts` | 2 | Low (pending tool calls) |
| Others | 12 | Low |
### High-Risk Examples (session.ts)
```typescript
// Line 915 - _mux could be null if startInteractive called during cleanup
`[Session] Starting interactive (with ${this._mux!.backend})`
// Line 954 - _muxSession could be null in race condition
this._muxSession!.muxName
```
### Fix
Add null guards before assertions, or document invariants:
```typescript
// Before:
this._mux!.backend
// After:
if (!this._mux) throw new Error('Invariant: _mux must be initialized before startInteractive');
this._mux.backend
```
### Positive Notes
- **0 instances of `as any`**
- **0 instances of `@ts-ignore` or `@ts-expect-error`**
- TypeScript overall score: 8.5/10
---
## 16. Low: Dead Utility Functions
**File**: `src/utils/string-similarity.ts`
**Severity**: LOW
**Impact**: Code clutter, confusion about what's actually used.
### Unused Functions
These are defined and exported but **never imported anywhere**:
- `isSimilar(a, b, threshold)` - similarity check with threshold
- `isSimilarByDistance(a, b, maxDistance)` - Levenshtein-based check
- `levenshteinDistance(a, b)` - raw edit distance
- `normalizePhrase(phrase)` - phrase normalization
### Actually Used
Only these are imported from the barrel:
- `stringSimilarity()` - used in ralph-tracker.ts
- `fuzzyPhraseMatch()` - used in ralph-tracker.ts
- `todoContentHash()` - used in ralph-tracker.ts
### Fix
Delete unused functions or mark as `@internal` if kept for future use.
---
## 17. Low: No Dependency Injection for File I/O
**Severity**: LOW (practical impact limited at current scale)
**Impact**: Can't mock filesystem for unit tests. 68+ hard-coded filesystem calls.
### Examples
```typescript
// state-store.ts - directly imports and uses fs
import { readFileSync, writeFileSync, existsSync, mkdirSync } from 'node:fs';
// push-store.ts - hard-coded paths
const KEYS_FILE = join(DATA_DIR, 'push-keys.json');
const SUBS_FILE = join(DATA_DIR, 'push-subscriptions.json');
// ai-checker-base.ts - direct execSync
execSync(`tmux kill-session -t "${this.checkMuxName}"`, { timeout: 3000 });
```
### Why This Is Lower Priority
- The codebase uses integration tests (spawning real processes/tmux sessions) rather than unit tests
- Most filesystem operations are in infrastructure code, not business logic
- Adding DI would be a large refactor with limited near-term benefit
---
## 18. Scorecard & Prioritized Roadmap
### Overall Scores (Post-Implementation)
| Category | Before | After | Notes |
|----------|--------|-------|-------|
| TypeScript Safety | 8.5/10 | 9/10 | 0 `any`, 0 `@ts-ignore`, Zod `z.infer` eliminates type drift |
| Error Handling | 8/10 | 8/10 | Unchanged — already strong |
| Async/Promise Safety | 9.5/10 | 9.5/10 | Unchanged — already strong |
| Resource Cleanup | 7/10 | 8/10 | CleanupManager adopted in server.ts, subagent-watcher, bash-tool-parser; Debouncer in 6 files. **Gaps**: respawn-controller (10+ manual timers) and ralph-tracker (2 manual timers) not migrated |
| Module Organization | 5/10 | 8/10 | Routes extracted (12 modules), types split (14 domain files), domain files split (ralph: 7, respawn: 5, session: 6) |
| Test Coverage | 6/10 | 7.5/10 | Shared mock infrastructure, 12 route test files, MockSession/MockStateStore consolidated |
| Config Centralization | 6/10 | 9/10 | 9 config files, ~65 constants centralized, 0 cross-file duplicates |
| Frontend Architecture | 4/10 | 7/10 | 8 extracted modules (3,453 LOC), app.js reduced 24% (15.2K → 11.5K), xterm-zerolag-input vendor build |
| Code Duplication | 5/10 | 8/10 | Debouncer utility, shared test mocks, barrel exports, config consolidation |
### Implementation Phases
**Phase 1 - Quick Wins (1-2 days)** ✅ COMPLETE
1. ✅ Export missing functions from utils barrel (~30 min) — `createAnsiPatternFull`, `createAnsiPatternSimple`, `SAFE_PATH_PATTERN`, `validateTokenCounts`, `validateTokensAndCost` all now exported from `src/utils/index.ts`
2. ✅ Delete dead utility functions (~15 min) — `isSimilar()` removed from `string-similarity.ts`; `levenshteinDistance()`, `isSimilarByDistance()`, `normalizePhrase()` made private (used internally by `fuzzyPhraseMatch`/`stringSimilarity`)
3. ✅ Consolidate duplicated `EXEC_TIMEOUT_MS` constant (~15 min) — Created `src/config/exec-timeout.ts` as single source of truth; `claude-cli-resolver.ts`, `opencode-cli-resolver.ts`, and `tmux-manager.ts` all import from it
4. ✅ Add `z.infer` to Zod schemas (~2 hours) — `src/web/schemas.ts` now has 36 `z.infer` type exports (lines 512-547) covering all schemas
5. ✅ Fix 10 weak "not.toThrow()" tests (~1 hour) — All `not.toThrow()` calls now have behavior assertions: `task-tracker.test.ts` (6 instances all followed by state checks), `image-watcher.test.ts` (1 instance followed by length check), `session-manager.test.ts` (1 instance followed by count check)
**Phase 2 - CleanupManager & Debounce (2-3 days)** ✅ COMPLETE
1. ✅ Create `Debouncer` utility class (~1 hour) — Created `src/utils/debouncer.ts` with `Debouncer` and `KeyedDebouncer` classes; exported from `src/utils/index.ts`
2. ✅ Migrate all 8 files from manual debounce to Debouncer — `state-store.ts` (2 Debouncers), `push-store.ts` (1 Debouncer), `bash-tool-parser.ts` (1 Debouncer), `image-watcher.ts` (1 KeyedDebouncer), `subagent-watcher.ts` (2 KeyedDebouncers), `server.ts` (1 KeyedDebouncer for persist timers), `ralph-tracker.ts` (2 Debouncers replacing 4 manual fields: `_todoUpdateTimer`, `_loopUpdateTimer`, `_todoUpdatePending`, `_loopUpdatePending`)
3. ✅ Migrate respawn-controller to CleanupManager — 10 manual timer fields replaced with single `CleanupManager` instance + `timerIds` Map. `startTrackedTimer()`/`cancelTrackedTimer()` preserved as wrappers for UI countdown display and timer events. `clearTimers()` uses dispose-and-recreate pattern for state transitions.
4. ✅ Migrate server.ts timer cleanup to CleanupManager (~2 hours) — `private cleanup = new CleanupManager()` present; terminal batch timers and pending respawn starts left as manual Maps (complex lifecycle)
5. ✅ Migrate remaining files — `bash-tool-parser.ts` (CleanupManager ✅), `subagent-watcher.ts` (CleanupManager ✅), `ralph-tracker.ts` (Debouncer ✅)
**Phase 3 - server.ts Route Extraction (3-4 days)** ✅ COMPLETE
1. ✅ Created `src/web/routes/` with 12 domain route modules + index barrel (4,090 LOC total): session (909), system (768), ralph (533), plan (459), respawn (315), case, file, hook-event, mux, push, scheduled, team
2. ✅ Created `src/web/middleware/auth.ts` (193 LOC) — Basic Auth, session cookies, rate limiting, security headers, CORS
3. ✅ Created `src/web/ports/` with 7 typed port interfaces (142 LOC) — SessionPort, EventPort, RespawnPort, ConfigPort, InfraPort, AuthPort; routes declare dependencies via intersection types
4. ✅ Created `src/web/route-helpers.ts` (154 LOC) — `findSessionOrFail()`, `formatUptime()`, `sanitizeHookData()`, `autoConfigureRalph()`
5. ✅ Reduced `server.ts` from 6,736 → 2,697 LOC (60% reduction). Remaining LOC is justified infrastructure: session lifecycle, SSE broadcast engine, terminal batching, respawn integration, resource cleanup
**Phase 4 - Domain File Splitting (2-3 days)** ✅ COMPLETE
1. ✅ Split `types.ts` into `src/types/` directory — 14 domain files (1,469 LOC total): common, session, task, app-state, respawn, ralph, api, lifecycle, run-summary, tools, teams, push, plan + index barrel. Original `types.ts` is now a 1-line re-export
2. ✅ Split `ralph-tracker.ts` into 7 files (exceeded plan of 4) — ralph-tracker (2,391), ralph-plan-tracker (477), ralph-status-parser (552), ralph-fix-plan-watcher (366), ralph-stall-detector (166), ralph-config (153), ralph-loop (522)
3. ✅ Split `respawn-controller.ts` into 5 files (exceeded plan of 3) — respawn-controller (3,228), respawn-health (229), respawn-metrics (229), respawn-patterns (131), respawn-adaptive-timing (134)
4. ✅ Split `session.ts` into 6 files (exceeded plan of 3) — session (2,168), session-manager (298), session-auto-ops (284), session-cli-builder (132), session-task-cache (101), session-lifecycle-log (114)
**Phase 5 - Frontend Modularization (3-4 days)** ✅ COMPLETE
1. ✅ Extracted `constants.js` (238 LOC) — shared constants, timing values, Z-index layers, `escapeHtml()`, `extractSyncSegments()`
2. ✅ Extracted `mobile-handlers.js` (449 LOC) — `MobileDetection`, `KeyboardHandler`, `SwipeHandler`
3. ✅ Extracted `voice-input.js` (853 LOC) — `DeepgramProvider`, `VoiceInput`
4. ✅ Extracted `notification-manager.js` (445 LOC) — `NotificationManager` class (5-layer system)
5. ✅ Extracted `keyboard-accessory.js` (279 LOC) — `KeyboardAccessoryBar`, `FocusTrap`
6. ✅ Extracted `api-client.js` (70 LOC) — `_api()`, `_apiJson()`, `_apiPost()`, `_apiPut()`
7. ✅ Extracted `subagent-windows.js` (1,119 LOC) — 13 subagent window methods
8. ✅ Removed inlined xterm-zerolag-input copy → built to `vendor/xterm-zerolag-input.js` from `packages/xterm-zerolag-input/`
9. ✅ Reduced `app.js` from ~15,200 → 11,473 LOC (24% reduction). All scripts loaded in correct dependency order in `index.html`
**Phase 6 - Config Consolidation (1 day)** ✅ COMPLETE
1. ✅ Created 6 new domain-focused config files (better than plan's 2 generic files): `server-timing.ts` (13 constants), `auth-config.ts` (5 constants), `tunnel-config.ts` (8 constants), `terminal-limits.ts` (4 constants), `ai-defaults.ts` (3 constants), `team-config.ts` (3 constants)
2. ✅ Total: 9 config files in `src/config/`, ~65 constants centralized
3. ✅ Eliminated all cross-file duplicates: `STATS_COLLECTION_INTERVAL_MS` (was in 2 files), `timeout: 10000` (was 6× inline in hooks-config.ts → `HOOK_TIMEOUT_MS`), AI model string (was in 5 files → `AI_CHECK_MODEL`), `MAX_TRACKED_AGENTS` (was shadowed in subagent-watcher.ts)
4. ✅ CLAUDE.md updated with config files table, import conventions, resource limits references
**Phase 7 - Test Infrastructure (2-3 days)** ✅ COMPLETE
1. ✅ Created `test/mocks/` directory with 5 files (541 LOC): `mock-session.ts` (312), `mock-state-store.ts` (60), `mock-route-context.ts` (121), `test-helpers.ts` (37), `index.ts` (11 — barrel export)
2. ✅ Consolidated MockSession into single shared definition — no duplicate class definitions remain (2 `vi.mock()`-based copies intentionally left in session-manager.test.ts and ralph-loop.test.ts)
3. ✅ `respawn-test-utils.ts` converted to backward-compatibility shim — re-exports from `test/mocks/`, retains respawn-specific utilities (MockAiIdleChecker, TimeController, etc.)
4. ✅ Created initial 3 route test files with 58 total tests: `session-routes.test.ts` (34 tests), `respawn-routes.test.ts` (13 tests), `system-routes.test.ts` (11 tests). Route test harness uses `app.inject()` — no real ports needed
5. ✅ All 12 route modules now have dedicated test files in `test/routes/`: session, respawn, system, ralph, plan, push, team, mux, file, scheduled, hook-event, case
---
## Appendix: File Size Inventory (Post-Implementation)
### Before vs After
| File | Before | After | Change |
|------|--------|-------|--------|
| `src/web/server.ts` | 6,736 | 2,697 | **−60%** (routes, auth, ports extracted) |
| `src/web/public/app.js` | 15,196 | 11,473 | **−24%** (8 modules extracted) |
| `src/ralph-tracker.ts` | 3,905 | 2,391 | **−39%** (6 companion files extracted) |
| `src/respawn-controller.ts` | 3,611 | 3,228 | **−11%** (4 companion files extracted) |
| `src/session.ts` | 2,418 | 2,168 | **−10%** (5 companion files extracted) |
| `src/types.ts` | 1,443 | 1 | **−99%** (14 domain files in `src/types/`) |
### New Infrastructure Created
| Directory | Files | Total LOC | Purpose |
|-----------|-------|-----------|---------|
| `src/web/routes/` | 13 | 4,090 | Domain route modules |
| `src/web/ports/` | 7 | 142 | Port interfaces for DI |
| `src/web/middleware/` | 1 | 193 | Auth middleware |
| `src/types/` | 14 | 1,469 | Domain type files |
| `src/config/` | 9 | ~450 | Centralized config |
| `test/mocks/` | 5 | 541 | Shared test mocks |
| `test/routes/` | 4 | ~500 | Route handler tests |
### Extracted Frontend Modules
| Module | Lines | Purpose |
|--------|-------|---------|
| `subagent-windows.js` | 1,119 | Subagent window management |
| `voice-input.js` | 853 | DeepgramProvider, VoiceInput |
| `mobile-handlers.js` | 449 | MobileDetection, KeyboardHandler, SwipeHandler |
| `notification-manager.js` | 445 | 5-layer notification system |
| `keyboard-accessory.js` | 279 | KeyboardAccessoryBar, FocusTrap |
| `constants.js` | 238 | Shared constants, timing, Z-index |
| `api-client.js` | 70 | API fetch wrapper |
### What's Working Well
These patterns should be **preserved, not refactored**:
- Clean one-way dependency graph (no circular deps)
- EventEmitter-based decoupling between domain models
- Proper `import type` usage (19 files, consistent)
- Utility type adoption (101 instances of Record, Partial, Omit, etc.)
- `assertNever()` for exhaustive switch checking
- `StaleExpirationMap` and `LRUMap` for bounded collections
- State persistence circuit breaker pattern
- TypeScript strict mode with all safety flags enabled
- `CleanupManager` for centralized timer/watcher disposal
- `Debouncer`/`KeyedDebouncer` for consistent debounce patterns
- Port interfaces for route module dependency injection
- `Object.assign(CodemanApp.prototype, ...)` for frontend module composition
-423
View File
@@ -1,423 +0,0 @@
# Performance Analysis & Optimization Opportunities
**Date**: 2026-03-07
**Scope**: Full-stack performance analysis — backend PTY handling, SSE broadcasting, frontend terminal rendering, local echo overlay, DOM updates, config/scaling limits.
**Constraint**: All recommendations preserve existing functionality including local echo, backpressure, anti-flicker pipeline, and mobile support.
---
## Executive Summary
The codebase is already well-optimized in critical paths. The multi-layer backpressure system, adaptive terminal batching, DEC 2026 sync markers, and incremental state serialization are strong. The main opportunities are in **reducing unnecessary work** (SSE filtering, DOM rebuilds, lazy terminal init) rather than algorithmic changes.
**Top 5 high-impact opportunities:**
| # | Optimization | Impact | Risk | Effort |
|---|-------------|--------|------|--------|
| 1 | Session-scoped SSE subscriptions | Bandwidth -60-80%, CPU -40% | Medium | Medium |
| 2 | Lazy xterm.js for minimized subagent windows | Memory -3.5MB at 50 agents | Low | Low |
| 3 | Targeted badge update (skip full tab rebuild) | Eliminates O(n) reflow on badge change | Low | Low |
| 4 | Conditional SSE padding (tunnel-only, terminal-only) | Bandwidth -70% when tunneled | Low | Low |
| 5 | Canvas renderer on mobile | GPU pressure reduction, battery savings | Low | Low |
---
## 1. SSE Broadcasting
### Current State
- **92 event types** broadcast to all connected clients (max 100)
- Single `JSON.stringify()` per event, shared across all clients (efficient)
- **No per-client filtering** — every client receives every event regardless of which session they're viewing
- 8KB padding appended to **every** event when tunnel is active (forces Cloudflare proxy flush)
- Backpressure: clients marked as backpressured if `reply.raw.write()` returns false; recovery via `session:needsRefresh`
### Bottlenecks
**B1: No session-scoped SSE subscriptions** (`server.ts:1986`)
- Client viewing session A still receives all events for sessions B through T
- With 20 active sessions, ~95% of terminal events are irrelevant to any given client
- Cost: wasted bandwidth, CPU for JSON parsing, and event handler dispatch on client
**B2: Unconditional 8KB padding** (`server.ts:1977`)
- Every event gets 8KB comment padding when tunnel is active
- A `task:updated` event (~200 bytes payload) becomes ~8.2KB
- High-frequency events like `session:terminal` need the padding; low-frequency events like `session:created` don't
### Recommendations
**R1: Session-scoped SSE subscriptions** (High impact)
- Add `?sessions=id1,id2` query param to `/api/events` SSE endpoint
- Server filters events by session ID before broadcasting
- Client subscribes to active session + "global" events (session lifecycle, system)
- Re-subscribes on tab switch (or subscribe to all with client-side filter as fallback)
- **Savings**: ~80% bandwidth reduction for single-session viewers; ~60% for multi-session dashboards
**R2: Tiered SSE padding** (Medium impact)
- Only pad `session:terminal` events and SSE heartbeats (the two that need proxy flush)
- Skip padding for low-frequency structural events (`session:created`, `task:updated`, etc.)
- **Savings**: ~70% padding overhead reduction; terminal events already large enough to flush
---
## 2. Terminal Rendering
### Current State (Well-Optimized)
- **6-layer anti-flicker pipeline**: Server batching (adaptive 16-50ms) → DEC 2026 sync wrap → single JSON serialize → client rAF batching → sync segment parser → chunked buffer loading (32KB/frame)
- **64KB/frame write budget** with DEC 2026 sync-segment awareness (prevents 141KB single-frame freezes)
- **3-layer backpressure**: SSE cap (128KB queued → drop + refresh), frame budget (64KB/frame), chunked restore (32KB/frame)
- WebGL renderer enabled by default with canvas fallback on context loss
- Typical latency: 16-32ms; worst case: ~115ms (50ms server batch + 50ms sync wait + 16ms rAF)
### Bottlenecks
**B3: WebGL on mobile** (`app.js:627-637`)
- Mobile GPUs are weaker; WebGL context loss more likely on low-end devices
- Canvas renderer is sufficient for mobile (typically 1 session, smaller viewport)
**B4: Static scrollback for all sessions** (`app.js:572`)
- Default 5000 lines scrollback for all sessions regardless of activity level
- Heavy output sessions (build logs, test runners) accumulate large scroll buffers
**B5: No addon lazy loading**
- FitAddon, Unicode11Addon, and WebGLAddon all loaded at terminal init
- Unicode11Addon only needed for CJK content; WebGLAddon is large
### Recommendations
**R3: Force canvas renderer on mobile** (Low risk)
- Detect `MobileDetection.isMobile()` and skip WebGL addon loading
- Reduces GPU memory pressure, prevents context loss crashes
- Mobile typically has 1-2 sessions — canvas performance is more than adequate
**R4: Dynamic scrollback based on session activity** (Low risk)
- Active sessions (working state): 5000 lines (current default)
- Inactive/idle sessions: reduce to 2000 lines
- Restore on session select (fetch from server buffer)
- **Savings**: ~60% scrollback memory for idle sessions
**R5: Lazy-load Unicode11Addon** (Low risk)
- Only load when CJK content is detected in terminal output
- Detection: check for characters in CJK Unicode ranges during ANSI stripping (already iterating)
- Most sessions never need it
---
## 3. DOM & Session Tab Rendering
### Current State
- Session tabs use **intelligent incremental updates** with debounced 100ms rendering
- Incremental path: only updates changed properties (classes, textContent, badges) when session list is stable
- Full rebuild path: triggered when sessions added/removed **or badge count changes**
- Subagent windows: per-window xterm.js instances, even when minimized
### Bottlenecks
**B6: Badge count change triggers full tab rebuild** (`app.js:3207-3209`)
- A single subagent badge increment on one tab triggers `_fullRenderSessionTabs()` — rebuilds entire sidebar HTML via `innerHTML =`
- With 20 sessions, this is an O(n) reflow for a single badge number change
- Badge changes are frequent during active subagent work
**B7: Minimized subagent windows retain xterm.js instances** (`subagent-windows.js`)
- 50 subagent windows × ~75KB per xterm.js instance = ~3.75MB DOM memory
- Minimized windows are invisible but their terminals remain in DOM
- xterm.js instances continue processing resize events even when hidden
**B8: `backdrop-filter: blur()` on overlays** (`styles.css:2246-2247, 3098`)
- Forces new stacking context, disables browser compositing optimizations
- 50-100ms layout thrashing on modal open/close
- Only 2 uses, but they're on frequently toggled overlays
### Recommendations
**R6: Targeted badge update without full rebuild** (Low risk)
- When badge count changes but session list is stable, update only the badge `<span>` textContent
- Keep incremental path for badge changes; only use full rebuild for structural changes (add/remove sessions)
- **Savings**: Eliminates O(n) reflow per badge change; reduces to O(1) targeted update
**R7: Lazy xterm.js initialization for subagent windows** (Medium impact)
- Only create xterm.js Terminal instance when window is restored/maximized
- On minimize: serialize terminal buffer, dispose Terminal instance, keep buffer in memory
- On restore: create new Terminal, write buffer back
- **Savings**: ~3.5MB DOM reduction at 50 minimized agents; eliminates hidden resize processing
- **Trade-off**: ~200-500ms restore delay (buffer write), mitigated by chunked loading
**R8: Replace `backdrop-filter: blur()` with `background: rgba()`** (Low risk)
- Use semi-transparent background instead of blur effect
- Or use `will-change: transform` hint if blur is kept
- **Savings**: Eliminates forced recomposition layer; 50-100ms faster overlay open
---
## 4. Backend PTY & State Management
### Current State (Excellent)
- **BufferAccumulator**: Array-based chunking with lazy join on read — avoids O(n) string concatenation
- **ANSI stripping**: Throttled at 150ms intervals with lazy evaluation (not per-chunk)
- **State persistence**: 500ms debounce + incremental JSON caching per session (only dirty sessions re-serialized)
- **Expensive parsers**: Throttled to 150ms window, accumulated data capped at 64KB
- **Memory**: All buffers have hard limits (2MB terminal, 1MB text, 1000 messages, 64KB line buffer)
### Bottlenecks
**B9: Pending clean data cap at 64KB** (`session.ts:1097-1133`)
- Between 150ms processing windows, raw PTY data accumulates in `_pendingCleanData`
- Capped at 64KB — excess data rolls off (old data discarded)
- During heavy output (large build logs), this means parsers may miss content
- Acceptable trade-off for performance, but worth documenting
**B10: `LRUMap.delete()` is O(n) worst case** (`utils/lru-map.ts:137-138`)
- When deleting the newest entry, iterates all keys to find new newest
- Rare in practice (delete is uncommon; set/get are hot paths)
- Could matter during mass cleanup of 500 agents
### Recommendations
**R9: Consider adaptive pending data cap** (Low priority)
- During idle detection (critical to get right), increase cap to 128KB
- During active working state, keep at 64KB (parsers less critical)
- **Benefit**: More accurate idle detection during heavy output
**R10: Track second-newest in LRUMap** (Low priority)
- Maintain a `_secondNewestKey` alongside `_newestKey`
- On delete of newest, promote second-newest without iteration
- Only matters at scale (500+ agents with frequent eviction)
---
## 5. Local Echo & Input Path
### Current State (Well-Designed)
- **DOM overlay approach** — `<span>` elements in `.xterm-screen` at z-index 7, completely independent of `terminal.write()`
- **Render caching**: `_lastRenderKey` includes text, position, column offsets — skips redundant re-renders
- **Input flow**: Char accumulation → Enter triggers flush → 80ms delay before `\r` (ensures text reaches PTY first)
- **Tab completion**: Baseline snapshot → detect buffer change → 300ms fallback timer
- **CJK support**: Per-character width detection with `terminal.unicode.getStringCellWidth()` preferred, manual fallback
- **Prompt detection**: Bottom-up line scan, O(rows) — cached position, column-lock prevents jitter
### Bottlenecks
**B11: tmux send-keys latency** (~50-100ms per input)
- Each `writeViaMux()` spawns a child process (`tmux send-keys`)
- Text and Enter sent separately with 50ms delay between
- For rapid typing: characters batch before Enter, so overhead is per-command not per-keystroke
- **Acceptable trade-off** for session persistence (tmux survives server restarts)
**B12: 80ms delay between text flush and Enter** (`app.js:872-875`)
- Intentional: ensures text reaches PTY before Enter, preventing Ink from processing empty input
- Adds 80ms to perceived Enter-to-response latency
- Could potentially be reduced with acknowledgment-based approach
**B13: Scroll listener on terminal viewport** (`zerolag-input-addon.ts:139`)
- 50ms debounced re-render on scroll — acceptable but fires frequently during heavy output
- Overlay hidden when scrolled up (correct behavior), shown when at bottom
### Recommendations
**R11: Reduce Enter delay from 80ms to 50ms** (Low risk, test carefully)
- The tmux `send-keys` already has 50ms internal delay
- Combined with network latency, 80ms client-side may be excessive
- Test with Ink-heavy sessions (Claude Code's status bar) — if text arrives before Enter at 50ms, reduce
- **Savings**: 30ms perceived latency reduction per command
**R12: Batch tmux send-keys via stdin pipe** (Medium effort, high impact for rapid input)
- Instead of spawning `tmux send-keys` per input, maintain a persistent connection
- Use `tmux -C` (control mode) for programmatic interaction without child process spawning
- **Savings**: Eliminate ~50-100ms process spawn overhead per input
- **Risk**: Control mode has different semantics; needs careful testing with session persistence
**R13: Skip overlay re-render during heavy output scroll** (Low risk)
- When terminal is receiving >10KB/s output, hide overlay entirely (user isn't typing during heavy output)
- Re-show overlay after 500ms of output silence
- **Savings**: Eliminates unnecessary DOM overlay re-renders during build logs / test output
---
## 6. Polling & File Watchers
### Current State
- **SubagentWatcher**: 1s base poll, full scan throttled to every 5s, fs.watch() on known directories
- **TranscriptWatcher**: 1 per session, fs.watch() primary with 1s poll fallback
- **ImageWatcher**: chokidar per session with 100ms stability poll, burst limit 20/10s
- **TeamWatcher**: chokidar primary with 30s poll fallback, LRU caches (50 teams, 200 tasks)
- **RalphTracker**: Todo cleanup every 5 minutes
### Scaling Profile (20 sessions)
| Component | Instances | Frequency | Total ops/sec |
|-----------|-----------|-----------|---------------|
| SubagentWatcher | 1 (global) | Full scan every 5s | 0.2/s |
| TranscriptWatcher | 20 | 1s poll (fallback) | 20/s max |
| ImageWatcher | 20 | 100ms poll (during writes only) | 200/s burst |
| TeamWatcher | 1 (global) | 30s poll (fallback) | 0.03/s |
| SSE heartbeat | 1 (global) | 15s | 0.07/s |
| SSE dead client check | 1 (global) | 30s | 0.03/s |
| Mux stats collection | 1 (global) | 2s | 0.5/s |
| **Total steady-state** | | | **~21/s** |
### Recommendations
**R14: Increase TranscriptWatcher poll interval to 2s** (Low risk)
- Transcript changes are infrequent (new messages every few seconds at most)
- fs.watch() is the primary mechanism; polling is fallback
- **Savings**: Halves fallback filesystem checks (20/s → 10/s for 20 sessions)
**R15: Share chokidar instances for co-located session directories** (Medium effort)
- Sessions in the same parent directory could share a single chokidar watcher with depth:3
- Common case: multiple sessions in `~/projects/foo/` — one watcher covers all
- **Savings**: Reduce chokidar instances from 20 to ~5-10 for typical workloads
---
## 7. Frontend Asset Delivery
### Current State
- **app.js**: 12,027 lines (source) → esbuild minified → gzip/brotli compressed (~30-40KB gzipped)
- **Static caching**: `maxAge: '1y'` via `@fastify/static`
- **Service worker**: Push notification handler only — no asset caching
- **No code splitting**: Single monolithic app.js bundle
### Bottlenecks
**B14: No cache-busting mechanism**
- `maxAge: '1y'` means browsers cache aggressively
- After deployment, users need `Ctrl+Shift+R` to see updates
- No content hash in filenames or ETags for automatic invalidation
**B15: Monolithic app.js**
- All 12K lines loaded on initial page load regardless of which features are used
- Ralph wizard, plan orchestrator UI, team management — all loaded upfront
- Mobile loads the same bundle as desktop
### Recommendations
**R16: Add content hash to asset filenames** (Medium impact)
- Build step: rename `app.js` → `app.[hash].js`
- Generate a manifest or inject hash into HTML template
- Keep `maxAge: '1y'` — cache invalidation happens via filename change
- **Savings**: Eliminates stale cache issues after deployment; removes need for manual hard refresh
**R17: Code-split app.js into core + feature modules** (High effort, medium impact)
- Core (~4K lines): terminal, SSE, session management, tabs, input handling
- Deferred (~8K lines): Ralph wizard, plan UI, team management, subagent windows, image viewer
- Load deferred modules on first use via dynamic `import()` or lazy `<script>` injection
- **Savings**: ~60% reduction in initial load size; faster time-to-interactive
- **Risk**: Complexity increase; need to handle loading states for deferred features
- **Note**: May not be worth the effort given the app is already gzipped to ~30-40KB
---
## 8. CSS Performance
### Current State
- **styles.css**: 7,153 lines with ~45 box-shadow uses, 2 backdrop-filter uses
- Animations: GPU-accelerated keyframes for pulsing alerts, loading spinners
- Z-index layering: well-organized (subagent 1000, plan 1100, log 2000, image 3000, overlay 7)
### Recommendations
**R18: Replace backdrop-filter with opaque overlay** (Low risk, covered in R8)
**R19: Use `contain: content` on subagent windows** (Low risk)
- Add CSS containment to subagent window containers
- Prevents layout changes inside windows from triggering reflow on parent
- Especially valuable with 50 windows: changes in one window won't invalidate others
- ```css
.subagent-window { contain: content; }
```
- **Savings**: Reduces layout recalculation scope from global to per-window
**R20: Use `content-visibility: auto` on off-screen subagent windows** (Low risk)
- Browser skips rendering of off-screen windows entirely
- Combined with `contain-intrinsic-size` to prevent layout shift
- ```css
.subagent-window.minimized { content-visibility: hidden; }
```
- **Savings**: Browser skips paint/layout for minimized windows; complements R7
---
## 9. Memory & Scaling Limits
### Current Budget (20 sessions)
| Component | Per Session | Total | Status |
|-----------|-----------|-------|--------|
| Terminal buffer | 2MB | 40MB | Hard-limited, auto-trim |
| Text output | 1MB | 20MB | Hard-limited, auto-trim |
| Messages | ~1MB | 20MB | Capped at 1000, trims to 800 |
| Respawn buffer | 1MB | 20MB | Hard-limited |
| **Buffers total** | | **100MB** | Acceptable |
| TranscriptWatcher | ~100KB | 2MB | |
| ImageWatcher | ~50KB | 1MB | |
| SubagentWatcher | ~500KB | 500KB | Global |
| Frontend terminal cache | ~256KB | 5MB | LRU, max 20 entries |
| **Total estimated** | | **~110MB** | Comfortable |
### At Max Scale (50 sessions)
- Buffers: ~250MB
- Watchers: ~5MB
- **Total: ~255MB** + Node.js overhead — acceptable on modern hardware
### Potential Leak Vectors (All Mitigated)
- `_shortIdCache` in server — unbounded Map, but entries are tiny (string→string); grows at O(sessions created), not O(events)
- All CleanupManager-registered resources tracked and disposed on session stop
- `isStopped` guard prevents new timers after session cleanup
---
## 10. Implementation Priority Matrix
### Phase 1 — Quick Wins (1-2 hours each, low risk)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R6 | Targeted badge update | `app.js` (3207-3209) |
| R3 | Canvas renderer on mobile | `app.js` (627-637) |
| R8 | Replace backdrop-filter blur | `styles.css` (2246, 3098) |
| R19 | CSS containment on subagent windows | `styles.css` |
| R20 | `content-visibility: hidden` on minimized windows | `styles.css` |
### Phase 2 — Medium Effort (half-day each)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R2 | Tiered SSE padding | `server.ts` (broadcast function) |
| R7 | Lazy xterm.js for minimized subagents | `subagent-windows.js` |
| R11 | Reduce Enter delay to 50ms | `app.js` (872-875), test with Ink |
| R14 | TranscriptWatcher 2s poll | `transcript-watcher.ts` |
| R16 | Content-hash asset filenames | `build.mjs`, `server.ts` |
### Phase 3 — Larger Initiatives (1-2 days each)
| # | Optimization | Files to Change |
|---|-------------|-----------------|
| R1 | Session-scoped SSE subscriptions | `server.ts`, `app.js` (SSE connect) |
| R5 | Lazy Unicode11Addon loading | `app.js`, build pipeline |
| R12 | Persistent tmux control mode | `tmux-manager.ts` |
| R17 | Code-split app.js | `app.js`, `build.mjs`, HTML template |
### Not Recommended (Low ROI or High Risk)
| # | Why Not |
|---|---------|
| R4 | Dynamic scrollback adds complexity; memory savings marginal vs total budget |
| R9 | Adaptive pending data cap adds state; current 64KB cap rarely matters |
| R10 | LRUMap.delete() O(n) is theoretical; never triggered at current scale |
| R15 | Shared chokidar instances add directory-matching complexity for minimal gain |
---
## Appendix: Key File Locations
| Area | File | Key Lines |
|------|------|-----------|
| SSE broadcast | `src/web/server.ts` | 1961-1989 (broadcast), 1934-1959 (backpressure) |
| Terminal batching | `src/web/server.ts` | 1994-2048 (per-session adaptive batching) |
| Frame budget | `src/web/public/app.js` | 1370-1478 (flushPendingWrites, 64KB cap) |
| Flicker filter | `src/web/public/app.js` | 1176-1255 (50ms sync wait, 256KB safety) |
| Tab rendering | `src/web/public/app.js` | 3108-3357 (incremental + full rebuild) |
| Tab switching | `src/web/public/app.js` | 3560-3760 (cache + chunked load + deferred UI) |
| Local echo | `packages/xterm-zerolag-input/src/` | All files (overlay, prompt, CJK) |
| Local echo integration | `src/web/public/app.js` | 640, 815-988 (input flow) |
| Subagent windows | `src/web/public/subagent-windows.js` | Full file (window mgmt, drag, minimize) |
| State persistence | `src/state-store.ts` | 161-250 (debounced save, incremental JSON) |
| Buffer accumulator | `src/utils/buffer-accumulator.ts` | Full file (array chunks, lazy join) |
| PTY handling | `src/session.ts` | 1046-1133 (data flow), 1173-1230 (parsing) |
| Config limits | `src/config/` | 9 files (buffer, map, timing, auth, etc.) |
| Anti-flicker docs | `docs/terminal-anti-flicker.md` | Architecture reference |
| CSS | `src/web/public/styles.css` | 2246 (backdrop-filter), full file |
| Build pipeline | `scripts/build.mjs` | 59-68 (minify + compress) |
@@ -1,168 +0,0 @@
# Performance & Responsiveness Optimization Plan
**Date**: 2026-02-28
**Status**: Phases 1–4 Complete. Phase 5 optional/deferred.
---
## Executive Summary
Three independent research passes analyzed the Codeman codebase for performance bottlenecks across frontend rendering, backend hot paths, and system-level resource usage. The codebase already has strong foundational optimizations (per-session adaptive batching, rAF terminal writes, DEC 2026 sync markers, backpressure handling). This plan targets the remaining high-impact opportunities.
**Key finding**: The biggest wins come from **skipping unnecessary work** — serializing unchanged state, processing output nobody is watching, and reducing broadcast volume.
---
## Phase 1: Quick Wins — COMPLETE
All Phase 1 items were found to already exist in the codebase during verification:
| # | Item | Status | Evidence |
|---|------|--------|----------|
| 1.1 | Skip terminal writes for hidden tabs | Done | SSE handler filters by `activeSessionId` (app.js:4076) |
| 1.2 | mobile.css media query | Done | `media="(max-width: 1023px)"` on link tag (index.html:13) |
| 1.3 | Deduplicate init API calls | Done | `_initGeneration` dedup + 3s fallback timer (app.js:2901-2904) |
| 1.4 | Remove cache-busting timestamps | Done | No `?_t=` patterns found anywhere |
| 1.5 | JS/CSS minification + compression | Done | esbuild minify + gzip + brotli in build.mjs (lines 42-51) |
---
## Phase 2: Frontend Responsiveness — COMPLETE
### 2.1 Batch `getBoundingClientRect()` in connection lines — DONE
- **Files**: `src/web/public/app.js` (`_updateConnectionLinesImmediate()`)
- **Change**: Refactored to batch all layout reads into Phase 1 (collect all rects into a Map), then perform all SVG writes in Phase 2 using cached values. Classic read-then-write pattern prevents interleaved forced reflows.
### 2.2 Clean up ResizeObservers — Already implemented
- `forceCloseSubagentWindow()` disconnects observers (app.js:12618-12620)
- `cleanupAllFloatingWindows()` disconnects all on reconnect (app.js:12649-12653)
- Observer refs stored on `windowData.resizeObserver` (app.js:12492)
### 2.3 Drag handler cleanup — Already implemented
- `makeWindowDraggable()` returns listener refs, stored in `windowData.dragListeners`
- `forceCloseSubagentWindow()` removes all document-level drag listeners (app.js:12622-12630)
- Panel drags add listeners on mousedown, remove on mouseup (app.js:10253-10284)
### 2.4 Mobile window position cache — Skipped
- O(n) loop over max ~20 windows; complexity of cached counter not justified
### 2.5 Lazy modal DOM — Skipped
- Large effort, marginal benefit for a vanilla JS app with fast DOM construction
---
## Phase 3: Backend Hot Paths — COMPLETE
### 3.1 State diff broadcasts — ALREADY OPTIMIZED
- `broadcastSessionStateDebounced()` already batches at 500ms intervals
- `toLightDetailedState()` excludes heavy buffers (textOutput, terminalBuffer)
- Per-session serialization is <1ms; with debouncing, only 1-3 sessions serialize per flush
- JSON.stringify happens once per broadcast (not per client) — serialization cost is negligible
- Full state diffs would add significant frontend complexity for marginal gain
### 3.2 Improve session list cache hit rate — DONE
- **Files**: `src/web/server.ts` (`broadcast()` method)
- **Change**: Cache now only invalidated on truly structural events (`session:created`, `session:deleted`, `session:updated`) instead of on every `session:*` and `respawn:*` event. High-frequency events like `session:working`, `session:idle`, `session:completion`, `respawn:stateChanged` no longer defeat the 1s TTL cache.
- **Impact**: Cache hit ratio from ~0% to ~80%+ during active sessions. The debounced `session:updated` still refreshes the cache within 500ms of any state change.
### 3.3 Skip PTY processing — ALREADY OPTIMIZED
- `_processExpensiveParsers()` is already throttled to every 150ms (not per-chunk)
- Lazy ANSI stripping via `getCleanData()` closure — only computed when a consumer needs it
- Quick pre-checks skip parsers when content is irrelevant (e.g., token parser only runs if data contains "token")
- OpenCode sessions skip all Claude-specific parsers entirely
- Further optimization would require visibility-aware processing, adding complexity for marginal gain
### 3.4 Batch subagent liveness checks — Deferred
- `/proc/{pid}` stat calls are ~0.1ms each; even with 500 agents, total is 50ms every 10s
- Current approach is simple and reliable; batching adds race condition risk
- Consider only if profiling shows this as a bottleneck
### 3.5 Deduplicate detection update emissions — DONE
- **Files**: `src/respawn-controller.ts` (`startDetectionUpdates()`)
- **Change**: Detection status now only emitted when key fields (confidenceLevel, statusText, controller state) actually change. Previously emitted every 2s regardless, broadcasting identical status to all SSE clients.
- **Impact**: For stable/idle sessions, eliminates ~100% of redundant detection broadcasts. For active sessions, reduces broadcasts to only meaningful state transitions.
---
## Phase 4: System-Level Improvements — COMPLETE
### 4.1 Incremental state persistence — DONE
- **Files**: `src/state-store.ts` (`assembleStateJson()`, `setSession()`)
- **Change**: Added `dirtySessions` Set and `cachedSessionJsons` Map. On persist, only dirty sessions are re-serialized; clean sessions reuse cached JSON fragments. `setSession()` marks sessions dirty; `assembleStateJson()` rebuilds only changed fragments.
- **Impact**: Serialization cost reduced from O(all sessions) to O(dirty sessions). Typical steady-state: 1-2 dirty sessions instead of 50.
### 4.2 Replace polling with fs watchers for team watcher — DONE
- **Files**: `src/team-watcher.ts` (`setupFsWatchers()`)
- **Change**: Added chokidar watchers on both `~/.claude/teams/` and `~/.claude/tasks/` directories for instant event-driven detection. Lock files ignored via chokidar config. Mtime-based dedup skips unchanged files. Polling interval relaxed from 5s to 30s as a fallback.
- **Impact**: Near-instant team detection; polling overhead eliminated for normal operation.
### 4.3 Consolidate subagent file watchers — DONE
- **Files**: `src/subagent-watcher.ts` (`setupDirectoryWatcher()`)
- **Change**: Replaced per-agent chokidar watchers with one `fs.watch()` per session subagent directory. Events are routed to the correct agent via filename. Per-file debouncing (100ms) prevents hammering on bulk discovery.
- **Impact**: Inotify watchers reduced from potentially 500 (one per agent) to ~50 (one per session directory).
### 4.4 Stream transcript files instead of full reads — DONE
- **Files**: `src/subagent-watcher.ts` (`tailFile()`, `findDescriptionInAgentFile()`, parent transcript lookup)
- **Change**: Multiple streaming strategies implemented:
- **Live monitoring**: Position-based `tailFile()` with `createReadStream({ start: fromPosition })` — only reads new content
- **Parent transcript lookup**: Streams only last 16KB (`createReadStream({ start: offset })`)
- **Description extraction**: Streams only first 8KB, exits early after 5 lines
- **Full read**: Only for on-demand transcript review panel (with optional `limit` parameter)
- **Impact**: File I/O for bulk agent discovery reduced from ~50MB to ~5MB.
---
## Phase 5: Long-Term Architectural (Optional) — NOT STARTED
These items are deferred until scaling demands justify the complexity.
### 5.1 Worker thread for PTY processing
- **Files**: `src/session.ts`
- **Problem**: ANSI stripping, Ralph tracking, and bash tool parsing all run on the main event loop. At scale (50 busy sessions), this consumes 300-500ms CPU/sec.
- **Fix**: Offload ANSI strip + line processing to a worker thread pool. Main thread receives clean text + parsed events.
- **Impact**: Frees event loop for I/O operations. Most impactful at 10+ concurrent busy sessions.
### 5.2 Per-session SSE subscriptions
- **Files**: `src/web/server.ts`
- **Problem**: Every SSE event is broadcast to all connected clients. A client watching session A still receives events for sessions B through Z.
- **Fix**: Clients subscribe to specific session IDs. Server only sends events to interested clients.
- **Impact**: Reduces SSE broadcast fan-out from N clients to ~1-2 per event. Major improvement at 100 SSE clients.
### 5.3 O(1) LRUMap via doubly-linked list
- **Files**: `src/utils/lru-map.ts` (~lines 98-110)
- **Problem**: `get()` uses delete + re-insert to refresh position — O(n) on Map iteration for delete.
- **Fix**: Implement classic LRU with doubly-linked list + Map for O(1) get/put/evict.
- **Impact**: Low — current sizes (max 500) make this barely measurable. Only worthwhile if LRUMap is used on hot paths.
---
## Completion Summary
| Phase | Scope | Status | Items |
|-------|-------|--------|-------|
| 1 | Quick Wins | **Complete** | 5/5 (all pre-existing) |
| 2 | Frontend Responsiveness | **Complete** | 3/3 actionable done, 2 skipped |
| 3 | Backend Hot Paths | **Complete** | 4/4 actionable done, 1 deferred |
| 4 | System-Level | **Complete** | 4/4 done |
| 5 | Long-Term Architectural | **Not started** | 0/3 — deferred until needed |
**Overall**: 16/16 actionable items complete. 3 optional items deferred.
---
## Measurement
Before starting Phase 5, establish baselines:
1. **Frontend**: Record Chrome DevTools Performance trace with 10 sessions open. Measure:
- Frame rate during rapid terminal output
- Long tasks (>50ms) count per 30s
- Heap size after 1h session
2. **Backend**: Add `performance.now()` instrumentation around:
- `flushSessionTerminalBatch()` — time per flush
- `broadcastSessionStateDebounced()` — serialization time
- `StateStore.save()` — persist time
- Event loop lag via `monitorEventLoopDelay()`
3. **First load**: Lighthouse score on desktop and mobile (simulated 3G)
-74
View File
@@ -1,74 +0,0 @@
# Codeman Performance Optimization Plan
## Current State
The backend is **already production-grade** — SSE broadcasting, state persistence, terminal batching, buffer management, and memory patterns are all well-optimized. The biggest gains are on the **frontend delivery** side.
## Implemented Optimizations
### 1. V8 Compile Cache (10-20% faster cold start)
**Files:** `scripts/codeman-web.service`, `package.json`
Node.js re-parses and compiles all JS on every cold start. `NODE_COMPILE_CACHE` caches V8 compiled bytecode to disk, reusing it on subsequent starts.
- Added `Environment=NODE_COMPILE_CACHE=/home/arkon/.codeman/compile-cache` to systemd service
- Added to `npm start` script for non-systemd usage
- Zero code changes, immediate win on every restart
### 2. WebGL Addon Lazy-Loading (244KB saved on mobile, non-blocking on desktop)
**Files:** `src/web/public/index.html`, `src/web/public/app.js`
`xterm-addon-webgl.min.js` (244KB) was loaded eagerly for all users via `<script defer>`, but only used on desktop with WebGL2 support.
- Removed `<script defer>` from `index.html`
- Added dynamic script loading in `app.js` — only downloads on desktop when WebGL is needed
- Mobile users never download the file at all (244KB saved)
- Desktop: loads in parallel with page rendering, addon initializes when ready
- Graceful fallback: canvas renderer used if WebGL unavailable or script fails
### 3. Preload Hints (~50-100ms faster perceived load)
**Files:** `src/web/public/index.html`
Browser discovers `<script defer>` tags only when the parser reaches them at the bottom of `<body>`. By then, the HTML parse has blocked for hundreds of lines.
- Added `<link rel="preload" as="script">` in `<head>` for `vendor/xterm.min.js`, `constants.js`, `app.js`
- Browser starts fetching critical scripts immediately during HTML parse (before reaching `<body>`)
- Zero runtime overhead — just hints for the browser's preload scanner
### 4. Batch Tmux Reconciliation (N subprocess calls → 1)
**Files:** `src/tmux-manager.ts`
`reconcileSessions()` previously called `tmux has-session` + `tmux display-message` per known session, plus `tmux list-sessions` for discovery, plus `tmux display-message` per discovered session. With 20 sessions: 41+ subprocess calls.
- Replaced with single `tmux list-panes -a -F '#{session_name}\t#{pane_pid}'` call
- Builds a Map from the result, then does O(1) lookups for both known and discovered sessions
- Also replaced inner O(n) `isKnown` scan with a Set lookup
- 20 sessions: 41 subprocess calls → 1, with faster lookups
### 5. Asset Hashing / Cache Busting (already implemented)
**Files:** `scripts/build.mjs` (pre-existing)
Content-hash cache busting was already implemented in the build script:
- All app JS/CSS files get content hashes (`app.abc123.js`)
- `index.html` rewritten to reference hashed filenames
- Pre-compressed with gzip + Brotli
- 1-year immutable cache works correctly — new deploys get new filenames
## Already Optimized (No Action Needed)
| Area | Why It's Fine |
|------|---------------|
| **SSE Broadcasting** | Single serialization per broadcast, preformatted frames, backpressure handling, session subscription filtering |
| **State Persistence** | 500ms debounce, incremental per-session JSON caching, async atomic writes, circuit breaker on failures |
| **Terminal Batching** | Adaptive intervals (16-50ms), per-session queues, immediate flush at 32KB, array-based accumulation |
| **Buffer Management** | BufferAccumulator (array-push, lazy join), auto-trim at 2MB/1MB, no string concatenation in hot paths |
| **ANSI Stripping** | Pre-compiled regex via factory functions, single-pass processing |
| **Static File Serving** | @fastify/static with 1-year cache, pre-compressed Brotli/gzip, no-cache for HTML |
| **Memory Management** | CleanupManager, LRUMap, StaleExpirationMap, bounded buffers, explicit listener cleanup |
| **Import Patterns** | Pure ESM, lazy web server import, no circular deps, no dynamic imports in hot paths |
| **Config Loading** | Small constant files, no I/O at import time, specific imports (no barrel) |
@@ -1,788 +0,0 @@
# Phase 4: Domain File Splitting — Implementation Plan
**Date**: 2026-03-01
**Prerequisites**: Phase 1-3 complete (utils cleanup, CleanupManager/Debouncer migration, route extraction)
**Goal**: Split 4 god files into focused modules with barrel exports for transparent migration.
---
## Table of Contents
1. [Split types.ts into types/ directory](#1-split-typests-into-types-directory)
2. [Split ralph-tracker.ts into focused modules](#2-split-ralph-trackerts-into-focused-modules)
3. [Split respawn-controller.ts into focused modules](#3-split-respawn-controllerts-into-focused-modules)
4. [Split session.ts into focused modules](#4-split-sessionts-into-focused-modules)
5. [Execution Order & Dependencies](#5-execution-order--dependencies)
6. [Validation Checklist](#6-validation-checklist)
---
## 1. Split types.ts into types/ directory
**Current**: 1,443 lines, 71 exports, imported by 36 files.
**Risk**: LOW — pure type refactor, no runtime behavior change.
### Target Structure
```
src/types/
├── index.ts (barrel re-export — transparent migration)
├── common.ts (Disposable, BufferConfig, CleanupResourceType, CleanupRegistration)
├── session.ts (SessionStatus, SessionMode, ClaudeMode, SessionConfig, SessionColor,
│ SessionState, OpenCodeConfig, SessionOutput)
├── task.ts (TaskStatus, TaskDefinition, TaskState)
├── app-state.ts (AppState, AppConfig, GlobalStats, TokenUsageEntry, TokenStats,
│ DEFAULT_CONFIG, createInitialState, createInitialGlobalStats)
├── respawn.ts (RespawnConfig, PersistedRespawnConfig, CycleOutcome,
│ RespawnCycleMetrics, RespawnAggregateMetrics, HealthStatus,
│ RalphLoopHealthScore, TimingHistory, RespawnPreset)
├── ralph.ts (RalphLoopStatus, RalphLoopState, RalphTodoStatus, RalphTodoPriority,
│ RalphTodoItem, RalphTodoProgress, RalphSessionState,
│ RalphStatusValue, RalphTestsStatus, RalphWorkType, RalphStatusBlock,
│ CompletionConfidence, RalphTrackerState,
│ CircuitBreakerState, CircuitBreakerReason, CircuitBreakerStatus,
│ createInitialCircuitBreakerStatus, createInitialRalphTrackerState,
│ createInitialRalphSessionState)
├── api.ts (ApiErrorCode, ApiResponse, HookEventType, QuickStartResponse,
│ CaseInfo, createErrorResponse, isError, getErrorMessage)
├── lifecycle.ts (LifecycleEventType, LifecycleEntry)
├── run-summary.ts (RunSummaryEventType, RunSummaryEventSeverity, RunSummaryEvent,
│ RunSummaryStats, RunSummary, createInitialRunSummaryStats)
├── tools.ts (ActiveBashToolStatus, ActiveBashTool, ImageDetectedEvent)
├── teams.ts (TeamConfig, TeamMember, TeamTask, InboxMessage, PaneInfo)
├── push.ts (PushSubscriptionRecord, VapidKeys)
└── plan.ts (PlanTaskStatus, TddPhase, PlanItem re-export, NiceConfig,
DEFAULT_NICE_CONFIG, ProcessStats)
```
### Steps
1. **Create `src/types/` directory** and each domain file above.
2. **Move types** from `src/types.ts` into their domain files. Preserve all JSDoc comments. Each file should import from siblings as needed (e.g., `ralph.ts` imports `CircuitBreakerState` within itself — no cross-file deps needed since they're in the same file).
3. **Create barrel `src/types/index.ts`** that re-exports everything:
```typescript
export * from './common.js';
export * from './session.js';
export * from './task.js';
export * from './app-state.js';
export * from './respawn.js';
export * from './ralph.js';
export * from './api.js';
export * from './lifecycle.js';
export * from './run-summary.js';
export * from './tools.js';
export * from './teams.js';
export * from './push.js';
export * from './plan.js';
```
4. **Delete old `src/types.ts`** and replace with a single-line re-export barrel:
```typescript
export * from './types/index.js';
```
This ensures `import { ... } from './types.js'` continues to work everywhere — zero changes to 36 import sites.
5. **Verify**: `tsc --noEmit` and `npm run lint` must pass. No runtime changes.
### Internal Dependencies Between Domain Files
Some types reference others across domains. Handle with imports:
| File | Imports From |
|------|-------------|
| `app-state.ts` | `session.ts` (SessionState), `task.ts` (TaskState), `ralph.ts` (RalphLoopState, RalphSessionState) |
| `respawn.ts` | None (self-contained) |
| `ralph.ts` | None (self-contained) |
| `run-summary.ts` | None (self-contained) |
| `api.ts` | None (self-contained) |
| `session.ts` | `respawn.ts` (RespawnConfig), `ralph.ts` (RalphTrackerState, RalphTodoItem, CircuitBreakerStatus, RalphSessionState, RunSummaryEvent) |
Wait — `SessionState` references `RespawnConfig`, `RalphTrackerState`, `CircuitBreakerStatus`, and `RunSummaryEvent`. This creates imports from `session.ts` → `respawn.ts`, `ralph.ts`, `run-summary.ts`. This is fine (one-way deps, no cycles).
---
## 2. Split ralph-tracker.ts into focused modules
**Current**: 3,868 lines, single `RalphTracker` class with 5 responsibilities.
**Risk**: MEDIUM — class has shared mutable state, but extractable modules are well-isolated.
### Coupling Analysis Summary
| Module | Coupling | Extractability |
|--------|----------|----------------|
| Plan task tracking | LOW | HIGH — only reads `cycleCount` |
| Fix-plan file watching | LOW | HIGH — callback-based todo replacement |
| Iteration stall detection | LOW | HIGH — notification-based |
| RALPH_STATUS block parsing + circuit breaker | MEDIUM | MEDIUM — callback for circuit breaker updates |
| Todo parsing, loop detection, completion | HIGH | LOW — deeply entangled shared state |
### Target Structure
```
src/
├── ralph-tracker.ts (~1,800 LOC — core: output parsing, loop state,
│ todo management, completion detection)
├── ralph-plan-tracker.ts (~600 LOC — plan tasks, checkpoints, history, rollback)
├── ralph-status-parser.ts (~300 LOC — RALPH_STATUS block parsing, circuit breaker)
├── ralph-fix-plan-watcher.ts (~150 LOC — @fix_plan.md file watching)
└── ralph-stall-detector.ts (~80 LOC — iteration stall detection)
```
### Step 2a: Extract `RalphPlanTracker` (~600 LOC)
**Why first**: Lowest coupling. Only dependency is `cycleCount` for checkpoint detection.
**Extract these from `RalphTracker`**:
Types to export:
- `EnhancedPlanTask` (interface, currently lines 56-87)
- `CheckpointReview` (interface, currently lines 90-139)
Properties to move:
- `_planVersion: number`
- `_planHistory: Array<{version, timestamp, tasks, summary}>`
- `_planTasks: Map<string, EnhancedPlanTask>`
- `_checkpointIterations: number[]`
- `_lastCheckpointIteration: number`
Methods to move:
- `initializePlanTasks(items)`
- `updatePlanTask(taskId, update)`
- `addPlanTask(params)`
- `getPlanTasks()`
- `generateCheckpointReview()`
- `getPlanHistory()`
- `rollbackToVersion(version)`
- `isCheckpointDue()`
- `planVersion` getter
- `_savePlanToHistory()` (private)
- `_unblockDependentTasks()` (private)
- `_checkForCheckpoint()` (private)
Events emitted (define in new class):
- `planInitialized`
- `planTaskUpdate`
- `taskBlocked`
- `taskUnblocked`
- `planCheckpoint`
**Interface with parent**:
```typescript
export class RalphPlanTracker extends EventEmitter {
constructor() { ... }
// Parent calls this when iteration changes (for checkpoint detection)
notifyCycleCount(cycleCount: number): void { ... }
// Full public API moves here unchanged
initializePlanTasks(items: PlanItem[]): void { ... }
updatePlanTask(taskId: string, update: { ... }): { ... } | null { ... }
// ...etc
}
```
**In `RalphTracker`**: Replace plan methods with delegation:
```typescript
readonly planTracker = new RalphPlanTracker();
// Forward plan events
this.planTracker.on('planInitialized', (...args) => this.emit('planInitialized', ...args));
// ...etc
// In detectLoopStatus(), when cycleCount changes:
this.planTracker.notifyCycleCount(this._loopState.cycleCount);
```
### Step 2b: Extract `RalphFixPlanWatcher` (~150 LOC)
**Extract these**:
Properties:
- `_workingDir: string | null`
- `_fixPlanPath: string | null`
- `_fixPlanWatcher: FSWatcher | null`
- `_fixPlanWatcherErrorHandler`
- `_fixPlanReloadDeb`
Methods:
- `setWorkingDir(workingDir)`
- `loadFixPlanFromDisk()`
- `startWatchingFixPlan()`
- `stopWatchingFixPlan()`
- `handleFixPlanChange()`
- `isFileAuthoritative` getter
**Interface with parent**:
```typescript
export class RalphFixPlanWatcher extends EventEmitter {
get isFileAuthoritative(): boolean { ... }
setWorkingDir(workingDir: string): void { ... }
stop(): void { ... }
}
// Events:
// 'todosLoaded' → (todos: Array<{id, content, status, priority}>) — parent replaces _todos
```
**In `RalphTracker`**:
```typescript
readonly fixPlanWatcher = new RalphFixPlanWatcher();
constructor() {
this.fixPlanWatcher.on('todosLoaded', (items) => {
// Replace _todos with file-based items
this._todos.clear();
for (const item of items) {
this.addOrUpdateTodo(item.id, item.content, item.status, item.priority);
}
});
}
// Delegate isFileAuthoritative
get isFileAuthoritative(): boolean {
return this.fixPlanWatcher.isFileAuthoritative;
}
```
### Step 2c: Extract `RalphStallDetector` (~80 LOC)
**Extract these**:
Properties:
- `_lastIterationChangeTime`
- `_lastObservedIteration`
- `_iterationStallTimerId`
- `_iterationStallWarningMs`
- `_iterationStallCriticalMs`
- `_iterationStallWarned`
Methods:
- `startIterationStallDetection()`
- `stopIterationStallDetection()`
- `checkIterationStall()`
- `getIterationStallMetrics()`
- `configureIterationStallThresholds(warningMs, criticalMs)`
**Interface with parent**:
```typescript
export class RalphStallDetector extends EventEmitter {
constructor(private cleanup: CleanupManager) { ... }
start(): void { ... }
stop(): void { ... }
// Parent calls when iteration changes
notifyIterationChanged(iteration: number): void {
this._lastIterationChangeTime = Date.now();
this._lastObservedIteration = iteration;
this._iterationStallWarned = false;
}
// Parent calls to check if loop is active
setLoopActive(active: boolean): void { ... }
getIterationStallMetrics(): { ... } { ... }
}
// Events: 'iterationStallWarning', 'iterationStallCritical'
```
### Step 2d: Extract `RalphStatusParser` (~300 LOC)
**Extract these**:
Properties:
- `_circuitBreaker: CircuitBreakerStatus`
- `_statusBlockBuffer: string[]`
- `_inStatusBlock: boolean`
- `_lastStatusBlock: RalphStatusBlock | null`
- `_completionIndicators: number`
- `_exitGateMet: boolean`
- `_totalFilesModified: number`
- `_totalTasksCompleted: number`
Methods:
- `processStatusBlockLine(line)`
- `parseStatusBlock(lines)`
- `detectCompletionIndicators(line)`
- `updateCircuitBreaker(hasProgress, testsStatus, status)`
- `resetCircuitBreaker()`
- `circuitBreakerStatus` getter
- `lastStatusBlock` getter
- `cumulativeStats` getter
- `exitGateMet` getter
Regex patterns to move:
- `RALPH_STATUS_START_PATTERN` through `RALPH_RECOMMENDATION_PATTERN`
- `COMPLETION_INDICATOR_PATTERNS`
**Interface with parent**:
```typescript
export class RalphStatusParser extends EventEmitter {
processLine(line: string): void { ... } // calls processStatusBlockLine + detectCompletionIndicators
get circuitBreakerStatus(): CircuitBreakerStatus { ... }
get lastStatusBlock(): RalphStatusBlock | null { ... }
get exitGateMet(): boolean { ... }
get cumulativeStats(): { ... } { ... }
resetCircuitBreaker(): void { ... }
reset(): void { ... }
}
// Events: 'statusBlockDetected', 'circuitBreakerUpdate', 'exitGateMet'
```
**In `RalphTracker.processLine()`**:
```typescript
// Replace inline status block handling with delegation
this.statusParser.processLine(line);
```
### Step 2e: Keep in `ralph-tracker.ts` (~1,800 LOC)
The core remains tightly coupled and stays together:
- Output parsing pipeline (`processTerminalData`, `processCleanData`, `processLine`)
- Loop state management (`_loopState`, `detectLoopStatus`, `enable/disable/startLoop/stopLoop`)
- Todo management (`_todos`, `detectTodoItems`, `addOrUpdateTodo`, `updateTodoStatus`, `getTodoStats`)
- Completion detection (`detectCompletionPhrase`, `handleCompletionPhrase`, `calculateCompletionConfidence`)
- All-tasks-complete detection (`detectAllTasksComplete`)
- Auto-enable logic (`shouldAutoEnable`)
- Lifecycle (`reset`, `fullReset`, `clear`, `restoreState`, `destroy`)
- Event debouncing and buffering
The class coordinates the extracted modules via composition:
```typescript
export class RalphTracker extends EventEmitter {
readonly planTracker = new RalphPlanTracker();
readonly fixPlanWatcher = new RalphFixPlanWatcher();
readonly stallDetector: RalphStallDetector;
readonly statusParser = new RalphStatusParser();
constructor() {
super();
this.stallDetector = new RalphStallDetector(this.cleanup);
this._wireSubModuleEvents();
}
private _wireSubModuleEvents(): void {
// Forward all sub-module events through RalphTracker
// so external consumers don't need to know about the split
for (const event of ['planInitialized', 'planTaskUpdate', ...]) {
this.planTracker.on(event, (...args) => this.emit(event, ...args));
}
// ...same for statusParser, stallDetector, fixPlanWatcher
}
}
```
### Migration Safety
- All events continue to be emitted from `RalphTracker` (forwarded from sub-modules)
- All public methods stay on `RalphTracker` (delegated to sub-modules)
- External consumers (`session.ts`, `case-routes.ts`) see zero API changes
- New sub-modules are exposed as `readonly` properties for direct access where needed
---
## 3. Split respawn-controller.ts into focused modules
**Current**: 3,611 lines, single `RespawnController` class with 6 responsibilities.
**Risk**: MEDIUM — health scoring and metrics are cleanly decoupled; detection is tightly coupled.
### Coupling Analysis Summary
| Module | Coupling | Extractability |
|--------|----------|----------------|
| Health scoring | NONE | HIGH — pure calculations from metrics |
| Cycle metrics | LOW | HIGH — standalone tracking |
| Adaptive timing | LOW | HIGH — standalone timing adjustments |
| Stuck-state detection | LOW | MEDIUM — needs state + config refs |
| Pattern detection utilities | NONE | HIGH — pure functions |
| State machine + idle detection + AI checkers | HIGH | LOW — deeply entangled |
### Target Structure
```
src/
├── respawn-controller.ts (~2,200 LOC — state machine, idle detection,
│ AI checkers, terminal handling, hook signals,
│ auto-accept, step execution)
├── respawn-health.ts (~250 LOC — health scoring + recommendations)
├── respawn-metrics.ts (~200 LOC — cycle metrics + aggregate stats)
├── respawn-adaptive-timing.ts (~100 LOC — adaptive timing with percentile calc)
└── respawn-patterns.ts (~50 LOC — terminal pattern detection utilities)
```
### Step 3a: Extract `RespawnPatterns` (~50 LOC)
**Pure utility functions, zero coupling**.
Move:
- `isCompletionMessage(data): boolean`
- `hasWorkingPattern(data, window): boolean`
- `extractTokenCount(data): number | null`
- `PROMPT_PATTERNS` array
- `WORKING_PATTERNS` array
```typescript
// src/respawn-patterns.ts
import { TOKEN_PATTERN, SPINNER_PATTERN } from './utils/index.js';
export const PROMPT_PATTERNS = ['❯', '>', '$', '%', '#'];
export const WORKING_PATTERNS = [/* 70+ patterns */];
export function isCompletionMessage(data: string): boolean { ... }
export function hasWorkingPattern(data: string, window: string): boolean { ... }
export function extractTokenCount(data: string): number | null { ... }
```
**In `RespawnController`**: Import and call:
```typescript
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
```
### Step 3b: Extract `RespawnAdaptiveTiming` (~100 LOC)
**Self-contained timing controller**.
Move properties:
- `timingHistory: TimingHistory`
Move methods:
- `recordTimingData(idleDetectionMs, cycleDurationMs)`
- `updateAdaptiveTiming()`
- `getTimingHistory()`
- `getAdaptiveCompletionConfirmMs()`
```typescript
export class RespawnAdaptiveTiming {
private timingHistory: TimingHistory;
constructor(private config: { adaptiveMinConfirmMs: number; adaptiveMaxConfirmMs: number }) {
this.timingHistory = { recentIdleDetectionMs: [], recentCycleDurationMs: [], ... };
}
recordTimingData(idleDetectionMs: number, cycleDurationMs: number): void { ... }
getAdaptiveCompletionConfirmMs(): number { ... }
getTimingHistory(): TimingHistory { ... }
reset(): void { ... }
}
```
### Step 3c: Extract `RespawnCycleMetrics` (~200 LOC)
**Standalone metrics tracker**.
Move properties:
- `currentCycleMetrics`
- `recentCycleMetrics[]`
- `aggregateMetrics`
- `MAX_CYCLE_METRICS_IN_MEMORY`
Move methods:
- `startCycleMetrics(idleReason)`
- `recordCycleStep(step)`
- `completeCycleMetrics(outcome, errorMessage?)`
- `updateAggregateMetrics(metrics)`
- `getAggregateMetrics()`
- `getRecentCycleMetrics(limit?)`
```typescript
export class RespawnCycleMetricsTracker {
private currentCycleMetrics: Partial<RespawnCycleMetrics> | null = null;
private recentCycleMetrics: RespawnCycleMetrics[] = [];
private aggregateMetrics: RespawnAggregateMetrics;
startCycle(sessionId: string, cycleNumber: number, idleReason: string): void { ... }
recordStep(step: string): void { ... }
completeCycle(outcome: CycleOutcome, errorMessage?: string): RespawnCycleMetrics | null { ... }
getAggregate(): RespawnAggregateMetrics { ... }
getRecent(limit?: number): RespawnCycleMetrics[] { ... }
reset(): void { ... }
}
```
**Callback**: `completeCycle()` returns the completed metrics so the controller can pass them to `adaptiveTiming.recordTimingData()`.
### Step 3d: Extract `RespawnHealthCalculator` (~250 LOC)
**Pure calculation — no state of its own**.
Move methods:
- `calculateHealthScore()`
- `calculateCycleSuccessScore()`
- `calculateCircuitBreakerScore()`
- `calculateIterationProgressScore()`
- `calculateAiCheckerScore()`
- `calculateStuckRecoveryScore()`
- `generateHealthRecommendations(components)`
- `generateHealthSummary(score, status, components)`
- `shouldSkipClear()` (belongs here since it's a pure calculation on token/config)
```typescript
export interface HealthInputs {
aggregateMetrics: RespawnAggregateMetrics;
circuitBreakerStatus: CircuitBreakerStatus;
iterationStallMetrics: { stallDurationMs: number; warningMs: number; criticalMs: number } | null;
aiCheckerState: { disabled: boolean; inCooldown: boolean; hasErrors: boolean };
stuckRecoveryCount: number;
maxStuckRecoveries: number;
}
export function calculateHealthScore(inputs: HealthInputs): RalphLoopHealthScore { ... }
export function shouldSkipClear(
lastTokenCount: number,
skipClearThresholdPercent: number,
maxContextTokens: number
): boolean { ... }
```
**Made as pure functions** (not a class) since they hold no state.
### Step 3e: Keep in `respawn-controller.ts` (~2,200 LOC)
The core state machine, idle detection, and AI checker integration stays:
- State machine transitions (`setState`, `start`, `stop`, `pause`, `resume`)
- Terminal data handling (`handleTerminalData`)
- All 5 idle detection layers + hook signals
- AI checker integration (`tryStartAiCheck`, `startAiCheck`, `startPlanCheck`)
- Auto-accept logic
- Step execution (`sendUpdateDocs`, `sendClear`, `sendInit`, `sendKickstart`)
- Timer management (`startTrackedTimer`, `cancelTrackedTimer`)
- Stuck-state detection and recovery
- Action logging
The class composes extracted modules:
```typescript
import { RespawnAdaptiveTiming } from './respawn-adaptive-timing.js';
import { RespawnCycleMetricsTracker } from './respawn-metrics.js';
import { calculateHealthScore, shouldSkipClear } from './respawn-health.js';
import { isCompletionMessage, hasWorkingPattern, extractTokenCount } from './respawn-patterns.js';
export class RespawnController extends EventEmitter {
private adaptiveTiming: RespawnAdaptiveTiming;
private cycleMetrics: RespawnCycleMetricsTracker;
calculateHealthScore(): RalphLoopHealthScore {
return calculateHealthScore({
aggregateMetrics: this.cycleMetrics.getAggregate(),
circuitBreakerStatus: this.session.ralphTracker.circuitBreakerStatus,
iterationStallMetrics: this.session.ralphTracker.getIterationStallMetrics(),
aiCheckerState: { ... },
stuckRecoveryCount: this.stuckRecoveryCount,
maxStuckRecoveries: this.config.maxStuckRecoveries ?? 3,
});
}
}
```
---
## 4. Split session.ts into focused modules
**Current**: 2,418 lines, single `Session` class.
**Risk**: LOW-MEDIUM — extractable pieces are utility-like with clear boundaries.
### Coupling Analysis Summary
| Module | Coupling | Extractability |
|--------|----------|----------------|
| CLI arg builder | NONE | HIGH — pure functions used at spawn time |
| Auto-compact/clear | LOW | HIGH — self-contained automation with config |
| Token tracking | LOW | MEDIUM — reads PTY output, writes state |
| Task description cache | LOW | HIGH — separate LRU cache |
| PTY + mux lifecycle | HIGH | KEEP — core of the class |
| Tracker integration | HIGH | KEEP — event forwarding plumbing |
### Target Structure
```
src/
├── session.ts (~1,600 LOC — PTY lifecycle, terminal I/O,
│ tracker integration, output processing,
│ token tracking, state management)
├── session-cli-builder.ts (~250 LOC — Claude/OpenCode CLI arg construction)
├── session-auto-ops.ts (~300 LOC — auto-compact, auto-clear automation)
└── session-task-cache.ts (~100 LOC — task description LRU cache)
```
### Step 4a: Extract `SessionCliBuilder` (~250 LOC)
**Pure functions — zero coupling to Session instance**.
Move:
- `buildClaudeArgs()` logic (currently inlined in `startInteractive` and `runPrompt`)
- `buildOpenCodeArgs()` logic
- Model mapping constants
- Claude mode to flag mapping
- Environment variable construction
```typescript
// src/session-cli-builder.ts
export interface CliBuilderConfig {
claudeMode: ClaudeMode;
model?: string;
workingDir: string;
sessionId: string;
niceConfig?: NiceConfig;
isOpenCode?: boolean;
openCodeConfig?: OpenCodeConfig;
}
export function buildInteractiveArgs(config: CliBuilderConfig): string[] { ... }
export function buildPromptArgs(config: CliBuilderConfig, prompt: string): string[] { ... }
export function buildShellArgs(shell?: string): string[] { ... }
export function buildClaudeEnv(config: CliBuilderConfig): Record<string, string> { ... }
```
### Step 4b: Extract `SessionAutoOps` (~300 LOC)
**Self-contained automation with config-based thresholds**.
Move properties:
- `_autoCompactThreshold`
- `_autoClearThreshold`
- `_isAutoCompacting`
- `_isAutoClearing`
- `_autoCompactCount`
- `_autoClearCount`
- `_lastAutoCompactTime`
- `_lastAutoClearTime`
Move methods:
- `checkAutoCompact(tokenCount)`
- `performAutoCompact()`
- `checkAutoClear(tokenCount)`
- `performAutoClear()`
- Auto-compact/clear threshold configuration
```typescript
export class SessionAutoOps extends EventEmitter {
constructor(
private writeCommand: (command: string) => Promise<void>,
private getTokenCount: () => number,
config: { compactThreshold: number; clearThreshold: number }
) { ... }
/** Called after token count updates. Checks thresholds and triggers if needed. */
checkThresholds(tokenCount: number): void { ... }
updateConfig(config: { compactThreshold?: number; clearThreshold?: number }): void { ... }
getStats(): { autoCompactCount: number; autoClearCount: number; ... } { ... }
}
// Events: 'autoCompact', 'autoClear'
```
**In `Session`**: Compose and wire:
```typescript
private autoOps = new SessionAutoOps(
(cmd) => this.writeViaMux(cmd),
() => this._state.tokenCount,
{ compactThreshold: 110_000, clearThreshold: 140_000 }
);
```
### Step 4c: Extract `SessionTaskCache` (~100 LOC)
**Isolated LRU cache for task descriptions**.
Move:
- `_taskDescriptionCache: LRUMap<number, { description: string; timestamp: number }>`
- `_taskDescriptionMaxAge`
- `findTaskDescriptionNear(lineNumber)`
- `cacheTaskDescription(lineNumber, description)`
```typescript
export class SessionTaskCache {
private cache: LRUMap<number, { description: string; timestamp: number }>;
private maxAgeMs: number;
constructor(maxSize: number = 50, maxAgeMs: number = 30_000) { ... }
find(lineNumber: number, searchRadius: number = 50): string | null { ... }
add(lineNumber: number, description: string): void { ... }
clear(): void { ... }
}
```
### Step 4d: Keep in `session.ts` (~1,600 LOC)
The core stays together:
- PTY process management (`spawn`, `kill`, `resize`, `writeViaMux`)
- Data streaming pipeline (PTY → buffer → ANSI strip → JSON parse → events)
- Tracker initialization and event forwarding (RalphTracker, BashToolParser, TaskTracker)
- Output processing (message extraction, completion detection)
- Token tracking (status line parsing)
- State management (`toState()`, `updateState()`)
- Session lifecycle (`startInteractive`, `startShell`, `runPrompt`)
- CLI info detection (version, model, account)
---
## 5. Execution Order & Dependencies
Execute in this order to minimize risk. Each step is independently deployable.
```
Step 1: types.ts split
↓ (no runtime change, just file reorganization)
Step 2a: RalphPlanTracker extraction
↓ (independent of types split)
Step 2b: RalphFixPlanWatcher extraction
Step 2c: RalphStallDetector extraction
Step 2d: RalphStatusParser extraction
↓ (ralph-tracker.ts now ~1,800 LOC)
Step 3a: RespawnPatterns extraction
Step 3b: RespawnAdaptiveTiming extraction
Step 3c: RespawnCycleMetrics extraction
Step 3d: RespawnHealthCalculator extraction
↓ (respawn-controller.ts now ~2,200 LOC)
Step 4a: SessionCliBuilder extraction
Step 4b: SessionAutoOps extraction
Step 4c: SessionTaskCache extraction
↓ (session.ts now ~1,600 LOC)
```
**Parallelization**: Steps 1, 2a-2d, 3a-3d, and 4a-4c can be done by separate agents in parallel since they touch different files. However, within each group, sequential execution is safer.
### Risk Mitigation
- **Barrel exports**: Every split uses delegation + barrel re-export so external consumers see zero API changes
- **Event forwarding**: Sub-modules emit events, parent class forwards them — no event contract changes
- **Incremental**: Each step can be verified independently with `tsc --noEmit` + `npm run lint`
- **No test changes needed**: External API stays identical; existing tests continue to pass
---
## 6. Validation Checklist
After each step, verify:
- [ ] `tsc --noEmit` passes (no type errors)
- [ ] `npm run lint` passes (no unused imports, etc.)
- [ ] `npm run format:check` passes
- [ ] `npx vitest run test/respawn-controller.test.ts` passes (for respawn splits)
- [ ] `npx vitest run test/ralph-tracker.test.ts` passes (for ralph splits)
- [ ] `npx vitest run test/session-manager.test.ts` passes (for session splits)
- [ ] Dev server starts: `npx tsx src/index.ts web`
- [ ] Existing sessions work (create, interact, delete)
- [ ] Respawn cycle works (enable respawn, verify idle detection fires)
- [ ] No new circular dependencies: `npx madge --circular src/`
### Size Targets
| File | Before | After |
|------|--------|-------|
| `src/types.ts` | 1,443 LOC | 1 LOC (re-export barrel) |
| `src/ralph-tracker.ts` | 3,868 LOC | ~1,800 LOC |
| `src/respawn-controller.ts` | 3,611 LOC | ~2,200 LOC |
| `src/session.ts` | 2,418 LOC | ~1,600 LOC |
| **Total new files** | — | 12 files |
| **Net LOC change** | — | ~0 (refactor only) |
-738
View File
@@ -1,738 +0,0 @@
# Phase 1 Implementation Plan: Quick Wins
**Source**: `docs/code-structure-findings.md` (Phase 1 - Quick Wins section)
**Estimated effort**: 1-2 days
**Tasks**: 5 independent tasks (can be done in parallel unless noted)
---
## Safety Constraints
Before starting ANY work, read and follow these rules:
1. **Never run `npx vitest run`** (full suite) -- it kills tmux sessions. You are running inside a Codeman-managed tmux session.
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
3. **Never test on port 3000** -- the live dev server runs there. Tests use ports 3150+.
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
6. **Never kill tmux sessions** -- check `echo $CODEMAN_MUX` first.
---
## Task Dependencies
All 5 tasks are independent and can be done in parallel. However:
- Task 1 (barrel exports) is a prerequisite if you want to update import sites to use the barrel after Task 3 (consolidate EXEC_TIMEOUT_MS). The EXEC_TIMEOUT_MS consolidation creates a new export that should be added to the barrel.
- Task 2 (delete dead functions) removes functions that Task 1 would otherwise need to add to the barrel. Do Task 2 first or simultaneously with Task 1 to avoid adding exports for dead code.
**Recommended order**: Task 2 -> Task 1 -> Task 3 -> Task 4 -> Task 5
---
## Task 1: Export Missing Functions from Utils Barrel
**File**: `src/utils/index.ts`
**Time**: ~30 minutes
### Problem
The barrel file (`src/utils/index.ts`) is missing exports for several functions that are defined in util modules, forcing consumers to use deep imports or preventing usage entirely.
### Missing Exports
From `src/utils/regex-patterns.ts`:
- `createAnsiPatternFull()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
- `createAnsiPatternSimple()` -- factory for fresh ANSI regex (documented in CLAUDE.md)
- `stripAnsi()` -- ANSI stripping utility
- `SAFE_PATH_PATTERN` -- regex for safe file paths (currently deep-imported by `schemas.ts` and `tmux-manager.ts`)
From `src/utils/token-validation.ts`:
- `validateTokenCounts()` -- token count validation (documented in CLAUDE.md)
- `validateTokensAndCost()` -- token + cost validation (documented in CLAUDE.md)
**Note**: Do NOT export `isSimilar`, `isSimilarByDistance`, `levenshteinDistance`, or `normalizePhrase` from `string-similarity.ts` -- these are dead code (see Task 2).
### Edit 1: Add missing regex-patterns exports
**File**: `src/utils/index.ts`
**Old code** (lines 13-18):
```typescript
export {
ANSI_ESCAPE_PATTERN_FULL,
ANSI_ESCAPE_PATTERN_SIMPLE,
TOKEN_PATTERN,
SPINNER_PATTERN,
} from './regex-patterns.js';
```
**New code**:
```typescript
export {
ANSI_ESCAPE_PATTERN_FULL,
ANSI_ESCAPE_PATTERN_SIMPLE,
TOKEN_PATTERN,
SPINNER_PATTERN,
createAnsiPatternFull,
createAnsiPatternSimple,
stripAnsi,
SAFE_PATH_PATTERN,
} from './regex-patterns.js';
```
### Edit 2: Add missing token-validation exports
**File**: `src/utils/index.ts`
**Old code** (line 19):
```typescript
export { MAX_SESSION_TOKENS } from './token-validation.js';
```
**New code**:
```typescript
export { MAX_SESSION_TOKENS, validateTokenCounts, validateTokensAndCost } from './token-validation.js';
```
### Optional follow-up: Update deep imports to use barrel
These files currently deep-import `SAFE_PATH_PATTERN` and could be updated to use the barrel instead:
- `src/web/schemas.ts` line 11: `import { SAFE_PATH_PATTERN } from '../utils/regex-patterns.js';` could become `import { SAFE_PATH_PATTERN } from '../utils/index.js';`
- `src/tmux-manager.ts` line 44: `import { SAFE_PATH_PATTERN } from './utils/regex-patterns.js';` could become part of existing barrel import
This is a low-priority cosmetic change. The barrel export itself is the important fix.
### Verification
```bash
tsc --noEmit
npm run lint
```
---
## Task 2: Delete Dead Utility Functions
**File**: `src/utils/string-similarity.ts`
**Time**: ~15 minutes
### Problem
Four exported functions in `string-similarity.ts` are never imported anywhere in the codebase:
- `levenshteinDistance()` (lines 27-69)
- `isSimilar()` (lines 106-108)
- `isSimilarByDistance()` (lines 123-125)
- `normalizePhrase()` (lines 139-144)
Only three functions are actually used (all by `ralph-tracker.ts` via the barrel):
- `stringSimilarity()` -- uses `levenshteinDistance()` internally
- `fuzzyPhraseMatch()` -- uses `normalizePhrase()` and `isSimilarByDistance()` internally
- `todoContentHash()`
### Strategy
`levenshteinDistance()` is called by `stringSimilarity()`, and `normalizePhrase()` and `isSimilarByDistance()` are called by `fuzzyPhraseMatch()`. So they cannot be deleted -- they just need to be un-exported (made private to the module).
`isSimilar()` is truly dead -- not called by anything. Delete it entirely.
### Edit 1: Remove `export` from `levenshteinDistance`
**File**: `src/utils/string-similarity.ts`
**Old code** (line 27):
```typescript
export function levenshteinDistance(a: string, b: string): number {
```
**New code**:
```typescript
function levenshteinDistance(a: string, b: string): number {
```
### Edit 2: Delete `isSimilar` function entirely
**File**: `src/utils/string-similarity.ts`
**Old code** (lines 94-108):
```typescript
/**
* Check if two strings are similar within a given threshold.
*
* @param a - First string
* @param b - Second string
* @param threshold - Minimum similarity ratio (default: 0.85 = 85% similar)
* @returns True if similarity >= threshold
*
* @example
* isSimilar('COMPLETE', 'COMPLET', 0.85) // true (87.5% similar)
* isSimilar('COMPLETE', 'DONE', 0.85) // false (0% similar)
*/
export function isSimilar(a: string, b: string, threshold = 0.85): boolean {
return stringSimilarity(a, b) >= threshold;
}
```
**New code**: (delete entirely -- replace with empty string)
### Edit 3: Remove `export` from `isSimilarByDistance`
**File**: `src/utils/string-similarity.ts`
**Old code** (line 123):
```typescript
export function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
```
**New code**:
```typescript
function isSimilarByDistance(a: string, b: string, maxDistance = 2): boolean {
```
### Edit 4: Remove `export` from `normalizePhrase`
**File**: `src/utils/string-similarity.ts`
**Old code** (line 139):
```typescript
export function normalizePhrase(phrase: string): string {
```
**New code**:
```typescript
function normalizePhrase(phrase: string): string {
```
### Verification
```bash
tsc --noEmit
npx vitest run test/string-utilities.test.ts
npm run lint
```
Note: If `test/string-utilities.test.ts` imports any of the now-unexported functions, those test imports will fail. Check the test file and remove tests for `isSimilar` (deleted) and update any direct tests for `levenshteinDistance`, `isSimilarByDistance`, `normalizePhrase` to test them indirectly through the public API (`stringSimilarity`, `fuzzyPhraseMatch`), or remove those tests.
---
## Task 3: Consolidate Duplicated `EXEC_TIMEOUT_MS` Constant
**Files**:
- `src/utils/claude-cli-resolver.ts` (line 17)
- `src/utils/opencode-cli-resolver.ts` (line 16)
- `src/tmux-manager.ts` (line 63) -- also has its own copy
**Time**: ~15 minutes
### Problem
`EXEC_TIMEOUT_MS = 5000` is defined identically in three files. Changes need to happen in all three places.
### Strategy
Create a shared constant and export it. The natural home is a new config file since the existing config files (`buffer-limits.ts`, `map-limits.ts`) follow this pattern. However, to keep it minimal, we can add it to an existing config file or create a small one.
**Recommended approach**: Add to `src/config/timing-config.ts` (new file) as a single constant. This file can grow later in Phase 6 to hold other timing constants.
Alternatively, the simplest approach: export from one of the existing utils and import in the others. Since both CLI resolvers are in `src/utils/`, the cleanest approach is to put it in a shared location.
### Option A: Add to existing config (simpler)
Create `src/config/exec-timeout.ts`:
**New file**: `src/config/exec-timeout.ts`
```typescript
/**
* Timeout for child process exec commands (e.g., `which claude`, `which opencode`, tmux commands).
* Used across CLI resolvers and tmux manager.
*/
export const EXEC_TIMEOUT_MS = 5000;
```
### Edit 1: Update `claude-cli-resolver.ts`
**File**: `src/utils/claude-cli-resolver.ts`
**Old code** (lines 11-17):
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { delimiter, dirname, join } from 'node:path';
import { homedir } from 'node:os';
/** Timeout for exec commands (5 seconds) */
const EXEC_TIMEOUT_MS = 5000;
```
**New code**:
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { delimiter, dirname, join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
```
### Edit 2: Update `opencode-cli-resolver.ts`
**File**: `src/utils/opencode-cli-resolver.ts`
**Old code** (lines 10-16):
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { homedir } from 'node:os';
/** Timeout for exec commands (5 seconds) */
const EXEC_TIMEOUT_MS = 5000;
```
**New code**:
```typescript
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
```
### Edit 3: Update `tmux-manager.ts`
**File**: `src/tmux-manager.ts`
**Old code** (line 63):
```typescript
const EXEC_TIMEOUT_MS = 5000;
```
**New code**:
```typescript
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
```
Note: `tmux-manager.ts` already has many imports at the top of the file. Add this import near the other local imports (around lines 43-56). The `const EXEC_TIMEOUT_MS = 5000;` on line 63 should be deleted entirely (replaced with the import).
### Verification
```bash
tsc --noEmit
npm run lint
```
---
## Task 4: Add `z.infer` to Zod Schemas
**Files**:
- `src/web/schemas.ts` (add type exports)
- `src/types.ts` (replace manual interfaces with `z.infer` re-exports where applicable)
**Time**: ~2 hours
### Problem
All 30+ Zod schemas in `schemas.ts` define validation rules, but zero use `z.infer` to derive TypeScript types. Instead, `types.ts` manually duplicates interfaces that match the schemas. When a schema changes, the type must be manually updated too.
### Strategy
Add `z.infer` type exports to `schemas.ts` for each exported schema. This creates derived types as the single source of truth. For schemas that have corresponding manual interfaces in `types.ts`, the manual interface can be replaced with a re-export of the inferred type.
**Important**: Not all schemas have matching interfaces in `types.ts`. The `RespawnConfig` interface in `types.ts` (line 395) has all required fields, while `RespawnConfigSchema` has all optional fields (it's for partial updates). These are NOT the same type and should NOT be unified.
### Edit 1: Add inferred type exports to `schemas.ts`
**File**: `src/web/schemas.ts`
After each schema definition, add a corresponding type export. Add the following lines at the **end of the file** (after line 509):
**Old code** (end of file, lines 506-509):
```typescript
.optional(),
});
```
Wait -- the end of file is actually at line 509 after the `RalphLoopStartSchema`. Add the type exports after the last schema:
**Append to end of file** `src/web/schemas.ts`:
```typescript
// ========== Inferred Types ==========
// Derive TypeScript types from Zod schemas (single source of truth)
export type CreateSessionInput = z.infer<typeof CreateSessionSchema>;
export type RunPromptInput = z.infer<typeof RunPromptSchema>;
export type ResizeInput = z.infer<typeof ResizeSchema>;
export type CreateCaseInput = z.infer<typeof CreateCaseSchema>;
export type QuickStartInput = z.infer<typeof QuickStartSchema>;
export type HookEventInput = z.infer<typeof HookEventSchema>;
export type RespawnConfigInput = z.infer<typeof RespawnConfigSchema>;
export type ConfigUpdateInput = z.infer<typeof ConfigUpdateSchema>;
export type SettingsUpdateInput = z.infer<typeof SettingsUpdateSchema>;
export type SessionInputWithLimitInput = z.infer<typeof SessionInputWithLimitSchema>;
export type SessionNameInput = z.infer<typeof SessionNameSchema>;
export type SessionColorInput = z.infer<typeof SessionColorSchema>;
export type RalphConfigInput = z.infer<typeof RalphConfigSchema>;
export type FixPlanImportInput = z.infer<typeof FixPlanImportSchema>;
export type RalphPromptWriteInput = z.infer<typeof RalphPromptWriteSchema>;
export type AutoClearInput = z.infer<typeof AutoClearSchema>;
export type AutoCompactInput = z.infer<typeof AutoCompactSchema>;
export type ImageWatcherInput = z.infer<typeof ImageWatcherSchema>;
export type FlickerFilterInput = z.infer<typeof FlickerFilterSchema>;
export type QuickRunInput = z.infer<typeof QuickRunSchema>;
export type ScheduledRunInput = z.infer<typeof ScheduledRunSchema>;
export type LinkCaseInput = z.infer<typeof LinkCaseSchema>;
export type GeneratePlanInput = z.infer<typeof GeneratePlanSchema>;
export type GeneratePlanDetailedInput = z.infer<typeof GeneratePlanDetailedSchema>;
export type CancelPlanInput = z.infer<typeof CancelPlanSchema>;
export type PlanTaskUpdateInput = z.infer<typeof PlanTaskUpdateSchema>;
export type PlanTaskAddInput = z.infer<typeof PlanTaskAddSchema>;
export type CpuLimitInput = z.infer<typeof CpuLimitSchema>;
export type SubagentWindowStatesInput = z.infer<typeof SubagentWindowStatesSchema>;
export type SubagentParentMapInput = z.infer<typeof SubagentParentMapSchema>;
export type InteractiveRespawnInput = z.infer<typeof InteractiveRespawnSchema>;
export type RespawnEnableInput = z.infer<typeof RespawnEnableSchema>;
export type PushSubscribeInput = z.infer<typeof PushSubscribeSchema>;
export type PushPreferencesUpdateInput = z.infer<typeof PushPreferencesUpdateSchema>;
export type RalphLoopStartInput = z.infer<typeof RalphLoopStartSchema>;
```
### What NOT to do
Do NOT replace the `RespawnConfig` interface in `types.ts` with `z.infer<typeof RespawnConfigSchema>`. The schema has all optional fields (for partial config updates), but the interface has required fields (for the full config object). These are intentionally different shapes.
Similarly, do NOT try to unify every interface in `types.ts` with a schema -- most interfaces in `types.ts` represent internal domain objects (SessionState, TaskState, etc.) that have no corresponding Zod schema. The schemas only exist for API request validation.
### Future opportunity
In a future phase, route handlers in `server.ts` can use these inferred types for request body typing:
```typescript
const body = CreateSessionSchema.parse(request.body) as CreateSessionInput;
```
This task only adds the type exports. Migrating route handlers to use them is out of scope.
### Verification
```bash
tsc --noEmit
npm run lint
npm run format:check
```
---
## Task 5: Fix Weak `not.toThrow()` Tests with Behavioral Assertions
**Files**:
- `test/task-tracker.test.ts` -- 6 instances
- `test/image-watcher.test.ts` -- 1 instance
- `test/task-queue.test.ts` -- 1 instance
- `test/hooks-config.test.ts` -- 1 instance
- `test/session-manager.test.ts` -- 1 instance
**Time**: ~1 hour
### Problem
10 tests only assert `not.toThrow()` without verifying the actual defensive behavior. These tests prove the code doesn't crash but don't verify it does the right thing.
### Fix Strategy
After each `not.toThrow()`, add a behavioral assertion that verifies the state is correct (e.g., no tasks were created, no side effects occurred).
### Edit 1: `task-tracker.test.ts` -- null message (line 566)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle null message', () => {
expect(() => tracker.processMessage(null)).not.toThrow();
});
```
**New code**:
```typescript
it('should handle null message', () => {
expect(() => tracker.processMessage(null)).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
expect(tracker.getRunningCount()).toBe(0);
});
```
### Edit 2: `task-tracker.test.ts` -- message without content (line 569-571)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle message without content', () => {
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
});
```
**New code**:
```typescript
it('should handle message without content', () => {
expect(() => tracker.processMessage({ message: {} })).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
});
```
### Edit 3: `task-tracker.test.ts` -- empty content array (line 573-575)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle empty content array', () => {
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
});
```
**New code**:
```typescript
it('should handle empty content array', () => {
expect(() => tracker.processMessage({ message: { content: [] } })).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
});
```
### Edit 4: `task-tracker.test.ts` -- tool_result for unknown task (lines 577-590)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle tool_result for unknown task', () => {
expect(() => {
tracker.processMessage({
message: {
content: [{
type: 'tool_result',
tool_use_id: 'unknown-task',
is_error: false,
content: 'Done',
}],
},
});
}).not.toThrow();
});
```
**New code**:
```typescript
it('should handle tool_result for unknown task', () => {
expect(() => {
tracker.processMessage({
message: {
content: [{
type: 'tool_result',
tool_use_id: 'unknown-task',
is_error: false,
content: 'Done',
}],
},
});
}).not.toThrow();
expect(tracker.getTask('unknown-task')).toBeUndefined();
expect(tracker.getAllTasks().size).toBe(0);
});
```
### Edit 5: `task-tracker.test.ts` -- empty terminal output (lines 592-595)
**File**: `test/task-tracker.test.ts`
**Old code**:
```typescript
it('should handle empty terminal output', () => {
expect(() => tracker.processTerminalOutput('')).not.toThrow();
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
});
```
**New code**:
```typescript
it('should handle empty terminal output', () => {
expect(() => tracker.processTerminalOutput('')).not.toThrow();
expect(() => tracker.processTerminalOutput(' ')).not.toThrow();
expect(tracker.getAllTasks().size).toBe(0);
expect(tracker.getRunningCount()).toBe(0);
});
```
### Edit 6: `image-watcher.test.ts` -- unwatchSession for non-watched session (line 123)
**File**: `test/image-watcher.test.ts`
**Old code**:
```typescript
it('should be safe to call for non-watched session', () => {
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
});
```
**New code**:
```typescript
it('should be safe to call for non-watched session', () => {
expect(() => watcher.unwatchSession('nonexistent')).not.toThrow();
expect(watcher.getWatchedSessions()).toHaveLength(0);
});
```
### Edit 7: `task-queue.test.ts` -- dependencies on non-existent tasks (lines 538-542)
**File**: `test/task-queue.test.ts`
**Old code**:
```typescript
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
// Dependencies on non-existent tasks are valid - they just won't be satisfied
expect(() => {
queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
}).not.toThrow();
});
```
**New code**:
```typescript
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
// Dependencies on non-existent tasks are valid - they just won't be satisfied
let task: ReturnType<typeof queue.addTask> | undefined;
expect(() => {
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
}).not.toThrow();
expect(task).toBeDefined();
expect(task!.dependencies).toEqual(['non-existent-id']);
// Task should be pending but blocked (dependency unsatisfied)
expect(queue.next()?.prompt).toBeUndefined();
});
```
Wait -- `queue.next()` returns `null` when no next task is available (all blocked). Let me adjust:
**New code** (corrected):
```typescript
it('should allow dependencies on non-existent tasks (just unsatisfied, not a cycle)', () => {
// Dependencies on non-existent tasks are valid - they just won't be satisfied
let task: ReturnType<typeof queue.addTask> | undefined;
expect(() => {
task = queue.addTask({ prompt: 'Task D', dependencies: ['non-existent-id'] });
}).not.toThrow();
expect(task).toBeDefined();
expect(task!.dependencies).toEqual(['non-existent-id']);
// Task exists but is blocked (dependency unsatisfied), so next() skips it
expect(queue.getAllTasks()).toHaveLength(1);
expect(queue.next()).toBeNull();
});
```
### Edit 8: `hooks-config.test.ts` -- valid JSON check (line 129)
**File**: `test/hooks-config.test.ts`
**Old code**:
```typescript
it('should write valid JSON', () => {
writeHooksConfig(testDir);
const settingsPath = join(testDir, '.claude', 'settings.local.json');
const content = readFileSync(settingsPath, 'utf-8');
expect(() => JSON.parse(content)).not.toThrow();
});
```
**New code**:
```typescript
it('should write valid JSON', () => {
writeHooksConfig(testDir);
const settingsPath = join(testDir, '.claude', 'settings.local.json');
const content = readFileSync(settingsPath, 'utf-8');
const parsed = JSON.parse(content);
expect(parsed).toBeDefined();
expect(typeof parsed).toBe('object');
expect(parsed.hooks).toBeDefined();
});
```
### Edit 9: `session-manager.test.ts` -- stopSession for non-existent (line 216)
**File**: `test/session-manager.test.ts`
**Old code**:
```typescript
it('should handle non-existent session gracefully', async () => {
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
});
```
**New code**:
```typescript
it('should handle non-existent session gracefully', async () => {
await expect(manager.stopSession('non-existent')).resolves.not.toThrow();
expect(manager.getSessionCount()).toBe(0);
});
```
### Verification
Run each test file individually:
```bash
npx vitest run test/task-tracker.test.ts
npx vitest run test/image-watcher.test.ts
npx vitest run test/task-queue.test.ts
npx vitest run test/hooks-config.test.ts
npx vitest run test/session-manager.test.ts
```
**Important**: `hooks-config.test.ts` and `session-manager.test.ts` spawn real servers on ports 3130-3131. Only run them if you are NOT running other tests that use those ports.
---
## Final Verification Checklist
After all 5 tasks are complete, run the following in order:
```bash
# 1. TypeScript type checking
tsc --noEmit
# 2. Linting
npm run lint
# 3. Formatting
npm run format:check
# 4. Run affected test files individually (NOT the full suite)
npx vitest run test/string-utilities.test.ts
npx vitest run test/task-tracker.test.ts
npx vitest run test/image-watcher.test.ts
npx vitest run test/task-queue.test.ts
npx vitest run test/session-manager.test.ts
npx vitest run test/hooks-config.test.ts
```
If any formatting issues arise, fix with:
```bash
npm run format
```
If any lint issues arise, fix with:
```bash
npm run lint:fix
```
### Summary of Changes
| Task | Files Modified | Files Created |
|------|---------------|---------------|
| 1. Barrel exports | `src/utils/index.ts` | -- |
| 2. Dead functions | `src/utils/string-similarity.ts` | -- |
| 3. EXEC_TIMEOUT_MS | `src/utils/claude-cli-resolver.ts`, `src/utils/opencode-cli-resolver.ts`, `src/tmux-manager.ts` | `src/config/exec-timeout.ts` |
| 4. z.infer types | `src/web/schemas.ts` | -- |
| 5. Weak tests | `test/task-tracker.test.ts`, `test/image-watcher.test.ts`, `test/task-queue.test.ts`, `test/hooks-config.test.ts`, `test/session-manager.test.ts` | -- |
**Total files modified**: 10
**Total files created**: 1
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -1,689 +0,0 @@
# Phase 6 Implementation Plan: Config Consolidation
**Source**: `docs/code-structure-findings.md` (Phase 6 — Config Consolidation)
**Estimated effort**: 1 day
**Tasks**: 8 tasks with dependencies (see dependency graph below)
---
## Safety Constraints
Before starting ANY work, read and follow these rules:
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
7. **Verify the dev server starts**: After each task, run `npx tsx src/index.ts web --port 3099 &` on a non-production port, confirm `curl -s http://localhost:3099/api/status | jq .status` returns `"ok"`, then kill the background process.
---
## Goal
Consolidate ~70 scattered numeric constants from 15+ source files into 6 new domain-focused config files, eliminating cross-file duplicates (including a 5x-duplicated AI model string) and making all tuning knobs discoverable in `src/config/`.
**Non-goal**: Moving every constant. Module-internal implementation details (like regex patterns, algorithm-specific magic numbers, or constants only used once in deeply coupled logic) stay where they are. The goal is discoverability of operational tuning knobs, not mechanical relocation.
---
## Design Decisions
### What gets centralized (and why)
Constants are candidates for centralization when they meet **any** of these criteria:
1. **Duplicated across files** — DRY violation (e.g., `STATS_COLLECTION_INTERVAL_MS` in `server.ts` and `mux-routes.ts`, AI model string in 5 files)
2. **Operational tuning knobs** — values an operator might want to adjust for performance, security, or behavior without understanding the implementation (e.g., SSE health check interval, auth session TTL, rate limits)
3. **Cross-cutting concerns** — values that establish system-wide contracts (e.g., max terminal dimensions used by both server routes and frontend)
### What stays in place (and why)
Constants that are **internal implementation details** of a single module stay where they are:
- **Algorithm parameters** — `TODO_SIMILARITY_THRESHOLD`, `adaptiveCompletionConfirmMs`, confidence weights. These are meaningless without understanding the algorithm.
- **Display/UI formatting** — `TEXT_PREVIEW_LENGTH`, `SMART_TITLE_MAX_LENGTH`, `COMMAND_DISPLAY_LENGTH` in `subagent-watcher.ts`. Only used locally, tightly coupled to rendering logic.
- **Module-internal timing** — `LINE_BUFFER_FLUSH_INTERVAL` in `session.ts`, `AI_CHECK_POLL_INTERVAL` in `ai-checker-base.ts`. Internal implementation of specific features.
- **Frontend constants** — `constants.js` already centralizes frontend values well. Don't mix frontend and backend config.
- **Respawn `DEFAULT_CONFIG`** — these are user-configurable defaults for the respawn config interface, not system constants. They live properly in `respawn-controller.ts`. The AI model/context defaults within it are replaced with imports from the new `ai-defaults.ts` (Task 5).
- **Session auto-ops thresholds** — `AUTO_RETRY_DELAY_MS`, `COMPACT_COOLDOWN_MS`, etc. in `session-auto-ops.ts` are internal to that module's retry logic and already well-documented in place.
### File organization: domain-based, not category-based
A single `timing-config.ts` with 70 unrelated timing values would be worse than the current state — developers would need to grep it just like they grep the whole codebase now. Instead, constants are grouped by **the system they configure**:
| New File | Domain | Developer Question It Answers |
|----------|--------|-------------------------------|
| `server-timing.ts` | Web server performance | "How do I tune SSE batching / terminal throughput?" |
| `auth-config.ts` | Authentication & security | "What are the rate limits and session TTLs?" |
| `tunnel-config.ts` | QR auth & Cloudflare tunnel | "What are the QR token rotation parameters?" |
| `terminal-limits.ts` | Terminal dimensions & input | "What are the max cols/rows/input size?" |
| `ai-defaults.ts` | AI checker model & context | "What model do the AI checkers use? What's the context limit?" |
| `team-config.ts` | Agent Teams polling & caching | "How often does team polling run? What are the cache limits?" |
---
## Task Dependencies
```
Task 1 (server-timing.ts)
Task 2 (auth-config.ts)
Task 3 (tunnel-config.ts)
Task 4 (terminal-limits.ts)
Task 5 (ai-defaults.ts)
Task 6 (team-config.ts)
└──> Task 7 (Fix remaining duplicates)
└──> Task 8 (Update CLAUDE.md + final verification)
```
**Tasks 1–6** are independent and can run in parallel.
**Task 7** depends on Tasks 1–6 (needs the new config files to exist).
**Task 8** depends on Task 7.
---
## Task 1: Create `src/config/server-timing.ts`
**Estimated effort**: 30 minutes
**Files created**: `src/config/server-timing.ts`
**Files modified**: `src/web/server.ts`, `src/web/routes/mux-routes.ts`
### Constants to extract from `src/web/server.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `TERMINAL_BATCH_INTERVAL` | `16` | Terminal data batching interval (60fps) |
| `TASK_UPDATE_BATCH_INTERVAL` | `100` | Task event batching interval (ms) |
| `STATE_UPDATE_DEBOUNCE_INTERVAL` | `500` | State persistence debounce (ms) |
| `SESSIONS_LIST_CACHE_TTL` | `1000` | Sessions list cache TTL (ms) |
| `SCHEDULED_CLEANUP_INTERVAL` | `300000` | Scheduled runs cleanup check (5 min) |
| `SCHEDULED_RUN_MAX_AGE` | `3600000` | Completed scheduled run max age (1 hour) |
| `SSE_HEALTH_CHECK_INTERVAL` | `30000` | SSE client health check (30s) |
| `SESSION_LIMIT_WAIT_MS` | `5000` | Session limit retry wait (5s) |
| `ITERATION_PAUSE_MS` | `2000` | Scheduled run iteration pause (2s) |
| `BATCH_FLUSH_THRESHOLD` | `32768` | Terminal batch immediate flush threshold (32KB) |
| `STATS_COLLECTION_INTERVAL_MS` | `2000` | Mux stats collection interval (2s) |
### Implementation
1. Create `src/config/server-timing.ts` with all 11 constants, preserving existing JSDoc comments.
2. In `src/web/server.ts`: Remove the 11 local constant declarations (lines ~92–121). Add `import { TERMINAL_BATCH_INTERVAL, ... } from '../config/server-timing.js'`.
3. In `src/web/routes/mux-routes.ts`: Remove the duplicate `STATS_COLLECTION_INTERVAL_MS` (line 10) and its comment. Add `import { STATS_COLLECTION_INTERVAL_MS } from '../../config/server-timing.js'`. This fixes a **duplicate constant** (finding #10).
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Web server performance and scheduling constants.
*
* Controls terminal batching throughput, SSE health checking,
* state persistence debouncing, and scheduled run timing.
*
* @module config/server-timing
*/
// ============================================================================
// Terminal & SSE Performance
// ============================================================================
/** Terminal data batching interval — targets 60fps (ms) */
export const TERMINAL_BATCH_INTERVAL = 16;
/** Immediate flush threshold for terminal batches (bytes).
* Set high (32KB) to allow effective batching; avg Ink events are ~14KB. */
export const BATCH_FLUSH_THRESHOLD = 32 * 1024;
/** Task event batching interval (ms) */
export const TASK_UPDATE_BATCH_INTERVAL = 100;
/** SSE client health check interval (ms) */
export const SSE_HEALTH_CHECK_INTERVAL = 30 * 1000;
// ============================================================================
// State Persistence
// ============================================================================
/** State update debounce — batches expensive toDetailedState() calls (ms) */
export const STATE_UPDATE_DEBOUNCE_INTERVAL = 500;
/** Sessions list cache TTL — avoids re-serializing on every SSE init (ms) */
export const SESSIONS_LIST_CACHE_TTL = 1000;
// ============================================================================
// Scheduled Runs
// ============================================================================
/** Scheduled runs cleanup check interval (ms) */
export const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
/** Completed scheduled run max age before cleanup (ms) */
export const SCHEDULED_RUN_MAX_AGE = 60 * 60 * 1000;
/** Session limit retry wait before retrying (ms) */
export const SESSION_LIMIT_WAIT_MS = 5000;
/** Pause between scheduled run iterations (ms) */
export const ITERATION_PAUSE_MS = 2000;
// ============================================================================
// Mux Stats
// ============================================================================
/** Mux stats collection interval (ms) */
export const STATS_COLLECTION_INTERVAL_MS = 2000;
```
### Verification
```bash
tsc --noEmit
npx tsx src/index.ts web --port 3099 &
curl -s http://localhost:3099/api/status | jq .status # "ok"
kill %1
```
---
## Task 2: Create `src/config/auth-config.ts`
**Estimated effort**: 20 minutes
**Files created**: `src/config/auth-config.ts`
**Files modified**: `src/web/middleware/auth.ts`, `src/hooks-config.ts`
### Constants to extract from `src/web/middleware/auth.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `AUTH_SESSION_TTL_MS` | `86400000` | Auth session cookie TTL (24h) |
| `MAX_AUTH_SESSIONS` | `100` | Max concurrent auth sessions |
| `AUTH_FAILURE_MAX` | `10` | Max failed auth attempts per IP |
| `AUTH_FAILURE_WINDOW_MS` | `900000` | Failed auth tracking window (15 min) |
### Constants to extract from `src/hooks-config.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `HOOK_TIMEOUT_MS` | `10000` | Timeout for Claude Code hook commands |
The `timeout: 10000` value is hardcoded 6 times in `hooks-config.ts` as inline literals. Extract to a single named constant.
### Implementation
1. Create `src/config/auth-config.ts` with the 5 constants.
2. In `src/web/middleware/auth.ts`: Remove the 4 local constant declarations (lines 17–25). Add import from `../../config/auth-config.js`. Keep `AUTH_COOKIE_NAME` in place — it's a string identifier, not a tunable numeric constant.
3. In `src/hooks-config.ts`: Replace all 6 inline `timeout: 10000` occurrences with `timeout: HOOK_TIMEOUT_MS`. Add import from `./config/auth-config.js`.
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Authentication, rate limiting, and hook security constants.
*
* Controls auth session lifecycle, brute-force protection,
* and Claude Code hook timeouts.
*
* @module config/auth-config
*/
// ============================================================================
// Session Cookies
// ============================================================================
/** Auth session cookie TTL — matches autonomous run length (ms) */
export const AUTH_SESSION_TTL_MS = 24 * 60 * 60 * 1000;
/** Max concurrent auth sessions per server */
export const MAX_AUTH_SESSIONS = 100;
// ============================================================================
// Rate Limiting
// ============================================================================
/** Max failed auth attempts per IP before 429 rejection */
export const AUTH_FAILURE_MAX = 10;
/** Failed auth attempt tracking window (ms) */
export const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
// ============================================================================
// Hooks
// ============================================================================
/** Timeout for Claude Code hook curl commands (ms) */
export const HOOK_TIMEOUT_MS = 10000;
```
### Verification
```bash
tsc --noEmit
npm run lint
```
---
## Task 3: Create `src/config/tunnel-config.ts`
**Estimated effort**: 20 minutes
**Files created**: `src/config/tunnel-config.ts`
**Files modified**: `src/tunnel-manager.ts`
### Constants to extract from `src/tunnel-manager.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `QR_TOKEN_TTL_MS` | `60000` | QR token auto-rotation interval (60s) |
| `QR_TOKEN_GRACE_MS` | `90000` | Grace period for previous token (90s) |
| `SHORT_CODE_LENGTH` | `6` | Length of QR short code |
| `QR_RATE_LIMIT_MAX` | `30` | Global QR attempt rate limit |
| `QR_RATE_LIMIT_WINDOW_MS` | `60000` | QR rate limit reset window (60s) |
| `URL_TIMEOUT_MS` | `30000` | Cloudflared URL fetch timeout (30s) |
| `RESTART_DELAY_MS` | `5000` | Tunnel restart delay after crash (5s) |
| `FORCE_KILL_MS` | `5000` | SIGTERM → SIGKILL escalation timeout (5s) |
### Implementation
1. Create `src/config/tunnel-config.ts` with all 8 constants.
2. In `src/tunnel-manager.ts`: Remove the 8 local constant declarations (lines ~39–75). Add `import { QR_TOKEN_TTL_MS, ... } from './config/tunnel-config.js'`.
3. Keep the `TUNNEL_URL_REGEX` in `tunnel-manager.ts` — it's a parsing detail, not a tuning knob.
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Cloudflare tunnel and QR authentication constants.
*
* Controls QR token rotation timing, rate limiting,
* and tunnel process lifecycle.
*
* @module config/tunnel-config
*/
// ============================================================================
// QR Token Rotation
// ============================================================================
/** QR token auto-rotation interval (ms) */
export const QR_TOKEN_TTL_MS = 60_000;
/** Grace period — previous token still valid during rotation (ms) */
export const QR_TOKEN_GRACE_MS = 90_000;
/** Length of the short code in QR URL path (chars) */
export const SHORT_CODE_LENGTH = 6;
// ============================================================================
// QR Rate Limiting
// ============================================================================
/** Global rate limit for QR auth attempts across all IPs */
export const QR_RATE_LIMIT_MAX = 30;
/** QR rate limit reset window (ms) */
export const QR_RATE_LIMIT_WINDOW_MS = 60_000;
// ============================================================================
// Tunnel Process Lifecycle
// ============================================================================
/** Max time to wait for cloudflared URL before timeout (ms) */
export const URL_TIMEOUT_MS = 30_000;
/** Restart delay after unexpected tunnel exit (ms) */
export const RESTART_DELAY_MS = 5_000;
/** SIGTERM → SIGKILL escalation timeout (ms) */
export const FORCE_KILL_MS = 5_000;
```
### Verification
```bash
tsc --noEmit
```
---
## Task 4: Create `src/config/terminal-limits.ts`
**Estimated effort**: 20 minutes
**Files created**: `src/config/terminal-limits.ts`
**Files modified**: `src/web/routes/session-routes.ts`
### Constants to extract from `src/web/routes/session-routes.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `MAX_INPUT_LENGTH` | `65536` | Max input length per request (64KB) |
| `MAX_TERMINAL_COLS` | `500` | Max terminal columns |
| `MAX_TERMINAL_ROWS` | `200` | Max terminal rows |
| `MAX_SESSION_NAME_LENGTH` | `128` | Max session name length (chars) |
### Why a separate file instead of adding to `buffer-limits.ts`
`buffer-limits.ts` covers memory buffer sizes (2MB terminal, 1MB text). These constants are **validation limits** for API inputs — different concern. A terminal resize request must not exceed `MAX_TERMINAL_COLS`; this has nothing to do with buffer trimming.
### Implementation
1. Create `src/config/terminal-limits.ts` with all 4 constants.
2. In `src/web/routes/session-routes.ts`: Remove the 4 local constant declarations (lines 45–48). Add `import { MAX_INPUT_LENGTH, MAX_TERMINAL_COLS, MAX_TERMINAL_ROWS, MAX_SESSION_NAME_LENGTH } from '../../config/terminal-limits.js'`.
3. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Terminal dimension and input validation limits.
*
* Used by API routes to validate resize, input, and session
* creation requests. Separate from buffer-limits.ts which
* controls memory buffer sizes.
*
* @module config/terminal-limits
*/
/** Max input length per API request (bytes) */
export const MAX_INPUT_LENGTH = 64 * 1024;
/** Max terminal columns for resize requests */
export const MAX_TERMINAL_COLS = 500;
/** Max terminal rows for resize requests */
export const MAX_TERMINAL_ROWS = 200;
/** Max session name length (chars) */
export const MAX_SESSION_NAME_LENGTH = 128;
```
### Verification
```bash
tsc --noEmit
```
---
## Task 5: Create `src/config/ai-defaults.ts`
**Estimated effort**: 30 minutes
**Files created**: `src/config/ai-defaults.ts`
**Files modified**: `src/respawn-controller.ts`, `src/ai-idle-checker.ts`, `src/ai-plan-checker.ts`, `src/web/routes/respawn-routes.ts`
### Problem: AI model string duplicated 5 times
The model identifier `'claude-opus-4-5-20251101'` appears in 5 places across 4 files. When the model changes, all 5 must be updated — a guaranteed source of bugs. The context limits (`16000`, `8000`) are similarly scattered across 3 files each.
| Constant | Current Value | Duplicated In |
|----------|---------------|---------------|
| `AI_CHECK_MODEL` | `'claude-opus-4-5-20251101'` | `respawn-controller.ts` (×2: idle + plan), `ai-idle-checker.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` (×2: idle + plan) |
| `AI_IDLE_CHECK_MAX_CONTEXT` | `16000` | `respawn-controller.ts`, `ai-idle-checker.ts`, `respawn-routes.ts` |
| `AI_PLAN_CHECK_MAX_CONTEXT` | `8000` | `respawn-controller.ts`, `ai-plan-checker.ts`, `respawn-routes.ts` |
### Implementation
1. Create `src/config/ai-defaults.ts` with the 3 constants.
2. In `src/respawn-controller.ts` `DEFAULT_CONFIG` (line 538): Replace `aiIdleCheckModel: 'claude-opus-4-5-20251101'` with `aiIdleCheckModel: AI_CHECK_MODEL`, `aiIdleCheckMaxContext: 16000` with `aiIdleCheckMaxContext: AI_IDLE_CHECK_MAX_CONTEXT`, `aiPlanCheckModel: 'claude-opus-4-5-20251101'` with `aiPlanCheckModel: AI_CHECK_MODEL`, `aiPlanCheckMaxContext: 8000` with `aiPlanCheckMaxContext: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
3. In `src/ai-idle-checker.ts` `DEFAULT_AI_CHECK_CONFIG` (line 46): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 16000` with `maxContextChars: AI_IDLE_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
4. In `src/ai-plan-checker.ts` `DEFAULT_PLAN_CHECK_CONFIG` (line 45): Replace `model: 'claude-opus-4-5-20251101'` with `model: AI_CHECK_MODEL`, `maxContextChars: 8000` with `maxContextChars: AI_PLAN_CHECK_MAX_CONTEXT`. Add import from `./config/ai-defaults.js`.
5. In `src/web/routes/respawn-routes.ts` config merge block (lines 173–179): Replace all 4 inline fallback values with imports from `../../config/ai-defaults.js`.
6. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Default model and context limits for AI-powered checkers.
*
* Centralizes the AI model identifier and context window sizes used by
* the idle checker, plan checker, respawn controller defaults, and
* respawn route fallbacks. Change the model here when upgrading.
*
* @module config/ai-defaults
*/
/** Default model for AI idle and plan checkers */
export const AI_CHECK_MODEL = 'claude-opus-4-5-20251101';
/** Max context chars for idle checker (~4k tokens) */
export const AI_IDLE_CHECK_MAX_CONTEXT = 16000;
/** Max context chars for plan checker (~2k tokens, plan mode UI is compact) */
export const AI_PLAN_CHECK_MAX_CONTEXT = 8000;
```
### Verification
```bash
tsc --noEmit
# Verify no remaining hardcoded model strings
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
```
---
## Task 6: Create `src/config/team-config.ts`
**Estimated effort**: 15 minutes
**Files created**: `src/config/team-config.ts`
**Files modified**: `src/team-watcher.ts`
### Constants to extract from `src/team-watcher.ts`
| Constant | Value | Purpose |
|----------|-------|---------|
| `TEAM_POLL_INTERVAL_MS` | `30000` | Team directory poll interval (30s) |
| `MAX_CACHED_TEAMS` | `50` | LRU cache size for team configs |
| `MAX_CACHED_TASKS` | `200` | LRU cache size for team tasks + inboxes |
### Why centralize these
Team polling frequency and cache sizes are operational knobs that affect both performance (polling too often wastes CPU) and responsiveness (polling too rarely means stale team state in the UI). They're also the kind of values a developer tuning for a large team deployment would want to find quickly. `MAX_CACHED_TASKS` is used for both the task cache and inbox cache — worth documenting.
### Implementation
1. Create `src/config/team-config.ts` with the 3 constants.
2. In `src/team-watcher.ts`: Remove the 3 local constants (lines 23–25). Add `import { TEAM_POLL_INTERVAL_MS, MAX_CACHED_TEAMS, MAX_CACHED_TASKS } from './config/team-config.js'`. Note: rename `POLL_INTERVAL_MS` → `TEAM_POLL_INTERVAL_MS` to avoid ambiguity with the identically-named constant in `subagent-watcher.ts`.
3. Update the usage site: `setInterval(... POLL_INTERVAL_MS)` → `setInterval(... TEAM_POLL_INTERVAL_MS)`.
4. Run `tsc --noEmit`.
### New file template
```typescript
/**
* @fileoverview Agent Teams polling and cache configuration.
*
* Controls how frequently TeamWatcher polls ~/.claude/teams/
* and how many teams/tasks are cached in memory.
*
* @module config/team-config
*/
/** Team directory poll interval (ms) */
export const TEAM_POLL_INTERVAL_MS = 30_000;
/** Max cached team configs (LRU eviction) */
export const MAX_CACHED_TEAMS = 50;
/** Max cached team tasks and inbox messages (LRU eviction).
* Used for both teamTasks and inboxCache maps. */
export const MAX_CACHED_TASKS = 200;
```
### Verification
```bash
tsc --noEmit
```
---
## Task 7: Fix remaining cross-file duplicates
**Estimated effort**: 30 minutes
**Files modified**: `src/index.ts`, `src/subagent-watcher.ts`
### Duplicate 1: `STATS_COLLECTION_INTERVAL_MS`
Already fixed in Task 1 — both `server.ts` and `mux-routes.ts` now import from `server-timing.ts`.
### Duplicate 2: AI model string
Already fixed in Task 5 — all 5 occurrences now import from `ai-defaults.ts`.
### Duplicate 3: `MAX_SCREENSHOT_SIZE` / `MAX_TEXT_FILE_SIZE` / `MAX_RAW_FILE_SIZE`
These file size limits in `file-routes.ts` and `system-routes.ts` are **API-specific validation limits**. They're only used in their respective route files and aren't duplicated. **Leave in place** — they're local to their route module and well-commented.
### Action A: Move `MAX_CONSECUTIVE_ERRORS` and `ERROR_RESET_MS` to config
`src/index.ts` has two process-level constants that are operational tuning knobs:
| Constant | Value | Purpose |
|----------|-------|---------|
| `MAX_CONSECUTIVE_ERRORS` | `5` | Max consecutive unhandled errors before process exit |
| `ERROR_RESET_MS` | `60000` | Error counter reset interval (1 min) |
These belong in a config file since they control server reliability behavior. Add them to `src/config/server-timing.ts` (they're server operational constants).
1. Add to `src/config/server-timing.ts`:
```typescript
// ============================================================================
// Process Error Recovery
// ============================================================================
/** Max consecutive unhandled errors before auto-restart */
export const MAX_CONSECUTIVE_ERRORS = 5;
/** Error counter reset interval — forgives errors after quiet period (ms) */
export const ERROR_RESET_MS = 60_000;
```
2. In `src/index.ts`: Remove lines 19–20, add import from `./config/server-timing.js`.
3. Run `tsc --noEmit`.
### Action B: Fix `MAX_TRACKED_AGENTS` shadow in `subagent-watcher.ts`
`subagent-watcher.ts` defines its own `MAX_TRACKED_AGENTS = 500` locally instead of importing the identical value from `config/map-limits.ts`. This is a latent bug — if someone changes the config value, the subagent watcher's copy stays stale.
1. In `src/subagent-watcher.ts`: Remove the local `MAX_TRACKED_AGENTS` constant. Add `import { MAX_TRACKED_AGENTS } from './config/map-limits.js'` (the value there is `MAX_TODOS_PER_SESSION = 500` — **verify** the map-limits constant is actually named `MAX_TRACKED_AGENTS` or if it needs to be added). If the constant doesn't exist in `map-limits.ts` under that name, add it.
2. Run `tsc --noEmit`.
### Verification
```bash
tsc --noEmit
npm run lint
npm run format:check
```
---
## Task 8: Update CLAUDE.md and final verification
**Estimated effort**: 20 minutes
**Files modified**: `CLAUDE.md`
### Updates to CLAUDE.md
1. **Config Files table** (`src/config/`): Add the 6 new files:
| File | Purpose |
|------|---------|
| `buffer-limits.ts` | Terminal/text buffer size limits |
| `map-limits.ts` | Global limits for Maps, sessions, watchers |
| `exec-timeout.ts` | Execution timeout configuration |
| `server-timing.ts` | Web server batching, SSE, scheduled run timing |
| `auth-config.ts` | Auth session TTL, rate limits, hook timeout |
| `tunnel-config.ts` | QR token rotation, tunnel process lifecycle |
| `terminal-limits.ts` | Terminal dimension and input validation limits |
| `ai-defaults.ts` | AI checker model and context limits |
| `team-config.ts` | Agent Teams polling and cache sizes |
2. **Import Conventions** section: Add:
```
- **Config**: Import from specific files: `import { MAX_TERMINAL_COLS } from './config/terminal-limits'`
```
3. **Phase 6 status** in `docs/code-structure-findings.md`: Mark as COMPLETE with summary of what was done.
### Final verification checklist
```bash
# Type checking
tsc --noEmit
# Linting
npm run lint
# Formatting
npm run format:check
# Dev server starts
npx tsx src/index.ts web --port 3099 &
curl -s http://localhost:3099/api/status | jq .status # "ok"
kill %1
# Verify no remaining duplicates
grep -rn 'STATS_COLLECTION_INTERVAL_MS' src/ # Should only appear in config + import sites
grep -rn 'timeout: 10000' src/hooks-config.ts # Should be 0 — all replaced with HOOK_TIMEOUT_MS
grep -rn 'claude-opus-4-5-20251101' src/ # Should only appear in config/ai-defaults.ts
```
---
## What is NOT in scope (and why)
These constants were considered but deliberately left in their current files:
### Respawn controller defaults (`src/respawn-controller.ts`)
The `DEFAULT_CONFIG` object (lines 538–578) contains ~30 default values for the `RespawnConfig` interface. These are **user-facing configuration defaults**, not system constants — they're the starting values for a config object that users can modify via the API and UI. Centralizing them would break the locality between the config interface definition and its defaults. They already have excellent JSDoc with `@default` tags. The only values extracted are the AI model/context constants (Task 5) which are duplicated in other files.
### Subagent watcher timing (`src/subagent-watcher.ts`)
The 18 constants at lines 129–158 are all internal to the subagent watcher's polling/lifecycle algorithm. Moving them to a config file would force developers to context-switch between two files to understand the polling logic. They're already grouped with clear comments. Exception: `MAX_TRACKED_AGENTS` is consolidated with `map-limits.ts` (Task 7B) since it duplicates a global limit.
### Session auto-ops timing (`src/session-auto-ops.ts`)
The 8 constants at lines 19–40 are internal to the auto-compact/clear retry state machine. They form a coherent group that's meaningless without the surrounding implementation context.
### Run summary constants (`src/run-summary.ts`)
`MAX_EVENTS`, `TRIM_TO_EVENTS`, `TOKEN_MILESTONE_INTERVAL`, `STATE_STUCK_WARNING_MS`, `STATE_STUCK_CHECK_INTERVAL` — all module-internal. The buffer-style limits (`MAX_EVENTS`/`TRIM_TO_EVENTS`) follow the same pattern as `buffer-limits.ts` but are only used in this one file.
### Frontend (`src/web/public/constants.js`)
Already well-centralized. Frontend and backend run in different environments — mixing them in TypeScript config files would create import problems. If frontend constants need expansion, do it in `constants.js`. Note: `app.js` has 2 inline uses of `256 * 1024` that should use the existing `TERMINAL_TAIL_SIZE` from `constants.js` — a minor cleanup that can be done opportunistically but is not worth a task here.
### Tmux manager timing (`src/tmux-manager.ts`)
The 6 constants (lines 65–78) are internal to tmux process lifecycle management. They're low-level retry/wait values that are meaningless without understanding the tmux spawn sequence.
### Process-internal constants
`image-watcher.ts`, `bash-tool-parser.ts`, `transcript-watcher.ts`, `ralph-tracker.ts`, `task-tracker.ts`, `file-stream-manager.ts`, `session-lifecycle-log.ts`, `session-task-cache.ts`, `respawn-metrics.ts`, `respawn-adaptive-timing.ts`, `ai-checker-base.ts` — all have module-local constants that are internal implementation details.
### `localhost:3000` default URL
The string `'http://localhost:3000'` or port `3000` appears as a fallback default in ~5 files (`session-cli-builder.ts`, `tmux-manager.ts`, `tunnel-manager.ts`, `server.ts`, CLI). While technically duplicated, extracting it provides little value — each usage has a different fallback chain (env var → config → hardcoded) and the port is also baked into systemd service files and documentation. The risk of a missed update is low since port 3000 is deeply conventional.
### `SAVE_DEBOUNCE_MS = 500` in `state-store.ts` / `push-store.ts`
Same value (500ms), but they debounce different persistence targets (state.json vs push-subscriptions.json). If one needed faster/slower debouncing, they'd diverge. Coupling them would be misleading.
---
## Summary
| Metric | Before | After |
|--------|--------|-------|
| Config files in `src/config/` | 3 | 9 |
| Constants centralized | ~25 | ~65 |
| Cross-file duplicates | 9+ (`STATS_COLLECTION_INTERVAL_MS`, `timeout: 10000` ×6, AI model ×5, context limits ×3 each, `MAX_TRACKED_AGENTS`) | 0 |
| Files with `timeout: 10000` inline | 1 (6 occurrences) | 0 |
| Files with hardcoded AI model string | 4 (5 occurrences) | 1 (config only) |
| Files modified | — | 11 |
| Files created | — | 6 |
@@ -1,953 +0,0 @@
# Phase 7 Implementation Plan: Test Infrastructure
**Source**: `docs/code-structure-findings.md` (Phase 7 — Test Infrastructure)
**Estimated effort**: 2–3 days
**Tasks**: 11 tasks with dependencies (see dependency graph below)
---
## Safety Constraints
Before starting ANY work, read and follow these rules:
1. **Never run `npx vitest run`** (full suite) — it kills tmux sessions. You are running inside a Codeman-managed tmux session.
2. **Run individual tests only**: `npx vitest run test/<file>.test.ts`
3. **Never test on port 3000** — the live dev server runs there. Tests use ports 3150+.
4. **After TypeScript changes**: Run `tsc --noEmit` to verify type checking passes.
5. **Before considering done**: Run `npm run lint` and `npm run format:check` to ensure CI passes.
6. **Never kill tmux sessions** — check `echo $CODEMAN_MUX` first.
7. **Port assignments for this phase**: New tests use ports 3220–3229 (see individual tasks for assignments).
---
## Goal
Eliminate duplicated test mocks, activate the unused `respawn-test-utils.ts` utilities, and add route-level test coverage for the server's 12 route modules — the single largest untested area in the codebase (162 route handlers, 0 dedicated tests).
**Non-goals**:
- Full end-to-end integration tests (those require real Claude CLI / tmux sessions)
- 100% route coverage in this phase — focus on the highest-value route modules first
- Refactoring test patterns in existing passing tests that don't use shared mocks
- Migrating `vi.mock()`-based module replacement mocks (different pattern, see Task 6/7)
---
## Current State
### Mock Duplication (Finding #9)
`MockSession` is defined **4 times** across test files with varying levels of completeness:
| File | Properties | Methods | EventEmitter | Notes |
|------|-----------|---------|-------------|-------|
| `test/respawn-test-utils.ts` | 6 | 20+ | Yes | **Most complete**. Includes terminal simulation, token count, ANSI output, plan mode prompts. **Never imported by any test.** |
| `test/respawn-controller.test.ts` | 6 | 9 | Yes | Subset of respawn-test-utils. Missing token simulation, ANSI helpers. |
| `test/respawn-team-awareness.test.ts` | ~6 | ~9 | Yes | Near-copy of respawn-controller.test.ts version. |
| `test/session-manager.test.ts` | 4 | 8 | Yes | **Inside `vi.mock()` factory** — replaces `../src/session.js` module. Different shape: `start()`/`stop()`/`toState()`/`sendInput()` for lifecycle testing. |
`MockStateStore` is defined **2 times** (both inside `vi.mock()` factories):
| File | Shape | Methods | Mock Pattern |
|------|-------|---------|-------------|
| `test/session-manager.test.ts` | `{ sessions, config }` | `getConfig`, `getSessions`, `getSession`, `setSession`, `removeSession` | `vi.mock('../src/state-store.js')` |
| `test/ralph-loop.test.ts` | `{ ralphLoop, tasks, config }` | `getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask` | `vi.mock('../src/state-store.js')` |
### Important: Two distinct mocking patterns
The codebase uses two different mocking patterns that require different migration strategies:
1. **Direct instantiation** (respawn-controller, respawn-team-awareness): `MockSession` is defined at file scope and instantiated directly in tests. These can be migrated to shared mocks via simple import replacement.
2. **Module replacement** (session-manager, ralph-loop): Mocks are defined inside `vi.mock()` factories that replace entire modules (`../src/session.js`, `../src/state-store.js`). These factories run in an isolated scope and return `{ Session: MockClass }` or `{ getStore: vi.fn(() => instance) }`. Migrating these requires either `vi.hoisted()` or restructuring the test's module mocking — higher risk for limited benefit.
### Unused Test Utilities
`test/respawn-test-utils.ts` exports these utilities that **no test file imports**:
- `TimeController` / `createTimeController()` — abstraction over vitest fake timers
- `MockAiIdleChecker` / `MockAiPlanChecker` — fully mocked AI checkers with result queueing
- `createStateTracker()` / `createEventRecorder()` — state transition and event recording
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — pre-configured RespawnConfig objects
- `waitForState()` / `waitForEvent()` / `createDeferred()` — async test helpers
- `terminalOutputs` — factory object for common terminal output patterns
### Route Test Coverage
Currently **zero** dedicated tests for the 12 route modules in `src/web/routes/`. The existing test files that touch API endpoints:
| Test File | What It Tests | Approach |
|-----------|--------------|----------|
| `test/api-responses.test.ts` | Response structure validation | Imports types, no HTTP calls |
| `test/api-generate-plan.test.ts` | Plan generation API | Mocks validation logic, Port 3191 declared |
| `test/auth-security.test.ts` | Auth middleware | Integration tests with WebServer, Ports 3160/3161 |
| `test/qr-auth.test.ts` | QR authentication | Integration + unit tests, Port 3162 |
None of these test the route handlers themselves with real HTTP requests against a running Fastify instance.
---
## Design Decisions
### Shared mocks: Superset strategy
Rather than creating a lowest-common-denominator mock, `MockSession` in `test/mocks/` will be the **superset** from `respawn-test-utils.ts` (the most complete version). Test files that need a simpler mock can just ignore the extra methods — having unused methods costs nothing, but missing methods forces local re-definition.
### vi.mock() tests: Don't migrate
The `session-manager.test.ts` and `ralph-loop.test.ts` tests define mocks inside `vi.mock()` factories. These use **module-level replacement** (replacing `../src/session.js` and `../src/state-store.js` entirely), which is fundamentally different from the direct-instantiation pattern. Migrating them would require `vi.hoisted()` or factory restructuring — high complexity for limited benefit since these mocks are already working. We leave these as-is and create the shared mocks for **new** tests and for the two direct-instantiation tests (Tasks 4–5).
### MockStateStore: Union of both shapes
The shared `MockStateStore` in `test/mocks/` will include methods from both existing definitions (session management + Ralph loop), so any **new** test can use it. Methods default to no-ops via `vi.fn()`. Existing `vi.mock()`-based tests are not migrated.
### Route testing strategy: Lightweight Fastify instances
Each route test file will:
1. Create a minimal `Fastify` instance
2. Register **only** the route module under test
3. Provide a mock context object satisfying the port interfaces
4. Use `app.inject()` (Fastify's built-in test helper) — no real HTTP, no port needed
This avoids port conflicts entirely and runs fast. Only tests that need SSE or WebSocket behavior will use a real listening server with assigned ports.
### Port assignments (for tests needing real servers)
| Port | Test File | Purpose |
|------|-----------|---------|
| 3220 | `test/routes/session-routes.test.ts` | SSE integration (if needed) |
| 3221 | `test/routes/system-routes.test.ts` | Status/stats endpoints |
| 3222 | `test/routes/respawn-routes.test.ts` | Respawn API |
| 3223 | `test/routes/ralph-routes.test.ts` | Ralph API |
| 3224–3229 | Reserved | Future route tests |
Most tests should NOT need real ports — `app.inject()` is preferred. Verified: ports 3220–3229 are completely unused by existing tests (highest used port is 3211 in `opencode-resize.test.ts`).
---
## Task Dependencies
```
Task 1 (Consolidate MockSession)
Task 2 (Consolidate MockStateStore)
└──> Task 3 (Create test/mocks/ barrel)
├──> Task 4 (Migrate respawn-controller.test.ts)
├──> Task 5 (Migrate respawn-team-awareness.test.ts)
└──> Task 6 (Route test scaffold + helpers)
├──> Task 7 (Session routes tests)
└──> Task 8 (System + respawn routes tests)
Task 9 (Slim down respawn-test-utils.ts) — depends on Tasks 4, 5
```
**Tasks 1–2** are independent and can run in parallel.
**Task 3** depends on Tasks 1–2.
**Tasks 4–6** depend on Task 3 and can run in parallel.
**Tasks 7–8** depend on Task 6 and can run in parallel.
**Task 9** depends on Tasks 4, 5 (must verify migrations work before removing duplicates from source).
---
## Task 1: Consolidate MockSession into `test/mocks/mock-session.ts`
**Estimated effort**: 2 hours
**Files created**: `test/mocks/mock-session.ts`
**Files modified**: None yet (consumers migrate in Tasks 4–5)
### Source
The canonical MockSession comes from `test/respawn-test-utils.ts` (lines 89–241). It is the most complete version with:
- All properties needed by `RespawnController`: `id`, `workingDir`, `status`, `writeBuffer`, `terminalBuffer`, `muxName`
- `write()` / `writeViaMux()` for input simulation
- Buffer inspection: `lastWrite`, `hasWritten(pattern)`, `clearWriteBuffer()`
- Terminal simulation: `simulateTerminalOutput()`, `simulatePrompt()`, `simulateReady()`, `simulateCompletionMessage()`, `simulateWorking()`, `simulateClearComplete()`, `simulateInitComplete()`, `simulatePlanModePrompt()`, `simulateElicitationDialog()`, `simulateTokenCount()`, `simulateAnsiOutput()`
- Lifecycle: `close()`
### Implementation
1. Create `test/mocks/` directory.
2. Create `test/mocks/mock-session.ts`:
- Copy the `MockSession` class **exactly** from `test/respawn-test-utils.ts` (lines 89–241)
- Copy `terminalOutputs` helper object (tightly coupled to mock)
- Copy `createMockSession()` factory function
- Export all three: `export { MockSession, createMockSession, terminalOutputs }`
- Ensure all `vi` imports come from `vitest`
**CRITICAL**: Copy the source verbatim — do NOT rewrite the simulation methods. The respawn controller's detection logic matches specific output patterns (e.g., `'\u276f '` for prompt, `'\u273b Worked for'` for completion). Using different patterns would cause test failures.
### Template
```typescript
/**
* Shared MockSession for tests that need terminal simulation.
*
* Copied from test/respawn-test-utils.ts (the canonical, most complete version).
* Used by respawn, route, and subagent tests.
*/
import { EventEmitter } from 'node:events';
// Copy MockSession class exactly from test/respawn-test-utils.ts lines 89–241
export class MockSession extends EventEmitter {
// ... (copy verbatim from respawn-test-utils.ts)
}
/**
* Factory for common terminal output strings.
* Must match the patterns used in MockSession's simulate* methods.
*/
export const terminalOutputs = {
// ... (copy verbatim from respawn-test-utils.ts)
};
/**
* Convenience factory.
*/
export function createMockSession(id?: string): MockSession {
return new MockSession(id);
}
```
### Verification
```bash
tsc --noEmit # Ensure file compiles
```
---
## Task 2: Consolidate MockStateStore into `test/mocks/mock-state-store.ts`
**Estimated effort**: 1 hour
**Files created**: `test/mocks/mock-state-store.ts`
**Files modified**: None (existing vi.mock()-based tests are NOT migrated; this is for new route tests)
### Source
Union of both existing definitions:
- From `test/session-manager.test.ts`: session CRUD methods (`getConfig`, `getSession`, `setSession`, `removeSession`, `getSessions`)
- From `test/ralph-loop.test.ts`: Ralph state methods (`getConfig`, `getRalphLoopState`, `setRalphLoopState`, `getTasks`, `setTask`, `removeTask`)
### Template
```typescript
/**
* Shared MockStateStore for tests.
*
* Includes methods for both session management and Ralph loop testing.
* All methods are vi.fn() spies — tests can override return values as needed.
*
* NOTE: This is for direct instantiation in new tests. Existing tests that
* use vi.mock('../src/state-store.js') keep their inline definitions.
*/
import { vi } from 'vitest';
export class MockStateStore {
state: Record<string, unknown> = {
sessions: {} as Record<string, unknown>,
config: { maxConcurrentSessions: 5 },
ralphLoop: { status: 'stopped' },
tasks: {} as Record<string, unknown>,
};
// Session methods
getConfig = vi.fn(() => this.state.config);
getSessions = vi.fn(() => this.state.sessions as Record<string, unknown>);
getSession = vi.fn((id: string) => (this.state.sessions as Record<string, unknown>)[id]);
setSession = vi.fn((id: string, state: unknown) => {
(this.state.sessions as Record<string, unknown>)[id] = state;
});
removeSession = vi.fn((id: string) => {
delete (this.state.sessions as Record<string, unknown>)[id];
});
// Ralph state methods
getRalphLoopState = vi.fn(() => this.state.ralphLoop);
setRalphLoopState = vi.fn((update: Record<string, unknown>) => {
this.state.ralphLoop = { ...(this.state.ralphLoop as Record<string, unknown>), ...update };
});
// Task methods
getTasks = vi.fn(() => this.state.tasks);
setTask = vi.fn();
removeTask = vi.fn();
// Settings methods
getSettings = vi.fn(() => ({}));
setSettings = vi.fn();
// Generic persistence
save = vi.fn();
load = vi.fn();
/** Reset all state and mocks for clean test isolation */
reset(): void {
this.state = {
sessions: {},
config: { maxConcurrentSessions: 5 },
ralphLoop: { status: 'stopped' },
tasks: {},
};
vi.clearAllMocks();
}
}
```
### Verification
```bash
tsc --noEmit
```
---
## Task 3: Create `test/mocks/index.ts` barrel export
**Estimated effort**: 30 minutes
**Depends on**: Tasks 1, 2
**Files created**: `test/mocks/index.ts`, `test/mocks/test-helpers.ts`
**Files modified**: None
### Implementation
1. Create `test/mocks/test-helpers.ts` with the async utilities from `respawn-test-utils.ts`:
```typescript
/**
* Reusable async test helpers.
* Extracted from respawn-test-utils.ts.
*/
/** Wait for an EventEmitter to emit a specific event, with timeout */
export function waitForEvent(
emitter: { once: (event: string, listener: (...args: unknown[]) => void) => void },
event: string,
timeoutMs = 5000,
): Promise<unknown> {
return new Promise((resolve, reject) => {
const timer = setTimeout(
() => reject(new Error(`Timed out waiting for event "${event}" after ${timeoutMs}ms`)),
timeoutMs,
);
emitter.once(event, (...args: unknown[]) => {
clearTimeout(timer);
resolve(args.length === 1 ? args[0] : args);
});
});
}
/** Create a deferred promise with external resolve/reject */
export function createDeferred<T = void>(): {
promise: Promise<T>;
resolve: (value: T) => void;
reject: (reason?: unknown) => void;
} {
let resolve!: (value: T) => void;
let reject!: (reason?: unknown) => void;
const promise = new Promise<T>((res, rej) => {
resolve = res;
reject = rej;
});
return { promise, resolve, reject };
}
```
2. Create `test/mocks/index.ts` barrel:
```typescript
/**
* Shared test mocks — import from here instead of defining inline.
*
* @example
* import { MockSession, MockStateStore, terminalOutputs } from './mocks/index.js';
*/
export { MockSession, createMockSession, terminalOutputs } from './mock-session.js';
export { MockStateStore } from './mock-state-store.js';
export { waitForEvent, createDeferred } from './test-helpers.js';
```
### Verification
```bash
tsc --noEmit
```
---
## Task 4: Migrate `respawn-controller.test.ts` to shared mocks
**Estimated effort**: 30 minutes
**Depends on**: Task 3
**Files modified**: `test/respawn-controller.test.ts`
### Steps
1. Remove the local `MockSession` class definition (approx. 50 lines).
2. Add: `import { MockSession } from './mocks/index.js';`
3. Verify all test methods still exist on the shared mock. The shared mock is a superset, so all existing usage should work.
4. If the local mock had any test-specific customizations (e.g., extra properties added in `beforeEach`), keep those in the test file as inline assignments on the shared instance.
5. Run the test to confirm it passes.
### Potential issues
- The local mock's `simulateCompletionMessage()` may have a slightly different output format than the shared mock's (from respawn-test-utils.ts). Verify the respawn controller's completion detection regex matches the shared mock's output pattern (`'\u273b Worked for ...'`).
- If the local mock adds `pid` or `isWorking` properties that the shared mock doesn't have, add inline assignments in `beforeEach`.
### Verification
```bash
npx vitest run test/respawn-controller.test.ts
```
---
## Task 5: Migrate `respawn-team-awareness.test.ts` to shared mocks
**Estimated effort**: 30 minutes
**Depends on**: Task 3
**Files modified**: `test/respawn-team-awareness.test.ts`
### Steps
1. Remove the local `MockSession` class definition.
2. Add: `import { MockSession } from './mocks/index.js';`
3. Keep `MockTeamWatcher` in this file — it's test-specific and extends the real `TeamWatcher`, not a general-purpose mock.
4. Run the test to confirm it passes.
### Verification
```bash
npx vitest run test/respawn-team-awareness.test.ts
```
---
## Task 6: Create route test scaffold and helpers
**Estimated effort**: 2 hours
**Depends on**: Task 3
**Files created**: `test/mocks/mock-route-context.ts`, `test/routes/` directory, `test/routes/_route-test-utils.ts`
### Problem
The 12 route modules in `src/web/routes/` have zero dedicated test coverage. Each route module takes `(app: FastifyInstance, ctx: PortIntersection)` — we need a reusable way to create mock context objects that satisfy the port interfaces.
### Design
Create a `MockRouteContext` factory that builds a mock object satisfying all port interfaces. Each port's methods are `vi.fn()` stubs. Tests can override specific methods as needed.
### Route registration signatures (verified)
Each route module requires a specific port intersection. The mock must satisfy all of them:
| Route Module | Required Ports |
|-------------|----------------|
| `registerSessionRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
| `registerSystemRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort & AuthPort` |
| `registerRespawnRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
| `registerRalphRoutes` | `SessionPort & EventPort & RespawnPort & ConfigPort & InfraPort` |
| `registerPlanRoutes` | `SessionPort & EventPort & ConfigPort & InfraPort` |
| `registerCaseRoutes` | `EventPort & ConfigPort` |
| `registerScheduledRoutes` | `SessionPort & EventPort & InfraPort` |
| `registerFileRoutes` | `SessionPort` |
| `registerMuxRoutes` | `InfraPort` |
| `registerPushRoutes` | `InfraPort` |
| `registerTeamRoutes` | `InfraPort` |
| `registerHookEventRoutes` | `EventPort & AuthPort` |
### Implementation
1. Create `test/mocks/mock-route-context.ts`:
```typescript
/**
* Mock context for route handler testing.
*
* Satisfies ALL port interfaces (SessionPort, EventPort, RespawnPort,
* ConfigPort, InfraPort, AuthPort) so any route module can be tested.
* Override specific methods in individual tests as needed.
*
* Verified against actual port interfaces in src/web/ports/:
* - SessionPort: 6 methods (sessions, addSession, cleanupSession,
* setupSessionListeners, persistSessionState, persistSessionStateNow,
* getSessionStateWithRespawn)
* - EventPort: 5 methods (broadcast, sendPushNotifications, batchTerminalData,
* broadcastSessionStateDebounced, batchTaskUpdate)
* - RespawnPort: 2 maps + 4 methods
* - ConfigPort: 5 readonly + 7 methods (incl getDefaultClaudeMdPath,
* getLightState, getLightSessionsState, stopTranscriptWatcher)
* - InfraPort: 7 readonly + 2 methods (startScheduledRun, stopScheduledRun)
* - AuthPort: 3 readonly (authSessions, qrAuthFailures, https)
*/
import { vi } from 'vitest';
import { MockSession, createMockSession } from './mock-session.js';
/**
* Creates a mock context that satisfies all port interfaces.
* Pre-populated with one session for convenience.
*/
export function createMockRouteContext(options?: { sessionId?: string }) {
const sessionId = options?.sessionId ?? 'test-session-1';
const session = createMockSession(sessionId);
const sessions = new Map<string, MockSession>();
sessions.set(sessionId, session);
return {
// -- SessionPort --
sessions,
addSession: vi.fn(),
cleanupSession: vi.fn(),
setupSessionListeners: vi.fn(),
persistSessionState: vi.fn(),
persistSessionStateNow: vi.fn(),
getSessionStateWithRespawn: vi.fn((s: unknown) => s),
// -- EventPort --
broadcast: vi.fn(),
sendPushNotifications: vi.fn(),
batchTerminalData: vi.fn(),
broadcastSessionStateDebounced: vi.fn(),
batchTaskUpdate: vi.fn(),
// -- RespawnPort --
respawnControllers: new Map(),
respawnTimers: new Map(),
setupRespawnListeners: vi.fn(),
setupTimedRespawn: vi.fn(),
restoreRespawnController: vi.fn(),
saveRespawnConfig: vi.fn(),
// -- ConfigPort --
store: {
getConfig: vi.fn(() => ({})),
getSessions: vi.fn(() => ({})),
getSession: vi.fn(),
setSession: vi.fn(),
removeSession: vi.fn(),
getSettings: vi.fn(() => ({})),
setSettings: vi.fn(),
getRalphLoopState: vi.fn(() => ({})),
setRalphLoopState: vi.fn(),
getTasks: vi.fn(() => ({})),
save: vi.fn(),
load: vi.fn(),
},
port: 3000,
https: false,
testMode: true,
serverStartTime: Date.now(),
getGlobalNiceConfig: vi.fn(async () => undefined),
getModelConfig: vi.fn(async () => null),
getClaudeModeConfig: vi.fn(async () => ({})),
getDefaultClaudeMdPath: vi.fn(async () => undefined),
getLightState: vi.fn(() => ({ sessions: [], status: 'ok' })),
getLightSessionsState: vi.fn(() => []),
startTranscriptWatcher: vi.fn(),
stopTranscriptWatcher: vi.fn(),
// -- InfraPort --
mux: {
createSession: vi.fn(),
killSession: vi.fn(),
listSessions: vi.fn(() => []),
getStats: vi.fn(() => ({})),
},
runSummaryTrackers: new Map(),
activePlanOrchestrators: new Map(),
scheduledRuns: new Map(),
teamWatcher: { getTeams: vi.fn(() => []), hasActiveTeammates: vi.fn(() => false) },
tunnelManager: null,
pushStore: null,
startScheduledRun: vi.fn(),
stopScheduledRun: vi.fn(),
// -- AuthPort --
authSessions: null,
qrAuthFailures: null,
// https already declared above in ConfigPort (shared property)
// Convenience accessors (not part of any port interface)
_session: session,
_sessionId: sessionId,
};
}
export type MockRouteContext = ReturnType<typeof createMockRouteContext>;
```
2. Add to `test/mocks/index.ts` barrel:
```typescript
export { createMockRouteContext, type MockRouteContext } from './mock-route-context.js';
```
3. Create `test/routes/` directory for route test files.
4. Create `test/routes/_route-test-utils.ts` with Fastify test helpers:
```typescript
/**
* Shared utilities for route testing.
*
* Creates minimal Fastify instances with just the route module under test
* and a mock context. Uses app.inject() for HTTP testing without real ports.
*/
import Fastify, { type FastifyInstance } from 'fastify';
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
export interface RouteTestHarness {
app: FastifyInstance;
ctx: MockRouteContext;
}
/**
* Creates a Fastify instance with a route module registered against a mock context.
*
* @param registerFn - The route registration function (e.g., registerSessionRoutes).
* Uses `any` for ctx parameter because route functions expect typed port intersections
* that MockRouteContext satisfies structurally but not nominally.
* @param ctxOptions - Optional overrides for the mock context
*/
export async function createRouteTestHarness(
// eslint-disable-next-line @typescript-eslint/no-explicit-any
registerFn: (app: FastifyInstance, ctx: any) => void,
ctxOptions?: { sessionId?: string },
): Promise<RouteTestHarness> {
const app = Fastify({ logger: false });
const ctx = createMockRouteContext(ctxOptions);
registerFn(app, ctx);
await app.ready();
return { app, ctx };
}
```
### Why `ctx: any` in the harness
Route registration functions like `registerSessionRoutes(app, ctx: SessionPort & EventPort & ConfigPort & InfraPort & AuthPort)` expect specific port intersection types. TypeScript won't accept `unknown` here because it's not assignable to the port types. The `MockRouteContext` satisfies the interfaces structurally (it has all the required properties and methods), but since it's not declared as implementing them, we need `any` at the call site. This is the standard pattern for test mocks in TypeScript.
### Verification
```bash
tsc --noEmit
```
---
## Task 7: Add session routes tests
**Estimated effort**: 4 hours
**Depends on**: Task 6
**Files created**: `test/routes/session-routes.test.ts`
**Port**: 3220 (only if SSE tests needed; prefer `app.inject()`)
### Coverage targets
`src/web/routes/session-routes.ts` is the largest route module (43 handlers). Focus on the most critical endpoints first:
#### Priority 1: Session CRUD (must test)
| Method | Path | What to test |
|--------|------|-------------|
| `GET` | `/api/sessions` | Returns session list; empty when no sessions |
| `GET` | `/api/sessions/:id` | Returns session state; 404 for unknown ID |
| `POST` | `/api/sessions` | Creates session; validates workingDir; rejects invalid paths |
| `DELETE` | `/api/sessions/:id` | Calls cleanupSession; 404 for unknown ID |
#### Priority 2: Session I/O
| Method | Path | What to test |
|--------|------|-------------|
| `POST` | `/api/sessions/:id/input` | Sends input to session; validates input length; 404 for unknown |
| `POST` | `/api/sessions/:id/resize` | Validates cols/rows bounds; 404 for unknown |
| `GET` | `/api/sessions/:id/buffer` | Returns terminal buffer; 404 for unknown |
#### Priority 3: Session actions
| Method | Path | What to test |
|--------|------|-------------|
| `POST` | `/api/sessions/:id/run` | Runs prompt on session |
| `POST` | `/api/sessions/:id/clear` | Clears session |
| `POST` | `/api/sessions/:id/compact` | Compacts session |
| `POST` | `/api/sessions/:id/interactive` | Starts interactive mode |
| `POST` | `/api/sessions/:id/quick-start` | Quick start flow |
### Test pattern
```typescript
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { createRouteTestHarness, type RouteTestHarness } from './_route-test-utils.js';
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
describe('session-routes', () => {
let harness: RouteTestHarness;
beforeEach(async () => {
harness = await createRouteTestHarness(registerSessionRoutes);
});
afterEach(async () => {
await harness.app.close();
});
describe('GET /api/sessions', () => {
it('returns empty array when no sessions', async () => {
harness.ctx.sessions.clear();
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
expect(res.statusCode).toBe(200);
expect(JSON.parse(res.body)).toEqual([]);
});
it('returns session list with one session', async () => {
const res = await harness.app.inject({ method: 'GET', url: '/api/sessions' });
expect(res.statusCode).toBe(200);
const sessions = JSON.parse(res.body);
expect(sessions).toHaveLength(1);
});
});
describe('GET /api/sessions/:id', () => {
it('returns 404 for unknown session', async () => {
const res = await harness.app.inject({
method: 'GET',
url: '/api/sessions/nonexistent',
});
expect(res.statusCode).toBe(404);
});
});
describe('POST /api/sessions/:id/input', () => {
it('rejects input exceeding max length', async () => {
const res = await harness.app.inject({
method: 'POST',
url: `/api/sessions/${harness.ctx._sessionId}/input`,
payload: { input: 'x'.repeat(65537) },
});
expect(res.statusCode).toBe(400);
});
});
describe('POST /api/sessions/:id/resize', () => {
it('rejects cols exceeding max', async () => {
const res = await harness.app.inject({
method: 'POST',
url: `/api/sessions/${harness.ctx._sessionId}/resize`,
payload: { cols: 501, rows: 24 },
});
expect(res.statusCode).toBe(400);
});
});
});
```
### Key assertions to include
- **404 for unknown sessions**: Every `:id` endpoint must return 404 for nonexistent IDs
- **Input validation**: Bad paths, oversized inputs, invalid resize dimensions
- **Side effects**: Verify `ctx.broadcast()` was called with correct event type after mutations
- **Response shape**: Verify response bodies match expected API types
### Verification
```bash
npx vitest run test/routes/session-routes.test.ts
```
---
## Task 8: Add system + respawn routes tests
**Estimated effort**: 4 hours
**Depends on**: Task 6
**Files created**: `test/routes/system-routes.test.ts`, `test/routes/respawn-routes.test.ts`
### System routes (`src/web/routes/system-routes.ts`)
Focus on status and configuration endpoints:
| Method | Path | What to test |
|--------|------|-------------|
| `GET` | `/api/status` | Returns server status with uptime, session count |
| `GET` | `/api/stats` | Returns mux stats |
| `GET` | `/api/config` | Returns current config |
| `PUT` | `/api/config` | Updates config; validates input |
| `GET` | `/api/settings` | Returns user settings |
| `PUT` | `/api/settings` | Updates settings; validates input |
| `GET` | `/api/subagents` | Returns subagent list |
| `GET` | `/api/screenshots` | Returns screenshot list |
### Respawn routes (`src/web/routes/respawn-routes.ts`)
| Method | Path | What to test |
|--------|------|-------------|
| `GET` | `/api/sessions/:id/respawn` | Returns respawn status; null when not configured |
| `POST` | `/api/sessions/:id/respawn/start` | Starts respawn; 404 for unknown session |
| `POST` | `/api/sessions/:id/respawn/stop` | Stops respawn; 404 for unknown session |
| `PUT` | `/api/sessions/:id/respawn/config` | Updates respawn config; validates |
| `POST` | `/api/sessions/:id/respawn/enable` | Enables respawn loop |
| `POST` | `/api/sessions/:id/respawn/disable` | Disables respawn loop |
### Test patterns
Same pattern as Task 7 — `createRouteTestHarness` with `registerSystemRoutes` / `registerRespawnRoutes`.
For respawn tests, pre-populate `ctx.respawnControllers` with a mock controller in `beforeEach`:
```typescript
beforeEach(async () => {
harness = await createRouteTestHarness(registerRespawnRoutes);
// Add a mock respawn controller for the default session
harness.ctx.respawnControllers.set(harness.ctx._sessionId, {
getState: vi.fn(() => 'idle'),
getConfig: vi.fn(() => ({})),
getStatus: vi.fn(() => ({ state: 'idle', health: 100 })),
start: vi.fn(),
stop: vi.fn(),
updateConfig: vi.fn(),
enable: vi.fn(),
disable: vi.fn(),
});
});
```
### Verification
```bash
npx vitest run test/routes/system-routes.test.ts
npx vitest run test/routes/respawn-routes.test.ts
```
---
## Task 9: Slim down `respawn-test-utils.ts`
**Estimated effort**: 30 minutes
**Depends on**: Tasks 4, 5
**Files modified**: `test/respawn-test-utils.ts`
After Tasks 4–5 are verified passing with shared mocks, slim down `respawn-test-utils.ts` to remove duplicates.
### Steps
1. **Remove** from `respawn-test-utils.ts` what has been moved to shared mocks:
- `MockSession` class → now in `test/mocks/mock-session.ts`
- `createMockSession()` → now in `test/mocks/mock-session.ts`
- `terminalOutputs` → now in `test/mocks/mock-session.ts`
- `waitForEvent()` / `createDeferred()` → now in `test/mocks/test-helpers.ts`
2. **Keep** respawn-specific utilities that don't belong in the general mocks:
- `TimeController` / `createTimeController()` — respawn-specific timer control
- `MockAiIdleChecker` / `MockAiPlanChecker` — respawn-specific AI mocks
- `createStateTracker()` / `createEventRecorder()` — respawn state tracking
- `FAST_TEST_CONFIG` / `AI_ENABLED_TEST_CONFIG` — respawn config presets
- `waitForState()` — respawn state machine waiter
3. **Update imports** in `respawn-test-utils.ts` to re-use shared mocks:
```typescript
import { MockSession, createMockSession, terminalOutputs } from './mocks/index.js';
import { waitForEvent, createDeferred } from './mocks/index.js';
export { MockSession, createMockSession, terminalOutputs, waitForEvent, createDeferred };
```
This preserves backward compatibility for any future tests that import from `respawn-test-utils.ts` directly while eliminating the duplication.
### Verification
```bash
tsc --noEmit
npx vitest run test/respawn-controller.test.ts
npx vitest run test/respawn-team-awareness.test.ts
```
---
## What is NOT in scope (and why)
### Migrating `session-manager.test.ts` and `ralph-loop.test.ts` mocks
Both files define mocks inside `vi.mock()` factories that replace entire modules:
```typescript
// session-manager.test.ts — mock replaces ../src/session.js
vi.mock('../src/session.js', () => {
class MockSession extends EventEmitter { ... }
return { Session: MockSession };
});
// ralph-loop.test.ts — mock replaces ../src/state-store.js
vi.mock('../src/state-store.js', () => {
class MockStateStore { ... }
return { getStore: vi.fn(() => instance), StateStore: MockStateStore };
});
```
These are fundamentally different from the direct-instantiation pattern:
- The `vi.mock()` factory runs in an isolated scope — outer imports are not available
- The mock class must be returned with the exact export names (`Session`, `getStore`, `StateStore`)
- The `session-manager.test.ts` MockSession auto-registers into a shared `mockState.sessions` Map (tight coupling with test setup)
Migrating would require `vi.hoisted()` to share the class between factory and test scope, plus restructuring the test's module-mocking setup. This is high-complexity, high-risk refactoring with limited benefit since these tests already work. The shared `MockStateStore` in `test/mocks/` is available for **new** tests (like route tests) that use direct instantiation instead.
### Full integration tests with real Fastify server
Route tests use `app.inject()` which simulates HTTP without opening ports. Full integration tests that spin up `WebServer`, create real sessions, and stream SSE would be valuable but are a separate effort requiring:
- A test WebServer factory
- Session lifecycle management in tests
- SSE client test utilities
- Significantly more setup/teardown complexity
### Testing auth middleware in route tests
Route tests bypass authentication (no auth middleware registered on the test Fastify instance). Auth middleware has its own dedicated tests in `auth-security.test.ts` and `qr-auth.test.ts`. Testing auth + routes together is a future integration test concern.
### Testing SSE event streaming
SSE integration requires a running server with `EventSource` client. This is significantly more complex than `app.inject()` tests and is deferred. The existing `sse-events.test.ts` covers SSE patterns.
### Complete route coverage for all 12 modules
This phase covers the 3 highest-value route modules (session, system, respawn — 98 of 162 handlers). The remaining 9 modules (ralph, plan, push, team, mux, file, scheduled, hook-event, case) should be added incrementally in follow-up work.
---
## Summary
| Metric | Before | After |
|--------|--------|-------|
| MockSession definitions | 4 (across 4 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
| MockStateStore definitions | 2 (across 2 files) | 1 shared (2 vi.mock() copies remain, intentionally) |
| Files importing from `respawn-test-utils.ts` | 0 | Utilities split into `test/mocks/` |
| Route test files | 0 | 3 (session, system, respawn) |
| Route handlers with dedicated tests | 0 | ~30 (highest-priority endpoints) |
| Shared mock directory | None | `test/mocks/` with 5 files + barrel |
### Final verification checklist
```bash
# Type checking
tsc --noEmit
# Linting
npm run lint
# Formatting
npm run format:check
# Run all affected tests individually
npx vitest run test/respawn-controller.test.ts
npx vitest run test/respawn-team-awareness.test.ts
npx vitest run test/routes/session-routes.test.ts
npx vitest run test/routes/system-routes.test.ts
npx vitest run test/routes/respawn-routes.test.ts
# Verify unchanged tests still pass
npx vitest run test/session-manager.test.ts
npx vitest run test/ralph-loop.test.ts
# Dev server still starts
npx tsx src/index.ts web --port 3099 &
curl -s http://localhost:3099/api/status | jq .status # "ok"
kill %1
```
-5
View File
@@ -300,11 +300,6 @@ For reference when writing browser tests:
.xterm // Terminal container
#helpModal // Help modal
#appSettingsModal // Settings modal
#sessionOptionsModal // Session Options (same set-* surface)
#createCaseModal // Add Case (same set-* surface)
.set-rail-item // Rail entry: scrolls in App Settings, switches in the other two
.set-section // A settings section (`.hidden` on the inactive ones outside App Settings)
.set-row // One setting: label + description left, control right
.modal-content // Modal content
.modal-close // Modal close button
.header-brand .logo // Logo text
+26 -106
View File
@@ -2,18 +2,14 @@
> Official documentation for Claude Code hooks system, extracted from [code.claude.com](https://code.claude.com/docs/en/hooks).
**Last Updated**: 2026-07-25
**Last Updated**: 2026-01-24
**Source**: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)
> This is a maintained summary, not an exhaustive copy of the upstream reference.
> Check the source link for event-specific schemas before adding a new hook.
---
## Overview
Hooks are automated scripts that execute at specific events during your Claude Code session. They allow you to:
- Validate, modify, or block tool usage
- Add context to prompts
- Implement custom workflows
@@ -25,12 +21,12 @@ Hooks are automated scripts that execute at specific events during your Claude C
Hooks are configured in settings files:
| File | Scope |
| ----------------------------- | -------------------------- |
| `~/.claude/settings.json` | User (global) |
| `.claude/settings.json` | Project |
| File | Scope |
|------|-------|
| `~/.claude/settings.json` | User (global) |
| `.claude/settings.json` | Project |
| `.claude/settings.local.json` | Local project (gitignored) |
| Plugin hook files | Plugin-specific |
| Plugin hook files | Plugin-specific |
### Basic Structure
@@ -53,9 +49,8 @@ Hooks are configured in settings files:
```
**Key Fields**:
- `matcher`: Pattern to match tool names (case-sensitive, supports regex like `Edit|Write` or `*` for all)
- `type`: `"command"`, `"http"`, `"mcp_tool"`, `"prompt"`, or `"agent"` where the event supports it
- `type`: `"command"` for bash or `"prompt"` for LLM-based evaluation
- `command`: Bash command to execute
- `prompt`: LLM prompt for evaluation (prompt-based hooks only)
- `timeout`: Optional timeout in seconds (default: 60)
@@ -64,10 +59,6 @@ Hooks are configured in settings files:
## Hook Events
Claude Code's current event surface is broader than the detailed subset below. In
particular, `TeammateIdle` and `TaskCompleted` are supported lifecycle events used
by Codeman; they are not stale or plugin-defined event names.
### PreToolUse
**When**: After Claude creates tool parameters, before processing the tool call.
@@ -75,17 +66,15 @@ by Codeman; they are not stale or plugin-defined event names.
**Use Cases**: Approval, denial, or modification of tool calls.
**Common Matchers**:
- `Bash` - Shell commands
- `Write` - File writing
- `Edit` - File editing
- `Read` - File reading
- `Agent` - Subagent tasks
- `Task` - Subagent tasks
- `WebFetch`, `WebSearch` - Web operations
- `mcp__<server>__<tool>` - MCP tools
**Output Control**:
```json
{
"hookSpecificOutput": {
@@ -107,14 +96,13 @@ by Codeman; they are not stale or plugin-defined event names.
**Use Cases**: Auto-approve or deny permissions.
**Output Control**:
```json
{
"hookSpecificOutput": {
"hookEventName": "PermissionRequest",
"decision": {
"behavior": "allow|deny",
"updatedInput": {},
"updatedInput": { },
"message": "deny reason",
"interrupt": false
}
@@ -129,7 +117,6 @@ by Codeman; they are not stale or plugin-defined event names.
**Use Cases**: Provide feedback, run formatters/linters, log operations.
**Output Control**:
```json
{
"decision": "block",
@@ -141,38 +128,15 @@ by Codeman; they are not stale or plugin-defined event names.
}
```
#### Asynchronous Rewake
Command hooks can set `"asyncRewake": true` to run asynchronously and wake an
idle Claude turn when the hook exits with code 2. The hook's stderr is delivered
to Claude as a system reminder. This implies `"async": true`; ordinary async
hooks do not wake an idle turn, and their output waits for the next interaction.
Codeman uses this on `PostToolUse(Bash)`: a self-contained Node helper extracts
the background task ID from the Bash result, watches the originating transcript
and, for subagents, the top-level parent transcript for the matching completion
notification, and exits 2. Claude records a subagent's Bash result in its
`subagents/agent-*.jsonl` file but queues completion in the lead session JSONL.
The task ID keeps each wake targeted. The helper does not send terminal input,
so it cannot submit a user's partially written prompt.
For script-dispatched Codex work, `codex-run.sh` writes the final response
between `CODEMAN_RESULT_BEGIN/END` markers in the background task output. The
rewake helper includes a maximum of 64 KiB of that report in its feedback. UI
subagent discovery and dispatcher result delivery are separate contracts.
### Notification
**When**: When Claude Code sends notifications.
**Matchers**:
- `permission_prompt`
- `idle_prompt`
- `auth_success`
- `elicitation_dialog`
- `elicitation_complete`
- `elicitation_response`
### UserPromptSubmit
@@ -181,7 +145,6 @@ subagent discovery and dispatcher result delivery are separate contracts.
**Use Cases**: Add context, validate, or block prompts.
**Output Control**:
```json
{
"decision": "block",
@@ -202,7 +165,6 @@ subagent discovery and dispatcher result delivery are separate contracts.
**Use Cases**: **Ralph Wiggum loops** - block exit and refeed prompt.
**Output Control**:
```json
{
"decision": "block",
@@ -211,7 +173,6 @@ subagent discovery and dispatcher result delivery are separate contracts.
```
Or to allow exit:
```json
{
"continue": true,
@@ -223,42 +184,15 @@ Or to allow exit:
### SubagentStop
**When**: When a subagent (Agent tool call) finishes responding.
**When**: When a subagent (Task tool call) finishes responding.
**Use Cases**: Control nested loops, verify subagent output.
The hook input includes `agent_id`, `agent_transcript_path`, and
`last_assistant_message`. Like `Stop`, a command hook can return
`{"decision":"block","reason":"..."}` to keep the subagent running and feed
the reason back to it.
Codeman uses this to prevent premature reports from workers that still own live
Monitor or background-Bash processes. It derives candidate task IDs from the
subagent transcript, but requires a matching live Linux process descriptor for
`tasks/<id>.output`; historical task text by itself is not treated as active.
### TeammateIdle
**When**: When an agent-team teammate is about to go idle.
**Use Cases**: Reassign work, continue a teammate loop, or notify an orchestrator.
**Matcher Support**: None. The hook fires for every occurrence.
### TaskCompleted
**When**: When a task is about to be marked completed.
**Use Cases**: Validate completion or forward team progress to an external UI.
**Matcher Support**: None. The hook fires for every occurrence.
### PreCompact
**When**: Before a compact operation.
**Matchers**:
- `manual` - Invoked from `/compact`
- `auto` - Invoked from auto-compact
@@ -267,7 +201,6 @@ subagent transcript, but requires a matching live Linux process descriptor for
**When**: When Claude Code starts or resumes a session.
**Matchers**:
- `startup` - Fresh start
- `resume` - From `--resume`, `--continue`, or `/resume`
- `clear` - From `/clear`
@@ -276,7 +209,6 @@ subagent transcript, but requires a matching live Linux process descriptor for
**Use Cases**: Load development context, set environment variables.
**Persisting Environment Variables**:
```bash
#!/bin/bash
if [ -n "$CLAUDE_ENV_FILE" ]; then
@@ -287,7 +219,6 @@ exit 0
```
**Output Control**:
```json
{
"hookSpecificOutput": {
@@ -302,7 +233,6 @@ exit 0
**When**: When a session ends.
**Reason Values**:
- `clear`
- `logout`
- `prompt_input_exit`
@@ -324,7 +254,7 @@ Hooks receive JSON via stdin with common fields:
"permission_mode": "default",
"hook_event_name": "PreToolUse",
"tool_name": "Bash",
"tool_input": {},
"tool_input": { },
"tool_use_id": "toolu_01ABC123..."
}
```
@@ -332,7 +262,6 @@ Hooks receive JSON via stdin with common fields:
### Tool-Specific Input
**Bash**:
```json
{
"tool_name": "Bash",
@@ -345,7 +274,6 @@ Hooks receive JSON via stdin with common fields:
```
**Write**:
```json
{
"tool_name": "Write",
@@ -357,7 +285,6 @@ Hooks receive JSON via stdin with common fields:
```
**Edit**:
```json
{
"tool_name": "Edit",
@@ -375,11 +302,11 @@ Hooks receive JSON via stdin with common fields:
### Exit Codes
| Code | Behavior |
| ----- | --------------------------------------------------------------------- |
| 0 | Success. `stdout` processed (shown in verbose or added as context) |
| 2 | Blocking error. Only `stderr` used. Blocks tool/prompt based on event |
| Other | Non-blocking error. `stderr` shown in verbose, execution continues |
| Code | Behavior |
|------|----------|
| 0 | Success. `stdout` processed (shown in verbose or added as context) |
| 2 | Blocking error. Only `stderr` used. Blocks tool/prompt based on event |
| Other | Non-blocking error. `stderr` shown in verbose, execution continues |
### JSON Output (Exit Code 0)
@@ -396,12 +323,7 @@ Hooks receive JSON via stdin with common fields:
## Prompt-Based Hooks
Prompt and agent handlers are supported by decision-oriented events including
`PreToolUse`, `PermissionRequest`, `PostToolUse`, `PostToolUseFailure`,
`PostToolBatch`, `UserPromptSubmit`, `Stop`, `SubagentStop`, `TaskCreated`, and
`TaskCompleted`. Check the upstream reference before choosing a handler type.
For example, a Stop event can use LLM-based evaluation:
For Stop and SubagentStop events, you can use LLM-based evaluation:
```json
{
@@ -422,7 +344,6 @@ For example, a Stop event can use LLM-based evaluation:
```
**LLM Response Format**:
```json
{
"ok": true,
@@ -441,18 +362,17 @@ Hooks can be defined in Skills, Agents, and Slash Commands using frontmatter:
name: secure-operations
hooks:
PreToolUse:
- matcher: 'Bash'
- matcher: "Bash"
hooks:
- type: command
command: './scripts/security-check.sh'
command: "./scripts/security-check.sh"
---
```
These hooks:
- Are scoped to the component's lifecycle
- Only run when that component is active
- Support all hook events; a subagent-scoped `Stop` is converted to `SubagentStop`
- Support: PreToolUse, PostToolUse, Stop
---
@@ -630,11 +550,11 @@ exit 0
## Environment Variables
| Variable | Description |
| -------------------- | ------------------------------------------------ |
| `CLAUDE_PROJECT_DIR` | Project root directory |
| `CLAUDE_CODE_REMOTE` | `"true"` for web, empty for CLI |
| `CLAUDE_ENV_FILE` | Path to write persistent env vars (SessionStart) |
| Variable | Description |
|----------|-------------|
| `CLAUDE_PROJECT_DIR` | Project root directory |
| `CLAUDE_CODE_REMOTE` | `"true"` for web, empty for CLI |
| `CLAUDE_ENV_FILE` | Path to write persistent env vars (SessionStart) |
---
@@ -673,4 +593,4 @@ Use `/hooks` command to view registered hooks and make changes.
---
_Source: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)_
*Source: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)*
-121
View File
@@ -1,121 +0,0 @@
# Claude voice dictation in Codeman
Wire Codeman's existing mic button to the same speech-to-text service Claude Code's own
`/voice` mode uses, so dictation works with **no third-party API key** for anyone already
signed in to Claude Code on the server.
## Why the CLI's own voice mode cannot be reused directly
Claude Code 2.1.x ships voice input: `/voice hold|tap|off` arms it, the CLI opens the
**host's** microphone (native `audio-capture-napi`, falling back to `sox`/`arecord` on Linux
after probing `/proc/asound/cards`), streams PCM upstream and types the transcript into its
own composer.
Every part of that is on the wrong machine for Codeman. The CLI runs inside a tmux pane on
the server, which is typically headless and has no sound card at all, while the human is in
a browser on a phone somewhere else. Toggling `/voice` in the pane from Codeman would arm a
microphone nobody is sitting in front of. So Codeman keeps capturing audio in the browser,
where the user actually is, and only borrows the CLI's **transcription backend**.
## The backend, as the CLI uses it
Extracted from the 2.1.226 binary (`connectVoiceStream`):
| | |
| --- | --- |
| URL | `wss://api.anthropic.com/api/ws/speech_to_text/voice_stream` |
| Query | `encoding=linear16`, `sample_rate=16000`, `channels=1`, `endpointing_ms=300`, `utterance_end_ms=1000`, `language=<lang>`, `use_conversation_engine=true`, `stt_provider=deepgram-nova3` |
| Headers | `Authorization: Bearer <Claude Code OAuth access token>`, `User-Agent`, `x-app: cli`, `anthropic-client-platform`, optional `x-config-keyterms` |
| Audio | raw binary frames, PCM signed 16-bit little-endian, 16 kHz, mono |
| Keepalive | `{"type":"KeepAlive"}` on open, then every 8 s |
| Finalize | `{"type":"CloseStream"}`, then wait for the endpoint frame |
| Downstream | `{"type":"TranscriptText"\|"TranscriptInterim","data":"…"}` (running interim), `{"type":"TranscriptEndpoint"}` (promotes the pending interim to final), `{"type":"TranscriptError",…}`, `{"type":"error","message":…}` |
Deepgram Nova-3 runs server-side, so the Deepgram-quality result arrives without a Deepgram
account. Verified against the live endpoint before this design was written: connect, stream
PCM, receive interims and an endpoint frame.
## Architecture
The browser cannot call that endpoint itself: it would need the OAuth bearer token in page
JavaScript (and CORS would refuse anyway). So the audio goes browser → Codeman → Anthropic,
and Codeman is the only thing that ever touches the token.
```
mic → AudioWorklet (Float32 → PCM16 @16 kHz)
→ wss://<codeman>/ws/voice/stream [cookie/basic auth, Origin+Host guarded]
→ VoiceStreamRelay (reads ~/.claude/.credentials.json per connect)
→ wss://api.anthropic.com/api/ws/speech_to_text/voice_stream
← {"t":"transcript","text":…,"final":…} → existing _insertText() path
```
Nothing about the insert path changes: the transcript lands in the same preview overlay,
the same direct/compose insert modes, the same green Send button.
### Server pieces
- **`src/claude-credentials.ts`** — locate and parse the Claude Code OAuth credentials.
`parseClaudeCredentials()` is pure (JSON string + `now` → status) and unit-tested;
`readClaudeOAuthToken()` wraps it with IO: `$CLAUDE_CONFIG_DIR/.credentials.json` or
`~/.claude/.credentials.json`, and on macOS the login keychain
(`security find-generic-password -s "Claude Code-credentials"`).
**Read-only, always.** Codeman never writes credentials and never refreshes the token: a
refresh rotates the refresh token, and racing Claude Code's own refresh could sign the
user out of their CLI. An expired token surfaces as a plain "run a Claude session to
refresh" error instead.
The token is never logged, never returned by any endpoint, and never sent to the browser.
- **`src/web/voice-stream.ts`** — pure `buildVoiceStreamUrl()` / `buildVoiceStreamHeaders()` /
`sanitizeKeyterms()` (ASCII-only, deduped, 1024-char cap, mirroring the CLI), plus
`VoiceStreamRelay`, which owns one upstream socket: keepalive timer, audio passthrough,
transcript translation, finalize, and the caps below.
- **`src/web/routes/voice-routes.ts`**
- `GET /api/voice/status` → `{ available, reason, subscriptionType?, expiresAt? }`. Never
the token. `available:false` with a machine-readable `reason` (`disabled`, `no-credentials`,
`expired`) is what the settings row and the provider resolver read.
- `GET /ws/voice/stream?language=&keyterms=` → the relay. Same upgrade guard as
`/ws/sessions/:id/terminal`: allowed Host, same-site Origin, and the global auth hook has
already run on the handshake.
Caps, because an open mic is an open pipe: one stream per connection, `MAX_VOICE_STREAMS`
concurrent server-wide, a hard `MAX_STREAM_MS` per stream, and a per-frame size cap. A tab
left recording cannot bill an unbounded amount of upstream audio.
### Frontend pieces
- **`voice-pcm-worklet.js`** — an `AudioWorkletProcessor` converting Float32 blocks to PCM16
and posting ~256 ms frames back. `MediaRecorder` cannot produce raw PCM, which is why the
existing Deepgram path (container audio, auto-detected) cannot be reused as-is. Falls back
to `ScriptProcessorNode` where AudioWorklet is unavailable.
- **`ClaudeVoiceProvider`** in `voice-input.js` — mirrors `DeepgramProvider`'s shape
(`start({language, keyterms, onStream, onResult, onError, onEnd})`) so `VoiceInput` treats
the three providers uniformly.
- **Provider resolution** — new `voiceSettings.provider`: `auto` (default) | `claude` |
`deepgram` | `webspeech`. `auto` picks Claude when `/api/voice/status` reports it
available, else Deepgram when a key is set, else Web Speech. Pinning a provider always
wins, so an existing Deepgram user can keep exactly what they have.
### Settings
- `claudeVoiceEnabled` — synced, **default OFF**, gating the whole server side. Off is the
honest default: turning it on means this machine's Claude subscription starts paying for
transcription for whoever can reach the UI, and the audio goes to Anthropic rather than to
wherever it went before. One switch in Settings → Voice, and the mic works with no key.
- `voiceSettings.provider` — per the resolution table above; joins the existing synced
`voiceSettings` object.
## Things worth knowing
- **This uses an undocumented endpoint with subscription credentials.** It is the user's own
token, on the user's own machine, driving the user's own Claude Code install, but it is not
a published API and Anthropic can change or restrict it. Default-OFF is deliberate; the
Deepgram and Web Speech paths stay untouched as the supported fallbacks.
- **Multi-user mode**: every user's dictation would run on the server owner's Claude
credentials, exactly as every user's *sessions* already run on them. Consistent, but worth
stating out loud in the settings copy.
- **Token lifetime** is about 8 hours, refreshed by Claude Code itself whenever it runs. The
relay re-reads the file on every connect rather than caching, so a refresh is picked up on
the next press of the mic.
- **HTTPS or localhost**: `getUserMedia` needs a secure context. Prod is HTTPS behind
`tailscale serve`, so this is already satisfied; the existing error copy covers the rest.
@@ -1,9 +1,3 @@
> **⚠️ ARCHIVED 2026-05-21 — superseded, kept for history.**
> The headline items here were verified resolved: the P0 `{WORKING_DIR}` placeholder
> is now replaced (`plan-orchestrator.ts:431`), and the "~66 dead functions in app.js"
> are gone (app.js was modularized 15K→3K LOC). A fresh `npm run knip` sweep on
> 2026-05-21 found only a handful of unused test helpers. Do not treat this as a live TODO.
# Codebase Cleanup Findings
Compiled from parallel analysis of the entire Codeman codebase by 3 research agents (2026-02-19).
-191
View File
@@ -1,191 +0,0 @@
# The CLI registry
Every run mode Codeman can launch — Claude Code, Terminal/Shell, OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek Harness and OMP — is a `CliEntry`: a data record describing how to find the binary, how to build its command line, what environment it needs, and what it can do. Code that used to ask "which CLI is this?" asks the entry instead.
## Where it lives
| File | What it holds |
| ------------- | ------------------------------------------------------------------------------------------------- |
| `types.ts` | The `CliEntry` interface and everything under it. Read this first. |
| `stock.ts` | The shipped catalog. **The only file allowed to name a CLI id.** |
| `schema.ts` | Zod validation, including the cross-field checks that reject an incoherent entry at LOAD time. |
| `argv.ts` | The argv engine: the only code that turns typed tokens into a command string. |
| `patterns.ts` | The NAMED value patterns (`model`, `uuid`, `path-segment`, …) and the regex-compilation guard. |
| `profiles.ts` | The names of behaviours that genuinely need code, kept import-free so `schema.ts` can validate one. |
| `registry.ts` | Loading, merging `~/.codeman/clis.json`, and the accessors (`getCli`, `enabledClis`). |
`src/session-cli-registry-bridge.ts` maps the legacy per-mode option bag onto the engine, and `src/utils/cli-resolver.ts` / `src/utils/cli-launcher.ts` do registry-driven binary resolution and launcher-profile dispatch.
## The override file
`~/.codeman/clis.json` (instance-scoped through `dataPath()`) holds overrides and custom entries only, never a copy of the stock catalog: `{ "clis": { "<id>": { ...partial entry... } } }`. Objects merge key-wise onto the stock entry, arrays replace wholesale. **The file must be mode 0600**; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you `chmod 600` it. Every reason a file was ignored or an entry dropped is logged once, prefixed `[cli-registry]`, on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read only on restart.
## The shape of an entry
```ts
interface CliEntry {
id: CliId; // 'codex'
label: string; // 'Codex' — shown in menus
shortBadge: string; // tab badge, e.g. 'CX'
accent: string; // single hex colour
enabled: boolean;
stock: boolean; // set by the loader; a custom entry can never claim it
order: number;
kind: 'agent' | 'shell';
discovery: CliDiscovery; // how to find and prove the binary
launch: CliLaunch; // the structured argv template
env: CliEnv; // exports, tmux setenv keys, the env-override allowlist
capabilities: CliCapabilities; // what every call site reads instead of the id
// .workDetect?: { promptGlyph, workingLine } — how this CLI's pane shows work
overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store
}
```
`capabilities` is the important part. It is what `isExternalCliMode()`, `isAltScreenStripMode()`, `hooksAvailableForMode()` and every other former per-mode branch actually read.
### Regexes that come from config
Two capability fields carry a regular expression an override file can set: `discovery.version.regex` and `capabilities.workDetect.workingLine`. Both go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
`workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused.
### Three capabilities that must stay independent
`external`, `hooks` and `altScreen` describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. `shell` has no hooks but is **not** an external CLI, so a hooks predicate written as `!isExternalCliMode()` accepted `until=stop` on a shell session and then blocked the caller for their entire timeout. `deepseek` is the mirror image: it IS external and it DOES have hooks.
`test/cli-capability-predicates.test.ts` asserts that no two of the three are equivalent across the catalog, so collapsing them fails the build rather than a user's session.
## Arg-template safety
The composed command line is interpolated into `bash -c "…"` inside tmux, which makes command construction a security boundary. Four independent layers keep config out of it:
1. **Config contains no shell text.** There is no `command: "..."` field anywhere in the schema. An entry declares a sequence of typed tokens; `argv.ts` is the only place that turns them into a string, and it owns every separator itself — one space between tokens, ` || ` between fallback variants. Neither can originate from config, because config has no field that could hold either.
2. **Every literal is validated at LOAD time** against a safe-word pattern (no space, quote, backtick, `$`, `;`, `&`, `|`, redirection, parens, braces, newline or backslash). A bad literal **rejects the whole entry** rather than being dropped, because a silently dropped flag would change security-relevant behaviour — losing `--no-approve` is not a cosmetic difference.
3. **Values resolve through NAMED patterns.** A value placeholder selects a `TokenPattern` (`model`, `uuid`, `slug`, `path-segment`, `tool-list`, …) from `patterns.ts`; config can never supply its own regex for a value, so a `clis.json` structurally cannot widen its own validation. A value that fails its pattern drops the whole argument, exactly as the hand-written builders did: an invalid `--model` omits `--model`, it never substitutes something else.
4. **Escaping is independent of validation.** `renderToken()` re-checks the resolved value before emitting it unquoted, and single-quotes anything else — so even a value that somehow bypassed validation is quoted, never concatenated raw.
The only config-supplied regexes are `discovery.version.regex` and `discovery.identity.regex`. Both run against **command output** rather than a shell token, both are compiled through `compileVersionRegex()` (length cap, nested-quantifier rejection, never the `g` flag), and the output they see is truncated first.
## Named profiles: the escape hatch
Some differences genuinely need to run code rather than be described. Those are **named profiles**: a capability field holds a profile NAME, and the implementation lives in one place keyed by that name — never by CLI id.
- `discovery.launcherProfile` — for a CLI whose binary is not the agent. `dsh` boots `$DSH_HOME/profiles/<name>`, so "installed" and "runnable" have different answers; the profile answers both, plus why a specifically-named target will not work. Implemented in `utils/cli-launcher.ts`.
- `env.setenvProfile` — per-CLI environment setup that is more than a list of keys, such as DeepSeek's status bridge.
- `capabilities.transcript` — which on-disk history reader understands this CLI (`claude-jsonl`, `codex-rollout`, `deepseek-zstd`, `omp-jsonl`, `none`).
- `capabilities.echo.predictProfile` — the predictive-echo model a composer needs.
The names live in `profiles.ts`, which is kept free of imports so `schema.ts` can validate a name at load time. A profile this build does not implement is a load-time error naming the field, rather than a CLI that silently looks permanently uninstalled.
## DeepSeek: the four assumptions it breaks
DeepSeek is worth reading before assuming an entry looks like its siblings — the schema carries four extensions because of it.
| What it breaks | How the registry expresses it |
| ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------ |
| `dsh` is a profile LAUNCHER, not the agent, so "installed" is not "runnable". | `discovery.launcherProfile` + `discovery.launcherTargetParam`. |
| Its permission switch is the **`DSH_PERMISSION_MODE` env var**, not a flag — the harness has none. | `env.configSetenv` (so the ordinary `privilegedParams` clamp still reaches it) **and** `capabilities.privilegedEnvKeys`. |
| It is the only non-claude mode with real hook signals, and for it that is a per-SESSION question. | `capabilities.hooks: 'supervised'` — a third state, not a boolean. |
| Its transcript is zstd session files, one frame per write. | `capabilities.transcript: 'deepseek-zstd'`. |
## Identity probes
`discovery.identity` asks the binary whether it is the program we meant, and it runs **before** the version probe, because a version probe cannot tell an impostor from the real thing. Debian ships an unrelated `dsh` (dancer's shell) that answers `--version` perfectly happily, and npm carries squatters for both `pi` and `grok`.
`discovery.version.requireVersionMatch` is the weaker companion: a binary whose version output has the wrong shape counts as ABSENT rather than present-with-unknown-version. That is what a short, generic binary name needs, and it is what keeps `codeman doctor` and the run mode from telling the user opposite things about the same binary — both read the same regex off the same entry.
## The no-id-branching rule
`test/cli-registry-no-id-branching.test.ts` fails the build if a CLI id comparison appears outside the stock catalog. It builds its id list from the live catalog, blanks comment lines before scanning (comments legitimately quote the pattern to explain why a branch was removed, and blanking rather than dropping is what keeps reported line numbers pointing at the real file), and keeps an allowlist in which **every entry carries its reason**.
It matches four shapes, not one: `mode === '<id>'`, `mode !== '<id>'`, `case '<id>':`, and `['<id>', …].includes(mode)`. The first version matched `===` only, and that gap was not academic — the refactor it guards converted the `===` sites and left the negated ones, so 36 `!==` branches survived it, including a seven-mode chain auto-enabling Ralph under a comment asking the next person to keep it in step with a predicate by hand while the sibling code path already read the capability. A guard that sees half the shapes reports a count measured over the half it happens to catch.
The allowlist is not a formality. If a branch is about what a CLI can DO it belongs in `CliCapabilities`; the entries that remain are things that are not CLI-behaviour branches at all — chiefly the legacy per-mode `<Mode>Config` objects on `POST /api/sessions`, which are a fact about the public HTTP API rather than about any CLI, plus a few documented cases where `mode === 'claude'` is genuinely the right question (Read My Mind reads Claude's _own_ transcript, so a capability there would be actively wrong).
## Two namespaces called `param`
`launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a **launch param**. The **legacy wire field** a param arrives as is a separate namespace, and `launch.legacyConfigAliases` is the only bridge between the two.
This matters because it is invisible when it is wrong. `capabilities.privilegedParams[].param` is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a name from the wrong namespace clamps **nothing**: no load error, no failing test, the clamp simply stops running. Codex is the entry where the two names differ (`bypassApprovals` as the param, `dangerouslyBypassApprovals` on the wire), so it is the one that catches a regression. `schema.ts` rejects any entry naming a param it never declared, on both `configSetenv.fromParam` and `privilegedParams.param`.
## Fields declared for later
`shortBadge`, `accent`, `capabilities.echo`, `capabilities.wheelForward`, `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` are **declared but not yet read**. They all describe frontend behaviour, and the frontend is deliberately untouched here: `app.js`, `terminal-ui.js` and `styles.css` keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
Treat those values as **transcribed, not authoritative** — nothing enforces that `echo.policy` matches `_updateLocalEchoState`'s fallthrough, or that `accent` matches the gradient CSS paints, so re-measure before wiring one up. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; `test/cli-registry-no-id-branching.test.ts` pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
`overlays.credStore` is in the same category, for a sharper reason: the Docker credential-seeding path still reads its own `CRED_STORES` table, because this shape allows ONE store per CLI and the live table needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek declares none here even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is simultaneously the worst thing here to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest.
Everything else in the interface is live, including `overlays.remote` / `overlays.docker`, which back `defaultRemoteCommandForMode()` and `defaultDockerCommandForMode()` directly. Those two used to be hardcoded `Record<…CommandMode, string>` tables duplicating the registry with nothing keeping the two in step; `test/location-overlay-commands.test.ts` pins every resulting command as a literal string.
## Consumers outside the server
Two things need the catalogue but cannot import TypeScript, so `npm run generate:cli-catalog`
(`scripts/generate-cli-catalog.mts`) emits two artifacts from `stock.ts`. Both are committed,
and `test/cli-catalog-sync.test.ts` fails if either drifts from a fresh generation.
| Artifact | Consumer | Why it exists |
| ------------------------------------ | ---------------------------------- | ---------------------------------------------------------------------------------- |
| `config/clis.stock.json` | `scripts/lib/cli-catalog.mjs` (Docker build args), tests | A `.mjs` cannot import the registry. |
| a marked block inside `install.sh` | the installer itself | It runs via `curl \| bash` before any checkout exists, so it can read neither. |
Only `id`, `label`, `shortBadge`, `enabled`, `order`, `kind` and `discovery` are exported.
`launch`, `env`, `capabilities` and `overlays` are spawn-time concerns the server alone
interprets, and a test asserts they never leak into the artifact — a second reading of the
launch model in a consumer that cannot be tested against a real spawn is exactly what this
registry exists to prevent.
The install.sh copy is **embedded, not fetched**, and is the FULL catalogue. An earlier design
fetched it and fell back to a hardcoded two-CLI list, which degraded silently on an empty
response; there is no degraded mode to fall into now, and no network fetch either — a `curl |
bash` from master already carries a catalogue exactly as fresh as the script itself, so there is
nothing a refresh would buy that isn't already true. An earlier draft added an opt-in refresh
with a `TRUSTED`/`DISPLAY` array split to keep it from ever writing the executed command; it was
dropped before merge rather than shipped half-verified — the split's only actual write was the
label, `DISPLAY` never diverged from `TRUSTED` in practice, and the added surface (a second
array, a fetch path, three failure shapes to warn on) bought nothing the embedded copy didn't
already have.
### The install-command trust boundary
Three rules, and the middle one is why the embed matters:
1. **The server never executes an entry's `install.command`.** Unchanged, and still enforced by nothing executing it: the field is display text (`CliDiscovery.install.command`).
2. **`install.sh` executes only commands embedded in itself.** Those arrive in the same file, over the same TLS fetch, in the same commit as the `curl \| bash` line that fetched the script — identical trust to the hardcoded vendor one-liners it replaces.
3. **Nothing fetched at install time is ever executed.** There is no second code path that fetches anything after the script itself has been fetched.
That is mechanical rather than a promise. `CLI_INSTALL_CMD_TRUSTED` is written only from the
generated block and is the only array the installer ever runs or displays — there is no second
array a refresh could rewrite, because there is no refresh. `test/cli-catalog-sync.test.ts`
asserts that the embedded commands are exactly the registry's, and
`test/install-sh-invariants.test.ts` that nothing in `install.sh` `eval`s.
### bash 3.2
macOS ships bash 3.2 and the documented install is `curl -fsSL <url> | bash` under
`set -euo pipefail`, so a bash-4 construct is not a warning there — it kills the install. The
generated block therefore uses parallel indexed arrays with **offset/length windows** into one
flat array instead of delimiters (a `$HOME` containing a space needs no `IFS` handling, and an
entry with nothing to contribute gets length 0 and is never iterated). CI runs `bash -n` and
executes the script inside a real `bash:3.2` container, because the empty-window case is a
runtime `set -u` abort that `bash -n` cannot see.
## Resolve at call time, never at import
Anything reading the registry must resolve it when it is asked, not when its module is first imported. `sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()` and each resolver's `searchDirs` thunk all re-read the catalog per call.
A module-level const freezes at first import, and the failure is asymmetric: a CLI enabled while the server is running moved the run menu but not the frozen surface, so validation rejected a mode the menu offered, or `codeman doctor` reported a catalog nobody had any more.
## Adding a CLI
1. Add a `CliEntry` to `stock.ts`.
2. Run `npm run generate:cli-catalog` and commit **both** artifacts (`config/clis.stock.json` and `install.sh`). The installer's detection, its install menu, its reminder text and the Docker agent image all follow from that one step — this is what makes upstream `b6d0f1fa` ("wire OMP into install.sh's CLI detection, it had none") impossible rather than merely fixed.
3. Add a golden spawn-command pin to `test/cli-registry-spawn-golden.test.ts`, a row to `test/cli-capability-predicates.test.ts`, its remote/docker commands to `test/location-overlay-commands.test.ts`, and its search paths to `test/install-sh-detection-parity.test.ts`.
4. Only if it cannot install with a plain `npm install -g <pkg>`: give it a layer in `docker/agent.Dockerfile` and set `discovery.install.agentImageLayer: { kind: 'dedicated', reason }` on its entry in `stock.ts`. `test/docker-agent-image-coverage.test.ts` requires both, so an exclusion cannot quietly become an omission. An entry with no `npmPackage` needs only the Dockerfile layer, since it never enters the shared npm layer in the first place.
5. That is usually all. If you find yourself wanting to add an `if` somewhere, the guard test will tell you — and the answer is a capability field, or a named profile if it genuinely needs to run code.
## See also
- [Agent CLIs](wiki/Agent-CLIs.md) — the user-facing per-CLI guide.
- `docs/architecture-invariants.md` — the mechanics and the history behind the rules above.
- `docs/deepseek-integration.md` — why DeepSeek is shaped the way it is.

Some files were not shown because too many files have changed in this diff Show More