Compare commits

...
Author SHA1 Message Date
Codeman maintainer a017e9a8e0 chore: version packages 2026-09-12 06:10:14 +02:00
Codeman maintainer 65ddedd1d4 fix: act on the 1.27.0 pre-release review
A Fable 5.1 reviewer read the whole release diff against 1.26.2 and returned
SHIP WITH FIXES. These are its findings, verified before acting on each.

**The changelog advertised a feature the code refuses (major).** The #401
changeset and docs/web-tabs.md both listed `*.localhost` in the loopback set.
The follow-up in 02b0e278 moved it out of the auto-route set on security
grounds and updated CLAUDE.md but neither of those, and that changeset becomes
the 1.27.0 CHANGELOG entry: a user would have read the release notes, tapped
`http://app.localhost:3000/` on a phone and got a connection error from a
documented feature. Both corrected, and the user guide now says why it is
excluded and that adding such a dashboard by hand still works.

**Dictation delivered its text twice (minor, #388).** `keydownSnapshot` started
`null`, so `keydownSnapshot ?? canonicalCount` at the input event read a counter
xterm had ALREADY bumped: on a fresh page load with no keydown yet, xterm's own
capture listener forwards the `insertText` itself (it is not gated behind a
keydown), then the snapshot equals the bumped count, `count > snapshot` is
false, and the controller emits the same text again. Reproduced directly
against the module: it emitted `hello` for input xterm had already delivered.
A `0` baseline restores that file's own invariant, that a missed recovery is
acceptable and a duplicated keystroke is not. Two regression tests, covering
both the xterm-already-delivered and genuinely-dropped halves.

**The sorted rail's arrow-key walk followed the DOM (minor).** `_tabKeydownHandler`
steps `querySelectorAll` order, which is `sessionOrder`, while a sorted rail
paints its rows with the flex `order` property, so ArrowDown from the top card
landed wherever that session happened to sit in the tab order. It now sorts its
node list by the COMPUTED order first: computed rather than inline, because web
tabs take their `order: 9999` from CSS and would otherwise read as 0 and lead
the walk. This is the one place that follows the paint; the Alt+N badge, the
drag model and the filter all still deliberately read the DOM.

**A trusted dashboard was auto-reused by a tapped link (minor, #401).** The
reuse loop skipped `managed` and direct-mode records but not `trusted`. A
trusted frame is mounted with `allow-same-origin`, i.e. on Codeman's origin
with the user's cookie, and these links come from agent output, which is the
threat model the loopback allowlist was just narrowed for. An agent that can
write into the dev server's tree could print a path that one tap opens inside
that privileged frame. Excluded from auto-reuse, with a test; opening it from
the Run dropdown is still an explicit action and unchanged.

**Two documentation claims that were no longer true.** CLAUDE.md said
test/location-overlay-commands.test.ts pins every remote pane command, but
remote claude and remote omp now have their own arm in `buildRemoteLaunchCommand`
and never reach `defaultRemoteCommandForMode`, which is what that test asserts,
so it pins nothing for them and changing either arm will not fail it. Named the
real pins instead. Also documented the arrow-key-walk exception in the rail
paragraph.

Left as follow-ups, deliberately: `POST /api/webviews` does not dedupe by URL
server-side, so two devices tapping one link concurrently can still save two
dashboards for one origin (pre-existing endpoint behaviour that #401 makes
reachable by a tap), and the location-overlay golden should assert the real
remote claude/omp commands rather than a branch neither reaches.

Full gate green: 359 files, 6869 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 06:09:43 +02:00
Codeman maintainer 8b23f3e260 feat(rail): sort the vertical tab rail by activity, and give its rows the home screen's card
The vertical rail lists exactly the sessions both home screens list, so it now
answers their question the same way instead of showing the raw tab order.

Order: new per-device `tabRailSort` (App Settings -> Appearance -> Tabs ->
Vertical Rail Order, default "By activity"). It runs `CodemanSessionOrder` over
rows classified by `_mobileOverviewState`, i.e. literally the home screens'
comparator, `lastSubmitAt`-anchored running group included.

It is applied as the flex `order` property, never by reordering the DOM.
`#sessionTabs` stays in `sessionOrder`, which is what keeps the Alt+N badge
honest (it names a shortcut, not a row position, so it deliberately does NOT
run 1,2,3 down a sorted rail), and keeps drag-and-drop, the arrow-key walk, the
sidebar filter and `_scrollActiveTabIntoView()` all reading the list they
always read. A session changing state then moves one inline style instead of
forcing the full rebuild that would restart every card's animation on every SSE
tick. The incremental render path re-applies it, since a state flip adds no tab
and never reaches the full rebuild, and an empty string is what clears it when
sorting stops. Web tabs are pinned past the cards by a CSS `order: 9999`, since
`renderWebviewTabs()` emits the same markup for every layout and the flex
default of 0 would interleave them. Drag is switched off while sorting (the
drop rewrites `sessionOrder` correctly and the sort puts the card straight
back, so the affordance would be a lie); 'manual' is the way back.

Cards: detailed rail rows become bordered cards on `--bg-card`, with the stamps
line on its own full-width row and the pill at its right end. The state dot
goes 6px to 9px, keeps its orbiting ring while working and gains the green
halo; idle mutes toward `--text-muted` as the home rail does. Needs/error/
waiting reuse `home-sessions-blink-red`/`-yellow` rather than a second copy.

These card rules are RAIL-SCOPED and deliberately absent from the comma-grouped
selectors that carry both vertical surfaces: the rail is an occasional,
resizable list you scan, while the sidebar is a permanently-docked nav column
where 20 stacked cards read as a wall. Every state-dot rule also excludes
`.tab-alert-action`/`.tab-alert-idle` by hand, because those alert rules are
only (0,3,0) and these are (0,5,1)+.

Lines: the lineage bracket already drew in the rail, but its track sat 6px from
the left edge, so half of its 11px outer glow was clipped by the window frame
and it read as a thread pinned to the frame. It now runs at 10px, mid-channel
in the gutter the rail already reserves.

Tests: test/tab-rail-order.test.ts (17) drives the real `isTabRailSorted()` and
`_tabRailSortOrder()` out of app.js, covering the row model (a WORKING row
ranked by `lastSubmitAt`, which would otherwise rank every running turn as
freshly started and fail no rendering test), the Alt+N badge, and the opt-out.

Verified in Chromium across sorted/manual/simple/header-strip/sidebar with the
setting flipped at runtime: no page errors, and the header strip and sidebar
render byte-identically to before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 06:09:28 +02:00
Codeman maintainer 02b0e27898 fix: merge-time follow-ups for #400, #401, #362 and #388
Each item is from the pre-merge review of the PR it names, applied on master
rather than by pushing to a contributor branch.

#400 (response viewer, shenlvkang-collab)
- The brief view opened at `scrollTop = 0`, right when it was a single card
  holding the last row. Now that it renders the whole turn, the top is the
  turn's first narration line and the answer can be screens below it, while
  loadFullContext already scrolls to the bottom of the same turn. A multi-row
  turn now opens at its newest text; a single card still opens at the top.

#401 (loopback links as web tabs, shenlvkang-collab)
- Drop `*.localhost` from the auto-route set. Every other member is an address
  literal that can only mean this box; a `*.localhost` DNS name is not one, and
  a resolver with a search domain retries `evil.localhost` as
  `evil.localhost.<search domain>`. The link source is agent-written terminal
  output, so that set is the whole confinement on a tap that makes Codeman
  fetch a URL server-side and persist it. The page-side test stays broader
  (`isOnBoxHostname`), where a false positive only declines to proxy.
- A link to the origin root navigated nothing: the path was flattened to '',
  which openWebview reads as "no deep link", leaving an open frame where it was.
- `this.webviews` being set does not mean it is loaded. initWebviews() assigns a
  truthy empty map and only then awaits the list, so a tap during page load
  found nothing to reuse and POSTed a duplicate record. Join the in-flight
  refresh instead.
- One dashboard per dev server rather than per host spelling, which is what the
  method's own comment already promised.
- Toast on the auto-create: it writes webviews.json, broadcasts over SSE and
  adds a Run-dropdown row on every signed-in device, with a new tab as its only
  previous signal.

#362 (remote omp continuation, timkjr)
- Accept the allowlisted `mode === 'omp'` arm as-is; a blanket registry render
  would hand deepseek a locally-resolved --profile and bypass claude's own
  overlay. A registry-declared switch is the follow-up if a third mode needs it.
- Revert the whole-file Prettier reformat of docs/remote-sessions.md (docs/ is
  hand-formatted and outside `npm run format`), keeping only the two new
  sections.
- Correct three stale passages: architecture-invariants' `exec claude
  --dangerously-skip-permissions`, the `exec <cli>` paragraph (claude and omp
  now have their own arms, and the claude pane's PID is the login shell), and
  omp-integration's `-c 'omp'`. RemoteCommandMode gains deepseek and omp.
- Add the missing `_maybeCaptureOmpSessionId` remote-guard test; the sibling
  guard in `_pinOmpRespawnId` had one and this path runs earlier, on the first
  idle turn.

#388 (keyCode 229 recovery, aakhter)
- Gate notifyCanonicalData on shouldSuppressTerminalQueryResponse and
  isTerminalFocusOrMouseReport. onData also carries the DA/DSR/CPR/OSC replies
  xterm answers during Ink redraws and its SGR mouse and focus reports; any of
  those landing between the keydown and the candidate's resolution was read as
  "xterm spoke for this keystroke", standing the recovery down and leaving the
  character dropped, worst on a busy agent pane. Reached through
  window.CodemanTerminalInput: the predicates live in a module IIFE that closes
  long before this call site, so bare references would throw into the
  surrounding try/catch and stop the notify from ever running.

Every fix has a test that fails without it (verified by reverting each).
Full gate green on the combined tree: 358 files, 6849 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 05:15:41 +02:00
Ark0N 9d664ffe01 Merge pull request #388 from aakhter/pr/keycode-229-input-recovery
fix(terminal): recover dropped keyCode 229 input (Android/IME)
2026-09-12 05:14:39 +02:00
Ark0N a28b04c368 Merge pull request #362 from timkjr/feat/omp-remote-continuation
fix(omp,remote): thread remote-omp resume/continue through respawn and reattach
2026-09-12 05:14:24 +02:00
Ark0N e35b68e253 Merge pull request #401 from shenlvkang-collab/pr/loopback-links-webtab
feat(webview): open localhost links through a proxied web tab from another device
2026-09-12 05:14:09 +02:00
Ark0N 77d9ad59f7 Merge pull request #400 from shenlvkang-collab/pr/claude-viewer-last-turn
fix(web): show the whole last turn in the Claude response viewer's brief view
2026-09-12 05:13:55 +02:00
shenlvkang-collabandClaude Fable 5.1 d9eeb039db feat(webview): open localhost links through a proxied web tab from another device
An agent prints `http://localhost:5173/` (a dev server, a preview it just
served) and the user taps it on a phone. That address only exists on the
Codeman box, so the link was a guaranteed connection error from any other
device — while the web-tab proxy fetches from the server, where it works.

A loopback link (`localhost`, `*.localhost`, 127/8, 0.0.0.0, ::1) activated
in the terminal or clicked in the Response Viewer now opens as a proxied
web tab whenever the Codeman page itself is not on that box. A saved
proxied dashboard on the same origin is reused, with the link's own path,
query and fragment opened inside it (a mounted frame is navigated, not torn
down, so its state survives); otherwise one is saved under its host:port,
sandboxed like any other web tab, so it is in the Run dropdown next time.

Only loopback is routed this way. A LAN or tailnet address may well be
reachable from the device (a VPN, the same Wi-Fi) and a direct open is the
cheaper, richer path, so those keep opening in a new browser tab; on the
box itself every link opens directly. The terminal link provider and the
viewer's click handler consult one hook and fall through to their existing
behaviour when it declines.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
2026-09-10 12:48:47 +08:00
shenlvkang-collabandClaude Fable 5.1 bd61735393 fix(web): show the whole last turn in the Claude response viewer's brief view
The eye button rendered `data.text`, which is one row: the last assistant
row of the transcript. A Claude answer is a median of 3 model messages
(p90 11) split around tool calls, so the brief view usually showed the tail
of an answer ("Done.", "Let me look.") and the substance appeared only after
More. The full view was fine, which is why the brief one read as broken by
comparison.

The brief view now asks `?context=turn`. The reader answers with the
assistant messages of the last ANSWERED turn (`selectLastAnsweredTurn`: the
highest `turn` that has an assistant row, so a prompt queued after the
answer does not blank the view) and the frontend renders them exactly as
the full view renders that turn: one badge, then continuation segments,
gated on the numeric `turn` as before.

`data.text` is unchanged in every context — still the last assistant row,
never `messages.at(-1)` — because agent pollers hash it. Readers that emit
no turns (Codex, the pane parser, DeepSeek, an older server) return `text`
only for `context=turn`, and the brief view keeps its single card for them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
2026-09-10 12:45:16 +08:00
Codeman maintainer 5b667264b4 chore: version packages 2026-09-10 03:20:58 +02:00
Ark0N e3d5fd90cd Merge pull request #392 from JDProfresh/fix/ios-safari-toolbar-gap
fix(mobile): lift the iOS Safari toolbar by the measured chrome overlap
2026-09-10 03:10:12 +02:00
Ark0N 713f632a64 Merge pull request #397 from irisitymichaelgrundberg/fix/detached-session-owns-its-pane-size
fix(terminal): let a detached session's own window own its pane size
2026-09-10 02:58:54 +02:00
Ark0N 92b5dfacb0 Merge pull request #396 from irisitymichaelgrundberg/fix/terminal-font-settle-before-fit
fix(terminal): fit the terminal only once the terminal font is measurable
2026-09-10 02:58:48 +02:00
Ark0N 890a1b0902 Merge pull request #395 from irisitymichaelgrundberg/fix/full-history-replay-row-alignment
fix(terminal): keep row alignment in the full-history pane replay
2026-09-10 02:58:42 +02:00
Ark0N a360763890 Merge pull request #394 from irisitymichaelgrundberg/fix/ctrl-v-pastes-twice
fix(paste): handle only the first paste event the Ctrl+V trap receives
2026-09-10 02:58:36 +02:00
Codeman maintainer 57899f879e feat(ui): add a Blur entrance animation on all four surfaces
An iOS-style focus pull: the thing arrives out of focus and the blur fades
off it as the opacity comes up. Opacity leads the blur (full opacity around
45%, blur still lifting), which is what separates it from a cross-fade.
Ships on tabs (440ms), agent windows (560ms), the terminal pane (520ms) and
connection lines (380ms), plus a `Soft focus` theme that sets all four.
Default stays `legacy`, so an untouched install is unchanged.

The terminal pane is the one surface that cannot blur itself the documented
way, and `blur` takes a deliberate exception to the "never a filter on
.terminal-container" rule. Every alternative was measured against a live
xterm and does not work: a backdrop-filter veil on ::before blurs perfectly
while STATIC, and Chrome silently drops the backdrop the moment ANY
animation runs on that pseudo-element (the veil computes blur(15.3px) while
the text behind it stays razor sharp); driving the radius from rAF buys the
same full-screen blur per frame plus main-thread work. The cost the rule
exists to avoid is inherent to blurring a terminal, so the style buys it
knowingly: opt-in, off by default, one ~520ms run per session open, class
straight back off, will-change still unset. Worst-case price, headless
SwiftShader with no GPU: frame deltas 16.7ms -> 33.3ms for the run, against
16.7ms flat for `fade`. cols x rows measured unchanged at 178x38 before,
during and after, so FitAddon never sees it.

The line entrance animates `filter` too, where each line already carried
its glow. Both kinds now hold it in --line-glow and both keyframes say
`blur(N) var(--line-glow)`, so the function lists match and interpolate
instead of the glow vanishing for the run and popping back (a lineage
line's glow is a different colour, set per element). Its 100% frame omits
`opacity` on purpose so the endpoint comes from the element's own resting
value: 0.9 subagent, 0.72 lineage, 0.95 working.

test/entrance-animations.test.ts is a new static guard over the whole
feature, not just this style: the rule -> keyframes -> theme-option chain a
style silently does nothing without, the terminal's paint-only property
allowlist (the FitAddon rule), the --line-glow contract, and reduced-motion
coverage. Mutation-checked both ways.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 02:57:40 +02:00
Codeman maintainer d4fe3afc9d feat(files): raise the download cap to 2GB and stream /api/download
The 50MB cap on file-raw, the attachment /raw route and /api/download was
memory protection for a `readFile()` that no longer exists: file-raw and
/raw were rewritten to stream through `sendFileBody()` and answer Range
requests, so size costs a read stream rather than RSS (measured: a 600MB
download moved peak RSS by ~37MB). All the cap still did was refuse
legitimate downloads of build artifacts, videos and archives.

It is now MAX_FILE_DOWNLOAD_BYTES in config/buffer-limits.ts, default 2GB,
env CODEMAN_MAX_DOWNLOAD_BYTES, 0 = unlimited. `parseByteLimitEnv()` is
separate from the `parseInt(...) || default` idiom used elsewhere in that
file precisely because that idiom reads 0 as falsy and would silently
restore the default for the one value that means "no limit".

/api/download was the last route that really did buffer the whole file. It
now shares sendFileBody() with the other two, so it streams, advertises
Accept-Ranges, and is resumable. Its Content-Disposition also goes through
buildContentDisposition() rather than raw interpolation.

Refusals move from 400 to 413 across all three, which is the correct status
for the case; with the cap at 2GB it is a path almost nothing reaches now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 02:57:07 +02:00
Michael Grundberg 77fcd65b4a fix(terminal): force the re-measure, bound the wait, and test both
Review of the previous commit found that waiting for the font does not, on its
own, do anything.

`FitAddon.proposeDimensions()` measures nothing — it divides the container by a
CACHED cell size, and xterm refreshes that cache only from `open()`, from a
resize that actually changed the grid, and on a device-pixel-ratio change.
Nothing in it listens for font loading. So a fit that runs after the font
arrives can still divide by the fallback cell, propose the grid it already has,
and short-circuit before anything re-measures. The wait now ends by calling
`_charSizeService.measure()` itself, which is the step that makes the following
fit see the real font. Private API, as FitAddon's own dependency on `_core` is,
and guarded because a terminal can be disposed mid-wait.

The wait was also unbounded, and it sat behind the buffer-load gate. Neither
`FontFaceSet.load()` nor `FontFaceSet.ready` has a deadline, so a font request
that never settled left the tab spinning with live output queued behind it —
permanently, and on every session, since they share one promise. The comment
claimed the opposite ("a font that never loads must not block the terminal, so
this always resolves"), which was true of the per-face loads and false of
`ready`. It is now raced against TERMINAL_FONT_WAIT_MS, and the await moved
ahead of `_beginBufferLoad` so a slow font cannot hold output back at all —
which also removes the stale-select interaction with `_restoringFlushedState`,
since that flag is not yet set when the wait runs.

The awaited set no longer includes faces that cannot move the measured cell.
The bundled symbols font is ~1.2MB of private-use-area glyphs and xterm
measures `W`, so awaiting it put a megabyte in front of the first frame for
nothing; the generic families match no FontFace at all.

A runtime font change had the same race the boot-time one did:
applyTerminalFontFamily wrote the new family and fit on the next line, against
a family the browser might not have loaded. It now re-arms the wait and fits
again when it settles.

The claim that this could not be tested was wrong: the repo's vm harness
reaches both halves. The new suite pins the family filter, the forced
re-measure, the deadline, a rejecting load, a browser with no font API, and a
terminal disposed mid-wait — plus the four ordering properties in
selectSession, including that iOS Safari's synchronous focus still precedes the
first await. Each assertion was checked by reverting its fix.

Also corrects the docstring's reason for calling `document.fonts.load` (the
stylesheet is render-blocking and long parsed by then; the real reason is that
the WebGL renderer rasterises through a canvas atlas, and canvas text never
triggers a CSS font fetch), restores the JSDoc block the previous commit
displaced from getTerminalDimensions, and fixes a comment that described the
first fit as already having run when the mobile-Safari branch defers it.
2026-09-09 16:57:42 +02:00
Michael Grundberg 2b57c595df fix(terminal): gate the row-preserving skips on a capture, not the query flag
Review of the previous commit found the guard inverted: the three skips keyed
on `?full=1`, which is only what the client asked for. When the capture comes
back null — ENOBUFS, a timeout, a vanished pane, or a session with no mux at
all — the reply falls back to the byte history, which IS a stream of
successive frames and still needs stripping. Gating on the request returned it
whole: measured at 82KB against 4KB for the same buffer without `full=1`. A
direct-PTY session takes that path on every first selection, not only during
an outage. The skips now key on `isFullCapture`, meaning a capture arrived.

Three further defects the same review surfaced, all on this path:

Keeping the trailing rows is only sound when a cursor move follows to count
back up from them. On the two branches where the cursor query fails there is
no move, so the caret was left at the bottom of the pane — worse than before.
The cursor is now read first and settles both decisions together.

The move is relative rather than absolute. `CUP` numbers rows from the top of
the browser's screen, so it is only right while the browser's row count equals
the pane's, and `resizeWindow` does not wait for tmux, so a capture can be
taken before a requested resize applies. Measured against real tmux with a
browser four rows shorter than the pane: the absolute move lands on a blank
row, the relative one lands on the caret's row.

An all-blank pane no longer reads as content. Retaining trailing rows and
appending a move made it non-empty, and the caller treats non-empty as "replay
this", so a blank screen would have replaced real history — the downgrade
`_replayWouldShrinkBuffer` refuses, arriving from the server side where that
guard cannot see it.

The documentation claimed one line per screen row. `-J` joins a hard-wrapped
row into its logical line, so that is false whenever any row wrapped: measured
at 10 lines for a 12-row pane. Both entries now say what actually holds, and
the stale "NOT repositioned" contract in the mux interface is updated too.

Tests: the byte-history fallback is stripped, an empty capture leaves history
intact, and the extracted helpers are unit-tested directly rather than through
source-text matching. The slice window in the capture test is bounded at the
next method, having overrun into its neighbours.
2026-09-09 15:35:02 +02:00
Michael Grundberg 070e8da81b fix(terminal): yield only the resize send, and take sizing back on redock
Review of the previous commit found four defects in it.

The guard sat above the local fit, so it suppressed a reflow as well as the
server write. tab-rail-resize performs its single settle-time refit through
sendResize and has no fallback for a truthy activeSessionId, so dragging the
rail stopped reflowing a detached session's terminal in the dashboard. The
mobile-keyboard guard fourteen lines below already draws the line correctly —
withhold the send, never the reflow — and the guard now sits after the fit.

_lastResizeDims is one value for the whole window, and both guards skip
updating it, so while a popup owns a session that value no longer describes
the PTY. _redock repaired it only for the active session. Pop out A, switch to
B, close the popup: selecting A later found unchanged dimensions, returned
"unchanged", and selectSession skipped its 400ms redraw wait — while the
server, comparing against the real pane, did resize and did raise SIGWINCH, so
the fetch painted the pre-redraw frame. _redock now clears the record on every
path, active or not.

_redock could also fire a resize for a session already gone: _onSessionDeleted
redocks before cleanup, so the id can be dead and the request is a guaranteed
404. It now checks the session still exists.

restoreTerminalSize — the header's redraw button and Ctrl+Shift+R — silently
did nothing for a detached session while still reporting success with
dimensions nothing was set to. It now says the session is sized by its own
window, where the same button works.

The `force` comment claimed a client-side dedupe that does not exist; the
deduplication is server-side against the real pane. Corrected to say what the
flag actually buys. The _redock doc comment now records that the function
writes to the server and is not idempotent.

Tests: _redock was the untested half and is the half three of these defects
sit in. It now has coverage for clearing the stale record on both the active
and inactive paths, re-asserting only for the session being shown, and staying
silent for a deleted session. The existing sendResize test now asserts the
local fit still runs.
2026-09-09 15:24:55 +02:00
Michael Grundberg 0e82443222 fix(terminal): fit the terminal only once the terminal font is measurable
Opening a session could render its frame with characters spliced into each
other, as though two frames were overlaid — a status-line fragment landing
in the middle of a file path, for instance. Resizing the browser window
cleared it.

The first fit runs while the browser is still painting with a fallback font.
A cell measured against that font has a different width and height from one
measured against the terminal font, so the fit produces the wrong column and
row count. Codeman sizes the pane to it and replays the capture. When the
font finishes loading the measurement changes, the pane is resized a second
time, and the CLI repaints for a shape that does not match the frame already
on screen. Its later partial updates then land on the wrong rows.

selectSession now waits for the font before it measures, so the pane is
sized once, at the size that sticks, and the capture is taken at that size.
The wait always resolves, so a font that never loads cannot block a
terminal, and it resolves immediately once the font is in, so a tab switch
pays nothing after the first load.

document.fonts.ready alone is not enough: it can resolve before the
stylesheet declaring @font-face has been parsed. document.fonts.load for
each family in the stack is what actually requests the faces.

Measured on a session opening at 2328px wide: the cell went from 8.43x16.00
to 8.00x21.00 roughly 900ms in, moving the grid from 112x36 to 118x28 after
the replay had already been painted.
2026-09-09 14:20:34 +02:00
Michael Grundberg 323730a29d fix(terminal): keep row alignment in the full-history pane replay
Switching to a session left the caret one row below the composer's input
line, on the box border, and every cursor-relative update the CLI sent
afterwards was measured from the wrong row. Any fresh output repaired it,
because the CLI then repainted the whole frame.

Two things were wrong with the full-history replay, and they compound.

The capture never restored the cursor. The visible-frame path ends with an
absolute cursor move back to the pane's position; the linear path returned
its text and left the caret wherever the last character landed, which for an
agent CLI is the bottom-most row carrying text — the status line.

The rows it addressed did not line up with the pane's rows either. Four
transforms ran over the capture and each can delete a line: the trailing
blank rows were stripped, redraw-bloat stripping ran, the trim that cuts
everything above the Claude banner ran, and leading whitespace was removed.
All four are right for a byte stream of successive frames. A capture is the
rendered pane, one line per screen row, so each deletion shifted the frame
out from under the restored cursor.

The full-history path now appends the pane's own cursor position and keeps
every row, so row N of the reply is row N of the pane. The visible-frame and
tail paths are untouched.

Restoring the cursor is what makes row alignment load-bearing here, and
neither CLAUDE.md nor the architecture invariants said so — which is how
four line-deleting transforms accumulated on the path. Both now record it.

Verified against a live 315x59 pane: the reply carries 59 rows, its row 55
is the composer's input line matching tmux, and it ends with the cursor move
that lands there.
2026-09-09 14:20:34 +02:00
Michael Grundberg 5ac516dd3b fix(terminal): let a detached session's own window own its pane size
Popping a session out left both windows sizing the same pane. The dashboard
keeps the session active and keeps measuring it, and its terminal is
narrower than the popup because the session rail takes width the popup does
not have. One PTY cannot hold two sizes, so the CLI drew frames that fit
neither window and the popup showed a garbled frame.

sendResize and the debounced window-resize handler now stand aside for a
session this window has marked detached. A solo window is exempt, since it
is the owner. _maybeRefetchFullHistory already stood aside on exactly this
condition, so the rule is not a new one.

Sizing has to come back when the popup closes: while it owned the session
the dashboard sent no resizes, so the PTY still holds the popup's geometry.
_redock now re-asserts, with force, because the dimensions the dashboard
last sent are the ones it is about to send again.

Reproduced with a dashboard and a popup on one session: before, the pane
sat at 315 columns while the popup rendered 289. After, both report the
same size and the popup's frame matches the pane exactly.
2026-09-09 14:20:34 +02:00
Codeman maintainer 5130ca6633 fix(pr-bot): fail fast when the review model's budget is spent
Claude Code answers an exhausted model budget INSIDE the turn ("You've
reached your Fable limit. Run /usage-credits to continue or switch models
with /model.") and then sits there with nothing to write. The reviewer never
produces a report, so `runTurn` waited out its full 40-minute deadline and
reported a bare "timed out after 40 min without a report", which reads as a
hung reviewer rather than an account that needs attention.

Measured on 2026-09-08: #388, #393, #394 and #377 each lost 40 minutes this
way, and because every attempt counted, all four reached MAX_AUTO_RETRIES and
would NOT have been picked up again once the budget returned. One spent
afternoon quietly took the whole queue out of service.

`findModelLimitNotice()` reads the notice off the pane and `runTurn` returns
a new `limit` outcome instead of waiting. It is consulted in exactly two
places, both of which mean "the turn produced nothing": on a stop where
`isDone()` is still false, and on each timed-out wait slice. A review that
merely discusses usage limits in its own findings therefore cannot be
mistaken for one that hit the wall, and the pattern matches neither the model
name nor a straight apostrophe, since the pane renders a typographic one and
every model prints the same sentence.

A spent budget is an account condition, not a bad PR, so it no longer spends
the per-head retry budget: the queue resumes by itself when the budget does.
Telegram now names the cause and the file to change.

Tests use the pane captured verbatim off the run that lost the 40 minutes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 19:27:43 +02:00
Michael GrundbergandClaude Opus 5 b87bc6871b fix(paste): handle only the first paste event the Ctrl+V trap receives
Ctrl+V in the terminal inserted the clipboard text twice. Right-click →
Paste inserted it once.

`_handleImagePaste()` appends a hidden contenteditable div, focuses it, and
reads the clipboard out of the paste event that lands there. Two separate
routes deliver that event for a single keypress. The function issues
`document.execCommand('paste')` itself, which in Firefox dispatches a
trusted paste event and then returns false, because the trap cancels the
event and the command never completes; Chromium and WebKit refuse that
command and dispatch nothing. The keydown's own default action delivers the
other, because xterm calls the custom key handler before its own `cancel()`,
so returning false never calls preventDefault. Firefox therefore ran the
trap's listener twice and both runs reached `terminal.paste()`. The
context-menu paste involves no keydown at all, which is why that path stayed
correct.

The trap now accepts the first paste event and cancels every later one, so
how many paste events a browser delivers no longer changes what the PTY
sees. Measured on a live install, one Ctrl+V each: Firefox two events and
two writes before this change, Chromium and WebKit one and one, and every
engine one write after it.

The `execCommand('paste')` call stays. Stripping it out also ends the
doubling, and all three engines still deliver one event without it, since
`trap.focus()` has already run when the key's default action resolves. It is
kept because the trap technique arrived in #84 for plain HTTP and for
mobile, and a desktop measurement says nothing about real iOS Safari or
Android Chrome: where a browser aims the default action at the element
focused when the keydown began, the command is the only route into the trap,
and the trap is the only place clipboard image blobs are read.

test/image-paste-trap.test.ts loads image-input.js into a `node:vm` context
with a fake document and fires two paste events at the trap. It covers text
and images, and fails on the old code with the text pasted twice and the
image uploaded twice.

Docs: the invariant goes into docs/architecture-invariants.md as a Terminal
paste section and into CLAUDE.md as a Frontend entry, both recording the
measured event counts and why the redundant call is still there. README.md
and the Keyboard Shortcuts and Input and Voice wiki pages gain a Ctrl+V row,
which all three tables were missing while listing every other clipboard
binding.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 16:49:50 +02:00
JD c367b12f77 fix(mobile): lift the iOS Safari toolbar by the measured chrome overlap, not 100vh minus the visual height
The phone block lifted the toolbar (and padded .main) by (100vh - --app-height) on iOS Safari to clear a bottom bar that position: fixed elements were assumed to sit behind. On iPhone Safari fixed elements already stop above the bar, and 100vh is the large viewport with the bar collapsed while --app-height is the visual viewport with it expanded, so the expression measures the bar's collapsible height and shows up as an empty band between the toolbar and the bar whenever the bar is expanded. The terminal was padded by the same amount.

The lift is now --chrome-overlap, set in updateAppHeight() as innerHeight minus the visual viewport height: the distance the layout viewport that anchors fixed elements extends past the visible area. That is 0 on iPhone Safari, so the toolbar meets the bar, and it is the overlap itself on any browser where fixed elements really do land behind the chrome, so those keep the lift. The keyboard-visible rules, which already override the toolbar offset, are unchanged.
2026-09-08 01:39:07 -04:00
timkjrandClaude Sonnet 5 797f0d387c fix(remote): address review feedback on omp/claude respawn continuity
- Remote omp command now renders through buildSpawnCommandFromRegistry
  (the mode-agnostic engine local/docker spawns use) instead of the
  buildOmpCommand() the CLI-registry refactor deleted.
- Session._pinOmpRespawnId()/_maybeCaptureOmpSessionId() now skip
  host-local ~/.omp resolution entirely for a remote session and fall
  back to --continue: that resolver only ever reads THIS host's
  filesystem, which is meaningless (and could wrongly alias an
  unrelated local conversation) for a conversation that lives on the
  remote host.
- Remote-claude launch now honors an explicit resumeSessionId distinct
  from sessionId (mirrors claudeDockerPaneCommand's shape), and
  validates sessionId the same way that sibling does before
  interpolating it into the remote shell command.
- Add the still-missing header-cwd half of the trailing-slash test,
  and document respawn/reattach continuation + auto-reconnect-vs-
  clean-exit in docs/remote-sessions.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 22:11:54 -05:00
timkjr 88243e9ffa fix(remote): never auto-revive a remote session after a clean agent exit
The COD-108 reconnect watcher treated any dead local pane as a dropped
transport and re-ran the pane command — so a normal ctrl-c/ctrl-d on a
remote omp/opencode/claude auto-spawned a FRESH agent (claude only
looked correct because its '--session-id || --resume' fallback resumed,
with a loud 'already in use' error first).

Distinguish a transport drop from an intentional exit: only reconnect
when the durable remote tmux session (codeman-ssh-*) is verifiably
still alive on the remote host. A clean exit tears that session down;
the watcher now probes it via ssh has-session and skips (remote-gone)
when it is gone OR unknown (fail closed). The probe is cached
per-session and fired async so the 5s tick never blocks on ssh.

Also thread ompConfig/resumeSessionId into the remote builders so a
dead-pane respawn of an omp session resumes (--resume <id>) or
continues (--continue) instead of launching bare omp.

Tests: 3 new cases pinning remote-gone / unknown / alive decisions;
remote omp resume + --continue fallback. Verified live: all three
remote CLIs stay dead after exit.
2026-09-07 21:30:31 -05:00
timkjr 0a5bc1ac2e fix(omp,remote): pin remote conversations on respawn so ctrl-d/ctrl-c resumes instead of relaunching fresh
Two independent defects made ANY clean exit from a remote SSH session (user
ctrl-d or ctrl-c, or a dropped pane) relaunch the agent as a NEW conversation:

1. SSH-remote claude was launched as a bare `claude --dangerously-skip-permissions`,
   so the remote-respawn path (COD-108 reattachRemote re-running the idempotent
   launch command) started a fresh conversation every time. Pin it to the
   deterministic Codeman session id, mirroring the docker-claude shape
   (claudeDockerPaneCommand): `--session-id <id>` to create, with the
   `|| --resume <id>` fallback so the idempotent re-run resumes instead of
   erroring with "already in use". A per-host commands.claude override still wins.

2. OMP --resume pinning silently degraded to ambiguous `--continue` whenever a
   case path ended in a trailing slash (e.g. remote `remotePath` stored verbatim
   as `/home/user/dotfiles/`): mangleOmpWorkingDir produced `-dotfiles-` while
   omp persists sessions under `-dotfiles`, readdirSync returned null for an
   existing dir, and findLatestOmpSessionId/resolveAndClaimOmpSessionId never
   matched. Normalize the trailing slash before mangling (new exported
   stripTrailingSlash) and compare the session header cwd against the same
   normalized value.

Both were found live 2026-08-29 on a remote OMP/Claude node: ctrl-c and ctrl-d
behaved identically, both relaunching a fresh session.
2026-09-07 21:20:41 -05:00
Aamer Akhter e8a93ada1f fix(terminal): forward the orphaned input event instead of replaying a guessed key
The previous shape guessed the character from `event.key` on keydown, re-emitted
it, and then tried to suppress a late canonical copy with a 250 ms
character-keyed dedupe. Review found three defects in that, all reproducible:
the dedupe matched on the character alone with nothing scoping a candidate to
the keydown that created it, so the same character typed twice inside the window
had its second, real byte swallowed; anything whose committed text differed from
`event.key` (Enter, IME punctuation) was delivered twice, because the dedupe
could never match it; and the trigger ignored `key === 'Unidentified'`, which is
what a soft keyboard reports, so it may never have fired where it was needed.

The input event already carries the committed text in `ev.data` — exactly what
xterm itself would have forwarded — so nothing has to be guessed. The controller
now only decides WHETHER to forward, by asking whether xterm produced canonical
data since the keydown that began the keystroke. No character-keyed matching
survives, so the first two defects are structurally impossible rather than
defended against, and nothing reads `key`/`keyCode`, so the third cannot recur.

Three details are load-bearing and each has a test that fails without it:

- The "did xterm speak?" snapshot is taken at KEYDOWN, not at the input event.
  `_keyPress` emits and sets `_keyPressHandled` before `input` fires, so a
  snapshot read at input time already contains that emission, reads it as
  silence, and delivers the character twice.
- Our `input` listener is registered with `capture: true`. The target is visited
  twice in the event path, so a capture listener calling `stopPropagation()`
  stops later BUBBLE listeners on that same target; xterm's `cancel()` runs
  exactly in the branch where it handled the input, so on bubble we would never
  observe handled events, and whether we observed them at all would hang off
  `options.cancelEvents`. Measured in jsdom and headless chromium; the table is
  in the module header.
- Enter is deliberately no longer special-cased. That mapping is what made the
  committed text differ from the re-emitted value in the first place.

The scope is also narrower than the old name suggests, and the browser test now
proves it rather than assuming it. For a keydown that reports keyCode 229 xterm
ALREADY self-rescues, via `CompositionHelper._handleAnyTextareaChanges()`
diffing the helper textarea on a 0 ms timer. A test asserting "we recovered it"
there passes while xterm does all the work, so the browser tests assert WHO
delivered the byte: zero canonical emissions for the genuinely orphaned case,
exactly one delivery for the case xterm rescues itself.

Also addresses review notes: the module gains an `@fileoverview` with
`@dependency`/`@loadorder` and an entry in the load-order list and module
inventory, and the wiring test moves out of the Ctrl+C smart-copy file into its
own. The keydown hook deliberately still runs for every key event rather than
moving behind the 229 gate: gating it would reinstate exactly the blindness
described above, and it is now a single counter assignment.
2026-09-07 19:11:20 -04:00
Codeman maintainer a164c07f92 chore: version packages 2026-09-07 22:54:36 +02:00
Codeman maintainer 4f2dfb4e6d fix(mobile): carry resumeId through the phone overview's past rows
#386 made Codex conversations resumable from Past Sessions, and
resumeMobileOverviewSession() correctly passes row.resumeId on to
resumeHistorySession(). The phone's own row projection never copied the
field off the unified-list item though, so row.resumeId was always
undefined there and a tapped Codex row started a FRESH session on a thread
that was already on disk. The desktop path worked; only the phone was blind.

The test fails without the projection line, and pins the other half too: a
claude row must not grow a resumeId, since the field is what distinguishes
"resume this conversation" from "start a new one".

Docs: CLAUDE.md and architecture-invariants both still described the unified
list as merging Claude transcript files. It has been three stores since this
PR (Claude's ~/.claude/projects, omp's ~/.omp/agent/sessions, codex's
~/.codex/sessions), the alias field keeps its Claude-era name without being
Claude-only, and the scanner-only rule behind resumeId was written down
nowhere.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 22:44:25 +02:00
Ark0N 344e93c824 Merge pull request #386 from irisitymichaelgrundberg/feat/codex-resume
Merging with the phone-overview resumeId fix and the two unified-list doc passages applied on master.
2026-09-07 22:43:46 +02:00
Codeman maintainer f1b7283393 fix(cli-registry): guard workDetect.workingLine like every other config regex
#385 made the composer glyph and the working status line per-CLI registry
data, which is right, but `workingLine` arrived as a config-supplied regex
validated with a bare `new RegExp()`. That skips `compileVersionRegex()`,
the helper the registry uses for exactly this: a `~/.codeman/clis.json`
override can set the field, the compiled pattern is run against every
accumulated PTY chunk and every pane capture, and a nested quantifier there
backtracks on the event loop for the whole server rather than one session.

Route it through the helper in both places, which are not redundant: the
schema refine rejects the entry at LOAD time so a bad pattern never reaches
a session, and `_workingLinePattern()` compiles through the same helper so
the runtime cannot hold a pattern the schema would have refused. The helper
returns null instead of throwing, so the Claude-pattern fallback stops being
a try/catch and becomes structural. Both shipped patterns compile unchanged,
and Claude's is behaviourally identical to CLAUDE_WORKING_LINE_PATTERN.

Also match the Codex footer case-insensitively on the E. It was
characterised against codex-cli 0.152.1, which prints a lowercase `esc`;
a version capitalising it would make the whole fix silently inert, since
the pane would simply never look like it was working.

Docs: CLAUDE.md, architecture-invariants and cli-registry.md all still
stated the Claude-mode-only rule this PR retires, and none of them named
the new capability or the regex guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 22:42:20 +02:00
Ark0N a49be03f96 Merge pull request #385 from irisitymichaelgrundberg/fix/work-detection-external-clis
Merging with follow-up fixes applied on master: workingLine routed through compileVersionRegex() in both the schema refine and _workingLinePattern(), the Codex footer matched case-insensitively on the E, plus the doc passages that stated the retired Claude-mode-only rule.
2026-09-07 22:41:25 +02:00
Codeman maintainer 7fde978ce8 chore: version packages 2026-09-07 19:11:56 +02:00
Codeman maintainer 8ee7926e27 feat(agent-cases): tag agent-spawned case dirs and sweep their leftovers
A long orchestration creates one case directory per worker and deleting the
sessions never removed them, so ~/codeman-cases accumulated scratch folders
that were indistinguishable from real projects. They are now labelled and
have a cleanup path.

- src/agent-case-marker.ts: a case dir quick-start CREATES for an agent-driven
  spawn gets a .codeman-agent-case.json marker (when, by whom, parent session,
  mode). Only the create branch writes it, so a linked case, a cloned repo or
  any pre-existing path is never labelled; reading is total, so a malformed
  marker means "not agent-created" rather than a half-trusted entry.
- The signal is the new X-Codeman-Agent-Origin header the skill preamble sets
  on its shared curl (preamble bumped to 1.22.0), or an agentOrigin body
  field, falling back to a resolved parentSessionId so a worker spawned by a
  stale skill copy is still labelled.
- GET /api/cases publishes it as agentCreated; GET /api/cases/agent-created is
  a read-only cleanup listing adding inUse and modifiedAt; Add Case -> Manage
  badges each case and offers a review-then-delete sweep that names every
  directory in its confirm and skips any case a live session is working in.
  Removal stays on the existing DELETE /api/cases/:name.
- Agent preamble caches are collected too: ~/.cache/codeman-agent-<id>.sh was
  written per claude session and never removed (236 leftovers measured on a
  working machine). Now deleted with the session and swept at boot, guarded by
  a live-session keep set plus a 7-day age floor.

Verified end to end on an isolated instance: marker written for header, body
and lineage-only spawns, absent with no agent signal and for a pre-existing
directory; inUse flipping on session end; badge, sticky bar, confirm and sweep
driven in a browser; preamble seeded on create, removed on delete, boot sweep
taking only the aged orphans.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 19:09:24 +02:00
Aamer Akhter 82b090c74a fix(terminal): recover dropped keyCode 229 input
Android/GBoard-style keyboards fire keydown with keyCode 229 and, on some
paths, never mutate xterm's helper textarea. xterm has nothing to diff, so
it emits no data and the typed character is silently dropped: it never
reaches the PTY and never appears on screen.

terminal-keycode229-recovery.js is a standalone controller that re-emits
exactly those keys, and only once. xterm stays authoritative throughout:

- Only an explicit keyCode 229 keydown carrying a single printable key (or
  Enter) is eligible; Process/Unidentified/Dead, modifiers, AltGraph and a
  live composition are all left alone.
- The re-emit is scheduled from a microtask and then a zero-delay timer, so
  xterm's own textarea diff always gets the first opportunity; canonical
  data for the same key cancels the pending fallback.
- compositionstart and blur drop every pending candidate, so a real IME
  composition lifecycle is never second-guessed.
- After a recovery, one late canonical value attributed to that key token
  (via beforeinput/input on the helper textarea) is suppressed so the
  character cannot be delivered twice; the record expires after 250ms and
  an unattributed byte is never suppressed.

terminal-ui.js wires it at the two existing choke points — the custom key
handler and the onData registration, the latter now a named handler so the
recovery path can re-enter it — with both hooks wrapped so a failure in the
fallback can never break canonical input.

Unit coverage drives the module directly in a vm; the wiring itself is
covered end-to-end in the (browser-only) terminal-copy-shortcut suite.
2026-09-07 12:27:36 -04:00
Michael GrundbergandClaude Opus 5 2f9663e389 Merge branch 'master' into feat/codex-resume
master and this branch both rewrote the two `_claudeSessionId` resets inside
`start()`, so `src/session.ts` conflicted at both of them.

master's commit ccfda623 puts `restoredConversation` at the head of each
fallback chain. A restored mux attach means the CLI never stopped, so a
`/clear` before the Codeman restart may already have moved it to a
conversation the launch id knows nothing about. The persisted chain's tail is
that conversation, and the CLI's own hook reported it first-hand.

This branch adds `this._codexConfig?.resumeSessionId` to the same two chains,
so a resumed codex session keeps its thread-id alias across every mux reattach
and boot recovery.

Both fixes belong. Each chain now reads restoredConversation, then
_resumeSessionId, then omp's alias, then codex's alias, then the launch id.
The comments from both sides are kept.

test/session-claude-conversation-chain.test.ts pins the shape of those two
assignments by matching the source text, and its pattern named omp's alias as
the last term before `this.id`. Codex's alias now sits between the two, so the
pattern widens to pin the ends of the chain and let the middle grow. A `[^;]`
run cannot cross a statement boundary, so each match is still one assignment.

Checked on the merged tree: typecheck, lint, prettier and the frontend syntax
check all pass, and the CI suite runs 6721 tests green across 349 files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 08:48:31 +02:00
Codeman maintainer 61d22eee1c chore: version packages
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-07 00:06:00 +02:00
Codeman maintainer 92af855ce4 fix(base-path): keep the crash beacon under the mount, strip CODEMAN_BASE_URL in tests, add the wiring test
The merge-time items from the #381 review. navigator.sendBeacon is not fetch,
so the base-aware wrapper never saw the two crash-diag beacons and a sub-path
install posted them to the origin root every two seconds. The test suite now
strips CODEMAN_BASE_URL like CODEMAN_GESTURE, since the constructor reads it
as a fallback and an operator who exports it would see the root-install
byte-identity assertions fail. test/base-path-server.test.ts boots a real
WebServer under /codeman and checks the ingress strip, the base injection,
the rebased redirects, the 404 envelope and a prefixed WebSocket upgrade.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 23:11:01 +02:00
Ark0N cfc8fe7e41 Merge pull request #381 from mtiller/feat/reverse-proxy-base-url
feat(web): support a reverse-proxy base URL
2026-09-06 23:10:23 +02:00
Codeman maintainer 80397fe140 fix(hooks,mobile): the merge-time items from the #367 and #368 reviews
#367 (UserPromptSubmit hook): `hook:prompt_submitted` went on the wire
unregistered; it is now in both SSE registries (158 = 158), and the hook only
lands in the run summary when the conversation actually moved, since one row
per prompt would evict useful rows from the 1000-event FIFO and clutter the
Summary timeline and /api/search.

#368 (Add Case header submit): the pending-state dimming targeted the footer
button, which the <=860px layout hides, so on a phone the only visible submit
control stayed at full brightness while a clone ran. The header button now
dims too, and a static test pins the header-submit contract so it cannot
silently disappear again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 23:05:26 +02:00
Ark0N 7991f481b6 Merge pull request #368 from shenlvkang-collab/pr/mobile-add-case-submit
fix(mobile): give the Add Case modal a reachable submit button
2026-09-06 23:02:45 +02:00
Ark0N bca1b764cc Merge pull request #367 from shenlvkang-collab/pr/claude-conversation-first-hand
fix(session): learn the live Claude conversation from the CLI's own hook
2026-09-06 23:02:33 +02:00
Ark0N 1c1773278f Merge pull request #369 from shenlvkang-collab/pr/claude-response-viewer-per-message
fix(web): render one Claude response-viewer message per model message
2026-09-06 23:02:13 +02:00
Michael Grundberg 327e440607 fix(codex): fold a codex session into its own rollout row
Review fixes for #386.

Duplicate rows. A codex conversation showed twice, once live and once as a
past rollout row, because nothing aliased a codex session to its thread id.
That is worse than cosmetic: the stale row still resumes, so clicking it
starts a second `codex resume` on a thread already open in another pane.

  - A RESUMED session knows its thread id up front, so it folds from its own
    side: add `codexConfig.resumeSessionId` to the `claudeSessionId` chain.
    Not only in the constructor — `start()` recomputes that id at two further
    points (the mux branch, and the unconditional "third reset point" whose
    own comment already warned that omitting omp's fallback there stomps the
    mux branch's resolved alias). Both listed Claude's and omp's ids only, so
    for codex every mux reattach and boot recovery reset the alias back to
    the Codeman id and the duplicate returned.
  - A FRESH session has no thread id until codex writes the rollout, so it is
    folded from the other side. The scanner now reports
    `session_meta.originator`, which is `codeman_<sessionId>` for every pane
    Codeman spawns, and `gatherUnifiedInputs()` stamps the matching live and
    persisted rows, newest rollout winning (`/new` inside the TUI leaves
    several rollouts sharing one originator).
  - Persisted rows read `codexConfig.resumeSessionId` too. A resumed session
    demoted to a persisted-only record would otherwise lose its alias, and
    the originator fallback cannot rescue that one: a resumed rollout keeps
    its ORIGINAL session_meta, so it still names the pane that created the
    thread rather than the pane that resumed it.

Identity cache. It was written as soon as the thread id was known, but codex
writes the first user message only when the user submits, so any scan in that
window pinned `firstPrompt: undefined` for the life of the process — and the
home screen, the command palette and the search-index refresh all scan.
`shouldCacheIdentity()` now keeps an identity only once the prompt is known or
the head read filled its whole window.

Also from review: both caps count emitted rows rather than file index, so a
store of sub-agent threads no longer spends the `lastPrompt` budget before the
first row that needed it; the cache is an `LRUMap` sized like the one beside
it; the unreachable filename fallback is gone; a rollout recording no cwd is
dropped rather than emitted with `workingDir: ''`; and the unified-session
module header names all three transcript stores.

Tests. The resume wiring now has cases for a row with a thread id, a row
without one, and a `resumeId` on a non-codex row; the "no continuation is
wired" case narrows to gemini/antigravity, which is no longer true of codex.
`codex-resume-alias-survives-start.test.ts` drives a real Session through
`start()` rather than asserting on pre-stamped inputs — that gap is why the
reset points went unnoticed. Plus the maintainer's own cache repro, the
tail-budget case, a no-cwd case, and merge cases for both folds.
2026-09-06 21:53:38 +02:00
Codeman maintainer a2aaea3c0e docs(file-picker): state the Home/cases nesting the right way round, and document the new fallback chain
The two merge-time edits the #383 review asked for. The comment above the
picker's fallback chain said Home is nested under Codeman Cases; on the
native default it is the other way round (~/codeman-cases sits inside ~).
And the "Filesystem path picker" paragraph in architecture-invariants still
said the picker falls back to /mnt/d, which #383 changed to: the session's
Current Folder, then the Codeman Cases root, then /mnt/d, then the first
root. No code behaviour changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 21:20:45 +02:00
Codeman maintainer 9f5010aa51 fix(pr-bot): announce a bot-made merge once, cap automatic retries, show why a review failed
Observed on the first live merge (#383 via the Telegram button): runConfirmed
announced the merge and the scan five seconds later announced it again as a
closed PR. The scan now stays quiet for PRs the bot itself merged or closed,
and a merge of a `merge-with-fixes` verdict reminds that merging applies none
of the listed fixes.

A failed review used to be re-queued on every scan with no limit (two PRs
failed once each and were retried fine, but a head that keeps failing would
cost a session every ten minutes forever): three failures on one head now
stop the automatic retries until /review N or a new push. The failure notice
carries the reviewer's last message, so "finished without writing
report.json" says what it wrote instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 21:09:47 +02:00
Codeman maintainer f33b37c008 feat(pr-bot): review open PRs in Codeman sessions and report over Telegram
Maintainer tooling in scripts/pr-bot/: a daemon (systemd user unit
codeman-pr-bot) that lists open PRs with gh, reviews each head commit once in
a Codeman claude session (`prbot-<n>`) running in a private `git clone
--shared`, and sends the verdict, ranked findings, checks and a recommendation
to Telegram with action buttons. Merge, close, post-comment and approve-CI
happen only from a Telegram command or button plus a confirmation tap; the
bot never writes to GitHub on its own. The Telegram token and chat id come
from the existing notifier bot's env file.

Verified live: three PRs reviewed end to end (383, 363, 368), reports
delivered with buttons, reviewer sessions on the pinned model. Findings
along the way, each fixed and documented: a linked worktree inherits the
main checkout's model pin (hence the shared clone), undici's 5-minute header
timeout cut off the first review, gh was missing from the service PATH, and
the periodic scan orphaned an in-flight review's record.

typecheck/lint/format now cover scripts/pr-bot; tests in
test/pr-bot-{report,state,commands}.test.ts; guide in docs/pr-bot.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 21:05:42 +02:00
Ark0N 097d585278 Merge pull request #383 from opticon454/fix/case-picker-default-root
fix(file-picker): default the case picker to Codeman Cases, not Home
2026-09-06 21:03:00 +02:00
Michael Grundberg 8285fff91c feat(codex): list codex conversations and resume them
Codex conversations never appeared in the session list, and the resume path
skipped codex, so picking one back up meant finding its thread id by hand and
POSTing codexConfig.resumeSessionId to /api/sessions.

Two gaps caused it:

- The unified list is built from ~/.claude/projects plus omp's own store.
  Codex writes to neither: its rollouts live in ~/.codex/sessions/<y>/<m>/<d>.
- terminal-ui.js sends a continuation only for the CLIs with a
  "continue most recent" flag. Codex has no such flag — it names a thread by an
  exact id — and nothing supplied one.

Add codex-transcript.ts, the codex analog of omp-transcript.ts, and wire it into
gatherUnifiedInputs() beside the omp scan. A rollout row carries `resumeId`, the
thread id `codex resume` takes, and the resume path sends it as
codexConfig.resumeSessionId.

`resumeId` is what keeps the two kinds of row apart: only a transcript scanner
sets it, so a LIVE codex row — whose sessionId is Codeman's own uuid — can never
ask codex for a thread that does not exist.

Three things measured against a real store of 519 rollouts rather than assumed:

- Rollouts are far too large to read whole (median 407 KiB, p90 1.3 MiB, max
  25 MiB, 381 MiB total), so this reads a 128 KiB head for the identity and the
  opening prompt and a bounded tail for the most recent one. session_meta is
  written once and never rewritten, so per-path identity is cached; a warm
  rescan of that store costs ~75ms against ~470ms cold.
- codex 0.152.1 emits no event_msg/user_message rows at all. It writes
  event_msg/item_completed carrying an item.type of UserMessage. Both shapes are
  read, plus response_item as a last resort.
- That last resort sees injected context, and the first such row is the repo's
  AGENTS.md every time, so injections are dropped rather than used as titles.

Sub-agent threads (thread_source: 'subagent') are left out; codex spawns them
for itself and on a real store they outnumber the resumable threads.
2026-09-06 19:45:36 +02:00
Michael Grundberg 51957e2ed4 fix(session): let each CLI declare how its own pane shows work
A Codex session reported `isWorking: false` for its entire life, including
mid-turn. Codeman has four paths that mark a session working, and all four were
inert for Codex:

- The spinner fast path tests eight braille frames, and Codex animates none.
- The activity-streak fallback was wrapped in `!isExternalCliMode(mode)`.
- The pane probe inside `_confirmIdle` would have matched, since Codex prints
  `esc to interrupt`, but arming it required the literal glyph `❯` and Codex
  draws `›` on its composer row.
- The text detector sat inside `_processExpensiveParsers`, whose first statement
  returns early for an external CLI.

Add an optional `workDetect: { promptGlyph, workingLine }` to CliCapabilities,
so the two strings that differ per CLI are registry data rather than constants
in the detector. Claude declares its existing pair and behaves as before. Codex
declares `›` and `esc to interrupt`. The text detector moves above the
external-CLI early return, guarded on the descriptor so a CLI without one still
skips the ANSI strip that the early return used to save it.

A CLI that declares no descriptor falls back to Claude's pair, and the
activity-streak gate now reads "has a descriptor, or is not external", so the
plain shell mode keeps the behaviour it had.

Rewrite the test that asserted the old premise in its own comment, so it makes
the same guarantee for a genuinely uncharacterised CLI, and add Codex coverage
built from verbatim pane captures on Codex CLI 0.152.1.
2026-09-06 17:41:01 +02:00
Codeman maintainer 8ad2215118 fix(docker): close the three adoption gaps the negative guarantee missed
Review follow-ups to #357. Each is a path that still touched, or still hid, a
container Codeman does not own.

**Export still mutated it.** The four fail-closed layers cover create/start/
stop/remove, but `POST /api/docker-cases/:name/export` reaches the container
twice through neither: a full export `docker commit`s it, and even a
workspace-only export `docker pause`s it first for snapshot consistency. Pause
freezes the owner's processes for as long as the tar takes, on a container we
promised not to touch. Full export is refused for an adopted case (it packages
someone else's container, with their logins, into a bundle Codeman hands out);
workspace-only keeps working and no longer pauses, accepting a live filesystem
the way `tar` does on any running host directory.

**A freshly linked OWNED case became unusable.** The run menu now probes the
container for its CLIs, and a failed probe hides every agent mode behind the
reason. For an adopted case that is right. For an owned one the container does
not exist until the first session launches it, so every newly linked Docker case
answered `container "codeman-case-x" not found (adoption never creates a
container — start it yourself first)` and offered nothing but Shell, for a
container the launch chain was about to create itself. A failed probe is
recorded only when the case is adopted; `CaseInfo.docker.owned` is on the wire
so the frontend can tell them apart. Verified in a browser: owned-with-no-
container offers all ten modes and no notice, adopted-but-stopped offers Shell
and says why.

**Multi-user gating.** Adoption is admin-only, unlike `docker-link` beside it.
Linking creates OUR container, whose sole bind mount `isWorkingDirAllowed` has
already confined to the caller's space; an adopted container's mounts are
whatever its owner gave it, so one mounting `/` hands the adopter a shell over
the whole host — exactly the workspace scoping multi-user mode exists to
enforce. Listing the engine's containers and browsing directories inside an
arbitrary one are machine-level reads and follow the docker-HOST policy for the
same reason. The preflight is deliberately not admin-only: the run menu fires it
for every docker case, so it admits a non-admin for a container already linked
to a case they can access, and nothing else.

Verified end to end against a real pre-existing root container (alpine + tmux,
no bind mounts): adopt, claude session inside it, workspace export, session
close and case unlink all left `StartedAt`, `RestartCount`, `Pid` and `Paused`
untouched; the pane ran the CONTAINER's claude, without
`--dangerously-skip-permissions`; a stopped container was refused at both
preflight and launch and was never started.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1
2026-09-05 16:22:54 +02:00
Codeman maintainer 3d8ffcb9a2 Merge pull request #357 from dignfei/feat/docker-adopt-existing-container
feat(docker): attach a case to an already-running container

Conflicts came from work that landed after the PR was opened, and each is
resolved onto the newer abstraction rather than by keeping the older code:

- `defaultDockerCommandForMode` is registry-driven since #347, so the PR's
  `runsAsRoot` arm became `overlays.docker.rootCommand` (claude only). Claude
  Code still refuses `--dangerously-skip-permissions` as root in 2.1.261 and the
  refusal is visible only inside the container, so an adopted root container
  otherwise just shows a dead pane. Which flag to drop is a per-CLI fact, and
  `test/cli-registry-no-id-branching.test.ts` forbids expressing it as a branch.

- The probe's mode list and its mode -> binary table both duplicated the
  registry. They now read `enabledCliIds()` / `discovery.binaries[0]`, which is
  also what fixes the merge's silent regression: the hand-written list predates
  `omp`, and the run menu gates every docker case on this probe, so owned
  containers would have lost that mode. `shell` needs no arm — it declares no
  binary, so it is dropped from the lookup and reported available regardless.

- The per-mode `mode === 'claude' && !cliDir` chain in `tmux-manager.ts` is one
  `missingCliMessage(mode)` gate since #347; the PR's docker exemption moved onto
  it. Its test now pins the single gate instead of counting seven arms.

- The create arm keeps #349's swap-limit warning filter, which the adopted arm
  never reaches; the run-mode list gains `omp` from #353.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1
2026-09-05 16:21:51 +02:00
DevvynandClaude Sonnet 5 06febfa032 fix(file-picker): default the case picker to Codeman Cases, not Home
The "Link Existing" case picker opens with an empty path and no
sessionId, so the browse endpoint's fallback root picked whichever
root happened to be first in the list — which was always `Home`.

On the native default that's harmless (~/codeman-cases nests inside
Home anyway), but a Docker deployment binds CODEMAN_APPDATA_PATH
(Home) and CODEMAN_CASES_PATH at unrelated host paths, so the picker
opened somewhere with no cases in sight. Worse: if CODEMAN_CASES_PATH
is ever changed after cases already exist, the old cases directory
lingers, still reachable, under Home — indistinguishable at a glance
from the real one under the new Codeman Cases root.

Prefer the Codeman Cases root in the fallback chain, ahead of the
generic roots[0].

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
2026-09-05 17:14:05 +08:00
Michael TillerandClaude Opus 4.8 7e4914d991 feat(web): support a reverse-proxy base URL (--base-url / CODEMAN_BASE_URL)
Codeman can now be mounted under a sub-path behind a reverse proxy that
forwards the prefix unchanged (e.g. https://host/codeman/). Default is `/`
(root), which is byte-identical to the historical behavior.

Design — few choke points, mirrored ingress/egress:
- src/config/base-path.ts: pure single-source normalize/validate/join/strip.
- Server ingress: stripBasePath() inside Fastify rewriteUrl, so routes stay
  declared prefix-agnostic; un-prefixed requests (hooks, health, docker bridge
  hitting the raw port) pass through unchanged.
- Server egress: one onSend hook prepends the base to root-absolute Location
  headers (covers all redirects).
- HTML: renderIndexHtml points <base href> at the mount and injects
  window.__CODEMAN_BASE__ — ONLY when a base is set (inert at root).
- Frontend runtime URLs: CodemanBase.url() route builder in constants.js,
  applied transparently by a fetch wrapper and explicitly at the
  EventSource/WebSocket/window.open/<img|iframe|a>-src sites.
- sw.js derives its base from self.location; manifest uses relative start_url/scope.
- Web-tab proxy: proxyPrefixFor(cap, basePath) is the single base-aware root that
  cascades to the injected <base>, HTML/attr rewrites, runtimeUrlShim, Set-Cookie
  Path and Location; capabilityFromReferer strips the base off the browser Referer,
  while the ingress parsers stay base-agnostic (rewriteUrl already stripped it).

--base-url rides the daemon relaunch (buildWebArgs) and the service unit
(resolveServicePlan). constants.js is guarded against a missing `window` for
isolated unit-test contexts.

Tests: test/base-path.test.ts (pure helpers), base-path coverage in
webview-proxy/render-index-html/daemon-control; CodemanBase stubbed in the
vm-isolated panels-ui test contexts. Docs: Remote-Access.md (sub-path section +
nginx example), security-architecture.md env table, CLAUDE.md pattern.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XUkPBxbumnct6qSrx4JDju
2026-09-04 16:08:57 -04:00
Codeman maintainer 6f7add7ce4 chore: version packages 2026-09-04 20:46:19 +02:00
Codeman maintainer eeb5f9d0b2 docs: web-tab egress guard, capability revocation and referrer policy
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:15 +02:00
Codeman maintainer 2ab21c1b32 fix(webview): revoke proxy capabilities on logout and stamp Referrer-Policy
WebviewCapabilityStore.revokeOwner() shipped for two releases with a docstring
claiming logout called it and no caller at all. The capability is a bearer
credential exempt from cookie auth with a rolling TTL refreshed on every use, so
a proxy URL that leaked (browser history, a screenshot, a dashboard with a loose
referrer policy) stayed valid for as long as anything kept polling it.

- POST /api/logout revokes the caller's capabilities (all of them in single-user
  mode), the admin forced logout revokes the target user's, and user deletion
  revokes whatever that user had open. revokeOwner returns the count for the
  admin audit line.
- Proxied responses carry `Referrer-Policy: same-origin` and the upstream's own
  policy is dropped: every URL inside the frame carries the capability, and a
  dashboard on no-referrer-when-downgrade or unsafe-url handed it to any
  third-party host it linked. Verified with Playwright that a sandboxed frame
  under an upstream `unsafe-url` sends no Referer to a third party while the
  root-absolute fetch and the CSS-triggered 404 fallback still reach the
  dashboard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:13 +02:00
Codeman maintainer 550e08a791 fix(webview): refuse link-local and cloud-metadata targets on the resolved address
The web-tab proxy, its Test probe and its WebSocket relay accepted any http(s)
host. A live PoC relayed an IMDSv2-shaped PUT with custom headers to a loopback
echo server through a capability and no cookie, and 169.254.169.254 (decimal,
hex, IPv6-mapped, or via a DNS name) was as valid a dashboard as any other.

Loopback and RFC1918 stay allowed on purpose: a localhost Grafana is the feature.
Only link-local and the fixed cloud-metadata addresses are refused
(169.254.0.0/16, fe80::/10, fd00:ec2::254, 168.63.129.16, 100.100.100.200,
metadata.google.internal), at three stages that are each load-bearing:

- the Zod schema, so a save gets a clear refusal;
- a synchronous hostname check at every connect site, because net.connect skips
  DNS for an IP literal and a lookup hook never sees one;
- a `lookup` hook on an undici Agent (webviewFetch) and on the ws client, which
  judges the RESOLVED addresses of a name and refuses when any is blocked. This
  is what closes DNS rebinding, which a hostname-string check cannot.

Adds undici@^6 so the proxy runs the package's own fetch with the package's own
Agent; a package Agent handed to Node's bundled fetch can mismatch protocols.

Verified live on an isolated beta: 169.254.169.254.nip.io (a real name resolving
to the metadata address) is refused by probe, proxy (403) and WS relay (4003),
while 127.0.0.1.nip.io still passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:12 +02:00
Codeman maintainer 99ad9cb236 fix(docker): never exit the server unless something is known to restart it
#373 restarts the Compose container by exiting the server, which is right for
the shipped deployment: `restart: unless-stopped` relaunches it. The updater
verified that policy through the Docker socket and, when it could not (no
socket mounted), failed open and exited anyway. Failing open is the correct
choice for the GATE, where refusing would block every install without a
socket, but not for the kill: a container the daemon does not restart goes
down for good, with no UI left to recover it from. That is exactly the case a
plain `docker run` of this image without `--restart` produces, and the image
sets CODEMAN_IN_CONTAINER=1 itself, so it takes the container path.

The decision now happens server-side, where both the socket and the Compose
env are reachable, and rides down to the script as `--restart-by-exit 0|1`.
It is 1 when the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (added there
and only there, since that file is what sets the restart policy; the image ENV
deliberately does not) or when the daemon confirmed an auto-restart policy.
Otherwise the build still lands, the status becomes
`completed-needs-manual-restart` with the `docker restart` hint, and the
server keeps running. The shipped deployment is unchanged in effect: with the
socket it was already confirmed, and without it the declaration now covers it.

Also: a root-run `Start-Codeman.sh` (common on Unraid) created the
fingerprint baseline's `.codeman` directory before the container's first start
and left it root-owned, which the unprivileged server could then never write
its own state into. It is chowned to PUID:PGID when running as root.

Verified with a real image build of the merged tree (classic builder; this
box's BuildKit lacks buildx): runs as uid 1000, tsc/esbuild and the toolchain
present, the four CLIs at their pins, docker/.env absent, and `docker inspect
$HOSTNAME` returns the restart policy through the mounted socket as that user.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 14:36:35 +02:00
Codeman maintainer 823f56a243 Merge pull request #373 from opticon454/feature/docker-self-update
feat(docker): restore in-app self-update in the Compose dep
2026-09-04 14:25:56 +02:00
Codeman maintainer 72fd231d11 test(setup): one answer for CODEMAN_DATA_DIR, the strip from #371
#356 and #371 fixed the same leak two ways. #356 pointed CODEMAN_DATA_DIR at a
second throwaway directory and cleaned it up in afterAll and on exit; #371
deletes the variable along with CODEMAN_INSTANCE and CODEMAN_TMUX_SOCKET, so
`getDataDir()` falls back to `homedir()`, which the temp HOME already redirects.
Merged as they were, setup.ts set the variable and deleted it a few lines
later, and the second directory was created for nothing.

The strip wins: same protection, one tree to clean up, and the isolation test
#371 adds pins the list statically. The extra directory, its restore and its
two rmSync calls go, the vitest config `env` entries that set the same variable
go (they were documented as inert and would now be contradicted by the setup
file either way), the two test comments that described the old mechanism are
reworded, and CLAUDE.md's testing paragraph names the three stripped variables
and why CODEMAN_INSTANCE has to be stripped in the setup file rather than a hook.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 14:20:11 +02:00
Codeman maintainer 65d19c725e Merge pull request #371 from opticon454/fix/test-env-instance-isolation
fix(test): strip the instance-selection env vars in test/setup.ts
2026-09-04 14:20:04 +02:00
Codeman maintainer 80626567b2 chore: version packages 2026-09-04 14:01:20 +02:00
Codeman maintainer a81e87f440 fix(cli-registry): log why clis.json was ignored, and say 0600 when that is the rule
The loader refuses a `clis.json` with any group/world permission bit, read bits
included, so a file created with a normal umask (0644) is ignored. That is a
defensible posture for a file that chooses the binaries Codeman spawns, but two
things around it made the override feature look dead: the warning said
"group/world-writable", which a 0644 file is not, and `LoadResult.warnings` was
returned to a caller nobody wired up, so nothing anywhere printed it. A user
following the docs got silence.

The message now names the rule and the command that satisfies it, the loader
logs every warning once on first load (the result is memoized, so once per
process), the module header stops claiming that nothing ever writes (the
quarantine rename of a malformed file is a write, on first use) and the
registry doc gains a short section on the override file with the 0600
requirement in it. Whether the check should relax to writable bits only is a
separate decision; this keeps the shipped behaviour and makes it visible.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:51:20 +02:00
Codeman maintainer 2e0129f1f8 docs(test): name the real reason the suite could reach ~/.codeman
#356 stopped a bare suite run from overwriting the production
`remote-hosts.json` by pointing `CODEMAN_DATA_DIR` at a throwaway dir, and it
gated every case-tree delete on the temp HOME. Both changes are right; the
explanation written next to them is not. It says `os.homedir()` reads
/etc/passwd rather than `$HOME` on Linux, which would mean the temp HOME in
test/setup.ts never worked. It does: libuv checks the env var before the passwd
entry (measured: `HOME=/tmp/x node -e 'console.log(os.homedir())'` prints
/tmp/x), and CLAUDE.md's testing section relies on exactly that.

What bypasses the temp HOME is `CODEMAN_DATA_DIR` itself. `getDataDir()` reads
it as an absolute override before it looks at `homedir()`, so one inherited from
the shell (a second instance, a beta run) sends the whole suite at the real data
dir. That is the case setup.ts now closes, and #371 names the same variable from
the other direction.

The comments in setup.ts, the `safeRmHomeTree` helper, the voice-routes and
case-clone tests now say that, and the containment gate is described as what it
is: defense in depth. CLAUDE.md's testing paragraph gets the same note so the
next reader does not chase a homedir() bug that does not exist.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:50:22 +02:00
Codeman maintainer 28b44237ae fix(remote): classify the has-session probe by exit status, and forget it once the pane is back
#355 made the remote auto-reconnect watcher revive a dead pane only when the
durable remote tmux session is verifiably still alive, which is the right rule:
a clean Ctrl-C / Ctrl-D / exit tears that session down and must never relaunch
a fresh agent. Its probe, though, read `has-session`'s stdout and treated an
empty string as "gone". `tmux has-session` prints NOTHING on success (measured
on a scratch socket: exit 0, empty stdout, the failure message goes to stderr),
so every live remote session classified as gone and transport-drop reconnects
were silently disabled along with the clean-exit revives.

The probe now goes by exit status through a pure, unit-tested mapping
(`classifyRemoteAliveExit`): 0 is alive; ssh's own 255, a timeout (`killed`,
no numeric code) and a spawn failure are unknown, which the watcher already
treats as do-not-revive; any other status is the remote command's and means
gone (tmux's 1 for a missing session, 127 when tmux is not installed there).

Two smaller things in the same area:

- The cached answer was never invalidated, so after one successful reattach a
  stale `true` would have revived the NEXT clean exit (the original bug back
  after the first transport drop), and a cached `false` from a clean exit would
  have left a manually restarted session with auto-reconnect permanently off.
  The tick now forgets the cache entry whenever the pane is seen alive.
- The fire-and-forget probe has a 15s timeout against a 5s tick, so an
  unreachable host stacked up to three ssh processes per dead session. An
  in-flight set caps it at one.

The probe command is pinned as a literal string, and the reattach-then-clean-exit
sequence is driven through the watcher in the tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:50:22 +02:00
Ark0N 96960785d2 Merge pull request #356 from timkjr/pr/test-isolation
fix(test): isolate route tests from the production ~/.codeman data dir
2026-09-04 13:50:03 +02:00
Ark0N ee6a7af1d1 Merge pull request #355 from timkjr/pr/remote-exit
fix(remote): never auto-revive a remote session after a clean agent exit
2026-09-04 13:49:49 +02:00
Ark0N 850b00572c Merge pull request #347 from opticon454/feature/cli-registry-core
PR A: CLI registry core as a pure internal refactor
2026-09-04 13:49:35 +02:00
timkjr 4a63ab1604 test: extend CASES_DIR containment guard to the rest of the suite
#356 introduced safeRmHomeTree/isUnderTestHome to stop tests from deleting
the PRODUCTION ~/codeman-cases tree on platforms where os.homedir() ignores
the $HOME override -- but only applied it to the one file caught doing it
live. CASES_DIR has no CODEMAN_DATA_DIR-style env override at all, so every
other test file's raw rmSync(join(CASES_DIR, ...)) was the same unguarded
pattern, just not yet triggered.

Routes every CASES_DIR delete in these 10 files through safeRmHomeTree:
cli-skill-target, edge-cases, integration-flows, operation-lightspeed,
ralph-integration, routes/case-clone-routes, routes/voice-routes,
session-cleanup, sse-events, sse-subscription-filter.

Also fixes one instance in case-clone-routes.test.ts that mkdirSync'd then
rmSync'd a CASES_DIR path directly with no guard at all -- the exact
clobbering pattern #356 exists to prevent, found by extending the sweep.

Held as a separate commit (and intended as a separate PR once #356 merges)
rather than folding into #356 -- keeps the already-checked skinny fix
reviewable on its own; this is the same bug class applied broadly, not new
functionality.

Verified: all 10 files pass (180 tests), npm run typecheck clean.
2026-09-02 20:41:24 -05:00
timkjr 4068c02b9e fix(test): write the remote-hosts fixture where the route actually reads it
The "never writes hooks for a remote attach" test stubbed CODEMAN_DATA_DIR
to a separate throwaway dir just for this write, but session-routes.ts's
CODEMAN_CONFIG_DIR is a module-load-time constant frozen at test/setup.ts's
sandboxed dir before this test ever runs. The fixture landed somewhere the
route handler could never read, so the remote-host lookup silently failed
(NOT_FOUND) and the test passed for the wrong reason -- createErrorResponse
never sets reply.code(), so Fastify's default 200 made the NOT_FOUND branch
and the intended success branch indistinguishable by status code alone.

Write straight to getDataDir() instead, matching the docker-hosts fixture
convention already used elsewhere in this file. Verified the fix actually
exercises the success path (host resolves, 200 with a real session), not
just an accidental 200 from the error branch.
2026-09-02 20:41:24 -05:00
timkjr ff88b6957e fix(test): guard the CASES_DIR delete + harden the data-dir teardown
PR #356 stopped the remote-hosts.json fixture write from clobbering prod.
Two holes in the same file remain:

1. The quick-start afterEach still ran rmSync(CASES_DIR, recursive).
   CASES_DIR is join(homedir(), 'codeman-cases'), and on Linux builds
   where os.homedir() reads /etc/passwd instead of $HOME it resolves to
   the PROD case tree - so a full-suite run deleted the real
   ~/codeman-cases. Add a shared safeRmHomeTree() containment gate that
   only deletes a path under the redirected test HOME.

2. setup.ts teardown did rmSync(process.env.CODEMAN_DATA_DIR ?? '') AFTER
   restoring the env - if a pre-existing prod CODEMAN_DATA_DIR was set,
   that deleted prod. Capture the throwaway dir in a const and clean that.

A broader test-isolation sweep (10 files: cli-skill-target, edge-cases,
integration-flows, operation-lightspeed, ralph-integration,
case-clone-routes, voice-routes, session-cleanup, sse-events,
sse-subscription-filter) also applies the same containment gates to every
per-case delete. It is intentionally NOT included here to keep this PR
skinny; it is identified and available on request.
2026-09-02 20:41:24 -05:00
timkjr 2694d3f74a fix(test): isolate route tests from the production ~/.codeman data dir
session-routes-workspace-hooks.test.ts wrote its h1/box/10.0.0.5 host
fixture into getDataDir()/remote-hosts.json. getDataDir() resolves via
homedir() → ~/.codeman (INSTANCE_SUFFIX='' by default), and overriding
HOME in test/setup.ts does NOT change os.homedir() on Linux — so every
full-suite run silently overwrote the PRODUCTION remote-hosts.json,
wiping user-defined remote hosts, emptying the launch-case dropdown and
breaking remote session creation (found live 2026-08-29).

The vitest v4 test.env config key is ignored (probe confirmed the
worker still saw CODEMAN_DATA_DIR=undefined), so the reliable fix is
stubbing the env inside the test: the fixture write now goes to a
throwaway /tmp dir via vi.stubEnv + finally unstub. Verified: prod
remote-hosts.json hash is identical before and after the suite run.
2026-09-02 20:40:58 -05:00
DevvynandClaude Opus 5 66eb01ba8f feat(docker): restore in-app self-update in the Compose deployment
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.

Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:

- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
  update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
  container on the new dist/. This is the one supervisor whose updater does NOT
  outlive the restart, which is safe only because the terminal "restarting"
  marker is written first.
- node_modules and dist are named volumes over the bind mount, so
  container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
  `npm run build` is tsc + esbuild and node-pty has no Linux prebuild.

An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.

The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.

Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.

Documented in docs/docker-self-update.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
2026-09-02 19:33:32 +08:00
Codeman maintainer 1e24817b51 chore: version packages 2026-09-02 10:49:36 +02:00
Ark0N f7cf15485e feat(models): offer Fable 5.1 in the model picker and task routing (#372)
Adds `claude-fable-5-1` to the App Settings model picker and the five task-routing selects, mirroring how Fable 5 is already offered: a base option with data-ctx="1" plus its [1m] companion row. No settings-ui.js logic change, since the cards and the 1M switch are built from those options.
2026-09-02 10:48:51 +02:00
DevvynandClaude Opus 5 1125f7c1c5 fix(test): strip the instance-selection env vars in test/setup.ts
`test/setup.ts` gives every test file a temp HOME so the suite cannot touch the
real Codeman tree, and strips the env vars that would leak past it — but the
list only covered auth and the gesture flag. The three vars
`src/config/instance.ts` derives the data dir and tmux socket from were missing,
and they reach past the temp HOME:

- **`CODEMAN_DATA_DIR` is the one that matters.** It is an ABSOLUTE override
  read in `getDataDir()`, so it bypasses HOME entirely: a developer who exports
  it — or a shell left over from `codeman web -d` — has the suite reading and
  WRITING their real `state.json`, `users.json`, `intents.json` and
  `hook-secret`.
- **`CODEMAN_INSTANCE`** moves the data dir to `~/.codeman-<name>` and the
  socket to `codeman-<name>`. Inside the temp HOME that is not data loss, but it
  silently changes the paths tests assert on — and `scripts/run-beta.sh` exports
  it, so any shell that has run a beta carries it.
- **`CODEMAN_TMUX_SOCKET`** renames the socket `resolveTmuxSocketName()`
  returns. `TmuxManager` no-ops its shell commands under vitest, so this is
  assertion drift rather than a stray `tmux -L` against prod — same class of
  leak, same one-line fix.

They are deleted in the setup file rather than in a hook because
`CODEMAN_INSTANCE` is captured into a module-level const the first time
`config/instance.ts` is imported; a `beforeEach` would already be too late.

`test/test-env-isolation.test.ts` pins the whole list in two halves, because the
obvious half is not enough: asserting the vars are unset passes trivially on a
machine that never set them, so a removed `delete` line would sail through on
almost every box and on CI. The static half reads `setup.ts` and asserts each
name is deleted there, which fails everywhere. An anti-drift check catches the
other direction — a var stripped in `setup.ts` but never given a reason in the
list — and is scoped to the strip section so the teardown's restores are not
mistaken for strips.

Verified by demonstrating the leak: with the `CODEMAN_DATA_DIR` line removed and
the var exported, the runtime assertion fails; with the line restored it passes.
Full suite: no new failures against an upstream/master baseline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 09:49:09 +08:00
DevvynandClaude Opus 5 c5b84fb5f4 docs(cli-registry): annotate overlays.credStore as declared-for-later
Review item 4 named THREE live tables duplicating registry data. Two are now
read from the entry (`defaultRemoteCommandForMode`, `defaultDockerCommandForMode`);
the third, `resolveDockerCredentialArtifacts`, is not — and it was left neither
wired nor annotated, which is the state that item explicitly rules out.

It is not wired because the shape cannot express the live table: `credStore` is
ONE store per CLI, and `CRED_STORES` needs two for gemini (`.gemini` for the
CLI's own auth plus `.config/gcloud` for Vertex), while deepseek's entry declares
none at all even though `.dsh` is seeded. Wiring it means making the field an
array and correcting those two entries — a change to credential seeding, which
is at once the worst thing in that file to get wrong and the least covered by
tests, since every docker IO path is no-op'd under vitest. It belongs in its own
change, measured against a real container.

So it is annotated instead, at the field, in the type's declared-for-later
header, in docs/cli-registry.md, and in the pinned DECLARED_FOR_LATER list — the
last of which means wiring it later makes a test fail rather than leaving a
stale comment behind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 09:35:55 +08:00
DevvynandClaude Opus 5 6acf0dea0f fix(cron): scope the launch pre-flight to launcher CLIs, not every mode
CI caught three cron-service failures. Both are mine, from converting cron's
per-mode ladders to capability reads without checking what each ladder's scope
actually was.

**The pre-flight.** cron only ever pre-flighted `deepseek` — dsh is a profile
LAUNCHER, so "installed" is not "runnable" and a bare `dsh` can boot a profile
that cannot drive a pane. I replaced that with an unscoped
`resolveCliLaunchError(mode)`, which pre-flights EVERY mode, so a claude cron
job on a box with no claude binary now failed with "Claude CLI not found"
instead of reaching tmux-manager's own throw. Three tests assert the latter.
It is now gated on `discovery.launcherProfile !== undefined`, which is
byte-identical to the `mode === 'deepseek'` check it replaces and generalises to
the next launcher. The equivalent HTTP-route conversion was already scoped (to
`capabilities.external`, matching what that route has always pre-flighted); I
simply failed to carry the same reasoning across.

**The model.** cron's ladder was `mode !== 'shell' && mode !== 'deepseek'`, and
I read it as `capabilities.model.source === 'claude-settings-file'` — which is
the HTTP route's question, not cron's. There, every external CLI reads its model
from its own config object earlier in the chain, so only claude reaches the
global default; cron has no such config, so the same expression silently
narrowed the default model from eight modes to one. Now `!== 'none'`, which is
exactly the two entries the ladder excluded. Not caught by a test — found by
re-deriving each ladder's scope after the first failure.

Also names a fourth deliberate behaviour change in the changeset, found while
tracing these: `session.ts` carried a hand-written list of modes with no
direct-PTY fallback and OMP was missing from it, though CLAUDE.md's own text
says "all eight require tmux". `requiresMux` comes off the entry now, so an omp
session whose mux creation fails refuses instead of silently starting outside
tmux.

Verified by diffing failing tests BY NAME against an upstream/master baseline,
rather than by file as before — which is how the regression slipped through: the
three new failures landed inside a file already failing for unrelated
Windows-path reasons, and the aggregate count happened to collide.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 08:49:04 +08:00
DevvynandClaude Opus 5 4830e662f9 refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery
(search dirs, version + identity probes), the launch argv template, env
handling, the `capabilities` flags that replace per-CLI branching, and the
`overlays` that back the remote/docker pane commands. Code that used to ask
"which CLI is this?" reads the entry instead.

Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every
spawn command as a literal string, captured from the hand-written builders
before they were deleted, and `test/location-overlay-commands.test.ts` does the
same for all 20 remote and in-container pane commands.

Config can never contain shell text: an entry declares typed argv tokens,
literals are validated against a safe-word pattern at LOAD time (a bad literal
rejects the whole entry — a silently dropped `--no-approve` is not cosmetic),
and values resolve through patterns NAMED in code, so a user `clis.json` cannot
widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only
in this release.

OMP is included as a registry entry rather than a tenth hand-written builder,
so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of
`buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen
and doctor ladders all drop out.

Guard rails:

- `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id
  branching reappears outside `stock.ts`, in any of its four shapes (`===`,
  `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the
  negated forms, which is how 36 of them survived an earlier pass. Every
  allowlisted branch carries its reason.
- `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities;
  deriving one from another shipped the `until=stop`-hangs-on-shell bug.
- `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and
  `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config`
  wire field is separate, bridged only by `legacyConfigAliases`. Getting
  `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass
  clamp's only handle on a CLI's privilege switch, and a wrong name clamps
  nothing with no error and no failing test — so `schema.ts` rejects an entry
  naming a param it never declared.
- Registry data resolves AT CALL TIME (`sessionModeSchema()`,
  `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs`
  thunks). A module-level const freezes at first import, so a CLI enabled while
  the server ran moved the run menu but not that surface.
- Six fields are annotated DECLARED-FOR-LATER and read by nothing
  (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/
  `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed
  rather than measured. A test pins the list so it cannot quietly grow.

Three user-visible changes, all deliberate and named:

- `probeDockerCliVersion()` derives the in-container binary from the registry
  rather than assuming it equals the mode name (`antigravity` runs `agy`).
- The remote CLI version probe now covers grok and deepseek, which the
  hardcoded map it replaces omitted while its own comment said the rule was
  "every mode except shell".
- `codeman doctor`'s CLI rows are generated from the entries, so Claude's
  install hint is the install command rather than a docs URL, five CLIs gain
  hints they never had, and the row order follows the catalog.

Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars
(matching the `cliId` pattern) before its failure message quotes the value
back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading
the hand-editable `clis.json`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 08:26:45 +08:00
Codeman maintainer 71ffbf18e4 chore: version packages 2026-09-01 21:55:00 +02:00
Codeman maintainer 826ddaa9aa build(docker): ship the Docker CLI in the Compose image, not the whole engine
docker/server.Dockerfile installed Debian's `docker.io` to get a client for the
socket mounted by Compose. That package is the full ENGINE: even with
--no-install-recommends it pulls 15 packages including containerd, runc, dmsetup
and iptables, none of which a container that only talks to a mounted socket can
use, and it ships Docker 20.10.24 (2023).

Copy the CLI and the buildx plugin from the official docker:29-cli image
instead. Measured on the same node:22-bookworm-slim base: 266 MB -> 108 MB, so
158 MB smaller with a current CLI (29.7.2) in place of a two-year-old one.

Three things verified rather than assumed, by building the real image and
running it:

- docker:cli is an ALPINE image, so copying a binary into this Debian one is
  only safe because the binaries are static Go builds (ldd: "Not a valid dynamic
  program"). In the built image, `docker --version`, `docker ps` and
  `docker build` all work against a mounted host socket as the unprivileged
  runtime user.
- buildx is copied on purpose. scripts/build-agent-image.mjs shells out to
  `docker build` and Codeman auto-builds the agent image on the first Docker
  case. Without the plugin that still works today — CLI 29 falls back to the
  classic builder, tested — but that builder is deprecated and will be dropped,
  so the plugin keeps the path supported.
- docker-compose is NOT copied: Codeman never shells out to it.

Pinned to the 29 major, matching how the base images here are pinned.
2026-09-01 21:54:56 +02:00
Codeman maintainer e2b72aafd7 chore: version packages 2026-09-01 11:32:50 +02:00
Codeman maintainer b15cc0eb1a fix(ui): keep the plan-usage chip's 5h slot when no session window is open
The header chip silently shrank from "5h 4% · 7d 52%" to a lone "7d 52%", which
reads as half the feature breaking rather than as an idle window.

Nothing was broken. Claude Code documents `rate_limits.five_hour` as "present
only while the API reports it and its resets_at has not passed", so between
5-hour session windows the key simply leaves the statusline payload. Codeman's
snapshot replaces the Claude half wholesale on every sample, so the segment
disappeared until usage opened a new window. Confirmed against a live 2.1.252
session by capturing real statusline payloads on an isolated tmux socket: the
boot render carries no `rate_limits` at all, and the first post-response render
carries both windows.

The slot now stays, with a dimmed em dash. Claude only: a missing CODEX bucket
means that plan has no such limit rather than an idle window, so those stay
omitted (pinned by the existing test). The placeholder can never stand alone
either — hasWindows() still gates the row, so a provider reporting nothing
renders nothing rather than a row of dashes. The tooltip says "5-hour limit: no
active session window" instead of dropping the line.

Verified in a real browser against a dev instance: the idle chip renders
"5h — · 7d 52%" with the dash at opacity 0.55 in --text-dim while the live
value keeps its green, and the chip holds its shape (100px idle vs 107px with
both windows).
2026-09-01 11:32:18 +02:00
Codeman maintainer 2a32b5064a Merge pull request #349 from opticon454/feature/docker-compose
Docker Compose deployment: Codeman runs in a container and spawns Docker cases
as SIBLING containers through the mounted host socket (Docker-outside-of-Docker).

Resolved the README conflict (master had grown to eight CLIs since the branch
was cut) and moved the Compose blurb out of the feature bullets into Quick
Start, next to the other ways of starting Codeman.

Three review findings from the PR discussion are fixed here rather than left
for a follow-up, because two of them are shipped-image problems:

- `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against
  the whole context-relative path, so `docker/.env` — which the deployment's own
  README tells the user to fill with CODEMAN_PASSWORD and provider API keys —
  was picked up by `COPY . .` and baked into the image at
  /opt/codeman/docker/.env. Verified in both directions against a real build
  context: with a canary secret in docker/.env, the unfixed ignore file lets
  /ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for
  the checked-in .env.example) leaves only the example behind.
- `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which
  still hardcoded ~/codeman-cases, so `codeman skill install --case <name>`
  reported "Case not found" on exactly the deployment the override exists for.
  Both now resolve through config/cases-dir.ts. state-store.ts keeps its own
  literal on purpose: that one migrates the historical ~/claudeman-cases
  directory by name and is about the old default, not the active location.
- CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the
  three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore
  joins the documented list of files that genuinely belong in the repo root.

The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug
fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/
deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's
lastClaudeSessionId was handed to every non-claude CLI.

Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax,
public assets, lockfile.
2026-09-01 11:32:03 +02:00
shenlvkang-collab ccfda623fe fix(session): learn the live Claude conversation from the CLI's own hook
Which conversation a pane is on was re-derived by correlating
~/.claude/history.jsonl against Session.lastSubmitAt — and lastSubmitAt is
bumped only by input that flows through Codeman's own write path
(Session.write / writeViaMux). A user who attaches to the pane's tmux session
directly never set it, so resolveActiveClaudeSessionIdFromHistory() returned at
its first line for that pane's whole life and the response viewer stayed pinned
to the launch conversation, showing a pre-/clear transcript indefinitely.

A UserPromptSubmit hook reports the live conversation id from inside the CLI
process, delivered under the pane's own $CODEMAN_SESSION_ID. That binding is a
fact rather than a correlation: it never consults workingDir, so it cannot be
claimed by a sibling pane on the same folder, a closed tab, or a bare `claude`
in the user's terminal. A pane holding such an id skips the correlation
entirely, so the number of prompts eligible for cwd-based guessing goes DOWN,
never up — the naive alternative (relax the guard, or synthesize an anchor from
PTY activity) is the reverted bug the resolver's own comment describes.

The hook also stamps lastSubmitAt, so it finally means "a prompt was submitted"
rather than "typed into Codeman's web terminal". Conversations vouched for
first-hand — and only those — extend a persisted claudeSessionChain, whose tail
re-pins the conversation when a surviving tmux session is re-attached after a
restart. ⚠️ start() resets the id at THREE points and the last one runs
unconditionally after the mux branch, so the tail is applied there too; patching
only the mux branch looks right and silently does nothing.

⚠️ The hook's stdout is discarded with curl's own -o /dev/null. Claude Code
injects a UserPromptSubmit hook's stdout into the model's context ("Exit code 0
- stdout shown to Claude"), and a trailing >/dev/null does NOT work: curlCmd
already ends `... 2>/dev/null || true`, and in `pipeline || true >/dev/null` the
shell binds the redirection to `true`, which never runs on the success path. The
discard is opt-in so the five SSE-fed events keep byte-identical command text
and no workspace's settings file is rewritten for them. The staleness marker is
quote-free for the matching reason: hooksJson is JSON.stringify'd, so a quoted
needle never matches and the gate would rewrite every workspace on every spawn.

Existing workspaces heal on their next Claude spawn through the staleness sweep.
2026-09-01 12:33:24 +08:00
shenlvkang-collab 3eff1feb5d fix(web): render one Claude response-viewer message per model message
The Claude reader concatenated every assistant row between two human prompts
into one card, fusing up to 74 distinct model messages into a single card, and
it never read the attachment rows that hold a prompt typed while the agent was
working. Measured over 57 real transcripts on 2026-09-01, the viewer shows
1,806 messages instead of 356 and 353 user cards instead of 178, with the
assistant text sequence unchanged row for row and the response without
?context=full byte-identical on all 57 files.

One assistant row IS one whole model message: in that corpus no assistant row
carries more than one content block and no message id carries more than one
text block, so there was nothing to reassemble. Each row becomes its own
message carrying an additive {kind, label, turn}, and the frontend renders a
same-role run inside one turn as badge-less continuation segments — which is
what keeps a p90 of 11 messages per turn from reading as card spam. A numeric
turn gates that rendering, so Codex, the external-CLI pane parser and an older
server keep one badge per card.

A prompt typed while Claude is working is recorded ONLY as an
attachment/queued_command row. Taking it when origin.kind is 'human' and
commandMode is 'prompt' recovers 162 user cards from 163 such rows — one is a
verbatim repeat inside an unanswered user run and is collapsed by the existing
dedup guard — and restores the turn boundary whose absence let the assistant
runs fuse. The CLI's own queue entries are cleanly separable: of 322
queued_command rows, 159 are commandMode 'task-notification' and not one of
them carries an origin key.

This narrows #169 rather than reverting it: sidechain exclusion, the
restored-<uuid8> rebind, replayed-snapshot dedup and synthetic-row filtering
are all unchanged and still asserted.
2026-09-01 12:30:48 +08:00
shenlvkang-collab 5969a1df96 fix(mobile): give the Add Case modal a reachable submit button
mobile.css hides #createCaseModal's .set-foot below 860px, and that modal's
header — unlike Settings' — carries no set-head-save. So on a phone the
Create/Link button existed nowhere and the modal could not be submitted at all.

Adds the header button and drives both together through switchCaseModalTab()
and submitCaseModal(), so whichever one is pressed the other shows the same
pending state and is equally unclickable. Following the Settings pattern also
means Add Case picks up the existing .set-head-actions:has(.set-head-save) tray
and .set-head-save sizing with no new CSS; the mobile.css comment that still
listed Add Case as a lone-× sheet is corrected to match.
2026-09-01 12:29:02 +08:00
Codeman maintainer 0da0c8219d chore: version packages
1.24.2. Also corrects two numbers in the CLAUDE.md trust-dialog paragraph that
was written while the fix was still uncommitted: the keystroke cap is 6, not 3,
and the scan now schedules its own follow-up read rather than waiting on PTY
output that a static dialog never produces.
2026-09-01 02:31:31 +02:00
Codeman maintainer aaa93d4252 fix(session): answer Claude Code 2.1.252's reversed folder-trust dialog
Every claude session in a directory claude had not seen before died about six
seconds after it started (`Pane is dead (status 1)`), before the agent drew a
composer. Reproduced on a fresh case and measured.

Claude Code 2.1.252 rewrote the dialog. It used to be

  ❯ 1. Yes, I trust this folder
    2. No, exit

and is now unnumbered, reversed, and highlights the option that quits:

  ❯ No, exit
    Yes, I trust this folder

Detection still worked (the confirm affordance carries the match once the
numbered option text is gone), so the failure was entirely in the answer: the
auto-accept pressed Enter on the highlighted default, which is now exit.

trustDialogNextKey() reads the ❯ marker off the rendered pane and returns ONE
keystroke at a time: an arrow while the cursor is on the wrong option, Enter
only once the screen shows it on the trust option, and null for a frame that
does not say. Both layouts are handled, and which way the trust option lies is
read from the frame rather than assumed, so a further reordering costs a
repaint instead of a session. The last marked option wins, because the
direct-PTY fallback reads an append-only buffer where an older frame must not
out-vote the freshest one.

Two things only a live pane showed:

- The scan ran solely from the PTY onData handler. The arrow that moves the
  cursor is the last output the pane produces, so the first fix parked every
  session with the cursor sitting on the right option and no Enter ever sent.
  It now schedules its own follow-up read (_trustDialogTimer, cleared in
  _clearAllTimers()), offset past the scan throttle so the chain cannot break
  on a boundary.
- The keystroke cap goes 3 -> 6, since answering is no longer one press.

The bundled codeman skill had the same blind \r as its bounded fallback, so
preamble 1.21.0 replaces it with _trust_key/_accept_trust: read
terminal?full=1, steer onto the trust option, re-read, then confirm. Those
keystrokes go out under their own clientId, because input sequence numbers are
monotonic per client and spending prompt numbers on dialog keys would make the
next send-and-wait look like a stale duplicate and vanish while reporting
success. The readiness recipes in docs/extending-codeman.md,
docs/api-reference.md and the skill's own reference carry the corrected answer,
plus a symptom-table entry for a worker whose pane is dead seconds after spawn.

Verified live on an isolated instance (own data dir and tmux socket): fresh
case -> arrow at 5 s -> Enter at 7 s -> composer, with hasTrustDialogAccepted
recorded. With the server-side auto-accept disabled in a throwaway copy, the
skill's fallback cleared a genuinely parked dialog in 1.1 s and spawn_worker
took a brand-new case to a live composer in 7.2 s; spawn_workers + sendwait +
last_text then ran end to end.
2026-09-01 02:31:26 +02:00
Codeman maintainer 3518af3a9f docs: correct CLAUDE.md drift and document four undocumented subsystems
Audit of CLAUDE.md against the tree. Verified still accurate: the 31-module
frontend load order (matches index.html exactly), SSE registry parity at
157 = 157 (confirmed by running the parity test), config/ 21 files, types/ 22
domain files, 136 mobile device profiles, the version line, and every Quick
Reference command.

Drift corrected: 24 route modules to 25, ~220 handlers to ~227, system-routes
51 to 56, app.js ~5K lines to ~6.7K and 30 modules to 31, install.sh 92KB to
104KB. Completed the CLI resolver inventory, which was missing
deepseek-cli-resolver and omp-cli-resolver even though both modes are
documented, and named the shared cli-executable-resolver lookup chain.

Filled the gaps found by sweeping every src module against the file:

- Owner tab layouts (COD-359) had 6 source modules, 7 test files, 2 routes, an
  SSE event and a state.json key, with zero mentions anywhere in CLAUDE.md or
  docs/. The paragraph records the four things a reader would otherwise get
  wrong: it is backend-only as of 1.24.1 with no frontend consumer, the service
  is the sole mutation boundary, it projects onto PUT /api/session-order rather
  than replacing it, and reconciliation is gated on a successful restore.
- codeman doctor and codeman users were undocumented top-level CLI commands.
- Four subsystems whose invariants lived only in their @fileoverview:
  the workspace-trust dialog recognizer, proc-tree's bounded walk (the
  2026-07-30 incident that took a machine down), deepseek-web-server (one
  child process, deliberately not a shell session), and the Files panel
  search matcher (globs are never compiled to a RegExp).

Also fixes a stale "156 event types" comment in constants.js (actual: 157) and
a contradiction in AGENTS.md, which still carried the retired "never run the
full suite inside a managed tmux session" rule against CLAUDE.md's current
"npm test is the gate and is safe to run bare".

Note: the trust-dialog paragraph documents trustDialogNextKey(), which is part
of a sibling session's in-flight fix for the Claude Code 2.1.252 layout change
(unnumbered, reversed options with "No, exit" highlighted, so a blind carriage
return picks exit and kills the pane). That fix was uncommitted in the shared
tree when this landed, so the doc leads the code until it is committed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NMN8UuvdBim3iM87reuQ9Z
2026-09-01 02:10:09 +02:00
Codeman maintainer e5c5d890aa chore: version packages 2026-08-31 22:34:57 +02:00
Codeman maintainer d5b5f8f618 fix(docker): make the dsh profile install survive pnpm's build-script gate
Follow-up to #350, which fixed the actual blocker (issue #352): `dsh plugin` is
a thin forwarder that `spawnSync`s a literal `pnpm` with no npm fallback, so an
image without pnpm dies at exit 127 and takes the whole build with it.

That PR also pinned an allowlist of the two packages whose lifecycle scripts
pnpm blocked at the time. Replace it with a policy that cannot go stale: pnpm,
unlike npm, refuses dependency build scripts by default and FAILS the install
over it (`ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on pnpm 11.24), and the
names to allow move between rebuilds because `@deepseek-harness-tui/dsh-tui` is
resolved by dist-tag, not pinned: 0.9.3 pulled `@google/genai` (whose script is
a literal `preinstall: no-op`), 0.10.0-beta.x does not. An allowlist of two
names would have let the next tree break the build the same way. Allowing them
wholesale is also the exposure this image already accepts three layers up,
where `npm install -g` runs the install scripts of every transitive dep of the
five CLIs above with no gate at all.

Also correct a comment in the `/api/deepseek/install-profile` route that
asserted the opposite of what #352 proved ("dsh bundles its own package
manager, so no system pnpm is required"). The route's behavior is already
right: dsh's own "pnpm not found on PATH" stderr reaches the caller as the
OPERATION_FAILED detail, so the UI's "add a terminal profile" button names the
fix. Documented the prerequisite in docs/deepseek-integration.md, and taught
the docker-cases image smoke test about `dsh`/`omp` plus the profile check that
`dsh --version` does NOT cover.
2026-08-31 22:26:09 +02:00
Codeman maintainer 7762809202 Merge pull request #350 from opticon454/bugfix/dsh-pnpm
fix(docker): install pnpm for DeepSeek profile
2026-08-31 22:22:30 +02:00
Codeman maintainer 02bbf13b3c chore: version packages 2026-08-30 16:29:14 +02:00
Codeman maintainer da91b4353b Merge pull request #353 from timkjr/omp-mode
feat: add OMP (Oh My Pi) as a new session backend
2026-08-30 16:16:26 +02:00
d fei 47ee49128c style: match the prettier version the lockfile pins
Format check failed twice, on different files each time, because three prettier
versions were in play: package.json says ^3.4.0, package-lock pins 3.8.3 (CI runs
npm ci, so that is the one CI uses), and the local node_modules had 3.9.6. Files
formatted with 3.9.6 were then "fixed" with 3.4.2, pushing session-routes and
system-routes onto a third style — every version change moved the failure to a
different set of files.

Line-break placement in `await import` and a union type only; no logic changes.
2026-08-29 23:49:04 -07:00
d fei 8e5e207386 fix(docker): send the probe body as an object, and explain an unreachable container
The run menu still offered every mode for an attached container. The browser's
actual request showed why:

  POST /api/docker-cases/adopt-preflight -> 400
  {"error":"Invalid input: expected object, received string"}

_api serializes `body` and sets Content-Type itself, and three call sites each
passed an already-stringified body, so it was encoded twice and the server saw a
JSON string where it expects an object. curl was fine throughout, so nothing in
the server logs pointed at it.

Also fixes the design defect underneath: a failed probe fell through to "do not
gate", which silently offered every mode. When the container has been recreated,
is stopped, or the engine is unreachable, the user sees claude, clicks it, and
it can only fail — with the reason visible nowhere. A failed probe now hides
every agent mode (Shell needs no CLI and stays) and shows the server's own
reason at the top of the menu.

Two static guards switched from a character window to brace matching. They
sliced between two call sites, and _loadRunModeHistory's call appears above its
definition, so the slice came out empty and the assertion verified nothing —
the same trap twice in one file.
2026-08-29 23:31:28 -07:00
d fei 5452ad5c5a feat(docker): add a folder picker to both path fields
Both paths in the adoption form had to be typed. Each gets a Browse button
using the same path-input-group markup Link Existing uses, so the two look and
behave alike.

What they can browse differs, and that is the point. The host workspace path
reuses the existing host picker. The container workdir cannot: an adopted
container has nothing mounted at a matching host path, so a host listing would
be a different filesystem — and getting this field wrong is the source of the
opaque OCI chdir error at launch, which makes it the field that most needs to
be clickable.

Adds a read-only POST /api/docker-cases/browse: one `ls` through docker exec, no
writes, no lifecycle, path shell-escaped like every other value. `ls -Ap` marks
directories with a trailing slash and keeps names with spaces intact.

PathPicker takes an optional fetchListing source rather than being forked: the
container variant only swaps where the rows come from, and reuses the rendering,
navigation, Up and Choose/Select unchanged.
2026-08-29 23:31:28 -07:00
d fei 23ab2e77fd fix(files): give the picker a root when the server runs as root
Link Existing's Browse did nothing: GET /api/filesystem/browse answered 403
"No filesystem browse roots are available".

Two rules were fighting. /root is a default blocked tree in the attachment
guard, and Codeman running as root — containers, plenty of servers — makes
homedir() exactly /root, so the picker's own allowlisted Home root was blocked;
the other candidates live under it or do not exist. The root list came out
empty and there was nothing the user could open.

The blocked trees exist to keep ~/.ssh and friends out of reach, not to seal off
the user's own home. Only trees that would swallow a configured root whole are
dropped now: /root goes when Home is it (or sits inside it), /etc holds no
configured root and is untouched. Secrets stay protected — isSensitivePath
independently matches .ssh/, .env and credentials* at any depth, and it is what
the directory probe asks about.

⚠️ Navigation must reuse the same narrowed list the roots were chosen with.
Handing the raw trees downstream admits a root and then refuses every path
inside it, which reads as a picker that opens and does nothing.
2026-08-29 23:31:28 -07:00
d fei 3685ad85bc fix(docker): stop requiring the CLI on the host for a container session
Attaching a container, picking claude and hitting Run gave one line —
`execvp(3) failed.: No such file or directory` — and the run-mode menu offered
every mode. Three separate defects, found on a real deployment.

TmuxManager.createSession resolved the CLI directory without distinguishing a
docker session, so a host with no claude threw, the catch fell back to a direct
PTY, and that PTY exec'd the CLI on the HOST. The failure surfaced as a bare
execvp error naming nothing. A docker session runs its CLI inside the container;
the host does not need it. All eight modes now sit behind a cliRunsInContainer
guard, and whether the container has the CLI is settled by the adoption
preflight or the image gate before launch.

The running check used a bare double quote and command substitution. The whole
chain is embedded in an outer `bash -c "…"`, so the unescaped quote closed that
string early and the remainder was re-tokenized. It is now a `grep -qx` pipeline
using only the single-quote form every other line in the builder already uses.

Claude Code refuses --dangerously-skip-permissions as root. Our base image runs
a non-root user, so an owned container never hit this; an adopted container's
user belongs to its owner and is frequently root, and keeping the flag killed
the pane with a message visible only inside the container. The preflight now
reports runsAsRoot and the launch chain drops the flag for it.

The menu also showed every mode because the container CLI probe only started
when the menu opened. It is warmed when the case is selected instead.
2026-08-29 23:31:28 -07:00
d fei 06e7cbe286 fix(docker): probe the container's CLIs live instead of trusting attach time
Storing the container's CLIs on the case at attach time left two gaps: a case
linked before that field existed has none at all, and a container's CLIs can be
installed or removed long after it was linked. A real deployment hit the first
one — the host had only codex, the container only claude, and with no stored
list the menu still gated on the host and hid the mode that actually worked.

The probe now runs when a container case is selected, reusing the existing
adopt-preflight endpoint, so there is no new backend surface. Results are cached
per case for the page's lifetime, since the menu opens often and the probe is a
`docker exec` round trip; a concurrent probe for the same case is deduplicated
with an in-flight marker.

A failed probe leaves the cache empty, which the caller reads as "unknown" and
therefore does not gate. Hiding every mode because one probe failed is worse
than offering one that turns out to be missing, which the launch path already
refuses with a specific message.

The repaint only happens while the menu is still open, so a late answer cannot
make the list jump under a user who already closed it.
2026-08-29 21:19:17 -07:00
d fei 8b20f5b1f8 fix(docker): probe container CLIs by their real binary name
The adoption preflight used the mode name as the binary name. claude, codex,
opencode, gemini and pi happen to match, so it never showed — but antigravity
ships as `agy` and deepseek as `dsh`, so a container that has either was
reported as not having it, and the mode was silently dropped from the case.

Adds a MODE_BINARIES map, single-sourced with defaultDockerCommandForMode, which
launches those same binaries. Probing and result filtering share one `binaryFor`
so the two cannot drift apart.
2026-08-29 21:02:19 -07:00
d fei 2f83a37c6d feat(docker): take run-mode availability from the container
The run-mode dropdown hides CLIs that are not installed on the HOST (#201). That
is right for local sessions and wrong for a container case, whose agents run
inside the container: a host with no claude installed hides the mode while the
container ships one, which is exactly what happened on a real deployment.

The adoption preflight already probes what the container has, so that result is
persisted on the case and surfaced through CaseInfo. Docker cases gate on it;
every other case keeps the host probe unchanged.

An absent list reads as "do not gate" rather than "nothing available": an owned
container runs our base image, which ships every CLI, and treating unknown as
empty would leave the menu with Shell alone.
2026-08-29 21:02:19 -07:00
d fei 34c12ca18b feat(docker): make the container field a picker you can also type into
Typing a container name from memory is error-prone. The field becomes a native
datalist: pick from the engine's containers, type to filter, or type a name that
is not listed (the engine may be remote, or the container may not exist yet).
A datalist gives all three natively, so no dropdown state machine is introduced.

Adds listDockerContainers and GET /api/docker-hosts/:hostId/containers, following
the listRemoteCodemanSessions discovery precedent: read-only and never throwing,
so an unreachable daemon returns an empty list and the field degrades to plain
text instead of erroring.

Stopped containers stay in the list, sorted after running ones and labelled.
Attaching does require a running container, but hiding stopped ones turns "my
container is not in the list" into a dead end, while showing
`Exited (137) 8 days ago` says exactly what to fix.
2026-08-29 21:02:19 -07:00
d fei e2f750cb30 i18n(docker): translate the attach panel, and unblock translation
The new strings were English only. Adding entries surfaced a deeper problem: the
translator matches whole text nodes and skips `code`/`pre`, so an inline `<code>`
mid-sentence splits a hint into fragments that can never match an entry — which is
why the panel's existing "Build it once with <code>...</code>" hint was never
translated either.

Drops the inline markup from the new hints so each is a single text node, then
adds the zh-CN entries. The brand name goes through the existing {name}
placeholder.

Server-side error bodies are deliberately not added: the client receives them
already interpolated with a concrete container name, so a template key could
never match.
2026-08-29 21:02:19 -07:00
d fei bb45909169 feat(docker): link to container attach from the Create New tab
Attaching lived only on the Docker tab, but the place users look for anything
container-shaped is the "Run in an isolated Docker container" checkbox on Create
New. A feature nobody can find is a feature nobody has.

Adds a one-click link there that switches to the Docker tab, turns the toggle on
and focuses the container field. Reuses switchCaseModalTab and the existing sync
helper; no new CSS.
2026-08-29 21:02:19 -07:00
d fei c98a59d709 fix(docker): verify the container workdir and end the probe with exit 0
Two defects that only a real container exposes.

The probe chained `command -v X && echo X` with semicolons, and a script's exit
status is its last command's. A container without the last probed CLI made the
whole `sh -lc` exit 1, so a perfectly healthy container with tmux and claude was
reported as "could not exec into the container". A missing CLI is data here, not
failure, so the script now ends with `exit 0`.

containerWorkdir defaulted to hostWorkspacePath. That default holds for an owned
container only because the create-time bind mount puts the host directory at that
exact path; attaching mounts nothing, so the two are independent facts. A host
path absent inside the container makes `docker exec --workdir` fail with an OCI
chdir error that surfaces in the pane as a bare "execvp failed". The preflight now
proves the directory exists inside the container and refuses at link time.
2026-08-29 21:02:19 -07:00
d fei bc55b6b0da feat(docker): add the attach-an-existing-container panel
The Docker tab gains an "Attach to an existing container" toggle. Ticking it
swaps the create-time fields (image, network, advanced) — which describe a
`docker create` attaching never runs — for the container name, and routes the
submit to the adopt endpoint.

Reuses the existing linkDockerCase flow end to end: only the final call differs.
The docker-host upsert still applies, since it is what resolves the
engine/context/daemon for `docker exec`; its create-time fields are simply never
read for an attached case.
2026-08-29 21:02:19 -07:00
d fei 15eebde832 feat(docker): attach a case to an already-running container
Docker cases could only run in a container Codeman created itself. Attaching to
one the user already built and runs means Codeman must leave that container's
lifecycle completely alone, which the launch chain could not do: it was
`image inspect` -> `inspect || create` -> `start` -> `exec`.

Adds `DockerCase.owned`, mirroring the `owned:false` contract remote-SSH already
uses for attached sessions. Absent (every existing case) means owned, so current
behaviour is byte-identical. `false` means the container belongs to the user and
Codeman may only exec into it.

The launch chain for an attached container only looks, then execs: no image gate
(the image is theirs), no create, and no `start` — starting a container we do not
own is the very mutation attaching promises not to perform. A missing or stopped
container fails closed with an actionable message instead. Credential seeding is
skipped too: those copies read from create-time read-only mounts that do not
exist here, and writing host credentials into someone's container is not ours to
do, so its CLIs must already be authenticated inside it.

Four fail-closed guards. buildDockerStopCommand and buildDockerRemoveCommand
throw during pure string construction, so no caller bug can turn into a
`docker stop`/`rm` on a container we do not own; removeDockerContainer refuses
again at the lowest layer; drift reports "none" for an attached container, which
carries no `codeman.confighash` label and would otherwise always look drifted and
409 the launch gate forever; and the orphan reaper skips attached containers
through a check deliberately independent of the two conditions already covering
them.

`owned` is applied AFTER the config hash is computed. dockerConfigHash takes an
explicit field list, so ownership can never shift an existing case's hash — if it
did, every pre-existing case would trip the drift gate at once, and the remedy
the UI offers is "recreate the container".

Adds POST /api/cases/docker-adopt and a read-only
POST /api/docker-cases/adopt-preflight. The preflight refuses at LINK time rather
than at session launch, where the only ways out would be a dead pane or starting
a container we do not own.

Tests assert the negative guarantee directly — that create, start, stop, rm,
restart and kill are absent from the generated commands while `docker exec -it`
and `new-session -A` remain — since it cannot be observed by using the feature.
2026-08-29 21:02:19 -07:00
timkjr da5f5447d0 fix(remote): never auto-revive a remote session after a clean agent exit
The COD-108 reconnect watcher treated any dead local pane as a dropped
transport and re-ran the pane command — so a normal ctrl-c/ctrl-d on a
remote claude/opencode/omp auto-spawned a FRESH agent (claude only
looked correct because its '--session-id || --resume' fallback resumed,
with a loud 'already in use' error first).

Distinguish a transport drop from an intentional exit: only reconnect
when the durable remote tmux session (codeman-ssh-*) is verifiably
still alive on the remote host. A clean exit tears that session down;
the watcher now probes it via ssh has-session and skips (remote-gone)
when it is gone OR unknown (fail closed). The probe is cached
per-session and fired async so the 5s tick never blocks on ssh.

Tests: 3 new cases pinning remote-gone / unknown / alive decisions.
Verified live: all remote CLIs stay dead after ctrl-c/ctrl-d.
2026-08-29 17:43:07 -05:00
timkjr b6d0f1fa32 fix(omp): wire OMP into install.sh's CLI detection (it had none)
Every other CLI (claude/opencode/codex/gemini/antigravity/pi/grok/dsh) has
a check_*/get_*_path pair wired into install.sh's detection loop and the
"no AI CLI found" aggregate checks. OMP had neither -- a user with only
omp installed would be told no CLI was found and offered to install
Claude Code or OpenCode.

Added OMP_SEARCH_PATHS (mirrors src/utils/omp-cli-resolver.ts's
OMP_SEARCH_DIRS) and check_omp()/get_omp_path(), wired into both
aggregate conditions (the interactive install-menu trigger and the
end-of-run reminder) and added omp's real vendor curl one-liner to the
reminder block. The DeepSeek Harness line was never in that reminder to
begin with -- confirmed it has no vendor one-liner (dsh installs via
Codeman's own API after the server is already up), so it stays out, with
an explanatory line instead.

Also fixed the "Skip" menu text, which was missing Gemini and DeepSeek
Harness from its example list independent of the omp gap, and the same
stale sibling-CLI-list bug (missing DeepSeek Harness and OMP, "the
eight"/"这七个") in README.md and the repo's existing README.zh-CN.md.
2026-08-28 15:19:19 -05:00
timkjr 65e994d29a fix(omp): correct docs/counts/URLs, resolver install-path order, stray comment + CSS
Small cleanup items from upstream review (Ark0N/Codeman#353):

- OMP_SEARCH_DIRS now leads with ~/.local/bin, matching omp.sh's real
  installer target (~/.omp/bin was an earlier unverified guess, confirmed
  wrong against a real --no-cache Docker build).
- docs/omp-integration.md: fixed the dead GitHub URL (can1357/omp ->
  can1357/oh-my-pi), corrected the CLI count (ninth backend, tenth
  SessionMode incl. shell -- not eighth), matched the install-path guidance
  to the resolver fix, updated the version example to the actually-tested
  18.0.8, and added a Docker-section caveat: --resume pinning does not
  currently reach an in-container omp process, since Docker panes never see
  ompConfig.
- docs/architecture-invariants.md: fixed a heading missing ", OMP" (CLAUDE.md
  already linked to the -omp anchor, so the link was dead) and added an OMP
  specifics paragraph -- the one external CLI missing an entry in this doc.
- .changeset/omp-backend.md: corrected the sibling-CLI list (was missing Pi,
  Grok, and DeepSeek Harness) and the backend count.
- Removed a stray orphaned comment fragment in the quick-start docker branch
  and split two CSS lines that had two declarations jammed onto one line.
2026-08-28 14:37:28 -05:00
timkjr f18dccace1 fix: don't discard codex/gemini/antigravity conversations on Resume; fix DELETE ownership dup + missing broadcast
resumeHistorySession() creates the resumed row in its own mode via a
modeConfigKey map (opencode/pi/grok/omp -> continueSession, deepseek ->
resumeSession) and retires the old row afterward. codex, gemini and
antigravity were missing from that map, so resuming one of their rows
started a brand-new session with NO continuation while still deleting
the row it came from -- silent data loss dressed as the duplicate-row
fix. Gate row retirement on continuesSomething (true only for modes that
actually got a continuation config) instead of wiring an unverified
sessionId->native-conversation-id assumption for the three affected CLIs.

DELETE /api/sessions/:id reimplemented the ownership 404 check inline in
two places instead of going through findSessionOrFail, and its
persisted-only-session branch never broadcast session:deleted, so other
open tabs kept the retired row until their next unrelated fetch. Extract
the shared 404 into sessionNotFoundError(), add findPersistedSessionOrFail()
alongside findSessionOrFail() in route-helpers.ts (same ownership
contract, returns a SessionState instead of a live Session), and use both
from the route instead of inline checks. Add the missing broadcast.
2026-08-28 14:03:07 -05:00
timkjr 2ee2eacb4b fix(omp): clamp OMP_AUTH_BROKER_URL/TOKEN, correct the env-allowlist docs
The docs claimed omp "has no documented vendor-key namespace of its own"
and "the multi-user clamp has nothing to gate" for omp — both false. Per
omp's own docs/environment-variables.md, it reads ~40 provider keys from
env (pi's known 34-key problem in the same shape), and its own knobs are
mostly PI_* (already globally allowlisted): PI_CONFIG_DIR,
PI_CODING_AGENT_DIR, PI_CODING_AGENT_SESSION_DIR, PI_SUBPROCESS_CMD,
PI_SHELL_PREFIX. The first three also move the ~/.omp tree
omp-session-resolver.ts/omp-transcript.ts hardcode, silently degrading
pinning/history — a known gap shared with pi, documented but not fixed
here.

The OMP_* prefix this PR adds brings in OMP_AUTH_BROKER_URL/
OMP_AUTH_BROKER_TOKEN, where omp resolves credentials from — the same
shape DEEPSEEK_BASE_URL is already dropped for in
clampEnvOverridesForOwner(). Add both to OWNER_CLAMPED_ENV_KEYS so a
non-granted owner in multi-user mode can't redirect them, and correct the
false claims in CLAUDE.md, docs/omp-integration.md, and the stale
resolveOmpHome() comment. Also documents omp's default
tools.approvalMode: yolo, which was previously unstated.
2026-08-28 13:45:12 -05:00
timkjr c4f6eb1e5e fix(omp): resolve and pin the respawn session id only at actual respawn time
findLatestOmpSessionId()'s newest-mtime pin ran eagerly inside
_buildRespawnPaneOptions(), which startInteractive() calls unconditionally
on every boot-recovery reattach — before anything checks whether the pane
is actually dead. With two omp tabs in the same case dir, this could pin
an ALIVE pane's session onto whichever sibling's file happened to be
newest on disk, purely as a side effect of building options that might
never lead to a respawn (reported in Ark0N/Codeman#353 review).

Move resolution out of the eager builder into _pinOmpRespawnId(), called
explicitly only where a respawn is actually confirmed: the dead-pane
branch in _setupOrAttachMuxSession() and reattachRemote(). Add
resolveAndClaimOmpSessionId(), which verifies each candidate's own file
header (cwd) rather than trusting the mangled-directory match alone, and
tracks claimed ids in a process-wide registry so two ambiguous resolutions
can't both pick the same sibling's conversation.
2026-08-28 13:18:25 -05:00
timkjrandClaude Sonnet 5 ab83d8ffec fix(omp): a fresh "Run OMP" click no longer silently resumes an old conversation
Found live 2026-08-27 by Tim: clicking Run OMP to start a brand-new session
in a case directory with prior omp history launched --resume <old-id>
instead of a clean `omp` invocation.

Root cause: Session._resolvedOmpRespawnConfig() resolves-and-pins the
newest on-disk omp conversation as a side effect on this._ompConfig. That
is correct when reattaching to an ALREADY-TRACKED mux session (a dead-pane
respawn, or a boot-recovery reattach - the constructor sets _muxSession
from persisted state before startInteractive() ever runs there), but it
ran unconditionally. startInteractive() computes
`respawnPaneOptions: this._buildRespawnPaneOptions()` eagerly in the same
object literal that builds `createSessionOptions.ompConfig: this._ompConfig`,
so for a genuinely brand-new session (no muxSession in its create config,
_muxSession still null) the resolve-and-pin side effect ran and poisoned
this._ompConfig before that field was even read.

Fix: gate the resolve-and-pin logic on `this._muxSession` already being
set. A fresh session has no muxSession yet and now passes through
untouched; a real reattach (muxSession present since construction) keeps
resolving and pinning exactly as before.

Verified live in production against the exact reported scenario (a fresh
omp session in a case dir with 8+ hours of prior omp history) - confirmed
both via the API (ompConfig stays empty, claudeSessionId equals the
session's own id) and visually in the GUI. Regression test constructs a
real Session + TmuxManager to exercise the actual private-method
interaction directly, since no existing test called startInteractive() at
all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjrandClaude Sonnet 5 1829fe91af docs(omp): add docs/omp-integration.md, matching sibling CLI docs
OMP was the one external CLI mode with no dedicated user-guide doc, unlike
opencode/pi/grok/deepseek which each have one. Covers install, auth (omp
owns its own entirely - no Codeman-side login flow or bypass switch),
what Codeman wires up (OmpConfig), the exact-id pinning mechanism and the
directory-mangling bug behind it, kill-survival via transcript scanning,
terminal behavior, Docker/remote-SSH cases, and known gaps (no idle hook,
mid-turn kill data loss, unverified symlinked-$HOME behavior).

Cross-referenced from README.md's Multi-CLI doc list and docs/docker-cases.md's
credential-seeding summary (which now also documents OMP's sessions/-is-shared
exception to the seed-everything pattern the other CLIs use).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjrandClaude Sonnet 5 d74cde759b feat(omp): install omp in the docker agent image, isolate its credentials
OMP had full routing at the Docker layer (default pane command, schema) but
was never actually installed in docker/agent.Dockerfile, and had no
credential-isolation entry in docker-hosts.ts's CRED_STORES - a Docker-mode
OMP session would have failed with "omp: command not found", and even with
the binary present would have had no config/auth seeded, despite the README
already claiming OMP has "seamless auth, isolated credentials" in Docker.

- docker/agent.Dockerfile: install omp via its own installer (standalone
  binary, same shape as grok/antigravity - not on npm). Verified against a
  real --no-cache build: the installer actually targets ~/.local/bin, not
  ~/.omp/bin as the resolver's OMP_SEARCH_DIRS ordering would suggest -
  confirmed omp/18.0.8 installs and runs correctly inside the image.
- src/docker-hosts.ts: add a .omp/agent CRED_STORES entry. Unlike every
  sibling CLI in this family, sessions/ is SHARED (RW), not seeded: Codeman
  reads ~/.omp/agent/sessions/**/*.jsonl host-side for history recovery and
  --resume pinning (omp-transcript.ts, omp-session-resolver.ts), the same
  reason codex's sessions/ is shared rather than seeded. Seeding it instead
  would silently break the kill-survival feature for Docker cases. Only the
  small config files (config.yml/mcp.json/models.yml/settings.yml) are
  seeded; the SQLite caches and terminal-sessions/ stay container-local.
- test/docker-hosts.test.ts: pin the new CRED_STORES entry's behavior.

Found in passing (NOT fixed here, unrelated and pre-existing on master): the
agent image's DeepSeek (dsh) plugin-install step currently fails on a fresh
build ("pnpm not found on PATH"), confirmed via git diff against
origin/master that this line is untouched by this branch. Worth a separate
issue/PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjrandClaude Sonnet 5 853681f970 harden(omp): resume-path test coverage, silent-fallback logging, cwd validation
Follow-up from a full-branch review pass (Opus) of the omp-mode integration:

- Add pinning tests for resolveOmpConfigForCreate() (session-routes.ts),
  exported to make it testable: the exact "resume this OMP row from
  history" pipeline that mangleOmpWorkingDir's earlier bug lived in had
  zero coverage despite being the resolver module's whole reason to exist.
- Log a warning when findLatestOmpSessionId() finds nothing on disk and
  continuation silently degrades to omp's own ambiguous --continue,
  in both call sites (session create and respawn pinning) - previously
  silent, making the degradation invisible to anyone debugging it.
- Require an absolute cwd before trusting a session file's working
  directory in omp-transcript.ts's parser, so a corrupted/malformed
  session file can't point a downstream resume at a relative or empty
  path.
- Document (don't speculatively fix) an unverified symlinked-$HOME edge
  case in mangleOmpWorkingDir(): the review's suggested realpath() fix
  assumes omp itself resolves symlinks before mangling, which is
  unconfirmed - guessing wrong there would trade one silent mismatch
  for a different one.
- Incidental: fixed unrelated pre-existing prettier drift in
  session-routes.ts (antigravity/opencode dynamic import line-wrapping)
  that was blocking the pre-commit formatting gate on this file.

Confirmed as a non-issue: the model-name regex allowing "/" is
intentional (provider/model ids like "crof/glm-5.2" were used
successfully in live testing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjrandClaude Sonnet 5 ed983f898b fix(omp): resolve claudeSessionId alias on boot-recovery reattach
Two bugs compounded to break continuation pinning on every real OMP
case (only /tmp-based manual testing happened to work by coincidence):

1. startInteractive() had a second, unconditional claudeSessionId
   assignment after the mux branch that clobbered its correctly
   resolved value back to the session's own id on every mux path.

2. mangleOmpWorkingDir() assumed omp mirrors Claude Code's directory
   naming (home prefix kept), but omp actually strips $HOME first.
   findLatestOmpSessionId() was silently returning null for every
   case under ~/codeman-cases/, so resumeSessionId never resolved for
   any real case dir - only /tmp paths (outside $HOME) worked, which
   is every dir this feature was previously tested against.

Verified live: killed and relaunched the omp-verify server process
mid-session (plain reattach, pane stayed alive) and confirmed
claudeSessionId now resolves to the real omp transcript uuid instead
of the Codeman session's own id.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjr 54a930c80e feat(omp): survive a full session kill by reading omp's own transcripts
Claude conversations survive "Kill Tmux & Claude" because Codeman reads
them back independently from ~/.claude/projects, not from its own
session bookkeeping. omp conversations had no equivalent: kill the
Codeman session and the conversation vanished from Past Sessions
entirely, even though omp itself never forgot it on disk.

Adds omp-transcript.ts, a scanner over omp's own
~/.omp/agent/sessions/<mangled-cwd>/<uuid>.jsonl files (the same shape
as Claude Code's own transcript scanner, but simpler -- these files are
small enough to read whole instead of doing head/tail windows). Each
file's own "session" header line carries the real cwd and session id
directly, so unlike Claude's mangled-directory-name decoding this
never has to guess. Wired into gatherUnifiedInputs() as a second
history source alongside the Claude scan, and HistoryInput/
mergeUnifiedSessions() now carry an optional `mode` so a non-claude
history-only row still gets a real mode badge.

Also fixes the ambiguity behind the "continue picks the wrong
conversation" report from this session's testing: omp mints its OWN
session uuid, unrelated to Codeman's, so a live/persisted row and its
own history-scan row would otherwise show up as two separate entries
for the same conversation the moment the id gets resolved. Reuses the
existing claudeSessionId alias field (mergeUnifiedSessions' fold-into-
owner mechanism) to point at the resolved omp id, threading it through
every place `_claudeSessionId` gets (re)computed -- the constructor,
_resolvedOmpRespawnConfig, and a new _maybeCaptureOmpSessionId() that
opportunistically resolves it the first time a brand-new omp session
(one that has never gone through a respawn) goes idle.

Also closes a THIRD instance of the "ompConfig never got wired in
here" gap this session kept finding: restoreMuxSessions() in server.ts
restores every sibling CLI's config from persisted state on boot except
omp's, so a boot-recovered omp session always lost its resolved resume
id and fell back to guessing again.

Verified live end-to-end: told a session a secret, killed it fully
(Kill Tmux equivalent, killMux=true -- the Codeman session AND its tmux
pane both gone), and the conversation still showed up in the unified
list as a history-sourced row with the real first prompt as its title
and an omp mode badge, keyed by omp's own session id.

Known remaining gap, not fixed here: the claudeSessionId alias doesn't
yet resolve reliably on every boot-recovery path for a session that
was never respawned while alive (e.g. a plain re-attach to a pane that
was never dead) -- worth a follow-up, but doesn't affect the two things
that matter most: the conversation surviving a kill, and continuation
correctness once an id has been resolved (which happens on the very
next respawn either way).
2026-08-28 11:32:30 -05:00
timkjr 4c332c6141 fix(omp): retire the old row on resume, and let DELETE remove persisted-only sessions
Every non-claude "Resume" click creates a brand-new Codeman session
(there is no id to reattach to), but the old row was never cleaned up
-- click resume on the same conversation a few times and the session
list fills up with duplicate rows sharing one name. resumeHistorySession
now retires the row it resumed from after the new one starts.

That retirement needs DELETE to actually work on a row that was never
live in the first place (the normal case for anything showing up in
"Resume Conversation"): findSessionOrFail only checks the in-memory
live-session map, so DELETE 404s on a persisted-only entry today. Give
the route a fallback: when the id isn't live, look it up in persisted
state instead and demote/remove it there (respecting the existing
pinned-session protection). Verified live against a real persisted-only
row via the API, and added route-test coverage for both the success
and still-truly-unknown-id cases (which needed a demoteOrRemoveSession
mock the route harness didn't have).

Also includes an unrelated pre-existing prettier drift fix picked up
by npm run format (omp-cli-resolver.ts, antigravity/opencode import
wrapping in session-routes.ts).
2026-08-28 11:32:30 -05:00
timkjr 253599ce9c fix(omp): wire ompConfig into respawnPane and default to --continue there
respawnPane() -- the path used when a session's pane died (crash, idle
respawn, or the user's own /exit) but the Codeman session object is
still tracked -- never had ompConfig wired through at all, in either
its options destructure or its inner buildSpawnCommand() call. This is
a gap in the original OMP patch, distinct from the resumeHistorySession
fix (which only covers a session that has been fully closed and shows
up as a history row): reselecting a tab whose CLI process just exited
goes through this path instead, and always launched a bare, contextless
`omp` no matter what.

Beyond the wiring, respawning a dead pane is semantically different
from creating a brand-new session: the conversation is still "this
session" to the user, so _buildRespawnPaneOptions() now defaults
ompConfig to continueSession:true unless the session already carries
an explicit resumeSessionId (which still wins in buildOmpCommand).

Verified live: told a session a secret, exited OMP so the pane died
(session and tmux both left alone), forced the exact dead-pane-respawn
path, and the new process replied with the secret -- confirming
`omp --continue` fired instead of a blank omp.
2026-08-28 11:32:30 -05:00
timkjr 3e1a0e679f fix(omp): resume by mode, not silently as claude, and support --continue
resumeHistorySession() never sent mode when recreating a session from a
history/session-manager row, so the server default silently opened a
plain Claude session for every non-claude row -- reproduced live: OMP
rows spawned Claude sessions on click. Thread the row's mode through
every call site (welcome list, session manager, mobile overview) and
only send the Claude-specific resumeSessionId for claude rows.

Codeman has no live PTY-reattach outside server boot, and it's moot for
OMP anyway (exiting it kills the pane's only process), so route the
non-claude relaunch through each CLI's own continue-most-recent flag
instead of a context-free fresh start. OMP never got one: buildOmpCommand
only implemented --model/--resume despite omp --help documenting
-c/--continue. Added continueSession to OmpConfig end-to-end (type,
schema, builder) mirroring the existing opencode/pi/grok/deepseek
fields, and wired resumeHistorySession to use it.

Verified live: told a real omp session a secret, exited it, closed the
tab without killing tmux, relaunched with --continue in the same
directory, and had it recall the secret.
2026-08-28 11:32:30 -05:00
timkjr 7ec48adcc8 fix(omp): keep external-CLI mode enumerations complete in skill docs
Two prose lists in skills/codeman/ named some but not all external CLI
modes after the omp-mode rebase, which is exactly the drift
test/agent-skill-mode-lists.test.ts exists to catch: SKILL.md's
no-hook-signals list was missing omp, and endpoints.md's version-probe
sentence named pi/grok/omp as a bare 3-mode run with no matching class.
2026-08-28 11:32:30 -05:00
timkjr 9841f4ffb9 refactor(omp): align omp resolver + doctor with upstream shared CLI resolver
- omp-cli-resolver.ts already uses createCliExecutableResolver; add dedicated
  test/omp-cli-resolver.test.ts mirroring pi's (version-probe accept/reject,
  negative-cache backoff, VITEST hermeticity gate)
- dependency-registry omp entry now requires OMP_VERSION_REGEX match like pi,
  so codeman doctor and the run-mode resolver agree on what counts as installed
- system-routes /api/omp/status surfaces version
2026-08-28 11:32:30 -05:00
Codeman maintainer d8688dc143 fix(web): drop the provider label from the plan-usage chip when there is only one
The chip prefixes every row with the provider name, so a machine that only
has Claude limits renders "CLAUDE 5H 60% 7D 23%" — a 46px label naming the
only thing it could possibly be. The name exists to tell two rows apart, so
it should only appear when there are two.

updatePlanUsageChip() now checks whether both Claude and Codex actually have
windows before building the rows, and emits the .pu-provider span only in
that case. The tooltip keeps naming the provider in both cases: it has the
room, and the chip no longer does.

Verified in a browser on an isolated beta instance: Claude-only renders bare
windows with no .pu-provider in the DOM, Codex-only the same, and the
two-provider chip is byte-identical to before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 13:59:40 +02:00
Codeman maintainer 23fae0c5af chore: version packages
Codex plan usage in the header chip (#346), a visible inline rename in
the session sidebar (#345), and the install.sh Tailscale re-run fix plus
the README network-access prompt description.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:25:32 +02:00
Ark0N e4699159e9 Merge pull request #346 from JackStuart/codex/show-codex-usage-limits
feat(web): show Codex plan usage in header
2026-08-28 00:14:54 +02:00
Ark0N da085f5f7f Merge pull request #345 from fibr/fix/sidebar-inline-rename
fix(ui): show inline rename text in session sidebar
2026-08-28 00:14:48 +02:00
Codeman maintainer 23e32b22d5 docs(readme): describe the actual three-way network-access prompt
The installer bullet still described a two-way choice with 0.0.0.0 as "the
default", which predates the Tailscale option. The prompt has offered three
choices for a while (Tailscale / any device on your network / this machine
only), and the default is computed from what is already on the machine rather
than being fixed at 0.0.0.0.

Now states all three options, that the Tailscale one is a loopback bind
fronted by `tailscale serve` with the tailnet as the login, and how the
highlighted default is chosen. Line 220 already documented the Tailscale
option correctly; this was the only stale spot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:02:33 +02:00
Devvyn 26b4ffbb0f fix(docker): install pnpm for DeepSeek profile 2026-08-27 20:47:57 +08:00
Devvyn e2179bd530 chore(docker): remove local handover references 2026-08-27 19:42:37 +08:00
Devvyn b85f7659b7 feat(docker): add Compose deployment support 2026-08-27 19:38:38 +08:00
timkjr e82380e14a fix(ui): close unclosed CSS blocks that killed the stylesheet tail
The rebase hand-repair dropped the closing brace of .welcome-btn-pi:hover
and .btn-toolbar.btn-run.mode-pi:hover before the inserted OMP rules.
The browser CSS parser drops every rule after an unclosed block, so the
deployed UI rendered as unstyled text bars (only ~456 of ~2583 rules
applied). Verified clean via esbuild --minify (no css-syntax-error) and
rebuilt dist.
2026-08-26 20:10:34 -05:00
timkjr c0423bf560 fix(omp): complete omp wiring in UI files, skill docs, and tests after rebase 2026-08-26 20:10:34 -05:00
timkjr 4f5678fac4 feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) 2026-08-26 20:05:48 -05:00
Codeman maintainer 7dfb4acf24 fix(install): offer Tailscale setup on re-run instead of losing it to a failed build
The network-access prompt, where Tailscale serve is configured, runs AFTER
the build step. A build failure therefore exits before the question is ever
asked, and a user who then finishes the build by hand (rather than re-running
install.sh) ends up with a healthy loopback-only Codeman, a connected
Tailscale, and no serve mapping — with nothing anywhere pointing at
`install.sh tailscale`, the command that fixes it. Reported from a fresh
Ubuntu 24 install that died on the node-pty compile.

- maybe_offer_tailscale_repair(): on the update/re-run path, detect exactly
  that state (loopback bind + tailscale Running + no serve mapping fronting
  Codeman) and offer the retrofit. Silent for a deliberate non-loopback bind,
  silent once a mapping exists, silent when tailscale is absent, and prints
  the command instead of prompting when non-interactive. Returns 0 even when
  setup fails so it can never abort an update.
- print_security_notice(): the loopback branch now names
  `install.sh tailscale` when Tailscale is installed on the box, rather than
  the generic "tailscale serve / cloudflared tunnel" advice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 18:29:15 +02:00
Codeman maintainer d3f851a5e5 chore: version packages
install.sh installs a build toolchain on Linux (node-pty has no Linux
prebuild, so a stock Ubuntu 24 server died inside node-gyp with
"not found: make"), plus review hardening for #339: the write-queue
reset paths now release the one-chunk-in-flight gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 18:04:44 +02:00
Ark0N 5f8d4de443 Merge pull request #340 from aakhter/pr/cod-341-file-viewer-search
feat(file-viewer): COD-341 search the full workspace
2026-08-26 18:03:21 +02:00
Ark0N 00b32ad2b8 Merge pull request #339 from dignfei/fix/terminal-live-write-backpressure
fix(terminal): bound live xterm backpressure
2026-08-26 18:03:14 +02:00
Jack Stuart b00ab3ceea feat(web): show Codex plan usage in header 2026-08-26 18:43:52 +08:00
Sergei Lupashin 134e200aec fix(ui): show inline rename text in session sidebar 2026-08-25 19:41:48 +02:00
Codeman maintainer a51563ce1f chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 19:14:41 +02:00
Ark0N ca5fe1ab3e Merge pull request #338 from Ark0N/feat/vertical-rail-detailed-rows
Vertical tab rail: detailed rows (created / working / status), plus a rename-cancel fix
2026-08-25 19:12:57 +02:00
Ark0N 9cd10afdc9 Merge pull request #341 from Ark0N/feat/deepseek-agent-workers
Spawn and drive DeepSeek Harness workers from the codeman agent skill
2026-08-25 19:12:45 +02:00
Ark0N 975705ad87 Merge pull request #337 from Ark0N/feat/deepseek-harness
feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
2026-08-25 19:12:01 +02:00
Codeman maintainer 93a1042bb3 Merge remote-tracking branch 'origin/feat/deepseek-harness' into feat/deepseek-agent-workers
# Conflicts:
#	CLAUDE.md
2026-08-25 19:02:34 +02:00
Codeman maintainer a628737d1f fix(deepseek): review-driven hardening across the harness integration
Fifteen review findings on the dsh mode, the serious ones first:

- Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys.
  _configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into
  every dsh pane and applyEnvOverrides() lands after it, so a non-granted
  owner who could redirect the base URL would have the operator's key sent
  as a bearer credential to a host of their choosing.
- Wait registry: until=stop/blocked is refused on docker and remote-SSH
  dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions).
  The HERDR triple is set via LOCAL tmux setenv, which crosses neither
  docker exec nor ssh, so such a session can never post a hook event and
  the wait burned its whole timeout on every turn.
- Approvals: a dsh item is an ALERT, not an answerable card. The answer
  route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the
  option parser cannot read a third-party TUI's frames, so an answer was a
  blind keystroke into a foreign composer), and the push notification
  carries no Approve/Deny actions for dsh sessions.
- Status shim (v3): --seq is forwarded and the server drops stale retried
  reports inside a 60s window (the TUI retries with backoff, so a retried
  'working' could land after 'blocked' and resolve an approval whose
  dialog was still on screen); 4xx responses exit 0 instead of retrying,
  so one misconfigured session cannot feed the auth rate-limit bucket
  until the hook endpoint 429s for the whole instance.
- Web-UI server: concurrent starts are serialized through a lock (two
  racing POSTs used to pick the same port and orphan the winner), and the
  readiness poll / timeout paths only clear or stop the singleton while it
  is still theirs. First click actually opens the tab now
  (refreshWebviews, not the nonexistent loadWebviews). DELETE
  /api/deepseek/web requires the privileged grant in multi-user mode.
- Cron: deepseek jobs run the same two-part launch gate as the HTTP
  create paths (impl moved into the resolver so all three share it) and no
  longer stamp a Claude default model on the session.
- Parity sweeps: quick-start's docker branch rejects deepSeekConfig like
  the remote branch; the Ralph auto-enable list gained deepseek;
  HookEventType gained agent_working; the phone overview run menu filters
  managed webview records like the desktop menu.
- install.sh: the dsh identity probe closes stdin (under curl|bash a
  child that reads stdin eats the rest of the script), bounds the exec
  with timeout where available, and is memoized to one scan per install.
- Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand
  identity (it rendered as an unstyled UA-grey button); stale markup
  comment about the web shortcut rewritten; clamp docs updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 19:01:29 +02:00
Codeman maintainer 015b865f56 fix(deepseek): review fixes for the transcript reader — docker/remote gate, poll memo, honest pairing docs
Four review findings on the worker-transcript feature:

- Docker and remote-SSH dsh sessions now keep the pane segmenter: their
  transcripts live in the container's / remote host's own ~/.dsh, which
  the local reader can never see, so the transcript path returned
  'nothing said yet' forever and an agent polling such a worker starved
  on an answer that existed. Gated on !session.docker && !session.remote
  (statically pinned) and documented in the integration guide.

- last-response reads are memoized on (path, mtime, size, blocks): the
  skill's last_text polls once per second, and each poll decompressed and
  reparsed the whole file on the event loop even when nothing had been
  appended. An unchanged poll now costs one stat.

- The pairing ladder's comment claimed /new is served by step 2; in truth
  the boot-window transcript wins for as long as it exists (deliberately:
  preferring newest-eligible would hand a worker its busier sibling's
  reply). The comment now states the real tradeoff instead of the
  aspirational one. Same for decodeZstdFrames' 'skipped' wording — a
  corrupt frame truncates the decode there, which is the safe behavior.

- stripReasoningPrefix no longer runs on user prompt text, so a prompt
  containing a literal </think> renders whole in blocks view.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 18:30:41 +02:00
Codeman maintainer 33f77c4680 fix(tabs): review fixes for the detailed rail — width-dialog default, compact wrap pass, rich-aware resets
Three review findings on the detailed-rows feature, all in its edge cases:

- The App Settings width select consulted the handheld defaults blob
  (tabRailWidth: 256) BEFORE the rich-aware default, which the renderer
  never reads — so a tablet's unsized rich rail rendered 320 while the
  dialog said 256, and a routine Save persisted the 256 (below the 288px
  tight threshold, permanently). The chain now mirrors
  applyTabRailWidth()'s actual resolution.

- _setTabRailWidth() re-rendered on a compact flip but never re-ran
  applyTabWrapSettings(), the one owner of the folder line, whose railRich
  input reads the compact class this function just toggled. A rich rail
  dragged below 240px kept emitting folder rows — persistently, for a
  stored width < 240, since the boot wrap pass runs before the class is
  first applied. The wrap pass now re-runs on the flip, with exactly one
  render either way.

- Both reset affordances (handle dblclick, Enter on the handle) reset to
  the hardcoded 256 even on a rich rail, landing it below the tight
  threshold; both now resolve the rich-aware default (320), via a new
  optional defaultWidth input on resolveTabRailKeyboardWidth().

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 18:26:07 +02:00
Codeman maintainer 6261b6f655 feat(skill): spawn and drive DeepSeek Harness workers
The agent skill could spawn a worker in any mode, but it could only
DRIVE a claude one: every other CLI has neither a real end-of-turn
signal nor an answer to read, so the recipes route them through output
markers.

dsh has both halves now -- its harness reports idle/working/blocked to
Codeman, and the previous commit reads its transcript -- so it joins
claude as a mode the four verbs work on unchanged. `spawn_workers alpha
beta:deepseek` is a mixed fleet in one call, and `sendwait` / `last_text`
/ `delete_session` need no per-mode variant.

Preamble 1.20.0 (SKILL.md's §0 heredoc regenerated from it):

- `spawn_worker` grows a deepseek branch that gates on the harness
  composer. ⚠️ Readiness there is NOT the stop signal: the harness
  reports idle at BOOT ~300 ms before its composer paints (measured
  2.26 s vs 2.56 s after spawn), so a send-and-wait fired straight after
  quick-start resolves on the boot edge, reports a turn that never ran,
  and strands the prompt in a pane not yet taking input. Waiting for the
  composer also spends that edge, since signals are edge-triggered.
- `spawn_workers` takes `name[:mode]`, so a mixed fleet stays one
  concurrent call. Case names still have to be unique -- the mode never
  disambiguates two workers that would share a directory.
- `sendwait` asks for `wait:"stop,exit"` instead of the `wait:true`
  default set. That set also carries `idle`, which for an external CLI is
  inferred from output stabilization: on a dsh worker whose TUI repaints
  rarely, the re-wait resolved in 0 ms with `signal:"idle"` on a turn
  with three minutes left to run. It also makes a wrong mode loud -- the
  modes that cannot deliver `stop` answer 400 before writing anything,
  instead of resolving on a flap.
- The self-heal resend carries `delivered:true` forward. The resend is a
  tagged duplicate, so the server truthfully reports `delivered:false`
  about a write it skipped, and §1's cleanup then read a completed turn
  as an undelivered one and kept a finished worker forever.
- dsh workers spawn with the permission posture the Run button sends,
  because the harness default still asks and a worker parked on an
  approval row cannot finish a fan-out. The multi-user clamp still
  applies.

Docs: a worked dsh flow in recipes.md, readiness and the signal rules in
verbs.md, and the corrections this makes necessary -- `stop`/`blocked`
are no longer claude-only, and `last-response` is no longer permanently
empty for deepseek. The integration guide gains a section on reading a
session back and driving one as a worker; its web-UI section was also
stale (that server moved out of a shell session).

The static guard that keeps those lists from naming some external CLIs but
not others is extended rather than exempted: it now knows the three real
classes inside that family (no transcript, no hook signals, and the
positive twin -- the modes whose answers can be read), with the hook class
derived from `hooksAvailableForMode()` so the predicate and the prose
cannot drift apart. Any other partial list still fails, and a new backend
belongs to none of the classes until someone says so.
2026-08-25 04:17:48 +02:00
Codeman maintainer d1bc0c517d feat(deepseek): read dsh session transcripts for last-response
`GET /api/sessions/:id/last-response` is how an agent (and the Response
Viewer) reads what a worker said. DeepSeek was falling through to the
pane segmenter with the other external CLIs, which for this mode is not
merely coarse but wrong: dsh-TUI paints a full-screen splash, so a
`last-response` call on a fresh dsh session answered with its ASCII-art
logo -- and anything polling for a worker's first reply reads that as a
reply.

dsh does not belong in that group. It writes a structured JSONL
transcript per session, so read it. Four things in that file shaped the
reader, all measured against real transcripts on disk:

1. dsh appends ONE ZSTD FRAME PER WRITE, and Node's zlib zstd decoder
   (one-shot and streaming alike) stops at the first frame end: a real
   56-line transcript decoded as 1 line / 158 bytes -- the session header
   alone, i.e. a silent truncation that reads as "nothing said yet"
   forever. `zstdFrameRanges()` walks frame and block headers to find
   exact boundaries; splitting on the 4-byte magic would corrupt
   everything after a magic sequence occurring inside compressed data.
   zstd is resolved at RUNTIME because it landed in Node 22.15 while the
   project floor is 22.0, so an older Node keeps the pane behaviour.
2. Every turn also records a plugin-sourced `user/message` (the runtime
   context snapshot), which must not render as the user's own words.
3. A turn that ends in an error carries the provider's message; it is
   surfaced as `Turn error: …` (and a non-error early stop as
   `Turn ended: …`) rather than as an empty string, which an agent reads
   as "still thinking" through fifteen polls.
4. Reply text is assembled per (turn, step): a finalized message wins and
   the streamed deltas fill in only for a step that never finalized, so a
   partial answer is readable mid-turn and never doubled. "Finalized" is
   tracked as a set of steps rather than as non-empty text, because a
   step whose whole reply was reasoning strips to '' at the `</think>`
   boundary and would otherwise resurrect the raw deltas in its place.

Session-to-transcript pairing is by the transcript's own header `cwd`
plus a boot window against the session's createdAt, never by
reproducing dsh's directory mangling (already two forms on disk) and
never by newest-mtime alone -- mtime alone handed a freshly spawned
worker its predecessor's answer in the same case directory.

An empty result still wins over the pane; only a Node that cannot decode
zstd falls back to it.
2026-08-25 04:14:26 +02:00
Codeman maintainer c30dfaf0e7 fix(deepseek): run the web UI server in the background, not in a shell tab
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a
shell tab running the server next to it. The shell was deliberate - the server
lived in an ordinary session so it was visible, scrollable, killable and died
with its tab, and nothing new had to supervise a long-lived HTTP server. That
reasoning was sound and the result was still wrong in use: opening a dashboard
should open one tab, and after the first launch the terminal is pure noise.

The server moves to a background child process owned by a new
`src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session
gave away for free is now explicit, which is most of the module:

- Exactly one server. A second click reuses the running one instead of racing
  it for a port; the session flow could not do this at all, because two clicks
  were simply two sessions.
- Restarted when the requested authority changes. `--trusted-host` fences dsh's
  own /api against the browser authority, and a Codeman reachable at both
  loopback and a tailnet name has two. Reusing a server fenced for the other
  origin renders a page whose every call 403s, which reads as a broken
  dashboard rather than a misconfigured one, so a mismatch restarts instead.
- Killed on shutdown. The child is detached so its whole plugin tree can be
  signalled at once, which also means it would outlive Codeman and hold its
  port against the next start - the exact EADDRINUSE this feature already got
  wrong once.
- Boot output captured and returned. With no shell tab there is nowhere else
  for a stack trace to land, so a failed spawn reports its own tail.

The endpoint is fenced at the same bar as the profile installer and for the
same reason: booting a dsh profile executes the plugin code in it, so this is a
privileged action even though it reads as "open a page". `authority` comes from
the client (`location.host`) because only the browser knows which origin is in
play, and it is regex-confined at the schema boundary - defence in depth behind
the argv-array spawn, admitting host:port in the shapes a browser authority can
take and nothing readable as a second argument.

`GET /api/deepseek/web-port` is gone; port selection moved into the supervisor,
which is the thing that knows whether a server is already running. The two
client-side probe helpers went with it, since the server now owns the wait.

Verified over the tailnet authority end to end: no session is created (session
count unchanged, one tab), the server runs on 3081 beside the user's own dsh
web on 3080, status reports the tailnet authority, and the proxied dashboard
renders with zero 4xx. Full gate green (6148 passed, +6).
2026-08-25 03:08:15 +02:00
Codeman maintainer 15ae5f5d81 fix(deepseek): make the web-UI shortcut pick a free port, verify it, and trust its frame
The `Run > DeepSeek web UI...` shortcut failed three ways at once against a real
install, and the three are independent.

1. It hardcoded `--port 3080`. That is dsh web's OWN default, which makes it
   precisely the port a DeepSeek user is most likely to be serving on already,
   so the launch died with EADDRINUSE against the user's own server. The port
   now comes from `GET /api/deepseek/web-port`, which walks 3080..3119 for a
   free loopback port by BINDING it (a connect probe cannot tell "free" from
   "listening but not answering yet").

2. It opened the tab unconditionally. The crashed server left a saved dashboard
   pointing at nothing, with the failure only visible in a shell tab nobody had
   a reason to look at. The launch now polls the existing webview probe until
   the URL answers, and on timeout reports the error naming the shell tab
   instead of persisting a dead dashboard.

3. The saved tab was untrusted, so the frame was sandboxed without
   `allow-same-origin` and the dashboard was broken twice over: the dsh
   client-runtime reads `localStorage` while loading its plugins and died there
   ("the document is sandboxed and lacks the 'allow-same-origin' flag"), and an
   opaque-origin frame sends `Origin: null`, so dsh's own trust fence 403'd
   every `/api` call no matter which authority `--trusted-host` named. Passing
   `location.host` only means anything once the frame actually carries that
   origin, so `--trusted-host` had never once done its job. The managed tab is
   now created `trusted: true`.

   That trade is real and deliberate: a trusted proxied frame is same-origin
   with Codeman and can reach Codeman's API. It is defensible only because this
   dashboard is an agent harness Codeman just started itself, on loopback, which
   can already run code as the user. It is not a precedent for trusting
   third-party dashboards, which is why it is set at this one call site rather
   than defaulted.

Separately, the shortcut listed its own dashboard twice: once as the menu entry
that starts it and once as the row that entry had written on the previous click.
Webviews now carry an optional `managed` marker, managed rows are filtered out
of the saved-dashboard list, and a relaunch repoints the existing row rather
than stacking one dead dashboard per restart (which the per-launch port would
otherwise guarantee). `managed` is declared in the schema because a plain
`z.object` strips undeclared keys, so an undeclared marker would never survive
the round trip.

`DEEPSEEK_WEB_PORT` is gone from constants.js; its doc comment asserted that a
hand-started `dsh web` and the shortcut "land on the same place and share one
saved tab", which is the bug stated as a feature.

Verified on a real install with the user's own `dsh web` holding 3080: the
shortcut takes 3081, the server answers, exactly one DeepSeek entry shows in the
run menu, and the proxied dashboard renders its workspaces and completes its own
API calls (the previously-403'd `api/settings.describe` now succeeds). Full gate
green (6142 passed), typecheck/lint/format/public-assets clean.
2026-08-25 02:39:57 +02:00
Codeman maintainer 14de2b7012 fix(tabs): keep the created stamp reachable on a tight rail, and do not skip the first render
Two review nits on the vertical rail's detailed rows.

1. The tab-rail-tight rule (below 288px) hides `.tab-meta-created`, and its
   comment claimed the value "survives in the row's title attribute either way".
   It did not: the only title carrying it lived ON that element, and a
   `display: none` element has no hover target, so the created stamp was not
   shrunk but gone with no way to ask for it. Rather than just correcting the
   comment, `_sidebarRichMetaHTML()` now puts BOTH absolute stamps on the
   `.tab-meta` line itself, so the pill and the gaps around the stamps remain as
   hover targets. An item's own title still wins where the item is visible.

2. applyTabOrientation() decided whether applyTabWrapSettings() had already
   re-rendered by comparing `_tallTabsEnabled` before and after. That reads an
   UNDEFINED previous value as "it rendered", but applyTabWrapSettings()
   deliberately renders nothing on its first call ever (it only establishes the
   baseline: `prevTallTabs !== undefined && prevTallTabs !== showFolder`). So on
   a first call that also flips the folder row, neither function rendered and the
   rows stayed stale. Reachable when the pre-paint script throws and leaves the
   layout attributes on their catch-branch fallbacks for applyTabOrientation() to
   correct. The guard now mirrors applyTabWrapSettings()'s own condition.

Both new tests were run against the unfixed code first and fail there, which is
the only thing that makes them regression tests. (The third, "does not render
twice", passes either way by design: it pins that fix 2 did not introduce a
double rebuild.)

Verified in a real browser against a live server with two sessions, driving the
narrowing through _setTabRailWidth() the way the resize drag does: at the 320
default the row reads "CREATED 2m ago · IDLE <1m" with the created element
displayed; at 256 the tight class is on, the created element computes to
display:none, the visible text drops to "IDLE <1m", and the meta line's title
still reads "First created: ...". At 220 the compact threshold drops rich rows
entirely. Screenshots confirm no truncation artifacts in either state.

Full gate green (6104 passed), typecheck, lint, format, frontend-syntax and
public-assets all clean.
2026-08-24 22:58:28 +02:00
Codeman maintainer cdceede33d fix(deepseek): atomic shim write, honest attribution comment, name-fallback profile classifier
The three smaller review nits, plus the first real test coverage for the status
shim (it had none: it is emitted as a STRING, so tsc never sees it).

1. The shim was written with a plain writeFileSync. The TUI can be exec'ing that
   exact path while an upgraded Codeman refreshes it, and a reader catching a
   half-written file gets a syntax error, exits non-zero, and is retried four
   times per state change for a file that will never parse. Now temp + rename
   (atomic within the directory), with the temp chmod'ed before the rename since
   writeFileSync's mode only applies on create, and removed if the write throws.
   SHIM_VERSION bumped to 2, because SHIM_SOURCE changed and an existing v1 shim
   would otherwise keep matching the embedded marker and never be refreshed.

2. The pane-id comment claimed the ambient env "cannot be spoofed by an argument
   the agent itself could influence". The agent runs IN that pane and can invoke
   the shim with CODEMAN_SESSION_ID unset and any argv it likes. It buys nothing
   it did not already have (the hook-secret file is readable from the same pane,
   so it can POST /api/hook-event directly), but the comment read like a security
   boundary. Rewritten to say what the preference actually buys: correct
   attribution when a TUI mangles or re-uses the pane argument. Accidents, not
   adversaries.

3. classifyProfile() folded the directory name into the same haystack as the
   bundles, but only the TUI arm could match a bare name, so a stock profile
   whose package.json has no dsh.profile.bundles (hand-edited, older layout,
   mid-install) classified as `unknown` -> launchable -> eligible as the DEFAULT
   pick, which is exactly the pane-dies-on-arrival failure the two-part
   availability gate exists to prevent. The stock names are now a LAST-resort
   fallback consulted after the bundle patterns, so real bundle evidence still
   wins over a name the user chose. The loose `tui` arm gained word boundaries:
   it decides which profile boots by default, and matching the middle of
   `intuition` is not a rule anyone could predict.

New test/deepseek-status-shim.test.ts runs the generated script the way the
harness does -- real node process, real argv, real env, real listener -- and
covers the exit-code contract that makes the retry behaviour safe: mapped states
post and exit 0, an unknown verb or unmapped state exits 0 WITHOUT posting (a
non-zero there would be four HTTP requests per state change forever), a rejecting
server or an unreachable one exits non-zero so the caller retries, the hook secret
is read at execution time, and `node --check` parses the file (a template-literal
typo in SHIM_SOURCE is invisible to tsc).

Trap worth recording, hit while writing it: the tests must spawn the shim
ASYNCHRONOUSLY. The listener lives in the test process, so spawnSync blocks the
event loop that has to accept the connection, the shim waits out its own 1500ms
socket timeout and exits 1, and it reads exactly like a broken shim (measured:
Socket._onTimeout in its --trace-exit output, server logging nothing).

Verified: full gate green (6142 passed, +10), typecheck/lint/format clean.
2026-08-24 18:00:52 +02:00
Codeman maintainer 2034719d61 fix(deepseek): close the env-var clamp hole, bound the profile install, make the hook gate per-session
Three review findings on the DeepSeek Harness mode, plus one the third exposed.

1. The multi-user clamp was bypassable by a sibling field on the same request.
   clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but
   DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER
   _configureDeepSeek(), so a non-granted owner sending
   envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated
   instance: a session created with permissionMode "read-only" and that override
   ran with DSH_PERMISSION_MODE=danger-full-access in its pane.

   Every other CLI's bypass is a command-line flag reachable only through the
   per-CLI config, which is why the config clamp alone is the whole gate for
   them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted
   owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to
   what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on
   that list because it aims the launcher at a profile tree whose plugin code
   runs at boot, before any approval row can apply. Verified end to end in real
   multi-user mode: a non-granted user sending both now gets workspace-write and
   no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched.

2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout`
   signals only the direct child, and a plugin install fans out into
   package-manager children that keep the inherited stdio pipes open, so `close`
   never fires and the held-open request leaks with no route-level deadline.
   Reproduced: with a 1.5s built-in timeout the promise was still unsettled after
   6s and both fan-out children were alive. Now detached: true plus negative-pid
   SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a
   last-resort reap for a grandchild that escaped the group. Same probe after the
   change: close fires, direct child and both grandchildren dead.

3. hooksAvailableForMode() promised more than a dsh session can deliver.
   deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that
   triple is the only reason a dsh session posts hook events, so `until=stop` was
   accepted and then blocked for the caller's whole timeout: the exact
   infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now
   takes HookCapabilityOptions and every call site passes sessionHookOptions(),
   with the deepseek arm reading `!== false` so a forgotten one degrades to the
   old behaviour. The refusal names the setting rather than saying "no Claude
   Code hooks", which would send the caller hunting a bug that is really a
   setting they chose. Profile conformance stays unknowable at request time and
   is documented as such. The stale "True for `claude` and nothing else" docblock
   is corrected.

4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a
   claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent
   capture read Claude's own transcript, and adding deepseek silently widened
   both to a mode that has none. They compare mode === 'claude' directly now, and
   a static check pins them there.

Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the
wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 --
bridge off plus explicit until=stop is a 400 naming the setting, bridge off with
no `until` still 200s on idle/exit, bridge on accepts stop.
2026-08-24 16:01:02 +02:00
d fei 7c62b16e5f fix(terminal): bound live xterm backpressure 2026-08-24 19:06:50 +08:00
Aamer Akhter d15d979a33 fix: COD-341 correct changeset package name 2026-08-23 22:55:10 -04:00
Aamer Akhter 858b15e3f5 chore: COD-341 add File Viewer search release note 2026-08-23 22:46:58 -04:00
Codeman maintainer b330f1d9e8 feat(tabs): give the vertical rail the home screen's per-session detail
The vertical tab rail (tabOrientation 'vertical') listed names and nothing
else, while the rich sidebar and both home screens already answered the
question a docked column exists to answer: which of these sessions wants me
next, and how long has it been like that. The rail is a docked column too, so
it now draws the same row.

- New per-device setting tabRailDetail ('rich' | 'simple', default rich),
  App Settings -> Appearance -> Tabs, in SettingsUpdateSchema + displayKeys and
  stamped as data-tab-rail-detail by the pre-paint script, so a detailed rail
  does not flash through simple rows on every load.
- ONE gate for both vertical surfaces: isRichTabRows() =
  isSessionSidebarRich() || isTabRailRich(). The row model, the markup and the
  20s in-place clock are the existing rich-sidebar ones, classified by
  _mobileOverviewState/_mobileOverviewSince, so the rail, the sidebar, the
  desktop home rail and the phone overview cannot disagree about what
  "working" means or which stamp measures it.
- Detail rides on its OWN attribute, exactly as the sidebar's does, so every
  existing [data-tab-orientation='vertical'] rule keeps matching both variants
  untouched. A flip of detail ALONE still forces a full render (the stamps line
  is emitted by the row template, not toggled by CSS) and re-runs
  applyTabWrapSettings(), which owns the folder line and is now rail-aware.
- CSS: every rich paint rule gains a rail twin as a COMMA-GROUPED selector,
  never :is() - an :is() list takes its most specific argument, which would
  lift the sidebar arm from (0,3,1) to the rail's (0,5,1) and let these rules
  outrank things they never used to.
- Width is why there are thresholds. At 256px the stamps line ellipsizes
  mid-word, the same reason the rich sidebar is 300px, so a rail that has never
  been sized defaults to 320 (RICH_DEFAULT_WIDTH, the existing Wide preset,
  which also keeps the settings select on a named choice). A width the user has
  chosen is never overridden: below 288px the created stamp is dropped rather
  than truncated (tab-rail-tight, CSS only) and below 240px the rows go back to
  simple (tab-rail-compact, which re-renders).
- The rich clock is armed and disarmed by applyTabOrientation() as well as
  applySessionListLayout(); a leaked interval would rewrite stamps in a list
  that no longer has any.

Also fixes a data-loss bug in the inline tab rename that predates the rail and
reproduces in every layout, header strip included: Escape set the input to ''
and blurred it, and the blur handler commits - so cancelling a rename PUT an
empty name, and the tab fell back to its folder label (measured against a live
server: ["rail-alpha","","rail-gamma"]). Escape now calls cancelRename(), which
invalidates the edit so the blur that follows the input's removal is a no-op.

Tests: rail-detail gate, the three ways it turns back off (simple, compact,
horizontal), sidebar-wins, render-on-detail-flip and the plumbing/CSS guards in
test/session-list-layout.test.ts; the rename cancel in test/inline-rename.test.ts
(browser suite), pinned by running it against the old code first. Verified live
against a real server on an isolated instance: detailed/simple/compact/header/
sidebar variants, click-select, the ... menu, inline rename, Alt+N, the in-place
stamp tick and a full settings-picker round-trip including reload.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 04:45:02 +02:00
Aamer Akhter c14171b534 fix(file-viewer): COD-341 finalize deferred navigation 2026-08-23 22:42:10 -04:00
Aamer Akhter acd9ffedc8 fix(file-viewer): COD-341 complete search transitions 2026-08-23 22:27:26 -04:00
Aamer Akhter 921933775b test(file-viewer): COD-341 execute session lifecycle path 2026-08-23 22:08:29 -04:00
Aamer Akhter f6a1f06633 fix(file-viewer): COD-341 synchronize session lifecycle 2026-08-23 21:58:45 -04:00
Aamer Akhter dab8e6643c fix(file-viewer): COD-341 deduplicate normal tree loads 2026-08-23 21:40:10 -04:00
Codeman maintainer 4cda150493 feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.

DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:

1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
   $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
   `base` -- the interactive terminal front door is always a third-party
   plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
   and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
   button gates on the latter, because reporting only the binary would spawn a
   pane that dies on arrival. When the binary is present but no profile is,
   the run menu offers to install one (POST /api/deepseek/install-profile).

2. The permission switch is an env var, not a flag. The harness has no
   command-line permission option; its sandbox/approval rows read
   DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
   Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
   own workspace-write, which still asks, so the multi-user clamp is the
   only-if-sent branch and clamps to workspace-write, never read-only.

3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
   earned that. The terminal front door reports idle/working/blocked to a
   supervising process over a generic env-gated contract; a generated shim
   (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
   report to /api/hook-event as stop / agent_working / permission_prompt. So a
   dsh session gets definitive respawn triggers, real wait-endpoint signals and
   real Approvals Inbox items instead of output-stabilization guesswork.
   `agent_working` is new (157th SSE constant) and joins
   APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
   alert at once.

The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.

Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.

Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.

Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 03:37:56 +02:00
Aamer Akhter 3af36f7c34 fix(file-viewer): COD-341 preserve search row layout 2026-08-23 21:23:17 -04:00
Aamer Akhter 49797e37dd fix(file-viewer): COD-341 gate stale search results 2026-08-23 21:09:52 -04:00
Aamer Akhter c614331d60 feat(file-viewer): COD-341 add server-side search 2026-08-23 21:01:39 -04:00
Codeman maintainer 9cfd8e8989 fix(docker): survive xAI installer's own /usr/local/bin/grok symlink
The agent-image grok step copied /root/.grok/bin/grok onto /usr/local/bin/grok
with cp -L. Newer versions of xAI's install.sh already create
/usr/local/bin/grok as a symlink to that same binary, so the copy failed with
'same file' and the --no-cache rebuild died at the grok layer (2026-08-24).
Stage the copy under a temp name, drop whatever the installer left at the
destination, then move into place - correct against both old and new
installers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 01:12:38 +02:00
Codeman maintainer 8fe393826b chore: version packages 2026-08-24 01:00:51 +02:00
Codeman maintainer 7a340fe7bc fix(tabs): review-driven hardening for the tab-layout foundation and vertical rail
Post-merge follow-ups from the deep review of #334 and #335, so they ship in
the same release as the features.

Tab-layout foundation (#335):
- PUT /api/session-order drops unknown/foreign ids again instead of 400ing
  the whole write, in both the owner and the admin path (single-user requests
  are the synthetic admin, so that path is the one the browser hits). The
  frontend debounces its reorder push and swallows errors, so a session
  deleted inside the debounce window silently cost the user the entire
  reorder - and the endpoint sits on the stable /api/v1 surface, where the
  pre-layout server merged leniently.
- A failed mux restore no longer locks explicit deletions into 500s for the
  process lifetime: runSessionDeletion and webviewDeleted degrade to
  best-effort without layout coordination, while the automated stale sweep
  (runStaleSessionCleanup) stays fail-closed.
- sse-events doc comment: no 'suppressed' hook event exists; hooks stay 8.
- registerSessionWithLayout resolves its owner through ownerLayoutKey()
  instead of a hardcoded '@single'.

Vertical rail (#334) - all rail-awareness gaps in sidebar-only predicates,
unified behind the new _isVerticalTabList() (sidebar OR rail):
- Drag-reorder read the insertion side from clientX in the rail, so
  before/after was effectively arbitrary on vertical rows; the drag-over
  indicators now draw as top/bottom edges there like the sidebar's.
- The active tab is scrolled into view in the rail (Alt+N/palette selection
  used to leave the row below the fold).
- Floating subagent/ultracode windows anchor to the RIGHT of rail tabs, and
  the connector redraw gates (render tail + strip scroll) cover the rail.
- Server-seeded tabOrientation is applied when the async settings load
  resolves, not only at boot, so a fresh device shows the rail immediately.
- The pre-paint script stamps data-tab-orientation and --tab-rail-width
  (sidebar-wins and solo carve-outs included), removing the flash of the
  header strip on every vertical-mode load.
- The session name font defaults to 12px, the sidebar's historical 0.75rem
  size, so installs that never touch the new slider are not restyled.

Also documents the rail in CLAUDE.md (second #sessionTabs host, mover
ordering, the axis-predicate rule) and gives tab-rail-resize.js its
@dependency/@loadorder header. Full gate green (6093 tests); the excluded
browser suite was run by hand - only the known environmental failures
(opencode/codex binaries) remain, identical to pristine master.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 00:59:48 +02:00
Codeman maintainer f3c615b669 fix(files): give the file preview a working detach button
The button next to the file preview's close icon was Copy Content, whose
overlapping-pages glyph reads as a pop-out control - and for a PDF or any
media/binary preview it was completely dead: those branches never fill
filePreviewContent, so the click hit an empty-content guard and did nothing,
with no feedback.

There is now a real detach button that opens the previewed file in a browser
tab (raw route for PDFs/images/media/text, the server-converted PDF preview
for docx/pptx), severs window.opener by hand so a blocked pop-up stays
detectable, closes the overlay on success (which also stops any playing
media), and disarms on close so it can never open a stale file. The copy
button now toasts 'Nothing to copy in this preview' instead of staying
silent.

Verified live with Playwright against an isolated instance: button visible
and armed on a PDF preview, file-raw answers 200, clicking opens the URL and
tears the overlay down, text previews keep a working copy buffer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 00:59:27 +02:00
Ark0N 82f81d21c4 Merge pull request #333 from Ark0N/feat/grok-mode
feat(grok): add Grok Build (xAI) as a seventh CLI run mode
2026-08-24 00:43:04 +02:00
Codeman maintainer c173ae0264 Merge remote-tracking branch 'origin/master' into worktree-grok-mode
# Conflicts:
#	src/web/public/app.js
2026-08-24 00:32:25 +02:00
Ark0N e9dd55e5fd Merge pull request #334 from aakhter/pr/cod-358-vertical-rail
feat(tabs): add a resizable vertical session rail
2026-08-24 00:14:10 +02:00
Ark0N dd96f252ea Merge pull request #335 from aakhter/pr/cod-359-tab-layout
feat(tabs): add owner-scoped tab layout foundation
2026-08-24 00:12:20 +02:00
Aamer Akhter 74194e4fc0 feat(tabs): COD-359 add owner-scoped tab layouts 2026-08-23 14:46:10 -04:00
Aamer Akhter c45c6c3846 merge upstream master into COD-358 2026-08-23 14:15:07 -04:00
Aamer Akhter 17b141dc25 test(workflows): keep recent-run fixture clock-independent 2026-08-23 14:12:38 -04:00
Codeman maintainer 57f326ab8f docs(grok): fix the mode counts and line refs the seventh mode invalidated
The agent skill's endpoints.md is what other agents read as ground truth, and
three of its facts went stale when grok landed:

- `/api/v1/grok/status` was added to the probe list, but the sentence after it
  still said only Pi's response carries `.data.version`. Grok's carries it for
  the same reason (a squatted binary name), and an agent that trusts the old
  wording has no way to tell a misresolved grok from an absent one.
- the `active-tools` bullet listed grok among the modes it stays empty for, then
  claimed in the same breath that `isExternalCliMode` "lists only those five".
- its three source line refs had all drifted: `isExternalCliMode` is now
  session.ts:174-183 (it was already wrong before this branch), the external-CLI
  early return is session.ts:2261, and TEXT_COMMAND_PATTERN is
  bash-tool-parser.ts:89.

CLAUDE.md and architecture-invariants.md counted modes in their Docker-cases and
Web-tabs paragraphs ("any of the five CLI backends", "never a sixth
SessionMode"). Both numbers were already stale before grok (antigravity and pi
had made it seven) and grok is now in the agent image, so the counts are gone
rather than incremented: the invariant those sentences carry is that Docker and
web tabs are not modes at all, which no number has ever helped state. The two
plan docs keep their original wording, being historical design records.
2026-08-23 19:57:13 +02:00
Codeman maintainer 6b0b6d10ad fix(test): anchor the workflow-run fixture to now instead of a pinned epoch
test/workflow-run-watcher.test.ts pinned its fixture's newest activity at
2026-06-14T20:06:40Z and then asked getRecentRunSummaries(100000) to return
it. That argument is MINUTES, so the window is 69.4 days: the assertion
expired at 2026-08-23T06:46:40Z and the file has failed on every branch
since, on a suite nobody had touched. The last green CI run finished at
06:47:43Z, about a minute inside the boundary, which is why it landed as a
surprise rather than a bisectable regression.

The fixture epochs now hang off a RUN_ANCHOR of Date.now() - 601s with every
offset preserved verbatim, so the parsed durations, the ordering and the
live-vs-done discriminators are all unchanged, and the recency filter is
still the thing under test. It just cannot rot again.
2026-08-23 19:54:15 +02:00
Aamer Akhter c9ea8bbac5 fix(tabs): COD-358 re-query tab after rename cancel 2026-08-23 13:40:42 -04:00
Aamer Akhter 1795a138b3 test(workflows): document clock-independent fixture fix 2026-08-23 12:47:39 -04:00
Aamer Akhter e3a2fb767f feat(tabs): COD-358 add resizable vertical session rail 2026-08-23 12:21:02 -04:00
Codeman maintainer 3f8c8e99d1 feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.

Grok mixes two existing shapes and the wiring follows from that:

- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
  (--always-approve, grok's bypassPermissions mode; config-level deny rules
  still apply on top). The Run button sends it true, like runAntigravity(),
  and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
  a bare grok spawn is grok's own ask-mode default, which is already safe,
  so only a sent config needs the flag forced off. Cron needs nothing for
  the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
  with mouse support, so it stays OUT of isAltScreenStripMode() and lands
  on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
  (unmeasured against an authenticated composer; documented fallback is the
  'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
  also installs a grok bin), so grok-cli-resolver.ts version-probes every
  candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
  GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
  shared with the dependency registry so doctor and run mode cannot drift.

Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.

Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.

Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).

Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 08:39:03 +02:00
Codeman maintainer 88bb98de43 chore: version packages 2026-08-22 14:45:08 +02:00
Codeman maintainer 687e9d7565 feat!: retire the sc tmux chooser in favour of codeman tui
scripts/tmux-chooser.sh is deleted. codeman tui replaces it and does the
job better: sc numbered its entries globally but only accepted a single
[1-9] keypress, so sessions 10+ were listed and unselectable, and it
inferred nothing about what an agent was doing. The tui carries the
server's real states, answers permission dialogs, and leaves an attach
with one key.

install.sh no longer creates the tmux-chooser symlink or the sc alias.
It now sweeps both up instead, on update AND uninstall, so an update
cannot leave a symlink pointing at a script this version stopped
shipping. The alias removal is marker-owned: it matches the exact line
the installer wrote, so someone's own 'alias sc=' for another tool is
never touched, and it rewrites through 'cat >' so the profile keeps its
mode and ownership. Verified against three profile shapes.

BREAKING CHANGE: the 'sc' command and the 'tmux-chooser' symlink are
gone. Use 'codeman tui' (and 'codeman tui --list' / 'codeman tui <n>').
2026-08-22 14:44:14 +02:00
Codeman maintainer 5d81cc01ca docs: point terminal users at codeman tui instead of the sc chooser
codeman tui supersedes the sc bash chooser: it reaches sessions 10+,
carries the server's real states instead of a static list, and leaves an
attach with one key. Every place that told a user to run sc now names the
tui equivalent, including the two wiki pages and install.sh's next-steps
banner. Both wiki pages also carried the wrong detach chord (Ctrl+A D;
the socket's prefix is C-b), which the tui makes moot.

Source comments that explained themselves as "the sc -l replacement" now
just say what they do. docs/tui-plan.md and CHANGELOG.md are historical
records and keep their references.
2026-08-22 14:36:23 +02:00
Ark0N f49249fb2f Merge pull request #312 from Ark0N/feat/tui
feat: codeman tui, a terminal dashboard with live agent states
2026-08-22 14:32:51 +02:00
Codeman maintainer bc3f9f8a37 docs: record the socket-resolution rule where the mechanism lives
CLAUDE.md gained "any new `tmux -L` caller through `resolveTmuxSocketName()`"
but architecture-invariants.md, which that bullet points at for the
mechanism, still described only the dataPath() half. Say why the rule
exists there too: the TUI is the first non-server process to shell out
to tmux.
2026-08-22 14:25:23 +02:00
Codeman maintainer 97acfc61c3 docs: stop telling sc users to press a key only the TUI binds
The sc chooser runs a plain `tmux attach-session` and binds nothing, so
F1 does not detach from it; only `codeman tui`'s attach claims that key,
and only for its own duration. The line it replaced was wrong too (the
socket's prefix is C-b, not C-a), so name the real chord and say which
command gives you the single key instead.
2026-08-22 14:19:31 +02:00
Codeman maintainer c27363459d docs(tui): stop telling people to press a key that does not work
The guide and the README both said to detach with `Ctrl+B D`. Beta testing
proved that wrong twice over: tmux binds lowercase `d` to `detach-client` and
capital `D` to `choose-client`, and even the correct letter fails for anyone who
keeps Ctrl held, because that sends `Ctrl+D`, which tmux leaves unbound. A
tester followed the documented instruction, stayed attached, and exited the
agent to escape.

Both now say `F1`, and the attach section describes what actually happens: the
session strip across the top of the pane, `Alt+1`..`Alt+9` switching without
returning to the dashboard, and `r` to resume a session whose pane has died.
Also corrected: `1-9` switches rather than jump-attaches, `x` confirms with `y`
rather than a typed name, and a new session opens straight into its pane.

`docs/tui-plan.md` is deliberately untouched — it is the design record of what
was planned, not a description of what shipped.
2026-08-22 14:13:58 +02:00
Codeman maintainer 737c2ed7f8 fix(tui): escape the separator in the switch binding, closing the sizing leak
The loose end from 777f974, now explained. Sessions came back from a detach on
`window-size latest` instead of `manual`, and the restore primitive round-tripped
correctly in isolation, so the corruption had to be upstream of it. It was: the
snapshot was taken from state this code had already broken.

`bindSwitchKey` passed a bare `;` between the two commands it wanted in one
binding. That is a command separator to tmux's OWN parser, not an argument: it
ended the `bind-key` and executed what followed immediately. So the binding kept
only `switch-client`, and `set-window-option ... window-size latest` RAN against
every switchable session at attach time — before the sizing snapshot was taken.
Every session was therefore snapshotted as `latest` and faithfully restored to
`latest`.

Proven against real tmux both ways before fixing: a bare `;` leaves the session
on `latest` and stores a one-command binding, while `\;` leaves it `manual` and
stores both commands.

Verified end to end: 7 sessions manual before, 1 latest + 6 manual during the
attach (the attached one follows the terminal, the rest are pre-sized), no dot
padding on a switch, and all 7 back to 120x40 manual after the detach.

This also means the "follow the terminal after switching" half of 777f974 never
actually worked — it was never in the binding.
2026-08-22 14:13:58 +02:00
Codeman maintainer bb24d2c256 fix(tui): stop the preview stacking every repaint of a session
The overview showed the same session twice, one frame above another, after
switching sessions (reported from the beta with a screenshot).

Claude repaints by ABSOLUTE CURSOR POSITIONING, not by clearing: a 198KB pane
tail carries 1142 `CSI r;c H` and exactly one `CSI 2J`. The replay honoured the
COLUMN of those sequences and ignored the ROW, so a repaint could never
overwrite what came before and was appended instead. That same tail replayed as
FIFTY stacked copies of one frame. The preview shows the last N lines, so on a
short terminal you saw the newest frame by luck and on a tall one you saw the
end of the previous frame above it.

A cursor HOME now starts the buffer over. That is not a heuristic but the
line-based equivalent of what a home means: a full-screen app announcing it is
repainting from the top, with everything on screen about to be overwritten in
place. Only row 1 column 1 counts — any other address is a write position
inside the frame being painted, and resetting on those would erase live
content.

Measured on the real tail that produced the screenshot: 198599 bytes and 50
copies of the welcome frame collapse to 40 lines carrying exactly one.

The old test pinned the append behaviour, including a spurious leading empty
line that the initial CUP produced; both are gone.
2026-08-22 14:13:58 +02:00
Codeman maintainer bc1821661f fix(tui): say alt+1-9 on the bar, and stop the dot grid when switching
The bar now reads "alt+1-9 switch · F1 back to the codeman dashboard", so the
switch keys are discoverable instead of secret. Shown only when those keys were
actually claimed, the same rule the way-out key follows: a bar naming a key
that does nothing is the bug this series started with.

THE DOT GRID. Switching landed in a pane occupying part of the terminal with
tmux's dot fill everywhere else. It was never a size mismatch — the window was
already the right size. `window-size latest` only resizes a window while a
client is ON it, and the sessions behind the tab strip have none until you
switch, so the resize happened AT the switch: tmux painted the newly-available
area with dots and an idle claude had no reason to redraw into it. Every
switchable session is now pre-sized to the attaching terminal, which moves that
repaint to attach time while the user is still looking at the first session,
and the switch binding restores `window-size latest` on arrival so a mid-attach
terminal resize still follows. Measured: 14 consecutive switches across 7
sessions, zero dot-padded rows, against 1-in-6 before.

⚠️ Known loose end, deliberately not papered over: after a detach the window
SIZE is restored exactly but the window-size MODE can come back as `latest`
rather than `manual`. The restore primitive round-trips correctly in isolation
(manual -> presize -> latest -> restore = manual) and no call site in the TUI
or the server sets `latest` afterwards, so the cause is not yet identified. The
practical effect is nil: the remaining client keeps the window at its own size
and Codeman re-pins `manual` on the browser's next resize.
2026-08-22 14:13:58 +02:00
Codeman maintainer 74c9879359 fix(tui): size every switchable session, not just the one being attached
Switching with Alt+N landed in a pane that filled part of the terminal with
tmux padding the rest as a dot grid — reported from the beta with a screenshot
showing the pane in the left half and dots everywhere else.

Codeman pins every window `window-size manual` at the BROWSER's size
(tmux-manager.ts), so no attaching client can resize it. The attach already
lifted that for the session it opened, which is why a plain attach looked
right; `switch-client` then moved the user into a session that had never been
lifted, and the old pin reasserted itself. `window-size latest` now goes on
every session the strip can reach, alongside the bar those sessions already
get, and each one's original sizing is snapshotted and restored on detach.

Verified by round-tripping a session pinned at 120x40 manual: latest 190x49
while attached, back to 120x40 manual after, with no dot rows at either step
and the bar intact at full width after a switch.
2026-08-22 14:13:58 +02:00
Codeman maintainer 499d3d6e4d fix(tui): finish the 1-9 rename in the fallback footer
The renderer's own FOOTER_KEYS table still said 'jump'. It is only reached when
the app layer supplies no footerKeys, so nothing visible was wrong, but a
fallback that contradicts the live footer is exactly the kind of drift that
turns into a bug report later.
2026-08-22 14:13:58 +02:00
Codeman maintainer 9b29666e03 fix(tui): keep the way out on the bar, and make Alt+1..9 actually switch
Four faults, all reported at once, and three of them were mine from the last
two commits.

THE HINT VANISHED. Two independent causes. First, a leaked F1 binding: an
attach whose TUI was killed leaves `F1 -> detach-client` in tmux's root table,
and the claim treated "already bound" as someone else's key, so every later
attach fell back to advertising the tmux chord — the bar stopped saying F1
while F1 still worked. A key already bound to `detach-client` now counts as
ours. Second, width: tmux truncates a status line that overflows and drops the
RIGHT-aligned segment, which is the hint. The strip now gets a budget measured
from the terminal's width minus the hint, and it drops tabs from the far end
until it fits. ⚠️ Measured on VISIBLE columns, not format bytes: `#[reverse]`
costs zero columns, and counting it made a strip that "fitted" still truncate
the hint at 80, 100, 120 and 176 columns on a real terminal.

ALT+N DID NOT SWITCH. On the dashboard, a bare digit meant jump AND ATTACH, and
a terminal sends Alt+N as ESC then N: when those land in separate reads —
routine over SSH — the chord decodes as Escape plus a bare digit, so "switch to
tab 2" threw the user into tab 2's pane. A digit now SELECTS, matching what
Alt+N means in the web UI; Enter is how you go in. Inside a pane the keys never
reached the TUI at all, since tmux owns the terminal, so the attach now binds
Alt+1..9 in tmux's root table to `switch-client` — the strip is usable rather
than decorative. ⚠️ The bar is applied to every session the strip can reach,
each highlighting its own tab: with it on the attached session only, switching
landed the user in a pane with no strip and no way out on screen.

⚠️ The leaked-state sweep was missing `status-position`, so it removed the
marker and left the position behind — and with no marker the leftover no longer
matched, making it permanently unsweepable. Found by diffing every session's
options after a detach.
2026-08-22 14:13:58 +02:00
Codeman maintainer aec6516638 feat(tui): keep the session tabs visible inside a pane, and move the way out to F1
Attaching made every other session disappear: the dashboard is gone, tmux owns
the terminal, and there is nothing left saying what else is running. The attach
bar now carries the session strip, numbered exactly as the dashboard numbers
them, with the session you are in inverted, and it sits at the TOP of the pane
where the web UI keeps its tabs.

The strip is a WINDOW around the active tab, not the whole list, with ellipses
marking each end that is actually cut. The bar is one line shared with the way
out, and that hint is the only instruction a user gets while tmux has the
terminal, so it must never be crowded off; a test drives 20 long-named sessions
through the bar and asserts it survives.

⚠️ The strip is a snapshot taken at attach time and never refreshed. The TUI is
blocked in `spawnSync` for the whole attach so there is no loop to update from,
and tmux's own format language cannot map a `codeman-<hex>` session name back
to a label a human recognises. Slightly stale beats absent.

The way out moves from F12 to F1, which sits beside Esc where a hand backing
out already goes. Verified against BOTH encodings a terminal sends for it:
xterm's SS3 (ESC O P) and PuTTY's default (ESC [ 1 1 ~).

`status-position` joins the snapshot, so a session that had its bar at the
bottom gets it back there on detach along with everything else.
2026-08-22 14:13:58 +02:00
Codeman maintainer c59f006bb6 fix(tui): stop drawing from the unicode blocks a plain terminal font lacks
Three separate "why are there boxes" reports, and I fixed them one glyph at a
time instead of as a class, so the next one was always waiting. Grouping the
tester's terminal by unicode block made the rule obvious:

  RENDERS   Latin-1 (·), Box Drawing (─ │), Block Elements (█ ▛ ▐),
            Geometric Shapes (○ ▶), General Punctuation (…), Arrows
  TOFU      Miscellaneous Technical (⏎ U+23CE, ⏵ U+23F5), the sparse end
            of Dingbats (❯ U+276F)

That is an ordinary font, not a broken one, so it is the profile to design
against. The working spinner moves off Dingbats and Math Operators onto
quadrant blocks (▖▘▝▗) — the same block as the `▛█▐` art claude itself draws,
which that font renders fine — and the blocked marker moves off `⚠`
(Misc Symbols, emoji presentation on many terminals) onto `▲`, the block that
already gives us `▶` and `○`.

The preview fold gains claude's own spinner dingbats (✢ ✳ ∗ ✻ ✽ ✴ → `*`) and
`⚠` → `!`. Its animated status line is exactly where a reader looks, so tofu
there is the most visible kind there is.

A test now enforces this as a CLASS: no glyph in the unicode set may come from
Misc Technical, Misc Symbols or Dingbats, with U+2714 the single documented
exception because it was observed rendering on the very font that failed the
others. Verified by scanning a live frame driven with the tester's exact
environment: zero glyphs from any of the three blocks.
2026-08-22 14:13:58 +02:00
Codeman maintainer b2ac6c1bd9 fix(tui): confirm a kill with y, and make the dialog say what it would kill
Killing demanded the session's NAME typed out in full. That is the right
ceremony for dropping a production database and the wrong one for closing a
pane you are looking at; the beta tester's verdict was "thats stupid, just make
me type Y to confirm". `x` then `y` is already two deliberate keystrokes on a
row the user selected, and the conversation lives in its transcript, which a
kill does not touch.

Everything that is not `y` CANCELS rather than being ignored, so a stray key
closes the dialog instead of leaving a destructive prompt armed and waiting for
whatever gets typed next. Enter cancels too: it is the key most likely to be
hit by reflex, and this is the one dialog that destroys something.

⚠️ Found while verifying the new dialog: it did not name the session. The label
was computed as `row.session.name ?? id.slice(0, 8)`, and `??` falls back only
on null or undefined, so every session the server left with an EMPTY name — all
of them, until the TUI started naming its own — sailed through and the box read
"Kill ?". A destructive prompt that cannot say what it will destroy is worse
than no prompt, and it is now a single keystroke. The caller passes the same
label the LIST shows, so the dialog names the row in front of the user.

The typed-name machinery goes with it: TuiConfirmState.typed, setConfirmInput(),
confirmAccepts() and the 'typing'/'reject' steps are all removed rather than
left as unreachable branches.
2026-08-22 14:13:58 +02:00
Codeman maintainer 7b1150ca4f fix(tui): start a session straight into it, and drop two unsafe glyphs
Two reports from the same beta screenshot.

Starting a session left the user on the dashboard next to the row they had just
asked for, which reads as the create having silently failed. Starting a session
is a request to WORK in it, so the terminal now goes there as soon as the pane
exists, and the CLI booting is worth watching. If the pane is slow the notice
says so and the row is left selected, exactly as the resume path does.

The footer's `↵` was drawing as an empty box: `⏎` (U+23CE) has poor font
coverage, on the same terminal that renders `·`, `─`, `│`, `○`, `▶` and `✔`
perfectly. It is now U+21B5, from the Arrows block every monospace font ships.

`✋` (U+270B) was worse than a coverage problem: it is East Asian WIDE, so the
renderer, which addresses cells by column, was reserving two cells for it. The
golden frames had the age column shifted a space left to match, which is how
long that had been wrong. It is now `!`, and the frames align correctly.

A test walks the whole unicode glyph set and fails on any entry wider than one
cell, so a glyph that shifts the layout cannot be added again. The comment on
the table spells out both bars a glyph has to clear, because the tier check
answers neither: it asks whether the LOCALE is UTF-8, which says nothing about
whether a font has the glyph or how wide it draws.
2026-08-22 14:13:58 +02:00
Codeman maintainer 95a1f540b5 feat(tui): switch sessions with the web UI's shortcuts
Alt+1..9 switches to that session, and `[` / `]` / Tab step through them, so the
muscle memory from the web UI carries over.

Alt+N SELECTS rather than attaches, which is what the web UI's Alt+N does:
switching which tab you look at is cheap and reversible, and the terminal
equivalent is moving the selection and its preview, not handing the whole
terminal to a pane. Bare 1-9 keeps its documented jump-and-attach meaning.

Two of the web UI's chords cannot cross into a terminal, so the nearest
transmittable keys carry them instead:

  Alt+[ / Alt+]  ESC+[ IS the CSI introducer every arrow key arrives on, and
                 ESC+] is OSC, so neither chord is distinguishable from a
                 sequence. Bare `[` and `]` do the job.
  Ctrl+Tab       a terminal cannot report the Ctrl, so plain Tab carries it.

⚠️ The parser now decodes ESC + a printable character in ONE read as an Alt
chord, and the app replays every chord it does not claim as `escape` then that
character. That fallback is load-bearing, not tidiness: a real Esc landing in
the same read as the next keystroke is byte-identical to a chord, and without
the replay "Esc then q" typed quickly decoded as Alt+Q, matched nothing and was
swallowed. The e2e suite caught exactly that as the dashboard refusing to quit.
A lone Esc is still held and flushed on the caller's timer, which is what keeps
the two separable at all.
2026-08-22 14:13:58 +02:00
Codeman maintainer e7b7e90a1b fix(tui): fold rare prompt glyphs in the preview so they stop rendering as boxes
A beta tester photographed claude's `❯` prompt and its `⏵⏵` bypass-permissions
marker rendering as empty boxes in the preview pane. Their font has no coverage
for those codepoints while drawing `·`, `─`, `│` and `▶` perfectly.

The glyph TIER cannot help here. It answers "can this terminal do Unicode at
all", which is a locale question, and it correctly says yes for exactly the
terminals this affects. Coverage is per-glyph and undetectable from inside the
process, so the handful of rare glyphs CLIs use as chrome are folded to the
ASCII arrows they already look like, and everything a plain font does render is
left alone.

Scoped tightly: the preview only, never the TUI's own chrome, and skipped
entirely at the `nerd` tier where the user has declared a font that can draw
anything. The table is short and every entry was seen as tofu in a real
terminal rather than guessed at. The fold is length-preserving, so the preview
pane's column arithmetic is unaffected.
2026-08-22 14:13:58 +02:00
Codeman maintainer 84f8a8a2fe feat(tui): offer r to resume a session whose pane has died
Refusing the attach stopped the freeze but told the user to throw the session
away (`x` to close, `n` for new), which loses the conversation. tmux's own
dead-pane screen already says what to do instead: `claude --resume "<name>"`.

The Error card now offers `r` when the row can actually be resumed (claude,
with a conversation id and a working directory), and the footer says so. One
press resumes into a fresh pane and attaches to it, so a dead end becomes
recovery.

⚠️ Three things keep this from becoming the resume runaway that once spawned 35
sessions in 40 seconds. The offer holds a session ID, not a row, and is
re-resolved from the model when the key is pressed: a row captured when the
card opened is stale by then. It disarms BEFORE anything async, so a second `r`
cannot start a second resume. And it routes through resumeSelected(), which
owns the `resuming` flag and ends in attachToSession() rather than the group
dispatch.

⚠️ The `r` branch has to run BEFORE the generic dismiss, because a message
overlay is dismissed by ANY key: without that ordering the offer is consumed as
"some key was pressed" and the card merely closes. `help` keeps the any-key
behaviour, so the two modes no longer share a case.

Verified end to end against a genuinely dead claude pane: card, footer, one
press, one new session, and F12 back to the dashboard.
2026-08-22 14:13:58 +02:00
Codeman maintainer 6ad9145417 feat(tui): leave an attach with ONE key, F12, and no modifier
Three beta rounds died on tmux's native way out, and the last one died on the
instruction rather than the mechanism: "press Ctrl+B, release Ctrl, then d" is,
in the tester's words, very unclear, and holding the modifier through both keys
silently does nothing.

So the way out stops being a chord. The attach claims F12 in tmux's prefix-less
`root` table for its own duration, and the bar reads "press F12 to get back to
the codeman dashboard" — one keystroke, nothing to hold, nothing to release,
no order to get right. F12 because stock tmux ships an empty root table apart
from mouse bindings, and none of the CLIs that run in these panes want the key.

⚠️ The bar names the one key ONLY when the claim succeeded, and falls back to
the chord wording otherwise. A bar advertising a key that does nothing is the
bug this whole series started with, and it must not come back in a new costume.
Same claim rules as the prefix alias: taken only when tmux reports the key
unbound, given back only while it still means `detach-client`.

The chord and the held-Ctrl alias both keep working; they are simply no longer
what the user is told to press.
2026-08-22 14:13:58 +02:00
Codeman maintainer ab4a868688 fix(tui): make the detach chord work when Ctrl is never released
Reported three times as "Ctrl+B and d is still not working", on a build whose
bar already named the right key. Measured against a live pane: of the three
ways a person types this, only one worked.

  Ctrl+B, release Ctrl, then d   detaches
  Ctrl+B then Ctrl+D (held)      nothing happens
  Ctrl+B then Shift+D            nothing happens

Holding Ctrl through both keys sends 0x02 then 0x04, and tmux ships `C-d`
unbound in the prefix table, so the keystroke is swallowed in silence and the
attach looks frozen. That is not a user error worth documenting around: holding
the modifier is how most people type a two-key chord.

The attach now claims the held-Ctrl form of whatever key detaches (`d` → `C-d`)
for its own duration and gives it back on restore, and the bar advertises it
only once the claim succeeded, so it can never name a key that does nothing.
⚠️ The key is claimed ONLY when tmux reports it unbound, and released only
while it still means `detach-client`, so a binding of the user's own is never
shadowed or removed. The alias is deliberately excluded from the leaked-state
sweep: key tables are server-global, so the sweep cannot tell a leak from a
second TUI's live claim, and a stray `C-d`→detach is harmless either way.

Ruled out along the way, with evidence rather than assumption: the encoding.
tmux negotiates no extended-key mode upstream on attach (no kitty CSI-u, no
modifyOtherKeys, no DECSET 2017), so Ctrl+B does arrive as a plain 0x02 even
from a Claude pane, which has its own keyboard protocol.
2026-08-22 14:13:58 +02:00
Codeman maintainer 35b2c1baa5 fix(tui): refuse to attach to a dead pane, and stop naming sessions after CLI noise
Two more from the same beta round, both reported as "basic things are broken".

Attaching to a DEAD pane trapped the user. Codeman sets `remain-on-exit on`, so
a session whose agent has exited does not disappear: the row looks ordinary,
the server still reports it idle, and Enter handed the terminal to a pane that
reads no input. With the detach chord also wrong at the time, that was a hard
freeze with no way out. Enter now probes `#{pane_dead}` first and refuses with
an Error card naming the session and what to do instead. The probe fails OPEN,
so it can never block an attach to a live pane. ⚠️ It also has to paint: the
keypress that reaches attachToSession() has already painted by the time an
awaited probe resolves, so message() alone left the refusal invisible and Enter
looked inert, which is the bug it was added to fix.

A session started from the TUI came out unnamed, because startSession() sent no
sessionName and rowLabel() then fell back to the transcript's first line. A
brand-new session has no prompt to be named after, so the list showed a
perfectly healthy session called "Login interrupted" — the CLI's startup
output, reading like a failure report. Sessions the TUI starts are now named
`w<n>-<case>` like the web UI's, and rowLabel() prefers the case directory over
a scraped prompt for any row with a mux name, since a LIVE pane is identified
by where it runs while a history row genuinely is its prompt.
2026-08-22 14:13:58 +02:00
Codeman maintainer bf860382a2 fix(tui): sweep an attach status bar a killed terminal left behind
restore() runs after spawnSync returns, which covers detaching and the agent
exiting inside the pane, but not the terminal dying while attached. Closing the
window or dropping the SSH kills the TUI where it stands, and the bar it
installed stays pinned on the session: the next attach wears a stale bar naming
a different session, and the pane is a row shorter for good. Seen on the beta,
where the tester closed the window instead of detaching.

One sweep at startup, fire-and-forget so it can neither delay the first frame
nor fail a start. Only a bar carrying our own marker is touched, and the marker
is now the single source of the bar's own wording so the two cannot drift; a
user's hand-written status bar on the same session is left exactly as it is.
The session goes back to `status off`, which is how Codeman creates every pane
it owns and the only state this bar is ever applied over.
2026-08-22 14:13:58 +02:00
Codeman maintainer aa487f13ce fix(tui): advertise the key that actually detaches, and stop tmux painting it green
Two things the attach status bar got wrong, both found in a beta test.

The bar read `Ctrl+B D`. tmux key tables are case-sensitive: lowercase `d` is
`detach-client`, capital `D` is `choose-client`. Pressing what the bar said
opened a client chooser and left the tester attached, with the way out on
screen and inert. The key is now READ from `list-keys -T prefix` the same way
the prefix already was, rather than hardcoded, so a rebound tmux is followed
too and the label cannot drift from the binding again. It never goes through
formatPrefixKey(), which uppercases.

The bar also rendered as a full-width bright green slab. Only `status-format[0]`
was styled, so tmux's stock `status-style` (`bg=green,fg=black`) stayed
underneath it and won; `#[reverse]` on top could not undo it. `status-style` is
now set explicitly to `bg=default,fg=default` and snapshotted/restored with the
rest, so the bar sits on the terminal's own background and reads as a hint
line.

Tests pin both: that the chord ends in lowercase `d` and never ` D`, that a
rebound key prints verbatim, that `status-style` is part of the banner, and
that parseDetachKey() picks `d` out of verbatim tmux 3.4 `list-keys` output
while ignoring `detach-client -a`/`-P`, which act on other clients.
2026-08-22 14:13:58 +02:00
Codeman maintainer 0919f9da62 fix(tui): make an attach fit the terminal, show the way out, and resume history
Three things the first beta test surfaced.

1. Attaching from a terminal of a different shape showed the pane clipped to the
   browser's size, with tmux's dot padding filling the rest. Codeman pins every
   window it owns to `window-size manual` at whatever the web client reports
   (tmux-manager.ts), so no attaching client can resize it. The handoff now
   brackets the attach with `window-size latest` and restores the snapshot on
   detach. `latest`, rather than a one-off resize to our own size, is also what
   lets a terminal resized MID-attach follow along: tmux recomputes on every
   SIGWINCH while the TUI is blocked in spawnSync and cannot.

2. Nothing on screen said how to get back out, because Codeman keeps the status
   bar off on its panes (the web UI carries that information around the terminal
   instead). The tester exited the agent looking for the exit, leaving a dead
   pane. An attach now wears a `status-format[0]` bar reading "<prefix> D
   detach, back to the codeman dashboard", with the prefix READ from tmux rather
   than assumed, and the session's options are put back exactly as they were on
   detach. One option, not status-left/status-right, so tmux draws no window
   list beside it; `reverse` so it inherits the terminal's own theme. Restoring
   an array option unsets the BASE name, since dropping the `[0]` index leaves
   an empty array, which renders as a blank bar on a session that had one. The
   help overlay names the chord, and the dashboard confirms the detach.

3. Enter on a RECENT row said resuming was not wired up. It now creates a
   session carrying that conversation (`resumeSessionId` plus `/interactive`,
   the path the web UI's Resume Conversation list already uses), in the
   directory it ran in and under its old name, then attaches to it.

   The attach mechanics deliberately sit in a method the group dispatch cannot
   reach, plus a re-entrancy flag: routing resume back through the Enter handler
   re-dispatched on "this row is RECENT" and spawned one session per pass, 35 in
   about 40 seconds on the beta before it was killed. test/tui/tui-e2e.test.ts
   pins one press to one session with a pane that never appears, which is
   exactly the case that looped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 75b272ff0a perf: pace the refetch and back the tail poll off a quiet pane
Both of the dashboard's periodic reads hit endpoints that are far more
expensive than their cadence assumed, and the cost lands on the SERVER's
event loop, so it is paid by every browser client too.

`GET /api/sessions/unified` is ~550ms against 11 live sessions: it scans
every Claude transcript plus the lifecycle log, uncached, and republishes
the search index. `scheduleRefresh()` was a 250ms trailing debounce with no
floor, and a queued refresh re-ran the instant the previous one returned
(by recursing, which also chained one pending promise per iteration), so a
stream of events paced the refetches at the endpoint's own latency: with
`session:updated` broadcast per session per 500ms while anything is
working, the scans ran back to back. `resyncDelayMs()` now keeps ambient
refetches 3s apart, measured start-to-start. The user's own actions call
`refresh()` directly and are unaffected, so what this paces is only
"notice what changed elsewhere".

`GET /api/sessions/:id/terminal` is ~80-100ms: two `execSync` tmux calls,
then the whole byte buffer normalized before the tail is taken. It was
polled every second for as long as a live row was selected. It now backs
off 1s, 2s, 4s, 5s while consecutive reads change nothing, and resets to 1s
on any change, when the selection moves, when this dashboard sends input or
answers a dialog, and on return from an attach. A pane that is printing is
still read every second; a pane at its composer is not.

The poll also kept running in three places it had nothing to draw for: the
whole time the user was attached in tmux (an attach can last hours), and
behind the message overlays that an async action opens (answered, killed,
started), which are not keystroke-driven and so never reached the
`afterInput()` path that stops it. `setInterval` becomes a chained
`setTimeout`, since the delay now varies.

Measured against the live server, same idle row selected, 25s window:
22 tail reads before, 5 after. With a working pane selected it stays at 22,
which is the intended cadence for a pane whose output you are watching.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 1385415e53 refactor: drop the two store members nothing consults
`TuiModelStore.confirmSatisfied()` and `approvalFor()` had no caller
outside their own tests. The first one mattered: it answered "does the
typed text authorize this kill?" with an exact name match, while the rule
actually consulted (`confirmAccepts()` in tui-app) also accepts the
8-character id prefix a mux name carries. Two divergent answers to one
question, the stricter one unreachable and waiting to be picked up by
mistake. knip cannot see class members, so the dead-code sweep never
flagged either.

The tests they existed for now assert observable state instead, and the
approvals one got stronger on the way: it checks that a session id coming
back does not inherit the dead session's dialog, which is the invariant
`removeSession()` is actually keeping.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 954a9ac26a fix: date a working row by its turn, not by the session age
`TuiSessionRow` declared `lastSubmitAt`/`inputTokens`/`outputTokens`,
`stateSince()` ordered the WORKING group by the first of them and
`renderRowLines()` painted the other two, but nothing ever filled any of
them in: the unified list carries none, and the `session:updated` payload
that does was discarded (an event only schedules a refetch).

So a running turn was dated by its SESSION's creation instead. Measured
against the live server before the fix: w65 (created 21h ago, turn started
one minute earlier) outranked w67 (created 15 minutes ago, turn started
five minutes earlier), the reverse of the rule docs/tui.md states, and the
elapsed column read `21h` for a turn a minute old. The token column was
unreachable code for the same reason.

`fetchLiveSessionMetrics()` reads the three fields from `GET /api/sessions`
and `applyLiveMetrics()` folds them onto the rows. That route answers from
the server's cached LIGHT state (no terminal buffers): 10-20ms measured,
against the ~550ms the unified list in the same `Promise.all` already
costs, so it is cheap enough to ride every refresh. It is best-effort like
the approvals and tmux reads beside it, because losing the anchor is
better than losing the list.

A ZERO is treated as unknown rather than merged: `stateSince()` reads
`lastSubmitAt ?? createdAt` and 0 is not nullish, so a merged 0 would date
every never-submitted session to the epoch.

The snapshot path gets the same merge, or `codeman tui --list` would number
the WORKING group differently from the dashboard that `codeman tui <n>`
indexes into.

Verified live: working rows now show 28m/8m (turn age, tokens 280.5k/65.2k)
where they showed 21h/34m and no tokens. The e2e assertion fails on master's
wiring with `[*] 10m` against a session that pressed Enter one minute ago.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 5fd6c5dd44 docs: extend the instance-isolation rule to tmux socket resolution
The data-dir half was already spelled out; the socket half only lived in
a function docstring, and the TUI is the first code that shells out to
`tmux -L` from a process that is not the server.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 8b5fc974b0 fix: cover tui in the CLI inventory and drop the em-dashes it printed
The inventory test predates the `tui` command, so a rename or an
accidental removal would have gone unnoticed: it now asserts the command,
its `-l`/`--list` flag and its optional position operand.

The digest and search-result lines joined their halves with an em-dash,
which the repo's own convention rules out, so both now use the middle dot
the surrounding lines already use. The one em-dash left in `src/tui/` is
load-bearing: `search-service.ts` builds a session snippet with it, and
the pattern that strips the repeated label has to match it.

Also moves `buildSearchEntries`'s doc comment back onto
`buildSearchEntries`; it had ended up stacked above a helper.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer a52abd9f96 fix: drop the two keymap and style entries nothing reaches
`mark()` had no callers (knip's only finding on this branch), and the
renderer's fallback help list advertised `r` resume, which is deferred
with the rest of phase 3: a help screen naming a verb the build does not
implement is worse than no help.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 008dfddc23 docs: document codeman tui
The user guide covers what the dashboard is (and is not), the two
non-interactive fast paths, the four groups and their ordering, the full
keymap, what answering an approval does server-side, and the SSH/narrow
and degraded cases. The example frame is a real 100x30 capture against
the E2E fake server, not a drawing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 32549789c7 fix: keep the plan-usage chip across a degraded-to-connected upgrade
A server that comes up mid-run was upgrading the header's hostname and
version but not its chip, which then stayed blank until the next telemetry
event. Also swaps a typographic apostrophe out of a preview error, which is
not renderable on the ASCII glyph tier.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 29d9a55eb6 fix: drop stale approvals and re-check the preview when the world changes
Two small honesty fixes at the edges: a server that goes down leaves the
dashboard holding prompts nothing can classify any more and whose answer
route is unreachable, so degraded mode clears them; and a resize can cross
the narrow breakpoint, where there is no preview pane to poll for.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer ef812236b0 fix: read a row-addressed repaint as lines in the preview
Measured against a live Claude pane: an Ink TUI paints by ROW and emits
almost no newlines, so dropping cursor-position sequences collapsed a whole
screen into one unreadable line, and a tail cut mid-sequence printed the
remains of it (";1H") as text. Now a jump to column 1 starts a display line,
a jump inside a row moves the write position (capped, since a stream may
address a column no terminal has), and a severed CSI head is dropped before
parsing.

The preview is readable against a real session as a result: tool calls, the
working line and the composer all land where they belong.

Also drop the repeated session name from a search row, whose snippet opens
with the name the row already shows in its first column.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer b51abe2c27 test: drive the phase-2 verbs end to end under a pty
The fake API server grows the routes the dashboard now calls (terminal tail,
input, approvals answer, search, away digest, plan usage on status), and the
new cases assert on what the server RECEIVED rather than on the frame: the
prompt arrives as one line ending in a carriage return, and the answers as
the exact action and option digit.

Also covered: the tail refreshing in place, the search overlay selecting a
live session, the digest rendering, one bell for an item announced twice,
and the 409 path reported as "no longer on screen".

The plan-usage chip is punctuated with the glyph tier's separator, so an
ASCII terminal no longer gets a stray middle dot in the header.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer eb8958ddd0 feat: answer approvals and send prompts from the dashboard
The dashboard stops being read-only. The selected session's tail is polled
once a second while the plain list has focus and the layout is wide, and an
unchanged tail never reaches the model, so a quiet session costs no repaint.
A row with no live buffer says so instead of polling forever.

Keys: y/n and the parsed digits answer the selected session's dialog through
`POST /api/approvals/:id/answer` (never a blind keystroke: that route
re-captures the pane and 409s when the dialog has moved on, which the TUI
reports as "no longer on screen"); `p` opens a one-line composer aimed at
the selected session; `/` searches with a 250ms debounce and Enter switches
to a live session result; `g` shows the away digest. A new prompt rings the
bell exactly once, tracked by item id so a repaint or a refetch cannot
stutter, and the plan-usage chip rides `GET /api/status` plus its telemetry
event.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 953a560eee feat: render the approval card, composer, search and digest
The preview pane now leads with the pending dialog when the selected session
has one: the question, the options with their digits, and the keys that
answer them, red for a dialog and yellow for a waiting prompt. The card is
capped at half the pane, because the tail is why the pane exists.

Around it: a header badge counting prompts that need a human, a preview
title that sacrifices the path rather than the state word, the footer
becoming the composer line while one is open (with the cell the terminal
cursor belongs in, so it can be shown there and hidden everywhere else), and
the search and digest panels as overlays with a stable width.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer dc89f05b14 feat: hold composer, search and digest state in the TUI model
The store gains the three overlays phase 2 needs, each taking the keyboard
when it is set and all of them cleared together by closeOverlay(), plus the
pure flattening of `GET /api/search`'s typed groups into rows a cursor can
move over: headers are chrome, and only a session that is on the list counts
as selectable, since a history hit has no row to move the cursor to.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer bb73400afa feat: add the TUI's editor, approval and digest pure cores
Three small pure modules the phase-2 verbs are built on:

- tui-composer: the single-line editor behind `p` and `/`, holding text as
  code points so a cursor can never split a surrogate pair, with the scroll
  window derived from the width rather than remembered.
- tui-approvals: what an approvals-inbox item's card says, which keys are
  live for it (a digit answers only when the server parsed that option, and
  an idle prompt answers to none of them), and which ids the bell has not
  rung for yet.
- tui-digest: the away digest as compact lines, counts first and one line
  per entry, with a capped tail per section.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 9d7dd2ab62 test: drive codeman tui end to end under a pty
Spawns the real command in a pseudo-terminal against a fake API server
(canned status/unified/approvals plus an SSE stream the test pushes
into), which is the only way to cover raw-mode key decoding, frames
reaching a terminal, SSE-driven refresh and the exit sequence that has to
restore the user's screen.

Two details the assertions depend on: frames are addressed absolutely
rather than newline-separated, so the parser takes the last COMPLETE
frame (the pty delivers one in several chunks, and reading a half-written
frame would be racy), and it reads the sidebar column only, or a name
echoed in the preview pane could answer for a row.

The child gets its own data dir and a tmux socket name nothing runs on,
so nothing here can see or touch the machine's real sessions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 3f88226d50 feat: register the tui command with its two fast paths
`codeman tui` opens the dashboard, `codeman tui --list` prints the
numbered list and exits (the `sc -l` replacement, plain when piped) and
`codeman tui <n>` attaches straight to a row (the `sc 2` replacement).
Both fast paths short-circuit before any screen setup, and both refuse
the numbers path without a terminal instead of half-opening a UI.

Bare `codeman` still prints help: the web UI stays the primary surface.
The TUI module is imported lazily so the other commands do not pay for it
at startup.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer e5ae2826c6 feat: add the codeman tui dashboard
The IO half of src/tui: it owns the terminal, the timers, stdin and the
tmux handoff, and every decision it makes that is a function of its
inputs is an exported pure helper with unit tests (attach planning, the
typed kill confirmation, keymap selection, the repaint test, degraded
rows).

What it does: live session list over the unified API with SSE-driven
resync (debounced, with a 2s poll fallback the client asks for), cursor
and 1-9 navigation, attach and return, kill behind a typed confirmation
that refuses history rows and the session hosting the TUI, a new-session
case and CLI picker over quick-start, and degraded mode straight from
tmux when no server answers, re-probing so a server that starts upgrades
the dashboard in place.

Restoring the terminal is the part that has to be bulletproof: leave() is
idempotent and runs from normal quit, SIGINT/SIGTERM, a process exit hook
and prepended fatal handlers (src/index.ts already handles those by
exiting, so a listener registered after it would never run).

Attach is a handoff, never a proxy: the screen is restored and tmux gets
the real terminal. Inside tmux on the same socket there is nothing to
hand off to, so it issues switch-client and exits.

The preview pane, approvals answering, the prompt composer, search and
the digest are the next step; the region renders a placeholder rather
than pretending to load something.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 85b69e0923 feat: render the TUI picker overlay and a caller-supplied keymap
The footer and the help overlay held the plan's full keymap, which would
advertise verbs (prompt, search, digest, answer, resume) that the build
does not implement yet and teach users that the TUI ignores keys. Both
now take their entries from the render options when the caller passes
them; the built-in lists stay as the fallback.

The picker overlay windows its items around the cursor rather than
clipping them, so the selected case stays visible in a long list.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 6682231d68 feat: give the TUI model a revision signal and picker state
The app layer repaints on state change, so the store has to be able to
say that something changed: `revision` is bumped by every mutating
method, and the repaint test compares it against the last painted frame.
Without it an idle dashboard would either redraw on a timer or go stale.

Three additions come with it, all optional so nothing existing changes
shape: `TuiSessionRow.muxName` (the unified list carries no mux name, so
the app fills it in from the local tmux enumeration and a row without one
cannot be attached), a `new-session` UI mode, and `TuiPickerState`, the
one-column chooser behind `n`.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer e140132e45 feat: add the TUI's API, SSE and degraded-mode client
Everything the dashboard needs from outside the process, behind one typed
surface, so the app loop stays a loop. It is a client of the running server and
nothing else: rows come from the unified list, blocked states from the
approvals inbox, and answering goes through the endpoint that re-captures the
pane and refuses with a 409 when the dialog has already been answered in tmux.
That refusal is a typed result rather than an exception, because a human
beating you to a prompt is normal operation.

Discovery mirrors the daemon probe (`CODEMAN_API_URL`, else loopback on
`CODEMAN_PORT`, self-signed TLS accepted) and credentials come from where
`codeman attach` already reads them. An explicit port outranks the ambient
`CODEMAN_API_URL`, which every managed session exports: a caller that named a
port must not be redirected at whatever server owns its shell.

Input is single-line and `\r`-terminated at this layer, so no caller can strand
text on an unsubmitted composer, and each send is tagged for the server's
exactly-once path. The event stream defaults to a `?sessions=` filter that
matches nothing, which drops the terminal firehose while lifecycle, hook and
approval events still arrive. A silent-but-open stream is caught by a watchdog
rather than a socket error, since that failure mode reports nothing at all.

With no server answering, sessions are listed from tmux on the instance socket
(argv, never a shell string) and decorated from a read-only peek at state.json,
which keeps the "the server died, get me to my sessions" path alive.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 6475c010a6 feat: decode the SSE wire format for the TUI
Node has no EventSource, so the live-update stream is read as raw bytes and
decoded here. Three details are what the parser exists for: a TCP read can end
between the CR and the LF of a CRLF, so a trailing CR is held back rather than
dispatched; the tunnel padding the server appends after a frame is a comment
with no blank line after it and must not split anything; and the keepalive is a
NAMED event, because an SSE comment is invisible to a browser client by spec.

Event classification lives here too, as a set rather than a prefix test:
`session:terminal` is most of the stream and the preview pane pulls its own
tail, so it is deliberately not a resync trigger.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 64c8048dda refactor: resolve the tmux socket from the instance config
The socket name was computed inside tmux-manager, which the TUI cannot import
just to learn which `-L` name its degraded-mode listing belongs on (that module
is the server's tmux driver, not a lookup table). The resolver moves next to
`dataPath()`, where the other half of the instance identity already lives, so
both processes agree by construction instead of by a copied default.

Behaviour is unchanged: the override still wins only when it is a name that can
be passed to `tmux -L` safely, and TmuxManager keeps warning about one that
cannot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 0566ea3453 docs: record why the key parser reads LF as Enter
Ctrl+J is unbindable as a result, which is worth knowing before someone tries
to bind it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 6ef3b2ba2e feat: render TUI frames from the model and layout
One absolutely-addressed line per row, each closed with an erase-to-end, so
nothing scrolls and a repaint cannot leave the previous frame's tail behind.
The caller wraps the result in synchronized-output brackets; that is an IO
decision and stays out of the renderer.

Color is passed in rather than detected. chalk's detection is right for the
one-shot CLI but would make a frame non-deterministic, so the palette is raw
SGR in the same semantic roles cli-style uses, and `color: false` emits nothing
but the cursor addressing, the session's own colors in the preview included.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 74fe2cad9f feat: add the TUI responsive layout math
Below 72 columns the preview pane is dropped and rows take two lines, the
constraint the `sc` chooser was built around and the reason it is still usable
on a phone; above it a clamped sidebar carries the list and the preview takes
the rest.

Every region is clamped to a non-negative size, so a 5x5 terminal degrades to a
header instead of handing the renderer negative widths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer aa2deea73e feat: add the TUI session model, classification and cursor
Rows are the ones GET /api/sessions/unified already returns and blocked states
are the items the approvals inbox already parsed, both imported as types only
so a CLI process pulls in neither the server nor node-pty. Classification
speaks the web UI's language (red blocked, yellow waiting, green working) so a
user with both surfaces open never has to translate between them.

Groups order by how long a session has been in its state, which is why WORKING
anchors on the pane's last Enter: a working pane repaints about once a second,
so its last-activity stamp always says "now".

Selection is tracked by session id, never by row index: rows re-sort under the
cursor whenever a session starts working or an approval lands, and an
index-tracked cursor would quietly move the selection to another session
between two keystrokes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 5d7fdb528b feat: add the TUI raw-mode key parser
Decodes printable UTF-8, the control keys, arrows in both CSI and SS3 forms and
SGR mouse reports out of a byte stream that can tear anywhere, so a sequence
split across two reads decodes the same as one that arrives whole.

A lone ESC cannot be told from the start of an arrow key by looking at bytes,
so the parser holds it and the caller resolves it with flush() once its
disambiguation timer fires. Unknown sequences are swallowed: a stray CSI must
never reach a prompt composer as typed text.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 64cf8384f2 feat: add the TUI's SGR-aware preview helpers
The preview pane shows a session's raw terminal stream, so it needs the tail
reconstructed rather than emulated: SGR survives, cursor steering and OSC do
not, and a carriage return returns to column 0 so a spinner that repaints its
line 200 times contributes one line instead of 200.

Widths count East Asian Wide characters as two columns, which the clip and pad
helpers rely on to never cut a wide character, a code point or an escape
sequence in half.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 596c08d20c chore: stop ignoring src/tui
The entry dates from an abandoned prototype (0.1427) and would have kept the
real TUI modules untracked while `git status` stayed silent about it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 1d0c3650f9 docs: fix the codeman attach description and the detach prefix
`codeman attach <path>` posts an attachment card for a local file; it
was described as attaching a Claude hook context. And Codeman never
overrides the tmux prefix for local sessions (only remote-SSH and docker
panes get C-q), so the detach hint is Ctrl+B D, matching the chooser.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer b9afd5a57e test: derive the CLI inventory from the real commander program
The file asserted against a hand-written fixture array with its own
argument parser, so it could not see a command being renamed, losing an
alias or disappearing, and it described a `tui` command that does not
exist. It now walks program.commands: names, aliases, subcommands,
option flags, operands, descriptions, and a guard against registering a
name or alias twice at one level.

Assertions are "at least this exists", so a new command (including the
tui one this plan adds later) passes without editing the test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 14a911b4f6 fix: color the server startup line and its security warning
The startup banner is now the only one (the CLI printed a duplicate) and
is painted like the rest of the CLI. The non-loopback-without-password
warning was plain console.warn while the CLI's copy of the same warning
was yellow; chalk degrades off a TTY, so journald and web.log stay free
of escape codes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 9b1d269943 feat: wire the CLI through the style kit (doctor colors, spinners, confirm)
- doctor is colorized through the ReportStyle hook: verdict glyph and
  failing status text painted, paths and hints muted, versions left
  alone. `doctor --json` still prints raw JSON.
- `codeman web -d`, `web --stop` and `service install` block for up to
  30s polling /api/status; each now runs under a spinner instead of a
  silent terminal.
- `codeman reset` asks a real y/N question on a TTY. Non-interactive
  callers keep the old "Use --force to confirm." refusal, so no script
  can be answered by a question it cannot see.
- `codeman list` was a drifted copy of `codeman session list`; both now
  call one renderer, with the shorthand opting out of the stopped and
  web-server sections.
- `web` no longer prints its own "running at" line: the server prints
  one, and unlike this one it also covers the daemon and service paths.
- every chalk call goes through the palette, so the CLI has one place
  where colors are decided.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 4e5d0dcbd6 fix: measure the doctor table columns and let the CLI paint them
"Antigravity CLI" is 15 characters and the hardcoded padEnd(14) pushed
that whole row one column right. Widths now come from the widest cell.

The header always said the CLI layer may colorize, but there was no way
to: renderTable now takes an optional ReportStyle whose hooks are
identity by default, so the module still decides nothing about color and
its output stays byte-stable. Padding is applied outside the paint, so a
row with no path detail ends at its status text instead of trailing
spaces inside a color run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer f9d6c4f0c3 feat: shared CLI style kit
One vocabulary for everything the codeman CLI prints: semantic palette,
the glyph set the commands already used, heading/rule/kv, width-aware
table layout, a stderr spinner and a y/N confirm.

Color detection stays chalk's, so NO_COLOR and non-TTY degradation keep
working with no second detector to disagree with it. The layout math and
glyph selection are pure and exported, which is what lets the dependency
report reuse them while staying color-free.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 09d6bb9eb0 docs: TUI rework plan (codeman tui, herdr research)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer a922d301b1 chore: version packages 2026-08-21 20:24:38 +02:00
Ark0N abca552676 Merge pull request #327 from dignfei/fix/terminal-ime-punctuation
fix(terminal): preserve IME punctuation input
2026-08-21 20:23:26 +02:00
Ark0N 12a996b107 Merge pull request #331 from dignfei/fix/shell-history-performance
fix(terminal): bound shell history replay
2026-08-21 20:23:17 +02:00
d fei 458e751a33 fix(terminal): keep shell history loading explicit 2026-08-22 01:55:50 +08:00
d fei dab432b3fd fix(terminal): bound shell history replay 2026-08-21 08:23:31 -04:00
d fei f744719650 fix(terminal): preserve IME punctuation input 2026-08-20 10:35:23 -04:00
348 changed files with 59416 additions and 4316 deletions
+18
View File
@@ -0,0 +1,18 @@
.git
.agents
.claude
.codex
# `**/` matters: a .dockerignore pattern is matched against the WHOLE
# context-relative path, so a bare `.env` excludes ONLY the root file and
# `COPY . .` would bake docker/.env -- CODEMAN_PASSWORD and any provider API
# keys -- into the published image at /opt/codeman/docker/.env (verified).
**/.env
**/.env.*
!**/.env.example
node_modules
dist
coverage
out
test-results
tmp
*.log
+12 -4
View File
@@ -69,10 +69,18 @@ shared-host, multi-user, or tunneled deployments.
- **Multi-instance tmux socket is process-wide.** Two Codeman instances on the same `CODEMAN_INSTANCE` share a tmux socket and can attach each other's live sessions — isolate with distinct `CODEMAN_INSTANCE` values.
- **The live log-tail route reads `/var/log` and `~/logs`** in addition to the session working directory (read-only) — a deliberate choice for tailing system/app logs. On a password-protected remote deployment an authenticated user can therefore read those roots outside their session. See `docs/security-architecture.md` §5.
Recent hardening (this release): web-push subscription endpoints are restricted
to https public hosts (SSRF guard — rejects internal/metadata IPs, validated at
subscribe and send time), and tmux session names discovered on the shared socket
are validated against the safe-name pattern before reaching any shell call site.
- **The web-tab proxy fetches from the server's network position.** Any authenticated user can save a dashboard URL on loopback or a private range and have Codeman relay to it; that is the feature. Link-local and cloud-metadata addresses are the only refused targets (see below). On a shared host, restrict who holds an account.
Recent hardening (2026-09-04): the web-tab proxy, its "Test" probe and its
WebSocket relay refuse link-local and cloud-metadata targets (`169.254.0.0/16`,
`fe80::/10`, `fd00:ec2::254`, `168.63.129.16`, `100.100.100.200`,
`metadata.google.internal`), judged on the RESOLVED address so a DNS name pointing
there is refused too; proxy capabilities are revoked on logout, admin logout and
user deletion; proxied responses carry `Referrer-Policy: same-origin`. Earlier:
web-push subscription endpoints are restricted to https public hosts (SSRF guard,
rejects internal/metadata IP literals, validated at subscribe and send time), and
tmux session names discovered on the shared socket are validated against the
safe-name pattern before reaching any shell call site.
For the detailed rationale, defenses, and recommended secure setups, see
[`docs/security-architecture.md`](../docs/security-architecture.md).
+3 -2
View File
@@ -2,6 +2,9 @@
.agents/
skills-lock.json
# In-session decision scratchpad (context-survival mechanism, not a deliverable)
DECISIONS.md
# Written by install.sh into end-user clones when setup finishes
.install-complete
@@ -93,8 +96,6 @@ packages/gesture-control/.vite/
# Claude Code plan tracking
plan.json
# Unfinished TUI (local development only)
src/tui/
.claude/
media-assets/
commands
+2 -1
View File
@@ -2,7 +2,8 @@
Canonical agent/contributor guidance for this repository lives in [CLAUDE.md](CLAUDE.md) —
project structure, build/test/lint commands, code style, testing safety rules
(never run the full suite inside a managed tmux session), security notes, and
(`npm test` is the CI gate and is safe to run bare; the three excluded suites
have their own runners), security notes, and
the deployment workflow are all maintained there. Please read it before making
changes, and keep it the single source of truth rather than duplicating
sections here.
+426
View File
@@ -1,5 +1,431 @@
# aicodeman
## 1.27.0
### Minor Changes
- Session lists that answer "which of these wants me next?", loopback links that work from a phone, and a batch of input and remote-session fixes.
**The vertical tab rail sorts by activity and wears the home screen's cards.** A new per-device setting (App Settings → Appearance → Tabs → **Vertical Rail Order**, default _By activity_) orders rail rows with the same comparator both home screens use: whatever is blocked on you first, then whatever has been running longest, then the most recently quiet. Detailed rail rows become cards, with the state dot keeping its working ring and gaining the home rail's green halo. ⚠️ Existing vertical-rail users get sorting on upgrade, and a self-sorting list cannot also be drag-reorderable: choose _Manual_ to get your own order and drag-reordering back. The lineage bracket also moves 4px further from the rail's left edge, where its glow was being clipped by the window frame.
**The Claude Response Viewer's brief view shows the whole last turn.** It used to render one row, so the eye button often showed the "Done." tail of an answer whose substance was in the rows above it. A multi-row turn now also opens at its newest text instead of its first narration line.
**A `localhost` link in agent output opens as a proxied web tab.** An agent prints `http://localhost:5173/` and you tap it on a phone: that address only exists on the Codeman box, so the link was a guaranteed connection error from any other device. It now opens through the proxy, reusing a saved dashboard for the same dev server (one tab per server, not per host spelling) or saving one under its `host:port`. LAN and tailnet addresses still open directly, and on the box itself every link opens directly. `*.localhost` is deliberately not auto-routed: it is the only spelling that is a DNS name rather than an address literal, and these links come from agent output; add such a dashboard by hand instead. Trusted (non-sandboxed) dashboards are likewise never auto-reused by a tapped link.
**Remote omp and remote claude sessions continue their conversation across a respawn or reattach.** Remote claude now launches an idempotent `--session-id || --resume` pair and remote omp respawns with `--continue`, instead of starting a fresh conversation each time. An omp session id is never resolved from the local `~/.omp` for a remote session, which would have pinned an unrelated local conversation.
**Android and IME keyboards no longer drop committed characters.** Chrome on Android delivers a `composed: true` input event preceded by a keydown, which is exactly the shape xterm refuses to forward, so the character vanished. A recovery controller forwards it when, and only when, xterm produced nothing for that keystroke, so dictation and soft-keyboard input cannot be delivered twice either.
### Thanks
- **@shenlvkang-collab** for the Response Viewer last-turn fix (#400) and for loopback links as web tabs (#401), both carefully measured, #400 against 285 real transcripts.
- **@timkjr** for remote-omp resume/continue through respawn and reattach (#362), including dropping a half that had already landed and verifying the merge kept none of it.
- **@aakhter** for the Android/IME input recovery (#388), and in particular for finding that an earlier version of their own browser test was passing vacuously, and saying so.
## 1.26.2
### Patch Changes
- Terminal rendering fixes, a Ctrl+V paste fix, an iOS Safari toolbar fix, a 2GB download cap, and a Blur entrance animation.
### Terminal rendering
Three independent causes behind #398, where opening a session rendered a frame with characters spliced into each other and left the caret on the composer's border instead of its input line, until the CLI next wrote anything:
- **The full-history replay now keeps row alignment** (#395). The linear capture path never restored the cursor, so every cursor-relative update the CLI sent afterwards was measured from the status line instead of the pane's real position, and four transforms that each can delete a line (trailing-blank stripping, redraw-bloat stripping, the pre-banner trim, leading-whitespace removal) shifted the frame out from under it. The full-history path now appends the pane's own cursor position and keeps every row, so row N of the reply is row N of the pane. The visible-frame and tail paths are untouched.
- **The first fit waits for the terminal font** (#396). A cell measured against a fallback font gives the wrong column and row count, so the pane was sized twice and the CLI repainted for a shape that no longer matched the frame on screen. `selectSession` now holds for the font before measuring, bounded at 2s so a font that never arrives cannot strand a session, and it ends by re-measuring explicitly — `FitAddon.proposeDimensions()` divides by a cached cell size and nothing in it listens for font loading, so waiting alone would still divide by the fallback cell.
- **A detached session's own window owns its pane size** (#397). Popping a session out left both windows sizing one PTY, and the dashboard's terminal is narrower than the popup because the session rail takes width the popup does not have, so the CLI drew frames that fit neither. The dashboard now withholds the resize send (never the local reflow) for a session showing in its own window, and takes sizing back on redock.
### Other fixes
- **Ctrl+V no longer pastes twice** (#394). One keypress delivered two paste events to the clipboard trap: Firefox dispatches a trusted event for `document.execCommand('paste')` and then returns `false`, and the key's own default action fires another, because xterm's custom key handler returns false without cancelling the keydown. Right-click → Paste has no keydown, which is why only the keyboard duplicated. The trap now consumes exactly one event per keypress.
- **iOS Safari: the phone toolbar sits on Safari's bottom bar** (#391, #392). The toolbar was lifted by `100vh - --app-height`, which on iPhone Safari measures the bar's collapsible height rather than an overlap — fixed elements there already stop above the bar — leaving an empty ~40px band and padding the terminal by the same amount. The lift is now `--chrome-overlap` (`innerHeight` minus the visual viewport height), which is 0 on iPhone Safari and equals the real overlap anywhere fixed elements do land behind the chrome.
### Downloads
`file-raw`, the attachment `/raw` route and `GET /api/download` now cap at **2GB** instead of 50MB, configurable via `CODEMAN_MAX_DOWNLOAD_BYTES` (`0` = unlimited). The old cap was memory protection for a `readFile()` that no longer exists: those bodies stream and answer `Range` requests, so size costs a read stream rather than RSS (measured: a 600MB download moved peak RSS by ~37MB), and all the cap still did was refuse legitimate downloads of build artifacts, videos and archives. `/api/download` was the last route that really did buffer the whole file; it now streams, advertises `Accept-Ranges` and is resumable. Refusals move from `400` to `413`, the correct status for the case.
### Blur entrance animation
A new opt-in `Blur` style on all four entrance surfaces (tabs, agent windows, the terminal pane, connection lines), plus a `Soft focus` theme that sets all four: an iOS-style focus pull where the thing arrives out of focus and the blur fades off it as the opacity comes up. App Settings → Appearance → Entrance Animations, or mix per surface at `?animlab=1`. Entrance animations stay off by default, so an untouched install is unchanged.
### Maintainer tooling
The PR bot now fails fast when the review model's budget is spent, instead of hanging a review for the full 40-minute timeout and burning its retry cap.
### Thanks
- @irisitymichaelgrundberg for #394, #395, #396 and #397, and for the #398 investigation that separated three causes behind one symptom
- @JDProfresh for reporting #391 and fixing it in #392
## 1.26.1
### Patch Changes
- Codex sessions no longer report idle for their entire life, and Codex conversations now appear in Past Sessions and can be resumed.
**Per-CLI work detection (#385, irisitymichaelgrundberg).** The composer glyph and the working status line are now registry data (`capabilities.workDetect`) rather than Claude constants. Claude keeps its exact current pair, Codex declares `›` plus its `esc to interrupt` footer, and any CLI that declares neither falls back to Claude's, which is what every session used before. Work detection had been gated Claude-mode-only on the reasoning that an external CLI has no `❯`, which was true and still left every Codex session reporting `idle` from the moment it started. `workingLine` is config-supplied and its compiled pattern runs on the PTY hot path, so it goes through `compileVersionRegex()` in both the schema refine and the runtime compile: a nested quantifier there would backtrack on the event loop for the whole server. The Codex footer is matched case-insensitively on the E, so a future version capitalising it cannot make the fix silently inert.
**Codex conversations in Past Sessions (#386, irisitymichaelgrundberg).** A bounded scanner reads codex's `~/.codex/sessions` rollout store, so the unified session list now merges three transcript stores rather than one (Claude's `~/.claude/projects`, omp's `~/.omp/agent/sessions`, codex's `~/.codex/sessions`). A scanned row carries a `resumeId`, the rollout's own thread id, which lets it resume through `codexConfig.resumeSessionId`; a live session never carries one, so a row without it stays a genuinely fresh session. Live and resumed Codex sessions fold into their rollout row through the existing alias map, including a `session_meta.originator` match for fresh panes, so a conversation never shows up twice. The phone overview carries `resumeId` through its own row projection, without which a tapped Codex past row started a fresh session on a thread already on disk.
### Thanks
- @irisitymichaelgrundberg for both PRs (#385, #386), and for turning a full review round on #386 in a day.
## 1.26.0
### Minor Changes
- Tag the case directories agent workers create, and clean up what they leave behind.
A long agent orchestration creates one case directory per worker, and deleting the
sessions never removed them, so `~/codeman-cases` filled with scratch folders that
looked exactly like real projects.
- A case directory `POST /api/quick-start` **creates** for an agent-driven spawn now
carries a `.codeman-agent-case.json` marker recording when it was made, by whom,
from which session, and in which mode. Only the branch that creates the directory
writes it, so a linked case, a cloned repo or any pre-existing path is never
labelled, and deleting the marker file adopts a scratch case as a real one.
- The label comes from the new `X-Codeman-Agent-Origin` header that the packaged agent
skill sets on its shared curl invocation (preamble 1.22.0), or an `agentOrigin` body
field, falling back to a resolved `parentSessionId` so workers spawned by an older
skill copy are still labelled.
- `GET /api/cases` publishes it as `agentCreated`, and the new read-only
`GET /api/cases/agent-created` lists the scratch cases with `inUse` (a live session
is still working in it) and `modifiedAt`.
- Add Case -> Manage badges every agent-created case and adds a sticky **Clean up**
entry point that names each directory in its confirmation and skips any case a
running session is using. Removal still goes through `DELETE /api/cases/:name`.
- The agent skill's per-session preamble cache (`~/.cache/codeman-agent-<id>.sh`) is
now removed with the session and swept at boot. One was written per Claude session
and nothing ever deleted them (236 orphans on a working machine); the sweep keeps
every live session's file and only takes orphans older than seven days.
## 1.25.0
### Minor Changes
- Codeman can be mounted under a sub-path behind a reverse proxy (#381, @mtiller). `--base-url /codeman` (or `CODEMAN_BASE_URL`) makes the server strip the prefix on the way in, rebase redirects on the way out, inject `<base>` and `window.__CODEMAN_BASE__` into the shell, and route web-tab proxying and WebSocket upgrades under the mount, so one TLS name can front several apps. A root install is byte-identical to before. Applied on top: the crash-diag beacon stays under the mount (sendBeacon is not fetch, so the base-aware wrapper never saw it), the test suite strips `CODEMAN_BASE_URL`, and a wiring test boots a real server under a prefix.
A case can attach to a container that is already running (#357, @dignfei). `DockerCase.owned:false` mirrors the remote-SSH attach contract: Codeman only execs into such a container, never creates, starts, stops, removes, pauses or commits it, with the refusal enforced at string-construction time so no caller bug can reach `docker stop`. The Add Case dialog gets an attach panel with a container picker, the run menu takes its mode availability from the CLIs actually present in the container, and adoption is admin-only in multi-user mode. Three gaps closed after review: export no longer pauses or commits an adopted container, a freshly linked owned case no longer hides every agent mode behind a probe of a container that does not exist yet, and multi-user gating is explicit.
The Claude response viewer renders one message per model message (#369, @shenlvkang-collab). The reader used to fuse every assistant row between two human prompts into one card and never read the attachment rows that hold a prompt typed mid-turn; measured over 57 real transcripts it now shows 1,806 messages instead of 356 and recovers 162 absorbed user prompts, with the assistant text unchanged row for row.
A Claude pane learns its live conversation from the CLI's own `UserPromptSubmit` hook (#367, @shenlvkang-collab). The conversation id used to be re-derived by correlating `~/.claude/history.jsonl` against a stamp only Codeman's own input path set, so a pane driven straight from tmux stayed pinned to its launch conversation forever. The hook reports the id first-hand, addressed by the pane's own `$CODEMAN_SESSION_ID`, and the chain of conversations is persisted so a restart re-pins the right one. The new `hook:prompt_submitted` SSE event is registered (158 = 158), and it lands in the run summary only when the conversation actually moved.
The Add Case modal can be submitted from a phone again (#368, @shenlvkang-collab). Since 1.16.4 the layout below 860px hid the modal footer, which held the only Create/Clone/Link button. A header submit button now sits beside the close button, dims while a submit is pending, and a static test pins the contract so it cannot silently disappear again.
The Link Existing case picker opens in the Codeman Cases directory instead of Home (#383, @opticon454). Under Docker the two are unrelated trees and Home holds nothing but dot directories, so the picker showed no cases at all. The fallback chain is now Current Folder, then Codeman Cases, then `/mnt/d`, then the first root.
A PR review bot for the maintainer (`scripts/pr-bot/`, guide in `docs/pr-bot.md`). It reviews every open pull request in its own Codeman session inside a private clone and reports the verdict, ranked findings and a recommendation to Telegram with action buttons; merge, close, post-comment and approve-CI happen only from a confirmed tap. Maintainer tooling, not part of the server or the CLI.
### Thanks
- @mtiller for the reverse-proxy base URL (#381).
- @dignfei for attaching cases to running containers (#357).
- @shenlvkang-collab for the response viewer fix (#369), the first-hand conversation hook (#367) and the phone Add Case fix (#368).
- @opticon454 for the case picker default (#383).
## 1.24.7
### Patch Changes
- The web-tab proxy refuses link-local and cloud-metadata targets. Its Test probe, the proxy itself and the WebSocket relay accepted any http(s) host, so a saved dashboard URL could reach `169.254.169.254` (in decimal, hex, IPv6-mapped or DNS-name form) through a capability and no cookie. Loopback and RFC1918 addresses stay allowed on purpose, since a localhost Grafana is the feature; only link-local and the fixed cloud-metadata addresses are refused, at the schema, at every connect site, and through a DNS lookup hook that judges the resolved addresses, which is what closes DNS rebinding. Adds `undici` so the proxy runs its fetch through its own agent.
Proxy capabilities are revoked on logout. `revokeOwner()` had shipped with no caller, so a leaked proxy URL stayed valid for as long as anything kept polling it. `POST /api/logout`, the admin forced logout and user deletion now revoke the capabilities they should, and proxied responses carry `Referrer-Policy: same-origin` with the upstream's own policy dropped, so a dashboard on a loose referrer policy cannot hand the capability to a third-party host it links to.
The Docker Compose deployment updates itself from App Settings again (#373, @opticon454). The checkout Compose builds from is bind-mounted at `/opt/codeman`, so an update's `git checkout` and rebuild land on the host and survive container recreation; build artefacts live in named volumes so container-compiled native modules never enter the host checkout; the image keeps devDependencies and a build toolchain; and the restart is the server exiting under `restart: unless-stopped`. An in-place update applies code only, so the updater refuses a release that changes `server.Dockerfile` or `docker-compose.yaml`, or that adds keys to `.env.example` the user's `.env` has no value for (Compose interpolates an unset variable to the empty string and starts anyway), and points at `docker/Start-Codeman.sh` on the host instead. The four global agent CLIs in the image are pinned. A follow-up makes the final step fail safe: the server exits only when the Compose file declares `CODEMAN_RESTART_BY_EXIT=1` or the daemon confirms an auto-restart policy, and otherwise the build is staged for a manual restart, so a container nothing would restart is never taken down. Details in `docs/docker-self-update.md`.
The test suite strips `CODEMAN_INSTANCE`, `CODEMAN_DATA_DIR` and `CODEMAN_TMUX_SOCKET` before any application module loads (#371, @opticon454), with a two-half test whose static half reads `test/setup.ts` so a dropped line fails everywhere. This replaces the throwaway data dir #356 had set for the same variable.
### Thanks
- @opticon454 for the Compose self-update (#373) and the test isolation fix (#371).
## 1.24.6
### Patch Changes
- CLI backends are now a data-driven registry (#347, @opticon454). Every run mode (Claude Code, Terminal/Shell, OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek Harness and OMP) is a `CliEntry` in `src/config/cli-registry/`: binary discovery (search dirs, version and identity probes), the launch argv template, environment handling, the multi-user privileged-parameter and privileged-env-key clamps, the remote and Docker pane commands, and the capability flags the rest of the app reads instead of branching on a CLI's name. `~/.codeman/clis.json` can override any stock entry or add a custom CLI; it is read-only in this release, must be mode 0600, and every reason it was ignored is now logged once on first load (`docs/cli-registry.md`). Config never contains shell text: entries declare typed argv tokens, literals are validated at load time, and values resolve through patterns named in code. Registry data resolves at call time rather than at module import, so a CLI enabled while the server runs moves every surface at once, and a guard test fails the build if per-CLI-id branching reappears outside the stock catalog.
This is an internal refactor. The spawn command every CLI receives is byte-identical to the previous hand-written builders, verified by pinned golden strings in the test suite and by diffing both implementations across 11,602 option combinations for all ten modes. Five small deliberate changes ride along: the in-container version probe derives the binary from the registry (`antigravity` runs `agy`), the remote version probe now covers Grok and DeepSeek, `codeman doctor`'s CLI rows are generated from the registry (Claude's install hint is the install command, five CLIs gain hints, the row order follows the catalog), OMP now requires tmux like its siblings instead of silently falling back to a direct PTY, and an OMP session's attach client now receives `COLORTERM=truecolor` like the other truecolor CLIs.
Remote sessions are no longer auto-revived after a clean agent exit (#355, @timkjr). The reconnect watcher could not tell a transport drop from a Ctrl-C, Ctrl-D or `exit` inside the remote CLI, so a clean exit relaunched a fresh agent (OpenCode and OMP started a new conversation every time; Claude only looked fine because its `--resume` fallback masked it). The watcher now revives a dead pane only when the durable remote tmux session is verifiably still alive, via a `has-session` probe over ssh, and an unreachable host means do not revive. A follow-up classifies that probe by exit status, since `tmux has-session` prints nothing on success and reading its stdout had marked every live session as gone, forgets the cached answer whenever the pane is seen alive again so a stale result cannot revive a later clean exit, and caps the probe at one in flight per session.
The test suite can no longer reach the production `~/.codeman` data dir (#356, @timkjr). `test/setup.ts` now points `CODEMAN_DATA_DIR` at a throwaway directory, which is the absolute override that bypasses the suite's temporary HOME when inherited from the shell, and every test that deletes a case tree goes through a containment gate that refuses paths outside the temporary HOME. A bare suite run had overwritten a real `remote-hosts.json` with a route test's fixture. The comments around it and CLAUDE.md's testing section now name that variable as the cause; `os.homedir()` itself does follow `$HOME`.
### Thanks
- @opticon454 for the CLI registry (#347), the phased resubmission of #343, and the review rounds that hardened it.
- @timkjr for the remote auto-revive fix (#355) and the test-isolation sweep (#356).
## 1.24.5
### Patch Changes
- Fable 5.1 is selectable in App Settings.
`claude-fable-5-1` is in Claude Code's model catalog (display name "Fable 5.1", June 2026 knowledge cutoff), but the model picker only went up to Fable 5, so pinning it meant hand-editing a case's `.claude/settings.local.json`. It now appears as a card under **App Settings -> Models -> New Claude sessions**, and as an option in **Task routing** (Default for tasks, plus the Explore / Implement / Test / Review overrides).
It is offered exactly the way Fable 5 already is: the "1M capable" badge, the 1M context window switch stays live for it, and base + switch compose into `claude-fable-5-1[1m]`. Both strings are accepted by the CLI.
Deliberately not claimed: that a 1M window is what sets Fable 5.1 apart. The CLI's model catalog marks both fable entries as natively 1M with the same window, so an always-on window for 5.1 next to a switchable one for 5 would encode a difference the models do not have.
### Thanks
- @shenlvkang-collab for #370, which surfaced that Fable 5.1 was missing from the picker.
## 1.24.4
### Patch Changes
- The Compose deployment image ships the Docker CLI instead of the whole Docker engine.
`docker/server.Dockerfile` installed Debian's `docker.io` to get a client for the mounted
host socket. That package is the full **engine**: even with `--no-install-recommends` it
pulls 15 packages including containerd, runc, dmsetup and iptables, none of which a
container that only talks to a socket can use. It also ships Docker 20.10.24, from 2023.
The CLI and the buildx plugin are now copied from the official `docker:29-cli` image
instead. Measured on the same `node:22-bookworm-slim` base: **266 MB → 108 MB**, a 158 MB
saving, with the current CLI (29.7.2) in place of a two-year-old one.
Verified by building the real image and running it: the binaries are static Go builds, so
they work on this glibc image even though they come from an Alpine one, and `docker
--version`, `docker ps` and `docker build` all succeed against a mounted host socket as
the unprivileged runtime user. buildx is copied deliberately — `scripts/build-agent-image.mjs`
shells out to `docker build` and Codeman auto-builds the agent image on the first Docker
case, which without the plugin falls back to the classic builder Docker has deprecated.
`docker-compose` is not copied; Codeman never shells out to it.
## 1.24.3
### Patch Changes
- Docker Compose deployment, and the plan-usage chip stops losing its 5-hour window.
**Run Codeman itself in a container** (#349, @opticon454). `docker/` now carries a
local-image Compose deployment: copy `docker/.env.example` to `docker/.env`, set
`CODEMAN_PASSWORD`, run `bash docker/Start-Codeman.sh`. Docker cases then start as
**sibling** containers through the mounted host socket rather than nested ones, which
inverts an assumption the bare-host path takes for granted: the daemon no longer shares
Codeman's filesystem, so a bind source that is valid inside Codeman means nothing to it.
`CODEMAN_DOCKER_HOST_HOME` translates sources under HOME into the daemon's namespace and
`CODEMAN_CASES_PATH` points the cases dir at a host-absolute bind mount, so a workspace
resolves to the same absolute path on both sides. `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1`
drops `--memory-swap` for hosts without swap accounting (`--memory` still applies) and
filters only that one kernel warning. Guides: `docs/docker-compose.md`, `docker/README.md`.
Three things were fixed while landing it:
- **`docker/.env` was being baked into the image.** A `.dockerignore` pattern matches the
whole context-relative path, so the bare `.env` line excluded only the root file while
`COPY . .` picked up `docker/.env` — the file the deployment's own README tells you to
fill with `CODEMAN_PASSWORD` and provider API keys — and left it at
`/opt/codeman/docker/.env`. Now excluded via `**/.env`, verified in both directions
against a real build context with a canary secret.
- **`codeman skill install --case <name>` could not find a case under Compose.**
`CODEMAN_CASES_PATH` moved the server's cases dir but not the CLI's, which still
hardcoded `~/codeman-cases`. Both now resolve through one place.
- **A Docker case handed its Claude conversation id to every other CLI.** `resumeOnStart`
seeded `dockerResumeId` from `lastClaudeSessionId` regardless of mode, and
`appendResumeFlag()` maps a resume id onto codex/gemini/pi/grok/deepseek/omp/antigravity.
This one is a plain master bug, unrelated to Compose.
**The plan-usage chip keeps its 5-hour slot.** It silently shrank from `5h 4% · 7d 52%`
to a lone `7d 52%`, which reads as half the feature breaking. Nothing was broken: Claude
Code ships `rate_limits.five_hour` "only while the API reports it and its resets_at has
not passed", so between 5-hour session windows the key simply leaves the statusline
payload. The slot now stays with a dimmed em dash and the tooltip says "no active session
window". Claude only — a missing Codex bucket means that plan has no such limit, so those
stay omitted.
### Thanks
- @opticon454 for #349, and for a write-up that made an infrastructure PR quick to review
## 1.24.2
### Patch Changes
- Fix every new claude session dying on Claude Code 2.1.252's rewritten folder-trust dialog.
That dialog used to offer `❯ 1. Yes, I trust this folder` / `2. No, exit`, so Codeman
answered it by pressing Enter on the highlighted default. 2.1.252 dropped the numbers,
reversed the options and highlights `No, exit`, so the same Enter now answers _exit_: a
session in any directory claude had not seen before died (`Pane is dead (status 1)`)
about six seconds after it started, before the agent ever drew a composer.
- `trustDialogNextKey()` (`src/session-trust-dialog.ts`) now reads the `❯` marker off
the rendered pane and returns ONE keystroke at a time: an arrow while the cursor is on
the wrong option, Enter only once the screen shows it on the trust option. A frame it
cannot read presses nothing. Both the 2.1.252 and the older numbered layout are
handled, and the direction is derived from the frame rather than assumed, so a further
reordering costs a repaint instead of a session.
- The scan schedules its own follow-up read. It had only ever run from the PTY data
handler, which was enough while one Enter answered the dialog; the arrow that moves the
cursor is the last output the pane produces, so a two-keystroke answer would otherwise
stall with the cursor sitting on the right option forever. The keystroke cap goes from
3 to 6 for the same reason.
- The bundled `codeman` agent skill gets the same treatment (preamble 1.21.0): its
`_accept_trust` fallback reads `terminal?full=1`, steers onto the trust option and
confirms only after re-reading, instead of posting a blind `\r`. It sends those
keystrokes under its own `clientId`, because input sequence numbers are monotonic per
client and spending prompt numbers on dialog keys would make the next send-and-wait
look like a stale duplicate and vanish silently.
- Readiness recipes in `docs/extending-codeman.md`, `docs/api-reference.md` and the
skill's own reference carry the corrected answer and a new symptom-table entry for a
worker whose pane is dead seconds after the spawn.
Also included: a CLAUDE.md audit against the tree, correcting counted drift (route
modules, handler counts, frontend module count and app.js size, install.sh size) and
documenting several subsystems that had no entry.
## 1.24.1
### Patch Changes
- The Docker agent base image builds again.
**`docker/agent.Dockerfile` could not be built from a fresh checkout** (#352, fix in #350): the DeepSeek Harness step died with `dsh: pnpm not found on PATH` and exit 127, which took the whole image with it and, because Codeman auto-builds this image on the first Docker case, left Docker mode unusable on a clean host. `dsh plugin` does not bundle a package manager; it spawns a literal `pnpm` with no npm fallback, so pnpm is now installed alongside `dsh` and the layer proves it with `pnpm --version`.
The profile install also passes `--config.dangerouslyAllowAllBuilds=true`, because pnpm, unlike npm, refuses dependency lifecycle scripts by default and fails the install over it (`ERR_PNPM_IGNORED_BUILDS`, exit 1). Which packages that hits moves between rebuilds, since the terminal profile is resolved by dist-tag rather than pinned: the tree that broke the build in August pulled `@google/genai`, today's does not. An allowlist of those names would have gone stale rather than prevented the next break, and running those scripts is the same exposure the image already accepts three layers up, where `npm install -g` runs the install scripts of every transitive dependency of the five CLIs above it with no gate at all.
Documentation caught up with two things it had wrong: the image smoke test in `docs/docker-cases.md` now covers `dsh` and `omp`, and checks the dsh **profile** rather than only the binary (`dsh` is a launcher, so `dsh --version` says nothing about whether a session can start), and `docs/deepseek-integration.md` names pnpm as a prerequisite for installing a terminal profile at all, by hand or through the UI button. A comment in the `/api/deepseek/install-profile` route claimed the opposite of what this bug proved, and is corrected; the route's behaviour was already right, surfacing dsh's own "pnpm not found on PATH" line as the install error.
### Thanks
- @opticon454 for #350, with a reproduction that made this a confirmation rather than a hunt
- @timkjr for reporting #352, and for finding it while verifying Docker support for someone else's PR
## 1.24.0
### Minor Changes
- OMP (Oh My Pi) as a tenth run mode, mode-faithful Resume for external CLIs, and a cleaner plan-usage chip.
**OMP (`omp`) run mode** (#353): Oh My Pi joins Claude Code, shell, OpenCode, Codex, Gemini, Antigravity, Pi, Grok Build and DeepSeek Harness as a run mode, in local, Docker and remote-SSH sessions: toolbar dropdown, welcome button, phone overview, command palette, clone-repo brain picker, cron agent types, tab badges and per-mode colours, plus `GET /api/omp/status`, a `codeman doctor` entry, install.sh detection and the docker agent image. The resolver leads with `~/.local/bin` (the upstream installer's real target) and demands `omp/<semver>` from `--version`, so an unrelated binary with the same three-letter name is never spawned. Past omp conversations appear in Past Sessions, read from omp's own session files (the header line carries the real working directory, so nothing has to reverse-engineer omp's directory mangling), and a respawned or resumed omp session is pinned to an exact conversation with `--resume <id>` instead of omp's newest-file `--continue`. Review hardening before merge: the pin is resolved only at the moment a respawn is actually confirmed (an eager resolve on boot recovery used to alias two omp tabs in one case directory onto one conversation), candidates are verified against their own header `cwd` and claimed process-wide so siblings cannot double-pin; `OMP_*` joins the env-override allowlist and `OMP_AUTH_BROKER_URL`/`OMP_AUTH_BROKER_TOKEN` are clamped for non-granted owners in multi-user mode, the same shape as `DEEPSEEK_BASE_URL`. Known and documented: omp's own knobs are mostly `PI_*` (it is a pi fork), its default `tools.approvalMode` is `yolo`, and in-container `--resume` pinning does not reach a Docker omp pane.
**Resume keeps the row's own CLI** (#353): clicking Resume on an OpenCode, Pi, Grok, DeepSeek or OMP row used to create a plain Claude session, since the create request never carried the row's mode. Resume now relaunches in the row's own mode with that CLI's continue flag, and retires the stale row it came from so three clicks no longer leave three copies of the same name. Codex, Gemini and Antigravity rows have no continuation wired yet, so their rows are deliberately left in place. `DELETE /api/sessions/:id` accepts a persisted-only session (ownership enforced through the same helper as live lookups, 404 rather than 403 so nothing leaks) and broadcasts `session_deleted` so other tabs drop the row too.
**Plan-usage chip drops the provider label when there is only one**: a machine with only Claude limits rendered `CLAUDE 5H 60% 7D 23%`, a 46px label naming the only thing it could be. The name exists to tell two rows apart, so it now appears only when both Claude and Codex have windows; the tooltip still names the provider either way.
### Thanks
- @timkjr for #353, and for turning every review finding around within a day
## 1.23.2
### Patch Changes
- Codex plan usage in the header chip, a visible inline rename in the session sidebar, and an installer that no longer loses Tailscale access on a re-run.
**Codex plan usage in the header chip** (#346): the plan-usage chip used to show Claude's 5-hour and weekly limits without saying they were Claude's, which stops being a detail the moment you run more than one CLI. It now renders one compact row per provider, Claude above Codex, each labelled and colour-coded by how much is used up. Claude's numbers still come from Codeman's marked `statusLine.command` exporter; Codex's come from the signed-in host CLI's read-only `account/rateLimits/read` app-server request at startup and every five minutes, so credentials stay inside the CLI and no auth material reaches the browser. Only the main `codex` bucket is read (model-specific buckets such as Spark are separate limits and are deliberately excluded), and the Codex row is omitted entirely when no 5-hour or weekly window is available, rather than inventing one.
**Inline rename is visible in the session sidebar** (#345): starting a rename on a sidebar row opened a focused input you could not see. The row's ellipsis clamp was still painting over the live editor, so text and caret went in blind. The sidebar now gets the same unclamped editor layout the vertical tab rail already had. Covered by a Chromium regression test that asserts the painted `overflow` and the input's measured width, not just the class name.
**install.sh keeps Tailscale access on a re-run**: a re-run whose build failed could drop a working Tailscale binding instead of preserving it. The installer now offers Tailscale setup again on re-run rather than losing it, and the README describes the three-way network-access prompt (Tailscale / LAN / local-only) as it actually behaves.
### Thanks
- @JackStuart for #346
- @fibr for #345
- @tailong-wu for #342, whose analysis of the terminal refresh replay loop matched a fix that had landed on master a few hours earlier
## 1.23.1
### Patch Changes
- Fix a fresh-Linux install failure, and bound the browser terminal's live write queue.
**install.sh now installs a build toolchain.** Reported against a stock Ubuntu 24 server: node-pty publishes prebuilt binaries for darwin and win32 only, so on Linux it is always compiled from source during `npm install`. The installer set up Node, tmux and git but never a compiler, so a machine without `build-essential` died deep inside node-gyp with `not found: make` — which reads like an npm bug rather than a missing system package. `make`, a C++ compiler and `python3` are now checked up front exactly like git and tmux, installed per distro (apt / dnf / pacman / apk / zypper) behind the same consent prompt, and re-verified afterwards rather than assumed. If `npm install` fails anyway — including on `install.sh update` — it now names the missing tools and the command that installs them instead of leaving a node-gyp stack trace as the last word.
**Bounded live xterm backpressure** (#339): live output is now one chunk in flight at a time, released by xterm's own parse callback, so xterm's private WriteBuffer can no longer hide an unbounded backlog behind the browser's 128 KiB render cap; queued, loading and incoming bytes all count against that cap. Automatic drop recovery for a shell stays on the bounded 1 MiB tail — a 100k-line shell capture is tens of MiB, and parsing it on the main thread is the freeze the cap exists to prevent — while TUI modes still recover full history behind the existing downgrade guard. Duplicate SSE terminal events are dropped before JSON parsing while WebSocket owns terminal I/O, and recovery is single-flight per active session. Follow-up hardening: the three write-queue reset paths now also release the in-flight gate, so a parse callback that never lands cannot leave live output permanently stalled.
**File Viewer searches the workspace** (#340): the File Viewer search box now queries the server-side file search endpoint with a 250 ms debounce and strict response validation, instead of filtering only the part of the tree already loaded. Tree and search state are scoped to the active session, the hidden-file preference and independent request epochs, so a stale response cannot repaint the panel; cached-tree restoration, directory results and reset behaviour survive session switches and both panel-hide paths.
### Thanks
- @dignfei for #339
- @aakhter for #340
- 858b15e: Search the full session workspace from File Viewer while keeping results scoped to the active session and hidden-file preference.
## 1.23.0
### Minor Changes
- DeepSeek Harness as a ninth run mode, DeepSeek agent workers, and detailed rows for the vertical tab rail.
**DeepSeek Harness (`dsh`) run mode** (#337): DeepSeek's plugin-native agent framework joins Claude Code, shell, OpenCode, Codex, Gemini, Antigravity, Pi and Grok as a run mode. The harness is a profile launcher rather than an agent, so availability is two questions (binary AND a pane-capable profile): the Run button gates on both, a missing terminal profile is offered as a one-click install (`POST /api/deepseek/install-profile`, the only endpoint in Codeman that installs third-party code, fenced accordingly), and the resolver demands the harness's own help banner so Debian's unrelated `dsh` (dancer's shell) can never be spawned. Its permission switch is the `DSH_PERMISSION_MODE` env export (the harness has no bypass flag), injected via tmux setenv and clamped for non-granted owners in multi-user mode, including the env-override path. The community TUI's supervisor-reporting contract makes deepseek the first non-Claude mode with REAL lifecycle signals: a generated status shim turns its idle/working/blocked reports into definitive `stop`/`permission_prompt`/`agent_working` hook events, so dsh sessions get real respawn triggers, real wait signals and red "needs you" alerts instead of output-stabilization guesswork. The vendor's browser UI opens as a managed web tab through a background `dsh web` fenced to Codeman's origin. Docker image support included.
**DeepSeek agent workers** (#341): the codeman agent skill can spawn and drive dsh workers like claude ones — tasked, waited on and read with the same calls. `GET /api/sessions/:id/last-response` reads the harness's real zstd transcript (one frame per append; the reader walks frame boundaries itself, since a naive decode silently truncates to the first frame), distinguishes real prompts from plugin-injected context, and reports a failed turn's provider error instead of an empty answer.
**Vertical tab rail: detailed rows** (#338): the vertical rail can now show the home screen's per-session line (created stamp, state duration, status pill) via the new per-device `tabRailDetail` setting (default detailed; `simple` restores the 1.22.0 rows). One shared row model and one gate (`isRichTabRows()`) keep the rail, the rich sidebar and both home screens in agreement about what "working" means. A never-sized rail opens at the 320px Wide preset; below 288px the created stamp is dropped, below 240px rows fall back to simple. Also fixes Escape during an inline tab rename committing an empty name (the session then displayed its folder name).
**Review hardening across all three** (post-review commits on each PR): multi-user owners without the bypass grant can no longer redirect the server's forwarded `DEEPSEEK_API_KEY` via a `DEEPSEEK_BASE_URL` override; waits on `stop`/`blocked` are refused for docker/remote dsh sessions (their status bridge cannot reach the harness) and docker/remote dsh sessions keep the pane reader (their transcripts are not local); dsh approvals are alerts answered in the terminal, never blind keystrokes into a third-party TUI; the status shim forwards the contract's `--seq` token (stale retried reports are dropped server-side) and treats 4xx as permanent so a misconfigured session cannot rate-limit the hook endpoint for the whole instance; concurrent DeepSeek web-UI starts are serialized; cron deepseek jobs run the same launch gate as the HTTP paths; the installer's dsh identity probe is stdin-closed, bounded and memoized; transcript reads are memoized per (path, mtime, size) so 1s polling stops decoding unchanged files; the rail's width dialog, compact-threshold folder rows and reset affordances are rich-aware.
### Patch Changes
- b330f1d: Vertical tab rail: detailed rows, and a rename cancel that no longer wipes the name.
The vertical rail (Tab Orientation → Vertical) now draws the same per-session
line the home screen and the rich sidebar draw — when the session was created,
how long it has been in the state it is in, the folder it runs in, and a status
pill — instead of just the name. New per-device setting **Vertical Rail Rows**
(`tabRailDetail`, App Settings → Appearance → Tabs) with `Detailed` as the
default and `Simple (name only)` as the opt-out. A rail that has never been
sized now opens at 320px (the existing Wide preset) so the line fits; a narrower
rail sheds the created stamp below 288px and falls back to simple rows below
240px.
Also fixes a data-loss bug in the inline tab rename that predates the rail:
pressing Escape cleared the input and blurred it, and the blur handler commits —
so cancelling a rename stored an EMPTY session name and the tab fell back to its
folder label. Escape now cancels without a request, in every layout.
## 1.22.0
### Minor Changes
- 3f8c8e9: Add Grok Build (xAI `grok`) as a seventh CLI run mode. SessionMode gains 'grok', with its own resolver (version-probed, since the name has npm squatters; GET /api/grok/status surfaces path + version), GrokConfig (model, alwaysApprove -> --always-approve, resume/continue), GROK*\*/XAI*\* env allowlist entries, the multi-user only-if-sent bypass clamp, Docker (own image step + per-file credential seeding) and remote-SSH command defaults, cron agentType, run-mode/welcome/tab UI with a charcoal identity, and docs (grok-integration.md + plan). Verified end to end against grok 1.0.5 on an isolated instance.
- 74194e4: Add the owner-scoped tab-layout model, persistence, API, lifecycle repair, and synchronized legacy ordering foundation.
- e3a2fb7: Add an optional resizable vertical session rail with responsive layout, complete labels, accessible controls, and stable inline rename.
### Patch Changes
- Fix the file preview's dead pop-out control: a real detach button now opens the previewed file in a browser tab (raw route for PDFs/images/media/text, converted-PDF preview for docx/pptx) and the copy button reports when a preview has no text to copy instead of silently doing nothing. Review-driven hardening for the new tab features: PUT /api/session-order drops unknown ids again instead of rejecting the whole write (a session deleted inside the browser's debounce window could silently lose the user's reorder), a failed mux restore no longer blocks explicit session/webview deletion for the process lifetime (the automated stale sweep stays fail-closed), and the vertical rail gains the axis-awareness the sidebar-only predicates missed: correct drag-reorder insertion, active-tab scroll-into-view, floating windows anchored beside rail tabs, connector redraws on rail scroll, server-seeded orientation applied on first load, a pre-paint stamp so vertical mode no longer flashes through the header strip, and a 12px session-name default matching the sidebar's historical size so untouched installs are not restyled.
## 1.21.0
### Minor Changes
- **`codeman tui`: a terminal dashboard for your sessions.** For the times you are in SSH or Termius instead of a browser. The web UI remains the primary surface and bare `codeman` still prints help, so the dashboard itself is strictly additive.
Sessions are grouped NEEDS YOU / WORKING / IDLE / RECENT in the same status language as the web tabs and the phone overview, and the states come from the server (hooks, idle confirmation, the approvals inbox) over the existing HTTP/SSE API rather than being screen-scraped. That is what lets the dashboard answer a permission dialog instead of only reporting one.
- `↑↓`/`j`/`k` select; `1`-`9`, `[`/`]` and `Tab` switch between sessions
- `Enter` attaches and hands the terminal to tmux; **`F1` comes back**, one key, no modifier. Inside the pane a bar across the top carries the session strip and `Alt+1`..`Alt+9` switch without returning to the dashboard first
- `Enter` on a RECENT row resumes that conversation; on a session whose pane has died it refuses and offers `r` to resume it in a fresh pane
- `y`/`n`/digits answer the selected session's pending permission or question card (the server re-captures the pane first, so a keystroke can never land in the composer)
- `p` sends a one-line prompt without attaching, `x` kills (`y` confirms), `n` starts a session and opens straight into it
- `/` cross-session search, `g` away digest, `?` help, live preview pane, plan-usage chip in the header, a terminal bell when a new approval arrives
- `codeman tui --list` and `codeman tui <n>` are scriptable fast paths; with no server running it lists panes straight from the instance's tmux socket, attach-only, and upgrades live when the server comes back
Narrow terminals (under 72 columns, a phone SSH client) drop the preview and get a single-column layout. `NO_COLOR`, non-UTF-8 glyph fallback and a non-TTY refusal are all handled. Zero new dependencies: hand-rolled ANSI over chalk and commander. User guide: `docs/tui.md`.
**Breaking: the `sc` tmux chooser is retired.** `scripts/tmux-chooser.sh` is deleted and `install.sh` no longer creates the `tmux-chooser` symlink or the `sc` alias; it sweeps both up instead, on update and on uninstall. `codeman tui` replaces it and does the job better: `sc` numbered its entries globally but only accepted a single `[1-9]` keypress, so sessions 10+ were listed and could not be selected, and it inferred nothing about what an agent was doing. The alias cleanup is marker-owned, matching the exact line the installer wrote, so a user's own `alias sc=` for another tool is untouched.
**CLI polish that came with it.**
- New shared style kit (`src/cli-style.ts`) used across the CLI: semantic palette, glyphs, width-aware table, spinner, confirm.
- `codeman doctor` is colorized and its table is measured, so the "Antigravity CLI" label no longer pushes its row out of column. `--json` output is unchanged.
- `codeman web` no longer prints its "running at" line twice, and the server's non-loopback security warning is painted like the CLI's (chalk degrades off a TTY, so journald and `web.log` stay free of escape codes).
- Spinners on the silent up-to-30s waits in `codeman web -d`, `codeman web --stop` and `codeman service install`.
- `codeman reset` asks a real y/N confirmation on a TTY; non-interactive callers keep the old `--force` refusal.
- `codeman list` and `codeman session list` share one renderer instead of drifting copies.
- `codeman attach` is described correctly in the README (it shows an attachment card for a local file).
- `test/cli-commands.test.ts` now derives its inventory from the real commander program instead of a hand-written fixture that had drifted.
**Internal.** New `tmux -L` callers resolve the socket through `resolveTmuxSocketName()`, now exported from `config/instance.ts`, so a second process can never point a beta instance at prod's panes. CLAUDE.md and `docs/architecture-invariants.md` both record the rule.
### Thanks
The TUI went through seven rounds of beta testing over PuTTY/SSH by **@Ark0N**, which is where the way out of an attach, the session strip, the preview repaint handling and the glyph set all came from.
## 1.20.1
### Patch Changes
- Terminal input and scrollback fixes (PRs #327, #331):
- IME punctuation preserved (#327): keyCode 229 / `Process` key events are now delegated to xterm's CompositionHelper instead of being suppressed, so an active Chinese IME committing numbers and full-width punctuation (,。!? and friends) reaches the terminal correctly. The CJK input field sends the browser's committed text instead of guessing from `KeyboardEvent.key`, and the redundant Android orphan-input fallback is removed so xterm is the single input owner.
- Shell history replay bounded (#331): selecting a Shell session loads a bounded 1 MiB tail instead of replaying the entire multi-megabyte tmux scrollback on xterm's main thread; full history stays available via the explicit "Load full history" action. tmux history limits now apply correctly on both legacy tmux (global default set in the same command queue before pane creation) and tmux 3.7+ (per-pane targeting that never resizes or trims unrelated live panes). Also adds `Server-Timing` and `[TERMINAL-PERF]` timing stages for terminal loads, fixes `scrollToLastNonEmptyLine` double-counting scrollback rows, and keeps live output ordered behind snapshot replays.
### Thanks
- @dignfei for both fixes: the IME punctuation root-cause fix (#327) and the bounded shell history replay with the tmux history-limit correctness work (#331).
## 1.20.0
### Minor Changes
+68 -38
View File
File diff suppressed because one or more lines are too long
+27 -20
View File
@@ -5,7 +5,7 @@
<h2 align="center">Mission control for AI coding agents</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; Terminal - One Dashboard &bull; Any Device</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; Grok &bull; OMP &bull; Terminal - One Dashboard &bull; Any Device</em>
</p>
<p align="center">
@@ -27,7 +27,7 @@
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — parallel subagent visualization" width="900">
</p>
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, or Pi inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
Get started in one line (macOS & Linux, Windows via WSL):
@@ -42,7 +42,7 @@ codeman web
The installer asks before every system change, and re-running the same line updates in place. Full details: [Quick Start - Installation](#quick-start---installation).
- **One dashboard, six CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, Gemini, or Pi](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **One dashboard, eight CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **Truly phone-friendly** - a [touch-optimized terminal](#mobile-optimized-web-ui) with instant local echo, QR login, swipe navigation, and push notifications
- **Runs while you sleep** - [idle detection + respawn cycling](#respawn-controller) and auto-resume when a subscription limit resets, for 24+ hour unattended runs
- **See your agents think** - [live floating windows](#live-agent-visualization) for every subagent and teammate, with real-time transcripts
@@ -61,14 +61,14 @@ The installer asks before every system change, and re-running the same line upda
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it. A few things worth knowing:
This installs Node.js, tmux and a build toolchain if missing (node-pty ships no Linux prebuilds, so it compiles from source), clones Codeman to `~/.codeman/app`, and builds it. A few things worth knowing:
- **It asks first.** Every system change (package installs, AI CLI download) is prompted, and a menu at the end lets you choose: run Codeman in this terminal, install it as a background service (systemd/launchd, auto-start on boot), or don't start yet. Nothing runs in the background unless you pick it.
- **Network or local-only, your choice.** The installer asks whether the dashboard should be reachable from other devices on your network (`0.0.0.0`, the default, with a strongly recommended password prompt) or from this machine only (`127.0.0.1`, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. A bare `codeman web` started by hand still defaults to loopback.
- **How it's reachable, your choice.** The installer offers three ways to reach the dashboard: **Tailscale** (loopback bind fronted by `tailscale serve`, so you get `https://<machine>.<tailnet>.ts.net` with a real certificate and your tailnet as the login, no password needed), **any device on your network** (`0.0.0.0`, with a strongly recommended password prompt), or **this machine only** (`127.0.0.1`, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. The highlighted default reflects what is already on the machine (Tailscale when it is already in use, your existing binding on a re-run), and a bare Enter never pulls in new software. A bare `codeman web` started by hand still defaults to loopback.
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh update` and `install.sh uninstall` also exist.
- **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), or [Pi](https://pi.dev) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the six is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), or [OMP](https://github.com/can1357/oh-my-pi) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the nine is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
```bash
codeman web
@@ -82,6 +82,8 @@ codeman users add alice --admin # create the first admin account
codeman web --multiuser # named logins + per-user case spaces
```
**Prefer Docker Compose?** A local-image Compose deployment ships in `docker/`: copy `docker/.env.example` to `docker/.env`, set `CODEMAN_PASSWORD`, then run `bash docker/Start-Codeman.sh` on Linux. Codeman runs in a container and spawns Docker cases as sibling containers through the host socket. See the [Docker deployment guide](docker/README.md) for direct Compose commands, storage and networking options.
Details in [Multi-User Mode](#multi-user-mode-opt-in) below.
<details>
@@ -171,7 +173,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), or [Pi](https://pi.dev)). After installing, `http://localhost:3000` is accessible from your Windows browser.
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), or [OMP](https://github.com/can1357/oh-my-pi)). After installing, `http://localhost:3000` is accessible from your Windows browser.
</details>
@@ -253,7 +255,7 @@ Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in
| Field | What it does |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. **Add Case** creates one from scratch, links an existing folder, or clones a GitHub repo straight into one (**Clone Repo**). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, `Pi`, or `Terminal` (plain shell). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, `Pi`, `Grok`, `OMP`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Models → New Claude sessions). A soft default — `/model` still works in-session. |
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
@@ -285,7 +287,7 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
- **Phone/tablet** — the UI is fully touch-optimized; scan the desktop **QR code** to log in without typing a password.
- **Outside your network** — `./scripts/tunnel.sh start` opens a Cloudflare tunnel (set `CODEMAN_PASSWORD` first).
- **SSH** — the `sc` chooser attaches to any session from a terminal (`sc` interactive, `sc 2` quick-attach, `sc -l` list).
- **SSH** — `codeman tui` is a full-screen dashboard in the terminal (`codeman tui --list` to list, `codeman tui 2` to attach straight to one).
### 7. Operate & maintain
@@ -437,7 +439,7 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
- **Background daemon & service install** — `codeman web -d` runs the server detached with a pidfile, `~/.codeman/web.log`, and verified startup (it polls the server until it answers, so a port clash never reads as success); `codeman service install` writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew `node`, `tmux` and `claude` are actually found. Secrets are never written into unit files
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → System → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Clone a GitHub repo as a case** — paste a repository URL into **Add Case → Clone Repo** and Codeman clones it into `~/codeman-cases/<name>` and registers it as a normal case, ready to run an agent in. It preflights the URL while you type (tells you whether it can be cloned anonymously and offers the repo's real branches and tags for the optional branch/tag field), fills the case name in from the URL, and lets you pick which CLI the Run button should use. Public repositories over `https://`; Codeman never collects or stores credentials
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, or **Pi** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md) and [`docs/pi-integration.md`](docs/pi-integration.md)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, **Pi**, **Grok**, or **OMP** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*` vs `GROK_*`/`XAI_*` vs `OMP_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md), [`docs/pi-integration.md`](docs/pi-integration.md), [`docs/grok-integration.md`](docs/grok-integration.md) and [`docs/omp-integration.md`](docs/omp-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
@@ -460,7 +462,7 @@ Run a case inside its own hardened Docker container instead of directly on your
- **Shared per-case container** — many sessions can `docker exec` into the same container; killing one session never tears the container out from under the others.
- **Hardened by default** — non-root, `--cap-drop ALL`, `no-new-privileges`, PID/memory caps, never `--privileged` or the docker socket; a **sealed** profile (no host credentials, network off) is one toggle away.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / Pi logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / OMP logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Durable** — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.
Prerequisite: just Docker (or Podman). The agent base image builds itself automatically on first use, with progress streamed to the UI (or pre-build it with `node scripts/build-agent-image.mjs`). Full guide: [`docs/docker-cases.md`](docs/docker-cases.md).
@@ -658,17 +660,19 @@ These run for **every** request — before auth, even on the default no-password
---
## SSH Alternative (`sc`)
## Terminal UI (`codeman tui`)
If you prefer SSH (Termius, Blink, etc.), the `sc` command is a thumb-friendly session chooser:
A full-screen dashboard for your sessions, in the terminal. Same states as the web UI, because it is a client of the same server:
```bash
sc # Interactive chooser
sc 2 # Quick attach to session 2
sc -l # List sessions
codeman tui # the dashboard
codeman tui --list # numbered session list, then exit (scriptable)
codeman tui 2 # attach straight to session 2 of that list
```
Single-digit selection (1-9), color-coded status, token counts, auto-refresh. Detach with `Ctrl+A D`.
Sessions are grouped **NEEDS YOU → WORKING → IDLE → RECENT**, longest-waiting first. `↑↓`/`j`/`k` select, `1`-`9` and `[`/`]` switch between sessions, `Enter` attaches into the tmux pane (**`F1`** to come back). Inside a pane the bar across the top keeps the session strip visible and `Alt+1`-`Alt+9` switch without leaving. `y`/`n`/digit answer a pending permission dialog right from the list, `p` sends a one-line prompt, `n` starts a session and opens straight into it, `x` kills one (`y` confirms), `/` searches, `g` shows the away digest, `?` is help, `q` quits. Below 72 columns it drops the preview pane and becomes a single-column list, so it stays usable in Termius on a phone. With no server running it still starts in attach-only degraded mode.
The web UI remains the primary surface; see **[docs/tui.md](docs/tui.md)** for the full guide.
---
@@ -687,6 +691,7 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move active tab left / right |
| `Ctrl/Cmd+C` | Copy selection, or interrupt when nothing is selected |
| `Ctrl+Shift+C` | Copy selection (never interrupts) |
| `Ctrl/Cmd+V` | Paste, or upload a clipboard image and paste its path |
| `Ctrl/Cmd+L` | Clear terminal |
| `Ctrl+Shift+R` | Restore terminal size |
| `Ctrl+Shift+V` | Toggle voice input |
@@ -793,7 +798,7 @@ When a CLI runs in a Codeman-managed session, these environment variables are se
5. **`/api/v1/*`** is a stable alias of `/api/*`.
6. **Wait instead of polling, and don't treat a timeout as an error.** The wait endpoints answer with HTTP `200` and `wait.timedOut: true` when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. `wait.timeoutMs` tells you the timeout the server actually applied after clamping (600s ceiling).
7. **Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity/pi) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
8. **Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
7. **Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity/omp) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.8. **Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
### Recipes
@@ -893,7 +898,9 @@ codeman session start -d /path/to/repo # (s) start a session
codeman session list # list sessions
codeman session logs <id> # tail output
codeman task add "fix the failing test" # (t) queue a task
codeman attach <path> # attach a Claude hook context
codeman attach <path> # show an attachment card for a local file
codeman tui --list # numbered session list (plain text when piped)
codeman tui 3 # attach to session 3 of that list
```
### Hooks (events flowing _back_ to Codeman)
@@ -1007,7 +1014,7 @@ flowchart TB
subgraph External["External"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / Pi</small>"]
BG["Background Agents<br/><small>(Task tool)</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / OMP</small>"] BG["Background Agents<br/><small>(Task tool)</small>"]
end
end
+6 -20
View File
@@ -5,7 +5,7 @@
<h2 align="center">AI 编程智能体的任务控制中心</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; Grok &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
</p>
<p align="center">
@@ -58,7 +58,7 @@ curl -fsSL https://getcodeman.com/install | bash
- **重跑即更新。** 再次运行同一条命令即可原地更新已完成的安装:`~/.codeman/app` 中的本地改动会被 stash(绝不丢弃),运行中的服务会自动重启并校验。若首次安装中途失败,重跑会继续完成完整的安装流程。也可以使用 `install.sh update` 与 `install.sh uninstall`。
- **CI / 无终端环境:** 没有终端时,涉及系统改动的步骤会带着说明中止,而不是静默执行;在自动化场景设置 `CODEMAN_NONINTERACTIVE=1` 即可批准这些步骤。
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli) 或 [Pi](https://pi.dev)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这六个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli)、[Pi](https://pi.dev)、[Grok Build](https://github.com/xai-org/grok-build)、[DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) 或 [OMP](https://github.com/can1357/oh-my-pi)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这九个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
```bash
codeman web
@@ -141,7 +141,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli) 或 [Pi](https://pi.dev))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli)、[Pi](https://pi.dev)、[Grok Build](https://github.com/xai-org/grok-build)、[DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) 或 [OMP](https://github.com/can1357/oh-my-pi))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
</details>
@@ -221,7 +221,7 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
| 字段 | 作用 |
| ---------------------- | ------------------------------------------------------------------------------------------- |
| **工作目录 / case** | 智能体操作的文件夹。「case」就是一个 Codeman 记住的命名工作目录。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini`、`Pi` 或 `Terminal`(普通 shell)。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini`、`Pi`、`Grok` 或 `Terminal`(普通 shell)。 |
| **模型** | 每会话模型(App Settings → Claude Model)。软默认值 —— 会话内 `/model` 依然有效。 |
| **Effort / Ultracode** | 推理力度(`low`–`max`),或用 `ultracode` 开启动态多智能体工作流。随时可用 `/effort` 切换。 |
@@ -253,7 +253,7 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
- **手机/平板** —— UI 完全触控优化;扫描桌面上的**二维码**即可免密码登录。
- **网络之外** —— `./scripts/tunnel.sh start` 打开一条 Cloudflare 隧道(先设置 `CODEMAN_PASSWORD`)。
- **SSH** —— `sc` 选择器可从终端附着任意会话(`sc` 交互式,`sc 2` 快速附着,`sc -l` 列表)。
- **SSH** —— `codeman tui` 是终端里的全屏会话面板(`codeman tui --list` 列出,`codeman tui 2` 直接附着到某个会话)。
### 7. 运维与维护
@@ -394,7 +394,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
## 更多特性
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity**、**Gemini** 或 **Pi**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*`、`PI_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md) 与 [`docs/pi-integration.md`](docs/pi-integration.md)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity**、**Gemini**、**Pi** 或 **Grok**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*`、`PI_*`、`GROK_*`/`XAI_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)、[`docs/pi-integration.md`](docs/pi-integration.md) 与 [`docs/grok-integration.md`](docs/grok-integration.md)
- **Docker 会话** —— 在隔离且加固的容器中运行案例。**Create New** 上勾选一个复选框即可用合理的默认值启动容器并在其中启动智能体;同一案例的多个会话共享一个容器;可将容器连同工作区导出为可移植的 `.tar.gz`,迁移到另一台机器。详见 [`docs/docker-cases.md`](docs/docker-cases.md)
- **远程 SSH 会话**:把案例指向另一台机器,让智能体在那里一个持久的远程 tmux 中运行:SSH 断连不中断任务、自动重连,还能发现并附着主机上已在运行的会话。详见 [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
@@ -615,20 +615,6 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
---
## SSH 替代方案(`sc`)
如果你更喜欢 SSH(Termius、Blink 等),`sc` 命令是一个便于拇指操作的会话选择器:
```bash
sc # 交互式选择器
sc 2 # 快速附着到会话 2
sc -l # 列出会话
```
单数字选择(1–9)、颜色编码的状态、token 计数、自动刷新。用 `Ctrl+A D` 分离。
---
## 键盘快捷键
> Ctrl 绑定在 macOS 上也接受 Cmd。
+1
View File
@@ -4,6 +4,7 @@
"scripts/*.mjs",
"scripts/*.js",
"scripts/watch-subagents.ts",
"scripts/pr-bot/main.ts",
"scripts/remotion/Root.tsx",
"scripts/remotion/index.ts",
"test/**/*.test.ts",
+4
View File
@@ -19,10 +19,14 @@
* why these are a runnable suite (`npm run test:browser`) rather than skipped.
*/
export const BROWSER_TEST_GLOBS = [
'test/tab-rail-resize.browser.test.ts',
'test/session-sidebar-ux.browser.test.ts',
'test/session-options-responsive.browser.test.ts',
'test/inline-rename.test.ts',
'test/opencode-resize.test.ts',
'test/webgl-fallback.test.ts',
'test/terminal-copy-shortcut.test.ts',
'test/terminal-keycode229-recovery.browser.test.ts',
'test/codex-predictive-echo.test.ts', // also needs a real codex binary
];
+11
View File
@@ -0,0 +1,11 @@
{
"extends": "../tsconfig.json",
"compilerOptions": {
"rootDir": "..",
"noEmit": true,
"declaration": false,
"declarationMap": false,
"sourceMap": false
},
"include": ["../scripts/pr-bot/**/*.ts"]
}
+76
View File
@@ -0,0 +1,76 @@
# =============================================================================
# Codeman Docker Compose environment template
# Copy this file to .env and set the values for the Docker host.
# =============================================================================
TZ=Australia/Perth
# Optional overrides for direct `docker compose` use. The Bash start script
# detects these values from CODEMAN_APPDATA_PATH automatically. Compose uses
# 1000:1000 when the variables are omitted.
# PUID=1000
# PGID=1000
# Name of the account that runs Codeman and all local CLI sessions. Changing
# this value rebuilds the image with a matching account.
CODEMAN_RUNTIME_USER=opencode
# Required. Persistent Codeman application data, CLI credentials, and session
# state are stored here on the host and mounted at the runtime account's home
# directory in the container.
CODEMAN_APPDATA_PATH=/mnt/user/appdata/Coding/codeman
# Optional. Absolute host path of this Codeman checkout, mounted at
# /opt/codeman so App Settings -> Updates can update Codeman in place. The Bash
# start script detects it from the compose file's own location, so it only needs
# setting for direct `docker compose` use or a checkout kept elsewhere. Point it
# at a directory that is not a git checkout and in-app updates are unavailable.
# CODEMAN_REPO_PATH=/mnt/user/appdata/Coding/codeman/app
# Required for Docker cases. This must be an absolute path on the Docker host.
# Codeman and each isolated case use this same path, so it cannot be a
# container-only path such as /home/opencode/codeman-cases.
CODEMAN_CASES_PATH=/mnt/user/appdata/Coding/codeman/codeman-cases
# Required. Network bind address, host port, and local image tag.
CODEMAN_HOST=0.0.0.0
CODEMAN_PORT=3000
CODEMAN_IMAGE=codeman:local
# Required for any network-accessible Codeman instance. Use a unique, strong
# password. This file is safe to commit; copy it to .env and set the value.
CODEMAN_PASSWORD=changeme
# Required. Username for Codeman HTTP Basic authentication.
CODEMAN_USERNAME=admin
# Optional: authenticate Gemini CLI without an interactive login.
GEMINI_API_KEY=
# Linux default. On Docker Desktop, use the socket path supported by your
# Docker installation when it differs from /var/run/docker.sock.
DOCKER_SOCKET=/var/run/docker.sock
# Optional override for direct `docker compose` use. The Bash start script
# detects this from DOCKER_SOCKET automatically. The direct Compose default is
# 999, but the correct value depends on the Docker host.
# DOCKER_SOCKET_GID=999
# Set to 1 only when Docker-case hook callbacks are required.
CODEMAN_DOCKER_BRIDGE_HOOKS=0
# Set to 1 when `docker info` reports `SwapLimit=false`. The case memory limit
# remains active; Codeman omits --memory-swap and filters the daemon's exact
# unsupported-swap warning while preserving all other Docker create errors.
CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=0
# Required only when applying the macvlan example in README.md.
CODEMAN_MACVLAN_NETWORK=br0.11
CODEMAN_IPV4_ADDRESS=10.10.11.236
CODEMAN_MAC_ADDRESS=02:10:11:00:00:EC
# Required only when creating a new managed macvlan network, rather than using
# the external-network macvlan example.
CODEMAN_MACVLAN_PARENT=br0.11
CODEMAN_MACVLAN_SUBNET=10.10.11.0/24
CODEMAN_MACVLAN_GATEWAY=10.10.11.1
+108
View File
@@ -0,0 +1,108 @@
# Codeman Docker deployment
This folder contains the Compose configuration, server image Dockerfile, and environment template for a locally built Codeman server.
## Start
From the repository root, create the runtime environment file and set the required values, especially `CODEMAN_PASSWORD`.
```sh
cp docker/.env.example docker/.env
bash docker/Start-Codeman.sh
```
On PowerShell, use the following command instead.
```powershell
Copy-Item docker/.env.example docker/.env
docker compose --env-file docker/.env -f docker/docker-compose.yaml up --build -d
```
Every required value is defined and explained in `.env.example`. `GEMINI_API_KEY` is intentionally optional and may remain blank.
On Linux, `Start-Codeman.sh` stops with an error when required paths are missing. It creates the application-data directory when safe, detects its numeric owner as `PUID:PGID`, and detects `DOCKER_SOCKET_GID` from the configured Docker socket. It rejects a root-owned application-data directory because Codeman and its local CLI sessions must remain unprivileged.
Codeman, Claude, OpenCode, and other local sessions run as the unprivileged account named by `CODEMAN_RUNTIME_USER`, which defaults to `opencode`. When Compose is run directly, `PUID` and `PGID` default to `1000:1000`; set them in `.env` when the application-data directory has a different owner. The Bash start script determines them automatically instead.
To retain Docker-case support without root when running Compose directly, set `DOCKER_SOCKET_GID` to the numeric group ID of the host socket. On a standard Linux Docker host, obtain it with `stat -c '%g' /var/run/docker.sock`. The Bash start script detects it automatically.
## Updating
Use **App Settings → Updates** in the web UI. The checkout Compose builds from is
also mounted at `/opt/codeman`, so an update's `git checkout` and rebuild persist
on the host, and the server exiting is what restarts the container onto the new
build.
Releases that change `server.Dockerfile`, `docker-compose.yaml`, or add a key to
`.env.example` cannot be applied that way — the updater detects them, names what
changed, and asks you to run `Start-Codeman.sh` here on the host instead. Details:
[`../docs/docker-self-update.md`](../docs/docker-self-update.md).
## Application data storage
The default configuration uses a host-folder bind mount:
```yaml
volumes:
- type: bind
source: ${CODEMAN_APPDATA_PATH}
target: /home/${CODEMAN_RUNTIME_USER}
```
Set `CODEMAN_APPDATA_PATH` in `.env` to a directory that the Docker daemon can access. The example value is `/mnt/user/appdata/Coding/codeman`.
`CODEMAN_CASES_PATH` is the separate host directory for managed case workspaces. It is mounted into Codeman at the same absolute path, allowing the host Docker daemon to bind it into an isolated case container. Set it to a child directory of `CODEMAN_APPDATA_PATH` unless you deliberately store workspaces elsewhere.
Compose also exposes `CODEMAN_APPDATA_PATH` to Codeman as `CODEMAN_DOCKER_HOST_HOME`. This lets Docker case seed files, CLI credentials and the hook secret be mounted using paths that exist in the host daemon's filesystem. Direct host installations do not set this variable and retain their existing behaviour.
Set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` when `docker info` reports `SwapLimit=false`. Codeman continues to apply the configured case memory limit, omits Docker's unsupported `--memory-swap` option, and filters only the daemon's exact swap-capability warning. Every other Docker create error and its exit status remain visible.
For an existing installation created by a root-running image, change ownership of the application-data directory before upgrading so the configured `PUID` and `PGID` can read the saved credentials and state:
```sh
chown -R 99:100 /mnt/user/appdata/Coding/codeman
```
Replace `99:100` and the path with the values from your `.env` file.
Do not replace this bind mount with a Docker-managed named volume when Docker cases are enabled. Codeman passes seed, credential, transcript and hook-secret bind sources to the host Docker daemon, so their source files must have stable paths in the daemon's filesystem. A named volume does not provide the required host path mapping.
## Static macvlan networking
The default configuration publishes a host port. It does not use `network_mode: host`. To attach Codeman directly to an existing external macvlan network with a static IP address and MAC address, remove the `ports:` section and add the following to the `codeman` service:
```yaml
mac_address: ${CODEMAN_MAC_ADDRESS}
networks:
codeman_lan:
ipv4_address: ${CODEMAN_IPV4_ADDRESS}
```
Then add this top-level network declaration:
```yaml
networks:
codeman_lan:
external: true
name: ${CODEMAN_MACVLAN_NETWORK}
```
Set `CODEMAN_MACVLAN_NETWORK`, `CODEMAN_IPV4_ADDRESS`, and `CODEMAN_MAC_ADDRESS` in `.env`. The values in `.env.example` match the supplied Unraid example network and should be changed for other hosts.
### Create a managed macvlan network
If an external macvlan network does not already exist, use this top-level declaration instead. Do not use it together with the external-network declaration.
```yaml
networks:
codeman_lan:
driver: macvlan
driver_opts:
parent: ${CODEMAN_MACVLAN_PARENT}
ipam:
config:
- subnet: ${CODEMAN_MACVLAN_SUBNET}
gateway: ${CODEMAN_MACVLAN_GATEWAY}
```
Macvlan containers are ordinarily not reachable from their Docker host without additional host-network routing. Confirm the selected address, MAC address, parent interface, and subnet are reserved and valid for the target network before starting the stack.
+129
View File
@@ -0,0 +1,129 @@
#!/usr/bin/env bash
set -euo pipefail
script_dir=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
env_file="$script_dir/.env"
compose_file="$script_dir/docker-compose.yaml"
if [[ ! -f "$env_file" ]]; then
printf 'Error: Docker environment file is missing: %s\n' "$env_file" >&2
printf 'Create it from %s/.env.example before starting Codeman.\n' "$script_dir" >&2
exit 1
fi
compose_command=(docker compose --env-file "$env_file" -f "$compose_file")
appdata_path=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "CODEMAN_APPDATA_PATH" { sub(/^[^=]*=/, ""); print; exit }'
)
docker_socket=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "DOCKER_SOCKET" { sub(/^[^=]*=/, ""); print; exit }'
)
if [[ -z "$appdata_path" ]]; then
printf 'Error: CODEMAN_APPDATA_PATH is not set in %s\n' "$env_file" >&2
exit 1
fi
if [[ ! -d "$appdata_path" ]]; then
if [[ "$EUID" == '0' ]]; then
printf 'Error: Refusing to create CODEMAN_APPDATA_PATH as root: %s\n' "$appdata_path" >&2
printf 'Create it as the unprivileged account that should run Codeman, then retry.\n' >&2
exit 1
fi
mkdir -p -- "$appdata_path"
fi
if owner_ids=$(stat -c '%u:%g' -- "$appdata_path" 2>/dev/null); then
:
elif owner_ids=$(stat -f '%u:%g' "$appdata_path" 2>/dev/null); then
:
else
printf 'Error: Cannot determine the owner of CODEMAN_APPDATA_PATH: %s\n' "$appdata_path" >&2
exit 1
fi
export PUID=${owner_ids%%:*}
export PGID=${owner_ids##*:}
if [[ "$PUID" == '0' ]]; then
printf 'Error: CODEMAN_APPDATA_PATH is owned by root: %s\n' "$appdata_path" >&2
printf 'Change the directory ownership to the unprivileged account that should run Codeman.\n' >&2
exit 1
fi
if [[ -z "$docker_socket" || ! -S "$docker_socket" ]]; then
printf 'Error: DOCKER_SOCKET is not a Unix socket: %s\n' "${docker_socket:-<unset>}" >&2
exit 1
fi
if socket_ids=$(stat -c '%u:%g' -- "$docker_socket" 2>/dev/null); then
:
elif socket_ids=$(stat -f '%u:%g' "$docker_socket" 2>/dev/null); then
:
else
printf 'Error: Cannot determine the owner of DOCKER_SOCKET: %s\n' "$docker_socket" >&2
exit 1
fi
export DOCKER_SOCKET_GID=${socket_ids##*:}
repo_path=${CODEMAN_REPO_PATH:-$(cd -- "$script_dir/.." && pwd)}
if [[ ! -d "$repo_path" ]]; then
printf 'Error: CODEMAN_REPO_PATH is not a directory: %s\n' "$repo_path" >&2
exit 1
fi
export CODEMAN_REPO_PATH="$repo_path"
# The in-app updater runs `git checkout` and `npm install` against this checkout
# as PUID:PGID. If the directory belongs to someone else, git refuses outright
# ("detected dubious ownership") and the update fails at the first step — so warn
# here, where the fix is obvious, rather than in a failed update hours later.
if repo_owner=$(stat -c '%u' -- "$repo_path" 2>/dev/null || stat -f '%u' "$repo_path" 2>/dev/null); then
if [[ "$repo_owner" != "$PUID" ]]; then
printf 'Warning: %s is owned by UID %s but Codeman runs as UID %s.\n' "$repo_path" "$repo_owner" "$PUID" >&2
printf 'In-app updates will fail until the ownership matches. Codeman itself still starts.\n' >&2
fi
fi
if [[ ! -d "$repo_path/.git" ]]; then
printf 'Note: %s is not a git checkout, so in-app updates are unavailable.\n' "$repo_path" >&2
fi
# Record what the container is about to be built and created FROM. The in-app
# updater compares these against the release it wants to apply: a release that
# changes either file cannot be applied by the container restarting itself (a
# restart reuses the existing image and config), so it is refused and the user
# is sent back here. Written on every start, so the baseline always describes
# the container that is actually running. See docs/docker-self-update.md.
if command -v sha256sum >/dev/null 2>&1; then
sha256_of() { sha256sum -- "$1" | cut -d' ' -f1; }
elif command -v shasum >/dev/null 2>&1; then
sha256_of() { shasum -a 256 -- "$1" | cut -d' ' -f1; }
else
sha256_of() { printf ''; }
fi
dockerfile_sha=$(sha256_of "$script_dir/server.Dockerfile")
compose_sha=$(sha256_of "$compose_file")
if [[ -n "$dockerfile_sha" && -n "$compose_sha" ]]; then
# $CODEMAN_APPDATA_PATH is mounted at the runtime account's home, so this is
# dataPath('docker-env-applied.json') as the server inside the container sees it.
state_dir="$appdata_path/.codeman"
mkdir -p -- "$state_dir"
printf '{\n "dockerfileSha256": "%s",\n "composeSha256": "%s"\n}\n' \
"$dockerfile_sha" "$compose_sha" >"$state_dir/docker-env-applied.json.tmp"
mv -- "$state_dir/docker-env-applied.json.tmp" "$state_dir/docker-env-applied.json"
# A root-run start (common on Unraid) would otherwise leave a root-owned
# `.codeman` on a FIRST start, before the container has created it as PUID,
# and the unprivileged server could then never write its own state there.
if [[ "$EUID" == '0' ]]; then
chown -- "$PUID:$PGID" "$state_dir" "$state_dir/docker-env-applied.json"
fi
else
printf 'Warning: no sha256 tool found; in-app updates will not detect environment changes.\n' >&2
fi
exec docker compose --env-file "$env_file" -f "$compose_file" up --build -d
+80 -2
View File
@@ -51,6 +51,58 @@ RUN npm install -g --ignore-scripts @earendil-works/pi-coding-agent \
&& npm cache clean --force \
&& pi --version
# Grok Build (`grok`, xAI) is NOT on npm: a standalone ~160MB Rust binary through
# xAI's installer, which targets $HOME/.grok/bin with no --dir override. At build
# time that is root's home and unreachable by the `agent` user, so copy the binary
# into /usr/local/bin and drop root's ~/.grok in the same layer so the image does
# not carry the download twice. The staging cp -T is what makes this survive the
# installer's own behavior EITHER way: newer installers already symlink
# /usr/local/bin/grok -> /root/.grok/bin/grok, and a direct `cp -L` onto that
# symlink fails with "same file" (2026-08-24 rebuild), while removing the link
# first and copying fresh works for both old and new installers.
RUN curl -fsSL https://x.ai/cli/install.sh | bash \
&& cp -L /root/.grok/bin/grok /usr/local/bin/grok.real \
&& rm -f /usr/local/bin/grok \
&& mv /usr/local/bin/grok.real /usr/local/bin/grok \
&& chmod 755 /usr/local/bin/grok \
&& rm -rf /root/.grok /root/.local/bin/grok /root/.local/bin/agent \
&& grok --version
# DeepSeek Harness (`dsh`). A normal npm package, but the ONLY entry here whose
# binary runs nothing on its own: `dsh` is a profile launcher, and DeepSeek ships
# only `web` and `headless`, so without an interactive profile a
# `mode: 'deepseek'` container would start a pane that dies on arrival. The
# profile itself is installed further down, into the `agent` HOME, because
# Codeman deliberately does NOT seed `profiles/` from the host: it is a
# per-profile node_modules tree, host-arch-specific and far too large to copy on
# every container start.
# ⚠️ `pnpm` is a HARD dependency of `dsh plugin`, not optional tooling: the
# subcommand is a thin forwarder that `spawnSync`s a literal `pnpm` with no
# fallback to npm, so on an image without it the profile install below dies
# with `dsh: pnpm not found on PATH` / exit 127 and takes the whole build with
# it (issue #352). It stays on PATH at runtime too, so a container user can run
# `dsh plugin add` themselves.
RUN npm install -g @deepseek-ai/dsh pnpm \
&& npm cache clean --force \
&& dsh --version \
&& pnpm --version
# OMP (Oh My Pi) is NOT on npm: a standalone binary via omp.sh's installer, which
# targets $HOME/.local/bin with no --dir override (verified 2026-08-27 — the
# resolver's OMP_SEARCH_DIRS lists ~/.omp/bin first, which turned out to be the
# WRONG guess for the installer's actual target; build this step for real
# rather than trust that ordering). At build time $HOME is root's home and
# unreachable by the `agent` user, so copy the binary into /usr/local/bin and
# drop root's ~/.local/bin/omp in the same layer so the image does not carry
# the download twice.
RUN curl -fsSL https://omp.sh/install | sh \
&& cp -L /root/.local/bin/omp /usr/local/bin/omp.real \
&& rm -f /usr/local/bin/omp \
&& mv /usr/local/bin/omp.real /usr/local/bin/omp \
&& chmod 755 /usr/local/bin/omp \
&& rm -f /root/.local/bin/omp \
&& omp --version
# `agent` user (gid 0) with an arbitrary-uid-writable HOME. The uid is
# auto-assigned (node:22-slim already occupies uid 1000 with its `node` user); at
# runtime Codeman overrides with `--user <hostUid>:0` on Linux, so the baked uid
@@ -68,11 +120,37 @@ ENV HOME=/home/agent
# transcript/rollout dirs (`.claude/projects`, `.codex/sessions`) are bind-mounted from
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir;
# Antigravity nests its state inside `.gemini/antigravity-cli`, so it rides that seed.)
# `.pi/agent` IS pre-created: pi is seeded per-FILE (auth/settings/trust/models), and a
# `.pi/agent` and `.grok` ARE pre-created: both are seeded per-FILE (pi:
# auth/settings/trust/models; grok: auth.json/config.toml/pager.toml), and a
# per-file seed copy, unlike a whole-dir one, does not create its parent directory.
# `.dsh` is pre-created for the same per-file reason (.env/settings.yaml/
# cordis.patch.yml), and the interactive profile is built into it HERE rather than
# after `USER agent`: this layer's closing chgrp/chmod is what makes the whole tree
# writable by the arbitrary uid the container actually runs as, and a profile
# installed after it would miss that fixup. DSH_HOME points the launcher at the
# agent's dir while this still runs as root.
# ⚠️ `dangerouslyAllowAllBuilds` is what keeps that profile install from becoming
# the next #352. pnpm (unlike npm) blocks dependency lifecycle scripts by default
# and FAILS the install over it — `ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on
# pnpm 11.24 — so any package in the tui's tree that ships one stops the build
# dead. An allowlist of the offenders rots: `@deepseek-harness-tui/dsh-tui` is
# resolved by dist-tag, not pinned, and 0.9.3 pulled `@google/genai` (a
# `preinstall: no-op`) where 0.10.0-beta.x does not, so the names to allow move
# under us between rebuilds. Allowing them wholesale is also the SAME exposure
# this image already accepts three layers up: `npm install -g` runs the install
# scripts of every transitive dep of the five CLIs above it, with no gate at all.
# `.omp/agent` is pre-created for the same reason `.codex` is: it is a MIXED
# store (per-file config seeds PLUS a shared `sessions/` RW bind mount for
# Codeman's own host-side history/resume reads), and neither kind of artifact
# creates its own parent directory.
RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
&& mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \
/home/agent/.claude/projects /home/agent/.codex/sessions /home/agent/.pi/agent \
/home/agent/.claude/projects /home/agent/.codex/sessions /home/agent/.pi/agent /home/agent/.grok \
/home/agent/.dsh /home/agent/.omp/agent \
&& DSH_HOME=/home/agent/.dsh HOME=/home/agent \
dsh plugin --profile dsh-tui add --config.dangerouslyAllowAllBuilds=true \
@deepseek-harness-tui/dsh-tui \
&& test -f /home/agent/.dsh/profiles/dsh-tui/package.json \
&& chgrp -R 0 /home/agent \
&& chmod -R g=u /home/agent
+111
View File
@@ -0,0 +1,111 @@
name: codeman
services:
codeman:
build:
context: ..
dockerfile: docker/server.Dockerfile
args:
CODEMAN_RUNTIME_USER: ${CODEMAN_RUNTIME_USER}
PGID: ${PGID:-1000}
PUID: ${PUID:-1000}
image: ${CODEMAN_IMAGE}
init: true
restart: unless-stopped
ports:
- "${CODEMAN_PORT}:${CODEMAN_PORT}"
environment:
# Tells the self-updater to restart by exiting (the restart policy below
# relaunches it) rather than by looking for an init system that is not
# here. Also set in the image; repeated so a container started without the
# image default still self-identifies.
CODEMAN_IN_CONTAINER: "1"
# This file sets `restart: unless-stopped` below, so the updater may restart
# the server by EXITING. Declared here and only here, never in the image: a
# container started by plain `docker run` has no restart policy unless the
# operator gave it one, and there the updater asks the daemon instead and
# stages the update for a manual restart when it cannot get an answer.
CODEMAN_RESTART_BY_EXIT: "1"
CODEMAN_DOCKER_BRIDGE_HOOKS: ${CODEMAN_DOCKER_BRIDGE_HOOKS}
# Host-side equivalent of the runtime user's HOME. Docker case seed,
# credential and hook mounts are translated into the daemon namespace.
CODEMAN_DOCKER_HOST_HOME: ${CODEMAN_APPDATA_PATH}
CODEMAN_DOCKER_DISABLE_SWAP_LIMIT: ${CODEMAN_DOCKER_DISABLE_SWAP_LIMIT}
CODEMAN_CASES_PATH: ${CODEMAN_CASES_PATH}
CODEMAN_HOST: ${CODEMAN_HOST}
CODEMAN_PASSWORD: ${CODEMAN_PASSWORD}
CODEMAN_PORT: ${CODEMAN_PORT}
CODEMAN_USERNAME: ${CODEMAN_USERNAME}
GEMINI_API_KEY: ${GEMINI_API_KEY}
PGID: ${PGID:-1000}
PUID: ${PUID:-1000}
TZ: ${TZ}
group_add:
# Retain access to the host Docker socket without running as root.
- ${DOCKER_SOCKET_GID:-999}
volumes:
# Application data and CLI credentials persist on the configured host
# path, rather than in a Docker-managed volume.
- type: bind
source: ${CODEMAN_APPDATA_PATH}
target: /home/${CODEMAN_RUNTIME_USER}
# Docker cases are sibling containers on the host daemon. Their workspace
# must be visible to Codeman at the same absolute path used by that daemon.
- type: bind
source: ${CODEMAN_CASES_PATH}
target: ${CODEMAN_CASES_PATH}
# Codeman uses the host daemon to create isolated Docker cases. This is
# Docker-outside-of-Docker, not Docker-in-Docker.
- type: bind
source: ${DOCKER_SOCKET}
target: /var/run/docker.sock
# The application source, so App Settings -> Updates can update in place.
# This is the SAME checkout used as the build context above, mounted over
# the image's baked copy: a `git checkout` performed inside the container
# then lands on the host and survives the container being recreated.
# Without it the pull would go to the container's writable layer and be
# silently discarded by the next `up`. See docs/docker-self-update.md.
# Defaults to `..` — the build context above — which Compose resolves
# against the project directory, so plain `docker compose up` works with
# no extra configuration. Set CODEMAN_REPO_PATH only to point elsewhere.
- type: bind
source: ${CODEMAN_REPO_PATH:-..}
target: /opt/codeman
# Build artefacts live in named volumes layered OVER the repo bind mount,
# so `npm install` and `npm run build` inside the container never write
# into the host checkout. That keeps container-compiled native modules
# (node-pty is built from source here) out of a checkout that may also be
# used to run Codeman natively, and keeps `git status` clean. Docker seeds
# an EMPTY named volume from the image, so the first start inherits the
# image's already-built node_modules and dist rather than paying for a
# bootstrap build.
- type: volume
source: codeman-node-modules
target: /opt/codeman/node_modules
- type: volume
source: codeman-dist
target: /opt/codeman/dist
extra_hosts:
- "host.docker.internal:host-gateway"
security_opt:
- no-new-privileges:true
cap_drop:
- ALL
healthcheck:
test:
- CMD-SHELL
- >-
node -e "fetch('http://127.0.0.1:${CODEMAN_PORT}/api/status').then((response) => process.exit(response.status < 500 ? 0 : 1)).catch(() => process.exit(1))"
interval: 30s
timeout: 5s
retries: 3
start_period: 30s
volumes:
# Container-owned build artefacts. They persist across container recreation,
# so an in-app update's `npm install` output is not thrown away by the next
# `up`, and they are seeded from the image on first use. Removing them (or
# `docker compose down -v`) is the supported reset: the next start rebuilds
# from the image.
codeman-node-modules:
codeman-dist:
+142
View File
@@ -0,0 +1,142 @@
# syntax=docker/dockerfile:1
# Build the application from the checkout supplied as the Docker build context.
# No published Codeman application image is required.
FROM node:22-bookworm-slim AS build
RUN apt-get update \
&& apt-get install -y --no-install-recommends python3 make g++ \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /opt/codeman
COPY . .
# devDependencies are deliberately KEPT (no `npm prune --omit=dev`). The in-app
# updater rebuilds from inside this container, and `npm run build` is tsc +
# esbuild — both devDependencies. Pruning them saves image size and takes the
# self-updater with it. See docs/docker-self-update.md.
RUN npm ci \
&& npm run build \
&& npm cache clean --force
# The Docker CLI talks to the host daemon through the socket mounted by
# docker/docker-compose.yaml. It does not run a Docker daemon in this container.
FROM node:22-bookworm-slim
ARG CODEMAN_RUNTIME_USER=opencode
ARG PUID=1000
ARG PGID=1000
# python3/make/g++ are here for the SELF-UPDATER, not for this build. An update
# runs `npm install` inside the running container, and node-pty ships no Linux
# prebuild, so a release that bumps it compiles from source right here. Without
# a toolchain that install fails and the update rolls back — every time, on the
# releases that need it most. Same reason install.sh installs one on bare hosts.
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
ca-certificates \
curl \
g++ \
git \
make \
openssh-client \
procps \
python3 \
ripgrep \
tmux \
&& rm -rf /var/lib/apt/lists/*
# The Docker CLI, taken from the official image rather than Debian's `docker.io`.
# That package is the full ENGINE: with --no-install-recommends it still pulls 15
# packages including containerd, runc, dmsetup and iptables, none of which a
# client that only talks to a mounted socket can use. Measured on top of this
# base image: `docker.io` costs 266 MB and ships Docker 20.10.24 (2023), while
# these two files cost 108 MB and ship the current CLI (493 MB vs 335 MB total).
#
# The binaries are STATIC Go builds, so they run on this glibc image even though
# the image they come from is Alpine (verified: `docker --version`, `docker ps`
# and `docker build` all work here against a mounted host socket).
#
# buildx is copied on purpose. `scripts/build-agent-image.mjs` shells out to
# `docker build` — Codeman auto-builds the agent image on the first Docker case —
# and without the plugin that silently falls back to the CLASSIC builder, which
# Docker has deprecated and will eventually drop. `docker-compose` is NOT copied:
# Codeman never shells out to it.
COPY --from=docker:29-cli /usr/local/bin/docker /usr/local/bin/docker
COPY --from=docker:29-cli \
/usr/local/libexec/docker/cli-plugins/docker-buildx \
/usr/local/libexec/docker/cli-plugins/docker-buildx
# Keep credentials out of the image. Users authenticate these CLIs at runtime
# through Codeman sessions, and the configured host bind mount retains state.
#
# ⚠️ PINNED ON PURPOSE. Unpinned, the agent CLI versions a user ends up with are
# a function of WHEN their image was built, not of any commit — so a Codeman
# release that depends on newer CLI behaviour (the trust-dialog handling is
# pinned to Claude Code 2.1.252's layout; wheel forwarding to >= 2.1.187) breaks
# on an older image with no diff anywhere to explain why. In-app updates make
# rebuilds RARER, which makes that drift worse. Pinning turns "this release needs
# a newer CLI" into a Dockerfile change, which the updater's environment gate
# already detects and refuses (docs/docker-self-update.md).
#
# Bump these deliberately, in a release. `--no-cache` is still needed to rebuild
# this layer when only the pins change upstream.
RUN npm install --global \
@anthropic-ai/claude-code@2.1.258 \
@google/gemini-cli@0.58.0 \
@openai/codex@0.152.1 \
opencode-ai@1.18.26 \
&& npm cache clean --force
# Keep the web server and every local Codeman session unprivileged. PUID and
# PGID match the host-owned application-data directory mounted by Compose. The
# requested GID may not exist in the base image, and a host UID such as 1000 may
# already belong to the baked `node` account, so handle both cases explicitly.
RUN set -eux; \
case "${PUID}" in ''|*[!0-9]*) echo "PUID must be numeric" >&2; exit 1;; esac; \
case "${PGID}" in ''|*[!0-9]*) echo "PGID must be numeric" >&2; exit 1;; esac; \
if [ "${PUID}" -eq 0 ]; then \
echo "PUID must identify an unprivileged account, not root" >&2; \
exit 1; \
fi; \
if ! getent group "${PGID}" >/dev/null; then \
groupadd --gid "${PGID}" codeman-runtime; \
fi; \
existing_user="$(getent passwd "${PUID}" | cut -d: -f1 || true)"; \
if [ -n "${existing_user}" ]; then \
usermod \
--login "${CODEMAN_RUNTIME_USER}" \
--gid "${PGID}" \
--home "/home/${CODEMAN_RUNTIME_USER}" \
--move-home \
--shell /bin/bash \
"${existing_user}"; \
else \
useradd \
--uid "${PUID}" \
--gid "${PGID}" \
--create-home \
--home-dir "/home/${CODEMAN_RUNTIME_USER}" \
--shell /bin/bash \
"${CODEMAN_RUNTIME_USER}"; \
fi
WORKDIR /opt/codeman
COPY --from=build /opt/codeman /opt/codeman
# CODEMAN_IN_CONTAINER tells the self-updater it must restart by exiting rather
# than by asking an init system that is not here (src/web/self-update.ts).
# NODE_ENV stays `production`; the updater passes `npm install --include=dev`
# explicitly, since that value would otherwise omit the build toolchain.
ENV CODEMAN_IN_CONTAINER=1 \
CODEMAN_PORT=3000 \
HOME=/home/${CODEMAN_RUNTIME_USER} \
NODE_ENV=production
EXPOSE 3000
USER ${CODEMAN_RUNTIME_USER}
CMD ["node", "dist/index.js", "web"]
+10 -5
View File
@@ -204,11 +204,16 @@ turn.
The reliable sequence is: poll `GET /api/v1/sessions/:id` until `.data.pid` is
non-null, then `wait-output` for the composer's own marker (`bypass`, the status
bar of a CLI spawned in bypass mode) with a short timeout, handling the trust
dialog only as the bounded fallback (`trust` matched → send `\r` → wait for
`bypass` again). Do not probe `trust` first and Enter blindly: the dialog text
stays in the terminal buffer for the life of the session, so a `trust` probe with
`from=buffer` keeps matching on every later run and the Enter lands in a ready
composer. A worked version is in
dialog only as the bounded fallback.
⚠️ **The fallback is not a bare `\r`.** Claude Code 2.1.252 unnumbered the dialog's
options, reversed them and highlights `No, exit`, so an Enter sent blind quits the
CLI and the pane dies seconds after the spawn. Read the `❯` marker off the current
frame (`GET /api/v1/sessions/:id/terminal?full=1`), send `ESC [ B` while it is on
`No, exit`, re-read, and confirm only once it is on `Yes, I trust this folder`.
Reading the current frame is also what keeps this correct on later runs: the dialog
text stays in the terminal buffer for the life of the session, so a `trust` probe
with `from=buffer` keeps matching long after the dialog is gone. A worked version is in
[`extending-codeman.md`](extending-codeman.md#seam-3-http-api-and-cli).
### `GET /api/v1/sessions/:id/wait`
File diff suppressed because one or more lines are too long
+137
View File
@@ -0,0 +1,137 @@
# The CLI registry
Every run mode Codeman can launch — Claude Code, Terminal/Shell, OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek Harness and OMP — is a `CliEntry`: a data record describing how to find the binary, how to build its command line, what environment it needs, and what it can do. Code that used to ask "which CLI is this?" asks the entry instead.
## Where it lives
| File | What it holds |
| ------------- | ------------------------------------------------------------------------------------------------- |
| `types.ts` | The `CliEntry` interface and everything under it. Read this first. |
| `stock.ts` | The shipped catalog. **The only file allowed to name a CLI id.** |
| `schema.ts` | Zod validation, including the cross-field checks that reject an incoherent entry at LOAD time. |
| `argv.ts` | The argv engine: the only code that turns typed tokens into a command string. |
| `patterns.ts` | The NAMED value patterns (`model`, `uuid`, `path-segment`, …) and the regex-compilation guard. |
| `profiles.ts` | The names of behaviours that genuinely need code, kept import-free so `schema.ts` can validate one. |
| `registry.ts` | Loading, merging `~/.codeman/clis.json`, and the accessors (`getCli`, `enabledClis`). |
`src/session-cli-registry-bridge.ts` maps the legacy per-mode option bag onto the engine, and `src/utils/cli-resolver.ts` / `src/utils/cli-launcher.ts` do registry-driven binary resolution and launcher-profile dispatch.
## The override file
`~/.codeman/clis.json` (instance-scoped through `dataPath()`) holds overrides and custom entries only, never a copy of the stock catalog: `{ "clis": { "<id>": { ...partial entry... } } }`. Objects merge key-wise onto the stock entry, arrays replace wholesale. **The file must be mode 0600**; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you `chmod 600` it. Every reason a file was ignored or an entry dropped is logged once, prefixed `[cli-registry]`, on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read only on restart.
## The shape of an entry
```ts
interface CliEntry {
id: CliId; // 'codex'
label: string; // 'Codex' — shown in menus
shortBadge: string; // tab badge, e.g. 'CX'
accent: string; // single hex colour
enabled: boolean;
stock: boolean; // set by the loader; a custom entry can never claim it
order: number;
kind: 'agent' | 'shell';
discovery: CliDiscovery; // how to find and prove the binary
launch: CliLaunch; // the structured argv template
env: CliEnv; // exports, tmux setenv keys, the env-override allowlist
capabilities: CliCapabilities; // what every call site reads instead of the id
// .workDetect?: { promptGlyph, workingLine } — how this CLI's pane shows work
overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store
}
```
`capabilities` is the important part. It is what `isExternalCliMode()`, `isAltScreenStripMode()`, `hooksAvailableForMode()` and every other former per-mode branch actually read.
### Regexes that come from config
Two capability fields carry a regular expression an override file can set: `discovery.version.regex` and `capabilities.workDetect.workingLine`. Both go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
`workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused.
### Three capabilities that must stay independent
`external`, `hooks` and `altScreen` describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. `shell` has no hooks but is **not** an external CLI, so a hooks predicate written as `!isExternalCliMode()` accepted `until=stop` on a shell session and then blocked the caller for their entire timeout. `deepseek` is the mirror image: it IS external and it DOES have hooks.
`test/cli-capability-predicates.test.ts` asserts that no two of the three are equivalent across the catalog, so collapsing them fails the build rather than a user's session.
## Arg-template safety
The composed command line is interpolated into `bash -c "…"` inside tmux, which makes command construction a security boundary. Four independent layers keep config out of it:
1. **Config contains no shell text.** There is no `command: "..."` field anywhere in the schema. An entry declares a sequence of typed tokens; `argv.ts` is the only place that turns them into a string, and it owns every separator itself — one space between tokens, ` || ` between fallback variants. Neither can originate from config, because config has no field that could hold either.
2. **Every literal is validated at LOAD time** against a safe-word pattern (no space, quote, backtick, `$`, `;`, `&`, `|`, redirection, parens, braces, newline or backslash). A bad literal **rejects the whole entry** rather than being dropped, because a silently dropped flag would change security-relevant behaviour — losing `--no-approve` is not a cosmetic difference.
3. **Values resolve through NAMED patterns.** A value placeholder selects a `TokenPattern` (`model`, `uuid`, `slug`, `path-segment`, `tool-list`, …) from `patterns.ts`; config can never supply its own regex for a value, so a `clis.json` structurally cannot widen its own validation. A value that fails its pattern drops the whole argument, exactly as the hand-written builders did: an invalid `--model` omits `--model`, it never substitutes something else.
4. **Escaping is independent of validation.** `renderToken()` re-checks the resolved value before emitting it unquoted, and single-quotes anything else — so even a value that somehow bypassed validation is quoted, never concatenated raw.
The only config-supplied regexes are `discovery.version.regex` and `discovery.identity.regex`. Both run against **command output** rather than a shell token, both are compiled through `compileVersionRegex()` (length cap, nested-quantifier rejection, never the `g` flag), and the output they see is truncated first.
## Named profiles: the escape hatch
Some differences genuinely need to run code rather than be described. Those are **named profiles**: a capability field holds a profile NAME, and the implementation lives in one place keyed by that name — never by CLI id.
- `discovery.launcherProfile` — for a CLI whose binary is not the agent. `dsh` boots `$DSH_HOME/profiles/<name>`, so "installed" and "runnable" have different answers; the profile answers both, plus why a specifically-named target will not work. Implemented in `utils/cli-launcher.ts`.
- `env.setenvProfile` — per-CLI environment setup that is more than a list of keys, such as DeepSeek's status bridge.
- `capabilities.transcript` — which on-disk history reader understands this CLI (`claude-jsonl`, `codex-rollout`, `deepseek-zstd`, `omp-jsonl`, `none`).
- `capabilities.echo.predictProfile` — the predictive-echo model a composer needs.
The names live in `profiles.ts`, which is kept free of imports so `schema.ts` can validate a name at load time. A profile this build does not implement is a load-time error naming the field, rather than a CLI that silently looks permanently uninstalled.
## DeepSeek: the four assumptions it breaks
DeepSeek is worth reading before assuming an entry looks like its siblings — the schema carries four extensions because of it.
| What it breaks | How the registry expresses it |
| ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------ |
| `dsh` is a profile LAUNCHER, not the agent, so "installed" is not "runnable". | `discovery.launcherProfile` + `discovery.launcherTargetParam`. |
| Its permission switch is the **`DSH_PERMISSION_MODE` env var**, not a flag — the harness has none. | `env.configSetenv` (so the ordinary `privilegedParams` clamp still reaches it) **and** `capabilities.privilegedEnvKeys`. |
| It is the only non-claude mode with real hook signals, and for it that is a per-SESSION question. | `capabilities.hooks: 'supervised'` — a third state, not a boolean. |
| Its transcript is zstd session files, one frame per write. | `capabilities.transcript: 'deepseek-zstd'`. |
## Identity probes
`discovery.identity` asks the binary whether it is the program we meant, and it runs **before** the version probe, because a version probe cannot tell an impostor from the real thing. Debian ships an unrelated `dsh` (dancer's shell) that answers `--version` perfectly happily, and npm carries squatters for both `pi` and `grok`.
`discovery.version.requireVersionMatch` is the weaker companion: a binary whose version output has the wrong shape counts as ABSENT rather than present-with-unknown-version. That is what a short, generic binary name needs, and it is what keeps `codeman doctor` and the run mode from telling the user opposite things about the same binary — both read the same regex off the same entry.
## The no-id-branching rule
`test/cli-registry-no-id-branching.test.ts` fails the build if a CLI id comparison appears outside the stock catalog. It builds its id list from the live catalog, blanks comment lines before scanning (comments legitimately quote the pattern to explain why a branch was removed, and blanking rather than dropping is what keeps reported line numbers pointing at the real file), and keeps an allowlist in which **every entry carries its reason**.
It matches four shapes, not one: `mode === '<id>'`, `mode !== '<id>'`, `case '<id>':`, and `['<id>', …].includes(mode)`. The first version matched `===` only, and that gap was not academic — the refactor it guards converted the `===` sites and left the negated ones, so 36 `!==` branches survived it, including a seven-mode chain auto-enabling Ralph under a comment asking the next person to keep it in step with a predicate by hand while the sibling code path already read the capability. A guard that sees half the shapes reports a count measured over the half it happens to catch.
The allowlist is not a formality. If a branch is about what a CLI can DO it belongs in `CliCapabilities`; the entries that remain are things that are not CLI-behaviour branches at all — chiefly the legacy per-mode `<Mode>Config` objects on `POST /api/sessions`, which are a fact about the public HTTP API rather than about any CLI, plus a few documented cases where `mode === 'claude'` is genuinely the right question (Read My Mind reads Claude's _own_ transcript, so a capability there would be actively wrong).
## Two namespaces called `param`
`launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a **launch param**. The **legacy wire field** a param arrives as is a separate namespace, and `launch.legacyConfigAliases` is the only bridge between the two.
This matters because it is invisible when it is wrong. `capabilities.privilegedParams[].param` is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a name from the wrong namespace clamps **nothing**: no load error, no failing test, the clamp simply stops running. Codex is the entry where the two names differ (`bypassApprovals` as the param, `dangerouslyBypassApprovals` on the wire), so it is the one that catches a regression. `schema.ts` rejects any entry naming a param it never declared, on both `configSetenv.fromParam` and `privilegedParams.param`.
## Fields declared for later
`shortBadge`, `accent`, `capabilities.echo`, `capabilities.wheelForward`, `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` are **declared but not yet read**. They all describe frontend behaviour, and the frontend is deliberately untouched here: `app.js`, `terminal-ui.js` and `styles.css` keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
Treat those values as **transcribed, not authoritative** — nothing enforces that `echo.policy` matches `_updateLocalEchoState`'s fallthrough, or that `accent` matches the gradient CSS paints, so re-measure before wiring one up. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; `test/cli-registry-no-id-branching.test.ts` pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
`overlays.credStore` is in the same category, for a sharper reason: the Docker credential-seeding path still reads its own `CRED_STORES` table, because this shape allows ONE store per CLI and the live table needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek declares none here even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is simultaneously the worst thing here to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest.
Everything else in the interface is live, including `overlays.remote` / `overlays.docker`, which back `defaultRemoteCommandForMode()` and `defaultDockerCommandForMode()` directly. Those two used to be hardcoded `Record<…CommandMode, string>` tables duplicating the registry with nothing keeping the two in step; `test/location-overlay-commands.test.ts` pins every resulting command as a literal string.
## Resolve at call time, never at import
Anything reading the registry must resolve it when it is asked, not when its module is first imported. `sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()` and each resolver's `searchDirs` thunk all re-read the catalog per call.
A module-level const freezes at first import, and the failure is asymmetric: a CLI enabled while the server is running moved the run menu but not the frozen surface, so validation rejected a mode the menu offered, or `codeman doctor` reported a catalog nobody had any more.
## Adding a CLI
1. Add a `CliEntry` to `stock.ts`.
2. Add a golden spawn-command pin to `test/cli-registry-spawn-golden.test.ts`, a row to `test/cli-capability-predicates.test.ts`, and its remote/docker commands to `test/location-overlay-commands.test.ts`.
3. That is usually all. If you find yourself wanting to add an `if` somewhere, the guard test will tell you — and the answer is a capability field, or a named profile if it genuinely needs to run code.
## See also
- [Agent CLIs](wiki/Agent-CLIs.md) — the user-facing per-CLI guide.
- `docs/architecture-invariants.md` — the mechanics and the history behind the rules above.
- `docs/deepseek-integration.md` — why DeepSeek is shaped the way it is.
+1 -1
View File
@@ -91,7 +91,7 @@ These map 1:1 to `CronJobSchema` (`src/web/schemas.ts`) and the `CronJob` type
| Field | Required | Values / limits | Notes |
| -------------------------- | ----------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` \| `pi` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. ⚠️ A `pi` job's readiness poll looks for `❯`/a token count, neither of which pi prints, so it burns the poll budget and then sends the prompt anyway (slower start, still works). |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` \| `pi` \| `grok` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. ⚠️ A `pi` or `grok` job's readiness poll looks for `❯`/a token count, which neither CLI prints, so it burns the poll budget and then sends the prompt anyway (slower start, still works). |
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
| `promptMode` | ✅ | `inline_text` \| `prompt_file_path` | See §5. |
+178
View File
@@ -0,0 +1,178 @@
# DeepSeek Harness (`dsh`) integration plan
> **Status**: Executed. This document records the plan, the decision behind each
> wiring point, and what was and was not verified. The user-facing guide is
> [`deepseek-integration.md`](./deepseek-integration.md); the per-decision
> invariants live in
> [`architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek`](./architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek).
> Template: the grok integration ([`grok-integration-plan.md`](./grok-integration-plan.md)),
> itself calibrated against pi. Every fact below was measured against a live
> **dsh 0.1.1-rc.2** install and **@deepseek-harness-tui/dsh-tui 0.9.0**, not read
> from documentation.
## 1. What the DeepSeek Harness is
[deepseek-ai/deepseek-harness](https://github.com/deepseek-ai/deepseek-harness)
(open-sourced 2026-08-13, MIT) is a plugin-native agent framework: tools, skills,
sessions, sandboxes and whole APPS are Cordis plugins composed into *profiles*.
`dsh` is the launcher — `dsh --profile <name>` boots
`$DSH_HOME/profiles/<name>`, an ordered stack of plugin-bundle patch layers under
the user's own overrides. State lives in `~/.dsh` (`.env` 0600, `settings.yaml`,
`cordis.patch.yml`, `profiles/`, `sessions/`, `storages/`).
## 2. Shape decisions (why DeepSeek is wired the way it is)
DeepSeek is a ninth run mode. Never a location overlay, never a web tab (the
browser UI is handled separately, §3). Three of its decisions have no precedent
in the six external CLIs before it.
| Question | Decision | Why |
| --- | --- | --- |
| What does a pane run? | `dsh --profile <name>`, profile discovered | **The decision that shapes everything else.** DeepSeek ships `web`, `headless` and `base` — no terminal agent. The interactive front door is always a third-party plugin, so Codeman resolves a binary AND a profile inventory, and "available" means both. `resolveDefaultDeepSeekProfile()` prefers a recognized TUI, then an UNRECOGNIZED profile (anyone can publish an app bundle; a classifier that has not heard of one must not hide it), and refuses `web`/`headless`, which cannot occupy a pane. |
| Which TUI? | none blessed; default for BOOTSTRAP only | `POST /api/deepseek/install-profile` defaults to `@deepseek-harness-tui/dsh-tui` (~27.5k weekly downloads, ~4x the next, MIT, and it speaks the status contract in §2.3), but accepts any npm name and the resolver never assumes that profile exists. Codeman offers a default; it does not pick a winner. |
| Permission bypass | `DSH_PERMISSION_MODE` env export, no flag | The harness has NO command-line permission option; its sandbox/approval rows read one env var with three presets (`read-only` / `workspace-write` / `danger-full-access`, read off `dsh --dump-default-config`). This is the one legitimate exception to the `CLAUDE_CODE_EFFORT_LEVEL` ban: that var hard-locks in-session switching, whereas the harness reads this with `??` as a boot-time DEFAULT, so it stays soft. Exported via `tmux setenv`, never on the command line. The Run button sends `danger-full-access`, matching every sibling Run button. |
| Multi-user clamp branch | only-if-sent, clamped to `workspace-write`, **plus an env-var half** | Omitting the export leaves the harness on `workspace-write`, which still ASKS, so an absent config is already safe (the codex/antigravity/grok shape, not pi's materialize). Clamping to `workspace-write` rather than `read-only` is deliberate: the clamp removes privilege, it must not break a session's ability to edit its own workspace. ⚠️ Unlike every sibling, clamping the CONFIG is only half the gate: the switch is an env var, `DSH_*` is an allowlisted `envOverrides` prefix, and `applyEnvOverrides()` runs AFTER `_configureDeepSeek()`, so `envOverrides: {DSH_PERMISSION_MODE: 'danger-full-access'}` on the same request would land last and win. `clampEnvOverridesForOwner()` drops `DSH_PERMISSION_MODE` and `DSH_HOME` for a non-granted owner (dropping falls through to the clamped export). `DSH_HOME` because it aims the launcher at a profile tree whose plugin code runs at BOOT, before any approval row. |
| `hooksAvailableForMode()` granularity | per SESSION for deepseek, per mode for everything else | `deepSeekConfig.statusReporting: false` disarms the `HERDR_*` export, and the triple is the only reason a dsh session posts anything, so a mode-only answer would accept `until=stop` where nothing can send one — the infinite-wait the predicate exists to prevent. Call sites pass `sessionHookOptions(session)`; the default stays permissive so a forgotten one degrades to the old behaviour. ⚠️ Profile conformance stays unknowable at request time (an unrecognized profile is deliberately launchable), so a non-conforming TUI still times out on an explicit `stop`; the default set keeps `idle`/`exit` for that. ⚠️ The predicate is NOT "is this claude": Read My Mind and intent capture read Claude's transcript and were silently widened by this change, so they compare `mode === 'claude'` directly now. |
| Profile install spawn | own process group, hand-rolled timeout | `dsh plugin add` fans out into package-manager children, and spawn's built-in `timeout` signals only the direct child: survivors keep the inherited stdio pipes open, `close` never fires, and the held-open request leaks with no route-level deadline. `detached: true` + negative-pid SIGTERM→SIGKILL, the same escalation `runGit()` uses for the same reason, plus a last-resort reap for a grandchild that escaped the group. |
| Idle detection | **real hook events via a status shim** | The standout decision. The TUI already reports its lifecycle to a supervising process through a generic env-gated contract inherited from Herdr: `HERDR_ENV=1` + `HERDR_BIN_PATH` + `HERDR_PANE_ID` make it run `<bin> pane report-agent <id> --state idle\|working\|blocked …` on every state change, exit 0 = delivered. `deepseek-status-shim.ts` generates a script into the data dir and points `HERDR_BIN_PATH` at it. So deepseek is the only non-claude mode that passes `hooksAvailableForMode()` — earned by emitting definitive signals, not granted. An interface implementation, not an impersonation: no real `herdr` binary is ever executed, and a TUI that ignores the contract simply falls back to output stabilization. |
| `agent_working` event | new, 157th SSE constant | The one hook event with no Claude Code hook behind it. A harness turn cannot run while its own modal approval is on screen, so "started working" proves a dialog was answered in the terminal. Without it a dsh red alert would survive until the next `stop` — the exact stuck-alert bug the claude path already fixed once, and its pane-capture staleness sweep is Claude-dialog-shaped and cannot help here. |
| Resolver | identity probe THEN version probe | Strictest of the family, and not by preference. `dsh` is not merely a squattable npm name: Debian ships an unrelated `dsh` (dancer's shell, `apt install dsh`) which would answer a version probe convincingly and then be handed a spawn line. `dsh --help` must match `DeepSeek Harness` first. `DEEPSEEK_VERSION_REGEX` keeps the prerelease tail (`0.1.1-rc.2`), since truncating it would report an rc as a release. |
| Env allowlist | `DSH_*` + `DEEPSEEK_*` | `DSH_*` covers the launcher's documented inputs (`DSH_HOME`, `DSH_PERMISSION_MODE`, `DSH_TELEMETRY_MODE`, the `DSH_TUI_*` knobs); `DEEPSEEK_*` is the vendor namespace holding `DEEPSEEK_API_KEY`/`DEEPSEEK_BASE_URL`, same reasoning that admitted `XAI_*` for grok. ⚠️ Pi's lesson repeats exactly: a dsh `settings.yaml` can nominate ANY env var as a provider credential (`apiKeyEnv`), and the allowlist is one GLOBAL list, so admitting those would widen every mode at once. They stay out. |
| Model | NOT a session field | The model is a composition entry (`agent-default-model`) in the profile's config tree, set in `~/.dsh/settings.yaml` + `cordis.patch.yml`. Both create paths deliberately resolve no model for this mode rather than inventing a flag. |
| Alt-screen strip | OUT of `isAltScreenStripMode()` | Third-party fullscreen TUIs with their own scrollback and mouse handling — the opencode case, not the Ink case. |
| Local echo | `'buffer'` via the `_updateLocalEchoState` fallthrough | UNMEASURED against a live authenticated session (see §5), same honest gap grok shipped with. The leading TUI's composer supports `@` completion and history search, which *may* make it per-keystroke reactive like codex; if so the fallback is the `'off'` branch. |
| Docker | image installs dsh AND a profile | Profiles are deliberately NOT seeded from the host: each is a per-profile `node_modules` tree, host-arch-specific and far too large to copy per container start. Only `~/.dsh/.env`, `settings.yaml`, `cordis.patch.yml` are seeded (auth + model composition). The profile install rides the `useradd` layer so the closing `chgrp`/`chmod g=u` covers it, which is what keeps it usable under the arbitrary uid the container runs as. |
| Remote SSH | `exec "$SHELL" -i -l -c 'dsh'` | Boots the remote box's default profile; a remote with several needs the per-host `commands.deepseek` override, since `deepSeekConfig` does not cross ssh. |
## 3. The web profile
The browser UI is the only interactive surface DeepSeek ships itself, so it gets
a **shortcut, not a run mode**: `Run ▸ DeepSeek web UI…` starts
`dsh web --no-open --host 127.0.0.1 --port <free> --trusted-host <codeman-authority>`
as a background process and opens the URL as an ordinary web tab.
The server was a **shell session** first, on the reasoning that Codeman already
supervises those (visible, scrollable, killable, dies with its tab) so nothing
new had to own a long-lived HTTP server. That version worked and was still
wrong in use: clicking "open the DeepSeek web UI" put a terminal tab on screen
next to the web tab actually asked for, every single time, and after the first
launch the terminal was pure noise. Opening a dashboard should open one tab.
So `POST /api/deepseek/web` owns it instead (`src/deepseek-web-server.ts`), and
what the session gave away for free is now explicit: exactly one server, reused
rather than raced on a second click; restarted when the requested authority
changes; killed on Codeman shutdown (a detached child would otherwise hold its
port against the next start — the very EADDRINUSE this feature already got
wrong once); and boot output captured, since with no shell tab there is nowhere
else for a stack trace to land. It is fenced at the same bar as the profile
installer: booting a dsh profile executes the plugin code in it, so it requires
the privileged grant in multi-user mode.
`--trusted-host` is load-bearing — dsh fences its `/api` behind a browser-trust
check on the request authority, and a Codeman web tab reaches it through
Codeman's own origin via the webview proxy, not directly. The authority comes
from the CLIENT (`location.host`) because only the browser knows which of a
multi-homed Codeman's origins is actually in play.
Three things about this shortcut are load-bearing and each came from it failing
in exactly that way against a real install:
- **The port is chosen, never hardcoded.** `GET /api/deepseek/web-port` walks
3080..3119 for a free loopback port. 3080 is dsh's own default, which makes it
precisely the port a DeepSeek user is most likely to already be serving on:
binding it unconditionally killed the launch with `EADDRINUSE` against the
user's own `dsh web`.
- **The tab is opened only after the server answers.** The launch polls
`POST /api/webviews/probe` until the URL responds, so a server that dies on
startup reports the failure and points at its shell tab, instead of silently
persisting a dashboard aimed at nothing.
- **The saved tab is `trusted: true`, and must be.** An untrusted webview is
sandboxed without `allow-same-origin`, which breaks this dashboard twice: the
dsh client-runtime reads `localStorage` while loading plugins and dies there,
and an opaque-origin frame sends `Origin: null`, so dsh's trust check 403s
every `/api` call regardless of what `--trusted-host` names. Passing
`location.host` only means anything once the frame actually carries that
origin. The trade is real — a trusted proxied frame is same-origin with
Codeman and can reach Codeman's API — and is defensible only because this
particular dashboard is an agent harness Codeman just started itself on
loopback, which can already run code as the user. It is not a precedent for
trusting third-party dashboards generally.
The record is marked `managed: 'deepseek-web'`, which keeps it out of the
saved-dashboard list: the shortcut that maintains it is already a menu entry, so
listing both showed the same dashboard twice. Being managed is also what lets a
relaunch repoint the existing row instead of stacking one dead dashboard per
restart, since the port is now chosen per launch.
The authority baked into `--trusted-host` is the one the launch was clicked
from, and reuse is conditional on it: a running server fenced for a *different*
origin is stopped and restarted rather than reused, because reusing it renders a
page whose every API call 403s — which reads as a broken dashboard rather than a
misconfigured one.
## 4. Touch points (the checklist)
Backend: `types/session.ts` (SessionMode + `DeepSeekConfig` + SessionState),
`utils/deepseek-cli-resolver.ts` (new) + barrel, `deepseek-status-shim.ts` (new),
`tmux-manager.ts` (`buildDeepSeekCommand`, dispatch, resume flag, PATH export,
truecolor, `_configureDeepSeek`, availability error, plumbing), `session.ts`
(external-mode gate, label, config plumbing, tmux-required error, attach env),
`mux-interface.ts`, `schemas.ts` (prefixes, `DeepSeekConfigSchema`,
`DeepSeekInstallProfileSchema`, both mode enums, remote overrides, cron agentType,
`agent_working`), `session-wait-registry.ts` (`hooksAvailableForMode`),
`hook-event-routes.ts` (`APPROVAL_RESOLVING_EVENTS`), `session-routes.ts` (clamp +
both create paths + `resolveDeepSeekLaunchError`), `system-routes.ts`
(`GET /api/deepseek/status`, `POST /api/deepseek/install-profile`), `server.ts`
(availability inject + mux restore), `sse-events.ts`, `docker-hosts.ts`,
`remote-hosts.ts`, `config/dependency-registry.ts`,
`response-viewer-transcript.ts`, `cron/cron-service.ts` (comment),
`tui/tui-client.ts` + `tui-app.ts`.
Frontend: `index.html` (welcome button, run-mode entry, install affordance, web-UI
shortcut, cron option, clone Brain option), `session-ui.js` (`runDeepSeek()`,
`runDeepSeekWeb()`, `installDeepSeekProfile()`, dispatch, availability, "Run DS"
label, external-CLI gates), `app.js` (label, `ds` tab badge, kill-menu, SSE map),
`settings-ui.js` (welcome gate + `_onHookAgentWorking`), `constants.js`,
`mobile-overview.js`, `home-sessions.js`, `panels-ui.js`, `i18n.js`,
`terminal-ui.js`, `styles.css` + `mobile.css` (brand-indigo identity; the non-og
skin block and the mobile `!important` pair are both load-bearing).
Meta: `docker/agent.Dockerfile`, `install.sh`, `package.json` keyword,
`skills/codeman/reference/*`, CLAUDE.md, `architecture-invariants.md`.
Tests: `test/deepseek-mode.test.ts` + `test/deepseek-cli-resolver.test.ts` (new);
`run-mode-ui`, `render-index-html`, `mobile-overview`, `agent-skill-mode-lists`
(extended).
## 5. Verification performed
See the summary at the end of the implementing session for the live run. In
short: the CI gate green; the resolver, profile inventory, spawn-line and clamp
behaviour covered by 31 new unit tests; and an isolated instance used to exercise
`GET /api/deepseek/status` and a real session against the live dsh install.
**Not verified (honest gaps):**
- The local-echo `'buffer'` policy against the TUI's real composer (§2). If it
turns out per-keystroke reactive like codex's, flip it to the `'off'` branch;
teaching `PredictiveEchoAddon` its composer row is the larger follow-up.
- Scrollback/repaint behaviour of a third-party fullscreen TUI under the narrow
strip during a long session.
- A Docker case with `mode: 'deepseek'` (needs a `--no-cache` agent-image
rebuild — see the `--no-cache` rule in CLAUDE.md).
- A remote-SSH deepseek case.
- The web-UI shortcut against a tunnel authority. Loopback and a tailnet name are
both verified end to end through the webview proxy (dashboard renders, its
`/api` calls succeed, no shell session created).
## 6. Follow-ups
- **Response viewer**: read `~/.dsh/sessions/**` (JSONL) the way codex rollouts
are read back. Highest-value follow-up, and very achievable.
- **`headless` as an execution backend** for Codeman's own AI checks
(`ai-idle-checker`, `ai-plan-checker`), today Claude-only.
- **Profile/model picker in Session Options**, reading `GET /api/deepseek/status`
`.profiles`.
- **`--patch` overlays per session**, which is the harness-native way to change
agent composition without touching the user's profile.
- Measure the local-echo policy and pin the result the way pi did.
+305
View File
@@ -0,0 +1,305 @@
# DeepSeek Harness (`dsh`) in Codeman
Codeman can run [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)
as a session backend, alongside Claude Code, OpenCode, Codex, Gemini,
Antigravity, Pi and Grok. It is the ninth run mode, and the one that is wired
least like the others, for two reasons worth understanding before you use it.
## 1. The agent is a profile, not the binary
`dsh` is a **launcher**, not an agent. It boots a *profile*: an ordered stack of
plugin-bundle patch layers under `$DSH_HOME/profiles/<name>` (`$DSH_HOME`
defaults to `~/.dsh`). DeepSeek ships three bundles and none of them is a
terminal agent:
| Profile | What it is | Can Codeman run it in a tab? |
| ------------ | --------------------------------- | ---------------------------- |
| `web` | the browser UI, served on :3080 | no — but see §6 |
| `headless` | answers one task and exits | no |
| (`base`) | the shared core, no app at all | no |
The interactive terminal front door is **always a third-party plugin**. So
"DeepSeek is installed" and "Codeman can start a DeepSeek session" are different
questions, and Codeman answers both separately:
```bash
curl -s localhost:3000/api/deepseek/status | jq
{
"available": true, # the `dsh` binary resolved and proved its identity
"runnable": false, # ...but nothing installed can drive a pane
"path": "/home/you/.local/bin",
"version": "0.1.1-rc.2",
"dshHome": "/home/you/.dsh",
"defaultProfile": null,
"profiles": [ { "name": "web", "kind": "web", "bundles": [...] } ]
}
```
### Installing a terminal profile
From the UI: open the **Run** dropdown. When `dsh` is installed but no
pane-capable profile is, the menu shows **DeepSeek — add a terminal profile…**.
One click installs one and the normal DeepSeek entry appears.
By hand, or to pick a different front door:
```bash
dsh plugin --profile dsh-tui add @deepseek-harness-tui/dsh-tui
```
⚠️ **`pnpm` has to be on PATH for either route.** `dsh plugin` is a thin forwarder
that spawns a literal `pnpm` with no npm fallback, so without one it exits 127 with
`dsh: pnpm not found on PATH` — both by hand and behind the UI button, which
surfaces that same line as the install error. `npm install -g pnpm` (or
`corepack enable pnpm`) is the fix. This is what broke the Docker agent image in
[#352](https://github.com/Ark0N/Codeman/issues/352); the image now installs pnpm
alongside `dsh`.
Codeman's default is `@deepseek-harness-tui/dsh-tui` because it is by a wide
margin the most used community TUI, it is MIT, and it implements the status
contract described in §3. It is a **default, not a requirement**: any profile
under `$DSH_HOME/profiles` that is not `web` or `headless` shows up in the
inventory and can be launched, including one you compose yourself. The endpoint
accepts any npm package name:
```bash
curl -sX POST localhost:3000/api/deepseek/install-profile \
-H 'Content-Type: application/json' \
-d '{"profile":"my-tui","package":"@someone/dsh-tui"}'
```
Installing a plugin is arbitrary code execution on the host, so in multi-user
mode this endpoint requires the can-bypass-permissions grant (the same bar as a
`shell` session). The request is held open while the package manager runs and is
bounded at five minutes; the install runs in its own process group, so hitting
that bound kills the whole tree rather than just the launcher.
> **`dsh` is also a Debian program.** `apt install dsh` gives you "dancer's
> shell", a distributed shell, which would answer `--version` convincingly.
> Codeman's resolver therefore demands the harness's own help banner before it
> will point a spawn line at a candidate, and `GET /api/deepseek/status` reports
> `path` and `version` so a misresolution is diagnosable rather than presenting
> as "the mode just doesn't work".
## 2. Permissions are an env var, not a flag
The harness has **no `--dangerously-skip-permissions` equivalent**. Its sandbox
and approval rows are configuration, driven by one documented input,
`DSH_PERMISSION_MODE`, with three presets (read off `dsh --dump-default-config`):
| `DSH_PERMISSION_MODE` | sandbox | approvals | notes |
| --------------------- | -------------------- | --------- | ------------------------- |
| `read-only` | `read-only` | ask | |
| `workspace-write` | `workspace-write` | ask | the harness's own default |
| `danger-full-access` | `danger-full-access` | **never** | what the Run button sends |
Codeman exports it via `tmux setenv`, never on the command line. Because the
harness reads it with `??`, it is a **soft default**: it sets the boot-time
preset and you can still change permission mode inside the session.
Omitting it entirely leaves the harness on `workspace-write`, which still asks —
which is why the multi-user clamp only needs to force a *sent* value down. A
non-granted owner's `danger-full-access` becomes `workspace-write`, not
`read-only`: the clamp removes privilege without breaking the session's ability
to edit its own workspace.
Because the switch is an env var rather than a flag, that clamp has a second half
no other CLI needs. `DSH_*` is an allowlisted `envOverrides` prefix (it has to be:
that is also how you set the harness's ordinary knobs), and env overrides are
applied *after* the permission export, so in multi-user mode a non-granted owner
sending
```json
{ "mode": "deepseek", "envOverrides": { "DSH_PERMISSION_MODE": "danger-full-access" } }
```
would otherwise hand back the privilege the config clamp just removed. For a
non-granted owner Codeman therefore **drops `DSH_PERMISSION_MODE` and `DSH_HOME`
from `envOverrides`**; dropping them falls through to the clamped config and the
server's own `DSH_HOME`. `DSH_HOME` is in that list because it points the
launcher at a profile tree, and a profile's plugin code runs at boot, before any
approval row can apply. Single-user installs and granted owners are unaffected.
## 3. Real idle detection (the interesting part)
Every other external CLI mode in Codeman is **readiness-guessed**: Codeman
watches the PTY go quiet and infers that a turn ended. Claude is the exception,
because Claude Code fires hooks.
DeepSeek is the second exception. The community terminal front door already
reports its own lifecycle to a supervising process through a generic,
env-var-gated contract (inherited from [Herdr](https://herdr.dev)): when
`HERDR_ENV=1`, `HERDR_BIN_PATH` and `HERDR_PANE_ID` are set, it shells out on
every state change with
```
"$HERDR_BIN_PATH" pane report-agent "$HERDR_PANE_ID" \
--source custom:dsh-tui --agent dsh-tui \
--state idle|working|blocked [--message ...] --seq N
```
Codeman points `HERDR_BIN_PATH` at a small generated shim
(`~/.codeman/dsh-status-shim.mjs`, written at session create) which forwards each
report to `POST /api/hook-event`. The mapping:
| Harness state | Codeman hook event | What you get |
| ------------- | ------------------ | -------------------------------------------------------- |
| `blocked` | `permission_prompt`| red "needs you" tab alert + an Approvals Inbox item |
| `idle` | `stop` | definitive end-of-turn: respawn triggers, `wait` returns |
| `working` | `agent_working` | clears an alert answered in the terminal, at once |
So a DeepSeek session gets Claude-grade signals: `GET /api/sessions/:id/wait`
really can block on `stop` and `blocked` for it, and it is the only non-Claude
mode for which that is true (`hooksAvailableForMode`).
That is a per-*session* answer, not a per-mode one. Turning the bridge off with
`deepSeekConfig.statusReporting: false` means nothing will ever post a hook event
for that session, so an explicit `until=stop` is refused up front (with a message
naming the setting) rather than blocking for your whole timeout. Omitting `until`
never fails: the hook-only signals are dropped from the default set and you still
get `idle` and `exit`.
One limit worth knowing: whether the *profile* implements the contract cannot be
known at request time (Codeman deliberately treats an unrecognized profile as
launchable). A dsh session running a non-conforming TUI therefore still accepts
`until=stop` and will time out on it. `idle`/`exit` are the reliable pair there.
This is an interface implementation, not an impersonation — nothing on your
machine executes a real `herdr` binary. If you use a terminal profile that does
*not* implement the contract, the shim is simply never called and the mode falls
back to output-stabilization readiness like its siblings. Turn it off per session
with `deepSeekConfig.statusReporting: false`.
## 4. Starting a session
From the UI, pick **DeepSeek** in the Run dropdown (or the **Run DeepSeek**
welcome button) and press Run. Over the API:
```bash
curl -sX POST localhost:3000/api/quick-start \
-H 'Content-Type: application/json' \
-d '{
"caseName": "myproject",
"mode": "deepseek",
"deepSeekConfig": {
"profile": "dsh-tui",
"permissionMode": "danger-full-access"
}
}'
```
`deepSeekConfig` fields: `profile`, `permissionMode`, `resumeSession`,
`resumeSessionId`, `statusReporting`. Resume prefers an explicit id over the
most-recent form, and both are passed through to the profile's app, which is
where `--resume` is understood.
**Models are not a session field.** The model is a composition entry in the
profile's config tree (`agent-default-model`), not a CLI flag, so Codeman does
not try to set one. Configure it where the harness does: `~/.dsh/settings.yaml`
plus a home-level `~/.dsh/cordis.patch.yml`, or a `--patch` overlay on the
profile. That is also how you point dsh at a local or third-party provider.
**Environment.** `DSH_*` and `DEEPSEEK_*` are allowlisted for `envOverrides`
(so `DSH_HOME`, `DSH_PERMISSION_MODE`, `DEEPSEEK_API_KEY`, `DEEPSEEK_BASE_URL`
all flow through). Provider keys with *other* names are deliberately not: a dsh
`settings.yaml` can nominate any env var as a credential via `apiKeyEnv`, and
Codeman's allowlist is global, so admitting them would widen it for every mode at
once. Authenticate those the way dsh does, from the file or the server's own
environment.
## 5. Reading a session back, and driving one as a worker
dsh writes a real transcript — `$DSH_HOME/sessions/<mangled-cwd>/<id>/session.jsonl.zstd`
— so `GET /api/sessions/:id/last-response` reads that rather than segmenting the
pane, and the Response Viewer shows a dsh conversation the way it shows a claude
or codex one (`?context=full` returns prompt / response / tool blocks).
Reading the pane instead is not merely coarse for this mode, it is wrong: dsh-TUI
paints a full-screen splash, so the segmenter answered a `last-response` call for
a fresh dsh session with its ASCII-art logo — which anything polling for a
worker's first answer reads as an answer. Three things about the file shaped the
reader (`src/deepseek-transcript.ts`):
- **It is one zstd FRAME per append, not one zstd stream.** `zstd -dc` decodes all
of them, Node's `zlib` zstd decoder stops at the first: a real 56-line
transcript came back as 1 line. The reader walks frame headers itself. On a Node
older than 22.15 (no zstd at all) the mode falls back to the pane, as before.
- **Not every `user/message` is the user.** Each turn also records a
plugin-sourced runtime-context snapshot; only `source.kind === 'user'` is a
prompt.
- **A failed turn is not an empty one.** `turn/end` carries the provider's error,
which is returned as `Turn error: …` (and an early stop such as `max-tokens` as
`Turn ended: …`) instead of an empty string that reads as "still thinking".
The transcript reader applies to **local** dsh sessions only. A Docker case's
harness writes its transcript inside the container's own `~/.dsh` (the workspace
bind mount does not cover it), and a remote-SSH case's lives on the remote host,
so the local reader could never find those files — such sessions keep the pane
segmenter, coarse but real. The splash caveat above applies to them accordingly.
### As an agent worker
Because dsh has both halves — a real end-of-turn signal and a real transcript — an
agent can drive a dsh session the same way it drives a claude one, and the bundled
`codeman` agent skill does. Spawning `beta:deepseek` in its worker list gives a
worker that is tasked, waited on and read with the same calls as its claude
siblings; no other external CLI mode qualifies. Two edges are worth repeating here:
- **Readiness is not the stop signal.** The harness reports `idle` at boot roughly
300 ms *before* the composer paints (measured 2.26 s vs 2.56 s after spawn), so a
send-and-wait fired immediately after create resolves on that boot report,
reports a turn that never ran, and leaves the prompt in a pane that was not yet
accepting input. Wait for the composer (`❯`) instead.
- **Wait on `stop`, not on the default signal set.** That set also carries `idle`,
which for every external CLI is inferred from output stabilization; a dsh TUI
that repaints rarely reads as idle mid-turn.
## 6. The web UI as a tab
The browser UI is the one interactive surface DeepSeek ships itself, so it gets a
shortcut rather than a run mode: **Run ▸ DeepSeek web UI…** starts
`dsh web --no-open --host 127.0.0.1 --port <free> --trusted-host <codeman-host>`
as a background child process (`src/deepseek-web-server.ts`, behind
`POST/GET/DELETE /api/deepseek/web`) and opens it as a Codeman web tab once the
server actually answers.
It is a child process rather than a shell session because the session version
opened a terminal tab nobody asked for on every click. What the session gave for
free is therefore explicit here: one instance with reuse, a restart when the
requested `--trusted-host` authority differs from the running one, a kill on
server stop, and captured boot output. The `--trusted-host` flag is load-bearing —
dsh fences its `/api` behind a browser-trust check on the request authority, and a
Codeman web tab reaches it through Codeman's own origin via the webview proxy, not
directly. Without it the page renders and every API call fails.
## 7. Docker and remote cases
Docker cases work: the agent image installs `dsh` and bootstraps a `dsh-tui`
profile into the container. Profiles are deliberately **not** seeded from the
host (each is a per-profile `node_modules` tree, host-arch-specific and far too
large to copy on every container start); only `~/.dsh/.env`, `settings.yaml` and
`cordis.patch.yml` are seeded, which is what carries auth and model composition
in. As with pi and grok, in-container sessions are invisible host-side:
`~/.dsh/sessions` inside a container is that container's own.
Remote SSH cases default to `dsh` through a login shell, which boots the remote
box's default profile. If the remote has several, name one with the per-host
`commands.deepseek` override — the local `deepSeekConfig` does not cross ssh.
## 8. What is not wired
Deliberately minimal, on the same reasoning as the grok integration: the harness
is a fast-moving developer preview and every flag added is a flag validated
forever.
- `--patch` overlays per session (the profile's own layers apply as normal).
- `dsh plugin` management beyond first-time profile install.
- The `headless` profile as a one-shot execution backend for Codeman's own
internal AI checks (today those are Claude-only).
- Model/provider selection from Session Options.
## Verified against
`dsh 0.1.1-rc.2` and `@deepseek-harness-tui/dsh-tui 0.9.0`. The permission
presets, the profile layout, and the supervisor contract above were all read off
the live install rather than from documentation.
+59 -4
View File
@@ -2,7 +2,7 @@
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` / `pi` all work inside the container.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` / `pi` / `grok` / `deepseek` / `omp` all work inside the container.
## One-time setup: build the base image
@@ -25,12 +25,24 @@ A zero exit code only proves the layers ran, not that the toolchain works. Verif
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy pi; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
'for c in claude codex gemini opencode agy pi grok dsh omp; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
Antigravity (`agy`) is the one CLI not installed from npm (Google ships a standalone binary), so it has its own Dockerfile step and adds roughly 190MB; a full image lands near 1.6GB. Pi also gets its own step, because upstream documents installing it with `--ignore-scripts` and that flag must not silently change how the other four npm CLIs install.
⚠️ `dsh --version` is the one line above that answers a different question than the
others: `dsh` is a profile launcher, so a working binary says nothing about whether
the image can actually run a DeepSeek session. Check the profile the Dockerfile
installs into the agent's HOME as well, or a `mode: 'deepseek'` case starts a pane
that dies on arrival:
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md).
```bash
docker run --rm codeman/agent:base ls ~/.dsh/profiles/dsh-tui/package.json
```
Building that profile is also why `pnpm` is in the image: `dsh plugin` forwards straight to a literal `pnpm` and exits 127 without it (issue #352), and pnpm — unlike npm — blocks dependency lifecycle scripts by default and fails the install over it, so the profile step passes `--config.dangerouslyAllowAllBuilds=true`.
Antigravity (`agy`) and Grok (`grok`) are the two CLIs not installed from npm (Google and xAI ship standalone binaries), so each has its own Dockerfile step, adding roughly 190MB and 160MB respectively. Pi also gets its own step, because upstream documents installing it with `--ignore-scripts` and that flag must not silently change how the other npm CLIs install.
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md). Grok is seeded per-file for the same reason (`auth.json`, `config.toml`, `pager.toml` out of `~/.grok`, which also holds `sessions/`, `memory/` and the ~160MB binary under `downloads/`), with the same consequence for `grok -c`. See [`grok-integration.md`](./grok-integration.md). OMP is the one CLI in this family where `sessions/` is the EXCEPTION rather than the rule: `~/.omp/agent/{config.yml,mcp.json,models.yml,settings.yml}` are seeded per-file (the dir also holds SQLite caches and `terminal-sessions/`), but `~/.omp/agent/sessions/` is shared RW like codex's, not seeded, because Codeman reads it host-side for history recovery and `--resume` pinning. See [`omp-integration.md`](./omp-integration.md).
## Quickest path: one-click "Run in Docker"
@@ -66,6 +78,49 @@ curl -X POST localhost:3000/api/cases/docker-link -d '{"name":"sandbox","hostId"
curl -X POST localhost:3000/api/quick-start -d '{"caseName":"sandbox","mode":"claude"}'
```
## Attach to a container you already run
The tab's **Attach to an existing container** toggle points a case at a container **you**
built and run. Codeman only ever `docker exec`s into it: it never creates, starts, stops,
restarts or removes it, and it seeds no credentials into it, so the CLIs inside must already
be installed and logged in. A missing or stopped container is an error to report, not a state
to fix — start it yourself and reopen the session.
- **Container Name** is a picker over the engine's containers that you can also type into
(the engine may be remote, or the container may not exist yet when you fill the form).
Stopped containers are listed too, sorted last and labelled, so "mine isn't here" is never
a dead end.
- **Container Workdir** is a path that must already exist **inside** the container. Adoption
mounts nothing, so it need not match the host workspace path; **Browse** lists directories
inside the container itself. Without this check, a wrong path fails at launch as a bare
`execvp failed` inside the pane.
- **Workspace Path** is still a real host directory. It backs file previews, attachments and
watchers exactly as it does for an owned case, but here it is only a mirror: nothing is
bind-mounted, so point it at whatever host directory your container already exposes.
- **Check container** runs a read-only preflight and reports what is inside before you commit
to a case name (running or not, tmux present, which CLIs resolved).
- **Run modes come from the container**, not the host: a host with no `claude` still offers
Claude if the container ships it, and a mode the container lacks is hidden.
- Claude is launched **without** `--dangerously-skip-permissions` when the container's exec
user is root, because Claude Code refuses that flag as root and the refusal is only visible
inside the container.
- Image, network and resource settings disappear from the form: they describe a
`docker create` that adoption never runs.
Recreate is refused for an adopted case, full-image export is refused (it would commit a
container that is not ours), unlinking the case leaves the container running, and the boot
reaper skips it. Workspace-only export still works and never pauses the container.
Equivalent API:
```bash
curl -X POST localhost:3000/api/docker-cases/adopt-preflight -d '{"hostId":"local","container":"my-dev-box","containerWorkdir":"/workspace"}'
curl -X POST localhost:3000/api/cases/docker-adopt -d '{"name":"devbox","hostId":"local","container":"my-dev-box","hostWorkspacePath":"/home/you/projects/devbox","containerWorkdir":"/workspace"}'
```
In multi-user mode adoption is **admin-only**, unlike `docker-link`: an adopted container's
mounts belong to whoever built it, so one mounting `/` would hand the adopter the whole host.
## Lifecycle
- **Reconnect after a Codeman restart** lands back in the same live agent (the in-container tmux survives).
+78
View File
@@ -0,0 +1,78 @@
# Docker Compose deployment
This configuration builds the Codeman application image locally from this checkout. It does not download or depend on a pre-built Codeman image.
For the Compose configuration, environment settings, storage migration, and macvlan networking examples, see the [Docker deployment guide](../docker/README.md).
The image includes Claude Code, Codex, Gemini CLI, and OpenCode. Authenticate a CLI from its Codeman session; credentials are never baked into the image.
## Prerequisites
- Docker Engine or Docker Desktop with Docker Compose v2
- A reachable Docker daemon
The application container mounts the Docker daemon socket so Codeman can create and manage its isolated Docker cases. Treat anyone who can administer this Compose project as having Docker-host-equivalent access.
## Start
Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/Coding/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes.
```sh
cp docker/.env.example docker/.env
```
On PowerShell, use the following command instead.
```powershell
Copy-Item docker/.env.example docker/.env
```
On Linux, run the stack with the start script. It determines `PUID` and `PGID` from the owner of `CODEMAN_APPDATA_PATH`, and `DOCKER_SOCKET_GID` from the configured Docker socket, before invoking Compose. A root-owned application-data directory is rejected so the runtime account cannot become UID 0.
```sh
bash docker/Start-Codeman.sh
```
On other platforms, run Compose directly. `PUID` and `PGID` default to `1000:1000`; set them in `docker/.env` when the application-data directory has a different owner.
```sh
docker compose --env-file docker/.env -f docker/docker-compose.yaml up --build -d
```
Open `http://localhost:3000` and sign in with the username and password from `docker/.env`.
## Operations
The local image is tagged `codeman:local` by default. Change `CODEMAN_IMAGE` in `docker/.env` if a different local tag suits your environment.
```sh
docker compose --env-file docker/.env -f docker/docker-compose.yaml logs -f codeman
bash docker/Start-Codeman.sh
docker compose --env-file docker/.env -f docker/docker-compose.yaml down
```
`CODEMAN_APPDATA_PATH` holds Codeman state and survives container recreation. Remove that host directory only when deliberately resetting the installation.
`CODEMAN_CASES_PATH` must be an absolute path on the Docker host. Compose mounts it at the same path inside Codeman, so the host daemon can bind the managed workspace into isolated Docker cases. Do not set it to `/home/${CODEMAN_RUNTIME_USER}/codeman-cases`.
Compose passes `CODEMAN_APPDATA_PATH` into Codeman as `CODEMAN_DOCKER_HOST_HOME`. Codeman uses that value to translate generated Docker seed, credential and hook-secret bind sources from the container's home path into paths visible to the host Docker daemon.
If `docker info` reports `SwapLimit=false`, set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1`. Isolated cases retain their configured memory limit. Codeman omits the unsupported swap-limit option and filters only the daemon's exact swap-capability warning while retaining every other Docker create error.
If that directory was created by an earlier root-running image, change its ownership to the configured `PUID:PGID` before starting this version. This preserves existing CLI credentials and session state while allowing the unprivileged runtime account to use them.
## Updating
Codeman updates itself from **App Settings → Updates**, as it does on a bare host. The checkout mounted at `/opt/codeman` is the same directory Compose builds from, so the update's `git checkout` and rebuild land on the host and survive container recreation; the restart is the server exiting, which `restart: unless-stopped` turns into a relaunch on the new build.
That applies application code only. A release that changes `docker/server.Dockerfile`, `docker/docker-compose.yaml`, or adds a key to `docker/.env.example` needs the image rebuilt or the container recreated, which a container cannot do to itself. The updater detects each case and refuses with a message naming what changed; run `docker/Start-Codeman.sh` on the host to apply those.
`CODEMAN_REPO_PATH` overrides which checkout is mounted. It defaults to the compose project's parent directory, so it normally needs no setting. Point it at a directory that is not a git checkout and in-app updates are reported as unavailable.
Full detail, including the fingerprint baseline and the troubleshooting table: [`docker-self-update.md`](docker-self-update.md).
## Docker cases
The default socket path is `/var/run/docker.sock`, which works with a standard Linux Docker Engine. The Bash start script detects its numeric group ID. When running Compose directly, set `DOCKER_SOCKET_GID`, for example using `stat -c '%g' /var/run/docker.sock`, so the unprivileged `CODEMAN_RUNTIME_USER` account can create Docker cases. Docker Desktop users should set `DOCKER_SOCKET` in `docker/.env` only when their Docker installation exposes a different compatible socket path.
Codeman Docker cases are sibling containers on the host daemon, not children of the application container. The Compose configuration handles their workspace bind mount through `CODEMAN_CASES_PATH`; the `/home/${CODEMAN_RUNTIME_USER}` application-data mapping is for Codeman state and ordinary in-container sessions, not sibling-case workspaces.
+218
View File
@@ -0,0 +1,218 @@
# Self-update in the Docker Compose deployment
Codeman running as a container updates itself from **App Settings → Updates**, the
same place and the same button as a bare-host install. This document explains how
that works, what it deliberately refuses to do, and how to recover when it stops.
The bare-host updater is documented in
[`architecture-invariants.md#self-update`](architecture-invariants.md#self-update);
this file covers only what the container changes.
## The short version
| Change in the release | Applied by |
| -------------------------------- | ------------------------------------------------ |
| Application code | The in-app updater |
| `docker/server.Dockerfile` | `docker/Start-Codeman.sh` on the host |
| `docker/docker-compose.yaml` | `docker/Start-Codeman.sh` on the host |
| New key in `docker/.env.example` | Add it to `docker/.env`, then `Start-Codeman.sh` |
The in-app updater detects all three of the bottom rows itself and refuses with a
message naming what changed, so you never have to work out which case you are in.
## Why the container needs its own path
The bare-host updater does `git checkout <tag> && npm install && npm run build`,
then asks systemd or launchd to restart the service. Two of those assumptions are
false in a container:
1. **There is no init system.** A container's supervisor is the Docker daemon,
which acts on the container, not on processes inside it.
2. **The image is immutable.** A `git pull` into the image's baked `/opt/codeman`
would land in the container's writable layer, survive `docker restart`, and be
silently discarded by the next `docker compose up`.
Both are solved by configuration rather than by a second updater:
- **The checkout is a host bind mount.** `docker-compose.yaml` mounts the repo
(the same directory used as the build context) over `/opt/codeman`, so the
updater's `git checkout` writes to the host filesystem and survives the
container being recreated.
- **The restart is the server exiting.** `restart: unless-stopped` relaunches the
container whenever its main process ends, including on a clean exit — so the
updater's final step is to signal the server, and Docker starts it again on the
freshly built `dist/`.
Everything else — the release-tag channel, the auto-stash, the atomic
`update-status.json` the browser polls across the connection drop, the boot-time
reconcile that flips `restarting` to `completed` — is the existing machinery,
unchanged. The container path is a new `SupervisorKind`, not a new updater.
## What the pieces are
| Piece | Role |
| ---------------------------------------------- | ------------------------------------------------------------------- |
| Repo bind mount at `/opt/codeman` | Makes the pull persistent. Without it, self-update is unavailable. |
| `codeman-node-modules`, `codeman-dist` volumes | Container-owned build artefacts, layered over the bind mount. |
| `CODEMAN_IN_CONTAINER=1` | Tells `detectSupervisor()` to restart by exiting. |
| `restart: unless-stopped` | Turns that exit into a restart. Verified before every update. |
| `CODEMAN_RESTART_BY_EXIT=1` | The Compose file's declaration of that policy, so the updater may exit even with no Docker socket. |
| Toolchain + devDependencies in the image | Lets `npm install` and `npm run build` run inside the container. |
| `docker-env-applied.json` | Fingerprint baseline, written by `Start-Codeman.sh` on every start. |
### Why build artefacts are in named volumes
`node_modules` and `dist` are mounted as named volumes **on top of** the repo bind
mount. Without that, an update's `npm install` would write into the host checkout,
leaving container-compiled native modules (node-pty builds from source here) in a
directory that may also be used to run Codeman natively, and leaving `git status`
permanently noisy.
Docker seeds an empty named volume from the image, so the first start inherits the
image's already-built `node_modules` and `dist` and pays no bootstrap cost.
`docker compose down -v` is the supported reset: the next start re-seeds them.
### Why the runtime image carries a build toolchain
`npm run build` is `tsc` plus `esbuild`, both devDependencies, so the image no
longer runs `npm prune --omit=dev`. And `npm install` may rebuild node-pty, which
ships no Linux prebuild, so `python3`, `make` and `g++` are installed as well.
This is the real cost of in-place updates: a noticeably larger image than a
runtime-only one. It buys an update that takes about a minute instead of a full
image rebuild, and it is why `NODE_ENV=production` is paired with an explicit
`npm install --include=dev` in the updater.
## The environment gate
An in-place update applies **code only**. A restarted container reuses its existing
image and configuration, so a release that changes the environment cannot take
effect that way — and would half-apply: new code against an old environment. The
updater therefore checks the **target release's own files**, read straight out of
git with `git show <tag>:<path>` before anything is checked out.
### 1. `server.Dockerfile` changed, so the image must be rebuilt
Compared by sha256 against the fingerprint `Start-Codeman.sh` recorded when the
running container was built.
### 2. `docker-compose.yaml` changed, so the container must be recreated
Same mechanism. A restart cannot pick up a new mount, port or environment
variable; only recreating the container can.
### 3. `.env.example` gained keys your `.env` has no value for
The check that matters most, because **Compose will not tell you**. An unset
`${VAR}` interpolates to the empty string; Compose prints a warning to a terminal
nobody is watching and starts anyway. A new required setting therefore arrives as
a silently blank environment variable and misbehaves later, far from the cause.
The updater names the missing keys instead.
Commented-out lines in `.env.example` are deliberately *not* keys — that is how
the file marks optional overrides such as `# PUID=1000`, and counting them would
block updates on settings you are meant to leave alone.
### 4. A restart policy that would not bring the container back
Before signalling the server, the updater asks the Docker daemon for its own
container's restart policy. If it is `no`, the update is refused: applying it
would take Codeman down and leave no UI to recover from.
If the policy cannot be read at all (no Docker socket mounted) the update is
still allowed, but the final step changes: the server exits only when the
Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (the shipped one does, because
it is the file that sets `restart: unless-stopped`) or the daemon confirmed an
auto-restart policy. Otherwise the build completes and the panel asks you to
restart the container by hand. A container started by plain `docker run` with no
restart policy therefore gets a staged update, never an outage.
### What the gate deliberately does not do
Every unknown fails **open**:
- A missing fingerprint baseline (a container started before this feature existed)
is not treated as a change, or those installs could never update at all.
- An unreadable `.env`, an unreachable Docker socket, or a target tag whose files
cannot be read all yield "no blocker" rather than a refusal.
The one place an unknown does NOT fail open is the kill itself: with neither the
Compose declaration nor a daemon answer, the updater stages the build and asks
for a manual restart rather than exiting a server nothing may bring back.
The gate catches a specific, detectable class of mistake; it is not a last line of
defence. It is also re-evaluated server-side on `POST /api/system/update`, so
hiding the button in the UI is a courtesy rather than the control.
## The one residual risk
The gate is derived from the diff, so it cannot see a release that needs a newer
environment **without changing any of those files** — for example, code that
depends on newer agent-CLI behaviour.
That is why the four global CLIs in `server.Dockerfile` are **pinned**. Unpinned,
the versions a user ends up with are a function of when their image was built
rather than of any commit, and in-app updates make rebuilds rarer, which makes
that drift worse over time. Pinned, "this release needs a newer CLI" becomes a
Dockerfile change, which check 1 already detects. Bump them deliberately, as part
of a release.
The complementary merge-side guard is `test/docker-compose-env-parity.test.ts`,
which fails CI when a variable is added to `docker-compose.yaml` without an entry
in `.env.example`, or the reverse.
## Sequence of an in-place update
1. **Check** — `GET /api/system/update/check` finds the latest release tag, fetches
that one ref so the gate can read the target's files, and returns any blockers.
2. **Start** — `POST /api/system/update` re-evaluates the gate, writes `queued` to
`update-status.json`, stages `self-update.sh` outside the repo and runs it.
3. **Apply** — stash if dirty, fetch the tag, check it out, `npm install
--include=dev`, `npm run build`. A failure at any step rolls back to the
previous commit, rebuilds it and reports `failed`; the server is never
restarted into a broken build.
4. **Restart** — write the terminal `restarting` marker, then signal the server.
The container exits and Docker restarts it.
5. **Reconcile** — the rebooted server compares its own version against the target
and flips the status to `completed` or `failed`. The browser, still polling,
picks that up.
Step 4 kills the updater script along with the container — unlike the systemd
path, it does not outlive the restart. That is safe only because the terminal
marker is written first, which is why nothing may be appended after the kill.
## Troubleshooting
**"This install can't update itself (unknown)"** — the repo bind mount is missing,
so the container is running the baked image copy. Check `CODEMAN_REPO_PATH` and
confirm the mounted directory really contains `.git`.
**The update fails immediately with a git ownership or permission error** — the
mounted checkout belongs to a different user than the one Codeman runs as
(`PUID`), so git refuses it as "dubious ownership". `Start-Codeman.sh` warns
about this at start; fix it by chowning the checkout to the same account that
owns `CODEMAN_APPDATA_PATH`.
**A rebuild is reported as required every time** — the fingerprint baseline does
not match the checkout. `Start-Codeman.sh` writes it on every start, so start
through that script rather than a bare `docker compose up` after either file
changes.
**Codeman does not come back after an update** — the build succeeded, since the
updater gates the restart on it, so read the container logs with `docker compose
logs codeman`. To roll back, check out the previous tag in the host checkout and
run `docker/Start-Codeman.sh`.
**The update failed during `npm install`** — most likely a native rebuild with no
toolchain, meaning the image predates the toolchain being added. Rebuild once from
the host and the in-app path works from then on.
**Resetting the build artefacts** — `docker compose down -v`, then
`Start-Codeman.sh`. This discards the named volumes and re-seeds them from a fresh
image.
## Disabling it
Set `CODEMAN_DISABLE_SELF_UPDATE=1` in `docker/.env` and pass it through in the
compose file's `environment:` block. The Updates panel then reports that in-app
updates are disabled, and the host-side script is the only way to update.
+31 -11
View File
@@ -228,6 +228,14 @@ window: the wait resolves on `idle` in a couple of seconds with `timedOut: false
indistinguishable from a finished turn. Wait for the pid, then wait for the
composer, answering the dialog only as the bounded fallback.
⚠️ **Answering it is not "press Enter".** Claude Code 2.1.252 dropped the options'
numbers, reversed them, and highlights `No, exit` by default, so a blind `\r` quits
the CLI and the pane is dead seconds after the spawn. Read the `❯` marker off the
rendered pane (`GET .../terminal?full=1`), send `ESC [ B` while it sits on `No, exit`,
re-read, and confirm only once the marker is on `Yes, I trust this folder`. Codeman's
own auto-accept (`trustDialogNextKey()` in `src/session-trust-dialog.ts`) does exactly
this, inside a 90 s startup window and a 6-keystroke cap.
A worked orchestration: start a worker, get it ready, prompt it, wait, clean up.
```bash
@@ -244,24 +252,36 @@ SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" \
[ -n "$SID" ] && [ "$SID" != null ] || { echo "quick-start failed"; exit 1; }
# 2. READINESS: composer marker first, trust dialog only as the bounded fallback.
# Skip this and step 3 reports a turn that never ran. Do NOT probe trust first
# and Enter blindly: the dialog text stays in the buffer for the life of the
# session, so on every later run that probe matches stale text and the Enter
# lands in a ready composer. Match single tokens only: TUI text can arrive
# without its spaces. Stage 1 is short on purpose (an already-trusted case
# matches in <1 s; a first-run case can never pass it and pays it in full).
# Skip this and step 3 reports a turn that never ran. Match single tokens only:
# TUI text can arrive without its spaces. Stage 1 is short on purpose (an
# already-trusted case matches in <1 s; a first-run case can never pass it and
# pays it in full).
# ⚠️ NEVER answer the dialog with a bare \r. Its highlighted option is `No, exit`
# (claude-cli 2.1.252), so a blind Enter quits the CLI; and the dialog text stays
# in the buffer for the life of the session, so a `from=buffer` probe for `trust`
# keeps matching long after it is gone. Read the CURRENT pane instead and steer.
until [ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ]
do sleep 1; done
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=5000') # composer's status bar = ready
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
ESC=$(printf '\033') # \x1b is GNU-sed only; this form also works on macOS
for _ in 1 2 3 4 5 6; do
# Which option the ❯ marker sits on, read off the CURRENT frame.
K=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$ESC\[[0-9;?]*[a-zA-Z]//g" -e "s/$ESC[()][AB0]//g" | tr -d ' \t' \
| grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/')
[ -n "$K" ] || break # no dialog on screen: nothing to answer
[ "$K" = confirm ] && IN="\r" || IN="$ESC[B"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' -d '{"input":"\r","useMux":true}' >/dev/null
-H 'Content-Type: application/json' \
-d "$(jq -nc --arg i "$IN" '{input:$i,useMux:true}')" >/dev/null
[ "$K" = confirm ] && break
sleep 1 # re-read: confirm the arrow landed
done
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=45000' >/dev/null
+106
View File
@@ -0,0 +1,106 @@
# Grok Build (xAI) integration plan
> **Status**: Executed. This document records the plan, the decision behind each wiring
> point, and what was and was not verified. The user-facing guide is
> [`grok-integration.md`](./grok-integration.md); the per-decision invariants live in
> [`architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok`](./architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok).
> Template: the pi integration (`c5b5963`, [`pi-integration-plan.md`](./pi-integration-plan.md)),
> which was itself calibrated against the four follow-up commits the antigravity
> integration needed. All of grok's facts below were verified against **grok 1.0.5**
> (`grok 1.0.5 (5115b46bc9)`), installed live during the work.
## 1. What Grok Build is
[xai-org/grok-build](https://github.com/xai-org/grok-build) is xAI's coding agent: a
Rust fullscreen-TUI binary named `grok`, installed by
`curl -fsSL https://x.ai/cli/install.sh | bash` into `~/.grok/bin` (with symlinks into
`~/.local/bin`; the installer also ships an `agent` alias). Config lives in
`~/.grok/config.toml`, TUI appearance in `~/.grok/pager.toml`, credentials in
`~/.grok/auth.json` (0600), sessions under `~/.grok/sessions/`. Auth is browser OAuth
on first launch, `grok login --device-auth` for SSH boxes, or `XAI_API_KEY` for
headless use. It has Claude-style permission modes (`default`/`acceptEdits`/`auto`/
`dontAsk`/`bypassPermissions`/`plan`), allow/deny rules, hooks, MCP, subagents, and a
headless `-p` mode.
## 2. Shape decisions (why grok is wired the way it is)
Grok is a seventh run mode, alongside Claude Code, shell, OpenCode, Codex, Gemini,
Antigravity and Pi. Never a location overlay, never a web tab. Its wiring mixes two
existing shapes:
| Question | Decision | Why |
| --- | --- | --- |
| Permission bypass | `GrokConfig.alwaysApprove` -> `--always-approve` | Grok's real flag (verified via `--help`): "Auto-approve all tool executions", i.e. its `bypassPermissions` mode. Config-level deny rules still apply on top. The Run button sends `true`, matching `runAntigravity()` and Claude's own `--dangerously-skip-permissions` default: Codeman sessions exist for autonomous work. |
| Multi-user clamp branch | only-if-sent (codex/antigravity branch) | A bare `grok` spawn is grok's own ask-mode default, which is already safe, so the clamp only needs to force a SENT `alwaysApprove` off. Contrast pi, whose absent default is an answerable prompt and therefore needs the materialize branch. Cron needs nothing for grok for the same reason (`clampCronExternalCliConfigs`). |
| Alt-screen strip | OUT of `isAltScreenStripMode()` | Grok is a fullscreen alternate-screen TUI with mouse support (its own scrollback pane, `pager.toml [terminal] alt_screen`), i.e. the opencode case, not the Ink repaint case. It falls through to the narrow tmux-attach strip like opencode/antigravity/pi. |
| Resolver | version probe, like pi | `grok` has npm squatters (the unrelated `@vibe-kit/grok-cli` installs a `grok` bin). Candidates must pass `grok --version`; `GROK_VERSION_REGEX` is exported and shared with the dependency registry so doctor and run mode cannot disagree. The probe cannot tell two version-printing `grok`s apart, so `GET /api/grok/status` surfaces path AND version. Search dirs: `~/.grok/bin` first (installer target), then `~/.local/bin`, `/usr/local/bin`, `~/bin`. |
| Env allowlist | `GROK_*` + `XAI_*` prefixes | `GROK_*` covers grok's documented inputs (`GROK_HOME`, `GROK_CONFIG`/`GROK_CONFIG_PATH`, `GROK_MEMORY`, `GROK_WORKFLOWS`, `GROK_SANDBOX`, `GROK_OIDC_*`, `GROK_AUTH_PROVIDER_COMMAND`). `XAI_*` is xAI's vendor namespace and carries `XAI_API_KEY`, grok's documented headless auth var: the same narrow-vendor-namespace reasoning that admitted `GOOGLE_*` for gemini. Foreign provider keys stay out, as always. |
| Resume | `--resume <id>` / `--continue`, id-regexed | Grok's `--resume` also matches session TITLES (arbitrary user strings, case-insensitive). The `^[a-zA-Z0-9._-]+$` regex doubles as the no-titles rule, so nothing free-form can reach the `bash -c` spawn line. A valid explicit id wins over `-c`, mirroring pi. |
| Local echo | `'buffer'` via the `_updateLocalEchoState` fallthrough | UNMEASURED against an authenticated session (see §4). If grok's composer turns out per-keystroke reactive like codex's, the fallback is one `'off'` branch; teaching `PredictiveEchoAddon` grok's composer row is the larger follow-up. |
| Truecolor | `COLORTERM=truecolor` + `unset NO_COLOR` | Rust TUI with themes; joins the codex/gemini/antigravity/pi list in `buildEnvExports()` and `buildMuxAttachEnv()`. |
| Docker credentials | per-file seed: `auth.json`, `config.toml`, `pager.toml` | `~/.grok` also holds `sessions/`, `memory/`, `completions/`, `docs/` and the ~160MB binary under `downloads/`; a whole-dir seed would copy all of it on every container start. Same trade-off as pi: in-container sessions are invisible host-side, so `grok -c` in a Docker case sees only that container's history. |
| Docker install | own Dockerfile step | Not an npm package. xAI's installer has no `--dir` override, so the step copies `/root/.grok/bin/grok` (through the symlink, `cp -L`) into `/usr/local/bin` and removes root's `~/.grok` in the same layer. |
| Remote SSH | `exec "$SHELL" -i -l -c 'grok'` | sshd's remote-command PATH does not include `~/.grok/bin`; same login-shell fix as every other agent CLI. |
| What is NOT wired | `--permission-mode`, `--allow`/`--deny`, `-p` headless, `--worktree`, `--sandbox`, `--reasoning-effort`, `-s/--session-id`, `--fork-session`, `--agent`, `--output-format` | Follow-ups. The flag surface is kept minimal on purpose; grok is pre-1.0-style fast-moving and every flag added is a flag validated forever. |
## 3. Touch points (the checklist)
Backend: `types/session.ts` (SessionMode + GrokConfig + SessionState), `utils/grok-cli-resolver.ts` (new)
+ barrel, `tmux-manager.ts` (`buildGrokCommand`, dispatch, resume flag, PATH export, truecolor,
availability error, plumbing), `session.ts` (external-mode gate, label, config plumbing,
tmux-required error, attach env), `mux-interface.ts`, `schemas.ts` (prefixes, `GrokConfigSchema`,
both mode enums, remote command overrides, cron agentType), `session-routes.ts` (clamp + both
create paths), `system-routes.ts` (`GET /api/grok/status`), `server.ts` (availability inject +
mux restore), `docker-hosts.ts`, `remote-hosts.ts`, `config/dependency-registry.ts`,
`cron/cron-service.ts` (comment), `response-viewer-transcript.ts`, `tui/tui-client.ts` + `tui-app.ts`.
Frontend: `index.html` (welcome button, run-mode entry, cron option, clone Brain option),
`session-ui.js` (`runGrok()`, dispatch, availability, "Run GK" label, external-CLI gates,
runMode setter), `app.js` (label, `gk` tab badge, kill-menu), `settings-ui.js`,
`mobile-overview.js`, `home-sessions.js`, `panels-ui.js`, `i18n.js`, `styles.css` +
`mobile.css` (charcoal monochrome identity; the non-og skin block and the mobile
`!important` pair are both load-bearing, see the pi plan's §2.9 cascade trap).
Meta: `docker/agent.Dockerfile`, `install.sh`, `package.json` keyword, changeset,
`skills/codeman/reference/*`, CLAUDE.md, READMEs, `architecture-invariants.md`,
`remote-sessions.md`, `security-architecture.md`, `docker-cases.md`, `cron-guide.md`.
Tests: `test/grok-mode.test.ts` + `test/grok-cli-resolver.test.ts` (new);
`external-cli-bypass-clamp`, `system-routes`, `render-index-html`, `run-mode-ui`,
`mobile-overview`, `local-echo-codex-gating` (extended).
## 4. Verification performed
On this box, with grok 1.0.5 really installed and an isolated
`CODEMAN_INSTANCE=grokwt` server (own data dir, own tmux socket, port 5077):
1. `npm test` (the CI gate): green, 5900+ tests. `typecheck`, `lint`, `format:check`,
`check:frontend-syntax`, `check:public-assets`, `check:lockfile`: green.
2. `GET /api/grok/status` -> `{available: true, path: "/home/arkon/.local/bin", version: "1.0.5"}`
through the real resolver and probe.
3. `POST /api/quick-start {mode: "grok", grokConfig: {alwaysApprove: true}}` -> session
created, tmux pane spawned, real spawn line verified to end in `grok --always-approve`,
and the actual grok TUI rendered its OAuth device-approval screen in the pane
(unauthenticated box, so sign-in is exactly where a first run lands).
4. `grokConfig` persisted into the instance's `state.json`.
5. Session deleted by exact id; instance data dir and throwaway case removed.
**Not verified (honest gaps, all requiring an xAI account or more hardware):**
an authenticated conversation end to end; the local-echo buffer policy against grok's
real composer (§2); scrollback/repaint behavior of the fullscreen TUI under the narrow
strip during a long session; a Docker case with `mode: 'grok'` (needs a `--no-cache`
agent-image rebuild); a remote-SSH grok case; cron readiness degradation (expected:
same slow-start-then-send as pi, documented in `cron-guide.md`).
## 5. Follow-ups
- Idle/completion signal: grok has a hooks system (user-guide `10-hooks.md`); a hook
POSTing to `/api/hook-event` could give grok sessions real idle detection instead of
output-stabilization. Highest-value follow-up, same slot as pi's `agent_settled` idea.
- Response viewer: sessions are ACP JSONL under `~/.grok/sessions/<encoded-cwd>/<id>/updates.jsonl`;
`grok -p ... --output-format json | jq -r '.sessionId'` exists for correlation.
- Permission-mode picker (`--permission-mode`, `--allow`/`--deny`) in Session Options.
- Measure the local-echo policy and the fullscreen-TUI scrollback behavior against an
authenticated session; pin the result in `local-echo-codex-gating` the way pi did.
- `grok doctor` is a built-in terminal-support check worth pointing users at when a
pane renders oddly.
+133
View File
@@ -0,0 +1,133 @@
# Grok Build (xAI) sessions
Codeman can drive [Grok Build](https://github.com/xai-org/grok-build) (xAI's `grok`
CLI, the agent behind docs.x.ai/build) as a session backend, alongside Claude Code,
OpenCode, Codex, Gemini, Antigravity and Pi. `grok` is a seventh **run mode**: its own
PTY, its own tmux session, its own tab identity (monochrome charcoal, `gk` badge). It
is not a location overlay like Docker or remote-SSH cases, and it is not a web tab.
The design rationale behind each decision below lives in
[`grok-integration-plan.md`](./grok-integration-plan.md). Everything here was verified
against grok 1.0.5.
## Install
```bash
curl -fsSL https://x.ai/cli/install.sh | bash
```
The installer places the binary in `~/.grok/bin` and symlinks it into `~/.local/bin`
(it also installs an `agent` alias Codeman ignores). `grok update` self-updates.
Codeman resolves the binary via the server PATH and then the usual install locations,
`~/.grok/bin` first. **`grok` is a name with known squatters** (the unrelated
`@vibe-kit/grok-cli` npm package also installs a `grok` bin), so like `pi` the
resolver does not trust a PATH hit on its own: it runs `grok --version` once and
requires version-shaped output (`grok 1.0.5 (5115b46bc9)`). Check what it resolved:
```bash
curl -s localhost:3000/api/grok/status | jq
# { "available": true, "path": "/home/you/.grok/bin", "version": "1.0.5" }
```
The endpoint carries `version` on top of the sibling `/api/*/status` shape precisely
so a misresolution is visible rather than presenting as "the mode just doesn't work".
## Authenticate
- **Browser OAuth (default)**: the first `grok` run opens a sign-in flow; in a
Codeman pane you get the device-code screen with a URL to open elsewhere.
Credentials land in `~/.grok/auth.json` (0600) and refresh automatically.
- **Device code**: `grok login --device-auth`, made for SSH boxes and headless hosts.
- **API key**: `export XAI_API_KEY="xai-..."` (console.x.ai). Used as a fallback when
no session token exists. As a per-session Codeman `envOverride` it flows through
socket-scoped `tmux setenv`, never the spawn command line.
- **Enterprise OIDC**: `GROK_OIDC_ISSUER` / `GROK_OIDC_CLIENT_ID`.
## What Codeman wires up
`GrokConfig` (per session, persisted in `state.json`, round-trips through respawn):
| Field | Flag | Notes |
| ----------------- | --------------------------- | --------------------------------------------------------------------- |
| `model` | `--model <v>` | e.g. `grok-4.5`, or a custom `[model.<name>]` from `config.toml` |
| `alwaysApprove` | `--always-approve` | Grok's `bypassPermissions` mode; deny rules still apply on top |
| `continueSession` | `--continue` | Most recent session for the working directory; skipped when resuming |
| `resumeSessionId` | `--resume <v>` | Ids only, never titles (grok's own `--resume` also matches titles) |
Every value is regex-validated and **dropped** (not escaped) if it fails, because the
result is interpolated into the pane's `bash -c "..."` command.
The Run button sends `grokConfig: { alwaysApprove: true }`, the same product decision
as Claude's `--dangerously-skip-permissions` default and Antigravity's
`--dangerously-skip-permissions`: Codeman sessions exist for autonomous work. Keep
hard limits as `deny` rules in `~/.grok/config.toml` (they apply in every mode), and
in **multi-user mode** a non-granted owner's `alwaysApprove` is forced off
server-side; a bare `grok` spawn is grok's own ask-mode default.
Env overrides: the `GROK_*` prefix (`GROK_HOME`, `GROK_CONFIG`, `GROK_MEMORY`,
`GROK_WORKFLOWS`, `GROK_SANDBOX`, `GROK_OIDC_*`, ...) plus the `XAI_*` vendor
namespace (`XAI_API_KEY`) are allowlisted. Foreign provider keys are not, as ever.
## What Codeman deliberately does NOT wire up
- **`--permission-mode`, `--allow`/`--deny`.** The boolean covers the autonomous
case; the full rule surface is a follow-up with UI.
- **`-p`/headless, `--output-format`, `--json-schema`.** Codeman drives the TUI.
- **`--worktree`, `--sandbox`, `--reasoning-effort`, `-s/--session-id`,
`--fork-session`, `--agent`/`--agents`.** Tracked as follow-ups in the plan doc.
## Terminal behavior
Grok renders a **fullscreen alternate-screen TUI** (scrollback pane + prompt, mouse
supported). Under Codeman it runs inside tmux like every external CLI, so the
fullscreen rendering stays inside the pane and the browser terminal shows tmux's
repaints; grok stays out of the alt-screen strip list on purpose (the opencode case,
not the Ink case). If a pane renders oddly, `grok doctor` checks terminal, color and
input support without starting a session, and `~/.grok/pager.toml` can force
`alt_screen = "inline"`.
On touch devices grok currently gets the buffered local-echo overlay like Claude,
Gemini, OpenCode and Pi. This is the fallthrough default and has not been measured
against an authenticated grok composer; if grok turns out per-keystroke reactive the
way codex was (issues #218/#219/#220/#222), the fix is the `'off'` branch in
`_updateLocalEchoState` (terminal-ui.js).
## Docker cases
The agent image installs grok in its own Dockerfile step (not npm; xAI's installer
targets `$HOME/.grok/bin` with no `--dir` override, so the binary is copied to
`/usr/local/bin`). Rebuild with the mandatory `--no-cache`:
```bash
node scripts/build-agent-image.mjs --no-cache
```
Credentials are **seeded**, not shared: `auth.json`, `config.toml` and `pager.toml`
are copied into the container's own `~/.grok`, so an in-container grok never writes
refreshed OAuth tokens back to the host and `docker commit` exports stay secret-free.
Only those three files, because `~/.grok` also holds `sessions/`, `memory/` and the
~160MB binary under `downloads/`. Trade-off, same as pi: in-container sessions are
invisible host-side, so `grok -c` inside a Docker case only sees that container's own
history.
## Remote SSH cases
`grok` mode is routed through an interactive login shell
(`exec "$SHELL" -i -l -c 'grok'`), because sshd's remote-command PATH does not include
`~/.grok/bin`. Per-session config and `envOverrides` do not cross ssh and are rejected
rather than silently ignored; use the per-host command override instead. For auth on
the remote host, `grok login --device-auth` exists for exactly this.
## Known gaps
- **No idle/completion hook yet.** Idle detection falls back to output-stabilization
like the other external CLIs. Grok has a hooks system, so a Codeman hook POSTing to
`/api/hook-event` is the highest-value follow-up.
- **No response viewer.** Grok writes ACP JSONL sessions under
`~/.grok/sessions/<encoded-cwd>/<session-id>/updates.jsonl`; nothing reads them yet.
- **Cron jobs mis-detect readiness.** The readiness poll looks for `❯` or a token
count, neither of which grok prints, so a grok cron job burns its poll budget and
then sends the prompt anyway. It works; it is just slower to start.
- **Ralph, respawn heuristics, token/CLI-info parsing and the `❯` readiness probe are
off** for grok, as for every external CLI.
+179
View File
@@ -0,0 +1,179 @@
# OMP (Oh My Pi) sessions
Codeman can drive [OMP](https://github.com/can1357/oh-my-pi) (`omp`, Oh My Pi) as a session
backend, alongside Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi, Grok and
DeepSeek Harness. `omp` is the ninth CLI backend (tenth `SessionMode`, counting
`shell`): its own PTY, its own tmux session, its own tab identity. It is not a
location overlay like Docker or remote-SSH cases, and it is not a web tab.
## Install
```bash
curl -fsSL https://omp.sh/install | sh
```
The installer places the binary in `~/.local/bin` (verified against a real
`--no-cache` Docker build — see `docker/agent.Dockerfile`; an earlier guess of
`~/.omp/bin` was wrong). Codeman resolves the binary via the server PATH and then
the usual install locations (`~/.local/bin` first, then `~/.omp/bin`,
`/usr/local/bin`, `~/.bun/bin`, `~/.npm-global/bin`, `~/bin`).
**`omp` is a short name**, so like `pi` and `grok` the resolver does not trust a PATH
hit on its own: it runs `omp --version` and requires `omp/<semver>`-shaped output
(e.g. `omp/18.0.8`) before accepting a candidate. Check what it resolved:
```bash
curl -s localhost:3000/api/omp/status | jq
# { "available": true, "path": "/home/you/.local/bin", "version": "18.0.8" }
```
## Authenticate
OMP owns its own auth and provider configuration entirely in `~/.omp` — there is
no Codeman-side login flow, API key field, or bypass switch to configure. Run `omp`
directly once outside Codeman to complete whatever onboarding the CLI itself asks
for; every session started through Codeman afterward inherits that config.
## What Codeman wires up
`OmpConfig` (per session, persisted in `state.json`, round-trips through respawn):
| Field | Flag | Notes |
| ------------------ | --------------- | ---------------------------------------------------------- |
| `model` | `--model <v>` | Regex-validated (`[a-zA-Z0-9._-/]+`); `provider/model` forms like `crof/glm-5.2` pass |
| `continueSession` | `--continue` | omp's own "most recent conversation in this directory" heuristic |
| `resumeSessionId` | `--resume <id>` | Ids only, id-regexed; wins over `--continue` when both are present |
Every value is regex-validated and **dropped** (not escaped) if it fails, because the
result is interpolated into the pane's spawn command.
**omp reads its own model routing and hooks from `~/.omp`, so no trust or
permission flags are needed** — unlike every sibling CLI in this family, there is no
bypass-permissions equivalent to wire up, so `buildOmpCommand()` only ever passes
`--model`/`--resume`/`--continue`. ⚠️ That does NOT mean omp is unrestricted: its
documented default `tools.approvalMode` is `yolo`, so an omp pane auto-approves exec
with no flag from Codeman — the CLI's own config, not Codeman, is what would need to
change that.
Env overrides: the `OMP_*` prefix is allowlisted, and per omp's own
`docs/environment-variables.md` it is not the narrow surface it looks like. omp reads
roughly 40 provider keys from the environment (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`,
`XAI_API_KEY`, `HF_TOKEN`, ...) — pi's 34-key problem in the same shape — which is why
none of those get a dedicated allowlist entry; a session authenticates from `~/.omp`
config or the server process's own env instead, like pi. omp's own documented knobs
are mostly `PI_*`, not `OMP_*` (`PI_CONFIG_DIR`, `PI_CODING_AGENT_DIR`,
`PI_CODING_AGENT_SESSION_DIR`, `PI_SUBPROCESS_CMD`, `PI_SHELL_PREFIX`,
`OMP_PROFILE`/`PI_PROFILE`), and `PI_*` is already allowlisted globally because pi
mode needs it — so an omp session today already accepts all of those. The first three
also move the tree `omp-session-resolver.ts` and `omp-transcript.ts` hardcode
(`resolveOmpHome()` assumes `~/.omp` unconditionally), so pinning and history quietly
stop working under a redirected config root; this is a known gap, not fixed here.
The `OMP_` prefix itself brings in `OMP_AUTH_BROKER_URL` / `OMP_AUTH_BROKER_TOKEN`,
where omp resolves credentials from — the same shape `DEEPSEEK_BASE_URL` is dropped
for in `clampEnvOverridesForOwner()` (session-routes.ts), so both are clamped there
for a non-granted owner in multi-user mode. None of this matters in single-user mode.
## Exact-id pinning: why `--resume`, not just `--continue`
`--continue` alone is ambiguous the moment **any** other omp conversation has
touched the same working directory more recently — it just picks the newest session
file on disk, silently. That happens routinely: a closed-then-resumed Codeman row
plus a still-running duplicate, two Codeman sessions pointed at the same case, or a
plain reattach after a server restart.
`src/utils/omp-session-resolver.ts` resolves and **pins** the exact conversation id
once (`findLatestOmpSessionId()` reads `~/.omp/agent/sessions/<mangled-workingDir>/`,
the newest `.jsonl` file's embedded uuid), then every later respawn reuses that
pinned id via `--resume` instead of re-guessing with `--continue`.
⚠️ **The directory mangling is NOT a straight `/` → `-` replace.** Unlike Claude
Code's `~/.claude/projects/*` convention (which keeps the full path, e.g.
`-home-user-codeman-cases-foo`), omp strips the `$HOME` prefix FIRST and only then
dash-replaces (`/home/user/codeman-cases/foo` → `-codeman-cases-foo`; a path outside
`$HOME`, like `/tmp/...`, is dash-replaced as-is with no stripping). Getting this
wrong doesn't error — `findLatestOmpSessionId()` just silently returns null for
every case under `$HOME` (virtually all real Codeman cases), so pinning quietly
degrades to omp's own ambiguous `--continue`. This was found and fixed 2026-08-27
after months of testing had only ever exercised `/tmp`-based working directories,
where the bug's wrong output happened to coincidentally match the right one.
## Surviving a full session kill
`src/omp-transcript.ts` scans `~/.omp/agent/sessions/**/*.jsonl` directly — a second,
independent history source alongside Codeman's own state. This means an OMP
conversation's history (working directory, first/last prompt, size) is recoverable
in the Past Sessions list even when **both** the Codeman session record and the
underlying tmux pane are gone — verified live against a full OS reboot, not just a
"Kill Tmux" button click.
## Terminal behavior
OMP renders inside tmux like every external CLI (narrow scrollback strip — alt-screen
toggles only, not the full Claude/Codex/Gemini strip). It stays out of the
alt-screen-strip list and lands on the `'buffer'` local-echo policy via the
`_updateLocalEchoState` fallthrough, same as grok and pi.
## Docker cases
The agent image installs omp in its own Dockerfile step (not npm; omp's installer
targets `$HOME/.local/bin` with no `--dir` override, the same shape as grok's
installer). Rebuild with the mandatory `--no-cache`:
```bash
node scripts/build-agent-image.mjs --no-cache
```
⚠️ **`--resume` pinning does not currently reach an in-container omp process.**
Docker panes are built from `defaultDockerCommandForMode`, which never sees
`ompConfig` — `appendResumeFlag()`'s `case 'omp'` keys off the top-level
`resumeSessionId` field, which nothing populates for omp today. Host-side history
recovery still works (the shared `sessions/` mount below), but a respawned
in-container omp pane falls back to its own ambiguous `--continue`, not a pinned
id. Flagged in upstream review, not yet fixed.
Credentials are **mostly seeded**, but `sessions/` is the one exception in this CLI
family: `~/.omp/agent/{config.yml,mcp.json,models.yml,settings.yml}` are seeded
(read-only mount, copied into the container's own `~/.omp/agent` once), so an
in-container omp never writes refreshed config back to the host and `docker commit`
exports stay secret-free. But `~/.omp/agent/sessions/` is **shared (RW)**, not
seeded — the same treatment as codex's `sessions/`, and for the identical reason:
Codeman reads it host-side (`omp-transcript.ts`, `omp-session-resolver.ts`) for
history recovery and `--resume` pinning. Seeding it instead of sharing it would make
an in-container OMP conversation invisible to Codeman's own history/resume logic,
silently breaking Docker support for the kill-survival feature above. The rest of
`~/.omp/agent` (`agent.db`/`history.db`/`models.db` SQLite caches,
`terminal-sessions/`, `blobs/`, `cache/`) stays container-local and is neither
shared nor seeded.
## Remote SSH cases
`omp` mode is routed through an interactive login shell
(`exec "$SHELL" -i -l -c 'omp'`), because sshd's remote-command PATH does not
include `~/.local/bin`. Per-session config and `envOverrides` do not cross ssh and are
rejected rather than silently ignored; use the per-host command override instead.
⚠️ A **respawn or reattach** of a remote omp session runs `omp --continue`, not a
bare `omp`, so it lands back in the same conversation. It is deliberately
`--continue` rather than the exact `--resume <id>` the local and docker paths
pin: `omp-session-resolver.ts` only ever reads THIS host's `~/.omp/agent/sessions/`,
and a remote conversation's session file lives on the remote host under the
remote user's home, so resolving locally would pin a stranger's id. See
[Respawn / reattach continuation](remote-sessions.md#respawn--reattach-continuation).
## Known gaps
- **No idle/completion hook.** Idle detection falls back to output-stabilization
like every other external CLI. If omp ever ships a hooks system, a Codeman hook
POSTing to `/api/hook-event` would be the highest-value follow-up.
- **Killing a pane mid-turn loses the conversation for real.** `tmux kill-session`
before an in-TUI `/exit` beats omp's own session-file flush — confirmed by direct
testing (kill after a clean `/exit` resumes correctly; kill without `/exit` first
does not). This is not something Codeman can compensate for from outside the
process; it would need an upstream omp fix (e.g. flush-on-SIGTERM).
- **Unverified: `$HOME` as a symlink.** The directory-mangling fix above compares
against the literal `homedir()` string, not a `realpath()`-resolved one. Whether
omp itself canonicalizes symlinks before mangling is unconfirmed — this has not
been tested against a symlinked-home setup.
- Ralph, respawn heuristics, token/CLI-info parsing and the `❯` readiness probe are
off for omp, as for every external CLI.
+1 -1
View File
@@ -139,7 +139,7 @@ set -g extended-keys-format csi-u
Codeman's browser input path sends `\r` for submit, so basic use works
unconfigured — what degrades is newline-in-editor, mostly when you attach to the
pane directly (`sc`).
pane directly (`codeman tui`).
⚠️ Upstream notes the setting may need a full `tmux kill-server` to take effect.
**Never run `tmux kill-server` on Codeman's socket** — it would kill every live
+145
View File
@@ -0,0 +1,145 @@
# PR bot: automatic pull-request reviews, reported over Telegram
The PR bot is maintainer tooling that lives in `scripts/pr-bot/`. It watches the
repository's open pull requests, reviews each one in a Codeman claude session running in
a private clone of the repository, and sends the verdict to a Telegram chat with the ranked
findings, a recommendation and action buttons. The maintainer decides what happens next
from the phone: merge, post the drafted review comment, close, approve a waiting CI run,
or ask the reviewer session a follow-up question.
It reviews on its own. It never writes to GitHub on its own.
## How a review runs
1. Every poll (default 10 minutes) the bot lists open PRs with `gh`. A PR is queued
when its head commit differs from the one last reviewed, so a push re-reviews and an
untouched PR is never reviewed twice. Draft PRs and bot PRs are skipped. The backlog
is ordered mergeable-and-small first, conflicting-and-huge last.
2. The PR head is fetched into a private ref (`refs/pr-bot/<n>`) of the main repository
and checked out (detached) in a private clone under
`~/.codeman/pr-bot/worktrees/pr-<n>`, made with `git clone --shared` so the object
store stays shared and nothing is duplicated. The maintainer's own checkout is never
checked out or reset by the bot. A clone rather than a linked worktree because Claude
Code reads a linked worktree's project settings from the MAIN checkout, whose model
pin would silently override the bot's. `node_modules` is a symlink to the main
checkout's tree when the PR itself leaves the dependency files untouched (judged
against the PR's merge base, not against current master), and a real `npm ci`
otherwise (the symlink is unlinked first, so npm can never write through it; an
install interrupted by a restart is discarded, never reused).
3. A review brief is written to `~/.codeman/pr-bot/jobs/pr-<n>/brief.md`: the PR
metadata, CI state, mergeability, the file list, the body verbatim, the ground rules
(nothing reaches GitHub, no installs, no builds, no services, never port 3000), the
review protocol (CLAUDE.md and CONTRIBUTING first, then correctness, security,
invariants, tests, contract, scope), the checks to run, the verdict vocabulary and
the exact JSON to produce.
4. A Codeman session named `prbot-<n>` is created in the clone over the HTTP API,
the composer is awaited (the folder-trust dialog is read off the screen and answered
one key at a time), and one prompt points the session at the brief. The bot waits on
the `stop`/`blocked`/`exit` hook signals, never on the heuristic `idle`, with a hard
timeout (default 40 minutes).
5. The session writes `report.json` and `report.md` next to the brief and replies
`REVIEW COMPLETE`. The bot parses the JSON leniently, records the Claude session id
for follow-ups, deletes the Codeman session, keeps the clone, and sends the
summary to Telegram. Reviews run one at a time.
Verdicts: `merge`, `merge-with-fixes`, `request-changes`, `close`, `needs-discussion`.
Findings are ranked `blocker` / `major` / `minor` / `nit`, each with file and line.
## The Telegram side
Each review arrives as one message: PR number and title, author, size, CI state,
mergeability, the verdict with confidence, the summary, the top findings, the checks
that were run, the recommendation, and buttons:
| Button / command | What it does |
| --- | --- |
| 📄 Full report · `/report N` | Sends `report.md` (as a file when long). |
| 💬 Draft comment · `/draft N` | Shows the comment drafted for the contributor. Nothing is posted. |
| 📮 Post comment · `/post N` | Shows the draft again and asks for confirmation, then posts it under your GitHub account. |
| ✅ Merge · `/merge N` | Re-checks mergeability and CI, lists warnings (red CI, new commits since the review, a non-merge verdict), asks for confirmation, then merges with a merge commit. Refuses a conflicting PR. |
| 🗑 Close · `/close N reason` | Asks for the closing comment if none was given, asks for confirmation, then closes with that comment. |
| ▶️ Approve CI run · `/approve N` | Approves a workflow run that GitHub holds for a first-time contributor. Shown only when one is waiting. |
| 🔁 Re-review · `/review N` | Queues a fresh review at the front of the queue. |
| `/ask N question`, or reply to any review message | Resumes the reviewer's Claude conversation in the same clone and relays the answer. It can inspect, run checks, or make uncommitted changes there; it still never pushes. |
| `/status` · `/scan` · `/pause` · `/resume` · `/help` | Housekeeping. |
Merge, close and post always take a second tap. Confirmations expire after 15 minutes.
Only messages from the configured chat are acted on; anyone else gets silence.
When a PR is merged or closed, the bot announces it, removes the clone and the
private ref, and keeps the record.
## Setup
Requirements on the machine that runs the bot: a running Codeman (the sessions are
spawned there), `gh` logged in as the account that should merge and comment, `git`,
Node 22, and the repository checkout with its `node_modules`.
Config is `~/.codeman/pr-bot.env` (`KEY=VALUE`, keep it mode 0600). The Telegram token
and chat id are read from the existing notifier bot's env file
(`~/codeman-cases/telegram/.env`) when present, so on the maintainer's machine no key
has to be copied; set them here to use a different bot.
| Key | Default | Meaning |
| --- | --- | --- |
| `TELEGRAM_BOT_TOKEN` | from the shared env file | BotFather token. |
| `TELEGRAM_CHAT_ID` | from the shared env file | The one chat that receives reports and may issue commands. |
| `GITHUB_REPO` | `Ark0N/Codeman` | `owner/name`. |
| `CODEMAN_API_URL` | `https://127.0.0.1:3000` | The Codeman that spawns the review sessions. A self-signed certificate is accepted. |
| `CODEMAN_USERNAME` / `CODEMAN_PASSWORD` | unset | Only when that Codeman has a password. |
| `PR_BOT_POLL_INTERVAL` | `600` | Seconds between GitHub polls (minimum 60). |
| `PR_BOT_MAIN_CHECKOUT` | the repo this script is in | The repository the clones share objects with and fetch from. |
| `PR_BOT_DATA_DIR` | `~/.codeman/pr-bot` | State, briefs, reports, clones. |
| `PR_BOT_MODEL` | unset (the session default) | Codeman `modelOverride` for the review sessions, e.g. `opus[1m]`. ⚠️ Pick a model whose budget can absorb a re-review of every open PR on every head commit: when it runs out, Claude Code answers the limit **inside the turn** and the reviewer has nothing to write. The bot now names that failure in seconds (`findModelLimitNotice`) instead of burning the whole `PR_BOT_REVIEW_TIMEOUT`, and a limit does not spend the per-head retry budget, so the queue resumes by itself once the budget does. |
| `PR_BOT_EFFORT` | unset | Codeman `effort` for the review sessions. |
| `PR_BOT_REVIEW_TIMEOUT` | `40` | Minutes before a review is abandoned. |
| `PR_BOT_FOLLOWUP_TIMEOUT` | `20` | Minutes before a follow-up is abandoned. |
| `PR_BOT_AUTO_REVIEW` | `1` | `0` reviews only on `/review N`. |
| `PR_BOT_REVIEW_DRAFTS` | `0` | `1` reviews draft PRs too. |
| `PR_BOT_TELEGRAM_ENV_FILE` | `~/codeman-cases/telegram/.env` | Where the shared token and chat id are read from. |
```bash
npm run pr-bot -- check # config, gh, git, Codeman, Telegram, open PR count
npm run pr-bot -- scan # the open PRs in review order, with what is new
npm run pr-bot -- review 383 --no-telegram # one review now, printed instead of sent
npm run pr-bot -- run # the daemon
npm run pr-bot -- install-service # systemd user unit codeman-pr-bot, enabled and started
npm run pr-bot -- status # what the state file knows
tail -f ~/.codeman/pr-bot/bot.log # the service logs to a file, not the journal
```
## Safety properties worth knowing before changing it
- **GitHub writes happen in exactly one place** (`runConfirmed` in `bot.ts`) and only
after a confirmation tap on a nonce that expires. The review session's brief forbids
`gh` writes, pushes and merges, and the session has no reason to have the token
anyway: it runs as the same user as the maintainer's own sessions, so the prompt rule
is the guard, and the clone's checkout is detached so an accidental push has no
branch to land on.
- **The maintainer's checkout is shared with other agent sessions**, so the bot never
runs `git checkout`, `reset`, `stash` or `clean` there. It only fetches into
`refs/pr-bot/*` there; everything else happens inside the per-PR clone.
- **The clones are `git clone --shared`.** Their objects live in the main checkout, so
the `refs/pr-bot/<n>` ref there is what keeps a PR's commits safe from `git gc`; it
is deleted together with the clone when the PR closes.
- **`node_modules` may be a symlink into the live checkout.** The brief forbids
installs, and `worktree.ts` unlinks the symlink before any `npm ci`. `src/web/public/vendor`
is copied per file, never linked, because postinstall regenerates it in place.
- **Sessions are named `prbot-<n>`** and tracked by id; the bot deletes only those, on
completion, on shutdown, and (by name) as a sweep at startup after a crash. It never
touches the maintainer's `w<n>-*` sessions.
- **Readiness and end-of-turn follow the codeman skill's rules**: composer first
(`shift+tab` in the pane), trust dialog read from the screen, `stop,blocked,exit`
signals rather than `idle`. A session that asks a question is reported as a failed
review with the pane's last lines, not left hanging.
- **Telegram input is data.** Command parsing is a fixed grammar; free text is only ever
relayed to a reviewer session as the maintainer's own follow-up, or used as a closing
comment after confirmation.
Tests: `test/pr-bot-report.test.ts` (parsing, formatting, CI classification, command
grammar, trust-dialog reader, config), `test/pr-bot-state.test.ts`, and
`test/pr-bot-commands.test.ts` (the command and confirmation flows against a stubbed
`gh` and Telegram: a GitHub write happens once, after the tap, never for a foreign chat
or a reused nonce). Type-checked by
`npm run typecheck` through `config/tsconfig.pr-bot.json`, linted and formatted with
the main sources.
+51 -2
View File
@@ -1,7 +1,7 @@
# Remote Sessions (SSH)
Codeman can run a session's agent on a **remote host over SSH** instead of the
local machine. The agent (Claude, OpenCode, Codex, Antigravity, Gemini, Pi, or a plain shell)
local machine. The agent (Claude, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or a plain shell)
runs inside a `tmux` server **on the remote host**, so it survives the SSH
connection dropping; Codeman attaches to it the same way it attaches to a local
managed session.
@@ -30,7 +30,7 @@ Types live in `src/types/session.ts`; persistence in `src/remote-hosts.ts`.
| `RemoteHost` (extends `RemoteSshOptions`) | A saved host: `id`, `label`, `host`, `username`, `port?`, `commands?` (per-mode launch command override). |
| `RemoteCase` | A working directory on a host: `name`, `type: 'remote'`, `hostId`, `remotePath`. |
| `SessionRemote` (extends `RemoteSshOptions`) | The resolved bundle stamped onto a live session: host coordinates + `remotePath` + `commands`, plus **`owned?`** and **`remoteSessionName?`** (COD-105 — see [Ownership](#ownership-launched-vs-discovered-and-attached-cod-105)). Built by `toSessionRemote(host, case)` (sets `owned: true`) for the launch path, or `toAttachedSessionRemote(host, name, path)` (sets `owned: false`) for the attach path. Both copy the advanced SSH options through so every connection is identical. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity' \| 'pi'>` — the modes that can run remotely. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity' \| 'pi' \| 'grok' \| 'deepseek' \| 'omp'>` — the modes that can run remotely. |
| `RemoteSessionInfo` (COD-105) | One discovered remote tmux session: `name` (always `codeman-*`), `attached` (a client is connected), `created` (epoch s), `windows`. Returned by `listRemoteCodemanSessions()`. |
Persistence is two flat JSON arrays in the instance data dir:
@@ -116,6 +116,11 @@ Key points:
the agent. The per-mode command comes from `remote.commands?.[mode]` or
`defaultRemoteCommandForMode(mode)` (`exec claude` / `exec opencode` /
`exec codex` / `exec gemini` / `exec agy` / `exec bash -l`).
⚠️ **claude and omp no longer take that path**: both have their own arm in
`buildRemoteLaunchCommand` so a respawn can continue the same conversation
(see [Respawn / reattach continuation](#respawn--reattach-continuation)), and
because the claude arm is an `a || b` pair under `-c`, its pane PID is the
**login shell**, not the agent.
- The **whole tmux invocation is a single shell-quoted ssh argument**, and the
pane command is independently quoted, so a `remotePath` with spaces is safe.
- Connection options come from the **same `buildSshConnectionArgs(remote)`** as
@@ -202,6 +207,50 @@ The early return is a structural guarantee that **no code path can ever issue a
remote `kill-session` for a session we don't own** — the only `kill-session` run is
on the local socket, which never reaches the remote socket.
## Respawn / reattach continuation
A dropped connection or a dead pane must reconnect to the **same conversation**,
not launch a fresh one — the whole point of a durable remote session.
- **Claude**: the launch command is idempotent — `claude --session-id <id> ||
claude --resume <id>` (see `buildRemoteLaunchCommand`'s claude branch). The
first run creates the conversation under the deterministic session id; every
later reattach/respawn re-runs the same line, `--session-id` fails
("already in use"), and the `||` fallback resumes it.
- **OMP**: `omp` has no equivalent idempotent single-line form, so
`Session._pinOmpRespawnId()` resolves and pins an explicit `--resume <id>`
before a respawn (mirroring the local/docker builders, rendered through the
same `buildSpawnCommandFromRegistry` engine — not a hand-rolled command and
not `appendResumeFlag()`, which is docker-only and cannot work here: appending
a flag after the quoted `-c 'omp'` hands the id to the login shell as `$0`
instead of to `omp`). ⚠️ **The resolver only ever reads THIS host's local
`~/.omp/agent/sessions/`**, which is meaningless for a remote session — the
conversation and its session file live on the remote host, under the remote
user's home. For a remote session, `_pinOmpRespawnId()` therefore skips local
resolution entirely and falls back to `omp`'s own ambiguous `--continue`
(`ompConfig.continueSession`), which the remote pane command already renders.
This is a known, accepted degradation versus the local/docker paths' exact
`--resume` pin — safe in practice because each remote respawn talks to
exactly one remote pane's own omp history, so "most recent" is normally
correct, but it can drift the same way `--continue` always could if two
remote sessions ever share one remote directory.
## Auto-reconnect vs. a clean agent exit
`remoteAutoReconnect` (default ON) watches for a dropped SSH connection and
reconnects with bounded backoff. It must **never** revive a session whose agent
exited cleanly (Ctrl-C, Ctrl-D, `exit`) — that tears down the durable remote
tmux session itself, and a transport-level `isPaneDead()` cannot tell that apart
from a plain network drop. `remoteTmuxSessionAlive()` (#355) resolves this by
probing the remote host directly: `tmux -L codeman-remote has-session -t
codeman-ssh-<id8>` over the same `buildSshConnectionArgs` as launch, classified
by **exit status alone** (`classifyRemoteAliveExit`: `0` = alive, ssh's `255` or
a timeout = unknown, anything else = gone) — `has-session` prints nothing on
success, so reading stdout would misclassify every live session as gone. An
unreachable host answers "unknown", which also means do not revive. The answer
is cached per session and cleared whenever the pane is next seen alive, so a
stale `true` from one transport drop can never revive the NEXT clean exit.
## API
Routes are registered in `src/web/routes/case-routes.ts`:
+6 -5
View File
@@ -312,8 +312,8 @@ TOCTOU window.
| Route | Cap | Notes |
|-------|-----|-------|
| `file-content` | 10 MB | text preview |
| `file-raw` | 50 MB | inline MIME map; **`X-Content-Type-Options: nosniff` on all responses**; streamed, `Range`-aware (206 slices come from the same validated path, and the cap is checked before the range) |
| `POST /api/download` | 50 MB | forced `attachment`; sensitive‑path blocklist |
| `file-raw` | 2 GB (`CODEMAN_MAX_DOWNLOAD_BYTES`, `0` = unlimited) | inline MIME map; **`X-Content-Type-Options: nosniff` on all responses**; streamed, `Range`-aware (206 slices come from the same validated path, and the cap is checked before the range) |
| `GET /api/download` | same cap | forced `attachment`; sensitive‑path blocklist; streamed, `Range`-aware |
### SVG / content‑type XSS
@@ -340,7 +340,7 @@ the attachment guard below.
Live external attachments (`src/attachment-registry.ts`) mint an `att_<uuid>` id
for a host file so browser requests carry the id, never an absolute path. Serving
is by id (`GET /api/sessions/:id/attachments/:attachmentId/raw`, 50 MB cap,
is by id (`GET /api/sessions/:id/attachments/:attachmentId/raw`, same download cap,
`nosniff`) and re‑resolves the symlink + re‑checks the **attachment guard**
(`src/config/attachment-guard.ts`: the shared sensitive‑path blocklist **plus**
the `/root` and `/etc` trees, extendable via `attachmentBlockedPaths` /
@@ -489,7 +489,7 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never** `--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, and five seeded files from `~/.pi/agent`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, five seeded files from `~/.pi/agent`, and three from `~/.grok`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
@@ -518,7 +518,7 @@ A saved dashboard URL renders as a tab, served through Codeman's own origin at `
- **The proxy is exempt from cookie auth and the Origin/CSRF guard, and that is deliberate.** The iframe is sandboxed without `allow-same-origin`, so it is opaque‑origin: its requests are cross‑site, meaning the `SameSite=lax` session cookie is never attached and its writes and WS upgrades arrive with `Origin: null`. The credential is instead a 192‑bit capability in the path, minted only by an authenticated `POST /api/webviews/:id/open`, held in memory (a restart invalidates every one), rolling TTL, bound to the minting user, and granting nothing but "relay bytes to this one saved URL". ⚠️ **The Host allowlist is NOT bypassed**, so DNS‑rebinding protection is unaffected. A second `Referer`‑keyed form exists for root‑absolute assets and is the only exemption decided by a request‑supplied header, so it is fenced to safe methods on non‑`/api`, non‑`/ws`, non‑`/q` paths. Edges pinned by `test/webview-auth-exemption.test.ts`.
- **Sandboxed by default; `allow-same-origin` is an explicit per‑dashboard opt‑in.** A proxied page is same‑origin with Codeman, so without the sandbox its JavaScript could read the Codeman document and call the agent‑spawning API. ⚠️ In BOTH modes the `Authorization` header and the `codeman_session` cookie are stripped before the upstream request, because a trusted (same‑origin) frame makes the browser attach Codeman's own Basic‑auth credentials to every proxied request; forwarding them would hand `CODEMAN_PASSWORD` to the dashboard.
- **Not an open relay, and not a privilege boundary.** `resolveUpstreamUrl()` refuses anything leaving the saved origin, and cross‑origin redirects are handed back unchanged rather than followed. The proxy does reach whatever the SERVER can reach, which is not an escalation for someone who already commands `--dangerously-skip-permissions` agents, but in multi‑user mode it means a non‑admin's dashboard is fetched from the server's network position. Saved URLs are validated to plain http(s) with no embedded credentials, and there is deliberately **no magic‑link path**: terminal output can never create a webview (the mistake the attachment scanner had to be walled off from).
- **Not an open relay, and not a privilege boundary.** `resolveUpstreamUrl()` refuses anything leaving the saved origin, and cross‑origin redirects are handed back unchanged rather than followed. The proxy does reach whatever the SERVER can reach, which is not an escalation for someone who already commands `--dangerously-skip-permissions` agents, but in multi‑user mode it means a non‑admin's dashboard is fetched from the server's network position. Saved URLs are validated to plain http(s) with no embedded credentials, and there is deliberately **no magic‑link path**: terminal output can never create a webview (the mistake the attachment scanner had to be walled off from). The one refused destination class is link‑local and cloud‑metadata addresses (`169.254.0.0/16`, `fe80::/10`, `fd00:ec2::254`, `168.63.129.16`, `100.100.100.200`, `metadata.google.internal`): `webview-egress-policy.ts` refuses them at save time, and `webview-egress.ts` re‑judges the RESOLVED address at connect time through a `lookup` hook on the proxy's undici Agent and on its WebSocket client, so a DNS name pointing into those ranges is refused as well. Loopback and RFC1918 stay allowed on purpose. Capabilities are revoked on logout, admin logout and user deletion, and proxied responses carry `Referrer-Policy: same-origin` so a dashboard cannot hand the capability‑bearing URL to a third‑party host it links.
---
@@ -529,6 +529,7 @@ A saved dashboard URL renders as a tab, served through Codeman's own origin at `
| `CODEMAN_PASSWORD` (+ `CODEMAN_USERNAME`) | Enable HTTP Basic auth |
| `--host` / `CODEMAN_HOST` | Bind host (default `127.0.0.1`) |
| `CODEMAN_ALLOWED_HOSTS` | Extra `Host`/`Origin` allowlist entries for reverse proxies (comma‑separated; exact host, or leading‑dot `.suffix` for subdomains) — see §3 |
| `--base-url` / `CODEMAN_BASE_URL` | Sub‑path prefix Codeman is mounted under behind a reverse proxy, e.g. `/codeman` (default `/`); the proxy must forward the prefix unchanged. Independent of `CODEMAN_ALLOWED_HOSTS` |
| `--allow-unauthenticated-network` / `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK` | Acknowledge an unauthenticated non‑loopback bind (downgrades the warning) |
| `--https` | Enable TLS (adds HSTS) |
| `CODEMAN_INSTANCE` | Scope tmux socket + data dir for isolation |
+219
View File
@@ -0,0 +1,219 @@
# Codeman TUI Rework Plan
Status: **phases 0-2 implemented** on `feat/tui`; phases 3-4 remain follow-ups. The user guide is [`docs/tui.md`](tui.md); this document stays the design record.
- Phase 0: `src/cli-style.ts` (palette, glyphs, `heading`/`kv`/`table`/`spinner`/`confirm`) plus the mechanical fixes of §5, and `test/cli-commands.test.ts` now derives its inventory from the real commander `program` instead of parsing a fixture.
- Phases 1-2: `src/tui/`. `tui-app.ts` (main loop, attach handoff, verbs) and `tui-client.ts` (API, SSE, degraded enumeration) are the only IO; `tui-model`, `tui-layout`, `tui-render`, `tui-keys`, `tui-ansi`, `tui-composer`, `tui-approvals`, `tui-digest`, `tui-sse` and `tui-types` are pure and unit-tested, with an E2E suite driving the real binary under node-pty.
- Deferred with the rest of phase 3: `r` (resume a RECENT row) is not wired up, so the help overlay does not advertise it.
- Not started: phase 3 (mouse, `--pick` popup switcher, opt-in attach status line, OSC 9) and phase 4 (retiring the bash choosers).
The goal: replace Codeman's scattered terminal surfaces with one first-class TUI, `codeman tui`, that gives SSH/terminal users the same at-a-glance awareness the web UI gives browsers. The reference point is herdr (herdr.dev), the trending Rust "agent multiplexer" whose defining feature is a live agent-state sidebar. Codeman can match and beat that sidebar in the terminal because the states herdr infers from screen-scraping heuristics are states our server already computes from hooks, pane probing, and the approvals inbox.
---
## 1. What we have today (inventory)
Three disconnected surfaces, three visual idioms, two data sources:
| Surface | What it is | Data source | Idiom |
| --- | --- | --- | --- |
| `codeman` CLI (`src/cli.ts`, 1214 lines) | commander + chalk, ~20 commands | HTTP API + state files | `✓`/`✗` line-per-fact, no interactivity |
| `sc` (`scripts/tmux-chooser.sh`, 663 lines) | bash number-menu chooser, mobile-tuned (44 cols) | `tmux -L codeman` + `state.json` via jq | 256-color, numbered, full repaint per key |
| `scripts/tmux-manager.sh` (529 lines) | bash cursor TUI with kill/info | `mux-sessions.json` (and writes it back) | 8-color, box-drawn, arrow keys |
Weaknesses found in the audit (file:line refs verified 2026-08-16):
1. **No interactive picker in the Node CLI at all.** Every `session stop`, `task status`, `session logs` requires a pasted UUID prefix. There is no `codeman attach <session>`; `codeman attach` is actually the attachment-card command (and `README.md:895` describes it wrongly).
2. **`sc` cannot reach sessions 10+ interactively**: entries are numbered globally (`tmux-chooser.sh:343`) but input accepts a single `[1-9]` keypress (`:487-493`). Page 2 shows items 8-14 that mostly cannot be selected.
3. **No cursor/selection concept in `sc`** (`BG_SEL` at `:90` is dead code); arrows only page.
4. The two bash tools can disagree about which sessions exist (different data files), and only `sc` is on PATH.
5. **Zero live feedback anywhere**: `codeman web -d` and `service install` block silently up to 30s (`daemon-control.ts:395-412`); no spinner exists in the codebase.
6. Styling drift: `doctor` is the only table and is deliberately monochrome with a colorize hook nobody wired up (`dependency-report.ts:5-7`); `codeman web` prints its "running at" line twice (colored `cli.ts:934`, plain `server.ts:2366`); the server's security warning is colorless `console.warn` while the CLI's version of the same warning is yellow; `tmux-manager.sh`'s header box is visibly misaligned; `padEnd(14)` overflows on "Antigravity CLI".
7. Bash TUIs emit raw escapes unconditionally (no TTY/NO_COLOR gate); `install.sh` and `postinstall.js` do it right.
8. Detach hint inconsistency: chooser says Ctrl+B D, `README.md:671` says Ctrl+A D.
9. Inside an attached session there is **no chrome at all**: Codeman turns the tmux status bar off (`tmux-manager.ts:1978`), so an SSH user in a pane has no session identity, no state, no way back to a picker except detach.
10. `test/cli-commands.test.ts` asserts against a hand-written fixture, not the real `program`, and that fixture already lists a `tui` command that does not exist (`:57-61`). The name is pre-approved by our own test file.
## 2. Research: how herdr does it
herdr (github.com/herdrdev/herdr, ~30k stars, single Rust binary, pre-1.0) is a background terminal multiplexer "your coding agents live on". What matters for us:
- **The agent-state sidebar is the product.** Every pane is classified live as `working` / `blocked` / `done` / `idle` and grouped in a sidebar, so you see who needs you without switching tabs. Reviews unanimously call this "the killer feature tmux can't match".
- **Detection is heuristic-first**: process-name matching + screen-manifest TOML rules parsing the visible frame; optional per-agent "integration install" adds lifecycle hooks over JSON-RPC on a unix socket for accurate states. Claude Code there is on the heuristic path and reviewers note blocked-state lag.
- **Model**: workspaces → tabs → panes, tmux-style prefix keys (Ctrl+B V split, arrows navigate, D detach), mouse-first (click select, drag resize, right-click menus, touch over SSH), adapts to narrow widths.
- **Agent-shaped API**: socket API with `pane read` (visible/recent/detection), `send-text`/`send-keys`/`run`, `agent start|prompt|wait|explain`, `pane wait-output` with regex, plugins placed as overlay/split/tab/popup.
- **Persistence**: sessions survive disconnects, reattach from any terminal / SSH.
- Weaknesses reviewers cite: pre-1.0 churn, bus factor 1, no session resurrection, rendering lag with many panes.
What is striking is how much of herdr Codeman already has, server-side: our hooks give exact `permission_prompt`/`stop`/`idle_prompt` events (herdr's "integration" path, but installed by default), `_confirmIdle()` does the screen-probe fallback, the approvals inbox parses the actual dialog options, and the agent skill + wait primitives are our socket API. What we lack is purely the presentation layer in the terminal.
Prior art for the architecture we want: **agent-deck** (Bubble Tea + tmux) proves the "TUI list + attach into tmux" model works great: session list with live glyphs (● ◐ ○ ✕), Enter attaches into a tmux pane, status polling, groups, fuzzy search. We take the shape, not the code.
Licensing note: herdr is reported variously as Apache-2.0/AGPL-3.0. Irrelevant either way: we copy concepts, never code.
### What we take / what we skip
Take: the four-state sidebar as the organizing principle; grouping by "needs you first"; narrow-width adaptation; mouse support; tmux-familiar keys; the "attention at a glance" framing.
Skip: being a multiplexer. tmux already backs every Codeman session and is a hard dependency; herdr had to build pane management because it owns terminals, we do not. Also skip (for now): plugin marketplace, split layouts, pane drag. Our TUI is a **dashboard + switchboard over tmux**, not a tmux replacement.
## 3. Design: `codeman tui`
One command, one full-screen client of the existing HTTP/SSE API.
**Positioning (owner decision, 2026-08-16): the web UI remains THE primary surface.** The TUI is strictly additive, for users who want a terminal workflow (SSH, Termius, tmux die-hards). Bare `codeman` keeps printing help; nothing existing changes behavior. The `sc` bash chooser also stays untouched for now; flipping its alias to `codeman tui` is deferred to a follow-up release once the TUI has mileage.
### Layout (≥100 cols)
```
codeman tnode · v1.19.0 · 6 sessions · 5h ▂▂▅ 32% wk 61% ? help q quit
────────────────────────────────────────────────────────────────────────────────────────────
NEEDS YOU ──────────────────────────┐ ┌ w4-api-refactor ── claude · ~/dev/api ────────────
▶ 1 w4-api-refactor ⚠ approval 2m │ │ ✻ Actualizing… (2m 14s · ↓ 12.3k tokens)
2 w6-docs ✋ waiting 11m │ │
│ │ ⚠ Claude requests: Bash(git push origin main)
WORKING ────────────────────────────┤ │ 1. Yes 2. Yes, don't ask again 3. No
3 w1-codeman ✻ 17m 45.2k │ │
4 w2-gallery ✻ 3m 8.1k │ │ [y] approve [n] deny [Enter] attach
IDLE ───────────────────────────────┤ │
5 w3-promo ○ 2h │ │ …live tail of the selected session's
RECENT ─────────────────────────────┤ │ terminal (ANSI colors preserved),
· api-hotfix ✔ done Fri │ │ updating while you browse the list…
────────────────────────────────────────────────────────────────────────────────────────────
↑↓ select · ⏎ attach · 1-9 jump · y/n answer · p prompt · n new · x kill · / search · g digest
```
- **Header**: hostname/instance, server version, session count, plan-usage chip (same telemetry that feeds the web chip, when available). Degrades gracefully when the server is down (see §3.6).
- **Sidebar**: sessions grouped `NEEDS YOU` → `WORKING` → `IDLE` → `RECENT` (past sessions from the unified list, resumable). Within groups, reuse the activity ordering already built for the home screens in PR #303 (blocked first, running longest, quiet newest); that logic is pure and shared.
- **Preview pane**: live tail of the selected session, SGR colors preserved, cursor-movement stripped. When the selected session has a pending approval, the parsed dialog is rendered as a card above the tail with one-key answer bindings.
- **Footer**: contextual keymap (changes when a dialog/confirm is active).
### States and vocabulary
Exactly the web's language so the two surfaces read the same:
| Group | Glyph | Color | Source |
| --- | --- | --- | --- |
| NEEDS YOU (question/permission) | `⚠` | red, blinking row | approvals inbox / `permission_prompt` |
| NEEDS YOU (waiting for input) | `✋` | yellow | `idle_prompt` / waiting classification |
| WORKING | `✻` animating through `· ✢ ✳ ∗ ✻ ✽` at 2Hz | green | working classification (the same glyph family Claude itself draws, a deliberate nod) |
| IDLE | `○` | muted | idle |
| RECENT / done | `✔` | muted green | unified list history rows |
Nerd-font/glyph fallback exactly like `sc` does today (`[!] [w] [*] [-] [ok]` when the terminal is not known-capable), plus full NO_COLOR / `tput colors` degradation (8-color and mono renderings are designed, not accidental).
### Keymap
- `↑/↓` or `j/k` select · `Enter` attach · `1-9` jump-attach (parity with `sc`, but now the cursor covers 10+)
- `y`/`n` (or the digit keys) answer the selected session's pending approval right from the dashboard, via `POST /api/approvals/:id/answer`. The server already re-captures the pane and 409s if the dialog is gone, so this is safe by construction.
- `p` send a one-line prompt to the selected session without attaching (`POST /input` with `\r`, the composer opens in the footer)
- `n` new session (case picker → mode picker, drives `POST /api/quick-start`) · `x` kill with typed confirm (never bulk; refuses the session hosting the TUI itself, like tmux-manager.sh does)
- `/` fuzzy search across sessions/history/attachments (`GET /api/search`) · `g` away digest (`GET /api/away-digest`) rendered as a panel
- `r` resume selected RECENT row (unified list `resume-session` flow) · `?` help overlay · `q` quit
- Mouse (phase 3): SGR mouse reporting, click selects, wheel scrolls list/preview, click on footer keys triggers them. Works over SSH, same as herdr's touch story.
### Responsive behavior
The `sc` design constraint survives: below ~72 cols (Termius, iPhone portrait) the preview pane drops and the TUI is a single-column list with two-line rows, nearly identical to today's `sc` but with a cursor, live states, and the answer/prompt/new/kill verbs. The layout switch is width-driven at draw time, no mode flag.
### Attach model
Enter suspends the TUI (restore main screen + cooked mode), then hands the terminal to `tmux -L <socket> attach-session -t <name>` with `stdio: inherit`. On tmux exit/detach, the TUI resumes and refreshes. Full fidelity (mouse, paste, colors) is tmux's, we never proxy bytes.
- Inside tmux already: same socket → `switch-client -t`; different socket → warn about nesting and offer detach-first. `$TMUX` + `CODEMAN_MUX` detection.
- **Return path**: a tmux binding installed for codeman sessions (opt-in) runs `codeman tui --pick` inside `tmux display-popup -E`, a minimal picker-only mode (list + jump, no preview) so switching sessions from inside a pane is one keystroke, fzf-style.
- Optional per-attach chrome (opt-in setting, default off since `status off` at `tmux-manager.ts:1978` is deliberate): a minimal codeman-styled tmux status line showing `name · state · alert`, set on attach, restored on detach.
### Notifications
While the TUI is open and a session flips to NEEDS YOU: flash the row, ring BEL, and optionally emit OSC 9 (desktop notification in kitty/WezTerm/iTerm2, and it traverses SSH). This is the herdr sidebar promise delivered even when the terminal is backgrounded.
### Degraded mode (server down)
`sc` works without the server today and the TUI must too: when no server answers, enumerate `tmux -L codeman list-sessions` + read `state.json` (read-only), show a "server not running" header line, and offer attach only (no states, no approvals). This keeps the "web server crashed, get me to my sessions" path alive.
## 4. Architecture
### A client of the server, not a second brain
Everything live comes from the API the web UI already uses:
| Need | Endpoint |
| --- | --- |
| Session list + history | `GET /api/sessions/unified` |
| Live updates | SSE `GET /api/events` (heartbeat `sse:heartbeat` already exists; fall back to 2s polling) |
| Pending approvals + parsed options | `GET /api/approvals`, answer via `POST /api/approvals/:id/answer` |
| Preview tail | `GET /api/sessions/:id/terminal?tail=N` (throttled to the selected session only) |
| Prompt send | `POST /api/sessions/:id/input` (single line + `\r`, per the composer contract) |
| New session | `POST /api/quick-start` (routes remote/docker cases correctly) |
| Search | `GET /api/search` |
| Away digest | `GET /api/away-digest` |
| Plan usage chip | latest status-telemetry snapshot (`plan-usage-latest`) |
Server discovery and auth reuse what exists: instance config from `src/config/instance.ts` (`CODEMAN_INSTANCE`, `CODEMAN_PORT`), the probe logic from `daemon-control.ts`, credentials from `~/.codeman/.env` (the established `codeman attach` pattern), self-signed HTTPS accepted for loopback probes (the hooks-on-HTTPS lesson). Multi-user scoping comes free: the API only returns what the authenticated user owns.
### Renderer: hand-rolled, zero new dependencies (decision)
Options considered:
- **Ink (React for CLIs)**: what Claude Code uses. Pros: layout engine, ecosystem. Cons: pulls React into a CLI that today ships only commander+chalk; rerender model fights the two things we care most about (a raw-ANSI preview region and 2Hz glyph animation without flicker); version-pins React for every `npm i -g aicodeman`.
- **blessed/neo-blessed**: unmaintained, skip.
- **Hand-rolled screen core** (recommended): this repo hand-rolls ANSI everywhere already and has the expertise (regex-patterns, stripAnsi, the xterm work). The core is small and boring: alt screen + raw mode + cursor-home full-frame repaint from an off-screen string buffer, throttled to state changes and the 2Hz animation tick, wrapped in DECSET 2026 (synchronized output) where supported so repaints are atomic in modern terminals (tmux, kitty, WezTerm, iTerm2). No diffing needed at these frame rates.
The one genuinely tricky pure function: SGR-aware line clipping for the preview (keep colors, strip cursor movement/OSC/DECSET, clip to width while carrying SGR state, reset at EOL). That is a pure module with exhaustive unit tests, and it is exactly the kind of function Ink would not have given us anyway.
### Module layout
```
src/tui/
tui-app.ts entry + main loop + attach handoff (IO)
tui-client.ts API + SSE client, degraded-mode enumeration (IO)
tui-model.ts pure: state store, grouping, ordering (reuses PR #303 helpers)
tui-layout.ts pure: responsive layout math, row building
tui-render.ts pure: model+layout -> frame string (palette, glyphs, fallbacks)
tui-keys.ts pure: byte stream -> key/mouse events (incl. SGR mouse decode)
tui-ansi.ts pure: SGR-aware clip/filter for the preview
```
Pure modules unit-test with no TTY. `cli.ts` gains one thin `tui` command registration (and `--list`/`<n>` fast paths for `sc -l` / `sc 2` parity, which must stay fast: they short-circuit before any screen setup).
## 5. CLI-wide polish (the rest of "make it much nicer")
A shared style kit, `src/cli-style.ts`: one palette (mirroring the web's status colors), one glyph set with fallback, `heading()`, `kv()`, `table()` (width-aware, fixes the Antigravity overflow), `spinner()` (finally: the 30s silent daemon/service waits get a live line), `confirm()` (used by `reset --force`'s missing prompt and `x` in the TUI). Then the mechanical fixes from §1: colorize `doctor` through the hook that already exists for it, dedupe the `codeman web` startup line, colorize the server's security warning, fix the README `codeman attach` description and the Ctrl+B/Ctrl+A detach drift, TTY/NO_COLOR gates everywhere.
## 6. Phasing
| Phase | Contents | Size |
| --- | --- | --- |
| 0 | `cli-style.ts` + mechanical fixes (§5), real CLI tests (retire the fixture parser in `test/cli-commands.test.ts`) | S |
| 1 | `codeman tui` core: list + states via SSE, cursor + 1-9, attach/return loop, kill w/ confirm, new session, narrow mode, degraded mode, `sc` alias flip + `--list`/`<n>` parity | M/L |
| 2 | Preview pane (SGR clip), approvals answering, prompt composer, search, digest, resume, plan-usage header | M |
| 3 | Mouse support, `--pick` popup switcher + tmux binding, opt-in attach status line, BEL/OSC 9 notifications | M |
| 4 | Retire `tmux-chooser.sh`/fold `tmux-manager.sh` (keep as thin wrappers for one release), docs/README/wiki, screenshots for promo | S |
Phases 0-1 are the useful minimum; 2 is where it beats herdr's sidebar (answering approvals from the dashboard); 3 is delight.
## 7. Testing
- Pure modules (`tui-model/layout/render/keys/ansi`): plain vitest, frame snapshots as stripped strings plus targeted ANSI assertions.
- Interactive E2E: spawn the built TUI under `node-pty` (already a dependency), feed keys, assert on captured frames; the vitest tmux mock (`IS_TEST_MODE`) keeps attach paths inert. Port rules per CLAUDE.md (3150+, `app.inject()` where possible by testing `tui-client` against injected routes).
- Manual: Termius/iPhone portrait (the 44-col case), tmux nesting, server-down mode, NO_COLOR, non-nerd-font terminal.
## 8. Invariants this plan respects
- tmux socket and data dir always via instance config (`dataPath()`, `-L codeman`); a beta instance TUI sees only its own world.
- Never bulk kill, always confirm, never touch another session implicitly, refuse killing the session the TUI runs in (w1/w2/w3 are sacred).
- Input is single-line with `\r`, via the server (never raw tmux send-keys from the TUI while the server owns the session).
- Approvals answering goes through the server's re-capture + 409 path, never blind keystrokes.
- `status off` on panes stays the default; any chrome is opt-in.
- No new runtime dependencies; the npm package stays light.
## 9. Decisions (resolved 2026-08-16)
1. **Bare `codeman` does NOT open the TUI** (owner decision): the web UI is the main thing, the TUI is additional. `codeman tui` only.
2. **`sc` stays the bash chooser for now**; the alias flip is a follow-up once the TUI has mileage. `codeman tui --list` / `codeman tui <n>` provide the same fast paths for people who want to switch.
3. Opt-in tmux status line: deferred to phase 3 along with the `--pick` popup switcher.
4. Preview tail goes over the API (auth/multi-user/remote-consistent); previews are simply unavailable in degraded server-down mode.
5. Name is `codeman tui` (the test fixture historically expected it).
Initial PR scope: phases 0-2. Phase 3 (mouse, popup switcher, status line, OSC 9) and phase 4 (bash chooser retirement) are follow-ups.
+278
View File
@@ -0,0 +1,278 @@
# Terminal UI (`codeman tui`)
`codeman tui` is a full-screen dashboard for your Codeman sessions, in the terminal.
It shows every session grouped by whether it needs you, lets you answer a permission
dialog or send a prompt without switching anywhere, and puts you inside a session's
tmux pane with one keystroke.
It is **additional, not a replacement**: the web UI stays the primary surface and
gets every feature first. The TUI exists for the terminal workflow (SSH, Termius,
a tmux window you keep open all day), and it is a *client* of the running server,
so the two surfaces can never disagree about what a session is doing. It is also
not a multiplexer: tmux still owns every pane, and attaching hands the terminal to
tmux rather than proxying bytes.
## Starting it
```bash
codeman tui # the dashboard
codeman tui --list # print the numbered session list and exit
codeman tui 2 # attach straight to session 2 of that list
```
The two fast paths are the scriptable ones.
Neither sets up a screen, so both are as quick as the one API call they make, and
`--list` prints plain text when piped, so it composes with `grep`/`awk`.
What it needs:
| Needs | What you get |
| --- | --- |
| **Full features** | A running Codeman server (states, approvals, preview, prompts, search, digest). The TUI finds it the way `codeman attach` does: `CODEMAN_API_URL`, else loopback on `CODEMAN_PORT` for this `CODEMAN_INSTANCE`. The self-signed certificate an `--https` install generates is accepted, as it is everywhere else in the CLI. |
| **Server down** | It still starts, in **degraded mode**: sessions are enumerated straight from `tmux -L codeman` plus a read-only peek at `state.json`, and attach is the only verb. See [Troubleshooting](#troubleshooting). |
| **A terminal** | `codeman tui` refuses to run when stdin/stdout are not a TTY, and says to use `--list` instead. A cron job or a pipe therefore fails loudly rather than emitting escape codes into a log. |
## What it looks like
A real frame at 100x30 (`NO_COLOR`, trailing blank rows trimmed). The selected
session has a pending permission dialog, so the preview pane leads with the card:
```
codeman ⚠ 2 tnode · v1.19.0 · 5 sessions · 5h 32% · wk 61% ? help q quit
NEEDS YOU ─────────────────────────│ w4-api-refactor · claude · /home/you/dev/api · blocked
1 w6-docs ✋ 11m│ ⚠ requests: Bash(git push origin main)
▶ 2 w4-api-refactor ⚠ 2m│ 1. Yes
WORKING ───────────────────────────│ 2. Yes, and do not ask again
3 w1-codeman ∗ 1h│ 3. No, tell Claude what to do
4 w2-gallery ∗ 15m│ y approve · n deny · digit chooses
IDLE ──────────────────────────────│
5 w3-promo shell ○ 2h│ > refactor the api routes onto the shared port interface
RECENT ────────────────────────────│
6 api-hotfix ✔ 3d│ Read src/web/ports/session-port.ts (48 lines)
│ Read src/api/routes.ts (312 lines)
│ Edit src/api/routes.ts
│ 1 -import { SessionManager } from "../session-manager.js";
│ 2 +import type { SessionPort } from "../web/ports/session-
│
│ Bash(npm run typecheck)
│ └ tsc --noEmit: no errors
│
│ ✻ Actualizing… (2m 14s · ↓ 12.3k tokens)
↑↓ select · ⏎ attach · y approve · n deny · 1-9 option · p prompt · x kill · / search · g digest ·
```
- **Header**: the machine, the server version, how many sessions are live, and the
plan-usage chip (the same statusLine telemetry that feeds the web chip, when the
server has a snapshot). A `⚠ n` badge counts pending approvals.
- **Sidebar**: every session, grouped and numbered.
- **Preview**: a live tail of the selected session, its own colors preserved, with
the parsed dialog card on top when that session is blocked.
- **Footer**: only the keys that work right now. `n` reads `n new` normally and
`n deny` when the selected session has a dialog, because it cannot be both.
The same world through `--list`:
```
1 waiting w6-docs /home/you/dev/docs
2 blocked w4-api-refactor /home/you/dev/api
3 working w1-codeman /home/you/dev/codeman
4 working w2-gallery /home/you/dev/gallery
5 idle w3-promo /home/you/dev/promo
6 done api-hotfix /home/you/dev/api
```
The numbers are the same on both surfaces, so `codeman tui --list` then
`codeman tui 4` is one thought.
## The four groups
Groups are always in this order, and a session is in exactly one of them:
| Group | Glyph | Means | Comes from |
| --- | --- | --- | --- |
| **NEEDS YOU** | `⚠` | A permission or question dialog is blocking the agent | The approvals inbox (`permission_prompt` hooks, with the on-screen options parsed) |
| | `✋` | Waiting for your next instruction, or errored | `idle_prompt`, or an errored session (equally something only a human clears) |
| **WORKING** | `✻` animating | A turn is running | The same working classification the web dashboard uses |
| **IDLE** | `○` | Live, but sitting there | |
| **RECENT** | `✔` | A past session from the unified list | History rows, no live pane |
Ordering inside a group is "the one that has waited longest, first": blocked
sessions sort by how long the dialog has been up, working sessions by when their
turn started (the pane's last Enter, since a working pane repaints every second
and would otherwise always look freshly started), and quiet ones by last activity.
That is the ordering the web home screens already use.
The cursor sticks to a **session**, not a row number, so a session that jumps to
NEEDS YOU does not drag your selection with it. The number beside each row is what
`1-9` and `codeman tui <n>` mean, and it is renumbered on every re-sort.
When a new dialog appears, the terminal bell rings once, for that dialog only: the
same item announced twice does not ring twice.
## Keymap
| Key | Does |
| --- | --- |
| `↑` `↓` or `j` `k` | Move the cursor. PageUp/PageDown jump five rows. |
| `Enter` | Attach to the selected session (see [Attaching](#attaching)) |
| `1`-`9` | Jump to that row and attach. When a dialog is on screen, a digit answers it instead (see below). |
| `y` | Approve the selected session's dialog |
| `n` | Deny it, or **start a new session** when there is no dialog |
| `p` | Send one line to the selected session without attaching |
| `x` | Kill the selected session; `y` confirms, any other key cancels |
| `/` | Search sessions, events and files |
| `g` | Away digest: what happened while you were gone |
| `?` | Help overlay |
| `Esc` | Close whatever overlay is open |
| `q` or `Ctrl+C` | Quit, restoring the screen you started with |
Inside the `p` composer and the `/` query: `←` `→` `Home` `End` `Delete`
`Backspace` plus `Ctrl+A` / `Ctrl+E` / `Ctrl+U` / `Ctrl+W`, `Enter` to send or open,
`Esc` (or `Ctrl+C`) to cancel. In the kill confirmation you retype the session name;
anything else cancels. In the `n` pickers, type to filter, `Enter` chooses.
Verbs that need the server (`y`/`n`/`p`/`x`/`/`/`g`) say so in degraded mode
instead of failing silently; `Enter` and `1-9` keep working.
### `p` sends exactly one line
The composer is a single line by design, ending in a carriage return: that is the
input contract every Codeman path follows, because multi-line text breaks the
agent's own composer. Pasted newlines become spaces rather than being rejected, so
a paste cannot silently run a different command than the one you read.
## Answering approvals
This is the thing the terminal could not do before. Select a blocked session and:
- `y` approves.
- `n` picks the parsed "No" option, or sends Esc when the dialog did not parse one.
- A digit picks that numbered option, **but only a digit the dialog actually
offers**. A digit with no matching option falls through to the list's own
jump-and-attach binding, so it can never be typed at whatever has focus.
The answer goes through `POST /api/approvals/:id/answer`, which **re-captures the
pane before it types anything**. If the dialog is no longer on screen (you answered
it in tmux a moment ago, or the agent moved on), the server refuses with a 409 and
the TUI says `that dialog is no longer on screen` rather than pressing a key into a
live composer. The answer is scoped to the options the server parsed off the actual
frame, never to a guess.
An idle prompt (`✋`) is not a dialog: there is nothing to approve, so `p` is the
reply path and the footer says `p reply` instead of `p prompt`.
## Attaching
`Enter` suspends the dashboard (main screen back, cooked mode back) and hands the
terminal to tmux with `stdio: inherit`. Colors, mouse and paste are tmux's, at full
fidelity.
**Press `F1` to come back.** One key, no modifier to hold or release, nothing to
type in a particular order. tmux's own way out is a chord — press the prefix, let
go, then a letter — and beta testing showed that is genuinely hard to convey: the
bar first named the wrong letter (tmux binds lowercase `d` to `detach-client` and
capital `D` to `choose-client`), and once corrected it still failed for anyone who
kept Ctrl held, because that sends `Ctrl+D`, which tmux leaves unbound. So the TUI
claims `F1` in tmux's prefix-less key table for the length of the attach and gives
it back afterwards. The chord still works; it is simply not what you are told to
press.
You do not have to remember any of it. For as long as the attach lasts the pane
wears a bar across the top:
```
1 w3-codeman-… 2 w4-codeman-… 3 testcase … alt+1-9 switch · F1 back to the codeman dashboard
```
That is the **session strip**: the other sessions stay visible from inside a pane,
numbered exactly as the dashboard numbers them, with the one you are in inverted.
`Alt+1`..`Alt+9` switch between them without going back to the dashboard first. With
more sessions than fit, the strip shows a window around the current one and marks
each cut end with `…`; the way-out hint is measured first and always keeps its space.
Codeman keeps the status bar off on its panes (the web UI carries that information
around the terminal instead), so the TUI turns it on for the attach and puts it back
exactly as it was on detach, along with each window's size. Every session the strip
can switch to is dressed and sized the same way, so switching is instant and lands
in a pane that already fills your terminal.
Detaching leaves the agent running; typing `exit` or pressing `Ctrl+D` would end it,
which is the difference the bar exists to make obvious. If an agent does exit, its
pane stays as a corpse: the TUI refuses to attach to a dead pane and offers `r` to
resume the conversation in a fresh one instead.
Three cases:
| Where you are | What happens |
| --- | --- |
| Not in tmux | `tmux -L codeman attach-session` |
| Already in tmux on Codeman's socket | `switch-client`, so you do not nest |
| In tmux on a **different** socket | Refused, with an explanation: detach from that tmux first, then run `codeman tui` again |
A direct-PTY session has no pane to attach to, and says so.
**`Enter` on a RECENT row resumes that conversation** instead: there is no pane to
attach to, so the TUI creates a new claude session carrying the old transcript
(`resumeSessionId`, exactly what the web UI's "Resume Conversation" list does), in
the directory it originally ran in and under its old name, then attaches to it. It
is claude-only, and a row with no working directory or no conversation id says why
rather than resuming something else.
`x` never bulk-kills: it kills one session, only after you retype its name, never a
history row, and never the session the TUI itself is running in.
## Over SSH, and on a phone
The TUI is an ordinary terminal program with no local dependencies beyond tmux, so
`ssh box` then `codeman tui` works exactly like running it locally. There is no
separate remote mode.
Below 72 columns (Termius, an iPhone in portrait) the preview pane is dropped and
rows take two lines each, keeping the cursor, the live states and the
answer/prompt/kill verbs. The switch is
width-driven at draw time, so unfolding a foldable or resizing a window re-lays out
immediately; there is no mode flag to set.
## Troubleshooting
**"The Codeman server rejected these credentials."** The server has
`CODEMAN_PASSWORD` set. Export `CODEMAN_PASSWORD` (and `CODEMAN_USERNAME` if it is
not `admin`), or put them in the data dir's `.env` (`~/.codeman/.env`), which is
where `codeman attach` already reads them from.
**`server not running: attach only`** in a yellow banner. Nothing answered on the
expected port, so the TUI fell back to enumerating tmux. You get names and attach;
you do not get states, approvals or previews, because those only exist on the
server. Start the server (`codeman web -d`, or `systemctl --user start codeman-web`)
and the banner clears on its own: the TUI keeps re-probing.
**It found the wrong server, or none.** Discovery is instance-scoped. A beta
instance (`CODEMAN_INSTANCE=beta`) has its own data dir *and* its own tmux socket,
so its TUI sees only its own sessions. Set `CODEMAN_PORT` or `CODEMAN_API_URL`
explicitly when you run more than one.
**"this terminal is already inside tmux on socket ..."** You are in a tmux session
on a socket that is not Codeman's, so attaching would nest two multiplexers whose
prefix keys collide. Detach from that tmux and run `codeman tui` from outside.
**Boxes and glyphs render as garbage.** The TUI picks a glyph tier from the
environment: no `TERM` (or `dumb`), or a non-UTF-8 locale, gets the ASCII set
(`[!] [w] [*] [-]`, `+`/`-`/`|` frames). Force it either way with
`CODEMAN_TUI_GLYPHS=ascii|unicode|nerd`.
**Colors.** Standard `NO_COLOR` / `FORCE_COLOR` handling (chalk's, the same as the
rest of the CLI). Under `NO_COLOR` the frame is cursor addressing and text only,
and the preview's own colors are stripped too, so a session's output cannot repaint
the dashboard.
**It refuses to open at all**, saying it needs an interactive terminal. stdout or
stdin is not a TTY. That is the guard: use `codeman tui --list`.
## Related
- [`docs/tui-plan.md`](tui-plan.md): the design record. Why hand-rolled ANSI, why a
client and not a second brain, and what is deliberately deferred.
- [`docs/approvals-inbox-plan.md`](approvals-inbox-plan.md): where the parsed
dialogs and the answer endpoint come from.
- [`docs/remote-sessions.md`](remote-sessions.md): remote-SSH cases, which the TUI
lists like any other session.
+46 -4
View File
@@ -52,6 +52,37 @@ sandbox, cookies, CORS, CSP, or any reverse proxy sitting in front of Codeman, s
passing Test does not guarantee the embedded page will render (see the
cookie-authenticated reverse proxy caveat below).
## Links to `localhost` from another device
An agent prints `http://localhost:5173/` (a dev server, a preview, a report it just
served) and you tap it on your phone. That address only exists on the Codeman box, so
the phone's browser can never load it — but the web-tab proxy fetches from the server,
where it works.
So a **loopback** link (`localhost`, `127.0.0.0/8`, `0.0.0.0`, `::1`) clicked
in the terminal or in the Response Viewer opens as a **proxied web tab** whenever the
Codeman page itself is not on that box. A saved proxied dashboard on the same origin is
reused (one tab per dev server, with the link's own path opened inside it, and one tab
per dev server rather than per host spelling, so `localhost:5173` and `127.0.0.1:5173`
share it); otherwise one is saved under its `host:port` so it is in the Run dropdown
next time, and a toast tells you it was saved. Sandboxed by default, like any other web
tab.
⚠️ **`*.localhost` is deliberately not auto-routed**, even though a browser treats it as
loopback. Every other name in that list is an address literal that can only mean this
box; a `*.localhost` DNS name is not one, and on a resolver with a search domain
configured `evil.localhost` can be retried as `evil.localhost.<search domain>`, which
someone else can control. Since the links come from agent output, one tap would then
make Codeman fetch an agent-chosen origin server-side and save it. If you really run
`api.localhost` dev hosts, add that dashboard by hand: doing so is an explicit action,
which is the difference that matters here. A **trusted** (non-sandboxed) dashboard is
likewise never auto-reused by a tapped link, for the same reason.
Only loopback is routed this way. A LAN or tailnet address (`192.168.…`, `100.…`,
`box.ts.net`) may well be reachable from the device — a VPN, the same Wi-Fi — and a
direct open is the cheaper, richer path, so those links still open in a new browser tab.
On the box itself (a browser on `localhost`) every link opens directly.
## The sandbox, and when to turn it off
Because a proxied dashboard is served from Codeman's own address, it is
@@ -161,10 +192,21 @@ then every API call fails, which looks like the dashboard being broken.
then streams the body without any time bound; a header timeout is logged
server-side and answered as a 502 that names the limit. WebSocket handshakes use
the separate `CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS` (default 30s).
- **Not a security boundary.** The proxy reaches whatever the Codeman server can
reach. That is not an escalation for someone who already commands
`--dangerously-skip-permissions` agents, but in multi-user mode it does mean a
non-admin user's dashboard is fetched from the server's network position.
- **Not a security boundary, with one carve-out.** The proxy reaches whatever the
Codeman server can reach (a `localhost` dashboard is the point), so it is not an
escalation for someone who already commands `--dangerously-skip-permissions`
agents, but in multi-user mode it does mean a non-admin user's dashboard is
fetched from the server's network position. The carve-out: link-local and
cloud-metadata addresses (`169.254.0.0/16`, `fe80::/10`, `fd00:ec2::254`,
Azure's `168.63.129.16`, Alibaba's `100.100.100.200`, the
`metadata.google.internal` alias) are refused at save time AND at connect
time, judged on the address a name actually resolves to. Nothing anyone embeds
as a dashboard lives there; an instance's IAM credentials do.
- **The proxy URL is a bearer credential.** `/webview/<cap>/...` needs no cookie,
so treat it like a password. It is revoked when you log out, when an admin logs
you out, and when your account is deleted, and it expires after 12 hours
without use. Proxied responses carry `Referrer-Policy: same-origin`, so a
dashboard that links to third-party sites does not hand them the URL.
## Where the code lives
+1 -1
View File
@@ -88,7 +88,7 @@ would. That indirection buys:
- **Real scrollback.** History is held by tmux, so reconnecting replays what happened while
you were gone instead of starting from blank.
- **Attach from anywhere else.** The same session is reachable from a terminal over SSH
with the `sc` chooser, or plain `tmux -L codeman attach`.
with `codeman tui`, or plain `tmux -L codeman attach`.
- **Secrets off the command line.** Environment overrides are injected with socket-scoped
`tmux setenv` rather than being visible in the spawn command.
+4 -2
View File
@@ -73,8 +73,10 @@ features are Claude-only; [Agent CLIs](Agent-CLIs) lists exactly which.
### Can I attach to a session from a terminal instead of the browser?
Yes. `sc` is an interactive chooser (`sc 2` attaches directly, `sc -l` lists), or use tmux
directly on the `codeman` socket. Detach with `Ctrl+A D`.
Yes. `codeman tui` is a full-screen dashboard of your sessions, with the same
NEEDS YOU / WORKING / IDLE grouping the web UI uses. `codeman tui --list` prints the
numbered list and exits, and `codeman tui 2` attaches straight to session 2. `Enter`
attaches, `F1` comes back. You can also use tmux directly on the `codeman` socket.
## Running unattended
+1 -1
View File
@@ -119,7 +119,7 @@ Shell and external CLI sessions accept `idle`, `working`, and `exit`.
## SSE
`GET /api/events` is the live event stream. 155 event names, kept in sync between server and
`GET /api/events` is the live event stream. 156 event names, kept in sync between server and
client with a test that fails on drift.
The heartbeat is a **named** `sse:heartbeat` event rather than an SSE comment, because
+1
View File
@@ -14,6 +14,7 @@ works, slash commands included.
| `Shift+Enter` / `Ctrl+Enter` | Newline without sending. |
| `Ctrl+C` | Copy if text is selected, otherwise interrupt. |
| `Ctrl+Shift+C` | Copy, never interrupts. |
| `Ctrl+V` | Paste. A clipboard image uploads instead. |
| `Ctrl+L` | Clear the terminal. |
### Exactly-once delivery
+1
View File
@@ -25,6 +25,7 @@ Press `Ctrl+?` in the app for the same list in a floating overlay.
| `Ctrl+Enter` | Same. |
| `Ctrl+C` | Copy the selection, or interrupt when nothing is selected. |
| `Ctrl+Shift+C` | Copy the selection. Never interrupts. |
| `Ctrl+V` | Paste. An image on the clipboard uploads and pastes its file path instead. |
| `Ctrl+L` | Clear the terminal. |
| `Ctrl+Shift+R` | Restore terminal size. |
| `Ctrl` `+` / `Ctrl` `-` | Font size. |
+51 -6
View File
@@ -167,6 +167,47 @@ and is not one.
Also make sure the proxy forwards WebSocket upgrades. The terminal is a WebSocket, and the
upgrade runs the same Host and Origin checks, closing with code `4003` on failure.
### Mounting under a sub-path
By default Codeman assumes it is served at the origin root (`/`). To mount it under a
sub-path — e.g. `https://example.com/codeman/` — start it with `--base-url` (or the
`CODEMAN_BASE_URL` env var):
```bash
codeman web --base-url /codeman
# or
CODEMAN_BASE_URL=/codeman codeman web
```
The value is a plain path prefix; `/` (the default) means "mounted at the root". With a
prefix set, Codeman emits every URL — the HTML shell and its assets, API/SSE/WebSocket
calls, redirects, the PWA manifest and the service worker — under that prefix, so a browser
loading `https://example.com/codeman/` stays inside the mount.
**Forward the prefix unchanged — do NOT strip it.** Codeman expects the proxy to pass the
full path (including `/codeman/`) straight through. A minimal nginx block:
```nginx
location /codeman/ {
proxy_pass http://127.0.0.1:3000; # note: no trailing slash — keep the /codeman/ prefix
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header Upgrade $http_upgrade; # WebSocket
proxy_set_header Connection "upgrade";
}
```
Notes and current limits:
- The prefix must still be paired with `CODEMAN_ALLOWED_HOSTS` for your domain, exactly as
above — the two are independent.
- Health checks, Claude Code hooks and the docker bridge connect to the raw port directly
(bypassing the proxy), so Codeman also keeps answering at the un-prefixed paths on the port
itself. Nothing about those flows changes.
- **Web-tab (dashboard) proxying** is base-path aware: proxied dashboards have their injected
`<base>` tag, root-absolute asset rewrites, runtime `fetch`/XHR shim, `Set-Cookie` paths, and
redirects all rebased onto the mount, so they load the same under `--base-url` as at the root.
## Session cookies and rate limits
The first request prompts for HTTP Basic credentials. On success the server issues an opaque
@@ -179,22 +220,26 @@ the same IP, which matters because all tunnel traffic arrives from one loopback
## Terminal alternatives
You do not have to use a browser. `sc` is a thumb-friendly session chooser for SSH clients
like Termius or Blink:
You do not have to use a browser. `codeman tui` is a full-screen session dashboard that
works well in SSH clients like Termius or Blink:
```bash
sc # interactive chooser
sc 2 # attach to session 2
sc -l # list
codeman tui # the dashboard
codeman tui 2 # attach straight to session 2
codeman tui --list # numbered list, then exit
```
Detach with `Ctrl+A D`. The sessions are the same ones the dashboard shows.
`Enter` attaches into the pane and `F1` comes back. Under 72 columns it drops the preview
and becomes a single-column list, so it stays usable on a phone. The sessions are the same
ones the dashboard shows. See [docs/tui.md](https://github.com/Ark0N/Codeman/blob/master/docs/tui.md)
for the full guide.
## Common problems
| Symptom | Cause and fix |
| ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
| `403 host not allowed` | Your domain is not in the allowlist. Set `CODEMAN_ALLOWED_HOSTS`. |
| Assets 404 / blank page under a sub-path | Start Codeman with `--base-url /<prefix>` and have the proxy forward the prefix unchanged (don't strip it). |
| Phone shows the login page but the terminal never connects | The proxy is not forwarding WebSocket upgrades. |
| Browser warns about the certificate | Expected with `--https` and its self-signed certificate. Tailscale gives you a real one instead. |
| LAN IP does not respond, but a tunnel to the same box works | The server is bound to loopback. That is the default. A tunnel reaches it; a LAN browser cannot. |
+4 -2
View File
@@ -133,8 +133,10 @@ TUIs render correctly.
Worth knowing:
- **Scrollback.** The first time you open a session, Codeman pulls the entire tmux
scrollback, not just the recent tail. Scrolling to the very top pulls again on demand.
- **Scrollback.** Agent/TUI sessions pull their entire tmux scrollback on first open.
Shell sessions open from a bounded recent tail so a large transcript cannot stall tab
switching; press **Load full history** to pull the rest explicitly. Ordinary Shell scrolling
and automatic output recovery stay within the bounded browser buffer.
- **Wheel and touch scrolling** are forwarded into Claude's own transcript on recent Claude
versions, so the wheel scrolls the conversation rather than the terminal. `Shift+Wheel` is
always local scrollback. Other CLIs scroll locally.
+3 -1
View File
@@ -19,7 +19,9 @@ It renders what it can:
| PDF and Office documents | Converted for preview when a converter is available. |
| Anything else | Download. |
Caps: 10 MB for text preview, 50 MB for raw and download. Sensitive paths (`.env`, anything
Caps: 10 MB for text preview, 2 GB for raw and download (set `CODEMAN_MAX_DOWNLOAD_BYTES`
to change it, `0` for no limit — these bodies are streamed, so a large file costs a read
stream rather than server memory). Sensitive paths (`.env`, anything
matching credentials, `~/.ssh`, AWS credentials) are blocked from download, and SVG and HTML
are served as downloads rather than rendered, so they cannot execute in the page.
+353 -27
View File
@@ -125,6 +125,22 @@ PI_SEARCH_PATHS=(
"$HOME/bin/pi"
)
# DeepSeek Harness search paths (from src/utils/deepseek-cli-resolver.ts)
DSH_SEARCH_PATHS=(
"$HOME/.local/bin/dsh"
"/usr/local/bin/dsh"
"$HOME/.npm-global/bin/dsh"
"$HOME/bin/dsh"
)
# Grok CLI search paths (from src/utils/grok-cli-resolver.ts)
GROK_SEARCH_PATHS=(
"$HOME/.grok/bin/grok"
"$HOME/.local/bin/grok"
"/usr/local/bin/grok"
"$HOME/bin/grok"
)
# Antigravity CLI search paths (from src/utils/antigravity-cli-resolver.ts)
ANTIGRAVITY_SEARCH_PATHS=(
"$HOME/.local/bin/agy"
@@ -133,6 +149,17 @@ ANTIGRAVITY_SEARCH_PATHS=(
"$HOME/bin/agy"
)
# OMP CLI search paths (from src/utils/omp-cli-resolver.ts's OMP_SEARCH_DIRS —
# ~/.local/bin leads, omp.sh's installer target; ~/.omp/bin is a fallback only)
OMP_SEARCH_PATHS=(
"$HOME/.local/bin/omp"
"$HOME/.omp/bin/omp"
"/usr/local/bin/omp"
"$HOME/.bun/bin/omp"
"$HOME/.npm-global/bin/omp"
"$HOME/bin/omp"
)
# ============================================================================
# Color Output
# ============================================================================
@@ -229,7 +256,11 @@ print_security_notice() {
echo -e " ${YELLOW}${BOLD}Security:${NC}"
echo -e " Codeman binds ${BOLD}127.0.0.1${NC} (this machine only) — no password needed by default."
echo -e " To reach it from another device, do ONE of:"
echo -e " ${CYAN}•${NC} tailscale serve / cloudflared tunnel ${DIM}(recommended)${NC}, or"
if check_tailscale; then
echo -e " ${CYAN}•${NC} ${CYAN}bash $INSTALL_DIR/install.sh tailscale${NC} ${DIM}(Tailscale is installed here; HTTPS, recommended)${NC}, or"
else
echo -e " ${CYAN}•${NC} tailscale serve / cloudflared tunnel ${DIM}(recommended)${NC}, or"
fi
echo -e " ${CYAN}•${NC} ${CYAN}codeman web --host 0.0.0.0${NC} AND set ${CYAN}CODEMAN_PASSWORD${NC}"
echo -e " A non-loopback bind without a password still starts, but warns loudly."
echo -e " ${DIM}Details: docs/security-architecture.md${NC}"
@@ -396,6 +427,27 @@ check_tmux() {
command -v tmux &>/dev/null
}
# node-pty ships prebuilt binaries for darwin and win32 ONLY, so on Linux it is
# always compiled from source during `npm install`. Without a toolchain that
# fails deep inside node-gyp with `not found: make`, which reads like an npm bug
# rather than a missing system package (issue: fresh Ubuntu 24 server install).
# So the toolchain is checked up front, exactly like git and tmux.
#
# Returns a human-readable list of what is missing, empty when all present.
missing_build_tools() {
local missing=""
command -v make &>/dev/null || missing="make"
if ! command -v c++ &>/dev/null && ! command -v g++ &>/dev/null && ! command -v clang++ &>/dev/null; then
missing="${missing:+$missing, }a C++ compiler (g++)"
fi
command -v python3 &>/dev/null || missing="${missing:+$missing, }python3"
printf '%s' "$missing"
}
check_build_tools() {
[[ -z "$(missing_build_tools)" ]]
}
check_claude() {
# Check PATH first
if command -v claude &>/dev/null; then
@@ -569,6 +621,119 @@ get_pi_path() {
done
}
# `grok` has known squatters too (the unrelated @vibe-kit/grok-cli), so the
# server-side resolver additionally probes `grok --version`. Detection here only
# feeds the "you have no AI CLI" hint, so a plain executable test is enough.
check_grok() {
if command -v grok &>/dev/null; then
return 0
fi
for path in "${GROK_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
return 0
fi
done
return 1
}
# `dsh` is the hardest name of the lot: Debian ships an unrelated `dsh`
# (dancer's shell). The server-side resolver settles it by demanding the
# harness's own help banner; detection here only feeds the "you have no AI CLI"
# hint, so the same banner grep is enough — but unlike every sibling probe it
# EXECUTES the candidate, so it must be bounded. </dev/null is load-bearing
# twice over: a foreign binary that blocks on stdin would hang the install, and
# under `curl | bash` a child that reads stdin EATS THE REST OF THIS SCRIPT.
# The timeout (where coreutils ships one; stock macOS has none) bounds a binary
# that ignores EOF, mirroring the server resolver's own EXEC_TIMEOUT_MS.
dsh_banner_probe() {
local runner=()
if command -v timeout &>/dev/null; then runner=(timeout 5); fi
"${runner[@]}" "$1" --help </dev/null 2>/dev/null | grep -qi "DeepSeek Harness"
}
# Resolved ONCE and memoized: the probe executes a possibly-foreign binary, and
# the check/get/reminder call sites together used to re-run the whole scan many
# times per install.
DSH_RESOLVE_DONE=""
DSH_RESOLVED_PATH=""
resolve_dsh() {
[[ -n "$DSH_RESOLVE_DONE" ]] && return 0
DSH_RESOLVE_DONE=1
local candidate path
if command -v dsh &>/dev/null; then
candidate="$(command -v dsh)"
if dsh_banner_probe "$candidate"; then
DSH_RESOLVED_PATH="$candidate"
return 0
fi
fi
for path in "${DSH_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]] && dsh_banner_probe "$path"; then
DSH_RESOLVED_PATH="$path"
return 0
fi
done
return 0
}
check_dsh() {
resolve_dsh
[[ -n "$DSH_RESOLVED_PATH" ]]
}
get_dsh_path() {
resolve_dsh
echo "$DSH_RESOLVED_PATH"
}
get_grok_path() {
if command -v grok &>/dev/null; then
command -v grok
return
fi
for path in "${GROK_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
echo "$path"
return
fi
done
}
# `omp` is a short name too, so like grok/pi the server-side resolver
# additionally probes `omp --version`. Detection here only feeds the
# "you have no AI CLI" hint, so a plain executable test is enough.
check_omp() {
if command -v omp &>/dev/null; then
return 0
fi
for path in "${OMP_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
return 0
fi
done
return 1
}
get_omp_path() {
if command -v omp &>/dev/null; then
command -v omp
return
fi
for path in "${OMP_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
echo "$path"
return
fi
done
}
check_cloudflared() {
# Check ~/.local/bin first (matches tunnel-manager.ts resolution order)
if [[ -x "$HOME/.local/bin/cloudflared" ]]; then
@@ -836,6 +1001,50 @@ install_git_suse() {
run_as_root zypper install -y git
}
# Build toolchain for node-pty's source compile (see missing_build_tools).
install_buildtools_debian() {
info "Installing build tools via apt (build-essential, python3)..."
ensure_sudo
run_as_root apt-get update -qq
run_as_root apt-get install -y -qq build-essential python3
}
install_buildtools_fedora() {
info "Installing build tools (gcc, gcc-c++, make, python3)..."
ensure_sudo
if command -v dnf &>/dev/null; then
run_as_root dnf install -y gcc gcc-c++ make python3
else
run_as_root yum install -y gcc gcc-c++ make python3
fi
}
install_buildtools_arch() {
info "Installing build tools via pacman (base-devel, python)..."
ensure_sudo
run_as_root pacman -Sy --noconfirm base-devel python
}
install_buildtools_alpine() {
info "Installing build tools via apk (build-base, python3)..."
ensure_sudo
run_as_root apk add --no-cache build-base python3
}
install_buildtools_suse() {
info "Installing build tools via zypper..."
ensure_sudo
run_as_root zypper install -y gcc gcc-c++ make python3
}
install_buildtools_macos() {
# macOS normally never gets here: node-pty ships darwin prebuilds. Only a
# forced source build needs a compiler, and Xcode CLT is its only supplier.
info "Requesting Xcode Command Line Tools..."
xcode-select --install 2>/dev/null || true
die "Finish the Xcode Command Line Tools install in the dialog, then re-run this installer."
}
install_cloudflared_macos() {
info "Installing cloudflared via Homebrew..."
ensure_homebrew
@@ -1072,21 +1281,28 @@ add_to_path() {
success "Added to $profile"
}
setup_sc_alias() {
# The `sc` bash chooser was retired in favour of `codeman tui`, which reaches
# sessions 10+, carries the server's real states and leaves an attach with one
# key. Older installers wrote this alias, so take it back out.
#
# Marker-owned on purpose: it matches the exact line WE wrote, so a user's own
# `alias sc=` for something entirely different is never touched. The rewrite
# goes through `cat >` rather than `mv` so the profile keeps its own mode and
# ownership.
remove_sc_alias() {
local profile
profile=$(detect_shell_profile)
[[ -f "$profile" ]] || return 0
grep -qE "^alias sc='tmux-chooser'\$" "$profile" 2>/dev/null || return 0
# Check if alias already exists
if [[ -f "$profile" ]] && grep -qE "^alias sc=" "$profile" 2>/dev/null; then
info "Alias 'sc' already configured in $profile"
return 0
local tmp
tmp=$(mktemp 2>/dev/null) || return 0
if sed -e "/^alias sc='tmux-chooser'\$/d" \
-e '/^# Codeman tmux session shortcut$/d' "$profile" > "$tmp" 2>/dev/null; then
cat "$tmp" > "$profile"
info "Removed the retired 'sc' alias from $profile (use: codeman tui)"
fi
echo "" >> "$profile"
echo "# Codeman tmux session shortcut" >> "$profile"
echo "alias sc='tmux-chooser'" >> "$profile"
info "Added 'sc' alias for tmux-chooser"
rm -f "$tmp"
}
# ============================================================================
@@ -1704,6 +1920,41 @@ setup_tailscale_access() {
return 0
}
# A loopback install with Tailscale already connected but nothing fronting
# Codeman is one command away from working remote access — and that is exactly
# where a user lands when the first install died BEFORE the network-access
# prompt (it runs after the build, so any build failure costs the network step
# too) or when they finished a broken build by hand instead of re-running the
# installer. Detect that state on re-run and offer the retrofit, rather than
# leaving them to discover `install.sh tailscale` on their own. Never nags a
# deliberate network bind, and never nags once a serve mapping already exists.
maybe_offer_tailscale_repair() {
# A non-loopback bind already has network access; leave that choice alone.
if [[ "$EXISTING_FOUND" == "1" && -n "$EXISTING_HOST" && "$EXISTING_HOST" != "127.0.0.1" ]]; then
return 0
fi
check_tailscale || return 0
command -v node &>/dev/null || return 0
[[ "$(ts_status_field 's.BackendState')" == "Running" ]] || return 0
# Already fronting Codeman: nothing to repair.
[[ -z "$(detect_tailscale_serve_url)" ]] || return 0
echo ""
info "Tailscale is connected here, but no serve mapping fronts Codeman yet."
if [[ "$NONINTERACTIVE" == "1" ]] || ! has_tty; then
echo -e " ${DIM}Enable HTTPS access from your tailnet with:${NC} ${CYAN}bash $INSTALL_DIR/install.sh tailscale${NC}"
return 0
fi
if ! prompt_yes_no "Set up Tailscale HTTPS access now? (your tailnet is the login; no password needed)" "y"; then
echo -e " ${DIM}Any time later:${NC} ${CYAN}bash $INSTALL_DIR/install.sh tailscale${NC}"
return 0
fi
if setup_tailscale_access; then
verify_tailscale_access || true
fi
return 0
}
# `install.sh tailscale`: retrofit Tailscale access onto an existing install
# (also the target of every "set it up later" hint above).
setup_tailscale_subcommand() {
@@ -1956,6 +2207,29 @@ setup_tunnel_service() {
# Installation Helpers
# ============================================================================
# npm install with an actionable message for the failure that actually happens
# on a fresh Linux box: no toolchain, so node-pty cannot compile.
npm_install_deps() {
if npm install --quiet --no-fund --no-audit 2>/dev/null; then
return 0
fi
if npm install --no-fund --no-audit; then
return 0
fi
error "npm install failed."
if [[ "$(detect_os)" == "linux" ]] && ! check_build_tools; then
error "Missing native build tools: $(missing_build_tools)"
error "node-pty has no Linux prebuilds, so it must compile from source."
error "Install them and re-run this installer:"
error " Debian/Ubuntu: sudo apt-get install -y build-essential python3"
error " Fedora/RHEL: sudo dnf install -y gcc gcc-c++ make python3"
error " Arch: sudo pacman -S --noconfirm base-devel python"
error " Alpine: sudo apk add build-base python3"
fi
exit 1
}
install_dependency() {
local dep_name="$1"
local os="$2"
@@ -2069,6 +2343,31 @@ main() {
fi
fi
# Native build toolchain. node-pty compiles from source on Linux, so this is
# a hard requirement there, not a nicety.
if [[ "$os" == "linux" ]]; then
info "Checking build tools (node-pty compiles from source on Linux)..."
local missing_tools
missing_tools="$(missing_build_tools)"
if [[ -z "$missing_tools" ]]; then
success "Build tools are installed"
else
warn "Missing build tools: $missing_tools"
headless_guard "install build tools (system package via sudo)"
if prompt_yes_no "Install the build tools now?"; then
install_dependency "buildtools" "$os" "$distro"
hash -r 2>/dev/null || true
missing_tools="$(missing_build_tools)"
if [[ -n "$missing_tools" ]]; then
die "Build tools still missing after install: $missing_tools. Install them manually and re-run."
fi
success "Build tools installed"
else
die "A build toolchain (make, g++, python3) is required: node-pty has no Linux prebuilds and compiles from source."
fi
fi
fi
# AI CLI (Codeman drives one of: Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi)
local has_claude=false
local has_opencode=false
@@ -2076,6 +2375,9 @@ main() {
local has_gemini=false
local has_antigravity=false
local has_pi=false
local has_grok=false
local has_dsh=false
local has_omp=false
info "Checking AI CLI tools..."
if check_claude; then
@@ -2102,17 +2404,29 @@ main() {
has_pi=true
success "Pi CLI found at $(get_pi_path)"
fi
if check_grok; then
has_grok=true
success "Grok CLI found at $(get_grok_path)"
fi
if check_dsh; then
has_dsh=true
success "DeepSeek Harness found at $(get_dsh_path)"
fi
if check_omp; then
has_omp=true
success "OMP CLI found at $(get_omp_path)"
fi
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" && "$has_antigravity" == "false" && "$has_pi" == "false" ]]; then
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" && "$has_antigravity" == "false" && "$has_pi" == "false" && "$has_grok" == "false" && "$has_dsh" == "false" && "$has_omp" == "false" ]]; then
echo ""
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, Antigravity, Gemini, or Pi."
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, DeepSeek Harness, or OMP."
headless_guard "install an AI CLI (curl | bash from its vendor)"
echo ""
echo -e " ${BOLD}Which AI CLI would you like to install?${NC}"
echo -e " ${CYAN}1)${NC} Claude Code (Anthropic)"
echo -e " ${CYAN}2)${NC} OpenCode (open-source)"
echo -e " ${CYAN}3)${NC} Both"
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex, Antigravity or Pi)"
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex, Antigravity, Gemini, Pi, Grok, DeepSeek Harness or OMP)"
echo ""
local cli_choice=""
@@ -2160,6 +2474,7 @@ main() {
info "Install one later, e.g.: npm install -g @openai/codex (Codex)"
info " or: curl -fsSL https://antigravity.google/cli/install.sh | bash (Antigravity)"
info " or: npm install -g --ignore-scripts @earendil-works/pi-coding-agent (Pi)"
info " or: curl -fsSL https://x.ai/cli/install.sh | bash (Grok)"
elif [[ "$has_claude" == "false" ]] && [[ "$has_opencode" == "false" ]]; then
die "The selected AI CLI failed to install. Install one manually and re-run the installer."
fi
@@ -2225,7 +2540,7 @@ main() {
# ========================================================================
info "Installing dependencies..."
npm install --quiet --no-fund --no-audit 2>/dev/null || npm install --no-fund --no-audit
npm_install_deps
info "Building..."
npm run build --quiet 2>/dev/null || npm run build
@@ -2243,13 +2558,14 @@ main() {
ln -sf "$INSTALL_DIR/dist/index.js" "$symlink_dir/codeman"
info "Created symlink: $symlink_dir/codeman"
# Install tmux-chooser as 'tmux-chooser' command
if [[ -f "$INSTALL_DIR/scripts/tmux-chooser.sh" ]]; then
ln -sf "$INSTALL_DIR/scripts/tmux-chooser.sh" "$symlink_dir/tmux-chooser"
info "Created symlink: $symlink_dir/tmux-chooser"
# Add 'sc' alias for quick access
setup_sc_alias
# tmux-chooser/`sc` is retired; `codeman tui` replaces it. Sweep up what
# an older installer left behind, so an update does not leave a symlink
# pointing at a script this version no longer ships.
if [[ -L "$symlink_dir/tmux-chooser" ]]; then
rm -f "$symlink_dir/tmux-chooser"
info "Removed the retired tmux-chooser symlink (use: codeman tui)"
fi
remove_sc_alias
# Add ~/.local/bin to PATH if not already there
if [[ ":$PATH:" != *":$symlink_dir:"* ]]; then
@@ -2450,23 +2766,28 @@ main() {
echo -e " ${BOLD}Mobile Access (Termius/SSH):${NC}"
echo ""
echo -e " ${CYAN}sc${NC} # Interactive tmux session chooser"
echo -e " ${CYAN}sc 2${NC} # Quick attach to session 2"
echo -e " ${CYAN}sc -h${NC} # Help"
echo -e " ${CYAN}codeman tui${NC} # Full-screen session dashboard"
echo -e " ${CYAN}codeman tui 2${NC} # Attach straight to session 2"
echo -e " ${CYAN}codeman tui -l${NC} # Numbered list, then exit"
echo ""
echo -e " ${BOLD}Documentation:${NC}"
echo -e " https://github.com/Ark0N/Codeman"
echo ""
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini && ! check_antigravity && ! check_pi; then
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini && ! check_antigravity && ! check_pi && ! check_grok && ! check_dsh && ! check_omp; then
echo -e " ${YELLOW}${BOLD}Reminder:${NC} Install at least one AI CLI to start using Codeman:"
echo -e " ${CYAN}curl -fsSL https://claude.ai/install.sh | bash${NC} # Claude Code"
echo -e " ${CYAN}curl -fsSL https://opencode.ai/install | bash${NC} # OpenCode"
echo -e " ${CYAN}npm install -g @openai/codex${NC} # Codex"
echo -e " ${CYAN}curl -fsSL https://antigravity.google/cli/install.sh | bash${NC} # Antigravity"
echo -e " ${CYAN}npm install -g --ignore-scripts @earendil-works/pi-coding-agent${NC} # Pi"
echo -e " ${CYAN}curl -fsSL https://x.ai/cli/install.sh | bash${NC} # Grok"
echo -e " ${CYAN}curl -fsSL https://omp.sh/install | sh${NC} # OMP"
echo ""
echo -e " DeepSeek Harness has no vendor one-liner — install it from within Codeman"
echo -e " once the server is up (Run dropdown → Install DeepSeek Profile, or see"
echo -e " docs/deepseek-integration.md)."
fi
# Security notice — last informational block so it stays visible (when not
@@ -2519,7 +2840,7 @@ update() {
git fetch --quiet origin
git reset --hard "origin/$BRANCH" --quiet
npm install --quiet --no-fund --no-audit 2>/dev/null || npm install --no-fund --no-audit
npm_install_deps
npm run build --quiet 2>/dev/null || npm run build
date -u +%Y-%m-%dT%H:%M:%SZ > "$INSTALL_DIR/.install-complete"
success "Updated to $(node -e "console.log(require('./package.json').version)")"
@@ -2556,6 +2877,10 @@ update() {
BIND_ACK="$EXISTING_ACK"
fi
# An update is the only place a half-configured install gets a second
# chance at remote access; the fresh-install path asks outright.
maybe_offer_tailscale_repair
print_security_notice
}
@@ -2620,6 +2945,7 @@ uninstall() {
rm -f "$symlink_dir/tmux-chooser"
success "Removed symlink: $symlink_dir/tmux-chooser"
fi
remove_sc_alias
# Remove install directory
if [[ -d "$INSTALL_DIR" ]]; then
+12 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.20.0",
"version": "1.27.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.20.0",
"version": "1.27.0",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
@@ -32,6 +32,7 @@
"jpeg-js": "^0.4.4",
"node-pty": "^1.1.0",
"qrcode": "^1.5.4",
"undici": "^6.28.0",
"uuid": "^14.0.0",
"web-push": "^3.6.7",
"ws": "^8.21.0",
@@ -11550,6 +11551,15 @@
"dev": true,
"license": "MIT"
},
"node_modules/undici": {
"version": "6.28.0",
"resolved": "https://registry.npmjs.org/undici/-/undici-6.28.0.tgz",
"integrity": "sha512-LIY910g9TI13YS95lrMFrs8Rm/u/irgHeTWoKCoteeJ04CUJ92eEfj0rVn+7VKMPBpUPiUoBKfhNyLI23EE/KA==",
"license": "MIT",
"engines": {
"node": ">=18.17"
}
},
"node_modules/undici-types": {
"version": "6.21.0",
"resolved": "https://registry.npmjs.org/undici-types/-/undici-types-6.21.0.tgz",
+11 -7
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.20.0",
"version": "1.27.0",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -28,18 +28,19 @@
"test:mobile": "vitest run --config test/mobile/vitest.config.ts",
"check:frontend-syntax": "node scripts/check-frontend-syntax.mjs",
"fix:node-pty": "node scripts/fix-node-pty.mjs",
"typecheck": "tsc --noEmit",
"lint": "eslint --config config/eslint.config.js 'src/**/*.ts'",
"lint:fix": "eslint --config config/eslint.config.js 'src/**/*.ts' --fix",
"format": "prettier --write 'src/**/*.ts' 'src/web/public/**/*.{js,css,html,json}'",
"format:check": "prettier --check 'src/**/*.ts' 'src/web/public/**/*.{js,css,html,json}'",
"typecheck": "tsc --noEmit && tsc -p config/tsconfig.pr-bot.json",
"lint": "eslint --config config/eslint.config.js 'src/**/*.ts' 'scripts/pr-bot/**/*.ts'",
"lint:fix": "eslint --config config/eslint.config.js 'src/**/*.ts' 'scripts/pr-bot/**/*.ts' --fix",
"format": "prettier --write 'src/**/*.ts' 'scripts/pr-bot/**/*.ts' 'src/web/public/**/*.{js,css,html,json}'",
"format:check": "prettier --check 'src/**/*.ts' 'scripts/pr-bot/**/*.ts' 'src/web/public/**/*.{js,css,html,json}'",
"check:public-assets": "node scripts/check-public-assets.mjs",
"capture:subagents": "node scripts/capture-subagent-screenshots.mjs",
"changeset": "changeset",
"version-packages": "changeset version && npm install --package-lock-only && node scripts/check-lockfile-sync.mjs",
"check:lockfile": "node scripts/check-lockfile-sync.mjs",
"knip": "npx --yes knip@latest --config config/knip.json",
"release": "changeset publish"
"release": "changeset publish",
"pr-bot": "tsx scripts/pr-bot/main.ts"
},
"prettier": {
"singleQuote": true,
@@ -62,6 +63,8 @@
"codex",
"antigravity",
"pi",
"grok",
"deepseek",
"gemini-cli",
"ai-agents",
"agent",
@@ -100,6 +103,7 @@
"jpeg-js": "^0.4.4",
"node-pty": "^1.1.0",
"qrcode": "^1.5.4",
"undici": "^6.28.0",
"uuid": "^14.0.0",
"web-push": "^3.6.7",
"ws": "^8.21.0",
+4
View File
@@ -83,9 +83,11 @@ appendFileSync(
// 4. Minify frontend assets
run('minify input-cjk.js', 'npx esbuild dist/web/public/input-cjk.js --minify --outfile=dist/web/public/input-cjk.js --allow-overwrite');
run('minify terminal-keycode229-recovery.js', 'npx esbuild dist/web/public/terminal-keycode229-recovery.js --minify --outfile=dist/web/public/terminal-keycode229-recovery.js --allow-overwrite');
run('minify i18n.js', 'npx esbuild dist/web/public/i18n.js --minify --outfile=dist/web/public/i18n.js --allow-overwrite');
run('minify sanitize-html.js', 'npx esbuild dist/web/public/sanitize-html.js --minify --outfile=dist/web/public/sanitize-html.js --allow-overwrite');
run('minify app.js', 'npx esbuild dist/web/public/app.js --minify --outfile=dist/web/public/app.js --allow-overwrite');
run('minify tab-rail-resize.js', 'npx esbuild dist/web/public/tab-rail-resize.js --minify --outfile=dist/web/public/tab-rail-resize.js --allow-overwrite');
run('minify terminal-ui.js', 'npx esbuild dist/web/public/terminal-ui.js --minify --outfile=dist/web/public/terminal-ui.js --allow-overwrite');
run('minify respawn-ui.js', 'npx esbuild dist/web/public/respawn-ui.js --minify --outfile=dist/web/public/respawn-ui.js --allow-overwrite');
run('minify ralph-panel.js', 'npx esbuild dist/web/public/ralph-panel.js --minify --outfile=dist/web/public/ralph-panel.js --allow-overwrite');
@@ -109,8 +111,10 @@ console.log('\n[build] content-hash cache busting');
'notification-manager.js',
'keyboard-accessory.js',
'input-cjk.js',
'terminal-keycode229-recovery.js',
'sanitize-html.js',
'app.js',
'tab-rail-resize.js',
'terminal-ui.js',
'respawn-ui.js',
'ralph-panel.js',
File diff suppressed because it is too large Load Diff
+316
View File
@@ -0,0 +1,316 @@
/**
* @fileoverview Codeman HTTP client for the PR bot: spawn a claude session in a
* directory, wait until its composer is up, run one prompt to the END of its turn,
* read the answer, delete the session.
*
* This is the `skills/codeman` §0 preamble translated to TypeScript, and it keeps
* the traps that preamble documents:
* - readiness is the rendered composer (`shift+tab` in the pane), never `idle`;
* - the folder-trust dialog is READ off the screen and answered one keystroke at a
* time (Claude Code 2.1.252 highlights "No, exit" by default, so a blind Enter kills
* the session);
* - send-and-wait waits on `stop,blocked,exit`, never on the flapping `idle`, with a
* short first wait, one Enter nudge for a stranded prompt, and tagged-duplicate
* resends that re-wait without retyping (the server treats an already-applied
* (clientId, seq) frame as "wait only");
* - the bot deletes only sessions it created, by exact id.
*
* The production server is HTTPS with a self-signed certificate on loopback, so the
* undici Agent skips certificate verification for that one connection.
*/
import { Agent, fetch as undiciFetch } from 'undici';
export interface CodemanClientOptions {
apiUrl: string;
username?: string;
password?: string;
}
export interface CreateSessionOptions {
workingDir: string;
name: string;
modelOverride?: string;
effort?: string;
resumeSessionId?: string;
}
export interface WaitResult {
ended: boolean;
timedOut: boolean;
signal?: string;
}
export interface SessionRecord {
id: string;
name: string;
status: string;
pid: number | null;
claudeSessionId?: string | null;
workingDir: string;
mode: string;
}
export type TurnOutcome =
| { kind: 'stop' }
| { kind: 'blocked' }
| { kind: 'exit' }
| { kind: 'timeout' }
| { kind: 'limit'; message: string };
/**
* Claude Code answers a spent model budget INSIDE the turn ("You've reached your Fable
* limit. Run /usage-credits to continue or switch models with /model.") and then simply
* sits there with nothing to write. Measured 2026-09-08: four reviews each burned their
* whole 40-minute deadline and reported a bare "timed out without a report", which reads
* as a hung reviewer rather than an account that needs attention, and the retries spent
* the per-head budget so the PRs would not have been picked up again once credits
* returned. Matching the notice turns 40 silent minutes into a named failure in seconds.
*
* Deliberately model-agnostic: the same sentence is printed for every model, and the
* apostrophe is typographic on the pane, so neither the model name nor `'` is matched.
*/
const MODEL_LIMIT_PATTERN = /reached your [^\n]{0,40}?\blimit\b|\/usage-credits/i;
/**
* Thrown instead of a plain Error when a review died on a spent model budget, so the
* caller can tell an account condition apart from a review that genuinely failed.
*/
export class ModelLimitError extends Error {
override readonly name = 'ModelLimitError';
}
/** The limit notice as one clean line, or undefined if the screen does not carry it. */
export function findModelLimitNotice(screen: string): string | undefined {
const line = stripAnsi(screen)
.split('\n')
.find((l) => MODEL_LIMIT_PATTERN.test(l));
return line?.replace(/^[\s>|]*(?:\u23bf|\u2514|\u256d|\u2570|\u23a2|\u2502|\u23bd)?\s*/u, '').trim() || undefined;
}
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
export function stripAnsi(text: string): string {
// eslint-disable-next-line no-control-regex
return text.replace(/\x1b\[[0-9;?]*[a-zA-Z]/g, '').replace(/\x1b[()][AB0]/g, '');
}
/** Which key answers the trust dialog right now, read from the rendered pane. */
export function trustDialogKey(screen: string): 'confirm' | 'move' | null {
const compact = stripAnsi(screen).replace(/\s+/g, '');
const matches = compact.match(/❯[0-9.]*(yes,itrustthisfolder|no,exit)/gi);
if (!matches || matches.length === 0) return null;
const last = matches[matches.length - 1].toLowerCase();
return last.includes('yes,') ? 'confirm' : 'move';
}
export class CodemanClient {
// headersTimeout/bodyTimeout default to 300 s in undici, which is shorter than one
// long-poll slice on the wait endpoints (up to 580 s): the first review died at
// exactly five minutes with a bare "fetch failed". The per-request AbortSignal is
// the only ceiling here.
private readonly agent = new Agent({ connect: { rejectUnauthorized: false }, headersTimeout: 0, bodyTimeout: 0 });
private readonly authHeader?: string;
constructor(private readonly opts: CodemanClientOptions) {
if (opts.password) {
this.authHeader = 'Basic ' + Buffer.from(`${opts.username || 'admin'}:${opts.password}`).toString('base64');
}
}
private async request<T>(
method: string,
path: string,
body?: unknown,
query?: Record<string, string | number | undefined>,
timeoutMs = 60_000
): Promise<T> {
const url = new URL(this.opts.apiUrl + path);
for (const [k, v] of Object.entries(query ?? {})) if (v !== undefined) url.searchParams.set(k, String(v));
const headers: Record<string, string> = { Accept: 'application/json' };
if (this.authHeader) headers.Authorization = this.authHeader;
if (body !== undefined) headers['Content-Type'] = 'application/json';
let res;
try {
res = await undiciFetch(url, {
method,
headers,
body: body === undefined ? undefined : JSON.stringify(body),
dispatcher: this.agent,
signal: AbortSignal.timeout(timeoutMs),
});
} catch (err) {
const cause = (err as { cause?: { message?: string; code?: string } }).cause;
const detail = cause ? ` (${cause.code ?? ''} ${cause.message ?? ''})`.replace(/\(\s+/, '(').trim() : '';
throw new Error(`${method} ${path}: ${(err as Error).message}${detail}`);
}
const text = await res.text();
let json: { success?: boolean; data?: T; error?: string; errorCode?: string } & Record<string, unknown> = {};
try {
json = text ? JSON.parse(text) : {};
} catch {
throw new Error(`${method} ${path}: non-JSON ${res.status} response: ${text.slice(0, 200)}`);
}
if (!res.ok || json.success === false) {
throw new Error(
`${method} ${path}: ${res.status} ${json.errorCode ?? ''} ${json.error ?? text.slice(0, 200)}`.trim()
);
}
// Most routes use the {success, data} envelope; a few legacy GETs return the raw shape.
return (json.success === true && json.data !== undefined ? json.data : json) as T;
}
async status(): Promise<{ version?: string }> {
return this.request<{ version?: string }>('GET', '/api/status');
}
async listSessions(): Promise<SessionRecord[]> {
const data = await this.request<SessionRecord[] | { sessions: SessionRecord[] }>('GET', '/api/sessions');
return Array.isArray(data) ? data : (data.sessions ?? []);
}
async getSession(id: string): Promise<SessionRecord> {
return this.request<SessionRecord>('GET', `/api/sessions/${id}`);
}
/** Create + start. Creation alone leaves pid null and no pane, so the two are one step here. */
async createInteractiveSession(opts: CreateSessionOptions): Promise<string> {
const created = await this.request<{ session: { id: string } }>('POST', '/api/sessions', {
workingDir: opts.workingDir,
mode: 'claude',
name: opts.name,
modelOverride: opts.modelOverride,
effort: opts.effort,
resumeSessionId: opts.resumeSessionId,
});
const id = created.session?.id;
if (!id) throw new Error('POST /api/sessions returned no session id');
await this.request('POST', `/api/sessions/${id}/interactive`, {});
return id;
}
async deleteSession(id: string): Promise<void> {
if (!id || id.length < 8) throw new Error(`refusing to delete session "${id}"`);
await this.request('DELETE', `/api/sessions/${id}`);
}
async waitOutput(id: string, match: string, from: 'now' | 'buffer', timeoutMs: number): Promise<boolean> {
const data = await this.request<{ wait?: { matched?: boolean } }>(
'GET',
`/api/sessions/${id}/wait-output`,
undefined,
{ match, from, timeout: timeoutMs },
timeoutMs + 15_000
);
return Boolean(data.wait?.matched);
}
async waitSignal(id: string, until: string, timeoutMs: number): Promise<WaitResult> {
const data = await this.request<{ wait?: WaitResult }>(
'GET',
`/api/sessions/${id}/wait`,
undefined,
{ until, timeout: timeoutMs },
timeoutMs + 15_000
);
return data.wait ?? { ended: false, timedOut: true };
}
async terminalText(id: string): Promise<string> {
const data = await this.request<{ terminalBuffer?: string }>('GET', `/api/sessions/${id}/terminal`, undefined, {
full: '1',
});
return data.terminalBuffer ?? '';
}
async sendKeys(id: string, input: string, clientId: string, seq: number): Promise<void> {
await this.request('POST', `/api/sessions/${id}/input`, { input, useMux: true, clientId, seq });
}
async lastResponse(id: string): Promise<string> {
const data = await this.request<{ text?: string }>('GET', `/api/sessions/${id}/last-response`);
return data.text ?? '';
}
/** Composer wait, trust-dialog fallback, composer wait again. Throws when the pane never gets there. */
async ensureReady(id: string, log: (m: string) => void): Promise<void> {
if (await this.waitOutput(id, 'shift+tab', 'buffer', 5000)) return;
for (let i = 1; i <= 6; i++) {
const key = trustDialogKey(await this.terminalText(id));
if (!key) break;
log(`trust dialog on screen: ${key === 'confirm' ? 'Enter' : 'arrow down'}`);
await this.sendKeys(id, key === 'confirm' ? '\r' : '\x1b[B', `prbot-trust-${id}`, i);
if (key === 'confirm') break;
await sleep(1000);
}
if (await this.waitOutput(id, 'shift+tab', 'buffer', 45_000)) return;
throw new Error('the session never drew its composer (no `shift+tab` in the pane after 50s)');
}
/**
* Send ONE prompt and block until the turn ends, the session blocks on a question,
* the pane exits, or `deadlineMs` passes. `isDone` lets the caller finish early on
* an out-of-band signal (the report file appearing), which also covers a stop edge
* that fired between two waits.
*/
async runTurn(
id: string,
prompt: string,
opts: { deadlineMs: number; isDone?: () => boolean; log: (m: string) => void }
): Promise<TurnOutcome> {
if (prompt.includes('\n'))
throw new Error('runTurn prompts must be single-line (embedded newlines are stripped by tmux)');
const clientId = `prbot-${id}`;
const seq = Math.floor(Date.now() / 1000);
const frame = { input: prompt + '\r', useMux: true, clientId, seq, wait: 'stop,blocked,exit', waitTimeout: 20_000 };
const started = Date.now();
const post = (body: unknown, timeout: number) =>
this.request<{ delivered?: boolean; wait?: WaitResult }>(
'POST',
`/api/sessions/${id}/input`,
body,
undefined,
timeout + 15_000
);
let r = await post(frame, 20_000);
if (!r.delivered) throw new Error('the prompt was not delivered (pane dead?)');
let wait = r.wait;
let nudged = false;
// Only consulted when the turn produced nothing, so a review that merely QUOTES the
// notice in its report cannot be mistaken for one that hit it.
const limitNotice = async (): Promise<string | undefined> =>
findModelLimitNotice(await this.terminalText(id).catch(() => ''));
while (true) {
if (wait && !wait.timedOut) {
const outcome = toOutcome(wait);
if (outcome.kind === 'stop' && !opts.isDone?.()) {
const limit = await limitNotice();
if (limit) return { kind: 'limit', message: limit };
}
return outcome;
}
if (opts.isDone?.()) return { kind: 'stop' };
const limit = await limitNotice();
if (limit) return { kind: 'limit', message: limit };
const remaining = opts.deadlineMs - (Date.now() - started);
if (remaining <= 0) return { kind: 'timeout' };
if (!nudged) {
// An Ink repaint occasionally eats the Enter: a bare \r is the missing key when
// the prompt is stranded and a no-op when the turn is genuinely running.
nudged = true;
await this.sendKeys(id, '\r', clientId, seq + 1);
}
const slice = Math.min(remaining, 580_000);
opts.log(`still working (${Math.round((Date.now() - started) / 60_000)} min)`);
r = await post({ ...frame, waitTimeout: slice }, slice);
wait = r.wait;
}
}
}
function toOutcome(wait: WaitResult): TurnOutcome {
const signal = wait.signal ?? '';
if (signal === 'blocked') return { kind: 'blocked' };
if (signal === 'exit') return { kind: 'exit' };
return { kind: 'stop' };
}
+194
View File
@@ -0,0 +1,194 @@
/**
* @fileoverview PR bot configuration.
*
* Read from `~/.codeman/pr-bot.env` (KEY=VALUE lines, mode 0600, the same shape as
* the data dir's `.env`) with the process environment layered on top, then validated
* into a typed config. `parseEnvFile` and `buildConfig` are pure so the validation
* rules are unit-testable without touching the filesystem.
*
* Nothing here reads Codeman's own settings: the bot is maintainer tooling that
* drives a running Codeman over HTTP, it is not part of the server.
*/
import { existsSync, readFileSync } from 'fs';
import { homedir } from 'os';
import { dirname, join, resolve } from 'path';
import { fileURLToPath } from 'url';
export interface PrBotConfig {
/** Telegram bot token from BotFather. */
telegramBotToken: string;
/** The ONE chat the bot talks to and accepts commands from. Everything else is ignored. */
telegramChatId: string;
/** `owner/name` of the repository whose PRs are reviewed. */
githubRepo: string;
/** Codeman server the review sessions are spawned on. */
codemanApiUrl: string;
codemanUsername?: string;
codemanPassword?: string;
/** How often open PRs are listed. */
pollIntervalMs: number;
/** The maintainer's checkout; worktrees are added from its git dir. Never checked out by the bot. */
mainCheckout: string;
/** State, reports and worktrees live under here. */
dataDir: string;
worktreesDir: string;
/** Optional model / effort for the review sessions (Codeman `modelOverride` / `effort`). */
model?: string;
effort?: string;
/** Hard ceiling for one review turn. */
reviewTimeoutMs: number;
/** Hard ceiling for one follow-up turn. */
followupTimeoutMs: number;
/** When false, PRs are only reviewed on an explicit `/review N`. */
autoReview: boolean;
/** Draft PRs are skipped unless this is on. */
reviewDrafts: boolean;
}
export const CONFIG_FILE_NAME = 'pr-bot.env';
/**
* The maintainer's existing Telegram notifier bot (a separate, send-only process)
* keeps its token and chat id here. The PR bot shares that bot identity by default,
* so it reads those two keys from the same file rather than making anyone copy a
* secret around. Override with `PR_BOT_TELEGRAM_ENV_FILE`.
*/
export const DEFAULT_TELEGRAM_ENV_FILE = join('codeman-cases', 'telegram', '.env');
const SHARED_TELEGRAM_KEYS = ['TELEGRAM_BOT_TOKEN', 'TELEGRAM_CHAT_ID'] as const;
/** The keys the env file understands, for `check` and the docs. */
export const CONFIG_KEYS = [
'TELEGRAM_BOT_TOKEN',
'TELEGRAM_CHAT_ID',
'GITHUB_REPO',
'CODEMAN_API_URL',
'CODEMAN_USERNAME',
'CODEMAN_PASSWORD',
'PR_BOT_POLL_INTERVAL',
'PR_BOT_MAIN_CHECKOUT',
'PR_BOT_DATA_DIR',
'PR_BOT_MODEL',
'PR_BOT_EFFORT',
'PR_BOT_REVIEW_TIMEOUT',
'PR_BOT_FOLLOWUP_TIMEOUT',
'PR_BOT_AUTO_REVIEW',
'PR_BOT_REVIEW_DRAFTS',
'PR_BOT_TELEGRAM_ENV_FILE',
] as const;
/** Parse `KEY=VALUE` lines. Comments, blanks, `export ` prefixes and matching quotes are handled. */
export function parseEnvFile(text: string): Record<string, string> {
const out: Record<string, string> = {};
for (const rawLine of text.split(/\r?\n/)) {
const line = rawLine.trim();
if (!line || line.startsWith('#')) continue;
const eq = line.indexOf('=');
if (eq <= 0) continue;
const key = line
.slice(0, eq)
.trim()
.replace(/^export\s+/, '');
let value = line.slice(eq + 1).trim();
if (value.length >= 2) {
const first = value[0];
const last = value[value.length - 1];
if ((first === '"' && last === '"') || (first === "'" && last === "'")) value = value.slice(1, -1);
}
if (/^[A-Z_][A-Z0-9_]*$/.test(key)) out[key] = value;
}
return out;
}
function intFrom(raw: string | undefined, fallback: number, min: number): number {
const n = parseInt(raw ?? '', 10);
if (!Number.isFinite(n) || n <= 0) return fallback;
return Math.max(min, n);
}
function flagFrom(raw: string | undefined, fallback: boolean): boolean {
if (raw === undefined || raw === '') return fallback;
return !['0', 'false', 'no', 'off'].includes(raw.trim().toLowerCase());
}
/** Build the typed config from an env map. Throws with every missing key named at once. */
export function buildConfig(
env: Record<string, string | undefined>,
defaults: { home: string; repoRoot: string }
): PrBotConfig {
const missing: string[] = [];
const telegramBotToken = env.TELEGRAM_BOT_TOKEN?.trim() ?? '';
const telegramChatId = env.TELEGRAM_CHAT_ID?.trim() ?? '';
if (!telegramBotToken) missing.push('TELEGRAM_BOT_TOKEN');
if (!telegramChatId) missing.push('TELEGRAM_CHAT_ID');
if (missing.length) throw new Error(`pr-bot config is missing: ${missing.join(', ')}`);
const githubRepo = env.GITHUB_REPO?.trim() || 'Ark0N/Codeman';
if (!/^[\w.-]+\/[\w.-]+$/.test(githubRepo)) throw new Error(`GITHUB_REPO must be owner/name, got "${githubRepo}"`);
const codemanApiUrl = (env.CODEMAN_API_URL?.trim() || 'https://127.0.0.1:3000').replace(/\/+$/, '');
if (!/^https?:\/\//.test(codemanApiUrl))
throw new Error(`CODEMAN_API_URL must be http(s)://..., got "${codemanApiUrl}"`);
const dataDir = resolve(env.PR_BOT_DATA_DIR?.trim() || join(defaults.home, '.codeman', 'pr-bot'));
const mainCheckout = resolve(env.PR_BOT_MAIN_CHECKOUT?.trim() || defaults.repoRoot);
return {
telegramBotToken,
telegramChatId,
githubRepo,
codemanApiUrl,
codemanUsername: env.CODEMAN_USERNAME?.trim() || undefined,
codemanPassword: env.CODEMAN_PASSWORD || undefined,
pollIntervalMs: intFrom(env.PR_BOT_POLL_INTERVAL, 600, 60) * 1000,
mainCheckout,
dataDir,
worktreesDir: join(dataDir, 'worktrees'),
model: env.PR_BOT_MODEL?.trim() || undefined,
effort: env.PR_BOT_EFFORT?.trim() || undefined,
reviewTimeoutMs: intFrom(env.PR_BOT_REVIEW_TIMEOUT, 40, 5) * 60_000,
followupTimeoutMs: intFrom(env.PR_BOT_FOLLOWUP_TIMEOUT, 20, 2) * 60_000,
autoReview: flagFrom(env.PR_BOT_AUTO_REVIEW, true),
reviewDrafts: flagFrom(env.PR_BOT_REVIEW_DRAFTS, false),
};
}
/** The repository this script lives in (scripts/pr-bot/ -> repo root). */
export function scriptRepoRoot(): string {
return resolve(dirname(fileURLToPath(import.meta.url)), '..', '..');
}
export function configFilePath(): string {
return join(process.env.CODEMAN_DATA_DIR || join(homedir(), '.codeman'), CONFIG_FILE_NAME);
}
export function telegramEnvFilePath(fromFile: Record<string, string>): string {
return resolve(
process.env.PR_BOT_TELEGRAM_ENV_FILE ||
fromFile.PR_BOT_TELEGRAM_ENV_FILE ||
join(homedir(), DEFAULT_TELEGRAM_ENV_FILE)
);
}
/**
* Layers, lowest first: the shared Telegram notifier's `.env` (token + chat id only),
* then `~/.codeman/pr-bot.env`, then the process environment, so a one-off
* `PR_BOT_MODEL=... npx tsx ...` wins over everything.
*/
export function loadConfig(): PrBotConfig {
const file = configFilePath();
const fromFile = existsSync(file) ? parseEnvFile(readFileSync(file, 'utf8')) : {};
const sharedFile = telegramEnvFilePath(fromFile);
const shared = existsSync(sharedFile) ? parseEnvFile(readFileSync(sharedFile, 'utf8')) : {};
const merged: Record<string, string | undefined> = {};
for (const key of SHARED_TELEGRAM_KEYS) if (shared[key]) merged[key] = shared[key];
Object.assign(merged, fromFile);
for (const key of CONFIG_KEYS) {
const v = process.env[key];
if (v !== undefined && v !== '') merged[key] = v;
}
try {
return buildConfig(merged, { home: homedir(), repoRoot: scriptRepoRoot() });
} catch (err) {
throw new Error(`${(err as Error).message} (config file: ${file}; shared Telegram env: ${sharedFile})`);
}
}
+230
View File
@@ -0,0 +1,230 @@
/**
* @fileoverview GitHub access for the PR bot, entirely through the `gh` CLI.
*
* `gh` carries the maintainer's own login, so the bot needs no token of its own and
* every write (merge, close, comment, CI approval) lands under that account. That is
* why every write here is only ever reached from an explicit, confirmed Telegram
* command (see bot.ts); nothing in this file is called on a timer.
*
* `classifyCi` and `latestRunPerWorkflow` are pure and unit-tested.
*/
import { execFile } from 'child_process';
import { promisify } from 'util';
const execFileAsync = promisify(execFile);
export interface PrSummary {
number: number;
title: string;
author: string;
headSha: string;
baseRef: string;
headRef: string;
isDraft: boolean;
mergeable: 'MERGEABLE' | 'CONFLICTING' | 'UNKNOWN';
mergeState: string;
additions: number;
deletions: number;
changedFiles: number;
updatedAt: string;
url: string;
isCrossRepository: boolean;
labels: string[];
}
export interface PrFile {
path: string;
additions: number;
deletions: number;
}
export interface PrDetail extends PrSummary {
body: string;
files: PrFile[];
authorAssociation: string;
linkedIssues: { number: number; title: string }[];
commitCount: number;
commentCount: number;
reviewDecision: string;
headRepo: string;
}
export interface WorkflowRun {
id: number;
name: string;
status: string;
conclusion: string | null;
}
export type CiState = 'passed' | 'failed' | 'pending' | 'awaiting-approval' | 'none';
export interface CiStatus {
state: CiState;
runs: WorkflowRun[];
}
const PR_LIST_FIELDS =
'number,title,author,headRefOid,baseRefName,headRefName,isDraft,mergeable,mergeStateStatus,additions,deletions,changedFiles,updatedAt,url,isCrossRepository,labels';
export async function gh(args: string[], opts: { timeoutMs?: number; input?: string } = {}): Promise<string> {
const child = execFileAsync('gh', args, {
maxBuffer: 32 * 1024 * 1024,
timeout: opts.timeoutMs ?? 60_000,
env: { ...process.env, GH_PROMPT_DISABLED: '1', GH_NO_UPDATE_NOTIFIER: '1' },
});
if (opts.input !== undefined && child.child.stdin) {
child.child.stdin.end(opts.input);
}
const { stdout } = await child;
return stdout;
}
interface RawPr {
number: number;
title: string;
author?: { login?: string };
headRefOid: string;
baseRefName: string;
headRefName: string;
isDraft: boolean;
mergeable: string;
mergeStateStatus: string;
additions: number;
deletions: number;
changedFiles: number;
updatedAt: string;
url: string;
isCrossRepository: boolean;
labels?: { name: string }[];
}
function toSummary(raw: RawPr): PrSummary {
const mergeable = raw.mergeable === 'MERGEABLE' || raw.mergeable === 'CONFLICTING' ? raw.mergeable : 'UNKNOWN';
return {
number: raw.number,
title: raw.title ?? '',
author: raw.author?.login ?? 'unknown',
headSha: raw.headRefOid,
baseRef: raw.baseRefName,
headRef: raw.headRefName,
isDraft: Boolean(raw.isDraft),
mergeable,
mergeState: raw.mergeStateStatus ?? 'UNKNOWN',
additions: raw.additions ?? 0,
deletions: raw.deletions ?? 0,
changedFiles: raw.changedFiles ?? 0,
updatedAt: raw.updatedAt ?? '',
url: raw.url,
isCrossRepository: Boolean(raw.isCrossRepository),
labels: (raw.labels ?? []).map((l) => l.name),
};
}
export async function listOpenPrs(repo: string): Promise<PrSummary[]> {
const out = await gh(['pr', 'list', '--repo', repo, '--state', 'open', '--limit', '100', '--json', PR_LIST_FIELDS]);
const raw = JSON.parse(out) as RawPr[];
return raw.map(toSummary);
}
export async function getPrDetail(repo: string, number: number): Promise<PrDetail> {
const fields = `${PR_LIST_FIELDS},body,files,commits,comments,reviewDecision,closingIssuesReferences,headRepository,headRepositoryOwner`;
const out = await gh(['pr', 'view', String(number), '--repo', repo, '--json', fields]);
const raw = JSON.parse(out) as RawPr & {
body?: string;
files?: { path: string; additions: number; deletions: number }[];
commits?: unknown[];
comments?: unknown[];
reviewDecision?: string;
closingIssuesReferences?: { number: number; title: string }[];
headRepository?: { name?: string };
headRepositoryOwner?: { login?: string };
};
let authorAssociation = 'NONE';
try {
const assoc = await gh(['api', `repos/${repo}/pulls/${number}`, '--jq', '.author_association']);
authorAssociation = assoc.trim() || 'NONE';
} catch {
// Metadata only; a failed lookup must not fail the review.
}
const owner = raw.headRepositoryOwner?.login;
const name = raw.headRepository?.name;
return {
...toSummary(raw),
body: raw.body ?? '',
files: (raw.files ?? []).map((f) => ({ path: f.path, additions: f.additions ?? 0, deletions: f.deletions ?? 0 })),
authorAssociation,
linkedIssues: (raw.closingIssuesReferences ?? []).map((i) => ({ number: i.number, title: i.title })),
commitCount: raw.commits?.length ?? 0,
commentCount: raw.comments?.length ?? 0,
reviewDecision: raw.reviewDecision ?? '',
headRepo: owner && name ? `${owner}/${name}` : '',
};
}
/** The API returns newest first; keep only the newest run of each workflow. */
export function latestRunPerWorkflow(runs: WorkflowRun[]): WorkflowRun[] {
const seen = new Set<string>();
const out: WorkflowRun[] = [];
for (const run of runs) {
if (seen.has(run.name)) continue;
seen.add(run.name);
out.push(run);
}
return out;
}
/**
* Collapse workflow runs into one word the report can show. `action_required` is
* the fork-PR case where GitHub waits for a maintainer to approve the run: the PR
* looks unchecked and stays that way until someone clicks, so it gets its own state.
*/
export function classifyCi(runs: WorkflowRun[]): CiState {
const latest = latestRunPerWorkflow(runs);
if (latest.length === 0) return 'none';
if (latest.some((r) => r.conclusion === 'action_required')) return 'awaiting-approval';
if (latest.some((r) => ['queued', 'in_progress', 'waiting', 'pending', 'requested'].includes(r.status)))
return 'pending';
if (latest.some((r) => ['failure', 'timed_out', 'cancelled', 'startup_failure'].includes(r.conclusion ?? '')))
return 'failed';
if (latest.every((r) => ['success', 'skipped', 'neutral'].includes(r.conclusion ?? ''))) return 'passed';
return 'pending';
}
export async function getCiStatus(repo: string, headSha: string): Promise<CiStatus> {
const out = await gh([
'api',
`repos/${repo}/actions/runs?head_sha=${headSha}&event=pull_request&per_page=30`,
'--jq',
'[.workflow_runs[] | {id, name, status, conclusion}]',
]);
const runs = JSON.parse(out) as WorkflowRun[];
return { state: classifyCi(runs), runs: latestRunPerWorkflow(runs) };
}
export async function approveWorkflowRun(repo: string, runId: number): Promise<void> {
await gh(['api', '-X', 'POST', `repos/${repo}/actions/runs/${runId}/approve`]);
}
/** Merge commits, matching the repository's history (`Merge pull request #N from ...`). */
export async function mergePr(repo: string, number: number): Promise<string> {
return gh(['pr', 'merge', String(number), '--repo', repo, '--merge'], { timeoutMs: 120_000 });
}
export async function closePr(repo: string, number: number, comment: string): Promise<string> {
const args = ['pr', 'close', String(number), '--repo', repo];
if (comment.trim()) args.push('--comment', comment);
return gh(args);
}
export async function commentPr(repo: string, number: number, body: string): Promise<string> {
return gh(['pr', 'comment', String(number), '--repo', repo, '--body-file', '-'], { input: body });
}
export async function ghAuthOk(): Promise<boolean> {
try {
await gh(['auth', 'status']);
return true;
} catch {
return false;
}
}
+281
View File
@@ -0,0 +1,281 @@
#!/usr/bin/env -S npx tsx
/**
* @fileoverview CLI entry for the PR bot.
*
* npx tsx scripts/pr-bot/main.ts run # the daemon (what the service runs)
* npx tsx scripts/pr-bot/main.ts check # config, gh, Codeman, Telegram, git
* npx tsx scripts/pr-bot/main.ts scan # list open PRs and what would be queued
* npx tsx scripts/pr-bot/main.ts review N [--no-telegram] # one review, now
* npx tsx scripts/pr-bot/main.ts status # what the state file knows
* npx tsx scripts/pr-bot/main.ts notify N # resend PR N's review message to Telegram
* npx tsx scripts/pr-bot/main.ts install-service # systemd user unit, enabled + started
* npx tsx scripts/pr-bot/main.ts uninstall-service
*
* User guide: docs/pr-bot.md
*/
import { execFileSync } from 'child_process';
import { existsSync, mkdirSync, writeFileSync } from 'fs';
import { homedir } from 'os';
import { join } from 'path';
import { PrBot, type TelegramLike } from './bot.js';
import { CodemanClient } from './codeman-client.js';
import { configFilePath, loadConfig, type PrBotConfig } from './config.js';
import { ghAuthOk, listOpenPrs } from './github.js';
import { orderBacklog } from './report.js';
import { StateStore } from './state.js';
import { TelegramClient } from './telegram.js';
const SERVICE_NAME = 'codeman-pr-bot';
function log(msg: string): void {
console.log(`${new Date().toISOString()} ${msg}`);
}
/** Prints what the bot would have sent; used by `review --no-telegram`. */
class ConsoleTelegram implements TelegramLike {
private nextId = 1;
isOurChat(): boolean {
return true;
}
async sendMessage(text: string): Promise<number> {
console.log(`\n--- telegram (html) ---\n${text}\n---`);
return this.nextId++;
}
async sendPlain(text: string): Promise<number> {
console.log(`\n--- telegram (plain) ---\n${text}\n---`);
return this.nextId++;
}
async editReplyMarkup(): Promise<void> {}
async deleteMessage(): Promise<void> {}
async answerCallback(): Promise<void> {}
async sendDocument(filename: string, content: string): Promise<void> {
console.log(`\n--- telegram document ${filename} (${content.length} chars) ---`);
}
async getUpdates(): Promise<[]> {
return [];
}
async setMyCommands(): Promise<void> {}
}
function makeCodeman(cfg: PrBotConfig): CodemanClient {
return new CodemanClient({ apiUrl: cfg.codemanApiUrl, username: cfg.codemanUsername, password: cfg.codemanPassword });
}
export function logFilePath(cfg: PrBotConfig): string {
return join(cfg.dataDir, 'bot.log');
}
function unitFile(cfg: PrBotConfig): string {
const tsx = join(cfg.mainCheckout, 'node_modules', '.bin', 'tsx');
// A user service gets a minimal PATH, which is where `gh` (and an nvm/Homebrew
// node) are not: the first run failed its scan with `spawn gh ENOENT`. Bake the
// installing shell's PATH in, as `codeman service install` does.
const seen = new Set<string>();
const path = (process.env.PATH || '/usr/local/bin:/usr/bin:/bin')
.split(':')
.filter((p) => p && !p.endsWith('/node_modules/.bin') && !seen.has(p) && seen.add(p))
.join(':');
return `[Unit]
Description=Codeman PR review bot (Telegram)
After=network-online.target
Wants=network-online.target
StartLimitIntervalSec=300
StartLimitBurst=5
[Service]
Type=simple
WorkingDirectory=${cfg.mainCheckout}
ExecStart=${tsx} scripts/pr-bot/main.ts run
Restart=always
RestartSec=15
Environment=HOME=${homedir()}
Environment=NODE_ENV=production
Environment=PATH=${path}
# A file rather than the journal: on some boxes \`journalctl --user\` cannot read
# the user journal at all, and a review bot whose logs cannot be found is not
# debuggable from a phone.
StandardOutput=append:${logFilePath(cfg)}
StandardError=append:${logFilePath(cfg)}
SyslogIdentifier=${SERVICE_NAME}
[Install]
WantedBy=default.target
`;
}
async function cmdCheck(): Promise<void> {
const cfg = loadConfig();
console.log(
`config file: ${configFilePath()}${existsSync(configFilePath()) ? '' : ' (absent, defaults + shared Telegram env)'}`
);
console.log(`repo: ${cfg.githubRepo}`);
console.log(`codeman: ${cfg.codemanApiUrl}`);
console.log(`main checkout: ${cfg.mainCheckout}`);
console.log(`data dir: ${cfg.dataDir}`);
console.log(`model: ${cfg.model ?? '(session default)'}, effort: ${cfg.effort ?? '(default)'}`);
console.log(
`poll: every ${cfg.pollIntervalMs / 60_000} min; review timeout ${cfg.reviewTimeoutMs / 60_000} min; auto-review ${cfg.autoReview}`
);
let ok = true;
const step = async (name: string, fn: () => Promise<string>) => {
try {
console.log(`✔ ${name}: ${await fn()}`);
} catch (err) {
ok = false;
console.log(`✘ ${name}: ${(err as Error).message}`);
}
};
await step('gh auth', async () =>
(await ghAuthOk()) ? 'logged in' : Promise.reject(new Error('run `gh auth login`'))
);
await step('git', async () =>
execFileSync('git', ['-C', cfg.mainCheckout, 'rev-parse', '--git-dir'], { encoding: 'utf8' }).trim()
);
await step('codeman', async () => {
const s = await makeCodeman(cfg).status();
return `up (version ${s.version ?? 'unknown'})`;
});
await step('telegram', async () => {
const me = await new TelegramClient(cfg.telegramBotToken, cfg.telegramChatId).getMe();
return `@${me.username ?? '?'} for chat ${cfg.telegramChatId}`;
});
await step('open PRs', async () => `${(await listOpenPrs(cfg.githubRepo)).length}`);
if (!ok) process.exit(1);
}
async function cmdScan(): Promise<void> {
const cfg = loadConfig();
const store = new StateStore(join(cfg.dataDir, 'state.json'));
const open = await listOpenPrs(cfg.githubRepo);
const rows = orderBacklog(open).map((pr) => {
const rec = store.pr(pr.number);
const state =
rec?.reviewedSha === pr.headSha ? `reviewed (${rec?.verdict ?? '?'})` : rec?.reviewedSha ? 'updated' : 'new';
const flags = [pr.isDraft ? 'draft' : '', pr.mergeable === 'CONFLICTING' ? 'conflicts' : '']
.filter(Boolean)
.join(', ');
return `#${pr.number}\t${state}\t+${pr.additions}/-${pr.deletions}\t${pr.author}\t${pr.title}${flags ? ` [${flags}]` : ''}`;
});
console.log(`${open.length} open PRs in review order:\n${rows.join('\n')}`);
}
async function cmdStatus(): Promise<void> {
const cfg = loadConfig();
const store = new StateStore(join(cfg.dataDir, 'state.json'));
console.log(`paused: ${store.state.paused}; telegram offset: ${store.state.telegramOffset}`);
for (const rec of Object.values(store.state.prs).sort((a, b) => b.number - a.number)) {
console.log(
`#${rec.number}\t${rec.status}\t${rec.verdict ?? '-'}\t${rec.reviewedSha?.slice(0, 8) ?? '-'}\t${rec.author}\t${rec.title}${
rec.lastError ? `\n\t${rec.lastError.split('\n')[0]}` : ''
}`
);
}
}
async function cmdReview(args: string[]): Promise<void> {
const number = parseInt(args.find((a) => /^\d+$/.test(a)) ?? '', 10);
if (!Number.isFinite(number)) throw new Error('usage: review <pr-number> [--no-telegram]');
const cfg = loadConfig();
const telegram = args.includes('--no-telegram')
? new ConsoleTelegram()
: new TelegramClient(cfg.telegramBotToken, cfg.telegramChatId);
const bot = new PrBot(cfg, { telegram, codeman: makeCodeman(cfg), log });
const rec = await bot.reviewPr(number);
console.log(
`\n#${number}: ${rec.status}${rec.verdict ? ` (${rec.verdict})` : ''}${rec.lastError ? `\n${rec.lastError}` : ''}`
);
if (rec.reportMdPath) console.log(`report: ${rec.reportMdPath}`);
process.exit(rec.status === 'reviewed' ? 0 : 1);
}
async function cmdNotify(args: string[]): Promise<void> {
const number = parseInt(args[0] ?? '', 10);
if (!Number.isFinite(number)) throw new Error('usage: notify <pr-number>');
const cfg = loadConfig();
const bot = new PrBot(cfg, {
telegram: new TelegramClient(cfg.telegramBotToken, cfg.telegramChatId),
codeman: makeCodeman(cfg),
log,
});
const rec = bot.store.pr(number);
if (!rec?.report) throw new Error(`no review of #${number} in ${cfg.dataDir}`);
await bot.sendSummary(rec);
console.log(`sent the review message for #${number}`);
}
async function cmdRun(): Promise<void> {
const cfg = loadConfig();
const bot = new PrBot(cfg, {
telegram: new TelegramClient(cfg.telegramBotToken, cfg.telegramChatId),
codeman: makeCodeman(cfg),
log,
});
let stopping = false;
const shutdown = (signal: string) => {
if (stopping) return;
stopping = true;
log(`${signal}: stopping`);
bot
.stop()
.catch((err) => log(`stop: ${(err as Error).message}`))
.finally(() => process.exit(0));
};
process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));
log(`starting: repo ${cfg.githubRepo}, codeman ${cfg.codemanApiUrl}, data ${cfg.dataDir}`);
await bot.start();
}
function cmdInstallService(): void {
const cfg = loadConfig();
const dir = join(homedir(), '.config', 'systemd', 'user');
mkdirSync(dir, { recursive: true });
const path = join(dir, `${SERVICE_NAME}.service`);
mkdirSync(cfg.dataDir, { recursive: true });
writeFileSync(path, unitFile(cfg));
execFileSync('systemctl', ['--user', 'daemon-reload'], { stdio: 'inherit' });
execFileSync('systemctl', ['--user', 'enable', SERVICE_NAME], { stdio: 'inherit' });
// `restart` rather than `enable --now`: a re-install must pick up the new unit.
execFileSync('systemctl', ['--user', 'restart', SERVICE_NAME], { stdio: 'inherit' });
console.log(`installed ${path}\nlogs: tail -f ${logFilePath(cfg)}`);
}
function cmdUninstallService(): void {
const path = join(homedir(), '.config', 'systemd', 'user', `${SERVICE_NAME}.service`);
execFileSync('systemctl', ['--user', 'disable', '--now', SERVICE_NAME], { stdio: 'inherit' });
if (existsSync(path)) execFileSync('rm', ['-f', path]);
execFileSync('systemctl', ['--user', 'daemon-reload'], { stdio: 'inherit' });
console.log(`removed ${SERVICE_NAME}`);
}
async function main(): Promise<void> {
const [cmd = 'run', ...rest] = process.argv.slice(2);
switch (cmd) {
case 'run':
return cmdRun();
case 'check':
return cmdCheck();
case 'scan':
return cmdScan();
case 'status':
return cmdStatus();
case 'review':
return cmdReview(rest);
case 'notify':
return cmdNotify(rest);
case 'install-service':
return cmdInstallService();
case 'uninstall-service':
return cmdUninstallService();
default:
console.error(
'usage: main.ts run | check | scan | status | review <N> [--no-telegram] | install-service | uninstall-service'
);
process.exit(2);
}
}
main().catch((err) => {
console.error((err as Error).stack ?? String(err));
process.exit(1);
});
+368
View File
@@ -0,0 +1,368 @@
/**
* @fileoverview Pure report handling: parse the reviewer's JSON (leniently, it is
* model output), render the Telegram summary (HTML, under the 4096-char cap), the
* status list, the inline keyboard, and the backlog order. Unit-tested.
*/
import type { CiState, PrSummary } from './github.js';
import { VERDICTS, type Verdict } from './review-task.js';
export type Severity = 'blocker' | 'major' | 'minor' | 'nit';
export interface Finding {
severity: Severity;
title: string;
file?: string;
line?: number;
detail: string;
invariant?: string;
}
export interface CheckResult {
name: string;
command?: string;
result: 'pass' | 'fail' | 'skipped';
notes?: string;
}
export interface ReviewReport {
verdict: Verdict;
confidence: 'high' | 'medium' | 'low';
summary: string;
changes: string[];
findings: Finding[];
checks: CheckResult[];
scope: 'focused' | 'mixed';
risk: string;
recommendation: string;
draftComment: string;
assumptions: string[];
}
export const TELEGRAM_MAX = 4096;
/** Leave room for HTML tags the counter cannot see and for the keyboard-less fallback. */
const SUMMARY_BUDGET = 3600;
const SEVERITY_ORDER: Severity[] = ['blocker', 'major', 'minor', 'nit'];
const SEVERITY_ICON: Record<Severity, string> = { blocker: '🔴', major: '🟠', minor: '🟡', nit: '⚪' };
const VERDICT_LABEL: Record<Verdict, string> = {
merge: '✅ MERGE',
'merge-with-fixes': '🟢 MERGE WITH FIXES',
'request-changes': '🟠 REQUEST CHANGES',
close: '❌ CLOSE',
'needs-discussion': '💬 NEEDS DISCUSSION',
};
const CI_LABEL: Record<CiState, string> = {
passed: 'CI ✅',
failed: 'CI ❌',
pending: 'CI ⏳',
'awaiting-approval': 'CI ⏸ needs your approval',
none: 'CI none',
};
export function escapeHtml(s: string): string {
return s.replace(/&/g, '&amp;').replace(/</g, '&lt;').replace(/>/g, '&gt;');
}
function str(v: unknown, fallback = ''): string {
return typeof v === 'string' ? v : fallback;
}
function strList(v: unknown): string[] {
if (!Array.isArray(v)) return [];
return v.filter((x): x is string => typeof x === 'string' && x.trim().length > 0);
}
/** Extract the first JSON object from text that may carry fences or prose around it. */
export function extractJsonObject(text: string): unknown {
const trimmed = text.trim();
try {
return JSON.parse(trimmed);
} catch {
// fall through
}
const fence = trimmed.match(/```(?:json)?\s*([\s\S]*?)```/);
if (fence) {
try {
return JSON.parse(fence[1]);
} catch {
// fall through
}
}
const start = trimmed.indexOf('{');
const end = trimmed.lastIndexOf('}');
if (start >= 0 && end > start) {
try {
return JSON.parse(trimmed.slice(start, end + 1));
} catch {
return null;
}
}
return null;
}
/** Normalize model output into a ReviewReport. Returns null only when there is no verdict at all. */
export function parseReport(raw: unknown): ReviewReport | null {
if (!raw || typeof raw !== 'object') return null;
const o = raw as Record<string, unknown>;
const verdictRaw = str(o.verdict).trim().toLowerCase().replace(/[_ ]/g, '-');
const verdict = (VERDICTS as readonly string[]).includes(verdictRaw) ? (verdictRaw as Verdict) : null;
if (!verdict) return null;
const confidenceRaw = str(o.confidence).trim().toLowerCase();
const confidence = confidenceRaw === 'high' || confidenceRaw === 'low' ? confidenceRaw : 'medium';
const findings: Finding[] = [];
if (Array.isArray(o.findings)) {
for (const f of o.findings) {
if (!f || typeof f !== 'object') continue;
const fo = f as Record<string, unknown>;
const sevRaw = str(fo.severity).trim().toLowerCase();
const severity = (SEVERITY_ORDER as string[]).includes(sevRaw) ? (sevRaw as Severity) : 'minor';
const title = str(fo.title).trim();
if (!title) continue;
const line = typeof fo.line === 'number' && Number.isFinite(fo.line) ? Math.trunc(fo.line) : undefined;
findings.push({
severity,
title,
file: str(fo.file).trim() || undefined,
line,
detail: str(fo.detail).trim(),
invariant: str(fo.invariant).trim() || undefined,
});
}
}
findings.sort((a, b) => SEVERITY_ORDER.indexOf(a.severity) - SEVERITY_ORDER.indexOf(b.severity));
const checks: CheckResult[] = [];
if (Array.isArray(o.checks)) {
for (const c of o.checks) {
if (!c || typeof c !== 'object') continue;
const co = c as Record<string, unknown>;
const name = str(co.name).trim();
if (!name) continue;
const resRaw = str(co.result).trim().toLowerCase();
const result = resRaw === 'pass' || resRaw === 'fail' ? resRaw : 'skipped';
checks.push({
name,
command: str(co.command).trim() || undefined,
result,
notes: str(co.notes).trim() || undefined,
});
}
}
return {
verdict,
confidence,
summary: str(o.summary).trim(),
changes: strList(o.changes),
findings,
checks,
scope: str(o.scope).trim().toLowerCase() === 'mixed' ? 'mixed' : 'focused',
risk: str(o.risk).trim(),
recommendation: str(o.recommendation).trim(),
draftComment: str(o.draftComment).trim(),
assumptions: strList(o.assumptions),
};
}
export function countBySeverity(findings: Finding[]): Record<Severity, number> {
const out: Record<Severity, number> = { blocker: 0, major: 0, minor: 0, nit: 0 };
for (const f of findings) out[f.severity]++;
return out;
}
function findingLine(f: Finding): string {
const where = f.file ? ` <code>${escapeHtml(f.file)}${f.line ? `:${f.line}` : ''}</code>` : '';
return `${SEVERITY_ICON[f.severity]} ${escapeHtml(f.title)}${where}`;
}
function checksLine(checks: CheckResult[]): string {
if (!checks.length) return '';
const parts = checks.map((c) => {
const icon = c.result === 'pass' ? '✅' : c.result === 'fail' ? '❌' : '⏭';
return `${escapeHtml(c.name)} ${icon}`;
});
return `<b>Checks:</b> ${parts.join(' · ')}`;
}
function truncate(text: string, max: number): string {
if (text.length <= max) return text;
return text.slice(0, Math.max(0, max - 1)).trimEnd() + '…';
}
export interface SummaryMeta {
ci: CiState;
/** Time the review took, for the footer. */
durationMin?: number;
}
/** The message the maintainer reads on the phone. HTML parse mode. */
export function formatTelegramSummary(pr: PrSummary, report: ReviewReport, meta: SummaryMeta): string {
const header =
`🔍 <b>PR #${pr.number}</b> · ${escapeHtml(truncate(pr.title, 120))}\n` +
`<i>by ${escapeHtml(pr.author)} · +${pr.additions}/−${pr.deletions} · ${pr.changedFiles} files · ${CI_LABEL[meta.ci]} · ${
pr.mergeable === 'CONFLICTING'
? 'conflicts ⚠️'
: pr.mergeable === 'MERGEABLE'
? 'mergeable'
: 'mergeability unknown'
}${pr.isDraft ? ' · draft' : ''}</i>\n` +
`<a href="${escapeHtml(pr.url)}">${escapeHtml(pr.url)}</a>\n`;
const verdict = `\n<b>${VERDICT_LABEL[report.verdict]}</b> <i>(confidence ${report.confidence}${report.scope === 'mixed' ? ', mixed scope' : ''})</i>\n`;
const summary = report.summary ? `\n${escapeHtml(report.summary)}\n` : '';
const counts = countBySeverity(report.findings);
const countStr = SEVERITY_ORDER.filter((s) => counts[s] > 0)
.map((s) => `${counts[s]} ${s}${counts[s] === 1 ? '' : 's'}`)
.join(', ');
const findingsHeader = report.findings.length ? `\n<b>Findings</b> (${countStr}):\n` : '\n<b>Findings:</b> none\n';
const checks = checksLine(report.checks);
const recommendation = report.recommendation ? `\n<b>Recommendation:</b> ${escapeHtml(report.recommendation)}\n` : '';
const footer = meta.durationMin !== undefined ? `\n<i>review took ${meta.durationMin} min</i>` : '';
const fixed = header + verdict + summary + findingsHeader;
const tail = (checks ? `\n${checks}\n` : '') + recommendation + footer;
let budget = SUMMARY_BUDGET - fixed.length - tail.length;
const lines: string[] = [];
let shown = 0;
for (const f of report.findings) {
const line = findingLine(f) + '\n';
if (line.length > budget) break;
lines.push(line);
budget -= line.length;
shown++;
}
const hidden = report.findings.length - shown;
const more = hidden > 0 ? `<i>… ${hidden} more in the full report</i>\n` : '';
return fixed + lines.join('') + more + tail;
}
export function formatReviewFailure(
pr: Pick<PrSummary, 'number' | 'title' | 'author' | 'url'>,
reason: string
): string {
return (
`⚠️ <b>PR #${pr.number}</b> · ${escapeHtml(truncate(pr.title, 120))}\n` +
`<i>by ${escapeHtml(pr.author)}</i>\n<a href="${escapeHtml(pr.url)}">${escapeHtml(pr.url)}</a>\n\n` +
`The review did not complete: ${escapeHtml(truncate(reason, 1500))}\n\n` +
`Use /review ${pr.number} to try again.`
);
}
/** Split on line boundaries so no chunk exceeds Telegram's cap. */
export function splitTelegramMessage(text: string, max = TELEGRAM_MAX): string[] {
if (text.length <= max) return [text];
const chunks: string[] = [];
let current = '';
for (const line of text.split('\n')) {
let piece = line;
while (piece.length > max) {
if (current) {
chunks.push(current);
current = '';
}
chunks.push(piece.slice(0, max));
piece = piece.slice(max);
}
const candidate = current ? `${current}\n${piece}` : piece;
if (candidate.length > max) {
chunks.push(current);
current = piece;
} else {
current = candidate;
}
}
if (current) chunks.push(current);
return chunks;
}
export interface InlineButton {
text: string;
callback_data: string;
}
/** Callback data is capped at 64 bytes by Telegram; these stay far under it. */
export function buildReportKeyboard(prNumber: number, opts: { ci: CiState; hasDraft: boolean }): InlineButton[][] {
const rows: InlineButton[][] = [
[
{ text: '📄 Full report', callback_data: `report:${prNumber}` },
...(opts.hasDraft ? [{ text: '💬 Draft comment', callback_data: `draft:${prNumber}` }] : []),
{ text: '🔁 Re-review', callback_data: `review:${prNumber}` },
],
[
{ text: '✅ Merge', callback_data: `merge:${prNumber}` },
...(opts.hasDraft ? [{ text: '📮 Post comment', callback_data: `post:${prNumber}` }] : []),
{ text: '🗑 Close', callback_data: `close:${prNumber}` },
],
];
if (opts.ci === 'awaiting-approval')
rows.push([{ text: '▶️ Approve CI run', callback_data: `approveci:${prNumber}` }]);
return rows;
}
export function confirmKeyboard(action: string, prNumber: number, nonce: string): InlineButton[][] {
return [
[
{ text: `Yes, ${action} #${prNumber}`, callback_data: `confirm:${action}:${prNumber}:${nonce}` },
{ text: 'Cancel', callback_data: `cancel:${action}:${prNumber}:${nonce}` },
],
];
}
export interface StatusRow {
number: number;
title: string;
author: string;
verdict?: Verdict;
status: string;
ci?: CiState;
mergeable: PrSummary['mergeable'];
isDraft: boolean;
}
export function formatStatusList(rows: StatusRow[], paused: boolean): string {
if (!rows.length) return 'No open pull requests.';
const lines = rows.map((r) => {
const v = r.verdict
? VERDICT_LABEL[r.verdict].split(' ')[0]
: r.status === 'reviewing'
? '⏳'
: r.status === 'queued'
? '🕓'
: '·';
const flags = [
r.ci ? CI_LABEL[r.ci].replace('CI ', '') : '',
r.mergeable === 'CONFLICTING' ? 'conflicts' : '',
r.isDraft ? 'draft' : '',
]
.filter(Boolean)
.join(', ');
return `${v} <b>#${r.number}</b> ${escapeHtml(truncate(r.title, 60))} <i>(${escapeHtml(r.author)}${flags ? `; ${flags}` : ''})</i>`;
});
return `${paused ? '⏸ auto-review paused\n' : ''}<b>Open PRs (${rows.length})</b>\n${lines.join('\n')}`;
}
/**
* Backlog order for a fresh sweep: the ones you can act on first (mergeable, small),
* conflicting and huge ones last. Ties keep the newer PR first.
*/
export function orderBacklog<T extends Pick<PrSummary, 'number' | 'mergeable' | 'additions' | 'deletions'>>(
prs: T[]
): T[] {
const size = (p: T) => p.additions + p.deletions;
return [...prs].sort((a, b) => {
const ca = a.mergeable === 'CONFLICTING' ? 1 : 0;
const cb = b.mergeable === 'CONFLICTING' ? 1 : 0;
if (ca !== cb) return ca - cb;
const sa = size(a);
const sb = size(b);
if (sa !== sb) return sa - sb;
return b.number - a.number;
});
}
export function verdictLabel(v: Verdict): string {
return VERDICT_LABEL[v];
}
+241
View File
@@ -0,0 +1,241 @@
/**
* @fileoverview The review brief handed to each reviewer session, and the follow-up
* brief. Pure: the bot writes the result to a file and sends the session one short
* line pointing at it (prompts are single-line over tmux, and a brief this size
* belongs on disk anyway).
*
* The brief is opinionated on purpose. It names the repository's own rules (CLAUDE.md,
* CONTRIBUTING.md), the checks to run, the verdict vocabulary, and the exact JSON the
* bot parses. Everything the maintainer would say out loud before delegating a
* review lives here.
*/
import type { CiStatus, PrDetail } from './github.js';
export const VERDICTS = ['merge', 'merge-with-fixes', 'request-changes', 'close', 'needs-discussion'] as const;
export type Verdict = (typeof VERDICTS)[number];
export interface ReviewBriefInput {
pr: PrDetail;
ci: CiStatus;
mergeBase: string;
worktreeDir: string;
mainCheckout: string;
reportJsonPath: string;
reportMdPath: string;
}
function ciLine(ci: CiStatus): string {
const detail = ci.runs.map((r) => `${r.name}: ${r.conclusion ?? r.status}`).join(', ');
switch (ci.state) {
case 'passed':
return `passed (${detail})`;
case 'failed':
return `FAILED (${detail}); read the failing job's log with \`gh run view <id> --log-failed\` before you trust or dismiss it`;
case 'pending':
return `still running (${detail})`;
case 'awaiting-approval':
return 'never ran: the workflow is waiting for a maintainer to approve it (first-time contributor), so run the checks yourself';
default:
return 'no workflow runs found for this head (a conflicting PR gets no CI at all); run the checks yourself';
}
}
export function buildReviewBrief(input: ReviewBriefInput): string {
const { pr, ci, mergeBase, worktreeDir, mainCheckout, reportJsonPath, reportMdPath } = input;
const files = pr.files.map((f) => `- \`${f.path}\` (+${f.additions}/-${f.deletions})`).join('\n');
const linked = pr.linkedIssues.length
? pr.linkedIssues.map((i) => `- #${i.number} ${i.title}`).join('\n')
: '- none linked';
const mergeability =
pr.mergeable === 'CONFLICTING'
? 'CONFLICTING with master. It cannot be merged as-is and GitHub runs no CI for it. Review the PR head as it stands, and say in the report whether the conflicts look mechanical or structural (`git merge-tree` against origin/master helps).'
: pr.mergeable === 'MERGEABLE'
? 'mergeable'
: 'unknown (GitHub has not computed it yet)';
return `# Review brief: PR #${pr.number} ${pr.title}
You are reviewing a pull request against Codeman on behalf of the maintainer. You are
in a private clone at \`${worktreeDir}\`, checked out (detached) at the PR head. The
maintainer reads your report on a phone and decides what happens next, so write for
someone who has not seen the diff.
## Ground rules (read twice)
- Nothing you do here reaches GitHub. Do NOT push, comment, merge, close, label, or
create anything with \`gh\`; \`gh\` is for READING only (\`gh pr view\`, \`gh run view\`,
\`gh api\` GETs).
- Do NOT run \`npm install\`, \`npm ci\`, \`npm update\` or \`npm run build\`: \`node_modules\`
may be a symlink into the maintainer's live checkout. Everything else in package.json
scripts is fine (\`npm run typecheck\`, \`npm run lint\`, \`npm test -- <file>\`, ...).
- Do NOT restart, stop or install any service, and never bind port 3000: the
maintainer's production Codeman runs there. Test ports are 3150 and up.
- \`${mainCheckout}\` is the maintainer's shared checkout. You may READ it for comparison;
never run a git command there that changes anything (no checkout, reset, stash, clean).
- Stay inside this clone for writes. Do not create files elsewhere except the two
report files named below.
- Do not ask questions. Nobody is watching this session. Where something is ambiguous,
decide, and list the assumption in the report.
## The pull request
- **#${pr.number}** ${pr.title}
- Author: ${pr.author} (${pr.authorAssociation.toLowerCase().replace(/_/g, ' ')})${pr.headRepo ? `, from \`${pr.headRepo}\`` : ''}
- URL: ${pr.url}
- Base: \`${pr.baseRef}\` at merge base \`${mergeBase.slice(0, 12)}\`; head: \`${pr.headSha.slice(0, 12)}\` (${pr.commitCount} commits)
- Size: +${pr.additions} / -${pr.deletions} across ${pr.changedFiles} files
- Mergeability: ${mergeability}
- CI: ${ciLine(ci)}
- Draft: ${pr.isDraft ? 'yes' : 'no'}; existing comments: ${pr.commentCount}${pr.labels.length ? `; labels: ${pr.labels.join(', ')}` : ''}
### Linked issues
${linked}
### Files changed
${files || '- (none reported)'}
### PR description, verbatim
\`\`\`text
${pr.body.trim() || '(empty)'}
\`\`\`
## How to review
1. Read \`CLAUDE.md\` at the root and \`.github/CONTRIBUTING.md\`. Most review feedback on
this repository traces back to a rule already written there, and a change that
contradicts one of those rules is a finding even when the code works. Open the
\`docs/architecture-invariants.md\` sections the change touches.
2. Understand the change: \`git log --oneline ${mergeBase.slice(0, 12)}..HEAD\` and
\`git diff ${mergeBase.slice(0, 12)}..HEAD\`. Read the surrounding code, not only the
hunks: the file's \`@fileoverview\` first, then the call sites of anything changed.
3. Look for, in this order: correctness bugs (wrong logic, races, missed error paths,
lost state across restart); security (auth and ownership checks, path confinement,
the env-prefix allowlist, shell/command injection, SSRF, secrets on the command
line or in state files); violations of CLAUDE.md rules (cite the rule); behaviour
changes without tests; contract changes (\`/api/v1\` paths, response envelope,
\`errorCode\` values, SSE event names are public and stable, see
\`docs/versioning-policy.md\`); scope (one change per PR: flag unrelated changes
bundled in); docs and registries that must move with the code (CLAUDE.md and
architecture-invariants when a rule changes, \`sse-events.ts\` and \`constants.js\`
parity, \`docs/api-reference.md\`); housekeeping that does not belong in a PR
(version bumps, CHANGELOG edits, files pulled back into Prettier's scope, committed
vendor bundles, changeset files are fine).
4. Run the checks and record what you ran and what came back:
\`npm run typecheck\`, \`npm run lint\`, \`npm run check:frontend-syntax\`,
\`npm run format:check\`, then the tests covering the touched areas
(\`npm test -- test/<file>.test.ts\`, several files at once is fine). Run the full
\`npm test\` when the change is broad or touches shared infrastructure (session,
tmux, routes, state); it takes minutes, which is acceptable. A red check that is
also red on origin/master is not the PR's fault: say so rather than blaming it.
Other test suites may be running on this machine at the same time and they share
the 3150+ port range, so re-run a failed file on its own (\`npm test -- <file>\`)
before you read an EADDRINUSE or a timeout as the PR's regression.
5. Verify before you report. A finding that could be a misread must be confirmed by
reading the full code path, by a tiny test, or by running it. Every finding names a
file and line. Rank: **blocker** (must be fixed before merge: data loss, security,
breaks a documented invariant, breaks the build or tests), **major** (should be
fixed: a real bug in an edge the PR introduces, a missing test for new behaviour),
**minor**, **nit**.
6. Judge the PR, not the author. Contributors here are volunteers and the maintainer
thanks them by name in every release; be exact and be kind.
## Verdict vocabulary
- \`merge\`: no blockers or majors, checks green; merge as-is.
- \`merge-with-fixes\`: mergeable, but with small things the maintainer would rather fix
at merge time than round-trip (list them so they can be applied on top).
- \`request-changes\`: blockers or majors the author should fix.
- \`close\`: wrong direction, superseded, or not wanted; say what should happen instead.
- \`needs-discussion\`: a design question the maintainer must answer before anyone
spends more time (name the question).
## Output, mandatory
Write BOTH files, then reply with exactly one line: \`REVIEW COMPLETE\`.
1. \`${reportJsonPath}\`: a single JSON object, no markdown fences, this shape:
\`\`\`json
{
"verdict": "merge | merge-with-fixes | request-changes | close | needs-discussion",
"confidence": "high | medium | low",
"summary": "Two or three sentences: what the PR does, and the review's bottom line.",
"changes": ["one bullet per thing the PR actually changes"],
"findings": [
{
"severity": "blocker | major | minor | nit",
"title": "one line",
"file": "path/from/repo/root.ts",
"line": 123,
"detail": "what is wrong, why it matters, what to do instead",
"invariant": "the CLAUDE.md / CONTRIBUTING rule it breaks, or omit"
}
],
"checks": [
{ "name": "typecheck", "command": "npm run typecheck", "result": "pass | fail | skipped", "notes": "" }
],
"_checks_note": "result is from the PR's point of view: a regression test you deliberately ran against master to prove it fails is a pass (say so in notes), a red run caused by another suite on the machine is skipped with the reason, only a genuine problem with the PR is fail",
"scope": "focused | mixed",
"risk": "One or two sentences naming the judgment calls a second reviewer should look at.",
"recommendation": "Two to four sentences for the maintainer: what to do next and why.",
"draftComment": "A comment to the contributor, in markdown, ready to post (rules below).",
"assumptions": ["anything you had to decide alone"]
}
\`\`\`
2. \`${reportMdPath}\`: the full report in markdown for the maintainer, in this order:
what the PR does; the verdict with the reasoning; findings in severity order with
file:line and the fix; checks run with results; CLAUDE.md rules touched; scope and
risk; recommendation; assumptions. Include the diff stat. No length limit, but no
padding either.
### Draft comment rules
The draft is written AS the maintainer TO the contributor and must stand alone: the
reader has not seen this brief. Open by thanking them and saying in one sentence what
the PR does. Then the findings that need action, each with file:line and the concrete
ask, blockers first. Close with what happens next (merge after fixes, will fix at merge
time, and so on). When the verdict is \`merge\`, the whole comment is a short thank-you
naming anything you would touch at merge time. Plain markdown. No em-dashes (use
commas, colons or parentheses). No emojis. No "Generated with Claude Code" or similar
attribution line. No hedging words. The maintainer reads it before it is posted and may
edit it.
`;
}
/** Sent as ONE line; the brief above is on disk. */
export function reviewKickoffLine(briefPath: string): string {
return `Read ${briefPath} and carry out the review it describes. Do not ask questions. Finish by writing both report files it names, then reply with exactly: REVIEW COMPLETE`;
}
export function followupKickoffLine(followupPath: string): string {
return `Read ${followupPath}: it holds a follow-up from the maintainer about the pull request you reviewed. Do what it asks within the ground rules of the original brief (no pushing, no gh writes, no npm install, no builds, no services), then answer in plain text. Do not ask questions.`;
}
export function buildFollowupBrief(input: {
prNumber: number;
title: string;
instruction: string;
worktreeDir: string;
reportMdPath: string;
briefPath: string;
}): string {
return `# Follow-up on PR #${input.prNumber} ${input.title}
The maintainer read your review report (\`${input.reportMdPath}\`; the original brief is
\`${input.briefPath}\`, and its ground rules still apply: nothing reaches GitHub, no
installs, no builds, no services, writes stay inside \`${input.worktreeDir}\`).
Their message:
\`\`\`text
${input.instruction.trim()}
\`\`\`
Answer concisely and concretely, for a phone screen: lead with the answer, then the
evidence (commands run, file:line). If the message asks you to change code, make the
change in this clone, run the relevant checks, and describe the diff (\`git diff
--stat\` plus the essential hunks). Keep the changes uncommitted unless asked to commit;
never push. If it asks for something outside the ground rules, say so and stop.
`;
}
+166
View File
@@ -0,0 +1,166 @@
/**
* @fileoverview The bot's persisted state: one record per PR (what was reviewed at
* which head, the parsed report, the Claude session to resume for follow-ups, the
* Telegram messages that belong to it), the Telegram update offset, pending
* confirmations, and the pause flag. One JSON file, written atomically (tmp + rename)
* with mode 0600, since reports quote code and draft comments.
*/
import { existsSync, mkdirSync, readFileSync, renameSync, writeFileSync } from 'fs';
import { dirname, join } from 'path';
import type { CiState, PrSummary } from './github.js';
import type { ReviewReport } from './report.js';
import type { Verdict } from './review-task.js';
export type PrStatus = 'new' | 'queued' | 'reviewing' | 'reviewed' | 'failed' | 'skipped' | 'closed';
export interface PrRecord {
number: number;
title: string;
author: string;
url: string;
headSha: string;
isDraft: boolean;
mergeable: PrSummary['mergeable'];
additions?: number;
deletions?: number;
changedFiles?: number;
status: PrStatus;
ci?: CiState;
reviewedSha?: string;
reviewedAt?: string;
reviewDurationMin?: number;
verdict?: Verdict;
report?: ReviewReport;
briefPath?: string;
reportJsonPath?: string;
reportMdPath?: string;
/** The Claude conversation to resume for follow-ups. */
claudeSessionId?: string;
/** The live Codeman session while a turn is running; cleared afterwards. */
activeSessionId?: string;
worktreeDir?: string;
telegramMessageId?: number;
lastError?: string;
/** Consecutive failed attempts at `failedSha`; the scan stops auto-retrying at MAX_AUTO_RETRIES. */
failedAttempts?: number;
failedSha?: string;
closedAs?: 'merged' | 'closed';
updatedAt: string;
}
export interface PendingConfirm {
action: 'merge' | 'close' | 'post';
prNumber: number;
createdAt: string;
messageId?: number;
/** Closing comment for `close`. */
reason?: string;
}
export interface BotState {
version: 1;
paused: boolean;
telegramOffset: number;
prs: Record<string, PrRecord>;
pending: Record<string, PendingConfirm>;
/** Telegram message id -> PR number, so a reply to any of the bot's messages finds its PR. */
messages: Record<string, number>;
/** Telegram message id -> PR number for "reply with the closing reason" prompts. */
reasonPrompts: Record<string, number>;
}
export function emptyState(): BotState {
return { version: 1, paused: false, telegramOffset: 0, prs: {}, pending: {}, messages: {}, reasonPrompts: {} };
}
const MAX_MESSAGE_MAP = 2000;
export class StateStore {
state: BotState;
constructor(private readonly path: string) {
this.state = emptyState();
if (existsSync(path)) {
try {
const parsed = JSON.parse(readFileSync(path, 'utf8')) as Partial<BotState>;
this.state = { ...emptyState(), ...parsed, version: 1 };
} catch (err) {
throw new Error(`state file ${path} is unreadable: ${(err as Error).message}`);
}
}
}
save(): void {
mkdirSync(dirname(this.path), { recursive: true });
this.pruneMessageMap();
const tmp = join(dirname(this.path), `.state.${process.pid}.${Date.now()}.tmp`);
writeFileSync(tmp, JSON.stringify(this.state, null, 2), { mode: 0o600 });
renameSync(tmp, this.path);
}
pr(number: number): PrRecord | undefined {
return this.state.prs[String(number)];
}
/**
* Refresh a PR's metadata, keeping its review. Mutates the EXISTING record in place:
* a review in flight holds a reference to it, and a scan that replaced the object
* with a copy made that review write its verdict into an orphan (first daemon run:
* PR 363 reported to Telegram, state still said `reviewing`).
*/
upsertPr(summary: PrSummary): PrRecord {
const key = String(summary.number);
const existing = this.state.prs[key];
const record: PrRecord = existing ?? {
number: summary.number,
title: summary.title,
author: summary.author,
url: summary.url,
headSha: summary.headSha,
isDraft: summary.isDraft,
mergeable: summary.mergeable,
status: 'new',
updatedAt: new Date().toISOString(),
};
record.title = summary.title;
record.author = summary.author;
record.url = summary.url;
record.headSha = summary.headSha;
record.isDraft = summary.isDraft;
record.mergeable = summary.mergeable;
record.additions = summary.additions;
record.deletions = summary.deletions;
record.changedFiles = summary.changedFiles;
if (record.status === 'closed') {
// Reopened.
record.status = record.reviewedSha ? 'reviewed' : 'new';
record.closedAs = undefined;
}
record.updatedAt = new Date().toISOString();
this.state.prs[key] = record;
return record;
}
openPrs(): PrRecord[] {
return Object.values(this.state.prs)
.filter((r) => r.status !== 'closed')
.sort((a, b) => b.number - a.number);
}
rememberMessage(messageId: number, prNumber: number): void {
this.state.messages[String(messageId)] = prNumber;
}
prForMessage(messageId: number | undefined): number | undefined {
if (messageId === undefined) return undefined;
return this.state.messages[String(messageId)];
}
private pruneMessageMap(): void {
const keys = Object.keys(this.state.messages);
if (keys.length <= MAX_MESSAGE_MAP) return;
// Message ids grow monotonically per chat; drop the oldest.
keys.sort((a, b) => Number(a) - Number(b));
for (const key of keys.slice(0, keys.length - MAX_MESSAGE_MAP)) delete this.state.messages[key];
}
}
+188
View File
@@ -0,0 +1,188 @@
/**
* @fileoverview Minimal Telegram Bot API client (long polling, no webhook: the box sits
* behind Tailscale) plus the pure command / callback parsers.
*
* Only updates from the configured chat are ever acted on; everything else is dropped
* without an answer, so a stranger who finds the bot gets silence, not a menu.
*/
export interface TelegramMessage {
message_id: number;
chat: { id: number | string };
from?: { id: number; username?: string };
text?: string;
reply_to_message?: { message_id: number; text?: string };
}
export interface TelegramCallbackQuery {
id: string;
from: { id: number; username?: string };
message?: TelegramMessage;
data?: string;
}
export interface TelegramUpdate {
update_id: number;
message?: TelegramMessage;
callback_query?: TelegramCallbackQuery;
}
export interface SendOptions {
replyMarkup?: unknown;
replyToMessageId?: number;
disablePreview?: boolean;
}
export class TelegramClient {
private readonly base: string;
constructor(
token: string,
private readonly chatId: string
) {
this.base = `https://api.telegram.org/bot${token}`;
}
private async call<T>(method: string, body?: Record<string, unknown>, timeoutMs = 30_000): Promise<T> {
const res = await fetch(`${this.base}/${method}`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(body ?? {}),
signal: AbortSignal.timeout(timeoutMs),
});
const json = (await res.json()) as { ok: boolean; result?: T; description?: string };
if (!json.ok) throw new Error(`telegram ${method}: ${json.description ?? res.status}`);
return json.result as T;
}
isOurChat(chatId: number | string | undefined): boolean {
return chatId !== undefined && String(chatId) === this.chatId;
}
async getMe(): Promise<{ username?: string }> {
return this.call<{ username?: string }>('getMe');
}
async sendMessage(text: string, opts: SendOptions = {}): Promise<number> {
const result = await this.call<{ message_id: number }>('sendMessage', {
chat_id: this.chatId,
text,
parse_mode: 'HTML',
disable_web_page_preview: opts.disablePreview ?? true,
reply_markup: opts.replyMarkup,
reply_to_message_id: opts.replyToMessageId,
});
return result.message_id;
}
/** Plain text, no parse mode: for content the bot did not write (reviewer answers, drafts). */
async sendPlain(text: string, opts: SendOptions = {}): Promise<number> {
const result = await this.call<{ message_id: number }>('sendMessage', {
chat_id: this.chatId,
text,
disable_web_page_preview: opts.disablePreview ?? true,
reply_markup: opts.replyMarkup,
reply_to_message_id: opts.replyToMessageId,
});
return result.message_id;
}
async editReplyMarkup(messageId: number, replyMarkup: unknown): Promise<void> {
try {
await this.call('editMessageReplyMarkup', {
chat_id: this.chatId,
message_id: messageId,
reply_markup: replyMarkup,
});
} catch (err) {
// "message is not modified" is Telegram's way of saying the keyboard already looks like that.
if (!String(err).includes('not modified')) throw err;
}
}
async deleteMessage(messageId: number): Promise<void> {
try {
await this.call('deleteMessage', { chat_id: this.chatId, message_id: messageId });
} catch {
// Already gone, or older than Telegram allows a bot to delete; the message was informational.
}
}
async answerCallback(callbackId: string, text?: string): Promise<void> {
await this.call('answerCallbackQuery', { callback_query_id: callbackId, text });
}
async sendDocument(filename: string, content: string, caption?: string): Promise<void> {
const form = new FormData();
form.set('chat_id', this.chatId);
if (caption) form.set('caption', caption);
form.set('document', new Blob([content], { type: 'text/markdown' }), filename);
const res = await fetch(`${this.base}/sendDocument`, {
method: 'POST',
body: form,
signal: AbortSignal.timeout(60_000),
});
const json = (await res.json()) as { ok: boolean; description?: string };
if (!json.ok) throw new Error(`telegram sendDocument: ${json.description ?? res.status}`);
}
async getUpdates(offset: number, timeoutSec: number): Promise<TelegramUpdate[]> {
return this.call<TelegramUpdate[]>(
'getUpdates',
{ offset, timeout: timeoutSec, allowed_updates: ['message', 'callback_query'] },
(timeoutSec + 15) * 1000
);
}
async setMyCommands(commands: { command: string; description: string }[]): Promise<void> {
await this.call('setMyCommands', { commands });
}
}
export interface ParsedCommand {
command: string;
prNumber?: number;
rest: string;
}
/** `/merge 381 force` -> {command:'merge', prNumber:381, rest:'force'}; `/help@botname` is handled. */
export function parseCommand(text: string | undefined): ParsedCommand | null {
if (!text) return null;
const m = text.trim().match(/^\/([a-zA-Z_]+)(?:@\w+)?(?:\s+([\s\S]*))?$/);
if (!m) return null;
const command = m[1].toLowerCase();
const argText = (m[2] ?? '').trim();
const numMatch = argText.match(/^#?(\d+)\b\s*([\s\S]*)$/);
if (numMatch) return { command, prNumber: parseInt(numMatch[1], 10), rest: numMatch[2].trim() };
return { command, rest: argText };
}
export interface ParsedCallback {
action: string;
prNumber: number;
nonce?: string;
/** For confirm/cancel: the action being confirmed. */
target?: string;
}
export function parseCallback(data: string | undefined): ParsedCallback | null {
if (!data) return null;
const parts = data.split(':');
if (parts[0] === 'confirm' || parts[0] === 'cancel') {
if (parts.length !== 4) return null;
const prNumber = parseInt(parts[2], 10);
if (!Number.isFinite(prNumber)) return null;
return { action: parts[0], target: parts[1], prNumber, nonce: parts[3] };
}
if (parts.length !== 2) return null;
const prNumber = parseInt(parts[1], 10);
if (!Number.isFinite(prNumber)) return null;
return { action: parts[0], prNumber };
}
/** Find the PR number a report message is about, from its first line (`🔍 PR #381 · ...`). */
export function prNumberFromMessageText(text: string | undefined): number | null {
if (!text) return null;
const m = text.match(/PR #(\d+)/);
return m ? parseInt(m[1], 10) : null;
}
+253
View File
@@ -0,0 +1,253 @@
/**
* @fileoverview Per-PR checkouts for the review sessions.
*
* The maintainer's checkout is SHARED with other agent sessions (CLAUDE.md, Session
* Safety), so the bot never runs `git checkout` there. It fetches the PR head into a
* private ref (`refs/pr-bot/<n>`) of the main repository, which anchors the objects,
* and checks the PR out in a private clone under the bot's own data dir; every
* in-tree git command runs with `-C <clone>`.
*
* Why a `git clone --shared` and not a linked worktree: Claude Code resolves a linked
* worktree's project settings through the git common dir, i.e. the MAIN checkout's
* `.claude/settings.local.json`, whose model pin then silently overrides anything
* written into the worktree (measured 2026-09-05: a worktree pinned to
* `claude-fable-5-1` reported `claude-opus-5[1m]`, the main checkout's pin). A shared clone has its own
* project root, so Codeman's `modelOverride` and hooks land where the CLI reads them,
* while `objects/info/alternates` keeps the object store shared (no duplication).
*
* Dependencies: a clone has no `node_modules`. When the PR leaves the lockfile
* untouched, `node_modules` is a SYMLINK to the main checkout's tree (read-only use:
* tsc, vitest, eslint). When the PR changes dependencies, the symlink is unlinked
* first and `npm ci` installs a real tree, so npm can never write through the link
* into the live server's modules. `src/web/public/vendor` is COPIED per file, never
* linked: postinstall regenerates it in place, and a link would let a PR's bundle
* overwrite the bundle the production server is serving.
*/
import { execFile } from 'child_process';
import {
cpSync,
existsSync,
lstatSync,
mkdirSync,
readdirSync,
rmSync,
statSync,
symlinkSync,
unlinkSync,
writeFileSync,
} from 'fs';
import { join } from 'path';
import { promisify } from 'util';
const execFileAsync = promisify(execFile);
export interface WorktreeInfo {
dir: string;
headSha: string;
mergeBase: string;
deps: 'linked' | 'installed' | 'kept';
}
export type Logger = (msg: string) => void;
async function git(args: string[], cwd: string, timeoutMs = 120_000): Promise<string> {
const { stdout } = await execFileAsync('git', args, { cwd, maxBuffer: 64 * 1024 * 1024, timeout: timeoutMs });
return stdout;
}
export function prRef(prNumber: number): string {
return `refs/pr-bot/${prNumber}`;
}
/** The upstream master, as fetched into the main repository, mirrored into the clone. */
const MASTER_REF = 'refs/remotes/origin/master';
export function worktreeDirFor(worktreesDir: string, prNumber: number): string {
return join(worktreesDir, `pr-${prNumber}`);
}
const DEP_FILES = [
'package.json',
'package-lock.json',
'packages/xterm-zerolag-input/package.json',
'packages/gesture-control/package.json',
];
async function originUrl(mainCheckout: string): Promise<string> {
return (await git(['remote', 'get-url', 'origin'], mainCheckout)).trim();
}
/** A linked worktree from the first version of this file: `.git` is a FILE there. */
function isLegacyWorktree(dir: string): boolean {
const dotGit = join(dir, '.git');
try {
return statSync(dotGit).isFile();
} catch {
return false;
}
}
function isOwnClone(dir: string): boolean {
try {
return statSync(join(dir, '.git')).isDirectory();
} catch {
return false;
}
}
/** Fetch the PR head, (re)create the clone at it, and make node_modules usable. */
export async function preparePrWorktree(opts: {
mainCheckout: string;
worktreesDir: string;
prNumber: number;
/** Reset a reused clone to the fetched head (drops edits a follow-up may have made). */
reset: boolean;
log: Logger;
}): Promise<WorktreeInfo> {
const { mainCheckout, worktreesDir, prNumber, log } = opts;
const ref = prRef(prNumber);
const dir = worktreeDirFor(worktreesDir, prNumber);
mkdirSync(worktreesDir, { recursive: true });
log(`fetching origin master + pull/${prNumber}/head`);
await git(
['fetch', '--quiet', 'origin', `+refs/heads/master:${MASTER_REF}`, `+refs/pull/${prNumber}/head:${ref}`],
mainCheckout,
300_000
);
const headSha = (await git(['rev-parse', ref], mainCheckout)).trim();
if (existsSync(dir) && isLegacyWorktree(dir)) {
log(`replacing the linked worktree at ${dir} with a clone`);
await git(['worktree', 'remove', '--force', dir], mainCheckout).catch(() =>
rmSync(dir, { recursive: true, force: true })
);
await git(['worktree', 'prune'], mainCheckout);
}
if (existsSync(dir) && !isOwnClone(dir)) {
log(`removing stale directory ${dir}`);
rmSync(dir, { recursive: true, force: true });
}
if (!existsSync(dir)) {
log(`cloning (shared objects) into ${dir}`);
await git(['clone', '--quiet', '--shared', '--no-checkout', mainCheckout, dir], mainCheckout, 300_000);
// `origin` of the clone should mean GitHub, like everywhere else, not the main
// checkout's path; the refs below are fetched from the main checkout by path.
await git(['remote', 'set-url', 'origin', await originUrl(mainCheckout)], dir);
}
// Mirror the two refs from the main repository (objects are already reachable via
// alternates, so this only moves refs). `+` because both can move backwards.
await git(['fetch', '--quiet', mainCheckout, `+${MASTER_REF}:${MASTER_REF}`, `+${ref}:${ref}`], dir);
const current = (await git(['rev-parse', '--verify', '--quiet', 'HEAD'], dir).catch(() => '')).trim();
if (current !== headSha) {
log(`checking out ${headSha.slice(0, 8)}${current ? ` (was ${current.slice(0, 8)})` : ''}`);
await git(['checkout', '--quiet', '--detach', ref], dir);
}
if (opts.reset) {
await git(['reset', '--hard', '--quiet', ref], dir);
}
const mergeBase = (await git(['merge-base', MASTER_REF, 'HEAD'], dir)).trim();
const deps = await ensureDependencies({ mainCheckout, dir, ref, mergeBase, log });
ensureVendorCopy(mainCheckout, dir, log);
return { dir, headSha, mergeBase, deps };
}
/** Written into a clone's own node_modules once `npm ci` has finished; its absence means a half install. */
const INSTALL_MARKER = '.pr-bot-installed';
async function ensureDependencies(opts: {
mainCheckout: string;
dir: string;
ref: string;
mergeBase: string;
log: Logger;
}): Promise<WorktreeInfo['deps']> {
const { mainCheckout, dir, ref, mergeBase, log } = opts;
const target = join(dir, 'node_modules');
// Against the MERGE BASE, not master: master's own version bumps since the PR
// branched would otherwise make every older PR look like a dependency change and
// cost a full npm ci each. Only what the PR itself did to the dependency files counts.
let depsChanged = false;
try {
await git(['diff', '--quiet', mergeBase, ref, '--', ...DEP_FILES], mainCheckout);
} catch {
depsChanged = true;
}
let existing = existsSync(target) || isSymlink(target) ? lstatSync(target) : null;
if (existing?.isDirectory() && !existsSync(join(target, INSTALL_MARKER))) {
// A real tree without the marker is an install that was interrupted (service
// restart mid `npm ci`); never trust it.
log('discarding an incomplete node_modules install');
rmSync(target, { recursive: true, force: true });
existing = null;
}
if (!depsChanged) {
if (existing?.isSymbolicLink()) return 'linked';
if (existing?.isDirectory()) return 'kept';
symlinkSync(join(mainCheckout, 'node_modules'), target, 'dir');
log('node_modules linked to the main checkout (dependencies unchanged by the PR)');
return 'linked';
}
// The PR changes dependencies: a real install, and NEVER through the symlink.
if (existing?.isSymbolicLink()) unlinkSync(target);
if (existing?.isDirectory()) return 'kept';
log('the PR changes dependencies: running npm ci in the clone (this can take minutes)');
await execFileAsync('npm', ['ci', '--no-audit', '--no-fund', '--loglevel=error'], {
cwd: dir,
timeout: 20 * 60_000,
maxBuffer: 64 * 1024 * 1024,
});
writeFileSync(join(target, INSTALL_MARKER), new Date().toISOString());
return 'installed';
}
function isSymlink(path: string): boolean {
try {
return lstatSync(path).isSymbolicLink();
} catch {
return false;
}
}
function ensureVendorCopy(mainCheckout: string, dir: string, log: Logger): void {
const rel = join('src', 'web', 'public', 'vendor');
const src = join(mainCheckout, rel);
const dst = join(dir, rel);
if (!existsSync(src)) return;
// Two of the vendor files are tracked in git, so the directory already exists in a
// fresh checkout; copy whatever is MISSING (the postinstall-built xterm bundles).
mkdirSync(dst, { recursive: true });
let copied = 0;
for (const entry of readdirSync(src)) {
const target = join(dst, entry);
if (existsSync(target)) continue;
cpSync(join(src, entry), target, { recursive: true });
copied++;
}
if (copied) log(`${copied} vendor bundle(s) copied from the main checkout`);
}
export async function removePrWorktree(opts: {
mainCheckout: string;
worktreesDir: string;
prNumber: number;
log: Logger;
}): Promise<void> {
const dir = worktreeDirFor(opts.worktreesDir, opts.prNumber);
if (existsSync(dir)) {
opts.log(`removing ${dir}`);
if (isLegacyWorktree(dir)) {
await git(['worktree', 'remove', '--force', dir], opts.mainCheckout).catch(() => undefined);
await git(['worktree', 'prune'], opts.mainCheckout).catch(() => undefined);
}
rmSync(dir, { recursive: true, force: true });
}
try {
await git(['update-ref', '-d', prRef(opts.prNumber)], opts.mainCheckout);
} catch {
// The ref may never have been created; nothing to delete.
}
}
+55 -7
View File
@@ -7,18 +7,25 @@
# the repo (the server stages it at ~/.codeman/self-update-runner.sh) — `git
# checkout` rewrites the in-repo copy and bash reads scripts lazily.
#
# ⚠️ The `docker-compose` supervisor is the exception to "outlives": there the
# restart IS the container exiting, which kills this script too. That is safe
# because the terminal "restarting" marker is written before the kill and the
# rebooted server reconciles it — but nothing may be added after that kill.
#
# Reports progress by writing ~/.codeman/update-status.json atomically; the
# browser polls GET /api/system/update/status across the restart drop. The
# freshly-booted server reconciles the final "restarting" → "completed"/"failed".
#
# Cross-platform: restarts via systemd (Linux), launchd (macOS), or prints a
# manual command (foreground installs). Linux launches inside a transient
# systemd scope so `systemctl restart codeman-web` can't kill it mid-build.
# Cross-platform: restarts via systemd (Linux), launchd (macOS), a container exit
# under Docker Compose (the restart policy relaunches it), or prints a manual
# command (foreground installs). Linux launches inside a transient systemd scope
# so `systemctl restart codeman-web` can't kill it mid-build.
#
# Args (all from the server, never user input — tag is validated server-side):
# --repo <dir> --tag <codeman@X.Y.Z> --supervisor <systemd|launchd|none>
# --repo <dir> --tag <codeman@X.Y.Z> --supervisor <systemd|launchd|docker-compose|none>
# --status-file <path> --update-id <uuid> --from-version <ver> --node <path>
# --log <path> [--prev-sha <sha>] [--stash]
# --log <path> [--prev-sha <sha>] [--stash] [--server-pid <pid>]
# [--restart-by-exit 0|1] (docker-compose only: may we exit the server?)
#
set -uo pipefail
@@ -32,6 +39,7 @@ REPO=""
TAG=""
SUPERVISOR="none"
SERVER_PID=""
RESTART_BY_EXIT="0"
STATUS_FILE=""
UPDATE_ID=""
FROM_VERSION=""
@@ -52,6 +60,7 @@ while [[ $# -gt 0 ]]; do
--log) LOG="$2"; shift 2 ;;
--prev-sha) PREV_SHA="$2"; shift 2 ;;
--server-pid) SERVER_PID="$2"; shift 2 ;;
--restart-by-exit) RESTART_BY_EXIT="$2"; shift 2 ;;
--stash) DO_STASH=1; shift ;;
*) shift ;;
esac
@@ -144,7 +153,7 @@ rollback_and_fail() {
echo "[self-update] $msg — rolling back to ${PREV_SHA:-<none>}"
if [[ -n "$PREV_SHA" ]]; then
git checkout --force "$PREV_SHA" >/dev/null 2>&1 || true
npm install --no-fund --no-audit >/dev/null 2>&1 || true
npm install --no-fund --no-audit --include=dev >/dev/null 2>&1 || true
npm run build >/dev/null 2>&1 || true
fi
fail "$msg — rolled back to the previous version" "$msg"
@@ -176,7 +185,9 @@ write_status "checkout" "Checking out $TAG…"
git -c advice.detachedHead=false checkout --force "$TAG" || rollback_and_fail "Could not check out $TAG"
# 4) Install dependencies (heartbeat keeps the UI live during this slow step).
run_step "installing" "Installing dependencies" npm install --no-fund --no-audit \
# --include=dev: tsc and esbuild are devDependencies, and the Compose image sets
# NODE_ENV=production, which would otherwise omit them and fail the build below.
run_step "installing" "Installing dependencies" npm install --no-fund --no-audit --include=dev \
|| rollback_and_fail "Dependency install failed"
# 5) Build (gate the restart on success — never restart into a torn dist/).
@@ -200,6 +211,43 @@ case "$SUPERVISOR" in
|| fail "Build succeeded but launchd restart failed" "launchctl"
}
;;
docker-compose)
# In the Compose deployment there is no init system to ask: the "restart" is
# the server EXITING, so the container's `restart: unless-stopped` policy
# relaunches it on the dist/ we just built. The repo and dist/ live on host
# mounts, so the new build survives the container being replaced.
#
# ⚠️ This script dies WITH the container it is restarting — it is a child of
# the server process, not a survivor like the systemd-scope path. That is
# fine, and load-bearing: the terminal "restarting" marker is already written
# above, and the freshly-booted server reconciles it. Nothing may be appended
# after the kill that the update depends on.
#
# ⚠️ The server is signalled by PID rather than `docker restart`: this
# container's own Docker CLI talks to the HOST daemon, and a self-directed
# restart there races the client's own death. Exiting is the one path that
# needs no cooperation from anything outside the container.
#
# ⚠️ Only when the SERVER said the container comes back (`--restart-by-exit 1`:
# the Compose file declared it, or the daemon reported an auto-restart policy).
# An unknown policy stages the build and asks for a restart instead. Exiting
# blind would take a container the daemon does not restart down for good,
# with no UI left to recover it from.
if [[ "$RESTART_BY_EXIT" != "1" ]]; then
MANUAL_CMD="docker restart \$(hostname) # from the Docker host"
write_status "completed-needs-manual-restart" "Update built — restart the Codeman container to apply v$TO_VERSION."
echo "[self-update] docker-compose: restart-by-exit not confirmed — not exiting; manual restart required"
exit 0
fi
if [[ -n "$SERVER_PID" ]] && kill "$SERVER_PID" 2>/dev/null; then
: # container exit + restart policy take it from here
else
MANUAL_CMD="docker restart \$(hostname) # from the Docker host"
write_status "completed-needs-manual-restart" "Update staged — restart the Codeman container to apply v$TO_VERSION."
echo "[self-update] docker-compose: could not signal server pid '$SERVER_PID' — manual restart required"
exit 0
fi
;;
launchd-daemon)
# System-level KeepAlive LaunchDaemon (headless Mac): kickstarting the system
# domain needs root, but we don't need it — kill the server and launchd
-662
View File
@@ -1,662 +0,0 @@
#!/bin/bash
# ============================================================================
# Codeman Sessions - Mobile-friendly Tmux Session Chooser
# Optimized for iPhone/Termius (portrait ~45 chars, landscape ~95 chars)
# ============================================================================
#
# Design principles:
# - Single-digit selection (1-9) for fast thumb typing
# - Compact display, no wasted space
# - Color-coded status for quick scanning
# - Names pulled from Codeman state.json
# - Minimal keystrokes to attach
#
# Usage:
# tmux-chooser # Interactive chooser
# tmux-chooser 1 # Quick attach to session 1
# tmux-chooser -l # List only (non-interactive)
# tmux-chooser -h # Help
#
# Alias: alias sc='tmux-chooser'
# Then: sc (interactive)
# sc 2 (attach session 2)
#
# ============================================================================
set -e
# ============================================================================
# Configuration
# ============================================================================
CODEMAN_STATE="$HOME/.codeman/state.json"
CODEMAN_SESSIONS="$HOME/.codeman/mux-sessions.json"
# Dedicated tmux socket all Codeman sessions live on. MUST match
# DEFAULT_CODEMAN_TMUX_SOCKET / CODEMAN_TMUX_SOCKET in src/tmux-manager.ts —
# otherwise list-sessions would enumerate the user's default tmux server
# (missing the real Codeman sessions, surfacing unrelated ones).
CODEMAN_TMUX_SOCKET="${CODEMAN_TMUX_SOCKET:-codeman}"
TMUX_CMD=(tmux -L "$CODEMAN_TMUX_SOCKET")
# iPhone 17 Pro portrait width (conservative)
MAX_WIDTH=44
MAX_NAME_LEN=28
# Page size for pagination (leave room for header/footer)
PAGE_SIZE=7
# Auto-refresh timeout (seconds) - 0 to disable
AUTO_REFRESH=60
# ============================================================================
# Icon Detection (Nerd Fonts vs ASCII)
# ============================================================================
detect_icons() {
if [[ "$TERM_PROGRAM" == "iTerm"* ]] || \
[[ "$TERM" == "xterm-kitty" ]] || \
[[ -n "$WEZTERM_PANE" ]] || \
[[ "$LC_TERMINAL" == "iTerm2" ]]; then
ICON_SESSION="󰆍"
ICON_ATTACHED="●"
ICON_DETACHED="○"
ICON_UNKNOWN="◌"
else
ICON_SESSION="[T]"
ICON_ATTACHED="*"
ICON_DETACHED="-"
ICON_UNKNOWN="?"
fi
}
detect_icons
# ============================================================================
# Colors - ANSI 256 for better Termius compatibility
# ============================================================================
R='\033[0m' # Reset
B='\033[1m' # Bold
D='\033[2m' # Dim
GREEN='\033[38;5;82m'
YELLOW='\033[38;5;220m'
BLUE='\033[38;5;75m'
CYAN='\033[38;5;87m'
RED='\033[38;5;203m'
GRAY='\033[38;5;245m'
WHITE='\033[38;5;255m'
BG_SEL='\033[48;5;236m'
# ============================================================================
# Utilities
# ============================================================================
truncate() {
local str="$1"
local max="$2"
local len=${#str}
if [ "$len" -le "$max" ]; then
echo "$str"
return
fi
if [[ "$str" == *"/"* ]]; then
echo "..${str: -$((max-2))}"
else
echo "${str:0:$((max-1))}…"
fi
}
find_full_session_id() {
local short_id="$1"
if [ -f "$CODEMAN_STATE" ]; then
local full_id
full_id=$(jq -r --arg short "$short_id" '
.sessions | keys[] | select(startswith($short))
' "$CODEMAN_STATE" 2>/dev/null | head -1)
if [ -n "$full_id" ]; then
echo "$full_id"
return
fi
fi
if [ -f "$CODEMAN_SESSIONS" ]; then
local full_id
full_id=$(jq -r --arg short "$short_id" '
.[] | select(.sessionId | startswith($short)) | .sessionId
' "$CODEMAN_SESSIONS" 2>/dev/null | head -1)
if [ -n "$full_id" ]; then
echo "$full_id"
return
fi
fi
echo "$short_id"
}
get_session_name() {
local session_id="$1"
local name=""
local workdir=""
if [ -f "$CODEMAN_SESSIONS" ]; then
local result
result=$(jq -r --arg id "$session_id" '
.[] | select(.sessionId | startswith($id)) | "\(.name // "")\t\(.workingDir // "")"
' "$CODEMAN_SESSIONS" 2>/dev/null | head -1)
if [ -n "$result" ]; then
name="${result%% *}"
workdir="${result#* }"
fi
fi
if [ -z "$name" ] && [ -f "$CODEMAN_STATE" ]; then
local result
result=$(jq -r --arg id "$session_id" '
.sessions | to_entries[] | select(.key | startswith($id)) | "\(.value.name // "")\t\(.value.workingDir // "")"
' "$CODEMAN_STATE" 2>/dev/null | head -1)
if [ -n "$result" ]; then
name="${result%% *}"
[ -z "$workdir" ] && workdir="${result#* }"
fi
fi
if [ -n "$name" ]; then
echo "$name"
return
fi
if [ -n "$workdir" ]; then
echo "${workdir##*/}"
return
fi
echo "${session_id:0:8}"
}
get_working_dir() {
local session_id="$1"
if [ -f "$CODEMAN_SESSIONS" ]; then
local dir
dir=$(jq -r --arg id "$session_id" '
.[] | select(.sessionId | startswith($id)) | .workingDir // empty
' "$CODEMAN_SESSIONS" 2>/dev/null | head -1)
if [ -n "$dir" ] && [ "$dir" != "null" ]; then
echo "${dir/#$HOME/~}"
return
fi
fi
if [ -f "$CODEMAN_STATE" ]; then
local dir
dir=$(jq -r --arg id "$session_id" '
.sessions | to_entries[] | select(.key | startswith($id)) | .value.workingDir // empty
' "$CODEMAN_STATE" 2>/dev/null | head -1)
if [ -n "$dir" ] && [ "$dir" != "null" ]; then
echo "${dir/#$HOME/~}"
return
fi
fi
echo ""
}
get_tokens() {
local session_id="$1"
if [ -f "$CODEMAN_STATE" ]; then
local tokens
tokens=$(jq -r --arg id "$session_id" '
.sessions | to_entries[] | select(.key | startswith($id)) |
((.value.inputTokens // 0) + (.value.outputTokens // 0))
' "$CODEMAN_STATE" 2>/dev/null | head -1)
if [ -n "$tokens" ] && [ "$tokens" != "null" ] && [ "$tokens" -gt 0 ] 2>/dev/null; then
if [ "$tokens" -gt 1000 ]; then
echo "$((tokens / 1000))k"
else
echo "${tokens}"
fi
return
fi
fi
echo ""
}
get_respawn_status() {
local session_id="$1"
if [ -f "$CODEMAN_SESSIONS" ]; then
local respawn_enabled
respawn_enabled=$(jq -r --arg id "$session_id" '
.[] | select(.sessionId | startswith($id)) | .respawnConfig.enabled // false
' "$CODEMAN_SESSIONS" 2>/dev/null | head -1)
if [ "$respawn_enabled" = "true" ]; then
echo "R"
return
fi
fi
echo ""
}
check_deps() {
if ! command -v jq &>/dev/null; then
echo -e "${YELLOW}Note: Install jq for session names${R}"
echo ""
fi
}
# ============================================================================
# Tmux Session Parser
# ============================================================================
declare -a SESSION_PIDS
declare -a MUX_NAMES
declare -a SESSION_STATES
declare -a SESSION_IDS
declare -a DISPLAY_NAMES
declare -a WORKING_DIRS
declare -a TOKEN_COUNTS
declare -a RESPAWN_STATUS
parse_sessions() {
SESSION_PIDS=()
MUX_NAMES=()
SESSION_STATES=()
SESSION_IDS=()
DISPLAY_NAMES=()
WORKING_DIRS=()
TOKEN_COUNTS=()
RESPAWN_STATUS=()
local i=0
# Parse tmux list-sessions output
while IFS= read -r line; do
local session_name="${line%%:*}"
# Only show codeman sessions
if [[ "$session_name" != codeman-* ]]; then
continue
fi
# Check if attached
local state="Detached"
if [[ "$line" == *"(attached)"* ]]; then
state="Attached"
fi
# Get PID from tmux
local pid
pid=$("${TMUX_CMD[@]}" display-message -t "$session_name" -p '#{pane_pid}' 2>/dev/null || echo "0")
SESSION_PIDS+=("$pid")
MUX_NAMES+=("$session_name")
SESSION_STATES+=("$state")
# Extract session ID from codeman session name
local session_id=""
local cm_regex='^codeman-(.+)$'
if [[ "$session_name" =~ $cm_regex ]]; then
session_id="${BASH_REMATCH[1]}"
fi
SESSION_IDS+=("$session_id")
# Get display name and metadata
if [ -n "$session_id" ]; then
DISPLAY_NAMES+=("$(get_session_name "$session_id")")
WORKING_DIRS+=("$(get_working_dir "$session_id")")
TOKEN_COUNTS+=("$(get_tokens "$session_id")")
RESPAWN_STATUS+=("$(get_respawn_status "$session_id")")
else
DISPLAY_NAMES+=("$session_name")
WORKING_DIRS+=("")
TOKEN_COUNTS+=("")
RESPAWN_STATUS+=("")
fi
i=$((i + 1))
done < <("${TMUX_CMD[@]}" list-sessions 2>/dev/null || true)
}
# ============================================================================
# Display Functions
# ============================================================================
clear_screen() {
printf '\033[2J\033[H'
}
print_header() {
local count=${#SESSION_PIDS[@]}
echo -e "${B}${CYAN}Codeman Sessions${R} ${D}($count)${R}"
echo -e "${D}$(printf '%.0s─' {1..32})${R}"
}
print_entry() {
local idx="$1"
local num=$((idx + 1))
local name="${DISPLAY_NAMES[$idx]}"
local state="${SESSION_STATES[$idx]}"
local dir="${WORKING_DIRS[$idx]}"
local tokens="${TOKEN_COUNTS[$idx]}"
local respawn="${RESPAWN_STATUS[$idx]}"
local name_max=$MAX_NAME_LEN
[ -n "$respawn" ] && name_max=$((name_max - 2))
[ -n "$tokens" ] && name_max=$((name_max - 4))
name=$(truncate "$name" $name_max)
local status_icon status_color
if [[ "$state" == *"Attached"* ]]; then
status_icon="$ICON_ATTACHED"
status_color="$GREEN"
elif [[ "$state" == *"Detached"* ]]; then
status_icon="$ICON_DETACHED"
status_color="$GRAY"
else
status_icon="$ICON_UNKNOWN"
status_color="$YELLOW"
fi
local num_str="${B}${WHITE}${num})${R}"
local name_str="${B}${WHITE}${name}${R}"
local status_str="${status_color}${status_icon}${R}"
local respawn_str=""
if [ -n "$respawn" ]; then
respawn_str=" ${GREEN}${respawn}${R}"
fi
local token_str=""
if [ -n "$tokens" ]; then
token_str=" ${D}${tokens}${R}"
fi
echo -e " ${num_str} ${name_str} ${status_str}${respawn_str}${token_str}"
if [ -n "$dir" ]; then
dir=$(truncate "$dir" $((MAX_NAME_LEN - 2)))
echo -e " ${D}${dir}${R}"
fi
}
print_footer() {
local page="$1"
local total_pages="$2"
echo ""
echo -e "${D}────────────────────────────────${R}"
if [ "$total_pages" -gt 1 ]; then
echo -e " ${D}Page $((page+1))/$total_pages${R} ${GRAY}[${WHITE}n${GRAY}]ext [${WHITE}p${GRAY}]rev${R}"
fi
echo -e " ${GRAY}[${WHITE}1-9${GRAY}]attach [${WHITE}r${GRAY}]efresh [${WHITE}q${GRAY}]uit${R}"
}
print_no_sessions() {
clear_screen
echo -e "${B}${CYAN}Codeman Sessions${R}"
echo -e "${D}$(printf '%.0s─' {1..32})${R}"
echo ""
echo -e " ${YELLOW}No tmux sessions found${R}"
echo ""
echo -e " ${D}Start one with:${R}"
echo -e " ${WHITE}codeman web${R}"
echo ""
echo -e "${D}$(printf '%.0s─' {1..32})${R}"
echo -e " ${GRAY}[${WHITE}r${GRAY}]efresh [${WHITE}q${GRAY}]uit${R}"
}
# ============================================================================
# Main Display Loop
# ============================================================================
current_page=0
render() {
clear_screen
parse_sessions
local count=${#SESSION_PIDS[@]}
if [ "$count" -eq 0 ]; then
print_no_sessions
return
fi
local total_pages=$(( (count + PAGE_SIZE - 1) / PAGE_SIZE ))
if [ "$current_page" -ge "$total_pages" ]; then
current_page=$((total_pages - 1))
fi
if [ "$current_page" -lt 0 ]; then
current_page=0
fi
local start=$((current_page * PAGE_SIZE))
local end=$((start + PAGE_SIZE))
if [ "$end" -gt "$count" ]; then
end=$count
fi
print_header
echo ""
for ((i = start; i < end; i++)); do
print_entry $i
done
print_footer $current_page $total_pages
}
attach_session() {
local idx="$1"
local mux_name="${MUX_NAMES[$idx]}"
if [ -z "$mux_name" ]; then
return 1
fi
clear_screen
echo -e "${GREEN}Attaching to ${B}${DISPLAY_NAMES[$idx]}${R}${GREEN}...${R}"
echo -e "${D}(Ctrl+B D to detach)${R}"
sleep 0.3
"${TMUX_CMD[@]}" attach-session -t "$mux_name"
return 0
}
# ============================================================================
# Input Handler
# ============================================================================
handle_input() {
local key="$1"
local count=${#SESSION_PIDS[@]}
local total_pages=$(( (count + PAGE_SIZE - 1) / PAGE_SIZE ))
case "$key" in
[1-9])
local idx=$((key - 1))
if [ "$idx" -lt "$count" ]; then
attach_session "$idx"
return 0
fi
;;
$'\e')
read -rsn2 -t 0.1 seq 2>/dev/null || true
case "$seq" in
'[A'|'[D')
if [ "$total_pages" -gt 1 ]; then
current_page=$(( (current_page - 1 + total_pages) % total_pages ))
fi
;;
'[B'|'[C')
if [ "$total_pages" -gt 1 ]; then
current_page=$(( (current_page + 1) % total_pages ))
fi
;;
esac
;;
n|N|j|J)
if [ "$total_pages" -gt 1 ]; then
current_page=$(( (current_page + 1) % total_pages ))
fi
;;
p|P|k|K)
if [ "$total_pages" -gt 1 ]; then
current_page=$(( (current_page - 1 + total_pages) % total_pages ))
fi
;;
r|R)
;;
q|Q)
clear_screen
exit 0
;;
'')
if [ "$count" -eq 1 ]; then
attach_session 0
return 0
fi
;;
esac
return 0
}
# ============================================================================
# List Mode
# ============================================================================
list_mode() {
parse_sessions
local count=${#SESSION_PIDS[@]}
if [ "$count" -eq 0 ]; then
echo "No tmux sessions"
exit 0
fi
for ((i = 0; i < count; i++)); do
local num=$((i + 1))
local name="${DISPLAY_NAMES[$i]}"
local state="${SESSION_STATES[$i]}"
local respawn="${RESPAWN_STATUS[$i]}"
local indicator="-"
[[ "$state" == *"Attached"* ]] && indicator="*"
[ -n "$respawn" ] && indicator="${indicator}R"
echo "$num) $name [$indicator]"
done
}
# ============================================================================
# Quick Attach
# ============================================================================
quick_attach() {
local num="$1"
parse_sessions
local count=${#SESSION_PIDS[@]}
local idx=$((num - 1))
if [ "$idx" -lt 0 ] || [ "$idx" -ge "$count" ]; then
echo -e "${RED}Invalid session: $num${R}"
echo "Available: 1-$count"
exit 1
fi
attach_session "$idx"
}
# ============================================================================
# Help
# ============================================================================
show_help() {
cat << 'EOF'
Codeman Sessions - Mobile-friendly Tmux Session Chooser
USAGE:
sc Interactive chooser
sc <number> Quick attach to session N
sc -l List sessions (non-interactive)
sc -h Show this help
INTERACTIVE KEYS:
1-9 Attach to session
n/j/↓ Next page
p/k/↑ Previous page
r Refresh
q Quit
INDICATORS:
* / ● Attached (someone connected)
- / ○ Detached (available)
R Respawn enabled
45k Token count
TIPS:
- Detach from tmux: Ctrl+B D
- Session names from Codeman state
- Optimized for Termius/iPhone
EOF
}
# ============================================================================
# Main
# ============================================================================
main() {
case "${1:-}" in
-h|--help)
show_help
exit 0
;;
-l|--list)
list_mode
exit 0
;;
[1-9]|[1-9][0-9])
quick_attach "$1"
exit $?
;;
esac
check_deps
render
while true; do
local timeout_opt=""
if [ "$AUTO_REFRESH" -gt 0 ]; then
timeout_opt="-t $AUTO_REFRESH"
fi
if read -rsn1 $timeout_opt key 2>/dev/null; then
handle_input "$key"
fi
render
done
}
trap 'clear_screen; exit 0' INT
main "$@"
+173 -48
View File
@@ -47,7 +47,7 @@ later call opens with, and your first REAL call performs them anyway:
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.19.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
```
⚠️ **Never spend a Bash call on this check alone.** §1's block opens with this same
@@ -75,8 +75,8 @@ PRE="${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh"
mkdir -p "$(dirname "$PRE")"
# Rewrite unless the file already ends with THIS version's stamp, so a stale or a
# half-written file self-heals here instead of costing you a round trip to rm it.
grep -qs '^CODEMAN_PREAMBLE=1.19.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.19.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
grep -qs '^CODEMAN_PREAMBLE=1.22.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -96,7 +96,10 @@ AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:
# draw the lineage. Set once here and every present and future create call carries it;
# it is ignored on every other endpoint. Purely cosmetic (see §5.1) and it can never
# fail a spawn, so there is no case where you would want to leave it off.
CURL=(curl -sk "${AUTH[@]}" -H "X-Codeman-Parent-Session: $SELF")
# X-Codeman-Agent-Origin: marks a case directory a spawn CREATES as agent scratch, so the
# user can find and delete it long after your workers are gone (§5.14). Same deal: set
# once, cosmetic, never fails a spawn, and it labels only directories Codeman creates.
CURL=(curl -sk "${AUTH[@]}" -H "X-Codeman-Parent-Session: $SELF" -H "X-Codeman-Agent-Origin: codeman-skill")
CID=codeman-agent-1 # FIXED literal, never "agent-$$": see below
# Fail-CLOSED session delete. The DELETE lives INSIDE the guard on purpose: the older
@@ -121,23 +124,90 @@ _composer_up() { # <sid> <timeoutMs> -> "true"/"false". `shift+tab` is the one
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' \
--data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
_dsh_up() { # <sid> <timeoutMs> -> "true"/"false". The DeepSeek Harness TUI's
# composer glyph. Override with DSH_READY_MARK for a profile that draws another one.
"${CURL[@]}" -G "$API/api/v1/sessions/$1/wait-output" \
--data-urlencode "match=${DSH_READY_MARK:-❯}" --data-urlencode 'from=buffer' \
--data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
# ---- the workspace-trust dialog: READ the screen, never press Enter blind ----
# Claude Code 2.1.252 dropped the option numbers, REVERSED them, and highlights
# "No, exit" by default:
# Security guide
# ❯ No, exit
# Yes, I trust this folder
# Enter to confirm . Esc to cancel
# so the bare \r that answered the old layout now answers *exit* and the pane is
# dead (`status 1`) seconds after the spawn -- measured on a live 2.1.252 case.
# These two read the rendered pane and steer onto the trust option instead.
_trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
# full=1 returns the RENDERED pane; a claude pane keeps no tmux history, so that
# is the current frame rather than every repaint since launch. tail -1 anyway,
# because the freshest marked row is the only one still true.
"${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
k=$(_trust_key "$sid")
[ -n "$k" ] || return 1 # no dialog on screen, or a layout this cannot read
# A SEPARATE clientId for these keys. seq is monotonic per clientId, so
# spending prompt numbers here would make the next sendwait -- whose default
# seq is the epoch second -- look like a stale duplicate and vanish silently.
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg k "$([ "$k" = confirm ] && printf '\r' || printf '\033[B')" \
--arg c "$CID-trust-$sid" --argjson s "$i" \
'{input:$k,useMux:true,clientId:$c,seq:$s}')" >/dev/null
[ "$k" = confirm ] && return 0
sleep 1; i=$((i+1)) # re-read: the arrow is CONFIRMED before Enter goes out
done
return 1
}
# spawn_worker <caseName> [mode] -> session id on stdout, diagnostics on stderr.
# quick-start AND readiness in one call, with a strict contract: NON-EMPTY stdout means
# a READY claude worker in a hook-carrying case. Anything less is rc 1 with EMPTY
# stdout, and the half-spawned session is deleted here rather than handed back, because
# a worker that never drew its composer would eat the task prompt with its trust
# dialog. There is deliberately no pid poll: wait-output already blocks until the
# composer draws, and pid!=null proved startup, never readiness.
# a READY worker whose end-of-turn signal can be trusted -- a claude worker in a
# hook-carrying case, or a `deepseek` worker whose harness TUI drew its composer.
# Anything less is rc 1 with EMPTY stdout, and the half-spawned session is deleted here
# rather than handed back, because a worker that never drew its composer would eat the
# task prompt with its trust dialog. There is deliberately no pid poll: wait-output
# already blocks until the composer draws, and pid!=null proved startup, never readiness.
spawn_worker() {
local name="${1:?spawn_worker needs a case name}" mode="${2:-claude}" q sid cp r
# parentSessionId doubles the CURL header, so a spawn_worker copied off the shared
# curl (or a body someone rebuilt from this recipe) still carries its lineage.
# deepseek: ask for the same permission posture the Run button sends, because the
# harness's own default (`workspace-write`) still ASKS, and a worker that stops on
# an approval row is a worker no fan-out can finish. It is not an escalation --
# claude workers already spawn with permissions skipped, and in multi-user mode the
# server clamps this back to `workspace-write` for an owner without the grant.
# Spawn by hand (§5.1) when you want a worker that asks.
q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg n "$name" --arg m "$mode" --arg p "$SELF" '{caseName:$n,mode:$m,parentSessionId:$p}')")
-d "$(jq -nc --arg n "$name" --arg m "$mode" --arg p "$SELF" \
'{caseName:$n,mode:$m,parentSessionId:$p}
+ (if $m == "deepseek" then {deepSeekConfig:{permissionMode:"danger-full-access"}} else {} end)')")
sid=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$q")
# NOT retryable in a loop: every quick-start failure code is terminal (§5.1).
[ -n "$sid" ] || { jq -c '{error,errorCode}' <<<"$q" >&2; return 1; }
[ "$mode" = claude ] || { printf '%s\n' "$sid"; return 0; } # only claude draws a composer
if [ "$mode" = deepseek ]; then
# The one non-claude mode with REAL end-of-turn signals: its TUI reports
# idle/working/blocked to Codeman, so sendwait, until=stop and the Approvals
# Inbox all work here exactly as they do for claude. No hook file to vet
# (the bridge is env-injected, not a workspace file) and no trust dialog.
# ⚠️ Readiness is still not optional, and NOT interchangeable with the stop
# signal: the harness's boot report lands ~300ms BEFORE the composer paints
# (measured 2.26s vs 2.56s after spawn), so a sendwait fired straight after
# quick-start returns on that BOOT signal, reports a turn that never ran, and
# strands the prompt in a pane that was not yet taking input.
r=$(_dsh_up "$sid" 45000)
[ "$r" = true ] || { echo "dsh worker $sid never drew a composer: no pane-capable profile, a profile whose composer is not '${DSH_READY_MARK:-❯}' (set DSH_READY_MARK), or a harness that failed to boot -- check GET /api/v1/deepseek/status. Deleted it" >&2
delete_session "$sid" >/dev/null; return 1; }
printf '%s\n' "$sid"; return 0
fi
[ "$mode" = claude ] || { printf '%s\n' "$sid"; return 0; } # no other mode draws a composer to wait on
# The server installs hooks into every claude workspace now, so this grep normally
# passes; it stays because the install is gated on a setting the operator can turn
# off, remote sessions never get hooks, and a session created by an older server
@@ -147,38 +217,41 @@ spawn_worker() {
grep -qs '/api/hook-event' "$cp/.claude/settings.local.json" || {
echo "case '$name' resolved to '$cp', which has no Codeman hooks (workspaceHooksEnabled off, remote, or an older server?): turn the setting on, or work §5.1+§5.5 by hand with markers" >&2
delete_session "$sid" >/dev/null; return 1; }
# Short composer wait FIRST, then the trust-dialog probe: a case still showing the
# dialog can never pass the composer wait, so probing early keeps a cold case from
# Short composer wait FIRST, then the trust dialog: a case still showing the
# dialog can never pass the composer wait, so acting early keeps a cold case from
# paying the whole long wait before the fallback even runs (§5.2). A warm case
# matches in under a second and never reaches the probe.
# matches in under a second and never reaches it, and _accept_trust returns in a
# blink when there is no dialog, so this costs nothing in the ordinary slow case.
r=$(_composer_up "$sid" 5000)
if [ "$r" != true ]; then
if "${CURL[@]}" -G "$API/api/v1/sessions/$sid/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000' \
| jq -e '.data.wait.matched' >/dev/null; then
# Codeman's own auto-accept gives up after 90 s / 3 tries; this is that bounded fallback.
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" '{input:"\r",useMux:true,clientId:$c,seq:1}')" >/dev/null
fi
# Codeman answers this dialog itself and normally wins the race; this is the
# bounded fallback for when its 90 s window / 6-keystroke cap has run out.
_accept_trust "$sid"
r=$(_composer_up "$sid" 45000)
fi
[ "$r" = true ] || { echo "worker $sid never drew a composer; deleted it. Retry by hand via the §5.2 ladder (its billed stage-4 probe included)" >&2
delete_session "$sid" >/dev/null; return 1; }
printf '%s\n' "$sid"
}
# spawn_workers <caseName>... -> one "<caseName> <sessionId>" line per worker, in order;
# the sessionId column is EMPTY for a spawn that failed (stderr has why). CONCURRENT:
# N workers cost about what one costs. Spawning them one Bash call at a time is the
# single biggest avoidable delay in this skill. Names must be UNIQUE: two workers in
# one case directory co-edit the same tree (§4), so a repeat is an error here, not a race.
# spawn_workers <caseName[:mode]>... -> one "<caseName> <sessionId>" line per worker, in
# order; the sessionId column is EMPTY for a spawn that failed (stderr has why).
# CONCURRENT: N workers cost about what one costs. Spawning them one Bash call at a time
# is the single biggest avoidable delay in this skill. A bare name is a claude worker;
# `beta:deepseek` makes that one a DeepSeek Harness worker, and a mixed fleet is one
# call. Case names must be UNIQUE: two workers in one case directory co-edit the same
# tree (§4), so a repeat is an error here, not a race (the mode never disambiguates two
# workers, since they would still share the directory).
spawn_workers() {
local d n i=0
local d spec n m i=0
[ "$#" -gt 0 ] || { echo "spawn_workers: no case names given" >&2; return 1; }
[ -z "$(printf '%s\n' "$@" | sort | uniq -d)" ] || { echo "spawn_workers: duplicate case names" >&2; return 1; }
[ -z "$(printf '%s\n' "$@" | sed 's/:.*//' | sort | uniq -d)" ] || { echo "spawn_workers: duplicate case names" >&2; return 1; }
d=$(mktemp -d "${TMPDIR:-/tmp}/codeman-spawn.XXXXXX") || return 1
for n in "$@"; do ( spawn_worker "$n" > "$d/$i" ) & i=$((i+1)); done
for spec in "$@"; do
n=${spec%%:*}; m=${spec#*:}; [ "$m" = "$spec" ] && m=claude
( spawn_worker "$n" "$m" > "$d/$i" ) & i=$((i+1))
done
wait
i=0; for n in "$@"; do printf '%s %s\n' "$n" "$(cat "$d/$i" 2>/dev/null)"; i=$((i+1)); done
i=0; for spec in "$@"; do printf '%s %s\n' "${spec%%:*}" "$(cat "$d/$i" 2>/dev/null)"; i=$((i+1)); done
rm -rf "$d"
}
# sendwait <sid> <prompt> [seq] -> blocks until that worker's turn ENDS (~10 min ceiling
@@ -194,27 +267,46 @@ spawn_workers() {
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy only for a claude
# worker spawn_worker handed back (hooks vetted); hook-less workspaces and other modes
# resolve on flapping idle: markers instead (§5.5).
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
# implement the status contract is the one case that LOOKS like claude but is not:
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
# is still answering, and the re-wait below then resolved in 0 ms with
# `signal:"idle"` on a turn that had another three minutes to run (measured).
# A wait named after the end of a turn should only end with the turn, or with
# the worker. ⚠️ This is also what makes a wrong mode LOUD: the modes that
# cannot deliver `stop` answer 400 (before writing anything) instead of
# resolving on a flap, which is the answer that sends you to markers (§5.5).
body=$(jq -nc --arg p "$p" --arg c "$CID-$sid" --argjson s "$seq" \
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:true,waitTimeout:20000}')
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:"stop,exit",waitTimeout:20000}')
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")")
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
fi
printf '%s\n' "$r"
}
# last_text <sid> [prev] -> that worker's last assistant message. Polled, because the
# transcript write LAGS the stop signal, and "some text exists" is not "THIS turn's
# text exists": right after a SECOND turn on the same worker the endpoint still serves
# last_text <sid> [prev] -> that worker's last assistant message (claude, codex and
# deepseek write a real transcript; the other modes have none, so read the terminal
# instead -- §5.4). Polled, because the transcript write LAGS the stop signal, and
# "some text exists" is not "THIS turn's text exists": right after a SECOND turn on the same worker the endpoint still serves
# the previous answer for a beat (observed live). When reading consecutive turns, pass
# the previous answer as [prev]: the poll then holds out for text that differs from it,
# falling back to whatever it last saw if the budget runs dry, so an honestly repeated
@@ -233,10 +325,10 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.19.0
CODEMAN_PREAMBLE=1.22.0
PREAMBLE
)
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.19.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
```
Every later Bash call that touches the API starts with the same two loader lines from
@@ -287,8 +379,9 @@ and no per-call body to hand-build.
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null # §0 loader
[ "${CODEMAN_PREAMBLE:-}" = 1.19.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
N=(alpha beta) # INVENT one fresh case name per worker; never list cases first
# (a name may carry a mode: `beta:deepseek`, see below)
T=('reply with one line: the absolute path of your working directory'
'reply with one line: your model name') # tasks, same order as N
@@ -351,7 +444,39 @@ Four things this block leans on, each one link away, no detour needed to run it:
strands the prompt on the composer until a bare `\r` follows: all three are reasons
to let `sendwait` build the call rather than hand-rolling it.
- Each `sendwait` costs that worker one billed turn, as does every prompt you send it.
- Deleting the sessions does **not** remove the case directories: §5.14.
- Deleting the sessions does **not** remove the case directories. They are marked as
agent-created, so `GET /api/v1/cases/agent-created` lists them for cleanup: §5.14.
### DeepSeek Harness workers
The block above spawns claude workers. Any entry in `N` may instead name a mode
(`beta:deepseek`), and **a `deepseek` worker is driven by the same four verbs, with no
change to the rest of the block**: `spawn_workers` waits for its composer, `sendwait`
blocks on its real end-of-turn signal, `last_text` reads its answer, `delete_session`
removes it.
That is true of no other non-claude mode, and it is worth knowing why: the DeepSeek
Harness TUI reports `idle`/`working`/`blocked` to Codeman over the supervisor contract it
implements, so dsh is the one external CLI with definitive `stop`/`blocked` signals
instead of guessed-from-silence ones — and it writes a structured transcript, which is
what `last-response` reads for it. `shell`, `opencode`, `codex`, `gemini`, `antigravity`,
`pi`, `grok` and `omp` have neither and still need markers ([§5.5](reference/verbs.md#55-markers-for-hook-less-workers)).
Three things to know before you spawn one:
- **It needs a pane-capable profile.** `dsh` ships only `web`/`headless`, so the terminal
agent is always an installed profile. `GET /api/v1/deepseek/status` answers both
questions separately (`available` = the binary, `runnable` = a profile that can drive a
pane); a spawn without one fails with `OPERATION_FAILED` rather than falling back.
- **Do not task it on the strength of a `stop` alone.** The harness reports `idle` at
boot ~300 ms *before* its composer paints (measured 2.26 s vs 2.56 s), so a `sendwait`
fired straight after `quick-start` resolves on that boot signal, reports a turn that
never ran, and leaves the prompt in a pane that was not yet taking input. Letting
`spawn_worker` gate on readiness is what steps past that edge; it is not optional.
- **A profile that does not implement the contract looks like a hang.** Codeman cannot
know at spawn time whether one does. The tell is a `sendwait` that times out on a
worker whose pane clearly finished: that profile is one of them, so drive it with
markers instead.
## 2. What do you want to do?
@@ -360,10 +485,10 @@ One row per job. Acting on this table alone is correct; the §5 links are the de
| I want to | Call | Detail |
|-----------|------|--------|
| start a worker **where the work is** | `POST /api/v1/quick-start {"caseName":…}`, which **creates** `~/codeman-cases/<name>` unless the name is already a case. Any other path (a git worktree): `POST /api/v1/sessions {"workingDir":…}` then `POST /api/v1/sessions/:id/interactive`. Both install hooks by default, so expect full signals in either, and **verify** rather than assume. N workers means N worktrees | [§5.1](reference/verbs.md#51-where-to-spawn) |
| know a new worker can accept a prompt | `GET .../wait-output?match=shift+tab&from=buffer` (urlencode the `+`) | [§5.2](reference/verbs.md#52-readiness) |
| deliver a task **and** know when it finished | `POST .../input` with `"input":"…\r"`, `clientId`, `seq`, `"wait":true`. Resolves on `stop`, so it is trustworthy only where the workspace **has hooks** (claude mode; installed by default, but the operator can disable it and remote sessions never get them). Costs the worker one billed turn | [§5.3](reference/verbs.md#53-send-a-task-and-wait) |
| know a new worker can accept a prompt | `GET .../wait-output?match=shift+tab&from=buffer` (urlencode the `+`); a `deepseek` worker draws `❯` instead, and its boot `stop` fires ~300 ms BEFORE that, so never read the signal as readiness | [§5.2](reference/verbs.md#52-readiness) |
| deliver a task **and** know when it finished | `POST .../input` with `"input":"…\r"`, `clientId`, `seq`, `"wait":true`. Resolves on `stop`, so it is trustworthy where the signal is real: claude mode with hooks (installed by default, but the operator can disable it and remote sessions never get them) and `deepseek` mode through its status bridge. Costs the worker one billed turn | [§5.3](reference/verbs.md#53-send-a-task-and-wait) |
| know a hook-less worker finished | it has no `stop`, and `wait:true` there resolves on flapping `idle` **without erroring**: make it print a split, unique marker and `wait-output` on that instead | [§5.5](reference/verbs.md#55-markers-for-hook-less-workers) |
| read the answer | `GET .../last-response`, **polled** (claude/codex only; empty for the other modes) | [§5.4](reference/verbs.md#54-read-the-answer) |
| read the answer | `GET .../last-response`, **polled** (claude, codex and deepseek write a transcript; empty for the other modes) | [§5.4](reference/verbs.md#54-read-the-answer) |
| know if it is alive | `GET .../wait?until=exit&timeout=1000`: an immediate `signal:"exit"` means dead. `status` and `pid` both lie | [§5.6](reference/verbs.md#56-alive-and-stuck) |
| know if it is stuck | `GET .../active-tools` and `GET .../run-summary` are structured and free; two `terminal?tail=` samples are the crude fallback | [§5.6](reference/verbs.md#56-alive-and-stuck) |
| make a runaway worker stop | `POST .../input {"input":"\u001b"}` (ESC, **no** `\r`). Deleting the session would destroy the conversation instead | [§5.7](reference/verbs.md#57-interrupt-without-destroying) |
@@ -373,7 +498,7 @@ One row per job. Acting on this table alone is correct; the §5 links are the de
| find yourself, list what exists | `GET /api/v1/sessions`, match your `$SELF` by **prefix** | [§5.11](reference/verbs.md#511-list-and-find-yourself) |
| read or record what the user wants | `GET/PUT .../intent`, and `POST .../readmymind` to predict | [§5.12](reference/verbs.md#512-read-my-mind) |
| talk to a claude worker directly | `ListAgents` / `SendMessage`, when the feature is on at both ends | [§5.13](reference/verbs.md#513-messaging-claude-workers) |
| clean up | `delete_session "$SID"` per id you created. Case directories and git worktrees are **not** removed with it | [§5.14](reference/verbs.md#514-clean-up) |
| clean up | `delete_session "$SID"` per id you created. Case directories and git worktrees are **not** removed with it; `GET /api/v1/cases/agent-created` lists the scratch case dirs your spawns left behind, for you to report | [§5.14](reference/verbs.md#514-clean-up) |
## 3. Rules digest
@@ -464,7 +589,7 @@ these**; open the one row you actually hit.
| [5.1 Where to spawn](reference/verbs.md#51-where-to-spawn) | the work is **not** a fresh scratch case: a linked case, a git worktree, any path that already existed. Hooks are absent there, which silently breaks send-and-wait. The costliest mistake in this skill |
| [5.2 Readiness](reference/verbs.md#52-readiness) | a worker never drew its composer, or you need the trust-dialog ladder by hand |
| [5.3 Send a task and wait](reference/verbs.md#53-send-a-task-and-wait) | the `sendwait` body, its signals, and the duplicate-resend loop |
| [5.4 Read the answer](reference/verbs.md#54-read-the-answer) | `last_text` came back empty, or the mode is not claude/codex |
| [5.4 Read the answer](reference/verbs.md#54-read-the-answer) | `last_text` came back empty, or the mode is not claude/codex/deepseek |
| [5.5 Markers for hook-less workers](reference/verbs.md#55-markers-for-hook-less-workers) | the worker has no `stop` hook: synchronize on a split, unique printed marker |
| [5.6 Alive and stuck](reference/verbs.md#56-alive-and-stuck) | is it dead or just slow? `status` and `pid` both lie |
| [5.7 Interrupt without destroying](reference/verbs.md#57-interrupt-without-destroying) | a runaway worker you want to stop but keep |
@@ -474,7 +599,7 @@ these**; open the one row you actually hit.
| [5.11 List and find yourself](reference/verbs.md#511-list-and-find-yourself) | enumerate sessions, or match `$SELF` by prefix |
| [5.12 Read My Mind](reference/verbs.md#512-read-my-mind) | read or record what the user wants for a case |
| [5.13 Messaging claude workers](reference/verbs.md#513-messaging-claude-workers) | `ListAgents` / `SendMessage` instead of the HTTP path |
| [5.14 Clean up](reference/verbs.md#514-clean-up) | what deleting a session does **not** remove |
| [5.14 Clean up](reference/verbs.md#514-clean-up) | what deleting a session does **not** remove, and how to list the case dirs you left |
## 6. Setup and auth
+129 -37
View File
@@ -1,4 +1,4 @@
# ---- Codeman agent preamble 1.19.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -18,7 +18,10 @@ AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:
# draw the lineage. Set once here and every present and future create call carries it;
# it is ignored on every other endpoint. Purely cosmetic (see §5.1) and it can never
# fail a spawn, so there is no case where you would want to leave it off.
CURL=(curl -sk "${AUTH[@]}" -H "X-Codeman-Parent-Session: $SELF")
# X-Codeman-Agent-Origin: marks a case directory a spawn CREATES as agent scratch, so the
# user can find and delete it long after your workers are gone (§5.14). Same deal: set
# once, cosmetic, never fails a spawn, and it labels only directories Codeman creates.
CURL=(curl -sk "${AUTH[@]}" -H "X-Codeman-Parent-Session: $SELF" -H "X-Codeman-Agent-Origin: codeman-skill")
CID=codeman-agent-1 # FIXED literal, never "agent-$$": see below
# Fail-CLOSED session delete. The DELETE lives INSIDE the guard on purpose: the older
@@ -43,23 +46,90 @@ _composer_up() { # <sid> <timeoutMs> -> "true"/"false". `shift+tab` is the one
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' \
--data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
_dsh_up() { # <sid> <timeoutMs> -> "true"/"false". The DeepSeek Harness TUI's
# composer glyph. Override with DSH_READY_MARK for a profile that draws another one.
"${CURL[@]}" -G "$API/api/v1/sessions/$1/wait-output" \
--data-urlencode "match=${DSH_READY_MARK:-❯}" --data-urlencode 'from=buffer' \
--data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
# ---- the workspace-trust dialog: READ the screen, never press Enter blind ----
# Claude Code 2.1.252 dropped the option numbers, REVERSED them, and highlights
# "No, exit" by default:
# Security guide
# ❯ No, exit
# Yes, I trust this folder
# Enter to confirm . Esc to cancel
# so the bare \r that answered the old layout now answers *exit* and the pane is
# dead (`status 1`) seconds after the spawn -- measured on a live 2.1.252 case.
# These two read the rendered pane and steer onto the trust option instead.
_trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
# full=1 returns the RENDERED pane; a claude pane keeps no tmux history, so that
# is the current frame rather than every repaint since launch. tail -1 anyway,
# because the freshest marked row is the only one still true.
"${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
k=$(_trust_key "$sid")
[ -n "$k" ] || return 1 # no dialog on screen, or a layout this cannot read
# A SEPARATE clientId for these keys. seq is monotonic per clientId, so
# spending prompt numbers here would make the next sendwait -- whose default
# seq is the epoch second -- look like a stale duplicate and vanish silently.
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg k "$([ "$k" = confirm ] && printf '\r' || printf '\033[B')" \
--arg c "$CID-trust-$sid" --argjson s "$i" \
'{input:$k,useMux:true,clientId:$c,seq:$s}')" >/dev/null
[ "$k" = confirm ] && return 0
sleep 1; i=$((i+1)) # re-read: the arrow is CONFIRMED before Enter goes out
done
return 1
}
# spawn_worker <caseName> [mode] -> session id on stdout, diagnostics on stderr.
# quick-start AND readiness in one call, with a strict contract: NON-EMPTY stdout means
# a READY claude worker in a hook-carrying case. Anything less is rc 1 with EMPTY
# stdout, and the half-spawned session is deleted here rather than handed back, because
# a worker that never drew its composer would eat the task prompt with its trust
# dialog. There is deliberately no pid poll: wait-output already blocks until the
# composer draws, and pid!=null proved startup, never readiness.
# a READY worker whose end-of-turn signal can be trusted -- a claude worker in a
# hook-carrying case, or a `deepseek` worker whose harness TUI drew its composer.
# Anything less is rc 1 with EMPTY stdout, and the half-spawned session is deleted here
# rather than handed back, because a worker that never drew its composer would eat the
# task prompt with its trust dialog. There is deliberately no pid poll: wait-output
# already blocks until the composer draws, and pid!=null proved startup, never readiness.
spawn_worker() {
local name="${1:?spawn_worker needs a case name}" mode="${2:-claude}" q sid cp r
# parentSessionId doubles the CURL header, so a spawn_worker copied off the shared
# curl (or a body someone rebuilt from this recipe) still carries its lineage.
# deepseek: ask for the same permission posture the Run button sends, because the
# harness's own default (`workspace-write`) still ASKS, and a worker that stops on
# an approval row is a worker no fan-out can finish. It is not an escalation --
# claude workers already spawn with permissions skipped, and in multi-user mode the
# server clamps this back to `workspace-write` for an owner without the grant.
# Spawn by hand (§5.1) when you want a worker that asks.
q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg n "$name" --arg m "$mode" --arg p "$SELF" '{caseName:$n,mode:$m,parentSessionId:$p}')")
-d "$(jq -nc --arg n "$name" --arg m "$mode" --arg p "$SELF" \
'{caseName:$n,mode:$m,parentSessionId:$p}
+ (if $m == "deepseek" then {deepSeekConfig:{permissionMode:"danger-full-access"}} else {} end)')")
sid=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$q")
# NOT retryable in a loop: every quick-start failure code is terminal (§5.1).
[ -n "$sid" ] || { jq -c '{error,errorCode}' <<<"$q" >&2; return 1; }
[ "$mode" = claude ] || { printf '%s\n' "$sid"; return 0; } # only claude draws a composer
if [ "$mode" = deepseek ]; then
# The one non-claude mode with REAL end-of-turn signals: its TUI reports
# idle/working/blocked to Codeman, so sendwait, until=stop and the Approvals
# Inbox all work here exactly as they do for claude. No hook file to vet
# (the bridge is env-injected, not a workspace file) and no trust dialog.
# ⚠️ Readiness is still not optional, and NOT interchangeable with the stop
# signal: the harness's boot report lands ~300ms BEFORE the composer paints
# (measured 2.26s vs 2.56s after spawn), so a sendwait fired straight after
# quick-start returns on that BOOT signal, reports a turn that never ran, and
# strands the prompt in a pane that was not yet taking input.
r=$(_dsh_up "$sid" 45000)
[ "$r" = true ] || { echo "dsh worker $sid never drew a composer: no pane-capable profile, a profile whose composer is not '${DSH_READY_MARK:-❯}' (set DSH_READY_MARK), or a harness that failed to boot -- check GET /api/v1/deepseek/status. Deleted it" >&2
delete_session "$sid" >/dev/null; return 1; }
printf '%s\n' "$sid"; return 0
fi
[ "$mode" = claude ] || { printf '%s\n' "$sid"; return 0; } # no other mode draws a composer to wait on
# The server installs hooks into every claude workspace now, so this grep normally
# passes; it stays because the install is gated on a setting the operator can turn
# off, remote sessions never get hooks, and a session created by an older server
@@ -69,38 +139,41 @@ spawn_worker() {
grep -qs '/api/hook-event' "$cp/.claude/settings.local.json" || {
echo "case '$name' resolved to '$cp', which has no Codeman hooks (workspaceHooksEnabled off, remote, or an older server?): turn the setting on, or work §5.1+§5.5 by hand with markers" >&2
delete_session "$sid" >/dev/null; return 1; }
# Short composer wait FIRST, then the trust-dialog probe: a case still showing the
# dialog can never pass the composer wait, so probing early keeps a cold case from
# Short composer wait FIRST, then the trust dialog: a case still showing the
# dialog can never pass the composer wait, so acting early keeps a cold case from
# paying the whole long wait before the fallback even runs (§5.2). A warm case
# matches in under a second and never reaches the probe.
# matches in under a second and never reaches it, and _accept_trust returns in a
# blink when there is no dialog, so this costs nothing in the ordinary slow case.
r=$(_composer_up "$sid" 5000)
if [ "$r" != true ]; then
if "${CURL[@]}" -G "$API/api/v1/sessions/$sid/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000' \
| jq -e '.data.wait.matched' >/dev/null; then
# Codeman's own auto-accept gives up after 90 s / 3 tries; this is that bounded fallback.
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" '{input:"\r",useMux:true,clientId:$c,seq:1}')" >/dev/null
fi
# Codeman answers this dialog itself and normally wins the race; this is the
# bounded fallback for when its 90 s window / 6-keystroke cap has run out.
_accept_trust "$sid"
r=$(_composer_up "$sid" 45000)
fi
[ "$r" = true ] || { echo "worker $sid never drew a composer; deleted it. Retry by hand via the §5.2 ladder (its billed stage-4 probe included)" >&2
delete_session "$sid" >/dev/null; return 1; }
printf '%s\n' "$sid"
}
# spawn_workers <caseName>... -> one "<caseName> <sessionId>" line per worker, in order;
# the sessionId column is EMPTY for a spawn that failed (stderr has why). CONCURRENT:
# N workers cost about what one costs. Spawning them one Bash call at a time is the
# single biggest avoidable delay in this skill. Names must be UNIQUE: two workers in
# one case directory co-edit the same tree (§4), so a repeat is an error here, not a race.
# spawn_workers <caseName[:mode]>... -> one "<caseName> <sessionId>" line per worker, in
# order; the sessionId column is EMPTY for a spawn that failed (stderr has why).
# CONCURRENT: N workers cost about what one costs. Spawning them one Bash call at a time
# is the single biggest avoidable delay in this skill. A bare name is a claude worker;
# `beta:deepseek` makes that one a DeepSeek Harness worker, and a mixed fleet is one
# call. Case names must be UNIQUE: two workers in one case directory co-edit the same
# tree (§4), so a repeat is an error here, not a race (the mode never disambiguates two
# workers, since they would still share the directory).
spawn_workers() {
local d n i=0
local d spec n m i=0
[ "$#" -gt 0 ] || { echo "spawn_workers: no case names given" >&2; return 1; }
[ -z "$(printf '%s\n' "$@" | sort | uniq -d)" ] || { echo "spawn_workers: duplicate case names" >&2; return 1; }
[ -z "$(printf '%s\n' "$@" | sed 's/:.*//' | sort | uniq -d)" ] || { echo "spawn_workers: duplicate case names" >&2; return 1; }
d=$(mktemp -d "${TMPDIR:-/tmp}/codeman-spawn.XXXXXX") || return 1
for n in "$@"; do ( spawn_worker "$n" > "$d/$i" ) & i=$((i+1)); done
for spec in "$@"; do
n=${spec%%:*}; m=${spec#*:}; [ "$m" = "$spec" ] && m=claude
( spawn_worker "$n" "$m" > "$d/$i" ) & i=$((i+1))
done
wait
i=0; for n in "$@"; do printf '%s %s\n' "$n" "$(cat "$d/$i" 2>/dev/null)"; i=$((i+1)); done
i=0; for spec in "$@"; do printf '%s %s\n' "${spec%%:*}" "$(cat "$d/$i" 2>/dev/null)"; i=$((i+1)); done
rm -rf "$d"
}
# sendwait <sid> <prompt> [seq] -> blocks until that worker's turn ENDS (~10 min ceiling
@@ -116,27 +189,46 @@ spawn_workers() {
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy only for a claude
# worker spawn_worker handed back (hooks vetted); hook-less workspaces and other modes
# resolve on flapping idle: markers instead (§5.5).
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
# implement the status contract is the one case that LOOKS like claude but is not:
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
# is still answering, and the re-wait below then resolved in 0 ms with
# `signal:"idle"` on a turn that had another three minutes to run (measured).
# A wait named after the end of a turn should only end with the turn, or with
# the worker. ⚠️ This is also what makes a wrong mode LOUD: the modes that
# cannot deliver `stop` answer 400 (before writing anything) instead of
# resolving on a flap, which is the answer that sends you to markers (§5.5).
body=$(jq -nc --arg p "$p" --arg c "$CID-$sid" --argjson s "$seq" \
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:true,waitTimeout:20000}')
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:"stop,exit",waitTimeout:20000}')
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")")
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
fi
printf '%s\n' "$r"
}
# last_text <sid> [prev] -> that worker's last assistant message. Polled, because the
# transcript write LAGS the stop signal, and "some text exists" is not "THIS turn's
# text exists": right after a SECOND turn on the same worker the endpoint still serves
# last_text <sid> [prev] -> that worker's last assistant message (claude, codex and
# deepseek write a real transcript; the other modes have none, so read the terminal
# instead -- §5.4). Polled, because the transcript write LAGS the stop signal, and
# "some text exists" is not "THIS turn's text exists": right after a SECOND turn on the same worker the endpoint still serves
# the previous answer for a beat (observed live). When reading consecutive turns, pass
# the previous answer as [prev]: the poll then holds out for text that differs from it,
# falling back to whatever it last saw if the budget runs dry, so an honestly repeated
@@ -155,4 +247,4 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.19.0
CODEMAN_PREAMBLE=1.22.0
+40 -19
View File
@@ -237,7 +237,10 @@ minutes, never retry the credential.
flushed slightly *after* the `stop` hook fires, so a read taken the instant the wait
returns is too early (verified live: empty on the first call, full prose seconds later).
It is also `""` before the worker's first completed turn, and permanently `""` for
`shell`, `opencode`, `gemini`, `antigravity` and `pi`, which write no Claude transcript.
`shell`, `opencode`, `gemini`, `antigravity`, `pi`, `grok` and `omp`, which write no transcript at
all. `deepseek` is NOT one of those — it is read from `$DSH_HOME/sessions/**` and lags
for the same reason claude does (the harness finalizes the assistant message just after
it reports `idle`), so poll it the same way.
**Fix** Poll it, bounded (10 tries, 1 s apart). If it is still empty on a hook-less mode,
that is expected, not a failure: read `terminal?tail=` and strip ANSI instead.
@@ -279,7 +282,9 @@ than into an existing checkout.
| start case + session in one call | `POST /api/v1/quick-start` |
| create a session in an arbitrary directory (no case, **no PTY**, id at `.data.session.id`) | `POST /api/v1/sessions`, then `POST /api/v1/sessions/:id/interactive` or `.../shell` to start it, see [Starting a worker](#starting-a-worker) |
| send input | `POST /api/v1/sessions/:id/input` |
| **read a worker's answer** (claude/codex) | `GET /api/v1/sessions/:id/last-response` → `.data.{text,timestamp}`, clean transcript text, no TUI noise. ⚠️ **Poll it**, see [symptom 7](#7-last-response-returns-an-empty-string-right-after-stop) |
| **read a worker's answer** (claude/codex/deepseek) | `GET /api/v1/sessions/:id/last-response` → `.data.{text,timestamp}`, clean transcript text, no TUI noise. ⚠️ **Poll it**, see [symptom 7](#7-last-response-returns-an-empty-string-right-after-stop) |
| read the whole conversation | `GET /api/v1/sessions/:id/last-response?context=full` → `.data.messages[]`. ⚠️ **Only `{role,text}` is present for every mode.** `kind`/`label` come from claude (`prompt`/`response`), deepseek and the pane parser (which also emit `status`/`tool`) but NOT from codex; `timestamp` from claude and codex but not deepseek/pane; `turn` and `queued:true` (a prompt typed while the agent was working) from claude only. `.data.text` is unchanged by `context=full` — it stays the last assistant message, never `messages[-1]` |
| read the last **answered turn** (claude only) | `GET /api/v1/sessions/:id/last-response?context=turn` → `.data.messages[]` holds every assistant message of the most recent turn that has one (the whole answer, not just its final row); `.data.text` is still the last assistant row. Other modes answer `text` only, with no `messages` |
| read terminal (tail is in **BYTES**, raw ANSI) | `GET /api/v1/sessions/:id/terminal?tail=3000` → `.data.terminalBuffer`, for *diagnosis* (unsubmitted prompt?), not for reading answers |
| full tmux scrollback (context bomb; post-mortems only) | `GET /api/v1/sessions/:id/terminal?full=1` |
| background agents, one session | `GET /api/v1/sessions/:id/subagents` |
@@ -321,7 +326,7 @@ on signals and markers for exactly this reason.
⚠️ `GET /api/v1/sessions/:id/output` → `.data.textOutput` looks like the obvious read
but stays **empty for interactive tmux-backed sessions** (it is fed only by the legacy
JSON-stream path). Verified empty on live claude and shell sessions. Use
`last-response` for claude/codex answers; only fall back to `terminal?tail=` for
`last-response` for claude/codex/deepseek answers; only fall back to `terminal?tail=` for
hook-less modes, or to diagnose a prompt that was never submitted, and strip ANSI:
```bash
@@ -336,19 +341,20 @@ ESC=$(printf '\033')
`POST /api/v1/quick-start` body (all optional):
`{"caseName":"worker-1","mode":"claude","sessionName":"w9-worker","effort":"high"}`
, `mode` ∈ `claude|shell|opencode|codex|gemini|antigravity|pi`; response is
, `mode` ∈ `claude|shell|opencode|codex|gemini|antigravity|pi|grok|deepseek|omp`; response is
`.data.{sessionId, caseName, casePath}`. Creates the case directory (a real directory
on the user's disk) if missing, do not retry it in a loop, and remember the name.
⚠️ A `mode` whose CLI is **not installed on the server** fails the spawn with
`OPERATION_FAILED`; it never falls back to claude. Probe first whenever you did not pick
the mode yourself: `GET /api/v1/claude/status`, `GET /api/v1/opencode/status`,
`GET /api/v1/codex/status`, `GET /api/v1/gemini/status`, `GET /api/v1/antigravity/status`
and `GET /api/v1/pi/status` each return `.data.{available, path}` (no session needed).
Pi's also carries `.data.version`, because `pi` is a short generic name that an unrelated
binary on `$PATH` can shadow: the resolver rejects one whose `--version` is not
semver-shaped, so `available:false` there can mean "a different `pi` is in front" rather
than "nothing is installed". `shell` has no CLI to probe.
`GET /api/v1/codex/status`, `GET /api/v1/gemini/status`, `GET /api/v1/antigravity/status`, `GET /api/v1/grok/status`, `GET /api/v1/deepseek/status`,
`GET /api/v1/pi/status` and `GET /api/v1/omp/status` each return `.data.{available, path}` (no session needed).
Pi's, grok's and OMP's also carry `.data.version`, because `pi` is a short generic name,
`grok` is a name with npm squatters, and `omp` is a similarly short name, so an unrelated
binary on `$PATH` can shadow any of them: the resolver rejects one whose `--version` is
not version-shaped, so `available:false` there can mean "a different program of the same
name is in front" rather than "nothing is installed". `shell` has no CLI to probe.
⚠️ **Branch on `.success` before reading `.data.sessionId`.** On any failure the field
is absent, `jq -r` prints the literal string `null`, and every later call then targets
@@ -359,6 +365,15 @@ the global 50, or the per-user 25 in multi-user mode, never the waiter cap),
`CONFLICT`, `OPERATION_FAILED` and `INVALID_INPUT`. None of them are retryable in a
loop.
⚠️ A case directory quick-start **creates** for you is labelled agent-created (a
`.codeman-agent-case.json` marker, written because the §0 preamble sends
`X-Codeman-Agent-Origin`), which is what lets the user find it afterwards:
`GET /api/v1/cases/agent-created` returns `.data.cases[]` of
`{name, path, createdAt, createdBy, parentSessionId, inUse, modifiedAt}`, newest first,
read-only, scoped to the caller's own case space. Report it when you finish; deleting is
`DELETE /api/v1/cases/:name` and is the user's call by name ([§5.14](verbs.md#514-clean-up)).
A directory that already existed is never labelled.
⚠️ `caseName` resolves through the linked-cases registry first, so a name that happens
to match a case the user linked in lands in that **real repo**, not a fresh scratch
directory. Pick distinctive scratch names, and use a linked name deliberately when you
@@ -462,10 +477,10 @@ Quirks that will bite you:
session answers with an empty timeline rather than a 404.
- ⚠️ **`active-tools` proves presence, never absence.** It is fed by the BashToolParser,
which reads Claude's rendered `● Bash(…)` lines, and `_processExpensiveParsers`
returns early for every external CLI mode (`session.ts:2136`), so it is permanently
`[]` on `opencode`/`codex`/`gemini`/`antigravity`/`pi`. ⚠️ **`shell` is NOT one of those**
(`isExternalCliMode`, `session.ts:165-167`, lists only those five), so the parser does
run on a shell worker, and `TEXT_COMMAND_PATTERN` (`bash-tool-parser.ts:88`) matches
returns early for every external CLI mode (`session.ts:~2225`), so it is permanently
`[]` on `opencode`/`codex`/`gemini`/`antigravity`/`pi`/`grok`/`deepseek`/`omp`. ⚠️ **`shell` is NOT one of those**
(`isExternalCliMode`, `session.ts:176-187`, lists only those seven), so the parser does
run on a shell worker, and `TEXT_COMMAND_PATTERN` (`bash-tool-parser.ts:89`) matches
bare `tail|cat|head|less|grep|watch|multitail <path>` lines with no `● Bash(` wrapper:
a shell worker running `cat build.log` really does populate this. In practice it stays
empty for most shell work. It also never sees non-Bash
@@ -639,10 +654,14 @@ block, so a linked case or a raw `workingDir` had no hooks at all. `POST
session-create path installs hooks regardless of how the directory got there. See
[symptom 8](#8-send-and-wait-resolves-instantly-with-signalidle-and-the-answer-is-last-turns).
Default `until` set: `stop,idle,exit`. On non-claude modes the server silently drops
`stop`/`blocked` from the *default* set (echoed back as `wait.until`, e.g.
Default `until` set: `stop,idle,exit`. On modes with no hook signals the server silently
drops `stop`/`blocked` from the *default* set (echoed back as `wait.until`, e.g.
`["idle","exit"]` on shell); requesting them *explicitly* there is a 400 naming the
mode. ⚠️ That 400 is about **mode**, so a hooks-less *claude* session accepts
mode. ⚠️ `deepseek` is not one of those: its harness reports its own lifecycle, so it
keeps the full default set and accepts an explicit `until=stop`. ⚠️ For dsh the answer is
per-SESSION rather than per-mode — a session created with `statusReporting: false` has no
bridge, and an explicit `until=stop` there is a 400 naming that setting. ⚠️ That 400 is
otherwise about **mode**, so a hooks-less *claude* session accepts
`until=stop` happily and then never resolves it. ⚠️ On hook-less modes the lifecycle
signals are also **coarse in practice**: a
short shell command produced **no** `idle` transition within 60 s (verified live), so
@@ -790,8 +809,10 @@ for environment and setup problems.
| `CODEMAN_MUX` unset but you seem to be in a session | remote-SSH case: the env vars are not exported there. Fail closed, refuse to act |
| connection refused from inside a container | a loopback-bound server is unreachable from a container, and `CODEMAN_DOCKER_BRIDGE_HOOKS=1` does **not** fix that: it opens a hooks-only listener, so hook events start flowing but `/api/v1/*` stays refused. Driving the API from inside a Docker case needs a reachable bind (an operator decision); report it, don't retry |
| wait routes 404 on a valid session id | read the `.error` text: a `Route ...` prefix means the server predates the wait endpoints (< 1.13.0; a dev build can serve them while reporting an older version, so probe, never version-compare), poll `terminal?tail=` and say so. `Session ... not found` means your id is wrong, not the server |
| wait on `stop` never resolves | non-claude mode, or hooks not reaching the server (Docker/remote), or a case created by Codeman < 1.13.0 against an `--https` install (its hook curls lacked `-k` and TLS-failed silently; a 1.13.0+ server rewrites them the next time a session starts in that case). Use markers or `idle,exit` |
| new claude worker ignores its first prompt | it was showing the first-run trust dialog and Codeman's auto-accept did not fire (it is bounded by a 90 s window and an attempt cap); use the readiness recipe in SKILL.md, wait for `shift+tab` first, accept the dialog only as the bounded fallback |
| wait on `stop` never resolves | a mode with no hook signals, or hooks not reaching the server (Docker/remote), or a case created by Codeman < 1.13.0 against an `--https` install (its hook curls lacked `-k` and TLS-failed silently; a 1.13.0+ server rewrites them the next time a session starts in that case). Use markers or `idle,exit` |
| wait on `stop` never resolves, on a **dsh** worker whose pane clearly finished | that profile does not implement the harness's supervisor contract, which Codeman cannot detect at request time (an unrecognized profile is treated as launchable on purpose). The wait is accepted and then times out. Drive that worker with markers, or switch to a profile that reports — `@deepseek-harness-tui/dsh-tui` does |
| new claude worker ignores its first prompt | it was showing the first-run trust dialog and Codeman's auto-accept did not fire (it is bounded by a 90 s window and a keystroke cap); use the readiness recipe in SKILL.md, wait for `shift+tab` first, answer the dialog only as the bounded fallback |
| a brand-new claude worker's pane is DEAD (`status 1`) seconds after the spawn | something pressed Enter at the first-run trust dialog. Since claude-cli 2.1.252 its options are unnumbered, reversed, and the highlighted default is `No, exit`, so a blind `\r` — an up-front Enter, or a task prompt typed into the dialog — quits the CLI. Answer it by reading the `❯` marker off `terminal?full=1` and arrowing onto `Yes, I trust this folder` first: `_accept_trust` in the §0 preamble |
| readiness burns its whole budget, then the worker answers fine anyway | you matched `bypass`, which is the statusline of ONE permission mode. Codeman spawns `--dangerously-skip-permissions` by default, but the server's `claudeMode` setting also has `auto` (`auto mode on`), `allowedTools` and `normal` (both `don't ask on`), and the effective per-session value is not exposed on `GET /api/v1/sessions/:id`. Match **`shift+tab`** instead: every mode's status bar ends `(shift+tab to cycle)` (measured per mode against claude-cli 2.1.226). Expect `blocked` signals mid-turn on the non-default modes |
| ANSI escapes survive the strip pipeline | `sed -e 's/\x1b…'` on macOS: `\x1b` is GNU-only, BSD sed matches nothing and strips nothing. Use the `ESC=$(printf '\033')` form above |
| `wait-output` times out although the pane shows the text | multi-word match against a TUI screen; the stream has no spaces there, match one token |
+2 -2
View File
@@ -56,7 +56,7 @@ own head: the worker enforcing the cap is the one who has to be told about it.
| synchronize on end of turn | HTTP `wait until=stop` (fires for message-initiated turns too, verified live) |
| liveness / death check | HTTP `wait?until=exit` |
| interrupt a running turn (break-glass) | HTTP input, a bare `\x1b` with no `\r` |
| non-claude modes (`shell`/`opencode`/`codex`/`gemini`/`antigravity`/`pi`) | HTTP only (no other CLI has messaging) |
| non-claude modes (`shell`/`opencode`/`codex`/`gemini`/`antigravity`/`pi`/`grok`/`deepseek`/`omp`) | HTTP only (no other CLI has messaging) |
| delete | HTTP, via SKILL.md's `delete_session` guard |
## Availability: probe, never assume
@@ -347,7 +347,7 @@ Without a break-glass, a pair with a bad brief is a token bonfire with no off sw
### Mixed fleets: the pairing matrix
Non-claude workers (`shell`, `opencode`, `codex`, `gemini`, `antigravity`, `pi`) cannot be peers
Non-claude workers (`shell`, `opencode`, `codex`, `gemini`, `antigravity`, `pi`, `grok`, `deepseek`, `omp`) cannot be peers
at all; no other CLI has this feature. Their tasks route over HTTP, and you never mention
messaging in their briefs. The claude half of the fleet can use messaging among itself,
subject to the namespace rule: **messaging works between two sessions that share one
+71 -17
View File
@@ -21,7 +21,7 @@ by sourcing the preamble file the §0 bootstrap wrote, and checking its version
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.18.3 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
```
Do **not** re-paste the preamble body into each call. Sourcing it is what retires the
@@ -69,9 +69,14 @@ SEQ=1 # $CID is the fixed literal from the preamble; never rebuild
# plus a two-marker screen match in session-trust-dialog.ts), not the output stream.
# It still misses two ways, and both leave the dialog up until someone answers it:
# it only scans in the first 90 s after the pane started (TRUST_DIALOG_WINDOW_MS),
# and it gives up after 3 Enter presses (TRUST_DIALOG_MAX_ATTEMPTS). So: composer
# marker first, dialog only as the bounded fallback (a blind Enter up front would
# land in an already-ready composer).
# and it gives up after 6 keystrokes (TRUST_DIALOG_MAX_ATTEMPTS). So: composer
# marker first, dialog only as the bounded fallback.
# ⚠️ The dialog is NOT answered with Enter. Since claude-cli 2.1.252 the options
# lost their numbers, swapped places, and the highlighted one is `No, exit`, so a
# blind \r quits the CLI and the pane is dead seconds after the spawn (measured).
# _accept_trust (§0 preamble) reads the ❯ marker off the rendered pane, arrows onto
# `Yes, I trust this folder`, re-reads to confirm the move landed, and only then
# presses Enter.
# Stage 1 is SHORT on purpose: an already-trusted case matches in <1 s, while a
# virgin case can never pass it (the dialog is up) and always pays it in full,
# the long budget belongs to stage 3, after the dialog is answered.
@@ -92,13 +97,8 @@ done
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
fi
_accept_trust "$SID" # reads the marker and steers; never a blind \r. Own clientId,
# so it spends none of $SEQ's numbers.
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
fi
@@ -106,8 +106,8 @@ if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
# stage 4, mode-agnostic and bounded: answering a trivial prompt IS readiness.
# COSTS THE WORKER ONE BILLED TURN, so it only runs when the fast marker missed.
# Split token (the typed line echoes into the stream) and unique per call. Must stay
# AFTER the dialog fallback: free text plus \r into a trust dialog still up answers
# it blind, the same footgun as an up-front Enter.
# AFTER the dialog fallback: the select widget swallows the text and the \r answers
# whatever is highlighted, which on a live dialog is `No, exit`.
TOK="${RANDOM}_$$"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"reply with the word READY immediately followed by _'"$TOK"' and nothing else\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
@@ -188,7 +188,7 @@ for _ in $(seq 1 10); do
done
printf '%s\n' "$TXT"
# (.data is {text,timestamp}; text is also "" before the first completed turn and
# always "" for shell/opencode/gemini/antigravity/pi, which have no transcript, use
# always "" for shell/opencode/gemini/antigravity/pi/grok/omp, which have no transcript, use
# the terminal tail there, and here only to diagnose an unsubmitted prompt.)
# 6. clean up: exact id, own list only, through the fail-closed preamble helper
@@ -198,6 +198,59 @@ delete_session "$SID"
Increment `SEQ` for every *new* input to the same worker. Reuse the same `SEQ` only to
re-ask about the same delivery (the duplicate-wait loop above).
## Flow 1b: DeepSeek Harness worker, end to end
A `deepseek` worker is driven with the same four verbs as a claude one, because the
harness reports its own lifecycle: its `stop` is a real end-of-turn signal, and its
answer comes from a real transcript. The differences are all at the edges.
```bash
# 0. Is there anything to spawn? `available` is the binary, `runnable` is a profile
# that can drive a pane -- dsh ships only web/headless, so the two differ.
"${CURL[@]}" "$API/api/v1/deepseek/status" | jq -c '{available:.data.available,runnable:.data.runnable,profile:.data.defaultProfile}'
# 1. Spawn. `deepSeekConfig` is optional: an absent profile picks the first
# pane-capable one, and an absent permissionMode leaves the harness on its own
# workspace-write default, which still ASKS before it acts.
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"dsh-worker","mode":"deepseek","deepSeekConfig":{"permissionMode":"danger-full-access"}}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
[ -n "$SID" ] || { jq -c '{error, errorCode}' <<<"$Q"; exit 1; } # OPERATION_FAILED = no runnable profile
CREATED+=("$SID")
# 2. Readiness, and ONLY readiness. ⚠️ Do not use the stop signal for this: the
# harness reports idle at BOOT, ~300 ms before the composer paints.
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=❯' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000' \
| jq -e '.data.wait.matched' >/dev/null || { echo "no composer"; delete_session "$SID"; exit 1; }
# 3. Task it. Identical to a claude worker, including the \r and the (clientId, seq).
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"Read calc.py and tell me in one sentence whether add() is correct.\r","useMux":true,"clientId":"codeman-dsh-1","seq":1,"wait":"stop,exit","waitTimeout":300000}' \
| jq -c '{delivered:.data.delivered,signal:.data.wait.signal,timedOut:.data.wait.timedOut}'
# 4. Read it. From $DSH_HOME/sessions/**, not the pane -- scraping a dsh pane returns
# its ASCII-art splash. Poll: the harness finalizes the message just after it
# reports idle. Two answers are not the model's words and say so:
# "Turn error: …" (the provider or harness failed) and "Turn ended: …" (early stop).
for _ in $(seq 1 15); do
TXT=$("${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
# 5. Full conversation, if you need the tool calls too:
# "${CURL[@]}" "$API/api/v1/sessions/$SID/last-response?context=full" | jq -r '.data.messages[]|"[\(.label)] \(.text)"'
delete_session "$SID"
```
⚠️ **`wait:"stop,exit"`, not `wait:true`.** The default set also carries `idle`, which
for an external CLI is inferred from output stabilization: a dsh TUI that repaints
rarely reads as idle mid-turn, and a wait carrying `idle` then resolves in 0 ms on a
turn with minutes left to run (measured). The same reason the preamble's `sendwait`
asks for `stop,exit` on every mode.
## Flow 2: shell worker, marker-synchronized
`shell` sessions have no hooks (`stop`/`blocked` are a 400 there), and their lifecycle
@@ -474,9 +527,10 @@ done
circuit breaker, which exists to stop a worker that crashes on every start from being
restarted in a loop; clearing it unasked re-arms that loop.
- Then run **Flow 1's readiness stages 1-3** on each SID. A path claude has never been
run in shows the trust dialog, and typing your task into a dialog answers it blind and
loses the task. Stages 1-3 cost no turn; stage 4, if it fires, costs that worker one
billed turn.
run in shows the trust dialog, and typing your task into it does not just lose the
task: the select widget swallows the text and the trailing `\r` answers the
highlighted option, which since claude-cli 2.1.252 is `No, exit`. Stages 1-3 cost no
turn; stage 4, if it fires, costs that worker one billed turn.
### 4. Hand out the tasks: markers, not send-and-wait
+97 -25
View File
@@ -152,17 +152,46 @@ It is **decoration, and resolved rather than trusted**, so treat it accordingly:
### 5.2 Readiness
A new session reports `idle` before its CLI has spawned, and a brand-new case shows a
**dsh workers first**, because their trap is the opposite of claude's: they have no
trust dialog and boot straight into a composer (`❯`, matched `from=buffer`), but the
harness reports `idle` — which reaches you as a `stop` signal — about 300 ms BEFORE that
composer paints (measured 2.26 s vs 2.56 s after spawn, twice). So the signal that means
"this worker finished its turn" is also the first thing it emits at boot, and a
send-and-wait fired straight after `quick-start` resolves on it, reports a turn that
never ran, and leaves the prompt in a pane that was not yet taking input. Wait for the
composer, not for the signal; `spawn_worker` does exactly that, and by the time it
returns the boot edge is spent (signals are edge-triggered, so nothing can catch it
later). A profile whose composer is not `❯` needs `DSH_READY_MARK` set to whatever it
does draw.
For claude: a new session reports `idle` before its CLI has spawned, and a brand-new case shows a
**trust dialog** first, so neither "wait for idle" nor "wait for ❯" means ready (the
trust dialog contains `❯` too, observed live). Codeman auto-accepts that dialog
itself, reliably enough that stage 1 usually just works: `_maybeAcceptTrustDialog()`
reads the **rendered pane** via `capturePaneText()` rather than the arriving chunk
(the per-chunk `includes()` version could never match, because tmux repaints the row
with cursor-forward escapes in place of spaces, and it is documented in-source as the
historical bug). The remaining miss modes are structural: the auto-accept only runs
inside a 90 s window after interactive start and gives up after 3 attempts. So keep
the dialog handling as a bounded fallback, and never send a blind Enter up front (if
auto-accept already fired, it lands in the composer).
historical bug).
⚠️ **The answer is no longer "press Enter".** Claude Code 2.1.252 dropped the option
numbers, reversed the two options, and highlights the one that quits:
```
❯ No, exit
Yes, I trust this folder
Enter to confirm · Esc to cancel
```
so a blind `\r` answers *exit*: the pane is dead (`Pane is dead (status 1)`) about six
seconds after the spawn, measured on a fresh case. Read the marker off the rendered
pane (`GET .../terminal?full=1`), send `ESC [ B` while it sits on `No, exit`, re-read,
and press Enter only once the marker is on the trust option. `_accept_trust` in the
§0 preamble is exactly that, and `trustDialogNextKey()` is the server-side twin.
The remaining miss modes are structural: the auto-accept only runs inside a 90 s window
after interactive start and gives up after 6 keystrokes. So keep the dialog handling as
a bounded fallback, and never send a blind Enter up front — landing in an already-ready
composer only wastes a turn, landing in this dialog ends the worker.
Stage 1 is short on purpose: an already-trusted case matches `shift+tab` in under a
second, while a case still showing the dialog cannot pass stage 1 at all and always
@@ -220,14 +249,11 @@ SEQ=1 # $CID came from the §0 preamble; do NOT rebuild it from $$
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
# composer never appeared, so the trust dialog is probably still up; accept it once
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
fi
# Composer never appeared, so the trust dialog is probably still up. NEVER a blind
# Enter here: the highlighted option is "No, exit". _accept_trust (§0 preamble) reads
# the marker off the pane, arrows onto the trust option, re-reads, then confirms. It
# carries its OWN clientId, so it spends none of $SEQ's numbers.
_accept_trust "$SID"
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
fi
@@ -236,8 +262,10 @@ if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
# of a broken worker, and answering is proof that it works. Split the token (your
# keystrokes echo into the stream) and keep it unique per call. This costs the worker
# one billed turn, so it runs only after the fast path missed. It must stay AFTER
# stage 2, which is the only thing that clears the trust dialog: free text plus \r
# into a dialog still up answers it blind, the same footgun as the up-front Enter.
# stage 2, which is the only thing that clears the trust dialog: the typed text is
# swallowed by the select widget and the \r then answers whatever is highlighted,
# which since 2.1.252 is "No, exit" -- the same footgun as the up-front Enter, except
# that it kills the worker rather than wasting a turn.
TOK="${RANDOM}_$$"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"reply with the word READY immediately followed by _'"$TOK"' and nothing else\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
@@ -341,17 +369,35 @@ If the loop exhausts its cap, do not keep looping: read the terminal, report wha
see, and remember that a still-typed-but-unsubmitted prompt (missing `\r`) can only be
recovered by submitting it with `{"input":"\r"}`.
⚠️ `stop` and `blocked` fire for `claude` sessions only (they are Claude Code hooks,
and only when the workspace actually has them, see [§5.1](#51-where-to-spawn)). On
`shell`/`opencode`/`codex`/`gemini`/`antigravity`/`pi`, requesting them explicitly is a
⚠️ `stop` and `blocked` fire for `claude` sessions (they are Claude Code hooks, and
only when the workspace actually has them, see [§5.1](#51-where-to-spawn)) **and for
`deepseek`** — the one external CLI that reports its own lifecycle, so its `stop` is a
real end-of-turn signal rather than a guess. On
`shell`/`opencode`/`codex`/`gemini`/`antigravity`/`pi`/`grok`/`omp`, requesting them explicitly is a
400, and lifecycle transitions there are coarse (a short shell command may emit **no**
`idle` transition at all, verified live), so synchronize those with markers.
⚠️ A dsh session can still refuse them for a per-SESSION reason: `statusReporting:
false` at create time disarms the bridge, and an explicit `until=stop` is then a 400
naming that setting. And a `stop` that is *accepted* is not proof it will ever fire —
whether the installed profile implements the supervisor contract cannot be known at
request time, so a non-conforming one accepts the wait and times out on it. One timeout
on a dsh worker whose pane clearly finished identifies that profile; switch it to
markers.
### 5.4 Read the answer
For `claude` and `codex` workers this is the read path: `last-response` returns the
agent's final message as clean text, taken from the transcript rather than the screen,
so it carries none of the TUI's box-drawing or repaint noise.
For `claude`, `codex` and `deepseek` workers this is the read path: `last-response`
returns the agent's final message as clean text, taken from the transcript rather than
the screen, so it carries none of the TUI's box-drawing or repaint noise.
⚠️ For `deepseek` it reads `$DSH_HOME/sessions/**`, and reading it is the ONLY way to
get that answer: dsh-TUI paints a full-screen splash, so scraping its pane returns the
ASCII-art logo (that is what `last-response` itself used to return for dsh). Two dsh
answers are not the model's words and say so: `Turn error: …` (the provider or the
harness failed the turn) and `Turn ended: …` (an early stop such as `max-tokens`). A
turn still streaming reads back as the partial answer so far, so a non-empty read is
not by itself proof the turn ended — that is what the `stop` signal is for.
```bash
for _ in $(seq 1 10); do # the transcript write LAGS the stop signal
@@ -361,7 +407,17 @@ done
printf '%s\n' "$TXT"
```
`.data` is `{text, timestamp}`. ⚠️ **On a hook-less workspace this reads the PREVIOUS
`.data` is `{text, timestamp}`. Add `?context=full` for the whole conversation in
`.data.messages[]`. ⚠️ **The four readers do not emit the same fields — only `{role, text}`
is guaranteed.** `kind`/`label` come from claude (`prompt`/`response`), deepseek and the pane
parser (the last two also emit `status`/`tool`), but **not** from codex; `timestamp` comes
from claude and codex but not from deepseek or the pane parser. A claude worker additionally
carries `turn` (a run of same-speaker messages inside one `turn` is one utterance split into
segments, not separate exchanges) and `queued: true` on a prompt the user typed while the
agent was still working. Filter on `role`, not on `kind`, unless you know the mode.
`.data.text` does not change under `context=full`: it stays the
last **assistant** message, so never read it as `messages[-1]`, which can be a prompt.
⚠️ **On a hook-less workspace this reads the PREVIOUS
turn.** `last-response` returns whatever the transcript last flushed, so it is only as
correct as your end-of-turn signal: pair it with a `stop` signal or a marker, never
with a bare `idle` ([§5.1](#51-where-to-spawn)). ⚠️ **Poll it, do not read it once.** `text` is written
@@ -369,9 +425,10 @@ from the transcript file, which is flushed slightly *after* the `stop` hook fire
single read taken the instant send-and-wait returns comes back `""` even though the
turn finished (verified live: empty on the first call, full text seconds later). `text`
is also `""` before the worker's first completed turn, and always `""` for modes with
no transcript (`shell`, `opencode`, `gemini`, `antigravity`, `pi`; the first four
no transcript (`shell`, `opencode`, `gemini`, `antigravity`, `pi`, `grok`, `omp`; the first four
verified live, pi from the same source path), which is
why the loop above is bounded rather than open-ended. Fall back to the terminal buffer
why the loop above is bounded rather than open-ended. A dsh worker lags too, for its own
reason: the harness finalizes the assistant message just after it reports `idle`. Fall back to the terminal buffer
there, tail in **bytes** (`textOutput` in `GET .../output` stays empty for interactive
sessions; don't use it):
@@ -454,7 +511,7 @@ turn), and both better than diffing terminal samples:
```
⚠️ `active-tools` is parsed out of Claude's own output format, so it is **empty for
`opencode`/`codex`/`gemini`/`antigravity`/`pi`** (those parsers are skipped wholesale) and
`opencode`/`codex`/`gemini`/`antigravity`/`pi`/`grok`/`deepseek`/`omp`** (those parsers are skipped wholesale) and
in practice empty for `shell`. Source-verified, not measured live.
Only if neither helps: sample `terminal?tail=` twice a few seconds apart. A changing
@@ -674,6 +731,21 @@ Deleting a session ends the agent and its pane. It does **not** remove:
it, and ask before running `git worktree remove`, which discards uncommitted work
inside it.
Those case directories are **labelled** rather than left anonymous. A directory
`quick-start` creates for a spawn carrying the preamble's `X-Codeman-Agent-Origin`
header gets a `.codeman-agent-case.json` marker, which is what puts it in the web UI's
agent-case cleanup list (Add Case → Manage) and in:
```bash
"${CURL[@]}" "$API/api/v1/cases/agent-created" | jq -r '.data.cases[] | "\(.name)\t\(.createdAt)\tinUse=\(.inUse)"'
```
Read-only, scoped to the user's own case space, and `inUse` is true while a live
session is still working in that directory. Report that list when you finish a run
with workers, so the user knows exactly what to sweep; the deletion is still theirs to
ask for by name. Only a directory Codeman **created** is ever labelled, so a linked
case, a cloned repo or a worktree never appears there.
Confirm cleanup with `GET /api/v1/sessions`, never with `/api/v1/sessions/unified`
(that one folds in transcript history from the whole machine and will keep showing
your worker forever).
+183
View File
@@ -0,0 +1,183 @@
/**
* @fileoverview The marker file that records a case directory as one Codeman scaffolded
* FOR an agent-spawned session, so scratch worker workspaces can be told apart from the
* user's real projects long after the sessions that created them are gone.
*
* Why a file in the case directory rather than a central registry in `~/.codeman`:
* the thing being labelled is a directory on the user's disk, and the label has to
* survive everything that can happen to Codeman's own state (a wiped data dir, a
* different instance, a hand-moved case). A registry would also need stale-entry
* pruning and owner scoping of its own, while a marker is deleted by the same `rm -rf`
* that deletes the case, and is discoverable by a user who just runs `ls -a`.
*
* ⚠️ Written ONLY on the path that CREATES the directory (`POST /api/quick-start`'s
* `!existsSync` branch). A linked case, a cloned repo, a git worktree or any other
* pre-existing directory must never be labelled agent-created: the label drives a
* cleanup affordance, and mislabelling someone's repo there is the one failure mode
* that costs real work. `POST /api/sessions` takes an existing `workingDir` and so
* writes no marker at all, by construction.
*
* ⚠️ Reading is strict and total: anything that does not parse as a version-1 marker
* (truncated write, hand-edited junk, a user's unrelated file of the same name) reads
* as "not agent-created" rather than as a partially-trusted entry. A marker is
* metadata; deleting the file is the supported way to adopt a scratch case as a real
* one, which is what the `note` field written into it tells the user.
*/
import { readFile, writeFile } from 'node:fs/promises';
import { join } from 'node:path';
/** Marker filename inside the case directory. Dot-prefixed so it stays out of the way. */
export const AGENT_CASE_MARKER_FILE = '.codeman-agent-case.json';
/** Current marker schema version. A marker of any other version reads as absent. */
export const AGENT_CASE_MARKER_VERSION = 1;
/**
* Origin recorded when a create request carried a resolvable spawning session but no
* explicit origin of its own (an agent driving the API by hand, or an older copy of
* the skill). Nothing in the browser UI sets lineage, so this really does mean "another
* session spawned this", not "a human clicked Run".
*/
export const AGENT_ORIGIN_SPAWNED_BY_SESSION = 'agent-session';
/** Origin the packaged agent skill sends on its shared curl invocation. */
export const AGENT_ORIGIN_CODEMAN_SKILL = 'codeman-skill';
/** Longest accepted origin token (the value is echoed into the UI and the marker). */
const MAX_ORIGIN_LENGTH = 32;
/** Longest accepted free-text field read back out of a marker. */
const MAX_MARKER_FIELD_LENGTH = 200;
/** Lowercase token: what an origin may look like on the wire and on disk. */
const AGENT_ORIGIN_PATTERN = /^[a-z0-9][a-z0-9._-]*$/;
/** Explains the file to whoever finds it in their case directory. */
const MARKER_NOTE =
'Created by a Codeman agent worker (see the Manage tab in Add Case). ' +
'Delete this file to keep the case out of the agent-case cleanup list; ' +
'deleting the whole directory removes the case.';
/**
* What a case directory records about the agent spawn that created it.
* Every field beyond `version`/`createdAt`/`createdBy` is decoration for the cleanup UI.
*/
export interface AgentCaseMarker {
version: typeof AGENT_CASE_MARKER_VERSION;
/** ISO timestamp of the spawn that created the directory. */
createdAt: string;
/** Who asked: `codeman-skill`, `agent-session`, or another caller's own token. */
createdBy: string;
/** Full id of the session that spawned the worker, when one resolved. */
parentSessionId?: string;
/** That session's display name at spawn time, so the user recognises it later. */
parentSessionName?: string;
/** Run mode the worker was started in (`claude`, `deepseek`, …). */
mode?: string;
/** Owner the case was created for, in multi-user mode. */
owner?: string;
}
/**
* Validate an origin token coming off the wire (`agentOrigin` body field or the
* `X-Codeman-Agent-Origin` header). Returns `undefined` for anything that is not a
* short lowercase token — the value reaches the UI and a JSON file, so it is
* allowlisted rather than escaped at each use.
*/
export function normalizeAgentOrigin(raw: unknown): string | undefined {
if (typeof raw !== 'string') return undefined;
const value = raw.trim().toLowerCase();
if (!value || value.length > MAX_ORIGIN_LENGTH) return undefined;
return AGENT_ORIGIN_PATTERN.test(value) ? value : undefined;
}
/** Trim an optional free-text marker field to something safe to store and render. */
function normalizeField(raw: unknown): string | undefined {
if (typeof raw !== 'string') return undefined;
const value = raw.trim();
return value ? value.slice(0, MAX_MARKER_FIELD_LENGTH) : undefined;
}
/**
* Build a marker from a spawn's details. Pure, so the route can hand it straight to
* the writer and the tests can assert on the shape without touching a disk.
*/
export function buildAgentCaseMarker(input: {
createdBy: string;
createdAt?: Date;
parentSessionId?: string;
parentSessionName?: string;
mode?: string;
owner?: string;
}): AgentCaseMarker {
const marker: AgentCaseMarker = {
version: AGENT_CASE_MARKER_VERSION,
createdAt: (input.createdAt ?? new Date()).toISOString(),
createdBy: normalizeAgentOrigin(input.createdBy) ?? AGENT_ORIGIN_SPAWNED_BY_SESSION,
};
const parentSessionId = normalizeField(input.parentSessionId);
const parentSessionName = normalizeField(input.parentSessionName);
const mode = normalizeField(input.mode);
const owner = normalizeField(input.owner);
if (parentSessionId) marker.parentSessionId = parentSessionId;
if (parentSessionName) marker.parentSessionName = parentSessionName;
if (mode) marker.mode = mode;
if (owner) marker.owner = owner;
return marker;
}
/**
* Parse marker JSON. Returns `null` for anything that is not a well-formed version-1
* marker, including a valid-JSON object of the wrong shape — see the strictness note
* in the file header.
*/
export function parseAgentCaseMarker(raw: string): AgentCaseMarker | null {
let value: unknown;
try {
value = JSON.parse(raw);
} catch {
return null;
}
if (!value || typeof value !== 'object' || Array.isArray(value)) return null;
const record = value as Record<string, unknown>;
if (record.version !== AGENT_CASE_MARKER_VERSION) return null;
const createdAt = normalizeField(record.createdAt);
const createdBy = normalizeAgentOrigin(record.createdBy);
if (!createdAt || !createdBy || Number.isNaN(Date.parse(createdAt))) return null;
return buildAgentCaseMarker({
createdBy,
createdAt: new Date(createdAt),
parentSessionId: normalizeField(record.parentSessionId),
parentSessionName: normalizeField(record.parentSessionName),
mode: normalizeField(record.mode),
owner: normalizeField(record.owner),
});
}
/**
* Write the marker into `casePath`. Best-effort by design: the marker is metadata for
* a later cleanup, and a failed write must never fail the worker spawn that is the
* point of the request. Returns whether it landed.
*/
export async function writeAgentCaseMarker(casePath: string, marker: AgentCaseMarker): Promise<boolean> {
try {
const body = JSON.stringify({ ...marker, note: MARKER_NOTE }, null, 2);
await writeFile(join(casePath, AGENT_CASE_MARKER_FILE), `${body}\n`, 'utf-8');
return true;
} catch {
return false;
}
}
/** Read the marker out of `casePath`, or `null` if there isn't a valid one. */
export async function readAgentCaseMarker(casePath: string): Promise<AgentCaseMarker | null> {
try {
return parseAgentCaseMarker(await readFile(join(casePath, AGENT_CASE_MARKER_FILE), 'utf-8'));
} catch {
return null;
}
}
+299
View File
@@ -0,0 +1,299 @@
/**
* @fileoverview One style vocabulary for everything the `codeman` CLI prints:
* palette, glyphs, the small block helpers (heading/rule/kv), width-aware table
* layout, a stderr spinner and a y/N confirm.
*
* Color detection is chalk's alone. It already honors NO_COLOR, FORCE_COLOR,
* TERM=dumb and TTY-ness, and a second detector here would disagree with it on
* some terminal with no way to tell which one was right.
*
* The layout math is pure and exported separately from anything that touches a
* terminal, which is what lets it be unit-tested with no TTY and reused by
* `utils/dependency-report.ts` while that file stays color-free.
*
* @module cli-style
*/
import chalk, { type ChalkInstance } from 'chalk';
import { createInterface } from 'node:readline';
// Direct import, not the `utils` barrel: the barrel pulls in node-pty and every
// CLI resolver, which a style module has no business loading.
import { stripAnsi } from './utils/regex-patterns.js';
// ─────────────────────────────────────────────────────────────────────────────
// Palette and glyphs
// ─────────────────────────────────────────────────────────────────────────────
/** Semantic roles, mirroring the web UI's status language (green fine, yellow waiting, red blocked). */
export const palette = {
ok: chalk.green,
warn: chalk.yellow,
err: chalk.red,
info: chalk.cyan,
muted: chalk.gray,
emph: chalk.bold,
accent: chalk.magenta,
} as const satisfies Record<string, ChalkInstance>;
/** The glyph vocabulary the CLI already used, in one place. */
export const GLYPH = {
ok: '✓',
fail: '✗',
warn: '⚠',
idle: '○',
dot: '●',
arrow: '→',
} as const;
/** Spinner frames (braille, one cell wide in every terminal we support). */
export const SPINNER_FRAMES = ['⠋', '⠙', '⠹', '⠸', '⠼', '⠴', '⠦', '⠧', '⠇', '⠏'] as const;
/** What a line is reporting, independent of how it is painted. */
export type Tone = 'ok' | 'warn' | 'err' | 'idle' | 'info';
const TONE_GLYPH: Record<Tone, string> = {
ok: GLYPH.ok,
warn: GLYPH.warn,
err: GLYPH.fail,
idle: GLYPH.idle,
info: GLYPH.dot,
};
const TONE_STYLE: Record<Tone, ChalkInstance> = {
ok: palette.ok,
warn: palette.warn,
err: palette.err,
idle: palette.muted,
info: palette.info,
};
/** Glyph for a tone. Pure, so the mapping is testable without a terminal. */
export function glyphFor(tone: Tone): string {
return TONE_GLYPH[tone];
}
/** Paint text in a tone's color. */
export function tint(tone: Tone, text: string): string {
return TONE_STYLE[tone](text);
}
// ─────────────────────────────────────────────────────────────────────────────
// Blocks
// ─────────────────────────────────────────────────────────────────────────────
/** Section heading. The blank line above it is part of the existing block idiom. */
export function heading(text: string): string {
return `\n${palette.emph(text)}`;
}
/** Horizontal rule under a title. */
export function rule(width = 40): string {
return palette.muted('─'.repeat(Math.max(0, width)));
}
/**
* Indented `Label: value` line. `pad` aligns the values of a block by padding
* the label column (including its colon), for blocks whose labels differ in
* length.
*/
export function kv(label: string, value: string, pad = 0): string {
const key = pad > 0 ? padCell(`${label}:`, pad) : `${label}:`;
return ` ${key} ${value}`;
}
// ─────────────────────────────────────────────────────────────────────────────
// Width-aware layout (pure)
// ─────────────────────────────────────────────────────────────────────────────
/** Printed width of a cell: ANSI sequences take no columns. */
export function displayWidth(text: string): number {
return stripAnsi(text).length;
}
export type CellAlign = 'left' | 'right';
/** Pad to `width` columns, measuring by display width so colored cells still align. */
export function padCell(text: string, width: number, align: CellAlign = 'left'): string {
const fill = ' '.repeat(Math.max(0, width - displayWidth(text)));
return align === 'right' ? `${fill}${text}` : `${text}${fill}`;
}
/**
* Pad AFTER the paint, so the fill stays outside the color run and a trailing
* empty column can be trimmed away instead of ending in a reset sequence with
* invisible spaces before it.
*/
export function padStyled(text: string, width: number, paint: (t: string) => string): string {
return `${paint(text)}${' '.repeat(Math.max(0, width - displayWidth(text)))}`;
}
/** Widest cell per column. Short rows count as empty cells, never as narrower columns. */
export function columnWidths(rows: readonly (readonly string[])[]): number[] {
const widths: number[] = [];
for (const row of rows) {
for (let i = 0; i < row.length; i++) {
widths[i] = Math.max(widths[i] ?? 0, displayWidth(row[i] ?? ''));
}
}
return widths;
}
export interface TableOptions {
/** Per-column alignment; missing entries are left-aligned. */
align?: readonly CellAlign[];
/** Spaces between columns. */
gap?: number;
/** Prefix for every row. */
indent?: string;
}
/**
* Lay rows out in columns sized to their widest cell. The last cell of a row is
* never padded, so no line carries trailing whitespace.
*/
export function layoutTable(rows: readonly (readonly string[])[], options: TableOptions = {}): string[] {
const { align = [], gap = 1, indent = '' } = options;
const widths = columnWidths(rows);
const separator = ' '.repeat(Math.max(0, gap));
return rows.map((row) => {
const cells = row.map((cell, i) => (i === row.length - 1 ? cell : padCell(cell, widths[i], align[i] ?? 'left')));
return `${indent}${cells.join(separator)}`;
});
}
/** `layoutTable()` as one printable block. */
export function table(rows: readonly (readonly string[])[], options: TableOptions = {}): string {
return layoutTable(rows, options).join('\n');
}
// ─────────────────────────────────────────────────────────────────────────────
// Spinner
// ─────────────────────────────────────────────────────────────────────────────
const HIDE_CURSOR = '\x1b[?25l';
const SHOW_CURSOR = '\x1b[?25h';
const CLEAR_LINE = '\x1b[K';
/** The slice of a stream a spinner needs; `process.stderr` satisfies it. */
export interface SpinnerStream {
isTTY?: boolean;
write(chunk: string): unknown;
}
export interface Spinner {
start(): Spinner;
/** Change the text mid-flight. Silent on a non-TTY, which prints once and stops. */
setText(text: string): void;
/** Clear the line, restore the cursor and optionally print a final line. */
stop(finalLine?: string): void;
}
export interface SpinnerOptions {
stream?: SpinnerStream;
intervalMs?: number;
}
/**
* In-place progress line on stderr, for the calls that block for tens of seconds
* (daemon start, service install). Only a TTY gets the animation: piped output
* and journald get the text once, so a log file never fills with `\r` frames.
*/
export function spinner(text: string, options: SpinnerOptions = {}): Spinner {
const stream = options.stream ?? process.stderr;
const intervalMs = options.intervalMs ?? 90;
const animated = Boolean(stream.isTTY);
let label = text;
let frame = 0;
let timer: NodeJS.Timeout | null = null;
let started = false;
let stopped = false;
const restoreCursor = () => {
if (animated) stream.write(SHOW_CURSOR);
};
const render = () => {
stream.write(`\r${palette.info(SPINNER_FRAMES[frame % SPINNER_FRAMES.length])} ${label}${CLEAR_LINE}`);
frame++;
};
const handle: Spinner = {
start() {
if (started || stopped) return handle;
started = true;
if (!animated) {
stream.write(`${label}\n`);
return handle;
}
stream.write(HIDE_CURSOR);
// A hidden cursor left behind by a Ctrl+C outlives the process, so the
// exit hook is not optional.
process.once('exit', restoreCursor);
render();
// Unref'd: a spinner must never be the reason the process stays alive.
timer = setInterval(render, intervalMs);
timer.unref();
return handle;
},
setText(next: string) {
label = next;
if (animated && started && !stopped) render();
},
stop(finalLine?: string) {
if (stopped) return;
stopped = true;
if (timer) {
clearInterval(timer);
timer = null;
}
if (animated && started) {
stream.write(`\r${CLEAR_LINE}`);
restoreCursor();
process.off('exit', restoreCursor);
}
if (finalLine && animated) stream.write(`${finalLine}\n`);
},
};
return handle;
}
/** Run `work` with a spinner up, stopping it however `work` ends. */
export async function withSpinner<T>(text: string, work: () => Promise<T>, options?: SpinnerOptions): Promise<T> {
const handle = spinner(text, options).start();
try {
return await work();
} finally {
handle.stop();
}
}
// ─────────────────────────────────────────────────────────────────────────────
// Confirm
// ─────────────────────────────────────────────────────────────────────────────
/** Is there a human on the other end of both halves of the terminal? */
export function isInteractive(): boolean {
return Boolean(process.stdin.isTTY && process.stdout.isTTY);
}
/**
* y/N prompt. Answers `false` immediately when stdin is not a TTY (a script
* piping into the CLI must never hang on an invisible question), so callers
* that support a `--force` flag can branch on `isInteractive()` to keep printing
* their "pass --force" hint instead.
*/
export async function confirm(question: string): Promise<boolean> {
if (!isInteractive()) return false;
const rl = createInterface({ input: process.stdin, output: process.stdout });
try {
const answer = await new Promise<string>((resolve) => {
rl.once('SIGINT', () => resolve(''));
rl.question(`${question} ${palette.muted('[y/N]')} `, resolve);
});
return /^y(es)?$/i.test(answer.trim());
} finally {
rl.close();
// readline resumes stdin; a still-flowing stdin keeps the process alive.
process.stdin.pause();
}
}
+300 -211
View File
@@ -8,7 +8,6 @@
*/
import { Command } from 'commander';
import chalk from 'chalk';
import { createRequire } from 'module';
import http from 'node:http';
import https from 'node:https';
@@ -16,6 +15,8 @@ import { existsSync, readFileSync } from 'node:fs';
import { isAbsolute, join } from 'node:path';
import { homedir } from 'node:os';
import { dataPath } from './config/instance.js';
import { casePath } from './config/cases-dir.js';
import { assertValidBasePath } from './config/base-path.js';
import { installAgentSkillInto, removeAgentSkillFrom, type AgentSkillApplyResult } from './hooks-config.js';
import { getSessionManager } from './session-manager.js';
import { getTaskQueue } from './task-queue.js';
@@ -26,6 +27,9 @@ import { isSupportedAttachmentExtension } from './attachment-registry.js';
import { daemonStatus, startDaemon, stopDaemon, type WebLaunchOptions } from './daemon-control.js';
import { installService, serviceStatus, uninstallService } from './service-installer.js';
import { isLoopbackBindHost, isUnauthenticatedNetworkAcknowledged } from './web/network-auth-policy.js';
import { confirm, heading, isInteractive, kv, palette, rule, tint, withSpinner, type Tone } from './cli-style.js';
import type { ToolResult } from './utils/dependency-checker.js';
import type { ReportStyle } from './utils/dependency-report.js';
const require = createRequire(import.meta.url);
const pkg = require('../package.json') as { version: string };
@@ -107,14 +111,14 @@ program
.action(async (filePath, options) => {
const extension = String(filePath).split('.').pop()?.toLowerCase() || '';
if (!isAbsolute(filePath) || !isSupportedAttachmentExtension(extension)) {
console.error(chalk.red('✗ attach requires an absolute path to a png, pdf, docx, pptx, md, or txt file'));
console.error(palette.err('✗ attach requires an absolute path to a png, pdf, docx, pptx, md, or txt file'));
process.exit(1);
}
const sessionId = options.session || process.env.CODEMAN_SESSION_ID;
const apiUrl = options.url || process.env.CODEMAN_API_URL || 'https://127.0.0.1:3000';
if (sessionId && (await postAttachment(apiUrl, sessionId, filePath))) {
console.log(chalk.green('✓ Attachment card requested'));
console.log(palette.ok('✓ Attachment card requested'));
return;
}
@@ -144,7 +148,9 @@ export function resolveCliCasePath(name: string): string {
} catch {
// no registry yet, or unreadable/invalid JSON: fall through to the cases dir
}
return join(homedir(), 'codeman-cases', name);
// Same resolver the server uses, so CODEMAN_CASES_PATH (Docker Compose) moves
// the CLI's idea of a case with it instead of leaving it on the home default.
return casePath(name);
}
/**
@@ -174,7 +180,7 @@ export function resolveSkillTargetPath(options: {
function resolveSkillTarget(options: { case?: string }): string {
const resolved = resolveSkillTargetPath(options);
if (resolved.missingCase !== undefined) {
console.error(chalk.red(`✗ Case not found: ${resolved.missingCase}`));
console.error(palette.err(`✗ Case not found: ${resolved.missingCase}`));
process.exit(1);
}
return resolved.target;
@@ -199,9 +205,9 @@ function reportSkillResult(result: AgentSkillApplyResult, target: string): void
};
const message = messages[result];
if (message.ok) {
console.log(chalk.green(`✓ ${message.text}`));
console.log(palette.ok(`✓ ${message.text}`));
} else {
console.error(chalk.red(`✗ ${message.text}`));
console.error(palette.err(`✗ ${message.text}`));
process.exit(1);
}
}
@@ -220,7 +226,7 @@ skillCmd
const target = resolveSkillTarget(options);
reportSkillResult(await installAgentSkillInto(target), target);
} catch (err) {
console.error(chalk.red(`✗ Failed to install agent skill: ${getErrorMessage(err)}`));
console.error(palette.err(`✗ Failed to install agent skill: ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -235,7 +241,7 @@ skillCmd
const target = resolveSkillTarget(options);
reportSkillResult(await removeAgentSkillFrom(target), target);
} catch (err) {
console.error(chalk.red(`✗ Failed to remove agent skill: ${getErrorMessage(err)}`));
console.error(palette.err(`✗ Failed to remove agent skill: ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -252,11 +258,11 @@ sessionCmd
try {
const manager = getSessionManager();
const session = await manager.createSession(options.dir);
console.log(chalk.green(`✓ Session started: ${session.id}`));
console.log(palette.ok(`✓ Session started: ${session.id}`));
console.log(` Working directory: ${session.workingDir}`);
console.log(` PID: ${session.pid}`);
} catch (err) {
console.error(chalk.red(`✗ Failed to start session: ${getErrorMessage(err)}`));
console.error(palette.err(`✗ Failed to start session: ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -268,70 +274,81 @@ sessionCmd
try {
const manager = getSessionManager();
await manager.stopSession(id);
console.log(chalk.green(`✓ Session stopped: ${id}`));
console.log(palette.ok(`✓ Session stopped: ${id}`));
} catch (err) {
console.error(chalk.red(`✗ Failed to stop session: ${getErrorMessage(err)}`));
console.error(palette.err(`✗ Failed to stop session: ${getErrorMessage(err)}`));
process.exit(1);
}
});
/** Session status in the shared vocabulary: idle is fine, busy is working, anything else is a problem. */
function sessionStatusLabel(status: string): string {
if (status === 'idle') return palette.ok('idle');
if (status === 'busy') return palette.warn('busy');
return palette.err(status);
}
/**
* The one session listing. `codeman list` used to be a copy of this that had
* drifted (it lost the stopped and web-server sections), so it now calls the
* same renderer and only opts out of those two sections.
*/
function printSessionList(options: { includeStored: boolean }): void {
const manager = getSessionManager();
const sessions = manager.getAllSessions();
const stored = manager.getStoredSessions();
if (sessions.length === 0 && Object.keys(stored).length === 0) {
console.log(palette.warn('No sessions found'));
return;
}
console.log(heading('Active Sessions:'));
if (sessions.length === 0) {
console.log(' (none)');
} else {
for (const session of sessions) {
console.log(
` ${palette.info(session.id.slice(0, 8))} ${sessionStatusLabel(session.status)} ${session.workingDir}`
);
}
}
if (options.includeStored) {
const stoppedSessions = Object.values(stored).filter((s) => s.status === 'stopped');
if (stoppedSessions.length > 0) {
console.log(heading('Stopped Sessions:'));
for (const session of stoppedSessions) {
const name = session.name ? ` (${session.name})` : '';
console.log(
` ${palette.muted(session.id.slice(0, 8))} ${palette.muted('stopped')}${name} ${session.workingDir}`
);
}
}
// Sessions the web server owns: this process has no PTY for them, so they
// only exist in the shared state file.
const activeSessions = Object.values(stored).filter((s) => s.status !== 'stopped');
if (sessions.length === 0 && activeSessions.length > 0) {
console.log(heading('Active Sessions (from web server):'));
for (const session of activeSessions) {
const name = session.name ? ` (${session.name})` : '';
const mode = session.mode === 'shell' ? palette.muted(' [shell]') : '';
const cost = session.totalCost ? palette.muted(` $${session.totalCost.toFixed(4)}`) : '';
console.log(
` ${palette.info(session.id.slice(0, 8))} ${sessionStatusLabel(session.status)}${name}${mode}${cost} ${session.workingDir}`
);
}
}
}
console.log('');
}
sessionCmd
.command('list')
.alias('ls')
.description('List all sessions')
.action(() => {
const manager = getSessionManager();
const sessions = manager.getAllSessions();
const stored = manager.getStoredSessions();
if (sessions.length === 0 && Object.keys(stored).length === 0) {
console.log(chalk.yellow('No sessions found'));
return;
}
console.log(chalk.bold('\nActive Sessions:'));
if (sessions.length === 0) {
console.log(' (none)');
} else {
for (const session of sessions) {
const status =
session.status === 'idle'
? chalk.green('idle')
: session.status === 'busy'
? chalk.yellow('busy')
: chalk.red(session.status);
console.log(` ${chalk.cyan(session.id.slice(0, 8))} ${status} ${session.workingDir}`);
}
}
const stoppedSessions = Object.values(stored).filter((s) => s.status === 'stopped');
if (stoppedSessions.length > 0) {
console.log(chalk.bold('\nStopped Sessions:'));
for (const session of stoppedSessions) {
const name = session.name ? ` (${session.name})` : '';
console.log(` ${chalk.gray(session.id.slice(0, 8))} ${chalk.gray('stopped')}${name} ${session.workingDir}`);
}
}
// Show active sessions from state (when web server manages them)
const activeSessions = Object.values(stored).filter((s) => s.status !== 'stopped');
if (sessions.length === 0 && activeSessions.length > 0) {
console.log(chalk.bold('\nActive Sessions (from web server):'));
for (const session of activeSessions) {
const status =
session.status === 'idle'
? chalk.green('idle')
: session.status === 'busy'
? chalk.yellow('busy')
: chalk.red(session.status);
const name = session.name ? ` (${session.name})` : '';
const mode = session.mode === 'shell' ? chalk.gray(' [shell]') : '';
const cost = session.totalCost ? chalk.gray(` $${session.totalCost.toFixed(4)}`) : '';
console.log(` ${chalk.cyan(session.id.slice(0, 8))} ${status}${name}${mode}${cost} ${session.workingDir}`);
}
}
console.log('');
});
.action(() => printSessionList({ includeStored: true }));
sessionCmd
.command('logs <id>')
@@ -342,12 +359,12 @@ sessionCmd
const output = options.errors ? manager.getSessionError(id) : manager.getSessionOutput(id);
if (output === null) {
console.log(chalk.yellow(`Session ${id} not found or not active`));
console.log(palette.warn(`Session ${id} not found or not active`));
return;
}
if (output === '') {
console.log(chalk.gray('(no output)'));
console.log(palette.muted('(no output)'));
return;
}
@@ -374,7 +391,7 @@ taskCmd
completionPhrase: options.completion,
timeoutMs: options.timeout ? parseInt(options.timeout, 10) : undefined,
});
console.log(chalk.green(`✓ Task added: ${task.id}`));
console.log(palette.ok(`✓ Task added: ${task.id}`));
console.log(` Prompt: ${prompt.slice(0, 50)}${prompt.length > 50 ? '...' : ''}`);
console.log(` Priority: ${task.priority}`);
});
@@ -393,26 +410,28 @@ taskCmd
}
if (tasks.length === 0) {
console.log(chalk.yellow('No tasks found'));
console.log(palette.warn('No tasks found'));
return;
}
const statusColors = {
pending: chalk.gray,
running: chalk.yellow,
completed: chalk.green,
failed: chalk.red,
pending: palette.muted,
running: palette.warn,
completed: palette.ok,
failed: palette.err,
};
console.log(chalk.bold('\nTasks:'));
console.log(palette.emph('\nTasks:'));
for (const task of tasks) {
const color = statusColors[task.status];
const prompt = task.prompt.slice(0, 40) + (task.prompt.length > 40 ? '...' : '');
console.log(` ${chalk.cyan(task.id.slice(0, 8))} ${color(task.status.padEnd(10))} [${task.priority}] ${prompt}`);
console.log(
` ${palette.info(task.id.slice(0, 8))} ${color(task.status.padEnd(10))} [${task.priority}] ${prompt}`
);
}
const counts = queue.getCount();
console.log(chalk.bold('\nSummary:'));
console.log(palette.emph('\nSummary:'));
console.log(
` Pending: ${counts.pending}, Running: ${counts.running}, Completed: ${counts.completed}, Failed: ${counts.failed}`
);
@@ -427,11 +446,11 @@ taskCmd
const task = queue.getTask(id);
if (!task) {
console.log(chalk.red(`Task ${id} not found`));
console.log(palette.err(`Task ${id} not found`));
return;
}
console.log(chalk.bold('\nTask Details:'));
console.log(palette.emph('\nTask Details:'));
console.log(` ID: ${task.id}`);
console.log(` Status: ${task.status}`);
console.log(` Priority: ${task.priority}`);
@@ -441,10 +460,10 @@ taskCmd
console.log(` Session: ${task.assignedSessionId}`);
}
if (task.error) {
console.log(` Error: ${chalk.red(task.error)}`);
console.log(` Error: ${palette.err(task.error)}`);
}
if (task.output) {
console.log(chalk.bold('\nOutput:'));
console.log(palette.emph('\nOutput:'));
console.log(task.output.slice(0, 500) + (task.output.length > 500 ? '...' : ''));
}
console.log('');
@@ -457,9 +476,9 @@ taskCmd
.action((id) => {
const queue = getTaskQueue();
if (queue.removeTask(id)) {
console.log(chalk.green(`✓ Task removed: ${id}`));
console.log(palette.ok(`✓ Task removed: ${id}`));
} else {
console.log(chalk.red(`Task ${id} not found`));
console.log(palette.err(`Task ${id} not found`));
}
});
@@ -474,13 +493,13 @@ taskCmd
if (options.all) {
count = queue.clearAll();
console.log(chalk.green(`✓ Cleared ${count} tasks`));
console.log(palette.ok(`✓ Cleared ${count} tasks`));
} else if (options.failed) {
count = queue.clearFailed();
console.log(chalk.green(`✓ Cleared ${count} failed tasks`));
console.log(palette.ok(`✓ Cleared ${count} failed tasks`));
} else {
count = queue.clearCompleted();
console.log(chalk.green(`✓ Cleared ${count} completed tasks`));
console.log(palette.ok(`✓ Cleared ${count} completed tasks`));
}
});
@@ -503,38 +522,38 @@ ralphCmd
}
if (loop.isRunning()) {
console.log(chalk.yellow('Ralph loop is already running'));
console.log(palette.warn('Ralph loop is already running'));
return;
}
loop.on('taskAssigned', (taskId, sessionId) => {
console.log(chalk.cyan(`→ Task ${taskId.slice(0, 8)} assigned to session ${sessionId.slice(0, 8)}`));
console.log(palette.info(`→ Task ${taskId.slice(0, 8)} assigned to session ${sessionId.slice(0, 8)}`));
});
loop.on('taskCompleted', (taskId) => {
console.log(chalk.green(`✓ Task ${taskId.slice(0, 8)} completed`));
console.log(palette.ok(`✓ Task ${taskId.slice(0, 8)} completed`));
});
loop.on('taskFailed', (taskId, error) => {
console.log(chalk.red(`✗ Task ${taskId.slice(0, 8)} failed: ${error}`));
console.log(palette.err(`✗ Task ${taskId.slice(0, 8)} failed: ${error}`));
});
loop.on('stopped', () => {
console.log(chalk.yellow('\nRalph loop stopped'));
console.log(palette.warn('\nRalph loop stopped'));
printStats(loop.getStats());
process.exit(0);
});
await loop.start();
console.log(chalk.green('✓ Ralph loop started'));
console.log(palette.ok('✓ Ralph loop started'));
if (options.minHours) {
console.log(` Minimum duration: ${options.minHours} hours`);
}
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
console.log(palette.muted(' Press Ctrl+C to stop\n'));
// Keep process running
process.on('SIGINT', () => {
console.log(chalk.yellow('\nStopping Ralph loop...'));
console.log(palette.warn('\nStopping Ralph loop...'));
loop.stop();
});
});
@@ -545,11 +564,11 @@ ralphCmd
.action(() => {
const loop = getRalphLoop();
if (!loop.isRunning()) {
console.log(chalk.yellow('Ralph loop is not running'));
console.log(palette.warn('Ralph loop is not running'));
return;
}
loop.stop();
console.log(chalk.green('✓ Ralph loop stopped'));
console.log(palette.ok('✓ Ralph loop stopped'));
});
ralphCmd
@@ -562,9 +581,10 @@ ralphCmd
});
function printStats(stats: ReturnType<ReturnType<typeof getRalphLoop>['getStats']>) {
const statusColor = stats.status === 'running' ? chalk.green : stats.status === 'paused' ? chalk.yellow : chalk.gray;
const statusColor =
stats.status === 'running' ? palette.ok : stats.status === 'paused' ? palette.warn : palette.muted;
console.log(chalk.bold('\nRalph Loop Status:'));
console.log(palette.emph('\nRalph Loop Status:'));
console.log(` Status: ${statusColor(stats.status)}`);
console.log(` Elapsed: ${stats.elapsedHours.toFixed(2)} hours`);
if (stats.minDurationMs) {
@@ -574,14 +594,14 @@ function printStats(stats: ReturnType<ReturnType<typeof getRalphLoop>['getStats'
);
}
console.log(chalk.bold('\nTasks:'));
console.log(palette.emph('\nTasks:'));
console.log(` Pending: ${stats.pending}`);
console.log(` Running: ${stats.running}`);
console.log(` Completed: ${stats.completed} (${stats.tasksCompleted} this session)`);
console.log(` Failed: ${stats.failed}`);
console.log(` Generated: ${stats.tasksGenerated}`);
console.log(chalk.bold('\nSessions:'));
console.log(palette.emph('\nSessions:'));
console.log(` Active: ${stats.activeSessions}`);
console.log(` Idle: ${stats.idleSessions}`);
console.log(` Busy: ${stats.busySessions}`);
@@ -695,20 +715,20 @@ program
}
}
console.log(chalk.bold('\nCodeman Status'));
console.log('─'.repeat(40));
console.log(heading('Codeman Status'));
console.log(rule(40));
console.log(chalk.bold('\nWeb Server:'));
console.log(heading('Web Server:'));
if (probe.reachable) {
const version = probe.version ? ` (v${probe.version})` : '';
console.log(` Status: ${chalk.green('running')}${version} at ${probe.url}`);
console.log(kv('Status', `${palette.ok('running')}${version} at ${probe.url}`));
if (probe.authRequired) {
console.log(chalk.gray(' (answers 401: set CODEMAN_PASSWORD/CODEMAN_USERNAME to see session details)'));
console.log(palette.muted(' (answers 401: set CODEMAN_PASSWORD/CODEMAN_USERNAME to see session details)'));
}
} else {
console.log(` Status: ${chalk.red('not reachable')} at ${candidates.join(' or ')}`);
console.log(kv('Status', `${palette.err('not reachable')} at ${candidates.join(' or ')}`));
console.log(
chalk.gray(' (start it with `codeman web`, or check your service: systemctl --user status codeman-web)')
palette.muted(' (start it with `codeman web`, or check your service: systemctl --user status codeman-web)')
);
}
@@ -716,26 +736,26 @@ program
// as such, so the numbers are never silently a different thing.
if (probe.sessions) {
const live = probe.sessions;
console.log(chalk.bold('\nSessions (live, from the server):'));
console.log(` Total: ${live.length}`);
console.log(` Idle: ${live.filter((s) => s.status === 'idle').length}`);
console.log(` Busy: ${live.filter((s) => s.status === 'busy').length}`);
console.log(heading('Sessions (live, from the server):'));
console.log(kv('Total', String(live.length)));
console.log(kv('Idle', String(live.filter((s) => s.status === 'idle').length)));
console.log(kv('Busy', String(live.filter((s) => s.status === 'busy').length)));
} else {
const manager = getSessionManager();
const storedValues = Object.values(manager.getStoredSessions());
console.log(chalk.bold('\nSessions (from saved state):'));
console.log(` Active: ${storedValues.filter((s) => s.status !== 'stopped').length}`);
console.log(` Idle: ${storedValues.filter((s) => s.status === 'idle').length}`);
console.log(` Busy: ${storedValues.filter((s) => s.status === 'busy').length}`);
console.log(heading('Sessions (from saved state):'));
console.log(kv('Active', String(storedValues.filter((s) => s.status !== 'stopped').length)));
console.log(kv('Idle', String(storedValues.filter((s) => s.status === 'idle').length)));
console.log(kv('Busy', String(storedValues.filter((s) => s.status === 'busy').length)));
}
const taskCounts = getTaskQueue().getCount();
console.log(chalk.bold('\nTasks:'));
console.log(` Total: ${taskCounts.total}`);
console.log(` Pending: ${taskCounts.pending}`);
console.log(` Running: ${taskCounts.running}`);
console.log(` Completed: ${taskCounts.completed}`);
console.log(` Failed: ${taskCounts.failed}`);
console.log(heading('Tasks:'));
console.log(kv('Total', String(taskCounts.total)));
console.log(kv('Pending', String(taskCounts.pending)));
console.log(kv('Running', String(taskCounts.running)));
console.log(kv('Completed', String(taskCounts.completed)));
console.log(kv('Failed', String(taskCounts.failed)));
console.log('');
});
@@ -745,9 +765,17 @@ program
.option('-f, --force', 'Skip confirmation')
.action(async (options) => {
if (!options.force) {
console.log(chalk.yellow('This will stop all sessions and clear all state.'));
console.log(chalk.yellow('Use --force to confirm.'));
return;
console.log(palette.warn('This will stop all sessions and clear all state.'));
// Non-interactive callers keep the old refusal: a script piping into the
// CLI must never be able to reset state by hanging on an unseen question.
if (!isInteractive()) {
console.log(palette.warn('Use --force to confirm.'));
return;
}
if (!(await confirm('Reset all Codeman state?'))) {
console.log(palette.muted('○ Cancelled, nothing was changed'));
return;
}
}
const manager = getSessionManager();
@@ -756,7 +784,7 @@ program
await manager.stopAllSessions();
store.reset();
console.log(chalk.green('✓ All state reset'));
console.log(palette.ok('✓ All state reset'));
});
// Shorthand commands at root level
@@ -767,38 +795,45 @@ program
.action(async (options) => {
const manager = getSessionManager();
const session = await manager.createSession(options.dir);
console.log(chalk.green(`✓ Session started: ${session.id}`));
console.log(palette.ok(`✓ Session started: ${session.id}`));
});
program
.command('list')
.alias('ls')
.description('List all sessions (shorthand)')
.action(() => {
const manager = getSessionManager();
const sessions = manager.getAllSessions();
const stored = manager.getStoredSessions();
.description('List active sessions (shorthand; `codeman session list` also shows stopped ones)')
.action(() => printSessionList({ includeStored: false }));
if (sessions.length === 0 && Object.keys(stored).length === 0) {
console.log(chalk.yellow('No sessions found'));
// ============ TUI ============
program
.command('tui')
.argument('[n]', 'attach straight to the nth session of `codeman tui --list`')
.description('Terminal dashboard for your sessions (the web UI remains the primary surface)')
.option('-l, --list', 'Print the numbered session list and exit, instead of opening the dashboard')
.action(async (position: string | undefined, options: { list?: boolean }) => {
// Imported here, not at the top: the dashboard pulls in the whole TUI core,
// and every other command would pay for it at startup.
const { runTui, runTuiAttach, runTuiList } = await import('./tui/tui-app.js');
if (options.list) {
process.exitCode = await runTuiList();
return;
}
console.log(chalk.bold('\nActive Sessions:'));
if (sessions.length === 0) {
console.log(' (none)');
} else {
for (const session of sessions) {
const status =
session.status === 'idle'
? chalk.green('idle')
: session.status === 'busy'
? chalk.yellow('busy')
: chalk.red(session.status);
console.log(` ${chalk.cyan(session.id.slice(0, 8))} ${status} ${session.workingDir}`);
if (position !== undefined) {
const n = Number.parseInt(position, 10);
if (!Number.isSafeInteger(n) || n < 1) {
console.error(palette.err(`"${position}" is not a session number.`));
console.error(`Run ${palette.info('codeman tui --list')} to see them.`);
process.exitCode = 1;
return;
}
process.exitCode = await runTuiAttach(n);
return;
}
console.log('');
// The dashboard owns the terminal until it quits; exiting explicitly keeps a
// stray handle (a socket mid-close) from stranding the user's shell.
process.exit(await runTui());
});
// ============ Web / daemon / service Commands ============
@@ -809,6 +844,11 @@ function addWebLaunchOptions(cmd: Command): Command {
.option('-H, --host <host>', 'Host to bind to', process.env.CODEMAN_HOST || '127.0.0.1')
.option('-p, --port <port>', 'Port to listen on (env: CODEMAN_PORT)', process.env.CODEMAN_PORT || '3000')
.option('--https', 'Enable HTTPS with self-signed certificate (only needed for remote access, not localhost)')
.option(
'--base-url <path>',
'Sub-path Codeman is mounted under behind a reverse proxy, e.g. /codeman (env: CODEMAN_BASE_URL)',
process.env.CODEMAN_BASE_URL || '/'
)
.option('--title-hostname <hostname>', 'Override the hostname shown in the browser title')
.option(
'--allow-unauthenticated-network',
@@ -825,19 +865,28 @@ function toWebLaunchOptions(options: {
host: string;
port: string;
https?: boolean;
baseUrl?: string;
titleHostname?: string;
allowUnauthenticatedNetwork?: boolean;
multiuser?: boolean;
}): WebLaunchOptions {
const port = parseInt(options.port, 10);
if (!Number.isInteger(port) || port <= 0 || port > 65535) {
console.error(chalk.red(`✗ Invalid port: ${options.port}`));
console.error(palette.err(`✗ Invalid port: ${options.port}`));
process.exit(1);
}
let basePath: string;
try {
basePath = assertValidBasePath(options.baseUrl);
} catch (err) {
console.error(palette.err(`✗ ${err instanceof Error ? err.message : String(err)}`));
process.exit(1);
}
return {
host: options.host,
port,
https: !!options.https,
basePath,
titleHostname: options.titleHostname,
allowUnauthenticatedNetwork: !!options.allowUnauthenticatedNetwork,
multiuser: !!options.multiuser,
@@ -852,11 +901,11 @@ function warnIfUnauthenticatedNetwork(launch: WebLaunchOptions): void {
if (isLoopbackBindHost(launch.host)) return;
if (isUnauthenticatedNetworkAcknowledged(launch.allowUnauthenticatedNetwork)) return;
console.log(
chalk.yellow(
palette.warn(
`⚠ Binding ${launch.host} without CODEMAN_PASSWORD: anyone who can reach this port gets terminal control.`
)
);
console.log(chalk.yellow(' Set CODEMAN_PASSWORD, or bind 127.0.0.1 and front it with tailscale serve.'));
console.log(palette.warn(' Set CODEMAN_PASSWORD, or bind 127.0.0.1 and front it with tailscale serve.'));
}
// Web interface command
@@ -872,17 +921,18 @@ webCmd.action(async (options) => {
const launch = toWebLaunchOptions(options);
if (options.stop) {
const result = await stopDaemon(launch);
// stopDaemon waits for the process to actually exit (up to 15s).
const result = await withSpinner('Stopping Codeman...', () => stopDaemon(launch));
if (result.ok && result.reason === 'not-running') {
console.log(chalk.gray(`○ ${result.message}`));
console.log(palette.muted(`○ ${result.message}`));
return;
}
if (result.ok) {
console.log(chalk.green(`✓ ${result.message ?? `Stopped Codeman (pid ${result.pid})`}`));
console.log(chalk.gray(' Your agents keep running in tmux.'));
console.log(palette.ok(`✓ ${result.message ?? `Stopped Codeman (pid ${result.pid})`}`));
console.log(palette.muted(' Your agents keep running in tmux.'));
return;
}
console.error(chalk.red(`✗ ${result.message ?? 'Could not stop the server'}`));
console.error(palette.err(`✗ ${result.message ?? 'Could not stop the server'}`));
process.exit(1);
}
@@ -890,31 +940,33 @@ webCmd.action(async (options) => {
const status = await daemonStatus(launch);
if (status.responding) {
const version = status.version ? ` (v${status.version})` : '';
console.log(chalk.green(`✓ Responding at ${status.url}${version}`));
console.log(palette.ok(`✓ Responding at ${status.url}${version}`));
} else {
console.log(chalk.yellow(`○ Nothing answering at ${status.url}`));
console.log(palette.warn(`○ Nothing answering at ${status.url}`));
}
console.log(` Daemon pid: ${status.running ? chalk.green(String(status.pid)) : chalk.gray('not running')}`);
console.log(chalk.gray(` Pidfile: ${status.pidFile}`));
console.log(chalk.gray(` Log: ${status.logPath}`));
console.log(kv('Daemon pid', status.running ? palette.ok(String(status.pid)) : palette.muted('not running'), 11));
console.log(palette.muted(kv('Pidfile', status.pidFile, 11)));
console.log(palette.muted(kv('Log', status.logPath, 11)));
if (!status.running && status.responding) {
console.log(chalk.gray(' (running, but not started with --daemon: probably a service or a foreground run)'));
console.log(palette.muted(' (running, but not started with --daemon: probably a service or a foreground run)'));
}
return;
}
if (options.daemon) {
warnIfUnauthenticatedNetwork(launch);
console.log(chalk.cyan('Starting Codeman in the background...'));
const result = await startDaemon(launch);
// The start polls /api/status for up to 30s; without this the shell just sits there.
const result = await withSpinner('Starting Codeman in the background, waiting for it to answer...', () =>
startDaemon(launch)
);
if (result.ok) {
console.log(chalk.green(`\n✓ Codeman is running at ${result.url} (pid ${result.pid})`));
console.log(chalk.gray(` Logs: ${result.logPath}`));
console.log(chalk.gray(' Stop it with: codeman web --stop'));
console.log(chalk.gray(' Want it back after a reboot? codeman service install'));
console.log(palette.ok(`\n✓ Codeman is running at ${result.url} (pid ${result.pid})`));
console.log(palette.muted(` Logs: ${result.logPath}`));
console.log(palette.muted(' Stop it with: codeman web --stop'));
console.log(palette.muted(' Want it back after a reboot? codeman service install'));
return;
}
console.error(chalk.red(`\n✗ ${result.message ?? 'Failed to start'}`));
console.error(palette.err(`\n✗ ${result.message ?? 'Failed to start'}`));
process.exit(1);
}
@@ -924,29 +976,36 @@ webCmd.action(async (options) => {
const https = launch.https;
const titleHostname = options.titleHostname;
const allowUnauthenticatedNetwork = launch.allowUnauthenticatedNetwork ?? false;
const protocol = https ? 'https' : 'http';
const basePath = launch.basePath ?? '';
// Single source of truth for subsystems that read it directly (e.g. renderers).
if (basePath) process.env.CODEMAN_BASE_URL = basePath;
const displayHost = host === '0.0.0.0' ? 'localhost' : host;
console.log(chalk.cyan(`Starting Codeman web interface on ${displayHost}:${port}${https ? ' (HTTPS)' : ''}...`));
console.log(
palette.info(
`Starting Codeman web interface on ${displayHost}:${port}${basePath ? basePath + '/' : ''}${https ? ' (HTTPS)' : ''}...`
)
);
try {
const server = await startWebServer(port, https, false, host, titleHostname, allowUnauthenticatedNetwork);
console.log(chalk.green(`\n✓ Web interface running at ${protocol}://${displayHost}:${port}`));
// The server prints its own "running at" line (it also covers the daemon and
// service launch paths), so this one used to be a duplicate of it.
const server = await startWebServer(port, https, false, host, titleHostname, allowUnauthenticatedNetwork, basePath);
if (https) {
console.log(chalk.yellow(' Note: Accept the self-signed certificate in your browser on first visit'));
console.log(palette.warn(' Note: Accept the self-signed certificate in your browser on first visit'));
}
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
console.log(palette.muted(' Press Ctrl+C to stop\n'));
// Graceful shutdown handler — flush state and clean up on SIGTERM/SIGINT
let shuttingDown = false;
const shutdown = async (signal: string) => {
if (shuttingDown) return;
shuttingDown = true;
console.log(chalk.yellow(`\n${signal} received, shutting down gracefully...`));
console.log(palette.warn(`\n${signal} received, shutting down gracefully...`));
try {
await server.stop();
} catch (err) {
console.error(chalk.red(`Error during shutdown: ${getErrorMessage(err)}`));
console.error(palette.err(`Error during shutdown: ${getErrorMessage(err)}`));
}
process.exit(0);
};
@@ -954,7 +1013,7 @@ webCmd.action(async (options) => {
process.on('SIGINT', () => shutdown('SIGINT'));
process.on('SIGHUP', () => shutdown('SIGHUP'));
} catch (err) {
console.error(chalk.red(`✗ Failed to start web server: ${getErrorMessage(err)}`));
console.error(palette.err(`✗ Failed to start web server: ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -970,20 +1029,23 @@ addWebLaunchOptions(
).action(async (options) => {
const launch = toWebLaunchOptions(options);
warnIfUnauthenticatedNetwork(launch);
console.log(chalk.cyan('Installing the Codeman service...'));
const result = await installService(launch);
for (const warning of result.warnings ?? []) console.log(chalk.yellow(`⚠ ${warning}`));
// Install polls the new unit's /api/status for up to 30s before it can honestly
// report success, so the wait needs a visible heartbeat.
const result = await withSpinner('Installing the Codeman service, waiting for it to answer...', () =>
installService(launch)
);
for (const warning of result.warnings ?? []) console.log(palette.warn(`⚠ ${warning}`));
if (!result.ok) {
console.error(chalk.red(`✗ ${result.message}`));
console.error(palette.err(`✗ ${result.message}`));
process.exit(1);
}
console.log(chalk.green(`✓ ${result.message}`));
console.log(chalk.gray(` Unit: ${result.unitPath}`));
console.log(palette.ok(`✓ ${result.message}`));
console.log(palette.muted(` Unit: ${result.unitPath}`));
if (process.env.CODEMAN_PASSWORD) {
console.log(
chalk.yellow(
palette.warn(
' Note: CODEMAN_PASSWORD was NOT copied into the unit file. Add it there yourself if the service needs auth.'
)
);
@@ -996,10 +1058,10 @@ serviceCmd
.action(() => {
const result = uninstallService();
if (!result.ok) {
console.error(chalk.red(`✗ ${result.message}`));
console.error(palette.err(`✗ ${result.message}`));
process.exit(1);
}
console.log(chalk.green(`✓ ${result.message}`));
console.log(palette.ok(`✓ ${result.message}`));
});
addWebLaunchOptions(
@@ -1007,15 +1069,15 @@ addWebLaunchOptions(
).action(async (options) => {
const status = await serviceStatus(toWebLaunchOptions(options));
if (!status.kind) {
console.log(chalk.yellow(`No supported supervisor on ${process.platform}. Use \`codeman web -d\` instead.`));
console.log(palette.warn(`No supported supervisor on ${process.platform}. Use \`codeman web -d\` instead.`));
return;
}
console.log(` Supervisor: ${status.kind} (${status.name})`);
console.log(` Unit file: ${status.installed ? chalk.green(status.unitPath) : chalk.gray('not installed')}`);
console.log(` Loaded: ${status.loaded ? chalk.green('yes') : chalk.gray('no')}`);
console.log(` Unit file: ${status.installed ? palette.ok(status.unitPath) : palette.muted('not installed')}`);
console.log(` Loaded: ${status.loaded ? palette.ok('yes') : palette.muted('no')}`);
const version = status.version ? ` (v${status.version})` : '';
console.log(
` Responding: ${status.responding ? chalk.green(`yes at ${status.url}${version}`) : chalk.gray(`no at ${status.url}`)}`
` Responding: ${status.responding ? palette.ok(`yes at ${status.url}${version}`) : palette.muted(`no at ${status.url}`)}`
);
});
@@ -1085,7 +1147,7 @@ usersCmd
.action(async (name, options) => {
const { createUser, isValidUsername } = await import('./user-store.js');
if (!isValidUsername(name)) {
console.error(chalk.red('✗ Username must be lowercase, start alphanumeric, 2-32 chars ([a-z0-9_-])'));
console.error(palette.err('✗ Username must be lowercase, start alphanumeric, 2-32 chars ([a-z0-9_-])'));
process.exit(1);
}
try {
@@ -1096,18 +1158,18 @@ usersCmd
password = await promptHiddenPassword('New password: ');
const confirm = await promptHiddenPassword('Confirm password: ');
if (password !== confirm) {
console.error(chalk.red('✗ Passwords do not match'));
console.error(palette.err('✗ Passwords do not match'));
process.exit(1);
}
}
if (!password || password.length < 8) {
console.error(chalk.red('✗ Password must be at least 8 characters'));
console.error(palette.err('✗ Password must be at least 8 characters'));
process.exit(1);
}
const user = await createUser({ username: name, role: options.admin ? 'admin' : 'user', password });
console.log(chalk.green(`✓ Created ${user.role} "${user.username}"`));
console.log(palette.ok(`✓ Created ${user.role} "${user.username}"`));
} catch (err) {
console.error(chalk.red(`✗ ${getErrorMessage(err)}`));
console.error(palette.err(`✗ ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -1126,14 +1188,14 @@ usersCmd
password = await promptHiddenPassword('New password: ');
const confirm = await promptHiddenPassword('Confirm password: ');
if (password !== confirm) {
console.error(chalk.red('✗ Passwords do not match'));
console.error(palette.err('✗ Passwords do not match'));
process.exit(1);
}
}
await setPassword(name, password, { mustChangePassword: false });
console.log(chalk.green(`✓ Password updated for "${name}"`));
console.log(palette.ok(`✓ Password updated for "${name}"`));
} catch (err) {
console.error(chalk.red(`✗ ${getErrorMessage(err)}`));
console.error(palette.err(`✗ ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -1146,17 +1208,17 @@ usersCmd
const { readUsers } = await import('./user-store.js');
const users = await readUsers(true);
if (users.length === 0) {
console.log(chalk.yellow('No users defined (run: codeman users add <name> --admin)'));
console.log(palette.warn('No users defined (run: codeman users add <name> --admin)'));
return;
}
console.log(chalk.bold('\nUsers:'));
console.log(palette.emph('\nUsers:'));
for (const u of users) {
const role = u.role === 'admin' ? chalk.magenta('admin') : chalk.cyan('user ');
const state = u.disabled ? chalk.red('disabled') : chalk.green('enabled ');
const role = u.role === 'admin' ? palette.accent('admin') : palette.info('user ');
const state = u.disabled ? palette.err('disabled') : palette.ok('enabled ');
const flags = [u.mustChangePassword ? 'must-change-pw' : '', u.canBypassPermissions ? 'can-bypass' : '']
.filter(Boolean)
.join(' ');
console.log(` ${role} ${state} ${u.username}${flags ? chalk.gray(` [${flags}]`) : ''}`);
console.log(` ${role} ${state} ${u.username}${flags ? palette.muted(` [${flags}]`) : ''}`);
}
console.log('');
});
@@ -1171,16 +1233,43 @@ usersCmd
await deleteUser(name);
if (options.deleteSpace) {
await deleteUserSpace(name);
console.log(chalk.green(`✓ Deleted user "${name}" and their space`));
console.log(palette.ok(`✓ Deleted user "${name}" and their space`));
} else {
console.log(chalk.green(`✓ Deleted user "${name}" (space left on disk)`));
console.log(palette.ok(`✓ Deleted user "${name}" (space left on disk)`));
}
} catch (err) {
console.error(chalk.red(`✗ ${getErrorMessage(err)}`));
console.error(palette.err(`✗ ${getErrorMessage(err)}`));
process.exit(1);
}
});
/**
* Missing REQUIRED tools are failures; a missing optional one or a skipped check
* is just absence, so it stays muted rather than shouting red at everyone
* without LibreOffice installed.
*/
function dependencyTone(result: ToolResult): Tone {
if (result.status === 'ok') return 'ok';
if (result.status === 'skipped') return 'idle';
return result.required ? 'err' : 'idle';
}
/**
* The colorize hook `dependency-report.ts` was written for. Versions stay in the
* default color (they are data, not a verdict); everything that IS a verdict is
* painted, and the supporting detail is muted so the glyph column reads first.
*/
const DOCTOR_STYLE: ReportStyle = {
title: (text) => palette.emph(text),
heading: (text) => palette.emph(palette.info(text)),
glyph: (result, glyph) => tint(dependencyTone(result), glyph),
label: (text) => text,
status: (result, text) => (result.status === 'ok' ? text : tint(dependencyTone(result), text)),
path: (text) => palette.muted(text),
meta: (text) => palette.muted(text),
summary: (text) => palette.emph(text),
};
program
.command('doctor')
.alias('check-deps')
@@ -1190,7 +1279,7 @@ program
.action(async (options) => {
const { createRealHost, checkAll } = await import('./utils/dependency-checker.js');
const { renderTable, renderJson, computeExitCode } = await import('./utils/dependency-report.js');
const { DEPENDENCY_REGISTRY, TOOL_CATEGORIES } = await import('./config/dependency-registry.js');
const { dependencyRegistry, TOOL_CATEGORIES } = await import('./config/dependency-registry.js');
if (options.category && !(TOOL_CATEGORIES as readonly string[]).includes(options.category)) {
console.error(`Unknown category "${options.category}". Valid categories: ${TOOL_CATEGORIES.join(', ')}`);
@@ -1198,15 +1287,15 @@ program
}
const host = createRealHost();
const registry = options.category
? DEPENDENCY_REGISTRY.filter((t) => t.category === options.category)
: DEPENDENCY_REGISTRY;
const allTools = dependencyRegistry();
const registry = options.category ? allTools.filter((t) => t.category === options.category) : allTools;
const results = checkAll(registry, host);
if (options.json) {
// Raw JSON, never styled: this output is parsed, not read.
console.log(JSON.stringify(renderJson(results, host.environment), null, 2));
} else {
console.log(renderTable(results, host.environment));
console.log(renderTable(results, host.environment, DOCTOR_STYLE));
}
process.exit(computeExitCode(results));
});
+393
View File
@@ -0,0 +1,393 @@
/**
* @fileoverview Scan `~/.codex/sessions/<yyyy>/<mm>/<dd>/rollout-*.jsonl` for Past
* Sessions rows, the codex analog of what `scanOmpSessionsHistory()`
* (omp-transcript.ts) does for omp and `scanProjectDir()` (session-routes.ts)
* does for Claude's own `~/.claude/projects` transcripts.
*
* Without this a codex conversation is invisible to Codeman the moment its
* session record goes away, even though codex itself never forgot it: the
* unified list is built from `~/.claude/projects` plus omp's own store, and
* codex writes to neither. A user who wanted to pick a codex thread back up had
* to find its id by hand and pass `codexConfig.resumeSessionId` to the API.
*
* ## Why this reads windows rather than whole files
*
* An omp session file is the conversation only, so its scanner reads each file
* whole. A codex rollout is not comparable: it carries every reasoning block and
* every tool call, and its `session_meta` line alone embeds the full base
* instructions. Measured on a real store of 519 rollouts, the median file is
* 407 KiB, the 90th percentile 1.3 MiB and the largest 25 MiB, for 381 MiB in
* total. So this reads a head window for the identity and the opening prompt,
* and a tail window for the most recent one.
*
* The head budget is 128 KiB because `session_meta` runs to roughly 19 KiB and
* the first real user message lands near 69 KiB behind it, both measured on
* codex 0.152.1.
*
* ## Where the prompt text comes from
*
* Codex has emitted user input under three shapes, and this reads all of them,
* preferring the ones that carry real input only:
*
* - `event_msg` / `item_completed` with an `item.type` of `UserMessage`, which
* is what codex 0.152.1 writes.
* - `event_msg` / `user_message`, which older versions wrote.
* - `response_item` rows with `role: 'user'`, the last resort. These mix real
* input with injected context (AGENTS.md, environment context, compaction
* summaries), so they are read only when neither shape above appears, and
* the obvious injections are dropped.
*
* @module codex-transcript
*/
import { open, readdir, stat } from 'node:fs/promises';
import { homedir } from 'node:os';
import { join } from 'node:path';
import { LRUMap } from './utils/lru-map.js';
/** Covers `session_meta` (~19 KiB) plus the first user message (~69 KiB behind it). */
const HEAD_BYTES = 131072;
/** Enough to hold the last few turns' worth of lines without re-reading the file. */
const TAIL_BYTES = 65536;
/**
* Newest rollouts to REPORT. Counted in emitted rows, not files scanned: the
* store is mostly sub-agent threads this never returns, so capping files first
* would spend the budget on rows nobody sees.
*/
const MAX_ROLLOUTS = 400;
/**
* How many emitted rows also get a tail read for `lastPrompt`. The head read is
* cached (see below) but the tail cannot be, because appending to a rollout is
* exactly what changes it, so this is the one genuinely per-request cost and it
* stays bounded. Counted in emitted rows for the same reason as above — against
* file index a store of sub-agent threads spends the whole budget before the
* first row that needed it.
*/
const MAX_TAIL_READS = 100;
/** Directory nesting under `sessions/` is year/month/day; stop well past that. */
const MAX_WALK_DEPTH = 5;
/** A rollout shorter than this cannot hold a complete `session_meta` line. */
const MIN_ROLLOUT_BYTES = 100;
export interface CodexHistorySession {
/** The rollout's own thread id — the token `codex resume <id>` expects. */
sessionId: string;
/**
* `session_meta.originator`, which codex stamps from
* CODEX_INTERNAL_ORIGINATOR_OVERRIDE — `codeman_<sessionId>` for every pane
* Codeman spawns. The only link between a FRESH codex pane and the rollout it
* is writing, since such a pane knows no thread id of its own.
*/
originator?: string;
workingDir: string;
sizeBytes: number;
/** ISO timestamp, from the file's own mtime. */
lastModified: string;
firstPrompt?: string;
lastPrompt?: string;
}
/** The half of a rollout that never changes once codex has written it. */
interface RolloutIdentity {
threadId?: string;
cwd?: string;
/** `'subagent'` marks a thread codex spawned for itself. */
threadSource?: string;
/** `codeman_<sessionId>` for a pane Codeman spawned; codex's own default otherwise. */
originator?: string;
firstPrompt?: string;
}
/**
* `session_meta` is written once and never rewritten — the same fact
* `readCodexRolloutMetaCached()` in session-routes.ts relies on — so a path's
* identity is cached, and a rescan costs a `stat` per file plus head reads for
* rollouts this process has not seen before.
*
* ⚠️ The first user message is NOT written up front: codex writes it when the
* user submits. Caching before then pins `firstPrompt: undefined` for the life
* of the process, and every scan of the home screen, the command palette and the
* search-index refresh can land in that window — so the row reads as having no
* prompt until a restart. `shouldCacheIdentity()` is the guard.
*
* Bounded, unlike a plain Map: this process runs for days and every sub-agent
* rollout adds an entry. Same reason and same size as `codexRolloutMetaCache`.
*/
const identityCache = new LRUMap<string, RolloutIdentity>({ maxSize: 4096 });
/**
* Is this identity settled enough to keep?
*
* A known `firstPrompt` settles it. So does a head read that FILLED its window,
* which means the prompt is genuinely not in the first `HEAD_BYTES` rather than
* not written yet. A short file with no prompt is the ambiguous case — codex is
* still to write one — so that one is re-read next scan.
*/
function shouldCacheIdentity(identity: RolloutIdentity, fileSize: number): boolean {
if (!identity.threadId) return false;
return identity.firstPrompt !== undefined || fileSize >= HEAD_BYTES;
}
function codexSessionsRoot(): string {
const home = process.env.CODEX_HOME || join(homedir(), '.codex');
return join(home, 'sessions');
}
/** Read at most `bytes` from the front of a file. Returns '' when unreadable. */
async function readHead(path: string, bytes: number): Promise<string> {
const fh = await open(path, 'r').catch(() => null);
if (!fh) return '';
try {
const buf = Buffer.alloc(bytes);
const { bytesRead } = await fh.read(buf, 0, bytes, 0);
return buf.subarray(0, bytesRead).toString('utf-8');
} catch {
return '';
} finally {
await fh.close().catch(() => {});
}
}
/**
* Read at most `bytes` from the end of a file, dropping the leading partial
* line so every line handed back parses.
*/
async function readTail(path: string, size: number, bytes: number): Promise<string> {
const fh = await open(path, 'r').catch(() => null);
if (!fh) return '';
try {
const want = Math.min(bytes, size);
const buf = Buffer.alloc(want);
const { bytesRead } = await fh.read(buf, 0, want, size - want);
const text = buf.subarray(0, bytesRead).toString('utf-8');
if (want >= size) return text; // whole file, nothing was cut
const nl = text.indexOf('\n');
return nl === -1 ? '' : text.slice(nl + 1);
} catch {
return '';
} finally {
await fh.close().catch(() => {});
}
}
/** Flatten codex's message content, which is a string or an array of text blocks. */
function contentText(content: unknown): string {
if (typeof content === 'string') return content.trim();
if (!Array.isArray(content)) return '';
return content
.filter(
(b): b is { text: string } => !!b && typeof b === 'object' && typeof (b as { text?: unknown }).text === 'string'
)
.map((b) => b.text)
.join('\n')
.trim();
}
/** One line's user-prompt text, whichever of the three shapes it is. */
function userPromptFromLine(entry: {
type?: string;
payload?: {
type?: string;
role?: string;
content?: unknown;
message?: unknown;
item?: { type?: string; content?: unknown };
};
}): { text: string; injectionProne: boolean } | null {
const p = entry.payload;
if (!p) return null;
if (entry.type === 'event_msg' && p.type === 'item_completed' && p.item?.type === 'UserMessage') {
const text = contentText(p.item.content);
return text ? { text, injectionProne: false } : null;
}
if (entry.type === 'event_msg' && p.type === 'user_message') {
const text = typeof p.message === 'string' ? p.message.trim() : contentText(p.message);
return text ? { text, injectionProne: false } : null;
}
if (entry.type === 'response_item' && p.role === 'user') {
const text = contentText(p.content);
return text ? { text, injectionProne: true } : null;
}
return null;
}
/**
* Injected context rather than something the user typed. Codex prepends the
* repository's AGENTS.md and wraps environment context in a tag, and both arrive
* as `response_item` user rows.
*/
function isInjectedContext(text: string): boolean {
return text.startsWith('#') || text.startsWith('<');
}
/** Collapse to one line and cap, so a row carries a title rather than an essay. */
function asPreview(text: string): string {
const flat = text.replace(/\s+/g, ' ').trim();
return flat.length > 200 ? `${flat.slice(0, 200)}…` : flat;
}
/** Parse a head window into the facts about a rollout that never change. */
function parseIdentity(head: string): RolloutIdentity {
const out: RolloutIdentity = {};
let fallback: string | undefined;
for (const line of head.split('\n')) {
if (!line) continue;
let entry: {
type?: string;
payload?: {
id?: string;
session_id?: string;
cwd?: string;
thread_source?: string;
originator?: string;
type?: string;
role?: string;
content?: unknown;
message?: unknown;
item?: { type?: string; content?: unknown };
};
};
try {
entry = JSON.parse(line);
} catch {
continue; // truncated tail of the window, or a malformed line
}
const p = entry.payload;
if (entry.type === 'session_meta' && p) {
out.threadId ??= p.id || p.session_id;
out.cwd ??= p.cwd;
out.threadSource ??= p.thread_source;
out.originator ??= p.originator;
} else if (entry.type === 'turn_context' && p) {
out.cwd ??= p.cwd;
}
if (out.firstPrompt) continue;
const prompt = userPromptFromLine(entry);
if (!prompt) continue;
if (!prompt.injectionProne) {
out.firstPrompt = asPreview(prompt.text);
} else if (!fallback && !isInjectedContext(prompt.text)) {
fallback = asPreview(prompt.text);
}
}
out.firstPrompt ??= fallback;
return out;
}
/** The most recent user prompt in a tail window, or undefined. */
function parseLastPrompt(tail: string): string | undefined {
let best: string | undefined;
let fallback: string | undefined;
for (const line of tail.split('\n')) {
if (!line) continue;
try {
const prompt = userPromptFromLine(JSON.parse(line));
if (!prompt) continue;
if (!prompt.injectionProne) best = asPreview(prompt.text);
else if (!isInjectedContext(prompt.text)) fallback = asPreview(prompt.text);
} catch {
// Malformed line — keep scanning.
}
}
return best ?? fallback;
}
/** Every rollout file under `sessions/`, newest first. */
async function listRollouts(root: string): Promise<Array<{ path: string; mtimeMs: number; size: number }>> {
const files: Array<{ path: string; mtimeMs: number; size: number }> = [];
const walk = async (dir: string, depth: number): Promise<void> => {
if (depth > MAX_WALK_DEPTH) return;
const entries = await readdir(dir, { withFileTypes: true }).catch(() => null);
if (!entries) return;
for (const entry of entries) {
const full = join(dir, entry.name);
if (entry.isDirectory()) {
await walk(full, depth + 1);
continue;
}
if (!entry.isFile() || !entry.name.endsWith('.jsonl')) continue;
const st = await stat(full).catch(() => null);
if (!st || st.size < MIN_ROLLOUT_BYTES) continue;
files.push({ path: full, mtimeMs: st.mtimeMs, size: st.size });
}
};
await walk(root, 0);
files.sort((a, b) => b.mtimeMs - a.mtimeMs);
return files;
}
/**
* Codex conversations on this host, newest first, for the unified session list.
*
* Sub-agent threads are left out: codex spawns them for itself, they are not
* something a person picks back up, and on a real store they outnumber the
* threads that are.
*/
export async function scanCodexSessionsHistory(): Promise<CodexHistorySession[]> {
const files = await listRollouts(codexSessionsRoot());
const out: CodexHistorySession[] = [];
for (const file of files) {
if (out.length >= MAX_ROLLOUTS) break;
let identity = identityCache.get(file.path);
if (!identity) {
identity = parseIdentity(await readHead(file.path, HEAD_BYTES));
if (shouldCacheIdentity(identity, file.size)) identityCache.set(file.path, identity);
}
if (!identity.threadId || identity.threadSource === 'subagent') continue;
// A row with no directory has nowhere to resume INTO, and emitting an empty
// one makes a click post `workingDir: ''`. omp drops such a row; so does this.
if (!identity.cwd) continue;
const lastPrompt =
out.length < MAX_TAIL_READS ? parseLastPrompt(await readTail(file.path, file.size, TAIL_BYTES)) : undefined;
out.push({
sessionId: identity.threadId,
originator: identity.originator,
workingDir: identity.cwd,
sizeBytes: file.size,
lastModified: new Date(file.mtimeMs).toISOString(),
firstPrompt: identity.firstPrompt,
lastPrompt: lastPrompt ?? identity.firstPrompt,
});
}
return out;
}
/**
* Which codex thread each Codeman-spawned pane is writing, keyed by Codeman
* session id.
*
* Codeman spawns every codex pane with
* CODEX_INTERNAL_ORIGINATOR_OVERRIDE=codeman_<sessionId>, and codex stamps that
* into `session_meta.originator`. That is the ONLY link between a fresh codex
* pane and the rollout it is writing: such a pane knows no thread id of its own,
* so it cannot be folded into its own Past-Sessions row from its own side.
*
* Newest wins. `/new` typed inside the codex TUI leaves several rollouts sharing
* one originator, and the pane is on the most recent — so this expects `rows`
* newest-first, as `scanCodexSessionsHistory()` returns them.
*/
export function codexThreadBySessionId(rows: CodexHistorySession[]): Map<string, string> {
const out = new Map<string, string>();
for (const row of rows) {
const owner = /^codeman_(.+)$/.exec(row.originator ?? '')?.[1];
if (owner && !out.has(owner)) out.set(owner, row.sessionId);
}
return out;
}
/** Test seam: drop the per-path identity cache. */
export function __clearCodexIdentityCache(): void {
identityCache.clear();
}
+101
View File
@@ -0,0 +1,101 @@
/**
* @fileoverview Reverse-proxy base-path support — the single source of truth for
* the URL prefix Codeman is mounted under.
*
* When Codeman runs behind a reverse proxy at a sub-path (e.g. `/codeman/`), the
* proxy forwards the FULL request path INCLUDING that prefix (it does not strip
* it). Every URL the server emits to the browser (the HTML shell, redirects,
* the manifest/service-worker) and every URL the browser builds (fetch/SSE/WS)
* must therefore carry the prefix too.
*
* This module normalizes the operator-supplied value (`--base-url` / the
* `CODEMAN_BASE_URL` env var) into ONE canonical form used everywhere:
* - `''` — mounted at the origin root (the default, `/`)
* - `/foo` — mounted at a sub-path (leading slash, NO trailing slash)
*
* Keeping the normalized form free of a trailing slash means `basePath + '/api/x'`
* and `basePath + '/'` both compose cleanly, and `''` degrades to the historical
* root behavior with no special-casing at the call sites.
*
* @module config/base-path
*/
/**
* A normalized base path is either empty (root) or one-or-more `/segment`
* groups, where a segment is a conservative, proxy-safe subset of path
* characters. This deliberately excludes anything that could change routing
* meaning (`?`, `#`, `:`, whitespace, `%`) so the prefix is a plain path.
*/
const VALID_BASE_PATH = /^(?:\/[A-Za-z0-9._~-]+)+$/;
/**
* Normalize an operator-supplied base path into the canonical form.
*
* Accepts loose input (`codeman`, `/codeman`, `/codeman/`, `//codeman//`) and
* returns `''` for root or `/codeman` otherwise. Does NOT validate the character
* set — call {@link assertValidBasePath} (or {@link isValidBasePath}) for that.
*/
export function normalizeBasePath(input: string | undefined | null): string {
if (input === undefined || input === null) return '';
let p = String(input).trim();
if (p === '' || p === '/') return '';
if (!p.startsWith('/')) p = '/' + p;
p = p.replace(/\/{2,}/g, '/'); // collapse duplicate slashes
p = p.replace(/\/+$/, ''); // drop trailing slash(es)
return p;
}
/** True if `normalized` is a legal canonical base path (`''` or `/seg[/seg...]`). */
export function isValidBasePath(normalized: string): boolean {
return normalized === '' || VALID_BASE_PATH.test(normalized);
}
/**
* Normalize AND validate, throwing a human-readable error on bad input. Used by
* the CLI so a typo (`--base-url /a b`, `--base-url ?x`) fails loudly at startup
* instead of silently producing broken URLs.
*/
export function assertValidBasePath(input: string | undefined | null): string {
const normalized = normalizeBasePath(input);
if (!isValidBasePath(normalized)) {
throw new Error(
`Invalid --base-url ${JSON.stringify(input)}: use a plain path like "/codeman" ` +
`(letters, digits, and ._~- in each segment).`
);
}
return normalized;
}
/**
* Join the base path onto a root-absolute application path (`/api/x` → `/base/api/x`).
*
* Leaves alone anything that is not a root-absolute app path: empty strings,
* protocol-relative (`//host`) and absolute URLs (`http://`, `ws://`, `data:`),
* fragments/queries, and paths already carrying the prefix. This is the one
* function the whole codebase routes URL construction through.
*/
export function joinBasePath(basePath: string, path: string): string {
if (!basePath) return path;
if (typeof path !== 'string' || path.length === 0) return path;
if (!path.startsWith('/')) return path; // relative / fragment / query — resolved against <base>
if (path.startsWith('//')) return path; // protocol-relative
if (path === basePath || path.startsWith(basePath + '/') || path.startsWith(basePath + '?')) {
return path; // already prefixed
}
return basePath + path;
}
/**
* Strip the base path off an INCOMING request URL so internal routing stays
* prefix-agnostic. Requests that arrive WITHOUT the prefix (health checks,
* hooks, the docker bridge — all of which hit the raw port, bypassing the proxy)
* are returned unchanged, so the server answers at both `/api/x` and
* `/base/api/x`.
*/
export function stripBasePath(basePath: string, url: string): string {
if (!basePath) return url;
if (url === basePath) return '/';
if (url.startsWith(basePath + '/')) return url.slice(basePath.length);
if (url.startsWith(basePath + '?')) return '/' + url.slice(basePath.length);
return url;
}
+47
View File
@@ -112,3 +112,50 @@ export const FILE_PEEK_BYTES = 8 * 1024 - 1; // 8KB (inclusive end offset)
* Override: CODEMAN_MAX_PASTE_IMAGE_BYTES (bytes)
*/
export const MAX_PASTE_IMAGE_BYTES = parseInt(process.env.CODEMAN_MAX_PASTE_IMAGE_BYTES || '') || 50 * 1024 * 1024; // 50MB
// ============================================================================
// File Download Limits
// ============================================================================
/**
* Parse a byte-limit env var, where `0` explicitly means "no limit".
*
* The `parseInt(...) || default` idiom used elsewhere in this file cannot
* express that: it treats 0 as falsy and silently restores the default.
*/
function parseByteLimitEnv(raw: string | undefined, fallback: number): number {
if (raw === undefined || raw.trim() === '') return fallback;
const parsed = Number.parseInt(raw, 10);
if (!Number.isFinite(parsed) || parsed < 0) return fallback;
return parsed;
}
/**
* Maximum size (bytes) of a file served by the raw/download file routes:
* `GET /api/sessions/:id/file-raw` (the Files panel's download link and the
* file-preview overlay), the attachment `/raw` route, and `GET /api/download`.
*
* ⚠️ This is a sanity bound, NOT memory protection. All three bodies are
* STREAMED and `Range`-aware (`sendFileBody` in file-routes.ts), so a large
* file costs one read stream rather than its size in RSS. The historical 50MB
* cap predates that streaming rewrite and its "prevent memory exhaustion"
* comment described a `readFile()` that no longer exists — all it did was
* refuse legitimate downloads of build artifacts, videos and archives.
*
* Set `CODEMAN_MAX_DOWNLOAD_BYTES=0` to remove the cap entirely.
* Override: CODEMAN_MAX_DOWNLOAD_BYTES (bytes)
*/
export const MAX_FILE_DOWNLOAD_BYTES = parseByteLimitEnv(
process.env.CODEMAN_MAX_DOWNLOAD_BYTES,
2 * 1024 * 1024 * 1024 // 2GB
);
/** True when `size` exceeds the download cap (a cap of 0 means unlimited). */
export function exceedsDownloadLimit(size: number): boolean {
return MAX_FILE_DOWNLOAD_BYTES > 0 && size > MAX_FILE_DOWNLOAD_BYTES;
}
/** Human-readable "File too large (…)" message for a refused download. */
export function downloadTooLargeMessage(size: number): string {
return `File too large (${Math.round(size / 1024 / 1024)}MB > ${Math.round(MAX_FILE_DOWNLOAD_BYTES / 1024 / 1024)}MB limit). Raise or remove it with CODEMAN_MAX_DOWNLOAD_BYTES (0 = unlimited).`;
}
+37
View File
@@ -0,0 +1,37 @@
/**
* @fileoverview Where case (project) folders live.
*
* Deliberately NOT instance-scoped, unlike `dataPath()`: `~/codeman-cases` is
* shared by every Codeman on the machine, the same way `~/codeman-users/<u>`
* user spaces are, so a beta instance sees the same projects as prod.
*
* `CODEMAN_CASES_PATH` overrides the location. Docker Compose deployments set
* it to a host-absolute bind mount so a Docker case's workspace resolves to the
* SAME absolute path inside Codeman and on the host daemon that mounts it.
*
* ⚠️ **One resolver, every caller.** This started life as three hardcoded
* `join(homedir(), 'codeman-cases')` copies. When only the web server's copy
* learned the override, `codeman skill install --case <name>` still looked in
* the home default and reported "Case not found" on exactly the deployment the
* override exists for. A new cases-dir consumer imports this; it does not
* rebuild the path.
*
* (`state-store.ts` keeps its own literal on purpose: that one migrates the
* historical `~/claudeman-cases` directory to `~/codeman-cases` by name, and is
* about the old default location rather than the active one.)
*
* @module config/cases-dir
*/
import { homedir } from 'node:os';
import { join } from 'node:path';
/** Absolute path to the shared cases directory. */
export function getCasesDir(): string {
return process.env.CODEMAN_CASES_PATH || join(homedir(), 'codeman-cases');
}
/** Absolute path to one case folder inside it. */
export function casePath(name: string): string {
return join(getCasesDir(), name);
}
+190
View File
@@ -0,0 +1,190 @@
/**
* @fileoverview The argv rendering engine — turns a `CliLaunch` spec plus a set of resolved
* parameter values into the shell command string that goes into `bash -c "..."`.
*
* SECURITY MODEL (read before touching this file):
*
* 1. Config contains no shell text. There is no `command: "..."` field anywhere in the
* schema. An entry declares a sequence of typed tokens (`ArgSpec`); this module is the
* ONLY place that turns them into a string, and it owns every separator itself: a single
* space between tokens, and ` || ` between fallback variants. Neither can originate from
* config, because config has no field that could hold either.
* 2. Every literal (`lit`, `flag`, `value`) is validated against `SAFE_BARE_TOKEN` — no
* space, quote, backtick, `$`, `;`, `&`, `|`, `<`, `>`, parens, braces, newline or
* backslash — at LOAD time (see schema.ts), so a bad literal fails registry validation
* rather than reaching this renderer.
* 3. Every `valueFrom` resolves through a declared `ParamSpec`, whose `token` variant names
* a PATTERN rather than accepting one — see patterns.ts. A value that fails its pattern
* causes the WHOLE ArgSpec to be dropped, exactly like the hand-written builders this
* replaces (an invalid `--model` value silently omits `--model`, it does not substitute
* something else).
* 4. Escaping and validation are independent. `renderToken()` always re-checks the resolved
* value against `SAFE_BARE_TOKEN` before emitting it unquoted; anything else is
* single-quote-escaped. So even a value that somehow bypassed pattern validation is still
* quoted, never concatenated raw.
*
* @module config/cli-registry/argv
*/
import type { ArgSpec, CliEntry, CliLaunch, Cond, EngineValue, ParamSpec, QuoteStyle } from './types.js';
import { matchesPattern } from './patterns.js';
import { SAFE_BARE_TOKEN } from './patterns.js';
/** Resolved parameter values, keyed by the name declared in `CliLaunch.params`. */
export type ParamValues = Record<string, string | boolean | undefined>;
/** Values the caller supplies for the reserved engine params. */
export type EngineValues = Partial<Record<EngineValue, string>>;
/**
* POSIX single-quote escaping: end-quote, escaped-literal-quote, restart-quote. Identical in
* shape to the three copies already in the codebase (tmux-manager.ts, remote-hosts.ts,
* docker-hosts.ts) — kept local rather than importing one of them so this module has no
* dependency on the files it is replacing.
*/
function singleQuoteEscape(value: string): string {
return `'${value.replace(/'/g, `'\\''`)}'`;
}
function doubleQuoteEscape(value: string): string {
// Escape the characters that are special inside a double-quoted bash string. SAFE_BARE_TOKEN
// already excludes all of them, so in practice this never fires; kept as defense in depth.
return `"${value.replace(/([$`"\\])/g, '\\$1')}"`;
}
/**
* Render a single resolved value per its requested quote style. `auto` (the default) emits
* bare only when the value is provably safe; every other case single-quotes.
*/
function renderToken(value: string, style: QuoteStyle | undefined): string {
const safe = SAFE_BARE_TOKEN.test(value);
switch (style) {
case 'double':
return doubleQuoteEscape(value);
case 'single':
return singleQuoteEscape(value);
case 'bare':
return safe ? value : singleQuoteEscape(value);
case 'auto':
default:
return safe ? value : singleQuoteEscape(value);
}
}
/** Resolve one parameter to a plain string, or undefined if it is unset / invalid. */
function resolveParam(
name: string,
spec: ParamSpec | undefined,
params: ParamValues,
engineValues: EngineValues
): string | undefined {
if (!spec) return undefined;
if (spec.type === 'engine') return engineValues[spec.source];
const raw = params[name];
if (raw === undefined) return spec.type === 'enum' ? spec.default : undefined;
if (spec.type === 'bool') return typeof raw === 'boolean' ? String(raw) : undefined;
if (spec.type === 'enum') {
const s = String(raw);
return spec.values.includes(s) ? s : spec.default;
}
// token
const s = String(raw);
return matchesPattern(spec.pattern, s) ? s : undefined;
}
/** Is the resolved value "set" for the purposes of a `state` condition? */
function isSet(name: string, params: ParamValues, resolved: (n: string) => string | undefined): boolean {
if (name in params) {
const raw = params[name];
if (typeof raw === 'boolean') return true; // a bool param is always "set" once declared
}
return resolved(name) !== undefined;
}
function evalCond(
cond: Cond | undefined,
params: ParamValues,
resolved: (n: string) => string | undefined,
gatesPassed: ReadonlySet<string>
): boolean {
if (!cond) return true;
if ('allOf' in cond) return cond.allOf.every((c) => evalCond(c, params, resolved, gatesPassed));
if ('anyOf' in cond) return cond.anyOf.some((c) => evalCond(c, params, resolved, gatesPassed));
if ('not' in cond) return !evalCond(cond.not, params, resolved, gatesPassed);
if ('capabilityGate' in cond) return gatesPassed.has(cond.capabilityGate);
if ('state' in cond) {
const set = isSet(cond.param, params, resolved);
return cond.state === 'set' ? set : !set;
}
// { param, is }
const raw = params[cond.param];
if (typeof cond.is === 'boolean') return raw === cond.is;
return resolved(cond.param) === cond.is;
}
function renderArg(
spec: ArgSpec,
params: ParamValues,
resolved: (n: string) => string | undefined,
gatesPassed: ReadonlySet<string>
): string | null {
if (!evalCond(spec.when, params, resolved, gatesPassed)) return null;
if ('lit' in spec) return spec.lit;
if ('flag' in spec && !('value' in spec) && !('valueFrom' in spec)) return spec.flag;
if ('flag' in spec && 'value' in spec) return `${spec.flag} ${renderToken(spec.value, spec.quote)}`;
if ('flag' in spec && 'valueFrom' in spec) {
const v = resolved(spec.valueFrom);
return v === undefined ? null : `${spec.flag} ${renderToken(v, spec.quote)}`;
}
// bare positional
const v = resolved((spec as { valueFrom: string }).valueFrom);
return v === undefined ? null : renderToken(v, (spec as { quote?: QuoteStyle }).quote);
}
/**
* Render one CLI's launch command. Returns the full `bash -c` payload — never a shell
* fragment with embedded newlines or unescaped separators, by construction (see file header).
*
* `gatesPassed` — the set of `capabilities.gates` keys whose version requirement is
* currently satisfied. Callers compute this once per spawn (it depends on a version probe),
* never inside the renderer, keeping this function pure and easy to test byte-for-byte.
*/
export function renderLaunch(
launch: CliLaunch,
params: ParamValues,
engineValues: EngineValues,
gatesPassed: ReadonlySet<string> = new Set()
): string {
const cache = new Map<string, string | undefined>();
const resolved = (name: string): string | undefined => {
if (cache.has(name)) return cache.get(name);
const v = resolveParam(name, launch.params[name], params, engineValues);
cache.set(name, v);
return v;
};
const passing = launch.variants.filter((variant) => evalCond(variant.when, params, resolved, gatesPassed));
const chosen = launch.chain === 'fallback' ? passing : passing.slice(0, 1);
const rendered = chosen.map((variant) =>
variant.args
.map((arg) => renderArg(arg, params, resolved, gatesPassed))
.filter((tok): tok is string => tok !== null)
.join(' ')
);
return rendered.join(' || ');
}
/** Convenience: render an entry's launch command straight from a `CliEntry`. */
export function renderCliCommand(
entry: CliEntry,
params: ParamValues,
engineValues: EngineValues,
gatesPassed?: ReadonlySet<string>
): string {
return renderLaunch(entry.launch, params, engineValues, gatesPassed);
}
+61
View File
@@ -0,0 +1,61 @@
/**
* @fileoverview Barrel for the CLI registry module.
* @module config/cli-registry
*/
export type {
ArgSpec,
CliCapabilities,
CliCredStore,
CliDiscovery,
CliEntry,
CliEnv,
CliId,
CliIdentityProbe,
CliLaunch,
CliOverlays,
CliRegistryFile,
CliVariant,
CliVersionProbe,
Cond,
EngineValue,
ParamSpec,
QuoteStyle,
} from './types.js';
export {
matchesPattern,
TOKEN_PATTERNS,
SAFE_BARE_TOKEN,
compileVersionRegex,
MAX_VERSION_OUTPUT,
} from './patterns.js';
export type { TokenPattern } from './patterns.js';
export { renderLaunch, renderCliCommand } from './argv.js';
export type { EngineValues, ParamValues } from './argv.js';
export { CliEntrySchema } from './schema.js';
export type { ValidatedCliEntry } from './schema.js';
export { STOCK_CLIS } from './stock.js';
export {
asCliId,
cliIds,
enabledCliIds,
enabledClis,
getCli,
listClis,
loadCliRegistry,
reloadCliRegistry,
resolveInstallCommandForPlatform,
resolveRegistry,
} from './registry.js';
export type { LoadResult } from './registry.js';
export {
COMPOSER_ANCHOR_KINDS,
isKnownLauncherProfile,
isKnownPredictProfile,
isKnownSetenvProfile,
LAUNCHER_PROFILE_NAMES,
PREDICT_PROFILES,
SETENV_PROFILE_NAMES,
TRANSCRIPT_READER_NAMES,
} from './profiles.js';
export type { LauncherProfileName, SetenvProfileName } from './profiles.js';
+121
View File
@@ -0,0 +1,121 @@
/**
* @fileoverview Named value patterns for the CLI registry's argv engine.
*
* Config entries select a pattern BY NAME; the regexes themselves live here, in code.
* That is deliberate and is the reason a user-editable `clis.json` cannot widen its own
* validation: there is no field anywhere in the schema that accepts a raw regex for a
* shell token, so no entry can supply `.*` (nor a catastrophically backtracking one).
*
* The sole user-supplied regex in the whole registry is `discovery.version.regex`, which
* is applied to `--version` OUTPUT rather than to a shell token, and goes through
* `compileVersionRegex()` below.
*
* Every pattern here is transcribed from the builder it replaces in tmux-manager.ts, so
* the argv engine accepts and rejects exactly the values the hand-written builders did.
*
* @module config/cli-registry/patterns
*/
/** Names a value pattern. Config may only reference these. */
export type TokenPattern =
| 'model'
| 'model-claude'
| 'model-pi'
| 'id'
| 'id-dotted'
| 'uuid'
| 'slug'
| 'path-segment'
| 'tool-list'
| 'config-kv';
/**
* The patterns, each traced to the builder it came from.
*
* ⚠️ These are ALLOWLISTS (`^...$` over a safe character class), never blocklists — with
* one deliberate exception, `tool-list`, which mirrors the existing `--allowedTools`
* sanitizer. That one is a metacharacter REJECTION because tool specs legitimately contain
* `(`, `)`, `*`, `:` and spaces (`Bash(git:*), Read`), so an allowlist of safe words cannot
* express it. Keeping it byte-identical to the original matters more than making it uniform.
*/
const PATTERNS: Record<TokenPattern, RegExp> = {
// buildOpenCodeCommand / buildCodexCommand / buildGeminiCommand / buildAntigravityCommand
model: /^[a-zA-Z0-9._\-/]+$/,
// buildSpawnCommand's claude branch — `[` and `]` for bracketed model aliases
'model-claude': /^[a-zA-Z0-9._\-[\]]+$/,
// buildPiCommand — `:` for a thinking suffix (`sonnet:high`), `/` for `provider/id`
'model-pi': /^[a-zA-Z0-9._\-/:]+$/,
// opencode --session, codex resume
id: /^[a-zA-Z0-9_-]+$/,
// gemini --resume, antigravity --conversation, pi --session
'id-dotted': /^[a-zA-Z0-9._-]+$/,
// claude --resume / --session-id
uuid: /^[a-f0-9-]+$/,
// pi --provider
slug: /^[a-z0-9-]+$/,
// dsh --profile. Deliberately STRICTER than `id-dotted`: a profile name is both
// interpolated into the shell line AND joined into a filesystem path, so it must be a
// single path segment. Requiring a leading alphanumeric is what rules out `.`, `..` and
// dotfile names, which `id-dotted` would happily accept.
'path-segment': /^[a-zA-Z0-9][a-zA-Z0-9._-]*$/,
// codex --config tui.animations=false
'config-kv': /^[A-Za-z0-9._-]+=[A-Za-z0-9._-]+$/,
// Placeholder; `tool-list` is handled by isSafeToolList() below, not by a match.
'tool-list': /^$/,
};
/**
* Shell metacharacters rejected in an `--allowedTools` value. Transcribed verbatim from
* buildClaudePermissionFlags so the accepted set does not move.
*/
const TOOL_LIST_DANGEROUS = /[;&|$`\\{}<>'"[\]\n\r]/;
/** Does `value` satisfy the named pattern? */
export function matchesPattern(pattern: TokenPattern, value: string): boolean {
if (pattern === 'tool-list') return value.length > 0 && !TOOL_LIST_DANGEROUS.test(value);
return PATTERNS[pattern].test(value);
}
/** Every pattern name, for schema validation and error messages. */
export const TOKEN_PATTERNS = Object.keys(PATTERNS) as TokenPattern[];
/**
* Characters a token may contain and still be emitted UNQUOTED into the `bash -c "..."`
* command string. Intentionally narrower than "what bash tolerates": anything outside it
* gets single-quoted, so the classification can only ever err toward more quoting.
*/
export const SAFE_BARE_TOKEN = /^[A-Za-z0-9._:@=+/,-]+$/;
/**
* Longest `--version` output we will run a user-supplied regex over. A version banner is a
* line or two; anything larger is a misconfiguration, and capping the input is what keeps a
* sloppy (not necessarily malicious) regex from becoming a stall.
*/
export const MAX_VERSION_OUTPUT = 200;
/** Longest permitted `discovery.version.regex` source. */
const MAX_VERSION_REGEX_SOURCE = 200;
/**
* Nested quantifiers — `(a+)+`, `(a*)*`, `(a+)*` and friends — the classic catastrophic
* backtracking shape. Rejected outright rather than analysed: this field exists to pull a
* semver out of a banner, and nothing legitimate for that job needs a nested quantifier.
*/
const NESTED_QUANTIFIER = /\([^)]*[+*][^)]*\)\s*[+*{]/;
/**
* Compile a user-supplied version regex, or return null if it is not one we are willing to
* run. Returning null (rather than throwing) lets the caller degrade to "version unknown",
* which every consumer already handles.
*/
export function compileVersionRegex(source: string): RegExp | null {
if (source.length > MAX_VERSION_REGEX_SOURCE) return null;
if (NESTED_QUANTIFIER.test(source)) return null;
try {
// No `g`: a global regex carries lastIndex state across calls, which is a documented
// footgun in this codebase (see utils/regex-patterns.ts).
return new RegExp(source);
} catch {
return null;
}
}
+100
View File
@@ -0,0 +1,100 @@
/**
* @fileoverview The NAMES of code profiles a `CliEntry` field may select, and the helpers
* that validate them.
*
* A profile is the escape hatch for behaviour that is genuinely code-shaped and cannot be
* expressed as data — codex's predictive write-through echo, deepseek's profile-launcher
* runnability check, deepseek's status bridge — without letting any of that code branch on
* a CLI's id. A registry field names a profile; the implementation lives beside whatever it
* needs, and looks its name up here.
*
* ⚠️ This module is PURE and must stay that way: names, types and predicates only, no
* imports outside this directory. The implementations pull in resolvers and the status
* shim, which in turn reach back into the registry, so holding them here would close an
* import cycle (profiles → deepseek-cli-resolver → cli-resolver → registry → schema →
* profiles). Keeping the names here and the implementations at their call sites is what
* lets `schema.ts` validate a profile name at LOAD time — a custom entry naming a profile
* this build does not implement fails loudly instead of silently failing closed later.
*
* The rule all of this enforces: `test/cli-registry-no-id-branching.test.ts` fails on any
* `mode === '<stock id>'` comparison outside `stock.ts`, so a NEW behavioural special case
* must be added here, named, and referenced from a registry field — never inlined as an id
* check at the call site.
*
* ⚠️ A profile is a LAST resort, not a convenience. Reach for one only when the behaviour
* needs to run code (a side effect, a computed value, a probe); anything that is a list, a
* flag, or a string belongs in the entry as data, where a custom CLI can also use it.
*
* @module config/cli-registry/profiles
*/
/**
* Predictive local-echo profiles, selected via `capabilities.echo.predictProfile`.
*
* Implementation: packages/xterm-zerolag-input/src/predictive-echo-addon.ts.
*
* ⚠️ Unlike the other two registries, an unknown name here degrades to the 'buffer' policy
* rather than failing. Echo is a comfort feature — a worse-but-working overlay beats a
* refused session — which is why `predictProfile` alone is not schema-validated below.
*/
export const PREDICT_PROFILES: Record<string, true> = {
codex: true,
};
/**
* Launcher profiles, selected via `discovery.launcherProfile`.
*
* For a CLI whose binary launches some further target, and so cannot answer two questions
* from the binary alone: is it RUNNABLE (stricter than "is the binary on disk?"), and what
* is the DEFAULT target when the caller names none? A CLI naming no profile is runnable
* exactly when its binary resolves, and has no default target.
*
* Implementation: `src/utils/cli-launcher.ts`.
*/
export const LAUNCHER_PROFILE_NAMES = [
// `dsh` is a launcher over $DSH_HOME/profiles/<name>, and the profiles DeepSeek itself
// ships (web, headless) cannot drive a terminal pane. Binary AND a pane-capable profile.
'deepseek-profile',
] as const;
/**
* Extra `tmux setenv` work, selected via `env.setenvProfile`.
*
* Implementation: `src/tmux-manager.ts`, which already owns every setenv call.
*
* ⚠️ Anything that is merely "forward this name from the server's own env" belongs in
* `env.tmuxSetenvKeys` as data and must NOT be given a profile.
*/
export const SETENV_PROFILE_NAMES = [
// DeepSeek's terminal front door reports idle/working/blocked to a supervisor over the
// generic env-gated Herdr contract; this makes Codeman that supervisor. It needs a
// profile rather than key names because it writes an executable shim to disk and then
// exports that shim's path along with the session's own pane id.
'deepseek-status-bridge',
] as const;
export type LauncherProfileName = (typeof LAUNCHER_PROFILE_NAMES)[number];
export type SetenvProfileName = (typeof SETENV_PROFILE_NAMES)[number];
/**
* Transcript readers, selected via `capabilities.transcript`. Unlike the profile registries
* above this one is closed over the schema enum itself rather than an open string, since
* transcript format is a small, genuinely fixed set — see CliCapabilities['transcript'].
*/
export const TRANSCRIPT_READER_NAMES = ['claude-jsonl', 'codex-rollout', 'deepseek-zstd', 'none'] as const;
/** Composer-row finders, selected via `capabilities.echo.anchor.kind`. Also schema-closed. */
export const COMPOSER_ANCHOR_KINDS = ['glyph', 'cursor', 'none'] as const;
/** True when `name` is a predictive-echo profile this build actually implements. */
export function isKnownPredictProfile(name: string | undefined): boolean {
return name !== undefined && Object.prototype.hasOwnProperty.call(PREDICT_PROFILES, name);
}
export function isKnownLauncherProfile(name: string): name is LauncherProfileName {
return (LAUNCHER_PROFILE_NAMES as readonly string[]).includes(name);
}
export function isKnownSetenvProfile(name: string): name is SetenvProfileName {
return (SETENV_PROFILE_NAMES as readonly string[]).includes(name);
}
+235
View File
@@ -0,0 +1,235 @@
/**
* @fileoverview Loads, merges and re-validates the CLI registry.
*
* `~/.codeman/clis.json` holds OVERRIDES and CUSTOM entries only — never a full copy of the
* stock catalog — so a shipped fix to a stock definition actually reaches an existing
* install, and the file stays small enough to hand-edit.
*
* Resolution: start from `STOCK_CLIS` → deep-merge each override by id (objects merge
* key-wise, arrays replace wholesale) → validate every resulting entry. A stock entry that
* fails validation after merge falls back to its pristine stock definition (a fat-fingered
* override cannot brick a shipped CLI); a custom entry that fails is dropped with a warning
* rather than failing the whole load. Stock entries are always emitted, so `shell` and
* `claude` can be disabled but can never go missing — large parts of the app assume at
* minimum that a shell fallback exists.
*
* ⚠️ READ-ONLY. Nothing in this module writes, creates or migrates the file. That is a
* deliberate property, not a missing feature: there is no settings UI and no write API yet,
* so there is nothing to persist, and it means importing the registry — which
* `src/web/schemas.ts` does, transitively, just to validate a request — performs no
* filesystem writes. A `seededStockIds` ratchet belongs with the write API that needs it.
* The one exception is the quarantine RENAME of a file that fails to parse (see
* `readRegistryFile`), which happens on first use rather than at import.
*
* ⚠️ The file must be mode 0600. `isUnsafePermissions` refuses ANY group/world bit, read
* bits included, so a file created with a normal umask (0644) is ignored. Every reason the
* file was ignored, or an entry in it dropped, is logged ONCE on first load: the warnings
* used to be returned to a caller that nobody wired up, so a normally-created file was
* ignored with no feedback anywhere (found reviewing #347).
*
* @module config/cli-registry/registry
*/
import { existsSync, readFileSync, renameSync, statSync } from 'node:fs';
import { dataPath } from '../instance.js';
import type { CliEntry, CliId, CliRegistryFile } from './types.js';
import { CliEntrySchema } from './schema.js';
import { STOCK_CLIS } from './stock.js';
/** Construct a validated CliId. Throws if `raw` is not a well-formed id — call at API boundaries. */
export function asCliId(raw: string): CliId {
if (!/^[a-z][a-z0-9-]{0,23}$/.test(raw)) {
throw new Error(`invalid CLI id: ${JSON.stringify(raw)}`);
}
return raw as CliId;
}
function filePath(): string {
return dataPath('clis.json');
}
/**
* Keys that must never be merged out of a hand-editable JSON file.
*
* `JSON.parse` produces `__proto__` as an ORDINARY own property, but `result[key] = …` on a
* plain object walks the setter chain and would set the merged object's PROTOTYPE instead.
* Not exploitable today — every merged entry is spread into `{ ...merged, id, stock }` and
* then Zod-parsed before anything reads it, which drops the effect — but "not exploitable
* because of what a caller happens to do afterwards" is a property that quietly stops
* holding. A `continue` in the loop that reads the file is the cheap end of that trade.
*/
const UNMERGEABLE_KEYS = new Set(['__proto__', 'constructor', 'prototype']);
/** Plain-object deep merge: nested objects merge key-wise, arrays and primitives replace. */
function deepMerge<T>(base: T, override: unknown): T {
if (override === null || typeof override !== 'object' || Array.isArray(override)) {
return (override === undefined ? base : (override as T)) ?? base;
}
if (base === null || typeof base !== 'object' || Array.isArray(base)) {
return override as T;
}
const result: Record<string, unknown> = { ...(base as Record<string, unknown>) };
for (const [key, value] of Object.entries(override as Record<string, unknown>)) {
if (UNMERGEABLE_KEYS.has(key)) continue;
result[key] = deepMerge((base as Record<string, unknown>)[key], value);
}
return result as T;
}
export interface LoadResult {
entries: CliEntry[];
warnings: string[];
}
/**
* Refuse a registry file with any group/world permission bit — same posture as the ssh-key
* discipline, so 0600 is the only accepted mode. This file selects the binaries Codeman
* spawns, so a writable one is a way to redirect every session.
*
* POSIX only: Windows has no meaningful group/world bits on NTFS (Node reports every file
* as mode 0o666 there regardless of its actual ACL), so this check would flag every file on
* Windows and silently ignore all user config. `win32` relies on NTFS ACLs instead, which
* this check cannot see and does not attempt to.
*/
function isUnsafePermissions(path: string): boolean {
if (process.platform === 'win32') return false;
try {
const mode = statSync(path).mode & 0o777;
return (mode & 0o077) !== 0;
} catch {
return false;
}
}
function readRegistryFile(path: string, warnings: string[]): CliRegistryFile | null {
if (!existsSync(path)) return null;
if (isUnsafePermissions(path)) {
warnings.push(
`${path} must be mode 0600 (no group/world permission bits; run \`chmod 600 ${path}\`); ignoring it and falling back to stock CLIs.`
);
return null;
}
let raw: string;
try {
raw = readFileSync(path, 'utf-8');
} catch (err) {
warnings.push(`Failed to read ${path}: ${(err as Error).message}. Falling back to stock CLIs.`);
return null;
}
try {
const parsed = JSON.parse(raw) as CliRegistryFile;
if (typeof parsed !== 'object' || parsed === null || typeof parsed.clis !== 'object') {
throw new Error('missing "clis" object');
}
return parsed;
} catch (err) {
// QUARANTINE, never overwrite: the file is hand-editable, so a syntax error is far more
// likely to be a half-finished edit than junk. Renaming keeps the user's work.
const quarantined = `${path}.invalid-${Date.now()}`;
try {
renameSync(path, quarantined);
warnings.push(`${path} was not valid JSON (${(err as Error).message}); moved to ${quarantined}.`);
} catch {
warnings.push(
`${path} was not valid JSON (${(err as Error).message}); left in place, falling back to stock CLIs.`
);
}
return null;
}
}
/**
* Merge the stock catalog with a (possibly absent) registry file. PURE — no IO, which is
* what lets the load tests drive every merge case directly.
*/
export function resolveRegistry(stock: CliEntry[], file: CliRegistryFile | null, warnings: string[]): LoadResult {
const stockById = new Map(stock.map((e) => [e.id as string, e]));
const overrides = file?.clis ?? {};
const entries: CliEntry[] = [];
for (const stockEntry of stock) {
const id = stockEntry.id as string;
const override = overrides[id];
const merged = override ? deepMerge(stockEntry, override) : stockEntry;
// `stock: true` is forced here rather than read from the merged object, so an override
// can never flip a custom entry's provenance or vice versa.
const parsed = CliEntrySchema.safeParse({ ...merged, id, stock: true });
if (parsed.success) {
entries.push(parsed.data as CliEntry);
} else {
warnings.push(
`Override for stock CLI "${id}" failed validation; using the shipped definition. ${parsed.error.message}`
);
entries.push(stockEntry);
}
}
for (const [id, raw] of Object.entries(overrides)) {
if (stockById.has(id)) continue; // already merged above
// Same forcing in the other direction: a custom entry claiming `stock: true` cannot
// shadow or impersonate a shipped one.
const parsed = CliEntrySchema.safeParse({ ...(raw as object), id, stock: false });
if (parsed.success) {
entries.push(parsed.data as CliEntry);
} else {
warnings.push(`Custom CLI "${id}" failed validation and was dropped. ${parsed.error.message}`);
}
}
entries.sort((a, b) => a.order - b.order);
return { entries, warnings };
}
let cache: LoadResult | null = null;
/**
* Load the effective registry (stock + user overrides). Memoized for the process lifetime;
* `reloadCliRegistry()` invalidates.
*/
export function loadCliRegistry(): LoadResult {
if (cache) return cache;
const warnings: string[] = [];
const existing = readRegistryFile(filePath(), warnings);
cache = resolveRegistry(STOCK_CLIS, existing, warnings);
// Once per process (the result is memoized): silence here is what made a 0644 file look
// like "the override feature does nothing".
for (const warning of warnings) console.warn(`[cli-registry] ${warning}`);
return cache;
}
/** Drop the memoized registry so the next `loadCliRegistry()` re-reads the file. */
export function reloadCliRegistry(): void {
cache = null;
}
export function listClis(): CliEntry[] {
return loadCliRegistry().entries;
}
export function enabledClis(): CliEntry[] {
return listClis().filter((e) => e.enabled);
}
export function getCli(id: string): CliEntry | undefined {
return listClis().find((e) => (e.id as string) === id);
}
export function cliIds(): string[] {
return listClis().map((e) => e.id as string);
}
/** Every enabled entry's id, in registry order. */
export function enabledCliIds(): string[] {
return enabledClis().map((e) => e.id as string);
}
/**
* Resolve the install command for the current platform, falling back to the linux one (the
* common case for a `curl | bash` or `npm install -g` line) and then to whatever is
* declared. Display text only — never executed. See CliDiscovery.install.command.
*/
export function resolveInstallCommandForPlatform(entry: CliEntry): string | undefined {
const { command } = entry.discovery.install;
const platform = process.platform as 'linux' | 'darwin' | 'win32';
return command[platform] ?? command.linux ?? Object.values(command)[0];
}
+445
View File
@@ -0,0 +1,445 @@
/**
* @fileoverview Zod validation for CLI registry entries.
*
* Every object here is `.strict()`: an unknown key is a hard validation error, not a
* silently-ignored one. That matters for a security-relevant schema — a typo in a field name
* must never degrade to "field absent, so the permissive default applies".
*
* The load-bearing rule enforced here is `SHELL_TOKEN`: it is what makes it impossible for a
* `clis.json` entry to smuggle shell metacharacters into the eventual `bash -c "..."` string
* (see argv.ts's file header for the full model).
*
* @module config/cli-registry/schema
*/
import { z } from 'zod';
import { compileVersionRegex, TOKEN_PATTERNS } from './patterns.js';
import { isKnownLauncherProfile, isKnownSetenvProfile } from './profiles.js';
/** A bare CLI id: lowercase, starts with a letter, at most 24 chars. Also used as a CSS/URL token. */
const cliId = z
.string()
.regex(/^[a-z][a-z0-9-]{0,23}$/, 'id must be lowercase, start with a letter, and be at most 24 chars');
/** An env var name. */
const envName = z
.string()
.regex(/^[A-Z_][A-Z0-9_]*$/, 'env var name must be UPPER_SNAKE_CASE')
.max(64);
/**
* A shell-safe bare word: no space, quote, backtick, `$`, `;`, `&`, `|`, `<`, `>`, parens,
* braces, newline or backslash. Every LITERAL in the launch spec (base command, flag names,
* fixed values) must satisfy this — see argv.ts's file header.
*/
const shellToken = z
.string()
.min(1)
.max(256)
.regex(/^[A-Za-z0-9._:@=+/,-]+$/, 'must be a plain word with no shell metacharacters');
const flagToken = z.string().regex(/^--?[A-Za-z0-9][A-Za-z0-9-]*$/, 'must look like -x or --long-flag');
const quoteStyle = z.enum(['auto', 'bare', 'double', 'single']);
const condSchema: z.ZodType<import('./types.js').Cond> = z.lazy(() =>
z.union([
z.object({ param: z.string(), is: z.union([z.string(), z.boolean()]) }).strict(),
z.object({ param: z.string(), state: z.enum(['set', 'unset']) }).strict(),
z.object({ allOf: z.array(condSchema).min(1).max(8) }).strict(),
z.object({ anyOf: z.array(condSchema).min(1).max(8) }).strict(),
z.object({ not: condSchema }).strict(),
z.object({ capabilityGate: z.string() }).strict(),
])
);
const paramSpecSchema = z.union([
z
.object({ type: z.literal('enum'), values: z.array(z.string()).min(1).max(16), default: z.string().optional() })
.strict(),
z.object({ type: z.literal('bool') }).strict(),
z.object({ type: z.literal('token'), pattern: z.enum(TOKEN_PATTERNS as [string, ...string[]]) }).strict(),
z
.object({
type: z.literal('engine'),
source: z.enum([
'sessionId',
'sessionName',
'muxName',
'effortLevel',
'effortSettingsJson',
'codemanPrefixedSessionId',
'launcherDefaultTarget',
]),
})
.strict(),
]);
const argSpecSchema = z.union([
z.object({ lit: shellToken, when: condSchema.optional() }).strict(),
z.object({ flag: flagToken, when: condSchema.optional() }).strict(),
z.object({ flag: flagToken, value: shellToken, quote: quoteStyle.optional(), when: condSchema.optional() }).strict(),
z
.object({ flag: flagToken, valueFrom: z.string(), quote: quoteStyle.optional(), when: condSchema.optional() })
.strict(),
z.object({ valueFrom: z.string(), quote: quoteStyle.optional(), when: condSchema.optional() }).strict(),
]);
const variantSchema = z
.object({
id: z.string().min(1).max(40),
when: condSchema.optional(),
// min(0): the `shell` entry declares a variant with no args — tmux-manager resolves the
// real login shell in code, since it varies per remote user's /etc/passwd entry.
args: z.array(argSpecSchema).max(32),
})
.strict();
const launchSchema = z
.object({
params: z.record(z.string(), paramSpecSchema),
chain: z.enum(['first', 'fallback']).optional(),
variants: z.array(variantSchema).min(1).max(4),
legacyConfigAliases: z.record(z.string(), z.string()).optional(),
legacyConfigField: z.string().min(1).max(40).optional(),
resumeAppend: z
.union([
z.object({ style: z.literal('flag'), flag: flagToken }).strict(),
z.object({ style: z.literal('positional'), token: shellToken }).strict(),
])
.optional(),
})
.strict()
.superRefine((launch, ctx) => {
const paramNames = new Set(Object.keys(launch.params));
const checkValueFrom = (name: string, path: (string | number)[]) => {
if (!paramNames.has(name)) {
ctx.addIssue({ code: 'custom', message: `valueFrom "${name}" is not a declared param`, path });
}
};
launch.variants.forEach((variant, vi) => {
variant.args.forEach((arg, ai) => {
if ('valueFrom' in arg) checkValueFrom(arg.valueFrom, ['variants', vi, 'args', ai, 'valueFrom']);
});
});
if (launch.chain === 'fallback') {
const last = launch.variants.at(-1);
if (last?.when) {
ctx.addIssue({
code: 'custom',
message: 'the last variant of a fallback chain must have no `when` (it must be the guaranteed terminal case)',
path: ['variants', launch.variants.length - 1, 'when'],
});
}
}
if (launch.legacyConfigAliases) {
for (const paramName of Object.keys(launch.legacyConfigAliases)) {
if (!paramNames.has(paramName)) {
ctx.addIssue({
code: 'custom',
message: `legacyConfigAliases key "${paramName}" is not a declared param`,
path: ['legacyConfigAliases', paramName],
});
}
}
}
});
const versionProbeSchema = z
.object({
arg: shellToken,
regex: z.string().max(200).optional(),
requireVersionMatch: z.boolean().optional(),
retryOnTransientFailure: z.boolean().optional(),
})
.strict();
const identityProbeSchema = z
.object({
arg: shellToken,
// Same 200-char cap as version.regex, and compiled through the same compileVersionRegex()
// guard at use time. This is the second and last config-supplied regex in the registry.
regex: z.string().min(1).max(200),
})
.strict();
const discoverySchema = z
.object({
// min(0): the `shell` entry has no binary of its own (it resolves the login shell in code).
binaries: z.array(shellToken).max(4),
searchDirs: z.array(z.string().max(300)).max(16),
version: versionProbeSchema.optional(),
identity: identityProbeSchema.optional(),
launcherProfile: z.string().max(40).optional(),
launcherTargetParam: z.string().max(40).optional(),
install: z
.object({
// z.record with an enum key type requires every enum member in Zod v4; the install
// command legitimately varies by platform and most entries only need one or two, so
// this is a plain object of optional platform keys instead.
command: z
.object({
linux: z.string().max(500).optional(),
darwin: z.string().max(500).optional(),
wsl: z.string().max(500).optional(),
win32: z.string().max(500).optional(),
})
.strict(),
npmPackage: z.string().max(200).optional(),
docsUrl: z.url().optional(),
})
.strict(),
})
.strict();
const envExportSchema = z
.object({
name: envName,
value: z.union([
shellToken,
z
.object({
engine: z.enum([
'sessionId',
'sessionName',
'muxName',
'effortLevel',
'effortSettingsJson',
'codemanPrefixedSessionId',
'launcherDefaultTarget',
]),
})
.strict(),
]),
when: condSchema.optional(),
})
.strict();
const envSchema = z
.object({
exports: z.array(envExportSchema).max(16),
unset: z.array(envName).max(16),
tmuxSetenvKeys: z.array(envName).max(32),
dockerExecEnvNames: z.array(envName).max(32),
configSetenv: z
.array(z.object({ name: envName, fromParam: z.string().min(1).max(40) }).strict())
.max(8)
.optional(),
allowedPrefixes: z
.array(
z
.string()
.min(3)
.max(32)
.regex(/^[A-Z][A-Z0-9_]*_$/)
)
.max(8),
allowedKeys: z.array(envName).max(8),
configContentVar: envName.optional(),
setenvProfile: z.string().max(40).optional(),
})
.strict();
const echoSchema = z
.object({
policy: z.enum(['buffer', 'predict', 'off']),
anchor: z.union([
z
.object({ kind: z.literal('glyph'), glyph: z.string().min(1).max(4), offset: z.number().int().min(0).max(16) })
.strict(),
z.object({ kind: z.literal('cursor') }).strict(),
z.object({ kind: z.literal('none') }).strict(),
]),
predictProfile: z.string().max(40).optional(),
})
.strict();
const capabilitiesSchema = z
.object({
external: z.boolean(),
requiresMux: z.boolean(),
hooks: z.enum(['none', 'always', 'supervised']),
transcript: z.enum(['claude-jsonl', 'codex-rollout', 'deepseek-zstd', 'omp-jsonl', 'none']),
altScreen: z.enum(['strip-full', 'strip-mux-only', 'preserve']),
echo: echoSchema,
wheelForward: z
.object({ mode: z.enum(['never', 'version-gated']), minVersion: z.string().max(20).optional() })
.strict(),
keyboardAccessory: z.enum(['agent', 'shell']),
privilegedCommandGate: z.boolean(),
startMode: z.enum(['interactive', 'shell']),
stripInkBloat: z.boolean(),
ralph: z.boolean(),
respawn: z.boolean(),
effort: z.boolean(),
agentSkillInjection: z.boolean(),
statusLineTelemetry: z.boolean(),
workDetect: z
.object({
promptGlyph: z.string().min(1).max(8),
// Config-supplied regex, so it goes through the same guard as `version.regex`:
// ~/.codeman/clis.json can set this, and the compiled pattern runs on the PTY
// hot path, where a nested quantifier would be a ReDoS against the event loop.
// A broken pattern must also fail at LOAD time rather than inside a data handler.
workingLine: z
.string()
.min(1)
.refine(
(src) => compileVersionRegex(src) !== null,
'workingLine must be a regex compileVersionRegex() accepts: at most 200 characters, no nested quantifiers'
),
})
.strict()
.optional(),
model: z
.object({ source: z.enum(['flag', 'claude-settings-file', 'none']), param: z.string().optional() })
.strict(),
privilegedParams: z
.array(
z
.object({
param: z.string(),
clampTo: z.union([z.boolean(), z.string()]),
materializeWhenAbsent: z.boolean().optional(),
})
.strict()
)
.max(8),
// Exact env var NAMES, not prefixes: this list is a targeted deny, and a prefix here
// would let one entry silently strip a whole namespace off every owner's overrides.
privilegedEnvKeys: z.array(envName).max(8),
gates: z.record(z.string(), z.object({ minVersion: z.string().max(20), failClosed: z.boolean() }).strict()),
maxFrameBytes: z.number().int().positive().optional(),
})
.strict();
const credStoreSchema = z
.object({
rel: z.string().min(1).max(100),
shareDirs: z.array(z.string().max(100)).optional(),
shareFiles: z.array(z.string().max(100)).optional(),
seedFiles: z.array(z.string().max(100)).optional(),
seedWhole: z.boolean().optional(),
})
.strict();
/**
* A remote/docker default pane command: space-separated bare words from the SAME safe
* charset as `shellToken` (no shell metacharacters), so `claude --dangerously-skip-permissions`
* is expressible while still excluding `;`, `|`, `$`, backticks and quotes — this is not an
* escape hatch into arbitrary shell text, it is one bare command plus bare flags.
*/
const commandLine = z
.string()
.min(1)
.max(200)
.regex(
/^[A-Za-z0-9._:@=+/,-]+( [A-Za-z0-9._:@=+/,-]+)*$/,
'must be space-separated bare words with no shell metacharacters'
);
const overlayTargetSchema = z.union([
z.object({ command: commandLine.optional(), rootCommand: commandLine.optional() }).strict(),
z.object({ disabled: z.literal(true) }).strict(),
]);
const overlaysSchema = z
.object({
remote: overlayTargetSchema.optional(),
docker: overlayTargetSchema.optional(),
credStore: credStoreSchema.optional(),
})
.strict();
export const CliEntrySchema = z
.object({
id: cliId,
label: z.string().min(1).max(60),
shortBadge: z.string().min(1).max(6),
accent: z.string().regex(/^#[0-9a-fA-F]{6}$/, 'accent must be a 6-digit hex colour'),
enabled: z.boolean(),
stock: z.boolean(),
order: z.number().int(),
kind: z.enum(['agent', 'shell']),
discovery: discoverySchema,
launch: launchSchema,
env: envSchema,
capabilities: capabilitiesSchema,
overlays: overlaysSchema,
})
.strict()
.superRefine((entry, ctx) => {
const gateNames = new Set(Object.keys(entry.capabilities.gates));
const walkConds = (cond: import('./types.js').Cond | undefined) => {
if (!cond) return;
if ('capabilityGate' in cond && !gateNames.has(cond.capabilityGate)) {
ctx.addIssue({
code: 'custom',
message: `capabilityGate "${cond.capabilityGate}" is not declared in capabilities.gates`,
});
}
if ('allOf' in cond) cond.allOf.forEach(walkConds);
if ('anyOf' in cond) cond.anyOf.forEach(walkConds);
if ('not' in cond) walkConds(cond.not);
};
for (const variant of entry.launch.variants) {
walkConds(variant.when);
for (const arg of variant.args) walkConds(arg.when);
}
// Reject a profile name this build does not implement, rather than letting it fail
// closed at use time. An unimplemented `launcherProfile` would make the CLI look
// permanently uninstalled, and an unimplemented `setenvProfile` would silently skip
// setup the CLI needs; both are far easier to diagnose as a load-time error naming the
// field. (`echo.predictProfile` is deliberately NOT checked here — see profiles.ts.)
const { launcherProfile } = entry.discovery;
if (launcherProfile !== undefined && !isKnownLauncherProfile(launcherProfile)) {
ctx.addIssue({
code: 'custom',
message: `discovery.launcherProfile "${launcherProfile}" is not a profile this build implements`,
path: ['discovery', 'launcherProfile'],
});
}
// An env var exported from a param that does not exist would silently export nothing,
// and for DSH_PERMISSION_MODE that means silently losing a permission clamp.
const declaredParams = new Set(Object.keys(entry.launch.params));
entry.env.configSetenv?.forEach((mapping, i) => {
if (!declaredParams.has(mapping.fromParam)) {
ctx.addIssue({
code: 'custom',
message: `configSetenv fromParam "${mapping.fromParam}" is not a declared launch param`,
path: ['env', 'configSetenv', i, 'fromParam'],
});
}
});
// Same class of silent failure on the OTHER privileged surface, and this one is a
// security control: `privilegedParams[].param` is the multi-user bypass clamp's only
// handle on a CLI's privilege switch, and a name that is not a declared param clamps
// NOTHING — no load error, no failing test, the clamp simply stops running. The clamp
// resolves the name through `legacyConfigAliases`, so this check is what keeps the two
// in ONE namespace rather than two that merely coincide today: they do not for codex
// (`bypassApprovals` vs `dangerouslyBypassApprovals`), and giving deepseek's
// `permissionMode` an alias later would otherwise have removed its clamp with nothing
// saying so.
entry.capabilities.privilegedParams.forEach((clamp, i) => {
if (!declaredParams.has(clamp.param)) {
ctx.addIssue({
code: 'custom',
message: `privilegedParams param "${clamp.param}" is not a declared launch param`,
path: ['capabilities', 'privilegedParams', i, 'param'],
});
}
});
const { setenvProfile } = entry.env;
if (setenvProfile !== undefined && !isKnownSetenvProfile(setenvProfile)) {
ctx.addIssue({
code: 'custom',
message: `env.setenvProfile "${setenvProfile}" is not a profile this build implements`,
path: ['env', 'setenvProfile'],
});
}
});
export type ValidatedCliEntry = z.infer<typeof CliEntrySchema>;
File diff suppressed because it is too large Load Diff
+559
View File
@@ -0,0 +1,559 @@
/**
* @fileoverview Type definitions for the CLI registry — the single source of truth for
* which agent CLIs Codeman supports and how each one is discovered, launched and treated.
*
* This replaces the hard-coded `SessionMode` union and the ~123 per-mode branches that grew
* out of it. The guiding rule: NO code may branch on a CLI's id. Behaviour that genuinely
* differs between CLIs is expressed either as data here, or as a named PROFILE selected by
* a capability field (see profiles.ts) — never as `mode === 'codex'`.
*
* @module config/cli-registry/types
*/
import type { TokenPattern } from './patterns.js';
/**
* A CLI identifier. Branded so an arbitrary string cannot be passed where a validated id is
* expected; construct with `asCliId()` at the API boundary.
*/
export type CliId = string & { readonly __cliId: unique symbol };
// ---------------------------------------------------------------------------
// Launch argv DSL
// ---------------------------------------------------------------------------
/** Values the ENGINE supplies. Config may reference these by name but never author them. */
export type EngineValue =
| 'sessionId'
| 'sessionName'
| 'muxName'
| 'effortLevel'
| 'effortSettingsJson'
/** `sessionId` prefixed `codeman_<id>` — codex's unique per-pane rollout originator. */
| 'codemanPrefixedSessionId'
/**
* For a launcher CLI (`discovery.launcherProfile`), the target to launch when the caller
* named none — deepseek's default `dsh` profile. Resolved at spawn time, never frozen
* into config, because it depends on what is installed on this machine right now.
*/
| 'launcherDefaultTarget';
/**
* A declared launch parameter. `token` params carry caller-supplied data and are therefore
* the only ones that need a pattern; `engine` params are produced in code.
*/
export type ParamSpec =
| { type: 'enum'; values: string[]; default?: string }
| { type: 'bool' }
| { type: 'token'; pattern: TokenPattern }
| { type: 'engine'; source: EngineValue };
/** A boolean guard over parameter state. */
export type Cond =
| { param: string; is: string | boolean }
| { param: string; state: 'set' | 'unset' }
| { allOf: Cond[] }
| { anyOf: Cond[] }
| { not: Cond }
/** Names an entry in `capabilities.gates`. Fail-closed gates omit when version is unknown. */
| { capabilityGate: string };
/**
* How a token is quoted when emitted into the bash command string.
*
* This exists ONLY to preserve byte-identical output with the hand-written builders being
* replaced (claude wraps its values in double quotes; the other builders emit bare words).
* It is never a safety lever: `renderToken()` verifies the value is metacharacter-free
* before honouring an explicit style, and falls back to single-quote escaping if it is not.
* So the worst a wrong `quote` can do is make output uglier, never unsafe.
*/
export type QuoteStyle = 'auto' | 'bare' | 'double' | 'single';
/** One argv element. */
export type ArgSpec =
/** A bare literal word, e.g. the base binary or codex's `resume` subcommand. */
| { lit: string; when?: Cond }
/** A valueless flag, e.g. `--no-approve`. */
| { flag: string; when?: Cond }
/** A flag with a fixed literal value. */
| { flag: string; value: string; quote?: QuoteStyle; when?: Cond }
/** A flag whose value comes from a declared param. */
| { flag: string; valueFrom: string; quote?: QuoteStyle; when?: Cond }
/** A bare positional value from a param, e.g. codex's `resume <id>`. */
| { valueFrom: string; quote?: QuoteStyle; when?: Cond };
/** One alternative command form. */
export interface CliVariant {
/** Stable name for diagnostics and tests, e.g. 'resume' / 'new'. */
id: string;
when?: Cond;
args: ArgSpec[];
}
export interface CliLaunch {
params: Record<string, ParamSpec>;
/**
* 'first' — emit the first variant whose `when` passes (the usual case).
* 'fallback' — emit EVERY passing variant joined by the engine's own ` || `, which is how
* claude's `--resume X || --session-id Y` shell fallback is expressed without
* config ever containing shell text. The engine owns the operator.
*/
chain?: 'first' | 'fallback';
variants: CliVariant[];
/**
* Maps a declared param name to the field name it arrives under on the legacy
* `POST /api/sessions` wire shape (`OpenCodeConfig.continueSession`, etc — the per-mode
* config objects predate this registry and stay on the wire for compatibility). A param
* with no entry here is looked up under its own name. This is what lets the spawn-command
* bridge (`session-cli-registry-bridge.ts`) stay generic: it reads the raw legacy config
* object through this DATA-declared alias table instead of a per-mode `if (mode === ...)`.
*/
legacyConfigAliases?: Record<string, string>;
/**
* The field on the legacy spawn option bag holding this CLI's `<Mode>Config` object
* (`openCodeConfig`, `codexConfig`, …). Those per-mode objects predate this registry and
* stay on the wire for API compatibility, so SOMETHING has to know which one to read —
* declaring it here as data is what keeps the bridge a generic reader instead of a
* `switch (mode)`.
*
* ABSENT means this CLI's launch fields live at the TOP LEVEL of the option bag rather
* than nested in a config object. That is claude, whose discrete `claudeMode` /
* `allowedTools` / `model` / `resumeSessionId` fields predate the `<Mode>Config` pattern
* entirely — so "read the option bag itself" is not a special case for it, it is just
* the other shape.
*/
legacyConfigField?: string;
/**
* How to APPEND a resume id onto an already-built base command, for the docker in-container
* "tmux was re-created, resume the surviving transcript" path (`appendResumeFlag` in
* tmux-manager.ts) — a narrower, append-only sibling of the full `variants` shape above,
* which builds a whole command from scratch. Absent = this CLI has no resume flag to
* append (shell, opencode: opencode's docker resume goes through its own config object).
*/
resumeAppend?: { style: 'flag'; flag: string } | { style: 'positional'; token: string };
}
// ---------------------------------------------------------------------------
// Discovery
// ---------------------------------------------------------------------------
export interface CliVersionProbe {
arg: string;
/** Serialized regex, applied to `--version` output only. See compileVersionRegex(). */
regex?: string;
/**
* Treat a binary whose version output does not match as ABSENT rather than as
* present-with-unknown-version. For CLIs with short, generic binary names (`pi`), where a
* `which` hit is not by itself evidence the right program is installed.
*/
requireVersionMatch?: boolean;
/** Retry a failed probe with backoff instead of caching the failure (claude's behaviour). */
retryOnTransientFailure?: boolean;
}
/**
* An identity probe: proof that the binary we found is the program we meant, not an
* unrelated one that happens to share the name.
*
* A version probe is not enough on its own. Debian ships a `dsh` (dancer's shell) that
* answers `--version` perfectly happily, and npm carries squatters for `pi` and `grok`.
* `requireVersionMatch` catches a binary whose version output has the WRONG SHAPE; this
* catches one whose output has the right shape but names the wrong program.
*
* Ordering matters and belongs to the resolver, not to config: identity is checked FIRST,
* so an impostor is rejected before its version string is ever parsed.
*/
export interface CliIdentityProbe {
/** Argument that makes the binary describe itself, e.g. `--help`. */
arg: string;
/**
* Serialized regex the output must match. Compiled through `compileVersionRegex()`, so
* it inherits the same length cap and nested-quantifier rejection — this is the second
* (and last) config-supplied regex in the registry, and it runs against truncated
* command output exactly like the first.
*/
regex: string;
}
export interface CliDiscovery {
/**
* Binary name(s), first hit wins.
*
* This is why the registry fixes a live bug: the mode name is NOT always the binary
* name (`antigravity` runs `agy`), and `probeDockerCliVersion` assumed it was.
*/
binaries: string[];
/** Extra directories probed after `which`. A leading `~` expands to homedir; nothing else. */
searchDirs: string[];
version?: CliVersionProbe;
/** Proof the binary is the right program, checked BEFORE the version probe. */
identity?: CliIdentityProbe;
/**
* Names a LAUNCHER profile (profiles.ts): this CLI's binary is a launcher over some
* further target, so two questions the registry normally answers from the binary alone
* have to be asked of that target instead.
*
* - Is it RUNNABLE? Stricter than "is the binary on disk?".
* - What is the DEFAULT target, when the caller names none?
*
* DeepSeek is why this exists and is its only user. `dsh` launches a profile from
* `$DSH_HOME/profiles/<name>`, and the profiles DeepSeek itself ships (`web`,
* `headless`) cannot drive a terminal pane — so a perfectly-installed `dsh` with no
* third-party TUI profile is installed-but-NOT-runnable. The Run button gates on
* runnability while the "add a profile" affordance gates on mere availability;
* collapsing the two would either hide the affordance that fixes the problem or offer a
* run that always fails.
*
* The default target reaches the launch spec as the `launcherDefaultTarget` engine
* value, so it stays a runtime lookup rather than a value frozen into config.
*
* Absent (the normal case) means the binary IS the program, and its presence IS
* runnability.
*/
launcherProfile?: string;
/**
* The launch param naming the target a caller asked for, so the launcher profile can say
* why THAT specific target will not start rather than only whether any will. Meaningless
* without `launcherProfile`.
*/
launcherTargetParam?: string;
install: {
/**
* DISPLAY TEXT ONLY. Shown verbatim in "CLI not found. Install with: ...".
*
* ⚠️ NEVER executed by the server. That is a documented invariant, not an oversight:
* running it would turn a config file into a code-execution surface. A proposal to
* execute this on enable is deliberately deferred to its own change so the trust
* model can be decided on its own merits rather than inside a refactor.
*/
command: Partial<Record<'linux' | 'darwin' | 'wsl' | 'win32', string>>;
/** Package name for an npm-installable CLI. Display/tooling metadata only. */
npmPackage?: string;
docsUrl?: string;
};
}
// ---------------------------------------------------------------------------
// Environment
// ---------------------------------------------------------------------------
export interface CliEnv {
/** `export K=V` in the bash prelude. Values are literals or engine values, never secrets. */
exports: Array<{ name: string; value: string | { engine: EngineValue }; when?: Cond }>;
/** `unset K` — e.g. claude's CLAUDECODE, the truecolor CLIs' NO_COLOR. */
unset: string[];
/**
* NAMES ONLY. Values are read from the server's own process.env and pushed via
* `tmux setenv`, so a secret is structurally unable to reach the command line.
*/
tmuxSetenvKeys: string[];
/** NAMES ONLY, forwarded as `docker exec -e NAME`. */
dockerExecEnvNames: string[];
/**
* Env vars set via `tmux setenv` from a LAUNCH PARAM rather than from the server's own
* environment — for a CLI whose switch is an env var instead of a flag.
*
* DeepSeek's `DSH_PERMISSION_MODE` is the case this exists for. Routing it through a
* declared param (rather than a bespoke configure step) is what lets the ordinary
* `privilegedParams` clamp apply to it: the clamp rewrites the param, and whatever the
* param ends up as is what gets exported.
*
* ⚠️ Values are read from a declared, schema-validated param, never from free text, and
* they reach the pane through `tmux setenv` rather than the command line.
*/
configSetenv?: Array<{ name: string; fromParam: string }>;
/** This entry's contribution to the env-override allowlist. Never widens BLOCKED_ENV_KEYS. */
allowedPrefixes: string[];
allowedKeys: string[];
/**
* Env var carrying a JSON config blob pushed via `tmux setenv` (opencode's
* OPENCODE_CONFIG_CONTENT). Generic so it is not an opencode special case.
*/
configContentVar?: string;
/**
* Names an entry in `SETENV_PROFILES` (profiles.ts): extra `tmux setenv` work that is
* genuinely code-shaped rather than a list of key names.
*
* DeepSeek's status bridge is the only current user. It has to write an executable shim
* to disk (`ensureDeepSeekStatusShim()`), then export the shim's path and this session's
* pane id — a side effect and two computed values, none of which `tmuxSetenvKeys` (a
* list of names forwarded from the server's own env) can express.
*
* Plain secret forwarding stays in `tmuxSetenvKeys` and must NOT move here.
*/
setenvProfile?: string;
}
// ---------------------------------------------------------------------------
// Capabilities
// ---------------------------------------------------------------------------
/**
* The closed set of behavioural switches. Each field replaces an id-check somewhere.
*
* `hooks`, `transcript` and `altScreen` are INDEPENDENT on purpose. The three predicates
* they back (`hooksAvailableForMode`, `isExternalCliMode`, `isAltScreenStripMode`) describe
* three different, deliberately unequal sets, and deriving any one from another has already
* caused a real bug — a `shell` session has no hooks but is not an "external CLI", so
* `!isExternalCliMode()` wrongly accepted `until=stop` on it and hung for the full timeout.
* Keeping them as separate fields makes that invariant structural rather than commented.
*/
export interface CliCapabilities {
/**
* Non-Claude run mode that uses its own TUI and output format (`isExternalCliMode`):
* no Claude transcript, no hooks, no Claude-format token/BashTool parsing. An explicit
* field rather than derived from `hooks`/`kind`, precisely because it must stay
* independent — see this interface's own doc comment.
*/
external: boolean;
/**
* How to read this CLI's own TUI for whether it is mid-turn.
*
* Codeman infers a working agent from the pane, so the two strings it needs are the
* ones that differ per CLI: the glyph on the composer row, and the status line the CLI
* draws while a turn runs. Holding them here is what lets a non-Claude CLI report work
* at all — `external` used to gate the whole detector, so every external CLI reported
* itself permanently idle even mid-turn.
*
* `promptGlyph` only ARMS the idle confirmation and is never on its own evidence that a
* turn ended, because a CLI redraws its composer throughout a turn. `workingLine` is
* the evidence, and `_confirmIdle` consults it before believing the pane went quiet.
*
* An entry that omits this field keeps Codeman's historical behaviour: the Claude glyph
* arms the confirmation and the Claude working line answers it. Leave it out for a CLI
* whose TUI nobody has characterised, and its sessions report work exactly as before.
*/
workDetect?: {
/** The glyph this CLI draws on its composer row, e.g. Claude's `❯`, Codex's `›`. */
promptGlyph: string;
/** Source of a regex matching the status line this CLI draws while a turn runs. */
workingLine: string;
};
/** No direct-PTY fallback: the CLI must run inside tmux (secrets ride tmux setenv). */
requiresMux: boolean;
/**
* Whether `stop`/`blocked` wait signals can ever fire for this CLI.
*
* ⚠️ A TRI-STATE, not a boolean, because for one CLI this is a per-SESSION question:
* 'none' — no hook signals, ever (every external CLI, and `shell`).
* 'always' — the CLI installs Codeman's hooks (claude).
* 'supervised' — the CLI REPORTS its own idle/working/blocked state to a supervisor
* over a generic env-gated contract, and Codeman is that supervisor
* (deepseek, via deepseek-status-shim.ts). Definitive rather than
* inferred, so it earns real signals — but the session can disarm the
* bridge (`deepSeekConfig.statusReporting: false`), and a docker or
* remote session cannot reach it at all.
*
* That last case is why `hooksAvailableForMode()` takes per-session options and why
* every call site must pass `sessionHookOptions(session)`. Answering from the mode alone
* would promise a `stop` that never arrives, which is the infinite-wait-dressed-as-a-
* timeout the predicate exists to prevent.
*/
hooks: 'none' | 'always' | 'supervised';
/**
* Which transcript reader, if any, understands this CLI's on-disk history.
*
* `deepseek-zstd` is the odd one out: dsh writes zstd-compressed session files and
* appends ONE FRAME PER WRITE, so it needs a reader that walks frame headers itself
* rather than the stock decoder. It exists because the pane segmenter served dsh's
* ASCII-art splash as the worker's first answer.
*/
transcript: 'claude-jsonl' | 'codex-rollout' | 'deepseek-zstd' | 'omp-jsonl' | 'none';
/**
* 'strip-full' — alt-screen + erase-scrollback + mouse DECSETs stripped (Ink TUIs).
* 'strip-mux-only' — only tmux's own attach-time smcup (the safe default).
* 'preserve' — leave everything (a direct-PTY shell running vim/less/htop).
*/
altScreen: 'strip-full' | 'strip-mux-only' | 'preserve';
echo: {
policy: 'buffer' | 'predict' | 'off';
/** How the local-echo overlay locates the composer row. */
anchor: { kind: 'glyph'; glyph: string; offset: number } | { kind: 'cursor' } | { kind: 'none' };
/** Names a PREDICT_PROFILES key. Unknown or absent degrades to 'buffer', never to broken. */
predictProfile?: string;
};
/** Forwarding the wheel to the CLI's own transcript. 'never' keeps local scrollback. */
wheelForward: { mode: 'never' | 'version-gated'; minVersion?: string };
keyboardAccessory: 'agent' | 'shell';
/** Multi-user: this CLI is a raw shell, so its commands need the privileged gate. */
privilegedCommandGate: boolean;
startMode: 'interactive' | 'shell';
stripInkBloat: boolean;
ralph: boolean;
respawn: boolean;
effort: boolean;
agentSkillInjection: boolean;
statusLineTelemetry: boolean;
/** Where a model override is delivered. Claude uniquely writes settings.local.json. */
model: { source: 'flag' | 'claude-settings-file' | 'none'; param?: string };
/**
* Params a non-granted multi-user owner may not set freely, and what they are forced to.
* Data-driven so a CUSTOM CLI's bypass flag is clampable exactly like codex's.
*
* `materializeWhenAbsent` distinguishes two real shapes, not one:
* - only-if-sent (false/omitted; codex, antigravity, grok): the CLI's own
* absent-config default already spawns safe, so the clamp should only touch
* a config the caller actually sent.
* - materialize (true; gemini, pi): the absent-config default is ITSELF unsafe
* for a non-granted owner (gemini defaults to `yolo`; pi's absent default is
* an interactive trust prompt the session user could just answer "yes" to),
* so the clamp must CREATE a config object even when none was sent.
*
* ⚠️ `param` names the LAUNCH PARAM, like every other `param` in this file — never the
* legacy wire field. The clamp translates it through `legacyConfigAliases` on the way out,
* the same hop `env.configSetenv` makes. The two names coincide for most entries and
* DELIBERATELY do not for codex (`bypassApprovals` here, `dangerouslyBypassApprovals` on
* the wire), which is what keeps the distinction visible. `schema.ts` rejects an entry
* naming a param it never declared, because getting this wrong is a SILENT no-op: no load
* error, no failing test, the clamp just stops clamping.
*/
privilegedParams: Array<{ param: string; clampTo: boolean | string; materializeWhenAbsent?: boolean }>;
/**
* Env var names a non-granted multi-user owner may not set at all, DROPPED from
* `envOverrides` before spawn.
*
* ⚠️ This is a second, structurally different privileged surface from `privilegedParams`
* above, and one cannot substitute for the other. `privilegedParams` clamps a field on a
* per-CLI config object, which reaches the CLI as an argv flag. These clamp env vars,
* which reach it through `tmux setenv` — a path no argv clamp can see.
*
* DeepSeek is why this exists. Its permission switch IS an env var
* (`DSH_PERMISSION_MODE`), not a flag, so a config-level clamp alone leaves a real
* multi-user control with nothing enforcing it. Worse, `DSH_*` is an allowlisted
* `envOverrides` prefix and `applyEnvOverrides()` runs AFTER the per-CLI env configure
* step, so a non-granted owner sending that key on the SAME request would land last and
* hand back exactly the privilege the config clamp just removed.
*
* Dropping (rather than rewriting) is deliberate: the value then falls through to what
* the CLI's own env configuration exports, which is already the clamped one.
*
* The other two DeepSeek keys are here for reasons worth keeping written down:
* - `DSH_HOME` points the launcher at a profile tree whose plugin code runs at BOOT,
* before any approval row could apply.
* - `DEEPSEEK_BASE_URL` would redirect the server's OWN forwarded `DEEPSEEK_API_KEY`
* to a host of the caller's choosing.
*
* Every other CLI's bypass is a command-line flag reachable only through its config
* object, which is why `privilegedParams` alone is the whole gate for them.
*/
privilegedEnvKeys: string[];
/** Version gates referenced by `capabilityGate` conditions. */
gates: Record<string, { minVersion: string; failClosed: boolean }>;
/** Cap on a single terminal frame, when this CLI needs a tighter one than the default. */
maxFrameBytes?: number;
}
// ---------------------------------------------------------------------------
// Location overlays (remote SSH / docker)
// ---------------------------------------------------------------------------
/** Docker credential seeding policy — which host dirs are copied or shared into a container. */
export interface CliCredStore {
rel: string;
shareDirs?: string[];
shareFiles?: string[];
seedFiles?: string[];
seedWhole?: boolean;
}
export interface CliOverlays {
/**
* The remote/docker DEFAULT pane command: just the CLI invocation (e.g. `claude
* --dangerously-skip-permissions`), independent of each location's own wrapping
* (remote: login-shell `-c`; docker: `exec`). Absent `command` = the bare
* `discovery.binaries[0]`. `disabled: true` = this location has no story for this CLI at
* all (docker for `shell`) — distinct from "no override", which still gets a default.
*/
remote?: { command?: string } | { disabled: true };
/**
* `rootCommand` is the same invocation for a container whose exec user is uid 0. Only
* declare it when the normal `command` would be REFUSED as root: claude's carries
* `--dangerously-skip-permissions`, which Claude Code rejects outright under root, and
* the rejection is visible only inside the container, so the pane dies with no clue on
* the outside. Codeman's own base image runs a non-root user and never selects this; an
* ADOPTED container belongs to its owner and is frequently root. Absent = use `command`.
*/
docker?: { command?: string; rootCommand?: string } | { disabled: true };
/**
* ⚠️ DECLARED-FOR-LATER, unlike `remote`/`docker` above, which are live.
*
* The Docker credential-seeding path still reads its own `CRED_STORES` table in
* `docker-hosts.ts`, because this shape cannot yet express that table: it allows ONE store
* per CLI, and the live table needs two for gemini (`.gemini` for the CLI's own auth plus
* `.config/gcloud` for Vertex), while deepseek's entry here declares none at all even
* though `.dsh` is seeded. Wiring it therefore means making this an ARRAY and correcting
* those two entries — a change to credential seeding, which is both the highest-consequence
* thing in this file to get wrong and the least covered by tests, since every docker IO
* path is no-op'd under vitest. It belongs in its own change, measured against a real
* container.
*/
credStore?: CliCredStore;
}
// ---------------------------------------------------------------------------
// The entry
// ---------------------------------------------------------------------------
/**
* ⚠️ DECLARED-FOR-LATER: fields no code reads yet.
*
* `shortBadge`, `accent`, `overlays.credStore`, `capabilities.echo`, `capabilities.wheelForward`,
* `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` all describe FRONTEND
* behaviour, and the frontend is deliberately untouched by the change that introduced this
* registry — `app.js`, `terminal-ui.js`, `styles.css` and friends keep their own
* hand-authored per-CLI rules, and moving them is its own piece of work with its own way of
* being verified (a mobile/browser suite the CI gate cannot see).
*
* They are declared now because each entry should describe its CLI completely, and because
* transcribing them while the hand-written source is still on screen is when the values are
* actually known. But an unread field is a promise, not a fact: nothing enforces that
* `echo.policy` here matches `_updateLocalEchoState`'s fallthrough, or that `accent` matches
* the gradient CSS paints. Treat every value in this group as TRANSCRIBED, not authoritative,
* and re-measure against the frontend before wiring one up.
*
* The rest of the interface is live: something reads it, and `test/cli-registry-*.test.ts`
* pins what it does with it.
*/
export interface CliEntry {
id: CliId;
label: string;
/** Two-ish character tab badge, e.g. 'OC'. */
shortBadge: string;
/** Single hex colour. CSS derives every per-CLI gradient from it via --cli-accent. */
accent: string;
enabled: boolean;
/** Set by the loader from the shipped catalog; a user entry can never claim it. */
stock: boolean;
order: number;
/** 'shell' unlocks the raw-shell code paths; everything else is an agent CLI. */
kind: 'agent' | 'shell';
discovery: CliDiscovery;
launch: CliLaunch;
env: CliEnv;
capabilities: CliCapabilities;
overlays: CliOverlays;
}
/**
* The on-disk shape of ~/.codeman/clis.json — overrides and custom entries only, never the
* full catalog. Small and hand-readable by design.
*
* ⚠️ READ-ONLY in this build. Nothing here writes this file: there is no settings UI and no
* write API yet, so there is nothing to persist. That also means importing the registry
* (and therefore `schemas.ts`, which validates against it) performs no filesystem writes —
* an import side effect worth not having.
*/
export interface CliRegistryFile {
schemaVersion: number;
/**
* Stock ids already introduced to this install — the ratchet that lets one file both gain
* newly-shipped CLIs on upgrade AND remember that the user disabled one.
*
* Read and IGNORED here, and never written: the ratchet only earns its keep once a CLI
* can be disabled, which needs the write API. Declared now purely so a file written by a
* later version still loads cleanly under this one instead of failing `.strict()`.
*/
seededStockIds?: string[];
/** Keyed by id: a partial override of a stock entry, or a complete custom entry. */
clis: Record<string, unknown>;
}
+159 -128
View File
@@ -7,7 +7,8 @@
* @module config/dependency-registry
*/
import { PI_VERSION_REGEX } from '../utils/pi-cli-resolver.js';
import { enabledClis } from './cli-registry/registry.js';
import { compileVersionRegex } from './cli-registry/patterns.js';
export type ProbeEnvironment = 'linux' | 'darwin' | 'win32' | 'wsl';
@@ -56,134 +57,164 @@ export interface ToolDependency {
const ALL: ProbeEnvironment[] = ['linux', 'darwin', 'wsl', 'win32'];
export const DEPENDENCY_REGISTRY: ToolDependency[] = [
{
id: 'node',
label: 'Node.js',
category: 'core',
required: true,
minVersion: '22.0.0',
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['node'], versionArg: '--version' } }],
installHint: { linux: 'https://nodejs.org', darwin: 'brew install node', wsl: 'https://nodejs.org' },
},
{
id: 'claude',
label: 'Claude CLI',
category: 'core',
required: false,
usedBy: ['Claude Code sessions (default backend)'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['claude'], versionArg: '--version' } }],
installHint: { linux: 'https://docs.claude.com/claude-code', darwin: 'https://docs.claude.com/claude-code' },
},
{
id: 'tmux',
label: 'tmux',
category: 'core',
required: true,
resolvers: [{ match: ['linux', 'darwin', 'wsl'], resolver: { kind: 'path', bins: ['tmux'], versionArg: '-V' } }],
installHint: { linux: 'sudo apt install tmux', darwin: 'brew install tmux', wsl: 'sudo apt install tmux' },
},
{
id: 'opencode',
label: 'OpenCode CLI',
category: 'core',
required: false,
usedBy: ['OpenCode sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['opencode'], versionArg: '--version' } }],
},
{
id: 'codex',
label: 'Codex CLI',
category: 'core',
required: false,
usedBy: ['Codex sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['codex'], versionArg: '--version' } }],
},
{
id: 'gemini',
label: 'Gemini CLI',
category: 'core',
required: false,
usedBy: ['Gemini sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['gemini'], versionArg: '--version' } }],
},
{
id: 'antigravity',
label: 'Antigravity CLI',
category: 'core',
required: false,
usedBy: ['Antigravity sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['agy'], versionArg: '--version' } }],
},
{
id: 'pi',
label: 'Pi CLI',
category: 'core',
required: false,
usedBy: ['Pi sessions'],
// The only entry that requires a version match, for the same reason
// pi-cli-resolver.ts probes: `pi` is a short generic name (Raspberry Pi tooling,
// personal scripts), so a `which pi` hit alone is not the coding agent. Both sides
// share PI_VERSION_REGEX, so the doctor and the run mode cannot drift into telling
// the user opposite things about the same binary.
resolvers: [
{
match: ALL,
resolver: {
kind: 'path',
bins: ['pi'],
versionArg: '--version',
versionRegex: PI_VERSION_REGEX,
requireVersionMatch: true,
/**
* The doctor's ROW IDENTITY for a CLI, where it differs from the registry id.
*
* These are two separate contracts and they have never been the same thing: `codeman doctor`
* prints a tool table whose ids predate the registry, and `dsh` names the BINARY while the
* run mode is `deepseek`. Keeping the historical id here means the doctor's output does not
* shift under a refactor that was supposed to change nothing a user can see.
*
* `usedBy` is likewise preserved verbatim rather than generated, because the strings are
* shown to the user and claude's does not follow the pattern.
*/
const DOCTOR_ROW_OVERRIDES: Record<string, { id?: string; label?: string; usedBy: string[] }> = {
claude: { usedBy: ['Claude Code sessions (default backend)'] },
opencode: { usedBy: ['OpenCode sessions'] },
codex: { usedBy: ['Codex sessions'] },
gemini: { usedBy: ['Gemini sessions'] },
antigravity: { usedBy: ['Antigravity sessions'] },
pi: { usedBy: ['Pi sessions'] },
grok: { usedBy: ['Grok sessions'] },
// Both the id and the label are historical: `dsh` names the binary, and the doctor has
// always spelled this row out in full rather than as `${label} CLI`.
deepseek: { id: 'dsh', label: 'DeepSeek Harness CLI', usedBy: ['DeepSeek sessions'] },
};
/**
* Build one `codeman doctor` row per enabled CLI, straight from its registry entry.
*
* This replaces eight hand-written rows that had to be kept in step with the run modes by
* hand — and were not: an earlier draft of this refactor silently dropped the Grok and
* DeepSeek rows, so `codeman doctor` stopped reporting two shipped CLIs at all. Deriving
* the list makes that class of omission impossible.
*
* ⚠️ The version regex is compiled through `compileVersionRegex()`, NOT `new RegExp()`. It
* is a config-supplied pattern, so it goes through the same length cap and
* nested-quantifier rejection the argv engine applies; the doctor runs it over command
* output exactly like the resolver does, and skipping the guard here would leave one
* unguarded path into a user-supplied regex.
*
* ⚠️ Sharing the entry's regex with the resolver is what stops the doctor and the run mode
* telling the user opposite things about the same binary — the Dependencies panel reporting
* "Pi CLI ✓" on a box where Run Pi stays hidden.
*/
function cliDependencyEntries(): ToolDependency[] {
const rows: ToolDependency[] = [];
for (const cli of enabledClis()) {
// `shell` has no binary of its own (the login shell is resolved at spawn time), so
// there is nothing for the doctor to probe.
const bin = cli.discovery.binaries[0];
if (!bin) continue;
const override = DOCTOR_ROW_OVERRIDES[cli.id as string];
const version = cli.discovery.version;
const versionRegex = version?.regex ? (compileVersionRegex(version.regex) ?? undefined) : undefined;
rows.push({
id: override?.id ?? (cli.id as string),
label: override?.label ?? `${cli.label} CLI`,
category: 'core',
required: false,
usedBy: override?.usedBy ?? [`${cli.label} sessions`],
resolvers: [
{
match: ALL,
resolver: {
kind: 'path',
bins: [bin],
versionArg: version?.arg ?? '--version',
versionRegex,
// Only meaningful for a CLI whose binary name is short, generic or squatted
// (pi, grok, dsh): a bare `which` hit there is not evidence of the right
// program, so a version mismatch means MISSING rather than unknown-version.
requireVersionMatch: version?.requireVersionMatch,
},
},
},
],
},
{
id: 'libreoffice',
label: 'LibreOffice',
category: 'office',
required: false,
usedBy: ['document preview', 'thumbnails'],
resolvers: [
{
match: ['linux', 'darwin', 'wsl'],
resolver: { kind: 'path', bins: ['libreoffice', 'soffice'], versionArg: '--version' },
},
],
installHint: { linux: 'sudo apt install libreoffice', darwin: 'brew install --cask libreoffice' },
},
{
id: 'pdftoppm',
label: 'pdftoppm',
category: 'office',
required: false,
usedBy: ['document preview', 'PDF/Office first-page thumbnails'],
// poppler's pdftoppm prints its version to stderr; presence is what matters here.
resolvers: [
{ match: ['linux', 'darwin', 'wsl'], resolver: { kind: 'path', bins: ['pdftoppm'], versionArg: '-v' } },
],
installHint: {
linux: 'sudo apt install poppler-utils',
darwin: 'brew install poppler',
wsl: 'sudo apt install poppler-utils',
],
installHint: cli.discovery.install.command,
});
}
return rows;
}
/**
* The tools `codeman doctor` probes, resolved AT CALL TIME.
*
* ⚠️ A FUNCTION, not a module-level const, and for the same reason `sessionModeSchema()` and
* `allowedEnvPrefixes()` are functions: `cliDependencyEntries()` reads the CLI registry, and
* a const would have frozen the doctor's rows at first import while every schema resolved
* per parse. A CLI enabled while the server was running — or a `reloadCliRegistry()` — then
* moved the run menu and the validation but never the doctor, which would keep reporting the
* catalog as it stood when something first imported this module. Building the array per call
* costs a handful of object literals on a command that shells out to probe binaries anyway.
*/
export function dependencyRegistry(): ToolDependency[] {
return [
{
id: 'node',
label: 'Node.js',
category: 'core',
required: true,
minVersion: '22.0.0',
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['node'], versionArg: '--version' } }],
installHint: { linux: 'https://nodejs.org', darwin: 'brew install node', wsl: 'https://nodejs.org' },
},
},
{
id: 'msoffice',
label: 'MS Office',
category: 'office',
required: false,
usedBy: ['document preview', 'thumbnails'],
resolvers: [
{
match: ['wsl', 'win32'],
resolver: {
kind: 'windows-side',
appDirs: ['Microsoft Office/root/Office16'],
exes: ['WINWORD.EXE', 'POWERPNT.EXE', 'EXCEL.EXE'],
{
id: 'tmux',
label: 'tmux',
category: 'core',
required: true,
resolvers: [{ match: ['linux', 'darwin', 'wsl'], resolver: { kind: 'path', bins: ['tmux'], versionArg: '-V' } }],
installHint: { linux: 'sudo apt install tmux', darwin: 'brew install tmux', wsl: 'sudo apt install tmux' },
},
...cliDependencyEntries(),
{
id: 'libreoffice',
label: 'LibreOffice',
category: 'office',
required: false,
usedBy: ['document preview', 'thumbnails'],
resolvers: [
{
match: ['linux', 'darwin', 'wsl'],
resolver: { kind: 'path', bins: ['libreoffice', 'soffice'], versionArg: '--version' },
},
],
installHint: { linux: 'sudo apt install libreoffice', darwin: 'brew install --cask libreoffice' },
},
{
id: 'pdftoppm',
label: 'pdftoppm',
category: 'office',
required: false,
usedBy: ['document preview', 'PDF/Office first-page thumbnails'],
// poppler's pdftoppm prints its version to stderr; presence is what matters here.
resolvers: [
{ match: ['linux', 'darwin', 'wsl'], resolver: { kind: 'path', bins: ['pdftoppm'], versionArg: '-v' } },
],
installHint: {
linux: 'sudo apt install poppler-utils',
darwin: 'brew install poppler',
wsl: 'sudo apt install poppler-utils',
},
],
},
];
},
{
id: 'msoffice',
label: 'MS Office',
category: 'office',
required: false,
usedBy: ['document preview', 'thumbnails'],
resolvers: [
{
match: ['wsl', 'win32'],
resolver: {
kind: 'windows-side',
appDirs: ['Microsoft Office/root/Office16'],
exes: ['WINWORD.EXE', 'POWERPNT.EXE', 'EXCEL.EXE'],
},
},
],
},
];
}
+15
View File
@@ -40,6 +40,21 @@ const INSTANCE_SUFFIX = CODEMAN_INSTANCE ? `-${CODEMAN_INSTANCE}` : '';
/** Default tmux socket for this instance. `CODEMAN_TMUX_SOCKET` still overrides. */
export const DEFAULT_TMUX_SOCKET = `codeman${INSTANCE_SUFFIX}`;
/** Characters tmux accepts in a `-L` socket name. */
export const SAFE_TMUX_SOCKET_PATTERN = /^[a-zA-Z0-9_.-]+$/;
/**
* This instance's tmux socket: the `CODEMAN_TMUX_SOCKET` override when it is a
* safe name, else the instance default. Every process that runs `tmux -L` has
* to resolve it through here (the server via TmuxManager, the TUI for its
* degraded-mode listing), or a beta instance ends up driving prod's sessions.
*/
export function resolveTmuxSocketName(): string {
const raw = process.env.CODEMAN_TMUX_SOCKET;
if (raw !== undefined && SAFE_TMUX_SOCKET_PATTERN.test(raw)) return raw;
return DEFAULT_TMUX_SOCKET;
}
let _ensured = false;
/**
+2 -1
View File
@@ -8,7 +8,8 @@
* src/web/public/constants.js and deliberately stays at 50k — 100k xterm lines per tab
* is a mobile-memory hazard — so DEFAULT_TERMINAL_SCROLLBACK_LINES stays 50,000 to match.
* The terminalScrollbackLines/terminalBufferMaxBytes/terminalBufferTrimBytes settings keys
* remain schema-validated but inert (a follow-up wires them); only tmuxHistoryLimit is live.
* remain schema-validated but inert (a follow-up wires them); only tmuxHistoryLimit is wired.
* tmux <3.7 applies it to new panes; tmux 3.7+ can also resize live panes.
* All values remain env- and settings-overridable and bounds-clamped via
* resolveTerminalHistoryConfig().
*/
+56 -7
View File
@@ -10,6 +10,8 @@
import { v4 as uuidv4 } from 'uuid';
import { readFile } from 'node:fs/promises';
import { statSync, realpathSync } from 'node:fs';
import { getCli } from '../config/cli-registry/registry.js';
import { resolveCliLaunchError } from '../utils/cli-launcher.js';
import { Session } from '../session.js';
import { applyWorkspaceHooks } from '../hooks-config.js';
import { SseEvent } from '../web/sse-events.js';
@@ -48,7 +50,9 @@ const delay = (ms: number): Promise<void> => new Promise((r) => setTimeout(r, ms
* answer "yes" to, which then loads and EXECUTES repo-local `.pi/extensions` TypeScript,
* so `approveProjectTrust: false` (`--no-approve`) is materialized. Omitting `--approve`
* is NOT a clamp.
* Codex and antigravity need nothing here: their absent config already spawns safe.
* Codex, antigravity, grok and deepseek need nothing here: their absent config already spawns safe
* (grok's bare spawn is its own ask-mode default and deepseek's omits DSH_PERMISSION_MODE
* entirely, leaving the harness on workspace-write, which asks; both switches are only ever sent).
* Granted/admin/single-user get undefined for both, i.e. upstream defaults untouched.
*/
export function clampCronExternalCliConfigs(
@@ -56,9 +60,26 @@ export function clampCronExternalCliConfigs(
ownerGranted: boolean
): { geminiConfig: GeminiConfig | undefined; piConfig: PiConfig | undefined } {
if (ownerGranted) return { geminiConfig: undefined, piConfig: undefined };
// A cron job carries no per-CLI config at all, so ONLY the materialize-when-absent params
// can apply here — an only-if-sent clamp has nothing to clamp. Reading them off the
// registry rather than naming gemini and pi means a future CLI whose bare spawn is unsafe
// is covered the moment its entry says so, instead of silently missing this path.
const entry = getCli(mode);
const aliases = entry?.launch.legacyConfigAliases ?? {};
const materialized: Record<string, unknown> = {};
for (const { param, clampTo, materializeWhenAbsent } of entry?.capabilities.privilegedParams ?? []) {
// Same registry-param → legacy-wire-field hop the HTTP clamp makes. Neither gemini's
// `approvalMode` nor pi's `approveProjectTrust` is aliased today, so this changes nothing
// now — but the two are DIFFERENT namespaces, and writing the raw param here would make
// this path stop clamping the moment one of them gained an alias, silently.
if (materializeWhenAbsent) materialized[aliases[param] ?? param] = clampTo;
}
const has = Object.keys(materialized).length > 0;
const field = entry?.launch.legacyConfigField;
return {
geminiConfig: mode === 'gemini' ? { approvalMode: 'auto_edit' } : undefined,
piConfig: mode === 'pi' ? { approveProjectTrust: false } : undefined,
geminiConfig: has && field === 'geminiConfig' ? (materialized as GeminiConfig) : undefined,
piConfig: has && field === 'piConfig' ? (materialized as PiConfig) : undefined,
};
}
@@ -385,7 +406,7 @@ export class CronService {
// Section 6.3: re-resolve the owner's grant at FIRE time (it may have been revoked
// since create). Gates shell/launchCommand AND clamps the external-CLI bypass below.
const ownerGranted = await canUsernameRunPrivilegedCommands(job.owner);
if ((job.agentType === 'shell' || job.launchCommand) && !ownerGranted) {
if ((getCli(job.agentType)?.capabilities.privilegedCommandGate || job.launchCommand) && !ownerGranted) {
return this.failRun(job, run, 'Owner lacks the can-bypass-permissions grant for shell/launchCommand jobs');
}
@@ -393,11 +414,37 @@ export class CronService {
let session: Session;
try {
const mode = job.agentType;
// A LAUNCHER CLI's binary is not its agent, so "installed" is not "runnable": without
// this, a job on a box carrying only dsh's stock web/headless profiles spawns a bare
// `dsh` that boots a profile unable to drive a pane, and the prompt is typed into a
// logging server or a dead pane instead of failing the run with an actionable message.
//
// ⚠️ Scoped to `discovery.launcherProfile`, which is byte-identical to the
// `mode === 'deepseek'` check this replaces (dsh is the only launcher today) and
// generalises to the next one. Deliberately NOT every CLI: cron has never pre-flighted
// a merely-missing binary, and doing so replaces tmux-manager's own not-found throw
// ("Session launch failed") with a different message for claude and shell. An earlier
// draft of this line was unscoped and did exactly that — three cron tests caught it.
if (getCli(mode)?.discovery.launcherProfile !== undefined) {
const cronLaunchError = await resolveCliLaunchError(mode);
if (cronLaunchError) return this.failRun(job, run, cronLaunchError);
}
const globalNice = await this.deps.getGlobalNiceConfig();
const modelConfig = await this.deps.getModelConfig();
const claudeModeConfig = await this.deps.getClaudeModeConfig();
const effectiveClaudeMode = await resolveClaudeModeForUsername(claudeModeConfig.claudeMode, job.owner);
const model = mode !== 'shell' ? modelConfig?.defaultModel || undefined : undefined;
// Cron carries no per-CLI config object, so the only model it can supply is the global
// default — and only to a CLI that takes a model at all.
//
// ⚠️ `!== 'none'` is the faithful reading of the `mode !== 'shell' && mode !== 'deepseek'`
// ladder this replaces: those two are exactly the entries declaring `model.source: 'none'`
// (shell has no model; deepseek's is a profile composition entry, not a session flag).
// NOT `=== 'claude-settings-file'`, which is the HTTP route's question — there, every
// external CLI reads its model from its own config object earlier in the chain, so only
// claude reaches the global default. Cron has no such config, so the same expression
// means something different here.
const model =
getCli(mode)?.capabilities.model.source !== 'none' ? modelConfig?.defaultModel || undefined : undefined;
// Section 6.3: materialize the safe default for a non-granted owner (see
// clampCronExternalCliConfigs — cron sends no per-CLI config, so the CLI's own
// spawn default is what would otherwise apply).
@@ -425,7 +472,7 @@ export class CronService {
piConfig,
owner: job.owner,
});
this.deps.addSession(session);
await this.deps.addSession(session);
this.store.incrementSessionsCreated();
this.deps.persistSessionState(session);
await this.deps.setupSessionListeners(session);
@@ -546,7 +593,9 @@ export class CronService {
private sendPromptWhenReady(sessionId: string, prompt: string, job: CronJob, run: CronJobRun): void {
setImmediate(() => {
const poll = async (): Promise<void> => {
if (job.agentType !== 'shell') {
// A shell pane is ready the moment it exists; an agent CLI has a TUI to paint
// first. That is the `kind` the registry already records, not a fact about shell.
if (getCli(job.agentType)?.kind !== 'shell') {
for (let attempt = 0; attempt < CRON_READY_MAX_ATTEMPTS; attempt++) {
await delay(500);
const s = this.deps.sessions.get(sessionId);
+3
View File
@@ -45,6 +45,8 @@ export interface WebLaunchOptions {
host: string;
port: number;
https: boolean;
/** Reverse-proxy sub-path prefix (normalized: '' for root, or '/foo'). */
basePath?: string;
titleHostname?: string;
allowUnauthenticatedNetwork?: boolean;
multiuser?: boolean;
@@ -87,6 +89,7 @@ export interface DaemonStatus {
export function buildWebArgs(options: WebLaunchOptions): string[] {
const args = ['web', '--host', options.host, '--port', String(options.port)];
if (options.https) args.push('--https');
if (options.basePath) args.push('--base-url', options.basePath);
if (options.titleHostname) args.push('--title-hostname', options.titleHostname);
if (options.allowUnauthenticatedNetwork) args.push('--allow-unauthenticated-network');
if (options.multiuser) args.push('--multiuser');
+271
View File
@@ -0,0 +1,271 @@
/**
* @fileoverview The DeepSeek Harness -> Codeman status bridge.
*
* ## Why this exists
*
* Every external CLI mode before this one (opencode, codex, gemini, antigravity,
* pi, grok) is READINESS-GUESSED: Codeman watches the PTY go quiet and infers a
* turn ended. Claude is the exception, because Claude Code fires real hooks. The
* DeepSeek Harness TUI gives us a third option, and a much better one than
* guessing: the community terminal front door already reports its own lifecycle
* to an owning supervisor, and it does so through a fully GENERIC, env-var-gated
* contract it inherited from Herdr (herdr.dev).
*
* When all three of `HERDR_ENV=1`, `HERDR_BIN_PATH` and `HERDR_PANE_ID` are set,
* the TUI shells out on every state change:
*
* "$HERDR_BIN_PATH" pane report-agent "$HERDR_PANE_ID" \
* --source custom:dsh-tui --agent dsh-tui \
* --state idle|working|blocked [--message <text>] --seq <n>
*
* and treats exit code 0 as "delivered" (retrying with backoff otherwise). So
* Codeman points `HERDR_BIN_PATH` at the script below and gets DEFINITIVE
* idle/working/blocked signals for dsh sessions: real respawn triggers, real
* `wait`/`wait-output` stop+blocked signals, and real Approvals Inbox items,
* on par with Claude's hooks rather than with output stabilization.
*
* This is an interface implementation, not an impersonation: we implement the
* one verb (`pane report-agent`) that the contract defines, and nothing on the
* machine ever executes a real `herdr` binary — `HERDR_BIN_PATH` is our own
* script, in our own data dir. `HERDR_ENV=1` is the flag the TUI checks to know
* a supervisor is present; a supervisor IS present, it is Codeman.
*
* ## Why it is generated rather than committed
*
* The shim must be an executable file at a stable absolute path in every
* install shape: a git clone (where `scripts/` exists), an `npm i -g aicodeman`
* (where `files` ships only `dist` plus two named scripts), and any
* `CODEMAN_INSTANCE`. Writing it into the data dir at session-create time makes
* one code path cover all of them, single-sources the content here in TS, and
* follows the precedent of `self-update-runner.sh`. It is rewritten whenever the
* embedded version marker changes, so an upgraded Codeman refreshes a stale shim
* without the user knowing it exists.
*
* @module deepseek-status-shim
*/
import { chmodSync, mkdirSync, readFileSync, renameSync, rmSync, writeFileSync } from 'node:fs';
import { dirname } from 'node:path';
import { dataPath } from './config/instance.js';
/**
* Bumped whenever SHIM_SOURCE changes. The marker is embedded in the generated
* file, so `ensureDeepSeekStatusShim()` can tell a current shim from one written
* by an older Codeman and rewrite only when needed (rather than rewriting on
* every session create, or — worse — leaving a stale one in place forever).
*/
const SHIM_VERSION = 3;
const SHIM_MARKER = `codeman-dsh-status-shim v${SHIM_VERSION}`;
/**
* Mapping from the harness's three lifecycle states to Codeman hook events.
*
* - `blocked` -> `permission_prompt`: the TUI reports blocked when a tool
* approval or an `ask_user_question` questionnaire is on screen, which is
* exactly the red "needs you" alert and an answerable Approvals Inbox item.
* - `idle` -> `stop`: the definitive end-of-turn signal, the one respawn and the
* wait endpoints care about.
* - `working` -> `agent_working`: a turn STARTED. Codeman infers "working" from
* PTY output well enough on its own, but the event is what RESOLVES a pending
* approval when the user answers a dialog in the terminal instead of in the
* inbox. Without it a dsh session's red alert would survive until the next
* `stop`, which is the exact stuck-alert bug the claude path already had to
* fix once (and the pane-capture staleness sweep that fixed it there is
* Claude-dialog-shaped, so it cannot help here).
*/
export const DEEPSEEK_STATE_TO_HOOK_EVENT: Readonly<Record<string, string>> = Object.freeze({
idle: 'stop',
blocked: 'permission_prompt',
working: 'agent_working',
});
/**
* The generated script.
*
* Constraints it must satisfy, each learned from an existing Codeman hook bug:
* - **TLS**: `CODEMAN_API_URL` is loopback HTTPS with a self-signed cert on
* `--https`/tailscale installs, so certificate verification is disabled for
* the request. Without this the whole bridge dies silently, exactly as the
* claude hook curls did before they grew `-k`.
* - **Secret**: the hook-secret file is read AT EXECUTION TIME, never baked in,
* so rotation needs no respawn and the value never lands on a command line.
* - **Exit codes**: 0 means delivered. Anything else makes the TUI retry with
* backoff, so transport failures self-heal, but an unknown verb or an
* unmapped state exits 0 to avoid a pointless retry storm over something that
* will never succeed.
* - **Timeout**: bounded below the caller's own 2s budget, so we lose the race
* deliberately rather than being killed mid-flight.
*/
const SHIM_SOURCE = `#!/usr/bin/env node
// ${SHIM_MARKER}
// GENERATED BY CODEMAN — do not edit. Rewritten from src/deepseek-status-shim.ts
// whenever its version marker changes.
//
// Implements the one verb the DeepSeek Harness TUI's supervisor contract uses:
// pane report-agent <paneId> --state <idle|working|blocked> [--message <t>] ...
// and forwards it to this Codeman instance as a hook event.
import { readFileSync } from 'node:fs'
import http from 'node:http'
import https from 'node:https'
const STATE_TO_EVENT = ${JSON.stringify(DEEPSEEK_STATE_TO_HOOK_EVENT)}
const TIMEOUT_MS = 1500
const argv = process.argv.slice(2)
const flag = (name) => {
const i = argv.indexOf(name)
return i >= 0 && i + 1 < argv.length ? argv[i + 1] : undefined
}
// Unknown verb: succeed silently. Retrying could never make it succeed, and a
// non-zero exit here would make the caller retry four times per state change.
if (argv[0] !== 'pane' || argv[1] !== 'report-agent') process.exit(0)
const event = STATE_TO_EVENT[String(flag('--state') ?? '')]
if (!event) process.exit(0)
// The pane id we hand the TUI IS the Codeman session id, but prefer the ambient
// env: it is set by the same code that set HERDR_PANE_ID, so a TUI that mangles,
// truncates or re-uses the pane argument still reports against the right session.
// NOT a security boundary, and do not read it as one: the agent runs IN this pane
// and can invoke the shim with CODEMAN_SESSION_ID unset and any argv it likes.
// That buys it nothing it did not already have, since the hook-secret file is
// readable from the same pane and any process there can POST /api/hook-event
// directly. Attribution here is about accidents, not adversaries.
const sessionId = process.env.CODEMAN_SESSION_ID || argv[2]
const apiUrl = process.env.CODEMAN_API_URL
if (!sessionId || !apiUrl) process.exit(1)
let secret = ''
try {
secret = readFileSync(process.env.CODEMAN_HOOK_SECRET_FILE || '', 'utf-8').trim()
} catch {
// Missing file: the loopback bypass still applies when no tunnel is running.
}
// The contract's ordering token: the TUI retries failed deliveries with
// backoff, so a stale report can land AFTER a newer one. Forwarded so the
// server can drop out-of-order arrivals instead of, say, resolving an
// approval with a retried 'working' while the harness sits blocked.
const seq = Number(flag('--seq'))
const body = JSON.stringify({
event,
sessionId,
data: {
source: 'dsh-status-shim',
agent: flag('--agent') || 'dsh',
...(Number.isFinite(seq) ? { seq } : {}),
...(flag('--message') ? { message: flag('--message') } : {}),
},
})
let url
try {
url = new URL('/api/hook-event', apiUrl)
} catch {
process.exit(1)
}
const transport = url.protocol === 'https:' ? https : http
const req = transport.request(
{
protocol: url.protocol,
hostname: url.hostname,
port: url.port,
path: url.pathname,
method: 'POST',
timeout: TIMEOUT_MS,
headers: {
'Content-Type': 'application/json',
'Content-Length': Buffer.byteLength(body),
'X-Codeman-Hook-Secret': secret,
},
// Loopback HTTPS with a self-signed cert (--https / tailscale installs).
rejectUnauthorized: false,
},
(res) => {
res.resume()
const status = res.statusCode ?? 0
// 2xx: delivered. 4xx: PERMANENT — a 401 (missing/rotated secret) or 429
// can never be fixed by retrying, and each retry feeds the auth-failure
// rate-limit bucket, so a single misconfigured dsh session could 429 the
// hook endpoint for the whole instance (killing every claude session's
// real hooks). Exit 0 so the TUI does not retry; only transport errors
// and 5xx stay retryable.
process.exit(status >= 200 && status < 500 ? 0 : 1)
}
)
req.on('timeout', () => {
req.destroy()
process.exit(1)
})
req.on('error', () => process.exit(1))
req.end(body)
`;
/** Absolute path of the generated shim for this instance. */
export function deepSeekStatusShimPath(): string {
return dataPath('dsh-status-shim.mjs');
}
let ensuredThisProcess = false;
/**
* Write the shim if it is missing or stale, and return its path.
*
* Idempotent and cheap: after the first call in a process it does nothing, and
* even the first call only rewrites when the on-disk marker differs. Never
* throws — a data dir that cannot be written is a degraded status bridge, not a
* failed session start, so callers fall back to output-stabilization readiness
* by receiving null.
*/
export function ensureDeepSeekStatusShim(): string | null {
const path = deepSeekStatusShimPath();
if (ensuredThisProcess) return path;
try {
let current = '';
try {
current = readFileSync(path, 'utf-8');
} catch {
// Missing — fall through to the write.
}
if (!current.includes(SHIM_MARKER)) {
mkdirSync(dirname(path), { recursive: true });
// Temp + rename, not a plain write: the TUI can be executing this exact
// path at the moment an upgraded Codeman refreshes it (every state change
// runs it, and session create is when the rewrite happens), and a reader
// that catches a half-written file gets a syntax error, exits non-zero,
// and is retried four times per state change for a file that will never
// parse. rename(2) is atomic within the directory, so a concurrent exec
// sees either the old shim or the new one, never a truncated one.
// Same reasoning as the state-store writes; pid-suffixed so two instances
// sharing a data dir cannot collide on the temp name.
const tempPath = `${path}.${process.pid}.tmp`;
try {
writeFileSync(tempPath, SHIM_SOURCE, { mode: 0o700 });
// The mode argument only applies when writeFileSync CREATES the file, so
// a leftover temp from a crashed run would keep its old permissions.
chmodSync(tempPath, 0o700);
renameSync(tempPath, path);
} catch (err) {
rmSync(tempPath, { force: true });
throw err;
}
}
// Re-assert the mode even when the content matched: a shim that lost its
// executable bit (a restored backup, a copied data dir) would make every
// report fail, and the TUI would retry four times per state change forever.
chmodSync(path, 0o700);
ensuredThisProcess = true;
return path;
} catch (err) {
console.warn(`[DeepSeek] Could not install the status shim at ${path}: ${(err as Error).message}`);
return null;
}
}
/** Test seam: forget the per-process memo so a fresh temp HOME is re-provisioned. */
export function resetDeepSeekStatusShimForTest(): void {
ensuredThisProcess = false;
}
+696
View File
@@ -0,0 +1,696 @@
/**
* @fileoverview Reading a DeepSeek Harness (`dsh`) session transcript off disk.
*
* ## Why this exists
*
* `GET /api/sessions/:id/last-response` is how an agent (and the Response
* Viewer) reads what a worker actually said. For Claude it comes from
* `~/.claude/projects/**`, for Codex from `~/.codex/sessions/**`, and for every
* other external CLI it comes from segmenting the terminal buffer, because
* those CLIs write nothing a reader could open.
*
* dsh is not in that last group: it writes a complete, structured JSONL
* transcript per session. Falling back to the pane for it was measurably wrong
* rather than merely coarse — dsh-TUI paints a full-screen splash, so the pane
* segmenter answered a `last-response` call for a fresh dsh session with the
* ASCII-art logo:
*
* {"text":"✦dsh-TUI v0.8.8█▀▀▀▄█▀▀▀▀█▀▀▀▀█▀▀▀▄█▀▀▀▀…","hasContext":true}
*
* which an agent polling for a worker's answer reads as an answer. This module
* is the real source: it locates the session's transcript, decodes it, and
* returns the last turn's text.
*
* ## The three things that make dsh transcripts unlike codex rollouts
*
* **1. One zstd FRAME per append, not one zstd stream.** The file is
* `session.jsonl.zstd`, and dsh appends by compressing each batch of lines into
* its own frame and writing it at the end. `zstd -dc` handles that (frames
* concatenate by definition), but Node's `zlib.zstdDecompress()` and
* `createZstdDecompress()` both stop at the first frame end: measured on a real
* 56-line transcript, Node returned 158 bytes / 1 line where the CLI returned
* 43,747 bytes / 56 lines. That is a silent truncation to the session header —
* every call would have reported "no answer yet" forever. `decodeZstdFrames()`
* below walks the frame headers itself and decompresses each frame, and
* `test/deepseek-transcript.test.ts` pins it against multi-frame fixtures.
*
* **2. The user's prompts are mixed with injected context.** Every turn also
* writes a `user/message` whose source is a plugin (the runtime-context
* snapshot: sandbox policy, approval policy, cwd). Those are `source.kind ===
* 'plugin'`; a real prompt is `source.kind === 'user'`. Rendering the plugin
* ones would show the agent its own boilerplate back as the user's words.
*
* **3. A failed turn is not an empty turn.** `turn/end` carries
* `reason.kind === 'error'` with the provider's message. Returning `""` there
* makes an agent poll `last-response` fifteen times and conclude the worker
* never answered, when the truth ("the provider rejected the request") was on
* disk the whole time. A turn that ends in an error and produced no text
* answers with that error, prefixed so it can never be mistaken for the model's
* own words.
*
* Verified against `dsh 0.1.1-rc.2` + `@deepseek-harness-tui/dsh-tui 0.8.8`.
*/
import { promises as fs } from 'node:fs';
import { homedir } from 'node:os';
import { join } from 'node:path';
import * as zlib from 'node:zlib';
/**
* One rendered block, in the shape the Response Viewer already speaks (see
* `web/response-viewer-transcript.ts`). Imported as a type only — this module
* must stay usable from the session layer without dragging web/ into it.
*/
export interface DeepSeekTranscriptBlock {
kind: 'prompt' | 'response' | 'status' | 'tool';
label: 'Prompt' | 'Response' | 'Status' | 'Tool';
role: 'user' | 'assistant';
text: string;
}
export interface DeepSeekTranscriptResult {
/** Last turn's answer (or its error, prefixed). Empty before the first turn. */
text: string;
/** ISO timestamp of the event `text` came from, or '' when unknown. */
timestamp: string;
/** Rendered blocks, oldest first. Only built when the caller asks for them. */
blocks: DeepSeekTranscriptBlock[];
/** dsh's own session id, from the header line. */
sessionId?: string;
/** Workspace the harness recorded for the session. */
cwd?: string;
}
/**
* zstd decompression is a RUNTIME capability here, not an import.
*
* Node grew `zlib` zstd support in 22.15 (and `@types/node` still does not
* declare it), while Codeman's floor is Node 22.0. So it is resolved through a
* narrow cast and checked before use: on an older 22.x a dsh session keeps the
* pane-segmenter behaviour it had before this module existed instead of
* throwing on every `last-response` call.
*/
type ZstdDecompressSync = (buf: Buffer) => Buffer;
const zstdDecompressSync: ZstdDecompressSync | undefined = (
zlib as unknown as { zstdDecompressSync?: ZstdDecompressSync }
).zstdDecompressSync;
/** Whether this Node can decode the compressed transcripts dsh writes. */
export function zstdSupported(): boolean {
return typeof zstdDecompressSync === 'function';
}
/** zstd frame magic (RFC 8878 §3.1.1). */
const ZSTD_MAGIC = 0xfd2fb528;
/** Skippable-frame magic range: 0x184D2A50..0x184D2A5F. */
const ZSTD_SKIPPABLE_LO = 0x184d2a50;
const ZSTD_SKIPPABLE_HI = 0x184d2a5f;
const DID_FIELD_SIZE = [0, 1, 2, 4];
const FCS_FIELD_SIZE = [0, 2, 4, 8];
/**
* Byte ranges of the zstd frames in `buf`, in order.
*
* Walks frame headers and block headers only — no decompression — so the cost
* is proportional to the number of blocks, not to the content. Stops (rather
* than throws) at the first thing it cannot parse, so a transcript still being
* appended to mid-write yields every whole frame before the torn tail instead
* of failing the whole read.
*
* ⚠️ Splitting on the magic bytes instead would be wrong: the 4-byte sequence
* can occur inside compressed data, and a false split corrupts everything after
* it. The block walk is what makes the boundaries exact.
*/
export function zstdFrameRanges(buf: Buffer): Array<[number, number]> {
const ranges: Array<[number, number]> = [];
let offset = 0;
while (offset + 4 <= buf.length) {
const magic = buf.readUInt32LE(offset);
if (magic >= ZSTD_SKIPPABLE_LO && magic <= ZSTD_SKIPPABLE_HI) {
if (offset + 8 > buf.length) break;
const end = offset + 8 + buf.readUInt32LE(offset + 4);
if (end > buf.length || end <= offset) break;
offset = end;
continue;
}
if (magic !== ZSTD_MAGIC) break;
let p = offset + 4;
if (p >= buf.length) break;
const descriptor = buf[p] as number;
p += 1;
const fcsFlag = descriptor >> 6;
const singleSegment = (descriptor >> 5) & 1;
const hasChecksum = (descriptor >> 2) & 1;
const dictIdFlag = descriptor & 3;
if (!singleSegment) p += 1; // window descriptor
p += DID_FIELD_SIZE[dictIdFlag] as number;
// FCS is absent for flag 0 UNLESS Single_Segment is set, where it is 1 byte.
p += fcsFlag === 0 ? (singleSegment ? 1 : 0) : (FCS_FIELD_SIZE[fcsFlag] as number);
if (p > buf.length) break;
let lastBlock = false;
let torn = false;
while (!lastBlock) {
if (p + 3 > buf.length) {
torn = true;
break;
}
const header = (buf[p] as number) | ((buf[p + 1] as number) << 8) | ((buf[p + 2] as number) << 16);
p += 3;
lastBlock = (header & 1) === 1;
const blockType = (header >> 1) & 3;
const blockSize = header >> 3;
if (blockType === 3) {
torn = true; // reserved: refuse rather than guess
break;
}
p += blockType === 1 ? 1 : blockSize; // RLE stores a single byte
if (p > buf.length) {
torn = true;
break;
}
}
if (torn) break;
if (hasChecksum) p += 4;
if (p > buf.length) break;
ranges.push([offset, p]);
offset = p;
}
return ranges;
}
/**
* Decode a possibly multi-frame zstd buffer. A buffer that does not start with
* a zstd magic is passed through unchanged, which is what lets the same reader
* open a plain `session.jsonl` (dsh writes one when compression is off).
*
* A frame that fails to decompress truncates the decode THERE rather than
* failing it: everything decoded before it is kept, so a half-written tail
* frame does not cost the caller the whole conversation. (Not "skipped" — a
* frame after a corrupt one is never reached, which is the safe reading: dsh
* appends, so a bad frame means everything after it is suspect too.)
*/
export function decodeZstdFrames(buf: Buffer): string {
if (buf.length < 4) return buf.toString('utf8');
const magic = buf.readUInt32LE(0);
if (magic !== ZSTD_MAGIC && (magic < ZSTD_SKIPPABLE_LO || magic > ZSTD_SKIPPABLE_HI)) {
return buf.toString('utf8');
}
if (!zstdDecompressSync) return '';
const parts: Buffer[] = [];
for (const [start, end] of zstdFrameRanges(buf)) {
try {
parts.push(zstdDecompressSync(buf.subarray(start, end)));
} catch {
// Torn or corrupt frame: keep what decoded before it.
break;
}
}
return Buffer.concat(parts).toString('utf8');
}
interface DshEvent {
type?: string;
seq?: number | null;
time?: number;
data?: Record<string, unknown>;
}
function asRecord(value: unknown): Record<string, unknown> | undefined {
return value && typeof value === 'object' && !Array.isArray(value) ? (value as Record<string, unknown>) : undefined;
}
function asArray(value: unknown): unknown[] {
return Array.isArray(value) ? value : [];
}
/**
* Strip a leaked reasoning prefix.
*
* Some providers stream reasoning into the same text block and close it with
* `</think>` without ever opening it (measured on a local deepseek-v4-flash
* route: `"I'll read the file first.</think>\n\nThe add function is…"`). The
* closing tag is the only reliable boundary, so everything up to the LAST one
* goes. A block with no tag is returned untouched.
*/
function stripReasoningPrefix(text: string): string {
const close = text.lastIndexOf('</think>');
return close === -1 ? text : text.slice(close + '</think>'.length);
}
/** `stripReasoning` is for ASSISTANT content only: a user prompt containing a
* literal `</think>` (someone pasting a transcript, say) must render whole. */
function textOfContent(content: unknown, stripReasoning = true): string {
const parts: string[] = [];
for (const entry of asArray(content)) {
const block = asRecord(entry);
if (!block) continue;
if (block.type === 'text' && typeof block.text === 'string') {
parts.push(stripReasoning ? stripReasoningPrefix(block.text) : block.text);
}
}
return parts.join('').trim();
}
function toolCallsOfContent(content: unknown): string[] {
const calls: string[] = [];
for (const entry of asArray(content)) {
const block = asRecord(entry);
if (!block || block.type !== 'tool-call') continue;
const name = typeof block.name === 'string' ? block.name : 'tool';
const args = typeof block.arguments === 'string' ? block.arguments : JSON.stringify(block.arguments ?? {});
calls.push(`${name}(${args})`);
}
return calls;
}
/** Flatten a `tool/result` message down to its text payload. */
function textOfToolResult(message: unknown): string {
const parts: string[] = [];
for (const entry of asArray(asRecord(message)?.content)) {
const block = asRecord(entry);
if (!block) continue;
if (block.type === 'text' && typeof block.text === 'string') parts.push(block.text);
if (block.type === 'tool-result') {
for (const inner of asArray(block.content)) {
const innerBlock = asRecord(inner);
if (innerBlock?.type === 'text' && typeof innerBlock.text === 'string') parts.push(innerBlock.text);
}
}
}
return parts.join('\n').trim();
}
function isoTime(time: unknown): string {
return typeof time === 'number' && Number.isFinite(time) ? new Date(time).toISOString() : '';
}
interface TurnAccumulator {
/** Finalized `assistant/message` text, in step order. */
finalized: Map<number, string>;
/** Steps that produced a finalized message AT ALL. ⚠️ Not the same as a
* non-empty entry in `finalized`: a step whose whole reply was reasoning
* strips to `''`, and without this the deltas — which are NOT stripped at
* write time — would be resurrected in its place, putting the model's raw
* `</think>` monologue in front of the caller (measured). */
finalizedSteps: Set<number>;
/** Streamed deltas per step, used only where no finalized message landed. */
streamed: Map<number, string>;
/** Step order as encountered, so a reply reads in the order it was produced. */
steps: number[];
timestamp: string;
/** Pre-rendered "Turn error: …" / "Turn ended: …" line, when the turn did not
* end with `completed`. */
ending?: string;
}
function ensureStep(turn: TurnAccumulator, step: number): void {
if (!turn.steps.includes(step)) turn.steps.push(step);
}
function turnText(turn: TurnAccumulator): string {
const parts: string[] = [];
for (const step of turn.steps) {
// Deltas are only consulted for a step the model never finalized — a step
// that has both would otherwise render its text twice.
const text = turn.finalizedSteps.has(step)
? (turn.finalized.get(step) ?? '')
: stripReasoningPrefix(turn.streamed.get(step) ?? '');
if (text.trim()) parts.push(text.trim());
}
return parts.join('\n\n').trim();
}
/**
* Parse a decoded dsh transcript.
*
* `text` is the LAST TURN's answer, not the last assistant message anywhere in
* the file: a turn that errored after an earlier turn answered must not hand
* back the earlier turn's text as though it were this turn's reply.
*/
export function parseDeepSeekTranscript(raw: string, options: { blocks?: boolean } = {}): DeepSeekTranscriptResult {
const wantBlocks = options.blocks === true;
const blocks: DeepSeekTranscriptBlock[] = [];
const turns = new Map<number, TurnAccumulator>();
const turnOrder: number[] = [];
let sessionId: string | undefined;
let cwd: string | undefined;
const getTurn = (n: number): TurnAccumulator => {
let turn = turns.get(n);
if (!turn) {
turn = { finalized: new Map(), finalizedSteps: new Set(), streamed: new Map(), steps: [], timestamp: '' };
turns.set(n, turn);
turnOrder.push(n);
}
return turn;
};
for (const line of raw.split('\n')) {
if (!line.trim()) continue;
let event: DshEvent;
try {
event = JSON.parse(line) as DshEvent;
} catch {
continue; // a torn tail line, or a frame we could not decode
}
const data = asRecord(event.data) ?? {};
const turnNo = typeof data.turn === 'number' ? data.turn : 0;
const stepNo = typeof data.step === 'number' ? data.step : 0;
switch (event.type) {
case 'session': {
const header = event as unknown as Record<string, unknown>;
if (typeof header.id === 'string') sessionId = header.id;
if (typeof header.cwd === 'string') cwd = header.cwd;
break;
}
case 'user/message': {
// ⚠️ Only a real prompt. The plugin-sourced twin is the runtime-context
// snapshot dsh injects every turn (sandbox policy, approvals, cwd).
if (asRecord(data.source)?.kind !== 'user') break;
if (!wantBlocks) break;
const text = textOfContent(data.content, false);
if (text) blocks.push({ kind: 'prompt', label: 'Prompt', role: 'user', text });
break;
}
case 'assistant/message': {
const message = asRecord(data.message);
const turn = getTurn(turnNo);
ensureStep(turn, stepNo);
const text = textOfContent(message?.content);
if (message) turn.finalizedSteps.add(stepNo);
if (text) {
turn.finalized.set(stepNo, text);
turn.timestamp = isoTime(event.time) || turn.timestamp;
if (wantBlocks) blocks.push({ kind: 'response', label: 'Response', role: 'assistant', text });
}
if (wantBlocks) {
for (const call of toolCallsOfContent(message?.content)) {
blocks.push({ kind: 'tool', label: 'Tool', role: 'assistant', text: call });
}
}
break;
}
case 'assistant/chunk': {
const chunk = asRecord(data.chunk);
if (chunk?.type !== 'text-delta' || typeof chunk.text !== 'string') break;
const turn = getTurn(turnNo);
ensureStep(turn, stepNo);
turn.streamed.set(stepNo, (turn.streamed.get(stepNo) ?? '') + chunk.text);
break;
}
case 'text-chunks': {
// The batched form of the same deltas (dsh coalesces once a stream gets
// going). ⚠️ These carry `seq: null`, so file order is the only order.
const turn = getTurn(turnNo);
ensureStep(turn, stepNo);
const texts = asArray(data.texts)
.filter((t): t is string => typeof t === 'string')
.join('');
if (texts) turn.streamed.set(stepNo, (turn.streamed.get(stepNo) ?? '') + texts);
break;
}
case 'tool/result': {
if (!wantBlocks) break;
const text = textOfToolResult(data.message);
if (text) blocks.push({ kind: 'tool', label: 'Tool', role: 'assistant', text });
break;
}
case 'turn/end': {
const turn = getTurn(turnNo);
const reason = asRecord(data.reason);
if (reason && reason.kind !== 'completed') {
// Two different things wear this field: a provider failure
// (`kind:'error'` with a message) and an ordinary early stop
// (`kind:'max-tokens'`, measured live). Calling the second one an
// error would misreport a truncated but real answer.
const error = asRecord(reason.error);
const message = typeof error?.message === 'string' ? error.message : undefined;
const kind = typeof reason.kind === 'string' ? reason.kind : 'unknown';
turn.ending = message ? `Turn error: ${message}` : `Turn ended: ${kind}`;
if (wantBlocks) {
blocks.push({ kind: 'status', label: 'Status', role: 'assistant', text: turn.ending });
}
}
turn.timestamp = isoTime(event.time) || turn.timestamp;
break;
}
default:
break;
}
}
const lastTurn = turnOrder.length > 0 ? turns.get(turnOrder[turnOrder.length - 1] as number) : undefined;
let text = lastTurn ? turnText(lastTurn) : '';
// A turn that failed and said nothing answers with its failure, labelled so
// it can never read as the model's own words. Without this an agent polls
// `last-response` fifteen times and concludes the worker never answered.
if (!text && lastTurn?.ending) text = lastTurn.ending;
return { text, timestamp: lastTurn?.timestamp ?? '', blocks, sessionId, cwd };
}
/**
* `$DSH_HOME` for one session: a per-session override wins (`DSH_HOME` is an
* allowlisted `envOverrides` prefix, and pointing a worker at its own profile
* tree is a documented thing to do), then the server's own environment, then
* `~/.dsh`. Reading the wrong tree does not fail loudly — it silently finds no
* transcript — so this must resolve exactly the way the spawn did.
*/
/* ⚠️ The override is EPHEMERAL: `envOverrides` is applied at spawn and exported
* through `tmux setenv`, but is deliberately not persisted to state.json (it can
* carry provider keys). A session that overrode `DSH_HOME` and then outlived a
* server restart therefore resolves to the default tree and finds no transcript
* — it reads as "nothing said yet" rather than as another session's answer,
* because every candidate is matched on its recorded `cwd`. */
export function resolveDeepSeekHome(session: { deepSeekHomeOverride?: string }): string {
const override = session.deepSeekHomeOverride;
if (override && override.trim()) return override.trim();
const fromEnv = process.env.DSH_HOME;
if (fromEnv && fromEnv.trim()) return fromEnv.trim();
return join(homedir(), '.dsh');
}
/**
* How far apart a session's start and its transcript's `createdAt` may be and
* still be the same session. dsh writes the header within ~2 s of pane start
* (measured); 60 s absorbs a cold profile boot without ever reaching a sibling
* started minutes later.
*/
const PAIRING_WINDOW_MS = 60_000;
/** Transcript file names dsh has used, newest convention first. */
const TRANSCRIPT_FILES = ['session.jsonl.zstd', 'session.jsonl'];
/**
* Locate the transcript for a session.
*
* dsh buckets sessions by a mangled cwd (`--home-you-code-app--`) and then by
* its own session id, and the id form has changed between versions (`<uuid>`
* and `session-<uuid>` both exist on disk here). ⚠️ So the mangling is NOT
* reproduced: every candidate's own header line carries `cwd`, which is
* authoritative, and matching on it is immune to the next naming change.
*
* Pairing a Codeman session with ITS transcript then has one hard rule and one
* ladder. The rule: a transcript created BEFORE this session started belongs to
* an earlier conversation in the same directory and is never eligible. Measured
* cost of getting that wrong — a freshly spawned worker answered its very first
* `last-response` with the PREVIOUS session's reply, which is worse than saying
* nothing, because an agent cannot tell a stale answer from a fresh one.
*
* The ladder, once the older ones are out:
*
* 1. a transcript whose header `createdAt` sits within `PAIRING_WINDOW_MS` of
* this session's start — that is this pane's own boot, and it stays right
* even when a sibling session is running in the same case directory;
* 2. otherwise the newest transcript created after this session started;
* 3. otherwise nothing.
*
* ⚠️ The boot transcript wins for as long as it exists on disk — deliberately,
* and even over a LATER transcript in the same workspace. Step 2 cannot tell a
* `/new` from a sibling session that started later in the same directory, so
* preferring newest-eligible would hand a worker its busier sibling's reply
* (the exact bug the hard rule above was measured against, one seat over).
* The cost of that choice: after an interactive `/new` in a dsh tab, this
* reader keeps serving the pre-`/new` conversation (the same session's own
* earlier turns — stale, never foreign); step 2 is reached only when no
* boot-window transcript exists. Worker fleets never `/new`, so they only
* ever see step 1.
*/
export async function findDeepSeekTranscript(options: {
dshHome: string;
workingDir: string;
startedAt?: number;
}): Promise<string | null> {
const sessionsDir = join(options.dshHome, 'sessions');
let buckets: string[];
try {
buckets = (await fs.readdir(sessionsDir, { withFileTypes: true }))
.filter((entry) => entry.isDirectory())
.map((entry) => entry.name);
} catch {
return null;
}
const candidates: Array<{ path: string; mtimeMs: number }> = [];
for (const bucket of buckets) {
const bucketPath = join(sessionsDir, bucket);
let sessions: string[];
try {
sessions = (await fs.readdir(bucketPath, { withFileTypes: true }))
.filter((entry) => entry.isDirectory())
.map((entry) => entry.name);
} catch {
continue;
}
for (const sessionDir of sessions) {
for (const file of TRANSCRIPT_FILES) {
const path = join(bucketPath, sessionDir, file);
const stat = await fs.stat(path).catch(() => null);
if (!stat || !stat.isFile() || stat.size === 0) continue;
candidates.push({ path, mtimeMs: stat.mtimeMs });
break;
}
}
}
if (candidates.length === 0) return null;
candidates.sort((a, b) => b.mtimeMs - a.mtimeMs);
const startedAt = options.startedAt ?? 0;
// Slack in both directions: the harness writes its header a beat after the
// pane starts, and mtimes on a shared clock are not worth trusting to the ms.
const floor = startedAt > 0 ? startedAt - PAIRING_WINDOW_MS : 0;
let laterMatch: string | null = null;
for (const candidate of candidates) {
const header = await readTranscriptHeader(candidate.path);
if (!header || header.cwd !== options.workingDir) continue;
// No usable header timestamp: fall back to the file's own mtime, which is
// still enough to keep a pre-session transcript out.
const createdAt = header.createdAt ?? candidate.mtimeMs;
if (createdAt < floor) continue;
if (startedAt > 0 && Math.abs(createdAt - startedAt) <= PAIRING_WINDOW_MS) return candidate.path;
if (!laterMatch) laterMatch = candidate.path;
}
return laterMatch;
}
/**
* Read only the first frame of a transcript, which is where the header line
* lives. Bounded: a candidate scan must never decompress every conversation on
* the box to answer one `last-response` call.
*/
async function readTranscriptHeader(path: string): Promise<{ cwd?: string; id?: string; createdAt?: number } | null> {
let handle;
try {
handle = await fs.open(path, 'r');
} catch {
return null;
}
try {
const head = Buffer.alloc(65536);
const { bytesRead } = await handle.read(head, 0, head.length, 0);
if (bytesRead === 0) return null;
const text = decodeZstdFrames(head.subarray(0, bytesRead));
const firstLine = text.split('\n').find((line) => line.trim());
if (!firstLine) return null;
const parsed = JSON.parse(firstLine) as { type?: string; cwd?: string; id?: string; createdAt?: number };
if (parsed.type !== 'session') return null;
return {
cwd: parsed.cwd,
id: parsed.id,
createdAt: typeof parsed.createdAt === 'number' ? parsed.createdAt : undefined,
};
} catch {
return null;
} finally {
await handle.close().catch(() => {});
}
}
/** Hard ceiling on a transcript read. A long agent run is a few hundred KB; a
* file past this is pathological and is not worth a synchronous decode. */
const MAX_TRANSCRIPT_BYTES = 64 * 1024 * 1024;
/**
* Memo of the last few decoded transcripts, keyed on (path, mtime, size,
* blocks). The skill's `last_text` polls once per second, and each poll used
* to zstdDecompressSync + reparse the WHOLE file on the event loop even when
* nothing had been appended — a multi-MB transcript made that a repeated
* ~100ms-class stall on the single-threaded server. A poll that finds the
* file unchanged now costs one stat. Insertion-order eviction; tiny, because
* an entry only earns its keep while a session is being actively polled.
*/
const parseMemo = new Map<string, DeepSeekTranscriptResult>();
const PARSE_MEMO_MAX = 16;
/** Test seam: a fixture that rewrites one path in place inside a single mtime
* tick would otherwise read its predecessor back out of the memo. */
export function resetDeepSeekTranscriptMemoForTest(): void {
parseMemo.clear();
}
/**
* Read one dsh session's last answer.
*
* ⚠️ The two empty outcomes are deliberately different, because the caller must
* treat them differently:
*
* - `null` means **this reader cannot run here** (a Node without zstd), and is
* the signal to fall back to the pane segmenter.
* - an empty `text` means **read fine, nothing said yet** — no transcript for
* this workspace, or a turn still in flight.
*
* Collapsing the two would put the ASCII-art splash back in front of an agent
* that is polling for a worker's first answer.
*/
export async function readDeepSeekLastResponse(
session: { workingDir: string; createdAt?: Date | number; deepSeekHomeOverride?: string },
options: { blocks?: boolean } = {}
): Promise<DeepSeekTranscriptResult | null> {
const createdAt = session.createdAt instanceof Date ? session.createdAt.getTime() : session.createdAt;
// dsh compresses by default, so a Node without zstd can read nothing here.
// That is the one case the pane is still the better answer.
if (!zstdSupported()) return null;
const empty: DeepSeekTranscriptResult = { text: '', timestamp: '', blocks: [] };
const path = await findDeepSeekTranscript({
dshHome: resolveDeepSeekHome(session),
workingDir: session.workingDir,
startedAt: typeof createdAt === 'number' ? createdAt : undefined,
});
if (!path) return empty;
const stat = await fs.stat(path).catch(() => null);
if (!stat || stat.size > MAX_TRANSCRIPT_BYTES) return empty;
const memoKey = `${path}|${stat.mtimeMs}|${stat.size}|${options.blocks ? 1 : 0}`;
const memoized = parseMemo.get(memoKey);
if (memoized) return memoized;
let buf: Buffer;
try {
buf = await fs.readFile(path);
} catch {
return empty;
}
const result = parseDeepSeekTranscript(decodeZstdFrames(buf), options);
if (parseMemo.size >= PARSE_MEMO_MAX) {
const oldest = parseMemo.keys().next().value;
if (oldest !== undefined) parseMemo.delete(oldest);
}
parseMemo.set(memoKey, result);
return result;
}
+283
View File
@@ -0,0 +1,283 @@
/**
* @fileoverview Supervises the one background `dsh web` process behind the Run
* menu's "DeepSeek web UI..." entry.
*
* The shortcut originally started the server inside an ordinary SHELL SESSION,
* on the reasoning that Codeman already knows how to supervise those: it was
* visible, scrollable, killable, and died with its tab, and nothing new had to
* own a long-lived HTTP server. That reasoning was sound and the result was
* still wrong in use — clicking "open the DeepSeek web UI" spawned a terminal
* tab the user never asked for, next to the web tab they did, and the terminal
* was noise every time after the first.
*
* So the server moves here instead: one child process, no session, no tab.
* What that buys back has to be paid for explicitly, which is what this module
* is:
*
* - **Exactly one.** A second click reuses the running server rather than
* racing it for a port. The old shell-session flow could not do this at all,
* because two clicks were simply two sessions.
* - **Restarted when the authority changes.** `--trusted-host` fences dsh's
* `/api` against the browser authority, and a Codeman reachable at both
* loopback and a tailnet name has two. Whoever asks last wins, because the
* asker is by definition the origin about to load the page.
* - **Killed on shutdown.** A detached child that outlived Codeman would hold
* its port against the next start, which is exactly the EADDRINUSE this
* feature already got wrong once.
* - **Failures reported, not swallowed.** The shell tab used to be where the
* stack trace landed. With no tab, the spawn's own output is captured and
* handed back to the caller instead.
*/
import { spawn, type ChildProcess } from 'node:child_process';
import { createServer } from 'node:net';
import { join } from 'node:path';
import { getErrorMessage } from './types.js';
/**
* Where the port search starts, and how far it walks.
*
* 3080 is `dsh web`'s own default, so it is the friendly first choice — and
* emphatically not a fixed port. DeepSeek's web UI is a thing users run
* themselves, which makes the default precisely the port most likely to be
* taken already; hardcoding it made this feature die with EADDRINUSE against
* the user's own server.
*/
const PORT_BASE = 3080;
const PORT_SPAN = 40;
/** How long a freshly spawned server gets to answer before we call it failed. */
const READY_TIMEOUT_MS = 30_000;
const READY_POLL_MS = 400;
/** Grace between SIGTERM and SIGKILL when stopping the tree. */
const KILL_GRACE_MS = 3_000;
/** Bound on captured child output, so a chatty boot cannot grow without limit. */
const OUTPUT_CAP = 16_384;
export interface DeepSeekWebStatus {
running: boolean;
port: number | null;
url: string | null;
/** Browser authority this server was started to trust (`--trusted-host`). */
authority: string | null;
}
interface RunningServer {
child: ChildProcess;
port: number;
authority: string;
output: () => string;
}
let current: RunningServer | null = null;
/**
* True when nothing holds `port` on loopback.
*
* Binding is the only honest test: a connect probe cannot tell "free" from
* "listening but not answering yet", and this runs moments before `dsh web`
* binds the same port. It is inherently racy, which is why the caller still
* waits for the server to actually answer before reporting success.
*/
async function isLoopbackPortFree(port: number): Promise<boolean> {
return new Promise((resolve) => {
const probe = createServer();
probe.once('error', () => resolve(false));
probe.once('listening', () => probe.close(() => resolve(true)));
probe.listen(port, '127.0.0.1');
});
}
async function findFreePort(): Promise<number | null> {
for (let port = PORT_BASE; port < PORT_BASE + PORT_SPAN; port++) {
if (await isLoopbackPortFree(port)) return port;
}
return null;
}
/** Does the server answer HTTP yet? Any status counts: dsh may 4xx a bare GET. */
async function answersHttp(port: number): Promise<boolean> {
try {
await fetch(`http://127.0.0.1:${port}/`, { signal: AbortSignal.timeout(2_000) });
return true;
} catch {
return false;
}
}
/**
* Signal the whole process group.
*
* `dsh web` boots a plugin tree and fans out, so signalling only the direct
* child leaves survivors holding the port. Same negative-pid escalation as
* `runGit()` in git-clone.ts and the profile installer.
*/
function killTree(child: ChildProcess, signal: NodeJS.Signals): void {
try {
if (child.pid) process.kill(-child.pid, signal);
} catch {
try {
child.kill(signal);
} catch {
/* already gone */
}
}
}
export function getDeepSeekWebStatus(): DeepSeekWebStatus {
if (!current) return { running: false, port: null, url: null, authority: null };
return {
running: true,
port: current.port,
url: `http://127.0.0.1:${current.port}`,
authority: current.authority,
};
}
/** Stop the background server, if one is running. Safe to call when none is. */
export async function stopDeepSeekWeb(): Promise<void> {
const running = current;
current = null;
if (!running) return;
await new Promise<void>((resolve) => {
let done = false;
const finish = () => {
if (done) return;
done = true;
clearTimeout(hard);
resolve();
};
running.child.once('exit', finish);
killTree(running.child, 'SIGTERM');
const hard = setTimeout(() => {
killTree(running.child, 'SIGKILL');
finish();
}, KILL_GRACE_MS);
});
}
type StartResult = { ok: true; port: number; url: string; reused: boolean } | { ok: false; error: string };
/**
* Serializes concurrent starts. Two POSTs racing (two devices, or a double
* click while the first boots) used to both see `current === null`, pick the
* SAME free port, and spawn twice: the loser died on EADDRINUSE while its exit
* handler nulled the singleton out from under the winner, leaving a live
* `dsh web` nothing tracked or killed — the exact orphan this module exists to
* prevent. The second caller now simply waits and reuses the first's server.
*/
let startLock: Promise<unknown> = Promise.resolve();
/**
* Start (or reuse) the background `dsh web` for `authority`.
*
* @param dshDir directory holding the resolved `dsh` binary.
* @param authority browser authority to pass as `--trusted-host`.
*/
export function startDeepSeekWeb(dshDir: string, authority: string): Promise<StartResult> {
const run = startLock.then(
() => startDeepSeekWebLocked(dshDir, authority),
() => startDeepSeekWebLocked(dshDir, authority)
);
startLock = run.then(
() => undefined,
() => undefined
);
return run;
}
async function startDeepSeekWebLocked(dshDir: string, authority: string): Promise<StartResult> {
// Reuse only when the running server is BOTH healthy and fenced for the
// authority now asking. A server trusting the other origin renders a page
// whose every API call 403s, which looks like a broken dashboard rather than
// a misconfigured one.
if (current) {
if (current.authority === authority && (await answersHttp(current.port))) {
return { ok: true, port: current.port, url: `http://127.0.0.1:${current.port}`, reused: true };
}
await stopDeepSeekWeb();
}
const port = await findFreePort();
if (port === null) {
return { ok: false, error: `No free port for the DeepSeek web UI in ${PORT_BASE}-${PORT_BASE + PORT_SPAN - 1}` };
}
let child: ChildProcess;
try {
child = spawn(
join(dshDir, 'dsh'),
['web', '--no-open', '--host', '127.0.0.1', '--port', String(port), '--trusted-host', authority],
{
stdio: ['ignore', 'pipe', 'pipe'],
// Own process group so the whole plugin tree can be signalled at once.
detached: true,
env: process.env,
}
);
} catch (err) {
return { ok: false, error: `Failed to start dsh web: ${getErrorMessage(err)}` };
}
// The pipes must be drained whether or not anyone reads them: a full pipe
// blocks the child. Storage is capped; draining is not.
let output = '';
const capture = (chunk: Buffer) => {
if (output.length < OUTPUT_CAP) output += chunk.toString('utf-8');
};
child.stdout?.on('data', capture);
child.stderr?.on('data', capture);
let exited = false;
child.once('exit', () => {
exited = true;
// Only clear if this is still the current server: a restart may have
// already replaced it, and clearing then would drop the live one.
if (current?.child === child) current = null;
});
child.once('error', () => {
exited = true;
if (current?.child === child) current = null;
});
const running: RunningServer = { child, port, authority, output: () => output };
current = running;
const deadline = Date.now() + READY_TIMEOUT_MS;
while (Date.now() < deadline) {
if (exited) {
// Guarded like the exit/error handlers: a concurrent stop (DELETE route,
// shutdown) may already have cleared or replaced the singleton, and an
// unconditional null here would drop a server this call does not own.
if (current === running) current = null;
const tail = output.trim().slice(-800);
return { ok: false, error: tail ? `dsh web exited during startup: ${tail}` : 'dsh web exited during startup' };
}
if (await answersHttp(port)) {
return { ok: true, port, url: `http://127.0.0.1:${port}`, reused: false };
}
await new Promise((r) => setTimeout(r, READY_POLL_MS));
}
// Timeout: kill OUR child. Only route through stopDeepSeekWeb() while the
// singleton is still ours — signalling `current` unconditionally here could
// SIGTERM a healthy server a concurrent actor now owns.
if (current === running) {
await stopDeepSeekWeb();
} else {
killTree(running.child, 'SIGKILL');
}
const tail = output.trim().slice(-800);
return {
ok: false,
error: tail
? `dsh web did not answer on port ${port} within ${READY_TIMEOUT_MS / 1000}s: ${tail}`
: `dsh web did not answer on port ${port} within ${READY_TIMEOUT_MS / 1000}s`,
};
}
/** Test seam: forget any tracked child without signalling it. */
export function resetDeepSeekWebForTest(): void {
current = null;
}
+8 -1
View File
@@ -30,6 +30,7 @@ import { spawn } from 'node:child_process';
import { pipeline } from 'node:stream/promises';
import type { DockerEngine, SessionDocker } from './types.js';
import { runWithConversionLimit } from './document-conversion-limiter.js';
import { isAdoptedContainer } from './docker-hosts.js';
const IS_TEST_MODE = !!process.env.VITEST;
@@ -287,7 +288,13 @@ export async function exportDockerCase(params: {
const bundlePath = join(exportsDir, exportBundleName(caseName, timestamp, mode));
const stageDir = join(exportsDir, `.stage-${caseName}-${timestamp}`);
mkdirSync(stageDir, { recursive: true });
const wasRunning = await isContainerRunning(argv, docker.containerName);
// ⚠️ NEVER pause an ADOPTED container. The freeze exists only to make the committed
// image and the workspace tar mutually consistent, and it is a lifecycle mutation on a
// container that belongs to the user — it stops their processes for however long the
// tar takes. A workspace-only export of an adopted case therefore accepts a live
// filesystem, the same guarantee `tar` gives on any running host directory. Full-image
// export is refused for an adopted case at the route, before reaching here.
const wasRunning = !isAdoptedContainer(docker) && (await isContainerRunning(argv, docker.containerName));
let commitTag: string | undefined;
try {
+410 -24
View File
@@ -23,7 +23,8 @@
import { existsSync, mkdirSync, readFileSync, writeFileSync } from 'node:fs';
import fs from 'node:fs/promises';
import { join, dirname } from 'node:path';
import { dirname, isAbsolute, join, relative, resolve } from 'node:path';
import { enabledCliIds, getCli } from './config/cli-registry/registry.js';
import { fileURLToPath } from 'node:url';
import { homedir } from 'node:os';
import { createHash } from 'node:crypto';
@@ -32,7 +33,6 @@ import { promisify } from 'node:util';
import { dataPath } from './config/instance.js';
import type {
DockerCase,
DockerCommandMode,
DockerEngine,
DockerHost,
DockerNetworkMode,
@@ -55,6 +55,30 @@ export const DEFAULT_AGENT_IMAGE = 'codeman/agent:base';
/** HOME inside the base image (the `agent` user). Cred mounts + hook-secret land under it. */
export const CONTAINER_HOME = '/home/agent';
/**
* Modes the adoption preflight probes for inside an existing container, derived from the
* CLI registry so a newly-enabled CLI is probed without a second list to remember.
*
* No arm for `shell` here: it declares no binary, so `probeAdoptableContainer` drops it
* from the `command -v` list and reports it available unconditionally, which is the same
* answer a special case would have produced.
*/
export function dockerAdoptProbeModes(): SessionMode[] {
return enabledCliIds() as SessionMode[];
}
/**
* The BINARY a mode looks for inside a container. ⚠️ NOT always the mode name:
* `antigravity` ships as `agy` and `deepseek` as `dsh`, so probing by mode name
* would report those two as missing on a container that has them. Same source
* `probeDockerCliVersion` reads, and the same one `defaultDockerCommandForMode`
* launches from — a local table here duplicated the registry with nothing
* keeping the two in step.
*/
function containerBinaryFor(mode: SessionMode): string | undefined {
return getCli(mode)?.discovery.binaries[0];
}
/** Per-case container name prefix. The `case` letters deliberately do NOT matter to
* tmux; this is a DOCKER name (`^[a-zA-Z0-9][a-zA-Z0-9_.-]+$`), and case names are
* already validated `^[a-zA-Z0-9_-]+$`, so `codeman-case-<name>` is always valid. */
@@ -134,19 +158,30 @@ export function dockerContainerName(caseName: string): string {
return `${CONTAINER_NAME_PREFIX}${caseName}`;
}
/** Default pane command per CLI mode (mirror of defaultRemoteCommandForMode). */
export function defaultDockerCommandForMode(mode: SessionMode): string {
const commands: Record<DockerCommandMode, string> = {
shell: 'exec bash -l',
// Mirror the LOCAL claude default so the in-container agent runs non-interactively.
claude: 'exec claude --dangerously-skip-permissions',
opencode: 'exec opencode',
codex: 'exec codex',
gemini: 'exec gemini',
antigravity: 'exec agy',
pi: 'exec pi',
};
return commands[mode as DockerCommandMode] || commands.shell;
/**
* Default in-container pane command per CLI mode (mirror of defaultRemoteCommandForMode).
*
* ⚠️ Read from the registry (`overlays.docker`), not from a hardcoded
* `Record<DockerCommandMode, string>`. That table duplicated the registry exactly with
* nothing keeping the two in step. `shell` is the one arm still written here, because it is
* the entry that declares `docker: { disabled: true }` — a container has no per-user login
* shell to resolve, so it gets a plain `bash -l` rather than a CLI invocation.
*
* ⚠️ `runsAsRoot` selects the overlay's `rootCommand` when it declares one. Claude Code
* REFUSES `--dangerously-skip-permissions` under uid 0 ("cannot be used with root/sudo
* privileges", still true in 2.1.261), and the refusal is only visible INSIDE the
* container, so the pane just dies. Our own base image runs a non-root user and never hits
* it; an ADOPTED container's user belongs to its owner and is frequently root. Which flag
* to drop is a per-CLI fact, so it lives in the registry rather than in a branch here.
*/
export function defaultDockerCommandForMode(mode: SessionMode, runsAsRoot = false): string {
const entry = getCli(mode);
const overlay = entry?.overlays.docker;
if (!entry || (overlay && 'disabled' in overlay)) return 'exec bash -l';
// Mirrors the LOCAL default for each CLI; claude's carries
// `--dangerously-skip-permissions` so the in-container agent runs non-interactively.
const cli = (runsAsRoot ? overlay?.rootCommand : undefined) ?? overlay?.command ?? entry.discovery.binaries[0];
return cli ? `exec ${cli}` : 'exec bash -l';
}
/** `container:/workdir` display string (mirror of remoteDisplayPath's `user@host:path`). */
@@ -249,7 +284,22 @@ export function toSessionDocker(host: DockerHost, dockerCase: DockerCase): Sessi
extraCreateArgs: host.extraCreateArgs,
extraExecArgs: host.extraExecArgs,
};
return { ...base, configHash: dockerConfigHash(base) };
// `owned` is deliberately applied AFTER the hash: dockerConfigHash() picks an
// explicit field list, so ownership can never shift an existing case's hash and
// mass-trip the drift gate.
const session: SessionDocker = { ...base, configHash: dockerConfigHash(base) };
if (dockerCase.owned === false) session.owned = false;
return session;
}
/**
* An ADOPTED container is one the user built and runs themselves. Codeman may
* only exec into it; it must never create, start, stop, restart or remove it.
* Every lifecycle branch routes through this one predicate so a new call site
* cannot silently opt out.
*/
export function isAdoptedContainer(docker: Pick<SessionDocker, 'owned'>): boolean {
return docker.owned === false;
}
// ========== Shell escaping ==========
@@ -274,6 +324,24 @@ export interface DockerMount {
readonly?: boolean;
}
/**
* Resolve a bind source into the Docker daemon's filesystem namespace.
*
* A bare-host Codeman process and its Docker daemon see the same HOME, so the
* source is returned unchanged. In Docker-outside-of-Docker deployments,
* `runtimeHome` is the path inside Codeman while `daemonHome` is the host path
* bind-mounted there. Sources beneath HOME must therefore be translated before
* they are sent through the Docker socket.
*/
export function resolveDockerDaemonMountSource(source: string, runtimeHome: string, daemonHome?: string): string {
const configuredDaemonHome = daemonHome?.trim();
if (!configuredDaemonHome) return source;
const relativeSource = relative(resolve(runtimeHome), resolve(source));
if (relativeSource.startsWith('..') || isAbsolute(relativeSource)) return source;
return resolve(configuredDaemonHome, relativeSource);
}
/**
* Resolved, IO-free context for buildDockerCreateArgs. The caller (tmux-manager)
* resolves the environment-dependent bits (host uid, existing cred mounts, the
@@ -297,6 +365,8 @@ export interface DockerCreateContext {
addHostGateway: boolean;
/** Engine host-gateway alias (host.docker.internal / host.containers.internal). */
gatewayAlias: string;
/** Omit --memory-swap when the host kernel cannot enforce swap limits. */
disableSwapLimit?: boolean;
}
/**
@@ -314,12 +384,15 @@ function mountSpec(m: DockerMount): string {
return `type=bind,src=${m.src},dst=${m.dst}${m.readonly ? ',readonly' : ''}`;
}
function resourceFlags(resources?: DockerResourceLimits): string[] {
function resourceFlags(resources?: DockerResourceLimits, disableSwapLimit = false): string[] {
if (!resources) return [];
const flags: string[] = [];
if (resources.memory) {
// memory-swap == memory disables swap, making --memory a REAL OOM cap.
flags.push('--memory', resources.memory, '--memory-swap', resources.memory);
flags.push('--memory', resources.memory);
// memory-swap == memory disables swap where the daemon supports swap
// accounting. Some kernels, including the deployed Unraid host, do not;
// requesting it there emits a warning and Docker ignores the value.
if (!disableSwapLimit) flags.push('--memory-swap', resources.memory);
}
if (resources.cpus) flags.push('--cpus', resources.cpus);
if (resources.pidsLimit) flags.push('--pids-limit', String(resources.pidsLimit));
@@ -384,7 +457,7 @@ export function buildDockerCreateArgs(ctx: DockerCreateContext): string[] {
if (addHostGateway) args.push('--add-host', `${gatewayAlias}:host-gateway`);
args.push(
...resourceFlags(docker.resources),
...resourceFlags(docker.resources, ctx.disableSwapLimit),
// GPU passthrough (needs the NVIDIA container toolkit on the host). No storage
// cap is set, so the container's writable layer + volumes grow elastically as
// data flows in (bounded only by host disk).
@@ -614,8 +687,45 @@ const CRED_STORES: CredStorePolicy[] = [
rel: '.pi/agent',
seedFiles: ['auth.json', 'settings.json', 'trust.json', 'models.json', 'models-store.json'],
},
// Grok (xAI) keeps auth + config in `~/.grok`, but that dir ALSO holds
// `sessions/`, `memory/`, `downloads/` (the ~100MB binary itself) and `bin/`,
// so seedWhole would copy all of it into every container start. Seed only what
// grok needs to authenticate and behave consistently. Same trade-off as pi:
// in-container grok sessions are invisible host-side, so `grok -c` inside a
// Docker case only sees that container's own history.
{
rel: '.grok',
seedFiles: ['auth.json', 'config.toml', 'pager.toml'],
},
// DeepSeek Harness keeps credentials in `~/.dsh/.env` (0600) and composition in
// `settings.yaml` / `cordis.patch.yml`. `profiles/` is deliberately NOT seeded:
// it is a pnpm workspace holding a full node_modules tree per profile, which is
// both enormous and host-arch-specific. An in-container dsh therefore needs its
// profile installed IN the image (see docker/agent.Dockerfile), and the seeded
// files only supply auth and model composition. Same host-invisibility trade-off
// as pi and grok: `~/.dsh/sessions` inside a container is that container's own.
{
rel: '.dsh',
seedFiles: ['.env', 'settings.yaml', 'cordis.patch.yml'],
},
{ rel: '.config/gcloud', seedWhole: true },
{ rel: '.config/opencode', seedWhole: true },
// OMP keeps its config in `~/.omp/agent` (config.yml/mcp.json/models.yml/
// settings.yml — small, no bigger than grok's config.toml/pager.toml), but
// that dir ALSO holds agent.db/history.db/models.db (SQLite caches) and
// terminal-sessions/blobs/cache (large, regenerable), so seed only the
// config files. UNLIKE pi/grok, `sessions/` is SHARED (RW), not
// host-invisible: Codeman reads `~/.omp/agent/sessions/**/*.jsonl`
// HOST-SIDE for history recovery and --resume pinning
// (omp-transcript.ts, omp-session-resolver.ts) — the same reason codex's
// `sessions/` is shared rather than seeded. Without this, an in-container
// OMP conversation would be invisible to Codeman's own history-scan/resume
// logic, silently breaking the kill-survival feature for Docker cases.
{
rel: '.omp/agent',
shareDirs: ['sessions'],
seedFiles: ['config.yml', 'mcp.json', 'models.yml', 'settings.yml'],
},
];
/**
@@ -690,9 +800,15 @@ export interface DockerDriftStatus {
* daemon down) means there is nothing to drift. No-op under VITEST.
*/
export async function checkDockerConfigDrift(
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost' | 'containerName' | 'configHash'>
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost' | 'containerName' | 'configHash' | 'owned'>
): Promise<DockerDriftStatus> {
if (IS_TEST_MODE) return { exists: false, running: false, drifted: false };
// An ADOPTED container carries no `codeman.confighash` label — it was never
// created from our config — so every comparison would report drift and the
// launch gate would demand a recreate we are not allowed to perform. Ownership
// of its configuration belongs to the user; report "no drift" and never offer
// to rebuild it.
if (isAdoptedContainer(docker)) return { exists: true, running: false, drifted: false };
const argv = dockerEngineArgv(docker);
try {
const { stdout } = await execFileAsync(
@@ -720,8 +836,15 @@ export async function checkDockerConfigDrift(
* case's lastClaudeSessionId. No-op under VITEST.
*/
export async function removeDockerContainer(
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost' | 'containerName'>
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost' | 'containerName' | 'owned'>
): Promise<void> {
// Fail CLOSED at the lowest layer: an adopted container is the user's, and no
// caller — recreate-on-drift, case delete, a future teardown — may remove it.
if (isAdoptedContainer(docker)) {
throw new Error(
`Refusing to remove adopted container "${docker.containerName}": Codeman does not own its lifecycle.`
);
}
if (IS_TEST_MODE) return;
const argv = dockerEngineArgv(docker);
await execFileAsync(argv[0], [...argv.slice(1), 'rm', '-f', docker.containerName], { timeout: 30_000 });
@@ -967,6 +1090,255 @@ export async function checkDockerTmuxAvailable(
}
}
/** Preflight facts about an ALREADY-RUNNING container the user wants to adopt. */
export interface AdoptedContainerProbe {
ok: boolean;
exists: boolean;
running: boolean;
/** The container's own image ref (informational — we never enforce ours on it). */
image?: string;
/** `command -v tmux` inside the container; required for durable sessions. */
tmuxPath?: string;
/** Modes whose CLI resolved inside the container (`command -v <mode>`). */
availableModes?: SessionMode[];
/** Whether the requested working directory exists INSIDE the container. */
workdirExists?: boolean;
/** Whether the container's exec user is root (uid 0). */
runsAsRoot?: boolean;
error?: string;
}
/** One container on the engine, as offered to the adoption picker. */
export interface DockerContainerInfo {
name: string;
image: string;
running: boolean;
/** Engine's own status string, e.g. "Up 3 hours" / "Exited (0) 2 days ago". */
status: string;
}
/**
* List the engine's containers for the adoption picker (mirror of
* `listRemoteCodemanSessions`). Read-only and NEVER throws: an unreachable
* daemon, a missing engine or zero containers all return `[]`, because this
* feeds a convenience picker whose input the user can always type by hand.
*
* Stopped containers ARE included, sorted after running ones and carrying their
* status: adoption requires a running container, but hiding a stopped one turns
* "my container is not in the list" into a dead end with no explanation, while
* showing `my-box (Exited (0) 2 days ago)` says exactly what to fix.
*/
export async function listDockerContainers(
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost'>
): Promise<DockerContainerInfo[]> {
if (IS_TEST_MODE) return [];
const argv = dockerEngineArgv(docker);
try {
const { stdout } = await execFileAsync(
argv[0],
[...argv.slice(1), 'ps', '-a', '--format', '{{.Names}}\t{{.Image}}\t{{.State}}\t{{.Status}}'],
{ timeout: DOCKER_PROBE_TIMEOUT_MS }
);
const rows = stdout
.split('\n')
.map((line) => line.split('\t'))
.filter((parts) => parts.length >= 4 && parts[0])
.map(([name, image, state, status]) => ({
name,
image: image || '',
running: state === 'running',
status: status || '',
}));
// Running first, then by name, so the containers a user can actually adopt
// are the ones at the top of the list.
return rows.sort((a, b) => Number(b.running) - Number(a.running) || a.name.localeCompare(b.name));
} catch {
return [];
}
}
/**
* Preflight an EXISTING container for adoption. Read-only by construction: it
* runs `inspect` plus one `exec` of `command -v`, and never creates, starts or
* modifies anything. Refusing here is what keeps the failure at link time — a
* clear message — instead of at session launch, where the only alternatives
* would be a dead pane or starting a container we do not own.
*
* `--pull=never` is irrelevant here: adoption never touches images. The image
* ref is reported only so the UI can show what the user is attaching to.
*/
export async function probeAdoptableContainer(
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost' | 'containerName'>,
modes: SessionMode[] = [],
containerWorkdir?: string
): Promise<AdoptedContainerProbe> {
if (IS_TEST_MODE) {
return {
ok: true,
exists: true,
running: true,
tmuxPath: '/usr/bin/tmux',
availableModes: modes,
workdirExists: true,
};
}
const argv = dockerEngineArgv(docker);
let running = false;
let image: string | undefined;
try {
const { stdout } = await execFileAsync(
argv[0],
[...argv.slice(1), 'inspect', '-f', '{{.State.Running}}\t{{.Config.Image}}', docker.containerName],
{ timeout: DOCKER_PROBE_TIMEOUT_MS }
);
const [state = '', img = ''] = stdout.trim().split('\t');
running = state === 'true';
image = img || undefined;
} catch {
return {
ok: false,
exists: false,
running: false,
error: `container "${docker.containerName}" not found (adoption never creates a container — start it yourself first)`,
};
}
if (!running) {
return {
ok: false,
exists: true,
running: false,
image,
error: `container "${docker.containerName}" exists but is not running (Codeman never starts a container it does not own — start it yourself, then retry)`,
};
}
// One exec resolves tmux plus every requested CLI, so adoption costs a single
// round trip. Binaries are fixed mode names, never user input.
// A mode with no binary of its own (`shell`) is dropped: there is nothing to look up,
// and `command -v ''` would make the whole probe meaningless.
const wanted = modes.filter((m) => !!containerBinaryFor(m));
const probes = ['tmux', ...wanted.map((m) => containerBinaryFor(m) as string)];
// `; exit 0` is load-bearing: the script's status is its LAST command's, so a
// missing final CLI made the whole `sh -lc` exit 1 and the probe reported
// "could not exec into the container" for a container that was perfectly fine.
// Absence of a CLI is data here, not failure — only a real exec error is.
const steps = probes.map((bin) => `command -v ${bin} >/dev/null 2>&1 && echo ${bin}`);
// The workdir is checked INSIDE the container, and that is a fact independent
// of hostWorkspacePath: an owned container gets the host dir bind-mounted at the
// same absolute path at create time, but adoption mounts nothing, so the two
// paths only coincide if the user mounted it there themselves. `docker exec
// --workdir <missing>` fails with an OCI chdir error the pane surfaces as a bare
// "execvp failed", so it is resolved here into an actionable message.
if (containerWorkdir) steps.push(`[ -d ${shellescape(containerWorkdir)} ] && echo __workdir__`);
// Claude Code REFUSES --dangerously-skip-permissions as root. Our own base
// image runs a non-root user so an owned container never hits it; an adopted
// container's user belongs to its owner and is frequently root.
steps.push(`[ "$(id -u)" = 0 ] && echo __root__`);
const script = `${steps.join('; ')}; exit 0`;
try {
const { stdout } = await execFileAsync(
argv[0],
[...argv.slice(1), 'exec', docker.containerName, 'sh', '-lc', script],
{ timeout: DOCKER_PROBE_TIMEOUT_MS }
);
const found = new Set(
stdout
.split('\n')
.map((line) => line.trim())
.filter(Boolean)
);
if (!found.has('tmux')) {
return {
ok: false,
exists: true,
running: true,
image,
error: `container "${docker.containerName}" has no tmux (required for durable sessions; install it inside the container)`,
};
}
const workdirExists = containerWorkdir ? found.has('__workdir__') : undefined;
if (containerWorkdir && !workdirExists) {
return {
ok: false,
exists: true,
running: true,
image,
workdirExists: false,
error: `"${containerWorkdir}" does not exist inside container "${docker.containerName}". Adoption mounts nothing, so the container workdir must already exist there — set it to a path inside the container (it need not match the host workspace path).`,
};
}
return {
ok: true,
exists: true,
running: true,
image,
tmuxPath: 'tmux',
availableModes: modes.filter((m) => {
const bin = containerBinaryFor(m);
return bin ? found.has(bin) : true; // `shell` needs no binary
}),
workdirExists,
runsAsRoot: found.has('__root__'),
};
} catch (err) {
const msg = err instanceof Error ? err.message : String(err);
return { ok: false, exists: true, running: true, image, error: `could not exec into the container: ${msg}` };
}
}
/** One directory listing from INSIDE a container, shaped like the host picker's. */
export interface DockerBrowseResult {
path: string;
parent: string | null;
entries: Array<{ name: string; path: string; type: 'directory' | 'file' }>;
error?: string;
}
/**
* List a directory INSIDE a container, for the adoption form's container-workdir
* picker. The host filesystem picker cannot serve this: the path lives in the
* container, and for an adopted container nothing is mounted at a matching host
* location, so the user would otherwise be typing a path blind.
*
* Read-only: one `ls` through `docker exec`, no writes, no lifecycle. The path
* is shell-escaped like every other value this module interpolates, and output
* is parsed as NUL-free lines with a leading type marker so a filename with
* spaces survives.
*/
export async function browseInContainer(
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost' | 'containerName'>,
path: string
): Promise<DockerBrowseResult> {
const target = path && path.startsWith('/') ? path : '/';
const parent = target === '/' ? null : target.replace(/\/+$/, '').split('/').slice(0, -1).join('/') || '/';
if (IS_TEST_MODE) return { path: target, parent, entries: [] };
const argv = dockerEngineArgv(docker);
// `-p` marks directories with a trailing slash; `-A` shows dotfiles but not
// the . and .. entries the picker navigates with its own Up control.
const script = `cd ${shellescape(target)} 2>/dev/null && ls -Ap 2>/dev/null || echo __ERR__`;
try {
const { stdout } = await execFileAsync(
argv[0],
[...argv.slice(1), 'exec', docker.containerName, 'sh', '-lc', script],
{ timeout: DOCKER_PROBE_TIMEOUT_MS, maxBuffer: 4 * 1024 * 1024 }
);
if (stdout.includes('__ERR__')) return { path: target, parent, entries: [], error: 'Not a readable directory' };
const base = target.endsWith('/') ? target : `${target}/`;
const entries = stdout
.split('\n')
.map((line) => line.trim())
.filter(Boolean)
.map((name) => {
const isDir = name.endsWith('/');
const clean = isDir ? name.slice(0, -1) : name;
return { name: clean, path: `${base}${clean}`, type: (isDir ? 'directory' : 'file') as 'directory' | 'file' };
})
.sort((a, b) => Number(b.type === 'directory') - Number(a.type === 'directory') || a.name.localeCompare(b.name));
return { path: target, parent, entries };
} catch (err) {
return { path: target, parent, entries: [], error: err instanceof Error ? err.message : String(err) };
}
}
/**
* Resolve the host's IP on the default docker bridge (the address a container
* reaches as `host.docker.internal`), so the server can bind a hooks-only listener
@@ -1029,9 +1401,18 @@ export async function reapOrphanedDockerContainers(
}
const cases = await readDockerCases(configDir);
const expected = new Set(cases.map((c) => c.container ?? dockerContainerName(c.name)));
// ADOPTED containers are never reapable, and this guard is deliberately
// independent of the two conditions that already cover them (we never applied
// the `codeman.managed=1` label filtered on above, and they are referenced by a
// live case so they are in `expected`). An adopted container is the user's
// property; it must survive even if a future edit narrows either condition.
const adopted = new Set(
cases.filter((item) => item.owned === false).map((item) => item.container ?? dockerContainerName(item.name))
);
const reaped: string[] = [];
for (const { name, inst } of rows) {
if (inst !== instance) continue; // only THIS instance's containers
if (adopted.has(name)) continue; // never reap a container we do not own
if (expected.has(name)) continue; // still referenced by a live case
try {
await execFileAsync(bin, ['rm', '-f', name], { timeout: DOCKER_PROBE_TIMEOUT_MS });
@@ -1054,8 +1435,13 @@ export async function probeDockerCliVersion(
mode: SessionMode
): Promise<string | undefined> {
if (IS_TEST_MODE) return undefined;
const bin = mode === 'shell' ? null : mode;
if (!bin) return undefined;
// ⚠️ The MODE NAME IS NOT ALWAYS THE BINARY NAME — `antigravity` runs `agy`. This used
// to pass the mode straight through as the command, which would have probed a binary that
// does not exist. Only claude reaches this today (it is the one CLI with a version gate),
// so nothing was actually broken, but the registry is what makes it correct for the next
// CLI that needs a version.
const bin = getCli(mode)?.discovery.binaries[0];
if (!bin) return undefined; // `shell` has no binary of its own
const argv = dockerEngineArgv(docker);
try {
const { stdout } = await execFileAsync(
+114 -6
View File
@@ -366,19 +366,33 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
// never lands in this config and rotation needs no respawn. If the var/file is
// missing the header is empty — the middleware then allows the request only on
// the plain loopback bypass (tunnel down), same as pre-secret behavior.
const curlCmd = (event: HookEventType) =>
const curlCmd = (event: HookEventType, options: { discardStdout?: boolean } = {}) =>
`HOOK_DATA=$(cat 2>/dev/null || echo '{}'); ` +
`printf '{"event":"${event}","sessionId":"%s","data":%s}' "$CODEMAN_SESSION_ID" "$HOOK_DATA" | ` +
// `-k`, same as the statusline exporter: CODEMAN_API_URL is loopback HTTPS with
// a self-signed cert on --https/tailscale installs. Without it curl exits 60,
// the `|| true` swallows it, and ALL SIX hook events die silently: respawn loses
// its definitive idle signals and the wait endpoints lose stop/blocked.
`curl -sk -X POST "$CODEMAN_API_URL/api/hook-event" ` +
`curl -sk ${options.discardStdout ? '-o /dev/null ' : ''}-X POST "$CODEMAN_API_URL/api/hook-event" ` +
`-H 'Content-Type: application/json' ` +
`-H "X-Codeman-Hook-Secret: $(cat "$CODEMAN_HOOK_SECRET_FILE" 2>/dev/null)" ` +
`--data @- ` +
`2>/dev/null || true`;
// The same POST with stdout DISCARDED, via curl's own `-o`. UserPromptSubmit is
// one of the hook events whose stdout Claude Code injects into the model's
// context (the CLI's own hook reference: "Exit code 0 - stdout shown to
// Claude"), so an undiscarded curl pastes Codeman's `{"success":true,…}`
// envelope into the user's prompt on every single turn.
// ⚠️ It MUST be curl's flag, not a trailing redirect. `curlCmd` already ends
// `… 2>/dev/null || true`, and in `pipeline || true >/dev/null` the shell binds
// the redirection to `true` — which never runs on the success path — so the
// envelope still reaches stdout. Verified in dash and bash.
// ⚠️ The flag is opt-in so the other events' command text stays byte-identical:
// their stdout feeds the SSE stream harmlessly, and changing it would rewrite
// every workspace's settings file for no gain.
const curlCmdSilent = (event: HookEventType) => curlCmd(event, { discardStdout: true });
return {
hooks: {
Notification: [
@@ -410,6 +424,16 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
hooks: [{ type: 'command', command: curlCmd('stop'), timeout: HOOK_TIMEOUT_SECONDS }],
},
],
// The pane's LIVE conversation id, reported by the CLI process itself.
// Without it the response viewer has to guess which `<uuid>.jsonl` a pane
// is on after a `/clear`, and the only anchor it can guess from is an
// Enter that went THROUGH Codeman — so a user who attaches to tmux
// directly never gets one and stays pinned to the launch conversation.
UserPromptSubmit: [
{
hooks: [{ type: 'command', command: curlCmdSilent('prompt_submitted'), timeout: HOOK_TIMEOUT_SECONDS }],
},
],
SubagentStop: [
{
hooks: [
@@ -735,9 +759,22 @@ export async function refreshStaleCodemanHooks(casePath: string): Promise<void>
// Approvals Inbox needs the elicitation_complete/elicitation_response
// matchers; their absence marks a pre-inbox hooks block.
const hasElicitationComplete = hooksJson.includes('elicitation_complete');
// The UserPromptSubmit event is what gives a tmux-driven pane a first-hand
// conversation id; its absence marks a pre-prompt_submitted hooks block.
// ⚠️ No surrounding quotes: `hooksJson` is JSON.stringify'd, so the marker
// inside the command reads \"prompt_submitted\" and a quoted needle never
// matches — which would make this gate permanently false and rewrite every
// workspace's settings file on every Claude spawn. The sibling markers are
// quote-free for the same reason.
const hasPromptSubmit = hooksJson.includes('prompt_submitted');
if (
!isOurs ||
(hasSecret && hasBackgroundWake && hasSubagentStopGuard && hasElicitationComplete && !hasTlsFlaglessCurl)
(hasSecret &&
hasBackgroundWake &&
hasSubagentStopGuard &&
hasElicitationComplete &&
hasPromptSubmit &&
!hasTlsFlaglessCurl)
)
return;
const generated = generateHooksConfig();
@@ -1029,9 +1066,80 @@ export async function installAgentSkillInto(skillDir: string): Promise<AgentSkil
*/
export async function seedAgentSessionPreamble(sessionId: string): Promise<void> {
const content = await readFile(join(agentSkillSourceDir(), 'preamble.sh'), 'utf-8');
const cacheDir = process.env.XDG_CACHE_HOME || join(homedir(), '.cache');
await mkdir(cacheDir, { recursive: true });
await writeFile(join(cacheDir, `codeman-agent-${sessionId}.sh`), content, { mode: 0o600 });
await mkdir(agentPreambleCacheDir(), { recursive: true });
await writeFile(agentPreamblePath(sessionId), content, { mode: 0o600 });
}
/** Where the preamble caches live. One formula, shared by seed / remove / prune. */
function agentPreambleCacheDir(): string {
return process.env.XDG_CACHE_HOME || join(homedir(), '.cache');
}
/** `codeman-agent-<sessionId>.sh` in that directory. */
function agentPreamblePath(sessionId: string): string {
return join(agentPreambleCacheDir(), `codeman-agent-${sessionId}.sh`);
}
/** Matches exactly what seedAgentSessionPreamble writes, and nothing else in ~/.cache. */
const AGENT_PREAMBLE_FILE_PATTERN = /^codeman-agent-(.+)\.sh$/;
/** How long a preamble cache with no live session behind it is kept before the sweep takes it. */
export const AGENT_PREAMBLE_MAX_AGE_MS = 7 * 24 * 60 * 60 * 1000;
/**
* Drop one session's preamble cache. Called when a session is deleted, which is the
* precise counterpart to seeding it at create: one file per claude session was being
* written and nothing ever removed them (236 leftovers measured on a working machine,
* the oldest three weeks old). Best-effort — a file that will not delete is litter,
* never a reason to fail a teardown.
*/
export async function removeAgentSessionPreamble(sessionId: string): Promise<void> {
await unlink(agentPreamblePath(sessionId)).catch(() => {});
}
/**
* Sweep preamble caches left by sessions that are gone: the delete path above covers
* an orderly teardown, and this covers everything else (a crash, a killed server, a
* session deleted by an older build, another instance's leftovers).
*
* ⚠️ Two guards, and both matter: a file whose session is in `keepSessionIds` is never
* touched however old it is, and everything else needs `maxAgeMs` of age on top. A live
* session's cache is load-bearing — remove it and the skill's two-line loader fails its
* version check mid-run — and the age floor is what keeps a session belonging to
* ANOTHER instance (whose ids this process cannot see) out of the blast radius. Losing
* one is degradation rather than breakage: the §0 fallback block rewrites it.
*
* Returns how many it removed. Best-effort throughout; a missing cache dir is 0.
*/
export async function pruneAgentSessionPreambles(
keepSessionIds: Iterable<string>,
maxAgeMs: number = AGENT_PREAMBLE_MAX_AGE_MS
): Promise<number> {
const cacheDir = agentPreambleCacheDir();
const keep = new Set(keepSessionIds);
const cutoff = Date.now() - maxAgeMs;
let removed = 0;
let entries: string[];
try {
entries = await readdir(cacheDir);
} catch {
return 0;
}
for (const entry of entries) {
const sessionId = AGENT_PREAMBLE_FILE_PATTERN.exec(entry)?.[1];
if (!sessionId || keep.has(sessionId)) continue;
const path = join(cacheDir, entry);
try {
if ((await lstat(path)).mtimeMs > cutoff) continue;
await unlink(path);
removed++;
} catch {
/* best-effort — a vanished or unreadable file is not our problem */
}
}
return removed;
}
/**
+18 -4
View File
@@ -19,6 +19,9 @@ import type {
GeminiConfig,
AntigravityConfig,
PiConfig,
GrokConfig,
DeepSeekConfig,
OmpConfig,
SessionRemote,
SessionDocker,
} from './types.js';
@@ -78,13 +81,16 @@ export interface CreateSessionOptions {
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
piConfig?: PiConfig;
grokConfig?: GrokConfig;
deepSeekConfig?: DeepSeekConfig;
ompConfig?: OmpConfig;
/** When restoring after reboot, resume a previous Claude conversation by its session ID */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (e.g., CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS). Ephemeral — not written to disk. */
envOverrides?: Record<string, string>;
/** Claude CLI effort level, injected as a `--settings` soft default (overridable via /effort in-session) */
effort?: EffortLevel;
/** tmux history-limit (scrollback lines) to set for this session. */
/** tmux history-limit (scrollback lines) allocated when this session is created. */
historyLimit?: number;
/** Remote execution metadata for local tmux sessions wrapping SSH */
remote?: SessionRemote;
@@ -110,13 +116,16 @@ export interface RespawnPaneOptions {
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
piConfig?: PiConfig;
grokConfig?: GrokConfig;
deepSeekConfig?: DeepSeekConfig;
ompConfig?: OmpConfig;
/** Resume a previous Claude conversation when respawning */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (preserved across respawns). */
envOverrides?: Record<string, string>;
/** Claude CLI effort level (preserved across respawns, injected via `--settings`) */
effort?: EffortLevel;
/** tmux history-limit (scrollback lines) to set for this session after respawn. */
/** Original tmux history-limit retained for config parity; respawn cannot resize the existing pane. */
historyLimit?: number;
/** Remote execution metadata for local tmux sessions wrapping SSH */
remote?: SessionRemote;
@@ -128,7 +137,12 @@ export interface RespawnPaneOptions {
/** Options for pane buffer capture (COD-47 full-history mode). */
export interface PaneCaptureOptions {
/** Capture the entire tmux scrollback instead of just the visible frame. */
/**
* Capture the entire scrollback instead of just the visible frame, as linear
* text ending with a cursor move back to the pane's caret position. An
* implementation returns '' when the pane holds nothing visible, which the
* caller reads as "nothing to replay" and keeps its existing history.
*/
fullHistory?: boolean;
/** Bound the full-history capture to this many scrollback lines (`-S -<N>`). */
historyLimitLines?: number;
@@ -216,7 +230,7 @@ export interface TerminalMultiplexer extends EventEmitter {
/** Update Ralph enabled state for a session */
updateRalphEnabled(sessionId: string, enabled: boolean): void;
/** Apply a tmux history-limit to all tracked sessions. */
/** Apply history-limit to live panes where tmux supports it, otherwise to future panes. */
setHistoryLimit(limit: number): Promise<void>;
// ========== Discovery ==========
+175
View File
@@ -0,0 +1,175 @@
/**
* @fileoverview Scan `~/.omp/agent/sessions/*&#47;*.jsonl` for Past Sessions rows,
* the omp analog of what `scanProjectDir()` (session-routes.ts) does for
* Claude's own `~/.claude/projects` transcripts.
*
* Without this, an omp conversation exists ONLY as a Codeman-level live/
* persisted session record — delete that (a "Kill Tmux" close, or any other
* cleanup) and the conversation vanishes from Past Sessions entirely, even
* though `omp` itself never forgot it. Claude conversations don't have that
* problem because Codeman already reads them back from Claude's own
* transcript files independent of its own session bookkeeping; this gives
* omp conversations the same treatment.
*
* Each omp session file's SECOND line is a `{"type":"session","id":...,
* "cwd":...}` header carrying the real (unmangled) working directory and the
* session's own id directly — no need to reverse-engineer the mangled
* directory name the way Claude Code's own scanner has to (see
* `decodeProjectKey()` in session-routes.ts and its "lossy" caveat). Prompt
* text comes from each `{"type":"message","message":{"role":"user",...}}`
* entry, giving a real first-message title instead of a bare case name.
*
* Unlike Claude's transcripts (which can run to tens of MB of tool-call
* output), an omp session file is the conversation only, so this reads each
* file whole rather than doing head/tail windows — bounded by a size cap so
* one unexpectedly huge file can't blow up memory.
*
* @module omp-transcript
*/
import { readFileSync, readdirSync, statSync } from 'node:fs';
import { homedir } from 'node:os';
import { join } from 'node:path';
function ompSessionsRoot(): string {
return join(homedir(), '.omp', 'agent', 'sessions');
}
/** Skip anything absurdly large rather than parsing it whole into memory. */
const MAX_OMP_SESSION_FILE_BYTES = 2 * 1024 * 1024;
/** Defensive cap on total files scanned across every directory, mirroring
* the Claude scanner's own instinct not to let one pathological tree stall
* a request — a real omp install has, at most, a few hundred of these. */
const MAX_OMP_SESSION_FILES = 2000;
export interface OmpHistorySession {
sessionId: string;
workingDir: string;
sizeBytes: number;
/** ISO timestamp, from the file's own mtime. */
lastModified: string;
firstPrompt?: string;
lastPrompt?: string;
}
function extractUserPromptText(message: unknown): string | undefined {
if (!message || typeof message !== 'object') return undefined;
const m = message as { role?: unknown; content?: unknown };
if (m.role !== 'user' || !Array.isArray(m.content)) return undefined;
const parts: string[] = [];
for (const block of m.content) {
if (block && typeof block === 'object' && (block as { type?: unknown }).type === 'text') {
const text = (block as { text?: unknown }).text;
if (typeof text === 'string') parts.push(text);
}
}
const joined = parts.join(' ').trim();
return joined || undefined;
}
/** Parse one omp session `.jsonl` file, or null when it's unreadable, empty, or has no session header. */
function parseOmpSessionFile(filePath: string): OmpHistorySession | null {
let stat: ReturnType<typeof statSync>;
try {
stat = statSync(filePath);
} catch {
return null;
}
if (stat.size === 0 || stat.size > MAX_OMP_SESSION_FILE_BYTES) return null;
let raw: string;
try {
raw = readFileSync(filePath, 'utf-8');
} catch {
return null;
}
let sessionId: string | undefined;
let workingDir: string | undefined;
let firstPrompt: string | undefined;
let lastPrompt: string | undefined;
for (const line of raw.split('\n')) {
if (!line) continue;
let entry: unknown;
try {
entry = JSON.parse(line);
} catch {
continue;
}
if (!entry || typeof entry !== 'object') continue;
const e = entry as Record<string, unknown>;
if (e.type === 'session' && typeof e.id === 'string' && typeof e.cwd === 'string' && e.cwd.startsWith('/')) {
// A corrupted or malformed session file could carry a relative or empty
// cwd; requiring an absolute path keeps a downstream resume attempt
// from being pointed at a nonsense working directory.
sessionId = e.id;
workingDir = e.cwd;
} else if (e.type === 'message') {
const prompt = extractUserPromptText(e.message);
if (prompt) {
if (!firstPrompt) firstPrompt = prompt;
lastPrompt = prompt;
}
}
}
if (!sessionId || !workingDir) return null;
return {
sessionId,
workingDir,
sizeBytes: stat.size,
lastModified: stat.mtime.toISOString(),
firstPrompt,
lastPrompt,
};
}
/**
* Scan every omp conversation on disk into Past-Sessions rows. Best-effort
* throughout: a missing `~/.omp` (never installed/used), an unreadable
* directory, or one corrupt file yields fewer rows rather than throwing —
* this feeds the same unified merge the Claude transcript scanner does, and
* one broken source must never blank the whole Past Sessions list.
*/
export function scanOmpSessionsHistory(): OmpHistorySession[] {
const root = ompSessionsRoot();
let dirEntries: string[];
try {
dirEntries = readdirSync(root);
} catch {
return [];
}
const out: OmpHistorySession[] = [];
for (const dirName of dirEntries) {
if (out.length >= MAX_OMP_SESSION_FILES) break;
const dirPath = join(root, dirName);
let dirStat: ReturnType<typeof statSync>;
try {
dirStat = statSync(dirPath);
} catch {
continue;
}
if (!dirStat.isDirectory()) continue;
let files: string[];
try {
files = readdirSync(dirPath);
} catch {
continue;
}
for (const file of files) {
if (out.length >= MAX_OMP_SESSION_FILES) break;
if (!file.endsWith('.jsonl')) continue;
try {
const parsed = parseOmpSessionFile(join(dirPath, file));
if (parsed) out.push(parsed);
} catch {
// One bad file must not sink the whole scan.
}
}
}
return out;
}
+142 -39
View File
@@ -4,9 +4,9 @@ import { join } from 'node:path';
import { homedir } from 'node:os';
import { exec } from 'node:child_process';
import { promisify } from 'node:util';
import { getCli } from './config/cli-registry/registry.js';
import type {
RemoteCase,
RemoteCommandMode,
RemoteHost,
RemoteSessionInfo,
RemoteSshOptions,
@@ -89,33 +89,54 @@ export function remoteLoginShellCommand(command: string): string {
return `exec ${REMOTE_LOGIN_SHELL} -i -l -c ${shellescape(command)}`;
}
/**
* The CLI text a location overlay should launch for `mode`, or null when this build has no
* entry for it. `overlays.<location>.command` when the entry names one, otherwise the bare
* binary — which is what every non-claude CLI wants, and why only claude declares a command.
*
* ⚠️ This returns the CLI INVOCATION only. Each location wraps it its own way (remote: a
* login-shell `-c`; docker: `exec`), which is exactly why the overlay stores the unwrapped
* form rather than a ready-made line.
*/
function overlayCliCommand(mode: SessionMode, location: 'remote' | 'docker'): string | null {
const entry = getCli(mode);
if (!entry) return null;
const overlay = entry.overlays[location];
if (overlay && 'disabled' in overlay) return null;
return overlay?.command ?? entry.discovery.binaries[0] ?? null;
}
/**
* The default remote pane command for `mode`.
*
* Agent CLIs (claude/opencode/codex/gemini/antigravity/…) are typically installed under
* per-user paths like ~/.local/bin or ~/.opencode/bin, added to PATH only by the remote
* user's interactive-login shell startup files (~/.zshrc etc.). ssh's remote-command
* execution is neither interactive nor login, so a bare `exec claude` sees only sshd's
* minimal default PATH and fails with "command not found" (exit 127) — confirmed via
* `tmux capture-pane` on the remain-on-exit-preserved dead pane. Route through
* `$SHELL -i -l -c`, the same fix shell mode uses, so PATH is fully resolved first.
*
* ⚠️ The per-CLI half is now READ FROM THE REGISTRY (`overlays.remote`), not from a
* hardcoded `Record<RemoteCommandMode, string>`. The table it replaces duplicated the
* registry exactly, with nothing keeping the two in step — a capability that is both wrong
* and unread is worse than an absent one, because the next person trusts it. Notes that were
* attached to individual rows and are still true:
* - claude carries `--dangerously-skip-permissions` so the remote agent runs
* non-interactively (no trust-folder prompt nothing on the remote can answer);
* `overlays.remote.command` on the claude entry is where that now lives.
* - `dsh` alone boots nothing — the launcher needs a profile, and the remote box's profile
* inventory is unknown here. The per-host `commands.deepseek` override names one.
* The per-host `commands.*` override remains the escape hatch for every mode.
*/
export function defaultRemoteCommandForMode(mode: SessionMode): string {
// Agent CLIs (claude/opencode/codex/gemini/antigravity) are typically installed
// under per-user paths like ~/.local/bin or ~/.opencode/bin, added to PATH only by
// the remote user's interactive-login shell startup files (~/.zshrc etc.). ssh's
// remote-command execution is neither interactive nor login, so a bare `exec
// claude` sees only sshd's minimal default PATH and fails with "command not
// found" (exit 127) — confirmed via `tmux capture-pane` on the
// remain-on-exit-preserved dead pane. Route through `$SHELL -i -l -c`, the same
// fix already used for shell mode below, so PATH is fully resolved before the
// CLI name is looked up.
const commands: Record<RemoteCommandMode, string> = {
// $SHELL, not a hardcoded bash: sshd sets it from the remote user's
// /etc/passwd entry, so this launches their actual login shell (zsh,
// fish, etc.). -i -l so it sources rc files (~/.zshrc etc.), matching
// the local shell-mode launch.
shell: `exec ${REMOTE_LOGIN_SHELL} -i -l`,
// Mirror the LOCAL claude default so the remote agent runs non-interactively
// (no trust-folder/permission prompt that nothing on the remote answers). The
// per-host `commands.claude` override stays the escape hatch.
claude: remoteLoginShellCommand('claude --dangerously-skip-permissions'),
opencode: remoteLoginShellCommand('opencode'),
codex: remoteLoginShellCommand('codex'),
gemini: remoteLoginShellCommand('gemini'),
antigravity: remoteLoginShellCommand('agy'),
pi: remoteLoginShellCommand('pi'),
};
return commands[mode as RemoteCommandMode] || commands.shell;
// $SHELL, not a hardcoded bash: sshd sets it from the remote user's /etc/passwd entry, so
// this launches their actual login shell (zsh, fish, …). `-i -l` so it sources rc files,
// matching the local shell-mode launch. Not templatable as overlay data: the shell is
// whatever the REMOTE passwd says, which is why `shell` is the one arm still written here.
const shellCommand = `exec ${REMOTE_LOGIN_SHELL} -i -l`;
const cli = overlayCliCommand(mode, 'remote');
return cli === null ? shellCommand : remoteLoginShellCommand(cli);
}
export function remoteSshTarget(host: Pick<RemoteHost, 'username' | 'host'>): string {
@@ -258,18 +279,22 @@ export async function checkRemoteTmuxAvailable(
}
/**
* The CLI binary each session mode runs on the remote host. Antigravity's
* binary is `agy` (the mode name is not the command); shell has no CLI to
* probe, so it is absent.
* The CLI binary a session mode runs on the remote host, read from the registry rather than
* from a hardcoded map. `shell` has no CLI to probe and resolves to undefined, which is what
* makes the probe return null for it.
*
* ⚠️ Deriving this CHANGES BEHAVIOUR, deliberately and in one direction. The map it replaces
* listed claude/opencode/codex/gemini/antigravity/pi/omp and simply omitted `grok` and
* `deepseek` — its own comment said the rule was "every mode except shell", so the two were
* an oversight from when those CLIs were added, not a decision. A remote grok or deepseek
* session therefore reported no version at all. It now probes `grok --version` /
* `dsh --version` through the same login-shell wrapper as its siblings.
*
* (`antigravity` is why this cannot be the mode name: its binary is `agy`.)
*/
const REMOTE_CLI_BIN: Partial<Record<SessionMode, string>> = {
claude: 'claude',
opencode: 'opencode',
codex: 'codex',
gemini: 'gemini',
antigravity: 'agy',
pi: 'pi',
};
function remoteCliBin(mode: SessionMode): string | undefined {
return getCli(mode)?.discovery.binaries[0];
}
/**
* Build the SSH command that reads the remote CLI's version (`claude --version`
@@ -285,7 +310,7 @@ export function buildRemoteCliVersionProbeCommand(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
mode: SessionMode
): string | null {
const bin = REMOTE_CLI_BIN[mode];
const bin = remoteCliBin(mode);
if (!bin) return null;
return [
...buildSshConnectionArgs(host),
@@ -321,6 +346,84 @@ export async function probeRemoteCliVersion(
}
}
/**
* COD-108 — build the SSH command that asks whether THIS Codeman's durable
* remote tmux session (`-L codeman-remote -s codeman-ssh-<id>`) is still alive
* on the remote host.
*
* `has-session` exits 0 when the session exists, non-zero otherwise (and
* stderr is swallowed). Connection options come from the shared
* `buildSshConnectionArgs` so this probe reaches exactly the hosts the launch
* can reach — same port/identity/proxy/jump-host as `buildRemoteLaunchCommand`.
*/
export function buildRemoteSessionAliveCommand(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
remoteSessionName: string
): string {
const [ssh, ...connectionArgs] = buildSshConnectionArgs(host);
const remoteCmd = `tmux -L codeman-remote has-session -t ${shellescape(remoteSessionName)} 2>/dev/null`;
return [ssh, ...connectionArgs, remoteSshTarget(host), shellescape(remoteCmd)].join(' ');
}
/**
* COD-108 — resolve whether THIS Codeman's durable remote tmux session is still
* alive on the remote host, for the auto-reconnect watcher.
*
* Returns:
* - `true` → the remote tmux session exists (the agent is still running
* on the remote; the LOCAL pane died from a transport drop →
* safe to auto-reconnect).
* - `false` → the remote session is gone (the agent exited cleanly and
* the remote tmux tore down; reviving would relaunch a fresh
* agent — must NOT auto-reconnect).
* - `undefined` → probe failed (host unreachable, ssh error, tmux missing).
* Callers MUST treat this as "do not reconnect": an
* unreachable host is not a reason to relaunch the agent.
*
* VITEST guard — returns `true` under test so a real ssh never runs; the
* command construction is covered by `buildRemoteSessionAliveCommand`.
*/
export async function remoteTmuxSessionAlive(
remote: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
remoteSessionName: string
): Promise<boolean | undefined> {
if (process.env.VITEST) return true;
const command = buildRemoteSessionAliveCommand(remote, remoteSessionName);
try {
await execAsync(command, { timeout: 15_000 });
return classifyRemoteAliveExit(0, false);
} catch (err) {
const e = err as { code?: unknown; killed?: boolean };
return classifyRemoteAliveExit(typeof e.code === 'number' ? e.code : null, e.killed === true);
}
}
/**
* Map the `has-session` probe's exit status onto the tri-state the watcher
* reads. Pure, so the mapping is unit-tested even though the probe itself is
* VITEST-guarded.
*
* ⚠️ `tmux has-session` prints NOTHING on success (measured: exit 0, empty
* stdout; the failure message goes to stderr), so the exit status is the ONLY
* signal. An earlier version read stdout and therefore classified every live
* remote session as gone, which silently disabled transport-drop reconnects.
*
* - exit 0 → the durable remote session exists → `true`.
* - exit 255 is ssh's own failure (unreachable host, auth, proxy/jump error)
* and a timeout arrives as `killed` with no numeric code: we learned
* nothing about the session → `undefined`, which the watcher treats as
* "do not revive".
* - any other non-zero status is the REMOTE command's: tmux's 1 for a missing
* session, or 127 when tmux is not installed there (no durable session can
* exist without it) → `false`.
*/
export function classifyRemoteAliveExit(code: number | null, killed: boolean): boolean | undefined {
if (killed) return undefined;
if (code === 0) return true;
if (code === null || code === 255) return undefined;
return false;
}
/**
* COD-105 — build the SSH command that lists `codeman-*` tmux sessions on a
* remote host's canonical `-L codeman` socket.
+19 -1
View File
@@ -108,6 +108,15 @@ export interface ReconnectSessionView {
isRemote: boolean;
/** Result of `isPaneDead(muxName)` for this session. */
paneDead: boolean;
/**
* Whether the DURABLE remote tmux session is still alive on the remote host.
* Tri-state: `true` = transport drop with the agent still running (safe to
* reattach); `false` = the remote session is gone (the agent exited cleanly
* via ctrl-c/ctrl-d/exit and the remote tmux tore down); `undefined` =
* unknown/unresolvable. The watcher must NOT revive when the remote session
* is gone or unknown — a clean exit must never auto-relaunch the agent.
*/
remoteAlive: boolean | undefined;
}
/**
@@ -130,7 +139,8 @@ export type ReconnectSkipReason =
| 'in-flight'
| 'not-due'
| 'exhausted'
| 'disabled';
| 'disabled'
| 'remote-gone';
export interface DecideReconnectInput {
session: ReconnectSessionView;
@@ -166,6 +176,14 @@ export function decideReconnect(input: DecideReconnectInput): ReconnectAction {
if (!session.paneDead) return { kind: 'skip', reason: 'pane-alive' };
// Intentional kill / detach must NEVER be auto-revived.
if (guarded) return { kind: 'skip', reason: 'guarded' };
// A clean exit tears down the durable remote tmux (the session's only pane
// exiting destroys it). Reviving is ONLY correct for a transport drop: the
// agent is still running on the remote, so the durable session must still
// exist. When it is gone (or status is unknown — probe failed/unreachable),
// the agent exited intentionally and must not be auto-relaunched (found
// live 2026-08-29: remote omp/opencode ctrl-c/ctrl-d auto-respawned fresh
// sessions; only claude's `|| --resume` accidentally masked it).
if (session.remoteAlive !== true) return { kind: 'skip', reason: 'remote-gone' };
const s = state ?? freshReconnectState();
+36 -6
View File
@@ -2,12 +2,16 @@
* @fileoverview Pure merge/filter logic for the unified session list (COD-121).
*
* Combines four read-only views of a session — live (in-memory `Session`),
* persisted (`state.json`), transcript history (`~/.claude/projects`), and the
* lifecycle audit log — plus mux process stats, into one de-duplicated list
* keyed by sessionId. Transcript-history rows are keyed by the Claude
* conversation UUID (the `.jsonl` filename stem), which diverges from the
* Codeman id for resumed sessions — an alias map (claudeSessionId → Codeman id,
* built from the live/persisted views) folds them into the owning session item.
* persisted (`state.json`), transcript history, and the lifecycle audit log —
* plus mux process stats, into one de-duplicated list keyed by sessionId.
*
* Transcript history is not one source but three, because the CLIs keep their
* conversations in their own stores: Claude's `~/.claude/projects`, omp's
* `~/.omp/agent/sessions` and codex's `~/.codex/sessions`. Each row is keyed by
* whatever id that CLI names the conversation with, which diverges from the
* Codeman id for a resumed session and for every non-Claude one — an alias map
* (claudeSessionId → Codeman id, built from the live/persisted views) folds them
* into the owning session item.
* Higher-precedence sources overwrite scalar fields when present
* (history < lifecycle < persisted < live), while the `sources` array
* always accumulates every contributing view. A "meaningfulness floor" drops
@@ -39,6 +43,11 @@ export type UnifiedSessionItem = {
/** Main repo root a worktree belongs to (#266). */
worktreeRepo?: string;
remote?: boolean;
/**
* Token this row's CLI resumes by, when that is not `sessionId`. Set only from
* a transcript scanner — see the field of the same name on `HistoryInput`.
*/
resumeId?: string;
/** Pinned to the top of the session manager list (COD-139). */
pinned?: boolean;
/** When the session was pinned (epoch ms) — orders the pinned group desc. */
@@ -99,6 +108,22 @@ export type HistoryInput = {
gitBranch?: string;
worktreeName?: string;
worktreeRepo?: string;
/**
* Set only by a non-claude transcript source (currently omp and codex); the
* Claude scanner never stamps this; the meaningfulness floor below still
* counts a row with a `mode` as real, since that also signals "not claude" —
* see where it's read below for the isReal check this touches.
*/
mode?: string;
/**
* The token this CLI's own resume command expects, when it is NOT the row's
* `sessionId`. Codex names a thread by an id of its own that lives in the
* rollout, and a live codex session's `sessionId` is Codeman's uuid instead —
* so a resume that reused `sessionId` would ask codex for a thread that does
* not exist. Only a transcript scanner sets this, which is what keeps the two
* kinds of row apart.
*/
resumeId?: string;
};
/** Mux process-stat view. */
@@ -175,6 +200,11 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
overwrite(item, 'gitBranch', h.gitBranch);
overwrite(item, 'worktreeName', h.worktreeName);
overwrite(item, 'worktreeRepo', h.worktreeRepo);
// Claude rows never set this (they're implicitly claude); a non-claude
// transcript source (currently only omp) does, so a history-only row
// still gets a mode badge instead of reading as claude by default.
overwrite(item, 'mode', h.mode);
overwrite(item, 'resumeId', h.resumeId);
const ms = Date.parse(h.lastModified);
if (!Number.isNaN(ms) && item.lastActivityAt === undefined) item.lastActivityAt = ms;
}
+202
View File
@@ -0,0 +1,202 @@
/**
* @fileoverview Bridges the legacy per-mode spawn options (`buildSpawnCommand`'s option bag
* in tmux-manager.ts, unchanged on the wire since before this registry existed) onto the CLI
* registry's generic argv engine (`renderLaunch`).
*
* The per-mode `<Mode>Config` objects on `POST /api/sessions` predate the registry and stay
* on the wire for API compatibility (`docs/versioning-policy.md`), so SOMETHING has to know
* which field holds which CLI's config. That knowledge is DATA — `launch.legacyConfigField`
* and `launch.legacyConfigAliases`, declared once per entry in `config/cli-registry/stock.ts`
* — which is what lets this file stay a generic reader rather than a `switch (mode)`.
*
* An entry declaring NO `legacyConfigField` reads its params straight off the top-level
* option bag. That is claude, whose discrete `claudeMode`/`allowedTools`/`model`/
* `resumeSessionId` fields predate the `<Mode>Config` pattern — not a special case for
* claude, just the other of the two shapes the wire has always had.
*
* @module session-cli-registry-bridge
*/
import type { CliEntry } from './config/cli-registry/types.js';
import { renderLaunch, type EngineValues, type ParamValues } from './config/cli-registry/argv.js';
import { matchesPattern } from './config/cli-registry/patterns.js';
import { buildEffortCliArgs, sanitizeCliSessionName } from './session-cli-builder.js';
import { compareVersions } from './utils/dependency-checker.js';
import { getClaudeCliVersion } from './utils/claude-cli-resolver.js';
import { launcherDefaultTarget } from './utils/cli-launcher.js';
import { getCli } from './config/cli-registry/registry.js';
import type {
AntigravityConfig,
ClaudeMode,
CodexConfig,
DeepSeekConfig,
EffortLevel,
GeminiConfig,
GrokConfig,
OmpConfig,
OpenCodeConfig,
PiConfig,
} from './types/session.js';
export interface SpawnBridgeOptions {
mode: string;
sessionId: string;
model?: string;
claudeMode?: ClaudeMode;
allowedTools?: string;
openCodeConfig?: OpenCodeConfig;
codexConfig?: CodexConfig;
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
piConfig?: PiConfig;
grokConfig?: GrokConfig;
deepSeekConfig?: DeepSeekConfig;
ompConfig?: OmpConfig;
resumeSessionId?: string;
effort?: EffortLevel;
sessionName?: string;
claudeCliVersion?: string | null;
}
/**
* The raw legacy config object this entry's params should be read from: the declared
* `<Mode>Config` field, or the option bag itself when none is declared.
*/
function legacyConfigFor(entry: CliEntry, options: SpawnBridgeOptions): Record<string, unknown> | undefined {
const field = entry.launch.legacyConfigField;
if (field === undefined) return options as unknown as Record<string, unknown>;
return (options as unknown as Record<string, unknown>)[field] as Record<string, unknown> | undefined;
}
/**
* Same lookup, addressed by mode rather than by entry, for callers holding only a mode and an
* option bag (tmux-manager's env configuration). Returns undefined for an unregistered mode.
*/
export function legacyConfigForMode(
mode: string,
options: Record<string, unknown>
): Record<string, unknown> | undefined {
const entry = getCli(mode);
if (!entry) return undefined;
return legacyConfigFor(entry, options as unknown as SpawnBridgeOptions);
}
/**
* Build `ParamValues` for every declared `token`/`bool`/`enum` param by reading it out of the
* legacy config object through `legacyConfigAliases` (falling back to the param's own name).
* `engine`-sourced params are skipped — those come from `EngineValues`, never legacy config.
*/
function buildParamsFromLegacyConfig(entry: CliEntry, rawConfig: Record<string, unknown> | undefined): ParamValues {
const params: ParamValues = {};
if (!rawConfig) return params;
const aliases = entry.launch.legacyConfigAliases ?? {};
for (const [paramName, spec] of Object.entries(entry.launch.params)) {
if (spec.type === 'engine') continue;
const legacyKey = aliases[paramName] ?? paramName;
const value = rawConfig[legacyKey];
if (value === undefined) continue;
// Anything that is not already a string or boolean is DROPPED rather than coerced: the
// wire shape is Zod-validated upstream, so a surprise here means something is wrong,
// and `String({})` would happily produce a token nobody intended.
if (typeof value === 'string' || typeof value === 'boolean') {
params[paramName] = value;
}
}
return params;
}
/**
* The env vars this CLI declares in `env.configSetenv`, resolved from its legacy config
* object — i.e. the ones whose value comes from the CALLER rather than the server's own
* environment.
*
* ⚠️ Re-validated here against the declared `ParamSpec` even though the wire shape is already
* Zod-checked upstream. These values reach `tmux setenv`, and for DeepSeek the value IS a
* permission level: a builder must never trust its caller on a security-relevant field, and
* the cost of re-checking an enum is nothing.
*
* A value that fails validation is DROPPED, not defaulted — which is the safe direction: the
* var goes unset, and the CLI falls back to its own default (for dsh, `workspace-write`,
* which asks) rather than to something we guessed.
*/
export function configSetenvValues(
entry: CliEntry,
rawConfig: Record<string, unknown> | undefined
): Record<string, string> {
const out: Record<string, string> = {};
const mappings = entry.env.configSetenv;
if (!mappings || !rawConfig) return out;
const aliases = entry.launch.legacyConfigAliases ?? {};
for (const { name, fromParam } of mappings) {
const spec = entry.launch.params[fromParam];
if (!spec) continue; // schema-validated at load; belt and braces
const raw = rawConfig[aliases[fromParam] ?? fromParam];
if (typeof raw !== 'string') continue;
if (spec.type === 'enum' && !spec.values.includes(raw)) continue;
if (spec.type === 'token' && !matchesPattern(spec.pattern, raw)) continue;
out[name] = raw;
}
return out;
}
/**
* Which `capabilities.gates` are currently satisfied. `resolveVersion` is called AT MOST
* ONCE, and only when the entry actually declares a gate — a `--version` subprocess probe
* has no reason to run for an entry with none.
*/
function resolveGatesPassed(entry: CliEntry, resolveVersion: () => string | null): Set<string> {
const passed = new Set<string>();
const gateEntries = Object.entries(entry.capabilities.gates);
if (gateEntries.length === 0) return passed;
const cliVersion = resolveVersion();
if (!cliVersion) return passed; // fail-closed: an unknown version satisfies no gate
for (const [name, gate] of gateEntries) {
if (compareVersions(cliVersion, gate.minVersion) >= 0) passed.add(name);
}
return passed;
}
/**
* Render the spawn command for `entry` from the legacy option bag. Returns `undefined` for a
* `shell`-kind entry (or any entry declaring no launch variants), which callers take as "fall
* back to the local login-shell resolution" — shell has no CLI to template.
*/
export function buildSpawnCommandFromRegistry(entry: CliEntry, options: SpawnBridgeOptions): string | undefined {
if (entry.kind === 'shell' || entry.launch.variants.length === 0) return undefined;
const params = buildParamsFromLegacyConfig(entry, legacyConfigFor(entry, options));
const engineValues: EngineValues = {
sessionId: options.sessionId,
// Allowlist-sanitized (Unicode letters/digits + ` . _ : -`, 64 chars), matching
// buildNameCliArgs exactly — sanitizeCliSessionName is the injection guard for this
// value, NOT the `quote: 'double'` escaping on the --name arg (which only makes an
// unsafe value inert, it does not launder one into something meaningful).
sessionName: sanitizeCliSessionName(options.sessionName),
};
// Only a launcher CLI has one, and resolving it means a filesystem scan of the launcher's
// profile tree, so skip the lookup entirely for the eight entries that declare no profile.
if (entry.discovery.launcherProfile !== undefined) {
engineValues.launcherDefaultTarget = launcherDefaultTarget(entry) ?? undefined;
}
// Mirrors buildEffortCliArgs exactly: ultracode carries a fixed settings blob, every other
// level rides a plain `--effort <level>` flag. Reusing the canonical builder here (rather
// than re-deriving the ultracode special case) keeps the EFFORT_LEVELS allowlist and the
// settings-JSON shape single-sourced in session-cli-builder.ts.
const [effortFlag, effortValue] = buildEffortCliArgs(options.effort);
if (effortFlag === '--settings') engineValues.effortSettingsJson = effortValue;
else if (effortFlag === '--effort') engineValues.effortLevel = effortValue;
// Preserves buildSpawnCommand's original fallback exactly: an EXPLICIT `undefined` probes
// the local claude CLI (getClaudeCliVersion, null under vitest); an explicit `null` means
// "known to be unresolvable" and must not probe. The probe only ever runs from
// resolveGatesPassed, and only for an entry that actually declares a gate, so this stays
// generic without spawning a stray `claude --version` for every other CLI's launch.
const gatesPassed = resolveGatesPassed(entry, () =>
options.claudeCliVersion !== undefined ? options.claudeCliVersion : getClaudeCliVersion()
);
return renderLaunch(entry.launch, params, engineValues, gatesPassed);
}
+93 -9
View File
@@ -1,16 +1,32 @@
/**
* @fileoverview Recognizing Claude Code's workspace-trust dialog on screen.
* @fileoverview Recognizing Claude Code's workspace-trust dialog on screen, and
* working out which keystroke answers it.
*
* Claude asks once per directory before it will read or edit anything:
* Claude asks once per directory before it will read or edit anything. The
* layout has changed under us at least twice; both of these are live shapes:
*
* Quick safety check: Is this a project you created or one you trust? ...
* Quick safety check: Is this a project you created or one you trust? ... (<= 2.1.220)
* ❯ 1. Yes, I trust this folder
* 2. No, exit
* Enter to confirm · Esc to cancel
*
* Quick safety check: Is this a project you created or one you trust? ... (2.1.252)
* Security guide
* ❯ No, exit
* Yes, I trust this folder
* Enter to confirm · Esc to cancel
*
* Codeman sessions run permission-skipping or classifier-guarded modes, so the
* answer is always yes, and a session parked on this dialog is simply stuck.
*
* ⚠️ **Never press Enter without reading the selection.** The options are now
* unnumbered, REVERSED, and the highlighted default is "No, exit" — so the blind
* `\r` that answered the old layout picks *exit* on the new one and the pane
* dies (`Pane is dead (status 1)`) seconds after the session starts, which is
* exactly what a fresh case did on Claude Code 2.1.252. `trustDialogNextKey()`
* reads the `❯` marker instead and moves the cursor onto the trust option before
* it confirms anything.
*
* **Why the text has to be compacted.** tmux repaints a row by writing each word
* and then a cursor-forward (`\x1b[C`) instead of a space, and Ink colours each
* word separately, so the wire carries `I\x1b[Ctrust\x1b[Cthis\x1b[Cfolder`.
@@ -39,6 +55,25 @@ const TRUST_PHRASES = [
/** The dialog's own affordances. Prose that quotes the question will not have these. */
const CONFIRM_PHRASES = ['entertoconfirm', 'esctocancel', '2.no,exit'];
/** The option that answers yes, compacted. Identical text in both layouts. */
const YES_OPTION = 'yes,itrustthisfolder';
/** The option that quits Claude. It is the highlighted DEFAULT since 2.1.252. */
const NO_OPTION = 'no,exit';
/** Ink's selection marker. The only marked row while the dialog is up. */
const SELECTION_MARK = '❯';
/** A numbered option's `1.` / `2.` prefix, which the 2.1.220 layout put after the marker. */
const OPTION_NUMBER_PREFIX = /^\d+\./;
/** Move the selection one row down / up. Literal, so `send-keys -l` carries them. */
export const TRUST_KEY_DOWN = '\x1b[B';
export const TRUST_KEY_UP = '\x1b[A';
/** Confirm the highlighted option. */
export const TRUST_KEY_CONFIRM = '\r';
/**
* Charset-select sequences (`ESC ( B`), which tmux emits around styled runs and
* `stripAnsi` does not cover. Left in, they would land inside a phrase as a
@@ -65,6 +100,50 @@ export function isTrustDialogScreen(text: string): boolean {
return TRUST_PHRASES.some((p) => compact.includes(p)) && CONFIRM_PHRASES.some((p) => compact.includes(p));
}
/**
* Which option the `❯` marker sits on, or null when this text does not say.
*
* The LAST marked option wins. A pane capture holds exactly one frame and so
* exactly one marker, but the direct-PTY fallback reads an append-only buffer
* where every repaint since launch is still present — there the freshest frame
* is the one at the end, and an older one must not out-vote it.
*/
function selectedTrustOption(compact: string): { at: number; option: 'yes' | 'no' } | null {
let selected: { at: number; option: 'yes' | 'no' } | null = null;
for (let at = compact.indexOf(SELECTION_MARK); at >= 0; at = compact.indexOf(SELECTION_MARK, at + 1)) {
const after = compact.slice(at + SELECTION_MARK.length).replace(OPTION_NUMBER_PREFIX, '');
if (after.startsWith(YES_OPTION)) selected = { at, option: 'yes' };
else if (after.startsWith(NO_OPTION)) selected = { at, option: 'no' };
}
return selected;
}
/**
* The single keystroke that moves this dialog one step closer to "yes", or null
* when the screen does not show clearly enough to touch.
*
* One step per call on purpose: the caller re-reads the screen between
* keystrokes, so a moved cursor is CONFIRMED before Enter is pressed rather than
* assumed. Firing arrow+Enter together would re-create the failure this exists
* to prevent whenever the arrow is dropped (Ink drops keystrokes while it is
* still mounting a widget) — the Enter would then land on "No, exit".
*
* Returning null is the safe answer, not a failure: an unreadable frame means
* wait for the next repaint, and a layout whose options this cannot name means
* leave the dialog to the human. The caller's startup window bounds the waiting.
*/
export function trustDialogNextKey(text: string): string | null {
const compact = compactScreenText(text);
if (!compact.includes(YES_OPTION)) return null; // no trust option to steer onto
const selected = selectedTrustOption(compact);
if (!selected) return null; // marker missing, or not on an option we recognize
if (selected.option === 'yes') return TRUST_KEY_CONFIRM;
// On "No, exit". Which way the trust option lies is read from THIS frame — it
// sits below in 2.1.252 and above in the numbered layout before it — so the
// order flipping again costs a repaint, not a killed session.
return compact.includes(YES_OPTION, selected.at) ? TRUST_KEY_DOWN : TRUST_KEY_UP;
}
/**
* How long after the pane starts the dialog is still plausible. It renders
* before the main UI, so this only has to cover a slow first launch; leaving it
@@ -72,16 +151,21 @@ export function isTrustDialogScreen(text: string): boolean {
*/
export const TRUST_DIALOG_WINDOW_MS = 90_000;
/** Minimum gap between two Enter presses, and between two screen reads. */
/** Minimum gap between two keystrokes, and between two screen reads. */
export const TRUST_DIALOG_RETRY_MS = 1500;
/**
* Attempts before giving up and leaving the dialog to the user. A keystroke can
* land while Ink is still mounting the widget and be dropped, which is the other
* half of why sessions got stuck here; retrying costs nothing, but retrying
* forever would hammer Enter into whatever came next.
* Keystrokes before giving up and leaving the dialog to the user. A keystroke
* can land while Ink is still mounting the widget and be dropped, which is the
* other half of why sessions got stuck here; retrying costs nothing, but
* retrying forever would hammer Enter into whatever came next.
*
* Six rather than three because answering is no longer one press: the 2.1.252
* layout needs an arrow onto the trust option and then Enter, each confirmed
* against a re-read of the screen, so a cap of three left only one dropped
* keystroke of slack.
*/
export const TRUST_DIALOG_MAX_ATTEMPTS = 3;
export const TRUST_DIALOG_MAX_ATTEMPTS = 6;
/**
* How much of the append-only terminal buffer to read on a direct-PTY session,

Some files were not shown because too many files have changed in this diff Show More