Commit Graph
1478 Commits
Author SHA1 Message Date
shenlvkang-collabandClaude Fable 5.1 d9eeb039db feat(webview): open localhost links through a proxied web tab from another device
An agent prints `http://localhost:5173/` (a dev server, a preview it just
served) and the user taps it on a phone. That address only exists on the
Codeman box, so the link was a guaranteed connection error from any other
device — while the web-tab proxy fetches from the server, where it works.

A loopback link (`localhost`, `*.localhost`, 127/8, 0.0.0.0, ::1) activated
in the terminal or clicked in the Response Viewer now opens as a proxied
web tab whenever the Codeman page itself is not on that box. A saved
proxied dashboard on the same origin is reused, with the link's own path,
query and fragment opened inside it (a mounted frame is navigated, not torn
down, so its state survives); otherwise one is saved under its host:port,
sandboxed like any other web tab, so it is in the Run dropdown next time.

Only loopback is routed this way. A LAN or tailnet address may well be
reachable from the device (a VPN, the same Wi-Fi) and a direct open is the
cheaper, richer path, so those keep opening in a new browser tab; on the
box itself every link opens directly. The terminal link provider and the
viewer's click handler consult one hook and fall through to their existing
behaviour when it declines.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
2026-09-10 12:48:47 +08:00
Ark0N e3d5fd90cd Merge pull request #392 from JDProfresh/fix/ios-safari-toolbar-gap
fix(mobile): lift the iOS Safari toolbar by the measured chrome overlap
2026-09-10 03:10:12 +02:00
Ark0N 713f632a64 Merge pull request #397 from irisitymichaelgrundberg/fix/detached-session-owns-its-pane-size
fix(terminal): let a detached session's own window own its pane size
2026-09-10 02:58:54 +02:00
Ark0N 92b5dfacb0 Merge pull request #396 from irisitymichaelgrundberg/fix/terminal-font-settle-before-fit
fix(terminal): fit the terminal only once the terminal font is measurable
2026-09-10 02:58:48 +02:00
Ark0N 890a1b0902 Merge pull request #395 from irisitymichaelgrundberg/fix/full-history-replay-row-alignment
fix(terminal): keep row alignment in the full-history pane replay
2026-09-10 02:58:42 +02:00
Ark0N a360763890 Merge pull request #394 from irisitymichaelgrundberg/fix/ctrl-v-pastes-twice
fix(paste): handle only the first paste event the Ctrl+V trap receives
2026-09-10 02:58:36 +02:00
Codeman maintainer 57899f879e feat(ui): add a Blur entrance animation on all four surfaces
An iOS-style focus pull: the thing arrives out of focus and the blur fades
off it as the opacity comes up. Opacity leads the blur (full opacity around
45%, blur still lifting), which is what separates it from a cross-fade.
Ships on tabs (440ms), agent windows (560ms), the terminal pane (520ms) and
connection lines (380ms), plus a `Soft focus` theme that sets all four.
Default stays `legacy`, so an untouched install is unchanged.

The terminal pane is the one surface that cannot blur itself the documented
way, and `blur` takes a deliberate exception to the "never a filter on
.terminal-container" rule. Every alternative was measured against a live
xterm and does not work: a backdrop-filter veil on ::before blurs perfectly
while STATIC, and Chrome silently drops the backdrop the moment ANY
animation runs on that pseudo-element (the veil computes blur(15.3px) while
the text behind it stays razor sharp); driving the radius from rAF buys the
same full-screen blur per frame plus main-thread work. The cost the rule
exists to avoid is inherent to blurring a terminal, so the style buys it
knowingly: opt-in, off by default, one ~520ms run per session open, class
straight back off, will-change still unset. Worst-case price, headless
SwiftShader with no GPU: frame deltas 16.7ms -> 33.3ms for the run, against
16.7ms flat for `fade`. cols x rows measured unchanged at 178x38 before,
during and after, so FitAddon never sees it.

The line entrance animates `filter` too, where each line already carried
its glow. Both kinds now hold it in --line-glow and both keyframes say
`blur(N) var(--line-glow)`, so the function lists match and interpolate
instead of the glow vanishing for the run and popping back (a lineage
line's glow is a different colour, set per element). Its 100% frame omits
`opacity` on purpose so the endpoint comes from the element's own resting
value: 0.9 subagent, 0.72 lineage, 0.95 working.

test/entrance-animations.test.ts is a new static guard over the whole
feature, not just this style: the rule -> keyframes -> theme-option chain a
style silently does nothing without, the terminal's paint-only property
allowlist (the FitAddon rule), the --line-glow contract, and reduced-motion
coverage. Mutation-checked both ways.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 02:57:40 +02:00
Codeman maintainer d4fe3afc9d feat(files): raise the download cap to 2GB and stream /api/download
The 50MB cap on file-raw, the attachment /raw route and /api/download was
memory protection for a `readFile()` that no longer exists: file-raw and
/raw were rewritten to stream through `sendFileBody()` and answer Range
requests, so size costs a read stream rather than RSS (measured: a 600MB
download moved peak RSS by ~37MB). All the cap still did was refuse
legitimate downloads of build artifacts, videos and archives.

It is now MAX_FILE_DOWNLOAD_BYTES in config/buffer-limits.ts, default 2GB,
env CODEMAN_MAX_DOWNLOAD_BYTES, 0 = unlimited. `parseByteLimitEnv()` is
separate from the `parseInt(...) || default` idiom used elsewhere in that
file precisely because that idiom reads 0 as falsy and would silently
restore the default for the one value that means "no limit".

/api/download was the last route that really did buffer the whole file. It
now shares sendFileBody() with the other two, so it streams, advertises
Accept-Ranges, and is resumable. Its Content-Disposition also goes through
buildContentDisposition() rather than raw interpolation.

Refusals move from 400 to 413 across all three, which is the correct status
for the case; with the cap at 2GB it is a path almost nothing reaches now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 02:57:07 +02:00
Michael Grundberg 77fcd65b4a fix(terminal): force the re-measure, bound the wait, and test both
Review of the previous commit found that waiting for the font does not, on its
own, do anything.

`FitAddon.proposeDimensions()` measures nothing — it divides the container by a
CACHED cell size, and xterm refreshes that cache only from `open()`, from a
resize that actually changed the grid, and on a device-pixel-ratio change.
Nothing in it listens for font loading. So a fit that runs after the font
arrives can still divide by the fallback cell, propose the grid it already has,
and short-circuit before anything re-measures. The wait now ends by calling
`_charSizeService.measure()` itself, which is the step that makes the following
fit see the real font. Private API, as FitAddon's own dependency on `_core` is,
and guarded because a terminal can be disposed mid-wait.

The wait was also unbounded, and it sat behind the buffer-load gate. Neither
`FontFaceSet.load()` nor `FontFaceSet.ready` has a deadline, so a font request
that never settled left the tab spinning with live output queued behind it —
permanently, and on every session, since they share one promise. The comment
claimed the opposite ("a font that never loads must not block the terminal, so
this always resolves"), which was true of the per-face loads and false of
`ready`. It is now raced against TERMINAL_FONT_WAIT_MS, and the await moved
ahead of `_beginBufferLoad` so a slow font cannot hold output back at all —
which also removes the stale-select interaction with `_restoringFlushedState`,
since that flag is not yet set when the wait runs.

The awaited set no longer includes faces that cannot move the measured cell.
The bundled symbols font is ~1.2MB of private-use-area glyphs and xterm
measures `W`, so awaiting it put a megabyte in front of the first frame for
nothing; the generic families match no FontFace at all.

A runtime font change had the same race the boot-time one did:
applyTerminalFontFamily wrote the new family and fit on the next line, against
a family the browser might not have loaded. It now re-arms the wait and fits
again when it settles.

The claim that this could not be tested was wrong: the repo's vm harness
reaches both halves. The new suite pins the family filter, the forced
re-measure, the deadline, a rejecting load, a browser with no font API, and a
terminal disposed mid-wait — plus the four ordering properties in
selectSession, including that iOS Safari's synchronous focus still precedes the
first await. Each assertion was checked by reverting its fix.

Also corrects the docstring's reason for calling `document.fonts.load` (the
stylesheet is render-blocking and long parsed by then; the real reason is that
the WebGL renderer rasterises through a canvas atlas, and canvas text never
triggers a CSS font fetch), restores the JSDoc block the previous commit
displaced from getTerminalDimensions, and fixes a comment that described the
first fit as already having run when the mobile-Safari branch defers it.
2026-09-09 16:57:42 +02:00
Michael Grundberg 2b57c595df fix(terminal): gate the row-preserving skips on a capture, not the query flag
Review of the previous commit found the guard inverted: the three skips keyed
on `?full=1`, which is only what the client asked for. When the capture comes
back null — ENOBUFS, a timeout, a vanished pane, or a session with no mux at
all — the reply falls back to the byte history, which IS a stream of
successive frames and still needs stripping. Gating on the request returned it
whole: measured at 82KB against 4KB for the same buffer without `full=1`. A
direct-PTY session takes that path on every first selection, not only during
an outage. The skips now key on `isFullCapture`, meaning a capture arrived.

Three further defects the same review surfaced, all on this path:

Keeping the trailing rows is only sound when a cursor move follows to count
back up from them. On the two branches where the cursor query fails there is
no move, so the caret was left at the bottom of the pane — worse than before.
The cursor is now read first and settles both decisions together.

The move is relative rather than absolute. `CUP` numbers rows from the top of
the browser's screen, so it is only right while the browser's row count equals
the pane's, and `resizeWindow` does not wait for tmux, so a capture can be
taken before a requested resize applies. Measured against real tmux with a
browser four rows shorter than the pane: the absolute move lands on a blank
row, the relative one lands on the caret's row.

An all-blank pane no longer reads as content. Retaining trailing rows and
appending a move made it non-empty, and the caller treats non-empty as "replay
this", so a blank screen would have replaced real history — the downgrade
`_replayWouldShrinkBuffer` refuses, arriving from the server side where that
guard cannot see it.

The documentation claimed one line per screen row. `-J` joins a hard-wrapped
row into its logical line, so that is false whenever any row wrapped: measured
at 10 lines for a 12-row pane. Both entries now say what actually holds, and
the stale "NOT repositioned" contract in the mux interface is updated too.

Tests: the byte-history fallback is stripped, an empty capture leaves history
intact, and the extracted helpers are unit-tested directly rather than through
source-text matching. The slice window in the capture test is bounded at the
next method, having overrun into its neighbours.
2026-09-09 15:35:02 +02:00
Michael Grundberg 070e8da81b fix(terminal): yield only the resize send, and take sizing back on redock
Review of the previous commit found four defects in it.

The guard sat above the local fit, so it suppressed a reflow as well as the
server write. tab-rail-resize performs its single settle-time refit through
sendResize and has no fallback for a truthy activeSessionId, so dragging the
rail stopped reflowing a detached session's terminal in the dashboard. The
mobile-keyboard guard fourteen lines below already draws the line correctly —
withhold the send, never the reflow — and the guard now sits after the fit.

_lastResizeDims is one value for the whole window, and both guards skip
updating it, so while a popup owns a session that value no longer describes
the PTY. _redock repaired it only for the active session. Pop out A, switch to
B, close the popup: selecting A later found unchanged dimensions, returned
"unchanged", and selectSession skipped its 400ms redraw wait — while the
server, comparing against the real pane, did resize and did raise SIGWINCH, so
the fetch painted the pre-redraw frame. _redock now clears the record on every
path, active or not.

_redock could also fire a resize for a session already gone: _onSessionDeleted
redocks before cleanup, so the id can be dead and the request is a guaranteed
404. It now checks the session still exists.

restoreTerminalSize — the header's redraw button and Ctrl+Shift+R — silently
did nothing for a detached session while still reporting success with
dimensions nothing was set to. It now says the session is sized by its own
window, where the same button works.

The `force` comment claimed a client-side dedupe that does not exist; the
deduplication is server-side against the real pane. Corrected to say what the
flag actually buys. The _redock doc comment now records that the function
writes to the server and is not idempotent.

Tests: _redock was the untested half and is the half three of these defects
sit in. It now has coverage for clearing the stale record on both the active
and inactive paths, re-asserting only for the session being shown, and staying
silent for a deleted session. The existing sendResize test now asserts the
local fit still runs.
2026-09-09 15:24:55 +02:00
Michael Grundberg 0e82443222 fix(terminal): fit the terminal only once the terminal font is measurable
Opening a session could render its frame with characters spliced into each
other, as though two frames were overlaid — a status-line fragment landing
in the middle of a file path, for instance. Resizing the browser window
cleared it.

The first fit runs while the browser is still painting with a fallback font.
A cell measured against that font has a different width and height from one
measured against the terminal font, so the fit produces the wrong column and
row count. Codeman sizes the pane to it and replays the capture. When the
font finishes loading the measurement changes, the pane is resized a second
time, and the CLI repaints for a shape that does not match the frame already
on screen. Its later partial updates then land on the wrong rows.

selectSession now waits for the font before it measures, so the pane is
sized once, at the size that sticks, and the capture is taken at that size.
The wait always resolves, so a font that never loads cannot block a
terminal, and it resolves immediately once the font is in, so a tab switch
pays nothing after the first load.

document.fonts.ready alone is not enough: it can resolve before the
stylesheet declaring @font-face has been parsed. document.fonts.load for
each family in the stack is what actually requests the faces.

Measured on a session opening at 2328px wide: the cell went from 8.43x16.00
to 8.00x21.00 roughly 900ms in, moving the grid from 112x36 to 118x28 after
the replay had already been painted.
2026-09-09 14:20:34 +02:00
Michael Grundberg 323730a29d fix(terminal): keep row alignment in the full-history pane replay
Switching to a session left the caret one row below the composer's input
line, on the box border, and every cursor-relative update the CLI sent
afterwards was measured from the wrong row. Any fresh output repaired it,
because the CLI then repainted the whole frame.

Two things were wrong with the full-history replay, and they compound.

The capture never restored the cursor. The visible-frame path ends with an
absolute cursor move back to the pane's position; the linear path returned
its text and left the caret wherever the last character landed, which for an
agent CLI is the bottom-most row carrying text — the status line.

The rows it addressed did not line up with the pane's rows either. Four
transforms ran over the capture and each can delete a line: the trailing
blank rows were stripped, redraw-bloat stripping ran, the trim that cuts
everything above the Claude banner ran, and leading whitespace was removed.
All four are right for a byte stream of successive frames. A capture is the
rendered pane, one line per screen row, so each deletion shifted the frame
out from under the restored cursor.

The full-history path now appends the pane's own cursor position and keeps
every row, so row N of the reply is row N of the pane. The visible-frame and
tail paths are untouched.

Restoring the cursor is what makes row alignment load-bearing here, and
neither CLAUDE.md nor the architecture invariants said so — which is how
four line-deleting transforms accumulated on the path. Both now record it.

Verified against a live 315x59 pane: the reply carries 59 rows, its row 55
is the composer's input line matching tmux, and it ends with the cursor move
that lands there.
2026-09-09 14:20:34 +02:00
Michael Grundberg 5ac516dd3b fix(terminal): let a detached session's own window own its pane size
Popping a session out left both windows sizing the same pane. The dashboard
keeps the session active and keeps measuring it, and its terminal is
narrower than the popup because the session rail takes width the popup does
not have. One PTY cannot hold two sizes, so the CLI drew frames that fit
neither window and the popup showed a garbled frame.

sendResize and the debounced window-resize handler now stand aside for a
session this window has marked detached. A solo window is exempt, since it
is the owner. _maybeRefetchFullHistory already stood aside on exactly this
condition, so the rule is not a new one.

Sizing has to come back when the popup closes: while it owned the session
the dashboard sent no resizes, so the PTY still holds the popup's geometry.
_redock now re-asserts, with force, because the dimensions the dashboard
last sent are the ones it is about to send again.

Reproduced with a dashboard and a popup on one session: before, the pane
sat at 315 columns while the popup rendered 289. After, both report the
same size and the popup's frame matches the pane exactly.
2026-09-09 14:20:34 +02:00
Michael GrundbergandClaude Opus 5 b87bc6871b fix(paste): handle only the first paste event the Ctrl+V trap receives
Ctrl+V in the terminal inserted the clipboard text twice. Right-click →
Paste inserted it once.

`_handleImagePaste()` appends a hidden contenteditable div, focuses it, and
reads the clipboard out of the paste event that lands there. Two separate
routes deliver that event for a single keypress. The function issues
`document.execCommand('paste')` itself, which in Firefox dispatches a
trusted paste event and then returns false, because the trap cancels the
event and the command never completes; Chromium and WebKit refuse that
command and dispatch nothing. The keydown's own default action delivers the
other, because xterm calls the custom key handler before its own `cancel()`,
so returning false never calls preventDefault. Firefox therefore ran the
trap's listener twice and both runs reached `terminal.paste()`. The
context-menu paste involves no keydown at all, which is why that path stayed
correct.

The trap now accepts the first paste event and cancels every later one, so
how many paste events a browser delivers no longer changes what the PTY
sees. Measured on a live install, one Ctrl+V each: Firefox two events and
two writes before this change, Chromium and WebKit one and one, and every
engine one write after it.

The `execCommand('paste')` call stays. Stripping it out also ends the
doubling, and all three engines still deliver one event without it, since
`trap.focus()` has already run when the key's default action resolves. It is
kept because the trap technique arrived in #84 for plain HTTP and for
mobile, and a desktop measurement says nothing about real iOS Safari or
Android Chrome: where a browser aims the default action at the element
focused when the keydown began, the command is the only route into the trap,
and the trap is the only place clipboard image blobs are read.

test/image-paste-trap.test.ts loads image-input.js into a `node:vm` context
with a fake document and fires two paste events at the trap. It covers text
and images, and fails on the old code with the text pasted twice and the
image uploaded twice.

Docs: the invariant goes into docs/architecture-invariants.md as a Terminal
paste section and into CLAUDE.md as a Frontend entry, both recording the
measured event counts and why the redundant call is still there. README.md
and the Keyboard Shortcuts and Input and Voice wiki pages gain a Ctrl+V row,
which all three tables were missing while listing every other clipboard
binding.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 16:49:50 +02:00
JD c367b12f77 fix(mobile): lift the iOS Safari toolbar by the measured chrome overlap, not 100vh minus the visual height
The phone block lifted the toolbar (and padded .main) by (100vh - --app-height) on iOS Safari to clear a bottom bar that position: fixed elements were assumed to sit behind. On iPhone Safari fixed elements already stop above the bar, and 100vh is the large viewport with the bar collapsed while --app-height is the visual viewport with it expanded, so the expression measures the bar's collapsible height and shows up as an empty band between the toolbar and the bar whenever the bar is expanded. The terminal was padded by the same amount.

The lift is now --chrome-overlap, set in updateAppHeight() as innerHeight minus the visual viewport height: the distance the layout viewport that anchors fixed elements extends past the visible area. That is 0 on iPhone Safari, so the toolbar meets the bar, and it is the overlap itself on any browser where fixed elements really do land behind the chrome, so those keep the lift. The keyboard-visible rules, which already override the toolbar offset, are unchanged.
2026-09-08 01:39:07 -04:00
Codeman maintainer 4f2dfb4e6d fix(mobile): carry resumeId through the phone overview's past rows
#386 made Codex conversations resumable from Past Sessions, and
resumeMobileOverviewSession() correctly passes row.resumeId on to
resumeHistorySession(). The phone's own row projection never copied the
field off the unified-list item though, so row.resumeId was always
undefined there and a tapped Codex row started a FRESH session on a thread
that was already on disk. The desktop path worked; only the phone was blind.

The test fails without the projection line, and pins the other half too: a
claude row must not grow a resumeId, since the field is what distinguishes
"resume this conversation" from "start a new one".

Docs: CLAUDE.md and architecture-invariants both still described the unified
list as merging Claude transcript files. It has been three stores since this
PR (Claude's ~/.claude/projects, omp's ~/.omp/agent/sessions, codex's
~/.codex/sessions), the alias field keeps its Claude-era name without being
Claude-only, and the scanner-only rule behind resumeId was written down
nowhere.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 22:44:25 +02:00
Ark0N 344e93c824 Merge pull request #386 from irisitymichaelgrundberg/feat/codex-resume
Merging with the phone-overview resumeId fix and the two unified-list doc passages applied on master.
2026-09-07 22:43:46 +02:00
Codeman maintainer f1b7283393 fix(cli-registry): guard workDetect.workingLine like every other config regex
#385 made the composer glyph and the working status line per-CLI registry
data, which is right, but `workingLine` arrived as a config-supplied regex
validated with a bare `new RegExp()`. That skips `compileVersionRegex()`,
the helper the registry uses for exactly this: a `~/.codeman/clis.json`
override can set the field, the compiled pattern is run against every
accumulated PTY chunk and every pane capture, and a nested quantifier there
backtracks on the event loop for the whole server rather than one session.

Route it through the helper in both places, which are not redundant: the
schema refine rejects the entry at LOAD time so a bad pattern never reaches
a session, and `_workingLinePattern()` compiles through the same helper so
the runtime cannot hold a pattern the schema would have refused. The helper
returns null instead of throwing, so the Claude-pattern fallback stops being
a try/catch and becomes structural. Both shipped patterns compile unchanged,
and Claude's is behaviourally identical to CLAUDE_WORKING_LINE_PATTERN.

Also match the Codex footer case-insensitively on the E. It was
characterised against codex-cli 0.152.1, which prints a lowercase `esc`;
a version capitalising it would make the whole fix silently inert, since
the pane would simply never look like it was working.

Docs: CLAUDE.md, architecture-invariants and cli-registry.md all still
stated the Claude-mode-only rule this PR retires, and none of them named
the new capability or the regex guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 22:42:20 +02:00
Ark0N a49be03f96 Merge pull request #385 from irisitymichaelgrundberg/fix/work-detection-external-clis
Merging with follow-up fixes applied on master: workingLine routed through compileVersionRegex() in both the schema refine and _workingLinePattern(), the Codex footer matched case-insensitively on the E, plus the doc passages that stated the retired Claude-mode-only rule.
2026-09-07 22:41:25 +02:00
Codeman maintainer 8ee7926e27 feat(agent-cases): tag agent-spawned case dirs and sweep their leftovers
A long orchestration creates one case directory per worker and deleting the
sessions never removed them, so ~/codeman-cases accumulated scratch folders
that were indistinguishable from real projects. They are now labelled and
have a cleanup path.

- src/agent-case-marker.ts: a case dir quick-start CREATES for an agent-driven
  spawn gets a .codeman-agent-case.json marker (when, by whom, parent session,
  mode). Only the create branch writes it, so a linked case, a cloned repo or
  any pre-existing path is never labelled; reading is total, so a malformed
  marker means "not agent-created" rather than a half-trusted entry.
- The signal is the new X-Codeman-Agent-Origin header the skill preamble sets
  on its shared curl (preamble bumped to 1.22.0), or an agentOrigin body
  field, falling back to a resolved parentSessionId so a worker spawned by a
  stale skill copy is still labelled.
- GET /api/cases publishes it as agentCreated; GET /api/cases/agent-created is
  a read-only cleanup listing adding inUse and modifiedAt; Add Case -> Manage
  badges each case and offers a review-then-delete sweep that names every
  directory in its confirm and skips any case a live session is working in.
  Removal stays on the existing DELETE /api/cases/:name.
- Agent preamble caches are collected too: ~/.cache/codeman-agent-<id>.sh was
  written per claude session and never removed (236 leftovers measured on a
  working machine). Now deleted with the session and swept at boot, guarded by
  a live-session keep set plus a 7-day age floor.

Verified end to end on an isolated instance: marker written for header, body
and lineage-only spawns, absent with no agent signal and for a pre-existing
directory; inUse flipping on session end; badge, sticky bar, confirm and sweep
driven in a browser; preamble seeded on create, removed on delete, boot sweep
taking only the aged orphans.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 19:09:24 +02:00
Michael GrundbergandClaude Opus 5 2f9663e389 Merge branch 'master' into feat/codex-resume
master and this branch both rewrote the two `_claudeSessionId` resets inside
`start()`, so `src/session.ts` conflicted at both of them.

master's commit ccfda623 puts `restoredConversation` at the head of each
fallback chain. A restored mux attach means the CLI never stopped, so a
`/clear` before the Codeman restart may already have moved it to a
conversation the launch id knows nothing about. The persisted chain's tail is
that conversation, and the CLI's own hook reported it first-hand.

This branch adds `this._codexConfig?.resumeSessionId` to the same two chains,
so a resumed codex session keeps its thread-id alias across every mux reattach
and boot recovery.

Both fixes belong. Each chain now reads restoredConversation, then
_resumeSessionId, then omp's alias, then codex's alias, then the launch id.
The comments from both sides are kept.

test/session-claude-conversation-chain.test.ts pins the shape of those two
assignments by matching the source text, and its pattern named omp's alias as
the last term before `this.id`. Codex's alias now sits between the two, so the
pattern widens to pin the ends of the chain and let the middle grow. A `[^;]`
run cannot cross a statement boundary, so each match is still one assignment.

Checked on the merged tree: typecheck, lint, prettier and the frontend syntax
check all pass, and the CI suite runs 6721 tests green across 349 files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 08:48:31 +02:00
Codeman maintainer 92af855ce4 fix(base-path): keep the crash beacon under the mount, strip CODEMAN_BASE_URL in tests, add the wiring test
The merge-time items from the #381 review. navigator.sendBeacon is not fetch,
so the base-aware wrapper never saw the two crash-diag beacons and a sub-path
install posted them to the origin root every two seconds. The test suite now
strips CODEMAN_BASE_URL like CODEMAN_GESTURE, since the constructor reads it
as a fallback and an operator who exports it would see the root-install
byte-identity assertions fail. test/base-path-server.test.ts boots a real
WebServer under /codeman and checks the ingress strip, the base injection,
the rebased redirects, the 404 envelope and a prefixed WebSocket upgrade.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 23:11:01 +02:00
Ark0N cfc8fe7e41 Merge pull request #381 from mtiller/feat/reverse-proxy-base-url
feat(web): support a reverse-proxy base URL
2026-09-06 23:10:23 +02:00
Codeman maintainer 80397fe140 fix(hooks,mobile): the merge-time items from the #367 and #368 reviews
#367 (UserPromptSubmit hook): `hook:prompt_submitted` went on the wire
unregistered; it is now in both SSE registries (158 = 158), and the hook only
lands in the run summary when the conversation actually moved, since one row
per prompt would evict useful rows from the 1000-event FIFO and clutter the
Summary timeline and /api/search.

#368 (Add Case header submit): the pending-state dimming targeted the footer
button, which the <=860px layout hides, so on a phone the only visible submit
control stayed at full brightness while a clone ran. The header button now
dims too, and a static test pins the header-submit contract so it cannot
silently disappear again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 23:05:26 +02:00
Ark0N 7991f481b6 Merge pull request #368 from shenlvkang-collab/pr/mobile-add-case-submit
fix(mobile): give the Add Case modal a reachable submit button
2026-09-06 23:02:45 +02:00
Ark0N bca1b764cc Merge pull request #367 from shenlvkang-collab/pr/claude-conversation-first-hand
fix(session): learn the live Claude conversation from the CLI's own hook
2026-09-06 23:02:33 +02:00
Ark0N 1c1773278f Merge pull request #369 from shenlvkang-collab/pr/claude-response-viewer-per-message
fix(web): render one Claude response-viewer message per model message
2026-09-06 23:02:13 +02:00
Michael Grundberg 327e440607 fix(codex): fold a codex session into its own rollout row
Review fixes for #386.

Duplicate rows. A codex conversation showed twice, once live and once as a
past rollout row, because nothing aliased a codex session to its thread id.
That is worse than cosmetic: the stale row still resumes, so clicking it
starts a second `codex resume` on a thread already open in another pane.

  - A RESUMED session knows its thread id up front, so it folds from its own
    side: add `codexConfig.resumeSessionId` to the `claudeSessionId` chain.
    Not only in the constructor — `start()` recomputes that id at two further
    points (the mux branch, and the unconditional "third reset point" whose
    own comment already warned that omitting omp's fallback there stomps the
    mux branch's resolved alias). Both listed Claude's and omp's ids only, so
    for codex every mux reattach and boot recovery reset the alias back to
    the Codeman id and the duplicate returned.
  - A FRESH session has no thread id until codex writes the rollout, so it is
    folded from the other side. The scanner now reports
    `session_meta.originator`, which is `codeman_<sessionId>` for every pane
    Codeman spawns, and `gatherUnifiedInputs()` stamps the matching live and
    persisted rows, newest rollout winning (`/new` inside the TUI leaves
    several rollouts sharing one originator).
  - Persisted rows read `codexConfig.resumeSessionId` too. A resumed session
    demoted to a persisted-only record would otherwise lose its alias, and
    the originator fallback cannot rescue that one: a resumed rollout keeps
    its ORIGINAL session_meta, so it still names the pane that created the
    thread rather than the pane that resumed it.

Identity cache. It was written as soon as the thread id was known, but codex
writes the first user message only when the user submits, so any scan in that
window pinned `firstPrompt: undefined` for the life of the process — and the
home screen, the command palette and the search-index refresh all scan.
`shouldCacheIdentity()` now keeps an identity only once the prompt is known or
the head read filled its whole window.

Also from review: both caps count emitted rows rather than file index, so a
store of sub-agent threads no longer spends the `lastPrompt` budget before the
first row that needed it; the cache is an `LRUMap` sized like the one beside
it; the unreachable filename fallback is gone; a rollout recording no cwd is
dropped rather than emitted with `workingDir: ''`; and the unified-session
module header names all three transcript stores.

Tests. The resume wiring now has cases for a row with a thread id, a row
without one, and a `resumeId` on a non-codex row; the "no continuation is
wired" case narrows to gemini/antigravity, which is no longer true of codex.
`codex-resume-alias-survives-start.test.ts` drives a real Session through
`start()` rather than asserting on pre-stamped inputs — that gap is why the
reset points went unnoticed. Plus the maintainer's own cache repro, the
tail-budget case, a no-cwd case, and merge cases for both folds.
2026-09-06 21:53:38 +02:00
Codeman maintainer a2aaea3c0e docs(file-picker): state the Home/cases nesting the right way round, and document the new fallback chain
The two merge-time edits the #383 review asked for. The comment above the
picker's fallback chain said Home is nested under Codeman Cases; on the
native default it is the other way round (~/codeman-cases sits inside ~).
And the "Filesystem path picker" paragraph in architecture-invariants still
said the picker falls back to /mnt/d, which #383 changed to: the session's
Current Folder, then the Codeman Cases root, then /mnt/d, then the first
root. No code behaviour changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-06 21:20:45 +02:00
Ark0N 097d585278 Merge pull request #383 from opticon454/fix/case-picker-default-root
fix(file-picker): default the case picker to Codeman Cases, not Home
2026-09-06 21:03:00 +02:00
Michael Grundberg 8285fff91c feat(codex): list codex conversations and resume them
Codex conversations never appeared in the session list, and the resume path
skipped codex, so picking one back up meant finding its thread id by hand and
POSTing codexConfig.resumeSessionId to /api/sessions.

Two gaps caused it:

- The unified list is built from ~/.claude/projects plus omp's own store.
  Codex writes to neither: its rollouts live in ~/.codex/sessions/<y>/<m>/<d>.
- terminal-ui.js sends a continuation only for the CLIs with a
  "continue most recent" flag. Codex has no such flag — it names a thread by an
  exact id — and nothing supplied one.

Add codex-transcript.ts, the codex analog of omp-transcript.ts, and wire it into
gatherUnifiedInputs() beside the omp scan. A rollout row carries `resumeId`, the
thread id `codex resume` takes, and the resume path sends it as
codexConfig.resumeSessionId.

`resumeId` is what keeps the two kinds of row apart: only a transcript scanner
sets it, so a LIVE codex row — whose sessionId is Codeman's own uuid — can never
ask codex for a thread that does not exist.

Three things measured against a real store of 519 rollouts rather than assumed:

- Rollouts are far too large to read whole (median 407 KiB, p90 1.3 MiB, max
  25 MiB, 381 MiB total), so this reads a 128 KiB head for the identity and the
  opening prompt and a bounded tail for the most recent one. session_meta is
  written once and never rewritten, so per-path identity is cached; a warm
  rescan of that store costs ~75ms against ~470ms cold.
- codex 0.152.1 emits no event_msg/user_message rows at all. It writes
  event_msg/item_completed carrying an item.type of UserMessage. Both shapes are
  read, plus response_item as a last resort.
- That last resort sees injected context, and the first such row is the repo's
  AGENTS.md every time, so injections are dropped rather than used as titles.

Sub-agent threads (thread_source: 'subagent') are left out; codex spawns them
for itself and on a real store they outnumber the resumable threads.
2026-09-06 19:45:36 +02:00
Michael Grundberg 51957e2ed4 fix(session): let each CLI declare how its own pane shows work
A Codex session reported `isWorking: false` for its entire life, including
mid-turn. Codeman has four paths that mark a session working, and all four were
inert for Codex:

- The spinner fast path tests eight braille frames, and Codex animates none.
- The activity-streak fallback was wrapped in `!isExternalCliMode(mode)`.
- The pane probe inside `_confirmIdle` would have matched, since Codex prints
  `esc to interrupt`, but arming it required the literal glyph `❯` and Codex
  draws `›` on its composer row.
- The text detector sat inside `_processExpensiveParsers`, whose first statement
  returns early for an external CLI.

Add an optional `workDetect: { promptGlyph, workingLine }` to CliCapabilities,
so the two strings that differ per CLI are registry data rather than constants
in the detector. Claude declares its existing pair and behaves as before. Codex
declares `›` and `esc to interrupt`. The text detector moves above the
external-CLI early return, guarded on the descriptor so a CLI without one still
skips the ANSI strip that the early return used to save it.

A CLI that declares no descriptor falls back to Claude's pair, and the
activity-streak gate now reads "has a descriptor, or is not external", so the
plain shell mode keeps the behaviour it had.

Rewrite the test that asserted the old premise in its own comment, so it makes
the same guarantee for a genuinely uncharacterised CLI, and add Codex coverage
built from verbatim pane captures on Codex CLI 0.152.1.
2026-09-06 17:41:01 +02:00
Codeman maintainer 8ad2215118 fix(docker): close the three adoption gaps the negative guarantee missed
Review follow-ups to #357. Each is a path that still touched, or still hid, a
container Codeman does not own.

**Export still mutated it.** The four fail-closed layers cover create/start/
stop/remove, but `POST /api/docker-cases/:name/export` reaches the container
twice through neither: a full export `docker commit`s it, and even a
workspace-only export `docker pause`s it first for snapshot consistency. Pause
freezes the owner's processes for as long as the tar takes, on a container we
promised not to touch. Full export is refused for an adopted case (it packages
someone else's container, with their logins, into a bundle Codeman hands out);
workspace-only keeps working and no longer pauses, accepting a live filesystem
the way `tar` does on any running host directory.

**A freshly linked OWNED case became unusable.** The run menu now probes the
container for its CLIs, and a failed probe hides every agent mode behind the
reason. For an adopted case that is right. For an owned one the container does
not exist until the first session launches it, so every newly linked Docker case
answered `container "codeman-case-x" not found (adoption never creates a
container — start it yourself first)` and offered nothing but Shell, for a
container the launch chain was about to create itself. A failed probe is
recorded only when the case is adopted; `CaseInfo.docker.owned` is on the wire
so the frontend can tell them apart. Verified in a browser: owned-with-no-
container offers all ten modes and no notice, adopted-but-stopped offers Shell
and says why.

**Multi-user gating.** Adoption is admin-only, unlike `docker-link` beside it.
Linking creates OUR container, whose sole bind mount `isWorkingDirAllowed` has
already confined to the caller's space; an adopted container's mounts are
whatever its owner gave it, so one mounting `/` hands the adopter a shell over
the whole host — exactly the workspace scoping multi-user mode exists to
enforce. Listing the engine's containers and browsing directories inside an
arbitrary one are machine-level reads and follow the docker-HOST policy for the
same reason. The preflight is deliberately not admin-only: the run menu fires it
for every docker case, so it admits a non-admin for a container already linked
to a case they can access, and nothing else.

Verified end to end against a real pre-existing root container (alpine + tmux,
no bind mounts): adopt, claude session inside it, workspace export, session
close and case unlink all left `StartedAt`, `RestartCount`, `Pid` and `Paused`
untouched; the pane ran the CONTAINER's claude, without
`--dangerously-skip-permissions`; a stopped container was refused at both
preflight and launch and was never started.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1
2026-09-05 16:22:54 +02:00
Codeman maintainer 3d8ffcb9a2 Merge pull request #357 from dignfei/feat/docker-adopt-existing-container
feat(docker): attach a case to an already-running container

Conflicts came from work that landed after the PR was opened, and each is
resolved onto the newer abstraction rather than by keeping the older code:

- `defaultDockerCommandForMode` is registry-driven since #347, so the PR's
  `runsAsRoot` arm became `overlays.docker.rootCommand` (claude only). Claude
  Code still refuses `--dangerously-skip-permissions` as root in 2.1.261 and the
  refusal is visible only inside the container, so an adopted root container
  otherwise just shows a dead pane. Which flag to drop is a per-CLI fact, and
  `test/cli-registry-no-id-branching.test.ts` forbids expressing it as a branch.

- The probe's mode list and its mode -> binary table both duplicated the
  registry. They now read `enabledCliIds()` / `discovery.binaries[0]`, which is
  also what fixes the merge's silent regression: the hand-written list predates
  `omp`, and the run menu gates every docker case on this probe, so owned
  containers would have lost that mode. `shell` needs no arm — it declares no
  binary, so it is dropped from the lookup and reported available regardless.

- The per-mode `mode === 'claude' && !cliDir` chain in `tmux-manager.ts` is one
  `missingCliMessage(mode)` gate since #347; the PR's docker exemption moved onto
  it. Its test now pins the single gate instead of counting seven arms.

- The create arm keeps #349's swap-limit warning filter, which the adopted arm
  never reaches; the run-mode list gains `omp` from #353.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1
2026-09-05 16:21:51 +02:00
DevvynandClaude Sonnet 5 06febfa032 fix(file-picker): default the case picker to Codeman Cases, not Home
The "Link Existing" case picker opens with an empty path and no
sessionId, so the browse endpoint's fallback root picked whichever
root happened to be first in the list — which was always `Home`.

On the native default that's harmless (~/codeman-cases nests inside
Home anyway), but a Docker deployment binds CODEMAN_APPDATA_PATH
(Home) and CODEMAN_CASES_PATH at unrelated host paths, so the picker
opened somewhere with no cases in sight. Worse: if CODEMAN_CASES_PATH
is ever changed after cases already exist, the old cases directory
lingers, still reachable, under Home — indistinguishable at a glance
from the real one under the new Codeman Cases root.

Prefer the Codeman Cases root in the fallback chain, ahead of the
generic roots[0].

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
2026-09-05 17:14:05 +08:00
Michael TillerandClaude Opus 4.8 7e4914d991 feat(web): support a reverse-proxy base URL (--base-url / CODEMAN_BASE_URL)
Codeman can now be mounted under a sub-path behind a reverse proxy that
forwards the prefix unchanged (e.g. https://host/codeman/). Default is `/`
(root), which is byte-identical to the historical behavior.

Design — few choke points, mirrored ingress/egress:
- src/config/base-path.ts: pure single-source normalize/validate/join/strip.
- Server ingress: stripBasePath() inside Fastify rewriteUrl, so routes stay
  declared prefix-agnostic; un-prefixed requests (hooks, health, docker bridge
  hitting the raw port) pass through unchanged.
- Server egress: one onSend hook prepends the base to root-absolute Location
  headers (covers all redirects).
- HTML: renderIndexHtml points <base href> at the mount and injects
  window.__CODEMAN_BASE__ — ONLY when a base is set (inert at root).
- Frontend runtime URLs: CodemanBase.url() route builder in constants.js,
  applied transparently by a fetch wrapper and explicitly at the
  EventSource/WebSocket/window.open/<img|iframe|a>-src sites.
- sw.js derives its base from self.location; manifest uses relative start_url/scope.
- Web-tab proxy: proxyPrefixFor(cap, basePath) is the single base-aware root that
  cascades to the injected <base>, HTML/attr rewrites, runtimeUrlShim, Set-Cookie
  Path and Location; capabilityFromReferer strips the base off the browser Referer,
  while the ingress parsers stay base-agnostic (rewriteUrl already stripped it).

--base-url rides the daemon relaunch (buildWebArgs) and the service unit
(resolveServicePlan). constants.js is guarded against a missing `window` for
isolated unit-test contexts.

Tests: test/base-path.test.ts (pure helpers), base-path coverage in
webview-proxy/render-index-html/daemon-control; CodemanBase stubbed in the
vm-isolated panels-ui test contexts. Docs: Remote-Access.md (sub-path section +
nginx example), security-architecture.md env table, CLAUDE.md pattern.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XUkPBxbumnct6qSrx4JDju
2026-09-04 16:08:57 -04:00
Codeman maintainer 2ab21c1b32 fix(webview): revoke proxy capabilities on logout and stamp Referrer-Policy
WebviewCapabilityStore.revokeOwner() shipped for two releases with a docstring
claiming logout called it and no caller at all. The capability is a bearer
credential exempt from cookie auth with a rolling TTL refreshed on every use, so
a proxy URL that leaked (browser history, a screenshot, a dashboard with a loose
referrer policy) stayed valid for as long as anything kept polling it.

- POST /api/logout revokes the caller's capabilities (all of them in single-user
  mode), the admin forced logout revokes the target user's, and user deletion
  revokes whatever that user had open. revokeOwner returns the count for the
  admin audit line.
- Proxied responses carry `Referrer-Policy: same-origin` and the upstream's own
  policy is dropped: every URL inside the frame carries the capability, and a
  dashboard on no-referrer-when-downgrade or unsafe-url handed it to any
  third-party host it linked. Verified with Playwright that a sandboxed frame
  under an upstream `unsafe-url` sends no Referer to a third party while the
  root-absolute fetch and the CSS-triggered 404 fallback still reach the
  dashboard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:13 +02:00
Codeman maintainer 550e08a791 fix(webview): refuse link-local and cloud-metadata targets on the resolved address
The web-tab proxy, its Test probe and its WebSocket relay accepted any http(s)
host. A live PoC relayed an IMDSv2-shaped PUT with custom headers to a loopback
echo server through a capability and no cookie, and 169.254.169.254 (decimal,
hex, IPv6-mapped, or via a DNS name) was as valid a dashboard as any other.

Loopback and RFC1918 stay allowed on purpose: a localhost Grafana is the feature.
Only link-local and the fixed cloud-metadata addresses are refused
(169.254.0.0/16, fe80::/10, fd00:ec2::254, 168.63.129.16, 100.100.100.200,
metadata.google.internal), at three stages that are each load-bearing:

- the Zod schema, so a save gets a clear refusal;
- a synchronous hostname check at every connect site, because net.connect skips
  DNS for an IP literal and a lookup hook never sees one;
- a `lookup` hook on an undici Agent (webviewFetch) and on the ws client, which
  judges the RESOLVED addresses of a name and refuses when any is blocked. This
  is what closes DNS rebinding, which a hostname-string check cannot.

Adds undici@^6 so the proxy runs the package's own fetch with the package's own
Agent; a package Agent handed to Node's bundled fetch can mismatch protocols.

Verified live on an isolated beta: 169.254.169.254.nip.io (a real name resolving
to the metadata address) is refused by probe, proxy (403) and WS relay (4003),
while 127.0.0.1.nip.io still passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:12 +02:00
Codeman maintainer 99ad9cb236 fix(docker): never exit the server unless something is known to restart it
#373 restarts the Compose container by exiting the server, which is right for
the shipped deployment: `restart: unless-stopped` relaunches it. The updater
verified that policy through the Docker socket and, when it could not (no
socket mounted), failed open and exited anyway. Failing open is the correct
choice for the GATE, where refusing would block every install without a
socket, but not for the kill: a container the daemon does not restart goes
down for good, with no UI left to recover it from. That is exactly the case a
plain `docker run` of this image without `--restart` produces, and the image
sets CODEMAN_IN_CONTAINER=1 itself, so it takes the container path.

The decision now happens server-side, where both the socket and the Compose
env are reachable, and rides down to the script as `--restart-by-exit 0|1`.
It is 1 when the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (added there
and only there, since that file is what sets the restart policy; the image ENV
deliberately does not) or when the daemon confirmed an auto-restart policy.
Otherwise the build still lands, the status becomes
`completed-needs-manual-restart` with the `docker restart` hint, and the
server keeps running. The shipped deployment is unchanged in effect: with the
socket it was already confirmed, and without it the declaration now covers it.

Also: a root-run `Start-Codeman.sh` (common on Unraid) created the
fingerprint baseline's `.codeman` directory before the container's first start
and left it root-owned, which the unprivileged server could then never write
its own state into. It is chowned to PUID:PGID when running as root.

Verified with a real image build of the merged tree (classic builder; this
box's BuildKit lacks buildx): runs as uid 1000, tsc/esbuild and the toolchain
present, the four CLIs at their pins, docker/.env absent, and `docker inspect
$HOSTNAME` returns the restart policy through the mounted socket as that user.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 14:36:35 +02:00
Codeman maintainer 823f56a243 Merge pull request #373 from opticon454/feature/docker-self-update
feat(docker): restore in-app self-update in the Compose dep
2026-09-04 14:25:56 +02:00
Codeman maintainer a81e87f440 fix(cli-registry): log why clis.json was ignored, and say 0600 when that is the rule
The loader refuses a `clis.json` with any group/world permission bit, read bits
included, so a file created with a normal umask (0644) is ignored. That is a
defensible posture for a file that chooses the binaries Codeman spawns, but two
things around it made the override feature look dead: the warning said
"group/world-writable", which a 0644 file is not, and `LoadResult.warnings` was
returned to a caller nobody wired up, so nothing anywhere printed it. A user
following the docs got silence.

The message now names the rule and the command that satisfies it, the loader
logs every warning once on first load (the result is memoized, so once per
process), the module header stops claiming that nothing ever writes (the
quarantine rename of a malformed file is a write, on first use) and the
registry doc gains a short section on the override file with the 0600
requirement in it. Whether the check should relax to writable bits only is a
separate decision; this keeps the shipped behaviour and makes it visible.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:51:20 +02:00
Codeman maintainer 28b44237ae fix(remote): classify the has-session probe by exit status, and forget it once the pane is back
#355 made the remote auto-reconnect watcher revive a dead pane only when the
durable remote tmux session is verifiably still alive, which is the right rule:
a clean Ctrl-C / Ctrl-D / exit tears that session down and must never relaunch
a fresh agent. Its probe, though, read `has-session`'s stdout and treated an
empty string as "gone". `tmux has-session` prints NOTHING on success (measured
on a scratch socket: exit 0, empty stdout, the failure message goes to stderr),
so every live remote session classified as gone and transport-drop reconnects
were silently disabled along with the clean-exit revives.

The probe now goes by exit status through a pure, unit-tested mapping
(`classifyRemoteAliveExit`): 0 is alive; ssh's own 255, a timeout (`killed`,
no numeric code) and a spawn failure are unknown, which the watcher already
treats as do-not-revive; any other status is the remote command's and means
gone (tmux's 1 for a missing session, 127 when tmux is not installed there).

Two smaller things in the same area:

- The cached answer was never invalidated, so after one successful reattach a
  stale `true` would have revived the NEXT clean exit (the original bug back
  after the first transport drop), and a cached `false` from a clean exit would
  have left a manually restarted session with auto-reconnect permanently off.
  The tick now forgets the cache entry whenever the pane is seen alive.
- The fire-and-forget probe has a 15s timeout against a 5s tick, so an
  unreachable host stacked up to three ssh processes per dead session. An
  in-flight set caps it at one.

The probe command is pinned as a literal string, and the reattach-then-clean-exit
sequence is driven through the watcher in the tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:50:22 +02:00
Ark0N ee6a7af1d1 Merge pull request #355 from timkjr/pr/remote-exit
fix(remote): never auto-revive a remote session after a clean agent exit
2026-09-04 13:49:49 +02:00
Ark0N 850b00572c Merge pull request #347 from opticon454/feature/cli-registry-core
PR A: CLI registry core as a pure internal refactor
2026-09-04 13:49:35 +02:00
DevvynandClaude Opus 5 66eb01ba8f feat(docker): restore in-app self-update in the Compose deployment
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.

Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:

- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
  update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
  container on the new dist/. This is the one supervisor whose updater does NOT
  outlive the restart, which is safe only because the terminal "restarting"
  marker is written first.
- node_modules and dist are named volumes over the bind mount, so
  container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
  `npm run build` is tsc + esbuild and node-pty has no Linux prebuild.

An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.

The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.

Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.

Documented in docs/docker-self-update.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
2026-09-02 19:33:32 +08:00
Ark0N f7cf15485e feat(models): offer Fable 5.1 in the model picker and task routing (#372)
Adds `claude-fable-5-1` to the App Settings model picker and the five task-routing selects, mirroring how Fable 5 is already offered: a base option with data-ctx="1" plus its [1m] companion row. No settings-ui.js logic change, since the cards and the 1M switch are built from those options.
2026-09-02 10:48:51 +02:00
DevvynandClaude Opus 5 c5b84fb5f4 docs(cli-registry): annotate overlays.credStore as declared-for-later
Review item 4 named THREE live tables duplicating registry data. Two are now
read from the entry (`defaultRemoteCommandForMode`, `defaultDockerCommandForMode`);
the third, `resolveDockerCredentialArtifacts`, is not — and it was left neither
wired nor annotated, which is the state that item explicitly rules out.

It is not wired because the shape cannot express the live table: `credStore` is
ONE store per CLI, and `CRED_STORES` needs two for gemini (`.gemini` for the
CLI's own auth plus `.config/gcloud` for Vertex), while deepseek's entry declares
none at all even though `.dsh` is seeded. Wiring it means making the field an
array and correcting those two entries — a change to credential seeding, which
is at once the worst thing in that file to get wrong and the least covered by
tests, since every docker IO path is no-op'd under vitest. It belongs in its own
change, measured against a real container.

So it is annotated instead, at the field, in the type's declared-for-later
header, in docs/cli-registry.md, and in the pinned DECLARED_FOR_LATER list — the
last of which means wiring it later makes a test fail rather than leaving a
stale comment behind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 09:35:55 +08:00
DevvynandClaude Opus 5 6acf0dea0f fix(cron): scope the launch pre-flight to launcher CLIs, not every mode
CI caught three cron-service failures. Both are mine, from converting cron's
per-mode ladders to capability reads without checking what each ladder's scope
actually was.

**The pre-flight.** cron only ever pre-flighted `deepseek` — dsh is a profile
LAUNCHER, so "installed" is not "runnable" and a bare `dsh` can boot a profile
that cannot drive a pane. I replaced that with an unscoped
`resolveCliLaunchError(mode)`, which pre-flights EVERY mode, so a claude cron
job on a box with no claude binary now failed with "Claude CLI not found"
instead of reaching tmux-manager's own throw. Three tests assert the latter.
It is now gated on `discovery.launcherProfile !== undefined`, which is
byte-identical to the `mode === 'deepseek'` check it replaces and generalises to
the next launcher. The equivalent HTTP-route conversion was already scoped (to
`capabilities.external`, matching what that route has always pre-flighted); I
simply failed to carry the same reasoning across.

**The model.** cron's ladder was `mode !== 'shell' && mode !== 'deepseek'`, and
I read it as `capabilities.model.source === 'claude-settings-file'` — which is
the HTTP route's question, not cron's. There, every external CLI reads its model
from its own config object earlier in the chain, so only claude reaches the
global default; cron has no such config, so the same expression silently
narrowed the default model from eight modes to one. Now `!== 'none'`, which is
exactly the two entries the ladder excluded. Not caught by a test — found by
re-deriving each ladder's scope after the first failure.

Also names a fourth deliberate behaviour change in the changeset, found while
tracing these: `session.ts` carried a hand-written list of modes with no
direct-PTY fallback and OMP was missing from it, though CLAUDE.md's own text
says "all eight require tmux". `requiresMux` comes off the entry now, so an omp
session whose mux creation fails refuses instead of silently starting outside
tmux.

Verified by diffing failing tests BY NAME against an upstream/master baseline,
rather than by file as before — which is how the regression slipped through: the
three new failures landed inside a file already failing for unrelated
Windows-path reasons, and the aggregate count happened to collide.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 08:49:04 +08:00
DevvynandClaude Opus 5 4830e662f9 refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery
(search dirs, version + identity probes), the launch argv template, env
handling, the `capabilities` flags that replace per-CLI branching, and the
`overlays` that back the remote/docker pane commands. Code that used to ask
"which CLI is this?" reads the entry instead.

Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every
spawn command as a literal string, captured from the hand-written builders
before they were deleted, and `test/location-overlay-commands.test.ts` does the
same for all 20 remote and in-container pane commands.

Config can never contain shell text: an entry declares typed argv tokens,
literals are validated against a safe-word pattern at LOAD time (a bad literal
rejects the whole entry — a silently dropped `--no-approve` is not cosmetic),
and values resolve through patterns NAMED in code, so a user `clis.json` cannot
widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only
in this release.

OMP is included as a registry entry rather than a tenth hand-written builder,
so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of
`buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen
and doctor ladders all drop out.

Guard rails:

- `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id
  branching reappears outside `stock.ts`, in any of its four shapes (`===`,
  `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the
  negated forms, which is how 36 of them survived an earlier pass. Every
  allowlisted branch carries its reason.
- `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities;
  deriving one from another shipped the `until=stop`-hangs-on-shell bug.
- `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and
  `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config`
  wire field is separate, bridged only by `legacyConfigAliases`. Getting
  `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass
  clamp's only handle on a CLI's privilege switch, and a wrong name clamps
  nothing with no error and no failing test — so `schema.ts` rejects an entry
  naming a param it never declared.
- Registry data resolves AT CALL TIME (`sessionModeSchema()`,
  `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs`
  thunks). A module-level const freezes at first import, so a CLI enabled while
  the server ran moved the run menu but not that surface.
- Six fields are annotated DECLARED-FOR-LATER and read by nothing
  (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/
  `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed
  rather than measured. A test pins the list so it cannot quietly grow.

Three user-visible changes, all deliberate and named:

- `probeDockerCliVersion()` derives the in-container binary from the registry
  rather than assuming it equals the mode name (`antigravity` runs `agy`).
- The remote CLI version probe now covers grok and deepseek, which the
  hardcoded map it replaces omitted while its own comment said the rule was
  "every mode except shell".
- `codeman doctor`'s CLI rows are generated from the entries, so Claude's
  install hint is the install command rather than a docs URL, five CLIs gain
  hints they never had, and the row order follows the catalog.

Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars
(matching the `cliId` pattern) before its failure message quotes the value
back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading
the hand-editable `clis.json`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 08:26:45 +08:00