The spec's as-built list, the wiki's Tile Grid page and the CLAUDE.md
tile grid paragraph: the button has no native title, its hover card
says the count and what a click and a right-click do, it is the
button's aria-describedby (always present, hidden, kept current), and it
hides in the capture phase on any press, click or right-click so the
count menu never opens beside it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Found live: a click on Tiles hides the card and keeps it hidden while
the pointer rests on the button, until the pointer leaves. With the
pointer left there, tabbing away and back onto the button showed no card
either, so a keyboard user whose mouse happened to sit on the button
never got the Shift+F10 hint.
Leaving the button (blur) now ends that suppression as well: a keyboard
focus that comes back later is a new arrival. A click that opens the
grid or a right-click that opens the menu still shows nothing, since no
pointerenter or focus follows while the pointer stays.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner feedback 1: "give me the hover info to right click over the tile
button to adjust it". The only hint was the native title, which the
browser shows late, small and unstyled.
The Tiles button now has its own hover card, under it in the count
menu's panel style:
- "Tiles · N" with the remembered count (live: a pick in the menu
changes it while the card is up), "Click: open the grid" ("close the
grid" while it is open), "Right-click: choose 2, 4 or 6 tiles", and
when the count does not fit the window what opens instead ("This
window fits 4 tiles: a click opens 4"; open, what fits). Shown from
the keyboard it adds "Shift+F10: the same menu from the keyboard".
- Shows 300 ms after a pointer that hovers rests on the button, or after
a keyboard focus (:focus-visible); never for a touch pointer or a
device that cannot hover (plus a CSS @media (hover: none) backstop),
never on a hidden button, never while the count menu is open. A short
fade on opacity and transform, none under reduced motion; it takes no
pointer.
- Hides on pointer leave, blur, Escape, a scroll and a resize, and on a
press, a click or a right-click on the button (capture phase, so it is
gone before the menu or the grid opens, and it stays gone while the
pointer rests there). openTileCountMenu hides it too, and the focus
the menu's Escape puts back on the button brings no card back.
- It replaces the button's native title (two tooltips never stack); the
aria-label stays and aria-describedby points at the card, which always
exists and is kept current, its keyboard line included while hidden, so
screen readers hear the same text without a hover. Text is diffed
against the last English, as the rest of the grid chrome, so the zh-CN
translator is not fought on every refresh.
- zh-CN for every line (平铺 · N, 单击, 右键单击, Shift+F10, the fits
note); no other header button changes (its styles are .tile-hint only).
Tests: tile-grid-hint.test.ts (install and aria wiring, the delay, each
show and hide path, content per state and count, the menu rule, the
translator, the CSS); the count menu and i18n tests read the accessible
name instead of the removed title, and the i18n harvest covers the card
(F10 joins the key names that stay Latin). The vm harness gains
removeAttribute.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A global `.btn-icon-header:hover { transform: rotate(45deg) }`, meant for
the settings gear, turned every header icon button on hover, so the folder,
Tiles, Split and the rest swung their rounded hover background into a
diamond. Three buttons had already cancelled it one by one (the font-size
buttons, notifications, the sidebar toggle).
The rule is gone, and with it those three overrides. Hover motion now moves
the icon only:
- the settings gear's icon turns 45 degrees (one tooth, so it lands on the
same shape);
- the Tiles button's four squares spread apart, each toward its corner;
- the folder cross-fades to an open folder (a second drawing in its SVG,
`.icon-folder-closed` / `.icon-folder-open`);
- every other icon just takes the hover colour.
Pointer devices only (`@media (hover: hover)`, so a tap cannot leave an icon
stuck mid-motion), and the transitions are off under reduced motion.
Owner request: the Tiles and folder buttons "weirdly turn" on hover.
Checked live on a dark and a light skin (rest, mid, end frames). Pinned by
test/header-icon-hover.test.ts, mutation-checked five ways (the button
rotation back, the open drawing missing, the motion not hover-gated, a
square spreading toward the wrong corner, reduced motion keeping its
transition). Gate: 507 files, 9798 tests passed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Found live: closing the grid with a click starts the single view's
selection, which focuses its terminal when its replay lands, a few
hundred milliseconds later. A right-click on Tiles in between opened the
count menu with the keyboard in it, and the late focus then moved the
keyboard into the terminal while the menu stayed open (3 of 3 tries), so
the arrows, Enter or Escape meant for the menu went to the session's
PTY instead (an Escape arrived there as an ESC byte).
The menu now closes when the keyboard leaves it for another element, as
any menu does. A focus going nowhere (a click on a button in Safari,
which does not focus it) does not count, so a click on a count still
picks it. Live afterwards: the menu either closes as the terminal takes
the keyboard, or keeps it when the replay landed first; never open with
the keyboard elsewhere.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- docs/tile-grid-plan.md: owner decision 10 (right-click Tiles is a
2 / 4 / 6 count menu, default 6, remembered per device; the session
picker is gone; decision 8's "picker on right-click" and "exactly the
stored set" superseded) with the owner's answers on the details; three
as-built bullets (the count menu, the animation, painting first); the
Tiles-button bullet rewritten for the count; the entry points, the
capacity note, the multi-user row and the Escape invariant follow.
- docs/architecture-invariants.md: the Opening paragraph rewritten (the
count, the stored grid in its cells, the menu owning its Escape, the
paced connect and focusOnConnect, the motion rules); the z-index list
names the count menu and the closing grid's still copy.
- CLAUDE.md: the tile grid paragraph and the z-index line.
- Wiki: Tile Grid (the click, the count menu, the animation, reduced
motion) and Keyboard Shortcuts (right-click Tiles).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner answer: "paced connect: in". Measured on the checkpoint build: of
the 140 ms between the click on Tiles and its first frame, about 100 ms
was the six xterms being built inside the click, so nothing moved on
screen for that long and the entrance could not start.
openTileGrid now mounts every tile and lays the grid out as before, then
_connectTilesPaced builds one terminal per animation frame, the focused
tile's first, then reading order. The click paints its empty tiles in
about 20 ms and the entrance plays while the terminals are built. The
time until every tile has painted does not change: the load queue serves
one capture at a time, so only the focused tile's connect is on its path,
one frame later. Each tile still connects once, into its final cell (one
fit, one PTY resize).
- openTileGrid returns before the terminals exist, so a selection that
focuses a tile whose terminal is not built yet hands the keyboard over
in _connectTile (focusOnConnect), never when focus: false was asked,
and to the newly focused tile when focus moved meanwhile.
- _connectTile connects a tile once (entry.connected; a remount after
Attach resets it), so a re-form or a remount before a tile's turn is
never connected twice.
- A grid closed or opened again meanwhile stops the old run (a run
token), and asks for no further frames.
Tests: tile-grid-paced-connect.test.ts (the order, the final cells, the
keyboard, close and reopen, removed and remounted tiles); the tests that
read connect or the terminal's focus right after openTileGrid now run
the queued frames first (flushFrames in the vm harness), every
assertion kept.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner request: "when clicking on the tile button first make this
animation nicer". Recorded before: the click froze the page, six empty
tiles cut in at once, each tile's history then scrolled in visibly as
its load landed, and closing showed an empty single view for about a
quarter of a second before its own replay scrolled in.
The grid's own motion, on by default (not an entrance-animations.js
theme, which are off by default):
- Opening: each tile fades and settles in (opacity, translateY 6px,
scale .97), 180 ms, 24 ms apart in reading order: the last of six is
done at 300 ms. Its terminal stays transparent until the load queue
reports its first capture done, then fades in whole (160 ms), so no
replay scrolls by; a 15 s backstop shows it should that never come. A
tile added later enters the same way.
- Closing with the toggle (button, Ctrl+Shift+G; owner answer: only
these): the close stays synchronous, and a still copy of the tiles
(clones: no xterm, socket or listener; inert, aria-hidden, no pointer)
dims at once over the stage, holds until the single view's
selectSession has replayed its session (at most 700 ms), then fades
out. No empty single view between the two. A reopen drops a copy
still showing; a web tab hides it.
- A re-form to another count fades the old grid's copy out at once while
the new tiles enter.
- The count menu fades in (140 ms).
Every one animates opacity and transform only, so FitAddon measures the
final cell and each tile still sends one PTY resize (#464); nothing
moves, and no copy is made, under prefers-reduced-motion.
Tests (tile-grid-motion.test.ts): the stagger, the reveal and its
backstop, the same fits and connects with and without motion, the held
copy (released on settle, by its cap, removed on the last tile-leave or
its fallback), only the toggle animates, a reopen purges, the zoomed
tile alone, the re-form, reduced motion, and a CSS guard that every new
keyframe touches only opacity and transform. The vm harness gains
style.setProperty, cloneNode, isConnected and lastElementChild.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner decision 10 ("give me then the option to choose only HOW many
tiles, 2,4,6 default is 6 so the menu is easier"; asked where it lives:
"Click opens 6"). The click still opens the grid at once; right-click,
Shift+F10 or the Menu key on the button opens a small menu of three
counts, each drawn as the grid's own layout (2x1, 2x2, 3x2), the
remembered one checked. It replaces the session picker, which is gone
(method, markup hook, CSS, zh-CN entries).
- A pick is remembered per device in codeman:tile-count (default 6;
codeman:tile-grid stays ids only) and opens that many tiles. The click
and Ctrl+Shift+G then open with it, at most what the window fits.
- Which sessions: the rule the click already had (the grid this tab
last had, else an open split's two, else tab order, the active one
included and focused), trimmed from the end (the focused one kept) or
filled from tab order to the count. A remembered grid comes back with
its tiles first, in their cells, holes filled first, then tab order
(owner answer, superseding decision 8's "exactly the stored set"). A
page-load restore still brings back exactly what was stored.
- With the grid open a pick re-forms it: a count change is a shape
change under the cell model's rule (reformTileCells), the focused tile
always kept, every joining tile mounted and laid out before any of
them connects, so each fits once and sends one PTY resize.
- Ctrl/Cmd+click on a tab with the grid closed opens the count in total,
that session among them and focused (owner answer: N, not N+1).
- A count the window cannot fit is greyed out with the reason; a
remembered one stays checked, the keyboard starts on the largest that
fits. Arrows, Home/End, Enter or Space; Escape closes the menu alone
(the global handler gives it the key first, like the tab-group menu)
and puts the keyboard back on the Tiles button; Tab and a click
elsewhere close it.
- zh-CN for every new string (N 个窗格, 窗格数量, the titles, the Help
modal row); the i18n test harvests the menu now, with a session named
"6 tiles" as the user-text trap.
Tests: tile-grid-picker.test.ts becomes tile-grid-count-menu.test.ts
(the button checks kept, every picker check carried over to the menu,
plus keyboard, remembered count, re-form, Ctrl/Cmd+click and the batch
connect); the open-set, cap, restore, shortcuts, split-coexistence and
i18n expectations follow the count; the Help modal test escapes its
label (the new one has parentheses). The vm harness tracks
document.activeElement and makes SVG elements.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Tiles button's right-click becomes a count menu (owner decision 10):
how many tiles, 2, 4 or 6, default 6, remembered per device. These are
its pure parts in constants.js (window.CodemanTileGrid):
- TILE_GRID_COUNTS / TILE_GRID_COUNT_DEFAULT and sanitizeTileCount: a
remembered count is one of 2, 4, 6, anything else reads as 6.
- tileGridSetForCount: what the grid opens (or an open grid shows)
trimmed or filled to N: trimmed from the end with the session to focus
always kept, filled from the open sessions in tab order; fewer
sessions than N give fewer tiles; never past the cap.
- tileCellCols: the column count a stored cell list was laid out with
(stored cells carry no shape of their own).
- reformTileCells: a count change is a shape change: the tiles that
stay keep their cells, the cell model's rule (fitTileCells) reshapes,
and the tiles that join fill the empty cells in reading order, holes
first.
Nothing uses them yet; the menu comes next.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner feedback 1: "so the empty tab doesnt always have to be the last
one! so I can move freely around and the empty tab can also be tab nr
4 or 3". This replaces the slot refusal of aecada8c.
grid.cells (a session id or null per cell) is now the one source of
truth; grid.ids is a getter deriving the tiles in reading order, so
everything that only wants the tiled sessions (focus neighbour,
cycling, the load queue's order, the picker, closeSession) is
unchanged. The shape still comes from the tile count and the cap
counts tiles, never empty cells.
- A tile dragged onto an empty cell moves there and leaves its own
cell empty, nothing else moving (_moveTileToCell, through
_reorderTiles: no remount, reconnect or reload; only a tile whose
cell size changed fits). A tiled session's tab does the same; a tab
of a session not tiled yet joins in the cell it is dropped on. Each
slot knows its cell and reads "Drop a tab or a tile here".
- Move Tile goes to the adjacent cell: into it when empty, a swap when
a tile is there (tileCellInDirection).
- Removing a tile leaves its cell empty; adding one takes the first
empty cell. A shape change goes through fitTileCells: each tile keeps
its row and column when all fit (2x2 growing to 3x2), else the tiles
pack in reading order.
- Focus never lands on an empty cell: Alt+Shift+Arrows run over the
cells, Ctrl+Tab and Alt+[ ] over the tiles.
- codeman:tile-grid stays ids only: its ids are the cells with null for
an empty one. A reload brings the holes back when the shape is the
same (a session gone since leaves its cell empty), another shape
packs, the old packed format reads unchanged, and a followed
#session= link keeps the holes.
Docs: the spec's as-built bullet (rewritten in place), the wiki's Tile
Grid page and Keyboard Shortcuts, the invariants and CLAUDE.md.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner: an empty cell need not be the last one ("the empty tab can also
be tab nr 4 or 3"). The helpers that let the grid hold cells instead of
a packed list, with no behaviour change on their own:
- fitTileCells: the cells after a shape change. The same shape keeps
every cell, holes included; a new shape keeps each tile at its row
and column when all fit (2x2 growing to 3x2: the four tiles stay
put), else the tiles pack in reading order.
- tileInDirection takes cells: focus never lands on an empty cell.
Left and right go along the row past a hole and never leave it; up
and down take the nearest row with a tile, the same column else the
nearest (lower on a tie). For a packed list this is exactly the old
rule, short last row included.
- tileCellInDirection: the adjacent cell a Move Tile chord moves into
or swaps with.
- sanitizeTileGridState reads the stored ids as cells (null for an
empty one) and returns them as `cells` beside the packed `ids`; a
dropped id becomes a hole, never a shift. The old packed format reads
as cells with no hole.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The owner's answers on moving tiles:
- "dont move the tile": an empty slot no longer takes a tile. A slot is
always the last cell, so a move there shifted every tile after it.
Each drop target now says what it accepts (_acceptTabDrops'
`accepts`): a tile takes any session but its own, an empty slot only a
session not tiled yet. A refused drag is still held (dropEffect none,
no highlight), and dropSessionOnSlot refuses a tiled session too, its
tab included. A session not tiled yet still joins on a slot.
- A cancelled drag changes nothing, focus included (best practice): the
header focuses its tile on click, never on press, so a drag that ends
with Escape or outside leaves focus and the idle alert alone. The body
keeps press-to-focus, so focus still moves before a press reaches
xterm. The rename input stops its own clicks.
- A tiled tab dropped on the zoomed tile stays refused (confirmed).
- The Alt+Shift+Arrow focus chords skip a text field too (best
practice), as the move chords already did: shifted arrows select
there. Toggle and zoom are not text-editing keys and are unchanged.
Docs: the wiki, the spec's as-built bullet (with the owner's answers),
the invariants and CLAUDE.md say so.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The wiki's Tile Grid page gets a "Moving tiles" section and the move
chords in its keys table (with what they leave to a text field and
take from a terminal editor inside a tile); Keyboard Shortcuts lists
the chords and the header drag. The spec records moving as an owner
request in its as-built list, with the reasoning behind the default
keys. The invariants and CLAUDE.md say every move goes through
_reorderTiles (no remount, reconnect or reload; only a tile whose cell
size changed fits) and that the header drag carries its own type and
is not draggedTabId.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Four rebindable registry chords in the Tiles group move the focused
tile: it trades places with the neighbour the Alt+Shift+Arrow focus
chords pick (tileInDirection), through the same _reorderTiles path as
the header drag, and keeps the focus. Nothing at an edge.
They go through tileShortcutFor()/runTileShortcut() like the other
tile chords, so every xterm key handler swallows them while they
apply: only while the grid is open, a zoomed grid included (a no-op
there, so the keys never reach the CLI), and never in a text field
other than xterm's own textarea, where Ctrl+Shift+Arrows select by
word.
Ctrl+Shift+Arrows because every other two-modifier arrow chord is
taken: Ctrl+Alt switches workspaces (GNOME, Xfce, some Windows
graphics drivers), Ctrl+Alt+Shift moves a window to another workspace
(GNOME, Cinnamon, Xfce), Super belongs to the desktop, Alt is the
browser's back and forward, Alt+Shift focuses tiles. No browser,
GNOME, KDE, macOS or Claude Code default uses Ctrl+Shift+Arrows; it
costs a terminal editor's word selection inside a tile while the grid
is open.
The Help modal lists the chords and the header drag; the shortcut
overlay lists the registry. zh-CN entries for every new string.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner request: "give me the option to move the tiles around". A tile's
header (its free area, not the buttons or the rename input) is now a
native drag handle: dropped on another tile the two trade places, dropped
on an empty slot it moves there (the last cell; the tiles after it close
up). Escape or a drop anywhere else is the browser's own cancel and
moves nothing.
Every move goes through one path, _reorderTiles: the header drag, a
tab of a tiled session dropped on a tile or a slot, and (next commit)
the Move Tile chords. Nothing is remounted, reconnected or reloaded.
Divider sizes belong to the cells, so a moved tile takes its new
cell's size: each tile whose cell size changed fits once (one PTY
resize, #464) and every other tile is left alone, in place of the
debounced refit of every tile the swap used to schedule.
The drag reuses the tab drop targets (capture phase, stopped before
xterm), carries a type of its own and never text, and is not
draggedTabId, so neither a text field nor the tab strip takes it.
Moving is off while a tile is zoomed (draggable off, and a tab drag
of a tiled session onto the zoomed tile is refused too) and with a
single tile. The handle shows a grab cursor and says it drags in its
tooltip (zh-CN included); the dragged tile is dimmed and the target
highlighted.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The digit rule from 21ae48a5 hid the official DeepSeek ids (`deepseek-chat`,
`deepseek-reasoner` carry no digit), so a session on the official route with
the model field on showed the logo alone (its bundle row pins a provider
alone, so the config had nothing either). It also still misread a folder name
with a digit when every field before it was off.
Now the captured field is rejected when it is what the field can be when it is
NOT the model, and read otherwise:
- capabilities.modelDetect.rejectWords (registry data, single tokens, compared
ignoring case; the schema bounds them and requires a screenLine). dsh lists
every effort id its adapters offer (pi-ai THINKING_LEVELS plus the DeepSeek
adapter's off/low/high/max) and the shipped mode ids, from dsh 0.1.1-rc.2 /
dsh-TUI 0.10.0-beta.1. A mode's drawn label (`plan mode`, `full access`,
CJK) can never be one captured field.
- In the shared screen reader, for every CLI: a field equal to the session's
own working-directory basename is the folder, never the model.
Fixtures: `deepseek-chat` and `deepseek-reasoner` with the model field on are
read; every effort id, `default`, `plan mode`, and the folder name first (with
and without a digit) are not; the live qwen footer still reads `qwen3.8-27b`.
Known gaps, all off by default, are named in stock.ts: a custom mode id drawn
raw, a git branch or a one-word session title first, and the non-compact
footer layout (nothing read there; the route config applies).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- api-reference: the `config` source in the displayModel table, read again at
every pane start, attach and relaunch rather than restored.
- cli-registry: `modelDetect.configResolver` (a named, read-only, bounded
reader), the stock `deepseek-route` reader and its rules, and why the dsh
footer pattern needs a digit.
- deepseek-integration §4: Codeman reads the route for display only, the way
dsh-TUI resolves it; the catalog check it cannot see.
- architecture-invariants (tile grid), tile-grid-plan "as built" (owner
feedback 1), the wiki's model row, CLAUDE.md's source order.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A session header's tooltip (tile and split) says where a model the CLI did
not report came from; the new `config` source reads
"DeepSeek · qwen3.8-27b (from config)", with the zh-CN pattern
"(来自配置)" (the harness and model names pass through). The header's model
text itself is unchanged. Covered in the chrome tooltip test and the i18n
harvester's exercise.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
dsh-TUI draws the model as its status line's first field only while the
status bar's model field is on (the default). Switched off, the first field
is the reasoning effort (` medium · th-scratch`), else the mode, else the
cwd's basename, and the footer pattern read that as the model, which would
also outrank the route config added in the previous commit.
The captured field must now carry a digit, as a model id does (a version) and
an effort word, a mode name or most folder names do not. A model id without
one (`deepseek-chat`) is not read from the screen and the session falls back
to its route config: silent, never wrong.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
displayModel gains a `config` source, ranked below any report from the running
CLI and above the launch model: custom endpoint, then statusline or screen,
then config, then launch, then nothing. The screen still wins whenever it
names a model, since that is what the running TUI uses.
- Registry data: capabilities.modelDetect gains `configResolver`, a NAMED
reader (src/model-config-resolvers.ts), like a launcher profile; dsh names
'deepseek-route' (the reader from the previous commit). `screenLine` becomes
optional; the schema refuses a modelDetect naming nothing, an unknown
reader, or screenLines without a screenLine.
- Session: the reader runs from _withPaneLifecycle's finally, so at every pane
start, attach and relaunch, with the session's own launch config
(legacyConfigForMode) and env (its clamped overrides, then the server's), so
a per-session DSH_HOME is the home read. Async; a read that lands after a
newer one or after the session stopped is dropped; a remote or docker
session reads nothing locally. A change emits displayModelChanged
(broadcast and persist). Not restored after a restart: the next attach
reads it again, and a restored screen value outranks it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A DeepSeek session whose screen names no model (dsh-TUI's status bar model
field off, or not drawn yet) can still name the model its route config pins
(owner request). src/deepseek-route-config.ts resolves it the way dsh and
dsh-TUI 0.10.0-beta.1 do, for the session's profile (else the one the launch
boots) under the session's dsh home:
- dsh composes a profile from patch layers: the bundles, then
profiles/<profile>/cordis.patch.yml, then $DSH_HOME/cordis.patch.yml (which
outranks it). A patch's `config` replaces the dsh-tui row's whole config, a
`name` mismatch skips it, `disabled` turns the row off.
- dsh-TUI takes its route from that config only when it names BOTH provider and
model (modelRoute.js); anything less falls back to state the config does not
hold, so the answer is nothing. The bundle row pins a provider alone by
design and is not read (it resolves outside the dsh home); settings.yaml's
agent-default-model is the headless default and is never read.
- Any doubt answers nothing: a profile without dsh-TUI, a half-pinned route,
an unreadable or oversized layer, a symlink out of the dsh home, a mount that
does not answer, a file beyond a narrow strict YAML subset (no dependency
added: plain keys, single-line string scalars for the values it needs; tags,
anchors, aliases, merge keys, multi-line scalars, flow or block-scalar
config, duplicate keys, typed scalars, a second document, or a nested row
re-defining dsh-tui all answer null).
- Bounded and read-only: every path is probed with probePathKind() first, read
async with a 64 KiB cap, and must realpath inside the dsh home. Only the
model id leaves the module.
The default-profile inventory reuses the resolver's classification through a
new pure deepSeekProfileFromManifest(), read with the same bounded rules.
Not wired to sessions yet; the next commit does.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- api-reference: the new `displayModel` session field, its sources in order
(custom-endpoint, statusline, screen, launch) and that it is untrusted
display text, persisted and restored when the CLI reported it.
- cli-registry: `capabilities.modelDetect` (one capture group, the last rows
of the probe's capture, anchored on chrome only that CLI draws), the two
stock patterns (dsh-TUI, codex) and the fifth config regex.
- architecture-invariants (tile grid): the header painter, the id as data, the
untrusted model text, no writes for an unchanged session, the truncation
order, Pane A's strip and its fits through syncTerminalGeometry.
- tile-grid-plan "as built", the wiki's Tile Grid page (logo and model rows,
Split's strips), and CLAUDE.md's tile grid and CLI registry paragraphs.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
"Each session in the split view" (owner request) gets the tile header's
strip: the harness logo, the name and the model, painted by the same
_paintSessionHarness.
- Pane B's header is built from nodes now (it was innerHTML with the name
escaped) and follows renames and model changes on every tab render, like a
tile header; its close button is a tile button (26px target, 19px glyph).
- Pane A is the main terminal, which has no header of its own: while the
split is open it gets the same strip, minus the close, as the first child of
.terminal-wrap, and it names the active session. The strip takes 28px from
the main terminal, so the opening resize fits it with the strip already in
place, and closing removes the strip before giving the height back. Both go
through sendResize / syncTerminalGeometry (#464): the close no longer calls
a bare fitAddon.fit(), and a close that skips the server resize (Pane A's
session ended) still refits through syncTerminalGeometry.
- The partial-history banner, which overlays the top of .terminal-wrap,
starts below Pane A's strip while it is there.
test/split-pane-headers.test.ts drives the real split code on the grid's vm
harness: both headers, text-only names and models, refresh on a tab render,
no writes for an unchanged session, and the opening/closing fits.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The tile header is now `● [logo] name · model ..... ⋯ ⤢ ×` (owner request):
- The logo is PR #532's `run-mode-dot <cliId>` slot, so the logos, the skins
and the plain dot of an id without a logo stay single-sourced in styles.css.
The id is data (a class and a catalog lookup), never a branch; the frontend
id-branching guard now also scans constants.js, terminal-split.js and
tile-grid.js (the one existing shell branch in tile-grid.js, the attach
route, is allowlisted with its reason).
- The model is the session's displayModel, as text in a data-i18n-skip span
inside a box whose tooltip may translate. Unknown means the logo alone.
- The logo's tooltip and accessible name say "<harness> · <model>", plus where
a model the CLI did not report came from ("set at launch", "custom
endpoint"; zh-CN patterns for both, the names pass through). The model's box
is aria-hidden so a screen reader hears the model once.
- One painter, _paintSessionHarness (terminal-split.js, shared with the split
panes next), diffs against what it last wrote, never the DOM: an unchanged
session writes nothing on a tab render.
- On a narrow header the model gives way first, then the name: the name does
not shrink at all and is capped at its box, since any shrink factor takes a
subpixel from a name that fits and ellipsizes it.
The chrome and zoom tests found header parts by child position; they now look
them up by class, with every assertion kept (the rename tests had been passing
against the new logo node by position). The i18n harvester files the logo's
labels as harness and model names that must stay as they are.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A session header can only name the model a session runs if the server knows
it, so SessionState gains `displayModel: { model, source }`, resolved in a pure
module (src/session-display-model.ts), strongest first:
- custom-endpoint: a Custom Model Endpoint Profile's modelId answers the
session, whatever alias the CLI prints;
- statusline / screen: the newest report from the running CLI itself.
Claude's statusLine exporter already posts model.display_name on every
render; the status-telemetry route now records it (only for a CLI with
capabilities.statusLineTelemetry). A CLI whose registry entry declares the
new capabilities.modelDetect has its footer read off the pane capture the
idle/working probe already takes (no extra tmux call), so an in-session
/model switch is followed at the next transition;
- launch: the model the session was launched with (claude's --model or the
app-wide default, another CLI's <cli>Config.model), read where the registry
says the model param lives;
- nothing known: no field, never a placeholder.
modelDetect is registry data, measured on live panes: dsh-TUI's status line
on the row under its composer (qwen3.8-27b on the owner's route) and codex's
`<model> <effort> ·` footer on its last row. Both anchor on chrome only that
CLI draws, over the last rows of the screen only; a transcript line shaped like
the footer is never taken (fixture tests). The pattern goes through
compileVersionRegex() with exactly one capture group, checked at load time.
An unreadable or covered footer keeps the last model (unlike the watching
label: a model does not stop running when something covers its row). Model
text is untrusted: escape sequences and control characters are stripped and it
is capped at 64 characters. A change emits displayModelChanged, broadcast
(session:updated) and persisted; a restart restores a CLI-reported model until
the next report.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Run menus (toolbar dropdown, phone overview picker, Custom Endpoint rows,
model picker) marked every backend with an 8px colour dot, so telling Codex
from DeepSeek meant reading the label. Each known backend now draws its own
logo in that slot. It is CSS only: every surface already renders
`.run-mode-dot <id>`, so no markup changes.
- Brand-coloured marks (Claude, Gemini, Antigravity, DeepSeek, OMP) paint as a
background image; monochrome ones (Codex, OpenCode, Pi, Grok, plus Shell and
web URLs) are masks over the row's text colour, so they follow every skin.
- Logos are inline SVG data URIs (img-src already allows data:), from
@lobehub/icons-static-svg 1.95.1 (MIT); the OMP mark is omp.sh's own.
- Drops the non-og skin overrides that re-tinted four dots with a
`background:` shorthand, which would have wiped the logo.
- An id with no logo (a clis.json addition) keeps a dot, now in --text-dim
instead of being transparent.
- test/run-menu-cli-logos.test.ts pins that every stock agent plus shell/web
has a logo in exactly one paint group and that nothing resets the slot.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit d00229ee29)
tile-grid-load-queue carried its own FakeSocket, FakeFit and FakeTerminal,
near-copies of terminal-tile-input's. It now imports
test/mocks/terminal-tile-fakes.ts, which gains what only it used:
FakeSocket.drop(code), the terminal's scrollToLine / scrollToTop, and the
replay-pace extension (an opt-in `holdParse` that keeps write callbacks
from running, as on a disposed xterm, and empty writes left out of
`writes`, since the replay queues one only to hear it was parsed).
One definition serves both files with no per-file switch:
terminal-tile-input passes unchanged with the extension in place, the
shared fit resizes to the default 80x24 the tile already has, and
FakeSocket.OPEN is the real value. Every assertion is unchanged.
Mutation-checked through the shared fakes: dropping destroy()'s replay
settle fails the destroy-while-parsing case, a queue that runs two loads
at once fails eleven cases, and a tile that never registers its input
socket fails six in terminal-tile-input.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
terminal-tile-input defined FakeSocket, FakeFit and FakeTerminal inline;
they move unchanged to test/mocks/terminal-tile-fakes.ts so the grid's
load-queue test can drive a real TerminalTile on the same fakes instead of
its own near-copies. No assertion changed (a tile that never registers its
input socket still fails six cases).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_sseFilterSessionId() and its doc comment landed between
_updateSseSubscription()'s doc comment and that function, so two doc
blocks sat back to back and _updateSseSubscription had none. Each block is
now above its own function; the moved one's em dash became a colon.
Comments only.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
f9709c3a added a clause to removeTile's doc comment without rewrapping it,
leaving one line far past the file's width. Comment only.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Five tile-grid tests defined the same `tileEl(id)` lookup and three more
inlined it; the harness (test/mocks/tile-grid-vm.ts) now exports it and
they import it. tile-grid-open-set built a second vm context just to read
constants.js, although the harness it already imports has loaded the same
file: it reads windowStub.CodemanTileGrid instead.
No assertion changed. Mutation-checked: tiles without their
data-session-id fail 34 tests across the seven files, and a broken
tileGridOpenSet fails the open-set cases.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The rename from SplitTerminalPane carried comments over that the PR 1
changes made untrue:
- constants.js buildSplitPickerSessions said TerminalTile._sendResize has
no detached check (it stands aside like the primary pane) and that Pane
B sends no `seq` (its input rides the exactly-once queue). The
conclusions stay: a detached session's window owns its PTY size, and a
session with no PTY has a pane nothing feeds or reads.
- terminal-split.js said TerminalTile has no "dims unchanged" skip (it has
_lastSentDims), and its @loadorder still ended at respawn-ui.js rather
than tile-grid.js.
- terminal-tile.js still called every pane "Pane B", said a shell load
lands in "a 50000-line xterm" (the scrollback is an option now) and
that a TUI session always gets a full replay (with boundedLoad, grid
tiles get the bounded window), and told some reasons as history ("an earlier draft", "used
to", "It LOOKED intermittent"). Those now give the reason in the
present tense, and the comments touched lose their em dashes.
writeChunked's doc comment is left as it is: the perf work rewrites
that function and owns its comment.
Comments only.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
terminal-tile-unit slices connect() out of terminal-tile.js up to
'async _loadBuffer()'. The load queue commit (fca7acd0) gave _loadBuffer a
`{ refresh }` parameter, so that anchor stopped matching, indexOf returned
-1 and the slice ran to the end of the file: every check in the test
passed against code outside connect(). The anchor is now
'async _loadBuffer(' and the test asserts both anchors resolve, so a rename
fails it instead of widening it (checked by putting the old anchor back).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_applyTileLayout carried a loop that set every tile's zoomed class and
its ⤢ button's pressed state and label. That loop is now _syncTileZoom,
beside _syncTileSlots and _syncTileDividers, which _applyTileLayout calls
the same way. Same order of writes, same last-English-label compare.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
openTilePicker set `position: fixed` inline, which .tile-picker-menu
already declares; the inline copy (carried over from the split picker,
whose menu has the same rule) is gone. The top/right offsets under the
Tiles button stay inline, since they are measured.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
.tile-attach-btn and .tile-picker-open repeated the same seven
declarations and the same :disabled opacity; they now share one rule, and
each keeps only what differs (the picker's narrower padding, the
in-flight attach's progress cursor). .tile.focused read --accent-color, an
alias of --accent, while every other grid rule reads --accent; it reads
--accent too.
Computed styles of 22 grid elements (headless Chromium, the default,
daylight-blue and og skins) are identical before and after.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`.tile-grid--zoomed .tile-divider` and `.tile-grid--zoomed .tile-slot` hid
elements that do not exist while a tile is zoomed: _applyTileLayout toggles
the zoomed class in the same pass that syncs zero dividers and zero slots,
and it is the only place either is created. Both rules are gone.
The dividers test asserted the CSS rule; it now asserts what the user sees
(no divider elements while zoomed, two again on restore), and
tile-grid-entry-points gains the same check for the empty slot of three
tiles in a 2x2. Both fail if the zoom stops removing them.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_tileGridCapacityNow had one caller, _tileGridLimit, and only measured the
area (the grid section, or the single view the grid would replace) for it.
The measurement now sits in _tileGridLimit, whose doc comment says what is
measured.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
computeTileLayout and tileGridCapacity took minTileW / minTileH, defaulted
to TILE_MIN_W / TILE_MIN_H, and no caller or test ever passed them. The
parameters are gone and both read the constants; the spec's signature line
says so.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
nextSession and prevSession each carried the same grid branch (cycleTile
from the active session, a human selection, skip the tab walk). It is now
_cycleTileFocus(delta) in tile-grid.js, next to the other focus moves;
both call it optionally, so a page or harness without tile-grid.js walks
the tabs as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Both put a new set of sessions on the grid in place of the open one: choose
the focus (the session in focus if the set holds it, else the first), close
an open grid forgotten, drop activeSessionId so re-parking snapshots nothing,
then open. They now call one _replaceTileGrid(ids), which carries the
reason for the order once.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the grid already open, openGroupAsTiles closed it (which sets
activeSessionId to null, so re-parking does not snapshot the parked
terminal) and only then chose the focus, so the active session was never
"in the group" and the group's first session always took focus. The
picker's Open chooses before closing. Now both do: the focused session
keeps focus when the group holds it, otherwise the group's first session
gets it.
tile-grid-entry-points covers both cases, with the grid open and closed;
the open-grid case failed before this change.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
closeTileGrid, removeTile, dropSessionOnTile and _remountTile each spelled
out `grid.queue?.drop(tile); tile.destroy()`. They now call
_destroyTerminalTile(tile), which keeps that order (the waiting loads are
resolved and the loading state cleared before the tile goes) and says why
once. The tile's element stays the caller's to remove.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
removeTile took `auto` for the neighbour it focuses, but no caller ever
passed anything but true: a tile leaving is never a human picking its
neighbour. The option is gone (the refocus passes `auto: true` itself, and
the doc comment says so), and the four call sites that spelled out the
defaults (the header's ×, remove-tile, a stopped socket, a popped-out
session) are plain removeTile(id). `refocus: false` callers are unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The grid holds at most TILE_GRID_MAX (6) tiles, but the file header of
tile-grid.js, the comment over its section in index.html and the grid
block in styles.css still said 1 to 9. The tile header descriptions in
_buildTileHeader, .tile-header and the chrome test listed `⋯ ×` without
the zoom button, and the glyph-size rule still spoke of four glyphs and
the plus that owner decision 9 removed. Comments only.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Tiles picker opens on right-click now. Three pieces only served the
left-click picker: stopPropagation on the opening event, the tick of delay
before the outside-click listener went in (so the opening click could not
close it), and the exception for clicks on the Tiles button. A right-click
fires no click event, and a left click on the button runs toggleTileGrid,
which closes the picker before the click reaches the document. The
preventDefault that keeps the browser menu away stays.
The outside-click close had no test: tile-grid-open-set now checks that a
click inside the picker leaves it open and one elsewhere closes it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
TerminalTile carried its own copy of the smart-copy branch (clean the
selection with this session's gutter, copy, clear, toast). PR 1 gave
cleanedTerminalSelection and copyTerminalSelection a `{ terminal, sessionId }`
target for exactly this, and nothing passed it. The tile now calls both with
its own terminal and session, and its copy code is gone.
Two things change for a tile, both to the primary pane's rule: a clipboard
write that fails keeps the selection (nothing was copied, so it stays for a
retry) instead of clearing it, and focus returns to the tile's xterm after
the copy, which matters when the execCommand fallback focused a temporary
textarea. Pinned in terminal-tile-input with the write failing and
succeeding, and the no-selection Ctrl+C / Ctrl+Shift+C split.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Opening the grid runs _cleanupPreviousSession() once to park the main
terminal, and for a non-shell session that serialized the terminal
(1000 lines of scrollback) into the snapshot cache and up to 256 KB of
localStorage. Closing the grid drops the main terminal's snapshot of
every tiled id (stale by then), so when the parked session is itself a
tile, which it is unless the grid opens on a set without it, that copy
was always thrown away. openTileGrid now passes skipSnapshot in exactly
that case; a parked session that stays out of the grid (Open group as
tiles from another session) keeps its snapshot as before.
Measured: the same serialize on the main terminal's buffer costs 32 to
42 ms per grid open (n=6, 35 KB) plus the localStorage write; shells
never took one, so the A/B runs (shells) show no difference. Snapshot
serializes per grid open with a non-shell session active and tiled:
1 -> 0.
Tests: the grid passes skipSnapshot only when the parked session is
tiled; the real _cleanupPreviousSession skips the serialize only when
asked; mutation-checked both ways.
Scope: PR 2 (tile-grid.js, the app.js seam).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The invariants for this performance pass: a tile's replay holds the load
queue only while xterm parses it, and destroy() settles a replay in
progress; the main terminal's resize timer refits the split's Pane B
only, leaving grid tiles to the grid's observer; the page's SSE filter
names TILE_GRID_SSE_FILTER while tiles own the terminal; grid tiles send
lines= on their full captures. Plus an "As built" note in the spec, whose
parking section still says the subscription stays [activeSessionId].
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
GET /api/sessions/:id/terminal?full=1 captured the whole tmux history
(capture-pane -S -<history limit>, 100,000 lines by default) and cut it to
`tail` only afterwards, all of it synchronous on the server's event loop.
A grid tile keeps TILE_SCROLLBACK lines plus its screen, so the rest was
captured to be thrown away, once per tile on every grid open, restore and
deploy reconnect. The route now takes an optional `lines=<n>` (an integer
of at least 1, clamped to the configured history limit) and passes it as
the capture's history bound (the existing historyLimitLines, so -S -<n>);
absent or malformed, the limit itself, so every existing caller gets the
same capture as before. Only full captures read it: the visible-frame path
(a shell tile's `tail=` load) reads no history and is untouched. The
capture still ends with its RELATIVE cursor move back to the caret, still
counts as a full capture (isFullCapture: the line-deleting transforms stay
off) and still reports captureCols/captureRows.
Grid tiles (boundedLoad) send lines=<scrollback + rows> on every full
capture of theirs: a TUI load and a shell history pull. The split's Pane B
asks for everything, as before.
Measured:
- A real haiku Claude pane on tileperf (about 3k lines of history):
bounded captures (lines=50, 500, 2000, 100000) against the unbounded
one, 4 PASS 0 FAIL: each a line-aligned suffix of it, ending in the
same relative cursor move (ESC[4A CR ESC[2C), same source
(mux-full-history) and capture geometry. Capture time there 72 ms both
ways, that history being shorter than the tile's bound. As a grid tile
(it sent lines=10047) its screen matched the pane row for row, 47 of
47 at the pane's own 77x47, caret on the composer.
- Six tiles restoring with Claude-style loads (full=1&tail=1MiB forced
on shells with about 19k lines of tmux history each, above the tile's
bound; n=3+3 interleaved, load 4.2 to 7.2): capture per tile med
219 ms [194 to 294] -> 155 ms [127 to 211]; server event-loop delay in
the capture window, max med 262 -> 201 ms; all painted 4.5 -> 3.6 s.
At checkpoint 1 a 30k-line history cost 713 ms per capture (event loop
blocked up to 765 ms each); the bound caps that at the tile's size.
Tests: the route passes lines= through, clamps it, ignores every malformed
form and leaves the visible-frame capture exactly as it was; a bounded
capture keeps its rows and ends in the cursor restore; grid tiles send it
on full captures and the split's Pane B does not. Mutation-checked six ways
(lines ignored, no clamp, a lenient parse, lines on the visible path, the
tile sending none, Pane B sending it). Documented in docs/api-reference.md
(/api/v1 is public).
Scope: PR 2 (the grid's loads; server route plus terminal-tile.js).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the grid open the page's SSE filter still named the focused tile's
session, so the server streamed that session's output over SSE as well.
The main terminal is parked (its socket closed, _wsReady false), so every
frame was JSON.parsed and then dropped by the park guard; the tile has the
same output over its own socket. The filter now names TILE_GRID_SSE_FILTER
(constants.js, a fixed id no session takes) while tiles own the terminal,
at both places that set it: the live re-subscribe every tile focus runs
(_updateSseSubscription) and the connect URL an SSE reconnect rebuilds
(connectSSE), through one helper, _sseFilterSessionId(). Leaving the grid
gives the filter back to the session shown (selectSession re-subscribes).
Why it is safe, server side: the filter is read in exactly one place,
SseStreamManager.flushSessionTerminalBatch. The connect route parses it,
POST /api/events/subscribe replaces it (updateClientFilter); broadcast(),
the multi-user ownership check (canDeliver), the heartbeat, the order and
tab-layout frames and the shutdown notice never read it, and no push,
viewing or acknowledgement logic does. The only page consumer of
session:terminal is _onSSETerminal -> _onSessionTerminal, a no-op while
tiles own the terminal.
Measured at checkpoint 1 (6 tiles, focused tile a printing shell): 16 to
18 frames/s, 2.2 to 2.4 KB/s parsed and dropped -> 0.
Live, this code (6 tiles, shells printing), SSE terminal frames per 5 s:
0 with the grid open; 0 after an SSE reconnect with the grid open (connect
URL sessions=tile-grid); 86 in the single view after closing the grid and
86 after a reload into it (that connect URL names no session, as before;
selectSession's re-subscribe names the shown one). Just before this commit:
18 frames/s, 2.4 KB/s. A tile focus runs no connectSSE and no handleInit;
it posts the grid id.
Tests: the page subscribes with the grid id on open and on every tile
focus, gives the session back on close and on a reload into the single
view, and connectSSE asks the same helper. Server, live, multi-user: the
id is taken on the connect query and on a re-subscribe, and then withholds
terminal output while session:updated and hook events still reach their
owner (and only their owner). Mutation-checked five ways (helper ignoring
the grid, connectSSE on the raw id, the server dropping non-UUID ids on
subscribe and on connect, the filter gating every event).
Scope: PR 2 (constants.js, the app.js SSE seam).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_selectTiledSession added a once animationend listener to the focused
tile's tab on every focus. On every skin but OG the glow is `animation:
none`, so animationend never fires: the listeners piled up on the tab (and
the class stayed). It now glows a tab only when it is not glowing already,
so a tab holds at most one; on OG the animation ends, the class goes and
the next focus glows again.
Measured (50 tile focus changes, daylight-blue): animationend listeners on
the tabs 0 -> 49 before, 0 -> 6 (one per tab) after. Over the leak run's
20 grid open/close cycles the page's listener count grew 1003 -> 1042;
this is the part CDP could attribute.
Live, this code: 50 tile focus changes leave 6 animationend listeners on
the tabs, one per tab.
The single view's copy of the same block (app.js selectSession) has the
same leak on every non-OG skin; it is a separate, pre-existing copy and is
left alone here.
Test: ten focus changes leave one listener; after animationend the next
focus glows again; mutation-checked.
Scope: PR 2 (tile-grid.js).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
xterm's DOM renderer rewrites its rows (`.xterm-rows > div`) on every frame
a pane changes, and with a grid of tiles that is every record the i18n
observer gets (about 4,600 a second for six printing tiles, all of them
row rewrites). Each paid one closest() over the whole skip selector list
(the per-record skip of 21beacf7, which stays). Every row of one terminal
shares its parent, so that parent's own shouldSkip() verdict is now kept
once it says skip: same verdict, one closest() per terminal instead of one
per record. A rows container outside any skipped surface keeps the full
check, so nothing that was translated stops being translated.
Measured (6 printing tiles, 30 s profiles, n=3 interleaved A/B, load 5.8
to 9.5, equivalent class-check variant): observer 391 to 445 ms -> 47 to
52 ms per 30 s in English, 473 -> 58 ms in zh-CN; main-thread script time
-0.35 s per 30 s (-14%). Frame share unchanged within noise.
Live, this code (6 printing tiles, 30 s profiles, interleaved against the
file at 00440c02, load 4.4 to 6.6): observer 417 to 430 ms -> 56 to 58 ms
per 30 s; script time 2.63 to 2.79 s -> 2.24 s in the undisturbed window.
The other window of this code was disturbed by CPU contention on the box
(every rendering cost 3 to 4 times higher, xterm's own included, 351 frames
in 30 s) and is not counted; its observer time was 56 ms all the same.
Test: 60 row rewrites in a terminal cost one closest(), the rows stay
untranslated, and a stray `.xterm-rows` outside any skip surface is still
translated (the verdict closest() gives); mutation-checked both ways.
Scope: PR 2 (i18n, on top of 21beacf7).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A window resize reached every grid tile twice: the grid's own
ResizeObserver refits them (tile-grid.js _scheduleTileGridRefit, 150 ms
trailing), then the main terminal's trailing resize timer (terminal-ui.js
throttledResize, 300 ms) ran _forEachTile(fit) over them again. The second
pass re-measured six panes and sent nothing (_lastSentDims dedupes the PTY
side). The timer now refits the split's Pane B only ({ grid: false }); grid
tiles exist only while the grid owns the terminal, and then its observer
already covers them.
Measured (6 tiles, 20-step window resize and back, headless, tileperf):
fit() 12 -> 6 per resize burst; PTY resizes 6 -> 6; browser layouts
unchanged within noise (85/75 -> 92/71), so this removes wasted calls only.
Live, this code (6 tiles, tileperf): a 20-step window resize, and the resize
back, each ran fit() 6 times and sent 6 PTY resizes (12 and 6 before).
Test: the timer's one _forEachTile call passes { grid: false } (it lives
inside initTerminal, so read from source like the #464 geometry tests);
mutation-checked.
Scope: PR 2. The line is in terminal-ui.js (a PR 1 seam), but on PR 1 alone
_forEachTile reaches only the split's Pane B: the double refit needs the grid.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A tile's replay (writeChunked) wrote its capture 32 KB per animation
frame, so a 1 MiB load took about a second of frames, and in the grid the
load queue's slot was held across all of it: tile N+1's capture waited for
tile N's last frame. xterm 6 already parses its write queue in 12 ms
slices and yields between them, so the slices now all go in at once (up
to a 1 MiB window, since xterm's queue throws past 50 MB and Pane B's
unbounded full=1 capture can reach the server's 32 MB) and the replay
resolves on the callback of an empty write queued behind them, i.e. once
xterm has parsed the last slice. The single-flight flag is still held for
the whole replay. A disposed xterm never runs that callback, so destroy()
now settles a replay in progress: a removed tile can no longer hold its
flag or the grid's one load queue. Queued up front, the capture also stays
in one piece during a refresh: live output written meanwhile lands after
it, not between two of its slices.
Measured (tileperf, 6 printing shells with 1 MiB histories, headless,
n=3 interleaved A/B against the starting file, load 8.6 to 11.8):
- grid fresh open, 6 tiles, all painted: 5.10 s -> 2.57 s (-50%);
restore after reload: 6.49 s -> 3.86 s (-41%); per-tile replay
669 to 734 ms -> 298 to 321 ms (median).
- Same work in half the time: frames over 20 ms 26% -> 40% of the
(shorter) load window, about 86 -> 62 slow frames in all; longest long
task on restore 304 -> 227 ms; server event-loop delay unchanged
(max 111 to 122 -> 122 to 134 ms, one capture in flight throughout).
- Split Pane B (the other TerminalTile) with the main terminal on WebGL
and its long-task guard armed: load 1.6 to 3.8 s -> 0.8 to 1.7 s over
15 loads each; 0 long tasks of 200 ms or more either way, the guard
never tripped. With an unbounded full=1 capture (about 21k lines):
2.5 to 3.4 s -> 1.9 to 2.6 s, 0 long tasks of 200 ms or more.
- At checkpoint 1 (equivalent patch, n=3 to 6): fresh 6.4 -> 2.7 s,
restore 8.9 -> 4.2 s, TUI-style reconnect 11.2 to 11.8 -> 6.2 s.
Tests: the replay queues every slice at once and holds the flag until
xterm has parsed it; a replay larger than the window goes one window at a
time; a pane destroyed mid-parse settles at once; in the grid, a tile
destroyed while xterm still parses its replay releases the queue and the
next tile loads (fake xterm whose callbacks never run). The rAF-driven
tests now hold the parse callbacks instead. All mutation-checked (no
settle in destroy, settle before the parse, no window). Browser
split-pane-terminal: same 1 failed / 2 passed as at the starting HEAD
(the failure is in the test's own setup, before connect).
Scope: PR 1 (terminal-tile.js writeChunked and destroy(); Pane B replays
the same way). Moving it onto PR 1 needs its two call sites adapted
(PR 1 has no _runLoad yet) and leaves the tile-grid-load-queue.test.ts
hunk with PR 2.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The spec records decision 9 and an as-built entry replacing the "+ / New
session in this case" one, and marks the target picture, the header line
and the + bullet as built without it. CLAUDE.md's header list, the
invariants' z-index line (the + menu's layer) and the wiki's Tile Grid
page (the header string, its table and the cap sentence) drop the +.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner: "remove the + button from these views". The tile header is now
● name ... ⋯ ⤢ ×. Gone with it, because nothing else used them: the +
menu (openTileAddMenu, closeTileAddMenu and their hooks in closeTileGrid
and the global Escape handler), its "New session in this case" entry and
runInCaseForTiles, the .tile-add-empty rules, the four i18n entries only
the menu showed, and buildTilePickerSessions' exclude argument (only the
menu passed it).
Every other way of adding tiles stays and needed nothing from the menu:
the Tiles button and its right-click picker, Ctrl/Cmd+click on a tab,
dragging a tab onto a tile or an empty slot, "Open group as tiles", and
Run joining the open grid (_joinTileGridFromRun). The picker list, the
cap helper _tileGridLimit and the user-text skip on names are shared and
kept.
Tests: the + menu cases (picker, cap, auto-join, the zh-CN harvest) are
removed; tile-grid-chrome pins the header as exactly ⋯ ⤢ × with no add
menu or runInCaseForTiles left on the app.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner request: translate "Run SH" and the rest of the Run family, plus
the two leftovers from the last report.
Run: one pattern turns "Run <code>" (Run CC, Run SH, Run OC, Run CX ...,
and any registry CLI's shortBadge) into "运行 <code>"; the mode codes and
product names stay, and exact entries still win ("Run Shell" was already
运行 Shell, "Run OMP" 运行 OMP, "Run PI" takes the "Run Pi" entry). New
entries: "Terminal / Shell" (the run menu's shell item) and "Send Enter"
(the phone toolbar's Enter button title). The phone overview's Run button
already showed 运行 beside a mode word kept as typed.
Help modal and shortcut overlay: "Tabs", "Toggle Session Sidebar", both
"Copy Selection" rows, "Focus Tabs", and "Wheel" (滚轮, a mouse input like
Click). Key names stay English: the Help modal's Home key, and every key
the overlay renders, now carry data-i18n-skip, because "Home" is also a
dictionary word (the Home button) and showed as 主页 in the key column.
The invariants' paneExit section gains the badge's translation rule
(from the previous commit): its updates compare with the remembered
English, never the DOM.
Tests: i18n-exit-run-help covers every Run label _applyRunMode can show
(its hard-coded ones and Run <shortBadge> for every stock CLI), the
toolbar titles, the Help modal through the real translator in JSDOM (no
English outside the key column, the Home key kept while the word Home
elsewhere still translates, Wheel translated), every shortcut registry
group and label, and that the overlay's keys are skipped.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner request: translate the EXITED badge. The badge carried
data-i18n-skip on purpose, because its in-place update compared the DOM
with the English label: a translated badge would never have matched, and
every incremental tab pass would have written English back for the
translator to redo. The tab's accessible name (which carries the exit,
the badge being aria-hidden) was set unconditionally on every pass, the
same trap once it is translated.
Now the badge is left to the translator. applyPaneExitBadge remembers the
last English badge text (data-label) and accessible name
(data-aria-source), both also seeded by the full render, and compares
with those, never the DOM. The tab strip's incremental path updates tabs
in place (no row-HTML comparison), so nothing else re-renders on a
translated badge.
i18n.js: "exited" -> 已退出, patterns for "exited (N)" -> 已退出(N) and
"exited (signal N)" -> 已退出(信号 N), and for the accessible name
"<name> session, agent exited ..." -> "<name> 会话,智能体已退出 ...", the
session name passed through untranslated. The header strip, the session
sidebar and the vertical rail all host the same tab markup, so this
covers all three. English reads exactly as before.
Tests: session-pane-exit-ui pins the new markup, the remembered English
and that a translated badge and accessible name survive an unchanged
pass; i18n-exit-run-help runs every paneExitLabel form (and the
accessible name, with names that are dictionary words) through the real
translator in both languages.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Found live in zh-CN: the + menu's session items stayed titled "The grid
holds at most 6 tiles". Each item carried data-i18n-skip on the whole
button to keep the session name as typed, and the translator skips an
element's attributes along with its text. Only the name is skipped now (a
child span, as the picker does), so the title is translated.
The zh-CN coverage test classified a title inside a skipped subtree as
user text, which is how this got past it; such a label is now a failure
of its own ("no UI label sits inside a skipped subtree").
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner request: with App Settings language set to 简体中文, the grid reads
fully in Chinese. Every string the grid puts on screen gets its own
ZH_CN entry, so none reaches the generic leading-verb fallback: the Tiles
button (both states, with the right-click hint), the Tiles and Split
chips, the grid region, the Help modal's Tiles rows, the shortcut
registry's Tiles group and labels (overlay and App Settings list, and
"not bound"), "Open group as tiles", the picker, a tile's +, the header
buttons, the Attach overlay (not attached, attaching, exited, ended, the
hint), the empty slot, the dividers, the Split button while tiles are
open, and the toasts. Strings with a count, an exit code or a duration
are translateDynamic patterns: the cap texts (both wordings, with and
without ": the new session opens on its own"), "This window fits N
tile(s)", the auto-zoom hint, "The agent exited (N)" / "(signal N)", the
crash-restart confirm (the existing confirm wrapper runs it through t();
the session name passes through untranslated, in the single view too),
and the tile header tooltip ("idle 3m"), which requires the duration so
bare state words stay out of the table (they collide with other
surfaces, see mobile-overview.js).
Wording: 平铺 for the feature, 窗格 for one tile, 附加 for attach, 智能体,
案例, as the table already has them. Key names stay; Click, Right-click
(mouse actions) and Arrows in the Help modal's key column are translated.
English reads exactly as before (only additions to the table).
test/tile-grid-i18n.test.ts drives the real tile code through every
state that writes text, harvests each string and requires Chinese with
no Latin word left beyond key names and durations, and the same English
in en; plus the markup through the real translator in JSDOM, and session
and group names (also when they equal a UI word) staying untranslated.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Three refreshes compared the DOM with the English source: the tile
header's tooltip, the Attach overlay's text and the zoom button's title.
With App Settings language set to 简体中文 the i18n observer writes the
translation into the DOM, so the comparison never matched again and every
chrome refresh (each session:updated, several a second with busy tiles)
wrote the English back for the observer to translate once more. Each now
remembers the last English value on the tile entry and compares with that.
English mode behaves exactly as before.
Also: tileShortcutFor's comment still called the inert chord a default
pending the owner's answer; it is owner decision 6.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the grid closed, Ctrl/Cmd+click on a tab opens what the Tiles button
would show (decision 8) plus that session; the wiki only said it opens
the grid.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner feedback: the Header buttons group's keywords named neither Split
nor Tiles. Both are now in the group's data-search. The filter
(_filterSettings) matches each chip by its own data-search and its label,
and never by the wrapper's keywords (which would light up every header
chip for "split"), so the two chips also get their own: "tiles tile grid
side by side several sessions" and "split pane side by side two
sessions". Before, only the label words matched; "tile grid" found
nothing. Pinned by running the real filter over the real markup (JSDOM).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner feedback: the tile header's ⋯ ⤢ + × read as tiny next to the
session name. The buttons inherited the header's 12px font. They are now
26px click targets (min-width, so a wider glyph still fits) with a 16px
glyph, the same as the app header's own icon buttons (.btn-icon-header);
the thin ellipsis and cross get 19px so all four read at one visual
size. The header grows from 24 to 28px to hold them, the inline rename
input to 22px. Checked live at DSF 1 on a dark (daylight-blue) and a
light (paper-gray) skin, focused and unfocused tiles, and a zoomed tile.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Owner decision 8 ("when I hit the tiles button, open the tiles already!").
A click with the grid closed now opens it straight away, and Ctrl+Shift+G
runs the same function (toggleTileGrid), so the two cannot drift. What
opens comes from one pure helper, tileGridOpenSet (constants.js):
a. the grid this tab last had, if any of its sessions survive, opened
exactly (an open split closes and its sessions do not join);
b. else an open split's two sessions, Pane A focused;
c. else the open sessions in tab order (the picker's list: no detached
ones), up to what the grid takes here (the cap of 6, fewer when the
window fits fewer), the active session always among them and focused.
A click with the grid open still closes it.
The picker moved to right-click (oncontextmenu, browser menu suppressed).
With the grid open it is preselected with the current tiles, and Open
replaces them. Ctrl/Cmd+click on a tab with the grid closed opens the
toggle's set plus that session. The button's title, the Help modal and
the wiki say right-click chooses which sessions.
Docs: decision 8 and an as-built entry in the spec (Entry points too),
CLAUDE.md, the invariants (#tile-grid, Opening), the wiki's Tile Grid and
Keyboard Shortcuts pages.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Six was tested smooth on the owner's desktop; nine missed the headless
frame bar and is untested on real hardware. TILE_GRID_MAX (constants.js)
is now 6 and stays the one cap every limit reads; the layout table gets
its own bound, TILE_LAYOUT_MAX = 9, so the 7 to 9 layouts keep working
(unreachable) and going back to nine is that one line.
Every way in stops at the cap: opening, addTile, a tile's +, a session
Run makes, Ctrl/Cmd+click, the picker, "Open group as tiles", and a stored
grid with more ids (it comes back as its first six, focus kept only if it
survives, a dropped zoom cleared, row fractions that no longer match the
3x2 reset). The limits now go through one helper, _tileGridLimit(), whose
texts say which limit binds: "Up to 6 tiles" / "The grid holds at most 6
tiles" when it is the cap, "This window fits N" when it is the window.
Docs: decision 7 and an as-built entry in the spec, CLAUDE.md, the
invariants, the wiki's Tile Grid and Dashboard pages.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
xterm's DOM renderer replaces terminal rows every frame (the split pane's
Pane B, every tile of the grid), and the translator's MutationObserver
walked each added row and ran closest(SKIP_SELECTOR) for every text node
and element in it, only to find each one inside .xterm and skip it. A CPU
profile of six printing tiles put about 2.2 s of 40 s there (closest,
translateNode, tree walks).
Now one shouldSkip(mutation.target) per record decides it: every node a
record adds or edits sits under that target, so both translators would
return on their own closest() check anyway, and the output is identical.
Measured in headless Chromium, six tiles printing 20 lines/s each, three
interleaved 40 s pairs at the same machine load: frames over 20 ms fell
from 9.1/11.5/11.8% to 7.3/8.1/8.0%; nine tiles 9.5% to 7.2%. Pinned in
i18n-branding.test: a burst of terminal rows causes no tree walk, terminal
text stays untranslated, application DOM beside it still translates.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The spec listed it as the default applied while the owner's answer was
pending. The owner has decided: with showTileGridButton off the chord is
inert and passes through like any unbound key; on, it toggles the grid.
Recorded as decision 6 in the spec (with a line under Gating), and as an
owner decision in CLAUDE.md and the invariants.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A session Run makes while the grid is open joins as a tile right away, so
the tile connects and sends its size before Run starts the pane. The
server only records a resize for a session with no PTY and spawns the pane
at 120x40, and Run's own resize step measures the parked main terminal
(display: none, so nothing). Measured live for Shell and Claude: a 97x17
tile over a 120x40 pane, for good (#464).
The chrome refresh now remembers each tile's last-seen pid and calls
TerminalTile.paneStarted() when it appears or changes. paneStarted()
forgets the sent size and sends it; a hidden tile (a zoomed neighbour)
sends nothing and keeps it forgotten, so its next fit() sends it, which a
plain fit({ force: true }) would lose. Keyed on the sessions map, so a
handleInit after an SSE drop counts too.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- CLAUDE.md: a Tile grid paragraph beside the split-pane one (parking, the
one load queue, the selection and close rules, chords, dividers, auto-join,
the exited-agent case), tile-grid.js (7.6) in the load order, the
desktop-gated header markers and the Tiles picker in the z-index stack.
- docs/architecture-invariants.md#tile-grid: the mechanisms and the reason
behind each rule; the split section now says where a waiting grid load
differs and that every capture carries a deadline.
- docs/wiki/Tile-Grid.md: the user manual page (turning it on, the ways in, a
tile's header, keys, leaving, persistence, Split), linked from the sidebar,
The Dashboard, Keyboard Shortcuts and Settings Reference.
- The Help modal lists the tile chords (pinned in help-modal-shortcuts.test).
- docs/tile-grid-plan.md: status updated, and an "as built" list of where PR 2
went another way than the spec.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every Run path makes each session it created visible through
_ensureCreatedSessionVisible and then selects the first one, a human
selection that used to leave the grid for the single view. That helper now
hands the new session to _joinTileGridFromRun: with the grid open it joins the
next free slot, so Run's selection focuses its tile. No Attach overlay flashes
on it while Run starts its pane. Sessions created elsewhere (agents, other
devices, cron) arrive only by session:created and never join; a grid already
holding what the window fits does not take it, and a hint says the new
session opens on its own.
A tile's + adds "New session in this case": the normal Run (current run mode)
for the case the tile's session belongs to, with the toolbar's case put back
afterwards; the session it creates joins the grid like any Run from this tab.
Disabled for a session outside every case.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
showTileGridButton gets the full per-device treatment Split has: a header
chip in App Settings beside Split, its load and save lines, OFF by default
(and in the handheld defaults), a member of the displayKeys merge policy,
stripped from the settings PUT and never declared in the .strict()
SettingsUpdateSchema (sending it would 400 the whole save).
The setting also gates the Ctrl+Shift+G toggle (the applied default while the
owner's answer is pending; one line in tileShortcutFor to change): OFF, the
chord is inert and reaches the terminal like any unbound key; ON, it opens and
closes the grid where one can open. A grid that is open however it was opened
(Ctrl/Cmd+click, a dropped tab, "Open group as tiles") keeps all its chords,
the toggle that closes it included.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The grid is stored in localStorage `codeman:tile-grid` as
{ v: 1, open, ids, focused, zoomed, colFr, rowFr }: ids, focus, a zoom the
user chose (an automatic one is worked out again from the window) and the
divider fractions, never content. It is written as it changes (layout, focus,
zoom, divider drags); closing the grid keeps it remembered as open: false for
one-click return, and the last tile leaving forgets it. That stored state is
now the only "remembered" grid, so the Tiles toggle, the picker's preselection
and Ctrl/Cmd+click all bring back the grid this device last had, across
reloads. Never written or read in a solo window.
The restore runs INSIDE handleInit, in place of its single-view
selectSession(restoreId, { auto: true }), so with a stored open grid the main
terminal never loads on that page load (its first select would pull a
whole-history capture only to be parked). The stored ids are sanitized against
the session list (deleted, detached and duplicate ids dropped, anything that
is not a v1 object ignored), and the fractions and zoom go back on. A window
too narrow for the grid keeps the single view and the stored grid waits; a
#session= link on load wins and leaves the grid remembered but closed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Found live: the attach and shell routes report a refusal in the envelope of
a 200 ({success: false}), and Attach read only res.ok, so a refused attach
remounted the tile as if it had worked. It now reads the envelope.
The refusal in question: an agent that exited in a live pane (paneExit, e.g.
a shell ended with `exit 3`) still has the pane's tmux client running, so
both routes refuse to start anything ("Session already has a running
process"), and the single view has no restart for it either. Its tile now
shows the exit with a pointer to Close session instead of an Attach button
that cannot work. A session with no PTY attached, or one whose socket closed
because it exited (4009), still gets Attach.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Three more ways into the grid:
- Drag a session tab from the strip onto a tile: a session not yet tiled
replaces that tile in place (the replaced session keeps running); one
already tiled swaps places with it. A layout that is not full (3 tiles in a
2x2, 5 in a 3x2) shows its empty cells as slots, and a tab dropped on one
joins the grid there. Tiles and slots handle the drag in the capture phase
and stop it, because its payload is the session id as text and xterm's
helper textarea would type it into the PTY; a drag that is not a tab (a
file) is left alone. The dropped session takes focus (a human selection).
A sorted or grouped rail does not offer tab dragging, so neither does this.
- Ctrl/Cmd+click on a tab puts that session in the grid and focuses it,
opening the grid on what the Tiles toggle would bring back if it was
closed; on a window too narrow for the grid it stays an ordinary click.
- "Open group as tiles" in the grouped rail's group menu makes the group's
live sessions (as many as the window fits) the grid.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The grid now places every tile explicitly (grid-column / grid-row, reading
order) with a 6px divider track between columns and between rows, instead of
relying on DOM order and a gap. Track sizes are fractions (grid-template fr
values) that reset to equal whenever the column or row count changes.
Dragging a divider trades size between the two tracks either side, each kept
at the minimum tile size (the pure dragTrackFractions in constants.js, always
computed from the fractions the drag started with, so it cannot drift). The
affected tiles reflow locally at most once per animation frame, with no PTY
resize; each hears exactly one fit (one PTY resize) at pointer-up, and tiles
in other tracks hear nothing. Pointer capture keeps the drag on the divider,
the body locks the resize cursor and text selection for its duration, and
closing the grid or removing a tile mid-drag tears it down, as the split's
divider does. A zoomed grid shows no dividers.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A Tiles button beside Split in the header, opt-in through the per-device
showTileGridButton setting (read in applyHeaderVisibilitySettings; the
settings checkbox, displayKeys membership and schema exclusion follow with
persistence) and hard-gated like Split: hidden by its --hidden marker, a JS
width check with a live media listener, a @media (max-width: 1179px) backstop
and never in a solo window. While the grid is open the button closes it and
reads as pressed.
Closed, it opens a picker: a checkbox per open session in tab order (never one
popped out to its own window; one with no PTY is offered, its tile shows the
Attach overlay), names as text, preselected with the grid this tab last left,
else the active session and an open split's two. Boxes past what the window
can fit are disabled with the count shown, and Open opens the grid on the
checked sessions, focusing the active one if checked. Escape and an outside
click close it; its close method is idempotent and the global Escape handler
calls it.
A tile's + lists the open sessions not yet tiled; picking one adds it and
focuses it (a human selection). A grid that already holds what the window fits
disables the entries. "New session in this case" waits for auto-join.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A tile has no live terminal when its session has no PTY attached (pid null,
e.g. restored after a server restart), when the agent exited in a live pane
(paneExit), or when the server closed the tile's socket because the session
exited (4009, which used to leave only the "session ended" marker). Its body
now says which, with an Attach button, in an overlay laid over the terminal
so the body and its xterm keep their size.
Attach is the single view's own re-attach: POST /interactive (or /shell for a
shell) with NO body, at most one in flight per session, since the route has
no in-flight guard of its own. A tripped PTY-exit breaker goes through the
same confirm before clearBreaker: true, and nothing automatic ever sends it.
On success the tile is remounted onto the new pane (a socket stopped for good
cannot reconnect), keeping the keyboard if it had it; the overlay stays away
while the server catches up, and a failed attach says so and keeps it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
⤢ in a tile's header, or Alt+Shift+Enter (registry entry zoom-tile, applies
only while the grid is open and is swallowed before the Shift+Enter newline
gate), makes that tile fill the grid like tmux zoom. The other tiles stay
connected but hidden, so they measure nothing and send no resize; pressing it
again restores the grid and refits every tile, since the hidden ones have a
stale size. Zooming a tile that is not focused focuses it first (a human
selection).
As in tmux, moving focus to another tile restores the grid, and so does
removing the zoomed tile or adding one while a tile is zoomed by hand.
When the grid area cannot fit the tiles' minimum size, the grid zooms the
focused tile itself with a hint; that zoom follows focus and lifts once the
window fits again. A zoom the user chose is left alone.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every tile gets a fixed-height header above its body: `● name ......... ⋯ ×`.
- The dot is the six-state classifier the tab rows and both home screens
share (_sidebarRichRow), with the existing .home-sessions-dot--* classes;
hovering the header says the state and for how long ("working 3m"). A
tile whose session waits on a permission prompt or question gets a pulsing
red border (box-shadow only, never layout; still under reduced motion).
- The name is text with data-i18n-skip; a double-click renames it through
the tab rename's own write queue (Enter or leaving the field commits,
Escape cancels, an IME composition owns Enter), and an in-flight name shows
as already applied, as on the tab.
- ⋯ is the tab rail's session menu (options, new window, close session with
its confirm); × removes the tile ONLY, the session keeps running. Neither
button focuses a tile that is not focused.
- The header is fixed at 24px so nothing in it can resize the body, and with
it the xterm and its PTY (#464).
Every tab render refreshes the headers, so they follow status and name
changes. Tabs of tiled sessions carry .in-tiles (both render paths), and the
tab strip re-renders when tiles come and go.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
tile-grid-park-guards.test.ts carried its own copy of the fake DOM, written
before test/mocks/tile-grid-vm.ts existed, including the remove() that spliced
the wrong element when a child was no longer listed (fixed in the shared copy
only). It now uses the shared harness, which gains what the guards need: a
settable clock behind performance.now, the PerformanceObserver callbacks the
code under test registers, and localFit on the fake tile. Same 35 cases.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
xterm paints its viewport black, and a tile's rows rarely fill it exactly, so
every tile showed a black strip between its last row and its bottom edge. The
main terminal's container already makes the viewport transparent; tiles get
the same rule, so the gap shows the terminal background.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Shortcut registry (DEFAULT_SHORTCUTS, group Tiles, all rebindable):
- Toggle Tile Grid, Ctrl+Shift+G: opens the grid this tab last left (one step
back after a selection outside it), else an open split as two tiles, else
the active session as one tile; pressed again, back to the single view of
the focused session. xterm emits nothing for a shifted Ctrl letter; the
browser's find-previous is overridden only where the grid can open.
- Focus Tile Left/Right/Up/Down, Alt+Shift+Arrows: a human selection of the
tile in that direction.
- Remove Focused Tile, unbound: the session keeps running.
tileShortcutFor() decides whether a chord applies (the toggle wherever a grid
could open, the rest only while one is open, so outside the grid
Alt+Shift+Arrows reach the terminal untouched) and is registry-aware. The
capture handler dispatches it, and the main terminal's and every tile's xterm
key handler return false for it, for every event type and before the
Shift+Enter gate, so a chord that applies never reaches a PTY.
Coexistence with the split pane: opening the grid over an open split closes
it (no wasted resize for the pane about to park) and seeds the grid with both
of its sessions, Pane A focused. While the grid is open openSplitPicker and
openSplitPane refuse and the Split button reads as unavailable
(aria-disabled); closing the grid never reopens a split.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
selectSession gets the tile branch, right after its "already active" early
return: a tiled session is focused in its tile (_selectTiledSession: an
activeSessionId change, the shared _refreshSessionPanels and xterm.focus(),
no cleanup, replay, resize or socket of the parked main terminal). Decision 1:
a USER-initiated pick of a session that is not tiled leaves the grid for the
single view (the grid is remembered), and so does an explicit leaveTiles; an
app-driven pick (auto) never collapses it. A followed #session= link passes
leaveTiles (navigation).
App-driven paths pick a tile instead of the first sessionOrder entry:
- closeSession on the focused tile focuses the neighbouring tile (next in grid
order, else previous), captured before the await like wasActive, since the
delete broadcast may already have removed the tile; the last tile closes the
grid and falls back to the normal pick.
- A tiled session deleted elsewhere loses its tile and a neighbour takes focus
with auto (the last one lands on the welcome screen as before); a close from
this tab only drops the tile and leaves the follow-up to closeSession.
- A tiled session popped out to its own window leaves the grid.
Focus rules: pressing a tile is a human selection (pointerdown, never
preventDefault); Ctrl+Tab and Alt+[ / Alt+] cycle through the tiles; only a
human selection acknowledges an idle alert. Home leaves the grid (remembered);
killing every session closes it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
tile-grid.js (load order 7.6) adds the grid to CodemanApp: openTileGrid,
closeTileGrid, addTile, removeTile and _selectTiledSession, over a
<section class="tile-grid"> that is a SIBLING of .terminal-wrap and takes its
place under .main.tiles-active. Every tile is a TerminalTile with the grid's
one load queue, TILE_SCROLLBACK, a bounded load and its own per-device font
size (codeman-tile-font-size; Ctrl +/- sizes the tiles while the grid is open).
Layout comes from computeTileLayout; one ResizeObserver on the section refits
each tile (xterm and PTY together) on the trailing edge.
Opening parks the main terminal: _cleanupPreviousSession runs once (its
snapshot is right at that moment, and it closes the main socket), and
activeSessionId always names the focused tile's session, so the panels follow
focus. With the main socket closed, every main-terminal path that would write
the focused tile's output into the hidden xterm, fetch a capture for it,
resize it or reopen its socket now stands aside through _tilesOwnTerminal():
the SSE terminal, clear and refresh handlers, the dropped-output recovery,
the completion/error writelns, retryConnection and handleInit (both re-arm the
tiles instead; handleInit keeps live tiles and drops dead ones), sendResize,
throttledResize, the history re-pull, and the WebGL long-task observer, which
watches the whole page and must not count tile renders toward the main
terminal's sticky WebGL disable. The header connection state comes from the
tile sockets.
Closing destroys every tile, invalidates the main terminal's cached content
(snapshot, codeman-xs key, buffer cache) for every tiled id, since it predates
the grid, and replays the focused session fresh in the single view.
_focusedPane() answers with the focused tile and _forEachTile reaches every
grid tile. A tile whose socket stops for good is removed (4003, 4004, 4010) or
keeps its "session ended" marker (4009).
No entry point yet: the grid is opened from the shortcut registry in a later
commit.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
GET /api/sessions/:id/terminal runs synchronous tmux calls on the server, so
N tiles loading at once would stall every WebSocket and SSE stream back to
back (and after a deploy restart all N reopen within the same second).
TerminalTile takes the options PR 1 deferred to the grid:
- scheduleLoad(tile, kind, run): every capture the tile fetches (initial
load, reconnect refresh, server {t:'r'} refresh, shell history pull) runs
when its owner says so. Absent (the split's Pane B), a load runs at once.
- scrollback (the grid passes TILE_SCROLLBACK) and fontSize.
- boundedLoad: a TUI tile loads the bounded full=1&tail= window, never its
whole history.
TileLoadQueue (terminal-tile.js, DOM-free) is that one queue: concurrency 1,
a history pull ahead of background refreshes, then the owner's rank (the grid
ranks the focused tile first, then reading order). A destroyed tile's waiting
loads are dropped unrun, and destroy() aborts the running fetch so the queue
moves on.
Also, for Pane B as well: the load now has a deadline covering the body
(CodemanFetchDeadline), so a capture that never answers cannot hold the
single-flight flag (or the queue) forever; a refresh clears the screen at its
turn rather than when it is asked for; and a close while a load only waits in
the queue writes the disconnected marker at once.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
window.CodemanTileGrid (constants.js), the grid's pure half:
- computeTileLayout: columns x rows by tile count (1x1, 2x1, 3x1 on a grid
area at least 1800px wide else 2x2, 2x2, 3x2, 3x3), capped at 9, and
whether every cell clears the minimum tile size (480x240).
- tileGridCapacity: how many tiles a grid area can hold.
- sanitizeTileGridState: a stored grid (ids only) made safe to apply;
unknown, deleted, detached and duplicate ids are dropped, focus and zoom
must name a kept tile, track fractions must be sane.
- tileNeighbor / tileInDirection / cycleTile: which tile takes focus when one
leaves, on a directional chord, and on Ctrl+Tab or Alt+[ ].
- TILE_SCROLLBACK (10,000 lines, not the primary pane's 50,000) and the tile
font default.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The block selectSession runs in an idle callback once the terminal content is
on screen (respawn banner and countdown, action log, task panel, Ralph state,
CLI info, project insights, subagent window visibility, file browser) moves
verbatim into its own method. The tile grid's focus change needs the same
refresh without the rest of selectSession, and one copy keeps the two from
drifting. No behavior change: the stale-generation guard moves with it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Close Session was bound to Ctrl+W by default. Ctrl+W is delete-word in
every shell, readline prompt and agent CLI, so muscle memory killed the
session (its tmux pane and CLI, with no confirm) mid-sentence, and with
the split pane open it was not even the pane being typed in.
Close Session now has no default key: the capture-phase handler lets
Ctrl+W through and xterm sends ^W to whichever pane is focused. The
action stays in the registry and can be bound in App Settings ->
Shortcuts; the shortcut overlay shows it as not bound. The Help modal,
CLAUDE.md, the split and tile-grid specs and three wiki pages stop
advertising Ctrl+W as kill. Owner decision (tile-grid decision 5).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When SSE comes back after a server restart, handleInit's reconnect
branch already re-opens the primary pane's socket; it now also calls
the split pane tile's reconnectNow(), so Pane B no longer waits out its
backoff (up to 10 s between tries) after every deploy. Live: Pane B was
back 4.6 s after the server process respawned, i.e. as soon as it
listened.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CLAUDE.md, architecture-invariants#split-pane-sessions and the split-pane
spec described Pane B as having no reconnect, seq-less input and xterm's
own Ctrl+V. Updated for TerminalTile (terminal-tile.js, load order 7.4):
reconnect and stop codes, the input-socket map and which input is kept
out of the persisted queue, image paste, the geometry rules, and
_focusedPane() with its Ctrl+W exception. The tile-grid spec now records
PR 1 as built (no key handler factory; scheduleLoad and scrollback move
to PR 2 with their first user).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
With the split open, every app-level terminal action resolved against
Pane A: Ctrl+L typed into Pane B cleared Pane A's display while xterm
sent the ^L into Pane B's PTY, and Ctrl+Shift+R restored Pane A's size.
_focusedPane() now answers with the terminal focused LAST (a mic or
header click moves DOM focus but not the user's pane): Pane B claims it
from its own textarea's focus, the primary terminal's focus gives it
back, and a destroyed pane never holds it. Ctrl+L, Ctrl+Shift+R, voice
dictation and image paste act on the focused pane. Ctrl+W deliberately
still closes the active session: it kills with no confirm, so moving it
is an owner decision (docs/tile-grid-plan.md, decision 5).
Also destroys every tile after each terminal-tile-input test: a real
reconnect timer from one test opened a socket in a later one and flaked
under full-suite load.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Pane B registers the primary pane's file-path/URL link provider on
its own terminal, so a path an agent prints there opens the file
preview or log viewer for Pane B's session (it was plain text).
- Ctrl+V/Cmd+V in Pane B goes through the primary pane's paste trap,
aimed at Pane B: a pasted image uploads to Pane B's session and its
path is typed there; text keeps its bracketed-paste markers. xterm's
default handled text only.
Tests pin both the tile wiring and the targeted link provider itself
(registered on the target terminal, opening against the target's
session at click time, the primary's tap-path provider untouched).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Font size, family and weight changes refit Pane B AND tell its PTY.
They used to reflow the xterm only, leaving the CLI wrapping at the old
column count, the garbled-redraw class #464 fixed for the primary pane.
- The resize frame reports the size the xterm actually holds, with no
40x10 floor (the divider's 20% clamp leaves about 28 columns), skips
an unchanged size, and is always re-sent on a fresh socket so it
re-registers as a desktop viewer.
- The server's {t:'zc'} geometry report is handled: a different column
count is adopted, rows stay local, using the primary pane's own
reconcilePtyGeometry verdict.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A Codeman restart (every deploy) or a network blip used to leave Pane B
dead, with a marker asking the user to close and reopen the split.
TerminalTile now reconnects:
- A transient close reconnects on the primary pane's backoff ladder
(CodemanWsReconnect) plus jitter; the attempt count resets only on a
successful open. The redelivery sweep's forced close (1005) counts as
transient.
- On reopen the closed state is cleared before the buffer refresh that
closes the output gap, so no stale marker lands under a healthy pane.
- 4003/4004/4009/4010 stop the pane for good and report once through a
new onExit(code) callback; the marker says why.
- Sockets are replaced race-free: the old one is detached before a new
one opens, and every handler ignores events from a socket that is no
longer current. destroy() cancels a pending reconnect.
- reconnectNow() lets an owner skip the backoff.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
TerminalTile used to send every xterm onData chunk as a bare {t:'i'}
frame: no seq, no ACK, silently dropped while its socket was down, and
typing never acknowledged the session's idle alert. Keystrokes and
pastes now go through app._sendInputAsync over the tile's own socket,
registered in the input-socket map while open (HTTP fallback while
not), so they are ACKed, persisted until delivered, redelivered after a
drop, and the ACK clears the idle alert.
What xterm generates on its own stays out of that persisted queue: a
query reply (DA/CPR/OSC) is dropped, as the primary pane drops it, and a
focus or mouse report goes out once via _sendInputEphemeral. The tile's
socket carries the tab identity with a :tile suffix so it can never
evict the primary pane's socket.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pure move and rename, no behavior change. The split pane's second
terminal (SplitTerminalPane) moves out of terminal-split.js into its own
terminal-tile.js (load order 7.4) as TerminalTile, so the tile grid can
reuse it. terminal-split.js keeps the split orchestration (picker,
divider, auto-collapse) and constructs a TerminalTile for Pane B.
Tests follow the class: split-pane-terminal-unit becomes
terminal-tile-unit, and the Shift+Enter guard and the two browser suites
read terminal-tile.js / window.TerminalTile. The browser suites match
master (one pre-existing environmental failure in both).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Both read activeSessionId at the END of an async gap, so switching tabs
in between sent the input to the wrong session:
- An image upload inserted its paths with sendInput(), which re-reads
activeSessionId after the uploads finish. It now inserts into the
session the batch was uploaded to, through the same durable queue.
- Voice dictation read the target when the transcript arrived and again
when the send button or the compose overlay's Send was pressed. The
target is now captured in start() (via _focusedPane()), the local-echo
overlay is only used when that target is the active session, and a
target that closed meanwhile gets a toast instead of a 404.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
No behavior change. Prepares the split pane's second terminal (and later
grid tiles) to share what today only the primary terminal has:
- _inputSocketFor/_registerInputSocket/_unregisterInputSocket: the
exactly-once input queue, its ACK handling and the redelivery sweep now
deliver over any registered socket bound to a session, not only
this._ws. ACKs are routed by the receiving socket's session; silence is
judged per socket; a stale handle cannot unregister its replacement.
- registerFilePathLinkProvider, cleanedTerminalSelection,
copyTerminalSelection and _handleImagePaste take an optional target
terminal and session (defaults: the primary pane).
- _focusedPane() is the one place to ask which pane the keyboard is in
(primary only, for now); _forEachTile() replaces the _splitPane special
cases in the font, family, weight, skin and resize paths.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Plans a grid of up to nine live sessions side by side. All tiles are
equal TerminalTiles, the main terminal is parked while the grid is
open, and activeSessionId follows the focused tile. The split pane stays
and shares the tile class. PR 1 builds the seams and TerminalTile, so
the split's second pane gains reconnect, exactly-once input, links,
image paste and focus-following shortcuts. PR 2 adds the grid.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- hooks-config: a probe the bulk cap refused gets ONE bounded re-probe past the
cap (probeBeforeTouching), and whatever is still unknown is skipped. The
per-spawn hook and statusLine helpers used to fall back to an unbounded
lstat/readFile there, which on a dead workspace never settled and could take
the last threadpool workers (and hang the boot hook sweep). New test: cap
engaged, stat/lstat/readFile hanging on two more paths; both helpers return.
- describeUnknownPath()/unknownPathReason(): POST /api/sessions, quick-start and
GET /api/cases/:name now say a folder was not checked (other mounts are still
not answering) instead of blaming a healthy folder at the stall ceiling.
errorCodes unchanged.
- #535 x #516: Create in a custom folder probes the parent through the bounded
probe before realpath/stat/lstat/readdir touch it; an unknown parent is 422
OPERATION_FAILED (UNREACHABLE) within the probe timeout. New test.
- Docs: MAX_STALLED default is 2 (follows UV_THREADPOOL_SIZE), CaseInfo
.unreachable covers a refused probe, the boot sweep skips an unanswering
workspace, a CLAUDE.md gotcha for bounded probes, verbs.md documents the 422
(plugin mirror synced), api-reference documents the custom-folder 422.
- Tests: the launcher case-lookup describe is no longer nested in the Grok
block, and the cap-below-ceiling test no longer depends on an inherited
UV_THREADPOOL_SIZE / CODEMAN_PATH_PROBE_MAX_STALLED.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- createEditCoordinator's finally block rebases the edits queued during a write; one that the write's 409 made inapplicable was dropped with no toast. It is now reported once, like the main loop and adoptExternal do (found by the PR bot's re-review; regression test fails without it).
- At the 32-group server cap the row and group menus no longer offer a new group, which could only fail with an untranslated 'group limit reached'. MAX_GROUPS is exported from tab-layout-browser.js.
- CLAUDE.md names the pagehide keepalive as the one deliberate exception to 'never PUT the layout outside the coordinator'.
- The Dashboard wiki page describes tab groups in the vertical rail row.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- A cached list of repositories below a folder is re-checked against the
Docker case workspaces as they are now, so a repository linked as a Docker
workspace within the 30 s list cache is no longer inspected.
- A diff past runGit's 8 MB output bound is cut short from git's partial
output instead of failing with a 500.
- The browser test waits for its slow route handler on unroute
(unrouteAll behavior 'wait'), so a late route.continue() cannot fail the run.
- "Upstream is gone" now reads "Upstream not on remote", true for a branch
that was never pushed as well as one deleted on the remote; docs mirrored.
- The diff route checks the repository against the workspace's own cached
repository list (findWorkspaceRepo) and refreshes only that repository,
instead of a fresh status of every repository in the folder.
- CLAUDE.md: a Key Patterns entry for the git read surface and its rules.
- The enclosing repository is identified with one cached rev-parse before
any full status, so an unrelated repository above the workspace costs one
process and its failure no longer hides the repositories below.
- Wiki: the bottom-bar indicator moves out of the header-controls table.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- The doctor now judges candidates like the run mode's resolver: the PATH
hit, then each search dir, each one version-checked on its own and
skipped on a mismatch (a wrong `pi`/`grok` on the PATH no longer hides
the real one in a search dir). A search-dir candidate must be an
absolute path to an executable regular file, so a relative dir or a
file without the x bit reads as missing, as it does in the Run menu.
`isExecutableRegularFile` is exported from cli-executable-resolver.ts
and reused rather than copied.
- Every doctor probe passes killSignal: 'SIGKILL'; a --version that
ignores SIGTERM held the probe for its full runtime (15 s vs 5 s
measured with a TERM-trapping script).
- README no longer claims parity with the Run menu or nvm prefixes.
- The Diagnostics panel marks a missing optional tool with ○, a missing
required one with ✗, as the terminal doctor does.
- expandSearchDir names its twin, expandHome() in cli-resolver.ts.
- test/doctor-cli-json.test.ts is hermetic: temp HOME, a PATH of only
`which` and `node`, and a clis.json that drops the registry's absolute
search dirs, so it never runs the machine's installed agent CLIs.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Route test hygiene: each test works in its own mkdtemp folder, every
deletion goes through safeRmHomeTree, and the suite refuses to start
outside test/setup.ts's temp HOME, so a raw `npx vitest` can no longer
delete a real ~/projects or the live linked-cases registry.
- Path policy: the symlink-resolved target is also judged against the
resolved home, data dir and system roots (home reached through a link,
macOS /etc -> /private/etc); test expectations are realpath-safe.
- Refuse a target equal to or inside the caller's or the shared cases
directory, pointing at plain Create New (it would list twice, and
deleting the local copy removes files).
- The registry re-read comment no longer claims to prevent the
lost-update race; documented as narrowing it, like /api/cases/link.
- UI: the success toast names the folder the server created, the
"under ~/codeman-cases" blurb and name hint change while a custom
folder is ticked, a "/" parent previews and sends /<name> instead of
an empty path, and the new labels have zh-CN entries.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The behavioural check for the detailed-rail rename clamp (#534, #526) lives in test/inline-rename.test.ts, a browser suite the CI gate does not run. This pins the cascade from styles.css itself, from computed selector specificity and source order, so a later clamp rule cannot silently out-rank the shared unclamp again. Mutation-checked: deleting the detailed-rail twin fails exactly that case.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
fix(cases): bound path probes for linked workspaces and session creation, so an unreachable mount cannot freeze the server
# Conflicts:
# src/web/routes/case-routes.ts
feat(ui): git status indicator in the bottom bar, with a panel of uncommitted and unpushed work
# Conflicts:
# config/test-suites.ts
# docs/api-reference.md
- applyWorkspaceHooks: an "unknown" probe that is not near a stalled path
(refused by the stall cap, or an unexpected stat error) no longer reads as
"go ahead". It checks existence with pathExistsForWrite first, so a deleted
workspace is not recreated by the mkdir -p in ensureCodemanHooks.
- pastCap gets a hard ceiling, PATH_PROBE_STALL_CEILING = UV_THREADPOOL_SIZE
(default 4) minus one, so explicit requests against several dead paths can
never take the last libuv worker. The bulk cap now defaults to one below the
ceiling (2 with the default pool), leaving a slot for an explicit request.
- A stall widens to its mount only for network and FUSE filesystem types read
from /proc/self/mounts; on a local mount (a path typed under a local /home
that reaches a NAS through a symlink) it narrows to the stalled path.
- GET /api/cases/:name probes CLAUDE.md with pastCap, like the folder probe.
- Comment in config/path-probe.ts describes the mount-scoped stall.
Reopening the editor over a rename still in flight filled it from the name
the server had not replaced yet, so dismissing it (blur commits) queued the
old name behind the new one and undid the rename. The queue now records the
newest queued name per session (_inlineRenamePending, cleared with the queue
entry), and a reopened editor takes its prefix, input and "unchanged"
comparison from it. An untouched confirm sends nothing more.
A failed write only toasted while its editor was still current. The queue
reports the failure itself now, and the editor only puts its label back.
One rejected task blocked every later rename of that session until reload.
Each task now chains from a settled predecessor, the local apply after a
successful PUT is guarded, and the queue entry is cleaned up on either
outcome.
The rail and sidebar editor's 4rem floor moves from a stylesheet
`!important` into the inline min-width startInlineRename already writes per
layout (0 in the header strip, 4rem in the rail and sidebar).
Tests: the reopened-editor case now expects only "First" to be sent; new
cases cover a 500 answered after the editor is gone and a throw in
updateSubagentParentNames; the long-prefix check runs in the sidebar and
detailed sidebar too and asserts the inline floor; the header strip editor
keeps min-width 0.
Two problems with renaming a tab in the vertical rail, both easier to hit now
that the grouped rail has its own inline editor beside the session one.
Writes. A committed rename PUT its name and only applied the answer if the
same editor was still open when it came back. Reopening the editor before the
PUT answered (F2 or right-click again, or starting a group rename, which
cancels the session editor) threw the confirmed name away, so the tab kept
showing the old name until an SSE frame happened to repaint it. Two quick
renames also raced as two concurrent PUTs. Inline renames now go through a
per-session queue: one PUT at a time in the order they were made, the
confirmed name applied to app.sessions whatever happened to the editor, and
the "already that name" check made when the write runs rather than when Enter
is pressed, so confirming the name still on screen over a write in flight is
a real write.
Layout. The editor (a flex row) could not shrink below the input's intrinsic
width, so a long w<n>-<case> prefix pushed the label past its row: the prefix
slid out of view in the detailed rows and the input was clipped mid-word in
the compact rail. The label now has min-width 0, the prefix gives way first
(down to 2rem, with an ellipsis), the input keeps 4rem, and in the compact
rail the row's adornments step aside while the name is edited. The detailed
rows' three-line clamp also outranked the shared unclamp rule, which is what
the existing "unclamped editor" browser test caught; it is restated there.
Header strip, sidebar and flat-rail markup are unchanged.
Tests (test/inline-rename.test.ts, browser suite): the unclamp check runs for
simple and detailed rows; a write-ordering describe covers ordering, a
reopened editor cancelled over a confirmed write, a re-sent unchanged name and
a group rename taking over; a long-prefix describe drives real rows from a
live session in simple, detailed and compact rails.
- never inspect a repository at or inside a Docker case workspace (walk-up, scan, diff route): git would run its clean filters on the host
- a branch whose upstream was deleted and pruned reports upstreamGone and falls back to commits on no remote, instead of green
- turning the setting off during a poll releases the in-flight flag
- log.showSignature=false; reword the docs: clean filters still run
- CLAUDE.md frontend load order, changeset names git-diff
- discovery reads a bounded, sorted directory listing; leading-dash paths allowed; diff 500 redacts credentials
- keyboard focus survives the poll re-render; panel stays on screen on narrow viewports; aria-expanded visible on light skins
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JrzFKEdBLwVfu6ev2ZscJS
The Git window shows each group's files under their folders, collapsed until clicked, with single-child folder chains merged and open folders surviving the refresh. App Settings → Bottom bar → 'Git status: group files by folder' (per device) switches back to the flat list.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JrzFKEdBLwVfu6ev2ZscJS
Rows open an in-panel diff (staged, not staged, untracked as additions, deleted as removals) via GET /api/sessions/:id/git-diff, with Back and Open file. The route matches repo and path against the current status, runs git diff read-only (--no-ext-diff --no-textconv), and caps output at 400 KB.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JrzFKEdBLwVfu6ev2ZscJS
- doctor probes each CLI's discovery.searchDirs when which misses and runs --version on the resolved path, so a service with a minimal PATH no longer reports installed CLIs as missing
- GET /api/doctor shares one in-flight run per category
- Diagnostics group hidden from non-admins in multi-user mode (_applyDoctorAdminGate)
- 500 uses INTERNAL_ERROR; a killed child reports 'timed out after 30 s'
- browser test blocks service workers so page.route() is reliable
- wiki: Diagnostics sentence
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JrzFKEdBLwVfu6ev2ZscJS
The bounded path probe answered "absent" both when a path did not exist and
when it simply did not answer, so a stalled linked case 404'd and the Run
button scaffolded a stray local case over it, and two stalled paths anywhere
made every unrelated path read as absent (hooks skipped, statusLine
overridden, the clone warning lost).
- probePath()/probePathKind() are tri-state: present (or directory/file),
absent (ENOENT/ENOTDIR only) and unknown (timeout, other errors, refusal).
boundedPathExists() stays as the display-only boolean.
- A stalled path takes only its own mount out of probing (deepest mount
point from /proc/self/mounts, never /; just the path itself when there is
no mount table). Unrelated paths keep probing. The process-wide cap is a
backstop that answers unknown, and a single-path user request can probe
past it ({ pastCap: true }), still bounded and still recorded as stalled.
One console.warn when a path first stalls and one when the cap engages.
- GET /api/cases/:name keeps NOT_FOUND for definite absence only. An
unreachable linked case answers with its registered path and
unreachable: true; a local one answers OPERATION_FAILED. runClaude and
runShell create a case only on errorCode NOT_FOUND. The case list keeps an
unreachable linked case, marked unreachable, instead of dropping it, and
fix-plan reports an unreadable plan as an error, not "no plan".
- applyWorkspaceHooks and the statusLine helpers skip only a workspace that
is absent or on the stalled mount; a capacity refusal no longer stops
hooks being installed elsewhere, and an unreadable settings file never
lets the exporter override a user's own statusLine.
- The clone flow's repo-settings warning is back on its synchronous check,
and stripCaseEnvKeys uses pathExistsForWrite.
- POST /api/sessions (workingDir) and POST /api/quick-start (case folder)
probe with the bounded probe instead of statSync/existsSync. Missing and
non-directory keep INVALID_INPUT; unknown is OPERATION_FAILED, and
quick-start never scaffolds over a folder that did not answer.
- PATH_PROBE_TIMEOUT_MS and MAX_STALLED_PATH_PROBES move to
src/config/path-probe.ts, overridable via CODEMAN_PATH_PROBE_TIMEOUT_MS
(default 1500) and CODEMAN_PATH_PROBE_MAX_STALLED (default 3), and are
documented in the Settings Reference.
- The probe is exported from the utils barrel and imported from there.
- Pointer drag: a press released outside the rail no longer lingers. The
release is heard on window while a press is pending, a move with the
primary button up cancels it, a new press cancels any previous drag, and
an existing Escape listener is removed before another is added, so no
orphaned capture listener can swallow Escape before the terminal.
- Inline group rename: a commit by blur leaves focus where the user put it;
Enter and Escape still return focus to the header.
- A failed layout read while edits are pending keeps the held layout and the
editor and re-reads once the write settles, so a 409 is still rebased.
Dropping unsaved work now always says so in a toast.
- "Move to <group>" quotes the group name (with a matching zh-CN pattern), so
a group named "New group" or "ungrouped" no longer reads or translates like
the fixed entries.
- The group menu glyph stays visible under (hover: none).
- The sessionStorage replay copy carries { owner, baseVersion, savedAt } and is
ignored for another owner, after 60 s, or against an older layout. A move
with no anchor carries no index, so a replay keeps the row last.
- A 400 that survives the re-read is reported as "Could not save tab groups."
- closeTabRailActionMenu() no longer removes the group menu's DOM.
- Cancelling "Delete group" returns focus to the header.
- Stale comments updated.
The grouped vertical rail can now be edited from the browser: groups are
created, renamed, reordered and deleted, and tabs are moved between them, by
menu, keyboard or pointer drag. Every edit is saved through the existing
PUT /api/tab-layout; there are no server changes.
Saving (tab-layout-browser.js, pure):
- Edits are named operations (createGroup, renameGroup, deleteGroup,
reorderGroup, moveRef) applied to the rail at once, mirroring the server
model: a moved session takes the sessions that still follow it, and a
hand-moved child is marked placement 'manual'. normalizeLayout now keeps
placement and updatedAt, since whole layouts are written back.
- createEditCoordinator keeps ONE PUT {baseVersion, layout} in flight. Edits
made in the same turn share a write; edits made while one is in flight go
out on the version it returns. A 409 replays the operations onto the
layout the server returned and retries (bounded); an operation that no
longer applies is dropped and reported. A 400 re-reads first; any other
failure reports and re-reads.
- dropOperation maps a finished drag to one operation, or null for a drop
that changes nothing.
Wiring (app.js, tab-rail-resize.js):
- The session row menu gains Move up/down, Move to <group>, Move to
Ungrouped and Move to new group in the vertical rail. Before the first
group exists it offers only "Move to new group", which is how a flat rail
becomes grouped; the header strip's menu is unchanged.
- A group header opens its menu with Shift+F10 / ContextMenu, right-click or
a hover glyph (a non-focusable aria-hidden span, so the treeitem still
holds no interactive child): Rename, New group, Move group up/down,
Delete. F2 renames inline. A web tab row's Shift+F10 opens its settings
plus the same moves.
- The menu closes on Escape (consumed before the global Escape handler, focus
back to its row or header), a pointer outside, Tab, focus leaving it, a
resize, a second open and any full re-render.
- Inline group rename shares the session rename's ownership handle, so only
the current editor releases the render guard. Enter or blur commits,
Escape cancels, IME composition keys are left to the IME, and the label
becomes a flex slot so the editor gets the full width while typing.
- Pointer drag (mouse and pen) in the grouped rail only: rows before/after a
row or into a group, a header drag reorders groups. Escape cancels; the
click that ends a drag neither selects nor toggles. The flat rail and the
header strip keep their HTML5 drag untouched.
- A tab:layoutChanged read is deferred while a write is in flight and run
once it settles; a read otherwise rebases unsaved edits. On pagehide,
unconfirmed edits go out in a keepalive PUT and into sessionStorage, and
replay after reload (a no-op when the keepalive landed).
- New strings have zh-CN entries; group names reach the DOM only as text.
Unchanged: the flat rail's markup when no group exists, the tree semantics
and single roving tab stop, sessionOrder and Alt+N.
Tests: test/tab-layout-editing.test.ts (operations, coordinator, drop
mapping, menus, rename, dismissal, SSE deferral, reload recovery, flat-rail
identity) and test/tab-layout-editing.browser.test.ts (real pointer drags,
editor paint, menu Escape), listed in BROWSER_TEST_GLOBS.
A linked case can live on a network mount. When that mount goes away, a
hard mount makes stat() wait indefinitely, and the existsSync() probes in
the case routes and the workspace hook/statusline helpers ran on the event
loop, so a single GET /api/cases (or a session create in that workspace)
froze the whole web server until the mount came back.
Add boundedPathExists() (src/utils/bounded-path-probe.ts): an async stat
that answers "absent" after 1.5 s, shares one in-flight probe per path,
remembers a timed-out path until its stat finally settles, and refuses to
start new probes while two stalled ones still hold libuv threadpool
workers. Route the read-side probes in case-routes.ts and hooks-config.ts
through it. The settings writers in hooks-config.ts use an async lstat
that treats only ENOENT as missing, so an unreachable workspace is never
mistaken for an empty one and has its settings recreated.
Optional and per-device (showGitStatus, default off). GET /api/sessions/:id/git-status is
read-only and offline (no fetch, --no-optional-locks), skips remote and Docker sessions, caps its
lists, and single-flights concurrent polls. The toolbar indicator shows uncommitted files,
commits not pushed, or a check; clicking opens a draggable panel in the style of the Files window.
Which repositories: the enclosing one when there is one; otherwise every repository up to two
levels below the working directory (capped, skipping dot-folders and node_modules, never
following symlinks), each in a collapsible section, with the indicator summing them. A repository
that merely sits above the workspace and is the home folder or higher (a dotfiles repo) is
ignored. Git-supplied text is only ever written with textContent.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JrzFKEdBLwVfu6ev2ZscJS
POST /api/cases takes an optional path; Add Case > Create New gets a 'Create in a
custom folder' option with Browse. The folder is created (or an empty one filled),
scaffolded like a normal case and registered as a linked case. System, home,
credential and Codeman folders are refused; a folder with files is Link Existing's
job; a failure after the first write undoes what this call created. Admin only in
multi-user mode, like Link Existing.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The card-row rule (line-clamp: 3) out-ranked the shared unclamp-while-renaming
override. Restate it at the same weight; the test now covers both rail layouts.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
One consolidated minor changeset with the Thanks block first; the four contributor changesets (#520, #521, #522, #523) are folded into it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge-time fixes for the three findings of the third review round of #499.
- minor: a composition on an empty prompt did not follow the prompt after
output or a resize. The post-write re-place in flushPendingWrites and the
resize observer both ran rerender() only when hasPending was true, and
hasPending deliberately excludes the composition, so the first word of a
prompt (an overlay holding only a composition) stayed on the old row over
whatever output moved there. Both sites now call rerender() unconditionally;
it already returns early when there is nothing to draw, so nothing changes
without a composition. New browser case drives the real
batchTerminalWrite/flushPendingWrites path against real xterm 6 and the
overlay built from source, moves the prompt from row 0 to row 3 and checks
the overlay follows (it fails on the old guard, overlay left on row 0), with
a parity case for pending text. The structure test pins the post-write site
through vm and the resize site, which is a closure inside initTerminal(), by
source.
- nit: removeChar() dropped the composition but did not repaint on its false
path, leaving a composition-only overlay on screen showing text the addon no
longer held. It now hides the overlay there when a composition was dropped.
Package tests cover that path and the flushed path repainting without the
tail.
- nit: the package README did not document setComposition() or the
composition getter and described hasPending as "any content". Added both to
the API tables plus a short IME composition section, reworded hasPending
(pending or flushed text, excludes the composition), and made the quick
start re-render unconditionally instead of teaching the hasPending guard.
The hasPending JSDoc says the same.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Maintainer merge-time fixes for the three PRs that landed together on the
session create / launch / persistence path.
#514 findings (bot verdict merge-with-fixes):
- minor, fixed: SessionState.model was published and persisted for every
mode, so a codex/opencode cron session reported the app-wide Claude
default it never ran on. toState() now emits it only where the new
cliTakesSessionModel() holds (registry capability model.source ===
'claude-settings-file', no CLI id branch). POST /api/sessions uses the
same helper for its non-claude refusal, so refusal and publication cannot
drift. Recovery then hands back undefined for other modes on its own.
- nit, fixed: the `model` schema admitted a leading dash (and '.', '[').
The first character must now be a letter or digit; still a subset of the
registry's model-claude pattern, so nothing accepted is refused at launch.
- nit, fixed (reject, the consistent choice): `model` with
attachRemoteSession was silently dropped. Now a 400 INVALID_INPUT, as
#514 does for non-claude CLIs and quick-start does for remote cases.
advisorModel (#530) gets the same refusal there. effort and envOverrides
keep their older silent ignore on that branch so no existing caller breaks.
#515 finding (bot verdict merge, one nit):
- nit, fixed: the types/session.ts @fileoverview described CodexConfig as
(model, resumeSessionId); it now lists reasoningEffort, bypass,
animations and renderMode too.
Audit of the merged combination (not reviewed before):
- The conflict resolutions in session.ts (toState), types/session.ts,
reboot-restore-routes.ts, server.ts (restoreMuxSessions), CLAUDE.md and
skills/codeman/reference/endpoints.md (+ plugin mirror) keep both sides
correctly; nothing was lost or doubled.
- A claude session with both `model` and `advisorModel` launches with
`--model <id>` and ONE merged `--settings` JSON (ultracode + advisorModel,
or advisorModel beside `--effort <level>`), on the tmux template
(including the resume || new variant and with the statusLine exporter)
and on the direct-PTY fallback. Both values (and effort) survive
restoreMuxSessions onto a dead pane, a reboot restore into a fresh pane,
and restartCli/dead-pane respawn via _buildRespawnPaneOptions.
- quick-start and ralph-loop take no per-session `model` (matching #514's
scope, POST /api/sessions only) and launch on the app-wide default, which
toState now persists for claude, so recovery stays consistent.
- No defect found in the combination beyond the findings above. Noted, not
changed: advisorModel is still published for any mode a caller sends it
with (launch-inert there; the UI and skill send it for claude only).
Tests: test/session-model-recovery.test.ts pins the pair through both
recovery shapes for effort ultracode/high/none, the recovery constructors'
fields, the tmux-manager builder hop, and the codex/opencode/shell
non-publication; test/advisor-model.test.ts pins the launch lines and a
real direct-PTY Session's pty.spawn argv; the route test covers flag-shaped
models, attach refusals and the published fields. Docs: SessionState.model
docstring, the reboot-restore-registry header, the golden test comment and
the CLAUDE.md model/advisor bullets.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Maintainer merge-time fixes for the grouped vertical rail (#517) and its
tree semantics (#519), from the two PR reviews.
#517 minors
- A collapsed group hid rows that need the user with no signal on its
header. The header now takes the most urgent alert among the session
rows its collapse hides, in the tab alert language (tab-alert-action
red ring, tab-alert-idle yellow ring, the existing ::before rules
extended to the header). New pure hiddenGroupAlerts() over a per-section
`hidden` list; _syncTabGroupHeaderAlerts() patches it on BOTH render
paths, since alerts change without a rebuild. The kept selection draws
its own ring and is not counted.
- Every layout read rebuilt the whole tab strip, and failed reads retried
every 5 s forever. _applyTabLayout() now rebuilds only when the
structure key changed. The key drops the layout version (bumped on
every session create/close and order PUT) and instead carries group
names and the rows each collapse hides, so a version bump that moves
nothing costs nothing and a rename still rebuilds. The load coordinator
backs off (5, 10, 20, 40 s, capped at 60 s) and stops after 4 retries;
the next SSE init or tab:layoutChanged tries again, a success resets.
- A malformed stored collapse value disabled collapse on that device for
good. A parse or shape error now reads as nothing collapsed and is
rewritten to []; ok:false stays reserved for a store that throws.
- Ctrl+Shift+{ / } still reordered across groups, where the server
re-ranks per group, sends no session:orderChanged and leaves this
client's sessionOrder and Alt+N targets diverged. The move is now a
no-op unless the neighbour is in the active session's own section
(_canSwapActiveTabWith, reading the projection's new sectionByRef, which
also covers rows a collapse hides). Within a group the swap still works
and the server agrees with it; the flat rail and the strip are
unchanged.
#517 nits
- Keyboard group toggle dropping focus: already fixed by #519's
focus-by-identity; the Enter toggle test now pins focus on the header.
- Header <button> inside role=tablist: moot, #519 made the header a
treeitem inside role=tree.
- Byte-identity test not comparing against master: skipped in the suite
(a test cannot read another revision's files portably). Checked by
hand instead: the flat strip and flat rail markup of this branch before
and after this commit are identical in all 16 cases (both orientations,
manual and activity sort, no layout and zero groups, full and
incremental paths).
- Doubled blank line in docs/architecture-invariants.md: removed.
#519 minors
- A tap on a tree header or unselected row dismissed the touch keyboard:
the roving tabindex parks those at -1, so the [tabindex] arm of
MOBILE_KEYBOARD_DISMISS_EXEMPT_SELECTOR missed them. The selector now
lists [role="treeitem"].
- The tree key handler acted on keys pressed on a focused control inside
a row (Enter on the overflow button re-selected and reloaded the active
session instead of reopening its menu). It now returns unless the key
landed on the treeitem itself.
#519 nits
- aria-posinset/setsize went stale when the activity-sorted grouped rail
re-sorted rows on the incremental path. The position pass is extracted
(_applyTabTreePositions) and re-run, with aria-selected and the header
alerts, at the end of the incremental branch while the rail is a tree.
- An expanded group with no open rows was announced as an expanded parent
owning an empty group. A group with no open rows is now a tree leaf: no
aria-expanded, no aria-owns, its rows container presentation; Left and
Right do nothing on it, and its chevron keys off the section's
collapsed class instead of aria-expanded.
Tests: tab-layout-browser (malformed storage, backoff with a bounded
drain, structure key, hidden alerts, leaf groups, sectionByRef),
tab-layout-rail (header alerts on both paths, render-on-change, backoff
without rebuilds, malformed storage, Ctrl+Shift section gate, in-row
control keys, leaf header keys, posinset after an incremental re-sort,
the dismiss selector matching tree items), and three new Chromium tests
in tab-activation.browser (Enter on a focused overflow button, the touch
keyboard staying up on tree taps, the collapsed header's red ring). Every
new test fails on the pre-fix sources. Docs: architecture-invariants
owner-tab-layouts and keyboard-dismissal sections, one clause in
CLAUDE.md's dismissal rule.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Maintainer merge-time fixes for the MCP server sync (opt-in mcpSyncEnabled, synced, default OFF).
M1, parse errors echoed config text (secrets included) into the HTTP response and Settings:
smol-toml's TomlError carries a code frame of the offending lines and V8's JSON "Unexpected
token" errors quote source. Both catch sites now go through describeMcpSyncError(): a parse
failure is reported by line/column only ("not valid TOML (line 3, column 21)", "not valid
JSON"), an errno failure by Node's own message (code, syscall, path), the module's own
messages via a McpConfigError class, anything else as "unexpected error". Tests put a secret
on the broken line (TOML, both JSON message shapes, and a write refused at the re-parse that
would have quoted a copied server's env) and assert it is absent from the result and from the
route's response body; they fail against the old code.
M2, CODEX_HOME / CLAUDE_CONFIG_DIR / XDG_CONFIG_HOME were ignored, so a sync could create a
file the CLI never reads and report success: new optional registry field
capabilities.mcpConfig.relocation { envVar, path } (registry data, no id branch; schema
reuses the env-name and no-traversal path rules). Declared for claude (CLAUDE_CONFIG_DIR,
checked in the 2.1.289 binary), codex (CODEX_HOME), opencode (XDG_CONFIG_HOME) and gemini
(GEMINI_CLI_HOME, gemini-cli paths.ts); antigravity follows $HOME only (agy 1.1.12 has no
relocation var). Resolved from the server process env at call time: absolute moves the file,
empty means unset, anything else reports the target with the new status "skipped" plus the
reason and writes nothing. Dedupe is now by resolved file. When a caller overrides `home`
without passing `env`, process.env is not consulted, and the route tests clear those vars so
a CI runner's XDG_CONFIG_HOME can never aim a write outside the temp HOME.
M3, feature undocumented: CLAUDE.md Key Patterns paragraph (opt-in, admin-only, additive
only, backups, re-parse validation, 0600 for copied secrets, names-only responses with
position-only parse errors, capabilities.mcpConfig and relocation), a Settings-Reference row
in the wiki, and docs/cli-registry.md + docs/api-reference.md updated for relocation, the
"skipped" status and the error policy.
Nits:
- N1 Preview/Sync before Save: the UI remembers the saved value on open and says "Save
settings to turn MCP sync on first" instead of calling the routes; the 403 message also
says to turn it on and save.
- N2 non-admins in multi-user mode: _applyMcpSyncAdminGate() hides the whole MCP group, called
from applyMcpSyncVisibility() and the codeman:me event like the CLI-management gate.
- N3 scope chip says "synced".
- N4 "(1 servers)" pluralised; the unsupported list only names installed CLIs (route test
pins it with a per-test installed set).
- N5 McpSyncResult / McpSyncTargetResult moved to src/types/mcp-sync.ts (barrel export); only
the route imported them, so no churn.
Verified with an isolated instance (throwaway HOME, own instance and tmux socket) and
Playwright: chip, save-first message, preview rendering and the admin gate.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge-time fixes for the webhook notification channel (ntfy, Slack, Discord, generic JSON).
Minor 1, App Settings Save silently dropped webhook edits: the modal's main Save now
persists the webhook group beside the settings PUT, the same way it already saves the
model config (saveModelConfigFromSettings), but only when the group differs from what
loadWebhook() put on screen (_webhookPending), so an untouched group never re-PUTs. A
refusal (bad URL, enabled with no URL) shows a warning toast, keeps the modal open and
scrolls to the group with the pasted URL still in the box, instead of a success toast.
Send test now saves pending edits first, so it never tests the old URL while the box
shows a new one. The row says so in one line.
Minor 2, no test for the server.ts glue: new test/webhook-push-glue.test.ts drives the
private sendPushNotifications on a real (never started) WebServer with an EMPTY push
store and webhook.json in the instance data dir, delivering through the real
egress-guarded fetch to a local receiver: a permission prompt arrives with the
host-prefixed ntfy Title and body while Web Push is never called, an immediate repeat is
deduped, "response complete" is skipped under scope attention and sent under all, and a
disabled config or a non-push event sends nothing. Verified it fails when the webhook
call is moved below the "no subscriptions" return.
Minor 3, docs: webhook.json added to CLAUDE.md State Files; a Webhooks section in
docs/wiki/Notifications-And-Approvals.md (setup, what is sent, the secret URL, public
ntfy topics, local targets allowed, dedupe, instance-wide reach in multi-user mode) plus
a table row, and a line in Settings-Reference; new section 10c in
docs/security-architecture.md for the second outbound channel through the web-tab
egress guard.
Nits:
- Orphaned JSDoc: the webhook schema moved below the push schemas, so
PushSubscribeSchema has its comment back.
- Duplicated enums: WebhookUpdateSchema uses z.enum(WEBHOOK_KINDS/WEBHOOK_SCOPES), so
the schema cannot accept a kind the store would coerce away.
- describeError classifies egress refusals with isEgressBlockedError (the
CODEMAN_EGRESS_BLOCKED code anywhere in the cause chain) instead of a message regex;
tests pin a deep cause chain and that matching words alone are not a refusal.
- Markup: the URL input uses set-input, the whitespace-only line is gone, and the switch
row hints to pick a long random topic on public ntfy.sh.
- Remove a saved URL: a "Remove URL" button (shown only while a URL is saved, with a
confirm) sends { url: "", enabled: false }.
- Types placement: WEBHOOK_KINDS/SCOPES and WebhookKind/Scope/Urgency/Config/Result/Status
moved to src/types/push.ts (the IO-side WebhookMessage/Request/Fetch stay in the module).
Browser test extended: main Save persists a pending edit, a refused URL keeps the modal
open with the URL, Send test saves a newly pasted URL first, Remove URL clears it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Minor: the Key tester "14 lines" test never pressed a key into the
tester (the previous test blurred it, so the presses landed on <body>
and the cap was never exercised). It now refocuses the field, asserts
the focus, clears the log, checks 2 presses accumulate to 6 lines, then
4 presses of another key cap the log at exactly 14 with the oldest 4
lines evicted in order, and the readonly field stays empty. Verified to
fail with the cap changed to 20.
- Nit: the split-pane invariant implied Ctrl+Enter could use the CLI's
declared newline chord. Reworded after checking the send-key route:
Ctrl+Enter is always a real 0x0a, Shift+Enter is the declared
capabilities.newline chord (0x0a unless the CLI declares another), sent
on keydown only. The same imprecision in the auto-named sessions
paragraph is corrected too.
- Nit: docs/wiki/Settings-Reference.md now lists the Key tester row in
the Terminal & Input table.
- Nit: test/shift-enter-keypress.browser.test.ts exercised a hand-copied
predicate named `shipped`. It now loads the real app from a real
WebServer and presses real keys into the handlers terminal-ui.js
(app.terminal, recording the real _sendInputAsync send path) and
terminal-split.js (a real SplitTerminalPane) attach, recording the
send-key POSTs through a fetch wrapper. It asserts no \r reaches either
send path for Shift/Ctrl+Enter, exactly one send-key per press for the
right session, and that Enter and Alt+Enter are untouched. The old
keydown-only gate stays as a labelled reproduction of xterm's keypress
behaviour on a bare Terminal. Verified to fail on both panes with the
gate narrowed back to keydown.
- Nit: the keypress trap is now written down beside the other key-gate
rules (Command palette and shortcut registry): xterm runs the custom
handler for keydown, keypress and keyup and drops only Ctrl/Alt/Meta
keypresses, so a gate on a chord that can carry Shift alone must
swallow every event type. The smart-copy keydown-only rule points at it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Stale second marker above a trailing refresh's replay: xterm parses
write() on a later tick while clear() is synchronous, so a marker stamped
in a load's finally, just before _endBufferLoad() starts the trailing
refresh, landed in the freshly cleared buffer above that refresh's replay.
_stampMarkerIfOwed() now returns early while a refresh is pending; that
refresh re-owes the marker on a closed socket and writes the one copy
below its own replay. Pinned by marker-count assertions on the two
existing trailing-refresh tests plus a new async-parse fake (writes
parsed on a later tick, clear() synchronous) for back-to-back refreshes
and a pull with a queued refresh and a close mid-pull; all four fail
without the guard. Also checked against a real @xterm/headless 6.0.0.
- Marker withheld for up to the 45 s request budget: kept the behaviour and
made the comment and the docs truthful. The pull's request phase holds no
live output, but it holds the single-flight flag, so a coalesced {t:'r'}
refresh and a close's owed marker wait for the response. Writing the
marker at once during that phase would need a separate "awaiting
response" state and, with a refresh pending, reopens the same
write-vs-clear() race as above; a Codeman restart resets the in-flight
request along with the socket, so that pull fails at once and stamps.
- Stale comments: _onSocketClosed() now says the deferral covers any load,
_writeDisconnectedMarker() points at _stampMarkerIfOwed(), and the pull's
finally comment describes the hand-off to a trailing refresh.
- Invariants doc: dropped "the initial load" from the loads a close can land
in (connect() awaits it before creating the socket), reworded the
"nested refresh stamps its own" sentence to describe the guard, and noted
what the request phase holds.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ensureStatusLineExporterScript() rewrites ~/.codeman/statusline-exporter.sh
via a temp file + rename whenever the script content changes (a fresh data
dir, or a release that changes it). The temp name was pid + Date.now(), so
claude sessions created in the same millisecond (spawn_workers, a multi-tab
Run) shared one temp path: the first rename consumed it and every other
writer failed with ENOENT on chmod or rename. createSession() treats that as
a mux failure and falls back to a direct PTY, so those sessions silently ran
outside tmux (no reattach after a server restart) while quick-start still
reported success.
Measured on a fresh isolated instance, 4 concurrent claude quick-starts:
master put 2 of 4 in tmux in both rounds; with this change 4 of 4, both
rounds. The temp suffix now comes from randomBytes, like the skill writer in
the same file and user-store.ts already do. The new test freezes Date.now()
and runs eight refreshes at once; it fails on master with the same ENOENT.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
spawn_worker builds its quick-start body itself ({caseName, mode,
parentSessionId}), so an agent driving the skill had no way to give a worker
the advisor without hand-building the call and losing the readiness ladder,
hooks vetting and trust-dialog fallback. Setting CODEMAN_WORKER_ADVISOR
(fable / opus / sonnet) now adds `advisorModel` for every claude worker that
spawn_worker or spawn_workers starts; other modes ignore it.
A refused value fails the spawn with the server's INVALID_INPUT message. A
server without advisor support drops the field silently (the schema is not
strict), so spawn_worker reads it back and says so on stderr.
The preamble changed, so CODEMAN_PREAMBLE is bumped to 1.33.4 and stale
cached copies are rewritten instead of silently ignoring the variable.
SKILL.md's heredoc and the plugin mirror are synced.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude Code's advisor tool (code.claude.com/docs/en/advisor) lets the session's
main model consult a second, stronger model at decision points: before
committing to an approach, on a recurring error, and before declaring a task
done. Codeman can now start claude sessions with one.
- `advisorModel` field on POST /api/sessions, /api/quick-start and
/api/ralph-loop/start (fable, opus, sonnet or a full model id in those
families; haiku cannot advise and is refused). Stored on the session and
persisted, so respawn, boot restore and reboot restore keep it. Remote and
docker quick-starts refuse it, as they refuse effort.
- App Settings, Models, "Advisor" segment (Default / Sonnet / Opus / Fable),
synced as `claudeAdvisorModel`. Run, resume and the Ralph wizard send it.
Default sends nothing, leaving the CLI's own /advisor choice in charge.
- Carried as the `advisorModel` key in the launch's single --settings JSON,
merged with ultracode and the statusLine exporter, never the --advisor
flag: `claude --advisor haiku` exits 1 at launch, which would leave a dead
pane on every respawn, while the settings key degrades to no advisor. A
launch without an advisor is byte-identical to before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- app.js: the shortcut dispatcher returns early for events aimed at a data-raw-keys
field, so Ctrl+W / Ctrl+L / Escape / Alt+1 / Ctrl+K pressed in the Key tester no
longer kill the session, clear the terminal or close Settings
- stock.ts: drop Codex's esc-enter (a line feed works); no stock CLI declares a chord.
The esc-enter path is tested through a clis.json override
- tests: unused port (3194), Ctrl+Enter asserts no keypress, shortcut-isolation test
(verified to fail without the guard)
- docs/comments point at capabilities.newline; set-input class, trailing whitespace
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Follow-ups from the #506 review.
The live-frame queue opened before the fetch, freezing Pane B for the
whole round trip. It now opens beside capturedAt; the request uses the
shared terminal fetch deadline and the body read a 10 s one.
A {t:'r'} refresh queued behind a pull ran its clear() after the
disconnected marker was written and wiped it, and a close during a
refresh load wrote the marker above the replay. The marker is now an
owed flag (_markerOwed) that each load settles in its own finally.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
capabilities.newline replaces choosing the Shift+Enter bytes in the send-key
route. Key tester shows the keydown/keypress/keyup a browser reports.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Opt-in (mcpSyncEnabled, default OFF; routes 403 until on). Review fixes:
- codex TOML read/validated with smol-toml: CRLF, inline tables and
command-less tables no longer yield a duplicate [mcp_servers.x]; the new
text is re-parsed before writing
- null-prototype tables and own-key checks; unsafe names ignored at every level
- servers switched off in their own CLI (codex/opencode/antigravity) are not copied
- only CLIs that are installed or already have a config file take part
- files receiving env/headers are left 0600; symlinked configs are written through
- one apply at a time (409), unique tmp files cleaned on failure, failed status
- routes set real HTTP status codes; api-reference section; format type single-sourced
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
xterm runs the custom key handler for keypress too and drops Ctrl/Alt
keypresses but not Shift-only ones, so the stray \r submitted the prompt
after the newline. Swallow every event type for Shift/Ctrl+Enter and send
only on keydown, in the primary pane and Pane B. Adds a static guard and a
real xterm + Chromium browser test.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Adds capabilities.mcpConfig to the CLI registry (Claude, Gemini, Codex,
OpenCode), an additive src/mcp-sync.ts, GET/POST /api/mcp-sync and a
Settings > Agents & CLIs control. Never edits or removes an existing
server; backs up each file it changes; reports conflicts.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The grouped vertical rail is now an ARIA tree with a tree keyboard model,
and tab rows are pinned as full-row activation targets whose controls keep
their own actions and stable hit targets.
Tree semantics (grouped vertical rail only):
- #sessionTabs becomes role=tree while grouped and returns to its shipped
role=tablist and label when grouping ends. The header strip, sidebar and
flat rail keep role=tablist / role=tab exactly as before (the flat rail's
markup is unchanged byte for byte).
- A named group's header is a level-1 treeitem with aria-expanded that
aria-owns its rows' role=group (rows are level 2). Ungrouped rows and the
row a collapsed group keeps showing are level-1 items; a collapsed header
owns nothing, and the "Ungrouped" heading is a visual divider hidden from
assistive tech. aria-level, aria-setsize and aria-posinset are set on every
item, and aria-selected follows the selection without a rebuild.
- Exactly one treeitem carries tabindex=0 (roving). Controls inside rows
leave the tab order, so Shift+F10 / ContextMenu open a row's actions
(session action menu, web tab settings).
- Up/Down walk visible items, Home/End jump, Right expands a header or enters
it, Left collapses a header or climbs from a row to its header, Enter/Space
select a row or toggle a header. With the activity sort on, the walk follows
painted order within each group; the flat list keeps its whole-list walk.
- Focus survives a full re-render by identity (a row a collapse just hid hands
focus to its header), but a render never pulls focus into the rail.
- The group header is the treeitem itself (no nested button), still toggled by
click through the same onclick and still the lineage proxy anchor.
Full-row activation:
- Clicking a row's status dot, mode chip, name or padding already selected it
upstream; that is now pinned in real Chromium for the strip, the flat rail
and the grouped rail, together with every control (gear, detach, close,
overflow, web tab gear and close) running only its own action.
- The close control now shows a pointer like its siblings instead of the
default arrow.
- Enter/Space on a focused web tab in the flat list opens it; it used to call
selectSession(undefined).
- The action controls are pinned to stay under the pointer when a row is
hovered (no reflow-on-hover moving the gear out from under a click).
New Chromium suite test/tab-activation.browser.test.ts is listed in
BROWSER_TEST_GLOBS (run with npm run test:browser).
The vertical tab rail now reads the owner's tab layout (GET /api/tab-layout)
and draws its groups as collapsible sections. This is the first frontend
consumer of the tab-layout backend and it is read-only: nothing in the
browser writes the layout yet.
- tab-layout-browser.js (new, pure, loaded before app.js): projects the
layout onto the live sessions and open web tabs, renders the grouped
markup, stores collapse per device, and sequences loads newest-wins with
a bounded retry on failure.
- app.js: loads the layout on init and on tab:layoutChanged, renders the
grouped rail from the same per-row markup the flat rail uses, falls
through to a full render whenever the grouping structure changes, and
withholds drag-reorder in the grouped rail.
- Grouping is opt-in by construction. With no layout, a failed read, a
layout without groups, or a horizontal strip, the rail renders exactly
as before (byte-identical markup).
- Grouping is a render layer only: sessionOrder, Alt+N, Ctrl+Tab and the
palette keep reading the server-projected order, and row badges keep
their Alt+N slot.
- A collapsed group still shows the active row; lineage arcs to a hidden
session anchor to its group header.
- webview-tabs.js: renderWebviewTab() extracted so a single web tab can be
placed into its group with unchanged markup.
SessionState now carries the model a session launched with, and both
recovery constructors (mux recovery and reboot restore) pass it back, so a
recovered session relaunches on the same --model rather than the account
default. A top-level `model` sent with any other CLI is refused, since
those take their model in their own config object, and an empty string
means no per-session model, as it does for modelOverride.
CLAUDE.md now describes both routes for a Claude model. The tests pin
which of `model` and `modelOverride` reaches the launch and which the
case file, and that a model opening with a dash renders as --model's value.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Both create schemas now refuse an unknown level, and a non-granted owner's
codexConfig keeps its reasoningEffort when the clamp forces bypass off.
docs/architecture-invariants.md lists the two --config values codex now
takes from codexConfig.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
codexConfig takes a `reasoningEffort`, one of the levels codex accepts,
and the session starts with `--config model_reasoning_effort=<level>`.
The registry declares one literal per level, gated on the enum, because
an argv token cannot splice a value into a literal.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
POST /api/sessions takes an optional `model`, and a Claude session
launches with `claude --model <id>`. It wins over the app-wide default
model and writes nothing to disk, unlike `modelOverride`, which stays as
it is and still writes the case's .claude/settings.local.json.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- A #session=<id> link whose session never appears (closed, a typo, or
another user's session in multi-user mode) is dropped after
URL_SESSION_WAIT_MS (30 s) with a "Session not found" toast instead of
waiting forever. One stored timer per link, cleared whenever the link is
followed, replaced by a newer link, or retired.
- goHome() and opening a web tab now retire a waiting link, so a session
that turns up later no longer takes the screen. App-made web tab opens
(frame self-recovery, the fallback after the active web tab closes) pass
auto: true and keep it, as selectSession() does.
- zh-CN translation for the new toast.
- selectSession's auto: true comment now lists the #session=<id> link.
- docs: the 30 s bound, a win.location.replace() tip that avoids piling up
history entries, and the fragment declared a stable SemVer surface in
versioning-policy.md.
- Tests: timeout drops and toasts, an early arrival is still selected, the
wait does not restart, goHome and a web tab retire it, an auto web tab
open keeps it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- A markdown preview opened by attachment id under a bare file name
(attachment cards, history drawer) no longer resolves relative refs
against the workspace root: filePreviewText carries attachmentId, and
the rebase pass turns those images into their alt text and unwraps
those links. Absolute-path and workspace previews are unchanged.
- _renderMarkdown(text, { breaks = true } = {}): the File Viewer passes
breaks: false, so a hard-wrapped paragraph renders as one paragraph;
the Response Viewer keeps a <br> per newline.
- Absolute paths linkified inside a rendered document now carry the
preview's data-session-id.
- CLAUDE.md, architecture-invariants and the Working-With-Files wiki page
now say that only an in-workspace path clicked in the terminal keeps
the tail viewer.
- Tests in test/file-preview-markdown.test.ts for all three fixes,
including an end-to-end run of the shipping app.js + marked + DOMPurify.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- terminal-split.js: move the socket's close into _onSocketClosed(), which
defers the marker while a history pull holds live output (_liveQueue);
_pullHistory() records closedBefore and its finally writes the marker
after the queue flush when the socket closed during the pull, replayed
or not, so it never lands above held frames or between replay chunks
- tests: drive the real close path for a close mid-fetch ending in a skip,
a downgrade or a failed fetch, a close during the chunked replay, and a
close with no pull running; pin the onclose wiring in the static guard;
describe the mid-fetch case on its own
- CLAUDE.md: turn the plain-text split-pane pointer into a link
- architecture-invariants.md: describe the deferred marker
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- _restoreOverlayFocus(key, modal) now leaves focus alone when something
outside the overlay already holds it (not <body>, not inside the modal).
The Session Manager's "Switch to session" and "Open folder" call
selectSession() before closeSessionManager(), and the restore was pulling
focus back from the terminal to the header button. Both close methods pass
their modal; a regression test drives that order.
- Test harness: focusHarness() routes getElementById through a local binding
instead of leaking globalThis.__els, and its modal stubs report their own
search box as contained, as the real DOM does.
- CLAUDE.md and docs/architecture-invariants.md: record that the global
Escape handler calls every close method on every Escape (capture phase),
so a close method with side effects must return early when not open.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review feedback. The global Escape handler in app.js calls both
`closeSessionManager()` and `closeCommandPalette()` on every Escape, whether or
not either overlay is open, in the capture phase. Nothing was saved in that
case, so `_restoreOverlayFocus()` fell through to `terminal.focus()` and moved
focus before the focused element's own Escape handler ran:
- split view: with focus in Pane B, keys typed after Escape went to Pane A
- any text field (File Viewer editor, search and history filters, case picker):
keys typed after Escape went into the terminal
- inline tab rename: the capture-phase focus fired the input's blur (which
commits) before its own Escape handler (which cancels), so Escape committed
the rename instead of cancelling it
Both close methods now bail out on `classList.contains('active')`.
Separately, gating the terminal fallback on `activeSessionId` alone only covered
the welcome screen. On a touch device with the keyboard down, focus sits on
`<body>`, so closing the Session Manager focused the terminal and brought the
keyboard up — `selectSession()` deliberately skips that focus, and this
overrode it. It now goes through `_shouldFocusTerminalForTabSwitch()`.
Tests: the Session Manager case's modal stub now uses the harness's
`makeClassList()` (without `contains` the new guard reads it as "not open" and
skips the restore the case is about), plus two new cases — closing either
overlay without opening it first with an active session asserts the terminal was
not focused, which is the path the global Escape chain takes and none of the
five existing cases covered, and a touch device with the keyboard down asserts
the same. Each was checked against the unguarded code: removing either guard
turns exactly its own case red.
Both the Command Palette and the Session Manager call `search.focus()` on
open, and both closed by removing the `active` class and nothing else. Hiding
a focused input does not hand focus back to anyone — the browser drops it on
`<body>` — so after Escape closed the overlay every keystroke went nowhere and
the user had to click the terminal before they could type again.
Measured in headless chromium against a real shell session, one overlay at a
time:
overlay activeElement after Esc can type afterwards
App Settings XTERM yes
Session Options XTERM yes
Token Stats XTERM yes
Monitor Panel XTERM yes
Session Manager BODY no <- fixed here
Command Palette BODY no <- fixed here
The four that worked did so because they use `FocusTrap`, whose `deactivate()`
restores focus to whatever held it before. These two never got one. Every close
path has the same hole — Escape, the close method, picking an item — so the
restore lives in the close functions rather than in the global Escape chain.
Deliberately only the save/restore half of `FocusTrap`, not the whole thing:
`FocusTrap.activate()` moves focus to the first focusable element, which in
neither overlay is the search box, so adopting it wholesale would trade "type a
filter the moment it opens" for "focus survives the close" — and the former is
the reason Cmd+K exists. The terminal fallback is gated on there being an
active session: an overlay opened from the welcome screen has no terminal to
return to, and focusing one on a phone summons the on-screen keyboard over a
screen with no input on it.
The five new cases were checked against the unfixed code first: four of them
fail without this change.
A page that keeps one Codeman window open, such as a task board, could only
show a session by sending that window to /session/<id>, which loads the whole
app again for every click. The dashboard now reads a #session=<id> fragment
when it loads and on hashchange, selects that session, and removes the
fragment with history.replaceState so the next identical link is still a
change. Re-pointing a window that already shows the dashboard changes only the
fragment, so the page stays loaded and the switch is a tab change.
A link can name a session the dashboard does not list yet, because the page
that created it may link before session:created arrives. The id waits until
that event names it, and picking another tab yourself retires it.
Following a link is an app selection (`auto: true`). The page that set the
fragment may be a script, so it must not spend the session's idle alert.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Address Ark0N's review on #506:
- The history pull's own `\x1bc` reset erased the "Pane B disconnected"
marker onclose wrote, painting a fresh, current-looking history while
onData kept silently dropping every keystroke on the dead socket — a
Codeman restart drops the socket while the tmux session (and so the HTTP
pull) survives, making this easy to hit. onclose now tracks the closure
via `_wsClosed` in addition to writing the marker (extracted into
`_writeDisconnectedMarker()`), and a replay re-stamps it in the pull's
`finally` block, after the live-frame flush, whichever order the close
and the pull land in.
- `_maybeLoadMoreHistory()` now stands aside for a detached session,
mirroring `_sendResize()`'s existing check and app.js's
`_maybeRefetchFullHistory()` — its own window already owns its PTY size
and scrollback.
- Wording: a non-shell CLI's history is out of scope for this pull, not
absent (codex and Claude's inline renderer do grow tmux history); the
alternate-screen skip only matters for a direct-PTY shell, since tmux
never surfaces the alt buffer to the browser xterm. CLAUDE.md points at
the invariants heading directly.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tmux repaints a burst of output instead of scrolling it, so a shell
pane's xterm keeps about one screen of scrollback while tmux holds every
line. The primary pane goes back for it when the wheel reaches the top;
Pane B is a separate xterm that loaded history once at connect and never
again, so after a `cat` its earlier output was unreachable.
Pane B now does the same for a shell session: wheel-up at the top of the
normal screen pulls ?full=1&tail=TERMINAL_TAIL_SIZE and holds the
reader's place across the replay. The wheel listener is capture-phase
because xterm stopPropagation()s the events it consumes.
It follows the primary pane's rules from #494 and its 1.33.2 merge-time
fixes: a window holding no more rows than the pane (which covers a
downgrade), or a pane already at its `scrollback + rows` cap, is skipped
without a rewrite. That skip backs off to 60 s when the window was
truncated or the pane is full, since each ask costs the server a
whole-history capture-pane; an untruncated window keeps the 4 s cooldown.
There is no truncation banner in Pane B, so the 'tail' relabel does not
apply.
Live frames, a {t:'c'} clear included, are held with their arrival time
while the replay runs and applied in order only if they arrived after the
capture. The fetch has a 10 s deadline since it holds live output while
it runs. The tail of _loadBuffer() becomes _endBufferLoad() so the pull
shares its single-flight bookkeeping.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review follow-up on #503. marked percent-encodes link and image destinations, and the rebase pass encoded them a second time, so a space or a CJK character in a file name made file-raw look for a file literally named my%20image.png; refs are now decoded once (a malformed escape is kept as written) and stripped of ?query along with #fragment. Root-relative refs resolve from the workspace root as on GitHub instead of falling through as Codeman URLs. Rebased links carry the preview's own session id and the response-viewer delegate prefers it, so a document opened from another session's attachment card opens its links in that workspace rather than the active tab's.
The sanitizer no longer allows name=: marked never emits it, and <img name="app"> made document.app that image, which every inline onclick="app.…()" handler resolves before the global, so one rendered README broke every viewer button until a reload. Adds the zh-CN strings for the three toolbar titles.
A cron job in "Paste (direct)" input mode wrote `<text>\r` into the pane
in one piece. Claude Code (measured on 2.1.283) takes a burst of about a
hundred characters as a paste, so the `\r` landed as a newline and the
prompt sat unsent on the composer while the run reported `prompt_sent`.
Delivery now lives in `deliverCronPrompt()`. Paste mode writes the text
raw, waits CRON_PASTE_ENTER_DELAY_MS (300 ms), sends `\r` as a separate
write down the same PTY (so it cannot overtake the text), and arms the
session's composer check through the new public
`Session.verifySubmitted()`, which re-presses Enter while the prompt is
still visibly unsent. A session with nothing to write to now fails the
run instead of reporting the prompt as sent. Typed mode is unchanged.
Verified on an isolated instance: a paste-mode job with a 104-character
prompt submitted on the first Enter and Claude answered.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A prompt posted to /api/sessions/:id/input without `useMux` was written
into the pane in one piece. Claude Code (measured on 2.1.283) takes a
`<text>\r` burst of about a hundred characters or more as a paste, so the
trailing `\r` landed as a newline in the composer and the prompt sat there
unsent while the route answered 200. A later raw `\r` did not recover it;
a tmux `send-keys Enter` did. Short prompts submitted, which is why it
looked random. The same stranding was seen with Codex and OpenCode.
A plain prompt (printable text plus exactly one trailing `\r`, detected by
`isPlainPromptInput()`) now goes through `writeViaMux` even without
`useMux`: the text is typed, Enter is pressed as its own key, and the
SubmitVerifier re-presses it while the prompt is still on the composer.
The write is awaited, since the browser's POST fallback sends frames one
at a time and a following keystroke must not overtake the Enter. Raw
frames (escape sequences, bracketed paste, a line feed, a bare `\r`) and
an explicit `useMux: false` keep the direct write.
Verified on an isolated instance: the 239- and 104-character prompts that
stranded (at +1 s, at +50 s on ultracode, and on a warm session) all
submitted on the first Enter with no `useMux`. The phone's local-echo
flow (a burst, then its `\r` as a separate write) was measured unaffected.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Folds the #490 and #492 contributor changesets (the latter said minor) into one patch changeset with the Thanks section, one paragraph per change and the fixes applied while landing.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- _logScrollRouting() reports cliMouseTracking, the gate's new input, in both
the de-dup signature and the console line (xterm's own mouseTracking stays
'none' for Claude, so it gave no reason for a no).
- Restore two guard tests the new gate made vacuous: the local-scrollback
opt-out footgun test and the codex/gemini "no version rescues it" fixtures
now set cliMouseTracking: true, so removing the opt-out or re-adding codex to
the gate fails again.
- Update the comments and architecture-invariants lines that still described
the version-only rule (wheel handler header, gate doc, the false paths of
_maybePageCliTranscript, "holds a tracking mode on continuously").
- Name both fullscreen switches (CLAUDE_CODE_NO_FLICKER=1 and "tui":
"fullscreen" in ~/.claude/settings.json) in the code comment, the invariants
and the two wiki pages.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Skip and latch a bounded Shell window once the browser is at xterm's
scrollback cap (scrollback + rows): a 1 MiB window of short lines can carry
more rows than the browser can ever hold, so it replayed and re-captured on
every scroll-to-top with no 60 s back-off.
- Label a replayed bounded window 'tail' even when the capture was byte-capped,
so the banner keeps offering Load full history instead of calling the rest
unrecoverable.
- Pin GET /terminal?full=1&tail=<n> in the route tests: full-history source,
truncationReason 'tail', and the closing relative cursor move survive the cut.
- Log the bounded skip via _logScrollRouting('repull-skipped-bounded').
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- pre-push hook: skip with a notice when npm is not on PATH (GUI git
clients and IDEs often run hooks with a minimal PATH), instead of
blocking every push on "npm: not found"; real-push test with a
stripped PATH
- test/git-hooks.test.ts: pin GIT_CONFIG_NOSYSTEM=1 and
GIT_CONFIG_GLOBAL=/dev/null around the resolveGitHooksDir tests, so
an exported global or a system core.hooksPath no longer fails them
- watch tsconfig.json, .prettierignore and .editorconfig too:
typecheck and format:check read them
- check:browser-excludes: fail loudly when the vitest list output and
the walked test/**/*.test.ts tree share no path (format drift would
otherwise pass vacuously)
- Reword the PRE_PUSH_MARKER comment: bumping its version would make every
installed v1 hook read as foreign and never refresh again.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- claude's declared-for-later wheelForward says the live rule in
_shouldForwardWheelToApp is the version AND the server-published
cliMouseTracking flag, so whoever wires the field up needs both
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- claude watchingLine: the lookahead keys on "Artifact" alone, so a
footer truncated mid-chip ("1 Artifact…", "1 Artifact comm…") is still
refused instead of reporting the shell beside it; comment follows
- test: both truncations return no watching label
- invariants: a chip that waits on a human never counts as watching, and
the ^ anchor is what stops the retry past the chip
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- test: the complete-identity case now checks the combined
agentImageBuildArgPairs() argv on both producers, so the manual
build-agent-image.mjs path cannot drop the identity unnoticed
- both producers: GIT_IDENTITY_BUILD_ARGS carries the mirror/parity
warning its gh/az neighbour has
- the partial-identity error names CODEMAN_AGENT_IMAGE_GIT_USER_NAME and
CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL; test regex follows
- wiki Docker-Cases: mention the identity variables next to the gh/az
switches
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- test: every ENV PATH= line in server.Dockerfile must start $PATH:, and
the ~/.local/bin append is pinned alongside /opt/codeman-cli/bin
- invariants + CLAUDE.md: the append-only PATH rule names ~/.local/bin too
- docker-compose.md: Settings-installed CLIs live in ~/.local on the
app-data mount; reinstall once after upgrading; hand-run npm installs
need --prefix ~/.local
- installEnv() JSDoc describes the in-container npm prefix redirect
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When Claude hands work to an ultracode workflow or background agents, it
ends its own turn and closes it with `✻ Waiting for 1 dynamic workflow to
finish` instead of `✻ Brewed for 1m 18s`, then resumes by itself when the
workers report back. The pane sits quiet with the composer up, so the idle
probe called the session idle for the whole wait. At phone width the
workflow's progress row also drops its ticking timer, so nothing on screen
changes for minutes.
A new optional registry field, `capabilities.workDetect.awaitingLine`,
names that closing row, and `_probePaneWorking()` counts it as work.
Claude renders the row once from a snapshot and never redraws it, so the
same words stay on screen after the workers finish. `isAwaitingWorkers()`
therefore tests only the newest column-0 row directly above the composer,
never the whole pane and never the PTY stream; a follow-up turn always
puts rows of its own there. The column-0 anchor also keeps an agent from
holding its own tab busy by printing the sentence.
Verified against the live Mac mini pane that reported the bug (2.1.283),
and end to end on an isolated instance: an ultracode session running a
90 s workflow at 46 columns stayed busy through the wait and the
follow-up turn, then went idle 6 s after that turn closed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On a phone every inactive tab rendered transparent: grey 11px text
floating in unmarked gaps, a boxed Alt+N digit in each tab (a phone has
no Alt key), names capped at 50px so a shared `w1-` prefix was most of
what showed, and the tab that did not fit was chopped mid-word against
the connection dot. The strip looked like a row of disabled labels.
Phone block of mobile.css only:
- Every header tab is a chip, filled and bordered from the skin's
--control-* tokens, name in --text at weight 500. Written
`:where(.header) .session-tab` so it stays at (0,1,0): the per-colour
left border still wins, and sidebar layout (where the list leaves the
header) is untouched.
- The Alt+N digit is hidden in the header; inactive tabs drop their
empty .tab-actions container, which padded the chip's right side.
- Name cap 50px -> 80px, status dot 4px -> 6px, strip gap 2px -> 6px.
- Scroll-driven edge fade: a mask on the strip whose widths follow its
own inline scroll timeline (registered @property lengths), so the
clipped tab dissolves into the edge. No JS; a strip that does not
overflow gets no mask, and browsers without scroll timelines keep the
old hard edge.
The tap-zone arithmetic comment is updated for the numberless phone
tabs and the bigger dot (the required reserve drops from 38px to 36px;
the 44px min-width stays). test/mobile-tab-strip-chips.test.ts pins the
(0,1,0) selector, the top-level @property registration and the
timeline-after-shorthand order, each of which fails silently otherwise.
test/mobile/tabs.test.ts follows the new name cap.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Clicking a .md in the Files panel showed wrapped source with an Edit
pencil and no way to see it rendered, although marked + DOMPurify were
already on the page for the Response Viewer. The viewer now renders
.md/.markdown through that same pipeline (one parser, one click
delegate) with an MD pill back to source, and the plain-text view gains
Lines (CSS-counter gutter) and Wrap toggles. All three persist per device
in their own localStorage keys.
- Relative images are rebased onto the workspace-confined file-raw route
under the document's directory, built inside a <template> so no fetch
fires before the rewrite; a failed load degrades to alt text. Relative
links become a.rv-path so the existing delegate opens them in the
viewer; fragment and http(s) links are untouched.
- The rendered container carries data-i18n-skip so the translator does
not rewrite the document's prose.
- Markdown fetches the route's 10000-line ceiling; other text keeps 500.
- avif renders inline (file-content image set, file-raw MIME map), and
avif/ico printed paths open the viewer instead of tailing bytes. .md
deliberately stays with the tail viewer for printed paths.
With local echo on, committed text sits in the LocalEchoOverlay and does not
reach the PTY before Enter, so the PTY cursor that places the preview span
stays at the prompt start. The span's z-index 6 only counts inside
.xterm-helpers (its own z-index 5 stacking context), and the overlay is a
z-index 7 layer whose first line is opaque from the prompt column, so every
composition after the first one in a prompt was drawn under the overlay.
- xterm-zerolag-input: add setComposition(text) and a composition getter.
The overlay draws the composition as an underlined, aria-hidden tail after
its pending text, through the same wrapping and grow-upward layout. It is
never part of pendingText, hasPending or anything sent; clear() and
removeChar() drop it, and rerender()/refreshFont() keep it.
- terminal-ui.js: while local echo shows typed text (on, and not handed back
to PTY echo by a nav key), render and clear the preview through
setComposition. The helper span stays for local echo off, and as the
fallback when the overlay cannot place the text (no prompt found).
- Browser test against real xterm 6, the overlay bundled from its source
and styles.css: a second composition after pending text is the topmost
element after that text, and the commit lands in the overlay once. Unit
tests for setComposition in the package and for the routing in the
structure test.
- CLAUDE.md and architecture-invariants: state the preview's effective layer.
- Observe keydown in the capture phase on terminal.element, an ancestor of
the helper textarea, so the controller sees it before xterm's own capture
listener finalizes the composition and emits the commit through onData.
Finalize on exactly the keys CompositionHelper.keydown does (every keyCode
except 20/229/16/17/18), ignoring isComposing and key as xterm does.
- Bound awaitingCommit with the same 2 s fallback as the committed phase, so
a composition whose commit never reaches onData cannot turn the next
unrelated keystroke or paste into an IME commit.
- pagehide resets the controller instead of destroying it, so a back-forward
cache restore keeps the preview working.
- Give the preview an opaque background from the terminal theme.
- Route an IME commit through the ordinary printable/paste local echo branch
and complete the commit afterwards; drop the send-on-throw fallback.
- Pin the event order with an xterm stand-in registered in the capture phase
ahead of the controller, and against real xterm in a browser test.
- CLAUDE.md: note the IME commit routing and the z-index 6 preview layer.
webview:changed carried only {action, id} and the SSE routing hint had no
webview: branch, so every connected client received it: in multi-user mode
any user saw the ids of other users' web-tab creates, edits and deletes.
The event now carries the web tab's owner (from the stored record) and is
routed to that owner plus admins. Single-user delivery is unchanged.
resolveGitHooksDir now returns a directory only when it is the repo's own
<git-common-dir>/hooks (compared on canonical paths), so a core.hooksPath
elsewhere, global or repo-local, is never written to by postinstall, while a
core.hooksPath pointing back at the repo's own .git/hooks still resolves.
The pre-push hook skips with a one-line notice when a pushed ref is not the
checked-out HEAD (tags peeled) or when git status shows uncommitted or
untracked changes under a path the checks read (src, config, scripts, test,
package.json, package-lock.json, install.sh), since the checks read the
working tree rather than the pushed commit.
Also: honest timing (~10-40s instead of ~15s), CLAUDE.md Session Safety note
on CODEMAN_SKIP_PREPUSH for another session's WIP, 14 (not 9) Playwright
tests, and a note that the browser-excludes check only sees direct imports.
npm run check:browser-excludes finds tests that import a browser driver and
asks `vitest list` whether the CI config still collects them; wired into CI.
npm install now also installs a marker-owned pre-push hook that runs the
static CI checks (~15s). Skip with CODEMAN_SKIP_PREPUSH=1; hand-written
hooks are left alone.
WebKit on iOS does not show text being composed by an IME inside the
terminal, so users type blind until it commits. Add mobile-ime-preview.js,
a visual-only controller that renders the composition in the xterm helper
layer and holds a committed chunk until local echo, parsed terminal output
or a 2s fallback shows it. Wire it into terminal-ui.js, the script order,
the build minify/hash lists and styles, with unit and wiring tests.
Claude 2.1.280 renders inline by default: no alt screen, no mouse tracking, transcript in real scrollback. The version-only gate still sent every wheel tick and touch swipe as SGR reports, which Claude ignores, so scrolling a Claude session was dead while codex (routed locally) worked. Gate forwarding on the server-recorded cliMouseTracking flag, which fullscreen mode (CLAUDE_CODE_NO_FLICKER=1) sets.
Adds a centered star call-to-action under the badge row in both the
English and Simplified Chinese READMEs.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A window cut at the tail size can be smaller than the browser's buffer
while tmux still holds more. The downgrade guard reads that as "tmux has
nothing more to give", which is true of an unbounded capture only, so a
bounded window reaching it marked the session exhausted and removed Load
full history from the banner.
The bounded skip now runs first, so such a window never reaches the
exhausted path, and it no longer writes banner state: relabelling it from
the bounded payload would call a terminal holding all of a Load full
history pull "the most recent 1 MiB".
A skipped window that came back truncated cannot reach anything older
than the browser shows, and every ask costs the server a synchronous
capture-pane of the whole history (tail is applied after the capture), so
it puts the session on the 60 s cooldown. An untruncated one keeps 4 s.
_replayWouldShrinkBuffer takes optional pre-estimated rows so a megabyte
capture is not scanned twice. CLAUDE.md's Full-scrollback replay entry no
longer says Shell never pulls on ordinary scroll.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A burst of output leaves a Shell pane with about one screen of browser
scrollback, because tmux repaints the burst instead of scrolling it,
while tmux itself keeps every line. Shell declined the scroll-to-top
re-pull other modes use, and the Load full history button renders only
once a replay was truncated, so a Shell tab under 1 MiB could not
scroll back at all.
The scroll gesture now pulls ?full=1&tail=TERMINAL_TAIL_SIZE, the same
bound a tab switch loads; the route's existing tail cut marks longer
histories 'tail', so the banner still offers the unbounded pull. A
window no longer than the browser's buffer is not rewritten.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An agent that publishes an artifact arms a monitor for its comments and
ends its turn. Claude Code shows that on the footer as `1 Artifact
comment monitor`, and #473 put that chip on the list of background work,
so the session counted as watching and its idle prompt opened already
acknowledged. Unlike every other chip on the list, that monitor waits
on the user: the agent hears nothing until somebody comments.
Claude's `watchingLine` now refuses any footer that carries the chip,
through a lookahead over the whole row, so a shell running beside the
monitor cannot report the session as watching either.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The image sets NPM_CONFIG_PREFIX=/opt/codeman-cli, which is image content, so
Update-Codeman.sh discarded every npm-installed CLI (dsh, pi). In the Compose
container, POST /api/clis/:id/install now installs into ~/.local on the
persistent home mount, and ~/.local/bin is appended to the image PATH.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
The phone case picker had no way to narrow a long case list, so finding
one meant scrolling a sheet that showed about six rows at a time.
- A search field filters rows by name (every typed word must match, any
order, case-insensitive), with a "No matching cases" state. Enter picks
the case when exactly one row is left; Escape clears, then closes.
- The field is not auto-focused, so opening the picker does not raise the
keyboard. The list holds its unfiltered height while searching so the
sheet does not jump, and the input is 16px so iOS Safari does not zoom.
- Layout: the sheet padded the home-indicator inset on top of the footer
already doing so, leaving a dead band under Create New Case; the sheet
now grows to 80dvh and the list fills it instead of a separate 50vh cap.
- Opening scrolls the list (its own box, not scrollIntoView) to the
currently selected case.
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
The Run bar's case picker only loaded /api/cases at page load, so folders
deleted or created on disk stayed listed until a reload. It now refetches
on open and every 5 seconds while open, repainting only when the list
changed and falling back to another case if the selected one was removed.
The Manage tab of Add Case gains a search box filtering by name or path.
Reorder arrows are disabled while a filter is active so a swap cannot
involve a hidden case.
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(cleanup): keep .claude-images while a sibling session uses the same dir
cleanupSession() recursively removes {workingDir}/.claude-images. That
directory belongs to the working directory rather than to the session, and
several sessions routinely share one case directory, so closing one session
deleted the pasted images a live sibling still referred to.
The removal now runs only when no other live session has the same working
directory. A session that is itself being cleaned up does not count as live,
so two sessions of one case closed together still remove the dir.
Split out ahead of the exited-agent sweep for Ark0N/Codeman#446, which closes
sessions unattended and would otherwise make the loss routine.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(session): close sessions whose agent exited cleanly (#446)
Part 2 of Ark0N/Codeman#446. Part 1 records an exited agent as
SessionState.paneExit. A session whose agent the user ended with /exit is
now closed through cleanupSession(), the same path the X button takes, so
finished sessions stop piling up on the board. The lifecycle log records
the reason as "agent exited cleanly (status 0)", and the conversation stays
resumable from the Resume list.
shouldCloseCleanlyExitedSession() in the new pure module pane-exit-sweep.ts
holds the rule. It closes a session only when all of these hold:
- The exit status is an explicit numeric 0 with no signal. An absent status
is how a SIGKILL presents on tmux 3.2a, so it counts as unknown and the
row stays. A non-zero status or any signal also keeps the row, with the
exit code on the tab.
- Two authoritative pane reads agreed on that exit.
TmuxManager.getPaneExitReadCount() counts them, and a failed, empty or
skipped read neither confirms nor resets the count.
- No start, attach or relaunch is running for the pane.
Session.paneLifecycleInFlight covers _setupOrAttachMuxSession(), whose
dead-pane branch revives an exited pane on purpose, and restartCli().
setPaneExit() already scopes paneExit to local mux-backed sessions, so
remote, docker and direct-PTY sessions are never closed.
planRebootRestore() now refuses a record whose persisted paneExit is a
clean exit. That covers an agent that exited just before a reboot, before
the sweep reached it. A crashed agent's record stays eligible, like its row.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(web): show "exited" on the phone overview and desktop home rail (#446)
Part 1 of Ark0N/Codeman#446 taught the tab strip and the rich rail rows
to say that a session's agent has exited. The phone overview and the
desktop home rail still said "idle", beside a green or pulsing dot.
_mobileOverviewExit() in mobile-overview.js is now the one rule for all
three surfaces, and _sidebarRichRow() uses it as well. It changes what a
row shows and leaves the row's state alone, because the state still picks
the section and the sort order. An exited row gets an "exited" pill, a
neutral dot and row accent, and a duration measured from when the server
first saw the pane dead. A pending permission prompt or question still
wins, as it does on the tab.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(cleanup): close the gaps review found in the #446 sweep and image guard
Four fixes from a dual review of Ark0N/Codeman#446 part 2.
- The .claude-images guard compares canonical paths, so a sibling that
reaches the same directory through a symlink keeps it. Its comment used to
say that case only missed a deletion; it caused one.
- A detached session counts as a live sibling. DELETE ?killMux=false removes
it from the server's map while its pane keeps running, so the guard now
reads persisted records too, and exempts only sessions being killed rather
than every session in cleaningUp.
- A session being closed refuses startInteractive() and startShell(). The
/interactive route awaits listener setup before the start, and a start
that raced the close could launch a CLI in a tmux session whose record was
then deleted. A failed close clears the mark again.
- The clean-exit sweep tries each exit once, keyed by session id and the
exit's at stamp, so a close that fails is not retried and logged every
two seconds.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(session): keep a clean exit that lands within 10 s of a pane start (#446)
A CLI that prints a startup error ("not logged in", a bad profile, a config
error) and exits 0 used to lose its tab, and the error with it, about 4 s
after launch. The sweep now keeps any clean exit that lands within
CLEAN_EXIT_MIN_PANE_LIFETIME_MS (10 s) of the last start, attach or relaunch
finishing (Session.paneStartedAt, stamped when _withPaneLifecycle ends). The
row stays as "exited (0)" for the user to read and close.
Verified on an isolated instance: a shell that ran `exit 0` 2 s after start
kept its row, one that exited after 13 s was closed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A single input over MAX_INPUT_LENGTH (64 KiB) was queued for reliable
delivery, refused by both transports (the WebSocket silently, POST with a
400), and never dropped: the client treated the 400 as transient, so the
frame was re-sent every 2 s forever, blocked every later input for that
session, and came back from localStorage on every reload.
- Client: a paste over the frame limit is split into in-limit frames
(never cutting a surrogate pair) delivered in seq order; over 1 MiB, or
an oversized mux write, it is refused with a toast and never queued.
- Client: the POST drain drops a frame answered 400/413; a WS error ACK
drops it too; frames over the limit persisted by an older build are
pruned on load.
- Server: the WebSocket answers an oversized sequenced frame with
{t:'ia',seq,err:'too_large',max} instead of silence (an older client
reads that as a plain ACK and drops it); the POST schema uses
MAX_INPUT_LENGTH instead of a second 100000 limit.
Verified end to end on an isolated instance: a 110 KB paste reached the
PTY byte-identical over both the WebSocket and the POST path, a poisoned
120 KB persisted frame was pruned on load, and a 2 MB paste showed the
refusal toast with nothing queued.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`dsh plugin` spawns a literal `pnpm` with no npm fallback, so the Run
menu's "DeepSeek - add a terminal profile" button failed with
`dsh: pnpm not found on PATH` (exit 127) on the server image. The agent
image already installs pnpm for the same reason (#352). Pin pnpm@12.6.0
in the runtime-writable CLI prefix and note it in the DeepSeek doc.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
Adds claude-opus-5-5 to the App Settings model picker (base option with
data-ctx="1" plus its [1m] companion row, since Opus 5.5 has a 1M window)
and to the five task-routing selects, mirroring how Fable 5.1 was added.
Co-authored-by: Claude <noreply@anthropic.com>
* fix(sessions): stop pinning the w1-myapp placeholder as Claude's /resume title
Local claude spawns passed the tab name as `--name`. That flag is not only the
cross-session peer name: it is also the prompt-box label, the `/resume` picker
entry and the terminal title, and a pinned title stops Claude generating its own
(`customTitle ?? aiTitle`). So every conversation of a case was listed in
`/resume` as the same `w1-myapp`, and none of them got a generated title. On one
workspace, 34 of 34 conversations spawned with `--name` had no ai-title, while
every conversation spawned without it had one.
Only a name the user chose is pinned now: `Session.cliPinnedName` is the name
when `nameSource === 'manual'`, carried to the builders as a separate `cliName`
so the tab/mux name is untouched. Placeholder and auto names let Claude title
the conversation again.
A rename in Codeman also reaches `/resume`: the new name is appended to the
conversation's transcript as the `custom-title` row `/rename` writes (never
creating the file, never writing an empty title). For a pane spawned without
`--name` this holds immediately; a pane spawned with one re-appends its own
title each turn, so there the new name holds from the next spawn.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(sessions): skip no-op renames and docker sessions when syncing the /resume title
A same-name PUT (the Session Options field saves on blur and recomposes the
unchanged placeholder) no longer flips nameSource to manual or appends a
custom-title row, and docker sessions skip the host transcript scan since their
transcript lives in the container. The skill pages no longer use a w<N>- name
as the peer-name example, and the changeset notes the re-append caveat.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs: record that nameSource decides --name and renames reach /resume
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: codeman-local <codeman@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
_loadSessionManagerList() re-projects each unified item into the
history-record shape _buildHistoryItem renders, and dropped these three
fields. The row's own onActivate still read them from the unified item, but
everything built from the record did not: the ⋯ menu's "Resume session"
relaunched a codex row as claude (no mode, no resumeId), a resumed session
lost its conversation id, and Cmd+K rows showed no mode badge. Same class
of bug as the worktree fields the re-projection already carries (#266).
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (da07b38c, db4557d9) with no
corresponding update to the plan doc itself — every checklist still read
Status: TODO and every box unchecked. Brings the doc in line with the tree:
- A new "Status as of 2026-09-22" section up top: what's actually
implemented (verified by grepping the routes/schema/UI, not just trusting
the commit messages), the availability-flag staleness bug found and fixed
in this session (commit 0c77dd0a) with its devbox verification record, and
one real outstanding gap.
- The outstanding gap: a custom CLI created via Phase 5's write API has no
way to actually be launched. The Run menu is static per-mode markup with
no consumer of window.__codemanCliCatalog, so Phase 6's own "create a
custom entry, confirm it can be launched" verify step was never actually
exercised against this. Documented with two candidate fixes, neither
started.
- Each phase's checklist flipped to [x] where confirmed present in the tree,
Status lines updated from TODO to DONE, and the two originally-open
questions (Phase 2's installed source, Phase 5's PUT endpoint shape)
marked resolved against what actually shipped.
No code changes in this commit — documentation only, so a future session
(or the one already mid-flight on a separate checkout of this same branch)
picks up accurate status instead of a stale plan.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs: add the CLI-registry deployment plan and the parked Copilot plan
Both were sitting as untracked scratch files in the master checkout,
never committed to any branch. Moving them here rather than leaving them
loose:
- DEPLOYMENT_PLAN.md is the live tracker for the CLI-registry follow-up
series (PR A #347 merged, PR B #380 merged, PR B2 merged as #458) and
is where PR C (this branch's own CLI-management work) belongs.
- docs/copilot-integration-plan.md is explicitly PARKED, referenced by
name in docs/cli-enable-disable-plan.md's own header as a sibling plan
tracked separately — kept for continuity, not active on this branch.
The other scratch files found alongside these (PRA.md, PRB.md, PR-B2.md
and their review-response counterparts) described PR A/B/B2, all now
merged — deleted from the master checkout as stale rather than committed
anywhere, since their content is superseded by the real merged PRs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* fix(cli-registry): render enabled CLIs in launch surfaces
* test(cli-registry): update frontend branch guard
* fix(test): isolate suite from deployment environment
* fix(cli-registry): revise Decision 4 - claude is toggleable, shell stays permanent
shell/claude were both structurally un-disableable in the original plan
(Decision 4). Revised: shell keeps the hard backend guarantee (it is the
one non-agent mode several code paths assume always exists as a raw-
terminal fallback), but claude is now a normal toggleable entry like any
other CLI.
Safe to do because internal session creation (tmux-manager.ts, session.ts,
Ralph, plan-orchestrator) resolves a CLI via getCli(), which does not
check `enabled` at all - only the Run menu and the HTTP-facing
sessionModeSchema() (new session requests through the normal API) key off
it. Disabling claude therefore behaves identically in kind to disabling
any other CLI: no internal fallback path breaks, it just stops being
offered for new sessions until re-enabled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): hide shell's toggle entirely instead of greying it out
A permanently-disabled switch next to every other row's working toggle
read as broken rather than intentional. shell now renders no switch at
all - a plain "Always available" label - so there is nothing to click
that could look like it should work but doesn't. Backend guard is
unchanged (UNDISABLEABLE_IDS still refuses shell unconditionally); this
is UI-only.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): sort the Installed CLIs list, installed-first then alphabetical
renderCliList() previously rendered in registry order (each entry's fixed
order field). Now sorts installed CLIs first, then not-installed, each
group alphabetical by label - matches how a user actually scans the list
(what's ready to use, then what needs installing).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* style: prettier fixes from the master merge
* fix(cli-registry): install/edit take effect immediately, confirm before install, phone labels
Four gaps found verifying #476 against the #343 review trail:
- Installed or edited CLIs kept reading as missing/stale. Every binary lookup
(the nine per-CLI resolvers and the generic registry one) caches in its own
closure, with a negative-cache backoff of up to 5 minutes, and nothing
cleared them. invalidateCliExecutableResolvers(binaries) now drops those
caches per binary; install (success or failure), create, edit and delete
call it plus invalidateCliResolverCache(id). Before this, a CLI installed
from Settings could fail to launch for minutes, and an edited custom entry
kept launching its old binary until a restart.
- The Settings "installed" badge for a custom entry used a private `which`,
ignoring the entry's searchDirs and the login-shell lookup that spawn and
the Run menu use; it now asks the same generic resolver they do.
- Install ran on a single click. The #343 review asked for auto-install to
sit behind an explicit confirm; the confirm now names the exact command,
which GET /api/clis returns for stock entries only (installCommand).
- The phone Run button showed the two-letter tab badge ("CC", "CX") instead
of the word ("Claude", "Codex"). It uses the registry label again, which is
identical to the old static table for every stock CLI (now pinned).
14 new tests; 9 of them fail against the previous head and pass here.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): address #476 review — safe serialized writes, no id branches, docs
Must-fix:
- registry-writer: start fresh only on ENOENT; refuse (409) a clis.json that
does not parse or has group/world permission bits instead of overwriting it
(isUnsafePermissions now exported from registry.ts)
- mutateRegistryFile(): one promise chain for every mutation, with the
existence/duplicate checks inside the serialized step, plus a unique tmp
name per write
- docs: CLAUDE.md, architecture-invariants, cli-registry (new Settings
section) and api-reference (the six /api/clis routes)
- drop DEPLOYMENT_PLAN.md and docs/copilot-integration-plan.md
Smaller:
- PUT /api/clis/custom/:id keeps the entry's current enabled state when the
body omits it
- runMode setter falls back to the first enabled catalogue entry, not 'claude'
- shell guard keyed on kind === 'shell' (routes + Settings list); stock probe
map shared with server.ts via utils/cli-installed-probes.ts
- stock claude label is now 'Claude Code', so the Run menu / phone overview
label rewrites are gone (doctor row keeps "Claude CLI" via its override)
- welcome buttons are translatable again and read "Run Claude Code" /
"Run Shell"; zh-CN gains "Run Codex" / "Run OMP"
- install: per-id in-flight guard (409) and CODEMAN_* stripped from its env
- fileoverview / CliEnableSchema comments no longer say stock-only
- test-env isolation changes moved to their own PR
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* test(cli-registry): pin the #343/#347 findings #476 makes reachable
A CLI toggled or created through the routes is accepted or rejected by
CreateSessionSchema with no restart (#343 finding 2), and a custom CLI created
through the API renders a real local, remote and docker launch command
(#347 finding 5: no more `cd <path> && undefined`).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(self-update): stop a stalled status from blocking every later update
A Homebrew node upgrade under a long-running server deletes the versioned
Cellar path the server passes as --node, so every status write from the
updater failed. The update itself still built and restarted (npm and the
build use node from PATH), but update-status.json stayed "queued" forever.
The boot reconcile ran one minute after the restart, inside its 15 min
window, and isInFlight() had no age limit, so "An update is already in
progress." blocked every later update until the next server restart.
- self-update.sh falls back to node on PATH when --node is not executable.
- expireStalledStatus() (pure) fails an in-flight status whose last write
is older than the stale window; applied on every read (start + status
poll) and persisted. The live updater heartbeats every few seconds, so a
running update never trips it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(self-update): a hung graceful shutdown no longer leaves a LaunchDaemon install down
On a KeepAlive LaunchDaemon (headless macOS) the updater restarts by sending
the server SIGTERM and letting launchd respawn it. launchd only respawns once
the process EXITS, and nothing escalates a stuck stop (systemd would SIGKILL
after TimeoutStopSec). Observed after an update to 1.32.1: the server closed
port 3000, server.stop() never resolved, the process stayed alive and the
service stayed down until it was killed by hand.
- cli.ts: the signal handler arms an unref'd 10s timer that force-exits if
server.stop() hangs.
- self-update.sh (launchd-daemon): wait up to 30s for the server pid to exit,
then SIGKILL it. tmux sessions live outside the server and survive.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
- docs/wiki/The-Dashboard.md: the tab-appearance table gains the exited
state (muted dot plus an `exited (137)` badge) and explains the bare
`exited` variant.
- The detailed sidebar and rail no longer pair the muted dot with an "idle"
pill: an exited session's pill reads "exited" (neutral styling) and its
since stamp measures from the observed exit. This is a label override on
the row model, not a new state, so SESSION_ACTIVITY_RANK and the home
screen order are untouched, and a pending alert still keeps its own pill.
The row signature includes the flag so the incremental path repaints it.
- The exited badge is aria-hidden like its sibling badges, and the exit is
appended to the tab's aria-label in both render paths through one helper.
- test/tmux-manager.test.ts re-adds the junk-trailing-field parser case
against parsePaneRows.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- test/setup.ts strips CLAUDE_CONFIG_DIR (pinned in test-env-isolation), so
transcript-fixture tests such as session-custom-model-restart no longer go
red on a machine that exports it for a separate Claude account (#255).
- The vanished-tmux-session branch of _setupOrAttachMuxSession() relaunches
the CLI through createSession() just like a failed respawn, so it now takes
the same resume pin. A genuinely new session is unaffected.
- After a dead-pane respawn of a fallback-chain CLI, _claudeSessionId names
the conversation the walk actually pinned instead of the chain tail, which
the walk may have passed over for lack of a transcript.
- _claudeConfigDir() trims the override like claudeProjectsDir() does.
- The remote-reattach test is labelled as documentation, since the pin
builder's own remote guard would make it pass either way.
- CLAUDE.md: the create-path pin persists through toState() as
resumeSessionId, and the end of the walk adds no pin rather than clearing
the launch seed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- The TERMINAL DROP crash-trail line moves behind the scheduler's debounce
guard, so it is written once per window rather than once per dropped
frame. At the server's 8ms batching, one second of drops evicted the whole
50-entry trail, including the recovery lines that explain it.
- A refresh that failed at the capture fetch deadline now returns
'deadline', and the scheduler does not retry it: that is a stalled link,
not contention, and each retry was another ?full=1 capture waiting out a
deadline of up to two minutes. The early-return retries are unchanged.
CLAUDE.md and the code comments no longer claim every skip reason is
transient contention.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- While another device holds the pane width (_paneWidthRefused), a resize
now asks for the container's width without applying it locally
(_geometryForResizeRequest: rows follow the container, columns stay at
the PTY's). Fitting first re-wrapped the whole buffer to the container
and back on every 30s mobile retry, and throttledResize ran the
scrollback clear for a resize that brings no redraw. selectSession
clears the flag, since it belongs to the previous pane. New unit tests
run the real mixin against a fake terminal and fail without the fix.
- Session seeds _ptyCols/_ptyRows at spawn (_notePtySpawnGeometry), so a
reattached pane reports its tmux window's real size through ptyGeometry.
- Session.resize's declined-branch comment names ptyGeometry, not the
deleted ptyCols/ptyRows getters.
- Delete the dead terminalGeometryAgrees() and its window export.
- test/xterm-private-api.test.ts header: it pins the exact locked version,
so any bump fails, not only a major.
- The main-terminal fit sweep also matches fitAddon?.fit?.(), and
CLAUDE.md names the modules it actually covers.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Clone Repo clears the credential helpers for a non-admin, but a non-admin's
Docker case with credential seeding on still receives a copy of the server
account's gh/az sign-in when the agent-image switches are on, the same as
the Claude and Codex credentials. Say so in the multi-user notes so the docs
do not read as a stronger guarantee than they are.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Remove exactly the codeman-node-modules/codeman-dist volumes by Compose
label after a plain `down`, instead of `down --volumes` (which also takes
any volume an override file declares while the message named two).
`down --volumes` remains only as a warned fallback when the project name
cannot be resolved.
- Report a failing first `docker compose config --format json` call with a
clear error instead of exiting silently under `set -e`.
- Filter empty label lines in the collision guard so an unlabelled container
cannot hide a real collision; name the moved-checkout exit in its error.
- Comments no longer cite a guard or incident in Start-Codeman.sh that does
not exist; the README states the real gap (a Node base-image bump leaves
codeman-node-modules stale because the lockfile did not move).
- docs: Update-Codeman.sh in the docker-self-update.md short-version table
and a mention in docker-compose.md; "Major updates" moved under "Updating"
in docker/README.md.
- test: smoke test covers the new sequence, the config failure and the
empty-line case; quiet stdio; @fileoverview names the fourth concern.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- session.ts: a pane capture that fails now CLEARS the watching label
(and emits watchingChanged so pages drop the badge) instead of keeping
the last one, so a failed capture degrades toward an alert rather than
pre-acknowledging the next real idle prompt. Test updated; invariant
noted in architecture-invariants.
- approvals-ui.js: the header bell counts only unacknowledged items
(pendingApprovalsCount), matching codeman tui's pendingApprovalCount();
pinned in watching-no-alert.test.ts.
- mobile-overview.js: move the orphaned "Pill copy per state" JSDoc back
onto MOBILE_OVERVIEW_PILL_LABEL.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- stock.ts: claude is no longer the only entry declaring transcriptGutter;
codex declares it too.
- architecture-invariants: the strip applies when the session's CLI declares
a margin (not detection), and a note that it keys on the session's launch
mode, not on what is running in the pane (a claude pane dropped to a shell
still loses up to two columns; copyStripMargin is the escape hatch).
- render-index-html test: the gutter map is injected for a solo
/session/:id render as well.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- mobile.css: gemini and antigravity run/gear rules get `!important` like
pi/omp/grok/deepseek, so the gear half no longer keeps the skin accent
while the body takes the mode colour (two-tone button on the default skin).
- test/skin-themes.test.ts: static guard that every run mode with a base
`.btn-toolbar.btn-run.mode-<id>` rule also has a resting rule inside the
`html:not([data-skin="og"])` block; ids are derived from the stylesheet.
- stock.ts: grok's accent comment names zinc-300 (border/badge colour);
gemini's accent is #8ab4f8 to match its tab badge and run-mode dot, noted
as the one exception to the border-colour method.
- types.ts: "(below)" -> "(above)".
- docs/cli-registry.md, CLAUDE.md: `accent` is now measured, not transcribed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CLAUDE.md loads into every session, and its Architecture section had grown
feature write-ups (history, measurements, rationale) that belong in
docs/architecture-invariants.md per the file's own header. Each long block
now keeps what the feature is, where it lives, its setting/default and the
rules that prevent real bugs, and links to its invariants section. Everything
removed was moved there: 29 new sections, extra facts appended to the
existing ones.
Also: hard-coded counts (SSE events, route handlers, module/file counts,
device profiles) replaced by pointers to the source of truth, and the
Debugging commands fixed to use the codeman tmux socket and HTTPS for prod.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Addresses the review on #472.
- CRED_STORES: `.config/gh` and `.azure` now carry `enabledByEnv`
(CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ), and resolveDockerCredentialArtifacts
skips a store unless that variable is exactly `1`, read at container
create. A host that merely has ~/.config/gh/hosts.yml or a plaintext MSAL
cache no longer copies them into every case container. Tests: the default
environment seeds neither even with the files present, and each store
follows only its own switch.
- Multi-user mode: a non-admin's Clone Repo clone and preflight run with
`git -c credential.helper=` (GIT_NO_CREDENTIAL_HELPERS, placed before the
subcommand), so the server account's helpers are never lent to them.
Verified against a real private repo that it also clears the URL-scoped
credential.<url>.helper entries, and that public clones still work.
Tests: the argv in test/git-clone.test.ts, and the route decision
(non-admin cleared; admin and single-user kept) in
test/routes/case-clone-credential-helpers.test.ts.
- Docs: recreate the case container to pick up seeds (docker/README.md,
Docker-Cases wiki, docker-cases.md); the multi-user behaviour in
docker/README.md and security-architecture.md; "functionally unchanged"
instead of "unchanged" for an image built with both switches off
(server.Dockerfile comment, README, changeset).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw
CONTRIBUTING says releases are handled by the maintainer via changesets after
merge, and every `.changeset/*.md` on master was written by him or by the
release bot — including the ones covering other people's pull requests. The
summary this file carried moves to the pull-request description, where it is
the maintainer's to reuse or rewrite.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claimed after a manual test that a session comes back from a server restart
without its badge until it next produces output. Measured instead of assumed,
and it is wrong: a codex session with a background terminal still running had
its label back within about 20 seconds of the restart, with no input from
anyone. Reconciliation re-attaches the pane, the attach repaint carries the
composer glyph, the idle confirmation arms on it, and the probe re-reads the
label — the ordinary path, doing the ordinary thing.
What produced the false claim was a session whose monitor had simply expired
while it sat there. Its footer carries no chip, so `watching: null` was the
right answer and there was nothing missing to restore.
Recorded at the field and in the invariants, because the shape of this invites
exactly one wrong fix: a polling timer to keep a value fresh that the pane
already refreshes by itself.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add Case -> Clone Repo could only reach public repositories in the Docker
deployment. This lets a deployment opt in to the GitHub CLI and the Azure
CLI (+ azure-devops extension) as git credential helpers. Codeman itself
still collects no credentials.
- server.Dockerfile / agent.Dockerfile: CODEMAN_INSTALL_GH /
CODEMAN_INSTALL_AZ build args (0 or 1, default 0; anything else stops the
build). Off leaves no apt repository, package, extension, helper script
or credential entry, so a default build is unchanged. On installs from
the vendors' apt repositories and configures system gitconfig helpers:
github.com / gist.github.com -> `gh auth git-credential`, dev.azure.com /
*.visualstudio.com -> new docker/git-credential-azure-cli (an Entra ID
token from `az account get-access-token`, or AZURE_DEVOPS_EXT_PAT).
A helper whose CLI is not signed in prints nothing, so a private clone
still fails fast.
- The extension lives in AZURE_EXTENSION_DIR outside HOME
(/opt/codeman-az-extensions, runtime-owned; /opt/az-extensions, gid-0
group-writable in the agent image).
- Hosts turn them on in docker-compose.override.yml: `build: args:` for the
server image, `environment:` CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ for the
agent image. build-agent-image.mjs and the in-app auto-build share one
env -> ARG table (pinned by the parity test) and pass nothing when unset.
docker-compose.yaml is untouched; .env.example only gains a comment, so
the self-updater's environment gate sees no new keys.
- Docker cases seed the gh sign-in (~/.config/gh/hosts.yml, config.yml) and
the az sign-in files from ~/.azure per file, read-only, like pi/grok.
- The Clone Repo AUTH_REQUIRED message says how to sign the server's git
in instead of claiming private repositories cannot be cloned.
- Docs: docker/README.md "Private repositories", docker-compose.md,
docker-cases.md, the Quick-Start / Core-Concepts / Docker-Cases wiki
pages, security-architecture.md, architecture-invariants.md, changeset.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw
Review fixes for #469.
The Ctrl+C branch cleaned the selection to decide whether to copy and then
passed that cleaned string to copyTerminalSelection(), which cleans again. The
trailing trim is a fixed point, so that was safe until this PR; the margin
strip is not, because it takes the lesser of the declared width and the run
every line shares, so a second pass takes up to `margin` columns more. The
branch now gates on the cleaned string and hands the raw one on. Verified in
chromium with a real drag, a real Ctrl+C and a real clipboard read on a live
claude pane: an on-screen ` fix(terminal): trim it` reaches the clipboard
as ` fix(terminal): trim it`, and reverting the branch reproduces the
reported ` fix(terminal): trim it`.
Pane B of a split resolves its own width. `_cliGutterColumns()` and
`_normalisedSelectionRange()` take the session and the terminal to read,
defaulting to the primary pane's, so Pane B looks its own run mode up instead
of keeping a margin Pane A drops on the same keystroke. Verified live with two
claude panes open side by side.
A detached session window (`/session/:id`) receives the gutter map. The
injection sat inside the block that skips the run menu's payloads for a solo
window, so the toggle worked in the main window and did nothing in the popup on
the same device. It needs no availability probe, so it moved below that block
and the solo window still carries none of the payloads it skipped before.
The settings description said the width is measured and named Codex as exempt.
Nothing is measured, and Codex is one of the two panes that are stripped.
docs/wiki/Settings-Reference.md gains the row every Terminal and Input toggle
carries. CLAUDE.md no longer says the clean touches trailing runs "and nothing
else" one sentence before the leading-margin rule, and both it and
docs/architecture-invariants.md record that the strip is not idempotent.
Two round-trip tests run on a mode that declares a gutter, which the existing
copyTerminalSelection cases could not, since they all use the harness default
mode that declares none. The Ctrl+C branch itself is pinned at the source,
because it lives inside initTerminal's attachCustomKeyEventHandler closure over
a real xterm the vm harness cannot build. Both pins fail on the reintroduced
bug. Gate: 7865 passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`_onSessionTerminal` drops an incoming frame once the app-owned render queues
already hold 128KB. That is the right call — the alternative is an unbounded
backlog — but a hole in a TUI byte stream is a desynced cursor, and a desynced
cursor is muffled text (#464). The drop was only half of it.
The recovery was a fire-and-forget timer: it nulled its own handle and then
called `_onSessionNeedsRefresh()`, which opens with four early returns. Two of
them — a buffer load already in flight, a refresh already owning this session —
are MOST likely to be true during exactly the output burst that caused the
drop. So the recovery was skipped precisely when it was needed, with nothing
left to retry it, and the dropped bytes were never replayed.
`_onSessionNeedsRefresh` reports whether it actually repainted now, and
`_scheduleDroppedOutputRecovery` re-arms while it has not. Bounded by
`DROP_RECOVERY_MAX_ATTEMPTS`, because every reason the refresh can be skipped is
transient contention that clears in seconds and a permanently failing refresh
must not become a loop against the API; giving up at the cap leaves exactly what
the old code left, so the floor is no worse than before. The same 2s debounce
still collapses a burst of drops into one attempt.
This is the principle Ark0N established reviewing #431 for the WebSocket
output-gap marker — only a repaint that actually happened settles the recovery —
applied to the one recovery path that still trusted a timer having fired.
The retry decision is a pure function in constants.js so the gate can reach it,
and the scheduler itself is driven from app.js under a fake clock. The retry
case and the no-retry case only pin the fix AS A PAIR: either alone passes
against something wrong, one against the old fire-and-forget timer and the other
against retrying forever. Checked by reverting app.js to the old shape, where
three of the twelve fail.
Two harness details that would otherwise have made the tests lie. The vm context
baked in the real `setTimeout`, so `vi.useFakeTimers()` could not reach the
scheduler and every case reported zero calls; it delegates lazily now. And
app.js reached `CodemanDroppedOutput` as a bare global, which resolves in a
browser but not in the vm — worth fixing beyond the test, because that call sits
inside a timer where a ReferenceError is swallowed and would take the recovery
with it. It reads through `window.` like terminal-ui.js does with its own
constants.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Reported from a manual test: a Codex session went on showing the watching badge
after its background terminal had finished. The server was right and the page
was stale — `Session.watching` changes while the session's status does not, and
nothing broadcast it.
The label is usually SET on the idle transition, which broadcasts anyway, so the
badge always appeared correctly. It CLEARS when the work ends, and a CLI can end
background work without taking a turn: codex repaints its background-terminal
row away and stays idle, so `_confirmIdle()` concludes without emitting `idle`
(that emit is guarded by `wasWorking || isInitialReady`) and no other event
fires. Every open page kept drawing a badge the server had already dropped.
`_readWatching()` now emits `watchingChanged` when, and only when, the label
really changes, and the wiring pushes the session state on it. No new SSE event:
the badge reads off the session payload every surface already has.
A/B measured on an isolated beta with the page loaded and then left untouched.
Without this commit the server dropped the label at t+50s and the page still
showed the badge at t+100s; with it, page and server cleared in the same
ten-second window.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Measured on the same beta, codex-cli 0.154.0: a sub-agent started without
waiting outlives the turn exactly as a background terminal does — the sandboxed
process was still running — and codex shows nothing for it. The last rows of the
pane are the composer and the status line, and `Sub-agents running` belongs to
the on-demand `/subagents` panel rather than to the row above the composer.
So there is no second codex label to add. A codex session waiting on a sub-agent
reads as plainly idle, which misfiles nothing (codex raises no idle prompts) and
simply leaves that one kind of quiet unexplained until codex pins a row for it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief)
found the trust boundary weaker than the comments around it claimed. Eleven
findings, all applied.
The two blockers were both about who can write the row the label is read from.
Claude's window covered two rows, and the second one is the status line, whose
command a session running with permissions bypassed can write into its own
`.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of
its own and silence its own idle alert. The default window is one row now, which
is the footer and nothing else, and the constant says why. Separately, the label
reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML,
which is an injection sink for any config-supplied pattern whose capture group is
permissive; it goes through escapeHtml() like every other untrusted string in
that file.
The Codex entry could not be fixed the same way, and now says so. Its row is
third from the bottom only while a terminal runs; with none running that slot
holds the last row of the transcript, so matching the complete row (with the
`/stop to close` tail, window narrowed to three) raises the bar without closing
it. What contains it is `hooks: 'none'`: no hook event from a codex session
reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an
alert. The registry comment, `docs/cli-registry.md` and the test all state that
rather than claiming a guarantee the code does not have.
Also from the review: the TUI header badge no longer counts an acknowledged
item, which was the same gate the classifier fix already went through and was
wrong for human acknowledgement too; the TUI approval card reads the quiet
reason and drops to a new `info` tone instead of asking for a reply; the badge
carries an aria-label, because the phone it was built for has no hover target;
the schema refuses `watchingLines` without a `watchingLine`; and the pattern and
its window are resolved together rather than one memoized and one not.
Documentation moved with it. The mechanism now lives in
`docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer,
`docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped
buzzing, and both that page and the changeset name the limitation neither did
before: a question asked in plain prose is not a dialog, so it is silenced along
with the false alarms while background work runs.
Verified live again after the narrowing, on an isolated beta: a Claude session
reported `1 monitor` and took its idle prompt acknowledged, and a Codex session
reported `1 background terminal` against the full-row anchor.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ten findings from a two-model review of this branch. Both reviewers cleared the
detection logic itself; everything here is a gap around it.
A route that starts a command in a pane now PERSISTS as well as broadcasts.
`/interactive` and `/shell` did neither before, and the pane-exit watcher cannot
cover for them: its next tick finds `paneExit` already cleared in memory,
reports no change and writes nothing, so `state.json` kept saying the agent had
exited for as long as the session stayed quiet. Nothing reads that record for a
decision yet, which is exactly why it had to be fixed now — part 2 is designed
to read it. The `clearPaneExitForNewPane()` docstring claimed its callers
already persisted; that claim was false for these two, and now says what the
caller owes instead.
The watcher's four guards were unreachable by any test. `refreshPaneExits()`
opened with `if (IS_TEST_MODE) return;`, so the read gate, the in-flight
suppression, the generation counter and the empty-read rule could each be
deleted with the whole suite green. The tmux call moves into `readPaneRows()`,
which a test subclass overrides — the shape `runRemoteReconnectTick` already
uses in this file for the same reason — and the test-mode gate moves with it, so
what a test cannot do is spawn a process rather than exercise the bookkeeping.
Each of the four guards now has a test that fails when it is deleted.
The muted status dot turned out to be a specificity fight on three surfaces, not
two. `.tab-status.error` was not excluded, so a session whose agent exited and
whose PTY-exit breaker then tripped lost its red dot to the mute — the state the
browser answers with a "restart it?" confirm, and a needs-you colour by the same
argument that protects the two alert classes. And mobile.css gives a `busy` dot
a 9px size and a green glow with `!important`, while `status` stays `busy` for a
pane whose agent died mid-turn, so a phone rendered a grey dot still wearing the
green halo beside a badge reading "exited". Both measured against the real
stylesheets, both now excluded, and the CSS test reads mobile.css too instead of
being structurally blind to half the problem.
Six comments said things that were not true. Two named the stats collector as
what replaces a restored reading, which is the opposite of the design. The
interval constant argued that 2000 ms keeps a read inside a tick, when the
5000 ms exec timeout means it cannot — which is why the in-flight guard exists.
`MuxSession.discovered` did not say the flag is permanent, though `saveSessions()`
serializes it. The empty-read docstring claimed a distinction that `|| true`
makes impossible. The invariants doc promised more than its drift test delivers.
And CLAUDE.md had no pointer at all, leaving its two hardest prohibitions
("never set `status: 'error'`", "never null the pid") only in the file it is
meant to route people to.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five items, two of which he could only see by running it, plus six smaller
ones. Taking the two blockers first, because both were wrong in ways the
existing tests could not catch.
**Adopting the PTY's rows put the CLI's input line off-screen.** A phone that
took a desktop's 43 rows into a viewport with room for 18 painted an
`.xterm-screen` far taller than its container; xterm's own viewport then had
nothing to scroll, so the bottom of the frame sat below the container with no
gesture able to reach it. Output visible, typing invisible, for as long as the
desktop kept the claim hot. `reconcilePtyGeometry` adopts COLUMNS ONLY now:
width is the axis Ink's wrap and `eraseLines` arithmetic depend on, and keeping
the local row count keeps the composer at the bottom of a viewport that
scrolls. Measured at his geometry — a 360x300 container against a 198x43 pane
now keeps 13 rows, takes 198 columns, paints 202px into a 210px container, and
the input line is inside the box.
**`capture-geometry-retry.browser.test.ts` failed, and CI could not see it**
because the file is in `BROWSER_TEST_GLOBS`. Its premise WAS the clamp —
`getTerminalDimensions()` floored while `fitAddon.fit()` did not — which this
work removes at the source, so it can never hold again at any viewport. The
case survives on its own terms: a pane already drawing at the requested size
must not be replayed. Its premise is now the #464 invariant itself, that the
floored report and the terminal agree, which is a stronger guard because the
clamp coming back fails it here rather than silently restoring the replay loop.
The helper docblock that repeated the old premise is corrected too.
**A session with no pane reported 120x40 and the client adopted it.**
`resize()` writes `_ptyCols`/`_ptyRows` only when `ptyProcess` is set and
nothing seeds them from the spawn geometry, so a dead-pane session still held
the constructor defaults — clicking that tab resized the browser terminal to
120x40 and, on anything narrower, claimed another device owned the pane when
none existed. `Session.ptyGeometry` returns null without a pane, the HTTP route
answers `{}` and the socket sends no frame at all. The raw `ptyCols`/`ptyRows`
getters are deleted rather than left available to be misused again.
**The 40-column floor clipped the pane with nothing able to reach it.** The
affordance keyed on a PTY mismatch, and the floor produces no mismatch — xterm
and the PTY agree throughout, the terminal is simply wider than the box. It
keys on what does not FIT now, MEASURED (`.xterm-screen` against the container,
on the next frame, because the screen takes its width with the render) rather
than derived from cell arithmetic. Measured at 360px: font 24 applies 40
columns and paints 560px, and all 200px of the overhang is reachable.
`.pty-oversized` is renamed `.term-overflows-x`, because after this the old
name describes only one of the two causes.
**"Scroll sideways" did not work on touch for the sessions it targets.**
`touch-action: pan-x` is cancelled before it starts by the `preventDefault()`
`touchstart` calls on every 'content' tap. The terminal's own touchmove handler
pans the container now, with the axis locked once per gesture so a diagonal
cannot pan and scroll at once, and the CSS grants no `touch-action` at all —
handing the browser a pan AS WELL would move the pane twice for one finger on
the taps where that preventDefault does not run. Measured under real touch
dispatch: a 140px swipe reaches `scrollLeft` 140 where it reached 0 before, the
buffer does not move with it, and a vertical swipe still scrolls the scrollback.
Three defects in the above, found while checking it rather than by being told:
- `canPanHorizontally` first tested `scrollWidth > clientWidth` alone, which is
true of a container that is not a scroller — a sideways swipe would have
locked the axis, done nothing, AND suppressed the vertical scroll it should
have been. Gated on the class as well.
- The notice advised scrolling sideways whenever the PTY was wider, including
when it still fitted and nothing scrolled. It is gated on measured overflow,
and on a comparison against the width this container WOULD request rather
than the one it currently holds — once adopted those are equal, so the second
question answers itself false while the condition is still true.
- `_syncTerminalOverflowAffordance` could throw out of `document.getElementById`
before reaching its try block. It runs off every geometry change, so a
cosmetic affordance could have taken the resize down with it.
The smaller items:
- `docs/architecture-invariants.md` no longer explains the equality guard as a
clamp signature; it records what the clamp used to do and why it cannot any
more. Edited by hand — that file is outside the Prettier glob, and letting
Prettier near it rewrote eleven unrelated emphasis markers.
- `throttledResize`'s HTTP fallback reads the reply. It is the path where a
declined resize is least likely to be noticed, because no socket means no
`{"t":"zc"}` frame either.
- The changeset covers the whole release: the geometry work, the queued replay
clear, the renderer watchdog, the body-covering fetch deadline, the WebSocket
output-gap reconcile, the build-generated service-worker precache and
per-build cache key, and the crash-trail hygiene.
- `@xterm/headless` is declared in the root devDependencies instead of being
reached through workspace hoisting.
- The output-gap marker is cleared after any response arrives, not only when
the capture was non-empty: a server that answers with an empty capture HAS
reconciled us, and leaving the marker set refetched on every reconnect.
- `e587d845`'s message claimed a test asserted the failed-load copy against the
built asset. It did not — that assertion lived in a probe deleted with the
other scratch scripts, so the claim was false when it was written. There is a
real test now, and it reads the source rather than `dist/`, because `dist/` is
not committed and a test that skips when it is absent would pass for the wrong
reason in CI.
`Session.ptyGeometry` gets behavioural coverage against the real class in
`session-resize-arbitration.test.ts` rather than a source guard, including the
contrast — a pane that does exist still reports, and still follows a resize —
so "always null" would fail it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copying a paragraph out of a Claude Code or Codex pane puts that pane's own
two-column transcript gutter on the clipboard, so every pasted line arrives
indented. #451 shipped the trailing half of the copy clean and left the leading
half out, because deriving the width from the selection fires on 73% of ordinary
indented text and cannot tell a margin from content.
The width is DECLARED rather than derived. `capabilities.transcriptGutter` on
the CLI registry is a bounded integer; claude and codex each declare 2, measured
on live panes, and no other stock entry declares any, so a CLI whose transcript
layout nobody has measured is never touched. The server publishes the map as
`window.__codemanTranscriptGutter`, built by filtering `enabledClis()` on the
capability rather than by listing ids, and `_activeCliGutterColumns()` looks the
active session's mode up in it. The copy path reads no terminal buffer at all.
The declared width is a CEILING, not the answer: `clean()` strips the lesser of
it and the run every selected line shares. A block can therefore only shift as a
unit, the structure inside a selection survives by construction, and a selection
reaching column 0 loses nothing. That is what keeps a `git log` body at its own
four-space indent inside an agent's two-column gutter.
Codex was measured separately, because it renders nothing like Claude: it draws
boxes narrower than the pane and pushes its transcript into ordinary scrollback.
On a live 0.154.0 answer its `•`/`›`/`⚠` markers sit in the gutter, prose
continuations sit at 2, and a nested YAML block the model wrote rendered at
2/4/6/8 for its own 0/2/4/6. Replayed at 100, 120, 160, 198, 235 and 282 columns
its indents were 0, 2, 4, 6 and 8 at every one, never 1. Copying that YAML out
of a live Codex pane now yields 0/2/4/6: gutter gone, nesting intact.
Two derived versions were built and measured first, and both are recorded in the
code because both looked correct:
- Painted trailing padding — a full-screen TUI writes real spaces across the
unused part of a row, a shell leaves them never-written for xterm to trim —
has no false positives and never over-stripped. It is also a function of pane
WIDTH: the padding exists only while a rendered line stops short of the CLI's
own layout width, and Claude's prose wraps to fill it. Dragging the same two
prose rows of one live transcript at five window sizes, the share of padded
rows ran 44%, 6%, 6%, 7% and 87% at 123, 160, 198, 235 and 298 columns, so the
strip silently did nothing at every ordinary size while a corpus captured
entirely at 282 columns said it worked.
- Taking the narrowest indent on the rows around the selection fires at every
width and over-strips about 1% of selections, because a file listing inside
the transcript can be the narrowest thing on screen.
Measured over 1,392,281 selections — every 1, 2, 3, 5, 10 and 20-row window of
real Claude screens replayed from live PTY streams at 100, 120, 160, 198, 235
and 282 columns — the declared width over-strips none, breaks no relative indent
and alters no text, and serves 100% of the selections whose own indent covers
the gutter. Verified end to end in a browser with a real mouse drag and a real
Ctrl+C: Claude and Codex panes paste flush at 123, 198 and 298 columns, a shell
pane is untouched at every one.
The strip sits behind `copyStripMargin` (App Settings, Selection & clipboard),
per-device and default ON: a display key, absent from the .strict()
SettingsUpdateSchema, read as `!== false` because the desktop branch of
getDefaultSettings() returns {}. The toggle is checked before the map.
Two review findings from #451, handled:
- The mid-row flag governs ONE line now. `range.start.x > 0` excludes only the
first selected line, the one whose margin the mousedown genuinely cut off, so
the same three rows no longer produce three different clipboard results.
- The reversed-drag finding does not reproduce on the pinned xterm.
`getSelectionPosition()` reads `_selectionService.selectionStart`, whose
getter returns `SelectionModel.finalSelectionStart`, and that swaps the pair
when `areSelectionValuesReversed()` says so. A real upward mouse drag through
chromium against xterm 6.0 reports the same range as the downward drag.
`_normalisedSelectionRange()` keeps the ordering as a guard, because the model
one layer down exposes the unnormalised fields under the same two names.
Tests: test/terminal-copy-clean.test.ts (64, up from 31), plus the injected
script stripped in test/server-index-title.test.ts. Every guard is pinned:
removing any one of seven reds at least one test, including declaring the wrong
gutter width. Full suite green, 7,861 passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A third pass in a real browser, at the widths this app actually renders at.
The notice a failed history load writes into the blanked pane was one
70-character sentence. At 430px that exactly filled the line; at 320px it
wrapped and left a lone '.' on a line of its own. The floor this app will
render at is 40 columns — reachable today by raising the font on a phone — so
the notice is three lines now, none over 25 columns, one fact each: what
failed, that the session is still alive, and what to do.
It says RELOAD rather than "reopen the tab" because `selectSession`
early-returns when the session is already active, so clicking the tab you are
already on retries nothing. The earlier wording named no next step at all,
which left a mostly-empty terminal and no way out of it.
CLAUDE.md no longer cites "758px reachable to the right" as evidence: that
figure is a property of the test content, not of the fix, and the file's value
is that a reader can trust a claim without re-deriving it. What is pinned
instead is the invariant that survives any content — the full pane width is
reachable, and removing the class returns scrollLeft to 0, so a resolved
mismatch cannot leave the pane parked off-screen.
Verified at 430, 360 and 320px against the shipped bundle, with the test
asserting the built asset carries the copy so an edit that never reached the
build fails rather than passing on the source's wording.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The gate caught fourteen failures the focused tests could not: every harness
that builds a partial app out of cherry-picked mixin methods, and every source
guard that named `fitAddon.fit()` by hand.
Most are wiring — `syncTerminalGeometry`, `_refitAfterCellSizeChange` and
`_resizeTerminalTo` added to the fakes so the real chain runs rather than a
stub of it. `file-browser-search` is the one that shows why it matters: without
the method on the fake, selectSession's unconditional call threw into its own
catch and every later assertion in the file measured a load that never
happened.
Two are not wiring.
`detached-session-pane-sizing` pinned the behaviour this change deliberately
reverses. It asserted the LOCAL fit still runs for a session owned by its own
window — "withhold the send, never the reflow" — so the assertion is restated
rather than patched, with the reason beside it and in the file's docblock: a
reflow the PTY is never told about leaves this xterm rendering a CLI's frames
against a shape that does not exist, and the popup that owns the pane is
drawing for its own width regardless. The old rule bought a garbled frame, not
a correct one.
`mobile-prompt-composer` sliced `_cleanupSessionData` as a fixed 1200-character
window, so the assertion depended on how much unrelated code sat above the line
it cared about. It reads the whole method now.
`terminal-scroll-intent` records `syncTerminalGeometry` rather than `fit`,
under its own name: recording a bare fit there would name the very thing the
subject was changed to stop doing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Issue #464, "text gets muffled sometimes, in both TUI default and fullscreen".
The screenshot is not a dropped frame or a frozen renderer — it is arithmetic.
Claude Code's TUI wraps its frame at the width the PTY reported and erases the
previous frame by walking the cursor up the rows it believes that frame took. A
browser terminal of a different width makes each logical line occupy more
physical rows than Ink counted, so `eraseLines(n)` clears too few and the new
frame paints over rows nothing erased: doubled lines, and short tool summaries
sitting inside longer prose rows with the prose's tail still visible.
Reproduced against this repo's own xterm before changing anything — a 120-column
PTY against a 62-column terminal renders every wrapped line twice. `test/
terminal-pty-geometry.test.ts` pins that, and pins the clean render at matching
widths beside it, so the assertion cannot be satisfied by code that fixes
nothing.
Four ways the two drifted apart, none of them observable from either end:
1. `fitAddon.fit()` resizes xterm to `proposeDimensions()` RAW while every
server-facing path reported those floored at 40x10. Measured in Chrome at
430px: font size 44 proposed 13 columns, the server was told 40, and xterm
stayed at 13. Three call sites each did their own fit-then-floor, and two
re-read the proposal after the fit — `_shrinkPaddingToFit()` runs exactly
there, so the container had moved.
2. `throttledResize` (keyboard up) and `sendResize` (session detached into its
own window) reflowed locally and withheld only the SIGWINCH. That is the one
combination that cannot be right: a reflow nothing is rendering for buys
nothing and costs correctness. Both now withhold everything, and the
keyboard's settle timer still sends the one resize that stops the PTY going
stale.
3. `setFontSize`/`setFontFamily`/`setFontWeight` move the cell size — a geometry
change — and told the server nothing at all, so raising the font on a phone
left the CLI wrapping at the old column count.
4. `Session.resize` DECLINES a small-viewport request while a desktop connection
holds an active sizing claim, and said nothing, because resize was write-only.
`syncTerminalGeometry()` is now the one function that may change the terminal's
size: it fits, floors and applies as a single step, so the numbers xterm holds
are the numbers the server is told. A test sweeps every module for a bare
`fit()` on the main terminal, and finds exactly one — the owner's own.
For (4) the client cannot win, so it is told the truth instead: both transports
answer a resize with `session.ptyCols`/`ptyRows` (`{"t":"zc"}` on the socket,
the body of the resize POST) and `_onPtyGeometryReport` adopts them. A terminal
that keeps a shape the PTY refused does not render "too narrow", it renders
garbled. Adopting can leave the pane wider than the screen and the container is
`overflow: hidden`, so `.pty-oversized` grants horizontal reach for exactly as
long as the mismatch lasts: correct-and-reachable beats correct-and-clipped
beats garbled. That rule sets both overflow axes and its own `touch-action`
because mobile.css loads later and sets `.terminal-container { overflow:
visible; touch-action: none }` — a bare `overflow-x` would leave overflow-y
computing to `auto` and hand the browser a vertical scroll container the
terminal's touch handler knows nothing about.
Verified in Chrome at 430px against a live server, with a desktop client holding
the claim: the phone adopts 198x43, gets `overflow-x: auto` / `overflow-y:
hidden` / `touch-action: pan-x`, 758px of reach to the right, and keeps its own
vertical scrolling. The pre-fix build was measured in the same harness for the
control.
Two things this deliberately does not do. It does not change who owns the pane
size — the desktop still wins, and `_startMobileResizeRetry` still takes it back
once that goes idle. And `throttledResize` still holds the PTY's shape for the
whole keyboard animation rather than sending a SIGWINCH per step; that decision
predates this and was not re-tested here.
Also in this commit, Ark0N's third-pass review items on #431:
- The response viewer's byte-buffer fallback and `_onSessionClearTerminal` both
used the no-param `/terminal` form, capped only by `terminalBufferMaxBytes`
(32MB) — the largest body the frontend asks for anywhere. One carried no
deadline at all and the other got the 15s tail budget. Both now take the
full-history budget.
- A `?full=1` capture that outruns its deadline falls back to the bounded tail.
The pane is blanked before that fetch, so an abort used to leave a black
rectangle, discard the queued live output and never reach `_connectWs`. A
failed load now still opens the socket, says one dim line where the content
would have been, and clears the tab's spinner — which nothing did, so a failed
select left `aria-busy="true"` set forever.
- `_wsOutputGapSession` is cleared at the repaint that settles it, not in a
`finally` that also ran on the catch. A reconcile that threw, or hit the new
deadline — the flaky link the marker exists for — dropped the gap with nothing
to retry it. `ws.onopen` no longer clears it up front either.
- The replay-clear invariant is pinned in the gate, which is the drift this PR
exists to fix: `_resetTerminalForReplay` must be a queued write and nothing
else, and no module may blank the terminal with a `clear()+reset()` pair.
- `DIAG_ENTRY_MAX_CHARS` replaces the hardcoded 300, bound through a local
first: `CodemanDiag?.x` still throws a ReferenceError when the identifier was
never declared, and that is the one function in the app that must not throw.
- panels-ui's two kill-all clears route through the same helper, and the
xterm-version guard's comment says "resolved lockfile version" rather than
"dependency RANGE", which is what it has pinned since the last round.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four changes the maintainer asked for on Ark0N/Codeman#446 before merging.
The pane-exit watcher stays always-on, but a tick now costs nothing when there
is nothing to observe. `hasObservablePaneSession()` skips the tmux exec while
every session on the manager is one of the shapes `Session.paneExitApplies`
already forces to UNKNOWN: a remote SSH session (its local pane holds the ssh
client), a docker case (a `docker exec` into the container's own tmux), and a
record rebuilt from the socket (no provenance at all). The timer is untouched.
Skipping retracts nothing, for the same reason a failed read does not: the map
still holds the last real reading, and every path that puts a new command in a
pane calls `clearPaneExit()` itself. The two copies of that rule are pinned
against each other in `test/session-pane-exit.test.ts`, because drift between
them is silent in both directions.
`DEFAULT_PANE_EXIT_INTERVAL_MS` was already a constant beside the stats and
remote-reconnect intervals; its comment now says why the watcher owns its own
cadence and why the number is what it is.
The never-default-an-absent-status rule is written where `PaneExit` is declared.
It names `status ?? 0` as the thing never to write, and says that an agent the
OOM killer took would otherwise read as a user typing `/exit` — which is what
absent-stays-absent keeps a later clean-exit sweep away from. Nothing fails when
somebody adds that `??`, which is why the sentence is there rather than a test.
Checking the dot's specificity found a second fight, and it was losing. On the
tab strip the alert rules win as intended: a session that exits with a
permission dialog pending still renders red, and yellow for an idle alert. On
the rich vertical tab rail they did not — that rail's own `tab-state-*` dot
rules are (0,9,1) against the strip's mute at (0,5,0), so an exited session
there kept a full green dot AND the working halo beside a badge reading
"exited". The rail twin matches that specificity exactly and therefore must stay
below those rules in source order; it clears the halo as well, which the strip's
rule never had to think about.
`test/session-pane-exit-ui.test.ts` now resolves the real stylesheet in jsdom
rather than matching selector text: postcss collects every rule that paints
`.tab-status`, a real engine decides, and the tests read back the answer. Two
mutations were run against it to prove it has teeth — dropping the hand-written
alert exclusions fails three cases, and moving the rail twin above the state
rules fails one.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Codex states background work too, and it says so in a different place. Claude
writes `· 1 monitor ·` on the last row of the screen; Codex pins
`1 background terminal running · /ps to view · /stop to close` ABOVE its
composer, which puts that row third from the bottom once the status line and
the composer are counted.
So how far up the screen to look is now per-CLI data as well:
`capabilities.workDetect.watchingLines`, bounded to 1..8 by the schema, and
defaulting to Claude's two. That bound is the point. The window is half the
injection guard, since every row it adds is another row the agent itself may be
able to write, and the label is what silences an idle alert. The other half is
the anchor, and Codex's is ` · /ps to view`: chrome naming a slash command only
the CLI can offer, so a session that writes "I left 1 background terminal
running for you" into its own output matches nothing.
Measured against a live codex-cli 0.154.0 pane rather than read out of a
binary. The row appears when the terminal starts, follows the composer down as
the conversation grows, and is gone after `/stop`. Verified end to end on an
isolated beta: the session payload carried `watching: "1 background terminal"`
and the badge rendered with it, and both cleared when the terminal stopped. The
fixtures in the tests are that capture verbatim.
Codex has no hook signals, so no idle prompt and no false NEEDS YOU row: for a
Codex session this is the badge alone, which is the case the maintainer said a
registry field could cover and a hook never could. Cross-CLI tests pin that
neither pattern fires on the other's screen, and that a CLI declaring nothing
still reports nothing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review fixes. Two of these are defects in the previous commit.
1. The fetch deadline only covered time-to-headers. `await fetch()` settles on
response headers, so clearing the abort timer in a finally around it left the
body — the multi-megabyte `?full=1` capture the deadline exists for —
completely unbounded; it only ever bounded a server that accepts a connection
and never replies. Measured against a server that sends headers immediately
and stalls the body 4s under a 1s deadline: fetch resolved at 30ms, timer
cleared there, body completed at 4026ms unaborted. Now the body is read
inside `_fetchTerminalCapture`, which returns {json, headers, headersAt} —
headers because two callers read server-timing, headersAt because those same
callers measure header-vs-body time and can no longer observe that moment.
`_terminalCaptureInflight` is scoped the same way, so a body still streaming
counts toward a capture starting beside it. Same test now aborts at 1005ms.
2. The precache could never be hit, and the previous commit made that expensive
rather than free. `renderIndexHtml` runs `cacheBustAssets`, which appends
`?v=<mtime>` to every same-origin .js/.css reference INCLUDING content-hashed
names — confirmed against a running instance:
`vendor/xterm-zerolag-input.6fee72f2.js?v=1789402869101`. `caches.match` is
query-sensitive, so entries keyed on the bare hashed path were unreachable;
deriving the list from the manifest turned cheap 404s into ~1.3MB downloaded
at every install that nothing could read back, once per deploy now that
CACHE_NAME rotates. The fallback match takes `{ ignoreSearch: true }`, which
also lets runtime-cached entries survive an mtime change.
3. `_wsOutputGapSession` was only cleared in ws.onopen, so paths that already
repaint the buffer left it set and the socket replayed everything a second
time. `selectSession` loads the buffer and only THEN calls `_connectWs`, so
neither the _isLoadingBuffer nor the _terminalRefreshOwner guard applied.
`_markTerminalBufferReconciled()` is now called from _onSessionNeedsRefresh's
finally, from selectSession after its load, and from _cleanupSessionData.
The scope claim was also wrong and is corrected in the comment: when the
network drops, SSE drops with it and handleInit's keepTerminal branch already
reconciles. The genuinely uncovered case is the WS dying while SSE stays up,
where _onSSETerminal discards SSE terminal frames until _wsReady flips in
onclose — up to the ping+pong window of output nothing writes.
4. CLAUDE.md said "all of them measured rather than reasoned", which the PR's
own "not verified" section contradicted. Split explicitly: the replay race is
measured, the watchdog mechanism is verified against xterm 6.0.0 under jsdom
(field path resolves, a forced stale handle makes refreshRows a no-op, the
kick schedules a fresh frame), and the iOS rAF-discard premise is reasoned
and still wants a device. Adds the two missing entries — the WebSocket
reconcile and the sw.js/build.mjs "keep these in sync or the build throws"
contract.
Also: test/xterm-private-api.test.ts pins the RESOLVED lockfile version instead
of the declared `^6.0.0` range, which was the wrong assertion in both directions
— a real upgrade to 6.4.0 can rename a private field while resolving inside the
range, and an innocuous range edit failed while changing nothing installed. And
test/sw-precache-manifest.test.ts now parses HASHABLE out of scripts/build.mjs
rather than hand-copying it, which was the same drift this PR exists to fix; the
parse is guarded against silently matching nothing.
The deadline fix has a behavioural test against a real socket plus a source
guard asserting `await res.json()` precedes the finally — verified to fail when
the helper is reverted to the old shape, so it is not vacuous.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four ways the terminal can silently stop being correct — in each case the
buffer keeps updating, nothing throws, and the only recourse is a reload.
1. Renderer freeze after backgrounding. iOS DISCARDS scheduled rAF callbacks
when a PWA backgrounds, and xterm's RenderDebouncer only clears its
`_animationFrame` handle from inside that callback — so one drop leaves it
permanently set and every later refresh() early-returns. Parsing is
decoupled from rendering, so bytes keep filling the buffer correctly while
nothing paints. Codeman has exactly ONE xterm for the whole page load, so a
single backgrounding wedges it until a reload. Adds a 2s liveness poll and
`_kickRenderer()`, which does what the dropped `_innerRefresh` would have.
2. Replay clears raced live output. xterm's write() is async-queued while
reset() is synchronous and, per upstream, "does not clear input buffers and
does not reset the parser" — so bytes queued before a reset are parsed after
it and fuse into the snapshot. Verified against the real xterm 6 here:
write('p8'); reset(); write('rmissions') renders "p8rmissions". The main
path was already safe via a queued erase; the needsRefresh and clearTerminal
paths were not. All three now share one queued `\x1bc` (RIS), which unlike
3J/H/2J also resets modes, charsets, scroll regions and SGR state.
3. Output lost on WebSocket reconnect. Input frames carry seq+cid and are
delivered exactly once; output frames carry nothing. ws.onopen re-sends dims
and flushes queued input, and needsRefresh only fires on external-CLI
startup and SSE backpressure drain — never on reconnect. Output produced
while offline was simply absent afterwards. Interim fix: reaching onclose
means the drop was unintentional, so the session is marked and the next open
reconciles from the server buffer. Sequencing output is the follow-up.
4. Terminal captures had no deadline. No AbortController anywhere in the
frontend, including `?full=1`, which the code itself calls "unbounded-ish
work: at the default history limit it can be megabytes". Adds a budget that
scales with full-vs-tail and with captures in flight, degrading to a plain
fetch where AbortController is missing.
Also: the service-worker precache was dead — the build content-hashes assets
but sw.js listed pre-hash names, so 15 of 23 entries 404'd (verified against a
running instance) and cache.add().catch() hid it. Offline still worked via
runtime caching, but CACHE_NAME was a constant so activate's cleanup never
deleted anything and every past release's assets accumulated. Both are now
derived from the build manifest. Crash-trail entries are flattened and capped,
since they are joined with \n into one value and one call site interpolates a
server-controlled WS close reason.
The watchdog reads xterm privates — there is no public API. Every access is
optional-chained so a shape change degrades to a no-op. `_renderService` only
exists after open(), which needs a real DOM, so the gate cannot assert the
field path; test/xterm-private-api.test.ts pins the dependency range instead.
Tests: 23 new (terminal-resilience, sw-precache-manifest, xterm-private-api),
all pure/static so they run in the gate, which excludes the mobile suite. One
static source guard in history-truncation-notice updated for the renamed call;
the behaviour it pins is unchanged.
Not verified: no browser available, so no runtime reproduction of the freeze
and no real-device test of the reconnect path. Both warrant a device pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The badge alone left the row in NEEDS YOU, which is the thing the issue was
about. The fix is the alert that does not fire.
An idle prompt from a session that is watching its own background work now
opens ALREADY acknowledged. `hook-event-routes` passes `Session.watching` to
`notePrompt()`, which sets `acknowledgedAt` and records why in a new
`acknowledgedReason`. Nothing new suppresses anything: `acknowledge()` has
always meant "the alert this prompt armed is spent", and the prompt itself
stays pending, answerable and available as Read My Mind context. A wrong label
therefore costs a card that does not blink, never an alert that was never
created.
Every surface follows from that. The broadcast carries the reason, so a live
page declines to arm the tab alert and raises no desktop notification. The push
is skipped, since a false alarm is hardest to ignore on a phone. A reloading
page reads `acknowledgedAt` in `seedApprovals()`, which it already did. And
`classifySession()` now reads it too, which is a pre-existing bug fixed here:
acknowledging on one device cleared the alert everywhere except `codeman tui`.
It re-arms for free, because the next idle prompt supersedes the item and is
built fresh. Only `idle` is eligible, so a dialog that blocks the agent still
goes red whatever else it started.
The label is pane-derived and therefore prompt-injectable, so it is now read
from the last two rows of the screen only, with Claude's pattern anchored on
the `·` its footer joins items with, ANSI-stripped and length-capped at the
source. An agent that prints `· 1 monitor ·` into its own output finds no
match.
Verified on an isolated beta: a session that armed a monitor took its idle
prompt acknowledged with no alert on any surface, wore the badge, and showed
"quiet, watching 1 monitor" on its still-answerable card; the same session with
the monitor killed alerted normally on the next prompt. `test/watching-no-alert.test.ts`
pins both directions across all four surfaces.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A single pin that failed its transcript gate returned the options untouched,
so `resumeSessionId` fell back to `_resumeSessionId` — undefined for an
ordinary session — and the renderer emitted the bare
`claude --dangerously-skip-permissions --session-id "<this.id>"`. Every
session prompted before its first `/clear` owns a transcript under that id,
so the dropped pin handed back exactly the refusal this branch removes, with
no `||` branch to catch it. It was also a regression against master on the
`restartCli()` path, which pinned `_claudeSessionId ?? this.id` and, since the
constructor seeds that field, could never land unpinned.
The pin now walks three candidates in priority order — the conversation
chain's tail, the launch seed, then the session's own id — and takes the first
one a transcript backs. A candidate that misses is passed over rather than
ending the walk.
Falling off the end pins nothing, which also settles the second half of the
problem: the old code skipped the transcript check whenever the pin was the
session's own id, so a genuinely new pane rendered the two-branch form after
all. That costs a brand-new session claude's "No conversation found" line in
its scrollback, and `wrapWithNice()` prefixes only the first branch of the
rendered `a || b`, so the branch that actually runs loses its priority for the
life of the session. With no transcript anywhere the bare `--session-id` is
the correct command, so the comment claiming an unchanged shape is now true.
The transcript lookup reads the server process's own `CLAUDE_CONFIG_DIR` when
a session declares none. A pane inherits the server environment through tmux,
so on an install that exports it the CLI writes its transcripts there and
every lookup under `~/.claude` was a false negative — which under the old code
meant the colliding command. `claudeCredentialsPath()` and
`realClaudeConfigDir()` resolve the same directory the same way. The header
sentence calling a skipped resume "the safe direction" described the opposite
of what happens at this call site, and says so now.
The create-path fallback writes `_resumeSessionId` alongside the create
options. That branch leaves `isRestored` false, so `_claudeSessionId` is
recomputed from the launch fields and settled on `this.id` while the CLI
resumed the chain tail; the response viewer, Read My Mind and the unified-list
alias map read that field until the next first-hand hook.
Four new tests: a chain tail with no transcript while the session id has one,
no transcript anywhere, the create path's alias, and the process-env lookup.
All four fail against the previous commit. Two existing tests move with the
gate — the guess-refusal test now backs the session's own id, and the
custom-model restart test gives its working pane the transcript that makes
`--session-id` collide in the first place, alongside a new one pinning the
no-transcript case.
CLAUDE.md described the pin as a `restartCli()`-only thing sourced from the
live conversation id. All three halves of that moved here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An agent that arms a monitor, backgrounds a shell or hands a task to a
cloud session is told to end its turn. The pane then falls quiet, Claude
Code's idle_prompt notification arrives a minute later, and every surface
files the session under NEEDS YOU with nothing for a human to answer.
Claude states what it is still running on the last row of its screen
(`⏵⏵ bypass permissions on · 1 monitor · ← for agents`). That row is now
`capabilities.workDetect.watchingLine` in the CLI registry, guarded by
compileVersionRegex() like every other config regex, and the idle probe
reads it off the capture it already takes: `watchingLabel()` in
session-activity.ts searches the last five lines only, so a session that
PRINTS "1 monitor" is not mistaken for one running it.
The label lands on Session.watching and rides toLightDetailedState() out
to every surface. The phone overview, the desktop home rail and the rich
sidebar rows wear it as a `watching` badge in the accent colour, beside
the state pill and never in place of it: an agent can arm a monitor and
ask a question in the same breath, and only the pill says which.
Verified end to end against a throwaway session on an isolated beta
instance: the payload carried `watching: "1 monitor"` once the turn
ended, the badge rendered next to a yellow `waiting` pill, and both
cleared when the monitor died.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A CLI that launches with `--session-id <id>` refuses an id that is already in
use (claude: `Error: Session ID ... is already in use.`), and every session
whose agent has been prompted owns a transcript under that id. The dead-pane
respawn in `_setupOrAttachMuxSession()` passed the bare launch line, so
recovering such a session relaunched a CLI that died on startup, the pane went
dead again at once, and the conversation was stranded behind a tab that looked
merely idle.
`restartCli()` has pinned a resume id against this since the custom-model work,
and its comment states the assumption that made the other path look safe:
"Unlike the dead-pane respawn, this one kills a WORKING pane whose conversation
already has a transcript". A pane whose agent exited has a transcript too.
Both relaunch paths now build options through
`_buildRespawnPaneOptionsWithResumePin()`, and so does the create-path fallback
after a failed respawn, which otherwise met the same refusal that made it the
fallback. Four gates guard the pin, each standing for a way of resuming the
WRONG conversation or of making a working relaunch fail.
A remote or docker session is never pinned. Unlike `restartCli()`, whose route
refuses both, the dead-pane respawn is reached by every session shape. Their
pane commands already render a self-healing `--session-id || --resume`, and
both flip to resume-first once the resume id differs; the conversation lives on
the far side, so a local id resolves to nothing there and the `--session-id`
fallback then collides with the transcript the far side does hold.
The id comes from the conversation CHAIN rather than `_claudeSessionId`, which
also holds history-correlated guesses keyed on the working directory.
`_recordClaudeSessionInChain()` refuses those so they cannot "write a foreign
conversation into this pane's permanent record", and launching from one is
worse than the display bug that rule prevents. The chain tail also outranks the
launch seed, which is written once at construction and never moves off a
`/clear`.
A pin no transcript backs is dropped, because the fallback branch keeps
`--session-id <this.id>` and would collide. A synthetic `restored-<fragment>`
id from socket discovery is dropped too, and logged: it fails claude's `uuid`
token pattern, so the renderer would emit the unpinned command while the caller
believed otherwise.
Tests cover each gate and the rendered command. Four of them fail against the
unfixed source; the remote and docker ones were separately checked against a
build with only that guard removed, since they pass on master for the wrong
reason.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Same guard as Start-Codeman.sh's own (docs/docker-self-update.md-adjacent
incident, 2026-09-21): docker-compose.yaml hard-codes `name: codeman`, so a
second checkout run without COMPOSE_PROJECT_NAME resolves to the SAME
Compose project as any other checkout on the host. It has to live here too,
not just in Start-Codeman.sh: this script's own --no-cache build and its
`down`/`down --volumes` both run BEFORE the handoff at the bottom of the
file, so Start-Codeman.sh's copy of the guard would only fire after this
script's own destructive calls already ran — and its default `down
--volumes` is more destructive than Start-Codeman.sh's own targeted
refresh, clearing every named volume the resolved project has.
Also fixes a real bug the same guard shipped with: under `set -o pipefail`,
`grep -v` legitimately exits 1 when nothing survives the filter (the
ordinary, no-collision case), and without `|| true` on the pipeline that
non-zero status propagates through the command substitution and `set -e`
aborts the WHOLE script at the guard — every time, collision or not. Caught
only by actually executing the guard end-to-end against a stub `docker`
(the existing smoke-test harness), never by a static text/regex check on
the source; the stub's `config --format json` response was also fixed to
pretty-print like real Compose does, since a compact one-liner silently
resolved project_name to empty and exercised neither script's guard the
way production output does.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
The tab now reads "exited (137)" beside the session name, drawn from the
`paneExit` field the server publishes. `applyPaneExitBadge()` owns the DOM
work, called from the incremental render path — the only path a live session
ever takes, since going from live to exited adds and removes no tab and so
never reaches the full rebuild.
An unknown answer draws nothing. A death tmux could not explain reads "exited"
with no number rather than "exited (0)", so an unexplained death and a clean
exit do not look alike. A signal death reads "exited (signal 9)".
The badge carries `data-i18n-skip`, like the status pills: it is generated
text, `i18n.js` walks inserted content, and a dictionary entry added later
would fight the renderer, whose in-place comparison is against English.
The tab also carries a `tab-agent-exited` class that mutes the status dot. That
dot is drawn from `status`, which stays `idle` or `busy` for an exited pane as
the issue requires, so without this a green or pulsing dot sits beside a badge
saying the agent is gone — the first thing a tester asked about. `status`
itself is untouched, so this is a rendering rule only. The CSS excludes the two
alert classes by hand, following the convention the rich-rail dot rules
document: a dot turning red or yellow because a session is blocked on a human
outranks "the agent exited".
The tab keeps its click behavior. X still closes it, and nothing here closes,
sweeps or restarts anything.
`docs/architecture-invariants.md` gains the mechanism under "Session data and
lifecycle", where every comparable one already lives: what the tri-state means,
the four shapes it is absent for, why the watcher cannot ride the stats
collector, why an absent `#{pane_dead_status}` is not 0, and the three things
that must never happen to an exited pane.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mux layer now knows a pane's agent has exited. This puts it on the session
record, where the board and, later, the reboot restore can see it.
`SessionState.paneExit` carries `{ status?, signal?, at }` and rides the
existing `session:updated` broadcast through `toState()`. No new SSE event. The
server pulls each answer from `mux.getPaneExit()` rather than off a broadcast
payload, so the raw reading never reaches a browser: for a remote or docker
session that reading is the death of an ssh client or a `docker exec`, not of
the agent.
The field is tri-state, and the third state is its absence: `undefined` means
Codeman does not know, and it never reads as alive. `Session.setPaneExit()`
forces that unknown for every shape a dead local pane does not describe. A
direct-PTY session owns no pane. A remote SSH session's local pane holds the
ssh client, whose death means a transport drop OR an exit, which is the
ambiguity PR #355 settled by not guessing. A docker case's local pane holds a
`docker exec` into the container's own tmux. And a session rebuilt from the
socket has no provenance at all: `reconcileSessions()` gives it a synthetic
`restored-<fragment>` id that matches no `state.json` entry, so a remote
session rediscovered after `mux-sessions.json` was lost arrives with no
`remote` field and looks local — `MuxSession.discovered` marks it, and absent
metadata there counts as unproven rather than as proof. The scoping lives on
`Session` rather than in `TmuxManager` so there is one copy of the rule.
`status` and `pid` are untouched. `status: 'error'` belongs to the PTY-exit
circuit breaker and makes the browser offer a restart, and a null `pid` is what
makes the browser re-attach and launch a fresh CLI. A reading that repeats the
previous answer writes nothing and broadcasts nothing.
An unknown answer never reads as alive, but a stale KNOWN one would keep
reading as exited, so `clearPaneExitForNewPane()` retracts it on every path
that puts a new command in the pane: the start/attach path, the `restartCli()`
relaunch behind a custom-model switch, and the remote reattach. Without the
second of those, switching an endpoint on an exited session launched a new
command and then persisted and broadcast the old exit straight back onto it.
`toState()` is also what `state.json` persists, so the record survives a
reboot, which is the only thing that does: a reboot takes the tmux server, and
with it every live signal and every `mux-sessions.json` entry. Nothing reads it
there yet — making the restore refuse such a session is a behavior change that
belongs with the part that closes them. Recovery threads the saved value back
through the constructor so the first persist after boot cannot blank it.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Codeman creates every tmux pane with `remain-on-exit on`. When the agent exits,
tmux keeps the pane, the tmux session, and the `tmux attach-session` process
Codeman records as the session's pid, so no PTY exit handler fires and nothing
writes the exit down. tmux itself knows: it marks the pane dead and reports the
exit status. This reads that.
`PANE_LIST_FORMAT` gains `#{pane_dead}`, `#{pane_dead_status}` and
`#{pane_dead_signal}`, and `startPaneExitWatcher()` refreshes a
muxName-to-observation map from ONE batched `tmux list-panes -a` per tick. Boot
reconciliation already ran that same call, so it now fills the map too and
recovery starts with a reading.
The watcher owns its own interval rather than riding `startStatsCollection()`,
which the issue suggested. That collector is armed when a browser opens the
Monitor panel and DISARMED when it closes it, and boot skips it entirely unless
recovery found a live session, so a session created on a freshly booted server
would publish nothing and one browser could turn detection off for every other.
Measured on an isolated instance: a dead pane with status 0 reported nothing
until `POST /api/mux-sessions/stats/start` was called by hand. It is still one
batched read per tick; only the timer changed.
Three rules keep a positive answer trustworthy. A session answers only when
tmux listed exactly one pane for it, because Codeman never splits a pane and a
session the user split by hand has none that speaks for the agent. A pane
answers only when `#{pane_dead}` said 1 or 0, because an empty field is a tmux
that did not answer. An absent status stays absent rather than becoming 0:
measured on tmux 3.2a, a SIGKILLed pane reports neither a status nor a signal,
and calling that a clean exit would be wrong in the direction that matters.
Two guards stop a slow read undoing a fast one. `EXEC_TIMEOUT_MS` is 5000 ms
against a 2000 ms interval, so a read can outlive two ticks: one already in
flight suppresses the next, and a generation counter that every
`clearPaneExit()` bumps discards a read that started before a respawn or a
kill. An observation also carries its pane pid, so a second command in the same
pane that exits the same way starts a new timestamp rather than inheriting the
first death's.
A non-empty read of `list-panes -a` is authoritative for the whole socket, so
sessions missing from it are pruned, which also bounds the map as tmux sessions
come and go outside `killSession()`. A failed or empty read retracts nothing.
The manager reports the raw pane reading and applies no session-shape scoping,
because the remote-reconnect watcher beside it needs exactly that raw reading.
`parsePaneList` becomes `parsePaneRows`, returning one row per pane instead of
a name-to-pid map; reconciliation builds its map from the rows. The parser's
existing cases carry over unchanged, including the launchd/systemd literal-tab
regression from PR #71.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three blockers, all fixed and verified by actually running the script (not
just string-matching it):
1. `exec "$script_dir/Start-Codeman.sh"` failed EACCES/exit 126 on every
checkout, since Start-Codeman.sh is committed non-executable (100644) —
the same fact my own second commit on this branch established. Fixed to
`exec bash "$script_dir/Start-Codeman.sh"`.
2. `down` ran before `build --no-cache`, so Codeman and every session it was
running were offline for the entire rebuild, and a build failure left the
stack down with nothing to bring it back — the exact ordering mistake
Start-Codeman.sh's own "Build BEFORE taking the stack down" comment exists
to prevent. Reordered to build, then down, then hand off.
3. The default path could throw the rebuild away: codeman-node-modules/
codeman-dist only re-seed from the image while EMPTY, Start-Codeman.sh
only clears them when it detects the checkout's HEAD or package-lock.json
moved, and neither condition is true for the Dockerfile-only change this
script exists for — so a plain `bash docker/Update-Codeman.sh` rebuilt an
image whose fresh node_modules/dist then sat unused behind the old
volumes. Made clearing them the default; `--keep-volumes` opts out
(replaces the old `--volumes`/`-v` flag, which is no longer needed since
clearing is now the default).
Smaller items from the same review, also fixed:
- The --no-cache build now derives PUID/PGID from CODEMAN_APPDATA_PATH's
owner first, via the identical owner_of() helper Start-Codeman.sh uses
(parity-tested) — without it, the build used Compose's default 1000:1000
regardless of the real appdata owner (99:100 on the Unraid layout
docker/README.md documents), and Start-Codeman.sh's own correctly-PUID'd
build during the handoff would then rebuild those layers anyway, so the
--no-cache image never actually shipped.
- docker/README.md's "rebuilds ... only when it detects ... moved" wrongly
described BOTH the rebuild and the volume-clearing as conditional;
Start-Codeman.sh rebuilds on every start, only the volume-clearing is
conditional. Corrected, and reworded around the new default.
- --help/-h now prints usage and exits 0 instead of falling into the
unrecognised-argument branch.
- "the ONLY named volumes this stack declares" now says docker-compose.yaml
specifically, since a docker-compose.override.yml could add more.
New tests: PUID/PGID derivation parity with Start-Codeman.sh's owner_of(),
--help handling, and — the one that actually catches blocker #1, which five
source-string-matching tests did not — a real end-to-end smoke test: a
synthetic deployment, a stub `docker` on PATH logging every invocation, the
real script executed via a real subprocess. Confirms the real command
sequence (build --no-cache, then down --volumes or plain down, then evidence
the handoff genuinely ran Start-Codeman.sh) and that a working handoff fails
honestly at Start-Codeman.sh's own later check rather than with EACCES.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Start-Codeman.sh, its sibling and the script it hands off to, is itself
committed non-executable (100644) upstream — it's documented and invoked
as `bash docker/Start-Codeman.sh`, never `./docker/Start-Codeman.sh`. The
"is executable" test I'd added for Update-Codeman.sh asserted the opposite
convention, which the file correctly does not follow.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
docker/README.md and docs/docker-self-update.md both already point operators
at "stop the stack, rebuild, restart" for anything the in-app updater refuses
to apply (a changed server.Dockerfile, a changed docker-compose.yaml, or a
new required .env key) — but that was a manual, hand-typed procedure with no
script of its own, unlike every other start/update path this deployment has.
docker/Update-Codeman.sh scripts it: `docker compose down`, then an
unconditional `docker compose build --no-cache` (a major update should be
certain of what actually ships, not reuse whatever layers happened to be
cached), then hands off to the existing Start-Codeman.sh for the same
careful PUID/PGID, override-file and fingerprint handling every other start
already goes through — rather than reimplementing any of that by hand and
risking it drifting out of step.
An optional --volumes/-v flag also removes the codeman-node-modules/
codeman-dist named volumes, the scripted form of the "Resetting the build
artefacts" procedure docs/docker-self-update.md already documents by hand.
Safe: those two are the only named volumes this stack declares; application
data and case workspaces are host bind mounts, never touched by
`docker compose down` either way.
Docs updated: a "Major updates" section in docker/README.md, and a pointer
from docs/docker-self-update.md's existing "Resetting the build artefacts"
troubleshooting entry.
Tests: extended test/docker-entrypoint.test.ts (the existing home for
Start-Codeman.sh's own static checks) with a bash -n parse check, the
down-before-build-before-handoff ordering, the --volumes flag's effect,
unrecognised-argument handling, and byte-for-byte agreement with
Start-Codeman.sh's own override-file resolution logic (so `down` here and
`up` there can never target different Compose files).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Two related fixes, found while re-measuring stock.ts's `accent` field
against the actual rendered UI (docs/cli-registry.md flags this field as
"transcribed, not authoritative — re-measure before wiring one up"):
1. A real, user-visible bug: `.btn-toolbar.btn-run.mode-gemini`,
`.mode-antigravity` and `.mode-omp` had no override rule inside the
`html:not([data-skin="og"])` block, unlike codex/pi/grok/deepseek, which
do. The generic `.btn-toolbar.btn-run` rule in that block resolves at
higher specificity than the base sheet's per-mode pair, so all three
rendered as plain claude-blue on every skin except `og` — including
`daylight-blue`, which is the actual DEFAULT skin for a fresh install
(index.html's pre-paint script), not an edge case. Added the three
missing rules, sourced from each CLI's own already-designed og-skin
colours (no new colours invented), mirroring the exact pattern
pi/grok/deepseek already use. Also corrected the stale comment on the
pi rule, which claimed this was still broken for gemini/antigravity.
2. `stock.ts`'s `accent` field was simply wrong for most CLIs — e.g. claude
was registered as Anthropic's brand orange (#d97757) while its button
renders blue, antigravity was registered purple while it renders cyan,
pi was registered green while it renders pink. Measured each CLI's real
`border-color` from its own `.mode-<id>` rule on the og skin (the
cleanest single representative hex each entry's gradient resolves
around) and corrected all 9 non-shell entries to match. `accent` has no
reader yet (confirmed via the DECLARED_FOR_LATER guard test), so this
changes no rendered output — it's a data-accuracy fix, matching the
registry's own "transcribed, not authoritative" warning taken literally.
Also fixed a false claim in types.ts's doc comment for the field
("CSS derives every per-CLI gradient from it via --cli-accent") — no
such CSS variable exists anywhere in the codebase.
Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format:check/
check:public-assets/check:frontend-syntax all clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
The maintainer's promised merge-time fixes from the final review of #453:
1. closeSplitPane() tears down a divider drag still in progress, so a split
that collapses mid-drag no longer leaves body.split-pane-resizing (the
page-wide col-resize cursor and user-select lock) set until a reload.
2. openSplitPane() re-applies the picker's own exclusions (detached session,
pid === null, no session record) for a row that went stale while the
menu sat open, refusing silently like its neighbouring gates.
3. architecture-invariants: the hard-hide of .btn-split is the
@media (max-width: 1179px) rule in styles.css, not mobile.css.
4. SplitTerminalPane.destroy() nulls onclose (and onerror) beside onopen
and onmessage.
5. Picker rows drop the data-session-id attribute nothing read.
6. The Pane-A-ends branch collapses with skipPrimaryResize, so the closing
resize is no longer aimed at the session the server just removed.
7. The {t:'r'} refresh path is single-flight across the fetch and the
chunked write, coalescing a mid-replay refresh into one trailing re-run.
Tests: split-pane-auto-collapse-unit gains the drag-teardown, exclusion and
skip-resize cases; the new split-pane-terminal-unit covers destroy() and the
refresh single-flight. All were run against the pre-fix module to confirm
they fail there.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit dbd39aed015ae5ae5870aba398bf4b4ab5118e47)
- styles.css: restate the composer overlay's own bottom gutter after the fold rules
(the generic .paste-overlay longhand erased it: 0px flat, hinge strip replacing it
folded) and subtract the fold strip from the dialog's max-height
- test/foldable-layout.test.ts: simulate the cascade for
.paste-overlay.prompt-composer-overlay (fails without the CSS fix); pin the palette
anchor by name instead of ELEMENTS.at(-1)
- keyboard-accessory.js: guard the app global in refreshForActiveSession() like the
rest of the file
- keyboard-accessory.js: a whitespace-only draft is empty (Send no longer submits
blank lines); the text still goes out untrimmed
- keyboard-accessory.js: derive _composerMaxLength and the frame refusal from one
64 KiB frame limit minus both bracketed-paste markers so they cannot drift
- keyboard-accessory.js: translate the textarea placeholder and label at build time,
since the DOM translator skips <textarea> subtrees
- i18n.js: zh-CN entries for the composer dialog copy
- docs/wiki/Mobile-Guide.md: describe the Compose key instead of a clipboard key
- CLAUDE.md: a "Mobile prompt composer" paragraph after the accessory bar one
- test/mobile-prompt-composer.test.ts: pin the whitespace rule and the derived budget
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit f6725ba52da17b0bdbee8be3b5011e7cae514f69)
- test/opencode-resize.test.ts: retarget the launcher guard at the real code (this.selectSession(firstSessionId), any this.activeSessionId assignment) with an anti-vacuity check; the old strings existed nowhere, so it could never fail
- session-ui.js: restore as comments the two invariants the merged bodies lost (deepseek leaves statusReporting unset, i.e. ON; no effort field for external CLIs, it is Claude-specific)
- docs/cli-registry.md: move the frontend-guard paragraph below the two backend-guard paragraphs so they keep their antecedent, and note the widened comparison shape
- test/frontend-cli-no-id-branching.test.ts: the comparison shape accepts any left-hand identifier (const m = this._runMode; m === 'codex' was invisible), normalized to `mode`; the two `m !== 'shell'` display filters are allowlisted and the remaining blind spots documented
- test/run-mode-dispatch.test.ts: table-driven pin of run() dispatch (claude to runClaude, each RUN_MODE_LAUNCH id to _runCliMode(id), shell to runShell, unknown to runClaude, lock held and released)
- CLAUDE.md: name the second CI-gated guard next to the backend one
- server.ts: every </head> injection passes a replacer function; a clis.json label containing $' re-injected the rest of the document past escapeScriptJson (two render tests pin it, proven failing on the string form)
- _isAltCliMode(): no reference anywhere in the tree, nothing to fix
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 1ea363ff808a62861559bc141e724b163cc1c56e)
The maintainer's promised follow-ups to opticon454's picker promotion,
applied on the landing branch after the merge (ecb95b5d):
- session-ui.js: the promotion tag ("Currently loaded" / "Last used") and
the "Default" pill are two separate spans, so a promoted row that is
also the endpoint's defaultModelId shows both instead of silently
losing its Default marking; two tests pin it (both fail on the old
exclusive-slot rendering).
- styles.css: a dedicated #customModelPickModal .set-scope rule, since
the pill was only styled inside the three settings modals and rendered
as plain body text here; same skin tokens, modal layout untouched.
- docs/wiki/Custom-Model-Endpoints.md: describe the promotion (currently
loaded, else last used per device), the separate Default pill, and
that nothing is ever auto-chosen.
- CLAUDE.md + docs/custom-model-endpoints.md: credit the real "Last used"
writers (_runCustomModelEntryViaRestart and
_quickStartWithCustomModelConfirm; runCustomModelEntry only dispatches
since 88e5b7b2) and drop the now-wrong "both defer to Default" sentence.
- Not done: moving the one-shot "last used" write into
_runCustomModelEntryOneShot, because the existing one-shot tests assert
that _quickStartWithCustomModelConfirm writes the key itself.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 0cb0f911adc14a852ba5c2951768a4aa87c25657)
The smart-copy gate only entered its selection-check block behind
hasSelection(), so a selection-less Ctrl+Shift+C skipped straight to
`return true` and ceded the keystroke to the browser's own handling
(e.g. Chrome's Inspect-Element binding) instead of matching Pane A's
"never falls through" contract for that chord.
Verified live in a real browser that this is a UX-parity fix, not an
interrupt-safety one: xterm's evaluateKeyboardEvent never emits PTY
data for a shifted ctrl-letter regardless of any gate (only "_" and
"@" get special-cased), so no accidental 0x03 was ever at risk. The
regression test added here asserts on the dispatched event's
defaultPrevented rather than the absence of a WS frame, since the
frame-count check passes vacuously for this exact key combo whether
or not the gate fires.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- buildSplitPickerSessions() now excludes any session with pid === null
(exited CLI, tripped PTY-exit breaker, a restore that never re-attached).
Pane B has no equivalent of selectSession()'s auto re-attach POST, so a
split opened onto one had nothing reading its tmux pane: no terminal
events ever arrived and Session.write() silently dropped every keystroke
with no ack either way, while the socket itself reported healthy.
- Fixed the hollow chord regression test: the synthetic keydowns carried no
keyCode, which is what xterm's evaluateKeyboardEvent switches on to
produce a data frame at all, so the assertion held regardless of whether
the gate fired. Adding real keyCodes surfaced a second, real bug in the
Alt+B case: the event bubbles to app.js's own document-level shortcut
dispatcher, which really toggles the sidebar and resets the layout
attribute the gate reads before Pane B's own (later, non-capture) handler
ever sees it — fixed by driving the app's real settings cache instead of
only the DOM attribute.
- Ported the two remaining primary-pane gates with real consequences:
Ctrl+Z (SIGTSTP) is swallowed for every non-shell session, matching
terminal-ui.js's reasoning (an Ink/TUI agent loop stops dead with no
visible output otherwise), and Shift/Ctrl+Enter now POSTs to
/api/sessions/:id/send-key for THIS pane's own session instead of
letting xterm send a bare \r, which used to submit an incomplete prompt
instead of inserting a newline. Smart-copy Ctrl+C is re-implemented
against Pane B's own terminal (copying app.copyTerminalSelection() would
have copied Pane A's selection instead).
- Updated docs/architecture-invariants.md and docs/split-pane-sessions-plan.md
to match, and added CLAUDE.md's missing .split-picker-menu z-index entry.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Pane B had no attachCustomKeyEventHandler of its own, so the document
capture-phase shortcut handler's preventDefault() (which does not stop
xterm) left Ctrl+K/Alt+1/Alt+B ALSO writing their raw byte/escape
sequence into Pane B's live PTY on top of whatever the app action did
to Pane A. Pane B now installs the same registry-aware gates the
primary pane's own attachCustomKeyEventHandler uses. Ctrl+V is left on
xterm's default paste — no image-paste trap to route it to.
Plus the rest of the review's smaller items:
- Narrowing the window past the desktop gate now closes an open split
instead of leaving it stranded on screen.
- Split is refused while a web tab is active (activeWebviewId), which
used to open Pane B's socket behind a hidden container.
- Pane B now handles the server's `{t:'r'}` refresh frame via a shared
_loadBuffer() helper (also used by connect()), instead of ignoring it.
- The divider drag now uses pointer events + setPointerCapture (mirrors
tab-rail-resize.js), a button!==0 guard, preventDefault, and a
body.split-pane-resizing cursor/selection lock — a plain mousedown
drag selected the text under the cursor as it crossed both terminals.
- Pane B's close control and the picker rows are real <button>s now
(keyboard-reachable), with matching CSS chrome resets.
- Dropped the redundant CodemanBase.base prefix on the buffer fetch
(the global fetch wrapper already applies it).
- data-preview-order for the Split settings chip moved from a collision
with Ultracode Agents (both 15/12) to 11.5, matching its real
position between Multi-monitor and Ultracode Agents in the header;
widened test/app-settings-structure.test.ts's regex to allow the
decimal (Number() already parses it fine for the preview sort).
- Added zh-CN i18n entries for the Split button and empty-picker text.
- Dropped the stray unused `vi` import Ark0N flagged as unrelated to
this feature (vitest's `globals: true` makes it ambient anyway).
- Documented the fix and the deliberate no-cid/seq choice in the
split-pane-sessions architecture-invariants entry.
Added a real-Chromium regression test asserting Ctrl+K/Alt+1/Alt+B
dispatched at Pane B's own textarea send no `{t:'i'}` frame over its
WebSocket. Full CI gate green (409 files, 7736 tests) plus all 8
split-pane browser tests.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
CLAUDE.md previously only mentioned split-pane in the load-order list,
with nothing in the Architecture/frontend prose the way every other
feature gets, and its own module count was one stale (34, should have
been bumped to 35 when terminal-split.js was added). Add a short
pointer-style paragraph next to the other terminal features, fix the
count.
docs/wiki/The-Dashboard.md's header button table and
docs/wiki/Settings-Reference.md's header chips list are the two
user-facing surfaces that never mention Split at all; added both, plus
a note that the feature is desktop-only regardless of the setting.
docs/split-pane-sessions-plan.md: recorded the one design note that
isn't a code change — the global capture-phase shortcut handler always
resolves against Pane A, so Ctrl+L/Ctrl+W typed into Pane B affects the
other session. Not fixed for v1, same reasoning as the rest of the
"deliberately plainer" section.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two majors:
- The divider drag was unthrottled: every mousemove did a full xterm
reflow on BOTH panes and sent Pane B a {t:'z'} resize frame with no
unchanged-dimensions skip, fanning out into a `tmux resize-window`
child plus a SIGWINCH per event — ~50 of each dragging across half a
wide viewport. SplitTerminalPane.fit() is now split into localFit()
(reflow only) and fit() (reflow + send); the drag coalesces moves
into one localFit() per animation frame via requestAnimationFrame,
and sends the real resize for both panes exactly once, at drag end,
matching the primary pane's own throttledResize convention.
- Pane B pulled the FULL scrollback unchunked for every session mode,
writing it in one terminal.write() call. Mirrors the primary pane's
own mode check (app.js's selectSession): shell sessions get a
bounded 1MiB ?tail= fetch instead of ?full=1, and the fetched buffer
is written through a minimal chunked writer (32KB slices, yielding a
frame between each) instead of one primary-pane chunkedTerminalWrite
this simpler, independently created/destroyed pane has no equivalent
of (no session-switch generation counters or live-output gate).
Smaller items from the same review:
- Pane B now follows live appearance changes (applyTerminalSkin,
applyTerminalFontFamily, applyTerminalFontWeights, setFontSize all
propagate to it, matching the teammateTerminals pattern) and reads
the real codeman-font-size/terminalFontFamily/weights/DEFAULT_SCROLLBACK
settings at construction instead of hardcoding fontSize 14 / scrollback 5000.
- The Pane-B-promotion path now skips selectSession() when
_closingSessions already owns this delete (the user closing Pane A's
own tab), matching _onSessionDeleted's own active-session-handoff guard.
- Detaching a session AFTER a split is already open now yields the PTY
size in _sendResize() too (not just at picker-open time), mirroring
sendResize's own detachedElsewhere guard.
- .btn-split joins the body.solo-mode hide list, next to .btn-multimonitor.
- The split row was 6px wider than its container (two flex-shrink:0
50% panes plus a 6px divider): both panes are now flex-shrink 1.
- Pane B's header and the split-picker rows are marked so i18n.js's
exact-string lookup skips them, matching .session-name elsewhere —
a session literally named e.g. "Sessions" was translatable on zh-CN.
- The Split button now reflects open/closed state via a `.split-open`
accent style, aria-pressed, and a title/aria-label that says which
behaviour the next click gets.
- _splitPane.connect() is no longer an unawaited call with no .catch().
- terminal-split.js's fileoverview pointed at a doc path that was
renamed away in the previous push; @dependency now credits
constants.js for CodemanTerminalFont, not terminal-ui.js.
- index.html's Split settings chip no longer reuses data-preview-order
"12" (already the Ultracode Agents chip's slot in the same "header"
preview group).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Blocker from Ark0N's second PR #453 pass: moving showSplitButton into
settings-ui.js's per-device displayKeys set was only half of making it
per-device. saveAppSettings() still put it in the object PUT to
/api/settings, SettingsUpdateSchema (.strict()) does not declare it,
the server answered 400 INVALID_INPUT, and because the call site never
checked res.ok the UI still reported "Settings saved" while NOTHING
persisted — workspaceHooksEnabled, agentSkillEnabled, tunnelEnabled,
claudeModel, every toggle, on every save, on every device. Strip it
out via the same destructure every other per-device key goes through
(`showSplitButton: _ssp,`), drop the stray mention from a schemas.ts
comment (a mention there reads as "this is a real field" to the next
grep), and add a static guard test mirroring
test/terminal-auto-copy.test.ts's three-way rule.
Also finishes the desktop gate the first pass only did in CSS at
599px: SPLIT_PANE_MIN_WIDTH (1180, matching HOME_SESSIONS_MIN_WIDTH)
now backs an actual JS width check in _applySplitButtonVisibility,
with a matchMedia listener so a live window resize hides/shows the
button without a reload — the CSS backstop in styles.css is the
reverse-direction guarantee for when JS hasn't run.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ark0N's PR #453 review: nothing gated this feature to desktop even
though the design called for it (two 240px min-width panes plus the
divider need ~486px, and the divider has no touch handlers), and
showSplitButton was a SYNCED setting, so turning it on at a desk also
put the button in the phone header.
- Hard-hide .btn-split on phones in mobile.css regardless of the
setting, matching the other desktop-oriented header buttons in the
same @media (max-width: 599px) block.
- Move showSplitButton into settings-ui.js's per-device displayKeys
set and drop it from SettingsUpdateSchema entirely, matching the
showFileViewerButton/skin precedent (CLAUDE.md's "per-device keys
... must NOT be added to SettingsUpdateSchema" rule) — a desktop
opt-in must never sync onto a phone that never asked for it. Removes
the now-invalid server-round-trip test for the setting.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Exclude popped-out (detached) sessions from the split picker:
SplitTerminalPane._sendResize() has no yield-to-detached-window check
the way the primary pane's sendResize() does, so splitting against a
detached session put its own window and Pane B in a fight over the
same PTY's dimensions. Simplest fix per the review: keep them out of
buildSplitPickerSessions() entirely.
- Show a visible dead state when Pane B's WebSocket drops. onData
already silently discards keystrokes while the socket isn't OPEN
(there is no reconnect for v1), so a dropped socket left the pane
looking normal while it quietly ate everything typed into it.
- openSplitPane() returns early with no active session, so a split
triggered from the home screen no longer creates and connects Pane B
behind the opaque welcome overlay with nothing to show for it.
- onMove() during a divider drag now bails when the split has
auto-collapsed mid-drag (the other pane's session ending) instead of
throwing on `divider.parentElement` being null.
- Promote Pane B via `selectSession(id, { auto: true })` when Pane A's
session ends — this is an app-driven selection, not the user clicking
a tab, so it must not spend the promoted session's idle alert.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ark0N's PR #453 review: fit() was only ever called from the divider
drag, and the trailing-edge ResizeObserver callback in terminal-ui.js
(throttledResize) only ever measured Pane A's own container. Split at
a wide viewport, shrink the window (or toggle the Alt+B sidebar, or
drag the tab rail), and Pane A's cols changed while Pane B silently
kept its stale PTY size in both xterm and the real pane.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Per Ark0N's review on PR #453: rename the design spec to
docs/split-pane-sessions-plan.md, matching every other feature's
*-plan.md convention, and drop the 957-line implementation task plan
(docs/superpowers/plans/2026-09-15-split-pane-sessions.md) — workflow
scaffolding for the subagent-driven-development run, not repo
documentation. Fixes the now-dangling link in architecture-invariants.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Real-browser regression coverage for the previous commit:
- SplitTerminalPane connects onto an already-quiet session and shows its
existing scrollback with no new output, proving the ?full=1 fetch (not
a live echo) populated the pane.
- openSplitPane() force-resizes Pane A synchronously as part of opening
a split.
- Dragging the divider force-resizes Pane A once, at drag end.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Pane B's SplitTerminalPane.connect() only opened a WebSocket and waited for
live output — ws-routes.ts's terminal socket sends nothing on connect, only
future 'terminal' events — so it stayed blank until the target session
happened to produce new output. It looked intermittent rather than
always-broken because a resize sent by _sendResize() often nudges the
session's real tmux window to a new size, and tmux repaints its current
screen on resize; that incidental repaint was what usually populated the
pane. When Pane B's computed dimensions already matched the session's
last-known size, Session.resize() skipped the resize as a no-op and the
pane stayed empty. Fetch the existing scrollback (?full=1) before opening
the socket, same as the primary pane does.
Pane A never told its own session's PTY/tmux about a size change at all,
relying purely on the passive 300ms-debounced ResizeObserver in
terminal-ui.js. openSplitPane() now force-resizes Pane A immediately on
entering split (mirroring closeSplitPane()'s existing symmetric call), and
the divider-drag handler force-resizes it once at drag end (matching the
codebase's established trailing-edge debounce convention rather than
flooding a resize per mousemove).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
test/split-pane-auto-collapse.browser.test.ts covers "Pane B's session ends"
in a real Chromium, but that suite is excluded from the npm test CI gate.
The "Pane A's session ends, Pane B gets promoted" branch had no coverage
anywhere, and it is the one branch whose correctness depends on exact
ordering: _splitSessionId must be captured BEFORE closeSplitPane() runs
(which nulls it) or the promoted session id is lost. Loads terminal-split.js
via `vm` against a minimal fake CodemanApp (same technique as
test/session-close-fallback.test.ts), and pins all three branches (Pane A
ends, Pane B ends, unrelated session ends) plus that the original
_onSessionDeleted always still fires.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Nothing stopped a stale picker click (opened before switching tabs) or
clicking Pane B's own session tab while split from landing on
openSplitPane(sessionId) with sessionId === activeSessionId, or from
selectSession() rebinding the primary pane onto the session Pane B was
already showing — either way, two live WebSockets to one session, each
independently claiming PTY dimensions via its own {t:'z',...} resize frame.
openSplitPane() now refuses early when the target is already the active
session, and a new selectSession() prototype patch (same top-level pattern
as the existing _onSessionDeleted patch) closes an active split BEFORE the
primary pane rebinds to the session Pane B holds.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The new "Split-pane sessions" section was inserted between the "Session
list layout (header strip vs. left sidebar)" heading and that section's own
body paragraphs, orphaning the heading from its content. Move "Split-pane
sessions" to after the Session list layout section's full body, before
"Gesture control: the setting" — no change to the Session list layout prose
itself, only where the new section sits relative to it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
.main.webview-active hid .terminal-wrap when a web tab became active, but
.terminal-wrap is reparented INSIDE .terminal-split-container while a split
is open, so Pane B and the divider stayed stranded on screen over the
dashboard iframe. Hide the whole split container as one unit, mirroring the
existing .terminal-wrap rule; no state is destroyed, so returning to the
session tab shows the split intact.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
.split-picker-menu/-item/-empty (created in openSplitPicker()) had zero CSS
and could not be dismissed except by picking an item — a default-path defect
since the Split button ships enabled to anyone who flips showSplitButton on.
Add CSS matching the sibling .run-mode-menu popover's look (floating-bg
backdrop blur, border, shadow, z-index 1000 above the header's 100), and
dismiss on outside click or Escape via the same one-shot listener pattern
session-ui.js already uses for its other transient popovers
(toggleCaseSettings(), toggleRunModeMenu()). Picking an item now routes
through the same _dismissSplitPicker() method as the outside-click/Escape
handlers, so the listeners never outlive the menu.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
_sendResize() clamped Pane B's proposed cols/rows to a 40/10 floor before
sending the {t:'z',...} resize frame, so the PTY was misinformed of Pane B's
real width at the divider's own reachable 20% position, causing real
output-wrapping bugs. The primary pane (terminal-ui.js's
getTerminalDimensions()) sends fitAddon.proposeDimensions() unclamped and
lets the server enforce its own valid range ([1,500]/[1,200] in
ws-routes.ts); Pane B now matches that convention.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
.btn-split--hidden had no matching CSS rule anywhere, so the opt-in Split
header button shipped visible to every user on every viewport regardless of
the setting. Add the `display: none !important` rule alongside its sibling
marker classes (.btn-multimonitor--hidden etc.), plus a static regression
guard (test/split-pane-hidden-button-css.test.ts) that fails if any future
"*--hidden" marker class in index.html is missing a matching CSS rule.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Task 4's implementer found two real bugs in this plan's browser-test
helpers: POST /api/sessions nests the id at data.session.id (not
data.id), and mode:'shell' needs a follow-up POST .../shell to actually
spawn a PTY. Fixed in Task 4's own snippet (documentation accuracy —
already fixed in the real committed code) and pre-emptively in Tasks
5/6's createShellSession() helper before either was dispatched, so
neither implementer has to rediscover it independently. Also corrected
the <script> tag snippet to defer, matching the real file's convention.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Monkey-patching the instance's _onSessionDeleted inside a
DOMContentLoaded listener races connectSSE()'s handler-wrapper cache,
which captures the function reference by value on first connect and
never re-reads it. Patching CodemanApp.prototype at module-evaluation
time (synchronous script-tag order) is unraceable: it completes before
any instance exists or connectSSE() ever runs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Preflight scan for SDD execution caught two classes of defect before
dispatch: Task 2's test invented a buildTestApp() helper and response
envelope that don't exist for /api/settings; Tasks 4-6 used
@playwright/test's runner against a test/browser/ directory that
doesn't exist in this codebase. Both corrected against real patterns
found in existing tests (system-routes-settings-partial-put.test.ts,
terminal-copy-shortcut.test.ts, tab-rail-resize.browser.test.ts).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Global Constraints previously read as if local-echo/CJK/accessory-bar
were desktop features; they are mobile-only, and split-pane is the
desktop-only side of that equation. Also names the exact spec section
instead of a loose paraphrase, and adds a worktree/branch note so an
executing subagent knows where this plan runs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
7 tasks: pure divider/picker helpers, showSplitButton header wiring,
split-container CSS, SplitTerminalPane (Pane B's independent xterm+WS),
open/close orchestration with picker and divider drag, auto-collapse on
either session ending, and an architecture-invariants entry.
Also folds in the "detach session" prior art discovered mid-brainstorm
into the spec's architecture section.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The Problem paragraph and the architecture section used "tab" where
"pane" was meant, colliding with the browser's own tab concept.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Scopes v1 of an in-app split view (two live session panes side-by-side,
draggable divider) after multi-monitor spanning turned out to solve a
different problem than showing multiple panes at once.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Three of Ark0N's four "will take at merge" items, applied instead since
they were straightforward to do properly:
1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on
<file>::<expression> (fixed last round) closed the line-shift problem
but opened a new one: every stock id was already allowlisted for
session-ui.js in the `mode === '<id>'` form, so a BRAND NEW branch
reusing that exact expression anywhere in the file passed unnoticed.
Reproduced live (`if (this.mode === 'codex')` injected into
runOpenCode()) — stayed green under the old version. Each allowlist
entry now carries the exact count of approved call sites, and a new
test asserts actual-vs-declared count for every key; a mismatch in
either direction is real (higher = new unreviewed branch riding in on
an existing approval, lower = a reviewed site was removed and the
entry is now stale). Reproduced again against the fix: same injection
now fails with an exact diagnostic (expected 2, found 3).
2. Added test/run-mode-launch-table-drift.test.ts. RUN_MODE_LAUNCH
restates four things stock.ts already owns (label, install command,
supportsCustomModel, the external-mode key set), and they agree today
with nothing enforcing it. supportsCustomModel is the dangerous one:
the Run-menu picker's rows come from the server-injected
window.__codemanCustomModelClis (built from
capabilities.customModelInjection.kind), so a CLI gaining a real
injection recipe later would be OFFERED in the picker while
_runCliMode silently drops the customModel field for it — the session
launches on the vendor's cloud while the UI claims the local endpoint.
Drives the real session-ui.js via JSDOM and compares RUN_MODE_LAUNCH
against STOCK_CLIS on all four axes.
3. Inlined the "Open Question 7 in PR-B2.md" references in the allowlist
reasons — PR-B2.md is a local planning doc, never part of the
committed tree, so the reference was dead on arrival for anyone
reading the repo. Points at the PR #458 review thread instead.
4. Added a sentence to docs/cli-registry.md naming the new frontend guard
alongside the backend one it mirrors.
Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format/
check:frontend-syntax all clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
Four things from the maintainer's review on PR #459, all fixed:
1. Bound _getCustomModelCurrentlyLoaded's probe client-side (~800ms via
Promise.race, on top of — never instead of — the route's own 5s
server-side timeout). Without it, an asleep/firewalled endpoint behind
a saved model list left the picker completely invisible for up to 5s
after the Run menu had already closed, with no spinner or toast.
`timeoutMs` is an optional param (default 800, real callers never pass
it) so a test can drive it in milliseconds, same pattern as
`_watchLlamaSwapLoading`'s own `pollIntervalMs` — this code runs in a
JSDOM window's own realm, whose setTimeout vi.useFakeTimers() cannot
patch.
2. "Last used" is now written only once a launch actually applies, never
on the mere click. It moved out of runCustomModelEntry (unconditional)
and into each path's own success point: _quickStartWithCustomModelConfirm
after the final post succeeds, and _runCustomModelEntryViaRestart right
after the apply's success check. A context-window-warning decline means
this exact model cannot work with this CLI at all, so the old
unconditional write would promote, next time the picker opened, the one
model guaranteed to fail again.
3. Documented the promotion/tag precedence and the new
codeman:customModelLastUsed:<mode>:<endpointId> localStorage key in both
CLAUDE.md's Custom Model Endpoint Profiles section and
docs/custom-model-endpoints.md's Run-menu picker section.
4. Added zh-CN entries for "Currently loaded" and "Last used" in i18n.js,
next to this modal's existing "Choose a model"/"Custom Endpoints" pair.
New tests: the client-side timeout (endpoint that never answers, one that
answers within the bound, and a rejected-after-timeout probe settling
quietly), and "last used" recording on success vs. NOT recording on either
confirmation's decline, for both the restart and one-shot paths.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Code review (high effort) on the previous commit found a real race: making
_openCustomModelPickModal async (it now awaits the currently-loaded-model
probe before rendering) meant a second, faster call for a different
endpoint could render first, only for the first call's slower probe to
resolve afterwards and overwrite the modal with the wrong endpoint's model
list — while _pendingCustomModelPick (set synchronously, before either
await) still named the second, correct endpoint. Picking a model in that
state would launch/apply the wrong model on the wrong endpoint.
Fixed with the same mutable-generation-counter guard
_watchLlamaSwapLoading already uses for an identical async-superseded-by-
newer-call shape: every DOM write, including _pendingCustomModelPick
itself, is deferred until after the awaited probe, and a call that finds
its generation already superseded bails out untouched instead of clobbering
whatever a newer call already rendered.
Added a regression test driving two overlapping opens with a controlled
promise so the earlier, slower probe resolves after the later, faster one
renders, asserting the late response is a no-op.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Custom Model Endpoint Profiles' "which model" picker (session-ui.js's
_openCustomModelPickModal) always listed models in their raw discovery
order, so on a host with several downloaded GGUFs the user had to
remember (or eyeball the "Default" tag) which one llama-swap actually
had hot before picking — the whole point of the picker being fast is
undone if it makes you think first.
The picker now promotes exactly one model to the top of the list:
- If llama-swap reports a model from this host's own list `ready`
right now (via the existing GET /api/model-endpoints/:id/running-status
route), that model is promoted and tagged "Currently loaded" — it's
what a launch attaches to with zero wait.
- Otherwise, the last model actually launched on this exact
(harness, endpoint) pair is promoted and tagged "Last used", read
from a new per-device localStorage key
(codeman:customModelLastUsed:<mode>:<endpointId>), written by
runCustomModelEntry on every launch attempt regardless of outcome.
- A plain (non-llama-swap) OpenAI-compatible server, an unreachable
endpoint, or a loaded-but-not-yet-ready model never promotes
anything — the rest of the list keeps its discovery order.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Two required fixes from Ark0N's review of #458:
1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on
<file>::<line>::<expression>. A single inserted line anywhere above an
entry shifted every subsequent line number, so all 21 entries went stale
simultaneously and the same 21 branches were reported as "new" — on a
file six other open PRs also touch. Dropped the line number from the key
(<file>::<expression>, matching the backend guard's own design), which
collapses 21 line-keyed entries to 11 or-collapse where the same
expression recurs at multiple call sites in the same file.
2. test/run-mode-ui.test.ts's terminal-ownership guard scanned method
bodies via `^ {2}async (run[A-Za-z]*)\(\) \{$`, which matched the 8
one-line run<Mode>() wrappers PR B2 introduced but not _runCliMode(mode),
where the real logic (and the actual risk the guard exists to catch) now
lives. Fixed the regex to `^ {2}async (_?run[A-Za-z]*)\(\w*\) \{$` and
added _runCliMode to the sanity list. Same-class fix in
test/opencode-resize.test.ts, which had the identical blind spot via
runOpenCode.toString().
Both reproduced live before fixing (inserted the same comment line; added
this.terminal.clear() to _runCliMode) to confirm the bug, then confirmed
the fix catches it and the suite stays green otherwise.
Also resolves Open Question 2 by dropping window.__codemanCliCatalog
entirely: nothing consumed it, and a registry DECLARED_FOR_LATER field
costs nothing until read while an unconsumed script tag on every page
render is a different trade. Reverts Phase 1 cleanly — server.ts's
injection, shortBadge back in types.ts's DECLARED_FOR_LATER list and the
pinned guard test, and the three associated render-index-html.test.ts /
server-index-title.test.ts assertions.
Full gate: 405 files / 7717 tests / 0 failures (net unchanged), typecheck/
lint/format clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
PR #380 (PR B) held back the frontend half of the CLI registry refactor,
explicitly deferring window.__codemanCliCatalog and making session-ui.js /
mobile-overview.js catalogue-driven as "PR B2".
- Inject window.__codemanCliCatalog in renderIndexHtml(), following the
existing __codemanCustomModelClis pattern (escapeScriptJson-guarded,
resolved per-request). Reading CliEntry.shortBadge here is what makes it
genuinely read, so it drops out of types.ts's DECLARED_FOR_LATER list.
- Consolidate session-ui.js's 8 near-duplicate run<Mode>() launch functions
(opencode/codex/gemini/antigravity/pi/omp/grok/deepseek) into one shared
_runCliMode() plus a local RUN_MODE_LAUNCH config table. The 8 method
names stay as thin wrappers (index.html calls them by name; tests assert
on the name). Also collapses a duplicated 8-way isAltMode/isExternalCli
OR-chain (same expression, copy-pasted twice in openSessionOptions) into
one EXTERNAL_CLI_MODES check.
- Add test/frontend-cli-no-id-branching.test.ts, a guard scoped to
session-ui.js/mobile-overview.js only (not the rest of src/web/public/,
which stays explicitly out of scope per CLAUDE.md), mirroring the
backend's own no-id-branching guard.
mobile-overview.js and the wiring of accent/echo/wheelForward/
keyboardAccessory were investigated and deliberately left alone: the first
is already a single, tested, gated table (not duplicated logic); the second
set belongs to terminal-ui.js/keyboard-accessory.js/styles.css, files
outside this PR's mandate.
Verified on a tmux-capable devbox (this sandbox has no tmux): full CI gate
at 405 files / 7717 tests / 0 failures, typecheck clean, 94 targeted tests
covering exact per-CLI wire-body shapes unmodified and passing, and a live
anti-vacuity check on the new guard (injected a real branch, confirmed it
fails, reverted, confirmed green).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
Installer v2. `curl -fsSL https://getcodeman.com/install | bash` now looks at the machine first, asks at most three questions up front (how the dashboard is reached, optionally what to call the machine on your tailnet, whether to run Codeman as a background service), does the install unattended behind progress spinners with the output in `~/.codeman/install.log`, and ends on the URL with a QR code to scan. One consent covers every missing package and sudo asks for your password once. Flags pipe through `bash -s --` (`--tailscale | --lan | --local`, `--name <n> | --no-rename`, `--service | --run | --no-start`, `--yes`, `--password`, `--port`), `install.sh status` prints the URL and the QR code again, and the cloudflared question moved out of the main flow into `install.sh cloudflared`. On the Tailscale route, a `:443` that already belongs to another app gets Codeman under `https://<node>/codeman` (or on a second port) instead of a dead end, the node can be renamed opt-in (`--name`, `install.sh name`, undone by uninstall), and the HTTPS-certificates toggle is polled with the admin page opened for you. Also fixed on the way: the installer's own `npm install` no longer lets the postinstall start a stray server on port 3000 (the service crash-looped on EADDRINUSE while the done screen said "running"), the LAN address comes from the default route rather than the first interface, a hand-written LaunchDaemon on a headless Mac is left alone, a flag re-run keeps an existing dashboard password, and the done screen's start command carries the sub-path and port it was installed with.
"description":"Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
@@ -28,12 +28,15 @@ The frontend is plain JS served from `src/web/public/` with no bundler in dev: e
CI runs all of these, so save yourself a round trip:
```bash
npm run typecheck # tsc --noEmit, strict mode
npm run typecheck # tsc --noEmit, strict mode
npm run lint
npm run format:check
npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules
npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules
npm run check:browser-excludes # every browser-driven test is kept out of `npm test`
```
`npm install` also installs a `pre-push` git hook that runs these static checks (about 10-40s, machine-dependent) and blocks the push if one fails. It skips itself when you push something other than the checked-out HEAD, or when the tree has uncommitted changes the checks would read. Skip it once with `CODEMAN_SKIP_PREPUSH=1 git push`; it never replaces a `pre-push` hook of your own.
- @opticon454 for four PRs in one night: the Git status indicator with its panel of uncommitted and unpushed work and per-file diffs (#537), `codeman doctor` in Settings (#536), creating a case in a custom folder (#535) and the detailed-rail rename clamp fix (#534). Both review rounds came back within half an hour with every item addressed.
- @aakhter for making the grouped vertical rail editable end to end (#525), the inline rename write queue and the long-prefix editor layout (#526), and bounded path probes so an unreachable network mount can no longer freeze the server (#516). Every round came back with tests that replay the exact sequences from the review, and the review nits were already fixed before landing.
**Edit tab groups in the vertical rail (#525).** The grouped rail is now editable from the browser: create, rename, reorder and delete groups, and move tabs between groups or back to Ungrouped, from the row menu, a group menu (Shift+F10, ContextMenu, right-click, or the header glyph, which stays visible on touch screens; F2 renames inline) or a mouse/pen drag. A flat rail offers "Move to new group" to make the first one. Every edit is a named operation saved through the existing `PUT /api/tab-layout`, one write in flight at a time; a version conflict replays the pending operations onto the server's layout and retries, so a concurrent edit from another device survives, and unsaved edits survive a reload. A drag released outside the rail leaves nothing behind, and committing a group rename by clicking elsewhere leaves focus where you clicked. No server changes.
**Inline rename that keeps up (#526).** Inline renames go through a per-session queue: one PUT at a time in the order they were made, the confirmed name applied even if the editor was reopened or cancelled meanwhile, an editor reopened over a rename in flight starts from that name, and a failed rename always shows its toast. A long `w<n>-<case>` prefix no longer pushes the editor out of its row in the rail or the sidebar, and the detailed rail no longer keeps its 3-line clamp around the editor (#534 found and fixed the same clamp independently).
**Git status in the bottom bar (#537).** Turn on Settings → Header & Panels → Bottom bar → "Git status" (per-device, off by default) and a small indicator shows the active session's repository at a glance (`● 3` uncommitted files, `↑ 2` commits not pushed, `✓` when everything is committed and pushed). Click it for a draggable window listing the uncommitted files (staged, not staged, untracked, conflicts; click one for its diff; grouped under collapsible folders) and the unpushed commits. A folder holding several projects gets a section per repository found up to two levels down, and an unrelated repository above the workspace (a dotfiles repo in your home folder) is ignored. Read-only and offline: Codeman never fetches or changes the repository, and git never runs on a repository a Docker case can write to. Not shown for Docker or remote sessions. New `GET /api/sessions/:id/git-status` and `GET /api/sessions/:id/git-diff`.
**Diagnostics in Settings (#536).** Settings → System → Diagnostics runs `codeman doctor` on the server (`GET /api/doctor`) and lists which agent CLIs, tmux, Node and the optional office tools are installed, with versions, paths and install hints. The probe runs in a child process, so a slow `--version` cannot freeze the server, and both the panel and the terminal `codeman doctor` now also look in each CLI's usual install directories, so a CLI installed outside a service's minimal PATH is found. Admin only in multi-user mode.
**Create a case in a custom folder (#535).** Add Case → Create New has a "Create in a custom folder" option with a Browse button: the case folder is created inside the parent you pick, scaffolded like any other case and listed alongside the rest (deleting it unlinks, never removes files). `POST /api/cases` accepts an optional `path` for the same thing. The folder must not exist or must be empty, and system folders, the home folder, credential folders and the cases directory itself are refused. Nothing is left behind if creation fails part-way. Admin only in multi-user mode.
**An unreachable mount no longer freezes the server (#516).** A linked case can live on a network mount, and when that mount goes away a hard mount makes `stat()` wait indefinitely; the synchronous probes in the case routes, the workspace hook and statusLine helpers and session creation used to freeze the whole server with it. Those probes now go through one bounded, tri-state probe (present, absent, or unknown when nothing answers in time): a stalled path costs one threadpool worker, paths on the same mount answer "unknown" without a new stat, and unrelated paths keep working. "Unknown" is never treated as "absent": `GET /api/cases/:name` reports an unreachable linked case with `unreachable: true` instead of NOT_FOUND, Run creates a case only on a real NOT_FOUND, and session creation answers OPERATION_FAILED for a folder that did not answer and never scaffolds over it. Tunable with `CODEMAN_PATH_PROBE_TIMEOUT_MS` (default 1500) and `CODEMAN_PATH_PROBE_MAX_STALLED`.
**Fixes applied while landing.** Tab groups: an edit made while an earlier save was still in flight, and made inapplicable by that save's conflict (its group deleted on another device), is no longer dropped silently but reported like every other dropped edit, and the menus stop offering a new group once the 32-group limit is reached instead of failing with an untranslated error. Rail: a static CI check now pins that no rail or sidebar clamp out-ranks the rename unclamp (the browser test that caught it is outside the gate). Git status: a cached repository list is re-checked against the current Docker workspaces on every poll, a diff larger than 8 MB is cut short instead of failing, a diff click refreshes only that repository, a dotfiles repository above the workspace is identified with one `rev-parse` before any full status (a failing status there no longer hides the repositories below), and "Upstream is gone" now reads "Upstream not on remote", which is also true for a branch that was never pushed. Doctor: candidates are judged like the Run menu's own resolver (a wrong binary on the PATH no longer hides the right one in an install directory, a non-executable file or a relative directory reads as missing), probes are killed with SIGKILL on timeout, a missing optional tool shows ○ instead of ✗, and the contract test no longer runs the machine's installed CLIs. Custom-folder cases: the symlink-resolved target is judged against resolved roots too (home reached through a link, macOS `/private/etc`), a target inside the cases directory is refused, the success toast names the folder the server created, the new labels have zh-CN translations, and the route test can no longer delete a real `~/projects` or the live linked-cases registry when run outside `npm test`.
## 1.34.0
### Minor Changes
- 6aecc3b: ### Thanks
- @opticon454 for four PRs in one batch: webhook notifications (#523), MCP server sync (#521), the Shift+Enter keypress fix (#520) and the newline chord plus Key tester (#522). Every review item was answered in one round, and the merge-order map across all four made landing them together easy.
- @aakhter for the grouped vertical rail (#517) and its ARIA tree and full-row activation (#519), which give the owner tab-layout API its first frontend, and for the iOS IME composition preview (#499), carried through three careful review rounds including the overlay rework in the zerolag package.
- @irisitymichaelgrundberg for per-session Claude models on `POST /api/sessions` (#514) and Codex reasoning effort per session (#515), both kept registry-driven with no CLI id branching.
- @timkjr for keeping Pane B painting during a history pull and its "disconnected" marker last in every interleaving (#524), with an old-versus-new table measured in real Chrome.
**Webhook notifications (#523).** Settings → Notifications → Webhook posts the same events as Web Push (permission prompts, questions, errors, idle) to ntfy, Slack, Discord or any JSON URL, so a headless server can reach a phone with no browser open. Off by default. The URL is a bearer secret: it lives in its own 0600 file (`~/.codeman/webhook.json`), is never returned by the API, and the routes (`GET`/`PUT /api/webhook`, `POST /api/webhook/test`) are admin only in multi-user mode. Delivery goes through the web-tab egress guard (link-local and cloud-metadata targets refused), does not follow redirects, times out after 5 s, dedupes repeats, and neutralises `@everyone`/Slack control characters in agent-supplied text.
**MCP server sync (#521).** Opt-in (`mcpSyncEnabled`, synced, off by default; `GET`/`POST /api/mcp-sync` answer 403 until it is on). Settings → Agents & CLIs → MCP servers previews or copies each installed, enabled CLI's MCP servers into the others' own config files (Claude, Gemini, Codex, OpenCode, Antigravity). It only adds missing servers, never edits or removes one, skips servers you switched off, keeps a `.codeman-bak` of every file it changes, re-parses the result before writing, writes through symlinked dotfiles, leaves files that receive env values or headers readable by you only, and reports same-name conflicts instead of overwriting. CLIs with no known MCP config (Pi, Grok, OMP, DeepSeek) are listed as unsupported. Adds the `smol-toml` dependency to read Codex's `config.toml` safely.
**Claude advisor tool.** Claude Code's experimental advisor (a stronger model the session's main model consults at decision points) can now be set per session: an `advisorModel` field on `POST /api/sessions`, `POST /api/quick-start` and `POST /api/ralph-loop/start` (`fable`, `opus`, `sonnet`, or a full id in those families), and a synced App Settings default under Models → Advisor. It rides the launch's one `--settings` JSON rather than the `--advisor` flag, because the flag exits at launch on any pairing the CLI refuses and would leave a dead pane on every respawn. It is persisted, so respawns and both restore paths keep it, and `/advisor` still switches it in-session. Agents using the codeman skill can give their claude workers one with `CODEMAN_WORKER_ADVISOR=opus`.
**Per-session Claude model (#514) and Codex reasoning effort (#515).**`POST /api/sessions` takes an optional `model` that launches that one Claude session with `claude --model <id>` and writes nothing to disk (`modelOverride` still writes the case default). It is persisted, so both recovery paths relaunch on it. `codexConfig.reasoningEffort` starts a codex session at a chosen effort (`--config model_reasoning_effort=<level>`), and it survives respawn and resume.
**Grouped vertical rail (#517, #519).** When the owner has tab groups (`/api/tab-layout`), the vertical rail draws them as collapsible sections, with collapse remembered per device, the active row always visible, and lineage arcs anchored to a collapsed group's header. The grouped rail is an ARIA tree with one tab stop and the standard arrow-key model. With no groups, the rail is unchanged byte for byte. Editing groups from the browser comes in a follow-up.
**iOS IME composition preview (#499).** On iOS Safari, the text an IME is composing (Japanese, Chinese, Korean, and the predictive composition on English keyboards) is now drawn in the terminal before it commits, inside the local-echo overlay when local echo is on. Inert on every other platform. The `xterm-zerolag-input` package gains `setComposition()`.
**Key tester and newline chord (#522).** Settings → Terminal & Input has a Key tester that shows the keydown/keypress/keyup events the browser reports, to diagnose a device where a shortcut behaves differently. Keys pressed in it never trigger app shortcuts. Shift+Enter's newline chord is now CLI registry data (`capabilities.newline`, line feed by default); no stock CLI changes.
**Fixes.** Shift+Enter no longer submits the prompt after inserting the newline: the key handler swallowed only `keydown`, so xterm's `keypress` still sent a bare `\r` (#520). Claude sessions created at the same moment (`spawn_workers`, a multi-tab Run) no longer fall out of tmux onto the direct-PTY fallback: the statusLine exporter's temp file name collided within one millisecond (#531). Pane B of the split view keeps painting during a history pull, and its "disconnected" marker stays the last line however a close, a pull and a refresh interleave (#524).
**Fixes applied while landing.** Webhooks: the App Settings Save button now saves webhook edits too (a refused URL keeps the dialog open with a warning), Send test saves pending edits first, and a Remove URL button clears a saved URL. MCP sync: a config file that fails to parse is reported by line and column only, never by quoting its content, which can hold API keys; the sync follows `CLAUDE_CONFIG_DIR`, `CODEX_HOME`, `XDG_CONFIG_HOME` and `GEMINI_CLI_HOME` from the server's environment and skips a target it cannot place instead of writing a file the CLI never reads; Preview before saving says to save first; and the MCP group is hidden from non-admins in multi-user mode. Grouped rail: a collapsed group's header shows the red or yellow ring of a hidden row that needs you; layout reads rebuild the rail only when something it draws changed, and failed reads back off (5, 10, 20, 40 s) instead of retrying every 5 s forever; a corrupted collapse preference resets instead of disabling collapse; Ctrl+Shift+{ / } only moves a tab within its own group; tapping a group header or row no longer dismisses the phone keyboard; keys pressed on a row's own buttons no longer move tree focus; and screen-reader positions stay correct after a re-sort. Sessions: `model` on `POST /api/sessions` refuses a value starting with a dash, and `model` or `advisorModel` together with `attachRemoteSession` is now a 400 instead of being ignored; non-Claude sessions no longer report or persist Claude's default model. Split view: a refresh queued behind a history pull no longer leaves a second, stale "disconnected" marker above its replay. iOS IME: a composition on an empty prompt now follows the prompt when output or a resize moves it, and the `xterm-zerolag-input` README documents `setComposition()`. The Shift+Enter and Key tester browser tests now drive the shipped handlers instead of copies.
## 1.33.3
### Patch Changes
- f776ad8: ### Thanks
- @JDProfresh for rendering Markdown in the File Viewer (#503) through the chat's existing markdown pipeline and sanitizer rather than a second one, plus the Lines and Wrap toggles and the sanitizer fix that stops a document from clobbering `document.app`.
- @timkjr for bringing the Shell scroll-to-top history pull to the split view's second pane (#506), following #494's rules down to the back-off, with tests that fail on the code before each fix.
- @irisitymichaelgrundberg for the `#session=<id>` dashboard link (#507), so a page that keeps one Codeman window open can switch it between sessions without reloading it.
- @dignfei for handing focus back when the Command Palette or the Session Manager closes (#509), and for the six-overlay measurement that showed exactly which two were broken.
**Markdown files render in the File Viewer (#503).** Opening a `.md` or `.markdown` file now shows it as a document: headings, tables, code blocks with the same copy buttons as the chat, images relative to the file, and links to other documents that open inside the viewer. An `MD` pill switches back to the source, and Edit works from either view. Plain text gets a `Lines` gutter (never part of a copy) and a `Wrap` toggle, all three remembered per device. `.avif` images preview inline, and printed `.avif`/`.ico` paths open the viewer instead of the tail view. An in-workspace file path clicked in the terminal still opens the live tail view.
**Link a dashboard window to a session (#507).** An outside page, such as a task board, that keeps one Codeman window open can now switch it to a session by pointing it at `/#session=<id>`. Only the fragment changes, so the page stays loaded and the switch is an ordinary tab selection. A link to a session the dashboard does not list yet waits up to 30 seconds for it to appear and then shows "Session not found"; picking another tab, going Home or opening a web tab cancels the wait. Following a link does not count as looking at the session, so its idle alert stays armed. The fragment is documented in `docs/extending-codeman.md` and is now a stable surface under `docs/versioning-policy.md`.
**Escape no longer strands the keyboard (#509).** Closing the Command Palette or the Session Manager now hands focus back to whatever held it before they opened, usually the terminal, so you can keep typing without clicking first. An Escape pressed while neither is open changes nothing.
**Split view: a Shell Pane B scrolls back into tmux history (#506).** Wheel up at the top of a Shell session in the split view's second pane now pulls the most recent 1 MiB of its tmux history and keeps your place, the same as the primary pane since 1.33.2.
**Fixes applied while landing.** Markdown opened from an attachment card no longer resolves relative images and links against the workspace root, where they could show a missing image or open a different file of the same name; they render as their alt text and link text instead. Rendered files no longer turn every source line break into a hard break the way chat messages do, so a README wrapped at 80 columns reads as flowing paragraphs. Absolute-path links inside a rendered document open in that document's session. A disconnected Pane B keeps its "disconnected" marker as the last line even when the socket closes in the middle of a history pull. Closing the Session Manager through a row's "Switch to session" or "Open folder" no longer pulls focus back from the terminal to the header button.
## 1.33.2
### Patch Changes
- e439cf0: ### Thanks
- @aakhter for keeping web-tab events private to their owner in multi-user mode (#501), with end-to-end isolation tests that fail without the fix, and for the browser-test exclusion check and pre-push hook (#500), including the hooks-dir resolution that never writes outside the repo's own `.git/hooks`.
- @opticon454 for keeping CLIs installed from Settings across Docker container updates (#490) and for the static Git identity for the Docker images (#492).
- @timkjr for letting a Shell pane's scroll-to-top reach tmux history (#494), with tests that fail on the commit before each fix.
- @JDProfresh for tracking down why wheel and touch scrolling did nothing in Claude's default inline view (#498), with the tmux measurements that proved it.
- @irisitymichaelgrundberg for the follow-up that makes an agent waiting on artifact comments raise its alert again (#491).
**A tab stays busy while Claude waits for its own workers.** When Claude hands work to an ultracode workflow or background agents, it ends its turn with `✻ Waiting for 1 dynamic workflow to finish` and resumes by itself when they report back. The idle probe used to call that session idle for the whole wait, and at phone width nothing on screen changes for minutes. A new optional registry field, `capabilities.workDetect.awaitingLine`, names that closing row, and only the newest column-0 row directly above the composer counts, so the session goes idle normally once the follow-up turn ends.
**Prompts sent through the API are no longer left unsent.** A prompt posted to `POST /api/sessions/:id/input` without `useMux` was written into the pane in one piece, and Claude Code (measured on 2.1.283) takes a burst of about a hundred characters or more as a paste, so the trailing `\r` became a newline and the prompt sat on the composer while the route answered 200. Short prompts went through, which is why it looked random; Codex and OpenCode showed the same thing. A plain prompt (printable text plus exactly one trailing `\r`) now goes through tmux: the text is typed, Enter is pressed as its own key, and the server presses it again while the prompt is still on the composer. Raw frames (escape sequences, a bracketed paste, a line feed, a bare `\r`) and an explicit `"useMux": false` keep the direct write. The same fix reaches cron jobs in "Paste (direct)" input mode, which reported `prompt_sent` for a prompt that never left the composer: the text is written raw, Enter follows as its own write 300 ms later, and the session presses it again while the prompt is still unsent. A cron run with no session to write to now fails instead of reporting the prompt as sent.
**Scrolling works again in Claude's default inline view (#498).** Wheel and touch gestures were forwarded to every Claude 2.1.187+ session as mouse reports, but only Claude's fullscreen renderer (`CLAUDE_CODE_NO_FLICKER=1`, or `"tui": "fullscreen"` in `~/.claude/settings.json`) listens for them, so in the default view scrolling did nothing. Codeman now forwards them only while Claude has mouse tracking switched on, and otherwise scrolls the terminal's own scrollback.
**Shell panes scroll back into tmux history (#494).** Scrolling to the top of a Shell pane now pulls the most recent 1 MiB of its tmux history, so output that arrived in a burst is reachable without pressing **Load full history**, which still loads the rest.
**An agent waiting on artifact comments alerts again (#491).** A session whose agent published an artifact and is waiting for somebody to comment on it now raises the normal idle alert and lands in NEEDS YOU, instead of being treated as busy with background work.
**Web-tab changes stay private in multi-user mode (#501).** The `webview:changed` event reached every connected user, exposing the ids of other users' web-tab creates, edits and deletes. It now carries the tab's owner and reaches that owner plus admins only. Single-user mode is unchanged apart from a new optional `owner` field on the event.
**Phone header tabs look like tabs (#504).** On phones every header tab is now a chip with a fill and a border, the Alt+N digit (a keyboard hint a phone cannot use) is hidden, names get 80px instead of 50px, and the strip fades at whichever edge still has tabs scrolled out of view.
**Docker: CLIs installed from Settings survive container updates (#490).** On the Compose deployment, CLIs installed from App Settings (DeepSeek, Pi and other npm-based CLIs) now go to `~/.local` on the persistent home mount, and `~/.local/bin` is on the image PATH, so recreating the container no longer discards them. Anything installed from Settings before this release has to be installed once more after the rebuild.
**Docker: a static Git identity for the server and agent images (#492).** Set `GIT_USER_NAME` and `GIT_USER_EMAIL` in `docker/.env` (or `CODEMAN_AGENT_IMAGE_GIT_USER_NAME` / `CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL` on a bare-host install) and the identity is written to `/etc/gitconfig` in the server image and the Docker-case agent image. A half-set pair is refused on every build path. Both Docker changes edit `server.Dockerfile`, so the in-app updater asks Compose deployments to rebuild with `docker/Start-Codeman.sh` instead of updating in place.
**Contributor tooling (#500).**`npm run check:browser-excludes`, now a CI step, fails when a test that drives a real browser is still collected by `npm test`. `npm install` also installs a pre-push hook that runs the static CI checks before a push; it steps aside when the pushed ref is not HEAD or the tree has uncommitted changes the checks would read, and `CODEMAN_SKIP_PREPUSH=1 git push` skips it once.
**Fixes applied while landing.** A Shell pane's scroll-to-top (#494) no longer re-pulls the same window on every gesture once the browser's 50,000-row scrollback is full, and a pull that hit the byte cap no longer claims the older history is gone. The scroll-routing diagnostics (#498) now log whether Claude has mouse tracking on. The artifact-comment check (#491) also refuses a footer cut off in the middle of the chip. The pre-push hook (#500) steps aside when `npm` is not on PATH, as in some GUI git clients, instead of blocking every push. The Docker Git identity error (#492) names the two variables to set. New tests pin the image PATH order, the identity on both agent-image build paths, and the `?full=1&tail=` terminal route.
## 1.33.1
### Patch Changes
- ### Thanks
- @irisitymichaelgrundberg for closing sessions whose agent exited cleanly (#486), built carefully around every way a pane exit can lie (a SIGKILL with no status, a single misread), with the `.claude-images` guard split into its own commit as asked.
- @opticon454 for the live-refreshing case picker and Manage search (#483), and for the uv/uvx, libsecret and pnpm additions to the Docker images (#487, #485).
**Finished sessions close themselves (#486).** A session whose agent you ended with `/exit` is now closed the same way the X button closes it, so finished sessions stop piling up on the board; the conversation stays resumable from the Resume list and the lifecycle log records "agent exited cleanly (status 0)". Only an explicit exit status 0 with no signal, confirmed by two pane reads, qualifies: a crashed or OOM-killed agent keeps its row with the exit code on the tab. The phone overview and desktop home rail now say `exited` instead of `idle`, reboot restore no longer offers to rebuild a session whose agent had exited, and closing one session no longer deletes the `.claude-images` directory that a sibling session in the same case still uses. Thanks @irisitymichaelgrundberg.
**Search in the phone Select Case sheet (#488).** The bottom sheet gains a "Search cases" field that filters by name (every word must match, any order, ignoring case), Enter picks the case when exactly one row is left, and Escape clears then closes. Also fixes a dead band under Create New Case and a list shorter than the sheet could show.
**An oversized paste no longer jams a session's input (#484).** A single input over the 64 KiB frame limit used to be refused by both transports, retried every 2 s forever, block every later input for that session and come back from localStorage on each reload. Pastes over the limit are now split into in-limit frames delivered in order (up to 1 MiB; larger ones are refused with a toast and never queued), a refused frame is dropped instead of retried, frames persisted by an older build are pruned on load, and the WebSocket answers an oversized frame with an explicit `too_large` error instead of silence.
- 8841bcc: Add a search box to the Manage tab of the Add Case dialog. It filters the case list by name or path, and the reorder arrows are disabled while a filter is active so a swap cannot involve a hidden case.
- 8841bcc: The case picker now refreshes its list from `/api/cases` when it opens and every 5 seconds while it stays open, so folders deleted or created on disk appear without a page reload. If the selected case has been removed, the picker falls back to another case without saving it as the last-used one.
Thanks @opticon454.
- 77ba41f: Install `uv` and `uvx` in the Compose server image and the agent image, so MCP servers launched with `uvx` (such as the Nginx Proxy Manager MCP) can be enabled by Codex instead of failing with `uvx` not found. Both images also install `libsecret-1-0`, the native library the `keytar` dependency of the Azure DevOps MCP (`@azure-devops/mcp`) needs; without it the server crashes before answering the MCP initialize handshake.
The Compose server image now also carries `pnpm`: `dsh plugin` spawns a literal `pnpm` with no npm fallback, so the Run menu's "DeepSeek - add a terminal profile" button failed with `dsh: pnpm not found on PATH` there. Because this release changes `server.Dockerfile`, the in-app updater asks Compose deployments to rebuild the image (`Update-Codeman.sh`) rather than applying it in place.
Thanks @opticon454 (#487, #485).
## 1.33.0
### Minor Changes
- CLI management from Settings (#476, finishing the CLI registry work from #343). `~/.codeman/clis.json` used to be hand-edit only; with the new opt-in `cliManagementEnabled` switch (synced, default OFF) App Settings → Agents & CLIs can enable or disable any CLI, install a missing stock CLI with its vetted install command, and add, edit or remove custom CLIs. Six new endpoints back it (`GET`/`POST /api/clis`, `PUT /api/clis/:id`, `POST /api/clis/:id/install`, `PUT /api/clis/custom/:id`, `DELETE /api/clis/:id`), documented in `docs/api-reference.md`. Every write is refused while the switch is off, is admin-only in multi-user mode, is serialized on one queue, and refuses to overwrite a `clis.json` that does not parse or has group/world permission bits. A custom entry is re-validated through the same schema as the stock ones and its install text is never executed. `shell` cannot be disabled. The Run menu and the welcome screen are now built from the enabled catalogue, so the welcome screen also offers Codex, Shell and any custom CLI, and the stock Claude entry is labelled "Claude Code".
Models: Opus 5.5 (`claude-opus-5-5`, 1M context capable) is offered in App Settings → Models and in task routing (#480).
Self-update: on a macOS `launchd-daemon` install, a Homebrew node upgrade could leave `update-status.json` stuck at `queued`, which made every later update fail with "An update is already in progress." The updater now falls back to `node` on PATH when the server's own node binary is gone, and an in-flight status that has not been written for 15 minutes is failed on the next read. A graceful shutdown that hangs is now force-exited after 10 s (and the launchd updater SIGKILLs a server that has not exited after 30 s), so launchd can start the new build instead of leaving the service down (#478). Both fixes protect updates that start FROM this release.
Session Manager (Cmd+K): rows keep their `mode`, `claudeSessionId` and `resumeId`, so the ⋯ menu's Resume session relaunches a Codex row as Codex on its own conversation, and the mode badge shows as it does on the home list (#477).
Maintainer fixes applied while landing #457: renaming a tab to the name it already has (the Session Options field saves on blur) is now a no-op, so it no longer pins the placeholder as the `/resume` title again; Docker sessions skip the transcript title sync, since their transcript lives in the container; and the agent skill's messaging examples no longer use a `w<N>-` name as the peer name.
Tests: the suite strips every inherited `CODEMAN_*` variable, so running it inside a Docker Compose deployment no longer writes into the deployment's real case root (#479).
### Thanks
- @opticon454 for CLI management (#476), the last piece of the CLI registry, with every review item answered in one round, and for splitting the test isolation fix out into #479.
- @shenlvkang-collab for the `/resume` title fix (#457) and the careful diagnosis behind it.
- @julian3xl for the Session Manager row fix (#477), their first contribution.
### Patch Changes
- 69a7128: fix(sessions): stop pinning the `w1-myapp` placeholder as Claude's session title. Local Claude spawns passed the tab name as `--name`, which is also the `/resume` picker entry and the terminal title, and a pinned title stops Claude generating its own, so every conversation of a case showed up in `/resume` as the same `w1-myapp` and none got a generated title. Only a name the user chose is pinned now; placeholder and auto-named tabs let Claude title the conversation again. Renaming a Claude tab also reaches `/resume`: the new name is appended to the conversation's transcript as the `custom-title` row `/rename` writes (a tab that was spawned with `--name` keeps re-appending its own title until its next respawn, so the rename wins from then on). Orchestrators that rely on a fixed peer name should give workers a descriptive `sessionName` rather than a `w<N>-` one.
## 1.32.1
### Patch Changes
- 13e652e: Terminal copy: copying text out of a Claude Code or Codex pane no longer puts the pane's two-column transcript gutter on the clipboard, so pasted lines arrive flush instead of indented (#469). The width comes from the CLI registry (`capabilities.transcriptGutter`, 2 for claude and codex, measured on live panes) and is only a ceiling: a selection only ever shifts as a block, so its own indentation survives. Other CLIs and shells are untouched. It works in split panes and detached session windows too, and can be turned off per device in App Settings under Selection & clipboard.
- 13e652e: Sessions: recovering a Claude session whose tmux pane had died relaunched `claude --session-id <id>`, which Claude refuses once that id has a transcript, so the pane died again straight away and the conversation was stranded. The relaunch now resumes the conversation (`--resume <id> || --session-id <id>`), including when tmux lost the whole session (#467).
- 00f022c: Terminal: when a burst of output overflows the render queue and a frame has to be dropped, the repaint that repairs it is now retried until it actually happens, instead of being scheduled once and silently skipped when another load was in flight (#470).
- 13e652e: Mobile: a long press on blank terminal space on Android Chrome no longer opens the keyboard and blanks the terminal (#471, fixes #360). The long-press guards are now armed before the press is checked for selectable text, so a press on empty space is swallowed the same way a press on a word already was.
- 13e652e: Sessions: a tab whose agent has exited (the CLI quit, but tmux kept the pane) now says so with a muted dot and an `exited (137)` badge, instead of looking like an idle session (#466, part 1 of #446). The state is published as `paneExit` on the session and survives a restart. Nothing closes such sessions yet; that is part 2.
- 13e652e: Docker: optional GitHub CLI and Azure CLI for private repositories (#472). Both are off by default. With `CODEMAN_INSTALL_GH=1` / `CODEMAN_INSTALL_AZ=1` as build args in `docker-compose.override.yml`, the server image gets `gh` and/or `az` (with the `azure-devops` extension) wired in as git credential helpers, so after one `gh auth login` or `az login` from a shell session, Add Case → Clone Repo can clone private GitHub and Azure DevOps repositories. `CODEMAN_AGENT_IMAGE_INSTALL_GH` / `_AZ` do the same for the Docker-case agent image, and only then are the sign-ins copied into new case containers. In multi-user mode a non-admin's clone runs with the credential helpers cleared. This changes `server.Dockerfile`, so Compose deployments need a `Start-Codeman.sh` rebuild rather than an in-app update.
- 13e652e: Run menu: the Gemini, Antigravity and OMP run buttons now show their own colours on every skin; they rendered in Claude blue on all skins except OG (#463). The CLI registry's `accent` values were also corrected to the colours the UI really paints, and a test now guards the stylesheet trap that caused it.
- 13e652e: Terminal: five ways the browser terminal could silently stop being correct are fixed (#431, #464). The browser terminal and the PTY can no longer disagree about their width, which is what produced doubled lines and half-overwritten text ("text gets muffled sometimes"): there is now one function that sizes the terminal, and every resize is answered with the geometry the PTY really holds. A replay clear goes through the terminal's own queue, so bytes written just before it no longer fuse into the next snapshot. A renderer that stops painting after an iOS PWA is backgrounded heals itself instead of needing a reload. Every terminal capture has a deadline that also covers the response body, and a capture that runs out of time during a tab switch falls back to the bounded tail instead of leaving a blank pane. Output lost to a half-open WebSocket is repainted on the next successful open. The service worker's precache list is now generated by the build and its cache is rotated per build, so old releases' assets no longer pile up.
- 13e652e: Docker: new `docker/Update-Codeman.sh` for the major-update path the docs used to describe by hand (#465). It rebuilds the image with `--no-cache` before taking the stack down, clears the build-artefact volumes, refuses to run when another checkout's Compose project already owns the same name, and then hands over to `Start-Codeman.sh`.
- 13e652e: Approvals: a session that is idle only because it is waiting on its own background work (Claude Code's `1 monitor` footer chip, or a Codex background terminal) no longer raises the yellow NEEDS YOU alert or a push (#473, fixes #468). Its idle item is opened already acknowledged, and the tab, the home screens and the rail show a small `watching` badge next to the state instead. The item still exists in the Approvals Inbox, and the TUI's pending count now leaves acknowledged items out.
- b404dac: Maintainer fixes applied while landing this batch:
- Terminal (#431): while another device holds the pane's width, a resize retry no longer re-fits xterm to the container and re-wraps the whole buffer every 30 s, and no longer clears scrollback for a redraw that never comes. The PTY's spawn geometry is now recorded at attach, so `ptyGeometry` never reports a size the PTY never held.
- Terminal (#470): the `TERMINAL DROP` crash-trail line is logged once per recovery window instead of once per dropped frame (which wiped the rest of the trail within a second), and a refresh that died at its fetch deadline is no longer retried.
- Sessions (#467): the resume pin also covers the branch where tmux lost the whole session, the conversation id Codeman reports follows what the relaunch actually resumed, and the test setup strips `CLAUDE_CONFIG_DIR` so the suite stays green for anyone running a separate Claude config dir.
- Sessions (#466): detailed sidebar and rail rows show an `exited` pill instead of `idle`, the exit is announced to screen readers, and the user manual's tab-appearance table lists the new state.
- Approvals (#473): a failed pane capture clears the `watching` badge rather than keeping a stale one (a failure now falls toward an alert, not toward silence), and the header bell's count leaves acknowledged items out, matching the TUI.
- Run menu (#463): the Gemini and Antigravity run buttons no longer render two-tone on phones, Gemini's registry accent matches its tab badge, and a test now guards the stylesheet trap for every run mode.
- Docker (#465): `Update-Codeman.sh` removes exactly the two build-artefact volumes it names instead of every named volume in the project, reports a failing `docker compose` instead of exiting silently, and its docs and comments were corrected. (#472): the multi-user notes say that a non-admin's seeded Docker case also receives the gh/az sign-in when those switches are on.
### Thanks
- @irisitymichaelgrundberg for four PRs in this release: the `watching` badge that stops background work from raising false alerts (#473, from their own report #468), the exited-agent badge (#466) and the dead-pane resume fix (#467), both from their report #446, and the transcript-gutter strip for copied text (#469), a follow-up to their #451.
- @rounakdatta for the terminal resilience work (#431) and the dropped-frame recovery (#470), both from their report #464, and for answering four rounds of review in full.
- @opticon454 for private-repository support in the Docker images (#472), the `Update-Codeman.sh` script (#465) and the run-button colour fix (#463).
- @DodgyBadger for the Android long-press fix (#471), from their own report #360.
## 1.32.0
### Minor Changes
- d47f93a: feat(custom-model): the model picker puts the ready model first
When a custom endpoint has more than one model, the Run menu's picker now promotes one row to the top instead of showing raw discovery order: the model llama-swap reports loaded and ready right now (tagged "Currently loaded", the one a launch attaches to with zero wait), else the model you last launched on that harness and endpoint (tagged "Last used", remembered per device). The endpoint's default keeps its own pill, nothing is ever auto-chosen, and a plain OpenAI-compatible server or an endpoint that does not answer within a second simply keeps the old order. The probe is bounded on the client too, so a GPU box that is off no longer holds the picker closed for five seconds.
- d47f93a: feat(split-pane): view two live sessions side by side
A new Split button in the header (opt-in in App Settings, off by default, desktop only at 1180px and wider) opens a picker and shows a second live session beside the active one: its own terminal, its own WebSocket, and a divider you can drag. When either session ends the view collapses back to one pane, with Pane B promoted to the primary when it is Pane A that ended. Nothing is persisted on purpose in this first cut, so a page reload always returns to a single pane. Pane B is deliberately plainer than the primary pane (no local-echo overlay, CJK input, touch handling or keyboard accessory bar); the design and the v2 boundaries are in discussion #452.
- 72d437a: Installer v2. `curl -fsSL https://getcodeman.com/install | bash` now looks at the machine first, asks at most three questions up front (how the dashboard is reached, optionally what to call the machine on your tailnet, whether to run Codeman as a background service), does the install unattended behind progress spinners with the output in `~/.codeman/install.log`, and ends on the URL with a QR code to scan. One consent covers every missing package and sudo asks for your password once. Flags pipe through `bash -s --` (`--tailscale | --lan | --local`, `--name <n> | --no-rename`, `--service | --run | --no-start`, `--yes`, `--password`, `--port`), `install.sh status` prints the URL and the QR code again, and the cloudflared question moved out of the main flow into `install.sh cloudflared`. On the Tailscale route, a `:443` that already belongs to another app gets Codeman under `https://<node>/codeman` (or on a second port) instead of a dead end, the node can be renamed opt-in (`--name`, `install.sh name`, undone by uninstall), and the HTTPS-certificates toggle is polled with the admin page opened for you. Also fixed on the way: the installer's own `npm install` no longer lets the postinstall start a stray server on port 3000 (the service crash-looped on EADDRINUSE while the done screen said "running"), the LAN address comes from the default route rather than the first interface, a hand-written LaunchDaemon on a headless Mac is left alone, a flag re-run keeps an existing dashboard password, and the done screen's start command carries the sub-path and port it was installed with.
- d47f93a: feat(mobile): a Compose key for writing prompts on a phone
The agent keyboard bars on phones replace their Paste key with Compose: a real multiline editor with autocorrect and spellcheck, per-session drafts kept in memory only, image attach that never writes into the terminal early, and a Send that delivers the text as one paste followed by Enter, so a long prompt no longer has to be typed blind into the terminal composer. Anything you had already typed into the terminal is picked up into the editor. Shell sessions keep the direct Paste key. This is the manual first slice from #359; the auto-open setting and terminal tap routing are a separate follow-up.
### Patch Changes
- d47f93a: refactor(run-menu): one table-driven launcher for every external CLI
The eight near-identical per-CLI launch functions in the Run menu collapsed into one launcher driven by a table that a CI test keeps in step with the CLI registry, and a second no-id-branching guard now covers the frontend the way the backend guard covers the server. No behaviour change: the refactor was verified byte-identical across 288 launch permutations against the previous code.
- e899af4: Maintainer fixes applied while landing the above. The model picker's promoted row keeps its Default pill (the promotion tag and the default marker are two pills now, and they render as pills in the picker rather than as plain text). The phone composer keeps its bottom gutter on folding devices (the generic fold rule used to erase it), a whitespace-only draft is no longer sent, and its dialog is translated on a zh-CN UI. A split that collapses mid-drag no longer leaves the page stuck in resize-cursor mode, Pane B refuses a session that has no live process, and a burst of refresh frames replays once instead of twice. The `</head>` script injections on the page render use replacer functions, so a CLI label containing `$'` can no longer splice the document into the inline script, and the frontend no-id-branching guard now catches comparisons on any variable name.
- 6ef71ec: ### Thanks
- @timkjr for split-pane sessions (#453): five review rounds turned around in two days, and the pointer-capture edge case measured in a real browser rather than reasoned about.
- @DodgyBadger for the mobile prompt composer (#444), a first contribution that took the scope back down to one slice when asked, and that verified the delivery path against a live tmux pane and a live Claude Code composer instead of trusting the diff.
- @opticon454 for putting the ready model first in the picker (#459) and for collapsing the eight Run-menu launch functions into one (#458), proven byte-identical across 288 launch permutations instead of argued.
⭐ <strong>Like Codeman? <a href="https://github.com/Ark0N/Codeman">Give it a star on GitHub!</a></strong> It takes one click and helps more people find the project. ⭐
- **Background daemon & service install** — `codeman web -d` runs the server detached with a pidfile, `~/.codeman/web.log`, and verified startup (it polls the server until it answers, so a port clash never reads as success); `codeman service install` writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew `node`, `tmux` and `claude` are actually found. Secrets are never written into unit files
- **Diagnostics in Settings** — **App Settings → System → Diagnostics → Run checks** runs `codeman doctor` on the server and lists Node, tmux, every agent CLI and the optional office tools with versions, paths and install hints. Besides the `PATH`, it also looks in each CLI's usual install directories (`~/.local/bin`, `~/.npm-global/bin` and the like), so most installs are found under a service with a minimal `PATH`. Admin only in multi-user mode.
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → System → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Git status in the bottom bar** — off by default (**App Settings → Header & Panels → Bottom bar → Git status**, per device). A small indicator at the right of the bottom bar shows the active session's repository at a glance: `● 3` uncommitted files, `↑ 2` commits not pushed, `⚠` merge conflicts, `✓` all committed and pushed. Click it for a draggable window listing the staged, not-staged, untracked and conflicted files (grouped under collapsed folders, or as a flat list if you turn that setting off) and the unpushed commits; **click a file to see its diff** (new files as all additions, deleted files as all removals), with **Open file** to jump to the viewer. A folder that holds several projects gets one collapsible section per repository found up to two levels down, all collapsed until you open them. Read-only and offline (Codeman never fetches or changes the repo); not shown for Docker or remote sessions.
- **Clone a GitHub repo as a case** — paste a repository URL into **Add Case → Clone Repo** and Codeman clones it into `~/codeman-cases/<name>` and registers it as a normal case, ready to run an agent in. It preflights the URL while you type (tells you whether it can be cloned anonymously and offers the repo's real branches and tags for the optional branch/tag field), fills the case name in from the URL, and lets you pick which CLI the Run button should use. Public repositories over `https://`; Codeman never collects or stores credentials
- **Create a case in a custom folder** — tick **Create in a custom folder** in **Add Case → Create New**, pick a parent folder (Browse included) and a name, and Codeman scaffolds the new case there instead of `~/codeman-cases`. The target must be a new or empty folder; system directories, your home folder itself, credential trees such as `~/.ssh`, and Codeman's own data folder are refused. Admin only in multi-user mode.
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, **Pi**, **Grok**, **DeepSeek Harness**, or **OMP** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*` vs `GROK_*`/`XAI_*` vs `DSH_*`/`DEEPSEEK_*` vs `OMP_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md), [`docs/pi-integration.md`](docs/pi-integration.md), [`docs/grok-integration.md`](docs/grok-integration.md), [`docs/deepseek-integration.md`](docs/deepseek-integration.md) and [`docs/omp-integration.md`](docs/omp-integration.md)
- **Custom model endpoints** _(new in 1.29.0, HTTP API for now)_ — point a session's CLI at any OpenAI-compatible endpoint instead of its native backend: a local llama.cpp, llama-swap, Ollama or vLLM box, or a cloud gateway such as Azure AI Foundry or OpenRouter. Save an endpoint once (`POST /api/model-endpoints`; its models are discovered from `/v1/models`), apply it to a session (`POST /api/sessions/:id/custom-model`), and the CLI restarts in place on that endpoint. Verified live for Claude, OpenCode, Pi, Grok and OMP; Codex, Gemini and DeepSeek have documented gaps, Antigravity has no mechanism. A toolbar picker is the follow-up. See [`docs/custom-model-endpoints.md`](docs/custom-model-endpoints.md)
- **Web tabs** — open Grafana, Uptime Kuma, a Vite dev server or any dashboard URL as a tab beside your sessions (Run dropdown → **Web / URL** → **Add URL**). Dashboards are proxied through Codeman's own origin, so an `http://` target works from a phone over HTTPS and through the tunnel, single-page apps route on their own paths, and a frame that reloads recovers itself. A `localhost` link an agent prints opens as a web tab automatically. See [`docs/web-tabs.md`](docs/web-tabs.md)
`Start-Codeman.sh` rebuilds the image on every start, but with the layer cache,
and it refreshes the build-artefact volumes selectively: `codeman-dist` when
the checkout's HEAD moved, `codeman-node-modules` only when `package-lock.json`
changed. That is right for an ordinary `git pull`. It is not enough when a
`server.Dockerfile` change bumps the Node base image without touching the
lockfile: `node-pty` is compiled from source (there is no Linux prebuild), so
the old `codeman-node-modules` volume would keep a build made for the previous
Node version. For that case, or whenever you want to be certain of what ships,
`docker/Update-Codeman.sh` force-rebuilds the image with no layer cache, stops
the stack, removes the `codeman-node-modules` and `codeman-dist` volumes, then
hands off to `Start-Codeman.sh` for the usual start:
```sh
bash docker/Update-Codeman.sh
```
Pass `--keep-volumes` to skip clearing them (safe only if you know the
rebuilt image's `node_modules`/`dist` did not change). The scripted default
is the "Resetting the build artefacts" procedure in
[`../docs/docker-self-update.md`](../docs/docker-self-update.md). Only those
two volumes are removed, by name within this Compose project; any volume a
`docker-compose.override.yml` adds is left alone, and application data and
case workspaces are host bind mounts, never touched either way.
## Git commit identity
Set `GIT_USER_NAME` and `GIT_USER_EMAIL` in `docker/.env` before rebuilding:
```sh
GIT_USER_NAME='Your Name'
GIT_USER_EMAIL='you@example.com'
```
Compose passes the values to the Codeman server build, and to the server process
when it builds Docker-case agent images. Both images write the pair to Git's
system configuration during their build, so commits retain the same identity
after a container or agent image is recreated. Set both values together; an
image build with only one value fails rather than using a partial identity. An
identity already present in `CODEMAN_APPDATA_PATH`'s `~/.gitconfig` overrides
the server image's system-level default.
Run `bash docker/Start-Codeman.sh` after changing the server values. Rebuild an
existing agent image with `node scripts/build-agent-image.mjs --no-cache` in the
server container, then recreate any Docker cases that should use it.
## Private repositories (GitHub and Azure DevOps)
The images can include the GitHub CLI (`gh`) and the Azure CLI (`az`, with the `azure-devops` extension), wired into the system Git configuration as credential helpers, so Codeman can clone private repositories. Both are **opt-in and off by default**, and are turned on per host in `docker-compose.override.yml`.
### Turning them on
Add the build arguments to `docker-compose.override.yml` (see [Local customisation](#local-customisation)), then rebuild with `Start-Codeman.sh`. Set only the one you need:
```yaml
services:
codeman:
build:
args:
CODEMAN_INSTALL_GH: '1'
CODEMAN_INSTALL_AZ: '1'
environment:
# The same two switches for the Docker-case agent image Codeman builds.
CODEMAN_AGENT_IMAGE_INSTALL_GH: '1'
CODEMAN_AGENT_IMAGE_INSTALL_AZ: '1'
```
The `build: args:` pair controls the Codeman server image. The `environment:` pair controls the agent image for [Docker cases](../docs/docker-cases.md), which Codeman builds on the first Docker case; an agent image that already exists is not rebuilt by this, so run `node scripts/build-agent-image.mjs --no-cache` inside the container afterwards. The same variables work in front of that command when building it by hand. Values must be `0` or `1`; anything else stops the build with an error naming the argument.
They are not `.env` settings: turning a CLI on is a per-host choice, which is what the override file is for, and a new `.env.example` key makes the in-app updater refuse to update every existing installation until its `.env` gains the key.
The Azure CLI is the large one, about 600 MB of the roughly 670 MB the pair adds. A CLI left off leaves nothing functional behind: no apt repository, no package, no `azure-devops` extension and no credential-helper entry, so git for that host behaves exactly as it does without this feature. With both off the image is functionally unchanged; it still carries the `AZURE_EXTENSION_DIR` variable, an empty extensions directory and one small layer that copies and then removes the helper script.
### Signing in
With a CLI on, the system Git configuration routes credentials through it:
Codeman itself still collects no Git credentials. Sign the container in once from a **Terminal / Shell** session (Run menu). The session runs as the runtime account, so the sign-in is stored under `CODEMAN_APPDATA_PATH` (`~/.config/gh`, `~/.azure`) and survives rebuilds and container recreation:
```sh
gh auth login # GitHub.com -> HTTPS -> "Login with a web browser" (device code)
az login --use-device-code # then: az devops configure --defaults organization=https://dev.azure.com/<org>
```
After that, **Add Case → Clone Repo** accepts private `https://` URLs on those hosts, and `git clone` works from any session. Until a CLI is signed in its helper prints nothing, so a private clone fails immediately with the usual authentication error rather than waiting on a prompt.
**Multi-user mode:** every Codeman user's git runs as the same server account, so these sign-ins would otherwise be shared. Clone Repo therefore runs a **non-admin**'s clone and preflight with every git credential helper cleared (`git -c credential.helper=`): a non-admin can clone public repositories and anything their own SSH setup allows, but not a private https repository through the admin's `gh`/`az` sign-in. Admins, and single-user mode, keep the helpers. A non-admin's own agent sessions still run as that same account, and with the agent-image `gh`/`az` switches on, a non-admin's Docker case with credential seeding on also receives the server account's `gh`/`az` sign-in, the same as the Claude and Codex credentials; see `docs/security-architecture.md`, multi-user mode.
Azure DevOps is authenticated with an Entra ID access token that the helper requests from `az` for each Git operation, so nothing is written to disk beyond `az`'s own sign-in. An account that has to use a personal access token can set `AZURE_DEVOPS_EXT_PAT` for the container instead (for example under `environment:` in `docker-compose.override.yml`); the helper prefers it when present. SSH remotes are unaffected by any of this and keep using the account's own keys.
Docker cases copy these sign-ins into a case container only when the matching agent-image switch is on (`CODEMAN_AGENT_IMAGE_INSTALL_GH=1` for `~/.config/gh/hosts.yml` and `config.yml`, `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` for the sign-in files from `~/.azure`) and the case has credential seeding on. With a switch off they are never copied, even when the files exist, because a GitHub token or an Azure refresh token is usable by anything in the container. The copies are made when the container is **created**, so an existing case container never picks them up: after turning a switch on, signing in, or rebuilding the agent image, **recreate the case container** (remove it; the next session in that case creates a fresh one).
The GitHub agent skill for `gh` installs into the runtime account's home in the same session:
```sh
gh skill install cli/cli gh --scope user
gh skill update gh # after a later gh release
```
### Versions
Both CLIs, and the extension, are installed from their vendors' repositories with no version pinned, so they arrive at whatever is current when that build step runs. Docker caches the step, though: `Start-Codeman.sh` rebuilds with the cache, which keeps the versions from the first build until the Dockerfile changes at or above that step or the image is rebuilt with `--no-cache`. They are apt packages owned by root, so they cannot be upgraded from a session; `az extension update --name azure-devops` is the exception and works without a rebuild.
## Local customisation
Compose merges `docker-compose.override.yml` on top of `docker-compose.yaml`. Keep host-specific changes there rather than editing `docker-compose.yaml`, so this repository can be updated without losing them. Both `docker-compose.override.yml` and `docker-compose.override.yaml` are ignored by Git.
What a session's terminal shows, for a client to replay: `data.terminalBuffer`,
with `source` (`mux-visible`, `mux-full-history` or `history`), `truncated`,
`truncationReason`, `fullSize`, and `captureCols`/`captureRows` when the pane's
geometry was read. The capture runs synchronous tmux calls on the server; the
`Server-Timing` header reports `capture`, `prepare` and `total`.
| Query | Meaning |
|---|---|
| `full=1` | tmux's scrollback, not only the visible frame (`source: 'mux-full-history'`), ending with a relative cursor move back to the pane's caret. |
| `tail=<bytes>` | Keep the newest `<bytes>` of the result (`truncationReason: 'tail'` when it cut). |
| `lines=<n>` | With `full=1` only: read at most `<n>` lines of tmux history above the visible frame. An integer of at least 1, clamped to the configured history limit; absent or malformed, the whole limit (100,000 lines by default), as before. `truncated` and `truncationReason` describe byte cuts only, not this bound. Without it a full capture reads all of that history before `tail` cuts it, so a client that keeps a fixed number of lines (the tile grid sends its xterm's scrollback plus its rows) should send it. |
## Session lineage (`parentSessionId`)
A create request may name the session that spawned it, which the web UI draws as a
@@ -461,6 +484,32 @@ also pure decoration: it confers no permission, and a child is unaffected by its
parent exiting. It appears on session state as `parentSessionId` (absent when
unresolved) and survives a server restart.
## Session model (`displayModel`)
Session state (`GET /api/v1/sessions`, the `session:updated` event) carries the model a
session runs as far as the server knows it, for the web UI's session headers:
| `custom-endpoint` | The session is pointed at a Custom Model Endpoint Profile; its `modelId` answers, whatever the CLI prints. |
| `statusline` | Claude's statusLine exporter reported it (`model.display_name`); follows an in-session `/model`. |
| `screen` | Read off the CLI's own footer (`capabilities.modelDetect`, today dsh and codex); follows a switch. |
| `config` | What the CLI's own config pins for the session (`capabilities.modelDetect.configResolver`, today dsh-TUI's route), while its screen names none. |
| `launch` | What the session was launched with (`--model`, the app-wide default, `<cli>Config.model`); nothing has reported since. |
Between `statusline` and `screen` the newest report wins. The field is absent when no
model is known (a shell, a CLI that reports none and was launched without one). `model`
is display text from a pane or a CLI report: control characters are stripped and it is at
most 64 characters, but treat it as untrusted text. A `statusline` or `screen` value is
persisted and restored after a server restart until the next report replaces it; a
`config` value is read again at every pane start, attach and relaunch instead.
## Approvals Inbox
Cross-session queue of prompts waiting on a human (permission dialogs,
@@ -691,6 +740,96 @@ normal `caseName`/`mode`/etc. body)
jarring than a full relaunch, and folding it into the one-shot path is
separate work — see `docs/custom-model-endpoints-plan.md`).
## Creating a case in a custom folder
`POST /api/cases` takes `{ name, description?, path? }`. Without `path` it creates `<cases dir>/<name>` as always. With `path` (absolute, or starting with `~`) the case folder is created at that exact path instead, scaffolded the same way (`CLAUDE.md`, `src/`, `.claude/settings.local.json`), and registered in the linked-cases registry, so it lists, resolves and deletes like a linked case (deleting unlinks; it never removes files). Response: `{ case: { name, path } }`, where `path` is the symlink-resolved folder.
The target is judged before anything is written:
- It must be absolute with no `..` and none of the shell metacharacters a session working directory is rejected for (spaces are fine). `400 INVALID_INPUT` otherwise.
- It must not be a system directory (`/etc`, `/usr`, `/proc`, ...), the home folder itself, Codeman's own data folder, or a credential/config tree (`~/.ssh`, `~/.aws`, `~/.claude`, ...). Judged on the path as typed and on its symlink-resolved form, against both the given and the symlink-resolved roots. `400`.
- It must not be, or be inside, the cases directory (the caller's own and the shared one): a case there is a plain create without `path`. `400`.
- Its parent must already exist (one folder is created, never a chain): `404 NOT_FOUND`. A parent that does not answer (an unreachable network mount) or cannot be read is `422 OPERATION_FAILED`, checked through the bounded path probe before anything else touches it.
- The folder must not exist, or must be an **empty** directory; a folder with contents is Link Existing's job: `409 ALREADY_EXISTS`. A symlink or a plain file at the target is `400`.
- `409 ALREADY_EXISTS` also for a case name already in use (in the cases dir or the registry) and for a folder that is already a case.
Admin only in multi-user mode (`403`), like `POST /api/cases/link`: it writes outside the cases directory and into the shared, ownerless registry. If anything fails after the first write, what this call created is removed (the whole folder if it created it, otherwise only the scaffold inside the empty folder you picked) and the response is `500`.
## Git status
`GET /api/sessions/:id/git-status` is what the bottom-bar Git indicator and its panel read (Settings → Header & Panels → Bottom bar, per-device, default off). It reports what the session's workspace has not committed or pushed. **Read-only and offline:** it never fetches, pulls, commits or writes (it runs `git status` with `--no-optional-locks`, so it does not even refresh the index), which is why `behind` is as of the last `git fetch`. The session is resolved like every session route (ownership via `findSessionOrFail`; another user's session is `404`). A repository whose root is, or is inside, a Docker case workspace is dropped (from the walk-up, the scan below a folder, and the diff route): a container can write there, and a repository's own clean filter or signature program would run on the host. When a branch's upstream does not exist on the remote (deleted and pruned, or never pushed, as after cloning an empty repository and committing), `upstreamGone` is `true` and the unpushed list falls back to commits on no remote-tracking ref at all.
`GET /api/sessions/:id/git-diff?repo=<repoRoot>&path=<path>&kind=staged|unstaged|untracked|conflicted` returns the unified diff of one file the panel lists (`{ diff, truncated, binary }`; staged is index vs HEAD, unstaged is working tree vs index, untracked is the whole file as additions). It is what opens when you click a file in the Git panel. `repo` and `path` are matched against the current status rather than trusted, so anything the status does not list is `404`. Read-only: it passes `--no-ext-diff --no-textconv` (no external diff or textconv driver runs), but a repository's clean filters still run, as they do for any `git diff`, which is why a repository a container can write to is never inspected (below). Capped at 400 KB, and refused (`400`) for remote and Docker sessions; a repository at or inside a Docker case workspace is not in the status, so it is `404` here.
**Which repositories.** git finds a repository by walking *up* from the session's working directory, so:
- Inside a repository (or at its root): that one repository, whole (a subfolder reports its enclosing repo, `path` says where it is, e.g. `../..`). A nested repo below it is just an untracked folder to the outer one and is not scanned; start the session inside it to see it.
- **Not** inside one (a folder that holds several projects): every repository found up to **two levels down**, nearest and alphabetical first, at most 12 (`reposTruncated` says when there were more). Dot-folders, `node_modules`, `dist`, `build`, `target`, `vendor`, `venv` and `__pycache__` are skipped, symlinks are never followed, and a repository's own contents are not searched. The list of repositories is re-scanned at most every 30 s; each repository's status is cached for 4 s.
- A repository that merely sits **above** the workspace and is the home folder or higher (a dotfiles repo in `$HOME`, or `/`) is ignored: its dirty files are not this session's work. A workspace that *is* that repository's root is not ignored.
- A worktree (whose `.git` is a file) counts as a repository. A submodule's own uncommitted files are not reported, only a changed submodule pointer.
`data` is `{ state, repos, reposTruncated, checkedAt }`:
- `state: 'ok'`: `repos[]`, each `{ name, path, status }` where `name` is the repository folder's name, `path` its root relative to the working directory, and `status` is:
`branch` (null when `detached`), `upstream`, `ahead`, `behind`, `hasRemote`, `counts` (`staged`, `unstaged`, `untracked`, `conflicted`, `uncommitted` = distinct paths, `stashes`), `files[]` (`path` relative to `repoRoot`, `origPath` for a rename, `index` and `worktree` status letters, `kind`: `staged` \| `unstaged` \| `untracked` \| `conflicted`; a file that is staged *and* modified again appears once per kind), `filesTruncated`, `unpushedCount` (exact) and `unpushed[]` (newest first: `hash`, `author`, `time` in epoch seconds, `subject`), `repoRoot`, `checkedAt`.
- `state: 'not-a-repo'`: no repository here, above (that counts) or within two levels below.
- `state: 'unsupported'` with `reason: 'remote' | 'docker'`: those sessions are never inspected (a Docker workspace is writable from inside its sandbox, and git here would run on the host).
- `state: 'error'` with a short `error` (git missing, timed out, or git's first stderr line with any `user:token@` credentials redacted).
Lists are capped (300 files and 50 commits per repository) while the counts stay exact. A branch with no upstream reports the commits no remote has (`HEAD --not --remotes`); a repository with no remote reports `unpushedCount: 0`, since there is nothing to push to. Concurrent polls of one folder share a single git invocation; `?fresh=1` (what the panel's Refresh button and opening the panel send) skips the short-lived caches, though it still joins a computation already running.
## CLI management
Read and write the CLI registry (`docs/cli-registry.md`). Every **write** route answers `403 FORBIDDEN` while `cliManagementEnabled` is off (the default), and for a non-admin in multi-user mode. A write that would overwrite a `clis.json` which does not parse, or which has group/world permission bits, is refused with `409 CONFLICT` and a message naming the fix; the file is left untouched.
| `GET` | `/api/clis` | none | Every entry, disabled ones included: `id`, `label`, `shortBadge`, `order`, `kind`, `enabled`, `stock`, `installed`, and `installCommand` for a stock entry. Not gated; a non-admin in multi-user mode gets `[]`. |
| `PUT` | `/api/clis/:id` | `{ enabled }` | Toggle an existing entry, stock or custom. `404` for an unknown id; `400 INVALID_INPUT` when disabling a `kind: 'shell'` entry. |
| `POST` | `/api/clis/:id/install` | none | Run a **stock** entry's install command (never a custom one: `400`). `409 CONFLICT` while an install for the same id is running; `422 OPERATION_FAILED` with the output tail when it fails. Never enables the entry. |
| `POST` | `/api/clis` | `{ id, label, shortBadge, binaries, argv, enabled? }` | Create a custom entry. `409 ALREADY_EXISTS` for a stock id or an existing custom id. `enabled` defaults to `true`. |
| `PUT` | `/api/clis/custom/:id` | `{ label, shortBadge, binaries, argv, enabled? }` | Replace an existing custom entry. An absent `enabled` keeps the entry's current state. `400` for a stock id, `404` for an unknown one. |
| `DELETE` | `/api/clis/:id` | none | Delete a custom entry. `400` for a stock id, `404` for an unknown one. |
## MCP server sync
Copies MCP servers between the agent CLIs' own user-level config files (`docs/cli-registry.md`, "MCP server sync"). **Opt-in:** both routes answer `403 FORBIDDEN` while the synced `mcpSyncEnabled` setting is off (the default), and for a non-admin in multi-user mode, because the routes write files in the server user's home. A second `POST` while one is running answers `409 CONFLICT`.
| `GET` | `/api/mcp-sync` | none | Dry run. Same result shape as `POST`, with `applied: false`; nothing is written. |
| `POST` | `/api/mcp-sync` | none | Adds each server a CLI is missing to that CLI's config file. Never edits or removes a server. `500` on an unexpected error. |
Result (`data`):
- `applied` — `false` for the dry run.
- `targets[]` — one per enabled CLI that declares an MCP config: `id`, `label`, `file`, `status`, `error?`, `servers` (names it already has), `added` (names added, or that would be), `skipped` (names its dialect cannot express, e.g. SSE for Codex and Antigravity).
- `status`: `ok`; `absent` (not installed and no config file, so not read or created); `skipped` (the CLI's relocation env var, e.g. `CODEX_HOME`, is set to a relative path in the server's environment, so its file cannot be located safely and is neither read nor written); `unreadable` (the file exists but cannot be parsed safely, so it is not written); `failed` (a read or write error, the file may be unchanged).
- `error` says why a target is not `ok`. A parse failure is reported by position only (`not valid TOML (line 3, column 21)`, `not valid JSON`), never with text from the file.
- `file` honours each CLI's own relocation env var as the server process sees it (`CLAUDE_CONFIG_DIR`, `CODEX_HOME`, `XDG_CONFIG_HOME`, `GEMINI_CLI_HOME`); see `docs/cli-registry.md`.
- `conflicts[]` — names defined differently by different CLIs. Existing definitions are kept; the first CLI's is copied where the name is missing.
- `disabled[]` — names left out because every definition is switched off in its own CLI (codex `enabled = false`, opencode `enabled: false`, antigravity `disabled: true`).
- `unsupported[]` — labels of enabled agent CLIs with no known MCP config file (nothing is guessed).
- Only installed CLIs are listed: one that is not installed is left out, as a supported CLI that is not installed reads `absent`.
The result carries server **names** only, never `env` values, `headers` or file content. Each changed file keeps its previous content as `<file>.codeman-bak` (overwritten by each sync); a file that receives servers carrying `env` or `headers` is left mode `0600`.
## Webhook notifications
Posts the Web Push events to ntfy, Slack, Discord or a generic JSON URL (Settings → Notifications). Off by default. The webhook URL is a bearer secret (anyone holding a Slack/Discord URL can post as it), so it lives in `~/.codeman/webhook.json` (0600), is **never returned**, and is kept out of `settings.json`. All three routes answer `403` for a non-admin in multi-user mode.
| `GET` | `/api/webhook` | none | `{ enabled, kind, scope, hasUrl, urlMasked, lastResult }`. `urlMasked` is scheme + host only. `lastResult` is the last delivery (`ok`, `status?`, `error?`, `at`) or `null`. |
| `PUT` | `/api/webhook` | `{ enabled?, kind?, scope?, url? }` (strict) | `kind`: `ntfy` \| `slack` \| `discord` \| `generic`. `scope`: `attention` (skip "response complete") \| `all`. An absent `url` keeps the saved one; `""` clears it. `400` for a non-http(s) URL, `user:pass@`, a link-local or cloud-metadata target, or enabling with no URL. |
| `POST` | `/api/webhook/test` | none | Sends one message with the saved config, even while disabled. `200` with `data.ok` telling whether the webhook accepted it; `400` if no URL is saved. |
Delivery goes through the same egress guard as web tabs (refused on the resolved address too), does not follow redirects, times out after 5 s, sends the same event for the same session at most once per 3 s, and has at most 5 requests in flight. Error text never contains the URL.
## Diagnostics
`GET /api/doctor[?category=core|office|other]` returns the `codeman doctor --json` report (`platform`, `summary`, `tools[]` with `status``ok` \| `missing` \| `outdated` \| `skipped` \| `error`, `version`, `path`, `installHint`). The probe engine is synchronous, so it runs in a child process of the same entry script, never on the server's event loop (30 s timeout). It names install paths and versions, so it is admin only in multi-user mode (`403`). `400` for an unknown category, `500` if the child produces no report.
## Voice dictation
Browser dictation transcribed through this server's Claude Code login, i.e. the
@@ -18,7 +18,17 @@ Every run mode Codeman can launch — Claude Code, Terminal/Shell, OpenCode, Cod
## The override file
`~/.codeman/clis.json` (instance-scoped through `dataPath()`) holds overrides and custom entries only, never a copy of the stock catalog: `{ "clis": { "<id>": { ...partial entry... } } }`. Objects merge key-wise onto the stock entry, arrays replace wholesale. **The file must be mode 0600**; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you `chmod 600` it. Every reason a file was ignored or an entry dropped is logged once, prefixed `[cli-registry]`, on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read only on restart.
`~/.codeman/clis.json` (instance-scoped through `dataPath()`) holds overrides and custom entries only, never a copy of the stock catalog: `{ "clis": { "<id>": { ...partial entry... } } }`. Objects merge key-wise onto the stock entry, arrays replace wholesale. **The file must be mode 0600**; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you `chmod 600` it. Every reason a file was ignored or an entry dropped is logged once, prefixed `[cli-registry]`, on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read after a change made through CLI management (below).
## Managing CLIs from Settings
App Settings → Agents & CLIs → **CLI management** (`cliManagementEnabled`, default OFF; admin-only in multi-user mode) lists every entry with an installed/not-installed badge and:
- toggles any entry on or off. A `kind: 'shell'` entry cannot be disabled, and the row shows no switch for it. A disabled CLI disappears from the Run menu, the welcome screen and the phone overview, and new session requests for it are rejected.
- installs a missing **stock** CLI by running its shipped install command, after a confirm that names the exact command. Only one install per CLI runs at a time, and the command runs without any `CODEMAN_*` variable in its environment. A custom entry's install command is never executed.
- adds, edits and deletes **custom** entries (id, label, badge, binaries, launch argv). The server re-validates the whole assembled entry through `CliEntrySchema`, so the form cannot bypass the load-time rules.
These are the only writes to `clis.json`. They are serialized, and a file that does not parse or has unsafe permissions is refused rather than overwritten; fix it (or `chmod 600` it) and retry. The HTTP routes are listed in `docs/api-reference.md` under *CLI management*.
## The shape of an entry
@@ -36,7 +46,11 @@ interface CliEntry {
launch: CliLaunch; // the structured argv template
env: CliEnv; // exports, tmux setenv keys, the env-override allowlist
capabilities: CliCapabilities; // what every call site reads instead of the id
// .workDetect?: { promptGlyph, workingLine } — how this CLI's pane shows work
// — how this CLI's pane shows work, work it started in the background, and a turn
// that ended waiting for workers it will resume from
// .modelDetect?: { screenLine, screenLines? }
// (where this CLI's own chrome names the model it runs: SessionState.displayModel)
overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store
}
```
@@ -45,16 +59,75 @@ interface CliEntry {
### Regexes that come from config
Two capability fields carry a regular expression an override file can set: `discovery.version.regex` and `capabilities.workDetect.workingLine`. Both go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
Five capability fields carry a regular expression an override file can set: `discovery.version.regex`, `capabilities.workDetect.workingLine`, `capabilities.workDetect.watchingLine`, `capabilities.workDetect.awaitingLine` and `capabilities.modelDetect.screenLine`. All five go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
`workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused.
`modelDetect.screenLine` names the model a session runs, for the tile grid's and the split pane's headers (`SessionState.displayModel`). It must have exactly ONE capture group, the model, which `schema.ts` checks at LOAD time, and it runs over the last `screenLines` (1 to 4, default 1) non-blank rows of the capture the idle/working probe already takes, joined with newlines so a pattern can anchor on the row above. Like `watchingLine`, the rows are pane text the agent writes most of, so a pattern must anchor on chrome only that CLI draws. The two stock ones, measured on live panes: dsh-TUI's status line on the row under its composer's rounded border (`╰─+╯\n ?(<model>)`, three rows), and codex's ` <model> <effort> · ` footer on its last row. A screen that does not match keeps the last model the session reported; a CLI without the field shows its launch model, if any. Claude needs none: its statusLine exporter reports `model.display_name` on every render. ⚠️ dsh-TUI's first field is the model only while its status bar's model field is on; switched off, it is the next field: the reasoning effort (` medium · <cwd>`), the session mode, or the folder name. So a captured field is not taken when it is one of the CLI's declared `modelDetect.rejectWords` (single tokens, compared ignoring case; dsh lists every effort id its adapters offer and the shipped mode ids) or the session's own working-directory basename (the shared reader's rule, for every CLI). Anything else the pattern captures is the model, so the official `deepseek-chat` / `deepseek-reasoner` ids are read.
`modelDetect.configResolver` names a READER in `src/model-config-resolvers.ts` (a name, never code in config, like a launcher profile) that resolves the model the CLI's own config pins for one session, for while its screen names none (the `config` source of `displayModel`, ranked below any report from the running CLI). It runs at every pane start, attach and relaunch, with the session's own launch config and env, and must be read-only, bounded (probe before read, no synchronous filesystem call) and return the model id alone. The one stock reader, `deepseek-route` (`src/deepseek-route-config.ts`), resolves dsh-TUI's route the way dsh composes it for the session's profile under the session's `DSH_HOME`: the last of `profiles/<profile>/cordis.patch.yml` and `$DSH_HOME/cordis.patch.yml` carrying `config` for the `dsh-tui` row counts, and only when it names both `provider` and `model`. Anything in doubt answers nothing: a half-pinned route, a profile without dsh-TUI, an unreadable, oversized or symlinked-out layer, a file beyond its narrow YAML subset.
`watchingLine` reads a different row of the same screen. A CLI draws it while work the agent
itself started is still running — Claude prints `⏵⏵ bypass permissions on · 1 monitor · ← for
agents` while a monitor, a backgrounded shell or a cloud session is live. Codeman turns that
into `Session.watching`, and an idle prompt from such a session opens already acknowledged,
so a pane waiting for its own background work never raises an alert a human cannot answer.
Group 1 is the label, and a CLI that declares no pattern reports no background work.
Claude's Artifact comment monitor is the one chip that does not count. It waits for a human
to comment on a page the agent published, so Claude's pattern refuses any footer that
carries it, and the idle alert goes out as usual.
Two CLIs declare such a row today, and they put it in different places. Claude writes its
chip on the last row of the screen, so it keeps the default one-row window and anchors on
the `·` its footer joins items with. Codex pins
`1 background terminal running · /ps to view · /stop to close` ABOVE its composer, which
puts the row third from the bottom once the status line and the composer are counted, so its
entry declares `watchingLines: 3` and matches that row end to end. Both were measured
against live panes rather than read out of a binary, which is the standard for adding a
third.
`awaitingLine` covers the quiet pane that is neither idle nor watching: a turn that ENDED
to wait for workers the CLI will resume from by itself. When background agents or an
ultracode workflow are still running at turn end, Claude closes the turn with
`✻ Waiting for 1 dynamic workflow to finish` instead of `✻ Brewed for 1m 18s`, and a pane
showing that row counts as working. ⚠️ Claude renders the row once and never redraws it, so
the words are still on screen after the workers report back and the follow-up turn ends.
The pattern is therefore never run over the whole pane: `isAwaitingWorkers()`
(`session-activity.ts`) walks up from the composer past blank, framed and indented rows and
tests only the first row that starts in column 0, which is the newest transcript row. Claude
starts its own rows in column 0 and the agent's prose never does, so the anchor also keeps an
agent from holding its own session busy.
That label is the one value in the registry that an AGENT can influence, because it comes off
the agent's own screen. Two things keep it honest, and both belong to whoever adds a pattern
for a new CLI. `watchingLabel()` in `session-activity.ts` searches only the last few
non-blank rows, which should be the part of the screen the CLI draws rather than the agent,
and the pattern should anchor on chrome only that CLI can produce. Keep the window as small
as the layout allows, since every row it adds is another row the agent may be able to write.
The label is also ANSI-stripped and length-capped at the source, and every interpolation of
it into markup goes through `escapeHtml()`, since it ends up on a badge and in an approval
card.
The two shipped entries do not sit equally well behind that rule, and the difference decides
what a pattern is allowed to do. Claude's chip is the last row, so its one-row window holds
nothing the agent can write — not even the status line above it, whose command a session
running with permissions bypassed can write into its own `.claude/settings.json`. Codex's row
shares its slot with the last row of the transcript whenever no terminal is running, so a
message ending in that exact line is matched. What keeps that harmless is `hooks: 'none'`: no
hook event from a codex session reaches the approvals inbox, so a forged label costs a wrong
badge and cannot silence an alert. Before giving a CLI both hook signals and a pattern, make
sure its row is one the agent cannot write.
### Three capabilities that must stay independent
`external`, `hooks` and `altScreen` describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. `shell` has no hooks but is **not** an external CLI, so a hooks predicate written as `!isExternalCliMode()` accepted `until=stop` on a shell session and then blocked the caller for their entire timeout. `deepseek` is the mirror image: it IS external and it DOES have hooks.
`test/cli-capability-predicates.test.ts` asserts that no two of the three are equivalent across the catalog, so collapsing them fails the build rather than a user's session.
## The newline chord
`capabilities.newline` (`'line-feed'` | `'esc-enter'`, absent = line feed) is the byte sequence the `send-key` route types into the pane for Shift+Enter. A line feed (`0x0a`, also Ctrl+Enter) is what Claude Code's Ink input reads as "insert a newline"; `esc-enter` (`ESC CR`, the Option/Alt+Enter chord) is there for a composer that ignores a bare line feed. No stock CLI declares it today: the bytes are typed by tmux on the server, so the browser's OS cannot change what a CLI reads, and Codex 0.147.0 was checked to take a line feed (a Shift+Enter that submits is the keypress leak fixed in #520, not a byte problem). A user `clis.json` can set it for a CLI that needs it. It is an enum rather than a byte string on purpose: config never carries bytes that get typed into a pane. Settings → Terminal & Input → **Key tester** prints what a browser reports for keydown/keypress/keyup, to see whether a device is sending what you think.
## Arg-template safety
The composed command line is interpolated into `bash -c "…"` inside tmux, which makes command construction a security boundary. Four independent layers keep config out of it:
@@ -102,6 +175,8 @@ It matches four shapes, not one: `mode === '<id>'`, `mode !== '<id>'`, `case '<i
The allowlist is not a formality. If a branch is about what a CLI can DO it belongs in `CliCapabilities`; the entries that remain are things that are not CLI-behaviour branches at all — chiefly the legacy per-mode `<Mode>Config` objects on `POST /api/sessions`, which are a fact about the public HTTP API rather than about any CLI, plus a few documented cases where `mode === 'claude'` is genuinely the right question (Read My Mind reads Claude's _own_ transcript, so a capability there would be actively wrong).
`test/frontend-cli-no-id-branching.test.ts` is the same guard for the two frontend files the CLI registry's Run-menu consolidation touches, `session-ui.js` and `mobile-overview.js` — deliberately not the rest of `src/web/public/`, whose per-CLI rules stay out of scope for now (see "Fields declared for later" below). Its allowlist keys on `<file>::<expression>` with no line number, since a single unrelated edit to a contended file would otherwise shift every subsequent line and make every entry go stale at once, and each entry additionally carries the exact number of approved call sites — a bare key would let a brand-new branch reusing an already-approved expression land unreviewed. Its comparison shape differs from the backend guard's in one respect: the left-hand side may be any identifier, not only one named `mode`, `id` or `agentType`, because the review of #458 found `const m = this._runMode; if (m === 'codex')` slipping past the named form while the scanned file already filters with `(m) => m !== 'shell'`.
## Two namespaces called `param`
`launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a **launch param**. The **legacy wire field** a param arrives as is a separate namespace, and `launch.legacyConfigAliases` is the only bridge between the two.
@@ -110,9 +185,9 @@ This matters because it is invisible when it is wrong. `capabilities.privilegedP
## Fields declared for later
`shortBadge`, `accent`, `capabilities.echo`, `capabilities.wheelForward`, `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` are **declared but not yet read**. They all describe frontend behaviour, and the frontend is deliberately untouched here: `app.js`, `terminal-ui.js` and `styles.css` keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
`accent`, `capabilities.echo`, `capabilities.wheelForward`, `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` are **declared but not yet read**. (`shortBadge` was on this list until the CLI management list in Settings started showing it.) They all describe frontend behaviour, and the frontend is deliberately untouched here: `app.js`, `terminal-ui.js` and `styles.css` keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
Treat those values as **transcribed, not authoritative** — nothing enforces that `echo.policy` matches `_updateLocalEchoState`'s fallthrough, or that `accent` matches the gradient CSS paints, so re-measure before wiring one up. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; `test/cli-registry-no-id-branching.test.ts` pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
Treat those values as **transcribed, not authoritative** — nothing enforces that `echo.policy` matches `_updateLocalEchoState`'s fallthrough, so re-measure before wiring one up. `accent` is the one exception: it was measured against styles.css on 2026-09-21 (method in the comment above `CLAUDE` in `stock.ts`), though nothing keeps it in step with the CSS either. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; `test/cli-registry-no-id-branching.test.ts` pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
`overlays.credStore` is in the same category, for a sharper reason: the Docker credential-seeding path still reads its own `CRED_STORES` table, because this shape allows ONE store per CLI and the live table needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek declares none here even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is simultaneously the worst thing here to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest.
@@ -184,6 +259,14 @@ A module-level const freezes at first import, and the failure is asymmetric: a C
4. Only if it cannot install with a plain `npm install -g <pkg>`: give it a layer in `docker/agent.Dockerfile` and set `discovery.install.agentImageLayer: { kind: 'dedicated', reason }` on its entry in `stock.ts`. `test/docker-agent-image-coverage.test.ts` requires both, so an exclusion cannot quietly become an omission. An entry with no `npmPackage` needs only the Dockerfile layer, since it never enters the shared npm layer in the first place.
5. That is usually all. If you find yourself wanting to add an `if` somewhere, the guard test will tell you — and the answer is a capability field, or a named profile if it genuinely needs to run code.
## MCP server sync
`capabilities.mcpConfig` (`{ path, format, relocation? }`, `path` relative to the home directory) names the file a CLI keeps its user-level MCP server list in and the dialect it is written in. `src/mcp-sync.ts` reads that list from every ENABLED CLI that declares one, and that is installed or already has the file (a CLI that is neither is reported `absent`, never created), and adds any server a CLI is missing from the others. It writes other tools' own config, so it is **opt-in**: `mcpSyncEnabled` (synced, default OFF) gates `GET`/`POST /api/mcp-sync` (403 while off) and the Settings → Agents & CLIs → MCP servers controls. Declared today for claude, gemini, codex, opencode and antigravity; every format was checked against what the CLI's own `mcp add` writes, except opencode's (documented, not installed to check). A CLI with no entry (pi, grok, omp, deepseek) is not guessed at: it is listed as `unsupported` in the result when enabled. Adding one is a registry entry plus a small adapter in `mcp-sync.ts`, and a verified fixture in `test/mcp-sync.test.ts`.
`relocation` (`{ envVar, path }`) names the env var the CLI itself reads to move that file: claude `CLAUDE_CONFIG_DIR` (`.claude.json` under it), codex `CODEX_HOME` (`config.toml`), opencode `XDG_CONFIG_HOME` (`opencode/opencode.json`) and gemini `GEMINI_CLI_HOME` (`.gemini/settings.json`); antigravity follows `$HOME` only, so it declares none. The var is read from the SERVER process env at call time, which is the env the CLIs Codeman spawns inherit. An absolute value moves the file to `<value>/<relocation.path>`, an empty one counts as unset (as it does for each CLI), and anything else reports the target `skipped` with the reason instead of writing a file the CLI never reads. A per-session relocation (a session's own `CLAUDE_CONFIG_DIR` in `envOverrides`) is not followed: the sync only knows the server's environment.
The rules the module keeps and the tests pin: it only ADDS (a name already defined, in any shape, is never edited or removed; a same-name difference is reported as a conflict); a server switched off in its own CLI is not copied; it never writes a file it could not parse (opencode JSONC with comments, a TOML file with a duplicate table) and re-parses the new text before writing; codex TOML is read with a real parser (`smol-toml`), so CRLF files and inline tables are handled; names such as `__proto__` are ignored and every table keyed by an untrusted name has no prototype; a symlinked config is written through, not replaced; a file that receives `env`/`headers` is left `0600`; only one apply runs at a time; and its result carries server names only, never env values or headers, and never file text: a parse failure is reported by line and column, not by the parser's message (smol-toml prints a code frame of the offending lines and V8's JSON errors quote source, either of which can hold a secret). The schema restricts `path` and `relocation.path` to a relative path without `..`, since sync writes to it.
## See also
- [Agent CLIs](wiki/Agent-CLIs.md) — the user-facing per-CLI guide.
@@ -81,6 +81,10 @@ Antigravity (`agy`) and Grok (`grok`) are the two CLIs not installed from npm (G
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md). Grok is seeded per-file for the same reason (`auth.json`, `config.toml`, `pager.toml` out of `~/.grok`, which also holds `sessions/`, `memory/` and the ~160MB binary under `downloads/`), with the same consequence for `grok -c`. See [`grok-integration.md`](./grok-integration.md). OMP is the one CLI in this family where `sessions/` is the EXCEPTION rather than the rule: `~/.omp/agent/{config.yml,mcp.json,models.yml,settings.yml}` are seeded per-file (the dir also holds SQLite caches and `terminal-sessions/`), but `~/.omp/agent/sessions/` is shared RW like codex's, not seeded, because Codeman reads it host-side for history recovery and `--resume` pinning. See [`omp-integration.md`](./omp-integration.md).
The image can also carry the GitHub CLI (`gh`) and the Azure CLI (`az` + the `azure-devops` extension, in `AZURE_EXTENSION_DIR=/opt/az-extensions` so it stays out of the seeded HOME), wired into the system git config as credential helpers for github.com and dev.azure.com / *.visualstudio.com, exactly as in `docker/server.Dockerfile`. Their sign-ins are seeded per-FILE like pi's: `~/.config/gh/{hosts.yml,config.yml}` and `~/.azure/{azureProfile.json,msal_token_cache.json,service_principal_entries.json,clouds.config,config}`, never `~/.azure`'s logs, command index or extensions. A token kept in a desktop keyring, or in the encrypted MSAL cache az uses on Windows/macOS, is not in those files and does not carry. None of the three is version-pinned; the `--no-cache` rebuild recommended above is also what refreshes them. Both CLIs are opt-in and OFF by default: `CODEMAN_AGENT_IMAGE_INSTALL_GH=1` / `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` in the environment of `scripts/build-agent-image.mjs`, or of the Codeman server for its own auto-build (in the Compose deployment, `environment:` in `docker-compose.override.yml`), become the `CODEMAN_INSTALL_GH` / `CODEMAN_INSTALL_AZ` build args and put that CLI, its extension and its helper entry into the image. Unset passes nothing, so a default build's argv is unchanged and the image has neither. The sign-in seeds follow the same switches, read when a case container is created: `.config/gh` only with `CODEMAN_AGENT_IMAGE_INSTALL_GH=1`, `.azure` only with `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` (`enabledByEnv` in `CRED_STORES`), never merely because the files exist. Seeds are create-time mounts and deliberately not part of the config hash (hashing them would trip the drift gate for every case), so an existing case container picks them up only when it is recreated.
Set `CODEMAN_AGENT_IMAGE_GIT_USER_NAME` and `CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL` together to configure the agent image's Git identity. Rebuild an existing `codeman/agent:base` with `node scripts/build-agent-image.mjs --no-cache`, then recreate Docker-case containers so they use the rebuilt image.
## Quickest path: one-click "Run in Docker"
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
@@ -6,6 +6,10 @@ For the Compose configuration, environment settings, storage migration, and macv
The image includes Claude Code, Codex, Gemini CLI, and OpenCode. Authenticate a CLI from its Codeman session; credentials are never baked into the image.
CLIs installed from **App Settings → Agents & CLIs → CLI management** (DeepSeek Harness, Pi, and the other npm-based ones) go to `~/.local` on the `CODEMAN_APPDATA_PATH` mount, so they survive an image rebuild and a container recreate. Releases up to 1.33.1 installed them into the image instead, so a CLI installed from Settings on one of those has to be installed again once after the rebuild. The same applies to a hand-run `npm install -g` inside a session: it writes to the image prefix (`/opt/codeman-cli`) and is lost on the next rebuild, so use `npm install -g --prefix ~/.local <package>` instead.
It can also include the GitHub CLI (`gh`) and the Azure CLI (`az`) with the `azure-devops` extension, wired in as Git credential helpers, so Clone Repo and `git clone` reach private GitHub and Azure DevOps repositories once they are signed in. Both are off by default; [Turning them on](../docker/README.md#turning-them-on) shows the `docker-compose.override.yml` settings.
## Prerequisites
- Docker Engine or Docker Desktop with Docker Compose v2
@@ -67,7 +71,7 @@ If that directory was created by an earlier root-running image, change its owner
Codeman updates itself from **App Settings → Updates**, as it does on a bare host. The checkout mounted at `/opt/codeman` is the same directory Compose builds from, so the update's `git checkout` and rebuild land on the host and survive container recreation; the restart is the server exiting, which `restart: unless-stopped` turns into a relaunch on the new build.
That applies application code only. A release that changes `docker/server.Dockerfile`, `docker/docker-compose.yaml`, or adds a key to `docker/.env.example` needs the image rebuilt or the container recreated, which a container cannot do to itself. The updater detects each case and refuses with a message naming what changed; run `docker/Start-Codeman.sh` on the host to apply those.
That applies application code only. A release that changes `docker/server.Dockerfile`, `docker/docker-compose.yaml`, or adds a key to `docker/.env.example` needs the image rebuilt or the container recreated, which a container cannot do to itself. The updater detects each case and refuses with a message naming what changed; run `docker/Start-Codeman.sh` on the host to apply those. For a major update, or a base-image change `Start-Codeman.sh` does not fully pick up, `docker/Update-Codeman.sh` rebuilds with no layer cache and clears the two build-artefact volumes before handing off to it (see "Major updates" in `docker/README.md`).
`CODEMAN_REPO_PATH` overrides which checkout is mounted. It defaults to the compose project's parent directory, so it normally needs no setting. Point it at a directory that is not a git checkout and in-app updates are reported as unavailable.
@@ -497,7 +497,7 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never**`--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, five seeded files from `~/.pi/agent`, and three from `~/.grok`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, five seeded files from `~/.pi/agent`, three from `~/.grok`, and, only when their opt-in switches `CODEMAN_AGENT_IMAGE_INSTALL_GH` / `_AZ` are `1`, `~/.config/gh/{hosts.yml,config.yml}` and the sign-in files from `~/.azure`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
@@ -516,6 +516,7 @@ Full feature guide: [`docker-cases.md`](docker-cases.md).
- **Auth is a parallel branch** (`middleware/auth.ts`) that leaves the single‑user path untouched: per‑user scrypt verify (`timingSafeEqual`, timing‑equalized against user enumeration), identity‑carrying cookies, a per‑username failure bucket (a botnet can't brute one account across IPs; one NATed user can't lock out the rest), and a `mustChangePassword` lockbox. The hook‑secret loopback bypass, host guard, and Origin/CSRF guard are unchanged (hooks authenticate the INSTANCE, not a user).
- **Ownership is enforced server‑side only** and fails closed: `req.authUser` (a synthetic admin in single‑user), `findSessionOrFail` returns NOT_FOUND (never 403) for a foreign session, list/SSE/WS/file‑preview/search all filter by `session.owner`, and SSE routing defaults session‑scoped events to their owner (unresolved owner → withheld). The load‑bearing rule is **non‑admin `workingDir` confinement**: a non‑admin's session/one‑shot working dir must realpath‑resolve inside `~/codeman-users/<name>/cases`, checked BEFORE any disk write.
- **Privileged actions are a one‑bit grant** (`canBypassPermissions`, default off): only granted users (and admins) get `--dangerously-skip-permissions` (others are silently downgraded to `--permission-mode auto`), shell‑mode sessions, cron `launchCommand`, and other CLIs' bypass flags. Machine‑level resources (remote/Docker host definitions, tunnel, self‑update, settings writes) are admin‑only.
- **Clone Repo does not lend the server's git sign-in to non-admins.** A clone writes only inside the caller's own case space, so it is not admin-gated, but the server account's git credential helpers (the Docker image's opt-in `gh`/`az` helpers, or any `gh auth setup-git`) are shared by every user. A non-admin's clone and preflight therefore run with `git -c credential.helper=`, which empties the helper list including the URL-scoped entries (`cloneWithoutCredentialHelpers` in `case-routes.ts`, argv pinned in `test/git-clone.test.ts`). This closes the Clone Repo path only: the account's SSH keys still apply to an `ssh://` URL, and a non-admin's agent sessions run as the same account, consistent with the first bullet above. Docker cases are a second route to the same sign-in: with `CODEMAN_AGENT_IMAGE_INSTALL_GH`/`_AZ` on, a non-admin's Docker case with credential seeding on (the default) receives a copy of the server account's `gh`/`az` sign-in, exactly as it receives the Claude and Codex credentials.
- **Admin actions are audited** append‑only to `~/.codeman/admin-audit.jsonl` (acting admin, action, target, IP). Passwords set by an admin create/reset are one‑time (returned once, force change). Under Basic auth, `logout` only truly ends QR‑issued sessions — to lock someone out, disable the account or reset the password (a proper login form is a deferred Phase 6).
---
@@ -531,6 +532,16 @@ A saved dashboard URL renders as a tab, served through Codeman's own origin at `
---
## 10c. Webhook notifications (outbound channel)
Opt-in and off by default: the server POSTs the Web Push events (permission prompts, questions, idle, errors, respawn blocked, crash-loop breaker, Ralph completion) to one URL an admin configures, formatted for ntfy, Slack, Discord or generic JSON. Source: `src/webhook-notify.ts`, routes in `src/web/routes/webhook-routes.ts`. User guide: [`wiki/Notifications-And-Approvals.md`](wiki/Notifications-And-Approvals.md).
- **A second server-side outbound channel through the web-tab egress guard (§10b).** Delivery goes through `webviewFetch`, so link-local and cloud-metadata targets are refused at save time and again on the RESOLVED address at connect time; redirects are not followed (`redirect: 'manual'`) and each send is bounded by a 5 s timeout. Loopback and RFC1918 stay allowed on purpose (a self-hosted ntfy is the point), so **Send test** works as a blind reachability probe (status, refused or timed out, never a response body) for whoever may call it. Web tabs already give that caller full LAN reach with bodies, so nothing new is exposed.
- **The URL is a bearer secret** (anyone holding a Slack or Discord webhook URL can post as it). It lives in `~/.codeman/webhook.json` (0600, tmp+rename), is kept out of `settings.json` (which every logged-in user reads through `GET /api/settings`), is never returned (`GET /api/webhook` gives scheme + host only), and never appears in a log line, a delivery result or an error message.
- **It carries session data to a third party.** Titles and bodies include session names, tool names and error text, all agent- or user-controlled, so Discord gets `allowed_mentions: { parse: [] }` and Slack's `& < >` are escaped: agent output cannot ping a channel. In multi-user mode all three routes are admin-only and the channel is instance-wide: it receives every user's session events, the same reach an admin's own Web Push has, which means non-admins' session details leave the box at the admin's choice.
| **Create New** | A fresh `~/codeman-cases/<name>` with a scaffolded `CLAUDE.md`. |
| **Clone Repo** | A public repo cloned into `~/codeman-cases/<name>` and registered as a case. |
| **Create New** | A fresh `~/codeman-cases/<name>` with a scaffolded `CLAUDE.md`, or with **Create in a custom folder**, a new folder inside a parent you choose, scaffolded the same way and registered in place like a linked case. |
| **Clone Repo** | A repo cloned into `~/codeman-cases/<name>` and registered as a case. Private repos need this machine's own git credentials (see below). |
| **Link Existing** | An existing folder anywhere on disk, registered in place. Nothing is copied or moved. |
Linked cases keep living where they are. Deleting a case in Codeman removes the
registration, and for a linked case that is all it removes.
**Clone Repo never asks for credentials.** It uses whatever the server's own git already has:
an ssh key, or a credential helper such as `gh auth setup-git`. The Docker image can include
helpers for GitHub (`gh`) and Azure DevOps (`az`), turned on in `docker-compose.override.yml`;
then signing those CLIs in once from a shell session is enough. See the private repositories
section of `docker/README.md`. Without credentials a private repo fails straight away with an
authentication error.
**Cases created from scratch are the only copy of that code.** Uninstalling Codeman does not
delete `~/codeman-cases/`, but treat that directory as real work, not scratch space.
| `Ctrl+K` (also `Cmd+K`, `Alt+K`)| Find an open session or start a new one. |
| `Ctrl+W` | Kill the active session. |
| `Ctrl+Tab` | Next session. |
| `Alt+[` / `Alt+]` | Previous / next tab. |
| `Alt+1` to `Alt+9` | Switch to tab N. Physical keys, so macOS Option layouts work. |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move the active tab left / right. |
| `Alt+B` | Collapse / expand the session sidebar, when that layout is on. |
`Ctrl+W` is not a Codeman shortcut: it goes to the terminal, where shells and agent CLIs
use it to delete the previous word. **Close Session** has no key by default; close a session
from its tab, or bind a key to it in App Settings → Shortcuts.
## Terminal
| Shortcut | Action |
@@ -36,6 +39,22 @@ Press `Ctrl+?` in the app for the same list in a floating overlay.
Anything you copy is cleaned on the way to the clipboard: each line loses the padding spaces a full-screen program paints across the rest of the row. Leading indentation is left exactly as it is, so indented code, a `git log` message body and `git diff` context lines paste back the way they looked on screen. An `Alt+drag` rectangular selection is copied exactly as it looks, so its columns stay lined up.
| **Create New** | Starting a fresh project. Creates `~/codeman-cases/<name>` and scaffolds a `CLAUDE.md` into it. |
| **Clone Repo** | Working on an existing public repo. Paste the URL; Codeman preflights it as you type, offers the repo's real branches and tags, and fills in the case name. |
| **Create New** | Starting a fresh project. Creates `~/codeman-cases/<name>` and scaffolds a `CLAUDE.md` into it. Tick **Create in a custom folder** to put it somewhere else instead. |
| **Clone Repo** | Working on an existing repo: public, or private once this machine's git can authenticate (the Docker image can include `gh`/`az` helpers for this). Paste the URL; Codeman preflights it as you type, offers the repo's real branches and tags, and fills in the case name. |
| **Link Existing** | The code is already on disk. Point at the folder, with **Browse** if you would rather click than type. |
The gear next to the picker holds two per-case toggles: **Agent Teams** and
@@ -131,7 +131,7 @@ the tmux server or rebooting the machine.
@@ -46,21 +46,34 @@ supervised by systemd or launchd; npm installs report as non-updatable. See
| Extended Keyboard Bar | Per device | Which accessory bar phones get. Shell sessions override it while they are active. |
| Wheel Scrolls Local History | Off | Keeps the wheel on the local buffer instead of forwarding it to the CLI. |
| Auto Copy Selection | Off | Copies highlighted terminal text to the clipboard the moment you finish selecting it. Ctrl+C still copies on demand. |
| Trim The Pane Margin On Copy | On | Takes the left margin a full-screen agent CLI paints down its own edge off a copy, so the text pastes flush. Each CLI declares its own width, and the strip never exceeds the indent every selected line shares, so nesting is kept. Claude Code and Codex declare a margin; a shell does not. |
| Normal / Bold font weight | xterm defaults | Per device, each slot from 100 to 900. The bundled JetBrains Mono renders every step, so a lighter normal weight makes Claude's bold headings stand out. Applies live to the terminal, both echo overlays and open team panes. |
| WebGL Renderer | On | With a GPU-stall watchdog that falls back to DOM rendering. |
| Gesture Control | Off | Camera hand tracking. Also needs `CODEMAN_GESTURE=1` on the server. |
| Key tester | n/a | A diagnostic that stores nothing. Click the box and press keys to see what this browser reports (key, code, modifiers) for keydown, keypress and keyup, for when a chord such as Shift+Enter behaves differently on one device. Keys pressed there reach no session and trigger no shortcut. |
### Header & Panels
Chips for every optional header control, with a live preview of the resulting header:
Run, Font Size, System Stats, Redraw Terminal, Response Viewer, Away Digest, Session
Manager, Attachments, File Viewer, Multi-monitor, Plan Usage, Lifecycle Log, Monitor,
**Bottom bar** (below the chips): **Git status** shows a small indicator at the right of the
bottom bar, off by default and per device. It reads `● N` uncommitted files, `↑ N` commits not
pushed, `⚠ N` merge conflicts, or `✓` when everything is committed and pushed. Click it for the
Git window; see [Working With Files](Working-With-Files#git-changes). **Git status: group files
by folder** (per device, on by default) shows changed files under collapsed folders in that
window; off lists every file by its full path.
Most default to off. The stock desktop header is system stats, File Viewer, and the gear.
New header controls never appear on phones.
New header controls never appear on phones. Split is desktop-only regardless of this
setting — the button and the feature both stay off below a ~1180px viewport, where two
resizable panes plus their divider have nowhere to go. **Tiles** is desktop-only the same
way; it also enables the `Ctrl+Shift+G` grid toggle on this device. See
[Tile Grid](Tile-Grid).
This section also holds background-agent tracking, including whether to track agents for
every session or only the active tab.
@@ -84,13 +97,22 @@ every session or only the active tab.
### Models
Claude model cards, the 1M context window switch, and the thinking effort segment. The cards
and the switch compose into one model choice, so there is no separate "which one wins"
question.
Claude model cards, the 1M context window switch, the thinking effort segment and the
advisor segment. The cards and the switch compose into one model choice, so there is no
separate "which one wins" question.
Model and effort are both**soft defaults**: the model is written into the case's
`.claude/settings.local.json` and effort is passed at start, so `/model` and `/effort`
inside a session override them at any time.
Model, effort and advisor are all**soft defaults**: the model is written into the case's
`.claude/settings.local.json` and effort and advisor are passed at start, so `/model`,
`/effort` and `/advisor`inside a session override them at any time.
**Advisor** gives new Claude sessions Claude Code's
[advisor tool](https://code.claude.com/docs/en/advisor): a second, stronger model that Claude
consults before committing to an approach, when an error keeps coming back, and before it
calls a task done. A common pairing is a Sonnet main model with an Opus or Fable advisor,
which costs less than running the stronger model all the time. **Default** leaves it to
whatever you picked with `/advisor` yourself. The advisor needs the Anthropic API (not
Bedrock or Vertex), and an advisor that ranks below the session's model is simply not
attached.
**Custom model endpoints** (off by default) adds a saved-endpoint list plus a matching
section to the Run dropdown, for pointing a harness at your own OpenAI-compatible server
@@ -109,11 +131,13 @@ instead of its native cloud backend. See [Custom Model Endpoints](Custom-Model-E
| Nice priority / value | Runs agent processes at a lower CPU priority. |
| Bypass approvals and sandbox | Pi's project trust. Read [Agent CLIs](Agent-CLIs) before enabling. |
| Animated status effects | Cosmetic. |
| MCP server sync | Copies the MCP servers each installed, enabled CLI (Claude, Codex, Gemini, OpenCode, Antigravity) has into the others' own config files. Synced, off by default, admin only in multi-user mode. Turn it on and save, then **Preview** shows what would change and **Sync now** applies it. It only adds missing servers, keeps the previous file as `.codeman-bak`, and leaves a file that receives env values or headers readable by you only. A config dir moved by `CODEX_HOME`, `CLAUDE_CONFIG_DIR`, `XDG_CONFIG_HOME` or `GEMINI_CLI_HOME` in Codeman's own environment is followed. |
### Notifications
Master toggle, browser notifications, push subscription, audio alerts, and the idle
threshold that decides when a quiet session counts as needing you. See
Master toggle, browser notifications, push subscription, audio alerts, the idle
threshold that decides when a quiet session counts as needing you, and the server-wide
webhook (ntfy, Slack, Discord or generic JSON; admins only in multi-user mode). See
[Notifications And Approvals](Notifications-And-Approvals).
### Voice
@@ -130,8 +154,10 @@ Rebinding for the shortcut registry. See [Keyboard Shortcuts](Keyboard-Shortcuts
### System
`CLAUDE.md` template for new cases, default working directory, the image watcher, and
Cloudflare tunnel controls including the tunnel and upload URLs. In multi-user mode, the
**Users** administration entry is injected here.
Cloudflare tunnel controls including the tunnel and upload URLs. The **Diagnostics** group runs
`codeman doctor` on the server and lists the agent CLIs, tmux, Node and the optional office
tools with their versions and install hints (admin only in multi-user mode). In multi-user
mode, the **Users** administration entry is injected here.
## Session Options
@@ -166,6 +192,8 @@ Some things are configured before the server starts, not in the UI:
| `CODEMAN_BASE_URL` | Mounts Codeman under a sub-path behind a reverse proxy that forwards the prefix unchanged. See [Remote Access](Remote-Access). |
| `CODEMAN_MAX_DOWNLOAD_BYTES` | Cap on raw file bodies and downloads. 2 GB by default, `0` for none. |
| `CODEMAN_MAX_REMOTE_FILE_SSH` | Concurrent ssh reads for files in remote cases. 4 by default. |
| `CODEMAN_PATH_PROBE_TIMEOUT_MS` | How long a linked case's folder may take to answer before it is shown as unreachable. 1500 ms by default; raise it for a slow but healthy mount. |
| `CODEMAN_PATH_PROBE_MAX_STALLED` | Unanswered folder checks allowed to pile up before new ones are refused. 2 by default: one below the threadpool size minus one, so it follows `UV_THREADPOOL_SIZE` (4 unless set), and it is never allowed above that ceiling. A check you start by opening one case or session may use the one slot left above it. |
| **Header tab strip** | The default. Wraps to a second row on desktop, scrolls sideways on a phone. |
| **Left sidebar** | A vertical list with a filter box and a live session count. `Alt+B` collapses it to a narrow rail that keeps the status dots and task badges visible. On a phone it is an off-canvas drawer rather than a docked rail. A detailed variant adds the home screen's per-session line (`created 3d ago · working 12m`) and a status pill. |
| **Vertical rail** | The strip turned vertical beside the terminal, resizable, with detailed rows by default. **Vertical Rail Order** sorts it by activity (blocked on you first, then longest running, then most recently quiet), the same order as the home screens; pick *Manual* to get your own order and drag-reordering back. Desktop and tablet only. |
| **Vertical rail** | The strip turned vertical beside the terminal, resizable, with detailed rows by default. **Vertical Rail Order** sorts it by activity (blocked on you first, then longest running, then most recently quiet), the same order as the home screens; pick *Manual* to get your own order and drag-reordering back.**Tab groups:** pick *Move to new group* from a row's ⋯ menu (or Shift+F10 on it) to make the first one; a group header's menu (right-click, Shift+F10 or its ⋯ glyph) renames it (also F2), reorders or deletes it, rows move between groups from their own menu or by dragging with a mouse or pen, and a collapsed group stays collapsed on that device. Desktop and tablet only. |
It is the same list either way, just re-hosted: tab order, drag-to-reorder, the `Alt+1`
to `Alt+9` numbers and every status colour below behave identically in both. The setting is
@@ -48,6 +48,7 @@ One tab per session, in your order, and that order syncs across your devices.
| Yellow tab, blinking | The agent is waiting for input from you. |
| Red tab, blinking | A question or permission prompt is blocking the session. |
| No dot | The session is not running. |
| Muted grey dot plus an `exited (137)` badge | The agent inside the pane has exited, with that exit code (or `exited (signal 9)`). A bare `exited` means tmux saw the pane die but did not report how, which is not the same as a clean `exited (0)`. Detailed sidebar and rail rows read `exited` in their pill. |
| `●` | The session's state: working, idle, waiting on you, needs you (red, and the tile's border pulses), error, ended. Hover the header for how long. |
| logo | Which agent runs in the tile (Claude Code, Codex, DeepSeek, Shell, ...). Hover it for the agent and the model by name. |
| name | Double-click to rename the session. |
| model | The model the session runs, when Codeman knows it: what the agent itself reports (it follows a `/model` switch), else the model its own config pins (DeepSeek's route, shown "from config"), else the model it was started with. Nothing when unknown. |
| `⋯` | The session menu: options, open in a new window, close the session. |
| `⤢` | Zoom: the tile fills the grid; press it again (or `Alt+Shift+Enter`) to get the grid back. |
| `×` | Remove the tile. The session keeps running; close it from `⋯` if you want it gone. |
Click a tile to focus it. The focused tile has the accent border, takes your keyboard, and
is the session every panel follows: files, git status, respawn and Ralph, subagent windows,
voice and image paste. Tabs of tiled sessions carry a small underline.
Drag the thin lines between tiles to resize columns and rows. A tile never gets smaller than
about 60 columns; when the window is too small for all the tiles, the grid shows the focused
one on its own until the window is big enough again.
A tile whose session is not running shows **Not attached** with an **Attach** button. A tile
whose agent exited inside its pane says so instead; close that session from `⋯`.
## Moving tiles
Drag a tile by its header (anywhere but its buttons) onto another tile and the two trade
places. Drop it on an empty slot and it moves there, leaving its old place empty; nothing else
moves, so the empty slot can be anywhere in the grid. The dropped tile takes the focus. Press
`Escape` or let go anywhere else and nothing changes, not even which tile has the focus: a
header focuses its tile when you click it, not when you press it.
With the keyboard, `Ctrl+Shift+Arrows` moves the focused tile one place left, right, up or
down: into the empty slot if that is the place, else trading places with the tile there. It
keeps the focus.
A moved tile takes the size of the place it lands in: column widths and row heights stay
where you dragged the dividers. Tiles do not move while one is zoomed. Where everything is,
the empty slot included, is saved with the grid and comes back on reload.
Closing a tile leaves its place empty when the grid keeps its shape (six tiles to five), and a
new tile takes the first empty place. When the number of tiles changes the grid's shape (four
tiles to five is two columns to three), the tiles keep their places if they still fit, or line
up again from the top left. `Alt+Shift+Arrows` and `Ctrl+Tab` never stop on an empty slot.
| Text and code | Syntax-aware preview. Long files are truncated in plain preview. |
| Text and code | Plain preview with Lines (line numbers) and Wrap toggles in the header. Long files are truncated in plain preview. |
| Markdown | Rendered by default: headings, tables, code blocks with copy buttons, images and links relative to the file (root-relative ones resolve from the workspace root, as on GitHub). Opened from an attachment card, where the file's folder is unknown, relative images show their alt text and relative links show as plain text. The MD pill in the header flips to source. |
| Images | Inline. |
| Audio and video | Inline with a working scrub bar, because range requests are supported. |
| PDF and Office documents | Converted for preview when a converter is available. |
@@ -85,8 +86,9 @@ File paths in a session are links. That works in two places:
render as underlined monospace links.
Clicking one opens it in the preview: images and PDFs render, video and audio play with a
working scrub bar, documents convert, text and Markdown show inline. Log-shaped files open in
the tail viewer instead, which follows a file that is still being written.
working scrub bar, documents convert, text shows inline and Markdown renders. The exception is
a text or Markdown file inside the workspace clicked in the terminal: that opens in the tail
viewer instead, which follows a file that is still being written.
Paths **outside** the session's workspace work too, which matters because that is where most
of an agent's output lands: a screenshot in `/tmp`, a capture in its own scratchpad, a file in
@@ -161,6 +163,41 @@ HEIC images from an iPhone are converted to JPEG on the way in.
When an agent produces a file the UI can show (a chart, a diagram, a document), it can
surface as an artifact attachment rather than a path you have to go and find.
## Git changes
Agents often leave work uncommitted or unpushed. Turn on **App Settings → Header & Panels →
Bottom bar → Git status** (per device, off by default) and the right of the bottom bar shows
the active session's repository: `● 3` uncommitted files, `↑ 2` commits not pushed, `⚠` merge
conflicts, `✓` when everything is committed and pushed.
Click it for a draggable window, in the style of the File Viewer:
- **Uncommitted changes**, grouped as staged, not staged, untracked and conflicted, each with a
status letter (`M` modified, `A` added, `D` deleted, `R` renamed, `?` new, `U` conflict).
- **Not pushed**: the commits no remote has. A branch with no upstream says so, and so does one whose
upstream does not exist on the remote, because it was never pushed or was deleted there ("Upstream
not on remote"), which counts every commit on no remote rather than showing a green tick.
- Files are grouped under their folders, collapsed until you click a folder (a chain of single-child
folders is one row, and the folders you opened stay open when the list refreshes). Turn off
**App Settings → Header & Panels → Bottom bar → Git status: group files by folder** for a flat
list of full paths instead.
- **Click a file** to see what changed in it, as a unified diff with added and removed lines
coloured. Staged files show index versus last commit, not-staged files show working tree
versus index, untracked files show as all additions and deleted files as all removals.
**Open file** jumps to the File Viewer; **Back** returns to the list. A binary file shows a
note instead, and a diff over 400 KB is cut short.
- A session folder that holds several projects gets one collapsible section per repository
found up to two levels down. They all start collapsed (each summary line shows its branch and
what is outstanding), and the ones you open stay open when the window refreshes; an unrelated repository above the workspace (a dotfiles repo
in your home folder) is ignored.
It is read-only and offline: Codeman never fetches, commits or changes the repository, so
"behind" is as of your last fetch. It is not shown for Docker or remote (SSH) sessions, and a repository at or inside a Docker case
workspace is skipped even from a local session (a container can write there, and git would run
that repository's own configuration on the host). The
data comes from `GET /api/sessions/:id/git-status` and `GET /api/sessions/:id/git-diff`
(see the [API reference](https://github.com/Ark0N/Codeman/blob/master/docs/api-reference.md)).
## Gotchas
- **The viewer follows the active session's workspace.** Switching tabs changes what you are
- @opticon454 for four PRs in one batch: webhook notifications (#523), MCP server sync (#521), the Shift+Enter keypress fix (#520) and the newline chord plus Key tester (#522). Every review item was answered in one round, and the merge-order map across all four made landing them together easy.
- @aakhter for the grouped vertical rail (#517) and its ARIA tree and full-row activation (#519), which give the owner tab-layout API its first frontend, and for the iOS IME composition preview (#499), carried through three careful review rounds including the overlay rework in the zerolag package.
- @irisitymichaelgrundberg for per-session Claude models on `POST /api/sessions` (#514) and Codex reasoning effort per session (#515), both kept registry-driven with no CLI id branching.
- @timkjr for keeping Pane B painting during a history pull and its "disconnected" marker last in every interleaving (#524), with an old-versus-new table measured in real Chrome.
**Webhook notifications (#523).** Settings → Notifications → Webhook posts the same events as Web Push (permission prompts, questions, errors, idle) to ntfy, Slack, Discord or any JSON URL, so a headless server can reach a phone with no browser open. Off by default. The URL is a bearer secret: it lives in its own 0600 file (`~/.codeman/webhook.json`), is never returned by the API, and the routes (`GET`/`PUT /api/webhook`, `POST /api/webhook/test`) are admin only in multi-user mode. Delivery goes through the web-tab egress guard (link-local and cloud-metadata targets refused), does not follow redirects, times out after 5 s, dedupes repeats, and neutralises `@everyone`/Slack control characters in agent-supplied text.
**MCP server sync (#521).** Opt-in (`mcpSyncEnabled`, synced, off by default; `GET`/`POST /api/mcp-sync` answer 403 until it is on). Settings → Agents & CLIs → MCP servers previews or copies each installed, enabled CLI's MCP servers into the others' own config files (Claude, Gemini, Codex, OpenCode, Antigravity). It only adds missing servers, never edits or removes one, skips servers you switched off, keeps a `.codeman-bak` of every file it changes, re-parses the result before writing, writes through symlinked dotfiles, leaves files that receive env values or headers readable by you only, and reports same-name conflicts instead of overwriting. CLIs with no known MCP config (Pi, Grok, OMP, DeepSeek) are listed as unsupported. Adds the `smol-toml` dependency to read Codex's `config.toml` safely.
**Claude advisor tool.** Claude Code's experimental advisor (a stronger model the session's main model consults at decision points) can now be set per session: an `advisorModel` field on `POST /api/sessions`, `POST /api/quick-start` and `POST /api/ralph-loop/start` (`fable`, `opus`, `sonnet`, or a full id in those families), and a synced App Settings default under Models → Advisor. It rides the launch's one `--settings` JSON rather than the `--advisor` flag, because the flag exits at launch on any pairing the CLI refuses and would leave a dead pane on every respawn. It is persisted, so respawns and both restore paths keep it, and `/advisor` still switches it in-session. Agents using the codeman skill can give their claude workers one with `CODEMAN_WORKER_ADVISOR=opus`.
**Per-session Claude model (#514) and Codex reasoning effort (#515).**`POST /api/sessions` takes an optional `model` that launches that one Claude session with `claude --model <id>` and writes nothing to disk (`modelOverride` still writes the case default). It is persisted, so both recovery paths relaunch on it. `codexConfig.reasoningEffort` starts a codex session at a chosen effort (`--config model_reasoning_effort=<level>`), and it survives respawn and resume.
**Grouped vertical rail (#517, #519).** When the owner has tab groups (`/api/tab-layout`), the vertical rail draws them as collapsible sections, with collapse remembered per device, the active row always visible, and lineage arcs anchored to a collapsed group's header. The grouped rail is an ARIA tree with one tab stop and the standard arrow-key model. With no groups, the rail is unchanged byte for byte. Editing groups from the browser comes in a follow-up.
**iOS IME composition preview (#499).** On iOS Safari, the text an IME is composing (Japanese, Chinese, Korean, and the predictive composition on English keyboards) is now drawn in the terminal before it commits, inside the local-echo overlay when local echo is on. Inert on every other platform. The `xterm-zerolag-input` package gains `setComposition()`.
**Key tester and newline chord (#522).** Settings → Terminal & Input has a Key tester that shows the keydown/keypress/keyup events the browser reports, to diagnose a device where a shortcut behaves differently. Keys pressed in it never trigger app shortcuts. Shift+Enter's newline chord is now CLI registry data (`capabilities.newline`, line feed by default); no stock CLI changes.
**Fixes.** Shift+Enter no longer submits the prompt after inserting the newline: the key handler swallowed only `keydown`, so xterm's `keypress` still sent a bare `\r` (#520). Claude sessions created at the same moment (`spawn_workers`, a multi-tab Run) no longer fall out of tmux onto the direct-PTY fallback: the statusLine exporter's temp file name collided within one millisecond (#531). Pane B of the split view keeps painting during a history pull, and its "disconnected" marker stays the last line however a close, a pull and a refresh interleave (#524).
**Fixes applied while landing.** Webhooks: the App Settings Save button now saves webhook edits too (a refused URL keeps the dialog open with a warning), Send test saves pending edits first, and a Remove URL button clears a saved URL. MCP sync: a config file that fails to parse is reported by line and column only, never by quoting its content, which can hold API keys; the sync follows `CLAUDE_CONFIG_DIR`, `CODEX_HOME`, `XDG_CONFIG_HOME` and `GEMINI_CLI_HOME` from the server's environment and skips a target it cannot place instead of writing a file the CLI never reads; Preview before saving says to save first; and the MCP group is hidden from non-admins in multi-user mode. Grouped rail: a collapsed group's header shows the red or yellow ring of a hidden row that needs you; layout reads rebuild the rail only when something it draws changed, and failed reads back off (5, 10, 20, 40 s) instead of retrying every 5 s forever; a corrupted collapse preference resets instead of disabling collapse; Ctrl+Shift+{ / } only moves a tab within its own group; tapping a group header or row no longer dismisses the phone keyboard; keys pressed on a row's own buttons no longer move tree focus; and screen-reader positions stay correct after a re-sort. Sessions: `model` on `POST /api/sessions` refuses a value starting with a dash, and `model` or `advisorModel` together with `attachRemoteSession` is now a 400 instead of being ignored; non-Claude sessions no longer report or persist Claude's default model. Split view: a refresh queued behind a history pull no longer leaves a second, stale "disconnected" marker above its replay. iOS IME: a composition on an empty prompt now follows the prompt when output or a resize moves it, and the `xterm-zerolag-input` README documents `setComposition()`. The Shift+Enter and Key tester browser tests now drive the shipped handlers instead of copies.
`setPrompt()` clears the cached prompt position and re-renders if anything is pending, so a mode switch cannot leave the overlay pinned to the old column.
`setPrompt()` clears the cached prompt position and re-renders if the overlay has anything to draw, so a mode switch cannot leave the overlay pinned to the old column.
---
@@ -202,8 +202,9 @@ Implements the xterm.js `ITerminalAddon` interface. It deliberately does **not**
|--------|---------|-------------|
| `addChar(char)` | `void` | Add a single printable character. Auto-detects existing buffer text on the first keystroke. |
| `removeChar()` | `'pending'` \| `'flushed'` \| `false` | Remove the last character. See [backspace handling](#backspace-handling). |
| `clear()` | `void` | Clear all state and hide the overlay. Call on Enter, Ctrl+C, Escape. |
| `removeChar()` | `'pending'` \| `'flushed'` \| `false` | Remove the last character and drop any IME composition. See [backspace handling](#backspace-handling). |
| `clear()` | `void` | Clear all state, the composition included, and hide the overlay. Call on Enter, Ctrl+C, Escape. |
| `setComposition(text)` | `void` | Show text an IME is still composing as an underlined tail after the typed text. Pass `''` to remove it. See [IME composition](#ime-composition). |
### Backspace handling
@@ -217,6 +218,19 @@ Implements the xterm.js `ITerminalAddon` interface. It deliberately does **not**
The cascade order is pending text, then flushed text, then auto-detected buffer text (which is what makes backspace work after tab completion). Backspace "just works" across any combination of typed, in-flight and completed text.
### IME composition
While an input method (Japanese kana, Chinese pinyin, Korean) is still composing, the text is not committed yet, so it is not in `pendingText` either. `setComposition(text)` draws it as an underlined, `aria-hidden` tail right after the pending and flushed text, using the same wrapping and on-screen layout as the rest of the overlay.
// xterm then emits the committed text through onData: add it with addChar()/appendText() as usual.
```
The composition is visual only: it is never part of `pendingText`, `hasPending` or `state`, so it can never be sent. Control characters and line breaks are stripped from it. `clear()` and `removeChar()` drop it. Because `hasPending` excludes it, re-place the overlay after output or a resize with an unconditional `rerender()`, not one gated on `hasPending`.
### Flushed text
"Flushed" means sent to the PTY but the echo has not arrived yet. This happens during tab switches and tab completion.
@@ -242,7 +256,7 @@ Finds text that exists after the prompt but was never typed through the overlay.
| Method | Description |
|--------|-------------|
| `rerender()` | Force a re-render. Call after buffer reloads, screen redraws, resizes and reconnects. |
| `rerender()` | Force a re-render. Call after buffer reloads, screen redraws, resizes and reconnects. A no-op when there is nothing to draw, so it needs no guard. |
| `refreshFont()` | Re-cache font and color properties from the terminal. Call after a font size or theme change. |
### Prompt
@@ -258,7 +272,8 @@ Finds text that exists after the prompt but was never typed through the overlay.
| Property | Type | Description |
|----------|------|-------------|
| `pendingText` | `string` | Unacknowledged text (read-only) |
| `hasPending` | `boolean` | `true` if the overlay has any content |
| `hasPending` | `boolean` | `true` if there is pending or flushed text. Excludes the IME composition, so it can be `false` while the overlay still shows one |
| `composition` | `string` | The text set by `setComposition()`, `''` when none (read-only) |
"description":"Instant keystroke feedback overlay for xterm.js: Mosh-inspired local echo that removes perceived input latency over SSH, tunnels and other high-RTT connections",
"description":"Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
// An empty list would make every browser test look excluded: fail rather than pass vacuously.
console.error('✗ `vitest list` reported no test files; refusing to pass on an empty CI set.');
process.exit(1);
}
// Same vacuous pass, one step removed: a listing whose paths never match the tree. This guard,
// not `vitest list --json`, is the answer to format drift: the JSON form prints absolute paths
// that would need canonicalizing against ROOT (symlinked checkouts), and its shape can drift too.
if(!listingMatchesTree(ciFiles,testFiles)){
constsample=[...ciFiles].slice(0,3).join(', ');
console.error(
`✗ none of the ${ciFiles.size} paths \`vitest list\` reported (e.g. ${sample}) is one of the ${testFiles.length} test/**/*.test.ts files; its output format has probably changed.`
);
process.exit(1);
}
constleaked=findLeaks(browserTests,ciFiles);
if(leaked.length>0){
console.error(`✗ ${leaked.length} browser-driven test file(s) are NOT excluded from ${CI_CONFIG}:\n`);
'That repository needs authentication. Codeman clones without credentials, so private repositories have to be cloned outside Codeman and added with Link Existing.',
"That repository needs authentication. Codeman never asks for credentials, so sign this server's git in first (for example `gh auth login` or `az login` from a shell session; the Docker image can include both, see docker/README.md), or clone it outside Codeman and add it with Link Existing.",
stderr: clean,
};
}
@@ -790,7 +813,8 @@ export function isGitAvailable(): boolean {
*/
exportasyncfunctionprobeGitRemote(
repository: string,
timeoutMs=GIT_LS_REMOTE_TIMEOUT_MS
timeoutMs=GIT_LS_REMOTE_TIMEOUT_MS,
opts:{withoutCredentialHelpers?: boolean}={}
):Promise<GitRemoteProbe>{
if(!isGitAvailable()){
return{
@@ -800,7 +824,7 @@ export async function probeGitRemote(
failure: classifyGitFailure('',false,'ENOENT: git not found'),
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.