Redraw (Ctrl+Shift+R and the header button, restoreTerminalSize) on a
focused tile or the split's Pane B called tile.fit({ force: true }) and
always toasted "Terminal restored to CxR". In TerminalTile._sendResize,
force only skipped the client-side dedupe: the frame carried no `f`, so
Session.resize skipped a size equal to the one it last applied and the
server did nothing. And _sendResize returned silently with the socket down
or the session popped out to its own window, while the toast still
claimed success.
Both halves are fixed, the first as parity with the primary pane:
- a forced fit now sends `f: true`, the flag the primary's sendResize sets,
which the server honours (ws-routes reads msg.f, Session.resize then runs
tmux resize-window and the PTY resize at the same size);
- fit() and _sendResize() return whether a frame went out, and
restoreTerminalSize toasts success only then. Otherwise it says why, as
the primary branch does: "sized by its own window" for a detached
session, "not connected" while the tile's socket is down (it announces
its size again on reopen), and the primary's "Could not determine
terminal size" for a pane that measured nothing.
What this does not claim: a forced resize to the size the PTY already has
changes no geometry, so it is not a cure for a garbled tile whose PTY
already matches; the primary pane's forced resize has the same limit. It
matters when the server's recorded size has drifted from the tmux window.
Tests: the forced frame carries f:true (and plain ones do not), fit()'s
return value on send, dedupe, detached and closed-socket paths, and Redraw
end to end on a real tile (sent, socket down, popped out), plus the three
no-success toasts in focused-pane-shortcuts.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A tile's refresh (a {t:'r'} or {t:'c'} frame, every reconnect) wiped the
pane with a synchronous xterm clear() at the load's turn, BEFORE its fetch,
and wrote live frames straight through the fetch and the replay. That is
the replay clear CLAUDE.md "Terminal resilience" forbids: bytes still
queued in xterm are parsed after a synchronous clear and fuse into the
snapshot, and clear() keeps the cursor's row, column, SGR and margins, so
the capture (raw rows, no home) started wherever the cursor sat. A failed
or empty fetch left the tile blank.
The refresh now runs in the primary pane's order (_onSessionNeedsRefresh,
_resetTerminalForReplay):
- fetch first, so the tile keeps its last frame through the round trip and
through a grid tile's wait in the load queue;
- from the response on, live frames are held in _liveQueue with their
arrival time, as _pullHistory already did, and the body read of a bounded
window (grid tile, shell) gets the pull's 10 s budget, while Pane B's
unbounded full=1 keeps the request's own budget;
- then the queued in-stream \x1bc immediately before the replay;
- then the held frames that arrived after the response (_flushLiveQueue,
now shared with _pullHistory), then the owed marker.
A failed, aborted or empty fetch writes nothing and resets nothing.
The _stampMarkerIfOwed guard for a pending trailing refresh stays (that
refresh settles the marker itself either way); only its rationale changed.
The fake xterm now treats an in-stream RIS like clear() in its row
emulation.
Tests: the ones that counted clear() calls on the refresh path now count
the in-stream reset instead, assert it sits right before the replay and
that clear() is never called (unit single-flight block, the marker
ordering tests, the reconnect test, the grid {t:'r'} and marker tests, and
the scroll test's server-clear overflow case, which now goes through a
refresh). New: the screen is untouched on a failed or empty fetch and on a
failed body read (held frames written in order), frames before the
response are written through and later ones held behind the replay, the
cutoff drops frames the capture covers, the body budgets, and a grid tile
keeps its last frame through its own capture's round trip.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The server sends {t:'c'} from one place only: a fresh Claude pane's first
prompt (Session.startInteractive), meaning "refresh after startup". The
primary pane answers it with a refetch and replay (_onSessionClearTerminal),
and stands aside while the grid is open, so the tile's own handling was the
only one that ran. That handling was a bare xterm clear(), which keeps only
the cursor's row and drops the banner, a resumed transcript and all
scrollback. An idle Claude never repaints static rows, so a Claude session
Run into the grid, or Attached in a tile, came up as a near-empty tile.
_onLiveClear() now calls _refreshBuffer(), the {t:'r'} path: single-flight,
coalesced into one trailing refresh behind a load already running (a shell
pull's held frames included), and paced by the grid's TileLoadQueue. The
queued {clear:true} entry and its branch in _pullHistory's flush are gone,
along with the _clearTerminal helper they used.
Tests: two unit tests pinned the bare clear (a clear frame queued in order
during a pull, and one applied at once before the capture); they are
replaced by tests that the frame coalesces behind the pull and refetches,
plus a socket-level {t:'c'} test, a coalescing test, and a grid test that
the frame waits its turn in the load queue.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Grid tiles and the split view's Pane B now get two fixes the primary pane already had in this release:
- opencode's hollow-buffer wheel paging and its click reports (#555)
- the Android soft-keyboard controller, so autocorrect no longer duplicates a line in a tile (#541, with the #441 drain)
Gate green on the branch tip 7409ad26: 525 files and 10243 tests, plus the touched and adjacent browser files.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two lines the #541 parity commit edited still called the split "desktop-only"
and listed "Phones and tablets" as a tile grid non-goal, right next to the new
note that a wide Android tablet clears the gate. The same commit documents
the gate as width alone in terminal-tile.js and architecture-invariants, and
that is what the code does: terminal-split.js and canOpenTileGrid in
tile-grid.js only compare window.innerWidth with SPLIT_PANE_MIN_WIDTH.
CLAUDE.md's Split-pane line now reads "desktop-only at 1180px (width alone,
so a wide Android tablet clears it)", in step with the Tile grid line, and the
tile-grid-plan non-goal names phones only and says a wide tablet or an
unfolded foldable in landscape can reach the grid, pointing at the keyboard
exception below it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The tile's hand-encoded click report went through _handleDesktopTerminalClick
and _sendSyntheticSgrTap to _sendInputAsync, so it took a seq, was persisted
and would be redelivered after a reload. The documented TerminalTile rule
(CLAUDE.md, Split-pane sessions) is that only typed input enters that queue
and focus/mouse reports go out ephemeral, and the tile's own _onTerminalData
says the same. Before #555 an opencode tile's click went through xterm's
encoder and that ephemeral path. A click still unacknowledged when the page
reloads, or sent during a server restart, could be replayed onto a later
screen, where a press+release can pick a dialog option.
_sendSyntheticSgrTap now takes an opt-in `ephemeral` field on its target and
sends through _sendInputEphemeral when it is set; _handleDesktopTerminalClick
passes the target through unchanged, and TerminalTile._installClickListener
sets it. Without the flag nothing changes, so the primary pane's own click
and touch tap reports stay on _sendInputAsync exactly as before (whether the
primary pane should also go ephemeral is a separate question, out of scope
here).
Tests: the tile case now requires a frame with no seq and nothing pending in
the reliable queue, and the targeted-click case in terminal-touch-tap spies on
both send paths: a target with the flag goes ephemeral, an untargeted click
and an untargeted tap stay durable. Dropping `ephemeral: true` from the tile,
or the branch in _sendSyntheticSgrTap, turns the matching test red.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A tile counts as hollow when every row above its screen is its own overflow
(baseY minus _overflowRows is 0), so unlike the primary pane, whose hollow
buffer has baseY 0, its viewport can sit above the bottom while it is hollow:
Shift+PageUp, a scrollbar drag or a wheel during the first replay leave it up
there. _maybePageCliTranscript never looked at the viewport, so every wheel,
wheel-down included, was turned into PageUp/PageDown and swallowed. xterm never
scrolled back, the stale rows stayed on screen while the CLI paged out of
view, and clicks were dropped too, because the click report refuses an
off-bottom viewport.
The tile now pages only while _terminalViewportAtBottom holds for its own
terminal, checked before the pending travel is touched. Off the bottom the
wheel stays with xterm, so a wheel-down brings the viewport home and paging
resumes from there. The primary pane is unchanged: its hollow test already
implies a viewport at the bottom, which the twin comment now says.
Tests: a unit case for a tile hollow by the discount with its viewport above
the bottom (no page key, no preventDefault, and no travel carried over once
back home), and the real-browser case now scrolls a hollow tile up and proves
a real wheel-down scrolls xterm home with no page key sent, then pages again.
Both go red with the gate removed, and the unit case also with the gate moved
below the pending-travel update.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
#541 fixed Android autocorrect duplicating the typed line in the primary
pane: xterm's keyCode-229 textarea diff is append-only, so an autocorrect on
space (delete a word, insert the corrected one) sent the whole line again.
The fix, an edit-based diff that sends one DEL per deleted code point and
then the inserted text, lives in terminal-keycode229-recovery.js together
with #441's next-keydown drain (a character committed in the same task as
Enter goes out ahead of the \r) and the original orphaned-insertText
recovery. Only the primary pane created that controller, so a grid tile or
the split's Pane B still ran xterm's stock behaviour. Both are gated on
width alone (1180 CSS px), which a wide Android tablet clears.
TerminalTile now creates its own controller in connect(), after the xterm
opens and before the first await, handed this tile's textarea, this tile's
CompositionHelper and _onTerminalData as the send path, so recovered bytes
go to the tile's own session through the exactly-once queue. As in the
primary pane, handleKeyEvent runs first in the custom key handler, above the
keyCode-229 early return, and notifyCanonicalData sits in the onData lambda,
gated on the same two CodemanTerminalInput predicates, never in
_onTerminalData, which the recovered bytes also take. destroy() tears the
controller down before disposing the xterm, which restores xterm's own diff
and removes the capture listeners. No mode or device gate, matching the
primary. The module itself is unchanged apart from its header; terminal-ui.js
gains only a comment naming the twin.
Tests: test/terminal-tile-input.test.ts now loads the real module into its
vm harness (with window timers, without which create() would silently throw
and every test would run against no controller) and drives a fake
CompositionHelper carrying xterm's own append-only diff. It covers install
and restore on the tile's own helper and textarea, autocorrect sent as an
edit (with a control reproducing the device-log duplicate), the last
character and an autocorrect each followed by Enter in one task, a
self-rescued 229 key delivered once, the onData gate ignoring query replies
and focus reports, two refused inserts after one keydown both recovered,
robustness when the controller throws, per-tile controllers, and a source pin
keeping the call above the early return. Removing the create, the
handleKeyEvent call, the notify, its gate, or the destroy each turns at least
one of them red, as does moving the notify into _onTerminalData. The browser
suite gains a TerminalTile block in
test/terminal-keycode229-recovery.browser.test.ts (real xterm, trusted
execCommand input, chunks asserted to address the tile's session, with a
destroyed-controller control).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
#555 made the primary pane page opencode's transcript with PageUp/PageDown
from the wheel, because opencode draws in place on the alternate screen and
leaves the browser's buffer with no scrollback. A TerminalTile (a grid tile,
the split's Pane B) left every wheel to xterm, so in an opencode tile the
wheel scrolled nothing, or only stale rows.
The tile now runs the primary pane's own gates aimed at itself (its terminal,
its session, never the active one): xterm's tracking mode, the Claude
forwarding gate, then the hollow-buffer test. A wheel that passes them is
consumed in the capture phase and turned into PageUp/PageDown through the
shared pageKeysForTravel math, coalesced per tile (40 ms, 512 bytes, the twin
of the primary pane's queue) and sent ephemeral on the tile's own socket.
Every other wheel stays with xterm as before, the shell history pull
included. The file names no CLI: the mode rules stay in terminal-ui.js, and
terminal-tile.js joins the frontend no-id-branching guard.
A plain port of the primary's baseY === 0 test would almost never fire in a
grid. A tile's first capture is taken at the PTY's previous size (usually the
taller primary pane's) and written into a shorter xterm, and its own
row-shrinking fits (zoom-out, divider drags, tile count changes) push more
rows above the screen. The tile counts those rows as its own overflow: all of
them after a load whose capture held a single screen (the server's
captureRows), plus whatever a local fit or a PTY geometry report pushes up,
reset by a clear and clamped to baseY. The paging gate gets baseY minus that
count. Output that scrolls real lines still counts as history, so the tile
stops paging there.
#555's other half, stripping opencode's mouse DECSETs so a drag selects text,
is server-side and already reached tile sockets. It also left the tile's
xterm unable to encode opencode's clicks, so the tile now installs the
primary pane's desktop click report (bubble phase, gated on the session's
cliMouseTracking, the tile's own link hover and selection). Both listeners,
the flush timer and the page-key state are torn down in destroy().
Still out of scope, as the fileoverview now says: touch paging (tiles have
no touch path) and SGR wheel forwarding to Claude's fullscreen renderer
(tile-grid-plan follow-up 4), so a fullscreen Claude tile keeps leaving the
wheel to xterm.
Tests: test/terminal-tile-scroll.test.ts drives a real tile in the vm
harness (session targeting, every no-page case, accumulation, the cap,
coalescing, byte parity with the primary pane, the overflow discount through
a load, a fit, a geometry report and a clear, the click report and destroy);
the discount cases fail with it removed. The fake xterm gains opt-in row
emulation. test/terminal-tile-scroll.browser.test.ts checks the same model
against a real xterm with trusted wheel events (browser suite, not the gate).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The primary pane's hollow-buffer paging (#555) and its desktop click report
read this.terminal and this.activeSessionId throughout, so a second pane (a
grid tile, the split's Pane B) could only get them by copying the gates and
their CLI rules. They now take an optional trailing target instead, the
pattern registerFilePathLinkProvider, copyTerminalSelection and
_handleImagePaste already use for tiles:
- _shouldForwardWheelToApp(ev, { terminal, sessionId })
- _localScrollbackIsHollow({ terminal, sessionId, localRows }), where
localRows stands in for baseY so a tile can discount rows it pushed above
the screen itself
- _handleDesktopTerminalClick(ev, { terminal, sessionId, linkHovered }),
_sendSyntheticSgrTap(x, y, target), _shouldReportMouseToCli(sessionId),
_terminalViewportAtBottom(terminal) and _clientPointToCell(x, y, terminal)
Every field left out means the primary pane's, and every existing caller
passes none, so the primary pane behaves exactly as before and its
grep-pinned call sites are unchanged. The mode list for hollow buffers and
the claude >= 2.1.187 forwarding gate stay in terminal-ui.js alone.
The stateless math moves into two pure exports on CodemanTerminalInput,
wheelDeltaLines and pageKeysForTravel, which _wheelScrollLinesFloat and
_maybePageCliTranscript now delegate to. Comments on both sides name the
tile's twins (the page-key pager and the 40 ms coalescer).
Tests: the exports agree with the primary pane's methods and bytes, the gates
read the target's session, buffer, rows and tracking mode rather than the
active ones, and a targeted click uses the target's geometry, selection,
scroll position and link hover.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Thirteen contributor PRs, each re-checked against its GitHub head, merged
with its own merge commit and landing fixes, reviewed, and gated together
(524 test files, 10194 tests): #559, #552, #550, #556, #551, #546, #542,
#540, #555, #543, #541, #502, #432.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Move the nightly browser-suite cron from 03:23 to 03:29 UTC, authored as
Ark0N. GitHub sends scheduled-run failure notices to whoever last modified
the cron line, but its docs do not say whether that means the commit author
or the pusher. The previous cron edit (02c65e98) was authored under the
maintainer identity, whose noreply@anthropic.com address GitHub resolves to
the unrelated login "claude", so under the author reading the nightly's
failure notices would never reach the maintainer. With this commit the
author, the committer and the pusher are all Ark0N, so every reading lands
on the maintainer. Only the minute changes; no doc or test names a clock
time.
After the first scheduled run on master, confirm with
gh api 'repos/Ark0N/Codeman/actions/runs?event=schedule&per_page=1'
--jq '.workflow_runs[0].actor.login', which should print Ark0N.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The same Tile Grid subsection and GIF as the English README, in the zh-CN
README under 多会话仪表盘, using the app's own zh-CN terms (平铺 for the
feature and its button, 窗格 for a tile, 平铺网格 for the grid).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A Tile Grid subsection under Multi-Session Dashboard with an 800px GIF of the
real Tiles button opening six live sessions side by side and closing back to a
single session, recorded from the 1.36.0 release candidate. Identity text in
two terminals is covered by bars; the GIF was OCR-checked for it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Run dropdown scroll browser test built WebServer on the fixed port 3290,
which the static port guard (test/test-ports-guard.test.ts, from #556) rejects
for any file outside its shrink-only legacy list, so the CI gate failed on the
landing branch. The test now binds an ephemeral port with new WebServer(0, ...)
and navigates to server.boundPort, and its header records the new convention.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pin the App Settings codexModel guard to the schema. The 28df21f4 landing fix
added a client check in saveAppSettings() that copies the pattern of
SettingsUpdateSchema.codexModel, so one bad character no longer 400s the whole
.strict() settings PUT behind a "Settings saved" toast. Nothing tied the copy
to the schema: a looser copy would bring the silent 400 back, and a stricter
one would refuse valid model ids.
The new test extracts the client pattern from the saveAppSettings() body,
checks it agrees with the schema on eight samples (empty, dotted, slashed,
colon, space, semicolon, leading dash, non-ASCII), and asserts the guard runs
before the localStorage write. Length is left out on purpose, since the
input's maxlength="100" covers .max(100). Both a loosened pattern and a guard
moved after the write turn the test red.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Applies the review's landing list for the native-wrapper window bridge, with the verifier corrections.
detachSession now refuses before asking the host when there is no window channel (no BroadcastChannel). Without the channel there is no roll-call liveness, so a hosted tab could never re-dock and would stay detached, and excluded from tiles and split, until the session ended. The guard sits before the host call so a channel-less host never gets a native window and a window.open as well.
A single resolver, tabDetachButtonEnabled(), now lives in app.js next to hasHostWindows() and decides the host-aware pop-out default for the tab icon, App Settings and the tab action menu (which also serves the tile grid's menu). Before this the menu read the raw setting and hid "Open in a new window" under a host. The menu and both settings-ui.js sites call it optionally with a fallback, because test/session-sidebar-ux.browser.test.ts loads tab-rail-resize.js onto a bare CodemanApp without app.js, and a bare call would throw before the menu is appended.
The "Close window" button on the solo session-gone overlay goes through _closeSoloWindow(), as the re-dock button already did, so it works in a host window.
openWebviewExternal no longer falls through to window.open when the host refuses (in a WebView that can replace the dashboard page); it toasts instead, like the session and file-preview paths.
The hasHostWindows and openInHostWindow JSDoc now say what the code does: anything but false counts as opened, and a saved web tab passes its own origin.
docs/versioning-policy.md lists the window.CodemanHost bridge under experimental surfaces, so it does not read as a stable contract until the wrapper docs section lands.
test/host-window-detach.test.ts gives the harness a live window channel (Object.create leaves it undefined, which the new guard would refuse) and pins the no-channel refusal.
The per-PR changeset is removed; the release writes one consolidated changeset.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Skip zero-size spreadsheet cells in renderTile. On sheets past the
8,000,000 px scroll cap, a cell clipped to nothing at the spacer's edge
(or a visible row or column the worker clamped to 0 px) was still created,
and the cell padding and border drew it as a 5 px box below the spacer
that grew the scroll area. The size is now computed before the element is
created and such cells are skipped, the same way the heading loops already
skip 0 px rows and columns. spreadsheet-preview.js is not an input of
SPREADSHEET_ASSET_VERSION, so the asset token stays valid.
Add zh-CN entries for the static spreadsheet preview strings (loading,
too large, no visible worksheets, empty worksheet, the warnings label,
timeout, failure, the four parser start and message failures, and the
unavailable message from panels-ui). The file-preview body is not a
skipped surface, so the exact-match entries apply with no code change.
The worker's admission refusal messages and the dynamic status message
stay English for a follow-up.
docs/security-architecture.md described the attachment gate as a
6-extension allowlist; it now names SUPPORTED_ATTACHMENT_EXTENSIONS in
src/attachment-registry.ts and what it covers, including the xlsx this
PR adds.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A composition that ends in the same task as an Enter keydown was sent
twice. The keydown settled the pending edit (sending the composed word
and setting _dataAlreadySent), then xterm's own keydown finalized the
composition synchronously through _finalizeComposition(false), which
ignores _dataAlreadySent and sent the word again. settleEdit() now takes
the keydown event and, while xterm has a composition in flight
(_isSendingComposition), leaves the text to xterm for any key that makes
it finalize synchronously. On 229, CapsLock and the modifiers xterm keeps
the composition on its async path, which honours _dataAlreadySent, so the
edit still applies there. The waiting timers are cleared before that
early return, so a timer cannot fire after Enter's textarea clear and
send a run of DELs.
The guard sits in settleEdit(), not in applyEdit() as the bot proposed.
In applyEdit() it would also silence the timer path, where xterm always
finalizes asynchronously and skips _dataAlreadySent, so a non-composing
character typed just before a composition (the x in xword) would be lost
where master and the PR head both deliver it.
Two unit tests pin it, both measured: one fails without the guard
(the Enter keydown sends 'ab word' instead of 'ab '), and one fails with
the guard moved into applyEdit() (the timer path sends 'ab ' instead of
'ab xword'; a 229 settle must also still send the edit).
The xterm private-API guard test now also checks the bundle still ships
_isSendingComposition, and names it in its failure message and comment.
CLAUDE.md: the surviving #441 sentence said a keydown decides before
xterm's 229 rescue has run and that Enter's clear makes the pending diff
emit nothing. Neither holds any more (the edit diff is settled first, and
master already sent one DEL there), so it now says the edit diff is
settled first at that keydown. The PR's sentence notes the composition
exception.
The PR's own changeset is removed; its text goes into the single
combined release changeset.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A lone repository git could not read rendered as a clean, empty one. The
panel took its single-repository view whenever the overview held one row,
and the error row only exists in the list view, so it showed "Nothing
uncommitted / No remote configured" with an empty header while the
indicator said "? 1". The single-repository view now needs a readable
repository and an untruncated overview; anything else takes the list view
(headed "1 repository"), and the tooltip names the unreadable repository
instead of saying "no branch".
The same condition covers a limit of 1 in a folder of several projects,
now that max repositories can go down to 1: the one row shown keeps the
"Showing the first" notice instead of looking like the only repository.
A repeated timeout query parameter reaches the route as an array, and
calling trim() on it answered 500 with an internal message, before the
ownership check. The route now treats a non-string timeout as "default",
like an empty or absent one, and the route test pins it.
The browser test gains the lone-unreadable-repository case (error row,
no "Nothing uncommitted", "? 1", tooltip names the repository) and the
truncated single-row case. Both fail against the unfixed panel.
The git timeout input steps by 1, not 5: the save accepts any whole
number of seconds and step 5 flagged values like 7 as invalid.
Docs: api-reference says repoLimit is only present in the
folder-of-projects case, the Settings Reference and Working With Files
glyph lists mention "? N", and the module header says the repository
count is the caller's maxRepos.
The PR's own changeset is removed; its text goes into the single
combined release changeset.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment and doc corrections that #555 made stale, no behaviour change.
- stock.ts: the grok and omp altScreen comments compared their strip to
opencode's, which is now strip-mux-and-mouse rather than the narrow
strip. Grok now says it shares antigravity's strip until measured, and
omp drops opencode from its comparison.
- terminal-ui.js: the touch-tap comment named Claude/Codex/Gemini as the
stripped modes, but the gate is now the cliMouseTracking flag alone and
covers opencode too, so it names the two stripping flavours instead.
- src/types/session.ts: the cliMouseTracking JSDoc (the flag the browser
now gates on exclusively) listed only claude/codex/gemini; it now names
the strip-full and strip-mux-and-mouse modes, including opencode under
tmux.
- src/session.ts: the usesMux getter doc now names isMuxMouseStripMode,
since the replay strip passes usesMux to it as well.
- docs/architecture-invariants.md: the narrow-strip list gains
grok/deepseek/omp (matching the PR's own CLAUDE.md line), the
"must REMEMBER" heading covers both DECSET-stripping flavours, and the
cliMouseTracking writer is described as the full-or-mouse branch it
really is.
- docs/wiki/The-Dashboard.md: the user manual said every non-Claude CLI
scrolls locally; opencode's wheel and swipes now page its conversation.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Move the nightly cron from 03:17 to 03:23 UTC. GitHub sends scheduled-run
failure notices to whoever last modified the cron line, and after the merge
that is the contributor, so a maintainer commit has to touch it. The
docs below give no clock time, so they cannot drift from the cron.
Drop the "Keep the failure artifacts" step and the blank line before it.
No browser test writes test-results/ or screenshots-echo-diag/ (only the
ignore files name them), and if-no-files-found: ignore made the step upload
nothing without a word. The run log already carries the failure output.
Reword the workflow header. Drop the claim that the skipped suite let two
semantically conflicting PRs merge green: that incident came from
test/mobile/keyboard.test.ts, which this job does not run. Correct the
codex-predictive-echo note: the test uses a fake key in a throwaway
CODEX_HOME and skips itself when codex is missing, so it needs a codex
binary, not an authenticated one.
opencode-resize: record WebSocket resize frames under the socket's own URL
instead of appending '#' + the session id. The URL already carries
/ws/sessions/<id>/terminal, and the suffix let toContain(sessionId) pass for
a resize sent on any session's socket, the bug this test exists to catch.
Reduce the six session-id extractions (opencode-resize and perf-browser) to
data.data?.session?.id. POST /api/sessions always answers in the
{ success, data: { session } } envelope, and the dead fallbacks are what
hid the original breakage.
split-pane-terminal: restore the browser config's 60 s test timeout (the
added 20000 ms override tightened it), and replace the comment that blamed
Codeman's post-create clear. Under vitest the session is an echo PTY, so
that clear comes back as text; the real fix is useMux:false, since a plain
prompt otherwise goes through tmux send-keys, which test mode no-ops.
CLAUDE.md: the CI note now says the gate excludes the Playwright tests in
BROWSER_TEST_GLOBS instead of a stale count of 14, and names
browser-suite.yml; the Testing warning says the browser suite runs nightly.
CONTRIBUTING.md gets the same one-line pointer under Tests.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Follow the keyboard: the menu's cap now uses var(--app-height, 100dvh)
instead of 100dvh. The viewport meta has no interactive-widget, so dvh does
not shrink for an on-screen keyboard, while MobileDetection always sets
--app-height to the visual viewport height and KeyboardHandler keeps it there
while the keyboard is up (body and .app already size the same way).
Subtract both safe areas from the cap: in the iPhone home-screen app the
phone header grows by the top inset and the toolbar sits above the bottom
inset, so without them the top of a full menu slid under the fixed header.
Add overflow-x: hidden. overflow-y: auto computes overflow-x to auto, so a
long nowrap custom-endpoint label showed a horizontal scrollbar that
touch-action: pan-y cannot pan (the same trap the file documents for
.run-mode-history).
Raise the toolbar while the Run menu is open, by adding
.toolbar:has(.run-mode-menu.active) to the existing popover raise rule. The
menu is trapped in the toolbar's stacking context, so on a touch device the
keyboard accessory bar (z 51) and the visible CJK input (z 52) covered its
last rows even when scrolled to the end. This follows the rule the case
settings popover and case combobox already use.
Reword the rule's comment: it claimed dvh follows the keyboard and that the
vh line is a fallback, and neither is true (a declaration carrying var() is
never dropped at parse time). The new text has no braces and no max-height
text, which the gate test's rule() slicer depends on.
Pin the fixes in the gate test (the --app-height and safe-area terms,
overflow-x: hidden, the toolbar raise) and retitle the cap test so it no
longer names dvh as the mechanism. The test now strips CSS comments before
reading the rule, since the rule's own comment names overflow-y: auto and
touch-action: pan-y and would otherwise keep those assertions green after the
declarations were deleted (checked by deleting them: the test now fails).
Drop .changeset/run-menu-scroll.md: it repeated the false dvh claim, and the
release writes one consolidated changeset at COM.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
App Settings now refuses a Default Codex model that the server would reject,
before anything is written to localStorage. SettingsUpdateSchema is .strict()
and checks codexModel with ^[a-zA-Z0-9._\-/]*$, so a value like gpt-oss:20b
400'd the whole settings PUT while the toast still said "Settings saved", and
because the bad value was already in the local blob every later save from that
device failed the same way. The client check uses the same pattern, shows an
error toast, focuses the field and keeps the modal open. The toast has a zh-CN
translation in i18n.js.
src/web/codex-launch-defaults.ts gets an @fileoverview (fill only unset fields,
re-validate persisted values, callers decide scope, never writes Codex config
files), as every module in src carries one.
Both new Codex rows in index.html carry has-field, like every other App
Settings field row, so on phones the input and the select stack under their
label instead of squeezing it into a narrow column.
The Agent CLIs wiki paragraph said the defaults apply to every local launch.
Scheduled (cron) codex jobs are built without a codexConfig and never get
them, while Resume goes through POST /api/sessions and does, so the sentence
now names the Run menu, Resume, POST /api/sessions and /api/quick-start, and
says cron jobs do not use them.
The Settings Reference lists the two new rows in the Agents & CLIs table. The
neighbouring "Bypass approvals and sandbox" row described Pi's project trust;
it is the Codex --dangerously-bypass-approvals-and-sandbox toggle, so its note
says that now.
The PR's own changeset is removed: the release writes one consolidated
changeset at COM, and the PR's text overstated the scope (it included cron).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pin the deprecation warning on every /api/screenshots route. The PR's test
sends two GET /api/screenshots requests and checks that exactly one warning
comes out, which proves the once-flag but not the individual calls: the
POST and GET /:name warn calls could be deleted and it would stay green. A
new it.each sends one request per route (GET list, a non-multipart POST that
reaches the handler before the content-type check, and GET /:name for a
missing file) on the fresh per-test harness, and each case asserts exactly
one warning naming POST /api/sessions/:id/paste-image. Removing any single
warn call now fails its own case (checked by deleting each call in turn).
The deprecation's CHANGELOG note, required by docs/versioning-policy.md for a
deprecated covered surface, rides the consolidated release changeset rather
than a file here, since the PR added none.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Move test/sse-tile-grid-filter.test.ts to an ephemeral port. The release
added it with #561 on fixed port 3287, after #556 was cut, so it is not on
the guard's LEGACY_FIXED_PORT_FILES and test/test-ports-guard.test.ts failed
on the merged tree. It now builds new WebServer(0, ...) and its url() helper
reads server.boundPort (only ever called inside tests, after beforeAll).
Converting it is preferred over listing it, since the legacy list is
shrink-only.
Update the five docs that still told contributors to pick a unique fixed
port, which the new guard now rejects for any WebServer test: CLAUDE.md
(Adding Features and Testing), AGENTS.md, .github/CONTRIBUTING.md and the
wiki's Contributing page (mirrored to the public GitHub wiki). They now say
to bind port 0 and read boundPort (or address().port for a raw server), and
note that the mobile suite keeps its fixed ports for now, because
test/mobile/helpers/server.ts caches servers by port, so createTestServer(0)
from two callers would share one server.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An opencode tile showed only the session name, while Claude Code, codex and
DeepSeek tiles add `· <model>`: opencode declared no `modelDetect`, so its
screen was never read for a model and it is launched without a model param.
opencode draws the model on its composer's agent row, directly above the
box's bottom edge: `┃ Build Big Pickle OpenCode Zen`. Read from its own
1.3.0 source, the row is the agent, the model's name, the provider's name and
`· <variant>` when the model has one, and only colour tells model from
provider. So the field is all of it, exactly what opencode itself shows (the
owner's choice over a short id that only appears after the first reply).
- The pattern anchors on that row sitting directly above the `╹` edge, ends
the field at a double space (where the 200-column layout's sidebar shares
the row), skips the `No provider selected` placeholder, and takes the LAST
such row in the window through a lookahead, so a composer-shaped row the
agent prints higher up can never stand in for it. A test with a forged pair
inside the window fails without the lookahead.
- It reads 8 rows: the home screen puts up to five rows of opencode's own
chrome under the composer (key hints, a tip, the cwd/version row). The
schema's `screenLines` bound goes from 4 to 8, the reader's own cap; the
comment there records why a taller window is only safe with such a pattern.
- A permission prompt or shell mode hides the row; the last model is kept.
Measured against every captured opencode 1.3.0 frame (home screen and in
session, 40/60/120/200 columns, mid-turn and at rest, permission prompt):
the model was read everywhere it is drawn and nowhere else. Live on an
isolated instance from this branch, a restored opencode session published
`displayModel: Big Pickle OpenCode Zen` (source: screen) and its tile header
rendered `oc-home · Big Pickle OpenCode Zen`. The owner's own home-screen pane
on the 1.36.0 beta reads the same.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A gemini session stayed "working" for good after its first turn. Its braille
spinner trips the generic SPINNER_PATTERN and marks the pane working, but only
a composer glyph arms the idle confirmation and gemini declared none, so it
fell back to Claude's `❯`, which gemini never draws.
Measured on live Gemini CLI 0.63.0 panes (capture-pane every 300 ms through
real turns with a shell call at 40, 120 and 200 columns, YOLO and default
approval mode, plus the raw PTY stream). The turns ran against a local
stand-in for the Gemini API (GOOGLE_GEMINI_BASE_URL, which Codeman's custom
endpoint support already sets), since the CLI's TUI does not depend on the
backend and no account is needed for it:
- The TUI repaints its whole bottom region every frame, composer included,
and the composer sits between a `▄` bar and a `▀` bar; the submitted prompt
is echoed between the same bars. The `▀` bar arms the idle check: every
repaint carries it, tmux's reattach repaint too. The composer's prompt
character is no good: it follows the approval mode (`*` in YOLO), and its
`>` also starts the echoed prompt, which would make the submit verifier
read a submitted prompt as stranded and press Enter again.
- While a turn runs a line `⠦ Thinking... (esc to cancel, 6s)` animates about
every 80 ms (largest gap mid-turn: 214 ms). The label can be any loading
phrase, so the working line is the `(esc to cancel, <n>` suffix, or a
spinner frame opening a line for when a long phrase wraps that suffix.
Nothing at rest matches either.
- A tool confirmation (default mode) replaces the composer, stops the
spinner and the pane goes silent (3.9 s gap), so it reads as idle.
Verified on an isolated instance from this branch: a YOLO turn emitted one
session:working (+170 ms) and one session:idle (2.5 s after the last output);
a default-mode turn went working -> idle while the confirmation waited ->
working once allowed -> idle at the end; after a server restart four restored
gemini panes went busy -> idle in about 4 s; a fresh launch settled in 3 s.
The launch-settle and uncharacterised-CLI tests that used gemini as their
example of a CLI without work detection now use grok and deepseek.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A codex session waiting on a background terminal read as plainly idle,
with no "1 background terminal" badge. Codex pins that row above its
composer, and the registry looked for it in the last three non-blank
rows. Codex 0.162.0 added a hint row under the status line at rest
(` ← for agents · ? for shortcuts`), which pushes the chip to FOURTH from
the bottom exactly when the idle probe reads it. Measured live on the
1.36.0 beta with `sleep 600` started as a background terminal: chip,
composer, status line, hint row; the server reported `watching: null`.
While a prompt is typed the hint goes away and the chip is third again.
The codex entry now declares `watchingLines: 4`. The trade is stated in
the entry: with no terminal running, the fourth row from the bottom is
the last transcript row (the last two while typing), which the agent
writes. As before, this is contained by codex having no hooks: a forged
row costs a wrong badge, never a silenced alert.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A pi session showed no model in its tab or tile unless one was passed at
launch, and none at all for the default route. pi draws its model in its
own footer, but its registry entry declared no modelDetect, so the pane
probe never read it.
The footer, from pi 1.1.0's footer code (0.84.4's is the same) and a live
pane (`0.8%/253k (auto) qwen3.8-27b-pi • xhigh`): usage and context
on the left, then at least two spaces and `[(provider) ]<model>`, with
` • <thinking>` for a reasoning model and ` → <routed model>` when
routed. The last two rows are read (an extension status row can sit
below), and the context field picks the stats row out of them.
pi truncates the right side to fit a narrow pane with no ellipsis,
leaving exactly two spaces of padding. A name with nothing after it is
therefore read only with three or more spaces in front, and with two only
when a following ` •`/` →` proves it whole, so a cut-off name is never
shown. `no-model`, pi's placeholder, is rejected.
The read rides the idle confirmation pi gained with its workDetect entry.
Verified on an isolated instance: a fresh pi session published
qwen3.8-27b-pi (source screen) about 8 s after launch.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every codex tile and tab showed no model, unlike claude and deepseek.
Codex reports its model only in the footer under the composer
(` GPT-6-Luna default · ~/codeman-cases/testcase`), and the registry read
the pane's LAST row for it. Codex 0.162.0 added a hint row under the
footer at rest (` ← for agents · ? for shortcuts`, or ` ? for
shortcuts`), so the last row was always the hint and `displayModel`
stayed null. Measured live on the 1.36.0 beta: the hint is there at rest
and after a turn, and gone while a prompt is being typed (the footer is
the last row again then).
The codex modelDetect window is now two rows, and the footer must be
either the last row or followed by exactly one more two-space-indented
row. Anchoring to the end of the window keeps the guard the one-row rule
had: with the footer hidden, the last two rows are a transcript line and
the `›` composer, so a footer-shaped line the agent printed is not read.
Tests cover both 0.162 hint variants, the typing layout, the hint never
read as a model, and a forged footer-plus-indented pair above the
composer.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An omp session stayed "working" for good once a turn started, the same
latch pi had: omp's braille spinner trips the SPINNER_PATTERN fast path,
and only a composer glyph arms the idle confirmation. omp declared none,
so it fell back to Claude's `❯`, which omp never draws once its setup
wizard is done.
Measured on live omp 18.8.6 and 18.0.11 panes, holding a turn open
against a local endpoint that never answers: the input row is `╰─ <text>`
and is redrawn at submit, at the end of a turn, at launch and on
reattach. While a turn runs, the status bar's leading `π` becomes a
braille spinner plus the elapsed time (` ⠼ 14s > ⬢ model > ...`; 18.0.11
pads it with two spaces, past a minute it reads `1m`), with a
`⎋ Working…` row above it. The registry entry now names the input row as
the glyph and either working signal as the working line.
The glyph also switches the submit verifier on for omp, which reads the
input row the way it reads Claude's composer. A prompt sent mid-turn goes
to omp's Steering queue and clears the row, so the verifier stands down.
Text left in the row after an Enter is the one case it re-presses.
Verified end to end on a sandboxed instance (own HOME and PATH, omp
18.8.6): session:idle at launch, session:working during a turn,
session:idle about 3 s after it ended, and a restored pane settled idle
about 3 s after a server restart.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An opencode session that ran a tool stayed "working" for good. The running
tool row draws a braille spinner (`⠋ Sleep for 12 seconds...`), which trips
the generic SPINNER_PATTERN and marks the pane working, but only a composer
glyph arms the idle confirmation and opencode declared none, so it fell back
to Claude's `❯`, which opencode never draws. Measured on an isolated
instance: a 16 s turn latched busy/isWorking for the rest of the session.
A text-only turn had the opposite problem and never showed as working.
Measured on live opencode 1.3.0 panes (capture-pane every 250-300 ms through
real turns at 40, 60, 120 and 200 columns, plus the raw PTY stream):
- Every composer row starts with a `┃` bar, and the submitted prompt lands in
the transcript with the same bar, so a turn's first repaint arms the idle
check, and tmux's reattach repaint does the same for a restored pane.
- While a turn runs the footer row starts with an 8-cell knight-rider
spinner, `⬝■■■■■■⬝ esc interrupt`, redrawn about every 40 ms (largest
gap mid-turn: 121 ms). At rest the TUI is silent and nothing on screen
draws a `⬝`/`■` run, the 200-column sidebar included.
- The working line is the spinner run, `[⬝■]{8}`, not the label: tmux ships
`esc` and `interrupt` as separate words joined by cursor moves, so the
label never reaches the stream detector, and at 40 columns the footer
wraps it to `esc` / `interr` / `upt`. All 344 spinner chunks of a turn
match the run after Codeman's ANSI strip.
- A pending permission prompt replaces the composer and stops the spinner,
so it reads as idle (waiting on the user).
- The last `┃` row on screen is the composer's agent/model row, or the
permission box's closing bar, never the prompt text, so the submit
verifier stands down and can never press Enter into a dialog.
Verified on an isolated instance from this branch: a 15 s tool turn emitted
exactly one session:working (+271 ms) and one session:idle (3 s after the
spinner stopped); a permission prompt read idle and the allowed turn went
working -> idle; after a server restart both restored opencode panes (one
at rest, one on a permission prompt) went busy -> idle in about 4 s; a
fresh launch reached an open page as idle in 3 s.
The launch-settle tests that used opencode as their example of a CLI
without work detection now use gemini and antigravity.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
40560ace made the 3 s launch settle announce its idle. For a CLI with
work detection that edge could now land in the middle of a turn: a
prompt sent ~2.9 s after launch has not been marked working yet (the
working line goes through the deferred parsers), so the settle called
the pane idle and a send-and-wait registered for that prompt resolved
before the turn even started. Before 40560ace the settle was silent, so
this edge is new.
The settle now also leaves a pane to `_confirmIdle()` when a prompt was
submitted since the timer was armed, the same way it already does for a
pane marked working; that confirmation reads the screen before it ends
the turn. A CLI without work detection still settles: nothing else ever
would.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A pi session stayed "working" for good once a turn started. pi's braille
spinner trips the SPINNER_PATTERN fast path, which marks the pane working,
but only a composer glyph arms the idle confirmation and pi declared none,
so it fell back to Claude's `❯`, which pi never draws. Measured on beta136:
an errored turn stayed busy/isWorking for 3+ minutes after pi was back at
rest.
pi has no composer glyph. Measured on a live pi 1.1.0 pane (capture-pane
every 250 ms through a turn): its composer sits between two `─` rules, and
while a turn runs it rewrites the top rule as `── ⠏ Working ───` on every
frame. The registry entry now names the rule as the glyph that arms the
check and a spinner frame inside it as the working line.
The same glyph settles a reattached pi pane (a restored pane gets no launch
timer): tmux's reattach repaint carries `─`, which arms the confirmation.
The submit verifier reads the last rule, finds no prompt text and stands
down, so it can never press Enter on a pi pane.
Verified on an isolated instance from this branch: a real pi turn emitted
session:working then session:idle, and after a server restart the restored
pi pane went busy -> idle in about 3 s.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A fresh codex, pi or opencode tile could spin "working" forever while the
pane sat at its composer. The external-CLI launch timer set the status to
idle WITHOUT an event, only `needsRefresh`. When the launch paint never
marked the pane working, the later idle confirmation found the status
already idle and announced nothing either, so every open browser kept the
`busy` from the spawn broadcast. A reload fixed it, which is why only open
pages were stuck. Measured on the 1.36.0 beta: w8 (codex) had
`lastPromptTime: 0`, i.e. no idle edge ever, and a page held a fresh pi
session at `busy` while GET /api/sessions said `idle`. pi and opencode hit
it on every launch (no work detection, so the timer is all they have);
codex only when its launch paint lost the race against the timer.
- `_concludeIdle()` is now the one place a pane is concluded idle: status,
working flag and prompt stamp change together and `idle` is emitted
(session:idle + state broadcast). `_confirmIdle()` uses it too.
- The launch settle (`_settlePaneStartup`) concludes a pane still in its
spawn-time `busy`. A pane already working is left to `_confirmIdle()`
only when its CLI declares `capabilities.workDetect`; for the rest the
timer is the only thing that can settle it, so it also clears a working
flag a launch spinner glyph latched.
- `_armPaneSettle()` also arms it for a RESTORED pane (Codeman restart,
auto-reattach, tile Attach) of a CLI without work detection (at this
commit shell, opencode, gemini, antigravity, pi, grok, deepseek and omp),
which used to stay `busy` server-side with nothing to clear it. Restored claude and
codex panes stay on their composer glyph, so a restart mid-turn is never
called idle. No `needsRefresh` there: an attach refetches by itself.
Pre-existing since the OpenCode integration, not a regression of this
release. Restore-path gap reported by the opencode and pi sessions working
the same symptom.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A claude tab drew no harness mark at all and every other agent CLI a
two-letter text pill (DS, CX, ...), so a claude tab read as "no harness"
next to its neighbours. Session tabs (header strip, side rail, sidebar,
phone chips) and the desktop home rail now draw the agent through PR
#532's run-mode-dot <id> slot, the id as data, the same mark the Run
menus and the tile and split headers use. The shell is not an agent and
keeps its SH pill; a CLI added through clis.json gets the slot's plain
dot instead of nothing.
The per-CLI tab pill colours and their light-skin ink overrides are gone
(the monochrome marks follow the tab's own text colour), and the logo
steps aside with the other adornments while a compact rail row renames.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>