The smart-copy gate only entered its selection-check block behind
hasSelection(), so a selection-less Ctrl+Shift+C skipped straight to
`return true` and ceded the keystroke to the browser's own handling
(e.g. Chrome's Inspect-Element binding) instead of matching Pane A's
"never falls through" contract for that chord.
Verified live in a real browser that this is a UX-parity fix, not an
interrupt-safety one: xterm's evaluateKeyboardEvent never emits PTY
data for a shifted ctrl-letter regardless of any gate (only "_" and
"@" get special-cased), so no accidental 0x03 was ever at risk. The
regression test added here asserts on the dispatched event's
defaultPrevented rather than the absence of a WS frame, since the
frame-count check passes vacuously for this exact key combo whether
or not the gate fires.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- buildSplitPickerSessions() now excludes any session with pid === null
(exited CLI, tripped PTY-exit breaker, a restore that never re-attached).
Pane B has no equivalent of selectSession()'s auto re-attach POST, so a
split opened onto one had nothing reading its tmux pane: no terminal
events ever arrived and Session.write() silently dropped every keystroke
with no ack either way, while the socket itself reported healthy.
- Fixed the hollow chord regression test: the synthetic keydowns carried no
keyCode, which is what xterm's evaluateKeyboardEvent switches on to
produce a data frame at all, so the assertion held regardless of whether
the gate fired. Adding real keyCodes surfaced a second, real bug in the
Alt+B case: the event bubbles to app.js's own document-level shortcut
dispatcher, which really toggles the sidebar and resets the layout
attribute the gate reads before Pane B's own (later, non-capture) handler
ever sees it — fixed by driving the app's real settings cache instead of
only the DOM attribute.
- Ported the two remaining primary-pane gates with real consequences:
Ctrl+Z (SIGTSTP) is swallowed for every non-shell session, matching
terminal-ui.js's reasoning (an Ink/TUI agent loop stops dead with no
visible output otherwise), and Shift/Ctrl+Enter now POSTs to
/api/sessions/:id/send-key for THIS pane's own session instead of
letting xterm send a bare \r, which used to submit an incomplete prompt
instead of inserting a newline. Smart-copy Ctrl+C is re-implemented
against Pane B's own terminal (copying app.copyTerminalSelection() would
have copied Pane A's selection instead).
- Updated docs/architecture-invariants.md and docs/split-pane-sessions-plan.md
to match, and added CLAUDE.md's missing .split-picker-menu z-index entry.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Pane B had no attachCustomKeyEventHandler of its own, so the document
capture-phase shortcut handler's preventDefault() (which does not stop
xterm) left Ctrl+K/Alt+1/Alt+B ALSO writing their raw byte/escape
sequence into Pane B's live PTY on top of whatever the app action did
to Pane A. Pane B now installs the same registry-aware gates the
primary pane's own attachCustomKeyEventHandler uses. Ctrl+V is left on
xterm's default paste — no image-paste trap to route it to.
Plus the rest of the review's smaller items:
- Narrowing the window past the desktop gate now closes an open split
instead of leaving it stranded on screen.
- Split is refused while a web tab is active (activeWebviewId), which
used to open Pane B's socket behind a hidden container.
- Pane B now handles the server's `{t:'r'}` refresh frame via a shared
_loadBuffer() helper (also used by connect()), instead of ignoring it.
- The divider drag now uses pointer events + setPointerCapture (mirrors
tab-rail-resize.js), a button!==0 guard, preventDefault, and a
body.split-pane-resizing cursor/selection lock — a plain mousedown
drag selected the text under the cursor as it crossed both terminals.
- Pane B's close control and the picker rows are real <button>s now
(keyboard-reachable), with matching CSS chrome resets.
- Dropped the redundant CodemanBase.base prefix on the buffer fetch
(the global fetch wrapper already applies it).
- data-preview-order for the Split settings chip moved from a collision
with Ultracode Agents (both 15/12) to 11.5, matching its real
position between Multi-monitor and Ultracode Agents in the header;
widened test/app-settings-structure.test.ts's regex to allow the
decimal (Number() already parses it fine for the preview sort).
- Added zh-CN i18n entries for the Split button and empty-picker text.
- Dropped the stray unused `vi` import Ark0N flagged as unrelated to
this feature (vitest's `globals: true` makes it ambient anyway).
- Documented the fix and the deliberate no-cid/seq choice in the
split-pane-sessions architecture-invariants entry.
Added a real-Chromium regression test asserting Ctrl+K/Alt+1/Alt+B
dispatched at Pane B's own textarea send no `{t:'i'}` frame over its
WebSocket. Full CI gate green (409 files, 7736 tests) plus all 8
split-pane browser tests.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
CLAUDE.md previously only mentioned split-pane in the load-order list,
with nothing in the Architecture/frontend prose the way every other
feature gets, and its own module count was one stale (34, should have
been bumped to 35 when terminal-split.js was added). Add a short
pointer-style paragraph next to the other terminal features, fix the
count.
docs/wiki/The-Dashboard.md's header button table and
docs/wiki/Settings-Reference.md's header chips list are the two
user-facing surfaces that never mention Split at all; added both, plus
a note that the feature is desktop-only regardless of the setting.
docs/split-pane-sessions-plan.md: recorded the one design note that
isn't a code change — the global capture-phase shortcut handler always
resolves against Pane A, so Ctrl+L/Ctrl+W typed into Pane B affects the
other session. Not fixed for v1, same reasoning as the rest of the
"deliberately plainer" section.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two majors:
- The divider drag was unthrottled: every mousemove did a full xterm
reflow on BOTH panes and sent Pane B a {t:'z'} resize frame with no
unchanged-dimensions skip, fanning out into a `tmux resize-window`
child plus a SIGWINCH per event — ~50 of each dragging across half a
wide viewport. SplitTerminalPane.fit() is now split into localFit()
(reflow only) and fit() (reflow + send); the drag coalesces moves
into one localFit() per animation frame via requestAnimationFrame,
and sends the real resize for both panes exactly once, at drag end,
matching the primary pane's own throttledResize convention.
- Pane B pulled the FULL scrollback unchunked for every session mode,
writing it in one terminal.write() call. Mirrors the primary pane's
own mode check (app.js's selectSession): shell sessions get a
bounded 1MiB ?tail= fetch instead of ?full=1, and the fetched buffer
is written through a minimal chunked writer (32KB slices, yielding a
frame between each) instead of one primary-pane chunkedTerminalWrite
this simpler, independently created/destroyed pane has no equivalent
of (no session-switch generation counters or live-output gate).
Smaller items from the same review:
- Pane B now follows live appearance changes (applyTerminalSkin,
applyTerminalFontFamily, applyTerminalFontWeights, setFontSize all
propagate to it, matching the teammateTerminals pattern) and reads
the real codeman-font-size/terminalFontFamily/weights/DEFAULT_SCROLLBACK
settings at construction instead of hardcoding fontSize 14 / scrollback 5000.
- The Pane-B-promotion path now skips selectSession() when
_closingSessions already owns this delete (the user closing Pane A's
own tab), matching _onSessionDeleted's own active-session-handoff guard.
- Detaching a session AFTER a split is already open now yields the PTY
size in _sendResize() too (not just at picker-open time), mirroring
sendResize's own detachedElsewhere guard.
- .btn-split joins the body.solo-mode hide list, next to .btn-multimonitor.
- The split row was 6px wider than its container (two flex-shrink:0
50% panes plus a 6px divider): both panes are now flex-shrink 1.
- Pane B's header and the split-picker rows are marked so i18n.js's
exact-string lookup skips them, matching .session-name elsewhere —
a session literally named e.g. "Sessions" was translatable on zh-CN.
- The Split button now reflects open/closed state via a `.split-open`
accent style, aria-pressed, and a title/aria-label that says which
behaviour the next click gets.
- _splitPane.connect() is no longer an unawaited call with no .catch().
- terminal-split.js's fileoverview pointed at a doc path that was
renamed away in the previous push; @dependency now credits
constants.js for CodemanTerminalFont, not terminal-ui.js.
- index.html's Split settings chip no longer reuses data-preview-order
"12" (already the Ultracode Agents chip's slot in the same "header"
preview group).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Blocker from Ark0N's second PR #453 pass: moving showSplitButton into
settings-ui.js's per-device displayKeys set was only half of making it
per-device. saveAppSettings() still put it in the object PUT to
/api/settings, SettingsUpdateSchema (.strict()) does not declare it,
the server answered 400 INVALID_INPUT, and because the call site never
checked res.ok the UI still reported "Settings saved" while NOTHING
persisted — workspaceHooksEnabled, agentSkillEnabled, tunnelEnabled,
claudeModel, every toggle, on every save, on every device. Strip it
out via the same destructure every other per-device key goes through
(`showSplitButton: _ssp,`), drop the stray mention from a schemas.ts
comment (a mention there reads as "this is a real field" to the next
grep), and add a static guard test mirroring
test/terminal-auto-copy.test.ts's three-way rule.
Also finishes the desktop gate the first pass only did in CSS at
599px: SPLIT_PANE_MIN_WIDTH (1180, matching HOME_SESSIONS_MIN_WIDTH)
now backs an actual JS width check in _applySplitButtonVisibility,
with a matchMedia listener so a live window resize hides/shows the
button without a reload — the CSS backstop in styles.css is the
reverse-direction guarantee for when JS hasn't run.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ark0N's PR #453 review: nothing gated this feature to desktop even
though the design called for it (two 240px min-width panes plus the
divider need ~486px, and the divider has no touch handlers), and
showSplitButton was a SYNCED setting, so turning it on at a desk also
put the button in the phone header.
- Hard-hide .btn-split on phones in mobile.css regardless of the
setting, matching the other desktop-oriented header buttons in the
same @media (max-width: 599px) block.
- Move showSplitButton into settings-ui.js's per-device displayKeys
set and drop it from SettingsUpdateSchema entirely, matching the
showFileViewerButton/skin precedent (CLAUDE.md's "per-device keys
... must NOT be added to SettingsUpdateSchema" rule) — a desktop
opt-in must never sync onto a phone that never asked for it. Removes
the now-invalid server-round-trip test for the setting.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Exclude popped-out (detached) sessions from the split picker:
SplitTerminalPane._sendResize() has no yield-to-detached-window check
the way the primary pane's sendResize() does, so splitting against a
detached session put its own window and Pane B in a fight over the
same PTY's dimensions. Simplest fix per the review: keep them out of
buildSplitPickerSessions() entirely.
- Show a visible dead state when Pane B's WebSocket drops. onData
already silently discards keystrokes while the socket isn't OPEN
(there is no reconnect for v1), so a dropped socket left the pane
looking normal while it quietly ate everything typed into it.
- openSplitPane() returns early with no active session, so a split
triggered from the home screen no longer creates and connects Pane B
behind the opaque welcome overlay with nothing to show for it.
- onMove() during a divider drag now bails when the split has
auto-collapsed mid-drag (the other pane's session ending) instead of
throwing on `divider.parentElement` being null.
- Promote Pane B via `selectSession(id, { auto: true })` when Pane A's
session ends — this is an app-driven selection, not the user clicking
a tab, so it must not spend the promoted session's idle alert.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ark0N's PR #453 review: fit() was only ever called from the divider
drag, and the trailing-edge ResizeObserver callback in terminal-ui.js
(throttledResize) only ever measured Pane A's own container. Split at
a wide viewport, shrink the window (or toggle the Alt+B sidebar, or
drag the tab rail), and Pane A's cols changed while Pane B silently
kept its stale PTY size in both xterm and the real pane.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Per Ark0N's review on PR #453: rename the design spec to
docs/split-pane-sessions-plan.md, matching every other feature's
*-plan.md convention, and drop the 957-line implementation task plan
(docs/superpowers/plans/2026-09-15-split-pane-sessions.md) — workflow
scaffolding for the subagent-driven-development run, not repo
documentation. Fixes the now-dangling link in architecture-invariants.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Real-browser regression coverage for the previous commit:
- SplitTerminalPane connects onto an already-quiet session and shows its
existing scrollback with no new output, proving the ?full=1 fetch (not
a live echo) populated the pane.
- openSplitPane() force-resizes Pane A synchronously as part of opening
a split.
- Dragging the divider force-resizes Pane A once, at drag end.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Pane B's SplitTerminalPane.connect() only opened a WebSocket and waited for
live output — ws-routes.ts's terminal socket sends nothing on connect, only
future 'terminal' events — so it stayed blank until the target session
happened to produce new output. It looked intermittent rather than
always-broken because a resize sent by _sendResize() often nudges the
session's real tmux window to a new size, and tmux repaints its current
screen on resize; that incidental repaint was what usually populated the
pane. When Pane B's computed dimensions already matched the session's
last-known size, Session.resize() skipped the resize as a no-op and the
pane stayed empty. Fetch the existing scrollback (?full=1) before opening
the socket, same as the primary pane does.
Pane A never told its own session's PTY/tmux about a size change at all,
relying purely on the passive 300ms-debounced ResizeObserver in
terminal-ui.js. openSplitPane() now force-resizes Pane A immediately on
entering split (mirroring closeSplitPane()'s existing symmetric call), and
the divider-drag handler force-resizes it once at drag end (matching the
codebase's established trailing-edge debounce convention rather than
flooding a resize per mousemove).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
test/split-pane-auto-collapse.browser.test.ts covers "Pane B's session ends"
in a real Chromium, but that suite is excluded from the npm test CI gate.
The "Pane A's session ends, Pane B gets promoted" branch had no coverage
anywhere, and it is the one branch whose correctness depends on exact
ordering: _splitSessionId must be captured BEFORE closeSplitPane() runs
(which nulls it) or the promoted session id is lost. Loads terminal-split.js
via `vm` against a minimal fake CodemanApp (same technique as
test/session-close-fallback.test.ts), and pins all three branches (Pane A
ends, Pane B ends, unrelated session ends) plus that the original
_onSessionDeleted always still fires.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Nothing stopped a stale picker click (opened before switching tabs) or
clicking Pane B's own session tab while split from landing on
openSplitPane(sessionId) with sessionId === activeSessionId, or from
selectSession() rebinding the primary pane onto the session Pane B was
already showing — either way, two live WebSockets to one session, each
independently claiming PTY dimensions via its own {t:'z',...} resize frame.
openSplitPane() now refuses early when the target is already the active
session, and a new selectSession() prototype patch (same top-level pattern
as the existing _onSessionDeleted patch) closes an active split BEFORE the
primary pane rebinds to the session Pane B holds.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The new "Split-pane sessions" section was inserted between the "Session
list layout (header strip vs. left sidebar)" heading and that section's own
body paragraphs, orphaning the heading from its content. Move "Split-pane
sessions" to after the Session list layout section's full body, before
"Gesture control: the setting" — no change to the Session list layout prose
itself, only where the new section sits relative to it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
.main.webview-active hid .terminal-wrap when a web tab became active, but
.terminal-wrap is reparented INSIDE .terminal-split-container while a split
is open, so Pane B and the divider stayed stranded on screen over the
dashboard iframe. Hide the whole split container as one unit, mirroring the
existing .terminal-wrap rule; no state is destroyed, so returning to the
session tab shows the split intact.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
.split-picker-menu/-item/-empty (created in openSplitPicker()) had zero CSS
and could not be dismissed except by picking an item — a default-path defect
since the Split button ships enabled to anyone who flips showSplitButton on.
Add CSS matching the sibling .run-mode-menu popover's look (floating-bg
backdrop blur, border, shadow, z-index 1000 above the header's 100), and
dismiss on outside click or Escape via the same one-shot listener pattern
session-ui.js already uses for its other transient popovers
(toggleCaseSettings(), toggleRunModeMenu()). Picking an item now routes
through the same _dismissSplitPicker() method as the outside-click/Escape
handlers, so the listeners never outlive the menu.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
_sendResize() clamped Pane B's proposed cols/rows to a 40/10 floor before
sending the {t:'z',...} resize frame, so the PTY was misinformed of Pane B's
real width at the divider's own reachable 20% position, causing real
output-wrapping bugs. The primary pane (terminal-ui.js's
getTerminalDimensions()) sends fitAddon.proposeDimensions() unclamped and
lets the server enforce its own valid range ([1,500]/[1,200] in
ws-routes.ts); Pane B now matches that convention.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
.btn-split--hidden had no matching CSS rule anywhere, so the opt-in Split
header button shipped visible to every user on every viewport regardless of
the setting. Add the `display: none !important` rule alongside its sibling
marker classes (.btn-multimonitor--hidden etc.), plus a static regression
guard (test/split-pane-hidden-button-css.test.ts) that fails if any future
"*--hidden" marker class in index.html is missing a matching CSS rule.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Task 4's implementer found two real bugs in this plan's browser-test
helpers: POST /api/sessions nests the id at data.session.id (not
data.id), and mode:'shell' needs a follow-up POST .../shell to actually
spawn a PTY. Fixed in Task 4's own snippet (documentation accuracy —
already fixed in the real committed code) and pre-emptively in Tasks
5/6's createShellSession() helper before either was dispatched, so
neither implementer has to rediscover it independently. Also corrected
the <script> tag snippet to defer, matching the real file's convention.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Monkey-patching the instance's _onSessionDeleted inside a
DOMContentLoaded listener races connectSSE()'s handler-wrapper cache,
which captures the function reference by value on first connect and
never re-reads it. Patching CodemanApp.prototype at module-evaluation
time (synchronous script-tag order) is unraceable: it completes before
any instance exists or connectSSE() ever runs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Preflight scan for SDD execution caught two classes of defect before
dispatch: Task 2's test invented a buildTestApp() helper and response
envelope that don't exist for /api/settings; Tasks 4-6 used
@playwright/test's runner against a test/browser/ directory that
doesn't exist in this codebase. Both corrected against real patterns
found in existing tests (system-routes-settings-partial-put.test.ts,
terminal-copy-shortcut.test.ts, tab-rail-resize.browser.test.ts).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Global Constraints previously read as if local-echo/CJK/accessory-bar
were desktop features; they are mobile-only, and split-pane is the
desktop-only side of that equation. Also names the exact spec section
instead of a loose paraphrase, and adds a worktree/branch note so an
executing subagent knows where this plan runs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
7 tasks: pure divider/picker helpers, showSplitButton header wiring,
split-container CSS, SplitTerminalPane (Pane B's independent xterm+WS),
open/close orchestration with picker and divider drag, auto-collapse on
either session ending, and an architecture-invariants entry.
Also folds in the "detach session" prior art discovered mid-brainstorm
into the spec's architecture section.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The Problem paragraph and the architecture section used "tab" where
"pane" was meant, colliding with the browser's own tab concept.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Scopes v1 of an in-app split view (two live session panes side-by-side,
draggable divider) after multi-monitor spanning turned out to solve a
different problem than showing multiple panes at once.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Three of Ark0N's four "will take at merge" items, applied instead since
they were straightforward to do properly:
1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on
<file>::<expression> (fixed last round) closed the line-shift problem
but opened a new one: every stock id was already allowlisted for
session-ui.js in the `mode === '<id>'` form, so a BRAND NEW branch
reusing that exact expression anywhere in the file passed unnoticed.
Reproduced live (`if (this.mode === 'codex')` injected into
runOpenCode()) — stayed green under the old version. Each allowlist
entry now carries the exact count of approved call sites, and a new
test asserts actual-vs-declared count for every key; a mismatch in
either direction is real (higher = new unreviewed branch riding in on
an existing approval, lower = a reviewed site was removed and the
entry is now stale). Reproduced again against the fix: same injection
now fails with an exact diagnostic (expected 2, found 3).
2. Added test/run-mode-launch-table-drift.test.ts. RUN_MODE_LAUNCH
restates four things stock.ts already owns (label, install command,
supportsCustomModel, the external-mode key set), and they agree today
with nothing enforcing it. supportsCustomModel is the dangerous one:
the Run-menu picker's rows come from the server-injected
window.__codemanCustomModelClis (built from
capabilities.customModelInjection.kind), so a CLI gaining a real
injection recipe later would be OFFERED in the picker while
_runCliMode silently drops the customModel field for it — the session
launches on the vendor's cloud while the UI claims the local endpoint.
Drives the real session-ui.js via JSDOM and compares RUN_MODE_LAUNCH
against STOCK_CLIS on all four axes.
3. Inlined the "Open Question 7 in PR-B2.md" references in the allowlist
reasons — PR-B2.md is a local planning doc, never part of the
committed tree, so the reference was dead on arrival for anyone
reading the repo. Points at the PR #458 review thread instead.
4. Added a sentence to docs/cli-registry.md naming the new frontend guard
alongside the backend one it mirrors.
Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format/
check:frontend-syntax all clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
Four things from the maintainer's review on PR #459, all fixed:
1. Bound _getCustomModelCurrentlyLoaded's probe client-side (~800ms via
Promise.race, on top of — never instead of — the route's own 5s
server-side timeout). Without it, an asleep/firewalled endpoint behind
a saved model list left the picker completely invisible for up to 5s
after the Run menu had already closed, with no spinner or toast.
`timeoutMs` is an optional param (default 800, real callers never pass
it) so a test can drive it in milliseconds, same pattern as
`_watchLlamaSwapLoading`'s own `pollIntervalMs` — this code runs in a
JSDOM window's own realm, whose setTimeout vi.useFakeTimers() cannot
patch.
2. "Last used" is now written only once a launch actually applies, never
on the mere click. It moved out of runCustomModelEntry (unconditional)
and into each path's own success point: _quickStartWithCustomModelConfirm
after the final post succeeds, and _runCustomModelEntryViaRestart right
after the apply's success check. A context-window-warning decline means
this exact model cannot work with this CLI at all, so the old
unconditional write would promote, next time the picker opened, the one
model guaranteed to fail again.
3. Documented the promotion/tag precedence and the new
codeman:customModelLastUsed:<mode>:<endpointId> localStorage key in both
CLAUDE.md's Custom Model Endpoint Profiles section and
docs/custom-model-endpoints.md's Run-menu picker section.
4. Added zh-CN entries for "Currently loaded" and "Last used" in i18n.js,
next to this modal's existing "Choose a model"/"Custom Endpoints" pair.
New tests: the client-side timeout (endpoint that never answers, one that
answers within the bound, and a rejected-after-timeout probe settling
quietly), and "last used" recording on success vs. NOT recording on either
confirmation's decline, for both the restart and one-shot paths.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Code review (high effort) on the previous commit found a real race: making
_openCustomModelPickModal async (it now awaits the currently-loaded-model
probe before rendering) meant a second, faster call for a different
endpoint could render first, only for the first call's slower probe to
resolve afterwards and overwrite the modal with the wrong endpoint's model
list — while _pendingCustomModelPick (set synchronously, before either
await) still named the second, correct endpoint. Picking a model in that
state would launch/apply the wrong model on the wrong endpoint.
Fixed with the same mutable-generation-counter guard
_watchLlamaSwapLoading already uses for an identical async-superseded-by-
newer-call shape: every DOM write, including _pendingCustomModelPick
itself, is deferred until after the awaited probe, and a call that finds
its generation already superseded bails out untouched instead of clobbering
whatever a newer call already rendered.
Added a regression test driving two overlapping opens with a controlled
promise so the earlier, slower probe resolves after the later, faster one
renders, asserting the late response is a no-op.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Custom Model Endpoint Profiles' "which model" picker (session-ui.js's
_openCustomModelPickModal) always listed models in their raw discovery
order, so on a host with several downloaded GGUFs the user had to
remember (or eyeball the "Default" tag) which one llama-swap actually
had hot before picking — the whole point of the picker being fast is
undone if it makes you think first.
The picker now promotes exactly one model to the top of the list:
- If llama-swap reports a model from this host's own list `ready`
right now (via the existing GET /api/model-endpoints/:id/running-status
route), that model is promoted and tagged "Currently loaded" — it's
what a launch attaches to with zero wait.
- Otherwise, the last model actually launched on this exact
(harness, endpoint) pair is promoted and tagged "Last used", read
from a new per-device localStorage key
(codeman:customModelLastUsed:<mode>:<endpointId>), written by
runCustomModelEntry on every launch attempt regardless of outcome.
- A plain (non-llama-swap) OpenAI-compatible server, an unreachable
endpoint, or a loaded-but-not-yet-ready model never promotes
anything — the rest of the list keeps its discovery order.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Two required fixes from Ark0N's review of #458:
1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on
<file>::<line>::<expression>. A single inserted line anywhere above an
entry shifted every subsequent line number, so all 21 entries went stale
simultaneously and the same 21 branches were reported as "new" — on a
file six other open PRs also touch. Dropped the line number from the key
(<file>::<expression>, matching the backend guard's own design), which
collapses 21 line-keyed entries to 11 or-collapse where the same
expression recurs at multiple call sites in the same file.
2. test/run-mode-ui.test.ts's terminal-ownership guard scanned method
bodies via `^ {2}async (run[A-Za-z]*)\(\) \{$`, which matched the 8
one-line run<Mode>() wrappers PR B2 introduced but not _runCliMode(mode),
where the real logic (and the actual risk the guard exists to catch) now
lives. Fixed the regex to `^ {2}async (_?run[A-Za-z]*)\(\w*\) \{$` and
added _runCliMode to the sanity list. Same-class fix in
test/opencode-resize.test.ts, which had the identical blind spot via
runOpenCode.toString().
Both reproduced live before fixing (inserted the same comment line; added
this.terminal.clear() to _runCliMode) to confirm the bug, then confirmed
the fix catches it and the suite stays green otherwise.
Also resolves Open Question 2 by dropping window.__codemanCliCatalog
entirely: nothing consumed it, and a registry DECLARED_FOR_LATER field
costs nothing until read while an unconsumed script tag on every page
render is a different trade. Reverts Phase 1 cleanly — server.ts's
injection, shortBadge back in types.ts's DECLARED_FOR_LATER list and the
pinned guard test, and the three associated render-index-html.test.ts /
server-index-title.test.ts assertions.
Full gate: 405 files / 7717 tests / 0 failures (net unchanged), typecheck/
lint/format clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
PR #380 (PR B) held back the frontend half of the CLI registry refactor,
explicitly deferring window.__codemanCliCatalog and making session-ui.js /
mobile-overview.js catalogue-driven as "PR B2".
- Inject window.__codemanCliCatalog in renderIndexHtml(), following the
existing __codemanCustomModelClis pattern (escapeScriptJson-guarded,
resolved per-request). Reading CliEntry.shortBadge here is what makes it
genuinely read, so it drops out of types.ts's DECLARED_FOR_LATER list.
- Consolidate session-ui.js's 8 near-duplicate run<Mode>() launch functions
(opencode/codex/gemini/antigravity/pi/omp/grok/deepseek) into one shared
_runCliMode() plus a local RUN_MODE_LAUNCH config table. The 8 method
names stay as thin wrappers (index.html calls them by name; tests assert
on the name). Also collapses a duplicated 8-way isAltMode/isExternalCli
OR-chain (same expression, copy-pasted twice in openSessionOptions) into
one EXTERNAL_CLI_MODES check.
- Add test/frontend-cli-no-id-branching.test.ts, a guard scoped to
session-ui.js/mobile-overview.js only (not the rest of src/web/public/,
which stays explicitly out of scope per CLAUDE.md), mirroring the
backend's own no-id-branching guard.
mobile-overview.js and the wiring of accent/echo/wheelForward/
keyboardAccessory were investigated and deliberately left alone: the first
is already a single, tested, gated table (not duplicated logic); the second
set belongs to terminal-ui.js/keyboard-accessory.js/styles.css, files
outside this PR's mandate.
Verified on a tmux-capable devbox (this sandbox has no tmux): full CI gate
at 405 files / 7717 tests / 0 failures, typecheck clean, 94 targeted tests
covering exact per-CLI wire-body shapes unmodified and passing, and a live
anti-vacuity check on the new guard (injected a real branch, confirmed it
fails, reverted, confirmed green).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
A full review of the release tree found seven things, and four of them were mine.
**The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext`
and `confirmedSwap` changed the wire field without moving three assertions that check
it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts`
(the swap modal and the context modal, each of which already receives exactly the right
per-question flag). Moved, with the titles.
**Worse, my own tests for the split never ran.** The four cases in
`session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`,
which was declared inside a sibling `describe`, so they threw a ReferenceError during
setup. The split would have shipped with no passing server-side coverage while the gate
reported the failure as four broken tests rather than as four tests that were never
written. `mockRunning` is hoisted to the outer describe.
**The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved
its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so
the other eight modes fell back to claude's `❯`. That is also starship's default shell
prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line
`❯ npm run build` sits on screen for as long as the command runs, the verifier reads it
as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to
nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N
prompt, `read -p`, an installer or a pager, where it takes the default. The module's own
fileoverview already stated the rule this broke. Now `?? ''`, which
`promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI
that actually declares a composer.
**My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing
three numbered rules.
The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared
in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told
an integrator to retry with `confirmed: true` for both questions, which is precisely the
thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md`
and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR`
multi-user consequence, both user-visible and both previously absent, and #454's gained
the one exception to its own claim: a Custom Endpoints launch ignores the Instance count
stepper and always starts one session.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CLAUDE.md gained the admin-only note when the key joined claude's privilegedEnvKeys;
architecture-invariants, which is where the exact-key allowlist rule is documented in
depth, still described the pre-change world. The reboot-restore half is the one worth
writing down: a non-granted owner's already-persisted CLAUDE_CONFIG_DIR is stripped on
restore, which moves that session back to the default Claude account with no error.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
More from the review of 5fc391a4, all documentation rather than behaviour.
The changeset was 1602 words of development log, written as the PR grew, with bullets
and loose paragraphs interleaved. That text becomes CHANGELOG.md and the GitHub release
body verbatim, so it is now one user-facing account of what the feature does and what
the real-server work bought, at roughly a fifth the length.
docs/api-reference.md promised a `cmd` field on running-status that the route
deliberately strips (it carries model paths and can carry --api-key).
Two places claimed the apply routes validate `modelId` against the endpoint's
discovered models. Neither does. Dropped the claim rather than adding the check:
discovery can be up to five minutes stale, so a 400 there would refuse a launch that
actually works, and a typo'd id already fails on the CLI's own first request. CLAUDE.md
now says so explicitly, since the absence is the surprising part.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two findings from the review of 5fc391a4, both fixed here rather than sent back.
**The API-key trust seed never matched a real key.** `seedApiKeyTrustFile()` wrote the
key verbatim into `customApiKeyResponses.approved`, but Claude Code stores and compares
only the last 20 characters (`key.trim().slice(-20)`, applied on both the write and the
lookup). For any real key the seed missed, so claude stopped at the interactive
"Detected a custom API key in your environment" prompt, whose default is
"No (recommended)": the launch hangs, or silently refuses the key this feature just
injected and falls through to an OAuth login the isolated config dir does not have. It
survived review because a keyless llama.cpp/llama-swap endpoint uses DEFAULT_API_KEY
('local-dummy-key', 15 chars), where slice(-20) returns the whole string and the seed
matches by accident, and every test used a key shorter than that. Now truncated through
`truncateApiKeyForTrustFile()`, with a test using a 57-character key that also asserts
the full credential never reaches that second file.
**One `confirmed` flag answered two different questions.** The context-floor warning
("this model's window is below what this CLI needs") and the swap-conflict warning
("loading this unloads the model another session is using") shared it, and the context
check runs first, so a user clicking "launch anyway" past the context warning silently
consented to evicting someone else's model. They are about different people, so an
answer to one is not consent to the other. Both routes now read `confirmedContext` and
`confirmedSwap` independently; the legacy `confirmed` still means both, because it
shipped in this feature's HTTP-API-only cut and an existing caller must keep working.
The frontend answers each question with its own flag and accumulates them, on the
one-shot path, the restart path and the batch carry-forward alike.
Also from the same review: the swap-confirm dialog no longer renders " are currently
using ..." when multi-user scoping leaves the affected-session list empty (the swap is
blocked regardless of ownership; only the NAMES are scoped), and the per-endpoint
llama-swap log tails are closed in `WebServer.stop()` instead of only by the idle sweep
whose interval that same teardown disposes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>