Commit Graph
402 Commits
Author SHA1 Message Date
Michael GrundbergandClaude Opus 5 05c788ce9d fix(watching): close the review findings on the label and its window
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief)
found the trust boundary weaker than the comments around it claimed. Eleven
findings, all applied.

The two blockers were both about who can write the row the label is read from.
Claude's window covered two rows, and the second one is the status line, whose
command a session running with permissions bypassed can write into its own
`.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of
its own and silence its own idle alert. The default window is one row now, which
is the footer and nothing else, and the constant says why. Separately, the label
reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML,
which is an injection sink for any config-supplied pattern whose capture group is
permissive; it goes through escapeHtml() like every other untrusted string in
that file.

The Codex entry could not be fixed the same way, and now says so. Its row is
third from the bottom only while a terminal runs; with none running that slot
holds the last row of the transcript, so matching the complete row (with the
`/stop to close` tail, window narrowed to three) raises the bar without closing
it. What contains it is `hooks: 'none'`: no hook event from a codex session
reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an
alert. The registry comment, `docs/cli-registry.md` and the test all state that
rather than claiming a guarantee the code does not have.

Also from the review: the TUI header badge no longer counts an acknowledged
item, which was the same gate the classifier fix already went through and was
wrong for human acknowledgement too; the TUI approval card reads the quiet
reason and drops to a new `info` tone instead of asking for a reply; the badge
carries an aria-label, because the phone it was built for has no hover target;
the schema refuses `watchingLines` without a `watchingLine`; and the pattern and
its window are resolved together rather than one memoized and one not.

Documentation moved with it. The mechanism now lives in
`docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer,
`docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped
buzzing, and both that page and the changeset name the limitation neither did
before: a question asked in plain prose is not a dialog, so it is silenced along
with the false alarms while background work runs.

Verified live again after the narrowing, on an isolated beta: a Claude session
reported `1 monitor` and took its idle prompt acknowledged, and a Codex session
reported `1 background terminal` against the full-row anchor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 17:11:33 +02:00
Michael GrundbergandClaude Opus 5 64c288a683 feat(codex): read Codex's own background-terminal row
Codex states background work too, and it says so in a different place. Claude
writes `· 1 monitor ·` on the last row of the screen; Codex pins
`1 background terminal running · /ps to view · /stop to close` ABOVE its
composer, which puts that row third from the bottom once the status line and
the composer are counted.

So how far up the screen to look is now per-CLI data as well:
`capabilities.workDetect.watchingLines`, bounded to 1..8 by the schema, and
defaulting to Claude's two. That bound is the point. The window is half the
injection guard, since every row it adds is another row the agent itself may be
able to write, and the label is what silences an idle alert. The other half is
the anchor, and Codex's is ` · /ps to view`: chrome naming a slash command only
the CLI can offer, so a session that writes "I left 1 background terminal
running for you" into its own output matches nothing.

Measured against a live codex-cli 0.154.0 pane rather than read out of a
binary. The row appears when the terminal starts, follows the composer down as
the conversation grows, and is gone after `/stop`. Verified end to end on an
isolated beta: the session payload carried `watching: "1 background terminal"`
and the badge rendered with it, and both cleared when the terminal stopped. The
fixtures in the tests are that capture verbatim.

Codex has no hook signals, so no idle prompt and no false NEEDS YOU row: for a
Codex session this is the badge alone, which is the case the maintainer said a
registry field could cover and a hook never could. Cross-CLI tests pin that
neither pattern fires on the other's screen, and that a CLI declaring nothing
still reports nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 08:58:52 +02:00
Michael GrundbergandClaude Opus 5 74884a20eb feat(approvals): let a watching session keep quiet, and fix the tui gate
The badge alone left the row in NEEDS YOU, which is the thing the issue was
about. The fix is the alert that does not fire.

An idle prompt from a session that is watching its own background work now
opens ALREADY acknowledged. `hook-event-routes` passes `Session.watching` to
`notePrompt()`, which sets `acknowledgedAt` and records why in a new
`acknowledgedReason`. Nothing new suppresses anything: `acknowledge()` has
always meant "the alert this prompt armed is spent", and the prompt itself
stays pending, answerable and available as Read My Mind context. A wrong label
therefore costs a card that does not blink, never an alert that was never
created.

Every surface follows from that. The broadcast carries the reason, so a live
page declines to arm the tab alert and raises no desktop notification. The push
is skipped, since a false alarm is hardest to ignore on a phone. A reloading
page reads `acknowledgedAt` in `seedApprovals()`, which it already did. And
`classifySession()` now reads it too, which is a pre-existing bug fixed here:
acknowledging on one device cleared the alert everywhere except `codeman tui`.
It re-arms for free, because the next idle prompt supersedes the item and is
built fresh. Only `idle` is eligible, so a dialog that blocks the agent still
goes red whatever else it started.

The label is pane-derived and therefore prompt-injectable, so it is now read
from the last two rows of the screen only, with Claude's pattern anchored on
the `·` its footer joins items with, ANSI-stripped and length-capped at the
source. An agent that prints `· 1 monitor ·` into its own output finds no
match.

Verified on an isolated beta: a session that armed a monitor took its idle
prompt acknowledged with no alert on any surface, wore the badge, and showed
"quiet, watching 1 monitor" on its still-answerable card; the same session with
the monitor killed alerted normally on the next prompt. `test/watching-no-alert.test.ts`
pins both directions across all four surfaces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 08:10:05 +02:00
Michael GrundbergandClaude Opus 5 3f2cde2db7 feat(session): say when a session is watching its own background work
An agent that arms a monitor, backgrounds a shell or hands a task to a
cloud session is told to end its turn. The pane then falls quiet, Claude
Code's idle_prompt notification arrives a minute later, and every surface
files the session under NEEDS YOU with nothing for a human to answer.

Claude states what it is still running on the last row of its screen
(`⏵⏵ bypass permissions on · 1 monitor · ← for agents`). That row is now
`capabilities.workDetect.watchingLine` in the CLI registry, guarded by
compileVersionRegex() like every other config regex, and the idle probe
reads it off the capture it already takes: `watchingLabel()` in
session-activity.ts searches the last five lines only, so a session that
PRINTS "1 monitor" is not mistaken for one running it.

The label lands on Session.watching and rides toLightDetailedState() out
to every surface. The phone overview, the desktop home rail and the rich
sidebar rows wear it as a `watching` badge in the accent colour, beside
the state pill and never in place of it: an agent can arm a monitor and
ask a question in the same breath, and only the pill says which.

Verified end to end against a throwaway session on an isolated beta
instance: the payload carried `watching: "1 monitor"` once the turn
ended, the badge rendered next to a yellow `waiting` pill, and both
cleared when the monitor died.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 18:00:30 +02:00
Codeman maintainer 299a21d5f5 fix(split-pane): merge-time fixes for split-pane sessions (#453)
The maintainer's promised merge-time fixes from the final review of #453:

1. closeSplitPane() tears down a divider drag still in progress, so a split
   that collapses mid-drag no longer leaves body.split-pane-resizing (the
   page-wide col-resize cursor and user-select lock) set until a reload.
2. openSplitPane() re-applies the picker's own exclusions (detached session,
   pid === null, no session record) for a row that went stale while the
   menu sat open, refusing silently like its neighbouring gates.
3. architecture-invariants: the hard-hide of .btn-split is the
   @media (max-width: 1179px) rule in styles.css, not mobile.css.
4. SplitTerminalPane.destroy() nulls onclose (and onerror) beside onopen
   and onmessage.
5. Picker rows drop the data-session-id attribute nothing read.
6. The Pane-A-ends branch collapses with skipPrimaryResize, so the closing
   resize is no longer aimed at the session the server just removed.
7. The {t:'r'} refresh path is single-flight across the fetch and the
   chunked write, coalescing a mid-replay refresh into one trailing re-run.

Tests: split-pane-auto-collapse-unit gains the drag-teardown, exclusion and
skip-resize cases; the new split-pane-terminal-unit covers destroy() and the
refresh single-flight. All were run against the pre-fix module to confirm
they fail there.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit dbd39aed015ae5ae5870aba398bf4b4ab5118e47)
2026-09-21 04:53:19 +02:00
Codeman maintainer d8e85285c9 fix(mobile): merge-time fixes for the prompt composer (#444)
- styles.css: restate the composer overlay's own bottom gutter after the fold rules
  (the generic .paste-overlay longhand erased it: 0px flat, hinge strip replacing it
  folded) and subtract the fold strip from the dialog's max-height
- test/foldable-layout.test.ts: simulate the cascade for
  .paste-overlay.prompt-composer-overlay (fails without the CSS fix); pin the palette
  anchor by name instead of ELEMENTS.at(-1)
- keyboard-accessory.js: guard the app global in refreshForActiveSession() like the
  rest of the file
- keyboard-accessory.js: a whitespace-only draft is empty (Send no longer submits
  blank lines); the text still goes out untrimmed
- keyboard-accessory.js: derive _composerMaxLength and the frame refusal from one
  64 KiB frame limit minus both bracketed-paste markers so they cannot drift
- keyboard-accessory.js: translate the textarea placeholder and label at build time,
  since the DOM translator skips <textarea> subtrees
- i18n.js: zh-CN entries for the composer dialog copy
- docs/wiki/Mobile-Guide.md: describe the Compose key instead of a clipboard key
- CLAUDE.md: a "Mobile prompt composer" paragraph after the accessory bar one
- test/mobile-prompt-composer.test.ts: pin the whitespace rule and the derived budget

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit f6725ba52da17b0bdbee8be3b5011e7cae514f69)
2026-09-21 04:39:28 +02:00
Codeman maintainer 0f955327b2 fix(cli-registry): merge-time fixes for the run-menu consolidation (#458)
- test/opencode-resize.test.ts: retarget the launcher guard at the real code (this.selectSession(firstSessionId), any this.activeSessionId assignment) with an anti-vacuity check; the old strings existed nowhere, so it could never fail
- session-ui.js: restore as comments the two invariants the merged bodies lost (deepseek leaves statusReporting unset, i.e. ON; no effort field for external CLIs, it is Claude-specific)
- docs/cli-registry.md: move the frontend-guard paragraph below the two backend-guard paragraphs so they keep their antecedent, and note the widened comparison shape
- test/frontend-cli-no-id-branching.test.ts: the comparison shape accepts any left-hand identifier (const m = this._runMode; m === 'codex' was invisible), normalized to `mode`; the two `m !== 'shell'` display filters are allowlisted and the remaining blind spots documented
- test/run-mode-dispatch.test.ts: table-driven pin of run() dispatch (claude to runClaude, each RUN_MODE_LAUNCH id to _runCliMode(id), shell to runShell, unknown to runClaude, lock held and released)
- CLAUDE.md: name the second CI-gated guard next to the backend one
- server.ts: every </head> injection passes a replacer function; a clis.json label containing $' re-injected the rest of the document past escapeScriptJson (two render tests pin it, proven failing on the string form)
- _isAltCliMode(): no reference anywhere in the tree, nothing to fix

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 1ea363ff808a62861559bc141e724b163cc1c56e)
2026-09-21 04:37:45 +02:00
Codeman maintainer d3f2ec0220 fix(custom-model): merge-time fixes for the promoted-model picker (#459)
The maintainer's promised follow-ups to opticon454's picker promotion,
applied on the landing branch after the merge (ecb95b5d):

- session-ui.js: the promotion tag ("Currently loaded" / "Last used") and
  the "Default" pill are two separate spans, so a promoted row that is
  also the endpoint's defaultModelId shows both instead of silently
  losing its Default marking; two tests pin it (both fail on the old
  exclusive-slot rendering).
- styles.css: a dedicated #customModelPickModal .set-scope rule, since
  the pill was only styled inside the three settings modals and rendered
  as plain body text here; same skin tokens, modal layout untouched.
- docs/wiki/Custom-Model-Endpoints.md: describe the promotion (currently
  loaded, else last used per device), the separate Default pill, and
  that nothing is ever auto-chosen.
- CLAUDE.md + docs/custom-model-endpoints.md: credit the real "Last used"
  writers (_runCustomModelEntryViaRestart and
  _quickStartWithCustomModelConfirm; runCustomModelEntry only dispatches
  since 88e5b7b2) and drop the now-wrong "both defer to Default" sentence.
- Not done: moving the one-shot "last used" write into
  _runCustomModelEntryOneShot, because the existing one-shot tests assert
  that _quickStartWithCustomModelConfirm writes the key itself.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 0cb0f911adc14a852ba5c2951768a4aa87c25657)
2026-09-21 04:30:01 +02:00
Codeman maintainer dcf9437308 Merge pull request #460 from Ark0N/feat/installer-v2
feat(install): three questions up front, an unattended build, and a URL you can scan
2026-09-21 04:23:15 +02:00
Codeman maintainer 9a2e14a93a Merge pull request #453 from timkjr/feat/split-pane-sessions
feat: split-pane sessions — view two live terminals side by side
2026-09-21 04:23:14 +02:00
Codeman maintainer ecb95b5d67 Merge pull request #459 from opticon454/feature/run-menu-picker-currently-loaded-model
feat(custom-model): promote the currently-loaded/last-used model in the Run-menu picker
2026-09-21 04:23:14 +02:00
Codeman maintainer 72d437ab63 fix(install): fold in both reviews of #460
The two reviews on the PR (DeepSeek Harness, then Claude) found one class of
bug twice and a list of smaller ones; all of them land here, each pinned in
test/install-sh-invariants.test.ts and, where it is bash logic, driven in the
bash:3.2 CI step as well.

The Start line the done screen prints is now composed in one place
(start_command_hint) from every non-default value, the same five the exec
branch exports through export_bind_env, so "do not start" under a sub-path or
a custom port no longer prints a bare `codeman web`. The --lan / --tailscale /
env preset paths read ${CODEMAN_PASSWORD:-$EXISTING_PASSWORD}: a flag re-run on
a unit that carried a password used to rewrite it without the password and
with the unauthenticated ack. --password and --port flip RECONFIGURE so they
reach the unit instead of taking the quiet update path, and `install.sh name`
re-syncs the unit's base URL after the mapping is re-added.

Also: the sudo keepalive is ended before the exec into the foreground server
(exec skips the EXIT trap, and the loop keys on $$); Ctrl+C in the HTTPS-toggle
poll is trapped for the poll only and skips Tailscale for the run instead of
killing the installer; uninstall asks before removing a LaunchDaemon this
installer never wrote; a foreign daemon gets a launchctl kickstart hint and the
done screen stops claiming the new build is running; the preflight summary
reads the Tailscale state with a line grep when node is not installed yet; the
LAN security notice uses the configured port; a bare re-run ends on the done
screen; a build failure after a rename names the install.sh tailscale
recovery; TS_JOINED_HERE (written, never read) is gone; the plan doc and
architecture-invariants say what the code does. A minor changeset is included.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 03:03:31 +02:00
Codeman maintainer 1ba0684438 docs(install): describe installer v2 and the Tailscale naming options
README, the Installation / Remote-Access / Running-As-A-Service wiki pages,
docs/security-architecture.md and CLAUDE.md describe the three-question flow,
the flags, the subcommands, the sub-path answer for an occupied :443 and why
the rename is opt-in. docs/installer-v2-plan.md is the design and the
verification record (what was measured, what still needs a fresh machine);
docs/tailscale-installer-plan.md points at it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 21:38:12 +02:00
timkjrandClaude Sonnet 5 46d8b92049 fix(split-pane): port Ctrl+Shift+C's never-falls-through guarantee to Pane B
The smart-copy gate only entered its selection-check block behind
hasSelection(), so a selection-less Ctrl+Shift+C skipped straight to
`return true` and ceded the keystroke to the browser's own handling
(e.g. Chrome's Inspect-Element binding) instead of matching Pane A's
"never falls through" contract for that chord.

Verified live in a real browser that this is a UX-parity fix, not an
interrupt-safety one: xterm's evaluateKeyboardEvent never emits PTY
data for a shifted ctrl-letter regardless of any gate (only "_" and
"@" get special-cased), so no accidental 0x03 was ever at risk. The
regression test added here asserts on the dispatched event's
defaultPrevented rather than the absence of a WS frame, since the
frame-count check passes vacuously for this exact key combo whether
or not the gate fires.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:34 -05:00
timkjrandClaude Sonnet 5 0b3e086334 fix(split-pane): address Ark0N's fourth pass — PTY-less picker exclusion, hollow chord test, remaining key gates
- buildSplitPickerSessions() now excludes any session with pid === null
  (exited CLI, tripped PTY-exit breaker, a restore that never re-attached).
  Pane B has no equivalent of selectSession()'s auto re-attach POST, so a
  split opened onto one had nothing reading its tmux pane: no terminal
  events ever arrived and Session.write() silently dropped every keystroke
  with no ack either way, while the socket itself reported healthy.
- Fixed the hollow chord regression test: the synthetic keydowns carried no
  keyCode, which is what xterm's evaluateKeyboardEvent switches on to
  produce a data frame at all, so the assertion held regardless of whether
  the gate fired. Adding real keyCodes surfaced a second, real bug in the
  Alt+B case: the event bubbles to app.js's own document-level shortcut
  dispatcher, which really toggles the sidebar and resets the layout
  attribute the gate reads before Pane B's own (later, non-capture) handler
  ever sees it — fixed by driving the app's real settings cache instead of
  only the DOM attribute.
- Ported the two remaining primary-pane gates with real consequences:
  Ctrl+Z (SIGTSTP) is swallowed for every non-shell session, matching
  terminal-ui.js's reasoning (an Ink/TUI agent loop stops dead with no
  visible output otherwise), and Shift/Ctrl+Enter now POSTs to
  /api/sessions/:id/send-key for THIS pane's own session instead of
  letting xterm send a bare \r, which used to submit an incomplete prompt
  instead of inserting a newline. Smart-copy Ctrl+C is re-implemented
  against Pane B's own terminal (copying app.copyTerminalSelection() would
  have copied Pane A's selection instead).
- Updated docs/architecture-invariants.md and docs/split-pane-sessions-plan.md
  to match, and added CLAUDE.md's missing .split-picker-menu z-index entry.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:34 -05:00
timkjrandClaude Sonnet 5 fafef0aa00 fix(split-pane): gate app-level chords out of Pane B, address Ark0N's third pass
Pane B had no attachCustomKeyEventHandler of its own, so the document
capture-phase shortcut handler's preventDefault() (which does not stop
xterm) left Ctrl+K/Alt+1/Alt+B ALSO writing their raw byte/escape
sequence into Pane B's live PTY on top of whatever the app action did
to Pane A. Pane B now installs the same registry-aware gates the
primary pane's own attachCustomKeyEventHandler uses. Ctrl+V is left on
xterm's default paste — no image-paste trap to route it to.

Plus the rest of the review's smaller items:
- Narrowing the window past the desktop gate now closes an open split
  instead of leaving it stranded on screen.
- Split is refused while a web tab is active (activeWebviewId), which
  used to open Pane B's socket behind a hidden container.
- Pane B now handles the server's `{t:'r'}` refresh frame via a shared
  _loadBuffer() helper (also used by connect()), instead of ignoring it.
- The divider drag now uses pointer events + setPointerCapture (mirrors
  tab-rail-resize.js), a button!==0 guard, preventDefault, and a
  body.split-pane-resizing cursor/selection lock — a plain mousedown
  drag selected the text under the cursor as it crossed both terminals.
- Pane B's close control and the picker rows are real <button>s now
  (keyboard-reachable), with matching CSS chrome resets.
- Dropped the redundant CodemanBase.base prefix on the buffer fetch
  (the global fetch wrapper already applies it).
- data-preview-order for the Split settings chip moved from a collision
  with Ultracode Agents (both 15/12) to 11.5, matching its real
  position between Multi-monitor and Ultracode Agents in the header;
  widened test/app-settings-structure.test.ts's regex to allow the
  decimal (Number() already parses it fine for the preview sort).
- Added zh-CN i18n entries for the Split button and empty-picker text.
- Dropped the stray unused `vi` import Ark0N flagged as unrelated to
  this feature (vitest's `globals: true` makes it ambient anyway).
- Documented the fix and the deliberate no-cid/seq choice in the
  split-pane-sessions architecture-invariants entry.

Added a real-Chromium regression test asserting Ctrl+K/Alt+1/Alt+B
dispatched at Pane B's own textarea send no `{t:'i'}` frame over its
WebSocket. Full CI gate green (409 files, 7736 tests) plus all 8
split-pane browser tests.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:33 -05:00
timkjrandClaude Sonnet 5 3152ec801d docs(split-pane): short CLAUDE.md rule, stale module count, wiki entries, shortcut-handler caveat
CLAUDE.md previously only mentioned split-pane in the load-order list,
with nothing in the Architecture/frontend prose the way every other
feature gets, and its own module count was one stale (34, should have
been bumped to 35 when terminal-split.js was added). Add a short
pointer-style paragraph next to the other terminal features, fix the
count.

docs/wiki/The-Dashboard.md's header button table and
docs/wiki/Settings-Reference.md's header chips list are the two
user-facing surfaces that never mention Split at all; added both, plus
a note that the feature is desktop-only regardless of the setting.

docs/split-pane-sessions-plan.md: recorded the one design note that
isn't a code change — the global capture-phase shortcut handler always
resolves against Pane A, so Ctrl+L/Ctrl+W typed into Pane B affects the
other session. Not fixed for v1, same reasoning as the rest of the
"deliberately plainer" section.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:33 -05:00
timkjrandClaude Sonnet 5 6c8bd6c606 fix(split-pane): gate the Split button to desktop, make it per-device
Ark0N's PR #453 review: nothing gated this feature to desktop even
though the design called for it (two 240px min-width panes plus the
divider need ~486px, and the divider has no touch handlers), and
showSplitButton was a SYNCED setting, so turning it on at a desk also
put the button in the phone header.

- Hard-hide .btn-split on phones in mobile.css regardless of the
  setting, matching the other desktop-oriented header buttons in the
  same @media (max-width: 599px) block.
- Move showSplitButton into settings-ui.js's per-device displayKeys
  set and drop it from SettingsUpdateSchema entirely, matching the
  showFileViewerButton/skin precedent (CLAUDE.md's "per-device keys
  ... must NOT be added to SettingsUpdateSchema" rule) — a desktop
  opt-in must never sync onto a phone that never asked for it. Removes
  the now-invalid server-round-trip test for the setting.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:32 -05:00
timkjrandClaude Sonnet 5 7fc66e8161 docs(split-pane): keep the design spec, drop the task-plan scaffolding
Per Ark0N's review on PR #453: rename the design spec to
docs/split-pane-sessions-plan.md, matching every other feature's
*-plan.md convention, and drop the 957-line implementation task plan
(docs/superpowers/plans/2026-09-15-split-pane-sessions.md) — workflow
scaffolding for the subagent-driven-development run, not repo
documentation. Fixes the now-dangling link in architecture-invariants.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:30 -05:00
timkjrandClaude Sonnet 5 f7852081b7 docs(split-pane): fix orphaned Session list layout section
The new "Split-pane sessions" section was inserted between the "Session
list layout (header strip vs. left sidebar)" heading and that section's own
body paragraphs, orphaning the heading from its content. Move "Split-pane
sessions" to after the Session list layout section's full body, before
"Gesture control: the setting" — no change to the Session list layout prose
itself, only where the new section sits relative to it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:28 -05:00
timkjr ba7b8b7bef docs: add split-pane sessions architecture-invariants entry 2026-09-20 13:10:26 -05:00
timkjrandClaude Sonnet 5 6f64e557e5 docs(plan): fix session-creation test bug found by Task 4's implementer
Task 4's implementer found two real bugs in this plan's browser-test
helpers: POST /api/sessions nests the id at data.session.id (not
data.id), and mode:'shell' needs a follow-up POST .../shell to actually
spawn a PTY. Fixed in Task 4's own snippet (documentation accuracy —
already fixed in the real committed code) and pre-emptively in Tasks
5/6's createShellSession() helper before either was dispatched, so
neither implementer has to rediscover it independently. Also corrected
the <script> tag snippet to defer, matching the real file's convention.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:25 -05:00
timkjrandClaude Sonnet 5 727817410c docs(plan): fix Task 6's SSE handler patch to target the prototype
Monkey-patching the instance's _onSessionDeleted inside a
DOMContentLoaded listener races connectSSE()'s handler-wrapper cache,
which captures the function reference by value on first connect and
never re-reads it. Patching CodemanApp.prototype at module-evaluation
time (synchronous script-tag order) is unraceable: it completes before
any instance exists or connectSSE() ever runs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:22 -05:00
timkjrandClaude Sonnet 5 28d3bd7da8 docs(plan): fix Task 2/4/5/6 tests against real test infrastructure
Preflight scan for SDD execution caught two classes of defect before
dispatch: Task 2's test invented a buildTestApp() helper and response
envelope that don't exist for /api/settings; Tasks 4-6 used
@playwright/test's runner against a test/browser/ directory that
doesn't exist in this codebase. Both corrected against real patterns
found in existing tests (system-routes-settings-partial-put.test.ts,
terminal-copy-shortcut.test.ts, tab-rail-resize.browser.test.ts).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:22 -05:00
timkjrandClaude Sonnet 5 d0f9bdd251 docs: fix plan wording and add execution-environment note
Global Constraints previously read as if local-echo/CJK/accessory-bar
were desktop features; they are mobile-only, and split-pane is the
desktop-only side of that equation. Also names the exact spec section
instead of a loose paraphrase, and adds a worktree/branch note so an
executing subagent knows where this plan runs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:22 -05:00
timkjrandClaude Sonnet 5 ff98006471 docs: add split-pane sessions implementation plan
7 tasks: pure divider/picker helpers, showSplitButton header wiring,
split-container CSS, SplitTerminalPane (Pane B's independent xterm+WS),
open/close orchestration with picker and divider drag, auto-collapse on
either session ending, and an architecture-invariants entry.

Also folds in the "detach session" prior art discovered mid-brainstorm
into the spec's architecture section.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:21 -05:00
timkjrandClaude Sonnet 5 5f1be90ae9 docs: fix tab/pane terminology in split-pane spec
The Problem paragraph and the architecture section used "tab" where
"pane" was meant, colliding with the browser's own tab concept.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:21 -05:00
timkjrandClaude Sonnet 5 3da8bb7046 docs: add split-pane sessions design spec
Scopes v1 of an in-app split view (two live session panes side-by-side,
draggable divider) after multi-monitor spanning turned out to solve a
different problem than showing multiple panes at once.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:20 -05:00
DevvynandClaude Sonnet 5 d9e6ebb20a fix(cli-registry): address round-2 review on #458 — count-based allowlist, RUN_MODE_LAUNCH drift guard
Three of Ark0N's four "will take at merge" items, applied instead since
they were straightforward to do properly:

1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on
   <file>::<expression> (fixed last round) closed the line-shift problem
   but opened a new one: every stock id was already allowlisted for
   session-ui.js in the `mode === '<id>'` form, so a BRAND NEW branch
   reusing that exact expression anywhere in the file passed unnoticed.
   Reproduced live (`if (this.mode === 'codex')` injected into
   runOpenCode()) — stayed green under the old version. Each allowlist
   entry now carries the exact count of approved call sites, and a new
   test asserts actual-vs-declared count for every key; a mismatch in
   either direction is real (higher = new unreviewed branch riding in on
   an existing approval, lower = a reviewed site was removed and the
   entry is now stale). Reproduced again against the fix: same injection
   now fails with an exact diagnostic (expected 2, found 3).

2. Added test/run-mode-launch-table-drift.test.ts. RUN_MODE_LAUNCH
   restates four things stock.ts already owns (label, install command,
   supportsCustomModel, the external-mode key set), and they agree today
   with nothing enforcing it. supportsCustomModel is the dangerous one:
   the Run-menu picker's rows come from the server-injected
   window.__codemanCustomModelClis (built from
   capabilities.customModelInjection.kind), so a CLI gaining a real
   injection recipe later would be OFFERED in the picker while
   _runCliMode silently drops the customModel field for it — the session
   launches on the vendor's cloud while the UI claims the local endpoint.
   Drives the real session-ui.js via JSDOM and compares RUN_MODE_LAUNCH
   against STOCK_CLIS on all four axes.

3. Inlined the "Open Question 7 in PR-B2.md" references in the allowlist
   reasons — PR-B2.md is a local planning doc, never part of the
   committed tree, so the reference was dead on arrival for anyone
   reading the repo. Points at the PR #458 review thread instead.

4. Added a sentence to docs/cli-registry.md naming the new frontend guard
   alongside the backend one it mirrors.

Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format/
check:frontend-syntax all clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-20 21:37:05 +08:00
DevvynandClaude Sonnet 5 88e5b7b200 fix(custom-model): address Ark0N's PR review — client-side probe timeout, defer "last used" past confirmation, docs, zh-CN
Four things from the maintainer's review on PR #459, all fixed:

1. Bound _getCustomModelCurrentlyLoaded's probe client-side (~800ms via
   Promise.race, on top of — never instead of — the route's own 5s
   server-side timeout). Without it, an asleep/firewalled endpoint behind
   a saved model list left the picker completely invisible for up to 5s
   after the Run menu had already closed, with no spinner or toast.
   `timeoutMs` is an optional param (default 800, real callers never pass
   it) so a test can drive it in milliseconds, same pattern as
   `_watchLlamaSwapLoading`'s own `pollIntervalMs` — this code runs in a
   JSDOM window's own realm, whose setTimeout vi.useFakeTimers() cannot
   patch.

2. "Last used" is now written only once a launch actually applies, never
   on the mere click. It moved out of runCustomModelEntry (unconditional)
   and into each path's own success point: _quickStartWithCustomModelConfirm
   after the final post succeeds, and _runCustomModelEntryViaRestart right
   after the apply's success check. A context-window-warning decline means
   this exact model cannot work with this CLI at all, so the old
   unconditional write would promote, next time the picker opened, the one
   model guaranteed to fail again.

3. Documented the promotion/tag precedence and the new
   codeman:customModelLastUsed:<mode>:<endpointId> localStorage key in both
   CLAUDE.md's Custom Model Endpoint Profiles section and
   docs/custom-model-endpoints.md's Run-menu picker section.

4. Added zh-CN entries for "Currently loaded" and "Last used" in i18n.js,
   next to this modal's existing "Choose a model"/"Custom Endpoints" pair.

New tests: the client-side timeout (endpoint that never answers, one that
answers within the bound, and a rejected-after-timeout probe settling
quietly), and "last used" recording on success vs. NOT recording on either
confirmation's decline, for both the restart and one-shot paths.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
2026-09-20 21:19:42 +08:00
Codeman maintainer 4205f6930f fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine.

**The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext`
and `confirmedSwap` changed the wire field without moving three assertions that check
it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts`
(the swap modal and the context modal, each of which already receives exactly the right
per-question flag). Moved, with the titles.

**Worse, my own tests for the split never ran.** The four cases in
`session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`,
which was declared inside a sibling `describe`, so they threw a ReferenceError during
setup. The split would have shipped with no passing server-side coverage while the gate
reported the failure as four broken tests rather than as four tests that were never
written. `mockRunning` is hoisted to the outer describe.

**The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved
its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so
the other eight modes fell back to claude's `❯`. That is also starship's default shell
prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line
`❯ npm run build` sits on screen for as long as the command runs, the verifier reads it
as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to
nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N
prompt, `read -p`, an installer or a pager, where it takes the default. The module's own
fileoverview already stated the rule this broke. Now `?? ''`, which
`promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI
that actually declares a composer.

**My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing
three numbered rules.

The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared
in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told
an integrator to retry with `confirmed: true` for both questions, which is precisely the
thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md`
and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR`
multi-user consequence, both user-visible and both previously absent, and #454's gained
the one exception to its own claim: a Custom Endpoints launch ignores the Instance count
stepper and always starts one session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:58:33 +02:00
Codeman maintainer 12de3c5164 docs(custom-model): record the CLAUDE_CONFIG_DIR clamp in architecture-invariants
CLAUDE.md gained the admin-only note when the key joined claude's privilegedEnvKeys;
architecture-invariants, which is where the exact-key allowlist rule is documented in
depth, still described the pre-change world. The reboot-restore half is the one worth
writing down: a non-granted owner's already-persisted CLAUDE_CONFIG_DIR is stripped on
restore, which moves that session back to the default Claude account with no error.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:35:32 +02:00
Codeman maintainer 9af12afb57 docs(custom-model): make the docs match the code, and trim the changeset
More from the review of 5fc391a4, all documentation rather than behaviour.

The changeset was 1602 words of development log, written as the PR grew, with bullets
and loose paragraphs interleaved. That text becomes CHANGELOG.md and the GitHub release
body verbatim, so it is now one user-facing account of what the feature does and what
the real-server work bought, at roughly a fifth the length.

docs/api-reference.md promised a `cmd` field on running-status that the route
deliberately strips (it carries model paths and can carry --api-key).

Two places claimed the apply routes validate `modelId` against the endpoint's
discovered models. Neither does. Dropped the claim rather than adding the check:
discovery can be up to five minutes stale, so a 400 there would refuse a launch that
actually works, and a typo'd id already fails on the CLI's own first request. CLAUDE.md
now says so explicitly, since the absence is the surprising part.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:34:24 +02:00
Codeman maintainer 1a99b5836c Merge pull request #430 from opticon454/custom-model-run-menu 2026-09-19 12:25:11 +02:00
Codeman maintainer 035bfbc2fe fix(remote): merge-time fixes for Wake-on-LAN
The MAC-count limit lived in two places that disagreed. RemoteHostSchema.wakeMac's
128-character cap admits seven comma-separated MACs while parseMacList takes at most
four, all-or-nothing, so a five-MAC value validated, was written to remote-hosts.json,
and then resolved to NO wake target: POST /api/sessions/:id/wake answered
"No wake-on-LAN target configured for this host" and the banner offered "Configure WoL"
for a host the user had just configured. MAX_WAKE_MACS now lives in
src/config/remote-wake-limits.ts and both sides refine against it. Its own module
because src/remote-wake.ts is import-fenced to session-routes.ts and server.ts (the
wiring guard that stops a watcher waking a host), and because schemas.ts must not drag
dgram/net/child_process into every request-validating module.

The documented 40 s request budget also omitted the wake's own cost. A `command` target
is bounded by REMOTE_WAKE_COMMAND_TIMEOUT_MS and runs BEFORE the readiness poll, so a
slow one pushed a wakeCommand host's worst case to ~68 s, past the 60 s
proxy_read_timeout the budget exists to stay under. _wakeAndWait now subtracts the
wake's measured elapsed time from the readiness budget, floored at one poll interval so
a wake that ate the whole budget still gets one probe. A magic packet is effectively
instant and is unaffected, which is why live testing never saw it.

Also: the two new endpoints are documented in docs/api-reference.md with the import
fence stated as the rule it is, CLAUDE.md's frontend module count moves to 34, and the
release changesets carry the Thanks section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:40 +02:00
Codeman maintainer 4c705094f7 fix(terminal): ship the copy clean as a trailing trim, without the shared dedent
#451 cleaned two things on copy. The trailing trim is right and every native
terminal does it. The shared leading-indent strip is this project's own rule,
and it is dropped here rather than shipped.

Measured against the shipped transform over 401,445 three-row windows across
1,010 tracked files in this repo, it fired on 73% of them: 92% inside a YAML
workflow, 76% over `git log` output, 48% in a TypeScript source. No width
threshold separates a margin from content because they are the same widths, a
live Claude Code pane's own margins measuring 2 and 5 columns while the most
common non-TUI shared run is 4. The failure modes are not symmetric either: a
wrong trailing trim costs nothing, while a wrong dedent silently deletes
information that was on the screen, with nothing in the clipboard to hint at
it, on git log bodies, on indented code read out of cat (semantic in Python),
on git diff context rows where the leading space is the marker, and on stack
traces.

It also could not be made self-consistent cheaply. Whether the first row joined
the measurement depended on the mousedown COLUMN, which the user never sees, so
one block of three rows produced three different clipboard results; and the
flag read getSelectionPosition().start, which is xterm's mousedown anchor and
is never normalised, so dragging UP through a block read it off the bottom row.
The PR's test stub hardcoded a downward drag, so its suite could not express
that case.

The transform, the wiring, the tests, the invariants, CLAUDE.md, the wiki page
and the changeset all move together. The test block now pins the ABSENCE as a
contract, with the git log, Python and git diff cases as its examples, so this
is not re-derived later. If it is ever revisited, the one qualification that
measured clean is painted trailing padding: zero false positives over all
401,445 windows.

Also from the review: the comments and invariant rule justifying the
padding-only clear described the pre-change code (the Ctrl+C gate reads the
CLEANED selection now, so such a selection falls through to the PTY on its own
and the clear is feedback rather than protection), the new 'Nothing to copy'
toast gained its zh-CN entry, and the invariants paragraph no longer repeats
its own opening sentence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:39 +02:00
Codeman maintainer c376534a50 fix(run,terminal): merge-time fixes for the Instance count stepper and capture geometry
#454: the behaviour the PR adds had no test, so a regression test drives
runGrok() at tabCount 3 and asserts three quick-start POSTs with sequential
w<n>-<case> names (verified to fail against master's session-ui.js). Each
caller now reads the count BEFORE its opening banner and announces it there,
the way runClaude() already did, so a launch no longer prints two headers and
a launch with another session already active still says how many are starting.
runClaude() calls the shared _readTabCount() instead of its own copy of the
1..20 clamp, and that helper optional-chains the element read, since hoisting
it above each caller's try block would otherwise let a missing #tabCount throw
where the launch-error path cannot report it.

#435: sizeMovedUnderLoad derived from data.source alone. `mux-visible` is not
sufficient: a failed display-message cursor query makes capturePaneBuffer skip
the snapshot repaint and return the raw capture, which the route still labels
mux-visible, so a size that moved during such a load bought a full forced
reload to repair a frame that was never positioned. It now tests
Number.isFinite(data.captureRows) like its two siblings.

Plus the invariants and CLAUDE.md lines promised on #435: a visible capture
reports its geometry and omits it when nothing was positioned, the comparison
runs on mux-visible only, and the replay is capped at one attempt and latches
per session when it cannot converge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:39 +02:00
Ark0N 2c3ccdf030 Merge pull request #439
feat(remote): wake a sleeping host (Wake-on-LAN) from input, banner and native magic packet
2026-09-19 12:18:18 +02:00
Ark0N 613b774bf1 Merge pull request #451
fix(terminal): trim the padding and shared indent out of a copied selection
2026-09-19 12:18:08 +02:00
RandalixandClaude Opus 5 5bb489addb fix(remote): authorize the attach wake first; tell the caller what happened to its bytes
Review round 3 on #439.

- The attachRemoteSession branch of POST /api/sessions ran `ensureHostAwake`
  before the multi-user gates, so a non-admin could have any configured
  host's `wakeCommand` spawned (or a packet broadcast) and the request held
  for the wake budget, then be refused for the workingDir. The admin gate
  now comes first, before the host is even looked up; remote hosts are
  admin-only infrastructure everywhere else. Route test: wake spy empty,
  403.
- The non-wait input route answers `{buffered:true}` when the registry took
  the chunk and `{buffered:true, dropped:true}` when it was over the cap
  and is gone (`RemoteInputOutcome` gains 'dropped'); additive to the bare
  `{}`.
- The send-and-wait path answers OPERATION_FAILED when the host never comes
  back, like create and attach, instead of writing into the stalled pane
  and reporting delivered:true plus a timeout.
- The flush writes with `fromUser: true`, so a first prompt buffered
  through a wake can still name the tab.

Docs: api-reference (input route), remote-sessions.md (two invariants),
CLAUDE.md key pattern.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-19 11:39:50 +02:00
RandalixandClaude Opus 5 1040f6c489 fix(remote): a proxied host is reachability-unknown; scope remote: SSE per session
Review round 2 on #439.

1. The bare TCP probe connects to host:port, which a host behind a jump host
   or SOCKS proxy does not answer even while ssh works. Acting on that
   verdict drew a permanent banner over a healthy session, replaced a real
   "needs tmux" error with "not reachable" in quick-start, and - with a wake
   target - buffered every HTTP input for the life of the session, since the
   readiness poll could never succeed. `WakeableRemote` now carries
   `jumpHost`/`socksProxy`/`extraSshOptions`, and `isProbeable()` turns such
   a host into reachability-UNKNOWN: input is delivered, `checkReachable` /
   `checkHostReachable` answer `null` (never `false`), `ensureHostAwake`
   returns `'unprobeable'` (handled like `'no-target'`), the quick-start gate
   fires on `=== false` only, and `GET …/reachability` reports
   `reachable: null, probeable: false` so the banner has nothing to key on.
   A wake target can still be fired for it, blind: no readiness poll, no
   reattach, no toast - the response says only whether the packet went out.

2. `'remote:'` joins the session-scoped SSE prefixes. The create/attach wake
   has no session yet, so the registry names the requesting user
   (`ensureHostAwake({ requestedBy })` -> `username` in the payload) and
   `deriveSseHint` routes on it; with neither it fails closed to admins.
   Single-user mode is unaffected.

Smaller, from the same review:

- A flush write that fails now drops the remaining buffer (logged) instead
  of retaining it: the wake still resolved and marked the host reachable, so
  the retained chunk waited for the NEXT wake and was replayed hours later,
  after everything typed since. Same policy as the oversized paste.
- The banner polls on tab activation (a user action) and on its 30 s timer
  only for a host with a wake target; a timer connecting to a host Codeman
  cannot wake is the traffic invariant #2 rejects keepalives for. A proxied
  host is never polled.
- `probeRemoteHostReachable`, `runRemoteWakeCommand` and the default UDP
  socket refuse under VITEST, as remote-files.ts does. The guard caught a
  leak on the spot: `createDefaultRemoteWakeDeps({ probe })` overrode the
  probe but still polled readiness with the real one, so the shutdown test
  had been connecting to a production address. The poll now uses the
  injected probe.
- docs/remote-sessions.md is additions only again (the reformatting is
  gone); the architecture-invariants overlap resolved itself in the merge.

Live, against a throwaway instance with a non-routable ghost host: proxied
-> no probe, no wake, the genuine ssh error after 10 s; direct (control) ->
probe, magic packet, "did not come back" after the 40 s budget.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-18 22:46:11 +02:00
RandalixandClaude Opus 5 e271a65e79 Merge origin/master into feat/remote-host-wake
Resolves CLAUDE.md count tables (route counts recounted on the merged
tree: 235 handlers, sessions 37) and keeps both the host-wake and the
reboot-restore banner in index.html.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-18 22:20:41 +02:00
Ark0N 492f8d8ddf Merge pull request #436 from irisitymichaelgrundberg/fix/replay-output-that-arrived-after-the-capture
fix(terminal): keep the output a pane capture could not contain
2026-09-18 21:34:42 +02:00
Devvyn 56209e7829 Merge remote-tracking branch 'upstream/master' into feature/run-menu-custom-model-picker 2026-09-19 03:25:26 +08:00
DevvynandClaude Sonnet 5 afb6754453 fix(custom-model): address third pre-merge review (Ark0N)
Blocker 1: the loading banner hides itself ~200ms after it reopens.

- _showCenterStatus reuses one shared DOM node; dismiss() scheduled
  el.hidden = true 200ms later with nothing to cancel it. On the
  Claude path, switchingToast.dismiss() is followed by one same-
  origin request (5-30ms locally) before _watchLlamaSwapLoading opens
  the new banner -- well inside that window -- so the stale timer
  fired against the shared node and hid the fresh banner, leaving the
  whole model-load wait with no progress text, no log line and no
  reachable Cancel button.
- Fixed by parking the pending timeout on the element and clearing it
  at the top of _showCenterStatus. Added a regression test that
  reproduces the exact repro (open, dismiss, reopen 20ms later,
  advance past 200ms) alongside the existing Cancel-button DOM tests;
  confirmed it fails without the fix and passes with it.

Blocker 2: the swap-conflict warning named other users' sessions.

- Both affectedSessions scans (POST .../custom-model and quick-start)
  walked the whole session map with no ownership filter, so in multi-
  user mode a non-admin pointing their own session at a shared
  endpoint learned another user's session name and id -- which with
  autoNameSessions on is that user's own prompt.
- The swap is still blocked pending confirmation regardless of
  ownership (a foreign session is just as real a disruption); only
  which ones get NAMED back to the caller is scoped, via the
  already-imported canAccessOwned. Added a two-owner test to
  test/routes/session-custom-model.test.ts covering both the
  foreign-owner (blocked, not named) and same-owner (named) cases.

Smaller ride-along fixes:

- server.ts boot recovery now passes contextLength into
  applyCustomModelInjection, so CLAUDE_CODE_MAX_CONTEXT_TOKENS is
  correctly rebuilt into _envOverrides after a restart instead of
  surviving only because tmux retains the old setenv.
- pumpLlamaSwapLogTail's finally now deletes by IDENTITY, not just by
  key, so an aborted pump finishing after a newer entry was created
  for the same endpoint can no longer delete that newer entry and
  orphan its connection.
- docs/custom-model-endpoints.md now notes that clearing a custom
  model removes injected keys by name, including CLAUDE_CONFIG_DIR --
  so a session that also had CLAUDE_CONFIG_DIR set via envOverrides
  (the per-client-account case) silently falls back to the default
  account on clear.

Left for later, as flagged in the review itself: the quick-start
case-scaffolding/cancel ordering (real behavioural reordering across
a large handler, too risky to make without a live re-test), and
retiring runCustomModelEntry's mode === 'claude' branch behind a
launchStrategy registry field (explicitly deferred by the reviewer to
"the next one").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-19 03:09:36 +08:00
Michael GrundbergandClaude Opus 5 f9edb33d15 fix(terminal): trim the padding and shared indent out of a copied selection
xterm hands back whole screen rows and trims only the cells that were
never written to, so the real spaces a full-screen TUI paints across the
unused part of a row count as content and reach the clipboard. Measured
against Claude Code in a 282-column pane, single lines arrived carrying
138 trailing spaces, and every line carried the two-space transcript
indent as well. Windows Terminal, iTerm2 and GNOME Terminal all trim that
for you, decideAutoCopy already calls a wall of spaces "never what the
gesture meant", and _selectTouchSelectionLine already treats those cells
as padding — the mouse and keyboard paths never had the same rule.

CodemanCopySelection.clean lives in constants.js beside decideAutoCopy,
its pure sibling. It drops the trailing run from each line, and removes
the leading run only where every selected row shares one. A selection of
a single row keeps its run, because one row shares nothing with anything
and stripping it would silently reindent one line of `git log` body text
or one line out of `less`. A drag that began inside a row keeps its
partial first line untouched and out of the measurement, which otherwise
pins the shared run to zero and leaves every following row indented.

Every pass over a line is a scan rather than a regex. `/[ \t]+(\r?)$/` is
quadratic on a line whose spaces are followed by a non-space character,
which is what right-aligned or centred TUI content looks like: measured
over 50 000 rows with a 280-column run it took 2.9s, against 1.3ms for
the scan, and a 2 000-column run took 16s. The scan is also the faster of
the two on an ordinary padded row.

cleanedTerminalSelection in terminal-ui.js is the half that needs the
live terminal. It returns a COLUMN selection untouched: Alt+drag makes
one, and a rectangle's rows lining up is the point of the gesture, so
both halves of the clean would destroy it. xterm exposes the mode nowhere
public, so the check reads terminal._core._selectionService, the way this
file already reads terminal._core for cell dimensions, and cleans
normally if a future xterm renames the field. A test pins that assumption
against the library rather than against a stub repeating the literal.

The Ctrl+C chord decides on the cleaned selection, not the raw one. A
drag across the blank part of a row selects real padding spaces, so the
raw text is truthy, and testing it would spend that press on a copy of
nothing and make the user press again to interrupt. A padding-only
selection is now dropped and the press falls through to the PTY, while
Ctrl+Shift+C still never falls through. copyTerminalSelection gates on
trim() for the same reason, since a multi-row drag across padding cleans
to line breaks alone and a bare newline pasted into a chat composer
submits it.

All four of the main terminal's copy paths go through it: the Ctrl+C
chord, right-click, the phone selection button and Auto Copy. The
browser's own Edit menu copy, a disabled copy shortcut and the subagent
windows still copy raw rows, as they did before, and the invariants doc
now says so rather than claiming every copy is cleaned. Auto Copy
resolves its own toggle before it reads the selection, since it is off by
default and a selection can run to the 50 000-row scrollback ceiling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 20:18:25 +02:00
Michael GrundbergandClaude Opus 5 75a028e825 fix(terminal): re-take the sticky-scroll baseline after a replay
A capture load now replays its queued tail, and that replay runs through
`batchTerminalWrite`, which samples `_wasAtBottomBeforeWrite` before it queues.
It runs inside `chunkedTerminalWrite`, before that promise resolves, with the
terminal freshly reset and rewritten — so the sample is always true. The caller
then restored the reader's position and the next `flushPendingWrites` scrolled
straight back to the bottom off the latched flag, undoing it. The only thing in
the way was `_hasRecentUserScrollUp()`, a 1500ms window a server-triggered
refresh is usually past.

`_syncStickyScrollBaseline()` re-takes the flag from wherever the viewport now
sits, and the two paths that restore a position call it right after doing so:
`_onSessionNeedsRefresh` and `_maybeRefetchFullHistory`. Those are the paths
#259 and #205 exist for, and they are also where a non-empty queue is most
likely, since a needsRefresh fires when output is flooding. Re-taking rather
than suppressing the sampling: suppressing leaves whatever stale value the flag
held from before the load, which on the full-history re-pull has no reason to
be false. `selectSession` and `_onSessionClearTerminal` deliberately end at the
bottom, so the sampled true is already the truth there and they do not call it.

`_bufferLoadFinishOpts` gains the coverage the CI gate can see: both mux
sources flush, `history` does not, and a payload naming no source does not.
Its only coverage was the browser suite, which CI does not run.

The JSDoc and the changeset now record the one duplicate window this cutoff
cannot close. The server appends output to the byte buffer in the same tick it
emits, but broadcasts on a batch timer — 8ms over WebSocket, 16 to 50ms over
SSE — so a batch pending when `capture-pane` ran leaves the server after the
reply and is replayed although the capture holds it. It is one batch interval
wide against a recovery window spanning the whole chunked write, and closing it
means flushing that batch server side before the capture.

The second browser test asserts its session was created, so a failed create
fails it instead of passing with zero hits.

docs/architecture-invariants.md no longer claims the replay leaves the
queued-event discard window alone. That clause now describes what decides how a
load ends, the baseline rule, the batch window, and the three covering tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:43:04 +02:00
DevvynandClaude Sonnet 5 9982a1325f fix(custom-model): address second pre-merge review (Ark0N)
Blocker: .center-status-banner never actually disappears.

- Add `.center-status-banner[hidden] { display: none; }`, same trap as
  `.home-sessions[hidden]`: the author-level `display: flex` beat the
  UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
  card stayed laid out at `opacity: 0` with its text/cancel/close
  children still `pointer-events: auto` -- an invisible 442x67 click
  blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
  banner (10001) and the swap-confirm/context-warning modals (10010)
  in CLAUDE.md's Z-index layers list.

Stale wording pointed at the reverted sticky-toast default:

- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
  `.toast-message` comment in styles.css all still said "toasts
  default to sticky" after 1f32128c put the flat 3s default back.
  Reworded all three to describe the actual behaviour: one call site
  passes an explicit `duration: 0`.

Smaller items from the same review:

- docs/api-reference.md said discovery failures answer
  `502 OPERATION_FAILED`; OPERATION_FAILED is 422 per src/types/api.ts
  and the error-code table earlier in the same file.
- The periodic re-discovery sweep (server.ts) never read
  customModelEndpointsEnabled, so turning the feature off left
  Codeman polling every saved endpoint forever. Added
  readCustomModelEndpointsEnabled() (custom-model-routes.ts, same
  shape as readPlanUsageTelemetryEnabled) and gated the interval
  callback on it.
- Reverted the formatting-only Prettier pass docs/api-reference.md
  picked up (table padding, *x* to _x_, JSON re-indent) by re-merging
  the new Custom Model Endpoints section onto the pre-PR file, so the
  diff is reviewable. No prose content was lost -- verified by diffing
  the result against the pre-revert file (formatting-only) and against
  the merge-base file (only the new section added).
- docs/custom-model-endpoints.md now states that a custom-model Claude
  session's isolated CLAUDE_CONFIG_DIR loses the user's global
  settings.json, user-level skills/agents/commands, and MCP servers
  from ~/.claude.json -- only `projects` is symlinked back.

Design question left open in the review (does `confirmed: true` need
to be two flags so "launch anyway" on the context warning doesn't also
skip the llama-swap displacement warning): keeping the single flag, as
offered. The 20s displacement sweep still catches a resulting swap
after the fact, so it's a surprise rather than a silent failure, and
splitting it is real behavioural surface I have no way to verify live
in this environment.

`npm run test:browser` could not be run in this environment (no tmux,
no downloaded Playwright browser binary) -- none of its suite's files
touch code this fix changes, but it still needs a real pass before
merge, same as any frontend change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-18 21:45:50 +08:00
Codeman maintainer bb8ada7e5f fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does.

1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred
   it from the name: a session the user renamed by hand to something shaped
   like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the
   next prompt overwrote their name. The route persists right after, so the
   loss went to disk. `restoreMuxSessions()` already passes it.

2. The already-live sets were snapshotted once before a loop that awaits a
   real `startInteractive()` per entry, so by the tenth entry the snapshot
   was tens of seconds old and a conversation resumed by hand from the
   Resume list in that window was invisible to it: two panes on one
   transcript, the exact thing the check exists to prevent. Both sets are
   now read per iteration, and the late case is spent rather than re-offered
   for the same reason the batch case is.

3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path.
   The stamp predates the reboot and the pane is new, so honouring it meant
   one click had every restored session type `continue` into itself about a
   minute later, unattended, against the route header's own promise that a
   restored session comes back idle and disarmed. The setting stays ENABLED,
   so it re-arms on the next real limit message. A Codeman restart still
   re-arms from the stamp, because the limit footer will not reprint on its
   own; the new option exists only to tell the two paths apart.

4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()`
   and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession`
   performs that it was missing. Cosmetic, but a run left open reads as
   still going in the away digest.

5. A restored claude session gets `seedAgentSessionPreamble()` like both
   create paths, so the agent skill's bootstrap stays a two-line loader.

6. The heuristic's container comment was wrong in one direction and quiet
   about the real gap: after a genuine host reboot a containerized Codeman
   sees the host's short uptime and the banner does appear. What it cannot
   see is a container-only restart, which is where this would help most.

7. The banner is hidden in a solo window, which shows one session and has
   no tab strip to put restored ones in.

Also reverts 17 of the 18 hunks in docs/api-reference.md, which were
Prettier reformatting of prose the PR does not otherwise touch (docs/ is
outside the format glob), keeping only the Reboot restore section and
repairing the two continuation lines that reformat de-indented; renumbers
reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which
loads after it; and gives the feature its CLAUDE.md entry plus a route
test for the multi-user workspace-forbidden branch, the only new rule that
had nothing behind it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:46:04 +02:00
DevvynandClaude Sonnet 5 e2034177c5 fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1)
DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.

Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
  published dependencies) into a scratch dir purely to read
  dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
  completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
  the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
  completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
  same endpoint. dsh's own error template ("DeepSeek API error (HTTP
  ${status})") reproduces the originally-reported
  "dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.

- New registry field `appendV1Suffix` (env kind only, deepseek's entry
  alone — claude/gemini must NOT get it, since claude was already
  confirmed working against the unmodified baseUrl). When set,
  buildCustomModelInjection runs endpoint.baseUrl through the same
  withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
  use, instead of writing it verbatim.

Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.

2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:41:19 +08:00