Commit Graph
376 Commits
Author SHA1 Message Date
timkjrandClaude Sonnet 5 d0f9bdd251 docs: fix plan wording and add execution-environment note
Global Constraints previously read as if local-echo/CJK/accessory-bar
were desktop features; they are mobile-only, and split-pane is the
desktop-only side of that equation. Also names the exact spec section
instead of a loose paraphrase, and adds a worktree/branch note so an
executing subagent knows where this plan runs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:22 -05:00
timkjrandClaude Sonnet 5 ff98006471 docs: add split-pane sessions implementation plan
7 tasks: pure divider/picker helpers, showSplitButton header wiring,
split-container CSS, SplitTerminalPane (Pane B's independent xterm+WS),
open/close orchestration with picker and divider drag, auto-collapse on
either session ending, and an architecture-invariants entry.

Also folds in the "detach session" prior art discovered mid-brainstorm
into the spec's architecture section.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:21 -05:00
timkjrandClaude Sonnet 5 5f1be90ae9 docs: fix tab/pane terminology in split-pane spec
The Problem paragraph and the architecture section used "tab" where
"pane" was meant, colliding with the browser's own tab concept.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:21 -05:00
timkjrandClaude Sonnet 5 3da8bb7046 docs: add split-pane sessions design spec
Scopes v1 of an in-app split view (two live session panes side-by-side,
draggable divider) after multi-monitor spanning turned out to solve a
different problem than showing multiple panes at once.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:20 -05:00
Codeman maintainer 4205f6930f fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine.

**The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext`
and `confirmedSwap` changed the wire field without moving three assertions that check
it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts`
(the swap modal and the context modal, each of which already receives exactly the right
per-question flag). Moved, with the titles.

**Worse, my own tests for the split never ran.** The four cases in
`session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`,
which was declared inside a sibling `describe`, so they threw a ReferenceError during
setup. The split would have shipped with no passing server-side coverage while the gate
reported the failure as four broken tests rather than as four tests that were never
written. `mockRunning` is hoisted to the outer describe.

**The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved
its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so
the other eight modes fell back to claude's `❯`. That is also starship's default shell
prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line
`❯ npm run build` sits on screen for as long as the command runs, the verifier reads it
as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to
nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N
prompt, `read -p`, an installer or a pager, where it takes the default. The module's own
fileoverview already stated the rule this broke. Now `?? ''`, which
`promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI
that actually declares a composer.

**My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing
three numbered rules.

The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared
in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told
an integrator to retry with `confirmed: true` for both questions, which is precisely the
thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md`
and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR`
multi-user consequence, both user-visible and both previously absent, and #454's gained
the one exception to its own claim: a Custom Endpoints launch ignores the Instance count
stepper and always starts one session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:58:33 +02:00
Codeman maintainer 12de3c5164 docs(custom-model): record the CLAUDE_CONFIG_DIR clamp in architecture-invariants
CLAUDE.md gained the admin-only note when the key joined claude's privilegedEnvKeys;
architecture-invariants, which is where the exact-key allowlist rule is documented in
depth, still described the pre-change world. The reboot-restore half is the one worth
writing down: a non-granted owner's already-persisted CLAUDE_CONFIG_DIR is stripped on
restore, which moves that session back to the default Claude account with no error.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:35:32 +02:00
Codeman maintainer 9af12afb57 docs(custom-model): make the docs match the code, and trim the changeset
More from the review of 5fc391a4, all documentation rather than behaviour.

The changeset was 1602 words of development log, written as the PR grew, with bullets
and loose paragraphs interleaved. That text becomes CHANGELOG.md and the GitHub release
body verbatim, so it is now one user-facing account of what the feature does and what
the real-server work bought, at roughly a fifth the length.

docs/api-reference.md promised a `cmd` field on running-status that the route
deliberately strips (it carries model paths and can carry --api-key).

Two places claimed the apply routes validate `modelId` against the endpoint's
discovered models. Neither does. Dropped the claim rather than adding the check:
discovery can be up to five minutes stale, so a 400 there would refuse a launch that
actually works, and a typo'd id already fails on the CLI's own first request. CLAUDE.md
now says so explicitly, since the absence is the surprising part.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:34:24 +02:00
Codeman maintainer 1a99b5836c Merge pull request #430 from opticon454/custom-model-run-menu 2026-09-19 12:25:11 +02:00
Codeman maintainer 035bfbc2fe fix(remote): merge-time fixes for Wake-on-LAN
The MAC-count limit lived in two places that disagreed. RemoteHostSchema.wakeMac's
128-character cap admits seven comma-separated MACs while parseMacList takes at most
four, all-or-nothing, so a five-MAC value validated, was written to remote-hosts.json,
and then resolved to NO wake target: POST /api/sessions/:id/wake answered
"No wake-on-LAN target configured for this host" and the banner offered "Configure WoL"
for a host the user had just configured. MAX_WAKE_MACS now lives in
src/config/remote-wake-limits.ts and both sides refine against it. Its own module
because src/remote-wake.ts is import-fenced to session-routes.ts and server.ts (the
wiring guard that stops a watcher waking a host), and because schemas.ts must not drag
dgram/net/child_process into every request-validating module.

The documented 40 s request budget also omitted the wake's own cost. A `command` target
is bounded by REMOTE_WAKE_COMMAND_TIMEOUT_MS and runs BEFORE the readiness poll, so a
slow one pushed a wakeCommand host's worst case to ~68 s, past the 60 s
proxy_read_timeout the budget exists to stay under. _wakeAndWait now subtracts the
wake's measured elapsed time from the readiness budget, floored at one poll interval so
a wake that ate the whole budget still gets one probe. A magic packet is effectively
instant and is unaffected, which is why live testing never saw it.

Also: the two new endpoints are documented in docs/api-reference.md with the import
fence stated as the rule it is, CLAUDE.md's frontend module count moves to 34, and the
release changesets carry the Thanks section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:40 +02:00
Codeman maintainer 4c705094f7 fix(terminal): ship the copy clean as a trailing trim, without the shared dedent
#451 cleaned two things on copy. The trailing trim is right and every native
terminal does it. The shared leading-indent strip is this project's own rule,
and it is dropped here rather than shipped.

Measured against the shipped transform over 401,445 three-row windows across
1,010 tracked files in this repo, it fired on 73% of them: 92% inside a YAML
workflow, 76% over `git log` output, 48% in a TypeScript source. No width
threshold separates a margin from content because they are the same widths, a
live Claude Code pane's own margins measuring 2 and 5 columns while the most
common non-TUI shared run is 4. The failure modes are not symmetric either: a
wrong trailing trim costs nothing, while a wrong dedent silently deletes
information that was on the screen, with nothing in the clipboard to hint at
it, on git log bodies, on indented code read out of cat (semantic in Python),
on git diff context rows where the leading space is the marker, and on stack
traces.

It also could not be made self-consistent cheaply. Whether the first row joined
the measurement depended on the mousedown COLUMN, which the user never sees, so
one block of three rows produced three different clipboard results; and the
flag read getSelectionPosition().start, which is xterm's mousedown anchor and
is never normalised, so dragging UP through a block read it off the bottom row.
The PR's test stub hardcoded a downward drag, so its suite could not express
that case.

The transform, the wiring, the tests, the invariants, CLAUDE.md, the wiki page
and the changeset all move together. The test block now pins the ABSENCE as a
contract, with the git log, Python and git diff cases as its examples, so this
is not re-derived later. If it is ever revisited, the one qualification that
measured clean is painted trailing padding: zero false positives over all
401,445 windows.

Also from the review: the comments and invariant rule justifying the
padding-only clear described the pre-change code (the Ctrl+C gate reads the
CLEANED selection now, so such a selection falls through to the PTY on its own
and the clear is feedback rather than protection), the new 'Nothing to copy'
toast gained its zh-CN entry, and the invariants paragraph no longer repeats
its own opening sentence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:39 +02:00
Codeman maintainer c376534a50 fix(run,terminal): merge-time fixes for the Instance count stepper and capture geometry
#454: the behaviour the PR adds had no test, so a regression test drives
runGrok() at tabCount 3 and asserts three quick-start POSTs with sequential
w<n>-<case> names (verified to fail against master's session-ui.js). Each
caller now reads the count BEFORE its opening banner and announces it there,
the way runClaude() already did, so a launch no longer prints two headers and
a launch with another session already active still says how many are starting.
runClaude() calls the shared _readTabCount() instead of its own copy of the
1..20 clamp, and that helper optional-chains the element read, since hoisting
it above each caller's try block would otherwise let a missing #tabCount throw
where the launch-error path cannot report it.

#435: sizeMovedUnderLoad derived from data.source alone. `mux-visible` is not
sufficient: a failed display-message cursor query makes capturePaneBuffer skip
the snapshot repaint and return the raw capture, which the route still labels
mux-visible, so a size that moved during such a load bought a full forced
reload to repair a frame that was never positioned. It now tests
Number.isFinite(data.captureRows) like its two siblings.

Plus the invariants and CLAUDE.md lines promised on #435: a visible capture
reports its geometry and omits it when nothing was positioned, the comparison
runs on mux-visible only, and the replay is capped at one attempt and latches
per session when it cannot converge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:39 +02:00
Ark0N 2c3ccdf030 Merge pull request #439
feat(remote): wake a sleeping host (Wake-on-LAN) from input, banner and native magic packet
2026-09-19 12:18:18 +02:00
Ark0N 613b774bf1 Merge pull request #451
fix(terminal): trim the padding and shared indent out of a copied selection
2026-09-19 12:18:08 +02:00
RandalixandClaude Opus 5 5bb489addb fix(remote): authorize the attach wake first; tell the caller what happened to its bytes
Review round 3 on #439.

- The attachRemoteSession branch of POST /api/sessions ran `ensureHostAwake`
  before the multi-user gates, so a non-admin could have any configured
  host's `wakeCommand` spawned (or a packet broadcast) and the request held
  for the wake budget, then be refused for the workingDir. The admin gate
  now comes first, before the host is even looked up; remote hosts are
  admin-only infrastructure everywhere else. Route test: wake spy empty,
  403.
- The non-wait input route answers `{buffered:true}` when the registry took
  the chunk and `{buffered:true, dropped:true}` when it was over the cap
  and is gone (`RemoteInputOutcome` gains 'dropped'); additive to the bare
  `{}`.
- The send-and-wait path answers OPERATION_FAILED when the host never comes
  back, like create and attach, instead of writing into the stalled pane
  and reporting delivered:true plus a timeout.
- The flush writes with `fromUser: true`, so a first prompt buffered
  through a wake can still name the tab.

Docs: api-reference (input route), remote-sessions.md (two invariants),
CLAUDE.md key pattern.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-19 11:39:50 +02:00
RandalixandClaude Opus 5 1040f6c489 fix(remote): a proxied host is reachability-unknown; scope remote: SSE per session
Review round 2 on #439.

1. The bare TCP probe connects to host:port, which a host behind a jump host
   or SOCKS proxy does not answer even while ssh works. Acting on that
   verdict drew a permanent banner over a healthy session, replaced a real
   "needs tmux" error with "not reachable" in quick-start, and - with a wake
   target - buffered every HTTP input for the life of the session, since the
   readiness poll could never succeed. `WakeableRemote` now carries
   `jumpHost`/`socksProxy`/`extraSshOptions`, and `isProbeable()` turns such
   a host into reachability-UNKNOWN: input is delivered, `checkReachable` /
   `checkHostReachable` answer `null` (never `false`), `ensureHostAwake`
   returns `'unprobeable'` (handled like `'no-target'`), the quick-start gate
   fires on `=== false` only, and `GET …/reachability` reports
   `reachable: null, probeable: false` so the banner has nothing to key on.
   A wake target can still be fired for it, blind: no readiness poll, no
   reattach, no toast - the response says only whether the packet went out.

2. `'remote:'` joins the session-scoped SSE prefixes. The create/attach wake
   has no session yet, so the registry names the requesting user
   (`ensureHostAwake({ requestedBy })` -> `username` in the payload) and
   `deriveSseHint` routes on it; with neither it fails closed to admins.
   Single-user mode is unaffected.

Smaller, from the same review:

- A flush write that fails now drops the remaining buffer (logged) instead
  of retaining it: the wake still resolved and marked the host reachable, so
  the retained chunk waited for the NEXT wake and was replayed hours later,
  after everything typed since. Same policy as the oversized paste.
- The banner polls on tab activation (a user action) and on its 30 s timer
  only for a host with a wake target; a timer connecting to a host Codeman
  cannot wake is the traffic invariant #2 rejects keepalives for. A proxied
  host is never polled.
- `probeRemoteHostReachable`, `runRemoteWakeCommand` and the default UDP
  socket refuse under VITEST, as remote-files.ts does. The guard caught a
  leak on the spot: `createDefaultRemoteWakeDeps({ probe })` overrode the
  probe but still polled readiness with the real one, so the shutdown test
  had been connecting to a production address. The poll now uses the
  injected probe.
- docs/remote-sessions.md is additions only again (the reformatting is
  gone); the architecture-invariants overlap resolved itself in the merge.

Live, against a throwaway instance with a non-routable ghost host: proxied
-> no probe, no wake, the genuine ssh error after 10 s; direct (control) ->
probe, magic packet, "did not come back" after the 40 s budget.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-18 22:46:11 +02:00
RandalixandClaude Opus 5 e271a65e79 Merge origin/master into feat/remote-host-wake
Resolves CLAUDE.md count tables (route counts recounted on the merged
tree: 235 handlers, sessions 37) and keeps both the host-wake and the
reboot-restore banner in index.html.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-18 22:20:41 +02:00
Ark0N 492f8d8ddf Merge pull request #436 from irisitymichaelgrundberg/fix/replay-output-that-arrived-after-the-capture
fix(terminal): keep the output a pane capture could not contain
2026-09-18 21:34:42 +02:00
Devvyn 56209e7829 Merge remote-tracking branch 'upstream/master' into feature/run-menu-custom-model-picker 2026-09-19 03:25:26 +08:00
DevvynandClaude Sonnet 5 afb6754453 fix(custom-model): address third pre-merge review (Ark0N)
Blocker 1: the loading banner hides itself ~200ms after it reopens.

- _showCenterStatus reuses one shared DOM node; dismiss() scheduled
  el.hidden = true 200ms later with nothing to cancel it. On the
  Claude path, switchingToast.dismiss() is followed by one same-
  origin request (5-30ms locally) before _watchLlamaSwapLoading opens
  the new banner -- well inside that window -- so the stale timer
  fired against the shared node and hid the fresh banner, leaving the
  whole model-load wait with no progress text, no log line and no
  reachable Cancel button.
- Fixed by parking the pending timeout on the element and clearing it
  at the top of _showCenterStatus. Added a regression test that
  reproduces the exact repro (open, dismiss, reopen 20ms later,
  advance past 200ms) alongside the existing Cancel-button DOM tests;
  confirmed it fails without the fix and passes with it.

Blocker 2: the swap-conflict warning named other users' sessions.

- Both affectedSessions scans (POST .../custom-model and quick-start)
  walked the whole session map with no ownership filter, so in multi-
  user mode a non-admin pointing their own session at a shared
  endpoint learned another user's session name and id -- which with
  autoNameSessions on is that user's own prompt.
- The swap is still blocked pending confirmation regardless of
  ownership (a foreign session is just as real a disruption); only
  which ones get NAMED back to the caller is scoped, via the
  already-imported canAccessOwned. Added a two-owner test to
  test/routes/session-custom-model.test.ts covering both the
  foreign-owner (blocked, not named) and same-owner (named) cases.

Smaller ride-along fixes:

- server.ts boot recovery now passes contextLength into
  applyCustomModelInjection, so CLAUDE_CODE_MAX_CONTEXT_TOKENS is
  correctly rebuilt into _envOverrides after a restart instead of
  surviving only because tmux retains the old setenv.
- pumpLlamaSwapLogTail's finally now deletes by IDENTITY, not just by
  key, so an aborted pump finishing after a newer entry was created
  for the same endpoint can no longer delete that newer entry and
  orphan its connection.
- docs/custom-model-endpoints.md now notes that clearing a custom
  model removes injected keys by name, including CLAUDE_CONFIG_DIR --
  so a session that also had CLAUDE_CONFIG_DIR set via envOverrides
  (the per-client-account case) silently falls back to the default
  account on clear.

Left for later, as flagged in the review itself: the quick-start
case-scaffolding/cancel ordering (real behavioural reordering across
a large handler, too risky to make without a live re-test), and
retiring runCustomModelEntry's mode === 'claude' branch behind a
launchStrategy registry field (explicitly deferred by the reviewer to
"the next one").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-19 03:09:36 +08:00
Michael GrundbergandClaude Opus 5 f9edb33d15 fix(terminal): trim the padding and shared indent out of a copied selection
xterm hands back whole screen rows and trims only the cells that were
never written to, so the real spaces a full-screen TUI paints across the
unused part of a row count as content and reach the clipboard. Measured
against Claude Code in a 282-column pane, single lines arrived carrying
138 trailing spaces, and every line carried the two-space transcript
indent as well. Windows Terminal, iTerm2 and GNOME Terminal all trim that
for you, decideAutoCopy already calls a wall of spaces "never what the
gesture meant", and _selectTouchSelectionLine already treats those cells
as padding — the mouse and keyboard paths never had the same rule.

CodemanCopySelection.clean lives in constants.js beside decideAutoCopy,
its pure sibling. It drops the trailing run from each line, and removes
the leading run only where every selected row shares one. A selection of
a single row keeps its run, because one row shares nothing with anything
and stripping it would silently reindent one line of `git log` body text
or one line out of `less`. A drag that began inside a row keeps its
partial first line untouched and out of the measurement, which otherwise
pins the shared run to zero and leaves every following row indented.

Every pass over a line is a scan rather than a regex. `/[ \t]+(\r?)$/` is
quadratic on a line whose spaces are followed by a non-space character,
which is what right-aligned or centred TUI content looks like: measured
over 50 000 rows with a 280-column run it took 2.9s, against 1.3ms for
the scan, and a 2 000-column run took 16s. The scan is also the faster of
the two on an ordinary padded row.

cleanedTerminalSelection in terminal-ui.js is the half that needs the
live terminal. It returns a COLUMN selection untouched: Alt+drag makes
one, and a rectangle's rows lining up is the point of the gesture, so
both halves of the clean would destroy it. xterm exposes the mode nowhere
public, so the check reads terminal._core._selectionService, the way this
file already reads terminal._core for cell dimensions, and cleans
normally if a future xterm renames the field. A test pins that assumption
against the library rather than against a stub repeating the literal.

The Ctrl+C chord decides on the cleaned selection, not the raw one. A
drag across the blank part of a row selects real padding spaces, so the
raw text is truthy, and testing it would spend that press on a copy of
nothing and make the user press again to interrupt. A padding-only
selection is now dropped and the press falls through to the PTY, while
Ctrl+Shift+C still never falls through. copyTerminalSelection gates on
trim() for the same reason, since a multi-row drag across padding cleans
to line breaks alone and a bare newline pasted into a chat composer
submits it.

All four of the main terminal's copy paths go through it: the Ctrl+C
chord, right-click, the phone selection button and Auto Copy. The
browser's own Edit menu copy, a disabled copy shortcut and the subagent
windows still copy raw rows, as they did before, and the invariants doc
now says so rather than claiming every copy is cleaned. Auto Copy
resolves its own toggle before it reads the selection, since it is off by
default and a selection can run to the 50 000-row scrollback ceiling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 20:18:25 +02:00
Michael GrundbergandClaude Opus 5 75a028e825 fix(terminal): re-take the sticky-scroll baseline after a replay
A capture load now replays its queued tail, and that replay runs through
`batchTerminalWrite`, which samples `_wasAtBottomBeforeWrite` before it queues.
It runs inside `chunkedTerminalWrite`, before that promise resolves, with the
terminal freshly reset and rewritten — so the sample is always true. The caller
then restored the reader's position and the next `flushPendingWrites` scrolled
straight back to the bottom off the latched flag, undoing it. The only thing in
the way was `_hasRecentUserScrollUp()`, a 1500ms window a server-triggered
refresh is usually past.

`_syncStickyScrollBaseline()` re-takes the flag from wherever the viewport now
sits, and the two paths that restore a position call it right after doing so:
`_onSessionNeedsRefresh` and `_maybeRefetchFullHistory`. Those are the paths
#259 and #205 exist for, and they are also where a non-empty queue is most
likely, since a needsRefresh fires when output is flooding. Re-taking rather
than suppressing the sampling: suppressing leaves whatever stale value the flag
held from before the load, which on the full-history re-pull has no reason to
be false. `selectSession` and `_onSessionClearTerminal` deliberately end at the
bottom, so the sampled true is already the truth there and they do not call it.

`_bufferLoadFinishOpts` gains the coverage the CI gate can see: both mux
sources flush, `history` does not, and a payload naming no source does not.
Its only coverage was the browser suite, which CI does not run.

The JSDoc and the changeset now record the one duplicate window this cutoff
cannot close. The server appends output to the byte buffer in the same tick it
emits, but broadcasts on a batch timer — 8ms over WebSocket, 16 to 50ms over
SSE — so a batch pending when `capture-pane` ran leaves the server after the
reply and is replayed although the capture holds it. It is one batch interval
wide against a recovery window spanning the whole chunked write, and closing it
means flushing that batch server side before the capture.

The second browser test asserts its session was created, so a failed create
fails it instead of passing with zero hits.

docs/architecture-invariants.md no longer claims the replay leaves the
queued-event discard window alone. That clause now describes what decides how a
load ends, the baseline rule, the batch window, and the three covering tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:43:04 +02:00
DevvynandClaude Sonnet 5 9982a1325f fix(custom-model): address second pre-merge review (Ark0N)
Blocker: .center-status-banner never actually disappears.

- Add `.center-status-banner[hidden] { display: none; }`, same trap as
  `.home-sessions[hidden]`: the author-level `display: flex` beat the
  UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
  card stayed laid out at `opacity: 0` with its text/cancel/close
  children still `pointer-events: auto` -- an invisible 442x67 click
  blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
  banner (10001) and the swap-confirm/context-warning modals (10010)
  in CLAUDE.md's Z-index layers list.

Stale wording pointed at the reverted sticky-toast default:

- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
  `.toast-message` comment in styles.css all still said "toasts
  default to sticky" after 1f32128c put the flat 3s default back.
  Reworded all three to describe the actual behaviour: one call site
  passes an explicit `duration: 0`.

Smaller items from the same review:

- docs/api-reference.md said discovery failures answer
  `502 OPERATION_FAILED`; OPERATION_FAILED is 422 per src/types/api.ts
  and the error-code table earlier in the same file.
- The periodic re-discovery sweep (server.ts) never read
  customModelEndpointsEnabled, so turning the feature off left
  Codeman polling every saved endpoint forever. Added
  readCustomModelEndpointsEnabled() (custom-model-routes.ts, same
  shape as readPlanUsageTelemetryEnabled) and gated the interval
  callback on it.
- Reverted the formatting-only Prettier pass docs/api-reference.md
  picked up (table padding, *x* to _x_, JSON re-indent) by re-merging
  the new Custom Model Endpoints section onto the pre-PR file, so the
  diff is reviewable. No prose content was lost -- verified by diffing
  the result against the pre-revert file (formatting-only) and against
  the merge-base file (only the new section added).
- docs/custom-model-endpoints.md now states that a custom-model Claude
  session's isolated CLAUDE_CONFIG_DIR loses the user's global
  settings.json, user-level skills/agents/commands, and MCP servers
  from ~/.claude.json -- only `projects` is symlinked back.

Design question left open in the review (does `confirmed: true` need
to be two flags so "launch anyway" on the context warning doesn't also
skip the llama-swap displacement warning): keeping the single flag, as
offered. The 20s displacement sweep still catches a resulting swap
after the fact, so it's a surprise rather than a silent failure, and
splitting it is real behavioural surface I have no way to verify live
in this environment.

`npm run test:browser` could not be run in this environment (no tmux,
no downloaded Playwright browser binary) -- none of its suite's files
touch code this fix changes, but it still needs a real pass before
merge, same as any frontend change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-18 21:45:50 +08:00
Codeman maintainer bb8ada7e5f fix(reboot-restore): the merge-time items from the #442 review
Seven things, none of which changes what the feature does.

1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred
   it from the name: a session the user renamed by hand to something shaped
   like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the
   next prompt overwrote their name. The route persists right after, so the
   loss went to disk. `restoreMuxSessions()` already passes it.

2. The already-live sets were snapshotted once before a loop that awaits a
   real `startInteractive()` per entry, so by the tenth entry the snapshot
   was tens of seconds old and a conversation resumed by hand from the
   Resume list in that window was invisible to it: two panes on one
   transcript, the exact thing the check exists to prevent. Both sets are
   now read per iteration, and the late case is spent rather than re-offered
   for the same reason the batch case is.

3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path.
   The stamp predates the reboot and the pane is new, so honouring it meant
   one click had every restored session type `continue` into itself about a
   minute later, unattended, against the route header's own promise that a
   restored session comes back idle and disarmed. The setting stays ENABLED,
   so it re-arms on the next real limit message. A Codeman restart still
   re-arms from the stamp, because the limit footer will not reprint on its
   own; the new option exists only to tell the two paths apart.

4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()`
   and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession`
   performs that it was missing. Cosmetic, but a run left open reads as
   still going in the away digest.

5. A restored claude session gets `seedAgentSessionPreamble()` like both
   create paths, so the agent skill's bootstrap stays a two-line loader.

6. The heuristic's container comment was wrong in one direction and quiet
   about the real gap: after a genuine host reboot a containerized Codeman
   sees the host's short uptime and the banner does appear. What it cannot
   see is a container-only restart, which is where this would help most.

7. The banner is hidden in a solo window, which shows one session and has
   no tab strip to put restored ones in.

Also reverts 17 of the 18 hunks in docs/api-reference.md, which were
Prettier reformatting of prose the PR does not otherwise touch (docs/ is
outside the format glob), keeping only the Reboot restore section and
repairing the two continuation lines that reformat de-indented; renumbers
reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which
loads after it; and gives the feature its CLAUDE.md entry plus a route
test for the multi-user workspace-forbidden branch, the only new rule that
had nothing behind it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 13:46:04 +02:00
DevvynandClaude Sonnet 5 e2034177c5 fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1)
DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.

Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
  published dependencies) into a scratch dir purely to read
  dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
  completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
  the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
  completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
  same endpoint. dsh's own error template ("DeepSeek API error (HTTP
  ${status})") reproduces the originally-reported
  "dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.

- New registry field `appendV1Suffix` (env kind only, deepseek's entry
  alone — claude/gemini must NOT get it, since claude was already
  confirmed working against the unmodified baseUrl). When set,
  buildCustomModelInjection runs endpoint.baseUrl through the same
  withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
  use, instead of writing it verbatim.

Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.

2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:41:19 +08:00
DevvynandClaude Sonnet 5 8520925e76 docs(custom-model): bring CLAUDE.md and api-reference.md up to date
Full documentation review pass across the branch's 30 commits.
CLAUDE.md's Custom Model Endpoint Profiles entry hadn't been touched
since the initial backend+picker cut (3 early commits) despite 27
follow-up commits adding real behavior — it described restart-in-place
as universal (now claude-only; 7 other CLIs launch one-shot) and
claimed codex's Responses-API gap as a flat protocol break (now
re-verified as a more precise tool-calling gap). Corrected both and
added a new paragraph covering everything landed since: the llama-swap
conflict check, the after-the-fact swap-displacement sweep, the
/running-cmd-based context-length fix, the context-window floor
warning, skipFirstRunPrompts, the real-time /api/events-based log
status, and the countdown-to-Cancel-button change.

docs/api-reference.md's custom-model-endpoints section was missing the
running-status route, the requiresConfirmation/requiresContextWarning
response shapes, and POST /api/quick-start's customModel field
entirely (the primary launch path for 7 of 8 supported CLIs) — added
all three. Also fixed a real markdown bug in custom-model-endpoints.md:
an inline code span (`POST <baseUrl>/v1/chat/completions`) split across
a line break, which CommonMark renders with the line ending collapsed
to a space, so it displayed as ".../v1/chat/ completions" with a
spurious space inside the path.

Verified: origin/master and upstream/master are both already an
ancestor of this branch (identical at bd286bf5, no new commits since
this branch was cut) — nothing to merge, no conflicts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:18:04 +08:00
DevvynandClaude Sonnet 5 db9729e1fc feat(custom-model): remove loading-banner countdown, add manual Cancel
Replaces the size-scaled expected-time estimate + matching auto-timeout
with a generic hardware/model-size disclaimer and a user-driven Cancel
button, per explicit request. Real load time depends on hardware this
feature has no way to know (VRAM, storage speed, GPU contention), so
the old estimate/timeout was a guess dressed up as a fact — worse, one
that could kill a genuinely slow load partway through on slower
hardware.

- _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline
  entirely — polls indefinitely until ready or cancelled, no automatic
  give-up. Message is now "Loading <model> (<size>) on <endpoint> —
  this can take a while depending on your hardware and the model
  size.", with the real llama.cpp log line still on its own second
  line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/
  _formatRemaining (dead code once the countdown is gone) —
  _lookupModelSizeGB is kept, the GB figure still shows.
- _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real
  "Cancel" button (distinct from the error-type "×" close button,
  since Cancel has a real consequence) that calls it on click. Caller
  owns what cancelling actually means, same split as the swap-confirm
  modal's promise-resolving buttons.
- Cancelling dismisses the banner, shows an info toast (not an error —
  this was deliberate), and closes the session, mirroring what the old
  timeout used to do automatically but now on the user's own call.
- New .center-status-cancel CSS (bordered pill button, distinct from
  the plain "×" close glyph).

Test changes: removed the now-invalid timeout-auto-close/estimate
tests, added cancel-flow tests (dismiss/toast-type/session-close,
never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests
for the new Cancel button (bootAppWithRealCenterStatus, evaluating
panels-ui.js instead of stubbing _showCenterStatus, since this button
is worth verifying for real rather than just through the stub every
other test in the file uses). Typecheck/lint/frontend-syntax clean;
full suite shows no new regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:00:50 +08:00
DevvynandClaude Sonnet 5 2d3fc65758 feat(custom-model): show real-time llama.cpp backend status in the loading banner
Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.

- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
  (custom-model-routes.ts): one persistent GET /api/events (SSE)
  connection held open per endpoint, parsing logData frames and
  keeping the latest source:"upstream" (backend llama-server) line —
  filtering out llama-swap's own source:"proxy" request-access lines.
  Idle-closed after 30s of no polling, same 20s sweep as the existing
  swap-displacement check.
- running-status route now returns logLine alongside the existing
  isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
  ("llama.cpp: <line>", bootlog timestamp/level/component prefix
  stripped for display) that stays on the last real thing llama.cpp
  said rather than clearing to blank between polls.

⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).

12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 12:22:45 +08:00
DevvynandClaude Sonnet 5 5ddc028a2f feat(custom-model): detect and notify when a session's model gets swapped out later
The llama-swap conflict check on the apply/create routes only ever runs
at THAT session's own launch/apply moment, and cannot see a swap caused
by a DIFFERENT session's later, ordinary use. Confirmed live: a second
Codex session picking a different model launched with no warning at
all — nothing conflicted at that exact instant — yet it silently
evicted the first session's model regardless (llama.cpp runs one model
at a time). Reproduced and root-caused via direct API calls against a
live test-picker instance rather than guessing.

- detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups
  live sessions with a customModel by endpointId, checks each group's
  endpoint via GET /running once, and flags a session whose own modelId
  is no longer in the running list. Read-only, best-effort per endpoint
  like refreshAllCustomModelHosts's sibling sweep.
- Notifies once per displacement via a caller-owned de-dupe Set: a
  session id is added when displaced, removed once its own model is
  loaded/ready again, so a later genuinely-new displacement can notify
  again.
- New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS,
  20s — much shorter than the 5-minute model-list refresh, since this
  is time-sensitive) broadcasts a new custom-model:swapped-out SSE
  event per displacement. De-dupe Set cleared per-session on session
  cleanup to avoid an unbounded leak.
- Frontend: global toast (not tied to the displaced session's tab,
  since the point is warning before the user types into it) naming the
  session, its previous model, and what's currently loaded.

Chose the "detect after the fact" scope (vs. checking before every
message send, which would add a round-trip to every turn on every
custom-model session) per explicit user decision after being presented
the trade-off.

9 new tests for the detection logic (flag/clear/re-flag cycle,
unreachable/deleted endpoints, non-llama-swap servers, multiple
sessions on one endpoint). SSE registry bumped 158->159, parity test
passing. Typecheck/lint/frontend-syntax clean; full suite shows no new
regressions (9 more passing than baseline, matching the new tests;
same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 11:09:24 +08:00
DevvynandClaude Sonnet 5 470f75b08c docs(custom-model): record live findings on codex's model-metadata warning
Investigated the user's report of "Model metadata for <id> not found.
Defaulting to fallback metadata..." on every custom-endpoint codex
launch, live against the test-picker's llama-swap deployment (codex
0.152.1):

- The warning is cosmetic. `codex exec 'reply with just OK'` against the
  isolated CODEX_HOME still printed the warning and still returned a
  real reply.
- The isolated CODEX_HOME never gets a models_cache.json written into
  it at all, even after extended real use (inspected a live, actively-
  used directory) — codex can't reach OpenAI's own hosted model catalog
  for this session and silently falls back every time, with no local
  file to create or clean up. There is also no config.toml override for
  a model's metadata.
- Fabricating a fake catalog entry to suppress it would mean copying the
  SHAPE of OpenAI's own proprietary models_cache.json schema, including
  real per-model system-prompt content visible in a genuine entry — not
  something to build for a warning confirmed to have no effect.
- More importantly: a real tool-call attempt against the same setup came
  back as agent_message TEXT (the tool-call JSON printed as the answer)
  rather than an executable function_call item, confirmed via
  `codex exec --json`'s raw event stream. Tool execution is what makes
  codex a coding agent, so it remains not usable for real work regardless
  of the metadata warning — a more precise, re-verified update to the
  existing "Responses API protocol gap" finding (which reported a harder
  Reconnecting/high-demand failure on a different llama-swap deployment;
  this one answers /v1/responses for plain chat but still can't execute
  tools).

No code changes — recipe/comment/confidence-table documentation only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 10:33:48 +08:00
DevvynandClaude Sonnet 5 211b872335 feat(custom-model): skip Claude Code's first-run wizard on custom-model launches
A fresh, isolated CLAUDE_CONFIG_DIR (used to keep an injected API key
away from a stored claude.ai OAuth login) looks like a brand-new Claude
Code profile to the CLI, so it replays its ENTIRE first-run sequence on
every single launch: the theme picker, the security-notes screen, the
per-project "trust this folder?" dialog, and (running with
--dangerously-skip-permissions) a one-time bypass-permissions warning —
confirmed live, none of which a real, already-onboarded profile shows
again.

- New registry-declared env-kind field `skipFirstRunPrompts` (alongside
  apiKeyTrustFile, which it reuses) — claude's entry only, carried
  through buildCustomModelInjection (pure) into
  applyCustomModelInjection (IO).
- seedFirstRunOnboardingState(): merges hasCompletedOnboarding: true and
  this session's own projects[workingDir].hasTrustDialogAccepted: true
  into the same <configDir>/.claude.json the API-key trust file already
  writes to — other projects and other fields on this session's own
  entry are left untouched.
- seedSkipBypassPermissionsPrompt(): merges
  skipDangerousModePermissionPrompt: true into <configDir>/settings.json,
  a separate file, same corrupt-tolerant merge behavior.
- applyCustomModelInjection() gains an optional workingDir parameter,
  threaded from session.workingDir (dedicated apply route) /
  resolvedCasePath (quick-start route) — boot recovery omits it
  (a dialog already answered once needs no re-seed on the same,
  persisted isolated directory).

Tests added at the pure-builder, IO-wrapper (including merge-preserves-
other-fields and corrupt-file-tolerance cases), and existing directory-
listing assertions updated for the new settings.json file. Typecheck/
lint/format clean; full suite shows no new regressions (baseline
pre-existing Windows-environment failures unchanged, 8 more passing
tests than before — the ones added here).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 09:21:36 +08:00
DevvynandClaude Sonnet 5 b45a96358e feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.

- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
  registry entry declares contextLengthVar (currently only claude) and
  the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
  (40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
  quick-start customModel path) check this before the swap-conflict
  check and before launching/restarting anything, returning
  {requiresContextWarning, modelId, contextLength, minSafeContextTokens}
  — skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
  _resolveContextWarningConfirm (session-ui.js), wired into both
  _quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
  (the path Claude actually uses) ahead of the swap-confirmation check.
  Explains the fix in-modal: give the model an explicit larger -c/
  --ctx-size in llama-swap instead of relying on --fit-ctx, which
  optimizes for the biggest model that fits rather than the biggest
  context.

Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 07:36:57 +08:00
Randalix acb8d4b0aa docs(remote): correct what the wake PR moved
- `host-wake-ui.js` joins the documented load order (12.2) and gets its
  `@dependency`/`@loadorder` tags; the frontend module count is 33, not 32.
- `remote-wake` is not "(pure)" — the module uses `dgram`/`net`/`child_process`.
- SSE counts: 160 constants, and the category is "Remote auto-reconnect / wake
  (5)"; the route table's per-file counts are refreshed (sessions 37, cases 34).
- The CLAUDE.md wake rule now names the create/attach wake, the 40 s request
  budget, the whole-chunk paste drop, the registry's lifetime (drop on cleanup,
  stop on shutdown) and the deliberately non-wake-aware WebSocket keystroke
  path — that paragraph is what the next person reads.
- Reverted the eight lines of unrelated Prettier markdown churn in
  `docs/architecture-invariants.md` (docs/ is not in the format glob, so it was
  an editor): only the new wake paragraph remains in the diff.
2026-09-16 20:44:48 +02:00
Randalix 7b947fa3f1 fix(remote): close the wake-state leaks and the dishonest wake budget
Review follow-up on the wake-on-LAN PR (five findings, all of them about the
state the feature keeps and the budgets it inherits):

- Wake state is dropped by `WebServer.cleanupSession` instead of the two delete
  routes, so it now goes with the session on EVERY cleanup path (cron, admin,
  scheduled-run teardown, error paths) instead of surviving with up to 4 KB of
  the user's buffered keystrokes. `registerSessionRoutes` returns the registry
  so the server can own its lifetime without the wake-capable code living in
  `server.ts`; the wiring guard is updated to allow that and gains a second
  assertion that `server.ts` calls nothing but `drop`/`stop` on it.
- `_effectiveRemote` returns before `_state`, so a LOCAL session no longer gets
  a wake-state entry — the input gate runs on every keystroke, so that entry
  used to be allocated for every session the user types in.
- An input chunk larger than the 4 KB cap is dropped OUTRIGHT instead of being
  head-trimmed and then written as a fragment: one paste is one `input` value
  and was never typed character by character, so its tail is a partial command
  the user never sent. The drop is logged.
- The manual wake button passes `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS` (40 s)
  like the create/attach paths, instead of inheriting the 90 s session default
  that the dashboard's reverse proxy cuts off at 60 s.
- `RemoteWakeRegistry.stop()` aborts in-flight readiness polls (abortable
  sleep) and refuses new wakes, and `WebServer.stop()` calls it, so a restart
  during a wake no longer waits the poll out.
- The banner/toast wording keys off a new `queuedInput` flag on the two SSE
  events, which is true only when the server actually holds bytes: browser
  keystrokes travel over the WebSocket, which never passes through the
  registry, so the wake BUTTON must not promise queued input. The failed-wake
  path also stops pattern-matching the error message (it re-asks the
  reachability route) and the WoL dialog says "admin-only" instead of "host not
  found" for a non-admin in multi-user mode.
2026-09-16 20:44:39 +02:00
Michael GrundbergandClaude Opus 5 71ed7b127c fix(sessions): make the discard a real inverse of the construction
Third review of the reboot-restore branch. The narrow discard the previous
commit introduced avoided everything cleanupSession() did wrongly, and in
dropping so much of it also dropped four things it had to keep.

The worst broke the retry the whole design rests on. setupSessionListeners()
returns early while sessionListenerRefs still holds the session id, and the
discard never cleared that entry. So the advertised flow — a rebuild fails
because the agent binary is missing, the user fixes their PATH and clicks
again — reused the same id, wired no listeners at all, and produced a tab
that never showed output, never updated its status and never persisted. That
is worse than the leak the discard was added to prevent. Three more
registrations leaked with it: a RunSummaryTracker and its interval, an image
watcher on the workspace, and the Ralph fix-plan watcher. The discard now
undoes each registration setupSessionListeners() makes, in its order, and
the per-session custom-model config directory, which holds the endpoint's
API key literally and which nothing else would ever remove.

The image-watcher flag was restored after the code that reads it, so a
session came back reporting the feature as on with nothing watching. It
moves to the before-spawn phase, and that phase now runs before the
listeners rather than after them.

The generation counter that lets a mid-restore dismiss win was global while
clear() is ownership-scoped, so one user's dismiss discarded another user's
unspent entries, permanently, because nothing rebuilds an in-memory plan. It
is now per owner. Bumping only the owners of entries the dismiss removed was
not enough either: take() has already emptied the plan by then, so a dismiss
landing mid-restore saw nothing of that owner's to remove and invalidated
nothing. The owners that matter are those with a restore in flight, filtered
by what the dismissing user may access, and that is what clear() now bumps.
Plan expiry bumps too, so a restore straddling the 24-hour boundary cannot
hand entries back and give an expired plan another full day.

Tests. discardPartiallyBuiltSession had no test at all: the only
implementation any test ran was the mock's one-line stub, which is why every
defect above was invisible. test/discard-partially-built-session.ts drives
the real WebServer, and the retry assertion fails if the listener refs are
left behind — verified by reverting the fix. The dismiss-race test drove the
registry by hand, so deleting the route's generation argument left it green;
it now goes through the route, and two further tests cover the multi-user
cases.

The mock context has now gone stale twice, because route tests pass it as
`ctx as never` and tsconfig.json includes only src, so nothing ever compares
it to the ports. A type-level guard is therefore inert — I wrote one and
confirmed it never fires. test/mocks/mock-route-context-completeness.ts
compares the mock's keys against WebServer.createRouteContext() at runtime
instead, and names what is missing.

Also: the API reference now says workspace-forbidden is judged against the
owner's grant, the banner's module header no longer claims Restore always
dismisses it, and the detail span gets the same min-width: 0 the phone rule
already needed.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 15:34:47 +02:00
DevvynandClaude Sonnet 5 993710263d fix(custom-model): stop trusting /props's n_ctx, parse the real context size from /running's cmd
Root cause of the context-overflow regression reported live: "API Error: 400
request (36437 tokens) exceeds the available context size (16384 tokens)".
Discovery had stored modelContextLengths.qwen3.8-27b-ud-q4_k_xl = 154112,
so CLAUDE_CODE_MAX_CONTEXT_TOKENS told Claude Code it had a huge window and
it never compacted - but the real llama-swap server was launched with
--fit-ctx 16384 (confirmed against /running's own cmd field) and refused
the request right at that real limit.

/props?model=<id>'s n_ctx (the field discovery read) is confirmed live to
be unreliable for a --fit-ctx-launched backend: it reported 154112 for the
same model /running says was launched with --fit-ctx 16384 - appears to
report the model's theoretical/trained maximum context, not the runtime-
configured one.

discoverModels() now parses the REAL configured size straight out of
llama-swap's own launch command instead (parseCtxFromCmd(), reading
/running's cmd field - --fit-ctx first, then the plain llama.cpp -c/
--ctx-size a hand-written command might use), and only falls back to the
old /props probe when cmd states no recognizable flag at all. One /running
call now covers every loaded model's context length in a single request,
same as it already did for the swap-conflict check and the load trigger.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 20:47:12 +08:00
DevvynandClaude Sonnet 5 7bbe408e44 feat(custom-model): live countdown on the loading banner; timeout is now an error
The loading banner now shows a live countdown against its own timeout
(updated every poll, so every second by default) instead of a static
"this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically
~1-3 min) on llama-swap - 47s remaining".

If the countdown reaches zero and the model still isn't ready, this is now
treated as a real failure rather than a "keep waiting" shrug:
- The banner turns into a sticky error (_showCenterStatus gains a `type`
  option - 'error' drops the spinner and adds a close button, since nothing
  is "in progress" anymore and a sticky message needs a way to dismiss it),
  naming the llama-swap server's own logs as where to look for detail.
- The session that load was for is closed automatically (closeSession) -
  requested explicitly: a console left open and pointed at a model that
  never finished loading is worse than no console at all. Both apply paths
  now thread the new session's id through to _watchLlamaSwapLoading for
  this (new required 3rd parameter, after endpointId/modelId).

_watchLlamaSwapGeneration's existing stale-call guard extends naturally to
this: a superseded call's own eventual timeout recognises it no longer owns
the banner and neither shows the error nor closes a session that may by
then belong to a different, newer launch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 20:22:02 +08:00
DevvynandClaude Sonnet 5 55dae31530 feat(custom-model): estimate model load time from its discovered size
Discovery now also parses a GB figure out of an auto-discovered model's own
description (llama-swap writes "Auto-discovered 16.35 GB - parameters
auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike
context length this needs no /props probe (the figure is right there in
/v1/models) so it is populated for every model regardless of loaded state.
A hand-configured profile's own description has no such figure and
correctly gets no entry.

The loading banner (_watchLlamaSwapLoading) now looks this up and, when
known, shows it plus a rough estimate from a small size->time matrix
(_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) -
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on
llama-swap... this can take a while" - and uses that same estimate's own
bracket to scale the banner's default give-up timeout for a very large
model, instead of a flat 5 minutes for everything. Explicitly labelled as
an UNMEASURED, typical-hardware estimate in every relevant comment - this
is not benchmarked against any real endpoint's actual storage/GPU, just a
reasonable expectation-setter. A model with no discoverable size (a
hand-configured profile) gets no size/estimate shown at all, matching the
"never a guess" convention modelContextLengths already established.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 19:58:46 +08:00
Michael GrundbergandClaude Opus 5 fbede5cd2a fix(sessions): act on the dual review of the reboot-restore route
Fifteen findings from two independent reviews of #442, three of them
blocking. Every one is addressed here.

The three blockers all sat in the restore route. A rebuild that threw after
addSession left a registered session with no pane behind it, visible on the
board, holding a layout slot and written to state.json, with its plan entry
already spent; the catch now cleans the session up and puts the entry back.
The loop checked neither the global nor the per-user session cap, so one
click could take a board past a documented limit; capacity is now re-checked
per iteration, because the loop is itself creating the sessions it counts.
Worst of the three, a rebuilt session carried none of the state its
constructor has no parameter for and then persisted itself over the record
that held it, zeroing token and cost totals and dropping the pin. The pin
matters most: pruning keeps a record only while it is pinned, so discarding
it handed the record to the next stale sweep. A new
reapplyPersistedSessionState() on the session port restores the pin, the
token totals, auto-compact, auto-clear, auto-resume, nice priority, the
flicker filter and the custom-model selection, and it runs before both
startInteractive and the first persist.

The rest, in the order they bite a user. Every rebuild failure was reported
as workspace-missing, so the banner told users their repo was gone when the
agent had simply failed to start; there are now distinct reasons, and the
toast names each one. The client read restored and skipped off the outer
response object rather than through the uniform envelope, so every count
came back zero and neither toast ever fired. A board left open across the
reboot never learned an offer existed, because the banner was seeded only on
the page-load path; it now re-reads on every SSE init. The workspace check
was existence-only, skipping the multi-user confinement that the create
route applies, so a withdrawn grant would not be noticed. The banner had no
phone breakpoint while its text was nowrap and its buttons could not shrink.

Smaller: a missing workspace is now re-offered rather than dropped, while an
already-open conversation is dropped rather than re-offered forever; a throw
anywhere in the route returns the unspent entries instead of discarding the
plan; the single flight is keyed by owner, since take() already stops two
callers receiving one entry; the env clamp's header no longer claims a
protection it cannot provide on this path today, and names the check that
does bite; the three endpoints are documented in docs/api-reference.md; and
the module header now says that os.uptime() reads the host's clock, so the
feature is effectively off inside a container.

The review also explained why the tests missed all of this: they proved the
construction claim through their own copy of the construction rather than
through the route, and the route tests used workspaces that did not exist,
so no Session was ever built. test/routes/reboot-restore-rebuild-failure.ts
mocks the Session module to drive the route's real path, and covers the
cleanup, the reason reported, the re-application ordering, the broadcast and
the caps. The mock route context gains the port method and the mux call the
route needs.

Refs #411

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 11:44:37 +02:00
DevvynandClaude Sonnet 5 0929694012 fix(custom-model): actually trigger the llama-swap load, not just watch for it
Root cause of "it doesn't look like llama-swap is actually switching the
model" (confirmed live: no load_model line in llama-swap's own logs after
applying a selection). llama-swap has no "switch model" admin endpoint - the
ONLY thing that starts a swap is a real inference request naming the model.
Every previous fix (the conflict check, the loading banner) assumed a swap
would start on its own; nothing ever actually asked llama-swap to load
anything until the launched CLI's first real prompt did, which could be
much later than "applying the selection" implied.

Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest
real request that will start a load - POST <baseUrl>/v1/chat/completions,
max_tokens: 1, one throwaway message - fire-and-forget (never awaited by
the caller; the frontend's own running-status polling is what actually
confirms readiness). Wired into both apply paths (the dedicated restart
route and the one-shot quick-start route), fired whenever the target model
isn't already the one loaded and ready - a broader condition than the
existing swapNeeded (which only gates the "this will evict another
session's model" confirmation ask and deliberately stays narrow to that).
modelSwapInProgress in both routes' responses now reflects this same
broader condition too, so the frontend's loading banner actually correlates
with a real in-flight load rather than only firing when something else
happened to be loaded already.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:53:34 +08:00
DevvynandClaude Sonnet 5 f865f74a0f feat(custom-model): launch directly on the endpoint, no restart, for 7 of 8 CLIs
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.

POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.

Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.

Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:04:10 +08:00
DevvynandClaude Sonnet 5 97464bfa27 fix(custom-model): pre-approve the injected API key in the isolated Claude config dir
The CLAUDE_CONFIG_DIR isolation from the previous commit fixed the cosmetic
auth warning but introduced a real regression: an otherwise-empty config
directory has none of a real profile's prior custom-API-key approvals, so
Claude Code stops at an interactive 'Detected a custom API key - use it?'
prompt on every single launch. Confirmed live. With nobody at a TTY to
answer, the prompt's own default ('No') silently refuses the very key this
feature just injected, which looks like the endpoint being ignored.

Adds apiKeyTrustFile to the env-kind customModelInjection capability shape
({relPath, shape: 'claude-api-key-responses'}), set on claude's entry to
{relPath: '.claude.json', shape: 'claude-api-key-responses'}. The apply step
merges customApiKeyResponses.approved: [apiKey] into
<isolatedConfigDir>/.claude.json - the exact field a real answered prompt
itself writes to (confirmed against a real ~/.claude.json after answering by
hand once), so this answers the prompt in advance rather than bypassing it.
Merges onto whatever the CLI already wrote into that file on an earlier
launch in the same isolated directory rather than overwriting it; a missing
or corrupt file is treated as empty rather than failing the apply.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 13:08:07 +08:00
DevvynandClaude Sonnet 5 0e8b1981af fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker:

1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still
   coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same
   config directory and warns about it (confirmed cosmetic - the API key
   wins for actual requests, verified via a real session's own API Usage
   Billing line). A custom-model claude session now gets an isolated
   CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty,
   no files written into it) so there is nothing to conflict with. projects
   is symlinked (junction on Windows) back into the real config dir so the
   response viewer, subagent windows and Read My Mind keep working for that
   session, best-effort.

2. Context-window overflow. Claude Code assumes a large default context
   window for a model id it doesn't recognise and never compacts, so a
   custom endpoint's real, much smaller context (verified live: a 400
   exceeding a 16384-token llama-swap model with a stock ~33.7K-token system
   prompt) silently overflows. Discovery now also learns each model's real
   context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx),
   but ONLY for a model llama-swap's own /v1/models response already marks
   status.value === 'loaded' - never an unloaded one, since llama-swap
   treats ?model= as a routing hint and probing an unloaded model risks
   triggering an actual, slow, GPU-swapping load as a side effect of
   read-only discovery. A server with no status field at all gets no
   enrichment rather than a guess; a model not probed this round keeps its
   previously-learned value until it disappears from the list entirely.
   Stored per model (CustomModelHost.modelContextLengths) and applied via a
   new contextLengthVar registry field, set to
   CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude.

Both new fields live on the existing env-kind customModelInjection
capability shape, declared only on claude's registry entry - every other
CLI's injection is unaffected (pinned by test).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 12:08:26 +08:00
DevvynandClaude Sonnet 5 5c25a52f95 fix(custom-model): wait for a freshly launched session to go idle before applying
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.

Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.

Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 10:38:37 +08:00
DevvynandClaude Sonnet 5 5a9ff07f57 feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real
llama.cpp server:

1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to
   apply the endpoint's defaultModelId (or the first discovered model)
   silently. Now, via the new selectCustomModelEntry() (session-ui.js):
   - exactly one discovered model launches straight away, same as before
   - two or more open a new #customModelPickModal listing every discovered
     model; defaultModelId (if set) is marked but never auto-chosen, since
     the point of asking is letting ONE launch deliberately differ from
     the saved default, not just confirming it
   The endpoint is re-fetched at click time rather than trusting anything
   cached from the dropdown's own render, since the model list can have
   changed (the sweep below, or a settings-panel edit) since it opened.
   runCustomModelEntry() itself — the actual launch, routed through run()
   for the in-flight lock, snapshot-guarded against applying to the wrong
   session — is unchanged; it now just always receives an explicit model
   id from one of these two paths instead of computing one itself.

2. Periodic re-discovery. Every saved endpoint's models now refresh
   automatically every 5 minutes in the background
   (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way
   as the Codex plan-usage poll it sits beside — this.cleanup.setInterval,
   off under testMode), so a model the server starts or stops serving shows
   up without another manual "Discover" click. The manual POST
   .../discover-models route and the new refreshAllCustomModelHosts()
   sweep (custom-model-routes.ts) now share one pure merge step
   (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId
   that no longer appears) rather than two copies that could drift. The
   sweep is best-effort per host — one endpoint being unreachable on a
   cycle never blocks the others — and re-reads the store before each
   host's write, keyed by id, so a concurrent edit or delete from the
   settings panel always wins over a sweep that started before it.

Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated
file for the sweep (kept separate from custom-model-routes.test.ts because
that file's data dir is shared across every test in it — one temp HOME per
FILE, not per test — which would make a sweep-touches-every-host assertion
meaningless there). test/custom-model-run-menu-ui.test.ts gained a new
describe block driving the real picker modal through JSDOM: single-model
bypass, multi-model dialog with the default marked-not-chosen, picking a
row closes the modal and launches with that exact model, the endpoint
re-fetch, and the two "vanished by click time" toast paths.

Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md,
docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated
— the last of these also caught up two sentences that had gone stale after
the draft-review fixes landed (the picker routes through run() now, not a
raw run*() call).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 08:50:18 +08:00
DevvynandClaude Sonnet 5 60e1bd52f7 fix(custom-model): act on the draft review — unparseable onclick, unwrapped envelope, wrong-session apply, missing lock, no tests
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.

Blockers:

1. Every generated inline onclick was unparseable. JSON.stringify's own
   double quotes terminated the double-quoted HTML attribute at the first
   one, leaving btn.onclick null on every picker entry and every Discover/
   Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
   argument, the same idiom deleteCase's onclick already uses four lines
   away in session-ui.js. This also closes the live-HTML-injection route
   through modelId (server-controlled, from the endpoint's own /v1/models
   reply): with quoting intact, a `>` inside it can no longer terminate the
   <button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
   like every other /api route (server.ts's preSerialization hook applies
   to arrays too), so Array.isArray(hosts) was always false in production
   and the picker/settings panel silently saw nothing. Both call sites now
   go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
   returns normally without ever changing activeSessionId, so the apply
   step used to silently re-point and restart whatever session the user was
   already looking at. runCustomModelEntry() now snapshots activeSessionId
   before the launch and requires it to have actually changed.

Majors:

4. Routes the launch through run() itself via a temporary _runMode swap
   (never persisted — setRunMode() would sync it to the server) instead of
   a parallel hardcoded dispatch table, so a custom-model launch now holds
   the same _runInFlight lock every other Run click gets. This also
   resolves the "hardcoded runners map contradicts the PR's own design"
   minor: dispatch is run()'s own, so a CLI whose customModelInjection
   recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
   against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
   parses markup this module generated itself) for exactly the DOM-level
   facts the review said needed no Playwright and no tmux: a generated
   button's onclick genuinely compiles and fires, a dangerous modelId never
   produces a live element, the envelope unwrap works, the session-changed
   guard holds, run() actually gets called (proving the in-flight lock
   engages), and _runMode is restored afterward. Confirmed against the
   pre-fix code first (reproduces btn.onclick === null exactly) so this
   isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
   render-index-html.test.ts for the other fixes below.

Minors:

- Generated entries now filter through isCliAvailable(), matching
  _refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
  (applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
  and to settings-modal open) instead of always rendering; the endpoint GET
  no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
  redactApiKey() replaces the field with a computed apiKeySet: boolean, and
  a PUT with no apiKey now keeps the stored one server-side
  (applyStoredApiKey()) instead of the client resending a value it was
  never given. New tests cover both directions (kept vs. replaced) by
  observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
  (_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
  event, since the real role can resolve after settings were first opened)
  — endpoint writes were already admin-only server-side, but the button
  used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
  toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
  route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
  that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
  (CLAUDE.md already records that exact literal turning the settings
  preview into a grey slab on light skins), .run-mode-custom-models gets
  the same gap: 2px .run-mode-menu's own flex gap only applies one level
  up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
  </script> (CliEntry.label is user-clis.json-settable, unlike
  __codemanCliAvailable's booleans-only payload) via a new exported
  escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
  docs/api-reference.md; left the "no zh-CN for the new Models-section
  group" minor unaddressed only insofar as the wider Models section (task
  routing, thinking effort, etc.) has never had zh-CN coverage either —
  everything this PR itself introduces (labels, hints, button text, the
  Run-menu's "Custom Endpoints" header) IS translated in i18n.js.

Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
DevvynandClaude Sonnet 5 98d26e14d9 docs(wiki): document Custom Model Endpoints and the Run-menu picker
New docs/wiki/Custom-Model-Endpoints.md (auto-synced to the live GitHub
wiki on push to master, per docs/wiki/Contributing.md) covers turning the
feature on, adding an endpoint, the Run-menu picker's one-off-run
behaviour, the per-harness confidence table, and what it deliberately does
not do yet (remote/Docker sessions, live hot-swap). Linked from the
sidebar, from Agent-CLIs.md's "Read next" list plus a short pointer
section, and from Settings-Reference.md's Models section.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
DevvynandClaude Sonnet 5 25fae9ad10 feat(custom-model): generate Run-menu entries from saved endpoint profiles
Follow-up to #393, picking up the work Ark0N invited in his merge comment:
"generate those entries from the saved profiles rather than a fixed
duplicate per harness, and put it in a follow-up PR so this one stays the
backend... The Run-menu picker is yours if you want it."

Adds the frontend surface the backend has been waiting on:

- Run menu: a "Custom Endpoints" section lists one entry per (harness that
  supports customModelInjection, saved endpoint) pair, e.g.
  "Claude Code (llama.cpp)". The harness list comes from
  window.__codemanCustomModelClis, injected at page render straight off the
  CLI registry's own capabilities (never a hardcoded id list in the
  frontend), so a CLI whose injection recipe lands later appears with no
  frontend change. Picking an entry runs that harness's own existing run*()
  function unmodified (case creation, env overrides, everything, forced to
  a single instance) and then applies the endpoint's default model to the
  session it creates via the existing POST /api/sessions/:id/custom-model
  route. Entries are hidden for a remote/docker active case, since that
  route already refuses both.
- Settings: App Settings -> Models gets a "Custom model endpoints" group
  wiring up the customModelEndpointsEnabled toggle (declared since #393,
  read by nothing until now) plus CRUD against the existing
  /api/model-endpoints routes: list, add/edit (inline form), delete,
  discover models.
- Backend: CustomModelHost gains an optional defaultModelId, the model the
  picker applies with no further choice per endpoint (one generated menu
  entry per CLI+endpoint pair, not per CLI+endpoint+model). The route
  refuses a value that isn't one of the endpoint's own discovered models,
  and a fresh discovery drops a default that no longer appears rather than
  carrying an invalid one forward.

Docs: docs/custom-model-endpoints.md describes the new picker and settings
panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the
"backend-only" status note and documents the picker's generation mechanism.

Tests: four new route tests cover defaultModelId validation, acceptance,
and the drop/keep behaviour across a re-discovery; a new render-index-html
test pins the __codemanCustomModelClis injection (present, agent CLIs
supporting the capability, antigravity and shell excluded) and its
solo-window skip. No browser test was added for the Run-menu picker itself
or the settings CRUD panel (this box has no tmux, so the live server used
by test:browser/test:mobile could not be exercised here) -- worth a
Playwright pass before merge, same as any other frontend PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
Randalix d0a5a583cd feat(remote): wake a sleeping host when a session is created or attached
Pressing Run on a remote case whose host was asleep failed with
`could not verify tmux on remote host 192.168.50.137: …` — an ssh error that
blames tmux for a machine that is merely suspended. The only wake paths were
typed input on an established session and the banner's Wake button, so OPENING a
session (the moment the user actually decides to use that host) had none.

`RemoteWakeRegistry.ensureHostAwake()` reuses the existing probe/wake/readiness
machinery for a host that has no session yet, and is wired into the two
user-initiated create paths: `POST /api/quick-start` for a remote case (before
the tmux prereq probe, which is what surfaced the misleading error) and
`POST /api/sessions` with `attachRemoteSession`. A host without a wake target is
not even probed, so its behavior and latency are byte-identical. The wake is
blocking — the caller gets the session or an error — but bounded by
REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS (40 s) instead of the 90 s session default,
because the dashboard sits behind a reverse proxy whose default
`proxy_read_timeout` is 60 s: a longer wait would be cut off at the proxy while
the session was still being created. The budget has to cover the whole request
(40 s wake + 1.5 s probe + the tmux probe's own 15 s = 56.5 s worst case), which
is why it is 40 s and not 45. A timeout now says the host did not come back, and
an unreachable host without a wake target says so instead of pointing at tmux.

The wiring is deliberately in the HTTP ROUTE, never in the shared session
service: `cron-service.ts` builds sessions there with nobody waiting on the
answer, and a wake on that path would power the host on for every schedule —
the timer-driven re-wake invariant #1 exists to prevent. Both halves are asserted
(importers of `remote-wake`, and `ensureHostAwake` having exactly one caller
file), so a future caller has to come through the guard test. A rejection from
the wake IO is caught too: a broken target must fail the wake, not the route.

`remote:hostWaking`/`remote:hostWakeFailed` now carry `forNewSession` for the
session-less case, where "input is queued" would be untrue; the toast then reads
"the session starts when it is back".

Live wake numbers are unchanged (this reuses the measured ~12 s S3 path); the
route behavior is covered by new tests in session-routes.test.ts with an injected
registry, so no test opens a real socket or ssh.
2026-09-15 22:37:37 +02:00
Randalix 8dfc965d13 fix(remote): stop the wake handlers shadowing each other; enforce the input cap
Two findings from a final review pass over the wake-on-LAN feature.

`_onRemoteHostWaking` / `_onRemoteHostWakeFailed` were defined in BOTH
`panels-ui.js` (toasts) and `host-wake-ui.js` (banner). Both files mix into
`CodemanApp.prototype` and `host-wake-ui.js` loads later, so the panels-ui copies
were silently shadowed: the toast never fired, and a wake started for a BACKGROUND
session (input on a non-active tab) produced no notification at all, since the
banner handler only acts on the active session. The handlers now live only in
`host-wake-ui.js`, show the toast unconditionally, and update the banner when the
woken session is the active one.

`appendBoundedPending` dropped only WHOLE chunks, so a single input value over the
cap (one large paste is one `input` value, up to the 100 KB input schema) was kept
in full: "bounded at 4 KB" held per chunk, not per session, and nothing was logged.
The surviving chunk's head is now trimmed too, code-point aware so a multi-byte
character is never split into a replacement char.

Adds the guard that would have caught the first one: every SSE dispatch handler must
be defined in exactly ONE frontend module. The existing test only asserts a handler
EXISTS somewhere, which two modules both satisfy while one is shadowed.
2026-09-15 21:01:53 +02:00
Codeman maintainer bd286bf502 docs(wiki): catch the manual up to 1.29.0 and add the three run modes it never had
The wiki was written for seven run modes and never received Grok Build, DeepSeek
Harness or OMP. They now appear everywhere the others do: the modes table and
per-CLI notes, install commands, environment prefixes, the Quick Start table, the
requirements rows, the vocabulary, and every "seven modes" count.

The 1.27 to 1.29.0 changes land on the pages that own them: attaching a case to an
existing container, multi-case adoption and the copy-a-case picker (Docker Cases);
file reads over ssh in remote cases and what stays unavailable (Remote SSH Sessions,
Working With Files, Security); single-page app routing, frame recovery, localhost
links as tabs and the egress guard (Web Tabs); DeepSeek as the one non-Claude mode
with real stop/blocked signals and Approvals items, Codex's own work detection,
last-response, the model-endpoint routes and refreshed counts (HTTP API, Driving
From An Agent, Hooks, Notifications, Keeping Agents Running, Core Concepts);
Shift+drag, right-click copy, Auto Copy, the Ctrl+Z guard, font weight, the vertical
rail and its activity sort (Keyboard Shortcuts, Input And Voice, The Dashboard,
Settings Reference); the 600px phone cutoff, Codex shift arrows and iPhone Duo
(Mobile Guide); the Docker Compose route and its update rule (Installation, Running
As A Service); four new symptom entries and a "which CLIs" question (Troubleshooting,
FAQ).

Custom model endpoints are deliberately left to #430, which adds that page and edits
Agent CLIs, Settings Reference and the sidebar; these edits stay out of the regions
#430, #428 and #376 touch, and all three still merge cleanly on top.

Both READMEs: the web-tab menu entry is labelled "Add URL" in the UI, not
"Add dashboard".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 19:05:59 +02:00