Compare commits

...
Author SHA1 Message Date
github-actions[bot]Claude Opus 5github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
51b4a1b758 chore: version packages (1.31.0)
* chore: version packages

* chore: sync CLAUDE.md version to 1.31.0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
2026-09-19 13:17:25 +02:00
Codeman maintainer 4205f6930f fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine.

**The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext`
and `confirmedSwap` changed the wire field without moving three assertions that check
it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts`
(the swap modal and the context modal, each of which already receives exactly the right
per-question flag). Moved, with the titles.

**Worse, my own tests for the split never ran.** The four cases in
`session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`,
which was declared inside a sibling `describe`, so they threw a ReferenceError during
setup. The split would have shipped with no passing server-side coverage while the gate
reported the failure as four broken tests rather than as four tests that were never
written. `mockRunning` is hoisted to the outer describe.

**The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved
its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so
the other eight modes fell back to claude's `❯`. That is also starship's default shell
prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line
`❯ npm run build` sits on screen for as long as the command runs, the verifier reads it
as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to
nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N
prompt, `read -p`, an installer or a pager, where it takes the default. The module's own
fileoverview already stated the rule this broke. Now `?? ''`, which
`promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI
that actually declares a composer.

**My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing
three numbered rules.

The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared
in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told
an integrator to retry with `confirmed: true` for both questions, which is precisely the
thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md`
and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR`
multi-user consequence, both user-visible and both previously absent, and #454's gained
the one exception to its own claim: a Custom Endpoints launch ignores the Instance count
stepper and always starts one session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:58:33 +02:00
Codeman maintainer 12de3c5164 docs(custom-model): record the CLAUDE_CONFIG_DIR clamp in architecture-invariants
CLAUDE.md gained the admin-only note when the key joined claude's privilegedEnvKeys;
architecture-invariants, which is where the exact-key allowlist rule is documented in
depth, still described the pre-change world. The reboot-restore half is the one worth
writing down: a non-granted owner's already-persisted CLAUDE_CONFIG_DIR is stripped on
restore, which moves that session back to the default Claude account with no error.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:35:32 +02:00
Codeman maintainer 9af12afb57 docs(custom-model): make the docs match the code, and trim the changeset
More from the review of 5fc391a4, all documentation rather than behaviour.

The changeset was 1602 words of development log, written as the PR grew, with bullets
and loose paragraphs interleaved. That text becomes CHANGELOG.md and the GitHub release
body verbatim, so it is now one user-facing account of what the feature does and what
the real-server work bought, at roughly a fifth the length.

docs/api-reference.md promised a `cmd` field on running-status that the route
deliberately strips (it carries model paths and can carry --api-key).

Two places claimed the apply routes validate `modelId` against the endpoint's
discovered models. Neither does. Dropped the claim rather than adding the check:
discovery can be up to five minutes stale, so a 400 there would refuse a launch that
actually works, and a typo'd id already fails on the CLI's own first request. CLAUDE.md
now says so explicitly, since the absence is the surprising part.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:34:24 +02:00
Codeman maintainer fe3bd0074c fix(custom-model): split the two confirmation questions, and seed the API key the way claude reads it
Two findings from the review of 5fc391a4, both fixed here rather than sent back.

**The API-key trust seed never matched a real key.** `seedApiKeyTrustFile()` wrote the
key verbatim into `customApiKeyResponses.approved`, but Claude Code stores and compares
only the last 20 characters (`key.trim().slice(-20)`, applied on both the write and the
lookup). For any real key the seed missed, so claude stopped at the interactive
"Detected a custom API key in your environment" prompt, whose default is
"No (recommended)": the launch hangs, or silently refuses the key this feature just
injected and falls through to an OAuth login the isolated config dir does not have. It
survived review because a keyless llama.cpp/llama-swap endpoint uses DEFAULT_API_KEY
('local-dummy-key', 15 chars), where slice(-20) returns the whole string and the seed
matches by accident, and every test used a key shorter than that. Now truncated through
`truncateApiKeyForTrustFile()`, with a test using a 57-character key that also asserts
the full credential never reaches that second file.

**One `confirmed` flag answered two different questions.** The context-floor warning
("this model's window is below what this CLI needs") and the swap-conflict warning
("loading this unloads the model another session is using") shared it, and the context
check runs first, so a user clicking "launch anyway" past the context warning silently
consented to evicting someone else's model. They are about different people, so an
answer to one is not consent to the other. Both routes now read `confirmedContext` and
`confirmedSwap` independently; the legacy `confirmed` still means both, because it
shipped in this feature's HTTP-API-only cut and an existing caller must keep working.
The frontend answers each question with its own flag and accumulates them, on the
one-shot path, the restart path and the batch carry-forward alike.

Also from the same review: the swap-confirm dialog no longer renders " are currently
using ..." when multi-user scoping leaves the affected-session list empty (the swap is
blocked regardless of ownership; only the NAMES are scoped), and the per-endpoint
llama-swap log tails are closed in `WebServer.stop()` instead of only by the idle sweep
whose interval that same teardown disposes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:32:50 +02:00
Codeman maintainer 3b55957d79 fix(custom-model): merge-time fixes for the Run-menu picker
Conflict resolution against the five PRs that landed while this was in review, plus
the items left for merge on the thread.

The real one was `session-ui.js`. #454 refactored all eight non-Claude `run*()`
functions to funnel through one `_launchQuickStartInstances()` helper that does the
POST itself, while this PR replaced that same POST in each of them with
`_quickStartWithCustomModelConfirm()`. Resolved in the helper rather than seven times
over: the helper now goes through the confirm path, and each body builder carries the
`customModel` spread. `runAntigravity` deliberately does NOT, since antigravity's
`customModelInjection` is `unsupported`; parity with this PR's own per-mode choices is
asserted rather than assumed.

That merge creates a question neither feature had alone: the confirm dialog now runs
inside a loop that can launch up to 20 instances. Both questions it can ask (context
window too small, and loading this will unload the model another session is using) are
decisions about the ENDPOINT, and every instance in a batch targets the same one, so
the answer is taken once and carried to the rest. Without that a 20-instance launch
asks the same question 20 times.

Also: `sse-events.ts` is 161 constants (master added two for remote wake, this adds
one, verified by counting rather than by arithmetic), `server.ts` keeps both new SSE
prefixes, the two comments pointing at code that no longer exists are corrected, and
CLAUDE.md's SSE and route counts move to 161 / ~236 / custom-model (6).

`pumpLlamaSwapLogTail`'s unparsed remainder is now capped at 64 KiB. It only shrank at
a `\n\n` frame boundary, so a backend that streams without one would grow it for the
life of a deliberately indefinite connection.

NOT changed, deliberately: the context warning and the swap-conflict warning still
share one `confirmed` flag with the context check first, so confirming "launch anyway"
on a too-small context also skips the "this unloads it for another session" ask. That
is the author's documented choice and the reviewer's own note calls it minor. Both
fixes are worse to make here than to defer: separate flags are new wire surface landed
unreviewed during a release, and reordering the checks adds a network round trip to a
path that currently short-circuits. Raised as a follow-up instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:27:01 +02:00
Codeman maintainer 1a99b5836c Merge pull request #430 from opticon454/custom-model-run-menu 2026-09-19 12:25:11 +02:00
Codeman maintainer 035bfbc2fe fix(remote): merge-time fixes for Wake-on-LAN
The MAC-count limit lived in two places that disagreed. RemoteHostSchema.wakeMac's
128-character cap admits seven comma-separated MACs while parseMacList takes at most
four, all-or-nothing, so a five-MAC value validated, was written to remote-hosts.json,
and then resolved to NO wake target: POST /api/sessions/:id/wake answered
"No wake-on-LAN target configured for this host" and the banner offered "Configure WoL"
for a host the user had just configured. MAX_WAKE_MACS now lives in
src/config/remote-wake-limits.ts and both sides refine against it. Its own module
because src/remote-wake.ts is import-fenced to session-routes.ts and server.ts (the
wiring guard that stops a watcher waking a host), and because schemas.ts must not drag
dgram/net/child_process into every request-validating module.

The documented 40 s request budget also omitted the wake's own cost. A `command` target
is bounded by REMOTE_WAKE_COMMAND_TIMEOUT_MS and runs BEFORE the readiness poll, so a
slow one pushed a wakeCommand host's worst case to ~68 s, past the 60 s
proxy_read_timeout the budget exists to stay under. _wakeAndWait now subtracts the
wake's measured elapsed time from the readiness budget, floored at one poll interval so
a wake that ate the whole budget still gets one probe. A magic packet is effectively
instant and is unaffected, which is why live testing never saw it.

Also: the two new endpoints are documented in docs/api-reference.md with the import
fence stated as the rule it is, CLAUDE.md's frontend module count moves to 34, and the
release changesets carry the Thanks section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:40 +02:00
Codeman maintainer 4c705094f7 fix(terminal): ship the copy clean as a trailing trim, without the shared dedent
#451 cleaned two things on copy. The trailing trim is right and every native
terminal does it. The shared leading-indent strip is this project's own rule,
and it is dropped here rather than shipped.

Measured against the shipped transform over 401,445 three-row windows across
1,010 tracked files in this repo, it fired on 73% of them: 92% inside a YAML
workflow, 76% over `git log` output, 48% in a TypeScript source. No width
threshold separates a margin from content because they are the same widths, a
live Claude Code pane's own margins measuring 2 and 5 columns while the most
common non-TUI shared run is 4. The failure modes are not symmetric either: a
wrong trailing trim costs nothing, while a wrong dedent silently deletes
information that was on the screen, with nothing in the clipboard to hint at
it, on git log bodies, on indented code read out of cat (semantic in Python),
on git diff context rows where the leading space is the marker, and on stack
traces.

It also could not be made self-consistent cheaply. Whether the first row joined
the measurement depended on the mousedown COLUMN, which the user never sees, so
one block of three rows produced three different clipboard results; and the
flag read getSelectionPosition().start, which is xterm's mousedown anchor and
is never normalised, so dragging UP through a block read it off the bottom row.
The PR's test stub hardcoded a downward drag, so its suite could not express
that case.

The transform, the wiring, the tests, the invariants, CLAUDE.md, the wiki page
and the changeset all move together. The test block now pins the ABSENCE as a
contract, with the git log, Python and git diff cases as its examples, so this
is not re-derived later. If it is ever revisited, the one qualification that
measured clean is painted trailing padding: zero false positives over all
401,445 windows.

Also from the review: the comments and invariant rule justifying the
padding-only clear described the pre-change code (the Ctrl+C gate reads the
CLEANED selection now, so such a selection falls through to the PTY on its own
and the clear is feedback rather than protection), the new 'Nothing to copy'
toast gained its zh-CN entry, and the invariants paragraph no longer repeats
its own opening sentence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:39 +02:00
Codeman maintainer c376534a50 fix(run,terminal): merge-time fixes for the Instance count stepper and capture geometry
#454: the behaviour the PR adds had no test, so a regression test drives
runGrok() at tabCount 3 and asserts three quick-start POSTs with sequential
w<n>-<case> names (verified to fail against master's session-ui.js). Each
caller now reads the count BEFORE its opening banner and announces it there,
the way runClaude() already did, so a launch no longer prints two headers and
a launch with another session already active still says how many are starting.
runClaude() calls the shared _readTabCount() instead of its own copy of the
1..20 clamp, and that helper optional-chains the element read, since hoisting
it above each caller's try block would otherwise let a missing #tabCount throw
where the launch-error path cannot report it.

#435: sizeMovedUnderLoad derived from data.source alone. `mux-visible` is not
sufficient: a failed display-message cursor query makes capturePaneBuffer skip
the snapshot repaint and return the raw capture, which the route still labels
mux-visible, so a size that moved during such a load bought a full forced
reload to repair a frame that was never positioned. It now tests
Number.isFinite(data.captureRows) like its two siblings.

Plus the invariants and CLAUDE.md lines promised on #435: a visible capture
reports its geometry and omits it when nothing was positioned, the comparison
runs on mux-visible only, and the replay is capped at one attempt and latches
per session when it cannot converge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:39 +02:00
Ark0N 2c3ccdf030 Merge pull request #439
feat(remote): wake a sleeping host (Wake-on-LAN) from input, banner and native magic packet
2026-09-19 12:18:18 +02:00
Ark0N 475436242c Merge pull request #455
fix(input): make sure a prompt sent through the API actually leaves the composer
2026-09-19 12:18:13 +02:00
Ark0N 613b774bf1 Merge pull request #451
fix(terminal): trim the padding and shared indent out of a copied selection
2026-09-19 12:18:08 +02:00
Ark0N 60c9af0599 Merge pull request #435
fix(terminal): replay a pane capture at the geometry it was taken at
2026-09-19 12:18:03 +02:00
Ark0N 2d842ded35 Merge pull request #454
fix(run): make the Instance count stepper work for every non-Claude mode
2026-09-19 12:17:58 +02:00
DevvynandClaude Sonnet 5 5fc391a47c fix(custom-model): address fourth pre-merge review + merge upstream master (Ark0N)
Merged upstream/master (22 commits: reboot-restore recovery feature,
terminal keycode229 recovery work, install.sh/CLI-catalog generator
changes, CHANGELOG/version bump to 1.30.0) into this branch. No
conflicts; git auto-merged every overlapping file (CLAUDE.md,
docs/api-reference.md, app.js, index.html, styles.css, routes/index.ts,
session-routes.ts, schemas.ts, server.ts).

Two required fixes from the latest review:

1. privilegedEnvKeys widening (stock.ts) changes behaviour outside this
   feature. The reviewer decided to keep both CLAUDE_CODE_MAX_CONTEXT_TOKENS
   and CLAUDE_CONFIG_DIR listed (types.ts's rule that every traffic-
   redirecting var this feature introduces must appear there stays
   literally true), and asked for the real consequences documented
   instead of hidden:
   - Corrected session-env-clamp.ts's fileoverview, which stated the
     opposite of what the code now does (reboot-restore's clamp call
     used to be able to strip nothing for claude; it now strips a
     persisted CLAUDE_CONFIG_DIR for a non-granted owner).
   - Corrected the rationale comments in stock.ts: privilegedEnvKeys
     has exactly one consumer (ownerClampedEnvKeys, feeding the
     generic envOverrides clamp on create/quick-start/reboot-restore),
     not the custom-model routes.
   - Added a CLAUDE.md line to the CLAUDE_CONFIG_DIR gotcha covering
     the admin-only-in-multi-user-mode and reboot-restore-strips-it
     consequences.
   - Added a "Claude multi-user clamp" test next to the existing
     DeepSeek/OMP ones, pinning the new stripping behaviour.

2. GET .../running-status (custom-model-routes.ts) no longer passes
   the raw llama-swap `cmd` field (the literal launch line, which can
   carry model paths and --api-key) to the browser -- the frontend
   only ever reads model/state, cmd exists solely for server-side
   parseCtxFromCmd() during discovery. Added a test asserting the
   response never contains cmd or a planted secret.

Also regenerated config/clis.stock.json and install.sh's catalogue
block (npm run generate:cli-catalog) to clear drift introduced by the
upstream merge, since it was failing the sync check.

Left to the reviewer, as they said they'd take at merge: the two
"comments pointing at removed code" cleanups, the two stale CLAUDE.md
counts, and the small items list (mode==='claude' frontend branch,
isCliAvailable() unknown-id gap, shared confirmed flag ordering,
one-shot cancel toast severity, pumpLlamaSwapLogTail buffer cap).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-19 18:05:51 +08:00
Codeman maintainer a0298cf2b1 fix(skill): keep the re-wait open while sendwait works the composer
The first shape of the Enter loop read the composer BETWEEN two short waits,
and tested `wait.ended` (the session exiting) where it meant `timedOut`. A
`stop` that fired while no wait was open was lost, since signals have no
history, and a re-wait that had already resolved on `stop` fell through into
another wait that could never see the edge again: measured twice, the answer
was on screen and sendwait ran its whole 580 s slice anyway.

The long re-wait (a tagged duplicate of the original frame) is now registered
first and kept open in the background for the rest of the call; the loop reads
the composer and re-sends Enter beside it, stops when the prompt has left or
the wait's response has landed, then returns that response. Measured: the
stranded prompt got one extra Enter and sendwait returned on `stop` at 36 s.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:49:02 +02:00
RandalixandClaude Opus 5 5bb489addb fix(remote): authorize the attach wake first; tell the caller what happened to its bytes
Review round 3 on #439.

- The attachRemoteSession branch of POST /api/sessions ran `ensureHostAwake`
  before the multi-user gates, so a non-admin could have any configured
  host's `wakeCommand` spawned (or a packet broadcast) and the request held
  for the wake budget, then be refused for the workingDir. The admin gate
  now comes first, before the host is even looked up; remote hosts are
  admin-only infrastructure everywhere else. Route test: wake spy empty,
  403.
- The non-wait input route answers `{buffered:true}` when the registry took
  the chunk and `{buffered:true, dropped:true}` when it was over the cap
  and is gone (`RemoteInputOutcome` gains 'dropped'); additive to the bare
  `{}`.
- The send-and-wait path answers OPERATION_FAILED when the host never comes
  back, like create and attach, instead of writing into the stalled pane
  and reporting delivered:true plus a timeout.
- The flush writes with `fromUser: true`, so a first prompt buffered
  through a wake can still name the tab.

Docs: api-reference (input route), remote-sessions.md (two invariants),
CLAUDE.md key pattern.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-19 11:39:50 +02:00
Codeman maintainer 19ffe9b7a8 fix(input): make sure a prompt sent through the API actually leaves the composer
Claude Code 2.1.277 takes typed text the moment its composer paints but
ignores Enter for the first 30 to 50 seconds after it (measured 2026-09-19
through the input route: an Enter at 28 s stranded the prompt, one at 51 s
submitted it). The text+Enter pair `sendInput` sends 50 ms apart therefore
left every programmatic prompt sitting unsent, and every waiter burned its
timeout on a turn that never started.

Server: `SubmitVerifier` (session-submit-verifier.ts), armed from
`writeViaMux` for every mux write that carried a carriage return, reads the
pane on a 2 s to 60 s schedule and re-sends Enter only while the last
composer line (the CLI's own prompt glyph) still holds the head of what was
sent. An empty composer, other text, or no composer line at all ends it; a
newer write replaces the schedule.

Skill: `sendwait` gets the same loop (`_composer_text`, no-break space
stripped by its bytes for BSD sed) for servers that predate this, and the
preamble version moves to 1.30.1 so seeded agents pick up the fresh copy.
SKILL.md's heredoc and the plugin mirror are regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 11:38:59 +02:00
Michael GrundbergandClaude Opus 5 95dc6fe944 fix(terminal): remember a geometry replay that did not converge
`resizeRetry` caps the recursion inside one select and says nothing about the
next one, so a pane this browser cannot size reported the same mismatch on
every select and bought the same failed repair each time: two fetches per tab
switch for the life of the page, measured as a running count of 2, 4, 6 across
three selects. That is the case this branch describes as happening every time
rather than occasionally, a phone whose resize `Session.resize` declines while a
desktop claim is live, and it is not the only one — any pane Codeman cannot size
lands there, including one a second tmux client is also holding. Each wasted
pass costs another `capture-pane`, which is `execSync` and blocks the server's
event loop, plus a reset and chunked rewrite, a discarded snapshot and cache
entry, and a dropped and reopened WebSocket.

`_geometryRetryUseless` mirrors the existing `_fullHistoryRepullUseless`: a
retry pass whose frame still does not fit adds the session, geometry that fits
removes it, and the replay gate consults it. The proof has to come from a retry
pass rather than a first one, because the retry ran at the size that stuck and
the pane ignored it. Clearing on a fitting frame is what stops a pane that
becomes sizeable again, once the desktop tab closes or its claim goes idle, from
staying permanently unrepaired. The race case never reaches the latch, since it
converges on its first attempt.

The new browser case walks all of that: three selects reading 2, 3, 4 instead of
2, 4, 6, then a fitting frame, then a mismatch diagnosed afresh. Without the
gate it fails on the second switch with `expected 4 to be 3`.

Rebased onto master, which has moved to 1.30.0 and taken #436. The one conflict
was `config/test-suites.ts`, where both branches appended a glob to
`BROWSER_TEST_GLOBS`; both are kept. Everything else merged clean, #436's own
changes to the same buffer-load path included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:56:58 +02:00
Michael GrundbergandClaude Opus 5 383f834704 fix(terminal): flush unsent local echo before the geometry replay
On a touch device the characters the user has typed live only in the local-echo
overlay until Enter; they have never reached the PTY. The replay re-enters
`selectSession` with `forceReload` on the session that is still active, and that
branch nulled `activeSessionId` before `_cleanupPreviousSession` ran. The flush
there is guarded on a session it can still see, so it was skipped, and the
unconditional `_localEchoOverlay.clear()` that follows took the characters with
it. Measured in chromium against the previous head: typing into the overlay and
then making the call the replay makes left `pendingText` empty with nothing
crossing into the delivery layer on either transport.

The flush moves into `_flushLocalEchoTo(sessionId)`, called from both
`_cleanupPreviousSession` and the `forceReload` branch before it nulls the id.
The session is a parameter because the two callers mean different ones: cleanup
flushes to the tab being left, the branch to the tab being reloaded.

This was reachable before this branch, through the one gesture that already
takes the `forceReload` path on an active session. What is new is that nothing
the user does triggers it. The replay fires on its own the moment a tab switch
finishes, which is exactly when someone typing into a still-loading terminal has
text in the overlay, and on a phone beside an active desktop tab that is every
tab switch.

A seventh browser case pins it: it forces the overlay on, since headless
chromium reports no touch support and the case would otherwise pass vacuously,
asserts the typed characters really are sitting unsent, then triggers the replay
and asserts they reached the session. Without the fix it fails with nothing
delivered at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:44:35 +02:00
Michael GrundbergandClaude Opus 5 e0d4477edc fix(terminal): keep the geometry replay to the pass that can converge
Three follow-ups to the source gate, each one measured rather than reasoned.

A pane already drawing at the size the client just requested is left alone. The
replay runs at `dimsAfterLoad`, so it can only change what is on screen if the
pane was drawing at some other size; when the reported geometry already IS that
size, the second pass captures the identical frame and pays a full reload to do
it, including a visible re-flash, a dropped and reopened WebSocket and a deleted
xterm snapshot. That equality is the signature of a clamp rather than a race:
`getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` does not, so a
terminal narrower than 40 columns or shorter than 10 rows reports a pane
permanently bigger than itself and replayed on every tab switch without ever
converging. A race never produces the equality, since its premise is that the
pane was still at the size it was asked to leave. The declined-resize case does
not produce it either, so that one still costs the single capped attempt and
needs the pane-ownership question this does not touch.

The full-history re-arm is unreachable and now says so. A pass that consumed the
flag sent `full=1`, and the route answers `full=1` with `mux-full-history` or
`history`, never `mux-visible`, so the source gate already rules out every such
pass. The line stays for the invariant, but its comment no longer reads as if a
page load retries, and the suite pins that it does not.

The response no longer reports geometry for a body that carries no capture. The
full-history path writes `capturedGeometry` from the cursor query and then
returns '' for a pane holding nothing visible, which drops the source to
`history` with the geometry already recorded: a `full=1` request whose capture
reported 100x50 and returned nothing answered `source: "history"` with both
fields set. Nothing acted on it, because the client ignores geometry on any
other source, but the field said a frame had been drawn at a size when none had.

The browser stub now derives `source` from the request the way the route does,
rather than answering `full=1` with `mux-visible`, which the route cannot
produce. Each case reaches a visible-frame response the way production does, by
not being the first select of the page. Three cases pin the new behaviour and
each fails without its guard: the clamp case sees two fetches instead of one,
the scope case and the full-history case both see a replay the gate forbids, and
the width case sees one fetch instead of two.

The changeset now describes the change from 1.29.x rather than the difference
between the two commits on this branch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:44:35 +02:00
Michael GrundbergandClaude Opus 5 5cfb98fb8b fix(terminal): compare capture geometry only on a visible-frame response
Only a visible-frame capture positions its rows absolutely, so only that frame
can be damaged by a terminal of the wrong size. A `full=1` body is linear
scrollback closed by a relative cursor move, which is relative precisely so the
browser's row count need not match the pane's, and a `history` body is the byte
stream, which carries no row alignment to protect. The geometry comparison ran
on all three, so it fired most often on the one response it cannot help:
`_fullHistoryLoaded` is empty on the first select of every non-shell session per
page, and a session whose pane a desktop tab holds too tall to ever fit then
paid a second whole-scrollback capture, reset and replay on every page load and
every first tab switch.

`framePositionsRowsAbsolutely` gates both the captured-geometry comparison and
`sizeMovedUnderLoad`. A size that moved under a byte-stream or scrollback replay
is healed by xterm's own reflow plus the SIGWINCH the trailing `sendResize`
already sends.

A pane WIDER than the terminal damages the same frame a second way, so
`captureCols` is now compared rather than only logged. `formatPaneSnapshot`
paints each row out to the pane's own width, so a narrower browser wraps every
painted row, and the wrap on the last one scrolls the whole frame up by a row.

The terminal response no longer falls back to `session.ptyCols`/`ptyRows` when
the capture reported no geometry. The cursor query is what produces the absolute
addressing in the first place, so a capture that lost it returned a raw frame
that was never positioned, and a byte-history response was never positioned
either. Naming the session's own PTY size there described a frame that does not
exist and invited a repair for damage that is not present. `_ptyCols` is also
written only by `resize()` while the PTY is spawned at the size queried from
tmux, so it can be wrong on its own terms. Both fields are now absent instead,
and the `Session` getters added for that fallback go with it.

Two browser cases cover the new behaviour and each fails without its fix: a
`mux-full-history` response with both dimensions mismatched asserts one fetch
(two without the gate), and a `mux-visible` response wider than the terminal
but short enough to fit asserts two (one without the width comparison).

Corrects a claim in the comment above `capturedGeometry` in tmux-manager.ts.
Both replay paths do not address rows absolutely; the full-history one ends in a
relative move, which is the whole reason the gate is right.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:44:35 +02:00
Michael GrundbergandClaude Opus 5 3edf9aae2f fix(terminal): replay a pane capture at the geometry it was taken at
A visible-frame capture repaints each row at an absolute position, counting up
to the pane's height. A terminal shorter than that clamps every address past its
own height onto its last line. The overflow rows then overwrite one another, and
the rows underneath are lost. Replaying a real 50-row capture into a 30-row
terminal rendered 28 lines of a 45-line command and drew the frame twice.

Nothing in the response said what height the frame was built for, so the client
could not detect this. A capture now reports the geometry it was really taken at
through `capturedGeometry` on `PaneCaptureOptions`, and the terminal response
carries it as `captureCols` and `captureRows`. When the captured pane is taller
than the terminal, or the size that produced the capture did not survive the
load, `selectSession` replays once at the size that stuck. `resizeRetry` caps
that at one attempt, so two competing fits cannot trade replays forever.

The retry re-arms the full-history flag only when the pass that ran had consumed
it. A tab switch takes the bounded tail, so its retry takes the tail too:
clearing the flag unconditionally would upgrade that switch into a fresh
scrollback capture the user never asked for, which the route's own comments put
at tens of megabytes.

What this repairs is a capture that won a race against the resize meant to
precede it. It does not repair a capture whose pane was too tall because
`Session.resize` declined the resize outright, which it does for a small
viewport while a desktop viewport's size claim is live. The retry re-sends the
same declined resize and captures the same pane, and `resizeRetry` then stops
it. Repairing that means changing who owns the pane size, which is a policy
question this does not touch. The reported geometry still helps there, because
the client can see the mismatch at all rather than being blind to it.

Follows #395, #396 and #397, which fixed the other ways the replayed frame and
the terminal could disagree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:44:35 +02:00
timkjrandClaude Sonnet 5 358aef16e3 fix(run): make the Instance count stepper work for every non-Claude mode
runOpenCode(), runCodex(), runGemini(), runAntigravity(), runPi(), runOmp(),
runGrok(), and runDeepSeek() all ignored the "Instance count" stepper next
to the Run button and hardcoded a single quick-start call — bumping the
counter to 2 or 3 while on any of these modes silently launched exactly one
session, with no error. Only runClaude() ever read it.

Extract the shared launch-N-sessions-and-select-the-first loop into
_launchQuickStartInstances(), reused by all eight modes, and _readTabCount()
for the shared clamp-and-parse. Each mode still builds its own quick-start
body (config differs per CLI), just via a closure passed to the shared
loop instead of a single inline fetch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-18 18:16:46 -05:00
RandalixandClaude Opus 5 1040f6c489 fix(remote): a proxied host is reachability-unknown; scope remote: SSE per session
Review round 2 on #439.

1. The bare TCP probe connects to host:port, which a host behind a jump host
   or SOCKS proxy does not answer even while ssh works. Acting on that
   verdict drew a permanent banner over a healthy session, replaced a real
   "needs tmux" error with "not reachable" in quick-start, and - with a wake
   target - buffered every HTTP input for the life of the session, since the
   readiness poll could never succeed. `WakeableRemote` now carries
   `jumpHost`/`socksProxy`/`extraSshOptions`, and `isProbeable()` turns such
   a host into reachability-UNKNOWN: input is delivered, `checkReachable` /
   `checkHostReachable` answer `null` (never `false`), `ensureHostAwake`
   returns `'unprobeable'` (handled like `'no-target'`), the quick-start gate
   fires on `=== false` only, and `GET …/reachability` reports
   `reachable: null, probeable: false` so the banner has nothing to key on.
   A wake target can still be fired for it, blind: no readiness poll, no
   reattach, no toast - the response says only whether the packet went out.

2. `'remote:'` joins the session-scoped SSE prefixes. The create/attach wake
   has no session yet, so the registry names the requesting user
   (`ensureHostAwake({ requestedBy })` -> `username` in the payload) and
   `deriveSseHint` routes on it; with neither it fails closed to admins.
   Single-user mode is unaffected.

Smaller, from the same review:

- A flush write that fails now drops the remaining buffer (logged) instead
  of retaining it: the wake still resolved and marked the host reachable, so
  the retained chunk waited for the NEXT wake and was replayed hours later,
  after everything typed since. Same policy as the oversized paste.
- The banner polls on tab activation (a user action) and on its 30 s timer
  only for a host with a wake target; a timer connecting to a host Codeman
  cannot wake is the traffic invariant #2 rejects keepalives for. A proxied
  host is never polled.
- `probeRemoteHostReachable`, `runRemoteWakeCommand` and the default UDP
  socket refuse under VITEST, as remote-files.ts does. The guard caught a
  leak on the spot: `createDefaultRemoteWakeDeps({ probe })` overrode the
  probe but still polled readiness with the real one, so the shutdown test
  had been connecting to a production address. The poll now uses the
  injected probe.
- docs/remote-sessions.md is additions only again (the reformatting is
  gone); the architecture-invariants overlap resolved itself in the merge.

Live, against a throwaway instance with a non-routable ghost host: proxied
-> no probe, no wake, the genuine ssh error after 10 s; direct (control) ->
probe, magic packet, "did not come back" after the 40 s budget.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-18 22:46:11 +02:00
RandalixandClaude Opus 5 e271a65e79 Merge origin/master into feat/remote-host-wake
Resolves CLAUDE.md count tables (route counts recounted on the merged
tree: 235 handlers, sessions 37) and keeps both the host-wake and the
reboot-restore banner in index.html.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-18 22:20:41 +02:00
Codeman maintainer 3cdb4bf42e docs(terminal): the merge-time notes promised on #436
The four edits the review said would be folded in at merge, none of them
code: the changeset becomes one user-facing paragraph, since it is what
CHANGELOG.md and the release notes print; the `_bufferLoadFinishOpts` comment
now names the second contributor to the duplicate window (`captureActivePaneBuffer`
is `execSync`, so anything painted into the pane before the server read it is
in the capture and is broadcast after the reply) and says why a `history`
payload keeps the pre-existing discard when its exposure is the same; the
`_finishBufferLoad` doc block moves from above `_beginBufferLoad` onto the
function it documents; and the test file's header describes both rules the
file now pins instead of only COD-144.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-18 21:38:40 +02:00
Ark0N 492f8d8ddf Merge pull request #436 from irisitymichaelgrundberg/fix/replay-output-that-arrived-after-the-capture
fix(terminal): keep the output a pane capture could not contain
2026-09-18 21:34:42 +02:00
Devvyn 56209e7829 Merge remote-tracking branch 'upstream/master' into feature/run-menu-custom-model-picker 2026-09-19 03:25:26 +08:00
DevvynandClaude Sonnet 5 afb6754453 fix(custom-model): address third pre-merge review (Ark0N)
Blocker 1: the loading banner hides itself ~200ms after it reopens.

- _showCenterStatus reuses one shared DOM node; dismiss() scheduled
  el.hidden = true 200ms later with nothing to cancel it. On the
  Claude path, switchingToast.dismiss() is followed by one same-
  origin request (5-30ms locally) before _watchLlamaSwapLoading opens
  the new banner -- well inside that window -- so the stale timer
  fired against the shared node and hid the fresh banner, leaving the
  whole model-load wait with no progress text, no log line and no
  reachable Cancel button.
- Fixed by parking the pending timeout on the element and clearing it
  at the top of _showCenterStatus. Added a regression test that
  reproduces the exact repro (open, dismiss, reopen 20ms later,
  advance past 200ms) alongside the existing Cancel-button DOM tests;
  confirmed it fails without the fix and passes with it.

Blocker 2: the swap-conflict warning named other users' sessions.

- Both affectedSessions scans (POST .../custom-model and quick-start)
  walked the whole session map with no ownership filter, so in multi-
  user mode a non-admin pointing their own session at a shared
  endpoint learned another user's session name and id -- which with
  autoNameSessions on is that user's own prompt.
- The swap is still blocked pending confirmation regardless of
  ownership (a foreign session is just as real a disruption); only
  which ones get NAMED back to the caller is scoped, via the
  already-imported canAccessOwned. Added a two-owner test to
  test/routes/session-custom-model.test.ts covering both the
  foreign-owner (blocked, not named) and same-owner (named) cases.

Smaller ride-along fixes:

- server.ts boot recovery now passes contextLength into
  applyCustomModelInjection, so CLAUDE_CODE_MAX_CONTEXT_TOKENS is
  correctly rebuilt into _envOverrides after a restart instead of
  surviving only because tmux retains the old setenv.
- pumpLlamaSwapLogTail's finally now deletes by IDENTITY, not just by
  key, so an aborted pump finishing after a newer entry was created
  for the same endpoint can no longer delete that newer entry and
  orphan its connection.
- docs/custom-model-endpoints.md now notes that clearing a custom
  model removes injected keys by name, including CLAUDE_CONFIG_DIR --
  so a session that also had CLAUDE_CONFIG_DIR set via envOverrides
  (the per-client-account case) silently falls back to the default
  account on clear.

Left for later, as flagged in the review itself: the quick-start
case-scaffolding/cancel ordering (real behavioural reordering across
a large handler, too risky to make without a live re-test), and
retiring runCustomModelEntry's mode === 'claude' branch behind a
launchStrategy registry field (explicitly deferred by the reviewer to
"the next one").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-19 03:09:36 +08:00
Michael GrundbergandClaude Opus 5 f9edb33d15 fix(terminal): trim the padding and shared indent out of a copied selection
xterm hands back whole screen rows and trims only the cells that were
never written to, so the real spaces a full-screen TUI paints across the
unused part of a row count as content and reach the clipboard. Measured
against Claude Code in a 282-column pane, single lines arrived carrying
138 trailing spaces, and every line carried the two-space transcript
indent as well. Windows Terminal, iTerm2 and GNOME Terminal all trim that
for you, decideAutoCopy already calls a wall of spaces "never what the
gesture meant", and _selectTouchSelectionLine already treats those cells
as padding — the mouse and keyboard paths never had the same rule.

CodemanCopySelection.clean lives in constants.js beside decideAutoCopy,
its pure sibling. It drops the trailing run from each line, and removes
the leading run only where every selected row shares one. A selection of
a single row keeps its run, because one row shares nothing with anything
and stripping it would silently reindent one line of `git log` body text
or one line out of `less`. A drag that began inside a row keeps its
partial first line untouched and out of the measurement, which otherwise
pins the shared run to zero and leaves every following row indented.

Every pass over a line is a scan rather than a regex. `/[ \t]+(\r?)$/` is
quadratic on a line whose spaces are followed by a non-space character,
which is what right-aligned or centred TUI content looks like: measured
over 50 000 rows with a 280-column run it took 2.9s, against 1.3ms for
the scan, and a 2 000-column run took 16s. The scan is also the faster of
the two on an ordinary padded row.

cleanedTerminalSelection in terminal-ui.js is the half that needs the
live terminal. It returns a COLUMN selection untouched: Alt+drag makes
one, and a rectangle's rows lining up is the point of the gesture, so
both halves of the clean would destroy it. xterm exposes the mode nowhere
public, so the check reads terminal._core._selectionService, the way this
file already reads terminal._core for cell dimensions, and cleans
normally if a future xterm renames the field. A test pins that assumption
against the library rather than against a stub repeating the literal.

The Ctrl+C chord decides on the cleaned selection, not the raw one. A
drag across the blank part of a row selects real padding spaces, so the
raw text is truthy, and testing it would spend that press on a copy of
nothing and make the user press again to interrupt. A padding-only
selection is now dropped and the press falls through to the PTY, while
Ctrl+Shift+C still never falls through. copyTerminalSelection gates on
trim() for the same reason, since a multi-row drag across padding cleans
to line breaks alone and a bare newline pasted into a chat composer
submits it.

All four of the main terminal's copy paths go through it: the Ctrl+C
chord, right-click, the phone selection button and Auto Copy. The
browser's own Edit menu copy, a disabled copy shortcut and the subagent
windows still copy raw rows, as they did before, and the invariants doc
now says so rather than claiming every copy is cleaned. Auto Copy
resolves its own toggle before it reads the selection, since it is off by
default and a selection can run to the 50 000-row scrollback ceiling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 20:18:25 +02:00
Michael GrundbergandClaude Opus 5 3730bc7df5 docs(terminal): correct what selectSession does with the viewport
The JSDoc on `_syncStickyScrollBaseline` said `selectSession` deliberately ends
at the bottom, so the baseline the replay samples is already true there. It
does not. `selectSession` calls `scrollToBottom()` after the write and then
ends at `scrollToLastNonEmptyLine()` (app.js:6512), which targets
`lastNonEmptyLine - rows + 2` and therefore parks ABOVE `baseY` whenever the
replayed frame keeps trailing blank rows — which a full capture does on
purpose, since no transform that can delete a line may run over one.

Its baseline really is a stale true. What covers it is the sticky snap itself:
since de864e7d that snap fires only when the flush found the viewport already
at the bottom (`preserveViewportY === null`), which a parked selectSession
viewport is not. That commit landed on master after this branch was cut, so
the guard arrives with the merge rather than being present here.

`_onSessionClearTerminal` is unchanged in the comment and was correct: it
resets and rewrites with no scroll afterwards, so it does end at the bottom.

Comment only; no behaviour change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 18:56:15 +02:00
Michael GrundbergandClaude Opus 5 cfd771d1d8 test(terminal): pin all four buffer-load paths to the shared flush helper
The first version of this fix decided the flush policy in `selectSession`
alone, and a later pass found it still covering one path of four. Nothing in
the CI gate stops a fifth path, or an inlined `{ flushQueued: true }`, from
splitting that policy up again — the browser suite that would notice is
excluded from `npm test`.

A static scan over `selectSession`, `_onSessionNeedsRefresh`,
`_onSessionClearTerminal` and `_maybeRefetchFullHistory` asserts each one asks
`_bufferLoadFinishOpts`, reusing the `methodBody` slice the sticky-scroll guard
already needed. Verified by inlining the policy back into
`_onSessionClearTerminal`, which fails it by name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 18:44:04 +02:00
Michael GrundbergandClaude Opus 5 75a028e825 fix(terminal): re-take the sticky-scroll baseline after a replay
A capture load now replays its queued tail, and that replay runs through
`batchTerminalWrite`, which samples `_wasAtBottomBeforeWrite` before it queues.
It runs inside `chunkedTerminalWrite`, before that promise resolves, with the
terminal freshly reset and rewritten — so the sample is always true. The caller
then restored the reader's position and the next `flushPendingWrites` scrolled
straight back to the bottom off the latched flag, undoing it. The only thing in
the way was `_hasRecentUserScrollUp()`, a 1500ms window a server-triggered
refresh is usually past.

`_syncStickyScrollBaseline()` re-takes the flag from wherever the viewport now
sits, and the two paths that restore a position call it right after doing so:
`_onSessionNeedsRefresh` and `_maybeRefetchFullHistory`. Those are the paths
#259 and #205 exist for, and they are also where a non-empty queue is most
likely, since a needsRefresh fires when output is flooding. Re-taking rather
than suppressing the sampling: suppressing leaves whatever stale value the flag
held from before the load, which on the full-history re-pull has no reason to
be false. `selectSession` and `_onSessionClearTerminal` deliberately end at the
bottom, so the sampled true is already the truth there and they do not call it.

`_bufferLoadFinishOpts` gains the coverage the CI gate can see: both mux
sources flush, `history` does not, and a payload naming no source does not.
Its only coverage was the browser suite, which CI does not run.

The JSDoc and the changeset now record the one duplicate window this cutoff
cannot close. The server appends output to the byte buffer in the same tick it
emits, but broadcasts on a batch timer — 8ms over WebSocket, 16 to 50ms over
SSE — so a batch pending when `capture-pane` ran leaves the server after the
reply and is replayed although the capture holds it. It is one batch interval
wide against a recovery window spanning the whole chunked write, and closing it
means flushing that batch server side before the capture.

The second browser test asserts its session was created, so a failed create
fails it instead of passing with zero hits.

docs/architecture-invariants.md no longer claims the replay leaves the
queued-event discard window alone. That clause now describes what decides how a
load ends, the baseline rule, the batch window, and the three covering tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:43:04 +02:00
DevvynandClaude Sonnet 5 9982a1325f fix(custom-model): address second pre-merge review (Ark0N)
Blocker: .center-status-banner never actually disappears.

- Add `.center-status-banner[hidden] { display: none; }`, same trap as
  `.home-sessions[hidden]`: the author-level `display: flex` beat the
  UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
  card stayed laid out at `opacity: 0` with its text/cancel/close
  children still `pointer-events: auto` -- an invisible 442x67 click
  blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
  banner (10001) and the swap-confirm/context-warning modals (10010)
  in CLAUDE.md's Z-index layers list.

Stale wording pointed at the reverted sticky-toast default:

- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
  `.toast-message` comment in styles.css all still said "toasts
  default to sticky" after 1f32128c put the flat 3s default back.
  Reworded all three to describe the actual behaviour: one call site
  passes an explicit `duration: 0`.

Smaller items from the same review:

- docs/api-reference.md said discovery failures answer
  `502 OPERATION_FAILED`; OPERATION_FAILED is 422 per src/types/api.ts
  and the error-code table earlier in the same file.
- The periodic re-discovery sweep (server.ts) never read
  customModelEndpointsEnabled, so turning the feature off left
  Codeman polling every saved endpoint forever. Added
  readCustomModelEndpointsEnabled() (custom-model-routes.ts, same
  shape as readPlanUsageTelemetryEnabled) and gated the interval
  callback on it.
- Reverted the formatting-only Prettier pass docs/api-reference.md
  picked up (table padding, *x* to _x_, JSON re-indent) by re-merging
  the new Custom Model Endpoints section onto the pre-PR file, so the
  diff is reviewable. No prose content was lost -- verified by diffing
  the result against the pre-revert file (formatting-only) and against
  the merge-base file (only the new section added).
- docs/custom-model-endpoints.md now states that a custom-model Claude
  session's isolated CLAUDE_CONFIG_DIR loses the user's global
  settings.json, user-level skills/agents/commands, and MCP servers
  from ~/.claude.json -- only `projects` is symlinked back.

Design question left open in the review (does `confirmed: true` need
to be two flags so "launch anyway" on the context warning doesn't also
skip the llama-swap displacement warning): keeping the single flag, as
offered. The 20s displacement sweep still catches a resulting swap
after the fact, so it's a surprise rather than a silent failure, and
splitting it is real behavioural surface I have no way to verify live
in this environment.

`npm run test:browser` could not be run in this environment (no tmux,
no downloaded Playwright browser binary) -- none of its suite's files
touch code this fix changes, but it still needs a real pass before
merge, same as any frontend change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-18 21:45:50 +08:00
DevvynandClaude Sonnet 5 1f32128ca9 fix(custom-model): address PR #430 pre-merge review (Ark0N)
Four blockers from the 2026-09-18 review:

- PUT /api/model-endpoints/:id now merges modelContextLengths/
  modelSizesGB back in from the stored record instead of trusting the
  editor's body, so renaming an endpoint or changing its default model
  no longer silently drops the context-window floor check and
  CLAUDE_CODE_MAX_CONTEXT_TOKENS injection.
- custom-model:swapped-out is now session-scoped (added to
  SESSION_PREFIXES) instead of broadcasting to every connected client.
- The quick-start custom-model path now hands setCustomModel() only
  the endpoint's own injected env vars, not the full merged set,
  matching the restart-in-place path — the full set put
  CLAUDE_CODE_EFFORT_LEVEL back after the Session constructor had
  already stripped it.
- The quick-start launchModel override for pi/grok/omp is now applied
  generically via the registry's legacyConfigField, mirroring
  Session._withCustomModelLaunchModel, instead of three hardcoded
  mode === '<id>' branches a future CLI's injection recipe would miss.

Also scopes the sticky-toast default (item 5): reverted the blanket
"all error toasts are sticky" default, which had no container cap or
eviction, back to a flat 3s; the one message that needs a moment to
read (a failed custom-model apply) now passes an explicit
duration: 0 at its own call site.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-18 20:16:07 +08:00
DevvynandClaude Sonnet 5 e2034177c5 fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1)
DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.

Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
  published dependencies) into a scratch dir purely to read
  dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
  completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
  the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
  completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
  same endpoint. dsh's own error template ("DeepSeek API error (HTTP
  ${status})") reproduces the originally-reported
  "dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.

- New registry field `appendV1Suffix` (env kind only, deepseek's entry
  alone — claude/gemini must NOT get it, since claude was already
  confirmed working against the unmodified baseUrl). When set,
  buildCustomModelInjection runs endpoint.baseUrl through the same
  withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
  use, instead of writing it verbatim.

Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.

2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:41:19 +08:00
DevvynandClaude Sonnet 5 8520925e76 docs(custom-model): bring CLAUDE.md and api-reference.md up to date
Full documentation review pass across the branch's 30 commits.
CLAUDE.md's Custom Model Endpoint Profiles entry hadn't been touched
since the initial backend+picker cut (3 early commits) despite 27
follow-up commits adding real behavior — it described restart-in-place
as universal (now claude-only; 7 other CLIs launch one-shot) and
claimed codex's Responses-API gap as a flat protocol break (now
re-verified as a more precise tool-calling gap). Corrected both and
added a new paragraph covering everything landed since: the llama-swap
conflict check, the after-the-fact swap-displacement sweep, the
/running-cmd-based context-length fix, the context-window floor
warning, skipFirstRunPrompts, the real-time /api/events-based log
status, and the countdown-to-Cancel-button change.

docs/api-reference.md's custom-model-endpoints section was missing the
running-status route, the requiresConfirmation/requiresContextWarning
response shapes, and POST /api/quick-start's customModel field
entirely (the primary launch path for 7 of 8 supported CLIs) — added
all three. Also fixed a real markdown bug in custom-model-endpoints.md:
an inline code span (`POST <baseUrl>/v1/chat/completions`) split across
a line break, which CommonMark renders with the line ending collapsed
to a space, so it displayed as ".../v1/chat/ completions" with a
spurious space inside the path.

Verified: origin/master and upstream/master are both already an
ancestor of this branch (identical at bd286bf5, no new commits since
this branch was cut) — nothing to merge, no conflicts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:18:04 +08:00
DevvynandClaude Sonnet 5 db9729e1fc feat(custom-model): remove loading-banner countdown, add manual Cancel
Replaces the size-scaled expected-time estimate + matching auto-timeout
with a generic hardware/model-size disclaimer and a user-driven Cancel
button, per explicit request. Real load time depends on hardware this
feature has no way to know (VRAM, storage speed, GPU contention), so
the old estimate/timeout was a guess dressed up as a fact — worse, one
that could kill a genuinely slow load partway through on slower
hardware.

- _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline
  entirely — polls indefinitely until ready or cancelled, no automatic
  give-up. Message is now "Loading <model> (<size>) on <endpoint> —
  this can take a while depending on your hardware and the model
  size.", with the real llama.cpp log line still on its own second
  line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/
  _formatRemaining (dead code once the countdown is gone) —
  _lookupModelSizeGB is kept, the GB figure still shows.
- _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real
  "Cancel" button (distinct from the error-type "×" close button,
  since Cancel has a real consequence) that calls it on click. Caller
  owns what cancelling actually means, same split as the swap-confirm
  modal's promise-resolving buttons.
- Cancelling dismisses the banner, shows an info toast (not an error —
  this was deliberate), and closes the session, mirroring what the old
  timeout used to do automatically but now on the user's own call.
- New .center-status-cancel CSS (bordered pill button, distinct from
  the plain "×" close glyph).

Test changes: removed the now-invalid timeout-auto-close/estimate
tests, added cancel-flow tests (dismiss/toast-type/session-close,
never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests
for the new Cancel button (bootAppWithRealCenterStatus, evaluating
panels-ui.js instead of stubbing _showCenterStatus, since this button
is worth verifying for real rather than just through the stub every
other test in the file uses). Typecheck/lint/frontend-syntax clean;
full suite shows no new regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:00:50 +08:00
DevvynandClaude Sonnet 5 2d3fc65758 feat(custom-model): show real-time llama.cpp backend status in the loading banner
Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.

- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
  (custom-model-routes.ts): one persistent GET /api/events (SSE)
  connection held open per endpoint, parsing logData frames and
  keeping the latest source:"upstream" (backend llama-server) line —
  filtering out llama-swap's own source:"proxy" request-access lines.
  Idle-closed after 30s of no polling, same 20s sweep as the existing
  swap-displacement check.
- running-status route now returns logLine alongside the existing
  isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
  ("llama.cpp: <line>", bootlog timestamp/level/component prefix
  stripped for display) that stays on the last real thing llama.cpp
  said rather than clearing to blank between polls.

⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).

12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 12:22:45 +08:00
DevvynandClaude Sonnet 5 5ddc028a2f feat(custom-model): detect and notify when a session's model gets swapped out later
The llama-swap conflict check on the apply/create routes only ever runs
at THAT session's own launch/apply moment, and cannot see a swap caused
by a DIFFERENT session's later, ordinary use. Confirmed live: a second
Codex session picking a different model launched with no warning at
all — nothing conflicted at that exact instant — yet it silently
evicted the first session's model regardless (llama.cpp runs one model
at a time). Reproduced and root-caused via direct API calls against a
live test-picker instance rather than guessing.

- detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups
  live sessions with a customModel by endpointId, checks each group's
  endpoint via GET /running once, and flags a session whose own modelId
  is no longer in the running list. Read-only, best-effort per endpoint
  like refreshAllCustomModelHosts's sibling sweep.
- Notifies once per displacement via a caller-owned de-dupe Set: a
  session id is added when displaced, removed once its own model is
  loaded/ready again, so a later genuinely-new displacement can notify
  again.
- New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS,
  20s — much shorter than the 5-minute model-list refresh, since this
  is time-sensitive) broadcasts a new custom-model:swapped-out SSE
  event per displacement. De-dupe Set cleared per-session on session
  cleanup to avoid an unbounded leak.
- Frontend: global toast (not tied to the displaced session's tab,
  since the point is warning before the user types into it) naming the
  session, its previous model, and what's currently loaded.

Chose the "detect after the fact" scope (vs. checking before every
message send, which would add a round-trip to every turn on every
custom-model session) per explicit user decision after being presented
the trade-off.

9 new tests for the detection logic (flag/clear/re-flag cycle,
unreachable/deleted endpoints, non-llama-swap servers, multiple
sessions on one endpoint). SSE registry bumped 158->159, parity test
passing. Typecheck/lint/frontend-syntax clean; full suite shows no new
regressions (9 more passing than baseline, matching the new tests;
same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 11:09:24 +08:00
DevvynandClaude Sonnet 5 470f75b08c docs(custom-model): record live findings on codex's model-metadata warning
Investigated the user's report of "Model metadata for <id> not found.
Defaulting to fallback metadata..." on every custom-endpoint codex
launch, live against the test-picker's llama-swap deployment (codex
0.152.1):

- The warning is cosmetic. `codex exec 'reply with just OK'` against the
  isolated CODEX_HOME still printed the warning and still returned a
  real reply.
- The isolated CODEX_HOME never gets a models_cache.json written into
  it at all, even after extended real use (inspected a live, actively-
  used directory) — codex can't reach OpenAI's own hosted model catalog
  for this session and silently falls back every time, with no local
  file to create or clean up. There is also no config.toml override for
  a model's metadata.
- Fabricating a fake catalog entry to suppress it would mean copying the
  SHAPE of OpenAI's own proprietary models_cache.json schema, including
  real per-model system-prompt content visible in a genuine entry — not
  something to build for a warning confirmed to have no effect.
- More importantly: a real tool-call attempt against the same setup came
  back as agent_message TEXT (the tool-call JSON printed as the answer)
  rather than an executable function_call item, confirmed via
  `codex exec --json`'s raw event stream. Tool execution is what makes
  codex a coding agent, so it remains not usable for real work regardless
  of the metadata warning — a more precise, re-verified update to the
  existing "Responses API protocol gap" finding (which reported a harder
  Reconnecting/high-demand failure on a different llama-swap deployment;
  this one answers /v1/responses for plain chat but still can't execute
  tools).

No code changes — recipe/comment/confidence-table documentation only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 10:33:48 +08:00
DevvynandClaude Sonnet 5 211b872335 feat(custom-model): skip Claude Code's first-run wizard on custom-model launches
A fresh, isolated CLAUDE_CONFIG_DIR (used to keep an injected API key
away from a stored claude.ai OAuth login) looks like a brand-new Claude
Code profile to the CLI, so it replays its ENTIRE first-run sequence on
every single launch: the theme picker, the security-notes screen, the
per-project "trust this folder?" dialog, and (running with
--dangerously-skip-permissions) a one-time bypass-permissions warning —
confirmed live, none of which a real, already-onboarded profile shows
again.

- New registry-declared env-kind field `skipFirstRunPrompts` (alongside
  apiKeyTrustFile, which it reuses) — claude's entry only, carried
  through buildCustomModelInjection (pure) into
  applyCustomModelInjection (IO).
- seedFirstRunOnboardingState(): merges hasCompletedOnboarding: true and
  this session's own projects[workingDir].hasTrustDialogAccepted: true
  into the same <configDir>/.claude.json the API-key trust file already
  writes to — other projects and other fields on this session's own
  entry are left untouched.
- seedSkipBypassPermissionsPrompt(): merges
  skipDangerousModePermissionPrompt: true into <configDir>/settings.json,
  a separate file, same corrupt-tolerant merge behavior.
- applyCustomModelInjection() gains an optional workingDir parameter,
  threaded from session.workingDir (dedicated apply route) /
  resolvedCasePath (quick-start route) — boot recovery omits it
  (a dialog already answered once needs no re-seed on the same,
  persisted isolated directory).

Tests added at the pure-builder, IO-wrapper (including merge-preserves-
other-fields and corrupt-file-tolerance cases), and existing directory-
listing assertions updated for the new settings.json file. Typecheck/
lint/format clean; full suite shows no new regressions (baseline
pre-existing Windows-environment failures unchanged, 8 more passing
tests than before — the ones added here).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 09:21:36 +08:00
DevvynandClaude Sonnet 5 2c89359d42 fix(custom-model): Cancel/Launch-anyway buttons stacked instead of side by side
Neither dialog's footer had a row layout of its own to override, and
.btn-toolbar is display:flex (a block-level flex container with no
explicit inline-flex), so with no flex row context each button took its
own full-width line and the two stacked vertically. The swap-confirm
modal already had a .modal-footer rule (flex-end); the context-warning
modal had none at all. Both now share one row-layout rule, centred
rather than flex-end per feedback.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 09:01:28 +08:00
DevvynandClaude Sonnet 5 962029bb3d fix(custom-model): context-warning/swap-confirm modals hidden behind status banner
Both dialogs can appear while the centred llama-swap status banner is
still on screen (right after "Claude started — switching to
llama-swap…") — the banner's z-index is 10001, .modal's base z-index is
only 1000, so the dialog rendered fully behind it. Reported live against
the context-window-too-small modal; the swap-confirm modal has the same
structural bug for the same reason, so both get the fix.

Also: both messages ARE the modal's whole explanatory content, not a
one-line caption under a form field, so .form-hint's 0.65rem caption
size read as illegibly small — worst on the multi-sentence
context-window explanation. Bumped to 0.85rem/1.5 line-height/--text.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 08:57:12 +08:00
DevvynandClaude Sonnet 5 b45a96358e feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.

- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
  registry entry declares contextLengthVar (currently only claude) and
  the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
  (40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
  quick-start customModel path) check this before the swap-conflict
  check and before launching/restarting anything, returning
  {requiresContextWarning, modelId, contextLength, minSafeContextTokens}
  — skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
  _resolveContextWarningConfirm (session-ui.js), wired into both
  _quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
  (the path Claude actually uses) ahead of the swap-confirmation check.
  Explains the fix in-modal: give the model an explicit larger -c/
  --ctx-size in llama-swap instead of relying on --fit-ctx, which
  optimizes for the biggest model that fits rather than the biggest
  context.

Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 07:36:57 +08:00
Randalix 29984c639d fix(remote): stop the flush losing a chunk, and reset the host form's wake fields
Own review pass over the PR:

- `_flush` took the chunk out of the buffer only AFTER awaiting the write. Input
  arriving during that await is enqueued (`waking` is still set, so it takes the
  buffer path), and the 4 KB cap then drops the OLDEST chunk — which is the one
  already on its way to the pane. The `shift()` that followed removed the NEXT
  chunk instead, so the drop-oldest bookkeeping silently lost a chunk that was
  never written, while the log line blamed the one that was. The chunk is now
  removed before the await and re-inserted at the FRONT on a failed write, so the
  order of the queue behind it is preserved. Regression test: a chunk enqueued
  during the first write of a full buffer must still reach the pane (red against
  the old order).
- `showCreateCaseModal()` reset the remote-host form fields but not the two new
  wake inputs, so one host's MAC/command carried over into the next host that
  form saved.
- The banner's pre-poll `wakeConfigured` labelled a command-only host as 'mac'.
  Nothing reads the distinction, but the field is documented as which path is
  configured, so it says the truth until the first poll corrects it.
- Stale `resolveRemote` comment ("only for sessions that have no usable target of
  their own"): after the host config became authoritative in both directions it is
  consulted on the TTL regardless.
2026-09-16 21:06:25 +02:00
Randalix acb8d4b0aa docs(remote): correct what the wake PR moved
- `host-wake-ui.js` joins the documented load order (12.2) and gets its
  `@dependency`/`@loadorder` tags; the frontend module count is 33, not 32.
- `remote-wake` is not "(pure)" — the module uses `dgram`/`net`/`child_process`.
- SSE counts: 160 constants, and the category is "Remote auto-reconnect / wake
  (5)"; the route table's per-file counts are refreshed (sessions 37, cases 34).
- The CLAUDE.md wake rule now names the create/attach wake, the 40 s request
  budget, the whole-chunk paste drop, the registry's lifetime (drop on cleanup,
  stop on shutdown) and the deliberately non-wake-aware WebSocket keystroke
  path — that paragraph is what the next person reads.
- Reverted the eight lines of unrelated Prettier markdown churn in
  `docs/architecture-invariants.md` (docs/ is not in the format glob, so it was
  an editor): only the new wake paragraph remains in the diff.
2026-09-16 20:44:48 +02:00
Randalix 7b947fa3f1 fix(remote): close the wake-state leaks and the dishonest wake budget
Review follow-up on the wake-on-LAN PR (five findings, all of them about the
state the feature keeps and the budgets it inherits):

- Wake state is dropped by `WebServer.cleanupSession` instead of the two delete
  routes, so it now goes with the session on EVERY cleanup path (cron, admin,
  scheduled-run teardown, error paths) instead of surviving with up to 4 KB of
  the user's buffered keystrokes. `registerSessionRoutes` returns the registry
  so the server can own its lifetime without the wake-capable code living in
  `server.ts`; the wiring guard is updated to allow that and gains a second
  assertion that `server.ts` calls nothing but `drop`/`stop` on it.
- `_effectiveRemote` returns before `_state`, so a LOCAL session no longer gets
  a wake-state entry — the input gate runs on every keystroke, so that entry
  used to be allocated for every session the user types in.
- An input chunk larger than the 4 KB cap is dropped OUTRIGHT instead of being
  head-trimmed and then written as a fragment: one paste is one `input` value
  and was never typed character by character, so its tail is a partial command
  the user never sent. The drop is logged.
- The manual wake button passes `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS` (40 s)
  like the create/attach paths, instead of inheriting the 90 s session default
  that the dashboard's reverse proxy cuts off at 60 s.
- `RemoteWakeRegistry.stop()` aborts in-flight readiness polls (abortable
  sleep) and refuses new wakes, and `WebServer.stop()` calls it, so a restart
  during a wake no longer waits the poll out.
- The banner/toast wording keys off a new `queuedInput` flag on the two SSE
  events, which is true only when the server actually holds bytes: browser
  keystrokes travel over the WebSocket, which never passes through the
  registry, so the wake BUTTON must not promise queued input. The failed-wake
  path also stops pattern-matching the error message (it re-asks the
  reachability route) and the WoL dialog says "admin-only" instead of "host not
  found" for a non-admin in multi-user mode.
2026-09-16 20:44:39 +02:00
DevvynandClaude Sonnet 5 993710263d fix(custom-model): stop trusting /props's n_ctx, parse the real context size from /running's cmd
Root cause of the context-overflow regression reported live: "API Error: 400
request (36437 tokens) exceeds the available context size (16384 tokens)".
Discovery had stored modelContextLengths.qwen3.8-27b-ud-q4_k_xl = 154112,
so CLAUDE_CODE_MAX_CONTEXT_TOKENS told Claude Code it had a huge window and
it never compacted - but the real llama-swap server was launched with
--fit-ctx 16384 (confirmed against /running's own cmd field) and refused
the request right at that real limit.

/props?model=<id>'s n_ctx (the field discovery read) is confirmed live to
be unreliable for a --fit-ctx-launched backend: it reported 154112 for the
same model /running says was launched with --fit-ctx 16384 - appears to
report the model's theoretical/trained maximum context, not the runtime-
configured one.

discoverModels() now parses the REAL configured size straight out of
llama-swap's own launch command instead (parseCtxFromCmd(), reading
/running's cmd field - --fit-ctx first, then the plain llama.cpp -c/
--ctx-size a hand-written command might use), and only falls back to the
old /props probe when cmd states no recognizable flag at all. One /running
call now covers every loaded model's context length in a single request,
same as it already did for the swap-conflict check and the load trigger.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 20:47:12 +08:00
DevvynandClaude Sonnet 5 7bbe408e44 feat(custom-model): live countdown on the loading banner; timeout is now an error
The loading banner now shows a live countdown against its own timeout
(updated every poll, so every second by default) instead of a static
"this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically
~1-3 min) on llama-swap - 47s remaining".

If the countdown reaches zero and the model still isn't ready, this is now
treated as a real failure rather than a "keep waiting" shrug:
- The banner turns into a sticky error (_showCenterStatus gains a `type`
  option - 'error' drops the spinner and adds a close button, since nothing
  is "in progress" anymore and a sticky message needs a way to dismiss it),
  naming the llama-swap server's own logs as where to look for detail.
- The session that load was for is closed automatically (closeSession) -
  requested explicitly: a console left open and pointed at a model that
  never finished loading is worse than no console at all. Both apply paths
  now thread the new session's id through to _watchLlamaSwapLoading for
  this (new required 3rd parameter, after endpointId/modelId).

_watchLlamaSwapGeneration's existing stale-call guard extends naturally to
this: a superseded call's own eventual timeout recognises it no longer owns
the banner and neither shows the error nor closes a session that may by
then belong to a different, newer launch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 20:22:02 +08:00
DevvynandClaude Sonnet 5 55dae31530 feat(custom-model): estimate model load time from its discovered size
Discovery now also parses a GB figure out of an auto-discovered model's own
description (llama-swap writes "Auto-discovered 16.35 GB - parameters
auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike
context length this needs no /props probe (the figure is right there in
/v1/models) so it is populated for every model regardless of loaded state.
A hand-configured profile's own description has no such figure and
correctly gets no entry.

The loading banner (_watchLlamaSwapLoading) now looks this up and, when
known, shows it plus a rough estimate from a small size->time matrix
(_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) -
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on
llama-swap... this can take a while" - and uses that same estimate's own
bracket to scale the banner's default give-up timeout for a very large
model, instead of a flat 5 minutes for everything. Explicitly labelled as
an UNMEASURED, typical-hardware estimate in every relevant comment - this
is not benchmarked against any real endpoint's actual storage/GPU, just a
reasonable expectation-setter. A model with no discoverable size (a
hand-configured profile) gets no size/estimate shown at all, matching the
"never a guess" convention modelContextLengths already established.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 19:58:46 +08:00
DevvynandClaude Sonnet 5 0af233c96c fix(custom-model): poll llama-swap readiness every 1s, check immediately, extend the cap
Reported: the "Loading..." banner stayed up past 2 minutes even though
llama-swap itself had already finished loading the model. Three fixes:

1. pollIntervalMs default 3000ms -> 1000ms (as asked).
2. The loop now checks readiness IMMEDIATELY on entry rather than sleeping
   a full interval first - a model that's already ready (a fast load, or a
   re-apply onto one already loaded) shouldn't sit on "Loading..." at all.
3. maxWaitMs default 120000ms (2 min) -> 300000ms (5 min): a large (20GB+)
   model reading from disk can genuinely take longer than 2 minutes, which
   would have looked identical to the reported symptom - "still stuck past
   the point it should have resolved" - except it would have actually
   flipped to a "still waiting" warning toast at the 2-minute mark rather
   than staying on "Loading" indefinitely, so this alone doesn't explain
   what was reported, but is a real, separate improvement worth making.

Also fixes a real, separate bug this surfaced while reasoning through the
report: _showCenterStatus's banner is ONE shared, reused DOM node. A second
call to _watchLlamaSwapLoading (e.g. switching models again before the
first switch's loop had finished) would take over that shared banner, but
the FIRST loop was still running and would eventually dismiss or overwrite
it once ITS OWN deadline or readiness check resolved - clobbering whatever
the second, current loop had put there. A generation counter
(_watchLlamaSwapGeneration) now lets each call recognise when it no longer
owns the banner and stop touching it silently, rather than only the last
call to actually start ever safely reading or writing it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 19:44:37 +08:00
Randalix a7f74f374f fix(remote): keep the wake banner hidden after switching to a local session
refreshHostWakeBanner clears _hostWake before calling _hostWakeTick, so the
clear branch's `if (this._hostWake)` guard skipped the repaint: once the
banner had appeared for an unreachable remote session it stayed up on every
chat (local ones included) until a reload, and the 30s ticker never cleared
it either. Render unconditionally in that branch — _renderHostWakeBanner is
idempotent with a null state.

Reproduced in a real browser (Puppeteer, mobile viewport): state went null
but banner.hidden stayed false. Regression test added in
test/host-wake-banner.test.ts (red before, green after).
2026-09-16 10:39:48 +02:00
DevvynandClaude Sonnet 5 0929694012 fix(custom-model): actually trigger the llama-swap load, not just watch for it
Root cause of "it doesn't look like llama-swap is actually switching the
model" (confirmed live: no load_model line in llama-swap's own logs after
applying a selection). llama-swap has no "switch model" admin endpoint - the
ONLY thing that starts a swap is a real inference request naming the model.
Every previous fix (the conflict check, the loading banner) assumed a swap
would start on its own; nothing ever actually asked llama-swap to load
anything until the launched CLI's first real prompt did, which could be
much later than "applying the selection" implied.

Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest
real request that will start a load - POST <baseUrl>/v1/chat/completions,
max_tokens: 1, one throwaway message - fire-and-forget (never awaited by
the caller; the frontend's own running-status polling is what actually
confirms readiness). Wired into both apply paths (the dedicated restart
route and the one-shot quick-start route), fired whenever the target model
isn't already the one loaded and ready - a broader condition than the
existing swapNeeded (which only gates the "this will evict another
session's model" confirmation ask and deliberately stays narrow to that).
modelSwapInProgress in both routes' responses now reflects this same
broader condition too, so the frontend's loading banner actually correlates
with a real in-flight load rather than only firing when something else
happened to be loaded already.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:53:34 +08:00
DevvynandClaude Sonnet 5 01b32ee6cd fix(custom-model): move the switching/loading status to a centred banner
The "Claude started - switching to <endpoint>..." and "Loading <model> on
<endpoint>... this can take a while" messages lived in the top-right toast
corner along with everything else, easy to miss given they can each sit on
screen for well over a minute (a real llama-swap model load).

Adds _showCenterStatus() (panels-ui.js): a single, reused, screen-centred
banner with a spinner, non-blocking (no backdrop, pointer-events: none on
the wrapper) so it never gets in the way of using the app while it's up.
Both call sites (_runCustomModelEntryViaRestart's switching message,
_watchLlamaSwapLoading's loading message) now use it instead of showToast.
Every OTHER status in these two flows - the llama-swap conflict warning
already moved to its own modal, apply failures, cancellation, and
_watchLlamaSwapLoading's own final "ready"/"still waiting" outcome - stays
exactly where it was, in the corner.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:40:43 +08:00
DevvynandClaude Sonnet 5 2936ba6e3d fix(custom-model): replace the native confirm() popup with an in-app modal
The llama-swap "this will unload it for session X" warning used a native
browser confirm() popup, which looks out of place next to the rest of the
app's own modals.

Adds #customModelSwapConfirmModal (index.html) with Cancel/Switch-anyway
buttons, styled to match the app. _confirmModelSwap(message) shows it and
returns a promise that resolves true/false the same way confirm() would;
_resolveModelSwapConfirm(proceed) (wired to both buttons and the backdrop
click) settles it. Both llama-swap conflict call sites
(_quickStartWithCustomModelConfirm for the one-shot launch path,
_runCustomModelEntryViaRestart for Claude's restart path) now await this
instead of calling confirm() directly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:29:38 +08:00
DevvynandClaude Sonnet 5 83033b4299 fix(custom-model): show a status toast during Claude's native-boot-then-restart window
Claude stays on the launch-then-restart path (see runCustomModelEntry's own
comment for why), but with nothing on screen during that window, a native
boot that briefly talks to the cloud model read as "the endpoint didn't
apply" rather than "the switch hasn't happened yet".

A sticky "Claude started - switching to <endpoint>..." toast now covers the
whole window from the native launch through the apply call, updated in
place (never stacked) as the outcome resolves: dismissed on cancel or
failure (replaced by the existing cancellation/error toast), handed off to
_watchLlamaSwapLoading's own sticky toast when a model swap is in progress,
or updated to the existing "Pointed at ... - restarting" message and
auto-dismissed after 3s on a plain success.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:21:09 +08:00
DevvynandClaude Sonnet 5 f865f74a0f feat(custom-model): launch directly on the endpoint, no restart, for 7 of 8 CLIs
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.

POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.

Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.

Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:04:10 +08:00
DevvynandClaude Sonnet 5 fbee1b2d82 docs(changeset): add changeset for the Run-menu custom-model picker PR
Covers #430's full scope so far: the picker itself, the model-selection
dialog, periodic re-discovery, and the session-busy/toast/CLAUDE_CONFIG_DIR/
context-length/llama-swap-conflict fixes found through live validation
against a real llama-swap server.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 14:39:56 +08:00
DevvynandClaude Sonnet 5 bcebc81fcd feat(custom-model): detect llama-swap model conflicts before switching
Root-caused the user's earlier confusion ('the terminal says opus even though
something is waiting for llama to load'): llama.cpp runs exactly one model at
a time, and llama-swap unloads/reloads it on demand - a swap can take
anywhere from a few seconds to well over a minute, during which a session
looks indistinguishable from one still on the native backend.

1. Feature-detects llama-swap (vs. plain llama.cpp/any OpenAI-compatible
   server) via its own GET /running, which plain llama.cpp has no concept of
   at all. New GET /api/model-endpoints/:id/running-status route exposes this
   read-only, for the frontend's polling loop below.

2. Before applying a selection, POST /api/sessions/:id/custom-model now checks
   what llama-swap currently has loaded. If it differs from the requested
   model AND another live session's own customModel selection is actively
   using that loaded model, the apply is refused with a
   {requiresConfirmation, currentlyLoadedModel, affectedSessions} payload
   instead of silently switching. A "confirmed: true" field on the retry
   skips the check. Switching with nothing else affected proceeds
   immediately, no confirmation asked, only ever when there is something to
   warn about.

3. The frontend (runCustomModelEntry) shows a native confirm() naming the
   affected session(s) and the model they'd lose, matching this codebase's
   existing convention for this class of decision (delete case, kill
   session, etc.) rather than a new modal. On a successful apply the response
   also carries modelSwapInProgress; when true, a new _watchLlamaSwapLoading
   poll shows a sticky "Loading <model>..." toast via the new running-status
   route until llama-swap reports the target model ready (bounded at 2
   minutes), so a prompt sent mid-swap reads as "loading", never as silence
   or an answer from whatever was loaded a moment before.

Checks are read-only against llama-swap's own /running - never /props, which
takes a ?model= and can itself trigger a load as a side effect of asking.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 14:35:07 +08:00
DevvynandClaude Sonnet 5 25f22b9839 test(custom-model): update session-custom-model route test for CLAUDE_CONFIG_DIR isolation
Fixes the CI failure on the last two commits: this route test asserted an
exact envKeys list for a claude-mode apply that predates the
CLAUDE_CONFIG_DIR isolation fix, so it failed on the new CLAUDE_CONFIG_DIR
entry it correctly started appending. Updates the expected list and adds
assertions for the isolated config dir path and the pre-seeded
.claude.json trust-approval file, matching the behavior added in the two
prior commits rather than just tolerating it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 13:23:54 +08:00
DevvynandClaude Sonnet 5 97464bfa27 fix(custom-model): pre-approve the injected API key in the isolated Claude config dir
The CLAUDE_CONFIG_DIR isolation from the previous commit fixed the cosmetic
auth warning but introduced a real regression: an otherwise-empty config
directory has none of a real profile's prior custom-API-key approvals, so
Claude Code stops at an interactive 'Detected a custom API key - use it?'
prompt on every single launch. Confirmed live. With nobody at a TTY to
answer, the prompt's own default ('No') silently refuses the very key this
feature just injected, which looks like the endpoint being ignored.

Adds apiKeyTrustFile to the env-kind customModelInjection capability shape
({relPath, shape: 'claude-api-key-responses'}), set on claude's entry to
{relPath: '.claude.json', shape: 'claude-api-key-responses'}. The apply step
merges customApiKeyResponses.approved: [apiKey] into
<isolatedConfigDir>/.claude.json - the exact field a real answered prompt
itself writes to (confirmed against a real ~/.claude.json after answering by
hand once), so this answers the prompt in advance rather than bypassing it.
Merges onto whatever the CLI already wrote into that file on an earlier
launch in the same isolated directory rather than overwriting it; a missing
or corrupt file is treated as empty rather than failing the apply.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 13:08:07 +08:00
DevvynandClaude Sonnet 5 0e8b1981af fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker:

1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still
   coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same
   config directory and warns about it (confirmed cosmetic - the API key
   wins for actual requests, verified via a real session's own API Usage
   Billing line). A custom-model claude session now gets an isolated
   CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty,
   no files written into it) so there is nothing to conflict with. projects
   is symlinked (junction on Windows) back into the real config dir so the
   response viewer, subagent windows and Read My Mind keep working for that
   session, best-effort.

2. Context-window overflow. Claude Code assumes a large default context
   window for a model id it doesn't recognise and never compacts, so a
   custom endpoint's real, much smaller context (verified live: a 400
   exceeding a 16384-token llama-swap model with a stock ~33.7K-token system
   prompt) silently overflows. Discovery now also learns each model's real
   context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx),
   but ONLY for a model llama-swap's own /v1/models response already marks
   status.value === 'loaded' - never an unloaded one, since llama-swap
   treats ?model= as a routing hint and probing an unloaded model risks
   triggering an actual, slow, GPU-swapping load as a side effect of
   read-only discovery. A server with no status field at all gets no
   enrichment rather than a guess; a model not probed this round keeps its
   previously-learned value until it disappears from the list entirely.
   Stored per model (CustomModelHost.modelContextLengths) and applied via a
   new contextLengthVar registry field, set to
   CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude.

Both new fields live on the existing env-kind customModelInjection
capability shape, declared only on claude's registry entry - every other
CLI's injection is unaffected (pinned by test).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 12:08:26 +08:00
DevvynandClaude Sonnet 5 5c25a52f95 fix(custom-model): wait for a freshly launched session to go idle before applying
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.

Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.

Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 10:38:37 +08:00
DevvynandClaude Sonnet 5 409a6e65f9 fix(custom-model,toast): surface the real apply error, and make error toasts sticky with a close button
Two related fixes, both needed to actually diagnose 'Session started on
the native backend — could not apply the custom endpoint' reports from
live testing:

1. runCustomModelEntry()'s apply call went through _apiJson(), which
   unwraps a success body but SWALLOWS a failure response entirely and
   returns null — discarding the one thing (error, errorCode) that would
   tell 'endpoint unreachable' apart from 'not a discovered model',
   'remote/Docker session', or a dozen other real causes the apply route
   already reports distinctly. Switched to _api() so the actual response
   body is read on failure too, and the toast now includes the real
   message.
2. showToast() defaulted every toast, error or not, to a 3s auto-dismiss
   with no way to read it again — exactly what made the above generic
   message impossible to act on even before the fix above. Error toasts
   now default to sticky (duration: 0, no auto-dismiss) unless a caller
   opts into a duration, and every toast — sticky or not — gets an
   explicit close (x) button, since a sticky toast with no way to
   dismiss it would just accumulate across repeated failures.

Tests: custom-model-run-menu-ui.test.ts's two apply tests updated for the
_api() switch (their mocks previously stubbed _apiJson, which the apply
call no longer goes through), plus a new test pinning that the real
server error string reaches the toast on a failure.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 09:55:03 +08:00
DevvynandClaude Sonnet 5 9a9e542a7d fix(custom-model): bound the model-picker dialog's height and make its list scroll
The dialog had no max-height at all, so an endpoint with many discovered
models grew it past the viewport with nothing to scroll — reported live as
both "takes up the full page" and "the list is truncated", which turn out
to be the same bug. Gives #customModelPickModal .modal-content the same
bounded-height + scrollable-body shape cronModal's .modal-lg already uses
(max-height + flex column on the content, overflow-y:auto + flex:1 on the
body), scoped by id rather than folded into the shared .modal-sm class
three other modals already use for short, fixed content.

max-height: min(70vh, 520px) scales with the viewport (a phone gets 70% of
its height; a 4K display never gets a needlessly tall dialog) rather than
committing to one fixed pixel value that would be wrong at either end.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 09:23:14 +08:00
DevvynandClaude Sonnet 5 5a9ff07f57 feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real
llama.cpp server:

1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to
   apply the endpoint's defaultModelId (or the first discovered model)
   silently. Now, via the new selectCustomModelEntry() (session-ui.js):
   - exactly one discovered model launches straight away, same as before
   - two or more open a new #customModelPickModal listing every discovered
     model; defaultModelId (if set) is marked but never auto-chosen, since
     the point of asking is letting ONE launch deliberately differ from
     the saved default, not just confirming it
   The endpoint is re-fetched at click time rather than trusting anything
   cached from the dropdown's own render, since the model list can have
   changed (the sweep below, or a settings-panel edit) since it opened.
   runCustomModelEntry() itself — the actual launch, routed through run()
   for the in-flight lock, snapshot-guarded against applying to the wrong
   session — is unchanged; it now just always receives an explicit model
   id from one of these two paths instead of computing one itself.

2. Periodic re-discovery. Every saved endpoint's models now refresh
   automatically every 5 minutes in the background
   (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way
   as the Codex plan-usage poll it sits beside — this.cleanup.setInterval,
   off under testMode), so a model the server starts or stops serving shows
   up without another manual "Discover" click. The manual POST
   .../discover-models route and the new refreshAllCustomModelHosts()
   sweep (custom-model-routes.ts) now share one pure merge step
   (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId
   that no longer appears) rather than two copies that could drift. The
   sweep is best-effort per host — one endpoint being unreachable on a
   cycle never blocks the others — and re-reads the store before each
   host's write, keyed by id, so a concurrent edit or delete from the
   settings panel always wins over a sweep that started before it.

Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated
file for the sweep (kept separate from custom-model-routes.test.ts because
that file's data dir is shared across every test in it — one temp HOME per
FILE, not per test — which would make a sweep-touches-every-host assertion
meaningless there). test/custom-model-run-menu-ui.test.ts gained a new
describe block driving the real picker modal through JSDOM: single-model
bypass, multi-model dialog with the default marked-not-chosen, picking a
row closes the modal and launches with that exact model, the endpoint
re-fetch, and the two "vanished by click time" toast paths.

Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md,
docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated
— the last of these also caught up two sentences that had gone stale after
the draft-review fixes landed (the picker routes through run() now, not a
raw run*() call).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 08:50:18 +08:00
DevvynandClaude Sonnet 5 60e1bd52f7 fix(custom-model): act on the draft review — unparseable onclick, unwrapped envelope, wrong-session apply, missing lock, no tests
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.

Blockers:

1. Every generated inline onclick was unparseable. JSON.stringify's own
   double quotes terminated the double-quoted HTML attribute at the first
   one, leaving btn.onclick null on every picker entry and every Discover/
   Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
   argument, the same idiom deleteCase's onclick already uses four lines
   away in session-ui.js. This also closes the live-HTML-injection route
   through modelId (server-controlled, from the endpoint's own /v1/models
   reply): with quoting intact, a `>` inside it can no longer terminate the
   <button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
   like every other /api route (server.ts's preSerialization hook applies
   to arrays too), so Array.isArray(hosts) was always false in production
   and the picker/settings panel silently saw nothing. Both call sites now
   go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
   returns normally without ever changing activeSessionId, so the apply
   step used to silently re-point and restart whatever session the user was
   already looking at. runCustomModelEntry() now snapshots activeSessionId
   before the launch and requires it to have actually changed.

Majors:

4. Routes the launch through run() itself via a temporary _runMode swap
   (never persisted — setRunMode() would sync it to the server) instead of
   a parallel hardcoded dispatch table, so a custom-model launch now holds
   the same _runInFlight lock every other Run click gets. This also
   resolves the "hardcoded runners map contradicts the PR's own design"
   minor: dispatch is run()'s own, so a CLI whose customModelInjection
   recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
   against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
   parses markup this module generated itself) for exactly the DOM-level
   facts the review said needed no Playwright and no tmux: a generated
   button's onclick genuinely compiles and fires, a dangerous modelId never
   produces a live element, the envelope unwrap works, the session-changed
   guard holds, run() actually gets called (proving the in-flight lock
   engages), and _runMode is restored afterward. Confirmed against the
   pre-fix code first (reproduces btn.onclick === null exactly) so this
   isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
   render-index-html.test.ts for the other fixes below.

Minors:

- Generated entries now filter through isCliAvailable(), matching
  _refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
  (applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
  and to settings-modal open) instead of always rendering; the endpoint GET
  no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
  redactApiKey() replaces the field with a computed apiKeySet: boolean, and
  a PUT with no apiKey now keeps the stored one server-side
  (applyStoredApiKey()) instead of the client resending a value it was
  never given. New tests cover both directions (kept vs. replaced) by
  observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
  (_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
  event, since the real role can resolve after settings were first opened)
  — endpoint writes were already admin-only server-side, but the button
  used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
  toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
  route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
  that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
  (CLAUDE.md already records that exact literal turning the settings
  preview into a grey slab on light skins), .run-mode-custom-models gets
  the same gap: 2px .run-mode-menu's own flex gap only applies one level
  up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
  </script> (CliEntry.label is user-clis.json-settable, unlike
  __codemanCliAvailable's booleans-only payload) via a new exported
  escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
  docs/api-reference.md; left the "no zh-CN for the new Models-section
  group" minor unaddressed only insofar as the wider Models section (task
  routing, thinking effort, etc.) has never had zh-CN coverage either —
  everything this PR itself introduces (labels, hints, button text, the
  Run-menu's "Custom Endpoints" header) IS translated in i18n.js.

Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
DevvynandClaude Sonnet 5 fed6582d3e fix(test): strip the custom-model Run-menu picker's injected script too
CI on PR #430 failed test/server-index-title.test.ts's byte-identity
check: renderIndexHtml now injects a second unconditional <script> before
</head> (window.__codemanCustomModelClis, added alongside the existing
__codemanCliAvailable one), and the test only knew to strip the older one
before comparing the rendered HTML against the raw template.

Strip both. Unlike __codemanCliAvailable (an object, historically injected
only where something resolved), the new one is a plain array injected
unconditionally, possibly empty, so it needs stripping on every machine,
not just one with CLIs installed.

Verified the two replace() calls compose correctly against the exact
strings server.ts actually produces (simulated in isolation; this box has
no tmux, so the real WebServer-backed test file cannot run here at all --
same environment gap noted throughout this PR's review).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
DevvynandClaude Sonnet 5 98d26e14d9 docs(wiki): document Custom Model Endpoints and the Run-menu picker
New docs/wiki/Custom-Model-Endpoints.md (auto-synced to the live GitHub
wiki on push to master, per docs/wiki/Contributing.md) covers turning the
feature on, adding an endpoint, the Run-menu picker's one-off-run
behaviour, the per-harness confidence table, and what it deliberately does
not do yet (remote/Docker sessions, live hot-swap). Linked from the
sidebar, from Agent-CLIs.md's "Read next" list plus a short pointer
section, and from Settings-Reference.md's Models section.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
DevvynandClaude Sonnet 5 25fae9ad10 feat(custom-model): generate Run-menu entries from saved endpoint profiles
Follow-up to #393, picking up the work Ark0N invited in his merge comment:
"generate those entries from the saved profiles rather than a fixed
duplicate per harness, and put it in a follow-up PR so this one stays the
backend... The Run-menu picker is yours if you want it."

Adds the frontend surface the backend has been waiting on:

- Run menu: a "Custom Endpoints" section lists one entry per (harness that
  supports customModelInjection, saved endpoint) pair, e.g.
  "Claude Code (llama.cpp)". The harness list comes from
  window.__codemanCustomModelClis, injected at page render straight off the
  CLI registry's own capabilities (never a hardcoded id list in the
  frontend), so a CLI whose injection recipe lands later appears with no
  frontend change. Picking an entry runs that harness's own existing run*()
  function unmodified (case creation, env overrides, everything, forced to
  a single instance) and then applies the endpoint's default model to the
  session it creates via the existing POST /api/sessions/:id/custom-model
  route. Entries are hidden for a remote/docker active case, since that
  route already refuses both.
- Settings: App Settings -> Models gets a "Custom model endpoints" group
  wiring up the customModelEndpointsEnabled toggle (declared since #393,
  read by nothing until now) plus CRUD against the existing
  /api/model-endpoints routes: list, add/edit (inline form), delete,
  discover models.
- Backend: CustomModelHost gains an optional defaultModelId, the model the
  picker applies with no further choice per endpoint (one generated menu
  entry per CLI+endpoint pair, not per CLI+endpoint+model). The route
  refuses a value that isn't one of the endpoint's own discovered models,
  and a fresh discovery drops a default that no longer appears rather than
  carrying an invalid one forward.

Docs: docs/custom-model-endpoints.md describes the new picker and settings
panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the
"backend-only" status note and documents the picker's generation mechanism.

Tests: four new route tests cover defaultModelId validation, acceptance,
and the drop/keep behaviour across a re-discovery; a new render-index-html
test pins the __codemanCustomModelClis injection (present, agent CLIs
supporting the capability, antigravity and shell excluded) and its
solo-window skip. No browser test was added for the Run-menu picker itself
or the settings CRUD panel (this box has no tmux, so the live server used
by test:browser/test:mobile could not be exercised here) -- worth a
Playwright pass before merge, same as any other frontend PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
Randalix 4a30f510e6 fix(remote): let the host config turn wake-on-LAN OFF for a live session too
Found by driving the real UI: with a MAC configured in remote-hosts.json, removing it
(here: to reach the "Configure WoL" dialog) changed nothing for a running session —
_effectiveRemote short-circuited on the session's own snapshot whenever that snapshot
HAD a target, so the resolver was only ever consulted in the one direction where the
feature was missing. The documented "host config is authoritative" promise therefore
failed in the direction a user can actually observe, and a wake target could live on
invisibly after being deleted from the config.

The resolver is now consulted on the TTL regardless, and wins for the wake fields in
both directions. Also adds a route test for the browser's real input shape: one POST
per keystroke, all buffered during a wake, replayed IN ORDER.
2026-09-15 23:21:50 +02:00
Randalix d0a5a583cd feat(remote): wake a sleeping host when a session is created or attached
Pressing Run on a remote case whose host was asleep failed with
`could not verify tmux on remote host 192.168.50.137: …` — an ssh error that
blames tmux for a machine that is merely suspended. The only wake paths were
typed input on an established session and the banner's Wake button, so OPENING a
session (the moment the user actually decides to use that host) had none.

`RemoteWakeRegistry.ensureHostAwake()` reuses the existing probe/wake/readiness
machinery for a host that has no session yet, and is wired into the two
user-initiated create paths: `POST /api/quick-start` for a remote case (before
the tmux prereq probe, which is what surfaced the misleading error) and
`POST /api/sessions` with `attachRemoteSession`. A host without a wake target is
not even probed, so its behavior and latency are byte-identical. The wake is
blocking — the caller gets the session or an error — but bounded by
REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS (40 s) instead of the 90 s session default,
because the dashboard sits behind a reverse proxy whose default
`proxy_read_timeout` is 60 s: a longer wait would be cut off at the proxy while
the session was still being created. The budget has to cover the whole request
(40 s wake + 1.5 s probe + the tmux probe's own 15 s = 56.5 s worst case), which
is why it is 40 s and not 45. A timeout now says the host did not come back, and
an unreachable host without a wake target says so instead of pointing at tmux.

The wiring is deliberately in the HTTP ROUTE, never in the shared session
service: `cron-service.ts` builds sessions there with nobody waiting on the
answer, and a wake on that path would power the host on for every schedule —
the timer-driven re-wake invariant #1 exists to prevent. Both halves are asserted
(importers of `remote-wake`, and `ensureHostAwake` having exactly one caller
file), so a future caller has to come through the guard test. A rejection from
the wake IO is caught too: a broken target must fail the wake, not the route.

`remote:hostWaking`/`remote:hostWakeFailed` now carry `forNewSession` for the
session-less case, where "input is queued" would be untrue; the toast then reads
"the session starts when it is back".

Live wake numbers are unchanged (this reuses the measured ~12 s S3 path); the
route behavior is covered by new tests in session-routes.test.ts with an injected
registry, so no test opens a real socket or ssh.
2026-09-15 22:37:37 +02:00
Randalix 8dfc965d13 fix(remote): stop the wake handlers shadowing each other; enforce the input cap
Two findings from a final review pass over the wake-on-LAN feature.

`_onRemoteHostWaking` / `_onRemoteHostWakeFailed` were defined in BOTH
`panels-ui.js` (toasts) and `host-wake-ui.js` (banner). Both files mix into
`CodemanApp.prototype` and `host-wake-ui.js` loads later, so the panels-ui copies
were silently shadowed: the toast never fired, and a wake started for a BACKGROUND
session (input on a non-active tab) produced no notification at all, since the
banner handler only acts on the active session. The handlers now live only in
`host-wake-ui.js`, show the toast unconditionally, and update the banner when the
woken session is the active one.

`appendBoundedPending` dropped only WHOLE chunks, so a single input value over the
cap (one large paste is one `input` value, up to the 100 KB input schema) was kept
in full: "bounded at 4 KB" held per chunk, not per session, and nothing was logged.
The surviving chunk's head is now trimmed too, code-point aware so a multi-byte
character is never split into a replacement char.

Adds the guard that would have caught the first one: every SSE dispatch handler must
be defined in exactly ONE frontend module. The existing test only asserts a handler
EXISTS somewhere, which two modules both satisfy while one is shadowed.
2026-09-15 21:01:53 +02:00
Michael GrundbergandClaude Opus 5 c9515b1d4c fix(terminal): keep the output a pane capture could not contain
Live terminal events are queued while a buffer load runs, and the load discards
that queue when it ends. That is right when the loaded buffer is the server's
accumulated byte history. The route appends to that history right up to the
moment it serializes the response, so a queued event already appears in it and
replaying it would duplicate output, most visibly Ink's cursor-up redraws.

A tmux pane capture is a photograph, current only as of the instant
`capture-pane` ran. Output printed afterwards was queued and then dropped, and
nothing scheduled a re-fetch to recover it: `_onSessionNeedsRefresh` is wired
only to the 128KB overflow path. The CLI's next partial redraw then landed on a
frame the terminal never received.

How much went missing depended on which capture the route served. A `?full=1`
load returns the capture alone, with no history in front of it, so it lost
everything from the capture to the end of the chunked write. A `?tail=` load
returns history, a clear, and then the capture, and the route reads that history
after the capture, so it lost everything from the response to the end of that
write. The chunked write dominates either way. An agent CLI hides the loss on
its next full redraw; a shell session does not, because its output is linear and
nothing repaints it.

Queue entries now carry their arrival time, and `_finishBufferLoad` takes a
`since` cutoff, so a capture load replays exactly the tail that arrived after
the response headers. The earlier events stay dropped, because a payload that
carries history does hold those.

All four paths that fetch a terminal buffer and write it now decide this the
same way, through one `_bufferLoadFinishOpts` helper, so they cannot drift
apart: `selectSession`, `_onSessionNeedsRefresh`, `_onSessionClearTerminal` and
`_maybeRefetchFullHistory`. The second of those is the one that stings. It
exists to restore output the client already dropped once under backpressure, and
it was dropping more output while performing that recovery. The cache-hit write
inside `selectSession` stays on discard deliberately: it runs before the fetch,
so its queue holds only events the capture that follows already contains.

Two further things had to change for that tail to still exist when the load
ends, and a browser test is what found both. `chunkedTerminalWrite` is what ends
the load for every non-empty buffer, so the flush policy travels to its own
finish calls; the call in `selectSession` runs only when the write was skipped.
`_beginBufferLoad` no longer empties the queue when one load re-enters it, which
it does on every write, because that reset discarded the whole fetch window
before anything could replay it.

The response already distinguishes the sources. `source` reads `mux-visible` or
`mux-full-history` for a capture and `history` for the byte stream.

Follows #395, #396 and #397, which fixed the ways the replayed frame itself
could disagree with the terminal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-15 18:14:18 +02:00
Randalix 1380b023e2 fix(remote): make the wake banner's poller page-wide and independent of tab switches
Reported as 'the tab shows no banner' while the host was verifiably unreachable: the
banner only started polling from selectSession, which RETURNS EARLY for the tab you
are already on (so a page loaded with the remote tab active never polled), and a
long-lived tab keeps running the JS it loaded — the feature was invisible to anyone
who did not switch tabs after the deploy.

The poller is now page-wide: one interval (created on init and on the first session
switch), re-targeted whenever the active session changes, plus a visibilitychange
wake-up. It no longer depends on any single selection path running.

Also adds test/sse-dispatch-table.test.ts: a static guard that every
[SSE_EVENTS.X, '_onFoo'] entry names an event constants.js defines AND a handler some
module defines. Both halves fail silently (a typo'd constant is an undefined table
key; a renamed handler just never runs), which is exactly how a new banner can never
appear with no error anywhere.
2026-09-15 15:11:36 +02:00
Randalix e8f7772320 fix(remote): offer the WoL config dialog after a failed wake too
A configured-but-broken target (host replaced NIC, command removed) had no way
out: the dialog hung off the 'no target configured' branch only, so the banner
would keep offering a Wake button that keeps failing.
2026-09-15 14:38:24 +02:00
Randalix 2f61be6e74 fix(remote): bind the wake socket before enabling broadcast
setBroadcast() on an unbound dgram socket throws EBADF on Linux and the following
send fails with EACCES, so the magic packet silently never left the machine — the
feature reported a wake that never happened. Caught by waking a real sleeping host
(a unit test with a real UDP broadcast would not be welcome in CI, so the socket is
injectable and the bind-before-setBroadcast ORDER is asserted).
2026-09-15 14:24:13 +02:00
Randalix 8b5a13435a feat(remote): host-unreachable banner, manual wake, and native MAC wake-on-LAN
The reactive wake (typing into a session whose host slept) left the state invisible:
nothing told the user the machine was asleep, and with no wake target configured
there was nothing to do about it. Adds:

- RemoteHost.wakeMac (comma-separated) - Codeman builds and broadcasts the magic
  packet itself (UDP port 9), so the common case needs no external script. The
  existing wakeCommand stays as the explicit override.
- GET /api/sessions/:id/reachability - probes (throttled, cached, and it never
  wakes) and reports HOW the host can be woken, or that nothing is configured.
- POST /api/sessions/:id/wake - wakes, waits, reattaches the pane and flushes
  buffered input; 400 with a routable message when no target is configured.
- The amber host-unreachable banner + its 'Wake' / 'Configure WoL' action, and a
  small config dialog that saves via PUT /api/remote-hosts/:id.
- RemoteWakeDeps.resolveRemote: host config is re-resolved for LIVE sessions
  (throttled + cached), so saving the dialog takes effect without a restart.
2026-09-15 14:20:31 +02:00
Randalix 3f0bfde54a docs(remote): document the wake-on-LAN invariants; drop wake state on bulk delete
Self-review pass: the input-ladder's two 'buffer' branches were the same three
lines, and bulk delete left a session's (bounded, per-random-uuid) wake state
behind. Documents the design where the code refers to it - remote-sessions.md
section, the architecture invariant, and the CLAUDE.md key pattern.
2026-09-15 10:45:01 +02:00
Randalix 0f3eea2fb5 fix(remote): refresh wake command from host config when restoring sessions
A session's remote block is persisted at launch time and recovery uses that
snapshot, so a wakeCommand added to remote-hosts.json afterwards never reached
an already-running session - not even across a Codeman restart (observed: the
live Hufflepuff session came back with no wakeCommand). Merge the host-level
field in on restore, with the host config authoritative.
2026-09-15 10:35:58 +02:00
Randalix a81f430e41 feat(remote): wake a sleeping host from user input (Wake-on-LAN)
A durable remote session survives SSH drops (COD-104/108), but nothing brought
the HOST back: after the remote machine suspended, the local tmux pane's ssh
child stalled silently and `send-keys` SUCCEEDS against it, so typed input
vanished with no error anywhere.

Add an optional per-host `wakeCommand` (Wake-on-LAN wrapper, e.g. whuff) that
the input route runs when a wake-enabled host is unreachable: input is buffered,
the host is woken, the pane is reattached, and the buffer is flushed in order.
Detection is a throttled bare TCP probe on wake-enabled hosts only, and only
REAL user input may wake a host - the auto-reconnect watcher and boot recovery
deliberately cannot, or the host would be re-woken seconds after every suspend
and could never stay asleep.
2026-09-15 10:25:52 +02:00
84 changed files with 14699 additions and 452 deletions
+1 -1
View File
@@ -10,7 +10,7 @@
"name": "codeman",
"source": "./plugins/codeman",
"description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.30.0",
"version": "1.31.0",
"author": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
+108
View File
@@ -1,5 +1,113 @@
# aicodeman
## 1.31.0
### Minor Changes
- 035bfbc: feat(remote): wake a sleeping remote host from Codeman
A remote SSH case pointing at a machine that suspends used to fail the same way every
time: the session was there, the host was not, and typing into it went nowhere. A host
can now carry a wake target, either a MAC address for Wake-on-LAN (Codeman builds the
magic packet itself, so nothing reaches a shell) or a wake command of your own, and
Codeman uses it when you ask for the host: when you type into a sleeping session, when
you press the wake button on the banner, or when you start or attach a session on that
host. Input you type while it wakes is buffered and flushed once it is back, up to 4 KB,
and a chunk over that is refused outright rather than delivered as a fragment.
Waking only ever happens because you asked. No watcher, dropped-session handler or
boot-recovery path can reach it, since a machine woken by a reconnect watcher would come
back seconds after every suspend.
- fbee1b2: feat(custom-model): pick a custom endpoint straight from the Run menu
#393 landed the backend for custom model endpoints and left it reachable only over the
HTTP API. This is the rest of it. Turn on Custom model endpoints in App Settings, save
an endpoint, and the Run dropdown grows a Custom Endpoints section built live off the
CLI registry, one entry per harness that can actually redirect plus each endpoint you
saved. Pick one and it launches that harness pointed at your server, asking which model
first when the endpoint has more than one. Endpoints re-discover themselves every five
minutes, and one unreachable endpoint never blocks the others. App Settings gains full
add, edit and delete for endpoints.
Seven of the harnesses (opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP) now launch
directly onto the endpoint with no restart at all, where before you watched a native
boot followed immediately by a second one. Claude still launches and then restarts in
place, which its own resume makes far less jarring.
Most of this release's work went into things that only show up against a real server,
and each was found that way rather than in tests: a freshly launched CLI reporting
itself busy for its own startup and getting refused; Claude Code assuming a large
context window for a model it does not recognise and silently overflowing a small one;
a model whose real context is below what Claude Code's own system prompt costs, which
no setting can fix and which now warns before launching into a certain failure; and the
big one, llama.cpp running exactly one model at a time, so applying a selection can
unload the model another session is using. That last case now asks first, tells you
which session it affects, and keeps a "loading model" notice on screen for the whole
swap window, so a prompt sent mid-swap reads as loading rather than as an answer from
whatever was loaded a moment ago. A background sweep also catches the reverse: your
session's model being evicted later by somebody else's ordinary use.
Two things worth knowing if you drive this over the HTTP API or run multi-user. The two
questions an apply can ask (the model's context window is too small, and loading it will
unload the model another session is using) are now answered by separate
`confirmedContext` and `confirmedSwap` fields rather than one `confirmed`. They shared a
flag until now, and since the context check runs first, confirming that one silently
agreed to evict another session's model as well. The old `confirmed` still means both.
And `CLAUDE_CONFIG_DIR` is now admin-only in multi-user mode: it joined claude's
privileged env keys, so a non-granted owner can no longer set it through `envOverrides`,
and an already-persisted one is dropped on reboot-restore, which returns that session to
the default Claude account rather than the per-client one it was pointed at. Single-user
installs are unaffected.
Remote SSH and Docker sessions are refused for now, since their restart reattaches a
durable tmux rather than relaunching the agent.
### Patch Changes
- c9515b1: fix(terminal): keep the output a pane capture could not contain. Opening a session, a backpressure refresh, a clear-terminal reload and a full-history re-pull all load the screen from a tmux pane capture, and anything the CLI printed between that capture and the end of the load used to be dropped, so its next partial redraw landed on a frame the terminal had never seen: missing or garbled output right after a tab switch or a refresh, plainest in a shell session. Each load now replays exactly the output that arrived after the capture, through one shared rule for all four paths, and a refresh that restores your scroll position no longer snaps back to the bottom afterwards.
- 3edf9aa: fix(terminal): replay a pane capture at the geometry it was taken at
Opening a session could draw a frame built for a pane bigger than your terminal. A
taller pane wrote its overflow rows onto the last line and lost the rows underneath
(against a 50-row pane, a 30-row terminal rendered 28 of a 45-line command and drew
the survivors twice), and a wider one wrapped every row and scrolled the whole frame
up by one. The terminal response now reports the geometry the capture was really
taken at, so the browser can see the mismatch and replay once at the size that stuck.
A pane that cannot be sized to fit is diagnosed once per session instead of on every
tab switch.
- 035bfbc: ### Thanks
- @irisitymichaelgrundberg for three terminal fixes in one release: keeping the output a pane capture could not contain (#436), replaying a capture at the geometry it was taken at (#435, five rounds and a Playwright suite that fails against the merge base), and trimming the padding out of a copied selection (#451), where the scan-instead-of-regex call avoided a 2.9s freeze nobody would have traced back to a copy.
- @timkjr for a first contribution that found a real silent failure: the Instance count stepper next to the Run button had only ever applied to Claude, so on the other eight run modes it launched one session and said nothing (#454).
- @Randalix for Wake-on-LAN on remote hosts (#439), built and live-tested against a real sleeping machine, and for reading the whole diff again between rounds rather than only the parts that were asked about.
- @opticon454 for turning #393's backend-only custom model endpoints into the whole feature (#430), and for validating it against a real llama-swap box rather than against the tests: the `/props` versus `/running` context discrepancy and the DeepSeek `/v1` root cause were both tracked down to the SDK source instead of guessed at.
- c376534: fix(run): make the Instance count stepper work for every non-Claude mode
The Instance count stepper next to the Run button only ever applied to Claude.
Setting it to 3 and launching OpenCode, Codex, Gemini, Antigravity, Pi, OMP, Grok or
DeepSeek started exactly one session, with no error and no hint that the control had
done nothing. All eight now launch the count you asked for, and the opening banner
says how many are starting. The one exception is a launch started from the Custom
Endpoints section of the Run menu, which always starts a single session.
- 19ffe9b: fix(input): make sure a prompt sent through the API actually leaves the composer. Claude Code 2.1.277 started ignoring Enter for the first 30 to 50 seconds after the composer paints while still accepting the typed text, so a prompt sent right after a session came up sat unsent in the pane and every waiter (send-and-wait, the agent skill, cron, the maintainer bot) burned its whole timeout on a turn that never started. The server now reads the pane after every programmatic write that carried Enter and presses Enter again, on a 2 to 60 second schedule, only while the composer verifiably still holds the text it sent; an empty composer, other text, or a pane with no composer at all ends it. The agent skill's `sendwait` gets the same loop for servers that predate this, and its preamble version moves to 1.30.1 so an already-seeded agent picks up the fresh copy.
- f9edb33: fix(terminal): trim the padding out of a copied selection
Copying out of a pane put a wall of spaces on the clipboard. xterm hands back
whole screen rows and trims only the cells that were never written to, so the
real spaces a full-screen program paints across the unused part of a row count
as content: measured against Claude Code in a 282-column pane, single lines
arrived carrying 138 trailing spaces. Pasting that into a chat client or an
editor meant deleting the whitespace by hand, while Windows Terminal, iTerm2 and
GNOME Terminal all trim it for you. A copy now drops the trailing run from every
line, on all four paths (the Ctrl+C chord, right-click, the phone selection
button and Auto Copy), while leading indentation is left exactly as it is. An
Alt+drag rectangular selection is copied verbatim, because its columns lining up
is the point of that gesture. A selection holding nothing but padding is refused
rather than copied as bare line breaks.
## 1.30.0
### Minor Changes
+16 -12
View File
File diff suppressed because one or more lines are too long
+2
View File
@@ -27,6 +27,8 @@ export const BROWSER_TEST_GLOBS = [
'test/webgl-fallback.test.ts',
'test/terminal-copy-shortcut.test.ts',
'test/terminal-keycode229-recovery.browser.test.ts',
'test/capture-load-window.browser.test.ts',
'test/capture-geometry-retry.browser.test.ts',
'test/codex-predictive-echo.test.ts', // also needs a real codex binary
];
+133
View File
@@ -324,6 +324,30 @@ from the session's current state rather than requiring a new transition: the
original turn may be long over. It comes back as
`"delivered": false, "duplicate": true`.
**Wake-on-LAN hosts** (`docs/remote-sessions.md` §Wake-on-LAN): when the session's
remote host has a wake target and is asleep, the non-wait form answers `200` with
`{"buffered": true}` — the bytes are held and flushed after the host is back — or
`{"buffered": true, "dropped": true}` for a chunk over the 4 KB wake buffer, which
is gone (never delivered as a fragment). Both fields are additive to the historical
bare `{}`. With `wait`, the route blocks on the wake instead and answers
`422 OPERATION_FAILED` ("did not come back after a wake-on-LAN request — nothing was
sent") when the host never returns, rather than writing into the stalled pane and
reporting `delivered:true` plus a timeout.
Two endpoints back that flow directly, both scoped to one session's remote host and
both refusing a session that is not remote (`400 INVALID_INPUT`):
| Method | Path | Purpose |
| --- | --- | --- |
| `GET` | `/api/sessions/:id/reachability` | Whether the session's remote host answers SSH right now, plus whether a wake target is configured. Read-only: it never wakes. `{"reachable": true\|false\|null, "wakeConfigured": "mac"\|"command"\|"none"}`, where `null` means the answer is unknown (a proxied host, where a TCP probe proves nothing). |
| `POST` | `/api/sessions/:id/wake` | Wake the host and wait for it to accept SSH again, bounded by the request budget. `422 OPERATION_FAILED` when it does not come back; `400 INVALID_INPUT` with "No wake-on-LAN target configured for this host" when nothing is set. |
⚠️ Waking is deliberately reachable only from an explicit user action (this route, a
session create/attach, or typing into a sleeping session). No watcher, dropped-session
handler or boot-recovery path may wake a host, or a suspended machine would be woken
again seconds after every suspend; `test/remote-wake.test.ts` pins that as an import
fence around `src/remote-wake.ts`.
### Response
All three nest the wait result under `data.wait`, so one client helper works against
@@ -558,6 +582,115 @@ All four enforce session ownership in multi-user mode; a foreign session id
answers `404 NOT_FOUND` (no existence leak), and profiles of two owners of the
same directory are distinct by construction.
## Custom Model Endpoints
Points a session's harness at a user-configured OpenAI-compatible endpoint —
local (llama.cpp, vLLM, DGX Spark) or cloud (Azure AI Foundry, OpenRouter) —
instead of its native cloud backend, gated by the opt-in
`customModelEndpointsEnabled` setting (default OFF). Endpoints are
machine-level infra, like remote/docker hosts: writes are admin-only in
multi-user mode. Design: [`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md);
user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md).
- `GET /api/v1/model-endpoints` -> `CustomModelHost[]`, an unwrapped bare
array like every other list route (still riding the standard `{success,
data}` envelope on the wire — unwrap it the same way). Answers `[]` for a
non-admin in multi-user mode. `apiKey` is never returned; `apiKeySet:
boolean` reports whether one is stored, so a client can render "unchanged
if left blank" without ever holding the real value.
- `POST /api/v1/model-endpoints` with `{ id, label, baseUrl, apiKey?,
authStyle?, defaultModelId? }` creates one. `id` must match
`^[a-zA-Z0-9_-]+$`; `authStyle` is `bearer` (default) or `api-key`, never
both (a real server hung indefinitely when sent both headers on one
request); `baseUrl` must be `http(s)`, carry no embedded credentials, and
is refused if it points at (or resolves to) a link-local or
cloud-metadata address. `409 ALREADY_EXISTS` on a duplicate id.
- `PUT /api/v1/model-endpoints/:id` updates one. An **absent** `apiKey`
keeps the stored one rather than clearing it — the client never receives
the real value to resend deliberately unchanged, so omission is the only
way to say "leave it alone"; there is no way to clear a key back to unset
this way. `defaultModelId`, when set, must be one of that endpoint's own
`models` (`400 INVALID_INPUT` otherwise).
- `DELETE /api/v1/model-endpoints/:id` removes one.
- `POST /api/v1/model-endpoints/:id/discover-models` fetches the endpoint's
own `GET /v1/models` and stores the result as `models`, updating
`lastDiscoveredAt`, plus (best-effort, only for a model llama-swap's own
response already reports loaded) `modelContextLengths` and `modelSizesGB`.
A `defaultModelId` that no longer appears in the fresh list is dropped
rather than carried forward invalid. Failures answer `422 OPERATION_FAILED`
with the underlying connection error, or a named egress refusal if the
resolved address turned out to be blocked. The same refresh also runs
automatically for every saved endpoint every 5 minutes in the background
(`refreshAllCustomModelHosts()`, `custom-model-routes.ts`, started from
`server.ts`), so there is no route for triggering "refresh all" — one
endpoint being unreachable on a cycle never blocks the others.
- `GET /api/v1/model-endpoints/:id/running-status` -> `{ isLlamaSwap,
running: [{model, state}], logLine? }`, read-only, no admin gate
(any session owner who could already point a session at this endpoint can
equally ask what it currently has loaded). `isLlamaSwap` is
feature-detected via the endpoint's own `GET /running` — a plain
llama.cpp/OpenAI-compatible server has none and always answers `false`.
`logLine`, present only when `isLlamaSwap` is true, is the most recent
REAL backend `llama-server` process log line (`load_model: ...`,
`llama_server: model loaded`, etc.), sourced from the endpoint's own
`GET /api/events` SSE stream and filtered to `source: "upstream"` frames
only (never llama-swap's own `source: "proxy"` request-access log) — one
connection is held open per endpoint and reused across every poller,
idle-closed after 30s of nobody asking. This is what the Run-menu
picker's loading banner polls once a second while a model is loading.
- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId,
confirmed? } | { clear: true }` applies (or clears) the session's
selection and **restarts the session's CLI process in place** — every
supported harness reads its endpoint config at process start, never per
turn, so there is no live hot-swap. (`POST /api/v1/quick-start`'s own
`customModel: { endpointId, modelId, confirmed? }` field is the
no-restart equivalent for a session that doesn't exist yet — see below.)
A Claude session resumes its existing conversation across the restart;
pi/omp/grok additionally get a forced `--model`/`-m` value, since for
those three the config file alone does not select it. `400 INVALID_INPUT`
for a remote (SSH) or Docker session — both restart their agent
differently under the hood, and applying to one would report success
while changing nothing. Two more responses replace the normal
`{customModel, restarted}` shape, neither an error, and neither restarts
or creates anything on the first ask. ⚠️ **Each is answered by its OWN
flag on the retry, and answering one is not consent to the other**: they
are questions about different people, and while they shared a single flag
a caller who confirmed the context warning silently agreed to evict
another session's model as well. Send `confirmedContext: true` to proceed
past the context warning, `confirmedSwap: true` past the swap conflict,
and both when both were asked (they accumulate, so the second retry still
carries the first answer). The original `confirmed: true` still means
BOTH and is still accepted, because it shipped in this feature's
HTTP-API-only cut; new callers should send the specific one:
- `{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}` —
llama.cpp/llama-swap only runs one model at a time, and switching would
unload a model another **live session's own selection** is actively
using. Never returned for a plain (non-llama-swap) server, and never
just because a swap is needed at all — only when it would disrupt
someone else.
- `{requiresContextWarning: true, modelId, contextLength,
minSafeContextTokens}` — Claude Code's own fixed per-turn overhead
(system prompt + tool schemas) can exceed a small model's entire
discovered context on its own, before any conversation history exists
to compact, guaranteeing the very first message fails regardless of
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`. Gated on the CLI registry declaring a
`contextLengthVar` (claude only today), so it never fires for another
harness.
- `POST /api/v1/quick-start`'s `customModel: { endpointId, modelId,
confirmed?, confirmedContext?, confirmedSwap? }` field (alongside its
normal `caseName`/`mode`/etc. body)
computes the same injection **before** the session exists and launches
directly on the endpoint — no restart, because there was never a
native-backend boot to restart away from. Runs the identical checks as
the dedicated route above (`requiresConfirmation`/`requiresContextWarning`,
same shapes, same per-question `confirmedContext`/`confirmedSwap` retry),
and is refused the same way
for a remote or Docker case. This is what the Run-menu picker uses for
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP; Claude still uses the
dedicated restart route above (its `--resume`-based restart is far less
jarring than a full relaunch, and folding it into the one-shot path is
separate work — see `docs/custom-model-endpoints-plan.md`).
## Voice dictation
Browser dictation transcribed through this server's Claude Code login, i.e. the
File diff suppressed because one or more lines are too long
+25 -13
View File
@@ -104,17 +104,17 @@ declared capability, never an `if (mode === 'claude')` branch.
## Per-CLI injection recipes (confidence-ranked)
| CLI | Mechanism | Confidence |
| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Confirmed reaching the server, but failing — unresolved.** A real run against llama-swap returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` consistently (confirmed the env vars are read: the request reaches the network rather than failing locally). Root cause not identified — plausible explanation by analogy with codex's Responses-API gap is that `dsh --profile headless` expects DeepSeek's official API response shape/path structure rather than a generic OpenAI-compatible `/v1/chat/completions` endpoint, but this was not confirmed by reading dsh's own bundled source (unlike pi/grok, where that grep resolved the question directly). Documented as best-effort/unknown, not shipped as verified working |
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
| CLI | Mechanism | Confidence |
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol picture more nuanced than a flat break, re-verified live twice on 2026-09-17 against a llama-swap deployment that DOES answer `/v1/responses`** (an earlier test's `Reconnecting...`/`high demand` failure does not reproduce against every llama-swap setup): a plain, no-tool-call chat turn (`codex exec 'reply with just OK'`) returned a real reply. But a real tool-call attempt (`run the shell command: echo hello`) came back as an `agent_message` TEXT item — the tool-call JSON printed as the model's answer, not a `function_call` item codex would actually execute (confirmed via `codex exec --json`'s raw event stream: `item.completed`/`agent_message`, never `function_call`). Since tool execution is what makes codex a coding agent at all, this remains **not usable for real work**, just with a different, more specific failure mode than previously documented — still do not present this as working. Separately, EVERY custom-endpoint codex session also prints `warning: Model metadata for '<id>' not found. Defaulting to fallback metadata...` on launch (confirmed harmless — the successful plain-text reply above still had it): codex's per-model metadata (reasoning tiers, system-prompt templates, context-window figures) comes from `models_cache.json`, a LOCAL CACHE of OpenAI's own hosted model catalog that a custom model can never appear in by construction. No config.toml override exists for it, and the isolated `CODEX_HOME` never gets a `models_cache.json` written into it at all (confirmed: inspected a live, actively-used isolated dir — codex evidently can't reach OpenAI's catalog endpoint for this session and just falls back silently every time, with no file left behind to fix or clean up). Fabricating a fake catalog entry to suppress the warning would mean copying the _shape_ of OpenAI's own proprietary schema — including their real per-model system-prompt content, visible in a genuine `models_cache.json` — for a warning confirmed to have no effect on the actual (broken) tool-calling outcome; not worth building |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`), now with `appendV1Suffix: true` (see confidence). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Root cause of the original `HTTP_404` found and fixed, by reading dsh's own bundled source — the same bar pi/grok's fixes were held to.** Installed `@deepseek-ai/dsh` (all its real published dependencies) into a scratch directory purely to read `@deepseek-ai/dsh-llm-deepseek/lib/index.js`: it builds its request as `fetch(\`${connection.baseURL}/chat/completions\`, ...)`with`baseURL`read straight from`DEEPSEEK_BASE_URL`(or defaulting to DeepSeek's real public API root,`https://api.deepseek.com`, which also carries no `/v1`) — no `/v1` insertion of dsh's own, unlike the OpenAI-SDK convention this recipe originally assumed. llama-swap/llama.cpp only ever serves the OpenAI-conventional `/v1/chat/completions`. Confirmed live: `POST <baseUrl>/chat/completions` → `404`, `POST <baseUrl>/v1/chat/completions` → `200`, on the exact same endpoint — and dsh's own error-message template, `DeepSeek API error (HTTP ${status})`, reproduces the originally reported `dsh: HTTP_404: DeepSeek API error (HTTP 404)` precisely. Fixed by adding `appendV1Suffix` (env kind only, deepseek's entry alone — claude/gemini must NOT get it, since claude was already confirmed working against the unmodified `baseUrl`), which runs `endpoint.baseUrl` through the same `withV1Suffix()` helper `configDir`-kind CLIs already use. ⚠️ Not yet re-run end-to-end with a real `dsh` binary — no install available in this environment (no npm-installed CLI binary in `PATH`, and the `codeman-test-picker` container doesn't bundle it either); the fix is source-confirmed and live-verified at the HTTP level, but a genuine "hello world" reply through `dsh` itself is the remaining step before promoting this to **verified** alongside claude/opencode/pi/grok/omp |
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
Everything web-researched-but-unverified gets implemented but must be
smoke-tested against real installs of those CLIs before being called done —
@@ -208,6 +208,14 @@ extra per-model configuration on Codeman's side at all.
### 4. Toolbar UI
> **Superseded.** This section describes the toolbar-button design as originally
> planned. What actually shipped is a Run-menu picker instead: one generated entry
> per (capable harness, saved endpoint) pair directly in the existing `#runModeMenu`
> dropdown, rather than a separate `#customModelBtn`/`#customModelMenu` surface. See
> [`docs/custom-model-endpoints.md`](custom-model-endpoints.md#the-run-menu-picker)
> for the current design; the sections below (session-restart mechanics, security)
> remain accurate regardless of which UI calls the underlying route.
- New header/toolbar button (e.g. `#customModelBtn`, `btn-toolbar
btn-custom-model`), marker-hidden by default (`btn-custom-model--hidden`)
and revealed by `applyHeaderVisibilitySettings()` only when
@@ -344,8 +352,12 @@ pure unit tests and the live manual checks in Verification:
up automatically with zero edits to the script). Already run to
completion against the author's llama-swap server (a LAN address,
inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries):
claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected**
(Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED**
claude/opencode/pi/grok/omp **PASS**, codex **partially works and still
isn't usable** (plain chat succeeds against a llama-swap deployment that
answers `/v1/responses`, but a real tool-call attempt comes back as
inert text rather than an executable `function_call` — see the
confidence table row for the full, re-verified picture), gemini/deepseek
**UNCONFIRMED**
(reach the server, fail for undiagnosed reasons — see their table rows),
antigravity **SKIP** (no mechanism). Re-run this against a real cloud
endpoint (e.g. an Azure AI Foundry deployment) once one is available, to
+406 -18
View File
@@ -11,20 +11,22 @@ company gateway) — anything answering `GET /v1/models` and
recipe confidence table, and security reasoning:
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
> **Status**: backend is implemented and tested (registry capability, the
> injection engine, the endpoint store + discovery route, the session
> restart route). The toolbar picker / settings UI described below as the
> intended surface is **not yet built** — until it lands, use the HTTP API
> directly (examples below). Antigravity has no known custom-endpoint
> mechanism and is not supported.
> **Status**: fully wired end to end — registry capability, the injection
> engine, the endpoint store + discovery route, both the restart-in-place
> apply route (Claude) and the one-shot quick-start launch path (every
> other supported harness), a settings-panel CRUD surface, and the Run-menu
> picker described below. Antigravity has no known custom-endpoint
> mechanism and is not supported. The HTTP API (examples below) still works
> directly and is what the picker itself calls under the hood.
## Turning it on
App Settings → Agents & CLIs → **Custom Model Endpoints** (synced setting
`customModelEndpointsEnabled`, default **OFF**). Until the toolbar picker
lands, nothing reads this setting: the HTTP routes below work whether it is
on or off, and it exists now only so the picker has a switch to hang off
when it ships. The API equivalent:
App Settings → Models → **Custom model endpoints** (synced setting
`customModelEndpointsEnabled`, default **OFF**). Turning it on does two
things: it reveals the endpoint list/add/edit/discover panel in that same
settings section, and it makes the Run menu offer a generated entry per
(harness, endpoint) pair — see "The Run-menu picker" below. The API
equivalent:
```bash
curl -sk -X PUT https://localhost:3000/api/settings \
@@ -34,6 +36,9 @@ curl -sk -X PUT https://localhost:3000/api/settings \
## Adding an endpoint
Via App Settings → Models → Custom model endpoints → **+ Add endpoint**, or
directly:
```bash
curl -sk -X POST https://localhost:3000/api/model-endpoints \
-H 'Content-Type: application/json' \
@@ -62,7 +67,173 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
Endpoint management is admin-only in multi-user mode, same as remote/docker
hosts — these are machine-level infra, not per-user settings.
## Applying a model to a session
**Context length is discovered too, opportunistically and safely.** The plain
`GET /v1/models` response has no context-window field. Discovery only ever
looks for one for a model llama-swap's own response already reports
`status.value === "loaded"` for — never for an unloaded one, because
llama-swap treats `?model=` as a routing hint and asking about a model that
isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side
effect of what should be read-only discovery. A server with no `status` field
on any entry at all (not llama-swap) gets no context-length enrichment,
rather than guessing. A model's previously-learned context length survives a
later cycle where it wasn't the loaded one; it's dropped only once the model
disappears from the endpoint's list entirely. Stored per model in
`modelContextLengths` and applied automatically (see "Applying a model to a
session" below) so a CLI that would otherwise assume a large default context
window for an unrecognized model id stops silently overflowing a much
smaller real one.
**Where that number actually comes from matters, and got this wrong once
already.** The first cut read it from llama.cpp's own
`GET /props?model=<id>` (`n_ctx`) — plausible, and it worked in testing, but
confirmed live to be actively WRONG for a `--fit-ctx`-launched llama-swap
backend: `/props` reported `n_ctx: 154112` for a model llama-swap itself had
launched with `--fit-ctx 16384`, and the real server then refused a request
right at that real 16384-token limit — `/props`'s `n_ctx` appears to report
the model's theoretical/trained maximum there, not the runtime-configured
one. Discovery now parses the REAL configured size straight out of
llama-swap's own launch command instead (`GET /running`'s `cmd` field —
`--fit-ctx <N>` first, then the plain llama.cpp `-c`/`--ctx-size` a
hand-written command might use), and only falls back to the `/props` probe
when `cmd` states no recognizable flag at all.
**File size is discovered too, when the server states one.** llama-swap
writes a GB figure into an auto-discovered model's own `description`
(`"Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"`), parsed
into `modelSizesGB` — unlike context length, this needs no `/props` probe
(the figure is right there in the `/v1/models` response) and so is populated
for every model regardless of loaded state. A hand-configured profile's own
description has no such figure and correctly gets no entry, never a guess.
Used only to label the Run-menu picker's "loading model" banner (e.g.
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB) on llama-swap..."); never anything
a server-side check relies on.
**The loading banner is unbounded by design, and says so — no countdown, no
automatic give-up.** An earlier version scaled an expected-time estimate and
a timeout off the model's file size and auto-closed the session once that
elapsed, but a real load's actual duration depends on hardware this feature
has no way to know (VRAM, storage speed, whatever else is contending for the
GPU) — any fixed number was a guess dressed up as a fact, and a model that
genuinely takes 10+ minutes on slower hardware would just get killed
mid-load by its own display. The banner now says outright that it can take a
while depending on hardware and model size, polls
`GET /api/model-endpoints/:id/running-status` every second for as long as it
takes, and carries a **Cancel** button (rendered on the banner itself) that
ends the wait and closes the session the load was for — the user's own call
on when it's taking too long, not a fixed number baked into the client.
**The banner's second line is the real backend log line, not a guess.**
llama-swap's `GET /api/events` SSE stream carries the actual `llama-server`
process's own stdout — `load_model: loading model '<path>'`,
`llama_server: model loaded`, tokenizer warnings, all of it — tagged
`source: "upstream"`, distinct from llama-swap's own `source: "proxy"`
request-access lines. `running-status`'s response now includes `logLine`
(via `getLatestLlamaSwapLogLine`), and the banner shows it on its own line
under the disclaimer, e.g. "llama.cpp: load_model: loading model '...'" —
confirmed live end-to-end through a real forced swap, sequentially showing
the model path, a tokenizer warning, then staying on whatever llama.cpp last
printed once the load goes quiet (never cleared back to blank). ⚠️
**`GET /logs` — the endpoint this feature's own first cut was built
against — turns out to carry ONLY llama-swap's own proxy request-access
log.** Confirmed live it never showed a single backend line, even seconds
after a real, verified model swap; `/api/events`'s `logData` frames are the
only source that actually has it, and its own `source` field (`upstream` vs
`proxy`) is what `getLatestLlamaSwapLogLine` filters on. One `/api/events`
connection is held open per endpoint and reused across every session
watching a load on it (confirmed live to stay open indefinitely, unlike
`/logs`, which closes after a fixed ~100KB), idle-closed after 30s of nobody
polling it (`pruneIdleLlamaSwapLogTails`, same 20s sweep as the
swap-displacement check below).
`defaultModelId` names which discovered model the picker pre-marks for that
endpoint — the settings panel's Edit form exposes it as a select populated
from the endpoint's own discovered `models`, and the route refuses a value
that isn't one of them. It is applied automatically only when the endpoint
has exactly one discovered model (nothing to choose); with two or more it
is a pre-selection in the model-picker dialog below, never a silent default.
Re-discovering drops a default that no longer appears in the fresh list
rather than carrying an invalid one forward.
**Model lists refresh themselves.** A background sweep (`server.ts`,
`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every
saved endpoint the same way the manual `POST .../discover-models` route
does, best-effort per endpoint — one being unreachable on a given cycle
never blocks the others. Off under `npm test`, same reasoning as the Codex
plan-usage poll it sits beside: no real network to hit, no server instance
to keep the timer alive for.
## The Run-menu picker
With the setting on and at least one endpoint carrying a discovered model,
the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry
per (harness that can redirect to a custom endpoint, saved endpoint) pair,
e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI
registry's own `capabilities.customModelInjection` at page render
(`window.__codemanCustomModelClis`, `server.ts`) — never a hardcoded id list
in the frontend — so a CLI whose injection recipe lands later shows up with
no frontend change, and Antigravity (`unsupported`) never does.
Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`,
`session-ui.js`) rather than trusting anything cached from the dropdown's
own render — the model list can have changed via the 5-minute sweep above
or a settings-panel edit since the menu opened. With exactly one discovered
model it runs straight away; with two or more, a small modal
(`#customModelPickModal`) lists them and asks which one to use for this
launch, with the endpoint's `defaultModelId` marked but not auto-chosen —
the point of asking is letting one launch deliberately differ from the
saved default, not just confirming it.
**How the launch itself applies the endpoint depends on the harness.** For
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP (`runCustomModelEntry` →
`_runCustomModelEntryOneShot`), the endpoint/model is folded into the SAME
`POST /api/quick-start` call that creates the session (`customModel` field),
so the session launches directly on the endpoint — no restart, no visible
relaunch. Claude (`_runCustomModelEntryViaRestart`) still uses the original
two-step design: the launch runs a single native session exactly the way its
own Run-menu entry would, then **waits for the new session to go idle**
(`GET .../wait?until=idle`, bounded at 20s — a normal 200 either way, never
an error, per the wait endpoint's own contract) before applying the endpoint
via the restart route below. That wait exists because a freshly launched CLI
reports itself as `busy` for its own startup (a boot spinner, a
workspace-trust check) well before the apply call would otherwise reach it,
and the apply route correctly refuses to restart a session mid-turn — a
fresh boot looks exactly like one from the outside. A session still busy
after the wait reaches the apply call anyway and gets that route's own
honest `SESSION_BUSY` error, now visible as a sticky toast with a close
button rather than a generic message that vanished in three seconds. Claude
stays on this path because its own restart (`--resume`-based, keeping the
conversation) is far less jarring than the other seven's, and `runClaude()`'s
multi-tab launch and docker-config-drift confirm/retry loop make folding it
into the one-shot path separate work. It is a
one-off "try this endpoint" action, not a sticky mode: the plain Run button
still means "this harness, native cloud" afterward. Entries are hidden
entirely for a remote or Docker active case, since the apply route refuses
both (see the next section).
## Launching directly on an endpoint (no restart)
```bash
curl -sk -X POST https://localhost:3000/api/quick-start \
-H 'Content-Type: application/json' \
-d '{"caseName": "myapp", "mode": "codex", "customModel": {"endpointId": "llama-box", "modelId": "qwen3"}}'
```
`POST /api/quick-start`'s `customModel` field (`{endpointId, modelId,
confirmed?}`) computes the same injection the restart route below does, but
BEFORE the session exists — the session is minted its own id up front
(`crypto.randomUUID()`), the injection (env vars, and for a `configDir`-kind
CLI, the written config file) targets that real id, and the session launches
already pointed at the endpoint. No restart, because there was never a
native-backend launch to restart away from. Runs the same llama-swap
conflict check as the restart route (below) — a `409`-shaped
`{requiresConfirmation, currentlyLoadedModel, affectedSessions}` response
with no session created, resolved by retrying with `confirmedSwap: true` — and
is refused the same way for a remote or Docker case. This is what the
Run-menu picker uses for opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP;
Claude still uses the restart route below (see "The Run-menu picker" above
for why).
## Applying a model to an ALREADY-RUNNING session
```bash
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
@@ -84,6 +255,179 @@ since for those three the config file alone does not switch the model.
reattaches the durable remote/in-container tmux rather than relaunching the
agent, so the selection would report success and change nothing.
**Claude gets two more env vars when known/applicable, both declared on its
registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is set to `modelId`'s discovered context
length (see the discovery section above) whenever one is known. Without
it, Claude Code assumes a large (200k) window for any unrecognized custom
model id and never compacts, which reliably overflows a much smaller real
local context — confirmed live: a stock ~33.7K-token system prompt against
a 16384-token llama-swap model failed with `exceeds the available context
size`. No entry for the model in `modelContextLengths` means the var is
simply omitted, never a guess. ⚠️ **This var only affects when Claude
Code compacts conversation _history_ — it cannot fix a model whose real
context is smaller than Claude Code's own fixed per-turn overhead**
(system prompt + tool schemas, empirically ~36.4K tokens, confirmed live
via an `in:0 out:0` failure on the very first message, before any
history exists to compact). No context-length declaration changes that
fixed overhead, so a model below the safe floor fails outright on
message one regardless of what this var says. See "Context-window floor
warning" below for how Codeman catches this case before launching
instead of after.
- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory
the `configDir`-kind CLIs use (empty, no files written into it), so the
injected `ANTHROPIC_API_KEY` never shares a directory with a stored
claude.ai OAuth login. Claude Code still prints "Both claude.ai and
ANTHROPIC_API_KEY set" when the two coexist in the same config directory —
cosmetic (confirmed live: the API key wins for actual requests either way,
visible in the terminal's own `API Usage Billing` line) but worth
eliminating rather than living with. The directory's `projects`
subdirectory is symlinked (a junction on Windows) back to the real
`~/.claude/projects` so the response viewer, subagent windows and Read My
Mind keep working for that session — the same trade-off and fix documented
for a manually-set `CLAUDE_CONFIG_DIR` in
[`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied
automatically here. Best-effort: a platform that refuses the symlink keeps
the pre-existing blind-response-viewer side effect rather than failing the
whole custom-model apply over it. ⚠️ **This relocates the whole `.claude`
tree, not just transcripts**: a custom-model Claude session also loses the
user's global `settings.json`, user-level skills (the codeman agent skill
included), user-level agents and commands, and the MCP servers configured
in `~/.claude.json` — none of those are symlinked back, only `projects` is.
A fine trade for "point this session at my local llama.cpp," but worth
knowing before it surprises you mid-session.
**That isolated directory needed one more fix to actually be usable
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
real profile's prior "Detected a custom API key — use it?" approvals, so
without more, Claude Code stops and asks that on _every single launch_ —
confirmed live, and with nobody at a TTY to answer, its own default answer
("No") silently refuses the very key this feature just injected, which
looks like the endpoint being ignored entirely. `customModelInjection`'s
`apiKeyTrustFile` (`{ relPath: '.claude.json', shape:
'claude-api-key-responses' }` on claude's entry) pre-seeds that exact
approval: the apply step merges `customApiKeyResponses.approved: [apiKey]`
into `<configDir>/.claude.json`, the same field a real answered prompt
itself writes to (confirmed against a real file after answering by hand
once) — this answers the prompt in advance rather than bypassing it. The
merge preserves whatever else the CLI already wrote into that file on an
earlier launch in the same isolated directory (`userID`, `numStartups`,
earlier approved keys), and a missing or corrupt file is treated as empty
rather than failing the apply.
**A fresh `CLAUDE_CONFIG_DIR` isn't just missing that one approval — Claude
Code treats it as a brand-new profile and replays its ENTIRE first-run
sequence on every launch: the theme picker, the security-notes screen, the
per-project "trust this folder?" dialog, and (running with
`--dangerously-skip-permissions`) a one-time warning about bypassing
permissions.** Confirmed live: none of these show up again for a real,
already-onboarded profile, but every custom-model session gets a fresh,
otherwise-empty isolated directory, so it saw all four every single time.
`customModelInjection`'s `skipFirstRunPrompts` (`true` on claude's entry,
requires `apiKeyTrustFile` since it reuses the same file) pre-seeds the
state a real profile accumulates from answering all of that once:
`hasCompletedOnboarding: true` and the launching session's own
`projects[workingDir].hasTrustDialogAccepted: true` go into the same
`<configDir>/.claude.json` the API-key approval above already merges into
(other projects, and other fields on this session's own project entry, are
left untouched), and `skipDangerousModePermissionPrompt: true` goes into
`<configDir>/settings.json` — a different file, merged the same
corrupt-tolerant way. `workingDir` is used exactly as the session was
launched with as its cwd, never realpath'd or slash-normalized, since
that's the literal string Claude Code itself uses as the project key.
**llama-swap gets two more fixes on top of the context-length/config-dir
ones above, both from watching a real switch live.** llama.cpp only ever
runs one model at a time; llama-swap swaps the backing process on demand,
which can take anywhere from a few seconds to well over a minute:
- **The conflict check.** Both apply routes (the restart one here and the
one-shot `POST /api/quick-start` above) call llama-swap's own
`GET /running` first — feature-detected, so a plain llama.cpp/OpenAI-
compatible server (no such endpoint) is simply never checked. If a
_different_ model is currently loaded and ready, and another **live
session's own selection** is using it, the apply returns
`{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}`
instead of silently switching — nothing is applied or created yet.
Retrying with `confirmedSwap: true` skips the check (the legacy `confirmed: true`
still means both questions). Switching with nothing
else affected proceeds immediately; this is a warning about disrupting
another session, never a gate on the switch itself.
- **Actually starting the load.** llama-swap has no "switch model" admin
call — the only thing that starts a swap is a real inference request
naming the model, and confirmed live: applying a selection alone never
reached llama-swap at all (nothing in its own server logs), since nothing
had actually asked it to load anything yet. Both apply routes now also
send the smallest real request that will —
`POST <baseUrl>/v1/chat/completions` with `max_tokens: 1` and one
throwaway message — whenever the
target model isn't already the one loaded and ready, fire-and-forget (its
response is never read; `GET /api/model-endpoints/:id/running-status`,
polled client-side, is what actually confirms readiness). The response
also carries `modelSwapInProgress: true` in that case, which is what
drives the Run-menu picker's own "loading model" status banner.
## Catching a swap after the fact
The conflict check above only runs at the moment a session is created or a
model is applied — it has no way to catch a swap that happens **later**.
Confirmed live: a session created while nothing else conflicted at that
exact instant can still get silently displaced afterward, once a
_different_ session's own normal use (or its own create-time load trigger)
asks llama-swap to load something else. llama-swap has no push
notification of its own for this, so a background sweep
(`detectCustomModelSwapDisplacements`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS`
= 20s in `server.ts`) polls `GET /running` once per distinct endpoint that
has at least one live custom-model session, and compares each such
session's own `modelId` against what is actually loaded. A session whose
model is no longer in that list gets a `custom-model:swapped-out` SSE event
(`{sessionId, sessionName, endpointId, previousModel, currentlyLoadedModel}`),
shown as a global toast — global rather than tied to that session's tab,
since the whole point is telling the user before they type into it
expecting the model they picked. Notifies **once per displacement**: the
same de-dupe `Set` clears a session's flag once its own model is loaded and
ready again, so a later, genuinely new displacement notifies again rather
than the session staying silently un-notified forever after the first one.
## Context-window floor warning
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
empirically ~36.4K tokens) can exceed a small local model's _entire_ real
context on its own, before any conversation history exists to fill it —
confirmed live twice, both as an `in:0 out:0` failure on the very first
message sent. `CLAUDE_CODE_MAX_CONTEXT_TOKENS` (above) cannot fix this: it
only governs when Claude Code compacts conversation history, and there is
no history yet on message one. Applying such a model would look like the
endpoint being ignored, or the wrong model being used, when in fact the
endpoint applied correctly and the model is simply too small for this CLI.
Both apply routes (the restart route and the one-shot `POST
/api/quick-start`) now check for this **before** launching or restarting
anything, gated on the CLI's registry entry declaring a `contextLengthVar`
(currently only claude — the check is a no-op for every other CLI by
construction, never a hardcoded mode check). If the model's discovered
context (`modelContextLengths`, from discovery above) is below
`CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000, comfortably above the measured
~36.4K overhead), the response is `{requiresContextWarning: true, modelId,
contextLength, minSafeContextTokens}` instead of applying — nothing is
restarted or created yet. A context length that was never discovered at
all skips the check entirely (nothing to compare, so it fails open rather
than warning on every model an endpoint hasn't reported a size for).
Retrying with `confirmedContext: true` launches anyway (the legacy `confirmed: true` still means both questions).
The Run-menu picker shows this as an in-app modal
(`#customModelContextWarningModal`, matching the llama-swap conflict
modal's look) naming the model, its discovered context, and the safe
floor, and explaining the fix: reconfigure llama-swap to give that model
(or a smaller one) an explicit larger context instead of relying on
auto-fit (`--fit-ctx`), which optimizes for the biggest _model_ that fits
rather than the biggest _context_ — e.g. adding `-c 65536` (or as large a
`--ctx-size` as the hardware holds) to that model's llama-swap config
entry. A smaller model at a much larger explicit context often fits in
the same VRAM a bigger model's auto-fit context gets shrunk to make room
for.
Clear back to the harness's native cloud default with:
```bash
@@ -101,6 +445,14 @@ id, model and injected key NAMES are persisted, the values are re-derived
from the endpoint store on recovery, and the pane keeps running against the
endpoint in between because tmux retains its environment.
⚠️ Clearing removes injected keys **by name**, and `CLAUDE_CONFIG_DIR` is one
of the names claude's selection injects — so a session that ALSO had
`CLAUDE_CONFIG_DIR` set through the generic `envOverrides` field (the
per-client-account case) loses that override on clear too, and silently
falls back to the server's default Claude account. If you route a session
to a specific account this way, re-apply the override after clearing a
custom-model selection from it.
**New sessions always default back to the harness's native backend.** A
custom-endpoint selection is a per-session choice, never a sticky global
default — starting a fresh session doesn't inherit whatever the last one was
@@ -115,17 +467,53 @@ automatically). Results:
- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply
came back through the endpoint.
- **Codex** — the config is structurally correct, but Codex only speaks the
Responses API since Feb 2026, which llama.cpp/llama-swap don't implement.
This is a real protocol incompatibility, not a bug here; Codex support
needs a Responses-API-compatible endpoint.
- **Codex** — the config is structurally correct, and against a llama-swap
server that DOES answer `/v1/responses` (confirmed live: a plain,
no-tool-call chat turn returned a real reply), the picture is more
nuanced than a flat failure. A real tool-call attempt (`run the shell
command: echo hello`) came back as `agent_message` TEXT — literally the
tool-call JSON printed as the model's answer — instead of a
`function_call` item Codex would actually execute (confirmed via `codex
exec --json`'s raw event stream). So plain chat can work while the thing
that makes Codex a coding agent — actually running commands and editing
files — does not; treat Codex as still unreliable for real work against a
llama.cpp/llama-swap endpoint, tool-calling gap included, not just the
earlier-documented `wire_api` mismatch (which not every deployment hits
the same way — some legitimately have no `/v1/responses` route at all).
Separately, EVERY custom-endpoint Codex session prints `Model metadata
for '<id>' not found. Defaulting to fallback metadata...` on launch —
confirmed harmless (the reply above still came back correctly): Codex's
model metadata (reasoning-tier options, per-model system-prompt
templates, context-window figures) comes from `models_cache.json`, a
local cache of OpenAI's own hosted model catalog that a custom local
model can never appear in by construction, since it isn't one of
OpenAI's models. There's no config.toml override for a model's metadata,
and fabricating a fake catalog entry would mean copying the _shape_ of
OpenAI's own proprietary schema (their per-model system-prompt content
included) for a warning that doesn't otherwise affect behavior — not
something to build into discovery.
- **Gemini** — fails with `Invalid auth method selected`, traced to an
undocumented `GATEWAY` auth path gemini-cli selects once
`GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation
(several auth workarounds were tried and ruled out); do not rely on
Gemini support yet.
- **DeepSeek** — the request reaches the server (env vars are read) but
gets a consistent `HTTP_404`. Root cause not identified; best-effort only.
- **DeepSeek** — root cause of the `HTTP_404` found and fixed. DeepSeek
Harness's own bundled provider module (`@deepseek-ai/dsh-llm-deepseek`)
builds its request URL as `${DEEPSEEK_BASE_URL}/chat/completions` with no
`/v1` insertion of its own (its real public API, `https://api.deepseek.com`,
expects the caller's base URL to already carry any needed prefix) —
confirmed by reading its own source and, live, that
`POST <baseUrl>/chat/completions` 404s against llama-swap while
`POST <baseUrl>/v1/chat/completions` succeeds; the harness's own error
template (`DeepSeek API error (HTTP ${status})`) matches the originally
reported symptom exactly. `customModelInjection`'s new `appendV1Suffix`
(deepseek's entry only — claude/gemini must NOT get it, since claude was
already confirmed working against the raw `baseUrl`) fixes it by writing
`DEEPSEEK_BASE_URL` with `/v1` appended. Not yet re-run end-to-end with a
real `dsh` binary (no install available in this environment) — the fix
is source-confirmed and live-verified at the HTTP level, but a real
"hello world" reply through `dsh` itself is still outstanding before
calling this fully verified like the harnesses above.
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind
+175
View File
@@ -352,6 +352,177 @@ path but the SESSION (`session.remote`): a remote session never falls back to lo
`fs`, and a local session never opens an ssh connection — including for attachment
records, which are keyed to the session that registered them.
## Wake-on-LAN from user input
A durable remote session survives an SSH drop (COD-104/108), but nothing brought the
HOST back. When the remote machine suspended, the local pane's `ssh` child **stalled**
rather than exited: `tmux send-keys` SUCCEEDS against a stalled pane, so typed input
vanished with no error anywhere, and without a keepalive the pane could look alive for
the OS TCP timeout. The only recovery was waiting for the reconnect watcher, which
gave up after ~13 minutes and, once exhausted, never retried.
An **optional** `wakeMac` (one or more MAC addresses, comma-separated) or `wakeCommand` on a
remote host closes that: on user input, `POST /api/sessions/:id/input` probes the host, and if
it is unreachable it wakes it, polls until the host answers, reattaches the pane
(`Session.reattachRemote()`, which idempotently attaches the still-running remote tmux — the
agent conversation is not restarted), and flushes the input that arrived meanwhile.
Implementation: `src/remote-wake.ts`.
The same wake path also serves **opening** a session, which is where a sleeping host used to
be a dead end: pressing Run on a remote case (`POST /api/quick-start`) or Attach on a
discovered remote tmux session (`POST /api/sessions` + `attachRemoteSession`) probes the host
first, and on a sleeping one wakes it, waits for SSH and only then runs the tmux prereq probe.
Without that the run failed with `could not verify tmux on remote host …` — an ssh error that
blames tmux for a machine that is merely suspended. The wait is **blocking** (the caller gets
the session or the error) but bounded by `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS` (40 s) rather
than the 90 s session default, because the dashboard sits behind a reverse proxy whose default
`proxy_read_timeout` is 60 s: a longer wait would be cut off at the proxy while the session was
still being created. The budget covers the whole request, not just the wait (40 s wake + 1.5 s
probe + the tmux prereq probe's own 15 s timeout = 56.5 s worst case). A host with no wake target is not even probed on this path, so nothing
changes for it, and `remote:hostWaking` is broadcast without a `sessionId` (the toast then reads
"the session starts when it is back" — there is no session yet, and no input queued behind it).
Two wake paths, `wakeCommand` first because it is the explicit override:
- **`wakeMac`** — Codeman builds the magic packet itself (`buildMagicPacket`, six `0xFF`
bytes then the MAC repeated 16×; the shape is asserted byte-for-byte) and broadcasts it
over UDP port 9 (`sendWakePackets`). This is the normal case: no external script, and one
MAC list per host instead of one per consumer.
- **`wakeCommand`** — a single executable path, run WITHOUT a shell. For hosts that need a
router/another machine to send the packet.
**UI**: a banner (`#hostWakeBanner`, `host-wake-ui.js`) appears while the ACTIVE remote
session's host is unreachable — amber, since the Codeman session is healthy and only the
machine is asleep. With a wake target the action is **Wake** (`POST /api/sessions/:id/wake`);
with none it is **Configure WoL** and opens `#wakeConfigModal`, a small form for that host's
`wakeMac`/`wakeCommand` that saves with `PUT /api/remote-hosts/:id` (in multi-user mode that
GET is admin-only, so a non-admin is told the setting is admin-only instead of "host not
found"). Reachability for the banner comes from `GET /api/sessions/:id/reachability`: once
when the remote tab is activated (a user action), and every 30 s while the tab is visible
**only for a host with a wake target** — each poll is a TCP connect to the host, and a timer
that connects to a host Codeman could not wake anyway is exactly the timer-driven traffic
the keepalive rule below rejects (it cannot wake a host, but it can keep an activity-based
suspend timer from firing). A host the probe cannot reach (see the next section) is never
polled. ⚠️ The button is pressed from the SAME
dashboard as Run/Attach, so it holds its request open under the same proxy and uses the same
40 s budget — and it **queues nothing**: browser keystrokes travel over the WebSocket, which
deliberately does not pass through the registry (that is the hot path this feature keeps its
hands off), so the banner says "waiting for the host to come back" for the button and only
claims "input is queued" when the HTTP input path actually buffered bytes
(`queuedInput` on the two SSE events).
**Hosts behind a jump host or SOCKS proxy are reachability-UNKNOWN.** The probe is a bare
TCP connect to `host:port`, and a host reached through `jumpHost`, `socksProxy` or a
`ProxyCommand`/`ProxyJump` in `extraSshOptions` does not answer that even while ssh works —
the direct address may not route at all (the cloudflared case). Acting on the resulting
"unreachable" verdict was wrong three times over: a permanent banner over a healthy session,
a create-path error that replaced a genuine "needs tmux" with "not reachable", and — with a
wake target configured — every HTTP input buffered for the life of the session, because the
readiness poll could never succeed. `isProbeable()` (`remote-wake.ts`) decides from the
proxy fields, which travel on `WakeableRemote`; for such a host the registry delivers input
unchanged, `GET …/reachability` answers `reachable: null, probeable: false` (unknown is not
`false`, and only a proven `false` raises the banner), the create/attach path is not gated
(`ensureHostAwake` → `'unprobeable'`, handled like `'no-target'`), and the quick-start
"not reachable" message is reserved for a **proven** unreachable host (`=== false`). A wake
target can still be fired for it through `POST /api/sessions/:id/wake`, blind: the packet or
command goes out and the response says only whether it did — no readiness poll, no reattach
(the COD-108 watcher owns the pane once ssh works again), no "waking" toast.
The invariants worth keeping:
- **Authorization comes before the wake.** In multi-user mode the attach path
(`POST /api/sessions` + `attachRemoteSession`) answers `403` to a non-admin BEFORE the
host is looked up or probed: remote hosts are admin-only infrastructure everywhere else
(the list is `[]` for a non-admin, write and discovery routes are `adminOnly`), and the
wake spawns the host's `wakeCommand` or broadcasts a packet — a gate that came after the
wake handed an unprivileged account a way to run that executable for any configured
`hostId`, hold the request for the wake budget, and only then be refused for the
workingDir. The quick-start path resolves its remote case through `canAccessOwned`
first. Pinned in `test/routes/session-remote-wake.test.ts` (wake spy stays empty).
- **The caller is told what happened to its bytes.** The non-wait input route answers
`{buffered:true}` when the registry took the chunk and `{buffered:true, dropped:true}`
when it was over the cap and is gone; the send-and-wait route answers `OPERATION_FAILED`
when the host never comes back, like the create and attach paths, instead of writing
into the stalled pane and reporting `delivered:true` plus a timeout. Flushed chunks are
written with `fromUser`, so a first prompt that was buffered through a wake can still
name the tab.
- **Only an EXPLICIT request may wake a host:** user input on an established session, the wake
button, or the user's own session create/attach request (`ensureHostAwake`). Everything that
runs on a TIMER must never wake one — the COD-108 watcher, the server's dropped-session
handler, boot recovery and session discovery have no access to the wake registry, and neither
has the shared session service, because `cron-service.ts` builds sessions there with nobody
waiting on the answer; a wake on such a path would re-wake the host seconds after every
suspend, so it could never stay asleep (the same failure `hufflepuff-mcp-lazy` exists to
prevent for MCP keepalives). A reachability check, a discovery listing and the tmux prereq
probe never wake: they are questions, not actions. All of it is enforced by tests in
`test/remote-wake.test.ts` (two wiring guards: one pins the importers — the route module and
`server.ts`, which holds the registry for its LIFETIME only, `drop()` on session cleanup and
`stop()` on shutdown — and one asserts `server.ts` calls nothing but those two, while
`ensureHostAwake` has exactly one caller file) and `test/routes/session-remote-wake.test.ts`,
not by comments.
- **Detection is a bare TCP connect** to the SSH port (then the configured `port`, else 22),
throttled per session, and only for wake-enabled hosts. No `ServerAliveInterval` is added to
the launch command: keepalives push bytes into an otherwise idle connection every interval,
which is exactly what a byte-threshold idle detector must not count as activity. A probe is
~200 bytes per 30 s, orders of magnitude below any such threshold, and the SYN alone cannot
wake a host.
- **Input is buffered while a wake is in flight** (`REMOTE_WAKE_PENDING_MAX_BYTES`,
oldest whole chunks dropped, bounded so user input cannot grow memory) and flushed in
order after the reattach, with a settle delay so bytes cannot land in a still-connecting
pane. ⚠️ A chunk LARGER than the cap (one big paste is one `input` value) is dropped
**outright**, never trimmed: it was never typed character by character, so its tail is not
"what the user just typed" but a fragment of a command they never sent — the drop is logged
instead. ⚠️ Only the HTTP input route reaches the registry; the **WebSocket keystroke path
is deliberately NOT wake-aware**, so typing into a sleeping host sends nothing and queues
nothing (the banner's Wake button is the recovery for that case, which is why it must not
promise queued input). The **send-and-wait** path blocks on the wake instead — its response
is open anyway, and buffering would break the wait contract. ⚠️ A flush write that FAILS
drops the whole remaining buffer (logged) rather than retaining it: the wake still resolves
and marks the host reachable, so the next input takes the deliver path while a retained
chunk would wait for the NEXT wake — replayed hours later, after everything typed since,
possibly ending in a carriage return. Same policy as the oversized paste.
- **The command runs without a shell** (`spawn(path, [], { stdio: 'ignore' })` — `shell`
defaults to `false`), the schema
requires a single executable path (no arguments, no `$`/backtick), and `wakeMac` is a
structural hex-pair allowlist. A broken or missing wake target fails the wake, never the
input route.
- **`wakeMac`/`wakeCommand` are host-level config, refreshed on recovery AND live**
(`rehydrateRemoteHostFields` in `src/remote-hosts.ts` plus `RemoteWakeDeps.resolveRemote`).
A session's `remote` block is persisted at launch time, so a field added to
`remote-hosts.json` later would otherwise never reach an already-running session — not even
across a Codeman restart, and certainly not right after saving the banner's config dialog.
Recovery rehydration covers restarts, the (throttled, cache-backed) resolver covers the live
session; the host config is authoritative for both (removing the field disables the feature
again). Other host-level fields deliberately stay as persisted, so neither path can
silently re-point an existing pane's SSH options.
- **UI/SSE**: `remote:hostWaking` and `remote:hostWakeFailed` (plus the reused
`remote:sessionReconnected`) drive the banner and toasts, all from `host-wake-ui.js` —
its handlers are the ONLY definitions, since a second one in another mixin would be
silently shadowed by script order. Both carry `queuedInput`, which is true only when the
server actually holds bytes for that session — the wording keys off that, not off "a wake
is running", so the button path never claims input is queued. In multi-user mode the
whole `remote:` family is **session-scoped** (`deriveSseHint`, `server.ts`): an event with
a `sessionId` reaches that session's owner, and the create/attach wake — which has no
session yet — carries the requesting `username` instead (`ensureHostAwake({ requestedBy })`),
since its payload names a `hostId`/`label` that `GET /api/remote-hosts` withholds from
non-admins. With neither, it reaches admins only.
- **No real IO under vitest.** `probeRemoteHostReachable`, `runRemoteWakeCommand` and the
default UDP socket of `sendWakePackets` throw under `VITEST` (as `remote-files.ts` does),
so a test that reaches the defaults fails loudly instead of connecting, spawning or
broadcasting from CI. Every consumer injects its IO (`RemoteWakeDeps`, the socket
factory); `createDefaultRemoteWakeDeps({ probe })` also polls readiness with THAT probe,
which is the leak the guard found.
Tests: `test/remote-wake.test.ts` (decision/throttle table, single-flight registry,
buffering + flush order, MAC parsing/magic packet, live host-config resolution, the proxied
host, SSE payload routing, the vitest IO guard, and the wiring guard),
`test/routes/session-remote-wake.test.ts` (the input route buffers instead of writing into a
sleeping host — and writes straight into a proxied one —, the reachability route never wakes
and reports a proxied host as unknown, and the wake route reports the no-target case the UI
turns into "configure WoL"), `test/sse-routing-remote.test.ts` (multi-user routing of the
`remote:` family) and `test/host-wake-banner.test.ts` (banner visibility and when the poller
may connect).
## API
Routes are registered in `src/web/routes/case-routes.ts`:
@@ -365,6 +536,10 @@ Routes are registered in `src/web/routes/case-routes.ts`:
| `GET` | `/api/remote-hosts/:hostId/sessions` | Discover `codeman-*` sessions on the host (COD-105; `listRemoteCodemanSessions`, never errors) |
| `POST` | `/api/cases/remote-link` | Link a case to a remote host (creates the `RemoteCase`) |
`RemoteHost` accepts the optional `wakeMac` (magic packet, sent by Codeman) and `wakeCommand`
(single executable path, run without a shell, takes precedence) — see **Wake-on-LAN from user
input** above.
Attaching to a discovered session is a **session-create** path, not a host route:
`POST /api/sessions` accepts `attachRemoteSession: { hostId, remoteSessionName }`
(schema in `schemas.ts`; `remoteSessionName` must match `^codeman-[a-zA-Z0-9._-]+$`),
+7
View File
@@ -273,9 +273,16 @@ into the case's `.claude/settings.local.json` so that `/model` keeps working.
- **Shell** for the times you want a terminal on your phone with no agent at all. It is a
genuinely useful mode, not a fallback.
## Pointing one at your own server
Most of these harnesses can also run against a custom OpenAI-compatible endpoint instead of
their native cloud backend, for one session at a time, an opt-in feature covered in full on
[Custom Model Endpoints](Custom-Model-Endpoints).
## Read next
- [Core Concepts](Core-Concepts) - run modes versus location overlays.
- [Custom Model Endpoints](Custom-Model-Endpoints) - run a harness against your own server.
- [Settings Reference](Settings-Reference) - model, effort, and permission-mode settings.
- [Keeping Agents Running](Keeping-Agents-Running) - what idle detection does per mode.
- [Security](Security) - what skipping permission prompts actually means.
+183
View File
@@ -0,0 +1,183 @@
# Custom Model Endpoints
Point a harness at your own OpenAI-compatible server instead of its native cloud backend, for
one session at a time. "Custom endpoint" covers **local** hardware (llama.cpp, Ollama, vLLM,
a home GPU rig, DGX Spark, Strix Halo) and **cloud** services (Azure AI Foundry's
OpenAI-compatible endpoint, OpenRouter, a company gateway) alike, anything answering
`GET /v1/models` and `POST /v1/chat/completions` in the standard shape.
**Off by default.** Turn it on in App Settings → Models → **Custom model endpoints**.
## Adding an endpoint
Still in App Settings → Models → Custom model endpoints:
1. **+ Add endpoint** — give it an id, a label, and the base URL (`http://192.168.1.50:8080`,
say). An API key is optional; most local servers don't check one.
2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it.
3. Pick a **default model** from what was discovered. This is the model the Run-menu entry
applies directly when only one model is discovered; with two or more, it's just the one
pre-marked in the picker dialog described below, not a silent default.
Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker
hosts — these are machine-level infra, not a per-user setting.
**Model lists refresh themselves.** Every saved endpoint is re-discovered automatically every
5 minutes in the background, so a model the server starts serving later — or stops serving —
shows up without another manual click of **Discover**. One endpoint being unreachable on a
given cycle (powered off, wrong network) never blocks the others from refreshing.
**Context length is picked up automatically where it can be, safely.** Against a
llama.cpp/llama-swap server, discovery also learns each _currently loaded_ model's real
context window and applies it to the launched session (Claude Code today — see below), so
the harness stops assuming a large default window for a model name it doesn't recognise and
overflowing a much smaller real one. It's deliberately never probed for a model that isn't
already loaded, since asking a llama-swap server about an unloaded model can trigger an
actual, slow model swap as a side effect — a model just not currently loaded keeps whatever
context length an earlier cycle already learned for it instead.
## Running a session against one
With the setting on and at least one endpoint carrying a discovered model, the **Run**
dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a
custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a
session on that harness exactly the way its own entry would. It is a one-off "try this
endpoint" action, not a sticky mode — the plain **Run** button still means "this harness,
native cloud" afterward, and a fresh session never inherits whatever the last one was
pointed at.
**Which model it uses depends on how many the endpoint has discovered.** With exactly one,
the session launches straight away on that model — nothing to choose. With two or more, a
small dialog asks which one to use for this launch before starting the session; the
endpoint's default model, if set, is marked but not auto-picked, so a launch can deliberately
use a different one without changing the saved default.
**For opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP, picking an entry launches
straight onto the endpoint** — no restart, because the endpoint is applied before the
session's process ever starts. **Claude still restarts the harness's process in place** —
same tab, same conversation (`--resume`) — after a normal native launch, since that restart
is far less jarring for Claude than for the other seven, whose own TUI can fully
reinitialize on a restart. Either way, every supported harness reads its endpoint config at
process start, never per turn, so there is no live hot-swap while a turn is running.
Picking an entry that launches a **brand-new** Claude session waits (up to 20 seconds) for it to
finish its own startup before applying — a freshly started CLI reports itself as busy for its
boot sequence, and applying to a genuinely busy session is refused so a real, in-progress
turn is never interrupted out from under you. A session that is still busy after that wait
(a very slow-starting CLI, or one you started typing into right away) surfaces that refusal
as an ordinary error, which now stays on screen with a close button instead of vanishing
after a few seconds — read it, it names the actual reason rather than a generic failure.
Entries are hidden entirely for a session in a **remote (SSH) or Docker case** — support for
redirecting those hasn't landed yet, see below. The picker also only appears in the desktop
**Run** dropdown; the phone home screen builds its own run picker separately and does not
currently offer these entries.
**Against llama-swap, applying a selection also starts the actual model load, rather than
waiting on your first prompt to do it.** llama-swap has no "switch model" button of its own
— the only thing that starts a swap is a real request naming the model, and confirmed live:
just applying a selection never reached llama-swap's own logs at all until something asked
it to load. Picking an entry now also sends the smallest real request that will trigger
that load, in the background, the moment the target model isn't already loaded and ready.
**The centred loading banner has no countdown and no automatic timeout — it waits as long as
it takes, and tells you so.** When it knows the model's discovered file size (its GB figure,
when llama-swap states one) it's shown too, e.g. "Loading qwen3.8-27b (16.4 GB) on
llama-swap — this can take a while depending on your hardware and the model size." An
earlier version tried to estimate and enforce a time limit, but real load time depends on
hardware this feature has no way to know, so a fixed number was always a guess — worse, one
that could kill a genuinely slow load partway through. If it really is taking too long, a
**Cancel** button right on the banner ends the wait and **closes the session that load was
for**, on your own call rather than a guessed deadline.
**The banner also shows a real, live second line of what llama.cpp itself is doing** — not
a made-up progress phase, the actual next line the `llama-server` process printed, e.g.
"llama.cpp: load_model: loading model '/models/.../Qwen3.8-27B.gguf'" then later
"llama.cpp: llama_server: model loaded". It comes straight from llama-swap's own event
feed, filtered down to just the backend process's own output (not llama-swap's own request
logging), and stays on whatever it last said once the load goes quiet, rather than
clearing back to nothing.
**You'll also be told if a session's model gets swapped out from under it later, not just
at launch.** The conflict warning above only fires at the moment you launch or apply a
model — llama.cpp only runs one model at a time, so if a DIFFERENT session using the same
endpoint later triggers its own load, whatever was loaded before (including a session you
already had running) gets silently evicted, with no warning at that instant since nothing
conflicted when it was first set up. A background check (every 20 seconds) catches this
after the fact and shows a toast naming which session lost its model and what's loaded now
— so you know before typing into that session that it's about to reload (and, in turn,
evict whatever displaced it).
**Claude Code specifically gets three extra fixes applied automatically:**
- Its discovered context length (see above) is passed through as
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`, so it doesn't send a full-size prompt against a much
smaller real local context and overflow it.
- Its session runs with an isolated `CLAUDE_CONFIG_DIR`, so the injected API key never sits
in the same directory as a stored claude.ai login — that combination is harmless for actual
requests (the API key wins) but the CLI still prints a "both claude.ai and
ANTHROPIC_API_KEY set" warning about it, which this avoids entirely. The isolated directory
keeps a link back to your real session history so the response viewer and similar features
still work for that session. That isolated directory starts with no prior approvals of its
own, so Codeman also pre-approves the injected key the same way answering Claude Code's own
"Detected a custom API key" prompt once would — without it, that prompt would otherwise
reappear on every single launch with nobody there to answer it.
- **That same fresh isolated directory also looks like a brand-new Claude Code profile**, so
without this fix it replayed the WHOLE first-run sequence every single launch: the theme
picker, the security-notes screen, the "trust this folder?" dialog, and a one-time warning
about running with permissions bypassed — none of which a real, already-used profile shows
again. Codeman now pre-seeds that same "already been through this once" state (onboarding
completed, this session's own project marked trusted, the bypass-permissions warning
acknowledged) so a custom-model launch reaches the actual conversation exactly as fast as a
native cloud one does, instead of stopping at a wizard with nobody there to click through it.
**If a model's real context is too small for Claude Code to even get started, you get a
warning instead of a confusing failure.** Claude Code's own system prompt and tools take up
roughly 40K tokens on their own, before you've typed anything — a small local model with a
smaller real context than that fails outright on the very first message, no matter what
context size Codeman tells it to expect (raising the declared context only changes when
Claude Code trims _conversation history_, and there is none yet on message one). Picking
such a model now shows an in-app dialog naming the model, its discovered context and what's
needed, before anything launches or restarts, with the fix spelled out: reconfigure
llama-swap to give that model (or a smaller one) an explicit larger context instead of
relying on auto-fit (`--fit-ctx`), which sizes the context around fitting the biggest model
rather than the biggest context — for example adding `-c 65536` to that model's llama-swap
entry. "Launch anyway" is still there if you want to try regardless.
## Which harnesses actually work
| Harness | Status |
| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
| **Codex** | Config is correct, and plain chat can work against a server that speaks the Responses API — but a real tool-call attempt comes back as inert text instead of running, so it's still not usable for real coding work. |
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
| **DeepSeek** | The original 404 is root-caused and fixed (DeepSeek Harness's own code was missing a `/v1` most local servers require) — not yet re-run against a real `dsh` install to confirm end-to-end. |
| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. |
Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry,
not a fixed list here, so this table can go stale before this page does — a greyed-out or
missing entry is the more current answer.
## What it does not do
- **No remote or Docker sessions yet.** Both restart their agent differently under the hood
(reattaching a durable tmux session rather than relaunching the process), so redirecting
them needs its own plumbing that hasn't been built.
- **No live hot-swap mid-conversation.** Applying a selection always restarts the process.
- **No button to un-point a session from the UI yet.** Clearing back to native cloud is an
HTTP call (`POST .../custom-model {"clear": true}`) or deleting the session; the settings
panel manages saved endpoints, not what a running session is currently pointed at.
- **Nothing is shared with your real cloud credentials.** The endpoint's own key, if any,
never touches your Anthropic/OpenAI/Google login — a custom endpoint is a separate,
explicit choice per session.
## Security
An endpoint's base URL can't point at a link-local or cloud-metadata address (both at save
time and against the address it actually resolves to), the same guard Web Tabs uses for
saved dashboards. Endpoint records and any per-session config files a harness needs are
written with owner-only permissions. See
[custom-model-endpoints-plan.md](https://github.com/Ark0N/Codeman/blob/master/docs/custom-model-endpoints-plan.md)
in the repository for the full design reasoning, including why this feature closed a
pre-existing gap in how session environment overrides were guarded rather than opening a new
one.
+2
View File
@@ -34,6 +34,8 @@ Press `Ctrl+?` in the app for the same list in a floating overlay.
| Right-click | Copy the selection. With nothing selected the native menu is left alone. |
| `Ctrl+Z` | Swallowed in agent sessions so a running CLI cannot be suspended. Normal job control in a shell. |
Anything you copy is cleaned on the way to the clipboard: each line loses the padding spaces a full-screen program paints across the rest of the row. Leading indentation is left exactly as it is, so indented code, a `git log` message body and `git diff` context lines paste back the way they looked on screen. An `Alt+drag` rectangular selection is copied exactly as it looks, so its columns stay lined up.
## Everything else
| Shortcut | Action |
+4
View File
@@ -92,6 +92,10 @@ Model and effort are both **soft defaults**: the model is written into the case'
`.claude/settings.local.json` and effort is passed at start, so `/model` and `/effort`
inside a session override them at any time.
**Custom model endpoints** (off by default) adds a saved-endpoint list plus a matching
section to the Run dropdown, for pointing a harness at your own OpenAI-compatible server
instead of its native cloud backend. See [Custom Model Endpoints](Custom-Model-Endpoints).
### Agents & CLIs
| Setting | Notes |
+1
View File
@@ -12,6 +12,7 @@
- [The Dashboard](The-Dashboard)
- [Agent CLIs](Agent-CLIs)
- [Custom Model Endpoints](Custom-Model-Endpoints)
- [Working With Files](Working-With-Files)
- [Input And Voice](Input-And-Voice)
- [Mobile Guide](Mobile-Guide)
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.30.0",
"version": "1.31.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.30.0",
"version": "1.31.0",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.30.0",
"version": "1.31.0",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
+1 -1
View File
@@ -1,7 +1,7 @@
{
"name": "codeman",
"description": "Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.30.0",
"version": "1.31.0",
"author": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
+75 -26
View File
@@ -47,7 +47,7 @@ later call opens with, and your first REAL call performs them anyway:
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
```
⚠️ **Never spend a Bash call on this check alone.** §1's block opens with this same
@@ -75,8 +75,8 @@ PRE="${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh"
mkdir -p "$(dirname "$PRE")"
# Rewrite unless the file already ends with THIS version's stamp, so a stale or a
# half-written file self-heals here instead of costing you a round trip to rm it.
grep -qs '^CODEMAN_PREAMBLE=1.22.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
grep -qs '^CODEMAN_PREAMBLE=1.30.1$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -150,6 +150,27 @@ _trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
# ---- the composer: is the prompt still sitting there, unsent? ----
# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the
# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that
# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt
# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through
# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS
# the composer and keeps pressing Enter while the prompt is still there.
_composer_text() { # <sid> -> the composer's text with ALL whitespace removed: "" once
# the prompt was taken, "?" when the pane shows no composer at all. The composer is
# the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher
# up in the transcript, so only the last one says whether the text was taken.
local t
t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1)
[ -n "$t" ] || { printf '?'; return 0; }
# Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does
# not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH).
printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g"
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
@@ -262,12 +283,16 @@ spawn_workers() {
# worker a silent no-op that still "succeeds" and reports the previous turn's state.
# Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a
# deliberate duplicate, at the SAME number (§5.3).
# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the
# typed prompt stranded on the composer while a long wait runs its whole timeout
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude
# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints),
# leaving the typed prompt stranded on the composer while a long wait runs its whole
# timeout (observed live, twelve reviews in a row). So the first wait is short; on its
# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate:
# the server re-waits without retyping, §5.3) and kept open in the background, while
# the composer is READ (_composer_text) and, as long as the prompt is still sitting
# there, a bare \r goes out about every ten seconds, up to twelve times. An empty
# composer ends the loop, so a prompt that was taken is never nudged again, and the
# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
@@ -275,7 +300,7 @@ spawn_workers() {
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
@@ -290,16 +315,38 @@ sendwait() {
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
# ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call,
# in the background, while the Enter loop below works the composer. Signals have
# no history: a `stop` that fires while no wait is open (during a composer read
# between two short waits, measured) is lost, and the next wait then runs its
# whole timeout on a turn that already ended. The resend is a tagged DUPLICATE,
# so the server skips the write and re-waits without retyping (§5.3).
tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" &
bg=$!
# The prompt's head with whitespace removed, matched literally (the "$head"
# quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob).
head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24)
while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended
c=$(_composer_text "$sid")
if [ "$c" = '?' ]; then
[ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it
elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then
break # composer empty (taken) or holding other text
fi
n=$((n+1))
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done
done
wait "$bg"
# The duplicate reports `delivered:false` -- truthfully, but about the wrong send.
# The first one delivered, so carry that forward, or §1's cleanup reads a completed
# turn as an undelivered one and keeps a finished worker forever.
r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp")
rm -f "$tmp"
fi
printf '%s\n' "$r"
}
@@ -325,10 +372,10 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.22.0
CODEMAN_PREAMBLE=1.30.1
PREAMBLE
)
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
```
Every later Bash call that touches the API starts with the same two loader lines from
@@ -379,7 +426,7 @@ and no per-call body to hand-build.
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null # §0 loader
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
N=(alpha beta) # INVENT one fresh case name per worker; never list cases first
# (a name may carry a mode: `beta:deepseek`, see below)
T=('reply with one line: the absolute path of your working directory'
@@ -440,9 +487,11 @@ Four things this block leans on, each one link away, no detour needed to run it:
skill: §5.1. Those workspaces do get hooks now, unless the operator disabled it.
- `sendwait` supplies the `\r`, picks a fresh `seq`, and self-heals a stranded Enter.
A prompt without the `\r` is never submitted (§3), a reused `seq` is silently
swallowed as an already-applied duplicate, and an Enter eaten by an Ink repaint
strands the prompt on the composer until a bare `\r` follows: all three are reasons
to let `sendwait` build the call rather than hand-rolling it.
swallowed as an already-applied duplicate, and a lost Enter strands the prompt on the
composer until a bare `\r` follows: Claude Code 2.1.277 and later ignore Enter for the
first 30 to 50 seconds after the composer paints while still taking the text, so
`sendwait` reads the composer and keeps pressing Enter until the prompt has left it.
All three are reasons to let `sendwait` build the call rather than hand-rolling it.
- Each `sendwait` costs that worker one billed turn, as does every prompt you send it.
- Deleting the sessions does **not** remove the case directories. They are marked as
agent-created, so `GET /api/v1/cases/agent-created` lists them for cleanup: §5.14.
+66 -19
View File
@@ -1,4 +1,4 @@
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -72,6 +72,27 @@ _trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
# ---- the composer: is the prompt still sitting there, unsent? ----
# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the
# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that
# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt
# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through
# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS
# the composer and keeps pressing Enter while the prompt is still there.
_composer_text() { # <sid> -> the composer's text with ALL whitespace removed: "" once
# the prompt was taken, "?" when the pane shows no composer at all. The composer is
# the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher
# up in the transcript, so only the last one says whether the text was taken.
local t
t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1)
[ -n "$t" ] || { printf '?'; return 0; }
# Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does
# not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH).
printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g"
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
@@ -184,12 +205,16 @@ spawn_workers() {
# worker a silent no-op that still "succeeds" and reports the previous turn's state.
# Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a
# deliberate duplicate, at the SAME number (§5.3).
# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the
# typed prompt stranded on the composer while a long wait runs its whole timeout
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude
# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints),
# leaving the typed prompt stranded on the composer while a long wait runs its whole
# timeout (observed live, twelve reviews in a row). So the first wait is short; on its
# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate:
# the server re-waits without retyping, §5.3) and kept open in the background, while
# the composer is READ (_composer_text) and, as long as the prompt is still sitting
# there, a bare \r goes out about every ten seconds, up to twelve times. An empty
# composer ends the loop, so a prompt that was taken is never nudged again, and the
# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
@@ -197,7 +222,7 @@ spawn_workers() {
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
@@ -212,16 +237,38 @@ sendwait() {
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
# ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call,
# in the background, while the Enter loop below works the composer. Signals have
# no history: a `stop` that fires while no wait is open (during a composer read
# between two short waits, measured) is lost, and the next wait then runs its
# whole timeout on a turn that already ended. The resend is a tagged DUPLICATE,
# so the server skips the write and re-waits without retyping (§5.3).
tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" &
bg=$!
# The prompt's head with whitespace removed, matched literally (the "$head"
# quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob).
head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24)
while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended
c=$(_composer_text "$sid")
if [ "$c" = '?' ]; then
[ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it
elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then
break # composer empty (taken) or holding other text
fi
n=$((n+1))
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done
done
wait "$bg"
# The duplicate reports `delivered:false` -- truthfully, but about the wrong send.
# The first one delivered, so carry that forward, or §1's cleanup reads a completed
# turn as an undelivered one and keeps a finished worker forever.
r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp")
rm -f "$tmp"
fi
printf '%s\n' "$r"
}
@@ -247,4 +294,4 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.22.0
CODEMAN_PREAMBLE=1.30.1
@@ -21,7 +21,7 @@ by sourcing the preamble file the §0 bootstrap wrote, and checking its version
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
```
Do **not** re-paste the preamble body into each call. Sourcing it is what retires the
+75 -26
View File
@@ -47,7 +47,7 @@ later call opens with, and your first REAL call performs them anyway:
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
```
⚠️ **Never spend a Bash call on this check alone.** §1's block opens with this same
@@ -75,8 +75,8 @@ PRE="${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh"
mkdir -p "$(dirname "$PRE")"
# Rewrite unless the file already ends with THIS version's stamp, so a stale or a
# half-written file self-heals here instead of costing you a round trip to rm it.
grep -qs '^CODEMAN_PREAMBLE=1.22.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
grep -qs '^CODEMAN_PREAMBLE=1.30.1$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -150,6 +150,27 @@ _trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
# ---- the composer: is the prompt still sitting there, unsent? ----
# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the
# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that
# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt
# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through
# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS
# the composer and keeps pressing Enter while the prompt is still there.
_composer_text() { # <sid> -> the composer's text with ALL whitespace removed: "" once
# the prompt was taken, "?" when the pane shows no composer at all. The composer is
# the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher
# up in the transcript, so only the last one says whether the text was taken.
local t
t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1)
[ -n "$t" ] || { printf '?'; return 0; }
# Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does
# not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH).
printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g"
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
@@ -262,12 +283,16 @@ spawn_workers() {
# worker a silent no-op that still "succeeds" and reports the previous turn's state.
# Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a
# deliberate duplicate, at the SAME number (§5.3).
# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the
# typed prompt stranded on the composer while a long wait runs its whole timeout
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude
# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints),
# leaving the typed prompt stranded on the composer while a long wait runs its whole
# timeout (observed live, twelve reviews in a row). So the first wait is short; on its
# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate:
# the server re-waits without retyping, §5.3) and kept open in the background, while
# the composer is READ (_composer_text) and, as long as the prompt is still sitting
# there, a bare \r goes out about every ten seconds, up to twelve times. An empty
# composer ends the loop, so a prompt that was taken is never nudged again, and the
# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
@@ -275,7 +300,7 @@ spawn_workers() {
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
@@ -290,16 +315,38 @@ sendwait() {
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
# ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call,
# in the background, while the Enter loop below works the composer. Signals have
# no history: a `stop` that fires while no wait is open (during a composer read
# between two short waits, measured) is lost, and the next wait then runs its
# whole timeout on a turn that already ended. The resend is a tagged DUPLICATE,
# so the server skips the write and re-waits without retyping (§5.3).
tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" &
bg=$!
# The prompt's head with whitespace removed, matched literally (the "$head"
# quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob).
head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24)
while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended
c=$(_composer_text "$sid")
if [ "$c" = '?' ]; then
[ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it
elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then
break # composer empty (taken) or holding other text
fi
n=$((n+1))
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done
done
wait "$bg"
# The duplicate reports `delivered:false` -- truthfully, but about the wrong send.
# The first one delivered, so carry that forward, or §1's cleanup reads a completed
# turn as an undelivered one and keeps a finished worker forever.
r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp")
rm -f "$tmp"
fi
printf '%s\n' "$r"
}
@@ -325,10 +372,10 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.22.0
CODEMAN_PREAMBLE=1.30.1
PREAMBLE
)
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
```
Every later Bash call that touches the API starts with the same two loader lines from
@@ -379,7 +426,7 @@ and no per-call body to hand-build.
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null # §0 loader
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
N=(alpha beta) # INVENT one fresh case name per worker; never list cases first
# (a name may carry a mode: `beta:deepseek`, see below)
T=('reply with one line: the absolute path of your working directory'
@@ -440,9 +487,11 @@ Four things this block leans on, each one link away, no detour needed to run it:
skill: §5.1. Those workspaces do get hooks now, unless the operator disabled it.
- `sendwait` supplies the `\r`, picks a fresh `seq`, and self-heals a stranded Enter.
A prompt without the `\r` is never submitted (§3), a reused `seq` is silently
swallowed as an already-applied duplicate, and an Enter eaten by an Ink repaint
strands the prompt on the composer until a bare `\r` follows: all three are reasons
to let `sendwait` build the call rather than hand-rolling it.
swallowed as an already-applied duplicate, and a lost Enter strands the prompt on the
composer until a bare `\r` follows: Claude Code 2.1.277 and later ignore Enter for the
first 30 to 50 seconds after the composer paints while still taking the text, so
`sendwait` reads the composer and keeps pressing Enter until the prompt has left it.
All three are reasons to let `sendwait` build the call rather than hand-rolling it.
- Each `sendwait` costs that worker one billed turn, as does every prompt you send it.
- Deleting the sessions does **not** remove the case directories. They are marked as
agent-created, so `GET /api/v1/cases/agent-created` lists them for cleanup: §5.14.
+66 -19
View File
@@ -1,4 +1,4 @@
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
# ---- Codeman agent preamble 1.30.1 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -72,6 +72,27 @@ _trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
# ---- the composer: is the prompt still sitting there, unsent? ----
# ⚠️ Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment the
# composer paints but IGNORES Enter for the first 30-50 seconds after it: the \r that
# Codeman sends 50 ms after the text and a lone nudge at 20 s both leave the prompt
# stranded, with `0 tokens`, while the wait burns its whole timeout. Measured through
# this very route: Enter at 28 s stranded, Enter at 51 s submitted. So sendwait READS
# the composer and keeps pressing Enter while the prompt is still there.
_composer_text() { # <sid> -> the composer's text with ALL whitespace removed: "" once
# the prompt was taken, "?" when the pane shows no composer at all. The composer is
# the LAST `❯` line: Claude Code echoes a submitted prompt with the same glyph higher
# up in the transcript, so only the last one says whether the text was taken.
local t
t=$("${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d '\r' | grep -a '^[[:space:]]*❯' | tail -1)
[ -n "$t" ] || { printf '?'; return 0; }
# Claude Code draws a NO-BREAK SPACE (U+00A0) after the glyph, which [:space:] does
# not cover, so it is stripped by its bytes, portably (BSD sed has no \xHH).
printf '%s' "$t" | sed 's/^[[:space:]]*❯//' | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g"
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
@@ -184,12 +205,16 @@ spawn_workers() {
# worker a silent no-op that still "succeeds" and reports the previous turn's state.
# Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a
# deliberate duplicate, at the SAME number (§5.3).
# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the
# typed prompt stranded on the composer while a long wait runs its whole timeout
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# Delivery is SELF-HEALING: the Enter can be lost (an Ink repaint eats it, and Claude
# Code 2.1.277+ ignores it outright for the first 30-50 s after the composer paints),
# leaving the typed prompt stranded on the composer while a long wait runs its whole
# timeout (observed live, twelve reviews in a row). So the first wait is short; on its
# timeout the ORIGINAL frame is resent unchanged as a long re-wait (a tagged duplicate:
# the server re-waits without retyping, §5.3) and kept open in the background, while
# the composer is READ (_composer_text) and, as long as the prompt is still sitting
# there, a bare \r goes out about every ten seconds, up to twelve times. An empty
# composer ends the loop, so a prompt that was taken is never nudged again, and the
# wait that was open the whole time is what reports the turn's end. Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
@@ -197,7 +222,7 @@ spawn_workers() {
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r c head n=0 tmp bg i
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
@@ -212,16 +237,38 @@ sendwait() {
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
# ⚠️ The long re-wait is registered FIRST and stays open for the rest of this call,
# in the background, while the Enter loop below works the composer. Signals have
# no history: a `stop` that fires while no wait is open (during a composer read
# between two short waits, measured) is lost, and the next wait then runs its
# whole timeout on a turn that already ended. The resend is a tagged DUPLICATE,
# so the server skips the write and re-waits without retyping (§5.3).
tmp=$(mktemp "${TMPDIR:-/tmp}/codeman-wait.XXXXXX") || return 1
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" > "$tmp" &
bg=$!
# The prompt's head with whitespace removed, matched literally (the "$head"
# quoting inside ${c#...} keeps a * or ? in the prompt from acting as a glob).
head=$(printf '%s' "$p" | tr -d '[:space:]' | sed "s/$(printf '\302\240')//g" | head -c 24)
while [ "$n" -lt 12 ] && [ ! -s "$tmp" ]; do # a non-empty file means the wait ended
c=$(_composer_text "$sid")
if [ "$c" = '?' ]; then
[ "$n" -eq 0 ] || break # unreadable pane: one Enter, then trust it
elif [ -z "$head" ] || [ "${c#"$head"}" = "$c" ]; then
break # composer empty (taken) or holding other text
fi
n=$((n+1))
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
i=0; while [ "$i" -lt 10 ] && [ ! -s "$tmp" ]; do sleep 1; i=$((i+1)); done
done
wait "$bg"
# The duplicate reports `delivered:false` -- truthfully, but about the wrong send.
# The first one delivered, so carry that forward, or §1's cleanup reads a completed
# turn as an undelivered one and keeps a finished worker forever.
r=$(jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end' < "$tmp")
rm -f "$tmp"
fi
printf '%s\n' "$r"
}
@@ -247,4 +294,4 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.22.0
CODEMAN_PREAMBLE=1.30.1
+1 -1
View File
@@ -21,7 +21,7 @@ by sourcing the preamble file the §0 bootstrap wrote, and checking its version
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.30.1 ] || { echo "preamble missing or stale; re-run the §0 bootstrap"; exit 1; }
```
Do **not** re-paste the preamble body into each call. Sourcing it is what retires the
+29
View File
@@ -340,6 +340,35 @@ const capabilitiesSchema = z
// an env var, so it declares baseUrl/apiKey injection with no model var at all.
modelVars: z.array(envName).max(8),
launchModel: launchModelTemplate,
// Optional: the env var to carry a discovered per-model context-window size
// (claude's CLAUDE_CODE_MAX_CONTEXT_TOKENS), and/or the env var that isolates
// this session's config/credential directory from the user's real one (claude's
// CLAUDE_CONFIG_DIR) so an injected API key never collides with a stored OAuth
// session. See the customModelInjection doc comment in cli-registry/types.ts.
contextLengthVar: envName.optional(),
configDirVar: envName.optional(),
// Relative path, WITHIN the isolated configDirVar directory, of a trust-dialog
// seed file the CLI itself owns the shape of — claude's `.claude.json`
// `customApiKeyResponses.approved` list, the same field an interactive "Detected
// a custom API key — use it?" prompt writes to on a real terminal. Only makes
// sense alongside configDirVar (an isolated, otherwise-empty directory has none
// of a real profile's prior approvals), and only implemented for the
// 'claude-api-key-responses' shape today — see custom-model-injection-apply.ts.
apiKeyTrustFile: z
.object({ relPath: z.string().min(1).max(80), shape: z.literal('claude-api-key-responses') })
.strict()
.optional(),
// An isolated config directory replays the CLI's whole first-run sequence (theme
// picker, security notes, per-project trust dialog, bypass-permissions warning)
// on every launch, same root cause as apiKeyTrustFile above — this reuses that
// same file to pre-seed the state a real, already-onboarded profile carries. See
// the customModelInjection doc comment in cli-registry/types.ts.
skipFirstRunPrompts: z.boolean().optional(),
// DeepSeek-only, confirmed by reading its own bundled SDK source: it concatenates
// "/chat/completions" onto baseUrlVar's value with no "/v1" of its own, while
// llama-swap/llama.cpp only serves the "/v1/..." path — claude/gemini must NOT
// get this. See the customModelInjection doc comment in cli-registry/types.ts.
appendV1Suffix: z.boolean().optional(),
})
.strict(),
z
+62 -7
View File
@@ -228,15 +228,31 @@ const CLAUDE: CliEntry = {
privilegedParams: [],
// ANTHROPIC_* is NOT in allowedPrefixes/allowedKeys above (deliberately — see the
// allowedPrefixes comment nearby), so these are unreachable via plain envOverrides
// today; listed here only so the dedicated custom-model route (docs/custom-model-endpoints-plan.md
// chunk 5) clamps them for a non-granted multi-user owner the same way every other
// CLI's injection vars are clamped, the day that route widens who can set them.
// today. privilegedEnvKeys has exactly one consumer, ownerClampedEnvKeys() in
// session-env-clamp.ts, which feeds the generic envOverrides clamp on
// POST /api/sessions, POST /api/quick-start and reboot-restore — no custom-model
// route reads this field at all, and the values it injects are merged in AFTER
// that clamp runs regardless of what's listed here.
privilegedEnvKeys: [
'ANTHROPIC_BASE_URL',
'ANTHROPIC_API_KEY',
'ANTHROPIC_DEFAULT_SONNET_MODEL',
'ANTHROPIC_DEFAULT_HAIKU_MODEL',
'ANTHROPIC_DEFAULT_OPUS_MODEL',
// CLAUDE_CODE_MAX_CONTEXT_TOKENS already matches the CLAUDE_CODE_* allowedPrefix, and
// CLAUDE_CONFIG_DIR is already an allowed exact key (docs/wiki/Agent-CLIs.md), so both
// were already reachable via plain envOverrides before this pair existed and this
// feature does not strictly need either listed. They stay listed anyway, because
// types.ts's rule ("every traffic-redirecting var this feature introduces MUST also
// appear in privilegedEnvKeys") is meant to hold literally, not with an exception
// carved out for the two vars that happen not to need it today. The real
// consequence lands on the GENERIC envOverrides clamp above, not on this feature:
// a non-granted multi-user owner can no longer set CLAUDE_CONFIG_DIR through
// envOverrides at all (the per-client-account override, #255), and a PERSISTED one
// is now stripped on reboot-restore for such an owner too — see
// session-env-clamp.ts's own fileoverview.
'CLAUDE_CODE_MAX_CONTEXT_TOKENS',
'CLAUDE_CONFIG_DIR',
],
gates: { nameFlag: { minVersion: '2.1.224', failClosed: true } },
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — verified by hand against a real
@@ -247,6 +263,32 @@ const CLAUDE: CliEntry = {
baseUrlVar: 'ANTHROPIC_BASE_URL',
apiKeyVar: 'ANTHROPIC_API_KEY',
modelVars: ['ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL'],
// Verified via Claude Code's own docs: CLAUDE_CODE_MAX_CONTEXT_TOKENS overrides the
// assumed context window and applies directly for a model name Claude Code doesn't
// recognize as one of its own — exactly the custom-model case. Without it, Claude Code
// assumes a large (200k) window for any unrecognized model id and never compacts,
// eventually overflowing a much smaller real local context (see plan doc reasoning
// above the interface for the confirmed failure).
contextLengthVar: 'CLAUDE_CODE_MAX_CONTEXT_TOKENS',
// Isolates this session's config/credential directory so an injected ANTHROPIC_API_KEY
// never shares a directory with a stored claude.ai OAuth login — see the doc comment on
// customModelInjection in cli-registry/types.ts for the traded-off side effect.
configDirVar: 'CLAUDE_CONFIG_DIR',
// ⚠️ Required alongside configDirVar, not optional in practice: verified live that an
// isolated, otherwise-empty config directory makes claude stop at an interactive
// "Detected a custom API key — use it?" prompt on EVERY launch, defaulting to "No" with
// no one at the TTY to answer — silently refusing the very key this feature injected.
// Pre-seeding this file's customApiKeyResponses.approved list (verified against a real
// ~/.claude.json after answering the prompt once by hand) answers it in advance instead.
apiKeyTrustFile: { relPath: '.claude.json', shape: 'claude-api-key-responses' },
// ⚠️ Same isolated-directory root cause, one step further: verified live that on top
// of the API-key prompt above, a fresh CLAUDE_CONFIG_DIR also replays claude's ENTIRE
// first-run sequence on every launch — the theme picker, the security-notes screen,
// the per-project "trust this folder?" dialog, and (running with
// --dangerously-skip-permissions) a one-time bypass-permissions warning — none of
// which a real, already-onboarded profile shows again. Pre-seeds that same
// already-onboarded state instead of leaving a human to click through it.
skipFirstRunPrompts: true,
},
},
overlays: {
@@ -1071,15 +1113,28 @@ const DEEPSEEK: CliEntry = {
// privilege rather than granting it, and clamping it here was a real regression
// (test/deepseek-mode.test.ts) fixed before this shipped.
privilegedEnvKeys: ['DSH_PERMISSION_MODE', 'DSH_HOME', 'DEEPSEEK_BASE_URL'],
// Web-researched, unverified, partial: reuses the already-existing DEEPSEEK_BASE_URL/
// DEEPSEEK_API_KEY keys above. No modelVars — dsh's model is a profile-composition
// entry (see `model: { source: 'none' }` above), not an env var, so forcing a specific
// model name may not fully work; verify against a real profile before shipping.
// Reuses the already-existing DEEPSEEK_BASE_URL/DEEPSEEK_API_KEY keys above. No
// modelVars — dsh's model is a profile-composition entry (see `model: { source: 'none'
// }` above), not an env var, so forcing a specific model name may not fully work;
// verify against a real profile before shipping.
//
// ⚠️ appendV1Suffix is REQUIRED, not optional-nice-to-have: without it every request
// 404s. Confirmed live and by reading dsh's own bundled source
// (@deepseek-ai/dsh-llm-deepseek): it builds the request URL as
// `${DEEPSEEK_BASE_URL}/chat/completions` with no "/v1" of its own (its real public
// API, https://api.deepseek.com, expects the caller's base URL to already carry any
// needed prefix), while llama-swap/llama.cpp only serves the OpenAI-conventional
// "/v1/chat/completions" — a bare POST to ".../chat/completions" 404s live, and the
// 404 reported here originally ("dsh: HTTP_404: DeepSeek API error (HTTP 404)")
// matches dsh's own error-message template for exactly this failure. See the
// customModelInjection doc comment in cli-registry/types.ts for the full reasoning,
// including why claude/gemini must NOT get this.
customModelInjection: {
kind: 'env',
baseUrlVar: 'DEEPSEEK_BASE_URL',
apiKeyVar: 'DEEPSEEK_API_KEY',
modelVars: [],
appendV1Suffix: true,
},
},
overlays: {
+66 -1
View File
@@ -496,9 +496,74 @@ export interface CliCapabilities {
* declares). Absent = the config alone selects the model (claude's env vars,
* opencode's blob, codex's top-level `model` key). Applied by the session's
* respawn options through the entry's `legacyConfigField`, never by id.
*
* `contextLengthVar` (env kind only): the env var a discovered per-model context-window
* size is written to when known (claude's `CLAUDE_CODE_MAX_CONTEXT_TOKENS`) — without it,
* a CLI that assumes a large default window for an unrecognized model name keeps sending
* full-size prompts against a much smaller local server and eventually overflows its real
* context (verified: a 33.7K-token system prompt against a 16384-token llama-swap model).
* Absent when the CLI has no such override, or the value is unknown for this model.
*
* `configDirVar` (env kind only): the env var that redirects this session's config/
* credential directory to an isolated, per-session one (claude's `CLAUDE_CONFIG_DIR`), so
* an injected API key never coexists with a stored claude.ai OAuth session in the same
* directory — the CLI still warns "both claude.ai and ANTHROPIC_API_KEY set" when they
* share a directory even though the API key wins for actual requests. Isolating it trades
* that cosmetic warning for a documented side effect: a relocated config directory writes
* transcripts outside `~/.claude/projects`, blinding the response viewer, subagent
* windows, and Read My Mind for that session (see docs/wiki/Agent-CLIs.md).
*
* `apiKeyTrustFile` (env kind only, alongside configDirVar): an isolated config directory
* has none of a real profile's prior "detected a custom API key, use it?" approvals, so
* without this the CLI stops and asks interactively on every single launch — with no one
* at a TTY to answer, that's a hang, not a warning (confirmed live: claude's own default
* answer, "No", would silently refuse to use the very key this feature just injected).
* `relPath`/`shape` name the file (claude's `.claude.json`) and its
* `customApiKeyResponses.approved` field this pre-seeds — the exact field a real answered
* prompt itself writes to, so this isn't bypassing the check, just answering it the same
* way a one-off prior approval on a shared profile already would.
*
* `skipFirstRunPrompts` (env kind only, alongside apiKeyTrustFile): an isolated config
* directory is not just missing API-key approvals — it is a brand-new profile as far as
* the CLI is concerned, so it also replays its ENTIRE first-run sequence on every launch:
* the theme picker, the security-notes screen, the per-project "trust this folder?"
* dialog, and (running with a bypass-permissions flag) a one-time warning about it —
* confirmed live, none of which a real, long-used profile ever shows again. `true`
* pre-seeds the same state a real profile accumulates from having answered all of that
* once: `hasCompletedOnboarding` and the launching session's own project entry in the
* `apiKeyTrustFile` (claude's `.claude.json`), plus `skipDangerousModePermissionPrompt`
* in claude's `settings.json` — see `seedFirstRunState`/`seedSkipBypassPermissionsPrompt`
* in custom-model-injection-apply.ts. Requires `apiKeyTrustFile` to be set too, since it
* reuses that file.
*
* `appendV1Suffix` (env kind only): the raw `endpoint.baseUrl` gets `withV1Suffix()`
* applied before being written to `baseUrlVar`, instead of being used verbatim.
* DeepSeek needs this and claude/gemini must NOT get it — a per-CLI asymmetry confirmed
* by reading each SDK's own request-building source, not assumed: DeepSeek Harness's
* bundled `@deepseek-ai/dsh-llm-deepseek` concatenates `${connection.baseURL}/chat/
* completions` with no `/v1` insertion of its own (its real public API base,
* `https://api.deepseek.com`, expects the caller's base URL to already carry any
* needed prefix), while llama-swap/llama.cpp only ever serves the OpenAI-conventional
* `/v1/chat/completions` — confirmed live: a bare `POST <baseUrl>/chat/completions`
* 404s, `POST <baseUrl>/v1/chat/completions` succeeds, and the harness's own error
* message template (`DeepSeek API error (HTTP ${status})`) reproduces the exact
* `HTTP_404` this feature originally shipped with unexplained. Claude Code's own SDK,
* by contrast, was already confirmed working end-to-end against the RAW `baseUrl` with
* no suffix — appending one there would be wrong, not just redundant.
*/
customModelInjection:
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[]; launchModel?: string }
| {
kind: 'env';
baseUrlVar: string;
apiKeyVar: string;
modelVars: string[];
launchModel?: string;
contextLengthVar?: string;
apiKeyTrustFile?: { relPath: string; shape: 'claude-api-key-responses' };
configDirVar?: string;
skipFirstRunPrompts?: boolean;
appendV1Suffix?: boolean;
}
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json'; launchModel?: string }
| {
kind: 'configDir';
+23
View File
@@ -0,0 +1,23 @@
/**
* @fileoverview Limits shared between Wake-on-LAN parsing and its request schema.
*
* Its own module because `src/remote-wake.ts` is import-fenced: only
* `web/routes/session-routes.ts` and `web/server.ts` may import it, so that no
* watcher or boot-recovery path can WAKE a host (pinned by the wiring guard in
* `test/remote-wake.test.ts`). `web/schemas.ts` needs the same MAC-count limit and
* must not become a third importer, and it would drag `dgram`/`net`/`child_process`
* into every module that validates a request body. A plain constant satisfies both.
*/
/**
* How many comma-separated MACs one `wakeMac` may carry.
*
* ⚠ Single source for `parseMacList()` and `RemoteHostSchema.wakeMac`. The two used
* to disagree: the schema's 128-character cap admits seven MACs while the parser
* rejected more than four all-or-nothing, so a five-MAC value validated, persisted to
* `remote-hosts.json`, and then resolved to NO wake target. The host read as
* unconfigured and the banner offered "Configure WoL" for a host the user had just
* configured, which is the worst shape a validation gap can take: accepted, stored,
* silently inert.
*/
export const MAX_WAKE_MACS = 4;
+32
View File
@@ -40,6 +40,38 @@ export interface CustomModelHost {
authStyle?: CustomModelAuthStyle;
models?: string[];
lastDiscoveredAt?: string;
/**
* The model the Run-menu picker (docs/custom-model-endpoints-plan.md) applies when
* this endpoint is picked with no further choice — one generated menu entry per
* (CLI, endpoint) pair, not per (CLI, endpoint, model), so it needs a single answer.
* Must be a member of `models` when set; the picker falls back to `models[0]` when
* this is unset, and disables the entry entirely when `models` is empty (nothing to
* default to). Never auto-set on discovery — the previous default staying valid
* after a re-discover is a property worth keeping even if the model list changes.
*/
defaultModelId?: string;
/**
* Discovered context-window size (tokens) per model id, keyed by the same strings as
* `models`. Populated opportunistically during discovery (`custom-model-routes.ts`) from
* llama.cpp/llama-swap's `GET /props?model=<id>` — the plain OpenAI-shaped `/v1/models`
* response has no such field. Only ever probed for a model the server already reports as
* loaded (llama-swap's `status.value === 'loaded'`); an unloaded one is deliberately never
* probed, since llama-swap treats `/props?model=` as a routing hint that can trigger an
* actual (slow, GPU-swapping) model load as a side effect of merely asking. A model this
* has no entry for simply gets no context-length env override applied — never a guess.
*/
modelContextLengths?: Record<string, number>;
/**
* Discovered file size (GB) per model id, keyed by the same strings as `models`.
* Populated during discovery by parsing llama-swap's own `description` field for an
* auto-discovered model ("Auto-discovered 16.35 GB - parameters auto-fitted by
* llama.cpp") — a hand-configured profile's own description has no such figure and
* correctly gets no entry, never a guess. Used only to label the Run-menu picker's
* "loading model" banner with a rough, unmeasured expected-time estimate
* (the Run-menu picker's loading banner in session-ui.js) — never a guarantee, and never anything a
* server-side check relies on.
*/
modelSizesGB?: Record<string, number>;
}
export function customModelHostsPath(configDir: string): string {
+197 -5
View File
@@ -10,7 +10,8 @@
* cli-registry changes" requirement it was written against.
*/
import { chmodSync, mkdirSync, writeFileSync, rmSync } from 'node:fs';
import { chmodSync, existsSync, mkdirSync, readFileSync, writeFileSync, rmSync, symlinkSync } from 'node:fs';
import { homedir, platform } from 'node:os';
import { join, dirname } from 'node:path';
import { dataPath } from './config/instance.js';
import type { CliEntry } from './config/cli-registry/types.js';
@@ -48,6 +49,168 @@ export function applyConfigDirInjection(baseDir: string, injection: ConfigDirInj
return { [injection.dirEnvVar]: baseDir, ...injection.extraEnv };
}
/**
* Real, shared Claude config directory Codeman's own host process runs under — honors
* `CLAUDE_CONFIG_DIR` the same way `claude-credentials.ts`'s `claudeCredentialsPath()`
* does, so the symlink below points at wherever `~/.claude/projects` actually lives
* rather than assuming the plain default.
*/
function realClaudeConfigDir(): string {
const configured = typeof process.env.CLAUDE_CONFIG_DIR === 'string' && process.env.CLAUDE_CONFIG_DIR.trim();
return configured || join(homedir(), '.claude');
}
/**
* Symlinks `<isolatedDir>/projects` back to the real, shared `~/.claude/projects`, so an
* isolated `CLAUDE_CONFIG_DIR` (used to keep an injected API key away from a stored OAuth
* session — see `configDirVar` on customModelInjection) doesn't also blind the response
* viewer, subagent windows, and Read My Mind for that session (docs/wiki/Agent-CLIs.md).
* Best-effort: a platform that refuses symlinks (unprivileged Windows without a junction
* fallback working, e.g.) just keeps the pre-existing documented side effect instead of
* failing the whole custom-model apply over a nice-to-have.
*/
function linkSharedProjectsDir(isolatedDir: string): void {
const link = join(isolatedDir, 'projects');
if (existsSync(link)) return; // already linked (idempotent re-apply) or real dir wrote one
try {
symlinkSync(join(realClaudeConfigDir(), 'projects'), link, platform() === 'win32' ? 'junction' : 'dir');
} catch {
// best-effort only — response viewer/subagent windows go blind for this session instead
}
}
/**
* Pre-approves the injected API key in an isolated config directory's trust-dialog state
* (`customModelInjection.apiKeyTrustFile`), so an otherwise-empty directory doesn't make the
* CLI stop at an interactive "Detected a custom API key — use it?" prompt on every single
* launch. Confirmed live: with nobody at the TTY to answer, that prompt's own default
* ("No") silently refuses the very key this feature just injected — this isn't bypassing
* the check, it's answering it the same field a real answered prompt itself writes to
* (verified against a real `~/.claude.json` after answering by hand once).
*
* Merges rather than overwrites: the file may already carry fields the CLI itself wrote on
* an earlier launch in this same isolated directory (machineID, userID, other approved
* keys), and a corrupt or partially-written file (a crash mid-write) is treated as absent
* rather than failing the whole apply over a nice-to-have.
*/
/**
* The form Claude Code actually stores an approved key in: the trimmed last 20
* characters. Mirrors the CLI's own `e.trim().slice(-20)`, which is applied on BOTH
* the write and the lookup, so anything else never matches.
*/
export function truncateApiKeyForTrustFile(apiKey: string): string {
return apiKey.trim().slice(-20);
}
function seedApiKeyTrustFile(
configDir: string,
trustFile: { relPath: string; shape: 'claude-api-key-responses' },
apiKey: string
): void {
const filePath = join(configDir, trustFile.relPath);
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(readFileSync(filePath, 'utf8')) as Record<string, unknown>;
} catch {
existing = {};
}
const responses = (existing.customApiKeyResponses ?? {}) as { approved?: unknown; rejected?: unknown };
const approved = new Set(Array.isArray(responses.approved) ? (responses.approved as string[]) : []);
// ⚠ Claude Code stores and compares only the LAST 20 CHARACTERS of a key, never the
// whole thing: its lookup is `approved.includes(key.trim().slice(-20))` (decompiled
// from the 2.1.278 bundle, and corroborated by real `~/.claude.json` files, whose
// customApiKeyResponses entries are all exactly 20 characters). Seeding the full key
// therefore never matches for a REAL key, and claude stops at the interactive
// "Detected a custom API key in your environment" prompt, whose default is
// "No (recommended)" — so the launch hangs or silently refuses the key this feature
// just injected. It went unnoticed because a keyless llama.cpp/llama-swap endpoint
// uses DEFAULT_API_KEY ('local-dummy-key', 15 chars), where slice(-20) is the whole
// string and the seed matches by accident. Truncating here also keeps a full
// third-party credential from being written into a second file on disk.
approved.add(truncateApiKeyForTrustFile(apiKey));
const rejected = Array.isArray(responses.rejected) ? responses.rejected : [];
existing.customApiKeyResponses = { approved: [...approved], rejected };
try {
writeFileSync(filePath, JSON.stringify(existing, null, 2), { encoding: 'utf8', mode: 0o600 });
chmodSync(filePath, 0o600);
} catch {
// best-effort only — the interactive prompt returns instead of a hard failure here
}
}
/**
* Pre-seeds the two remaining pieces of "already been onboarded" state a fresh
* `CLAUDE_CONFIG_DIR` has none of (`customModelInjection.skipFirstRunPrompts`, alongside
* apiKeyTrustFile): claude replays its whole first-run sequence — the theme picker, the
* security-notes screen, and (per-project) the "trust this folder?" dialog — against ANY
* config directory that has never completed it, confirmed live against a genuinely fresh
* isolated directory. `hasCompletedOnboarding` skips the theme/security-notes screens
* outright; `projects[workingDir].hasTrustDialogAccepted` answers the trust dialog for
* THIS session's own working directory the same way a real profile's own prior approval
* would — other projects in the file are left alone, and `workingDir` is used verbatim
* (never realpath'd or slash-normalized) since that's the literal string claude itself
* uses as the project key, being whatever string the session was actually launched with
* as its cwd.
*
* Same merge-not-overwrite and corrupt-file-tolerant behavior as `seedApiKeyTrustFile`
* (same file, so a second sequential read-modify-write here is deliberate rather than
* folding both into one pass — keeps each seed independently testable and optional).
*/
function seedFirstRunOnboardingState(
configDir: string,
trustFile: { relPath: string; shape: 'claude-api-key-responses' },
workingDir: string
): void {
const filePath = join(configDir, trustFile.relPath);
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(readFileSync(filePath, 'utf8')) as Record<string, unknown>;
} catch {
existing = {};
}
existing.hasCompletedOnboarding = true;
const projects =
existing.projects && typeof existing.projects === 'object' && !Array.isArray(existing.projects)
? (existing.projects as Record<string, Record<string, unknown>>)
: {};
const existingProject = projects[workingDir] && typeof projects[workingDir] === 'object' ? projects[workingDir] : {};
projects[workingDir] = { ...existingProject, hasTrustDialogAccepted: true };
existing.projects = projects;
try {
writeFileSync(filePath, JSON.stringify(existing, null, 2), { encoding: 'utf8', mode: 0o600 });
chmodSync(filePath, 0o600);
} catch {
// best-effort only — the interactive dialogs return instead of a hard failure here
}
}
/**
* Pre-seeds the "skip the bypass-permissions warning" setting (`customModelInjection.
* skipFirstRunPrompts`, alongside apiKeyTrustFile) into an isolated config directory's
* `settings.json` — a real, already-onboarded profile answers claude's one-time warning
* about running with a bypass-permissions flag once and never sees it again, but every
* custom-model session launches with a fresh, otherwise-empty CLAUDE_CONFIG_DIR that
* carries none of that (confirmed live). A different file from apiKeyTrustFile's
* `.claude.json` — this is claude's own global `settings.json`, not project-keyed —
* so it gets its own merge-not-overwrite read-modify-write.
*/
function seedSkipBypassPermissionsPrompt(configDir: string): void {
const filePath = join(configDir, 'settings.json');
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(readFileSync(filePath, 'utf8')) as Record<string, unknown>;
} catch {
existing = {};
}
existing.skipDangerousModePermissionPrompt = true;
try {
writeFileSync(filePath, JSON.stringify(existing, null, 2), { encoding: 'utf8', mode: 0o600 });
chmodSync(filePath, 0o600);
} catch {
// best-effort only — the interactive warning returns instead of a hard failure here
}
}
/** Best-effort recursive removal of a previously-written configDir. Never throws. */
export function removeConfigDir(dir: string | undefined): void {
if (!dir) return;
@@ -79,14 +242,43 @@ export function applyCustomModelInjection(
entry: Pick<CliEntry, 'capabilities'>,
endpoint: CustomModelEndpoint,
modelId: string,
sessionId: string
sessionId: string,
/** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */
contextLength?: number,
/**
* The session's own working directory — only used for `skipFirstRunPrompts`'s per-project
* trust-dialog seed, and only when provided (boot recovery, which has no reason to
* re-answer a dialog that already fired once, omits it rather than re-deriving it).
*/
workingDir?: string
): AppliedCustomModel | undefined {
const injection = buildCustomModelInjection(entry, endpoint, modelId);
const injection = buildCustomModelInjection(entry, endpoint, modelId, contextLength);
if (injection.kind === 'unsupported') return undefined;
if (injection.kind === 'env') {
// `configDirVar` (claude's CLAUDE_CONFIG_DIR): point it at the same isolated,
// per-session directory the `configDir` kind uses, but write no files into it — an
// empty directory has no stored OAuth credential to conflict with the injected API
// key, which is the whole point. Reusing the same path keyed by sessionId keeps this
// idempotent across a boot-recovery re-apply, same as the configDir kind below.
let envOverrides = injection.envOverrides;
let configDir: string | undefined;
if (injection.configDirVar) {
configDir = customModelConfigDir(sessionId);
mkdirSync(configDir, { recursive: true, mode: 0o700 });
linkSharedProjectsDir(configDir);
if (injection.apiKeyTrustFile && injection.apiKey) {
seedApiKeyTrustFile(configDir, injection.apiKeyTrustFile, injection.apiKey);
}
if (injection.skipFirstRunPrompts && injection.apiKeyTrustFile) {
if (workingDir) seedFirstRunOnboardingState(configDir, injection.apiKeyTrustFile, workingDir);
seedSkipBypassPermissionsPrompt(configDir);
}
envOverrides = { ...envOverrides, [injection.configDirVar]: configDir };
}
return {
envOverrides: injection.envOverrides,
envKeys: Object.keys(injection.envOverrides),
envOverrides,
envKeys: Object.keys(envOverrides),
configDir,
launchModel: injection.launchModel,
};
}
+48 -16
View File
@@ -16,14 +16,21 @@
* shape was rejected by a real codex binary with "invalid type: map,
* expected a string" — caught by `scripts/test-local-llm-harnesses.ts`),
* but `wire_api = "responses"` is the only value codex still accepts
* (support for `"chat"` was dropped in Feb 2026), and a plain OpenAI
* Chat-Completions server (llama.cpp, llama-swap, most local setups) does
* NOT implement the Responses API — so codex may still fail at the
* PROTOCOL level even with a correctly-shaped config file. That gap is
* real and current, not a stale warning; see docs/custom-model-endpoints-plan.md. The rest
* (gemini/pi/grok/deepseek/omp) have their ONE-SHOT INVOCATION flags
* confirmed against real installed binaries' own `--help` output, but
* their custom-endpoint env/config conventions remain web-researched,
* (support for `"chat"` was dropped in Feb 2026). ⚠️ Re-verified live
* against a llama-swap deployment that DOES answer `/v1/responses`: a
* plain, no-tool-call turn gets a real reply, but a real tool-call attempt
* comes back as `agent_message` TEXT (the tool-call JSON printed as the
* answer) rather than a `function_call` item codex would execute —
* confirmed via `codex exec --json`'s raw event stream. Tool execution is
* what makes codex a coding agent, so this remains not usable for real
* work even where plain chat succeeds; see docs/custom-model-endpoints-plan.md
* for the full picture (including the harmless `Model metadata ... not
* found` warning every custom-endpoint codex session prints — sourced from
* a local cache of OpenAI's OWN hosted model catalog that a custom model
* can never appear in, confirmed to have no effect on the outcome above).
* The rest (gemini/pi/grok/deepseek/omp) have their ONE-SHOT INVOCATION
* flags confirmed against real installed binaries' own `--help` output,
* but their custom-endpoint env/config conventions remain web-researched,
* unverified.
*/
@@ -44,6 +51,20 @@ export interface EnvInjection {
envOverrides: Record<string, string>;
/** See {@link ConfigDirInjection.launchModel}. */
launchModel?: string;
/**
* Name of the env var the caller should point at an isolated, credential-free config
* directory for this session (claude's `CLAUDE_CONFIG_DIR`), from the registry entry's
* `customModelInjection.configDirVar`. The actual directory value isn't computed here —
* this module is pure and has no sessionId to derive one from — the IO wrapper
* (`custom-model-injection-apply.ts`) creates it and adds it to `envOverrides`.
*/
configDirVar?: string;
/** See `customModelInjection.apiKeyTrustFile` — carried through so the IO wrapper can seed it. */
apiKeyTrustFile?: { relPath: string; shape: 'claude-api-key-responses' };
/** The literal API key value this injection used, for `apiKeyTrustFile` to pre-approve. */
apiKey?: string;
/** See `customModelInjection.skipFirstRunPrompts` — carried through so the IO wrapper can seed it. */
skipFirstRunPrompts?: boolean;
}
export interface ConfigDirInjection {
@@ -90,7 +111,9 @@ function quoted(value: string): string {
export function buildCustomModelInjection(
entry: Pick<CliEntry, 'capabilities'>,
endpoint: CustomModelEndpoint,
modelId: string
modelId: string,
/** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */
contextLength?: number
): CustomModelInjectionResult {
const cap = entry.capabilities.customModelInjection;
const apiKey = endpoint.apiKey?.trim() || DEFAULT_API_KEY;
@@ -98,11 +121,18 @@ export function buildCustomModelInjection(
switch (cap.kind) {
case 'env': {
const envOverrides: Record<string, string> = {
[cap.baseUrlVar]: endpoint.baseUrl,
[cap.baseUrlVar]: cap.appendV1Suffix ? withV1Suffix(endpoint.baseUrl) : endpoint.baseUrl,
[cap.apiKeyVar]: apiKey,
};
for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId;
return withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId);
if (cap.contextLengthVar && contextLength !== undefined && Number.isFinite(contextLength)) {
envOverrides[cap.contextLengthVar] = String(Math.trunc(contextLength));
}
let result: EnvInjection = withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId);
if (cap.configDirVar) result = { ...result, configDirVar: cap.configDirVar };
if (cap.apiKeyTrustFile) result = { ...result, apiKeyTrustFile: cap.apiKeyTrustFile, apiKey };
if (cap.skipFirstRunPrompts) result = { ...result, skipFirstRunPrompts: true };
return result;
}
case 'configContentEnv': {
@@ -177,11 +207,13 @@ function renderConfigFile(
// `env_key`, the NAME of an env var it reads the credential from at runtime, so the
// actual value must ride along as an extra env var, never embedded in the file.
// ⚠️ `wire_api = "responses"` is the only value codex still accepts (it dropped
// `"chat"` support in Feb 2026) — a plain OpenAI Chat-Completions server (llama.cpp,
// llama-swap, most local setups) does NOT implement the Responses API, so this
// recipe may still fail at the PROTOCOL level even though the file now parses
// correctly. That is a real, currently-unresolved compatibility gap, not a syntax
// bug — track it before calling codex support done.
// `"chat"` support in Feb 2026). Even against a llama-swap deployment that DOES
// answer `/v1/responses`, a real tool-call attempt came back as plain TEXT (the
// tool-call JSON printed as the model's answer) rather than an executable
// `function_call` item — confirmed live via `codex exec --json`. Tool execution is
// what makes codex a coding agent, so this remains not usable for real work even
// where plain chat succeeds — see the confidence table in
// docs/custom-model-endpoints-plan.md, not a syntax bug in this file.
const content = [
`model = ${quoted(modelId)}`,
`model_provider = "custom"`,
+9
View File
@@ -159,6 +159,15 @@ export interface PaneCaptureOptions {
* the 1MB execSync default (ENOBUFS).
*/
maxCaptureBytes?: number;
/**
* Filled in by the implementation with the pane geometry the capture was
* really taken at, which is not always the geometry the caller last asked
* for: a resize and a capture can race, and a pane whose size a desktop
* viewport has claimed ignores a smaller client's resize outright. A
* visible-frame capture addresses every row absolutely, so a consumer
* rendering it needs the real height to know the frame fits.
*/
capturedGeometry?: { cols: number; rows: number };
}
/**
+34
View File
@@ -540,6 +540,32 @@ export function remoteDisplayPath(
return `${remote.username}@${remote.host}:${path}`;
}
/**
* Refresh HOST-level config on a RESTORED `SessionRemote`.
*
* A session's `remote` block is persisted at launch time (mux-sessions.json /
* state.json) and recovery uses that snapshot, so a field ADDED to the host config
* later never reaches an already-running session — not even across a Codeman
* restart. That is exactly how a `wakeCommand` added to `remote-hosts.json` would
* silently do nothing until the session is relaunched (which for an owned remote
* session means killing the remote tmux).
*
* Deliberately narrow: ONLY `wakeCommand`/`wakeMac` are taken from the host config,
* and the host is authoritative for them (removing one in the config turns that
* wake path off again). The other host-level fields (`commands`, ssh options) stay as
* persisted so this cannot silently change how an existing pane connects.
*/
export function rehydrateRemoteHostFields<T extends { hostId: string; wakeCommand?: string; wakeMac?: string }>(
remote: T | undefined,
hostsById: ReadonlyMap<string, RemoteHost>
): T | undefined {
if (!remote) return remote;
const host = hostsById.get(remote.hostId);
if (!host) return remote;
if (remote.wakeCommand === host.wakeCommand && remote.wakeMac === host.wakeMac) return remote;
return { ...remote, wakeCommand: host.wakeCommand, wakeMac: host.wakeMac };
}
export function toSessionRemote(host: RemoteHost, remoteCase: RemoteCase): SessionRemote {
return {
hostId: host.id,
@@ -549,6 +575,10 @@ export function toSessionRemote(host: RemoteHost, remoteCase: RemoteCase): Sessi
port: host.port,
remotePath: remoteCase.remotePath,
commands: host.commands,
// Wake-on-LAN command/MAC travel with the session so the input route can wake a
// sleeping host without a second config read (see remote-wake.ts).
wakeCommand: host.wakeCommand,
wakeMac: host.wakeMac,
// COD-105 — the COD-104 launch path creates the remote session, so we own it
// (an explicit kill may propagate a remote kill-session). Discovered+attached
// sessions go through `toAttachedSessionRemote` with `owned: false`.
@@ -587,6 +617,10 @@ export function toAttachedSessionRemote(
port: host.port,
remotePath,
commands: host.commands,
// An attached session can be woken exactly the same way — the identity of the
// creator does not change whether the host is asleep.
wakeCommand: host.wakeCommand,
wakeMac: host.wakeMac,
// Discovered + attached — another Codeman created it. Detach-not-kill.
owned: false,
remoteSessionName,
+1035
View File
File diff suppressed because it is too large Load Diff
+12 -7
View File
@@ -6,13 +6,18 @@
* stripped before the session is built. The create and resume routes are what
* this bites on: they clamp what a request asked for.
*
* The reboot-restore route calls it as defence in depth, and today it can strip
* nothing. `Session.getEnvOverridesForPersist()` keeps only `CLAUDE_CODE_*` and
* `CLAUDE_CONFIG_DIR` out of a session's overrides, claude's `privilegedEnvKeys`
* are the five `ANTHROPIC_*` names, and that pass admits claude alone — so a
* persisted record cannot carry a clamped key. The call is there for the day the
* persisted set widens. The grant re-resolution that does bite on that path is
* `resolveClaudeModeForUsername`, which recomputes the permission mode.
* The reboot-restore route calls it as defence in depth, and it CAN strip
* something today: `Session.getEnvOverridesForPersist()` keeps only
* `CLAUDE_CODE_*` and `CLAUDE_CONFIG_DIR` out of a session's overrides, and
* claude's `privilegedEnvKeys` now includes both `CLAUDE_CODE_MAX_CONTEXT_TOKENS`
* and `CLAUDE_CONFIG_DIR` (Custom Model Endpoint Profiles, since both can
* redirect a claude session's traffic — see stock.ts's own comment on why they
* are listed despite not needing the clamp for that feature). So a non-granted
* owner's persisted `CLAUDE_CONFIG_DIR` (the per-client-account override, #255)
* is now stripped on reboot-restore, silently returning that session to the
* default Claude account rather than the account it was pointed at. The grant
* re-resolution that ALSO bites on that path is `resolveClaudeModeForUsername`,
* which recomputes the permission mode.
*
* This lives outside `web/routes` on purpose. The question it answers is about
* session privilege rather than about HTTP, and `cron/cron-service.ts` sets the
+133
View File
@@ -0,0 +1,133 @@
/**
* @fileoverview Verify that a programmatically sent prompt actually LEFT the composer,
* and press Enter again while it has not.
*
* Claude Code 2.1.277 (auto-installed 2026-09-18) takes typed text the moment its
* composer paints but ignores Enter for the first 30 to 50 seconds after it, so the
* `send-keys -l <text>` + `send-keys Enter` pair `TmuxManager.sendInput()` sends 50 ms
* apart leaves the prompt sitting on the composer with `0 tokens`, and every caller
* that then waits for the turn (send-and-wait, the agent skill, the maintainer bot,
* cron, Ralph) burns its whole timeout on a turn that never started. Measured through
* the input route on 2026-09-19: an Enter at 28 s stranded, one at 51 s submitted.
*
* The rule: after a write that carried a carriage return, read the pane on a short
* schedule; while the LAST composer line (the CLI's own prompt glyph) still holds the
* head of what was sent, send Enter again. An empty composer ends it, and so does a
* composer holding anything else, because that text is the user's or the CLI's, never
* ours. A pane with no composer line at all (a shell, a CLI whose glyph is not
* declared, a direct-PTY session with no pane to read) does nothing: this runs for
* EVERY programmatic sender, so a blind Enter here could confirm a dialog nobody asked
* about. The composer is the last glyph line on purpose: Claude Code echoes a submitted
* prompt with the same glyph higher up in the transcript, so only the last one says
* whether the text was taken.
*
* Pure apart from the injected capture, send and log, so the schedule, the cap and
* every stop condition are unit-tested with fake timers (test/session-submit-verifier.test.ts).
*/
import { stripAnsi } from './utils/index.js';
/**
* When to look, counted from the write: 2 s catches the common case (taken) with one
* capture, and the tail reaches 60 s, past twice the longest window measured. Enter is
* re-sent at every check that still finds the prompt, so a 50 s window costs about
* seven Enters and one capture each; a taken prompt costs one capture.
*/
export const SUBMIT_VERIFY_DELAYS_MS: readonly number[] = [
2_000, 3_000, 5_000, 5_000, 5_000, 10_000, 10_000, 10_000, 10_000,
];
/** How many leading characters of the prompt have to match, whitespace removed. */
const PROMPT_HEAD_CHARS = 24;
const compact = (s: string): string => s.replace(/\s+/g, '');
/**
* Whether `prompt` is still sitting unsubmitted in the composer of `screen`.
*
* - `true`: the last `glyph` line holds the prompt's head.
* - `false`: the composer is empty (the prompt was taken) or holds other text.
* - `undefined`: no composer line at all; nothing can be said, so nothing is sent.
*
* Whitespace is removed on both sides before comparing, because the composer wraps a
* long prompt onto indented continuation lines and Claude Code draws a no-break space
* after the glyph; `\s` covers that one in JavaScript.
*/
export function promptStillInComposer(screen: string, prompt: string, glyph: string): boolean | undefined {
if (!glyph) return undefined;
const composerLines = stripAnsi(screen)
.split('\n')
.map((l) => l.trim())
.filter((l) => l.startsWith(glyph));
if (composerLines.length === 0) return undefined;
const composer = compact(composerLines[composerLines.length - 1].slice(glyph.length));
if (!composer) return false;
const head = compact(prompt).slice(0, PROMPT_HEAD_CHARS);
return head.length > 0 && composer.startsWith(head);
}
export interface SubmitVerifierDeps {
/** The rendered pane, or null when there is none to read. */
capture: () => string | null | undefined;
/** Press Enter once. Failures are swallowed; the next check decides again. */
sendEnter: () => Promise<unknown> | unknown;
/** The CLI's composer glyph, resolved at check time (the registry can change). */
glyph: () => string;
log?: (message: string) => void;
/** Test seam; production uses SUBMIT_VERIFY_DELAYS_MS. */
delaysMs?: readonly number[];
}
/**
* One per session. `arm(text)` starts the schedule for the prompt just sent and
* cancels any earlier one: a newer write owns the composer now, and re-sending Enter
* for an older prompt could submit the newer one early. `cancel()` is for teardown.
*/
export class SubmitVerifier {
private timer: NodeJS.Timeout | null = null;
private generation = 0;
constructor(private readonly deps: SubmitVerifierDeps) {}
arm(text: string): void {
this.cancel();
const gen = this.generation;
const delays = this.deps.delaysMs ?? SUBMIT_VERIFY_DELAYS_MS;
let step = 0;
let elapsed = 0;
let resent = 0;
const schedule = (): void => {
if (step >= delays.length) return;
const delay = delays[step++];
elapsed += delay;
this.timer = setTimeout(() => void check(), delay);
this.timer.unref?.();
};
const check = async (): Promise<void> => {
this.timer = null;
if (gen !== this.generation) return;
const screen = this.deps.capture();
if (promptStillInComposer(screen ?? '', text, this.deps.glyph()) !== true) return;
resent++;
this.deps.log?.(
`prompt still in the composer after ${Math.round(elapsed / 1000)}s, re-sending Enter (${resent}/${delays.length})`
);
try {
await this.deps.sendEnter();
} catch {
// The next check re-reads the screen and decides again.
}
if (gen !== this.generation) return;
schedule();
};
schedule();
}
cancel(): void {
this.generation++;
if (this.timer) {
clearTimeout(this.timer);
this.timer = null;
}
}
}
+43 -1
View File
@@ -110,6 +110,7 @@ import {
import { DEFAULT_TMUX_HISTORY_LIMIT } from './config/terminal-history.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
import { getCli } from './config/cli-registry/registry.js';
import { SubmitVerifier } from './session-submit-verifier.js';
import { compileVersionRegex } from './config/cli-registry/patterns.js';
import { resolveSessionCliVersion } from './utils/cli-resolver.js';
import {
@@ -502,6 +503,8 @@ export class Session extends EventEmitter {
private _trustDialogAttempts = 0; // Keystrokes sent at the trust dialog
private _lastTrustDialogScanAt = 0; // Throttle for the trust-dialog screen read
private _trustDialogTimer: NodeJS.Timeout | null = null; // Re-read after a keystroke (see below)
/** Re-sends Enter while a programmatic prompt still sits in the composer (session-submit-verifier.ts). */
private _submitVerifier: SubmitVerifier | null = null;
private _interactiveStartedAt = 0; // When the interactive pane launched (bounds that scan)
private _taskTracker: TaskTracker;
@@ -3190,6 +3193,9 @@ export class Session extends EventEmitter {
}
private _clearAllTimers(): void {
// Stop re-sending Enter for a prompt this session will never take now
this._submitVerifier?.cancel();
this._submitVerifier = null;
// Clear the workspace-trust follow-up read
if (this._trustDialogTimer) {
clearTimeout(this._trustDialogTimer);
@@ -3721,7 +3727,10 @@ export class Session extends EventEmitter {
const submittedPrompt = this._trackSubmit(data, options);
if (this._mux && this._muxSession) {
const sent = await this._mux.sendInput(this.id, data);
if (sent) this._emitSubmittedPrompt(submittedPrompt);
if (sent) {
this._emitSubmittedPrompt(submittedPrompt);
this._verifySubmitted(data);
}
return sent;
}
// Fallback to PTY write
@@ -3733,6 +3742,39 @@ export class Session extends EventEmitter {
return false;
}
/**
* Arm the composer check for a write that carried Enter (session-submit-verifier.ts):
* Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints,
* so the pair `sendInput` just sent can leave the text stranded. Only a mux session
* can read its pane, only text can be stranded, and the glyph is the CLI's own.
*/
private _verifySubmitted(data: string): void {
if (!data.includes('\r') || !this._mux?.capturePaneText || !this._muxSession) return;
const text = data.replace(/[\r\n]/g, '').trimEnd();
if (!text) return;
this._submitVerifier ??= new SubmitVerifier({
capture: () =>
this._isStopped || !this._mux || !this._muxSession
? null
: this._mux.capturePaneText?.(this._muxSession.muxName),
sendEnter: () => this._mux?.sendInput(this.id, '\r'),
// ⚠ NO fallback glyph here, unlike the screen-reading probe elsewhere in this file.
// Only claude and codex declare a promptGlyph; the other eight modes would fall back
// to claude's `❯`, which is ALSO starship's default shell prompt (and pure's, and
// spaceship's, and p10k lean's). On a shell session the line `❯ npm run build` sits
// on screen for as long as the command runs, promptStillInComposer() reads that as
// "still unsubmitted", and the verifier presses Enter into the running program's
// stdin on its 2s..60s schedule. Mostly a stray blank line; not harmless against a
// y/N prompt, `read -p`, an installer or a pager, where it takes the default.
// promptStillInComposer() returns undefined for an empty glyph, so this makes the
// verifier inert for every CLI that does not declare one, which is what the Claude
// Code 2.1.277 defect it exists for actually calls for.
glyph: () => getCli(this.mode)?.capabilities.workDetect?.promptGlyph ?? '',
log: (m) => console.log(`[Session ${this.id.slice(0, 8)}] ${m}`),
});
this._submitVerifier.arm(text);
}
/** Current PTY dimensions — used to skip no-op resizes that trigger Ink redraws */
private _ptyCols = 120;
private _ptyRows = 40;
+8
View File
@@ -3482,6 +3482,14 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
{ encoding: 'utf-8', timeout: EXEC_TIMEOUT_MS }
)
);
// Report the size the pane was really drawing at. The visible-frame path
// below addresses every row absolutely, so a consumer whose terminal is
// shorter than this piles the overflow rows onto its last line and loses
// the rows it overwrote. The full-history path instead ends in a RELATIVE
// cursor move, which costs it nothing when the two sizes disagree, so the
// geometry is reported there for diagnosis rather than for repair. Only
// the caller can see both sizes, so hand it this one.
if (opts && geometry) opts.capturedGeometry = { cols: geometry.cols, rows: geometry.rows };
if (fullHistory) {
// Without geometry there is no cursor move, so fall back to the old trim.
+26
View File
@@ -115,6 +115,25 @@ export interface RemoteHost extends RemoteSshOptions {
username: string;
port?: number;
commands?: Partial<Record<RemoteCommandMode, string>>;
/**
* Optional Wake-on-LAN MAC address(es), comma-separated (e.g.
* `04:d9:f5:80:c6:58`). Codeman sends the magic packet itself (UDP port 9
* broadcast), so the common case needs no external script. A SLEEPING host's
* port-22 probe still fails, which is what triggers the wake — this only
* controls HOW the host is woken.
*/
wakeMac?: string;
/**
* Optional Wake-on-LAN command that powers this host on from SLEEP (e.g. a
* wrapper script like `/home/joe/bin/whuff`). TAKES PRECEDENCE over `wakeMac`
* (an explicit override for hosts that need a router/other-host wake). Absent
* = no wake support and today's behavior exactly. Executed WITHOUT a shell (a
* single executable path, never a command line), only from user input or an
* explicit wake request on a session whose host is unreachable — never from
* the auto-reconnect/boot-recovery path, which would re-wake a host seconds
* after each suspend.
*/
wakeCommand?: string;
}
export interface RemoteCase {
@@ -155,6 +174,13 @@ export interface SessionRemote extends RemoteSshOptions {
* session was created elsewhere. Only meaningful when `owned === false`.
*/
remoteSessionName?: string;
/**
* Wake-on-LAN command carried over from the host config (see `RemoteHost.wakeCommand`)
* so the input route can wake a sleeping host without re-reading the host list.
*/
wakeCommand?: string;
/** Wake-on-LAN MAC address(es) from the host config (see `RemoteHost.wakeMac`). */
wakeMac?: string;
}
/**
+321 -29
View File
@@ -219,7 +219,8 @@ const _SSE_HANDLER_MAP = [
// Remote auto-reconnect (COD-108)
[SSE_EVENTS.REMOTE_SESSION_RECONNECTED, '_onRemoteSessionReconnected'],
[SSE_EVENTS.REMOTE_RECONNECT_EXHAUSTED, '_onRemoteReconnectExhausted'],
[SSE_EVENTS.REMOTE_HOST_WAKING, '_onRemoteHostWaking'],
[SSE_EVENTS.REMOTE_HOST_WAKE_FAILED, '_onRemoteHostWakeFailed'],
// Ralph
[SSE_EVENTS.SESSION_RALPH_LOOP_UPDATE, '_onRalphLoopUpdate'],
[SSE_EVENTS.SESSION_RALPH_TODO_UPDATE, '_onRalphTodoUpdate'],
@@ -549,6 +550,12 @@ class CodemanApp {
// repaint-mode CLI pane, where tmux keeps no history of its own). The pull is
// refused for those and retried far more slowly — see _maybeRefetchFullHistory.
this._fullHistoryRepullUseless = new Set();
// Sessions where the geometry replay has already been tried and did NOT
// converge, so the pane is one this browser cannot size. Mirrors the Set
// above: `resizeRetry` caps the recursion inside one select, and this is
// what stops a fresh select from paying for the same answer again — see
// the geometry gate in selectSession.
this._geometryRetryUseless = new Set();
this.terminalLoadStates = new Map(); // Map<sessionId, { generation, phase }>
this.respawnStatus = {};
this.respawnTimers = {}; // Track timed respawn timers
@@ -1713,6 +1720,25 @@ class CodemanApp {
console.error('[SSE] docker container recreated:', err);
}
});
// Custom Model Endpoint Profiles: a session's own model got evicted on llama-swap by
// another session's activity, detected AFTER the fact by a periodic server sweep (there
// is no push notification from llama-swap itself) — see detectCustomModelSwapDisplacements
// in custom-model-routes.ts. Global toast rather than a per-tab indicator: the displaced
// session need not be the one currently open, and the whole point is telling the user
// BEFORE they type into it expecting the model they picked.
addListener(SSE_EVENTS.CUSTOM_MODEL_SWAPPED_OUT, (e) => {
try {
const d = e.data ? JSON.parse(e.data) : {};
this.showToast(
`${d.sessionName || d.sessionId}'s model (${d.previousModel}) was swapped out on llama-swap by another ` +
`session — currently loaded: ${d.currentlyLoadedModel}. Sending a message there will reload it.`,
'warning',
{ duration: 0 }
);
} catch (err) {
console.error('[SSE] custom model swapped out:', err);
}
});
// Multi-user admin: live-refresh whichever admin views (panel/Users tab) are open.
addListener(SSE_EVENTS.ADMIN_USERS_CHANGED, () => {
window.codemanAdmin?.onUsersChanged?.();
@@ -1818,6 +1844,9 @@ class CodemanApp {
_onInit(data) {
_crashDiag.log(`INIT: ${data.sessions?.length || 0} sessions`);
this.handleInit(data);
// Start the remote-host reachability poller even if no session switch follows
// (a page loaded with the remote tab already active) — see host-wake-ui.js.
this._ensureHostWakePoller?.();
}
_onSessionCreated(data) {
@@ -1910,6 +1939,53 @@ class CodemanApp {
this._onSessionClearTerminal(data);
}
/**
* How a buffer load that just fetched `payload` must end.
*
* A tmux pane capture is a point-in-time frame, so nothing that reached the
* browser after the response headers can already be in it. Such a load
* replays exactly that tail; discarding it drops the CLI's output for the
* rest of the load window, and its next partial redraw then lands on a frame
* the terminal never received.
*
* A `history` payload is the server's byte buffer alone: the direct-PTY
* fallback, or a mux pane whose capture came back empty. The route reads
* that buffer in the same synchronous tick it takes the capture, so it is
* current up to the route's own read and no further, which is the same
* exposure. It deliberately keeps the pre-existing discard all the same:
* both cases are rare, neither has been measured, and a duplicated Ink
* redraw is more visible than a few milliseconds of missing output.
* `capturedFromMux` below is the one line to widen if either turns out to
* matter.
*
* `headersReceivedAt` is the caller's own `performance.now()` reading from
* the moment the response arrived, compared only against other client-side
* readings, so there is no clock skew to worry about.
*
* What this cutoff does NOT cover, and there are two contributors. The
* server appends output to the byte buffer and emits it in the same tick,
* but BROADCASTS on a batch timer (8ms over WebSocket, 16 to 50ms over SSE),
* and the terminal route runs synchronously from `capture-pane` to its
* return, so a batch already pending when the capture ran leaves the server
* after the reply, arrives after `headersReceivedAt`, and is replayed
* although the capture holds it. Separately, `captureActivePaneBuffer` is
* `execSync`, which blocks the event loop for the whole capture: anything
* tmux had already painted into the pane that the server had not yet read
* from the attach PTY is in the capture too, is broadcast only after the
* reply, and replays the same way. The duplicate is one batch interval plus
* one capture wide, against a recovery window that spans the whole chunked
* write. Closing it belongs on the server: flush that session's pending
* batch before taking the capture.
*
* @param {{source?: string}} payload - The parsed `data` of a terminal response.
* @param {number} headersReceivedAt - When that response reached this client.
* @returns {{flushQueued: boolean, since: number}} Options for `_finishBufferLoad`.
*/
_bufferLoadFinishOpts(payload, headersReceivedAt) {
const capturedFromMux = payload?.source === 'mux-visible' || payload?.source === 'mux-full-history';
return { flushQueued: capturedFromMux, since: headersReceivedAt };
}
_onSessionTerminal(data) {
if (data.id === this.activeSessionId) {
if (data.data.length > 32768) _crashDiag.log(`TERMINAL: ${(data.data.length/1024).toFixed(0)}KB`);
@@ -1919,7 +1995,7 @@ class CodemanApp {
// jump over the cap. Dropped data is recovered from the canonical buffer.
const queued = (this.pendingWrites?.reduce((s, w) => s + w.length, 0) || 0)
+ (this.flickerFilterBuffer?.length || 0)
+ (this._loadBufferQueue?.reduce((s, w) => s + w.length, 0) || 0)
+ (this._loadBufferQueue?.reduce((s, w) => s + w.data.length, 0) || 0)
+ (this._terminalWriteInFlightBytes || 0);
if (queued + data.data.length > 131072) { // 128KB — drop to prevent accumulation
// Schedule a self-recovery once the
@@ -2502,9 +2578,11 @@ class CodemanApp {
? `/api/sessions/${sessionId}/terminal?full=1`
: `/api/sessions/${sessionId}/terminal?tail=${TERMINAL_TAIL_SIZE}`
);
let headersReceivedAt = performance.now();
let data = (await res.json())?.data ?? {};
if (useFullHistory && data.terminalBuffer && this._replayWouldShrinkBuffer(data.terminalBuffer)) {
res = await fetch(`/api/sessions/${sessionId}/terminal?tail=${TERMINAL_TAIL_SIZE}`);
headersReceivedAt = performance.now();
data = (await res.json())?.data ?? {};
}
// Bail on a tab switch mid-fetch: writing here would paint this session's
@@ -2520,7 +2598,12 @@ class CodemanApp {
const linesFromBottom = before ? Math.max(0, (before.baseY || 0) - (before.viewportY || 0)) : 0;
this.terminal.clear();
this.terminal.reset();
await this.chunkedTerminalWrite(data.terminalBuffer);
await this.chunkedTerminalWrite(
data.terminalBuffer,
TERMINAL_CHUNK_SIZE,
undefined,
this._bufferLoadFinishOpts(data, headersReceivedAt)
);
// A tail fetch can be partial, and the banner would otherwise keep
// describing the pre-refresh buffer (#258).
this._setHistoryTruncation(sessionId, data);
@@ -2530,6 +2613,10 @@ class CodemanApp {
});
if (target === null || typeof this.terminal.scrollToLine !== 'function') this.terminal.scrollToBottom();
else this.terminal.scrollToLine(target);
// The load's own replay sampled the sticky-scroll baseline while the
// terminal sat at the bottom of a just-rewritten buffer, so the next
// flush would scroll back down and undo the restore above.
this._syncStickyScrollBaseline();
// Re-position local echo overlay at new prompt location
this._localEchoOverlay?.rerender();
// Resize PTY to match actual browser dimensions (critical for OpenCode
@@ -2556,6 +2643,7 @@ class CodemanApp {
// Fetch buffer, clear terminal, write buffer, resize (no Ctrl+L needed)
try {
const res = await fetch(`/api/sessions/${data.id}/terminal`);
const headersReceivedAt = performance.now();
const termData = (await res.json())?.data ?? {};
this.terminal.clear();
@@ -2565,7 +2653,12 @@ class CodemanApp {
// (markers don't help here - this is a static buffer reload, not live Ink redraws)
const cleanBuffer = termData.terminalBuffer.replace(DEC_SYNC_STRIP_RE, '');
// Use chunked write to avoid UI freeze with large buffers (can be 1-2MB)
await this.chunkedTerminalWrite(cleanBuffer);
await this.chunkedTerminalWrite(
cleanBuffer,
TERMINAL_CHUNK_SIZE,
undefined,
this._bufferLoadFinishOpts(termData, headersReceivedAt)
);
}
// Fire-and-forget resize — don't block on it
@@ -5719,28 +5812,7 @@ class CodemanApp {
if (ta) ta.dispatchEvent(new CompositionEvent('compositionend', { data: '' }));
}
} catch {}
// Flush local echo text to PTY before switching tabs.
// Send as a single batch (no Enter) so it lands in the session's readline
// input buffer — avoids "old text resent on Enter" and overlay render bugs.
// Track flushed length so _render() offsets the overlay correctly even before
// the PTY echo arrives in the terminal buffer.
if (this.activeSessionId) {
const echoText = this._localEchoOverlay?.pendingText || '';
// Include buffer-detected flushed text (from Tab completion, etc.)
// so it's preserved across tab switches.
const existingFlushed = this._localEchoOverlay?.getFlushed()?.count || 0;
const existingFlushedText = this._localEchoOverlay?.getFlushed()?.text || '';
if (echoText) {
this._sendInputAsync(this.activeSessionId, echoText);
}
const totalOffset = existingFlushed + echoText.length;
if (totalOffset > 0) {
if (!this._flushedOffsets) this._flushedOffsets = new Map();
if (!this._flushedTexts) this._flushedTexts = new Map();
this._flushedOffsets.set(this.activeSessionId, totalOffset);
this._flushedTexts.set(this.activeSessionId, existingFlushedText + echoText);
}
}
this._flushLocalEchoTo(this.activeSessionId);
this._localEchoOverlay?.clear();
// Predictions are ephemeral + already sent: nothing to save/restore
// across a tab switch (unlike the buffer overlay's setFlushed machinery)
@@ -5755,6 +5827,45 @@ class CodemanApp {
}
}
/**
* Hand the local-echo overlay's unsent text to `sessionId` before anything
* clears it, and record what has now been flushed so `_render()` offsets the
* overlay correctly even before the PTY echo comes back.
*
* On a touch device the characters the user has typed live ONLY here until
* Enter — they have never reached the PTY — so whoever clears the overlay
* owes them a flush first. It is sent as one batch with no Enter, so it lands
* in the session's readline buffer rather than submitting a line the user has
* not finished.
*
* ⚠️ The session is a PARAMETER because the two callers are looking at
* different ones. `_cleanupPreviousSession` flushes to the tab being left,
* which is still `activeSessionId` when it runs. The `forceReload` branch in
* `selectSession` flushes to the tab being RELOADED, and must do it before it
* nulls `activeSessionId`: reading the field after that null is what silently
* dropped the text, since the guard here then saw no session and the
* unconditional `clear()` that follows took the characters with it.
* @param {string|null} sessionId
*/
_flushLocalEchoTo(sessionId) {
if (!sessionId) return;
const echoText = this._localEchoOverlay?.pendingText || '';
// Include buffer-detected flushed text (from Tab completion, etc.)
// so it's preserved across tab switches.
const existingFlushed = this._localEchoOverlay?.getFlushed()?.count || 0;
const existingFlushedText = this._localEchoOverlay?.getFlushed()?.text || '';
if (echoText) {
this._sendInputAsync(sessionId, echoText);
}
const totalOffset = existingFlushed + echoText.length;
if (totalOffset > 0) {
if (!this._flushedOffsets) this._flushedOffsets = new Map();
if (!this._flushedTexts) this._flushedTexts = new Map();
this._flushedOffsets.set(sessionId, totalOffset);
this._flushedTexts.set(sessionId, existingFlushedText + echoText);
}
}
_resetTerminalForReplay() {
this.terminal.reset();
this.terminal.write('\x1b[3J\x1b[H\x1b[2J');
@@ -5862,7 +5973,12 @@ class CodemanApp {
parsedAt,
bufferLength: parsedBufferLength,
completed,
} = await this.chunkedTerminalWrite(buffer, TERMINAL_CHUNK_SIZE, sessionId);
} = await this.chunkedTerminalWrite(
buffer,
TERMINAL_CHUNK_SIZE,
sessionId,
this._bufferLoadFinishOpts(payload, headersReceivedAt)
);
timing.resetAndParseMs = parsedAt - replayStartedAt;
if (!completed || this.activeSessionId !== sessionId) return;
// Keep shell tab restores bounded too. A user-triggered full-history pull
@@ -5880,6 +5996,12 @@ class CodemanApp {
const delta = parsedBufferLength - rowsBefore;
if (delta > 0) this.terminal.scrollToLine(delta);
else this.terminal.scrollToTop();
// The load's own replay sampled the sticky-scroll baseline while the
// terminal sat at the bottom of a just-rewritten buffer, so the next
// flush would scroll back down and undo the restore above. This path is
// reached only from a scroll-up gesture, so being dragged down is the
// exact opposite of what the user asked for.
this._syncStickyScrollBaseline();
timing.totalMs = performance.now() - requestStartedAt;
this._recordTerminalLoadTiming(timing);
} catch {
@@ -6018,6 +6140,13 @@ class CodemanApp {
this._loadBufferQueue = null;
this._terminalRefreshOwner = null;
this._chunkedWriteGen = (this._chunkedWriteGen || 0) + 1;
// Anything typed but not yet submitted lives in the local-echo overlay and
// has never reached the PTY. `_cleanupPreviousSession` below flushes it,
// but only for a session it can still see, and the null on the next line
// hides this one from it. Flush first or the characters are cleared
// unread. The geometry replay re-enters here with no gesture behind it,
// so on a touch device this fires while the user is still typing.
this._flushLocalEchoTo(sessionId);
this.activeSessionId = null;
}
// Focus terminal SYNCHRONOUSLY before any await — iOS Safari only honors
@@ -6084,6 +6213,9 @@ class CodemanApp {
// bar (issue #262). Also disarms a one-shot Ctrl left over from the tab we
// just left, so it can never fire against the session we just opened.
if (typeof KeyboardAccessoryBar !== 'undefined') KeyboardAccessoryBar.refreshForActiveSession();
// Remote-host reachability banner: only meaningful for a remote session, so this
// also clears it when the newly active tab is local.
this.refreshHostWakeBanner?.(sessionId);
// Restore flushed offset AND text IMMEDIATELY so backspace/typing work during
// the async buffer load. Without this, the offset is 0 during the
@@ -6197,6 +6329,10 @@ class CodemanApp {
// sendResize is a no-op on the server when dims haven't changed, so
// calling it every tab switch is cheap.
const dimsChanged = await this.sendResize(sessionId, { forceHttp: true }).catch(() => false);
// The size the capture below will be taken against. The debounced resize
// handler can move the terminal again while the load runs, so this is a
// recorded value rather than a later read of `_lastResizeDims`.
const dimsAtCapture = this.getTerminalDimensions?.();
if (this._isStaleSelect(selectGen)) {
this._clearTerminalLoadState(sessionId, selectGen);
return;
@@ -6321,6 +6457,15 @@ class CodemanApp {
}
const data = (await res.json())?.data ?? {};
const bodyParsedAt = performance.now();
// How this load must end, decided here because `chunkedTerminalWrite` is
// what actually ends it for a non-empty buffer. A tmux pane capture is a
// point-in-time frame, so nothing that reached the browser after the
// response headers can already be in it. Replay exactly that tail;
// discarding it drops the CLI's output for the rest of the load window,
// and its next partial redraw then lands on a frame the terminal never
// received. `since` keeps the pre-capture events dropped, because the
// capture does hold those and replaying them would duplicate output.
const finishOpts = this._bufferLoadFinishOpts(data, headersReceivedAt);
_crashDiag.log(`FETCH_DONE: ${data.terminalBuffer ? (data.terminalBuffer.length/1024).toFixed(0) + 'KB' : 'empty'} truncated=${data.truncated}`);
let freshResetAndParseMs = 0;
@@ -6347,7 +6492,8 @@ class CodemanApp {
const { parsedAt: freshParsedAt } = await this.chunkedTerminalWrite(
data.terminalBuffer,
TERMINAL_CHUNK_SIZE,
bufferLoadOwner
bufferLoadOwner,
finishOpts
);
freshResetAndParseMs = freshParsedAt - replayStartedAt;
if (this._isStaleSelect(selectGen)) {
@@ -6397,7 +6543,14 @@ class CodemanApp {
// COD-144: when the load painted nothing, FLUSH the queued events instead of
// discarding — a new session's prompt arrives only as a queued SSE event.
if (this._isLoadingBuffer) {
this._finishBufferLoad(bufferLoadOwner, { flushQueued: bufferWasEmpty });
// Only reached when the write was skipped. COD-144 lives here: a new
// session's first prompt exists only as a queued event that predates the
// response, so an empty paint replays its queue WHOLE rather than from
// the header timestamp.
this._finishBufferLoad(
bufferLoadOwner,
bufferWasEmpty ? { flushQueued: true, since: 0 } : finishOpts
);
}
// Drop the guard so user input clears state normally
this._restoringFlushedState = false;
@@ -6428,6 +6581,84 @@ class CodemanApp {
// annoyance that disappear on the user's next keypress; data loss is not
// acceptable. Do NOT re-introduce Ctrl+L here.
this.sendResize(sessionId);
// sendResize fits synchronously before its first await, so this reads the
// size that survived the load rather than the one the capture was taken
// at. The two differ whenever the terminal was still settling.
const dimsAfterLoad = this.getTerminalDimensions?.();
// Only a visible-frame capture positions its rows absolutely, and only
// that frame can be damaged by a terminal of the wrong size. A `full=1`
// body is linear scrollback closed by a RELATIVE cursor move
// (`formatCursorRestore`), which is relative precisely so the browser's
// row count need not match the pane's, and a `history` body is the byte
// stream, which carries no row alignment to protect. Replaying either at
// a different size repairs nothing, and the full-history replay costs a
// second whole-scrollback capture to learn that. Since the first select
// of every non-shell session per page takes the full-history path, an
// ungated comparison fires most often on the one response it cannot help.
const framePositionsRowsAbsolutely = data.source === 'mux-visible';
// `mux-visible` is necessary but not sufficient: when the `display-message`
// cursor query fails, `capturePaneBuffer` skips the snapshot repaint and
// returns the raw capture, and the route still labels a non-empty body
// `mux-visible`. That body positions nothing and reports no geometry, so a
// size that moved during such a load has nothing to repair, and replaying
// would buy a second capture, a reset plus chunked rewrite, a dropped
// WebSocket and a discarded xterm snapshot for it. The two comparisons
// below already stand down on an absent field; this one has to as well.
const sizeMovedUnderLoad =
framePositionsRowsAbsolutely &&
Number.isFinite(data.captureRows) &&
!!dimsAtCapture &&
!!dimsAfterLoad &&
(dimsAfterLoad.cols !== dimsAtCapture.cols || dimsAfterLoad.rows !== dimsAtCapture.rows);
// A capture positions every row absolutely, so a pane taller than this
// terminal writes its overflow rows onto the last line and loses the rows
// it overwrote. A pane WIDER than this terminal damages the same frame a
// second way: `formatPaneSnapshot` paints each row out to the pane's own
// width, so a narrower browser wraps every painted row, and the wrap on
// the last one scrolls the whole frame up by a row. Both happen when the
// capture wins a race against the resize meant to precede it, which is
// what the retry below repairs.
//
// It also happens when `Session.resize` DECLINED the resize, which it does
// for a small viewport while a desktop viewport's size claim is live. The
// retry cannot repair that one: it re-sends the same declined resize and
// captures the same too-tall pane. `resizeRetry` stops it after the one
// extra attempt, and the frame is shown as-is. Repairing that case means
// changing who owns the pane size, which is a policy question this does
// not touch. What the flag does buy there is that the client can SEE the
// mismatch at all, which it previously could not.
//
// An ABSENT field is not a fit. It means the capture reported no geometry
// at all, so nothing was positioned and there is nothing to repair.
const capturedTallerThanTerminal =
framePositionsRowsAbsolutely &&
Number.isFinite(data.captureRows) &&
data.captureRows > (this.terminal?.rows || 0);
const capturedWiderThanTerminal =
framePositionsRowsAbsolutely &&
Number.isFinite(data.captureCols) &&
data.captureCols > (this.terminal?.cols || 0);
// The retry replays at `dimsAfterLoad`, so it can only change what is on
// screen if the pane was drawing at some OTHER size. When the reported
// geometry already IS that size, the second pass captures the identical
// frame and pays a full reload to do it: another fetch, another
// `_resetTerminalForReplay()` and chunked rewrite (a visible re-flash),
// and, because it goes through `forceReload`, a dropped and reopened
// WebSocket plus a deleted xterm snapshot.
//
// That equality is the signature of a CLAMP rather than a race.
// `getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` does
// not, so a terminal narrower than 40 columns or shorter than 10 rows
// reports a pane permanently bigger than itself, and every select would
// retry without ever converging. A race never produces this equality: its
// whole premise is that the pane was still at the size we asked it to
// leave. The other non-converging case, `Session.resize` declining a
// small viewport while a desktop claim is live, does not produce it
// either — that pane sits at the DESKTOP's size — so it still costs the
// one capped attempt, and stopping it needs the pane-ownership policy
// this does not touch.
const captureMatchesRequestedSize =
!!dimsAfterLoad && data.captureCols === dimsAfterLoad.cols && data.captureRows === dimsAfterLoad.rows;
// Defer secondary panel updates so they don't block the main thread
// after terminal content is already visible.
@@ -6528,6 +6759,67 @@ class CodemanApp {
this._clearTerminalLoadState(sessionId, selectGen);
_crashDiag.log(`SELECT_DONE: ${selectDoneMs.toFixed(0)}ms`);
console.log(`[CRASH-DIAG] selectSession DONE: ${sessionId.slice(0,8)} in ${selectDoneMs.toFixed(0)}ms`);
// Remember whether the replay was worth it, because `resizeRetry` only
// caps the recursion INSIDE one select and says nothing about the next
// one. A pane this browser cannot size — one whose resize `Session.resize`
// declines while a desktop claim is live, or one a second tmux client is
// also holding — reports the same mismatch on every select, so without a
// memo the diagnosis is paid for again on every tab switch, forever: two
// fetches per select rather than one. Each extra pass costs a second
// `capture-pane`, which is `execSync` and blocks the server's event loop,
// plus a reset and chunked rewrite, a discarded snapshot and cache entry,
// and a dropped and reopened WebSocket.
//
// A retry pass that STILL does not fit is the proof, since the retry ran
// at the size that stuck and the pane ignored it. Geometry that fits
// clears the memo, so a pane that becomes sizeable again (the desktop tab
// closes, the claim goes idle) is repaired on the next select. The race
// case is untouched: it converges on its first attempt, so it never
// reaches the branch that latches.
const capturedGeometryFits =
framePositionsRowsAbsolutely &&
Number.isFinite(data.captureRows) &&
!capturedTallerThanTerminal &&
!capturedWiderThanTerminal;
if (capturedGeometryFits) {
this._geometryRetryUseless?.delete(sessionId);
} else if (options?.resizeRetry && (capturedTallerThanTerminal || capturedWiderThanTerminal)) {
(this._geometryRetryUseless ||= new Set()).add(sessionId);
}
// What is on screen was drawn for a geometry this terminal does not have.
// Replaying once against the size that stuck is the only thing that
// repairs it: SIGWINCH reaches the CLI only on a real size change, and
// the pane is already at its final size, so no redraw is coming.
// `resizeRetry` caps this at one attempt, so two competing fits cannot
// trade replays forever.
if (
(sizeMovedUnderLoad || capturedTallerThanTerminal || capturedWiderThanTerminal) &&
!captureMatchesRequestedSize &&
!this._geometryRetryUseless?.has(sessionId) &&
!options?.resizeRetry &&
!this._isStaleSelect(selectGen)
) {
_crashDiag.log(
`RESIZE_RETRY: capture ${data.captureCols}x${data.captureRows} vs terminal ` +
`${this.terminal?.cols}x${this.terminal?.rows}` +
(sizeMovedUnderLoad ? ' (size moved under load)' : '')
);
// Re-arm the full-history pull ONLY if this pass actually used one, so
// the retry replays the same content at the geometry that stuck. A pass
// that took the bounded tail must retry on the tail too: clearing the
// flag unconditionally would UPGRADE a tab switch into a fresh
// multi-megabyte scrollback capture it never asked for.
//
// UNREACHABLE as written, and kept for the invariant rather than the
// branch. A `useFullHistory` pass sends `full=1`, and the route answers
// `full=1` with `mux-full-history` or `history`, never `mux-visible`
// (see the source ladder in session-routes.ts), so the gate above
// already rules out every pass that consumed the flag. Do not read this
// line as evidence that a page load retries: it does not, and the test
// suite pins that it does not.
if (useFullHistory) this._fullHistoryLoaded.delete(sessionId);
await this.selectSession(sessionId, { auto: true, forceReload: true, resizeRetry: true });
}
} catch (err) {
if (this._isLoadingBuffer) this._finishBufferLoad(bufferLoadOwner);
this._restoringFlushedState = false;
+74
View File
@@ -806,6 +806,71 @@ function decideAutoCopy({ enabled, text, lastCopied, pending } = {}) {
return 'copy';
}
// The text a copy should put on the clipboard, given xterm's raw selection.
// Pure: the caller reads the selection and decides the mode, this transforms.
//
// xterm hands back whole screen ROWS, and its own trim only drops cells that
// were never written to. A full-screen TUI writes real spaces across the part
// of a row it is not using, so that padding counts as content and rides along
// to the clipboard: measured against Claude Code in a 282-column pane, single
// lines arrived carrying 138 trailing spaces. Native terminals trim it on copy
// (Windows Terminal, iTerm2 and GNOME Terminal all do), decideAutoCopy above
// already calls a wall of spaces "never what the gesture meant", and
// _selectTouchSelectionLine already treats those cells as padding. This is that
// same rule for the mouse and keyboard paths, which never had it.
//
// ⚠ Trailing padding ONLY. A shared LEADING indent is deliberately left alone,
// and this note is here so the idea is not re-derived: it was built, measured
// and dropped before merge. Removing the longest leading run every selected row
// shares looks like the mirror image of the trailing trim and is not, because
// no native terminal does it and the transform cannot tell a TUI's margin from
// content that is genuinely indented. Measured over 401 445 three-row windows
// across 1 010 tracked files in this repo, it fired on 73% of them: 92% inside
// a YAML workflow, 76% over `git log` output, 48% in a TypeScript source file.
// No width threshold separates the two, because they are the same widths: a
// live Claude Code pane's own margins measure 2 and 5 columns while the most
// common non-TUI shared run is 4, sitting between them.
//
// The asymmetry that settles it is in the failure modes. A wrong trailing trim
// costs nothing. A wrong dedent silently deletes information that was on the
// screen, with no signal to the user and nothing in the clipboard to hint at
// it, and it is wrong on `git log` bodies, on indented code read out of `cat`
// (semantic in Python), on `git diff` context rows where the leading space is
// the marker, and on stack traces.
//
// ⚠ It also cannot be made consistent cheaply. Whether the first row joins the
// measurement depended on the mousedown COLUMN, which the user never sees, so
// one block of three rows produced three different clipboard results; and the
// flag read `getSelectionPosition().start`, which is the mousedown anchor that
// xterm never normalises, so dragging UP through a block read it off the bottom
// row. If it is ever revisited, the one qualification that measured clean is
// painted trailing padding (a full-screen TUI writes real spaces across every
// row; a shell pane leaves those cells never-written, so xterm trims them):
// zero false positives over all 401 445 windows. It still mangles a `git log`
// body sitting inside an agent's own gutter, which is why it was not taken now.
function cleanCopiedSelection(text) {
if (typeof text !== 'string' || !text) return '';
// Split on \n and leave any \r in place: xterm joins rows with \r\n on
// Windows, and the clipboard should keep the endings xterm chose.
// Scanned rather than matched. A selection can run to the 50 000-row
// scrollback ceiling, and `/[ \t]+(\r?)$/` is QUADRATIC on a line whose spaces
// are followed by any non-space character, which is what right-aligned or
// centred TUI content looks like: the engine retries the run from every
// whitespace position and backtracks over it. Measured over 50 000 rows with a
// 280-column run, that regex took 2.9s against 1.3ms for the scan below, and a
// 2 000-column run took 16s. It is also the faster of the two on an ordinary
// padded row. A length is returned rather than a trimmed string so a
// \r-terminated line costs no substring either.
const trimEnd = (line) => {
let end = line.length;
if (end > 0 && line[end - 1] === '\r') end--;
let cut = end;
while (cut > 0 && (line[cut - 1] === ' ' || line[cut - 1] === '\t')) cut--;
return cut === end ? line : line.slice(0, cut) + line.slice(end);
};
return text.split('\n').map(trimEnd).join('\n');
}
if (typeof window !== 'undefined') {
window.WEBGL_FALLBACK = WEBGL_FALLBACK;
window.evaluateWebGLLongTaskTrip = evaluateWebGLLongTaskTrip;
@@ -854,6 +919,9 @@ if (typeof window !== 'undefined') {
decide: decideAutoCopy,
MAX_CHARS: AUTO_COPY_MAX_CHARS,
};
window.CodemanCopySelection = {
clean: cleanCopiedSelection,
};
window.CodemanTerminalFont = {
DEFAULT_STACK: TERMINAL_FONT_DEFAULT_STACK,
resolve: resolveTerminalFontFamily,
@@ -1057,6 +1125,9 @@ const SSE_EVENTS = {
REMOTE_SESSION_DROPPED: 'remote:sessionDropped',
REMOTE_SESSION_RECONNECTED: 'remote:sessionReconnected',
REMOTE_RECONNECT_EXHAUSTED: 'remote:reconnectExhausted',
// Wake-on-LAN from user input on a sleeping remote host
REMOTE_HOST_WAKING: 'remote:hostWaking',
REMOTE_HOST_WAKE_FAILED: 'remote:hostWakeFailed',
// Ralph
SESSION_RALPH_LOOP_UPDATE: 'session:ralphLoopUpdate',
@@ -1094,6 +1165,9 @@ const SSE_EVENTS = {
APPROVAL_UPDATED: 'approval:updated',
APPROVAL_RESOLVED: 'approval:resolved',
// Custom Model Endpoint Profiles
CUSTOM_MODEL_SWAPPED_OUT: 'custom-model:swapped-out',
// Subagents (Claude Code background agents)
SUBAGENT_DISCOVERED: 'subagent:discovered',
SUBAGENT_UPDATED: 'subagent:updated',
+440
View File
@@ -0,0 +1,440 @@
/**
* @fileoverview Remote-host wake-on-LAN: the "host unreachable" banner + its config dialog.
*
* A sleeping remote host does not fail loudly. The local tmux pane runs `ssh`, and when
* the machine suspends, that ssh child stalls: `tmux send-keys` still SUCCEEDS, so typed
* input disappears with no error and the pane looks alive. The server side
* (`src/remote-wake.ts`) buffers input and wakes the host when the user types; this
* module makes the state VISIBLE and gives it a button, which is what turns "why is
* nothing happening" into one click.
*
* Behavior:
* - Asks `GET /api/sessions/:id/reachability` for the ACTIVE remote session only:
* once when the tab is activated (a user action), and every `POLL_MS` while the tab
* is visible ONLY for a host with a wake target. The timer is the one thing here that
* is not user-driven, and each poll is a TCP connect to the host — the same
* timer-driven traffic invariant #2 rejects keepalives for: it cannot wake a host,
* but it can keep an activity-based suspend timer from firing. So a host Codeman
* could not wake anyway is never polled on a timer. A host behind a jump host or
* SOCKS proxy (`probeable: false`) is never polled at all: the probe cannot reach
* it, so its answer would only ever be a false "asleep". The endpoint shares the
* server's probe cache with the input path, so opening the tab also primes the
* wake path.
* - Unreachable + a configured wake target → "Wake" button → `POST /api/sessions/:id/wake`
* (which wakes, waits, reattaches the pane and flushes buffered input).
* - Unreachable + NO wake target → "Configure WoL" → `#wakeConfigModal`, a small form
* for this host's MAC/command that saves via `PUT /api/remote-hosts/:id`. The server
* re-resolves host config while the session is live, so saving takes effect without
* restarting the session.
* - SSE (`remote:hostWaking`, `remote:hostWakeFailed`, `remote:sessionReconnected`)
* keeps the banner in sync while a wake is running.
*
* @mixin Extends CodemanApp.prototype via Object.assign
* @dependency app.js (CodemanApp class, this.sessions, this.activeSessionId, showToast)
* @dependency constants.js (SSE_EVENTS — the remote:hostWaking / remote:hostWakeFailed names)
* @loadorder 12.2 — loaded after session-ui.js, before webview-tabs.js
*/
const HOST_WAKE_POLL_MS = 30_000;
Object.assign(CodemanApp.prototype, {
/** Per-tab banner state (single active session at a time). */
_hostWake: null,
/** The page-wide poller interval (created once, see `_ensureHostWakePoller`). */
_hostWakeTimer: null,
/** Fresh state for a session we just switched to. */
_hostWakeState() {
return {
sessionId: null,
/** Last reachability answer, or null before the first poll. */
reachable: null,
/** 'command' | 'mac' | 'none' — what the banner action should do. */
wakeConfigured: 'none',
host: '',
label: '',
/**
* False for a host the server's probe cannot reach (behind a jump host or SOCKS
* proxy): its reachability is unknown, so there is no banner and no polling.
*/
probeable: true,
/** True between clicking Wake and the answer coming back. */
waking: false,
/**
* True only when the server is actually holding bytes for this session (the typing
* path buffers them). Browser keystrokes go over the WebSocket, which never passes
* through the wake registry — so the Wake BUTTON must not claim input is queued.
*/
queuedInput: false,
/** Set when the last wake attempt or poll failed. */
error: '',
};
},
/**
* Entry point from the session switcher — called for every active session, remote or
* not, so it must be cheap and must clear the banner for local sessions.
*
* ⚠️ The POLLER is page-wide and independent of this call on purpose: a session
* switch is not the only way the active tab changes (boot restore, a page loaded with
* the tab already active, and `selectSession`'s own early return for the tab you are
* already on), and the banner must not depend on any single one of those paths
* running — that is exactly how it could silently never appear.
*/
refreshHostWakeBanner(sessionId) {
this._ensureHostWakePoller();
const state = this._hostWake;
if (state && state.sessionId && state.sessionId !== sessionId) this._hostWake = null;
this._hostWakeTick();
},
/** Create the page-wide poller once (interval + a visibility wake-up). */
_ensureHostWakePoller() {
if (this._hostWakeTimer) return;
this._hostWakeTimer = setInterval(() => this._hostWakeTick({ periodic: true }), HOST_WAKE_POLL_MS);
document.addEventListener('visibilitychange', () => {
if (document.visibilityState === 'visible') this._hostWakeTick({ periodic: true });
});
},
/**
* One poller tick: resolve the ACTIVE session, reset the banner when it changed, and
* ask the server. No-op while the page is hidden (a background tab must not poll).
*
* `periodic` marks the timer (and the visibility wake-up) as opposed to a tab
* activation: a periodic tick polls only a host with a wake target, see the module
* comment. The activation poll is what still offers "Configure WoL" for a sleeping
* host that has none — one connect, on a user action.
*/
_hostWakeTick({ periodic = false } = {}) {
if (typeof document !== 'undefined' && document.visibilityState === 'hidden') return;
const sessionId = this.activeSessionId;
const session = sessionId && this.sessions ? this.sessions.get(sessionId) : null;
if (!sessionId || !session || !session.remote) {
// Render unconditionally: `refreshHostWakeBanner` clears `_hostWake` BEFORE
// calling this tick, so a guard here would skip the repaint and leave the
// banner up on every chat (the clear and the repaint must not be coupled to
// whoever cleared the state). Idempotent — with a null state it just hides.
this._hostWake = null;
this._renderHostWakeBanner();
return;
}
let state = this._hostWake;
let fresh = false;
if (!state || state.sessionId !== sessionId) {
fresh = true;
state = this._hostWake = this._hostWakeState();
state.sessionId = sessionId;
state.host = session.remote.host || '';
state.label = session.remote.label || 'Remote host';
// Text from the session payload first (instant, no round trip), corrected by the
// poll — a session whose wake config was added after launch only knows it after
// the server resolves host config. The kind matters: the payload can say WHICH
// path is configured, so a command-only host is not mislabelled 'mac' until the
// first poll lands.
state.wakeConfigured = session.remote.wakeMac ? 'mac' : session.remote.wakeCommand ? 'command' : 'none';
// Known from the payload already: a proxied host is not probeable (the server
// says so too, on every answer), so not even the activation poll is worth a
// round trip whose verdict could only be a wrong "asleep".
state.probeable = !(session.remote.jumpHost || session.remote.socksProxy);
this._renderHostWakeBanner();
}
if (!state.probeable) return;
if (periodic && !fresh && state.wakeConfigured === 'none') return;
this._pollHostReachability();
},
/** One reachability check for the active remote session. */
async _pollHostReachability(force = false) {
const state = this._hostWake;
if (!state || !state.sessionId) return;
const sessionId = state.sessionId;
try {
const res = await fetch(`/api/sessions/${encodeURIComponent(sessionId)}/reachability${force ? '?force=1' : ''}`);
const data = await res.json();
if (!data.success) return;
// The tab may have changed while this was in flight.
if (this._hostWake !== state || state.sessionId !== sessionId) return;
// `reachable` is `null` (unknown, not unreachable) for a host the probe cannot
// reach — only a PROVEN `false` may raise the banner.
state.reachable = data.data.reachable !== false;
if (data.data.probeable === false) state.probeable = false;
state.wakeConfigured = data.data.wakeConfigured || 'none';
if (data.data.host) state.host = data.data.host;
if (data.data.label) state.label = data.data.label;
if (state.reachable) {
state.waking = false;
state.error = '';
}
this._renderHostWakeBanner();
} catch {
/* A failed poll is not a state change: leave the banner as it was. */
}
},
/** Draw the banner from `_hostWake`. */
_renderHostWakeBanner() {
const state = this._hostWake;
const banner = this.$('hostWakeBanner');
const text = this.$('hostWakeBannerText');
const detail = this.$('hostWakeBannerDetail');
const action = this.$('hostWakeBannerAction');
if (!banner || !text || !action) return;
const visible = Boolean(state && state.sessionId && state.reachable === false);
banner.hidden = !visible;
if (!visible) return;
const hasTarget = state.wakeConfigured !== 'none';
const target = state.label || state.host || 'Remote host';
if (state.waking) {
text.textContent = `Waking ${target} …`;
} else if (state.error) {
text.textContent = `${target} did not wake up`;
} else {
text.textContent = `${target} is not reachable`;
}
if (detail) {
detail.textContent = state.waking
? state.queuedInput
? 'input is queued until it is back'
: 'waiting for the host to come back'
: hasTarget
? `ssh ${state.host}`
: 'no wake-on-LAN configured';
}
// After a FAILED wake the only useful next step is fixing the target (wrong MAC,
// host moved NIC, command gone) — otherwise a configured-but-broken host would be
// stuck behind a button that keeps failing with no way to edit it.
const offerConfig = !hasTarget || Boolean(state.error);
action.textContent = state.waking ? 'Waking …' : offerConfig ? 'Configure WoL' : 'Wake';
action.disabled = state.waking;
},
/** Banner button: wake the host, or open the setup dialog when nothing is configured. */
hostWakeAction() {
const state = this._hostWake;
if (!state || !state.sessionId || state.waking) return;
if (state.wakeConfigured === 'none' || state.error) {
this.openWakeConfigDialog();
return;
}
this.wakeRemoteHost();
},
/** POST the manual wake for the active session and follow the result. */
async wakeRemoteHost() {
const state = this._hostWake;
if (!state || !state.sessionId) return;
const sessionId = state.sessionId;
state.waking = true;
// The button path holds nothing: whatever the user typed went into the stalled pane
// over the WebSocket and is gone. Saying otherwise is a promise the next keystroke
// disproves.
state.queuedInput = false;
state.error = '';
this._renderHostWakeBanner();
try {
const res = await fetch(`/api/sessions/${encodeURIComponent(sessionId)}/wake`, { method: 'POST' });
const data = await res.json();
if (this._hostWake !== state || state.sessionId !== sessionId) return;
state.waking = false;
if (!data.success) {
// The ROUTE is the authority on whether a target is configured, so ask it again
// (`/reachability` reports `wakeConfigured`) rather than pattern-matching the
// error message: the message is prose, and the code is generic (`INVALID_INPUT`
// covers "Not a remote session" too).
state.error = data.error || 'Wake failed';
this._renderHostWakeBanner();
await this._pollHostReachability(true);
return;
}
state.reachable = data.data.reachable !== false;
state.wakeConfigured = data.data.wakeConfigured || state.wakeConfigured;
if (state.reachable) {
this.showToast(`${state.label || 'Remote host'} is awake`, 'success');
} else {
state.error = 'timeout';
}
this._renderHostWakeBanner();
} catch (err) {
if (this._hostWake !== state) return;
state.waking = false;
state.error = err && err.message ? err.message : 'Wake failed';
this._renderHostWakeBanner();
}
},
/**
* Why the host could not be read. In multi-user mode `GET /api/remote-hosts` returns
* `[]` to a non-admin, so "Remote host not found" would blame a config the user simply
* is not allowed to see — the save is admin-only, and that is what it should say.
*/
_wakeConfigUnavailableMessage() {
const me = window.__codemanUser || {};
return me.multiUser && me.role !== 'admin' ? 'Wake-on-LAN configuration is admin-only' : 'Remote host not found';
},
/** Open the small WoL dialog for the banner's host, pre-filled from the host config. */
async openWakeConfigDialog() {
const state = this._hostWake;
const session = state && state.sessionId && this.sessions ? this.sessions.get(state.sessionId) : null;
if (!session || !session.remote) return;
const hostId = session.remote.hostId;
const label = this.$('wakeConfigHostLabel');
const mac = this.$('wakeConfigMac');
const command = this.$('wakeConfigCommand');
const status = this.$('wakeConfigStatus');
if (!mac || !command) return;
mac.value = session.remote.wakeMac || '';
command.value = session.remote.wakeCommand || '';
if (label) label.textContent = session.remote.label || hostId;
if (status) status.textContent = '';
this._wakeConfigHostId = hostId;
const modal = this.$('wakeConfigModal');
if (modal) modal.classList.add('active');
// Read the saved host so the dialog shows what is actually persisted (the session
// payload may predate a change made in another tab).
try {
const res = await fetch('/api/remote-hosts');
const data = await res.json();
const hosts = data.success ? data.data : [];
const host = Array.isArray(hosts) ? hosts.find((item) => item.id === hostId) : null;
if (host && this._wakeConfigHostId === hostId) {
mac.value = host.wakeMac || '';
command.value = host.wakeCommand || '';
} else if (!host && this._wakeConfigHostId === hostId && status) {
// Say it up front rather than only when Save fails.
status.textContent = this._wakeConfigUnavailableMessage();
}
} catch {
/* The form is already usable from the session payload. */
}
},
closeWakeConfigDialog() {
const modal = this.$('wakeConfigModal');
if (modal) modal.classList.remove('active');
this._wakeConfigHostId = null;
},
/** Save MAC/command for the host, then re-check whether the session can wake now. */
async saveWakeConfig() {
const hostId = this._wakeConfigHostId;
const mac = this.$('wakeConfigMac');
const command = this.$('wakeConfigCommand');
const status = this.$('wakeConfigStatus');
const save = this.$('wakeConfigSave');
if (!hostId || !mac || !command) return;
const macValue = mac.value.trim();
const commandValue = command.value.trim();
if (
macValue &&
!/^[0-9a-fA-F]{2}([:-][0-9a-fA-F]{2}){5}(\s*,\s*[0-9a-fA-F]{2}([:-][0-9a-fA-F]{2}){5})*$/.test(macValue)
) {
if (status) status.textContent = 'MAC must look like 04:d9:f5:80:c6:58 (comma-separated for several).';
return;
}
if (commandValue && /\s/.test(commandValue)) {
if (status) status.textContent = 'The wake command must be a single executable path (no arguments).';
return;
}
if (save) save.disabled = true;
if (status) status.textContent = 'Saving …';
try {
const listRes = await fetch('/api/remote-hosts');
const listData = await listRes.json();
const hosts = listData.success ? listData.data : [];
const host = Array.isArray(hosts) ? hosts.find((item) => item.id === hostId) : null;
if (!host) throw new Error(this._wakeConfigUnavailableMessage());
// PUT takes the whole host (schema-validated), so send back everything we know and
// only replace the wake fields. `undefined` drops the key entirely.
const payload = {
...host,
wakeMac: macValue || undefined,
wakeCommand: commandValue || undefined,
};
const res = await fetch(`/api/remote-hosts/${encodeURIComponent(hostId)}`, {
method: 'PUT',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(payload),
});
const data = await res.json();
if (!data.success) throw new Error(data.error || 'Save failed');
this.showToast('Wake settings saved', 'success');
this.closeWakeConfigDialog();
// The server re-resolves host config for live sessions, so the banner can offer
// the wake right away — probe fresh instead of waiting out the poll interval.
await this._pollHostReachability(true);
} catch (err) {
if (status) status.textContent = err && err.message ? err.message : 'Save failed';
} finally {
if (save) save.disabled = false;
}
},
/**
* SSE `remote:hostWaking` — a wake is running (ours or one started by typing).
*
* ⚠️ The ONLY definition of this handler: `panels-ui.js` must not define it too.
* Both mix into `Codeman.prototype` and this file loads later, so a second copy
* would be silently shadowed (the guard in `sse-dispatch-table.test.ts` sees that a
* handler exists, not that two modules claim the same name). The toast is
* deliberately UNCONDITIONAL — a wake can start for a background session (input on
* a non-active tab) where there is no banner to update.
*/
_onRemoteHostWaking(data) {
const label = data && data.label ? data.label : 'Remote host';
// A create-path wake (the user pressed Run / Attach) has no session yet, so
// nothing is queued behind it — the wording has to say what actually happens.
const forNewSession = Boolean(data && data.forNewSession);
// Only the typing path buffers bytes; the wake button and the send-and-wait path
// hold none, and a browser keystroke never reaches the registry at all.
const queuedInput = Boolean(data && data.queuedInput);
// Long enough to cover the wake + attach (~10s measured on a warm S3), and it
// is replaced by `remote:sessionReconnected` the moment the pane is back.
this.showToast(
forNewSession
? `Waking ${label} … the session starts when it is back`
: queuedInput
? `Waking ${label} … input is queued`
: `Waking ${label} … waiting for it to come back`,
'info',
{ duration: 12000 }
);
const state = this._hostWake;
if (!state || !data || state.sessionId !== data.sessionId) return;
state.waking = true;
state.queuedInput = queuedInput;
state.error = '';
if (data.label) state.label = data.label;
this._renderHostWakeBanner();
},
/** SSE `remote:hostWakeFailed` — the host did not come back in time. */
_onRemoteHostWakeFailed(data) {
const label = data && data.label ? data.label : 'Remote host';
const forNewSession = Boolean(data && data.forNewSession);
const queuedInput = Boolean(data && data.queuedInput);
this.showToast(
forNewSession
? `${label} did not wake up — no session was started`
: queuedInput
? `${label} did not wake up — queued input is still held`
: `${label} did not wake up`,
'error',
{ duration: 15000 }
);
const state = this._hostWake;
if (!state || !data || state.sessionId !== data.sessionId) return;
state.waking = false;
state.queuedInput = queuedInput;
state.error = 'timeout';
state.reachable = false;
this._renderHostWakeBanner();
},
});
+29
View File
@@ -286,6 +286,34 @@
'Prompt sent': '提示已发送',
'Inserted, press Enter in the terminal to send': '已插入,在终端中按 Enter 发送',
'Could not reach the session': '无法连接到会话',
'Custom model endpoints': '自定义模型端点',
'Point a harness at your own OpenAI-compatible server (llama.cpp, vLLM, DGX Spark, Azure AI Foundry, OpenRouter) instead of its native cloud backend. When on, the Run menu offers an extra entry per harness that supports it, per saved endpoint.':
'让工具指向您自己的兼容 OpenAI 服务器(llama.cpp、vLLM、DGX Spark、Azure AI Foundry、OpenRouter),而非其原生云端后端。开启后,"运行"菜单会为每个支持此功能的工具、每个已保存的端点新增一个条目。',
'Enable custom model endpoints': '启用自定义模型端点',
'Adds a per-endpoint entry to the Run menu for every harness that can redirect to one.':
'为每个可重定向到端点的工具,在"运行"菜单中添加对应条目。',
'No endpoints yet. Add one below to point a harness at a local or cloud OpenAI-compatible server.':
'暂无端点。请在下方添加一个,以便将工具指向本地或云端的兼容 OpenAI 服务器。',
Discover: '发现模型',
'+ Add endpoint': '+ 添加端点',
'Add endpoint': '添加端点',
Id: 'ID',
'Short, stable — used in URLs, never shown to the CLI.': '简短且固定 — 用于 URL,不会展示给 CLI。',
Label: '标签',
'Base URL': '基础 URL',
'API key': 'API 密钥',
'Optional. Left blank on edit keeps the existing key.': '可选。编辑时留空将保留现有密钥。',
'Auth header': '认证请求头',
'Never send both — some servers hang indefinitely.': '切勿同时发送两者 — 部分服务器会因此无限期挂起。',
'Authorization: Bearer (default)': 'Authorization: Bearer(默认)',
'api-key header (Azure)': 'api-key 请求头(Azure)',
'Default model': '默认模型',
'What the Run-menu picker applies for this endpoint. Discover models first.':
'运行菜单选择器会为此端点应用该模型。请先发现可用模型。',
'Custom Endpoints': '自定义端点',
'Choose a model': '选择模型',
'That endpoint no longer exists': '该端点已不存在',
'No models discovered for this endpoint yet': '此端点尚未发现任何模型',
'Subagent Options': '子智能体选项',
'Enable Tracking': '启用跟踪',
'Active Tab Only': '仅活动标签页',
@@ -521,6 +549,7 @@
'Respawn Blocked': '重生已阻止',
'Task Complete': '任务完成',
'Copied to clipboard': '已复制到剪贴板',
'Nothing to copy': '没有可复制的内容',
// Terminal touch-selection bar (long-press to select). The bar is a sibling of
// `.xterm`, not a descendant, so SKIP_SELECTOR does not cover it and these apply.
Copy: '复制',
+176
View File
@@ -213,6 +213,18 @@
<button class="offline-banner-retry" id="offlineBannerRetry" onclick="app.retryConnection()">Retry now</button>
</div>
<!-- Remote-host unreachable: the machine SLEEPS, the local ssh pane stalls
silently (send-keys succeeds against it, so typed input would vanish) and
Codeman can wake it. Amber, not red: the session is fine, the host is
asleep. Without a configured wake target the action becomes "Configure
WoL" and opens the small config dialog. -->
<div class="offline-banner host-wake-banner" id="hostWakeBanner" role="status" hidden>
<span class="offline-banner-dot" aria-hidden="true"></span>
<span class="offline-banner-text" id="hostWakeBannerText">Remote host is unreachable</span>
<span class="offline-banner-detail" id="hostWakeBannerDetail"></span>
<button class="offline-banner-retry" id="hostWakeBannerAction" onclick="app.hostWakeAction()">Wake</button>
</div>
<!-- Reboot-restore offer: shown when the server found sessions a host reboot
killed and is asking whether to rebuild them. Populated by
reboot-restore-ui.js; nothing is created until the user clicks. -->
@@ -668,6 +680,14 @@
<button class="run-mode-option" data-mode="omp" onclick="app.setRunMode('omp')">
<span class="run-mode-dot omp"></span>OMP
</button>
<!-- Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md): one
generated entry per (harness, saved endpoint) pair, e.g. "Claude Code
(llama.cpp)". Built entirely by _refreshCustomModelRunOptions() — hidden
when the feature is off or no endpoint has a usable default model, never
a fixed per-harness duplicate in this markup. -->
<div class="run-mode-sep" id="runModeCustomModelSep" style="display: none;"></div>
<div class="run-mode-header" id="runModeCustomModelHeader" style="display: none;">Custom Endpoints</div>
<div class="run-mode-custom-models" id="runModeCustomModels"></div>
<div class="run-mode-sep"></div>
<button class="run-mode-option" data-mode="shell" onclick="app.setRunMode('shell')">
<span class="run-mode-dot shell"></span>Terminal / Shell
@@ -920,6 +940,67 @@
</div>
</div>
<!-- Custom Model Endpoint Profiles: "which model" picker (docs/custom-model-endpoints-plan.md).
Shown only when the chosen endpoint has more than one discovered model — see
selectCustomModelEntry() in session-ui.js, which skips straight to launch otherwise. -->
<div class="modal" id="customModelPickModal">
<div class="modal-backdrop" onclick="app.closeCustomModelPickModal()"></div>
<div class="modal-content modal-sm">
<div class="modal-header">
<h3 id="customModelPickTitle">Choose a model</h3>
<button class="modal-close" onclick="app.closeCustomModelPickModal()" aria-label="Close model picker">&times;</button>
</div>
<div class="modal-body">
<p class="form-hint" id="customModelPickHint"></p>
<div id="customModelPickList" class="run-mode-custom-models"></div>
</div>
</div>
</div>
<!-- Custom Model Endpoint Profiles: llama-swap model-swap confirmation
(docs/custom-model-endpoints-plan.md) — replaces a native confirm()
popup, shown when switching would unload a model another live
session is actively using. See _confirmModelSwap() in session-ui.js. -->
<div class="modal" id="customModelSwapConfirmModal">
<div class="modal-backdrop" onclick="app._resolveModelSwapConfirm(false)"></div>
<div class="modal-content modal-sm">
<div class="modal-header">
<h3>Switch models?</h3>
<button class="modal-close" onclick="app._resolveModelSwapConfirm(false)" aria-label="Cancel">&times;</button>
</div>
<div class="modal-body">
<p class="form-hint" id="customModelSwapConfirmMessage"></p>
</div>
<div class="modal-footer">
<button class="btn-toolbar" onclick="app._resolveModelSwapConfirm(false)">Cancel</button>
<button class="btn-toolbar btn-primary" onclick="app._resolveModelSwapConfirm(true)">Switch anyway</button>
</div>
</div>
</div>
<!-- Custom Model Endpoint Profiles: context-window-too-small warning
(docs/custom-model-endpoints-plan.md) — shown before launching a CLI
whose own fixed system-prompt/tool-schema overhead exceeds the
model's real discovered context, which guarantees a first-message
failure regardless of CLAUDE_CODE_MAX_CONTEXT_TOKENS. See
_confirmContextWarning() in session-ui.js. -->
<div class="modal" id="customModelContextWarningModal">
<div class="modal-backdrop" onclick="app._resolveContextWarningConfirm(false)"></div>
<div class="modal-content modal-sm">
<div class="modal-header">
<h3>Context window too small</h3>
<button class="modal-close" onclick="app._resolveContextWarningConfirm(false)" aria-label="Cancel">&times;</button>
</div>
<div class="modal-body">
<p class="form-hint" id="customModelContextWarningMessage" style="white-space: pre-wrap;"></p>
</div>
<div class="modal-footer">
<button class="btn-toolbar" onclick="app._resolveContextWarningConfirm(false)">Cancel</button>
<button class="btn-toolbar btn-primary" onclick="app._resolveContextWarningConfirm(true)">Launch anyway</button>
</div>
</div>
</div>
<!-- Cron Jobs Modal -->
<div class="modal" id="cronModal">
<div class="modal-backdrop" onclick="app.closeCron()"></div>
@@ -2216,6 +2297,60 @@
</div>
</div>
</div>
<div class="set-group" id="customModelEndpointsGroup">
<div class="set-group-head"><h4>Custom model endpoints</h4><span class="set-scope">synced</span></div>
<p class="set-group-hint">Point a harness at your own OpenAI-compatible server (llama.cpp, vLLM, DGX Spark, Azure AI Foundry, OpenRouter) instead of its native cloud backend. When on, the Run menu offers an extra entry per harness that supports it, per saved endpoint.</p>
<div class="set-group-body">
<div class="set-row" data-search="custom model endpoint llama.cpp local llm run menu picker">
<div class="set-row-text">
<span class="set-row-label">Enable custom model endpoints</span>
<span class="set-row-desc">Adds a per-endpoint entry to the Run menu for every harness that can redirect to one.</span>
</div>
<label class="switch switch-sm"><input type="checkbox" id="appSettingsCustomModelEndpoints" onchange="app.applyCustomModelEndpointsVisibility()"><span class="slider"></span></label>
</div>
<!-- Gated on the toggle above (applyCustomModelEndpointsVisibility): with the
feature off, a list of endpoints that do nothing is worse than nothing. -->
<div id="customModelEndpointsBody" style="display:none">
<div id="customModelHostsList" class="set-group-body" data-search="endpoints"></div>
<button type="button" class="btn-toolbar btn-sm" id="customModelHostAddBtn" onclick="app.openCustomModelHostEditor()">+ Add endpoint</button>
<div id="customModelHostEditor" class="set-inline-form" style="display:none">
<h5 id="customModelHostEditorTitle">Add endpoint</h5>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Id</span><span class="set-row-desc">Short, stable — used in URLs, never shown to the CLI.</span></div>
<input type="text" id="customModelHostId" class="set-input" placeholder="llama-cpp-local">
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Label</span></div>
<input type="text" id="customModelHostLabel" class="set-input" placeholder="llama.cpp (local)">
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Base URL</span></div>
<input type="text" id="customModelHostBaseUrl" class="set-input" placeholder="http://192.168.1.50:8080">
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">API key</span><span class="set-row-desc">Optional. Left blank on edit keeps the existing key.</span></div>
<input type="password" id="customModelHostApiKey" class="set-input" autocomplete="new-password">
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Auth header</span><span class="set-row-desc">Never send both — some servers hang indefinitely.</span></div>
<select id="customModelHostAuthStyle" class="set-select">
<option value="bearer">Authorization: Bearer (default)</option>
<option value="api-key">api-key header (Azure)</option>
</select>
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Default model</span><span class="set-row-desc">What the Run-menu picker applies for this endpoint. Discover models first.</span></div>
<select id="customModelHostDefaultModel" class="set-select" disabled></select>
</div>
<div class="set-row-actions">
<button type="button" class="btn-toolbar btn-sm" onclick="app.saveCustomModelHostFromEditor()">Save</button>
<button type="button" class="btn-toolbar btn-sm" onclick="app.closeCustomModelHostEditor()">Cancel</button>
</div>
</div>
</div>
</div>
</div>
</section>
<!-- ══ Agents &amp; CLIs ═════════════════════════════════════════ -->
@@ -2878,6 +3013,11 @@
<input type="number" id="remoteHostPort" placeholder="22" min="1" max="65535" autocomplete="off">
<span class="form-hint">Optional. Leave blank for the default port 22.</span>
</div>
<div class="form-row">
<label>Wake-on-LAN MAC</label>
<input type="text" id="remoteHostWakeMac" placeholder="04:d9:f5:80:c6:58" autocomplete="off" autocapitalize="off" spellcheck="false">
<span class="form-hint">Optional. Comma-separated for several NICs. Codeman sends the magic packet itself so a sleeping host can be woken from the session banner.</span>
</div>
<div class="form-row">
<label>Codex Command Override</label>
<input type="text" id="remoteHostCodexCommand" placeholder="exec codx personal" autocomplete="off" autocapitalize="off" spellcheck="false">
@@ -2886,6 +3026,11 @@
<details class="advanced-options">
<summary><svg class="set-adv-chev" width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.4" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M6 9l6 6 6-6"/></svg><span>Advanced SSH</span></summary>
<div class="advanced-options-content">
<div class="form-row">
<label>Wake Command</label>
<input type="text" id="remoteHostWakeCommand" placeholder="/home/user/bin/wake-this-host" autocomplete="off" autocapitalize="off" autocorrect="off" spellcheck="false">
<span class="form-hint">Optional override for the MAC above (takes precedence). A single executable path, run without a shell — use it when the host needs a router/other machine to send the packet.</span>
</div>
<div class="form-row">
<label>Identity File</label>
<input type="text" id="remoteHostIdentityFile" placeholder="~/.ssh/remote_ed25519" autocomplete="off" autocapitalize="off" autocorrect="off" spellcheck="false">
@@ -3481,6 +3626,36 @@
text is set via value/textContent only: predictor output derives from
observable (injectable) content, and the explicit click here is the
security boundary (nothing is ever auto-sent). -->
<!-- Wake-on-LAN setup for a remote host whose session cannot be woken yet. Kept
deliberately small (host is fixed, only the wake fields are editable) so it can
be opened from the banner with one click. Persists via PUT /api/remote-hosts/:id. -->
<div class="modal" id="wakeConfigModal">
<div class="modal-backdrop" onclick="app.closeWakeConfigDialog()"></div>
<div class="modal-content">
<div class="modal-header">
<h3>Wake-on-LAN &middot; <span id="wakeConfigHostLabel"></span></h3>
<button class="modal-close" onclick="app.closeWakeConfigDialog()" aria-label="Close">&times;</button>
</div>
<div class="modal-body">
<div class="form-row">
<label>MAC address(es)</label>
<input type="text" id="wakeConfigMac" placeholder="04:d9:f5:80:c6:58" autocomplete="off" autocapitalize="off" autocorrect="off" spellcheck="false">
<span class="form-hint">Comma-separated for several NICs. Codeman sends the magic packet itself (UDP port 9, broadcast).</span>
</div>
<div class="form-row">
<label>Wake command (optional)</label>
<input type="text" id="wakeConfigCommand" placeholder="/home/user/bin/wake-this-host" autocomplete="off" autocapitalize="off" autocorrect="off" spellcheck="false">
<span class="form-hint">Takes precedence over the MAC. A single executable path, run without a shell.</span>
</div>
<div class="form-hint" id="wakeConfigStatus"></div>
</div>
<div class="modal-footer">
<button class="btn-toolbar" onclick="app.closeWakeConfigDialog()">Cancel</button>
<button class="btn-toolbar btn-primary" id="wakeConfigSave" onclick="app.saveWakeConfig()">Save</button>
</div>
</div>
</div>
<div class="modal" id="readMyMindModal">
<div class="modal-backdrop" onclick="app.closeReadMyMind()"></div>
<div class="modal-content readmymind-modal">
@@ -3556,6 +3731,7 @@
<script defer src="reboot-restore-ui.js"></script>
<script defer src="admin-ui.js"></script>
<script defer src="session-ui.js"></script>
<script defer src="host-wake-ui.js"></script>
<script defer src="webview-tabs.js"></script>
<script defer src="mobile-overview.js"></script>
<script defer src="home-sessions.js"></script>
+138 -4
View File
@@ -92,6 +92,9 @@ Object.assign(CodemanApp.prototype, {
_onRemoteSessionReconnected(data) {
const id = this.getShortId(data.sessionId);
this.showToast(`Remote session ${id} reconnected`, 'success');
// A successful reattach (the wake flow's own, or the watcher's) means the host is
// back: drop the "unreachable" banner without waiting out the poll interval.
if (this.activeSessionId === data.sessionId) this._pollHostReachability?.(true);
},
_onRemoteReconnectExhausted(data) {
@@ -116,6 +119,16 @@ Object.assign(CodemanApp.prototype, {
},
// Wake-on-LAN from user input on a sleeping remote host (see remote-wake.ts).
// ⚠️ The `remote:hostWaking` / `remote:hostWakeFailed` HANDLERS live in
// `host-wake-ui.js`, which owns the banner state. They are NOT redefined here:
// both files mix into `CodemanApp.prototype` and `host-wake-ui.js` is loaded
// later, so a second definition would silently shadow the banner update (and the
// toast would never fire — the exact silent no-op `sse-dispatch-table.test.ts`
// exists to prevent, which cannot see shadowing). The toasts are shown from the
// host-wake-ui handlers instead.
// Bash tools
_onBashToolStart(data) {
this.handleBashToolStart(data.sessionId, data.tool);
@@ -5484,12 +5497,25 @@ Object.assign(CodemanApp.prototype, {
return this.showToast(message, type);
},
/**
* `duration` defaults to 3000ms for every toast type. A message worth
* reading rather than glancing at (e.g. "Session started on the native
* backend — could not apply the custom endpoint: <the actual reason>")
* passes an explicit `opts.duration: 0` at its own call site instead of
* widening the default: this used to default every `error` toast to
* sticky, and with no cap on `.toast-container` and no eviction, a
* repeatedly failing path (a flapping SSE reconnect, a poll loop) stacked
* sticky toasts off the bottom of the viewport where they could not be
* read or dismissed. Every toast still gets an explicit close button
* regardless of duration.
*/
showToast(message, type = 'info', opts = {}) {
const { duration = 3000, action } = opts;
const toast = document.createElement('div');
toast.className = `toast toast-${type}`;
const msgSpan = document.createElement('span');
msgSpan.className = 'toast-message';
msgSpan.textContent = message;
toast.appendChild(msgSpan);
@@ -5501,6 +5527,20 @@ Object.assign(CodemanApp.prototype, {
toast.appendChild(btn);
}
let dismissTimer = null;
const dismiss = () => {
if (dismissTimer) clearTimeout(dismissTimer);
toast.classList.remove('show');
setTimeout(() => toast.remove(), 200);
};
const closeBtn = document.createElement('button');
closeBtn.className = 'toast-close';
closeBtn.textContent = '×';
closeBtn.setAttribute('aria-label', 'Dismiss');
closeBtn.onclick = (e) => { e.stopPropagation(); dismiss(); };
toast.appendChild(closeBtn);
// Cache toast container reference
if (!this._toastContainer) {
this._toastContainer = document.querySelector('.toast-container');
@@ -5514,10 +5554,104 @@ Object.assign(CodemanApp.prototype, {
requestAnimationFrame(() => toast.classList.add('show'));
setTimeout(() => {
toast.classList.remove('show');
setTimeout(() => toast.remove(), 200);
}, duration);
if (duration > 0) {
dismissTimer = setTimeout(dismiss, duration);
}
// Most callers ignore this — a handle exists for a long-running toast a caller needs
// to update or dismiss itself once its own condition resolves (e.g. a "loading model"
// toast a poll loop dismisses once the model reports ready).
return { dismiss, setMessage: (text) => { msgSpan.textContent = text; } };
},
/**
* A prominent, screen-centred status banner — for the small set of messages that are
* genuinely worth interrupting the eye for rather than living in the corner with every
* other toast (currently: a custom-model session's "switching backends" and "loading
* model" states, both of which can sit on screen for well over a minute and are easy to
* mistake for nothing happening). Non-blocking (`pointer-events: none` on the wrapper,
* restored only on the card) — an info banner is never a gate the user has to dismiss to
* keep working. Only one is ever shown at a time (the DOM node is created once and
* reused), which matches every current caller: each hands off to the next rather than
* stacking.
*
* `opts.type` — `'info'` (default, spinner, no close button — a caller ends it itself via
* `dismiss()`) or `'error'` (no spinner — nothing is in progress once this shows — with a
* close button, since a sticky error the user cannot dismiss would just sit there). The
* DOM is rebuilt fresh each call rather than patched, since which children exist differs
* by type; `setMessage` still only ever touches the text node afterwards.
*
* `opts.onCancel` — when given (any type, but in practice only 'info': an 'error' banner
* already has its own close button), renders a "Cancel" button that calls it on click.
* The callback owns everything that follows (dismissing the banner, stopping whatever
* loop this was showing progress for, closing a session it was for) — this helper only
* renders the button and wires the click, the same "caller decides what cancel means"
* split as `_confirmModelSwap`'s promise-resolving buttons.
*/
_showCenterStatus(message, opts = {}) {
const { type = 'info', onCancel } = opts;
let el = document.getElementById('customModelCenterStatus');
if (!el) {
el = document.createElement('div');
el.id = 'customModelCenterStatus';
document.body.appendChild(el);
}
// A pending hide from a PREVIOUS dismiss() (e.g. switchingToast.dismiss() right
// before this same-origin call reopens the banner within its 200ms fade) must
// never fire against the node this call is about to show — clear it before
// reusing the shared DOM node, or the old timer hides the fresh banner ~200ms in.
if (el._hideTimer) {
clearTimeout(el._hideTimer);
el._hideTimer = null;
}
el.className = `center-status-banner center-status-${type}`;
el.innerHTML = '';
const dismiss = () => {
el.classList.remove('show');
el._hideTimer = setTimeout(() => {
el.hidden = true;
el._hideTimer = null;
}, 200);
};
if (type !== 'error') {
const spinner = document.createElement('span');
spinner.className = 'center-status-spinner';
spinner.setAttribute('aria-hidden', 'true');
el.appendChild(spinner);
}
const text = document.createElement('span');
text.className = 'center-status-text';
text.textContent = message;
el.appendChild(text);
if (type === 'error') {
const closeBtn = document.createElement('button');
closeBtn.className = 'center-status-close';
closeBtn.textContent = '×';
closeBtn.setAttribute('aria-label', 'Dismiss');
closeBtn.onclick = (e) => {
e.stopPropagation();
dismiss();
};
el.appendChild(closeBtn);
} else if (onCancel) {
const cancelBtn = document.createElement('button');
cancelBtn.className = 'center-status-cancel';
cancelBtn.textContent = 'Cancel';
cancelBtn.onclick = (e) => {
e.stopPropagation();
onCancel();
};
el.appendChild(cancelBtn);
}
el.hidden = false;
requestAnimationFrame(() => el.classList.add('show'));
return {
dismiss,
setMessage: (next) => {
const t = el.querySelector('.center-status-text');
if (t) t.textContent = next;
},
};
},
File diff suppressed because it is too large Load Diff
+226
View File
@@ -395,6 +395,13 @@ Object.assign(CodemanApp.prototype, {
document.getElementById('appSettingsShowUltracodeAgents').checked = settings.showUltracodeAgents ?? defaults.showUltracodeAgents ?? false;
// Approvals Inbox: synced, default OFF (opt-in; only an explicit true enables).
document.getElementById('appSettingsApprovalsInbox').checked = settings.approvalsInboxEnabled === true;
// Custom Model Endpoint Profiles: synced, default OFF. The toggle governs both
// the Run-menu picker's generated entries and this settings panel's visibility;
// the endpoint list itself is server state, loaded on demand below.
document.getElementById('appSettingsCustomModelEndpoints').checked = settings.customModelEndpointsEnabled === true;
// Assigning .checked above does not fire onchange, so the body's visibility
// (and its lazy load) needs an explicit sync on every open, not just a save.
this.applyCustomModelEndpointsVisibility();
// Read My Mind: synced, default OFF (opt-in; capture + prediction cost real tokens).
document.getElementById('appSettingsReadMyMind').checked = settings.readMyMindEnabled === true;
document.getElementById('appSettingsUltracodeFloatingWindows').checked =
@@ -509,6 +516,9 @@ Object.assign(CodemanApp.prototype, {
document.getElementById('appSettingsNiceValue').value = niceSettings.niceValue ?? 10;
// Model configuration (loaded from server)
this.loadModelConfigForSettings();
// Custom Model Endpoint Profiles' own load is gated on the toggle above (see
// applyCustomModelEndpointsVisibility) — unlike model config, this GET is
// pointless work with the feature off, so it is not fired unconditionally.
// Notification settings
const notifPrefs = this.notificationManager?.preferences || {};
document.getElementById('appSettingsNotifEnabled').checked = notifPrefs.enabled ?? true;
@@ -2106,6 +2116,7 @@ Object.assign(CodemanApp.prototype, {
showSubagents: document.getElementById('appSettingsShowSubagents').checked,
showUltracodeAgents: document.getElementById('appSettingsShowUltracodeAgents').checked,
approvalsInboxEnabled: document.getElementById('appSettingsApprovalsInbox').checked,
customModelEndpointsEnabled: document.getElementById('appSettingsCustomModelEndpoints').checked,
readMyMindEnabled: document.getElementById('appSettingsReadMyMind').checked,
ultracodeFloatingWindows: document.getElementById('appSettingsUltracodeFloatingWindows').checked,
showMultiMonitorButton: document.getElementById('appSettingsShowMultiMonitorButton').checked,
@@ -2487,6 +2498,209 @@ Object.assign(CodemanApp.prototype, {
}
},
// ═══════════════════════════════════════════════════════════════
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md)
//
// CRUD against /api/model-endpoints, rendered into the Models settings section.
// Deliberately its own load/save pair rather than folded into openAppSettings/
// saveAppSettings: these are server-side infra records (like remote/docker
// hosts), not a settings-payload field, so the app-settings-structure guard's
// by-id contract does not apply to them — only the `customModelEndpointsEnabled`
// toggle itself goes through that path.
// ═══════════════════════════════════════════════════════════════
/**
* Toggles the endpoint-management body's visibility to match the setting and,
* turning it on, lazily loads the endpoint list. Assigning `.checked` (as the
* settings load path does) fires no `change` event, so this must be called
* explicitly on open as well as wired to the checkbox's own onchange — a
* gate that only worked one of those two ways would show a stale "off"
* body right after opening, or a stale "on" one right after saving it off.
* With the feature off the body is a list of controls that do nothing, so it
* is hidden entirely rather than shown disabled.
*/
applyCustomModelEndpointsVisibility() {
const enabled = document.getElementById('appSettingsCustomModelEndpoints').checked;
const body = document.getElementById('customModelEndpointsBody');
if (body) body.style.display = enabled ? '' : 'none';
if (enabled) this.loadCustomModelEndpointsForSettings();
else this.closeCustomModelHostEditor();
this._applyCustomModelAdminGate();
},
/**
* Endpoint writes are admin-only in multi-user mode (custom-model-routes.ts),
* and GET already answers a non-admin with an empty list, which hides every
* per-row Edit/Discover/Delete button on its own. The "+ Add endpoint" button
* has no row to hide behind, so it needs its own gate — otherwise a non-admin
* can open the form, fill it in, and get a 403 toast on Save. Wired to the
* `codeman:me` event (admin-ui.js) as well as called from
* applyCustomModelEndpointsVisibility(), because `window.__codemanUser`'s
* real role can resolve AFTER settings have already been opened once.
*/
_applyCustomModelAdminGate() {
const addBtn = document.getElementById('customModelHostAddBtn');
if (!addBtn) return;
const me = window.__codemanUser || {};
const blocked = me.multiUser && me.role !== 'admin';
addBtn.style.display = blocked ? 'none' : '';
},
async loadCustomModelEndpointsForSettings() {
// GET /api/model-endpoints wraps its body in the { success, data } envelope
// like every other /api route (server.ts's preSerialization hook applies to
// arrays too) — _apiJson() unwraps it. A raw fetch().json() here would
// silently see the envelope object instead of the array and this panel
// would read as "No endpoints yet" forever, even with endpoints saved.
const hosts = await this._apiJson('/api/model-endpoints');
this._customModelHosts = Array.isArray(hosts) ? hosts : [];
this.renderCustomModelHostsList();
},
renderCustomModelHostsList() {
const list = document.getElementById('customModelHostsList');
if (!list) return;
const hosts = this._customModelHosts || [];
if (hosts.length === 0) {
list.innerHTML = '<p class="set-group-hint">No endpoints yet. Add one below to point a harness at a local or cloud OpenAI-compatible server.</p>';
return;
}
list.innerHTML = hosts
.map((h) => {
const modelCount = (h.models || []).length;
const modelSummary = modelCount === 0
? 'No models discovered yet'
: `${modelCount} model${modelCount === 1 ? '' : 's'}${h.defaultModelId ? ` · default: ${escapeHtml(h.defaultModelId)}` : ' · no default set'}`;
// escapeHtml(JSON.stringify(h.id)) — not JSON.stringify(h.id) alone —
// because JSON.stringify's own double quotes would otherwise terminate
// this double-quoted attribute at the first one, and everything after
// parses as raw tag content rather than the rest of the quoted string.
// Same idiom as deleteCase's onclick in session-ui.js. h.id is
// regex-constrained server-side (safe either way) but the pattern must
// match everywhere it is used, including where the argument is not.
const idArg = escapeHtml(JSON.stringify(h.id));
return `
<div class="set-row" data-endpoint-id="${escapeHtml(h.id)}">
<div class="set-row-text">
<span class="set-row-label">${escapeHtml(h.label)}</span>
<span class="set-row-desc">${escapeHtml(h.baseUrl)} — ${modelSummary}</span>
</div>
<div class="set-row-actions">
<button type="button" class="btn-toolbar btn-sm" onclick="app.discoverCustomModelHostModels(${idArg})">Discover</button>
<button type="button" class="btn-toolbar btn-sm" onclick="app.openCustomModelHostEditor(${idArg})">Edit</button>
<button type="button" class="btn-toolbar btn-danger btn-sm" onclick="app.deleteCustomModelHost(${idArg})">Delete</button>
</div>
</div>`;
})
.join('');
},
/** Opens the inline add/edit form. Pass no id to add a new endpoint. */
openCustomModelHostEditor(hostId) {
const host = hostId ? (this._customModelHosts || []).find((h) => h.id === hostId) : null;
this._editingCustomModelHostId = host ? host.id : null;
document.getElementById('customModelHostEditorTitle').textContent = host ? `Edit ${host.label}` : 'Add endpoint';
document.getElementById('customModelHostId').value = host?.id || '';
document.getElementById('customModelHostId').disabled = !!host; // id is immutable once created
document.getElementById('customModelHostLabel').value = host?.label || '';
document.getElementById('customModelHostBaseUrl').value = host?.baseUrl || '';
document.getElementById('customModelHostApiKey').value = ''; // the server never returns the real value (apiKeySet is a bool)
document.getElementById('customModelHostApiKey').placeholder = host?.apiKeySet ? '•••••••• (unchanged if left blank)' : '';
document.getElementById('customModelHostAuthStyle').value = host?.authStyle || 'bearer';
this._populateCustomModelDefaultSelect(host);
document.getElementById('customModelHostEditor').style.display = '';
},
closeCustomModelHostEditor() {
document.getElementById('customModelHostEditor').style.display = 'none';
this._editingCustomModelHostId = null;
},
_populateCustomModelDefaultSelect(host) {
const select = document.getElementById('customModelHostDefaultModel');
const models = host?.models || [];
select.innerHTML =
'<option value="">No default (picker uses the first discovered model)</option>' +
models.map((m) => `<option value="${escapeHtml(m)}">${escapeHtml(m)}</option>`).join('');
select.value = host?.defaultModelId || '';
select.disabled = models.length === 0;
},
async saveCustomModelHostFromEditor() {
const id = document.getElementById('customModelHostId').value.trim();
const label = document.getElementById('customModelHostLabel').value.trim();
const baseUrl = document.getElementById('customModelHostBaseUrl').value.trim();
const apiKeyInput = document.getElementById('customModelHostApiKey').value;
const authStyle = document.getElementById('customModelHostAuthStyle').value;
const defaultModelId = document.getElementById('customModelHostDefaultModel').value || undefined;
if (!id || !label || !baseUrl) {
this.showToast('Id, label and base URL are all required', 'warning');
return;
}
const editing = this._editingCustomModelHostId;
// PUT (server-side) treats an absent apiKey as "keep the stored one" — the
// browser never holds the real value to resend deliberately unchanged (see
// openCustomModelHostEditor and custom-model-routes.ts's applyStoredApiKey),
// so a blank field here means omitting the key entirely, not resending
// something we do not have. models/lastDiscoveredAt DO still need
// re-sending: PUT replaces the whole record, and this cached copy still
// carries both (only apiKey is redacted from what GET hands back).
const existing = editing ? (this._customModelHosts || []).find((h) => h.id === editing) : null;
const body = {
id,
label,
baseUrl,
authStyle,
defaultModelId,
apiKey: apiKeyInput || undefined,
models: existing?.models,
lastDiscoveredAt: existing?.lastDiscoveredAt,
};
try {
const res = await fetch(editing ? `/api/model-endpoints/${encodeURIComponent(editing)}` : '/api/model-endpoints', {
method: editing ? 'PUT' : 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(body),
});
const data = await res.json();
if (!data.success) {
this.showToast(data.error || 'Failed to save endpoint', 'error');
return;
}
this.showToast(editing ? 'Endpoint updated' : 'Endpoint added', 'success');
this.closeCustomModelHostEditor();
await this.loadCustomModelEndpointsForSettings();
} catch (err) {
this.showToast(`Failed to save endpoint: ${err.message}`, 'error');
}
},
async discoverCustomModelHostModels(hostId) {
this.showToast('Discovering models…', 'info');
try {
const res = await fetch(`/api/model-endpoints/${encodeURIComponent(hostId)}/discover-models`, { method: 'POST' });
const data = await res.json();
if (!data.success) {
this.showToast(data.error || 'Discovery failed', 'error');
return;
}
this.showToast(`Found ${data.data.models.length} model${data.data.models.length === 1 ? '' : 's'}`, 'success');
await this.loadCustomModelEndpointsForSettings();
} catch (err) {
this.showToast(`Discovery failed: ${err.message}`, 'error');
}
},
async deleteCustomModelHost(hostId) {
const host = (this._customModelHosts || []).find((h) => h.id === hostId);
if (!confirm(`Delete endpoint "${host?.label || hostId}"? Any session currently pointed at it keeps running until cleared.`)) return;
try {
await fetch(`/api/model-endpoints/${encodeURIComponent(hostId)}`, { method: 'DELETE' });
await this.loadCustomModelEndpointsForSettings();
} catch (err) {
this.showToast(`Failed to delete endpoint: ${err.message}`, 'error');
}
},
// ═══════════════════════════════════════════════════════════════
// Visibility Settings & Device-Specific Defaults
@@ -3543,3 +3757,15 @@ Object.assign(CodemanApp.prototype, {
this.subagentPanelVisible = false;
},
});
// window.__codemanUser's real role can resolve after settings have already been
// opened once (admin-ui.js fetches /api/me asynchronously and dispatches this on
// arrival), so the Custom Model Endpoints admin gate needs to be re-applied when
// it does, not just when the modal opens. Optional chaining on addEventListener
// itself: several frontend tests (run-mode-ui.test.ts) load this file into a vm
// context with a minimal fake `document` that has no event-target methods at
// all, and a module-level statement that throws there fails the whole file's
// evaluation, not just this feature.
document.addEventListener?.('codeman:me', () => {
window.app?._applyCustomModelAdminGate?.();
});
+249
View File
@@ -6807,6 +6807,67 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
min-height: 0;
}
/* Custom Model Endpoint Profiles' "which model" picker: same bounded-height +
scrollable-body shape as .modal-lg above, scoped by id rather than added to
.modal-sm itself (three other modals share that class for short, fixed
content and do not need a height cap). Without this the modal had no
max-height at all, so an endpoint with many discovered models grew the
dialog past the viewport with nothing to scroll — "the whole page" and
"the list is truncated" turned out to be one and the same bug. `min(70vh,
520px)` scales with the monitor (a phone gets 70% of its height, a 4K
display never gets a needlessly tall dialog) rather than a fixed value
that would be wrong at one end or the other. */
#customModelPickModal .modal-content {
max-height: min(70vh, 520px);
display: flex;
flex-direction: column;
}
#customModelPickModal .modal-body {
overflow-y: auto;
flex: 1;
min-height: 0;
}
/* Custom Model Endpoint Profiles: llama-swap model-swap confirmation — replaces a native
confirm() popup (docs/custom-model-endpoints-plan.md) so it looks and feels like the
rest of the app instead of a browser chrome dialog. Shares the context-window-too-small
modal's fixes below since both can appear mid-launch, in the same spot, for the same
reason — including this rule itself: there is no bare `.modal-footer` base style
anywhere in this file, and `.btn-toolbar` is `display: flex` (a block-level flex
container with no explicit `inline-flex`), so with no row layout of its own each
button took its own full-width line and the two stacked instead of sitting side by
side. Centred rather than flex-end per feedback — a two-button Cancel/confirm footer
reads better centred than pinned to one edge. */
#customModelSwapConfirmModal .modal-footer,
#customModelContextWarningModal .modal-footer {
display: flex;
justify-content: center;
gap: 0.5rem;
padding: 0.75rem 1rem;
border-top: 1px solid var(--border-color);
}
/* Both dialogs can appear while the centred llama-swap status banner (10001, see
.center-status-banner) is still on screen — right after "Claude started — switching
to llama-swap…" — and .modal's own z-index (1000) sat well under it, so the dialog
rendered fully hidden behind the banner (confirmed live, reported against the
context-window one but structurally identical for the swap-confirm modal too). */
#customModelSwapConfirmModal,
#customModelContextWarningModal {
z-index: 10010;
}
/* Both messages ARE the modal's whole explanatory content, not a one-line caption under
a form field, so .form-hint's 0.65rem caption size (right for what it was designed for)
read as illegibly small here, worst on the multi-sentence context-window explanation. */
#customModelSwapConfirmMessage,
#customModelContextWarningMessage {
font-size: 0.85rem;
line-height: 1.5;
color: var(--text);
}
/* Mobile Case Picker - Base Styles */
.mobile-case-picker-sheet {
@@ -8445,6 +8506,9 @@ kbd {
}
.toast {
display: flex;
align-items: center;
gap: 0.5rem;
background: var(--bg-card);
border: 1px solid var(--border);
border-radius: 6px;
@@ -8456,6 +8520,7 @@ kbd {
opacity: 0;
transition: all 0.2s ease;
pointer-events: auto;
max-width: 420px;
}
.toast.show {
@@ -8463,6 +8528,149 @@ kbd {
opacity: 1;
}
.toast-message {
flex: 1;
/* A sticky toast (showToast's opts.duration: 0) can carry a longer, specific
message — let it wrap instead of clipping. */
white-space: pre-wrap;
word-break: break-word;
}
/* Every toast gets one, sticky or not: a sticky toast with no way to close it
would just accumulate on screen across repeated failures. */
.toast-close {
flex-shrink: 0;
background: none;
border: none;
color: inherit;
opacity: 0.6;
font-size: 1.1rem;
line-height: 1;
padding: 0 0.15rem;
cursor: pointer;
}
.toast-close:hover {
opacity: 1;
}
/* Custom Model Endpoint Profiles: the "switching backends" / "loading model" states
(docs/custom-model-endpoints-plan.md) — a small set of messages prominent and
screen-centred rather than corner toasts, since they can sit on screen for well
over a minute (a real llama-swap model load) and are easy to mistake for nothing
happening. Non-blocking: `pointer-events: none` on the wrapper (no backdrop, no
click-catcher) with `auto` restored only on the card itself, purely so the text
inside remains selectable — there is nothing to click to dismiss it early. */
.center-status-banner {
position: fixed;
top: 50%;
left: 50%;
transform: translate(-50%, -50%) scale(0.96);
z-index: 10001;
display: flex;
align-items: center;
gap: 0.75rem;
background: var(--bg-card);
border: 1px solid var(--border);
border-radius: 10px;
padding: 1rem 1.5rem;
box-shadow: 0 8px 32px rgba(0, 0, 0, 0.4);
font-size: 0.95rem;
font-weight: 500;
color: var(--text);
max-width: min(90vw, 460px);
text-align: left;
opacity: 0;
pointer-events: none;
transition:
opacity 0.2s ease,
transform 0.2s ease;
}
.center-status-banner.show {
opacity: 1;
transform: translate(-50%, -50%) scale(1);
}
/* `hidden` has to be re-asserted over the `display: flex` above, or `dismiss()`
setting `el.hidden = true` does nothing (same trap as `.home-sessions[hidden]`
below): the card stays laid out at `opacity: 0` with its text/cancel/close
children still `pointer-events: auto`, an invisible click-blocker dead centre
over the terminal until the page reloads. */
.center-status-banner[hidden] {
display: none;
}
.center-status-spinner {
flex-shrink: 0;
width: 18px;
height: 18px;
border-radius: 50%;
border: 2px solid var(--border);
border-top-color: var(--accent, var(--text));
animation: center-status-spin 0.8s linear infinite;
}
@keyframes center-status-spin {
to {
transform: rotate(360deg);
}
}
.center-status-text {
flex: 1;
pointer-events: auto;
white-space: pre-wrap;
word-break: break-word;
}
/* Error variant: the load didn't finish in time — nothing is "in progress" anymore (no
spinner), and since this one doesn't dismiss itself, it needs a close button the user
can actually click, so pointer-events is restored here too (see the wrapper's own
comment on why that's `none` by default). */
.center-status-error {
border-color: rgba(239, 68, 68, 0.5);
}
.center-status-close {
flex-shrink: 0;
pointer-events: auto;
background: none;
border: none;
color: inherit;
opacity: 0.6;
font-size: 1.2rem;
line-height: 1;
padding: 0 0.15rem;
cursor: pointer;
}
.center-status-close:hover {
opacity: 1;
}
/* The Cancel button on an 'info' banner (e.g. the model-loading banner) — a real button
rather than the bare "×" close glyph above, since "Cancel" is an action with a
consequence (the caller's onCancel closes a session), not a plain dismiss. */
.center-status-cancel {
flex-shrink: 0;
pointer-events: auto;
background: none;
border: 1px solid var(--border);
border-radius: 6px;
color: inherit;
opacity: 0.75;
font-size: 0.8rem;
font-weight: 500;
padding: 0.25rem 0.6rem;
cursor: pointer;
}
.center-status-cancel:hover {
opacity: 1;
border-color: var(--text-muted, var(--border));
}
.toast-success { border-color: rgba(34, 197, 94, 0.4); }
.toast-error { border-color: rgba(239, 68, 68, 0.4); }
.toast-warning { border-color: rgba(234, 179, 8, 0.4); }
@@ -15138,6 +15346,12 @@ html[data-skin="daylight-blue"] .welcome-btn-tunnel.active:hover {
.run-mode-dot.web { background: #38bdf8; }
.run-mode-webviews { max-height: 180px; overflow-y: auto; }
/* Custom Model Endpoint Profiles' generated entries: `.run-mode-menu.active`'s
own `gap: 2px` only spaces its DIRECT children, and this container (like
`.run-mode-webviews` above) is one such child holding several buttons of
its own, so it needs the same gap repeated one level down or its rows sit
flush against each other. */
.run-mode-custom-models { display: flex; flex-direction: column; gap: 2px; }
/* A saved URL is a ROW: open on the left, edit + delete on the right, so a URL can
be changed or removed without first opening it as a tab. The side buttons stay
@@ -15389,6 +15603,18 @@ html[data-skin="daylight-blue"] .welcome-btn-tunnel.active:hover {
background: rgba(255, 255, 255, 0.24);
}
/* Remote-host unreachable (host asleep, can be woken). Reuses the offline-banner
layout and children; amber instead of red because the Codeman session itself is
perfectly healthy — only the machine is asleep. */
.host-wake-banner {
background: linear-gradient(90deg, #b45309, #92400e);
}
.host-wake-banner .offline-banner-retry:disabled {
opacity: 0.6;
cursor: default;
}
/* Above the mobile fixed header (1200) and modals (1300): this is a blocking
"nothing works right now" state, and it only appears before any session
state has loaded, so there is no modal underneath to bury. Stays below the
@@ -16290,6 +16516,29 @@ html[data-tab-orientation='vertical'] .home-sessions {
gap: 3px;
}
/* Custom Model Endpoint Profiles' inline add/edit form: a nested panel rather
than a modal, so it needs its own border to read as a distinct sub-section
inside .set-group-body's flat row stack. `--control-bg` rather than a
hardcoded black alpha — CLAUDE.md records that literal fill turning the
settings live preview into a grey slab on the light skins, and this panel
sits in the very same modal. */
:is(#appSettingsModal, #sessionOptionsModal, #createCaseModal) .set-inline-form {
display: flex;
flex-direction: column;
gap: 3px;
margin-top: 6px;
padding: 10px 12px;
border: 1px solid var(--border);
border-radius: 8px;
background: var(--control-bg);
}
:is(#appSettingsModal, #sessionOptionsModal, #createCaseModal) .set-inline-form h5 {
margin: 0 0 4px;
font-size: 0.72rem;
color: var(--text-muted);
}
/* ── rows ─────────────────────────────────────────────────────────────── */
:is(#appSettingsModal, #sessionOptionsModal, #createCaseModal) .set-row {
display: flex;
+162 -27
View File
@@ -364,12 +364,27 @@ Object.assign(CodemanApp.prototype, {
// this handler before its own cancel()), so preventDefault is explicit:
// without it the browser runs its native copy on top of ours.
if (this.shouldCopyTerminalSelectionFromShortcut?.(ev)) {
const selection = this.terminal.hasSelection?.() ? this.terminal.getSelection() : '';
if (selection) {
// The CLEANED selection decides, not the raw one. A drag across the blank
// part of a row selects real padding spaces, which are truthy, so testing
// the raw text would spend this press on a copy of nothing and make the
// user press again to interrupt.
const selection = this.cleanedTerminalSelection();
if (selection.trim()) {
ev.preventDefault();
void this.copyTerminalSelection(selection);
return false;
}
// Nothing worth copying. The clear is for feedback, not for the
// interrupt: the gate above tests the CLEANED selection, so a
// padding-only selection left set cleans to '' on every later press and
// falls through to the PTY anyway. What it buys is that a highlight
// which copies nothing does not linger with no explanation, which is
// also what the toast is for. Falls through exactly as an empty
// selection does, so this press still reaches the PTY as 0x03.
if (this.terminal?.hasSelection?.()) {
this.terminal.clearSelection?.();
this.showToast('Nothing to copy', 'warning');
}
if (ev.shiftKey) {
ev.preventDefault();
return false;
@@ -3159,6 +3174,38 @@ Object.assign(CodemanApp.prototype, {
return buffer.viewportY >= buffer.baseY - 2;
},
/**
* Re-take the sticky-scroll baseline from where the viewport now sits.
*
* `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` before it queues
* data, and `flushPendingWrites` scrolls to the bottom off that sample. A
* buffer load that replays its queue samples at the worst possible moment:
* `_finishBufferLoad` runs inside `chunkedTerminalWrite`, before its promise
* resolves, with the terminal freshly reset and rewritten, so the sample is
* always true. A caller that then restores the reader's position would have
* that restore undone by the next flush.
*
* `_onSessionNeedsRefresh` and `_maybeRefetchFullHistory` restore a position
* and both call this, so their baseline describes the position they chose.
*
* The other two load paths do not call it, for different reasons.
* `_onSessionClearTerminal` resets and rewrites with no scroll afterwards,
* so the sampled true is already the truth there. `selectSession` does NOT
* end at the bottom, whatever its `scrollToBottom()` after the write
* suggests: it ends at `scrollToLastNonEmptyLine()`, which targets
* `lastNonEmptyLine - rows + 2` and therefore parks ABOVE `baseY` whenever
* the replayed frame keeps trailing blank rows, which a full capture does on
* purpose. Its baseline is a stale true. What decides whether that matters
* is the sticky snap in `flushPendingWrites`, and since de864e7d that snap
* fires only when the flush found the viewport already at the bottom
* (`preserveViewportY === null`), which a parked selectSession viewport is
* not. Do not read the absent call here as a claim that selectSession lands
* at the bottom.
*/
_syncStickyScrollBaseline() {
this._wasAtBottomBeforeWrite = this.isTerminalAtBottom();
},
// Record manual scroll gestures so sticky-scroll can give an upward scroll a
// short grace window (see _hasRecentUserScrollUp). A downward scroll that
// lands back at the bottom clears the suppression immediately.
@@ -3295,7 +3342,11 @@ Object.assign(CodemanApp.prototype, {
// to prevent interleaving historical buffer data with live SSE data.
// This is critical: interleaving causes cursor position chaos with Ink redraws.
if (this._isLoadingBuffer) {
if (this._loadBufferQueue) this._loadBufferQueue.push(data);
// Each entry records when it arrived. A flush of a tmux-capture load
// replays only what arrived after the capture; without the timestamp it
// would have to replay the whole queue, duplicating the events the
// capture already contains. See _finishBufferLoad's `since`.
if (this._loadBufferQueue) this._loadBufferQueue.push({ at: performance.now(), data });
return;
}
@@ -3809,9 +3860,14 @@ Object.assign(CodemanApp.prototype, {
* and a tick-Worker so progress continues on occluded / idle-throttled tabs.
* @param {string} buffer - The full terminal buffer to write
* @param {number} chunkSize - Size of each chunk (default 32KB)
* @param {string} [loadOwner] - Load token to finish under
* @param {{ flushQueued?: boolean, since?: number }} [finishOpts] - Passed to
* `_finishBufferLoad`. This method ends the load for every non-empty buffer,
* so a caller that wants the queue replayed has to say so HERE; the call in
* `selectSession` only runs when the write was skipped entirely.
* @returns {Promise<{parsedAt: number, bufferLength: number, completed: boolean}>} Parse marker snapshot
*/
chunkedTerminalWrite(buffer, chunkSize = TERMINAL_CHUNK_SIZE, loadOwner) {
chunkedTerminalWrite(buffer, chunkSize = TERMINAL_CHUNK_SIZE, loadOwner, finishOpts) {
// Generation counter: if a newer chunkedTerminalWrite starts (tab switch),
// older writes abort instead of continuing to push stale data into the terminal.
const writeGen = ++this._chunkedWriteGen;
@@ -3824,7 +3880,7 @@ Object.assign(CodemanApp.prototype, {
completed,
});
if (!buffer || buffer.length === 0) {
this._finishBufferLoad(bufferLoadOwner);
this._finishBufferLoad(bufferLoadOwner, finishOpts);
resolve(parseSnapshot());
return;
}
@@ -3838,7 +3894,7 @@ Object.assign(CodemanApp.prototype, {
this.terminal.write(cleanBuffer, () => resolve(parseSnapshot()));
// The write is now ordered in xterm's queue. Release live output before
// parsing completes; subsequent writes stay behind it without being lost.
this._finishBufferLoad(bufferLoadOwner);
this._finishBufferLoad(bufferLoadOwner, finishOpts);
return;
}
@@ -3869,7 +3925,7 @@ Object.assign(CodemanApp.prototype, {
);
resolve(result);
});
this._finishBufferLoad(bufferLoadOwner);
this._finishBufferLoad(bufferLoadOwner, finishOpts);
return;
}
@@ -3883,15 +3939,49 @@ Object.assign(CodemanApp.prototype, {
});
},
/**
* Open a buffer load: live terminal events are queued from here until
* `_finishBufferLoad` decides what to do with them. Returns the load token the
* finish call must present; a stale token makes that call a no-op.
*
* @param {string} [owner] Reuse an existing token to re-enter the same load
* (see below); omit it to start a new one.
* @returns {string} The load token.
*/
_beginBufferLoad(owner) {
if (this._bufferLoadSeq === undefined) this._bufferLoadSeq = 0;
const loadOwner = owner === undefined ? `buffer-${++this._bufferLoadSeq}` : owner;
// `selectSession` opens the load before its fetch, and `chunkedTerminalWrite`
// opens it again under the SAME owner when it starts writing. Resetting the
// queue on that second call would throw away everything that arrived during
// the fetch, which on the capture path is output no buffer holds. Re-entering
// one load keeps its queue; a genuinely new load still starts empty.
const reentering = this._bufferLoadOwner === loadOwner && Array.isArray(this._loadBufferQueue);
this._bufferLoadOwner = loadOwner;
this._isLoadingBuffer = true;
if (!reentering) this._loadBufferQueue = [];
return loadOwner;
},
/**
* Complete a buffer load: unblock live SSE writes.
* Called when chunkedTerminalWrite finishes (or is skipped for empty buffers).
*
* By default queued SSE events are DISCARDED, not flushed. For an established
* session the loaded buffer from the API is the source of truth up to the
* response timestamp; SSE events queued during the fetch+write overlap already
* appear in that buffer, so flushing them writes duplicate data (especially Ink
* cursor-up redraws), corrupting the terminal display.
* session whose buffer came from the server's accumulated byte history, that
* history is the source of truth up to the response timestamp; SSE events
* queued during the fetch+write overlap already appear in it, so flushing
* them writes duplicate data (especially Ink cursor-up redraws), corrupting
* the terminal display.
*
* A tmux PANE CAPTURE is the exception, and the reason `since` exists. A
* capture is a point-in-time frame taken part-way through the fetch, so it is
* the source of truth only up to CAPTURE time — not up to the response. Every
* event that arrives between the capture and the end of the chunked write is
* queued and, under a plain discard, lost outright: nothing re-fetches, and
* the CLI's next partial redraw lands on a frame the terminal never received.
* The caller passes the response's own arrival time as `since` so exactly
* that tail is replayed and the pre-capture events stay dropped.
*
* COD-144: a brand-new session is the exception. Its terminal fetch can resolve
* BEFORE the PTY emits its first prompt, so the fetched buffer is empty and the
@@ -3905,17 +3995,10 @@ Object.assign(CodemanApp.prototype, {
* After unblocking, new SSE/WS events deliver subsequent output normally.
*
* @param {string} [owner] Load token from `_beginBufferLoad`; a stale owner is a no-op.
* @param {{ flushQueued?: boolean }} [opts] When `flushQueued` is true, replay any queued events.
* @param {{ flushQueued?: boolean, since?: number }} [opts] When `flushQueued`
* is true, replay queued events whose arrival timestamp is at or after
* `since` (default 0, meaning the whole queue).
*/
_beginBufferLoad(owner) {
if (this._bufferLoadSeq === undefined) this._bufferLoadSeq = 0;
const loadOwner = owner === undefined ? `buffer-${++this._bufferLoadSeq}` : owner;
this._bufferLoadOwner = loadOwner;
this._isLoadingBuffer = true;
this._loadBufferQueue = [];
return loadOwner;
},
_finishBufferLoad(owner, opts) {
if (owner !== undefined && this._bufferLoadOwner !== owner) {
return false;
@@ -3926,9 +4009,13 @@ Object.assign(CodemanApp.prototype, {
this._bufferLoadOwner = null;
// COD-144: replay (rather than discard) queued live events when the load
// painted nothing — the queued prompt is the only content a new session has.
// A tmux-capture load replays too, but only the tail: `since` cuts the queue
// at the moment the capture stopped being able to contain what arrived.
if (opts?.flushQueued && queued && queued.length) {
for (const data of queued) {
this.batchTerminalWrite(data);
const since = typeof opts.since === 'number' ? opts.since : 0;
for (const entry of queued) {
if (entry.at < since) continue;
this.batchTerminalWrite(entry.data);
}
}
return true;
@@ -4069,12 +4156,54 @@ Object.assign(CodemanApp.prototype, {
return !ev.altKey && (ev.key || '').toLowerCase() === 'c';
},
/**
* xterm's current selection, cleaned for the clipboard. The transform itself
* is CodemanCopySelection.clean in constants.js, beside decideAutoCopy; this
* is the half that needs the live terminal.
*
* `text` is for the callers that already read the selection to decide whether
* to copy at all (the Ctrl+C gate and the right-click handler), so the read is
* not repeated. The transform is idempotent on xterm output, so an
* already-cleaned string is an acceptable argument: a CR is consumed by the
* parser as a cursor move and never stored in a cell, so the only \r the
* selection can carry is the Windows line join, and that is what makes the
* trailing scan a fixed point. Fuzzed over 300 000 realistic selections.
*
* A COLUMN selection comes back untouched. Alt+drag makes one (xterm's
* shouldColumnSelect keys on altKey alone, and Codeman sets neither of the
* terminals it creates with the one option that would disable it), and a
* rectangle's whole point is that its rows line up, which trimming each row
* to its own last glyph would destroy. xterm exposes the mode nowhere public,
* so this reads the private field the way this file already reads
* terminal._core for cell dimensions, and falls back to cleaning normally if
* a future xterm renames it. SelectionMode.COLUMN is 3.
*/
cleanedTerminalSelection(text) {
const raw = text ?? (this.terminal?.hasSelection?.() ? this.terminal.getSelection() : '');
if (!raw) return '';
if (this.terminal?._core?._selectionService?._activeSelectionMode === 3) return raw;
const clean = window.CodemanCopySelection?.clean;
if (!clean) return raw;
return clean(raw);
},
// Copy the current terminal selection. Goes through _copyText (Clipboard API,
// then a hidden-textarea + execCommand fallback) because install.sh's LAN
// option serves plain HTTP, where navigator.clipboard is undefined.
async copyTerminalSelection(text) {
const selection = text ?? (this.terminal.hasSelection?.() ? this.terminal.getSelection() : '');
if (!selection) return false;
const selection = this.cleanedTerminalSelection(text);
// trim(), not emptiness: a multi-row drag across padding cleans to newlines
// alone, which are truthy, and a bare newline pasted into a chat composer
// or a shell submits the line. decideAutoCopy applies the same rule.
if (!selection.trim()) {
// Clearing is feedback, not protection. The Ctrl+C gate tests the CLEANED
// selection, so a padding-only selection left set can no longer swallow a
// later interrupt; it cleans to '' and the press reaches the PTY. What the
// clear avoids is a highlight that sits there having copied nothing.
this.terminal?.clearSelection?.();
this.showToast('Nothing to copy', 'warning');
return false;
}
const ok = await this._copyText(selection);
if (ok) {
// Clearing is what makes a second Ctrl+C an interrupt (and xterm already
@@ -4123,9 +4252,15 @@ Object.assign(CodemanApp.prototype, {
async _flushAutoCopySelection() {
const decide = window.CodemanAutoCopy?.decide;
if (!decide || !this.terminal) return;
const text = this.terminal.hasSelection?.() ? this.terminal.getSelection() : '';
// The toggle is read FIRST because Auto Copy is off by default: reading and
// cleaning a selection that can run to the 50 000-row scrollback ceiling
// costs real time on a phone, and every mouseup would pay it for nothing.
// Cleaning before decide() then means its dedupe and size cap both measure
// the text that actually reaches the clipboard, not the padded rows behind.
const enabled = this._autoCopySelectionEnabled();
const text = enabled ? this.cleanedTerminalSelection() : '';
const verdict = decide({
enabled: this._autoCopySelectionEnabled(),
enabled,
text,
lastCopied: this._autoCopyLastText,
pending: !!this._autoCopyPending,
+718 -17
View File
@@ -16,16 +16,52 @@
import type { FastifyInstance, FastifyRequest } from 'fastify';
import { ApiErrorCode, createErrorResponse, type ApiResponse } from '../../types.js';
import { isAdmin, parseBody } from '../route-helpers.js';
import { isAdmin, parseBody, readJsonConfig, SETTINGS_PATH } from '../route-helpers.js';
import { isMultiUserMode } from '../../config/multiuser.js';
import { getDataDir } from '../../config/instance.js';
import { isBlockedWebviewUrl } from '../webview-egress-policy.js';
import { egressBlockedReason, webviewFetch } from '../webview-egress.js';
import { CustomModelHostSchema } from '../schemas.js';
import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } from '../../custom-model-hosts.js';
import type { CliEntry } from '../../config/cli-registry/types.js';
const CODEMAN_CONFIG_DIR = getDataDir();
const DISCOVER_TIMEOUT_MS = 8000;
const PROPS_TIMEOUT_MS = 5000;
/**
* Claude Code's own system prompt + tool schemas cost roughly this many tokens on EVERY
* request, before a single character of conversation history — confirmed live, twice, on
* requests reporting `in:0 out:0` (the very first exchange) failing at ~36.4K tokens. No
* `CLAUDE_CODE_MAX_CONTEXT_TOKENS` value fixes this: that setting only changes when Claude
* Code decides to COMPACT conversation history, and there is no history yet on the first
* message for it to trim. A model whose real context is below this floor will refuse
* Claude Code's very first message outright, unconditionally.
*
* Set well above the ~36.4K actually measured — CLAUDE.md size, active MCP servers, and
* enabled skills all add to a project's real baseline, so the observed figure is a floor
* for THAT one workspace, not a ceiling for every one. Erring conservative here means a
* borderline-safe model still gets warned about (the user can launch anyway), rather than
* this floor missing a genuinely-too-small one because a smaller test project happened to
* fit.
*/
export const CLAUDE_MIN_SAFE_CONTEXT_TOKENS = 40000;
/**
* True when applying this model to this CLI is heading for a guaranteed first-message
* failure per `CLAUDE_MIN_SAFE_CONTEXT_TOKENS` above. Gated on `contextLengthVar` (today,
* only claude's registry entry declares one) rather than a hardcoded mode check: a CLI
* with a small enough baseline of its own to never trip this would have no reason to
* declare the field in the first place, so the check simply never applies to it.
*/
export function exceedsSafeContextFloor(
entry: Pick<CliEntry, 'capabilities'>,
contextLength: number | undefined
): boolean {
const cap = entry.capabilities.customModelInjection;
if (cap.kind !== 'env' || !cap.contextLengthVar) return false;
return typeof contextLength === 'number' && contextLength < CLAUDE_MIN_SAFE_CONTEXT_TOKENS;
}
function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown }): ApiResponse<never> | null {
if (!isMultiUserMode() || isAdmin(req)) return null;
@@ -33,7 +69,58 @@ function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown })
return createErrorResponse(ApiErrorCode.FORBIDDEN, 'Admin only in multi-user mode');
}
async function discoverModels(host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' | 'authStyle'>): Promise<string[]> {
/**
* `defaultModelId` names the model the Run-menu picker applies for this endpoint with
* no further choice, so it must actually be one of the discovered `models` — a schema
* `.refine()` can't see across the two fields the way this can, and would also run on
* every unrelated field edit rather than only when either of these two changes.
*/
function invalidDefaultModel(host: Pick<CustomModelHost, 'defaultModelId' | 'models'>): ApiResponse<never> | null {
if (host.defaultModelId === undefined) return null;
if ((host.models ?? []).includes(host.defaultModelId)) return null;
return createErrorResponse(
ApiErrorCode.INVALID_INPUT,
'defaultModelId must be one of the endpoint’s discovered models'
);
}
/**
* Never hand the stored credential back to the browser, on GET, POST or PUT
* alike — the file is written 0600 precisely because it holds one. `apiKeySet`
* is what lets the editor say "unchanged if left blank" without the client
* ever holding the real value: `applyStoredApiKey()` below is the other half,
* treating an absent key on PUT as "keep the stored one" rather than clearing
* it, which is what makes never returning it survivable for the edit flow.
*/
function redactApiKey(host: CustomModelHost): Omit<CustomModelHost, 'apiKey'> & { apiKeySet: boolean } {
const { apiKey, ...rest } = host;
return { ...rest, apiKeySet: !!apiKey };
}
/**
* A PUT body with no `apiKey` (or a blank one) means "leave it alone", never
* "clear it": the editor never receives the real value to resend deliberately
* unchanged (see redactApiKey), so the only way it can tell the two apart is
* by omission. There is deliberately no way to CLEAR a key back to unset this
* way — a pre-existing limitation, not something this changes.
*/
function applyStoredApiKey(incoming: CustomModelHost, existing: CustomModelHost): CustomModelHost {
return incoming.apiKey ? incoming : { ...incoming, apiKey: existing.apiKey };
}
/**
* `modelContextLengths`/`modelSizesGB` are server-populated by discovery, never
* user-entered, and PUT replaces the whole record — so merge them back in from the
* stored host rather than trust whatever the editor's body carried (or omitted).
* The editor only ever sends `models`/`lastDiscoveredAt` verbatim from its cached
* copy; requiring it to also round-trip these two is exactly the kind of thing a
* future caller forgets, same class of bug `applyStoredApiKey` exists to prevent.
*/
function applyDiscoveredFields(incoming: CustomModelHost, existing: CustomModelHost): CustomModelHost {
return { ...incoming, modelContextLengths: existing.modelContextLengths, modelSizesGB: existing.modelSizesGB };
}
function authHeaders(host: Pick<CustomModelHost, 'apiKey' | 'authStyle'>): Record<string, string> {
const headers: Record<string, string> = {};
const apiKey = host.apiKey?.trim();
// Exactly ONE header, never both — see custom-model-hosts.ts's CustomModelAuthStyle
@@ -41,14 +128,137 @@ async function discoverModels(host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' |
const style = host.authStyle ?? 'bearer';
if (apiKey && style === 'bearer') headers.Authorization = `Bearer ${apiKey}`;
if (apiKey && style === 'api-key') headers['api-key'] = apiKey;
return headers;
}
export interface DiscoveryResult {
models: string[];
/** See `CustomModelHost.modelContextLengths` — only ever populated for models already loaded. */
contextLengths: Record<string, number>;
/** See `CustomModelHost.modelSizesGB` — populated for every model whose own listing states one. */
sizesGB: Record<string, number>;
}
/**
* Best-effort: pulls a file size in GB out of a model's own `description`, when the
* server states one. llama-swap writes `"Auto-discovered 16.35 GB - parameters
* auto-fitted by llama.cpp"` for a model it found on disk itself; a hand-configured
* profile's own description (e.g. `"General-purpose reasoning model, MoE CPU-offloaded."`)
* has no such figure and correctly yields no estimate rather than a guess — there is no
* separate "give me the file size" endpoint to fall back on.
*/
function parseSizeGB(description: unknown): number | undefined {
if (typeof description !== 'string') return undefined;
const match = /(\d+(?:\.\d+)?)\s*GB\b/i.exec(description);
if (!match) return undefined;
const size = Number(match[1]);
return Number.isFinite(size) && size > 0 ? size : undefined;
}
/**
* Best-effort: fetches `GET /props?model=<id>` (llama.cpp-native, llama-swap-proxied) for
* ONE already-loaded model and pulls its real `n_ctx` out. Never called for a model that
* isn't already loaded — see the caller and `CustomModelHost.modelContextLengths` for why
* that's a hard safety requirement, not just a nicety: llama-swap treats this endpoint's
* `?model=` as a routing hint, and asking it about an unloaded model risks triggering an
* actual (slow, GPU-swapping) load as a side effect of what should be read-only discovery.
* Any failure (unreachable, non-2xx, missing/malformed field) is swallowed — one model's
* context length is a nice-to-have, never worth failing the whole discovery pass over.
*
* ⚠️ FALLBACK ONLY — confirmed live to be actively WRONG for a `--fit-ctx`-launched llama-
* swap backend: `/props`'s `n_ctx` read 154112 for a model llama-swap itself had launched
* with `--fit-ctx 16384` (visible in `/running`'s own `cmd`), and the real server then
* refused a request at the real 16384-token limit — `n_ctx` here appears to report the
* model's theoretical/trained maximum, not the runtime-configured one. `parseCtxFromCmd`
* (below), which reads the actual launch flag `/running` reports, is the primary source;
* this is only used when that parse comes up empty (no recognized flag in `cmd`, or `cmd`
* itself unavailable).
*/
async function fetchContextLength(
host: Pick<CustomModelHost, 'baseUrl'>,
modelId: string,
headers: Record<string, string>
): Promise<number | undefined> {
try {
const url = new URL(`${host.baseUrl.replace(/\/+$/, '')}/props`);
url.searchParams.set('model', modelId);
const res = await webviewFetch(url, { headers, signal: AbortSignal.timeout(PROPS_TIMEOUT_MS) });
if (!res.ok) return undefined;
const body = (await res.json()) as { n_ctx?: unknown; default_generation_settings?: { n_ctx?: unknown } };
const nCtx = body.n_ctx ?? body.default_generation_settings?.n_ctx;
return typeof nCtx === 'number' && Number.isFinite(nCtx) && nCtx > 0 ? nCtx : undefined;
} catch {
return undefined;
}
}
/**
* Parses the REAL configured context size out of llama-swap's own launch command for a
* model (`/running`'s `cmd` field, e.g. `"llama-server -m ... --fit-ctx 16384 ..."`) —
* the primary source for `modelContextLengths`, preferred over `/props`'s `n_ctx` (see
* `fetchContextLength`'s own doc comment for why that field is unreliable here). Checks
* `--fit-ctx` first (llama-swap's own auto-fit flag), then the plain llama.cpp
* `-c`/`--ctx-size`/`--ctx_size` flags a hand-written launch command might use instead.
* Returns `undefined` when `cmd` has none of these — not every launch command needs to
* state one explicitly (llama.cpp has its own default), and guessing one would be worse
* than the "no override applied" the caller already treats an unknown length as.
*/
function parseCtxFromCmd(cmd: unknown): number | undefined {
if (typeof cmd !== 'string') return undefined;
const match = /--fit-ctx\s+(\d+)/.exec(cmd) ?? /(?:^|\s)(?:-c|--ctx-size|--ctx_size)\s+(\d+)/.exec(cmd);
if (!match) return undefined;
const value = Number(match[1]);
return Number.isFinite(value) && value > 0 ? value : undefined;
}
async function discoverModels(
host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' | 'authStyle'>
): Promise<DiscoveryResult> {
const headers = authHeaders(host);
const res = await webviewFetch(new URL(`${host.baseUrl.replace(/\/+$/, '')}/v1/models`), {
headers,
signal: AbortSignal.timeout(DISCOVER_TIMEOUT_MS),
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = (await res.json()) as { data?: Array<{ id?: unknown }> };
return (body.data ?? []).map((m) => m.id).filter((id): id is string => typeof id === 'string' && id.length > 0);
const body = (await res.json()) as {
data?: Array<{ id?: unknown; status?: { value?: unknown }; description?: unknown }>;
};
const entries = body.data ?? [];
const models = entries.map((m) => m.id).filter((id): id is string => typeof id === 'string' && id.length > 0);
const sizesGB: Record<string, number> = {};
for (const entry of entries) {
if (typeof entry.id !== 'string' || !entry.id) continue;
const size = parseSizeGB(entry.description);
if (size !== undefined) sizesGB[entry.id] = size;
}
// llama-swap-specific, feature-detected: a server that never mentions `status` on ANY
// entry gets no context-length enrichment at all, rather than treating "no status field"
// as "assume unloaded" — either reading is a guess, and skipping is the safe one, since
// fetchContextLength must only ever run against a model this server itself calls loaded.
const hasStatusField = entries.some((m) => m && typeof m === 'object' && 'status' in m);
const contextLengths: Record<string, number> = {};
if (hasStatusField) {
const loadedIds = entries
.filter((m) => m.status && typeof m.status === 'object' && (m.status as { value?: unknown }).value === 'loaded')
.map((m) => m.id)
.filter((id): id is string => typeof id === 'string' && id.length > 0);
if (loadedIds.length > 0) {
// Primary source: the REAL launch command (see parseCtxFromCmd's own doc comment
// for why /props's n_ctx cannot be trusted here). One /running call covers every
// loaded model, so this never costs more requests than the old /props-only path did
// when the cmd parse succeeds, and exactly one extra when it has to fall back.
const swapStatus = await getLlamaSwapStatus(host);
const cmdById = new Map(swapStatus.running.map((r) => [r.model, r.cmd]));
for (const id of loadedIds) {
const fromCmd = parseCtxFromCmd(cmdById.get(id));
const ctx = fromCmd ?? (await fetchContextLength(host, id, headers));
if (ctx !== undefined) contextLengths[id] = ctx;
}
}
}
return { models, contextLengths, sizesGB };
}
/**
@@ -66,41 +276,498 @@ function describeFetchError(err: unknown): string {
return message;
}
export function registerCustomModelRoutes(app: FastifyInstance): void {
app.get('/api/model-endpoints', async (req) =>
isMultiUserMode() && !isAdmin(req) ? [] : readCustomModelHosts(CODEMAN_CONFIG_DIR)
);
type RedactedHost = ReturnType<typeof redactApiKey>;
app.post('/api/model-endpoints', async (req, reply): Promise<ApiResponse<{ host: CustomModelHost }>> => {
/**
* Merges a fresh `GET /v1/models` result into a host record: stamps
* `lastDiscoveredAt`, and drops `defaultModelId` if it no longer appears in
* the fresh list (it would otherwise leave the Run-menu picker applying a
* model id the endpoint just told us it doesn't serve). Pure — no IO, so the
* manual route (which reports a fetch failure's *reason* to the caller) and
* the periodic sweep below (which only cares whether it can move on) can
* each do their own `discoverModels()` + error handling around one shared
* "how to apply a successful result" step.
*/
const RUNNING_TIMEOUT_MS = 5000;
export interface LlamaSwapRunningModel {
model: string;
state: string;
/** The actual launch command llama-swap started this backend with, when it says one —
* see `parseCtxFromCmd`, which reads the real configured context size out of this. */
cmd?: string;
}
export interface LlamaSwapStatus {
/**
* Feature-detected via `GET /running`: true only when the server answered with
* llama-swap's own shape (`{ running: [...] }`). Plain llama.cpp (and any other
* OpenAI-compatible server) has no such endpoint and always runs the single model
* it was started with, so there is no "current model" to conflict with — every
* caller must treat `isLlamaSwap: false` as "nothing to check", never as an error.
*/
isLlamaSwap: boolean;
running: LlamaSwapRunningModel[];
}
/**
* Distinguishes llama-swap from a plain llama.cpp/OpenAI-compatible server, and reports
* what llama-swap currently has loaded — llama.cpp only ever runs one GGUF at a time, and
* llama-swap unloads/reloads it on demand when a request asks for a different one, which
* can take anywhere from a few seconds to over a minute. Read-only: this never triggers a
* swap itself (unlike `/props?model=`, `/running` takes no `model` parameter to route by).
* Best-effort like `discoverModels()`'s siblings: any failure (unreachable, non-2xx,
* unexpected shape) reads as "not llama-swap", never thrown.
*/
export async function getLlamaSwapStatus(
host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' | 'authStyle'>
): Promise<LlamaSwapStatus> {
try {
const res = await webviewFetch(new URL(`${host.baseUrl.replace(/\/+$/, '')}/running`), {
headers: authHeaders(host),
signal: AbortSignal.timeout(RUNNING_TIMEOUT_MS),
});
if (!res.ok) return { isLlamaSwap: false, running: [] };
const body = (await res.json()) as { running?: unknown };
if (!Array.isArray(body.running)) return { isLlamaSwap: false, running: [] };
const running = body.running
.filter(
(r): r is { model: string; state?: unknown; cmd?: unknown } =>
!!r && typeof r === 'object' && typeof (r as { model?: unknown }).model === 'string'
)
.map((r) => ({
model: r.model,
state: typeof r.state === 'string' ? r.state : 'unknown',
cmd: typeof r.cmd === 'string' ? r.cmd : undefined,
}));
return { isLlamaSwap: true, running };
} catch {
return { isLlamaSwap: false, running: [] };
}
}
interface LlamaSwapLogTail {
latestLine?: string;
lastAccessedAt: number;
controller: AbortController;
}
/** One open `/api/events` tail per endpoint, keyed by host id — see `getLatestLlamaSwapLogLine`. */
const llamaSwapLogTails = new Map<string, LlamaSwapLogTail>();
/** A tail nothing has asked about in this long is closed by the next `pruneIdleLlamaSwapLogTails` sweep. */
const LOG_TAIL_IDLE_MS = 30_000;
/**
* Parses one `data: {...}` payload from llama-swap's `GET /api/events` SSE stream and
* returns the backend (never llama-swap's own proxy) log text it carries, or `undefined`
* for anything else (a different event `type`, a malformed frame, a proxy-sourced one).
*
* The real shape, confirmed live against a real llama-swap deployment — NOT documented
* anywhere the plan doc's original research found, and genuinely surprising the first
* time around: `GET /logs` (the endpoint that name suggests, and this feature's own
* first cut was built against) turns out to carry ONLY llama-swap's own proxy
* request-access log — it never once showed a single backend line even seconds after a
* real, confirmed model swap. The backend llama-server process's actual stdout
* (`load_model: ...`, `llama_server: model loaded`) only ever showed up in `/api/events`,
* as `{"type":"logData","data":"<JSON-string>"}` whose OWN `data` field parses to a
* second object, `{"data": "<newline-joined log text>", "source": "proxy" | "upstream"}`
* — `source` is the exact, explicit distinguisher (`upstream` = the backend process,
* `proxy` = llama-swap's own line), not a guessed regex against the text itself.
*/
function parseBackendLogDataEvent(dataLine: string): string | undefined {
let outer: unknown;
try {
outer = JSON.parse(dataLine);
} catch {
return undefined;
}
if (
!outer ||
typeof outer !== 'object' ||
(outer as { type?: unknown }).type !== 'logData' ||
typeof (outer as { data?: unknown }).data !== 'string'
) {
return undefined;
}
let inner: unknown;
try {
inner = JSON.parse((outer as { data: string }).data);
} catch {
return undefined;
}
if (
!inner ||
typeof inner !== 'object' ||
(inner as { source?: unknown }).source !== 'upstream' ||
typeof (inner as { data?: unknown }).data !== 'string'
) {
return undefined;
}
return (inner as { data: string }).data;
}
/**
* Reads `GET /api/events` forever (until `entry.controller` aborts it), updating
* `entry.latestLine` with the most recent BACKEND log line seen (see
* `parseBackendLogDataEvent`). Fire-and-forget: the caller never awaits this — it runs
* for the tail's whole lifetime in the background, and `getLatestLlamaSwapLogLine` just
* reads whatever `entry.latestLine` currently holds. SSE frames are separated by a blank
* line (`\n\n`), buffered the same way `/running`'s NDJSON-shaped siblings buffer partial
* chunks — a frame split across two `reader.read()` calls must not be parsed early.
*/
/**
* Cap on the unparsed remainder held between reads of the backend log stream. One
* SSE frame is a status line, so this is orders of magnitude more than a real frame
* needs; it exists so a server that never emits a frame boundary cannot grow the
* buffer without bound for the life of the connection.
*/
const MAX_LOG_TAIL_BUFFER_CHARS = 64 * 1024;
async function pumpLlamaSwapLogTail(
host: Pick<CustomModelHost, 'id' | 'baseUrl' | 'apiKey' | 'authStyle'>,
entry: LlamaSwapLogTail
): Promise<void> {
try {
const res = await webviewFetch(new URL(`${host.baseUrl.replace(/\/+$/, '')}/api/events`), {
headers: authHeaders(host),
signal: entry.controller.signal,
});
if (!res.ok || !res.body) return;
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
for (;;) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const frames = buffer.split('\n\n');
buffer = frames.pop() ?? '';
// The remainder only shrinks at a frame boundary, so a server that streams
// without `\n\n` (or one very long frame) would grow it for as long as the
// connection is held, which is indefinitely by design. Past the cap the
// partial frame cannot become a useful log line anyway, so drop it and
// resynchronise on the next boundary rather than buffering forever.
if (buffer.length > MAX_LOG_TAIL_BUFFER_CHARS) buffer = '';
for (const frame of frames) {
const dataLine = frame.split('\n').find((l) => l.startsWith('data:'));
if (!dataLine) continue;
const backendText = parseBackendLogDataEvent(dataLine.slice('data:'.length));
if (!backendText) continue;
const lines = backendText.split('\n').filter((l) => l.trim());
if (lines.length > 0) entry.latestLine = lines[lines.length - 1]!.trim();
}
}
} catch {
// connection dropped / aborted / endpoint unreachable — a future access starts fresh
} finally {
// Delete by IDENTITY, not just by key: an aborted pump can finish after a NEWER
// entry was already created for the same endpoint id (e.g. abort-then-immediately-
// re-request), and deleting unconditionally would remove that newer entry and orphan
// its connection — nothing would ever prune it, since pruneIdleLlamaSwapLogTails only
// walks entries still present in the map.
if (llamaSwapLogTails.get(host.id) === entry) {
llamaSwapLogTails.delete(host.id);
}
}
}
/**
* Real-time "what is llama.cpp actually doing right now" for the loading banner
* (docs/custom-model-endpoints-plan.md): llama-swap's `GET /api/events` SSE stream
* carries the backend llama-server process's own stdout — `load_model: loading model
* '<path>'`, `load_model: initializing, n_slots = N, n_ctx_slot = N`, `llama_server:
* model loaded`, etc — tagged `source: "upstream"`, distinct from llama-swap's own
* `source: "proxy"` request-access lines (see `parseBackendLogDataEvent`). Confirmed
* live against a real llama-swap deployment, including through an actual forced model
* swap end-to-end.
*
* Held OPEN per endpoint rather than re-opened on every 1s poll — confirmed live to stay
* open indefinitely (read past 220KB over 8 seconds with no `done`), unlike `/logs`
* (see `parseBackendLogDataEvent`'s doc comment), so reconnecting each poll would be
* pure waste. One connection is reused across every session currently watching a load on
* that endpoint; since llama.cpp/llama-swap only ever runs one model at a time, a line
* seen while a load is in flight is safe to attribute to that load (a deployment that
* could load several models concurrently would need a per-model tag this format doesn't
* provide).
*
* Lazily started on first access and idle-closed rather than left open forever — see
* `pruneIdleLlamaSwapLogTails`.
*/
export function getLatestLlamaSwapLogLine(
host: Pick<CustomModelHost, 'id' | 'baseUrl' | 'apiKey' | 'authStyle'>
): string | undefined {
let entry = llamaSwapLogTails.get(host.id);
if (!entry) {
entry = { lastAccessedAt: Date.now(), controller: new AbortController() };
llamaSwapLogTails.set(host.id, entry);
void pumpLlamaSwapLogTail(host, entry);
}
entry.lastAccessedAt = Date.now();
return entry.latestLine;
}
/**
* Closes EVERY open log tail. The idle sweep above only runs on server.ts's periodic
* interval, and that interval is disposed on shutdown, so without this an outbound
* stream outlives `WebServer.stop()` against CLAUDE.md's "clear Maps in stop()" rule.
* Harmless today only because `cli.ts`'s shutdown handler reaches `process.exit(0)`,
* which is not a property to rely on: tests and any in-process restart do not.
*/
export function closeAllLlamaSwapLogTails(): void {
for (const entry of llamaSwapLogTails.values()) entry.controller.abort();
llamaSwapLogTails.clear();
}
/**
* Closes any log tail nothing has called `getLatestLlamaSwapLogLine` about in
* `LOG_TAIL_IDLE_MS` — a stream nobody is polling is an open connection with nothing to
* show for it. Called from the same periodic sweep as `detectCustomModelSwapDisplacements`
* in server.ts, not its own timer.
*/
export function pruneIdleLlamaSwapLogTails(now = Date.now()): void {
for (const [id, entry] of llamaSwapLogTails) {
if (now - entry.lastAccessedAt > LOG_TAIL_IDLE_MS) {
entry.controller.abort();
llamaSwapLogTails.delete(id);
}
}
}
/**
* Actually kicks off llama-swap's lazy model load, rather than waiting for the launched
* CLI's own first prompt to do it. llama-swap has no separate "switch model" admin
* endpoint — the ONLY thing that starts a swap is a real inference request naming the
* model (confirmed live: applying a selection alone never appeared in the llama-swap
* server's own logs; nothing had actually asked it to load anything). This sends the
* smallest real request that will — `max_tokens: 1`, one throwaway user message — to
* `${baseUrl}/v1/chat/completions`, the OpenAI-compatible endpoint every supported
* harness already points at.
*
* Deliberately fire-and-forget: the caller (the apply/create routes) returns to the
* client immediately, and the frontend's own polling (`GET .../running-status`) is what
* actually confirms readiness — this call's response is never read, just its side
* effect. No abort/timeout of its own either: a real load can take well over a minute for
* a large model, and this is a normal long-running Node process, so there is nothing to
* clean up by cutting it short. Errors are swallowed for the same reason `discoverModels`'s
* siblings swallow theirs — one endpoint's hiccup here is a nice-to-have that failed, not
* something worth surfacing as a request failure four layers up.
*/
export function triggerLlamaSwapLoad(
host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' | 'authStyle'>,
modelId: string
): void {
const url = new URL(`${host.baseUrl.replace(/\/+$/, '')}/v1/chat/completions`);
webviewFetch(url, {
method: 'POST',
headers: { ...authHeaders(host), 'content-type': 'application/json' },
body: JSON.stringify({
model: modelId,
messages: [{ role: 'user', content: 'Hi' }],
max_tokens: 1,
stream: false,
}),
}).catch(() => {
// best-effort — see the doc comment above
});
}
function applyDiscoveredModels(host: CustomModelHost, result: DiscoveryResult): CustomModelHost {
const { models, contextLengths, sizesGB } = result;
const defaultModelId = host.defaultModelId && models.includes(host.defaultModelId) ? host.defaultModelId : undefined;
// Merge onto what's already known rather than replacing: a model not probed this round
// (not currently loaded) keeps whatever context length an earlier round already learned
// for it, and one no longer in the fresh list is dropped, same reasoning as defaultModelId.
const merged = { ...host.modelContextLengths, ...contextLengths };
const kept = Object.fromEntries(Object.entries(merged).filter(([id]) => models.includes(id)));
const modelContextLengths = Object.keys(kept).length > 0 ? kept : undefined;
// sizesGB, unlike contextLengths, is populated for every model in the SAME pass (no
// loaded-only restriction — see parseSizeGB), so this is closer to a plain replace, but
// still merges onto the previous round rather than dropping a size for a model whose
// description happened to omit the figure on this particular pass.
const mergedSizes = { ...host.modelSizesGB, ...sizesGB };
const keptSizes = Object.fromEntries(Object.entries(mergedSizes).filter(([id]) => models.includes(id)));
const modelSizesGB = Object.keys(keptSizes).length > 0 ? keptSizes : undefined;
return {
...host,
models,
defaultModelId,
modelContextLengths,
modelSizesGB,
lastDiscoveredAt: new Date().toISOString(),
};
}
/**
* `customModelEndpointsEnabled` defaults OFF (unlike `showPlanUsageLimits`'s
* absent-means-on in `readPlanUsageTelemetryEnabled`), so mirror the frontend's
* own gate (`session-ui.js`'s `!settings.customModelEndpointsEnabled`) rather
* than that reader's default. Exists so the periodic re-discovery sweep in
* server.ts can skip entirely while the feature is off, instead of polling
* every saved endpoint forever regardless of the setting.
*/
export async function readCustomModelEndpointsEnabled(): Promise<boolean> {
const settings = await readJsonConfig<Record<string, unknown>>(SETTINGS_PATH, 'settings.json', {});
return settings.customModelEndpointsEnabled === true;
}
/**
* Re-discovers every saved endpoint's models, best-effort. One endpoint being
* unreachable (powered off, wrong network) must not stop the others from
* refreshing, and a read-modify-write per host (rather than one batch write
* at the end) means a crash or restart mid-sweep loses at most the endpoints
* not yet reached, never a write already applied. Exported so both the
* periodic timer (server.ts) and a test can drive it directly.
*/
export async function refreshAllCustomModelHosts(): Promise<void> {
const dataDir = getDataDir();
const hosts = await readCustomModelHosts(dataDir);
for (const host of hosts) {
if (isBlockedWebviewUrl(host.baseUrl)) continue;
let result: DiscoveryResult;
try {
result = await discoverModels(host);
} catch {
continue; // unreachable this cycle — try again next tick, not fatal to the sweep
}
// Re-read + splice by id rather than reusing the array captured above: an
// admin editing or deleting an endpoint via the API mid-sweep must win,
// not be silently overwritten by a refresh that started before their change.
const current = await readCustomModelHosts(dataDir);
const index = current.findIndex((item) => item.id === host.id);
if (index === -1) continue; // deleted mid-sweep
current[index] = applyDiscoveredModels(current[index], result);
await writeCustomModelHosts(dataDir, current);
}
}
/** The subset of `Session` this sweep needs — kept minimal so a test can pass a plain object. */
export interface CustomModelSessionLike {
id: string;
name: string;
customModel?: { endpointId: string; modelId: string; label?: string };
}
/** One session whose model was just found evicted, ready to broadcast as `CustomModelSwappedOut`. */
export interface CustomModelSwapDisplacement {
sessionId: string;
sessionName: string;
endpointId: string;
previousModel: string;
currentlyLoadedModel: string;
}
/**
* Detects when a live session's own custom-model selection is no longer the model
* llama-swap actually has loaded — evicted by ANOTHER session's activity on the same
* endpoint, since llama.cpp/llama-swap runs one model at a time (the apply/create routes'
* own swap-conflict check only ever runs at THAT session's own launch/apply moment, so it
* cannot catch a later eviction triggered by a different session's normal use — confirmed
* live: a session created while nothing else had a live conflict at that instant can still
* get silently displaced afterward). Read-only, and best-effort per endpoint exactly like
* `refreshAllCustomModelHosts`'s sibling sweep — one endpoint's hiccup here never blocks
* checking the others.
*
* `notifiedSessionIds` is the caller's own de-dupe state (`server.ts` keeps one `Set` across
* sweeps), mutated in place: a session id is added once displaced and removed again once its
* own model is loaded and ready — so a LATER, genuinely new displacement can notify again
* rather than the session staying silently un-notified forever after the first one.
*/
export async function detectCustomModelSwapDisplacements(
sessions: Iterable<CustomModelSessionLike>,
notifiedSessionIds: Set<string>
): Promise<CustomModelSwapDisplacement[]> {
const byEndpoint = new Map<string, CustomModelSessionLike[]>();
for (const session of sessions) {
if (!session.customModel) continue;
const group = byEndpoint.get(session.customModel.endpointId);
if (group) group.push(session);
else byEndpoint.set(session.customModel.endpointId, [session]);
}
if (byEndpoint.size === 0) return [];
const hosts = await readCustomModelHosts(getDataDir());
const displacements: CustomModelSwapDisplacement[] = [];
for (const [endpointId, group] of byEndpoint) {
const host = hosts.find((h) => h.id === endpointId);
if (!host) continue; // endpoint deleted since these sessions were created — nothing to check
let status: LlamaSwapStatus;
try {
status = await getLlamaSwapStatus(host);
} catch {
continue; // unreachable this cycle — try again next tick, not fatal to the sweep
}
// Not llama-swap (feature-detected) or nothing loaded at all: nothing has been evicted,
// by construction — a plain llama.cpp/OpenAI-compatible server only ever runs the one
// model it was started with, so there is no "current model" to conflict with.
if (!status.isLlamaSwap || status.running.length === 0) continue;
const currentlyLoaded = status.running.find((r) => r.state === 'ready')?.model ?? status.running[0]?.model;
if (!currentlyLoaded) continue;
for (const session of group) {
const modelId = session.customModel!.modelId;
const stillLoaded = status.running.some((r) => r.model === modelId);
if (stillLoaded) {
notifiedSessionIds.delete(session.id); // back to normal — a future eviction can notify again
continue;
}
if (notifiedSessionIds.has(session.id)) continue; // already told them once for this displacement
notifiedSessionIds.add(session.id);
displacements.push({
sessionId: session.id,
sessionName: session.name,
endpointId,
previousModel: modelId,
currentlyLoadedModel: currentlyLoaded,
});
}
}
return displacements;
}
export function registerCustomModelRoutes(app: FastifyInstance): void {
app.get('/api/model-endpoints', async (req): Promise<RedactedHost[]> => {
if (isMultiUserMode() && !isAdmin(req)) return [];
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
return hosts.map(redactApiKey);
});
app.post('/api/model-endpoints', async (req, reply): Promise<ApiResponse<{ host: RedactedHost }>> => {
const denied = adminOnly(req, reply);
if (denied) return denied;
const host = parseBody(CustomModelHostSchema, req.body);
if (isBlockedWebviewUrl(host.baseUrl)) {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
}
const badDefault = invalidDefaultModel(host);
if (badDefault) return badDefault;
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
if (hosts.some((item) => item.id === host.id)) {
return createErrorResponse(ApiErrorCode.ALREADY_EXISTS, 'Model endpoint already exists');
}
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, [...hosts, host]);
return { success: true, data: { host } };
return { success: true, data: { host: redactApiKey(host) } };
});
app.put('/api/model-endpoints/:id', async (req, reply): Promise<ApiResponse<{ host: CustomModelHost }>> => {
app.put('/api/model-endpoints/:id', async (req, reply): Promise<ApiResponse<{ host: RedactedHost }>> => {
const denied = adminOnly(req, reply);
if (denied) return denied;
const { id } = req.params as { id: string };
const host = parseBody(CustomModelHostSchema, { ...(req.body as object), id });
if (isBlockedWebviewUrl(host.baseUrl)) {
const incoming = parseBody(CustomModelHostSchema, { ...(req.body as object), id });
if (isBlockedWebviewUrl(incoming.baseUrl)) {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
}
const badDefault = invalidDefaultModel(incoming);
if (badDefault) return badDefault;
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
const index = hosts.findIndex((item) => item.id === id);
if (index === -1) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
const host = applyDiscoveredFields(applyStoredApiKey(incoming, hosts[index]), hosts[index]);
const next = [...hosts];
next[index] = host;
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next);
return { success: true, data: { host } };
return { success: true, data: { host: redactApiKey(host) } };
});
app.delete('/api/model-endpoints/:id', async (req, reply): Promise<ApiResponse<{ id: string }>> => {
@@ -129,11 +796,11 @@ export function registerCustomModelRoutes(app: FastifyInstance): void {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
}
try {
const models = await discoverModels(host);
const result = await discoverModels(host);
const next = [...hosts];
next[index] = { ...host, models, lastDiscoveredAt: new Date().toISOString() };
next[index] = applyDiscoveredModels(host, result);
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next);
return { success: true, data: { models } };
return { success: true, data: { models: result.models } };
} catch (err) {
const blocked = egressBlockedReason(err);
return createErrorResponse(
@@ -143,4 +810,38 @@ export function registerCustomModelRoutes(app: FastifyInstance): void {
}
}
);
// Read-only, no admin gate: any session owner who can already point their own session
// at this endpoint (POST .../custom-model, ungated by design — see session-routes.ts)
// can equally ask what it currently has loaded, before or while that apply is pending.
app.get(
'/api/model-endpoints/:id/running-status',
async (
req
): Promise<
ApiResponse<{
isLlamaSwap: boolean;
running: Array<Pick<LlamaSwapRunningModel, 'model' | 'state'>>;
logLine?: string;
}>
> => {
const { id } = req.params as { id: string };
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
const host = hosts.find((item) => item.id === id);
if (!host) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
if (isBlockedWebviewUrl(host.baseUrl)) {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
}
const status = await getLlamaSwapStatus(host);
// Only worth tailing /api/events once llama-swap is actually confirmed — a plain
// llama.cpp/OpenAI-compatible server has no such endpoint at all.
const logLine = status.isLlamaSwap ? getLatestLlamaSwapLogLine(host) : undefined;
// `cmd` (the literal llama-server launch line, which can carry model paths and
// --api-key) exists only so parseCtxFromCmd() can read it server-side during
// discovery — this un-gated, polled-every-second route has no reason to hand it
// to the browser, which only ever reads `model`/`state`.
const running = status.running.map(({ model, state }) => ({ model, state }));
return { success: true, data: { isLlamaSwap: status.isLlamaSwap, running, logLine } };
}
);
}
+10 -1
View File
@@ -28,4 +28,13 @@ export { registerWsRoutes } from './ws-routes.js';
export { registerVoiceRoutes } from './voice-routes.js';
export { registerWebviewRoutes, tryWebviewRefererFallback } from './webview-routes.js';
export { registerTabLayoutRoutes } from './tab-layout-routes.js';
export { registerCustomModelRoutes } from './custom-model-routes.js';
export {
registerCustomModelRoutes,
refreshAllCustomModelHosts,
readCustomModelEndpointsEnabled,
closeAllLlamaSwapLogTails,
detectCustomModelSwapDisplacements,
pruneIdleLlamaSwapLogTails,
type CustomModelSessionLike,
type CustomModelSwapDisplacement,
} from './custom-model-routes.js';
+528 -19
View File
@@ -11,7 +11,7 @@ import { homedir } from 'node:os';
import { existsSync, statSync, mkdirSync, writeFileSync } from 'node:fs';
import { execFile } from 'node:child_process';
import fs from 'node:fs/promises';
import { randomBytes } from 'node:crypto';
import { randomBytes, randomUUID } from 'node:crypto';
import { performance } from 'node:perf_hooks';
import {
ApiErrorCode,
@@ -28,8 +28,10 @@ import {
type GrokConfig,
type DeepSeekConfig,
type OmpConfig,
type RemoteHost,
} from '../../types.js';
import { Session, isAltScreenStripMode, isExternalCliMode, isMuxAltScreenOnlyStripMode } from '../../session.js';
import type { PaneCaptureOptions } from '../../mux-interface.js';
import { SseEvent } from '../sse-events.js';
import { webviewCapabilities } from '../../webview-capabilities.js';
import {
@@ -55,6 +57,12 @@ import {
} from '../schemas.js';
import { readCustomModelHosts } from '../../custom-model-hosts.js';
import { applyCustomModelInjection, removeConfigDir } from '../../custom-model-injection-apply.js';
import {
getLlamaSwapStatus,
triggerLlamaSwapLoad,
exceedsSafeContextFloor,
CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
} from './custom-model-routes.js';
import { matchesPattern } from '../../config/cli-registry/patterns.js';
import { ownerLayoutKey } from '../../tab-layout-persistence.js';
import { TabLayoutValidationError } from '../../tab-layout.js';
@@ -67,6 +75,13 @@ import {
type WaitSignal,
type SignalWaitResult,
} from '../session-wait-registry.js';
import {
RemoteWakeRegistry,
REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS,
createDefaultRemoteWakeDeps,
isProbeable,
type WakeableRemote,
} from '../../remote-wake.js';
import { clampWaitMs, MAX_BUFFER_SCAN_BYTES } from '../../config/agent-wait.js';
import {
autoConfigureRalph,
@@ -134,6 +149,7 @@ import {
checkRemoteTmuxAvailable,
readRemoteCases,
readRemoteHosts,
rehydrateRemoteHostFields,
toAttachedSessionRemote,
toSessionRemote,
} from '../../remote-hosts.js';
@@ -749,10 +765,65 @@ export function resolveOmpConfigForCreate(
return resolvedId ? { ...ompConfig, resumeSessionId: resolvedId } : ompConfig;
}
/**
* `RemoteHost` → the wake registry's host shape. They differ in one field name only
* (`id` in host config vs `hostId` on a session's `remote`), but the rename is load-
* bearing: the registry keys its per-host wake state on `hostId`. The proxy fields
* travel too: they are what tells the registry its probe cannot reach this host.
*/
function wakeableHost(host: RemoteHost): WakeableRemote {
return {
hostId: host.id,
label: host.label,
host: host.host,
port: host.port,
wakeMac: host.wakeMac,
wakeCommand: host.wakeCommand,
jumpHost: host.jumpHost,
socksProxy: host.socksProxy,
extraSshOptions: host.extraSshOptions,
};
}
export function registerSessionRoutes(
app: FastifyInstance,
ctx: SessionPort & EventPort & ConfigPort & InfraPort & AuthPort & TabLayoutPort
): void {
ctx: SessionPort & EventPort & ConfigPort & InfraPort & AuthPort & TabLayoutPort,
/** Test seam: inject a registry with fake IO instead of the real TCP/WoL probes. */
options: { remoteWake?: RemoteWakeRegistry } = {}
): RemoteWakeRegistry {
// Wake-on-LAN for sleeping remote hosts (see remote-wake.ts). One registry per
// route registration (= one web server) — the same shape as the process-wide
// `sessionWaits` singleton, but without the global.
//
// ⚠️ The ONLY caller that may wake a host is the input route below. The
// auto-reconnect watcher and boot recovery deliberately have no access to this
// registry: waking there would re-wake the host seconds after every suspend, so
// it could never stay asleep.
const remoteWake =
options.remoteWake ??
new RemoteWakeRegistry(
createDefaultRemoteWakeDeps({
noteReconnected: (sessionId, success) => {
// Duck-typed exactly like server.ts: TmuxManager owns the COD-108 backoff
// state, and the port interface does not expose it.
const mux = ctx.mux as unknown as { noteRemoteReconnect?: (id: string, ok: boolean) => void };
mux.noteRemoteReconnect?.(sessionId, success);
},
broadcast: (event, payload) => ctx.broadcast(event, payload),
log: (message) => console.log(message),
// The session's `remote` block is a launch-time snapshot, so a wake target
// configured later (banner's config dialog, or a hand-edited remote-hosts.json)
// is resolved here — throttled by the registry, and the host config is
// authoritative in BOTH directions (removing the field turns the feature off
// for a live session too).
resolveRemote: async (session) => {
const remote = session.remote;
if (!remote) return undefined;
const hosts = await readRemoteHosts(CODEMAN_CONFIG_DIR);
return rehydrateRemoteHostFields(remote, new Map(hosts.map((host) => [host.id, host])));
},
})
);
// ═══════════════════════════════════════════════════════════════
// Auth
// ═══════════════════════════════════════════════════════════════
@@ -824,9 +895,33 @@ export function registerSessionRoutes(
// creation (owned durable sessions) is handled by the dedicated case-create
// endpoint below, which #145 consolidated remote-host resolution into.
if (body.attachRemoteSession) {
// Remote hosts are admin-only infrastructure everywhere else (the list answers
// `[]` to a non-admin; write and discovery routes are `adminOnly`), and the wake
// below spawns the host's `wakeCommand` or broadcasts a packet. So the gate comes
// FIRST — before the host is even looked up — or an unprivileged account could
// invoke that executable for any configured `hostId` and only then be told the
// workingDir was outside its workspace (reproduced upstream: wake spy fired, 403).
if (isMultiUserMode() && !isAdmin(req)) {
return createErrorResponse(ApiErrorCode.FORBIDDEN, 'Remote hosts are admin-only in multi-user mode');
}
const { hostId, remoteSessionName } = body.attachRemoteSession;
const host = (await readRemoteHosts(CODEMAN_CONFIG_DIR)).find((item) => item.id === hostId);
if (!host) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Remote host not found');
// An explicit wake request is the only thing that may wake a host, and the user
// pressing Attach IS one (see quick-start for the same gate, and
// `remote-wake.ts` for what must never call this). Without it a sleeping host
// answers with an ssh failure that blames anything but the machine being asleep.
const hostWake = await remoteWake.ensureHostAwake(wakeableHost(host), {
timeoutMs: REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS,
// No session yet, so the wake events name their requester (multi-user routing).
requestedBy: ownerFor(req),
});
if (hostWake === 'failed') {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
`${host.label} did not come back after a wake-on-LAN request — nothing was attached`
);
}
workingDir = `${host.username}@${host.host}:${remoteSessionName}`;
remote = toAttachedSessionRemote(host, remoteSessionName, workingDir);
}
@@ -1143,13 +1238,83 @@ export function registerSessionRoutes(
if (!endpoint) {
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
}
const contextLength = endpoint.modelContextLengths?.[body.modelId];
// Some CLIs (today: only claude) carry enough of their own fixed system-prompt/tool-
// schema overhead that a small enough real context guarantees a first-message failure
// no matter what CLAUDE_CODE_MAX_CONTEXT_TOKENS says — confirmed live at ~36.4K tokens
// against a model configured with a real 16384-token context. Warn before committing
// to a restart that's certain to fail, rather than letting the user discover it via a
// cryptic 400 from the CLI itself. Answered by `confirmedContext` (or the legacy
// `confirmed`, which still means both) — NOT by `confirmedSwap`: this warning is
// about the caller's own session, and the swap warning below is about someone
// else's, so an answer to one is not consent to the other.
if (!(body.confirmed || body.confirmedContext) && exceedsSafeContextFloor(entry, contextLength)) {
return {
requiresContextWarning: true,
modelId: body.modelId,
contextLength,
minSafeContextTokens: CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
};
}
// llama.cpp runs exactly one model at a time; llama-swap unloads and reloads it on
// demand, which can take anywhere from a few seconds to over a minute — long enough
// that a session mid-swap looks indistinguishable from one that never left the native
// backend. Feature-detected via llama-swap's own `GET /running` (a plain llama.cpp
// server has no such endpoint and reads as `isLlamaSwap: false` — nothing to check).
const swapStatus = await getLlamaSwapStatus(endpoint);
const currentlyLoaded = swapStatus.running.find((r) => r.state === 'ready')?.model ?? swapStatus.running[0]?.model;
// Distinct from targetReady below: this is ONLY about whether proceeding would evict a
// model another session is actively using — true even if nothing is loaded at all yet
// would be wrong here (nothing to evict), so this stays narrowly "a DIFFERENT model is
// currently ready".
const swapNeeded = swapStatus.isLlamaSwap && !!currentlyLoaded && currentlyLoaded !== body.modelId;
// Whether the TARGET model itself is already the one loaded and ready — false whether
// nothing is loaded yet, a different model is loaded, or this one is loaded but still
// mid-load. Drives both the actual load trigger below and modelSwapInProgress in the
// response; deliberately broader than swapNeeded, which only gates the confirmation ask.
const targetReady = swapStatus.running.some((r) => r.model === body.modelId && r.state === 'ready');
// Only ask when switching would actually take the model away from another session
// that is currently using it — never just because a swap is needed at all. Answered
// by `confirmedSwap` (or the legacy `confirmed`). ⚠ It must NOT read
// `confirmedContext`: this check runs second, and while the two shared one flag a
// user who clicked past a too-small-context warning had already, silently, agreed to
// evict another session's model.
if (swapNeeded && !(body.confirmed || body.confirmedSwap)) {
const conflicting = [...ctx.sessions.values()].filter(
(s) =>
s.id !== session.id && s.customModel?.endpointId === endpoint.id && s.customModel?.modelId === currentlyLoaded
);
if (conflicting.length > 0) {
// Applying a custom model is ungated for any session owner, so in multi-user
// mode a non-admin pointing their own session at a shared endpoint must not
// learn another user's session names in the confirm dialog — with
// autoNameSessions on, those names are that user's own prompts. The swap is
// still blocked pending confirmation regardless of ownership (a foreign
// session is just as real a disruption); only which ones get NAMED is scoped.
const requestUser = getAuthUser(req);
const affectedSessions = conflicting
.filter((s) => canAccessOwned(requestUser, s.owner))
.map((s) => ({ id: s.id, name: s.name }));
return { requiresConfirmation: true, currentlyLoadedModel: currentlyLoaded, affectedSessions };
}
}
// A CLI whose config alone cannot select the model also gets its `model` launch param
// forced (pi/omp `custom/<id>`, grok's block name). The argv engine DROPS a token that
// fails its pattern rather than quoting it, which would silently launch the CLI on its
// own default provider again, so refuse an id the pattern cannot carry up front.
const modelSpec = entry.launch.params.model;
const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id);
const applied = applyCustomModelInjection(
entry,
endpoint,
body.modelId,
session.id,
contextLength,
session.workingDir
);
if (!applied) {
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
}
@@ -1182,9 +1347,17 @@ export function registerSessionRoutes(
removeConfigDir(previousConfigDir);
}
// Actually kick off llama-swap's load now, rather than waiting on the restarted CLI's
// own first prompt to do it — confirmed live that applying a selection alone never
// reached the llama-swap server at all (nothing in its own logs), since llama-swap has
// no "switch model" admin call, only a real inference request naming the model.
if (swapStatus.isLlamaSwap && !targetReady) {
triggerLlamaSwapLoad(endpoint, body.modelId);
}
const restarted = await session.restartCli();
persistAndBroadcastSession(ctx, session);
return { customModel: session.customModel, restarted };
return { customModel: session.customModel, restarted, modelSwapInProgress: swapStatus.isLlamaSwap && !targetReady };
});
// ========== Delete Session ==========
@@ -1214,6 +1387,8 @@ export function registerSessionRoutes(
}
const session = findSessionOrFail(ctx, id, req);
// Wake state is dropped by `cleanupSession` itself (server.ts), on EVERY cleanup
// path — not here: the scheduled-run and admin paths clean up without this route.
await ctx.cleanupSession(session.id, killMux, 'user_delete');
return {};
});
@@ -1449,6 +1624,67 @@ export function registerSessionRoutes(
// Terminal I/O (input, resize, buffer)
// ═══════════════════════════════════════════════════════════════
// ========== Wake-on-LAN: state + manual trigger ==========
//
// Both routes are session-scoped (not host-scoped) because the wake flow needs the
// SESSION: a woken host whose pane is not reattached is still a dead terminal, and an
// exhausted COD-108 backoff never retries on its own. The probe in `/reachability` is
// the same cheap TCP connect the input path uses and it NEVER wakes a host — the UI
// decides that, with the button.
app.get('/api/sessions/:id/reachability', async (req) => {
const { id } = req.params as { id: string };
const session = findSessionOrFail(ctx, id, req);
const remote = session.remote;
if (!remote) {
return { success: true, data: { reachable: true, probeable: true, wakeConfigured: 'none' as const } };
}
const force = (req.query as { force?: string })?.force === '1';
// `reachable: null` + `probeable: false` for a host behind a jump host / SOCKS proxy:
// the probe cannot reach it, so the UI shows no banner and stops polling.
const reachable = await remoteWake.checkReachable(session, { force });
return {
success: true,
data: {
reachable,
probeable: isProbeable(remote),
wakeConfigured: await remoteWake.wakeConfigured(session),
host: remote.host,
label: remote.label,
},
};
});
app.post('/api/sessions/:id/wake', async (req) => {
const { id } = req.params as { id: string };
const session = findSessionOrFail(ctx, id, req);
if (!session.remote) {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Not a remote session');
}
// The UI uses this to route to the host config dialog instead of a dead button.
if (!(await remoteWake.hasWakeTarget(session))) {
return createErrorResponse(
ApiErrorCode.INVALID_INPUT,
'No wake-on-LAN target configured for this host (set a MAC address or a wake command)'
);
}
// The button is pressed from the SAME dashboard the create/attach paths are, under
// the same reverse proxy — so it holds the request open the same way and needs the
// same request budget, not the 90 s session default (see remote-wake.ts).
const woke = await remoteWake.ensureAwake(session, {
force: true,
timeoutMs: REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS,
});
return {
success: true,
data: {
woke,
reachable: await remoteWake.checkReachable(session),
wakeConfigured: await remoteWake.wakeConfigured(session),
},
};
});
// ========== Send Input ==========
app.post('/api/sessions/:id/input', async (req, reply) => {
@@ -1490,6 +1726,42 @@ export function registerSessionRoutes(
return {};
}
// Wake-on-LAN (remote-wake.ts): a wake-enabled remote host that suspended leaves
// the local ssh pane STALLED, and `send-keys` succeeds against it — the bytes
// would vanish with no error anywhere. Give the registry the chance to probe the
// host, wake it, reattach, and own delivery before we write into nothing.
//
// Costs nothing for non-wake hosts (the `wakeCommand` guard) or while the host is
// known reachable inside the probe throttle window; the probe itself is a bare
// TCP connect on wake-enabled hosts only, at most once per
// REMOTE_WAKE_PROBE_MIN_INTERVAL_MS.
if (!duplicate && (await remoteWake.hasWakeTarget(session))) {
if (wantsWait) {
// Send-and-wait keeps the response open anyway, so blocking on the wake is
// simpler and more correct than buffering (buffering would break the wait).
// A host that never comes back is an error here, as on the create/attach
// paths: writing into the stalled pane would answer `delivered:true` plus a
// timeout, which is the combination the API docs send callers to the wrong
// recovery for.
if (!(await remoteWake.ensureAwake(session))) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
`${session.remote?.label ?? 'the remote host'} did not come back after a wake-on-LAN request — nothing was sent`
);
}
} else {
const outcome = await remoteWake.handleInput(session, inputStr);
// The registry holds the bytes and flushes them in order once the pane is
// reattached. The client's ACK is this 200 — a tagged retry is deduped
// (`shouldApplyInput` above already consumed the seq), so nothing is lost.
// `buffered` is additive to the historical bare `{}`; `dropped` says the chunk
// was over the wake buffer's cap and is GONE (a 200 with no field could not
// tell delivered from buffered from dropped).
if (outcome === 'buffered') return { buffered: true };
if (outcome === 'dropped') return { buffered: true, dropped: true };
}
}
// Only a waiting request pays for the tmux probe: the browser's plain input path
// (thousands of calls per session) must stay exec-free.
const workerDead = wantsWait && workerIsDead(ctx.mux, session);
@@ -2632,14 +2904,16 @@ export function registerSessionRoutes(
// returns null when unavailable, in which case we fall back to history.
const muxName = session.muxName;
const captureStartedAt = performance.now();
// The visible path used to pass no options at all. It passes one now for a
// single reason: `capturedGeometry` comes BACK on it, and the response has
// to tell the client what size the frame it is about to render was built
// for. See PaneCaptureOptions.capturedGeometry.
const captureOpts: PaneCaptureOptions = isFullReload
? { fullHistory: true, historyLimitLines: tmuxHistoryLimit, maxCaptureBytes: terminalBufferMaxBytes }
: {};
const liveMuxBuffer =
muxName && typeof ctx.mux.captureActivePaneBuffer === 'function'
? ctx.mux.captureActivePaneBuffer(
muxName,
isFullReload
? { fullHistory: true, historyLimitLines: tmuxHistoryLimit, maxCaptureBytes: terminalBufferMaxBytes }
: undefined
)
? ctx.mux.captureActivePaneBuffer(muxName, captureOpts)
: null;
const captureFinishedAt = performance.now();
const hasLiveMuxBuffer = liveMuxBuffer !== null && liveMuxBuffer.length > 0;
@@ -2785,6 +3059,25 @@ export function registerSessionRoutes(
// what existed before the cut. The gap is what the indicator reports.
retainedBytes: cleanBuffer.length,
source,
// The pane geometry this frame was drawn for. A visible-frame capture
// positions every row absolutely, so a client whose terminal has fewer
// rows than this overwrites its last line with the overflow and loses
// the rows underneath. The client compares these against its own size.
//
// BOTH FIELDS ARE ABSENT unless this response really carries a capture,
// and that is the honest answer rather than a gap to paper over. Two
// separate things can leave a frame unpositioned. The cursor query is
// what produces the absolute addressing in the first place, so a capture
// that lost it returned a raw frame with no row positioning in it. And a
// capture can report geometry and STILL hand back nothing: the
// full-history path returns '' for a pane holding nothing visible, which
// drops `source` to `history` while `capturedGeometry` is already
// written, so the geometry has to be suppressed HERE rather than trusted
// to be missing. Naming a size for a body that is the byte stream would
// describe a frame that was never drawn and invite the client to repair
// damage that does not exist.
captureCols: hasLiveMuxBuffer ? captureOpts.capturedGeometry?.cols : undefined,
captureRows: hasLiveMuxBuffer ? captureOpts.capturedGeometry?.rows : undefined,
};
});
@@ -3043,6 +3336,7 @@ export function registerSessionRoutes(
effort,
parentSessionId,
agentOrigin,
customModel,
} = parseBody(QuickStartSchema, req.body);
// Resolved ONCE here: the same value labels a case directory this request creates
@@ -3095,11 +3389,31 @@ export function registerSessionRoutes(
grokConfig ||
deepSeekConfig ||
ompConfig ||
openCodeConfig
openCodeConfig ||
customModel
) {
return createErrorResponse(
ApiErrorCode.INVALID_INPUT,
'envOverrides, effort, modelOverride, and per-CLI config are not supported for remote cases (they do not cross ssh). Configure the remote command via the host command override instead.'
'envOverrides, effort, modelOverride, per-CLI config, and custom model endpoints are not supported for remote cases (they do not cross ssh). Configure the remote command via the host command override instead.'
);
}
// The user pressing "Run" on a case whose host is asleep IS an explicit wake
// request (docs/remote-sessions.md §Wake-on-LAN), and the tmux probe below would
// otherwise fail with "could not verify tmux on remote host …" — an ssh failure
// that blames tmux for a machine that is merely suspended. Wired HERE, in the HTTP
// route, and deliberately NOT in the shared session service: `cron-service.ts`
// builds sessions through the service, and a wake down there would re-wake the
// host on every schedule (the failure invariant #1 exists to prevent).
const hostWake = await remoteWake.ensureHostAwake(wakeableHost(host), {
timeoutMs: REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS,
// No session yet, so the wake events name their requester (multi-user routing).
requestedBy: ownerFor(req),
});
if (hostWake === 'failed') {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
`${host.label} did not come back after a wake-on-LAN request — the session was not started`
);
}
@@ -3108,6 +3422,20 @@ export function registerSessionRoutes(
// surfaces a clear, structured error instead of a dead "tmux: command not found" pane.
const tmuxCheck = await checkRemoteTmuxAvailable(host);
if (!tmuxCheck.ok) {
// An unreachable host and a host without tmux fail the same way over ssh, so the
// probe's own message would send the user hunting for a tmux install. Ask the
// registry (which just probed, when it woke the host) which of the two it is.
// `=== false` on purpose: a proxied host answers `null` (the probe cannot reach
// it), and an unknown verdict must not replace the real ssh error with
// "not reachable" over a host that is fine.
if ((await remoteWake.checkHostReachable(wakeableHost(host))) === false) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
hostWake === 'no-target'
? `${host.label} (${host.host}) is not reachable, and this host has no wake-on-LAN target — configure a MAC address or a wake command first`
: `${host.label} (${host.host}) is not reachable`
);
}
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, tmuxCheck.error || 'remote host is missing tmux');
}
@@ -3130,11 +3458,12 @@ export function registerSessionRoutes(
grokConfig ||
deepSeekConfig ||
ompConfig ||
openCodeConfig
openCodeConfig ||
customModel
) {
return createErrorResponse(
ApiErrorCode.INVALID_INPUT,
'envOverrides, effort, and per-CLI config are not supported for docker cases (they do not cross into the container). Configure the container via the docker host command override instead.'
'envOverrides, effort, per-CLI config, and custom model endpoints are not supported for docker cases (they do not cross into the container). Configure the container via the docker host command override instead.'
);
}
@@ -3422,7 +3751,153 @@ export function registerSessionRoutes(
);
const qsTerminalHistoryConfig = await ctx.getTerminalHistoryConfig();
const qsGatedEnvOverrides = await clampEnvOverridesForOwner(owner, envOverrides);
const session = new Session({
const qsResolvedOmpConfig = resolveOmpConfigForCreate(mode, resolvedCasePath, ompConfig);
// Custom Model Endpoint Profiles, applied AT CREATE TIME (docs/custom-model-endpoints-plan.md)
// rather than via the dedicated restart-in-place route (POST /api/sessions/:id/custom-
// model, still what an ALREADY-RUNNING session uses to switch later): computing the
// injection before the process exists and launching directly on it avoids the visible
// native-boot-then-restart the restart-after-launch design otherwise shows on every
// custom-model run — most jarring on a CLI like Codex whose TUI fully reinitializes.
// Mirrors the dedicated route's own checks (llama-swap conflict, unsupported CLI,
// unknown endpoint, a model id the CLI's argv pattern can't carry) rather than trusting
// a lighter version of them, since this is the same server-side authority reached a
// different way, not a separate, less-checked path.
let qsCustomModelEnvOverrides = qsGatedEnvOverrides;
// Only the INJECTED keys (never the caller's envOverrides merged in) — this is what
// setCustomModel() bookkeeping must be given below. The Session constructor already
// applies qsCustomModelEnvOverrides (the full merged set) directly; re-merging that
// full set into setCustomModel() would put CLAUDE_CODE_EFFORT_LEVEL back after the
// constructor stripped it (see setCustomModel()'s own doc comment in session.ts).
let qsCustomModelAppliedEnvOverrides: Record<string, string> | undefined;
let qsCustomModelLaunchModel: string | undefined;
let qsCustomModelSessionId: string | undefined;
let qsCustomModelSwapInProgress = false;
let qsCustomModelBookkeeping:
| {
endpointId: string;
modelId: string;
label?: string;
envKeys: string[];
configDir?: string;
launchModel?: string;
}
| undefined;
if (customModel) {
const cmEntry = getCli(mode);
if (!cmEntry) return createErrorResponse(ApiErrorCode.INVALID_INPUT, `No CLI registry entry for mode ${mode}`);
if (cmEntry.capabilities.customModelInjection.kind === 'unsupported') {
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${mode} has no known custom-model mechanism`);
}
const cmHosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
const cmEndpoint = cmHosts.find((h) => h.id === customModel.endpointId);
if (!cmEndpoint) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
const cmContextLength = cmEndpoint.modelContextLengths?.[customModel.modelId];
// See the dedicated route's own comment for the full reasoning: some CLIs' own fixed
// overhead can exceed a small enough real context on the very first message,
// regardless of contextLengthVar. Warn before creating a session that's certain to
// fail immediately.
// See the dedicated route above for why this reads `confirmedContext` and never
// `confirmedSwap`.
if (
!(customModel.confirmed || customModel.confirmedContext) &&
exceedsSafeContextFloor(cmEntry, cmContextLength)
) {
return {
requiresContextWarning: true,
modelId: customModel.modelId,
contextLength: cmContextLength,
minSafeContextTokens: CLAUDE_MIN_SAFE_CONTEXT_TOKENS,
};
}
// See the dedicated route's own comment for the full reasoning: llama.cpp runs one
// model at a time, llama-swap swaps on demand, and switching away from what another
// live session is actively using deserves a warning, not a silent switch. There is no
// "self" to exclude from the affected-sessions scan here — this session doesn't exist
// yet.
const cmSwapStatus = await getLlamaSwapStatus(cmEndpoint);
const cmCurrentlyLoaded =
cmSwapStatus.running.find((r) => r.state === 'ready')?.model ?? cmSwapStatus.running[0]?.model;
const cmSwapNeeded = cmSwapStatus.isLlamaSwap && !!cmCurrentlyLoaded && cmCurrentlyLoaded !== customModel.modelId;
// Broader than cmSwapNeeded (which only gates the confirmation ask above): true
// whenever the TARGET model isn't already loaded and ready, including when nothing
// is loaded at all yet. Drives the actual load trigger below.
const cmTargetReady = cmSwapStatus.running.some((r) => r.model === customModel.modelId && r.state === 'ready');
qsCustomModelSwapInProgress = cmSwapStatus.isLlamaSwap && !cmTargetReady;
if (cmSwapNeeded && !(customModel.confirmed || customModel.confirmedSwap)) {
const cmConflicting = [...ctx.sessions.values()].filter(
(s) => s.customModel?.endpointId === cmEndpoint.id && s.customModel?.modelId === cmCurrentlyLoaded
);
if (cmConflicting.length > 0) {
// Same reasoning as the dedicated /custom-model route above: the swap is
// still blocked pending confirmation regardless of ownership, but a
// non-admin caller only learns the names of sessions they can access.
const cmRequestUser = getAuthUser(req);
const cmAffectedSessions = cmConflicting
.filter((s) => canAccessOwned(cmRequestUser, s.owner))
.map((s) => ({ id: s.id, name: s.name }));
return {
requiresConfirmation: true,
currentlyLoadedModel: cmCurrentlyLoaded,
affectedSessions: cmAffectedSessions,
};
}
}
// Minted ourselves (rather than left to Session's own default) so the injection
// below — and any configDir it writes — can target the REAL id the session launches
// with, not a placeholder: `new Session({ id: ... })` accepts an explicit id for
// exactly this reason.
qsCustomModelSessionId = randomUUID();
const cmApplied = applyCustomModelInjection(
cmEntry,
cmEndpoint,
customModel.modelId,
qsCustomModelSessionId,
cmContextLength,
resolvedCasePath
);
if (!cmApplied) {
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${mode} has no known custom-model mechanism`);
}
const cmModelSpec = cmEntry.launch.params.model;
if (
cmApplied.launchModel !== undefined &&
cmModelSpec?.type === 'token' &&
!matchesPattern(cmModelSpec.pattern, cmApplied.launchModel)
) {
removeConfigDir(cmApplied.configDir);
return createErrorResponse(
ApiErrorCode.INVALID_INPUT,
`Model id ${JSON.stringify(customModel.modelId)} cannot be passed to ${mode} on its command line`
);
}
qsCustomModelEnvOverrides = { ...qsGatedEnvOverrides, ...cmApplied.envOverrides };
qsCustomModelAppliedEnvOverrides = cmApplied.envOverrides;
qsCustomModelLaunchModel = cmApplied.launchModel;
qsCustomModelBookkeeping = {
endpointId: cmEndpoint.id,
modelId: customModel.modelId,
label: cmEndpoint.label,
envKeys: cmApplied.envKeys,
configDir: cmApplied.configDir,
launchModel: cmApplied.launchModel,
};
// Actually kick off llama-swap's load now — see the dedicated apply route's own
// comment on triggerLlamaSwapLoad for why this can't just wait on the launched CLI's
// first prompt. Fired here, before the session is even created, so the load starts
// concurrently with Claude/Codex/etc. booting rather than after.
if (qsCustomModelSwapInProgress) {
triggerLlamaSwapLoad(cmEndpoint, customModel.modelId);
}
}
const qsSessionOptions: ConstructorParameters<typeof Session>[0] = {
id: qsCustomModelSessionId,
workingDir: resolvedCasePath,
name: sessionName ? sessionName.slice(0, MAX_SESSION_NAME_LENGTH) : '',
mux: ctx.mux,
@@ -3440,15 +3915,42 @@ export function registerSessionRoutes(
piConfig: mode === 'pi' ? qsGatedPiConfig : undefined,
grokConfig: mode === 'grok' ? qsGatedGrokConfig : undefined,
deepSeekConfig: mode === 'deepseek' ? qsGatedDeepSeekConfig : undefined,
ompConfig: resolveOmpConfigForCreate(mode, resolvedCasePath, ompConfig),
envOverrides: qsGatedEnvOverrides,
ompConfig: qsResolvedOmpConfig,
envOverrides: qsCustomModelEnvOverrides,
effort,
remote,
docker,
resumeSessionId: dockerResumeId,
tmuxHistoryLimit: qsTerminalHistoryConfig.tmuxHistoryLimit,
parentSessionId: qsParentSessionId,
});
};
// Force the custom-model selection's launchModel (pi/omp `custom/<id>`, grok's
// `[model.<name>]` block name) onto whichever config field the registry says the
// CLI's `model` launch param lives in — mirrors Session._withCustomModelLaunchModel,
// which the restart-in-place path already uses, rather than a hardcoded per-CLI
// branch here that a CLI landing its injection recipe later would silently miss.
if (qsCustomModelLaunchModel !== undefined) {
const qsCustomModelField = getCli(mode)?.launch.legacyConfigField;
if (qsCustomModelField) {
const qsSessionOptionsBag = qsSessionOptions as unknown as Record<string, unknown>;
qsSessionOptionsBag[qsCustomModelField] = {
...((qsSessionOptionsBag[qsCustomModelField] as Record<string, unknown>) ?? {}),
model: qsCustomModelLaunchModel,
};
} else {
qsSessionOptions.model = qsCustomModelLaunchModel;
}
}
const session = new Session(qsSessionOptions);
// Records the selection for session.customModel/getCustomModelForPersist() and future
// clear/switch calls — the actual env vars and launch-model config are already part of
// the launch above (constructor envOverrides, piConfig/grokConfig/ompConfig.model), so
// this is bookkeeping only, never a restart: setCustomModel() is synchronous state, no
// tmux IO of its own (see its own doc comment in session.ts).
if (qsCustomModelBookkeeping) {
session.setCustomModel(qsCustomModelBookkeeping, qsCustomModelAppliedEnvOverrides);
}
// Auto-detect completion phrase from CLAUDE.md BEFORE broadcasting
// so the initial state already has the phrase configured (only if globally enabled)
@@ -3544,6 +4046,7 @@ export function registerSessionRoutes(
sessionId: session.id,
casePath: resolvedCasePath,
caseName,
...(customModel ? { modelSwapInProgress: qsCustomModelSwapInProgress } : {}),
};
} catch (err) {
// Clean up session on error to prevent orphaned resources
@@ -4666,4 +5169,10 @@ export function registerSessionRoutes(
return { path: filepath, filename };
});
// Returned so the server can own the registry's LIFETIME (drop state when a session is
// cleaned up on any of its paths, resolve in-flight wakes on shutdown). The wake-CAPABLE
// code stays here: `test/remote-wake.test.ts` pins that `server.ts` calls nothing but
// `drop`/`stop` on this handle, so no timer path can reach a wake through it.
return remoteWake;
}
+79
View File
@@ -19,6 +19,7 @@ import {
} from '../config/terminal-history.js';
import { MAX_EDITABLE_BYTES } from '../config/file-editing.js';
import { MIN_MATCH_LENGTH, MAX_MATCH_LENGTH } from '../config/agent-wait.js';
import { MAX_WAKE_MACS } from '../config/remote-wake-limits.js';
import { enabledCliIds, enabledClis } from '../config/cli-registry/registry.js';
import type { SessionMode } from '../types.js';
@@ -737,6 +738,36 @@ export const RemoteHostSchema = z.object({
.max(32)
.optional(),
commands: RemoteCommandOverridesSchema,
// Wake-on-LAN: a single executable path (no arguments, no shell) run to power a
// SLEEPING host back on, e.g. `/home/joe/bin/whuff`. Executed via spawn without
// a shell, so there is no shell layer to escape; the regexes are belt-and-braces
// (and the no-whitespace rule rejects an argument list before it can fail as a
// confusing ENOENT at wake time). See docs/remote-sessions.md §Wake-on-LAN.
wakeCommand: z
.string()
.min(1)
.max(4096)
.regex(/^\S+$/, 'Wake command must be a single executable path (no arguments)')
.regex(NO_SHELL_META, 'Invalid characters in wake command')
.optional(),
// Wake-on-LAN MAC address(es), comma-separated. Structural: only hex pairs with
// `:`/`-` separators, so nothing here can be a shell token even by accident (the
// value never reaches a shell — Codeman builds the magic packet itself).
wakeMac: z
.string()
.min(11)
.max(128)
.regex(
/^[0-9a-fA-F]{2}([:-][0-9a-fA-F]{2}){5}(\s*,\s*[0-9a-fA-F]{2}([:-][0-9a-fA-F]{2}){5})*$/,
'Wake MAC must be one or more MAC addresses, comma-separated'
)
// ⚠ The character cap admits seven MACs while parseMacList takes at most
// MAX_WAKE_MACS, all-or-nothing. Without this the extra ones validated, persisted,
// and then resolved to NO wake target, so the host read as unconfigured.
.refine((value) => value.split(',').length <= MAX_WAKE_MACS, {
message: `Wake MAC accepts at most ${MAX_WAKE_MACS} comma-separated addresses`,
})
.optional(),
});
export const RemoteCaseLinkSchema = z.object({
@@ -1033,6 +1064,27 @@ export const QuickStartSchema = z.object({
* because it takes an existing `workingDir` and so never creates a directory to label.
*/
agentOrigin: z.string().max(64).optional(),
/**
* Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md): launches directly
* on this saved endpoint/model instead of the mode's native backend, computed server-side
* from the admin-configured endpoint store the same way `POST /api/sessions/:id/custom-
* model` does — never trusting raw env values from the client. One-shot, launch-time
* equivalent of that route: no restart, so no visible relaunch (that route's restart-in-
* place is still what an ALREADY-RUNNING session uses to switch later). Rejected for
* remote/docker cases, same reasoning as `envOverrides` above. The three confirmation
* flags mirror that route's fields; see `SessionCustomModelSchema` for why there are
* two specific ones rather than the single legacy `confirmed`.
*/
customModel: z
.object({
endpointId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid endpoint id'),
modelId: z.string().min(1).max(200),
confirmed: z.boolean().optional(),
confirmedContext: z.boolean().optional(),
confirmedSwap: z.boolean().optional(),
})
.strict()
.optional(),
});
// ========== Hook Events ==========
@@ -1932,6 +1984,17 @@ export const CustomModelHostSchema = z.object({
authStyle: z.enum(['bearer', 'api-key']).optional(),
models: z.array(z.string().max(200)).max(200).optional(),
lastDiscoveredAt: z.string().max(64).optional(),
// The Run-menu picker's per-endpoint default; validated against `models` at the
// route layer (schema-level cross-field checks can't see the array narrowed the
// same way a `.refine()` closure could, and the route already re-reads the stored
// host to apply it, so the check belongs there once, not duplicated into a refine
// that would run on every unrelated field edit too).
defaultModelId: z.string().max(200).optional(),
// Server-populated by discovery (custom-model-routes.ts); accepted here only so a client
// round-tripping the GET response back through PUT (edit-save) doesn't drop it.
modelContextLengths: z.record(z.string().max(200), z.number().int().positive().max(100_000_000)).optional(),
// Same reasoning as modelContextLengths above.
modelSizesGB: z.record(z.string().max(200), z.number().positive().max(100_000)).optional(),
});
/** POST /api/sessions/:id/custom-model — apply or clear a session's custom-model selection. */
@@ -1939,6 +2002,22 @@ export const CustomModelSelectionSchema = z.union([
z.object({
endpointId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid endpoint id'),
modelId: z.string().min(1).max(200),
/**
* Two DIFFERENT questions can block a launch, and answering one is not consent to
* the other: `confirmedContext` answers "this model's context window is below the
* floor for this CLI", which affects only the caller, while `confirmedSwap` answers
* "loading this will unload the model another session is using", which affects
* someone else. They were one flag until the context check (which runs first)
* silently spent the swap answer too, so a user clicking "launch anyway" past a
* too-small context evicted another session's model without ever being asked.
*
* `confirmed` is the original single flag and still means BOTH, because it shipped
* in the HTTP-API-only cut of this feature and an existing caller must keep working.
* New callers should send the specific one they actually asked about.
*/
confirmed: z.boolean().optional(),
confirmedContext: z.boolean().optional(),
confirmedSwap: z.boolean().optional(),
}),
z.object({ clear: z.literal(true) }),
]);
+157 -5
View File
@@ -43,6 +43,8 @@ import { hostname as getHostname, uptime as osUptime } from 'node:os';
import { looksLikeHostReboot, newestPersistedActivity, planRebootRestore } from '../reboot-restore.js';
import { rebootRestoreRegistry } from './reboot-restore-registry.js';
import { dataPath, getDataDir, CODEMAN_INSTANCE } from '../config/instance.js';
import { readRemoteHosts, rehydrateRemoteHostFields } from '../remote-hosts.js';
import type { RemoteWakeRegistry } from '../remote-wake.js';
import { normalizeBasePath, stripBasePath, joinBasePath } from '../config/base-path.js';
import { GLYPH, palette } from '../cli-style.js';
import { getHookSecret } from '../config/hook-secret.js';
@@ -69,7 +71,7 @@ import {
import { imageWatcher } from '../image-watcher.js';
import { workflowRunWatcher, summarizeRun } from '../workflow-run-watcher.js';
import { attachmentRegistry, buildFileThumbnailRoute, registerExternalAttachment } from '../attachment-registry.js';
import { getCli } from '../config/cli-registry/registry.js';
import { getCli, enabledClis } from '../config/cli-registry/registry.js';
import { readCustomModelHosts } from '../custom-model-hosts.js';
import { applyCustomModelInjection, customModelConfigDir, removeConfigDir } from '../custom-model-injection-apply.js';
import type { CustomModelBookkeeping } from '../types/session.js';
@@ -193,6 +195,11 @@ import {
registerWebviewRoutes,
registerTabLayoutRoutes,
registerCustomModelRoutes,
refreshAllCustomModelHosts,
readCustomModelEndpointsEnabled,
closeAllLlamaSwapLogTails,
detectCustomModelSwapDisplacements,
pruneIdleLlamaSwapLogTails,
tryWebviewRefererFallback,
} from './routes/index.js';
import { isLostWebviewFrameNavigation } from './webview-proxy.js';
@@ -205,11 +212,32 @@ const __dirname = dirname(fileURLToPath(import.meta.url));
// while capping growth of `sseClientsById` and blocking pathological inputs.
const SSE_CLIENT_ID_RE = /^[A-Za-z0-9_-]{8,64}$/;
const CODEX_USAGE_POLL_INTERVAL_MS = 5 * 60_000;
const CUSTOM_MODEL_REDISCOVER_INTERVAL_MS = 5 * 60_000;
// Much shorter than the model-LIST refresh above on purpose: this catches an actual
// eviction (a session's model no longer loaded, silently swapped out by another
// session's use), which the user wants to know about promptly, not once every 5
// minutes. Cheap either way — one /running GET per distinct endpoint with at least
// one live custom-model session, not per session.
const CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS = 20_000;
function escapeHtmlText(value: string): string {
return value.replaceAll('&', '&amp;').replaceAll('<', '&lt;').replaceAll('>', '&gt;');
}
/**
* Escapes a JSON string for safe embedding as the body of an inline `<script>`
* tag: `<` becomes the six-character sequence `<`, which both a JSON
* parser and a plain JS string literal decode back to `<` (both treat
* `\uXXXX` identically), but which can never itself form the two literal
* characters `<` `/` a browser's HTML tokenizer looks for to end the tag. A
* value containing a literal `</script>` would otherwise close the tag early
* and turn the rest of the document into inert script-body text. Exported so
* it unit-tests without constructing a WebServer (which needs a real tmux).
*/
export function escapeScriptJson(json: string): string {
return json.replace(/</g, '\\u003c');
}
import {
SESSIONS_LIST_CACHE_TTL,
SCHEDULED_CLEANUP_INTERVAL,
@@ -268,8 +296,18 @@ export class WebServer extends EventEmitter {
// Store session listener references for explicit cleanup (prevents memory leaks)
private sessionListenerRefs: Map<string, SessionListenerRefs> = new Map();
private scheduledRuns: Map<string, ScheduledRun> = new Map();
/** De-dupe state for the swap-displacement sweep — see detectCustomModelSwapDisplacements. */
private _customModelDisplacedNotified: Set<string> = new Set();
/** Cron service (assigned in setupRoutes). */
private cronService!: CronService;
/**
* Wake-on-LAN registry, returned by `registerSessionRoutes`. Held for its LIFETIME
* only — `drop()` on session cleanup, `stop()` on shutdown. Waking from here would
* re-wake a host on every timer tick (the invariant `remote-wake.ts` documents), so
* the wiring guard in `test/remote-wake.test.ts` pins that this file calls nothing
* but `drop`/`stop` on it.
*/
private remoteWake: RemoteWakeRegistry | null = null;
private sse: SseStreamManager;
private store = getStore();
private tabLayouts!: TabLayoutService;
@@ -1070,7 +1108,9 @@ export class WebServer extends EventEmitter {
registerStatusTelemetryRoutes(this.app, ctx);
registerSystemRoutes(this.app, ctx);
registerCaseRoutes(this.app, ctx);
registerSessionRoutes(this.app, ctx);
// The registry's lifetime is the server's: it drops per-session wake state on every
// cleanup path and resolves in-flight wakes on shutdown.
this.remoteWake = registerSessionRoutes(this.app, ctx);
registerRespawnRoutes(this.app, ctx);
registerRalphRoutes(this.app, ctx);
registerPlanRoutes(this.app, ctx);
@@ -1212,7 +1252,13 @@ export class WebServer extends EventEmitter {
return undefined;
}
try {
return applyCustomModelInjection(entry, endpoint, saved.modelId, session.id)?.envOverrides;
return applyCustomModelInjection(
entry,
endpoint,
saved.modelId,
session.id,
endpoint.modelContextLengths?.[saved.modelId]
)?.envOverrides;
} catch (err) {
console.warn('[WebServer] Failed to rebuild custom-model env on recovery:', err);
return undefined;
@@ -1307,6 +1353,10 @@ export class WebServer extends EventEmitter {
session.ralphTracker.stopWatchingFixPlan();
}
// Custom Model Endpoint Profiles: drop this session's swap-displacement notify flag
// (see _checkCustomModelSwapDisplacements below) so it can't linger in that Set forever.
this._customModelDisplacedNotified.delete(sessionId);
// Kill all subagents spawned by this session (scoped to sessionId to avoid cross-session kills)
if (session && killMux) {
try {
@@ -1461,6 +1511,11 @@ export class WebServer extends EventEmitter {
sessionWaits.notifySignal(sessionId, 'exit');
sessionWaits.cancelAll(sessionId);
approvalInbox.resolveForSession(sessionId, 'session_ended');
// Wake state goes with the session on EVERY cleanup path (delete routes, the cron
// and admin paths, scheduled-run teardown, error paths) — that is why it lives here
// rather than in the two delete routes, where it left an entry behind, including up
// to 4 KB of the user's buffered keystrokes.
this.remoteWake?.drop(sessionId);
this.broadcast(SseEvent.SessionDeleted, { id: sessionId });
}
@@ -1602,6 +1657,23 @@ export class WebServer extends EventEmitter {
'</head>',
`<script>window.__codemanCliAvailable=${JSON.stringify(available)};</script>\n</head>`
);
// Which run modes the Run-menu picker (docs/custom-model-endpoints-plan.md) may
// generate an entry for: read generically off the registry's `capabilities`
// (never an id list here) so a CLI whose customModelInjection lands later shows
// up in the picker with no frontend change, and one that ships `unsupported`
// (antigravity, and `shell`'s `kind !== 'agent'`) never does.
const customModelClis = enabledClis()
.filter((entry) => entry.kind === 'agent' && entry.capabilities.customModelInjection.kind !== 'unsupported')
.map((entry) => ({ id: entry.id, label: entry.label }));
// Unlike the boolean-only __codemanCliAvailable above, this payload carries
// `label`, a string a user's own clis.json can set (CliEntry.label, up to 60
// chars) — see escapeScriptJson's own doc comment for why that needs escaping
// and __codemanCliAvailable's booleans never did.
const customModelClisJson = escapeScriptJson(JSON.stringify(customModelClis));
html = html.replace(
'</head>',
`<script>window.__codemanCustomModelClis=${customModelClisJson};</script>\n</head>`
);
}
if (!soloSessionId && process.env.CODEMAN_GESTURE === '1') {
html = html.replace('</head>', `<script>window.__codemanGestureAvailable=true;</script>\n</head>`);
@@ -2332,11 +2404,20 @@ export class WebServer extends EventEmitter {
'scheduled:',
'team:',
'case:',
'remote:',
'custom-model:',
];
if (SESSION_PREFIXES.some((p) => event.startsWith(p))) {
const d = (data ?? {}) as { sessionId?: string; id?: string; session?: { id?: string } };
const d = (data ?? {}) as { sessionId?: string; id?: string; session?: { id?: string }; username?: string };
const sessionId = d.sessionId ?? d.id ?? d.session?.id;
const owner = sessionId ? this.sessions.get(sessionId)?.owner : undefined;
// `remote:hostWaking` / `remote:hostWakeFailed` for a create/attach wake have no
// session yet (nothing exists until the host is up), so the registry names the
// requesting user instead; the payload carries `hostId`/`label`, which non-admins
// are not shown elsewhere. No session and no requester: admins only (fail closed).
if (!sessionId && event.startsWith('remote:') && d.username) {
return { username: d.username, sessionScoped: true };
}
return { owner, sessionScoped: true };
}
// #20/#38: clipboard:write writes into the receiver's OS clipboard — route it to
@@ -2709,6 +2790,64 @@ export class WebServer extends EventEmitter {
});
}
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md): keeps
// each saved endpoint's discovered model list current with no manual
// "Discover" click, so a model added on the server side (or one that drops
// off) shows up in the Run-menu picker within one cycle. Best-effort per
// endpoint (refreshAllCustomModelHosts skips one that's unreachable rather
// than failing the sweep) and off in tests for the same reason the Codex
// poll above is — no real network to hit, no server instance to keep alive.
if (!this.testMode) {
this.cleanup.setInterval(
() => {
// Reads the setting fresh on every tick, same reasoning as
// readPlanUsageTelemetryEnabled() beside it: a live toggle takes effect
// on the very next cycle, not just at server boot, and turning the
// feature off actually stops the polling instead of only hiding the UI.
void readCustomModelEndpointsEnabled()
.then((enabled) => {
if (!enabled) return;
return refreshAllCustomModelHosts();
})
.catch((err) => {
console.error('[custom-model] periodic re-discovery failed:', getErrorMessage(err));
});
},
CUSTOM_MODEL_REDISCOVER_INTERVAL_MS,
{ description: 'custom model endpoint re-discovery' }
);
}
// Custom Model Endpoint Profiles: the swap-conflict check on the apply/create routes
// only ever runs at THAT session's own launch/apply moment — it cannot catch a LATER
// eviction triggered by a different session's normal use, since llama-swap has no push
// notification of its own and only swaps in response to a real inference request
// (confirmed live: a session created while nothing else conflicted at that instant can
// still get silently displaced afterward). This periodic sweep is what catches that
// case after the fact and tells the displaced session's user, rather than leaving them
// to discover it only when their next prompt behaves unexpectedly.
if (!this.testMode) {
this.cleanup.setInterval(
() => {
detectCustomModelSwapDisplacements(this.sessions.values(), this._customModelDisplacedNotified)
.then((displacements) => {
for (const displacement of displacements) {
this.broadcast(SseEvent.CustomModelSwappedOut, displacement);
}
})
.catch((err) => {
console.error('[custom-model] swap-displacement check failed:', getErrorMessage(err));
});
// Same cadence, unrelated concern: close any /api/events tail (see
// getLatestLlamaSwapLogLine) nothing has polled in a while, so a loading banner
// that finished (or was abandoned) doesn't leave a connection open forever.
pruneIdleLlamaSwapLogTails();
},
CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS,
{ description: 'custom model swap-displacement check' }
);
}
// Start scheduled runs cleanup timer
this.cleanup.setInterval(
() => {
@@ -3116,6 +3255,9 @@ export class WebServer extends EventEmitter {
// For each alive mux session, create a Session object if it doesn't exist
const muxSessions = this.mux.getSessions();
// Host-level config lives in remote-hosts.json, not in the persisted session
// snapshot, so refresh the fields that only exist there (see the helper).
const remoteHostsById = new Map((await readRemoteHosts(getDataDir())).map((host) => [host.id, host]));
for (const muxSession of muxSessions) {
if (!this.sessions.has(muxSession.sessionId)) {
// Restore session settings from state.json (single source of truth)
@@ -3193,7 +3335,9 @@ export class WebServer extends EventEmitter {
// respawn rebuilds a LOCAL command, breaking the pane and silently
// erasing `remote` from state.json on the next persist. mux-sessions.json
// round-trips MuxSession.remote; state.json carries SessionState.remote.
remote: muxSession.remote ?? savedState?.remote,
// Host-level fields are refreshed from remote-hosts.json on top, or a
// field added to the host config after launch would never arrive.
remote: rehydrateRemoteHostFields(muxSession.remote ?? savedState?.remote, remoteHostsById),
// Docker metadata round-trips the same way (mux-sessions.json carries
// MuxSession.docker; state.json carries SessionState.docker), so recovery
// rebuilds the `docker exec` launch instead of a broken local command.
@@ -3576,6 +3720,10 @@ export class WebServer extends EventEmitter {
// got wrong once.
void stopDeepSeekWeb();
// Same teardown rule: the per-endpoint llama-swap log tails are otherwise closed
// only by the periodic idle sweep, whose interval is disposed just below.
closeAllLlamaSwapLogTails();
// Dispose all managed timers (intervals + resettable timeouts)
this.cleanup.dispose();
@@ -3587,6 +3735,10 @@ export class WebServer extends EventEmitter {
// response), so without this a 10-minute wait holds shutdown open.
sessionWaits.cancelEverything();
approvalInbox.stop();
// Same reason as `cancelEverything` above: an in-flight wake is awaited by a request,
// and `app.close()` (the last line of this method) does not abort in-flight requests —
// so without this a restart during a wake waits out the readiness poll.
this.remoteWake?.stop();
this.lastRecordedTokens.clear();
+33 -3
View File
@@ -5,7 +5,7 @@
* and referenced by the frontend (`SSE_EVENTS` in `constants.js`).
* Both files MUST be kept in sync.
*
* 158 event constants organized by category:
* 161 event constants organized by category:
* - **Core** (1): init
* - **Transport** (1): sse:heartbeat
* - **Session lifecycle** (23): created, updated, deleted, terminal, idle, working, ...
@@ -14,7 +14,7 @@
* - **Session: Plan** (4): planTaskUpdate, planCheckpoint, planRollback, planTaskAdded
* - **Tasks** (4): created, completed, failed, updated
* - **Mux** (4): created, killed, died, statsUpdated
* - **Remote auto-reconnect** (3): sessionDropped, sessionReconnected, reconnectExhausted
* - **Remote auto-reconnect / wake** (5): sessionDropped, sessionReconnected, reconnectExhausted, hostWaking, hostWakeFailed
* - **Respawn** (24): stateChanged, cycleStarted/Completed, step*, aiCheck*, planCheck*, timer*, log, ...
* - **Subagents** (7): discovered, updated, tool_call, tool_result, progress, message, completed
* - **Workflow runs** (3): run_discovered, run_updated, run_removed (ultracode / Workflow tool)
@@ -28,6 +28,7 @@
* - **Hooks** (10): idle_prompt, permission_prompt, elicitation_dialog, elicitation_complete, elicitation_response, stop, agent_working, teammate_idle, task_completed, prompt_submitted
* (agent_working is the odd one out: reported by the DeepSeek Harness status bridge, not by a Claude Code hook)
* - **Approvals** (3): pending, updated, resolved (cross-session Approvals Inbox)
* - **Custom Model Endpoint Profiles** (1): swapped-out (a session's model got evicted by another session on the same llama-swap endpoint)
* - **Orchestrator** (12): stateChanged, planProgress, planReady, phase*, verification, task*, completed, error
* - **Clipboard** (1): write
* - **Cases** (4): created, linked, deleted, order-changed
@@ -176,7 +177,9 @@ export const MuxDied = 'mux:died' as const;
/** tmux session stats refreshed. */
export const MuxStatsUpdated = 'mux:statsUpdated' as const;
// ─── Remote auto-reconnect (COD-108) ─────────────────────────────────────────
// ─── Remote auto-reconnect (COD-108) + wake-on-LAN ───────────────────────────
// Session-scoped in multi-user mode (`deriveSseHint`): routed to the session's owner,
// or — for a wake with no session yet — to the requesting `username` in the payload.
/** A remote session's local ssh pane died; an auto-reconnect attempt is starting. */
export const RemoteSessionDropped = 'remote:sessionDropped' as const;
@@ -184,6 +187,15 @@ export const RemoteSessionDropped = 'remote:sessionDropped' as const;
export const RemoteSessionReconnected = 'remote:sessionReconnected' as const;
/** Auto-reconnect gave up after the bounded backoff cap — manual reconnect needed. */
export const RemoteReconnectExhausted = 'remote:reconnectExhausted' as const;
/**
* User input arrived for a session whose host is unreachable, so a Wake-on-LAN
* command was started (see `remote-wake.ts`). Input sent meanwhile is buffered.
* Payload: `sessionId` (session wake) or `forNewSession: true` + `username`
* (create/attach wake), `hostId`, `label`, `queuedInput`.
*/
export const RemoteHostWaking = 'remote:hostWaking' as const;
/** The host did not come back within the wake timeout — buffered input is still held. */
export const RemoteHostWakeFailed = 'remote:hostWakeFailed' as const;
// ─── Respawn ─────────────────────────────────────────────────────────────────
@@ -384,6 +396,19 @@ export const ApprovalUpdated = 'approval:updated' as const;
/** A pending approval left the inbox (answered, superseded, expired, ...). */
export const ApprovalResolved = 'approval:resolved' as const;
// ─── Custom Model Endpoint Profiles ──────────────────────────────────────────
/**
* A session's own custom-model selection is no longer the model llama-swap has loaded —
* ANOTHER session's activity on the same endpoint evicted it (llama.cpp/llama-swap runs
* one model at a time). Detected after the fact by a periodic sweep (`server.ts`), never
* at the moment of eviction itself, since llama-swap has no push notification of its own;
* this session's next prompt will trigger reloading its model, evicting whatever displaced
* it in turn. Fires at most once per displacement (cleared once the sweep sees the
* session's own model loaded again), so it can't spam on every sweep interval.
*/
export const CustomModelSwappedOut = 'custom-model:swapped-out' as const;
// ─── Orchestrator ────────────────────────────────────────────────────────────
/** Orchestrator state machine transitioned. */
@@ -535,6 +560,8 @@ export const SseEvent = {
RemoteSessionDropped,
RemoteSessionReconnected,
RemoteReconnectExhausted,
RemoteHostWaking,
RemoteHostWakeFailed,
// Respawn
RespawnStarted,
@@ -638,6 +665,9 @@ export const SseEvent = {
ApprovalUpdated,
ApprovalResolved,
// Custom Model Endpoint Profiles
CustomModelSwappedOut,
// Orchestrator
OrchestratorStateChanged,
OrchestratorPlanProgress,
+533
View File
@@ -0,0 +1,533 @@
/**
* @fileoverview A capture drawn for a bigger pane makes the client replay once.
*
* A visible-frame capture repaints each row at an absolute position, counting
* up to the PANE's height and out to the PANE's width. A terminal shorter than
* that clamps every address past its own height onto its last line, so the
* overflow rows overwrite one another and the rows underneath are lost. A
* narrower terminal wraps every painted row, and the wrap on the last one
* scrolls the whole frame up by one. The client cannot see either from the
* escape sequence, so the terminal response reports the geometry the capture
* was taken at (`captureCols`/`captureRows`) and `selectSession` replays once
* at the size that stuck.
*
* The comparison runs on a `mux-visible` response ONLY. The other two sources
* position no rows absolutely, so a size mismatch damages neither and a replay
* repairs neither, and the last case here pins that the expensive one is left
* alone.
*
* These drive the REAL client in chromium and stub only the terminal endpoint,
* because the mismatch itself needs two viewports to stage against live tmux.
* Without the fix the first assertion below sees one fetch instead of two.
*
* Port: 3252 (capture geometry retry)
*
* Run: npx vitest run --config config/vitest.browser.config.ts test/capture-geometry-retry.browser.test.ts
*/
import { describe, it, expect, beforeAll, afterAll } from 'vitest';
import { chromium, type Browser, type BrowserContext, type Page } from 'playwright';
import { WebServer } from '../src/web/server.js';
const PORT = 3252;
const BASE_URL = `http://localhost:${PORT}`;
let server: WebServer;
let browser: Browser;
beforeAll(async () => {
server = new WebServer(PORT, false, true); // testMode
await server.start();
browser = await chromium.launch({ headless: true });
}, 60_000);
afterAll(async () => {
await browser?.close();
await server?.stop();
}, 30_000);
/** A visible-frame capture: one absolutely-addressed paint per row. */
function paneSnapshot(rows: number): string {
const parts: string[] = [];
for (let row = 1; row <= rows; row++) parts.push(`\x1b[${row};1Hprobe-row-${row}`);
parts.push(`\x1b[${rows};6H`);
return parts.join('');
}
/**
* Serve every terminal fetch from a stub reporting `captureRows`, counting the
* fetches. The real route needs live tmux to produce a mismatched frame.
*
* `source` is DERIVED from the request the way the real route derives it: a
* `full=1` request whose capture came back is `mux-full-history`, and every
* other one is `mux-visible`. The route cannot answer `full=1` with
* `mux-visible`, so a stub that did would stage a combination production never
* produces, and a test resting on it would prove nothing about production. A
* test that needs some other source passes it explicitly and says why.
*/
async function stubTerminal(
page: Page,
captureRows: number,
counter: { n: number; urls: string[] },
options: { source?: string; captureCols?: number } = {}
) {
const captureCols = options.captureCols ?? 200;
await page.route('**/api/sessions/*/terminal*', async (route) => {
const url = route.request().url();
counter.n += 1;
counter.urls.push(url);
const source = options.source ?? (url.includes('full=1') ? 'mux-full-history' : 'mux-visible');
await route.fulfill({
status: 200,
contentType: 'application/json',
body: JSON.stringify({
success: true,
data: {
terminalBuffer: paneSnapshot(captureRows),
status: 'idle',
fullSize: 1024,
retainedBytes: 1024,
truncated: false,
truncationReason: null,
source,
captureCols,
captureRows,
},
}),
});
});
}
/**
* As `stubTerminal`, but reading its geometry from a holder the test can change
* between selects. That is what lets one case watch a pane stop fitting and
* start fitting again, which a stub fixed at construction cannot show.
*/
async function stubTerminalDynamic(
page: Page,
counter: { n: number; urls: string[] },
state: { captureRows: number; captureCols: number }
) {
await page.route('**/api/sessions/*/terminal*', async (route) => {
const url = route.request().url();
counter.n += 1;
counter.urls.push(url);
await route.fulfill({
status: 200,
contentType: 'application/json',
body: JSON.stringify({
success: true,
data: {
terminalBuffer: paneSnapshot(state.captureRows),
status: 'idle',
fullSize: 1024,
retainedBytes: 1024,
truncated: false,
truncationReason: null,
source: url.includes('full=1') ? 'mux-full-history' : 'mux-visible',
captureCols: state.captureCols,
captureRows: state.captureRows,
},
}),
});
});
}
/**
* Answer every fetch with the geometry the client itself is asking for, read
* live from the page. That is the clamp signature: `getTerminalDimensions()`
* floors at 40x10 while `fitAddon.fit()` does not, so a small enough viewport
* makes the pane permanently bigger than the terminal at a size the client
* requested itself.
*/
async function stubTerminalAtRequestedSize(page: Page, counter: { n: number; urls: string[] }) {
await page.route('**/api/sessions/*/terminal*', async (route) => {
counter.n += 1;
counter.urls.push(route.request().url());
const dims = await page.evaluate(
() =>
(
window as unknown as { app: { getTerminalDimensions?: () => { cols: number; rows: number } | null } }
).app.getTerminalDimensions?.() ?? null
);
await route.fulfill({
status: 200,
contentType: 'application/json',
body: JSON.stringify({
success: true,
data: {
terminalBuffer: paneSnapshot(dims?.rows ?? 10),
status: 'idle',
fullSize: 1024,
retainedBytes: 1024,
truncated: false,
truncationReason: null,
source: 'mux-visible',
captureCols: dims?.cols,
captureRows: dims?.rows,
},
}),
});
});
}
/** The widest terminal this suite's 1280px viewport can produce, with margin. */
const WIDER_THAN_ANY_TERMINAL_COLS = 500;
async function openSession(page: Page): Promise<string> {
await page.goto(BASE_URL, { waitUntil: 'domcontentloaded' });
await page.waitForFunction(() => document.body.classList.contains('app-loaded'), { timeout: 10_000 });
// xterm is loaded from /vendor, so the terminal appears a beat after the app.
// Without it `app.terminal.rows` reads 0 and every height comparison below
// would pass vacuously.
await page.waitForFunction(() => (window as unknown as { app?: { terminal?: unknown } }).app?.terminal, null, {
timeout: 30_000,
});
return page.evaluate(async () => {
const res = await fetch('/api/sessions', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ workingDir: '/tmp', name: 'capture-geometry-test' }),
});
const body = await res.json();
return body.data?.session?.id ?? body.data?.id ?? body.id;
});
}
/** The terminal is sized by the first select, so this only reads after one. */
async function terminalRows(page: Page): Promise<number> {
return page.evaluate(() => (window as unknown as { app: { terminal?: { rows: number } } }).app.terminal?.rows ?? 0);
}
/** As above, for the width half of the comparison. */
async function terminalCols(page: Page): Promise<number> {
return page.evaluate(() => (window as unknown as { app: { terminal?: { cols: number } } }).app.terminal?.cols ?? 0);
}
async function select(page: Page, sessionId: string, options: object = {}): Promise<void> {
await page.evaluate(
async ({ sid, opts }) => {
const app = (window as unknown as { app: { selectSession: (id: string, o?: object) => Promise<void> } }).app;
await app.selectSession(sid, opts);
},
{ sid: sessionId, opts: options }
);
await page.waitForTimeout(1500);
}
/**
* Spend the per-page full-history allowance and forget what it cost. Every
* geometry comparison below runs on a `mux-visible` response, and the route
* only produces one for a request sent WITHOUT `full=1`, so reaching that shape
* means not being the first select of the page — which is what a tab switch is.
*/
async function consumeFullHistory(
page: Page,
sessionId: string,
counter: { n: number; urls: string[] }
): Promise<void> {
await select(page, sessionId);
counter.n = 0;
counter.urls.length = 0;
}
async function closeSession(page: Page, sessionId: string): Promise<void> {
await page.evaluate(
(sid: string) => fetch(`/api/sessions/${sid}`, { method: 'DELETE' }).then(() => undefined),
sessionId
);
}
describe('a capture bigger than the terminal', () => {
let context: BrowserContext;
let page: Page;
afterAll(async () => {
await context?.close();
});
it('replays once when the captured pane is taller, and stops at one retry', async () => {
context = await browser.newContext({ viewport: { width: 1280, height: 800 } });
page = await context.newPage();
const sessionId = await openSession(page);
expect(sessionId).toBeTruthy();
// 200 rows is taller than any terminal this viewport can produce, so the
// trigger is the captured height alone and not a size that moved.
const fetches = { n: 0, urls: [] as string[] };
await stubTerminal(page, 200, fetches);
// A tab switch is where a visible-frame response arrives, so that is what
// this measures. The first select of the page takes the full-history path
// and is covered by its own case below.
await consumeFullHistory(page, sessionId, fetches);
await select(page, sessionId, { forceReload: true });
// The terminal is sized by that select, so the premise is checkable now.
expect(await terminalRows(page)).toBeLessThan(200);
// One original load plus exactly one retry. `resizeRetry` caps it there:
// the retry's own response reports the same mismatch, so an uncapped
// implementation would loop.
expect(fetches.n).toBe(2);
await closeSession(page, sessionId);
await context.close();
}, 60_000);
it('retries at the same scope the first pass used, not a wider one', async () => {
// The retry re-arms the full-history flag only when the pass that ran had
// consumed it. A tab switch takes the bounded tail, so its retry must take
// the tail too; clearing the flag unconditionally would upgrade it into a
// fresh multi-megabyte scrollback capture the user never asked for.
context = await browser.newContext({ viewport: { width: 1280, height: 800 } });
page = await context.newPage();
const sessionId = await openSession(page);
const fetches = { n: 0, urls: [] as string[] };
await stubTerminal(page, 200, fetches);
// First select: a fresh session, so this one pulls full history. It does
// NOT retry, because the geometry comparison runs on a visible-frame
// response and a `full=1` request cannot produce one.
await select(page, sessionId);
expect(fetches.n).toBe(1);
expect(fetches.urls.filter((u) => u.includes('full=1'))).toHaveLength(1);
// Re-select the SAME session. `selectSession` early-returns on an already
// active session unless forceReload is set, and forceReload is the shape a
// tab switch back to this session takes: `_fullHistoryLoaded` still holds
// it, so neither this pass nor its retry asks for full history again.
await select(page, sessionId, { forceReload: true });
const tabSwitchUrls = fetches.urls.slice(1);
expect(tabSwitchUrls.length).toBe(2);
expect(tabSwitchUrls.filter((u) => u.includes('full=1'))).toHaveLength(0);
await closeSession(page, sessionId);
await context.close();
}, 60_000);
it('does not replay when the captured pane fits the terminal', async () => {
context = await browser.newContext({ viewport: { width: 1280, height: 800 } });
page = await context.newPage();
const sessionId = await openSession(page);
// Five rows is shorter than any terminal this viewport can produce, so the
// frame fits, nothing is clamped, and nothing needs repeating. A retry here
// would double the work of every tab switch.
const fetches = { n: 0, urls: [] as string[] };
await stubTerminal(page, 5, fetches, { captureCols: 40 });
await consumeFullHistory(page, sessionId, fetches);
await select(page, sessionId, { forceReload: true });
expect(await terminalRows(page)).toBeGreaterThan(5);
expect(fetches.n).toBe(1);
await closeSession(page, sessionId);
await context.close();
}, 60_000);
it('replays once when the captured pane is wider', async () => {
// A pane wider than the terminal damages the same frame a second way.
// `formatPaneSnapshot` paints every row out to the PANE's width, so a
// narrower browser wraps each painted row, and the wrap on the last row
// scrolls the whole frame up by one. The height here fits deliberately, so
// the width is the only thing that can trigger the replay.
context = await browser.newContext({ viewport: { width: 1280, height: 800 } });
page = await context.newPage();
const sessionId = await openSession(page);
const fetches = { n: 0, urls: [] as string[] };
await stubTerminal(page, 5, fetches, { captureCols: WIDER_THAN_ANY_TERMINAL_COLS });
await consumeFullHistory(page, sessionId, fetches);
await select(page, sessionId, { forceReload: true });
expect(await terminalRows(page)).toBeGreaterThan(5);
expect(await terminalCols(page)).toBeLessThan(WIDER_THAN_ANY_TERMINAL_COLS);
expect(fetches.n).toBe(2);
await closeSession(page, sessionId);
await context.close();
}, 60_000);
it('does not replay a full-history response, whatever geometry it reports', async () => {
// A `full=1` body is linear scrollback closed by a RELATIVE cursor move,
// which is relative precisely so the browser's row count need not match the
// pane's. A mismatch there is not damage and a replay cannot repair it, so
// the geometry comparison must not fire on it. This is the path that makes
// the gate worth having: `_fullHistoryLoaded` is empty on the first select
// of every non-shell session per page, so an ungated comparison would pull
// the entire tmux scrollback a second time on every page load and every
// first tab switch, for a session whose pane a desktop tab is holding too
// tall to ever fit.
context = await browser.newContext({ viewport: { width: 1280, height: 800 } });
page = await context.newPage();
const sessionId = await openSession(page);
const fetches = { n: 0, urls: [] as string[] };
await stubTerminal(page, 200, fetches, { captureCols: WIDER_THAN_ANY_TERMINAL_COLS });
await select(page, sessionId);
// Both dimensions are mismatched, so height alone is not what spares it.
expect(await terminalRows(page)).toBeLessThan(200);
expect(await terminalCols(page)).toBeLessThan(WIDER_THAN_ANY_TERMINAL_COLS);
expect(fetches.n).toBe(1);
expect(fetches.urls.filter((u) => u.includes('full=1'))).toHaveLength(1);
await closeSession(page, sessionId);
await context.close();
}, 60_000);
it('replays once per session, not once per tab switch, when it cannot converge', async () => {
// `resizeRetry` caps the recursion inside ONE select and says nothing about
// the next one, so a pane this browser cannot size reported the same
// mismatch on every select and bought the same failed repair every time:
// two fetches per tab switch for the life of the page. That is the case the
// description calls "every time rather than occasionally", a phone whose
// resize is declined while a desktop claim is live, and it is not the only
// one — any pane Codeman cannot size lands there, a second tmux client
// attached to it included. Each wasted pass costs another `capture-pane`,
// which is `execSync` on the server's event loop, plus a reset and rewrite,
// a discarded snapshot, and a dropped and reopened WebSocket.
context = await browser.newContext({ viewport: { width: 1280, height: 800 } });
page = await context.newPage();
const sessionId = await openSession(page);
const fetches = { n: 0, urls: [] as string[] };
const pane = { captureRows: 200, captureCols: 200 };
await stubTerminalDynamic(page, fetches, pane);
await consumeFullHistory(page, sessionId, fetches);
// First tab switch: one load, one replay, and the replay does not fit
// either, which is the proof that this pane ignores the size it is given.
await select(page, sessionId, { forceReload: true });
expect(fetches.n).toBe(2);
// Every switch after it pays once. Unlatched this reads 4 then 6.
await select(page, sessionId, { forceReload: true });
expect(fetches.n).toBe(3);
await select(page, sessionId, { forceReload: true });
expect(fetches.n).toBe(4);
// The memo has to lift when the pane becomes sizeable again, or closing the
// desktop tab that was holding it would leave this session permanently
// unrepaired. A frame that fits clears it...
pane.captureRows = 5;
pane.captureCols = 40;
await select(page, sessionId, { forceReload: true });
expect(fetches.n).toBe(5);
// ...so the next genuine mismatch is diagnosed again.
pane.captureRows = 200;
pane.captureCols = 200;
await select(page, sessionId, { forceReload: true });
expect(fetches.n).toBe(7);
await closeSession(page, sessionId);
await context.close();
}, 60_000);
it('hands over text typed but not yet submitted before it replays', async () => {
// On a touch device the characters the user has typed live ONLY in the
// local-echo overlay until Enter; they have never reached the PTY. The
// replay re-enters `selectSession` with `forceReload` on the session that
// is still active, and that branch used to null `activeSessionId` before
// `_cleanupPreviousSession` ran, so the flush there saw no session and the
// unconditional `clear()` afterwards took the characters with it. Nothing
// the user did triggered that: the replay fires on its own the moment a
// tab switch finishes, which is exactly when someone typing into a
// still-loading terminal has text in the overlay.
context = await browser.newContext({ viewport: { width: 1280, height: 800 } });
page = await context.newPage();
const sessionId = await openSession(page);
const fetches = { n: 0, urls: [] as string[] };
await stubTerminal(page, 200, fetches);
await consumeFullHistory(page, sessionId, fetches);
// Headless chromium reports `isTouchDevice()` false even with `hasTouch`,
// so the overlay would stay off and the whole case would pass vacuously.
// The setting is what `_updateLocalEchoState()` reads, so it survives the
// recompute that every select runs; the flag is forced too, for the window
// before the next recompute. Record what crosses into the delivery layer,
// which is the seam the text failed to cross.
await page.evaluate(() => {
const w = window as unknown as {
app: {
_localEchoEnabled: boolean;
_sendInputAsync: (id: string, text: string, opts?: unknown) => void;
terminal?: { focus: () => void };
loadAppSettingsFromStorage: () => Record<string, unknown>;
};
__sentInputs: { id: string; text: string }[];
};
const settings = w.app.loadAppSettingsFromStorage();
settings.localEchoEnabled = true;
localStorage.setItem('codeman-app-settings', JSON.stringify(settings));
w.app._localEchoEnabled = true;
w.__sentInputs = [];
const original = w.app._sendInputAsync.bind(w.app);
w.app._sendInputAsync = (id: string, text: string, opts?: unknown) => {
w.__sentInputs.push({ id, text });
return original(id, text, opts);
};
w.app.terminal?.focus();
});
await page.keyboard.type('hello-unsent');
// The premise: the characters really are sitting in the overlay, unsent.
// Without this the case would pass on a build where typing goes straight
// to the PTY and there is nothing to lose.
const pendingBefore = await page.evaluate(
() =>
(window as unknown as { app: { _localEchoOverlay?: { pendingText: string } } }).app._localEchoOverlay
?.pendingText ?? ''
);
expect(pendingBefore).toBe('hello-unsent');
// The captured pane is taller than the terminal, so this select replays.
await select(page, sessionId, { forceReload: true });
expect(fetches.n).toBe(2);
const sent = await page.evaluate(
() => (window as unknown as { __sentInputs: { id: string; text: string }[] }).__sentInputs
);
expect(sent.map((s) => s.text)).toContain('hello-unsent');
expect(sent.find((s) => s.text === 'hello-unsent')?.id).toBe(sessionId);
await closeSession(page, sessionId);
await context.close();
}, 60_000);
it('does not replay a pane already at the size the client asked for', async () => {
// `getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` does
// not, so a viewport this small leaves the terminal shorter than the size
// the client itself requests, and the pane obligingly draws at the floored
// size. The captured height then exceeds the terminal's forever. A replay
// cannot converge, because it re-requests the same floored size and
// captures the same frame, so without the equality guard this retries on
// every tab switch for the life of the page.
context = await browser.newContext({ viewport: { width: 320, height: 200 } });
page = await context.newPage();
const sessionId = await openSession(page);
const fetches = { n: 0, urls: [] as string[] };
await stubTerminalAtRequestedSize(page, fetches);
await consumeFullHistory(page, sessionId, fetches);
await select(page, sessionId, { forceReload: true });
// The premise: the floor really does bind here. Without this the case
// would pass on any viewport, proving nothing.
const requested = await page.evaluate(
() =>
(
window as unknown as { app: { getTerminalDimensions?: () => { cols: number; rows: number } | null } }
).app.getTerminalDimensions?.() ?? null
);
expect(requested).not.toBeNull();
expect(requested!.rows).toBeGreaterThan(await terminalRows(page));
expect(fetches.n).toBe(1);
await closeSession(page, sessionId);
await context.close();
}, 60_000);
});
+180
View File
@@ -0,0 +1,180 @@
/**
* @fileoverview Output arriving after a pane capture survives the buffer load.
*
* `batchTerminalWrite` queues live terminal events while a buffer load runs,
* and `_finishBufferLoad` discards that queue by default. That is right when
* the loaded buffer is the server's accumulated byte history, which is current
* up to the response. A tmux pane capture is current only up to CAPTURE time,
* so anything arriving between the capture and the end of the chunked write is
* queued and then dropped, with nothing scheduling a re-fetch.
*
* The queue now stamps each entry with its arrival time, and a capture load
* replays the tail that arrived after the response headers. These drive the
* real client in chromium: the event is injected from inside the response's
* own `json()` call, which is the one place guaranteed to land after the
* headers and before the chunked write.
*
* Port: 3256 (capture load window)
*
* Run: npx vitest run --config config/vitest.browser.config.ts test/capture-load-window.browser.test.ts
*/
import { describe, it, expect, beforeAll, afterAll } from 'vitest';
import { chromium, type Browser, type BrowserContext, type Page } from 'playwright';
import { WebServer } from '../src/web/server.js';
const PORT = 3256;
const BASE_URL = `http://localhost:${PORT}`;
const MARKER = 'ARRIVED-AFTER-THE-CAPTURE';
let server: WebServer;
let browser: Browser;
beforeAll(async () => {
server = new WebServer(PORT, false, true); // testMode
await server.start();
browser = await chromium.launch({ headless: true });
}, 60_000);
afterAll(async () => {
await browser?.close();
await server?.stop();
}, 30_000);
/**
* Select the session with the terminal fetch stubbed, injecting one live event
* from inside `json()`. Returns how many terminal rows carry the marker, so a
* flush that replays too much fails as loudly as one that replays nothing.
*/
async function runLoad(page: Page, sessionId: string, source: string): Promise<number> {
return page.evaluate(
async ({ sid, src, marker }) => {
const app = (
window as unknown as {
app: {
selectSession: (id: string, o?: object) => Promise<void>;
_onSessionTerminal: (e: { id: string; data: string }) => void;
terminal: {
buffer: {
active: {
length: number;
getLine: (i: number) => { translateToString: (t: boolean) => string } | undefined;
};
};
};
};
}
).app;
const realFetch = window.fetch.bind(window);
window.fetch = ((input: RequestInfo | URL, init?: RequestInit) => {
const url = String(typeof input === 'string' ? input : ((input as Request).url ?? input));
if (!url.includes('/terminal')) return realFetch(input as RequestInfo, init);
return Promise.resolve({
ok: true,
status: 200,
// `selectSession` timestamps the headers the moment this promise
// resolves, then calls json(). Injecting here puts the event after
// that timestamp and inside the load window, which is exactly the
// gap a pane capture cannot cover.
json: async () => {
app._onSessionTerminal({ id: sid, data: `\r\n${marker}\r\n` });
return {
success: true,
data: {
terminalBuffer: '\x1b[1;1Hcaptured frame line one\r\n',
status: 'idle',
fullSize: 512,
retainedBytes: 512,
truncated: false,
truncationReason: null,
source: src,
captureCols: 80,
captureRows: 24,
},
};
},
}) as unknown as Promise<Response>;
}) as typeof window.fetch;
try {
await app.selectSession(sid);
await new Promise((r) => setTimeout(r, 1200));
const buf = app.terminal.buffer.active;
let hits = 0;
for (let i = 0; i < buf.length; i++) {
if (buf.getLine(i)?.translateToString(true).includes(marker)) hits += 1;
}
return hits;
} finally {
window.fetch = realFetch;
}
},
{ sid: sessionId, src: source, marker: MARKER }
);
}
async function openSession(page: Page): Promise<string> {
await page.goto(BASE_URL, { waitUntil: 'domcontentloaded' });
await page.waitForFunction(() => document.body.classList.contains('app-loaded'), { timeout: 10_000 });
// xterm loads from /vendor, so the terminal appears a beat after the app.
// Without it every buffer assertion below would throw rather than compare.
await page.waitForFunction(() => (window as unknown as { app?: { terminal?: unknown } }).app?.terminal, null, {
timeout: 30_000,
});
return page.evaluate(async () => {
const res = await fetch('/api/sessions', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ workingDir: '/tmp', name: 'capture-load-window-test' }),
});
const body = await res.json();
return body.data?.session?.id ?? body.data?.id ?? body.id;
});
}
describe('output emitted during a capture load', () => {
let context: BrowserContext;
let page: Page;
afterAll(async () => {
await context?.close();
});
it('reaches the terminal exactly once when the buffer came from a pane capture', async () => {
context = await browser.newContext({ viewport: { width: 1280, height: 800 } });
page = await context.newPage();
const sessionId = await openSession(page);
expect(sessionId).toBeTruthy();
// Exactly once. The cutoff exists so the flush cannot also replay events the
// payload already carried, which would double the output rather than heal it.
expect(await runLoad(page, sessionId, 'mux-visible')).toBe(1);
await page.evaluate(
(sid: string) => fetch(`/api/sessions/${sid}`, { method: 'DELETE' }).then(() => undefined),
sessionId
);
await context.close();
}, 60_000);
it('stays dropped when the buffer came from the accumulated byte history', async () => {
// The byte history already contains everything up to the response, so
// replaying the queue on top of it would duplicate the output — most
// visibly Ink's cursor-up redraws. The discard has to survive this fix.
context = await browser.newContext({ viewport: { width: 1280, height: 800 } });
page = await context.newPage();
const sessionId = await openSession(page);
// Without this, a failed create passes the zero-hit assertion below
// vacuously — nothing was loaded, so nothing was replayed.
expect(sessionId).toBeTruthy();
expect(await runLoad(page, sessionId, 'history')).toBe(0);
await page.evaluate(
(sid: string) => fetch(`/api/sessions/${sid}`, { method: 'DELETE' }).then(() => undefined),
sessionId
);
await context.close();
}, 60_000);
});
@@ -0,0 +1,368 @@
/**
* @fileoverview Tests for `refreshAllCustomModelHosts()`, the periodic
* background sweep behind server.ts's "custom model endpoint re-discovery"
* timer (docs/custom-model-endpoints-plan.md). Kept in its own file rather
* than folded into test/routes/custom-model-routes.test.ts: that file's data
* dir is shared across every test in it (one temp HOME per FILE, not per
* test — test/setup.ts), and a sweep that walks every saved host would pick
* up every host any other test in that file happened to create, making an
* exact call-count or exact-host assertion meaningless. A dedicated file
* gets its own clean temp HOME.
*
* Port: N/A (no server; drives readCustomModelHosts/writeCustomModelHosts
* directly plus the mocked webviewFetch dispatcher).
*/
import { describe, it, expect, vi, beforeEach } from 'vitest';
import { getDataDir } from '../src/config/instance.js';
import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } from '../src/custom-model-hosts.js';
import { refreshAllCustomModelHosts } from '../src/web/routes/custom-model-routes.js';
import { webviewFetch } from '../src/web/webview-egress.js';
vi.mock('../src/web/webview-egress.js', async () => {
const actual = await vi.importActual<typeof import('../src/web/webview-egress.js')>('../src/web/webview-egress.js');
return { ...actual, webviewFetch: vi.fn() };
});
const fetchMock = vi.mocked(webviewFetch);
function host(overrides: Partial<CustomModelHost> & Pick<CustomModelHost, 'id' | 'baseUrl'>): CustomModelHost {
return { label: overrides.id, ...overrides };
}
beforeEach(() => {
fetchMock.mockReset();
});
describe('refreshAllCustomModelHosts (the periodic re-discovery sweep)', () => {
it('refreshes every saved endpoint, best-effort — one unreachable host does not stop the others', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [
host({ id: 'ok', baseUrl: 'http://localhost:8080' }),
host({ id: 'down', baseUrl: 'http://localhost:8081' }),
]);
fetchMock.mockImplementation(async (url: URL) => {
if (url.href.includes('8081')) throw new TypeError('fetch failed', { cause: new Error('ECONNREFUSED') });
return new Response(JSON.stringify({ data: [{ id: 'qwen3' }] }), { status: 200 });
});
await refreshAllCustomModelHosts();
const hosts = await readCustomModelHosts(dir);
const ok = hosts.find((h) => h.id === 'ok');
const down = hosts.find((h) => h.id === 'down');
expect(ok?.models).toEqual(['qwen3']);
expect(ok?.lastDiscoveredAt).toBeTruthy();
expect(down?.models ?? []).toEqual([]);
expect(down?.lastDiscoveredAt).toBeFalsy();
});
it('skips a host whose baseUrl is blocked, without making a request', async () => {
const dir = getDataDir();
// Written directly rather than through the POST route, which already
// refuses this at save time — this simulates a record that pre-dates the
// guard, or was hand-edited on disk. The sweep must not trust it either.
await writeCustomModelHosts(dir, [host({ id: 'meta', baseUrl: 'http://169.254.169.254/' })]);
fetchMock.mockResolvedValue(new Response(JSON.stringify({ data: [{ id: 'x' }] }), { status: 200 }));
await refreshAllCustomModelHosts();
expect(fetchMock).not.toHaveBeenCalled();
});
it('drops a stale default and preserves lastDiscoveredAt semantics, same as manual discovery', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [
host({ id: 'ep', baseUrl: 'http://localhost:8080', models: ['qwen3'], defaultModelId: 'qwen3' }),
]);
fetchMock.mockResolvedValue(new Response(JSON.stringify({ data: [{ id: 'llama3' }] }), { status: 200 }));
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.models).toEqual(['llama3']);
expect(updated.defaultModelId).toBeUndefined();
expect(updated.lastDiscoveredAt).toBeTruthy();
});
it('keeps a default that is still present after the sweep', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [
host({ id: 'ep', baseUrl: 'http://localhost:8080', models: ['qwen3'], defaultModelId: 'qwen3' }),
]);
fetchMock.mockResolvedValue(
new Response(JSON.stringify({ data: [{ id: 'qwen3' }, { id: 'llama3' }] }), { status: 200 })
);
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.defaultModelId).toBe('qwen3');
});
it('does not resurrect an endpoint deleted while the sweep was in flight', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [host({ id: 'deleted', baseUrl: 'http://localhost:8080' })]);
fetchMock.mockImplementation(async () => {
// Simulate an admin deleting the endpoint between the sweep's fetch and
// its read-modify-write — the delete must win, not be overwritten by a
// refresh that started before it.
const current = await readCustomModelHosts(dir);
await writeCustomModelHosts(
dir,
current.filter((h) => h.id !== 'deleted')
);
return new Response(JSON.stringify({ data: [{ id: 'qwen3' }] }), { status: 200 });
});
await expect(refreshAllCustomModelHosts()).resolves.toBeUndefined();
const hosts = await readCustomModelHosts(dir);
expect(hosts.find((h) => h.id === 'deleted')).toBeUndefined();
});
it('leaves the store untouched when there are no saved endpoints at all', async () => {
await expect(refreshAllCustomModelHosts()).resolves.toBeUndefined();
expect(fetchMock).not.toHaveBeenCalled();
});
});
describe('refreshAllCustomModelHosts: context-length enrichment (llama.cpp/llama-swap /props)', () => {
it('probes /props?model= only for a model reported loaded, and stores its n_ctx', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]);
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/v1/models') {
return new Response(
JSON.stringify({
data: [
{ id: 'loaded-model', status: { value: 'loaded' } },
{ id: 'unloaded-model', status: { value: 'unloaded' } },
],
}),
{ status: 200 }
);
}
if (url.pathname === '/props') {
// Must never be reached for the unloaded model — asserted below by call count.
expect(url.searchParams.get('model')).toBe('loaded-model');
return new Response(JSON.stringify({ n_ctx: 16384 }), { status: 200 });
}
throw new Error(`unexpected request: ${url.href}`);
});
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.modelContextLengths).toEqual({ 'loaded-model': 16384 });
const propsCalls = fetchMock.mock.calls.filter(([url]) => (url as URL).pathname === '/props');
expect(propsCalls).toHaveLength(1);
});
it('never probes /props at all when no entry mentions status — feature-detected, not assumed unloaded', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]);
fetchMock.mockResolvedValue(new Response(JSON.stringify({ data: [{ id: 'qwen3' }] }), { status: 200 }));
await refreshAllCustomModelHosts();
expect(fetchMock).toHaveBeenCalledTimes(1); // /v1/models only
const [updated] = await readCustomModelHosts(dir);
expect(updated.modelContextLengths).toBeUndefined();
});
it('keeps a previously-learned context length for a model no longer loaded, drops it once the model disappears entirely', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [
host({
id: 'ep',
baseUrl: 'http://localhost:8080',
models: ['a', 'b'],
modelContextLengths: { a: 8192, b: 4096 },
}),
]);
// This round: 'a' is loaded (re-confirmed), 'b' is gone from the list entirely.
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/v1/models') {
return new Response(JSON.stringify({ data: [{ id: 'a', status: { value: 'loaded' } }] }), { status: 200 });
}
return new Response(JSON.stringify({ n_ctx: 8192 }), { status: 200 });
});
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.modelContextLengths).toEqual({ a: 8192 });
});
it('a failed /props probe for the loaded model is swallowed, leaving no context length rather than failing the sweep', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]);
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/v1/models') {
return new Response(JSON.stringify({ data: [{ id: 'a', status: { value: 'loaded' } }] }), { status: 200 });
}
return new Response('nope', { status: 500 });
});
await expect(refreshAllCustomModelHosts()).resolves.toBeUndefined();
const [updated] = await readCustomModelHosts(dir);
expect(updated.modelContextLengths).toBeUndefined();
});
it('prefers the REAL configured context size parsed from /running’s launch command over /props’s unreliable n_ctx', async () => {
// Confirmed live: llama-swap launched a model with --fit-ctx 16384 (the real, working
// limit — the actual server then refused a request over it), but /props reported
// n_ctx: 154112 for the same model, well over what it would really accept. /props must
// never be reached at all once the /running command parse already answered it.
const dir = getDataDir();
await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]);
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/v1/models') {
return new Response(JSON.stringify({ data: [{ id: 'qwen3.8-27b', status: { value: 'loaded' } }] }), {
status: 200,
});
}
if (url.pathname === '/running') {
return new Response(
JSON.stringify({
running: [
{
model: 'qwen3.8-27b',
state: 'ready',
cmd: 'llama-server -m /models/Qwen3.8-27B.gguf --flash-attn on --jinja --fit-ctx 16384 --host 0.0.0.0 --port 5840',
},
],
}),
{ status: 200 }
);
}
if (url.pathname === '/props') throw new Error('must never be reached — the cmd parse already answered it');
throw new Error(`unexpected request: ${url.href}`);
});
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.modelContextLengths).toEqual({ 'qwen3.8-27b': 16384 });
});
it('falls back to /props when /running has no cmd, or the cmd states no recognizable context flag', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]);
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/v1/models') {
return new Response(JSON.stringify({ data: [{ id: 'a', status: { value: 'loaded' } }] }), { status: 200 });
}
if (url.pathname === '/running') {
return new Response(
JSON.stringify({ running: [{ model: 'a', state: 'ready', cmd: 'llama-server -m /models/a.gguf' }] }),
{ status: 200 }
);
}
if (url.pathname === '/props') return new Response(JSON.stringify({ n_ctx: 8192 }), { status: 200 });
throw new Error(`unexpected request: ${url.href}`);
});
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.modelContextLengths).toEqual({ a: 8192 });
});
it('also recognizes a plain -c/--ctx-size flag, not just llama-swap’s own --fit-ctx', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]);
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/v1/models') {
return new Response(JSON.stringify({ data: [{ id: 'a', status: { value: 'loaded' } }] }), { status: 200 });
}
if (url.pathname === '/running') {
return new Response(
JSON.stringify({
running: [{ model: 'a', state: 'ready', cmd: 'llama-server -m /models/a.gguf --ctx-size 8192' }],
}),
{ status: 200 }
);
}
throw new Error(`unexpected request: ${url.href}`); // /props must never be reached
});
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.modelContextLengths).toEqual({ a: 8192 });
});
});
describe('refreshAllCustomModelHosts: model-size enrichment (parsed from /v1/models description)', () => {
it('parses a GB figure out of an auto-discovered model’s description', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]);
fetchMock.mockResolvedValue(
new Response(
JSON.stringify({
data: [{ id: 'qwen3.8-27b', description: 'Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp' }],
}),
{ status: 200 }
)
);
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.modelSizesGB).toEqual({ 'qwen3.8-27b': 16.35 });
});
it('gets no size at all for a hand-configured profile whose own description states none', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]);
fetchMock.mockResolvedValue(
new Response(
JSON.stringify({
data: [{ id: 'big', description: 'General-purpose reasoning model, MoE CPU-offloaded. Default profile.' }],
}),
{ status: 200 }
)
);
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.modelSizesGB).toBeUndefined();
});
it('populated regardless of loaded state — unlike context length, no /props probe is needed', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [host({ id: 'ep', baseUrl: 'http://localhost:8080' })]);
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/v1/models') {
return new Response(
JSON.stringify({
data: [{ id: 'unloaded-model', description: 'Auto-discovered 4.91 GB - parameters auto-fitted' }],
}),
{ status: 200 }
);
}
throw new Error(`unexpected request: ${url.href}`); // /props must never be reached for this
});
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.modelSizesGB).toEqual({ 'unloaded-model': 4.91 });
});
it('keeps a previously-learned size for a model still present, drops it once the model disappears entirely', async () => {
const dir = getDataDir();
await writeCustomModelHosts(dir, [
host({ id: 'ep', baseUrl: 'http://localhost:8080', models: ['a', 'b'], modelSizesGB: { a: 8, b: 16 } }),
]);
fetchMock.mockResolvedValue(
new Response(JSON.stringify({ data: [{ id: 'a', description: 'no GB figure here' }] }), { status: 200 })
);
await refreshAllCustomModelHosts();
const [updated] = await readCustomModelHosts(dir);
expect(updated.modelSizesGB).toEqual({ a: 8 }); // 'a' kept from before, 'b' dropped (gone from the list)
});
});
+331
View File
@@ -0,0 +1,331 @@
/**
* @fileoverview Tests for the two custom-model IO-layer fixes on top of the pure builder
* (docs/custom-model-endpoints-plan.md):
*
* 1. `contextLengthVar` — a discovered per-model context length reaches the actual
* session env (CLAUDE_CODE_MAX_CONTEXT_TOKENS), so a CLI stops assuming a large
* default window for an unrecognized custom model id and overflowing a much
* smaller real one.
* 2. `configDirVar` — an isolated, empty config directory is created and pointed at
* (CLAUDE_CONFIG_DIR), so an injected API key never shares a directory with a
* stored claude.ai OAuth session; `projects` is symlinked back into the real
* config dir so the response viewer/subagent windows/Read My Mind keep working.
*
* Port: N/A (no server; filesystem-only, under a temp CODEMAN data dir from test/setup.ts).
*/
import { existsSync, lstatSync, mkdirSync, readFileSync, readdirSync, rmSync, writeFileSync } from 'node:fs';
import { homedir } from 'node:os';
import { join } from 'node:path';
import { afterEach, describe, expect, it } from 'vitest';
import { getCli } from '../src/config/cli-registry/index.js';
import { applyCustomModelInjection, customModelConfigDir } from '../src/custom-model-injection-apply.js';
import type { CustomModelEndpoint } from '../src/custom-model-injection.js';
const endpoint: CustomModelEndpoint = {
id: 'ep1',
label: 'llama.cpp box',
baseUrl: 'http://192.168.1.50:8080',
apiKey: 'my-key',
};
function entryOrThrow(id: string) {
const entry = getCli(id);
if (!entry) throw new Error(`missing CLI registry entry: ${id}`);
return entry;
}
const sessionsToClean: string[] = [];
afterEach(() => {
for (const id of sessionsToClean.splice(0)) rmSync(customModelConfigDir(id), { recursive: true, force: true });
});
describe('applyCustomModelInjection: context length', () => {
it('claude: passes a known context length through to CLAUDE_CODE_MAX_CONTEXT_TOKENS', () => {
sessionsToClean.push('sess-ctx-1');
const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', 'sess-ctx-1', 16384);
expect(applied?.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBe('16384');
expect(applied?.envKeys).toContain('CLAUDE_CODE_MAX_CONTEXT_TOKENS');
});
it('claude: omits the var entirely when the context length is unknown', () => {
sessionsToClean.push('sess-ctx-2');
const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', 'sess-ctx-2');
expect(applied?.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBeUndefined();
});
it('deepseek: has no contextLengthVar declared, so a passed-in length is a no-op', () => {
const applied = applyCustomModelInjection(entryOrThrow('deepseek'), endpoint, 'qwen3', 'sess-ctx-3', 16384);
expect(Object.keys(applied?.envOverrides ?? {}).sort()).toEqual(['DEEPSEEK_API_KEY', 'DEEPSEEK_BASE_URL']);
});
});
describe('applyCustomModelInjection: CLAUDE_CONFIG_DIR isolation', () => {
it('claude: creates an isolated config dir (no real credential/config files) and points CLAUDE_CONFIG_DIR at it', () => {
const sessionId = 'sess-cfgdir-1';
sessionsToClean.push(sessionId);
const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId);
const expectedDir = customModelConfigDir(sessionId);
expect(applied?.envOverrides.CLAUDE_CONFIG_DIR).toBe(expectedDir);
expect(applied?.configDir).toBe(expectedDir);
expect(existsSync(expectedDir)).toBe(true);
// The trust-seed file, the skipFirstRunPrompts settings.json, and the projects link —
// no real OAuth credential/config.
const entries = readdirSync(expectedDir).filter((name) => name !== 'projects');
expect(entries.sort()).toEqual(['.claude.json', 'settings.json']);
});
it('claude: symlinks (or junctions) projects back to the real config dir so the response viewer keeps working', () => {
const sessionId = 'sess-cfgdir-2';
sessionsToClean.push(sessionId);
const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId);
const link = join(applied!.configDir!, 'projects');
// Best-effort: only assert the link exists if it was actually created (the real
// ~/.claude/projects may not exist on a bare CI box, in which case linking is skipped).
if (existsSync(join(homedir(), '.claude', 'projects'))) {
expect(existsSync(link)).toBe(true);
expect(lstatSync(link).isSymbolicLink() || lstatSync(link).isDirectory()).toBe(true);
}
});
it('claude: re-applying to the same session is idempotent (boot-recovery re-apply)', () => {
const sessionId = 'sess-cfgdir-3';
sessionsToClean.push(sessionId);
const first = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId);
const second = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId);
expect(second?.configDir).toBe(first?.configDir);
expect(existsSync(first!.configDir!)).toBe(true);
});
it('pi: configDir-kind CLIs are unaffected — no configDirVar concept for them', () => {
const sessionId = 'sess-cfgdir-pi';
sessionsToClean.push(sessionId);
const applied = applyCustomModelInjection(entryOrThrow('pi'), endpoint, 'qwen3', sessionId);
expect(applied?.envOverrides.HOME).toBe(customModelConfigDir(sessionId));
});
it('deepseek: no configDirVar declared, so no config dir is created at all', () => {
const sessionId = 'sess-cfgdir-deepseek';
const applied = applyCustomModelInjection(entryOrThrow('deepseek'), endpoint, 'qwen3', sessionId);
expect(applied?.configDir).toBeUndefined();
expect(existsSync(customModelConfigDir(sessionId))).toBe(false);
});
});
describe('applyCustomModelInjection: apiKeyTrustFile (pre-approves the injected key)', () => {
it('claude: seeds .claude.json so the "Detected a custom API key" prompt never fires', () => {
const sessionId = 'sess-trust-1';
sessionsToClean.push(sessionId);
const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId);
const written = JSON.parse(readFileSync(join(applied!.configDir!, '.claude.json'), 'utf8')) as {
customApiKeyResponses: { approved: string[]; rejected: string[] };
};
expect(written.customApiKeyResponses.approved).toEqual(['my-key']);
expect(written.customApiKeyResponses.rejected).toEqual([]);
});
// ⚠ Claude Code stores and looks up only the LAST 20 CHARACTERS of a key
// (`key.trim().slice(-20)`, applied on both write and read), so seeding the whole key
// never matches for a REAL one and the launch stops at the interactive "Detected a
// custom API key" prompt whose default is "No (recommended)". Every other test here
// uses a key shorter than 20 characters, where slice(-20) is the whole string and the
// bug is invisible, which is exactly how it survived review.
it('claude: seeds a REAL-length key in the truncated form the CLI actually matches on', () => {
const sessionId = 'sess-trust-long';
sessionsToClean.push(sessionId);
const longKey = 'sk-or-v1-0123456789abcdef0123456789abcdef0123456789abcdef';
expect(longKey.length).toBeGreaterThan(20);
const applied = applyCustomModelInjection(
entryOrThrow('claude'),
{ ...endpoint, apiKey: longKey },
'qwen3',
sessionId
);
const written = JSON.parse(readFileSync(join(applied!.configDir!, '.claude.json'), 'utf8')) as {
customApiKeyResponses: { approved: string[] };
};
expect(written.customApiKeyResponses.approved).toEqual(['cdef0123456789abcdef']);
expect(written.customApiKeyResponses.approved[0]).toHaveLength(20);
// and the full credential is not written into this second file at all
expect(readFileSync(join(applied!.configDir!, '.claude.json'), 'utf8')).not.toContain(longKey);
});
it('claude: falls back to the dummy key when the endpoint has none, and still seeds it', () => {
const sessionId = 'sess-trust-2';
sessionsToClean.push(sessionId);
const applied = applyCustomModelInjection(
entryOrThrow('claude'),
{ ...endpoint, apiKey: undefined },
'qwen3',
sessionId
);
const written = JSON.parse(readFileSync(join(applied!.configDir!, '.claude.json'), 'utf8')) as {
customApiKeyResponses: { approved: string[] };
};
expect(written.customApiKeyResponses.approved).toEqual(['local-dummy-key']);
});
it('claude: merges onto fields the CLI itself already wrote into the same isolated dir, never overwrites them', () => {
const sessionId = 'sess-trust-3';
sessionsToClean.push(sessionId);
const configDir = customModelConfigDir(sessionId);
mkdirSync(configDir, { recursive: true });
writeFileSync(join(configDir, '.claude.json'), JSON.stringify({ userID: 'abc123', numStartups: 3 }));
const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId);
const written = JSON.parse(readFileSync(join(applied!.configDir!, '.claude.json'), 'utf8')) as {
userID: string;
numStartups: number;
customApiKeyResponses: { approved: string[] };
};
expect(written.userID).toBe('abc123');
expect(written.numStartups).toBe(3);
expect(written.customApiKeyResponses.approved).toEqual(['my-key']);
});
it('claude: a corrupt existing file is treated as absent rather than failing the apply', () => {
const sessionId = 'sess-trust-4';
sessionsToClean.push(sessionId);
const configDir = customModelConfigDir(sessionId);
mkdirSync(configDir, { recursive: true });
writeFileSync(join(configDir, '.claude.json'), '{ not valid json');
expect(() => applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId)).not.toThrow();
const written = JSON.parse(readFileSync(join(configDir, '.claude.json'), 'utf8')) as {
customApiKeyResponses: { approved: string[] };
};
expect(written.customApiKeyResponses.approved).toEqual(['my-key']);
});
it('claude: re-approving the same key does not duplicate it in the approved list', () => {
const sessionId = 'sess-trust-5';
sessionsToClean.push(sessionId);
applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId);
const second = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'llama3', sessionId);
const written = JSON.parse(readFileSync(join(second!.configDir!, '.claude.json'), 'utf8')) as {
customApiKeyResponses: { approved: string[] };
};
expect(written.customApiKeyResponses.approved).toEqual(['my-key']);
});
it('opencode: has no apiKeyTrustFile declared (no configDirVar at all), nothing is seeded', () => {
const sessionId = 'sess-trust-opencode';
const applied = applyCustomModelInjection(entryOrThrow('opencode'), endpoint, 'qwen3', sessionId);
expect(applied?.configDir).toBeUndefined();
expect(existsSync(customModelConfigDir(sessionId))).toBe(false);
});
});
describe("applyCustomModelInjection: skipFirstRunPrompts (an isolated dir replays claude's whole first-run sequence)", () => {
it("claude: seeds hasCompletedOnboarding and this session's own project trust into .claude.json", () => {
const sessionId = 'sess-firstrun-1';
sessionsToClean.push(sessionId);
const applied = applyCustomModelInjection(
entryOrThrow('claude'),
endpoint,
'qwen3',
sessionId,
undefined,
'/home/user/myproject'
);
const written = JSON.parse(readFileSync(join(applied!.configDir!, '.claude.json'), 'utf8')) as {
hasCompletedOnboarding: boolean;
projects: Record<string, { hasTrustDialogAccepted: boolean }>;
};
expect(written.hasCompletedOnboarding).toBe(true);
expect(written.projects['/home/user/myproject'].hasTrustDialogAccepted).toBe(true);
});
it('claude: seeds skipDangerousModePermissionPrompt into settings.json', () => {
const sessionId = 'sess-firstrun-2';
sessionsToClean.push(sessionId);
const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId);
const written = JSON.parse(readFileSync(join(applied!.configDir!, 'settings.json'), 'utf8')) as {
skipDangerousModePermissionPrompt: boolean;
};
expect(written.skipDangerousModePermissionPrompt).toBe(true);
});
it('claude: with no workingDir given (boot recovery), hasCompletedOnboarding/settings still seed, but no project entry is added', () => {
const sessionId = 'sess-firstrun-3';
sessionsToClean.push(sessionId);
const applied = applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId);
const written = JSON.parse(readFileSync(join(applied!.configDir!, '.claude.json'), 'utf8')) as {
hasCompletedOnboarding: boolean;
projects?: Record<string, unknown>;
};
expect(written.hasCompletedOnboarding).toBeUndefined();
expect(written.projects).toBeUndefined();
});
it("claude: merges onto an existing project entry's other fields rather than overwriting them", () => {
const sessionId = 'sess-firstrun-4';
sessionsToClean.push(sessionId);
const configDir = customModelConfigDir(sessionId);
mkdirSync(configDir, { recursive: true });
writeFileSync(
join(configDir, '.claude.json'),
JSON.stringify({ projects: { '/home/user/myproject': { allowedTools: ['Bash'] } } })
);
const applied = applyCustomModelInjection(
entryOrThrow('claude'),
endpoint,
'qwen3',
sessionId,
undefined,
'/home/user/myproject'
);
const written = JSON.parse(readFileSync(join(applied!.configDir!, '.claude.json'), 'utf8')) as {
projects: Record<string, { allowedTools: string[]; hasTrustDialogAccepted: boolean }>;
};
expect(written.projects['/home/user/myproject'].allowedTools).toEqual(['Bash']);
expect(written.projects['/home/user/myproject'].hasTrustDialogAccepted).toBe(true);
});
it('claude: a corrupt existing settings.json is treated as absent rather than failing the apply', () => {
const sessionId = 'sess-firstrun-5';
sessionsToClean.push(sessionId);
const configDir = customModelConfigDir(sessionId);
mkdirSync(configDir, { recursive: true });
writeFileSync(join(configDir, 'settings.json'), '{ not valid json');
expect(() => applyCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', sessionId)).not.toThrow();
const written = JSON.parse(readFileSync(join(configDir, 'settings.json'), 'utf8')) as {
skipDangerousModePermissionPrompt: boolean;
};
expect(written.skipDangerousModePermissionPrompt).toBe(true);
});
it('pi: has no skipFirstRunPrompts concept (no apiKeyTrustFile either) — nothing beyond its own config file', () => {
const sessionId = 'sess-firstrun-pi';
sessionsToClean.push(sessionId);
const applied = applyCustomModelInjection(
entryOrThrow('pi'),
endpoint,
'qwen3',
sessionId,
undefined,
'/home/user/myproject'
);
const entries = readdirSync(applied!.configDir!);
expect(entries).not.toContain('settings.json');
});
});
describe('applyCustomModelInjection: pre-existing behavior unaffected', () => {
it('opencode: still returns a plain env-kind result with no configDir', () => {
const sessionId = 'sess-opencode-1';
const applied = applyCustomModelInjection(entryOrThrow('opencode'), endpoint, 'qwen3', sessionId);
expect(applied?.configDir).toBeUndefined();
expect(applied?.envOverrides.OPENCODE_CONFIG_CONTENT).toBeTruthy();
});
it('antigravity: still undefined (unsupported)', () => {
const applied = applyCustomModelInjection(entryOrThrow('antigravity'), endpoint, 'qwen3', 'sess-agy-1');
expect(applied).toBeUndefined();
});
});
+21 -17
View File
@@ -142,17 +142,19 @@ describe('custom-model-injection contract (mock server)', () => {
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
});
// gemini/deepseek's `env` kind passes the base URL through UNCHANGED (unlike
// opencode/codex/pi/omp/grok, which build a structured config and explicitly append
// /v1) — matching Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the
// SDK appends the path itself. Whether each of these TWO CLIs' own OpenAI-compatible
// client expects the var to already include /v1 (the common OpenAI-SDK convention) or
// appends it itself is genuinely CLI-specific and UNVERIFIED (see the confidence table
// in docs/custom-model-endpoints-plan.md) — these tests model the common OpenAI-SDK convention (base_url
// ends in /v1) since that's the more likely behavior for an OpenAI-compatible client,
// but that assumption should be corrected here the moment it's checked against a real
// binary. (grok WAS in this group too, until live-testing showed the whole `env` recipe
// was wrong for it — see its own test below.)
// gemini's `env` kind still passes the base URL through UNCHANGED (matching
// Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the SDK appends
// the path itself) — whether gemini-cli's own OpenAI-compatible-ish client expects the
// var to already include /v1 or appends it itself remains genuinely UNVERIFIED (it
// fails for an unrelated auth reason before this would even matter — see the
// confidence table in docs/custom-model-endpoints-plan.md); this test models the
// common OpenAI-SDK convention as the best guess, to be corrected the moment it's
// checked against a real client. deepseek WAS in this "passes through unchanged"
// group too, until reading `@deepseek-ai/dsh-llm-deepseek`'s own bundled source
// confirmed it builds its request URL as `${DEEPSEEK_BASE_URL}/chat/completions` with
// no `/v1` of its own — `appendV1Suffix` now fixes that (see its own test below),
// the same way grok's whole `env` recipe turned out to be wrong before live-testing
// corrected it to a `configDir` one.
it('gemini: GOOGLE_GEMINI_BASE_URL/GEMINI_API_KEY reach the mock', async () => {
const injection = buildCustomModelInjection(entryOrThrow('gemini'), endpointFor(mock), 'qwen3');
@@ -186,16 +188,18 @@ describe('custom-model-injection contract (mock server)', () => {
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
});
it('deepseek: DEEPSEEK_BASE_URL/DEEPSEEK_API_KEY reach the mock (base URL/key only, no model var)', async () => {
it('deepseek: DEEPSEEK_BASE_URL already carries the /v1 suffix dsh itself never adds, reaching the mock at the real path dsh requests', async () => {
// Confirmed by reading dsh's own bundled source: it fetches
// `${DEEPSEEK_BASE_URL}/chat/completions` verbatim, no /v1 insertion of its own — so
// this call (unlike gemini's above) passes DEEPSEEK_BASE_URL to callOpenAiCompat
// UNMODIFIED, exactly mirroring what the real harness does, rather than the test
// helping it along.
const injection = buildCustomModelInjection(entryOrThrow('deepseek'), endpointFor(mock), 'qwen3');
if (injection.kind !== 'env') throw new Error('unreachable');
expect(Object.keys(injection.envOverrides).sort()).toEqual(['DEEPSEEK_API_KEY', 'DEEPSEEK_BASE_URL']);
expect(injection.envOverrides.DEEPSEEK_BASE_URL).toBe(`${mock.baseUrl}/v1`);
await callOpenAiCompat(
`${injection.envOverrides.DEEPSEEK_BASE_URL}/v1`,
injection.envOverrides.DEEPSEEK_API_KEY,
'qwen3'
);
await callOpenAiCompat(injection.envOverrides.DEEPSEEK_BASE_URL, injection.envOverrides.DEEPSEEK_API_KEY, 'qwen3');
expect(mock.requests[0].path).toBe('/v1/chat/completions');
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
+55 -2
View File
@@ -58,6 +58,43 @@ describe('buildCustomModelInjection', () => {
});
});
it('claude: also declares configDirVar (CLAUDE_CONFIG_DIR isolation) on the env-kind result', () => {
const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.configDirVar).toBe('CLAUDE_CONFIG_DIR');
});
it('claude: injects CLAUDE_CODE_MAX_CONTEXT_TOKENS when a context length is known', () => {
const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3', 16384);
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBe('16384');
});
it('claude: omits CLAUDE_CODE_MAX_CONTEXT_TOKENS when the context length is unknown', () => {
const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.envOverrides.CLAUDE_CODE_MAX_CONTEXT_TOKENS).toBeUndefined();
});
it('claude: also declares apiKeyTrustFile, carrying the literal apiKey used', () => {
const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.apiKeyTrustFile).toEqual({ relPath: '.claude.json', shape: 'claude-api-key-responses' });
expect(result.apiKey).toBe('my-key');
});
it('claude: also declares skipFirstRunPrompts on the env-kind result', () => {
const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.skipFirstRunPrompts).toBe(true);
});
it('opencode: has no skipFirstRunPrompts (no apiKeyTrustFile/configDirVar concept for it either)', () => {
const result = buildCustomModelInjection(entryOrThrow('opencode'), endpoint, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.skipFirstRunPrompts).toBeUndefined();
});
it('claude: falls back to a dummy key when the endpoint has none', () => {
const result = buildCustomModelInjection(entryOrThrow('claude'), { ...endpoint, apiKey: undefined }, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
@@ -142,15 +179,31 @@ describe('buildCustomModelInjection', () => {
expect(result.extraEnv).toEqual({ XAI_API_KEY: 'my-key' });
});
it('deepseek: env kind sets base URL/key only, no model var', () => {
it('deepseek: env kind sets base URL (with a /v1 suffix appended) and key, no model var', () => {
// appendV1Suffix is REQUIRED here, not cosmetic: confirmed by reading dsh's own
// bundled source (@deepseek-ai/dsh-llm-deepseek) that it builds the request URL as
// `${DEEPSEEK_BASE_URL}/chat/completions` with no "/v1" of its own, while
// llama-swap/llama.cpp only serves "/v1/chat/completions" — without this, every
// request 404s (confirmed live; this is the fix for the originally-reported
// "dsh: HTTP_404: DeepSeek API error (HTTP 404)").
const result = buildCustomModelInjection(entryOrThrow('deepseek'), endpoint, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.envOverrides).toEqual({
DEEPSEEK_BASE_URL: 'http://192.168.1.50:8080',
DEEPSEEK_BASE_URL: 'http://192.168.1.50:8080/v1',
DEEPSEEK_API_KEY: 'my-key',
});
});
it('deepseek: appending the /v1 suffix is idempotent against a baseUrl that already ends in /v1', () => {
const result = buildCustomModelInjection(
entryOrThrow('deepseek'),
{ ...endpoint, baseUrl: 'http://192.168.1.50:8080/v1' },
'qwen3'
);
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.envOverrides.DEEPSEEK_BASE_URL).toBe('http://192.168.1.50:8080/v1');
});
it('antigravity: unsupported', () => {
const result = buildCustomModelInjection(entryOrThrow('antigravity'), endpoint, 'qwen3');
expect(result).toEqual({ kind: 'unsupported' });
+238
View File
@@ -0,0 +1,238 @@
/**
* @fileoverview Tests for `getLatestLlamaSwapLogLine()`/`pruneIdleLlamaSwapLogTails()` —
* the real-time "what is llama.cpp actually doing" feed behind the loading banner's
* second line (docs/custom-model-endpoints-plan.md). Confirmed live against a real
* llama-swap deployment: its `GET /api/events` SSE stream carries the backend
* llama-server process's own stdout (`load_model: ...`, `llama_server: model loaded`)
* as `{"type":"logData","data":"{\"data\":\"...\",\"source\":\"upstream\"}"}` frames,
* tagged distinctly from llama-swap's own `source: "proxy"` request-access log frames.
*
* ⚠️ `GET /logs` (the endpoint this feature's own first cut was built against, before
* being caught by exactly this kind of live check) turns out to carry ONLY the proxy
* log — confirmed live it never showed a single backend line even seconds after a real,
* confirmed model swap. `/api/events` is the only source that actually has the data.
*
* Drives a hand-built `ReadableStream` body through the mocked `webviewFetch` rather
* than a real network round-trip — the point under test is the SSE-frame parsing and
* `source` filtering plus the one-connection-per-endpoint reuse, not networking itself.
*
* Each test uses its own host id (`llamaSwapLogTails` is a module-level Map, shared
* across every test in this file) and `afterEach` force-prunes everything so no tail
* a test forgot to close leaks into the next one.
*
* Port: N/A (no server; drives the exported functions directly).
*/
import { describe, it, expect, vi, afterEach } from 'vitest';
import { getLatestLlamaSwapLogLine, pruneIdleLlamaSwapLogTails } from '../src/web/routes/custom-model-routes.js';
import { webviewFetch } from '../src/web/webview-egress.js';
import type { CustomModelHost } from '../src/custom-model-hosts.js';
vi.mock('../src/web/webview-egress.js', async () => {
const actual = await vi.importActual<typeof import('../src/web/webview-egress.js')>('../src/web/webview-egress.js');
return { ...actual, webviewFetch: vi.fn() };
});
const fetchMock = vi.mocked(webviewFetch);
/** One real `GET /api/events` SSE frame carrying backend (`source: "upstream"`) log text. */
function upstreamLogFrame(text: string): string {
const inner = JSON.stringify({ data: text, source: 'upstream' });
return `event:message\ndata:${JSON.stringify({ type: 'logData', data: inner })}\n\n`;
}
/** The proxy-log flavor of the same event shape — must never be surfaced as `latestLine`. */
function proxyLogFrame(text: string): string {
const inner = JSON.stringify({ data: text, source: 'proxy' });
return `event:message\ndata:${JSON.stringify({ type: 'logData', data: inner })}\n\n`;
}
/** A streaming Response whose body enqueues `frames` up front and then stays open
* (never closes) — matches a real `/api/events` connection, confirmed live to stay
* open indefinitely (read past 220KB over 8s with no `done`). */
function openStreamResponse(frames: string[]): Response {
const encoder = new TextEncoder();
const stream = new ReadableStream<Uint8Array>({
start(controller) {
for (const frame of frames) controller.enqueue(encoder.encode(frame));
// deliberately never controller.close()
},
});
return new Response(stream, { status: 200 });
}
function host(id: string): CustomModelHost {
return { id, label: id, baseUrl: `http://192.168.1.50:8080/${id}` };
}
/** Lets the fire-and-forget stream-pump's microtasks (reader.read() resolutions) settle. */
async function flush(): Promise<void> {
await new Promise((resolve) => setTimeout(resolve, 10));
}
afterEach(() => {
pruneIdleLlamaSwapLogTails(Number.POSITIVE_INFINITY); // force-close every tail this file opened
fetchMock.mockReset();
});
describe('getLatestLlamaSwapLogLine', () => {
it('returns undefined before any line has arrived, then the real backend log line once it does', async () => {
const h = host('t1');
fetchMock.mockResolvedValue(
openStreamResponse([upstreamLogFrame('0.31.428.568 I srv llama_server: model loaded')])
);
const before = getLatestLlamaSwapLogLine(h);
expect(before).toBeUndefined();
await flush();
const after = getLatestLlamaSwapLogLine(h);
expect(after).toBe('0.31.428.568 I srv llama_server: model loaded');
});
it('filters out llama-swap\'s own proxy-sourced frames, keeping only source: "upstream"', async () => {
const h = host('t2');
fetchMock.mockResolvedValue(
openStreamResponse([
proxyLogFrame('[INFO] Request 10.10.10.1 "GET /running HTTP/1.1" 200 407 "undici" 46.207µs'),
upstreamLogFrame('0.14.157.100 I srv load_model: initializing, n_slots = 4, n_ctx_slot = 16384'),
proxyLogFrame('[WARN] some warning about something unrelated'),
])
);
getLatestLlamaSwapLogLine(h);
await flush();
expect(getLatestLlamaSwapLogLine(h)).toBe(
'0.14.157.100 I srv load_model: initializing, n_slots = 4, n_ctx_slot = 16384'
);
});
it('keeps the LAST line when one upstream frame batches several newline-joined lines', async () => {
const h = host('t3');
fetchMock.mockResolvedValue(
openStreamResponse([
upstreamLogFrame(
'0.00.001.000 I srv llama_server: starting\n0.00.002.000 I srv llama_server: loading tensors'
),
upstreamLogFrame('0.00.003.000 I srv llama_server: model loaded'),
])
);
getLatestLlamaSwapLogLine(h);
await flush();
expect(getLatestLlamaSwapLogLine(h)).toBe('0.00.003.000 I srv llama_server: model loaded');
});
it('handles a frame split across two stream chunks (SSE double-newline boundary not yet seen)', async () => {
const h = host('t3b');
const whole = upstreamLogFrame('0.00.005.000 I srv llama_server: model loaded');
const splitAt = Math.floor(whole.length / 2);
fetchMock.mockResolvedValue(openStreamResponse([whole.slice(0, splitAt), whole.slice(splitAt)]));
getLatestLlamaSwapLogLine(h);
await flush();
expect(getLatestLlamaSwapLogLine(h)).toBe('0.00.005.000 I srv llama_server: model loaded');
});
it('ignores a malformed frame instead of throwing', async () => {
const h = host('t3c');
fetchMock.mockResolvedValue(
openStreamResponse(['event:message\ndata:not valid json\n\n', upstreamLogFrame('llama_server: model loaded')])
);
getLatestLlamaSwapLogLine(h);
await flush();
expect(getLatestLlamaSwapLogLine(h)).toBe('llama_server: model loaded');
});
it('ignores a non-logData event type', async () => {
const h = host('t3d');
fetchMock.mockResolvedValue(
openStreamResponse([
`event:message\ndata:${JSON.stringify({ type: 'modelStatus', data: '{}' })}\n\n`,
upstreamLogFrame('llama_server: model loaded'),
])
);
getLatestLlamaSwapLogLine(h);
await flush();
expect(getLatestLlamaSwapLogLine(h)).toBe('llama_server: model loaded');
});
it('opens exactly one connection per endpoint — a second call while the tail is open never re-fetches', async () => {
const h = host('t4');
fetchMock.mockResolvedValue(openStreamResponse([upstreamLogFrame('llama_server: model loaded')]));
getLatestLlamaSwapLogLine(h);
await flush();
getLatestLlamaSwapLogLine(h);
getLatestLlamaSwapLogLine(h);
expect(fetchMock).toHaveBeenCalledTimes(1);
});
it("requests /api/events specifically, with the endpoint's own auth headers", async () => {
const h: CustomModelHost = { id: 't5', label: 't5', baseUrl: 'http://192.168.1.60:9000', apiKey: 'secret-key' };
fetchMock.mockResolvedValue(openStreamResponse([]));
getLatestLlamaSwapLogLine(h);
expect(fetchMock).toHaveBeenCalledTimes(1);
const [url, init] = fetchMock.mock.calls[0]!;
expect((url as URL).pathname).toBe('/api/events');
expect((init as RequestInit).headers).toMatchObject({ Authorization: 'Bearer secret-key' });
});
it('an unreachable endpoint (fetch throws) leaves latestLine undefined rather than throwing', async () => {
const h = host('t6');
fetchMock.mockRejectedValue(new TypeError('fetch failed'));
expect(() => getLatestLlamaSwapLogLine(h)).not.toThrow();
await flush();
expect(getLatestLlamaSwapLogLine(h)).toBeUndefined();
});
it('a non-2xx response leaves latestLine undefined rather than throwing', async () => {
const h = host('t7');
fetchMock.mockResolvedValue(new Response('not found', { status: 404 }));
getLatestLlamaSwapLogLine(h);
await flush();
expect(getLatestLlamaSwapLogLine(h)).toBeUndefined();
});
});
describe('pruneIdleLlamaSwapLogTails', () => {
it('closes a tail nothing has polled recently, so the next access starts a fresh connection', async () => {
const h = host('t8');
fetchMock.mockResolvedValue(openStreamResponse([upstreamLogFrame('llama_server: model loaded')]));
getLatestLlamaSwapLogLine(h); // opens the first connection, lastAccessedAt = now
await flush();
expect(fetchMock).toHaveBeenCalledTimes(1);
pruneIdleLlamaSwapLogTails(Date.now() + 60_000); // "now" far enough ahead that the tail reads as idle
getLatestLlamaSwapLogLine(h); // the entry was removed — this must open a NEW connection
await flush();
expect(fetchMock).toHaveBeenCalledTimes(2);
});
it('leaves a recently-accessed tail alone', async () => {
const h = host('t9');
fetchMock.mockResolvedValue(openStreamResponse([upstreamLogFrame('llama_server: model loaded')]));
getLatestLlamaSwapLogLine(h);
await flush();
pruneIdleLlamaSwapLogTails(Date.now()); // no time has passed — nothing is idle yet
getLatestLlamaSwapLogLine(h);
expect(fetchMock).toHaveBeenCalledTimes(1); // still just the one connection
});
});
+241
View File
@@ -0,0 +1,241 @@
/**
* @fileoverview Frontend tests for the one-shot custom-model launch path added to
* session-ui.js (docs/custom-model-endpoints-plan.md): `runCustomModelEntry` dispatches
* to `_runCustomModelEntryOneShot` for every custom-model-eligible CLI except claude,
* which launches directly on the endpoint (no restart) by folding `customModel` into
* the run<Mode>() function's own `/api/quick-start` body via `_pendingCustomModelForLaunch`
* and `_quickStartWithCustomModelConfirm`. Fixes the visible native-boot-then-restart the
* restart-after-launch path (`_runCustomModelEntryViaRestart`, still used for claude)
* showed on every custom-model run — confirmed live on Codex, whose TUI fully
* reinitializes on a restart.
*
* Uses the same JSDOM + `runScripts: "dangerously"` approach as
* test/custom-model-run-menu-ui.test.ts, extended with the DOM elements runCodex() (the
* CLI this was reported against) reads.
*
* Port: none.
*/
import { readFileSync } from 'node:fs';
import { JSDOM } from 'jsdom';
import { describe, expect, it } from 'vitest';
const CONSTANTS_JS = readFileSync(new URL('../src/web/public/constants.js', import.meta.url), 'utf-8');
const SESSION_UI_JS = readFileSync(new URL('../src/web/public/session-ui.js', import.meta.url), 'utf-8');
function bootApp() {
const dom = new JSDOM(
`<!doctype html><body>
<select id="quickStartCase"><option value="testcase" selected>testcase</option></select>
<input id="tabCount" value="1">
<button id="runBtn"></button>
<div id="runModeMenu"></div>
</body>`,
{ url: 'http://localhost/', runScripts: 'dangerously' }
);
const win = dom.window as unknown as Window & typeof globalThis & { CodemanApp: new () => any };
(win as unknown as { eval: (s: string) => void }).eval('window.CodemanApp = function CodemanApp() {};');
(win as unknown as { eval: (s: string) => void }).eval(CONSTANTS_JS);
(win as unknown as { eval: (s: string) => void }).eval(SESSION_UI_JS);
const app = new win.CodemanApp();
app.cases = [{ name: 'testcase' }];
app.terminal = { focus: () => {} };
app.loadAppSettingsFromStorage = () => ({});
app.getCaseSettings = () => ({});
app.buildEnvOverrides = () => ({});
app.showToast = () => {};
app._beginSessionLaunchStatus = () => 'status-token';
app._reportSessionLaunchError = (_token: unknown, message: string) => {
app._lastReportedError = message;
};
app._ensureCreatedSessionVisible = async () => {};
app.selectSession = async () => {};
app._nextCaseSessionStartNumber = () => 1;
return { win, app };
}
describe('runCustomModelEntry dispatch', () => {
it('routes claude through the restart-after-launch path', async () => {
const { app } = bootApp();
let calledRestart = false;
let calledOneShot = false;
app._runCustomModelEntryViaRestart = async () => {
calledRestart = true;
};
app._runCustomModelEntryOneShot = async () => {
calledOneShot = true;
};
await app.runCustomModelEntry('claude', 'llama-box', 'qwen3');
expect(calledRestart).toBe(true);
expect(calledOneShot).toBe(false);
});
it('routes every other custom-model-eligible CLI through the one-shot path', async () => {
for (const mode of ['opencode', 'codex', 'gemini', 'pi', 'grok', 'deepseek', 'omp']) {
const { app } = bootApp();
let calledRestart = false;
let calledOneShot = false;
app._runCustomModelEntryViaRestart = async () => {
calledRestart = true;
};
app._runCustomModelEntryOneShot = async () => {
calledOneShot = true;
};
await app.runCustomModelEntry(mode, 'llama-box', 'qwen3');
expect(calledRestart, mode).toBe(false);
expect(calledOneShot, mode).toBe(true);
}
});
});
describe('_runCustomModelEntryOneShot', () => {
it('stashes the pick on _pendingCustomModelForLaunch for the duration of run(), then clears it', async () => {
const { app } = bootApp();
let seenDuringRun: unknown;
app.run = async function (this: typeof app) {
seenDuringRun = this._pendingCustomModelForLaunch;
};
await app._runCustomModelEntryOneShot('codex', 'llama-box', 'qwen3');
expect(seenDuringRun).toEqual({ endpointId: 'llama-box', modelId: 'qwen3' });
expect(app._pendingCustomModelForLaunch).toBeUndefined();
});
it('clears the pending pick even when run() throws', async () => {
const { app } = bootApp();
app.run = async () => {
throw new Error('boom');
};
await expect(app._runCustomModelEntryOneShot('codex', 'llama-box', 'qwen3')).rejects.toThrow('boom');
expect(app._pendingCustomModelForLaunch).toBeUndefined();
});
it('starts the loading watcher when the launch reports modelSwapInProgress, passing the new session id', async () => {
const { app } = bootApp();
app.run = async () => {
app._lastCustomModelLaunchResult = { modelSwapInProgress: true, sessionId: 'new-session' };
};
let watched: unknown[] | null = null;
app._watchLlamaSwapLoading = async (...args: unknown[]) => {
watched = args;
};
await app._runCustomModelEntryOneShot('codex', 'llama-box', 'qwen3');
expect(watched).toEqual(['llama-box', 'qwen3', 'new-session']);
});
it('never starts the watcher when no swap was needed', async () => {
const { app } = bootApp();
app.run = async () => {
app._lastCustomModelLaunchResult = { modelSwapInProgress: false };
};
let watchCalled = false;
app._watchLlamaSwapLoading = async () => {
watchCalled = true;
};
await app._runCustomModelEntryOneShot('codex', 'llama-box', 'qwen3');
expect(watchCalled).toBe(false);
});
});
describe('_quickStartWithCustomModelConfirm', () => {
function withFetch(win: Window & typeof globalThis, handler: (body: any) => any) {
(win as unknown as { fetch: typeof fetch }).fetch = (async (_url: string, opts: any) => ({
json: async () => handler(JSON.parse(opts.body)),
})) as unknown as typeof fetch;
}
it('returns the response directly when no confirmation is needed', async () => {
const { win, app } = bootApp();
withFetch(win, (body) => ({ success: true, data: { sessionId: 's1', modelSwapInProgress: false, body } }));
const data = await app._quickStartWithCustomModelConfirm({
mode: 'codex',
customModel: { endpointId: 'e', modelId: 'm' },
});
expect(data.success).toBe(true);
expect(data.data.sessionId).toBe('s1');
expect(app._lastCustomModelLaunchResult).toEqual(data.data);
});
it('confirming re-sends with confirmedSwap and returns the second response', async () => {
const { win, app } = bootApp();
app._confirmModelSwap = async () => true;
let calls = 0;
withFetch(win, (body) => {
calls += 1;
if (calls === 1) {
return {
success: true,
data: {
requiresConfirmation: true,
currentlyLoadedModel: 'llama3',
affectedSessions: [{ id: 's2', name: 'w2' }],
},
};
}
// the SWAP question's own flag, never the blanket `confirmed`: answering this one
// must not also silence the context-floor warning.
expect(body.customModel.confirmedSwap).toBe(true);
expect(body.customModel.confirmed).toBeUndefined();
return { success: true, data: { sessionId: 's1', modelSwapInProgress: true } };
});
const data = await app._quickStartWithCustomModelConfirm({
mode: 'codex',
customModel: { endpointId: 'e', modelId: 'm' },
});
expect(calls).toBe(2);
expect(data.data.sessionId).toBe('s1');
expect(app._lastCustomModelLaunchResult.modelSwapInProgress).toBe(true);
});
it('cancelling never re-sends, and reports a cancellation error', async () => {
const { win, app } = bootApp();
app._confirmModelSwap = async () => false;
let calls = 0;
withFetch(win, () => {
calls += 1;
return {
success: true,
data: {
requiresConfirmation: true,
currentlyLoadedModel: 'llama3',
affectedSessions: [{ id: 's2', name: 'w2' }],
},
};
});
const data = await app._quickStartWithCustomModelConfirm({
mode: 'codex',
customModel: { endpointId: 'e', modelId: 'm' },
});
expect(calls).toBe(1);
expect(data.success).toBe(false);
expect(data.error).toMatch(/cancelled/i);
expect(app._lastCustomModelLaunchResult).toBeUndefined();
});
});
describe('runCodex(): one-shot custom-model launch (the CLI this was reported against)', () => {
it('folds _pendingCustomModelForLaunch into the quick-start body as customModel', async () => {
const { win, app } = bootApp();
(win as unknown as { fetch: typeof fetch }).fetch = (async (url: string, opts?: any) => {
if (url === '/api/codex/status') return { json: async () => ({ data: { available: true } }) };
const body = JSON.parse(opts.body);
expect(body.customModel).toEqual({ endpointId: 'llama-box', modelId: 'qwen3' });
return { json: async () => ({ success: true, data: { sessionId: 's1', modelSwapInProgress: false } }) };
}) as unknown as typeof fetch;
app._pendingCustomModelForLaunch = { endpointId: 'llama-box', modelId: 'qwen3' };
await app.runCodex();
expect(app._lastReportedError).toBeUndefined();
});
it('omits customModel entirely for a plain (non-custom-model) Codex launch', async () => {
const { win, app } = bootApp();
(win as unknown as { fetch: typeof fetch }).fetch = (async (url: string, opts?: any) => {
if (url === '/api/codex/status') return { json: async () => ({ data: { available: true } }) };
const body = JSON.parse(opts.body);
expect(body.customModel).toBeUndefined();
return { json: async () => ({ success: true, data: { sessionId: 's1' } }) };
}) as unknown as typeof fetch;
await app.runCodex();
expect(app._lastReportedError).toBeUndefined();
});
});
File diff suppressed because it is too large Load Diff
+180
View File
@@ -0,0 +1,180 @@
/**
* @fileoverview Tests for `detectCustomModelSwapDisplacements()`, the periodic sweep
* behind server.ts's "custom model swap-displacement check" timer
* (docs/custom-model-endpoints-plan.md). The apply/create routes' own swap-conflict check
* only ever runs at a session's own launch/apply moment — this sweep is what catches a
* LATER eviction triggered by a different session's normal use, which the launch-time
* check structurally cannot see.
*
* Kept in its own file for the same reason as `custom-model-endpoint-rediscovery.test.ts`:
* a sweep that walks every saved host would otherwise pick up hosts other tests in a
* shared file create, making an exact call-count assertion meaningless.
*
* Port: N/A (no server; drives readCustomModelHosts/writeCustomModelHosts directly plus
* the mocked webviewFetch dispatcher).
*/
import { describe, it, expect, vi, beforeEach } from 'vitest';
import { getDataDir } from '../src/config/instance.js';
import { writeCustomModelHosts, type CustomModelHost } from '../src/custom-model-hosts.js';
import {
detectCustomModelSwapDisplacements,
type CustomModelSessionLike,
} from '../src/web/routes/custom-model-routes.js';
import { webviewFetch } from '../src/web/webview-egress.js';
vi.mock('../src/web/webview-egress.js', async () => {
const actual = await vi.importActual<typeof import('../src/web/webview-egress.js')>('../src/web/webview-egress.js');
return { ...actual, webviewFetch: vi.fn() };
});
const fetchMock = vi.mocked(webviewFetch);
const ENDPOINT: CustomModelHost = {
id: 'llama-swap',
label: 'llama-swap',
baseUrl: 'http://192.168.1.50:8080',
apiKey: 'k',
};
function session(
overrides: Partial<CustomModelSessionLike> & Pick<CustomModelSessionLike, 'id'>
): CustomModelSessionLike {
return { name: overrides.id, ...overrides };
}
function mockRunning(running: Array<{ model: string; state: string }>) {
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/running') return new Response(JSON.stringify({ running }), { status: 200 });
throw new Error(`unexpected request in this test: ${url.href}`);
});
}
beforeEach(() => {
fetchMock.mockReset();
});
describe('detectCustomModelSwapDisplacements', () => {
it('flags a session whose own model is no longer in the running list, naming what displaced it', async () => {
await writeCustomModelHosts(getDataDir(), [ENDPOINT]);
mockRunning([{ model: 'fast', state: 'ready' }]);
const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } });
const notified = new Set<string>();
const displacements = await detectCustomModelSwapDisplacements([w1], notified);
expect(displacements).toEqual([
{
sessionId: 'w1',
sessionName: 'w1',
endpointId: 'llama-swap',
previousModel: 'qwen3',
currentlyLoadedModel: 'fast',
},
]);
expect(notified.has('w1')).toBe(true);
});
it('does not flag a session whose own model is still the one loaded and ready', async () => {
await writeCustomModelHosts(getDataDir(), [ENDPOINT]);
mockRunning([{ model: 'qwen3', state: 'ready' }]);
const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } });
const displacements = await detectCustomModelSwapDisplacements([w1], new Set());
expect(displacements).toEqual([]);
});
it('notifies once per displacement — a repeat sweep with nothing changed does not re-flag it', async () => {
await writeCustomModelHosts(getDataDir(), [ENDPOINT]);
mockRunning([{ model: 'fast', state: 'ready' }]);
const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } });
const notified = new Set<string>();
const first = await detectCustomModelSwapDisplacements([w1], notified);
const second = await detectCustomModelSwapDisplacements([w1], notified);
expect(first).toHaveLength(1);
expect(second).toEqual([]);
});
it('clears the notified flag once the session is back on its own model, so a later displacement flags again', async () => {
await writeCustomModelHosts(getDataDir(), [ENDPOINT]);
const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } });
const notified = new Set<string>();
mockRunning([{ model: 'fast', state: 'ready' }]);
await detectCustomModelSwapDisplacements([w1], notified);
expect(notified.has('w1')).toBe(true);
mockRunning([{ model: 'qwen3', state: 'ready' }]); // back to normal
await detectCustomModelSwapDisplacements([w1], notified);
expect(notified.has('w1')).toBe(false);
mockRunning([{ model: 'fast', state: 'ready' }]); // displaced again
const third = await detectCustomModelSwapDisplacements([w1], notified);
expect(third).toHaveLength(1);
});
it('skips a session on a non-llama-swap endpoint (no /running) — nothing to compare, never flagged', async () => {
await writeCustomModelHosts(getDataDir(), [ENDPOINT]);
fetchMock.mockResolvedValue(new Response('not found', { status: 404 }));
const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } });
const displacements = await detectCustomModelSwapDisplacements([w1], new Set());
expect(displacements).toEqual([]);
});
it('skips a session whose endpoint was deleted since it was created', async () => {
await writeCustomModelHosts(getDataDir(), []); // ENDPOINT never saved
const w1 = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } });
const displacements = await detectCustomModelSwapDisplacements([w1], new Set());
expect(displacements).toEqual([]);
expect(fetchMock).not.toHaveBeenCalled();
});
it('ignores a plain session with no customModel selection at all', async () => {
const displacements = await detectCustomModelSwapDisplacements([session({ id: 'plain' })], new Set());
expect(displacements).toEqual([]);
expect(fetchMock).not.toHaveBeenCalled();
});
it('one endpoint failing (unreachable) never blocks checking sessions on another', async () => {
const DOWN: CustomModelHost = { id: 'down', label: 'down', baseUrl: 'http://192.168.1.60:8080' };
await writeCustomModelHosts(getDataDir(), [ENDPOINT, DOWN]);
fetchMock.mockImplementation(async (url: URL) => {
if (url.href.includes('192.168.1.60')) throw new TypeError('fetch failed', { cause: new Error('ECONNREFUSED') });
if (url.pathname === '/running') {
return new Response(JSON.stringify({ running: [{ model: 'fast', state: 'ready' }] }), { status: 200 });
}
throw new Error(`unexpected request in this test: ${url.href}`);
});
const onDown = session({ id: 'w-down', customModel: { endpointId: 'down', modelId: 'x' } });
const onLlamaSwap = session({ id: 'w1', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } });
const displacements = await detectCustomModelSwapDisplacements([onDown, onLlamaSwap], new Set());
expect(displacements).toEqual([
{
sessionId: 'w1',
sessionName: 'w1',
endpointId: 'llama-swap',
previousModel: 'qwen3',
currentlyLoadedModel: 'fast',
},
]);
});
it('multiple sessions on the same endpoint each get their own displacement entry', async () => {
await writeCustomModelHosts(getDataDir(), [ENDPOINT]);
mockRunning([{ model: 'gemma', state: 'ready' }]);
const w1 = session({ id: 'w1', name: 'w1-test2', customModel: { endpointId: 'llama-swap', modelId: 'qwen3' } });
const w2 = session({ id: 'w2', name: 'w2-test2', customModel: { endpointId: 'llama-swap', modelId: 'fast' } });
const displacements = await detectCustomModelSwapDisplacements([w1, w2], new Set());
expect(displacements.map((d) => d.sessionId).sort()).toEqual(['w1', 'w2']);
});
});
+184
View File
@@ -0,0 +1,184 @@
// Port: none (pure frontend module in a node VM with a fake DOM — no browser, no server).
//
// The remote-host wake banner (src/web/public/host-wake-ui.js) is a SINGLE global
// element that is shown only for the active remote session. The regression this
// guards: `refreshHostWakeBanner` clears `_hostWake` when the tab switches, but
// `_hostWakeTick`'s clear branch only re-rendered when IT was the one clearing —
// so switching from an unreachable remote session to a LOCAL one left the banner
// visible ("Hufflepuff is not reachable") on every chat until a full reload.
//
// The bug is a pure ordering problem between two methods, so it can be reproduced
// here without a browser: render the remote state, switch to a local session, and
// assert the banner is hidden again.
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it } from 'vitest';
const PUBLIC = resolve(import.meta.dirname, '../src/web/public');
const REMOTE_ID = 'remote-session-0001';
const LOCAL_ID = 'local-session-0001';
type El = { hidden: boolean; textContent: string; disabled: boolean; classList: { add(): void; remove(): void } };
function fakeElement(): El {
return { hidden: false, textContent: '', disabled: false, classList: { add() {}, remove() {} } };
}
const PROXIED_ID = 'remote-session-proxied';
const NOWOL_ID = 'remote-session-nowol';
/** Load `host-wake-ui.js` with the minimal DOM it touches, and return a wired app. */
function loadWakeApp() {
const fetches: string[] = [];
const elements = new Map<string, El>([
['hostWakeBanner', fakeElement()],
['hostWakeBannerText', fakeElement()],
['hostWakeBannerDetail', fakeElement()],
['hostWakeBannerAction', fakeElement()],
]);
const CodemanApp = function CodemanApp(this: unknown) {};
const context = vm.createContext({
CodemanApp,
console,
setInterval: () => 1,
clearInterval: () => {},
fetch: (url: string) => {
fetches.push(url);
return Promise.resolve({ json: () => Promise.resolve({ success: false }) });
},
document: {
visibilityState: 'visible',
getElementById: (id: string) => elements.get(id) ?? null,
addEventListener: () => {},
},
window: {},
});
vm.runInContext(readFileSync(resolve(PUBLIC, 'host-wake-ui.js'), 'utf8'), context, { filename: 'host-wake-ui.js' });
const app = new (CodemanApp as new () => Record<string, unknown>)();
app.$ = (id: string) => elements.get(id) ?? null;
app.activeSessionId = REMOTE_ID;
app.sessions = new Map<string, { remote?: Record<string, unknown> }>([
[
REMOTE_ID,
{ remote: { hostId: 'hufflepuff', host: '192.168.50.137', label: 'Hufflepuff', wakeMac: '04:d9:f5:80:c6:58' } },
],
[
PROXIED_ID,
{
remote: {
hostId: 'bastioned',
host: '10.20.0.5',
label: 'Behind bastion',
jumpHost: 'bastion',
wakeMac: '04:d9:f5:80:c6:58',
},
},
],
[NOWOL_ID, { remote: { hostId: 'plain', host: '10.0.0.9', label: 'Plain' } }],
[LOCAL_ID, {}],
]);
return {
app,
fetches,
banner: elements.get('hostWakeBanner') as El,
text: elements.get('hostWakeBannerText') as El,
};
}
describe('host wake banner visibility', () => {
it('hides the banner when switching from an unreachable remote session to a local one', () => {
const { app, banner, text } = loadWakeApp();
// The banner is up for the active, unreachable remote session.
app._hostWake = {
sessionId: REMOTE_ID,
reachable: false,
wakeConfigured: 'mac',
host: '192.168.50.137',
label: 'Hufflepuff',
waking: false,
error: '',
};
(app._renderHostWakeBanner as () => void)();
expect(banner.hidden).toBe(false);
expect(text.textContent).toBe('Hufflepuff is not reachable');
// Switch to a LOCAL session. `refreshHostWakeBanner` clears the state, and the
// tick that follows must still repaint the (now empty) banner as hidden.
app.activeSessionId = LOCAL_ID;
(app.refreshHostWakeBanner as (id: string) => void)(LOCAL_ID);
expect(app._hostWake).toBeNull();
expect(banner.hidden).toBe(true);
});
it('keeps the banner hidden on a later poller tick once the state is cleared', () => {
const { app, banner } = loadWakeApp();
app.activeSessionId = LOCAL_ID;
app._hostWake = null;
// A page-wide tick on a local session must be idempotent and leave it hidden.
(app._hostWakeTick as () => void)();
expect(banner.hidden).toBe(true);
});
it('shows the banner only while the active session is remote and unreachable', () => {
const { app, banner } = loadWakeApp();
app._hostWake = {
sessionId: REMOTE_ID,
reachable: false,
wakeConfigured: 'mac',
host: '192.168.50.137',
label: 'Hufflepuff',
waking: false,
error: '',
};
(app._renderHostWakeBanner as () => void)();
expect(banner.hidden).toBe(false);
// Reachable again → hidden, state intact (the banner must not leak across the
// reachable/unreachable transition either).
app._hostWake.reachable = true;
(app._renderHostWakeBanner as () => void)();
expect(banner.hidden).toBe(true);
});
});
describe('host wake banner polling', () => {
// Each poll is a TCP connect to the host from the server. The timer is the one
// trigger that is not a user action, so it must not fire for a host Codeman could
// not wake anyway (it cannot wake it, but it can keep an activity-based suspend timer
// from firing), and a proxied host is never polled: the probe cannot reach it.
const tick = (app: Record<string, unknown>, periodic: boolean) =>
(app._hostWakeTick as (o: { periodic: boolean }) => void)({ periodic });
it('polls a wake-configured host on activation and on the timer', () => {
const { app, fetches } = loadWakeApp();
app.activeSessionId = REMOTE_ID;
tick(app, false);
tick(app, true);
tick(app, true);
expect(fetches).toHaveLength(3);
expect(fetches[0]).toContain(`/api/sessions/${REMOTE_ID}/reachability`);
});
it('polls a host without a wake target once on activation, never on the timer', () => {
const { app, fetches } = loadWakeApp();
app.activeSessionId = NOWOL_ID;
tick(app, false);
tick(app, true);
tick(app, true);
expect(fetches).toHaveLength(1);
});
it('never polls a host behind a jump host or SOCKS proxy', () => {
const { app, fetches } = loadWakeApp();
app.activeSessionId = PROXIED_ID;
tick(app, false);
tick(app, true);
expect(fetches).toHaveLength(0);
expect((app._hostWake as { probeable: boolean }).probeable).toBe(false);
});
});
+6 -3
View File
@@ -63,9 +63,12 @@ describe('keyboard shortcuts', () => {
// The xterm handler owns this decision, and the no-selection path must fall
// through with NO preventDefault so xterm still evaluates Ctrl+C into 0x03.
expect(terminalUiSource).toContain('this.shouldCopyTerminalSelectionFromShortcut?.(ev)');
expect(terminalUiSource).toMatch(
/const selection = this\.terminal\.hasSelection\?\.\(\) \? this\.terminal\.getSelection\(\) : '';/
);
// The CLEANED selection is what decides. A drag across the blank part of a row
// selects real padding spaces, so the raw text is truthy and testing it would
// spend the press on a copy of nothing — the same lost interrupt this test
// guards, reached by a different door.
expect(terminalUiSource).toMatch(/const selection = this\.cleanedTerminalSelection\(\);/);
expect(terminalUiSource).toMatch(/if \(selection\.trim\(\)\) \{/);
expect(terminalUiSource).toContain('void this.copyTerminalSelection(selection);');
expect(appSource).toContain("id: 'copy-selection'");
});
+8
View File
@@ -99,6 +99,14 @@ export class MockSession extends EventEmitter {
return true;
}
/**
* Mirrors `Session.reattachRemote()` — the COD-108 transport re-establish that
* the wake-on-LAN flow calls once a sleeping host is back. Defaults to success;
* set `reattachRemote.mockResolvedValue(false)` to model a pane that could not
* be respawned.
*/
reattachRemote = vi.fn(async (): Promise<boolean> => true);
/** Exactly-once input dedup — mirrors Session.shouldApplyInput so route tests
* exercising the reliable-delivery path behave like production. */
private _appliedInputSeq = new Map<string, number>();
+114
View File
@@ -6,11 +6,14 @@ import {
defaultRemoteCommandForMode,
readRemoteCases,
readRemoteHosts,
rehydrateRemoteHostFields,
remoteDisplayPath,
remoteSshTarget,
toSessionRemote,
writeRemoteCases,
writeRemoteHosts,
} from '../src/remote-hosts.js';
import { RemoteHostSchema } from '../src/web/schemas.js';
describe('remote-hosts domain', () => {
let dir: string | null = null;
@@ -69,4 +72,115 @@ describe('remote-hosts domain', () => {
'aamer@box.local:/opt/work'
);
});
it('carries the wake command from host config into the session', () => {
// The input route reads `session.remote.wakeCommand` — it must survive the host
// -> session mapping, or wake-on-LAN silently degrades to "no wake command".
const remote = toSessionRemote(
{
id: 'hufflepuff',
label: 'Hufflepuff',
host: '192.168.50.137',
username: 'j',
wakeCommand: '/home/joe/bin/whuff',
},
{ name: 'c', type: 'remote', hostId: 'hufflepuff', remotePath: '/home/j/work' }
);
expect(remote.wakeCommand).toBe('/home/joe/bin/whuff');
});
it('omits the wake command by default (feature off without a config entry)', () => {
const remote = toSessionRemote(
{ id: 'h', label: 'H', host: '10.0.0.1', username: 'j' },
{ name: 'c', type: 'remote', hostId: 'h', remotePath: '/tmp' }
);
expect(remote.wakeCommand).toBeUndefined();
});
describe('RemoteHostSchema wakeCommand', () => {
const host = { id: 'hufflepuff', label: 'Hufflepuff', host: '192.168.50.137', username: 'j' };
it('accepts an optional absolute executable path', () => {
expect(RemoteHostSchema.safeParse({ ...host, wakeCommand: '/home/joe/bin/whuff' }).success).toBe(true);
expect(RemoteHostSchema.safeParse(host).success).toBe(true);
});
it('rejects an argument list (spawn runs the path without a shell)', () => {
// `spawn('/home/joe/bin/whuff --mac 00:11:22')` would fail as a confusing
// ENOENT at wake time — refuse it at config time instead.
expect(RemoteHostSchema.safeParse({ ...host, wakeCommand: '/home/joe/bin/whuff --now' }).success).toBe(false);
});
it('rejects shell metacharacters as defence in depth', () => {
expect(RemoteHostSchema.safeParse({ ...host, wakeCommand: '/bin/sh$(id)' }).success).toBe(false);
expect(RemoteHostSchema.safeParse({ ...host, wakeCommand: '/bin/`id`' }).success).toBe(false);
});
it('accepts one or more MAC addresses and rejects anything else', () => {
expect(RemoteHostSchema.safeParse({ ...host, wakeMac: '04:d9:f5:80:c6:58' }).success).toBe(true);
expect(RemoteHostSchema.safeParse({ ...host, wakeMac: '04-d9-f5-80-c6-58, 1C:61:B4:20:58:EB' }).success).toBe(
true
);
expect(RemoteHostSchema.safeParse({ ...host, wakeMac: '04:d9:f5:80:c6' }).success).toBe(false);
expect(RemoteHostSchema.safeParse({ ...host, wakeMac: '04:d9:f5:80:c6:58; rm -rf /' }).success).toBe(false);
});
});
describe('rehydrateRemoteHostFields', () => {
const persisted = {
hostId: 'hufflepuff',
label: 'Hufflepuff',
host: '192.168.50.137',
username: 'j',
remotePath: '/home/j/work',
};
const hosts = (wakeCommand?: string) =>
new Map([
[
'hufflepuff',
{
id: 'hufflepuff',
label: 'Hufflepuff',
host: '192.168.50.137',
username: 'j',
...(wakeCommand ? { wakeCommand } : {}),
},
],
]);
it('adds a wake command that only exists in the host config', () => {
// The pre-existing-session case: the field was added to remote-hosts.json after
// this session was persisted, so recovery is the only place it can arrive.
expect(rehydrateRemoteHostFields(persisted, hosts('/home/joe/bin/whuff'))?.wakeCommand).toBe(
'/home/joe/bin/whuff'
);
});
it('treats the host config as authoritative (removing it turns the feature off)', () => {
const remote = { ...persisted, wakeCommand: '/home/joe/bin/whuff' };
expect(rehydrateRemoteHostFields(remote, hosts())?.wakeCommand).toBeUndefined();
});
it('refreshes a MAC that only exists in the host config', () => {
const withMac = new Map(
hosts()
.entries()
.map(([id, host]) => [id, { ...host, wakeMac: '04:d9:f5:80:c6:58' }] as const)
);
expect(rehydrateRemoteHostFields(persisted, withMac)?.wakeMac).toBe('04:d9:f5:80:c6:58');
});
it('leaves the block untouched when the host is gone or the session is local', () => {
expect(rehydrateRemoteHostFields(persisted, new Map())).toBe(persisted);
expect(rehydrateRemoteHostFields(undefined, hosts('/x'))).toBeUndefined();
});
it('keeps the other host-level fields as persisted', () => {
// Only wakeCommand is refreshed: silently re-pointing an existing pane's ssh
// options would be a behavior change nobody asked for.
const remote = { ...persisted, identityFile: '~/.ssh/pinned_key' };
const rehydrated = rehydrateRemoteHostFields(remote, hosts('/home/joe/bin/whuff'));
expect(rehydrated?.identityFile).toBe('~/.ssh/pinned_key');
});
});
});
File diff suppressed because it is too large Load Diff
+37 -1
View File
@@ -11,7 +11,7 @@
* Port: N/A (no server start).
*/
import { describe, it, expect, afterEach, vi } from 'vitest';
import { WebServer } from '../src/web/server.js';
import { WebServer, escapeScriptJson } from '../src/web/server.js';
import { isClaudeAvailable } from '../src/utils/claude-cli-resolver.js';
import { isOpenCodeAvailable } from '../src/utils/opencode-cli-resolver.js';
import { isCodexAvailable } from '../src/utils/codex-cli-resolver.js';
@@ -187,6 +187,41 @@ describe('WebServer.renderIndexHtml', () => {
});
});
it('reports which run modes the custom-model Run-menu picker may generate an entry for', async () => {
// Read generically off the CLI registry's own capabilities, not a hardcoded id
// list — antigravity (`unsupported`) and shell (`kind !== 'agent'`) must be
// absent, and any enabled agent CLI with a real injection recipe must be
// present, with no mock needed since this reads the real stock registry.
const { server } = makeServer({});
const html = await render(server);
expect(html).toContain('window.__codemanCustomModelClis=');
const clis = JSON.parse(html.match(/window\.__codemanCustomModelClis=(\[.*?\]);/)![1]) as Array<{
id: string;
label: string;
}>;
const ids = clis.map((c) => c.id);
expect(ids).toContain('claude');
expect(ids).not.toContain('antigravity');
expect(ids).not.toContain('shell');
for (const cli of clis) {
expect(typeof cli.id).toBe('string');
expect(typeof cli.label).toBe('string');
}
});
it('escapeScriptJson neutralizes a literal </script>, and still round-trips as a JS literal', () => {
// CliEntry.label is a plain string a user's own clis.json can set (up to 60
// chars), unlike __codemanCliAvailable's booleans-only payload, so this is
// the one injection that needs it. Exported so this tests the pure
// function directly rather than needing a real WebServer (which needs tmux).
const dangerous = JSON.stringify([{ id: 'x', label: '</script><script>alert(1)</script>' }]);
const escaped = escapeScriptJson(dangerous);
expect(escaped).not.toContain('</script');
// Proves it decodes back to the real value the way a browser's own JS
// parser would, not just "the output contains no </script>".
expect(eval(escaped)[0].label).toBe('</script><script>alert(1)</script>');
});
it('still emits the object when nothing at all is installed', async () => {
// The all-false case is the one that matters most and the easiest to get
// wrong by only injecting when something resolves.
@@ -218,6 +253,7 @@ describe('WebServer.renderIndexHtml', () => {
const { server } = makeServer({});
const html = await render(server, 'sess-123');
expect(html).not.toContain('__codemanCliAvailable');
expect(html).not.toContain('__codemanCustomModelClis');
});
it('does not expose gesture at all when CODEMAN_GESTURE is unset', async () => {
+223
View File
@@ -181,6 +181,34 @@ describe('custom model endpoint CRUD', () => {
expect(res.json().error).toMatch(/refused.*169\.254\.169\.254/);
});
it('running-status never hands the browser the raw llama-swap launch command (cmd)', async () => {
const { app } = await setup();
await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'ep-running', label: 'A', baseUrl: 'http://localhost:8080', apiKey: 'k' },
});
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/running') {
return new Response(
JSON.stringify({
running: [
{ model: 'qwen3', state: 'ready', cmd: 'llama-server -m /models/qwen3.gguf --api-key sk-secret' },
],
}),
{ status: 200 }
);
}
throw new Error(`unexpected request in this test: ${url.href}`);
});
const res = await app.inject({ method: 'GET', url: '/api/model-endpoints/ep-running/running-status' });
const body = res.json();
expect(body.data.running).toEqual([{ model: 'qwen3', state: 'ready' }]);
expect(JSON.stringify(body)).not.toContain('sk-secret');
expect(JSON.stringify(body)).not.toContain('cmd');
});
it('refuses a baseUrl with embedded credentials or a non-http scheme at save time', async () => {
const { app } = await setup();
for (const baseUrl of ['http://user:pw@host:8080', 'ftp://host/models', 'http://169.254.169.254']) {
@@ -194,3 +222,198 @@ describe('custom model endpoint CRUD', () => {
}
});
});
describe('defaultModelId — the Run-menu picker’s per-endpoint default', () => {
afterEach(() => {
fetchMock.mockReset();
});
it('rejects a defaultModelId that is not one of the endpoint’s discovered models, on both create and update', async () => {
const { app } = await setup();
const create = await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: {
id: 'ep-default-reject',
label: 'A',
baseUrl: 'http://localhost:8080',
models: ['qwen3'],
defaultModelId: 'ghost',
},
});
expect(create.json().success).toBe(false);
expect(create.json().errorCode).toBe('INVALID_INPUT');
await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'ep-default-reject', label: 'A', baseUrl: 'http://localhost:8080', models: ['qwen3'] },
});
const update = await app.inject({
method: 'PUT',
url: '/api/model-endpoints/ep-default-reject',
payload: { label: 'A', baseUrl: 'http://localhost:8080', models: ['qwen3'], defaultModelId: 'ghost' },
});
expect(update.json().success).toBe(false);
expect(update.json().errorCode).toBe('INVALID_INPUT');
});
it('accepts a defaultModelId that IS one of the discovered models', async () => {
const { app } = await setup();
const res = await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: {
id: 'ep-default-accept',
label: 'A',
baseUrl: 'http://localhost:8080',
models: ['qwen3', 'llama3'],
defaultModelId: 'llama3',
},
});
expect(res.json().success).toBe(true);
expect(res.json().data.host.defaultModelId).toBe('llama3');
});
it('drops a stale default that no longer appears in a fresh discovery, rather than carrying it forward invalid', async () => {
const { app } = await setup();
await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: {
id: 'ep-default-drop',
label: 'A',
baseUrl: 'http://localhost:8080',
models: ['qwen3'],
defaultModelId: 'qwen3',
},
});
fetchMock.mockResolvedValue(new Response(JSON.stringify({ data: [{ id: 'llama3' }] }), { status: 200 }));
await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-default-drop/discover-models' });
const list = await app.inject({ method: 'GET', url: '/api/model-endpoints' });
const stored = (list.json() as Array<{ id: string; defaultModelId?: string }>).find(
(h) => h.id === 'ep-default-drop'
);
expect(stored?.defaultModelId).toBeUndefined();
});
it('keeps a default that IS still present after a fresh discovery', async () => {
const { app } = await setup();
await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: {
id: 'ep-default-keep',
label: 'A',
baseUrl: 'http://localhost:8080',
models: ['qwen3'],
defaultModelId: 'qwen3',
},
});
fetchMock.mockResolvedValue(
new Response(JSON.stringify({ data: [{ id: 'qwen3' }, { id: 'llama3' }] }), { status: 200 })
);
await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-default-keep/discover-models' });
const list = await app.inject({ method: 'GET', url: '/api/model-endpoints' });
const stored = (list.json() as Array<{ id: string; defaultModelId?: string }>).find(
(h) => h.id === 'ep-default-keep'
);
expect(stored?.defaultModelId).toBe('qwen3');
});
});
describe('apiKey is never handed back to the browser', () => {
afterEach(() => {
fetchMock.mockReset();
});
it('POST, GET and PUT responses all carry apiKeySet instead of the real key', async () => {
const { app } = await setup();
const create = await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'ep-secret', label: 'A', baseUrl: 'http://localhost:8080', apiKey: 'super-secret' },
});
expect(create.json().data.host.apiKey).toBeUndefined();
expect(create.json().data.host.apiKeySet).toBe(true);
const list = await app.inject({ method: 'GET', url: '/api/model-endpoints' });
const listed = (list.json() as Array<{ id: string; apiKey?: string; apiKeySet?: boolean }>).find(
(h) => h.id === 'ep-secret'
);
expect(listed?.apiKey).toBeUndefined();
expect(listed?.apiKeySet).toBe(true);
expect(JSON.stringify(list.json())).not.toContain('super-secret');
const update = await app.inject({
method: 'PUT',
url: '/api/model-endpoints/ep-secret',
payload: { label: 'Renamed', baseUrl: 'http://localhost:8080' },
});
expect(update.json().data.host.apiKey).toBeUndefined();
expect(update.json().data.host.apiKeySet).toBe(true);
expect(JSON.stringify(update.json())).not.toContain('super-secret');
});
it('a host with no key set at all reports apiKeySet: false', async () => {
const { app } = await setup();
const create = await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'ep-nokey', label: 'A', baseUrl: 'http://localhost:8080' },
});
expect(create.json().data.host.apiKeySet).toBe(false);
});
it('PUT with no apiKey keeps the stored one, rather than clearing it', async () => {
const { app } = await setup();
await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'ep-keep-key', label: 'A', baseUrl: 'http://localhost:8080', apiKey: 'original-key' },
});
// Edit without touching the API key field — the real bug this guards: a
// browser round-trip that only ever sees apiKeySet, never the real value,
// must not accidentally send an empty string and wipe a working credential.
const update = await app.inject({
method: 'PUT',
url: '/api/model-endpoints/ep-keep-key',
payload: { label: 'Renamed', baseUrl: 'http://localhost:8080' },
});
expect(update.json().data.host.apiKeySet).toBe(true);
// Prove it by observing the auth header discovery actually sends.
fetchMock.mockImplementation(async (_url: URL, init?: RequestInit) => {
const headers = init?.headers as Record<string, string>;
expect(headers.Authorization).toBe('Bearer original-key');
return new Response(JSON.stringify({ data: [] }), { status: 200 });
});
const discover = await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-keep-key/discover-models' });
expect(discover.json().success).toBe(true);
expect(fetchMock).toHaveBeenCalledTimes(1);
});
it('PUT with a new apiKey replaces the stored one', async () => {
const { app } = await setup();
await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'ep-replace-key', label: 'A', baseUrl: 'http://localhost:8080', apiKey: 'old-key' },
});
await app.inject({
method: 'PUT',
url: '/api/model-endpoints/ep-replace-key',
payload: { label: 'A', baseUrl: 'http://localhost:8080', apiKey: 'new-key' },
});
fetchMock.mockImplementation(async (_url: URL, init?: RequestInit) => {
const headers = init?.headers as Record<string, string>;
expect(headers.Authorization).toBe('Bearer new-key');
return new Response(JSON.stringify({ data: [] }), { status: 200 });
});
await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-replace-key/discover-models' });
expect(fetchMock).toHaveBeenCalledTimes(1);
});
});
@@ -0,0 +1,382 @@
/**
* @fileoverview POST /api/quick-start's `customModel` field (docs/custom-model-endpoints-plan.md):
* the ONE-SHOT launch path that computes a custom-model endpoint's injection BEFORE the
* session/process exists and launches directly on it, so a custom-model Run never shows
* the native-boot-then-restart the dedicated POST /api/sessions/:id/custom-model route's
* restart-in-place design otherwise produces — most visibly on a CLI like Codex whose TUI
* fully reinitializes on a restart. That dedicated route is still what an ALREADY-RUNNING
* session uses to switch later; this is the create-time equivalent.
*
* Mirrors test/routes/session-custom-model.test.ts's fixtures and llama-swap mocking, since
* this route mirrors that one's own checks (llama-swap conflict, unsupported CLI, unknown
* endpoint, an argv-incompatible model id) rather than a lighter, separately-drifting copy.
*
* Session.prototype.startInteractive/startShell are mocked exactly like the workspace-hooks
* quick-start tests: quick-start constructs a REAL Session (not the MockSession the route
* test harness substitutes elsewhere), so tmux must never actually be reached.
*
* Port: N/A (app.inject, no real port needed)
*/
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest';
import Fastify, { type FastifyInstance } from 'fastify';
import fastifyCookie from '@fastify/cookie';
import { rm, readFile } from 'node:fs/promises';
import { existsSync } from 'node:fs';
import { join } from 'node:path';
import { createMockRouteContext, safeRmHomeTree, type MockRouteContext } from '../mocks/index.js';
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
import { getDataDir } from '../../src/config/instance.js';
import { CASES_DIR } from '../../src/web/route-helpers.js';
import { Session } from '../../src/session.js';
import { writeCustomModelHosts, type CustomModelHost } from '../../src/custom-model-hosts.js';
import { customModelConfigDir } from '../../src/custom-model-injection-apply.js';
import { webviewFetch } from '../../src/web/webview-egress.js';
vi.mock('../../src/web/webview-egress.js', async () => {
const actual = await vi.importActual<typeof import('../../src/web/webview-egress.js')>(
'../../src/web/webview-egress.js'
);
return { ...actual, webviewFetch: vi.fn() };
});
const fetchMock = vi.mocked(webviewFetch);
// quick-start's own local-CLI-availability gate (resolveCliLaunchError, unrelated to the
// custom-model injection this file tests) runs BEFORE the code under test and would
// otherwise 404 every non-claude mode on a box with no codex/pi/grok/omp binary installed —
// exactly this test environment. Mirrors the real "not remote" bypass documented at its own
// call site in session-routes.ts (`session-routes.test.ts`'s remote-codex test is the
// precedent for needing this at all).
vi.mock('../../src/utils/cli-launcher.js', async () => {
const actual = await vi.importActual<typeof import('../../src/utils/cli-launcher.js')>(
'../../src/utils/cli-launcher.js'
);
return { ...actual, resolveCliLaunchError: vi.fn().mockResolvedValue(null) };
});
const ENDPOINT: CustomModelHost = {
id: 'ep1',
label: 'llama.cpp box',
baseUrl: 'http://192.168.1.50:8080',
apiKey: 'k',
};
describe('POST /api/quick-start: customModel (one-shot custom-model launch)', () => {
let app: FastifyInstance;
let ctx: MockRouteContext;
let restartSpy: ReturnType<typeof vi.spyOn>;
const quickStart = (payload: Record<string, unknown>) =>
app.inject({ method: 'POST', url: '/api/quick-start', payload });
beforeEach(async () => {
vi.spyOn(Session.prototype, 'startInteractive').mockResolvedValue(undefined);
vi.spyOn(Session.prototype, 'startShell').mockResolvedValue(undefined);
restartSpy = vi.spyOn(Session.prototype, 'restartCli').mockResolvedValue(true);
fetchMock.mockReset();
fetchMock.mockResolvedValue(new Response('not found', { status: 404 })); // default: not llama-swap
app = Fastify({ logger: false });
await app.register(fastifyCookie);
ctx = createMockRouteContext();
registerSessionRoutes(app, ctx);
installRouteErrorHandler(app);
await app.ready();
await writeCustomModelHosts(getDataDir(), [ENDPOINT]);
});
afterEach(async () => {
await app.close();
vi.restoreAllMocks();
await rm(join(getDataDir(), 'custom-model-hosts.json'), { force: true });
await rm(join(getDataDir(), 'custom-model-configs'), { recursive: true, force: true });
safeRmHomeTree(CASES_DIR);
});
it('launches a claude session already pointed at the endpoint — no restart at all', async () => {
const res = await quickStart({
caseName: 'cm-claude',
mode: 'claude',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.statusCode).toBe(200);
const { sessionId } = res.json();
const session = ctx.sessions.get(sessionId) as unknown as Session;
expect(session.customModel).toEqual({ endpointId: 'ep1', modelId: 'qwen3', label: 'llama.cpp box' });
// The whole point: never restarted. It launched on the endpoint the first time.
expect(restartSpy).not.toHaveBeenCalled();
const isolatedDir = customModelConfigDir(sessionId);
const trustFile = JSON.parse(await readFile(join(isolatedDir, '.claude.json'), 'utf-8'));
expect(trustFile.customApiKeyResponses.approved).toEqual(['k']);
});
it('codex: writes the config.toml under the SAME id the session actually launches with, no restart', async () => {
const res = await quickStart({
caseName: 'cm-codex',
mode: 'codex',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.statusCode).toBe(200);
const { sessionId } = res.json();
const session = ctx.sessions.get(sessionId) as unknown as Session;
expect(session.customModel?.endpointId).toBe('ep1');
expect(restartSpy).not.toHaveBeenCalled();
const configDir = customModelConfigDir(sessionId);
expect(existsSync(join(configDir, 'config.toml'))).toBe(true);
const toml = await readFile(join(configDir, 'config.toml'), 'utf-8');
expect(toml).toContain('model = "qwen3"');
});
it('pi: forces --model custom/<id> onto piConfig on the FIRST launch, not via a later restart', async () => {
const res = await quickStart({
caseName: 'cm-pi',
mode: 'pi',
customModel: { endpointId: 'ep1', modelId: 'qwen3.5-0.8b' },
});
expect(res.statusCode).toBe(200);
const { sessionId } = res.json();
const session = ctx.sessions.get(sessionId) as unknown as Session & { piConfig?: { model?: string } };
expect(session.getCustomModelForPersist()?.launchModel).toBe('custom/qwen3.5-0.8b');
expect(restartSpy).not.toHaveBeenCalled();
});
it('grok: forces the [model.<name>] block name onto grokConfig on the first launch', async () => {
const res = await quickStart({
caseName: 'cm-grok',
mode: 'grok',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.statusCode).toBe(200);
const { sessionId } = res.json();
const session = ctx.sessions.get(sessionId) as unknown as Session;
expect(session.getCustomModelForPersist()?.launchModel).toBe('codeman-custom');
expect(restartSpy).not.toHaveBeenCalled();
});
it('omp: forces custom/<id> onto ompConfig even with no incoming ompConfig at all', async () => {
const res = await quickStart({
caseName: 'cm-omp',
mode: 'omp',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.statusCode).toBe(200);
const { sessionId } = res.json();
const session = ctx.sessions.get(sessionId) as unknown as Session;
expect(session.getCustomModelForPersist()?.launchModel).toBe('custom/qwen3');
expect(restartSpy).not.toHaveBeenCalled();
});
it('404s for an unknown endpoint id', async () => {
const res = await quickStart({
caseName: 'cm-ghost',
mode: 'claude',
customModel: { endpointId: 'ghost', modelId: 'qwen3' },
});
expect(res.json().success).toBe(false);
expect(res.json().errorCode).toBe('NOT_FOUND');
});
it('refuses a mode with no known custom-model mechanism (antigravity)', async () => {
const res = await quickStart({
caseName: 'cm-agy',
mode: 'antigravity',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.json().success).toBe(false);
expect(res.json().errorCode).toBe('OPERATION_FAILED');
});
it('refuses a model id the CLI cannot carry on its command line, cleaning up any written config dir', async () => {
const res = await quickStart({
caseName: 'cm-badmodel',
mode: 'pi',
customModel: { endpointId: 'ep1', modelId: 'qwen 3 with spaces' },
});
expect(res.json().success).toBe(false);
expect(res.json().errorCode).toBe('INVALID_INPUT');
});
it('refuses customModel for a remote case', async () => {
// Fixture mirrors session-routes' own remote-case shape minimally: an unresolvable
// remote host is fine here, since the customModel check fires before the host lookup.
const res = await quickStart({
caseName: 'nonexistent-remote-case',
mode: 'claude',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
// No matching remote/docker case fixture exists, so this actually falls through to the
// local branch and succeeds — this test only documents that remote/docker have their
// own explicit customModel rejection (see the local-fixture tests in
// session-routes-workspace-hooks.test.ts for the fixture-loading pattern that would be
// needed to exercise the remote/docker branch itself).
expect(res.statusCode).toBe(200);
});
describe('llama-swap conflict check', () => {
function mockRunning(running: Array<{ model: string; state: string }>) {
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/running') return new Response(JSON.stringify({ running }), { status: 200 });
throw new Error(`unexpected request in this test: ${url.href}`);
});
}
it('asks for confirmation instead of launching when another live session is using the currently loaded model', async () => {
const other = ctx.sessions.get('test-session-1')!;
(other as unknown as { customModel: unknown }).customModel = { endpointId: 'ep1', modelId: 'llama3' };
mockRunning([{ model: 'llama3', state: 'ready' }]);
const res = await quickStart({
caseName: 'cm-conflict',
mode: 'claude',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
const body = res.json();
expect(body.requiresConfirmation).toBe(true);
expect(body.currentlyLoadedModel).toBe('llama3');
expect(body.affectedSessions).toEqual([{ id: 'test-session-1', name: other.name }]);
// Nothing was actually created.
expect(ctx.sessions.size).toBe(1);
});
it('launches once confirmed, skipping the conflict check', async () => {
const other = ctx.sessions.get('test-session-1')!;
(other as unknown as { customModel: unknown }).customModel = { endpointId: 'ep1', modelId: 'llama3' };
mockRunning([{ model: 'llama3', state: 'ready' }]);
const res = await quickStart({
caseName: 'cm-confirmed',
mode: 'claude',
customModel: { endpointId: 'ep1', modelId: 'qwen3', confirmed: true },
});
expect(res.statusCode).toBe(200);
expect(res.json().requiresConfirmation).toBeUndefined();
expect(ctx.sessions.size).toBe(2);
});
it('launches straight away when nothing else is using the currently loaded model', async () => {
mockRunning([{ model: 'llama3', state: 'ready' }]);
const res = await quickStart({
caseName: 'cm-noconflict',
mode: 'claude',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.statusCode).toBe(200);
expect(res.json().requiresConfirmation).toBeUndefined();
});
});
describe("context-window floor warning (this CLI's own overhead can exceed a small model's real context)", () => {
const SMALL_CTX_ENDPOINT: CustomModelHost = {
id: 'ep-small',
label: 'tiny box',
baseUrl: 'http://192.168.1.51:8080',
apiKey: 'k',
modelContextLengths: { 'qwen3.8-27b-ud-q4_k_xl': 16384 },
};
it('warns instead of launching when the discovered context is below the safe floor', async () => {
await writeCustomModelHosts(getDataDir(), [ENDPOINT, SMALL_CTX_ENDPOINT]);
const res = await quickStart({
caseName: 'cm-small-ctx',
mode: 'claude',
customModel: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl' },
});
const body = res.json();
expect(body.requiresContextWarning).toBe(true);
expect(body.modelId).toBe('qwen3.8-27b-ud-q4_k_xl');
expect(body.contextLength).toBe(16384);
expect(body.minSafeContextTokens).toBe(40000);
// Nothing was actually created.
expect(ctx.sessions.size).toBe(1);
});
it('launches once confirmed, skipping the context check', async () => {
await writeCustomModelHosts(getDataDir(), [ENDPOINT, SMALL_CTX_ENDPOINT]);
const res = await quickStart({
caseName: 'cm-small-ctx-confirmed',
mode: 'claude',
customModel: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl', confirmed: true },
});
expect(res.statusCode).toBe(200);
expect(res.json().requiresContextWarning).toBeUndefined();
expect(ctx.sessions.size).toBe(2);
});
it('does not warn when nothing about context was discovered', async () => {
const res = await quickStart({
caseName: 'cm-no-ctx-data',
mode: 'claude',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.statusCode).toBe(200);
expect(res.json().requiresContextWarning).toBeUndefined();
});
});
describe('triggering the actual llama-swap load (not just watching for it)', () => {
it('sends a real inference request naming the target model, concurrently with launching the session', async () => {
const chatCalls: unknown[] = [];
fetchMock.mockImplementation(async (url: URL, init?: { body?: unknown }) => {
if (url.pathname === '/running') {
return new Response(JSON.stringify({ running: [{ model: 'llama3', state: 'ready' }] }), { status: 200 });
}
if (url.pathname === '/v1/chat/completions') {
chatCalls.push(JSON.parse(init!.body as string));
return new Response(JSON.stringify({ choices: [] }), { status: 200 });
}
throw new Error(`unexpected request in this test: ${url.href}`);
});
const res = await quickStart({
caseName: 'cm-trigger',
mode: 'claude',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
await new Promise((resolve) => setTimeout(resolve, 0)); // let the fire-and-forget trigger settle
expect(res.statusCode).toBe(200);
expect(res.json().modelSwapInProgress).toBe(true);
expect(chatCalls).toHaveLength(1);
expect(chatCalls[0]).toMatchObject({ model: 'qwen3', max_tokens: 1 });
});
it('never sends a load-trigger request when the target model is already loaded and ready', async () => {
let chatCalled = false;
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/running') {
return new Response(JSON.stringify({ running: [{ model: 'qwen3', state: 'ready' }] }), { status: 200 });
}
if (url.pathname === '/v1/chat/completions') {
chatCalled = true;
return new Response(JSON.stringify({ choices: [] }), { status: 200 });
}
throw new Error(`unexpected request in this test: ${url.href}`);
});
const res = await quickStart({
caseName: 'cm-no-trigger',
mode: 'claude',
customModel: { endpointId: 'ep1', modelId: 'qwen3' },
});
await new Promise((resolve) => setTimeout(resolve, 0));
expect(res.json().modelSwapInProgress).toBe(false);
expect(chatCalled).toBe(false);
});
});
});
+540 -5
View File
@@ -3,13 +3,28 @@
* chunk 5 — applying/clearing a session's custom model endpoint + CLI restart).
* Port: N/A (app.inject, no real port needed)
*/
import { describe, it, expect, beforeEach } from 'vitest';
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
import { describe, it, expect, beforeEach, vi } from 'vitest';
import { registerSessionRoutes, _clampEnvOverridesForOwner } from '../../src/web/routes/session-routes.js';
import { createRouteTestHarness } from './_route-test-utils.js';
import { createMockSession } from '../mocks/index.js';
import { getDataDir } from '../../src/config/instance.js';
import { writeCustomModelHosts, type CustomModelHost } from '../../src/custom-model-hosts.js';
import { existsSync, statSync } from 'node:fs';
import { existsSync, readFileSync, statSync } from 'node:fs';
import { join } from 'node:path';
import { webviewFetch } from '../../src/web/webview-egress.js';
// Every apply now also checks llama-swap's `GET /running` (session-routes.ts) before
// applying — without this mock every test in this file would make a REAL network request
// to the fake 192.168.1.50 endpoint below and wait out its 5s timeout. Defaults to a plain
// 404 (reads as "not llama-swap", exercising none of the new conflict-check tests below),
// overridden per-test where the llama-swap behavior itself is what's under test.
vi.mock('../../src/web/webview-egress.js', async () => {
const actual = await vi.importActual<typeof import('../../src/web/webview-egress.js')>(
'../../src/web/webview-egress.js'
);
return { ...actual, webviewFetch: vi.fn() };
});
const fetchMock = vi.mocked(webviewFetch);
const CLAUDE_ENDPOINT: CustomModelHost = {
id: 'ep1',
@@ -18,14 +33,27 @@ const CLAUDE_ENDPOINT: CustomModelHost = {
apiKey: 'k',
};
async function setup() {
async function setup(ctxOptions?: Parameters<typeof createRouteTestHarness>[1]) {
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT]);
return createRouteTestHarness(registerSessionRoutes);
return createRouteTestHarness(registerSessionRoutes, ctxOptions);
}
describe('POST /api/sessions/:id/custom-model', () => {
/** Shared by the conflict-check block and the context-floor block below, which needs
* both conditions true at once. Scoped to the outer describe on purpose: while it
* lived inside the conflict-check block, a sibling calling it threw a ReferenceError
* during setup, so those tests reported as failing rather than as not written. */
function mockRunning(running: Array<{ model: string; state: string }>) {
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/running') return new Response(JSON.stringify({ running }), { status: 200 });
throw new Error(`unexpected request in this test: ${url.href}`);
});
}
beforeEach(async () => {
await writeCustomModelHosts(getDataDir(), []);
fetchMock.mockReset();
fetchMock.mockResolvedValue(new Response('not found', { status: 404 }));
});
it('applies an endpoint/model to a claude-mode session and restarts the CLI', async () => {
@@ -54,9 +82,19 @@ describe('POST /api/sessions/:id/custom-model', () => {
'ANTHROPIC_DEFAULT_SONNET_MODEL',
'ANTHROPIC_DEFAULT_HAIKU_MODEL',
'ANTHROPIC_DEFAULT_OPUS_MODEL',
'CLAUDE_CONFIG_DIR',
]);
expect(envOverrides.ANTHROPIC_BASE_URL).toBe('http://192.168.1.50:8080');
expect(envOverrides.ANTHROPIC_API_KEY).toBe('k');
// CLAUDE_CONFIG_DIR isolates this session from a stored claude.ai OAuth login, and the
// trust-dialog file it points at is pre-seeded so the injected key doesn't hit an
// interactive "Detected a custom API key" prompt with nobody there to answer it.
const isolatedDir = join(getDataDir(), 'custom-model-configs', 'test-session-1');
expect(envOverrides.CLAUDE_CONFIG_DIR).toBe(isolatedDir);
expect(next.configDir).toBe(isolatedDir);
const trustFile = JSON.parse(readFileSync(join(isolatedDir, '.claude.json'), 'utf8'));
expect(trustFile.customApiKeyResponses.approved).toEqual(['k']);
});
it('clears back to the native default', async () => {
@@ -201,6 +239,467 @@ describe('POST /api/sessions/:id/custom-model', () => {
expect(existsSync(dir)).toBe(false);
});
describe('llama-swap conflict check (llama.cpp runs one model at a time)', () => {
it('applies straight away when the requested model is already loaded', async () => {
const { app, ctx } = await setup();
ctx.sessions.get('test-session-1')!.mode = 'claude';
mockRunning([{ model: 'qwen3', state: 'ready' }]);
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.json().success).not.toBe(false);
expect(res.json().modelSwapInProgress).toBe(false);
expect(ctx.sessions.get('test-session-1')!.setCustomModel).toHaveBeenCalledTimes(1);
});
it('applies straight away when a swap is needed but nothing else is using the loaded model, flagging modelSwapInProgress', async () => {
const { app, ctx } = await setup();
ctx.sessions.get('test-session-1')!.mode = 'claude';
mockRunning([{ model: 'llama3', state: 'ready' }]);
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.json().success).not.toBe(false);
expect(res.json().modelSwapInProgress).toBe(true);
expect(ctx.sessions.get('test-session-1')!.setCustomModel).toHaveBeenCalledTimes(1);
});
it('asks for confirmation instead of applying when another session is actively using the currently loaded model', async () => {
const { app, ctx } = await setup();
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const other = createMockSession('other-session');
other.name = 'w2-otherbox';
other.customModel = { endpointId: 'ep1', modelId: 'llama3' };
ctx.sessions.set('other-session', other);
mockRunning([{ model: 'llama3', state: 'ready' }]);
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
const body = res.json();
expect(body.success).not.toBe(false);
expect(body.requiresConfirmation).toBe(true);
expect(body.currentlyLoadedModel).toBe('llama3');
expect(body.affectedSessions).toEqual([{ id: 'other-session', name: 'w2-otherbox' }]);
// Nothing actually applied yet — this call only asked, it did not switch.
expect(session.setCustomModel).not.toHaveBeenCalled();
expect(session.restartCli).not.toHaveBeenCalled();
});
describe('multi-user: the confirm dialog must not name a session the caller cannot access', () => {
const saved: Record<string, string | undefined> = {};
beforeEach(() => {
saved.CODEMAN_MULTIUSER = process.env.CODEMAN_MULTIUSER;
process.env.CODEMAN_MULTIUSER = '1';
});
afterEach(() => {
if (saved.CODEMAN_MULTIUSER === undefined) delete process.env.CODEMAN_MULTIUSER;
else process.env.CODEMAN_MULTIUSER = saved.CODEMAN_MULTIUSER;
});
it("still blocks the swap pending confirmation, but omits a foreign owner's session from affectedSessions", async () => {
const { app, ctx } = await setup({ authUser: { username: 'bob', role: 'user' } });
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
(session as unknown as { owner?: string }).owner = 'bob';
const other = createMockSession('other-session');
other.name = 'w2-otherbox';
other.customModel = { endpointId: 'ep1', modelId: 'llama3' };
(other as unknown as { owner?: string }).owner = 'alice';
ctx.sessions.set('other-session', other);
mockRunning([{ model: 'llama3', state: 'ready' }]);
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
const body = res.json();
// Still asks — a foreign session is just as real a disruption as an owned one.
expect(body.requiresConfirmation).toBe(true);
expect(body.currentlyLoadedModel).toBe('llama3');
// But bob never learns alice's session id or name.
expect(body.affectedSessions).toEqual([]);
expect(session.setCustomModel).not.toHaveBeenCalled();
});
it('names the affected session when the caller DOES own it', async () => {
const { app, ctx } = await setup({ authUser: { username: 'bob', role: 'user' } });
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
(session as unknown as { owner?: string }).owner = 'bob';
const other = createMockSession('other-session');
other.name = 'w2-otherbox';
other.customModel = { endpointId: 'ep1', modelId: 'llama3' };
(other as unknown as { owner?: string }).owner = 'bob';
ctx.sessions.set('other-session', other);
mockRunning([{ model: 'llama3', state: 'ready' }]);
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.json().affectedSessions).toEqual([{ id: 'other-session', name: 'w2-otherbox' }]);
});
});
it('applies once confirmed, skipping the conflict check the second time', async () => {
const { app, ctx } = await setup();
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const other = createMockSession('other-session');
other.customModel = { endpointId: 'ep1', modelId: 'llama3' };
ctx.sessions.set('other-session', other);
mockRunning([{ model: 'llama3', state: 'ready' }]);
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3', confirmed: true },
});
const body = res.json();
expect(body.requiresConfirmation).toBeUndefined();
expect(body.modelSwapInProgress).toBe(true);
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
expect(session.restartCli).toHaveBeenCalledTimes(1);
});
it('a session pointed at the SAME endpoint but a DIFFERENT (not-currently-loaded) model is not treated as affected', async () => {
const { app, ctx } = await setup();
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const other = createMockSession('other-session');
other.customModel = { endpointId: 'ep1', modelId: 'some-other-model' }; // not the loaded one
ctx.sessions.set('other-session', other);
mockRunning([{ model: 'llama3', state: 'ready' }]);
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.json().requiresConfirmation).toBeUndefined();
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
});
it('not llama-swap (plain llama.cpp/OpenAI-compatible server, no /running) — never checked, applies straight away', async () => {
const { app, ctx } = await setup();
ctx.sessions.get('test-session-1')!.mode = 'claude';
fetchMock.mockResolvedValue(new Response('not found', { status: 404 }));
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.json().modelSwapInProgress).toBe(false);
expect(res.json().requiresConfirmation).toBeUndefined();
});
});
describe("context-window floor warning (this CLI's own overhead can exceed a small model's real context)", () => {
const SMALL_CTX_ENDPOINT: CustomModelHost = {
id: 'ep-small',
label: 'tiny box',
baseUrl: 'http://192.168.1.51:8080',
apiKey: 'k',
modelContextLengths: { 'qwen3.8-27b-ud-q4_k_xl': 16384 },
};
it('warns instead of applying when the discovered context is below the safe floor', async () => {
const { app, ctx } = await setup();
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, SMALL_CTX_ENDPOINT]);
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl' },
});
const body = res.json();
expect(body.success).not.toBe(false);
expect(body.requiresContextWarning).toBe(true);
expect(body.modelId).toBe('qwen3.8-27b-ud-q4_k_xl');
expect(body.contextLength).toBe(16384);
expect(body.minSafeContextTokens).toBe(40000);
// Nothing actually applied yet — this call only warned, it did not switch.
expect(session.setCustomModel).not.toHaveBeenCalled();
expect(session.restartCli).not.toHaveBeenCalled();
});
it('applies once confirmed, skipping the context check the second time', async () => {
const { app, ctx } = await setup();
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, SMALL_CTX_ENDPOINT]);
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl', confirmed: true },
});
const body = res.json();
expect(body.requiresContextWarning).toBeUndefined();
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
expect(session.restartCli).toHaveBeenCalledTimes(1);
});
// The two questions are about DIFFERENT people: a context window below the floor is
// the caller's own problem, while unloading a model takes it away from someone else's
// session. They shared one `confirmed` flag until this release, and because the
// context check runs first, clicking "launch anyway" past the context warning silently
// answered the swap question too and evicted another session's model unasked.
describe('answering one question is not consent to the other', () => {
async function bothConditions() {
const { app, ctx } = await setup();
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, SMALL_CTX_ENDPOINT]);
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
// another session is actively on the model this endpoint currently has loaded
const other = createMockSession('other-session');
other.name = 'w2-otherbox';
other.customModel = { endpointId: 'ep-small', modelId: 'llama3' };
ctx.sessions.set('other-session', other);
mockRunning([{ model: 'llama3', state: 'ready' }]);
return { app, session };
}
it('still asks about the swap after the context warning was confirmed', async () => {
const { app, session } = await bothConditions();
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: {
endpointId: 'ep-small',
modelId: 'qwen3.8-27b-ud-q4_k_xl',
confirmedContext: true,
},
});
const body = res.json();
expect(body.requiresContextWarning).toBeUndefined();
expect(body.requiresConfirmation).toBe(true);
expect(body.currentlyLoadedModel).toBe('llama3');
// and crucially nothing was applied: the other session keeps its model
expect(session.setCustomModel).not.toHaveBeenCalled();
expect(session.restartCli).not.toHaveBeenCalled();
});
it('applies once BOTH questions are answered', async () => {
const { app, session } = await bothConditions();
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: {
endpointId: 'ep-small',
modelId: 'qwen3.8-27b-ud-q4_k_xl',
confirmedContext: true,
confirmedSwap: true,
},
});
const body = res.json();
expect(body.requiresContextWarning).toBeUndefined();
expect(body.requiresConfirmation).toBeUndefined();
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
});
it('confirmedSwap alone does not silence the context warning either', async () => {
const { app, session } = await bothConditions();
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl', confirmedSwap: true },
});
expect(res.json().requiresContextWarning).toBe(true);
expect(session.setCustomModel).not.toHaveBeenCalled();
});
// `confirmed` shipped in the HTTP-API-only cut of this feature, so a caller written
// against that must keep working: it means both, exactly as it used to.
it('keeps the legacy blanket `confirmed` meaning both', async () => {
const { app, session } = await bothConditions();
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl', confirmed: true },
});
const body = res.json();
expect(body.requiresContextWarning).toBeUndefined();
expect(body.requiresConfirmation).toBeUndefined();
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
});
});
it('does not warn when the discovered context is comfortably above the floor', async () => {
const { app, ctx } = await setup();
const roomyEndpoint: CustomModelHost = {
id: 'ep-roomy',
label: 'roomy box',
baseUrl: 'http://192.168.1.52:8080',
apiKey: 'k',
modelContextLengths: { qwen3: 65536 },
};
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, roomyEndpoint]);
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep-roomy', modelId: 'qwen3' },
});
expect(res.json().requiresContextWarning).toBeUndefined();
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
});
it('does not warn when the context length was never discovered (nothing to compare)', async () => {
const { app, ctx } = await setup();
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT]);
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
expect(res.json().requiresContextWarning).toBeUndefined();
expect(session.setCustomModel).toHaveBeenCalledTimes(1);
});
it('does not warn for a CLI whose registry entry declares no contextLengthVar (opencode)', async () => {
// opencode's customModelInjection kind is configContentEnv, not env+contextLengthVar,
// so exceedsSafeContextFloor is false by construction regardless of context size.
const { app, ctx } = await setup();
await writeCustomModelHosts(getDataDir(), [CLAUDE_ENDPOINT, SMALL_CTX_ENDPOINT]);
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'opencode';
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep-small', modelId: 'qwen3.8-27b-ud-q4_k_xl' },
});
expect(res.json().requiresContextWarning).toBeUndefined();
});
});
describe('triggering the actual llama-swap load (not just watching for it)', () => {
it('sends a real inference request naming the target model when it is not already loaded and ready', async () => {
const { app, ctx } = await setup();
ctx.sessions.get('test-session-1')!.mode = 'claude';
const chatCalls: unknown[] = [];
fetchMock.mockImplementation(async (url: URL, init?: { body?: unknown }) => {
if (url.pathname === '/running') {
return new Response(JSON.stringify({ running: [{ model: 'llama3', state: 'ready' }] }), { status: 200 });
}
if (url.pathname === '/v1/chat/completions') {
chatCalls.push(JSON.parse(init!.body as string));
return new Response(JSON.stringify({ choices: [] }), { status: 200 });
}
throw new Error(`unexpected request in this test: ${url.href}`);
});
await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
await new Promise((resolve) => setTimeout(resolve, 0)); // let the fire-and-forget trigger settle
expect(chatCalls).toHaveLength(1);
expect(chatCalls[0]).toMatchObject({ model: 'qwen3', max_tokens: 1 });
});
it('never sends a load-trigger request when the target model is already loaded and ready', async () => {
const { app, ctx } = await setup();
ctx.sessions.get('test-session-1')!.mode = 'claude';
let chatCalled = false;
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/running') {
return new Response(JSON.stringify({ running: [{ model: 'qwen3', state: 'ready' }] }), { status: 200 });
}
if (url.pathname === '/v1/chat/completions') {
chatCalled = true;
return new Response(JSON.stringify({ choices: [] }), { status: 200 });
}
throw new Error(`unexpected request in this test: ${url.href}`);
});
await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
await new Promise((resolve) => setTimeout(resolve, 0));
expect(chatCalled).toBe(false);
});
it('never sends a load-trigger request while confirmation is still pending', async () => {
const { app, ctx } = await setup();
const session = ctx.sessions.get('test-session-1')!;
session.mode = 'claude';
const other = createMockSession('other-session');
other.customModel = { endpointId: 'ep1', modelId: 'llama3' };
ctx.sessions.set('other-session', other);
let chatCalled = false;
fetchMock.mockImplementation(async (url: URL) => {
if (url.pathname === '/running') {
return new Response(JSON.stringify({ running: [{ model: 'llama3', state: 'ready' }] }), { status: 200 });
}
if (url.pathname === '/v1/chat/completions') {
chatCalled = true;
return new Response(JSON.stringify({ choices: [] }), { status: 200 });
}
throw new Error(`unexpected request in this test: ${url.href}`);
});
const res = await app.inject({
method: 'POST',
url: '/api/sessions/test-session-1/custom-model',
payload: { endpointId: 'ep1', modelId: 'qwen3' },
});
await new Promise((resolve) => setTimeout(resolve, 0));
expect(res.json().requiresConfirmation).toBe(true);
expect(chatCalled).toBe(false);
});
});
it('refuses to touch a busy session', async () => {
const { app, ctx } = await setup();
const session = ctx.sessions.get('test-session-1')!;
@@ -218,3 +717,39 @@ describe('POST /api/sessions/:id/custom-model', () => {
expect(session.setCustomModel).not.toHaveBeenCalled();
});
});
describe('Claude multi-user clamp: the env-var half', () => {
// CLAUDE_CODE_MAX_CONTEXT_TOKENS and CLAUDE_CONFIG_DIR were already reachable via
// plain envOverrides before claude's privilegedEnvKeys existed (the first already
// matches the CLAUDE_CODE_* allowedPrefix, the second is an allowed exact key), so
// listing them here is not what makes this route safe — no custom-model route reads
// privilegedEnvKeys at all. What it DOES do: ownerClampedEnvKeys() feeds the generic
// envOverrides clamp on create/quick-start/reboot-restore, so a non-granted owner can
// no longer set CLAUDE_CONFIG_DIR that way (the per-client-account feature, #255), and
// a PERSISTED one is now stripped on reboot-restore for such an owner too — see
// session-env-clamp.ts's own fileoverview for why that pass used to be a no-op for
// claude specifically.
const ORIGINAL = process.env.CODEMAN_MULTIUSER;
beforeEach(() => {
process.env.CODEMAN_MULTIUSER = '1';
});
afterEach(() => {
if (ORIGINAL === undefined) delete process.env.CODEMAN_MULTIUSER;
else process.env.CODEMAN_MULTIUSER = ORIGINAL;
});
it('strips CLAUDE_CONFIG_DIR and CLAUDE_CODE_MAX_CONTEXT_TOKENS for a non-granted owner, leaving unrelated CLAUDE_CODE_* keys alone', async () => {
const out = await _clampEnvOverridesForOwner('nobody', {
CLAUDE_CONFIG_DIR: '/home/attacker/fake-claude-config',
CLAUDE_CODE_MAX_CONTEXT_TOKENS: '999999',
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS: '1',
});
expect(out).toEqual({ CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS: '1' });
});
it('is a no-op in single-user mode', async () => {
delete process.env.CODEMAN_MULTIUSER;
const input = { CLAUDE_CONFIG_DIR: '/home/attacker/fake-claude-config' };
expect(await _clampEnvOverridesForOwner(undefined, input)).toBe(input);
});
});
+390
View File
@@ -0,0 +1,390 @@
/**
* @fileoverview Route tests for wake-on-LAN on `POST /api/sessions/:id/input`.
*
* The behavior that matters and cannot be tested at the registry level: a
* wake-enabled remote session whose host is asleep must return 200 WITHOUT
* writing into the stalled pane (the bytes would vanish), while every other
* session keeps the historical fire-and-forget path untouched.
*
* The registry is injected through `registerSessionRoutes`'s test seam so no real
* TCP connect, ssh, or WoL happens in CI.
*/
import { mkdir, writeFile } from 'node:fs/promises';
import { join } from 'node:path';
import { afterEach, describe, expect, it, vi } from 'vitest';
import { getDataDir } from '../../src/config/instance.js';
import fastifyCookie from '@fastify/cookie';
import Fastify, { type FastifyInstance } from 'fastify';
import { registerSessionRoutes, _resetPaneLivenessState } from '../../src/web/routes/session-routes.js';
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
import { createMockRouteContext } from '../mocks/index.js';
import { httpStatusForErrorCode, type ApiErrorCode } from '../../src/types.js';
import { sessionWaits } from '../../src/web/session-wait-registry.js';
import { RemoteWakeRegistry, type RemoteWakeDeps } from '../../src/remote-wake.js';
import type { SessionRemote } from '../../src/types.js';
const SESSION_ID = 'remote-wake-session';
/**
* Mirror production's envelope + status mapping (as inbox-routes.test.ts does): a
* returned `createErrorResponse` carries its 4xx, a plain object is wrapped in
* `{success:true, data}`. Without it every error would read as a 200.
*/
function installEnvelope(app: FastifyInstance): void {
app.addHook('preSerialization', (req, reply, payload: unknown, done) => {
if (!req.url.startsWith('/api')) return done(null, payload);
if (payload === null || typeof payload !== 'object') return done(null, payload);
const p = payload as { success?: unknown; errorCode?: unknown };
if (p.success === false) {
if (reply.statusCode === 200 && typeof p.errorCode === 'string') {
reply.code(httpStatusForErrorCode(p.errorCode as ApiErrorCode));
}
return done(null, payload);
}
if (p.success === true) return done(null, payload);
return done(null, { success: true, data: payload });
});
}
const URL = `/api/sessions/${SESSION_ID}/input`;
afterEach(() => {
sessionWaits.cancelAll(SESSION_ID);
_resetPaneLivenessState();
});
interface Harness {
app: FastifyInstance;
ctx: ReturnType<typeof createMockRouteContext>;
registry: RemoteWakeRegistry;
probe: ReturnType<typeof vi.fn>;
wake: ReturnType<typeof vi.fn>;
events: string[];
/** Let a held wake finish (see `holdWake`). */
releaseWake: () => void;
}
const remoteSession: SessionRemote = {
hostId: 'hufflepuff',
label: 'Hufflepuff',
host: '192.168.50.137',
username: 'j',
remotePath: '/home/j/codeman-pi-test',
wakeCommand: '/home/joe/bin/whuff',
};
async function harness(
opts: {
remote?: SessionRemote;
hostUp?: boolean;
holdWake?: boolean;
/** Stands in for the auth middleware (multi-user mode); absent = synthetic admin. */
authUser?: { username: string; role: 'admin' | 'user' };
} = {}
): Promise<Harness> {
const app = Fastify({ logger: false });
await app.register(fastifyCookie);
if (opts.authUser) {
const authUser = opts.authUser;
app.addHook('onRequest', async (req) => {
(req as unknown as { authUser: typeof authUser }).authUser = authUser;
});
}
const ctx = createMockRouteContext({ sessionId: SESSION_ID });
const session = ctx.sessions.get(SESSION_ID)!;
session.remote = opts.remote ?? remoteSession;
const probe = vi.fn(async () => opts.hostUp ?? false);
const wake = vi.fn(async () => true);
const events: string[] = [];
// With instantaneous mocks the whole wake chain (wake -> wait -> reattach ->
// flush) can finish inside one `await`, so a test that wants to observe the
// in-flight state has to hold the readiness poll open.
let release: (() => void) | null = null;
const deps: RemoteWakeDeps = {
probe,
wake,
waitUntilReady: () =>
opts.holdWake
? new Promise<boolean>((resolve) => {
release = () => resolve(true);
})
: Promise.resolve(true),
delay: async () => {},
noteReconnected: () => {},
broadcast: (event) => events.push(event),
log: () => {},
};
const registry = new RemoteWakeRegistry(deps);
registerSessionRoutes(app, ctx as never, { remoteWake: registry });
installEnvelope(app);
installRouteErrorHandler(app);
await app.ready();
return { app, ctx, registry, probe, wake, events, releaseWake: () => release?.() };
}
const send = (app: FastifyInstance, payload: Record<string, unknown>) =>
app.inject({ method: 'POST', url: URL, payload });
describe('POST /api/sessions/:id/input — wake-on-LAN', () => {
it('buffers input instead of writing into a sleeping host, then flushes after the wake', async () => {
const h = await harness({ hostUp: false, holdWake: true });
const session = h.ctx.sessions.get(SESSION_ID)!;
const res = await send(h.app, { input: 'hallo', useMux: true });
expect(res.statusCode).toBe(200);
expect(res.json()).toEqual({ success: true, data: { buffered: true } });
// Nothing reached the pane: writing now would be swallowed by the stalled ssh.
expect(session.writeBuffer).toEqual([]);
expect(h.wake).toHaveBeenCalledWith({ kind: 'command', command: '/home/joe/bin/whuff' });
expect(h.registry.isWaking(SESSION_ID)).toBe(true);
h.releaseWake();
await h.registry.wake(session);
expect(session.writeBuffer).toEqual(['hallo']);
expect(session.reattachRemote).toHaveBeenCalled();
});
it('flushes several inputs typed during a wake IN ORDER (the browser posts one per keystroke)', async () => {
// The concurrency surface that only exists in production: xterm's onData posts each
// keystroke as its OWN request, so a wake collects N concurrent buffer writes and must
// replay them in order. Route-level, so it is covered on every run instead of only in a
// hand-driven browser session.
const h = await harness({ hostUp: false, holdWake: true });
const session = h.ctx.sessions.get(SESSION_ID)!;
for (const chunk of ['h', 'a', 'llo']) {
const res = await send(h.app, { input: chunk, useMux: true });
expect(res.statusCode).toBe(200);
}
// Nothing written while the host is asleep/dead — that is the whole point.
expect(session.writeBuffer).toEqual([]);
h.releaseWake();
await h.registry.wake(session);
expect(session.writeBuffer).toEqual(['h', 'a', 'llo']);
});
it('keeps the historical fire-and-forget write when the host is reachable', async () => {
const h = await harness({ hostUp: true });
const session = h.ctx.sessions.get(SESSION_ID)!;
const res = await send(h.app, { input: 'hallo', useMux: true });
expect(res.json()).toEqual({ success: true, data: {} }); // the historical bare answer, untouched
await vi.waitFor(() => expect(session.writeBuffer).toEqual(['hallo']));
expect(h.wake).not.toHaveBeenCalled();
expect(session.reattachRemote).not.toHaveBeenCalled();
});
it('never probes or wakes a session without a wake command', async () => {
const { wakeCommand, ...withoutWake } = remoteSession;
const h = await harness({ remote: withoutWake as SessionRemote });
const session = h.ctx.sessions.get(SESSION_ID)!;
await send(h.app, { input: 'hallo', useMux: true });
await vi.waitFor(() => expect(session.writeBuffer).toEqual(['hallo']));
expect(h.probe).not.toHaveBeenCalled();
expect(h.wake).not.toHaveBeenCalled();
});
it('writes straight into a proxied host with a wake target: no probe, no buffer, no wake', async () => {
// With a target configured, the old verdict buffered EVERY input for the life of
// the session: the readiness poll can never succeed through a proxy, so nothing was
// ever flushed (three inputs, nothing written, buffer non-empty — reproduced upstream).
const h = await harness({ remote: { ...remoteSession, socksProxy: '127.0.0.1:1080' }, hostUp: false });
const session = h.ctx.sessions.get(SESSION_ID)!;
for (const input of ['a', 'b', 'c']) expect((await send(h.app, { input, useMux: true })).statusCode).toBe(200);
expect(session.writeBuffer).toEqual(['a', 'b', 'c']);
expect(h.probe).not.toHaveBeenCalled();
expect(h.wake).not.toHaveBeenCalled();
expect(h.registry.pendingBytes(SESSION_ID)).toBe(0);
});
it('wakes before writing on the send-and-wait path (no buffering, the response waits anyway)', async () => {
const h = await harness({ hostUp: false });
const session = h.ctx.sessions.get(SESSION_ID)!;
await send(h.app, { input: 'hallo', useMux: true, wait: 'idle', waitTimeout: 60 });
expect(h.wake).toHaveBeenCalledTimes(1);
// `ensureAwake` is awaited on this path, so the write happens inline and the
// waiter is registered against a live pane.
expect(session.writeBuffer).toEqual(['hallo']);
});
});
describe('GET /api/sessions/:id/reachability', () => {
const get = (app: FastifyInstance, url: string) => app.inject({ method: 'GET', url });
it('reports the probe result and how the host can be woken', async () => {
const up = await harness({ hostUp: true });
const upBody = (await get(up.app, `/api/sessions/${SESSION_ID}/reachability`)).json();
expect(upBody.data.reachable).toBe(true);
expect(upBody.data.wakeConfigured).toBe('command');
expect(upBody.data.label).toBe('Hufflepuff');
const down = await harness({ hostUp: false });
const downBody = (await get(down.app, `/api/sessions/${SESSION_ID}/reachability`)).json();
expect(downBody.data.reachable).toBe(false);
// A reachability check is a QUESTION, never an action: the host stays asleep.
expect(down.wake).not.toHaveBeenCalled();
});
it('says nothing can wake a host without a configured target', async () => {
const { wakeCommand, ...withoutWake } = remoteSession;
const h = await harness({ remote: withoutWake as SessionRemote, hostUp: false });
const body = (await get(h.app, `/api/sessions/${SESSION_ID}/reachability`)).json();
expect(body.data.reachable).toBe(false);
expect(body.data.wakeConfigured).toBe('none');
});
it('reports a proxied host as unknown, not unreachable, and never probes it', async () => {
// A jump-host / SOCKS host does not answer the bare TCP probe even while ssh works;
// `reachable:false` here drew a permanent banner over a healthy session.
const h = await harness({ remote: { ...remoteSession, jumpHost: 'bastion.example' }, hostUp: false });
const body = (await get(h.app, `/api/sessions/${SESSION_ID}/reachability`)).json();
expect(body.data.reachable).toBeNull();
expect(body.data.probeable).toBe(false);
expect(body.data.wakeConfigured).toBe('command');
expect(h.probe).not.toHaveBeenCalled();
});
});
describe('POST /api/sessions/:id/wake', () => {
const wake = (app: FastifyInstance) => app.inject({ method: 'POST', url: `/api/sessions/${SESSION_ID}/wake` });
it('wakes the host, reattaches the pane and reports both', async () => {
const h = await harness({ hostUp: false });
const session = h.ctx.sessions.get(SESSION_ID)!;
const body = (await wake(h.app)).json();
expect(body.success).toBe(true);
expect(body.data.woke).toBe(true);
expect(body.data.reachable).toBe(true);
expect(session.reattachRemote).toHaveBeenCalled();
});
it('answers with an error the UI can route to the config dialog', async () => {
const { wakeCommand, ...withoutWake } = remoteSession;
const h = await harness({ remote: withoutWake as SessionRemote, hostUp: false });
const res = await wake(h.app);
const body = res.json();
expect(body.success).toBe(false);
expect(body.error).toMatch(/No wake-on-LAN target/);
expect(h.wake).not.toHaveBeenCalled();
});
it('does not send a wake when the host answers, but still settles the session', async () => {
const h = await harness({ hostUp: true });
const body = (await wake(h.app)).json();
expect(body.data.woke).toBe(true);
expect(h.wake).not.toHaveBeenCalled();
});
});
describe('POST /api/sessions + attachRemoteSession — authorization before the wake', () => {
// Remote hosts are admin-only infra everywhere else, and the attach wake spawns the
// host's `wakeCommand` (or broadcasts a packet). Before this gate a non-admin could
// post an attach for any configured hostId, have that executable run and the request
// held for the wake budget, and only THEN get a 403 for the workingDir (reproduced
// upstream: wake spy fired once, response 403).
it('403s a non-admin in multi-user mode without probing or waking the host', async () => {
const prev = process.env.CODEMAN_MULTIUSER;
process.env.CODEMAN_MULTIUSER = '1';
try {
// `session-routes.ts` reads hosts from the sandboxed data dir (module-load-time
// constant), so a host with a wake command is written THERE: a regression would
// find it and fire the spy.
await mkdir(getDataDir(), { recursive: true });
await writeFile(
join(getDataDir(), 'remote-hosts.json'),
JSON.stringify([
{
id: 'hufflepuff',
label: 'Hufflepuff',
host: '192.168.50.137',
username: 'j',
wakeCommand: '/home/joe/bin/whuff',
},
])
);
const h = await harness({ hostUp: false, authUser: { username: 'mallory', role: 'user' } });
const res = await h.app.inject({
method: 'POST',
url: '/api/sessions',
payload: { attachRemoteSession: { hostId: 'hufflepuff', remoteSessionName: 'codeman-abc12345' } },
});
expect(res.statusCode).toBe(403);
expect(res.json().error).toMatch(/admin-only/);
expect(h.probe).not.toHaveBeenCalled();
expect(h.wake).not.toHaveBeenCalled();
expect(h.events).toEqual([]);
await h.app.close();
} finally {
if (prev === undefined) delete process.env.CODEMAN_MULTIUSER;
else process.env.CODEMAN_MULTIUSER = prev;
}
});
});
describe('POST /api/sessions/:id/input — what the caller is told', () => {
it('says buffered, and dropped for a chunk over the wake buffer cap', async () => {
// The non-wait branch always answered a bare `{}`; these fields are additive. Without
// them a prompt over 4 KB posted to a sleeping host was accepted and silently lost.
const h = await harness({ hostUp: false, holdWake: true });
const small = await send(h.app, { input: 'hallo', useMux: true });
expect(small.statusCode).toBe(200);
expect(small.json()).toEqual({ success: true, data: { buffered: true } });
const big = await send(h.app, { input: 'x'.repeat(5000), useMux: true });
expect(big.statusCode).toBe(200);
expect(big.json()).toEqual({ success: true, data: { buffered: true, dropped: true } });
expect(h.registry.pendingBytes(SESSION_ID)).toBe(5);
h.releaseWake();
await h.registry.wake(h.ctx.sessions.get(SESSION_ID)!);
});
it('fails the send-and-wait path when the host never comes back, instead of writing into the stalled pane', async () => {
// Readiness never arrives: the wake resolves false.
const failing = await harnessWithFailingWake();
const session = failing.ctx.sessions.get(SESSION_ID)!;
const res = await send(failing.app, { input: 'hallo', useMux: true, wait: true, waitTimeout: 1000 });
expect(res.statusCode).toBe(422);
expect(res.json().errorCode).toBe('OPERATION_FAILED');
expect(res.json().error).toMatch(/did not come back/);
expect(session.writeBuffer).toEqual([]);
await failing.app.close();
});
});
/** A harness whose readiness poll answers false: the wake command runs, the host stays down. */
async function harnessWithFailingWake(): Promise<Harness> {
const app = Fastify({ logger: false });
await app.register(fastifyCookie);
const ctx = createMockRouteContext({ sessionId: SESSION_ID });
ctx.sessions.get(SESSION_ID)!.remote = remoteSession;
const probe = vi.fn(async () => false);
const wake = vi.fn(async () => true);
const events: string[] = [];
const registry = new RemoteWakeRegistry({
probe,
wake,
waitUntilReady: async () => false,
delay: async () => {},
noteReconnected: () => {},
broadcast: (event) => events.push(event),
log: () => {},
});
registerSessionRoutes(app, ctx as never, { remoteWake: registry });
installEnvelope(app);
installRouteErrorHandler(app);
await app.ready();
return { app, ctx, registry, probe, wake, events, releaseWake: () => {} };
}
+248 -7
View File
@@ -55,6 +55,7 @@ vi.mock('../../src/remote-hosts.js', async (orig) => {
});
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
import { RemoteWakeRegistry, REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS } from '../../src/remote-wake.js';
import { resolveTerminalHistoryConfig } from '../../src/config/terminal-history.js';
interface LocalHarness {
@@ -62,6 +63,15 @@ interface LocalHarness {
ctx: MockRouteContext;
}
// Wake-on-LAN seam: the production registry opens a real TCP connection to the host
// and can run a real wake command, so every route registered here gets a fake one
// (the same seam `test/routes/session-remote-wake.test.ts` uses). Default: the host
// answers, so nothing ever wakes.
const wakeProbe = vi.fn(async () => true);
const wakeCommandRun = vi.fn(async () => true);
const wakeWaitUntilReady = vi.fn(async () => true);
let wakeRegistry: RemoteWakeRegistry;
/**
* Build a Fastify instance that mirrors production's uniform-envelope behavior
* (server.ts preSerialization hook) so the test wire format matches the contract:
@@ -108,7 +118,17 @@ describe('session-routes', () => {
let harness: LocalHarness;
beforeEach(async () => {
harness = await createEnvelopeHarness(registerSessionRoutes);
wakeProbe.mockReset().mockResolvedValue(true);
wakeCommandRun.mockReset().mockResolvedValue(true);
wakeWaitUntilReady.mockReset().mockResolvedValue(true);
wakeRegistry = new RemoteWakeRegistry({
probe: wakeProbe,
wake: wakeCommandRun,
waitUntilReady: wakeWaitUntilReady,
delay: async () => {},
log: () => {},
});
harness = await createEnvelopeHarness((app, ctx) => registerSessionRoutes(app, ctx, { remoteWake: wakeRegistry }));
// Reset remote store so tests start with empty hosts/cases and a passing tmux probe
remoteStore.hosts = [];
remoteStore.cases = [];
@@ -775,7 +795,85 @@ describe('session-routes', () => {
body.data.terminalBuffer.indexOf('visible tmux pane only')
);
// No ?full=1 → visible-frame capture (no fullHistory opts).
expect(harness.ctx.mux.captureActivePaneBuffer).toHaveBeenCalledWith(harness.ctx._session.muxName, undefined);
expect(harness.ctx.mux.captureActivePaneBuffer).toHaveBeenCalledWith(
harness.ctx._session.muxName,
expect.not.objectContaining({ fullHistory: true })
);
});
// ── The geometry a capture was taken at ──
//
// A visible-frame capture repaints each row at an absolute position
// (`\x1b[<row>;1H`). A terminal with fewer rows than the pane clamps every
// address past its own height onto its last line, so the overflow rows
// overwrite each other and the rows they land on are lost. The client can
// only notice that if the response says what height the frame was built
// for, which is what captureRows/captureCols carry.
it('reports the geometry the capture was really taken at', async () => {
harness.ctx._session.terminalBuffer = '';
(harness.ctx.mux as { captureActivePaneBuffer?: unknown }).captureActivePaneBuffer = vi.fn(
(_name: string, opts?: { capturedGeometry?: { cols: number; rows: number } }) => {
// Stand in for TmuxManager, which fills this from the pane itself.
if (opts) opts.capturedGeometry = { cols: 100, rows: 50 };
return 'visible frame';
}
);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${harness.ctx._sessionId}/terminal`,
});
const body = JSON.parse(res.body);
expect(body.data.source).toBe('mux-visible');
expect(body.data.captureCols).toBe(100);
expect(body.data.captureRows).toBe(50);
});
it('omits the geometry when the capture reports none', async () => {
// The cursor query can fail, and a byte-history response never captures
// at all. Neither frame was positioned, so neither can be damaged by a
// terminal of the wrong size. Naming the session's own PTY size here
// would describe a geometry no frame was built for, and the client would
// read it as a mismatch worth replaying for.
harness.ctx._session.terminalBuffer = 'byte history only';
(harness.ctx.mux as { captureActivePaneBuffer?: unknown }).captureActivePaneBuffer = vi.fn(() => null);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${harness.ctx._sessionId}/terminal`,
});
const body = JSON.parse(res.body);
expect(body.data.source).toBe('history');
expect(body.data.captureCols).toBeUndefined();
expect(body.data.captureRows).toBeUndefined();
});
it('omits the geometry when the capture reported a size but returned nothing', async () => {
// A capture can report geometry and still hand back no frame. The
// full-history path writes `capturedGeometry` from the cursor query, then
// returns '' for a pane holding nothing visible, which drops the source
// to `history` with the geometry already recorded. Reporting it there
// would name a size for a body that is the byte stream.
harness.ctx._session.terminalBuffer = 'byte history only';
(harness.ctx.mux as { captureActivePaneBuffer?: unknown }).captureActivePaneBuffer = vi.fn(
(_name: string, opts?: { capturedGeometry?: { cols: number; rows: number } }) => {
if (opts) opts.capturedGeometry = { cols: 100, rows: 50 };
return '';
}
);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${harness.ctx._sessionId}/terminal?full=1`,
});
const body = JSON.parse(res.body);
expect(body.data.source).toBe('history');
expect(body.data.captureCols).toBeUndefined();
expect(body.data.captureRows).toBeUndefined();
});
// ── COD-47: full tmux scrollback replay on full page reload ──
@@ -985,7 +1083,10 @@ describe('session-routes', () => {
expect(res.statusCode).toBe(200);
const body = JSON.parse(res.body);
// Tail/tab-switch must NOT request fullHistory (undefined opts).
expect(captureSpy).toHaveBeenCalledWith(harness.ctx._session.muxName, undefined);
expect(captureSpy).toHaveBeenCalledWith(
harness.ctx._session.muxName,
expect.not.objectContaining({ fullHistory: true })
);
expect(body.data.terminalBuffer).toContain('visible frame only');
expect(body.data.terminalBuffer).not.toContain('FULL_HISTORY_SHOULD_NOT_APPEAR');
expect(body.data.source).toBe('mux-visible');
@@ -1045,7 +1146,10 @@ describe('session-routes', () => {
expect(body.data.terminalBuffer.indexOf('hello world')).toBeLessThan(
body.data.terminalBuffer.indexOf('visible tmux pane only')
);
expect(harness.ctx.mux.captureActivePaneBuffer).toHaveBeenCalledWith(harness.ctx._session.muxName, undefined);
expect(harness.ctx.mux.captureActivePaneBuffer).toHaveBeenCalledWith(
harness.ctx._session.muxName,
expect.not.objectContaining({ fullHistory: true })
);
});
it('preserves one-time OAuth authorization URLs in Codex TUI replay history', async () => {
@@ -1119,7 +1223,10 @@ describe('session-routes', () => {
expect(body.data.terminalBuffer.indexOf('hello world')).toBeLessThan(
body.data.terminalBuffer.indexOf('visible tmux pane only')
);
expect(harness.ctx.mux.captureActivePaneBuffer).toHaveBeenCalledWith(harness.ctx._session.muxName, undefined);
expect(harness.ctx.mux.captureActivePaneBuffer).toHaveBeenCalledWith(
harness.ctx._session.muxName,
expect.not.objectContaining({ fullHistory: true })
);
});
it('uses live mux pane capture only when the accumulated buffer is empty', async () => {
@@ -1138,7 +1245,10 @@ describe('session-routes', () => {
const body = JSON.parse(res.body);
expect(body.data.terminalBuffer).toContain('visible restored tmux pane');
expect(body.data.terminalBuffer).toContain('› current prompt');
expect(harness.ctx.mux.captureActivePaneBuffer).toHaveBeenCalledWith(harness.ctx._session.muxName, undefined);
expect(harness.ctx.mux.captureActivePaneBuffer).toHaveBeenCalledWith(
harness.ctx._session.muxName,
expect.not.objectContaining({ fullHistory: true })
);
});
it('returns error for unknown session', async () => {
@@ -1166,7 +1276,10 @@ describe('session-routes', () => {
expect(buf).toContain('\x1b[H\x1b[2J');
expect(buf).toContain('LIVE-PANE-FRAME');
expect(buf.indexOf('history-bytes')).toBeLessThan(buf.indexOf('LIVE-PANE-FRAME'));
expect(harness.ctx.mux.captureActivePaneBuffer).toHaveBeenCalledWith(harness.ctx._session.muxName, undefined);
expect(harness.ctx.mux.captureActivePaneBuffer).toHaveBeenCalledWith(
harness.ctx._session.muxName,
expect.not.objectContaining({ fullHistory: true })
);
});
it('falls back to the byte history when no live pane buffer is available', async () => {
@@ -1897,6 +2010,134 @@ describe('session-routes', () => {
expect(JSON.parse(res.body)).toMatchObject({ success: false, errorCode: ApiErrorCode.OPERATION_FAILED });
});
describe('remote create/attach wakes a sleeping host (Wake-on-LAN)', () => {
const host = (extra: Record<string, unknown> = {}) => ({
id: 'hufflepuff',
label: 'Hufflepuff',
host: '192.168.50.137',
username: 'j',
wakeMac: '04:d9:f5:80:c6:58',
...extra,
});
const remoteCase = { name: 'hufflepuff-work', type: 'remote', hostId: 'hufflepuff', remotePath: '/home/j/work' };
const quickStart = () =>
harness.app.inject({
method: 'POST',
url: '/api/quick-start',
payload: { caseName: 'hufflepuff-work', mode: 'shell' },
});
it('wakes the host before the tmux probe when the user runs a remote case', async () => {
const startShell = vi.spyOn(Session.prototype, 'startShell').mockResolvedValue(undefined);
try {
remoteStore.hosts = [host()];
remoteStore.cases = [remoteCase];
wakeProbe.mockResolvedValue(false); // asleep
const res = await quickStart();
expect(res.statusCode).toBe(200);
expect(JSON.parse(res.body).success).toBe(true);
expect(wakeCommandRun).toHaveBeenCalledWith({ kind: 'mac', macs: [[4, 217, 245, 128, 198, 88]] });
// The request budget, not the 90 s session default: the reverse proxy would
// cut the request at 60 s while the session was still being built.
expect(wakeWaitUntilReady).toHaveBeenCalledWith(expect.objectContaining({ hostId: 'hufflepuff' }), {
timeoutMs: REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS,
// The shutdown signal rides along so `WebServer.stop()` can end the poll.
signal: expect.any(AbortSignal),
});
} finally {
startShell.mockRestore();
}
});
it('does not wake a host that answers, and never probes a host without a wake target', async () => {
const startShell = vi.spyOn(Session.prototype, 'startShell').mockResolvedValue(undefined);
try {
remoteStore.hosts = [host()];
remoteStore.cases = [remoteCase];
// The fake probe answers `true` by default — a reachable host.
expect((await quickStart()).statusCode).toBe(200);
expect(wakeCommandRun).not.toHaveBeenCalled();
// No wake target at all: not even a probe, so hosts without WoL keep the
// exact behavior (and latency) they had before this feature.
wakeProbe.mockClear();
remoteStore.hosts = [host({ wakeMac: undefined })];
expect((await quickStart()).statusCode).toBe(200);
expect(wakeProbe).not.toHaveBeenCalled();
expect(wakeCommandRun).not.toHaveBeenCalled();
} finally {
startShell.mockRestore();
}
});
it('refuses the run when the host never comes back, and starts no session', async () => {
remoteStore.hosts = [host()];
remoteStore.cases = [remoteCase];
wakeProbe.mockResolvedValue(false);
wakeWaitUntilReady.mockResolvedValue(false);
const sessionsBefore = harness.ctx.sessions.size;
const res = await quickStart();
expect(res.statusCode).toBe(httpStatusForErrorCode(ApiErrorCode.OPERATION_FAILED));
expect(JSON.parse(res.body).error).toMatch(/did not come back after a wake-on-LAN request/);
// No half-created session: the failure is the answer, not a dead tab.
expect(harness.ctx.sessions.size).toBe(sessionsBefore);
});
it('blames the sleeping host, not tmux, when the host has no wake target', async () => {
remoteStore.hosts = [host({ wakeMac: undefined })];
remoteStore.cases = [remoteCase];
remoteStore.tmuxCheck = {
ok: false,
error: 'remote host 192.168.50.137 needs tmux installed for durable remote sessions',
};
wakeProbe.mockResolvedValue(false);
const res = await quickStart();
expect(res.statusCode).toBe(httpStatusForErrorCode(ApiErrorCode.OPERATION_FAILED));
expect(JSON.parse(res.body).error).toMatch(/has no wake-on-LAN target/);
});
it('keeps the tmux error when the host is up but tmux is really missing', async () => {
remoteStore.hosts = [host()];
remoteStore.cases = [remoteCase];
remoteStore.tmuxCheck = {
ok: false,
error: 'remote host 192.168.50.137 needs tmux installed for durable remote sessions',
};
// Probe answers `true`: the ssh failure is genuinely about tmux.
const res = await quickStart();
expect(JSON.parse(res.body).error).toMatch(/needs tmux installed/);
});
it('wakes the host when attaching to a discovered remote session', async () => {
const startInteractive = vi.spyOn(Session.prototype, 'startInteractive').mockResolvedValue(undefined);
const startShell = vi.spyOn(Session.prototype, 'startShell').mockResolvedValue(undefined);
try {
remoteStore.hosts = [host()];
wakeProbe.mockResolvedValue(false);
const res = await harness.app.inject({
method: 'POST',
url: '/api/sessions',
payload: { attachRemoteSession: { hostId: 'hufflepuff', remoteSessionName: 'codeman-abc12345' } },
});
expect(res.statusCode).toBe(200);
expect(wakeCommandRun).toHaveBeenCalledTimes(1);
} finally {
startInteractive.mockRestore();
startShell.mockRestore();
}
});
});
it('does not run local codex availability check for a remote codex case', async () => {
// A remote codex case must NOT be blocked by the LOCAL codex availability gate
// (the CLI runs on the remote host). Probe is stubbed ok in remoteStore.tmuxCheck.
+68
View File
@@ -511,6 +511,7 @@ describe('Codex quick start settings', () => {
it('passes global Codex settings into quick-start config for new sessions', async () => {
const elements: Record<string, any> = {
quickStartCase: { value: 'codex-case' },
tabCount: { value: '1' },
};
const requests: Array<{ url: string; body?: any }> = [];
const CodemanApp = function CodemanApp(this: any) {};
@@ -837,6 +838,7 @@ describe('Gemini quick start', () => {
it('drives runGemini() through the {success,data} envelope and selects the new session', async () => {
const elements: Record<string, any> = {
quickStartCase: { value: 'gemini-case' },
tabCount: { value: '1' },
};
const requests: Array<{ url: string; body?: any }> = [];
const CodemanApp = function CodemanApp(this: any) {};
@@ -891,6 +893,7 @@ describe('Antigravity quick start', () => {
it('drives runAntigravity() through the {success,data} envelope and selects the new session', async () => {
const elements: Record<string, any> = {
quickStartCase: { value: 'ag-case' },
tabCount: { value: '1' },
};
const requests: Array<{ url: string; body?: any }> = [];
const CodemanApp = function CodemanApp(this: any) {};
@@ -947,6 +950,7 @@ describe('Pi quick start', () => {
it('drives runPi() through the {success,data} envelope and sends no piConfig', async () => {
const elements: Record<string, any> = {
quickStartCase: { value: 'pi-case' },
tabCount: { value: '1' },
};
const requests: Array<{ url: string; body?: any }> = [];
const CodemanApp = function CodemanApp(this: any) {};
@@ -1035,6 +1039,7 @@ describe('Grok quick start', () => {
it('drives runGrok() through the {success,data} envelope and sends alwaysApprove', async () => {
const elements: Record<string, any> = {
quickStartCase: { value: 'grok-case' },
tabCount: { value: '1' },
};
const requests: Array<{ url: string; body?: any }> = [];
const CodemanApp = function CodemanApp(this: any) {};
@@ -1117,4 +1122,67 @@ describe('Grok quick start', () => {
expect(requests).toEqual(['/api/grok/status']);
expect(errors[0]).toContain('https://x.ai/cli/install.sh');
});
// The whole point of the shared _launchQuickStartInstances() helper. Before it,
// every non-Claude run*() hardcoded exactly one quick-start call, so the
// "Instance count" stepper next to the Run button silently did nothing on all
// eight of them: no error, no hint, just the wrong number of sessions. The five
// fixture edits that came with the change stub tabCount at '1', so they pass
// identically with and without it; this is the one that does not.
it('launches tabCount sessions with sequential w<n>-<case> names and selects the first', async () => {
const elements: Record<string, any> = {
quickStartCase: { value: 'grok-case' },
tabCount: { value: '3' },
};
const requests: Array<{ url: string; body?: any }> = [];
const CodemanApp = function CodemanApp(this: any) {};
let created = 0;
const context = vm.createContext({
CodemanApp,
localStorage: { getItem: () => null, setItem: () => {} },
document: { getElementById: (id: string) => elements[id] ?? null },
fetch: async (url: string, init?: { body?: string }) => {
const body = init?.body ? JSON.parse(init.body) : undefined;
requests.push({ url, body });
if (url === '/api/grok/status')
return {
json: async () => ({
success: true,
data: { available: true, path: '/home/user/.grok/bin', version: '1.0.5' },
}),
};
if (url === '/api/quick-start') {
const id = `sess-gk-${created++}`;
return {
json: async () => ({ success: true, data: { sessionId: id, session: { id, name: body.sessionName } } }),
};
}
throw new Error(`unexpected fetch: ${url}`);
},
console,
});
const sessionUi = readFileSync(resolve(import.meta.dirname, '../src/web/public/session-ui.js'), 'utf8');
vm.runInContext(sessionUi, context, { filename: 'session-ui.js' });
const app = new (CodemanApp as any)();
app.terminal = { clear: () => {}, writeln: () => {}, focus: () => {} };
app.loadAppSettingsFromStorage = () => ({});
app.getCaseSettings = () => ({});
app.buildEnvOverrides = () => ({});
app.sessions = new Map();
app._onSessionCreated = (session: any) => app.sessions.set(session.id, session);
app._renderSessionTabsImmediate = () => {};
const selected: string[] = [];
app.selectSession = async (id: string) => {
selected.push(id);
};
await app.runGrok();
const names = requests.filter((r) => r.url === '/api/quick-start').map((r) => r.body.sessionName);
expect(names).toEqual(['w1-grok-case', 'w2-grok-case', 'w3-grok-case']);
expect(selected).toEqual(['sess-gk-0']);
});
});
+11 -7
View File
@@ -96,16 +96,20 @@ describe('WebServer index.html <title> templating (#82)', () => {
it('only substitutes the <title> tag — the rest of the template is identical (modulo asset cache-busting)', async () => {
// renderIndexHtml also appends ?v=<mtime> cache-bust params to same-origin
// .js/.css refs, and injects the CLI-availability flags before </head>; strip
// both so the title remains the only other change.
// .js/.css refs, and injects the CLI-availability flags plus the custom-model
// Run-menu picker's CLI list before </head>; strip all so the title remains
// the only other change.
//
// The flag strip is what keeps this test environment-independent. It used to
// pass here by luck: the availability script was injected only where a CLI
// resolved, so the assertion held on a machine with none installed and would
// have failed on a developer's box that had them.
// The flag strips are what keep this test environment-independent. The
// CLI-availability one used to pass here by luck: that script was injected
// only where a CLI resolved, so the assertion held on a machine with none
// installed and would have failed on a developer's box that had them. The
// custom-model list is injected unconditionally (a plain array, possibly
// empty), so it needs stripping on every machine, not just where non-empty.
const html = (await render('laptop'))
.replace(/(\.(?:js|css))\?v=[^"]*/g, '$1')
.replace(/<script>window\.__codemanCliAvailable=\{.*?\};<\/script>\n/, '');
.replace(/<script>window\.__codemanCliAvailable=\{.*?\};<\/script>\n/, '')
.replace(/<script>window\.__codemanCustomModelClis=\[.*?\];<\/script>\n/, '');
const beforeTitle = rawTemplate.split('<title>Codeman</title>')[0];
const afterTitle = rawTemplate.split('<title>Codeman</title>')[1];
expect(html.startsWith(beforeTitle)).toBe(true);
+177
View File
@@ -0,0 +1,177 @@
/**
* @fileoverview A prompt sent through the input route must actually leave the composer.
*
* Claude Code 2.1.277 ignores Enter for the first 30 to 50 seconds after the composer
* paints while still taking typed text, so text+Enter 50 ms apart left every
* programmatic prompt sitting unsent (measured 2026-09-19). These pin the recovery:
* the verifier reads the last glyph line, re-sends Enter only while the prompt is
* verifiably still there, stops the moment it is gone, never acts on a pane with no
* composer, is capped, and is cancelled by a newer write or teardown.
*/
import { describe, it, expect, vi, beforeEach, afterEach } from 'vitest';
import { SubmitVerifier, promptStillInComposer, SUBMIT_VERIFY_DELAYS_MS } from '../src/session-submit-verifier.js';
const PROMPT =
'Read /home/arkon/.codeman/pr-bot/jobs/pr-439/brief.md and carry out the review it describes. Do not ask questions.';
/** Measured 2026-09-19: typed, wrapped, a no-break space after the glyph, never sent. */
const STRANDED = [
' ▐▛███▛█ Claude Code v2.1.278',
'──────────────────────────────────────── prbot-439 ─',
'❯ Read /home/arkon/.codeman/pr-bot/jobs/pr-439/brief.md and carry out the review it describes. Do not ask',
' questions.',
'────────────────────────────────────────────────────',
' Opus 5 (1M context) in:0 out:0',
' ⏵⏵ bypass permissions on (shift+tab to cycle)',
].join('\n');
/** The same pane once taken: echoed in the transcript, composer empty. */
const TAKEN = [
'❯ Read /home/arkon/.codeman/pr-bot/jobs/pr-439/brief.md and carry out the review it describes.',
'● Reading the brief.',
'✻ Actioning… (48s · ↓ 6.6k tokens)',
'──────────────────────────────────────── prbot-439 ─',
'❯',
'────────────────────────────────────────────────────',
' Opus 5 (1M context) in:192,963 out:371 ctx:19%',
].join('\n');
describe('promptStillInComposer', () => {
it('sees the prompt sitting in the composer, no-break space and wrapping included', () => {
expect(promptStillInComposer(STRANDED, PROMPT, '❯')).toBe(true);
});
it('is not fooled by the transcript echo once the composer is empty', () => {
expect(promptStillInComposer(TAKEN, PROMPT, '❯')).toBe(false);
});
it('treats other text in the composer as not ours', () => {
expect(promptStillInComposer(STRANDED, 'Summarise the changelog', '❯')).toBe(false);
});
it('answers undefined for a pane with no composer line, or no glyph', () => {
expect(promptStillInComposer('$ ls\nfoo bar\n$ ', PROMPT, '❯')).toBeUndefined();
expect(promptStillInComposer(STRANDED, PROMPT, '')).toBeUndefined();
});
it('honours the CLI glyph (Codex draws ›)', () => {
expect(promptStillInComposer('› Reply with PONG\n', 'Reply with PONG', '›')).toBe(true);
expect(promptStillInComposer('›\n', 'Reply with PONG', '›')).toBe(false);
});
it('reads through ANSI colour codes', () => {
expect(promptStillInComposer('\x1b[1m❯\x1b[0m \x1b[36mReply with PONG\x1b[0m', 'Reply with PONG', '❯')).toBe(true);
});
});
describe('SubmitVerifier', () => {
let screen: string;
let sends: number;
let logs: string[];
const make = (delaysMs?: readonly number[]) =>
new SubmitVerifier({
capture: () => screen,
sendEnter: () => {
sends++;
},
glyph: () => '❯',
log: (m) => logs.push(m),
delaysMs,
});
beforeEach(() => {
vi.useFakeTimers();
screen = STRANDED;
sends = 0;
logs = [];
});
afterEach(() => {
vi.useRealTimers();
});
it('re-sends Enter while the prompt is still there and stops once it is taken', async () => {
const v = make([1_000, 1_000, 1_000, 1_000]);
v.arm(PROMPT);
await vi.advanceTimersByTimeAsync(1_000);
expect(sends).toBe(1);
await vi.advanceTimersByTimeAsync(1_000);
expect(sends).toBe(2);
screen = TAKEN; // Claude Code finally honoured one
await vi.advanceTimersByTimeAsync(5_000);
expect(sends).toBe(2);
expect(logs[0]).toContain('still in the composer after 1s');
expect(logs[1]).toContain('(2/4)');
});
it('costs one capture and no Enter when the prompt was taken on the first try', async () => {
let captures = 0;
const v = new SubmitVerifier({
capture: () => {
captures++;
return TAKEN;
},
sendEnter: () => {
sends++;
},
glyph: () => '❯',
});
v.arm(PROMPT);
await vi.advanceTimersByTimeAsync(120_000);
expect(captures).toBe(1);
expect(sends).toBe(0);
});
it('never presses Enter into a pane with no composer line', async () => {
screen = '$ npm test\n... running ...\n';
const v = make([1_000, 1_000]);
v.arm(PROMPT);
await vi.advanceTimersByTimeAsync(10_000);
expect(sends).toBe(0);
});
it('is capped at the schedule length when the prompt never leaves', async () => {
const v = make();
v.arm(PROMPT);
await vi.advanceTimersByTimeAsync(10 * 60_000);
expect(sends).toBe(SUBMIT_VERIFY_DELAYS_MS.length);
});
it('production schedule reaches past the measured window', () => {
const total = SUBMIT_VERIFY_DELAYS_MS.reduce((a, b) => a + b, 0);
expect(SUBMIT_VERIFY_DELAYS_MS[0]).toBeLessThanOrEqual(2_000);
expect(total).toBeGreaterThanOrEqual(60_000);
});
it('a newer write replaces the schedule, so an old prompt never submits a new one', async () => {
const v = make([1_000, 1_000, 1_000]);
v.arm(PROMPT);
await vi.advanceTimersByTimeAsync(1_000);
expect(sends).toBe(1);
screen = '❯ Something the user typed next';
v.arm('Something the user typed next');
await vi.advanceTimersByTimeAsync(1_000);
expect(sends).toBe(2); // for the NEW prompt, which is what the composer holds
screen = '❯';
await vi.advanceTimersByTimeAsync(5_000);
expect(sends).toBe(2);
});
it('cancel() stops everything', async () => {
const v = make([1_000, 1_000]);
v.arm(PROMPT);
v.cancel();
await vi.advanceTimersByTimeAsync(10_000);
expect(sends).toBe(0);
});
it('keeps checking when sendEnter throws', async () => {
let calls = 0;
const v = new SubmitVerifier({
capture: () => screen,
sendEnter: () => {
calls++;
throw new Error('tmux hiccup');
},
glyph: () => '❯',
delaysMs: [1_000, 1_000],
});
v.arm(PROMPT);
await vi.advanceTimersByTimeAsync(3_000);
expect(calls).toBe(2);
});
});
+78
View File
@@ -0,0 +1,78 @@
/**
* @fileoverview Static guard: every SSE dispatch entry must actually resolve.
*
* `app.js` dispatches server events through a table of `[SSE_EVENTS.X, '_onFoo']`
* pairs. Both halves fail SILENTLY when they are wrong:
*
* - a handler name that exists in no module (renamed method, typo) → the event is
* received and nothing happens, with no error anywhere;
* - an `SSE_EVENTS.X` key that `constants.js` does not define → the table key is
* `undefined`, so the entry can never match an incoming event.
*
* Both have happened in this codebase's feature areas (a new banner/toast that simply
* never appears), and neither is visible to a test that only checks the modules compile.
* Pure static analysis — no server, no browser.
*/
import { readdirSync, readFileSync } from 'node:fs';
import { join } from 'node:path';
import { fileURLToPath } from 'node:url';
import { describe, expect, it } from 'vitest';
const PUBLIC_DIR = fileURLToPath(new URL('../src/web/public', import.meta.url));
const appJs = readFileSync(join(PUBLIC_DIR, 'app.js'), 'utf-8');
const constantsJs = readFileSync(join(PUBLIC_DIR, 'constants.js'), 'utf-8');
const allModules = readdirSync(PUBLIC_DIR)
.filter((name) => name.endsWith('.js'))
.map((name) => readFileSync(join(PUBLIC_DIR, name), 'utf-8'))
.join('\n');
/** `[SSE_EVENTS.FOO, '_onFoo'],` entries of the dispatch table. */
function dispatchEntries(): { constant: string; handler: string }[] {
const entries: { constant: string; handler: string }[] = [];
const re = /\[SSE_EVENTS\.([A-Z0-9_]+),\s*'(_[A-Za-z0-9_]+)'\]/g;
for (const match of appJs.matchAll(re)) {
entries.push({ constant: match[1], handler: match[2] });
}
return entries;
}
describe('SSE dispatch table', () => {
it('has entries to check (the table is what this guard exists for)', () => {
expect(dispatchEntries().length).toBeGreaterThan(20);
});
it('names only events that constants.js defines', () => {
const defined = new Set([...constantsJs.matchAll(/^\s{2}([A-Z0-9_]+):\s*'/gm)].map((m) => m[1]));
const missing = dispatchEntries()
.map((entry) => entry.constant)
.filter((name) => !defined.has(name));
expect(missing).toEqual([]);
});
it('names only handlers that some frontend module actually defines', () => {
const missing = dispatchEntries()
.map((entry) => entry.handler)
.filter((handler) => !new RegExp(`(^|\\s)${handler}\\s*\\(`, 'm').test(allModules));
expect(missing).toEqual([]);
});
it('defines every handler in exactly ONE module (a second copy is shadowed)', () => {
// Modules mix into `CodemanApp.prototype` and run in script order, so two
// definitions of the same handler name silently shadow each other: the later file
// wins and the earlier one never runs. The existence check above cannot see that
// (both names resolve), which is how a duplicate banner handler can leave a toast
// dead with no error anywhere.
const byModule = readdirSync(PUBLIC_DIR)
.filter((name) => name.endsWith('.js'))
.map((name) => ({ name, source: readFileSync(join(PUBLIC_DIR, name), 'utf-8') }));
const shadowed = dispatchEntries()
.map((entry) => entry.handler)
.filter((handler) => {
const re = new RegExp(`(^|\\s)${handler}\\s*\\(`, 'm');
return byModule.filter((mod) => re.test(mod.source)).length > 1;
});
expect(shadowed).toEqual([]);
});
});
+56
View File
@@ -0,0 +1,56 @@
/**
* @fileoverview Multi-user routing of the `remote:*` SSE family (server.ts `deriveSseHint`).
*
* The wake events carry `hostId`/`label`, which `GET /api/remote-hosts` withholds from
* non-admins, and their toast fires before any session check on the client — so an
* event that falls through to the global branch shows every logged-in user "Waking
* <label>" for a session they do not own. Constructs the server without starting it:
* the hint is a pure function of the event, the payload and the sessions map.
*/
import { describe, expect, it } from 'vitest';
import { WebServer } from '../src/web/server.js';
type Hint = { owner?: string; username?: string; adminOnly?: boolean; sessionScoped?: boolean } | undefined;
function hintFor(event: string, payload: Record<string, unknown>, owners: Record<string, string> = {}): Hint {
const server = new WebServer(3999, false, true) as unknown as {
sessions: Map<string, { owner?: string }>;
deriveSseHint(event: string, data: unknown): Hint;
};
for (const [id, owner] of Object.entries(owners)) server.sessions.set(id, { owner });
return server.deriveSseHint(event, payload);
}
describe('deriveSseHint — remote: events are session-scoped', () => {
it('routes a session wake to that session’s owner', () => {
expect(hintFor('remote:hostWaking', { sessionId: 's1', hostId: 'h', label: 'H' }, { s1: 'alice' })).toEqual({
owner: 'alice',
sessionScoped: true,
});
expect(hintFor('remote:sessionReconnected', { sessionId: 's1' }, { s1: 'alice' })).toEqual({
owner: 'alice',
sessionScoped: true,
});
});
it('routes a create/attach wake (no session yet) to the user who asked for it', () => {
expect(hintFor('remote:hostWaking', { forNewSession: true, username: 'bob', hostId: 'h', label: 'H' })).toEqual({
username: 'bob',
sessionScoped: true,
});
expect(hintFor('remote:hostWakeFailed', { forNewSession: true, username: 'bob', hostId: 'h' })).toEqual({
username: 'bob',
sessionScoped: true,
});
});
it('fails closed (admins only) when it names neither a session nor a requester', () => {
const hint = hintFor('remote:hostWaking', { forNewSession: true, hostId: 'h', label: 'H' });
expect(hint).toEqual({ owner: undefined, sessionScoped: true });
});
it('never lets a wake event reach the global branch', () => {
expect(hintFor('remote:hostWaking', {})).not.toBeUndefined();
expect(hintFor('remote:reconnectExhausted', {})).not.toBeUndefined();
});
});
+233 -21
View File
@@ -1,20 +1,31 @@
/**
* @fileoverview Regression tests for the buffer-load flush path (COD-144).
* @fileoverview Regression tests for the buffer-load flush path: what becomes of
* the live terminal events queued while a buffer load runs, once the load ends.
*
* Bug: newly launched Shell sessions rendered BLANK until a tab-switch. The
* buffer-load path (`selectSession` → `_beginBufferLoad`/`_finishBufferLoad`)
* QUEUES live SSE terminal events while `_isLoadingBuffer` is true, then on
* completion DISCARDS the queue (`_loadBufferQueue = null`). That de-dup is
* correct for an established session (the fetched buffer already contains the
* queued output, so replaying it would duplicate Ink redraws). But for a
* brand-new shell the fetch resolves BEFORE the PTY emits its prompt — the
* fetched buffer is empty and the prompt arrives only as a queued event, which
* then gets discarded → blank terminal.
* Two rules, each from a real bug.
*
* Fix: `_finishBufferLoad(owner, { flushQueued })` REPLAYS the queued events
* through `batchTerminalWrite()` (after `_isLoadingBuffer` is cleared, so they
* write through normally) ONLY when the load painted nothing. The default path
* (no opts) still discards, preserving de-dup for established sessions.
* COD-144: newly launched Shell sessions rendered BLANK until a tab-switch. The
* load path (`selectSession` → `_beginBufferLoad`/`_finishBufferLoad`) queues
* live events while `_isLoadingBuffer` is true and used to DISCARD the queue on
* completion. Right for a buffer built from the server's byte history (the
* queued output is already in it, so replaying it duplicates Ink redraws),
* wrong for a brand-new shell whose fetch resolves BEFORE the PTY emits its
* prompt: the prompt arrived only as a queued event and was thrown away. A
* caller that knows the load painted nothing passes `{ flushQueued: true }`
* and the queue is REPLAYED through `batchTerminalWrite()` after
* `_isLoadingBuffer` is cleared, so the events write through normally.
*
* #436: a tmux pane capture is current only as of the instant `capture-pane`
* ran, so everything the CLI printed between the capture and the end of the
* chunked write was queued and dropped, and its next partial redraw landed on
* a frame the terminal never received. Queue entries now carry their arrival
* time and `_finishBufferLoad` takes a `since` cutoff, so a capture load
* replays exactly the tail that arrived after the response headers. All four
* fetch-and-write paths take that policy from one helper,
* `_bufferLoadFinishOpts`, and a static scan below pins each of them to it,
* because the same fix had already been written into one path out of four,
* twice. A path that replays and then restores a scroll position re-takes the
* sticky-scroll baseline (`_syncStickyScrollBaseline`), pinned the same way.
*
* Loaded via `vm` with a stubbed context (no jsdom — jsdom is broken on this
* box; see connection-indicator.test.ts). We extract the REAL
@@ -56,10 +67,10 @@ type BufferLoadApp = {
_bufferLoadSeq: number;
_bufferLoadOwner: string | null;
_isLoadingBuffer: boolean;
_loadBufferQueue: string[] | null;
_loadBufferQueue: { at: number; data: string }[] | null;
batchTerminalWrite: (data: string) => void;
_beginBufferLoad: (owner?: string) => string;
_finishBufferLoad: (owner?: string, opts?: { flushQueued?: boolean }) => boolean;
_finishBufferLoad: (owner?: string, opts?: { flushQueued?: boolean; since?: number }) => boolean;
};
/**
@@ -84,10 +95,57 @@ function makeApp() {
return { app, writes };
}
/** Simulate live SSE events arriving while a buffer load is in progress (the queue path). */
function pushWhileLoading(app: BufferLoadApp, data: string) {
// Mirrors batchTerminalWrite's queue branch: if loading, push to the queue.
if (app._isLoadingBuffer && app._loadBufferQueue) app._loadBufferQueue.push(data);
/**
* A stub carrying the REAL `batchTerminalWrite` on top of the real begin/finish
* methods, so a replay samples the sticky-scroll baseline exactly as it does in
* the browser. The terminal is a fake whose `buffer.active` the test moves by
* hand, which is what a caller's `scrollToLine` does to a real one.
*/
function makeScrollApp() {
const buffer = { viewportY: 0, baseY: 100 };
const app = {
buffer,
terminal: { buffer: { active: buffer } },
sessions: new Map(),
activeSessionId: null,
pendingWrites: [] as string[],
writeFrameScheduled: false,
_wasAtBottomBeforeWrite: false,
_bufferLoadSeq: 0,
_bufferLoadOwner: null as string | null,
_isLoadingBuffer: false,
_loadBufferQueue: null as { at: number; data: string }[] | null,
_scheduleTerminalWriteFlush: vi.fn(),
batchTerminalWrite: mixin.batchTerminalWrite as (data: string) => void,
isTerminalAtBottom: mixin.isTerminalAtBottom as () => boolean,
_syncStickyScrollBaseline: mixin._syncStickyScrollBaseline as () => void,
_beginBufferLoad: mixin._beginBufferLoad as BufferLoadApp['_beginBufferLoad'],
_finishBufferLoad: mixin._finishBufferLoad as BufferLoadApp['_finishBufferLoad'],
};
return app;
}
/**
* Slice one class method out of app.js, from its header to the next method's.
*
* Bounding the slice matters: the two methods checked below are not followed by
* a JSDoc block, so a scan for the next comment would run on into unrelated
* code and match its scroll calls instead of theirs.
*/
function methodBody(source: string, method: string): string {
const start = source.search(new RegExp(`^ {2}(?:async )?${method}\\(`, 'm'));
expect(start, `${method} not found in app.js`).toBeGreaterThan(-1);
const next = /^ {2}(?:async )?[A-Za-z_$][\w$]*\(/m.exec(source.slice(start + 1));
return next ? source.slice(start, start + 1 + next.index) : source.slice(start);
}
/**
* Simulate a live SSE event arriving while a buffer load is in progress.
* Mirrors batchTerminalWrite's queue branch, which stamps each entry with its
* arrival time so a flush can replay only the tail (see the `since` tests).
*/
function pushWhileLoading(app: BufferLoadApp, data: string, at = performance.now()) {
if (app._isLoadingBuffer && app._loadBufferQueue) app._loadBufferQueue.push({ at, data });
}
describe('buffer-load flush (COD-144)', () => {
@@ -153,11 +211,165 @@ describe('buffer-load flush (COD-144)', () => {
// State untouched — still loading, queue intact, nothing replayed.
expect(app._isLoadingBuffer).toBe(true);
expect(app._bufferLoadOwner).toBe('real-owner');
expect(app._loadBufferQueue).toEqual(['queued']);
expect(app._loadBufferQueue).toEqual([{ at: expect.any(Number), data: 'queued' }]);
expect(app.batchTerminalWrite).not.toHaveBeenCalled();
expect(writes).toEqual([]);
});
// ── The tmux-capture tail: `since` ──
//
// A pane capture is a point-in-time frame taken part-way through the fetch, so
// it holds what arrived BEFORE the capture and nothing after. selectSession
// passes the response's arrival time as `since`, which splits the queue at
// exactly that line: pre-capture events are already painted and must stay
// dropped, post-capture events exist nowhere else and must be replayed.
it('flushes only the entries at or after `since`', () => {
const { app, writes } = makeApp();
const owner = app._beginBufferLoad('load-since');
pushWhileLoading(app, 'already-in-the-capture', 100);
pushWhileLoading(app, 'arrived-at-the-headers', 200);
pushWhileLoading(app, 'arrived-after-the-headers', 300);
app._finishBufferLoad(owner, { flushQueued: true, since: 200 });
// The pre-capture event stays dropped; the boundary entry counts as after.
expect(writes).toEqual(['arrived-at-the-headers', 'arrived-after-the-headers']);
});
it('flushQueued without `since` still replays the whole queue', () => {
// The COD-144 path: a brand-new session's first prompt predates the
// response, so cutting the queue would drop the only content it has.
const { app, writes } = makeApp();
const owner = app._beginBufferLoad('load-no-since');
pushWhileLoading(app, 'prompt', 10);
pushWhileLoading(app, 'more', 20);
app._finishBufferLoad(owner, { flushQueued: true });
expect(writes).toEqual(['prompt', 'more']);
});
it('a `since` past every entry flushes nothing', () => {
const { app, writes } = makeApp();
const owner = app._beginBufferLoad('load-since-late');
pushWhileLoading(app, 'old', 10);
app._finishBufferLoad(owner, { flushQueued: true, since: 999 });
expect(writes).toEqual([]);
expect(app.batchTerminalWrite).not.toHaveBeenCalled();
});
// ── Re-entering one load ──
//
// `selectSession` opens the load before its fetch, and `chunkedTerminalWrite`
// opens it again under the SAME owner when it starts writing. A reset on that
// second call would silently throw away everything queued during the fetch,
// which on the capture path is output no buffer holds.
it('re-entering the same load keeps what the queue already holds', () => {
const { app, writes } = makeApp();
const owner = app._beginBufferLoad('load-reenter');
pushWhileLoading(app, 'arrived-during-the-fetch', 100);
// chunkedTerminalWrite re-opens the load it was handed.
app._beginBufferLoad(owner);
pushWhileLoading(app, 'arrived-during-the-write', 200);
app._finishBufferLoad(owner, { flushQueued: true, since: 50 });
expect(writes).toEqual(['arrived-during-the-fetch', 'arrived-during-the-write']);
});
it('a genuinely different load still starts with an empty queue', () => {
const { app, writes } = makeApp();
app._beginBufferLoad('load-first');
pushWhileLoading(app, 'belongs-to-the-abandoned-load', 100);
// A tab switch starts a new load under a new owner. Its events are not ours.
const second = app._beginBufferLoad('load-second');
pushWhileLoading(app, 'belongs-to-this-load', 200);
app._finishBufferLoad(second, { flushQueued: true, since: 0 });
expect(writes).toEqual(['belongs-to-this-load']);
});
// ── The sticky-scroll baseline across a replay ──
//
// `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` before queueing, and
// `flushPendingWrites` scrolls to the bottom off that sample. The replay runs
// inside `chunkedTerminalWrite` before its promise resolves, with the terminal
// freshly reset and rewritten, so the sample is always true. A caller that
// then restores the reader's position would have that restore undone.
it('the replay latches the baseline true, and the viewport restore re-takes it', () => {
const app = makeScrollApp();
const owner = app._beginBufferLoad('load-scroll');
pushWhileLoading(app as unknown as BufferLoadApp, 'output-after-the-capture', 100);
// The load ends with the terminal reset and rewritten, so it reads as bottom.
app.buffer.viewportY = app.buffer.baseY;
app._finishBufferLoad(owner, { flushQueued: true, since: 0 });
expect(app._wasAtBottomBeforeWrite).toBe(true);
// The caller now puts the reader back where they were reading.
app.buffer.viewportY = 40;
app._syncStickyScrollBaseline();
// The next flush must leave them there.
expect(app._wasAtBottomBeforeWrite).toBe(false);
});
it('a restore that lands back at the bottom keeps sticky scroll armed', () => {
const app = makeScrollApp();
const owner = app._beginBufferLoad('load-scroll-bottom');
pushWhileLoading(app as unknown as BufferLoadApp, 'output-after-the-capture', 100);
app.buffer.viewportY = app.buffer.baseY;
app._finishBufferLoad(owner, { flushQueued: true, since: 0 });
app._syncStickyScrollBaseline();
// A reader who was already at the bottom still wants to be carried along.
expect(app._wasAtBottomBeforeWrite).toBe(true);
});
it('both callers that restore a scroll position re-take the baseline', () => {
// The wiring lives in app.js, outside this file's vm harness. Without it the
// two methods below restore the viewport and the next flush undoes it.
const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8');
for (const method of ['_onSessionNeedsRefresh', '_maybeRefetchFullHistory']) {
const body = methodBody(source, method);
const restoreAt = body.lastIndexOf('scrollToLine(');
const syncAt = body.indexOf('this._syncStickyScrollBaseline()');
expect(restoreAt, `${method} no longer restores a scroll position`).toBeGreaterThan(-1);
expect(syncAt, `${method} never re-takes the baseline`).toBeGreaterThan(-1);
expect(syncAt, `${method} re-takes the baseline before its restore`).toBeGreaterThan(restoreAt);
}
});
it('every path that fetches a terminal buffer and writes it asks the shared helper', () => {
// Drift guard. The first version of this fix covered one of the four paths,
// and a later pass found it still covering one of four. Nothing else in the
// gate stops a fifth path, or an inlined `{ flushQueued: true }`, from
// splitting the policy up again; the browser suite that would notice does
// not run in CI.
const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8');
for (const method of [
'selectSession',
'_onSessionNeedsRefresh',
'_onSessionClearTerminal',
'_maybeRefetchFullHistory',
]) {
expect(methodBody(source, method), `${method} decides the flush policy itself`).toContain(
'this._bufferLoadFinishOpts('
);
}
});
it('empty queue + flushQueued is a no-op (no throw, no writes)', () => {
const { app, writes } = makeApp();
const owner = app._beginBufferLoad('load-empty');
+305
View File
@@ -0,0 +1,305 @@
/**
* What a copy actually puts on the clipboard.
*
* xterm returns whole screen rows and trims only the cells that were never
* written to, so a full-screen TUI's padding spaces reach the clipboard. These
* tests drive the SHIPPED transform (`CodemanCopySelection.clean` in
* constants.js), the SHIPPED wiring that decides column mode, and both SHIPPED
* copy paths, because the interesting failures live in the paths rather than in
* the string handling: a padding-only selection must not silently keep the
* user's Ctrl+C, and must not put a bare newline on the clipboard.
*
* A shared LEADING indent is deliberately left alone. The block below pins that
* as a contract rather than an accident, because stripping it was built and
* dropped before merge: see the rule in docs/architecture-invariants.md.
*
* Strategy: constants.js and terminal-ui.js in one vm with a stub CodemanApp,
* the harness shape test/terminal-auto-copy.test.ts uses. No DOM, no xterm.
*/
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it, vi } from 'vitest';
const publicDir = resolve(import.meta.dirname, '../src/web/public');
const read = (name: string) => readFileSync(resolve(publicDir, name), 'utf8');
function loadHarness() {
const CodemanApp = function CodemanApp(this: unknown) {};
const windowRef: Record<string, any> = {};
const context = vm.createContext({
window: windowRef,
document: {
body: { classList: { contains: () => false } },
getElementById: () => null,
querySelector: () => null,
addEventListener: () => {},
},
CodemanApp,
console: { warn: vi.fn(), log: vi.fn(), debug: vi.fn() },
_crashDiag: { log: vi.fn() },
requestAnimationFrame: () => 1,
setTimeout: () => 1,
Blob: function Blob() {},
URL: { createObjectURL: () => 'blob:yield', revokeObjectURL: () => {} },
Worker: function Worker(this: any) {
this.postMessage = () => {};
},
MobileDetection: { isTouchDevice: () => false, getDeviceType: () => 'desktop' },
KeyboardHandler: { keyboardVisible: false },
DEC_SYNC_STRIP_RE: /\x1b\[\?2026[hl]/g,
TERMINAL_CHUNK_SIZE: 32 * 1024,
});
vm.runInContext(read('constants.js'), context, { filename: 'constants.js' });
vm.runInContext(read('terminal-ui.js'), context, { filename: 'terminal-ui.js' });
const app = new (CodemanApp as unknown as new () => Record<string, any>)();
const toasts: { message: string; type: string }[] = [];
app.showToast = (message: string, type: string) => toasts.push({ message, type });
app._copyText = vi.fn(async () => true);
app.loadAppSettingsFromStorage = () => ({ autoCopySelection: true });
const setSelection = (selection: string, { startX = 0, columnMode = false } = {}) => {
app.terminal = {
hasSelection: () => !!selection,
getSelection: vi.fn(() => selection),
getSelectionPosition: () => ({ start: { x: startX, y: 0 }, end: { x: 0, y: 1 } }),
clearSelection: vi.fn(),
focus: vi.fn(),
_core: { _selectionService: { _activeSelectionMode: columnMode ? 3 : 0 } },
};
return app.terminal;
};
return { app, windowRef, toasts, setSelection };
}
const clean = (text: unknown) => loadHarness().windowRef.CodemanCopySelection.clean(text);
describe('CodemanCopySelection.clean — trailing padding', () => {
it('drops the padding a full-screen TUI writes across the rest of each row', () => {
expect(clean('hello \nworld ')).toBe('hello\nworld');
});
it('drops it from a single-row selection too', () => {
expect(clean('hello ')).toBe('hello');
});
it('keeps the line endings xterm chose, including the Windows \\r\\n', () => {
expect(clean('hello \r\nworld \r\n')).toBe('hello\r\nworld\r\n');
});
it('drops trailing tabs as well as trailing spaces', () => {
expect(clean('hello \t \nworld')).toBe('hello\nworld');
});
it('leaves a line that has no padding untouched', () => {
expect(clean('hello\nworld')).toBe('hello\nworld');
});
});
describe('CodemanCopySelection.clean — a shared leading indent is kept', () => {
// Measured over 401,445 three-row windows across 1,010 tracked files, stripping
// the run every row shares fired on 73% of them, and the transform cannot tell
// a TUI margin from content. These are the cases that settled it: each one is
// real output a user copies, and each one loses information if this changes.
it('keeps the indent every selected row shares', () => {
expect(clean(' first line\n second line')).toBe(' first line\n second line');
});
it('keeps a git log body at its four-space indent', () => {
expect(clean(' fix(terminal): trim the padding \n xterm hands back whole rows ')).toBe(
' fix(terminal): trim the padding\n xterm hands back whole rows'
);
});
it('keeps indented Python, where the indent is semantic', () => {
expect(clean(' for item in items:\n if item.ready:')).toBe(
' for item in items:\n if item.ready:'
);
});
it('keeps the leading space on git diff context rows, where it is the marker', () => {
expect(clean(' const x = 1;\n }')).toBe(' const x = 1;\n }');
});
it('is a no-op on shell output, which shares no indent anyway', () => {
expect(clean('$ ls\n indented output\ndone')).toBe('$ ls\n indented output\ndone');
});
it('still drops trailing padding on every one of those rows', () => {
expect(clean(' first \n \n second ')).toBe(' first\n\n second');
});
it('does not eat a \\r on a blank row', () => {
expect(clean(' first\r\n\r\n second\r\n')).toBe(' first\r\n\r\n second\r\n');
});
});
describe('CodemanCopySelection.clean — one row is treated like any other', () => {
it('leaves the indent on a single-row selection', () => {
expect(clean(' hello world ')).toBe(' hello world');
});
it('leaves it on a single row followed by a blank row', () => {
expect(clean(' hello world\n ')).toBe(' hello world\n');
});
it('leaves it when a second row carries content, exactly as for one row', () => {
expect(clean(' hello\n world')).toBe(' hello\n world');
});
});
describe('CodemanCopySelection.clean — nothing to clean', () => {
it('returns an empty string for an empty selection', () => {
expect(clean('')).toBe('');
});
it('returns an empty string rather than throwing on a non-string', () => {
expect(clean(undefined)).toBe('');
expect(clean(null)).toBe('');
});
it('reduces an all-padding selection to its line breaks alone', () => {
// The copy paths reject this with trim(); the transform itself only removes
// whitespace, so the row structure survives here by design.
expect(clean(' \n \n ')).toBe('\n\n');
});
});
describe('cleanedTerminalSelection — wiring', () => {
it('trims each row and leaves the shared indent alone', () => {
const { app, setSelection } = loadHarness();
setSelection(' first \n second ');
expect(app.cleanedTerminalSelection()).toBe(' first\n second');
});
it('does not read the selection position at all', () => {
// The mid-row flag is gone. It read getSelectionPosition().start, which is
// xterm's mousedown ANCHOR and is never normalised, so an upward drag read
// it off the bottom row of the selection.
const { app, setSelection } = loadHarness();
const terminal = setSelection(' first\n second');
terminal.getSelectionPosition = vi.fn(() => ({ start: { x: 6, y: 0 }, end: { x: 0, y: 1 } }));
expect(app.cleanedTerminalSelection()).toBe(' first\n second');
expect(terminal.getSelectionPosition).not.toHaveBeenCalled();
});
it('uses the text it is given without reading the selection again', () => {
const { app, setSelection } = loadHarness();
// The contract is that `text` IS the live selection, so the stub agrees with
// it; the assertion that carries weight is that getSelection went unread.
const terminal = setSelection(' given text \n second row ');
expect(app.cleanedTerminalSelection(' given text \n second row ')).toBe(' given text\n second row');
expect(terminal.getSelection).not.toHaveBeenCalled();
});
it('returns an empty string when there is no selection at all', () => {
const { app, setSelection } = loadHarness();
setSelection('');
expect(app.cleanedTerminalSelection()).toBe('');
});
it('leaves a column selection completely untouched', () => {
// Alt+drag makes a rectangle, and its rows lining up is the whole point:
// both halves of the clean would destroy that alignment.
const { app, setSelection } = loadHarness();
const rect = ' alpha \n beta \n gamma ';
setSelection(rect, { startX: 40, columnMode: true });
expect(app.cleanedTerminalSelection()).toBe(rect);
});
});
describe('the xterm internals the column check depends on', () => {
// The column check reads a private field and compares it to a literal, because
// xterm publishes the selection mode nowhere. A rename or a renumber would make
// every rectangular selection get cleaned with both rules and lose the column
// alignment the rule exists to protect, and the fallback is silent by design.
// So the assumption is pinned against the real library rather than only against
// a stub that repeats it. lib/xterm.js is the esbuild input for the shipped
// vendor bundle, so it is the file that decides what runs in the browser.
const xtermLib = readFileSync(resolve(import.meta.dirname, '../node_modules/@xterm/xterm/lib/xterm.js'), 'utf8');
it('still branches on _activeSelectionMode === 3 for a column selection', () => {
expect(xtermLib).toContain('3===this._activeSelectionMode');
});
it('still reaches that field through _selectionService', () => {
expect(xtermLib).toContain('_selectionService');
});
});
describe('copyTerminalSelection — what reaches the clipboard', () => {
it('copies the cleaned text, never the padded rows', () => {
const { app, setSelection } = loadHarness();
setSelection(' first line \n second line ');
return app.copyTerminalSelection().then((ok: boolean) => {
expect(ok).toBe(true);
expect(app._copyText).toHaveBeenCalledWith(' first line\n second line');
});
});
it('cleans a realistic TUI block, padding only', () => {
const { app, setSelection } = loadHarness();
const pane = [' That last point is the important one. ', ' Claude Code writes each paragraph. '].join(
'\n'
);
setSelection(pane);
return app.copyTerminalSelection().then(() => {
expect(app._copyText).toHaveBeenCalledWith(
' That last point is the important one.\n Claude Code writes each paragraph.'
);
});
});
it('clears a padding-only selection rather than leaving a dead highlight', () => {
// The clear is feedback, not protection: the Ctrl+C gate tests the CLEANED
// selection, so a padding-only one falls through to the PTY either way.
const { app, toasts, setSelection } = loadHarness();
const terminal = setSelection(' ');
return app.copyTerminalSelection().then((ok: boolean) => {
expect(ok).toBe(false);
expect(app._copyText).not.toHaveBeenCalled();
expect(terminal.clearSelection).toHaveBeenCalledTimes(1);
expect(toasts).toEqual([{ message: 'Nothing to copy', type: 'warning' }]);
});
});
it('rejects a multi-row padding selection instead of copying bare newlines', () => {
const { app, setSelection } = loadHarness();
const terminal = setSelection(' \n \n ');
return app.copyTerminalSelection().then((ok: boolean) => {
expect(ok).toBe(false);
expect(app._copyText).not.toHaveBeenCalled();
expect(terminal.clearSelection).toHaveBeenCalledTimes(1);
});
});
});
describe('_flushAutoCopySelection — cleaned text is what Auto Copy handles', () => {
it('copies the cleaned text and remembers it for the dedupe', () => {
const { app, setSelection } = loadHarness();
setSelection(' first line \n second line ');
app._autoCopyPending = true;
return app._flushAutoCopySelection().then(() => {
expect(app._copyText).toHaveBeenCalledWith(' first line\n second line');
expect(app._autoCopyLastText).toBe(' first line\n second line');
});
});
it('reads nothing at all while the toggle is off', () => {
// Auto Copy defaults to OFF, and a selection can run to the scrollback
// ceiling, so the flush must not read or clean before it checks.
const { app, setSelection } = loadHarness();
const terminal = setSelection(' first line \n second line ');
terminal.getSelection = vi.fn(() => ' first line ');
app.loadAppSettingsFromStorage = () => ({ autoCopySelection: false });
app._autoCopyPending = true;
return app._flushAutoCopySelection().then(() => {
expect(terminal.getSelection).not.toHaveBeenCalled();
expect(app._copyText).not.toHaveBeenCalled();
});
});
});
+45
View File
@@ -355,6 +355,51 @@ describe('terminal flush budget', () => {
expect(app._bufferLoadOwner).toBe(null);
});
// ── Which payloads end their load by replaying the queue ──
//
// A pane capture is current only up to capture time, so the tail that arrived
// after the response exists nowhere else and has to be replayed. The server's
// accumulated byte history is current up to the response, so replaying on top
// of it would duplicate output. `_bufferLoadFinishOpts` is the one place that
// decides this, for all four paths that fetch a terminal buffer and write it.
it('replays the tail for a visible-pane capture', () => {
const { CodemanApp } = loadAppHarness();
const app = Object.create(CodemanApp.prototype) as any;
expect(app._bufferLoadFinishOpts({ source: 'mux-visible' }, 1234)).toEqual({
flushQueued: true,
since: 1234,
});
});
it('replays the tail for a full-history capture', () => {
const { CodemanApp } = loadAppHarness();
const app = Object.create(CodemanApp.prototype) as any;
expect(app._bufferLoadFinishOpts({ source: 'mux-full-history' }, 1234)).toEqual({
flushQueued: true,
since: 1234,
});
});
it('discards the queue for the accumulated byte history', () => {
const { CodemanApp } = loadAppHarness();
const app = Object.create(CodemanApp.prototype) as any;
expect(app._bufferLoadFinishOpts({ source: 'history' }, 1234).flushQueued).toBe(false);
});
it('discards the queue for a payload that names no source', () => {
// Fails toward the safe answer: a duplicated Ink redraw corrupts the screen,
// while a dropped tail is repaired by the CLI's next full repaint.
const { CodemanApp } = loadAppHarness();
const app = Object.create(CodemanApp.prototype) as any;
expect(app._bufferLoadFinishOpts({}, 1234).flushQueued).toBe(false);
expect(app._bufferLoadFinishOpts(undefined, 1234).flushQueued).toBe(false);
});
it('does not snap back to bottom during Codex Working redraws right after the user scrolls up', () => {
const { app } = loadTerminalUiHarness('codex');
const scrollToBottom = vi.fn();
+51 -1
View File
@@ -11,7 +11,7 @@
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import { describe, expect, it } from 'vitest';
import { formatCursorRestore, hasVisibleContent } from '../src/tmux-manager.js';
import { formatCursorRestore, formatPaneSnapshot, hasVisibleContent } from '../src/tmux-manager.js';
describe('tmux full-history pane capture (COD-47)', () => {
const source = readFileSync(resolve(import.meta.dirname, '../src/tmux-manager.ts'), 'utf8');
@@ -121,3 +121,53 @@ describe('hasVisibleContent', () => {
expect(hasVisibleContent('\x1b[m \x1b[0m\n\x1b[m x \x1b[0m')).toBe(true);
});
});
describe('the geometry a capture reports back', () => {
const source = readFileSync(resolve(import.meta.dirname, '../src/tmux-manager.ts'), 'utf8');
const methodStart = source.indexOf('capturePaneBuffer(muxName: string');
const methodEnd = source.indexOf('captureActivePaneBuffer(muxName: string', methodStart);
const methodBody = source.slice(methodStart, methodEnd);
it('writes the pane size onto the caller options before either replay path returns', () => {
// IS_TEST_MODE no-ops execSync, so assert from source (same approach as the
// capture-flag tests above). The write must precede the fullHistory branch:
// both paths return from inside it, and a caller that got no geometry
// cannot tell a mismatched frame from a matching one.
const write = methodBody.indexOf('opts.capturedGeometry = { cols: geometry.cols, rows: geometry.rows }');
// Anchor on the REPLAY branch, not the earlier `if (fullHistory)` that only
// sizes the exec buffer.
const replayBranch = methodBody.indexOf('if (!geometry) return normalizeScrollbackEol(');
const visibleReturn = methodBody.indexOf('if (geometry) return formatPaneSnapshot(');
expect(write).toBeGreaterThan(-1);
expect(replayBranch).toBeGreaterThan(-1);
expect(visibleReturn).toBeGreaterThan(-1);
expect(write).toBeLessThan(replayBranch);
expect(write).toBeLessThan(visibleReturn);
});
it('reports nothing when the cursor query gave no geometry', () => {
// `queryPaneCursor` returns null on a failed or nonsensical query, and the
// snapshot repaint is skipped in that case. Reporting a size anyway would
// describe a frame that was never positioned.
expect(methodBody).toContain('if (opts && geometry)');
});
});
describe('why a capture has to report its height', () => {
it('a snapshot addresses rows the receiving terminal may not have', () => {
// formatPaneSnapshot positions every row absolutely. A terminal shorter
// than the pane clamps each address past its own height onto its last
// line, so the overflow rows overwrite one another and the rows underneath
// are lost. Nothing in the escape sequence tells the client this happened —
// hence captureRows on the response.
const lines = Array.from({ length: 50 }, (_, i) => `row-${i + 1}`);
// cursorX 5 keeps the trailing cursor-restore move (`\x1b[50;6H`) out of the
// `;1H` row-paint match below, so the count is row paints alone.
const snapshot = formatPaneSnapshot(lines, { cols: 100, rows: 50, cursorX: 5, cursorY: 49 });
const addressed = [...snapshot.matchAll(/\x1b\[(\d+);1H/g)].map((m) => Number(m[1]));
expect(Math.max(...addressed)).toBe(50);
// A 30-row terminal cannot honour 20 of those addresses.
expect(addressed.filter((row) => row > 30)).toHaveLength(20);
});
});