Compare commits

...
Author SHA1 Message Date
Codeman maintainer cbb1435a46 refactor(toolbar): one instance stepper, not two
The desktop toolbar carried two identical "minus 1 plus" instance steppers side
by side, one after Run and one after Run Shell. The second (#shellCount) is
gone for a cleaner strip.

Run Shell keeps the capability: both launch paths now read the remaining
#tabCount control through _toolbarInstanceCount(), which also makes an absent
stepper read as 1 instead of throwing. That matters because the group is
display:none on phones and tablets, and because the Run dropdown's
Terminal / Shell entry routes through runShell() too, where the visible counter
was previously ignored.

Desktop only: both steppers were already hidden under 1024px.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-15 00:50:09 +02:00
Codeman maintainer 8e4606c57b feat(mobile): search the case picker
The phone case sheet listed every case with no way to narrow it, while the
desktop toolbar combobox has filtered for a while. The sheet now carries a
search field that runs the same matcher (filterCasePickerOptions), so both
pickers answer a query identically: every term has to appear in the option's
searchText, which already carries the name, the rendered label, the path and
the remote/docker fields.

Details worth keeping:

- The filter resets on every open. The sheet is a one-shot picker, and a
  leftover query would present a truncated list as the whole one.
- No autofocus. Focusing raises the keyboard over a sheet anchored to the
  bottom of the screen, so the user asks for it.
- The sheet is a third position:fixed bottom-anchored surface, so it joins the
  toolbar and the accessory bar in KeyboardHandler's keyboard lift. iOS does
  not shrink the layout viewport, so an unlifted sheet would sit behind the
  keyboard with its own search box out of sight. resetLayout() clears the
  offset unscoped, or a sheet closed while the keyboard was up would slide in
  already displaced next time.
- Enter takes a single remaining match and otherwise just dismisses the
  keyboard; Escape drops the filter before it closes the sheet.
- No match renders an empty state rather than a blank sheet.
- The clear button needs an explicit [hidden] rule: the UA's display:none is
  specificity (0,0,0) and loses to the button's own display:flex.
- The input drops the global input:focus-visible ring, which inside an already
  bordered row drew a second border a few pixels in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-15 00:49:56 +02:00
Codeman maintainer 88e3faa456 chore: version packages 2026-09-15 00:07:19 +02:00
Codeman maintainer 70fc6b32d5 docs: record the dup/last input ACK, Shift+drag and right-click copy, and multi-case adopted containers
Three behaviours landed from #375 without their doc entries: the
duplicate input ACK now carries `dup:true` and the server's watermark
(`docs/reliable-input-delivery.md` still described a bare ACK), Shift+drag
and right-click copy in the terminal (the shortcut list did not know
them), and one adopted container backing several cases at different
in-container directories (the Docker cases paragraph still implied one
case per container for adopted containers too).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:56:19 +02:00
Codeman maintainer 9591b973cf fix(docker): carry the owned flag on the wire the way master already does
The cherry-picked "copy an existing case" commit declared a second
`CaseInfo.docker.owned` and emitted `owned: true|false` on every docker
case, while master had meanwhile shipped the same field from the
adopted-container work with a narrower wire shape: `owned` is present
only when false, absent means owned. Two declarations failed typecheck,
and two emit styles on one response would have made the picker's answer
depend on which read path filled it.

Keep master's shape at both response sites (the case list and the
single-case lookup, which lacked the field entirely), fold the picker's
reason for the field into the existing doc comment, and repoint the test
that pinned "set on exactly two sites" at the surviving form, adding a
negative pin so the duplicate style cannot come back.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:56:19 +02:00
d fei 025f061383 fix(docker): pre-fill the copied case instead of blanking two fields
The previous version cleared the case name and the in-container directory on the
grounds that they must differ. That left a form with three fields mysteriously
filled and two empty, and turned the most common operation — changing
/srv/app/api to /srv/app/web — into retyping a long path.

Both are now pre-filled, with focus on the in-container directory and the caret
at the end, since the tail is what changes. What stops an unmodified submit is no
longer an empty field but a guard: the values applied are recorded, compared at
submit time, and if nothing changed the reason is stated next to the field and
focus moves to it, without sending a request that is certain to be refused.

The server refuses these anyway (a duplicate case name, a twin case on the same
container and directory) and its errors are clear; but making a round trip to be
told "you forgot to edit the field you are looking at" is worse than saying so on
the spot. The guard only applies when a source case was actually selected, so
filling the adopt form from scratch is unaffected.

⚠️ The status text is written into dockerLinkStatus. My first version referenced
an id that does not exist (dockerAdoptStatus), which made the explanation vanish
silently and left only a toast. The test now extracts that id from the code and
looks it up in index.html, pinning that it must really exist.

(cherry picked from commit ba21ae11f4)
2026-09-14 23:56:19 +02:00
d fei 7a5543da09 feat(docker): add "copy an existing case" to the adopt panel
The backend already lets one adopted container back several cases pointing at
different in-container directories, but using it meant retyping the container
name, host and workspace one by one — exactly the friction that leaves a
capability unused. Picking an existing case from a dropdown now carries those
three over, leaving only the two fields that must differ: the case name and the
in-container directory.

Clearing those two is the point of the feature, not a convenience: keeping the
old name is refused by the server as "case already exists", and keeping the old
directory is refused as "a twin case on the same container and directory". Both
errors are clear, but a form pre-filled with values that are guaranteed to be
rejected is a trap. Focus lands on the in-container directory — the thing the
user came here to change.

⚠️ Only adopted containers are listed (docker.owned === false). A Codeman-built
container's lifecycle belongs to its one case — a second case would be torn out
by that case's recreate or delete — so the server refuses it anyway, and listing
it here would only manufacture a baffling error. `owned` may be absent and absent
means owned, so the test is `!== false`, not truthiness.

CaseInfo.docker gains containerWorkdir and owned for this: the former is the
"which directory does this case use" half of the picker, without which the user
cannot tell what to change it to; the latter backs the filter above. ⚠️ Both
places that build a docker CaseInfo (the list endpoint and the single-case query)
must set them — filling in only one makes the picker work or not depending on
which read path was taken, and a test pins "exactly two".

(cherry picked from commit f1ed3a58e1)
2026-09-14 23:56:19 +02:00
d fei cbb7f635ff feat(docker): let one adopted container back several cases in different dirs
Once a container is adopted, it could not be adopted a second time. But a
container usually holds more than one project directory, and opening a case for
another one had no path forward except starting a second container — precisely
what adoption exists to avoid.

The original reason was in a comment: two cases sharing an adopted container
would make one case's teardown race the other's launch on the same tmux server.
That reason does not hold. The in-container tmux session name is
dockerTmuxSessionName(sessionId), i.e. codeman-dkr-<id8>, keyed by SESSION and
not by case, and buildDockerKillCommand tears down exactly that name, so killing
A never touches B — hosting multiple sessions is what a tmux server is for.

The other three routes into an adopted container's lifecycle do not pass through
here either, confirmed one by one: the stop and remove builders throw outright;
recreate refuses `owned === false` before it even resolves the container name;
and orphan reaping filters on `label=codeman.managed=1`, which a user-built
container does not carry — a structural exclusion.

That leaves exactly three cases worth refusing, none of them tmux-related, split
into the pure, unit-tested classifyAdoptContainerConflict:
- owned-case   the container belongs to a Codeman-created case, whose lifecycle
               Codeman manages: one recreate or delete there would pull the
               container out from under the adopting case.
               ⚠️ `owned` may be absent and absent means owned (cases predate
               the field), so the test is `!== false`, not truthiness.
- other-owner  already adopted by a different user. Adoption hands out a shell
               inside someone else's container.
- duplicate    same container, same directory. The second case would behave
               identically to the first, so name the existing one rather than
               silently minting a twin. A different in-container directory is
               the case this change exists to support and passes.

(cherry picked from commit 1cb6bde891)
2026-09-14 23:56:19 +02:00
d fei e5684d0bba fix(ui): don't create a compositing layer for a hidden full-screen overlay
`backdrop-filter` promotes an element to its own compositing layer. A
position:fixed full-screen layer that is created and then hidden was measured to
leave a stale hit-test region behind in Chrome: the page renders perfectly, but
pointer events across the viewport go nowhere.

The report came from a long-lived tab connected to a remote server, where a
connection blip shows and then hides #offlineOverlay. The symptoms were a
terminal that would not scroll and, at the same time, an unrelated
click-to-expand that also stopped responding, while a freshly opened tab was
fine; a read-only console command (getComputedStyle + elementFromPoint, both of
which force a hit-test recomputation) then cured it. Two unrelated features
dying together and one read-only command fixing both points at hit-testing
itself rather than at either feature.

So the `backdrop-filter` moves onto the actually-visible selector and the layer
is never created while hidden. Only the two persistent overlays change:
offline-overlay (toggled with [hidden]) and file-preview-overlay (toggled with
.visible). path-picker and path-preview are created and removed by JS, leave
nothing behind, and are untouched.

⚠️ This is an evidence-based inference, not a fix verified by reproduction:
reproducing it needs a long-lived page that has been through a connection blip,
which I could not manufacture in a controlled environment. The guard test pins
both halves — no such property while hidden, and a real blur while shown — so a
later cleanup cannot quietly delete the effect.

(cherry picked from commit 08442dfee1)
2026-09-14 23:56:19 +02:00
d fei c7cc8e28d5 fix(sse): stop reloading the whole terminal when a reconnect lands on the same session
handleInit() did not distinguish a first load from an SSE reconnect: it always
cleared the terminal caches in _resetAllAppState() and re-ran selectSession() for
the session that was already on screen. Every reconnect therefore refetched up to
1 MiB of buffer and reset+rewrote xterm. On a link that drops a connection about
once a minute (measured at ~57s intervals against a healthy server) that reads as
the page refreshing itself and throwing away your reading position.

A reconnect that lands back on the still-open session now keeps the terminal
caches and activeSessionId and resyncs through _onSessionNeedsRefresh(). That
path still reloads the buffer, so output produced during the outage is not lost,
but it preserves distance-from-bottom — the same rule #259 established for a
refresh the server triggered rather than the user. The WS is reconnected
explicitly when it is not already on that session, since skipping selectSession()
skips its _connectWs() call.

First load (gen === 1) takes exactly the path it took before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rv24Pk4qzrsDYdVyDyJQmT
(cherry picked from commit 435569c76e)
2026-09-14 23:56:19 +02:00
d fei 01da577053 fix(input): recover when the seq counter falls behind the server watermark
Browser input is delivered exactly once by (clientId, seq). The server records a
watermark per clientId and discards anything not above it as a duplicate — but
acknowledged it with an ACK indistinguishable from "applied". The client then
dropped the record from its queue, the UI looked perfectly normal, and the
terminal received nothing at all.

The counter is persisted to localStorage through a debounced write. Kill the page
between "sent" and "persisted" and the restored counter is below the server's
watermark, after which every keystroke lands under it, is discarded, and is
ACKed. Reloading does not help: the clientId is restored from localStorage
alongside that stale counter. Measured on a real session — typing into the same
session from a fresh browser (new clientId, no watermark on the server) worked
perfectly, which is what localised the fault to client state.

Three changes:
- on rejection the server replies {"t":"ia",seq,"dup":true,"last":<watermark>}.
  It still ACKs, so the client can drop the record from its queue, but it now
  says the input was not applied and supplies the number needed to climb out.
- on `dup` the client lifts its counter above the watermark and re-queues.
  ⚠️ Only records whose FIRST delivery is being retried are re-sent: a retry
  judged duplicate means the mechanism is working (the original did arrive), and
  re-sending would type the same text twice.
- the counter is now persisted synchronously. The queue payload can stay
  debounced, but the counter is the thing that has to survive a crash, and
  leaving it on the lossiest path cancels the only guarantee there is.

⚠️ Reading the watermark is defensive: the session arrives through a structured
port, and a port missing that method must not take the whole input path down —
a throw inside the handler means the ACK is never sent and the record is stuck in
the client queue forever, which is worse than the ambiguity being fixed. A mock
port's test timeout is what exposed this.

(cherry picked from commit 05bb7081cc)
2026-09-14 23:56:19 +02:00
d fei 631386d3f7 fix(cjk): forward Ctrl/Alt-modified navigation keys to the CLI
claude advertises "Jump to bottom (ctrl+End)", so that chord has to actually
reach it. But PASSTHROUGH_KEYS carried only the bare forms (End -> \x1b[F) and
CTRL_KEYS held just six letters (c/d/l/z/a/e), which cannot express End. Ctrl+End
therefore failed in both directions:

- with an empty composer it went out as a bare \x1b[F, the modifier silently
  dropped, so the CLI received a plain End;
- with text in the composer the forwarding branch requires empty, so nothing was
  forwarded and the browser default applied — the caret jumped to the end of the
  draft, which is the "the shortcut now edits my input box" the user saw.

Encode them as CSI 1;<mod><final> instead, and forward Ctrl/Alt-modified
navigation keys whether or not the composer is empty: they are commands for the
CLI, and the composer has no editing semantics for them worth preserving (bare
Home/End still use the old table and edit locally).

⚠️ Bare Shift is deliberately excluded: Shift+arrow selects text in the composer,
a real editing gesture that must stay local. Shift held together with Ctrl/Alt is
still encoded into the modifier mask.

(cherry picked from commit 3fbaadadfb)
2026-09-14 23:56:19 +02:00
d fei b3a6ba2eb6 feat(terminal): make Shift+drag select, and right-click copy the selection
In a native terminal running a TUI with mouse tracking on (claude, codex), Shift
is the "let me select text" modifier: it bypasses the application's mouse
reporting so the emulator selects locally. Users bring that habit here, where it
did nothing — measured, `hasSelection` was already false during a Shift+drag and
no clearSelection call ran at all, because there was never a selection to clear.

The mismatch is that the two Shifts mean different things. xterm reads Shift as
"force selection", but that path is only taken when the application really has
mouse tracking on. The server strips the mouse DECSETs for claude/codex/gemini
(isAltScreenStripMode), so xterm's mouseTrackingMode is permanently `none`, that
branch is unreachable, and Shift instead lands in _onIncrementalClick — which
EXTENDS an existing selection. Extension is a no-op while selectionStart is
empty, so the drag had no anchor.

So plant the anchor xterm is missing. The listener sits on the capture phase of
the `.xterm` root, an ancestor of the `.xterm-screen` that SelectionService binds
to, and therefore runs before xterm's own mousedown; xterm then extends from our
anchor and the drag behaves like any other. Length is 0 so a Shift+click without
a drag does not select a stray character. An existing selection is left alone —
that is a genuine extend gesture, and xterm handles it correctly.

Right-click copies the selection (the mintty/PuTTY convention), completing the
gesture: until now there was nowhere for a finished selection to go. With no
selection the native menu is not hijacked — taking it away while offering
nothing in return is a pure loss.

(cherry picked from commit 7ab5015737)
2026-09-14 23:56:19 +02:00
Codeman maintainer 897a63183f chore(typecheck): include the local-LLM harness smoke script
scripts/test-local-llm-harnesses.ts (#393) sits outside tsconfig.json's
include, so nothing type-checked it. config/tsconfig.scripts.json pulls
it in; npm run typecheck now runs both projects, the way the pr-bot
config used to be chained before the bot moved out of the repo.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:47:32 +02:00
Codeman maintainer 942bf37e48 fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a
custom OpenAI-compatible endpoint by injecting env vars or a config file
and restarting the CLI in place. Review of the apply path found four
things, two of them destructive. This lands all four plus the smaller
items from the same review.

1. Clearing a selection did not clear it. The injected vars reach the CLI
   via `tmux setenv`, which persists at the tmux-session level and is
   inherited by `respawn-pane` (measured: `setenv FOO bar` survived two
   successive `respawn-pane -k`), so deleting the keys from the session's
   envOverrides relaunched the CLI still pointed at the old endpoint, and
   for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just
   been deleted. `Session.setCustomModel()` now reports the removed keys,
   queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys`
   carries them into `applyEnvOverrides()`, which `setenv -u`s them before
   re-applying the live overrides, on the same path that already unsets
   the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket
   that `setenv -u HOME` hands the next respawn the global HOME back.

2. Applying a model to a local claude session killed the pane. The
   relaunch was `claude --session-id <id>` and Claude refuses an id that
   already has a transcript, and unlike the dead-pane respawn this one
   kills a working pane first. `restartCli()` now pins the live
   conversation id as the resume id for that respawn when the CLI's launch
   declares a `fallback` chain, which renders the same
   `--resume <id> || --session-id <id>` shape the docker and remote pane
   commands use. Gated on the registry shape, not the CLI id: an entry
   whose resume id is minted by the CLI itself never declares that chain.

3. pi, omp and grok wrote their config file and then launched without the
   `--model` that selects it, so the file was ignored. The registry entry
   now declares `customModelInjection.launchModel` (`custom/{modelId}` for
   pi and omp, grok's `[model.codeman-custom]` block name), the builder
   renders it, and `_withCustomModelLaunchModel()` applies it onto the
   respawn options through `legacyConfigField`, leaving the stored
   <Mode>Config untouched so a clear falls back to the user's own model.
   A model id the CLI's `model` token pattern cannot carry is refused
   with a 400 rather than silently dropped by the argv engine.

4. Remote (SSH) and Docker sessions reported `restarted: true` and changed
   nothing: their `restartCli()` reattaches the durable tmux rather than
   relaunching the agent, and the env lands on the local pane. Both are
   refused with a 400 until those paths are plumbed.

Smaller items from the same review:

- The selection survives a Codeman restart as the disk-only `__customModel`
  bookkeeping (endpoint, model, injected key NAMES, config dir, launch
  model; never the values, which carry the API key). Recovery re-derives
  the values from the endpoint store through the same apply path the route
  uses and keeps the bookkeeping even when the endpoint is gone, so a
  later clear still has keys to unset.
- Discovery goes through `webviewFetch()`, so the RESOLVED address is
  judged by the same egress guard the web-tab proxy uses, and `baseUrl`
  reuses `webviewUrlSchema` (http(s) only, no embedded credentials,
  link-local and cloud-metadata addresses refused). undici's `fetch failed`
  wrapper is unwrapped so the user sees the ECONNREFUSED underneath.
- `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session
  config dir 0700/0600 (pi and omp embed the key literally), and that dir
  is removed with the session.
- `PR.md` is gone from the repo root and the design doc moved to
  `docs/custom-model-endpoints-plan.md` with the LAN address and the
  personal name scrubbed; every reference follows. The guide's `authStyle`
  text matches the shipped schema (`bearer | api-key`, default `bearer`)
  and says that `customModelEndpointsEnabled` is read by nothing until
  the picker lands.
- `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts`
  (four real type errors fixed). It is not yet wired into `npm run typecheck`
  because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json`
  there is the one-line follow-up.

Tests: `test/session-custom-model-restart.test.ts` drives a real Session and
fails on the unfixed code for items 1 to 3; the route suite covers item 4
and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets
run before the overrides and that a shell-metachar key never reaches tmux.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:46:28 +02:00
Codeman maintainer 1e42cb4e2d Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses)
2026-09-14 23:46:27 +02:00
Codeman maintainer e49c48145b fix(files): fail closed on remote symlinks, guard PUT for remote cases, bound ssh fan-out
Follow-up to #421 (remote-case file reads over ssh), addressing the review.

Symlink escape on a host without `readlink -f` (blocker). The probe's
portable fallback canonicalized only the directory chain and returned the
final component unresolved, so on macOS < 12.3 `ws/notes.txt -> ~/.ssh/id_rsa`
came back as `.../ws/notes.txt` (with the target's size), passed every
containment and blocklist check that runs on `realPath`, and `cat` followed
the link. The fallback now walks the directory chain with `cd -P`/`pwd -P`
and follows the LAST component with plain `readlink` for a bounded number of
hops, and anything it cannot fully resolve (a loop, a readlink failure, the
hop cap) is reported with an `x` marker that parses as null, i.e. 404. It
never returns the unresolved string. Measured on a real /bin/sh with
`readlink -f` shadowed: the pre-fix script reports `/ws/notes.txt`, the fixed
one `/secret/id_rsa`; both branches (native and fallback) now agree.

`PUT /api/sessions/:id/file-content` never had the remote guard the PR
described. It sits ahead of `validateSessionFilePath`, which resolves against
the LOCAL filesystem, because with a same-named directory on the Codeman host
(an sshfs mount of the remote tree, the documented stop-gap) the write landed
on the local twin while the viewer believed it edited the remote file.

ssh fan-out is bounded. `src/remote-ssh-limiter.ts` is a
document-conversion-limiter-shaped semaphore (default 4, env
`CODEMAN_MAX_REMOTE_FILE_SSH`) around every probe and buffered read; the
attachment-history list resolves its whole history in ONE batched probe
(`probeRemoteAttachmentHistory`, threaded into
`registerExternalAttachment({remoteProbes})` so the guards run unchanged)
instead of one handshake per entry; and probes chunk at 40 paths because the
whole script is one argv string. Terminal output in a remote session is
written on the remote host, so a prompt-injected agent printing hundreds of
`codeman://attach` links forked one ssh per link, each holding a 20 s
timeout, and a 100-entry history re-listed on every attachment:detected
tripped OpenSSH's default MaxStartups. Streams are deliberately not counted
(one per browser request, held for a whole playback, and gated behind a
counted probe anyway).

Smaller items from the same review: probe records are NUL-terminated and
index-keyed after a leading NUL (a newline in a filename can no longer shift
the alignment, and the banner is fenced off without last-N-lines guessing);
size comes from `stat -c %s || stat -f %z`; the three IO functions refuse
under VITEST instead of opening a connection; an unreachable host now reads
as unknown (missing: false) for detected AND external history entries, where
external used to fold its 502 into missing; a client that aborted during the
guard probe has its body's ssh child reaped (`reply.raw.destroyed` is checked
before the close listener is attached); `describeExecError` never returns
Node's `Command failed: <ssh line>` message, which carried the identity path
and the probe script into a 502 body; and the docs note that
`isSensitivePath`'s three home-anchored entries resolve against the Codeman
host's home, not the remote one.

Tests: the probe script runs on a real /bin/sh with a `readlink` shim that
rejects `-f` (the escape, a relative chain through a symlinked directory, a
loop, a newline filename, banner chatter that itself looks like a record),
the limiter's cap and FIFO order, and route tests for the PUT guard (local
twin untouched, no connection), the single batched history probe, the
unreachable-host alignment and the aborted-client reap. All four route tests
fail against the pre-fix file-routes.ts.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:42:06 +02:00
Codeman maintainer 792a251e35 Merge pull request #421 from Randalix/fix/remote-file-access
fix(files): read remote-case previews, downloads and attachments over ssh
2026-09-14 23:42:06 +02:00
Codeman maintainer 6dc27ae727 docs(webview): record the lost-frame page as the third unauthenticated 200, and the inline-style limit
The lost-frame recovery page is answered ahead of the credential checks in both
auth hooks, which makes it the third unauthenticated 200 beside the two hook
routes, and the only one decided by request headers alone. CLAUDE.md's security
table listed exactly two, and docs/web-tabs.md is not where anyone auditing that
looks, so it now has a row in the table and a fourth property in
docs/security-architecture.md section 10b, including the `/` carve-out and its
credential-free condition. Both state the property that comes with it: a
non-browser client can set those headers, so an unauthenticated caller can tell a
registered route (401) from a non-route (200) and enumerate the route table,
accepted because the routes are public in docs/api-reference.md.

docs/web-tabs.md gains the landing-page case in layer 6 and a Known limits entry:
masking trades away the Referer safety net, only HTML is rewritten server-side,
and a root-absolute url() inside an inline <style> block has the masked document
as its Referer, so it 404s where the Referer fallback used to rescue it. External
stylesheets are unaffected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:38:54 +02:00
Codeman maintainer 1306f731cf fix(webview): recover a proxied dashboard that reloads on its landing page
The runtime shim masks `/webview/<cap>/` off a proxied page's URL so its router
boots on the path it expects, and the landing page masks to exactly `/`. A
`location.reload()` there (a Vite dev server on a config change or a failed HMR
update, the likeliest case in the feature's own motivating scenario) therefore
asks for Codeman's root as an iframe navigation. `serveLostWebviewFrame()`
returned early for `/`, so on a passwordless install the frame received Codeman's
own app shell and rendered it inside the web tab, and with a password it got a
401 in the frame. Either way no `codeman:webview-lost` message was posted, and
because the document loaded fine the load handler cleared the failed-frame panel,
so the Reload / Open in new tab affordances never appeared. Before masking the
frame's URL was the prefixed one, so a reload worked; this was a regression.

`/` is the one lost-frame path a registered route also serves, so the route
table cannot tell that reload from a real navigation. Credentials can: nothing
in Codeman frames its own root, and a sandboxed frame is opaque-origin with no
cookie and no Authorization header. `carriesAuthCredentials()` (pure, in
webview-proxy.ts) makes that test, and `/` is now admitted by the auth hook only
when it fails; a framed `/` that does carry credentials still gets the shell.
Without a password no auth hook runs at all, so the index route applies the
same test itself (`isLostWebviewRootFrame`) before rendering the shell, and the
three places that emitted the recovery page share `sendLostWebviewFramePage()`.

Tests: the password form in webview-auth-exemption (recovery page for a
credential-free framed `/`, shell with valid Basic auth, 401 with a stale cookie
or a top-level navigation), the passwordless form against a real WebServer in
webview-lost-root-frame (port 3198), and the credential predicate in
webview-proxy. All three fail without the fix. Verified against a live isolated
instance as well: a framed `GET /` with no credentials answers the 470-byte
recovery page, a top-level `GET /` and a framed one carrying a cookie answer the
shell.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:38:54 +02:00
Codeman maintainer d9364f52e1 fix(webview): refuse a backslash or tab-led recovery path, which the URL parser reads as an origin
The lost-frame handler in webview-tabs.js remounts a web-tab frame at the path
the frame reports it lost. It promised "path only, never an origin" and collapsed
a leading run of slashes so `//host/x` could not jump the frame off the proxy,
but it left two spellings through that the WHATWG URL parser treats the same way:
a backslash, which is read as `/` for http(s) schemes, and an ASCII tab or
newline, which the parser deletes before it looks at anything, so `/\host/x` and
`/<tab>/host/x` both resolve to `https://host/x`. That mattered only in
direct-mode tabs, where `POST /api/webviews/:id/open` returns no embedUrl and the
recovered path is resolved with `new URL(path, src)` straight into the frame's
src; a page in such a tab could remount its own frame on a foreign origin.

Not an escalation (the page can already navigate itself anywhere, and the remount
carries no Codeman-origin access), but the comment did not hold and the existing
test only covered the form that already worked. The handler now strips tab, CR
and LF, collapses any leading run of `/` or `\` to one `/`, and refuses whatever
still opens a second separator. The new test drives the reachable direct-mode
branch with all four spellings and fails without the fix.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:38:54 +02:00
Codeman maintainer b0dddc9c57 Merge pull request #402 from shenlvkang-collab/pr/webview-route-masking
fix(webview): let a proxied single-page app route on its own path, and recover a frame that reloads
2026-09-14 23:38:54 +02:00
Codeman maintainer f5f399a8b7 test(docker): pin cap_add against the entrypoint, the PATH order and git_head_commit
The capability list is DERIVED from what the scripts do (chown => CHOWN +
DAC_OVERRIDE, a setpriv uid/gid drop => SETUID + SETGID, `init: true` next to a
uid drop => KILL) and compared to docker-compose.yaml's cap_add, the
entrypoint's own required_caps diagnosis, and the lists quoted in docker/README.md
and CLAUDE.md, so the drift that shipped the missing CAP_KILL fails here rather
than on someone's server. Also pinned: the CLI prefix is appended to PATH in
server.Dockerfile and entrypoint.sh pins its PATH before its first command;
Start-Codeman.sh derives PUID/PGID before creating the cases dir, builds before
`down`, writes the source marker only after a refresh, and never aborts on a
failed volume removal.

git_head_commit is run as the script defines it, extracted by its own
delimiters into a real bash, against temp repos made with real git: a symbolic
ref with a loose ref file, a detached HEAD, packed refs after `git pack-refs`,
a linked worktree (which must resolve nothing rather than something wrong) and
a directory that is not a checkout.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:37:04 +02:00
Codeman maintainer 2bda191471 docs(docker): describe the root start and drop, and keep the override file out of the image
docker/README.md and docs/docker-compose.md now say that the container starts
as root, corrects a daemon-created bind source and drops to PUID:PGID with
setpriv, which capabilities that needs, and that a compose file written
elsewhere must carry them. The README's PowerShell example runs Compose from
inside docker/ so the override file is discovered, instead of the `-f
docker/docker-compose.yaml` form its own Local customisation section warns
silently drops it, and the reverse-proxy section no longer asks for an override
file now that docker-compose.yaml forwards CODEMAN_ALLOWED_HOSTS itself.

.dockerignore excludes docker-compose.override.* everywhere: it is the
documented home for host-specific settings and rode `COPY . .` into the image,
the same shape as the docker/.env exclusion above it (verified with a scratch
build context: the override files and docker/.env are absent, .env.example and
the compose file present).

CLAUDE.md's Compose paragraph carries the corrected cap list, the writability
probe, and the two traps behind it (KILL is for tini, the CLI prefix is
appended to PATH).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:37:04 +02:00
Codeman maintainer f92883704e fix(docker): create the cases dir with the runtime owner and record a refresh only when it happened
Start-Codeman.sh created CODEMAN_CASES_PATH with a plain `mkdir -p` BEFORE it
derived PUID/PGID from the appdata directory, so the new directory landed as
the invoking user's uid and primary gid. On a host set up the way the README
suggests (`chown -R 99:100 <appdata>`) that gid is not PGID, and the container
refused to start on a directory the script had just made. PUID/PGID are now
derived first and the directory is chowned to them right after creation, with
a clear host-side error when that is not possible. As root this always works,
which also retires the old "refusing to create as root" branch for this path.

The build-artefact volume refresh had three holes. The docker-build-source.json
marker was written whether or not a volume had actually been removed, and the
project name came from a sed over `docker compose config --format json` keyed
on two-space indentation: an empty name made the label filter match nothing,
nothing was removed, and the marker recorded the new HEAD, so the check never
fired again while the stale volume kept serving old code. The name is now
parsed indentation-agnostically, an empty result falls back to `down --volumes`
(the documented reset; both volumes re-seed from the image by a plain copy),
the marker is written only after a successful refresh, and a failed `docker
volume rm` warns and leaves the marker alone instead of aborting under set -e
with the stack down. The image is also built BEFORE `down`, so the deployment
is offline only for the recreate rather than for the whole rebuild.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:37:04 +02:00
Codeman maintainer 1851d80f3a fix(docker): keep SIGTERM reaching the server, pin root's PATH, probe writability
Three changes to how the Compose container starts as root and drops to
PUID:PGID, each reproduced on Docker 29.1.3 / Compose v5.5.0 with a minimal
image of the same shape as server.Dockerfile.

- cap_add gains KILL. `init: true` makes tini PID 1, and tini stays root while
  the entrypoint drops the server to PUID. Signalling a process of a different
  uid needs CAP_KILL, and `cap_drop: ALL` had removed it, so every `docker
  compose down`/`restart` ended in `[FATAL tini (1)] Unexpected error when
  forwarding signal: 'Operation not permitted'` and the server being SIGKILLed
  instead of running `server.stop()`. Measured: without KILL the trap never
  fires, with it the child logs `GOT SIGTERM`.
- /opt/codeman-cli/bin is appended to PATH, never prepended, and entrypoint.sh
  pins its own PATH to the system directories before its first command. The
  prefix is chowned to the runtime account so sessions can update the agent
  CLIs in place, and the root entrypoint resolved stat/chown/setpriv by bare
  name through it: a `setpriv` planted there by the unprivileged uid ran as
  uid 0 at the next start. The image's full PATH is handed back to the server
  at the exec (`env PATH=...`), since Codeman resolves the CLIs through it.
- The ownership gate becomes a writability probe. A directory owned by neither
  root nor PUID:PGID is no longer refused on ownership alone; it is tested with
  `setpriv --reuid PUID --regid PGID --groups <same groups> test -w`, the exact
  identity the server gets, so a group-writable tree, an ACL or a CIFS/NFS
  mount reporting some unrelated uid all pass, and the refusal names path,
  owner and PUID:PGID. Root-owned directories are still chowned first.

Also: a pre-flight runs the drop before touching anything and, when it fails,
prints the cap_add list the compose file needs, so an out-of-tree compose file
(Unraid's Compose Manager) gets a one-line diagnosis instead of a restart loop;
`--bounding-set -all` is gone, since it is a silent no-op without CAP_SETPCAP;
a root:root Docker socket now produces a warning that Docker cases will not
work rather than silently losing group 0 at the drop; and CODEMAN_ALLOWED_HOSTS
is forwarded from .env with an empty default (documented as a commented entry
in .env.example so the parity test and the updater's env gate both stay quiet).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:37:04 +02:00
Codeman maintainer a29e1f61ef Merge pull request #377 from opticon454/bugfix-docker-user-perms
fix(docker): bind-mount ownership, Compose override discovery, and the default runtime account
2026-09-14 23:37:04 +02:00
Codeman maintainer 653e3cdf96 Merge pull request #423 from Ark0N/fix/xterm6-selection-background
fix(terminal): name the selection colour the way xterm 6 does
2026-09-14 23:35:43 +02:00
Codeman maintainer e54a8b1189 Merge pull request #422 from Ark0N/test/install-dsh-probe-bash32
test(ci): exercise the dsh identity probe with timeout missing (bash 3.2)
2026-09-14 23:35:43 +02:00
Codeman maintainer 44a754ea73 test(ci): exercise the dsh identity probe with timeout missing
The bash 3.2 job added with #380 cannot reach dsh_banner_probe, which is
the function #382 was filed against: this image ships `timeout`, so the
optional-prefix array is never empty, and with no `dsh` binary anywhere on
PATH the probe is not called at all. The fix landed in 1.28.2 with nothing
guarding it, and the failure mode is a runtime abort under `set -u` that
`bash -n` cannot see, which is precisely why the reporter had to find it by
reading the source rather than by running anything.

So call the probe directly, with `timeout` hidden behind a narrowed PATH,
and refuse to pass if `timeout` is still reachable (a guard that silently
stops exercising its branch is worse than no guard). Both directions are
asserted: a real DeepSeek Harness banner is accepted, and Debian's unrelated
`dsh` is refused, so the check covers the identity half too.

Verified by reverting install.sh to the pre-fix expansion, where the step
fails with the exact error from the issue, `runner[@]: unbound variable`.

Refs #382

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 23:33:41 +02:00
Codeman maintainer 9acc5aad50 fix(terminal): name the selection colour the way xterm 6 does
Every per-skin xterm palette declared its selection layer as `selection`,
the key xterm.js renamed to `selectionBackground` in v5. An ITheme is a
plain object handed straight to the terminal, so an unknown key is not an
error, it is dropped: all seven skins have been drawing xterm's built-in
default, rgba(255,255,255,0.3), rather than the colour sitting next to it
in the palette.

Nobody saw it on the dark skins, where white at 30% is close to what those
palettes asked for. On the four light skins it is white over a near-white
background: blended, Paper Gray's selection differs from its own background
by 3/255. That is not a subtle highlight, it is no highlight, and it looks
exactly like a selection gesture that failed, which is part of what #360
reports on Android Chrome.

test/skin-themes.test.ts pins both halves: the key name, and that the
blended selection stays at least 16/255 from the background on every skin,
plus the light-skin fallback landing under that floor, which is what makes
this a fix rather than a rename.

Refs #360

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 23:33:36 +02:00
Codeman maintainer 7c3c5b8f72 fix(mobile): show the Codex shift-arrow keys only on codex sessions
The two keys #408 adds to the mobile keyboard accessory bar send
Shift+Left and Shift+Right, which are Codex bindings (edit the last
queued message, step back through the prompt stack). They shipped on
both agent layouts, so a claude, pi, grok, omp, deepseek or gemini
session got two keys that do nothing. That was not only cosmetic: a tap
goes through sendNavKey(), which adds the session to
_echoPassthroughSessions and hands editing to plain PTY echo until Enter
or Ctrl+C, so on a phone a dead key also switched off the local echo
that makes typing feel instant there.

The reveal now follows the shape the 🧠 key already uses. The buttons
stay in both templates, carry an accessory-btn-codex marker class, and
are display:none in styles.css until the bar element carries
codex-enabled. The class has to live on the bar rather than on the keys
because setMode() rebuilds the buttons' innerHTML on every layout
switch. syncCodexKeys() toggles it from the active session's mode
(the same lookup _isShellSession() uses) and is called at init and from
refreshForActiveSession(), which selectSession() already invokes on
every switch. A session's mode is readonly on the server and fixed at
create, so no other event can change the answer; the welcome screen
(no active session) reads as not codex and hides the keys.

The frontend id-branching guard (test/cli-registry-no-id-branching.test.ts)
scans only src/**/*.ts, so the mode comparison in a public JS file is
in bounds, the same as the existing shell check beside it.

Tests: the new describe block in test/mobile-shell-keyboard.test.ts pins
the marker class in both templates, the CSS pair, the class for a codex
session in both layouts, its absence for claude/shell/pi/omp/deepseek,
the re-sync in both directions on a session switch, the no-session case,
and the init + refresh wiring. All six positive assertions fail without
the source change. README and the changeset now say the keys are
Codex-only.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 23:30:14 +02:00
Codeman maintainer 1e5a53830f Merge pull request #408 from shenlvkang-collab/feat/codex-shift-arrow-keys
feat(mobile): add Shift arrows for Codex queued input and prompt navigation
2026-09-14 23:30:14 +02:00
Randalix 63aafdf274 fix(files): serve remote-case attachments, the path a click takes outside the case
A clicked path that points OUTSIDE the case directory goes through the attachment
routes (the frontend's `_isExternalPreviewPath` sends every absolute path not under
`workingDir` to `POST /attachments`), and those had the same local-`fs` assumption
as file-raw: `realpathSync`/`fs.stat` on a path that only exists on the remote host,
so the file never opened — the case the #415 report was actually about.

- `registerExternalAttachment()` accepts `remote` and resolves through
  `remoteProbePaths` (canonical path, size/mtime, kind, plus the workspace root for
  the confinement check). Everything around it — blocklist, extension allowlist,
  workspace confinement, registry/dedupe — is now shared by both branches, so the
  remote path cannot drift from the local one.
- The by-id routes (`raw`, `preview`, `thumbnail`), the metadata poll and the
  attachment history list resolve over ssh too. `raw` streams with the same
  Range contract as file-raw; `preview` (office) and `thumbnail` answer 400 for a
  remote record; an unreachable host answers 502, a vanished file 404.
- Which host a record is read from follows the SESSION, never the path string: the
  same absolute path is a different file on each host, and a remote session never
  falls back to a local file with that name.
- Codex generated artifacts keep force-workspace confinement for a remote case: the
  well-known artifact directories are anchored at THIS host's home, so only a file
  inside the remote workspace is trusted.

Still local-only by design: writes, office conversion, thumbnails, the file
tree/picker and tail-file.
2026-09-14 17:06:42 +02:00
Codeman maintainer c03714eb74 chore: version packages
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 16:21:09 +02:00
Codeman maintainer 1ca0a33830 chore: version packages
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 16:20:36 +02:00
Ark0N 0c00a40530 Merge pull request #407 from Ark0N/feat/iphone-duo
iPhone Duo support: fold-aware dialogs, and a fold is no longer mistaken for the keyboard
2026-09-14 16:10:38 +02:00
Codeman maintainer e46089bc7f fix(statusline): print nothing instead of the bare word codeman
Ported from #416 (discussion #405): a statusline reading just `codeman`
is what a hand-run claude in a managed repo showed, and it reads as a
broken config rather than a footer. Three paths produced it and all
three now yield an empty footer: the exporter's `|| echo codeman`
fallback (now `curl -sfk ... || true`, with -f keeping an HTTP error
body off stdout), the unknown-session answer of POST /api/status-telemetry,
and formatSessionStatusText() with nothing to show. The exporter script
marker moves to V4 so live installs pick the new content up on the next
spawn.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:59:00 +02:00
Codeman maintainer fa1ea8d9fe fix(statusline): unset a stale user statusline var, write the exporter script atomically
Three small follow-ups from the #361 review.

A tmux setenv survives respawn-pane, so _configureStatusLineUserCommand
returning early when the user has no statusline left a previously
exported CODEMAN_USER_STATUSLINE_CMD in place: a user who deleted their
own statusline kept getting the stale one wrapped, and lost Codeman's
footer print-through, until the tmux session was recreated. It now
issues `setenv -u` in that case, the same shape as the effort-level
cleanup in applyEnvOverrides.

ensureStatusLineExporterScript truncated and rewrote a script that live
sessions execute on every statusline render, and chmod'd it after the
write. It now writes a temp file next to the target, chmods that, and
rename()s it into place.

The non-tmux direct-PTY fallback carries no exporter; that is now stated
at the spawn site and in the architecture-invariants paragraph rather
than left as a silent gap.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:59:00 +02:00
Codeman maintainer 707ea345eb fix(statusline): GET /api/settings never writes, and a save sends the collection switch only on a flip
Two follow-ups to #361's sticky telemetry switch.

GET /api/settings reconciled an absent showPlanUsageLimits by persisting
true, but readJsonConfig() answers {} for ANY read failure (a parse
error, EACCES, EMFILE, a read landing inside PUT's non-atomic write), not
only ENOENT, and every page load calls this route, so one unlucky read
replaced the whole settings file with a one-key file. The route is a
plain read again and the default moved into the reader:
readPlanUsageTelemetryEnabled() treats an absent key as ON, the same way
readWorkspaceHooksEnabled() does, which is what the desktop chip already
shows for an install that never touched the setting.

saveAppSettings() sent showPlanUsageLimits on every save. The chip
defaults OFF on handhelds, so a phone saving its font size persisted
false and switched collection off for every desktop, whose chip then
went stale with no error anywhere. The key is now stripped like the
other per-device display keys and re-added only when the save FLIPS the
chip relative to what the device had (planUsageCollectionFlip), so an
explicit toggle on any device still writes it in either direction.

Tests pin both: the GET route with a mocked filesystem (absent, missing,
EACCES, garbage, explicit), the reader default, and the flip helper plus
its wiring in saveAppSettings.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 15:59:00 +02:00
Ark0N b2b2c767ea Merge pull request #361 from timkjr/fix/statusline-injection-opt-out
fix(statusline): inject plan-usage telemetry via ephemeral CLI flag, never disk
2026-09-14 15:58:48 +02:00
Randalix 013a5d9cc8 fix(files): read remote-case file previews and downloads over ssh
A remote case's workingDir is an absolute path on the remote host, but the
file read routes resolved it with local `fs`: `validateSessionFilePath`'s
realpathSync fails for a path that does not exist on the Codeman host, so
every preview of an agent-written file answered "File not found" (#415).

Add src/remote-files.ts as the single remote-read layer, built on the same
buildSshConnectionArgs() the launch uses:

- remoteProbePaths(): ONE round trip returning realpath + stat for the
  requested path AND the workspace root, so containment is checked against a
  remotely canonicalized root (a symlinked remotePath is ordinary).
- remoteCreateReadStream(): streams the body (cat, or tail -c +N | head -c L
  for a Range) with nothing buffered in memory, and reaps the ssh child when
  the response ends so an aborted download cannot orphan it.
- remoteReadFile(): bounded read for file-content.

file-raw, file-content, file-preview and file-thumbnail now share one local/
remote target resolution. Guards keep their local strength: lexical pre-check,
remote realpath, workspace containment, sensitive-path blocklist, and the size
cap applied to the remote size before any bytes are read. An unreachable host
answers 502 with the remote reason instead of a misleading 404. Nothing is ever
copied to the Codeman host and there is NO local fallback (an sshfs mount of
the same tree must not shadow the remote bytes).

Deliberately unchanged: writes (edit=1 / PUT now answer 400 explicitly while
the viewer hides its Edit affordance), office previews, thumbnails, file tree,
picker, external attachment registration and tail-file stay local-only.
2026-09-14 14:54:41 +02:00
DevvynandClaude Sonnet 5 b6f75b87f5 fix(custom-model): don't clamp DEEPSEEK_API_KEY as a privileged env key
CI caught a real regression: DEEPSEEK_API_KEY was added to deepseek's
privilegedEnvKeys alongside DEEPSEEK_BASE_URL on the theory that "the pair
travels together," but that contradicts the documented and tested design
(clampEnvOverridesForOwner()'s own docstring in session-routes.ts) — a
non-granted owner supplying their OWN DeepSeek key removes privilege
rather than granting it, since the exfiltration vector is the BASE URL
(which redirects the server's own forwarded key to a foreign host), not
the key itself. Removed it from the list; test/deepseek-mode.test.ts's
existing two clamp tests now pass again.

Also swapped that test's "unrelated override" example off CODEX_HOME,
which the earlier commit in this same PR legitimately made privileged
(closing a real pre-existing gap, documented in PR.md) — so it stopped
being a valid "unrelated" example the moment that fix landed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00
DevvynandClaude Sonnet 5 e18499aa67 docs(pr): drop the draft/WIP framing now that the PR is submitted
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00
DevvynandClaude Sonnet 5 61779745aa test(custom-model): make the harness smoke test dynamic, verify all 9 CLIs end-to-end
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.

Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.

Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:

- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
  speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
  don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
  undocumented GATEWAY AuthType gemini-cli selects once
  GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
  tried
- deepseek: reaches the server (env vars are read) but gets a consistent
  HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)

Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).

deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00
DevvynandClaude Sonnet 5 41416566aa feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its
native cloud backend, for a given session. Covers local hardware (llama.cpp,
Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter).
Off by default (customModelEndpointsEnabled, synced, default OFF).

- Registry: capabilities.customModelInjection per CLI entry (env /
  configContentEnv / configDir / unsupported kinds)
- Pure injection builder (custom-model-injection.ts) turning an endpoint +
  model id into the real env vars / config content per CLI
- Endpoint store + CRUD routes (custom-model-hosts.ts,
  custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded
- Session integration: Session.setCustomModel()/restartCli()
  (POST /api/sessions/:id/custom-model), reusing the existing
  respawn-pane -k primitive to restart the CLI process with new env
- Multi-user hardening: every new redirect-capable env var added to its
  CLI's privilegedEnvKeys, closing a pre-existing gap where several were
  already reachable via the generic envOverrides field's prefix allowlist
- Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries
  against a real endpoint outside the web UI, independent of tmux/sessions
- Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying
  every CLI's injected values through a real HTTP shape

Real end-to-end validation against a live llama-swap server (inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries) found and
fixed three real bugs before they shipped:
- Codex's config.toml schema was wrong ([model].default table instead of
  a top-level model string + [model_providers.custom]); fixing it then
  surfaced a genuine, documented protocol incompatibility (Codex only
  speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
  don't implement)
- Claude Code's async session-title-generation call validates
  ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and
  hangs the whole -p invocation on an unrecognized name; documented for
  chunk 6, worked around in the standalone script only (--bare is NOT
  safe for a real interactive session, which needs hooks)
- The discovery route's authStyle: 'both' option (send both Authorization
  and api-key headers) reliably hung a real server; removed the option
  entirely rather than just changing the default

Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see
PR.md and deployment_plan.md for the full chunk breakdown and confidence
table.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
2026-09-13 17:42:35 +08:00
DevvynandClaude Sonnet 5 c179daf869 fix(docker): re-assert /opt/codeman-cli ownership every start, not just at build
/opt/codeman-cli is chowned to PUID:PGID once, at image build time, from
the PUID/PGID build args. That bake only happens when the image is
actually rebuilt (`docker compose up --build`, which Start-Codeman.sh
always does) — a deployment that runs the compose file directly instead
(Unraid's Compose Manager, a native systemd unit, any plain
`docker compose up`/`restart`) can change PUID/PGID in .env and restart
without ever rebuilding. The container then runs as the NEW uid via
entrypoint's setpriv (Linux needs no /etc/passwd entry to setuid to an
arbitrary number) while the CLI directory is still owned by the OLD one
baked into the image layer — silently breaking the self-update-a-CLI-
in-place fix that directory exists for.

Unlike HOME/CODEMAN_CASES_PATH, this one is pure image content Codeman
itself populated, never host data that might legitimately belong to
someone else, so there is no ownership to be careful about — it is
always correct for it to be owned by whoever the container is about to
run as. Re-assert it unconditionally on every start.

Verified live: built an image with PUID=99/PGID=100, ran it with
PUID=1234/PGID=4321 (no rebuild, simulating a changed .env restarted
directly), confirmed /opt/codeman-cli ends up 1234:4321-owned and is
genuinely writable by the running process.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
2026-09-13 17:41:31 +08:00
DevvynandClaude Sonnet 5 ae32daf135 fix(docker): address maintainer review on #377
Two real bugs the review caught, both verified live against a real
build on the Unraid host:

1. entrypoint.sh's chown fired on ANY ownership mismatch, not just a
   directory the daemon itself created root-owned. A host tree
   legitimately owned by some other account - an existing
   CODEMAN_CASES_PATH the README already allows pointing at a normal
   projects directory, or appdata under a different PUID/PGID
   convention than the one in use - got silently recursively re-owned
   with one log line to explain it. Now gated on the target actually
   being root-owned; anything else is a clean refusal naming the
   directory, its owner, and PUID/PGID. Start-Codeman.sh also now
   pre-creates CODEMAN_CASES_PATH the same way it already did
   CODEMAN_APPDATA_PATH, so Compose never has to materialise a missing
   bind source as root in the first place - the in-container chown
   becomes a safety net, not the primary mechanism.

2. The CLI-update chown (chown -R .../node_modules /usr/local/bin)
   handed the runtime account write access to entrypoint.sh itself
   (root-owned, executed as root on every container start with
   CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node binary - owning the
   DIRECTORY is enough to rename it aside and drop a replacement, which
   would let a compromised session arrange for its own script to run
   as root at the next restart. The four CLIs now install into a
   dedicated /opt/codeman-cli prefix (NPM_CONFIG_PREFIX); only that
   directory is chowned, /usr/local stays root-owned throughout.

Smaller fixes from the same review:

- Start-Codeman.sh's volume-refresh label filter wasn't project-scoped:
  a second Compose stack on the same host sharing the `codeman-dist`
  volume KEY could have had ITS volume deleted. Added a
  com.docker.compose.project filter, resolved from this stack's own
  `compose config --format json`.
- Override-file precedence was backwards (checked .yaml before .yml;
  Compose actually prefers .yml) - swapped, plus a warning when both
  exist.
- entrypoint.sh's setpriv now also passes --bounding-set -all, so
  CapBnd actually clears post-drop rather than just CapPrm/CapEff.
- A comment on git_head_commit() noting it returns nothing for a
  worktree checkout (.git as a file), consistent with the script's
  existing -d .git convention elsewhere.
- Doc drift: CLAUDE.md's Docker Compose section still described the
  old pre-created-and-chowned-by-hand model and didn't mention the
  root-then-drop entrypoint; the state-files list was missing
  docker-build-source.json; docs/docker-compose.md and
  docker/.env.example still had the pre-rename `Coding/codeman` path
  in one place each.

Verified end to end against a real build on the Unraid host: a
root-owned bind source is corrected as before; a directory owned by
neither root nor PUID:PGID is refused rather than silently rewritten;
a correctly-owned directory is left alone entirely; the four CLIs
resolve via PATH from /opt/codeman-cli while /usr/local/bin,
/usr/local/lib/node_modules and entrypoint.sh itself stay root-owned;
CapBnd is fully cleared post-drop.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
2026-09-13 17:41:31 +08:00
DevvynandClaude Sonnet 5 8fe3f34fc5 fix(docker): detect and refresh stale build-artefact volumes
codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded
from the image only while empty, so a rebuilt image's fresh dist/
node_modules sat unused behind old volume content until something
cleared it. The in-app self-updater never hit this (it rebuilds INSIDE
the running container, into the very volume already in use), but a
`docker compose build` triggered from outside it — Start-Codeman.sh,
after a manual `git pull` — did: the container came back up looking
unchanged, serving stale compiled routes against current source.

Start-Codeman.sh now compares the checkout's HEAD commit and
package-lock.json hash against a recorded marker
(docker-build-source.json) and clears just the affected volume(s)
before its own --build when either moved.

The in-place self-update path writes that same marker after a
successful build, so the two mechanisms agree on what the volumes
currently reflect — without it, the next plain Start-Codeman.sh run
would see the HEAD self-update just checked out, not recognise it as
already accounted for, and wipe the volumes self-update just correctly
rebuilt right back to the older baked image.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
2026-09-13 17:41:31 +08:00
DevvynandClaude Sonnet 5 89e2cb5814 fix(docker): let the runtime account update its own global CLIs
The four CLIs (claude, gemini, codex, opencode) are npm-installed
globally as root during the image build, before the unprivileged
runtime account exists. A session running as that account (e.g. a
codex-mode terminal) then hits EACCES the moment it tries to update
one in place, because npm renames the old package directory aside
before installing the new one, which needs write access to the
parent (/usr/local/lib/node_modules), not just the target package.

Chown that tree plus /usr/local/bin's CLI symlinks to PUID:PGID in
the same step that creates/renames the runtime account.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
2026-09-13 17:41:31 +08:00
DevvynandClaude Sonnet 5 d38bf33a69 docs(docker): document the reverse-proxy host allowlist
CODEMAN_ALLOWED_HOSTS is a real, documented application setting (the Host-
header allowlist in network-auth-policy.ts), but docker-compose.yaml does not
forward it from .env into the container - Compose only passes through
variables explicitly listed under environment:, and this is not one of them.
Set without that passthrough, any request through a reverse proxy is rejected
with 403 Forbidden: host not allowed before it reaches any handler, and
nothing in the Docker deployment docs said why.

Document the variable and the override needed to forward it, using the
Local customisation mechanism already described above it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-13 17:41:31 +08:00
DevvynandClaude Opus 5 9702126046 chore(docker): name the default runtime account codeman
CODEMAN_RUNTIME_USER defaulted to `opencode`, which no longer matches the
project and is confusing in a deployment whose every other identifier is
codeman. Rename the default in .env.example and in the Dockerfile ARG that
mirrors it, and correct the example comment that referred to
/home/opencode/codeman-cases.

Also drop the `Coding/` component from the example application-data path.
CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH now suggest /mnt/user/appdata/codeman
and its codeman-cases child, matching the account name and removing a directory
level that meant nothing outside the original author's host. README.md is
updated to match, including the chown example.

The npm package `opencode-ai` and the references to the OpenCode CLI are
deliberately left alone: those name a different tool, not this account.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 17:41:31 +08:00
DevvynandClaude Opus 5 748bbf5423 fix(docker): honour docker-compose.override.yml in Start-Codeman.sh
Naming a Compose file with -f disables Compose's automatic discovery of the
override file, so Start-Codeman.sh silently ignored docker-compose.override.yml.
Any local customisation placed in the conventional override file was dropped
without warning, and the only way to notice was to inspect the running
container.

Collect the -f arguments into an array, append the override file when one is
present, and reuse that array for the final launch so the two cannot drift
apart again. Both .yml and .yaml are checked, in Compose's own precedence
order, and the chosen file is reported on startup.

Document the override file in docker/README.md, including the two things that
are easy to get wrong: it is ignored when -f is passed without naming it, and
it cannot remove a key such as ports, which Compose concatenates. Add the
override file to .gitignore.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 17:41:30 +08:00
DevvynandClaude Opus 5 10876aa440 fix(docker): correct bind-mount ownership before dropping privileges
Compose binds CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH from the host. When
either path does not exist yet - a first run, a cleared application-data
directory, a restored backup - the Docker daemon creates it owned by root. The
server runs unprivileged as CODEMAN_RUNTIME_USER, so it cannot create its own
state directory, and the container restarts forever on:

  Failed to start web server: EACCES: permission denied, mkdir '/home/<user>/.codeman'

Start-Codeman.sh already worked around this by preparing the directory on the
host, so the failure only appears when Compose is run directly, which the README
documents as a supported path.

Add docker/entrypoint.sh, which starts as root, corrects the ownership of both
bind mounts, then drops to PUID:PGID with setpriv. The Dockerfile's USER
instruction is replaced by that entrypoint and CMD is unchanged.
docker-compose.yaml adds back only the four capabilities the chown and the
privilege drop require, so cap_drop: ALL continues to remove everything else.

Two guards keep existing deployments working:

- A container started with an explicit `user:` is left alone. The entrypoint
  execs straight through, with no elevation and no chown.
- A chown that fails is a warning, not an error. Bind mounts backed by NFS,
  CIFS or a rootless daemon can refuse chown while remaining perfectly
  writable, and those deployments must keep starting.

PUID and PGID are also exported as runtime environment defaults so the image
behaves correctly when run without Compose, rather than depending on build args
alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 17:41:30 +08:00
codeman-local b357fe832e feat(mobile): add Shift arrow keys for Codex prompt navigation 2026-09-12 20:45:05 +08:00
timkjrandClaude Sonnet 5 aeb55c92b0 fix(settings): reconcile showPlanUsageLimits default on first read
planUsageChipEnabled() (settings-ui.js) shows the header chip and the App
Settings checkbox as already ON whenever showPlanUsageLimits has never been
set — a discoverability default from 1.9.3. readPlanUsageTelemetryEnabled()
(hooks-config.ts) deliberately treats an absent key as "no telemetry" — a
privacy default, pinned by its own unit tests (never POST usage data
without an explicit persisted yes). Nothing reconciled those two
independent guesses, so a fresh install showed a checked box that silently
collected nothing until the user opened Settings and hit Save at least
once.

Verified live: an install that had never touched this setting had no
showPlanUsageLimits key in settings.json at all, and its running Claude
process's argv carried no --settings flag — zero telemetry ever collected
despite the chip rendering as enabled.

GET /api/settings now persists the resolved default (true) the first time
the key is truly absent — not explicit false — so "chip visible" and
"telemetry collected" become the same fact. readPlanUsageTelemetryEnabled's
own absent-means-false contract is untouched; after this runs once the key
is never absent again, so that branch stays correct in isolation while
being unreachable in practice for any install that has ever called this
route. An explicit false set afterward is respected forever.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 19:10:25 -05:00
shenlvkang-collabandClaude Fable 5.1 349a89ec3b fix(webview): let a proxied single-page app route on its own path, and recover a frame that reloads
A dashboard served through a web tab saw `/webview/<cap>/` as its
`location.pathname`, and no app has a route for that: a React Router, Vue
Router or Vite dev-server page painted its HTML and CSS and then replaced
them with its own "page not found" the moment its script ran (reproduced
with a minimal history-routed page).

The proxy's runtime shim now rewrites the history entry to the path the
page would see on its own origin, before any page script runs. The base
element still resolves relative URLs inside the prefix and every root-
absolute sink is rewritten back into it, so only what the page READS
changes. With the document URL masked the Referer-keyed 404 rescue can no
longer help a request the shim misses, so the remaining URL-taking entry
points (`Worker`, `SharedWorker`, `navigator.sendBeacon`, `window.open`)
are covered by the shim as well.

A navigation the page starts itself afterwards — `location.reload()`
(a dev server's full-reload HMR), a root-absolute `location.href` — lands
on Codeman's root with no capability anywhere: no prefix in the path, no
cookie in an opaque-origin frame, a Referer naming the masked page. It is
recognised by shape (a top-level iframe navigation asking for HTML, for a
path Codeman does not serve) and answered with a static page whose only
script posts `{type:'codeman:webview-lost', path}` to the parent; the tab
that owns the frame (matched by `event.source`, never by the payload)
remounts it inside the prefix at that path, bounded per frame. The
unauthenticated form is answered in the auth middleware before the
credential checks, so a dev server that reloads on every save cannot
rate-limit its own user out of Codeman; the authenticated form (Basic
auth, trusted mode) is answered by the 404 handler.

Verified end to end against a history-routed page: boots on `/`, its
API call succeeds, a reload inside the frame comes back routed on the
path it had pushed, `location.href = '/about'` comes back on `/about`,
and a deep link opens on its path.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-10 14:13:15 +08:00
timkjrandClaude Sonnet 5 d5b75af628 fix(statusline): sticky telemetry collection, footer print-through, EOF fix
Responds to Ark0N's review round on the ephemeral-CLI-flag statusline
injection rework:

- Rebase-detail fixes: registry-gated telemetry eligibility via
  getCli(mode)?.capabilities.statusLineTelemetry instead of a hardcoded
  mode === 'claude' check, using the capability flag master's CLI-registry
  refactor already declares for exactly this purpose.

- Design question settled: sticky (a). Rather than persisting the toggle
  as a new field and threading it through every session-creation path
  (cron, Ralph Loop API, quick-start), eliminated the per-session field
  entirely. readPlanUsageTelemetryEnabled() (hooks-config.ts) reads the
  existing showPlanUsageLimits setting fresh from settings.json at every
  claude create/respawn (TmuxManager.createSession/respawnPane) - no
  per-session state to survive a restart, and it applies uniformly to
  every creation path for free, since they all flow through the same
  TmuxManager methods.

  This required fixing a real bug found along the way: showPlanUsageLimits
  was not actually round-tripping through settings.json on save -
  settings-ui.js explicitly excluded it from the PUT body as a pure
  per-device display key. It now flows through normally (both true and
  false); the load-side per-device merge behavior is unchanged.

  Removed entirely as a result: the statusLineTelemetry field from
  CreateSessionSchema/SettingsUpdateSchema, CreateSessionOptions/
  RespawnPaneOptions, Session._statusLineTelemetry (this is what makes
  the restart-persistence bug moot rather than patched), and the
  frontend send sites.

- Footer print-through restored: the no-user-statusline branch of the
  exporter script now runs the telemetry POST in the foreground so its
  own stdout becomes the in-terminal footer, falling back to a plain
  "codeman" marker only on curl failure.

- Background-subshell EOF fix: the wrap-a-real-statusline branch closes
  stdin too, not just stdout/stderr (`>/dev/null 2>&1 </dev/null &`) -
  the un-redirected subshell process itself, not curl, was what held a
  reader-to-EOF's pipe open for however long curl took to finish. Added
  curl --max-time 5 so a hung (not just refused) Codeman cannot wedge
  the render.

Tests: real-shell-execution tests for the footer/EOF fixes (fake curl
stand-in on PATH, real sh subprocess spawns, real elapsed-time
measurements - verified non-vacuous against a hand-reconstructed
old-style script), unit tests for readPlanUsageTelemetryEnabled.
Adapted two existing tests whose payloads referenced the removed field.
Fixed during independent code review: a stray indentation break and a
test exercising the wrong (legacy) exporter code path.

Docs synced: CLAUDE.md, docs/usage-limits-display-plan.md (old
disk-based section marked superseded, kept for history),
docs/architecture-invariants.md.

Full suite green: 352 files, 6780 passed, 12 skipped, 0 failed.
tsc/lint/format:check/frontend-syntax all clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 20:37:29 -05:00
timkjrandClaude Sonnet 5 e15e8e43e8 feat(statusline): wrap the user's own real statusline instead of skipping it
Now that the exporter no longer lives in a fixed per-case file, it can
compose with the user's actual configured statusline rather than just
backing off when one is found.

findEffectiveUserStatusLineCommand() walks Claude Code's own settings
precedence for a workspace: project-local .claude/settings.local.json
> project-shared .claude/settings.json > the user's global
~/.claude/settings.json. A legacy Codeman-marked entry left behind in
the project's own settings.local.json is never treated as a real user
command — it's skipped and precedence continues to the next layer.

The shared exporter script (bumped to a V2 marker so stale copies
self-heal) now fires the telemetry POST in a background subshell —
its own stdout/stderr discarded so nothing leaks into the visible
statusline, and confirmed non-blocking (~4ms, even against an
unreachable endpoint) — then, if the pane's environment carries
CODEMAN_USER_STATUSLINE_CMD, feeds it the same stdin blob and relays
its stdout as ours. Otherwise it falls back to the plain "codeman"
marker as before.

The discovered command is threaded to the pane via `tmux setenv
CODEMAN_USER_STATUSLINE_CMD` (_configureStatusLineUserCommand) rather
than embedded in the spawn command line, for the same
premature-shell-expansion reason as the parent commit: tmux stores a
setenv value verbatim and never re-parses it, so once shellescape()d
for that one command, the command's own $/quotes survive untouched
into the pane's environment.

Verified live via direct shell execution of the generated script
(both branches: fallback and user-command wrapping) before deploy.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015GyMnFWnUzc41TDeHg9juW
2026-09-07 19:18:36 -05:00
timkjrandClaude Sonnet 5 d4aa3c8cca fix(statusline): inject plan-usage telemetry via ephemeral CLI flag, never disk
Codeman's plan-usage chip wrote a statusLine.command into the case's
.claude/settings.local.json to receive Claude Code's rate_limits blob.
That file-based statusLine took precedence over the user's own
global/project statusline for ANY `claude` run in that directory,
including entirely outside Codeman, with no disclosure in the App
Settings UI (labeled only as a header-display toggle) and no way to
remove it once written (the removal code path was unreachable dead
code — nothing ever called it with false).

Replace the disk write with an EPHEMERAL `claude --settings
'{"statusLine":{...}}'` CLI flag, resolved fresh at spawn time
(resolveStatusLineCliCommand in hooks-config.ts) and merged with
effort/ultracode into one --settings object (buildClaudeSettingsFlag
in tmux-manager.ts, since Claude Code accepts only one --settings
flag). Never touches disk, so a plain `claude` run outside Codeman is
untouched. Self-healing: any legacy disk-written exporter from an
older build is stripped the first time a session starts in that
workspace again. Still respects a user's own hand-authored statusLine
(skips the flag entirely rather than overriding it).

Mid-fix bug found and fixed: the exporter's command legitimately
depends on $CODEMAN_SESSION_ID/$CODEMAN_API_URL/$CODEMAN_HOOK_SECRET_FILE
and an internal $INPUT, all meant to be expanded only when Claude Code
itself executes the statusline, using the pane's tmux-setenv'd
environment. Passing that text through --settings routed it through
execSync's own implicit /bin/sh -c first (tmux respawn-pane's
`bash -c "..."` wrapper) — POSIX double quotes don't suppress $
expansion, so those vars got expanded prematurely against the
server's own environment (unset there), producing malformed JSON that
printed as literal error text in the statusline. Fixed by writing the
exporter as a real, shared script file (ensureStatusLineExporterScript,
marker-versioned so stale copies self-heal) and passing only its bare
path via --settings — nothing for any intermediate shell to mangle.
Verified against a real Claude CLI on an isolated tmux socket, and via
direct execSync reproduction of the exact nested wrapping
createSession/respawnPane use.

A hard "never inject, even ephemerally" kill-switch was added and then
removed in the same pass: with the disk-leak fixed, disabling
injection only cost the plan-usage telemetry the feature exists to
provide, for no remaining benefit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015GyMnFWnUzc41TDeHg9juW
2026-09-07 19:15:54 -05:00
119 changed files with 10908 additions and 437 deletions
-48
View File
@@ -1,48 +0,0 @@
---
'aicodeman': minor
---
`install.sh` and the Docker agent image now read the shipped CLI catalogue instead of
hand-maintaining their own lists.
Adding a CLI to `src/config/cli-registry/stock.ts` and running
`npm run generate:cli-catalog` wires it into the installer's detection, its install menu and
its closing reminder, and into the agent image's npm layer. Previously each of those was a
separate hand-written list that had to be kept in step and was not: upstream `b6d0f1fa` is
"wire OMP into install.sh's CLI detection (it had none)", where a user with only `omp`
installed was told no AI CLI was found and offered Claude Code, and the section comment above
that code named six of the nine CLIs.
The generator emits two committed artifacts, because neither consumer can import TypeScript:
`config/clis.stock.json` for the Docker build, and a marked block inside `install.sh` itself,
which runs via `curl | bash` before any checkout exists. The embedded copy is the FULL
catalogue: an earlier attempt fetched it and fell back to a hardcoded two-CLI list, degrading
silently on an empty response, and there is no degraded mode to fall into now — nor a network
fetch at all, since a `curl | bash` from master already carries a catalogue exactly as fresh as
the script itself.
**Trust model is unchanged and now mechanical.** The server still never executes an entry's
install command. `install.sh` executes only commands embedded in itself — same file, same TLS
fetch, same commit as the `curl | bash` line that fetched it — and nothing pulled from the
network at install time is ever run, because nothing is fetched at install time at all.
**The agent image respects `enabled`.** The generated catalogue carries that flag, so a CLI
shipping disabled is no longer baked into every image. It reads the stock catalogue rather than
the merged registry, so a user's `~/.codeman/clis.json` cannot change what is inside an image
tagged `codeman/agent:base`.
User-visible changes, all in the installer:
- The install menu is built from the catalogue, so it offers every enabled CLI with an install command that can drive a pane on its own — eight today, rather than the previous fixed two. Gemini had a command in the registry and appeared in no list in the script at all. DeepSeek is the one enabled CLI with a registry command that is deliberately NOT offered: `npm install -g @deepseek-ai/dsh` installs only the launcher, which ships no profile that can drive a terminal on its own, so choosing it used to leave the user with an AI CLI the installer considered "found" but that could not actually run anything. It still gets a hint pointing at its docs.
- Its entries use the registry's labels ("Claude" rather than "Claude Code"), the same trade already made for `codeman doctor` rows. A suffix map would just be the hand-maintained list again.
- On a `wget`-only host, only the menu entries that actually need `curl` are held back (still shown as copy-paste hints); the `npm install -g` entries, which never needed it, are unaffected. Rewriting `curl` to `wget` inside a string about to be executed is the wrong instinct either way.
- `CODEMAN_NONINTERACTIVE=1` still defaults to Claude Code, unchanged.
`install.sh` remains bash 3.2 compatible (macOS ships it): parallel indexed arrays with
offset/length windows instead of delimiters, no associative arrays, namerefs, `mapfile` or
here-strings. CI now runs `bash -n`, executes the script inside a real `bash:3.2` container —
which is what catches expanding an empty array under `set -u`, a runtime abort `bash -n` cannot
see — and checks the generated artifacts are in sync.
`docker/server.Dockerfile` is deliberately untouched; its narrower CLI list is now asserted as
a declared omission list so the divergence is visible rather than accidental.
+1 -1
View File
@@ -10,7 +10,7 @@
"name": "codeman",
"source": "./plugins/codeman",
"description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.28.1",
"version": "1.29.0",
"author": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
+3
View File
@@ -9,6 +9,9 @@
**/.env
**/.env.*
!**/.env.example
# Same shape: docker/docker-compose.override.yml is the documented home for
# host-specific settings, so it must not ride COPY . . into the image either.
**/docker-compose.override.*
node_modules
dist
coverage
+26
View File
@@ -74,6 +74,32 @@ jobs:
offer_ai_cli_install >/dev/null 2>&1
echo "bash $BASH_VERSION: skipping the AI CLI install menu continues"
'
# Issue #382: the dsh identity probe builds an OPTIONAL `timeout` prefix as an
# array, and on stock macOS there is no `timeout`, so the array is empty and the
# expansion aborts the whole installer under `set -u`. The step above cannot
# reach that branch: this image HAS `timeout`, and with no `dsh` on PATH the
# probe is never called at all. So hide `timeout` and call it directly.
docker run --rm -v "$PWD":/w -w /w -e CODEMAN_INSTALL_SH_LIB=1 bash:3.2 bash -c '
set -euo pipefail
. /w/install.sh
printf "#!/bin/sh\necho \"DeepSeek Harness 0.1\"\n" > /tmp/dsh
printf "#!/bin/sh\necho \"dancer shell (Debian dsh)\"\n" > /tmp/not-dsh
chmod 755 /tmp/dsh /tmp/not-dsh
# A PATH the probe can still work on, minus the binary under test.
mkdir -p /tmp/nobin
for b in grep sh; do ln -sf "$(command -v $b)" "/tmp/nobin/$b"; done
export PATH=/tmp/nobin
if command -v timeout >/dev/null 2>&1; then
echo "timeout is still on PATH, so this is NOT exercising the empty-array branch" >&2
exit 1
fi
dsh_banner_probe /tmp/dsh
if dsh_banner_probe /tmp/not-dsh; then
echo "identity probe accepted a foreign dsh" >&2
exit 1
fi
echo "bash $BASH_VERSION: dsh identity probe survives a missing timeout"
'
- name: CLI catalogue artifacts are in sync with stock.ts
run: npm run generate:cli-catalog -- --check
+8
View File
@@ -48,6 +48,10 @@ Thumbs.db
.env.local
.env.*.local
# Local Compose customisation (host-specific, not part of the project)
docker-compose.override.yml
docker-compose.override.yaml
# State files (local to each machine)
.claude/ralph-loop.local.md
@@ -105,3 +109,7 @@ readme-preview.mjs
# Uploaded images land here under each session working dir (runtime artifact)
.claude-images/
# Local-LLM harness smoke-test config (real IPs/keys) — see the .example.json
# alongside it in scripts/, which IS tracked as the template.
scripts/local-llm-test.config.json
+148 -6
View File
@@ -1,5 +1,153 @@
# aicodeman
## 1.29.0
### Minor Changes
- **Custom model endpoints, HTTP API first** (#393). Any run mode that has a mechanism for it can be pointed at a custom OpenAI-compatible endpoint (a local llama.cpp, llama-swap, Ollama or vLLM, or a cloud gateway) instead of its native backend, per session. Endpoints are stored in `~/.codeman/custom-model-hosts.json` (`GET/POST/PUT/DELETE /api/model-endpoints`, admin-only in multi-user mode), their model lists are discovered from the endpoint's own `/v1/models`, and `POST /api/sessions/:id/custom-model` applies one to a session by restarting its CLI in place. The mechanism is per-CLI registry data (`capabilities.customModelInjection`): env vars for Claude, Gemini, Grok and DeepSeek, `OPENCODE_CONFIG_CONTENT` for opencode, an isolated config dir for Codex, Pi and OMP, unsupported for Antigravity. Verified live against a llama-swap server for claude, opencode, pi, grok and omp; gemini and deepseek reach the server and fail for reasons not yet understood, and codex only speaks the Responses API, so a plain chat-completions server cannot serve it. Those three are documented as gaps rather than shipped as working. The toolbar picker is a follow-up; until it lands the feature is HTTP-API only (`docs/custom-model-endpoints.md`), and the `customModelEndpointsEnabled` setting is declared but read by nothing yet. Merged with maintainer follow-ups: clearing a selection now actually clears it (the injected vars are delivered by `tmux setenv`, which `respawn-pane` inherits, so the relaunched CLI came back still pointed at the endpoint; retired keys are now `setenv -u`'d before the respawn), applying a model to a local claude session no longer kills the pane (the relaunch pins `--resume <id>` with the `--session-id` fallback, since Claude Code refuses a session id that already has a transcript), pi, omp and grok now select the generated model through a registry-declared `launchModel` (`custom/<id>`, `-m codeman-custom`) instead of writing a config the CLI then ignored, remote and Docker sessions are refused with a clear 400 until those paths are plumbed, the selection survives a Codeman restart, discovery goes through the egress-guarded `webviewFetch()`, key-bearing files are written 0600 and the per-session config dir is removed with the session, and the design plan moved from the repo root to `docs/custom-model-endpoints-plan.md`. Along the way the multi-user clamp learned about `GOOGLE_GEMINI_BASE_URL`, `GROK_BASE_URL`, `CODEX_HOME`, `PI_CONFIG_DIR` and `OPENCODE_CONFIG_CONTENT`, which were already reachable through `envOverrides` and now count as privileged keys.
**Single-page apps work as web tabs, and a frame that reloads comes back** (#402). A history-routed dashboard (React Router, Vue Router, a Vite dev server) read `/webview/<cap>/` as its `location.pathname` and rendered its own "page not found" the moment its script ran. The proxy's runtime shim now masks the prefix off the document URL before any page script runs, while every URL the page emits still goes through the rewrite layers (now including `Worker`, `SharedWorker`, `sendBeacon` and `window.open`). A navigation the page starts itself afterwards (a dev server's full reload, a root-absolute `location.href`) used to land on Codeman's root with no capability; it is now recognised by shape, answered with a static recovery page that posts the lost path to the owning tab, and the frame is remounted inside the prefix at that path, bounded to five recoveries a minute per frame. Merged with maintainer follow-ups: the recovery path is sanitised properly (a leading backslash, or a tab/newline the URL parser deletes before parsing, resolved `/\evil.com` to a foreign origin in a direct-mode tab); a reload on the dashboard's landing page is recovered too, on password-protected and passwordless installs alike (it used to render Codeman's own shell inside the web tab); and the recovery page is written down as the third unauthenticated 200 in the security table and `docs/security-architecture.md`, with the route-enumeration property it implies stated rather than left to be discovered.
**Shift arrows for Codex on the phone keyboard bar** (#408). Two keys, `⇧←` and `⇧→`, send the Shift-modified arrows Codex binds to editing the last queued message and walking the prompt stack (verified against Codex 0.154.0's `/keymap`). Merged with a maintainer follow-up: the keys are shown only on Codex sessions (a `codex-enabled` class on the bar, the same shape as the Read My Mind key), because tapping one in any other session did nothing except hand that session to plain PTY echo for the rest of the prompt.
**Remote (SSH) cases can finally show you their files** (#421, fixes #415). File previews, downloads, text reads and the out-of-workspace attachment path resolved every path against the Codeman host's own filesystem, so in a remote case every click ended in "File not found" while the file plainly existed on the other machine. A single new ssh read layer (`src/remote-files.ts`, built on the same `buildSshConnectionArgs()` the launch uses) probes realpath and stat for the file and the workspace root in one round trip, then streams the body with `cat` (or a `tail`/`head` slice for a `Range`), so the 200/206/416 contract holds and nothing is buffered on the server. Symlinks are resolved on the host that can resolve them, containment is checked against the resolved remote root, the size cap applies to the remote size before a byte is requested, an unreachable host is a 502 rather than a 404, and there is deliberately no local fallback: a same-named file on the Codeman host is never served under a remote name. Writes, Office previews and generated thumbnails answer 400 for a remote case instead of a misleading 404. Merged with maintainer follow-ups: the `readlink -f` fallback resolved only the directory chain, so on a host without it a symlink's final component was returned unresolved and `ws/notes.txt -> ~/.ssh/id_rsa` passed containment while `cat` served the key; it now follows the last component with plain `readlink` for a bounded number of hops and fails closed (404) on a loop or the cap; `PUT /api/sessions/:id/file-content` answers 400 for a remote case as the PR already claimed (it still validated against the local filesystem, so a same-named local directory took the write); ssh children are bounded by a small semaphore (`CODEMAN_MAX_REMOTE_FILE_SSH`, default 4) covering the attachment-history fan-out, which now probes the whole history in one batched call, and the fire-and-forget magic-link registrations an injected agent could use to fork hundreds of `ssh` processes; probe records are NUL-delimited and index-keyed so a newline in a filename cannot shift one path's result onto the next; and a 502 body never carries the ssh command line.
**Docker Compose: bind-mount ownership, override files, a `codeman` runtime account, and no more stale volumes** (#377). A missing bind source (first run, cleared appdata, restored backup) is created root-owned by the daemon, and the unprivileged server crash-looped on `EACCES` when Compose was run directly; the image now starts through an entrypoint that corrects a root-owned bind mount and drops to `PUID:PGID` with `setpriv`, and the compose file adds back only the capabilities that needs. `Start-Codeman.sh` honours `docker-compose.override.yml` (naming a Compose file with `-f` silently disables Compose's own discovery of it), pre-creates the cases directory like it already did for appdata, and detects when the checkout's HEAD or lockfile moved under the `codeman-node-modules`/`codeman-dist` volumes and refreshes them, which used to leave a `docker compose build` serving stale compiled routes. The default runtime account is named `codeman` (it was `opencode`), the four global agent CLIs live in their own `/opt/codeman-cli` prefix so the runtime account can update them in place without owning `/usr/local/bin`, and `CODEMAN_ALLOWED_HOSTS` is documented and forwarded. Merged with maintainer follow-ups: `cap_add` gains `KILL` (with `init: true` tini runs as root while the server runs as `PUID`, and without CAP_KILL its SIGTERM forward failed and the server was SIGKILLed on every `compose down`/`restart`); the CLI prefix is appended to `PATH` rather than prepended and the root entrypoint pins its own `PATH`, since a `PUID`-writable directory ahead of `/usr/bin` let the runtime account plant a `setpriv` that ran as root on the next start; the entrypoint decides with a real writability probe as the runtime identity instead of an owner comparison, so ACLs, group-writable trees and NFS/CIFS mounts work and only a genuinely unwritable directory is refused, by name; the cases directory is created with the runtime owner after `PUID`/`PGID` are known; the build-source marker is written only when a refresh actually happened, an empty Compose project name falls back to `down --volumes`, the build runs before the `down` so the stack is offline only for the recreate, `docker-compose.override.*` stays out of the image, and `test/docker-entrypoint.test.ts` pins `cap_add` against what the entrypoint needs. ⚠️ Compose users: run `Start-Codeman.sh` once for this release rather than a plain `docker compose up`, so the rebuilt image, the refreshed volumes and the new entrypoint arrive together.
**Selected text is visible again on the light skins** (#423, part of #360). Every skin palette named its selection layer `selection`, the key xterm renamed to `selectionBackground` in v5, so all seven skins had been painting xterm's default white at 30% instead of the colour next to it in the palette. Dark skins hid it; on the four light skins a selection was white on near-white. The key is renamed and `test/skin-themes.test.ts` pins it. CI additionally exercises `install.sh`'s dsh identity probe with `timeout` missing under bash 3.2 (#422), the guard #382's fix shipped without.
**Eight fixes salvaged from #375** (dignfei; landed with the author's commits preserved, the rest of that PR is covered below). Shift+drag starts a text selection in a pane whose mouse reports go to the CLI, and right-click copies the selection. Ctrl- and Alt-modified navigation keys typed through the CJK composer reach the CLI as the modified sequences instead of plain arrows. A browser whose reliable-input sequence counter fell behind the server's watermark (a restored tab, a cleared localStorage) now recovers: the duplicate ACK carries `dup: true` plus the watermark, the client lifts its counter and re-sends, so a session that had silently stopped accepting typed prompts accepts them again. An SSE reconnect that lands on the session you are already looking at keeps its terminal buffer and resyncs instead of resetting the whole terminal. The hidden offline overlay and the file-preview overlay only apply `backdrop-filter` while shown, which removes a stale compositing layer that swallowed clicks. One adopted Docker container can back several cases at different in-container directories, and the adopt panel gains a "copy an existing case" picker. Of the PR's 27 commits, 14 had already shipped through #357, the selection theme key rename shipped as #423, and foreign tmux adoption plus SSH password auth stay with the author.
### Thanks
- **@opticon454** for custom model endpoints (#393), including the part nobody enjoys: working out each CLI's real endpoint mechanism against real binaries and writing down which ones do not work yet instead of claiming they do; and for the Docker Compose deployment fixes (#377), rebased and reworked through three review rounds.
- **@shenlvkang-collab** for making single-page apps route inside web tabs and recovering a frame that reloads (#402), the best-engineered PR of this batch, and for the Codex Shift arrows on the phone keyboard bar (#408), verified against Codex's own keymap.
- **@dignfei** for the eight fixes salvaged from #375 (terminal selection and copy, CJK navigation keys, input recovery, SSE reconnect, overlay compositing, multi-case adopted containers), landed under their own name.
- **@Randalix** for reporting #415 and then fixing it themselves with the whole missing ssh read side for remote cases (#421), with a real-shell test for the probe script and a full route suite.
### Patch Changes
- 349a89e: fix(webview): let a proxied single-page app route on its own path, and recover a frame that reloads
A dashboard served through a web tab saw `/webview/<cap>/` as its `location.pathname`, and
no app has a route for that: a React Router, Vue Router or Vite dev-server page painted its
HTML and CSS and then replaced them with its own "page not found" the moment its script ran.
The proxy's runtime shim now rewrites the history entry to the path the page would see on its
own origin before any page script runs, while every URL the page emits still goes through
the existing rewrite layers (plus `Worker`, `sendBeacon` and `window.open`, which the masked
Referer can no longer rescue). A navigation the page starts itself afterwards — a dev
server's full-reload HMR, a root-absolute `location.href` — lands on Codeman's root with no
capability; it is recognised by shape (an iframe navigation asking for HTML for a path Codeman
does not serve), answered with a static page that tells the owning tab which path was lost,
and the tab remounts the frame inside the prefix at that path. That answer is served before
the credential checks, so it never counts as a failed login.
- 013a5d9: File previews, downloads and text reads now work in a **remote (SSH) case**.
A remote case's working directory is an absolute path on the _remote_ host, but the
file routes resolved it with local `fs` — so a clicked path (or the File Viewer) always
failed as "File not found" even though the file existed and the session was clearly
working in that directory. `GET /api/sessions/:id/file-raw`, `file-content`,
`file-preview` and `file-thumbnail` now resolve and read through the same
`buildSshConnectionArgs()` connection the launch uses (`src/remote-files.ts`, one
`realpath`+`stat` probe per request returning both the file and the workspace root).
Clicked paths that point OUTSIDE the case directory (a remote `/tmp` scratchpad capture,
a screenshot elsewhere in the remote home) go through the attachment routes, which had
the same local-`fs` assumption: registration, the by-id `raw` stream, the metadata poll
and the attachment history list now resolve over ssh as well, so the click-path works
whether the file sits inside or outside the case. Which host a record is read from
follows the SESSION, never the path string — the same absolute path means a different
file on each host, and a remote session never falls back to a local file.
The guards are unchanged in strength: the workspace boundary is still enforced (now
resolved on the host that can actually resolve it), the sensitive-path blocklist and
the size cap (`CODEMAN_MAX_DOWNLOAD_BYTES`) still apply before any bytes are read, and
`Range` requests keep working, so remote `<video>`/`<audio>` seeking behaves like a
local file. An unreachable host is reported as `502` with the remote reason instead of
a misleading 404. Nothing is ever copied to the Codeman host.
Still not available for remote cases, and now said explicitly instead of 404-ing:
editing a file (`edit=1` / `PUT` answer 400, the viewer hides its Edit affordance),
office-document previews and generated thumbnails (both need the bytes on the server's
disk), the file tree / path picker, and `tail-file`. Docker cases are unaffected (their
workspace is bind-mounted at the same absolute path).
- b357fe8: Add Shift+Left and Shift+Right buttons to the default and extended mobile agent keyboard bars, shown only on Codex sessions, enabling Codex queued-message editing and prompt-stack navigation. Flush locally buffered drafts before navigation and keep terminal focus after taps.
- 9acc5aa: Fix an invisible terminal text selection on the light skins (#360). Every xterm palette declared its selection colour under the key `selection`, which xterm.js renamed to `selectionBackground` in v5. An `ITheme` is a plain object, so the unknown key was dropped without an error and every skin fell back to xterm's own default of `rgba(255,255,255,0.3)`: unnoticeable on the dark skins, which wanted roughly that anyway, and effectively invisible on Paper Gray, Solarized Light, Catppuccin Latte and Rosé Pine Dawn, where white at 30% over a near-white background moves a channel by about 3/255. Selecting text on those skins now highlights it, with desktop drag-select and the mobile long-press both fixed by the same rename.
## 1.28.2
### Patch Changes
- **Terminal font weight** (#417, from discussion #403). App Settings → Terminal → Font gains two
per-device rows, Normal font weight and Bold font weight, each a select from Default plus 100 to 900. Claude Code marks bold with a bare `ESC[1m` and no colour change, so with a family that ships
only a regular and a bold face a bold heading reads as body text; setting normal to 300 turns that
one small step into an obvious one. Both slots resolve against their own xterm default (an unset
bold never inherits normal), apply live to the terminal, both echo overlays and open Agent Teams
panes, and the bundled JetBrains Mono `@font-face` is declared over the font's real 100 to 800 axis
instead of 400 to 700, without which every weight below 400 rendered identically to 400 on a stock
install.
**Phones up to 599px get the phone layout** (#390, fixes #389). The phone tier's cutoff moves
from 430px to 600px in the JS classifier, mobile.css and every test and doc that pins it, so the
iPhone Plus and Pro Max sizes, the Pixel Pro and the Z Fold cover display (430 to 460px) get the
phone header, the Enter key and the accessory bar instead of the tablet layout. Verified on a real
iPhone 17 Pro Max; a Safari page zoom below 100% widens the reported viewport, which is why the
cutoff is 600 rather than 480.
**The plan-usage statusline exporter no longer touches your settings files** (#361, diagnosed in
#405). Codeman used to write its exporter into a workspace's `.claude/settings.local.json`, which
Claude Code ranks above `~/.claude/settings.json`, so it replaced your own statusline for ANY
`claude` run in that directory, including outside Codeman, and rendered the bare word `codeman`
when run by hand. The exporter is now passed to `claude` as an ephemeral `--settings` flag when
Codeman spawns it and is never written to disk; your own statusline (project-local, project, then
`~/.claude/settings.json`) is wrapped and printed through inside Codeman sessions, and a hand-run
`claude` sees nothing of Codeman. Workspaces an older Codeman wrote to self-heal the first time a
session starts there. Telemetry collection follows the Plan Usage chip setting, read fresh at every
Claude session create and respawn; an absent setting means on, and a device writes the switch only
when it flips the chip, so a phone (chip off by default) saving its font size can no longer switch
collection off for the desktop. The exporter prints nothing when it cannot reach Codeman, the
telemetry route answers an unknown session with an empty body, and the footer is empty rather than
a brand word. Known limit: sessions inside a Docker case do not feed the chip yet (the flag rides
local spawns only; the chip is account-wide, so any local Claude session covers it).
**`install.sh` and the Docker agent image read the CLI catalogue** (#380). Adding a CLI to
`src/config/cli-registry/stock.ts` and running `npm run generate:cli-catalog` wires it into the
installer's detection, install menu and closing reminder, and into the agent image's npm layer;
each of those was a separate hand-kept list before, and OMP had been missing from the installer's
detection entirely. The install menu offers every enabled CLI that can drive a pane (eight, rather
than the fixed two), DeepSeek is deliberately withheld because `npm install -g @deepseek-ai/dsh`
installs only a launcher with no runnable profile, a wget-only host keeps the entries that never
needed curl, and the agent image respects `enabled`. The script stays bash 3.2 compatible and CI
now executes it inside a real `bash:3.2` container. Choosing "s" (Skip) in the menu continues to
the clone and build instead of aborting.
**iPhone Duo support** (#407). A visual-viewport resize that changes the WIDTH is the device
changing shape and is never read as the virtual keyboard: closing an iPhone Duo (626 to 466pt wide)
or rotating any phone used to latch the keyboard layout with no keyboard on screen, sticky until the
device was opened again. The seven centred overlays keep their dialogs out of the hinge through the
CSS Viewport Segments variables (inert on devices that do not fold), the phone path picker and
preview stay flush under 600px, and a shape change with the keyboard up baselines to the layout
viewport so the settle event after a rotation no longer closes the keyboard layout. Two Duo device
profiles join the test matrix.
**Codeman is its own Claude Code plugin marketplace.** `/plugin marketplace add Ark0N/Codeman`
followed by `/plugin install codeman@codeman` installs the codeman agent skill as a plugin, from
`plugins/codeman/` (a mirror of `skills/codeman/` kept byte-identical by a test), which is a small
separate directory on purpose: a plugin root carrying a `package.json` gets an `npm install` on
every installer's machine. A Claude Code holding both the plugin and a user-level or per-case copy
lists the skill twice; pick one route.
Housekeeping: the maintainer's Telegram PR bot moved out of this repository (it is a client of the
HTTP API like any other), the COM flow gained a Discussions announcement step, and the changelog's
Thanks sections were backfilled for 1.22.0 to 1.28.1.
### Thanks
- @irisitymichaelgrundberg for the font-weight analysis in #403 that this release implements, and the statusline diagnosis in #405
- @JDProfresh for the phone breakpoint fix (#390)
- @timkjr for moving the statusline exporter off disk (#361)
- @opticon454 for driving the installer and the agent image from the CLI catalogue (#380)
## 1.28.1
### Patch Changes
@@ -28,7 +176,6 @@
### Thanks
1.28.1 is a same-day follow-on to 1.28.0, so the thanks for this pair belong here too:
- **@shenlvkang-collab** for the path picker's typed-path jump and name/date sort (#399), and for the care in the edges: the retry is bounded to one parent level, a typo keeps the listing you had instead of resetting to the root, and a full file path lands in its folder with the entry already selected.
- **@irisitymichaelgrundberg** for Claude truecolor in panes (#409), and above all for flagging the one reading they could not prove: that suppressing truecolor may have made Claude's block collapse into the background rather than fixing anything. That paragraph is why this got measured instead of taken on trust, and the measurement changed the changelog.
- **@timkjr** for trapping Ctrl+Z in agent sessions (#404), for finding that Caps Lock flips `ev.key` to `'Z'` without setting `shiftKey` so a plain `=== 'z'` check misses exactly the keystroke the guard exists for, and for stating up front that an agent CLI already holds its tty with ISIG off rather than overselling the fix.
@@ -209,7 +356,6 @@
### Thanks
1.26.0 carries no contributor PRs of its own. It lands the day after 1.25.0, so the thanks for that pair belong here too:
- @mtiller for the reverse-proxy base URL (#381).
- @dignfei for attaching cases to running containers (#357).
- @shenlvkang-collab for the response viewer fix (#369), the first-hand conversation hook (#367) and the phone Add Case fix (#368).
@@ -311,7 +457,6 @@
### Thanks
1.24.4 is a same-day follow-on to 1.24.3, so the thanks for that pair belong here too:
- @opticon454 for #349, and for a write-up that made an infrastructure PR quick to review
## 1.24.3
@@ -397,7 +542,6 @@
### Thanks
1.24.2 is a hotfix on top of 1.24.1, so the thanks for that pair belong here too:
- @opticon454 for #350, with a reproduction that made this a confirmation rather than a hunt
- @timkjr for reporting #352, and for finding it while verifying Docker support for someone else's PR
@@ -503,7 +647,6 @@
### Thanks
1.23.0 carries no contributor PRs of its own. It lands the day after 1.22.0, so the thanks for that pair belong here too:
- **@aakhter** built both halves of the new tab experience: the owner-scoped, server-authoritative tab-layout foundation with recipient-safe SSE publication and an unusually deep test suite (#335), and the resizable vertical session rail with accessible pointer/keyboard sizing and careful FitAddon handoff (#334). Fifth and sixth merged PRs, and the layout work also fixed real multi-user ordering leaks along the way.
## 1.22.0
@@ -519,7 +662,6 @@
- Fix the file preview's dead pop-out control: a real detach button now opens the previewed file in a browser tab (raw route for PDFs/images/media/text, converted-PDF preview for docx/pptx) and the copy button reports when a preview has no text to copy instead of silently doing nothing. Review-driven hardening for the new tab features: PUT /api/session-order drops unknown ids again instead of rejecting the whole write (a session deleted inside the browser's debounce window could silently lose the user's reorder), a failed mux restore no longer blocks explicit session/webview deletion for the process lifetime (the automated stale sweep stays fail-closed), and the vertical rail gains the axis-awareness the sidebar-only predicates missed: correct drag-reorder insertion, active-tab scroll-into-view, floating windows anchored beside rail tabs, connector redraws on rail scroll, server-seeded orientation applied on first load, a pre-paint stamp so vertical mode no longer flashes through the header strip, and a 12px session-name default matching the sidebar's historical size so untouched installs are not restyled.
### Thanks
- **@aakhter** built both halves of the new tab experience: the owner-scoped, server-authoritative tab-layout foundation with recipient-safe SSE publication and an unusually deep test suite (#335), and the resizable vertical session rail with accessible pointer/keyboard sizing and careful FitAddon handoff (#334). Fifth and sixth merged PRs, and the layout work also fixed real multi-user ordering leaks along the way.
## 1.21.0
+10 -7
View File
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -209,7 +209,7 @@ The most responsive AI coding agent experience on any phone. Full xterm.js termi
<tr><td>Password typing on phone</td><td><b>QR code scan — instant auth</b></td></tr>
</table>
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard; destructive commands require a double-press to confirm, so you never fire one by accident
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard; destructive commands require a double-press to confirm, so you never fire one by accident; on Codex sessions the bar also shows `⇧←` / `⇧→` (Shift+Left / Shift+Right: edit the last queued message / return through the prompt stack)
- **Dedicated Enter button** — replays the keypress through the terminal, so text buffered by local echo is flushed first rather than stranded
- **Swipe navigation & smart keyboard handling** — swipe left/right to switch sessions; toolbar and terminal shift up when the keyboard opens (`visualViewport` API)
- **Built for phones** — safe-area insets for notch and home indicator, 44px touch targets, bottom-sheet case picker, native momentum scrolling
+11
View File
@@ -0,0 +1,11 @@
{
"extends": "../tsconfig.json",
"compilerOptions": {
"rootDir": "..",
"noEmit": true,
"declaration": false,
"declarationMap": false,
"sourceMap": false
},
"include": ["../scripts/test-local-llm-harnesses.ts"]
}
+11 -5
View File
@@ -13,24 +13,24 @@ TZ=Australia/Perth
# Name of the account that runs Codeman and all local CLI sessions. Changing
# this value rebuilds the image with a matching account.
CODEMAN_RUNTIME_USER=opencode
CODEMAN_RUNTIME_USER=codeman
# Required. Persistent Codeman application data, CLI credentials, and session
# state are stored here on the host and mounted at the runtime account's home
# directory in the container.
CODEMAN_APPDATA_PATH=/mnt/user/appdata/Coding/codeman
CODEMAN_APPDATA_PATH=/mnt/user/appdata/codeman
# Optional. Absolute host path of this Codeman checkout, mounted at
# /opt/codeman so App Settings -> Updates can update Codeman in place. The Bash
# start script detects it from the compose file's own location, so it only needs
# setting for direct `docker compose` use or a checkout kept elsewhere. Point it
# at a directory that is not a git checkout and in-app updates are unavailable.
# CODEMAN_REPO_PATH=/mnt/user/appdata/Coding/codeman/app
# CODEMAN_REPO_PATH=/mnt/user/appdata/codeman/app
# Required for Docker cases. This must be an absolute path on the Docker host.
# Codeman and each isolated case use this same path, so it cannot be a
# container-only path such as /home/opencode/codeman-cases.
CODEMAN_CASES_PATH=/mnt/user/appdata/Coding/codeman/codeman-cases
# container-only path such as /home/codeman/codeman-cases.
CODEMAN_CASES_PATH=/mnt/user/appdata/codeman/codeman-cases
# Required. Network bind address, host port, and local image tag.
CODEMAN_HOST=0.0.0.0
@@ -44,6 +44,12 @@ CODEMAN_PASSWORD=changeme
# Required. Username for Codeman HTTP Basic authentication.
CODEMAN_USERNAME=admin
# Optional. Extra Host-header allowlist entries for a reverse-proxied domain
# (comma-separated; a bare `.suffix` matches every subdomain). Without it a
# proxied request is rejected with `403 Forbidden: host not allowed`. See
# README.md, "Reverse-proxy host allowlist".
# CODEMAN_ALLOWED_HOSTS=codeman.example.com,.internal.example.com
# Optional: authenticate Gemini CLI without an interactive login.
GEMINI_API_KEY=
+46 -6
View File
@@ -11,18 +11,21 @@ cp docker/.env.example docker/.env
bash docker/Start-Codeman.sh
```
On PowerShell, use the following command instead.
On PowerShell, use the following commands instead. Running Compose from inside `docker/` with no `-f` lets it discover `docker-compose.override.yml` on its own (see [Local customisation](#local-customisation)); naming the file with `-f docker/docker-compose.yaml` from the repository root silently drops the override unless it is named too.
```powershell
Copy-Item docker/.env.example docker/.env
docker compose --env-file docker/.env -f docker/docker-compose.yaml up --build -d
Set-Location docker
docker compose --env-file .env up --build -d
```
Every required value is defined and explained in `.env.example`. `GEMINI_API_KEY` is intentionally optional and may remain blank.
The container starts as root so `entrypoint.sh` can correct the ownership of a bind source the Docker daemon created (it creates a missing one as `root:root`), then drops to `PUID:PGID` with `setpriv` before the server starts, so Codeman itself never runs privileged. That drop needs `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against the file's `cap_drop: ALL`; a compose file written elsewhere (Unraid's Compose Manager, a hand-written unit) must carry the same additions, and the entrypoint names them when they are missing. A directory owned by neither root nor `PUID:PGID` is never re-owned: it is probed for writability as the runtime account and refused with a clear message if that fails. Setting `user:` in Compose skips the whole step.
On Linux, `Start-Codeman.sh` stops with an error when required paths are missing. It creates the application-data directory when safe, detects its numeric owner as `PUID:PGID`, and detects `DOCKER_SOCKET_GID` from the configured Docker socket. It rejects a root-owned application-data directory because Codeman and its local CLI sessions must remain unprivileged.
Codeman, Claude, OpenCode, and other local sessions run as the unprivileged account named by `CODEMAN_RUNTIME_USER`, which defaults to `opencode`. When Compose is run directly, `PUID` and `PGID` default to `1000:1000`; set them in `.env` when the application-data directory has a different owner. The Bash start script determines them automatically instead.
Codeman, Claude, OpenCode, and other local sessions run as the unprivileged account named by `CODEMAN_RUNTIME_USER`, which defaults to `codeman`. When Compose is run directly, `PUID` and `PGID` default to `1000:1000`; set them in `.env` when the application-data directory has a different owner. The Bash start script determines them automatically instead.
To retain Docker-case support without root when running Compose directly, set `DOCKER_SOCKET_GID` to the numeric group ID of the host socket. On a standard Linux Docker host, obtain it with `stat -c '%g' /var/run/docker.sock`. The Bash start script detects it automatically.
@@ -38,6 +41,43 @@ Releases that change `server.Dockerfile`, `docker-compose.yaml`, or add a key to
changed, and asks you to run `Start-Codeman.sh` here on the host instead. Details:
[`../docs/docker-self-update.md`](../docs/docker-self-update.md).
## Local customisation
Compose merges `docker-compose.override.yml` on top of `docker-compose.yaml`. Keep host-specific changes there rather than editing `docker-compose.yaml`, so this repository can be updated without losing them. Both `docker-compose.override.yml` and `docker-compose.override.yaml` are ignored by Git.
`Start-Codeman.sh` names the Compose file explicitly, which disables Compose's automatic discovery of the override file, so the script adds it back when one is present and prints the file it used. Running `docker compose` from this folder without any `-f` option finds it automatically. When passing `-f docker/docker-compose.yaml` from the repository root, add `-f docker/docker-compose.override.yml` as well, or the override is silently ignored.
An override file adds to and replaces individual settings. It cannot delete a key from `docker-compose.yaml`, and Compose concatenates rather than replaces `ports`, so removing a published port still requires editing `docker-compose.yaml`. The example below replaces the restart policy and adds a mount, leaving every other setting in place:
```yaml
services:
codeman:
restart: always
volumes:
- /srv/projects:/srv/projects
```
### Reverse-proxy host allowlist
Codeman rejects any request whose `Host` header is not on its own allowlist - a
DNS-rebinding guard, not a Compose or Docker concern. Loopback, any IP literal,
the configured `--host`, and a few tunnel-provider suffixes are allowed by
default; a reverse-proxied domain is not, and is rejected with
`403 Forbidden: host not allowed` before the request reaches any handler.
Add the domain with `CODEMAN_ALLOWED_HOSTS` in `.env`:
```sh
CODEMAN_ALLOWED_HOSTS='codeman.example.com,.internal.example.com'
```
`docker-compose.yaml` forwards it into the container (Compose only passes
through the environment keys it explicitly lists, and this is one of them, with
an empty default so the line is optional in `.env`).
See the application's own `docs/wiki/Remote-Access.md` for the full allowlist
format and the tunnel providers it accepts by default.
## Application data storage
The default configuration uses a host-folder bind mount:
@@ -49,7 +89,7 @@ volumes:
target: /home/${CODEMAN_RUNTIME_USER}
```
Set `CODEMAN_APPDATA_PATH` in `.env` to a directory that the Docker daemon can access. The example value is `/mnt/user/appdata/Coding/codeman`.
Set `CODEMAN_APPDATA_PATH` in `.env` to a directory that the Docker daemon can access. The example value is `/mnt/user/appdata/codeman`.
`CODEMAN_CASES_PATH` is the separate host directory for managed case workspaces. It is mounted into Codeman at the same absolute path, allowing the host Docker daemon to bind it into an isolated case container. Set it to a child directory of `CODEMAN_APPDATA_PATH` unless you deliberately store workspaces elsewhere.
@@ -60,7 +100,7 @@ Set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` when `docker info` reports `SwapLimit=
For an existing installation created by a root-running image, change ownership of the application-data directory before upgrading so the configured `PUID` and `PGID` can read the saved credentials and state:
```sh
chown -R 99:100 /mnt/user/appdata/Coding/codeman
chown -R 99:100 /mnt/user/appdata/codeman
```
Replace `99:100` and the path with the values from your `.env` file.
@@ -69,7 +109,7 @@ Do not replace this bind mount with a Docker-managed named volume when Docker ca
## Static macvlan networking
The default configuration publishes a host port. It does not use `network_mode: host`. To attach Codeman directly to an existing external macvlan network with a static IP address and MAC address, remove the `ports:` section and add the following to the `codeman` service:
The default configuration publishes a host port. It does not use `network_mode: host`. To attach Codeman directly to an existing external macvlan network with a static IP address and MAC address, remove the `ports:` section from `docker-compose.yaml` and add the following to the `codeman` service. The service and network additions can instead be placed in `docker-compose.override.yml`, but the `ports:` removal cannot, as described under [Local customisation](#local-customisation):
```yaml
mac_address: ${CODEMAN_MAC_ADDRESS}
+187 -7
View File
@@ -12,11 +12,34 @@ if [[ ! -f "$env_file" ]]; then
exit 1
fi
compose_command=(docker compose --env-file "$env_file" -f "$compose_file")
# Naming a Compose file explicitly disables Compose's automatic discovery of
# the override file, so it has to be added back by hand. Without this, local
# customisation in docker-compose.override.yml is silently ignored. The
# candidates are checked in Compose's own precedence order - measured on
# Compose v5.5.0 with both present: it uses `.yml` and ignores `.yaml`.
override_yml="$script_dir/docker-compose.override.yml"
override_yaml="$script_dir/docker-compose.override.yaml"
if [[ -f "$override_yml" && -f "$override_yaml" ]]; then
printf 'Warning: both %s and %s exist; Compose uses .yml and ignores .yaml.\n' \
"$override_yml" "$override_yaml" >&2
fi
compose_files=(-f "$compose_file")
for override_file in "$override_yml" "$override_yaml"; do
if [[ -f "$override_file" ]]; then
compose_files+=(-f "$override_file")
printf 'Using Compose override file: %s\n' "$override_file"
break
fi
done
compose_command=(docker compose --env-file "$env_file" "${compose_files[@]}")
appdata_path=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "CODEMAN_APPDATA_PATH" { sub(/^[^=]*=/, ""); print; exit }'
)
cases_path=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "CODEMAN_CASES_PATH" { sub(/^[^=]*=/, ""); print; exit }'
)
docker_socket=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "DOCKER_SOCKET" { sub(/^[^=]*=/, ""); print; exit }'
@@ -36,11 +59,18 @@ if [[ ! -d "$appdata_path" ]]; then
mkdir -p -- "$appdata_path"
fi
if owner_ids=$(stat -c '%u:%g' -- "$appdata_path" 2>/dev/null); then
:
elif owner_ids=$(stat -f '%u:%g' "$appdata_path" 2>/dev/null); then
:
else
if [[ -z "$cases_path" ]]; then
printf 'Error: CODEMAN_CASES_PATH is not set in %s\n' "$env_file" >&2
exit 1
fi
# `stat -c` is GNU, `stat -f` is BSD/macOS; the bind sources live on the Docker
# host, so both need to work.
owner_of() {
stat -c '%u:%g' -- "$1" 2>/dev/null || stat -f '%u:%g' "$1" 2>/dev/null
}
if ! owner_ids=$(owner_of "$appdata_path"); then
printf 'Error: Cannot determine the owner of CODEMAN_APPDATA_PATH: %s\n' "$appdata_path" >&2
exit 1
fi
@@ -54,6 +84,34 @@ if [[ "$PUID" == '0' ]]; then
exit 1
fi
# Pre-creating this here, exactly like CODEMAN_APPDATA_PATH above, means Compose
# never has to materialise a missing bind source itself - which it does as
# root:root - so the in-container entrypoint's chown never has to run for this
# path at all. It happens AFTER PUID/PGID are known (they come from the appdata
# directory just above) so the new directory can be given that exact owner: a
# plain `mkdir -p` lands as the invoking user's uid and PRIMARY gid, and on a
# host set up the way the README suggests (`chown -R 99:100 <appdata>`) that gid
# is not PGID, which the container would then refuse to run on. Unlike appdata,
# an EXISTING cases directory is left exactly as it is: the README explicitly
# allows pointing this at a normal projects directory the host account already
# owns, and the container checks that it is WRITABLE as PUID:PGID rather than
# who owns it.
if [[ ! -d "$cases_path" ]]; then
mkdir -p -- "$cases_path"
if [[ "$(owner_of "$cases_path")" != "$PUID:$PGID" ]]; then
# As root this always succeeds; as a member of PGID a chgrp does; anyone
# else gets the clear error here, where the fix is obvious, rather than a
# restart loop from the container.
if ! chown -- "$PUID:$PGID" "$cases_path" 2>/dev/null; then
printf 'Error: created CODEMAN_CASES_PATH (%s) but could not make it %s:%s (the owner of CODEMAN_APPDATA_PATH).\n' \
"$cases_path" "$PUID" "$PGID" >&2
printf 'Run `chown %s:%s %s` as root, or create the directory as that account, then retry.\n' \
"$PUID" "$PGID" "$cases_path" >&2
exit 1
fi
fi
fi
if [[ -z "$docker_socket" || ! -S "$docker_socket" ]]; then
printf 'Error: DOCKER_SOCKET is not a Unix socket: %s\n' "${docker_socket:-<unset>}" >&2
exit 1
@@ -92,6 +150,31 @@ if [[ ! -d "$repo_path/.git" ]]; then
printf 'Note: %s is not a git checkout, so in-app updates are unavailable.\n' "$repo_path" >&2
fi
# Reads HEAD without requiring a `git` binary on the host — this script
# otherwise checks the checkout only by testing for `.git` as a directory, and
# resolving refs by hand keeps that the same "no host git needed" guarantee.
# ⚠️ A worktree checkout has `.git` as a FILE (`gitdir: <path>`), not a
# directory, so this returns nothing there and the volume-refresh check below
# silently no-ops — consistent with the `-d .git` test used everywhere else in
# this script, not a special case, but worth knowing if a worktree checkout
# stops picking up a stale-volume refresh it should have caught.
git_head_commit() {
local git_dir="$1/.git" head_ref ref_path
[[ -d "$git_dir" ]] || return 1
head_ref=$(cat -- "$git_dir/HEAD" 2>/dev/null) || return 1
if [[ "$head_ref" == ref:* ]]; then
ref_path="${head_ref#ref: }"
if [[ -f "$git_dir/$ref_path" ]]; then
cat -- "$git_dir/$ref_path"
else
# Packed after a `git gc`; the loose ref file above is gone.
awk -v ref="$ref_path" '$2 == ref { print $1; exit }' "$git_dir/packed-refs" 2>/dev/null
fi
else
printf '%s' "$head_ref"
fi
}
# Record what the container is about to be built and created FROM. The in-app
# updater compares these against the release it wants to apply: a release that
# changes either file cannot be applied by the container restarting itself (a
@@ -126,4 +209,101 @@ else
printf 'Warning: no sha256 tool found; in-app updates will not detect environment changes.\n' >&2
fi
exec docker compose --env-file "$env_file" -f "$compose_file" up --build -d
# codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded from
# the image only while EMPTY, so a rebuilt image's fresh output sits unused
# behind old volume content until something clears it. The in-app self-updater
# never hits this — it rebuilds INSIDE the running container, into the very
# volume already in use — but a `docker compose build` triggered from outside
# it (this script, after a `git pull`) does: the container comes back up
# looking unchanged. Detect that here and clear just the affected volume(s) so
# the build below actually takes effect. Best-effort: with no sha256 tool this
# quietly does nothing, same as the environment-gate block above.
volumes_to_refresh=()
if [[ -n "$dockerfile_sha" ]]; then
repo_head=$(git_head_commit "$repo_path" || true)
lockfile_sha=$(sha256_of "$repo_path/package-lock.json" 2>/dev/null || true)
source_state_file="$state_dir/docker-build-source.json"
prev_head=''
prev_lockfile_sha=''
if [[ -f "$source_state_file" ]]; then
prev_head=$(sed -n 's/.*"headCommit": *"\([^"]*\)".*/\1/p' "$source_state_file")
prev_lockfile_sha=$(sed -n 's/.*"lockfileSha256": *"\([^"]*\)".*/\1/p' "$source_state_file")
fi
[[ -n "$repo_head" && "$repo_head" != "$prev_head" ]] && volumes_to_refresh+=('codeman-dist')
[[ -n "$lockfile_sha" && "$lockfile_sha" != "$prev_lockfile_sha" ]] && volumes_to_refresh+=('codeman-node-modules')
fi
if [[ ${#volumes_to_refresh[@]} -eq 0 ]]; then
exec "${compose_command[@]}" up --build -d
fi
# Runs even on this script's very first invocation against an EXISTING
# deployment, deliberately: that deployment's volumes may already be stale
# (there was no earlier version of this check to have caught it), and clearing
# an already-empty or nonexistent volume is a harmless no-op, so there is no
# fresh-install case this needs to avoid.
printf 'Source changed since the last start; refreshing: %s\n' "${volumes_to_refresh[*]}"
# Build BEFORE taking the stack down: the image build is the slow part and needs
# no container stopped, so the deployment is offline only for the recreate.
"${compose_command[@]}" build
# `com.docker.compose.volume` is the volume KEY, not a project-qualified name -
# a second stack on the same host (a beta instance started with a different
# COMPOSE_PROJECT_NAME, say) that also declares a volume keyed `codeman-dist`
# shares that label, and `head -n1` would pick whichever the daemon happens to
# list first. Scope the lookup to THIS stack's own resolved project name so it
# can only ever match this stack's volume. The name is read from the resolved
# config's top-level `name` key, indentation-agnostic (the formatting is not a
# contract), and the FIRST `name` in the output is the project's: nested ones
# (a network's `name:`) come later. `--format json` needs Compose v2.3+.
project_name=$(
"${compose_command[@]}" config --format json 2>/dev/null |
sed -n 's/^[[:space:]]*"name":[[:space:]]*"\([^"]*\)".*$/\1/p' | head -n1
)
"${compose_command[@]}" down
# Track whether the volumes were actually cleared. The marker below is written
# ONLY on success: with an unresolvable project name the label filter would
# match nothing, nothing would be removed, and a marker recording the new HEAD
# would stop this check from ever firing again while the stale volume kept
# serving old code. A failed removal likewise leaves the marker alone, so the
# next start retries, and the stack is brought back up regardless rather than
# left down.
refreshed=1
if [[ -z "$project_name" ]]; then
# The documented reset (docs/docker-self-update.md): both volumes re-seed from
# the image by a plain copy, so clearing the extra one costs a copy, not data.
printf 'Warning: could not resolve the Compose project name; clearing both build-artefact volumes with `down --volumes` instead.\n' >&2
"${compose_command[@]}" down --volumes || refreshed=0
else
for key in "${volumes_to_refresh[@]}"; do
volume_name=$(
docker volume ls -q \
--filter "label=com.docker.compose.volume=$key" \
--filter "label=com.docker.compose.project=$project_name" |
head -n1
)
if [[ -n "$volume_name" ]] && ! docker volume rm -- "$volume_name"; then
printf 'Warning: could not remove volume %s; it will be retried on the next start.\n' "$volume_name" >&2
refreshed=0
fi
done
fi
if [[ "$refreshed" == '1' ]]; then
printf '{\n "headCommit": "%s",\n "lockfileSha256": "%s"\n}\n' \
"$repo_head" "$lockfile_sha" >"$source_state_file.tmp"
mv -- "$source_state_file.tmp" "$source_state_file"
if [[ "$EUID" == '0' ]]; then
chown -- "$PUID:$PGID" "$source_state_file"
fi
else
printf 'Warning: the build-artefact volumes were NOT refreshed; the container may serve stale code until the next successful start.\n' >&2
fi
# Already built above, so no --build here: a second build would only re-check
# the cache.
exec "${compose_command[@]}" up -d
+21
View File
@@ -32,6 +32,10 @@ services:
CODEMAN_DOCKER_HOST_HOME: ${CODEMAN_APPDATA_PATH}
CODEMAN_DOCKER_DISABLE_SWAP_LIMIT: ${CODEMAN_DOCKER_DISABLE_SWAP_LIMIT}
CODEMAN_CASES_PATH: ${CODEMAN_CASES_PATH}
# Extra Host-header allowlist entries for a reverse-proxied deployment
# (docker/README.md, "Reverse-proxy host allowlist"). Optional, so it
# defaults to empty rather than requiring a line in every .env.
CODEMAN_ALLOWED_HOSTS: ${CODEMAN_ALLOWED_HOSTS:-}
CODEMAN_HOST: ${CODEMAN_HOST}
CODEMAN_PASSWORD: ${CODEMAN_PASSWORD}
CODEMAN_PORT: ${CODEMAN_PORT}
@@ -91,6 +95,23 @@ services:
- no-new-privileges:true
cap_drop:
- ALL
cap_add:
# The entrypoint corrects bind-mount ownership as root before dropping to
# PUID:PGID. Everything not listed here remains dropped by cap_drop above.
# test/docker-entrypoint.test.ts pins this list against what the
# entrypoint and `init: true` actually need, so a capability cannot go
# missing silently again.
- CHOWN
- DAC_OVERRIDE
# `init: true` makes tini PID 1, and tini stays ROOT while the entrypoint
# drops the server to PUID. Signalling a process of a different uid needs
# CAP_KILL; without it tini's SIGTERM forward fails ("Unexpected error
# when forwarding signal: 'Operation not permitted'"), tini dies, and the
# PID namespace teardown SIGKILLs the server instead of letting
# `server.stop()` flush state on every `docker compose down`/`restart`.
- KILL
- SETGID
- SETUID
healthcheck:
test:
- CMD-SHELL
+165
View File
@@ -0,0 +1,165 @@
#!/bin/sh
# Corrects ownership - host bind mounts, and the image-baked CLI prefix -
# then drops to PUID:PGID.
#
# Compose binds CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH from the host. When
# either path does not exist yet - a first run, a cleared application-data
# directory, a restored backup - the Docker daemon creates it owned by root,
# and an unprivileged server cannot then create its own state directory. The
# result is a container that restarts forever on:
#
# Failed to start web server: EACCES: permission denied, mkdir '/home/<user>/.codeman'
#
# Running this as root and dropping afterwards removes that failure mode without
# leaving the server privileged. The same root start also lets it re-assert
# /opt/codeman-cli's ownership on every start, not just at image build time -
# see the comment at that chown below for why that matters for anyone who
# runs the compose file directly rather than through Start-Codeman.sh.
#
# Capabilities this script needs against the compose file's `cap_drop: ALL`
# (test/docker-entrypoint.test.ts pins the list against docker-compose.yaml):
# CHOWN + DAC_OVERRIDE the chown of a root-owned bind source below
# SETUID + SETGID the setpriv drop itself
# KILL NOT used here, but required by the container: with
# `init: true` tini is PID 1 and runs as root while the
# server runs as PUID, and signalling a process of a
# different uid needs CAP_KILL. Without it every
# `docker compose down`/`restart` ends in tini dying with
# "Unexpected error when forwarding signal" and the
# server being SIGKILLed instead of stopping cleanly.
set -eu
# Honour an explicit `user:` in Compose: when the container was not started as
# root there is nothing to correct and no privilege to drop.
if [ "$(id -u)" -ne 0 ]; then
exec "$@"
fi
# Everything below runs as root and calls stat, chown, id, setpriv and friends
# by bare name, so the lookup path must not contain a directory the runtime
# account can write to. /opt/codeman-cli/bin is exactly that (it is chowned to
# PUID:PGID so sessions can update the agent CLIs in place), and the image
# appends it to PATH for the server's sake. Resolve root's commands through the
# system directories only, and hand the image's full PATH back to the server at
# the exec below, since Codeman resolves the agent CLIs through it.
runtime_path=$PATH
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
export PATH
: "${PUID:=1000}"
: "${PGID:=1000}"
# The capabilities the compose file must grant, named in the diagnosis below so
# an out-of-tree compose file (Unraid's Compose Manager, a hand-written unit)
# fails with a one-line fix instead of a restart loop.
required_caps='CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID'
# Pre-flight the drop itself before touching anything. A container started with
# `cap_drop: ALL` and none of the additions above fails here, and would otherwise
# die at the final exec with a bare "setpriv: setresuid failed: Operation not
# permitted" after chown had already failed, or worse, misreport a perfectly
# writable directory as unwritable because the probe below could not drop
# privileges to test it.
if ! setpriv --reuid "$PUID" --regid "$PGID" --clear-groups true 2>/dev/null; then
printf 'entrypoint: cannot drop privileges to PUID:PGID (%s:%s).\n' "$PUID" "$PGID" >&2
printf 'entrypoint: this image starts as root and drops with setpriv, which needs\n' >&2
printf 'entrypoint: cap_add: [%s]\n' "$required_caps" >&2
printf 'entrypoint: on top of cap_drop: ALL (see docker/docker-compose.yaml). Add them to the\n' >&2
printf 'entrypoint: compose file that started this container, or set `user:` to skip the drop entirely.\n' >&2
exit 1
fi
# Preserve the supplementary groups Compose granted through group_add - that is
# how the Docker socket stays reachable - while discarding root's own group.
supplementary=$(id -G | tr ' ' '\n' | grep -vx 0 | paste -sd, -)
[ -n "$supplementary" ] || supplementary="$PGID"
# Writable as the account the server is about to become? A real probe, run as
# exactly the identity the final exec below produces (PUID, PGID, the same
# supplementary groups, capabilities dropped), rather than a comparison of
# owners: ownership is not writability. A group-writable tree owned by another
# account, an ACL, or a CIFS/NFS mount that reports some unrelated uid are all
# fine to run on and would all fail an owner check.
writable_as_runtime() {
setpriv --reuid "$PUID" --regid "$PGID" --groups "$supplementary" test -w "$1" 2>/dev/null
}
for target in "${HOME:-}" "${CODEMAN_CASES_PATH:-}"; do
[ -n "$target" ] && [ -d "$target" ] || continue
owner=$(stat -c '%u:%g' "$target")
[ "$owner" = "${PUID}:${PGID}" ] && continue
# Only ever correct a directory the DAEMON created: root-owned, because
# neither PUID nor PGID existed yet when it materialised the missing bind
# source. Anything else - a host tree that legitimately belongs to some
# OTHER account, such as an existing CODEMAN_CASES_PATH the README already
# allows pointing at a normal project directory - is not this container's
# to reassign; recursively chowning it on every mismatch silently rewrote
# a credentials tree or a projects directory to PUID:PGID with one log
# line to explain it. Such a directory is left alone and only PROBED below.
#
# The chown is deliberately not fatal. A bind mount backed by NFS, CIFS or a
# rootless daemon can refuse chown while still being perfectly writable, and
# the probe below is what decides whether the server can run on it.
if [ "${owner%%:*}" = '0' ]; then
if chown -R "${PUID}:${PGID}" "$target" 2>/dev/null; then
printf 'entrypoint: corrected ownership of %s to %s:%s\n' "$target" "$PUID" "$PGID"
else
printf 'entrypoint: warning: cannot change ownership of %s to %s:%s; checking whether it is writable anyway\n' \
"$target" "$PUID" "$PGID" >&2
fi
fi
if writable_as_runtime "$target"; then
if [ "${owner%%:*}" != '0' ]; then
printf 'entrypoint: %s is owned by %s, not %s:%s, but is writable as the runtime account; leaving its ownership alone\n' \
"$target" "$owner" "$PUID" "$PGID"
fi
continue
fi
printf 'entrypoint: %s is not writable as PUID:PGID (%s:%s); it is owned by %s.\n' \
"$target" "$PUID" "$PGID" "$owner" >&2
printf 'entrypoint: refusing to change ownership of a directory this container did not create.\n' >&2
printf 'entrypoint: either chown it on the host, make it writable to %s:%s, or set PUID/PGID to match its owner.\n' \
"$PUID" "$PGID" >&2
exit 1
done
# /opt/codeman-cli (the four agent CLIs) is chowned to PUID:PGID once, at
# image BUILD time, from the PUID/PGID build args - server.Dockerfile's own
# comment on that RUN step explains why it lives in its own prefix rather than
# /usr/local. Unlike HOME/CODEMAN_CASES_PATH above, that bake happens only
# when the image is actually rebuilt (`docker compose up --build`, which
# Start-Codeman.sh always does) - a deployment that instead runs the compose
# file directly (Unraid's Compose Manager, a native Debian systemd unit, any
# `docker compose up`/`restart` with no --build) can change PUID/PGID in .env
# and restart without ever rebuilding, at which point the container runs as
# the NEW uid while the CLI directory is still owned by the OLD one baked into
# the image layer - silently breaking the very "self-update a CLI in place"
# fix this directory exists for. Re-assert it here, every start, unconditionally:
# unlike the host bind mounts above, this is pure image content Codeman itself
# populated, never host data that might legitimately belong to someone else,
# so there is no ownership to be careful about - it is always correct for it
# to be owned by whoever this container is about to run as.
if [ -d /opt/codeman-cli ] && [ "$(stat -c '%u:%g' /opt/codeman-cli)" != "${PUID}:${PGID}" ]; then
chown -R "${PUID}:${PGID}" /opt/codeman-cli
fi
# Discarding group 0 is right for root's own group, but it also discards a
# `group_add: 0` that was there to reach a Docker socket owned by root:root.
# The previous image ran as PUID with that group kept, so say so rather than
# letting Docker-case support vanish silently on such a host.
if [ -S /var/run/docker.sock ] && [ "$(stat -c '%g' /var/run/docker.sock)" = '0' ]; then
printf 'entrypoint: warning: /var/run/docker.sock is owned by group 0, which is dropped along with root;\n' >&2
printf 'entrypoint: warning: Docker cases will not work from this container. Give the socket a dedicated\n' >&2
printf 'entrypoint: warning: group on the host and set DOCKER_SOCKET_GID to it.\n' >&2
fi
# No `--bounding-set -all` here: it is a silent no-op without CAP_SETPCAP, which
# the compose file deliberately does not grant, and `no-new-privileges` already
# makes the bounding set moot. The reuid/regid drop leaves CapPrm/CapEff empty.
# The image's full PATH goes back to the server here; see the top of the file.
exec setpriv --reuid "$PUID" --regid "$PGID" --groups "$supplementary" \
env PATH="$runtime_path" "$@"
+47 -3
View File
@@ -24,7 +24,7 @@ RUN npm ci \
# docker/docker-compose.yaml. It does not run a Docker daemon in this container.
FROM node:22-bookworm-slim
ARG CODEMAN_RUNTIME_USER=opencode
ARG CODEMAN_RUNTIME_USER=codeman
ARG PUID=1000
ARG PGID=1000
@@ -71,6 +71,24 @@ COPY --from=docker:29-cli \
# Keep credentials out of the image. Users authenticate these CLIs at runtime
# through Codeman sessions, and the configured host bind mount retains state.
#
# Installed into a DEDICATED prefix, /opt/codeman-cli, not the base image's
# default /usr/local. A session needs write access to wherever these CLIs live
# so it can self-update one in place (observed via Codex's own
# `npm install -g @openai/codex`, which renames the old package directory
# aside before installing the new one — a rename needs write access to the
# PARENT directory, not just the target, so the runtime account needs that
# access at the directory level). Chowning /usr/local/bin and
# /usr/local/lib/node_modules directly to get it would ALSO hand away
# entrypoint.sh (COPY'd to /usr/local/bin below, root-owned, executed as root
# on every container start with CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node
# binary: owning the DIRECTORY is enough to rename it aside and drop a
# replacement, even though the file itself stays root-owned, which would let a
# compromised session arrange for its own script to run as root at the next
# restart — undoing the "the server itself never runs privileged" guarantee
# the entrypoint exists to provide. /opt/codeman-cli holds nothing else to
# escalate through, so owning it is exactly the CLI-update access it needs and
# no more.
#
# ⚠️ PINNED ON PURPOSE. Unpinned, the agent CLI versions a user ends up with are
# a function of WHEN their image was built, not of any commit — so a Codeman
# release that depends on newer CLI behaviour (the trust-dialog handling is
@@ -82,6 +100,15 @@ COPY --from=docker:29-cli \
#
# Bump these deliberately, in a release. `--no-cache` is still needed to rebuild
# this layer when only the pins change upstream.
# The prefix is APPENDED to PATH, never prepended: it is chowned to the runtime
# account below, and entrypoint.sh runs as root calling stat/chown/setpriv by
# bare name. A prefix ahead of /usr/bin would let a session drop a `setpriv`
# there and have it run as root at the next container start (measured with a
# minimal image of this exact shape). The four CLIs live only in this prefix,
# so they still resolve; entrypoint.sh additionally pins its own PATH to the
# system directories for the root part of the start.
ENV NPM_CONFIG_PREFIX=/opt/codeman-cli
ENV PATH=$PATH:/opt/codeman-cli/bin
RUN npm install --global \
@anthropic-ai/claude-code@2.1.258 \
@google/gemini-cli@0.58.0 \
@@ -93,6 +120,11 @@ RUN npm install --global \
# PGID match the host-owned application-data directory mounted by Compose. The
# requested GID may not exist in the base image, and a host UID such as 1000 may
# already belong to the baked `node` account, so handle both cases explicitly.
#
# The trailing chown hands the CLI prefix (/opt/codeman-cli, populated above)
# to that same account, so a session can self-update one of the CLIs in place.
# /usr/local stays root-owned throughout — see the comment on the npm install
# above for why that boundary matters.
RUN set -eux; \
case "${PUID}" in ''|*[!0-9]*) echo "PUID must be numeric" >&2; exit 1;; esac; \
case "${PGID}" in ''|*[!0-9]*) echo "PGID must be numeric" >&2; exit 1;; esac; \
@@ -120,7 +152,8 @@ RUN set -eux; \
--home-dir "/home/${CODEMAN_RUNTIME_USER}" \
--shell /bin/bash \
"${CODEMAN_RUNTIME_USER}"; \
fi
fi; \
chown -R "${PUID}:${PGID}" /opt/codeman-cli
WORKDIR /opt/codeman
@@ -135,8 +168,19 @@ ENV CODEMAN_IN_CONTAINER=1 \
HOME=/home/${CODEMAN_RUNTIME_USER} \
NODE_ENV=production
# Runtime defaults for the entrypoint, matching the account created above.
ENV PGID=${PGID} PUID=${PUID}
EXPOSE 3000
USER ${CODEMAN_RUNTIME_USER}
# The container starts as root so the entrypoint can correct the ownership of
# the host bind mounts, which the daemon creates as root whenever they do not
# already exist. The entrypoint then drops to PUID:PGID with setpriv, so the
# server itself never runs privileged. Setting `user:` in Compose bypasses both
# steps, leaving the caller in full control.
COPY docker/entrypoint.sh /usr/local/bin/entrypoint.sh
RUN chmod 0755 /usr/local/bin/entrypoint.sh
ENTRYPOINT ["/usr/local/bin/entrypoint.sh"]
CMD ["node", "dist/index.js", "web"]
+3 -3
View File
@@ -54,7 +54,7 @@ Model is NOT a session field: it is a composition entry in the profile's config
### Remote SSH cases
**Remote SSH cases** (COD-94/#145): cases can point at a **remote host** (`~/.codeman/remote-hosts.json` + `remote-cases.json` via `src/remote-hosts.ts`; CRUD under `/api/cases` — cases route file). A remote session launches a LOCAL tmux pane running `ssh <host>` that creates a durable REMOTE tmux session on a **dedicated socket** `-L codeman-remote` with name `codeman-ssh-<id>` — deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so a Codeman instance on the target host never adopts it; no `-g` global tmux options are set remotely. `remotePath`/`identityFile` are schema-guarded against shell injection (backticks/`$` rejected — same approach as `extraSshOptions`); remote tmux availability is probed via `checkRemoteTmuxAvailable()` in quick-start (ssh args carry `-o ConnectTimeout=10`). Remote claude defaults to an idempotent `claude --session-id <id> || claude --resume <id>` pair under a login shell, so a respawn or reattach continues the SAME conversation rather than starting a fresh one (remote omp gets the same treatment via `--continue`; ⚠️ because the claude arm is an `a || b` pair under `-c`, that pane's PID is the login shell, not the agent); per-host `commands.*` override. Session kill best-effort kills the remote tmux too. `SessionState.remote`/`MuxSession.remote` round-trip through recovery (`restoreMuxSessions` passes `remote` back into the Session constructor). ⚠️ Run flows must route remote cases through `POST /api/quick-start` (which resolves the remote case and skips LOCAL CLI availability gates) — `POST /api/sessions` stat-validates `workingDir` locally and has no `caseName`. `envOverrides`/`effort`/`modelOverride`/`codexConfig`/`geminiConfig` are rejected for remote quick-starts (not silently dropped). UI: Create Case modal → Remote tab. Tests: `test/remote-hosts.test.ts`, `test/remote-ssh-options.test.ts`.
**Remote SSH cases** (COD-94/#145): cases can point at a **remote host** (`~/.codeman/remote-hosts.json` + `remote-cases.json` via `src/remote-hosts.ts`; CRUD under `/api/cases` — cases route file). A remote session launches a LOCAL tmux pane running `ssh <host>` that creates a durable REMOTE tmux session on a **dedicated socket** `-L codeman-remote` with name `codeman-ssh-<id>` — deliberately failing the remote Codeman's `SAFE_MUX_NAME_PATTERN` so a Codeman instance on the target host never adopts it; no `-g` global tmux options are set remotely. `remotePath`/`identityFile` are schema-guarded against shell injection (backticks/`$` rejected — same approach as `extraSshOptions`); remote tmux availability is probed via `checkRemoteTmuxAvailable()` in quick-start (ssh args carry `-o ConnectTimeout=10`). Remote claude defaults to an idempotent `claude --session-id <id> || claude --resume <id>` pair under a login shell, so a respawn or reattach continues the SAME conversation rather than starting a fresh one (remote omp gets the same treatment via `--continue`; ⚠️ because the claude arm is an `a || b` pair under `-c`, that pane's PID is the login shell, not the agent); per-host `commands.*` override. Session kill best-effort kills the remote tmux too. `SessionState.remote`/`MuxSession.remote` round-trip through recovery (`restoreMuxSessions` passes `remote` back into the Session constructor). ⚠️ Run flows must route remote cases through `POST /api/quick-start` (which resolves the remote case and skips LOCAL CLI availability gates) — `POST /api/sessions` stat-validates `workingDir` locally and has no `caseName`. `envOverrides`/`effort`/`modelOverride`/`codexConfig`/`geminiConfig` are rejected for remote quick-starts (not silently dropped). UI: Create Case modal → Remote tab. Tests: `test/remote-hosts.test.ts`, `test/remote-ssh-options.test.ts`. ⚠️ **Reading a file in a remote case goes over ssh too** (#415): `src/remote-files.ts` is the single remote-READ layer (`buildRemoteFileCommand` = `buildSshConnectionArgs` + one shellescaped remote command; `remoteProbePaths` returns remote realpath + stat; `remoteCreateReadStream` streams a `Range` via `tail -c +N | head -c L` and its `close()` must be wired to the response's `close` or the ssh child outlives an aborted download). The guard order matches the local path exactly (`validateSessionFilePathLexical` → remote realpath of BOTH file and workspace root → containment → sensitive-path → size cap on the REMOTE size), a request path arrives from the browser and is only ever interpolated as a `shellescape`d token, and an unreachable host answers **502**, never a 404. ⚠️ The probe's symlink resolution FAILS CLOSED: `readlink -f` where it exists, otherwise a `cd -P`/`pwd -P` directory walk plus a bounded plain-`readlink` loop over the last component, and anything it cannot fully resolve is reported unresolvable (404), never as the unresolved string — the first version resolved the directory chain only, so on a host without `readlink -f` a `ws/notes.txt -> ~/.ssh/id_rsa` link passed containment under its own path while `cat` served the key. Records are NUL-separated and index-keyed so a newline in a filename cannot shift the mapping. ⚠️ ssh children are BOUNDED: probes and buffered reads go through `src/remote-ssh-limiter.ts` (a `document-conversion-limiter`-shaped semaphore, default 4), the attachment-history list probes its whole history in ONE batched call (`probeRemoteAttachmentHistory`, threaded into `registerExternalAttachment({remoteProbes})`), and probes chunk at 40 paths — a prompt-injected agent printing `codeman://attach` links in a remote session used to fork one `ssh` per link. `describeExecError` never returns Node's `Command failed: <ssh line>` message (identity path + probe script in a 502 body). The `PUT /file-content` guard sits AHEAD of `validateSessionFilePath`, which resolves LOCALLY, or a same-named local directory (an sshfs mount) takes the write. Under `VITEST` the three IO functions refuse rather than connect. This covers the ATTACHMENT routes too, which is the half a clicked path needs when the file is OUTSIDE the case directory (`_isExternalPreviewPath` sends it to `POST …/attachments`): registration, by-id `raw`, metadata and the history list all resolve over ssh (`registerExternalAttachment({remote})`, `resolveServableRemoteAttachment`), and what decides the host is the SESSION, never the path string — the same absolute path means a different file on each host. Deliberately NOT supported over ssh: writes (`edit=1`/`PUT` answer 400, `editable` is always false), office previews/thumbnails, the file tree/picker, `tail-file`. Tests: `test/remote-files.test.ts`, `test/routes/file-routes-remote.test.ts`.
### Docker cases
@@ -66,7 +66,7 @@ Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/do
### Input delivery and WS resilience
**Input**: `session.writeViaMux()` for programmatic/curl input — tmux `send-keys -l` (literal) + `send-keys Enter`. Single-line only (fire-and-once). Interactive **browser** input goes through a durable **exactly-once** layer: each frame carries a stable `clientId` + monotonic per-session `seq`, persisted to localStorage until the server ACKs (`{t:'ia',seq}` over WS, or HTTP 2xx), so a dropped link/reconnect can't lose or double-deliver a prompt. **WS resilience** (#149): the upgrade URL carries `cid = clientId + ':' + perTabNonce`, and `ws-connection-registry.ts` supersedes only same-TAB reconnects (two tabs on one session coexist; input frames keep the bare `clientId` for seq dedup); reconnects back off exponentially (attempts preserved across `_connectWs`), and the header connection chip renders from a real `_wsState` lifecycle (`connecting`/`connected`/`fallback`/`reconnecting`/`disconnected`).
**Input**: `session.writeViaMux()` for programmatic/curl input — tmux `send-keys -l` (literal) + `send-keys Enter`. Single-line only (fire-and-once). Interactive **browser** input goes through a durable **exactly-once** layer: each frame carries a stable `clientId` + monotonic per-session `seq`, persisted to localStorage until the server ACKs (`{t:'ia',seq}` over WS, or HTTP 2xx), so a dropped link/reconnect can't lose or double-deliver a prompt. A duplicate is ACKed as `{t:'ia',seq,dup:true,last}` so a client whose persisted counter rolled back below the server's watermark can lift itself out instead of typing into a silently dead terminal (`docs/reliable-input-delivery.md`). **WS resilience** (#149): the upgrade URL carries `cid = clientId + ':' + perTabNonce`, and `ws-connection-registry.ts` supersedes only same-TAB reconnects (two tabs on one session coexist; input frames keep the bare `clientId` for seq dedup); reconnects back off exponentially (attempts preserved across `_connectWs`), and the header connection chip renders from a real `_wsState` lifecycle (`connecting`/`connected`/`fallback`/`reconnecting`/`disconnected`).
### Per-session env overrides: exact-key allowlist and CLAUDE_CONFIG_DIR
@@ -92,7 +92,7 @@ Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/do
### Plan-usage chip (statusLine telemetry)
**Plan-usage chip** (`showPlanUsageLimits`, per-device: desktop default **ON** since 1.9.3, handhelds OFF) renders compact Claude and Codex provider rows. Claude Code (v2.1.80+) pipes a JSON blob to a configured `statusLine.command` on each render; on Pro/Max it carries a `rate_limits` object (`five_hour`/`seven_day` windows only — no Opus weekly field — each `{used_percentage 0-100, resets_at epoch-SECONDS}`). Codeman injects its OWN statusLine exporter (`generateStatusLineCommand()` in `hooks-config.ts`, identified by the `/api/status-telemetry` marker — it only ever adds/updates/removes a statusLine that is _ours_, never a user's hand-authored one) that POSTs the blob to `POST /api/status-telemetry`. That route (auth-exempt like `/api/hook-event` — localhost-only, hook-secret-gated whenever auth is active, COD-91) parses via `usage-telemetry.ts` (pure, unit-tested), broadcasts SSE `session:statusTelemetry` (de-duped per session by `telemetrySignature` since the statusline fires on every assistant message), and returns a compact plain-text footer for the exporter to **print-through**. Main Codex subscription usage comes from the signed-in host CLI's read-only app-server `account/rateLimits/read` request at startup and every 5 minutes; `usage-telemetry.ts` selects only the main `codex` bucket (never model-specific buckets such as Spark), maps whatever 5-hour/7-day windows it supplies, and omits the provider row when unavailable. Credentials stay inside the CLI and no auth material is sent to the browser. `plan-usage-latest.ts` merges both process-wide sources and replays them in the SSE init snapshot (`getLightState`) so `#planUsageChip` renders immediately on page load/reconnect. `planUsageChipEnabled()` remains the single resolver behind the checkbox, chip visibility, and Claude create-time exporter flag. **Distinct from auto-resume** (which reacts to the Claude limit _message_; this proactively shows live percentages). Design: `docs/usage-limits-display-plan.md`. Tests: `test/usage-telemetry.test.ts`, `test/codex-plan-usage.test.ts`, `test/plan-usage-chip.test.ts`, `test/plan-usage-latest.test.ts`.
**Plan-usage chip** (`showPlanUsageLimits`, per-device: desktop default **ON** since 1.9.3, handhelds OFF) renders compact Claude and Codex provider rows. Claude Code (v2.1.80+) pipes a JSON blob to a configured `statusLine.command` on each render; on Pro/Max it carries a `rate_limits` object (`five_hour`/`seven_day` windows only — no Opus weekly field — each `{used_percentage 0-100, resets_at epoch-SECONDS}`). ⚠️ **Injected as an EPHEMERAL `claude --settings` CLI flag at spawn (2026-09-07), never written to disk** — `resolveStatusLineCliCommand()`/`ensureStatusLineExporterScript()` in `hooks-config.ts` (`generateStatusLineCommand()`/`applyStatusLineConfig()` remain, but only as the legacy disk-write self-heal path: a workspace an older Codeman build touched gets its stale `.claude/settings.local.json` entry stripped the first time a session starts there again). The exporter WRAPS a user's own real statusline (`findEffectiveUserStatusLineCommand()`, walking Claude Code's own settings precedence) rather than replacing it, and POSTs the `rate_limits` blob to `POST /api/status-telemetry`. That route (auth-exempt like `/api/hook-event` — localhost-only, hook-secret-gated whenever auth is active, COD-91) parses via `usage-telemetry.ts` (pure, unit-tested), broadcasts SSE `session:statusTelemetry` (de-duped per session by `telemetrySignature` since the statusline fires on every assistant message), and returns a compact plain-text footer for the exporter to **print-through** (foreground POST in the no-wrap branch so its own stdout becomes the footer, printing NOTHING on failure, `curl -sfk` plus `|| true`, since a bare brand word on the statusline is what discussion #405 opened with; backgrounded — `>/dev/null 2>&1 </dev/null &`, closing stdin too — only in the wrap branch, where the user's own command owns the footer; `curl --max-time 5` bounds a hung, not just refused, Codeman). Main Codex subscription usage comes from the signed-in host CLI's read-only app-server `account/rateLimits/read` request at startup and every 5 minutes; `usage-telemetry.ts` selects only the main `codex` bucket (never model-specific buckets such as Spark), maps whatever 5-hour/7-day windows it supplies, and omits the provider row when unavailable. Credentials stay inside the CLI and no auth material is sent to the browser. `plan-usage-latest.ts` merges both process-wide sources and replays them in the SSE init snapshot (`getLightState`) so `#planUsageChip` renders immediately on page load/reconnect. `planUsageChipEnabled()` remains the single resolver behind the checkbox and chip visibility (DISPLAY only) — the SAME `showPlanUsageLimits` setting also doubles as the server-side telemetry COLLECTION switch, read FRESH from `settings.json` by `readPlanUsageTelemetryEnabled()` at every claude session create/respawn (`TmuxManager.createSession`/`respawnPane`), never cached, with no per-session field and no per-request wire field — applies uniformly across every claude-creation path (interactive Run, cron, Ralph Loop API, quick-start) and survives a Codeman restart by construction (nothing per-session to lose). ⚠️ An ABSENT key reads as ON, the same way an absent `workspaceHooksEnabled` does: the desktop chip already shows as on for an install that never touched the setting, and the exporter posts only to this Codeman over loopback. Resolving the default in the reader is what keeps `GET /api/settings` a plain read. It briefly reconciled the key on first read (persisting `true` when absent), but `readJsonConfig()` answers `{}` for ANY read failure, not only ENOENT, and every page load hits that route, so one unlucky read replaced the whole settings file with a one-key file; pinned by `test/routes/system-routes-settings-get-plan-usage-default.test.ts`. ⚠️ The client sends `showPlanUsageLimits` in a settings save ONLY when that save FLIPS the chip relative to what the device had (`planUsageCollectionFlip()` in settings-ui.js): the chip defaults OFF on handhelds, so sending it on every save let a phone saving its font size persist `false` and switch collection off for every desktop, whose chip then went stale with no error anywhere. An explicit toggle on any device still writes the switch. Injection covers LOCAL tmux-spawned claude sessions only: the non-tmux direct-PTY fallback (`Session.startInteractive` when tmux is unavailable) and the remote/docker pane builders do not carry the flag. Registry-gated on `getCli(mode)?.capabilities.statusLineTelemetry` rather than a hardcoded mode string. **Distinct from auto-resume** (which reacts to the Claude limit _message_; this proactively shows live percentages). Design: `docs/usage-limits-display-plan.md`. Tests: `test/usage-telemetry.test.ts`, `test/codex-plan-usage.test.ts`, `test/plan-usage-chip.test.ts`, `test/plan-usage-latest.test.ts`, `test/hooks-config.test.ts` (statusline exporter script + `readPlanUsageTelemetryEnabled`), `test/statusline-cli-flag.test.ts`.
### Cron jobs
+363
View File
@@ -0,0 +1,363 @@
# Custom Model Endpoint Profiles (all harnesses, local or cloud)
## Context
The author pays for Claude Code but also runs a capable local model behind an
OpenAI-compatible server (llama.cpp) — and wants the same mechanism to work
against a **cloud** OpenAI-compatible endpoint too (e.g. Azure AI Foundry's
OpenAI-compatible inference endpoint, OpenRouter, a self-hosted gateway).
Right now every Codeman session mode defaults to its native cloud backend
with no way to redirect a session at any other endpoint from the UI — the
closest existing precedent is DeepSeek's server-env-sourced
`DEEPSEEK_BASE_URL`, which isn't user-facing.
**Scope note**: this plan originally said "local LLM." It now covers any
OpenAI-compatible endpoint the user configures — local (llama.cpp, Ollama,
vLLM) or cloud (Azure AI Foundry, OpenRouter, a company gateway). The
mechanism is identical (a base URL Codeman probes via `GET /v1/models`); the
only real differences are auth-header convention (cloud endpoints often want
an `api-key` header, e.g. Azure, rather than `Authorization: Bearer`) and
that a cloud "model" may actually be a deployment name distinct from the
underlying model family (Azure AI Foundry deployments) — both are called out
where they matter below. Naming throughout this plan is **"custom model
endpoint,"** not "local model," to keep that scope explicit.
### Additional use case: on-premises AI hardware
"Local" isn't limited to a desktop running llama.cpp — a growing category of
purpose-built, on-premises AI hardware exists specifically to run a serious
model on-site with an OpenAI-compatible server, and this feature is exactly
the on-ramp for pointing Codeman at one:
- **NVIDIA DGX Spark** (and the DGX Spark-class "Spark" mini-supercomputer
line) — a compact on-prem inference/training box aimed at running large
local models with an OpenAI-compatible API surface.
- **AMD "Strix Halo" (Ryzen AI Max)** on-prem AI mini-PCs — unified-memory
APU hardware marketed for local LLM inference, typically fronted by
llama.cpp/Ollama/vLLM the same way a home server would be.
Neither needs anything new from this design: both present a standard
`/v1/models` + `/v1/chat/completions` OpenAI-compatible surface once the
inference server is running, so they're just another `baseUrl` entry in the
custom-model-hosts store, same as llama.cpp or a cloud endpoint. The
justification for building this generically (rather than hardcoding "point
Claude at my llama.cpp box") is precisely this: **the same endpoint registry
and per-CLI injection mechanism should work unmodified for any current or
future OpenAI-compatible box or service** — a home GPU rig today, a Spark or
Strix Halo appliance tomorrow, a company's on-prem inference cluster after
that — without Codeman needing to know or care what's actually serving the
model on the other end of that URL.
A concrete example worth naming: **[Ark0N/Qwen5090](https://github.com/Ark0N/Qwen5090)**
(from the same GitHub account as this project's owner) is a one-click
Windows / one-command Linux installer that stands up Qwen3.8-27B locally on
an RTX 5090 (or another RTX 50-series card with ≥24GB) behind an
OpenAI-compatible API, served by any of vLLM, NInfer, or llama.cpp — MIT-
licensed tooling over Apache-2.0 Qwen weights. It's a direct, ready-made
target for this feature: point a custom-model-hosts entry at whichever
backend it's running, and it needs nothing further from Codeman's side. It's
also notable for already wiring up DeepSeek Harness and Claude Code as
coding agents against that local server itself, which is effectively the
same "point a Codeman-supported harness at a local endpoint" idea this
feature is generalizing — worth using as a real-world reference/test target
once chunk 5 (session integration) exists, alongside the author's own llama.cpp
box.
Each harness has its own (different-shaped) mechanism for pointing at a
custom OpenAI-compatible base URL + model — env vars for Claude, a JSON
config blob for opencode, a TOML file for Codex, etc. The author gave the
starting recipes for those three; the rest (Gemini, Pi, Grok, DeepSeek, OMP,
Antigravity) were researched for this plan and are flagged by confidence
below. A real end-to-end pass against the author's own llama-swap server
(`scripts/test-local-llm-harnesses.ts`, inside a `codeman/agent:llm-test`
Docker image with all 9 CLIs installed) then confirmed **claude and
opencode work end-to-end**, corrected a real Codex config.toml schema bug
the given recipe had (see the Codex row below), and surfaced that Codex's
_protocol_ — not just its config shape — does not work against a plain
OpenAI-Chat-Completions server like llama.cpp/llama-swap at all. Confidence
below reflects what was actually observed, not just what was planned.
The feature must be:
- **Off by default**, one settings toggle turns it on.
- Endpoint entry: user gives a base URL — a LAN address or a cloud URL —
plus an optional API key, and Codeman calls `GET <baseUrl>/v1/models` to
discover and store the available model (or deployment) list.
- A **new toolbar selector** (separate from the existing Run-mode menu, since
it's a modifier on top of whichever harness is already selected/running)
lets the user pick "Cloud (default)" — the harness's own native backend —
or a model discovered from one of the configured custom endpoints.
- Picking a custom-endpoint model for an **already-running session restarts
that session's CLI process** with the injected env/config pointed at that
endpoint (confirmed with the maintainer — these harnesses read endpoint config at
process start, not per-turn, so a live hot-swap isn't possible).
- **New sessions always default back to the harness's native cloud backend.**
A custom-endpoint selection is a per-session override, not a sticky global
default — starting a fresh CLI (any mode) always launches against its
native backend unless the user explicitly picks a custom endpoint for that
new session too. The toolbar selector is scoped to "this session," never
carried forward as the default for future sessions.
This follows the repo's existing data-driven CLI-registry philosophy
(`test/cli-registry-no-id-branching.test.ts`): per-CLI behavior is a
declared capability, never an `if (mode === 'claude')` branch.
## Per-CLI injection recipes (confidence-ranked)
| CLI | Mechanism | Confidence |
| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Confirmed reaching the server, but failing — unresolved.** A real run against llama-swap returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` consistently (confirmed the env vars are read: the request reaches the network rather than failing locally). Root cause not identified — plausible explanation by analogy with codex's Responses-API gap is that `dsh --profile headless` expects DeepSeek's official API response shape/path structure rather than a generic OpenAI-compatible `/v1/chat/completions` endpoint, but this was not confirmed by reading dsh's own bundled source (unlike pi/grok, where that grep resolved the question directly). Documented as best-effort/unknown, not shipped as verified working |
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
Everything web-researched-but-unverified gets implemented but must be
smoke-tested against real installs of those CLIs before being called done —
call this out explicitly when implementing, don't just ship on faith.
**Cloud-endpoint specifics** to keep in mind per recipe above: an Azure AI
Foundry-style endpoint typically wants the API key in an `api-key` header
rather than (or in addition to) `Authorization: Bearer`, and its "model" is
often a deployment name rather than the underlying model family name — the
discovery step (`GET /v1/models`) still works the same way against Azure AI
Foundry's OpenAI-compatible endpoint shape, but a user may need to type the
deployment name manually if it isn't returned as expected.
## Architecture
### 1. Registry: new `capabilities.customModelInjection` field
Extend `src/config/cli-registry/types.ts` / `schema.ts` with a discriminated
union on each `CliEntry.capabilities`:
```ts
type CustomModelInjection =
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[] }
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json' }
| {
kind: 'configDir';
dirEnvVar: string;
fileName: string;
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml';
}
| { kind: 'unsupported' };
```
Declared per stock.ts entry per the table above. A pure function in a new
`src/custom-model-injection.ts` (`buildCustomModelInjection(entry, endpoint, modelId)`)
turns `(CliEntry, endpoint, modelId)` into either an `envOverrides` object
(kind `env`/`configContentEnv`) or a `{ dirEnvVar, files: [{path, content}] }`
descriptor (kind `configDir`) — unit-testable with no IO, mirroring how
`session-cli-builder.ts` is pure. The `configDir` kind additionally needs an
IO wrapper that writes those files under
`dataPath('custom-model-configs/<sessionId>/')` (new dir, cleaned up on
session delete — same lifecycle as other per-session generated state).
### 2. Endpoint registry: `src/custom-model-hosts.ts`
Same read-array/write-array shape as `src/remote-hosts.ts` /
`src/webview-store.ts`: `~/.codeman/custom-model-hosts.json` holding
`CustomModelEndpoint[] = { id, label, baseUrl, apiKey?, authStyle?: 'bearer'|'api-key'|'both', models?: string[], lastDiscoveredAt? }`.
`authStyle` defaults to `'both'` (send both header conventions on the
discovery probe, same approach the smoke-test script below uses) so one
endpoint entry works whether it's llama.cpp or Azure without the user having
to know which header their box wants in advance.
New route file `src/web/routes/custom-model-routes.ts` (registered in the
routes barrel), mirroring `case-routes.ts`'s remote/docker-host CRUD
(`GET/POST/PUT/DELETE /api/model-endpoints`, admin-gated in multi-user mode
the same way) plus:
- `POST /api/model-endpoints/:id/discover-models` — fetches
`${baseUrl}/v1/models`, stores the `data[].id` list, returns it. Bounded
timeout, and run the target through the **same SSRF egress guard already
used for web tabs** (`webview-egress-policy.ts` — reject link-local/cloud
metadata addresses) — this still matters for a cloud URL too, since the
guard is about preventing a redirect to internal infra, not about
local-vs-cloud.
**Why discovery rather than a free-text model field**: it removes the one
piece of configuration most likely to trip a user up — hand-typing the
exact model identifier a given inference server expects, which varies by
server and is an easy source of a silent "model not found" failure with no
useful error surfaced back through a CLI's own startup. Discovery also
means this design is not limited to a single-model box: a **multi-model
gateway** such as **[llama-swap](https://github.com/mostlygeek/llama-swap)**
(hot-swaps between several loaded llama.cpp model configs behind one
OpenAI-compatible endpoint) or a vLLM/LiteLLM/Ollama instance serving
several models advertises ALL of them through the same `/v1/models` call —
so one endpoint entry surfaces every model that gateway can serve, with no
extra per-model configuration on Codeman's side at all.
### 3. Settings
- New synced boolean `customModelEndpointsEnabled` in `SettingsUpdateSchema`
(`src/web/schemas.ts`), default `false`, documented inline like
`readMyMindEnabled`/`workspaceHooksEnabled`.
- New `.set-group` "Custom Model Endpoints" inside the **Agents & CLIs**
section (`settings-clis`, `index.html:2150+`) with the enable toggle plus
a list-editor (add/refresh-models/delete rows) for endpoints — closest
existing precedent is the respawn-presets array editor
(`schemas.ts:1285-1305`, `index.html:1243-1244`) for add/apply/delete-by-id
semantics, backed by the new CRUD routes above.
### 4. Toolbar UI
- New header/toolbar button (e.g. `#customModelBtn`, `btn-toolbar
btn-custom-model`), marker-hidden by default (`btn-custom-model--hidden`)
and revealed by `applyHeaderVisibilitySettings()` only when
`customModelEndpointsEnabled` is on — same pattern as the File
Viewer/Cron buttons.
- Clicking opens a dropdown (`#customModelMenu`, same `.run-mode-menu`-style
markup as the existing Run-mode gear menu) listing "Cloud (default)" plus
every discovered model, grouped by endpoint. An entry is disabled with a
tooltip when the active session's CLI has `customModelInjection.kind ===
'unsupported'` (Antigravity) or none declared.
- Selecting an entry calls a new route:
`POST /api/sessions/:id/custom-model { endpointId, modelId } | { clear: true }`.
Server: resolve the CLI entry for `session.mode`, build the injection via
§1, persist it as a new `session.customModel` state field (surfaced in
`toState()`/SSE so the tab can show a small badge, e.g. "🖥 qwen3 (local)"
or "☁ gpt-4o-mini (azure)", and the choice survives reload), merge into
the session's `envOverrides`, and **respawn the pane's CLI process**
through the same respawn/interactive-restart path
`session.ts`/`tmux-manager.ts` already use for effort/model changes
(`_configureCliEnv()` + `applyEnvOverrides()` at spawn time) — reuse,
don't reinvent, the existing kill-and-relaunch-in-pane machinery.
- New-session creation deliberately does **not** inherit a prior custom-
endpoint choice: `buildEnvOverrides()` (session-ui.js) never carries the
toolbar selection forward to the next `run()` call. Every new session
starts on its native backend; picking a custom endpoint in the toolbar for
a session applies only to that session (and, if done before Run is
clicked, to the one session about to be created — not to sessions created
afterward).
### 5. Multi-user security clamp
Every new env var this feature introduces that can redirect a session's
traffic (and thus wherever its credentials go) — `ANTHROPIC_BASE_URL`,
`GOOGLE_GEMINI_BASE_URL`, `GROK_BASE_URL`, the `CODEX_HOME`/`PI_CONFIG_DIR`
dir-redirects, plus the already-privileged `DEEPSEEK_BASE_URL` — must be
added to each CLI's `capabilities.privilegedEnvKeys` so
`clampEnvOverridesForOwner()` strips them for a non-granted multi-user
owner, exactly the precedent already documented for `DEEPSEEK_BASE_URL`/
`OMP_AUTH_BROKER_URL`. This matters _more_, not less, now that endpoints can
be cloud URLs: redirecting a non-granted user's session to an attacker's
cloud endpoint is a credential-exfiltration path, not just a mischief
redirect to a LAN box. Endpoint CRUD itself stays admin-only in multi-user
mode, same as remote/docker hosts.
## Files touched (representative, not exhaustive)
- `src/config/cli-registry/types.ts`, `schema.ts`, `stock.ts` — new capability + per-entry declarations
- `src/custom-model-injection.ts` (new) — pure per-CLI descriptor builder + unit tests
- `src/custom-model-hosts.ts` (new) — endpoint store
- `src/web/routes/custom-model-routes.ts` (new) — CRUD + discovery route
- `src/web/routes/session-routes.ts` — `POST /api/sessions/:id/custom-model`, clamp wiring
- `src/web/schemas.ts` — `customModelEndpointsEnabled`, endpoint/discover payload schemas, privileged-key updates
- `src/session.ts` — `customModel` state field, `toState()` surface
- `src/web/public/index.html`, `settings-ui.js`, `session-ui.js`, `styles.css` — settings group, toolbar button/menu, badge, accent CSS
- `src/web/sse-events.ts` + `constants.js` — if a dedicated SSE event is warranted for the badge (or just ride existing session-update broadcasts)
- `test/fixtures/mock-openai-server.ts` (new) + `test/custom-model-injection-contract.test.ts` (new) — see Mock-server validation below
- `scripts/test-local-llm-harnesses.ts` (already added, this branch; run via `npx tsx`) — the standalone real-CLI-and-real-endpoint smoke test, supporting any `--base-url` (local or cloud). Dynamic: derives its harness list and every env var/config it injects from the live CLI registry + `buildCustomModelInjection()` rather than a second hand-maintained copy — only the one-shot invocation flags (`ONE_SHOT` table) are CLI-specific info the registry doesn't model and stay hand-maintained
- `docs/custom-model-endpoints.md` (new) + a CLAUDE.md pointer bullet under External CLI modes / envOverrides
## Mock-server validation strategy (CI-runnable, no real CLI binaries needed)
Spawning nine real CLI binaries in CI isn't realistic, and neither the author's
llama.cpp box nor a real cloud subscription can be a CI dependency. So the
injection _logic_ gets a tier of automated coverage that sits between the
pure unit tests and the live manual checks in Verification:
1. **`test/fixtures/mock-openai-server.ts`** — a small in-process HTTP
server (plain `http.createServer`, no external deps, port picked per the
existing `const PORT = 3150+` convention) that:
- Serves `GET /v1/models` → a fixed fake model list (`{data:[{id:'qwen3'},...]}`),
for testing the discovery route.
- Serves `POST /v1/chat/completions` (OpenAI shape) **and**
`POST /v1/messages` (Anthropic Messages-API shape, since that's what
`ANTHROPIC_BASE_URL` traffic looks like) and records every request it
receives (headers, body, path) into an array the test can assert on —
including which auth header style it saw, so the `authStyle: 'both'`
default and Azure's `api-key` convention both get real coverage.
- Returns a minimal valid completion so a client library doesn't choke
on the response shape.
2. **`test/custom-model-injection-contract.test.ts`** — for every CLI with a
`customModelInjection` capability (i.e. every row in the table above
except `antigravity`):
- Point a fixture `CustomModelEndpoint` at the mock server's URL.
- Call `buildCustomModelInjection(entry, endpoint, modelId)` (the pure
function from §1) to get the real env vars / config-file content that
would be injected into that CLI's session.
- Replay those exact values through a minimal HTTP request shaped the
way that CLI is documented to send it (Anthropic Messages shape for
claude; OpenAI chat-completions shape for opencode/codex/pi/grok/omp;
`GOOGLE_GEMINI_BASE_URL`'s OpenAI-compat shape for gemini; dsh's
provider call for deepseek) against the mock server.
- Assert the mock server received the request **at the injected
`baseUrl`**, with **the injected API key** in the expected header, and
**the injected model id** in the body/path — i.e. prove the values
Codeman computes are internally consistent and would reach the right
place with the right identifiers, end to end, in CI, on every push.
- Also cover the `configDir` kind (codex/pi/omp): assert the written
`config.toml`/`models.json`/`models.yml` file parses and contains the
same base URL/key/model, and that it's written under the isolated
per-session dir rather than the user's real config path.
3. **Explicit, stated limitation** (goes in the test file's `@fileoverview`
and in this doc, not left implicit): this proves _"if the CLI honors its
documented env/config contract, it will hit the right endpoint with the
right model."_ It does **not** prove the real CLI binary actually reads
that env var / config file the way its docs say — that's still the job
of the live manual checks in Verification step 4-5 below, and is exactly
why the confidence table above did not stop at "researched" — every CLI
except antigravity (no mechanism at all) has since been run against a
real llama-swap server via `scripts/test-local-llm-harnesses.ts`:
claude/opencode/pi/grok/omp are confirmed PASS end-to-end, codex is
confirmed FAIL for a real documented protocol reason (Responses-API-only
since Feb 2026), and gemini/deepseek are confirmed reaching the server
but failing for reasons not yet root-caused (see their table rows). The
mock-server suite catches regressions in Codeman's own logic; it cannot
catch a CLI changing its env-var name in a future release, or a real
cloud endpoint behaving differently from a local llama.cpp box.
## Verification
1. `npm run typecheck && npm test` after each slice — this now includes the
mock-server contract suite from above, so injection-logic regressions
are caught automatically without touching real infrastructure.
2. Unit tests for `buildCustomModelInjection()` per CLI kind (pure, no IO).
3. Route tests (`app.inject`) for the new CRUD + discover-models endpoint
(mock `fetch` for `/v1/models`), and for the multi-user clamp on the new
privileged keys (mirror `test/routes/external-cli-bypass-clamp.test.ts`).
4. **Standalone real-binary smoke test**: `scripts/test-local-llm-harnesses.ts`
exercises every harness the CLI registry declares `customModelInjection`
support for against a real `--base-url` — local or cloud — outside of
Codeman's UI entirely, and is DYNAMIC (reads `enabledClis()` + calls the
real `buildCustomModelInjection()`, so a future registry change is picked
up automatically with zero edits to the script). Already run to
completion against the author's llama-swap server (a LAN address,
inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries):
claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected**
(Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED**
(reach the server, fail for undiagnosed reasons — see their table rows),
antigravity **SKIP** (no mechanism). Re-run this against a real cloud
endpoint (e.g. an Azure AI Foundry deployment) once one is available, to
prove the `authStyle`/deployment-name handling holds up outside llama.cpp.
5. Once the full feature (not just the standalone script) is built: add an
endpoint via the real UI, hit discover-models, confirm the returned model
list, pick Claude + the model on a real session, confirm via
`tmux -L codeman capture-pane`/`tmux showenv -t <pane>` that
`ANTHROPIC_BASE_URL`/`ANTHROPIC_API_KEY`/`ANTHROPIC_DEFAULT_*_MODEL` are
set post-restart, and confirm the endpoint's own logs show the next
prompt actually landing there. Repeat for opencode and Codex at minimum
before considering this shippable; spot-check the web-researched CLIs
and correct the plan's confidence table with what's actually observed.
6. `npm run lint && npm run format:check`.
7. Update `CHANGELOG.md`/changeset per the COM workflow when shipping.
+147
View File
@@ -0,0 +1,147 @@
# Custom Model Endpoint Profiles
Point any Codeman-supported harness — Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, or OMP — at a custom OpenAI-compatible endpoint instead of
its native cloud backend, for a given session. "Custom endpoint" covers both
**local** hardware (llama.cpp, Ollama, vLLM, a home GPU rig, or purpose-built
boxes like NVIDIA DGX Spark or AMD Strix Halo mini-PCs) and **cloud**
services (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
company gateway) — anything answering `GET /v1/models` and
`POST /v1/chat/completions` in the standard shape. Design doc, per-CLI
recipe confidence table, and security reasoning:
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
> **Status**: backend is implemented and tested (registry capability, the
> injection engine, the endpoint store + discovery route, the session
> restart route). The toolbar picker / settings UI described below as the
> intended surface is **not yet built** — until it lands, use the HTTP API
> directly (examples below). Antigravity has no known custom-endpoint
> mechanism and is not supported.
## Turning it on
App Settings → Agents & CLIs → **Custom Model Endpoints** (synced setting
`customModelEndpointsEnabled`, default **OFF**). Until the toolbar picker
lands, nothing reads this setting: the HTTP routes below work whether it is
on or off, and it exists now only so the picker has a switch to hang off
when it ships. The API equivalent:
```bash
curl -sk -X PUT https://localhost:3000/api/settings \
-H 'Content-Type: application/json' \
-d '{"customModelEndpointsEnabled": true}'
```
## Adding an endpoint
```bash
curl -sk -X POST https://localhost:3000/api/model-endpoints \
-H 'Content-Type: application/json' \
-d '{"id": "llama-box", "label": "Home llama.cpp", "baseUrl": "http://192.168.1.50:8080"}'
```
`apiKey` is optional (most local servers don't check it). `authStyle`
(`bearer` | `api-key`, default `bearer`) controls which auth header
convention discovery uses: `bearer` is `Authorization: Bearer <key>`
(llama.cpp, OpenAI-compatible servers, most gateways), `api-key` is the
`api-key: <key>` header Azure AI Foundry wants. There is deliberately no
"send both" option: measured against a real llama-swap server, a request
carrying both headers hung indefinitely. `baseUrl` must be `http(s)`, carry
no embedded credentials, and may not point at a link-local or cloud-metadata
address; discovery re-checks the address the name actually resolves to.
Discover its available models:
```bash
curl -sk -X POST https://localhost:3000/api/model-endpoints/llama-box/discover-models
```
This calls the endpoint's own `GET /v1/models` and stores the returned list
on the endpoint record; `GET /api/model-endpoints` lists everything
configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
Endpoint management is admin-only in multi-user mode, same as remote/docker
hosts — these are machine-level infra, not per-user settings.
## Applying a model to a session
```bash
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
-H 'Content-Type: application/json' \
-d '{"endpointId": "llama-box", "modelId": "qwen3"}'
```
This computes the CLI-specific env vars / config for that session's mode
(see the recipe table in `custom-model-endpoints-plan.md`) and **restarts the session's
CLI process in place** — same pane, same tmux session, fresh env. That
restart is necessary, not incidental: every supported harness reads its
endpoint config at process start, not per-turn, so there is no live
hot-swap. A Claude session is relaunched with `--resume <conversation> ||
--session-id <id>`, so it continues the conversation it was on; pi, omp and
grok are relaunched with the `--model` value that selects the injected
provider (`custom/<modelId>` for pi and omp, `codeman-custom` for grok),
since for those three the config file alone does not switch the model.
**Remote (SSH) and Docker sessions are refused** (400) for now: their restart
reattaches the durable remote/in-container tmux rather than relaunching the
agent, so the selection would report success and change nothing.
Clear back to the harness's native cloud default with:
```bash
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
-H 'Content-Type: application/json' -d '{"clear": true}'
```
Clearing also removes the env vars the selection injected from the tmux
session (they persist there and would otherwise be inherited by the
relaunched CLI) and deletes the per-session config directory
(`~/.codeman/custom-model-configs/<sessionId>`, written 0600 because pi and
omp embed the API key in it). That directory is also removed when the
session is deleted. The selection survives a Codeman restart: the endpoint
id, model and injected key NAMES are persisted, the values are re-derived
from the endpoint store on recovery, and the pane keeps running against the
endpoint in between because tmux retains its environment.
**New sessions always default back to the harness's native backend.** A
custom-endpoint selection is a per-session choice, never a sticky global
default — starting a fresh session doesn't inherit whatever the last one was
pointed at.
## Confidence per harness
Every harness except Antigravity has now been run end-to-end against a real
llama-swap server via `scripts/test-local-llm-harnesses.ts` (a dynamic
script that reads the live CLI registry, so a registry change is picked up
automatically). Results:
- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply
came back through the endpoint.
- **Codex** — the config is structurally correct, but Codex only speaks the
Responses API since Feb 2026, which llama.cpp/llama-swap don't implement.
This is a real protocol incompatibility, not a bug here; Codex support
needs a Responses-API-compatible endpoint.
- **Gemini** — fails with `Invalid auth method selected`, traced to an
undocumented `GATEWAY` auth path gemini-cli selects once
`GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation
(several auth workarounds were tried and ruled out); do not rely on
Gemini support yet.
- **DeepSeek** — the request reaches the server (env vars are read) but
gets a consistent `HTTP_404`. Root cause not identified; best-effort only.
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind
each result. `scripts/test-local-llm-harnesses.ts` is the standalone script
used to check a harness against a real endpoint outside the web UI
entirely; see its own `--help` for usage.
## Security note
Every env var this feature can set that redirects a session's traffic
(`ANTHROPIC_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`, `CODEX_HOME`, etc.) is
listed in that CLI's `privilegedEnvKeys` in the CLI registry, so a
non-granted multi-user owner cannot set one directly via the generic
`envOverrides` API field — only through this feature's own route, which
computes the value from an admin-configured, SSRF-guarded endpoint rather
than trusting arbitrary client input. See the "Multi-user security
hardening" section of `custom-model-endpoints-plan.md` for the full reasoning; several
of these were reachable via the generic `envOverrides` field even before
this feature existed, and building this surfaced and closed that gap.
+4 -2
View File
@@ -15,7 +15,7 @@ The application container mounts the Docker daemon socket so Codeman can create
## Start
Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/Coding/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes.
Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes.
```sh
cp docker/.env.example docker/.env
@@ -33,12 +33,14 @@ On Linux, run the stack with the start script. It determines `PUID` and `PGID` f
bash docker/Start-Codeman.sh
```
On other platforms, run Compose directly. `PUID` and `PGID` default to `1000:1000`; set them in `docker/.env` when the application-data directory has a different owner.
On other platforms, run Compose directly. `PUID` and `PGID` default to `1000:1000`; set them in `docker/.env` when the application-data directory has a different owner. Naming the file with `-f` disables Compose's own discovery of `docker/docker-compose.override.yml`, so add a second `-f` for it when you keep one (see `docker/README.md`, Local customisation).
```sh
docker compose --env-file docker/.env -f docker/docker-compose.yaml up --build -d
```
The container starts as root, corrects the ownership of a bind source the daemon had to create, and drops to `PUID:PGID` with `setpriv` before Codeman starts; the capabilities that needs are declared in `docker/docker-compose.yaml` and named by the entrypoint when a compose file written elsewhere lacks them.
Open `http://localhost:3000` and sign in with the username and password from `docker/.env`.
## Operations
+15
View File
@@ -59,6 +59,7 @@ unchanged. The container path is a new `SupervisorKind`, not a new updater.
| `CODEMAN_RESTART_BY_EXIT=1` | The Compose file's declaration of that policy, so the updater may exit even with no Docker socket. |
| Toolchain + devDependencies in the image | Lets `npm install` and `npm run build` run inside the container. |
| `docker-env-applied.json` | Fingerprint baseline, written by `Start-Codeman.sh` on every start. |
| `docker-build-source.json` | What HEAD/`package-lock.json` the build artefact volumes currently reflect. Written by both `Start-Codeman.sh` and this in-place update, so the two agree on whether those volumes are stale. |
### Why build artefacts are in named volumes
@@ -72,6 +73,20 @@ Docker seeds an empty named volume from the image, so the first start inherits t
image's already-built `node_modules` and `dist` and pays no bootstrap cost.
`docker compose down -v` is the supported reset: the next start re-seeds them.
That seeding-only-while-empty behaviour has a second, less obvious edge: it also
means a plain `docker compose build` triggered from OUTSIDE the container (for
example `Start-Codeman.sh`, after a `git pull` done by hand rather than through
this in-app updater) produces a fresh image whose freshly-built `dist`/
`node_modules` then sit unused behind the volumes' OLD content — the container
comes back up looking unchanged. `Start-Codeman.sh` detects this by comparing the
checkout's current HEAD and `package-lock.json` hash against `docker-build-source.json`,
and clears just the affected volume(s) before its own `--build` if they moved.
This in-place update writes that same file after a successful build precisely so
that comparison does not fire on stale information: without it, the next plain
`Start-Codeman.sh` run would see the HEAD this update just checked out, not
recognise it as already accounted for, and wipe the volumes this update just
correctly rebuilt right back to the OLDER image.
### Why the runtime image carries a build toolchain
`npm run build` is `tsc` plus `esbuild`, both devDependencies, so the image no
+6 -3
View File
@@ -333,9 +333,12 @@ Out of scope per the issue, and the current behavior already degrades correctly:
- **Docker cases**: the workspace is a host directory bind-mounted at the same absolute path, so a host-side
write is visible in the container immediately. Edit mode works and needs nothing special. Worth one line
in the docs.
- **Remote SSH cases**: `workingDir` is a path on the remote host. `validateSessionFilePath` realpaths it
locally, which fails, so the write returns 404 exactly like the read routes do today. Confirm the viewer
shows a clean empty/error state rather than an unexplained failure, and do not attempt an SFTP path.
- **Remote SSH cases**: `workingDir` is a path on the remote host, and the READ routes now
resolve it over ssh (`src/remote-files.ts`, same `buildSshConnectionArgs` discipline as the
launch path — #415). What stays unsupported is the WRITE side: an `edit=1` / `PUT` answers
`400` "editing is not supported for files in a remote (SSH) case", `editable` is always
`false`, office previews and generated thumbnails answer `400`, and no remote file is ever
copied to the server's disk. Do not attempt an SFTP write path.
---
+11 -2
View File
@@ -50,8 +50,17 @@ each `(clientId, seq)` at most once, so a resend can't type the prompt twice.
last-applied is seen. A replayed/lower seq returns `false`. Bounded MRU map
(`MAX_INPUT_DEDUP_CLIENTS = 256`).
- **WS route** (`ws-routes.ts`) — parses optional `cid`/`seq` on `{t:'i'}`; applies
via `shouldApplyInput` (skips a duplicate, still ACKs with `{t:'ia',seq}` so the
client drops it). Untagged frames apply unconditionally (no behavior change).
via `shouldApplyInput`. An applied frame is ACKed with `{t:'ia',seq}`; a duplicate is
ACKed as `{t:'ia',seq,dup:true,last:<watermark>}`, where `last` is the server's
highest applied seq for that `clientId` (`Session.lastInputSeq`). The client drops
the record either way, and on `dup` it lifts its own counter to `last` first and
re-sends a FIRST-attempt record (a retry being called a duplicate is the mechanism
working: the original landed). Without `last`, a tab killed between a send and the
persisted counter write came back counting BELOW the server's watermark, and every
later keystroke was dropped-but-ACKed: a silently dead terminal a reload could not
fix, since the stale counter was restored from localStorage too. The client now
persists the counter synchronously on every send for the same reason. Untagged
frames apply unconditionally (no behavior change).
- **POST route** (`/api/sessions/:id/input`) — optional `seq`/`clientId` in
`SessionInputWithLimitSchema`; a deduped duplicate returns 200 without writing
(the 200 is the client's ACK). `curl`/legacy callers omit the fields and always
+101
View File
@@ -251,6 +251,107 @@ unreachable host answers "unknown", which also means do not revive. The answer
is cached per session and cleared whenever the pane is next seen alive, so a
stale `true` from one transport drop can never revive the NEXT clean exit.
## File access over SSH
A remote case's `workingDir` is an absolute path on the **remote** host
(`Session.workingDir = RemoteCase.remotePath`), so the file routes cannot use local
`fs`: a local `realpathSync` on a remote-only path fails by construction, which is why
previewing a file used to answer `404 File not found` for a case that was working
perfectly (#415). `src/remote-files.ts` is the one module that reads remote bytes,
and it follows the same rule as the launch path: every ssh command line comes from
`buildSshConnectionArgs()` — **never** a hand-built ssh line.
| Request | What happens |
|---------|--------------|
| `GET /api/sessions/:id/file-raw` | Streamed over `ssh` (`cat`, or `tail -c +N \| head -c L` for a `Range`); the same 200/206/416 contract as a local file, so `<video>`/`<audio>` seeking works |
| `GET /api/sessions/:id/file-content` | `cat` into memory, capped by the existing text limit; `edit=1` answers `400` (see below) and `editable` is always `false` |
| `PUT /api/sessions/:id/file-content` | `400` before any path is looked at: the guard sits AHEAD of the local path validation, because with a same-named directory on the Codeman host (an `sshfs` mount) the write would otherwise land on the local twin |
| `GET /api/sessions/:id/file-preview` | Non-office files redirect to `file-raw` (which works remotely); docx/pptx answer `400` |
| `GET /api/sessions/:id/file-thumbnail` | `400` for remote files |
| `POST /api/sessions/:id/attachments` | Registers an absolute path that lives on the **remote** host (a clicked link pointing outside the case directory) by probing it there |
| `GET /api/sessions/:id/attachments/:attachmentId/raw` | Streams the registered remote file over ssh, same 200/206/416 contract; `preview` (office) and `thumbnail` answer `400` |
| `GET /api/sessions/:id/attachments/:attachmentId`, `GET …/attachments` (history) | Size/mtime/existence resolved over ssh, so a remote entry is not reported `missing`; the history list resolves EVERY entry in one batched probe, never one connection per entry |
⚠️ The attachment route is the one a clicked path takes when it is **outside** the case
directory (a remote `/tmp` scratchpad capture, a screenshot elsewhere in the home dir):
the frontend's `_isExternalPreviewPath()` sends every absolute path that is not under
`workingDir` there, so fixing only `file-raw` would leave exactly that half broken.
Guard order is deliberately **the same as locally**, and the checks are not weakened
by the transport:
1. Ownership (`findSessionOrFail` / the scope helper) — unchanged.
2. Lexical containment of `workingDir + path` — a `../` escape is refused before any
connection is opened.
3. ONE ssh round trip that returns `realpath` **and** `stat` for the path **and** the
workspace root (`remoteProbePaths`). Resolving the root remotely is what keeps the
boundary honest for a symlinked `remotePath`. The probe uses `readlink -f` when
available; on a host without it (macOS before 12.3) a POSIX fallback canonicalizes
the directory chain with `cd -P`/`pwd -P` and then follows the LAST component with
plain `readlink` for a bounded number of hops. ⚠️ **The fallback fails closed**: a
path it cannot fully resolve (a loop, a `readlink` failure, the hop cap) is reported
as unresolvable and answers 404, never as its own unresolved string. An earlier
version resolved only the directory chain, so `ws/notes.txt -> ~/.ssh/id_rsa` passed
containment under the link's own path while `cat` followed it to the key.
Records come back NUL-separated and index-keyed (`<index>|kind|size|mtime|realPath`,
after a leading NUL that fences off any login banner), so a filename containing a
newline cannot shift the alignment.
4. Containment of the remote realpath against the remote root. The sensitive-path
blocklist then applies on whichever routes already apply it locally (`/api/download`,
attachment registration, edit mode — where resolving symlinks first is what makes it
meaningful); the remote branch neither drops a guard the local path has nor invents a
stricter one. One entry of that blocklist is host-bound by construction: the three
home-anchored members (`~/.claude.json`, `~/.claude/settings.json`,
`~/.claude/settings.local.json`) are compared against the **Codeman host's** home
directory, so they do not match a remote home at a different path. Everything else in
the list is depth-anchored (`/.ssh/`, `/.aws/credentials`, `/.claude/.credentials.json`,
`/etc/shadow`, ...) and applies to a remote path unchanged.
5. Size cap (`CODEMAN_MAX_DOWNLOAD_BYTES`) applied to the **remote** size, before the
body is requested.
The path arrives from the browser (`?path=`) and is interpolated as a single
`shellescape`-quoted token, in a command that is itself shellescaped into the ssh
line; `BatchMode=yes` means a host needing a passphrase fails fast instead of hanging.
A failed connection is reported as **502** with the remote reason — never a 404, which
used to make an unreachable host look like a typo in the agent's output. The reason is
the first stderr line, the timeout, or the exit code; never Node's `Command failed: …`
message, which would carry the identity-file path and the probe script into the body.
**Connections are bounded.** Every probe and buffered read runs through a small global
semaphore (`src/remote-ssh-limiter.ts`, default 4, `CODEMAN_MAX_REMOTE_FILE_SSH`), the
attachment-history list resolves its whole history in one batched probe instead of one
handshake per entry, and probes are chunked at 40 paths per round trip. Terminal output
in a remote session is written on the remote host, so a prompt-injected agent printing
hundreds of `codeman://attach` links used to make the server fork one `ssh` per link,
each holding a 20 s probe timeout, and a 100-entry history re-listed on every
`attachment:detected` event tripped OpenSSH's default `MaxStartups 10:30:100`. Streams
(`file-raw`, by-id `raw`) are not counted: one is held per browser request for the life
of a playback, and each is gated behind a counted probe anyway.
⚠️ **There is deliberately NO local fallback.** A remote case reads the remote bytes or
fails, even when a file with the same absolute name exists on the Codeman host — which
is the ordinary case for the documented stop-gap workaround, an `sshfs` mount of the
remote tree at the identical path. Serving the local twin instead would silently hand
back a DIFFERENT filesystem's bytes under a name the user believes is the remote file
(a stale mount, a different checkout, a leftover file), and the failure would be
invisible. An existing mount therefore stops being load-bearing for previews and
downloads but is harmless, and a missing remote file stays a 404 even if the mount
still has it.
**Not available over ssh (by choice, not by accident):** editing a file (writes would
need SFTP; `docs/file-viewer-edit-plan.md` §6), office-document previews and
generated thumbnails (both need the bytes on the server's disk — no remote file is ever
spilled onto the server), the file-tree/picker listings, and `tail-file`. Those routes
are still local-only, so with an `sshfs` mount in place they read the mounted copy —
the two views can only disagree when that mount is stale. Docker cases are unaffected:
their workspace is bind-mounted at the same absolute path, so local `fs` reads real bytes.
⚠️ A remote record stores the **remote** path, and the same absolute path STRING means a
different file on each host. What decides which host to read is therefore never the
path but the SESSION (`session.remote`): a remote session never falls back to local
`fs`, and a local session never opens an ssh connection — including for attachment
records, which are keyed to the session that registered them.
## API
Routes are registered in `src/web/routes/case-routes.ts`:
+5 -2
View File
@@ -125,7 +125,9 @@ loopback bind matters. The auth pipeline (`src/web/middleware/auth.ts`,
`onRequest` hook) runs in this order:
1. **Localhost‑only exemptions** (always first): `POST /api/hook-event` and the QR
`/q/` short‑code path are exempt when `req.ip` is loopback (see §3). While the
`/q/` short‑code path are exempt when `req.ip` is loopback (see §3). The three
web‑tab exemptions (§10b: the capability in the path, the `Referer` form, and
the lost‑frame recovery page) sit in this same slot, ahead of the credential checks. While the
**managed tunnel is running**, the hook‑event exemption additionally requires
the per‑instance `X-Codeman-Hook-Secret` header (COD‑54); failed presentations
are rate‑limited in a **dedicated bucket** (separate from Basic‑Auth failures)
@@ -514,9 +516,10 @@ Full feature guide: [`docker-cases.md`](docker-cases.md).
## 10b. Web tabs (dashboard proxy)
A saved dashboard URL renders as a tab, served through Codeman's own origin at `/webview/<capability>/`. User guide: [`web-tabs.md`](web-tabs.md). Three properties carry the security weight:
A saved dashboard URL renders as a tab, served through Codeman's own origin at `/webview/<capability>/`. User guide: [`web-tabs.md`](web-tabs.md). Four properties carry the security weight:
- **The proxy is exempt from cookie auth and the Origin/CSRF guard, and that is deliberate.** The iframe is sandboxed without `allow-same-origin`, so it is opaque‑origin: its requests are cross‑site, meaning the `SameSite=lax` session cookie is never attached and its writes and WS upgrades arrive with `Origin: null`. The credential is instead a 192‑bit capability in the path, minted only by an authenticated `POST /api/webviews/:id/open`, held in memory (a restart invalidates every one), rolling TTL, bound to the minting user, and granting nothing but "relay bytes to this one saved URL". ⚠️ **The Host allowlist is NOT bypassed**, so DNS‑rebinding protection is unaffected. A second `Referer`‑keyed form exists for root‑absolute assets and is the only exemption decided by a request‑supplied header, so it is fenced to safe methods on non‑`/api`, non‑`/ws`, non‑`/q` paths. Edges pinned by `test/webview-auth-exemption.test.ts`.
- **The lost‑frame recovery page is the third unauthenticated 200, and the only one decided by request headers alone.** The proxy's runtime shim masks `/webview/<cap>/` off the page's own URL so a single‑page app routes on the path it expects; a navigation the page then starts itself (`location.reload()`, a root‑absolute `location.href`) lands on Codeman's root with no capability anywhere, no cookie (opaque origin) and a Referer naming the masked page. `serveLostWebviewFrame()` in `middleware/auth.ts` recognises it by shape (`GET`/`HEAD`, `Sec-Fetch-Dest: iframe` or `frame`, `Accept: text/html`, `Sec-Fetch-Mode: navigate` or absent) and answers, BEFORE the credential checks and without counting an auth failure, with a static page whose only content is a `postMessage` of the lost path to the parent tab (`default-src 'none'` plus the hash of that one script, `no-store`, `referrer: no-referrer`, no reflected input). It is fenced to paths that are NOT registered routes and never `/api/`, `/ws/` or `/q/`, with one carve‑out: `/` itself, because the landing page masks to exactly `/` and its reload otherwise rendered Codeman's app shell inside the web tab. `/` is admitted only when the request carries neither the `codeman_session` cookie nor an `Authorization` header: nothing in Codeman frames its own root and a sandboxed frame has neither, while a framed `/` that does carry credentials still gets the shell. On a passwordless install no auth hook runs, so the index route applies the same test itself (`isLostWebviewRootFrame`). ⚠️ Known property, accepted rather than mitigated: those headers are trivially set by a non‑browser client, so an unauthenticated caller can distinguish a registered route (401) from a non‑route (200) and enumerate the route table; the routes are public in `docs/api-reference.md`, so nothing is learned. Pinned by `test/webview-auth-exemption.test.ts` (password) and `test/webview-lost-root-frame.test.ts` (passwordless).
- **Sandboxed by default; `allow-same-origin` is an explicit per‑dashboard opt‑in.** A proxied page is same‑origin with Codeman, so without the sandbox its JavaScript could read the Codeman document and call the agent‑spawning API. ⚠️ In BOTH modes the `Authorization` header and the `codeman_session` cookie are stripped before the upstream request, because a trusted (same‑origin) frame makes the browser attach Codeman's own Basic‑auth credentials to every proxied request; forwarding them would hand `CODEMAN_PASSWORD` to the dashboard.
- **Not an open relay, and not a privilege boundary.** `resolveUpstreamUrl()` refuses anything leaving the saved origin, and cross‑origin redirects are handed back unchanged rather than followed. The proxy does reach whatever the SERVER can reach, which is not an escalation for someone who already commands `--dangerously-skip-permissions` agents, but in multi‑user mode it means a non‑admin's dashboard is fetched from the server's network position. Saved URLs are validated to plain http(s) with no embedded credentials, and there is deliberately **no magic‑link path**: terminal output can never create a webview (the mistake the attachment scanner had to be walled off from). The one refused destination class is link‑local and cloud‑metadata addresses (`169.254.0.0/16`, `fe80::/10`, `fd00:ec2::254`, `168.63.129.16`, `100.100.100.200`, `metadata.google.internal`): `webview-egress-policy.ts` refuses them at save time, and `webview-egress.ts` re‑judges the RESOLVED address at connect time through a `lookup` hook on the proxy's undici Agent and on its WebSocket client, so a DNS name pointing into those ranges is refused as well. Loopback and RFC1918 stay allowed on purpose. Capabilities are revoked on logout, admin logout and user deletion, and proxied responses carry `Referrer-Policy: same-origin` so a dashboard cannot hand the capability‑bearing URL to a third‑party host it links.
+12 -10
View File
@@ -2,6 +2,8 @@
> **Status: SHIPPED — deployed to prod + pushed to master, not yet released (2026-06-14).** App Settings → Display → **Plan Usage Limits** (`showPlanUsageLimits`). **Default changed in 1.9.3: desktop now defaults ON, handhelds stay OFF, resolved via `planUsageChipEnabled()`.** The per-device notes further down describing it as opt-in/synced record the original 2026-06-14 shape, not current behavior. Commits `c82f6c8` (feature) → `4d9d93d` (end-to-end fixes) → `eae225b` (per-user reconcile) → `95fb5fc` (init-snapshot replay). Full suite green (2869), CI green. No changeset/version bump yet.
>
> **2026-09-07 rework — the "Injection lifecycle" section below (disk-write reconcile via `applyStatusLineConfig`) is SUPERSEDED and describes the OLD mechanism, kept for history.** That disk write let a Codeman-marked `statusLine.command` in `.claude/settings.local.json` take precedence over the user's own global/project statusline for ANY `claude` run in that directory — including entirely outside Codeman — with no disclosure and no way to undo it (real bug, found 2026-08-31). The exporter is now injected as an EPHEMERAL `claude --settings` CLI flag at spawn (`resolveStatusLineCliCommand`/`ensureStatusLineExporterScript`, hooks-config.ts) — never written to disk — and it WRAPS the user's own real statusline (`findEffectiveUserStatusLineCommand`) rather than replacing it. `showPlanUsageLimits` now doubles as the telemetry COLLECTION switch too: `readPlanUsageTelemetryEnabled()` reads it fresh from `settings.json` at every claude session create/respawn (`TmuxManager.createSession`/`respawnPane`), so it applies uniformly across every claude-creation path — interactive Run, cron, the Ralph Loop API, quick-start — with no per-session state (a Codeman restart cannot silently kill it) and no per-request field on the wire at all. An absent key reads as ON (the reader resolves the default; `GET /api/settings` never writes), and a settings save carries the key only when it flips the chip on that device, so a handheld with the chip off cannot switch collection off for a desktop by saving something unrelated. The exporter prints nothing on failure rather than the bare word `codeman` (discussion #405).
>
> Two surfaces from one `statusLine` callback:
> - **Header chip** (top-right) — account-wide **plan limits**: `5h 35% · 7d 38%`, per-window green/yellow/red.
> - **In-terminal statusline footer** — the **current session's** status: `Opus 4.8 (1M context) in:562,411 out:1,188 ctx:56%`.
@@ -110,33 +112,33 @@ Fixed path (sessionId in the **body**, not the URL) so the auth exemption is an
2. **Fresh load / reconnect:** server stores the latest in `plan-usage-latest.ts`; `getLightState()` includes it as `planUsage`; the per-connection **init snapshot** replays it; `handleInit` paints the chip immediately (authoritative over localStorage). Null until the first telemetry of the process.
3. **Offline / cross-restart:** `restorePlanUsageChip()` reads `localStorage` on load (12h freshness guard).
### 5. Injection lifecycle — works for *any* user, never self-destructs
### 5. Injection lifecycle (SUPERSEDED 2026-09-07 — see header note; kept for history)
The setting `showPlanUsageLimits` is **synced** (in `settings.json`, not a per-device `displayKey`).
- **On toggle** (`PUT /api/settings`, `system-routes.ts`): reconcile the exporter across **all active Claude sessions' working dirs** — inject on enable, remove on disable. Server-side and authoritative, so existing sessions get the footer + feed the chip *immediately*, no new session needed, no dependency on a client's synced localStorage.
- **On session create** (`session-routes.ts`): **ADD-ONLY** — inject when `statusLineTelemetry` is true; **never remove**. Sessions in a repo share one `settings.local.json`, so a single create-with-false (e.g. a client whose synced setting hadn't loaded) must not yank the statusLine out from under other live sessions. Removal happens only via the explicit toggle.
- `applyStatusLineConfig()` is **`isOurs`-guarded** (matches `/api/status-telemetry`), so a user's own hand-authored statusLine is never touched, and it **updates an out-of-date ours-command** so fixes (e.g. `-k`) propagate. **No `CASES_DIR` gate** — runs for linked cases / real repos (where sessions actually run), mirroring `updateCaseModel`.
- ~~**On toggle** (`PUT /api/settings`, `system-routes.ts`): reconcile the exporter across **all active Claude sessions' working dirs** — inject on enable, remove on disable.~~ There is nothing to (re)inject into an already-running session under the new CLI-flag mechanism — the NEXT respawn (a Ralph cycle, `/clear`, a PTY-exit restart) already reads the setting fresh.
- ~~**On session create** (`session-routes.ts`): **ADD-ONLY** — inject when `statusLineTelemetry` is true; **never remove**.~~ There is no `statusLineTelemetry` request field anymore. `TmuxManager.createSession`/`respawnPane` read `readPlanUsageTelemetryEnabled()` fresh at spawn instead, uniformly across every claude-creation path.
- ~~`applyStatusLineConfig()` is **`isOurs`-guarded**~~ — `applyStatusLineConfig` still exists but only for the SELF-HEAL path now (`resolveStatusLineCliCommand` strips a legacy disk-written exporter the first time a session starts in a workspace an older Codeman build touched).
## Codeman-specific considerations
1. **Account-global limits.** The 5h/7d pools are shared across all sessions on the account → one shared header chip (freshest sample wins), not a per-tab bar.
2. **The footer is owned, by necessity.** A statusLine command always replaces Claude's default footer. Since `rate_limits` *only* arrives via statusLine, we reconstruct a useful **session-status** footer (model · tokens · ctx %) from the same payload rather than showing the limits there.
3. **`isOurs`-guarded.** Never removes/overwrites a user's own statusLine on disable; only manages the Codeman exporter.
3. **Never overwrites, now WRAPS.** The exporter composes with a user's own real statusline (`findEffectiveUserStatusLineCommand`) rather than replacing it; `applyStatusLineConfig`'s `isOurs`-guard now only backs the legacy self-heal removal path.
4. **Security envelope unchanged.** The exporter runs arbitrary shell every render — same trust model as the hook curls (localhost + `$CODEMAN_HOOK_SECRET_FILE`); reuses the hook-secret gate.
5. **Claude-only.** OpenCode/Codex emit no `rate_limits` JSON; injection is gated to `mode === 'claude'`.
5. **Claude-only, registry-gated.** Injection is gated on `getCli(mode)?.capabilities.statusLineTelemetry` (currently `true` only for claude) rather than a hardcoded `mode === 'claude'` string.
6. **Future — auto-resume synergy.** Live percentages would let `SessionAutoOps` pre-arm *before* the wall instead of reacting to the stall footer. Not built.
## Files shipped
- `src/usage-telemetry.ts` — pure parse/format (`parseStatusTelemetry`, `parseSessionStatus`, `formatSessionStatusText`, `telemetrySignature`) + `test/usage-telemetry.test.ts`.
- `src/hooks-config.ts` — `generateStatusLineCommand()` (`curl -sk`), `applyStatusLineConfig()` (add/update/remove, `isOurs`-guarded).
- `src/hooks-config.ts` — `resolveStatusLineCliCommand()`/`ensureStatusLineExporterScript()` (ephemeral CLI-flag injection, never disk), `findEffectiveUserStatusLineCommand()` (wrap the user's real statusline), `readPlanUsageTelemetryEnabled()` (fresh global-setting read), `applyStatusLineConfig()` (legacy self-heal removal only now).
- `src/session-cli-registry-bridge.ts` — merges the exporter path into the SAME `--settings` JSON object as effort/ultracode (Claude Code accepts only one `--settings` flag per invocation).
- `src/web/routes/status-telemetry-routes.ts` — `POST /api/status-telemetry`.
- `src/web/plan-usage-latest.ts` — process-wide last-known store for init replay.
- `src/web/schemas.ts` — `StatusTelemetrySchema` + `showPlanUsageLimits` + create-payload `statusLineTelemetry`.
- `src/web/schemas.ts` — `StatusTelemetrySchema` + `showPlanUsageLimits` (no separate create-payload or action field anymore).
- `src/web/middleware/auth.ts` — exemption extended to `/api/status-telemetry`.
- `src/web/routes/session-routes.ts` — add-only create-time injection.
- `src/web/routes/system-routes.ts` — settings-toggle reconcile.
- `src/tmux-manager.ts` — `createSession`/`respawnPane` read `readPlanUsageTelemetryEnabled()` fresh at spawn.
- `src/web/server.ts` — `getLightState().planUsage` (init snapshot).
- `src/web/sse-events.ts` + `constants.js` — `session:statusTelemetry`.
- Frontend: `app.js` (`_onSessionStatusTelemetry`, `updatePlanUsageChip`, `restorePlanUsageChip`, `handleInit`), `settings-ui.js` (toggle + `applyHeaderVisibilitySettings`), `index.html` (chip + toggle row), `styles.css` (chip + colors), `session-ui.js` (create payload).
+34 -4
View File
@@ -159,6 +159,24 @@ layers cooperate so a dashboard talking to its own backend just works:
using its `Referer` to identify the dashboard. This only fires for a request
that already missed every Codeman route, and never for one that resolves to a
real route, which is what keeps it from being an authentication bypass.
5. The same script **masks the proxy prefix off the page's own URL** before any
of the page's code runs (`history.replaceState` to the path the page would see
on its own origin). A single-page app routes on `location.pathname` at boot,
and `/webview/<cap>/` is a path no app has a route for: without this, a React
Router / Vue Router / Next dev server painted its HTML and CSS and then replaced
them with its own "page not found" the moment its script ran. The page only
*reads* the masked path; every URL it emits still goes through the layers above.
6. A navigation the page starts **itself** after that — `location.reload()` (a dev
server's full-reload HMR), a root-absolute `location.href = '/login'` — now
targets Codeman's root with no capability anywhere on it. Codeman recognises
that request by shape (a top-level `<iframe>` navigation asking for HTML, for a
path it does not serve) and answers a static page that does nothing but tell
the owning tab which path was lost; the tab remounts the frame inside the
prefix at that path. It never counts as a failed login, so a dev server that
reloads on every save cannot rate-limit its user out of Codeman. The landing
page is the one served path that gets the same answer: it masks to exactly
`/`, and a reload there is admitted as long as the request carries no Codeman
credentials, which a sandboxed frame never does.
On top of that, the proxy answers those requests with CORS headers. That sounds
wrong for same-host requests, but a sandboxed iframe has an *opaque* origin, so the
@@ -172,10 +190,22 @@ then every API call fails, which looks like the dashboard being broken.
EventSource, normal markup, the DOM sinks a page uses to build markup at runtime,
and `url()` inside stylesheets. Something that constructs requests by an unusual
route can still slip through. Symptom: the page renders but a panel stays empty.
- **Root-absolute `location` navigation.** A dashboard that navigates itself with
`location.href = '/login'` escapes the prefix, because `Location.href` is
unforgeable and cannot be patched the way the other sinks are. A relative
`location.href = 'login'` is fine (`<base>` covers it).
- **A root-absolute `url()` inside an inline `<style>` is not rescued.** Masking the
page's URL (layer 5) trades away the `Referer` safety net of layer 4 for
requests the shim cannot see, and only HTML is rewritten server-side. An
external stylesheet is fine: a `url()` it references is fetched with the
stylesheet's own URL as `Referer`, which is still inside the prefix. A
root-absolute `url(/img.png)` written directly into a `<style>` block in the
document has the masked document as its `Referer`, so it 404s where the
fallback used to rescue it. Symptom: one background image missing while
everything else renders. Narrow, and a `url()` the page sets from script is
still covered by layer 3.
- **Root-absolute `location` navigation is recovered, not prevented.** `Location`
is unforgeable, so `location.href = '/login'` or `location.reload()` really does
leave the prefix; the frame comes back through the recovery hop in layer 6 above,
which needs a browser that sends `Sec-Fetch-Dest` (every current one; iOS Safari
since 16.4). Older browsers show Codeman's 404 in the frame; the tab's **Reload**
button puts it back.
- **Cross-origin redirects are not followed.** If a dashboard bounces to a different
host (an external SSO provider, say), the proxy hands the redirect back unchanged
rather than relaying it, because relaying would make this an open proxy. Use
+1 -1
View File
@@ -15,7 +15,7 @@ page says so and names the setting.
| **Header, left** | The "C" logo (goes home) and the session list, unless you moved it to the sidebar. |
| **Header, right** | Status chips and panel buttons, most of them off by default. |
| **Center** | The terminal for the active session, or the home screen when nothing is selected. |
| **Bottom toolbar** | Run, Stop, Run Shell, the case picker, and the instance counters. |
| **Bottom toolbar** | Run, Stop, Run Shell, the case picker, and the instance counter. |
| **Overlays** | Panels and modals: Respawn, Cron, Subagents, File Viewer, Settings. |
## Session list layout
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.28.1",
"version": "1.29.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.28.1",
"version": "1.29.0",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
+2 -2
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.28.1",
"version": "1.29.0",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -29,7 +29,7 @@
"test:mobile": "vitest run --config test/mobile/vitest.config.ts",
"check:frontend-syntax": "node scripts/check-frontend-syntax.mjs",
"fix:node-pty": "node scripts/fix-node-pty.mjs",
"typecheck": "tsc --noEmit",
"typecheck": "tsc --noEmit && tsc -p config/tsconfig.scripts.json",
"lint": "eslint --config config/eslint.config.js 'src/**/*.ts'",
"lint:fix": "eslint --config config/eslint.config.js 'src/**/*.ts' --fix",
"format": "prettier --write 'src/**/*.ts' 'src/web/public/**/*.{js,css,html,json}'",
+1 -1
View File
@@ -1,7 +1,7 @@
{
"name": "codeman",
"description": "Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.28.1",
"version": "1.29.0",
"author": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
@@ -0,0 +1,9 @@
{
"_comment": "Copy this file to local-llm-test.config.json (gitignored) and fill in your own values. CLI flags on scripts/test-local-llm-harnesses.mjs always override these. Any field can be omitted. apiKey is OPTIONAL — omit it entirely (or delete this line) for an endpoint like llama.cpp that doesn't check one; it defaults to a harmless placeholder either way.",
"baseUrl": "http://192.168.1.50:8080",
"model": "qwen3",
"apiKey": "",
"prompt": "Reply with exactly: hello world",
"timeout": 30000,
"only": []
}
+23
View File
@@ -193,6 +193,29 @@ run_step "installing" "Installing dependencies" npm install --no-fund --no-audit
# 5) Build (gate the restart on success — never restart into a torn dist/).
run_step "building" "Building" npm run build || rollback_and_fail "Build failed"
# Docker Compose only: record what HEAD/package-lock.json the freshly-built
# codeman-dist/codeman-node-modules volumes now reflect. `Start-Codeman.sh`
# reads this same file (`$appdata_path/.codeman/…`, i.e. this container's own
# $HOME/.codeman since that path IS the appdata bind mount) to detect source
# changes an EXTERNAL `docker compose build` made and refresh those volumes —
# without this, the next plain `Start-Codeman.sh` run would see the HEAD this
# update just checked out, not recognise it as already accounted for, and wipe
# the volumes this update just correctly rebuilt right back to the OLDER image.
if [[ "$SUPERVISOR" == "docker-compose" ]]; then
build_source_file="$HOME/.codeman/docker-build-source.json"
mkdir -p -- "$HOME/.codeman"
build_head=$(git rev-parse HEAD 2>/dev/null || true)
build_lockfile_sha=''
if command -v sha256sum >/dev/null 2>&1; then
build_lockfile_sha=$(sha256sum -- package-lock.json 2>/dev/null | cut -d' ' -f1)
elif command -v shasum >/dev/null 2>&1; then
build_lockfile_sha=$(shasum -a 256 package-lock.json 2>/dev/null | cut -d' ' -f1)
fi
printf '{\n "headCommit": "%s",\n "lockfileSha256": "%s"\n}\n' \
"$build_head" "$build_lockfile_sha" >"$build_source_file.tmp" \
&& mv -- "$build_source_file.tmp" "$build_source_file"
fi
# 6) Restart the service so the new code loads. Write the terminal pre-restart
# marker FIRST so the freshly-booted server can reconcile it deterministically.
write_status "restarting" "Restarting Codeman…"
+699
View File
@@ -0,0 +1,699 @@
#!/usr/bin/env -S npx tsx
/**
* Standalone smoke-test for pointing each Codeman-supported harness CLI at a
* custom OpenAI-compatible endpoint — local (llama.cpp, Ollama, vLLM, ...) or
* cloud (Azure AI Foundry's OpenAI-compatible endpoint, OpenRouter, a
* self-hosted gateway, ...). Anything that answers GET /v1/models and POST
* /v1/chat/completions in the standard shape qualifies; --base-url is not
* assumed to be a LAN address.
*
* This is intentionally OUTSIDE the npm test suite and outside Codeman's own
* session/tmux machinery: it spawns each real CLI binary directly, one-shot,
* with the env vars / config files that CLI's own docs say redirect it to a
* custom endpoint, and checks it can answer "hello world".
*
* DYNAMIC BY DESIGN: this file imports the SAME `enabledClis()` registry and
* `buildCustomModelInjection()` builder the production feature uses (see
* ../src/config/cli-registry/, ../src/custom-model-injection.ts,
* ../src/custom-model-injection-apply.ts) rather than keeping a second,
* hand-maintained copy of each CLI's env vars/config shape. A registry
* change (a new CLI, an edited env var name, a fixed config template) is
* picked up here automatically with zero edits to this file. Only the
* ONE-SHOT INVOCATION FLAGS (how to make each CLI answer one prompt and
* exit — information the registry doesn't model at all, since it only knows
* how to launch the interactive TUI) stay in the small ONE_SHOT table below;
* a CLI newly added to the registry with no ONE_SHOT entry is reported
* UNKNOWN rather than silently skipped or guessed at.
*
* Cloud endpoints often differ from a bare llama.cpp box in two ways this
* script accounts for: (1) auth may be an `api-key` header (Azure's
* convention) rather than `Authorization: Bearer` — see --auth-style below.
* (2) a cloud endpoint's "model" may actually be a deployment name distinct
* from the model family (Azure AI Foundry deployments) — always pass
* --model explicitly for those rather than relying on GET /v1/models
* discovery.
*
* IMPORTANT CONFIDENCE NOTE: claude and opencode are verified end-to-end
* against a real llama-swap server. codex's config STRUCTURE is verified,
* but it only speaks the Responses API (dropped Chat-Completions support
* Feb 2026) — expect it to fail against a plain OpenAI-compatible server,
* that's a real protocol gap, not a bug here. gemini/pi/grok/omp have their
* ONE-SHOT INVOCATION flags confirmed against real installed binaries'
* `--help` output, but their custom-endpoint env/config conventions remain
* web-researched, unverified. deepseek (dsh) is a profile launcher with no
* documented one-shot prompt flag at all — best-effort only. antigravity
* has no known CLI/env/config mechanism (GUI-only per public docs) — its
* registry entry declares `customModelInjection: { kind: 'unsupported' }`,
* which this script picks up dynamically and always skips.
*
* Usage:
* npx tsx scripts/test-local-llm-harnesses.ts --base-url http://192.168.1.50:8080 [options]
* npx tsx scripts/test-local-llm-harnesses.ts --base-url https://<resource>.services.ai.azure.com/openai/v1 --model <deployment-name> --api-key $AZURE_AI_KEY
*
* Options:
* --base-url <url> Required. Root URL of the OpenAI-compatible endpoint (local or cloud).
* --model <name> Model/deployment id to request. Default: first from GET /v1/models.
* --api-key <key> API key to send. Default: local-dummy-key (fine for llama.cpp; required for most cloud endpoints).
* --auth-style <style> "bearer" (default, Authorization: Bearer) or "api-key" (the
* `api-key` header some cloud gateways, e.g. Azure, want).
* NEVER send both — live-tested against a real server, doing
* so reliably HANGS the request indefinitely.
* --prompt <text> Prompt to send. Default: "Reply with exactly: hello world".
* --only <id,id,...> Restrict to these harness ids (comma-separated).
* --timeout <ms> Per-harness spawn timeout. Default: 30000.
* --probe-help Instead of testing, resolve each installed binary and print --help.
* --keep-temp Don't delete generated per-harness config dirs afterward.
* --list Dry run: print the resolved plan per harness, execute nothing.
* -h, --help Show this help.
*/
import { execFileSync, spawn } from 'node:child_process';
import { mkdtempSync, rmSync, readFileSync, existsSync } from 'node:fs';
import { tmpdir, homedir } from 'node:os';
import { join, delimiter, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
import { enabledClis } from '../src/config/cli-registry/index.js';
import type { CliEntry } from '../src/config/cli-registry/types.js';
import {
buildCustomModelInjection,
GROK_CUSTOM_MODEL_NAME,
type CustomModelEndpoint,
} from '../src/custom-model-injection.js';
import { applyConfigDirInjection } from '../src/custom-model-injection-apply.js';
const TAG = '[test-local-llm-harnesses]';
const SCRIPT_DIR = dirname(fileURLToPath(import.meta.url));
const CONFIG_PATH = join(SCRIPT_DIR, 'local-llm-test.config.json');
const CONFIG_EXAMPLE_PATH = join(SCRIPT_DIR, 'local-llm-test.config.example.json');
type AuthStyle = 'bearer' | 'api-key';
interface ConfigDefaults {
baseUrl?: string | null;
model?: string | null;
apiKey?: string;
authStyle?: AuthStyle;
prompt?: string;
only?: string[] | null;
timeout?: number;
}
/**
* Loads scripts/local-llm-test.config.json (gitignored — real IP/model/key,
* per-machine) if present, so you don't have to retype --base-url every run.
* See local-llm-test.config.example.json (tracked) for the shape. CLI flags
* always override whatever this file sets; this only supplies defaults.
*/
function loadConfigFile(): ConfigDefaults {
if (!existsSync(CONFIG_PATH)) return {};
try {
const raw = JSON.parse(readFileSync(CONFIG_PATH, 'utf8'));
return {
baseUrl: raw.baseUrl ?? null,
model: raw.model ?? null,
apiKey: raw.apiKey || undefined, // empty string counts as "not set", not a real key
authStyle: raw.authStyle === 'api-key' ? 'api-key' : undefined, // never 'both'
prompt: raw.prompt ?? undefined,
only: Array.isArray(raw.only) && raw.only.length ? raw.only : null,
timeout: typeof raw.timeout === 'number' ? raw.timeout : undefined,
};
} catch (err) {
console.error(`${TAG} failed to parse ${CONFIG_PATH}: ${(err as Error).message} (ignoring it)`);
return {};
}
}
interface Opts {
baseUrl: string | null;
model: string | null;
apiKey: string;
authStyle: AuthStyle;
prompt: string;
only: string[] | null;
timeout: number;
probeHelp: boolean;
keepTemp: boolean;
list: boolean;
help: boolean;
}
function parseArgs(argv: string[], configDefaults: ConfigDefaults): Opts {
const opts: Opts = {
baseUrl: configDefaults.baseUrl ?? null,
model: configDefaults.model ?? null,
apiKey: configDefaults.apiKey ?? 'local-dummy-key',
authStyle: configDefaults.authStyle ?? 'bearer',
prompt: configDefaults.prompt ?? 'Reply with exactly: hello world',
only: configDefaults.only ?? null,
timeout: configDefaults.timeout ?? 30000,
probeHelp: false,
keepTemp: false,
list: false,
help: false,
};
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
switch (a) {
case '--base-url':
opts.baseUrl = argv[++i];
break;
case '--model':
opts.model = argv[++i];
break;
case '--api-key':
opts.apiKey = argv[++i];
break;
case '--auth-style':
opts.authStyle = argv[++i] as AuthStyle;
if (opts.authStyle !== 'bearer' && opts.authStyle !== 'api-key') {
console.error(`${TAG} --auth-style must be "bearer" or "api-key"`);
opts.help = true;
}
break;
case '--prompt':
opts.prompt = argv[++i];
break;
case '--only':
opts.only = argv[++i]
.split(',')
.map((s) => s.trim())
.filter(Boolean);
break;
case '--timeout':
opts.timeout = Number(argv[++i]);
break;
case '--probe-help':
opts.probeHelp = true;
break;
case '--keep-temp':
opts.keepTemp = true;
break;
case '--list':
opts.list = true;
break;
case '-h':
case '--help':
opts.help = true;
break;
default:
console.error(`${TAG} unknown argument: ${a}`);
opts.help = true;
}
}
return opts;
}
function printUsage(): void {
console.log(`Usage: npx tsx scripts/test-local-llm-harnesses.ts [--base-url <url>] [options]
Reads defaults from scripts/local-llm-test.config.json if it exists (copy
scripts/local-llm-test.config.example.json to create it — gitignored, since
it holds a real IP/model/key). CLI flags always override the config file.
--base-url becomes optional once that file supplies one.
Works against any custom OpenAI-compatible endpoint, local or cloud
(llama.cpp, Ollama, vLLM, Azure AI Foundry, OpenRouter, a self-hosted
gateway, ...) — anything answering GET /v1/models and POST
/v1/chat/completions in the standard shape.
Options:
--base-url <url> Required. Root URL of the OpenAI-compatible endpoint.
--model <name> Model/deployment id to request. Default: first from GET /v1/models.
--api-key <key> API key to send. Default: local-dummy-key (required for most cloud endpoints).
--auth-style <style> "bearer" (default) or "api-key" (Azure-style). Never both — sending
both headers together reliably hangs some real servers.
--prompt <text> Prompt to send. Default: "Reply with exactly: hello world".
--only <id,id,...> Restrict to these harness ids.
--timeout <ms> Per-harness spawn timeout. Default: 30000.
--probe-help Print each installed binary's --help instead of testing.
--keep-temp Keep generated per-harness config dirs afterward.
--list Dry run: print the resolved plan, execute nothing.
-h, --help Show this help.
Harness ids are read from the CLI registry at run time — pass an unknown
one and the error message lists what's actually enabled right now.
Examples:
npx tsx scripts/test-local-llm-harnesses.ts --base-url http://192.168.1.50:8080
npx tsx scripts/test-local-llm-harnesses.ts --base-url https://<resource>.services.ai.azure.com/openai/v1 --model <deployment-name> --api-key $AZURE_AI_KEY`);
}
const HOME = homedir();
/** Expands a leading `~` the way the CLI registry's own search dirs are written. */
function expandHome(p: string): string {
if (p === '~') return HOME;
if (p.startsWith('~/')) return join(HOME, p.slice(2));
return p;
}
function pathWithExtraDirs(extraDirs: string[]): string {
return [...extraDirs.map(expandHome), '/usr/local/bin', process.env.PATH ?? ''].join(delimiter);
}
/** Resolve a binary by trying `<bin> --version` with the CLI's own registry search dirs prefixed onto PATH. */
function resolveBinary(bin: string, searchDirs: string[]): string | null {
try {
execFileSync(bin, ['--version'], {
timeout: 5000,
stdio: 'pipe',
env: { ...process.env, PATH: pathWithExtraDirs(searchDirs) },
});
return bin;
} catch (err) {
// Some CLIs (e.g. dsh) don't support --version cleanly for identity but
// still exist on PATH; a non-ENOENT failure still counts as "found".
if (err && (err as NodeJS.ErrnoException).code === 'ENOENT') return null;
return bin;
}
}
function printHelp(bin: string, searchDirs: string[]): void {
try {
const out = execFileSync(bin, ['--help'], {
timeout: 5000,
stdio: 'pipe',
env: { ...process.env, PATH: pathWithExtraDirs(searchDirs) },
});
console.log(out.toString());
} catch (err) {
const e = err as { stdout?: Buffer; message?: string };
console.log((e.stdout ?? e.message ?? String(err)).toString());
}
}
// --- one-shot invocation table (NOT in the registry — genuinely separate info) ---
type Confidence = 'verified' | 'researched' | 'unknown';
interface OneShot {
/** `modelId` is the RAW model/deployment id (e.g. "qwen3.5-0.8b-...") — CLIs whose
* config wraps it under a provider/block name (pi/omp's "custom/<id>", grok's fixed
* block name) build the full `--model` value here, not in the injection layer. */
argv: (prompt: string, modelId: string) => string[];
confidence: Confidence;
note?: string;
}
/**
* How to make each CLI answer ONE prompt and exit. The registry has no concept
* of this (it only knows the interactive TUI launch line), so this table is
* necessarily hand-maintained — but it is the ONLY hand-maintained part left;
* everything about WHERE the prompt goes (env vars, config files) comes from
* the real registry + `buildCustomModelInjection()` above.
*
* A CLI enabled in the registry with no entry here reports UNKNOWN rather
* than being silently skipped or guessed at — see `resolveOneShot()`.
*/
const ONE_SHOT: Record<string, OneShot> = {
claude: {
confidence: 'verified',
// Claude Code's async session-title-generation call also uses
// ANTHROPIC_DEFAULT_HAIKU_MODEL and validates it against Claude's OWN internal
// recognized-model list, printing [claude-code:unrecognized_model] to stderr for
// a local model name. Confirmed live: `--settings '{"autoTitle":false}'` does NOT
// stop it (still hung the whole run); `--bare` does — the warning still prints,
// but the actual prompt now runs and returns the real answer. Confirmed against
// a real llama-swap server. ⚠️ `--bare` also disables hooks/LSP/plugin sync/
// CLAUDE.md auto-discovery — fine for this ISOLATED one-shot test, never safe to
// apply to a real interactive Codeman session (which needs hooks).
argv: (prompt) => ['--dangerously-skip-permissions', '--bare', '-p', prompt],
},
opencode: { confidence: 'verified', argv: (prompt) => ['run', prompt] },
codex: {
confidence: 'verified',
note: 'config STRUCTURE verified; codex only speaks the Responses API (dropped Chat-Completions Feb 2026) — expect FAIL against a plain OpenAI-compatible server, that is a protocol gap, not a bug here.',
argv: (prompt) => ['exec', '--dangerously-bypass-approvals-and-sandbox', prompt],
},
gemini: {
confidence: 'researched',
// --skip-trust: without it, an untrusted-folder check silently overrides
// --approval-mode yolo back to 'default' (confirmed live: "Approval mode
// overridden to 'default' because the current folder is not trusted").
argv: (prompt) => ['-p', prompt, '--approval-mode', 'yolo', '--skip-trust'],
},
pi: {
confidence: 'verified',
// --model custom/<id>: without an explicit --model, pi uses its own default
// provider (not our injected "custom" one) and fails with "No API key found
// for the selected model" — confirmed live. "custom" matches the provider name
// pi-models-json writes in custom-model-injection.ts. Verified end-to-end
// against a real llama-swap server after two real bugs were found and fixed:
// pi's `models` field must be an ARRAY of `{id}` objects (an object keyed by
// id silently loaded zero models), and PI_CONFIG_DIR does nothing for pi at
// all (grepped pi's own bundled source — not present anywhere); the actual
// working redirect is the CHILD PROCESS's `HOME` itself, since pi hardcodes
// `~/.pi/agent/models.json` with no dedicated override.
argv: (prompt, modelId) => ['--approve', '--model', `custom/${modelId}`, '-p', prompt],
},
grok: {
confidence: 'verified',
// -m <block name>: grok's config.toml (grok-toml template) declares the custom
// model under a fixed [model.<name>] block; GROK_CUSTOM_MODEL_NAME is that same
// name, imported from custom-model-injection.ts so the two can never drift apart.
// Verified end-to-end against a real llama-swap server after correcting the
// ORIGINAL recipe, which was wrong (env vars, not a config file — see the
// customModelInjection comment on grok's registry entry).
argv: (prompt) => ['--always-approve', '-m', GROK_CUSTOM_MODEL_NAME, '-p', prompt],
},
deepseek: {
confidence: 'unknown',
note: 'dsh is a profile launcher, not a documented one-shot prompt flag. Best-effort only.',
argv: (prompt) => ['--profile', 'headless', prompt],
},
omp: {
confidence: 'verified',
// --model custom/<id>: same reasoning as pi — omp's own default model has no
// credential, so without an explicit --model it never reaches our injected
// provider at all. Verified end-to-end against a real llama-swap server after
// the same two fixes as pi (array-shaped `models`, HOME-redirect instead of
// PI_CONFIG_DIR — omp hardcodes `~/.omp/agent/models.yml`).
argv: (prompt, modelId) => ['--model', `custom/${modelId}`, '-p', prompt],
},
};
// --- baseline server check ---------------------------------------------------
async function baselineCheck(
baseUrl: string,
apiKey: string,
authStyle: AuthStyle,
model: string | null,
prompt: string,
timeoutMs: number
): Promise<string> {
console.log(`\n=== Step 0: baseline check against ${baseUrl} (auth: ${authStyle}) ===`);
// Exactly ONE header, never both. An earlier version sent both auth conventions
// (Bearer + api-key) on the theory that an unused header is harmless — live-
// tested against a real llama-swap server, sending both reliably HUNG the
// request indefinitely (reproduced 3x: Bearer alone ~500ms, api-key alone
// ~600ms, both together no response inside a 15s timeout). Use --auth-style
// api-key for endpoints that specifically want that header (e.g. Azure AI
// Foundry); default 'bearer' covers everything else.
const authHeaders: Record<string, string> =
authStyle === 'api-key' ? { 'api-key': apiKey } : { Authorization: `Bearer ${apiKey}` };
let discoveredModel = model;
try {
const res = await fetch(`${baseUrl}/v1/models`, {
headers: authHeaders,
signal: AbortSignal.timeout(timeoutMs),
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = (await res.json()) as { data?: Array<{ id: string }> };
const ids: string[] = (body.data ?? []).map((m) => m.id);
console.log(`GET /v1/models -> ${ids.length ? ids.join(', ') : '(empty list)'}`);
if (!discoveredModel && ids.length) discoveredModel = ids[0];
} catch (err) {
console.error(`${TAG} GET /v1/models failed: ${(err as Error).message}`);
console.error(`${TAG} Is the server actually running at ${baseUrl}? Aborting.`);
process.exit(1);
}
if (!discoveredModel) {
console.error(`${TAG} No --model given and none discovered from /v1/models. Aborting.`);
process.exit(1);
}
// Live-tested against a real llama-swap server: a POST issued right after a GET on
// the same Node process reliably HANGS indefinitely (reproduced repeatedly — GET
// alone ~30ms, POST alone ~1-2s, GET-then-immediate-POST times out completely; a
// 2s pause between them fixed it every time). This looks like Node's fetch (undici)
// reusing a pooled keep-alive connection the server doesn't handle cleanly for a
// second request right behind a first. A short pause is the simplest portable fix
// (no extra deps, no need for undici's Agent/dispatcher API).
await new Promise((resolve) => setTimeout(resolve, 2000));
try {
const res = await fetch(`${baseUrl}/v1/chat/completions`, {
method: 'POST',
headers: { 'Content-Type': 'application/json', ...authHeaders },
body: JSON.stringify({
model: discoveredModel,
messages: [{ role: 'user', content: prompt }],
}),
signal: AbortSignal.timeout(timeoutMs),
});
if (!res.ok) throw new Error(`HTTP ${res.status}: ${await res.text()}`);
const body = (await res.json()) as { choices?: Array<{ message?: { content?: string } }> };
const reply: string = body.choices?.[0]?.message?.content ?? '';
if (!reply.trim()) throw new Error('empty reply');
console.log(`POST /v1/chat/completions -> "${reply.trim().slice(0, 200)}"`);
console.log('Server baseline: PASS\n');
} catch (err) {
console.error(`${TAG} POST /v1/chat/completions failed: ${(err as Error).message}`);
console.error(`${TAG} Server responded to /v1/models but not to a chat request. Aborting.`);
process.exit(1);
}
return discoveredModel;
}
// --- per-harness run ----------------------------------------------------------
interface ChildResult {
code: number | null;
stdout: string;
stderr: string;
timedOut: boolean;
}
function runChild(bin: string, argv: string[], env: Record<string, string>, searchDirs: string[], timeoutMs: number) {
return new Promise<ChildResult>((resolve) => {
let stdout = '';
let stderr = '';
let settled = false;
const child = spawn(bin, argv, {
env: { ...process.env, ...env, PATH: pathWithExtraDirs(searchDirs) },
stdio: ['ignore', 'pipe', 'pipe'],
});
const timer = setTimeout(() => {
if (settled) return;
settled = true;
child.kill('SIGKILL');
resolve({ code: null, stdout, stderr, timedOut: true });
}, timeoutMs);
child.stdout.on('data', (d) => (stdout += d.toString()));
child.stderr.on('data', (d) => (stderr += d.toString()));
child.on('error', (err) => {
if (settled) return;
settled = true;
clearTimeout(timer);
resolve({ code: null, stdout, stderr: `${stderr}\n${err.message}`, timedOut: false });
});
child.on('close', (code) => {
if (settled) return;
settled = true;
clearTimeout(timer);
resolve({ code, stdout, stderr, timedOut: false });
});
});
}
interface HarnessResult {
id: string;
confidence: Confidence | 'unsupported' | 'no-one-shot-recipe';
status: 'PASS' | 'FAIL' | 'UNCONFIRMED' | 'SKIP' | 'LIST';
detail: string;
}
async function runHarness(
entry: CliEntry,
opts: Opts,
model: string,
endpoint: CustomModelEndpoint
): Promise<HarnessResult> {
const id = entry.id;
const injectionCap = entry.capabilities.customModelInjection;
// Dynamic: driven by the REGISTRY's own capability, not a hardcoded id check.
// A future CLI declared unsupported is skipped automatically, same as antigravity today.
if (injectionCap.kind === 'unsupported') {
return {
id,
confidence: 'unsupported',
status: 'SKIP',
detail: 'no known custom-model mechanism (registry: unsupported)',
};
}
const oneShot = ONE_SHOT[id];
if (!oneShot) {
return {
id,
confidence: 'no-one-shot-recipe',
status: 'SKIP',
detail:
'registry supports custom-model injection for this CLI, but this script has no ONE_SHOT invocation entry yet — add one to test it',
};
}
const binary = entry.discovery.binaries[0] ?? id;
const searchDirs = entry.discovery.searchDirs;
const resolved = resolveBinary(binary, searchDirs);
if (!resolved) {
return {
id,
confidence: oneShot.confidence,
status: 'SKIP',
detail: `binary "${binary}" not found on PATH or search dirs`,
};
}
// The REAL injection logic — same function the production route calls.
const injection = buildCustomModelInjection(entry, endpoint, model);
let env: Record<string, string> = {};
let tempDir: string | null = null;
if (injection.kind === 'env') {
env = injection.envOverrides;
} else if (injection.kind === 'configDir') {
tempDir = mkdtempSync(join(tmpdir(), `codeman-local-llm-test-${id}-`));
env = applyConfigDirInjection(tempDir, injection);
}
// injection.kind === 'unsupported' already handled via injectionCap above.
const argv = oneShot.argv(opts.prompt, model);
if (opts.list) {
const detail = `${binary} ${argv.join(' ')} | env: ${Object.keys(env).join(', ')}${tempDir ? ` | configDir: ${tempDir}` : ''}`;
if (tempDir && !opts.keepTemp) rmSync(tempDir, { recursive: true, force: true });
return { id, confidence: oneShot.confidence, status: 'LIST', detail };
}
const { code, stdout, stderr, timedOut } = await runChild(binary, argv, env, searchDirs, opts.timeout);
let detailSuffix = '';
if (tempDir && !opts.keepTemp) rmSync(tempDir, { recursive: true, force: true });
else if (tempDir) detailSuffix = ` [config kept at ${tempDir}]`;
if (timedOut) {
return {
id,
confidence: oneShot.confidence,
status: 'FAIL',
detail: `timed out after ${opts.timeout}ms. stderr: ${stderr.slice(-300)}${detailSuffix}`,
};
}
const reply = stdout.trim();
const matched = /hello/i.test(reply) && /world/i.test(reply);
const softStatus: HarnessResult['status'] = oneShot.confidence === 'verified' ? 'FAIL' : 'UNCONFIRMED';
if (code !== 0) {
return {
id,
confidence: oneShot.confidence,
status: softStatus,
detail: `exit ${code}. stderr: ${stderr.trim().slice(-300) || '(empty)'}${detailSuffix}`,
};
}
if (!reply) {
return { id, confidence: oneShot.confidence, status: softStatus, detail: `exit 0 but empty stdout${detailSuffix}` };
}
if (matched) {
return { id, confidence: oneShot.confidence, status: 'PASS', detail: `${reply.slice(0, 200)}${detailSuffix}` };
}
return {
id,
confidence: oneShot.confidence,
status: 'UNCONFIRMED',
detail: `reply didn't match heuristic, judge by eye: "${reply.slice(0, 300)}"${detailSuffix}`,
};
}
// --- main ---------------------------------------------------------------------
async function main(): Promise<void> {
const configDefaults = loadConfigFile();
const opts = parseArgs(process.argv.slice(2), configDefaults);
if (opts.help) {
printUsage();
process.exit(0);
}
// Dynamic: pulled from the live registry, not a hardcoded id list. `kind === 'agent'`
// excludes 'shell' (no model/endpoint concept). Antigravity stays in this list (it IS
// an enabled agent CLI) — it's the `unsupported` capability check in runHarness that
// skips it, not an exclusion here.
const allEntries = enabledClis().filter((e) => e.kind === 'agent');
const byId = new Map<string, CliEntry>(allEntries.map((e) => [e.id as string, e]));
const ids: string[] = opts.only ?? [...byId.keys()];
const unknownIds = ids.filter((id) => !byId.has(id));
if (unknownIds.length) {
console.error(`${TAG} unknown harness id(s): ${unknownIds.join(', ')}`);
console.error(`${TAG} known ids (from the live CLI registry): ${[...byId.keys()].join(', ')}`);
process.exit(1);
}
const entries = ids.map((id) => byId.get(id)!);
// --probe-help never touches the network — no --base-url needed for it.
if (opts.probeHelp) {
for (const entry of entries) {
const binary = entry.discovery.binaries[0] ?? entry.id;
const resolved = resolveBinary(binary, entry.discovery.searchDirs);
console.log(`\n=== ${entry.id} (${binary}) ===`);
if (!resolved) {
console.log('(not found on PATH or search dirs)');
continue;
}
printHelp(binary, entry.discovery.searchDirs);
}
process.exit(0);
}
if (!opts.baseUrl) {
console.error(`${TAG} --base-url is required (pass it, or set "baseUrl" in ${CONFIG_PATH}).`);
console.error(`${TAG} See ${CONFIG_EXAMPLE_PATH} for the config file shape.\n`);
printUsage();
process.exit(1);
}
opts.baseUrl = opts.baseUrl.replace(/\/+$/, '');
const endpoint: CustomModelEndpoint = {
id: 'standalone-test',
label: 'standalone test',
baseUrl: opts.baseUrl,
apiKey: opts.apiKey,
};
// --list is a pure dry run: never touch the network, even if --model was given.
let model: string;
if (opts.list) {
model = opts.model ?? 'local-model';
console.log(`\n=== Step 0 skipped (--list never hits the network; using placeholder "${model}") ===\n`);
} else {
model = await baselineCheck(opts.baseUrl, opts.apiKey, opts.authStyle, opts.model, opts.prompt, opts.timeout);
}
console.log(`=== Testing ${entries.length} harness(es) ===`);
const results: HarnessResult[] = [];
for (const entry of entries) {
process.stdout.write(`\n--- ${entry.id} ---\n`);
const result = await runHarness(entry, opts, model, endpoint);
results.push(result);
console.log(`${result.status}: ${result.detail}`);
}
console.log('\n=== Summary ===');
const width = Math.max(...results.map((r) => r.id.length)) + 2;
for (const r of results) {
console.log(`${r.id.padEnd(width)} [${r.confidence.padEnd(20)}] ${r.status.padEnd(11)} ${r.detail.slice(0, 100)}`);
}
const hardFail = results.some((r) => r.status === 'FAIL' && r.confidence === 'verified');
if (hardFail) {
console.error(
`\n${TAG} at least one VERIFIED harness FAILed — that's a real regression, not just an unconfirmed guess.`
);
process.exit(1);
}
process.exit(0);
}
main().catch((err) => {
console.error(`${TAG} unexpected error:`, err);
process.exit(1);
});
+124 -12
View File
@@ -10,10 +10,12 @@ import { randomUUID } from 'node:crypto';
import { realpathSync } from 'node:fs';
import fs from 'node:fs/promises';
import { basename, extname, isAbsolute } from 'node:path';
import { isBlockedAttachmentPath, loadAttachmentGuardConfig } from './config/attachment-guard.js';
import { isBlockedAttachmentPath, isUnderTree, loadAttachmentGuardConfig } from './config/attachment-guard.js';
import { EDITABLE_EXTENSIONS } from './config/file-editing.js';
import { validateSessionFilePath } from './web/route-helpers.js';
import { remoteProbePaths, RemoteFileAccessError, type RemoteProbe } from './remote-files.js';
import type { AttachmentDetectedEvent, AttachmentDetectedType } from './types.js';
import type { SessionRemote } from './types/session.js';
/**
* Playable media extensions, single-sourced here because the WORKSPACE preview
@@ -215,6 +217,106 @@ export interface RegisterExternalAttachmentOptions {
* `codeman attach` CLI (which POSTs directly when a session id is known).
*/
forceWorkspaceConfinement?: boolean;
/**
* Remote (SSH) case: the path exists on the REMOTE host, so it is resolved and
* stat'ed there (`remoteProbePaths`) instead of with local `realpathSync`/`fs.stat`,
* which cannot see it at all (#415). A file outside the case directory is
* unreachable exactly like a file inside it.
*
* `sessionWorkingDir` must then be the REMOTE path too, and the workspace
* confinement check (when active) compares against the remotely canonicalized root,
* so a symlinked `remotePath` does not refuse every registration.
*/
remote?: SessionRemote;
/**
* Remote only: `[file, workspaceRoot]` probes a caller already resolved in a BATCHED
* `remoteProbePaths` call (the attachment-history list does one round trip for the
* whole history). Skips this registration's own ssh probe; every guard below still
* runs on the same resolved path it would have produced itself.
*/
remoteProbes?: readonly [RemoteProbe | null, RemoteProbe | null];
}
/**
* A path an attachment request resolved to, on whichever host it lives — the local
* filesystem or the remote host of a remote-SSH case. The rest of
* {@link registerExternalAttachment} (guards, extension allowlist, registry) is then
* host-agnostic: it only ever sees canonical absolute paths and numbers.
*/
interface ResolvedAttachmentFile {
resolvedPath: string;
size: number;
mtimeMs: number;
isFile: boolean;
extension: string;
/** Remote only: the workspace root, with symlinks resolved on the remote host. */
workspaceRoot?: string;
}
/** `extension` the way the attachment registry defines it (no dot, lowercased). */
function attachmentExtensionOf(path: string): string {
return extname(path).toLowerCase().replace(/^\./, '');
}
/** Local resolution: the historical realpath + stat. */
async function resolveLocalAttachment(requestedPath: string): Promise<ResolvedAttachmentFile> {
let resolvedPath: string;
try {
resolvedPath = realpathSync(requestedPath);
} catch {
throw new AttachmentRegistrationError('Attachment file not found', 404);
}
const stat = await fs.stat(resolvedPath);
return {
resolvedPath,
size: stat.size,
mtimeMs: stat.mtimeMs ?? 0,
isFile: typeof stat.isFile === 'function' ? stat.isFile() : true,
extension: attachmentExtensionOf(resolvedPath),
};
}
/**
* Remote resolution for a remote-SSH case: ONE ssh round trip returns the
* symlink-resolved path, the size/mtime and the kind, for the file AND (when a
* workspace is known) its root, which the confinement check compares against.
*/
async function resolveRemoteAttachment(
requestedPath: string,
remote: SessionRemote,
sessionWorkingDir?: string,
preResolved?: readonly [RemoteProbe | null, RemoteProbe | null]
): Promise<ResolvedAttachmentFile> {
const paths = sessionWorkingDir ? [requestedPath, sessionWorkingDir] : [requestedPath];
let probes: ReadonlyArray<RemoteProbe | null>;
if (preResolved) {
probes = preResolved;
} else {
try {
probes = await remoteProbePaths(remote, paths);
} catch (err) {
// 502 marks the TRANSPORT as the failure, distinct from the file's own 404/403,
// so a history listing can report the entry as unknown rather than missing.
throw new AttachmentRegistrationError(
err instanceof RemoteFileAccessError ? err.message : 'remote host unreachable',
502
);
}
}
const [probe, rootProbe] = probes;
if (!probe) {
throw new AttachmentRegistrationError('Attachment file not found', 404);
}
return {
resolvedPath: probe.realPath,
size: probe.size,
mtimeMs: probe.mtimeMs,
isFile: probe.kind === 'file',
extension: attachmentExtensionOf(probe.realPath),
workspaceRoot: rootProbe?.realPath,
};
}
export async function registerExternalAttachment(
@@ -226,12 +328,9 @@ export async function registerExternalAttachment(
throw new AttachmentRegistrationError('Attachment path must be an absolute local path');
}
let resolvedPath: string;
try {
resolvedPath = realpathSync(requestedPath);
} catch {
throw new AttachmentRegistrationError('Attachment file not found', 404);
}
const resolved = await (options.remote
? resolveRemoteAttachment(requestedPath, options.remote, options.sessionWorkingDir, options.remoteProbes)
: resolveLocalAttachment(requestedPath));
// COD-53: enforce the active attachment-guard policy on the symlink-resolved
// path before doing anything else.
@@ -243,7 +342,10 @@ export async function registerExternalAttachment(
// the caller forces it for this registration (the magic-link scanner — see
// forceWorkspaceConfinement). Strictly more restrictive than the blocklist.
const workingDir = options.sessionWorkingDir;
if (!workingDir || !validateSessionFilePath(workingDir, resolvedPath)) {
const confined = options.remote
? !!workingDir && isUnderTree(resolved.resolvedPath, resolved.workspaceRoot ?? workingDir)
: !!workingDir && !!validateSessionFilePath(workingDir, resolved.resolvedPath);
if (!confined) {
throw new AttachmentRegistrationError('Access to this file is blocked', 403);
}
}
@@ -253,20 +355,30 @@ export async function registerExternalAttachment(
// operator-configured extra trees. Symlinks are already resolved above.
// Cross-workspace attachment of non-blocked files stays allowed, so
// codeman-publish and the ~/.codeman review loop keep working.
if (isBlockedAttachmentPath(resolvedPath, guard.blockedTrees)) {
//
// The list is a pattern list over ABSOLUTE paths, so it is host-agnostic and holds
// for a remote path exactly as it does for a local one, with ONE exception worth
// knowing: `isSensitivePath`'s three home-anchored members (`~/.claude.json`,
// `~/.claude/settings.json`, `~/.claude/settings.local.json`) resolve against THIS
// host's `homedir()`, so on a remote host with a different home they do not match.
// Everything else in that list is depth-anchored (`/.ssh/`, `/.aws/credentials`,
// `/.claude/.credentials.json`, ...) and applies unchanged.
if (isBlockedAttachmentPath(resolved.resolvedPath, guard.blockedTrees)) {
throw new AttachmentRegistrationError('Access to this file is blocked', 403);
}
const extension = extname(resolvedPath).toLowerCase().replace(/^\./, '');
const resolvedPath = resolved.resolvedPath;
const extension = resolved.extension;
if (!isSupportedAttachmentExtension(extension)) {
throw new AttachmentRegistrationError('Unsupported attachment type');
}
const stat = await fs.stat(resolvedPath);
if (typeof stat.isFile === 'function' && !stat.isFile()) {
if (!resolved.isFile) {
throw new AttachmentRegistrationError('Attachment path is not a file');
}
const stat = { size: resolved.size, mtimeMs: resolved.mtimeMs };
const existing = attachmentRegistry.findByFilePath(sessionId, resolvedPath);
if (existing) {
existing.size = stat.size;
+44
View File
@@ -261,6 +261,19 @@ const echoSchema = z
})
.strict();
/**
* `capabilities.customModelInjection.launchModel`: the `model` launch-param value that
* selects the injected provider, with `{modelId}` standing for the chosen id. Bounded to
* the characters the `model`/`model-pi` token patterns accept plus the placeholder braces,
* so a template can never smuggle a token the argv engine would have to quote.
*/
const launchModelTemplate = z
.string()
.min(1)
.max(120)
.regex(/^[a-zA-Z0-9._\-/:{}]+$/)
.optional();
const capabilitiesSchema = z
.object({
external: z.boolean(),
@@ -317,6 +330,37 @@ const capabilitiesSchema = z
privilegedEnvKeys: z.array(envName).max(8),
gates: z.record(z.string(), z.object({ minVersion: z.string().max(20), failClosed: z.boolean() }).strict()),
maxFrameBytes: z.number().int().positive().optional(),
customModelInjection: z.discriminatedUnion('kind', [
z
.object({
kind: z.literal('env'),
baseUrlVar: envName,
apiKeyVar: envName,
// Empty is valid: deepseek's model routing is a profile-composition concern, not
// an env var, so it declares baseUrl/apiKey injection with no model var at all.
modelVars: z.array(envName).max(8),
launchModel: launchModelTemplate,
})
.strict(),
z
.object({
kind: z.literal('configContentEnv'),
envVar: envName,
template: z.literal('opencode-json'),
launchModel: launchModelTemplate,
})
.strict(),
z
.object({
kind: z.literal('configDir'),
dirEnvVar: envName,
fileName: z.string().min(1).max(80),
template: z.enum(['codex-toml', 'pi-models-json', 'omp-models-yml', 'grok-toml']),
launchModel: launchModelTemplate,
})
.strict(),
z.object({ kind: z.literal('unsupported') }).strict(),
]),
})
.strict();
+157 -2
View File
@@ -189,6 +189,12 @@ const CLAUDE: CliEntry = {
unset: ['CLAUDECODE'],
tmuxSetenvKeys: [],
dockerExecEnvNames: [],
// Deliberately excludes ANTHROPIC_* (base URL / API key / default-model overrides):
// custom-model-injection.ts's claude recipe uses those names, but they must reach a
// session ONLY through the admin-configured, SSRF-guarded custom-model route, never
// through a plain client-supplied envOverrides field. Widening this prefix would let
// any session-create caller redirect a session's Anthropic traffic and credentials to
// an arbitrary, unvalidated URL.
allowedPrefixes: ['CLAUDE_CODE_'],
allowedKeys: ['CLAUDE_CONFIG_DIR'],
},
@@ -220,8 +226,28 @@ const CLAUDE: CliEntry = {
statusLineTelemetry: true,
model: { source: 'claude-settings-file' },
privilegedParams: [],
privilegedEnvKeys: [],
// ANTHROPIC_* is NOT in allowedPrefixes/allowedKeys above (deliberately — see the
// allowedPrefixes comment nearby), so these are unreachable via plain envOverrides
// today; listed here only so the dedicated custom-model route (docs/custom-model-endpoints-plan.md
// chunk 5) clamps them for a non-granted multi-user owner the same way every other
// CLI's injection vars are clamped, the day that route widens who can set them.
privilegedEnvKeys: [
'ANTHROPIC_BASE_URL',
'ANTHROPIC_API_KEY',
'ANTHROPIC_DEFAULT_SONNET_MODEL',
'ANTHROPIC_DEFAULT_HAIKU_MODEL',
'ANTHROPIC_DEFAULT_OPUS_MODEL',
],
gates: { nameFlag: { minVersion: '2.1.224', failClosed: true } },
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — verified by hand against a real
// llama.cpp server. Claude reads these at process start only, so switching requires a
// respawn, never a live hot-swap.
customModelInjection: {
kind: 'env',
baseUrlVar: 'ANTHROPIC_BASE_URL',
apiKeyVar: 'ANTHROPIC_API_KEY',
modelVars: ['ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL'],
},
},
overlays: {
// Mirrors the local default so the remote/in-container agent runs non-interactively
@@ -287,6 +313,7 @@ const SHELL: CliEntry = {
privilegedParams: [],
privilegedEnvKeys: [],
gates: {},
customModelInjection: { kind: 'unsupported' }, // a raw shell has no "model" concept
},
overlays: {
// No `remote` entry: defaultRemoteCommandForMode special-cases kind==='shell' directly
@@ -366,6 +393,15 @@ const OPENCODE: CliEntry = {
...agentDefaults(),
altScreen: 'strip-mux-only',
echo: { policy: 'buffer', anchor: { kind: 'cursor' }, predictProfile: undefined },
// Verified by hand against a real llama.cpp server. Reuses the SAME env var opencode's
// own `env.configContentVar` already declares — the builder in custom-model-injection.ts
// must merge into whatever opencode config Codeman would otherwise send, not clobber it.
customModelInjection: { kind: 'configContentEnv', envVar: 'OPENCODE_CONFIG_CONTENT', template: 'opencode-json' },
// OPENCODE_CONFIG_CONTENT already matches the OPENCODE_ allowedPrefix above, so it was
// ALREADY reachable via plain envOverrides before this feature existed — it replaces
// opencode's whole config, provider api keys included, so a non-granted multi-user owner
// sending it is a pre-existing credential-redirection gap, not one this feature opens.
privilegedEnvKeys: ['OPENCODE_CONFIG_CONTENT'],
},
overlays: {
credStore: { rel: '.config/opencode', seedWhole: true },
@@ -455,6 +491,23 @@ const CODEX: CliEntry = {
// `dangerouslyBypassApprovals` on the wire), so it is the one that would have caught a
// regression; `schema.ts` now rejects a name that is not a declared param.
privilegedParams: [{ param: 'bypassApprovals', clampTo: false }],
// Verified by hand against a real llama.cpp server. Written to an isolated CODEX_HOME
// so the user's real ~/.codex/config.toml is never touched.
customModelInjection: {
kind: 'configDir',
dirEnvVar: 'CODEX_HOME',
fileName: 'config.toml',
template: 'codex-toml',
},
// CODEX_HOME already matches the CODEX_ allowedPrefix above, so it was ALREADY
// reachable via plain envOverrides before this feature existed. It is arguably
// MORE sensitive than a bare base-url var: a redirected CODEX_HOME points codex at a
// config.toml a non-granted owner fully controls, which can restate sandbox/approval
// policy INSIDE that file — a path the argv-level `bypassApprovals` clamp above
// cannot see or stop.
// CODEMAN_CUSTOM_MODEL_API_KEY: the credential config.toml's env_key references
// (see custom-model-injection.ts) — same reasoning as CODEX_HOME above.
privilegedEnvKeys: ['CODEX_HOME', 'CODEMAN_CUSTOM_MODEL_API_KEY'],
},
overlays: {
credStore: {
@@ -538,6 +591,20 @@ const GEMINI: CliEntry = {
// MATERIALIZE a config (not just touch an already-sent one) or a non-granted owner who
// sends no geminiConfig at all would still get yolo for free.
privilegedParams: [{ param: 'approvalMode', clampTo: 'auto_edit', materializeWhenAbsent: true }],
// Web-researched, unverified — needs a restart to pick up (CLI reads these at process
// start). Confirm the exact model-override env var name against the installed
// gemini-cli version before shipping.
customModelInjection: {
kind: 'env',
baseUrlVar: 'GOOGLE_GEMINI_BASE_URL',
apiKeyVar: 'GEMINI_API_KEY',
modelVars: ['GEMINI_MODEL'],
},
// All three already match the GEMINI_/GOOGLE_ allowedPrefixes above, so they were
// ALREADY reachable via plain envOverrides before this feature existed — a non-granted
// multi-user owner redirecting a gemini session's endpoint/credentials is a
// pre-existing gap this feature's analysis surfaced, not one it opens.
privilegedEnvKeys: ['GOOGLE_GEMINI_BASE_URL', 'GEMINI_API_KEY', 'GEMINI_MODEL'],
},
overlays: {
credStore: { rel: '.gemini', seedWhole: true }, // also covers antigravity — see its own entry
@@ -603,6 +670,10 @@ const ANTIGRAVITY: CliEntry = {
// Like codex: an ABSENT config already defaults safe (no bypass flag), so only a
// SENT config needs the flag forced off — nothing is materialized.
privilegedParams: [{ param: 'dangerouslySkipPermissions', clampTo: false }],
// No known CLI/env/config mechanism — Antigravity's own docs describe a GUI-only
// custom-endpoint setting and explicitly say it "cannot currently" become the core
// reasoning model. Toolbar entry stays disabled for this mode.
customModelInjection: { kind: 'unsupported' },
},
overlays: {
// No credStore of its own: agy nests its whole state under ~/.gemini/antigravity-cli/,
@@ -693,6 +764,34 @@ const PI: CliEntry = {
// just answer "yes" to, so omitting --approve is not itself a clamp — MATERIALIZE
// approveProjectTrust:false so buildPiCommand emits --no-approve outright.
privilegedParams: [{ param: 'approveProjectTrust', clampTo: false, materializeWhenAbsent: true }],
// CORRECTED after live-testing: `PI_CONFIG_DIR` does NOT exist anywhere in pi's own
// bundled source (grepped the installed package directly) — it does nothing for pi
// itself, despite being a real Codeman env var that OTHER things (omp) read. The
// confirmed working redirect is `HOME` itself: pi hardcodes `~/.pi/agent/models.json`
// with no dedicated override, so redirecting the CHILD PROCESS's HOME is what
// actually relocates it (verified: a model written under an isolated HOME's
// `.pi/agent/models.json` shows up in `pi --list-models` and answers a real prompt
// against a real llama-swap server; PI_CONFIG_DIR alone left it silently unable to
// see any provider). ⚠️ This is a bigger blast radius than a dedicated config-dir
// var: it also redirects pi's real sessions/auth/extensions for the DURATION of a
// custom-model session, not just its provider config — document this trade-off
// wherever this capability is surfaced.
customModelInjection: {
kind: 'configDir',
dirEnvVar: 'HOME',
fileName: '.pi/agent/models.json',
template: 'pi-models-json',
// Writing models.json is not enough: without `--model custom/<id>` pi stays on its
// own default provider and fails with "No API key found for the selected model"
// (confirmed live). `custom` is the provider name pi-models-json declares.
launchModel: 'custom/{modelId}',
},
// HOME is not `PI_`-prefixed, so unlike the old (wrong) PI_CONFIG_DIR guess this was
// never reachable via the generic envOverrides allowlist at all — listed here anyway,
// matching the documented pattern for every other CLI's dir-redirect var, since a
// redirected HOME is at least as sensitive as CODEX_HOME/GROK_HOME (pi executes
// repo-local .pi/extensions TypeScript — see the External CLI modes note in CLAUDE.md).
privilegedEnvKeys: ['HOME'],
},
overlays: {
credStore: {
@@ -788,6 +887,28 @@ const GROK: CliEntry = {
// already its safe interactive ask-mode, so the multi-user clamp only needs to force an
// EXPLICITLY-SENT bypass flag back off — nothing is materialized when config is absent.
privilegedParams: [{ param: 'alwaysApprove', clampTo: false }],
// CORRECTED after live-testing against a real grok binary: the original `env` kind
// (GROK_BASE_URL/GROK_MODEL/XAI_API_KEY) produced "Not signed in" — those env vars
// are NOT grok's real custom-endpoint mechanism. The real one (verified against
// xAI's own docs) is a `[model.<name>]` block in a config.toml under GROK_HOME,
// the same configDir shape as codex/pi/omp. `api_backend = "chat_completions"` is
// explicitly supported (unlike codex, which dropped it) — grok CAN talk to a plain
// OpenAI Chat-Completions server directly.
customModelInjection: {
kind: 'configDir',
dirEnvVar: 'GROK_HOME',
fileName: 'config.toml',
template: 'grok-toml',
// The `[model.<name>]` block the grok-toml template writes; `--model <name>` is what
// selects it (GROK_CUSTOM_MODEL_NAME in custom-model-injection.ts, pinned equal by
// test/custom-model-injection.test.ts so the two cannot drift).
launchModel: 'codeman-custom',
},
// GROK_HOME already matches the GROK_ allowedPrefix above, so it was ALREADY
// reachable via plain envOverrides before this feature existed — same reasoning
// as CODEX_HOME: a redirected config dir can restate policy the argv-level
// `alwaysApprove` clamp above cannot see.
privilegedEnvKeys: ['GROK_HOME'],
},
overlays: {
// ~/.grok also holds sessions/, memory/, downloads/ (the ~160MB binary), completions/,
@@ -943,7 +1064,23 @@ const DEEPSEEK: CliEntry = {
// The half no other CLI needs. `DSH_*` is an allowlisted envOverrides prefix and
// applyEnvOverrides() runs LAST, so without this a non-granted owner could send
// DSH_PERMISSION_MODE on the same request and land after the config clamp.
// ⚠️ DEEPSEEK_API_KEY deliberately stays OUT of this list (see the docstring on
// clampEnvOverridesForOwner() in session-routes.ts): _configureCliEnv() forwards the
// SERVER's own key into every dsh pane, so DEEPSEEK_BASE_URL is the exfiltration
// vector, not the key itself — a non-granted owner supplying THEIR OWN key removes
// privilege rather than granting it, and clamping it here was a real regression
// (test/deepseek-mode.test.ts) fixed before this shipped.
privilegedEnvKeys: ['DSH_PERMISSION_MODE', 'DSH_HOME', 'DEEPSEEK_BASE_URL'],
// Web-researched, unverified, partial: reuses the already-existing DEEPSEEK_BASE_URL/
// DEEPSEEK_API_KEY keys above. No modelVars — dsh's model is a profile-composition
// entry (see `model: { source: 'none' }` above), not an env var, so forcing a specific
// model name may not fully work; verify against a real profile before shipping.
customModelInjection: {
kind: 'env',
baseUrlVar: 'DEEPSEEK_BASE_URL',
apiKeyVar: 'DEEPSEEK_API_KEY',
modelVars: [],
},
},
overlays: {
// No credStore: dsh keeps everything under $DSH_HOME (default ~/.dsh), which is
@@ -1045,7 +1182,25 @@ const OMP: CliEntry = {
// Where omp resolves its auth from. No known concrete exfiltration path today (omp
// forwards no operator-held key into a pane), but a non-granted owner redirecting where
// a shared multi-tenant deployment resolves auth is not something to allow silently.
privilegedEnvKeys: ['OMP_AUTH_BROKER_URL', 'OMP_AUTH_BROKER_TOKEN'],
// HOME added for custom-model-injection.ts's omp recipe (see below). Unlike pi,
// PI_CONFIG_DIR genuinely IS one of the env vars omp reads (per the DeepSeek/OMP
// note in CLAUDE.md) — but live-testing this feature found it did NOT relocate
// omp's model config the way expected, while redirecting HOME itself (like pi)
// worked immediately (verified end-to-end: a real "hello world" reply came back).
privilegedEnvKeys: ['OMP_AUTH_BROKER_URL', 'OMP_AUTH_BROKER_TOKEN', 'HOME'],
// Verified end-to-end against a real llama-swap server (live-tested, not just
// researched — a real "hello world" reply came back). Same HOME-redirect mechanism
// as pi (see its customModelInjection comment for the full reasoning) — omp hardcodes
// `~/.omp/agent/models.yml` with no dedicated config-dir override either.
customModelInjection: {
kind: 'configDir',
dirEnvVar: 'HOME',
fileName: '.omp/agent/models.yml',
template: 'omp-models-yml',
// Same as pi: omp's own default model has no credential, so without an explicit
// `--model custom/<id>` it never reaches the injected provider at all.
launchModel: 'custom/{modelId}',
},
},
overlays: {
// `~/.omp/agent` also holds agent.db/history.db/models.db (SQLite caches) and
+51
View File
@@ -457,6 +457,57 @@ export interface CliCapabilities {
gates: Record<string, { minVersion: string; failClosed: boolean }>;
/** Cap on a single terminal frame, when this CLI needs a tighter one than the default. */
maxFrameBytes?: number;
/**
* How this CLI is pointed at a user-supplied custom OpenAI-compatible
* endpoint (local, e.g. llama.cpp, or cloud, e.g. Azure AI Foundry) — the
* Custom Model Endpoint Profiles feature (`docs/custom-model-endpoints-plan.md`). Declared
* per entry, never branched on id, same as every other capability here.
*
* `env`: plain env vars (claude's `ANTHROPIC_BASE_URL`/`ANTHROPIC_API_KEY`/
* `ANTHROPIC_DEFAULT_*_MODEL`). `configContentEnv`: a full config blob
* carried in one env var (opencode's `OPENCODE_CONFIG_CONTENT`).
* `configDir`: a generated config file under an isolated, dir-redirect-env-
* pointed directory so the user's real CLI config is never touched
* (codex's `CODEX_HOME`/`config.toml`, pi/omp's `PI_CONFIG_DIR`, grok's
* `GROK_HOME`/`config.toml`). `unsupported`: no known mechanism
* (antigravity) — the toolbar entry stays disabled for this CLI.
*
* ⚠️ grok was ORIGINALLY declared as `env` kind (`GROK_BASE_URL`/
* `GROK_MODEL`/`XAI_API_KEY`) — that recipe was WRONG, not just unverified:
* live-tested against a real grok binary, it produced "Not signed in",
* because those env vars are not grok's real custom-endpoint mechanism at
* all. The real one is a `[model.<name>]` block in a `config.toml` under
* `GROK_HOME` (verified against xAI's own docs), same shape as codex/pi/
* omp — this is why the confidence table in docs/custom-model-endpoints-plan.md exists:
* "researched" web docs can still be plausible-sounding and wrong.
*
* Every env var name this introduces that can redirect a session's
* traffic MUST also appear in `privilegedEnvKeys` above, exactly like
* `DEEPSEEK_BASE_URL` — a non-granted multi-user owner redirecting a
* session to their own endpoint is a credential-exfiltration path, not
* just a mischief redirect.
*
* `launchModel` is the value the entry's own `model` launch param must carry
* for the CLI to SELECT the injected provider, as a template where
* `{modelId}` is the chosen model id. Writing the config file is not enough
* for pi and omp (`--model custom/<id>`, or the CLI stays on its own default
* provider and reports "No API key found for the selected model") or for
* grok (`--model codeman-custom`, the `[model.<name>]` block the config
* declares). Absent = the config alone selects the model (claude's env vars,
* opencode's blob, codex's top-level `model` key). Applied by the session's
* respawn options through the entry's `legacyConfigField`, never by id.
*/
customModelInjection:
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[]; launchModel?: string }
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json'; launchModel?: string }
| {
kind: 'configDir';
dirEnvVar: string;
fileName: string;
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' | 'grok-toml';
launchModel?: string;
}
| { kind: 'unsupported' };
}
// ---------------------------------------------------------------------------
+65
View File
@@ -0,0 +1,65 @@
/**
* @fileoverview Read/write-array store for user-configured custom OpenAI-compatible
* model endpoints (local or cloud — docs/custom-model-endpoints-plan.md). Same
* shape as `remote-hosts.ts` / `webview-store.ts`: `~/.codeman/custom-model-hosts.json`
* holding a plain array, read/written whole. The file can hold API keys, so it is
* written 0600 via tmp+rename like `intents.json` (`mode` on `writeFile` applies only
* to a file being created; the rename is what keeps an existing file's bytes and
* mode from ever being observable half-written or world-readable).
*/
import { existsSync, mkdirSync } from 'node:fs';
import fs from 'node:fs/promises';
import { join } from 'node:path';
const CUSTOM_MODEL_HOSTS_FILE = 'custom-model-hosts.json';
export type CustomModelAuthStyle = 'bearer' | 'api-key';
export interface CustomModelHost {
id: string;
label: string;
/** Root URL, local or cloud — e.g. "http://192.168.1.50:8080" or an Azure AI Foundry URL. */
baseUrl: string;
apiKey?: string;
/**
* Defaults to 'bearer' (the common `Authorization: Bearer` convention — matches
* llama.cpp, OpenAI-compatible servers, and most gateways). Pick 'api-key' for
* endpoints that specifically want the `api-key` header, e.g. Azure AI Foundry.
*
* ⚠️ There is deliberately NO 'both' option. An earlier design sent BOTH headers
* on every discovery request on the theory that an unused header is harmless —
* live-tested against a real llama-swap server, sending both reliably HUNG the
* request indefinitely (reproduced 3× — Bearer alone: ~500ms, api-key alone:
* ~600ms, both together: no response inside a 15s timeout). Whatever auth
* middleware some servers run apparently does not handle two simultaneous
* credential conventions gracefully, so "send everything and let the server
* ignore what it doesn't need" is not a safe default — it can silently turn a
* working endpoint into one that always times out.
*/
authStyle?: CustomModelAuthStyle;
models?: string[];
lastDiscoveredAt?: string;
}
export function customModelHostsPath(configDir: string): string {
return join(configDir, CUSTOM_MODEL_HOSTS_FILE);
}
export async function readCustomModelHosts(configDir: string): Promise<CustomModelHost[]> {
try {
const raw = await fs.readFile(customModelHostsPath(configDir), 'utf-8');
const parsed = JSON.parse(raw);
return Array.isArray(parsed) ? (parsed as CustomModelHost[]) : [];
} catch {
return [];
}
}
export async function writeCustomModelHosts(configDir: string, hosts: CustomModelHost[]): Promise<void> {
if (!existsSync(configDir)) mkdirSync(configDir, { recursive: true });
const target = customModelHostsPath(configDir);
const tmp = `${target}.${process.pid}.tmp`;
await fs.writeFile(tmp, JSON.stringify(hosts, null, 2), { mode: 0o600 });
await fs.rename(tmp, target);
}
+96
View File
@@ -0,0 +1,96 @@
/**
* @fileoverview The one IO wrapper around `custom-model-injection.ts`'s pure
* `ConfigDirInjection` output — deliberately split out so that file, the
* discovery routes, and `scripts/test-local-llm-harnesses.ts` (via tsx) can
* all share EXACTLY one "write these files, merge this env" implementation.
* Before this existed, the route and the standalone script each carried
* their own copy of this logic, which is exactly the kind of drift the CLI
* registry's "declare once, consume everywhere" design exists to prevent —
* see docs/custom-model-endpoints-plan.md and the "dynamic to support
* cli-registry changes" requirement it was written against.
*/
import { chmodSync, mkdirSync, writeFileSync, rmSync } from 'node:fs';
import { join, dirname } from 'node:path';
import { dataPath } from './config/instance.js';
import type { CliEntry } from './config/cli-registry/types.js';
import {
buildCustomModelInjection,
type ConfigDirInjection,
type CustomModelEndpoint,
} from './custom-model-injection.js';
/** Where a session's isolated `configDir`-kind files live: never the user's real CLI config path. */
export function customModelConfigDir(sessionId: string): string {
return join(dataPath('custom-model-configs'), sessionId);
}
/**
* Writes a `ConfigDirInjection`'s files under `baseDir` and returns the full
* envOverrides object a caller should merge into the session/process env
* (the dir-redirect var plus any `extraEnv` the config file references by
* name). Never touches anything outside `baseDir` — the caller is
* responsible for choosing an isolated directory (never the user's real
* `~/.codex`, `~/.pi`, etc.).
*
* pi and omp embed the API key literally in the file, so the tree is written
* 0700/0600 like every other secret-bearing file under `~/.codeman`; the chmod
* covers a re-apply onto a file that already exists (`mode` only applies at
* creation).
*/
export function applyConfigDirInjection(baseDir: string, injection: ConfigDirInjection): Record<string, string> {
for (const file of injection.files) {
const filePath = join(baseDir, file.relPath);
mkdirSync(dirname(filePath), { recursive: true, mode: 0o700 });
writeFileSync(filePath, file.content, { encoding: 'utf8', mode: 0o600 });
chmodSync(filePath, 0o600);
}
return { [injection.dirEnvVar]: baseDir, ...injection.extraEnv };
}
/** Best-effort recursive removal of a previously-written configDir. Never throws. */
export function removeConfigDir(dir: string | undefined): void {
if (!dir) return;
try {
rmSync(dir, { recursive: true, force: true });
} catch {
// best-effort cleanup only
}
}
/** What applying an endpoint to a session yields, ready for `Session.setCustomModel()`. */
export interface AppliedCustomModel {
envOverrides: Record<string, string>;
envKeys: string[];
configDir?: string;
launchModel?: string;
}
/**
* Compute (and for the `configDir` kind, write) everything a session needs to run
* against `endpoint`/`modelId`. Returns undefined for a CLI with no mechanism.
*
* Idempotent on purpose: the boot-recovery path calls it again for a session that
* was already pointed at an endpoint, so the config files are rewritten in place
* (same content) and the env values, which are never persisted because they carry
* the API key, are re-derived from the endpoint store instead.
*/
export function applyCustomModelInjection(
entry: Pick<CliEntry, 'capabilities'>,
endpoint: CustomModelEndpoint,
modelId: string,
sessionId: string
): AppliedCustomModel | undefined {
const injection = buildCustomModelInjection(entry, endpoint, modelId);
if (injection.kind === 'unsupported') return undefined;
if (injection.kind === 'env') {
return {
envOverrides: injection.envOverrides,
envKeys: Object.keys(injection.envOverrides),
launchModel: injection.launchModel,
};
}
const configDir = customModelConfigDir(sessionId);
const envOverrides = applyConfigDirInjection(configDir, injection);
return { envOverrides, envKeys: Object.keys(envOverrides), configDir, launchModel: injection.launchModel };
}
+252
View File
@@ -0,0 +1,252 @@
/**
* @fileoverview Pure builder for the Custom Model Endpoint Profiles feature
* (docs/custom-model-endpoints-plan.md): turns a CLI registry entry's
* `capabilities.customModelInjection` declaration, a configured endpoint,
* and a chosen model id into the concrete env vars / config-file content
* that would redirect that CLI's session at the endpoint.
*
* No IO here on purpose (mirrors `session-cli-builder.ts`) — a caller
* writes `ConfigDirInjection.files` to disk under an isolated per-session
* directory and points `dirEnvVar` at it; this module only computes what
* those files/env vars should contain.
*
* Confidence: `claude` and `opencode` are verified end-to-end against a real
* llama-swap server (a real "hello world" reply came back). `codex`'s
* config.toml STRUCTURE is now verified (an earlier `[model].default` table
* shape was rejected by a real codex binary with "invalid type: map,
* expected a string" — caught by `scripts/test-local-llm-harnesses.ts`),
* but `wire_api = "responses"` is the only value codex still accepts
* (support for `"chat"` was dropped in Feb 2026), and a plain OpenAI
* Chat-Completions server (llama.cpp, llama-swap, most local setups) does
* NOT implement the Responses API — so codex may still fail at the
* PROTOCOL level even with a correctly-shaped config file. That gap is
* real and current, not a stale warning; see docs/custom-model-endpoints-plan.md. The rest
* (gemini/pi/grok/deepseek/omp) have their ONE-SHOT INVOCATION flags
* confirmed against real installed binaries' own `--help` output, but
* their custom-endpoint env/config conventions remain web-researched,
* unverified.
*/
import type { CliEntry } from './config/cli-registry/types.js';
export interface CustomModelEndpoint {
id: string;
label: string;
/** Root URL, no trailing slash required — e.g. "http://192.168.1.50:8080" or an Azure AI Foundry URL. */
baseUrl: string;
/** Falls back to a harmless placeholder for endpoints (llama.cpp) that don't check it. */
apiKey?: string;
}
export interface EnvInjection {
kind: 'env';
/** Ready to merge into a session's envOverrides. */
envOverrides: Record<string, string>;
/** See {@link ConfigDirInjection.launchModel}. */
launchModel?: string;
}
export interface ConfigDirInjection {
kind: 'configDir';
/** Env var that must be set to the directory the caller writes `files` under. */
dirEnvVar: string;
files: Array<{ relPath: string; content: string }>;
/**
* Env vars the written config file REFERENCES by name rather than embedding a
* literal value (codex's `env_key = "..."` convention: config.toml never carries
* the API key itself, only the name of an env var codex reads it from). Merge
* these into the session's envOverrides alongside `dirEnvVar` — never skip them,
* or the config points at a credential that was never actually set.
*/
extraEnv?: Record<string, string>;
/**
* The value the CLI's `model` launch param must carry for it to SELECT the injected
* provider (pi/omp: `custom/<modelId>`; grok: the `[model.<name>]` block name). Absent
* when the config alone selects the model. Rendered from the registry entry's
* `customModelInjection.launchModel` template, never hand-built per CLI.
*/
launchModel?: string;
}
export interface UnsupportedInjection {
kind: 'unsupported';
}
export type CustomModelInjectionResult = EnvInjection | ConfigDirInjection | UnsupportedInjection;
const DEFAULT_API_KEY = 'local-dummy-key';
/** Normalizes a base URL to end in exactly one trailing `/v1`, for CLIs whose config expects the OpenAI-style suffix. */
export function withV1Suffix(baseUrl: string): string {
const trimmed = baseUrl.replace(/\/+$/, '');
return /\/v1$/.test(trimmed) ? trimmed : `${trimmed}/v1`;
}
/** JSON-escapes a string for embedding in a TOML/YAML double-quoted scalar — a safe superset of both grammars' basic escapes. */
function quoted(value: string): string {
return JSON.stringify(value);
}
export function buildCustomModelInjection(
entry: Pick<CliEntry, 'capabilities'>,
endpoint: CustomModelEndpoint,
modelId: string
): CustomModelInjectionResult {
const cap = entry.capabilities.customModelInjection;
const apiKey = endpoint.apiKey?.trim() || DEFAULT_API_KEY;
switch (cap.kind) {
case 'env': {
const envOverrides: Record<string, string> = {
[cap.baseUrlVar]: endpoint.baseUrl,
[cap.apiKeyVar]: apiKey,
};
for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId;
return withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId);
}
case 'configContentEnv': {
const content = renderConfigContent(cap.template, endpoint, modelId, apiKey);
return withLaunchModel({ kind: 'env', envOverrides: { [cap.envVar]: content } }, cap.launchModel, modelId);
}
case 'configDir': {
const { content, extraEnv } = renderConfigFile(cap.template, endpoint, modelId, apiKey);
return withLaunchModel(
{ kind: 'configDir', dirEnvVar: cap.dirEnvVar, files: [{ relPath: cap.fileName, content }], extraEnv },
cap.launchModel,
modelId
);
}
case 'unsupported':
return { kind: 'unsupported' };
}
}
/** Render a `launchModel` template (`{modelId}` = the chosen id) onto an injection result. */
function withLaunchModel<T extends EnvInjection | ConfigDirInjection>(
result: T,
template: string | undefined,
modelId: string
): T {
if (!template) return result;
return { ...result, launchModel: template.split('{modelId}').join(modelId) };
}
function renderConfigContent(
template: 'opencode-json',
endpoint: CustomModelEndpoint,
modelId: string,
apiKey: string
): string {
switch (template) {
case 'opencode-json':
return JSON.stringify({
$schema: 'https://opencode.ai/config.json',
provider: {
custom: {
options: { baseURL: withV1Suffix(endpoint.baseUrl), apiKey },
models: { [modelId]: {} },
},
},
model: `custom/${modelId}`,
});
}
}
const CODEX_API_KEY_ENV_VAR = 'CODEMAN_CUSTOM_MODEL_API_KEY';
/** The `[model.<name>]` block name grok's config.toml uses for the injected model — also
* what `-m <name>` in the standalone script's ONE_SHOT argv must reference to select it. */
export const GROK_CUSTOM_MODEL_NAME = 'codeman-custom';
function renderConfigFile(
template: 'codex-toml' | 'pi-models-json' | 'omp-models-yml' | 'grok-toml',
endpoint: CustomModelEndpoint,
modelId: string,
apiKey: string
): { content: string; extraEnv?: Record<string, string> } {
const baseUrl = withV1Suffix(endpoint.baseUrl);
switch (template) {
case 'codex-toml': {
// Verified against real codex (>= Feb 2026): `model` is a top-level STRING, never
// a `[model].default` table — codex rejects that with "invalid type: map, expected
// a string" (caught by scripts/test-local-llm-harnesses.ts against a real llama-swap
// server). The API key is NEVER a literal TOML field: codex's schema only supports
// `env_key`, the NAME of an env var it reads the credential from at runtime, so the
// actual value must ride along as an extra env var, never embedded in the file.
// ⚠️ `wire_api = "responses"` is the only value codex still accepts (it dropped
// `"chat"` support in Feb 2026) — a plain OpenAI Chat-Completions server (llama.cpp,
// llama-swap, most local setups) does NOT implement the Responses API, so this
// recipe may still fail at the PROTOCOL level even though the file now parses
// correctly. That is a real, currently-unresolved compatibility gap, not a syntax
// bug — track it before calling codex support done.
const content = [
`model = ${quoted(modelId)}`,
`model_provider = "custom"`,
'',
'[model_providers.custom]',
`name = "Custom Endpoint"`,
`base_url = ${quoted(baseUrl)}`,
`env_key = ${quoted(CODEX_API_KEY_ENV_VAR)}`,
`wire_api = "responses"`,
'',
].join('\n');
return { content, extraEnv: { [CODEX_API_KEY_ENV_VAR]: apiKey } };
}
case 'pi-models-json':
// Verified against pi's OWN bundled docs (models.md): `models` is an ARRAY of
// `{id: "..."}` objects, NOT an object keyed by model id — the earlier shape here
// silently loaded zero models ("No models available"), confirmed live. `authHeader:
// true` is required too: pi does not automatically send `Authorization: Bearer
// <apiKey>` just because `apiKey` is set (per the same doc) — without it, a real
// (non-llama.cpp) endpoint that actually checks the key would reject every request.
return {
content: JSON.stringify(
{
providers: {
custom: {
baseUrl,
apiKey,
api: 'openai-completions',
authHeader: true,
models: [{ id: modelId }],
},
},
},
null,
2
),
};
case 'omp-models-yml':
// Mirrors the pi-models-json fix above (omp shares pi's config lineage per
// CLAUDE.md — it reads several of pi's own env vars): a flat list of bare model
// name strings under `models` is UNCONFIRMED against real omp docs (none are
// bundled with the binary) — this now matches pi's `{id: "..."}` object-list
// shape and adds `authHeader: true` on the same reasoning, but has not itself
// been live-tested the way pi's fix was. Verify before raising its confidence.
return {
content: `providers:\n custom:\n baseUrl: ${quoted(baseUrl)}\n apiKey: ${quoted(apiKey)}\n api: openai-completions\n authHeader: true\n models:\n - id: ${quoted(modelId)}\n`,
};
case 'grok-toml': {
// Verified against xAI's own docs (docs.x.ai/build/settings/reference): a
// `[model.<name>]` block, NOT plain env vars — an earlier `env`-kind recipe for
// grok was wrong, not just unverified (see the customModelInjection doc comment
// in cli-registry/types.ts). `api_backend = "chat_completions"` is explicitly
// supported (unlike codex, which dropped it after Feb 2026), so this one CAN
// talk to a plain OpenAI-compatible server directly. `env_key` reuses grok's own
// documented fallback var name (XAI_API_KEY) rather than inventing a new one.
const content = [
`[model.${GROK_CUSTOM_MODEL_NAME}]`,
`model = ${quoted(modelId)}`,
`base_url = ${quoted(baseUrl)}`,
`name = "Custom Endpoint"`,
`env_key = "XAI_API_KEY"`,
`api_backend = "chat_completions"`,
'',
].join('\n');
return { content, extraEnv: { XAI_API_KEY: apiKey } };
}
}
}
+58
View File
@@ -293,6 +293,64 @@ export function toSessionDocker(host: DockerHost, dockerCase: DockerCase): Sessi
return session;
}
/**
* Which existing case, if any, blocks adopting `container` at `containerWorkdir`.
*
* One container may back SEVERAL adopted cases, each pointing at a different
* directory inside it — that is the whole reason to adopt the same container
* twice, and it is safe because the in-container tmux session is named per
* SESSION (`dockerTmuxSessionName`, `codeman-dkr-<id8>`) and not per case, so a
* session teardown kills exactly one session and its siblings on the shared
* in-container tmux server are untouched. Nothing else reaches an adopted
* container's lifecycle either: stop/remove throw at the builder, recreate
* refuses `owned === false`, and the orphan reaper filters on the
* `codeman.managed=1` label that only Codeman-created containers carry.
*
* So the conflicts that remain are NOT about the tmux server:
* - `owned-case` the container backs a case Codeman CREATED, whose lifecycle
* it owns; a recreate or delete there would destroy the
* adopted case's container out from under it.
* - `other-owner` already adopted by a different user. Adoption hands out a
* shell inside someone else's container, so it stays scoped.
* - `duplicate` same container AND same directory: the second case would
* behave identically to the first, so name the first instead
* of silently creating a twin. A DIFFERENT directory is the
* supported case and returns null.
*/
export type AdoptContainerConflict =
| { kind: 'owned-case'; caseName: string }
| { kind: 'other-owner'; caseName: string }
| { kind: 'duplicate'; caseName: string }
| null;
export function classifyAdoptContainerConflict(params: {
container: string;
/** Directory inside the container this adoption targets (already defaulted). */
containerWorkdir: string;
existing: ReadonlyArray<
Pick<DockerCase, 'name' | 'container' | 'containerWorkdir' | 'hostWorkspacePath' | 'owned' | 'owner'>
>;
/** Owner visibility test (canAccessOwned bound to the caller). */
canAccess: (owner?: string) => boolean;
}): AdoptContainerConflict {
const { container, containerWorkdir, existing, canAccess } = params;
const sharing = existing.filter((item) => (item.container ?? dockerContainerName(item.name)) === container);
if (sharing.length === 0) return null;
// `owned` is optional and an ABSENT flag means owned (legacy cases predate the
// field), so this must test `!== false` rather than truthiness.
const owned = sharing.find((item) => item.owned !== false);
if (owned) return { kind: 'owned-case', caseName: owned.name };
const foreign = sharing.find((item) => !canAccess(item.owner));
if (foreign) return { kind: 'other-owner', caseName: foreign.name };
const twin = sharing.find((item) => (item.containerWorkdir ?? item.hostWorkspacePath) === containerWorkdir);
if (twin) return { kind: 'duplicate', caseName: twin.name };
return null;
}
/**
* An ADOPTED container is one the user built and runs themselves. Codeman may
* only exec into it; it must never create, start, stop, restart or remove it.
+21 -7
View File
@@ -13,11 +13,14 @@ import { realpathSync } from 'node:fs';
import { homedir } from 'node:os';
import { join, normalize, sep } from 'node:path';
import { registerExternalAttachment, type AttachmentRegistrationResult } from './attachment-registry.js';
import type { SessionRemote } from './types/session.js';
export interface GeneratedArtifactRegistrationOptions {
sessionId: string;
filePath: string;
sessionWorkingDir: string;
/** Remote (SSH) case: the path lives on the remote host (see attachment-registry). */
remote?: SessionRemote;
}
export async function registerGeneratedArtifactAttachment(
@@ -26,19 +29,30 @@ export async function registerGeneratedArtifactAttachment(
// Decide trust on the symlink-resolved path. If it can't be resolved, fall
// back to the strict force-confined policy (registration will 404 a missing
// file anyway).
let forceWorkspaceConfinement = true;
try {
const resolvedPath = realpathSync(options.filePath);
forceWorkspaceConfinement = !isAllowedGeneratedArtifactPath(resolvedPath, options.sessionWorkingDir);
} catch {
// Keep force confinement.
}
//
// A remote case keeps that strict policy unconditionally: the well-known Codex
// artifact directories are anchored at THIS host's home, which says nothing about
// a remote home, so only a file inside the remote workspace is trusted here.
const resolvedPath = options.remote ? undefined : tryRealpath(options.filePath);
const forceWorkspaceConfinement = !resolvedPath
? true
: !isAllowedGeneratedArtifactPath(resolvedPath, options.sessionWorkingDir);
return registerExternalAttachment(options.sessionId, options.filePath, {
sessionWorkingDir: options.sessionWorkingDir,
forceWorkspaceConfinement,
remote: options.remote,
});
}
/** `realpathSync` without the throw — undefined when the path does not resolve. */
function tryRealpath(path: string): string | undefined {
try {
return realpathSync(path);
} catch {
return undefined;
}
}
/** Well-known Codex generated-artifact directories, anchored at the user's home. */
function codexGeneratedDirs(): string[] {
const home = homedir();
+218 -6
View File
@@ -31,7 +31,7 @@
import { randomBytes } from 'node:crypto';
import { existsSync } from 'node:fs';
import { readFile, writeFile, mkdir, lstat, readdir, realpath, rename, unlink, rmdir } from 'node:fs/promises';
import { readFile, writeFile, mkdir, lstat, readdir, realpath, rename, unlink, rmdir, chmod } from 'node:fs/promises';
import { homedir } from 'node:os';
import { join, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
@@ -39,6 +39,7 @@ import { fileURLToPath } from 'node:url';
import type { HookEventType } from './types.js';
import { HOOK_TIMEOUT_SECONDS } from './config/auth-config.js';
import { dataPath } from './config/instance.js';
import { readJsonConfig, SETTINGS_PATH } from './web/route-helpers.js';
/**
* Serializes read-modify-write access to a `settings.local.json` path. Every
@@ -855,17 +856,19 @@ const STATUSLINE_MARKER = '/api/status-telemetry';
* (present in every managed session via tmux setenv), so the config is static.
*/
export function generateStatusLineCommand(): string {
// `curl -sk`: CODEMAN_API_URL is loopback HTTPS with a self-signed cert in the
// `curl -sfk`: CODEMAN_API_URL is loopback HTTPS with a self-signed cert in the
// production setup; without -k curl returns 000 and the statusline shows
// nothing. -k is safe here (loopback only). Falls back to a brand string so the
// footer is never blank if Codeman is unreachable.
// nothing. -k is safe here (loopback only); -f keeps an HTTP error body off
// the statusline. On any failure it prints NOTHING: the old `|| echo codeman`
// is the bare word that a hand-run `claude` in a managed repo rendered, and
// that reads as a broken config (discussion #405).
return (
`INPUT=$(cat 2>/dev/null || echo '{}'); ` +
`printf '{"sessionId":"%s","data":%s}' "$CODEMAN_SESSION_ID" "$INPUT" | ` +
`curl -sk -X POST "$CODEMAN_API_URL${STATUSLINE_MARKER}" ` +
`curl -sfk -X POST "$CODEMAN_API_URL${STATUSLINE_MARKER}" ` +
`-H 'Content-Type: application/json' ` +
`-H "X-Codeman-Hook-Secret: $(cat "$CODEMAN_HOOK_SECRET_FILE" 2>/dev/null)" ` +
`--data @- 2>/dev/null || echo codeman`
`--data @- 2>/dev/null || true`
);
}
@@ -906,6 +909,215 @@ export async function applyStatusLineConfig(casePath: string, enabled: boolean):
});
}
/**
* Version-agnostic marker embedded as a comment in the generated exporter
* SCRIPT (see ensureStatusLineExporterScript) — bump the numeric suffix
* whenever the script content changes so `ensureStatusLineExporterScript`'s
* content comparison rewrites stale copies on next use.
*/
const STATUSLINE_EXPORTER_SCRIPT_MARKER = 'CODEMAN_STATUSLINE_EXPORTER_V4';
function statusLineExporterScriptContent(): string {
// Where the telemetry POST runs depends on who owns the footer. When the pane's
// env carries CODEMAN_USER_STATUSLINE_CMD (set via tmux setenv by TmuxManager
// when findEffectiveUserStatusLineCommand found the user's own REAL statusLine —
// see that function's doc comment), the user's command owns the footer, so the
// POST runs in a BACKGROUND subshell with stdin/stdout/stderr all closed
// (`>/dev/null 2>&1 </dev/null &`) — closing stdout/stderr keeps it from adding
// latency or leaking into the visible statusline, and closing stdin too is what
// lets a host reading this script's own stdout to EOF (`sh script | cat`) see
// that EOF promptly: without it the backgrounded curl keeps the pipe's write end
// open until IT exits, so the reader blocks for however long curl takes (measured
// ~5s with a stand-in) instead of the ~9ms it takes once stdin is closed too.
// Absent a user statusline, NOTHING else will print the footer, so the POST runs
// in the FOREGROUND and ITS OWN stdout becomes the footer — `/api/status-telemetry`
// returns formatSessionStatusText(...) (model/tokens/context %) precisely so this
// can happen. If curl itself fails (refused/unreachable Codeman, or an HTTP
// error, which `-f` keeps off stdout) the footer is simply EMPTY (`|| true`):
// the old `|| echo codeman` rendered a bare brand word that reads as a broken
// config, the symptom discussion #405 opened with. `--max-time` bounds a
// HUNG (not just refused) Codeman so it cannot wedge the render indefinitely.
const post =
`printf '{"sessionId":"%s","data":%s}' "$CODEMAN_SESSION_ID" "$INPUT" | ` +
`curl -sfk --max-time 5 -X POST "$CODEMAN_API_URL${STATUSLINE_MARKER}" ` +
`-H 'Content-Type: application/json' ` +
`-H "X-Codeman-Hook-Secret: $(cat "$CODEMAN_HOOK_SECRET_FILE" 2>/dev/null)" ` +
`--data @-`;
return (
`#!/bin/sh\n` +
`# ${STATUSLINE_EXPORTER_SCRIPT_MARKER} — auto-generated by Codeman; safe to delete, regenerated on demand.\n` +
`INPUT=$(cat 2>/dev/null || echo '{}')\n` +
`if [ -n "$CODEMAN_USER_STATUSLINE_CMD" ]; then\n` +
` ( ${post} ) >/dev/null 2>&1 </dev/null &\n` +
` printf '%s' "$INPUT" | sh -c "$CODEMAN_USER_STATUSLINE_CMD"\n` +
`else\n` +
` ${post} 2>/dev/null || true\n` +
`fi\n`
);
}
async function readStatusLineCommandFromFile(settingsPath: string): Promise<string | undefined> {
if (!existsSync(settingsPath)) return undefined;
try {
const parsed = JSON.parse(await readFile(settingsPath, 'utf-8'));
const current = parsed.statusLine as { command?: unknown } | undefined;
return current && typeof current.command === 'string' ? current.command : undefined;
} catch {
return undefined; // Malformed — treat as absent, same posture as applyStatusLineConfig.
}
}
/**
* Walk Claude Code's OWN settings precedence for `workingDir` to find whatever
* statusLine command is ACTUALLY effective there right now: project-local
* `.claude/settings.local.json` > project-shared `.claude/settings.json` >
* the user's global `~/.claude/settings.json`. Returns undefined when none of
* the three configures one.
*
* A legacy Codeman-marked entry in the project's OWN settings.local.json
* (written by an older build's disk-based mechanism) is never treated as a
* real user command — resolveStatusLineCliCommand strips it before this ever
* runs, so ordinarily this function never even sees one; the marker check
* here is a second, defensive guard in case something else wrote a copy in
* between, and precedence simply continues to the next layer instead of
* stopping on it.
*/
export async function findEffectiveUserStatusLineCommand(workingDir: string): Promise<string | undefined> {
const projectLocal = await readStatusLineCommandFromFile(join(workingDir, '.claude', 'settings.local.json'));
if (projectLocal && !projectLocal.includes(STATUSLINE_MARKER)) return projectLocal;
const projectShared = await readStatusLineCommandFromFile(join(workingDir, '.claude', 'settings.json'));
if (projectShared) return projectShared;
return readStatusLineCommandFromFile(join(homedir(), '.claude', 'settings.json'));
}
/**
* Write (or refresh) the SHARED, single exporter script every claude session
* points its ephemeral --settings statusLine flag at, and return its absolute
* path. Idempotent: only rewrites when the marker-versioned content differs.
*
* This is the fix for a real bug found live 2026-08-31: the exporter's
* command string legitimately depends on `$CODEMAN_SESSION_ID`,
* `$CODEMAN_API_URL`, `$CODEMAN_HOOK_SECRET_FILE`, and its own internal
* `$INPUT` — all meant to be expanded ONLY when Claude Code itself finally
* executes the statusLine command, using the PANE's tmux-setenv'd
* environment. Passing that command as literal TEXT through
* `--settings '...'` routes it through this server's OWN spawn-time shell
* layers first (tmux respawn-pane's `bash -c "..."`, itself invoked via
* execSync's implicit `/bin/sh -c`) — and POSIX double quotes do NOT
* suppress `$` expansion, so those vars got expanded there and then, against
* the SERVER process's environment (where they are unset), producing a
* mangled curl call that posted malformed JSON and printed the server's raw
* error response as the statusline text itself. A bare file PATH has no `$`,
* quotes, or pipes for any of those intermediate shells to mangle — the
* script's own content (containing the real `$VAR`s) is never touched by a
* shell until Claude Code executes the file itself, at which point the
* pane's real environment is in scope. This mirrors the existing #208 fix in
* tmux-manager.ts (never embed a literal `$SHELL` meant for later
* expansion — resolve it, or in this case reference a file, instead).
*/
export async function ensureStatusLineExporterScript(): Promise<string> {
const scriptPath = dataPath('statusline-exporter.sh');
const desired = statusLineExporterScriptContent();
let current: string | null = null;
try {
current = await readFile(scriptPath, 'utf-8');
} catch {
// Doesn't exist yet.
}
if (current !== desired) {
// Temp file + rename: live sessions execute this script on every statusline
// render, and a truncate-then-write (plus a chmod AFTER the write) opened two
// windows in which Claude Code could run an empty or non-executable file.
// rename() swaps the complete, already-executable file in atomically.
const tmpPath = `${scriptPath}.${process.pid}.${Date.now()}.tmp`;
await writeFile(tmpPath, desired);
await chmod(tmpPath, 0o755);
await rename(tmpPath, scriptPath);
}
return scriptPath;
}
/**
* Whether plan-usage telemetry collection is CURRENTLY wanted — read FRESH
* from the persisted `showPlanUsageLimits` setting on every call, never
* cached and never per-session. Reusing that setting rather than inventing a
* second persisted flag: it's the SAME boolean the App Settings chip checkbox
* already writes (see `planUsageChipEnabled()` in settings-ui.js).
*
* This is what lets the on/off decision survive a Codeman restart (there is
* no per-session state to lose — see the now-removed `Session._statusLineTelemetry`,
* which WAS such a per-session field and went stale on every restart) and
* apply uniformly across every claude session-creation path — interactive
* create, cron, the Ralph Loop API, quick-start — with none of them needing
* to thread a request-time flag through: they all already construct a
* session via TmuxManager.createSession/respawnPane, which reads this at
* spawn time.
*
* An ABSENT key means ON, mirroring readWorkspaceHooksEnabled() above: the
* client shows the chip and its checkbox as already on for a desktop that has
* never touched the setting (planUsageChipEnabled() in settings-ui.js), and
* the exporter only ever posts to THIS Codeman over loopback, so the honest
* default for an install that never said otherwise is the one the user can
* see. Resolving the default here, in the reader, is what lets
* `GET /api/settings` stay a plain read: a reconcile write there ran on every
* page load and could replace an unreadable settings.json with a one-key
* file. Only an explicit `false` (a save that flipped the chip off on some
* device) turns collection off.
*/
export async function readPlanUsageTelemetryEnabled(): Promise<boolean> {
const settings = await readJsonConfig<Record<string, unknown>>(SETTINGS_PATH, 'settings.json', {});
return settings.showPlanUsageLimits !== false;
}
/**
* Resolve the statusLine command to pass as an EPHEMERAL `claude --settings`
* CLI flag for this one process (see buildSpawnCommandFromRegistry in
* session-cli-registry-bridge.ts) — never written to disk. This supersedes
* the old applyStatusLineConfig(path, true) disk-write: a file-based
* statusLine leaked into any plain `claude` run in that directory outside
* Codeman entirely (it took precedence over the user's own global/project
* statusline with no disclosure and no way to remove it — found live
* 2026-08-31).
*
* Also self-heals: if an OLDER Codeman build already wrote its marked
* exporter into this workspace's settings.local.json, it is stripped here
* (isOurs-guarded, same as applyStatusLineConfig's removal branch) so every
* workspace migrates off the disk-based mechanism the first time a session
* starts there again — no manual cleanup required. This self-heal runs
* regardless of `telemetryEnabled`, so a legacy leftover is cleaned up even
* while the setting is currently off.
*
* Returns undefined when telemetry isn't currently enabled (see
* readPlanUsageTelemetryEnabled), or when the workspace already has its OWN
* hand-configured statusLine (never override a real one).
*/
export async function resolveStatusLineCliCommand(
casePath: string,
telemetryEnabled: boolean
): Promise<string | undefined> {
const settingsPath = join(casePath, '.claude', 'settings.local.json');
let userHasOwnStatusLine = false;
if (existsSync(settingsPath)) {
try {
const existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
const current = existing.statusLine as { command?: unknown } | undefined;
if (current && typeof current.command === 'string') {
if (current.command.includes(STATUSLINE_MARKER)) {
await applyStatusLineConfig(casePath, false); // strip legacy disk-written exporter
} else {
userHasOwnStatusLine = true;
}
}
} catch {
// Malformed — leave it alone, same guard applyStatusLineConfig itself uses.
}
}
if (!telemetryEnabled || userHasOwnStatusLine) return undefined;
return ensureStatusLineExporterScript();
}
// ─── Agent skill injection ───────────────────────────────────────────────────
/**
+7
View File
@@ -123,6 +123,13 @@ export interface RespawnPaneOptions {
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (preserved across respawns). */
envOverrides?: Record<string, string>;
/**
* Env vars to REMOVE from the tmux session (`setenv -u`) before `envOverrides` is
* applied. `setenv` persists at the tmux-session level and is inherited by
* `respawn-pane`, so a key that merely disappears from `envOverrides` stays set
* for the relaunched CLI; clearing a custom-model selection has to name it.
*/
unsetEnvKeys?: string[];
/** Claude CLI effort level (preserved across respawns, injected via `--settings`) */
effort?: EffortLevel;
/** Original tmux history-limit retained for config parity; respawn cannot resize the existing pane. */
+398
View File
@@ -0,0 +1,398 @@
/**
* @fileoverview Remote (SSH) file access for remote-SSH cases.
*
* A remote case's `workingDir` is an absolute path on ANOTHER host
* (`Session.workingDir = RemoteCase.remotePath`, see docs/remote-sessions.md). Every
* file route used to read it with local `fs`, which cannot work: the local
* `realpathSync` in `validateSessionFilePath` fails first, so the request died as a
* 404 "File not found" before a byte was read (#415). This module is the ONE place
* that reads remote bytes, mirroring how `remote-hosts.ts` is the one place that
* builds an ssh command line.
*
* Connection options come from `buildSshConnectionArgs()` — never a hand-built ssh
* line (the COD-107 discipline in docs/remote-sessions.md) — so a proxied,
* custom-port or jump-hosted case reaches its files with exactly the credentials the
* launch used, and `BatchMode=yes` means a host that needs a passphrase fails fast
* instead of hanging on a prompt nothing can answer.
*
* ⚠️ The path is the injection surface: it arrives from the browser (`?path=`). It is
* always interpolated as a single `shellescape`d token, and the whole remote command
* is itself shellescaped into the ssh line, so the local shell and the remote shell
* each see one opaque argument. Never build a command here by concatenating a raw
* path into the string.
*
* Read-only by design: previews, text reads and streaming. Writing to a remote file
* is deliberately NOT implemented (docs/file-viewer-edit-plan.md §6), nor are the
* office-conversion/thumbnail paths that would need the bytes on the server's disk.
*/
import { exec, spawn } from 'node:child_process';
import { promisify } from 'node:util';
import { PassThrough, type Readable } from 'node:stream';
import type { SessionRemote } from './types/session.js';
import { buildSshConnectionArgs, remoteSshTarget, shellescape } from './remote-hosts.js';
import { runWithRemoteSshLimit } from './remote-ssh-limiter.js';
const execAsync = promisify(exec);
/**
* Bound on the probe (realpath + stat) round trip. The connect itself is already
* bounded by `buildSshConnectionArgs`'s default `-o ConnectTimeout=10`; this covers
* a host that accepts the TCP connection and then never answers.
*/
const REMOTE_PROBE_TIMEOUT_MS = 20_000;
/** Bound on a buffered remote read (`cat`), on top of the caller's own size cap. */
const REMOTE_READ_TIMEOUT_MS = 30_000;
/** Slack over the caller's byte cap so a file exactly at the limit still fits. */
const READ_BUFFER_SLACK_BYTES = 64 * 1024;
/** Marker a probe prints when the path does not exist on the remote host. */
const NOT_FOUND_MARKER = 'n';
/**
* Marker a probe prints when the path exists but could NOT be canonicalized (no
* `readlink -f`, and the portable fallback hit its hop cap or a `readlink` failure).
* Parsed as `null`, i.e. 404: a path whose real target is unknown must never be
* served, because every containment and blocklist check runs on the resolved path.
*/
const UNRESOLVABLE_MARKER = 'x';
/**
* Paths per ssh round trip. The whole remote script is ONE shellescaped argument,
* and Linux caps a single argv string at 128 KiB, so a 100-entry attachment history
* of long paths is split rather than risking `E2BIG` on the local `sh`.
*/
const REMOTE_PROBE_CHUNK_SIZE = 40;
/** Symlink hops the portable resolver follows before giving up (Linux uses 40). */
const REMOTE_SYMLINK_MAX_HOPS = 40;
/**
* Under vitest no real ssh connection may ever be opened (mirrors
* `checkRemoteTmuxAvailable` and friends in remote-hosts.ts). The route tests mock
* this module, so nothing reaches here today; this is what keeps the NEXT
* remote-session test that touches a file route from opening a connection from CI.
* A clear 502-shaped error, never a fake success: there are no fake bytes to return.
*/
function assertNotUnderTest(): void {
if (process.env.VITEST) {
throw new RemoteFileAccessError('remote file access is disabled under test');
}
}
/** What a remote path turned out to be. `other` = symlink/socket/fifo/device. */
export type RemotePathKind = 'file' | 'directory' | 'other';
export interface RemoteProbe {
/** The path with symlinks resolved on the REMOTE host. */
realPath: string;
kind: RemotePathKind;
/** Size in bytes (0 for anything that is not a regular file). */
size: number;
/** mtime in ms since epoch (0 when the remote `stat` reported none). */
mtimeMs: number;
}
/**
* A remote file access failed for a reason that is NOT "the file is missing" —
* unreachable host, timeout, ssh error, unexpected probe output. Callers map this to
* a 5xx with the remote reason in the message; a missing file is reported separately
* as `null`/404 so the two cannot be confused.
*/
export class RemoteFileAccessError extends Error {
constructor(message: string) {
super(message);
this.name = 'RemoteFileAccessError';
}
}
/**
* Wrap a remote shell command in the shared, shellescaped ssh line.
*
* The single entry point for "run this on the remote host": connection args (port,
* identity, jump host, SOCKS ProxyCommand, extra `-o`) all come from
* `buildSshConnectionArgs`, and the command is ONE shellescaped token, so a path with
* spaces, quotes or `$(…)` cannot escape into the ssh command line.
*/
export function buildRemoteFileCommand(remote: SessionRemote, shellCommand: string): string {
return [...buildSshConnectionArgs(remote), remoteSshTarget(remote), shellescape(shellCommand)].join(' ');
}
/**
* `realpath + stat + existence` for one or more paths, in a SINGLE ssh round trip.
*
* One call instead of three matters: without a shared connection (no ControlMaster)
* every extra `ssh` is a fresh handshake, and the file routes need the path AND the
* workspace root canonicalized to compare them.
*
* Output format: the script first prints a lone NUL, then one NUL-terminated record
* per path, `<index>|n` (missing), `<index>|x` (exists but cannot be canonicalized) or
* `<index>|kind|size|mtime|realPath`. Records are keyed by INDEX and separated by NUL
* rather than newline so that a remote filename containing a newline cannot shift the
* alignment, and the leading NUL is what separates a login banner or an eager rc-file
* `echo` (which land before the script runs) from the records without any "last N
* lines" guesswork. `realPath` is the last field, so a `|` in a path still parses.
*
* Symlink resolution is portable AND fails closed. `readlink -f` where available
* (Linux, macOS >= 12.3); otherwise the fallback canonicalizes the directory chain
* with `cd -P`/`pwd -P` and then follows the LAST component with plain `readlink`
* (which the systems lacking `-f` do have) for a bounded number of hops. A path the
* fallback cannot resolve prints `x`, never the unresolved string: every containment
* and blocklist check downstream runs on `realPath`, and an earlier version of this
* fallback returned the directory-resolved path with the final symlink still in it,
* so `ws/notes.txt -> ~/.ssh/id_rsa` passed containment while `cat` served the key.
*/
export function buildRemoteProbeCommand(paths: readonly string[]): string {
const probes = paths.map((path, index) => `probe ${index} ${shellescape(path)}`).join('\n');
return [
'resolve_last() {',
' q=$1',
' hops=0',
' while :; do',
' d=$(cd -P "$(dirname "$q")" 2>/dev/null && pwd -P) || return 1',
' q=$d/$(basename "$q")',
' [ -L "$q" ] || break',
' hops=$((hops + 1))',
` [ "$hops" -le ${REMOTE_SYMLINK_MAX_HOPS} ] || return 1`,
' l=$(readlink "$q" 2>/dev/null) || return 1',
' [ -n "$l" ] || return 1',
' case $l in /*) q=$l ;; *) q=$d/$l ;; esac',
' done',
' if [ -d "$q" ]; then q=$(cd -P "$q" 2>/dev/null && pwd -P) || return 1; fi',
' printf %s "$q"',
'}',
'probe() {',
' i=$1',
' p=$2',
` if [ ! -e "$p" ]; then printf '%s|${NOT_FOUND_MARKER}\\0' "$i"; return; fi`,
` r=$(readlink -f "$p" 2>/dev/null) || r=$(resolve_last "$p") || { printf '%s|${UNRESOLVABLE_MARKER}\\0' "$i"; return; }`,
` [ -n "$r" ] || { printf '%s|${UNRESOLVABLE_MARKER}\\0' "$i"; return; }`,
' if [ -d "$r" ]; then t=d; elif [ -f "$r" ]; then t=f; else t=o; fi',
' s=0',
' if [ "$t" = f ]; then s=$(stat -c %s "$r" 2>/dev/null || stat -f %z "$r" 2>/dev/null); [ -n "$s" ] || s=0; fi',
' m=$(stat -c %Y "$r" 2>/dev/null || stat -f %m "$r" 2>/dev/null || printf 0)',
` printf '%s|%s|%s|%s|%s\\0' "$i" "$t" "$s" "$m" "$r"`,
'}',
"printf '\\0'",
probes,
].join('\n');
}
/**
* Parse one probe record (index prefix already stripped). `null` for the not-found
* and unresolvable markers or anything malformed.
*/
export function parseRemoteProbeRecord(record: string): RemoteProbe | null {
if (!record || record === NOT_FOUND_MARKER || record === UNRESOLVABLE_MARKER) return null;
const parts = record.split('|');
if (parts.length < 4) return null;
const [kindRaw, sizeRaw, mtimeRaw] = parts;
const kind: RemotePathKind | null =
kindRaw === 'f' ? 'file' : kindRaw === 'd' ? 'directory' : kindRaw === 'o' ? 'other' : null;
if (!kind) return null;
const realPath = parts.slice(3).join('|');
if (!realPath) return null;
const size = Number.parseInt(sizeRaw, 10);
const mtimeSeconds = Number.parseInt(mtimeRaw, 10);
return {
realPath,
kind,
size: Number.isFinite(size) && size > 0 ? size : 0,
mtimeMs: Number.isFinite(mtimeSeconds) && mtimeSeconds > 0 ? mtimeSeconds * 1000 : 0,
};
}
/**
* Parse the output of {@link buildRemoteProbeCommand} into one entry per requested
* path, in order. Throws when a path's record is missing: that means the transport
* or the remote shell did something unexpected, and silently treating it as "not
* found" would turn an infrastructure failure into a wrong 404.
*
* Everything before the first NUL is the remote shell's own chatter (banner, rc-file
* output) and is discarded; records are matched by their index prefix, so neither
* extra output nor a newline inside a filename can shift the mapping.
*/
export function parseRemoteProbeOutput(stdout: string, paths: readonly string[]): Array<RemoteProbe | null> {
const records = stdout.split('\0').slice(1);
const byIndex = new Map<number, string>();
for (const record of records) {
const match = /^(\d+)\|([\s\S]*)$/.exec(record);
if (!match) continue;
const index = Number.parseInt(match[1], 10);
if (!byIndex.has(index)) byIndex.set(index, match[2]);
}
return paths.map((_, index) => {
const record = byIndex.get(index);
if (record === undefined) {
throw new RemoteFileAccessError('remote host returned no usable file information');
}
return parseRemoteProbeRecord(record);
});
}
/**
* Probe one or more remote paths. Entry is `null` for a path that does not exist (or
* could not be canonicalized, which is refused the same way).
*
* Large batches are split into round trips of {@link REMOTE_PROBE_CHUNK_SIZE}, each
* counted against the global ssh limiter, so an attachment history of 100 entries
* costs three connections in sequence rather than 100 at once.
*/
export async function remoteProbePaths(
remote: SessionRemote,
paths: readonly string[]
): Promise<Array<RemoteProbe | null>> {
assertNotUnderTest();
const results: Array<RemoteProbe | null> = [];
for (let offset = 0; offset < paths.length; offset += REMOTE_PROBE_CHUNK_SIZE) {
const chunk = paths.slice(offset, offset + REMOTE_PROBE_CHUNK_SIZE);
const command = buildRemoteFileCommand(remote, buildRemoteProbeCommand(chunk));
let stdout: string;
try {
const result = await runWithRemoteSshLimit(() =>
execAsync(command, { timeout: REMOTE_PROBE_TIMEOUT_MS, maxBuffer: 256 * 1024 })
);
stdout = result.stdout;
} catch (err) {
throw new RemoteFileAccessError(
`remote host ${remote.label || remote.host} unreachable: ${describeExecError(err)}`
);
}
results.push(...parseRemoteProbeOutput(stdout, chunk));
}
return results;
}
/** Read a whole remote file into memory, capped by `maxBytes`. */
export async function remoteReadFile(remote: SessionRemote, remotePath: string, maxBytes: number): Promise<Buffer> {
assertNotUnderTest();
const command = buildRemoteFileCommand(remote, `cat ${shellescape(remotePath)}`);
try {
const result = await runWithRemoteSshLimit(() =>
execAsync(command, {
timeout: REMOTE_READ_TIMEOUT_MS,
maxBuffer: maxBytes + READ_BUFFER_SLACK_BYTES,
encoding: 'buffer',
})
);
return Buffer.isBuffer(result.stdout) ? result.stdout : Buffer.from(result.stdout);
} catch (err) {
throw new RemoteFileAccessError(`failed to read remote file: ${describeExecError(err)}`);
}
}
/**
* Command that writes a remote file's bytes to stdout.
*
* ⚠️ Range reads use `tail -c +N | head -c L` (both POSIX, constant memory) because
* the alternative — `dd bs=1` — issues one read syscall per byte and would make video
* seeking unusable. The trade-off is that a `tail` failure (the file vanished
* mid-request) reports `head`'s exit status, i.e. a short body on an already-sent
* 206; the client retries. The uncompressed path (`cat`) reports its own failure
* correctly, so the streaming error path is still covered by the normal case.
*/
export function buildRemoteReadCommand(remotePath: string, range?: { start: number; end: number }): string {
const quoted = shellescape(remotePath);
if (!range) return `cat ${quoted}`;
const length = range.end - range.start + 1;
return `tail -c +${range.start + 1} ${quoted} | head -c ${length}`;
}
export interface RemoteFileStream {
/** The remote file's bytes, streamed from the ssh child's stdout. */
stream: Readable;
/**
* Abort the transfer and reap the ssh child. The caller MUST call this when the
* HTTP request ends — especially on a client disconnect — or the `ssh` process
* keeps running (and holding a connection open) after nobody is reading it.
*/
close(): void;
}
/**
* Stream a remote file (optionally a byte range) as a Node Readable.
*
* Nothing is buffered in server memory: the bytes go from `ssh`'s stdout straight to
* the HTTP response, which is what makes a multi-GB remote video cost one pipe.
*/
export function remoteCreateReadStream(
remote: SessionRemote,
remotePath: string,
range?: { start: number; end: number }
): RemoteFileStream {
if (process.env.VITEST) {
// Same rule as the buffered calls, in stream form: the consumer sees the error
// through the stream's normal failure path instead of a connection attempt.
const stream = new PassThrough();
process.nextTick(() => stream.destroy(new RemoteFileAccessError('remote file access is disabled under test')));
return { stream, close: () => stream.destroy() };
}
const command = buildRemoteFileCommand(remote, buildRemoteReadCommand(remotePath, range));
const child = spawn(command, { shell: true, stdio: ['ignore', 'pipe', 'pipe'] });
let stderr = '';
child.stderr?.on('data', (chunk: Buffer) => {
if (stderr.length < 2000) stderr += chunk.toString();
});
const stream = child.stdout;
let ended = false;
stream.on('end', () => {
ended = true;
});
stream.on('error', () => {
ended = true;
});
child.on('error', (err: Error) => {
stream.destroy(err);
});
child.on('close', (code: number | null) => {
// Only a truncated transfer is an error. A non-zero exit AFTER the body finished
// (e.g. a signal delivered as the last byte was flushed) must not destroy an
// already-complete response, or the browser reports a broken body for a file it
// received in full.
if (ended || code === 0 || code === null) return;
const detail = stderr.trim().split('\n')[0];
stream.destroy(new RemoteFileAccessError(`remote read failed (ssh exit ${code})${detail ? `: ${detail}` : ''}`));
});
return {
stream,
close(): void {
if (!stream.destroyed) stream.destroy();
child.kill('SIGTERM');
},
};
}
/**
* First useful line of an exec/stderr error, for a user-facing message.
*
* ⚠️ Never Node's `err.message`: for a failed `exec` it is `Command failed: <the whole
* ssh line>`, which carries the identity-file path and the probe script, and this
* string goes out in a 502 body. stderr, the timeout flag and the exit/spawn code are
* everything a user can act on.
*/
function describeExecError(err: unknown): string {
if (typeof err === 'object' && err !== null) {
const record = err as { stderr?: unknown; code?: unknown; killed?: unknown };
const stderr =
typeof record.stderr === 'string' ? record.stderr : Buffer.isBuffer(record.stderr) ? String(record.stderr) : '';
const line = stderr
.split('\n')
.map((entry) => entry.trim())
.find((entry) => entry.length > 0);
if (line) return line.slice(0, 300);
if (record.killed) return 'timed out';
if (typeof record.code === 'number') return `ssh exit ${record.code}`;
if (typeof record.code === 'string') return `ssh could not be started (${record.code})`;
}
return 'unknown error';
}
+7 -1
View File
@@ -147,8 +147,14 @@ export function remoteSshTarget(host: Pick<RemoteHost, 'username' | 'host'>): st
* POSIX single-quote shell-escaping (end-quote, escaped-quote, restart-quote).
* Mirrors the helper in tmux-manager.ts so a value with spaces/metachars stays a
* single shell token. Used here for identity paths and `-o KEY=VALUE` options.
*
* EXPORTED for `remote-files.ts` (#415, remote file access): that module wraps a
* remote shell command in the ssh line built by `buildSshConnectionArgs()`, so it
* needs the same escaping discipline for the remote command itself and for every
* path interpolated into it. A third private copy of this function is exactly how
* two escaping implementations drift apart.
*/
function shellescape(str: string): string {
export function shellescape(str: string): string {
return "'" + str.replace(/'/g, "'\\''") + "'";
}
+89
View File
@@ -0,0 +1,89 @@
/**
* @fileoverview Global concurrency limiter for the short-lived `ssh` children that
* remote-case file access spawns (`src/remote-files.ts`: the realpath+stat probe and
* the buffered text read).
*
* Two paths can fan those out without a human behind each one:
*
* - `GET /api/sessions/:id/attachments` resolves every history entry (up to
* `ATTACHMENT_HISTORY_LIMIT`, 100), and the attachments drawer re-runs it on every
* `attachment:detected` event while it is open, which is exactly when an agent is
* writing files. The route now batches the probes, but a burst of drawers is still
* a burst.
* - A `codeman://attach?path=` magic link in terminal output registers the path
* fire-and-forget, once per distinct link per PTY chunk. In a remote session that
* output is written by a process on the remote host, so a prompt-injected agent can
* print hundreds of links and have the server fork one `ssh` per link, each holding
* a 20s probe timeout.
*
* Without a cap that is the fork-bomb shape `document-conversion-limiter.ts` exists to
* prevent, and it also trips OpenSSH's default `MaxStartups 10:30:100`, which starts
* dropping connections at ten unauthenticated handshakes. This is that limiter for
* ssh: a small fixed pool, FIFO queueing, and a slot handed straight to the next
* waiter on release so the active count can never exceed the cap under interleaved
* async resumption.
*
* Streams (`remoteCreateReadStream`) are deliberately NOT counted: one is opened per
* browser request and held for the life of a media playback, so four open videos
* would otherwise block every preview and the history list. They are already gated
* behind a counted probe (the guard re-probe runs first), so their spawn RATE is
* bounded here even though their concurrency is bounded by the browser.
*
* NOT re-entrant: never acquire from inside a task already holding a slot.
*/
/**
* Max remote probes/reads allowed to run concurrently across the whole process.
* Override with CODEMAN_MAX_REMOTE_FILE_SSH (clamped to >= 1). Four keeps a burst
* well under OpenSSH's ten-handshake default.
*/
const MAX_CONCURRENT_REMOTE_SSH = (() => {
const raw = Number(process.env.CODEMAN_MAX_REMOTE_FILE_SSH);
return Number.isFinite(raw) && raw >= 1 ? Math.floor(raw) : 4;
})();
let active = 0;
const waiters: Array<() => void> = [];
/** Test/diagnostic hook: remote calls currently holding a slot. */
export function getActiveRemoteSshCount(): number {
return active;
}
/** Test/diagnostic hook: remote calls queued behind the cap. */
export function getQueuedRemoteSshCount(): number {
return waiters.length;
}
/** The configured cap, so a test can assert against the real number. */
export function getRemoteSshLimit(): number {
return MAX_CONCURRENT_REMOTE_SSH;
}
function acquire(): Promise<void> {
if (active < MAX_CONCURRENT_REMOTE_SSH) {
active++;
return Promise.resolve();
}
return new Promise<void>((resolve) => waiters.push(resolve));
}
function release(): void {
const next = waiters.shift();
if (next) {
// Hand the slot straight to the next waiter; `active` stays at the cap.
next();
} else {
active--;
}
}
/** Run `task` once an ssh slot is free, releasing the slot afterward. */
export async function runWithRemoteSshLimit<T>(task: () => Promise<T>): Promise<T> {
await acquire();
try {
return await task();
} finally {
release();
}
}
+25 -2
View File
@@ -56,6 +56,14 @@ export interface SpawnBridgeOptions {
effort?: EffortLevel;
sessionName?: string;
claudeCliVersion?: string | null;
/**
* Resolved by resolveStatusLineCliCommand (hooks-config.ts) — undefined skips the
* exporter. Claude only. Rides the SAME `--settings` JSON object as `effortSettingsJson`
* (see buildSpawnCommandFromRegistry): Claude Code accepts only one `--settings` flag
* per invocation, so the two must be merged before reaching the argv engine rather than
* rendered as two independent params.
*/
statusLineCommand?: string;
}
/**
@@ -186,8 +194,23 @@ export function buildSpawnCommandFromRegistry(entry: CliEntry, options: SpawnBri
// than re-deriving the ultracode special case) keeps the EFFORT_LEVELS allowlist and the
// settings-JSON shape single-sourced in session-cli-builder.ts.
const [effortFlag, effortValue] = buildEffortCliArgs(options.effort);
if (effortFlag === '--settings') engineValues.effortSettingsJson = effortValue;
else if (effortFlag === '--effort') engineValues.effortLevel = effortValue;
if (effortFlag === '--effort') {
engineValues.effortLevel = effortValue;
}
// Fold the ephemeral plan-usage statusLine exporter (see resolveStatusLineCliCommand in
// hooks-config.ts) into the SAME `--settings` JSON object as ultracode/ effort, since Claude
// Code accepts only one `--settings` flag per invocation — rendering them as two independent
// params would let the second one silently win. Claude-only in practice (statusLineCommand
// is resolved claude-mode-only upstream), but this merge is mode-agnostic.
if ((effortFlag === '--settings' && effortValue) || options.statusLineCommand) {
const settingsObj: Record<string, unknown> =
effortFlag === '--settings' && effortValue ? JSON.parse(effortValue) : {};
if (options.statusLineCommand) {
settingsObj.statusLine = { type: 'command', command: options.statusLineCommand };
}
engineValues.effortSettingsJson = JSON.stringify(settingsObj);
}
// Preserves buildSpawnCommand's original fallback exactly: an EXPLICIT `undefined` probes
// the local claude CLI (getClaudeCliVersion, null under vitest); an explicit `null` means
+8 -1
View File
@@ -254,7 +254,14 @@ export class SessionManager extends EventEmitter {
// future reader of state.json.
const state = session.toState();
const envOverrides = session.getEnvOverridesForPersist();
const toStore = envOverrides ? { ...state, __envOverrides: envOverrides } : state;
// __customModel: same convention, the disk-only bookkeeping of a custom-model
// selection (env KEYS, config dir, launch model; never the injected values).
const customModel = session.getCustomModelForPersist();
const toStore = {
...state,
...(envOverrides ? { __envOverrides: envOverrides } : {}),
...(customModel ? { __customModel: customModel } : {}),
};
this.store.setSession(session.id, toStore as SessionState);
}
+179 -3
View File
@@ -48,6 +48,8 @@ import {
type OpenCodeConfig,
type CodexConfig,
type EffortLevel,
type CustomModelBookkeeping,
type CustomModelSelection,
type GeminiConfig,
type AntigravityConfig,
type PiConfig,
@@ -577,6 +579,20 @@ export class Session extends EventEmitter {
// the CLAUDE_CODE_EFFORT_LEVEL env var, which would hard-lock the session.
private _effort: EffortLevel | undefined;
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md). `envKeys`,
// `configDir` and `launchModel` are internal bookkeeping ONLY (never surfaced via
// toState()/the customModel getter): they are what setCustomModel() needs to undo a
// previous injection (remove exactly the env keys it added, delete a previous isolated
// config dir) without guessing what it once wrote. Persisted disk-only (`__customModel`).
private _customModel: CustomModelBookkeeping | undefined;
// Env keys a retired custom-model selection injected that the NEXT respawn must
// `tmux setenv -u`. Deleting a key from `_envOverrides` alone does nothing to the
// tmux session, which keeps every `setenv` and hands it to `respawn-pane`, so the
// relaunched CLI would come back still pointed at the old endpoint (measured, see
// TmuxManager.applyEnvOverrides). Drained after a successful respawn.
private _pendingEnvUnsets = new Set<string>();
// tmux history-limit (scrollback lines) allocated when this session's pane is created.
private readonly _tmuxHistoryLimit: number;
@@ -1238,6 +1254,65 @@ export class Session extends EventEmitter {
}
}
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — public-safe
// subset only (never envKeys/configDir/launchModel, the bookkeeping for setCustomModel).
get customModel(): CustomModelSelection | undefined {
if (!this._customModel) return undefined;
const { endpointId, modelId, label } = this._customModel;
return { endpointId, modelId, label };
}
/**
* The full selection incl. bookkeeping, for state.json ONLY (`__customModel`, the
* same disk-only convention as `getEnvOverridesForPersist()`). Carries no env values,
* so nothing secret lands on disk; recovery re-derives them from the endpoint store.
* Without this a Codeman restart left the pane on the custom endpoint (tmux keeps
* its `setenv`s) while `customModel` came back undefined, so the state was wrong and
* clearing had nothing to unset. Must NOT be included in any API-bound serializer.
*/
getCustomModelForPersist(): CustomModelBookkeeping | undefined {
return this._customModel ? { ...this._customModel, envKeys: [...this._customModel.envKeys] } : undefined;
}
/**
* Update this session's custom-model selection and merge the endpoint's injected env
* vars into `_envOverrides` — first UNDOING whatever the previous selection injected
* (removing exactly those env keys), so switching endpoints, or clearing back to the
* harness's native cloud default, never leaves a stale key behind. Synchronous and
* side-effect-free beyond mutating state, matching `setNice`/`setColor` above — this
* class does no file IO, so it reports the PREVIOUS `configDir` (if any) for the
* caller to clean up on disk (custom-model-injection.ts's configDir kind).
*
* Keys the previous selection injected that the new one does not re-set are queued
* for `tmux setenv -u` on the next respawn (`_pendingEnvUnsets`, threaded through
* `_buildRespawnPaneOptions().unsetEnvKeys`): the tmux session inherits every
* `setenv` into `respawn-pane`, so dropping them from the map alone would relaunch
* the CLI still pointed at the old endpoint — and for the `configDir` kinds, at a
* `HOME`/`CODEX_HOME`/`GROK_HOME` the caller has just deleted.
*/
setCustomModel(
next: CustomModelBookkeeping | undefined,
envOverrides?: Record<string, string>
): { removedEnvKeys: string[]; previousConfigDir: string | undefined } {
const previousConfigDir = this._customModel?.configDir;
const removedEnvKeys: string[] = [];
if (this._customModel) {
for (const key of this._customModel.envKeys) {
if (this._envOverrides) delete this._envOverrides[key];
removedEnvKeys.push(key);
this._pendingEnvUnsets.add(key);
}
}
this._customModel = next ? { ...next, envKeys: [...next.envKeys] } : undefined;
if (envOverrides && Object.keys(envOverrides).length > 0) {
this._envOverrides = { ...(this._envOverrides ?? {}), ...envOverrides };
// A key the new selection sets again does not need an unset (applyEnvOverrides
// would set it right back anyway); keep the list to what actually goes away.
for (const key of Object.keys(envOverrides)) this._pendingEnvUnsets.delete(key);
}
return { removedEnvKeys, previousConfigDir };
}
// Token tracking getters and setters
get totalTokens(): number {
return this._totalInputTokens + this._totalOutputTokens;
@@ -1478,6 +1553,7 @@ export class Session extends EventEmitter {
ompConfig: this._ompConfig,
resumeSessionId: this._resumeSessionId,
effort: this._effort,
customModel: this.customModel,
// COD-118: runtime-only — surfaced so the frontend can require explicit user
// intent before restarting a crash-looped session. Deliberately NOT restored
// by the constructor: a Codeman restart starts with a fresh breaker so boot
@@ -1617,6 +1693,7 @@ export class Session extends EventEmitter {
console.error('[Session] Failed to respawn pane, will create new session');
needsNewSession = true;
} else {
this._pendingEnvUnsets.clear();
// Wait a moment for the respawned process to fully start
await new Promise((resolve) => setTimeout(resolve, MUX_STARTUP_DELAY_MS));
}
@@ -1710,14 +1787,68 @@ export class Session extends EventEmitter {
return true;
}
/**
* Kill and relaunch this session's CLI process IN PLACE — same pane, same tmux
* session, fresh env/args from current state. Custom Model Endpoint Profiles
* (docs/custom-model-endpoints-plan.md) is the first caller: after `setCustomModel()` merges new
* env vars into `_envOverrides`, the running CLI process still has the OLD env
* (inherited at its own process start, not live-reloaded), so switching a
* session's model/endpoint requires this restart to actually take effect.
*
* A GENERALIZED {@link reattachRemote} with the `!this._remote` guard dropped —
* `_buildRespawnPaneOptions()` already passes `remote: this._remote` through
* unconditionally, so `mux.respawnPane()` builds the right command either way
* (a local session gets `respawn-pane -k` + the real launch line, which is the
* kill-and-relaunch this method exists for; a remote session gets the existing
* reattach-to-durable-tmux behavior). Deliberately does NOT check `isBusy()` —
* that's the caller's job (mirrors `/interactive`'s guard), since a raw restart
* primitive shouldn't itself decide when it's safe to use.
*
* @returns true if the pane was respawned, false otherwise (no mux session, or
* the mux session is gone — see {@link reattachRemote} for that reasoning).
*/
async restartCli(): Promise<boolean> {
if (!this._useMux || !this._mux || !this._muxSession) return false;
const mux = this._mux;
if (!mux.muxSessionExists(this._muxSession.muxName)) {
console.log('[Session] restartCli: mux session gone, skipping:', this._muxSession.muxName);
return false;
}
this._pinOmpRespawnId();
const options = this._buildRespawnPaneOptions();
// Unlike the dead-pane respawn, this one kills a WORKING pane whose conversation
// already has a transcript, and a CLI that launches with `--session-id <id>` refuses
// an id that is already in use (claude: `Error: Session ID ... is already in use.`),
// which turned an endpoint switch into a dead pane and a lost session. A launch that
// declares a `fallback` chain renders `resume || new` once a resume id is set, the
// same `--resume <id> || --session-id <id>` shape the docker and remote pane commands
// already use, so pin the live conversation id for THIS respawn only. The registry
// shape is the gate, not the CLI's name: an entry whose resume id is minted by the
// CLI itself (codex/pi/omp/grok) never declares that chain, and its resume field is
// read from its own `<Mode>Config` rather than this top-level one anyway.
if (!options.resumeSessionId && getCli(this.mode)?.launch.chain === 'fallback') {
options.resumeSessionId = this._claudeSessionId ?? this.id;
}
const newPid = await mux.respawnPane(options);
if (!newPid) {
console.error('[Session] restartCli: respawnPane failed for', this._muxSession.muxName);
return false;
}
this._pendingEnvUnsets.clear();
console.log('[Session] restartCli: restarted CLI for', this._muxSession.muxName, 'pid', newPid);
return true;
}
/**
* Assemble the {@link RespawnPaneOptions} for this session. Single source of
* truth shared by interactive start, shell start (via their inline copies),
* and {@link reattachRemote} so the remote reattach path can never drift from
* {@link reattachRemote}, and {@link restartCli} so no respawn path can drift from
* the spawn path.
*/
private _buildRespawnPaneOptions(): import('./mux-interface.js').RespawnPaneOptions {
return {
const options: import('./mux-interface.js').RespawnPaneOptions = {
sessionId: this.id,
workingDir: this.workingDir,
mode: this.mode,
@@ -1745,12 +1876,37 @@ export class Session extends EventEmitter {
ompConfig: this._ompConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
unsetEnvKeys: this._pendingEnvUnsets.size > 0 ? [...this._pendingEnvUnsets] : undefined,
effort: this._effort,
historyLimit: this._tmuxHistoryLimit,
remote: this._remote,
docker: this._docker,
owner: this._owner,
};
return this._withCustomModelLaunchModel(options);
}
/**
* Force the custom-model selection's `launchModel` (pi/omp `custom/<id>`, grok's
* `[model.<name>]` block name) onto the CLI's `model` launch param. Where that param
* lives is registry DATA — the entry's `legacyConfigField` (`piConfig`, `grokConfig`,
* ...) or the top-level `model` for an entry that declares none — so this stays a
* generic reader rather than a branch per CLI. Applied on the OPTIONS only: the stored
* `<Mode>Config` keeps whatever model the user chose at create, which is exactly what a
* later clear must fall back to.
*/
private _withCustomModelLaunchModel(
options: import('./mux-interface.js').RespawnPaneOptions
): import('./mux-interface.js').RespawnPaneOptions {
const launchModel = this._customModel?.launchModel;
if (!launchModel) return options;
const entry = getCli(this.mode);
if (!entry) return options;
const field = entry.launch.legacyConfigField;
if (!field) return { ...options, model: launchModel };
const bag = options as unknown as Record<string, unknown>;
const existing = (bag[field] ?? {}) as Record<string, unknown>;
return { ...options, [field]: { ...existing, model: launchModel } };
}
/**
@@ -2174,7 +2330,12 @@ export class Session extends EventEmitter {
}
try {
// Pass --session-id to use the SAME ID as the Codeman session
// This ensures subagents can be directly matched to the correct tab
// This ensures subagents can be directly matched to the correct tab.
// No plan-usage statusLine exporter on this path: the ephemeral
// `--settings` injection (resolveStatusLineCliCommand, hooks-config.ts)
// is wired into the tmux spawn builders only, so a direct-PTY session
// has no Claude telemetry in the header chip. Documented in
// architecture-invariants (Plan-usage chip); tmux is the supported path.
const args = buildInteractiveArgs(
this.id,
this._claudeMode,
@@ -3411,6 +3572,21 @@ export class Session extends EventEmitter {
* half-open socket silently drops frames with no error) would type a prompt
* twice whenever an ACK is lost after the write landed.
*/
/**
* The highest input seq recorded for `clientId`, or 0 when this session has
* never seen it.
*
* Reported back on a REJECTED (duplicate) frame so the client can lift its own
* counter above this watermark. Without that number a client whose persisted
* counter fell behind ours has no way to find its way out: every fresh
* keystroke it sends lands at or below the watermark, is dropped as a
* duplicate, and is ACKed anyway — so the UI looks healthy while nothing is
* delivered, and a reload restores the same stale counter from localStorage.
*/
lastInputSeq(clientId: string): number {
return this._appliedInputSeq.get(clientId) ?? 0;
}
shouldApplyInput(clientId: string, seq: number): boolean {
const last = this._appliedInputSeq.get(clientId);
if (last !== undefined && seq <= last) return false;
+89 -11
View File
@@ -67,6 +67,11 @@ import {
legacyConfigForMode,
} from './session-cli-registry-bridge.js';
import type { CliEntry } from './config/cli-registry/types.js';
import {
resolveStatusLineCliCommand,
readPlanUsageTelemetryEnabled,
findEffectiveUserStatusLineCommand,
} from './hooks-config.js';
import {
buildSshConnectionArgs,
defaultRemoteCommandForMode,
@@ -728,6 +733,8 @@ export function buildSpawnCommand(options: {
ompConfig?: OmpConfig;
resumeSessionId?: string;
effort?: EffortLevel;
/** Resolved by resolveStatusLineCliCommand (hooks-config.ts) — undefined skips the exporter. Claude only. */
statusLineCommand?: string;
/** Codeman session name, passed to claude as `--name` (version-gated, sanitized; local spawns only). */
sessionName?: string;
/**
@@ -1731,21 +1738,34 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
* Key validation is strict (`/^[A-Z_][A-Z0-9_]*$/`) as defense-in-depth against
* shell-metachar injection even if upstream schema check is bypassed.
*/
private applyEnvOverrides(muxName: string, envOverrides?: Record<string, string>): void {
private applyEnvOverrides(muxName: string, envOverrides?: Record<string, string>, unsetKeys?: string[]): void {
const VALID_KEY = /^[A-Z_][A-Z0-9_]*$/;
// Legacy cleanup: pre-0.7.2 set CLAUDE_CODE_EFFORT_LEVEL via setenv, which persists
// on the tmux session and hard-locks /effort switching in every respawned pane.
// Effort now flows as a `--settings` soft default (see buildEffortSettingsFlag),
// so unconditionally unset the stale var before applying current overrides.
try {
execSync(`${this.tmux()} setenv -t ${shellescape(muxName)} -u CLAUDE_CODE_EFFORT_LEVEL`, {
timeout: EXEC_TIMEOUT_MS,
stdio: ['pipe', 'pipe', 'pipe'],
});
} catch {
/* Non-critical — var may not exist */
//
// The caller's own unsets ride the same path, and run BEFORE the overrides are
// (re)applied: a key that is both unset and present in `envOverrides` ends up set,
// so a stale unset can never clobber a live value. Removing a key from the map is
// not enough on its own — `setenv` persists at the tmux-session level and is
// inherited by `respawn-pane`, measured: `setenv FOO bar` survived two successive
// `respawn-pane -k`. Clearing a custom-model selection is what needs this.
for (const key of ['CLAUDE_CODE_EFFORT_LEVEL', ...(unsetKeys ?? [])]) {
if (!VALID_KEY.test(key)) {
console.warn(`[TmuxManager] Skipping invalid env unset key: ${JSON.stringify(key)}`);
continue;
}
try {
execSync(`${this.tmux()} setenv -t ${shellescape(muxName)} -u ${key}`, {
timeout: EXEC_TIMEOUT_MS,
stdio: ['pipe', 'pipe', 'pipe'],
});
} catch {
/* Non-critical — var may not exist */
}
}
if (!envOverrides) return;
const VALID_KEY = /^[A-Z_][A-Z0-9_]*$/;
for (const [key, value] of Object.entries(envOverrides)) {
if (!value) continue; // Skip empty — nothing to set
if (!VALID_KEY.test(key)) {
@@ -1832,6 +1852,38 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
}
/**
* Export the user's own REAL statusLine command (found by
* findEffectiveUserStatusLineCommand) via tmux setenv, so the shared
* exporter script (statusLineExporterScriptContent in hooks-config.ts) can
* wrap it. Via setenv rather than embedding it in the spawn command line:
* tmux stores a setenv value verbatim and never re-parses it as shell
* syntax, so once safely escaped for THIS one command, the command's own
* `$`/quotes survive untouched into the claude process's environment — the
* same reasoning that made the exporter script itself necessary (see
* ensureStatusLineExporterScript's doc comment). Only this ONE line needs
* shellescape(); the stored value itself is opaque to tmux from then on.
*
* With NO user command the variable is UNSET rather than left alone: a tmux
* setenv survives respawn-pane, so a user who deleted their own statusline
* would otherwise keep getting the stale one wrapped (and lose Codeman's
* footer print-through) until the tmux session was recreated. Same shape as
* the CLAUDE_CODE_EFFORT_LEVEL cleanup in applyEnvOverrides.
*/
private _configureStatusLineUserCommand(muxName: string, command: string | undefined): void {
const setOrUnset = command
? `CODEMAN_USER_STATUSLINE_CMD ${shellescape(command)}`
: '-u CODEMAN_USER_STATUSLINE_CMD';
try {
execSync(`${this.tmux()} setenv -t ${shellescape(muxName)} ${setOrUnset}`, {
timeout: EXEC_TIMEOUT_MS,
stdio: 'ignore',
});
} catch {
// Non-critical: the exporter prints its own footer, or nothing.
}
}
/**
* Creates a new tmux session wrapping Claude CLI or a shell.
* In test mode: creates an in-memory session only (no real tmux session).
@@ -1915,6 +1967,19 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const envExportsStr = this.buildEnvExports(sessionId, muxName, mode).join(' && ');
// Registry-gated (capabilities.statusLineTelemetry — claude only today), local
// spawns only (remote/docker have their own separate command builders — out of
// scope here). Also self-heals: strips any legacy disk-written exporter from an
// older Codeman build the first time a session starts in that workspace again.
const statusLineCommand =
getCli(mode)?.capabilities.statusLineTelemetry && !remote && !docker
? await resolveStatusLineCliCommand(workingDir, await readPlanUsageTelemetryEnabled())
: undefined;
// The user's own REAL statusLine, if any (walked via Claude Code's own
// settings precedence) — exported below so the shared exporter script
// can wrap it. Only worth discovering when we're actually injecting.
const userStatusLineCommand = statusLineCommand ? await findEffectiveUserStatusLineCommand(workingDir) : undefined;
const baseCmd = buildSpawnCommand({
mode,
sessionId,
@@ -1931,6 +1996,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
ompConfig,
resumeSessionId,
effort,
statusLineCommand,
sessionName: name,
});
@@ -1993,6 +2059,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
mode,
legacyConfigForMode(mode, options as unknown as Record<string, unknown>)
);
this._configureStatusLineUserCommand(muxName, userStatusLineCommand);
// Apply user-supplied env overrides (e.g., CLAUDE_CODE_EFFORT_LEVEL) via tmux setenv
// so secret values stay off the bash command line. Must run before respawn-pane.
@@ -2154,6 +2221,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
ompConfig,
resumeSessionId,
envOverrides,
unsetEnvKeys,
effort,
remote,
docker,
@@ -2170,6 +2238,13 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const envExportsStr = this.buildEnvExports(sessionId, muxName, mode).join(' && ');
// See createSession()'s identical resolution for rationale.
const statusLineCommand =
getCli(mode)?.capabilities.statusLineTelemetry && !remote && !docker
? await resolveStatusLineCliCommand(workingDir, await readPlanUsageTelemetryEnabled())
: undefined;
const userStatusLineCommand = statusLineCommand ? await findEffectiveUserStatusLineCommand(workingDir) : undefined;
const baseCmd = buildSpawnCommand({
mode,
sessionId,
@@ -2186,6 +2261,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
ompConfig,
resumeSessionId,
effort,
statusLineCommand,
sessionName: name,
});
const config = niceConfig || DEFAULT_NICE_CONFIG;
@@ -2205,9 +2281,11 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
mode,
legacyConfigForMode(mode, options as unknown as Record<string, unknown>)
);
this._configureStatusLineUserCommand(muxName, userStatusLineCommand);
// Re-apply user env overrides before respawn so the new shell inherits them.
this.applyEnvOverrides(muxName, envOverrides);
// Re-apply user env overrides before respawn so the new shell inherits them,
// dropping the ones the caller retired first (see applyEnvOverrides).
this.applyEnvOverrides(muxName, envOverrides, unsetEnvKeys);
// -c /tmp + cd bounce — see createSession() for rationale (stale FUSE state).
const launchCmd = remote || docker ? fullCmd : `cd ${JSON.stringify(workingDir)} && ${fullCmd}`;
+6
View File
@@ -184,6 +184,8 @@ export interface CaseInfo {
container: string;
image?: string;
path: string;
/** Directory INSIDE the container (defaults to `path` when unset). */
containerWorkdir?: string;
network?: string;
/**
* CLIs available INSIDE the container. A container case runs its agents in
@@ -201,6 +203,10 @@ export interface CaseInfo {
* first session: the container is created on demand by the launch chain, so treating
* "not found" as a fault there hid every agent mode behind an error telling the user
* to start a container Codeman was about to create itself.
*
* It also gates the Add Case panel's "copy an existing case" picker: only an ADOPTED
* container may back several cases at once (`classifyAdoptContainerConflict`), since an
* owned container's lifecycle belongs to its one case.
*/
owned?: boolean;
};
+32
View File
@@ -558,6 +558,29 @@ export interface SessionAttachmentHistoryItem {
/**
* Current state of a session
*/
/** The public half of a session's custom-model selection (on the wire, in `SessionState`). */
export interface CustomModelSelection {
endpointId: string;
modelId: string;
label?: string;
}
/**
* The full custom-model selection a session keeps: the public selection plus the
* bookkeeping `Session.setCustomModel()` needs to UNDO it later without guessing what
* it once wrote. Persisted to state.json only as the disk-only `__customModel` field
* (never broadcast); the injected env VALUES are not in here at all, since they carry
* the endpoint's API key, and are re-derived from the endpoint store on recovery.
*/
export interface CustomModelBookkeeping extends CustomModelSelection {
/** Env keys the selection injected into the session's envOverrides / tmux session. */
envKeys: string[];
/** Isolated per-session config directory written for a `configDir`-kind CLI. */
configDir?: string;
/** Value forced onto the CLI's `model` launch param (pi/omp `custom/<id>`, grok's block name). */
launchModel?: string;
}
export interface SessionState {
/** Unique session identifier */
id: string;
@@ -677,6 +700,15 @@ export interface SessionState {
resumeSessionId?: string;
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
effort?: EffortLevel;
/**
* Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md): the custom
* OpenAI-compatible endpoint (local or cloud) this session's CLI is currently pointed
* at, if any. Undefined = the harness's native cloud default. No secrets here — the
* endpoint's base URL/api key live only in Session._envOverrides, never in this public
* state. The internal half (which env keys were injected, which config dir was
* written) is {@link CustomModelBookkeeping}, persisted disk-only like `__envOverrides`.
*/
customModel?: CustomModelSelection;
/** Sanitized per-session attachment history. */
attachmentHistory?: SessionAttachmentHistoryItem[];
/**
+6 -3
View File
@@ -173,10 +173,13 @@ export function parseSessionStatus(data: RawStatuslinePayload | undefined): Sess
* Format the in-terminal statusline footer: the CURRENT SESSION's status —
* `Opus 4.8 (1M context) in:562,411 out:1,188 ctx:56%` — NOT the plan limits,
* which live in the Codeman header chip. Claude requires a statusLine command to
* emit the rate_limits JSON at all, so this is what that command prints back.
* emit the rate_limits JSON at all, so this is what that command prints back
* when it has no statusline of the user's own to wrap. With nothing to show it
* returns '' rather than a brand word: a bare `codeman` on the statusline is
* the symptom discussion #405 opened with.
*/
export function formatSessionStatusText(s: SessionStatus | null): string {
if (!s) return 'codeman';
if (!s) return '';
const groups: string[] = [];
if (s.modelDisplayName) groups.push(s.modelDisplayName);
const tok: string[] = [];
@@ -184,7 +187,7 @@ export function formatSessionStatusText(s: SessionStatus | null): string {
if (s.outputTokens != null) tok.push(`out:${withCommas(s.outputTokens)}`);
if (tok.length) groups.push(tok.join(' '));
if (s.contextUsedPercentage != null) groups.push(`ctx:${Math.round(clampPct(s.contextUsedPercentage))}%`);
return groups.length ? groups.join(' ') : 'codeman';
return groups.length ? groups.join(' ') : '';
}
/**
+71 -1
View File
@@ -23,7 +23,14 @@ import { getHookSecret, HOOK_SECRET_HEADER } from '../../config/hook-secret.js';
import { isMultiUserMode } from '../../config/multiuser.js';
import { findUser, setPassword, touchLastLogin, verifyPassword } from '../../user-store.js';
import { webviewCapabilities } from '../../webview-capabilities.js';
import { capabilityFromProxyPath, capabilityFromReferer } from '../webview-proxy.js';
import {
capabilityFromProxyPath,
capabilityFromReferer,
carriesAuthCredentials,
isLostWebviewFrameNavigation,
lostWebviewFramePage,
LOST_FRAME_PAGE_CSP,
} from '../webview-proxy.js';
import { ApiErrorCode, createErrorResponse, type AuthUser } from '../../types.js';
// Request-scoped identity (multi-user). Single-user leaves it undefined and the
@@ -176,6 +183,65 @@ function hasValidWebviewCapability(req: FastifyRequest, basePath = ''): boolean
return !!fromReferer && webviewCapabilities.resolve(fromReferer) !== undefined;
}
/**
* A web-tab frame that navigated itself off the proxy prefix (see
* isLostWebviewFrameNavigation). It cannot authenticate: opaque origin, no cookie,
* no capability left in the URL. Answer with the static recovery page here, BEFORE
* the credential checks, so the reload of a proxied dashboard neither shows a
* login challenge inside the tab nor counts as a failed attempt against the
* caller's IP — a dev server that full-reloads on every save would otherwise
* rate-limit its own user out of Codeman. Fenced like the Referer exemption: a
* path that resolves to a real route (/api, /q, a registered handler) is never
* answered this way, so a genuine unauthenticated navigation still gets the 401.
*
* `/` is the one registered route that IS answered here, and only when the
* request carries neither the session cookie nor an Authorization header. The
* shim maps `/webview/<cap>/` to exactly `/`, so a dashboard that reloads on its
* landing page (a Vite dev server on a config change) asks for Codeman's root
* as an iframe navigation; answering that with the app shell put Codeman inside
* its own web tab, and with a password it was a 401 in the frame. Nothing in
* Codeman frames its own root and the sandboxed frame has no credentials, so the
* credential-free form can only be that frame; a framed `/` WITH credentials is
* still the shell. Property worth knowing: a non-browser client can set these
* headers too, so an unauthenticated caller can tell a registered route (401)
* from a non-route (200) and enumerate the route table. Accepted, because the
* routes are public in docs/api-reference.md.
*
* @returns true when the reply was sent.
*/
function serveLostWebviewFrame(req: FastifyRequest, reply: FastifyReply): boolean {
if (!isLostWebviewFrameNavigation(req)) return false;
const url = (req.url ?? '').split('?')[0];
if (url.startsWith('/api/') || url.startsWith('/ws/') || url.startsWith('/q/')) return false;
if (url === '/') {
if (carriesAuthCredentials(req.headers, AUTH_COOKIE_NAME)) return false;
} else if (matchesRegisteredRoute(req, url)) {
return false;
}
sendLostWebviewFramePage(reply);
return true;
}
/**
* The landing-page case of serveLostWebviewFrame, for the index route. Without
* CODEMAN_PASSWORD no auth hook runs at all, so a lost frame's reload of `/`
* reaches `GET /` directly and the route asks this before rendering the shell.
* Under a password the hook has already answered a credential-free lost frame,
* so here it only ever sees the credentialed form, which stays the shell.
*/
export function isLostWebviewRootFrame(req: FastifyRequest): boolean {
if (!isLostWebviewFrameNavigation(req)) return false;
if ((req.url ?? '').split('?')[0] !== '/') return false;
return !carriesAuthCredentials(req.headers, AUTH_COOKIE_NAME);
}
/** Send the static recovery page (lostWebviewFramePage) with its own CSP, uncached. */
export function sendLostWebviewFramePage(reply: FastifyReply): FastifyReply {
reply.header('content-security-policy', LOST_FRAME_PAGE_CSP);
reply.header('cache-control', 'no-store');
return reply.type('text/html; charset=utf-8').send(lostWebviewFramePage());
}
/**
* Whether `url` resolves to a route Codeman actually registered.
*
@@ -302,6 +368,8 @@ export function registerAuthMiddleware(app: FastifyInstance, https: boolean, bas
done();
return;
}
// A web-tab frame that lost its prefix: hand it back to its tab, no credentials involved.
if (serveLostWebviewFrame(req, reply)) return;
const clientIp = req.ip;
@@ -439,6 +507,8 @@ function registerMultiUserAuthHook(
// ownership against the identity BOUND TO THE CAPABILITY, which is stricter
// than re-deriving it from a request that carries no credentials.
if (hasValidWebviewCapability(req, basePath)) return;
// A web-tab frame that lost its prefix: hand it back to its tab, no credentials involved.
if (serveLostWebviewFrame(req, reply)) return;
const clientIp = req.ip;
+88 -18
View File
@@ -2895,7 +2895,7 @@ class CodemanApp {
} else if (msg.t === 'ia') {
// Input ACK — the server applied (or deduped) this seq; drop it from
// the durable queue so it can never be re-delivered/lost.
this._onWsInputAck(msg.seq);
this._onWsInputAck(msg.seq, msg);
}
} catch {
// Ignore malformed messages
@@ -3073,7 +3073,11 @@ class CodemanApp {
this._pendingDeliveries.set(sessionId, list);
}
list.push(rec);
this._persistReliableState();
// ⚠️ SYNCHRONOUS, not the debounced writer: the seq counter is precisely the
// thing that must survive a crash, and a debounce puts it on the path most
// likely to be lost. A counter that comes back BELOW the server's watermark
// makes every later keystroke a silently-dropped duplicate (see _onWsInputAck).
this._persistReliableNow();
this._updateConnectionIndicator();
this._drainSession(sessionId);
}
@@ -3179,9 +3183,40 @@ class CodemanApp {
this.markIdleAlertSeen?.(sessionId);
}
/** Server input-ACK frame ({t:'ia',seq}) over the WebSocket. */
_onWsInputAck(seq) {
if (this._wsSessionId && Number.isInteger(seq)) this._ackDelivery(this._wsSessionId, seq);
/**
* Server input-ACK frame ({t:'ia',seq}) over the WebSocket.
*
* `dup:true` means the server REJECTED the frame as already-seen rather than
* applying it, and `last` is its watermark for this clientId. That combination
* is the escape hatch from a rolled-back counter: our seqs persist on a
* debounced write, so a tab killed between a send and that write comes back
* counting from BELOW the server's watermark, and from then on every keystroke
* is dropped-but-ACKed — a silently dead terminal that a reload cannot fix,
* because the stale counter is restored from localStorage too.
*
* ⚠️ Only a FIRST-attempt record is re-queued. A retry (`tries > 1`) being
* called a duplicate is the mechanism working as designed — the original did
* land — and re-sending it would type the same thing twice.
*/
_onWsInputAck(seq, msg) {
const sessionId = this._wsSessionId;
if (!sessionId || !Number.isInteger(seq)) return;
if (msg && msg.dup) {
const list = this._pendingDeliveries.get(sessionId);
const rec = list && list.find((r) => r.seq === seq);
const watermark = Number.isInteger(msg.last) ? msg.last : seq;
// Lift the counter clear of the server's watermark before anything else, so
// the re-queue below (and every later keystroke) gets an acceptable seq.
if ((this._seqCounters.get(sessionId) || 0) <= watermark) {
this._seqCounters.set(sessionId, watermark);
this._persistReliableNow();
}
const lost = rec && rec.tries <= 1 ? rec.data : null;
this._ackDelivery(sessionId, seq);
if (lost !== null) this._reliableSend(sessionId, lost, rec.useMux);
return;
}
this._ackDelivery(sessionId, seq);
}
/** Called from ws.onopen — flush everything pending over the fresh socket. */
@@ -3601,13 +3636,20 @@ class CodemanApp {
* Reset all app state maps, timers, and handlers to a clean baseline.
* Called by handleInit() on SSE reconnect / page reload to prevent
* memory leaks and stale data.
*
* @param {boolean} [preserveTerminal] Keep the terminal caches. Set when an SSE
* RECONNECT lands back on the session already on screen: the buffers still
* describe that session, and dropping them forces a full refetch + xterm
* reset that throws away the user's scroll position (see handleInit).
*/
_resetAllAppState() {
_resetAllAppState(preserveTerminal = false) {
this.sessions.clear();
this.ralphStates.clear();
this.terminalBuffers.clear();
this.terminalBufferCache.clear();
this._xtermSnapshots?.clear();
if (!preserveTerminal) {
this.terminalBuffers.clear();
this.terminalBufferCache.clear();
this._xtermSnapshots?.clear();
}
this.projectInsights.clear();
this.teams.clear();
this.teamTasks.clear();
@@ -3734,7 +3776,23 @@ class CodemanApp {
// Stop any active voice recording on reconnect
VoiceInput.cleanup();
this._resetAllAppState();
// A RECONNECT that lands back on the same session must not become a full
// reload. This used to clear the terminal caches and re-run selectSession()
// unconditionally, so every SSE reconnect refetched the buffer (up to 1 MiB)
// and reset+rewrote xterm. On a link that drops a connection about once a
// minute that reads as the page refreshing itself and losing your place.
// Keep the caches and the active id here; the restore block below resyncs
// through _onSessionNeedsRefresh(), which still reloads the buffer (so
// output produced during the outage is not lost) but preserves the reading
// position.
const activeBefore = this.activeSessionId;
const keepTerminal =
gen > 1 &&
!!activeBefore &&
Array.isArray(data.sessions) &&
data.sessions.some((s) => s.id === activeBefore);
this._resetAllAppState(keepTerminal);
data.sessions.forEach(s => {
this.sessions.set(s.id, s);
@@ -3864,20 +3922,32 @@ class CodemanApp {
}
const previousActiveId = this.activeSessionId;
this.activeSessionId = null;
if (this.sessionOrder.length > 0) {
if (this.sessionOrder.length === 0) {
this.activeSessionId = null;
} else {
// Priority: current active > localStorage > first session
let restoreId = previousActiveId;
if (!restoreId || !this.sessions.has(restoreId)) {
try { restoreId = localStorage.getItem('codeman-active-session'); } catch {}
}
// `auto`: the app is restoring a session on load, not a human opening
// one, so a pending idle alert on that tab stays armed until it is
// actually tapped (see the userInitiated note in selectSession).
if (restoreId && this.sessions.has(restoreId)) {
this.selectSession(restoreId, { auto: true });
if (keepTerminal && restoreId === previousActiveId && this.sessions.has(restoreId)) {
// Reconnect onto the session already on screen. renderSessionTabs() ran
// above and activeSessionId never changed, so the tab strip is already
// correct; only the buffer needs to catch up. The WS has its own
// backoff reconnect, but if it is not on this session (dead socket, or
// a give-up) nothing else would re-establish it from here.
if (this._wsSessionId !== restoreId) this._connectWs(restoreId);
void this._onSessionNeedsRefresh({ id: restoreId });
} else {
this.selectSession(this.sessionOrder[0], { auto: true });
this.activeSessionId = null;
// `auto`: the app is restoring a session on load, not a human opening
// one, so a pending idle alert on that tab stays armed until it is
// actually tapped (see the userInitiated note in selectSession).
if (restoreId && this.sessions.has(restoreId)) {
this.selectSession(restoreId, { auto: true });
} else {
this.selectSession(this.sessionOrder[0], { auto: true });
}
}
}
}
+2
View File
@@ -102,6 +102,8 @@
'Manage AI Coding tools in persistent tmux sessions.': '在持久化 tmux 会话中管理 AI 编程工具。',
'Select case': '选择案例',
'Select Case': '选择案例',
'Search cases': '搜索案例',
'No cases match': '没有匹配的案例',
'All cases': '全部案例',
'No directory': '未选择目录',
Run: '运行',
+37 -5
View File
@@ -692,11 +692,9 @@
<button class="btn-toolbar btn-enter" onclick="app.sendEnterKey()" title="Send Enter">
Enter
</button>
<div class="tab-count-group" title="Instance count">
<button class="tab-count-btn" onclick="app.decrementShellCount()">−</button>
<input type="number" id="shellCount" class="tab-count-input" value="1" min="1" max="20" readonly>
<button class="tab-count-btn" onclick="app.incrementShellCount()">+</button>
</div>
<!-- Run Shell had a second, identical instance-count stepper here. The
toolbar carried two of them side by side, so it is gone and Run
Shell reads the one above (#tabCount) like the Run button does. -->
<div class="case-select-group">
<div class="case-combobox" id="quickStartCasePicker">
<input
@@ -2906,6 +2904,13 @@
<label class="checkbox-row"><input type="checkbox" id="dockerAdoptExisting"> Attach to an existing container</label>
<span class="form-hint">On: Codeman only runs docker exec into a container you already built and run — it never creates, starts, stops or removes it. The CLIs must already be installed and logged in inside it.</span>
</div>
<div class="form-row docker-adopt-only" id="dockerAdoptCloneRow" hidden>
<label>Duplicate an Existing Case</label>
<select id="dockerAdoptCloneFrom" onchange="app.applyDockerCloneSource()">
<option value="">Start from scratch</option>
</select>
<span class="form-hint">Same container, another directory inside it. Picks up the container, host and workspace below &mdash; you only set a new name and container workdir. Adopted containers only: an owned container&#39;s lifecycle belongs to its one case.</span>
</div>
<div class="form-row docker-adopt-only">
<label>Container Name</label>
<input type="text" id="dockerContainerName" list="dockerContainerList" placeholder="my-dev-box" pattern="[a-zA-Z0-9][a-zA-Z0-9_.-]+" autocomplete="off" autocapitalize="off" spellcheck="false">
@@ -3010,6 +3015,33 @@
<h3>Select Case</h3>
<button class="modal-close" onclick="app.closeMobileCasePicker()" aria-label="Close case picker">&times;</button>
</div>
<!-- Search: the desktop toolbar has had a filtering combobox for a while
(#quickStartCaseSearch); this is the same matcher on the phone sheet,
where a long case list is otherwise a long scroll. -->
<div class="mobile-case-picker-search">
<span class="mobile-case-search-icon" aria-hidden="true">
<svg width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
<circle cx="11" cy="11" r="7"/><line x1="21" y1="21" x2="16.65" y2="16.65"/>
</svg>
</span>
<input
type="search"
id="mobileCaseSearch"
class="mobile-case-search-input"
placeholder="Search cases"
aria-label="Search cases"
aria-controls="mobileCaseList"
autocomplete="off"
autocapitalize="none"
autocorrect="off"
spellcheck="false"
enterkeyhint="search"
oninput="app.filterMobileCaseList()"
onkeydown="app.handleMobileCaseSearchKeydown(event)"
>
<button type="button" class="mobile-case-search-clear" id="mobileCaseSearchClear"
onclick="app.clearMobileCaseSearch()" aria-label="Clear search" hidden>&times;</button>
</div>
<div class="mobile-case-picker-body">
<div class="mobile-case-list" id="mobileCaseList">
<!-- Cases populated by JS -->
+46
View File
@@ -120,6 +120,41 @@ const CjkInput = (() => {
c: '\x03', d: '\x04', l: '\x0c', z: '\x1a', a: '\x01', e: '\x05',
};
/** CSI final byte per navigation key, for the modifier-carrying forms below. */
const CSI_NAV_FINAL = {
ArrowUp: 'A',
ArrowDown: 'B',
ArrowRight: 'C',
ArrowLeft: 'D',
End: 'F',
Home: 'H',
};
/**
* The `CSI 1 ; <mod> <final>` form for a Ctrl/Alt-modified navigation key, or
* null when this key is not one.
*
* A modified navigation key is a terminal COMMAND, not text editing — claude's
* own "Jump to bottom (ctrl+End)" is one. PASSTHROUGH_KEYS carries only the
* plain forms, so Ctrl+End used to fail in BOTH directions: with an empty
* field it was sent as a bare `\x1b[F` (the modifier silently dropped, so the
* CLI saw a plain End), and with any text in the field it was not forwarded at
* all and the browser's default moved the caret to the end of the composer,
* which is what the user sees as "the shortcut does something to the input box
* instead".
*
* ⚠️ Shift ALONE is deliberately excluded: Shift+arrow selects text inside the
* composer, which is a real editing gesture worth keeping local. Shift is still
* encoded when it accompanies Ctrl or Alt.
*/
function _modifiedNavSequence(e) {
const final = CSI_NAV_FINAL[e.key];
if (!final) return null;
if (!e.ctrlKey && !e.altKey) return null;
const mod = 1 + (e.shiftKey ? 1 : 0) + (e.altKey ? 2 : 0) + (e.ctrlKey ? 4 : 0);
return `\x1b[1;${mod}${final}`;
}
function _strip(str) {
return str.replace(/​/g, '');
}
@@ -321,6 +356,17 @@ const CjkInput = (() => {
return;
}
// Ctrl/Alt-modified navigation keys go to the PTY REGARDLESS of whether
// the field has text: they are commands for the CLI, and the composer has
// no editing behaviour for them worth preserving (plain Home/End still
// edit locally through the table below).
const modNav = _modifiedNavSequence(e);
if (modNav) {
e.preventDefault();
_send(modNav);
return;
}
// Arrow/function keys: forward to PTY when no real text
if (PASSTHROUGH_KEYS[e.key] && _isEffectivelyEmpty()) {
e.preventDefault();
+38 -3
View File
@@ -5,7 +5,9 @@
*
* - KeyboardAccessoryBar (singleton object) — Quick action buttons shown above the virtual
* keyboard on mobile: arrow up/down, /init, Tab, paste, Esc, and dismiss (the extended
* bar adds /clear, /compact, Shift+Tab and more). Tab flushes any locally-buffered
* bar adds /clear, /compact, Shift+Tab and more). Shift+Left/Right ship in both agent
* layouts but are revealed only on Codex sessions (`codex-enabled` marker class on the
* bar, synced on every session switch), since they are Codex bindings. Tab flushes any locally-buffered
* prompt text to the PTY before sending \t, so completion applies to what was typed.
* The paste button opens a dialog that handles both text paste and image attach
* (native picker + best-effort image paste, routed through app._uploadAndInsertImages).
@@ -661,6 +663,8 @@ const KeyboardAccessoryBar = {
</button>
<button class="accessory-btn" data-action="init" title="/init">/init</button>
<button class="accessory-btn" data-action="tab" title="Tab">Tab</button>
<button class="accessory-btn accessory-btn-codex" data-action="shift-left" title="Shift+Left (Codex: edit queued message)" aria-label="Shift+Left (Codex: edit queued message)">⇧←</button>
<button class="accessory-btn accessory-btn-codex" data-action="shift-right" title="Shift+Right (Codex: prompt stack back)" aria-label="Shift+Right (Codex: prompt stack back)">⇧→</button>
<button class="accessory-btn" data-action="paste" title="Paste from clipboard">
<svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
<path d="M16 4h2a2 2 0 0 1 2 2v14a2 2 0 0 1-2 2H6a2 2 0 0 1-2-2V6a2 2 0 0 1 2-2h2"/>
@@ -746,6 +750,8 @@ const KeyboardAccessoryBar = {
<button class="accessory-btn" data-action="clear-input" title="Clear the current unsent input">&#x232B; All</button>
<button class="accessory-btn accessory-btn-rmm" data-action="readmymind" title="Read My Mind: predict your next prompt">🧠</button>
<button class="accessory-btn" data-action="tab" title="Tab">Tab</button>
<button class="accessory-btn accessory-btn-codex" data-action="shift-left" title="Shift+Left (Codex: edit queued message)" aria-label="Shift+Left (Codex: edit queued message)">⇧←</button>
<button class="accessory-btn accessory-btn-codex" data-action="shift-right" title="Shift+Right (Codex: prompt stack back)" aria-label="Shift+Right (Codex: prompt stack back)">⇧→</button>
<button class="accessory-btn" data-action="shift-tab" title="Shift+Tab">⇧Tab</button>
<button class="accessory-btn" data-action="effort-max" title="/effort max">Max</button>
<button class="accessory-btn" data-action="ctrl-o" title="Ctrl+O">⌃O</button>
@@ -772,6 +778,9 @@ const KeyboardAccessoryBar = {
// The 🧠 key is opt-in (`readMyMindEnabled`, synced): it ships in both
// templates but stays display:none until the bar carries the marker class.
this.syncReadMyMind();
// The ⇧←/⇧→ keys are Codex bindings: same shape, gated on the active
// session's mode instead of a setting.
this.syncCodexKeys();
// Add click handlers — preventDefault stops event from reaching terminal
this.element.addEventListener('click', (e) => {
@@ -784,7 +793,7 @@ const KeyboardAccessoryBar = {
this.handleAction(action, btn);
// Refocus terminal so keyboard stays open (tap blurs terminal → keyboard dismisses → toolbar shifts)
const refocusActions = new Set(['scroll-up', 'scroll-down', 'arrow-left', 'arrow-right', 'tab', 'shift-tab', 'ctrl', 'ctrl-o', 'opt-enter', 'esc', 'effort-max', 'clear-input']);
const refocusActions = new Set(['scroll-up', 'scroll-down', 'arrow-left', 'arrow-right', 'tab', 'shift-tab', 'shift-left', 'shift-right', 'ctrl', 'ctrl-o', 'opt-enter', 'esc', 'effort-max', 'clear-input']);
if (refocusActions.has(action) ||
((action === 'clear' || action === 'compact') && this._confirmAction)) {
if (typeof app !== 'undefined' && app.terminal) {
@@ -815,6 +824,7 @@ const KeyboardAccessoryBar = {
refreshForActiveSession() {
this.clearCtrl();
this._applyLayout(this._resolveMode());
this.syncCodexKeys();
},
/** Which layout the current state calls for. */
@@ -827,6 +837,11 @@ const KeyboardAccessoryBar = {
return app.sessions?.get(app.activeSessionId)?.mode === 'shell';
},
_isCodexSession() {
if (typeof app === 'undefined' || !app.activeSessionId) return false;
return app.sessions?.get(app.activeSessionId)?.mode === 'codex';
},
/** Swap the button set in the DOM. */
_applyLayout(mode) {
if (!this.element || mode === this._mode) return;
@@ -913,6 +928,12 @@ const KeyboardAccessoryBar = {
case 'arrow-right':
this.sendNavKey('\x1b[C');
break;
case 'shift-left':
this.sendNavKey('\x1b[1;2D');
break;
case 'shift-right':
this.sendNavKey('\x1b[1;2C');
break;
case 'esc':
this.sendKey('\x1b');
break;
@@ -1009,6 +1030,20 @@ const KeyboardAccessoryBar = {
this.element.classList.toggle('rmm-enabled', enabled === true);
},
/** Reveal the ⇧←/⇧→ keys only while the active session runs Codex. They are
* Codex bindings (edit the last queued message / prompt stack back) and do
* nothing in any other CLI, yet a tap still goes through sendNavKey(), which
* hands the session to plain PTY echo for the rest of the prompt, so on a
* phone a dead key would also switch off local echo. Same marker-class
* shape as syncReadMyMind(): the class lives on the BAR because setMode()
* rebuilds the buttons' innerHTML. Synced at init and on every session
* switch (refreshForActiveSession); a session's mode is fixed at create, so
* no other event can change the answer. */
syncCodexKeys() {
if (!this.element) return;
this.element.classList.toggle('codex-enabled', this._isCodexSession());
},
/** Send a slash command to the active session.
* Sends text and Enter separately so Ink processes them as distinct events. */
sendCommand(command) {
@@ -1044,7 +1079,7 @@ const KeyboardAccessoryBar = {
},
/**
* A composer nav key (the four arrows) from the bar, under the SAME contract
* A composer nav key (arrows, including Shift+Left/Right) from the bar, under the SAME contract
* as pressing one on a hardware keyboard (the `isComposerNavKey` branch of
* terminal-ui.js's onData): flush the unsent draft so the key edits the real
* composer, then hand the session to plain PTY echo until Enter or Ctrl+C,
+20
View File
@@ -410,6 +410,12 @@ const KeyboardHandler = {
const keyboardHeight = this.initialViewportHeight - (window.visualViewport.height || window.innerHeight);
const accessoryBar = document.querySelector('.keyboard-accessory-bar');
// The mobile case picker is a third position:fixed bottom-anchored
// surface, and since it gained a search field the keyboard can open over
// it. iOS does not shrink the layout viewport, so an unlifted sheet sits
// BEHIND the keyboard with its own search box out of sight.
const caseSheet = document.querySelector('.mobile-case-picker.active .mobile-case-picker-sheet');
if (isSmallMedium) {
// Phones/small tablets: toolbar and accessory bar are position:fixed
// via CSS. Use translateY to lift them above the keyboard.
@@ -426,6 +432,9 @@ const KeyboardHandler = {
if (accessoryBar) {
accessoryBar.style.transform = keyboardOffset > 0 ? `translateY(${-keyboardOffset}px)` : '';
}
if (caseSheet) {
caseSheet.style.transform = keyboardOffset > 0 ? `translateY(${-keyboardOffset}px)` : '';
}
if (main && keyboardHeight > 0) {
const cjkInputHeight = cjkInput?.classList.contains('cjk-input-visible') ? 44 : 0;
main.style.paddingBottom = `${84 + cjkInputHeight}px`;
@@ -436,6 +445,9 @@ const KeyboardHandler = {
if (accessoryBar) {
accessoryBar.style.bottom = `${keyboardHeight}px`;
}
if (caseSheet) {
caseSheet.style.bottom = `${keyboardHeight}px`;
}
}
// CJK textarea positioning (always position:fixed on touch devices).
@@ -464,6 +476,10 @@ const KeyboardHandler = {
const accessoryBar = document.querySelector('.keyboard-accessory-bar');
const cjkInput = document.getElementById('cjkInput');
const main = document.querySelector('.main');
// Not scoped to `.active`, unlike the lift above: a sheet closed while the
// keyboard was still up must still have its inline offset cleared, or the
// next open slides in already displaced.
const caseSheet = document.querySelector('.mobile-case-picker-sheet');
if (toolbar) {
toolbar.style.transform = '';
@@ -476,6 +492,10 @@ const KeyboardHandler = {
cjkInput.style.transform = '';
cjkInput.style.bottom = '';
}
if (caseSheet) {
caseSheet.style.transform = '';
caseSheet.style.bottom = '';
}
if (main) {
main.style.paddingBottom = '';
}
+12
View File
@@ -2164,6 +2164,18 @@ html.mobile-init .file-browser-panel {
font-size: 1.5rem;
}
/* With the keyboard up the sheet is lifted above it (mobile-handlers.js), so
what is left to fit is much shorter than 60vh of the layout viewport. Cap
the list rather than the sheet, so the search row and the Create button
stay on screen and only the rows scroll. */
.keyboard-visible .mobile-case-picker-sheet {
max-height: 45vh;
}
.keyboard-visible .mobile-case-picker-body {
max-height: 28vh;
}
.mobile-case-picker-footer {
padding-bottom: calc(12px + var(--safe-area-bottom));
}
+209 -23
View File
@@ -869,17 +869,17 @@ Object.assign(CodemanApp.prototype, {
input.value = Math.max(1, current - 1);
},
// Shell count stepper functions
incrementShellCount() {
const input = document.getElementById('shellCount');
const current = parseInt(input.value) || 1;
input.value = Math.min(20, current + 1);
},
decrementShellCount() {
const input = document.getElementById('shellCount');
const current = parseInt(input.value) || 1;
input.value = Math.max(1, current - 1);
/**
* How many sessions the next launch creates, from the toolbar's single
* instance stepper. Run Shell used to carry a second, identical `− 1 +` group
* of its own (`#shellCount`); that one is gone, so both launch paths read
* this control. An absent stepper reads as 1 rather than throwing: the group
* is display:none on phones and tablets, and the vm-based unit tests stub
* only the elements they exercise.
*/
_toolbarInstanceCount() {
const raw = document.getElementById('tabCount')?.value;
return Math.min(20, Math.max(1, parseInt(raw, 10) || 1));
},
// Next free <prefix><n> index for a case's session tabs (e.g. w1-<case>,
@@ -930,7 +930,7 @@ Object.assign(CodemanApp.prototype, {
async runClaude() {
const caseName = document.getElementById('quickStartCase').value || 'testcase';
const tabCount = Math.min(20, Math.max(1, parseInt(document.getElementById('tabCount').value) || 1));
const tabCount = this._toolbarInstanceCount();
const ownsLaunchTerminal = this._beginSessionLaunchStatus(
`Starting ${tabCount} Claude session(s) in ${caseName}...`
@@ -1062,13 +1062,6 @@ Object.assign(CodemanApp.prototype, {
...(hasEnvOverrides ? { envOverrides } : {}),
...(effort ? { effort } : {}),
...(modelOverride !== undefined ? { modelOverride } : {}),
// Plan-usage statusLine exporter (App Settings → Display). The server
// ADDS our exporter on create when true; when false it intentionally
// leaves any existing exporter in place (a per-repo settings.local.json
// is shared by sibling sessions, so create-with-false must not yank it
// — see the comment in session-routes create). Disabling the setting
// removes it via the App Settings toggle path (system-routes), not here.
statusLineTelemetry: this.planUsageChipEnabled(globalSettings),
})
}).then(r => r.json())
);
@@ -1149,7 +1142,7 @@ Object.assign(CodemanApp.prototype, {
async runShell() {
const caseName = document.getElementById('quickStartCase').value || 'testcase';
const shellCount = Math.min(20, Math.max(1, parseInt(document.getElementById('shellCount').value) || 1));
const shellCount = this._toolbarInstanceCount();
const ownsLaunchTerminal = this._beginSessionLaunchStatus(
`Starting ${shellCount} Shell session(s) in ${caseName}...`,
@@ -3160,7 +3153,10 @@ Object.assign(CodemanApp.prototype, {
const adopting = document.getElementById('dockerAdoptExisting')?.checked;
if (adopting) modal.setAttribute('data-docker-adopt', '1');
else modal.removeAttribute('data-docker-adopt');
if (adopting) void this._loadDockerContainerOptions();
if (adopting) {
void this._loadDockerContainerOptions();
void this._loadDockerCloneOptions();
}
},
/**
@@ -3172,6 +3168,120 @@ Object.assign(CodemanApp.prototype, {
* Best-effort by design — the endpoint returns [] for an unreachable daemon,
* and an empty list simply leaves the field as plain text input.
*/
/**
* Fill the "Duplicate an Existing Case" picker with the ADOPTED docker cases.
*
* One adopted container can back several cases, each pointing at a different
* directory inside it (classifyAdoptContainerConflict) — but re-typing the
* container, host and workspace by hand for every directory is exactly the
* friction that makes the capability go unused. Picking a case here fills those
* three and leaves only the two fields that MUST differ: the case name and the
* container workdir.
*
* ⚠️ Adopted cases only (`docker.owned === false`). An owned container's
* lifecycle belongs to its one case — a second case on it would be destroyed
* out from under itself by that case's recreate or delete — and the server
* refuses it, so offering it here would only produce a confusing error.
*/
async _loadDockerCloneOptions() {
const select = document.getElementById('dockerAdoptCloneFrom');
const row = document.getElementById('dockerAdoptCloneRow');
if (!select || !row) return;
let cases = [];
try {
const res = await fetch('/api/cases');
const data = await res.json();
cases = (Array.isArray(data) ? data : data?.data || []).filter(
(c) => c?.docker && c.docker.owned === false
);
} catch {
cases = [];
}
select.textContent = '';
const blank = document.createElement('option');
blank.value = '';
blank.textContent = 'Start from scratch';
select.appendChild(blank);
for (const c of cases) {
const option = document.createElement('option');
option.value = c.name;
// Server-supplied strings: textContent, never markup.
option.textContent = `${c.name} — ${c.docker.container}:${c.docker.containerWorkdir || c.docker.path}`;
option.dataset.container = c.docker.container;
option.dataset.hostId = c.docker.hostId;
option.dataset.path = c.docker.path;
option.dataset.workdir = c.docker.containerWorkdir || c.docker.path;
select.appendChild(option);
}
// Nothing to duplicate yet: an empty picker is noise on the first adoption.
row.hidden = cases.length === 0;
},
/**
* Apply the picked case: carry over what STAYS the same, clear what must not.
*
* The two cleared fields are the point of the feature — a duplicate that kept
* the original's name would be rejected as an existing case, and one that kept
* its container workdir would be rejected as an exact twin (both by the server,
* with a clear message, but a form that pre-fills a value it knows will be
* refused is just a trap).
*/
applyDockerCloneSource() {
const select = document.getElementById('dockerAdoptCloneFrom');
const option = select?.selectedOptions?.[0];
if (!option || !option.value) return;
const set = (id, value) => {
const el = document.getElementById(id);
if (el) el.value = value || '';
};
set('dockerContainerName', option.dataset.container);
set('dockerHostId', option.dataset.hostId);
set('dockerWorkspacePath', option.dataset.path);
// Pre-filled, NOT cleared: these two must differ from the source, but editing
// `/srv/app/api` into `/srv/app/web` beats retyping a long path, and the same
// goes for the name. What keeps a duplicate from being submitted unchanged is
// the guard below (dockerCloneGuard), which is a better trade than an empty
// field: the form stays a starting point instead of a blank form with three
// fields mysteriously filled in.
set('dockerCaseName', option.value);
set('dockerAdoptWorkdir', option.dataset.workdir);
// Remembered so the guard can tell "unchanged" from "happens to look similar".
select.dataset.appliedName = option.value;
select.dataset.appliedWorkdir = option.dataset.workdir || '';
const workdir = document.getElementById('dockerAdoptWorkdir');
workdir?.focus();
// Caret at the end: the tail is the part that changes.
if (workdir) workdir.setSelectionRange(workdir.value.length, workdir.value.length);
},
/**
* Refuse a duplicate that still carries the source case's name or directory.
*
* Both are pre-filled so they can be EDITED, which means both can also be left
* alone by accident. The server refuses either (an existing case name, or an
* exact same-container-same-directory twin) with a clear message, but a
* round-trip to be told "you forgot to change the field you were looking at" is
* worse than saying so here, next to the field, before anything is sent.
*
* Returns the offending element, or null when the form is fine.
*/
dockerCloneGuard() {
const select = document.getElementById('dockerAdoptCloneFrom');
if (!select || !select.value) return null;
const name = document.getElementById('dockerCaseName');
const workdir = document.getElementById('dockerAdoptWorkdir');
if (name && name.value.trim() === (select.dataset.appliedName || '')) {
return { el: name, message: `"${name.value.trim()}" is the case you copied from — give this one a new name.` };
}
if (workdir && workdir.value.trim() === (select.dataset.appliedWorkdir || '')) {
return {
el: workdir,
message: 'Same container and same directory as the case you copied from — point this one at another directory.',
};
}
return null;
},
async _loadDockerContainerOptions() {
const list = document.getElementById('dockerContainerList');
if (!list) return;
@@ -3272,6 +3382,17 @@ Object.assign(CodemanApp.prototype, {
this.showToast('Enter the name of the running container to attach to', 'error');
return;
}
// A duplicate that still carries the source's name or directory: say so here,
// beside the field, rather than sending a request certain to come back refused.
const cloneIssue = adopting ? this.dockerCloneGuard() : null;
if (cloneIssue) {
this.showToast(cloneIssue.message, 'error');
const statusEl = document.getElementById('dockerLinkStatus');
if (statusEl) statusEl.textContent = cloneIssue.message;
cloneIssue.el.focus();
cloneIssue.el.select?.();
return;
}
try {
if (statusEl) {
@@ -3795,13 +3916,42 @@ Object.assign(CodemanApp.prototype, {
showMobileCasePicker() {
const modal = document.getElementById('mobileCasePickerModal');
const search = document.getElementById('mobileCaseSearch');
// Every open starts unfiltered: the sheet is a one-shot picker, and a query
// left over from last time would present a truncated case list as the whole
// one. Deliberately no autofocus: focusing raises the keyboard over a sheet
// that is anchored to the bottom of the screen, so the user asks for it.
this._mobileCaseFilter = '';
if (search) search.value = '';
this.renderMobileCaseList();
modal.classList.add('active');
},
/** Re-render the sheet's list for the current search text. */
renderMobileCaseList() {
const listContainer = document.getElementById('mobileCaseList');
const select = document.getElementById('quickStartCase');
if (!listContainer || !select) return;
const currentCase = select.value;
const clearBtn = document.getElementById('mobileCaseSearchClear');
const filter = this._mobileCaseFilter || '';
if (clearBtn) clearBtn.hidden = filter.length === 0;
// Same matcher the desktop combobox uses (every term must appear in the
// option's searchText, which carries the name, label, path and the
// remote/docker fields), so both pickers answer a query identically.
const allCases = this.filterCasePickerOptions(this.getCasePickerOptions(), filter);
if (allCases.length === 0) {
listContainer.innerHTML = '<div class="mobile-case-empty">No cases match</div>';
return;
}
// Build case list HTML
let html = '';
const allCases = this.getCasePickerOptions();
for (const c of allCases) {
const isSelected = c.name === currentCase;
@@ -3829,7 +3979,43 @@ Object.assign(CodemanApp.prototype, {
}
listContainer.innerHTML = html;
modal.classList.add('active');
},
/** oninput on the sheet's search field. */
filterMobileCaseList() {
const search = document.getElementById('mobileCaseSearch');
this._mobileCaseFilter = search?.value || '';
this.renderMobileCaseList();
},
clearMobileCaseSearch() {
const search = document.getElementById('mobileCaseSearch');
if (search) search.value = '';
this._mobileCaseFilter = '';
this.renderMobileCaseList();
search?.focus();
},
handleMobileCaseSearchKeydown(event) {
if (event.key === 'Enter') {
// A search that narrowed to one case is an unambiguous choice, so Enter
// takes it instead of leaving the user to reach past the keyboard for a
// single row. Several matches just dismiss the keyboard.
event.preventDefault();
const matches = this.filterCasePickerOptions(this.getCasePickerOptions(), this._mobileCaseFilter || '');
if (matches.length === 1) {
this.selectMobileCase(matches[0].name);
} else {
event.target?.blur?.();
}
} else if (event.key === 'Escape') {
// Swallowed: the document-level Escape handler closes the whole sheet, and
// the first Escape here means "drop the filter", not "give up on picking".
event.preventDefault();
event.stopPropagation();
if (this._mobileCaseFilter) this.clearMobileCaseSearch();
else this.closeMobileCasePicker();
}
},
closeMobileCasePicker() {
+38 -15
View File
@@ -2310,15 +2310,23 @@ Object.assign(CodemanApp.prototype, {
// Save to server (includes notification prefs for cross-browser persistence).
// Strip device-specific DISPLAY keys so they never sync across devices —
// localEcho/cjk/extendedKeyboard/skin are per-platform, and showPlanUsageLimits
// is per-device too (desktop can show the usage chip while mobile stays hidden).
// localEcho/cjk/extendedKeyboard/skin are per-platform.
// webglRendererEnabled is per-device as well (renderer choice is GPU-specific,
// and syncing would leak mobile's hidden-checkbox false onto desktop); it's
// also absent from SettingsUpdateSchema, which is .strict() — sending it
// would 400 the whole settings PUT.
// Telemetry COLLECTION is requested out-of-band via statusLineTelemetry (sent on
// ENABLE only, so a device with the chip OFF never strips the exporter that
// another device's chip depends on — see system-routes settings handler).
// showPlanUsageLimits is per-device for DISPLAY (loadAppSettingsFromServer
// only seeds it into localStorage when a device has no value yet, like every
// other display key) but ALSO doubles as the server-side plan-usage telemetry
// COLLECTION switch (readPlanUsageTelemetryEnabled in hooks-config.ts, read
// fresh at every claude session create/respawn). So it is stripped here like
// the others and re-added below ONLY when this save FLIPS it on this device
// (planUsageCollectionFlip): the chip defaults OFF on handhelds, so sending
// it on every save let a phone saving its font size persist `false` and
// switch collection off for every desktop, whose chip then went stale with
// no error anywhere. An explicit toggle on any device still writes it, in
// either direction.
const _chipFlip = this.planUsageCollectionFlip(_prev, settings.showPlanUsageLimits);
const {
localEchoEnabled: _leo,
cjkInputEnabled: _cjk,
@@ -2360,7 +2368,7 @@ Object.assign(CodemanApp.prototype, {
try {
const res = await this._apiPut('/api/settings', {
...serverSettings,
...(settings.showPlanUsageLimits ? { statusLineTelemetry: true } : {}),
...(_chipFlip !== undefined ? { showPlanUsageLimits: _chipFlip } : {}),
notificationPreferences: notifPrefsToSave,
voiceSettings,
});
@@ -2624,15 +2632,28 @@ Object.assign(CodemanApp.prototype, {
// Resolved per-device state of the plan-usage chip. Desktop defaults ON,
// handhelds default OFF (the mobile block in getDefaultSettings() sets false,
// and the mobile-header-buttons-policy guard depends on that staying false).
// Single source of truth for THREE call sites that must never disagree: the
// App Settings checkbox, the chip's visibility, and the statusLineTelemetry
// flag sent on session create. A chip shown without telemetry renders "—"
// forever, which is exactly the drift this helper prevents.
// Single source of truth for the two call sites that must never disagree:
// the App Settings checkbox and the chip's visibility. Telemetry COLLECTION
// no longer has a THIRD client-side call site here at all — the server reads
// this same persisted setting directly (readPlanUsageTelemetryEnabled in
// hooks-config.ts), fresh, at every claude session create/respawn.
planUsageChipEnabled(settings = null) {
const s = settings ?? this.loadAppSettingsFromStorage();
return s.showPlanUsageLimits ?? this.getDefaultSettings().showPlanUsageLimits ?? true;
},
// What a settings save tells the server about plan-usage COLLECTION: the new
// chip value when this save FLIPS it relative to what this device resolved
// before (stored value, else the per-device default), otherwise undefined,
// meaning "say nothing". The server reads an absent key as ON, so a device
// that never touched the chip leaves collection alone, and a handheld (chip
// default OFF) cannot switch it off for every desktop by saving its font
// size. Pure so test/plan-usage-collection-flip.test.ts can drive it.
planUsageCollectionFlip(prevSettings, now) {
const before = this.planUsageChipEnabled(prevSettings ?? {});
return now === before ? undefined : now;
},
applyHeaderVisibilitySettings() {
const settings = this.loadAppSettingsFromStorage();
const defaults = this.getDefaultSettings();
@@ -3127,11 +3148,13 @@ Object.assign(CodemanApp.prototype, {
'sessionLineageLines',
]);
// The plan-usage chip is a PER-DEVICE display setting (desktop default ON,
// handheld default OFF): desktop can show it while mobile stays hidden. It
// used to sync, so an older server.json may still carry a value — drop it
// so the server value is NEVER
// seeded into a device that didn't explicitly enable it (collection is handled
// separately via the statusLineTelemetry action, not this display flag).
// handheld default OFF): desktop can show it while mobile stays hidden. Drop
// the server's stored value here so it is NEVER seeded into a device that
// didn't explicitly enable it — even though this SAME setting also drives
// server-side telemetry collection now (readPlanUsageTelemetryEnabled in
// hooks-config.ts), that's a read the server does directly from settings.json
// at spawn time; it has nothing to do with what gets merged into THIS
// device's local display preference.
delete appSettings.showPlanUsageLimits;
// Merge settings: non-display keys always sync from server,
// display keys only seed from server when localStorage has no value
+130 -3
View File
@@ -6834,6 +6834,99 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
color: #fff;
}
/* Search row. Tokens only (no literal dark glass) so the light skins need no
override of their own; the sheet itself is already re-pointed at
var(--floating-bg) up in the skin block. */
.mobile-case-picker-search {
display: flex;
align-items: center;
gap: 8px;
margin: 12px 20px 4px;
padding: 0 10px;
background: var(--bg-input);
border: 1px solid var(--border-light);
border-radius: 10px;
}
.mobile-case-picker-search:focus-within {
border-color: var(--accent);
}
.mobile-case-search-icon {
display: flex;
align-items: center;
justify-content: center;
color: var(--text-dim);
flex-shrink: 0;
}
.mobile-case-search-input {
flex: 1;
min-width: 0;
border: none;
background: transparent;
color: var(--text);
/* 16px: anything smaller makes iOS Safari zoom the page on focus, which
leaves the sheet scrolled off-centre when the field is blurred again. */
font-size: 16px;
font-family: inherit;
padding: 11px 0;
outline: none;
}
.mobile-case-search-input::placeholder {
color: var(--text-dim);
}
/* The global input:focus-visible rule paints a 1px accent ring, which inside
an already-bordered row draws a second border a few pixels in. The row's
:focus-within border is the focus cue here, so the input drops its own. */
.mobile-case-search-input:focus-visible {
box-shadow: none;
}
/* The native affordance sits in a different spot per engine and is absent on
Android, so the sheet ships its own clear button and hides this one. */
.mobile-case-search-input::-webkit-search-cancel-button,
.mobile-case-search-input::-webkit-search-decoration {
-webkit-appearance: none;
appearance: none;
}
.mobile-case-search-clear {
flex-shrink: 0;
width: 28px;
height: 28px;
display: flex;
align-items: center;
justify-content: center;
border: none;
border-radius: 6px;
background: transparent;
color: var(--text-dim);
font-size: 1.2rem;
line-height: 1;
cursor: pointer;
}
.mobile-case-search-clear:active {
background: var(--bg-hover);
color: var(--text);
}
/* The UA's [hidden] rule is display:none at specificity (0,0,0) and loses to the
display:flex above, so the button has to hide itself explicitly. */
.mobile-case-search-clear[hidden] {
display: none;
}
.mobile-case-empty {
padding: 22px 20px;
text-align: center;
color: var(--text-dim);
font-size: 0.9rem;
}
.mobile-case-picker-body {
flex: 1;
overflow-y: auto;
@@ -10495,14 +10588,19 @@ kbd {
position: fixed;
inset: 0;
background: var(--modal-backdrop);
backdrop-filter: blur(6px);
-webkit-backdrop-filter: blur(6px);
z-index: 5100;
display: none;
align-items: center;
justify-content: center;
}
/* Same stale-hit-test reasoning as .offline-overlay above: this one is also a
persistent full-screen fixed element, shown by adding `.visible`. */
.file-preview-overlay.visible {
backdrop-filter: blur(8px);
-webkit-backdrop-filter: blur(8px);
}
.file-preview-overlay.visible {
display: flex;
}
@@ -11969,6 +12067,18 @@ kbd {
display: inline-flex;
}
/* Keyboard-accessory ⇧←/⇧→ keys: Codex bindings (edit the last queued
message / prompt stack back), so they are revealed only while the active
session runs Codex. Same marker-class shape as the 🧠 key above, for the
same reason (the bar's innerHTML is rebuilt on every layout switch); the
class is synced from the active session's mode (keyboard-accessory.js). */
.keyboard-accessory-bar .accessory-btn-codex {
display: none;
}
.keyboard-accessory-bar.codex-enabled .accessory-btn-codex {
display: inline-flex;
}
.approvals-badge {
position: absolute;
top: 2px;
@@ -15304,9 +15414,26 @@ html[data-skin="daylight-blue"] .welcome-btn-tunnel.active:hover {
padding-top: calc(20px + var(--safe-area-top));
padding-bottom: calc(20px + var(--safe-area-bottom));
background: rgba(6, 8, 12, 0.93);
overflow-y: auto;
}
/* ⚠️ `backdrop-filter` is applied ONLY while the overlay is actually shown.
It promotes the element to its own compositing layer, and a full-screen
`position: fixed` layer that is created and then hidden has been observed to
leave a STALE HIT-TEST REGION behind in Chrome: the page keeps rendering
correctly while every pointer event over the viewport lands on nothing.
Symptom (reported on a long-lived tab against a remote server, where a
connection blip shows and then hides #offlineOverlay): the terminal stops
scrolling AND unrelated click-to-expand controls stop responding at the same
time, while a freshly opened tab is fine — and a console one-liner that only
READS layout (getComputedStyle + elementFromPoint, both of which force a
hit-test recompute) restores it. Two unrelated features dying together, and a
read-only command curing them, is what points at hit-testing rather than at
either feature. Keeping the property off the hidden state means the layer is
never created while invisible. */
.offline-overlay:not([hidden]) {
backdrop-filter: blur(6px);
-webkit-backdrop-filter: blur(6px);
overflow-y: auto;
}
.offline-overlay[hidden] {
+80 -7
View File
@@ -179,14 +179,23 @@
// theme, so default behavior is unchanged. Shared at module scope and exported on the
// global so both terminal-ui.js (main terminal) and panels-ui.js (teammate terminals,
// a separate IIFE) can read the current skin's palette.
//
// ⚠️ The selection key is `selectionBackground`, xterm's name for it since v5 (#360).
// An ITheme is a plain object handed straight to xterm, so an unknown key is not an
// error, it is silently dropped: every palette here carried `selection`, so every skin
// drew xterm's built-in default instead, rgba(255,255,255,0.3). On the four light skins
// that is white at 30% over a near-white background, a delta of about 3/255 — the
// highlight was effectively invisible, which is what a long-press selection that
// "did nothing" actually looked like. A key only works here if xterm knows its name;
// test/skin-themes.test.ts pins the name AND that the blend stays visible.
const CODEMAN_XTERM_THEMES = {
og: { background: '#0d0d0d', foreground: '#e0e0e0', cursor: '#e0e0e0', cursorAccent: '#0d0d0d', selection: 'rgba(255,255,255,0.3)', black: '#0d0d0d', red: '#ff6b6b', green: '#51cf66', yellow: '#ffd43b', blue: '#339af0', magenta: '#cc5de8', cyan: '#22b8cf', white: '#e0e0e0', brightBlack: '#495057', brightRed: '#ff8787', brightGreen: '#69db7c', brightYellow: '#ffe066', brightBlue: '#5c7cfa', brightMagenta: '#da77f2', brightCyan: '#66d9e8', brightWhite: '#ffffff' },
'daylight-green': { background: '#161b23', foreground: '#dfe6ef', cursor: '#2fd3aa', cursorAccent: '#161b23', selection: 'rgba(47,211,170,0.22)', black: '#161b23', red: '#ff8585', green: '#34d8a0', yellow: '#f0c25a', blue: '#5cc6e8', magenta: '#c79af2', cyan: '#2bcbbb', white: '#dfe6ef', brightBlack: '#5b6675', brightRed: '#ffa0a0', brightGreen: '#5fe6b8', brightYellow: '#ffd884', brightBlue: '#82d4ee', brightMagenta: '#d6b3f7', brightCyan: '#5ee0d4', brightWhite: '#f3f6fa' },
'daylight-blue': { background: '#161b23', foreground: '#dfe6ef', cursor: '#38b6f0', cursorAccent: '#161b23', selection: 'rgba(56,182,240,0.22)', black: '#161b23', red: '#ff8585', green: '#34d8a0', yellow: '#f0c25a', blue: '#5cc6e8', magenta: '#c79af2', cyan: '#2bcbbb', white: '#dfe6ef', brightBlack: '#5b6675', brightRed: '#ffa0a0', brightGreen: '#5fe6b8', brightYellow: '#ffd884', brightBlue: '#82d4ee', brightMagenta: '#d6b3f7', brightCyan: '#5ee0d4', brightWhite: '#f3f6fa' },
'paper-gray': { background: '#f6f8fa', foreground: '#1f2328', cursor: '#0969da', cursorAccent: '#ffffff', selection: 'rgba(9,105,218,0.2)', black: '#24292f', red: '#cf222e', green: '#1a7f37', yellow: '#9a6700', blue: '#0969da', magenta: '#8250df', cyan: '#1b7c83', white: '#59636e', brightBlack: '#6e7781', brightRed: '#a40e26', brightGreen: '#116329', brightYellow: '#7d4e00', brightBlue: '#0550ae', brightMagenta: '#6639ba', brightCyan: '#116b75', brightWhite: '#1f2328' },
'solarized-light': { background: '#fdf6e3', foreground: '#586e75', cursor: '#147ba3', cursorAccent: '#fdf6e3', selection: 'rgba(38,139,210,0.2)', black: '#eee8d5', red: '#dc322f', green: '#758600', yellow: '#9b7800', blue: '#147ba3', magenta: '#d33682', cyan: '#2a9189', white: '#073642', brightBlack: '#93a1a1', brightRed: '#cb4b16', brightGreen: '#657b83', brightYellow: '#586e75', brightBlue: '#268bd2', brightMagenta: '#6c71c4', brightCyan: '#2aa198', brightWhite: '#002b36' },
'catppuccin-latte': { background: '#eff1f5', foreground: '#4c4f69', cursor: '#1e66f5', cursorAccent: '#ffffff', selection: 'rgba(30,102,245,0.18)', black: '#5c5f77', red: '#d20f39', green: '#3b8f2b', yellow: '#a86605', blue: '#1e66f5', magenta: '#8839ef', cyan: '#177f86', white: '#6c6f85', brightBlack: '#7c7f93', brightRed: '#b50930', brightGreen: '#2f7622', brightYellow: '#8b5604', brightBlue: '#174fbf', brightMagenta: '#6f2bc5', brightCyan: '#116b71', brightWhite: '#4c4f69' },
'rose-pine-dawn': { background: '#faf4ed', foreground: '#575279', cursor: '#286983', cursorAccent: '#fffaf3', selection: 'rgba(40,105,131,0.2)', black: '#575279', red: '#b4637a', green: '#286983', yellow: '#96681f', blue: '#477f91', magenta: '#907aa9', cyan: '#3f7f8b', white: '#6e6a86', brightBlack: '#797593', brightRed: '#984d66', brightGreen: '#1f5266', brightYellow: '#7d5417', brightBlue: '#386b7c', brightMagenta: '#765f90', brightCyan: '#326b76', brightWhite: '#575279' },
og: { background: '#0d0d0d', foreground: '#e0e0e0', cursor: '#e0e0e0', cursorAccent: '#0d0d0d', selectionBackground: 'rgba(255,255,255,0.3)', black: '#0d0d0d', red: '#ff6b6b', green: '#51cf66', yellow: '#ffd43b', blue: '#339af0', magenta: '#cc5de8', cyan: '#22b8cf', white: '#e0e0e0', brightBlack: '#495057', brightRed: '#ff8787', brightGreen: '#69db7c', brightYellow: '#ffe066', brightBlue: '#5c7cfa', brightMagenta: '#da77f2', brightCyan: '#66d9e8', brightWhite: '#ffffff' },
'daylight-green': { background: '#161b23', foreground: '#dfe6ef', cursor: '#2fd3aa', cursorAccent: '#161b23', selectionBackground: 'rgba(47,211,170,0.22)', black: '#161b23', red: '#ff8585', green: '#34d8a0', yellow: '#f0c25a', blue: '#5cc6e8', magenta: '#c79af2', cyan: '#2bcbbb', white: '#dfe6ef', brightBlack: '#5b6675', brightRed: '#ffa0a0', brightGreen: '#5fe6b8', brightYellow: '#ffd884', brightBlue: '#82d4ee', brightMagenta: '#d6b3f7', brightCyan: '#5ee0d4', brightWhite: '#f3f6fa' },
'daylight-blue': { background: '#161b23', foreground: '#dfe6ef', cursor: '#38b6f0', cursorAccent: '#161b23', selectionBackground: 'rgba(56,182,240,0.22)', black: '#161b23', red: '#ff8585', green: '#34d8a0', yellow: '#f0c25a', blue: '#5cc6e8', magenta: '#c79af2', cyan: '#2bcbbb', white: '#dfe6ef', brightBlack: '#5b6675', brightRed: '#ffa0a0', brightGreen: '#5fe6b8', brightYellow: '#ffd884', brightBlue: '#82d4ee', brightMagenta: '#d6b3f7', brightCyan: '#5ee0d4', brightWhite: '#f3f6fa' },
'paper-gray': { background: '#f6f8fa', foreground: '#1f2328', cursor: '#0969da', cursorAccent: '#ffffff', selectionBackground: 'rgba(9,105,218,0.2)', black: '#24292f', red: '#cf222e', green: '#1a7f37', yellow: '#9a6700', blue: '#0969da', magenta: '#8250df', cyan: '#1b7c83', white: '#59636e', brightBlack: '#6e7781', brightRed: '#a40e26', brightGreen: '#116329', brightYellow: '#7d4e00', brightBlue: '#0550ae', brightMagenta: '#6639ba', brightCyan: '#116b75', brightWhite: '#1f2328' },
'solarized-light': { background: '#fdf6e3', foreground: '#586e75', cursor: '#147ba3', cursorAccent: '#fdf6e3', selectionBackground: 'rgba(38,139,210,0.2)', black: '#eee8d5', red: '#dc322f', green: '#758600', yellow: '#9b7800', blue: '#147ba3', magenta: '#d33682', cyan: '#2a9189', white: '#073642', brightBlack: '#93a1a1', brightRed: '#cb4b16', brightGreen: '#657b83', brightYellow: '#586e75', brightBlue: '#268bd2', brightMagenta: '#6c71c4', brightCyan: '#2aa198', brightWhite: '#002b36' },
'catppuccin-latte': { background: '#eff1f5', foreground: '#4c4f69', cursor: '#1e66f5', cursorAccent: '#ffffff', selectionBackground: 'rgba(30,102,245,0.18)', black: '#5c5f77', red: '#d20f39', green: '#3b8f2b', yellow: '#a86605', blue: '#1e66f5', magenta: '#8839ef', cyan: '#177f86', white: '#6c6f85', brightBlack: '#7c7f93', brightRed: '#b50930', brightGreen: '#2f7622', brightYellow: '#8b5604', brightBlue: '#174fbf', brightMagenta: '#6f2bc5', brightCyan: '#116b71', brightWhite: '#4c4f69' },
'rose-pine-dawn': { background: '#faf4ed', foreground: '#575279', cursor: '#286983', cursorAccent: '#fffaf3', selectionBackground: 'rgba(40,105,131,0.2)', black: '#575279', red: '#b4637a', green: '#286983', yellow: '#96681f', blue: '#477f91', magenta: '#907aa9', cyan: '#3f7f8b', white: '#6e6a86', brightBlack: '#797593', brightRed: '#984d66', brightGreen: '#1f5266', brightYellow: '#7d5417', brightBlue: '#386b7c', brightMagenta: '#765f90', brightCyan: '#326b76', brightWhite: '#575279' },
};
const CODEMAN_LIGHT_SKINS = new Set(['paper-gray', 'solarized-light', 'catppuccin-latte', 'rose-pine-dawn']);
function currentSkin() {
@@ -299,6 +308,7 @@ Object.assign(CodemanApp.prototype, {
const container = document.getElementById('terminalContainer');
this.terminal.open(container);
this._installMobileTapMouseGuard();
this._installShiftDragSelection();
this._installTouchSelectionFocusGuard();
// Let xterm's CompositionHelper own IME key events. In particular, a
@@ -902,7 +912,25 @@ Object.assign(CodemanApp.prototype, {
container.addEventListener('contextmenu', (ev) => {
if (longPressTimer !== null || this._touchSelecting || this._touchSelectionActive) {
ev.preventDefault();
return;
}
// Right-click COPIES the selection, the mintty/PuTTY convention, because
// the browser's own menu structurally cannot offer it here: xterm paints
// glyphs into a canvas, so a terminal selection is not a DOM selection
// and the native "Copy" item has nothing to act on (it is absent or
// inert). This is the second half of the habit users bring from a native
// terminal running a mouse-tracking TUI — Shift+drag to select (see
// _installShiftDragSelection), right-click to copy — and without it that
// gesture dead-ends after the selection is made.
//
// With NOTHING selected the native menu is left alone: it still carries
// the browser-level items (reload, inspect) and suppressing it there
// would take them away to offer nothing in return.
if (!this.terminal?.hasSelection?.()) return;
const selection = this.terminal.getSelection();
if (!selection) return;
ev.preventDefault();
void this.copyTerminalSelection(selection);
});
container.addEventListener(
@@ -4844,6 +4872,51 @@ Object.assign(CodemanApp.prototype, {
this._sendSyntheticSgrTap(ev.clientX, ev.clientY);
},
/**
* Make Shift+drag START a selection instead of trying to extend one.
*
* In a native terminal running a mouse-tracking TUI (claude, codex), Shift is
* the "let me select text" modifier: it bypasses the app's mouse reporting so
* the emulator selects locally. Users bring that habit here, and here it did
* NOTHING — Shift+drag selected no text at all (measured).
*
* The reason is that the habit and xterm's Shift mean different things once
* the DECSETs are stripped. xterm reads Shift as "force selection" ONLY while
* the app actually has mouse tracking on; the server strips those DECSETs for
* claude/codex/gemini (isAltScreenStripMode), so xterm's mouseTrackingMode is
* permanently `none`, that branch is unreachable, and Shift instead falls into
* `_onIncrementalClick` — EXTEND an existing selection. Extending is a no-op
* when `selectionStart` is null, so the drag never anchors and no selection is
* ever built (this is why nothing gets cleared: there was nothing to clear).
*
* So plant the anchor xterm is missing. Runs in the CAPTURE phase on the
* `.xterm` root, an ancestor of the `.xterm-screen` element SelectionService
* binds to, so it lands before xterm's own mousedown; xterm's incremental
* handler then extends from our anchor and the drag behaves like a plain one.
* A Shift+drag with a selection ALREADY up is left alone — that is a genuine
* extend gesture and xterm already does it right.
*/
_installShiftDragSelection() {
const el = this.terminal?.element;
if (!el || el._codemanShiftDragInstalled) return;
el._codemanShiftDragInstalled = true;
el.addEventListener(
'mousedown',
(ev) => {
if (!ev.isTrusted || ev.button !== 0 || !ev.shiftKey) return;
if (ev.altKey || ev.ctrlKey || ev.metaKey) return;
if (this.terminal?.hasSelection?.()) return;
const pos = this._clientPointToCell(ev.clientX, ev.clientY);
if (!pos) return;
// _clientPointToCell is 1-based and viewport-relative; select() takes a
// 0-based column and an ABSOLUTE buffer row.
const viewportY = this.terminal.buffer?.active?.viewportY ?? 0;
this.terminal.select(pos.col - 1, pos.row - 1 + viewportY, 0);
},
true
);
},
_installMobileTapMouseGuard() {
const el = this.terminal?.element;
if (!el || el._codemanTapMouseGuardInstalled) return;
+56 -4
View File
@@ -194,6 +194,7 @@ Object.assign(CodemanApp.prototype, {
this._webviewFrameLru = this._webviewFrameLru || [];
await this.refreshWebviews();
this._installWebviewLostListener();
// Restore the previously open web tabs (per device: which dashboards you keep
// open is a workspace-layout choice, not something to sync across machines).
@@ -231,6 +232,55 @@ Object.assign(CodemanApp.prototype, {
this.renderSessionTabs();
},
/**
* Take back a frame that navigated itself off its proxy prefix.
*
* The proxy's runtime shim masks `/webview/<cap>/` off the document URL so a
* single-page app routes on the path it expects. A navigation the page then
* starts itself — `location.reload()` (a dev server's full-reload HMR), a
* root-absolute `location.href = '/login'` — lands on Codeman's root with no
* capability, where the server answers a static page that does nothing but
* post `{type:'codeman:webview-lost', path}` here. The frame is identified by
* `event.source` against the iframes this tab mounted (never by the payload),
* and remounted inside the prefix at that path. Bounded per frame so a page
* that reloads itself on every boot cannot spin.
*/
_installWebviewLostListener() {
if (this._webviewLostListener) return;
this._webviewLostListener = (event) => {
const data = event.data;
if (!data || typeof data !== 'object' || data.type !== 'codeman:webview-lost') return;
if (typeof data.path !== 'string' || !event.source) return;
const layer = document.getElementById('webviewLayer');
if (!layer) return;
for (const wrap of layer.querySelectorAll('.webview-frame')) {
const frame = wrap.querySelector('iframe');
if (!frame || frame.contentWindow !== event.source) continue;
const id = wrap.dataset.webviewId;
if (!id || !this.webviews?.has(id)) return;
const now = Date.now();
this._webviewRecoveries = this._webviewRecoveries || new Map();
const recent = (this._webviewRecoveries.get(id) || []).filter((at) => now - at < 60000);
if (recent.length >= 5) return;
recent.push(now);
this._webviewRecoveries.set(id, recent);
// Path only, never an origin. Three spellings would resolve to a foreign
// origin (in direct mode `new URL(path, src)` is the frame's src, so the
// frame would remount there): the protocol-relative `//host/x`; a
// backslash, which the WHATWG parser treats as `/` for http(s), so
// `/\host/x` too; and an ASCII tab or newline, which the parser deletes
// before it looks at anything, so `/<tab>/host/x` IS `//host/x` by the time
// it resolves. Drop the invisible ones, collapse the leading separators to
// one `/`, and refuse whatever still opens a second one. The proxied form
// is refused server-side as well (resolveUpstreamUrl).
const path = data.path.replace(/[\t\n\r]/g, '').replace(/^[/\\]+/, '/');
void this.openWebview(id, { path: path.startsWith('/') && !/^\/[/\\]/.test(path) ? path : '/' });
return;
}
};
window.addEventListener('message', this._webviewLostListener);
},
_persistWebviewOrder() {
try {
localStorage.setItem('codeman-webview-order', JSON.stringify(this.webviewOrder || []));
@@ -328,13 +378,15 @@ Object.assign(CodemanApp.prototype, {
if (data.webview) this.webviews.set(id, data.webview);
let src = data.embedUrl || data.webview?.url || webview.url;
const path = typeof options.path === 'string' ? options.path : '';
// A string `path` (even '') means "go there": the proxy prefix is
// `/webview/<cap>/` and the wildcard rides after it; in direct mode the deep
// link resolves against the dashboard's own origin. No `path` means "show
// the tab", leaving a mounted frame on whatever page it reached.
const path = typeof options.path === 'string' ? options.path : null;
if (path) {
// The proxy prefix is `/webview/<cap>/`; a wildcard rides after it. In
// direct mode the deep link resolves against the dashboard's own origin.
src = data.embedUrl ? `${data.embedUrl.replace(/\/?$/, '/')}${path.replace(/^\//, '')}` : new URL(path, src).href;
}
this._mountWebviewFrame(id, src, data.webview || webview, { navigate: !!path });
this._mountWebviewFrame(id, src, data.webview || webview, { navigate: path !== null });
this.activeWebviewId = id;
this.hideWelcome?.();
document.querySelector('.main')?.classList.add('webview-active');
+34 -2
View File
@@ -88,11 +88,43 @@ export function validateSessionFilePath(
} catch {
return null;
}
const relativePath = relative(resolvedWorkingDir, resolvedPath);
return confineToRoot(resolvedWorkingDir, resolvedPath);
}
/**
* The lexical half of {@link validateSessionFilePath}: same containment rule, but
* WITHOUT touching the filesystem.
*
* Needed for remote-SSH cases (`src/remote-files.ts`), where `workingDir` is an
* absolute path on the REMOTE host and any local `realpathSync` fails by
* construction — which is how every file-raw/file-content request in a remote case
* used to end up as a 404 before a single byte was read. The caller follows this
* pre-check with a remote realpath + the same containment rule, so escapes are
* refused exactly as they are locally; what changes is only WHICH filesystem
* resolves the symlinks.
*
* A lexical check alone would follow nothing, so it must never be the last word for
* a path that can contain a symlink — it is the cheap reject in front of the real
* (local or remote) resolution, not a replacement for it.
*/
export function validateSessionFilePathLexical(
sessionWorkingDir: string,
filePath: string
): { resolvedPath: string; relativePath: string } | null {
return confineToRoot(resolve(sessionWorkingDir), resolve(sessionWorkingDir, filePath));
}
/**
* Shared containment rule: `candidate` must sit inside `root` (both already
* canonical for their filesystem). `relative()` is the whole test — a `..` or an
* absolute result means the candidate escaped.
*/
function confineToRoot(root: string, candidate: string): { resolvedPath: string; relativePath: string } | null {
const relativePath = relative(root, candidate);
if (relativePath.startsWith('..') || isAbsolute(relativePath)) {
return null;
}
return { resolvedPath, relativePath };
return { resolvedPath: candidate, relativePath };
}
// Maximum hook data size (prevents oversized SSE broadcasts)
+32 -5
View File
@@ -79,6 +79,7 @@ import {
dockerContainerName,
dockerDisplayPath,
probeAdoptableContainer,
classifyAdoptContainerConflict,
listDockerContainers,
browseInContainer,
dockerAdoptProbeModes,
@@ -327,6 +328,7 @@ export function registerCaseRoutes(app: FastifyInstance, ctx: EventPort & Config
container,
image: host.image,
path: dockerCase.hostWorkspacePath,
containerWorkdir: dockerCase.containerWorkdir ?? dockerCase.hostWorkspacePath,
network: host.network ?? 'bridge',
...(dockerCase.availableModes ? { availableModes: dockerCase.availableModes } : {}),
...(dockerCase.owned === false ? { owned: false } : {}),
@@ -906,12 +908,35 @@ export function registerCaseRoutes(app: FastifyInstance, ctx: EventPort & Config
) {
return createErrorResponse(ApiErrorCode.ALREADY_EXISTS, 'Case already exists');
}
// Two cases must never share one adopted container: session close kills the
// in-container tmux by session id, but a shared adoption would let one case's
// teardown and another's launch race over the same tmux server.
// One container may back SEVERAL adopted cases, each pointing at a different
// directory inside it. What still blocks it, and why, lives in
// classifyAdoptContainerConflict — note that none of it is about the shared
// in-container tmux server, which is safe precisely because sessions there
// are named per SESSION id (`codeman-dkr-<id8>`), never per case.
const container = dockerCase.container;
if (dockerCases.some((item) => (item.container ?? dockerContainerName(item.name)) === container)) {
return createErrorResponse(ApiErrorCode.ALREADY_EXISTS, `Container "${container}" is already linked to a case`);
const conflict = classifyAdoptContainerConflict({
container,
containerWorkdir: dockerCase.containerWorkdir ?? dockerCase.hostWorkspacePath,
existing: dockerCases,
canAccess: (owner) => canAccessOwned(getAuthUser(req), owner),
});
if (conflict?.kind === 'owned-case') {
return createErrorResponse(
ApiErrorCode.ALREADY_EXISTS,
`Container "${container}" belongs to case "${conflict.caseName}", which Codeman created and whose lifecycle it manages. Adopt a container you started yourself, or open that case directly.`
);
}
if (conflict?.kind === 'other-owner') {
return createErrorResponse(
ApiErrorCode.FORBIDDEN,
`Container "${container}" is already adopted by another user.`
);
}
if (conflict?.kind === 'duplicate') {
return createErrorResponse(
ApiErrorCode.ALREADY_EXISTS,
`Case "${conflict.caseName}" already adopts "${container}" at that same directory. Point this one at another directory inside the container.`
);
}
if (!isWorkingDirAllowed(getAuthUser(req), dockerCase.hostWorkspacePath)) {
@@ -1583,7 +1608,9 @@ export function registerCaseRoutes(app: FastifyInstance, ctx: EventPort & Config
container,
image: host.image,
path: dockerCase.hostWorkspacePath,
containerWorkdir: dockerCase.containerWorkdir ?? dockerCase.hostWorkspacePath,
network: host.network ?? 'bridge',
...(dockerCase.owned === false ? { owned: false } : {}),
},
};
}
+146
View File
@@ -0,0 +1,146 @@
/**
* @fileoverview Custom Model Endpoint Profiles CRUD + discovery
* (docs/custom-model-endpoints-plan.md). Endpoints are machine-level infra,
* like remote/docker hosts, so writes are admin-only in multi-user mode
* (`case-routes.ts`'s `/api/remote-hosts` is the pattern this mirrors).
*
* Discovery (`POST /:id/discover-models`) fetches `${baseUrl}/v1/models`
* through `webviewFetch()` (`webview-egress.ts`), the same guarded dispatcher
* the web-tab proxy uses: `baseUrl` is refused at save time by the schema's
* hostname check (link-local / cloud-metadata literals and names), and the
* undici lookup hook refuses a name that RESOLVES into one of those ranges at
* connect time, redirects included — a save-time hostname check alone would
* let `models.example` resolve to 169.254.169.254 later. The endpoint is
* admin-configured, so this is defence in depth rather than the only gate.
*/
import type { FastifyInstance, FastifyRequest } from 'fastify';
import { ApiErrorCode, createErrorResponse, type ApiResponse } from '../../types.js';
import { isAdmin, parseBody } from '../route-helpers.js';
import { isMultiUserMode } from '../../config/multiuser.js';
import { getDataDir } from '../../config/instance.js';
import { isBlockedWebviewUrl } from '../webview-egress-policy.js';
import { egressBlockedReason, webviewFetch } from '../webview-egress.js';
import { CustomModelHostSchema } from '../schemas.js';
import { readCustomModelHosts, writeCustomModelHosts, type CustomModelHost } from '../../custom-model-hosts.js';
const CODEMAN_CONFIG_DIR = getDataDir();
const DISCOVER_TIMEOUT_MS = 8000;
function adminOnly(req: FastifyRequest, reply: { code: (n: number) => unknown }): ApiResponse<never> | null {
if (!isMultiUserMode() || isAdmin(req)) return null;
reply.code(403);
return createErrorResponse(ApiErrorCode.FORBIDDEN, 'Admin only in multi-user mode');
}
async function discoverModels(host: Pick<CustomModelHost, 'baseUrl' | 'apiKey' | 'authStyle'>): Promise<string[]> {
const headers: Record<string, string> = {};
const apiKey = host.apiKey?.trim();
// Exactly ONE header, never both — see custom-model-hosts.ts's CustomModelAuthStyle
// doc comment for why: sending both reliably HANGS some real servers.
const style = host.authStyle ?? 'bearer';
if (apiKey && style === 'bearer') headers.Authorization = `Bearer ${apiKey}`;
if (apiKey && style === 'api-key') headers['api-key'] = apiKey;
const res = await webviewFetch(new URL(`${host.baseUrl.replace(/\/+$/, '')}/v1/models`), {
headers,
signal: AbortSignal.timeout(DISCOVER_TIMEOUT_MS),
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = (await res.json()) as { data?: Array<{ id?: unknown }> };
return (body.data ?? []).map((m) => m.id).filter((id): id is string => typeof id === 'string' && id.length > 0);
}
/**
* undici reports every network failure as `TypeError('fetch failed', { cause })`, with the
* useful part (`connect ECONNREFUSED 127.0.0.1:8080`) one level down; surface the deepest
* message so the user sees the refused connection, not the wrapper.
*/
function describeFetchError(err: unknown): string {
let message = err instanceof Error ? err.message : String(err);
let current: unknown = err;
for (let depth = 0; depth < 5 && current instanceof Error && current.cause !== undefined; depth++) {
current = current.cause;
if (current instanceof Error && current.message) message = current.message;
}
return message;
}
export function registerCustomModelRoutes(app: FastifyInstance): void {
app.get('/api/model-endpoints', async (req) =>
isMultiUserMode() && !isAdmin(req) ? [] : readCustomModelHosts(CODEMAN_CONFIG_DIR)
);
app.post('/api/model-endpoints', async (req, reply): Promise<ApiResponse<{ host: CustomModelHost }>> => {
const denied = adminOnly(req, reply);
if (denied) return denied;
const host = parseBody(CustomModelHostSchema, req.body);
if (isBlockedWebviewUrl(host.baseUrl)) {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
}
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
if (hosts.some((item) => item.id === host.id)) {
return createErrorResponse(ApiErrorCode.ALREADY_EXISTS, 'Model endpoint already exists');
}
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, [...hosts, host]);
return { success: true, data: { host } };
});
app.put('/api/model-endpoints/:id', async (req, reply): Promise<ApiResponse<{ host: CustomModelHost }>> => {
const denied = adminOnly(req, reply);
if (denied) return denied;
const { id } = req.params as { id: string };
const host = parseBody(CustomModelHostSchema, { ...(req.body as object), id });
if (isBlockedWebviewUrl(host.baseUrl)) {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
}
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
const index = hosts.findIndex((item) => item.id === id);
if (index === -1) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
const next = [...hosts];
next[index] = host;
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next);
return { success: true, data: { host } };
});
app.delete('/api/model-endpoints/:id', async (req, reply): Promise<ApiResponse<{ id: string }>> => {
const denied = adminOnly(req, reply);
if (denied) return denied;
const { id } = req.params as { id: string };
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
await writeCustomModelHosts(
CODEMAN_CONFIG_DIR,
hosts.filter((item) => item.id !== id)
);
return { success: true, data: { id } };
});
app.post(
'/api/model-endpoints/:id/discover-models',
async (req, reply): Promise<ApiResponse<{ models: string[] }>> => {
const denied = adminOnly(req, reply);
if (denied) return denied;
const { id } = req.params as { id: string };
const hosts = await readCustomModelHosts(CODEMAN_CONFIG_DIR);
const index = hosts.findIndex((item) => item.id === id);
if (index === -1) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
const host = hosts[index];
if (isBlockedWebviewUrl(host.baseUrl)) {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, 'Endpoint base URL is not allowed');
}
try {
const models = await discoverModels(host);
const next = [...hosts];
next[index] = { ...host, models, lastDiscoveredAt: new Date().toISOString() };
await writeCustomModelHosts(CODEMAN_CONFIG_DIR, next);
return { success: true, data: { models } };
} catch (err) {
const blocked = egressBlockedReason(err);
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
blocked ? `Endpoint refused: ${blocked}` : `Could not reach endpoint: ${describeFetchError(err)}`
);
}
}
);
}
File diff suppressed because it is too large Load Diff
+1
View File
@@ -27,3 +27,4 @@ export { registerWsRoutes } from './ws-routes.js';
export { registerVoiceRoutes } from './voice-routes.js';
export { registerWebviewRoutes, tryWebviewRefererFallback } from './webview-routes.js';
export { registerTabLayoutRoutes } from './tab-layout-routes.js';
export { registerCustomModelRoutes } from './custom-model-routes.js';
+113 -22
View File
@@ -51,7 +51,11 @@ import {
SessionOrderUpdateSchema,
SessionWaitQuerySchema,
SessionWaitOutputQuerySchema,
CustomModelSelectionSchema,
} from '../schemas.js';
import { readCustomModelHosts } from '../../custom-model-hosts.js';
import { applyCustomModelInjection, removeConfigDir } from '../../custom-model-injection-apply.js';
import { matchesPattern } from '../../config/cli-registry/patterns.js';
import { ownerLayoutKey } from '../../tab-layout-persistence.js';
import { TabLayoutValidationError } from '../../tab-layout.js';
import {
@@ -94,7 +98,6 @@ import {
writeHooksConfig,
updateCaseModel,
stripCaseEnvKeys,
applyStatusLineConfig,
applyAgentSkill,
refreshUserAgentSkill,
seedAgentSessionPreamble,
@@ -949,27 +952,19 @@ export function registerSessionRoutes(
await updateCaseModel(workingDir, body.modelOverride || null);
}
// Plan-usage statusLine exporter (App Settings → Display → "Plan Usage
// Limits"). Claude-only; runs for ANY working dir (linked cases / real repos,
// where most sessions live), mirroring updateCaseModel above.
//
// ADD-ONLY: we never remove on create. Sessions in a repo share one
// settings.local.json, so a single create-with-false (e.g. a client whose
// synced setting hadn't loaded yet) must NOT yank the statusLine out from
// under other live sessions in that repo — that breaks their footer + the
// chip's data feed for everyone. The exporter is benign when the chip is off
// (the footer just shows session status). isOurs-guarded so a user's own
// statusLine is never touched.
//
// Same guard as the hooks call below (499d355): never for a remote attach
// (workingDir is a user@host:session pseudo-path — the mkdir inside
// applyStatusLineConfig would create it as a junk local dir), and only when
// the caller named a workingDir — the process-cwd fallback is $HOME under
// installer-created services, and a statusLine materializing in
// ~/.claude/settings.local.json was never asked for.
if (!remote && body.workingDir && (body.mode ?? 'claude') === 'claude' && body.statusLineTelemetry === true) {
await applyStatusLineConfig(workingDir, true);
}
// Plan-usage telemetry (App Settings → header chip): no request-time field
// here anymore, and NO disk write — a settings.local.json statusLine used
// to take precedence over the user's own global/project statusLine for ANY
// `claude` run in that directory, including entirely outside Codeman, with
// no disclosure and no way to undo it (real bug, found 2026-08-31).
// TmuxManager.createSession reads the persisted `showPlanUsageLimits`
// setting FRESH at spawn (readPlanUsageTelemetryEnabled in hooks-config.ts)
// and resolves it into an EPHEMERAL `claude --settings` CLI flag — never
// written to disk, so a plain `claude` run outside Codeman is untouched —
// and applies uniformly to every claude creation path (this route, cron,
// the Ralph Loop API, quick-start), not just this one. That resolution
// also self-heals: it strips any legacy disk-written exporter an older
// Codeman build left behind.
// Hooks for the workspace this session runs in (install vs refresh-only is the
// `workspaceHooksEnabled` setting; see applyWorkspaceHooks). Never for a remote
@@ -1161,6 +1156,102 @@ export function registerSessionRoutes(
return { color: session.color };
});
// ========== Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) ==========
//
// Applies (or clears) a session's custom OpenAI-compatible endpoint selection and
// RESTARTS the pane's CLI process — these harnesses read endpoint config at process
// start, not per-turn, so a live hot-swap isn't possible (confirmed with the
// maintainer). Endpoints come from the admin-configured custom-model-hosts store
// (chunk 3's CRUD routes), never raw client-supplied env — that's what keeps this
// route safe to let any session owner call for their own session, unlike the
// generic envOverrides field the privilegedEnvKeys clamp exists to guard.
//
// ⚠️ Local sessions only for now. A remote session's `restartCli()` renders
// `ssh ... tmux new-session -A`, which reattaches the durable remote tmux rather than
// restarting the agent, and the env lands on the LOCAL pane running ssh, which
// forwards nothing; docker is the same attach-or-create shape. Both used to answer
// `restarted: true` and change nothing, so they are refused until those paths are
// plumbed (the env would have to ride the remote/in-container launch command).
app.post('/api/sessions/:id/custom-model', async (req) => {
const { id } = req.params as { id: string };
const body = parseBody(CustomModelSelectionSchema, req.body, 'Invalid request body');
const session = findSessionOrFail(ctx, id, req);
if (session.remote || session.docker) {
return createErrorResponse(
ApiErrorCode.INVALID_INPUT,
'Custom model endpoints are not supported for remote (SSH) or Docker sessions yet'
);
}
if (session.isBusy()) {
return createErrorResponse(ApiErrorCode.SESSION_BUSY, 'Session is busy');
}
if ('clear' in body) {
const { previousConfigDir } = session.setCustomModel(undefined);
removeConfigDir(previousConfigDir);
const restarted = await session.restartCli();
persistAndBroadcastSession(ctx, session);
return { customModel: session.customModel, restarted };
}
const entry = getCli(session.mode);
if (!entry) {
return createErrorResponse(ApiErrorCode.INVALID_INPUT, `No CLI registry entry for mode ${session.mode}`);
}
if (entry.capabilities.customModelInjection.kind === 'unsupported') {
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
}
const hosts = await readCustomModelHosts(getDataDir());
const endpoint = hosts.find((h) => h.id === body.endpointId);
if (!endpoint) {
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Model endpoint not found');
}
// A CLI whose config alone cannot select the model also gets its `model` launch param
// forced (pi/omp `custom/<id>`, grok's block name). The argv engine DROPS a token that
// fails its pattern rather than quoting it, which would silently launch the CLI on its
// own default provider again, so refuse an id the pattern cannot carry up front.
const modelSpec = entry.launch.params.model;
const applied = applyCustomModelInjection(entry, endpoint, body.modelId, session.id);
if (!applied) {
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, `${session.mode} has no known custom-model mechanism`);
}
if (
applied.launchModel !== undefined &&
modelSpec?.type === 'token' &&
!matchesPattern(modelSpec.pattern, applied.launchModel)
) {
removeConfigDir(applied.configDir);
return createErrorResponse(
ApiErrorCode.INVALID_INPUT,
`Model id ${JSON.stringify(body.modelId)} cannot be passed to ${session.mode} on its command line`
);
}
const { previousConfigDir } = session.setCustomModel(
{
endpointId: endpoint.id,
modelId: body.modelId,
label: endpoint.label,
envKeys: applied.envKeys,
configDir: applied.configDir,
launchModel: applied.launchModel,
},
applied.envOverrides
);
// Clean up the OLD config dir on disk, unless the new one happens to reuse the same
// path (same session, configDir kind again) — never delete the dir we just wrote.
if (previousConfigDir && previousConfigDir !== applied.configDir) {
removeConfigDir(previousConfigDir);
}
const restarted = await session.restartCli();
persistAndBroadcastSession(ctx, session);
return { customModel: session.customModel, restarted };
});
// ========== Delete Session ==========
app.delete('/api/sessions/:id', async (req) => {
+6 -4
View File
@@ -8,8 +8,9 @@
* (localhost-only; hook-secret-gated while a tunnel runs — see middleware/auth).
*
* Returns a compact plain-text status string for the exporter to print as the
* in-terminal footer (print-through), so injecting our statusLine doesn't leave
* the terminal footer blank.
* in-terminal footer (print-through) when it has no statusline of the user's
* own to wrap. An unknown session gets an EMPTY body: the old brand-word
* answer rendered as the statusline itself (discussion #405).
*/
import { FastifyInstance } from 'fastify';
@@ -36,10 +37,11 @@ export function registerStatusTelemetryRoutes(app: FastifyInstance, ctx: Session
reply.type('text/plain; charset=utf-8');
// Unknown session — minimal footer, no broadcast.
// Unknown session: nothing to broadcast and nothing to print. Never a brand
// word here, it would render as the statusline.
if (!ctx.sessions.has(sessionId)) {
lastSig.delete(sessionId);
return 'codeman';
return '';
}
const payload = data as RawStatuslinePayload | undefined;
+32 -22
View File
@@ -5,7 +5,6 @@
*/
import { FastifyInstance } from 'fastify';
import { getCli } from '../../config/cli-registry/registry.js';
import { join, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
import { existsSync, mkdirSync, readdirSync } from 'node:fs';
@@ -33,7 +32,6 @@ import {
import { subagentWatcher } from '../../subagent-watcher.js';
import { imageWatcher } from '../../image-watcher.js';
import { workflowRunWatcher } from '../../workflow-run-watcher.js';
import { applyStatusLineConfig } from '../../hooks-config.js';
import { getLifecycleLog } from '../../session-lifecycle-log.js';
import {
buildAwayDigest,
@@ -938,6 +936,17 @@ export function registerSystemRoutes(
// ========== Settings ==========
app.get('/api/settings', async () => {
// A plain read. This route must NEVER write settings.json: readJsonConfig()
// answers `{}` for ANY read failure (a parse error, EACCES, EMFILE, a read
// that lands inside PUT's non-atomic write), not only for a missing file,
// and every page load calls this route, so a "persist the default when the
// key is absent" reconcile here replaced a whole settings file with one key
// on the first unlucky read. The plan-usage default is resolved by the
// READERS instead: an absent `showPlanUsageLimits` means ON to
// readPlanUsageTelemetryEnabled() (hooks-config.ts), the same way an absent
// `workspaceHooksEnabled` means ON, and the client resolves its own display
// default through planUsageChipEnabled(). Pinned by
// test/routes/system-routes-settings-get-plan-usage-default.test.ts.
return readJsonConfig(SETTINGS_PATH, 'settings', {});
});
@@ -992,9 +1001,9 @@ export function registerSystemRoutes(
} catch {
/* ignore */
}
// statusLineTelemetry and acknowledgeUnauthTunnel are ACTION fields (not stored
// settings) — strip them before persisting so settings.json stays clean.
const { statusLineTelemetry, acknowledgeUnauthTunnel, ...settingsToStore } = settings;
// acknowledgeUnauthTunnel is an ACTION field (not a stored setting) — strip
// it before persisting so settings.json stays clean.
const { acknowledgeUnauthTunnel, ...settingsToStore } = settings;
const merged = { ...existing, ...settingsToStore };
await fs.writeFile(SETTINGS_PATH, JSON.stringify(merged, null, 2));
@@ -1007,7 +1016,7 @@ export function registerSystemRoutes(
// Service toggles resolve from `merged` (existing + incoming), NEVER from the
// raw request body. A PARTIAL PUT omits keys it does not intend to change, and
// reading the body directly turned every omission into "apply the default":
// a body of just `{statusLineTelemetry:true}` would START the subagent watcher
// a body of just `{showPlanUsageLimits:true}` would START the subagent watcher
// (`?? true`) and STOP the workflow + image watchers (`?? false`), silently
// undoing the user's persisted config. Reading `merged` makes any PUT reconcile
// services to the effective stored settings instead, which also self-heals
@@ -1033,22 +1042,23 @@ export function registerSystemRoutes(
}
});
// Plan-usage chip: its DISPLAY is per-device (client-side, see settings-ui.js).
// Telemetry COLLECTION is server-side and enable-sticky — when a client turns
// the chip ON it sends statusLineTelemetry:true and we (re)inject our exporter
// into every ACTIVE Claude session's working dir so the live % starts flowing
// immediately (no new session needed). We deliberately never auto-REMOVE here:
// the exporter is benign/print-through and a per-repo settings.local.json is
// shared by sibling sessions, so one device's "off" must not yank the exporter
// another device's chip depends on. Each dir handled once.
if (statusLineTelemetry === true) {
const dirs = new Set<string>();
for (const session of ctx.sessions.values()) {
if (getCli(session.mode)?.capabilities.statusLineTelemetry && session.workingDir)
dirs.add(session.workingDir);
}
await Promise.all([...dirs].map((dir) => applyStatusLineConfig(dir, true).catch(() => {})));
}
// Plan-usage chip: its DISPLAY is per-device (client-side, see settings-ui.js),
// but `showPlanUsageLimits` ALSO doubles as the telemetry COLLECTION switch,
// persisted here in settingsToStore like any other setting (no special-casing
// needed — see readPlanUsageTelemetryEnabled's doc comment in hooks-config.ts).
// Telemetry COLLECTION used to be a SEPARATE, action-only, sticky mechanism
// here: toggling the chip ON re-injected a statusLine.command into every
// ACTIVE Claude session's settings.local.json so live % started flowing
// without a new session. That disk write was the bug fixed 2026-08-31 (it
// took precedence over the user's own statusline for ANY `claude` run in
// that directory, including outside Codeman, with no way to undo it).
// Collection is now decided by TmuxManager.createSession/respawnPane reading
// `showPlanUsageLimits` FRESH from settings.json at spawn time — no
// per-session field, no per-request threading through cron/Ralph-loop/
// quick-start/interactive-create (they all reach the same read), and no
// (re)injection into an already-running session needed here: the NEXT
// respawn (a Ralph cycle, `/clear`, a PTY-exit restart) already picks up
// whatever this PUT just persisted.
// Handle tunnel toggle dynamically
if ('tunnelEnabled' in settings) {
+25 -2
View File
@@ -192,8 +192,31 @@ export function registerWsRoutes(app: FastifyInstance, ctx: SessionPort, getHost
// a duplicate: the input was lost for good.
if (!delivered && cid && seq !== null) session.forgetInputSeq(cid, seq);
}
if (delivered && seq !== null && socket.readyState === 1) {
socket.send(`{"t":"ia","seq":${seq}}`);
if (seq !== null && socket.readyState === 1) {
if (apply) {
if (delivered) socket.send(`{"t":"ia","seq":${seq}}`);
} else {
// REJECTED as a duplicate. ACK it — the client must still drop it
// from its durable queue — but say so, and hand back our watermark.
//
// A plain ACK here is indistinguishable from "applied", which is
// what made a client with a rolled-back counter unrecoverable: its
// seqs persist to localStorage on a DEBOUNCED write, so a tab killed
// between a send and that write comes back with a counter BELOW this
// watermark, every later keystroke lands at or under it, and each one
// is dropped-but-ACKed. The UI stays clean, nothing is delivered, and
// a reload restores the same stale counter. `last` is what lets the
// client lift itself out.
// ⚠️ Defensive: the session arrives through a structural port, and an
// implementation without this method must not take the whole input
// path down with it — a throw here aborts the message handler and the
// frame is never ACKed at all, which strands it in the client's queue.
const watermark =
typeof (session as { lastInputSeq?: (c: string) => number }).lastInputSeq === 'function'
? (session as { lastInputSeq: (c: string) => number }).lastInputSeq(cid as string)
: seq;
socket.send(`{"t":"ia","seq":${seq},"dup":true,"last":${watermark}}`);
}
}
} else if (
msg.t === 'z' &&
+39 -8
View File
@@ -524,8 +524,6 @@ export const CreateSessionSchema = z.object({
effort: effortLevelSchema,
/** Model override to write to .claude/settings.local.json (e.g., "opus[1m]"). Empty string clears. */
modelOverride: z.string().max(50).optional(),
/** Inject the Claude statusLine source for the shared plan-usage chip. Claude sessions only; Codex is host-polled. */
statusLineTelemetry: z.boolean().optional(),
openCodeConfig: OpenCodeConfigSchema,
codexConfig: CodexConfigSchema,
geminiConfig: GeminiConfigSchema,
@@ -1239,6 +1237,13 @@ export const SettingsUpdateSchema = z
* stored profiles stay until DELETE /api/sessions/:id/intent.
*/
readMyMindEnabled: z.boolean().optional(),
/**
* Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md): the toolbar picker that lets a
* session point at a user-configured custom OpenAI-compatible endpoint (local or
* cloud) instead of its native cloud backend. SYNCED, default OFF — endpoint entry,
* discovery, and the extra toolbar surface are all opt-in.
*/
customModelEndpointsEnabled: z.boolean().optional(),
/**
* Read My Mind predictor model override. Empty/absent = the AI-checker
* default (opus: prediction quality is the product and it runs only on an
@@ -1298,13 +1303,13 @@ export const SettingsUpdateSchema = z
showFileBrowser: z.boolean().optional(),
showSubagents: z.boolean().optional(),
showMultiMonitorButton: z.boolean().optional(),
// Doubles as the plan-usage telemetry COLLECTION switch, read fresh from
// disk by readPlanUsageTelemetryEnabled() (hooks-config.ts) at every claude
// session create/respawn — not just the chip's DISPLAY preference. See that
// function's doc comment for why one persisted field serves both. Absent
// means ON there, and the client sends it only on a save that flips the
// chip (planUsageCollectionFlip in settings-ui.js), never on every save.
showPlanUsageLimits: z.boolean().optional(),
// Action field (NOT persisted as a setting): when true, (re)injects the
// plan-usage statusLine exporter into active Claude sessions so live usage %
// starts flowing. Sent on ENABLE only — the chip's DISPLAY is per-device
// (client-side), but telemetry COLLECTION is server-side, so the per-device
// toggle signals it out-of-band here rather than via showPlanUsageLimits.
statusLineTelemetry: z.boolean().optional(),
showRedrawButton: z.boolean().optional(),
// Input
gestureControlEnabled: z.boolean().optional(),
@@ -1890,3 +1895,29 @@ export const WebviewUpdateSchema = WebviewBaseSchema.partial();
/** POST /api/webviews/probe: reachability + framing check for the editor's Test button. */
export const WebviewProbeSchema = z.object({ url: webviewUrlSchema });
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — a
// user-configured custom OpenAI-compatible endpoint, local (llama.cpp) or cloud
// (Azure AI Foundry, etc.). Lives below `webviewUrlSchema` because `baseUrl` IS that
// schema: http(s) only, a real hostname, no embedded credentials, and the link-local /
// cloud-metadata refusal, the same bar a saved dashboard URL has to clear.
export const CustomModelHostSchema = z.object({
id: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid endpoint id'),
label: z.string().min(1).max(100),
baseUrl: webviewUrlSchema,
apiKey: z.string().max(4096).optional(),
// No 'both': live-tested against a real server, sending both auth header
// conventions on one request reliably HANGS it — see custom-model-hosts.ts.
authStyle: z.enum(['bearer', 'api-key']).optional(),
models: z.array(z.string().max(200)).max(200).optional(),
lastDiscoveredAt: z.string().max(64).optional(),
});
/** POST /api/sessions/:id/custom-model — apply or clear a session's custom-model selection. */
export const CustomModelSelectionSchema = z.union([
z.object({
endpointId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid endpoint id'),
modelId: z.string().min(1).max(200),
}),
z.object({ clear: z.literal(true) }),
]);
+79 -3
View File
@@ -67,6 +67,10 @@ import {
import { imageWatcher } from '../image-watcher.js';
import { workflowRunWatcher, summarizeRun } from '../workflow-run-watcher.js';
import { attachmentRegistry, buildFileThumbnailRoute, registerExternalAttachment } from '../attachment-registry.js';
import { getCli } from '../config/cli-registry/registry.js';
import { readCustomModelHosts } from '../custom-model-hosts.js';
import { applyCustomModelInjection, customModelConfigDir, removeConfigDir } from '../custom-model-injection-apply.js';
import type { CustomModelBookkeeping } from '../types/session.js';
import { registerGeneratedArtifactAttachment } from '../generated-artifact-attachments.js';
import {
buildDetectedAttachmentHistoryItem,
@@ -148,7 +152,13 @@ import { getLatestPlanUsage, setLatestCodexPlanUsage } from './plan-usage-latest
import { telemetrySignature } from '../usage-telemetry.js';
import { readCodexPlanUsage, resolveCodexBinaryPath } from '../utils/codex-cli-resolver.js';
import type { ScheduledRun } from './ports/index.js';
import { registerAuthMiddleware, registerSecurityHeaders, registerHostGuard } from './middleware/auth.js';
import {
registerAuthMiddleware,
registerSecurityHeaders,
registerHostGuard,
isLostWebviewRootFrame,
sendLostWebviewFramePage,
} from './middleware/auth.js';
import { isMultiUserMode } from '../config/multiuser.js';
import { bootstrapInitialAdmin, hasUsers, resolveClaudeModeForUsername } from '../user-store.js';
import { installRouteErrorHandler } from './route-error-handler.js';
@@ -179,8 +189,10 @@ import {
registerVoiceRoutes,
registerWebviewRoutes,
registerTabLayoutRoutes,
registerCustomModelRoutes,
tryWebviewRefererFallback,
} from './routes/index.js';
import { isLostWebviewFrameNavigation } from './webview-proxy.js';
import { CronService } from '../cron/cron-service.js';
const __dirname = dirname(fileURLToPath(import.meta.url));
@@ -804,7 +816,14 @@ export class WebServer extends EventEmitter {
// Security headers + CORS
registerSecurityHeaders(this.app, this.https, this.basePath);
this.app.get('/', async (_req, reply) => {
this.app.get('/', async (req, reply) => {
// A web-tab frame that reloaded on its dashboard's landing page. The proxy's
// runtime shim maps `/webview/<cap>/` to exactly `/`, so that reload asks for
// Codeman's own root as an iframe navigation, and it used to get the app
// shell rendered inside the web tab. Only the credential-free form is taken
// (nothing in Codeman frames its root; the sandboxed frame has no cookie and
// no Authorization); under a password the auth hook has answered it already.
if (isLostWebviewRootFrame(req)) return sendLostWebviewFramePage(reply);
return reply
.header('Cache-Control', 'no-cache')
.type('text/html; charset=utf-8')
@@ -976,6 +995,11 @@ export class WebServer extends EventEmitter {
// and the relay declines unless the Referer carries a live capability, so
// genuinely unknown `/api` paths still get the envelope below.
if (await tryWebviewRefererFallback(req, reply, this.basePath)) return reply;
// An authenticated web-tab frame (Basic auth, or trusted mode with a cookie)
// that navigated itself off its proxy prefix: the runtime shim masks the
// prefix so the page's router sees its own path, and a reload of that page
// lands here. The unauthenticated form is answered in the auth middleware.
if (!req.url.startsWith('/api') && isLostWebviewFrameNavigation(req)) return sendLostWebviewFramePage(reply);
if (req.url.startsWith('/api')) {
return reply.code(404).send(createErrorResponse(ApiErrorCode.NOT_FOUND, notFound));
}
@@ -1051,6 +1075,7 @@ export class WebServer extends EventEmitter {
registerOrchestratorRoutes(this.app, ctx);
registerWebviewRoutes(this.app, ctx, this.basePath);
registerTabLayoutRoutes(this.app, ctx);
registerCustomModelRoutes(this.app);
// Cron: build the service from the same context, recompute
// due times for any persisted jobs, then expose it to its routes.
@@ -1160,6 +1185,34 @@ export class WebServer extends EventEmitter {
});
}
/**
* Recovery half of Custom Model Endpoint Profiles: the env values a selection injects
* are never persisted (they carry the API key), so they are computed again from the
* endpoint store, through the SAME apply path the route uses. Undefined when the
* endpoint is gone or the CLI is unregistered: the bookkeeping is still restored so
* the selection can be cleared, and the pane keeps running on tmux's retained env.
*/
private async _rebuildCustomModelEnv(
session: Session,
saved: CustomModelBookkeeping
): Promise<Record<string, string> | undefined> {
const entry = getCli(session.mode);
if (!entry) return undefined;
const endpoint = (await readCustomModelHosts(getDataDir())).find((h) => h.id === saved.endpointId);
if (!endpoint) {
console.warn(
`[WebServer] custom-model endpoint ${saved.endpointId} no longer exists; selection kept for clearing`
);
return undefined;
}
try {
return applyCustomModelInjection(entry, endpoint, saved.modelId, session.id)?.envOverrides;
} catch (err) {
console.warn('[WebServer] Failed to rebuild custom-model env on recovery:', err);
return undefined;
}
}
/** Persists full session state including respawn config to state.json */
private _persistSessionStateNow(session: Session): void {
// See session-manager.updateSessionState: __envOverrides is an internal disk-only
@@ -1169,10 +1222,14 @@ export class WebServer extends EventEmitter {
// __attachmentHistory keeps the private (externalPath-bearing) history on disk,
// separate from the sanitized public attachmentHistory in toState().
const attachmentHistory = session.getAttachmentHistoryForPersist();
// __customModel keeps the selection's bookkeeping (injected env KEYS, config dir,
// launch model; never the values) so recovery can restore and later clear it.
const customModel = session.getCustomModelForPersist();
const state = {
...base,
...(envOverrides ? { __envOverrides: envOverrides } : {}),
...(attachmentHistory ? { __attachmentHistory: attachmentHistory } : {}),
...(customModel ? { __customModel: customModel } : {}),
} as SessionState;
const controller = this.respawnControllers.get(session.id);
if (controller) {
@@ -1375,6 +1432,9 @@ export class WebServer extends EventEmitter {
// come back to a loader whose file we deleted.
if (killMux) {
void removeAgentSessionPreamble(sessionId);
// The per-session custom-model config dir carries the endpoint's API key (pi and
// omp embed it literally); it must not outlive the session it was written for.
removeConfigDir(customModelConfigDir(sessionId));
}
await session.stop(killMux);
this.sessions.delete(sessionId);
@@ -1450,7 +1510,10 @@ export class WebServer extends EventEmitter {
// PER-DEVICE by the client (settings-ui.js applyHeaderVisibilitySettings). It
// used to be server-revealed from a synced setting, but that leaked the desktop
// choice onto mobile — display is now per-device only (like the response viewer).
// Telemetry collection stays server-side via the statusLineTelemetry action.
// Telemetry collection stays server-side, reading `showPlanUsageLimits` fresh
// from settings.json at every claude session create/respawn (see
// readPlanUsageTelemetryEnabled in hooks-config.ts) — the same setting this
// display-visibility check reads, doing double duty.
// Detached single-session ("solo") window: inject the target session id so
// the client can enter solo mode even if a (network-first) service worker
// later serves a cached shell. The client primarily detects solo mode from
@@ -1697,10 +1760,12 @@ export class WebServer extends EventEmitter {
sessionId,
filePath,
sessionWorkingDir: session.workingDir,
remote: session.remote,
})
: await registerExternalAttachment(sessionId, filePath, {
sessionWorkingDir: session.workingDir,
forceWorkspaceConfinement: true,
remote: session.remote,
});
const record = attachmentRegistry.get(sessionId, event.attachmentId);
if (record) {
@@ -2827,6 +2892,7 @@ export class WebServer extends EventEmitter {
// Note: a legacy CLAUDE_CODE_EFFORT_LEVEL entry is auto-migrated to `effort`
// by the Session constructor (env var would hard-lock /effort switching).
const savedEnvOverrides = (savedState as { __envOverrides?: Record<string, string> })?.__envOverrides;
const savedCustomModel = (savedState as { __customModel?: CustomModelBookkeeping })?.__customModel;
// Prefer the private (externalPath-bearing) history; fall back to the
// sanitized public copy for sessions persisted before that split.
const savedAttachmentHistory =
@@ -2889,6 +2955,16 @@ export class WebServer extends EventEmitter {
parentSessionId: savedState?.parentSessionId,
});
// Custom-model selection survives the restart. The tmux session still carries
// the injected `setenv`s (that is what kept the pane on the endpoint across the
// restart), but `_envOverrides` is rebuilt from a persist that deliberately
// excludes them, so re-derive the values from the endpoint store and re-write
// the isolated config dir; an endpoint that has since been deleted still gets
// the bookkeeping restored, which is what a later clear needs to unset.
if (savedCustomModel) {
session.setCustomModel(savedCustomModel, await this._rebuildCustomModelEnv(session, savedCustomModel));
}
// Update session name if it was a "Restored:" placeholder or doesn't match saved name
if (savedState?.name && muxSession.name !== savedState.name) {
this.mux.updateSessionName(muxSession.sessionId, savedState.name);
+112 -1
View File
@@ -38,6 +38,7 @@
* that everything else here works to preserve.
*/
import { createHash } from 'node:crypto';
import { WEBVIEW_PROXY_PREFIX } from '../config/webview-limits.js';
import { stripBasePath } from '../config/base-path.js';
@@ -425,6 +426,21 @@ export function runtimeUrlShim(prefix: string): string {
// and a throw here would break the dashboard rather than fix it.
return `<script>(function(){try{
var P=${JSON.stringify(prefix)};
// Route masking. A single-page app reads location.pathname on boot and routes
// on it; through the proxy that path starts with /webview/<cap>/, which no app
// has a route for, so it rendered its own "page not found" the moment its
// script ran — after the HTML and CSS had already painted. Replace the entry
// with the path the page would see on its own origin. The base element still resolves
// relative URLs inside the prefix, and every root-absolute sink below is
// rewritten back into it, so only what the page READS changes. The parent
// tab remounts the frame if the page ever navigates itself off the prefix
// (see lostWebviewFramePage), which is what makes a masked reload survivable.
try{
var L=location.pathname;
if(L.indexOf(P)===0&&window.history&&typeof history.replaceState==='function'){
history.replaceState(history.state,'',L.slice(P.length-1)+location.search+location.hash);
}
}catch(e){}
function rw(u){
try{
if(u==null)return u;
@@ -457,13 +473,24 @@ if(window.XMLHttpRequest&&XMLHttpRequest.prototype.open){
var a=[].slice.call(arguments);a[1]=rw(u);return oo.apply(this,a);
};
}
['WebSocket','EventSource'].forEach(function(k){
['WebSocket','EventSource','Worker','SharedWorker'].forEach(function(k){
var C=window[k];if(!C)return;
function W(u,p){return p===undefined?new C(rw(u)):new C(rw(u),p);}
W.prototype=C.prototype;
['CONNECTING','OPEN','CLOSING','CLOSED'].forEach(function(s){if(s in C)W[s]=C[s];});
window[k]=W;
});
// With the document URL masked, a request the shim misses can no longer be
// rescued by its Referer (that carried the prefix), so the remaining
// URL-taking entry points are covered here rather than left to the fallback.
if(window.navigator&&typeof navigator.sendBeacon==='function'){
var ob=navigator.sendBeacon;
navigator.sendBeacon=function(u,d){return ob.call(navigator,rw(u),d);};
}
if(typeof window.open==='function'){
var ow=window.open;
window.open=function(u){var a=[].slice.call(arguments);a[0]=rw(u);return ow.apply(this,a);};
}
var A=['src','href','action','poster','data','formaction','srcset'];
function rwSet(v){
try{
@@ -680,3 +707,87 @@ export function upstreamWebSocketUrl(target: URL): string {
ws.protocol = ws.protocol === 'https:' ? 'wss:' : 'ws:';
return ws.href;
}
// ───────────────────────── Lost-frame recovery ─────────────────────────
/**
* The script the recovery page runs. Kept as a constant so its CSP hash below
* is computed from the exact bytes that are served.
*/
const LOST_FRAME_SCRIPT = `(function(){try{
var path=location.pathname+location.search+location.hash;
if(window.parent&&window.parent!==window){window.parent.postMessage({type:'codeman:webview-lost',path:path},'*');}
}catch(e){}})();`;
const LOST_FRAME_SCRIPT_HASH = createHash('sha256').update(LOST_FRAME_SCRIPT, 'utf8').digest('base64');
/** CSP for the recovery page: nothing but its own hashed inline script. */
export const LOST_FRAME_PAGE_CSP = `default-src 'none'; script-src 'sha256-${LOST_FRAME_SCRIPT_HASH}'; style-src 'unsafe-inline'`;
/**
* Whether this request is a web-tab frame that has navigated off its proxy prefix.
*
* The runtime shim masks `/webview/<cap>/` off the document URL so a single-page
* app routes on the path it expects. The price is that a navigation the page
* starts ITSELF — `location.reload()` (a dev server's full-reload HMR), a
* root-absolute `location.href = '/login'` — now targets Codeman's own root with
* no capability anywhere on it: no prefix in the path, no cookie in an
* opaque-origin frame, and a Referer that names the masked page. Such a request
* is recognisable by shape alone: a top-level navigation of an `<iframe>`
* (`Sec-Fetch-Dest`), asking for HTML, for a path Codeman does not serve. The
* one served path that still qualifies is `/` itself, which the callers admit
* only when the request carries no credentials (see carriesAuthCredentials).
*
* The answer is `lostWebviewFramePage()`, a static page whose only content is a
* `postMessage` to the parent naming the path; the Codeman tab that owns the
* frame remounts it inside the prefix at that path. Nothing is exempted from
* auth by this except that static page, which carries no data.
*/
export function isLostWebviewFrameNavigation(req: {
method: string;
headers: Record<string, string | string[] | undefined>;
}): boolean {
if (req.method !== 'GET' && req.method !== 'HEAD') return false;
const dest = req.headers['sec-fetch-dest'];
if (dest !== 'iframe' && dest !== 'frame') return false;
const mode = req.headers['sec-fetch-mode'];
if (mode !== undefined && mode !== 'navigate') return false;
const accept = req.headers.accept;
return typeof accept === 'string' && accept.includes('text/html');
}
/**
* Whether a request carries something Codeman's auth would recognise: the
* session cookie, or an `Authorization` header (Basic auth, which a browser
* re-sends on every request to the realm once it has been accepted).
*
* `/` is the one lost-frame path a registered route also serves (the app shell),
* so the route table cannot tell a landing-page reload of a proxied dashboard
* (the runtime shim maps `/webview/<cap>/` to exactly `/`) from a genuine
* navigation. Credentials can: nothing in Codeman frames its own root, and a
* sandboxed web-tab frame is opaque-origin and carries neither, so an `<iframe>`
* navigation of `/` with NEITHER credential can only be that frame. A framed
* `/` that does carry credentials is left to the shell.
*/
export function carriesAuthCredentials(
headers: Record<string, string | string[] | undefined>,
sessionCookieName: string
): boolean {
const authorization = headers.authorization;
if (Array.isArray(authorization) ? authorization.length > 0 : (authorization ?? '').trim() !== '') return true;
const cookie = headers.cookie;
const cookies = Array.isArray(cookie) ? cookie.join('; ') : cookie;
if (typeof cookies !== 'string' || cookies === '') return false;
return cookies.split(';').some((part) => part.trim().startsWith(`${sessionCookieName}=`));
}
/** The static page that hands a lost frame back to its owning tab. */
export function lostWebviewFramePage(): string {
return (
'<!doctype html><html><head><meta charset="utf-8"><title>Reconnecting</title>' +
'<meta name="referrer" content="no-referrer"></head>' +
'<body style="margin:0;font:14px system-ui,sans-serif;color:#888;padding:16px">' +
'Reconnecting this web tab…' +
`<script>${LOST_FRAME_SCRIPT}</script></body></html>`
);
}
+20
View File
@@ -35,6 +35,26 @@ function expectRejected(mutate: (entry: Record<string, unknown>) => void, becaus
expect(result.success, `expected rejection: ${because}`).toBe(false);
}
describe('customModelInjection.launchModel', () => {
it('rejects a template with characters the argv engine would have to quote', () => {
expectRejected((e) => {
const caps = e.capabilities as Record<string, unknown>;
caps.customModelInjection = { ...(caps.customModelInjection as object), launchModel: 'custom/{modelId} --yolo' };
}, 'a space in the launch-model template');
expectRejected((e) => {
const caps = e.capabilities as Record<string, unknown>;
caps.customModelInjection = { ...(caps.customModelInjection as object), launchModel: '' };
}, 'an empty launch-model template');
});
it('accepts the placeholder form the stock entries use', () => {
const entry = baseEntry();
const caps = entry.capabilities as Record<string, unknown>;
caps.customModelInjection = { ...(caps.customModelInjection as object), launchModel: 'custom/{modelId}' };
expect(CliEntrySchema.safeParse(entry).success).toBe(true);
});
});
describe('the shipped catalog', () => {
it('validates every stock entry exactly as shipped', () => {
// If this fails, the catalog cannot load at all — every other test here is downstream.
@@ -0,0 +1,215 @@
/**
* @fileoverview Contract tests for Custom Model Endpoint Profiles
* (docs/custom-model-endpoints-plan.md chunk 7): for every CLI with a `customModelInjection`
* capability, build the real injection via `buildCustomModelInjection()`,
* then replay those exact values through an HTTP request shaped the way that
* CLI is documented to send it, against the in-process mock server
* (`test/fixtures/mock-openai-server.ts`). Asserts the mock received the
* request at the injected base URL, with the injected API key in the
* expected header, and the injected model id in the body.
*
* LIMITATION (stated here and in docs/custom-model-endpoints-plan.md, not left implicit): this
* proves "if the CLI honors its documented env/config contract, it will hit
* the right endpoint with the right model." It does NOT prove the real CLI
* binary actually reads that env var / config file the way its docs say —
* that's still the job of `scripts/test-local-llm-harnesses.ts` against a
* real endpoint and real binaries. This suite catches regressions in
* Codeman's own injection logic; it cannot catch a CLI changing its env-var
* name in a future release.
*
* Port: N/A (mock server binds a random free port, not a fixed one)
*/
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { getCli } from '../src/config/cli-registry/index.js';
import { buildCustomModelInjection, type CustomModelEndpoint } from '../src/custom-model-injection.js';
import { startMockOpenAiServer, type MockOpenAiServer } from './fixtures/mock-openai-server.js';
let mock: MockOpenAiServer;
beforeEach(async () => {
mock = await startMockOpenAiServer();
});
afterEach(async () => {
await mock.close();
});
function entryOrThrow(id: string) {
const entry = getCli(id);
if (!entry) throw new Error(`missing CLI registry entry: ${id}`);
return entry;
}
function endpointFor(mock: MockOpenAiServer): CustomModelEndpoint {
return { id: 'ep1', label: 'mock', baseUrl: mock.baseUrl, apiKey: 'contract-test-key' };
}
/** Replays an OpenAI-shaped chat-completions call using the given base URL/key/model. */
async function callOpenAiCompat(baseUrl: string, apiKey: string, model: string) {
return fetch(`${baseUrl}/chat/completions`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` },
body: JSON.stringify({ model, messages: [{ role: 'user', content: 'hello world' }] }),
});
}
describe('custom-model-injection contract (mock server)', () => {
it('claude: ANTHROPIC_BASE_URL/API_KEY reach a real Anthropic-shaped /v1/messages call', async () => {
const injection = buildCustomModelInjection(entryOrThrow('claude'), endpointFor(mock), 'qwen3');
if (injection.kind !== 'env') throw new Error('unreachable');
await fetch(`${injection.envOverrides.ANTHROPIC_BASE_URL}/v1/messages`, {
method: 'POST',
headers: { 'content-type': 'application/json', 'x-api-key': injection.envOverrides.ANTHROPIC_API_KEY },
body: JSON.stringify({
model: injection.envOverrides.ANTHROPIC_DEFAULT_SONNET_MODEL,
messages: [{ role: 'user', content: 'hello world' }],
}),
});
expect(mock.requests).toHaveLength(1);
expect(mock.requests[0].path).toBe('/v1/messages');
expect(mock.requests[0].headers['x-api-key']).toBe('contract-test-key');
expect((mock.requests[0].body as { model: string }).model).toBe('qwen3');
});
it('opencode: OPENCODE_CONFIG_CONTENT decodes to a baseURL/apiKey that reach the mock', async () => {
const injection = buildCustomModelInjection(entryOrThrow('opencode'), endpointFor(mock), 'qwen3');
if (injection.kind !== 'env') throw new Error('unreachable');
const config = JSON.parse(injection.envOverrides.OPENCODE_CONFIG_CONTENT);
const { baseURL, apiKey } = config.provider.custom.options;
expect(baseURL).toBe(`${mock.baseUrl}/v1`);
await callOpenAiCompat(baseURL, apiKey, 'qwen3');
expect(mock.requests[0].path).toBe('/v1/chat/completions');
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
expect((mock.requests[0].body as { model: string }).model).toBe('qwen3');
});
it('codex: config.toml decodes to a base_url/model, and env_key/extraEnv reach the mock over /v1/responses', async () => {
const injection = buildCustomModelInjection(entryOrThrow('codex'), endpointFor(mock), 'qwen3');
if (injection.kind !== 'configDir') throw new Error('unreachable');
const toml = injection.files[0].content;
const baseUrl = /base_url = "([^"]+)"/.exec(toml)?.[1];
const model = /^model = "([^"]+)"/m.exec(toml)?.[1];
const envKeyName = /env_key = "([^"]+)"/.exec(toml)?.[1];
expect(baseUrl).toBe(`${mock.baseUrl}/v1`);
expect(model).toBe('qwen3');
expect(toml).toContain('wire_api = "responses"');
expect(toml).not.toContain('api_key ='); // never a literal TOML field
expect(envKeyName).toBe('CODEMAN_CUSTOM_MODEL_API_KEY');
expect(injection.extraEnv).toEqual({ CODEMAN_CUSTOM_MODEL_API_KEY: 'contract-test-key' });
// The real credential rides as an env var (env_key names it) — replay it, not a
// value read from the file, since the file itself never carries the secret.
const apiKey = injection.extraEnv!.CODEMAN_CUSTOM_MODEL_API_KEY;
await fetch(`${baseUrl}/responses`, {
method: 'POST',
headers: { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` },
body: JSON.stringify({ model, input: 'hello world' }),
});
expect(mock.requests[0].path).toBe('/v1/responses');
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
});
it('pi: models.json decodes to a baseUrl/apiKey that reach the mock', async () => {
const injection = buildCustomModelInjection(entryOrThrow('pi'), endpointFor(mock), 'qwen3');
if (injection.kind !== 'configDir') throw new Error('unreachable');
const parsed = JSON.parse(injection.files[0].content);
const { baseUrl, apiKey } = parsed.providers.custom;
expect(baseUrl).toBe(`${mock.baseUrl}/v1`);
await callOpenAiCompat(baseUrl, apiKey, 'qwen3');
expect(mock.requests[0].path).toBe('/v1/chat/completions');
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
});
it('omp: models.yml decodes to a baseUrl/apiKey that reach the mock', async () => {
const injection = buildCustomModelInjection(entryOrThrow('omp'), endpointFor(mock), 'qwen3');
if (injection.kind !== 'configDir') throw new Error('unreachable');
const yml = injection.files[0].content;
const baseUrl = JSON.parse(/baseUrl: (".*")\n/.exec(yml)![1]);
const apiKey = JSON.parse(/apiKey: (".*")\n/.exec(yml)![1]);
expect(baseUrl).toBe(`${mock.baseUrl}/v1`);
await callOpenAiCompat(baseUrl, apiKey, 'qwen3');
expect(mock.requests[0].path).toBe('/v1/chat/completions');
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
});
// gemini/deepseek's `env` kind passes the base URL through UNCHANGED (unlike
// opencode/codex/pi/omp/grok, which build a structured config and explicitly append
// /v1) — matching Anthropic's own convention for claude's ANTHROPIC_BASE_URL, where the
// SDK appends the path itself. Whether each of these TWO CLIs' own OpenAI-compatible
// client expects the var to already include /v1 (the common OpenAI-SDK convention) or
// appends it itself is genuinely CLI-specific and UNVERIFIED (see the confidence table
// in docs/custom-model-endpoints-plan.md) — these tests model the common OpenAI-SDK convention (base_url
// ends in /v1) since that's the more likely behavior for an OpenAI-compatible client,
// but that assumption should be corrected here the moment it's checked against a real
// binary. (grok WAS in this group too, until live-testing showed the whole `env` recipe
// was wrong for it — see its own test below.)
it('gemini: GOOGLE_GEMINI_BASE_URL/GEMINI_API_KEY reach the mock', async () => {
const injection = buildCustomModelInjection(entryOrThrow('gemini'), endpointFor(mock), 'qwen3');
if (injection.kind !== 'env') throw new Error('unreachable');
await callOpenAiCompat(
`${injection.envOverrides.GOOGLE_GEMINI_BASE_URL}/v1`,
injection.envOverrides.GEMINI_API_KEY,
injection.envOverrides.GEMINI_MODEL
);
expect(mock.requests[0].path).toBe('/v1/chat/completions');
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
expect((mock.requests[0].body as { model: string }).model).toBe('qwen3');
});
it('grok: config.toml [model.<name>] block base_url/env_key + extraEnv reach the mock over /v1/chat/completions', async () => {
const injection = buildCustomModelInjection(entryOrThrow('grok'), endpointFor(mock), 'qwen3');
if (injection.kind !== 'configDir') throw new Error('unreachable');
const toml = injection.files[0].content;
const baseUrl = /base_url = "([^"]+)"/.exec(toml)?.[1];
const model = /^model = "([^"]+)"/m.exec(toml)?.[1];
expect(baseUrl).toBe(`${mock.baseUrl}/v1`);
expect(model).toBe('qwen3');
expect(toml).toContain('api_backend = "chat_completions"');
expect(injection.extraEnv).toEqual({ XAI_API_KEY: 'contract-test-key' });
await callOpenAiCompat(baseUrl!, injection.extraEnv!.XAI_API_KEY, model!);
expect(mock.requests[0].path).toBe('/v1/chat/completions');
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
});
it('deepseek: DEEPSEEK_BASE_URL/DEEPSEEK_API_KEY reach the mock (base URL/key only, no model var)', async () => {
const injection = buildCustomModelInjection(entryOrThrow('deepseek'), endpointFor(mock), 'qwen3');
if (injection.kind !== 'env') throw new Error('unreachable');
expect(Object.keys(injection.envOverrides).sort()).toEqual(['DEEPSEEK_API_KEY', 'DEEPSEEK_BASE_URL']);
await callOpenAiCompat(
`${injection.envOverrides.DEEPSEEK_BASE_URL}/v1`,
injection.envOverrides.DEEPSEEK_API_KEY,
'qwen3'
);
expect(mock.requests[0].path).toBe('/v1/chat/completions');
expect(mock.requests[0].headers.authorization).toBe('Bearer contract-test-key');
});
it('antigravity: unsupported, never reaches the mock', () => {
const injection = buildCustomModelInjection(entryOrThrow('antigravity'), endpointFor(mock), 'qwen3');
expect(injection).toEqual({ kind: 'unsupported' });
expect(mock.requests).toHaveLength(0);
});
it('mock server also answers GET /v1/models for the discovery route', async () => {
const res = await fetch(`${mock.baseUrl}/v1/models`);
const body = await res.json();
expect(body.data.map((m: { id: string }) => m.id)).toEqual(['qwen3', 'llama3']);
});
});
+188
View File
@@ -0,0 +1,188 @@
/**
* @fileoverview Tests for the Custom Model Endpoint Profiles pure builder.
* Uses the real CLI registry entries (getCli) rather than hand-rolled
* fixtures, so a change to a real entry's customModelInjection declaration
* is exercised here automatically instead of silently diverging.
*
* Port: N/A (no server needed)
*/
import { describe, it, expect } from 'vitest';
import { getCli } from '../src/config/cli-registry/index.js';
import {
buildCustomModelInjection,
withV1Suffix,
GROK_CUSTOM_MODEL_NAME,
type CustomModelEndpoint,
} from '../src/custom-model-injection.js';
const endpoint: CustomModelEndpoint = {
id: 'ep1',
label: 'llama.cpp box',
baseUrl: 'http://192.168.1.50:8080',
apiKey: 'my-key',
};
function entryOrThrow(id: string) {
const entry = getCli(id);
if (!entry) throw new Error(`missing CLI registry entry: ${id}`);
return entry;
}
describe('withV1Suffix', () => {
it('appends /v1 when missing', () => {
expect(withV1Suffix('http://host:8080')).toBe('http://host:8080/v1');
});
it('is idempotent when already present', () => {
expect(withV1Suffix('http://host:8080/v1')).toBe('http://host:8080/v1');
expect(withV1Suffix('http://host:8080/v1/')).toBe('http://host:8080/v1');
});
it('strips a trailing slash with no /v1', () => {
expect(withV1Suffix('http://host:8080/')).toBe('http://host:8080/v1');
});
});
describe('buildCustomModelInjection', () => {
it('claude: env kind sets base URL, api key, and all three tier model vars', () => {
const result = buildCustomModelInjection(entryOrThrow('claude'), endpoint, 'qwen3');
expect(result.kind).toBe('env');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.envOverrides).toEqual({
ANTHROPIC_BASE_URL: 'http://192.168.1.50:8080',
ANTHROPIC_API_KEY: 'my-key',
ANTHROPIC_DEFAULT_SONNET_MODEL: 'qwen3',
ANTHROPIC_DEFAULT_HAIKU_MODEL: 'qwen3',
ANTHROPIC_DEFAULT_OPUS_MODEL: 'qwen3',
});
});
it('claude: falls back to a dummy key when the endpoint has none', () => {
const result = buildCustomModelInjection(entryOrThrow('claude'), { ...endpoint, apiKey: undefined }, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.envOverrides.ANTHROPIC_API_KEY).toBe('local-dummy-key');
});
it('opencode: configContentEnv carries a JSON blob in OPENCODE_CONFIG_CONTENT', () => {
const result = buildCustomModelInjection(entryOrThrow('opencode'), endpoint, 'qwen3');
expect(result.kind).toBe('env');
if (result.kind !== 'env') throw new Error('unreachable');
const parsed = JSON.parse(result.envOverrides.OPENCODE_CONFIG_CONTENT);
expect(parsed.model).toBe('custom/qwen3');
expect(parsed.provider.custom.options.baseURL).toBe('http://192.168.1.50:8080/v1');
expect(parsed.provider.custom.options.apiKey).toBe('my-key');
expect(parsed.provider.custom.models.qwen3).toEqual({});
});
it('codex: configDir writes an isolated config.toml with model/base_url, and the key rides as extraEnv (never a literal TOML field)', () => {
const result = buildCustomModelInjection(entryOrThrow('codex'), endpoint, 'qwen3');
expect(result.kind).toBe('configDir');
if (result.kind !== 'configDir') throw new Error('unreachable');
expect(result.dirEnvVar).toBe('CODEX_HOME');
expect(result.files).toHaveLength(1);
expect(result.files[0].relPath).toBe('config.toml');
expect(result.files[0].content).toContain('model = "qwen3"');
expect(result.files[0].content).toContain('base_url = "http://192.168.1.50:8080/v1"');
expect(result.files[0].content).toContain('wire_api = "responses"');
expect(result.files[0].content).not.toContain('api_key ='); // never a literal TOML field
expect(result.files[0].content).toContain('env_key = "CODEMAN_CUSTOM_MODEL_API_KEY"');
expect(result.extraEnv).toEqual({ CODEMAN_CUSTOM_MODEL_API_KEY: 'my-key' });
});
it('codex: escapes a quote in the model id so it cannot break out of the TOML string', () => {
const result = buildCustomModelInjection(entryOrThrow('codex'), endpoint, 'weird"model');
if (result.kind !== 'configDir') throw new Error('unreachable');
expect(result.files[0].content).toContain('model = "weird\\"model"');
});
it('pi: configDir writes .pi/agent/models.json, redirected via HOME (verified live — PI_CONFIG_DIR does nothing for pi)', () => {
const result = buildCustomModelInjection(entryOrThrow('pi'), endpoint, 'qwen3');
if (result.kind !== 'configDir') throw new Error('unreachable');
expect(result.dirEnvVar).toBe('HOME');
expect(result.files[0].relPath).toBe('.pi/agent/models.json');
const parsed = JSON.parse(result.files[0].content);
expect(parsed.providers.custom.baseUrl).toBe('http://192.168.1.50:8080/v1');
expect(parsed.providers.custom.authHeader).toBe(true);
expect(parsed.providers.custom.models).toEqual([{ id: 'qwen3' }]); // array, NOT keyed by id
});
it('omp: configDir writes .omp/agent/models.yml, redirected via HOME (verified live end-to-end)', () => {
const result = buildCustomModelInjection(entryOrThrow('omp'), endpoint, 'qwen3');
if (result.kind !== 'configDir') throw new Error('unreachable');
expect(result.dirEnvVar).toBe('HOME');
expect(result.files[0].relPath).toBe('.omp/agent/models.yml');
expect(result.files[0].content).toContain('baseUrl: "http://192.168.1.50:8080/v1"');
expect(result.files[0].content).toContain('authHeader: true');
expect(result.files[0].content).toContain('- id: "qwen3"');
});
it('gemini: env kind sets GOOGLE_GEMINI_BASE_URL/GEMINI_API_KEY/GEMINI_MODEL', () => {
const result = buildCustomModelInjection(entryOrThrow('gemini'), endpoint, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.envOverrides).toEqual({
GOOGLE_GEMINI_BASE_URL: 'http://192.168.1.50:8080',
GEMINI_API_KEY: 'my-key',
GEMINI_MODEL: 'qwen3',
});
});
it('grok: configDir writes a config.toml [model.<name>] block, key rides as extraEnv (XAI_API_KEY)', () => {
const result = buildCustomModelInjection(entryOrThrow('grok'), endpoint, 'qwen3');
expect(result.kind).toBe('configDir');
if (result.kind !== 'configDir') throw new Error('unreachable');
expect(result.dirEnvVar).toBe('GROK_HOME');
expect(result.files).toHaveLength(1);
expect(result.files[0].relPath).toBe('config.toml');
expect(result.files[0].content).toContain('model = "qwen3"');
expect(result.files[0].content).toContain('base_url = "http://192.168.1.50:8080/v1"');
expect(result.files[0].content).toContain('api_backend = "chat_completions"');
expect(result.files[0].content).toContain('env_key = "XAI_API_KEY"');
expect(result.files[0].content).not.toContain('api_key ='); // never a literal TOML field
expect(result.extraEnv).toEqual({ XAI_API_KEY: 'my-key' });
});
it('deepseek: env kind sets base URL/key only, no model var', () => {
const result = buildCustomModelInjection(entryOrThrow('deepseek'), endpoint, 'qwen3');
if (result.kind !== 'env') throw new Error('unreachable');
expect(result.envOverrides).toEqual({
DEEPSEEK_BASE_URL: 'http://192.168.1.50:8080',
DEEPSEEK_API_KEY: 'my-key',
});
});
it('antigravity: unsupported', () => {
const result = buildCustomModelInjection(entryOrThrow('antigravity'), endpoint, 'qwen3');
expect(result).toEqual({ kind: 'unsupported' });
});
it('shell: unsupported', () => {
const result = buildCustomModelInjection(entryOrThrow('shell'), endpoint, 'qwen3');
expect(result).toEqual({ kind: 'unsupported' });
});
});
describe('launchModel (the model launch param that selects the injected provider)', () => {
it('pi and omp get --model custom/<modelId>: the config file alone leaves them on their default provider', () => {
for (const id of ['pi', 'omp']) {
const result = buildCustomModelInjection(entryOrThrow(id), endpoint, 'qwen3.5-0.8b');
if (result.kind !== 'configDir') throw new Error('unreachable');
expect(result.launchModel, id).toBe('custom/qwen3.5-0.8b');
}
});
it('grok gets the [model.<name>] block name, pinned to the constant the config template writes', () => {
const result = buildCustomModelInjection(entryOrThrow('grok'), endpoint, 'qwen3');
if (result.kind !== 'configDir') throw new Error('unreachable');
expect(result.launchModel).toBe(GROK_CUSTOM_MODEL_NAME);
expect(result.files[0].content).toContain(`[model.${GROK_CUSTOM_MODEL_NAME}]`);
});
it('CLIs whose config selects the model on its own declare none', () => {
for (const id of ['claude', 'opencode', 'codex', 'gemini', 'deepseek']) {
const result = buildCustomModelInjection(entryOrThrow(id), endpoint, 'qwen3');
if (result.kind === 'unsupported') throw new Error('unreachable');
expect(result.launchModel, id).toBeUndefined();
}
});
});
+4 -1
View File
@@ -372,7 +372,10 @@ describe('DeepSeek multi-user clamp: the env-var half', () => {
});
it('leaves unrelated overrides alone, and returns the same object when there is nothing to strip', async () => {
const input = { DEEPSEEK_API_KEY: 'sk-test', CODEX_HOME: '/tmp/cx' };
// CODEX_HOME is a poor "unrelated" example here — it is itself a privileged key
// (codex's own registry entry), so a genuinely non-privileged one is needed to
// prove the identity-return fast path, not just that DEEPSEEK_API_KEY is exempt.
const input = { DEEPSEEK_API_KEY: 'sk-test', OPENCODE_LOG_LEVEL: 'debug' };
const out = await _clampEnvOverridesForOwner('nobody', input);
expect(out).toBe(input);
expect(await _clampEnvOverridesForOwner('nobody', undefined)).toBeUndefined();
+138
View File
@@ -0,0 +1,138 @@
/**
* @fileoverview "Duplicate an existing case" in the container-adoption form.
*
* The server already allows one ADOPTED container to back several cases, each
* pointing at a different directory inside it (classifyAdoptContainerConflict).
* Re-typing the container, host and workspace by hand for every directory is the
* friction that would leave that capability unused, so the form carries them over
* and clears only the two fields that MUST differ.
*/
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import { describe, it, expect } from 'vitest';
const html = readFileSync(resolve(import.meta.dirname, '../src/web/public/index.html'), 'utf8');
const ui = readFileSync(resolve(import.meta.dirname, '../src/web/public/session-ui.js'), 'utf8');
const routes = readFileSync(resolve(import.meta.dirname, '../src/web/routes/case-routes.ts'), 'utf8');
const apiTypes = readFileSync(resolve(import.meta.dirname, '../src/types/api.ts'), 'utf8');
describe('the API exposes what the picker needs', () => {
it('reports each docker case s directory inside the container', () => {
// Without it the picker cannot show WHICH directory a case already uses, which
// is the one thing the user needs to see before choosing a different one.
expect(apiTypes).toMatch(/containerWorkdir\?: string;/);
expect(routes).toContain('containerWorkdir: dockerCase.containerWorkdir ?? dockerCase.hostWorkspacePath');
});
it('reports whether the container is owned, on EVERY case-shaped response', () => {
// Two sites build a docker CaseInfo (the list and the single-case lookup);
// filling only one leaves the picker blind depending on which the UI read.
expect(routes.match(/\.\.\.\(dockerCase\.owned === false \? \{ owned: false \} : \{\}\),/g) ?? []).toHaveLength(2);
});
it('treats an ABSENT owned flag as owned, so legacy cases are not offered', () => {
// `owned` is optional and predates this field; the wire carries it ONLY when
// false (absent = owned, the shape master already used), so the picker must
// test `=== false` rather than truthiness, or a legacy case would read as
// adopted and be offered a duplicate the server then refuses.
expect(routes).toContain('...(dockerCase.owned === false ? { owned: false } : {}),');
expect(routes).not.toContain('owned: dockerCase.owned !== false');
});
});
describe('the picker only offers what the server would accept', () => {
const fn = ui.slice(ui.indexOf('async _loadDockerCloneOptions()'), ui.indexOf('applyDockerCloneSource()'));
it('filters to ADOPTED cases only', () => {
expect(fn).toMatch(/c\.docker\.owned === false/);
});
it('hides the row entirely when there is nothing to duplicate', () => {
expect(fn).toMatch(/row\.hidden = cases\.length === 0/);
});
it('builds options with textContent, never markup', () => {
// Case names and container names are user- and engine-supplied strings.
expect(fn).toContain('option.textContent =');
expect(fn).not.toContain('innerHTML');
});
});
describe('applying a source fills every field, including the two that must differ', () => {
const fn = ui.slice(ui.indexOf('applyDockerCloneSource()'), ui.indexOf('dockerCloneGuard()'));
it('carries over container, host and workspace', () => {
for (const id of ['dockerContainerName', 'dockerHostId', 'dockerWorkspacePath']) {
expect(fn).toContain(`set('${id}', option.dataset.`);
}
});
it('PRE-FILLS the case name and container workdir rather than clearing them', () => {
// Editing `/srv/app/api` into `/srv/app/web` beats retyping a long path, and a
// form with three fields mysteriously filled and two blank reads as broken.
// What stops an unchanged submit is the guard, not an empty field.
expect(fn).toContain("set('dockerCaseName', option.value)");
expect(fn).toContain("set('dockerAdoptWorkdir', option.dataset.workdir)");
});
it('remembers what it applied, so the guard can tell unchanged from similar', () => {
expect(fn).toContain('select.dataset.appliedName = option.value');
expect(fn).toContain('select.dataset.appliedWorkdir =');
});
it('focuses the workdir with the caret at the END, where the edit happens', () => {
expect(fn).toMatch(/setSelectionRange\(workdir\.value\.length, workdir\.value\.length\)/);
});
it('does nothing for the blank "start from scratch" option', () => {
expect(fn).toMatch(/if \(!option \|\| !option\.value\) return;/);
});
});
describe('the guard refuses a duplicate that was never edited', () => {
const fn = ui.slice(ui.indexOf('dockerCloneGuard()'), ui.indexOf('dockerCloneGuard()') + 1400);
it('flags an unchanged case name', () => {
expect(fn).toMatch(/appliedName/);
expect(fn).toContain('give this one a new name');
});
it('flags an unchanged container workdir', () => {
expect(fn).toMatch(/appliedWorkdir/);
expect(fn).toContain('another directory');
});
it('stays silent when no source was picked', () => {
// Typing a fresh adoption by hand must not be second-guessed.
expect(fn).toMatch(/if \(!select \|\| !select\.value\) return null;/);
});
it('runs BEFORE the request, and focuses the offending field', () => {
const submit = ui.slice(ui.indexOf('const cloneIssue'), ui.indexOf('const cloneIssue') + 500);
expect(submit).toContain('cloneIssue.el.focus()');
expect(submit).toContain('return;');
});
it('reports into a status element that actually exists', () => {
// A dead id would silently drop the explanation next to the field.
const submit = ui.slice(ui.indexOf('const cloneIssue'), ui.indexOf('const cloneIssue') + 500);
const id = /getElementById\('([^']+)'\)/.exec(submit)?.[1];
expect(id).toBeTruthy();
expect(html).toContain(`id="${id}"`);
});
});
describe('the row is wired into the adoption panel', () => {
it('lives in the adopt-only block and starts hidden', () => {
expect(html).toMatch(/id="dockerAdoptCloneRow"[^>]*hidden/);
expect(html).toMatch(/class="form-row docker-adopt-only" id="dockerAdoptCloneRow"/);
});
it('loads its options whenever adopt mode turns on', () => {
// Slice from the DEFINITION, not the first call site.
const start = ui.indexOf('_syncDockerAdoptMode() {');
expect(start).toBeGreaterThan(-1);
const sync = ui.slice(start, start + 900);
expect(sync).toContain('_loadDockerCloneOptions()');
});
});
+112
View File
@@ -19,6 +19,8 @@ import {
checkDockerConfigDrift,
dockerConfigHash,
dockerAdoptProbeModes,
classifyAdoptContainerConflict,
dockerContainerName,
} from '../src/docker-hosts.js';
import { enabledCliIds, getCli } from '../src/config/cli-registry/index.js';
import {
@@ -451,3 +453,113 @@ describe('adopted container: a missing container means different things per owne
expect(routes).toContain('...(dockerCase.owned === false ? { owned: false } : {}),');
});
});
describe('adopted container: one container may back several cases', () => {
const base = {
type: 'docker' as const,
hostId: 'h1',
hostWorkspacePath: '/srv/work',
};
const mk = (over: Record<string, unknown>) => ({ ...base, ...over }) as never;
const mine = () => true;
it('allows a second adoption of the same container at a DIFFERENT directory', () => {
// The whole point of the feature: one container, two folders, two cases.
const conflict = classifyAdoptContainerConflict({
container: 'devbox',
containerWorkdir: '/app/api',
existing: [mk({ name: 'web', container: 'devbox', containerWorkdir: '/app/web', owned: false })],
canAccess: mine,
});
expect(conflict).toBeNull();
});
it('refuses an exact twin (same container AND same directory) and names the first case', () => {
const conflict = classifyAdoptContainerConflict({
container: 'devbox',
containerWorkdir: '/app/web',
existing: [mk({ name: 'web', container: 'devbox', containerWorkdir: '/app/web', owned: false })],
canAccess: mine,
});
expect(conflict).toEqual({ kind: 'duplicate', caseName: 'web' });
});
it('falls back to hostWorkspacePath when containerWorkdir is absent on either side', () => {
// containerWorkdir defaults to hostWorkspacePath, so an absent field on the
// stored case must compare equal to an incoming adoption that omits it too —
// otherwise the twin check silently stops firing for the default case.
const conflict = classifyAdoptContainerConflict({
container: 'devbox',
containerWorkdir: '/srv/work',
existing: [mk({ name: 'web', container: 'devbox', owned: false })],
canAccess: mine,
});
expect(conflict).toEqual({ kind: 'duplicate', caseName: 'web' });
});
it('still refuses a container backing a case Codeman CREATED', () => {
// Codeman owns that container's lifecycle: a recreate or case-delete there
// would destroy the adopted case's container out from under it.
const conflict = classifyAdoptContainerConflict({
container: 'codeman-case-web',
containerWorkdir: '/app/api',
existing: [mk({ name: 'web', container: 'codeman-case-web', owned: true })],
canAccess: mine,
});
expect(conflict).toEqual({ kind: 'owned-case', caseName: 'web' });
});
it('treats an ABSENT owned flag as owned, so legacy cases keep the old refusal', () => {
const conflict = classifyAdoptContainerConflict({
container: 'legacy',
containerWorkdir: '/app/api',
existing: [mk({ name: 'old', container: 'legacy' })],
canAccess: mine,
});
expect(conflict).toEqual({ kind: 'owned-case', caseName: 'old' });
});
it('derives the container name from the case name when the field is absent', () => {
const conflict = classifyAdoptContainerConflict({
container: dockerContainerName('web'),
containerWorkdir: '/app/api',
existing: [mk({ name: 'web', owned: true })],
canAccess: mine,
});
expect(conflict).toEqual({ kind: 'owned-case', caseName: 'web' });
});
it('refuses a container another user already adopted', () => {
const conflict = classifyAdoptContainerConflict({
container: 'devbox',
containerWorkdir: '/app/api',
existing: [mk({ name: 'theirs', container: 'devbox', owned: false, owner: 'bob' })],
canAccess: (owner) => owner === 'alice',
});
expect(conflict).toEqual({ kind: 'other-owner', caseName: 'theirs' });
});
it('an owned case outranks a foreign adoption, so the message names the real blocker', () => {
const conflict = classifyAdoptContainerConflict({
container: 'devbox',
containerWorkdir: '/app/api',
existing: [
mk({ name: 'theirs', container: 'devbox', owned: false, owner: 'bob' }),
mk({ name: 'built', container: 'devbox', owned: true }),
],
canAccess: (owner) => owner === 'alice',
});
expect(conflict).toEqual({ kind: 'owned-case', caseName: 'built' });
});
it('leaves an unrelated container alone', () => {
expect(
classifyAdoptContainerConflict({
container: 'fresh',
containerWorkdir: '/app',
existing: [mk({ name: 'web', container: 'devbox', owned: false })],
canAccess: mine,
})
).toBeNull();
});
});
+256
View File
@@ -0,0 +1,256 @@
/**
* @fileoverview Static and fixture checks for the Docker Compose deployment's
* privilege handling: `docker/entrypoint.sh` starts as root, corrects bind-mount
* ownership and drops to PUID:PGID, which only works while three files agree.
*
* 1. The capabilities `docker-compose.yaml` adds back on top of `cap_drop: ALL`
* must be exactly what the entrypoint and `init: true` need. This is the
* drift that shipped once already: the `USER` instruction became a root
* entrypoint, tini stayed root while the server became PUID, and with no
* CAP_KILL every `docker compose down` ended in tini failing to forward
* SIGTERM and the server being SIGKILLed. The list is derived here from what
* the scripts actually do, not copied.
* 2. The runtime-owned CLI prefix must never sit ahead of the system
* directories on the PATH the root entrypoint resolves commands through: a
* planted `setpriv` in a PUID-writable prefix ran as uid 0 (measured with a
* minimal image of the same shape).
* 3. `Start-Codeman.sh` derives PUID/PGID BEFORE it creates
* `CODEMAN_CASES_PATH`, so the directory it creates has the owner the
* container will accept, and its `git_head_commit` helper (a pure function
* over `.git`) resolves the three ref layouts a checkout can have.
*/
import { describe, it, expect, beforeAll, afterAll } from 'vitest';
import { readFileSync, mkdtempSync, rmSync, writeFileSync } from 'node:fs';
import { execFileSync } from 'node:child_process';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
const ROOT = process.cwd();
const read = (rel: string) => readFileSync(join(ROOT, rel), 'utf-8');
const compose = read('docker/docker-compose.yaml');
const entrypoint = read('docker/entrypoint.sh');
const dockerfile = read('docker/server.Dockerfile');
const startScript = read('docker/Start-Codeman.sh');
/** The `- NAME` entries under `cap_add:` (the block ends at the next key at the same indent). */
function composeCapAdd(text: string): string[] {
const m = text.match(/^(\s*)cap_add:\n((?:\1\s+.*\n)*)/m);
if (!m) return [];
return m[2]
.split('\n')
.map((l) => l.trim())
.filter((l) => l.startsWith('- '))
.map((l) => l.slice(2).trim())
.sort();
}
/**
* What the deployment needs, derived from the scripts. Each rule names the
* line that needs it, so a capability cannot be added or removed here without
* the reason changing too.
*/
function requiredCaps(): string[] {
const caps = new Set<string>();
if (/\bchown\b/.test(entrypoint)) {
// chown of a root-owned bind source, and traversing trees root cannot
// otherwise read on a mount with restrictive modes.
caps.add('CHOWN');
caps.add('DAC_OVERRIDE');
}
if (/setpriv .*--reuid/.test(entrypoint)) caps.add('SETUID');
if (/setpriv .*--(regid|groups|clear-groups)/.test(entrypoint)) caps.add('SETGID');
const dropsUid = /setpriv .*--reuid/.test(entrypoint);
if (/^\s*init:\s*true\s*$/m.test(compose) && dropsUid) {
// tini is PID 1 and stays root; signalling the PUID server needs CAP_KILL.
caps.add('KILL');
}
return [...caps].sort();
}
describe('docker-compose.yaml cap_add covers what entrypoint.sh and init:true need', () => {
it('the compose file adds back exactly the derived capability set', () => {
expect(composeCapAdd(compose)).toEqual(requiredCaps());
});
it('cap_drop: ALL is still the baseline', () => {
expect(compose).toMatch(/^\s*cap_drop:\n\s*- ALL\s*$/m);
});
it("the entrypoint's own diagnosis names the same list, so a missing cap gets a one-line fix", () => {
const m = entrypoint.match(/^required_caps='([^']+)'/m);
expect(m, 'entrypoint.sh must declare required_caps').not.toBeNull();
const named = m![1]
.split(',')
.map((c) => c.trim())
.sort();
expect(named).toEqual(composeCapAdd(compose));
});
it('the user-facing docs quote the same cap_add list', () => {
for (const rel of ['docker/README.md', 'CLAUDE.md']) {
const text = read(rel);
const quoted = [...text.matchAll(/cap_add: \[([^\]]+)\]/g)].map((m) =>
m[1]
.split(',')
.map((c) => c.trim())
.sort()
);
expect(quoted.length, `${rel} should quote the cap_add list at least once`).toBeGreaterThan(0);
for (const list of quoted) expect(list, rel).toEqual(composeCapAdd(compose));
}
});
});
describe('the runtime-owned CLI prefix never shadows root commands', () => {
it('server.Dockerfile appends /opt/codeman-cli/bin to PATH rather than prepending it', () => {
const pathLines = dockerfile.split('\n').filter((l) => /^ENV PATH=/.test(l));
expect(pathLines.length).toBeGreaterThan(0);
for (const line of pathLines) {
expect(line, 'a writable prefix ahead of $PATH lets a planted setpriv run as root').not.toMatch(
/^ENV PATH=\/opt\/codeman-cli/
);
}
expect(pathLines).toContain('ENV PATH=$PATH:/opt/codeman-cli/bin');
});
it('entrypoint.sh pins PATH to the system directories before its first command', () => {
const lines = entrypoint.split('\n');
const pinIdx = lines.findIndex((l) => l === 'PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin');
expect(pinIdx, 'the PATH pin must exist').toBeGreaterThan(-1);
const firstToolIdx = lines.findIndex((l) => !l.trim().startsWith('#') && /\b(setpriv|chown|stat)\b/.test(l));
expect(firstToolIdx).toBeGreaterThan(pinIdx);
// The only thing allowed before the pin is the `user:` short-circuit.
const before = lines
.slice(0, pinIdx)
.filter((l) => l.trim() && !l.trim().startsWith('#') && !/^(set -eu|runtime_path=\$PATH)$/.test(l.trim()));
expect(before).toEqual(['if [ "$(id -u)" -ne 0 ]; then', ' exec "$@"', 'fi']);
});
it("entrypoint.sh hands the image's full PATH back to the server at the drop", () => {
expect(entrypoint).toMatch(/exec setpriv [^\n]*\\\n\s*env PATH="\$runtime_path" "\$@"/);
});
it('entrypoint.sh no longer passes --bounding-set (a silent no-op without CAP_SETPCAP)', () => {
const code = entrypoint
.split('\n')
.filter((l) => !l.trim().startsWith('#'))
.join('\n');
expect(code).not.toMatch(/--bounding-set/);
expect(composeCapAdd(compose)).not.toContain('SETPCAP');
});
});
describe('Start-Codeman.sh', () => {
it('parses under bash -n', () => {
execFileSync('bash', ['-n', join(ROOT, 'docker/Start-Codeman.sh')]);
execFileSync('sh', ['-n', join(ROOT, 'docker/entrypoint.sh')]);
});
it('derives PUID/PGID before creating CODEMAN_CASES_PATH, so the new directory gets that owner', () => {
const puid = startScript.indexOf('export PUID=');
const mkdirCases = startScript.indexOf('mkdir -p -- "$cases_path"');
expect(puid).toBeGreaterThan(-1);
expect(mkdirCases).toBeGreaterThan(puid);
expect(startScript).toMatch(/chown -- "\$PUID:\$PGID" "\$cases_path"/);
});
it('builds before taking the stack down, and writes the source marker only after a refresh', () => {
const build = startScript.indexOf('"${compose_command[@]}" build');
const down = startScript.indexOf('"${compose_command[@]}" down');
const marker = startScript.indexOf('>"$source_state_file.tmp"');
expect(build).toBeGreaterThan(-1);
expect(down).toBeGreaterThan(build);
expect(marker).toBeGreaterThan(down);
expect(startScript).toMatch(/if \[\[ "\$refreshed" == '1' \]\]; then\n\s*printf '\{\\n {2}"headCommit"/);
// A failed volume removal must not abort under set -e with the stack down.
expect(startScript).not.toMatch(/\[\[ -n "\$volume_name" \]\] && docker volume rm/);
expect(startScript).toMatch(/&& ! docker volume rm -- "\$volume_name"; then/);
});
it('falls back to `down --volumes` when the Compose project name cannot be resolved', () => {
expect(startScript).toMatch(/if \[\[ -z "\$project_name" \]\]; then[\s\S]*down --volumes/);
});
});
describe('git_head_commit resolves every ref layout a checkout can have', () => {
let base: string;
const git = (cwd: string, ...args: string[]) =>
execFileSync('git', args, {
cwd,
encoding: 'utf-8',
env: {
...process.env,
GIT_AUTHOR_NAME: 't',
GIT_AUTHOR_EMAIL: 't@example.com',
GIT_COMMITTER_NAME: 't',
GIT_COMMITTER_EMAIL: 't@example.com',
},
}).trim();
/** Runs the function exactly as the script defines it, extracted by its own delimiters. */
const headCommit = (repo: string): { out: string; status: number } => {
const script = [`eval "$(sed -n '/^git_head_commit() {/,/^}/p' "$1")"`, 'git_head_commit "$2"'].join('\n');
try {
const out = execFileSync('bash', ['-c', script, '_', join(ROOT, 'docker/Start-Codeman.sh'), repo], {
encoding: 'utf-8',
});
return { out: out.trim(), status: 0 };
} catch (err) {
const e = err as { stdout?: string; status?: number };
return { out: (e.stdout ?? '').trim(), status: e.status ?? 1 };
}
};
const makeRepo = (name: string): string => {
const dir = join(base, name);
git(base, 'init', '-q', '-b', 'master', dir);
writeFileSync(join(dir, 'f'), 'x');
git(dir, 'add', 'f');
git(dir, 'commit', '-q', '-m', 'one');
return dir;
};
beforeAll(() => {
base = mkdtempSync(join(tmpdir(), 'codeman-head-commit-'));
});
afterAll(() => {
rmSync(base, { recursive: true, force: true });
});
it('symbolic ref with a loose ref file', () => {
const dir = makeRepo('loose');
expect(headCommit(dir)).toEqual({ out: git(dir, 'rev-parse', 'HEAD'), status: 0 });
});
it('detached HEAD', () => {
const dir = makeRepo('detached');
const sha = git(dir, 'rev-parse', 'HEAD');
git(dir, 'checkout', '-q', '--detach', sha);
expect(headCommit(dir)).toEqual({ out: sha, status: 0 });
});
it('packed refs after gc', () => {
const dir = makeRepo('packed');
const sha = git(dir, 'rev-parse', 'HEAD');
git(dir, 'pack-refs', '--all');
expect(readFileSync(join(dir, '.git/packed-refs'), 'utf-8')).toContain('refs/heads/master');
expect(headCommit(dir)).toEqual({ out: sha, status: 0 });
});
it('a linked worktree (.git is a file) resolves nothing rather than something wrong', () => {
const dir = makeRepo('main');
const wt = join(base, 'wt');
git(dir, 'worktree', 'add', '-q', wt);
const result = headCommit(wt);
expect(result.out).toBe('');
expect(result.status).not.toBe(0);
});
it('a directory that is not a checkout fails', () => {
const result = headCommit(base);
expect(result.out).toBe('');
expect(result.status).not.toBe(0);
});
});
+118
View File
@@ -0,0 +1,118 @@
/**
* @fileoverview In-process fake OpenAI/Anthropic-compatible HTTP server for the
* Custom Model Endpoint Profiles contract tests (docs/custom-model-endpoints-plan.md chunk 7).
*
* No external deps — plain `node:http`. Captures every request it receives
* (method, path, headers, parsed JSON body) so a test can assert the injected
* base URL / API key / model actually reached the right place, with the right
* auth header, in the shape a real llama.cpp/Azure/etc. endpoint would see it.
*
* Serves the request shapes this feature's recipes produce: OpenAI-style
* `POST /v1/chat/completions` (opencode/pi/grok/omp/gemini's compat
* endpoint), Anthropic-style `POST /v1/messages` (claude's ANTHROPIC_BASE_URL
* traffic), OpenAI's newer `POST /v1/responses` (codex's actual wire protocol
* as of Feb 2026 — it dropped chat-completions support), plus `GET /v1/models`
* for the discovery route's own tests.
*/
import { createServer, type IncomingMessage, type Server } from 'node:http';
import { AddressInfo } from 'node:net';
export interface CapturedRequest {
method: string;
path: string;
headers: Record<string, string | string[] | undefined>;
body: unknown;
}
export interface MockOpenAiServer {
baseUrl: string;
requests: CapturedRequest[];
close(): Promise<void>;
}
function readJsonBody(req: IncomingMessage): Promise<unknown> {
return new Promise((resolve) => {
const chunks: Buffer[] = [];
req.on('data', (c) => chunks.push(c));
req.on('end', () => {
const raw = Buffer.concat(chunks).toString('utf8');
if (!raw) return resolve(undefined);
try {
resolve(JSON.parse(raw));
} catch {
resolve(raw);
}
});
});
}
/** Starts the mock server on a random free port and resolves once it's listening. */
export async function startMockOpenAiServer(): Promise<MockOpenAiServer> {
const requests: CapturedRequest[] = [];
const server: Server = createServer((req, res) => {
void (async () => {
const body = await readJsonBody(req);
const path = (req.url ?? '').split('?')[0];
requests.push({ method: req.method ?? 'GET', path, headers: req.headers, body });
res.setHeader('content-type', 'application/json');
if (path === '/v1/models' && req.method === 'GET') {
res.writeHead(200);
res.end(JSON.stringify({ data: [{ id: 'qwen3' }, { id: 'llama3' }] }));
return;
}
if (path === '/v1/chat/completions' && req.method === 'POST') {
res.writeHead(200);
res.end(
JSON.stringify({
id: 'mock-completion',
choices: [{ index: 0, message: { role: 'assistant', content: 'hello world' } }],
})
);
return;
}
if (path === '/v1/messages' && req.method === 'POST') {
res.writeHead(200);
res.end(
JSON.stringify({
id: 'mock-message',
role: 'assistant',
content: [{ type: 'text', text: 'hello world' }],
})
);
return;
}
// Codex's real wire protocol (verified against a live binary: it dropped
// wire_api="chat" support in Feb 2026, so its config.toml always says
// wire_api="responses") — a different shape from OpenAI's chat-completions.
if (path === '/v1/responses' && req.method === 'POST') {
res.writeHead(200);
res.end(
JSON.stringify({
id: 'mock-response',
output: [{ type: 'message', role: 'assistant', content: [{ type: 'output_text', text: 'hello world' }] }],
})
);
return;
}
res.writeHead(404);
res.end(JSON.stringify({ error: 'not found in mock server', path }));
})();
});
await new Promise<void>((resolve) => server.listen(0, '127.0.0.1', resolve));
const { port } = server.address() as AddressInfo;
return {
baseUrl: `http://127.0.0.1:${port}`,
requests,
close: () => new Promise<void>((resolve, reject) => server.close((err) => (err ? reject(err) : resolve()))),
};
}
+310 -2
View File
@@ -6,17 +6,34 @@
*/
import { describe, it, expect, beforeAll, beforeEach, afterAll, afterEach } from 'vitest';
import { closeSync, existsSync, openSync, readFileSync, writeFileSync, mkdirSync, rmSync, symlinkSync } from 'node:fs';
import {
chmodSync,
closeSync,
existsSync,
openSync,
readFileSync,
writeFileSync,
mkdirSync,
rmSync,
symlinkSync,
statSync,
readdirSync,
} from 'node:fs';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import { SETTINGS_PATH } from '../src/web/route-helpers.js';
import { tmpdir, homedir } from 'node:os';
import { spawn } from 'node:child_process';
import {
applyStatusLineConfig,
ensureCodemanHooks,
findEffectiveUserStatusLineCommand,
generateBackgroundWakeScript,
generateHooksConfig,
generateStatusLineCommand,
generateSubagentStopGuardScript,
readPlanUsageTelemetryEnabled,
refreshStaleCodemanHooks,
resolveStatusLineCliCommand,
settingsWriteBlocker,
stripCaseEnvKeys,
updateCaseEnvVars,
@@ -1305,3 +1322,294 @@ describe('Hook Config Generation - Extended', () => {
expect(stopHooks[0].hooks[0].command).toContain('stop');
});
});
describe('readPlanUsageTelemetryEnabled', () => {
const backup = existsSync(SETTINGS_PATH) ? readFileSync(SETTINGS_PATH, 'utf-8') : null;
afterEach(() => {
if (backup !== null) {
writeFileSync(SETTINGS_PATH, backup);
} else {
rmSync(SETTINGS_PATH, { force: true });
}
});
it('reads true fresh from the persisted showPlanUsageLimits setting', async () => {
mkdirSync(join(SETTINGS_PATH, '..'), { recursive: true });
writeFileSync(SETTINGS_PATH, JSON.stringify({ showPlanUsageLimits: true }));
expect(await readPlanUsageTelemetryEnabled()).toBe(true);
});
it('reads false when the setting is explicitly false', async () => {
mkdirSync(join(SETTINGS_PATH, '..'), { recursive: true });
writeFileSync(SETTINGS_PATH, JSON.stringify({ showPlanUsageLimits: false }));
expect(await readPlanUsageTelemetryEnabled()).toBe(false);
});
it('defaults to true when the setting is absent or the file is missing (mirrors readWorkspaceHooksEnabled)', async () => {
// The desktop chip shows as ON for an install that never touched the
// setting, so collection must agree with it. Resolving the default HERE
// is what keeps GET /api/settings a plain read (see its route test).
rmSync(SETTINGS_PATH, { force: true });
expect(await readPlanUsageTelemetryEnabled()).toBe(true);
mkdirSync(join(SETTINGS_PATH, '..'), { recursive: true });
writeFileSync(SETTINGS_PATH, JSON.stringify({ someOtherSetting: true }));
expect(await readPlanUsageTelemetryEnabled()).toBe(true);
});
it('only an explicit false turns collection off; junk values read as on', async () => {
mkdirSync(join(SETTINGS_PATH, '..'), { recursive: true });
writeFileSync(SETTINGS_PATH, JSON.stringify({ showPlanUsageLimits: 'no' }));
expect(await readPlanUsageTelemetryEnabled()).toBe(true);
});
it('never caches — a change on disk is visible on the very next call', async () => {
mkdirSync(join(SETTINGS_PATH, '..'), { recursive: true });
writeFileSync(SETTINGS_PATH, JSON.stringify({ showPlanUsageLimits: false }));
expect(await readPlanUsageTelemetryEnabled()).toBe(false);
writeFileSync(SETTINGS_PATH, JSON.stringify({ showPlanUsageLimits: true }));
expect(await readPlanUsageTelemetryEnabled()).toBe(true);
});
});
describe('resolveStatusLineCliCommand', () => {
const testDir = join(tmpdir(), 'codeman-statusline-cli-test-' + Date.now());
beforeEach(() => {
mkdirSync(testDir, { recursive: true });
});
afterEach(() => {
rmSync(testDir, { recursive: true, force: true });
});
it('returns undefined when telemetry was not requested', async () => {
expect(await resolveStatusLineCliCommand(testDir, false)).toBeUndefined();
});
it('returns a bare exporter SCRIPT PATH (never the inline command) when requested', async () => {
// A bare path has no `$`, quotes, or pipes for any intermediate shell
// layer to mangle — see ensureStatusLineExporterScript's doc comment for
// the real bug this guards against.
const cmd = await resolveStatusLineCliCommand(testDir, true);
expect(cmd).toBeDefined();
expect(cmd).not.toContain('$');
expect(cmd).not.toContain("'");
expect(cmd).toMatch(/^\/.*statusline-exporter\.sh$/);
expect(existsSync(cmd!)).toBe(true);
const stat = statSync(cmd!);
expect(stat.mode & 0o111).not.toBe(0); // executable
expect(readFileSync(cmd!, 'utf-8')).toContain('CODEMAN_STATUSLINE_EXPORTER_V');
});
it('refreshes a stale exporter script atomically: executable on arrival, no temp file left behind', async () => {
const scriptPath = (await resolveStatusLineCliCommand(testDir, true))!;
// Simulate a script an older build wrote (different marker suffix).
writeFileSync(scriptPath, '#!/bin/sh\n# CODEMAN_STATUSLINE_EXPORTER_V0\necho stale\n');
chmodSync(scriptPath, 0o644);
const again = await resolveStatusLineCliCommand(testDir, true);
expect(again).toBe(scriptPath);
expect(readFileSync(scriptPath, 'utf-8')).not.toContain('echo stale');
expect(statSync(scriptPath).mode & 0o111).not.toBe(0);
const siblings = readdirSync(join(scriptPath, '..')).filter((f) => f.startsWith('statusline-exporter.sh.'));
expect(siblings).toEqual([]);
});
it('never overrides a real, hand-authored statusLine', async () => {
const claudeDir = join(testDir, '.claude');
mkdirSync(claudeDir, { recursive: true });
writeFileSync(
join(claudeDir, 'settings.local.json'),
JSON.stringify({ statusLine: { type: 'command', command: 'echo my-own-prompt' } }, null, 2)
);
expect(await resolveStatusLineCliCommand(testDir, true)).toBeUndefined();
// The user's own config is untouched — this is a read-only decision, not a write.
const parsed = JSON.parse(readFileSync(join(claudeDir, 'settings.local.json'), 'utf-8'));
expect(parsed.statusLine.command).toBe('echo my-own-prompt');
});
it('self-heals: strips a legacy disk-written exporter from an older Codeman build', async () => {
// Simulate a workspace touched by the pre-fix applyStatusLineConfig(dir, true).
await applyStatusLineConfig(testDir, true);
const settingsPath = join(testDir, '.claude', 'settings.local.json');
expect(JSON.parse(readFileSync(settingsPath, 'utf-8')).statusLine).toBeDefined();
const cmd = await resolveStatusLineCliCommand(testDir, true);
// Cleaned off disk...
expect(JSON.parse(readFileSync(settingsPath, 'utf-8')).statusLine).toBeUndefined();
// ...and telemetry still flows, via the ephemeral CLI flag instead.
expect(cmd).toMatch(/statusline-exporter\.sh$/);
});
it('does not resurrect the legacy exporter when telemetry is off during cleanup', async () => {
await applyStatusLineConfig(testDir, true);
const settingsPath = join(testDir, '.claude', 'settings.local.json');
const cmd = await resolveStatusLineCliCommand(testDir, false);
expect(cmd).toBeUndefined();
expect(JSON.parse(readFileSync(settingsPath, 'utf-8')).statusLine).toBeUndefined();
});
});
describe('statusline exporter script (real shell execution)', () => {
const testDir = join(tmpdir(), 'codeman-statusline-script-exec-test-' + Date.now());
const binDir = join(tmpdir(), 'codeman-statusline-script-exec-bin-' + Date.now());
beforeEach(() => {
mkdirSync(testDir, { recursive: true });
mkdirSync(binDir, { recursive: true });
});
afterEach(() => {
rmSync(testDir, { recursive: true, force: true });
rmSync(binDir, { recursive: true, force: true });
});
// A stand-in for the real `curl` binary, placed FIRST on PATH — same technique
// the exporter's own review used ("an arg-echoing stand-in"). It ignores every
// arg curl would have received; only its own scripted behavior matters here.
function writeFakeCurl(script: string): void {
const curlPath = join(binDir, 'curl');
writeFileSync(curlPath, `#!/bin/sh\n${script}\n`);
chmodSync(curlPath, 0o755);
}
function runExporter(
env: Record<string, string>
): Promise<{ code: number | null; stdout: string; durationMs: number }> {
return resolveStatusLineCliCommand(testDir, true).then(
(scriptPath) =>
new Promise((resolve, reject) => {
const start = Date.now();
const child = spawn('sh', [scriptPath!], {
env: { ...env, PATH: `${binDir}:${process.env.PATH}` },
stdio: ['pipe', 'pipe', 'ignore'],
});
let stdout = '';
child.stdout.setEncoding('utf8');
child.stdout.on('data', (chunk) => {
stdout += chunk;
});
child.on('error', reject);
child.on('close', (code) => resolve({ code, stdout, durationMs: Date.now() - start }));
child.stdin.end('{}');
})
);
}
const baseEnv = {
CODEMAN_SESSION_ID: 'x',
CODEMAN_API_URL: 'http://127.0.0.1:1',
CODEMAN_HOOK_SECRET_FILE: '/dev/null',
};
it('no-user-statusline branch: the POST runs in the foreground and its OWN stdout becomes the footer', async () => {
writeFakeCurl(`echo 'model: opus | 42% used'`);
const result = await runExporter(baseEnv);
expect(result.stdout.trim()).toBe('model: opus | 42% used');
});
it('no-user-statusline branch: prints NOTHING when curl fails (never a bare brand word)', async () => {
writeFakeCurl(`exit 1`);
const result = await runExporter(baseEnv);
expect(result.stdout).toBe('');
expect(result.code).toBe(0);
});
it('asks curl to fail on HTTP errors (-f) so an error body never becomes the footer', async () => {
const scriptPath = await resolveStatusLineCliCommand(testDir, true);
expect(readFileSync(scriptPath!, 'utf-8')).toContain('curl -sfk');
expect(readFileSync(scriptPath!, 'utf-8')).not.toContain('echo codeman');
});
it('wrap branch: never blocks a reader-to-EOF on a slow/hung curl (background subshell closes stdin too)', async () => {
writeFakeCurl(`sleep 3`);
const result = await runExporter({ ...baseEnv, CODEMAN_USER_STATUSLINE_CMD: 'echo my-own-statusline' });
expect(result.stdout.trim()).toBe('my-own-statusline');
expect(result.durationMs).toBeLessThan(1000);
}, 10000);
it('curl is bounded with --max-time so a HUNG (not just refused) Codeman cannot wedge the render', async () => {
const scriptPath = await resolveStatusLineCliCommand(testDir, true);
expect(readFileSync(scriptPath!, 'utf-8')).toContain('--max-time');
});
});
describe('findEffectiveUserStatusLineCommand', () => {
const testDir = join(tmpdir(), 'codeman-statusline-precedence-test-' + Date.now());
const userSettingsPath = join(homedir(), '.claude', 'settings.json');
beforeEach(() => {
mkdirSync(testDir, { recursive: true });
});
afterEach(() => {
rmSync(testDir, { recursive: true, force: true });
rmSync(userSettingsPath, { force: true }); // don't leak into other tests sharing this HOME
});
it('returns undefined when nothing is configured anywhere', async () => {
expect(await findEffectiveUserStatusLineCommand(testDir)).toBeUndefined();
});
it('finds the user global ~/.claude/settings.json when nothing else is set', async () => {
const userClaudeDir = join(homedir(), '.claude');
mkdirSync(userClaudeDir, { recursive: true });
writeFileSync(
join(userClaudeDir, 'settings.json'),
JSON.stringify({ statusLine: { type: 'command', command: 'echo user-global' } })
);
expect(await findEffectiveUserStatusLineCommand(testDir)).toBe('echo user-global');
});
it('project-SHARED settings.json wins over user-global', async () => {
const userClaudeDir = join(homedir(), '.claude');
mkdirSync(userClaudeDir, { recursive: true });
writeFileSync(
join(userClaudeDir, 'settings.json'),
JSON.stringify({ statusLine: { type: 'command', command: 'echo user-global' } })
);
const projectClaudeDir = join(testDir, '.claude');
mkdirSync(projectClaudeDir, { recursive: true });
writeFileSync(
join(projectClaudeDir, 'settings.json'),
JSON.stringify({ statusLine: { type: 'command', command: 'echo project-shared' } })
);
expect(await findEffectiveUserStatusLineCommand(testDir)).toBe('echo project-shared');
});
it('project-LOCAL settings.local.json wins over everything', async () => {
const projectClaudeDir = join(testDir, '.claude');
mkdirSync(projectClaudeDir, { recursive: true });
writeFileSync(
join(projectClaudeDir, 'settings.json'),
JSON.stringify({ statusLine: { type: 'command', command: 'echo project-shared' } })
);
writeFileSync(
join(projectClaudeDir, 'settings.local.json'),
JSON.stringify({ statusLine: { type: 'command', command: 'echo project-local' } })
);
expect(await findEffectiveUserStatusLineCommand(testDir)).toBe('echo project-local');
});
it('skips a legacy Codeman-marked entry in project settings.local.json and falls through', async () => {
await applyStatusLineConfig(testDir, true); // simulates a pre-fix disk-written exporter
const projectClaudeDir = join(testDir, '.claude');
writeFileSync(
join(projectClaudeDir, 'settings.json'),
JSON.stringify({ statusLine: { type: 'command', command: 'echo project-shared' } })
);
expect(await findEffectiveUserStatusLineCommand(testDir)).toBe('echo project-shared');
});
});
+66
View File
@@ -146,6 +146,72 @@ describe('CJK input module', () => {
expect(sent).toEqual(['中文', ...committed]);
});
it('forwards Ctrl/Alt-modified navigation keys to the PTY, modifier intact', () => {
// claude prints "Jump to bottom (ctrl+End)" and the shortcut has to REACH it.
// PASSTHROUGH_KEYS carries only the plain forms, so Ctrl+End used to fail in
// both directions: with an empty field it went out as a bare `\x1b[F` (a
// plain End), and with any text in the field it was not forwarded at all —
// the browser default then moved the caret to the end of the composer, which
// is what the user sees as "the shortcut acts on the input box instead".
const { textarea, sent } = loadCjkHarness();
const preventDefault = vi.fn();
textarea.fire('keydown', {
key: 'End',
ctrlKey: true,
altKey: false,
shiftKey: false,
metaKey: false,
preventDefault,
});
expect(preventDefault).toHaveBeenCalled();
expect(sent).toEqual(['\x1b[1;5F']);
});
it('forwards a modified navigation key even when the composer has text', () => {
// The empty-field rule belongs to PLAIN navigation (which really is local
// editing); a Ctrl-modified one is a command for the CLI either way.
const { textarea, sent } = loadCjkHarness();
textarea.value = PHANTOM + '未发送的草稿';
textarea.fire('keydown', {
key: 'Home',
ctrlKey: true,
altKey: false,
shiftKey: false,
metaKey: false,
preventDefault: vi.fn(),
});
expect(sent).toEqual(['\x1b[1;5H']);
});
it('encodes the modifier bitmask, Shift included when it rides along', () => {
const { textarea, sent } = loadCjkHarness();
textarea.fire('keydown', {
key: 'ArrowUp',
ctrlKey: true,
shiftKey: true,
altKey: false,
metaKey: false,
preventDefault: vi.fn(),
});
expect(sent).toEqual(['\x1b[1;6A']); // 1 + shift(1) + ctrl(4)
});
it('leaves Shift-ALONE navigation local, so selecting in the composer still works', () => {
const { textarea, sent } = loadCjkHarness();
textarea.value = PHANTOM + '草稿';
const preventDefault = vi.fn();
textarea.fire('keydown', {
key: 'ArrowLeft',
shiftKey: true,
ctrlKey: false,
altKey: false,
metaKey: false,
preventDefault,
});
expect(sent).toEqual([]);
expect(preventDefault).not.toHaveBeenCalled();
});
it('recovers committed text when compositionend never fires (stuck composition)', () => {
const { textarea, sent } = loadCjkHarness();
+100
View File
@@ -597,6 +597,34 @@ describe('composer nav keys from the bar', () => {
const sentKeys = (fetchMock: { mock: { calls: unknown[][] } }) =>
fetchMock.mock.calls.map((call) => JSON.parse((call[1] as { body: string }).body).input);
it.each(['simple', 'extended'])('exposes Shift arrows in the %s agent layout', (mode) => {
const { bar, barElement } = loadBar('codex');
bar.setMode(mode);
expect(barElement.actions).toContain('shift-left');
expect(barElement.actions).toContain('shift-right');
});
it.each([
['shift-left', '\x1b[1;2D'],
['shift-right', '\x1b[1;2C'],
])('%s flushes the draft before navigation and hands editing to the PTY', (action, sequence) => {
const { bar, app, overlay, fetchMock } = barWithDraft('unfinished follow-up');
const events: string[] = [];
app.sendInput = vi.fn(() => events.push('draft'));
fetchMock.mockImplementation(() => {
events.push('key');
return Promise.resolve({ ok: true, catch: () => {} });
});
bar.handleAction(action);
expect(app.sendInput).toHaveBeenCalledWith('unfinished follow-up');
expect(events).toEqual(['draft', 'key']);
expect(sentKeys(fetchMock)).toEqual([sequence]);
expect(overlay.pendingText).toBe('');
expect([...(app._echoPassthroughSessions as Set<string>)]).toEqual(['session-1']);
});
it('flushes the unsent draft before sending the arrow', () => {
// On a phone the typed text lives in the overlay and has NEVER reached the
// PTY, so an arrow sent on its own arrives at a composer the CLI still
@@ -640,3 +668,75 @@ describe('composer nav keys from the bar', () => {
expect(app.sendInput).not.toHaveBeenCalled();
});
});
describe('Codex shift-arrow keys are gated on the active session', () => {
// ⇧←/⇧→ are Codex bindings. They ship in both agent templates, but a tap in
// any other CLI would do nothing AND hand the session to PTY echo
// (sendNavKey adds it to _echoPassthroughSessions), so on a phone a dead key
// would also switch local echo off for the rest of the prompt. The reveal
// follows the 🧠 key's shape: a marker class on the BAR element, because
// setMode() rebuilds the buttons' innerHTML on every layout switch.
const stylesSource = readFileSync(resolve('src/web/public/styles.css'), 'utf8');
it('marks both shift keys in both agent templates so one CSS rule can hide them', () => {
const simple = keyboardSource.match(/_simpleButtons\s*:\s*`([\s\S]*?)`/)?.[1] ?? '';
const extended = keyboardSource.match(/_extendedButtons\s*:\s*`([\s\S]*?)`/)?.[1] ?? '';
for (const template of [simple, extended]) {
expect(template).toMatch(/accessory-btn-codex[^>]*data-action="shift-left"/);
expect(template).toMatch(/accessory-btn-codex[^>]*data-action="shift-right"/);
}
});
it('hides the keys in styles.css until the bar carries codex-enabled', () => {
expect(stylesSource).toMatch(/\.keyboard-accessory-bar \.accessory-btn-codex \{\s*display: none;/);
expect(stylesSource).toMatch(
/\.keyboard-accessory-bar\.codex-enabled \.accessory-btn-codex \{\s*display: inline-flex;/
);
});
it.each(['simple', 'extended'])('carries codex-enabled for a codex session in the %s layout', (mode) => {
const { bar, barElement } = loadBar('codex');
bar.setMode(mode);
expect(barElement.classList.contains('codex-enabled')).toBe(true);
// The buttons themselves are still in the DOM; the class is what reveals them.
expect(barElement.actions).toContain('shift-left');
expect(barElement.actions).toContain('shift-right');
});
it.each(['claude', 'shell', 'pi', 'omp', 'deepseek'])(
'does not carry codex-enabled for a %s session',
(sessionMode) => {
const { bar, barElement } = loadBar(sessionMode);
bar.setMode('extended');
expect(barElement.classList.contains('codex-enabled')).toBe(false);
}
);
it('re-syncs the class on a session switch, in both directions', () => {
const { app, bar, barElement } = loadBar('codex');
expect(barElement.classList.contains('codex-enabled')).toBe(true);
app.sessions.set('session-2', { mode: 'claude' });
app.activeSessionId = 'session-2';
bar.refreshForActiveSession();
expect(barElement.classList.contains('codex-enabled')).toBe(false);
app.activeSessionId = 'session-1';
bar.refreshForActiveSession();
expect(barElement.classList.contains('codex-enabled')).toBe(true);
});
it('drops the class when no session is active (welcome screen)', () => {
const { app, bar, barElement } = loadBar('codex');
app.activeSessionId = '';
bar.refreshForActiveSession();
expect(barElement.classList.contains('codex-enabled')).toBe(false);
});
it('is wired at init and on every session switch, like the 🧠 key', () => {
const initBody = keyboardSource.match(/\n init\(\) \{([\s\S]*?)\n \},/)?.[1] ?? '';
const refreshBody = keyboardSource.match(/\n refreshForActiveSession\(\) \{([\s\S]*?)\n \},/)?.[1] ?? '';
expect(initBody).toContain('this.syncCodexKeys();');
expect(refreshBody).toContain('this.syncCodexKeys();');
});
});
+12 -1
View File
@@ -471,7 +471,18 @@ describe('Virtual Keyboard', () => {
});
// Tab replaced /clear in the simple bar; /clear and /compact live in the
// extended bar only.
expect(actions).toEqual(['scroll-up', 'scroll-down', 'init', 'tab', 'paste', 'esc', 'dismiss']);
expect(actions).toEqual([
'scroll-up',
'scroll-down',
'init',
'tab',
'shift-left',
'shift-right',
'paste',
'readmymind',
'esc',
'dismiss',
]);
});
it('double-tap confirm on /clear button', async () => {
+50 -1
View File
@@ -4,7 +4,7 @@
*/
import { EventEmitter } from 'node:events';
import { vi } from 'vitest';
import type { SessionStatus } from '../../src/types.js';
import type { SessionAttachmentHistoryItem, SessionStatus, SessionRemote } from '../../src/types.js';
/**
* Enhanced mock session for testing RespawnController.
@@ -13,6 +13,18 @@ import type { SessionStatus } from '../../src/types.js';
export class MockSession extends EventEmitter {
id: string;
workingDir: string = '/tmp/test-workdir';
/**
* Mirrors `Session.remote` — set to a `SessionRemote` to model a remote-SSH case,
* whose `workingDir` is an absolute path on ANOTHER host. File routes must read it
* over ssh instead of with local `fs` (#415).
*/
remote?: SessionRemote;
/** Mirrors Session.attachmentHistory (the attachment panel's source of truth). */
attachmentHistory: SessionAttachmentHistoryItem[] = [];
/** Mirrors Session.getAttachmentHistoryForPersist(). */
getAttachmentHistoryForPersist(): SessionAttachmentHistoryItem[] {
return this.attachmentHistory;
}
/**
* The REAL union, deliberately. This used to be `'idle' | 'working'`, and
* `'working'` is not a `SessionStatus` at all — so `signalForStatus()` fell to its
@@ -318,6 +330,43 @@ export class MockSession extends EventEmitter {
this.color = c;
});
/** Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) */
customModel: { endpointId: string; modelId: string; label?: string } | undefined = undefined;
remote: unknown = undefined;
docker: unknown = undefined;
private _mockCustomModel:
| {
endpointId: string;
modelId: string;
label?: string;
envKeys: string[];
configDir?: string;
launchModel?: string;
}
| undefined;
setCustomModel = vi.fn(
(
next:
| {
endpointId: string;
modelId: string;
label?: string;
envKeys: string[];
configDir?: string;
launchModel?: string;
}
| undefined,
_envOverrides?: Record<string, string>
): { removedEnvKeys: string[]; previousConfigDir: string | undefined } => {
const previous = this._mockCustomModel;
this._mockCustomModel = next;
this.customModel = next ? { endpointId: next.endpointId, modelId: next.modelId, label: next.label } : undefined;
return { removedEnvKeys: previous?.envKeys ?? [], previousConfigDir: previous?.configDir };
}
);
restartCli = vi.fn(async () => true);
getCustomModelForPersist = vi.fn(() => this._mockCustomModel);
/** Stub for sendInput */
sendInput = vi.fn();
+53
View File
@@ -0,0 +1,53 @@
/**
* A persistent full-screen overlay must not carry `backdrop-filter` while it is
* hidden.
*
* The property promotes the element to its own compositing layer, and a
* full-screen `position: fixed` layer that is created and then hidden has been
* observed to leave a stale HIT-TEST region behind in Chrome: the page renders
* correctly while pointer events over the viewport land on nothing. Reported on
* a long-lived tab against a remote server (where a connection blip shows and
* then hides #offlineOverlay): terminal scrolling AND unrelated click-to-expand
* controls died together, a freshly opened tab was fine, and a console
* one-liner doing nothing but READING layout restored it.
*/
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import { describe, expect, it } from 'vitest';
const css = readFileSync(resolve(import.meta.dirname, '../src/web/public/styles.css'), 'utf8');
function ruleBody(selector: string): string {
const escaped = selector.replace(/[.*+?^${}()|[\]\\]/g, '\\$&');
const m = new RegExp(`(?:^|\\})\\s*${escaped}\\s*\\{([^}]*)\\}`, 'm').exec(css);
if (!m) throw new Error(`no rule for ${selector}`);
return m[1];
}
/** Persistent, full-screen, fixed overlays and the selector that shows each. */
const PERSISTENT_OVERLAYS: Array<{ base: string; shown: string }> = [
{ base: '.offline-overlay', shown: '.offline-overlay:not([hidden])' },
{ base: '.file-preview-overlay', shown: '.file-preview-overlay.visible' },
];
describe('persistent full-screen overlays do not composite while hidden', () => {
for (const { base, shown } of PERSISTENT_OVERLAYS) {
it(`${base} keeps backdrop-filter off its base rule`, () => {
expect(ruleBody(base)).not.toMatch(/backdrop-filter/);
});
it(`${base} still blurs once shown, via ${shown}`, () => {
// Moving the property must not silently DELETE the effect: the overlay is
// meant to blur what is behind it while it is up.
const body = ruleBody(shown);
expect(body).toMatch(/(^|\s)backdrop-filter:\s*blur\(/m);
expect(body).toMatch(/-webkit-backdrop-filter:\s*blur\(/);
});
}
it('the offline overlay still forces display:none when hidden', () => {
// The base rule is `display: flex`, so [hidden] alone would not hide it —
// this is the guard that rule stays put while the block is edited.
expect(ruleBody('.offline-overlay[hidden]')).toMatch(/display:\s*none\s*!important/);
});
});
+78
View File
@@ -0,0 +1,78 @@
/**
* `planUsageCollectionFlip()` in settings-ui.js: the one place that decides
* whether a settings save carries `showPlanUsageLimits` to the server.
*
* The chip is per-device for DISPLAY (desktop default ON, handhelds OFF) but
* the same persisted key is the server-side telemetry COLLECTION switch, read
* at every claude spawn. Sending it on every save let a phone saving its font
* size persist `false` and turn collection off for every desktop. So the save
* sends the key ONLY when it flips the chip relative to what the device had,
* and the server reads an absent key as ON.
*/
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it } from 'vitest';
const SOURCE = readFileSync(resolve(import.meta.dirname, '../src/web/public/settings-ui.js'), 'utf8');
function loadSettingsUi(defaultChip: boolean) {
const CodemanApp = function CodemanApp(this: unknown) {};
const context = vm.createContext({
CodemanApp,
VoiceInput: {},
localStorage: { getItem: () => null, setItem: () => {} },
document: { getElementById: () => null },
console,
});
vm.runInContext(SOURCE, context, { filename: 'settings-ui.js' });
const app = Object.create(CodemanApp.prototype) as {
getDefaultSettings: () => { showPlanUsageLimits: boolean };
planUsageCollectionFlip: (prev: Record<string, unknown> | null, now: boolean) => boolean | undefined;
};
app.getDefaultSettings = () => ({ showPlanUsageLimits: defaultChip });
return app;
}
describe('planUsageCollectionFlip', () => {
it('says nothing when a desktop that never touched the chip saves with it still on', () => {
const desktop = loadSettingsUi(true);
expect(desktop.planUsageCollectionFlip({}, true)).toBeUndefined();
expect(desktop.planUsageCollectionFlip(null, true)).toBeUndefined();
});
it('says nothing when a handheld (chip default OFF) saves an unrelated setting', () => {
const phone = loadSettingsUi(false);
expect(phone.planUsageCollectionFlip({ terminalFontSize: 14 }, false)).toBeUndefined();
expect(phone.planUsageCollectionFlip({ showPlanUsageLimits: false }, false)).toBeUndefined();
});
it('sends false only on the save that turned the chip off', () => {
const desktop = loadSettingsUi(true);
expect(desktop.planUsageCollectionFlip({}, false)).toBe(false);
expect(desktop.planUsageCollectionFlip({ showPlanUsageLimits: true }, false)).toBe(false);
expect(desktop.planUsageCollectionFlip({ showPlanUsageLimits: false }, false)).toBeUndefined();
});
it('sends true when any device, a handheld included, turns the chip on', () => {
const phone = loadSettingsUi(false);
expect(phone.planUsageCollectionFlip({}, true)).toBe(true);
expect(phone.planUsageCollectionFlip({ showPlanUsageLimits: false }, true)).toBe(true);
expect(phone.planUsageCollectionFlip({ showPlanUsageLimits: true }, true)).toBeUndefined();
});
});
describe('saveAppSettings wiring', () => {
it('strips showPlanUsageLimits from the synced payload and re-adds it only through the flip', () => {
const save = SOURCE.slice(
SOURCE.indexOf('async saveAppSettings()'),
SOURCE.indexOf('closeAppSettings()', SOURCE.indexOf('async saveAppSettings()'))
);
// Stripped from serverSettings like the other per-device display keys.
expect(save).toMatch(/showPlanUsageLimits: _pul,/);
// Decided once against the device's prior settings, before they are overwritten.
expect(save).toMatch(/const _chipFlip = this\.planUsageCollectionFlip\(_prev, settings\.showPlanUsageLimits\);/);
// And only a real flip reaches the PUT body.
expect(save).toMatch(/\.\.\.\(_chipFlip !== undefined \? \{ showPlanUsageLimits: _chipFlip \} : \{\}\),/);
});
});
+36
View File
@@ -71,3 +71,39 @@ describe('Session.shouldApplyInput (exactly-once input dedup)', () => {
expect(s.shouldApplyInput('recent', 2)).toBe(true);
});
});
describe('Session.lastInputSeq (the watermark a stuck client needs)', () => {
it('reports 0 for a client it has never seen', () => {
expect(makeSession().lastInputSeq('c-new')).toBe(0);
});
it('reports the highest seq applied for that client', () => {
const s = makeSession();
s.shouldApplyInput('c-1', 7);
expect(s.lastInputSeq('c-1')).toBe(7);
});
it('is what a rolled-back client must clear to be heard again', () => {
// The failure this exists for: the client's seq counter persists on a
// DEBOUNCED write, so a tab killed between a send and that write comes back
// counting from below the watermark. Every later keystroke then lands at or
// under it and is rejected — silently, because a rejected frame is ACKed too.
const s = makeSession();
for (let i = 1; i <= 40; i++) s.shouldApplyInput('c-1', i);
// Restored counter starts over at 1: dropped, and every subsequent one too.
expect(s.shouldApplyInput('c-1', 1)).toBe(false);
expect(s.shouldApplyInput('c-1', 2)).toBe(false);
// The watermark it is handed back is exactly what makes it recoverable.
const watermark = s.lastInputSeq('c-1');
expect(watermark).toBe(40);
expect(s.shouldApplyInput('c-1', watermark + 1)).toBe(true);
});
it('does not resurrect a seq that forgetInputSeq rolled back', () => {
const s = makeSession();
s.shouldApplyInput('c-1', 5);
s.forgetInputSeq('c-1', 5);
expect(s.lastInputSeq('c-1')).toBe(4);
expect(s.shouldApplyInput('c-1', 5)).toBe(true);
});
});
+78
View File
@@ -0,0 +1,78 @@
/**
* @fileoverview A client whose seq counter rolled back must heal itself.
*
* The failure: the browser tags input with (clientId, seq) and persists the
* counters to localStorage on a DEBOUNCED write. A tab killed between a send and
* that write comes back counting from BELOW the server's watermark, so every
* later keystroke is rejected as a duplicate — and, because a rejected frame was
* ACKed exactly like an applied one, the client dropped it from its queue and the
* UI looked perfectly healthy while the terminal took no input at all. Reloading
* could not help: clientId and the stale counter both come back from localStorage.
*
* Observed live on a server session: a fresh browser (new clientId, no watermark)
* typed into the same session fine, which is what isolated it to client state.
*/
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import { describe, it, expect } from 'vitest';
const appSource = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8');
const wsSource = readFileSync(resolve(import.meta.dirname, '../src/web/routes/ws-routes.ts'), 'utf8');
describe('the duplicate ACK carries what the client needs', () => {
it('marks a rejected frame as dup and reports the watermark', () => {
// A bare ACK is indistinguishable from "applied" — that ambiguity is the bug.
expect(wsSource).toMatch(/"dup":true,"last":\$\{watermark\}/);
expect(wsSource).toContain('lastInputSeq');
});
it('reads the watermark defensively, so a port without it still ACKs', () => {
// The session arrives through a structural port. A throw inside the message
// handler aborts it before the ACK is sent, stranding the frame in the
// client's durable queue — which is worse than the ambiguity being fixed here.
expect(wsSource).toMatch(/typeof \(session as \{ lastInputSeq\?/);
});
it('still ACKs a rejected frame, so the client can drop it from its queue', () => {
// Silence would strand the record and the redelivery sweep would spin on it.
const block = wsSource.slice(wsSource.indexOf('if (seq !== null && socket.readyState === 1)'));
expect(block.slice(0, 1200)).toContain('"t":"ia"');
});
});
describe('the client lifts itself over the watermark', () => {
const handler = appSource.slice(
appSource.indexOf('_onWsInputAck(seq, msg)'),
appSource.indexOf('/** Called from ws.onopen')
);
it('raises the counter to the watermark it was handed', () => {
expect(handler).toMatch(/_seqCounters\.set\(sessionId, watermark\)/);
});
it('re-queues a FIRST-attempt frame, whose input was genuinely lost', () => {
expect(handler).toMatch(/rec\.tries <= 1/);
expect(handler).toMatch(/this\._reliableSend\(sessionId, lost/);
});
it('does NOT re-queue a retry, which the dedup correctly suppressed', () => {
// A retry called a duplicate means the original DID land; re-sending it would
// type the same thing twice — the exact thing exactly-once delivery prevents.
expect(handler).toMatch(/const lost = rec && rec\.tries <= 1 \? rec\.data : null;/);
});
it('persists the raised counter immediately, not on the debounce', () => {
expect(handler).toContain('this._persistReliableNow()');
});
});
describe('the seq counter is persisted synchronously on every send', () => {
it('_reliableSend uses the immediate writer, never the debounced one', () => {
// The counter is precisely what must survive a crash, so it cannot ride the
// path most likely to be lost. (The queue PAYLOAD may still be debounced.)
const send = appSource.slice(appSource.indexOf('_reliableSend(sessionId, data, useMux)'));
const body = send.slice(0, send.indexOf('_nextSeq(sessionId) {'));
expect(body).toContain('this._persistReliableNow();');
expect(body).not.toMatch(/list\.push\(rec\);\s*\n\s*this\._persistReliableState\(\);/);
});
});
+402
View File
@@ -0,0 +1,402 @@
/**
* @fileoverview Tests for remote (SSH) file access (`src/remote-files.ts`).
*
* Two layers are covered:
*
* 1. PURE builders/parsers — command construction, escaping and probe parsing, no
* connection involved.
* 2. The probe SCRIPT itself, executed by a real `/bin/sh` against a real temp
* directory. The remote shell is the one place where a quoting mistake becomes an
* injection, and it cannot be exercised by an ssh-less unit test any other way: the
* script IS the remote command, so `sh -c <script>` reproduces exactly what sshd
* runs on the other end.
*
* Port: N/A (no HTTP server).
*/
import { describe, it, expect, beforeAll, afterAll } from 'vitest';
import { execFileSync } from 'node:child_process';
import { mkdtempSync, mkdirSync, rmSync, writeFileSync, existsSync, statSync, chmodSync } from 'node:fs';
import { join } from 'node:path';
import { homedir, tmpdir } from 'node:os';
import {
RemoteFileAccessError,
buildRemoteFileCommand,
buildRemoteProbeCommand,
buildRemoteReadCommand,
parseRemoteProbeRecord,
parseRemoteProbeOutput,
remoteProbePaths,
remoteReadFile,
remoteCreateReadStream,
} from '../src/remote-files.js';
import type { SessionRemote } from '../src/types/session.js';
/**
* Run a shell line through a real `/bin/sh` and return its `$@` as an argv array,
* WITHOUT executing anything. This is how the tests see the exact argument vector a
* command line would hand to the process — the local-shell half of the escaping chain.
*/
function shellArgv(command: string): string[] {
const out = execFileSync('sh', ['-c', `set -- ${command}; printf '%s\\0' "$@"`]);
// The trailing empty element is the printf format terminator.
return out.toString().split('\0').slice(0, -1);
}
/** A remote session fixture; every field is optional in production, so keep it minimal. */
function remoteFixture(overrides: Partial<SessionRemote> = {}): SessionRemote {
return {
hostId: 'host-1',
label: 'testhost',
host: '192.0.2.10',
username: 'j',
remotePath: '/srv/case',
...overrides,
};
}
describe('buildRemoteFileCommand', () => {
it('builds the ssh line from the shared connection args and one shellescaped command', () => {
const argv = shellArgv(buildRemoteFileCommand(remoteFixture(), 'cat /etc/hostname'));
// buildSshConnectionArgs returns tokens, and the shell re-splits them into the
// flags ssh actually wants (`-o` + `BatchMode=yes`), which is what this pins.
expect(argv.slice(0, 3)).toEqual(['ssh', '-o', 'BatchMode=yes']);
expect(argv).toContain('ConnectTimeout=10');
expect(argv).toContain('j@192.0.2.10');
// The remote command is ONE argument, whatever it contains.
expect(argv[argv.length - 1]).toBe('cat /etc/hostname');
expect(argv[argv.length - 2]).toBe('j@192.0.2.10');
});
it('routes port, identity, jump host and extra options through buildSshConnectionArgs', () => {
const argv = shellArgv(
buildRemoteFileCommand(
remoteFixture({
port: 2222,
identityFile: '~/.ssh/id_ed25519',
jumpHost: 'bastion.example.com',
extraSshOptions: ['StrictHostKeyChecking=accept-new'],
}),
'true'
)
);
expect(argv).toContain('-p');
expect(argv).toContain('2222');
expect(argv).toContain('-J');
expect(argv).toContain('bastion.example.com');
expect(argv).toContain('StrictHostKeyChecking=accept-new');
// `~` is expanded before escaping: ssh does not expand it inside -i.
expect(argv).toContain(join(homedir(), '.ssh/id_ed25519'));
});
it('keeps a shell-metacharacter command as a single opaque argument', () => {
const command = "cat '/tmp/it''s here' ; rm -rf ~ #";
const argv = shellArgv(buildRemoteFileCommand(remoteFixture(), command));
expect(argv[argv.length - 1]).toBe(command);
expect(argv).not.toContain('rm');
expect(argv).not.toContain('-rf');
});
});
describe('buildRemoteProbeCommand', () => {
it('probes every path exactly once, each as its own shell-quoted token', () => {
const script = buildRemoteProbeCommand(['/srv/case/a.png', '/srv/case']);
const probeCalls = script.split('\n').filter((line) => line.startsWith('probe '));
// The index is what the parser keys records on, so it is part of the call.
expect(probeCalls).toEqual(["probe 0 '/srv/case/a.png'", "probe 1 '/srv/case'"]);
});
it('quotes a path with spaces, quotes and a command substitution', () => {
const nasty = "/srv/case/it's $(touch /tmp/pwned).txt";
const script = buildRemoteProbeCommand([nasty]);
expect(script).toContain(`probe 0 '/srv/case/it'\\''s $(touch /tmp/pwned).txt'`);
expect(shellArgv(buildRemoteFileCommand(remoteFixture(), script)).at(-1)).toBe(script);
});
});
/**
* Run the probe script through a real `/bin/sh`. With `shadowReadlinkF` the PATH is
* fronted by a `readlink` that rejects `-f` the way macOS < 12.3 does (`illegal
* option -- f`) and otherwise defers to the real one, which forces the portable
* fallback branch on a host that natively has `readlink -f`.
*/
function runProbe(paths: string[], options: { cwd?: string; shadowReadlinkF?: boolean; shimDir?: string } = {}) {
const env =
options.shadowReadlinkF && options.shimDir
? { ...process.env, PATH: `${options.shimDir}:${process.env.PATH}` }
: process.env;
const stdout = execFileSync('sh', ['-c', buildRemoteProbeCommand(paths)], { cwd: options.cwd, env }).toString();
return parseRemoteProbeOutput(stdout, paths);
}
describe('the probe script on a real shell', () => {
let root: string;
let shimDir: string;
beforeAll(() => {
root = mkdtempSync(join(tmpdir(), 'codeman-remote-probe-'));
shimDir = join(root, 'shim-bin');
mkdirSync(shimDir);
const realReadlink = execFileSync('sh', ['-c', 'command -v readlink']).toString().trim();
writeFileSync(
join(shimDir, 'readlink'),
`#!/bin/sh\ncase "$1" in -f) echo 'readlink: illegal option -- f' >&2; exit 1;; esac\nexec ${realReadlink} "$@"\n`
);
chmodSync(join(shimDir, 'readlink'), 0o755);
});
afterAll(() => {
rmSync(root, { recursive: true, force: true });
});
it('resolves the fallback branch on a shell whose readlink has no -f', () => {
// Sanity check on the shim itself: without it this whole describe would be
// exercising the native branch twice.
expect(() =>
execFileSync('sh', ['-c', 'readlink -f / 2>/dev/null'], {
env: { ...process.env, PATH: `${shimDir}:${process.env.PATH}` },
})
).toThrow();
});
it.each([
['readlink -f', false],
['portable fallback', true],
])('refuses to report a symlink by its own path (%s): the target is what is served', (_label, shadow) => {
// The reviewer's exact reproduction: ws/notes.txt -> secret/id_rsa. The old
// fallback resolved only the DIRECTORY chain, returned `ws/notes.txt` as the
// realpath (with the TARGET's size), containment passed, and `cat` served the key.
const ws = join(root, `escape-${shadow ? 'fallback' : 'native'}`);
const secret = join(root, `secret-${shadow ? 'fallback' : 'native'}`);
mkdirSync(ws);
mkdirSync(secret);
writeFileSync(join(secret, 'id_rsa'), 'KEYKEYKEYKEY1');
execFileSync('ln', ['-s', join(secret, 'id_rsa'), join(ws, 'notes.txt')]);
const [probe] = runProbe([join(ws, 'notes.txt')], { shadowReadlinkF: shadow, shimDir });
expect(probe?.realPath).toBe(join(secret, 'id_rsa'));
expect(probe?.size).toBe(13);
});
it('follows a relative symlink chain through a symlinked directory on the fallback branch', () => {
const ws = join(root, 'chain');
mkdirSync(join(ws, 'sub'), { recursive: true });
writeFileSync(join(ws, 'sub', 'real.txt'), 'inside');
execFileSync('ln', ['-s', 'real.txt', join(ws, 'sub', 'hop1.txt')]);
execFileSync('ln', ['-s', 'hop1.txt', join(ws, 'sub', 'hop2.txt')]);
execFileSync('ln', ['-s', 'sub', join(ws, 'subl')]);
const probes = runProbe([join(ws, 'subl', 'hop2.txt'), join(ws, 'subl')], { shadowReadlinkF: true, shimDir });
expect(probes[0]).toMatchObject({ kind: 'file', size: 6, realPath: join(ws, 'sub', 'real.txt') });
expect(probes[1]).toMatchObject({ kind: 'directory', realPath: join(ws, 'sub') });
});
it.each([
['readlink -f', false],
['portable fallback', true],
])('fails CLOSED on a symlink loop (%s), never reporting the unresolved path', (_label, shadow) => {
const ws = join(root, `loop-${shadow ? 'fallback' : 'native'}`);
mkdirSync(ws);
execFileSync('ln', ['-s', 'b', join(ws, 'a')]);
execFileSync('ln', ['-s', 'a', join(ws, 'b')]);
const [probe] = runProbe([join(ws, 'a')], { shadowReadlinkF: shadow, shimDir });
expect(probe).toBeNull();
});
it('reports kind, size and realpath for a file, a directory and a missing path', () => {
const filePath = join(root, 'image.png');
writeFileSync(filePath, 'fake png bytes');
const probes = parseRemoteProbeOutput(
execFileSync('sh', ['-c', buildRemoteProbeCommand([filePath, root, join(root, 'nope.png')])]).toString(),
[filePath, root, join(root, 'nope.png')]
);
expect(probes[0]).toMatchObject({ kind: 'file', size: 14, realPath: filePath });
expect(probes[0]?.mtimeMs).toBeGreaterThan(0);
expect(probes[1]).toMatchObject({ kind: 'directory', size: 0, realPath: root });
expect(probes[2]).toBeNull();
});
it('resolves a symlink to its target', () => {
const target = join(root, 'target.txt');
const link = join(root, 'link.txt');
writeFileSync(target, 'x');
execFileSync('ln', ['-s', target, link]);
const [probe] = parseRemoteProbeOutput(execFileSync('sh', ['-c', buildRemoteProbeCommand([link])]).toString(), [
link,
]);
expect(probe?.realPath).toBe(target);
});
it('treats a hostile filename as data, never as a command', () => {
// No slashes in the payload: it has to be a legal FILENAME on this host while
// still being a command substitution to a shell.
const marker = `codeman_pwned_${process.pid}`;
const hostile = join(root, `it's; touch ${marker}; $(id).txt`);
writeFileSync(hostile, 'hostile');
const [probe] = parseRemoteProbeOutput(
execFileSync('sh', ['-c', buildRemoteProbeCommand([hostile])], { cwd: root }).toString(),
[hostile]
);
expect(probe?.realPath).toBe(hostile);
expect(existsSync(join(root, marker))).toBe(false);
});
it('keeps a filename containing a newline aligned with its own index', () => {
// One record per LINE would have made this two lines, shifting every record
// after it by one; records are NUL-terminated and index-keyed instead.
const weird = join(root, 'a\nb.txt');
writeFileSync(weird, 'nl');
const after = join(root, 'after.txt');
writeFileSync(after, 'after');
const probes = runProbe([weird, after, join(root, 'nope')]);
expect(probes[0]).toMatchObject({ kind: 'file', size: 2, realPath: weird });
expect(probes[1]).toMatchObject({ kind: 'file', size: 5, realPath: after });
expect(probes[2]).toBeNull();
});
it('discards a login banner and rc-file chatter printed before the records', () => {
const filePath = join(root, 'banner.txt');
writeFileSync(filePath, 'b');
const stdout = execFileSync('sh', [
'-c',
`echo 'Welcome to box'; printf '0|f|9|9|/etc/shadow\\n'; ${buildRemoteProbeCommand([filePath])}`,
]).toString();
// The chatter even LOOKS like a record; the leading NUL is what fences it off.
expect(parseRemoteProbeOutput(stdout, [filePath])[0]).toMatchObject({ realPath: filePath, size: 1 });
});
it('handles a path containing the field separator', () => {
const pipePath = join(root, 'a|b.txt');
writeFileSync(pipePath, 'xy');
const [probe] = parseRemoteProbeOutput(execFileSync('sh', ['-c', buildRemoteProbeCommand([pipePath])]).toString(), [
pipePath,
]);
expect(probe?.realPath).toBe(pipePath);
expect(probe?.size).toBe(2);
});
it('walks into a nested directory that exists', () => {
const nested = join(root, 'sub');
mkdirSync(nested, { recursive: true });
writeFileSync(join(nested, 'f.txt'), 'abc');
const [probe] = parseRemoteProbeOutput(
execFileSync('sh', ['-c', buildRemoteProbeCommand([join(nested, 'f.txt')])]).toString(),
[join(nested, 'f.txt')]
);
expect(probe?.size).toBe(3);
expect(statSync(join(nested, 'f.txt')).size).toBe(3);
});
});
describe('parseRemoteProbeRecord', () => {
it('parses a file record and converts mtime to milliseconds', () => {
expect(parseRemoteProbeRecord('f|1234|1700000000|/srv/case/a.png')).toEqual({
realPath: '/srv/case/a.png',
kind: 'file',
size: 1234,
mtimeMs: 1700000000 * 1000,
});
});
it('keeps a path that itself contains the separator', () => {
expect(parseRemoteProbeRecord('f|7|0|/srv/ca|se/a b.txt')?.realPath).toBe('/srv/ca|se/a b.txt');
});
it('maps directories, other kinds, the not-found and the unresolvable markers', () => {
expect(parseRemoteProbeRecord('d|0|5|/srv/case')?.kind).toBe('directory');
expect(parseRemoteProbeRecord('o|0|0|/srv/case/sock')?.kind).toBe('other');
expect(parseRemoteProbeRecord('n')).toBeNull();
// Exists but could not be canonicalized: refused like a missing file, never
// served under a path whose real target is unknown.
expect(parseRemoteProbeRecord('x')).toBeNull();
expect(parseRemoteProbeRecord('')).toBeNull();
});
it('rejects malformed lines instead of inventing a path', () => {
expect(parseRemoteProbeRecord('f|1|2')).toBeNull();
expect(parseRemoteProbeRecord('x|1|2|/p')).toBeNull();
expect(parseRemoteProbeRecord('f|1|2|')).toBeNull();
});
});
describe('parseRemoteProbeOutput', () => {
it('keys records by index after the leading NUL, so a login banner cannot shift the mapping', () => {
const stdout = 'welcome to the remote box\n\x000|f|3|1|/srv/a.txt\x001|n\x00';
expect(parseRemoteProbeOutput(stdout, ['/srv/a.txt', '/srv/b.txt'])).toEqual([
{ realPath: '/srv/a.txt', kind: 'file', size: 3, mtimeMs: 1000 },
null,
]);
});
it('accepts records in any order and ignores duplicates of an index', () => {
const stdout = '\x001|d|0|0|/srv\x000|f|3|1|/srv/a.txt\x000|f|9|9|/evil\x00';
expect(parseRemoteProbeOutput(stdout, ['/srv/a.txt', '/srv'])).toEqual([
{ realPath: '/srv/a.txt', kind: 'file', size: 3, mtimeMs: 1000 },
{ realPath: '/srv', kind: 'directory', size: 0, mtimeMs: 0 },
]);
});
it('throws when a requested path has no record (transport or shell failure, never a 404)', () => {
expect(() => parseRemoteProbeOutput('\x000|f|3|1|/srv/a.txt\x00', ['/a', '/b'])).toThrow(RemoteFileAccessError);
expect(() => parseRemoteProbeOutput('', ['/a'])).toThrow(RemoteFileAccessError);
// No leading NUL at all: the script never ran, whatever the shell printed.
expect(() => parseRemoteProbeOutput('0|f|3|1|/srv/a.txt', ['/srv/a.txt'])).toThrow(RemoteFileAccessError);
});
});
describe('under vitest', () => {
const remote = remoteFixture();
it('never opens a connection: probes and reads reject with a clear error', async () => {
// Mirrors checkRemoteTmuxAvailable's guard. The route tests mock this module, so
// this is the backstop for the next test that reaches the real one.
await expect(remoteProbePaths(remote, ['/srv/case'])).rejects.toThrow(/disabled under test/);
await expect(remoteReadFile(remote, '/srv/case/a.txt', 1024)).rejects.toThrow(/disabled under test/);
});
it('never opens a connection: a stream fails through its own error path', async () => {
const { stream, close } = remoteCreateReadStream(remote, '/srv/case/a.mp4');
const failure = await new Promise<Error>((resolveError) => stream.on('error', resolveError));
expect(failure).toBeInstanceOf(RemoteFileAccessError);
expect(() => close()).not.toThrow();
});
});
describe('buildRemoteReadCommand', () => {
it('streams the whole file with cat', () => {
expect(buildRemoteReadCommand("/srv/case/it's.mp4")).toBe("cat '/srv/case/it'\\''s.mp4'");
});
it('turns a byte range into a constant-memory tail | head', () => {
expect(buildRemoteReadCommand('/srv/case/v.mp4', { start: 2, end: 5 })).toBe(
"tail -c +3 '/srv/case/v.mp4' | head -c 4"
);
});
it('covers the first byte of the file (tail -c +1, not +0)', () => {
expect(buildRemoteReadCommand('/f', { start: 0, end: 0 })).toBe("tail -c +1 '/f' | head -c 1");
});
});
+67
View File
@@ -0,0 +1,67 @@
/**
* @fileoverview Tests for the remote-file ssh concurrency limiter
* (`src/remote-ssh-limiter.ts`): the cap holds under interleaved async resumption,
* waiters are served FIFO, and a task that throws still releases its slot.
*
* Port: N/A (no HTTP server).
*/
import { describe, it, expect } from 'vitest';
import {
getActiveRemoteSshCount,
getQueuedRemoteSshCount,
getRemoteSshLimit,
runWithRemoteSshLimit,
} from '../src/remote-ssh-limiter.js';
function deferred(): { promise: Promise<void>; resolve: () => void } {
let resolve!: () => void;
const promise = new Promise<void>((r) => {
resolve = r;
});
return { promise, resolve };
}
describe('runWithRemoteSshLimit', () => {
it('never lets more than the cap run at once, and queues the rest FIFO', async () => {
const cap = getRemoteSshLimit();
const gates = Array.from({ length: cap + 3 }, () => deferred());
const started: number[] = [];
let peak = 0;
const runs = gates.map((gate, index) =>
runWithRemoteSshLimit(async () => {
started.push(index);
peak = Math.max(peak, getActiveRemoteSshCount());
await gate.promise;
return index;
})
);
await Promise.resolve();
expect(started).toEqual(Array.from({ length: cap }, (_, i) => i));
expect(getActiveRemoteSshCount()).toBe(cap);
expect(getQueuedRemoteSshCount()).toBe(3);
// Releasing one hands the slot to the OLDEST waiter; the count stays at the cap.
gates[0].resolve();
await runs[0];
await Promise.resolve();
expect(started).toEqual([...Array.from({ length: cap }, (_, i) => i), cap]);
expect(getActiveRemoteSshCount()).toBe(cap);
for (const gate of gates) gate.resolve();
expect(await Promise.all(runs)).toEqual(gates.map((_, i) => i));
expect(peak).toBe(cap);
expect(getActiveRemoteSshCount()).toBe(0);
expect(getQueuedRemoteSshCount()).toBe(0);
});
it('releases the slot when the task throws', async () => {
await expect(runWithRemoteSshLimit(async () => Promise.reject(new Error('ssh exit 255')))).rejects.toThrow(
'ssh exit 255'
);
expect(getActiveRemoteSshCount()).toBe(0);
expect(await runWithRemoteSshLimit(async () => 'after')).toBe('after');
});
});
+196
View File
@@ -0,0 +1,196 @@
/**
* @fileoverview Route tests for Custom Model Endpoint Profiles CRUD + discovery.
*
* Discovery is mocked at `webviewFetch()` (webview-egress.ts), NOT at the global
* `fetch`: the route deliberately goes through the guarded undici dispatcher whose
* lookup hook refuses a name that resolves into a link-local / cloud-metadata range,
* so a global-fetch stub that still satisfied these tests would mean the guard had
* been bypassed.
* Port: N/A (app.inject, no real port needed)
*/
import { describe, it, expect, vi, afterEach } from 'vitest';
import { registerCustomModelRoutes } from '../../src/web/routes/custom-model-routes.js';
import { webviewFetch } from '../../src/web/webview-egress.js';
import { createRouteTestHarness } from './_route-test-utils.js';
vi.mock('../../src/web/webview-egress.js', async () => {
const actual = await vi.importActual<typeof import('../../src/web/webview-egress.js')>(
'../../src/web/webview-egress.js'
);
return { ...actual, webviewFetch: vi.fn() };
});
const fetchMock = vi.mocked(webviewFetch);
async function setup() {
return createRouteTestHarness(registerCustomModelRoutes);
}
describe('custom model endpoint CRUD', () => {
afterEach(() => {
fetchMock.mockReset();
});
it('starts empty', async () => {
const { app } = await setup();
const res = await app.inject({ method: 'GET', url: '/api/model-endpoints' });
expect(res.json()).toEqual([]);
});
it('creates, lists, updates, and deletes an endpoint', async () => {
const { app } = await setup();
const create = await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'ep1', label: 'llama.cpp box', baseUrl: 'http://192.168.1.50:8080' },
});
expect(create.statusCode).toBe(200);
expect(create.json().data.host.id).toBe('ep1');
const list = await app.inject({ method: 'GET', url: '/api/model-endpoints' });
expect(list.json()).toHaveLength(1);
const update = await app.inject({
method: 'PUT',
url: '/api/model-endpoints/ep1',
payload: { label: 'Renamed', baseUrl: 'http://192.168.1.50:8080' },
});
expect(update.statusCode).toBe(200);
expect(update.json().data.host.label).toBe('Renamed');
const del = await app.inject({ method: 'DELETE', url: '/api/model-endpoints/ep1' });
expect(del.statusCode).toBe(200);
const listAfter = await app.inject({ method: 'GET', url: '/api/model-endpoints' });
expect(listAfter.json()).toEqual([]);
});
it('rejects a duplicate id on create', async () => {
const { app } = await setup();
const payload = { id: 'dup', label: 'A', baseUrl: 'http://localhost:8080' };
await app.inject({ method: 'POST', url: '/api/model-endpoints', payload });
const second = await app.inject({ method: 'POST', url: '/api/model-endpoints', payload });
expect(second.json().success).toBe(false);
expect(second.json().errorCode).toBe('ALREADY_EXISTS');
});
it('404s updating/deleting an id that does not exist', async () => {
const { app } = await setup();
const update = await app.inject({
method: 'PUT',
url: '/api/model-endpoints/ghost',
payload: { label: 'A', baseUrl: 'http://localhost:8080' },
});
expect(update.json().errorCode).toBe('NOT_FOUND');
});
it('rejects a link-local/cloud-metadata base URL', async () => {
const { app } = await setup();
const res = await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'meta', label: 'A', baseUrl: 'http://169.254.169.254/' },
});
expect(res.json().success).toBe(false);
expect(res.json().errorCode).toBe('INVALID_INPUT');
});
it('discovers models via GET /v1/models and stores the result', async () => {
const { app } = await setup();
await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'ep1', label: 'A', baseUrl: 'http://localhost:8080', apiKey: 'k' },
});
fetchMock.mockImplementation(async (url: URL, init?: RequestInit) => {
expect(url.href).toBe('http://localhost:8080/v1/models');
const headers = init?.headers as Record<string, string>;
// Exactly ONE auth header — never both (a real server hung when sent both).
expect(headers.Authorization).toBe('Bearer k');
expect(headers['api-key']).toBeUndefined();
return new Response(JSON.stringify({ data: [{ id: 'qwen3' }, { id: 'llama3' }] }), { status: 200 });
});
const res = await app.inject({ method: 'POST', url: '/api/model-endpoints/ep1/discover-models' });
expect(res.json().data.models).toEqual(['qwen3', 'llama3']);
expect(fetchMock).toHaveBeenCalledTimes(1);
// Data dir is shared across this WHOLE test file (one temp HOME per file, not per
// test — test/setup.ts), so find by id rather than assuming index 0.
const list = await app.inject({ method: 'GET', url: '/api/model-endpoints' });
const stored = (list.json() as Array<{ id: string }>).find((h) => h.id === 'ep1');
expect(stored?.models).toEqual(['qwen3', 'llama3']);
expect(stored?.lastDiscoveredAt).toBeTruthy();
});
it('discovers models with authStyle "api-key" using only that header, never Authorization', async () => {
const { app } = await setup();
await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'ep-azure', label: 'A', baseUrl: 'http://localhost:8080', apiKey: 'k', authStyle: 'api-key' },
});
fetchMock.mockImplementation(async (_url: URL, init?: RequestInit) => {
const headers = init?.headers as Record<string, string>;
expect(headers['api-key']).toBe('k');
expect(headers.Authorization).toBeUndefined();
return new Response(JSON.stringify({ data: [] }), { status: 200 });
});
await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-azure/discover-models' });
expect(fetchMock).toHaveBeenCalledTimes(1);
});
it('reports a clear error when the endpoint is unreachable', async () => {
const { app } = await setup();
await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'ep-err', label: 'A', baseUrl: 'http://localhost:8080' },
});
// undici's shape: a bare `fetch failed` with the real reason one level down.
fetchMock.mockRejectedValue(
new TypeError('fetch failed', { cause: new Error('connect ECONNREFUSED 127.0.0.1:8080') })
);
const res = await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-err/discover-models' });
expect(res.json().success).toBe(false);
expect(res.json().errorCode).toBe('OPERATION_FAILED');
expect(res.json().error).toContain('ECONNREFUSED');
});
it('names the egress refusal when the endpoint resolves into a blocked range', async () => {
const { app } = await setup();
await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'ep-meta', label: 'A', baseUrl: 'http://models.example:8080' },
});
const { WebviewEgressBlockedError } = await vi.importActual<typeof import('../../src/web/webview-egress.js')>(
'../../src/web/webview-egress.js'
);
fetchMock.mockRejectedValue(
new TypeError('fetch failed', { cause: new WebviewEgressBlockedError('resolves to 169.254.169.254') })
);
const res = await app.inject({ method: 'POST', url: '/api/model-endpoints/ep-meta/discover-models' });
expect(res.json().success).toBe(false);
expect(res.json().error).toMatch(/refused.*169\.254\.169\.254/);
});
it('refuses a baseUrl with embedded credentials or a non-http scheme at save time', async () => {
const { app } = await setup();
for (const baseUrl of ['http://user:pw@host:8080', 'ftp://host/models', 'http://169.254.169.254']) {
const res = await app.inject({
method: 'POST',
url: '/api/model-endpoints',
payload: { id: 'bad', label: 'A', baseUrl },
});
expect(res.json().success, baseUrl).toBe(false);
expect(res.json().errorCode, baseUrl).toBe('INVALID_INPUT');
}
});
});
+736
View File
@@ -0,0 +1,736 @@
/**
* @fileoverview Route tests for file READ routes in a remote (SSH) case (#415).
*
* The mirror image of `test/routes/file-routes.test.ts`: every request here resolves
* against a path that exists only on another host, so the local `fs` layer must never
* be the thing that answers. The ssh layer (`src/remote-files.ts`) is mocked — a test
* never opens a connection — but the REAL module is kept alongside the mocks so
* `RemoteFileAccessError` and the command builders stay authentic.
*
* Port: N/A (app.inject doesn't open ports)
*/
import { describe, it, expect, beforeEach, afterEach, vi } from 'vitest';
import { Readable } from 'node:stream';
import { createRouteTestHarness, type RouteTestHarness } from './_route-test-utils.js';
import { registerFileRoutes } from '../../src/web/routes/file-routes.js';
import { attachmentRegistry } from '../../src/attachment-registry.js';
import { RemoteFileAccessError } from '../../src/remote-files.js';
import type { RemoteProbe } from '../../src/remote-files.js';
import type { SessionRemote } from '../../src/types/session.js';
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
import { MAX_FILE_DOWNLOAD_BYTES } from '../../src/config/buffer-limits.js';
// Keep the pure builders + the error class real; replace only the IO.
vi.mock('../../src/remote-files.js', async (importOriginal) => {
const actual = await importOriginal<typeof import('../../src/remote-files.js')>();
return {
...actual,
remoteProbePaths: vi.fn(),
remoteReadFile: vi.fn(),
remoteCreateReadStream: vi.fn(),
};
});
import { remoteProbePaths, remoteReadFile, remoteCreateReadStream } from '../../src/remote-files.js';
const mockedProbePaths = vi.mocked(remoteProbePaths);
const mockedReadFile = vi.mocked(remoteReadFile);
const mockedCreateReadStream = vi.mocked(remoteCreateReadStream);
const REMOTE_DIR = '/srv/remote/case';
const remote: SessionRemote = {
hostId: 'host-1',
label: 'testhost',
host: '192.0.2.10',
username: 'j',
remotePath: REMOTE_DIR,
};
function fileProbe(realPath: string, size: number): RemoteProbe {
return { realPath, kind: 'file', size, mtimeMs: 1_700_000_000_000 };
}
const dirProbe: RemoteProbe = { realPath: REMOTE_DIR, kind: 'directory', size: 0, mtimeMs: 0 };
describe('file routes in a remote (SSH) case', () => {
let harness: RouteTestHarness;
let sessionId: string;
let closeSpy: ReturnType<typeof vi.fn>;
beforeEach(async () => {
harness = await createRouteTestHarness(registerFileRoutes);
sessionId = harness.ctx._sessionId;
harness.ctx._session.attachmentHistory = [];
// The whole point of the fixture: the workspace is a path on ANOTHER host.
harness.ctx._session.workingDir = REMOTE_DIR;
harness.ctx._session.remote = { ...remote };
closeSpy = vi.fn();
mockedProbePaths.mockResolvedValue([fileProbe(`${REMOTE_DIR}/img.png`, 9), dirProbe]);
mockedReadFile.mockResolvedValue(Buffer.from('remote text'));
mockedCreateReadStream.mockReturnValue({
stream: Readable.from([Buffer.from('remote bytes')]),
close: closeSpy,
} as never);
});
afterEach(() => {
// The registry is process-global: a record left behind would leak into the next
// test's by-id requests.
attachmentRegistry.clearSession(sessionId);
vi.clearAllMocks();
});
describe('GET /api/sessions/:id/file-raw', () => {
it('streams the remote file and probes the path AND the workspace in one call', async () => {
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-raw?path=img.png`,
});
expect(res.statusCode).toBe(200);
expect(res.headers['content-type']).toBe('image/png');
expect(res.body).toBe('remote bytes');
// Both paths in one ssh round trip: the workspace root is needed to check
// containment against a REMOTELY canonicalized root.
expect(mockedProbePaths).toHaveBeenCalledWith(expect.objectContaining({ host: '192.0.2.10' }), [
`${REMOTE_DIR}/img.png`,
REMOTE_DIR,
]);
expect(mockedCreateReadStream).toHaveBeenCalledWith(
expect.objectContaining({ host: '192.0.2.10' }),
`${REMOTE_DIR}/img.png`,
undefined
);
});
it('serves a byte range as a 206 from the remote host', async () => {
mockedProbePaths.mockResolvedValue([fileProbe(`${REMOTE_DIR}/clip.mp4`, 100), dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-raw?path=clip.mp4`,
headers: { range: 'bytes=10-19' },
});
expect(res.statusCode).toBe(206);
expect(res.headers['content-range']).toBe('bytes 10-19/100');
expect(res.headers['content-length']).toBe('10');
expect(res.headers['accept-ranges']).toBe('bytes');
expect(mockedCreateReadStream).toHaveBeenCalledWith(expect.anything(), `${REMOTE_DIR}/clip.mp4`, {
start: 10,
end: 19,
});
});
it('reaps the ssh stream when the response is done', async () => {
await harness.app.inject({ method: 'GET', url: `/api/sessions/${sessionId}/file-raw?path=img.png` });
// The cleanup is registered on the raw response's lifecycle; without it an
// aborted download would leave the ssh child running.
expect(closeSpy).toHaveBeenCalled();
});
it('refuses a path that escapes the workspace lexically, without connecting', async () => {
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-raw?path=../../etc/shadow`,
});
expect(res.statusCode).toBe(404);
expect(mockedProbePaths).not.toHaveBeenCalled();
});
it('refuses a symlink that resolves outside the workspace on the remote host', async () => {
mockedProbePaths.mockResolvedValue([fileProbe('/etc/shadow', 10), dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-raw?path=innocent.png`,
});
expect(res.statusCode).toBe(404);
expect(mockedCreateReadStream).not.toHaveBeenCalled();
});
it('accepts a workspace reached through a remote symlink (both sides canonicalized)', async () => {
// remotePath is a symlinked mount: the file's realpath is genuinely inside the
// workspace's realpath, so refusing it would break the whole case.
harness.ctx._session.workingDir = '/mnt/link/case';
mockedProbePaths.mockResolvedValue([
fileProbe('/srv/real/case/img.png', 3),
{ realPath: '/srv/real/case', kind: 'directory', size: 0, mtimeMs: 0 },
]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-raw?path=img.png`,
});
expect(res.statusCode).toBe(200);
});
it('404s a file that does not exist on the remote host', async () => {
mockedProbePaths.mockResolvedValue([null, dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-raw?path=gone.png`,
});
expect(res.statusCode).toBe(404);
expect(mockedCreateReadStream).not.toHaveBeenCalled();
});
it('reports an unreachable host as a gateway failure, not a 404 or a 500', async () => {
mockedProbePaths.mockRejectedValue(
new RemoteFileAccessError('remote host testhost unreachable: Connection refused')
);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-raw?path=img.png`,
});
expect(res.statusCode).toBe(502);
expect(JSON.parse(res.body).error).toContain('Connection refused');
});
it('applies the size cap to the REMOTE size, before reading', async () => {
mockedProbePaths.mockResolvedValue([fileProbe(`${REMOTE_DIR}/huge.mp4`, MAX_FILE_DOWNLOAD_BYTES + 1), dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-raw?path=huge.mp4`,
});
expect(res.statusCode).toBe(413);
expect(mockedCreateReadStream).not.toHaveBeenCalled();
});
it('refuses a directory', async () => {
mockedProbePaths.mockResolvedValue([dirProbe, dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-raw?path=.`,
});
expect(res.statusCode).toBe(400);
});
it('does not touch the ssh layer for a local session', async () => {
delete harness.ctx._session.remote;
await harness.app.inject({ method: 'GET', url: `/api/sessions/${sessionId}/file-raw?path=img.png` });
expect(mockedProbePaths).not.toHaveBeenCalled();
expect(mockedCreateReadStream).not.toHaveBeenCalled();
});
describe('when a path with the same absolute name ALSO exists on this host', () => {
// The case that really happens in practice: the remote tree is mounted on the
// Codeman host at the identical absolute path (an sshfs mount, which is the
// documented stop-gap workaround for this very bug). The remote host stays the
// source of truth: there is deliberately no local fallback, because a fallback
// would silently serve the OTHER filesystem's bytes under the same path.
let shadowRoot: string;
let shadowFile: string;
beforeEach(() => {
shadowRoot = mkdtempSync(join(tmpdir(), 'codeman-remote-shadow-'));
shadowFile = join(shadowRoot, 'img.png');
writeFileSync(shadowFile, 'LOCAL BYTES');
harness.ctx._session.workingDir = shadowRoot;
harness.ctx._session.remote = { ...remote, remotePath: shadowRoot };
mockedProbePaths.mockResolvedValue([fileProbe(shadowFile, 12), { ...dirProbe, realPath: shadowRoot }]);
});
afterEach(() => {
rmSync(shadowRoot, { recursive: true, force: true });
});
it('serves the REMOTE bytes, never the local copy at the same path', async () => {
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-raw?path=img.png`,
});
expect(res.statusCode).toBe(200);
expect(res.body).toBe('remote bytes');
expect(res.body).not.toBe('LOCAL BYTES');
// The local file is untouched, proving the local side was never the source.
expect(readFileSync(shadowFile, 'utf8')).toBe('LOCAL BYTES');
});
it('still 404s when the remote host does not have the file, even though a local one exists', async () => {
mockedProbePaths.mockResolvedValue([null, { ...dirProbe, realPath: shadowRoot }]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-raw?path=img.png`,
});
expect(res.statusCode).toBe(404);
expect(mockedCreateReadStream).not.toHaveBeenCalled();
});
it('reads text from the remote host, not from the local twin', async () => {
const localText = join(shadowRoot, 'notes.txt');
writeFileSync(localText, 'local text');
mockedProbePaths.mockResolvedValue([fileProbe(localText, 11), { ...dirProbe, realPath: shadowRoot }]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-content?path=notes.txt`,
});
expect(JSON.parse(res.body).data.content).toBe('remote text');
expect(mockedReadFile).toHaveBeenCalledWith(expect.anything(), localText, expect.any(Number));
expect(readFileSync(localText, 'utf8')).toBe('local text');
});
});
});
describe('GET /api/sessions/:id/file-content', () => {
it('returns remote text content and never advertises the editor', async () => {
mockedProbePaths.mockResolvedValue([fileProbe(`${REMOTE_DIR}/notes.txt`, 11), dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-content?path=notes.txt`,
});
expect(res.statusCode).toBe(200);
const body = JSON.parse(res.body);
expect(body.success).toBe(true);
expect(body.data.content).toBe('remote text');
expect(body.data.editable).toBe(false);
expect(mockedReadFile).toHaveBeenCalledWith(expect.anything(), `${REMOTE_DIR}/notes.txt`, expect.any(Number));
});
it('classifies remote media by extension without reading it', async () => {
mockedProbePaths.mockResolvedValue([fileProbe(`${REMOTE_DIR}/logo.png`, 1024), dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-content?path=logo.png`,
});
const body = JSON.parse(res.body);
expect(body.data.type).toBe('image');
expect(body.data.url).toContain('file-raw');
expect(mockedReadFile).not.toHaveBeenCalled();
});
it('turns an edit request into an explicit 400 instead of a misleading 404', async () => {
mockedProbePaths.mockResolvedValue([fileProbe(`${REMOTE_DIR}/notes.txt`, 11), dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-content?path=notes.txt&edit=1`,
});
expect(res.statusCode).toBe(400);
const body = JSON.parse(res.body);
expect(body.success).toBe(false);
expect(body.error).toContain('not supported for files in a remote');
});
it('reports an unreachable host as a real 502 with the remote reason', async () => {
mockedProbePaths.mockRejectedValue(new RemoteFileAccessError('remote host testhost unreachable: timed out'));
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-content?path=notes.txt`,
});
expect(res.statusCode).toBe(502);
const body = JSON.parse(res.body);
expect(body.success).toBe(false);
expect(body.error).toContain('timed out');
});
it('rejects a path that escapes the workspace', async () => {
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-content?path=../../../etc/passwd`,
});
expect(JSON.parse(res.body).success).toBe(false);
expect(mockedProbePaths).not.toHaveBeenCalled();
});
});
describe('GET /api/sessions/:id/file-preview and file-thumbnail', () => {
it('redirects a non-office remote file to file-raw', async () => {
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-preview?path=scan.pdf`,
});
expect(res.statusCode).toBe(302);
expect(res.headers.location).toContain('/file-raw');
});
it('says office previews are unavailable rather than 404-ing', async () => {
mockedProbePaths.mockResolvedValue([fileProbe(`${REMOTE_DIR}/doc.docx`, 10), dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-preview?path=doc.docx`,
});
expect(res.statusCode).toBe(400);
expect(JSON.parse(res.body).error).toContain('not available for files in a remote');
});
it('says thumbnails are unavailable for a remote file', async () => {
mockedProbePaths.mockResolvedValue([fileProbe(`${REMOTE_DIR}/doc.pdf`, 10), dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-thumbnail?path=doc.pdf`,
});
expect(res.statusCode).toBe(400);
expect(JSON.parse(res.body).error).toContain('not available for files in a remote');
});
it('reports an unreachable host for previews too', async () => {
mockedProbePaths.mockRejectedValue(new RemoteFileAccessError('remote host testhost unreachable: no route'));
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-preview?path=doc.docx`,
});
expect(res.statusCode).toBe(502);
});
});
describe('attachments — a file click OUTSIDE the case directory (#415)', () => {
const outsidePath = '/tmp/agent-output/shot.png';
async function publish(path: string): Promise<{ statusCode: number; body: unknown }> {
const res = await harness.app.inject({
method: 'POST',
url: `/api/sessions/${sessionId}/attachments`,
payload: { path, notify: false },
});
return { statusCode: res.statusCode, body: JSON.parse(res.body) };
}
it('registers an out-of-workspace remote path by probing the remote host', async () => {
mockedProbePaths.mockResolvedValue([fileProbe(outsidePath, 42), dirProbe]);
const res = await publish(outsidePath);
expect(res.statusCode).toBe(200);
const data = (res.body as { data: { attachmentId: string; size: number; fileName: string } }).data;
expect(data.attachmentId).toMatch(/^att_/);
expect(data.fileName).toBe('shot.png');
expect(data.size).toBe(42);
// The path is outside the workspace, so a workspace-relative resolution could
// never have found it — the probe is what makes this work at all.
expect(mockedProbePaths).toHaveBeenCalledWith(expect.objectContaining({ host: '192.0.2.10' }), [
outsidePath,
REMOTE_DIR,
]);
});
it("serves the registered remote attachment's bytes by id, with range support", async () => {
mockedProbePaths.mockResolvedValue([fileProbe(outsidePath, 42), dirProbe]);
const published = await publish(outsidePath);
const attachmentId = (published.body as { data: { attachmentId: string } }).data.attachmentId;
// The by-id route re-probes (guard defense-in-depth) before streaming.
mockedProbePaths.mockResolvedValue([fileProbe(outsidePath, 42), dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/attachments/${attachmentId}/raw`,
});
expect(res.statusCode).toBe(200);
expect(res.headers['content-type']).toBe('image/png');
expect(res.body).toBe('remote bytes');
expect(mockedCreateReadStream).toHaveBeenCalledWith(expect.anything(), outsidePath, undefined);
const ranged = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/attachments/${attachmentId}/raw`,
headers: { range: 'bytes=1-3' },
});
expect(ranged.statusCode).toBe(206);
expect(ranged.headers['content-range']).toBe('bytes 1-3/42');
});
it('reports the remote size in the attachment metadata poll', async () => {
mockedProbePaths.mockResolvedValue([fileProbe(outsidePath, 42), dirProbe]);
const published = await publish(outsidePath);
const attachmentId = (published.body as { data: { attachmentId: string } }).data.attachmentId;
mockedProbePaths.mockResolvedValue([fileProbe(outsidePath, 84), dirProbe]);
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/attachments/${attachmentId}`,
});
expect(res.statusCode).toBe(200);
expect(JSON.parse(res.body).data.size).toBe(84);
});
it('404s a remote path that does not exist instead of reporting it as unreadable', async () => {
mockedProbePaths.mockResolvedValue([null, dirProbe]);
const res = await publish('/tmp/agent-output/gone.png');
expect(res.statusCode).toBe(404);
expect(JSON.stringify(res.body)).toContain('Attachment file not found');
});
it('reports an unreachable host as 502 for the click path too', async () => {
mockedProbePaths.mockRejectedValue(new RemoteFileAccessError('remote host testhost unreachable: timed out'));
const res = await publish(outsidePath);
expect(res.statusCode).toBe(502);
});
it('still refuses a blocked remote path (the blocklist is host-agnostic)', async () => {
mockedProbePaths.mockResolvedValue([fileProbe('/etc/shadow', 10), dirProbe]);
const res = await publish('/etc/shadow');
// 403 from the guard (the same answer the local path gives for a blocked tree).
expect(res.statusCode).toBe(403);
expect(JSON.stringify(res.body)).toMatch(/blocked/i);
});
it('does not offer office previews or thumbnails for a remote attachment', async () => {
mockedProbePaths.mockResolvedValue([fileProbe('/tmp/agent-output/report.docx', 10), dirProbe]);
const published = await publish('/tmp/agent-output/report.docx');
const attachmentId = (published.body as { data: { attachmentId: string } }).data.attachmentId;
mockedProbePaths.mockResolvedValue([fileProbe('/tmp/agent-output/report.docx', 10), dirProbe]);
const preview = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/attachments/${attachmentId}/preview`,
});
const thumbnail = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/attachments/${attachmentId}/thumbnail`,
});
expect(preview.statusCode).toBe(400);
expect(thumbnail.statusCode).toBe(400);
});
it('lists an out-of-workspace remote history entry without marking it missing', async () => {
harness.ctx._session.attachmentHistory = [
{
id: 'hist-1',
sessionId,
fileName: 'shot.png',
extension: 'png',
attachmentType: 'image',
size: 1,
mtimeMs: 1,
timestamp: 1,
source: 'external',
externalPath: outsidePath,
},
];
mockedProbePaths.mockImplementation(async (_remote, paths) =>
paths.map((path) => (path === outsidePath ? fileProbe(outsidePath, 42) : path === REMOTE_DIR ? dirProbe : null))
);
const res = await harness.app.inject({ method: 'GET', url: `/api/sessions/${sessionId}/attachments` });
expect(res.statusCode).toBe(200);
const [item] = JSON.parse(res.body).data.items;
expect(item.missing).toBe(false);
expect(item.size).toBe(42);
expect(item.attachmentId).toBeTruthy();
});
it('resolves a workspace-relative history entry over ssh', async () => {
harness.ctx._session.attachmentHistory = [
{
id: 'hist-2',
sessionId,
fileName: 'out.png',
extension: 'png',
attachmentType: 'image',
size: 1,
mtimeMs: 1,
timestamp: 1,
source: 'detected',
relativePath: 'out.png',
},
];
mockedProbePaths.mockImplementation(async (_remote, paths) =>
paths.map((path) =>
path === `${REMOTE_DIR}/out.png`
? fileProbe(`${REMOTE_DIR}/out.png`, 7)
: path === REMOTE_DIR
? dirProbe
: null
)
);
const res = await harness.app.inject({ method: 'GET', url: `/api/sessions/${sessionId}/attachments` });
const [item] = JSON.parse(res.body).data.items;
expect(item.missing).toBe(false);
expect(item.size).toBe(7);
expect(item.rawUrl).toContain('file-raw');
});
describe('the history list probes the whole history in ONE ssh round trip', () => {
// One connection per entry (up to ATTACHMENT_HISTORY_LIMIT, re-run on every
// attachment:detected while the drawer is open) tripped OpenSSH's default
// MaxStartups 10:30:100, which drops most of a burst that size.
const history = () => [
{
id: 'hist-a',
sessionId,
fileName: 'out.png',
extension: 'png',
attachmentType: 'image' as const,
size: 1,
mtimeMs: 1,
timestamp: 1,
source: 'detected' as const,
relativePath: 'out.png',
},
{
id: 'hist-b',
sessionId,
fileName: 'shot.png',
extension: 'png',
attachmentType: 'image' as const,
size: 1,
mtimeMs: 1,
timestamp: 1,
source: 'external' as const,
externalPath: outsidePath,
},
{
id: 'hist-c',
sessionId,
fileName: 'gone.png',
extension: 'png',
attachmentType: 'image' as const,
size: 1,
mtimeMs: 1,
timestamp: 1,
source: 'detected' as const,
relativePath: 'gone.png',
},
];
it('issues a single batched probe covering every entry plus the workspace root', async () => {
harness.ctx._session.attachmentHistory = history();
mockedProbePaths.mockImplementation(async (_remote, paths) =>
paths.map((path) =>
path === `${REMOTE_DIR}/out.png`
? fileProbe(path, 7)
: path === outsidePath
? fileProbe(outsidePath, 42)
: path === REMOTE_DIR
? dirProbe
: null
)
);
const res = await harness.app.inject({ method: 'GET', url: `/api/sessions/${sessionId}/attachments` });
expect(res.statusCode).toBe(200);
expect(mockedProbePaths).toHaveBeenCalledTimes(1);
const [, probed] = mockedProbePaths.mock.calls[0];
expect([...probed].sort()).toEqual(
[REMOTE_DIR, `${REMOTE_DIR}/gone.png`, `${REMOTE_DIR}/out.png`, outsidePath].sort()
);
// (ids are re-minted for external entries by the sanitizer, so key on the name)
const items = JSON.parse(res.body).data.items as Array<{ fileName: string; missing: boolean; size: number }>;
expect(items.map((item) => [item.fileName, item.missing, item.size])).toEqual([
['out.png', false, 7],
['shot.png', false, 42],
['gone.png', true, 1],
]);
});
it('reports every entry as unknown (missing: false), detected AND external alike, when the host is unreachable', async () => {
harness.ctx._session.attachmentHistory = history();
mockedProbePaths.mockRejectedValue(new RemoteFileAccessError('remote host testhost unreachable: timed out'));
const res = await harness.app.inject({ method: 'GET', url: `/api/sessions/${sessionId}/attachments` });
expect(res.statusCode).toBe(200);
const items = JSON.parse(res.body).data.items as Array<{ id: string; missing: boolean }>;
// The two branches used to disagree here: detected kept missing:false while
// external's 502 was folded into missing:true.
expect(items.map((item) => item.missing)).toEqual([false, false, false]);
});
});
});
describe('PUT /api/sessions/:id/file-content', () => {
// The remote guard has to come BEFORE the local path validation: with a
// directory of the same absolute name on this host (an sshfs mount of the remote
// tree, the documented stop-gap for #415) the write would land on the local twin
// while the viewer believes it edited the remote file.
let shadowRoot: string;
let shadowFile: string;
beforeEach(() => {
shadowRoot = mkdtempSync(join(tmpdir(), 'codeman-remote-put-'));
shadowFile = join(shadowRoot, 'notes.txt');
writeFileSync(shadowFile, 'LOCAL TEXT');
harness.ctx._session.workingDir = shadowRoot;
harness.ctx._session.remote = { ...remote, remotePath: shadowRoot };
});
afterEach(() => {
rmSync(shadowRoot, { recursive: true, force: true });
});
it('answers 400 for a remote case and never touches the local file of the same name', async () => {
const res = await harness.app.inject({
method: 'PUT',
url: `/api/sessions/${sessionId}/file-content`,
payload: { path: 'notes.txt', content: 'OVERWRITTEN', baseHash: 'whatever', force: true },
});
expect(res.statusCode).toBe(400);
expect(JSON.parse(res.body).error).toMatch(/not supported for files in a remote/);
expect(readFileSync(shadowFile, 'utf8')).toBe('LOCAL TEXT');
expect(mockedProbePaths).not.toHaveBeenCalled();
});
});
describe('a client that gives up during the guard probe', () => {
it('still has its ssh body child reaped', async () => {
// The probe is an ssh round trip; a client that aborted during it has already
// closed the response, so a `close` listener attached afterwards never fires.
const controller = new AbortController();
mockedProbePaths.mockImplementation(async () => {
controller.abort();
await new Promise((resolveDelay) => setTimeout(resolveDelay, 20));
return [fileProbe(`${REMOTE_DIR}/img.png`, 9), dirProbe];
});
await harness.app
.inject({ method: 'GET', url: `/api/sessions/${sessionId}/file-raw?path=img.png`, signal: controller.signal })
.catch(() => undefined);
await new Promise((resolveDelay) => setTimeout(resolveDelay, 50));
// The body WAS opened (the route ran to completion against an already-closed
// response), which is exactly the window the guard covers.
expect(mockedCreateReadStream).toHaveBeenCalledTimes(1);
expect(closeSpy).toHaveBeenCalled();
});
});
});

Some files were not shown because too many files have changed in this diff Show More