flushPendingWrites() captured the viewport of a user who was reading
scrollback, called terminal.write(), and restored the anchor on the next
line. xterm parses on its own schedule, so at that point the buffer has not
moved: the guard `viewportY !== preserveViewportY` was false, scrollToLine
was never called at all, and the Codex redraw landed a tick later and took
the viewport to the live bottom with nothing left to pull it back. Scrolling
up during a stream still got dragged down, which is what #358 reports, and a
refresh was the only way back to a coherent view.
The restore moves inside xterm's write callback, the first moment the
redraw's effect exists, and runs before _scheduleTerminalWriteFlush() so a
deferred remainder re-captures the restored anchor rather than the bottom.
Two things follow from it running later:
- A live anchor now wins over the sticky scroll-to-bottom. The two are
captured at different moments (_wasAtBottomBeforeWrite at the frame's
first batchTerminalWrite, the anchor at flush time), so a scroll-up in
between leaves both set, and running both would jump to the bottom and
come back a frame later instead of staying put.
- The anchor is dropped if the active session changed or a buffer load
started while the write was in flight. It indexes the buffer it was
captured from, and selectSession() resets the terminal and chunk-loads a
different scrollback.
The existing regression passed throughout, because its write mock moved the
viewport synchronously, which real xterm never does. The harness now models
an asynchronous parse (redraw lands, then the callback fires), and all five
of the anchor tests fail against the old code.
Fixes#358
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three behaviours landed from #375 without their doc entries: the
duplicate input ACK now carries `dup:true` and the server's watermark
(`docs/reliable-input-delivery.md` still described a bare ACK), Shift+drag
and right-click copy in the terminal (the shortcut list did not know
them), and one adopted container backing several cases at different
in-container directories (the Docker cases paragraph still implied one
case per container for adopted containers too).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The cherry-picked "copy an existing case" commit declared a second
`CaseInfo.docker.owned` and emitted `owned: true|false` on every docker
case, while master had meanwhile shipped the same field from the
adopted-container work with a narrower wire shape: `owned` is present
only when false, absent means owned. Two declarations failed typecheck,
and two emit styles on one response would have made the picker's answer
depend on which read path filled it.
Keep master's shape at both response sites (the case list and the
single-case lookup, which lacked the field entirely), fold the picker's
reason for the field into the existing doc comment, and repoint the test
that pinned "set on exactly two sites" at the surviving form, adding a
negative pin so the duplicate style cannot come back.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The previous version cleared the case name and the in-container directory on the
grounds that they must differ. That left a form with three fields mysteriously
filled and two empty, and turned the most common operation — changing
/srv/app/api to /srv/app/web — into retyping a long path.
Both are now pre-filled, with focus on the in-container directory and the caret
at the end, since the tail is what changes. What stops an unmodified submit is no
longer an empty field but a guard: the values applied are recorded, compared at
submit time, and if nothing changed the reason is stated next to the field and
focus moves to it, without sending a request that is certain to be refused.
The server refuses these anyway (a duplicate case name, a twin case on the same
container and directory) and its errors are clear; but making a round trip to be
told "you forgot to edit the field you are looking at" is worse than saying so on
the spot. The guard only applies when a source case was actually selected, so
filling the adopt form from scratch is unaffected.
⚠️ The status text is written into dockerLinkStatus. My first version referenced
an id that does not exist (dockerAdoptStatus), which made the explanation vanish
silently and left only a toast. The test now extracts that id from the code and
looks it up in index.html, pinning that it must really exist.
(cherry picked from commit ba21ae11f4)
The backend already lets one adopted container back several cases pointing at
different in-container directories, but using it meant retyping the container
name, host and workspace one by one — exactly the friction that leaves a
capability unused. Picking an existing case from a dropdown now carries those
three over, leaving only the two fields that must differ: the case name and the
in-container directory.
Clearing those two is the point of the feature, not a convenience: keeping the
old name is refused by the server as "case already exists", and keeping the old
directory is refused as "a twin case on the same container and directory". Both
errors are clear, but a form pre-filled with values that are guaranteed to be
rejected is a trap. Focus lands on the in-container directory — the thing the
user came here to change.
⚠️ Only adopted containers are listed (docker.owned === false). A Codeman-built
container's lifecycle belongs to its one case — a second case would be torn out
by that case's recreate or delete — so the server refuses it anyway, and listing
it here would only manufacture a baffling error. `owned` may be absent and absent
means owned, so the test is `!== false`, not truthiness.
CaseInfo.docker gains containerWorkdir and owned for this: the former is the
"which directory does this case use" half of the picker, without which the user
cannot tell what to change it to; the latter backs the filter above. ⚠️ Both
places that build a docker CaseInfo (the list endpoint and the single-case query)
must set them — filling in only one makes the picker work or not depending on
which read path was taken, and a test pins "exactly two".
(cherry picked from commit f1ed3a58e1)
Once a container is adopted, it could not be adopted a second time. But a
container usually holds more than one project directory, and opening a case for
another one had no path forward except starting a second container — precisely
what adoption exists to avoid.
The original reason was in a comment: two cases sharing an adopted container
would make one case's teardown race the other's launch on the same tmux server.
That reason does not hold. The in-container tmux session name is
dockerTmuxSessionName(sessionId), i.e. codeman-dkr-<id8>, keyed by SESSION and
not by case, and buildDockerKillCommand tears down exactly that name, so killing
A never touches B — hosting multiple sessions is what a tmux server is for.
The other three routes into an adopted container's lifecycle do not pass through
here either, confirmed one by one: the stop and remove builders throw outright;
recreate refuses `owned === false` before it even resolves the container name;
and orphan reaping filters on `label=codeman.managed=1`, which a user-built
container does not carry — a structural exclusion.
That leaves exactly three cases worth refusing, none of them tmux-related, split
into the pure, unit-tested classifyAdoptContainerConflict:
- owned-case the container belongs to a Codeman-created case, whose lifecycle
Codeman manages: one recreate or delete there would pull the
container out from under the adopting case.
⚠️ `owned` may be absent and absent means owned (cases predate
the field), so the test is `!== false`, not truthiness.
- other-owner already adopted by a different user. Adoption hands out a shell
inside someone else's container.
- duplicate same container, same directory. The second case would behave
identically to the first, so name the existing one rather than
silently minting a twin. A different in-container directory is
the case this change exists to support and passes.
(cherry picked from commit 1cb6bde891)
`backdrop-filter` promotes an element to its own compositing layer. A
position:fixed full-screen layer that is created and then hidden was measured to
leave a stale hit-test region behind in Chrome: the page renders perfectly, but
pointer events across the viewport go nowhere.
The report came from a long-lived tab connected to a remote server, where a
connection blip shows and then hides #offlineOverlay. The symptoms were a
terminal that would not scroll and, at the same time, an unrelated
click-to-expand that also stopped responding, while a freshly opened tab was
fine; a read-only console command (getComputedStyle + elementFromPoint, both of
which force a hit-test recomputation) then cured it. Two unrelated features
dying together and one read-only command fixing both points at hit-testing
itself rather than at either feature.
So the `backdrop-filter` moves onto the actually-visible selector and the layer
is never created while hidden. Only the two persistent overlays change:
offline-overlay (toggled with [hidden]) and file-preview-overlay (toggled with
.visible). path-picker and path-preview are created and removed by JS, leave
nothing behind, and are untouched.
⚠️ This is an evidence-based inference, not a fix verified by reproduction:
reproducing it needs a long-lived page that has been through a connection blip,
which I could not manufacture in a controlled environment. The guard test pins
both halves — no such property while hidden, and a real blur while shown — so a
later cleanup cannot quietly delete the effect.
(cherry picked from commit 08442dfee1)
handleInit() did not distinguish a first load from an SSE reconnect: it always
cleared the terminal caches in _resetAllAppState() and re-ran selectSession() for
the session that was already on screen. Every reconnect therefore refetched up to
1 MiB of buffer and reset+rewrote xterm. On a link that drops a connection about
once a minute (measured at ~57s intervals against a healthy server) that reads as
the page refreshing itself and throwing away your reading position.
A reconnect that lands back on the still-open session now keeps the terminal
caches and activeSessionId and resyncs through _onSessionNeedsRefresh(). That
path still reloads the buffer, so output produced during the outage is not lost,
but it preserves distance-from-bottom — the same rule #259 established for a
refresh the server triggered rather than the user. The WS is reconnected
explicitly when it is not already on that session, since skipping selectSession()
skips its _connectWs() call.
First load (gen === 1) takes exactly the path it took before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rv24Pk4qzrsDYdVyDyJQmT
(cherry picked from commit 435569c76e)
Browser input is delivered exactly once by (clientId, seq). The server records a
watermark per clientId and discards anything not above it as a duplicate — but
acknowledged it with an ACK indistinguishable from "applied". The client then
dropped the record from its queue, the UI looked perfectly normal, and the
terminal received nothing at all.
The counter is persisted to localStorage through a debounced write. Kill the page
between "sent" and "persisted" and the restored counter is below the server's
watermark, after which every keystroke lands under it, is discarded, and is
ACKed. Reloading does not help: the clientId is restored from localStorage
alongside that stale counter. Measured on a real session — typing into the same
session from a fresh browser (new clientId, no watermark on the server) worked
perfectly, which is what localised the fault to client state.
Three changes:
- on rejection the server replies {"t":"ia",seq,"dup":true,"last":<watermark>}.
It still ACKs, so the client can drop the record from its queue, but it now
says the input was not applied and supplies the number needed to climb out.
- on `dup` the client lifts its counter above the watermark and re-queues.
⚠️ Only records whose FIRST delivery is being retried are re-sent: a retry
judged duplicate means the mechanism is working (the original did arrive), and
re-sending would type the same text twice.
- the counter is now persisted synchronously. The queue payload can stay
debounced, but the counter is the thing that has to survive a crash, and
leaving it on the lossiest path cancels the only guarantee there is.
⚠️ Reading the watermark is defensive: the session arrives through a structured
port, and a port missing that method must not take the whole input path down —
a throw inside the handler means the ACK is never sent and the record is stuck in
the client queue forever, which is worse than the ambiguity being fixed. A mock
port's test timeout is what exposed this.
(cherry picked from commit 05bb7081cc)
claude advertises "Jump to bottom (ctrl+End)", so that chord has to actually
reach it. But PASSTHROUGH_KEYS carried only the bare forms (End -> \x1b[F) and
CTRL_KEYS held just six letters (c/d/l/z/a/e), which cannot express End. Ctrl+End
therefore failed in both directions:
- with an empty composer it went out as a bare \x1b[F, the modifier silently
dropped, so the CLI received a plain End;
- with text in the composer the forwarding branch requires empty, so nothing was
forwarded and the browser default applied — the caret jumped to the end of the
draft, which is the "the shortcut now edits my input box" the user saw.
Encode them as CSI 1;<mod><final> instead, and forward Ctrl/Alt-modified
navigation keys whether or not the composer is empty: they are commands for the
CLI, and the composer has no editing semantics for them worth preserving (bare
Home/End still use the old table and edit locally).
⚠️ Bare Shift is deliberately excluded: Shift+arrow selects text in the composer,
a real editing gesture that must stay local. Shift held together with Ctrl/Alt is
still encoded into the modifier mask.
(cherry picked from commit 3fbaadadfb)
In a native terminal running a TUI with mouse tracking on (claude, codex), Shift
is the "let me select text" modifier: it bypasses the application's mouse
reporting so the emulator selects locally. Users bring that habit here, where it
did nothing — measured, `hasSelection` was already false during a Shift+drag and
no clearSelection call ran at all, because there was never a selection to clear.
The mismatch is that the two Shifts mean different things. xterm reads Shift as
"force selection", but that path is only taken when the application really has
mouse tracking on. The server strips the mouse DECSETs for claude/codex/gemini
(isAltScreenStripMode), so xterm's mouseTrackingMode is permanently `none`, that
branch is unreachable, and Shift instead lands in _onIncrementalClick — which
EXTENDS an existing selection. Extension is a no-op while selectionStart is
empty, so the drag had no anchor.
So plant the anchor xterm is missing. The listener sits on the capture phase of
the `.xterm` root, an ancestor of the `.xterm-screen` that SelectionService binds
to, and therefore runs before xterm's own mousedown; xterm then extends from our
anchor and the drag behaves like any other. Length is 0 so a Shift+click without
a drag does not select a stray character. An existing selection is left alone —
that is a genuine extend gesture, and xterm handles it correctly.
Right-click copies the selection (the mintty/PuTTY convention), completing the
gesture: until now there was nowhere for a finished selection to go. With no
selection the native menu is not hijacked — taking it away while offering
nothing in return is a pure loss.
(cherry picked from commit 7ab5015737)
scripts/test-local-llm-harnesses.ts (#393) sits outside tsconfig.json's
include, so nothing type-checked it. config/tsconfig.scripts.json pulls
it in; npm run typecheck now runs both projects, the way the pr-bot
config used to be chained before the bot moved out of the repo.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Custom Model Endpoint Profiles (#393) let a session point its CLI at a
custom OpenAI-compatible endpoint by injecting env vars or a config file
and restarting the CLI in place. Review of the apply path found four
things, two of them destructive. This lands all four plus the smaller
items from the same review.
1. Clearing a selection did not clear it. The injected vars reach the CLI
via `tmux setenv`, which persists at the tmux-session level and is
inherited by `respawn-pane` (measured: `setenv FOO bar` survived two
successive `respawn-pane -k`), so deleting the keys from the session's
envOverrides relaunched the CLI still pointed at the old endpoint, and
for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just
been deleted. `Session.setCustomModel()` now reports the removed keys,
queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys`
carries them into `applyEnvOverrides()`, which `setenv -u`s them before
re-applying the live overrides, on the same path that already unsets
the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket
that `setenv -u HOME` hands the next respawn the global HOME back.
2. Applying a model to a local claude session killed the pane. The
relaunch was `claude --session-id <id>` and Claude refuses an id that
already has a transcript, and unlike the dead-pane respawn this one
kills a working pane first. `restartCli()` now pins the live
conversation id as the resume id for that respawn when the CLI's launch
declares a `fallback` chain, which renders the same
`--resume <id> || --session-id <id>` shape the docker and remote pane
commands use. Gated on the registry shape, not the CLI id: an entry
whose resume id is minted by the CLI itself never declares that chain.
3. pi, omp and grok wrote their config file and then launched without the
`--model` that selects it, so the file was ignored. The registry entry
now declares `customModelInjection.launchModel` (`custom/{modelId}` for
pi and omp, grok's `[model.codeman-custom]` block name), the builder
renders it, and `_withCustomModelLaunchModel()` applies it onto the
respawn options through `legacyConfigField`, leaving the stored
<Mode>Config untouched so a clear falls back to the user's own model.
A model id the CLI's `model` token pattern cannot carry is refused
with a 400 rather than silently dropped by the argv engine.
4. Remote (SSH) and Docker sessions reported `restarted: true` and changed
nothing: their `restartCli()` reattaches the durable tmux rather than
relaunching the agent, and the env lands on the local pane. Both are
refused with a 400 until those paths are plumbed.
Smaller items from the same review:
- The selection survives a Codeman restart as the disk-only `__customModel`
bookkeeping (endpoint, model, injected key NAMES, config dir, launch
model; never the values, which carry the API key). Recovery re-derives
the values from the endpoint store through the same apply path the route
uses and keeps the bookkeeping even when the endpoint is gone, so a
later clear still has keys to unset.
- Discovery goes through `webviewFetch()`, so the RESOLVED address is
judged by the same egress guard the web-tab proxy uses, and `baseUrl`
reuses `webviewUrlSchema` (http(s) only, no embedded credentials,
link-local and cloud-metadata addresses refused). undici's `fetch failed`
wrapper is unwrapped so the user sees the ECONNREFUSED underneath.
- `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session
config dir 0700/0600 (pi and omp embed the key literally), and that dir
is removed with the session.
- `PR.md` is gone from the repo root and the design doc moved to
`docs/custom-model-endpoints-plan.md` with the LAN address and the
personal name scrubbed; every reference follows. The guide's `authStyle`
text matches the shipped schema (`bearer | api-key`, default `bearer`)
and says that `customModelEndpointsEnabled` is read by nothing until
the picker lands.
- `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts`
(four real type errors fixed). It is not yet wired into `npm run typecheck`
because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json`
there is the one-line follow-up.
Tests: `test/session-custom-model-restart.test.ts` drives a real Session and
fails on the unfixed code for items 1 to 3; the route suite covers item 4
and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets
run before the overrides and that a shell-metachar key never reaches tmux.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Follow-up to #421 (remote-case file reads over ssh), addressing the review.
Symlink escape on a host without `readlink -f` (blocker). The probe's
portable fallback canonicalized only the directory chain and returned the
final component unresolved, so on macOS < 12.3 `ws/notes.txt -> ~/.ssh/id_rsa`
came back as `.../ws/notes.txt` (with the target's size), passed every
containment and blocklist check that runs on `realPath`, and `cat` followed
the link. The fallback now walks the directory chain with `cd -P`/`pwd -P`
and follows the LAST component with plain `readlink` for a bounded number of
hops, and anything it cannot fully resolve (a loop, a readlink failure, the
hop cap) is reported with an `x` marker that parses as null, i.e. 404. It
never returns the unresolved string. Measured on a real /bin/sh with
`readlink -f` shadowed: the pre-fix script reports `/ws/notes.txt`, the fixed
one `/secret/id_rsa`; both branches (native and fallback) now agree.
`PUT /api/sessions/:id/file-content` never had the remote guard the PR
described. It sits ahead of `validateSessionFilePath`, which resolves against
the LOCAL filesystem, because with a same-named directory on the Codeman host
(an sshfs mount of the remote tree, the documented stop-gap) the write landed
on the local twin while the viewer believed it edited the remote file.
ssh fan-out is bounded. `src/remote-ssh-limiter.ts` is a
document-conversion-limiter-shaped semaphore (default 4, env
`CODEMAN_MAX_REMOTE_FILE_SSH`) around every probe and buffered read; the
attachment-history list resolves its whole history in ONE batched probe
(`probeRemoteAttachmentHistory`, threaded into
`registerExternalAttachment({remoteProbes})` so the guards run unchanged)
instead of one handshake per entry; and probes chunk at 40 paths because the
whole script is one argv string. Terminal output in a remote session is
written on the remote host, so a prompt-injected agent printing hundreds of
`codeman://attach` links forked one ssh per link, each holding a 20 s
timeout, and a 100-entry history re-listed on every attachment:detected
tripped OpenSSH's default MaxStartups. Streams are deliberately not counted
(one per browser request, held for a whole playback, and gated behind a
counted probe anyway).
Smaller items from the same review: probe records are NUL-terminated and
index-keyed after a leading NUL (a newline in a filename can no longer shift
the alignment, and the banner is fenced off without last-N-lines guessing);
size comes from `stat -c %s || stat -f %z`; the three IO functions refuse
under VITEST instead of opening a connection; an unreachable host now reads
as unknown (missing: false) for detected AND external history entries, where
external used to fold its 502 into missing; a client that aborted during the
guard probe has its body's ssh child reaped (`reply.raw.destroyed` is checked
before the close listener is attached); `describeExecError` never returns
Node's `Command failed: <ssh line>` message, which carried the identity path
and the probe script into a 502 body; and the docs note that
`isSensitivePath`'s three home-anchored entries resolve against the Codeman
host's home, not the remote one.
Tests: the probe script runs on a real /bin/sh with a `readlink` shim that
rejects `-f` (the escape, a relative chain through a symlinked directory, a
loop, a newline filename, banner chatter that itself looks like a record),
the limiter's cap and FIFO order, and route tests for the PUT guard (local
twin untouched, no connection), the single batched history probe, the
unreachable-host alignment and the aborted-client reap. All four route tests
fail against the pre-fix file-routes.ts.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The lost-frame recovery page is answered ahead of the credential checks in both
auth hooks, which makes it the third unauthenticated 200 beside the two hook
routes, and the only one decided by request headers alone. CLAUDE.md's security
table listed exactly two, and docs/web-tabs.md is not where anyone auditing that
looks, so it now has a row in the table and a fourth property in
docs/security-architecture.md section 10b, including the `/` carve-out and its
credential-free condition. Both state the property that comes with it: a
non-browser client can set those headers, so an unauthenticated caller can tell a
registered route (401) from a non-route (200) and enumerate the route table,
accepted because the routes are public in docs/api-reference.md.
docs/web-tabs.md gains the landing-page case in layer 6 and a Known limits entry:
masking trades away the Referer safety net, only HTML is rewritten server-side,
and a root-absolute url() inside an inline <style> block has the masked document
as its Referer, so it 404s where the Referer fallback used to rescue it. External
stylesheets are unaffected.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The runtime shim masks `/webview/<cap>/` off a proxied page's URL so its router
boots on the path it expects, and the landing page masks to exactly `/`. A
`location.reload()` there (a Vite dev server on a config change or a failed HMR
update, the likeliest case in the feature's own motivating scenario) therefore
asks for Codeman's root as an iframe navigation. `serveLostWebviewFrame()`
returned early for `/`, so on a passwordless install the frame received Codeman's
own app shell and rendered it inside the web tab, and with a password it got a
401 in the frame. Either way no `codeman:webview-lost` message was posted, and
because the document loaded fine the load handler cleared the failed-frame panel,
so the Reload / Open in new tab affordances never appeared. Before masking the
frame's URL was the prefixed one, so a reload worked; this was a regression.
`/` is the one lost-frame path a registered route also serves, so the route
table cannot tell that reload from a real navigation. Credentials can: nothing
in Codeman frames its own root, and a sandboxed frame is opaque-origin with no
cookie and no Authorization header. `carriesAuthCredentials()` (pure, in
webview-proxy.ts) makes that test, and `/` is now admitted by the auth hook only
when it fails; a framed `/` that does carry credentials still gets the shell.
Without a password no auth hook runs at all, so the index route applies the
same test itself (`isLostWebviewRootFrame`) before rendering the shell, and the
three places that emitted the recovery page share `sendLostWebviewFramePage()`.
Tests: the password form in webview-auth-exemption (recovery page for a
credential-free framed `/`, shell with valid Basic auth, 401 with a stale cookie
or a top-level navigation), the passwordless form against a real WebServer in
webview-lost-root-frame (port 3198), and the credential predicate in
webview-proxy. All three fail without the fix. Verified against a live isolated
instance as well: a framed `GET /` with no credentials answers the 470-byte
recovery page, a top-level `GET /` and a framed one carrying a cookie answer the
shell.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The lost-frame handler in webview-tabs.js remounts a web-tab frame at the path
the frame reports it lost. It promised "path only, never an origin" and collapsed
a leading run of slashes so `//host/x` could not jump the frame off the proxy,
but it left two spellings through that the WHATWG URL parser treats the same way:
a backslash, which is read as `/` for http(s) schemes, and an ASCII tab or
newline, which the parser deletes before it looks at anything, so `/\host/x` and
`/<tab>/host/x` both resolve to `https://host/x`. That mattered only in
direct-mode tabs, where `POST /api/webviews/:id/open` returns no embedUrl and the
recovered path is resolved with `new URL(path, src)` straight into the frame's
src; a page in such a tab could remount its own frame on a foreign origin.
Not an escalation (the page can already navigate itself anywhere, and the remount
carries no Codeman-origin access), but the comment did not hold and the existing
test only covered the form that already worked. The handler now strips tab, CR
and LF, collapses any leading run of `/` or `\` to one `/`, and refuses whatever
still opens a second separator. The new test drives the reachable direct-mode
branch with all four spellings and fails without the fix.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The capability list is DERIVED from what the scripts do (chown => CHOWN +
DAC_OVERRIDE, a setpriv uid/gid drop => SETUID + SETGID, `init: true` next to a
uid drop => KILL) and compared to docker-compose.yaml's cap_add, the
entrypoint's own required_caps diagnosis, and the lists quoted in docker/README.md
and CLAUDE.md, so the drift that shipped the missing CAP_KILL fails here rather
than on someone's server. Also pinned: the CLI prefix is appended to PATH in
server.Dockerfile and entrypoint.sh pins its PATH before its first command;
Start-Codeman.sh derives PUID/PGID before creating the cases dir, builds before
`down`, writes the source marker only after a refresh, and never aborts on a
failed volume removal.
git_head_commit is run as the script defines it, extracted by its own
delimiters into a real bash, against temp repos made with real git: a symbolic
ref with a loose ref file, a detached HEAD, packed refs after `git pack-refs`,
a linked worktree (which must resolve nothing rather than something wrong) and
a directory that is not a checkout.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
docker/README.md and docs/docker-compose.md now say that the container starts
as root, corrects a daemon-created bind source and drops to PUID:PGID with
setpriv, which capabilities that needs, and that a compose file written
elsewhere must carry them. The README's PowerShell example runs Compose from
inside docker/ so the override file is discovered, instead of the `-f
docker/docker-compose.yaml` form its own Local customisation section warns
silently drops it, and the reverse-proxy section no longer asks for an override
file now that docker-compose.yaml forwards CODEMAN_ALLOWED_HOSTS itself.
.dockerignore excludes docker-compose.override.* everywhere: it is the
documented home for host-specific settings and rode `COPY . .` into the image,
the same shape as the docker/.env exclusion above it (verified with a scratch
build context: the override files and docker/.env are absent, .env.example and
the compose file present).
CLAUDE.md's Compose paragraph carries the corrected cap list, the writability
probe, and the two traps behind it (KILL is for tini, the CLI prefix is
appended to PATH).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Start-Codeman.sh created CODEMAN_CASES_PATH with a plain `mkdir -p` BEFORE it
derived PUID/PGID from the appdata directory, so the new directory landed as
the invoking user's uid and primary gid. On a host set up the way the README
suggests (`chown -R 99:100 <appdata>`) that gid is not PGID, and the container
refused to start on a directory the script had just made. PUID/PGID are now
derived first and the directory is chowned to them right after creation, with
a clear host-side error when that is not possible. As root this always works,
which also retires the old "refusing to create as root" branch for this path.
The build-artefact volume refresh had three holes. The docker-build-source.json
marker was written whether or not a volume had actually been removed, and the
project name came from a sed over `docker compose config --format json` keyed
on two-space indentation: an empty name made the label filter match nothing,
nothing was removed, and the marker recorded the new HEAD, so the check never
fired again while the stale volume kept serving old code. The name is now
parsed indentation-agnostically, an empty result falls back to `down --volumes`
(the documented reset; both volumes re-seed from the image by a plain copy),
the marker is written only after a successful refresh, and a failed `docker
volume rm` warns and leaves the marker alone instead of aborting under set -e
with the stack down. The image is also built BEFORE `down`, so the deployment
is offline only for the recreate rather than for the whole rebuild.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Three changes to how the Compose container starts as root and drops to
PUID:PGID, each reproduced on Docker 29.1.3 / Compose v5.5.0 with a minimal
image of the same shape as server.Dockerfile.
- cap_add gains KILL. `init: true` makes tini PID 1, and tini stays root while
the entrypoint drops the server to PUID. Signalling a process of a different
uid needs CAP_KILL, and `cap_drop: ALL` had removed it, so every `docker
compose down`/`restart` ended in `[FATAL tini (1)] Unexpected error when
forwarding signal: 'Operation not permitted'` and the server being SIGKILLed
instead of running `server.stop()`. Measured: without KILL the trap never
fires, with it the child logs `GOT SIGTERM`.
- /opt/codeman-cli/bin is appended to PATH, never prepended, and entrypoint.sh
pins its own PATH to the system directories before its first command. The
prefix is chowned to the runtime account so sessions can update the agent
CLIs in place, and the root entrypoint resolved stat/chown/setpriv by bare
name through it: a `setpriv` planted there by the unprivileged uid ran as
uid 0 at the next start. The image's full PATH is handed back to the server
at the exec (`env PATH=...`), since Codeman resolves the CLIs through it.
- The ownership gate becomes a writability probe. A directory owned by neither
root nor PUID:PGID is no longer refused on ownership alone; it is tested with
`setpriv --reuid PUID --regid PGID --groups <same groups> test -w`, the exact
identity the server gets, so a group-writable tree, an ACL or a CIFS/NFS
mount reporting some unrelated uid all pass, and the refusal names path,
owner and PUID:PGID. Root-owned directories are still chowned first.
Also: a pre-flight runs the drop before touching anything and, when it fails,
prints the cap_add list the compose file needs, so an out-of-tree compose file
(Unraid's Compose Manager) gets a one-line diagnosis instead of a restart loop;
`--bounding-set -all` is gone, since it is a silent no-op without CAP_SETPCAP;
a root:root Docker socket now produces a warning that Docker cases will not
work rather than silently losing group 0 at the drop; and CODEMAN_ALLOWED_HOSTS
is forwarded from .env with an empty default (documented as a commented entry
in .env.example so the parity test and the updater's env gate both stay quiet).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The bash 3.2 job added with #380 cannot reach dsh_banner_probe, which is
the function #382 was filed against: this image ships `timeout`, so the
optional-prefix array is never empty, and with no `dsh` binary anywhere on
PATH the probe is not called at all. The fix landed in 1.28.2 with nothing
guarding it, and the failure mode is a runtime abort under `set -u` that
`bash -n` cannot see, which is precisely why the reporter had to find it by
reading the source rather than by running anything.
So call the probe directly, with `timeout` hidden behind a narrowed PATH,
and refuse to pass if `timeout` is still reachable (a guard that silently
stops exercising its branch is worse than no guard). Both directions are
asserted: a real DeepSeek Harness banner is accepted, and Debian's unrelated
`dsh` is refused, so the check covers the identity half too.
Verified by reverting install.sh to the pre-fix expansion, where the step
fails with the exact error from the issue, `runner[@]: unbound variable`.
Refs #382
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every per-skin xterm palette declared its selection layer as `selection`,
the key xterm.js renamed to `selectionBackground` in v5. An ITheme is a
plain object handed straight to the terminal, so an unknown key is not an
error, it is dropped: all seven skins have been drawing xterm's built-in
default, rgba(255,255,255,0.3), rather than the colour sitting next to it
in the palette.
Nobody saw it on the dark skins, where white at 30% is close to what those
palettes asked for. On the four light skins it is white over a near-white
background: blended, Paper Gray's selection differs from its own background
by 3/255. That is not a subtle highlight, it is no highlight, and it looks
exactly like a selection gesture that failed, which is part of what #360
reports on Android Chrome.
test/skin-themes.test.ts pins both halves: the key name, and that the
blended selection stays at least 16/255 from the background on every skin,
plus the light-skin fallback landing under that floor, which is what makes
this a fix rather than a rename.
Refs #360
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two keys #408 adds to the mobile keyboard accessory bar send
Shift+Left and Shift+Right, which are Codex bindings (edit the last
queued message, step back through the prompt stack). They shipped on
both agent layouts, so a claude, pi, grok, omp, deepseek or gemini
session got two keys that do nothing. That was not only cosmetic: a tap
goes through sendNavKey(), which adds the session to
_echoPassthroughSessions and hands editing to plain PTY echo until Enter
or Ctrl+C, so on a phone a dead key also switched off the local echo
that makes typing feel instant there.
The reveal now follows the shape the 🧠 key already uses. The buttons
stay in both templates, carry an accessory-btn-codex marker class, and
are display:none in styles.css until the bar element carries
codex-enabled. The class has to live on the bar rather than on the keys
because setMode() rebuilds the buttons' innerHTML on every layout
switch. syncCodexKeys() toggles it from the active session's mode
(the same lookup _isShellSession() uses) and is called at init and from
refreshForActiveSession(), which selectSession() already invokes on
every switch. A session's mode is readonly on the server and fixed at
create, so no other event can change the answer; the welcome screen
(no active session) reads as not codex and hides the keys.
The frontend id-branching guard (test/cli-registry-no-id-branching.test.ts)
scans only src/**/*.ts, so the mode comparison in a public JS file is
in bounds, the same as the existing shell check beside it.
Tests: the new describe block in test/mobile-shell-keyboard.test.ts pins
the marker class in both templates, the CSS pair, the class for a codex
session in both layouts, its absence for claude/shell/pi/omp/deepseek,
the re-sync in both directions on a session switch, the no-session case,
and the init + refresh wiring. All six positive assertions fail without
the source change. README and the changeset now say the keys are
Codex-only.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A clicked path that points OUTSIDE the case directory goes through the attachment
routes (the frontend's `_isExternalPreviewPath` sends every absolute path not under
`workingDir` to `POST /attachments`), and those had the same local-`fs` assumption
as file-raw: `realpathSync`/`fs.stat` on a path that only exists on the remote host,
so the file never opened — the case the #415 report was actually about.
- `registerExternalAttachment()` accepts `remote` and resolves through
`remoteProbePaths` (canonical path, size/mtime, kind, plus the workspace root for
the confinement check). Everything around it — blocklist, extension allowlist,
workspace confinement, registry/dedupe — is now shared by both branches, so the
remote path cannot drift from the local one.
- The by-id routes (`raw`, `preview`, `thumbnail`), the metadata poll and the
attachment history list resolve over ssh too. `raw` streams with the same
Range contract as file-raw; `preview` (office) and `thumbnail` answer 400 for a
remote record; an unreachable host answers 502, a vanished file 404.
- Which host a record is read from follows the SESSION, never the path string: the
same absolute path is a different file on each host, and a remote session never
falls back to a local file with that name.
- Codex generated artifacts keep force-workspace confinement for a remote case: the
well-known artifact directories are anchored at THIS host's home, so only a file
inside the remote workspace is trusted.
Still local-only by design: writes, office conversion, thumbnails, the file
tree/picker and tail-file.
Rebased over #390, which moved the phone tier's cutoff from 430px to
600px. The palette's compound fold rule now lives in the 600-768px band
mobile.css pads, the cascade samples the palette inside that band, and
the closed iPhone Duo (466pt) is a phone rather than a small tablet while
the open one (626pt) stays a tablet. Comments in both stylesheets, the
device registry, CLAUDE.md and architecture-invariants say 600.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Ported from #416 (discussion #405): a statusline reading just `codeman`
is what a hand-run claude in a managed repo showed, and it reads as a
broken config rather than a footer. Three paths produced it and all
three now yield an empty footer: the exporter's `|| echo codeman`
fallback (now `curl -sfk ... || true`, with -f keeping an HTTP error
body off stdout), the unknown-session answer of POST /api/status-telemetry,
and formatSessionStatusText() with nothing to show. The exporter script
marker moves to V4 so live installs pick the new content up on the next
spawn.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Three small follow-ups from the #361 review.
A tmux setenv survives respawn-pane, so _configureStatusLineUserCommand
returning early when the user has no statusline left a previously
exported CODEMAN_USER_STATUSLINE_CMD in place: a user who deleted their
own statusline kept getting the stale one wrapped, and lost Codeman's
footer print-through, until the tmux session was recreated. It now
issues `setenv -u` in that case, the same shape as the effort-level
cleanup in applyEnvOverrides.
ensureStatusLineExporterScript truncated and rewrote a script that live
sessions execute on every statusline render, and chmod'd it after the
write. It now writes a temp file next to the target, chmods that, and
rename()s it into place.
The non-tmux direct-PTY fallback carries no exporter; that is now stated
at the spawn site and in the architecture-invariants paragraph rather
than left as a silent gap.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Two follow-ups to #361's sticky telemetry switch.
GET /api/settings reconciled an absent showPlanUsageLimits by persisting
true, but readJsonConfig() answers {} for ANY read failure (a parse
error, EACCES, EMFILE, a read landing inside PUT's non-atomic write), not
only ENOENT, and every page load calls this route, so one unlucky read
replaced the whole settings file with a one-key file. The route is a
plain read again and the default moved into the reader:
readPlanUsageTelemetryEnabled() treats an absent key as ON, the same way
readWorkspaceHooksEnabled() does, which is what the desktop chip already
shows for an install that never touched the setting.
saveAppSettings() sent showPlanUsageLimits on every save. The chip
defaults OFF on handhelds, so a phone saving its font size persisted
false and switched collection off for every desktop, whose chip then
went stale with no error anywhere. The key is now stripped like the
other per-device display keys and re-added only when the save FLIPS the
chip relative to what the device had (planUsageCollectionFlip), so an
explicit toggle on any device still writes it in either direction.
Tests pin both: the GET route with a mocked filesystem (absent, missing,
EACCES, garbage, explicit), the reader default, and the flip helper plus
its wiring in saveAppSettings.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
CLAUDE.md's folding-devices rule gains the two new invariants (a shape change
with the keyboard up baselines to window.innerHeight; a base gutter overridden
by a later @media block needs its own zero-base fold restatement, and a
compound rule written against a mobile.css shorthand is scoped to that band)
plus the architecture-invariants pointer it lacked; the new Folding devices
section there carries the mechanisms and the measurements. The device count
is 138 since the two Duo profiles landed (68 Playwright + 70 custom).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Three cascade problems in the fold reserved-region CSS, each measured by
computed style in headless Chromium (styles.css + mobile.css in index.html
link order):
- The unconditional .path-picker-overlay / .path-preview-overlay fold rules at
the end of the file beat the `padding: 0` both overlays set under 600px, so
every phone got a 16px and 18px gutter on dialogs built flush (393 and 500px:
edges floating off the screen). The fold strip is now restated on a ZERO
base inside the same media query: 0/0 without a fold, the strip alone with
one, 16/18 plus the strip from 626px up as before.
- .modal.command-palette-modal was unscoped, so outside the 430-768px band
(where mobile.css pads the palette with a shorthand) it ADDED 0.75rem with
no gutter to compose with and pushed the shell 6px off centre at 393, 900
and 1400px, while inside the band the shorthand beat the generic .modal rule
on the bottom side and the palette lost its block-end gutter. The compound
rule now lives inside that band and restates both sides.
- The tabletop cap on .response-viewer lost to mobile.css's `max-height:
92dvh` under 430px (same specificity, later file). mobile.css now carries an
identical twin at its end.
test/foldable-layout.test.ts simulates the padding cascade across both files
at every breakpoint, with and without the fold rules, and requires the two to
differ by exactly the fold strip; it also pins the palette rule to the band
mobile.css keys on and the response-viewer twin to the styles.css value. Its
model reproduces the Chromium numbers, and against the pre-fix stylesheets it
fails on all three problems.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A shape change with the keyboard up re-baselined initialViewportHeight to the
SHRUNK visual height, so heightDiff was 0 and the settle event the OS fires at
the new width (or any later address-bar drift) satisfied the hide branch and
ran onKeyboardHide() with the keyboard still on screen: accessory bar hidden,
toolbar lift dropped, main's padding cleared. It could not recover, since no
further 150px drop re-arms the show branch against a baseline already sitting
at the shrunk height.
Baseline to window.innerHeight instead when the keyboard is up: the page sets
no interactive-widget, so the keyboard shrinks only the visual viewport and
the layout viewport stays the display's full height on both engines, the same
fact updateLayoutForKeyboard() relies on.
The vm harness now models the two heights separately (resizeTo takes an
optional layout height) and pins the fold flavour (626x590, 466x378, 466x378),
the rotation flavour (393x359, 852x150, 852x160) and the eventual close. All
three fail against the old line.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Apple's "Designing for iPhone Duo" asks an app to adapt to both displays,
to stay continuous as the device opens and closes, and to treat the band a
partly-open display folds through as a reserved region. Three things here.
1. A visual-viewport resize that changes the WIDTH is the device changing
shape (a rotation, or a foldable opening or closing) and is never the
virtual keyboard, which only ever takes height. handleViewportResize()
read any height drop over 150px as the keyboard appearing, so closing a
Duo (890 to 678pt tall) latched keyboardVisible with no keyboard on
screen: the accessory bar appeared, main grew 84px of dead padding, and
updateAppHeight() stopped refreshing --app-height. The latch was sticky,
because clearing it needs the height back within 100px of a baseline
belonging to a display the user is no longer looking at. Rotating any
phone hit the same latch. The shape branch re-baselines instead, which
is also what lets a keyboard opened after the fold be detected.
2. The hinge is now a reserved region in CSS. --fold-inline-end and
--fold-block-end measure the strip to keep clear from the Viewport
Segments env() variables, and are 0px everywhere else, so the seven
centred overlays are inert by construction off a foldable. Each shrinks
its content box with padding rather than the box itself, so the backdrop
still covers the far side of the fold and still swallows taps there.
3. iPhone Duo (outer) and iPhone Duo (inner) join the mobile device
registry, derived from Apple's published pixel specs at 3x.
Verified in Chromium: flat, a dialog stays centred at 313 of a 626pt
viewport; in book pose it centres at 153 inside the 0-305 leading segment
with its right edge at 293, while the backdrop still spans all 626. The
3-term calc on the offline overlay resolves to 367px in tabletop pose and
20px flat.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three instructions a future contributor would follow literally were stale after
the last review round: the "Adding a CLI" checklist sent the agent-image reason to
AGENT_IMAGE_SPECIAL_CASES, a constant that no longer exists (it is
discovery.install.agentImageLayer on the entry in stock.ts), the trust-boundary
paragraph credited the embedded-commands pin to the invariants test when it is
test/cli-catalog-sync.test.ts, and install.sh claimed "the parity test" pinned the
DeepSeek Harness banner when no test did. That pin now exists: the invariants test
asserts the script's grep literal and the registry's discovery.identity.regex agree
on "DeepSeek Harness", and the comment names it.
docs/docker-cases.md separated the two reasons a CLI stays out of the shared npm
layer (no npmPackage at all versus an agentImageLayer entry), which it had folded
into one, and architecture-invariants no longer lists the agent image's CLI set by
hand.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Choosing "s" (Skip) in the new catalogue-driven install menu warned, printed the
install hints and then fell into the shared "The selected AI CLI failed to install"
gate one line below, because CLI_FOUND_COUNT is 0 by construction inside that block
and skipping does not change it. The AI CLI check runs before the clone and the
build, so a user who picked the documented skip option ended up with nothing
installed. The code this replaced guarded the gate with an elif on the skip choice.
The menu moves out of main() into offer_ai_cli_install() and the gate moves inside
the install branch: skipping continues to the clone, a chosen install that leaves
nothing behind is still fatal. Being a function, the interactive path can now be
driven with a stubbed read_reply, which is what nothing reached before: two
behavioural tests in test/install-sh-invariants.test.ts run the real function in a
real bash (skip continues with exit 0, a failed install dies with exit 1), and the
bash 3.2 CI step drives the skip path in the container as well.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Follow-up to #390. PHONE_MAX had become 599, an inclusive bound, while
three of its four consumers still read it as exclusive (width < PHONE_MAX
for phone); the one site that switched to <= disagreed with
getDeviceType(). It is 600 again with < at every site. The breakpoint
table in docs/mobile-testing-report.md says 600, and the three 430px
visual baselines are removed: they depict the tablet tier now, and the
visual suite recreates a missing baseline on its next run on the machine
that owns them. device-matrix.test.ts is also run through Prettier, which
the commit hook demanded and the format gate (src/ only) never did.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Runs the mirror/version drift check and both strict validations. The paths
are spelled out because the documented pair ended in a bare `.` that reads as
a full stop when copied out of prose, which surfaced as `missing required
argument 'path'` on first use. Needs the `claude` CLI, so it is a local check
rather than a CI step.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`scripts/pr-bot/` was maintainer tooling, not part of the server, the CLI or the
npm package: a Telegram bot that reviews open pull requests in Codeman sessions
and reports to the maintainer. It now lives in its own private repository and
keeps running unchanged, as a client of Codeman's HTTP API like any other.
It moved because it grew a second watcher, for GitHub Discussions, and shipping
that here would mean publishing the briefs it hands its review sessions, the
judgement calls in them and its safety model. None of that helps anyone
installing Codeman, and all of it is easier to change when it is not a public
interface. The move cost nothing structurally: the whole tree depended on one
external package plus Node builtins.
What this removes from the repo, and nothing else: the sources, their three test
files, `config/tsconfig.pr-bot.json`, `docs/pr-bot.md`, the `pr-bot` npm script,
the bot's globs in the typecheck/lint/format scripts, and its knip entry. CLAUDE.md
keeps a short pointer in place of the section, because the bot still constrains
work in here: it takes the `prbot-<n>` and `dscbot-<n>` session names on the local
Codeman, holds clones under `~/.codeman/pr-bot/`, and fetches pull-request heads
into `refs/pr-bot/*` of this checkout, which it must never check out or reset.
The CHANGELOG entries from 1.25.0 and earlier still describe it. That is history
rather than drift, and is left alone.
Verified after the removal: typecheck, lint and format:check clean, and the suite
passes 6843 tests across 357 files, which is the previous run minus exactly the
70 tests that moved out with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A remote case's workingDir is an absolute path on the remote host, but the
file read routes resolved it with local `fs`: `validateSessionFilePath`'s
realpathSync fails for a path that does not exist on the Codeman host, so
every preview of an agent-written file answered "File not found" (#415).
Add src/remote-files.ts as the single remote-read layer, built on the same
buildSshConnectionArgs() the launch uses:
- remoteProbePaths(): ONE round trip returning realpath + stat for the
requested path AND the workspace root, so containment is checked against a
remotely canonicalized root (a symlinked remotePath is ordinary).
- remoteCreateReadStream(): streams the body (cat, or tail -c +N | head -c L
for a Range) with nothing buffered in memory, and reaps the ssh child when
the response ends so an aborted download cannot orphan it.
- remoteReadFile(): bounded read for file-content.
file-raw, file-content, file-preview and file-thumbnail now share one local/
remote target resolution. Guards keep their local strength: lexical pre-check,
remote realpath, workspace containment, sensitive-path blocklist, and the size
cap applied to the remote size before any bytes are read. An unreachable host
answers 502 with the remote reason instead of a misleading 404. Nothing is ever
copied to the Codeman host and there is NO local fallback (an sshfs mount of
the same tree must not shadow the remote bytes).
Deliberately unchanged: writes (edit=1 / PUT now answer 400 explicitly while
the viewer hides its Edit affordance), office previews, thumbnails, file tree,
picker, external attachment registration and tail-file stay local-only.
Measured with both installed: a fresh Claude Code lists `codeman` (the
user-level or per-case copy) and `codeman:codeman` (the plugin). Neither
shadows the other and both work, so this is noise rather than breakage, but
the README, the wiki and the plugin README now say to choose one.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
With the repo root as the plugin root, `claude plugin install codeman@codeman`
copied the whole checkout into its cache and, because that root carries a
package.json, ran an npm install there: 832 MB, 511 packages and this repo's
postinstall build on every installer's machine (measured from a clean worktree
of the previous commit). A plugin root must be a directory without one.
The plugin is now `plugins/codeman/`: its manifest, a README, and a MIRROR of
`skills/codeman/`. A mirror rather than a symlink because the install copies
the plugin directory and a link pointing outside it would dangle; a mirror
rather than the source because every install path, injector and doc already
names `skills/codeman/`. `scripts/sync-plugin.mjs` (replacing
sync-plugin-version.mjs) mirrors the skill and syncs both manifest versions
inside `version-packages`; `test/plugin-manifest.test.ts` pins byte-identity,
the versions, the absence of a package.json in the plugin root and that the
repo root `.claude-plugin/` holds only the marketplace manifest.
`claude plugin validate --strict` now passes for both the plugin and the repo
root. Install commands are unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`.claude-plugin/marketplace.json` at the repo root makes
`/plugin marketplace add Ark0N/Codeman` work, and the one plugin it lists is
the repo itself (`source: "./"`), whose one component is `skills/codeman/`.
So `/plugin install codeman@codeman` is a third install route next to
`npx skills add` and `codeman skill install`, and the skill shows up in the
plugin directories that index Claude Code marketplaces.
Both manifests carry package.json's version: `scripts/sync-plugin-version.mjs`
rewrites them inside `version-packages`, right after `changeset version`, and
`test/plugin-manifest.test.ts` pins the equality, the skill's frontmatter name
(without it the installed skill would be named after a versioned cache dir),
and that no other plugin component (`commands/`, `agents/`, `hooks/`,
`.mcp.json`, `settings.json`) appears at the repo root, since an install would
silently ship it.
Verified with `claude plugin validate` (one expected warning: CLAUDE.md at a
plugin root is not plugin context) and a local marketplace add, install,
details, uninstall cycle against a clean checkout of this commit.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every release now gets an Announcements post shaped like #418 and #302:
features first, contributor mentions inline, Thanks at the end. The
step records the GraphQL command and ids, and why it exists: posts
stopped at 1.18 while ten releases shipped unannounced.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Bold text on the theme's default foreground carries exactly ONE cue, the
weight step. Claude Code marks its markdown bold with a bare ESC[1m and
changes no colour, and xterm substitutes a bright colour for bold only
when the foreground is a palette index 0-7, so the substitution never
fires for default-foreground text. A family shipping only a regular and
a bold face keeps that step small (measured on Consolas: glyph ink rises
from 14.25% to 16.57%), and picking a different family does not help,
because 400 stays 400 whatever the family. Lowering the NORMAL weight is
the only way to widen the gap.
Two per-device settings beside "Terminal font" in the Font group, each
defaulting to xterm's own value for its slot, so an untouched install
renders exactly as it did before. Both thread into the main terminal and
the Agent Teams panes, and apply on save without a reload.
The bundled face had to be unclamped in the same change or the settings
would look broken on a stock install. fonts/jetbrains-mono-variable.woff2
carries a wght axis of 100 to 800, but styles.css declared the face
`400 700`, and the descriptor is what the browser synthesizes from: at
that range 100, 200 and 300 rendered identically to 400 and 800
identically to 700 (measured in headless Chromium, both directions).
The two families ahead of it in the default stack, Fira Code and Cascadia
Code, exist only if the user installed them, so for most installs
"normal = 300" would have been a no-op. Declared `100 800`, every step is
distinct: 61%, 77% and 90% of the ink at 400, and 800 adds ~14% over 700.
Nothing in the stylesheets asks for a monospace weight outside 400-700,
so widening it changes nothing that rendered before.
Details that are easy to get wrong and are pinned by tests:
- Each slot falls back to its OWN xterm default, so an unset bold weight
can never inherit `normal` and become a visible change.
- A live save refreshes both echo overlays. They cache
terminal.options.fontWeight and paint it into their spans, so without
it the characters being typed keep the old weight while the rest of the
screen changes. Most visible on a phone, where local echo is on by
default.
- A live save reaches open Agent Teams panes, which read their options at
construction, exactly as applyTerminalSkin() propagates its own.
- A stored weight the picker does not list (a hand-set 350) is added to
the select rather than dropped, so merely opening App Settings cannot
reset it.
- _awaitTerminalFont() is untouched. CharSizeService measures through the
CSS `font` shorthand, which resets the weight, so the measured face is
always the 400 one and a weighted descriptor would request nothing new.
Verified end to end in a headless browser against a live server: the save
reaches the running terminal with no reload, the settings PUT stays 200
(both keys are display keys and are stripped before it, since
SettingsUpdateSchema is strict), the value survives a reload, and the
painted terminal really changes weight with the bundled font (lit-pixel
ink 0.83 / 0.95 / 1.00 / 1.13 / 1.21 at 100 / 300 / default / 700 / 800).
Proposed and analysed by @irisitymichaelgrundberg in discussion #403.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Auditing every release after the previous backfill turned up two more. 1.23.0
had no Thanks in either artifact; its three PRs (#337, #341, #338) are authored
by the maintainer, so like the others it credits the release it follows.
1.22.0 had the section on its GitHub release but never in CHANGELOG.md, which
is the drift that happens whenever the block is added post-hoc instead of in
the changeset.
Every release from 1.21.0 forward now carries a Thanks section in both places.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Those three shipped with maintainer-only commits and no Thanks section, on the
reasoning that a release with no contributor PRs has nobody to credit. That is
the wrong test: the newest tag is what GitHub marks Latest, so a contributor
who shipped in the release next door lands on a page acknowledging nobody.
Each now credits the release it follows and says so, rather than claiming work
its contributors did not do: 1.24.2 the hotfix on 1.24.1, 1.24.4 the same-day
follow-on to 1.24.3, 1.26.0 the day after 1.25.0. Wording is carried over
verbatim from those releases. The matching GitHub release bodies were edited to
match, since the two are separate artifacts once version-packages has run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1.28.1 is a same-day follow-on to 1.28.0 and is the release people land on as
"Latest", so it credits the same three contributors rather than showing no
acknowledgement at all. Matches the section just added to its GitHub release.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every release credits its contributors in both places: a "### Thanks" block
and a comment on each merged PR. The PR comments went out, this did not.
Past releases carry it because the block was written INTO the changeset, which
is what feeds both CHANGELOG.md and the GitHub release body; mine went only on
the GitHub release, so the changelog was short a section. Put it in the
changeset next time rather than patching both by hand afterwards.
1.28.1 gets none on purpose: every commit in it is a maintainer commit, the
same call as 1.26.0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
COM step 4. Staged as a single hunk: the shared checkout also holds another
session's in-progress pr-bot discussions work in this file, which is left
untouched and uncommitted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The fixed 120px/96px caps on the two wrapped layouts were row counts in disguise: a
third row was clipped into a ~4px scroller, hiding tabs inside a container nothing
invites you to scroll, while the header had the page below it to grow into. Both
layouts now share one rule capped at var(--tab-strip-max-height, 40vh), a safety net
for an absurd session count rather than a row limit.
Verified before shipping: .header is min-height + flex-shrink: 0 so it can grow, and
terminal-ui's ResizeObserver refits the terminal when it does; updateTabOverflowMode()
returns early for any non-desktop viewport, and below 1024px mobile.css pins the header
to max-height: 48px, so this is desktop-only in effect; the selector is comma-grouped
rather than :is(), so each arm keeps (0,2,0) and mobile.css's overrides still win on
source order. PostCSS parses the file cleanly (prettier ignores styles.css).
Authored in a parallel session against this shared checkout.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
House style, and these land in the changelog. Only the sentences added in the
previous commit are touched; the em-dashes in contributor text and in the
pre-existing COD-54/COD-115 comments are left alone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#409 (Claude truecolor). The changeset becomes the changelog, and its premise
does not hold on tmux 3.2 or newer. Measured here on tmux 3.4: `default-terminal`
sits at its compiled default of `tmux-256color`, a live claude pane reports
`TERM=tmux-256color`, and supports-color reads that as 256 colors, where
rgb(55,55,55) lands on ESC[48;5;237m — visible, just not the color the theme
named. The invisible block the PR describes needs TERM to resolve to a 16-color
entry: tmux older than 3.2, or a ~/.tmux.conf setting `default-terminal screen`,
which Codeman's own tmux server does read (it passes no -f). Both the changeset
and the invariants paragraph now say that, so the next report here gets paired
with the reporter's tmux -V instead of being read as universal. The change itself
stands on the simpler argument: claude was one of two entries not asking for
truecolor while twelve do.
Also reorders buildClaudeEnv(). It applied the registry's unset/exports AFTER the
whole env was built, so a clis.json entry naming CODEMAN_HOOK_SECRET_FILE or PATH
would strip it on the direct-PTY path while the tmux pane kept it — buildEnvExports()
emits `...cliEnv` ahead of `export CODEMAN_MUX=1` and cannot. The block now runs
first and Codeman's own keys are assigned on top, matching the pane.
#404 (Ctrl+Z trap). Adds the missing changeset, and records what the trap does
not cover: an agent CLI already holds its tty with ISIG off (verified on three
live panes: `susp = ^Z -isig -icanon`), so this is defence for the startup window
rather than a fix for the steady state, and two input paths still reach the PTY
unfiltered — the mobile accessory bar's one-shot Ctrl and the CJK textarea.
#399 (path picker sort). The server sorts by name and cuts at 500, so the client
sorting those 500 by date gives "the newest of the first 500 by name", which is
wrong in exactly the >500-entry folder the date sort exists for. The status line
now says "(first 500 by name)" so the cut is legible, with the reasoning parked
on _sortEntries.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
buildClaudeEnv(), the direct-PTY fallback taken when mux creation fails, now
reads getCli('claude').env and applies its unset and exports lists. It used to
delete COLORTERM and CLAUDECODE from a hand-maintained list of its own, which
left it contradicting the registry entry that the tmux pane and the attach
client both read. An engine value needing a mux name has nothing to resolve
against on this path, so it is skipped rather than guessed.
Claude no longer unsets NO_COLOR. The invisible-background bug does not need
it, and unsetting it overrides a preference the user set deliberately, so a
user who exports NO_COLOR globally keeps monochrome panes. The other seven
truecolor CLIs still unset it; that inconsistency is intentional and the
comment on the entry says so.
The invariants doc gains a Terminal colour env paragraph under Session launch
modes, where a reader looking up Claude will find it — the previous sentence
sat under a heading that lists only the non-Claude CLIs. It now says the lists
are the stock catalog and a clis.json override replaces them wholesale, and
that the declarations reach the tmux pane, its attach client and the direct
PTY but not a remote pane, whose command carries no env exports at all. Docker
hands COLORTERM=truecolor to every mode, including the two the registry says
must unset it.
The changeset named six peer CLIs and there are seven: deepseek also exports
truecolor. A test beside the existing OpenCode assertion pins the new
behaviour, so a future registry edit cannot make the backgrounds vanish again
in silence.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude draws the user's own messages as a block of background color, and
inside a Codeman pane that block was invisible. tmux hands each pane
TERM=screen, which supports-color reads as 16 colors, and Claude's registry
entry deleted COLORTERM on top of that. Claude therefore quantized every RGB
color its theme asked for down to the basic palette, where rgb(55, 55, 55)
and every other dark background becomes ESC[40m, the terminal's own black.
Changing the color in a custom Claude theme moved nothing on screen.
Claude now exports COLORTERM=truecolor and unsets NO_COLOR, matching codex,
gemini, antigravity, pi, grok and omp. CLAUDECODE stays unset, because Claude
reads it as a signal that it is running nested inside itself. Both the tmux
session and the attach client read this one registry entry, so they cannot
disagree.
PR #3 introduced the unset in February, citing xterm.js#484 for the claim
that xterm.js mishandles truecolor. xterm.js closed that issue in April 2019,
Codeman now depends on @xterm/xterm 6, and TmuxManager already sets
terminal-overrides ",*:Tc" on its own tmux server, so 24-bit color reaches
the browser today for every CLI that asks for it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Rebased onto current master (the one real conflict was the import line
in docker-hosts.ts Ark0N flagged; kept both), then addressed every
point from the review:
**1. Rebase.** Done — this branch now sits on current upstream/master.
**2. Agent-image special cases are data now, not an id-keyed table
outside stock.ts.** `AGENT_IMAGE_SPECIAL_CASE_IDS`/`AGENT_IMAGE_SPECIAL_CASES`
are gone. `CliDiscovery.install.agentImageLayer?: { kind: 'dedicated';
reason: string }` is a field on the registry entry itself (pi,
deepseek), `reason` is required by schema.ts, both producers
(docker-hosts.ts and cli-catalog.mjs) filter on its presence instead
of an id, and the coverage test reads it from the generated catalogue.
Also added the npm-package-name validation to the TS producer, which
only the .mjs one had — same SAFE_PACKAGE regex, duplicated
(necessarily, one side can't import the other) and now pinned
byte-identical by a new parity test.
**3. Changeset said five, it's eight.** (Not nine — see the DeepSeek
point below, which changes the true count.) Reworded to state it
structurally rather than pin a number that will go stale again.
Then the four behavior-changing findings:
- **DeepSeek was offered as a normal install option but can't actually
drive a pane.** `npm install -g @deepseek-ai/dsh` installs the
launcher only; DeepSeek ships no profile that can run standalone.
The generator now emits an empty install command for any
`launcherProfile` entry, so install.sh's menu (which requires a
non-empty command) skips it and falls through to its docs URL hint
instead — matching what the old hand-written code did before this
PR replaced it.
- **wget-only hosts lost every automatic install, including the npm
ones that never needed curl.** The menu-building loop now filters
PER ENTRY (only a command starting with `curl ` is held back) rather
than wiping the whole menu when DOWNLOADER != curl.
- **The DISPLAY/TRUSTED split and the catalogue refresh didn't hold up
under review** (refresh's only real write was the label; it ran
before the Node existence check; its own eval-detection test was
tripped by the word "eval'd" in a comment). Dropped entirely per
your own recommendation — embedded catalogue only, no network
fetch, no second array. install-sh-invariants.test.ts now asserts
the refresh/DISPLAY machinery does not exist rather than testing its
internals.
The three take-or-leave items, applied:
- `dsh_banner_probe`'s bash 3.2 empty-array bug: `${runner[@]}` →
`${runner[@]+"${runner[@]}"}`. Verified live in a real `bash:3.2.57`
container with `timeout` removed from PATH — crashed before, clean
now, full `detect_all_clis` path exercised end to end.
- `docker-agent-image-coverage.test.ts` now anchors on each layer's
`<binary> --version` proof line instead of `Dockerfile.includes(binary)`,
which stayed true if a layer were deleted but its comment survived.
- Doc drift: docs/docker-cases.md (four → five, and now describes the
data field), docker/agent.Dockerfile's "other four CLIs" comment (no
longer a magic number — CLI_NPM_PACKAGES is generated and can grow),
CLAUDE.md's install.sh size (104KB → ~112KB) and its stale mention of
the now-dropped refresh.
Verified: tsc clean, prettier clean, the full targeted suite (142
tests across the 8 affected files) green, and the full `npm test` gate
diffed BY TEST NAME against a clean upstream/master baseline run on
this same machine — identical 201-name failure set both sides (168
tests / 67 files, all pre-existing Windows-environment noise: symlinks,
PTY spawning, POSIX permission bits — none of it touching anything
this PR changes), zero new failures either side of the diff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
Adds a "Consumers outside the server" section covering the two generated
artifacts, why each exists (neither install.sh nor a .mjs can import
TypeScript), what is deliberately NOT exported and why, the three-rule install
command trust boundary, and the bash 3.2 constraint with the offset/length
window shape it forces.
The adding-a-CLI checklist gains the regenerate step, since forgetting it is how
the installer would keep detecting the old set while the server offers the new
one — the drift this change removes, one level out.
docs/docker-cases.md gains how CLI_NPM_PACKAGES is derived, why it reads the
stock catalogue and not the merged registry, and a table of the four documented
Dockerfile special cases with their reasons. CLAUDE.md gains a command row and
names the generated block, the bash 3.2 rule and the trust boundary in its
install.sh paragraph.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12
docker/agent.Dockerfile hardcoded the four npm-published CLIs it installs, one
of the several lists that had to be kept in step with the registry by hand.
It now takes them as `ARG CLI_NPM_PACKAGES`, supplied by
scripts/build-agent-image.mjs from config/clis.stock.json, with the default set
to today's list so a bare `docker build` still produces the same image. The arg
is expanded unquoted because word splitting is what turns the list into several
arguments, which is exactly why every token is validated against
^[@A-Za-z0-9][@A-Za-z0-9/._-]*$ on the producing side; a package name carrying a
space or a metacharacter is refused rather than reaching the RUN line. Verified
by building the layer: four packages in, four arguments out, and the default
still applies with no arg.
The list is filtered on each entry's `enabled` flag — the field whose absence
was the maintainer's §3 finding, where a CLI shipping disabled still got baked
into every image. No stock entry is disabled today, so that assertion would pass
vacuously; a unit test feeds the pure helper a fabricated disabled entry so the
fix is covered now rather than the first time someone ships one.
⚠️ It reads the STOCK catalogue, never the merged registry. A user's
~/.codeman/clis.json must not change what is inside an image tagged
codeman/agent:base, or two machines holding that tag hold different images.
Four CLIs keep hand-written layers because the registry cannot describe what
makes them special: pi's --ignore-scripts, deepseek's pnpm companion and dsh-tui
profile, and the three standalone installers. Rather than extend the schema for
a Docker-only benefit, the coverage test requires each to carry a written reason
AND still be present, so an exclusion cannot quietly become an omission.
There are two producers of this command line and there have to be — the .mjs
cannot import TypeScript, and src/docker-hosts.ts builds the same argv for the
in-app auto-build — so a parity test pins them together, package list, arg pairs
and rendered argv. Their order is pinned too: a different order is a different
RUN string and so a needless cache miss between the two build paths.
docker/server.Dockerfile is deliberately NOT edited (PRs #373 and #377 both
modify it); its narrower list is asserted as a declared omission list instead, so
the divergence is reviewable without touching the file.
Also fixes the in-app hint at index.html, which the new coverage test caught
still omitting omp.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12
install.sh carried nine search-path arrays, eighteen near-identical
check_<cli>/get_<cli>_path functions, and three separately hand-maintained
enumerations of all nine CLIs. They had to agree and did not: upstream b6d0f1fa
is "wire OMP into install.sh's CLI detection (it had none)", and the section
comment above the roll-call named six of the nine.
All of it now reads the generated catalogue. `detect_all_clis` resolves every
CLI in one memoized pass into CLI_FOUND_PATH/CLI_FOUND_COUNT; `check_cli` and
`get_cli_path` replace the eighteen pairs; the roll-call, the "no AI CLI found"
gate and the closing reminder become loops. Probe order per CLI is unchanged and
`test/install-sh-detection-parity.test.ts` proves it against the literals
transcribed from the arrays this deletes.
Behaviour changes worth naming:
- The install menu is built from the catalogue, so it offers every enabled CLI
that is not installed and ships a command — five instead of two. Gemini had a
command in the registry and appeared in NO list in this script.
- Its labels are now the registry's ("Claude" rather than "Claude Code"), the
same trade PR A made for `codeman doctor` rows. A suffix map would just be the
hand-maintained list again.
- On a wget-only host the menu prints commands instead of running them. The
registry's commands call curl, whereas the two literals this replaces went
through download_to_stdout; rewriting curl to wget inside a string we are
about to execute is the wrong instinct.
The trust boundary is mechanical, not a promise: CLI_INSTALL_CMD_TRUSTED is
written only from the generated per-platform arrays and is the only thing ever
executed; CLI_INSTALL_CMD_DISPLAY is what the optional, opt-in refresh may
rewrite. The refresh warns on all three failure shapes — empty body, unparseable
content, failed fetch — which is the silent-degradation bug from the review, and
it parses with node into tab-separated records read by `read`, never eval.
Bash 3.2 throughout (macOS ships it): parallel indexed arrays, offset/length
windows instead of delimiters, no associative arrays, namerefs, mapfile or
here-strings. Verified by executing the script under a real bash 3.2 container,
which is also now a CI step alongside `bash -n` and a catalogue `--check` — the
empty-window case (`shell` has no binaries) is a runtime `set -u` abort that
`bash -n` cannot see. Running it that way caught `detect_os` being called inside
the platform loop: ten forks, and ten copies of one error, since a `die` inside
`$( )` can only exit the subshell.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12
Two consumers of the registry cannot import TypeScript: `install.sh`, which runs
via `curl | bash` before any checkout exists, and `scripts/build-agent-image.mjs`.
Both currently hand-maintain their own CLI lists, and both have already drifted.
`scripts/generate-cli-catalog.mts` (`npm run generate:cli-catalog`, plus a
`--check` mode) emits from `STOCK_CLIS`:
- `config/clis.stock.json` for the `.mjs` and the tests. It carries `enabled` —
the field the earlier attempt omitted, which is how a disabled CLI's npm
package still got baked into every agent image.
- a marker-delimited block inside `install.sh`, embedded rather than fetched.
The embedded copy is the FULL catalogue on purpose: the earlier design fetched
it and fell back to a hardcoded two-CLI list, degrading silently on an empty
response. There is no degraded mode to fall into now.
The block is bash 3.2 safe: parallel indexed arrays, no associative arrays, no
namerefs, no mapfile. Variable-length lists use OFFSET/LENGTH windows into one
flat array rather than a delimiter, so a $HOME containing a space needs no IFS
handling and `shell` (no binaries) gets length 0 and is never iterated. Search
paths are emitted dir-major, matching the probe order the hand-written arrays
use and `test/install-sh-detection-parity.test.ts` pins.
Only fields the two consumers need are exported. `launch`/`env`/`capabilities`/
`overlays` are spawn-time concerns the server alone interprets, and a test
asserts they never leak into the artifact.
`main()` sits behind an `isMainModule()` guard so the sync test can import the
renderers. Without it, importing the module would rewrite the artifacts as a
side effect of checking them — passing always, guarding never.
This commit adds the block; it does not yet delete the hand-written arrays, so
the detection pin keeps measuring both against each other.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12
PR B replaces nine hand-written `*_SEARCH_PATHS` arrays in install.sh with one
block generated from `STOCK_CLIS`. This lands FIRST, against the hand-written
arrays, so the replacement has something to be measured against.
The arrays are not uniform, which is why "generate them from the registry" is a
claim rather than an obvious truth: claude alone has `~/.claude/local`, opencode
alone has `~/go/bin`, opencode/codex/gemini/pi/omp carry `~/.bun/bin` while
dsh/grok/agy do not, and omp's `~/.omp/bin` sits second rather than first. A
generated list that silently narrows leaves a user with that CLI installed being
told no AI CLI was found — upstream `b6d0f1fa` is that bug, fixed for omp by
hand after it shipped.
The test asserts a three-way identity: the pinned literals equal what install.sh
contains today, AND equal `searchDirs x binaries` from the registry, dir-major so
the probe ORDER is pinned too and not just the set. Both halves were verified to
fail independently — dropping one path from install.sh fails the first, changing
one `searchDirs` entry fails the second — because a pin that cannot fail is
worse than no pin. A fourth case asserts every stock CLI with a binary is
covered, which is the omp bug restated so it cannot recur silently.
No production code changes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12
CI caught a real regression: DEEPSEEK_API_KEY was added to deepseek's
privilegedEnvKeys alongside DEEPSEEK_BASE_URL on the theory that "the pair
travels together," but that contradicts the documented and tested design
(clampEnvOverridesForOwner()'s own docstring in session-routes.ts) — a
non-granted owner supplying their OWN DeepSeek key removes privilege
rather than granting it, since the exfiltration vector is the BASE URL
(which redirects the server's own forwarded key to a foreign host), not
the key itself. Removed it from the list; test/deepseek-mode.test.ts's
existing two clamp tests now pass again.
Also swapped that test's "unrelated override" example off CODEX_HOME,
which the earlier commit in this same PR legitimately made privileged
(closing a real pre-existing gap, documented in PR.md) — so it stopped
being a valid "unrelated" example the moment that fix landed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.
Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.
Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:
- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
undocumented GATEWAY AuthType gemini-cli selects once
GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
tried
- deepseek: reaches the server (env vars are read) but gets a consistent
HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)
Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).
deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its
native cloud backend, for a given session. Covers local hardware (llama.cpp,
Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter).
Off by default (customModelEndpointsEnabled, synced, default OFF).
- Registry: capabilities.customModelInjection per CLI entry (env /
configContentEnv / configDir / unsupported kinds)
- Pure injection builder (custom-model-injection.ts) turning an endpoint +
model id into the real env vars / config content per CLI
- Endpoint store + CRUD routes (custom-model-hosts.ts,
custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded
- Session integration: Session.setCustomModel()/restartCli()
(POST /api/sessions/:id/custom-model), reusing the existing
respawn-pane -k primitive to restart the CLI process with new env
- Multi-user hardening: every new redirect-capable env var added to its
CLI's privilegedEnvKeys, closing a pre-existing gap where several were
already reachable via the generic envOverrides field's prefix allowlist
- Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries
against a real endpoint outside the web UI, independent of tmux/sessions
- Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying
every CLI's injected values through a real HTTP shape
Real end-to-end validation against a live llama-swap server (inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries) found and
fixed three real bugs before they shipped:
- Codex's config.toml schema was wrong ([model].default table instead of
a top-level model string + [model_providers.custom]); fixing it then
surfaced a genuine, documented protocol incompatibility (Codex only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement)
- Claude Code's async session-title-generation call validates
ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and
hangs the whole -p invocation on an unrecognized name; documented for
chunk 6, worked around in the standalone script only (--bare is NOT
safe for a real interactive session, which needs hooks)
- The discovery route's authStyle: 'both' option (send both Authorization
and api-key headers) reliably hung a real server; removed the option
entirely rather than just changing the default
Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see
PR.md and deployment_plan.md for the full chunk breakdown and confidence
table.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
/opt/codeman-cli is chowned to PUID:PGID once, at image build time, from
the PUID/PGID build args. That bake only happens when the image is
actually rebuilt (`docker compose up --build`, which Start-Codeman.sh
always does) — a deployment that runs the compose file directly instead
(Unraid's Compose Manager, a native systemd unit, any plain
`docker compose up`/`restart`) can change PUID/PGID in .env and restart
without ever rebuilding. The container then runs as the NEW uid via
entrypoint's setpriv (Linux needs no /etc/passwd entry to setuid to an
arbitrary number) while the CLI directory is still owned by the OLD one
baked into the image layer — silently breaking the self-update-a-CLI-
in-place fix that directory exists for.
Unlike HOME/CODEMAN_CASES_PATH, this one is pure image content Codeman
itself populated, never host data that might legitimately belong to
someone else, so there is no ownership to be careful about — it is
always correct for it to be owned by whoever the container is about to
run as. Re-assert it unconditionally on every start.
Verified live: built an image with PUID=99/PGID=100, ran it with
PUID=1234/PGID=4321 (no rebuild, simulating a changed .env restarted
directly), confirmed /opt/codeman-cli ends up 1234:4321-owned and is
genuinely writable by the running process.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
Two real bugs the review caught, both verified live against a real
build on the Unraid host:
1. entrypoint.sh's chown fired on ANY ownership mismatch, not just a
directory the daemon itself created root-owned. A host tree
legitimately owned by some other account - an existing
CODEMAN_CASES_PATH the README already allows pointing at a normal
projects directory, or appdata under a different PUID/PGID
convention than the one in use - got silently recursively re-owned
with one log line to explain it. Now gated on the target actually
being root-owned; anything else is a clean refusal naming the
directory, its owner, and PUID/PGID. Start-Codeman.sh also now
pre-creates CODEMAN_CASES_PATH the same way it already did
CODEMAN_APPDATA_PATH, so Compose never has to materialise a missing
bind source as root in the first place - the in-container chown
becomes a safety net, not the primary mechanism.
2. The CLI-update chown (chown -R .../node_modules /usr/local/bin)
handed the runtime account write access to entrypoint.sh itself
(root-owned, executed as root on every container start with
CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node binary - owning the
DIRECTORY is enough to rename it aside and drop a replacement, which
would let a compromised session arrange for its own script to run
as root at the next restart. The four CLIs now install into a
dedicated /opt/codeman-cli prefix (NPM_CONFIG_PREFIX); only that
directory is chowned, /usr/local stays root-owned throughout.
Smaller fixes from the same review:
- Start-Codeman.sh's volume-refresh label filter wasn't project-scoped:
a second Compose stack on the same host sharing the `codeman-dist`
volume KEY could have had ITS volume deleted. Added a
com.docker.compose.project filter, resolved from this stack's own
`compose config --format json`.
- Override-file precedence was backwards (checked .yaml before .yml;
Compose actually prefers .yml) - swapped, plus a warning when both
exist.
- entrypoint.sh's setpriv now also passes --bounding-set -all, so
CapBnd actually clears post-drop rather than just CapPrm/CapEff.
- A comment on git_head_commit() noting it returns nothing for a
worktree checkout (.git as a file), consistent with the script's
existing -d .git convention elsewhere.
- Doc drift: CLAUDE.md's Docker Compose section still described the
old pre-created-and-chowned-by-hand model and didn't mention the
root-then-drop entrypoint; the state-files list was missing
docker-build-source.json; docs/docker-compose.md and
docker/.env.example still had the pre-rename `Coding/codeman` path
in one place each.
Verified end to end against a real build on the Unraid host: a
root-owned bind source is corrected as before; a directory owned by
neither root nor PUID:PGID is refused rather than silently rewritten;
a correctly-owned directory is left alone entirely; the four CLIs
resolve via PATH from /opt/codeman-cli while /usr/local/bin,
/usr/local/lib/node_modules and entrypoint.sh itself stay root-owned;
CapBnd is fully cleared post-drop.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded
from the image only while empty, so a rebuilt image's fresh dist/
node_modules sat unused behind old volume content until something
cleared it. The in-app self-updater never hit this (it rebuilds INSIDE
the running container, into the very volume already in use), but a
`docker compose build` triggered from outside it — Start-Codeman.sh,
after a manual `git pull` — did: the container came back up looking
unchanged, serving stale compiled routes against current source.
Start-Codeman.sh now compares the checkout's HEAD commit and
package-lock.json hash against a recorded marker
(docker-build-source.json) and clears just the affected volume(s)
before its own --build when either moved.
The in-place self-update path writes that same marker after a
successful build, so the two mechanisms agree on what the volumes
currently reflect — without it, the next plain Start-Codeman.sh run
would see the HEAD self-update just checked out, not recognise it as
already accounted for, and wipe the volumes self-update just correctly
rebuilt right back to the older baked image.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
The four CLIs (claude, gemini, codex, opencode) are npm-installed
globally as root during the image build, before the unprivileged
runtime account exists. A session running as that account (e.g. a
codex-mode terminal) then hits EACCES the moment it tries to update
one in place, because npm renames the old package directory aside
before installing the new one, which needs write access to the
parent (/usr/local/lib/node_modules), not just the target package.
Chown that tree plus /usr/local/bin's CLI symlinks to PUID:PGID in
the same step that creates/renames the runtime account.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
CODEMAN_ALLOWED_HOSTS is a real, documented application setting (the Host-
header allowlist in network-auth-policy.ts), but docker-compose.yaml does not
forward it from .env into the container - Compose only passes through
variables explicitly listed under environment:, and this is not one of them.
Set without that passthrough, any request through a reverse proxy is rejected
with 403 Forbidden: host not allowed before it reaches any handler, and
nothing in the Docker deployment docs said why.
Document the variable and the override needed to forward it, using the
Local customisation mechanism already described above it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
CODEMAN_RUNTIME_USER defaulted to `opencode`, which no longer matches the
project and is confusing in a deployment whose every other identifier is
codeman. Rename the default in .env.example and in the Dockerfile ARG that
mirrors it, and correct the example comment that referred to
/home/opencode/codeman-cases.
Also drop the `Coding/` component from the example application-data path.
CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH now suggest /mnt/user/appdata/codeman
and its codeman-cases child, matching the account name and removing a directory
level that meant nothing outside the original author's host. README.md is
updated to match, including the chown example.
The npm package `opencode-ai` and the references to the OpenCode CLI are
deliberately left alone: those name a different tool, not this account.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Naming a Compose file with -f disables Compose's automatic discovery of the
override file, so Start-Codeman.sh silently ignored docker-compose.override.yml.
Any local customisation placed in the conventional override file was dropped
without warning, and the only way to notice was to inspect the running
container.
Collect the -f arguments into an array, append the override file when one is
present, and reuse that array for the final launch so the two cannot drift
apart again. Both .yml and .yaml are checked, in Compose's own precedence
order, and the chosen file is reported on startup.
Document the override file in docker/README.md, including the two things that
are easy to get wrong: it is ignored when -f is passed without naming it, and
it cannot remove a key such as ports, which Compose concatenates. Add the
override file to .gitignore.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Compose binds CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH from the host. When
either path does not exist yet - a first run, a cleared application-data
directory, a restored backup - the Docker daemon creates it owned by root. The
server runs unprivileged as CODEMAN_RUNTIME_USER, so it cannot create its own
state directory, and the container restarts forever on:
Failed to start web server: EACCES: permission denied, mkdir '/home/<user>/.codeman'
Start-Codeman.sh already worked around this by preparing the directory on the
host, so the failure only appears when Compose is run directly, which the README
documents as a supported path.
Add docker/entrypoint.sh, which starts as root, corrects the ownership of both
bind mounts, then drops to PUID:PGID with setpriv. The Dockerfile's USER
instruction is replaced by that entrypoint and CMD is unchanged.
docker-compose.yaml adds back only the four capabilities the chown and the
privilege drop require, so cap_drop: ALL continues to remove everything else.
Two guards keep existing deployments working:
- A container started with an explicit `user:` is left alone. The entrypoint
execs straight through, with no elevation and no chown.
- A chown that fails is a warning, not an error. Bind mounts backed by NFS,
CIFS or a rootless daemon can refuse chown while remaining perfectly
writable, and those deployments must keep starting.
PUID and PGID are also exported as runtime environment defaults so the image
behaves correctly when run without Compose, rather than depending on build args
alone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A Fable 5.1 reviewer read the whole release diff against 1.26.2 and returned
SHIP WITH FIXES. These are its findings, verified before acting on each.
**The changelog advertised a feature the code refuses (major).** The #401
changeset and docs/web-tabs.md both listed `*.localhost` in the loopback set.
The follow-up in 02b0e278 moved it out of the auto-route set on security
grounds and updated CLAUDE.md but neither of those, and that changeset becomes
the 1.27.0 CHANGELOG entry: a user would have read the release notes, tapped
`http://app.localhost:3000/` on a phone and got a connection error from a
documented feature. Both corrected, and the user guide now says why it is
excluded and that adding such a dashboard by hand still works.
**Dictation delivered its text twice (minor, #388).** `keydownSnapshot` started
`null`, so `keydownSnapshot ?? canonicalCount` at the input event read a counter
xterm had ALREADY bumped: on a fresh page load with no keydown yet, xterm's own
capture listener forwards the `insertText` itself (it is not gated behind a
keydown), then the snapshot equals the bumped count, `count > snapshot` is
false, and the controller emits the same text again. Reproduced directly
against the module: it emitted `hello` for input xterm had already delivered.
A `0` baseline restores that file's own invariant, that a missed recovery is
acceptable and a duplicated keystroke is not. Two regression tests, covering
both the xterm-already-delivered and genuinely-dropped halves.
**The sorted rail's arrow-key walk followed the DOM (minor).** `_tabKeydownHandler`
steps `querySelectorAll` order, which is `sessionOrder`, while a sorted rail
paints its rows with the flex `order` property, so ArrowDown from the top card
landed wherever that session happened to sit in the tab order. It now sorts its
node list by the COMPUTED order first: computed rather than inline, because web
tabs take their `order: 9999` from CSS and would otherwise read as 0 and lead
the walk. This is the one place that follows the paint; the Alt+N badge, the
drag model and the filter all still deliberately read the DOM.
**A trusted dashboard was auto-reused by a tapped link (minor, #401).** The
reuse loop skipped `managed` and direct-mode records but not `trusted`. A
trusted frame is mounted with `allow-same-origin`, i.e. on Codeman's origin
with the user's cookie, and these links come from agent output, which is the
threat model the loopback allowlist was just narrowed for. An agent that can
write into the dev server's tree could print a path that one tap opens inside
that privileged frame. Excluded from auto-reuse, with a test; opening it from
the Run dropdown is still an explicit action and unchanged.
**Two documentation claims that were no longer true.** CLAUDE.md said
test/location-overlay-commands.test.ts pins every remote pane command, but
remote claude and remote omp now have their own arm in `buildRemoteLaunchCommand`
and never reach `defaultRemoteCommandForMode`, which is what that test asserts,
so it pins nothing for them and changing either arm will not fail it. Named the
real pins instead. Also documented the arrow-key-walk exception in the rail
paragraph.
Left as follow-ups, deliberately: `POST /api/webviews` does not dedupe by URL
server-side, so two devices tapping one link concurrently can still save two
dashboards for one origin (pre-existing endpoint behaviour that #401 makes
reachable by a tap), and the location-overlay golden should assert the real
remote claude/omp commands rather than a branch neither reaches.
Full gate green: 359 files, 6869 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The vertical rail lists exactly the sessions both home screens list, so it now
answers their question the same way instead of showing the raw tab order.
Order: new per-device `tabRailSort` (App Settings -> Appearance -> Tabs ->
Vertical Rail Order, default "By activity"). It runs `CodemanSessionOrder` over
rows classified by `_mobileOverviewState`, i.e. literally the home screens'
comparator, `lastSubmitAt`-anchored running group included.
It is applied as the flex `order` property, never by reordering the DOM.
`#sessionTabs` stays in `sessionOrder`, which is what keeps the Alt+N badge
honest (it names a shortcut, not a row position, so it deliberately does NOT
run 1,2,3 down a sorted rail), and keeps drag-and-drop, the arrow-key walk, the
sidebar filter and `_scrollActiveTabIntoView()` all reading the list they
always read. A session changing state then moves one inline style instead of
forcing the full rebuild that would restart every card's animation on every SSE
tick. The incremental render path re-applies it, since a state flip adds no tab
and never reaches the full rebuild, and an empty string is what clears it when
sorting stops. Web tabs are pinned past the cards by a CSS `order: 9999`, since
`renderWebviewTabs()` emits the same markup for every layout and the flex
default of 0 would interleave them. Drag is switched off while sorting (the
drop rewrites `sessionOrder` correctly and the sort puts the card straight
back, so the affordance would be a lie); 'manual' is the way back.
Cards: detailed rail rows become bordered cards on `--bg-card`, with the stamps
line on its own full-width row and the pill at its right end. The state dot
goes 6px to 9px, keeps its orbiting ring while working and gains the green
halo; idle mutes toward `--text-muted` as the home rail does. Needs/error/
waiting reuse `home-sessions-blink-red`/`-yellow` rather than a second copy.
These card rules are RAIL-SCOPED and deliberately absent from the comma-grouped
selectors that carry both vertical surfaces: the rail is an occasional,
resizable list you scan, while the sidebar is a permanently-docked nav column
where 20 stacked cards read as a wall. Every state-dot rule also excludes
`.tab-alert-action`/`.tab-alert-idle` by hand, because those alert rules are
only (0,3,0) and these are (0,5,1)+.
Lines: the lineage bracket already drew in the rail, but its track sat 6px from
the left edge, so half of its 11px outer glow was clipped by the window frame
and it read as a thread pinned to the frame. It now runs at 10px, mid-channel
in the gutter the rail already reserves.
Tests: test/tab-rail-order.test.ts (17) drives the real `isTabRailSorted()` and
`_tabRailSortOrder()` out of app.js, covering the row model (a WORKING row
ranked by `lastSubmitAt`, which would otherwise rank every running turn as
freshly started and fail no rendering test), the Alt+N badge, and the opt-out.
Verified in Chromium across sorted/manual/simple/header-strip/sidebar with the
setting flipped at runtime: no page errors, and the header strip and sidebar
render byte-identically to before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Each item is from the pre-merge review of the PR it names, applied on master
rather than by pushing to a contributor branch.
#400 (response viewer, shenlvkang-collab)
- The brief view opened at `scrollTop = 0`, right when it was a single card
holding the last row. Now that it renders the whole turn, the top is the
turn's first narration line and the answer can be screens below it, while
loadFullContext already scrolls to the bottom of the same turn. A multi-row
turn now opens at its newest text; a single card still opens at the top.
#401 (loopback links as web tabs, shenlvkang-collab)
- Drop `*.localhost` from the auto-route set. Every other member is an address
literal that can only mean this box; a `*.localhost` DNS name is not one, and
a resolver with a search domain retries `evil.localhost` as
`evil.localhost.<search domain>`. The link source is agent-written terminal
output, so that set is the whole confinement on a tap that makes Codeman
fetch a URL server-side and persist it. The page-side test stays broader
(`isOnBoxHostname`), where a false positive only declines to proxy.
- A link to the origin root navigated nothing: the path was flattened to '',
which openWebview reads as "no deep link", leaving an open frame where it was.
- `this.webviews` being set does not mean it is loaded. initWebviews() assigns a
truthy empty map and only then awaits the list, so a tap during page load
found nothing to reuse and POSTed a duplicate record. Join the in-flight
refresh instead.
- One dashboard per dev server rather than per host spelling, which is what the
method's own comment already promised.
- Toast on the auto-create: it writes webviews.json, broadcasts over SSE and
adds a Run-dropdown row on every signed-in device, with a new tab as its only
previous signal.
#362 (remote omp continuation, timkjr)
- Accept the allowlisted `mode === 'omp'` arm as-is; a blanket registry render
would hand deepseek a locally-resolved --profile and bypass claude's own
overlay. A registry-declared switch is the follow-up if a third mode needs it.
- Revert the whole-file Prettier reformat of docs/remote-sessions.md (docs/ is
hand-formatted and outside `npm run format`), keeping only the two new
sections.
- Correct three stale passages: architecture-invariants' `exec claude
--dangerously-skip-permissions`, the `exec <cli>` paragraph (claude and omp
now have their own arms, and the claude pane's PID is the login shell), and
omp-integration's `-c 'omp'`. RemoteCommandMode gains deepseek and omp.
- Add the missing `_maybeCaptureOmpSessionId` remote-guard test; the sibling
guard in `_pinOmpRespawnId` had one and this path runs earlier, on the first
idle turn.
#388 (keyCode 229 recovery, aakhter)
- Gate notifyCanonicalData on shouldSuppressTerminalQueryResponse and
isTerminalFocusOrMouseReport. onData also carries the DA/DSR/CPR/OSC replies
xterm answers during Ink redraws and its SGR mouse and focus reports; any of
those landing between the keydown and the candidate's resolution was read as
"xterm spoke for this keystroke", standing the recovery down and leaving the
character dropped, worst on a busy agent pane. Reached through
window.CodemanTerminalInput: the predicates live in a module IIFE that closes
long before this call site, so bare references would throw into the
surrounding try/catch and stop the notify from ever running.
Every fix has a test that fails without it (verified by reverting each).
Full gate green on the combined tree: 358 files, 6849 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
planUsageChipEnabled() (settings-ui.js) shows the header chip and the App
Settings checkbox as already ON whenever showPlanUsageLimits has never been
set — a discoverability default from 1.9.3. readPlanUsageTelemetryEnabled()
(hooks-config.ts) deliberately treats an absent key as "no telemetry" — a
privacy default, pinned by its own unit tests (never POST usage data
without an explicit persisted yes). Nothing reconciled those two
independent guesses, so a fresh install showed a checked box that silently
collected nothing until the user opened Settings and hit Save at least
once.
Verified live: an install that had never touched this setting had no
showPlanUsageLimits key in settings.json at all, and its running Claude
process's argv carried no --settings flag — zero telemetry ever collected
despite the chip rendering as enabled.
GET /api/settings now persists the resolved default (true) the first time
the key is truly absent — not explicit false — so "chip visible" and
"telemetry collected" become the same fact. readPlanUsageTelemetryEnabled's
own absent-means-false contract is untouched; after this runs once the key
is never absent again, so that branch stays correct in isolation while
being unreachable in practice for any install that has ever called this
route. An explicit false set afterward is respected forever.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ctrl+Z (SIGTSTP) suspends the foreground job on the pane's tty. In a plain
shell session that's the user's own job-control tool (suspend, fg back),
but in claude/omp/pi/codex/etc. sessions it stops an unattended agent loop
dead with no visible output — the same failure shape as an XOFF freeze,
just via job control instead of tty flow control. Ink-based TUIs usually
run in raw mode (ISIG off) where ^Z is inert, but that only holds once the
CLI is actually running and stays in raw mode; it's live at the shell
prompt before launch and during any raw-mode toggle.
Swallow it client-side in attachCustomKeyEventHandler, mode-gated so shell
sessions keep normal job control, mirroring the existing Ctrl+V/Ctrl+Backspace
interception pattern in the same handler. Case-insensitive key match (Caps
Lock flips ev.key to 'Z' without setting shiftKey, so a plain === 'z' check
let the exact suspend keystroke this exists to catch slip through).
Also cover subagent/teammate terminal windows (panels-ui.js's
initTeammateTerminal), which render a separate xterm instance with no
custom key handler at all and are always running an agent CLI — never a
shell — so the trap there is unconditional.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A dashboard served through a web tab saw `/webview/<cap>/` as its
`location.pathname`, and no app has a route for that: a React Router, Vue
Router or Vite dev-server page painted its HTML and CSS and then replaced
them with its own "page not found" the moment its script ran (reproduced
with a minimal history-routed page).
The proxy's runtime shim now rewrites the history entry to the path the
page would see on its own origin, before any page script runs. The base
element still resolves relative URLs inside the prefix and every root-
absolute sink is rewritten back into it, so only what the page READS
changes. With the document URL masked the Referer-keyed 404 rescue can no
longer help a request the shim misses, so the remaining URL-taking entry
points (`Worker`, `SharedWorker`, `navigator.sendBeacon`, `window.open`)
are covered by the shim as well.
A navigation the page starts itself afterwards — `location.reload()`
(a dev server's full-reload HMR), a root-absolute `location.href` — lands
on Codeman's root with no capability anywhere: no prefix in the path, no
cookie in an opaque-origin frame, a Referer naming the masked page. It is
recognised by shape (a top-level iframe navigation asking for HTML, for a
path Codeman does not serve) and answered with a static page whose only
script posts `{type:'codeman:webview-lost', path}` to the parent; the tab
that owns the frame (matched by `event.source`, never by the payload)
remounts it inside the prefix at that path, bounded per frame. The
unauthenticated form is answered in the auth middleware before the
credential checks, so a dev server that reloads on every save cannot
rate-limit its own user out of Codeman; the authenticated form (Basic
auth, trusted mode) is answered by the 404 handler.
Verified end to end against a history-routed page: boots on `/`, its
API call succeeds, a reload inside the frame comes back routed on the
path it had pushed, `location.href = '/about'` comes back on `/about`,
and a deep link opens on its path.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
An agent prints `http://localhost:5173/` (a dev server, a preview it just
served) and the user taps it on a phone. That address only exists on the
Codeman box, so the link was a guaranteed connection error from any other
device — while the web-tab proxy fetches from the server, where it works.
A loopback link (`localhost`, `*.localhost`, 127/8, 0.0.0.0, ::1) activated
in the terminal or clicked in the Response Viewer now opens as a proxied
web tab whenever the Codeman page itself is not on that box. A saved
proxied dashboard on the same origin is reused, with the link's own path,
query and fragment opened inside it (a mounted frame is navigated, not torn
down, so its state survives); otherwise one is saved under its host:port,
sandboxed like any other web tab, so it is in the Run dropdown next time.
Only loopback is routed this way. A LAN or tailnet address may well be
reachable from the device (a VPN, the same Wi-Fi) and a direct open is the
cheaper, richer path, so those keep opening in a new browser tab; on the
box itself every link opens directly. The terminal link provider and the
viewer's click handler consult one hook and fall through to their existing
behaviour when it declines.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
The eye button rendered `data.text`, which is one row: the last assistant
row of the transcript. A Claude answer is a median of 3 model messages
(p90 11) split around tool calls, so the brief view usually showed the tail
of an answer ("Done.", "Let me look.") and the substance appeared only after
More. The full view was fine, which is why the brief one read as broken by
comparison.
The brief view now asks `?context=turn`. The reader answers with the
assistant messages of the last ANSWERED turn (`selectLastAnsweredTurn`: the
highest `turn` that has an assistant row, so a prompt queued after the
answer does not blank the view) and the frontend renders them exactly as
the full view renders that turn: one badge, then continuation segments,
gated on the numeric `turn` as before.
`data.text` is unchanged in every context — still the last assistant row,
never `messages.at(-1)` — because agent pollers hash it. Readers that emit
no turns (Codex, the pane parser, DeepSeek, an older server) return `text`
only for `context=turn`, and the brief view keeps its single card for them.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
The picker's current-folder line was a read-only breadcrumb, so reaching a
deep folder meant tapping through every level, and the listing was fixed to
name order, so the file an agent had just written was somewhere in a
500-entry list.
The current folder is now an editable field: Enter or Go jumps there, a full
file path lands in its folder with that file selected, and a path that does
not resolve keeps the listing you had and says so, instead of the reset to
the root that a stale initialPath gets. A Sort control orders the listing by
name or modified time in either direction, folders always first, and the
choice is remembered per device like the hidden toggle. Each entry shows a
compact modified time (time of day today, month-day this year, else the
date).
GET /api/filesystem/browse stamps every entry with mtimeMs to make that
possible; the stat that already fetched a file's size now serves both, so
it is still one stat per entry. Entries without an mtime (an older server,
the in-container listing) sort after dated ones and then by name.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
An iOS-style focus pull: the thing arrives out of focus and the blur fades
off it as the opacity comes up. Opacity leads the blur (full opacity around
45%, blur still lifting), which is what separates it from a cross-fade.
Ships on tabs (440ms), agent windows (560ms), the terminal pane (520ms) and
connection lines (380ms), plus a `Soft focus` theme that sets all four.
Default stays `legacy`, so an untouched install is unchanged.
The terminal pane is the one surface that cannot blur itself the documented
way, and `blur` takes a deliberate exception to the "never a filter on
.terminal-container" rule. Every alternative was measured against a live
xterm and does not work: a backdrop-filter veil on ::before blurs perfectly
while STATIC, and Chrome silently drops the backdrop the moment ANY
animation runs on that pseudo-element (the veil computes blur(15.3px) while
the text behind it stays razor sharp); driving the radius from rAF buys the
same full-screen blur per frame plus main-thread work. The cost the rule
exists to avoid is inherent to blurring a terminal, so the style buys it
knowingly: opt-in, off by default, one ~520ms run per session open, class
straight back off, will-change still unset. Worst-case price, headless
SwiftShader with no GPU: frame deltas 16.7ms -> 33.3ms for the run, against
16.7ms flat for `fade`. cols x rows measured unchanged at 178x38 before,
during and after, so FitAddon never sees it.
The line entrance animates `filter` too, where each line already carried
its glow. Both kinds now hold it in --line-glow and both keyframes say
`blur(N) var(--line-glow)`, so the function lists match and interpolate
instead of the glow vanishing for the run and popping back (a lineage
line's glow is a different colour, set per element). Its 100% frame omits
`opacity` on purpose so the endpoint comes from the element's own resting
value: 0.9 subagent, 0.72 lineage, 0.95 working.
test/entrance-animations.test.ts is a new static guard over the whole
feature, not just this style: the rule -> keyframes -> theme-option chain a
style silently does nothing without, the terminal's paint-only property
allowlist (the FitAddon rule), the --line-glow contract, and reduced-motion
coverage. Mutation-checked both ways.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The 50MB cap on file-raw, the attachment /raw route and /api/download was
memory protection for a `readFile()` that no longer exists: file-raw and
/raw were rewritten to stream through `sendFileBody()` and answer Range
requests, so size costs a read stream rather than RSS (measured: a 600MB
download moved peak RSS by ~37MB). All the cap still did was refuse
legitimate downloads of build artifacts, videos and archives.
It is now MAX_FILE_DOWNLOAD_BYTES in config/buffer-limits.ts, default 2GB,
env CODEMAN_MAX_DOWNLOAD_BYTES, 0 = unlimited. `parseByteLimitEnv()` is
separate from the `parseInt(...) || default` idiom used elsewhere in that
file precisely because that idiom reads 0 as falsy and would silently
restore the default for the one value that means "no limit".
/api/download was the last route that really did buffer the whole file. It
now shares sendFileBody() with the other two, so it streams, advertises
Accept-Ranges, and is resumable. Its Content-Disposition also goes through
buildContentDisposition() rather than raw interpolation.
Refusals move from 400 to 413 across all three, which is the correct status
for the case; with the cap at 2GB it is a path almost nothing reaches now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review of the previous commit found that waiting for the font does not, on its
own, do anything.
`FitAddon.proposeDimensions()` measures nothing — it divides the container by a
CACHED cell size, and xterm refreshes that cache only from `open()`, from a
resize that actually changed the grid, and on a device-pixel-ratio change.
Nothing in it listens for font loading. So a fit that runs after the font
arrives can still divide by the fallback cell, propose the grid it already has,
and short-circuit before anything re-measures. The wait now ends by calling
`_charSizeService.measure()` itself, which is the step that makes the following
fit see the real font. Private API, as FitAddon's own dependency on `_core` is,
and guarded because a terminal can be disposed mid-wait.
The wait was also unbounded, and it sat behind the buffer-load gate. Neither
`FontFaceSet.load()` nor `FontFaceSet.ready` has a deadline, so a font request
that never settled left the tab spinning with live output queued behind it —
permanently, and on every session, since they share one promise. The comment
claimed the opposite ("a font that never loads must not block the terminal, so
this always resolves"), which was true of the per-face loads and false of
`ready`. It is now raced against TERMINAL_FONT_WAIT_MS, and the await moved
ahead of `_beginBufferLoad` so a slow font cannot hold output back at all —
which also removes the stale-select interaction with `_restoringFlushedState`,
since that flag is not yet set when the wait runs.
The awaited set no longer includes faces that cannot move the measured cell.
The bundled symbols font is ~1.2MB of private-use-area glyphs and xterm
measures `W`, so awaiting it put a megabyte in front of the first frame for
nothing; the generic families match no FontFace at all.
A runtime font change had the same race the boot-time one did:
applyTerminalFontFamily wrote the new family and fit on the next line, against
a family the browser might not have loaded. It now re-arms the wait and fits
again when it settles.
The claim that this could not be tested was wrong: the repo's vm harness
reaches both halves. The new suite pins the family filter, the forced
re-measure, the deadline, a rejecting load, a browser with no font API, and a
terminal disposed mid-wait — plus the four ordering properties in
selectSession, including that iOS Safari's synchronous focus still precedes the
first await. Each assertion was checked by reverting its fix.
Also corrects the docstring's reason for calling `document.fonts.load` (the
stylesheet is render-blocking and long parsed by then; the real reason is that
the WebGL renderer rasterises through a canvas atlas, and canvas text never
triggers a CSS font fetch), restores the JSDoc block the previous commit
displaced from getTerminalDimensions, and fixes a comment that described the
first fit as already having run when the mobile-Safari branch defers it.
Review of the previous commit found the guard inverted: the three skips keyed
on `?full=1`, which is only what the client asked for. When the capture comes
back null — ENOBUFS, a timeout, a vanished pane, or a session with no mux at
all — the reply falls back to the byte history, which IS a stream of
successive frames and still needs stripping. Gating on the request returned it
whole: measured at 82KB against 4KB for the same buffer without `full=1`. A
direct-PTY session takes that path on every first selection, not only during
an outage. The skips now key on `isFullCapture`, meaning a capture arrived.
Three further defects the same review surfaced, all on this path:
Keeping the trailing rows is only sound when a cursor move follows to count
back up from them. On the two branches where the cursor query fails there is
no move, so the caret was left at the bottom of the pane — worse than before.
The cursor is now read first and settles both decisions together.
The move is relative rather than absolute. `CUP` numbers rows from the top of
the browser's screen, so it is only right while the browser's row count equals
the pane's, and `resizeWindow` does not wait for tmux, so a capture can be
taken before a requested resize applies. Measured against real tmux with a
browser four rows shorter than the pane: the absolute move lands on a blank
row, the relative one lands on the caret's row.
An all-blank pane no longer reads as content. Retaining trailing rows and
appending a move made it non-empty, and the caller treats non-empty as "replay
this", so a blank screen would have replaced real history — the downgrade
`_replayWouldShrinkBuffer` refuses, arriving from the server side where that
guard cannot see it.
The documentation claimed one line per screen row. `-J` joins a hard-wrapped
row into its logical line, so that is false whenever any row wrapped: measured
at 10 lines for a 12-row pane. Both entries now say what actually holds, and
the stale "NOT repositioned" contract in the mux interface is updated too.
Tests: the byte-history fallback is stripped, an empty capture leaves history
intact, and the extracted helpers are unit-tested directly rather than through
source-text matching. The slice window in the capture test is bounded at the
next method, having overrun into its neighbours.
Review of the previous commit found four defects in it.
The guard sat above the local fit, so it suppressed a reflow as well as the
server write. tab-rail-resize performs its single settle-time refit through
sendResize and has no fallback for a truthy activeSessionId, so dragging the
rail stopped reflowing a detached session's terminal in the dashboard. The
mobile-keyboard guard fourteen lines below already draws the line correctly —
withhold the send, never the reflow — and the guard now sits after the fit.
_lastResizeDims is one value for the whole window, and both guards skip
updating it, so while a popup owns a session that value no longer describes
the PTY. _redock repaired it only for the active session. Pop out A, switch to
B, close the popup: selecting A later found unchanged dimensions, returned
"unchanged", and selectSession skipped its 400ms redraw wait — while the
server, comparing against the real pane, did resize and did raise SIGWINCH, so
the fetch painted the pre-redraw frame. _redock now clears the record on every
path, active or not.
_redock could also fire a resize for a session already gone: _onSessionDeleted
redocks before cleanup, so the id can be dead and the request is a guaranteed
404. It now checks the session still exists.
restoreTerminalSize — the header's redraw button and Ctrl+Shift+R — silently
did nothing for a detached session while still reporting success with
dimensions nothing was set to. It now says the session is sized by its own
window, where the same button works.
The `force` comment claimed a client-side dedupe that does not exist; the
deduplication is server-side against the real pane. Corrected to say what the
flag actually buys. The _redock doc comment now records that the function
writes to the server and is not idempotent.
Tests: _redock was the untested half and is the half three of these defects
sit in. It now has coverage for clearing the stale record on both the active
and inactive paths, re-asserting only for the session being shown, and staying
silent for a deleted session. The existing sendResize test now asserts the
local fit still runs.
Opening a session could render its frame with characters spliced into each
other, as though two frames were overlaid — a status-line fragment landing
in the middle of a file path, for instance. Resizing the browser window
cleared it.
The first fit runs while the browser is still painting with a fallback font.
A cell measured against that font has a different width and height from one
measured against the terminal font, so the fit produces the wrong column and
row count. Codeman sizes the pane to it and replays the capture. When the
font finishes loading the measurement changes, the pane is resized a second
time, and the CLI repaints for a shape that does not match the frame already
on screen. Its later partial updates then land on the wrong rows.
selectSession now waits for the font before it measures, so the pane is
sized once, at the size that sticks, and the capture is taken at that size.
The wait always resolves, so a font that never loads cannot block a
terminal, and it resolves immediately once the font is in, so a tab switch
pays nothing after the first load.
document.fonts.ready alone is not enough: it can resolve before the
stylesheet declaring @font-face has been parsed. document.fonts.load for
each family in the stack is what actually requests the faces.
Measured on a session opening at 2328px wide: the cell went from 8.43x16.00
to 8.00x21.00 roughly 900ms in, moving the grid from 112x36 to 118x28 after
the replay had already been painted.
Switching to a session left the caret one row below the composer's input
line, on the box border, and every cursor-relative update the CLI sent
afterwards was measured from the wrong row. Any fresh output repaired it,
because the CLI then repainted the whole frame.
Two things were wrong with the full-history replay, and they compound.
The capture never restored the cursor. The visible-frame path ends with an
absolute cursor move back to the pane's position; the linear path returned
its text and left the caret wherever the last character landed, which for an
agent CLI is the bottom-most row carrying text — the status line.
The rows it addressed did not line up with the pane's rows either. Four
transforms ran over the capture and each can delete a line: the trailing
blank rows were stripped, redraw-bloat stripping ran, the trim that cuts
everything above the Claude banner ran, and leading whitespace was removed.
All four are right for a byte stream of successive frames. A capture is the
rendered pane, one line per screen row, so each deletion shifted the frame
out from under the restored cursor.
The full-history path now appends the pane's own cursor position and keeps
every row, so row N of the reply is row N of the pane. The visible-frame and
tail paths are untouched.
Restoring the cursor is what makes row alignment load-bearing here, and
neither CLAUDE.md nor the architecture invariants said so — which is how
four line-deleting transforms accumulated on the path. Both now record it.
Verified against a live 315x59 pane: the reply carries 59 rows, its row 55
is the composer's input line matching tmux, and it ends with the cursor move
that lands there.
Popping a session out left both windows sizing the same pane. The dashboard
keeps the session active and keeps measuring it, and its terminal is
narrower than the popup because the session rail takes width the popup does
not have. One PTY cannot hold two sizes, so the CLI drew frames that fit
neither window and the popup showed a garbled frame.
sendResize and the debounced window-resize handler now stand aside for a
session this window has marked detached. A solo window is exempt, since it
is the owner. _maybeRefetchFullHistory already stood aside on exactly this
condition, so the rule is not a new one.
Sizing has to come back when the popup closes: while it owned the session
the dashboard sent no resizes, so the PTY still holds the popup's geometry.
_redock now re-asserts, with force, because the dimensions the dashboard
last sent are the ones it is about to send again.
Reproduced with a dashboard and a popup on one session: before, the pane
sat at 315 columns while the popup rendered 289. After, both report the
same size and the popup's frame matches the pane exactly.
Claude Code answers an exhausted model budget INSIDE the turn ("You've
reached your Fable limit. Run /usage-credits to continue or switch models
with /model.") and then sits there with nothing to write. The reviewer never
produces a report, so `runTurn` waited out its full 40-minute deadline and
reported a bare "timed out after 40 min without a report", which reads as a
hung reviewer rather than an account that needs attention.
Measured on 2026-09-08: #388, #393, #394 and #377 each lost 40 minutes this
way, and because every attempt counted, all four reached MAX_AUTO_RETRIES and
would NOT have been picked up again once the budget returned. One spent
afternoon quietly took the whole queue out of service.
`findModelLimitNotice()` reads the notice off the pane and `runTurn` returns
a new `limit` outcome instead of waiting. It is consulted in exactly two
places, both of which mean "the turn produced nothing": on a stop where
`isDone()` is still false, and on each timed-out wait slice. A review that
merely discusses usage limits in its own findings therefore cannot be
mistaken for one that hit the wall, and the pattern matches neither the model
name nor a straight apostrophe, since the pane renders a typographic one and
every model prints the same sentence.
A spent budget is an account condition, not a bad PR, so it no longer spends
the per-head retry budget: the queue resumes by itself when the budget does.
Telegram now names the cause and the file to change.
Tests use the pane captured verbatim off the run that lost the 40 minutes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ctrl+V in the terminal inserted the clipboard text twice. Right-click →
Paste inserted it once.
`_handleImagePaste()` appends a hidden contenteditable div, focuses it, and
reads the clipboard out of the paste event that lands there. Two separate
routes deliver that event for a single keypress. The function issues
`document.execCommand('paste')` itself, which in Firefox dispatches a
trusted paste event and then returns false, because the trap cancels the
event and the command never completes; Chromium and WebKit refuse that
command and dispatch nothing. The keydown's own default action delivers the
other, because xterm calls the custom key handler before its own `cancel()`,
so returning false never calls preventDefault. Firefox therefore ran the
trap's listener twice and both runs reached `terminal.paste()`. The
context-menu paste involves no keydown at all, which is why that path stayed
correct.
The trap now accepts the first paste event and cancels every later one, so
how many paste events a browser delivers no longer changes what the PTY
sees. Measured on a live install, one Ctrl+V each: Firefox two events and
two writes before this change, Chromium and WebKit one and one, and every
engine one write after it.
The `execCommand('paste')` call stays. Stripping it out also ends the
doubling, and all three engines still deliver one event without it, since
`trap.focus()` has already run when the key's default action resolves. It is
kept because the trap technique arrived in #84 for plain HTTP and for
mobile, and a desktop measurement says nothing about real iOS Safari or
Android Chrome: where a browser aims the default action at the element
focused when the keydown began, the command is the only route into the trap,
and the trap is the only place clipboard image blobs are read.
test/image-paste-trap.test.ts loads image-input.js into a `node:vm` context
with a fake document and fires two paste events at the trap. It covers text
and images, and fails on the old code with the text pasted twice and the
image uploaded twice.
Docs: the invariant goes into docs/architecture-invariants.md as a Terminal
paste section and into CLAUDE.md as a Frontend entry, both recording the
measured event counts and why the redundant call is still there. README.md
and the Keyboard Shortcuts and Input and Voice wiki pages gain a Ctrl+V row,
which all three tables were missing while listing every other clipboard
binding.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The phone block lifted the toolbar (and padded .main) by (100vh - --app-height) on iOS Safari to clear a bottom bar that position: fixed elements were assumed to sit behind. On iPhone Safari fixed elements already stop above the bar, and 100vh is the large viewport with the bar collapsed while --app-height is the visual viewport with it expanded, so the expression measures the bar's collapsible height and shows up as an empty band between the toolbar and the bar whenever the bar is expanded. The terminal was padded by the same amount.
The lift is now --chrome-overlap, set in updateAppHeight() as innerHeight minus the visual viewport height: the distance the layout viewport that anchors fixed elements extends past the visible area. That is 0 on iPhone Safari, so the toolbar meets the bar, and it is the overlap itself on any browser where fixed elements really do land behind the chrome, so those keep the lift. The keyboard-visible rules, which already override the toolbar offset, are unchanged.
The phone tier stopped at innerWidth < 430 and @media (max-width: 430px), so every current large phone landed in the tablet layout: the 430pt iPhone 14 Pro Max, 15 Plus, 15 Pro Max and 16 Plus, the 440pt iPhone 16 Pro Max and 17 Pro Max, Pixel 6 Pro, 7 Pro and OnePlus 12 Pro, the 448pt Pixel 8 Pro and 9 Pro XL, and the Galaxy Z Fold 5 cover screen at 460. On those devices the header icon row replaced the session pill, the toolbar kept the desktop Run Shell button instead of Enter and the mic, the keyboard accessory bar could never become visible because its .visible rule lives inside the phone block, and the toolbar jumped to the top of the page when the keyboard opened.
The new cutoff is 600, the line test/mobile/devices.ts already draws between large phones (430-599) and small tablets (600-767). No physical device sits between 480 and 600, but a phone zoomed out one or two steps in Safari does: a 440pt iPhone at 85% or 75% page zoom reports 518px or 587px and still needs the phone controls, which a 480 cutoff would have taken away. The phone block is max-width: 599px and the tablet block starts at min-width: 600px, so a 600px device is a tablet in CSS and in getDeviceType() alike instead of straddling the boundary the way 430pt phones did.
The number changes everywhere it is encoded: JS, CSS, comments, CLAUDE.md, the CI tests that pin the phone block, and the test:mobile helpers. Measurement history that names 430px stays as written.
- Remote omp command now renders through buildSpawnCommandFromRegistry
(the mode-agnostic engine local/docker spawns use) instead of the
buildOmpCommand() the CLI-registry refactor deleted.
- Session._pinOmpRespawnId()/_maybeCaptureOmpSessionId() now skip
host-local ~/.omp resolution entirely for a remote session and fall
back to --continue: that resolver only ever reads THIS host's
filesystem, which is meaningless (and could wrongly alias an
unrelated local conversation) for a conversation that lives on the
remote host.
- Remote-claude launch now honors an explicit resumeSessionId distinct
from sessionId (mirrors claudeDockerPaneCommand's shape), and
validates sessionId the same way that sibling does before
interpolating it into the remote shell command.
- Add the still-missing header-cwd half of the trailing-slash test,
and document respawn/reattach continuation + auto-reconnect-vs-
clean-exit in docs/remote-sessions.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The COD-108 reconnect watcher treated any dead local pane as a dropped
transport and re-ran the pane command — so a normal ctrl-c/ctrl-d on a
remote omp/opencode/claude auto-spawned a FRESH agent (claude only
looked correct because its '--session-id || --resume' fallback resumed,
with a loud 'already in use' error first).
Distinguish a transport drop from an intentional exit: only reconnect
when the durable remote tmux session (codeman-ssh-*) is verifiably
still alive on the remote host. A clean exit tears that session down;
the watcher now probes it via ssh has-session and skips (remote-gone)
when it is gone OR unknown (fail closed). The probe is cached
per-session and fired async so the 5s tick never blocks on ssh.
Also thread ompConfig/resumeSessionId into the remote builders so a
dead-pane respawn of an omp session resumes (--resume <id>) or
continues (--continue) instead of launching bare omp.
Tests: 3 new cases pinning remote-gone / unknown / alive decisions;
remote omp resume + --continue fallback. Verified live: all three
remote CLIs stay dead after exit.
Two independent defects made ANY clean exit from a remote SSH session (user
ctrl-d or ctrl-c, or a dropped pane) relaunch the agent as a NEW conversation:
1. SSH-remote claude was launched as a bare `claude --dangerously-skip-permissions`,
so the remote-respawn path (COD-108 reattachRemote re-running the idempotent
launch command) started a fresh conversation every time. Pin it to the
deterministic Codeman session id, mirroring the docker-claude shape
(claudeDockerPaneCommand): `--session-id <id>` to create, with the
`|| --resume <id>` fallback so the idempotent re-run resumes instead of
erroring with "already in use". A per-host commands.claude override still wins.
2. OMP --resume pinning silently degraded to ambiguous `--continue` whenever a
case path ended in a trailing slash (e.g. remote `remotePath` stored verbatim
as `/home/user/dotfiles/`): mangleOmpWorkingDir produced `-dotfiles-` while
omp persists sessions under `-dotfiles`, readdirSync returned null for an
existing dir, and findLatestOmpSessionId/resolveAndClaimOmpSessionId never
matched. Normalize the trailing slash before mangling (new exported
stripTrailingSlash) and compare the session header cwd against the same
normalized value.
Both were found live 2026-08-29 on a remote OMP/Claude node: ctrl-c and ctrl-d
behaved identically, both relaunching a fresh session.
Responds to Ark0N's review round on the ephemeral-CLI-flag statusline
injection rework:
- Rebase-detail fixes: registry-gated telemetry eligibility via
getCli(mode)?.capabilities.statusLineTelemetry instead of a hardcoded
mode === 'claude' check, using the capability flag master's CLI-registry
refactor already declares for exactly this purpose.
- Design question settled: sticky (a). Rather than persisting the toggle
as a new field and threading it through every session-creation path
(cron, Ralph Loop API, quick-start), eliminated the per-session field
entirely. readPlanUsageTelemetryEnabled() (hooks-config.ts) reads the
existing showPlanUsageLimits setting fresh from settings.json at every
claude create/respawn (TmuxManager.createSession/respawnPane) - no
per-session state to survive a restart, and it applies uniformly to
every creation path for free, since they all flow through the same
TmuxManager methods.
This required fixing a real bug found along the way: showPlanUsageLimits
was not actually round-tripping through settings.json on save -
settings-ui.js explicitly excluded it from the PUT body as a pure
per-device display key. It now flows through normally (both true and
false); the load-side per-device merge behavior is unchanged.
Removed entirely as a result: the statusLineTelemetry field from
CreateSessionSchema/SettingsUpdateSchema, CreateSessionOptions/
RespawnPaneOptions, Session._statusLineTelemetry (this is what makes
the restart-persistence bug moot rather than patched), and the
frontend send sites.
- Footer print-through restored: the no-user-statusline branch of the
exporter script now runs the telemetry POST in the foreground so its
own stdout becomes the in-terminal footer, falling back to a plain
"codeman" marker only on curl failure.
- Background-subshell EOF fix: the wrap-a-real-statusline branch closes
stdin too, not just stdout/stderr (`>/dev/null 2>&1 </dev/null &`) -
the un-redirected subshell process itself, not curl, was what held a
reader-to-EOF's pipe open for however long curl took to finish. Added
curl --max-time 5 so a hung (not just refused) Codeman cannot wedge
the render.
Tests: real-shell-execution tests for the footer/EOF fixes (fake curl
stand-in on PATH, real sh subprocess spawns, real elapsed-time
measurements - verified non-vacuous against a hand-reconstructed
old-style script), unit tests for readPlanUsageTelemetryEnabled.
Adapted two existing tests whose payloads referenced the removed field.
Fixed during independent code review: a stray indentation break and a
test exercising the wrong (legacy) exporter code path.
Docs synced: CLAUDE.md, docs/usage-limits-display-plan.md (old
disk-based section marked superseded, kept for history),
docs/architecture-invariants.md.
Full suite green: 352 files, 6780 passed, 12 skipped, 0 failed.
tsc/lint/format:check/frontend-syntax all clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Now that the exporter no longer lives in a fixed per-case file, it can
compose with the user's actual configured statusline rather than just
backing off when one is found.
findEffectiveUserStatusLineCommand() walks Claude Code's own settings
precedence for a workspace: project-local .claude/settings.local.json
> project-shared .claude/settings.json > the user's global
~/.claude/settings.json. A legacy Codeman-marked entry left behind in
the project's own settings.local.json is never treated as a real user
command — it's skipped and precedence continues to the next layer.
The shared exporter script (bumped to a V2 marker so stale copies
self-heal) now fires the telemetry POST in a background subshell —
its own stdout/stderr discarded so nothing leaks into the visible
statusline, and confirmed non-blocking (~4ms, even against an
unreachable endpoint) — then, if the pane's environment carries
CODEMAN_USER_STATUSLINE_CMD, feeds it the same stdin blob and relays
its stdout as ours. Otherwise it falls back to the plain "codeman"
marker as before.
The discovered command is threaded to the pane via `tmux setenv
CODEMAN_USER_STATUSLINE_CMD` (_configureStatusLineUserCommand) rather
than embedded in the spawn command line, for the same
premature-shell-expansion reason as the parent commit: tmux stores a
setenv value verbatim and never re-parses it, so once shellescape()d
for that one command, the command's own $/quotes survive untouched
into the pane's environment.
Verified live via direct shell execution of the generated script
(both branches: fallback and user-command wrapping) before deploy.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015GyMnFWnUzc41TDeHg9juW
Codeman's plan-usage chip wrote a statusLine.command into the case's
.claude/settings.local.json to receive Claude Code's rate_limits blob.
That file-based statusLine took precedence over the user's own
global/project statusline for ANY `claude` run in that directory,
including entirely outside Codeman, with no disclosure in the App
Settings UI (labeled only as a header-display toggle) and no way to
remove it once written (the removal code path was unreachable dead
code — nothing ever called it with false).
Replace the disk write with an EPHEMERAL `claude --settings
'{"statusLine":{...}}'` CLI flag, resolved fresh at spawn time
(resolveStatusLineCliCommand in hooks-config.ts) and merged with
effort/ultracode into one --settings object (buildClaudeSettingsFlag
in tmux-manager.ts, since Claude Code accepts only one --settings
flag). Never touches disk, so a plain `claude` run outside Codeman is
untouched. Self-healing: any legacy disk-written exporter from an
older build is stripped the first time a session starts in that
workspace again. Still respects a user's own hand-authored statusLine
(skips the flag entirely rather than overriding it).
Mid-fix bug found and fixed: the exporter's command legitimately
depends on $CODEMAN_SESSION_ID/$CODEMAN_API_URL/$CODEMAN_HOOK_SECRET_FILE
and an internal $INPUT, all meant to be expanded only when Claude Code
itself executes the statusline, using the pane's tmux-setenv'd
environment. Passing that text through --settings routed it through
execSync's own implicit /bin/sh -c first (tmux respawn-pane's
`bash -c "..."` wrapper) — POSIX double quotes don't suppress $
expansion, so those vars got expanded prematurely against the
server's own environment (unset there), producing malformed JSON that
printed as literal error text in the statusline. Fixed by writing the
exporter as a real, shared script file (ensureStatusLineExporterScript,
marker-versioned so stale copies self-heal) and passing only its bare
path via --settings — nothing for any intermediate shell to mangle.
Verified against a real Claude CLI on an isolated tmux socket, and via
direct execSync reproduction of the exact nested wrapping
createSession/respawnPane use.
A hard "never inject, even ephemerally" kill-switch was added and then
removed in the same pass: with the disk-leak fixed, disabling
injection only cost the plan-usage telemetry the feature exists to
provide, for no remaining benefit.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015GyMnFWnUzc41TDeHg9juW
The previous shape guessed the character from `event.key` on keydown, re-emitted
it, and then tried to suppress a late canonical copy with a 250 ms
character-keyed dedupe. Review found three defects in that, all reproducible:
the dedupe matched on the character alone with nothing scoping a candidate to
the keydown that created it, so the same character typed twice inside the window
had its second, real byte swallowed; anything whose committed text differed from
`event.key` (Enter, IME punctuation) was delivered twice, because the dedupe
could never match it; and the trigger ignored `key === 'Unidentified'`, which is
what a soft keyboard reports, so it may never have fired where it was needed.
The input event already carries the committed text in `ev.data` — exactly what
xterm itself would have forwarded — so nothing has to be guessed. The controller
now only decides WHETHER to forward, by asking whether xterm produced canonical
data since the keydown that began the keystroke. No character-keyed matching
survives, so the first two defects are structurally impossible rather than
defended against, and nothing reads `key`/`keyCode`, so the third cannot recur.
Three details are load-bearing and each has a test that fails without it:
- The "did xterm speak?" snapshot is taken at KEYDOWN, not at the input event.
`_keyPress` emits and sets `_keyPressHandled` before `input` fires, so a
snapshot read at input time already contains that emission, reads it as
silence, and delivers the character twice.
- Our `input` listener is registered with `capture: true`. The target is visited
twice in the event path, so a capture listener calling `stopPropagation()`
stops later BUBBLE listeners on that same target; xterm's `cancel()` runs
exactly in the branch where it handled the input, so on bubble we would never
observe handled events, and whether we observed them at all would hang off
`options.cancelEvents`. Measured in jsdom and headless chromium; the table is
in the module header.
- Enter is deliberately no longer special-cased. That mapping is what made the
committed text differ from the re-emitted value in the first place.
The scope is also narrower than the old name suggests, and the browser test now
proves it rather than assuming it. For a keydown that reports keyCode 229 xterm
ALREADY self-rescues, via `CompositionHelper._handleAnyTextareaChanges()`
diffing the helper textarea on a 0 ms timer. A test asserting "we recovered it"
there passes while xterm does all the work, so the browser tests assert WHO
delivered the byte: zero canonical emissions for the genuinely orphaned case,
exactly one delivery for the case xterm rescues itself.
Also addresses review notes: the module gains an `@fileoverview` with
`@dependency`/`@loadorder` and an entry in the load-order list and module
inventory, and the wiring test moves out of the Ctrl+C smart-copy file into its
own. The keydown hook deliberately still runs for every key event rather than
moving behind the 229 gate: gating it would reinstate exactly the blindness
described above, and it is now a single counter assignment.
#386 made Codex conversations resumable from Past Sessions, and
resumeMobileOverviewSession() correctly passes row.resumeId on to
resumeHistorySession(). The phone's own row projection never copied the
field off the unified-list item though, so row.resumeId was always
undefined there and a tapped Codex row started a FRESH session on a thread
that was already on disk. The desktop path worked; only the phone was blind.
The test fails without the projection line, and pins the other half too: a
claude row must not grow a resumeId, since the field is what distinguishes
"resume this conversation" from "start a new one".
Docs: CLAUDE.md and architecture-invariants both still described the unified
list as merging Claude transcript files. It has been three stores since this
PR (Claude's ~/.claude/projects, omp's ~/.omp/agent/sessions, codex's
~/.codex/sessions), the alias field keeps its Claude-era name without being
Claude-only, and the scanner-only rule behind resumeId was written down
nowhere.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#385 made the composer glyph and the working status line per-CLI registry
data, which is right, but `workingLine` arrived as a config-supplied regex
validated with a bare `new RegExp()`. That skips `compileVersionRegex()`,
the helper the registry uses for exactly this: a `~/.codeman/clis.json`
override can set the field, the compiled pattern is run against every
accumulated PTY chunk and every pane capture, and a nested quantifier there
backtracks on the event loop for the whole server rather than one session.
Route it through the helper in both places, which are not redundant: the
schema refine rejects the entry at LOAD time so a bad pattern never reaches
a session, and `_workingLinePattern()` compiles through the same helper so
the runtime cannot hold a pattern the schema would have refused. The helper
returns null instead of throwing, so the Claude-pattern fallback stops being
a try/catch and becomes structural. Both shipped patterns compile unchanged,
and Claude's is behaviourally identical to CLAUDE_WORKING_LINE_PATTERN.
Also match the Codex footer case-insensitively on the E. It was
characterised against codex-cli 0.152.1, which prints a lowercase `esc`;
a version capitalising it would make the whole fix silently inert, since
the pane would simply never look like it was working.
Docs: CLAUDE.md, architecture-invariants and cli-registry.md all still
stated the Claude-mode-only rule this PR retires, and none of them named
the new capability or the regex guard.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Merging with follow-up fixes applied on master: workingLine routed through compileVersionRegex() in both the schema refine and _workingLinePattern(), the Codex footer matched case-insensitively on the E, plus the doc passages that stated the retired Claude-mode-only rule.
A long orchestration creates one case directory per worker and deleting the
sessions never removed them, so ~/codeman-cases accumulated scratch folders
that were indistinguishable from real projects. They are now labelled and
have a cleanup path.
- src/agent-case-marker.ts: a case dir quick-start CREATES for an agent-driven
spawn gets a .codeman-agent-case.json marker (when, by whom, parent session,
mode). Only the create branch writes it, so a linked case, a cloned repo or
any pre-existing path is never labelled; reading is total, so a malformed
marker means "not agent-created" rather than a half-trusted entry.
- The signal is the new X-Codeman-Agent-Origin header the skill preamble sets
on its shared curl (preamble bumped to 1.22.0), or an agentOrigin body
field, falling back to a resolved parentSessionId so a worker spawned by a
stale skill copy is still labelled.
- GET /api/cases publishes it as agentCreated; GET /api/cases/agent-created is
a read-only cleanup listing adding inUse and modifiedAt; Add Case -> Manage
badges each case and offers a review-then-delete sweep that names every
directory in its confirm and skips any case a live session is working in.
Removal stays on the existing DELETE /api/cases/:name.
- Agent preamble caches are collected too: ~/.cache/codeman-agent-<id>.sh was
written per claude session and never removed (236 leftovers measured on a
working machine). Now deleted with the session and swept at boot, guarded by
a live-session keep set plus a 7-day age floor.
Verified end to end on an isolated instance: marker written for header, body
and lineage-only spawns, absent with no agent signal and for a pre-existing
directory; inUse flipping on session end; badge, sticky bar, confirm and sweep
driven in a browser; preamble seeded on create, removed on delete, boot sweep
taking only the aged orphans.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Android/GBoard-style keyboards fire keydown with keyCode 229 and, on some
paths, never mutate xterm's helper textarea. xterm has nothing to diff, so
it emits no data and the typed character is silently dropped: it never
reaches the PTY and never appears on screen.
terminal-keycode229-recovery.js is a standalone controller that re-emits
exactly those keys, and only once. xterm stays authoritative throughout:
- Only an explicit keyCode 229 keydown carrying a single printable key (or
Enter) is eligible; Process/Unidentified/Dead, modifiers, AltGraph and a
live composition are all left alone.
- The re-emit is scheduled from a microtask and then a zero-delay timer, so
xterm's own textarea diff always gets the first opportunity; canonical
data for the same key cancels the pending fallback.
- compositionstart and blur drop every pending candidate, so a real IME
composition lifecycle is never second-guessed.
- After a recovery, one late canonical value attributed to that key token
(via beforeinput/input on the helper textarea) is suppressed so the
character cannot be delivered twice; the record expires after 250ms and
an unattributed byte is never suppressed.
terminal-ui.js wires it at the two existing choke points — the custom key
handler and the onData registration, the latter now a named handler so the
recovery path can re-enter it — with both hooks wrapped so a failure in the
fallback can never break canonical input.
Unit coverage drives the module directly in a vm; the wiring itself is
covered end-to-end in the (browser-only) terminal-copy-shortcut suite.
master and this branch both rewrote the two `_claudeSessionId` resets inside
`start()`, so `src/session.ts` conflicted at both of them.
master's commit ccfda623 puts `restoredConversation` at the head of each
fallback chain. A restored mux attach means the CLI never stopped, so a
`/clear` before the Codeman restart may already have moved it to a
conversation the launch id knows nothing about. The persisted chain's tail is
that conversation, and the CLI's own hook reported it first-hand.
This branch adds `this._codexConfig?.resumeSessionId` to the same two chains,
so a resumed codex session keeps its thread-id alias across every mux reattach
and boot recovery.
Both fixes belong. Each chain now reads restoredConversation, then
_resumeSessionId, then omp's alias, then codex's alias, then the launch id.
The comments from both sides are kept.
test/session-claude-conversation-chain.test.ts pins the shape of those two
assignments by matching the source text, and its pattern named omp's alias as
the last term before `this.id`. Codex's alias now sits between the two, so the
pattern widens to pin the ends of the chain and let the middle grow. A `[^;]`
run cannot cross a statement boundary, so each match is still one assignment.
Checked on the merged tree: typecheck, lint, prettier and the frontend syntax
check all pass, and the CI suite runs 6721 tests green across 349 files.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The merge-time items from the #381 review. navigator.sendBeacon is not fetch,
so the base-aware wrapper never saw the two crash-diag beacons and a sub-path
install posted them to the origin root every two seconds. The test suite now
strips CODEMAN_BASE_URL like CODEMAN_GESTURE, since the constructor reads it
as a fallback and an operator who exports it would see the root-install
byte-identity assertions fail. test/base-path-server.test.ts boots a real
WebServer under /codeman and checks the ingress strip, the base injection,
the rebased redirects, the 404 envelope and a prefixed WebSocket upgrade.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
#367 (UserPromptSubmit hook): `hook:prompt_submitted` went on the wire
unregistered; it is now in both SSE registries (158 = 158), and the hook only
lands in the run summary when the conversation actually moved, since one row
per prompt would evict useful rows from the 1000-event FIFO and clutter the
Summary timeline and /api/search.
#368 (Add Case header submit): the pending-state dimming targeted the footer
button, which the <=860px layout hides, so on a phone the only visible submit
control stayed at full brightness while a clone ran. The header button now
dims too, and a static test pins the header-submit contract so it cannot
silently disappear again.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Review fixes for #386.
Duplicate rows. A codex conversation showed twice, once live and once as a
past rollout row, because nothing aliased a codex session to its thread id.
That is worse than cosmetic: the stale row still resumes, so clicking it
starts a second `codex resume` on a thread already open in another pane.
- A RESUMED session knows its thread id up front, so it folds from its own
side: add `codexConfig.resumeSessionId` to the `claudeSessionId` chain.
Not only in the constructor — `start()` recomputes that id at two further
points (the mux branch, and the unconditional "third reset point" whose
own comment already warned that omitting omp's fallback there stomps the
mux branch's resolved alias). Both listed Claude's and omp's ids only, so
for codex every mux reattach and boot recovery reset the alias back to
the Codeman id and the duplicate returned.
- A FRESH session has no thread id until codex writes the rollout, so it is
folded from the other side. The scanner now reports
`session_meta.originator`, which is `codeman_<sessionId>` for every pane
Codeman spawns, and `gatherUnifiedInputs()` stamps the matching live and
persisted rows, newest rollout winning (`/new` inside the TUI leaves
several rollouts sharing one originator).
- Persisted rows read `codexConfig.resumeSessionId` too. A resumed session
demoted to a persisted-only record would otherwise lose its alias, and
the originator fallback cannot rescue that one: a resumed rollout keeps
its ORIGINAL session_meta, so it still names the pane that created the
thread rather than the pane that resumed it.
Identity cache. It was written as soon as the thread id was known, but codex
writes the first user message only when the user submits, so any scan in that
window pinned `firstPrompt: undefined` for the life of the process — and the
home screen, the command palette and the search-index refresh all scan.
`shouldCacheIdentity()` now keeps an identity only once the prompt is known or
the head read filled its whole window.
Also from review: both caps count emitted rows rather than file index, so a
store of sub-agent threads no longer spends the `lastPrompt` budget before the
first row that needed it; the cache is an `LRUMap` sized like the one beside
it; the unreachable filename fallback is gone; a rollout recording no cwd is
dropped rather than emitted with `workingDir: ''`; and the unified-session
module header names all three transcript stores.
Tests. The resume wiring now has cases for a row with a thread id, a row
without one, and a `resumeId` on a non-codex row; the "no continuation is
wired" case narrows to gemini/antigravity, which is no longer true of codex.
`codex-resume-alias-survives-start.test.ts` drives a real Session through
`start()` rather than asserting on pre-stamped inputs — that gap is why the
reset points went unnoticed. Plus the maintainer's own cache repro, the
tail-budget case, a no-cwd case, and merge cases for both folds.
The two merge-time edits the #383 review asked for. The comment above the
picker's fallback chain said Home is nested under Codeman Cases; on the
native default it is the other way round (~/codeman-cases sits inside ~).
And the "Filesystem path picker" paragraph in architecture-invariants still
said the picker falls back to /mnt/d, which #383 changed to: the session's
Current Folder, then the Codeman Cases root, then /mnt/d, then the first
root. No code behaviour changes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Observed on the first live merge (#383 via the Telegram button): runConfirmed
announced the merge and the scan five seconds later announced it again as a
closed PR. The scan now stays quiet for PRs the bot itself merged or closed,
and a merge of a `merge-with-fixes` verdict reminds that merging applies none
of the listed fixes.
A failed review used to be re-queued on every scan with no limit (two PRs
failed once each and were retried fine, but a head that keeps failing would
cost a session every ten minutes forever): three failures on one head now
stop the automatic retries until /review N or a new push. The failure notice
carries the reviewer's last message, so "finished without writing
report.json" says what it wrote instead.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Maintainer tooling in scripts/pr-bot/: a daemon (systemd user unit
codeman-pr-bot) that lists open PRs with gh, reviews each head commit once in
a Codeman claude session (`prbot-<n>`) running in a private `git clone
--shared`, and sends the verdict, ranked findings, checks and a recommendation
to Telegram with action buttons. Merge, close, post-comment and approve-CI
happen only from a Telegram command or button plus a confirmation tap; the
bot never writes to GitHub on its own. The Telegram token and chat id come
from the existing notifier bot's env file.
Verified live: three PRs reviewed end to end (383, 363, 368), reports
delivered with buttons, reviewer sessions on the pinned model. Findings
along the way, each fixed and documented: a linked worktree inherits the
main checkout's model pin (hence the shared clone), undici's 5-minute header
timeout cut off the first review, gh was missing from the service PATH, and
the periodic scan orphaned an in-flight review's record.
typecheck/lint/format now cover scripts/pr-bot; tests in
test/pr-bot-{report,state,commands}.test.ts; guide in docs/pr-bot.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Codex conversations never appeared in the session list, and the resume path
skipped codex, so picking one back up meant finding its thread id by hand and
POSTing codexConfig.resumeSessionId to /api/sessions.
Two gaps caused it:
- The unified list is built from ~/.claude/projects plus omp's own store.
Codex writes to neither: its rollouts live in ~/.codex/sessions/<y>/<m>/<d>.
- terminal-ui.js sends a continuation only for the CLIs with a
"continue most recent" flag. Codex has no such flag — it names a thread by an
exact id — and nothing supplied one.
Add codex-transcript.ts, the codex analog of omp-transcript.ts, and wire it into
gatherUnifiedInputs() beside the omp scan. A rollout row carries `resumeId`, the
thread id `codex resume` takes, and the resume path sends it as
codexConfig.resumeSessionId.
`resumeId` is what keeps the two kinds of row apart: only a transcript scanner
sets it, so a LIVE codex row — whose sessionId is Codeman's own uuid — can never
ask codex for a thread that does not exist.
Three things measured against a real store of 519 rollouts rather than assumed:
- Rollouts are far too large to read whole (median 407 KiB, p90 1.3 MiB, max
25 MiB, 381 MiB total), so this reads a 128 KiB head for the identity and the
opening prompt and a bounded tail for the most recent one. session_meta is
written once and never rewritten, so per-path identity is cached; a warm
rescan of that store costs ~75ms against ~470ms cold.
- codex 0.152.1 emits no event_msg/user_message rows at all. It writes
event_msg/item_completed carrying an item.type of UserMessage. Both shapes are
read, plus response_item as a last resort.
- That last resort sees injected context, and the first such row is the repo's
AGENTS.md every time, so injections are dropped rather than used as titles.
Sub-agent threads (thread_source: 'subagent') are left out; codex spawns them
for itself and on a real store they outnumber the resumable threads.
A Codex session reported `isWorking: false` for its entire life, including
mid-turn. Codeman has four paths that mark a session working, and all four were
inert for Codex:
- The spinner fast path tests eight braille frames, and Codex animates none.
- The activity-streak fallback was wrapped in `!isExternalCliMode(mode)`.
- The pane probe inside `_confirmIdle` would have matched, since Codex prints
`esc to interrupt`, but arming it required the literal glyph `❯` and Codex
draws `›` on its composer row.
- The text detector sat inside `_processExpensiveParsers`, whose first statement
returns early for an external CLI.
Add an optional `workDetect: { promptGlyph, workingLine }` to CliCapabilities,
so the two strings that differ per CLI are registry data rather than constants
in the detector. Claude declares its existing pair and behaves as before. Codex
declares `›` and `esc to interrupt`. The text detector moves above the
external-CLI early return, guarded on the descriptor so a CLI without one still
skips the ANSI strip that the early return used to save it.
A CLI that declares no descriptor falls back to Claude's pair, and the
activity-streak gate now reads "has a descriptor, or is not external", so the
plain shell mode keeps the behaviour it had.
Rewrite the test that asserted the old premise in its own comment, so it makes
the same guarantee for a genuinely uncharacterised CLI, and add Codex coverage
built from verbatim pane captures on Codex CLI 0.152.1.
Review follow-ups to #357. Each is a path that still touched, or still hid, a
container Codeman does not own.
**Export still mutated it.** The four fail-closed layers cover create/start/
stop/remove, but `POST /api/docker-cases/:name/export` reaches the container
twice through neither: a full export `docker commit`s it, and even a
workspace-only export `docker pause`s it first for snapshot consistency. Pause
freezes the owner's processes for as long as the tar takes, on a container we
promised not to touch. Full export is refused for an adopted case (it packages
someone else's container, with their logins, into a bundle Codeman hands out);
workspace-only keeps working and no longer pauses, accepting a live filesystem
the way `tar` does on any running host directory.
**A freshly linked OWNED case became unusable.** The run menu now probes the
container for its CLIs, and a failed probe hides every agent mode behind the
reason. For an adopted case that is right. For an owned one the container does
not exist until the first session launches it, so every newly linked Docker case
answered `container "codeman-case-x" not found (adoption never creates a
container — start it yourself first)` and offered nothing but Shell, for a
container the launch chain was about to create itself. A failed probe is
recorded only when the case is adopted; `CaseInfo.docker.owned` is on the wire
so the frontend can tell them apart. Verified in a browser: owned-with-no-
container offers all ten modes and no notice, adopted-but-stopped offers Shell
and says why.
**Multi-user gating.** Adoption is admin-only, unlike `docker-link` beside it.
Linking creates OUR container, whose sole bind mount `isWorkingDirAllowed` has
already confined to the caller's space; an adopted container's mounts are
whatever its owner gave it, so one mounting `/` hands the adopter a shell over
the whole host — exactly the workspace scoping multi-user mode exists to
enforce. Listing the engine's containers and browsing directories inside an
arbitrary one are machine-level reads and follow the docker-HOST policy for the
same reason. The preflight is deliberately not admin-only: the run menu fires it
for every docker case, so it admits a non-admin for a container already linked
to a case they can access, and nothing else.
Verified end to end against a real pre-existing root container (alpine + tmux,
no bind mounts): adopt, claude session inside it, workspace export, session
close and case unlink all left `StartedAt`, `RestartCount`, `Pid` and `Paused`
untouched; the pane ran the CONTAINER's claude, without
`--dangerously-skip-permissions`; a stopped container was refused at both
preflight and launch and was never started.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1
feat(docker): attach a case to an already-running container
Conflicts came from work that landed after the PR was opened, and each is
resolved onto the newer abstraction rather than by keeping the older code:
- `defaultDockerCommandForMode` is registry-driven since #347, so the PR's
`runsAsRoot` arm became `overlays.docker.rootCommand` (claude only). Claude
Code still refuses `--dangerously-skip-permissions` as root in 2.1.261 and the
refusal is visible only inside the container, so an adopted root container
otherwise just shows a dead pane. Which flag to drop is a per-CLI fact, and
`test/cli-registry-no-id-branching.test.ts` forbids expressing it as a branch.
- The probe's mode list and its mode -> binary table both duplicated the
registry. They now read `enabledCliIds()` / `discovery.binaries[0]`, which is
also what fixes the merge's silent regression: the hand-written list predates
`omp`, and the run menu gates every docker case on this probe, so owned
containers would have lost that mode. `shell` needs no arm — it declares no
binary, so it is dropped from the lookup and reported available regardless.
- The per-mode `mode === 'claude' && !cliDir` chain in `tmux-manager.ts` is one
`missingCliMessage(mode)` gate since #347; the PR's docker exemption moved onto
it. Its test now pins the single gate instead of counting seven arms.
- The create arm keeps #349's swap-limit warning filter, which the adopted arm
never reaches; the run-mode list gains `omp` from #353.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1
The "Link Existing" case picker opens with an empty path and no
sessionId, so the browse endpoint's fallback root picked whichever
root happened to be first in the list — which was always `Home`.
On the native default that's harmless (~/codeman-cases nests inside
Home anyway), but a Docker deployment binds CODEMAN_APPDATA_PATH
(Home) and CODEMAN_CASES_PATH at unrelated host paths, so the picker
opened somewhere with no cases in sight. Worse: if CODEMAN_CASES_PATH
is ever changed after cases already exist, the old cases directory
lingers, still reachable, under Home — indistinguishable at a glance
from the real one under the new Codeman Cases root.
Prefer the Codeman Cases root in the fallback chain, ahead of the
generic roots[0].
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
Codeman can now be mounted under a sub-path behind a reverse proxy that
forwards the prefix unchanged (e.g. https://host/codeman/). Default is `/`
(root), which is byte-identical to the historical behavior.
Design — few choke points, mirrored ingress/egress:
- src/config/base-path.ts: pure single-source normalize/validate/join/strip.
- Server ingress: stripBasePath() inside Fastify rewriteUrl, so routes stay
declared prefix-agnostic; un-prefixed requests (hooks, health, docker bridge
hitting the raw port) pass through unchanged.
- Server egress: one onSend hook prepends the base to root-absolute Location
headers (covers all redirects).
- HTML: renderIndexHtml points <base href> at the mount and injects
window.__CODEMAN_BASE__ — ONLY when a base is set (inert at root).
- Frontend runtime URLs: CodemanBase.url() route builder in constants.js,
applied transparently by a fetch wrapper and explicitly at the
EventSource/WebSocket/window.open/<img|iframe|a>-src sites.
- sw.js derives its base from self.location; manifest uses relative start_url/scope.
- Web-tab proxy: proxyPrefixFor(cap, basePath) is the single base-aware root that
cascades to the injected <base>, HTML/attr rewrites, runtimeUrlShim, Set-Cookie
Path and Location; capabilityFromReferer strips the base off the browser Referer,
while the ingress parsers stay base-agnostic (rewriteUrl already stripped it).
--base-url rides the daemon relaunch (buildWebArgs) and the service unit
(resolveServicePlan). constants.js is guarded against a missing `window` for
isolated unit-test contexts.
Tests: test/base-path.test.ts (pure helpers), base-path coverage in
webview-proxy/render-index-html/daemon-control; CodemanBase stubbed in the
vm-isolated panels-ui test contexts. Docs: Remote-Access.md (sub-path section +
nginx example), security-architecture.md env table, CLAUDE.md pattern.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XUkPBxbumnct6qSrx4JDju
WebviewCapabilityStore.revokeOwner() shipped for two releases with a docstring
claiming logout called it and no caller at all. The capability is a bearer
credential exempt from cookie auth with a rolling TTL refreshed on every use, so
a proxy URL that leaked (browser history, a screenshot, a dashboard with a loose
referrer policy) stayed valid for as long as anything kept polling it.
- POST /api/logout revokes the caller's capabilities (all of them in single-user
mode), the admin forced logout revokes the target user's, and user deletion
revokes whatever that user had open. revokeOwner returns the count for the
admin audit line.
- Proxied responses carry `Referrer-Policy: same-origin` and the upstream's own
policy is dropped: every URL inside the frame carries the capability, and a
dashboard on no-referrer-when-downgrade or unsafe-url handed it to any
third-party host it linked. Verified with Playwright that a sandboxed frame
under an upstream `unsafe-url` sends no Referer to a third party while the
root-absolute fetch and the CSS-triggered 404 fallback still reach the
dashboard.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
The web-tab proxy, its Test probe and its WebSocket relay accepted any http(s)
host. A live PoC relayed an IMDSv2-shaped PUT with custom headers to a loopback
echo server through a capability and no cookie, and 169.254.169.254 (decimal,
hex, IPv6-mapped, or via a DNS name) was as valid a dashboard as any other.
Loopback and RFC1918 stay allowed on purpose: a localhost Grafana is the feature.
Only link-local and the fixed cloud-metadata addresses are refused
(169.254.0.0/16, fe80::/10, fd00:ec2::254, 168.63.129.16, 100.100.100.200,
metadata.google.internal), at three stages that are each load-bearing:
- the Zod schema, so a save gets a clear refusal;
- a synchronous hostname check at every connect site, because net.connect skips
DNS for an IP literal and a lookup hook never sees one;
- a `lookup` hook on an undici Agent (webviewFetch) and on the ws client, which
judges the RESOLVED addresses of a name and refuses when any is blocked. This
is what closes DNS rebinding, which a hostname-string check cannot.
Adds undici@^6 so the proxy runs the package's own fetch with the package's own
Agent; a package Agent handed to Node's bundled fetch can mismatch protocols.
Verified live on an isolated beta: 169.254.169.254.nip.io (a real name resolving
to the metadata address) is refused by probe, proxy (403) and WS relay (4003),
while 127.0.0.1.nip.io still passes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
#373 restarts the Compose container by exiting the server, which is right for
the shipped deployment: `restart: unless-stopped` relaunches it. The updater
verified that policy through the Docker socket and, when it could not (no
socket mounted), failed open and exited anyway. Failing open is the correct
choice for the GATE, where refusing would block every install without a
socket, but not for the kill: a container the daemon does not restart goes
down for good, with no UI left to recover it from. That is exactly the case a
plain `docker run` of this image without `--restart` produces, and the image
sets CODEMAN_IN_CONTAINER=1 itself, so it takes the container path.
The decision now happens server-side, where both the socket and the Compose
env are reachable, and rides down to the script as `--restart-by-exit 0|1`.
It is 1 when the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (added there
and only there, since that file is what sets the restart policy; the image ENV
deliberately does not) or when the daemon confirmed an auto-restart policy.
Otherwise the build still lands, the status becomes
`completed-needs-manual-restart` with the `docker restart` hint, and the
server keeps running. The shipped deployment is unchanged in effect: with the
socket it was already confirmed, and without it the declaration now covers it.
Also: a root-run `Start-Codeman.sh` (common on Unraid) created the
fingerprint baseline's `.codeman` directory before the container's first start
and left it root-owned, which the unprivileged server could then never write
its own state into. It is chowned to PUID:PGID when running as root.
Verified with a real image build of the merged tree (classic builder; this
box's BuildKit lacks buildx): runs as uid 1000, tsc/esbuild and the toolchain
present, the four CLIs at their pins, docker/.env absent, and `docker inspect
$HOSTNAME` returns the restart policy through the mounted socket as that user.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
#356 and #371 fixed the same leak two ways. #356 pointed CODEMAN_DATA_DIR at a
second throwaway directory and cleaned it up in afterAll and on exit; #371
deletes the variable along with CODEMAN_INSTANCE and CODEMAN_TMUX_SOCKET, so
`getDataDir()` falls back to `homedir()`, which the temp HOME already redirects.
Merged as they were, setup.ts set the variable and deleted it a few lines
later, and the second directory was created for nothing.
The strip wins: same protection, one tree to clean up, and the isolation test
#371 adds pins the list statically. The extra directory, its restore and its
two rmSync calls go, the vitest config `env` entries that set the same variable
go (they were documented as inert and would now be contradicted by the setup
file either way), the two test comments that described the old mechanism are
reworded, and CLAUDE.md's testing paragraph names the three stripped variables
and why CODEMAN_INSTANCE has to be stripped in the setup file rather than a hook.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
The loader refuses a `clis.json` with any group/world permission bit, read bits
included, so a file created with a normal umask (0644) is ignored. That is a
defensible posture for a file that chooses the binaries Codeman spawns, but two
things around it made the override feature look dead: the warning said
"group/world-writable", which a 0644 file is not, and `LoadResult.warnings` was
returned to a caller nobody wired up, so nothing anywhere printed it. A user
following the docs got silence.
The message now names the rule and the command that satisfies it, the loader
logs every warning once on first load (the result is memoized, so once per
process), the module header stops claiming that nothing ever writes (the
quarantine rename of a malformed file is a write, on first use) and the
registry doc gains a short section on the override file with the 0600
requirement in it. Whether the check should relax to writable bits only is a
separate decision; this keeps the shipped behaviour and makes it visible.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
#356 stopped a bare suite run from overwriting the production
`remote-hosts.json` by pointing `CODEMAN_DATA_DIR` at a throwaway dir, and it
gated every case-tree delete on the temp HOME. Both changes are right; the
explanation written next to them is not. It says `os.homedir()` reads
/etc/passwd rather than `$HOME` on Linux, which would mean the temp HOME in
test/setup.ts never worked. It does: libuv checks the env var before the passwd
entry (measured: `HOME=/tmp/x node -e 'console.log(os.homedir())'` prints
/tmp/x), and CLAUDE.md's testing section relies on exactly that.
What bypasses the temp HOME is `CODEMAN_DATA_DIR` itself. `getDataDir()` reads
it as an absolute override before it looks at `homedir()`, so one inherited from
the shell (a second instance, a beta run) sends the whole suite at the real data
dir. That is the case setup.ts now closes, and #371 names the same variable from
the other direction.
The comments in setup.ts, the `safeRmHomeTree` helper, the voice-routes and
case-clone tests now say that, and the containment gate is described as what it
is: defense in depth. CLAUDE.md's testing paragraph gets the same note so the
next reader does not chase a homedir() bug that does not exist.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
#355 made the remote auto-reconnect watcher revive a dead pane only when the
durable remote tmux session is verifiably still alive, which is the right rule:
a clean Ctrl-C / Ctrl-D / exit tears that session down and must never relaunch
a fresh agent. Its probe, though, read `has-session`'s stdout and treated an
empty string as "gone". `tmux has-session` prints NOTHING on success (measured
on a scratch socket: exit 0, empty stdout, the failure message goes to stderr),
so every live remote session classified as gone and transport-drop reconnects
were silently disabled along with the clean-exit revives.
The probe now goes by exit status through a pure, unit-tested mapping
(`classifyRemoteAliveExit`): 0 is alive; ssh's own 255, a timeout (`killed`,
no numeric code) and a spawn failure are unknown, which the watcher already
treats as do-not-revive; any other status is the remote command's and means
gone (tmux's 1 for a missing session, 127 when tmux is not installed there).
Two smaller things in the same area:
- The cached answer was never invalidated, so after one successful reattach a
stale `true` would have revived the NEXT clean exit (the original bug back
after the first transport drop), and a cached `false` from a clean exit would
have left a manually restarted session with auto-reconnect permanently off.
The tick now forgets the cache entry whenever the pane is seen alive.
- The fire-and-forget probe has a 15s timeout against a 5s tick, so an
unreachable host stacked up to three ssh processes per dead session. An
in-flight set caps it at one.
The probe command is pinned as a literal string, and the reattach-then-clean-exit
sequence is driven through the watcher in the tests.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
#356 introduced safeRmHomeTree/isUnderTestHome to stop tests from deleting
the PRODUCTION ~/codeman-cases tree on platforms where os.homedir() ignores
the $HOME override -- but only applied it to the one file caught doing it
live. CASES_DIR has no CODEMAN_DATA_DIR-style env override at all, so every
other test file's raw rmSync(join(CASES_DIR, ...)) was the same unguarded
pattern, just not yet triggered.
Routes every CASES_DIR delete in these 10 files through safeRmHomeTree:
cli-skill-target, edge-cases, integration-flows, operation-lightspeed,
ralph-integration, routes/case-clone-routes, routes/voice-routes,
session-cleanup, sse-events, sse-subscription-filter.
Also fixes one instance in case-clone-routes.test.ts that mkdirSync'd then
rmSync'd a CASES_DIR path directly with no guard at all -- the exact
clobbering pattern #356 exists to prevent, found by extending the sweep.
Held as a separate commit (and intended as a separate PR once #356 merges)
rather than folding into #356 -- keeps the already-checked skinny fix
reviewable on its own; this is the same bug class applied broadly, not new
functionality.
Verified: all 10 files pass (180 tests), npm run typecheck clean.
The "never writes hooks for a remote attach" test stubbed CODEMAN_DATA_DIR
to a separate throwaway dir just for this write, but session-routes.ts's
CODEMAN_CONFIG_DIR is a module-load-time constant frozen at test/setup.ts's
sandboxed dir before this test ever runs. The fixture landed somewhere the
route handler could never read, so the remote-host lookup silently failed
(NOT_FOUND) and the test passed for the wrong reason -- createErrorResponse
never sets reply.code(), so Fastify's default 200 made the NOT_FOUND branch
and the intended success branch indistinguishable by status code alone.
Write straight to getDataDir() instead, matching the docker-hosts fixture
convention already used elsewhere in this file. Verified the fix actually
exercises the success path (host resolves, 200 with a real session), not
just an accidental 200 from the error branch.
PR #356 stopped the remote-hosts.json fixture write from clobbering prod.
Two holes in the same file remain:
1. The quick-start afterEach still ran rmSync(CASES_DIR, recursive).
CASES_DIR is join(homedir(), 'codeman-cases'), and on Linux builds
where os.homedir() reads /etc/passwd instead of $HOME it resolves to
the PROD case tree - so a full-suite run deleted the real
~/codeman-cases. Add a shared safeRmHomeTree() containment gate that
only deletes a path under the redirected test HOME.
2. setup.ts teardown did rmSync(process.env.CODEMAN_DATA_DIR ?? '') AFTER
restoring the env - if a pre-existing prod CODEMAN_DATA_DIR was set,
that deleted prod. Capture the throwaway dir in a const and clean that.
A broader test-isolation sweep (10 files: cli-skill-target, edge-cases,
integration-flows, operation-lightspeed, ralph-integration,
case-clone-routes, voice-routes, session-cleanup, sse-events,
sse-subscription-filter) also applies the same containment gates to every
per-case delete. It is intentionally NOT included here to keep this PR
skinny; it is identified and available on request.
session-routes-workspace-hooks.test.ts wrote its h1/box/10.0.0.5 host
fixture into getDataDir()/remote-hosts.json. getDataDir() resolves via
homedir() → ~/.codeman (INSTANCE_SUFFIX='' by default), and overriding
HOME in test/setup.ts does NOT change os.homedir() on Linux — so every
full-suite run silently overwrote the PRODUCTION remote-hosts.json,
wiping user-defined remote hosts, emptying the launch-case dropdown and
breaking remote session creation (found live 2026-08-29).
The vitest v4 test.env config key is ignored (probe confirmed the
worker still saw CODEMAN_DATA_DIR=undefined), so the reliable fix is
stubbing the env inside the test: the fixture write now goes to a
throwaway /tmp dir via vi.stubEnv + finally unstub. Verified: prod
remote-hosts.json hash is identical before and after the suite run.
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.
Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:
- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
container on the new dist/. This is the one supervisor whose updater does NOT
outlive the restart, which is safe only because the terminal "restarting"
marker is written first.
- node_modules and dist are named volumes over the bind mount, so
container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
`npm run build` is tsc + esbuild and node-pty has no Linux prebuild.
An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.
The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.
Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.
Documented in docs/docker-self-update.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
Adds `claude-fable-5-1` to the App Settings model picker and the five task-routing selects, mirroring how Fable 5 is already offered: a base option with data-ctx="1" plus its [1m] companion row. No settings-ui.js logic change, since the cards and the 1M switch are built from those options.
`test/setup.ts` gives every test file a temp HOME so the suite cannot touch the
real Codeman tree, and strips the env vars that would leak past it — but the
list only covered auth and the gesture flag. The three vars
`src/config/instance.ts` derives the data dir and tmux socket from were missing,
and they reach past the temp HOME:
- **`CODEMAN_DATA_DIR` is the one that matters.** It is an ABSOLUTE override
read in `getDataDir()`, so it bypasses HOME entirely: a developer who exports
it — or a shell left over from `codeman web -d` — has the suite reading and
WRITING their real `state.json`, `users.json`, `intents.json` and
`hook-secret`.
- **`CODEMAN_INSTANCE`** moves the data dir to `~/.codeman-<name>` and the
socket to `codeman-<name>`. Inside the temp HOME that is not data loss, but it
silently changes the paths tests assert on — and `scripts/run-beta.sh` exports
it, so any shell that has run a beta carries it.
- **`CODEMAN_TMUX_SOCKET`** renames the socket `resolveTmuxSocketName()`
returns. `TmuxManager` no-ops its shell commands under vitest, so this is
assertion drift rather than a stray `tmux -L` against prod — same class of
leak, same one-line fix.
They are deleted in the setup file rather than in a hook because
`CODEMAN_INSTANCE` is captured into a module-level const the first time
`config/instance.ts` is imported; a `beforeEach` would already be too late.
`test/test-env-isolation.test.ts` pins the whole list in two halves, because the
obvious half is not enough: asserting the vars are unset passes trivially on a
machine that never set them, so a removed `delete` line would sail through on
almost every box and on CI. The static half reads `setup.ts` and asserts each
name is deleted there, which fails everywhere. An anti-drift check catches the
other direction — a var stripped in `setup.ts` but never given a reason in the
list — and is scoped to the strip section so the teardown's restores are not
mistaken for strips.
Verified by demonstrating the leak: with the `CODEMAN_DATA_DIR` line removed and
the var exported, the runtime assertion fails; with the line restored it passes.
Full suite: no new failures against an upstream/master baseline.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
Review item 4 named THREE live tables duplicating registry data. Two are now
read from the entry (`defaultRemoteCommandForMode`, `defaultDockerCommandForMode`);
the third, `resolveDockerCredentialArtifacts`, is not — and it was left neither
wired nor annotated, which is the state that item explicitly rules out.
It is not wired because the shape cannot express the live table: `credStore` is
ONE store per CLI, and `CRED_STORES` needs two for gemini (`.gemini` for the
CLI's own auth plus `.config/gcloud` for Vertex), while deepseek's entry declares
none at all even though `.dsh` is seeded. Wiring it means making the field an
array and correcting those two entries — a change to credential seeding, which
is at once the worst thing in that file to get wrong and the least covered by
tests, since every docker IO path is no-op'd under vitest. It belongs in its own
change, measured against a real container.
So it is annotated instead, at the field, in the type's declared-for-later
header, in docs/cli-registry.md, and in the pinned DECLARED_FOR_LATER list — the
last of which means wiring it later makes a test fail rather than leaving a
stale comment behind.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
CI caught three cron-service failures. Both are mine, from converting cron's
per-mode ladders to capability reads without checking what each ladder's scope
actually was.
**The pre-flight.** cron only ever pre-flighted `deepseek` — dsh is a profile
LAUNCHER, so "installed" is not "runnable" and a bare `dsh` can boot a profile
that cannot drive a pane. I replaced that with an unscoped
`resolveCliLaunchError(mode)`, which pre-flights EVERY mode, so a claude cron
job on a box with no claude binary now failed with "Claude CLI not found"
instead of reaching tmux-manager's own throw. Three tests assert the latter.
It is now gated on `discovery.launcherProfile !== undefined`, which is
byte-identical to the `mode === 'deepseek'` check it replaces and generalises to
the next launcher. The equivalent HTTP-route conversion was already scoped (to
`capabilities.external`, matching what that route has always pre-flighted); I
simply failed to carry the same reasoning across.
**The model.** cron's ladder was `mode !== 'shell' && mode !== 'deepseek'`, and
I read it as `capabilities.model.source === 'claude-settings-file'` — which is
the HTTP route's question, not cron's. There, every external CLI reads its model
from its own config object earlier in the chain, so only claude reaches the
global default; cron has no such config, so the same expression silently
narrowed the default model from eight modes to one. Now `!== 'none'`, which is
exactly the two entries the ladder excluded. Not caught by a test — found by
re-deriving each ladder's scope after the first failure.
Also names a fourth deliberate behaviour change in the changeset, found while
tracing these: `session.ts` carried a hand-written list of modes with no
direct-PTY fallback and OMP was missing from it, though CLAUDE.md's own text
says "all eight require tmux". `requiresMux` comes off the entry now, so an omp
session whose mux creation fails refuses instead of silently starting outside
tmux.
Verified by diffing failing tests BY NAME against an upstream/master baseline,
rather than by file as before — which is how the regression slipped through: the
three new failures landed inside a file already failing for unrelated
Windows-path reasons, and the aggregate count happened to collide.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery
(search dirs, version + identity probes), the launch argv template, env
handling, the `capabilities` flags that replace per-CLI branching, and the
`overlays` that back the remote/docker pane commands. Code that used to ask
"which CLI is this?" reads the entry instead.
Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every
spawn command as a literal string, captured from the hand-written builders
before they were deleted, and `test/location-overlay-commands.test.ts` does the
same for all 20 remote and in-container pane commands.
Config can never contain shell text: an entry declares typed argv tokens,
literals are validated against a safe-word pattern at LOAD time (a bad literal
rejects the whole entry — a silently dropped `--no-approve` is not cosmetic),
and values resolve through patterns NAMED in code, so a user `clis.json` cannot
widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only
in this release.
OMP is included as a registry entry rather than a tenth hand-written builder,
so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of
`buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen
and doctor ladders all drop out.
Guard rails:
- `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id
branching reappears outside `stock.ts`, in any of its four shapes (`===`,
`!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the
negated forms, which is how 36 of them survived an earlier pass. Every
allowlisted branch carries its reason.
- `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities;
deriving one from another shipped the `until=stop`-hangs-on-shell bug.
- `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and
`privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config`
wire field is separate, bridged only by `legacyConfigAliases`. Getting
`privilegedParams[].param` wrong is SILENT — it is the multi-user bypass
clamp's only handle on a CLI's privilege switch, and a wrong name clamps
nothing with no error and no failing test — so `schema.ts` rejects an entry
naming a param it never declared.
- Registry data resolves AT CALL TIME (`sessionModeSchema()`,
`allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs`
thunks). A module-level const freezes at first import, so a CLI enabled while
the server ran moved the run menu but not that surface.
- Six fields are annotated DECLARED-FOR-LATER and read by nothing
(`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/
`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed
rather than measured. A test pins the list so it cannot quietly grow.
Three user-visible changes, all deliberate and named:
- `probeDockerCliVersion()` derives the in-container binary from the registry
rather than assuming it equals the mode name (`antigravity` runs `agy`).
- The remote CLI version probe now covers grok and deepseek, which the
hardcoded map it replaces omitted while its own comment said the rule was
"every mode except shell".
- `codeman doctor`'s CLI rows are generated from the entries, so Claude's
install hint is the install command rather than a docs URL, five CLIs gain
hints they never had, and the row order follows the catalog.
Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars
(matching the `cliId` pattern) before its failure message quotes the value
back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading
the hand-editable `clis.json`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
docker/server.Dockerfile installed Debian's `docker.io` to get a client for the
socket mounted by Compose. That package is the full ENGINE: even with
--no-install-recommends it pulls 15 packages including containerd, runc, dmsetup
and iptables, none of which a container that only talks to a mounted socket can
use, and it ships Docker 20.10.24 (2023).
Copy the CLI and the buildx plugin from the official docker:29-cli image
instead. Measured on the same node:22-bookworm-slim base: 266 MB -> 108 MB, so
158 MB smaller with a current CLI (29.7.2) in place of a two-year-old one.
Three things verified rather than assumed, by building the real image and
running it:
- docker:cli is an ALPINE image, so copying a binary into this Debian one is
only safe because the binaries are static Go builds (ldd: "Not a valid dynamic
program"). In the built image, `docker --version`, `docker ps` and
`docker build` all work against a mounted host socket as the unprivileged
runtime user.
- buildx is copied on purpose. scripts/build-agent-image.mjs shells out to
`docker build` and Codeman auto-builds the agent image on the first Docker
case. Without the plugin that still works today — CLI 29 falls back to the
classic builder, tested — but that builder is deprecated and will be dropped,
so the plugin keeps the path supported.
- docker-compose is NOT copied: Codeman never shells out to it.
Pinned to the 29 major, matching how the base images here are pinned.
The header chip silently shrank from "5h 4% · 7d 52%" to a lone "7d 52%", which
reads as half the feature breaking rather than as an idle window.
Nothing was broken. Claude Code documents `rate_limits.five_hour` as "present
only while the API reports it and its resets_at has not passed", so between
5-hour session windows the key simply leaves the statusline payload. Codeman's
snapshot replaces the Claude half wholesale on every sample, so the segment
disappeared until usage opened a new window. Confirmed against a live 2.1.252
session by capturing real statusline payloads on an isolated tmux socket: the
boot render carries no `rate_limits` at all, and the first post-response render
carries both windows.
The slot now stays, with a dimmed em dash. Claude only: a missing CODEX bucket
means that plan has no such limit rather than an idle window, so those stay
omitted (pinned by the existing test). The placeholder can never stand alone
either — hasWindows() still gates the row, so a provider reporting nothing
renders nothing rather than a row of dashes. The tooltip says "5-hour limit: no
active session window" instead of dropping the line.
Verified in a real browser against a dev instance: the idle chip renders
"5h — · 7d 52%" with the dash at opacity 0.55 in --text-dim while the live
value keeps its green, and the chip holds its shape (100px idle vs 107px with
both windows).
Docker Compose deployment: Codeman runs in a container and spawns Docker cases
as SIBLING containers through the mounted host socket (Docker-outside-of-Docker).
Resolved the README conflict (master had grown to eight CLIs since the branch
was cut) and moved the Compose blurb out of the feature bullets into Quick
Start, next to the other ways of starting Codeman.
Three review findings from the PR discussion are fixed here rather than left
for a follow-up, because two of them are shipped-image problems:
- `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against
the whole context-relative path, so `docker/.env` — which the deployment's own
README tells the user to fill with CODEMAN_PASSWORD and provider API keys —
was picked up by `COPY . .` and baked into the image at
/opt/codeman/docker/.env. Verified in both directions against a real build
context: with a canary secret in docker/.env, the unfixed ignore file lets
/ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for
the checked-in .env.example) leaves only the example behind.
- `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which
still hardcoded ~/codeman-cases, so `codeman skill install --case <name>`
reported "Case not found" on exactly the deployment the override exists for.
Both now resolve through config/cases-dir.ts. state-store.ts keeps its own
literal on purpose: that one migrates the historical ~/claudeman-cases
directory by name and is about the old default, not the active location.
- CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the
three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore
joins the documented list of files that genuinely belong in the repo root.
The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug
fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/
deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's
lastClaudeSessionId was handed to every non-claude CLI.
Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax,
public assets, lockfile.
Which conversation a pane is on was re-derived by correlating
~/.claude/history.jsonl against Session.lastSubmitAt — and lastSubmitAt is
bumped only by input that flows through Codeman's own write path
(Session.write / writeViaMux). A user who attaches to the pane's tmux session
directly never set it, so resolveActiveClaudeSessionIdFromHistory() returned at
its first line for that pane's whole life and the response viewer stayed pinned
to the launch conversation, showing a pre-/clear transcript indefinitely.
A UserPromptSubmit hook reports the live conversation id from inside the CLI
process, delivered under the pane's own $CODEMAN_SESSION_ID. That binding is a
fact rather than a correlation: it never consults workingDir, so it cannot be
claimed by a sibling pane on the same folder, a closed tab, or a bare `claude`
in the user's terminal. A pane holding such an id skips the correlation
entirely, so the number of prompts eligible for cwd-based guessing goes DOWN,
never up — the naive alternative (relax the guard, or synthesize an anchor from
PTY activity) is the reverted bug the resolver's own comment describes.
The hook also stamps lastSubmitAt, so it finally means "a prompt was submitted"
rather than "typed into Codeman's web terminal". Conversations vouched for
first-hand — and only those — extend a persisted claudeSessionChain, whose tail
re-pins the conversation when a surviving tmux session is re-attached after a
restart. ⚠️ start() resets the id at THREE points and the last one runs
unconditionally after the mux branch, so the tail is applied there too; patching
only the mux branch looks right and silently does nothing.
⚠️ The hook's stdout is discarded with curl's own -o /dev/null. Claude Code
injects a UserPromptSubmit hook's stdout into the model's context ("Exit code 0
- stdout shown to Claude"), and a trailing >/dev/null does NOT work: curlCmd
already ends `... 2>/dev/null || true`, and in `pipeline || true >/dev/null` the
shell binds the redirection to `true`, which never runs on the success path. The
discard is opt-in so the five SSE-fed events keep byte-identical command text
and no workspace's settings file is rewritten for them. The staleness marker is
quote-free for the matching reason: hooksJson is JSON.stringify'd, so a quoted
needle never matches and the gate would rewrite every workspace on every spawn.
Existing workspaces heal on their next Claude spawn through the staleness sweep.
The Claude reader concatenated every assistant row between two human prompts
into one card, fusing up to 74 distinct model messages into a single card, and
it never read the attachment rows that hold a prompt typed while the agent was
working. Measured over 57 real transcripts on 2026-09-01, the viewer shows
1,806 messages instead of 356 and 353 user cards instead of 178, with the
assistant text sequence unchanged row for row and the response without
?context=full byte-identical on all 57 files.
One assistant row IS one whole model message: in that corpus no assistant row
carries more than one content block and no message id carries more than one
text block, so there was nothing to reassemble. Each row becomes its own
message carrying an additive {kind, label, turn}, and the frontend renders a
same-role run inside one turn as badge-less continuation segments — which is
what keeps a p90 of 11 messages per turn from reading as card spam. A numeric
turn gates that rendering, so Codex, the external-CLI pane parser and an older
server keep one badge per card.
A prompt typed while Claude is working is recorded ONLY as an
attachment/queued_command row. Taking it when origin.kind is 'human' and
commandMode is 'prompt' recovers 162 user cards from 163 such rows — one is a
verbatim repeat inside an unanswered user run and is collapsed by the existing
dedup guard — and restores the turn boundary whose absence let the assistant
runs fuse. The CLI's own queue entries are cleanly separable: of 322
queued_command rows, 159 are commandMode 'task-notification' and not one of
them carries an origin key.
This narrows #169 rather than reverting it: sidechain exclusion, the
restored-<uuid8> rebind, replayed-snapshot dedup and synthetic-row filtering
are all unchanged and still asserted.
mobile.css hides #createCaseModal's .set-foot below 860px, and that modal's
header — unlike Settings' — carries no set-head-save. So on a phone the
Create/Link button existed nowhere and the modal could not be submitted at all.
Adds the header button and drives both together through switchCaseModalTab()
and submitCaseModal(), so whichever one is pressed the other shows the same
pending state and is equally unclickable. Following the Settings pattern also
means Add Case picks up the existing .set-head-actions:has(.set-head-save) tray
and .set-head-save sizing with no new CSS; the mobile.css comment that still
listed Add Case as a lone-× sheet is corrected to match.
1.24.2. Also corrects two numbers in the CLAUDE.md trust-dialog paragraph that
was written while the fix was still uncommitted: the keystroke cap is 6, not 3,
and the scan now schedules its own follow-up read rather than waiting on PTY
output that a static dialog never produces.
Every claude session in a directory claude had not seen before died about six
seconds after it started (`Pane is dead (status 1)`), before the agent drew a
composer. Reproduced on a fresh case and measured.
Claude Code 2.1.252 rewrote the dialog. It used to be
❯ 1. Yes, I trust this folder
2. No, exit
and is now unnumbered, reversed, and highlights the option that quits:
❯ No, exit
Yes, I trust this folder
Detection still worked (the confirm affordance carries the match once the
numbered option text is gone), so the failure was entirely in the answer: the
auto-accept pressed Enter on the highlighted default, which is now exit.
trustDialogNextKey() reads the ❯ marker off the rendered pane and returns ONE
keystroke at a time: an arrow while the cursor is on the wrong option, Enter
only once the screen shows it on the trust option, and null for a frame that
does not say. Both layouts are handled, and which way the trust option lies is
read from the frame rather than assumed, so a further reordering costs a
repaint instead of a session. The last marked option wins, because the
direct-PTY fallback reads an append-only buffer where an older frame must not
out-vote the freshest one.
Two things only a live pane showed:
- The scan ran solely from the PTY onData handler. The arrow that moves the
cursor is the last output the pane produces, so the first fix parked every
session with the cursor sitting on the right option and no Enter ever sent.
It now schedules its own follow-up read (_trustDialogTimer, cleared in
_clearAllTimers()), offset past the scan throttle so the chain cannot break
on a boundary.
- The keystroke cap goes 3 -> 6, since answering is no longer one press.
The bundled codeman skill had the same blind \r as its bounded fallback, so
preamble 1.21.0 replaces it with _trust_key/_accept_trust: read
terminal?full=1, steer onto the trust option, re-read, then confirm. Those
keystrokes go out under their own clientId, because input sequence numbers are
monotonic per client and spending prompt numbers on dialog keys would make the
next send-and-wait look like a stale duplicate and vanish while reporting
success. The readiness recipes in docs/extending-codeman.md,
docs/api-reference.md and the skill's own reference carry the corrected answer,
plus a symptom-table entry for a worker whose pane is dead seconds after spawn.
Verified live on an isolated instance (own data dir and tmux socket): fresh
case -> arrow at 5 s -> Enter at 7 s -> composer, with hasTrustDialogAccepted
recorded. With the server-side auto-accept disabled in a throwaway copy, the
skill's fallback cleared a genuinely parked dialog in 1.1 s and spawn_worker
took a brand-new case to a live composer in 7.2 s; spawn_workers + sendwait +
last_text then ran end to end.
Audit of CLAUDE.md against the tree. Verified still accurate: the 31-module
frontend load order (matches index.html exactly), SSE registry parity at
157 = 157 (confirmed by running the parity test), config/ 21 files, types/ 22
domain files, 136 mobile device profiles, the version line, and every Quick
Reference command.
Drift corrected: 24 route modules to 25, ~220 handlers to ~227, system-routes
51 to 56, app.js ~5K lines to ~6.7K and 30 modules to 31, install.sh 92KB to
104KB. Completed the CLI resolver inventory, which was missing
deepseek-cli-resolver and omp-cli-resolver even though both modes are
documented, and named the shared cli-executable-resolver lookup chain.
Filled the gaps found by sweeping every src module against the file:
- Owner tab layouts (COD-359) had 6 source modules, 7 test files, 2 routes, an
SSE event and a state.json key, with zero mentions anywhere in CLAUDE.md or
docs/. The paragraph records the four things a reader would otherwise get
wrong: it is backend-only as of 1.24.1 with no frontend consumer, the service
is the sole mutation boundary, it projects onto PUT /api/session-order rather
than replacing it, and reconciliation is gated on a successful restore.
- codeman doctor and codeman users were undocumented top-level CLI commands.
- Four subsystems whose invariants lived only in their @fileoverview:
the workspace-trust dialog recognizer, proc-tree's bounded walk (the
2026-07-30 incident that took a machine down), deepseek-web-server (one
child process, deliberately not a shell session), and the Files panel
search matcher (globs are never compiled to a RegExp).
Also fixes a stale "156 event types" comment in constants.js (actual: 157) and
a contradiction in AGENTS.md, which still carried the retired "never run the
full suite inside a managed tmux session" rule against CLAUDE.md's current
"npm test is the gate and is safe to run bare".
Note: the trust-dialog paragraph documents trustDialogNextKey(), which is part
of a sibling session's in-flight fix for the Claude Code 2.1.252 layout change
(unnumbered, reversed options with "No, exit" highlighted, so a blind carriage
return picks exit and kills the pane). That fix was uncommitted in the shared
tree when this landed, so the doc leads the code until it is committed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NMN8UuvdBim3iM87reuQ9Z
Follow-up to #350, which fixed the actual blocker (issue #352): `dsh plugin` is
a thin forwarder that `spawnSync`s a literal `pnpm` with no npm fallback, so an
image without pnpm dies at exit 127 and takes the whole build with it.
That PR also pinned an allowlist of the two packages whose lifecycle scripts
pnpm blocked at the time. Replace it with a policy that cannot go stale: pnpm,
unlike npm, refuses dependency build scripts by default and FAILS the install
over it (`ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on pnpm 11.24), and the
names to allow move between rebuilds because `@deepseek-harness-tui/dsh-tui` is
resolved by dist-tag, not pinned: 0.9.3 pulled `@google/genai` (whose script is
a literal `preinstall: no-op`), 0.10.0-beta.x does not. An allowlist of two
names would have let the next tree break the build the same way. Allowing them
wholesale is also the exposure this image already accepts three layers up,
where `npm install -g` runs the install scripts of every transitive dep of the
five CLIs above with no gate at all.
Also correct a comment in the `/api/deepseek/install-profile` route that
asserted the opposite of what #352 proved ("dsh bundles its own package
manager, so no system pnpm is required"). The route's behavior is already
right: dsh's own "pnpm not found on PATH" stderr reaches the caller as the
OPERATION_FAILED detail, so the UI's "add a terminal profile" button names the
fix. Documented the prerequisite in docs/deepseek-integration.md, and taught
the docker-cases image smoke test about `dsh`/`omp` plus the profile check that
`dsh --version` does NOT cover.
Format check failed twice, on different files each time, because three prettier
versions were in play: package.json says ^3.4.0, package-lock pins 3.8.3 (CI runs
npm ci, so that is the one CI uses), and the local node_modules had 3.9.6. Files
formatted with 3.9.6 were then "fixed" with 3.4.2, pushing session-routes and
system-routes onto a third style — every version change moved the failure to a
different set of files.
Line-break placement in `await import` and a union type only; no logic changes.
The run menu still offered every mode for an attached container. The browser's
actual request showed why:
POST /api/docker-cases/adopt-preflight -> 400
{"error":"Invalid input: expected object, received string"}
_api serializes `body` and sets Content-Type itself, and three call sites each
passed an already-stringified body, so it was encoded twice and the server saw a
JSON string where it expects an object. curl was fine throughout, so nothing in
the server logs pointed at it.
Also fixes the design defect underneath: a failed probe fell through to "do not
gate", which silently offered every mode. When the container has been recreated,
is stopped, or the engine is unreachable, the user sees claude, clicks it, and
it can only fail — with the reason visible nowhere. A failed probe now hides
every agent mode (Shell needs no CLI and stays) and shows the server's own
reason at the top of the menu.
Two static guards switched from a character window to brace matching. They
sliced between two call sites, and _loadRunModeHistory's call appears above its
definition, so the slice came out empty and the assertion verified nothing —
the same trap twice in one file.
Both paths in the adoption form had to be typed. Each gets a Browse button
using the same path-input-group markup Link Existing uses, so the two look and
behave alike.
What they can browse differs, and that is the point. The host workspace path
reuses the existing host picker. The container workdir cannot: an adopted
container has nothing mounted at a matching host path, so a host listing would
be a different filesystem — and getting this field wrong is the source of the
opaque OCI chdir error at launch, which makes it the field that most needs to
be clickable.
Adds a read-only POST /api/docker-cases/browse: one `ls` through docker exec, no
writes, no lifecycle, path shell-escaped like every other value. `ls -Ap` marks
directories with a trailing slash and keeps names with spaces intact.
PathPicker takes an optional fetchListing source rather than being forked: the
container variant only swaps where the rows come from, and reuses the rendering,
navigation, Up and Choose/Select unchanged.
Link Existing's Browse did nothing: GET /api/filesystem/browse answered 403
"No filesystem browse roots are available".
Two rules were fighting. /root is a default blocked tree in the attachment
guard, and Codeman running as root — containers, plenty of servers — makes
homedir() exactly /root, so the picker's own allowlisted Home root was blocked;
the other candidates live under it or do not exist. The root list came out
empty and there was nothing the user could open.
The blocked trees exist to keep ~/.ssh and friends out of reach, not to seal off
the user's own home. Only trees that would swallow a configured root whole are
dropped now: /root goes when Home is it (or sits inside it), /etc holds no
configured root and is untouched. Secrets stay protected — isSensitivePath
independently matches .ssh/, .env and credentials* at any depth, and it is what
the directory probe asks about.
⚠️ Navigation must reuse the same narrowed list the roots were chosen with.
Handing the raw trees downstream admits a root and then refuses every path
inside it, which reads as a picker that opens and does nothing.
Attaching a container, picking claude and hitting Run gave one line —
`execvp(3) failed.: No such file or directory` — and the run-mode menu offered
every mode. Three separate defects, found on a real deployment.
TmuxManager.createSession resolved the CLI directory without distinguishing a
docker session, so a host with no claude threw, the catch fell back to a direct
PTY, and that PTY exec'd the CLI on the HOST. The failure surfaced as a bare
execvp error naming nothing. A docker session runs its CLI inside the container;
the host does not need it. All eight modes now sit behind a cliRunsInContainer
guard, and whether the container has the CLI is settled by the adoption
preflight or the image gate before launch.
The running check used a bare double quote and command substitution. The whole
chain is embedded in an outer `bash -c "…"`, so the unescaped quote closed that
string early and the remainder was re-tokenized. It is now a `grep -qx` pipeline
using only the single-quote form every other line in the builder already uses.
Claude Code refuses --dangerously-skip-permissions as root. Our base image runs
a non-root user, so an owned container never hit this; an adopted container's
user belongs to its owner and is frequently root, and keeping the flag killed
the pane with a message visible only inside the container. The preflight now
reports runsAsRoot and the launch chain drops the flag for it.
The menu also showed every mode because the container CLI probe only started
when the menu opened. It is warmed when the case is selected instead.
Storing the container's CLIs on the case at attach time left two gaps: a case
linked before that field existed has none at all, and a container's CLIs can be
installed or removed long after it was linked. A real deployment hit the first
one — the host had only codex, the container only claude, and with no stored
list the menu still gated on the host and hid the mode that actually worked.
The probe now runs when a container case is selected, reusing the existing
adopt-preflight endpoint, so there is no new backend surface. Results are cached
per case for the page's lifetime, since the menu opens often and the probe is a
`docker exec` round trip; a concurrent probe for the same case is deduplicated
with an in-flight marker.
A failed probe leaves the cache empty, which the caller reads as "unknown" and
therefore does not gate. Hiding every mode because one probe failed is worse
than offering one that turns out to be missing, which the launch path already
refuses with a specific message.
The repaint only happens while the menu is still open, so a late answer cannot
make the list jump under a user who already closed it.
The adoption preflight used the mode name as the binary name. claude, codex,
opencode, gemini and pi happen to match, so it never showed — but antigravity
ships as `agy` and deepseek as `dsh`, so a container that has either was
reported as not having it, and the mode was silently dropped from the case.
Adds a MODE_BINARIES map, single-sourced with defaultDockerCommandForMode, which
launches those same binaries. Probing and result filtering share one `binaryFor`
so the two cannot drift apart.
The run-mode dropdown hides CLIs that are not installed on the HOST (#201). That
is right for local sessions and wrong for a container case, whose agents run
inside the container: a host with no claude installed hides the mode while the
container ships one, which is exactly what happened on a real deployment.
The adoption preflight already probes what the container has, so that result is
persisted on the case and surfaced through CaseInfo. Docker cases gate on it;
every other case keeps the host probe unchanged.
An absent list reads as "do not gate" rather than "nothing available": an owned
container runs our base image, which ships every CLI, and treating unknown as
empty would leave the menu with Shell alone.
Typing a container name from memory is error-prone. The field becomes a native
datalist: pick from the engine's containers, type to filter, or type a name that
is not listed (the engine may be remote, or the container may not exist yet).
A datalist gives all three natively, so no dropdown state machine is introduced.
Adds listDockerContainers and GET /api/docker-hosts/:hostId/containers, following
the listRemoteCodemanSessions discovery precedent: read-only and never throwing,
so an unreachable daemon returns an empty list and the field degrades to plain
text instead of erroring.
Stopped containers stay in the list, sorted after running ones and labelled.
Attaching does require a running container, but hiding stopped ones turns "my
container is not in the list" into a dead end, while showing
`Exited (137) 8 days ago` says exactly what to fix.
The new strings were English only. Adding entries surfaced a deeper problem: the
translator matches whole text nodes and skips `code`/`pre`, so an inline `<code>`
mid-sentence splits a hint into fragments that can never match an entry — which is
why the panel's existing "Build it once with <code>...</code>" hint was never
translated either.
Drops the inline markup from the new hints so each is a single text node, then
adds the zh-CN entries. The brand name goes through the existing {name}
placeholder.
Server-side error bodies are deliberately not added: the client receives them
already interpolated with a concrete container name, so a template key could
never match.
Attaching lived only on the Docker tab, but the place users look for anything
container-shaped is the "Run in an isolated Docker container" checkbox on Create
New. A feature nobody can find is a feature nobody has.
Adds a one-click link there that switches to the Docker tab, turns the toggle on
and focuses the container field. Reuses switchCaseModalTab and the existing sync
helper; no new CSS.
Two defects that only a real container exposes.
The probe chained `command -v X && echo X` with semicolons, and a script's exit
status is its last command's. A container without the last probed CLI made the
whole `sh -lc` exit 1, so a perfectly healthy container with tmux and claude was
reported as "could not exec into the container". A missing CLI is data here, not
failure, so the script now ends with `exit 0`.
containerWorkdir defaulted to hostWorkspacePath. That default holds for an owned
container only because the create-time bind mount puts the host directory at that
exact path; attaching mounts nothing, so the two are independent facts. A host
path absent inside the container makes `docker exec --workdir` fail with an OCI
chdir error that surfaces in the pane as a bare "execvp failed". The preflight now
proves the directory exists inside the container and refuses at link time.
The Docker tab gains an "Attach to an existing container" toggle. Ticking it
swaps the create-time fields (image, network, advanced) — which describe a
`docker create` attaching never runs — for the container name, and routes the
submit to the adopt endpoint.
Reuses the existing linkDockerCase flow end to end: only the final call differs.
The docker-host upsert still applies, since it is what resolves the
engine/context/daemon for `docker exec`; its create-time fields are simply never
read for an attached case.
Docker cases could only run in a container Codeman created itself. Attaching to
one the user already built and runs means Codeman must leave that container's
lifecycle completely alone, which the launch chain could not do: it was
`image inspect` -> `inspect || create` -> `start` -> `exec`.
Adds `DockerCase.owned`, mirroring the `owned:false` contract remote-SSH already
uses for attached sessions. Absent (every existing case) means owned, so current
behaviour is byte-identical. `false` means the container belongs to the user and
Codeman may only exec into it.
The launch chain for an attached container only looks, then execs: no image gate
(the image is theirs), no create, and no `start` — starting a container we do not
own is the very mutation attaching promises not to perform. A missing or stopped
container fails closed with an actionable message instead. Credential seeding is
skipped too: those copies read from create-time read-only mounts that do not
exist here, and writing host credentials into someone's container is not ours to
do, so its CLIs must already be authenticated inside it.
Four fail-closed guards. buildDockerStopCommand and buildDockerRemoveCommand
throw during pure string construction, so no caller bug can turn into a
`docker stop`/`rm` on a container we do not own; removeDockerContainer refuses
again at the lowest layer; drift reports "none" for an attached container, which
carries no `codeman.confighash` label and would otherwise always look drifted and
409 the launch gate forever; and the orphan reaper skips attached containers
through a check deliberately independent of the two conditions already covering
them.
`owned` is applied AFTER the config hash is computed. dockerConfigHash takes an
explicit field list, so ownership can never shift an existing case's hash — if it
did, every pre-existing case would trip the drift gate at once, and the remedy
the UI offers is "recreate the container".
Adds POST /api/cases/docker-adopt and a read-only
POST /api/docker-cases/adopt-preflight. The preflight refuses at LINK time rather
than at session launch, where the only ways out would be a dead pane or starting
a container we do not own.
Tests assert the negative guarantee directly — that create, start, stop, rm,
restart and kill are absent from the generated commands while `docker exec -it`
and `new-session -A` remain — since it cannot be observed by using the feature.
The COD-108 reconnect watcher treated any dead local pane as a dropped
transport and re-ran the pane command — so a normal ctrl-c/ctrl-d on a
remote claude/opencode/omp auto-spawned a FRESH agent (claude only
looked correct because its '--session-id || --resume' fallback resumed,
with a loud 'already in use' error first).
Distinguish a transport drop from an intentional exit: only reconnect
when the durable remote tmux session (codeman-ssh-*) is verifiably
still alive on the remote host. A clean exit tears that session down;
the watcher now probes it via ssh has-session and skips (remote-gone)
when it is gone OR unknown (fail closed). The probe is cached
per-session and fired async so the 5s tick never blocks on ssh.
Tests: 3 new cases pinning remote-gone / unknown / alive decisions.
Verified live: all remote CLIs stay dead after ctrl-c/ctrl-d.
Every other CLI (claude/opencode/codex/gemini/antigravity/pi/grok/dsh) has
a check_*/get_*_path pair wired into install.sh's detection loop and the
"no AI CLI found" aggregate checks. OMP had neither -- a user with only
omp installed would be told no CLI was found and offered to install
Claude Code or OpenCode.
Added OMP_SEARCH_PATHS (mirrors src/utils/omp-cli-resolver.ts's
OMP_SEARCH_DIRS) and check_omp()/get_omp_path(), wired into both
aggregate conditions (the interactive install-menu trigger and the
end-of-run reminder) and added omp's real vendor curl one-liner to the
reminder block. The DeepSeek Harness line was never in that reminder to
begin with -- confirmed it has no vendor one-liner (dsh installs via
Codeman's own API after the server is already up), so it stays out, with
an explanatory line instead.
Also fixed the "Skip" menu text, which was missing Gemini and DeepSeek
Harness from its example list independent of the omp gap, and the same
stale sibling-CLI-list bug (missing DeepSeek Harness and OMP, "the
eight"/"这七个") in README.md and the repo's existing README.zh-CN.md.
Small cleanup items from upstream review (Ark0N/Codeman#353):
- OMP_SEARCH_DIRS now leads with ~/.local/bin, matching omp.sh's real
installer target (~/.omp/bin was an earlier unverified guess, confirmed
wrong against a real --no-cache Docker build).
- docs/omp-integration.md: fixed the dead GitHub URL (can1357/omp ->
can1357/oh-my-pi), corrected the CLI count (ninth backend, tenth
SessionMode incl. shell -- not eighth), matched the install-path guidance
to the resolver fix, updated the version example to the actually-tested
18.0.8, and added a Docker-section caveat: --resume pinning does not
currently reach an in-container omp process, since Docker panes never see
ompConfig.
- docs/architecture-invariants.md: fixed a heading missing ", OMP" (CLAUDE.md
already linked to the -omp anchor, so the link was dead) and added an OMP
specifics paragraph -- the one external CLI missing an entry in this doc.
- .changeset/omp-backend.md: corrected the sibling-CLI list (was missing Pi,
Grok, and DeepSeek Harness) and the backend count.
- Removed a stray orphaned comment fragment in the quick-start docker branch
and split two CSS lines that had two declarations jammed onto one line.
resumeHistorySession() creates the resumed row in its own mode via a
modeConfigKey map (opencode/pi/grok/omp -> continueSession, deepseek ->
resumeSession) and retires the old row afterward. codex, gemini and
antigravity were missing from that map, so resuming one of their rows
started a brand-new session with NO continuation while still deleting
the row it came from -- silent data loss dressed as the duplicate-row
fix. Gate row retirement on continuesSomething (true only for modes that
actually got a continuation config) instead of wiring an unverified
sessionId->native-conversation-id assumption for the three affected CLIs.
DELETE /api/sessions/:id reimplemented the ownership 404 check inline in
two places instead of going through findSessionOrFail, and its
persisted-only-session branch never broadcast session:deleted, so other
open tabs kept the retired row until their next unrelated fetch. Extract
the shared 404 into sessionNotFoundError(), add findPersistedSessionOrFail()
alongside findSessionOrFail() in route-helpers.ts (same ownership
contract, returns a SessionState instead of a live Session), and use both
from the route instead of inline checks. Add the missing broadcast.
The docs claimed omp "has no documented vendor-key namespace of its own"
and "the multi-user clamp has nothing to gate" for omp — both false. Per
omp's own docs/environment-variables.md, it reads ~40 provider keys from
env (pi's known 34-key problem in the same shape), and its own knobs are
mostly PI_* (already globally allowlisted): PI_CONFIG_DIR,
PI_CODING_AGENT_DIR, PI_CODING_AGENT_SESSION_DIR, PI_SUBPROCESS_CMD,
PI_SHELL_PREFIX. The first three also move the ~/.omp tree
omp-session-resolver.ts/omp-transcript.ts hardcode, silently degrading
pinning/history — a known gap shared with pi, documented but not fixed
here.
The OMP_* prefix this PR adds brings in OMP_AUTH_BROKER_URL/
OMP_AUTH_BROKER_TOKEN, where omp resolves credentials from — the same
shape DEEPSEEK_BASE_URL is already dropped for in
clampEnvOverridesForOwner(). Add both to OWNER_CLAMPED_ENV_KEYS so a
non-granted owner in multi-user mode can't redirect them, and correct the
false claims in CLAUDE.md, docs/omp-integration.md, and the stale
resolveOmpHome() comment. Also documents omp's default
tools.approvalMode: yolo, which was previously unstated.
findLatestOmpSessionId()'s newest-mtime pin ran eagerly inside
_buildRespawnPaneOptions(), which startInteractive() calls unconditionally
on every boot-recovery reattach — before anything checks whether the pane
is actually dead. With two omp tabs in the same case dir, this could pin
an ALIVE pane's session onto whichever sibling's file happened to be
newest on disk, purely as a side effect of building options that might
never lead to a respawn (reported in Ark0N/Codeman#353 review).
Move resolution out of the eager builder into _pinOmpRespawnId(), called
explicitly only where a respawn is actually confirmed: the dead-pane
branch in _setupOrAttachMuxSession() and reattachRemote(). Add
resolveAndClaimOmpSessionId(), which verifies each candidate's own file
header (cwd) rather than trusting the mangled-directory match alone, and
tracks claimed ids in a process-wide registry so two ambiguous resolutions
can't both pick the same sibling's conversation.
Found live 2026-08-27 by Tim: clicking Run OMP to start a brand-new session
in a case directory with prior omp history launched --resume <old-id>
instead of a clean `omp` invocation.
Root cause: Session._resolvedOmpRespawnConfig() resolves-and-pins the
newest on-disk omp conversation as a side effect on this._ompConfig. That
is correct when reattaching to an ALREADY-TRACKED mux session (a dead-pane
respawn, or a boot-recovery reattach - the constructor sets _muxSession
from persisted state before startInteractive() ever runs there), but it
ran unconditionally. startInteractive() computes
`respawnPaneOptions: this._buildRespawnPaneOptions()` eagerly in the same
object literal that builds `createSessionOptions.ompConfig: this._ompConfig`,
so for a genuinely brand-new session (no muxSession in its create config,
_muxSession still null) the resolve-and-pin side effect ran and poisoned
this._ompConfig before that field was even read.
Fix: gate the resolve-and-pin logic on `this._muxSession` already being
set. A fresh session has no muxSession yet and now passes through
untouched; a real reattach (muxSession present since construction) keeps
resolving and pinning exactly as before.
Verified live in production against the exact reported scenario (a fresh
omp session in a case dir with 8+ hours of prior omp history) - confirmed
both via the API (ompConfig stays empty, claudeSessionId equals the
session's own id) and visually in the GUI. Regression test constructs a
real Session + TmuxManager to exercise the actual private-method
interaction directly, since no existing test called startInteractive() at
all.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
OMP was the one external CLI mode with no dedicated user-guide doc, unlike
opencode/pi/grok/deepseek which each have one. Covers install, auth (omp
owns its own entirely - no Codeman-side login flow or bypass switch),
what Codeman wires up (OmpConfig), the exact-id pinning mechanism and the
directory-mangling bug behind it, kill-survival via transcript scanning,
terminal behavior, Docker/remote-SSH cases, and known gaps (no idle hook,
mid-turn kill data loss, unverified symlinked-$HOME behavior).
Cross-referenced from README.md's Multi-CLI doc list and docs/docker-cases.md's
credential-seeding summary (which now also documents OMP's sessions/-is-shared
exception to the seed-everything pattern the other CLIs use).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
OMP had full routing at the Docker layer (default pane command, schema) but
was never actually installed in docker/agent.Dockerfile, and had no
credential-isolation entry in docker-hosts.ts's CRED_STORES - a Docker-mode
OMP session would have failed with "omp: command not found", and even with
the binary present would have had no config/auth seeded, despite the README
already claiming OMP has "seamless auth, isolated credentials" in Docker.
- docker/agent.Dockerfile: install omp via its own installer (standalone
binary, same shape as grok/antigravity - not on npm). Verified against a
real --no-cache build: the installer actually targets ~/.local/bin, not
~/.omp/bin as the resolver's OMP_SEARCH_DIRS ordering would suggest -
confirmed omp/18.0.8 installs and runs correctly inside the image.
- src/docker-hosts.ts: add a .omp/agent CRED_STORES entry. Unlike every
sibling CLI in this family, sessions/ is SHARED (RW), not seeded: Codeman
reads ~/.omp/agent/sessions/**/*.jsonl host-side for history recovery and
--resume pinning (omp-transcript.ts, omp-session-resolver.ts), the same
reason codex's sessions/ is shared rather than seeded. Seeding it instead
would silently break the kill-survival feature for Docker cases. Only the
small config files (config.yml/mcp.json/models.yml/settings.yml) are
seeded; the SQLite caches and terminal-sessions/ stay container-local.
- test/docker-hosts.test.ts: pin the new CRED_STORES entry's behavior.
Found in passing (NOT fixed here, unrelated and pre-existing on master): the
agent image's DeepSeek (dsh) plugin-install step currently fails on a fresh
build ("pnpm not found on PATH"), confirmed via git diff against
origin/master that this line is untouched by this branch. Worth a separate
issue/PR.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Follow-up from a full-branch review pass (Opus) of the omp-mode integration:
- Add pinning tests for resolveOmpConfigForCreate() (session-routes.ts),
exported to make it testable: the exact "resume this OMP row from
history" pipeline that mangleOmpWorkingDir's earlier bug lived in had
zero coverage despite being the resolver module's whole reason to exist.
- Log a warning when findLatestOmpSessionId() finds nothing on disk and
continuation silently degrades to omp's own ambiguous --continue,
in both call sites (session create and respawn pinning) - previously
silent, making the degradation invisible to anyone debugging it.
- Require an absolute cwd before trusting a session file's working
directory in omp-transcript.ts's parser, so a corrupted/malformed
session file can't point a downstream resume at a relative or empty
path.
- Document (don't speculatively fix) an unverified symlinked-$HOME edge
case in mangleOmpWorkingDir(): the review's suggested realpath() fix
assumes omp itself resolves symlinks before mangling, which is
unconfirmed - guessing wrong there would trade one silent mismatch
for a different one.
- Incidental: fixed unrelated pre-existing prettier drift in
session-routes.ts (antigravity/opencode dynamic import line-wrapping)
that was blocking the pre-commit formatting gate on this file.
Confirmed as a non-issue: the model-name regex allowing "/" is
intentional (provider/model ids like "crof/glm-5.2" were used
successfully in live testing).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two bugs compounded to break continuation pinning on every real OMP
case (only /tmp-based manual testing happened to work by coincidence):
1. startInteractive() had a second, unconditional claudeSessionId
assignment after the mux branch that clobbered its correctly
resolved value back to the session's own id on every mux path.
2. mangleOmpWorkingDir() assumed omp mirrors Claude Code's directory
naming (home prefix kept), but omp actually strips $HOME first.
findLatestOmpSessionId() was silently returning null for every
case under ~/codeman-cases/, so resumeSessionId never resolved for
any real case dir - only /tmp paths (outside $HOME) worked, which
is every dir this feature was previously tested against.
Verified live: killed and relaunched the omp-verify server process
mid-session (plain reattach, pane stayed alive) and confirmed
claudeSessionId now resolves to the real omp transcript uuid instead
of the Codeman session's own id.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude conversations survive "Kill Tmux & Claude" because Codeman reads
them back independently from ~/.claude/projects, not from its own
session bookkeeping. omp conversations had no equivalent: kill the
Codeman session and the conversation vanished from Past Sessions
entirely, even though omp itself never forgot it on disk.
Adds omp-transcript.ts, a scanner over omp's own
~/.omp/agent/sessions/<mangled-cwd>/<uuid>.jsonl files (the same shape
as Claude Code's own transcript scanner, but simpler -- these files are
small enough to read whole instead of doing head/tail windows). Each
file's own "session" header line carries the real cwd and session id
directly, so unlike Claude's mangled-directory-name decoding this
never has to guess. Wired into gatherUnifiedInputs() as a second
history source alongside the Claude scan, and HistoryInput/
mergeUnifiedSessions() now carry an optional `mode` so a non-claude
history-only row still gets a real mode badge.
Also fixes the ambiguity behind the "continue picks the wrong
conversation" report from this session's testing: omp mints its OWN
session uuid, unrelated to Codeman's, so a live/persisted row and its
own history-scan row would otherwise show up as two separate entries
for the same conversation the moment the id gets resolved. Reuses the
existing claudeSessionId alias field (mergeUnifiedSessions' fold-into-
owner mechanism) to point at the resolved omp id, threading it through
every place `_claudeSessionId` gets (re)computed -- the constructor,
_resolvedOmpRespawnConfig, and a new _maybeCaptureOmpSessionId() that
opportunistically resolves it the first time a brand-new omp session
(one that has never gone through a respawn) goes idle.
Also closes a THIRD instance of the "ompConfig never got wired in
here" gap this session kept finding: restoreMuxSessions() in server.ts
restores every sibling CLI's config from persisted state on boot except
omp's, so a boot-recovered omp session always lost its resolved resume
id and fell back to guessing again.
Verified live end-to-end: told a session a secret, killed it fully
(Kill Tmux equivalent, killMux=true -- the Codeman session AND its tmux
pane both gone), and the conversation still showed up in the unified
list as a history-sourced row with the real first prompt as its title
and an omp mode badge, keyed by omp's own session id.
Known remaining gap, not fixed here: the claudeSessionId alias doesn't
yet resolve reliably on every boot-recovery path for a session that
was never respawned while alive (e.g. a plain re-attach to a pane that
was never dead) -- worth a follow-up, but doesn't affect the two things
that matter most: the conversation surviving a kill, and continuation
correctness once an id has been resolved (which happens on the very
next respawn either way).
Every non-claude "Resume" click creates a brand-new Codeman session
(there is no id to reattach to), but the old row was never cleaned up
-- click resume on the same conversation a few times and the session
list fills up with duplicate rows sharing one name. resumeHistorySession
now retires the row it resumed from after the new one starts.
That retirement needs DELETE to actually work on a row that was never
live in the first place (the normal case for anything showing up in
"Resume Conversation"): findSessionOrFail only checks the in-memory
live-session map, so DELETE 404s on a persisted-only entry today. Give
the route a fallback: when the id isn't live, look it up in persisted
state instead and demote/remove it there (respecting the existing
pinned-session protection). Verified live against a real persisted-only
row via the API, and added route-test coverage for both the success
and still-truly-unknown-id cases (which needed a demoteOrRemoveSession
mock the route harness didn't have).
Also includes an unrelated pre-existing prettier drift fix picked up
by npm run format (omp-cli-resolver.ts, antigravity/opencode import
wrapping in session-routes.ts).
respawnPane() -- the path used when a session's pane died (crash, idle
respawn, or the user's own /exit) but the Codeman session object is
still tracked -- never had ompConfig wired through at all, in either
its options destructure or its inner buildSpawnCommand() call. This is
a gap in the original OMP patch, distinct from the resumeHistorySession
fix (which only covers a session that has been fully closed and shows
up as a history row): reselecting a tab whose CLI process just exited
goes through this path instead, and always launched a bare, contextless
`omp` no matter what.
Beyond the wiring, respawning a dead pane is semantically different
from creating a brand-new session: the conversation is still "this
session" to the user, so _buildRespawnPaneOptions() now defaults
ompConfig to continueSession:true unless the session already carries
an explicit resumeSessionId (which still wins in buildOmpCommand).
Verified live: told a session a secret, exited OMP so the pane died
(session and tmux both left alone), forced the exact dead-pane-respawn
path, and the new process replied with the secret -- confirming
`omp --continue` fired instead of a blank omp.
resumeHistorySession() never sent mode when recreating a session from a
history/session-manager row, so the server default silently opened a
plain Claude session for every non-claude row -- reproduced live: OMP
rows spawned Claude sessions on click. Thread the row's mode through
every call site (welcome list, session manager, mobile overview) and
only send the Claude-specific resumeSessionId for claude rows.
Codeman has no live PTY-reattach outside server boot, and it's moot for
OMP anyway (exiting it kills the pane's only process), so route the
non-claude relaunch through each CLI's own continue-most-recent flag
instead of a context-free fresh start. OMP never got one: buildOmpCommand
only implemented --model/--resume despite omp --help documenting
-c/--continue. Added continueSession to OmpConfig end-to-end (type,
schema, builder) mirroring the existing opencode/pi/grok/deepseek
fields, and wired resumeHistorySession to use it.
Verified live: told a real omp session a secret, exited it, closed the
tab without killing tmux, relaunched with --continue in the same
directory, and had it recall the secret.
Two prose lists in skills/codeman/ named some but not all external CLI
modes after the omp-mode rebase, which is exactly the drift
test/agent-skill-mode-lists.test.ts exists to catch: SKILL.md's
no-hook-signals list was missing omp, and endpoints.md's version-probe
sentence named pi/grok/omp as a bare 3-mode run with no matching class.
- omp-cli-resolver.ts already uses createCliExecutableResolver; add dedicated
test/omp-cli-resolver.test.ts mirroring pi's (version-probe accept/reject,
negative-cache backoff, VITEST hermeticity gate)
- dependency-registry omp entry now requires OMP_VERSION_REGEX match like pi,
so codeman doctor and the run-mode resolver agree on what counts as installed
- system-routes /api/omp/status surfaces version
The chip prefixes every row with the provider name, so a machine that only
has Claude limits renders "CLAUDE 5H 60% 7D 23%" — a 46px label naming the
only thing it could possibly be. The name exists to tell two rows apart, so
it should only appear when there are two.
updatePlanUsageChip() now checks whether both Claude and Codex actually have
windows before building the rows, and emits the .pu-provider span only in
that case. The tooltip keeps naming the provider in both cases: it has the
room, and the chip no longer does.
Verified in a browser on an isolated beta instance: Claude-only renders bare
windows with no .pu-provider in the DOM, Codex-only the same, and the
two-provider chip is byte-identical to before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Codex plan usage in the header chip (#346), a visible inline rename in
the session sidebar (#345), and the install.sh Tailscale re-run fix plus
the README network-access prompt description.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The installer bullet still described a two-way choice with 0.0.0.0 as "the
default", which predates the Tailscale option. The prompt has offered three
choices for a while (Tailscale / any device on your network / this machine
only), and the default is computed from what is already on the machine rather
than being fixed at 0.0.0.0.
Now states all three options, that the Tailscale one is a loopback bind
fronted by `tailscale serve` with the tailnet as the login, and how the
highlighted default is chosen. Line 220 already documented the Tailscale
option correctly; this was the only stale spot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The rebase hand-repair dropped the closing brace of .welcome-btn-pi:hover
and .btn-toolbar.btn-run.mode-pi:hover before the inserted OMP rules.
The browser CSS parser drops every rule after an unclosed block, so the
deployed UI rendered as unstyled text bars (only ~456 of ~2583 rules
applied). Verified clean via esbuild --minify (no css-syntax-error) and
rebuilt dist.
The network-access prompt, where Tailscale serve is configured, runs AFTER
the build step. A build failure therefore exits before the question is ever
asked, and a user who then finishes the build by hand (rather than re-running
install.sh) ends up with a healthy loopback-only Codeman, a connected
Tailscale, and no serve mapping — with nothing anywhere pointing at
`install.sh tailscale`, the command that fixes it. Reported from a fresh
Ubuntu 24 install that died on the node-pty compile.
- maybe_offer_tailscale_repair(): on the update/re-run path, detect exactly
that state (loopback bind + tailscale Running + no serve mapping fronting
Codeman) and offer the retrofit. Silent for a deliberate non-loopback bind,
silent once a mapping exists, silent when tailscale is absent, and prints
the command instead of prompting when non-interactive. Returns 0 even when
setup fails so it can never abort an update.
- print_security_notice(): the loopback branch now names
`install.sh tailscale` when Tailscale is installed on the box, rather than
the generic "tailscale serve / cloudflared tunnel" advice.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
install.sh installs a build toolchain on Linux (node-pty has no Linux
prebuild, so a stock Ubuntu 24 server died inside node-gyp with
"not found: make"), plus review hardening for #339: the write-queue
reset paths now release the one-chunk-in-flight gate.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fifteen review findings on the dsh mode, the serious ones first:
- Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys.
_configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into
every dsh pane and applyEnvOverrides() lands after it, so a non-granted
owner who could redirect the base URL would have the operator's key sent
as a bearer credential to a host of their choosing.
- Wait registry: until=stop/blocked is refused on docker and remote-SSH
dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions).
The HERDR triple is set via LOCAL tmux setenv, which crosses neither
docker exec nor ssh, so such a session can never post a hook event and
the wait burned its whole timeout on every turn.
- Approvals: a dsh item is an ALERT, not an answerable card. The answer
route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the
option parser cannot read a third-party TUI's frames, so an answer was a
blind keystroke into a foreign composer), and the push notification
carries no Approve/Deny actions for dsh sessions.
- Status shim (v3): --seq is forwarded and the server drops stale retried
reports inside a 60s window (the TUI retries with backoff, so a retried
'working' could land after 'blocked' and resolve an approval whose
dialog was still on screen); 4xx responses exit 0 instead of retrying,
so one misconfigured session cannot feed the auth rate-limit bucket
until the hook endpoint 429s for the whole instance.
- Web-UI server: concurrent starts are serialized through a lock (two
racing POSTs used to pick the same port and orphan the winner), and the
readiness poll / timeout paths only clear or stop the singleton while it
is still theirs. First click actually opens the tab now
(refreshWebviews, not the nonexistent loadWebviews). DELETE
/api/deepseek/web requires the privileged grant in multi-user mode.
- Cron: deepseek jobs run the same two-part launch gate as the HTTP
create paths (impl moved into the resolver so all three share it) and no
longer stamp a Claude default model on the session.
- Parity sweeps: quick-start's docker branch rejects deepSeekConfig like
the remote branch; the Ralph auto-enable list gained deepseek;
HookEventType gained agent_working; the phone overview run menu filters
managed webview records like the desktop menu.
- install.sh: the dsh identity probe closes stdin (under curl|bash a
child that reads stdin eats the rest of the script), bounds the exec
with timeout where available, and is memoized to one scan per install.
- Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand
identity (it rendered as an unstyled UA-grey button); stale markup
comment about the web shortcut rewritten; clamp docs updated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Four review findings on the worker-transcript feature:
- Docker and remote-SSH dsh sessions now keep the pane segmenter: their
transcripts live in the container's / remote host's own ~/.dsh, which
the local reader can never see, so the transcript path returned
'nothing said yet' forever and an agent polling such a worker starved
on an answer that existed. Gated on !session.docker && !session.remote
(statically pinned) and documented in the integration guide.
- last-response reads are memoized on (path, mtime, size, blocks): the
skill's last_text polls once per second, and each poll decompressed and
reparsed the whole file on the event loop even when nothing had been
appended. An unchanged poll now costs one stat.
- The pairing ladder's comment claimed /new is served by step 2; in truth
the boot-window transcript wins for as long as it exists (deliberately:
preferring newest-eligible would hand a worker its busier sibling's
reply). The comment now states the real tradeoff instead of the
aspirational one. Same for decodeZstdFrames' 'skipped' wording — a
corrupt frame truncates the decode there, which is the safe behavior.
- stripReasoningPrefix no longer runs on user prompt text, so a prompt
containing a literal </think> renders whole in blocks view.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three review findings on the detailed-rows feature, all in its edge cases:
- The App Settings width select consulted the handheld defaults blob
(tabRailWidth: 256) BEFORE the rich-aware default, which the renderer
never reads — so a tablet's unsized rich rail rendered 320 while the
dialog said 256, and a routine Save persisted the 256 (below the 288px
tight threshold, permanently). The chain now mirrors
applyTabRailWidth()'s actual resolution.
- _setTabRailWidth() re-rendered on a compact flip but never re-ran
applyTabWrapSettings(), the one owner of the folder line, whose railRich
input reads the compact class this function just toggled. A rich rail
dragged below 240px kept emitting folder rows — persistently, for a
stored width < 240, since the boot wrap pass runs before the class is
first applied. The wrap pass now re-runs on the flip, with exactly one
render either way.
- Both reset affordances (handle dblclick, Enter on the handle) reset to
the hardcoded 256 even on a rich rail, landing it below the tight
threshold; both now resolve the rich-aware default (320), via a new
optional defaultWidth input on resolveTabRailKeyboardWidth().
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The agent skill could spawn a worker in any mode, but it could only
DRIVE a claude one: every other CLI has neither a real end-of-turn
signal nor an answer to read, so the recipes route them through output
markers.
dsh has both halves now -- its harness reports idle/working/blocked to
Codeman, and the previous commit reads its transcript -- so it joins
claude as a mode the four verbs work on unchanged. `spawn_workers alpha
beta:deepseek` is a mixed fleet in one call, and `sendwait` / `last_text`
/ `delete_session` need no per-mode variant.
Preamble 1.20.0 (SKILL.md's §0 heredoc regenerated from it):
- `spawn_worker` grows a deepseek branch that gates on the harness
composer. ⚠️ Readiness there is NOT the stop signal: the harness
reports idle at BOOT ~300 ms before its composer paints (measured
2.26 s vs 2.56 s after spawn), so a send-and-wait fired straight after
quick-start resolves on the boot edge, reports a turn that never ran,
and strands the prompt in a pane not yet taking input. Waiting for the
composer also spends that edge, since signals are edge-triggered.
- `spawn_workers` takes `name[:mode]`, so a mixed fleet stays one
concurrent call. Case names still have to be unique -- the mode never
disambiguates two workers that would share a directory.
- `sendwait` asks for `wait:"stop,exit"` instead of the `wait:true`
default set. That set also carries `idle`, which for an external CLI is
inferred from output stabilization: on a dsh worker whose TUI repaints
rarely, the re-wait resolved in 0 ms with `signal:"idle"` on a turn
with three minutes left to run. It also makes a wrong mode loud -- the
modes that cannot deliver `stop` answer 400 before writing anything,
instead of resolving on a flap.
- The self-heal resend carries `delivered:true` forward. The resend is a
tagged duplicate, so the server truthfully reports `delivered:false`
about a write it skipped, and §1's cleanup then read a completed turn
as an undelivered one and kept a finished worker forever.
- dsh workers spawn with the permission posture the Run button sends,
because the harness default still asks and a worker parked on an
approval row cannot finish a fan-out. The multi-user clamp still
applies.
Docs: a worked dsh flow in recipes.md, readiness and the signal rules in
verbs.md, and the corrections this makes necessary -- `stop`/`blocked`
are no longer claude-only, and `last-response` is no longer permanently
empty for deepseek. The integration guide gains a section on reading a
session back and driving one as a worker; its web-UI section was also
stale (that server moved out of a shell session).
The static guard that keeps those lists from naming some external CLIs but
not others is extended rather than exempted: it now knows the three real
classes inside that family (no transcript, no hook signals, and the
positive twin -- the modes whose answers can be read), with the hook class
derived from `hooksAvailableForMode()` so the predicate and the prose
cannot drift apart. Any other partial list still fails, and a new backend
belongs to none of the classes until someone says so.
`GET /api/sessions/:id/last-response` is how an agent (and the Response
Viewer) reads what a worker said. DeepSeek was falling through to the
pane segmenter with the other external CLIs, which for this mode is not
merely coarse but wrong: dsh-TUI paints a full-screen splash, so a
`last-response` call on a fresh dsh session answered with its ASCII-art
logo -- and anything polling for a worker's first reply reads that as a
reply.
dsh does not belong in that group. It writes a structured JSONL
transcript per session, so read it. Four things in that file shaped the
reader, all measured against real transcripts on disk:
1. dsh appends ONE ZSTD FRAME PER WRITE, and Node's zlib zstd decoder
(one-shot and streaming alike) stops at the first frame end: a real
56-line transcript decoded as 1 line / 158 bytes -- the session header
alone, i.e. a silent truncation that reads as "nothing said yet"
forever. `zstdFrameRanges()` walks frame and block headers to find
exact boundaries; splitting on the 4-byte magic would corrupt
everything after a magic sequence occurring inside compressed data.
zstd is resolved at RUNTIME because it landed in Node 22.15 while the
project floor is 22.0, so an older Node keeps the pane behaviour.
2. Every turn also records a plugin-sourced `user/message` (the runtime
context snapshot), which must not render as the user's own words.
3. A turn that ends in an error carries the provider's message; it is
surfaced as `Turn error: …` (and a non-error early stop as
`Turn ended: …`) rather than as an empty string, which an agent reads
as "still thinking" through fifteen polls.
4. Reply text is assembled per (turn, step): a finalized message wins and
the streamed deltas fill in only for a step that never finalized, so a
partial answer is readable mid-turn and never doubled. "Finalized" is
tracked as a set of steps rather than as non-empty text, because a
step whose whole reply was reasoning strips to '' at the `</think>`
boundary and would otherwise resurrect the raw deltas in its place.
Session-to-transcript pairing is by the transcript's own header `cwd`
plus a boot window against the session's createdAt, never by
reproducing dsh's directory mangling (already two forms on disk) and
never by newest-mtime alone -- mtime alone handed a freshly spawned
worker its predecessor's answer in the same case directory.
An empty result still wins over the pane; only a Node that cannot decode
zstd falls back to it.
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a
shell tab running the server next to it. The shell was deliberate - the server
lived in an ordinary session so it was visible, scrollable, killable and died
with its tab, and nothing new had to supervise a long-lived HTTP server. That
reasoning was sound and the result was still wrong in use: opening a dashboard
should open one tab, and after the first launch the terminal is pure noise.
The server moves to a background child process owned by a new
`src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session
gave away for free is now explicit, which is most of the module:
- Exactly one server. A second click reuses the running one instead of racing
it for a port; the session flow could not do this at all, because two clicks
were simply two sessions.
- Restarted when the requested authority changes. `--trusted-host` fences dsh's
own /api against the browser authority, and a Codeman reachable at both
loopback and a tailnet name has two. Reusing a server fenced for the other
origin renders a page whose every call 403s, which reads as a broken
dashboard rather than a misconfigured one, so a mismatch restarts instead.
- Killed on shutdown. The child is detached so its whole plugin tree can be
signalled at once, which also means it would outlive Codeman and hold its
port against the next start - the exact EADDRINUSE this feature already got
wrong once.
- Boot output captured and returned. With no shell tab there is nowhere else
for a stack trace to land, so a failed spawn reports its own tail.
The endpoint is fenced at the same bar as the profile installer and for the
same reason: booting a dsh profile executes the plugin code in it, so this is a
privileged action even though it reads as "open a page". `authority` comes from
the client (`location.host`) because only the browser knows which origin is in
play, and it is regex-confined at the schema boundary - defence in depth behind
the argv-array spawn, admitting host:port in the shapes a browser authority can
take and nothing readable as a second argument.
`GET /api/deepseek/web-port` is gone; port selection moved into the supervisor,
which is the thing that knows whether a server is already running. The two
client-side probe helpers went with it, since the server now owns the wait.
Verified over the tailnet authority end to end: no session is created (session
count unchanged, one tab), the server runs on 3081 beside the user's own dsh
web on 3080, status reports the tailnet authority, and the proxied dashboard
renders with zero 4xx. Full gate green (6148 passed, +6).
The `Run > DeepSeek web UI...` shortcut failed three ways at once against a real
install, and the three are independent.
1. It hardcoded `--port 3080`. That is dsh web's OWN default, which makes it
precisely the port a DeepSeek user is most likely to be serving on already,
so the launch died with EADDRINUSE against the user's own server. The port
now comes from `GET /api/deepseek/web-port`, which walks 3080..3119 for a
free loopback port by BINDING it (a connect probe cannot tell "free" from
"listening but not answering yet").
2. It opened the tab unconditionally. The crashed server left a saved dashboard
pointing at nothing, with the failure only visible in a shell tab nobody had
a reason to look at. The launch now polls the existing webview probe until
the URL answers, and on timeout reports the error naming the shell tab
instead of persisting a dead dashboard.
3. The saved tab was untrusted, so the frame was sandboxed without
`allow-same-origin` and the dashboard was broken twice over: the dsh
client-runtime reads `localStorage` while loading its plugins and died there
("the document is sandboxed and lacks the 'allow-same-origin' flag"), and an
opaque-origin frame sends `Origin: null`, so dsh's own trust fence 403'd
every `/api` call no matter which authority `--trusted-host` named. Passing
`location.host` only means anything once the frame actually carries that
origin, so `--trusted-host` had never once done its job. The managed tab is
now created `trusted: true`.
That trade is real and deliberate: a trusted proxied frame is same-origin
with Codeman and can reach Codeman's API. It is defensible only because this
dashboard is an agent harness Codeman just started itself, on loopback, which
can already run code as the user. It is not a precedent for trusting
third-party dashboards, which is why it is set at this one call site rather
than defaulted.
Separately, the shortcut listed its own dashboard twice: once as the menu entry
that starts it and once as the row that entry had written on the previous click.
Webviews now carry an optional `managed` marker, managed rows are filtered out
of the saved-dashboard list, and a relaunch repoints the existing row rather
than stacking one dead dashboard per restart (which the per-launch port would
otherwise guarantee). `managed` is declared in the schema because a plain
`z.object` strips undeclared keys, so an undeclared marker would never survive
the round trip.
`DEEPSEEK_WEB_PORT` is gone from constants.js; its doc comment asserted that a
hand-started `dsh web` and the shortcut "land on the same place and share one
saved tab", which is the bug stated as a feature.
Verified on a real install with the user's own `dsh web` holding 3080: the
shortcut takes 3081, the server answers, exactly one DeepSeek entry shows in the
run menu, and the proxied dashboard renders its workspaces and completes its own
API calls (the previously-403'd `api/settings.describe` now succeeds). Full gate
green (6142 passed), typecheck/lint/format/public-assets clean.
Two review nits on the vertical rail's detailed rows.
1. The tab-rail-tight rule (below 288px) hides `.tab-meta-created`, and its
comment claimed the value "survives in the row's title attribute either way".
It did not: the only title carrying it lived ON that element, and a
`display: none` element has no hover target, so the created stamp was not
shrunk but gone with no way to ask for it. Rather than just correcting the
comment, `_sidebarRichMetaHTML()` now puts BOTH absolute stamps on the
`.tab-meta` line itself, so the pill and the gaps around the stamps remain as
hover targets. An item's own title still wins where the item is visible.
2. applyTabOrientation() decided whether applyTabWrapSettings() had already
re-rendered by comparing `_tallTabsEnabled` before and after. That reads an
UNDEFINED previous value as "it rendered", but applyTabWrapSettings()
deliberately renders nothing on its first call ever (it only establishes the
baseline: `prevTallTabs !== undefined && prevTallTabs !== showFolder`). So on
a first call that also flips the folder row, neither function rendered and the
rows stayed stale. Reachable when the pre-paint script throws and leaves the
layout attributes on their catch-branch fallbacks for applyTabOrientation() to
correct. The guard now mirrors applyTabWrapSettings()'s own condition.
Both new tests were run against the unfixed code first and fail there, which is
the only thing that makes them regression tests. (The third, "does not render
twice", passes either way by design: it pins that fix 2 did not introduce a
double rebuild.)
Verified in a real browser against a live server with two sessions, driving the
narrowing through _setTabRailWidth() the way the resize drag does: at the 320
default the row reads "CREATED 2m ago · IDLE <1m" with the created element
displayed; at 256 the tight class is on, the created element computes to
display:none, the visible text drops to "IDLE <1m", and the meta line's title
still reads "First created: ...". At 220 the compact threshold drops rich rows
entirely. Screenshots confirm no truncation artifacts in either state.
Full gate green (6104 passed), typecheck, lint, format, frontend-syntax and
public-assets all clean.
The three smaller review nits, plus the first real test coverage for the status
shim (it had none: it is emitted as a STRING, so tsc never sees it).
1. The shim was written with a plain writeFileSync. The TUI can be exec'ing that
exact path while an upgraded Codeman refreshes it, and a reader catching a
half-written file gets a syntax error, exits non-zero, and is retried four
times per state change for a file that will never parse. Now temp + rename
(atomic within the directory), with the temp chmod'ed before the rename since
writeFileSync's mode only applies on create, and removed if the write throws.
SHIM_VERSION bumped to 2, because SHIM_SOURCE changed and an existing v1 shim
would otherwise keep matching the embedded marker and never be refreshed.
2. The pane-id comment claimed the ambient env "cannot be spoofed by an argument
the agent itself could influence". The agent runs IN that pane and can invoke
the shim with CODEMAN_SESSION_ID unset and any argv it likes. It buys nothing
it did not already have (the hook-secret file is readable from the same pane,
so it can POST /api/hook-event directly), but the comment read like a security
boundary. Rewritten to say what the preference actually buys: correct
attribution when a TUI mangles or re-uses the pane argument. Accidents, not
adversaries.
3. classifyProfile() folded the directory name into the same haystack as the
bundles, but only the TUI arm could match a bare name, so a stock profile
whose package.json has no dsh.profile.bundles (hand-edited, older layout,
mid-install) classified as `unknown` -> launchable -> eligible as the DEFAULT
pick, which is exactly the pane-dies-on-arrival failure the two-part
availability gate exists to prevent. The stock names are now a LAST-resort
fallback consulted after the bundle patterns, so real bundle evidence still
wins over a name the user chose. The loose `tui` arm gained word boundaries:
it decides which profile boots by default, and matching the middle of
`intuition` is not a rule anyone could predict.
New test/deepseek-status-shim.test.ts runs the generated script the way the
harness does -- real node process, real argv, real env, real listener -- and
covers the exit-code contract that makes the retry behaviour safe: mapped states
post and exit 0, an unknown verb or unmapped state exits 0 WITHOUT posting (a
non-zero there would be four HTTP requests per state change forever), a rejecting
server or an unreachable one exits non-zero so the caller retries, the hook secret
is read at execution time, and `node --check` parses the file (a template-literal
typo in SHIM_SOURCE is invisible to tsc).
Trap worth recording, hit while writing it: the tests must spawn the shim
ASYNCHRONOUSLY. The listener lives in the test process, so spawnSync blocks the
event loop that has to accept the connection, the shim waits out its own 1500ms
socket timeout and exits 1, and it reads exactly like a broken shim (measured:
Socket._onTimeout in its --trace-exit output, server logging nothing).
Verified: full gate green (6142 passed, +10), typecheck/lint/format clean.
Three review findings on the DeepSeek Harness mode, plus one the third exposed.
1. The multi-user clamp was bypassable by a sibling field on the same request.
clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but
DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER
_configureDeepSeek(), so a non-granted owner sending
envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated
instance: a session created with permissionMode "read-only" and that override
ran with DSH_PERMISSION_MODE=danger-full-access in its pane.
Every other CLI's bypass is a command-line flag reachable only through the
per-CLI config, which is why the config clamp alone is the whole gate for
them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted
owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to
what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on
that list because it aims the launcher at a profile tree whose plugin code
runs at boot, before any approval row can apply. Verified end to end in real
multi-user mode: a non-granted user sending both now gets workspace-write and
no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched.
2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout`
signals only the direct child, and a plugin install fans out into
package-manager children that keep the inherited stdio pipes open, so `close`
never fires and the held-open request leaks with no route-level deadline.
Reproduced: with a 1.5s built-in timeout the promise was still unsettled after
6s and both fan-out children were alive. Now detached: true plus negative-pid
SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a
last-resort reap for a grandchild that escaped the group. Same probe after the
change: close fires, direct child and both grandchildren dead.
3. hooksAvailableForMode() promised more than a dsh session can deliver.
deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that
triple is the only reason a dsh session posts hook events, so `until=stop` was
accepted and then blocked for the caller's whole timeout: the exact
infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now
takes HookCapabilityOptions and every call site passes sessionHookOptions(),
with the deepseek arm reading `!== false` so a forgotten one degrades to the
old behaviour. The refusal names the setting rather than saying "no Claude
Code hooks", which would send the caller hunting a bug that is really a
setting they chose. Profile conformance stays unknowable at request time and
is documented as such. The stale "True for `claude` and nothing else" docblock
is corrected.
4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a
claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent
capture read Claude's own transcript, and adding deepseek silently widened
both to a mode that has none. They compare mode === 'claude' directly now, and
a static check pins them there.
Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the
wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 --
bridge off plus explicit until=stop is a 400 naming the setting, bridge off with
no `until` still 200s on idle/exit, bridge on accepts stop.
The vertical tab rail (tabOrientation 'vertical') listed names and nothing
else, while the rich sidebar and both home screens already answered the
question a docked column exists to answer: which of these sessions wants me
next, and how long has it been like that. The rail is a docked column too, so
it now draws the same row.
- New per-device setting tabRailDetail ('rich' | 'simple', default rich),
App Settings -> Appearance -> Tabs, in SettingsUpdateSchema + displayKeys and
stamped as data-tab-rail-detail by the pre-paint script, so a detailed rail
does not flash through simple rows on every load.
- ONE gate for both vertical surfaces: isRichTabRows() =
isSessionSidebarRich() || isTabRailRich(). The row model, the markup and the
20s in-place clock are the existing rich-sidebar ones, classified by
_mobileOverviewState/_mobileOverviewSince, so the rail, the sidebar, the
desktop home rail and the phone overview cannot disagree about what
"working" means or which stamp measures it.
- Detail rides on its OWN attribute, exactly as the sidebar's does, so every
existing [data-tab-orientation='vertical'] rule keeps matching both variants
untouched. A flip of detail ALONE still forces a full render (the stamps line
is emitted by the row template, not toggled by CSS) and re-runs
applyTabWrapSettings(), which owns the folder line and is now rail-aware.
- CSS: every rich paint rule gains a rail twin as a COMMA-GROUPED selector,
never :is() - an :is() list takes its most specific argument, which would
lift the sidebar arm from (0,3,1) to the rail's (0,5,1) and let these rules
outrank things they never used to.
- Width is why there are thresholds. At 256px the stamps line ellipsizes
mid-word, the same reason the rich sidebar is 300px, so a rail that has never
been sized defaults to 320 (RICH_DEFAULT_WIDTH, the existing Wide preset,
which also keeps the settings select on a named choice). A width the user has
chosen is never overridden: below 288px the created stamp is dropped rather
than truncated (tab-rail-tight, CSS only) and below 240px the rows go back to
simple (tab-rail-compact, which re-renders).
- The rich clock is armed and disarmed by applyTabOrientation() as well as
applySessionListLayout(); a leaked interval would rewrite stamps in a list
that no longer has any.
Also fixes a data-loss bug in the inline tab rename that predates the rail and
reproduces in every layout, header strip included: Escape set the input to ''
and blurred it, and the blur handler commits - so cancelling a rename PUT an
empty name, and the tab fell back to its folder label (measured against a live
server: ["rail-alpha","","rail-gamma"]). Escape now calls cancelRename(), which
invalidates the edit so the blur that follows the input's removal is a no-op.
Tests: rail-detail gate, the three ways it turns back off (simple, compact,
horizontal), sidebar-wins, render-on-detail-flip and the plumbing/CSS guards in
test/session-list-layout.test.ts; the rename cancel in test/inline-rename.test.ts
(browser suite), pinned by running it against the old code first. Verified live
against a real server on an isolated instance: detailed/simple/compact/header/
sidebar variants, click-select, the ... menu, inline rename, Alt+N, the in-place
stamp tick and a full settings-picker round-trip including reload.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.
DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:
1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
$DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
`base` -- the interactive terminal front door is always a third-party
plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
button gates on the latter, because reporting only the binary would spawn a
pane that dies on arrival. When the binary is present but no profile is,
the run menu offers to install one (POST /api/deepseek/install-profile).
2. The permission switch is an env var, not a flag. The harness has no
command-line permission option; its sandbox/approval rows read
DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
own workspace-write, which still asks, so the multi-user clamp is the
only-if-sent branch and clamps to workspace-write, never read-only.
3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
earned that. The terminal front door reports idle/working/blocked to a
supervising process over a generic env-gated contract; a generated shim
(deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
report to /api/hook-event as stop / agent_working / permission_prompt. So a
dsh session gets definitive respawn triggers, real wait-endpoint signals and
real Approvals Inbox items instead of output-stabilization guesswork.
`agent_working` is new (157th SSE constant) and joins
APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
alert at once.
The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.
Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.
Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.
Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The agent-image grok step copied /root/.grok/bin/grok onto /usr/local/bin/grok
with cp -L. Newer versions of xAI's install.sh already create
/usr/local/bin/grok as a symlink to that same binary, so the copy failed with
'same file' and the --no-cache rebuild died at the grok layer (2026-08-24).
Stage the copy under a temp name, drop whatever the installer left at the
destination, then move into place - correct against both old and new
installers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Post-merge follow-ups from the deep review of #334 and #335, so they ship in
the same release as the features.
Tab-layout foundation (#335):
- PUT /api/session-order drops unknown/foreign ids again instead of 400ing
the whole write, in both the owner and the admin path (single-user requests
are the synthetic admin, so that path is the one the browser hits). The
frontend debounces its reorder push and swallows errors, so a session
deleted inside the debounce window silently cost the user the entire
reorder - and the endpoint sits on the stable /api/v1 surface, where the
pre-layout server merged leniently.
- A failed mux restore no longer locks explicit deletions into 500s for the
process lifetime: runSessionDeletion and webviewDeleted degrade to
best-effort without layout coordination, while the automated stale sweep
(runStaleSessionCleanup) stays fail-closed.
- sse-events doc comment: no 'suppressed' hook event exists; hooks stay 8.
- registerSessionWithLayout resolves its owner through ownerLayoutKey()
instead of a hardcoded '@single'.
Vertical rail (#334) - all rail-awareness gaps in sidebar-only predicates,
unified behind the new _isVerticalTabList() (sidebar OR rail):
- Drag-reorder read the insertion side from clientX in the rail, so
before/after was effectively arbitrary on vertical rows; the drag-over
indicators now draw as top/bottom edges there like the sidebar's.
- The active tab is scrolled into view in the rail (Alt+N/palette selection
used to leave the row below the fold).
- Floating subagent/ultracode windows anchor to the RIGHT of rail tabs, and
the connector redraw gates (render tail + strip scroll) cover the rail.
- Server-seeded tabOrientation is applied when the async settings load
resolves, not only at boot, so a fresh device shows the rail immediately.
- The pre-paint script stamps data-tab-orientation and --tab-rail-width
(sidebar-wins and solo carve-outs included), removing the flash of the
header strip on every vertical-mode load.
- The session name font defaults to 12px, the sidebar's historical 0.75rem
size, so installs that never touch the new slider are not restyled.
Also documents the rail in CLAUDE.md (second #sessionTabs host, mover
ordering, the axis-predicate rule) and gives tab-rail-resize.js its
@dependency/@loadorder header. Full gate green (6093 tests); the excluded
browser suite was run by hand - only the known environmental failures
(opencode/codex binaries) remain, identical to pristine master.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The button next to the file preview's close icon was Copy Content, whose
overlapping-pages glyph reads as a pop-out control - and for a PDF or any
media/binary preview it was completely dead: those branches never fill
filePreviewContent, so the click hit an empty-content guard and did nothing,
with no feedback.
There is now a real detach button that opens the previewed file in a browser
tab (raw route for PDFs/images/media/text, the server-converted PDF preview
for docx/pptx), severs window.opener by hand so a blocked pop-up stays
detectable, closes the overlay on success (which also stops any playing
media), and disarms on close so it can never open a stale file. The copy
button now toasts 'Nothing to copy in this preview' instead of staying
silent.
Verified live with Playwright against an isolated instance: button visible
and armed on a PDF preview, file-raw answers 200, clicking opens the URL and
tears the overlay down, text previews keep a working copy buffer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The agent skill's endpoints.md is what other agents read as ground truth, and
three of its facts went stale when grok landed:
- `/api/v1/grok/status` was added to the probe list, but the sentence after it
still said only Pi's response carries `.data.version`. Grok's carries it for
the same reason (a squatted binary name), and an agent that trusts the old
wording has no way to tell a misresolved grok from an absent one.
- the `active-tools` bullet listed grok among the modes it stays empty for, then
claimed in the same breath that `isExternalCliMode` "lists only those five".
- its three source line refs had all drifted: `isExternalCliMode` is now
session.ts:174-183 (it was already wrong before this branch), the external-CLI
early return is session.ts:2261, and TEXT_COMMAND_PATTERN is
bash-tool-parser.ts:89.
CLAUDE.md and architecture-invariants.md counted modes in their Docker-cases and
Web-tabs paragraphs ("any of the five CLI backends", "never a sixth
SessionMode"). Both numbers were already stale before grok (antigravity and pi
had made it seven) and grok is now in the agent image, so the counts are gone
rather than incremented: the invariant those sentences carry is that Docker and
web tabs are not modes at all, which no number has ever helped state. The two
plan docs keep their original wording, being historical design records.
test/workflow-run-watcher.test.ts pinned its fixture's newest activity at
2026-06-14T20:06:40Z and then asked getRecentRunSummaries(100000) to return
it. That argument is MINUTES, so the window is 69.4 days: the assertion
expired at 2026-08-23T06:46:40Z and the file has failed on every branch
since, on a suite nobody had touched. The last green CI run finished at
06:47:43Z, about a minute inside the boundary, which is why it landed as a
surprise rather than a bisectable regression.
The fixture epochs now hang off a RUN_ANCHOR of Date.now() - 601s with every
offset preserved verbatim, so the parsed durations, the ordering and the
live-vs-done discriminators are all unchanged, and the recency filter is
still the thing under test. It just cannot rot again.
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
scripts/tmux-chooser.sh is deleted. codeman tui replaces it and does the
job better: sc numbered its entries globally but only accepted a single
[1-9] keypress, so sessions 10+ were listed and unselectable, and it
inferred nothing about what an agent was doing. The tui carries the
server's real states, answers permission dialogs, and leaves an attach
with one key.
install.sh no longer creates the tmux-chooser symlink or the sc alias.
It now sweeps both up instead, on update AND uninstall, so an update
cannot leave a symlink pointing at a script this version stopped
shipping. The alias removal is marker-owned: it matches the exact line
the installer wrote, so someone's own 'alias sc=' for another tool is
never touched, and it rewrites through 'cat >' so the profile keeps its
mode and ownership. Verified against three profile shapes.
BREAKING CHANGE: the 'sc' command and the 'tmux-chooser' symlink are
gone. Use 'codeman tui' (and 'codeman tui --list' / 'codeman tui <n>').
codeman tui supersedes the sc bash chooser: it reaches sessions 10+,
carries the server's real states instead of a static list, and leaves an
attach with one key. Every place that told a user to run sc now names the
tui equivalent, including the two wiki pages and install.sh's next-steps
banner. Both wiki pages also carried the wrong detach chord (Ctrl+A D;
the socket's prefix is C-b), which the tui makes moot.
Source comments that explained themselves as "the sc -l replacement" now
just say what they do. docs/tui-plan.md and CHANGELOG.md are historical
records and keep their references.
CLAUDE.md gained "any new `tmux -L` caller through `resolveTmuxSocketName()`"
but architecture-invariants.md, which that bullet points at for the
mechanism, still described only the dataPath() half. Say why the rule
exists there too: the TUI is the first non-server process to shell out
to tmux.
The sc chooser runs a plain `tmux attach-session` and binds nothing, so
F1 does not detach from it; only `codeman tui`'s attach claims that key,
and only for its own duration. The line it replaced was wrong too (the
socket's prefix is C-b, not C-a), so name the real chord and say which
command gives you the single key instead.
The guide and the README both said to detach with `Ctrl+B D`. Beta testing
proved that wrong twice over: tmux binds lowercase `d` to `detach-client` and
capital `D` to `choose-client`, and even the correct letter fails for anyone who
keeps Ctrl held, because that sends `Ctrl+D`, which tmux leaves unbound. A
tester followed the documented instruction, stayed attached, and exited the
agent to escape.
Both now say `F1`, and the attach section describes what actually happens: the
session strip across the top of the pane, `Alt+1`..`Alt+9` switching without
returning to the dashboard, and `r` to resume a session whose pane has died.
Also corrected: `1-9` switches rather than jump-attaches, `x` confirms with `y`
rather than a typed name, and a new session opens straight into its pane.
`docs/tui-plan.md` is deliberately untouched — it is the design record of what
was planned, not a description of what shipped.
The loose end from 777f974, now explained. Sessions came back from a detach on
`window-size latest` instead of `manual`, and the restore primitive round-tripped
correctly in isolation, so the corruption had to be upstream of it. It was: the
snapshot was taken from state this code had already broken.
`bindSwitchKey` passed a bare `;` between the two commands it wanted in one
binding. That is a command separator to tmux's OWN parser, not an argument: it
ended the `bind-key` and executed what followed immediately. So the binding kept
only `switch-client`, and `set-window-option ... window-size latest` RAN against
every switchable session at attach time — before the sizing snapshot was taken.
Every session was therefore snapshotted as `latest` and faithfully restored to
`latest`.
Proven against real tmux both ways before fixing: a bare `;` leaves the session
on `latest` and stores a one-command binding, while `\;` leaves it `manual` and
stores both commands.
Verified end to end: 7 sessions manual before, 1 latest + 6 manual during the
attach (the attached one follows the terminal, the rest are pre-sized), no dot
padding on a switch, and all 7 back to 120x40 manual after the detach.
This also means the "follow the terminal after switching" half of 777f974 never
actually worked — it was never in the binding.
The overview showed the same session twice, one frame above another, after
switching sessions (reported from the beta with a screenshot).
Claude repaints by ABSOLUTE CURSOR POSITIONING, not by clearing: a 198KB pane
tail carries 1142 `CSI r;c H` and exactly one `CSI 2J`. The replay honoured the
COLUMN of those sequences and ignored the ROW, so a repaint could never
overwrite what came before and was appended instead. That same tail replayed as
FIFTY stacked copies of one frame. The preview shows the last N lines, so on a
short terminal you saw the newest frame by luck and on a tall one you saw the
end of the previous frame above it.
A cursor HOME now starts the buffer over. That is not a heuristic but the
line-based equivalent of what a home means: a full-screen app announcing it is
repainting from the top, with everything on screen about to be overwritten in
place. Only row 1 column 1 counts — any other address is a write position
inside the frame being painted, and resetting on those would erase live
content.
Measured on the real tail that produced the screenshot: 198599 bytes and 50
copies of the welcome frame collapse to 40 lines carrying exactly one.
The old test pinned the append behaviour, including a spurious leading empty
line that the initial CUP produced; both are gone.
The bar now reads "alt+1-9 switch · F1 back to the codeman dashboard", so the
switch keys are discoverable instead of secret. Shown only when those keys were
actually claimed, the same rule the way-out key follows: a bar naming a key
that does nothing is the bug this series started with.
THE DOT GRID. Switching landed in a pane occupying part of the terminal with
tmux's dot fill everywhere else. It was never a size mismatch — the window was
already the right size. `window-size latest` only resizes a window while a
client is ON it, and the sessions behind the tab strip have none until you
switch, so the resize happened AT the switch: tmux painted the newly-available
area with dots and an idle claude had no reason to redraw into it. Every
switchable session is now pre-sized to the attaching terminal, which moves that
repaint to attach time while the user is still looking at the first session,
and the switch binding restores `window-size latest` on arrival so a mid-attach
terminal resize still follows. Measured: 14 consecutive switches across 7
sessions, zero dot-padded rows, against 1-in-6 before.
⚠️ Known loose end, deliberately not papered over: after a detach the window
SIZE is restored exactly but the window-size MODE can come back as `latest`
rather than `manual`. The restore primitive round-trips correctly in isolation
(manual -> presize -> latest -> restore = manual) and no call site in the TUI
or the server sets `latest` afterwards, so the cause is not yet identified. The
practical effect is nil: the remaining client keeps the window at its own size
and Codeman re-pins `manual` on the browser's next resize.
Switching with Alt+N landed in a pane that filled part of the terminal with
tmux padding the rest as a dot grid — reported from the beta with a screenshot
showing the pane in the left half and dots everywhere else.
Codeman pins every window `window-size manual` at the BROWSER's size
(tmux-manager.ts), so no attaching client can resize it. The attach already
lifted that for the session it opened, which is why a plain attach looked
right; `switch-client` then moved the user into a session that had never been
lifted, and the old pin reasserted itself. `window-size latest` now goes on
every session the strip can reach, alongside the bar those sessions already
get, and each one's original sizing is snapshotted and restored on detach.
Verified by round-tripping a session pinned at 120x40 manual: latest 190x49
while attached, back to 120x40 manual after, with no dot rows at either step
and the bar intact at full width after a switch.
The renderer's own FOOTER_KEYS table still said 'jump'. It is only reached when
the app layer supplies no footerKeys, so nothing visible was wrong, but a
fallback that contradicts the live footer is exactly the kind of drift that
turns into a bug report later.
Four faults, all reported at once, and three of them were mine from the last
two commits.
THE HINT VANISHED. Two independent causes. First, a leaked F1 binding: an
attach whose TUI was killed leaves `F1 -> detach-client` in tmux's root table,
and the claim treated "already bound" as someone else's key, so every later
attach fell back to advertising the tmux chord — the bar stopped saying F1
while F1 still worked. A key already bound to `detach-client` now counts as
ours. Second, width: tmux truncates a status line that overflows and drops the
RIGHT-aligned segment, which is the hint. The strip now gets a budget measured
from the terminal's width minus the hint, and it drops tabs from the far end
until it fits. ⚠️ Measured on VISIBLE columns, not format bytes: `#[reverse]`
costs zero columns, and counting it made a strip that "fitted" still truncate
the hint at 80, 100, 120 and 176 columns on a real terminal.
ALT+N DID NOT SWITCH. On the dashboard, a bare digit meant jump AND ATTACH, and
a terminal sends Alt+N as ESC then N: when those land in separate reads —
routine over SSH — the chord decodes as Escape plus a bare digit, so "switch to
tab 2" threw the user into tab 2's pane. A digit now SELECTS, matching what
Alt+N means in the web UI; Enter is how you go in. Inside a pane the keys never
reached the TUI at all, since tmux owns the terminal, so the attach now binds
Alt+1..9 in tmux's root table to `switch-client` — the strip is usable rather
than decorative. ⚠️ The bar is applied to every session the strip can reach,
each highlighting its own tab: with it on the attached session only, switching
landed the user in a pane with no strip and no way out on screen.
⚠️ The leaked-state sweep was missing `status-position`, so it removed the
marker and left the position behind — and with no marker the leftover no longer
matched, making it permanently unsweepable. Found by diffing every session's
options after a detach.
Attaching made every other session disappear: the dashboard is gone, tmux owns
the terminal, and there is nothing left saying what else is running. The attach
bar now carries the session strip, numbered exactly as the dashboard numbers
them, with the session you are in inverted, and it sits at the TOP of the pane
where the web UI keeps its tabs.
The strip is a WINDOW around the active tab, not the whole list, with ellipses
marking each end that is actually cut. The bar is one line shared with the way
out, and that hint is the only instruction a user gets while tmux has the
terminal, so it must never be crowded off; a test drives 20 long-named sessions
through the bar and asserts it survives.
⚠️ The strip is a snapshot taken at attach time and never refreshed. The TUI is
blocked in `spawnSync` for the whole attach so there is no loop to update from,
and tmux's own format language cannot map a `codeman-<hex>` session name back
to a label a human recognises. Slightly stale beats absent.
The way out moves from F12 to F1, which sits beside Esc where a hand backing
out already goes. Verified against BOTH encodings a terminal sends for it:
xterm's SS3 (ESC O P) and PuTTY's default (ESC [ 1 1 ~).
`status-position` joins the snapshot, so a session that had its bar at the
bottom gets it back there on detach along with everything else.
Three separate "why are there boxes" reports, and I fixed them one glyph at a
time instead of as a class, so the next one was always waiting. Grouping the
tester's terminal by unicode block made the rule obvious:
RENDERS Latin-1 (·), Box Drawing (─ │), Block Elements (█ ▛ ▐),
Geometric Shapes (○ ▶), General Punctuation (…), Arrows
TOFU Miscellaneous Technical (⏎ U+23CE, ⏵ U+23F5), the sparse end
of Dingbats (❯ U+276F)
That is an ordinary font, not a broken one, so it is the profile to design
against. The working spinner moves off Dingbats and Math Operators onto
quadrant blocks (▖▘▝▗) — the same block as the `▛█▐` art claude itself draws,
which that font renders fine — and the blocked marker moves off `⚠`
(Misc Symbols, emoji presentation on many terminals) onto `▲`, the block that
already gives us `▶` and `○`.
The preview fold gains claude's own spinner dingbats (✢ ✳ ∗ ✻ ✽ ✴ → `*`) and
`⚠` → `!`. Its animated status line is exactly where a reader looks, so tofu
there is the most visible kind there is.
A test now enforces this as a CLASS: no glyph in the unicode set may come from
Misc Technical, Misc Symbols or Dingbats, with U+2714 the single documented
exception because it was observed rendering on the very font that failed the
others. Verified by scanning a live frame driven with the tester's exact
environment: zero glyphs from any of the three blocks.
Killing demanded the session's NAME typed out in full. That is the right
ceremony for dropping a production database and the wrong one for closing a
pane you are looking at; the beta tester's verdict was "thats stupid, just make
me type Y to confirm". `x` then `y` is already two deliberate keystrokes on a
row the user selected, and the conversation lives in its transcript, which a
kill does not touch.
Everything that is not `y` CANCELS rather than being ignored, so a stray key
closes the dialog instead of leaving a destructive prompt armed and waiting for
whatever gets typed next. Enter cancels too: it is the key most likely to be
hit by reflex, and this is the one dialog that destroys something.
⚠️ Found while verifying the new dialog: it did not name the session. The label
was computed as `row.session.name ?? id.slice(0, 8)`, and `??` falls back only
on null or undefined, so every session the server left with an EMPTY name — all
of them, until the TUI started naming its own — sailed through and the box read
"Kill ?". A destructive prompt that cannot say what it will destroy is worse
than no prompt, and it is now a single keystroke. The caller passes the same
label the LIST shows, so the dialog names the row in front of the user.
The typed-name machinery goes with it: TuiConfirmState.typed, setConfirmInput(),
confirmAccepts() and the 'typing'/'reject' steps are all removed rather than
left as unreachable branches.
Two reports from the same beta screenshot.
Starting a session left the user on the dashboard next to the row they had just
asked for, which reads as the create having silently failed. Starting a session
is a request to WORK in it, so the terminal now goes there as soon as the pane
exists, and the CLI booting is worth watching. If the pane is slow the notice
says so and the row is left selected, exactly as the resume path does.
The footer's `↵` was drawing as an empty box: `⏎` (U+23CE) has poor font
coverage, on the same terminal that renders `·`, `─`, `│`, `○`, `▶` and `✔`
perfectly. It is now U+21B5, from the Arrows block every monospace font ships.
`✋` (U+270B) was worse than a coverage problem: it is East Asian WIDE, so the
renderer, which addresses cells by column, was reserving two cells for it. The
golden frames had the age column shifted a space left to match, which is how
long that had been wrong. It is now `!`, and the frames align correctly.
A test walks the whole unicode glyph set and fails on any entry wider than one
cell, so a glyph that shifts the layout cannot be added again. The comment on
the table spells out both bars a glyph has to clear, because the tier check
answers neither: it asks whether the LOCALE is UTF-8, which says nothing about
whether a font has the glyph or how wide it draws.
Alt+1..9 switches to that session, and `[` / `]` / Tab step through them, so the
muscle memory from the web UI carries over.
Alt+N SELECTS rather than attaches, which is what the web UI's Alt+N does:
switching which tab you look at is cheap and reversible, and the terminal
equivalent is moving the selection and its preview, not handing the whole
terminal to a pane. Bare 1-9 keeps its documented jump-and-attach meaning.
Two of the web UI's chords cannot cross into a terminal, so the nearest
transmittable keys carry them instead:
Alt+[ / Alt+] ESC+[ IS the CSI introducer every arrow key arrives on, and
ESC+] is OSC, so neither chord is distinguishable from a
sequence. Bare `[` and `]` do the job.
Ctrl+Tab a terminal cannot report the Ctrl, so plain Tab carries it.
⚠️ The parser now decodes ESC + a printable character in ONE read as an Alt
chord, and the app replays every chord it does not claim as `escape` then that
character. That fallback is load-bearing, not tidiness: a real Esc landing in
the same read as the next keystroke is byte-identical to a chord, and without
the replay "Esc then q" typed quickly decoded as Alt+Q, matched nothing and was
swallowed. The e2e suite caught exactly that as the dashboard refusing to quit.
A lone Esc is still held and flushed on the caller's timer, which is what keeps
the two separable at all.
A beta tester photographed claude's `❯` prompt and its `⏵⏵` bypass-permissions
marker rendering as empty boxes in the preview pane. Their font has no coverage
for those codepoints while drawing `·`, `─`, `│` and `▶` perfectly.
The glyph TIER cannot help here. It answers "can this terminal do Unicode at
all", which is a locale question, and it correctly says yes for exactly the
terminals this affects. Coverage is per-glyph and undetectable from inside the
process, so the handful of rare glyphs CLIs use as chrome are folded to the
ASCII arrows they already look like, and everything a plain font does render is
left alone.
Scoped tightly: the preview only, never the TUI's own chrome, and skipped
entirely at the `nerd` tier where the user has declared a font that can draw
anything. The table is short and every entry was seen as tofu in a real
terminal rather than guessed at. The fold is length-preserving, so the preview
pane's column arithmetic is unaffected.
Refusing the attach stopped the freeze but told the user to throw the session
away (`x` to close, `n` for new), which loses the conversation. tmux's own
dead-pane screen already says what to do instead: `claude --resume "<name>"`.
The Error card now offers `r` when the row can actually be resumed (claude,
with a conversation id and a working directory), and the footer says so. One
press resumes into a fresh pane and attaches to it, so a dead end becomes
recovery.
⚠️ Three things keep this from becoming the resume runaway that once spawned 35
sessions in 40 seconds. The offer holds a session ID, not a row, and is
re-resolved from the model when the key is pressed: a row captured when the
card opened is stale by then. It disarms BEFORE anything async, so a second `r`
cannot start a second resume. And it routes through resumeSelected(), which
owns the `resuming` flag and ends in attachToSession() rather than the group
dispatch.
⚠️ The `r` branch has to run BEFORE the generic dismiss, because a message
overlay is dismissed by ANY key: without that ordering the offer is consumed as
"some key was pressed" and the card merely closes. `help` keeps the any-key
behaviour, so the two modes no longer share a case.
Verified end to end against a genuinely dead claude pane: card, footer, one
press, one new session, and F12 back to the dashboard.
Three beta rounds died on tmux's native way out, and the last one died on the
instruction rather than the mechanism: "press Ctrl+B, release Ctrl, then d" is,
in the tester's words, very unclear, and holding the modifier through both keys
silently does nothing.
So the way out stops being a chord. The attach claims F12 in tmux's prefix-less
`root` table for its own duration, and the bar reads "press F12 to get back to
the codeman dashboard" — one keystroke, nothing to hold, nothing to release,
no order to get right. F12 because stock tmux ships an empty root table apart
from mouse bindings, and none of the CLIs that run in these panes want the key.
⚠️ The bar names the one key ONLY when the claim succeeded, and falls back to
the chord wording otherwise. A bar advertising a key that does nothing is the
bug this whole series started with, and it must not come back in a new costume.
Same claim rules as the prefix alias: taken only when tmux reports the key
unbound, given back only while it still means `detach-client`.
The chord and the held-Ctrl alias both keep working; they are simply no longer
what the user is told to press.
Reported three times as "Ctrl+B and d is still not working", on a build whose
bar already named the right key. Measured against a live pane: of the three
ways a person types this, only one worked.
Ctrl+B, release Ctrl, then d detaches
Ctrl+B then Ctrl+D (held) nothing happens
Ctrl+B then Shift+D nothing happens
Holding Ctrl through both keys sends 0x02 then 0x04, and tmux ships `C-d`
unbound in the prefix table, so the keystroke is swallowed in silence and the
attach looks frozen. That is not a user error worth documenting around: holding
the modifier is how most people type a two-key chord.
The attach now claims the held-Ctrl form of whatever key detaches (`d` → `C-d`)
for its own duration and gives it back on restore, and the bar advertises it
only once the claim succeeded, so it can never name a key that does nothing.
⚠️ The key is claimed ONLY when tmux reports it unbound, and released only
while it still means `detach-client`, so a binding of the user's own is never
shadowed or removed. The alias is deliberately excluded from the leaked-state
sweep: key tables are server-global, so the sweep cannot tell a leak from a
second TUI's live claim, and a stray `C-d`→detach is harmless either way.
Ruled out along the way, with evidence rather than assumption: the encoding.
tmux negotiates no extended-key mode upstream on attach (no kitty CSI-u, no
modifyOtherKeys, no DECSET 2017), so Ctrl+B does arrive as a plain 0x02 even
from a Claude pane, which has its own keyboard protocol.
Two more from the same beta round, both reported as "basic things are broken".
Attaching to a DEAD pane trapped the user. Codeman sets `remain-on-exit on`, so
a session whose agent has exited does not disappear: the row looks ordinary,
the server still reports it idle, and Enter handed the terminal to a pane that
reads no input. With the detach chord also wrong at the time, that was a hard
freeze with no way out. Enter now probes `#{pane_dead}` first and refuses with
an Error card naming the session and what to do instead. The probe fails OPEN,
so it can never block an attach to a live pane. ⚠️ It also has to paint: the
keypress that reaches attachToSession() has already painted by the time an
awaited probe resolves, so message() alone left the refusal invisible and Enter
looked inert, which is the bug it was added to fix.
A session started from the TUI came out unnamed, because startSession() sent no
sessionName and rowLabel() then fell back to the transcript's first line. A
brand-new session has no prompt to be named after, so the list showed a
perfectly healthy session called "Login interrupted" — the CLI's startup
output, reading like a failure report. Sessions the TUI starts are now named
`w<n>-<case>` like the web UI's, and rowLabel() prefers the case directory over
a scraped prompt for any row with a mux name, since a LIVE pane is identified
by where it runs while a history row genuinely is its prompt.
restore() runs after spawnSync returns, which covers detaching and the agent
exiting inside the pane, but not the terminal dying while attached. Closing the
window or dropping the SSH kills the TUI where it stands, and the bar it
installed stays pinned on the session: the next attach wears a stale bar naming
a different session, and the pane is a row shorter for good. Seen on the beta,
where the tester closed the window instead of detaching.
One sweep at startup, fire-and-forget so it can neither delay the first frame
nor fail a start. Only a bar carrying our own marker is touched, and the marker
is now the single source of the bar's own wording so the two cannot drift; a
user's hand-written status bar on the same session is left exactly as it is.
The session goes back to `status off`, which is how Codeman creates every pane
it owns and the only state this bar is ever applied over.
Two things the attach status bar got wrong, both found in a beta test.
The bar read `Ctrl+B D`. tmux key tables are case-sensitive: lowercase `d` is
`detach-client`, capital `D` is `choose-client`. Pressing what the bar said
opened a client chooser and left the tester attached, with the way out on
screen and inert. The key is now READ from `list-keys -T prefix` the same way
the prefix already was, rather than hardcoded, so a rebound tmux is followed
too and the label cannot drift from the binding again. It never goes through
formatPrefixKey(), which uppercases.
The bar also rendered as a full-width bright green slab. Only `status-format[0]`
was styled, so tmux's stock `status-style` (`bg=green,fg=black`) stayed
underneath it and won; `#[reverse]` on top could not undo it. `status-style` is
now set explicitly to `bg=default,fg=default` and snapshotted/restored with the
rest, so the bar sits on the terminal's own background and reads as a hint
line.
Tests pin both: that the chord ends in lowercase `d` and never ` D`, that a
rebound key prints verbatim, that `status-style` is part of the banner, and
that parseDetachKey() picks `d` out of verbatim tmux 3.4 `list-keys` output
while ignoring `detach-client -a`/`-P`, which act on other clients.
Three things the first beta test surfaced.
1. Attaching from a terminal of a different shape showed the pane clipped to the
browser's size, with tmux's dot padding filling the rest. Codeman pins every
window it owns to `window-size manual` at whatever the web client reports
(tmux-manager.ts), so no attaching client can resize it. The handoff now
brackets the attach with `window-size latest` and restores the snapshot on
detach. `latest`, rather than a one-off resize to our own size, is also what
lets a terminal resized MID-attach follow along: tmux recomputes on every
SIGWINCH while the TUI is blocked in spawnSync and cannot.
2. Nothing on screen said how to get back out, because Codeman keeps the status
bar off on its panes (the web UI carries that information around the terminal
instead). The tester exited the agent looking for the exit, leaving a dead
pane. An attach now wears a `status-format[0]` bar reading "<prefix> D
detach, back to the codeman dashboard", with the prefix READ from tmux rather
than assumed, and the session's options are put back exactly as they were on
detach. One option, not status-left/status-right, so tmux draws no window
list beside it; `reverse` so it inherits the terminal's own theme. Restoring
an array option unsets the BASE name, since dropping the `[0]` index leaves
an empty array, which renders as a blank bar on a session that had one. The
help overlay names the chord, and the dashboard confirms the detach.
3. Enter on a RECENT row said resuming was not wired up. It now creates a
session carrying that conversation (`resumeSessionId` plus `/interactive`,
the path the web UI's Resume Conversation list already uses), in the
directory it ran in and under its old name, then attaches to it.
The attach mechanics deliberately sit in a method the group dispatch cannot
reach, plus a re-entrancy flag: routing resume back through the Enter handler
re-dispatched on "this row is RECENT" and spawned one session per pass, 35 in
about 40 seconds on the beta before it was killed. test/tui/tui-e2e.test.ts
pins one press to one session with a pane that never appears, which is
exactly the case that looped.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both of the dashboard's periodic reads hit endpoints that are far more
expensive than their cadence assumed, and the cost lands on the SERVER's
event loop, so it is paid by every browser client too.
`GET /api/sessions/unified` is ~550ms against 11 live sessions: it scans
every Claude transcript plus the lifecycle log, uncached, and republishes
the search index. `scheduleRefresh()` was a 250ms trailing debounce with no
floor, and a queued refresh re-ran the instant the previous one returned
(by recursing, which also chained one pending promise per iteration), so a
stream of events paced the refetches at the endpoint's own latency: with
`session:updated` broadcast per session per 500ms while anything is
working, the scans ran back to back. `resyncDelayMs()` now keeps ambient
refetches 3s apart, measured start-to-start. The user's own actions call
`refresh()` directly and are unaffected, so what this paces is only
"notice what changed elsewhere".
`GET /api/sessions/:id/terminal` is ~80-100ms: two `execSync` tmux calls,
then the whole byte buffer normalized before the tail is taken. It was
polled every second for as long as a live row was selected. It now backs
off 1s, 2s, 4s, 5s while consecutive reads change nothing, and resets to 1s
on any change, when the selection moves, when this dashboard sends input or
answers a dialog, and on return from an attach. A pane that is printing is
still read every second; a pane at its composer is not.
The poll also kept running in three places it had nothing to draw for: the
whole time the user was attached in tmux (an attach can last hours), and
behind the message overlays that an async action opens (answered, killed,
started), which are not keystroke-driven and so never reached the
`afterInput()` path that stops it. `setInterval` becomes a chained
`setTimeout`, since the delay now varies.
Measured against the live server, same idle row selected, 25s window:
22 tail reads before, 5 after. With a working pane selected it stays at 22,
which is the intended cadence for a pane whose output you are watching.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`TuiModelStore.confirmSatisfied()` and `approvalFor()` had no caller
outside their own tests. The first one mattered: it answered "does the
typed text authorize this kill?" with an exact name match, while the rule
actually consulted (`confirmAccepts()` in tui-app) also accepts the
8-character id prefix a mux name carries. Two divergent answers to one
question, the stricter one unreachable and waiting to be picked up by
mistake. knip cannot see class members, so the dead-code sweep never
flagged either.
The tests they existed for now assert observable state instead, and the
approvals one got stronger on the way: it checks that a session id coming
back does not inherit the dead session's dialog, which is the invariant
`removeSession()` is actually keeping.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`TuiSessionRow` declared `lastSubmitAt`/`inputTokens`/`outputTokens`,
`stateSince()` ordered the WORKING group by the first of them and
`renderRowLines()` painted the other two, but nothing ever filled any of
them in: the unified list carries none, and the `session:updated` payload
that does was discarded (an event only schedules a refetch).
So a running turn was dated by its SESSION's creation instead. Measured
against the live server before the fix: w65 (created 21h ago, turn started
one minute earlier) outranked w67 (created 15 minutes ago, turn started
five minutes earlier), the reverse of the rule docs/tui.md states, and the
elapsed column read `21h` for a turn a minute old. The token column was
unreachable code for the same reason.
`fetchLiveSessionMetrics()` reads the three fields from `GET /api/sessions`
and `applyLiveMetrics()` folds them onto the rows. That route answers from
the server's cached LIGHT state (no terminal buffers): 10-20ms measured,
against the ~550ms the unified list in the same `Promise.all` already
costs, so it is cheap enough to ride every refresh. It is best-effort like
the approvals and tmux reads beside it, because losing the anchor is
better than losing the list.
A ZERO is treated as unknown rather than merged: `stateSince()` reads
`lastSubmitAt ?? createdAt` and 0 is not nullish, so a merged 0 would date
every never-submitted session to the epoch.
The snapshot path gets the same merge, or `codeman tui --list` would number
the WORKING group differently from the dashboard that `codeman tui <n>`
indexes into.
Verified live: working rows now show 28m/8m (turn age, tokens 280.5k/65.2k)
where they showed 21h/34m and no tokens. The e2e assertion fails on master's
wiring with `[*] 10m` against a session that pressed Enter one minute ago.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The data-dir half was already spelled out; the socket half only lived in
a function docstring, and the TUI is the first code that shells out to
`tmux -L` from a process that is not the server.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The inventory test predates the `tui` command, so a rename or an
accidental removal would have gone unnoticed: it now asserts the command,
its `-l`/`--list` flag and its optional position operand.
The digest and search-result lines joined their halves with an em-dash,
which the repo's own convention rules out, so both now use the middle dot
the surrounding lines already use. The one em-dash left in `src/tui/` is
load-bearing: `search-service.ts` builds a session snippet with it, and
the pattern that strips the repeated label has to match it.
Also moves `buildSearchEntries`'s doc comment back onto
`buildSearchEntries`; it had ended up stacked above a helper.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`mark()` had no callers (knip's only finding on this branch), and the
renderer's fallback help list advertised `r` resume, which is deferred
with the rest of phase 3: a help screen naming a verb the build does not
implement is worse than no help.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The user guide covers what the dashboard is (and is not), the two
non-interactive fast paths, the four groups and their ordering, the full
keymap, what answering an approval does server-side, and the SSH/narrow
and degraded cases. The example frame is a real 100x30 capture against
the E2E fake server, not a drawing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A server that comes up mid-run was upgrading the header's hostname and
version but not its chip, which then stayed blank until the next telemetry
event. Also swaps a typographic apostrophe out of a preview error, which is
not renderable on the ASCII glyph tier.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two small honesty fixes at the edges: a server that goes down leaves the
dashboard holding prompts nothing can classify any more and whose answer
route is unreachable, so degraded mode clears them; and a resize can cross
the narrow breakpoint, where there is no preview pane to poll for.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured against a live Claude pane: an Ink TUI paints by ROW and emits
almost no newlines, so dropping cursor-position sequences collapsed a whole
screen into one unreadable line, and a tail cut mid-sequence printed the
remains of it (";1H") as text. Now a jump to column 1 starts a display line,
a jump inside a row moves the write position (capped, since a stream may
address a column no terminal has), and a severed CSI head is dropped before
parsing.
The preview is readable against a real session as a result: tool calls, the
working line and the composer all land where they belong.
Also drop the repeated session name from a search row, whose snippet opens
with the name the row already shows in its first column.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The fake API server grows the routes the dashboard now calls (terminal tail,
input, approvals answer, search, away digest, plan usage on status), and the
new cases assert on what the server RECEIVED rather than on the frame: the
prompt arrives as one line ending in a carriage return, and the answers as
the exact action and option digit.
Also covered: the tail refreshing in place, the search overlay selecting a
live session, the digest rendering, one bell for an item announced twice,
and the 409 path reported as "no longer on screen".
The plan-usage chip is punctuated with the glyph tier's separator, so an
ASCII terminal no longer gets a stray middle dot in the header.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The dashboard stops being read-only. The selected session's tail is polled
once a second while the plain list has focus and the layout is wide, and an
unchanged tail never reaches the model, so a quiet session costs no repaint.
A row with no live buffer says so instead of polling forever.
Keys: y/n and the parsed digits answer the selected session's dialog through
`POST /api/approvals/:id/answer` (never a blind keystroke: that route
re-captures the pane and 409s when the dialog has moved on, which the TUI
reports as "no longer on screen"); `p` opens a one-line composer aimed at
the selected session; `/` searches with a 250ms debounce and Enter switches
to a live session result; `g` shows the away digest. A new prompt rings the
bell exactly once, tracked by item id so a repaint or a refetch cannot
stutter, and the plan-usage chip rides `GET /api/status` plus its telemetry
event.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The preview pane now leads with the pending dialog when the selected session
has one: the question, the options with their digits, and the keys that
answer them, red for a dialog and yellow for a waiting prompt. The card is
capped at half the pane, because the tail is why the pane exists.
Around it: a header badge counting prompts that need a human, a preview
title that sacrifices the path rather than the state word, the footer
becoming the composer line while one is open (with the cell the terminal
cursor belongs in, so it can be shown there and hidden everywhere else), and
the search and digest panels as overlays with a stable width.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The store gains the three overlays phase 2 needs, each taking the keyboard
when it is set and all of them cleared together by closeOverlay(), plus the
pure flattening of `GET /api/search`'s typed groups into rows a cursor can
move over: headers are chrome, and only a session that is on the list counts
as selectable, since a history hit has no row to move the cursor to.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three small pure modules the phase-2 verbs are built on:
- tui-composer: the single-line editor behind `p` and `/`, holding text as
code points so a cursor can never split a surrogate pair, with the scroll
window derived from the width rather than remembered.
- tui-approvals: what an approvals-inbox item's card says, which keys are
live for it (a digit answers only when the server parsed that option, and
an idle prompt answers to none of them), and which ids the bell has not
rung for yet.
- tui-digest: the away digest as compact lines, counts first and one line
per entry, with a capped tail per section.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Spawns the real command in a pseudo-terminal against a fake API server
(canned status/unified/approvals plus an SSE stream the test pushes
into), which is the only way to cover raw-mode key decoding, frames
reaching a terminal, SSE-driven refresh and the exit sequence that has to
restore the user's screen.
Two details the assertions depend on: frames are addressed absolutely
rather than newline-separated, so the parser takes the last COMPLETE
frame (the pty delivers one in several chunks, and reading a half-written
frame would be racy), and it reads the sidebar column only, or a name
echoed in the preview pane could answer for a row.
The child gets its own data dir and a tmux socket name nothing runs on,
so nothing here can see or touch the machine's real sessions.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`codeman tui` opens the dashboard, `codeman tui --list` prints the
numbered list and exits (the `sc -l` replacement, plain when piped) and
`codeman tui <n>` attaches straight to a row (the `sc 2` replacement).
Both fast paths short-circuit before any screen setup, and both refuse
the numbers path without a terminal instead of half-opening a UI.
Bare `codeman` still prints help: the web UI stays the primary surface.
The TUI module is imported lazily so the other commands do not pay for it
at startup.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The IO half of src/tui: it owns the terminal, the timers, stdin and the
tmux handoff, and every decision it makes that is a function of its
inputs is an exported pure helper with unit tests (attach planning, the
typed kill confirmation, keymap selection, the repaint test, degraded
rows).
What it does: live session list over the unified API with SSE-driven
resync (debounced, with a 2s poll fallback the client asks for), cursor
and 1-9 navigation, attach and return, kill behind a typed confirmation
that refuses history rows and the session hosting the TUI, a new-session
case and CLI picker over quick-start, and degraded mode straight from
tmux when no server answers, re-probing so a server that starts upgrades
the dashboard in place.
Restoring the terminal is the part that has to be bulletproof: leave() is
idempotent and runs from normal quit, SIGINT/SIGTERM, a process exit hook
and prepended fatal handlers (src/index.ts already handles those by
exiting, so a listener registered after it would never run).
Attach is a handoff, never a proxy: the screen is restored and tmux gets
the real terminal. Inside tmux on the same socket there is nothing to
hand off to, so it issues switch-client and exits.
The preview pane, approvals answering, the prompt composer, search and
the digest are the next step; the region renders a placeholder rather
than pretending to load something.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The footer and the help overlay held the plan's full keymap, which would
advertise verbs (prompt, search, digest, answer, resume) that the build
does not implement yet and teach users that the TUI ignores keys. Both
now take their entries from the render options when the caller passes
them; the built-in lists stay as the fallback.
The picker overlay windows its items around the cursor rather than
clipping them, so the selected case stays visible in a long list.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The app layer repaints on state change, so the store has to be able to
say that something changed: `revision` is bumped by every mutating
method, and the repaint test compares it against the last painted frame.
Without it an idle dashboard would either redraw on a timer or go stale.
Three additions come with it, all optional so nothing existing changes
shape: `TuiSessionRow.muxName` (the unified list carries no mux name, so
the app fills it in from the local tmux enumeration and a row without one
cannot be attached), a `new-session` UI mode, and `TuiPickerState`, the
one-column chooser behind `n`.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Everything the dashboard needs from outside the process, behind one typed
surface, so the app loop stays a loop. It is a client of the running server and
nothing else: rows come from the unified list, blocked states from the
approvals inbox, and answering goes through the endpoint that re-captures the
pane and refuses with a 409 when the dialog has already been answered in tmux.
That refusal is a typed result rather than an exception, because a human
beating you to a prompt is normal operation.
Discovery mirrors the daemon probe (`CODEMAN_API_URL`, else loopback on
`CODEMAN_PORT`, self-signed TLS accepted) and credentials come from where
`codeman attach` already reads them. An explicit port outranks the ambient
`CODEMAN_API_URL`, which every managed session exports: a caller that named a
port must not be redirected at whatever server owns its shell.
Input is single-line and `\r`-terminated at this layer, so no caller can strand
text on an unsubmitted composer, and each send is tagged for the server's
exactly-once path. The event stream defaults to a `?sessions=` filter that
matches nothing, which drops the terminal firehose while lifecycle, hook and
approval events still arrive. A silent-but-open stream is caught by a watchdog
rather than a socket error, since that failure mode reports nothing at all.
With no server answering, sessions are listed from tmux on the instance socket
(argv, never a shell string) and decorated from a read-only peek at state.json,
which keeps the "the server died, get me to my sessions" path alive.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Node has no EventSource, so the live-update stream is read as raw bytes and
decoded here. Three details are what the parser exists for: a TCP read can end
between the CR and the LF of a CRLF, so a trailing CR is held back rather than
dispatched; the tunnel padding the server appends after a frame is a comment
with no blank line after it and must not split anything; and the keepalive is a
NAMED event, because an SSE comment is invisible to a browser client by spec.
Event classification lives here too, as a set rather than a prefix test:
`session:terminal` is most of the stream and the preview pane pulls its own
tail, so it is deliberately not a resync trigger.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The socket name was computed inside tmux-manager, which the TUI cannot import
just to learn which `-L` name its degraded-mode listing belongs on (that module
is the server's tmux driver, not a lookup table). The resolver moves next to
`dataPath()`, where the other half of the instance identity already lives, so
both processes agree by construction instead of by a copied default.
Behaviour is unchanged: the override still wins only when it is a name that can
be passed to `tmux -L` safely, and TmuxManager keeps warning about one that
cannot.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One absolutely-addressed line per row, each closed with an erase-to-end, so
nothing scrolls and a repaint cannot leave the previous frame's tail behind.
The caller wraps the result in synchronized-output brackets; that is an IO
decision and stays out of the renderer.
Color is passed in rather than detected. chalk's detection is right for the
one-shot CLI but would make a frame non-deterministic, so the palette is raw
SGR in the same semantic roles cli-style uses, and `color: false` emits nothing
but the cursor addressing, the session's own colors in the preview included.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Below 72 columns the preview pane is dropped and rows take two lines, the
constraint the `sc` chooser was built around and the reason it is still usable
on a phone; above it a clamped sidebar carries the list and the preview takes
the rest.
Every region is clamped to a non-negative size, so a 5x5 terminal degrades to a
header instead of handing the renderer negative widths.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Rows are the ones GET /api/sessions/unified already returns and blocked states
are the items the approvals inbox already parsed, both imported as types only
so a CLI process pulls in neither the server nor node-pty. Classification
speaks the web UI's language (red blocked, yellow waiting, green working) so a
user with both surfaces open never has to translate between them.
Groups order by how long a session has been in its state, which is why WORKING
anchors on the pane's last Enter: a working pane repaints about once a second,
so its last-activity stamp always says "now".
Selection is tracked by session id, never by row index: rows re-sort under the
cursor whenever a session starts working or an approval lands, and an
index-tracked cursor would quietly move the selection to another session
between two keystrokes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Decodes printable UTF-8, the control keys, arrows in both CSI and SS3 forms and
SGR mouse reports out of a byte stream that can tear anywhere, so a sequence
split across two reads decodes the same as one that arrives whole.
A lone ESC cannot be told from the start of an arrow key by looking at bytes,
so the parser holds it and the caller resolves it with flush() once its
disambiguation timer fires. Unknown sequences are swallowed: a stray CSI must
never reach a prompt composer as typed text.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The preview pane shows a session's raw terminal stream, so it needs the tail
reconstructed rather than emulated: SGR survives, cursor steering and OSC do
not, and a carriage return returns to column 0 so a spinner that repaints its
line 200 times contributes one line instead of 200.
Widths count East Asian Wide characters as two columns, which the clip and pad
helpers rely on to never cut a wide character, a code point or an escape
sequence in half.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The entry dates from an abandoned prototype (0.1427) and would have kept the
real TUI modules untracked while `git status` stayed silent about it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`codeman attach <path>` posts an attachment card for a local file; it
was described as attaching a Claude hook context. And Codeman never
overrides the tmux prefix for local sessions (only remote-SSH and docker
panes get C-q), so the detach hint is Ctrl+B D, matching the chooser.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The file asserted against a hand-written fixture array with its own
argument parser, so it could not see a command being renamed, losing an
alias or disappearing, and it described a `tui` command that does not
exist. It now walks program.commands: names, aliases, subcommands,
option flags, operands, descriptions, and a guard against registering a
name or alias twice at one level.
Assertions are "at least this exists", so a new command (including the
tui one this plan adds later) passes without editing the test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The startup banner is now the only one (the CLI printed a duplicate) and
is painted like the rest of the CLI. The non-loopback-without-password
warning was plain console.warn while the CLI's copy of the same warning
was yellow; chalk degrades off a TTY, so journald and web.log stay free
of escape codes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- doctor is colorized through the ReportStyle hook: verdict glyph and
failing status text painted, paths and hints muted, versions left
alone. `doctor --json` still prints raw JSON.
- `codeman web -d`, `web --stop` and `service install` block for up to
30s polling /api/status; each now runs under a spinner instead of a
silent terminal.
- `codeman reset` asks a real y/N question on a TTY. Non-interactive
callers keep the old "Use --force to confirm." refusal, so no script
can be answered by a question it cannot see.
- `codeman list` was a drifted copy of `codeman session list`; both now
call one renderer, with the shorthand opting out of the stopped and
web-server sections.
- `web` no longer prints its own "running at" line: the server prints
one, and unlike this one it also covers the daemon and service paths.
- every chalk call goes through the palette, so the CLI has one place
where colors are decided.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
"Antigravity CLI" is 15 characters and the hardcoded padEnd(14) pushed
that whole row one column right. Widths now come from the widest cell.
The header always said the CLI layer may colorize, but there was no way
to: renderTable now takes an optional ReportStyle whose hooks are
identity by default, so the module still decides nothing about color and
its output stays byte-stable. Padding is applied outside the paint, so a
row with no path detail ends at its status text instead of trailing
spaces inside a color run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One vocabulary for everything the codeman CLI prints: semantic palette,
the glyph set the commands already used, heading/rule/kv, width-aware
table layout, a stderr spinner and a y/N confirm.
Color detection stays chalk's, so NO_COLOR and non-TTY degradation keep
working with no second detector to disagree with it. The layout math and
glyph selection are pure and exported, which is what lets the dependency
report reuse them while staying color-free.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Post-merge follow-ups for PR #329 (shared CLI executable resolution):
- Negative-cache resolution misses with a doubling backoff (1min -> 5min
cap, cliResolveRetryDelayMs, mirroring claudeVersionRetryDelayMs): the
shared resolver cached success only, so a missing CLI re-ran the whole
chain - ending in a synchronous interactive login-shell spawn bounded by
the 5s EXEC_TIMEOUT_MS - on every /api/<cli>/status request and Run
attempt, stalling the event loop each time, forever. Success still caches
for the process lifetime, so an installed CLI is picked up within minutes
without a restart. Tests drive the backoff via an injectable clock
(createCliExecutableResolver `now` option, threaded through the
createPiResolverForTest / createAntigravityResolverForTest wrappers).
- Pass killSignal: 'SIGKILL' on the resolver's login-shell spawn and on the
pi/claude --version probes: execFileSync's timeout only SENDS the kill
signal and then keeps waiting for the child to exit, and interactive bash
ignores SIGTERM, so a login shell stuck in a blocking .bash_profile
survived the timeout and blocked the server permanently.
- Restore test hermeticity (PR #329 deleted pi's VITEST guards, and one
test pinned the deletion): under vitest the production resolver host now
replaces un-injected IO primitives with inert stubs - no real PATH
scanning, no login-shell spawns - and probePiVersion never executes a
`pi` candidate again (`pi` is a generic binary name, so route tests
hitting /api/pi/status executed whatever binary the machine carried).
Tests opt in through the runCommand/isExecutableFile injection hooks or
allowRealIoUnderVitest for real-filesystem fixtures. The deletion-pinning
test is replaced by behavioral pins, including a real-executable fixture
in the new test/pi-cli-resolver.test.ts that fails loudly if the pi gate
is ever removed again.
- Wire the six get*NotFoundMessage() exports (previously dead) into their
intended call sites: the createSession throws in tmux-manager and the
availability gates on POST /api/sessions and POST /api/quick-start in
session-routes, replacing a third hardcoded copy of the text. A not-found
error now names where resolution looked (server PATH, login shell,
checked directories). npm run knip no longer reports any unused export
from the resolver modules.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Post-merge follow-ups for #328 (GET /api/system/repo-status):
- Event-loop blocking: every git invocation in repo-status.ts is now async
(promisified execFile), never execFileSync — the per-remote ls-remote +
fetch could hold the event loop (SSE, PTY streaming) for up to ~60s per
request. The whole computation is single-flight with a 45s TTL cache
(createSingleFlightCache): concurrent requests share one in-flight
promise, a fresh result is served without spawning git, and a rejected
compute is never cached. Route handler shape and response fields
unchanged; remotes still processed sequentially (concurrent fetches in
one repo contend on ref locks).
- Credential disclosure: the redaction from git-clone.ts is extracted as
exported redactGitCredentials() (sanitizeGitOutput now uses it) and
applied via redactRemoteStatus() to every remote card's url and error
string, so a scheme://user:token@host remote URL (or git stderr echoing
it) never reaches a client.
- Non-interactive env: runGit() now uses the shared gitNonInteractiveEnv()
instead of a partial GIT_TERMINAL_PROMPT/BatchMode env, also closing the
GIT_ASKPASS/SSH_ASKPASS/SSH_ASKPASS_REQUIRE/DISPLAY/GCM_INTERACTIVE
prompt paths.
- Upstream parse bug: a local-branch upstream (@{upstream} with no slash,
e.g. after `git branch -u otherbranch`) made slice(0, indexOf('/')) into
slice(0, -1) and yielded garbage like "maste". parseTrackingRemote()
(pure, unit-tested) returns null for it, and the bare ref is dropped so
it cannot be mistaken for a remote-tracking ref downstream.
Tests extended in test/repo-status.test.ts (parseTrackingRemote,
redactGitCredentials/redactRemoteStatus, createSingleFlightCache
single-flight/TTL/rejection semantics).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three post-merge fixes for the external-CLI response viewer:
- ?context=full blocks now carry role ('user' for prompts, 'assistant'
for response/status/tool). The frontend's loadFullContext() renders
via msg.role, so the roleless blocks lost the "You" badge and every
turn rendered as the agent. kind/label/text are unchanged and the
frontend needs no change.
- normalizeDividerStatusLine() dropped its backtracking regex
(/^[─-]+\s*(.+?)\s*[─-]{3,}$/): the lazy middle went catastrophic on
a long dash run without a 3-dash tail (measured 15.5s at 4,000 chars,
minutes at 10,000), and pane text is agent-controlled with buffers up
to 32MB. Replaced by a linear counter walk with the identical accept
set and captured content, pinned char-for-char against the old regex
by a brute-force corpus test plus a hostile-input regression test
that fails by timeout with the RegExp version (same approach as the
glob-matcher hardening in 68ae9a8).
- 'pi' joins EXTERNAL_CLI_MODES: pi sessions had the identical
empty-viewer symptom the transcript branch exists to fix. The list
stays a local duplicate of isExternalCliMode() (importing session.ts
would drag node-pty into the pure module); a new exhaustive parity
test asserts the two mode sets can no longer drift.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
#325 renamed _sessionUsesServerMouseStrip to _shouldReportMouseToCli and
added the server-observed cliMouseTracking half of the gate, which also
turned codex tap reports from measured no-ops into not-sent-at-all. The
invariants paragraph still described the old name and the old behavior.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The reschedule guard is `this._status === 'running' && this.loopTimer === null`,
but the timer callback never nulls `loopTimer`. So the handle stays non-null from
the first fire onward, the guard is false on every subsequent pass, and the Ralph
loop silently stops polling after exactly two ticks.
It stops without changing status: `status` stays `running`, `stop()` is never
called, and no error is raised — the loop just quietly never runs again, which is
what makes it hard to notice on a long autonomous run.
Null the handle inside the callback before re-entering `runLoop()`, which is the
pattern `orchestrator-loop.ts` already uses for its own reschedule.
Test: a regression case in test/ralph-loop.test.ts that runs a real 5ms-interval
loop for ~16 intervals and asserts it ticks at least 3 times. Against the unfixed
source it reports exactly 2.
A CLI installed by nvm, Homebrew or a user-level npm prefix lives on a PATH that
only a login shell sets up. Codeman running under systemd or launchd does not get
that PATH — launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` — so every
resolver reported the CLI as unavailable on installs where it is plainly there
and works from a terminal.
Each of the six resolvers had its own hand-rolled copy of the same PATH walk, so
the fix is factored into one shared `createCliExecutableResolver()` with an
explicit lookup order: the server process PATH, then common install directories in
order, then an interactive login shell as the last resort. Only the last step
spawns anything, and only when the cheap lookups have already missed.
Also adds `formatCliNotFoundMessage()`, so a failure explains where it looked
instead of just asserting the CLI is missing. Its diagnostics are bounded and
control characters are flattened, so a not-found message cannot dump arbitrary
environment data.
Success is cached and failure is retried, so installing a CLI while the server is
running is picked up without a restart.
Net -103 lines across the six resolvers. Behaviour is unchanged wherever the CLI
was already on the process PATH: that remains the first thing checked.
Tests: 20 cases in test/cli-executable-resolver.test.ts covering the precedence
order, login-shell-only resolution, the caching rule, unsafe-name rejection, and
the bounded diagnostics.
`GET /api/system/update/check` answers "is there a newer published release
tag?", which is the right question for an npm install but not for a git clone
that tracks a branch. Such an install can be many commits behind its own remote
while the latest tag says it is current, and nothing surfaces that.
Adds `GET /api/system/repo-status`: an informational companion that reports what
this CHECKOUT looks like against its own remotes — current branch and commit,
ahead/behind counts per remote, the remote's role (tracking / upstream / other),
and a bounded list of incoming commits.
Read-only and defensive: every git invocation is `execFileSync` with an argv
array and a timeout, a non-git or remote-less install reports a structured
`error` rather than throwing, and nothing here mutates the working tree or
touches the updater's own state.
Tests: 24 cases in test/repo-status.test.ts.
`GET /api/sessions/:id/last-response` branches to a Codex-specific reader, then
falls through to scanning `~/.claude/projects` for a transcript. OpenCode, Gemini
and Antigravity render their own TUIs and never write one, so that scan finds
nothing and the response viewer is permanently empty for all three modes.
For these CLIs the pane IS the transcript, so segment it. `response-viewer-transcript.ts`
is a pure, dependency-free parser that splits a terminal buffer into prompt /
response / status / tool blocks, keying off the `›` prompt marker, status
dividers and `• Calling|Called` tool-activity lines. The route uses it to answer
with the LAST response, and to carry the parsed blocks under `?context=full`.
Codex keeps its existing branch: it has real rollout files, which are a better
source than scraped pane text.
The response shape is unchanged for every other mode, and Claude panes are
explicitly pinned to the Claude transcript path so a real transcript can never
be shadowed by scraped text.
Tests: 14 parser cases plus a route suite covering all three modes, the
`?context=full` payload, an empty pane, and the Claude regression guard.
Confirming an AskUserQuestion left its tab flowing red for the rest of
the turn (owner report: ~8 minutes on a running session, with no dialog
anywhere on screen). Two separate bugs, both live-verified.
The re-capture erased the evidence the staleness check runs on. Claude
Code fires the Notification behind the dialog (measured 6-7s on v2.1.237,
documented up to ~30s), so the 600ms re-capture routinely lands on a
frame the user has ALREADY answered, parses nothing, and applyCapture
overwrote item.options with undefined. A MISSING options is how "we never
could read this dialog" is expressed, and those items stay answerable by
design, so a cleared field was indistinguishable from a never-parsed one
and the item became permanently unsweepable: it survived every
GET /api/approvals and every page reload, cleared only on `stop`, and
still accepted an answer, sending a bare `1` into a composer with no
dialog under it. applyCapture is now ADD-ONLY for options.
Nothing ran the staleness check while a page was open. It lived only in
GET /api/approvals, which seedApprovals() calls on init and reconnect, so
`stop` was the first thing that ever cleared an answered dialog. The
`working` signal now runs the pane-VERIFIED variant (resolveIfDialogGone
-> verifyStillAnswerable): the heuristic only decides when to look, the
screen decides the outcome, so the existing "working can flap" rule is
respected.
A frame that parses no options is now conclusive in two cases, and only
those, so an unreadable capture still keeps the alert: the item once
parsed options, or the frame shows Claude actively running a turn. A
modal dialog BLOCKS the turn, so the two cannot coexist - measured, a
live-dialog frame carries neither the elapsed-timer spinner nor the
"esc to interrupt" footer, which the dialog replaces with "Enter to
select". That second signal is reached by a delayed staleness pass (3s)
scheduled alongside the re-capture, which closes the late-hook case where
the prompt is answered before the hook lands: nothing ever parses, `stop`
may have gone by already, and the alert outlived reloads until the 12h
TTL. The pass is deliberately later than RECAPTURE_DELAY_MS, whose whole
reason for existing is that the hook can beat Ink to the screen.
Frontend: _onHookElicitationComplete cleared only the elicitation entry,
but an AskUserQuestion arrives as permission_prompt, so it was clearing
the wrong alert; it now clears both, matching the server's kind-agnostic
APPROVAL_RESOLVING_EVENTS.
Verified end to end on an isolated beta instance, not just in unit tests:
before, resolution could only come from the stop route (approval:resolved
always immediately preceding hook:stop); after, it arrives from the new
paths, and a simulated late hook resolves at +3.12s with no stop, no
working signal and no GET, while the pane is still working. Tests use
frames captured off a live pane and each new one was confirmed to fail
against the old behaviour.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Found while verifying Auto Copy in a browser: a plain left click in a
claude/codex/gemini pane sent a synthetic SGR mouse report into the PTY
whether or not the program in that pane had ever enabled mouse tracking.
When the pane holds a plain shell (the CLI exited, or a shell was started
inside a session of that mode) readline prints the report as literal text
and it garbles the next line typed:
$ [<0;88;20Mecho hello
bash: 0: No such file or directory
The cause is that the browser could not know. The full strip
(isAltScreenStripMode) removes the mouse DECSETs from the stream, so
xterm's modes.mouseTrackingMode is permanently 'none' for those modes and
_sendSyntheticSgrTap() hand-encodes reports to stand in for xterm's own
encoder. With no state to consult it had to do that on every click.
What the strip removes, the server now remembers.
_recordStrippedMouseMode() records each sequence as it is stripped,
toState() publishes it as cliMouseTracking, and the browser's
_shouldReportMouseToCli() (renamed from _sessionUsesServerMouseStrip)
requires it at all three report sites: the desktop click, the touchend
tap, and the mobile tap classifier.
Details that are easy to get wrong:
* Only the tracking modes count (1000/1001/1002/1003). 1005/1006 select
an encoding and 1007 is alt-scroll; a CLI that picks SGR encoding
without turning tracking on is not asking about clicks, and counting
those would put the stray reports straight back.
* Modes are held in a Set, so a TUI disabling a mode it never enabled
cannot clear the ones that are really on.
* The change broadcasts immediately instead of through
broadcastSessionStateDebounced: the flag flips when a dialog opens, and
the user can click that dialog well inside the 500ms debounce window.
* It fails toward silence. After a server restart the flag is false until
the CLI re-emits its DECSET, which tmux does at client attach.
Verified against a live claude 2.x session: the CLI holds a tracking mode
on continuously, so its clicks are still reported byte for byte as
before, while a bash prompt in the same stripped mode now reports
nothing and types cleanly. The flag also propagates live over SSE in both
directions, checked by toggling ?1002h/?1002l from inside the pane.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
App Settings > Terminal & Input > Selection & clipboard > Auto Copy
Selection (`autoCopySelection`, per-device, default OFF). With it on,
highlighting text in the terminal copies it: mouse drag, double-click
word, triple-click line, and the phone long-press selection. Ctrl+C is
untouched and still copies on demand.
Three things decide the shape of it:
* It fires at the END of a gesture, never in onSelectionChange. That
callback runs for every cell a drag crosses, so copying there would be
one clipboard write per mouse move. It only arms a pending flag; a
document-level mouseup listener flushes, and the touch path calls the
flush itself because it preventDefaults its touchend and no mouseup
ever arrives there.
* The flush is synchronous inside the handler, because both clipboard
paths need user activation: Firefox gates navigator.clipboard
.writeText on it, and execCommand('copy'), the fallback the plain-HTTP
LAN install lands on, has to run in the gesture's own task. A timer or
a wait for onSelectionChange loses it, invisibly in Chrome.
* It deliberately does NOT do what copyTerminalSelection() does. That
one clears the selection (so a second Ctrl+C is an interrupt) and
focuses the terminal. Clearing would make text vanish under the cursor
that just highlighted it, and focusing opens the on-screen keyboard
over it on a phone. Focus is instead restored to whatever held it,
which only matters for the execCommand fallback.
Guards are pure in decideAutoCopy() (constants.js): off, blank or
whitespace-only text, and a 1M-char cap, since a drag off the top of the
viewport autoscrolls and one gesture can sweep the whole 50k-line
scrollback. Past the cap the copy is refused rather than truncated, with
a toast pointing at Ctrl+C.
Feedback is silent on success except once per page load, so a feature
that works by doing nothing visible can still be told from a dead
toggle; failures and refusals toast, throttled to 10s.
Per-device on both counts the settings rule requires: in `displayKeys`
and absent from the .strict() SettingsUpdateSchema, because clipboard
access differs by device and by origin.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The last two items of #322: copyTerminal() copied the entire buffer but
was wired to no button, shortcut or call site anywhere, and it wrote
through navigator.clipboard directly, which is undefined on the
plain-HTTP LAN install, so it would have failed there even if it were
reachable. Everything that actually copies goes through
copyTerminalSelection() and _copyText's execCommand fallback; whole-
buffer copy, should anyone want it, is a selectAll() away from that
same working path.
Closes#322
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Files search compiled the user's query into a backtracking RegExp:
'*a*a*a...' became '^.*a.*a.*a...$', the classic blowup, evaluated
synchronously against every walked path — a pathological query could
freeze the event loop for the whole server (and every user of it in
multi-user mode). /api/search stays regex-free for exactly this reason.
Globs now match through a two-pointer wildcard walk, O(text · pattern)
worst case, with a 256-char query cap bounding the pattern side; an
overlong query compiles to null, the same answer as an empty one.
Semantics are unchanged (anchored, case-insensitive, * spans slashes)
and the existing tests pass untouched; the pathological pattern gets a
test that fails by timeout with the RegExp version.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The new local-echo-overlay gotcha landed as a list item but left the
xterm-zerolag-input entry below it without its leading '- ', splitting
the Common Gotchas bullet list in two.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
An agent's numbered list wraps its URL, and the link opened a PREFIX of it:
1. https://github.com/users/someone/packages/container/p
ackage/thing
opened `…/container/p`. The provider already stitched hard wraps — Ink emits a real
newline, so nothing is flagged `isWrapped` and a row that fills the last column is
taken as continuing — but it joined the row texts VERBATIM, and the continuation
carries the list's own three-space indent. That whitespace lands in the middle of
the token, which is exactly where the URL pattern stops. Flush-left wrapped URLs
(Claude Code's own `/login`) worked, which is why this survived.
The touch-selection helpers had the shallower version of the same bug: they walked
`isWrapped` only, so `Line` grabbed the single row on screen rather than the
logical line, and a long-press on a wrapped token selected only its visible half.
So the reconstruction now lives in ONE place, `terminalLogicalLine` in
constants.js, and both consumers use it — the link provider matching patterns over
its text and the selection helpers measuring words and lines with it. A link that
spans a wrap and a `Line` that stops at the screen edge were the same bug twice.
The helper drops the leading whitespace of a HARD continuation (the program's
indent) and keeps that of a SOFT one (the emulator inserts nothing, so it is real
content), records the dropped width per segment so the offset↔cell mapping stays
exact in both directions, trims only the final row so earlier offsets stay aligned
to cells, and keeps the 12-row bound that stops a screenful of full-width output
from being re-scanned on every hover.
⚠️ Selection spans are computed in CELLS, not text offsets: an xterm selection is
one contiguous run, so a token spanning a hard wrap also covers the indent cells
between its halves. A run that skipped them cannot be expressed, and would not
match what is highlighted.
Tests: `test/terminal-logical-line.test.ts` (8 cases: the indent drop, resolving
from either row, both mapping directions, soft continuations kept verbatim, no
over-reach past a short row, the row bound, final-row trimming, a missing row) and
5 in `terminal-touch-tap.test.ts` (the whole URL from either row, a token selected
across the wrap, `Line` spanning both rows, no reach into the next line). Removing
either half of the fix reds 5 and 8 of them respectively.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The three fixes in this branch change what a tap and a long-press MEAN on a
phone, and add a UI surface with its own z-index — all of which this repo keeps
written down rather than discoverable only by reading the handlers.
- `docs/wiki/Mobile-Guide.md` (the published user manual): a new "Tapping, links
and copying" section, and the long-prompt behaviour in the keyboard section
where the existing scroll/tap rules live.
- `CLAUDE.md`: the touch-gesture invariants next to the scrollback/wheel material
(why the caret line is the boundary rather than the tap intent; why all three
selection guards exist), the overlay's new bottom bound alongside the
single-source note, and the selection bar in the z-index registry — 900, above
terminal content and the local-echo overlay and deliberately below floating
agent windows so it can never cover their controls.
- `i18n.js`: zh-CN for the bar's `Copy` / `Line` / `Clear selection`. The bar is a
SIBLING of `.xterm`, not a descendant, so `SKIP_SELECTOR` does not cover it and
the entries actually apply.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Typing a prompt long enough to wrap ran the text off the bottom of the screen: the
tail — the part being typed, where the cursor is — sat behind the on-screen
keyboard, so the user was typing blind. Two independent causes.
**The overlay had no bottom bound.** On touch devices keystrokes are buffered in
the local-echo overlay and do not reach the PTY until Enter, so the CLI never
learns the prompt is long and nothing scrolls or reflows to make room. Meanwhile
the renderer lays its wrapped lines out straight DOWNWARD from the prompt row
(`top = promptRow * cellH`, each line at `i * cellH`) with nothing clamping it to
the visible rows — and with the keyboard up there are only a handful of those.
The block now grows UPWARD once it would pass the last visible row: it is lifted
so its final line lands ON that row. Every line div is opaque, so it covers
transcript above rather than vanishing under the keyboard below — the same thing a
real terminal does when a composer expands. A prompt taller than the whole
viewport keeps its TAIL, for the same reason the fix exists: the end is what the
user is looking at. `startCol` indents only the line that starts at the prompt
marker, so it is dropped along with that line when only the tail fits, and the
cursor follows the last VISIBLE line.
`rows` joins the render key: the layout depends on it, so a keyboard opening —
which changes rows without changing the text — must not be skipped as a redundant
render.
**`_shrinkPaddingToFit()` was reclaiming the bars' own space.** On phones the
toolbar and accessory bar are `position: fixed`, so they occupy no layout space
and `main`'s padding-bottom is the ONLY thing reserving room for them. Shrinking
it by the full sub-row slack pulled the terminal's bottom edge down underneath
them, and the row the following re-fit gained was painted behind them — clipping
the last line of a long prompt. The shrink now has a floor: the MEASURED height of
the currently-visible fixed bars, so genuine over-reservation of the hard-coded
84px is still reclaimed while a device that needs those pixels keeps them. The
floor is `Math.min(currentPadding, measured)`, so it can only ever prevent a
shrink, never cause a grow that would resize the terminal as a side effect.
Overlay behaviour lives in `packages/xterm-zerolag-input/` (single-source; the
vendor bundles are generated), so the fix is in the package with the row count
passed in as an optional `totalRows` — absent, the layout is exactly as before.
Tests: 7 cases in the package's `overlay-renderer.test.ts` (upward lift, tail
retention, indent drop, cursor on the last visible line, and the unclamped
fallbacks) and 7 in a new `test/mobile-keyboard-bottom-padding.test.ts` (reclaim,
floor, partial reclaim, no-grow, hidden bars, CJK strip, whole-row slack). 5 and 4
of them respectively fail without the fix. Package suite 238 pass, including the
codex byte-identity and replay tests.
Verified on Android + Chrome against a live instance: a ~460-character prompt
wrapping ~12 rows stays on screen while typing and arrives at the PTY intact.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was no way to copy terminal text from a phone at all, and three layers
ruled it out independently: `user-select: none` across the whole terminal subtree
on touch devices (taps are cursor gestures there, so the OS callout had to go),
the WebGL renderer drawing glyphs as pixels with only the accessibility tree
behind them, and xterm's own selection being a mouse DRAG while the touch path
dispatches a zero-movement mousedown/mouseup pair — a click. `copyTerminal()`
exists but is wired to no button and calls `navigator.clipboard` directly, which
is undefined on the plain-HTTP LAN install the installer offers.
So the gesture drives xterm's `select()` directly: public API, renderer-
independent, and the highlight is drawn by xterm itself. Long-press is free real
estate — tap and swipe are taken, long-press and double-tap are used by nothing.
- **Long-press** (350ms, finger still within the shared tap slop) selects the
run of non-whitespace under the finger. Whitespace is the only delimiter on
purpose: every punctuation-aware word rule cuts a path, URL or hash in half,
which is what you came to copy.
- **Drag** while held extends the selection; touchmove diverts from scrolling.
- **Tap** while the bar is up extends it too. That is the ergonomic core:
picking up a 4px handle with a fingertip is a coin flip, tapping the other end
is not. Dismissal stays explicit (✕ or Copy), so no tap is spent leaving a mode
the user is still using.
- **Copy** goes through the existing `copyTerminalSelection()`, so it inherits
the execCommand fallback that is the only route that works on plain HTTP.
- **Line** takes the whole logical line, wraps included, trailing pad trimmed.
Three guards are what make the gesture survive contact with a real phone, and
each fixes a symptom measured on Android Chrome:
1. **The compat mouse pair after touchend.** xterm focuses from its screen-element
mousedown and SelectionService resets the model there, so lifting your finger
popped the keyboard and dissolved the selection in one go. The tap path already
had a guard for those events; the selection path simply never armed it. Armed
now, and the touchend is `preventDefault`ed so the synthesis is stopped at the
source (that listener is no longer passive).
2. **The platform's own long-press.** Android Chrome runs its handling at ~500ms
and focuses the nearest editable element — xterm's helper textarea, parked at
the cursor — which no touch handler can preventDefault because it never sees an
event. A focus guard blurs the terminal input for the duration of the gesture,
whatever focused it, bounded by a self-expiring deadline so a stuck flag can
never leave the keyboard unreachable. `contextmenu` is suppressed for the same
window, and the threshold sits at 350ms so it lands clear of the platform's.
3. **Copy re-focusing the terminal.** `copyTerminalSelection()` ends with
`terminal.focus()`, which is right on a desktop and wrong on a phone: the
keyboard covers what was just copied with nothing waiting to be typed.
The bar is built in JS because index.html is read once at server start, and its
styles live in styles.css rather than mobile.css because the gesture is
touch-driven, not width-driven — a touch tablet in landscape gets the gesture and
would otherwise have no bar to copy from.
12 tests in `terminal-touch-tap.test.ts` cover the word rule, forward and
backward extension, cross-row selection, Line, tap-to-extend, the copy path, and
each of the three guards including the focus guard's expiry.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
On a phone no link was openable, on either surface, for two unrelated reasons.
**Terminal.** xterm resolves the link under the pointer on `mousemove` and
activates it on `mouseup` over its SCREEN element. A touch tap delivers neither:
`touch-action: none` on the terminal subtree plus touchstart's preventDefault for
a 'content' tap suppress the browser's compatibility mouse events,
`_installMobileTapMouseGuard` drops the trusted ones that still arrive inside the
450ms tap window, and the synthetic mousedown/mouseup pair dispatched for mouse
REPORTING goes to the `.xterm` root — an ancestor of the node the linkifier
listens on, so it cannot reach it — and carries no mousemove either way. Every
URL and file path in the terminal was therefore inert on phones and tablets,
Claude Code's own `/login` URL included.
The tap path now activates the link itself, through the SAME provider that feeds
the hover linkifier (`_terminalLinkAtPoint`), so a tap and a desktop click can
never disagree about what is a link or where it ends — containment mirrors
xterm's own `_linkAtPosition`. It runs synchronously inside the touchend handler,
which is what keeps the user gesture that lets `window.open` past the popup
blocker, and before any mouse report, exactly as `_handleDesktopTerminalClick`
already skips the SGR tap for a hovered link.
Two kinds of row keep their existing meaning: the caret's logical line, where a
tap places the cursor and a URL the user typed must stay editable, and TUI-owned
rows, where a numbered choice or an expandable readback is answering a dialog and
routinely carries the very path the tap would otherwise open. The caret line is
the boundary rather than the tap intent, because a plain shell classifies EVERY
tap as 'input' and gating on that would leave every URL in shell output inert.
**Chat.** `marked` emits a bare `<a href>` and the markdown sanitizer's allowlist
carries no `target`, so a tap in the response viewer navigated the current tab
away: on a phone that unloads the whole dashboard — SSE, terminal buffers, unsent
composer text — and there is no middle-click or open-in-new-tab affordance to
work around it. `_renderMarkdown` now decorates anchors in the template pass it
already makes for code blocks. That pass runs AFTER sanitizing, so it is the only
source of both attributes: an agent-authored `target`/`rel` is already stripped,
and `rel="noopener noreferrer"` is set on the same element in the same breath, so
no page Codeman opens gets a `window.opener` handle back. Fragment links stay
in-page; mailto:/tel: are left to the OS rather than stranding an empty tab.
Tests: 10 cases in `terminal-touch-tap.test.ts` (URL, file path, log path,
scrollback, no-double-report, composer, shell mode, dialog row, no provider) and
a new `response-viewer-external-links.test.ts` driving the shipped marked +
DOMPurify + app.js. 7 of them fail without the fix.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
GET /api/sessions/:id/files gains an optional `q`. With one, the endpoint
answers a FLAT match list instead of a nested tree; without one, the response is
exactly what it was, so every existing caller is untouched.
compileFileQuery() (src/utils/file-query.ts) turns the query string into a
reusable predicate, so the walk prunes as it goes rather than streaming the
whole tree to the client to be filtered there. An empty or whitespace-only
query compiles to null, which is what makes "no query" and "blank query" the
same thing.
The search walk deliberately recurses past directories that do not match — a
file whose ancestors don't match is exactly what people are searching for — so
it carries its own maxMatches cap on top of the existing maxFiles and maxDepth
ones, and reports `truncated` when it stops early. Hidden-file and
excluded-directory rules are the same ones tree mode already applies.
Tests: file-query.test.ts covers the matcher; routes/file-search-mode.test.ts
drives the endpoint against a real temp tree and pins the two properties worth
having — that the walk reaches a match under non-matching parents, and that an
absent or whitespace query leaves the tree response alone. Gating the recursion
on a match turns those red.
Shell prompts using Nerd Font glyphs (powerline, p10k/starship folder and
git icons) rendered as missing-glyph boxes: the built-in xterm stack has no
private-use-area symbols, and phones have no Nerd Fonts installed at all.
- Bundle Symbols Nerd Font Mono (icons-only, MIT, 1.2MB woff2) served from
fonts/ and appended to the terminal stack before monospace — browsers fall
back per glyph, so icons render everywhere while text stays in the text
fonts. font-display: block + preload keep tofu out of xterm's glyph atlas.
- New per-device terminalFontFamily setting (App Settings > Terminal &
Input > Font): prepended to the built-in stack, never a replacement, so
the symbols fallback and final monospace always survive. Applied live on
save (refit + echo-overlay refreshFont, mirroring setFontSize).
- Single source for both xterm surfaces: TERMINAL_FONT_DEFAULT_STACK +
resolveTerminalFontFamily() in constants.js, unit-tested.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VjnbbZRBuvR5E3SDouwXr9
Resolves the four advisories that reach the production dependency tree. The
other 16 npm audit reports are devDependencies-only (Remotion, Puppeteer,
postcss, the eslint/tsx toolchain) and never ship to users.
- @fastify/static 9.1.3 -> 10.1.3 GHSA-8pvw-jcv7-9cmj (authz bypass via
non-canonical URL paths). Covers <=10.1.1, so all of 9.x is affected and
the fix exists only on the 10.x line.
- find-my-way 9.6.0 -> 9.8.0 GHSA-c96f-x56v-gq3h (HTTP/2 DDoS)
- fast-uri 3.1.2 -> 3.1.5 GHSA-v2hh-gcrm-f6hx (host confusion)
- brace-expansion -> 5.0.9/1.1.18 GHSA-3jxr-9vmj-r5cp (expansion DoS)
The last three are transitive and needed only a lockfile re-resolve, so no
overrides were introduced.
The @fastify/static major changes setHeaders' first argument from a Node
ServerResponse to a FastifyReply. Two consequences:
1. res.setHeader() -> reply.header(). The v9 body throws TypeError from
inside the plugin on every static request.
2. Precedence flips, silently. The callback used to write to the raw
response and lose to the route's staged reply headers; it now writes to
the reply and wins. That gave /sw.js a year of immutable in place of the
no-cache, no-store its route sets, pinning a service worker on every
client with no server-side recovery. A route that already set
Cache-Control now keeps it.
Verified against v9 to confirm the sw.js behaviour is a regression and not
a pre-existing bug.
ws appears in npm audit but production is on 8.21.0, outside the vulnerable
range; the only affected copy is bundled under @remotion/renderer (dev-only,
and remotion is pinned at 4.0.473 because the compositor refuses to start on
a version mismatch).
Adds test/static-cache-headers.test.ts, which drives a real server and covers
a caching contract that had no test at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`npm test` ran config/vitest.config.ts, which includes the browser, visual and
perf suites. On any machine without chromium, a free port and per-machine PNG
baselines that fails ~87 tests on a clean master, so the repo's most obvious
command could not be used as a pass/fail signal. The workaround had spread into
four docs as "never run bare `npm test`" warnings.
`npm test` now runs config/vitest.ci.config.ts — byte-for-byte what CI runs — so
local green means CI green. Verified: 264 files, 5248 tests, exit 0.
The suites it leaves out are not abandoned; each has a command:
test:browser 5 Playwright files (chromium + a live server; codex-predictive-echo
also needs a real codex binary)
test:mobile unchanged — the above plus per-machine PNG baselines
test:perf 2 wall-clock benchmarks; need an otherwise idle machine
test:all the old everything-behaviour, kept reachable
test:ci is untouched (CI still calls it). test:watch and test:coverage follow
test onto the gate's config.
The more important half is the hole this closes. The exclusion list lived as
literals in one config and pointed one way only: a file excluded from CI and
added to no runner would be tested by NOTHING, silently, with every command
still green — vitest counts "no files matched a filter" as success. That is the
same shape as the #279/#280 blind spot already documented in CLAUDE.md.
So the globs moved to config/test-suites.ts, one array per REASON a suite cannot
run in CI, and all three configs derive from it. test/test-suite-partition.test.ts
then checks the arithmetic against the files on disk: it fails if any test file
is reachable by no runner, or by two. Confirmed it fires by orphaning a file and
watching it name it. The partition is exact today:
gate 264 + browser 5 + perf 2 + mobile 9 = 280 = every *.test.ts in the repo
⚠️ One sharp edge, deliberate and documented: a file filter must match its
runner. `npm test -- test/mobile/keyboard.test.ts` now matches nothing and exits
GREEN having run zero tests, because the gate's config excludes that path.
CLAUDE.md recommended exactly that command in the on-screen-keyboard note; that
line now says `npm run test:mobile -- <file>`, and the Testing section calls out
the trap, since a green run of zero tests is worse than a red one.
Docs synced: CLAUDE.md, AGENTS.md, .github/CONTRIBUTING.md, README.md,
README.zh-CN.md, and two ci.yml comments that claimed only test/mobile/** was
excluded — it is three suites, and 5 Playwright files rather than 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Each cycle step (kickstart, update, /clear, /init) checks for `stopped` before
`await session.writeViaMux(...)`, then emits `stepSent` and calls
`setState('waiting_*')` after it.
stop() is asynchronous with respect to that await. One that lands while the
write is in flight has already passed the guard that ran, so the post-await
setState() puts a stopped controller back into a waiting state — re-arming its
step timers against a session the user asked to stop.
Re-check after the await, before emitting and setting state.
The guard reads the public `state` getter rather than `_state` on purpose:
TypeScript narrows `_state` across the await from the pre-await check and cannot
see that stop() mutated it, so `this._state === 'stopped'` is rejected as a
comparison with no overlap (TS2367) at all four sites.
Adds test/respawn-stop-race.test.ts, which drives the interleaving
deterministically by calling stop() from inside the mocked write rather than
relying on timing. All four steps go red without these guards.
validateSessionFilePath realpath-resolves the candidate path but compared it
against the raw sessionWorkingDir. When the workspace is itself reached through
a symlink the two sides live in different namespaces, so relative() reports a
spurious `../` and every file in that workspace is judged an escape — reads and
writes in the session are refused wholesale.
That is not an exotic setup: os.tmpdir() hands back a symlinked path on macOS
(/tmp -> /private/tmp), and symlinked project directories and bind-mounted case
paths hit it too.
Resolve both sides and compare canonical to canonical. This only makes the
comparison honest — it does not widen it. The candidate keeps its own realpath,
so a symlink pointing out of the workspace and a ../ traversal are still
refused, and a workspace that cannot be resolved now fails closed.
Three stubs in file-routes.test.ts used a blanket
realpathSync.mockReturnValue(escapeTarget), which answers the same path for the
workspace and the candidate; with both sides resolved that makes an escape look
contained. They now use the input-aware mockImplementation idiom the rest of
that file already uses, so the workspace resolves to itself and only the
candidate escapes. Verified they still bite: removing the confinement check
turns all of them red.
Adds test/route-helpers-symlink-confinement.test.ts, which exercises the
function against a real symlinked workspace on disk and pins the negative cases
(../ escape, symlink-out, missing file) alongside the fix.
Session List Layout gains a third option. The old "Left sidebar" becomes
"Left sidebar simple" and is unchanged down to the byte; the new "Left sidebar"
puts on each row what the desktop home rail and the phone overview already show:
when the session was first created, how long it has been in the state it is in,
and a status pill naming that state.
A docked column is not a tab strip. It has width to spare and a row per session
either way, and "name + folder" is the whole story a TAB can tell, not the whole
story there is. This is the information that was missing, and it already existed
one surface over.
Both sidebar values are the same layout, and both set data-session-list="sidebar";
the row detail rides on a separate data-sidebar-detail attribute. That split is
the load-bearing decision here: every one of the ~25 isSessionSidebarActive()
call sites and every html[data-session-list="sidebar"] rule in styles.css and
mobile.css keeps matching both variants without being touched. A third
data-session-list value would have meant auditing and editing all of them.
- Stored values: 'header', 'sidebar' (simple), 'sidebar-rich'. Anyone already on
'sidebar' keeps exactly the layout they picked — the rename is label-only.
- State classification and the "how long has it been like this" anchor come from
_mobileOverviewState() / _mobileOverviewSince(), not re-derived, so the three
surfaces cannot disagree about what "working" means. A working pane repaints
~1/s, so its duration is measured from the turn's last Enter: a running turn
reads "working 12m", not "0m".
- Stamps refresh in place on a 20s clock rather than by re-rendering — a rebuild
would restart every load spinner and alert animation in the list, twice a
minute. The clock runs only while rich rows are on screen, and is stopped from
both render paths and from applySessionListLayout().
- The incremental render path updates the pill, the accent class and the since
anchor; a tick alone cannot see a state change, and a new turn re-stamps
lastSubmitAt without changing state.
- applySessionListLayout() now re-renders on a DETAIL change too. simple <-> rich
leaves data-session-list on 'sidebar' both times, and the meta line is emitted
by the row template rather than toggled by CSS, so the old layout-only test
would have flipped the setting and repainted nothing.
- Width: 300px for the extra line. The collapsed 44px rail and the handheld
drawer are both explicitly held back from it — the desktop rule is (0,3,1) and
would otherwise out-specify mobile.css's (0,2,1) drawer base and pin a 320px
phone's drawer to 300px.
- Missing/stale mobile-overview.js degrades to a row with no meta line rather
than throwing and taking the whole tab strip down.
15 new tests cover the attribute split, the solo-window override, the
detail-change re-render, the row model, both render paths, the clock lifecycle
and the mobile width guard.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The frontend load order omitted session-lineage.js (29 modules listed, 30
loaded), and several inventory counts had drifted from the tree: route handlers
~200 to ~217 with system, files and approvals each understated, src/config 20 to
21 files, install.sh 69KB to 92KB, and the Prettier exemption list, which also
never mentioned mobile.css. Two of the missing handlers are endpoints CLAUDE.md
already documents in prose but never counted.
postcss is imported by two tests but was only present transitively via vite, so
knip reported it as an unlisted dependency. Declared at the version already
resolved in the lockfile.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Lineage arcs were coloured per child, so one tab's own workers each got a
different colour, which is the distinction the colours exist to make. The colour
is now keyed on the parent: every arc leaving one tab is the same colour however
many workers it spawns, so the strip reads as "these five came from w1, those
two came from w2". A child that spawns in turn is a parent in its own right and
gets its own colour for the arcs below it, so a chain changes colour at each
generation while each generation's fan-out stays uniform.
The new tests drive the real _appendLineageConnectionLines() and assert the
painted custom property, because testing the colour function alone passes just
as happily with the child id passed back in.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The active tab is the only one that grows a gear and a close button, and with a
short session name they were eating it: "w1" rendered a 13px label while gear
plus close took 50px of a 116px tab, so the tab's geometric centre landed on the
gear and a thumb aiming at the middle of the tab opened Session Options instead
of switching sessions. Reserving a minimum label width on the active tab widens
the tab by the difference instead.
The floor is set by the 10th tab onward, which renders no number badge and so
sits 10px further right; a numbered tab clears the icons at 20px but a
numberless one needs 40px. The test recomputes that inequality from the
stylesheet rather than pinning the pixel.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
30 pages covering install, concepts, the dashboard, the agent CLIs, unattended
runs, remote and Docker cases, security and the HTTP API, plus a sidebar and a
footer. The wiki repo has no CI and no review, so docs/wiki is the source of
truth and .github/workflows/wiki-sync.yml mirrors it on every push to master.
The workflow refuses to mirror when docs/wiki is missing or holds no pages,
because it deletes before it copies and would otherwise publish the deletion of
every page. The footer carries a {{VERSION}} placeholder stamped at publish
time rather than a hand-written version, which went stale on every release.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Give the active-session handoff one owner: closeSession captures wasActive before its await and the session_deleted handler stands down for a close this tab started.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Gate the idle-alert acknowledgement to human selections: the boot restore, a solo window opening its target, and the post-close fallback no longer spend a yellow tab alert.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Red tab alerts track the dialog, not the keyboard: typing no longer clears them, and a dialog answered in the terminal resolves itself on the next listing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Persist the 'I checked it' state of yellow idle tab alerts across reloads and devices.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five post-merge review items from PRs #306 (clickable file paths) and
#307 (session sidebar):
- constants.js FILE_PREVIEW_EXTENSIONS gains the media extensions it was
missing vs the single-source sets in attachment-registry.ts (m4v ogv
ogg oga m4a aac flac opus), so an in-workspace .m4a opens the preview
player instead of the log viewer; new test/media-extension-parity.test.ts
pins all three copies (constants.js, panels-ui.js, attachment-registry.ts)
against each other.
- FILE_PATH_LINK_PATTERN drops `etc` from its root alternation: /etc is
unconditionally in DEFAULT_BLOCKED_TREES, so every /etc link 403'd.
Negative cases added to the link-provider and response-viewer tests.
- updateSidebarCount() counts the rows actually on the sidebar list
(session rows + web-tab rows, minus filtered-out ones) instead of
this.sessions.size, and applySidebarFilter() refreshes it so the count
follows the filter box per keystroke.
- The incremental-render connection-line gate now also fires in sidebar
layout (this._lineageEdgeCount is permanently 0 there), matching the
strip-scroll listener widened in #307, so a badge changing row heights
redraws subagent/ultracode connectors.
- isSensitivePath() blocks ~/.claude.json, ~/.claude/settings.json and
~/.claude/settings.local.json (credential-bearing by schema), anchored
to homedir() read at check time so case-level .claude/settings*.json
files stay servable in the File Viewer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Root cause of the reviewer's mass-bump measurement (17 of 17 sessions with an
identical lastActivityAt): every restart restamps all sessions in the
constructor loop, and the boot auto-attach's repaint re-bumps the rest within
the same second. A 12-minute steady-state sample shows NO ambient mass bump,
so restarts are the whole story, and Codeman restarts on every deploy.
The stamp now has a display twin: recovery threads the previous run's
lastActivityAt from state.json into the wire-visible stamp (getter + toState),
and a 15s settle window keeps the attach repaint from overwriting it. Real
actions (input, task assignment, respawn) always write through. The private
stamp keeps its boot-anchored semantics untouched, because the idle
confirmation reads it as how long the pane has been quiet.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Post-#304 follow-ups. The install-vs-refresh decision (workspaceHooksEnabled,
default ON) moved from a session-routes-local helper into hooks-config.ts as
applyWorkspaceHooks(workspace, install?), and the claude session-create sites
that bypassed it now go through it: cron job fires (cron-service), legacy
scheduled-run iterations (runScheduledLoop), and the plan-orchestrator research
and planner one-shots. A cron or scheduled run firing in a linked case that
never had an interactive session ran hook-blind (no stop for completion
detection, no tab alert on a blocking dialog).
The shared core also carries the two guards every caller needs: a workspace
that no longer exists is skipped (ensureCodemanHooks mkdir -p's, so the boot
recovery sweep used to resurrect a deleted repo as an empty tree holding only
.claude/settings.local.json), and all errors are swallowed since a create must
never fail on hooks. Route handlers keep resolving the setting through their
ConfigPort and pass it in; non-route callers omit it and the core reads
settings.json itself (absent key or unreadable file = ON).
Two adjacent gates tightened in session-routes:
- the docker quick-start hooks branch excluded the five external CLIs but let
`shell` through, contradicting its own rule that only claude reads .claude
hooks; it is now gated on mode === 'claude'
- the statusLine exporter call in POST /api/sessions got the same
!remote && body.workingDir guard the hooks call got in 499d355 (it mkdirs the
same way, so a remote attach created a junk user@host:session dir locally and
a cwd-fallback create wrote into $HOME)
plan-routes' one-shot deliberately stays out: its workingDir is process.cwd(),
exactly the target 499d355 forbids writing into. restoreMuxSessions stays out
too: the boot sweep already covers recovered workspaces.
Tests: quick-start existing-case install, docker claude-installs/shell-does-not,
and the core directly (default-ON install, OFF add-nothing, OFF still heals a
stale block, malformed file untouched, vanished workspace skipped); the remote
and cwd-fallback regressions now also send statusLineTelemetry:true to pin the
statusLine guard.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three follow-ups from the 1.19.0 review of the activity-ordered home screens:
- Hook events now ride the same debounced session state broadcast the
working/idle handlers use. The blocked group ranks on lastActivityAt, and
without this a permission prompt raised after page load kept ranking by
whatever stamp the browser loaded with.
- A working row with no submit stamp now shows the lastActivityAt fallback
its sort anchor already uses: a row must never be ranked by a number it
does not display.
- Alt+digit resolves through the live-session projection the render paints
(sessionOrder minus dead ids), so a stale id cannot shift every painted
number off its target, web tabs included. New tests pin both surfaces to
one shared order and the numbering to the live projection.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Widening the servable extensions to EDITABLE_EXTENSIONS made ~/.codeman
JSON previewable for the first time, and the blocklist named only
state.json. But settings.json holds a credential BY SCHEMA
(voiceSettings.apiKey), push-keys.json holds the VAPID PRIVATE key, and
intents.json is written 0600 precisely because captured prompts can carry
secrets — all three were one authenticated click away once an agent
printed the path. Blocked alongside state.json, whose rule now also
catches state-* siblings.
The never-re-cuts-inside-an-anchor test used an unmatchable URL tail, so
it passed with the guard deleted; the fixture now carries a matchable
/tmp path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The running group sorts on lastSubmitAt, but nothing pushed a session:updated
when a turn STARTS — the browser kept whatever stamp it loaded with, so a
30-second-old turn could rank (and read) as an hour-long one. The working
handler now rides the same debounced state broadcast idle already uses.
And both call sites of window.CodemanSessionOrder now degrade to tab order
when the global is missing (iOS Safari's documented stale-cached-JS after a
deploy) instead of TypeErroring the whole home screen away.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
endpoints.md named 1.18.x as the version where workspace hooks became a
setting, but 1.18.x servers do NOT have this behavior — an agent driving
one would falsely conclude its workspace has hooks. The feature ships in
1.19.0. And preamble.sh changed content this PR without bumping its
CODEMAN_PREAMBLE stamp, so a cache stamped 1.18.3 would pass the
staleness check while holding old bytes; stamp bumped to 1.19.0 in
preamble.sh and the SKILL.md heredoc together (byte-identity pin).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A claude-mode attachRemoteSession create overwrites workingDir with the
user@host:session pseudo-path, which is a RELATIVE path locally — the old
refresh-only call no-op'd on it, but ensureCodemanHooks mkdirs, so it
created a junk local directory. And with workingDir omitted the cwd
fallback reaches the hooks write unvalidated; under installer-created
services cwd is $HOME, so hooks materialized in ~/.claude/settings.local.json.
Both guarded at the applyWorkspaceHooks call site; regression tests prove
the remote attach leaves no junk dir and the no-workingDir create leaves
the server cwd untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Brings in https://github.com/christianhaberl/Codeman/pull/4 (three commits,
authorship preserved) and adapts it across the 211 commits master gained
since the branch was cut:
- App Settings control re-authored for the set-* surface (PR #278): a
set-row in Layout -> Tabs, replacing the old settings-item markup the
branch targeted. i18n description synced.
- Lineage arcs (PR #291, post-branch) are SKIPPED in sidebar layout:
computeLineagePath()'s U-bridge geometry hangs from the horizontal
strip's bottom edge and has no meaning against a vertical list. The
lineage strip-scroll listener now also redraws subagent/ultracode
connectors while the sidebar scrolls vertically.
- The desktop home tab rail (post-branch) defers to the sidebar: both dock
the session list flush left, and the rail would render z-ordered under it.
- Active-row reveal unified into _scrollActiveTabIntoView() (#257 landed on
master after the branch): sidebar mode branches to scrollIntoView
block:'nearest', and _fullRenderSessionTabs() restores scrollTop alongside
the #257 scrollLeft restore so ambient rebuilds cannot yank a mid-scroll
sidebar back to the top.
- Mobile active-tab hoisting the branch guarded against no longer exists on
master (removed by #257); kept master's order-stable render.
Verified: typecheck, lint, format:check, check:frontend-syntax,
check:public-assets, PostCSS parse of both merged stylesheets, the 26 new
jsdom tests, the structural guard suites, and the headless-Chromium harness
(scripts/verify-session-sidebar.mts) green across all seven layout states
at 1600/1000/393px against current master.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A .json/.log/.yaml/code path outside the session workspace was refused as an
unsupported type, and clicking one in the terminal made it worse: text goes to
the log viewer, which spawns `tail -f` and allows only the workspace, /var/log
and ~/logs, so it answered "Path must be within working directory or allowed
log directories" while the same path clicked in the response viewer previewed
fine. Two surfaces, two answers, for a file the session can already cat.
- TEXT_ATTACHMENT_EXTENSIONS IS EDITABLE_EXTENSIONS (config/file-editing.ts),
not a second curated list that would drift from it. The rule reads: if the
viewer would open a file for editing inside the workspace, the same file
outside it can be read. The suffix was never the confidentiality gate here,
the path guard is (sensitive-file blocklist, /root and /etc trees, realpath
before the check), and it still runs on every registration.
- Widening what can be READ must not widen what can RUN. html/htm join svg in
serveRawFile's download-only branch, so markup is never served with a
renderable type on our own origin; other text goes out as inert
text/plain; charset=utf-8 with nosniff, matching what the path picker does.
The preview reads through fetch(), which ignores the disposition, so a
clicked .html still shows its source.
- ~/.codeman*/state.json joins isSensitivePath. It persists
SessionState.envOverrides and the env allowlist admits key-shaped names
(GEMINI_API_KEY, CLAUDE_CODE_*), so it can hold a live credential. Same
treatment as hook-secret and users.json, and the rest of the tree stays
attachable.
- The terminal sends an out-of-workspace path to the preview instead of the log
viewer. In-workspace text keeps the tail viewer, which is the point of it, and
file-stream-manager's allowlist is untouched: no `tail -f` on arbitrary host
paths.
- The by-id text preview is bounded like the workspace one: a Range request for
the first 512KB (a real partial read, not a discarded 50MB download) plus a
500-line cap, with the footer saying so.
Verified on an isolated instance: a 1.1MB external log opens in ~1.8s showing
500 lines with "showing first 500 lines" in the footer; json, yaml and code
preview; an .html carrying a script tag renders as source and does not execute;
.svg is still refused; a terminal click on an external .yaml opens the preview
with no log viewer and no attachment card; an in-workspace .log still opens the
streaming tail viewer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A clip an agent wrote inside the workspace played with a working scrub bar,
while the same file in /tmp was refused as an unsupported type. The workspace
preview classified media with its own inline extension sets and the attachment
allowlist had no media at all, so the two paths disagreed about what a video is.
- VIDEO_ATTACHMENT_EXTENSIONS and AUDIO_ATTACHMENT_EXTENSIONS now live in
attachment-registry.ts and are imported by file-content's classification, so
both paths answer the same. mp4/webm/mov/m4v/ogv and
mp3/wav/ogg/oga/m4a/aac/flac/opus join the attachment allowlist.
- Real MIME types for those extensions. Without one the raw route falls back to
application/octet-stream, which a <video> refuses to decode: the player
renders and then does nothing.
- getAttachmentType() gained the video and audio members of
AttachmentDetectedType. Attachment cards have no per-type CSS and their
thumbnail falls back to the type label, since the thumbnailer has no media
branch and answers 204 rather than spawning a converter.
- The preview overlay's by-id branch renders <video>/<audio> with the same
markup as the workspace branch, playsinline included. Serving was already
range-aware, so seeking works.
The image-watcher keeps its own narrow detection list (png/pdf/docx/pptx), so
this does not start popping cards for every video an agent writes. Text types
that are not md or txt (.json, .log, code files) remain out of the allowlist by
choice and still report what is previewable instead.
Verified on an isolated instance: an external mp4 and mp3 both play, seek, and
report the right duration, matching the in-workspace clip exactly, and a click
on an external mp4 in the terminal opens the player with no attachment card.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A path an agent prints was already underlined in the terminal, but clicking
one opened the preview overlay on "File not found": file-content/file-raw
resolve against the session workingDir and refuse anything outside it, and the
paths agents print most (a /tmp capture, Claude's own scratchpad, another
checkout) are outside it by definition. In the response viewer those paths were
not links at all.
- openFilePreview() detects an out-of-workspace path and registers it through
POST /api/sessions/:id/attachments first, rendering by attachment id. That is
the surface built for live external files, so the server-side guard is
unchanged: secret trees blocked, symlinks resolved, extension allowlist. The
workspace routes keep refusing escapes exactly as before.
- New optional `notify` field on that route. `notify: false` suppresses only the
attachment:detected broadcast, so a click does not also pop a card announcing
the file already filling the screen. Default stays true for the CLI and
publish callers.
- _linkifyFilePaths() links paths in rendered response-viewer markdown. It walks
text nodes and builds anchors with DOM APIs (the source is model output; never
a string rebuild of sanitized markup), skips subtrees already inside an <a>,
and keeps the message text byte-identical so copy-code is unaffected.
- One path pattern in constants.js now feeds both the xterm link provider and
the chat linkifier, a fresh instance per call since lastIndex is per-object
state. It picks up /Users and /mnt roots (nothing was clickable on macOS or
WSL), plus docx/pptx and video/audio extensions.
- .file-preview-overlay moves to z-index 5100, above the response viewer at
5000. At its old 2000 a path clicked in the chat opened the overlay behind the
panel it was launched from.
Verified end to end on an isolated instance, desktop and phone viewport: real
clicks in the terminal and the chat both render the image, external md and pdf
render, /etc/hosts is still refused, workspace previews unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The workspace-hooks install makes the skill's central hooks rule wrong in the
cautious direction. Six places told a worker that a linked case or a raw
workingDir has no `stop`/`blocked` and that send-and-wait cannot be trusted
there, so an agent would hand-roll output-marker synchronization in exactly the
workspaces where `wait:true` now works.
Rewritten against the setting rather than directory provenance:
- verbs.md §5.1: the where-to-spawn table, the rule paragraph (now naming
`workspaceHooksEnabled`, default ON, the add-only merge, and the boot sweep of
recovered sessions), and the silent-failure warning. The three cases that stay
hook-less regardless are called out: remote SSH sessions, docker cases that
opted out, and a workspace Codeman cannot write to.
- verbs.md §5.3: the send-and-wait precondition is "the workspace has the hooks
block", not "a case Codeman created".
- endpoints.md: the Signals-by-mode table is now keyed on the setting, with rows
for OFF, for remote/docker-opt-out, and for a session from an older server.
The old create-path grep list becomes a "before 1.18.x" note.
- SKILL.md §2 + the cost list, recipes.md Flow-1 contrast, messaging.md step 1.
"Check, do not assume" is kept and promoted to the load-bearing habit, because
the setting is not visible from the call and a session created by an older server
that has not restarted still has nothing.
The `spawn_worker` hooks grep STAYS: it guards the setting being off, remote
sessions, and older servers. Only its diagnostic changes, since "pick an unused
name" is no longer the fix. That text lives in both the §0 heredoc and
`preamble.sh`, which `test/agent-skill.test.ts` pins byte-identical, so both are
patched with the same bytes.
Docs only, no behavior change. 23 skill tests green, full test:ci 5109 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Installing hooks into any workspace a Claude session runs in is the right
default, but it takes a decision away from a user who deliberately removed
them: nothing on disk distinguishes "removed on purpose" from "never had any",
so they would come back on the next session create.
Adds the synced workspaceHooksEnabled setting (App Settings -> Agents & CLIs ->
Claude), default ON. OFF restores the older behavior exactly: a Codeman hooks
block that is already present is still refreshed when stale (COD-91), but one
is never added.
Every create path routes through one applyWorkspaceHooks() helper so the gate
cannot apply to some paths only, and the boot-time recovery sweep honours it too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The phone overview and the desktop tab rail list the same sessions, so
they now share one order (CodemanSessionOrder in constants.js, pure and
unit-tested): blocked on you first (longest-blocked at the top), then
running longest-turn-first, then quiet most-recently-quiet first.
The tiebreak flips direction halfway down on purpose: for a state a
session is still in, longer is more urgent; for a state it has stopped
in, more recent is more relevant. The running group keys off the pane's
last Enter (lastSubmitAt), never lastActivityAt, because a working pane
repaints about once a second and would rank every turn as freshly
started. A 0 stamp means "unknown" and sorts last within its state.
The desktop rail was previously in raw tab order. Its number badge stays
the Alt+1..9 index, so on a sorted rail it deliberately no longer runs
1,2,3 downward: it names a shortcut, not a row position. Its second
stamp changes from "active 3m ago" to the state duration the order is
computed from ("created 1d ago . working 40m"), since both working rows
otherwise read "active just now" and the order looked arbitrary.
The tab strip itself is untouched: still user-ordered and drag-sortable.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A session in a linked case (or any pre-existing repo) ran with no hooks block
at all: writeHooksConfig only fires when Codeman CREATES the case directory,
and refreshStaleCodemanHooks deliberately never adds one. Every hook-driven
surface was therefore dead in exactly the place most sessions run: no tab
alert or phone-overview NEEDS YOU row when a dialog blocks the pane, no
Approvals Inbox item, no push, no definitive stop/idle_prompt for respawn,
and no stop/blocked for the agent wait endpoints.
Both session-create paths and restoreMuxSessions() now call
ensureCodemanHooks(), an add-only merge that keeps a user's own handlers and
leaves a malformed settings file untouched. Claude Code re-reads
settings.local.json, so a session already running in the workspace starts
firing hooks without a restart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Captured live from an isolated instance running this branch: a regular
active tab beside a yellow waiting-for-input tab and a red needs-decision
tab. The gif covers one full 17.5s loop (LCM of the 2.5s red and 3.5s
yellow pulse cycles), so it loops cleanly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- SKILL.md: forbid the standalone preamble check and pre-spawn recon turns
(measured: two wasted model turns cost ~12s of a 28s two-worker run; the
hardened flow measured 20.2s cold / 12.8s warm end to end)
- Lineage lines: dip now hangs from the strip's bottom edge (cap 104 -> 64,
no stacked row offsets), fixing the deep bow in wrapped strips and keeping
row-1 arcs off row-2 tab labels; per-child color palette (skin blue first,
then matrix green, pink, violet, red, turquoise, orange) via an inline
--lineage-color custom property
- Session Options -> Session: per-TAB pop-out (open-in-window) button override
on top of the general showTabDetachButton setting; per-device localStorage
map rendered as the tab-show-detach class
- Tab alerts: seed the pending-hook state machine from GET /api/approvals
regardless of the approvals-inbox setting (reloads used to lose the red tab
entirely with the inbox off), clear unconditionally on approval_resolved,
and repaint the alert as a steady red/yellow ring + glow + status dot on a
::before overlay so it stays visible on the selected (active) tab until the
permission is actually resolved
- docs: worker warm-pool design sketch (verified numbers baked in)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two live failures from one root cause: Claude Code loads a same-named
user-level skill (~/.claude/skills/codeman, written once by `codeman skill
install`) over the fresh per-case copy, and nothing ever refreshed it. A
stale Aug-9 copy (pre fast-path, pre lineage header) made every agent-driven
spawn run the old recipes: workers spawned serially with pid polls and
without X-Codeman-Parent-Session, so the web UI drew no lineage arcs.
- refreshUserAgentSkill(): session create now refreshes a marker-owned
user-level copy (refresh-only: absent copies are not installed,
foreign/symlink copies stay untouched).
- seedAgentSessionPreamble(): local claude session create pre-seeds the
skill's preamble into ${XDG_CACHE_HOME:-~/.cache}/codeman-agent-<id>.sh,
single-sourced from the new skills/codeman/preamble.sh, so the skill's §0
bootstrap collapses to a two-line loader instead of a ~150-line paste the
model has to type out (measured ~47s of generation per run).
- SKILL.md: §0 now leads with the loader and keeps the full block as the
stale/missing fallback; explicit verbatim-paste warning (a hand-assembled
preamble is how the header and the fast-path functions got lost);
spawn_worker also sends parentSessionId in the body as defense in depth;
preamble stamp bumped to 1.18.3 so pre-fix cached preambles self-heal.
- test/agent-skill.test.ts pins preamble.sh byte-identical to the SKILL.md
heredoc and covers seeding (XDG + HOME fallback, 0600) and the user-level
refresh (absent/stale/foreign).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The changeset ships this as 1.18.2, so the stamp, the bootstrap's grep/write
condition, both re-source guards and the recipes guard all carry 1.18.2 now
instead of a version that would never exist.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fifteen review findings on the fast-path rewrite plus one caught live, all
verified against a real 1.18.1 server before landing:
- sendwait picks a fresh seq (the epoch second) instead of a fixed 2, so a
second prompt to the same worker is typed instead of silently swallowed as
an already-applied duplicate; explicit seq remains for deliberate resends
- sendwait self-heals stranded delivery: an Ink repaint occasionally eats the
Enter (observed live), so a timed-out short first wait sends one bare \r and
re-waits by resending the identical frame as a tagged duplicate
- spawn_worker verifies the resolved casePath carries Codeman hooks (the same
/api/hook-event marker the server checks), refusing names that resolve to
linked or pre-existing hook-less directories instead of running the job in
what may be the user's real repo
- spawn_worker probes the trust dialog after a short 5s composer wait, not the
full 45s, restoring the ladder staging verbs.md documents; on a readiness
miss it deletes the half-spawned session and returns 1 with empty stdout,
so a prompt can never be typed blind into a trust dialog
- spawn_workers refuses duplicate case names and empty argument lists, and
keys result files by index
- section 1 is bash 3.2 compatible (indexed arrays, no declare -A), prints the
full delivered/timedOut/signal tuple per worker with an explicit line for a
missing result, deletes only workers whose turn really ended (a timeout
means still working), cleans up spawned siblings when any spawn fails, and
guards its mktemp
- last_text takes the previous answer as an optional second argument for
consecutive-turn reads (the transcript briefly serves the prior answer
after a stop, observed live)
- the stale duplicate bullets in section 1's closing list are gone
- reference/verbs.md joins the mode-list drift guard's file list
- README's skill inventory covers verbs.md and the new SKILL.md shape
- the changeset is minor so the shipped release matches the 1.19.0 stamp
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured against a live 1.18.1 server, the API does the whole job in about ten
seconds: two cold claude workers spawned and ready in 6.3s, both tasked and both
answers read in 4.0s more. The slowness users reported was agent-side.
Three causes, all of them things the skill taught:
- It taught serial spawning. Nothing in the main document showed `&`/`wait`, so
"spawn two workers" read as "do the readiness ladder twice", which is one model
turn per worker.
- It had no spawn primitive. The happy path had to be reassembled on every run from
where-to-spawn, a four-stage readiness ladder, send-and-wait, the fan-out caveats
and a recipe with two variants. Each is a decision, and most carry a warning.
- It cost ~16k tokens before the first call, at 3.6:1 prose to code, with 25 warning
glyphs and 55 occurrences of "never". A document that is mostly failure modes
teaches caution, and caution bills as thinking tokens.
The preamble now defines the verbs rather than describing them: spawn_worker,
spawn_workers (concurrent), sendwait, last_text. Section 1 composes them into the
whole job in one Bash call and says to stop reading there.
Two ceremonies the measurements retired: the pid poll (one iteration, 33ms, and
wait-output already blocks on the composer) and reading settings.local.json to check
hooks for a case quick-start creates, which always has them. That check stays
required for linked cases and raw paths, where its absence silently breaks
send-and-wait.
The bootstrap's write condition now greps the version stamp, so a stale or truncated
preamble self-heals rather than failing and asking for a manual rm. The stamp line is
kept bare because the grep anchors on it with $; an inline comment there would rewrite
the file on every bootstrap.
Section 5 moved to reference/verbs.md behind an index, cutting the always-paid
SKILL.md from ~16.4k to ~7.6k tokens. Section numbers and anchor slugs are unchanged,
so existing references still resolve; all 201 anchors across the five files were
checked, with the checker positive-controlled against an injected bad link.
Verified by extracting the code blocks from the shipped file and running them against
the live server: bootstrap plus full fast path, two workers resolving on the
definitive stop signal, answers read and sessions deleted, in 6.8s.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to #285. Violet sits close to the terminal's own dim foreground,
so the arcs lost contrast exactly where they cross text, which is most of
their length. Blue reads at a glance on the dark skins and on the light
ones.
Colour still comes from a token every skin block already defines and tunes
for its own background (--session-blue instead of --session-purple), so it
stays one rule for all seven skins with no per-skin override, and the two
blues are not even the same: --session-blue is per palette while the
subagent rule hardcodes #3b82f6.
Hue no longer separates this layer from the subagent lines, so the
separation now rests entirely on shape (a lineage arc hangs under the strip
and never reaches a window), weight and dash pattern. Noted in the rule.
CSS only: no geometry, no markup, no settings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The refresh can now issue two fetches (full history, then the tail as a
downgrade fallback), which widens an existing window where the user switches
tabs mid-flight and this session's history gets painted into the terminal they
are now looking at. Guard it the way _maybeRefetchFullHistory already does.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two further instances of the same root cause, both in _onSessionNeedsRefresh,
which is SERVER-triggered (it fires after SSE backpressure clears) so the user
has no gesture to blame the result on.
1. It ended in an unconditional scrollToBottom, so a user quietly reading
scrollback was dropped to the live output by a background event. It now
holds their place. The rewrite REPLACES the buffer, so an absolute viewportY
captured beforehand is meaningless afterwards; distance from the bottom is
the anchor that survives, via computeRewriteScrollLine().
2. It rebuilt the terminal from a 1MB TAIL. Measured end to end on a 900-line
shell pane: an 869-row buffer came back as 158 rows, so the refresh meant to
REPAIR the display was destroying most of the scrollback every time it ran.
It now asks for full history, and falls back to the tail only when
_replayWouldShrinkBuffer refuses the capture, which keeps repaint-mode panes
(tmux holds roughly one frame for them) exactly as they were.
Also records truncation state here, so the #258 banner stops describing the
pre-refresh buffer.
Verified in a real browser against a live session: baseY 869 -> 869 where it
used to be 869 -> 158, a reader 200 lines up stays 200 lines up, and a follower
stays pinned to the bottom.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The lines that join a tab to the workers its codeman skill spawned were
drawn with numbers tuned against two tabs sitting side by side, and they
degraded in exactly the two situations the feature is actually used in.
1. A spawned worker is appended to the END of the strip, so the real span
between a lead and its worker is 800-1500px. With the dip clamped at
44px that is a 33px sag: the arc reads as a straight line drawn across
the terminal instead of a bracket hanging under the strip. The dip now
grows at 0.085/px and clamps at 104.
2. When the desktop strip wraps (tabs-two-rows / tabs-auto-wrap), a parent
on row 1 and its child on row 2 are ~14px apart, and the cross-row
branch drew parent-bottom to child-TOP: a flat line hidden inside the
row gap, with siblings overprinting each other. Both ends now anchor on
the tab BOTTOM with the control points below the LOWER row, so a wrapped
pair gets the same bracket a flat strip gets. That deletes the branch:
one shape covers both.
Visibility, at 1:1 rather than in a zoomed mockup: 2 -> 2.5px stroke,
4 4 -> 5 5 dashes (lineage-flow moves with them, -16 -> -20), opacity
.55 -> .72, and a second wider glow so the contrast comes from the halo
rather than from more weight, keeping the line under the subagent lines'
3px. A working child is bright (.95) outside the reduced-motion block, so
turning motion off no longer also dims every worker's arc. Sibling nesting
6 -> 8px and the direction dot 3 -> 3.5px to match the heavier stroke.
Verified at 1:1 in a harness driving the real styles.css and the real
computeLineagePath over three layouts (adjacent workers, workers at the
far end of a full strip, wrapped two-row strip) on a dark and a light
skin. test/session-lineage-lines.test.ts pins both regressions.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two bugs in the File Viewer's media player, both reproduced in a real
browser against an 18MB mp4 before and after the fix.
1. Closing the preview left the video playing. closeFilePreview() only
dropped the overlay's `visible` class, which is display:none and
nothing else, so the audio kept going with no visible player to pause.
Detaching the element is not a fix either: a detached HTMLMediaElement
plays on until it is garbage collected. _stopFilePreviewMedia() now
pauses, drops src and load()s every media element (also on re-open,
where overwriting innerHTML had the same effect), which additionally
aborts the in-flight download.
2. The scrub bar was inert. file-raw read the whole file and answered
200 with no Accept-Ranges, so Chrome reported video.seekable as
[0, 0] and silently reverted `currentTime = x`; Safari refuses to
start such media at all. Raw bodies are now streamed and range-aware:
Accept-Ranges: bytes on every response, 206 + Content-Range for a
Range request, 416 for one past EOF, and a malformed spec ignored
(200) per RFC 9110. Parsing is pure in src/web/http-range.ts.
Measured on tmp/codeman-crt-v5-66s.mp4 (18MB, 66.6s):
before seekable [0, 0] seek to 56.6s reverted to 3.9s close: still playing
after seekable [0, 66.56] seek to 56.6s landed at 60.2s close: paused, NETWORK_EMPTY
Range slices are byte-identical to `dd`, the full-file path is
byte-identical to the file, and the SVG octet-stream/attachment
hardening and the 50MB cap are unchanged (the cap is still checked
before the range).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four documentation defects found while analysing the agent skill against the
code it drives.
The lineage section attributed "deletes its session as soon as the one-shot
prompt returns" to `POST /api/v1/sessions/:id/run`. That is true of
`POST /api/v1/run`, which creates a throwaway session and calls cleanupSession
on both the success and the error path; the per-session route deletes nothing.
Name the right endpoint, and give the real reason the per-session one carries
no lineage: it is not a create call.
While verifying that, the per-session route turned out to be a sharper trap
than documented. `runPrompt()` rejects whenever a PTY already exists, which is
every interactive session, but the route has already returned `{}` with HTTP
200 by then and routes the rejection only to SSE. An agent calling it against
a live worker reads the 200 as delivery. Document it.
`Flow 3b` never existed in recipes.md. The real mapping is Flow 3 = shell
fan-out, Flow 4 = claude fan-out, Flow 5 = worker blocked on a prompt, so the
same sentence was also mislabelling Flow 4. Fixed in SKILL.md and in the
endpoints.md reference to it; every other Flow reference audited and correct.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes#259, closes#258. Both bottom out in the same gap: nothing tracked
whether the user was following live output or reading history.
#259 — the keyboard path forced the terminal to the bottom unconditionally
(onKeyboardShow/onKeyboardHide passed scrollToBottom:true, applied with no
check), so opening the keyboard while scrolled up yanked the user down. The
settle cycle now captures intent on its FIRST event, before any fit() has
reflowed the buffer, and returns to that anchor when the user was reading.
A later capture would read an already-moved viewportY, which is why the
capture point matters. The param is renamed restoreScroll to match.
Separately, flushPendingWrites gated viewport preservation on
_hasRecentUserScrollUp(), a 1500ms decay window, so a user who scrolled up and
then actually READ for longer lost protection mid-read. Being scrolled up IS
the intent however long ago it was expressed, so it now keys off position.
The recency window stays as a race guard on the sticky scroll-to-bottom.
The full-history repull already held the user's place and is unchanged.
#258 — truncation was reported by a grey line written INTO the terminal
("earlier output truncated"), which scrolls away with the output it describes,
cannot be acted on, and said the same thing whether the rest was one click away
or gone forever. The server set one `truncated` boolean at two sites meaning
opposite things, and the client discarded fullSize and source entirely.
The route now reports truncationReason ('tail' = intentional partial replay,
the rest is retained; 'capped' = the byte ceiling dropped it) plus
retainedBytes, and 'capped' is not downgraded by a later tail cut. The client
renders a dismissible banner outside terminal output with three honest states:
recoverable (offers Load full history), at-ceiling, and exhausted. The Load
button forces past the scroll cooldown but NOT past _replayWouldShrinkBuffer,
which still refuses a downgrade for repaint-mode panes.
The banner is an overlay, not a flex child: FitAddon derives rows/cols from the
terminal parent's computed height, so occupying real layout space would SIGWINCH
the CLI on every truncation-state change.
Verified in a real browser on the 7 skins: banner text and button clear 4.5:1
contrast on all of them, and terminal height is byte-identical with the banner
shown. The first cut used --bg-elevated and --accent-muted, which do not exist,
so light skins rendered a hardcoded dark bar under dark text; it now uses only
tokens every skin redefines.
test/terminal-scroll-intent.test.ts lives outside test/mobile/ deliberately —
that suite is excluded from test:ci, so a guard placed there is invisible to CI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PR #282 added pi across the prominent surfaces but left the enumerations
that read as exhaustive: the env-prefix allowlist (missing PI_*), the
external-CLI list for stop/blocked, cron's agent types (also missing
antigravity), the narrow-strip mode list, and the claude-only caveats in the
cron and Read My Mind guides. Both READMEs and the four affected docs now agree
with the schema.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Heal a stalled SSE stream: the server's :keepalive comment becomes a named
sse:heartbeat event (comments are invisible to EventSource by spec), and the
client gains a staleness watchdog that forces a reconnect after three missed
beats. Also applies a confirmed rename locally instead of waiting on SSE.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`GET /api/pi/status` shipped undocumented in the agent skill, and only a human
reading the doc noticed. Turns out none of its five siblings were documented
either, so this adds the whole family in one place: spawning with a mode whose
CLI is absent fails with OPERATION_FAILED rather than falling back, which is
exactly what an agent picking a backend it did not choose needs to know. Pi's
extra `.data.version` is called out, since a false `available:false` there means
an unrelated `pi` is in front on PATH.
On whether the endpoint scanner should also check registered-to-documented:
measured, and NO for the general case. The skill documents 34 of 217 registered
endpoints deliberately (it is an agent guide, not an API reference), so a blanket
reverse check needs a 183-entry allowlist that would fail CI on unrelated route
work and get appended to mechanically, which is worse than the gap it closes.
Grouping by path shape does not save it either: the families that yields are
things like `DELETE /api/<any>/:id`, lumping cases, webviews and docker hosts
together, and it would not have caught this gap anyway (the family had zero
documented members).
What IS cheap is a family the schema can enumerate with no allowlist: the new
assertion derives the agent modes from the Zod enum and requires each one's
`/api/<mode>/status` to be documented, so a seventh backend fails here until it
is. The sibling scanner still proves the other direction, that nothing documented
is a 404. Both mutation-checked: dropping pi's probe fails the new guard, and
documenting a nonexistent probe fails the old one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Second review pass on #282, the three items left open after f4dcfbe.
1. `codeman doctor` and the run mode disagreed about pi. The registry entry
accepted a bare `which pi` hit while pi-cli-resolver demanded semver-shaped
`--version` output, so the Dependencies panel could report an installed Pi CLI
on a box where Run Pi stays hidden, which reads as a broken mode rather than a
missing install. Both sides now share one exported PI_VERSION_REGEX, and
PathResolver gains an opt-in `requireVersionMatch` so a binary that fails the
shape check is reported MISSING instead of installed-with-unknown-version.
Only pi sets it; every other tool keeps its current behaviour.
2. The isAltScreenStripMode comment justified excluding pi with "the alt screen
is load-bearing for its fullscreen TUI". That is not what exclusion does: pi
is tmux-backed, so it falls through to isMuxAltScreenOnlyStripMode, which
strips the alt-screen toggles anyway. What exclusion actually preserves is
`\x1b[3J` and the mouse DECSETs, which is the real reason (pi renders into the
main screen and is mouse-aware). Comment and changeset now say that, and state
the consequence: fullscreen pi paints into the main buffer, like vim in a tmux
shell session.
3. skills/codeman still enumerated the five pre-pi modes in nine places, telling
agents a backend does not exist and understating class-wide caveats by one
mode. All updated, plus stale session.ts line references refreshed.
Tests: a new static guard derives the mode set from the Zod schema (not a copy)
and fails when a skill enumeration lists a partial set of external CLIs, verified
by mutation. It also documents the one legitimate exception it found: the "writes
no transcript" lists drop codex, which does write a rollout Codeman reads back.
Plus doctor cases for an unrelated `pi` on PATH and registry/resolver regex parity.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An EventSource that stops delivering does not always error. A proxy that
idle-closed the connection, a laptop resumed from sleep, a tailnet reconnect:
`onerror` never fires, the header dot stays green, and every SSE-driven surface
(tab status dots, sessions created on another device, renames) freezes until the
user reloads. Nothing on the client tracked stream liveness at all.
The server already wrote a keepalive every 15s, but as an SSE `:keepalive`
COMMENT, and comments are invisible to `EventSource` by spec, so there was
nothing a client could observe.
Server:
- `sse:heartbeat` under a new Transport category in the event registry
(155 constants now, both counts updated).
- `cleanupDeadClients()` writes that named frame (`{"t":<epoch ms>}`) instead of
the comment. Interval, tunnel padding and dead-socket eviction are unchanged.
The write stays per-client rather than going through `broadcast()`: the frame
carries no session data, so it needs no multi-user owner routing.
Client:
- `computeSseStale()` in constants.js, a pure policy beside
`computeConnectionLossUi`. Stale only when the transport believes it is
`connected`, the device is online, and no frame has arrived for 45s (three
missed heartbeats). The `connected`-only guard is also the loop breaker: a
forced reconnect leaves that state immediately, so the watchdog cannot re-fire
while one is in flight.
- The liveness stamp is applied inside `addListener` itself, so the
`_SSE_HANDLER_MAP` wrappers and the directly-registered listeners all feed it
from one place instead of three that can drift. The heartbeat's own listener
is a no-op that exists only to be registered, since `EventSource` drops named
events nobody listens for.
- A 5s watchdog forces `connectSSE()` when the policy says stale, and is cleared
at the top of `connectSSE()` and nowhere else (its only teardown path).
Recovery needs no new sync path: the reconnect re-runs `handleInit`, which
already rebuilds from the server. `visibilitychange` -> visible checks too,
riding the existing listener, since a background tab's timers are throttled
and a wake is exactly when a stream comes back zombie.
- The forced reconnect logs one diagnostic line: if a middlebox ever strips or
delays heartbeats, the failure mode is "silently reconnects every 45s", which
is undebuggable from a field report without it.
Tests: `test/sse-staleness.test.ts` (node VM over constants.js, threshold
boundaries and every not-stale guard) and `test/sse-heartbeat.test.ts` (drives
`cleanupDeadClients()` with fake replies: named frame not a comment, parseable
payload, padding only with a tunnel, dead clients still evicted).
Verified end to end on an isolated instance: with the stream closed client-side
(no `onerror`), a rename sticks, an out-of-band session stays invisible, then
the watchdog reconnects on its own and it appears without a reload.
Event names are part of the stable API contract, so this is a MINOR bump.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review follow-ups on #282. All four are the same failure shape: a list that
enumerates run modes, missed by the sweep that added 'pi'.
1. Cron ignored pi's project-trust clamp. The PR widened CronJobBaseSchema's
agentType to accept 'pi' but not the matching clamp beside gemini's, so a
non-granted multi-user owner's cron pi job spawned bare `pi` (pi's own
defaultProjectTrust, an interactive prompt they can answer "yes" to, which
loads and EXECUTES repo-local .pi/extensions TypeScript) while the same
user's UI/API launch was forced to --no-approve. The clamp is now a pure
exported helper, clampCronExternalCliConfigs(), so both it and gemini's
previously untested materialization are pinned.
2. POST /api/sessions/:id/interactive auto-enabled the Ralph tracker for pi:
its denylist covered opencode/codex/gemini/antigravity only. The tracker is
never fed for an external CLI (_processExpensiveParsers returns early), so a
pi session reported ralphEnabled and Ralph UI state no sibling backend shows.
3. REMOTE_CLI_BIN had no pi entry, so buildRemoteCliVersionProbeCommand()
returned null and Session.cliVersion stayed blank for every remote-SSH pi
session, even though the PR wired the remote launch command and the
per-mode override schema field.
4. The desktop home rail's badge map had no pi entry, and its lookup falls back
to '', which is what claude renders. A pi session read as Claude there while
the tab strip and phone overview badged it correctly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Renaming a tab appeared to do nothing: the new name only showed after a full
page reload. The PUT always succeeded; what was broken is how the tab strip
learns the result. `finishRename()` re-renders the strip from the client-side
`app.sessions` map, and nothing wrote the new name into that map, so the rename
depended on the `session:updated` SSE frame to carry its own write back. On a
page whose stream has gone quiet without erroring, that frame never lands and
the re-render repaints the stale label.
- `_applyLocalSessionName()` writes the confirmed name into `this.sessions` and
refreshes cached subagent parent names, mirroring `_onSessionUpdated`.
- `_putSessionName()` returns the stored name or null. `_apiPut` turns a network
error into a null Response and an API failure into a non-ok status, so a
rejected rename previously read as success and silently dropped the edit (the
old try/catch could never fire).
- Both surfaces use them: `startInlineRename()`'s `finishRename` and
`saveSessionName()`.
Two regression tests: the commit applies the name with no SSE frame dispatched,
and a 500 restores the old label, leaves the map untouched, and toasts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SessionMode gains 'pi', a first-class backend alongside Claude Code,
OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose
tab identity, welcome button, run-mode entry, cron agentType, Docker and
remote-SSH command defaults, and clone-repo Brain option.
Pi is a different shape of CLI from the other four, and three decisions
follow from that:
- It has NO permission prompts and no sandbox, so there is no
--dangerously-skip-permissions analog and none was invented. The
privilege-shaped knob is the tri-state approveProjectTrust, which makes
pi load and EXECUTE repo-local .pi/extensions TypeScript and install
missing project packages. clampExternalCliBypassForOwner() therefore
puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets
--no-approve even when no config was sent, because pi's own default is
a prompt the session user could answer themselves. That helper had zero
test coverage; it now has coverage for all four CLIs.
- Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars
share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode
context, so admitting them would widen the allowlist for every mode at
once. Auth goes through pi's /login or the server's own environment.
--api-key is deliberately never wired: it would put a provider secret on
the spawn command line.
- pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the
main screen with terminal-owned scrollback, and its 0.84.0 fullscreen
mode is runtime-switchable via /settings; that flip was measured to put
the pane into the alt screen, which the strip would have corrupted.
pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires
semver-shaped output, because `pi` is a short generic name a stray binary
can shadow; GET /api/pi/status surfaces path and version so a
misresolution is diagnosable rather than presenting as a broken mode.
Docker installs pi in its own --ignore-scripts step so that flag cannot
affect the other four CLIs, and seeds its credentials per-file rather than
whole-dir (~/.pi/agent also holds sessions, extensions and package trees).
Verified end to end against pi 0.84.1 on an isolated instance: resolver
search-dir fallback, flag construction, piConfig persistence across a full
server restart, the trust prompt and its --no-approve suppression, the
rose Run button on the default daylight-blue skin (the nested skin block
eats per-mode gradients unless the rule lives inside it), and the buffer
local-echo policy, which pi tolerates where codex did not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Release 1.16.6: phone overview started/idle stamps, plus fixes for the
selection-dialog keyboard lockout, the accessory bar arrows bypassing the
local-echo overlay, and recovered sessions being restamped as newly created
on every server restart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Below 860px Save moves into the header (a bottom action bar would cost 60px
of a phone sheet), which left the two ways OUT of the sheet sitting side by
side in mismatched shapes: a fat accent pill next to a bare 1.5rem glyph
with no box at all. They are the same decision (save-and-close vs
discard-and-close), hit in the same corner with the same thumb, so they now
share a recessed tray and matching pill geometry and read as one cluster.
- 36px on both, so the tray comes out at 44px including its 3px padding and
1px border — the same height as the phone header it sits in.
- `.modal-close` gets a real box (36x36, radius 9) only inside the tray; its
bare-glyph form is still right in a plain modal header.
- Tray colors come from skin tokens (--border/--bg-input). A hardcoded black
alpha would render as a grey slab on the four light skins, the same trap
the layout preview frame hit.
- `:has(.set-head-save)` keeps the tray off the sheets that carry a lone x:
Session Options and Add Case save from inside their own forms.
- The shared focus ring offsets OUTWARD, which inside the tray would draw on
top of the tray border, so it is inset to ring the button instead.
DOM order stays close-then-save so the focus trap still lands on Close;
row-reverse paints Save to its left.
Verified at 390x844: tray 44px tall, Save 36px, Close 36x36, both radius 9
inside a 12-radius tray. PostCSS-parsed (prettier does not catch an unclosed
CSS block, and styles.css is prettier-ignored by design).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#279 and #280 auto-merge cleanly, but the merged result was red: neither
branch could see the other, and CI cannot see either, because the only test
covering #279 lives in test/mobile/** which test:ci excludes.
Two problems, both in #279's test:
1. The in-terminal case tapped the terminal's top-left corner, i.e. an inert
transcript row, and asserted focus was retained. That is precisely the
gesture #280 redefines, so #280 turned it red. Aim it at the PROMPT row
instead: the one in-terminal tap whose outcome neither PR claims, so it
still proves the #terminalContainer exemption without asserting the
toggle's behaviour.
2. The "a real control is exempt" case was VACUOUS. It picked the first
button measuring >8px, which is .welcome-ralph-link inside the welcome
overlay hideWelcome() had already hidden: the rect still measures, but
elementFromPoint at that point returns .xterm-screen, so the case tapped
the TERMINAL and passed for the wrong reason. It only surfaced because
#280 changed what a terminal tap does. Require the sampled point to
actually resolve to the button, and fail loudly when no control is
usable rather than silently asserting nothing.
Mutation-checked: removing the install, the #terminalContainer exemption,
the control exemption or the `if (moved) return` scroll guard each turns
the test red on its own. The control exemption had no coverage before.
Also fold the duplicated tap slop into one constant: initTerminal's
TAP_THRESHOLD now reads MOBILE_KEYBOARD_DISMISS_TAP_SLOP instead of
re-declaring 8, since a drift between them is exactly the bug the second
#279 commit fixed. And restore the comment the slop constant was inserted
into the middle of, which left "Regions where a tap must NOT dismiss"
sitting above the slop rather than the selector it documents.
test/mobile/keyboard.test.ts: 5 failed | 47 passed (52). Master is
5 failed | 46 passed (51) — the same five pre-existing failures.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every terminal tap re-focuses the hidden textarea, so once the on-screen keyboard
is open the only way to close it is the accessory bar's dismiss chevron. Tapping
the transcript to get the screen back is the obvious gesture and it did nothing.
A tap on INERT content with the keyboard already up now dismisses it. Nothing
else claims that gesture: an inert row has no action to trigger, so by that point
the tap has already done its only other job (the mouse report).
Scoped to 'content' ON PURPOSE. The prompt row ('input') keeps
focus-then-position, so a second tap there still places the caret — that is real
capability and trading it away would be a worse deal than the bug. A separate
test pins it rather than leaving it to the reader.
Actionable rows are unchanged: readbacks, "esc to interrupt" status rows and menu
selections still blur via _isActionableMobileTerminalTap, which runs first.
`keeps the hidden keyboard input focused after an inert Claude transcript tap`
asserted the OLD behaviour and is renamed and inverted, since revising that
behaviour is the point of this change. Its setup already focused the terminal
before tapping, so it was always exercising the second-tap case.
test/terminal-touch-tap.test.ts: 28 tests. The two new ones fail on master —
`closes the keyboard on a second tap of INERT transcript content` behaviourally,
by asserting blur where master re-focuses.
test/mobile/keyboard.test.ts: 51 tests, 5 failed | 46 passed — the same five
pre-existing failures as master, untouched here.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The dismiss handler fired on any touchend, so a scroll closed the keyboard too —
a regression the original test could not see, because it only ever dispatched a
stationary tap.
The helper now takes an optional travel distance and emits touchmove steps, and
the test asserts a 120px scroll leaves the terminal input focused. Removing the
`if (moved) return` guard fails this assertion, so it genuinely pins the fix.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Regression from the dismiss handler in #279: it fired on any touchend,
and a scroll ends in touchend too. Scrolling to read something while composing
closed the keyboard and dropped the composer — worse than the bug it fixed.
Track finger travel from touchstart and only treat a near-stationary gesture as
a tap, using the same 8px TAP_THRESHOLD the terminal's own touch handling uses
so both agree on tap-vs-scroll. Multi-touch is never a dismissing tap.
All three listeners stay passive; nothing calls preventDefault.
Measured on a Pixel-class viewport with a Firefox UA:
tap -> dismissed
scroll (120px) -> keyboard kept
micro-drift (4px) -> dismissed, so an imprecise tap still works
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
On a phone the terminal holds focus on a hidden textarea, and nothing ever
released it. Once the keyboard was up, tapping the header, the tab strip or any
empty page chrome left it up — covering roughly half the screen with no in-app
way to dismiss it.
Repro, iPhone-class viewport (390x844), claude-mode session, focus the terminal
then tap the header logo:
| | document.activeElement after the tap |
| --- | --- |
| master | textarea.xterm-helper-textarea (keyboard stays up) |
| this branch | body (keyboard closes) |
A document-level touchend handler blurs the terminal input, deliberately scoped
so focus is never stolen from something that wants it:
- only when the terminal input actually holds focus;
- never inside #terminalContainer — _handleMobileTerminalTap already classifies
and routes those taps and owns that decision;
- never on a control. Anything focusable or clickable is about to take focus
itself, and the keyboard accessory bar exists to be used WHILE the keyboard is
open, so dismissing there would fight the user.
Bound to touchend rather than click: a tap meant to dismiss usually is not meant
to activate what sits underneath, and touchend fires before the synthesized
click so the blur lands first. The listener is passive — it never calls
preventDefault.
Test: `dismisses the on-screen keyboard when a tap lands outside the terminal`
in test/mobile/keyboard.test.ts. It fails on master with a BEHAVIOURAL assertion
(`expected 'xterm-helper-textarea' not to contain 'xterm-helper-textarea'`),
not a TypeError, and passes here. It drives real dispatched touch events rather
than calling the helper, because the handler is bound on document and a direct
call would bypass the routing under test.
test/mobile/keyboard.test.ts: 52 tests, 5 failed | 47 passed. Master is 51 tests,
5 failed | 46 passed — the same five pre-existing failures (stale layout and
accessory-bar expectations, a CJK timeout), untouched here.
Full suite: 4944 passed | 12 skipped, 0 failed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PR #274 lifted the Run menu's 250px cap to `calc(100vw - 24px)` so a
recent-session row would have room for its worktree pill and parent path.
The rows never took it: `.run-mode-history` is a block scroller, so its
<button> rows are shrink-to-fit and stayed at ~250px inside a 1376px menu,
leaving ~1100px of empty dropdown and no space for `.hist-dir`'s
`flex: 1` + `text-align: right` to expand into.
Rows now fill the menu, and the menu is capped at the 760px one full row
actually costs rather than the whole window.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Addresses the review on #244.
BLOCKING (item 1). selectSession() ends with scrollToLastNonEmptyLine(), which
parks the viewport above the bottom for any session taller than the screen, so
after a tab switch every tap classified as 'history' — touchstart ran
preventDefault() + blur, and touchend's early return skipped focus. Both routes
to focus closed on one gesture, the same mechanism as #173.
Suppressing the mouse REPORT while scrolled up is right and is kept; suppressing
FOCUS is not. touchstart now only preventDefaults 'content' taps (a scrolled-up
viewport sends nothing, so there is no compatibility click worth cancelling), and
the 'history' branch focuses instead of blurring.
Verified against the maintainer's own test, which was already on master and red:
`keeps the terminal input focusable after a tab switch parks the viewport
off-bottom` fails without this change and passes with it.
Item 2: dropped both `terminal-action-pending` guards. The class exists nowhere
in the repo, so both branches were permanently false and the comment promised
coverage that did not exist.
Item 3: removed the `Working` literals. Live claude 2.1.226 prints
"Cooked for 2m 6s" with a different bullet and a randomised verb, so they were
dead code. The status row is matched by its affordance ("esc to interrupt")
instead, which is what makes it actionable. The affordance regex is also
tightened to require a key or gesture name, so prose like "click here to open
the file" no longer dismisses the keyboard.
Item 4: removed _shouldForwardTouchScrollToApp and its test. It was never called,
and wiring it as written would have restricted forwarding to claude only,
dropping gemini from the path #205 established — a behaviour change this PR has
no reason to make.
Smaller items: the touchstart classification is cached and reused for the
touchend of the same gesture (keyed on exact coordinates, so a moved finger
re-classifies), removing two of the three full-viewport scans per gesture; the
duplicated touchLastX assignment is gone; and the no-touch bail-out returns null
rather than claiming 'history'.
test/mobile/keyboard.test.ts: 51 tests, 5 failed | 46 passed — the same 5
pre-existing failures as master, unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A mid-terminal tap on a claude-mode session left document.activeElement on
<body>, so the on-screen keyboard could not be raised and there was no way to
type — the blocker reduced upstream in #173.
_classifyMobileTerminalTap returns 'content' for any non-prompt row, and
_handleMobileTerminalTap blurred on every 'content' tap while touchstart's
preventDefault had already cancelled the compatibility click that would
otherwise focus xterm. Both routes to focus were closed on the same gesture.
Blur now applies only to rows that are actually TUI-owned. The distinguishing
signal is the affordance a CLI prints on or beside the row ("ctrl+r to expand",
"tap to collapse", "esc to interrupt"), not the row's title text — a readback's
title row carries no hint of its own, so the adjacent row is consulted too.
Keying on titles would recognise only the exact strings a fixture happens to
use and would let a real readback keep the keyboard open.
Measured with a real touchstart/touchend gesture, iPhone-class viewport,
claude-mode session, tapping mid-transcript:
before document.activeElement = body
after document.activeElement = xterm-helper-textarea
Note: upstream master already passes this assertion, so the added test is a
regression guard for this branch, not a test that fails on master.
test/mobile/keyboard.test.ts: 40 tests, 5 failed | 35 passed — the same 5
pre-existing failures as master (stale layout/accessory-bar expectations and a
CJK timeout), unchanged by this commit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both are read by every engine (the Claude path sends the language as its base
tag and the keyterms as a recognition hint), but they sat under the "Deepgram
Nova-3" heading, which read as if they only applied to Deepgram. That group now
holds just the API key.
Ids are unchanged, so the getElementById load/save contract in settings-ui.js is
untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mic button previously needed a Deepgram API key, or fell back to the
browser's Web Speech engine. It can now transcribe through the same
speech-to-text service Claude Code's own /voice mode uses, so anyone signed
in to Claude Code on the server gets dictation with no third-party account.
Claude Code's voice mode cannot be driven directly: it opens the HOST's
microphone (sox/arecord), and the CLI runs in a headless tmux pane while the
human is in a browser somewhere else. So capture stays in the browser and only
the transcription backend is borrowed.
Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the
page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since
MediaRecorder cannot emit raw PCM) and receives text.
- GET /api/voice/status reports readiness and never the token
- GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade
guard as the terminal socket, plus caps on concurrency, stream length and
frame size
- credentials are read-only: Codeman never refreshes them, since a refresh
rotates the refresh token and could sign the user out of their own CLI
- claudeVoiceEnabled (synced, default OFF) gates the whole server side
- voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers
Claude, then a configured Deepgram key, then the browser
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The rule "show ~/project rather than /home/<user>/project" had three
implementations in the frontend, two of them platform-specific in opposite
directions, so each looked correct to whoever wrote it.
- The Run menu's Recent Sessions rows matched /home/<user>/ only. On macOS
nothing was stripped, so every row spent its first ~19 characters on an
identical /Users/<user>/ prefix and the left-to-right ellipsis removed the
tail that identifies the row. That is #273, reported by @jordan8037310, who
also traced why the menu's 250px cap made it worse: the width was chosen on
the assumption the abbreviation had run.
- The case-manage list matched /Users/<user> only, the mirror image, so on a
Linux host no case path was ever abbreviated there. Unreported.
Both now call _shortenHomePath(), which was already correct for both layouts
and already used by the Resume list, Cmd+K, the desktop home rail and the phone
overview. Its regex collapses to one alternation with a lookahead, so a path
that is exactly $HOME renders "~" instead of being left raw, matching what the
case-manage list used to do on macOS.
test/home-path-abbreviation.test.ts pins the helper on both layouts and the
rendered case-manage label, and fails if a fourth copy of the pattern appears in
src/web/public. The Run-menu guard counts helper calls rather than pinning a
source line, so it survives the row restructure in #274.
test/run-mode-ui.test.ts gains a _shortenHomePath stub: its harness loads
session-ui.js without terminal-ui.js, which the real app never does.
Verified against an isolated instance with 27 real cases and 50 history rows:
27 of 27 case paths and 17 of 20 Run menu rows abbreviate, the other 3 are
/tmp paths that correctly stay raw, tooltips keep the full path, no page errors.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The footer buttons shipped with class="btn btn-secondary/primary", but no
.btn or .btn-secondary rule exists in this codebase, so all four rendered
as unstyled UA buttons. Moved them to the btn-toolbar convention every
other modal footer uses, with a scoped flex-row footer rule (btn-toolbar
is display:flex, block-level) mirroring the runSummaryModal footer.
Send's accent needs a (0,4,0) re-assert: the skin block's bare
.btn-toolbar rule is (0,2,1) under html:not([data-skin="og"]) and beats
.btn-toolbar.btn-primary (0,2,0), the same specificity trap CLAUDE.md
documents for mobile.css. Scoped to this modal; the repo-wide greying of
btn-primary on non-OG skins is pre-existing and left as a design call.
The empty-result copy now points at the steer note sitting right below
it ("Add a steer note and Rethink to try again"), zh-CN updated.
Verified with the steer E2E (still green) plus desktop, phone (390px),
and error-phase screenshots; static guards extended to pin the footer
convention and the accent re-assert.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds the optional free-text steer note to the Read My Mind modal: a
dashed input under the suggestions ("no, I meant the mobile bug") that
rides along as `steer` on every Rethink. The API already accepted it;
this wires the frontend end of the contract.
- Shown whenever Rethink is live (ready AND empty-result phases),
hidden only while a prediction runs; typed text survives re-runs.
- Enter in the field triggers Rethink, mirroring the prompt field's
Enter-to-send; a fresh open clears it with the rethink memory.
- Trimmed and capped to the schema's 2000 chars on the way out; a
plain open still sends an empty body (neither steer nor rejected).
- zh-CN strings for the placeholder and aria-label, phone-sized
touch target in mobile.css, static guards in the phase-3 test.
Verified with a browser E2E against a live dev server (stubbed predict
endpoint): payload contents, phase visibility, Enter wiring, and
reset-on-reopen all asserted with real keystrokes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes#273. Every row in the Run dropdown's Recent Sessions list rendered as
`/Users/<user>/co…`, indistinguishable from every other row.
The width was the symptom. The cause is that the home-prefix abbreviation
matched `/home/<user>/` only:
s.workingDir.replace(/^\/home\/[^/]+\//, '~/')
On macOS the prefix is `/Users/<user>/`, so nothing was stripped and every row
spent its first ~19 characters on an identical prefix, with left-to-right
ellipsis cutting the only part that identifies it. The 250px menu cap was
chosen, per its own comment, as "the width at which the common `~/<dir>/<repo>`
+ timestamp recent-session row still fits whole" — sizing that assumes the
abbreviation ran. On Linux it does. On macOS the menu was permanently too
narrow for content it was never actually shortening, which is why this reads
as fine on one platform and broken on the other.
Changes:
- the regex matches `/home/` and `/Users/`
- the row leads with the identifying folder in semibold, with the parent path
trailing, dimmed and right-aligned, so truncation removes context instead of
identity
- the menu goes full width above 769px and the history list grows 200px -> 320px.
Phones keep the compact popover deliberately: mobile.css positions this menu
itself and a viewport-wide drawer there would cover the composer
- a worktree pill renders from the fields /api/history/sessions already returns
unprojected (#266/#269), since a worktree's directory basename is often just
the worktree name and rows stayed ambiguous without it
- a trailing `/.claude/worktrees` is trimmed from the displayed parent path once
the pill states it, so the repo name stays visible
Verified in a browser at 1440px against a real 38-session history: menu 1416px,
0 of 34 rows clip their project name (was: all of them), 9 worktree pills
render, no page errors.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uTqt8ttmsBLXbm5JFHis3
The synthetic session names end up in the harness screenshots, so shipping one
contributor's project list into everyone else's review reads oddly. The mix of
CLI modes is what the fixture actually needs — each renders a different badge —
and that is unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The script carried two absolute paths from the machine it was written on: a full
scratchpad path including a session UUID, and /home/chaberl/projects as the
synthetic sessions' working directory. This branch is pushed to a public fork, so
they were visible to anyone.
Screenshot output now defaults to tmpdir() and is overridable via
SIDEBAR_SHOTS_DIR; the synthetic working directories are tmpdir()-based too, which
also makes the harness run for anyone who checks the branch out.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The header tab strip stops working past roughly a dozen sessions: it wraps
into two or three rows, eats vertical space and still cannot be scanned.
This adds a vertical session list in a left <aside> as an ALTERNATIVE
layout — a filter box, a live count, and a 44px collapsed rail that keeps
the ambient signal (status dot, task badge) visible.
The strip is not removed. Settings -> Display -> Tab Bar -> Session List
Layout switches between them and the default stays 'header', so existing
users see no change until they opt in.
Structure: one #sessionTabs element, two mount points. applySessionListLayout()
re-parents the SAME node between #sessionTabsHost and #sessionSidebarList,
which is why there is no second renderer and no duplicated wiring — app.$()
caches getElementById results and never invalidates them, so a moved node
keeps every existing consumer (settings-ui, webview-tabs, the generated
gesture bundle, the mobile tests) working untouched.
Notable integration points:
- Below 1024px the sidebar is an off-canvas drawer overlaying the terminal;
closed it gets inert + aria-hidden so it cannot be tabbed into, and touch
swipes over it no longer switch sessions.
- Subagent and ultracode windows anchor to the right edge of a sidebar row
instead of its bottom, connector curves follow.
- Alt+B toggles; the chord is gated out of the PTY so xterm cannot also
write ESC b into a live session.
- Collapse state lives in its own localStorage key (the settings blob is
rebuilt from DOM controls on every save) and falls back to in-memory
intent where storage throws.
Verified: frontend syntax + public asset checks, tsc, eslint, 26 new jsdom
tests, and a headless-Chromium harness (scripts/verify-session-sidebar.mts)
that renders a synthetic 25-session fleet in both layouts at 1600/1000/393px
and asserts mount point, widths, inert/aria state and row count.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:42:21 +02:00
535 changed files with 104110 additions and 6040 deletions
Keep the terminal anchored where you are reading while an agent streams (#358). Scrolling up during a Codex response could still be dragged back to the live bottom by the next redraw: the flush captured the viewport before writing and restored it immediately after, but xterm parses asynchronously, so at that moment the buffer had not moved yet, the restore compared the anchor against itself and did nothing, and the redraw landed a tick later with nothing left to pull the view back. The restore now runs inside xterm's own write callback, which is the first point at which the redraw's effect exists, and it holds across consecutive and chunked redraws. It is dropped if you switch sessions or a history replay starts before the write parses, since the anchor indexes the buffer it was captured from.
"description":"Codeman, self-hosted mission control for AI coding agents. Ships the codeman agent skill: let one Claude Code session spawn, prompt, wait on and read other sessions.",
"plugins":[
{
"name":"codeman",
"source":"./plugins/codeman",
"description":"Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
Thanks for wanting to help! Codeman is a small project with a fast loop: issues usually get a response within a day, good PRs get reviewed quickly, and every release credits its contributors and bug reporters by name in the release notes. This guide gets you from clone to merged PR without stepping on the traps.
## The short version
1.**Bugs**: open an issue with your OS, install method (installer / npm / git clone), browser, and which CLI + version the session was running.
2.**Questions and ideas**: use [Discussions](https://github.com/Ark0N/Codeman/discussions), not issues.
3.**Small fixes** (docs, typos, a new skin, a translation): just send the PR.
4.**Anything bigger**: open an issue or Discussion first and get a nod before building. Codeman has strong architectural invariants, and a design chat up front is what turns a big idea into a merged PR instead of a stalled one. This flow works: features like Clone Repo (#236) went idea, then design discussion, then review, then shipped.
5.**Security issues**: never a public issue. See [SECURITY.md](SECURITY.md).
## Dev setup
Requirements: Node.js 22+ (see `.nvmrc`), tmux, and at least one supported agent CLI on your PATH (Claude Code is the primary one).
```bash
git clone https://github.com/Ark0N/Codeman.git
cd Codeman
npm install # postinstall builds the vendored xterm addon bundles
npm run dev # dev server on http://localhost:3000
```
The frontend is plain JS served from `src/web/public/` with no bundler in dev: edit a `.js`/`.css` file and reload the page. The one exception is `index.html`, which is read once at server start, so markup changes need a server restart.
## Before you push
CI runs all of these, so save yourself a round trip:
```bash
npm run typecheck # tsc --noEmit, strict mode
npm run lint
npm run format:check
npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules
```
### Tests
```bash
npm test# the gate — exactly what CI runs
npm test -- test/<file>.test.ts # one file
```
`npm test` is the same suite CI runs, so a green run locally means a green run there. It leaves out three suites that cannot pass on an arbitrary machine, each with its own command:
```bash
npm run test:browser # Playwright + chromium (+ a live server; codex-predictive-echo needs a real codex binary)
npm run test:mobile # the above plus environment-specific PNG baselines
npm run test:perf # wall-clock benchmarks — run on an otherwise idle machine
npm run test:all # literally everything, environmental failures included
```
Expect `test:browser`/`test:mobile`/`test:perf` to fail where the machine cannot provide what they need; read that as "not runnable here", not as a regression. `config/test-suites.ts` holds the globs, and both configs derive from it, so the exclusions and those runners cannot drift apart.
If you add a test that binds a port, pick a unique one at 3150 or above (search the repo for `const PORT =` first). Never 3000.
Tests are tmux-safe by design: under vitest, the tmux layer becomes an in-memory mock, so tests cannot touch real sessions.
## Finding your way around
- Every source file starts with a `@fileoverview` JSDoc block. Read it before diving into the file, it is the map.
- [`CLAUDE.md`](../CLAUDE.md) at the repo root is the densest architecture primer in the repo. It is written for AI coding agents, but the invariants and gotchas in it apply to humans exactly the same, and most review feedback on PRs traces back to something already written there.
- Deep mechanisms and the history behind each rule live in [`docs/architecture-invariants.md`](../docs/architecture-invariants.md).
- Third-party extension surfaces are documented in [`docs/extending-codeman.md`](../docs/extending-codeman.md).
## Great first contributions
These are well-fenced areas where a first PR is genuinely easy to get right:
- **A new theme skin.** A skin is four things kept in sync: the `html[data-skin="…"]` token block in `styles.css`, the xterm ANSI palette in `terminal-ui.js`, the pre-paint allowlist and the settings picker (both in `index.html`). `test/skin-themes.test.ts` statically checks the sync, so if the test passes, your skin works.
- **A new language.** `src/web/public/i18n.js` is dependency-free, English is the canonical source, and `zh-CN` is a complete example to copy. Add your language's entries and register it in `SUPPORTED_LANGUAGES`.
- **Docs.** If you got stuck on something and then figured it out, the sentence that would have unstuck you is a PR.
- Anything labeled [`good first issue`](https://github.com/Ark0N/Codeman/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22).
Bigger extension points worth discussing first: new CLI backends (the pluggable resolver pattern has absorbed six CLIs so far; `docs/extending-codeman.md` and `docs/opencode-integration.md` show the shape), and real-device testing reports, especially mobile, which always find things emulation cannot.
## PR expectations
- **One change per PR.** Small and focused reviews fast; a grab-bag stalls.
- Target the `master` branch.
- **Keep your branch mergeable.** A PR with conflicts silently gets no CI runs at all (GitHub quirk), so rebase or merge master when conflicts appear.
- Include or update tests when you change behavior. Route handlers have a lightweight pattern in `test/routes/` using `app.inject()` (no live server needed).
- Formatting is Prettier with a deliberately narrow scope (`npm run format`), several frontend files are hand-formatted on purpose and excluded via `.prettierignore`. Don't "fix" a file by adding it back into Prettier's scope.
- Don't bump versions or touch `CHANGELOG.md`; releases are handled by the maintainer via changesets after merge.
- AI-assisted contributions are welcome (much of Codeman is built that way), with one condition: you must understand what you're submitting and have actually run it. "The model said it works" is not a test.
## Conduct
Be kind, be direct, assume good faith. Report unacceptable behavior privately via the contact in [SECURITY.md](SECURITY.md).
@@ -69,10 +69,18 @@ shared-host, multi-user, or tunneled deployments.
- **Multi-instance tmux socket is process-wide.** Two Codeman instances on the same `CODEMAN_INSTANCE` share a tmux socket and can attach each other's live sessions — isolate with distinct `CODEMAN_INSTANCE` values.
- **The live log-tail route reads `/var/log` and `~/logs`** in addition to the session working directory (read-only) — a deliberate choice for tailing system/app logs. On a password-protected remote deployment an authenticated user can therefore read those roots outside their session. See `docs/security-architecture.md` §5.
Recent hardening (this release): web-push subscription endpoints are restricted
to https public hosts (SSRF guard — rejects internal/metadata IPs, validated at
subscribe and send time), and tmux session names discovered on the shared socket
are validated against the safe-name pattern before reaching any shell call site.
- **The web-tab proxy fetches from the server's network position.** Any authenticated user can save a dashboard URL on loopback or a private range and have Codeman relay to it; that is the feature. Link-local and cloud-metadata addresses are the only refused targets (see below). On a shared host, restrict who holds an account.
Recent hardening (2026-09-04): the web-tab proxy, its "Test" probe and its
WebSocket relay refuse link-local and cloud-metadata targets (`169.254.0.0/16`,
if ! git clone "https://x-access-token:${WIKI_TOKEN}@github.com/${GITHUB_REPOSITORY}.wiki.git" wiki 2>"${RUNNER_TEMP}/clone-err.txt"; then
cat "${RUNNER_TEMP}/clone-err.txt"
echo "::error::Could not clone ${GITHUB_REPOSITORY}.wiki.git. If this says 'Repository not found', the wiki has never had a page: save one at https://github.com/${GITHUB_REPOSITORY}/wiki/_new and re-run. If it says 403, add a WIKI_TOKEN secret."
exit 1
fi
- name:Mirror pages
run:|
set -euo pipefail
# The mirror deletes before it copies, so an empty source would wipe
# every published page and the commit step would happily push that. A
# MISSING directory already fails safely (cp aborts under set -e); an
# empty one does not, so check explicitly. This is the one failure mode
# here that destroys something a browser edit cannot get back.
if [ ! -d docs/wiki ]; then
echo "::error::docs/wiki does not exist. Refusing to mirror, which would delete the entire published wiki."
git commit -m "docs: sync wiki from docs/wiki @ ${GITHUB_SHA:0:7}"
if ! git push 2>"${RUNNER_TEMP}/push-err.txt"; then
cat "${RUNNER_TEMP}/push-err.txt"
echo "::error::Could not push to ${GITHUB_REPOSITORY}.wiki.git. A 403 here means the token can read the wiki but not write it, which is the usual GITHUB_TOKEN case: add a fine-grained PAT with wiki write access as the WIKI_TOKEN secret."
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, or Gemini CLI inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
Get started in one line (macOS & Linux, Windows via WSL):
@@ -42,7 +42,7 @@ codeman web
The installer asks before every system change, and re-running the same line updates in place. Full details: [Quick Start - Installation](#quick-start---installation).
- **One dashboard, five CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, or Gemini](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **One dashboard, eight CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **Truly phone-friendly** - a [touch-optimized terminal](#mobile-optimized-web-ui) with instant local echo, QR login, swipe navigation, and push notifications
- **Runs while you sleep** - [idle detection + respawn cycling](#respawn-controller) and auto-resume when a subscription limit resets, for 24+ hour unattended runs
- **See your agents think** - [live floating windows](#live-agent-visualization) for every subagent and teammate, with real-time transcripts
@@ -61,14 +61,14 @@ The installer asks before every system change, and re-running the same line upda
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it. A few things worth knowing:
This installs Node.js, tmux and a build toolchain if missing (node-pty ships no Linux prebuilds, so it compiles from source), clones Codeman to `~/.codeman/app`, and builds it. A few things worth knowing:
- **It asks first.** Every system change (package installs, AI CLI download) is prompted, and a menu at the end lets you choose: run Codeman in this terminal, install it as a background service (systemd/launchd, auto-start on boot), or don't start yet. Nothing runs in the background unless you pick it.
- **Network or local-only, your choice.** The installer asks whether the dashboard should be reachable from other devices on your network (`0.0.0.0`, the default, with a strongly recommended password prompt) or from this machine only (`127.0.0.1`, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. A bare `codeman web` started by hand still defaults to loopback.
- **How it's reachable, your choice.** The installer offers three ways to reach the dashboard: **Tailscale** (loopback bind fronted by `tailscale serve`, so you get `https://<machine>.<tailnet>.ts.net` with a real certificate and your tailnet as the login, no password needed), **any device on your network** (`0.0.0.0`, with a strongly recommended password prompt), or **this machine only** (`127.0.0.1`, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. The highlighted default reflects what is already on the machine (Tailscale when it is already in use, your existing binding on a re-run), and a bare Enter never pulls in new software. A bare `codeman web` started by hand still defaults to loopback.
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh update` and `install.sh uninstall` also exist.
- **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), or [Gemini CLI](https://github.com/google-gemini/gemini-cli) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the five is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), or [OMP](https://github.com/can1357/oh-my-pi) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the nine is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
```bash
codeman web
@@ -82,6 +82,8 @@ codeman users add alice --admin # create the first admin account
codeman web --multiuser # named logins + per-user case spaces
```
**Prefer Docker Compose?** A local-image Compose deployment ships in `docker/`: copy `docker/.env.example` to `docker/.env`, set `CODEMAN_PASSWORD`, then run `bash docker/Start-Codeman.sh` on Linux. Codeman runs in a container and spawns Docker cases as sibling containers through the host socket. See the [Docker deployment guide](docker/README.md) for direct Compose commands, storage and networking options.
Details in [Multi-User Mode](#multi-user-mode-opt-in) below.
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), or [Gemini CLI](https://github.com/google-gemini/gemini-cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), or [OMP](https://github.com/can1357/oh-my-pi)). After installing, `http://localhost:3000` is accessible from your Windows browser.
</details>
@@ -207,7 +209,7 @@ The most responsive AI coding agent experience on any phone. Full xterm.js termi
<tr><td>Password typing on phone</td><td><b>QR code scan — instant auth</b></td></tr>
</table>
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard; destructive commands require a double-press to confirm, so you never fire one by accident
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard; destructive commands require a double-press to confirm, so you never fire one by accident; on Codex sessions the bar also shows `⇧←` / `⇧→` (Shift+Left / Shift+Right: edit the last queued message / return through the prompt stack)
- **Dedicated Enter button** — replays the keypress through the terminal, so text buffered by local echo is flushed first rather than stranded
- **Swipe navigation & smart keyboard handling** — swipe left/right to switch sessions; toolbar and terminal shift up when the keyboard opens (`visualViewport` API)
- **Built for phones** — safe-area insets for notch and home indicator, 44px touch targets, bottom-sheet case picker, native momentum scrolling
@@ -253,7 +255,7 @@ Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. **Add Case** creates one from scratch, links an existing folder, or clones a GitHub repo straight into one (**Clone Repo**). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, or `Terminal` (plain shell). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`,`Pi`, `Grok`, `OMP`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Models → New Claude sessions). A soft default — `/model` still works in-session. |
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
@@ -285,7 +287,7 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
- **Phone/tablet** — the UI is fully touch-optimized; scan the desktop **QR code** to log in without typing a password.
- **Outside your network** — `./scripts/tunnel.sh start` opens a Cloudflare tunnel (set `CODEMAN_PASSWORD` first).
- **SSH** — the `sc` chooser attaches to any session from a terminal (`sc` interactive, `sc 2` quick-attach, `sc -l` list).
- **SSH** — `codeman tui` is a full-screen dashboard in the terminal (`codeman tui --list` to list, `codeman tui 2` to attach straight to one).
### 7. Operate & maintain
@@ -406,6 +408,14 @@ The title is templated into the served HTML on first byte, so it's correct from
| **110k tokens** | Auto `/compact` | Context summarized, work continues |
| **140k tokens** | Auto `/clear` | Fresh start with `/init` |
### Tab Alerts
<p align="center">
<img src="docs/images/tab-alerts-glow-20260815.gif" alt="Session tabs: a regular active tab beside a yellow waiting-for-input tab and a red needs-decision tab, both with a breathing glow" width="900">
</p>
Every tab tells you its state at a glance. A running session keeps its green status dot. When a session stops and waits for input, its tab turns **yellow**: steady ring, tinted background, yellow dot, with a slow breathing glow on top. When a permission prompt or question is **blocking** the agent, the tab turns **red** with a faster pulse. The base tint never blinks off, so even a split-second glance (or a screenshot) reads the true state; the ring stays visible while the tab is selected, and a page reload re-arms pending alerts from the server, so a blocked session can never hide behind a fresh-looking tab.
### Notifications
Real-time desktop alerts when sessions need attention — `permission_prompt` and `elicitation_dialog` trigger critical red tab blinks, `idle_prompt` triggers yellow blinks. Click any notification to jump directly to the affected session. Hooks auto-configured per case directory.
- **Background daemon & service install** — `codeman web -d` runs the server detached with a pidfile, `~/.codeman/web.log`, and verified startup (it polls the server until it answers, so a port clash never reads as success); `codeman service install` writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew `node`, `tmux` and `claude` are actually found. Secrets are never written into unit files
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → System → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Clone a GitHub repo as a case** — paste a repository URL into **Add Case → Clone Repo** and Codeman clones it into `~/codeman-cases/<name>` and registers it as a normal case, ready to run an agent in. It preflights the URL while you type (tells you whether it can be cloned anonymously and offers the repo's real branches and tags for the optional branch/tag field), fills the case name in from the URL, and lets you pick which CLI the Run button should use. Public repositories over `https://`; Codeman never collects or stores credentials
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, or**Gemini** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, **Pi**, **Grok**, or **OMP** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*` vs `GROK_*`/`XAI_*` vs `OMP_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md), [`docs/pi-integration.md`](docs/pi-integration.md), [`docs/grok-integration.md`](docs/grok-integration.md) and [`docs/omp-integration.md`](docs/omp-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
@@ -451,8 +461,8 @@ Run a case inside its own hardened Docker container instead of directly on your
- **Resource templates** — expand the checkbox for a **Small / Medium / Large / GPU** preset (memory, CPUs, GPU), or set your own. **Disk is elastic** — storage grows as data flows in, no fixed cap.
- **Shared per-case container** — many sessions can `docker exec` into the same container; killing one session never tears the container out from under the others.
- **Hardened by default** — non-root, `--cap-drop ALL`, `no-new-privileges`, PID/memory caps, never `--privileged` or the docker socket; a **sealed** profile (no host credentials, network off) is one toggle away.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / Pi logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
-**Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / OMP logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Durable** — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.
Prerequisite: just Docker (or Podman). The agent base image builds itself automatically on first use, with progress streamed to the UI (or pre-build it with `node scripts/build-agent-image.mjs`). Full guide: [`docs/docker-cases.md`](docs/docker-cases.md).
@@ -637,7 +647,7 @@ These run for **every** request — before auth, even on the default no-password
### Input, files & headers
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` env-prefix allowlist gates which settings each CLI can receive
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` / `PI_*` env-prefix allowlist gates which settings each CLI can receive
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
- **Security headers** — `Content-Security-Policy` (`default-src 'self'`, every exception enumerated), `X-Content-Type-Options: nosniff`, `X-Frame-Options: SAMEORIGIN`, HSTS over HTTPS, and CORS reflected **only** for `localhost` / `127.0.0.1` / `::1`
@@ -650,17 +660,19 @@ These run for **every** request — before auth, even on the default no-password
---
## SSH Alternative (`sc`)
## Terminal UI (`codeman tui`)
If you prefer SSH (Termius, Blink, etc.), the `sc` command is a thumb-friendly session chooser:
A full-screen dashboard for your sessions, in the terminal. Same states as the web UI, because it is a client of the same server:
```bash
sc # Interactive chooser
sc 2# Quick attach to session 2
sc -l # List sessions
codeman tui# the dashboard
codeman tui --list # numbered session list, then exit (scriptable)
codeman tui 2# attach straight to session 2 of that list
Sessions are grouped **NEEDS YOU → WORKING → IDLE → RECENT**, longest-waiting first. `↑↓`/`j`/`k` select, `1`-`9` and `[`/`]` switch between sessions, `Enter` attaches into the tmux pane (**`F1`** to come back). Inside a pane the bar across the top keeps the session strip visible and `Alt+1`-`Alt+9` switch without leaving. `y`/`n`/digit answer a pending permission dialog right from the list, `p` sends a one-line prompt, `n` starts a session and opens straight into it, `x` kills one (`y` confirms), `/` searches, `g` shows the away digest, `?` is help, `q` quits. Below 72 columns it drops the preview pane and becomes a single-column list, so it stays usable in Termius on a phone. With no server running it still starts in attach-only degraded mode.
The web UI remains the primary surface; see **[docs/tui.md](docs/tui.md)** for the full guide.
For AI agents and automation that control Codeman without a browser: an agent that spins up worker sessions, a CI bot, or **Claude Code running _inside_ a Codeman session orchestrating other sessions**. Everything the UI does is HTTP + a CLI, so an agent can do it too.
> **Shortcut: install the packaged agent skill.** Everything below (plus worked multi-worker recipes) ships as a Claude Code skill in [`skills/codeman`](skills/codeman/SKILL.md), so an agent inside a session can drive Codeman without you pasting docs into the prompt. Three ways to get it:
>
> - `npx skills add Ark0N/Codeman --skill codeman -g`: global, works for any skills-aware agent
> - `codeman skill install` (global) or `codeman skill install --case <name>`: for npm installs that never cloned the repo; `codeman skill uninstall` reverses it
> - **App Settings → Agents & CLIs → Claude → Agent Skill** (`agentSkillEnabled`, default off): Codeman then injects the skill into each case on Claude session create; a user-authored `skills/codeman` in the case is never overwritten
>
> A global install (`codeman skill install`, or `npx skills add`) is picked up by **every new Claude Code session on the machine**, inside Codeman or not. The skill self-gates: outside a Codeman session (`CODEMAN_MUX` unset) it refuses to act, so a global install costs an idle session nothing.
>
> ⚠️ Turning `agentSkillEnabled` back off **does not remove already-injected copies** (a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` dir). Remove them per case with `codeman skill uninstall --case <name>`.
### The agent skill (start here)
Everything in this section also ships as a **Claude Code skill** in [`skills/codeman`](skills/codeman/SKILL.md). Install it once and you never paste API docs into a prompt again. You ask for what you want in plain English, and the agent already sitting inside a Codeman session loads the recipes and drives the API itself.
| Skills CLI | `npx skills add Ark0N/Codeman --skill codeman -g` | Global, works for any skills-aware agent |
| Claude Code plugin | `/plugin marketplace add Ark0N/Codeman` then `/plugin install codeman@codeman` | Global, through Claude Code's plugin manager; `/plugin update codeman` follows releases. Pick this OR a `codeman skill install`, not both: a Claude Code with both lists the skill twice (`codeman` and `codeman:codeman`) |
| Bundled CLI | `codeman skill install` | Global (`~/.claude/skills/codeman`), for npm installs that never cloned the repo |
| Bundled CLI | `codeman skill install --case <name>` | One case only |
| Web UI | App Settings → Agents & CLIs → Claude → **Agent Skill** | Auto-injects into each case on Claude session create (`agentSkillEnabled`, SYNCED, default off) |
`codeman skill uninstall [--case <name>]` reverses the CLI installs, and never touches a `skills/codeman` you wrote yourself.
#### Step 2: ask for things
That is the entire interface. No curl, no endpoint names, no session ids. These prompts work as written:
| _"What sessions are running right now?"_ | Lists them with name, mode and status. Read-only, safe to ask anytime. |
| _"Start a shell worker on the `myapp` case, run the test suite, tell me if it passes."_ | Spawns, waits on a split completion marker, reads back the exit code, cleans up. |
| _"Spin up 3 workers for lint, typecheck and tests. Run them in parallel, report failures."_ | The fan-out flow: one session per task, all started first, then gathered as each finishes. |
| _"Have a claude worker on `refactor-auth` summarize `src/session.ts`, then close it."_ | Spawns, runs the readiness ladder (first-run trust dialog included), send-and-wait, reads the clean transcript answer, deletes. |
| _"Watch session w4 and tell me if it gets stuck on a permission prompt."_ | Blocks on the `blocked` signal and surfaces the question to **you**. It never answers another session's prompt itself. |
#### Step 3: nothing
The agent deletes every session it started. Watch the tabs appear and disappear in the dashboard while it works.
#### A real run, start to finish
> **You:** spin up 3 shell workers, run lint / typecheck / the frontend syntax check in parallel, and tell me which failed.
```text
lint -> 9f2d8e5f dispatched
typecheck -> aff9c691 dispatched 3 tabs appear in the dashboard
syntax -> be9f1f15 dispatched
lint DONE_lint_17909 rc=0
typecheck DONE_typecheck_3409 rc=0 gathered as each one finishes
Those `DONE_<task>_<random>` strings are the skill's **split marker** trick, and they are why the fan-out is reliable on hook-less `shell` sessions: the typed line contains `${M}_17909`, so only the command's real *output* ever contains `DONE_17909`. An unsplit marker would match the echo of your own keystrokes before the command had even run.
| [`SKILL.md`](skills/codeman/SKILL.md) | Safety rules, the ready-made fast path (spawn N workers, task them, collect), and the verb index. Always loaded. |
| [`reference/verbs.md`](skills/codeman/reference/verbs.md) | The 14 verbs in detail: readiness, send-and-wait, markers, interrupts, cleanup. On demand. |
| [`reference/recipes.md`](skills/codeman/reference/recipes.md) | 6 worked multi-worker flows (fan-out, blocked-worker watch, messaging fan-out). On demand. |
| [`reference/endpoints.md`](skills/codeman/reference/endpoints.md) | Full endpoint tables, error codes, per-mode signal table, capacity limits. On demand. |
| [`reference/messaging.md`](skills/codeman/reference/messaging.md) | Talking to claude workers directly via Claude Code cross-session messaging. On demand. |
Every recipe in there was verified against a live server, and the comments record the failure modes that were measured rather than guessed.
#### Two things worth knowing
- **It self-gates.** Outside a Codeman session (`CODEMAN_MUX` unset) the skill refuses to act and does not guess an API URL, so a global install costs an unrelated Claude Code session nothing.
- **It is deliberately conservative.** Unprompted, it may only spawn sessions, prompt them, and delete ones **it created in that same conversation, by exact id**, through a fail-closed guard that refuses to delete the agent's own session. Deleting a case (which erases a real directory of your code), bulk kills, respawn/ralph/cron/orchestrator changes and settings writes all require you to ask, naming the target.
⚠️ Turning `agentSkillEnabled` back off **does not remove already-injected copies** (a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` dir). Remove them per case with `codeman skill uninstall --case <name>`.
---
**The rest of this section is the manual path**: the same operations as raw HTTP, for a CI bot, a shell script, or any agent without skill support.
### Detect that you're inside Codeman
@@ -723,8 +798,8 @@ When a CLI runs in a Codeman-managed session, these environment variables are se
4.**Response envelope.** Most endpoints return `{ "success": true, "data": … }` (errors: `{ "success": false, "error", "errorCode" }`). A few legacy GETs return bare bodies — **handle both** (`body.data ?? body`).
5.**`/api/v1/*`** is a stable alias of `/api/*`.
6.**Wait instead of polling, and don't treat a timeout as an error.** The wait endpoints answer with HTTP `200` and `wait.timedOut: true` when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. `wait.timeoutMs` tells you the timeout the server actually applied after clamping (600s ceiling).
7.**Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
8.**Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
7.**Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity/pi) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
7.**Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity/omp) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.8.**Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
### Recipes
@@ -824,7 +899,9 @@ codeman session start -d /path/to/repo # (s) start a session
codeman session list # list sessions
codeman session logs <id> # tail output
codeman task add "fix the failing test"# (t) queue a task
codeman attach <path> # attach a Claude hook context
codeman attach <path> # show an attachment card for a local file
codeman tui --list # numbered session list (plain text when piped)
npm run test:ci# Run tests (the CI suite; browser suites need extra setup)
npm test# Run tests (same suite CI runs; browser/mobile/perf suites have their own commands)
```
See [CLAUDE.md](./CLAUDE.md) for full documentation.
---
## Community
Questions, setup help, and ideas live in [GitHub Discussions](https://github.com/Ark0N/Codeman/discussions): the [Q&A section](https://github.com/Ark0N/Codeman/discussions/categories/q-a) answers the most common ones (phone access, overnight runs, updating), and the roadmap gets decided in [Ideas](https://github.com/Ark0N/Codeman/discussions/categories/ideas). Bugs go to [issues](https://github.com/Ark0N/Codeman/issues); reports usually get a response within a day, and every release credits its reporters and contributors by name. Want to contribute? [CONTRIBUTING.md](.github/CONTRIBUTING.md) has the map: skins, translations, and docs make great first PRs, and bigger features start life as a Discussion. And if you're proud of your rig, post it in [Show and tell](https://github.com/Ark0N/Codeman/discussions/300).
---
## Codebase Quality
The codebase went through a comprehensive 7-phase refactoring that eliminated god objects, centralized configuration, and established modular architecture:
This folder contains the Compose configuration, server image Dockerfile, and environment template for a locally built Codeman server.
## Start
From the repository root, create the runtime environment file and set the required values, especially `CODEMAN_PASSWORD`.
```sh
cp docker/.env.example docker/.env
bash docker/Start-Codeman.sh
```
On PowerShell, use the following commands instead. Running Compose from inside `docker/` with no `-f` lets it discover `docker-compose.override.yml` on its own (see [Local customisation](#local-customisation)); naming the file with `-f docker/docker-compose.yaml` from the repository root silently drops the override unless it is named too.
```powershell
Copy-Itemdocker/.env.exampledocker/.env
Set-Locationdocker
dockercompose--env-file.envup--build-d
```
Every required value is defined and explained in `.env.example`. `GEMINI_API_KEY` is intentionally optional and may remain blank.
The container starts as root so `entrypoint.sh` can correct the ownership of a bind source the Docker daemon created (it creates a missing one as `root:root`), then drops to `PUID:PGID` with `setpriv` before the server starts, so Codeman itself never runs privileged. That drop needs `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against the file's `cap_drop: ALL`; a compose file written elsewhere (Unraid's Compose Manager, a hand-written unit) must carry the same additions, and the entrypoint names them when they are missing. A directory owned by neither root nor `PUID:PGID` is never re-owned: it is probed for writability as the runtime account and refused with a clear message if that fails. Setting `user:` in Compose skips the whole step.
On Linux, `Start-Codeman.sh` stops with an error when required paths are missing. It creates the application-data directory when safe, detects its numeric owner as `PUID:PGID`, and detects `DOCKER_SOCKET_GID` from the configured Docker socket. It rejects a root-owned application-data directory because Codeman and its local CLI sessions must remain unprivileged.
Codeman, Claude, OpenCode, and other local sessions run as the unprivileged account named by `CODEMAN_RUNTIME_USER`, which defaults to `codeman`. When Compose is run directly, `PUID` and `PGID` default to `1000:1000`; set them in `.env` when the application-data directory has a different owner. The Bash start script determines them automatically instead.
To retain Docker-case support without root when running Compose directly, set `DOCKER_SOCKET_GID` to the numeric group ID of the host socket. On a standard Linux Docker host, obtain it with `stat -c '%g' /var/run/docker.sock`. The Bash start script detects it automatically.
## Updating
Use **App Settings → Updates** in the web UI. The checkout Compose builds from is
also mounted at `/opt/codeman`, so an update's `git checkout` and rebuild persist
on the host, and the server exiting is what restarts the container onto the new
build.
Releases that change `server.Dockerfile`, `docker-compose.yaml`, or add a key to
`.env.example` cannot be applied that way — the updater detects them, names what
changed, and asks you to run `Start-Codeman.sh` here on the host instead. Details:
Compose merges `docker-compose.override.yml` on top of `docker-compose.yaml`. Keep host-specific changes there rather than editing `docker-compose.yaml`, so this repository can be updated without losing them. Both `docker-compose.override.yml` and `docker-compose.override.yaml` are ignored by Git.
`Start-Codeman.sh` names the Compose file explicitly, which disables Compose's automatic discovery of the override file, so the script adds it back when one is present and prints the file it used. Running `docker compose` from this folder without any `-f` option finds it automatically. When passing `-f docker/docker-compose.yaml` from the repository root, add `-f docker/docker-compose.override.yml` as well, or the override is silently ignored.
An override file adds to and replaces individual settings. It cannot delete a key from `docker-compose.yaml`, and Compose concatenates rather than replaces `ports`, so removing a published port still requires editing `docker-compose.yaml`. The example below replaces the restart policy and adds a mount, leaving every other setting in place:
```yaml
services:
codeman:
restart:always
volumes:
- /srv/projects:/srv/projects
```
### Reverse-proxy host allowlist
Codeman rejects any request whose `Host` header is not on its own allowlist - a
DNS-rebinding guard, not a Compose or Docker concern. Loopback, any IP literal,
the configured `--host`, and a few tunnel-provider suffixes are allowed by
default; a reverse-proxied domain is not, and is rejected with
`403 Forbidden: host not allowed` before the request reaches any handler.
Add the domain with `CODEMAN_ALLOWED_HOSTS` in `.env`:
`docker-compose.yaml` forwards it into the container (Compose only passes
through the environment keys it explicitly lists, and this is one of them, with
an empty default so the line is optional in `.env`).
See the application's own `docs/wiki/Remote-Access.md` for the full allowlist
format and the tunnel providers it accepts by default.
## Application data storage
The default configuration uses a host-folder bind mount:
```yaml
volumes:
- type:bind
source:${CODEMAN_APPDATA_PATH}
target:/home/${CODEMAN_RUNTIME_USER}
```
Set `CODEMAN_APPDATA_PATH` in `.env` to a directory that the Docker daemon can access. The example value is `/mnt/user/appdata/codeman`.
`CODEMAN_CASES_PATH` is the separate host directory for managed case workspaces. It is mounted into Codeman at the same absolute path, allowing the host Docker daemon to bind it into an isolated case container. Set it to a child directory of `CODEMAN_APPDATA_PATH` unless you deliberately store workspaces elsewhere.
Compose also exposes `CODEMAN_APPDATA_PATH` to Codeman as `CODEMAN_DOCKER_HOST_HOME`. This lets Docker case seed files, CLI credentials and the hook secret be mounted using paths that exist in the host daemon's filesystem. Direct host installations do not set this variable and retain their existing behaviour.
Set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` when `docker info` reports `SwapLimit=false`. Codeman continues to apply the configured case memory limit, omits Docker's unsupported `--memory-swap` option, and filters only the daemon's exact swap-capability warning. Every other Docker create error and its exit status remain visible.
For an existing installation created by a root-running image, change ownership of the application-data directory before upgrading so the configured `PUID` and `PGID` can read the saved credentials and state:
```sh
chown -R 99:100 /mnt/user/appdata/codeman
```
Replace `99:100` and the path with the values from your `.env` file.
Do not replace this bind mount with a Docker-managed named volume when Docker cases are enabled. Codeman passes seed, credential, transcript and hook-secret bind sources to the host Docker daemon, so their source files must have stable paths in the daemon's filesystem. A named volume does not provide the required host path mapping.
## Static macvlan networking
The default configuration publishes a host port. It does not use `network_mode: host`. To attach Codeman directly to an existing external macvlan network with a static IP address and MAC address, remove the `ports:` section from `docker-compose.yaml` and add the following to the `codeman` service. The service and network additions can instead be placed in `docker-compose.override.yml`, but the `ports:` removal cannot, as described under [Local customisation](#local-customisation):
```yaml
mac_address:${CODEMAN_MAC_ADDRESS}
networks:
codeman_lan:
ipv4_address:${CODEMAN_IPV4_ADDRESS}
```
Then add this top-level network declaration:
```yaml
networks:
codeman_lan:
external:true
name:${CODEMAN_MACVLAN_NETWORK}
```
Set `CODEMAN_MACVLAN_NETWORK`, `CODEMAN_IPV4_ADDRESS`, and `CODEMAN_MAC_ADDRESS` in `.env`. The values in `.env.example` match the supplied Unraid example network and should be changed for other hosts.
### Create a managed macvlan network
If an external macvlan network does not already exist, use this top-level declaration instead. Do not use it together with the external-network declaration.
```yaml
networks:
codeman_lan:
driver:macvlan
driver_opts:
parent:${CODEMAN_MACVLAN_PARENT}
ipam:
config:
- subnet:${CODEMAN_MACVLAN_SUBNET}
gateway:${CODEMAN_MACVLAN_GATEWAY}
```
Macvlan containers are ordinarily not reachable from their Docker host without additional host-network routing. Confirm the selected address, MAC address, parent interface, and subnet are reserved and valid for the target network before starting the stack.
@@ -60,7 +60,7 @@ Module-level singleton in the style of `session-wait-registry.ts` (pure, no `Ses
Normal authed API (NOT the hook-secret bypass), `ApiResponse` envelope, Zod schemas in `schemas.ts`:
-`GET /api/approvals` → pending items, multi-user filtered by `canAccessOwned` (same policy as session lists).
-`GET /api/approvals` → pending items, multi-user filtered by `canAccessOwned` (same policy as session lists). Also sweeps the caller's own items for staleness through `verifyStillAnswerable()`: Claude Code fires no "permission answered" hook, so a dialog answered in the terminal used to sit pending until `stop` and re-arm a red tab alert on the next page load. Only items whose original frame parsed options can be dropped this way, so an unreadable capture keeps the alert.
-`approve` → `writeViaMux('1')` (option 1 is always plain Yes; no Enter, menus react to the digit).
-`deny` → `writeViaMux('\x1b')` (Esc is the official No/cancel; precedent: auto-resume sends Esc the same way).
@@ -68,6 +68,7 @@ Normal authed API (NOT the hook-secret bypass), `ApiResponse` envelope, Zod sche
-`text` → `idle` items only: single line, embedded newlines stripped, sent as `text\r` (the `\r` discipline from CLAUDE.md).
- Guards: item still pending (404 otherwise), session exists + ownership via `findSessionOrFail`, session mode installs hooks. **Answer-time re-capture**: for items whose frame parsed options, the pane is re-captured before sending; if the dialog no longer parses, the item resolves and the answer is refused with 409 (the keystroke would land in whatever now has focus). Marks `answered` BEFORE the write so a double-tap cannot double-send; rolls back to pending if the write fails.
-`POST /api/approvals/:id/dismiss` → remove without keystrokes.
-`POST /api/approvals/session/:sessionId/viewed` → acknowledge the session's pending **idle** item (`acknowledgedAt`, emitted as `approval:updated`). Added after the owner reported that a yellow tab clicked and checked went yellow again on reload: the view-clears-idle rule lived in one browser's memory, so the seed re-armed it and other devices never saw the clear. Acknowledgement is deliberately **not** resolution (the prompt is still unanswered, so it stays in the inbox and stays available as Read My Mind context), and deliberately **idle-only** (looking at a permission/question dialog does not answer it, so the red alert survives being viewed).
### SSE
@@ -84,7 +85,7 @@ Normal authed API (NOT the hook-secret bypass), `ApiResponse` envelope, Zod sche
New module `approvals-ui.js` (@loadorder 11.2, after panels-ui.js), prettier-formatted (not added to `.prettierignore`).
- **Seed on connect**: `GET /api/approvals` on init and SSE reconnect; each pending item re-feeds `setPendingHook(...)` so tab alerts and the phone overview survive reload (fixes problem 2 with zero changes to the alert state machine).
- **Seed on connect**: `GET /api/approvals` on init and SSE reconnect; each pending item re-feeds `setPendingHook(...)` so tab alerts and the phone overview survive reload (fixes problem 2 with zero changes to the alert state machine). Items carrying `acknowledgedAt` are skipped, and `markIdleAlertSeen()` (app.js) is what sets it: viewing a session clears its yellow locally and POSTs `.../viewed`, so "I checked it" survives the reload and reaches the user's other devices through `approval:updated`.
- **Desktop**: header bell `btn-approvals` with count badge. Ships default-hidden via marker class `btn-approvals--hidden` (same policy as the attachments button, so `test/mobile-header-buttons-policy.test.ts` excludes it from the default-visible enumeration); JS shows it only while count > 0. Click toggles a drawer of cards: session name + kind, tool/message summary, mono context block, buttons rendered from parsed options (else Approve/Deny), plus Dismiss and Open session. Esc closes; existing z-index layers respected.
- **Phone**: header button stays hidden (`mobile.css`); the phone surface is the overview's NEEDS YOU section, whose rows gain inline ✓/✗ buttons for permission items (tap-through to the session remains the row's main action). Toolbar classes/status language rules from the mobile-overview section of CLAUDE.md apply.
- **i18n**: new strings registered in i18n.js (en + zh-CN); status words carry `data-i18n-skip` where they would collide (mirroring the overview pills).
Every run mode Codeman can launch — Claude Code, Terminal/Shell, OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek Harness and OMP — is a `CliEntry`: a data record describing how to find the binary, how to build its command line, what environment it needs, and what it can do. Code that used to ask "which CLI is this?" asks the entry instead.
| `types.ts` | The `CliEntry` interface and everything under it. Read this first. |
| `stock.ts` | The shipped catalog. **The only file allowed to name a CLI id.** |
| `schema.ts` | Zod validation, including the cross-field checks that reject an incoherent entry at LOAD time. |
| `argv.ts` | The argv engine: the only code that turns typed tokens into a command string. |
| `patterns.ts` | The NAMED value patterns (`model`, `uuid`, `path-segment`, …) and the regex-compilation guard. |
| `profiles.ts` | The names of behaviours that genuinely need code, kept import-free so `schema.ts` can validate one. |
| `registry.ts` | Loading, merging `~/.codeman/clis.json`, and the accessors (`getCli`, `enabledClis`). |
`src/session-cli-registry-bridge.ts` maps the legacy per-mode option bag onto the engine, and `src/utils/cli-resolver.ts` / `src/utils/cli-launcher.ts` do registry-driven binary resolution and launcher-profile dispatch.
## The override file
`~/.codeman/clis.json` (instance-scoped through `dataPath()`) holds overrides and custom entries only, never a copy of the stock catalog: `{ "clis": { "<id>": { ...partial entry... } } }`. Objects merge key-wise onto the stock entry, arrays replace wholesale. **The file must be mode 0600**; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you `chmod 600` it. Every reason a file was ignored or an entry dropped is logged once, prefixed `[cli-registry]`, on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read only on restart.
## The shape of an entry
```ts
interfaceCliEntry{
id: CliId;// 'codex'
label: string;// 'Codex' — shown in menus
shortBadge: string;// tab badge, e.g. 'CX'
accent: string;// single hex colour
enabled: boolean;
stock: boolean;// set by the loader; a custom entry can never claim it
order: number;
kind:'agent'|'shell';
discovery: CliDiscovery;// how to find and prove the binary
launch: CliLaunch;// the structured argv template
env: CliEnv;// exports, tmux setenv keys, the env-override allowlist
capabilities: CliCapabilities;// what every call site reads instead of the id
// .workDetect?: { promptGlyph, workingLine } — how this CLI's pane shows work
overlays: CliOverlays;// remote-SSH / Docker pane commands, credential store
}
```
`capabilities` is the important part. It is what `isExternalCliMode()`, `isAltScreenStripMode()`, `hooksAvailableForMode()` and every other former per-mode branch actually read.
### Regexes that come from config
Two capability fields carry a regular expression an override file can set: `discovery.version.regex` and `capabilities.workDetect.workingLine`. Both go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
`workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused.
### Three capabilities that must stay independent
`external`, `hooks` and `altScreen` describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. `shell` has no hooks but is **not** an external CLI, so a hooks predicate written as `!isExternalCliMode()` accepted `until=stop` on a shell session and then blocked the caller for their entire timeout. `deepseek` is the mirror image: it IS external and it DOES have hooks.
`test/cli-capability-predicates.test.ts` asserts that no two of the three are equivalent across the catalog, so collapsing them fails the build rather than a user's session.
## Arg-template safety
The composed command line is interpolated into `bash -c "…"` inside tmux, which makes command construction a security boundary. Four independent layers keep config out of it:
1.**Config contains no shell text.** There is no `command: "..."` field anywhere in the schema. An entry declares a sequence of typed tokens; `argv.ts` is the only place that turns them into a string, and it owns every separator itself — one space between tokens, ` || ` between fallback variants. Neither can originate from config, because config has no field that could hold either.
2.**Every literal is validated at LOAD time** against a safe-word pattern (no space, quote, backtick, `$`, `;`, `&`, `|`, redirection, parens, braces, newline or backslash). A bad literal **rejects the whole entry** rather than being dropped, because a silently dropped flag would change security-relevant behaviour — losing `--no-approve` is not a cosmetic difference.
3.**Values resolve through NAMED patterns.** A value placeholder selects a `TokenPattern` (`model`, `uuid`, `slug`, `path-segment`, `tool-list`, …) from `patterns.ts`; config can never supply its own regex for a value, so a `clis.json` structurally cannot widen its own validation. A value that fails its pattern drops the whole argument, exactly as the hand-written builders did: an invalid `--model` omits `--model`, it never substitutes something else.
4.**Escaping is independent of validation.**`renderToken()` re-checks the resolved value before emitting it unquoted, and single-quotes anything else — so even a value that somehow bypassed validation is quoted, never concatenated raw.
The only config-supplied regexes are `discovery.version.regex` and `discovery.identity.regex`. Both run against **command output** rather than a shell token, both are compiled through `compileVersionRegex()` (length cap, nested-quantifier rejection, never the `g` flag), and the output they see is truncated first.
## Named profiles: the escape hatch
Some differences genuinely need to run code rather than be described. Those are **named profiles**: a capability field holds a profile NAME, and the implementation lives in one place keyed by that name — never by CLI id.
-`discovery.launcherProfile` — for a CLI whose binary is not the agent. `dsh` boots `$DSH_HOME/profiles/<name>`, so "installed" and "runnable" have different answers; the profile answers both, plus why a specifically-named target will not work. Implemented in `utils/cli-launcher.ts`.
-`env.setenvProfile` — per-CLI environment setup that is more than a list of keys, such as DeepSeek's status bridge.
-`capabilities.transcript` — which on-disk history reader understands this CLI (`claude-jsonl`, `codex-rollout`, `deepseek-zstd`, `omp-jsonl`, `none`).
-`capabilities.echo.predictProfile` — the predictive-echo model a composer needs.
The names live in `profiles.ts`, which is kept free of imports so `schema.ts` can validate a name at load time. A profile this build does not implement is a load-time error naming the field, rather than a CLI that silently looks permanently uninstalled.
## DeepSeek: the four assumptions it breaks
DeepSeek is worth reading before assuming an entry looks like its siblings — the schema carries four extensions because of it.
| What it breaks | How the registry expresses it |
| `dsh` is a profile LAUNCHER, not the agent, so "installed" is not "runnable". | `discovery.launcherProfile` + `discovery.launcherTargetParam`. |
| Its permission switch is the **`DSH_PERMISSION_MODE` env var**, not a flag — the harness has none. | `env.configSetenv` (so the ordinary `privilegedParams` clamp still reaches it) **and**`capabilities.privilegedEnvKeys`. |
| It is the only non-claude mode with real hook signals, and for it that is a per-SESSION question. | `capabilities.hooks: 'supervised'` — a third state, not a boolean. |
| Its transcript is zstd session files, one frame per write. | `capabilities.transcript: 'deepseek-zstd'`. |
## Identity probes
`discovery.identity` asks the binary whether it is the program we meant, and it runs **before** the version probe, because a version probe cannot tell an impostor from the real thing. Debian ships an unrelated `dsh` (dancer's shell) that answers `--version` perfectly happily, and npm carries squatters for both `pi` and `grok`.
`discovery.version.requireVersionMatch` is the weaker companion: a binary whose version output has the wrong shape counts as ABSENT rather than present-with-unknown-version. That is what a short, generic binary name needs, and it is what keeps `codeman doctor` and the run mode from telling the user opposite things about the same binary — both read the same regex off the same entry.
## The no-id-branching rule
`test/cli-registry-no-id-branching.test.ts` fails the build if a CLI id comparison appears outside the stock catalog. It builds its id list from the live catalog, blanks comment lines before scanning (comments legitimately quote the pattern to explain why a branch was removed, and blanking rather than dropping is what keeps reported line numbers pointing at the real file), and keeps an allowlist in which **every entry carries its reason**.
It matches four shapes, not one: `mode === '<id>'`, `mode !== '<id>'`, `case '<id>':`, and `['<id>', …].includes(mode)`. The first version matched `===` only, and that gap was not academic — the refactor it guards converted the `===` sites and left the negated ones, so 36 `!==` branches survived it, including a seven-mode chain auto-enabling Ralph under a comment asking the next person to keep it in step with a predicate by hand while the sibling code path already read the capability. A guard that sees half the shapes reports a count measured over the half it happens to catch.
The allowlist is not a formality. If a branch is about what a CLI can DO it belongs in `CliCapabilities`; the entries that remain are things that are not CLI-behaviour branches at all — chiefly the legacy per-mode `<Mode>Config` objects on `POST /api/sessions`, which are a fact about the public HTTP API rather than about any CLI, plus a few documented cases where `mode === 'claude'` is genuinely the right question (Read My Mind reads Claude's _own_ transcript, so a capability there would be actively wrong).
## Two namespaces called `param`
`launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a **launch param**. The **legacy wire field** a param arrives as is a separate namespace, and `launch.legacyConfigAliases` is the only bridge between the two.
This matters because it is invisible when it is wrong. `capabilities.privilegedParams[].param` is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a name from the wrong namespace clamps **nothing**: no load error, no failing test, the clamp simply stops running. Codex is the entry where the two names differ (`bypassApprovals` as the param, `dangerouslyBypassApprovals` on the wire), so it is the one that catches a regression. `schema.ts` rejects any entry naming a param it never declared, on both `configSetenv.fromParam` and `privilegedParams.param`.
## Fields declared for later
`shortBadge`, `accent`, `capabilities.echo`, `capabilities.wheelForward`, `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` are **declared but not yet read**. They all describe frontend behaviour, and the frontend is deliberately untouched here: `app.js`, `terminal-ui.js` and `styles.css` keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
Treat those values as **transcribed, not authoritative** — nothing enforces that `echo.policy` matches `_updateLocalEchoState`'s fallthrough, or that `accent` matches the gradient CSS paints, so re-measure before wiring one up. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; `test/cli-registry-no-id-branching.test.ts` pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
`overlays.credStore` is in the same category, for a sharper reason: the Docker credential-seeding path still reads its own `CRED_STORES` table, because this shape allows ONE store per CLI and the live table needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek declares none here even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is simultaneously the worst thing here to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest.
Everything else in the interface is live, including `overlays.remote` / `overlays.docker`, which back `defaultRemoteCommandForMode()` and `defaultDockerCommandForMode()` directly. Those two used to be hardcoded `Record<…CommandMode, string>` tables duplicating the registry with nothing keeping the two in step; `test/location-overlay-commands.test.ts` pins every resulting command as a literal string.
## Consumers outside the server
Two things need the catalogue but cannot import TypeScript, so `npm run generate:cli-catalog`
(`scripts/generate-cli-catalog.mts`) emits two artifacts from `stock.ts`. Both are committed,
and `test/cli-catalog-sync.test.ts` fails if either drifts from a fresh generation.
| `config/clis.stock.json` | `scripts/lib/cli-catalog.mjs` (Docker build args), tests | A `.mjs` cannot import the registry. |
| a marked block inside `install.sh` | the installer itself | It runs via `curl \| bash` before any checkout exists, so it can read neither. |
Only `id`, `label`, `shortBadge`, `enabled`, `order`, `kind` and `discovery` are exported.
`launch`, `env`, `capabilities` and `overlays` are spawn-time concerns the server alone
interprets, and a test asserts they never leak into the artifact — a second reading of the
launch model in a consumer that cannot be tested against a real spawn is exactly what this
registry exists to prevent.
The install.sh copy is **embedded, not fetched**, and is the FULL catalogue. An earlier design
fetched it and fell back to a hardcoded two-CLI list, which degraded silently on an empty
response; there is no degraded mode to fall into now, and no network fetch either — a `curl |
bash` from master already carries a catalogue exactly as fresh as the script itself, so there is
nothing a refresh would buy that isn't already true. An earlier draft added an opt-in refresh
with a `TRUSTED`/`DISPLAY` array split to keep it from ever writing the executed command; it was
dropped before merge rather than shipped half-verified — the split's only actual write was the
label, `DISPLAY` never diverged from `TRUSTED` in practice, and the added surface (a second
array, a fetch path, three failure shapes to warn on) bought nothing the embedded copy didn't
already have.
### The install-command trust boundary
Three rules, and the middle one is why the embed matters:
1.**The server never executes an entry's `install.command`.** Unchanged, and still enforced by nothing executing it: the field is display text (`CliDiscovery.install.command`).
2.**`install.sh` executes only commands embedded in itself.** Those arrive in the same file, over the same TLS fetch, in the same commit as the `curl \| bash` line that fetched the script — identical trust to the hardcoded vendor one-liners it replaces.
3.**Nothing fetched at install time is ever executed.** There is no second code path that fetches anything after the script itself has been fetched.
That is mechanical rather than a promise. `CLI_INSTALL_CMD_TRUSTED` is written only from the
generated block and is the only array the installer ever runs or displays — there is no second
array a refresh could rewrite, because there is no refresh. `test/cli-catalog-sync.test.ts`
asserts that the embedded commands are exactly the registry's, and
`test/install-sh-invariants.test.ts` that nothing in `install.sh``eval`s.
### bash 3.2
macOS ships bash 3.2 and the documented install is `curl -fsSL <url> | bash` under
`set -euo pipefail`, so a bash-4 construct is not a warning there — it kills the install. The
generated block therefore uses parallel indexed arrays with **offset/length windows** into one
flat array instead of delimiters (a `$HOME` containing a space needs no `IFS` handling, and an
entry with nothing to contribute gets length 0 and is never iterated). CI runs `bash -n` and
executes the script inside a real `bash:3.2` container, because the empty-window case is a
runtime `set -u` abort that `bash -n` cannot see.
## Resolve at call time, never at import
Anything reading the registry must resolve it when it is asked, not when its module is first imported. `sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()` and each resolver's `searchDirs` thunk all re-read the catalog per call.
A module-level const freezes at first import, and the failure is asymmetric: a CLI enabled while the server is running moved the run menu but not the frozen surface, so validation rejected a mode the menu offered, or `codeman doctor` reported a catalog nobody had any more.
## Adding a CLI
1. Add a `CliEntry` to `stock.ts`.
2. Run `npm run generate:cli-catalog` and commit **both** artifacts (`config/clis.stock.json` and `install.sh`). The installer's detection, its install menu, its reminder text and the Docker agent image all follow from that one step — this is what makes upstream `b6d0f1fa` ("wire OMP into install.sh's CLI detection, it had none") impossible rather than merely fixed.
3. Add a golden spawn-command pin to `test/cli-registry-spawn-golden.test.ts`, a row to `test/cli-capability-predicates.test.ts`, its remote/docker commands to `test/location-overlay-commands.test.ts`, and its search paths to `test/install-sh-detection-parity.test.ts`.
4. Only if it cannot install with a plain `npm install -g <pkg>`: give it a layer in `docker/agent.Dockerfile` and set `discovery.install.agentImageLayer: { kind: 'dedicated', reason }` on its entry in `stock.ts`. `test/docker-agent-image-coverage.test.ts` requires both, so an exclusion cannot quietly become an omission. An entry with no `npmPackage` needs only the Dockerfile layer, since it never enters the shared npm layer in the first place.
5. That is usually all. If you find yourself wanting to add an `if` somewhere, the guard test will tell you — and the answer is a capability field, or a named profile if it genuinely needs to run code.
## See also
- [Agent CLIs](wiki/Agent-CLIs.md) — the user-facing per-CLI guide.
-`docs/architecture-invariants.md` — the mechanics and the history behind the rules above.
-`docs/deepseek-integration.md` — why DeepSeek is shaped the way it is.
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` \| `pi` \| `grok` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. ⚠️ A `pi` or `grok` job's readiness poll looks for `❯`/a token count, which neither CLI prints, so it burns the poll budget and then sends the prompt anyway (slower start, still works). |
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not**`PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
| `deepseek` | Reuse the **existing**`DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Confirmed reaching the server, but failing — unresolved.** A real run against llama-swap returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` consistently (confirmed the env vars are read: the request reaches the network rather than failing locally). Root cause not identified — plausible explanation by analogy with codex's Responses-API gap is that `dsh --profile headless` expects DeepSeek's official API response shape/path structure rather than a generic OpenAI-compatible `/v1/chat/completions` endpoint, but this was not confirmed by reading dsh's own bundled source (unlike pi/grok, where that grep resolved the question directly). Documented as best-effort/unknown, not shipped as verified working |
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
Everything web-researched-but-unverified gets implemented but must be
smoke-tested against real installs of those CLIs before being called done —
call this out explicitly when implementing, don't just ship on faith.
**Cloud-endpoint specifics** to keep in mind per recipe above: an Azure AI
Foundry-style endpoint typically wants the API key in an `api-key` header
rather than (or in addition to) `Authorization: Bearer`, and its "model" is
often a deployment name rather than the underlying model family name — the
discovery step (`GET /v1/models`) still works the same way against Azure AI
Foundry's OpenAI-compatible endpoint shape, but a user may need to type the
deployment name manually if it isn't returned as expected.
## Architecture
### 1. Registry: new `capabilities.customModelInjection` field
Extend `src/config/cli-registry/types.ts` / `schema.ts` with a discriminated
- `scripts/test-local-llm-harnesses.ts` (already added, this branch; run via `npx tsx`) — the standalone real-CLI-and-real-endpoint smoke test, supporting any `--base-url` (local or cloud). Dynamic: derives its harness list and every env var/config it injects from the live CLI registry + `buildCustomModelInjection()` rather than a second hand-maintained copy — only the one-shot invocation flags (`ONE_SHOT` table) are CLI-specific info the registry doesn't model and stay hand-maintained
- `docs/custom-model-endpoints.md` (new) + a CLAUDE.md pointer bullet under External CLI modes / envOverrides
## Mock-server validation strategy (CI-runnable, no real CLI binaries needed)
Spawning nine real CLI binaries in CI isn't realistic, and neither the author's
llama.cpp box nor a real cloud subscription can be a CI dependency. So the
injection _logic_ gets a tier of automated coverage that sits between the
pure unit tests and the live manual checks in Verification:
1. **`test/fixtures/mock-openai-server.ts`** — a small in-process HTTP
server (plain `http.createServer`, no external deps, port picked per the
existing `const PORT = 3150+` convention) that:
- Serves `GET /v1/models` → a fixed fake model list (`{data:[{id:'qwen3'},...]}`),
## 2. Shape decisions (why DeepSeek is wired the way it is)
DeepSeek is a ninth run mode. Never a location overlay, never a web tab (the
browser UI is handled separately, §3). Three of its decisions have no precedent
in the six external CLIs before it.
| Question | Decision | Why |
| --- | --- | --- |
| What does a pane run? | `dsh --profile <name>`, profile discovered | **The decision that shapes everything else.** DeepSeek ships `web`, `headless` and `base` — no terminal agent. The interactive front door is always a third-party plugin, so Codeman resolves a binary AND a profile inventory, and "available" means both. `resolveDefaultDeepSeekProfile()` prefers a recognized TUI, then an UNRECOGNIZED profile (anyone can publish an app bundle; a classifier that has not heard of one must not hide it), and refuses `web`/`headless`, which cannot occupy a pane. |
| Which TUI? | none blessed; default for BOOTSTRAP only | `POST /api/deepseek/install-profile` defaults to `@deepseek-harness-tui/dsh-tui` (~27.5k weekly downloads, ~4x the next, MIT, and it speaks the status contract in §2.3), but accepts any npm name and the resolver never assumes that profile exists. Codeman offers a default; it does not pick a winner. |
| Permission bypass | `DSH_PERMISSION_MODE` env export, no flag | The harness has NO command-line permission option; its sandbox/approval rows read one env var with three presets (`read-only` / `workspace-write` / `danger-full-access`, read off `dsh --dump-default-config`). This is the one legitimate exception to the `CLAUDE_CODE_EFFORT_LEVEL` ban: that var hard-locks in-session switching, whereas the harness reads this with `??` as a boot-time DEFAULT, so it stays soft. Exported via `tmux setenv`, never on the command line. The Run button sends `danger-full-access`, matching every sibling Run button. |
| Multi-user clamp branch | only-if-sent, clamped to `workspace-write`, **plus an env-var half** | Omitting the export leaves the harness on `workspace-write`, which still ASKS, so an absent config is already safe (the codex/antigravity/grok shape, not pi's materialize). Clamping to `workspace-write` rather than `read-only` is deliberate: the clamp removes privilege, it must not break a session's ability to edit its own workspace. ⚠️ Unlike every sibling, clamping the CONFIG is only half the gate: the switch is an env var, `DSH_*` is an allowlisted `envOverrides` prefix, and `applyEnvOverrides()` runs AFTER `_configureDeepSeek()`, so `envOverrides: {DSH_PERMISSION_MODE: 'danger-full-access'}` on the same request would land last and win. `clampEnvOverridesForOwner()` drops `DSH_PERMISSION_MODE` and `DSH_HOME` for a non-granted owner (dropping falls through to the clamped export). `DSH_HOME` because it aims the launcher at a profile tree whose plugin code runs at BOOT, before any approval row. |
| `hooksAvailableForMode()` granularity | per SESSION for deepseek, per mode for everything else | `deepSeekConfig.statusReporting: false` disarms the `HERDR_*` export, and the triple is the only reason a dsh session posts anything, so a mode-only answer would accept `until=stop` where nothing can send one — the infinite-wait the predicate exists to prevent. Call sites pass `sessionHookOptions(session)`; the default stays permissive so a forgotten one degrades to the old behaviour. ⚠️ Profile conformance stays unknowable at request time (an unrecognized profile is deliberately launchable), so a non-conforming TUI still times out on an explicit `stop`; the default set keeps `idle`/`exit` for that. ⚠️ The predicate is NOT "is this claude": Read My Mind and intent capture read Claude's transcript and were silently widened by this change, so they compare `mode === 'claude'` directly now. |
| Profile install spawn | own process group, hand-rolled timeout | `dsh plugin add` fans out into package-manager children, and spawn's built-in `timeout` signals only the direct child: survivors keep the inherited stdio pipes open, `close` never fires, and the held-open request leaks with no route-level deadline. `detached: true` + negative-pid SIGTERM→SIGKILL, the same escalation `runGit()` uses for the same reason, plus a last-resort reap for a grandchild that escaped the group. |
| Idle detection | **real hook events via a status shim** | The standout decision. The TUI already reports its lifecycle to a supervising process through a generic env-gated contract inherited from Herdr: `HERDR_ENV=1` + `HERDR_BIN_PATH` + `HERDR_PANE_ID` make it run `<bin> pane report-agent <id> --state idle\|working\|blocked …` on every state change, exit 0 = delivered. `deepseek-status-shim.ts` generates a script into the data dir and points `HERDR_BIN_PATH` at it. So deepseek is the only non-claude mode that passes `hooksAvailableForMode()` — earned by emitting definitive signals, not granted. An interface implementation, not an impersonation: no real `herdr` binary is ever executed, and a TUI that ignores the contract simply falls back to output stabilization. |
| `agent_working` event | new, 157th SSE constant | The one hook event with no Claude Code hook behind it. A harness turn cannot run while its own modal approval is on screen, so "started working" proves a dialog was answered in the terminal. Without it a dsh red alert would survive until the next `stop` — the exact stuck-alert bug the claude path already fixed once, and its pane-capture staleness sweep is Claude-dialog-shaped and cannot help here. |
| Resolver | identity probe THEN version probe | Strictest of the family, and not by preference. `dsh` is not merely a squattable npm name: Debian ships an unrelated `dsh` (dancer's shell, `apt install dsh`) which would answer a version probe convincingly and then be handed a spawn line. `dsh --help` must match `DeepSeek Harness` first. `DEEPSEEK_VERSION_REGEX` keeps the prerelease tail (`0.1.1-rc.2`), since truncating it would report an rc as a release. |
| Env allowlist | `DSH_*` + `DEEPSEEK_*` | `DSH_*` covers the launcher's documented inputs (`DSH_HOME`, `DSH_PERMISSION_MODE`, `DSH_TELEMETRY_MODE`, the `DSH_TUI_*` knobs); `DEEPSEEK_*` is the vendor namespace holding `DEEPSEEK_API_KEY`/`DEEPSEEK_BASE_URL`, same reasoning that admitted `XAI_*` for grok. ⚠️ Pi's lesson repeats exactly: a dsh `settings.yaml` can nominate ANY env var as a provider credential (`apiKeyEnv`), and the allowlist is one GLOBAL list, so admitting those would widen every mode at once. They stay out. |
| Model | NOT a session field | The model is a composition entry (`agent-default-model`) in the profile's config tree, set in `~/.dsh/settings.yaml` + `cordis.patch.yml`. Both create paths deliberately resolve no model for this mode rather than inventing a flag. |
| Alt-screen strip | OUT of `isAltScreenStripMode()` | Third-party fullscreen TUIs with their own scrollback and mouse handling — the opencode case, not the Ink case. |
| Local echo | `'buffer'` via the `_updateLocalEchoState` fallthrough | UNMEASURED against a live authenticated session (see §5), same honest gap grok shipped with. The leading TUI's composer supports `@` completion and history search, which *may* make it per-keystroke reactive like codex; if so the fallback is the `'off'` branch. |
| Docker | image installs dsh AND a profile | Profiles are deliberately NOT seeded from the host: each is a per-profile `node_modules` tree, host-arch-specific and far too large to copy per container start. Only `~/.dsh/.env`, `settings.yaml`, `cordis.patch.yml` are seeded (auth + model composition). The profile install rides the `useradd` layer so the closing `chgrp`/`chmod g=u` covers it, which is what keeps it usable under the arbitrary uid the container runs as. |
| Remote SSH | `exec "$SHELL" -i -l -c 'dsh'` | Boots the remote box's default profile; a remote with several needs the per-host `commands.deepseek` override, since `deepSeekConfig` does not cross ssh. |
## 3. The web profile
The browser UI is the only interactive surface DeepSeek ships itself, so it gets
a **shortcut, not a run mode**: `Run ▸ DeepSeek web UI…` starts
`dsh web --no-open --host 127.0.0.1 --port <free> --trusted-host <codeman-authority>`
as a background process and opens the URL as an ordinary web tab.
The server was a **shell session** first, on the reasoning that Codeman already
supervises those (visible, scrollable, killable, dies with its tab) so nothing
new had to own a long-lived HTTP server. That version worked and was still
wrong in use: clicking "open the DeepSeek web UI" put a terminal tab on screen
next to the web tab actually asked for, every single time, and after the first
launch the terminal was pure noise. Opening a dashboard should open one tab.
So `POST /api/deepseek/web` owns it instead (`src/deepseek-web-server.ts`), and
what the session gave away for free is now explicit: exactly one server, reused
rather than raced on a second click; restarted when the requested authority
changes; killed on Codeman shutdown (a detached child would otherwise hold its
port against the next start — the very EADDRINUSE this feature already got
wrong once); and boot output captured, since with no shell tab there is nowhere
else for a stack trace to land. It is fenced at the same bar as the profile
installer: booting a dsh profile executes the plugin code in it, so it requires
the privileged grant in multi-user mode.
`--trusted-host` is load-bearing — dsh fences its `/api` behind a browser-trust
check on the request authority, and a Codeman web tab reaches it through
Codeman's own origin via the webview proxy, not directly. The authority comes
from the CLIENT (`location.host`) because only the browser knows which of a
multi-homed Codeman's origins is actually in play.
Three things about this shortcut are load-bearing and each came from it failing
in exactly that way against a real install:
- **The port is chosen, never hardcoded.**`GET /api/deepseek/web-port` walks
3080..3119 for a free loopback port. 3080 is dsh's own default, which makes it
precisely the port a DeepSeek user is most likely to already be serving on:
binding it unconditionally killed the launch with `EADDRINUSE` against the
user's own `dsh web`.
- **The tab is opened only after the server answers.** The launch polls
`POST /api/webviews/probe` until the URL responds, so a server that dies on
startup reports the failure and points at its shell tab, instead of silently
persisting a dashboard aimed at nothing.
- **The saved tab is `trusted: true`, and must be.** An untrusted webview is
sandboxed without `allow-same-origin`, which breaks this dashboard twice: the
dsh client-runtime reads `localStorage` while loading plugins and dies there,
and an opaque-origin frame sends `Origin: null`, so dsh's trust check 403s
every `/api` call regardless of what `--trusted-host` names. Passing
`location.host` only means anything once the frame actually carries that
origin. The trade is real — a trusted proxied frame is same-origin with
Codeman and can reach Codeman's API — and is defensible only because this
particular dashboard is an agent harness Codeman just started itself on
loopback, which can already run code as the user. It is not a precedent for
trusting third-party dashboards generally.
The record is marked `managed: 'deepseek-web'`, which keeps it out of the
saved-dashboard list: the shortcut that maintains it is already a menu entry, so
listing both showed the same dashboard twice. Being managed is also what lets a
relaunch repoint the existing row instead of stacking one dead dashboard per
restart, since the port is now chosen per launch.
The authority baked into `--trusted-host` is the one the launch was clicked
from, and reuse is conditional on it: a running server fenced for a *different*
origin is stopped and restarted rather than reused, because reusing it renders a
page whose every API call 403s — which reads as a broken dashboard rather than a
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` all work inside the container.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` / `pi` / `grok` / `deepseek` / `omp` all work inside the container.
## One-time setup: build the base image
@@ -21,14 +21,65 @@ The image is **secret-free**: credentials are delivered at runtime (bind mounts
node scripts/build-agent-image.mjs --no-cache
```
### Which CLIs the image contains
The npm-published CLIs come from `ARG CLI_NPM_PACKAGES`, which `scripts/build-agent-image.mjs`
fills from `config/clis.stock.json` (generated from `src/config/cli-registry/stock.ts`). Adding
a stock CLI that installs with a plain `npm install -g` needs no Dockerfile edit. The ARG
defaults to the same list in the same order, so a bare `docker build` produces a byte-identical
layer — a different order would be a different `RUN` string and so a needless cache miss.
⚠️ It reads the **stock** catalogue, never the merged registry. A user's `~/.codeman/clis.json`
must not change what is inside an image tagged `codeman/agent:base`, or two machines holding
that tag hold different images and every cache decision downstream is a lie. Each entry's
`enabled` flag IS honoured, so a CLI that ships disabled is never baked in.
Five CLIs keep hand-written layers, for two different reasons that are easy to conflate.
`antigravity`, `grok` and `omp` declare no `npmPackage` at all, so they never enter the shared
npm layer and each gets a vendor-installer layer instead. `pi` and `deepseek` ARE on npm but
carry `discovery.install.agentImageLayer` in `stock.ts` (a REGISTRY field, rather than an
id-keyed table duplicated between the two producers of the image's build args), which pulls
them out of the shared layer because a plain `npm install -g` is not enough for them:
| `pi` | Installs with `--ignore-scripts`, kept in its own layer so the flag cannot leak to the others. |
| `deepseek` | Needs `pnpm` alongside it (`dsh plugin`, issue #352) plus a `dsh-tui` profile install. |
| `antigravity` | Not on npm — Google ships a standalone binary (~190MB, the largest layer). |
| `grok`, `omp` | Not on npm — standalone vendor installers. |
`test/docker-agent-image-coverage.test.ts` requires every special case to carry a written
reason AND still be present in the Dockerfile, so an exclusion cannot silently become an
omission — which is the same failure upstream `b6d0f1fa` hit in `install.sh`.
Two things build this image: `scripts/build-agent-image.mjs` (a human) and
`ensureAgentBaseImage()` in `src/docker-hosts.ts` (the app, on the first Docker case). They
assemble the argv independently, because a `.mjs` cannot import TypeScript, so
`test/agent-image-build-args-parity.test.ts` pins them together. Without it, an image built by
hand and one built by the app could hold different CLIs under the same tag.
A zero exit code only proves the layers ran, not that the toolchain works. Verify by actually executing each CLI in the image, and check the build log for `Using cache` lines:
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
'for c in claude codex gemini opencode agy pi grok dsh omp; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
Antigravity (`agy`) is the one CLI not installed from npm (Google ships a standalone binary), so it has its own Dockerfile step and adds roughly 190MB; a full image lands near 1.6GB.
⚠️ `dsh --version` is the one line above that answers a different question than the
others: `dsh` is a profile launcher, so a working binary says nothing about whether
the image can actually run a DeepSeek session. Check the profile the Dockerfile
installs into the agent's HOME as well, or a `mode: 'deepseek'` case starts a pane
that dies on arrival:
```bash
docker run --rm codeman/agent:base ls ~/.dsh/profiles/dsh-tui/package.json
```
Building that profile is also why `pnpm` is in the image: `dsh plugin` forwards straight to a literal `pnpm` and exits 127 without it (issue #352), and pnpm — unlike npm — blocks dependency lifecycle scripts by default and fails the install over it, so the profile step passes `--config.dangerouslyAllowAllBuilds=true`.
Antigravity (`agy`) and Grok (`grok`) are the two CLIs not installed from npm (Google and xAI ship standalone binaries), so each has its own Dockerfile step, adding roughly 190MB and 160MB respectively. Pi also gets its own step, because upstream documents installing it with `--ignore-scripts` and that flag must not silently change how the other npm CLIs install.
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md). Grok is seeded per-file for the same reason (`auth.json`, `config.toml`, `pager.toml` out of `~/.grok`, which also holds `sessions/`, `memory/` and the ~160MB binary under `downloads/`), with the same consequence for `grok -c`. See [`grok-integration.md`](./grok-integration.md). OMP is the one CLI in this family where `sessions/` is the EXCEPTION rather than the rule: `~/.omp/agent/{config.yml,mcp.json,models.yml,settings.yml}` are seeded per-file (the dir also holds SQLite caches and `terminal-sessions/`), but `~/.omp/agent/sessions/` is shared RW like codex's, not seeded, because Codeman reads it host-side for history recovery and `--resume` pinning. See [`omp-integration.md`](./omp-integration.md).
## Quickest path: one-click "Run in Docker"
@@ -64,6 +115,49 @@ curl -X POST localhost:3000/api/cases/docker-link -d '{"name":"sandbox","hostId"
curl -X POST localhost:3000/api/quick-start -d '{"caseName":"sandbox","mode":"claude"}'
```
## Attach to a container you already run
The tab's **Attach to an existing container** toggle points a case at a container **you**
built and run. Codeman only ever `docker exec`s into it: it never creates, starts, stops,
restarts or removes it, and it seeds no credentials into it, so the CLIs inside must already
be installed and logged in. A missing or stopped container is an error to report, not a state
to fix — start it yourself and reopen the session.
- **Container Name** is a picker over the engine's containers that you can also type into
(the engine may be remote, or the container may not exist yet when you fill the form).
Stopped containers are listed too, sorted last and labelled, so "mine isn't here" is never
a dead end.
- **Container Workdir** is a path that must already exist **inside** the container. Adoption
mounts nothing, so it need not match the host workspace path; **Browse** lists directories
inside the container itself. Without this check, a wrong path fails at launch as a bare
`execvp failed` inside the pane.
- **Workspace Path** is still a real host directory. It backs file previews, attachments and
watchers exactly as it does for an owned case, but here it is only a mirror: nothing is
bind-mounted, so point it at whatever host directory your container already exposes.
- **Check container** runs a read-only preflight and reports what is inside before you commit
to a case name (running or not, tmux present, which CLIs resolved).
- **Run modes come from the container**, not the host: a host with no `claude` still offers
Claude if the container ships it, and a mode the container lacks is hidden.
- Claude is launched **without**`--dangerously-skip-permissions` when the container's exec
user is root, because Claude Code refuses that flag as root and the refusal is only visible
inside the container.
- Image, network and resource settings disappear from the form: they describe a
`docker create` that adoption never runs.
Recreate is refused for an adopted case, full-image export is refused (it would commit a
container that is not ours), unlinking the case leaves the container running, and the boot
reaper skips it. Workspace-only export still works and never pauses the container.
Equivalent API:
```bash
curl -X POST localhost:3000/api/docker-cases/adopt-preflight -d '{"hostId":"local","container":"my-dev-box","containerWorkdir":"/workspace"}'
curl -X POST localhost:3000/api/cases/docker-adopt -d '{"name":"devbox","hostId":"local","container":"my-dev-box","hostWorkspacePath":"/home/you/projects/devbox","containerWorkdir":"/workspace"}'
```
In multi-user mode adoption is **admin-only**, unlike `docker-link`: an adopted container's
mounts belong to whoever built it, so one mounting `/` would hand the adopter the whole host.
## Lifecycle
- **Reconnect after a Codeman restart** lands back in the same live agent (the in-container tmux survives).
This configuration builds the Codeman application image locally from this checkout. It does not download or depend on a pre-built Codeman image.
For the Compose configuration, environment settings, storage migration, and macvlan networking examples, see the [Docker deployment guide](../docker/README.md).
The image includes Claude Code, Codex, Gemini CLI, and OpenCode. Authenticate a CLI from its Codeman session; credentials are never baked into the image.
## Prerequisites
- Docker Engine or Docker Desktop with Docker Compose v2
- A reachable Docker daemon
The application container mounts the Docker daemon socket so Codeman can create and manage its isolated Docker cases. Treat anyone who can administer this Compose project as having Docker-host-equivalent access.
## Start
Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes.
```sh
cp docker/.env.example docker/.env
```
On PowerShell, use the following command instead.
```powershell
Copy-Item docker/.env.example docker/.env
```
On Linux, run the stack with the start script. It determines `PUID` and `PGID` from the owner of `CODEMAN_APPDATA_PATH`, and `DOCKER_SOCKET_GID` from the configured Docker socket, before invoking Compose. A root-owned application-data directory is rejected so the runtime account cannot become UID 0.
```sh
bash docker/Start-Codeman.sh
```
On other platforms, run Compose directly. `PUID` and `PGID` default to `1000:1000`; set them in `docker/.env` when the application-data directory has a different owner. Naming the file with `-f` disables Compose's own discovery of `docker/docker-compose.override.yml`, so add a second `-f` for it when you keep one (see `docker/README.md`, Local customisation).
```sh
docker compose --env-file docker/.env -f docker/docker-compose.yaml up --build -d
```
The container starts as root, corrects the ownership of a bind source the daemon had to create, and drops to `PUID:PGID` with `setpriv` before Codeman starts; the capabilities that needs are declared in `docker/docker-compose.yaml` and named by the entrypoint when a compose file written elsewhere lacks them.
Open `http://localhost:3000` and sign in with the username and password from `docker/.env`.
## Operations
The local image is tagged `codeman:local` by default. Change `CODEMAN_IMAGE` in `docker/.env` if a different local tag suits your environment.
docker compose --env-file docker/.env -f docker/docker-compose.yaml down
```
`CODEMAN_APPDATA_PATH` holds Codeman state and survives container recreation. Remove that host directory only when deliberately resetting the installation.
`CODEMAN_CASES_PATH` must be an absolute path on the Docker host. Compose mounts it at the same path inside Codeman, so the host daemon can bind the managed workspace into isolated Docker cases. Do not set it to `/home/${CODEMAN_RUNTIME_USER}/codeman-cases`.
Compose passes `CODEMAN_APPDATA_PATH` into Codeman as `CODEMAN_DOCKER_HOST_HOME`. Codeman uses that value to translate generated Docker seed, credential and hook-secret bind sources from the container's home path into paths visible to the host Docker daemon.
If `docker info` reports `SwapLimit=false`, set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1`. Isolated cases retain their configured memory limit. Codeman omits the unsupported swap-limit option and filters only the daemon's exact swap-capability warning while retaining every other Docker create error.
If that directory was created by an earlier root-running image, change its ownership to the configured `PUID:PGID` before starting this version. This preserves existing CLI credentials and session state while allowing the unprivileged runtime account to use them.
## Updating
Codeman updates itself from **App Settings → Updates**, as it does on a bare host. The checkout mounted at `/opt/codeman` is the same directory Compose builds from, so the update's `git checkout` and rebuild land on the host and survive container recreation; the restart is the server exiting, which `restart: unless-stopped` turns into a relaunch on the new build.
That applies application code only. A release that changes `docker/server.Dockerfile`, `docker/docker-compose.yaml`, or adds a key to `docker/.env.example` needs the image rebuilt or the container recreated, which a container cannot do to itself. The updater detects each case and refuses with a message naming what changed; run `docker/Start-Codeman.sh` on the host to apply those.
`CODEMAN_REPO_PATH` overrides which checkout is mounted. It defaults to the compose project's parent directory, so it normally needs no setting. Point it at a directory that is not a git checkout and in-app updates are reported as unavailable.
Full detail, including the fingerprint baseline and the troubleshooting table: [`docker-self-update.md`](docker-self-update.md).
## Docker cases
The default socket path is `/var/run/docker.sock`, which works with a standard Linux Docker Engine. The Bash start script detects its numeric group ID. When running Compose directly, set `DOCKER_SOCKET_GID`, for example using `stat -c '%g' /var/run/docker.sock`, so the unprivileged `CODEMAN_RUNTIME_USER` account can create Docker cases. Docker Desktop users should set `DOCKER_SOCKET` in `docker/.env` only when their Docker installation exposes a different compatible socket path.
Codeman Docker cases are sibling containers on the host daemon, not children of the application container. The Compose configuration handles their workspace bind mount through `CODEMAN_CASES_PATH`; the `/home/${CODEMAN_RUNTIME_USER}` application-data mapping is for Codeman state and ordinary in-container sessions, not sibling-case workspaces.
| Repo bind mount at `/opt/codeman` | Makes the pull persistent. Without it, self-update is unavailable. |
| `codeman-node-modules`, `codeman-dist` volumes | Container-owned build artefacts, layered over the bind mount. |
| `CODEMAN_IN_CONTAINER=1` | Tells `detectSupervisor()` to restart by exiting. |
| `restart: unless-stopped` | Turns that exit into a restart. Verified before every update. |
| `CODEMAN_RESTART_BY_EXIT=1` | The Compose file's declaration of that policy, so the updater may exit even with no Docker socket. |
| Toolchain + devDependencies in the image | Lets `npm install` and `npm run build` run inside the container. |
| `docker-env-applied.json` | Fingerprint baseline, written by `Start-Codeman.sh` on every start. |
| `docker-build-source.json` | What HEAD/`package-lock.json` the build artefact volumes currently reflect. Written by both `Start-Codeman.sh` and this in-place update, so the two agree on whether those volumes are stale. |
### Why build artefacts are in named volumes
`node_modules` and `dist` are mounted as named volumes **on top of** the repo bind
mount. Without that, an update's `npm install` would write into the host checkout,
leaving container-compiled native modules (node-pty builds from source here) in a
directory that may also be used to run Codeman natively, and leaving `git status`
permanently noisy.
Docker seeds an empty named volume from the image, so the first start inherits the
image's already-built `node_modules` and `dist` and pays no bootstrap cost.
`docker compose down -v` is the supported reset: the next start re-seeds them.
That seeding-only-while-empty behaviour has a second, less obvious edge: it also
means a plain `docker compose build` triggered from OUTSIDE the container (for
example `Start-Codeman.sh`, after a `git pull` done by hand rather than through
this in-app updater) produces a fresh image whose freshly-built `dist`/
`node_modules` then sit unused behind the volumes' OLD content — the container
comes back up looking unchanged. `Start-Codeman.sh` detects this by comparing the
checkout's current HEAD and `package-lock.json` hash against `docker-build-source.json`,
and clears just the affected volume(s) before its own `--build` if they moved.
This in-place update writes that same file after a successful build precisely so
that comparison does not fire on stale information: without it, the next plain
`Start-Codeman.sh` run would see the HEAD this update just checked out, not
recognise it as already accounted for, and wipe the volumes this update just
correctly rebuilt right back to the OLDER image.
### Why the runtime image carries a build toolchain
`npm run build` is `tsc` plus `esbuild`, both devDependencies, so the image no
longer runs `npm prune --omit=dev`. And `npm install` may rebuild node-pty, which
ships no Linux prebuild, so `python3`, `make` and `g++` are installed as well.
This is the real cost of in-place updates: a noticeably larger image than a
runtime-only one. It buys an update that takes about a minute instead of a full
image rebuild, and it is why `NODE_ENV=production` is paired with an explicit
`npm install --include=dev` in the updater.
## The environment gate
An in-place update applies **code only**. A restarted container reuses its existing
image and configuration, so a release that changes the environment cannot take
effect that way — and would half-apply: new code against an old environment. The
updater therefore checks the **target release's own files**, read straight out of
git with `git show <tag>:<path>` before anything is checked out.
### 1. `server.Dockerfile` changed, so the image must be rebuilt
Compared by sha256 against the fingerprint `Start-Codeman.sh` recorded when the
running container was built.
### 2. `docker-compose.yaml` changed, so the container must be recreated
Same mechanism. A restart cannot pick up a new mount, port or environment
variable; only recreating the container can.
### 3. `.env.example` gained keys your `.env` has no value for
The check that matters most, because **Compose will not tell you**. An unset
`${VAR}` interpolates to the empty string; Compose prints a warning to a terminal
nobody is watching and starts anyway. A new required setting therefore arrives as
a silently blank environment variable and misbehaves later, far from the cause.
The updater names the missing keys instead.
Commented-out lines in `.env.example` are deliberately *not* keys — that is how
the file marks optional overrides such as `# PUID=1000`, and counting them would
block updates on settings you are meant to leave alone.
### 4. A restart policy that would not bring the container back
Before signalling the server, the updater asks the Docker daemon for its own
container's restart policy. If it is `no`, the update is refused: applying it
would take Codeman down and leave no UI to recover from.
If the policy cannot be read at all (no Docker socket mounted) the update is
still allowed, but the final step changes: the server exits only when the
Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (the shipped one does, because
it is the file that sets `restart: unless-stopped`) or the daemon confirmed an
auto-restart policy. Otherwise the build completes and the panel asks you to
restart the container by hand. A container started by plain `docker run` with no
restart policy therefore gets a staged update, never an outage.
### What the gate deliberately does not do
Every unknown fails **open**:
- A missing fingerprint baseline (a container started before this feature existed)
is not treated as a change, or those installs could never update at all.
- An unreadable `.env`, an unreachable Docker socket, or a target tag whose files
cannot be read all yield "no blocker" rather than a refusal.
The one place an unknown does NOT fail open is the kill itself: with neither the
Compose declaration nor a daemon answer, the updater stages the build and asks
for a manual restart rather than exiting a server nothing may bring back.
The gate catches a specific, detectable class of mistake; it is not a last line of
defence. It is also re-evaluated server-side on `POST /api/system/update`, so
hiding the button in the UI is a courtesy rather than the control.
## The one residual risk
The gate is derived from the diff, so it cannot see a release that needs a newer
environment **without changing any of those files** — for example, code that
depends on newer agent-CLI behaviour.
That is why the four global CLIs in `server.Dockerfile` are **pinned**. Unpinned,
the versions a user ends up with are a function of when their image was built
rather than of any commit, and in-app updates make rebuilds rarer, which makes
that drift worse over time. Pinned, "this release needs a newer CLI" becomes a
Dockerfile change, which check 1 already detects. Bump them deliberately, as part
of a release.
The complementary merge-side guard is `test/docker-compose-env-parity.test.ts`,
which fails CI when a variable is added to `docker-compose.yaml` without an entry
> Template: the pi integration (`c5b5963`, [`pi-integration-plan.md`](./pi-integration-plan.md)),
> which was itself calibrated against the four follow-up commits the antigravity
> integration needed. All of grok's facts below were verified against **grok 1.0.5**
> (`grok 1.0.5 (5115b46bc9)`), installed live during the work.
## 1. What Grok Build is
[xai-org/grok-build](https://github.com/xai-org/grok-build) is xAI's coding agent: a
Rust fullscreen-TUI binary named `grok`, installed by
`curl -fsSL https://x.ai/cli/install.sh | bash` into `~/.grok/bin` (with symlinks into
`~/.local/bin`; the installer also ships an `agent` alias). Config lives in
`~/.grok/config.toml`, TUI appearance in `~/.grok/pager.toml`, credentials in
`~/.grok/auth.json` (0600), sessions under `~/.grok/sessions/`. Auth is browser OAuth
on first launch, `grok login --device-auth` for SSH boxes, or `XAI_API_KEY` for
headless use. It has Claude-style permission modes (`default`/`acceptEdits`/`auto`/
`dontAsk`/`bypassPermissions`/`plan`), allow/deny rules, hooks, MCP, subagents, and a
headless `-p` mode.
## 2. Shape decisions (why grok is wired the way it is)
Grok is a seventh run mode, alongside Claude Code, shell, OpenCode, Codex, Gemini,
Antigravity and Pi. Never a location overlay, never a web tab. Its wiring mixes two
existing shapes:
| Question | Decision | Why |
| --- | --- | --- |
| Permission bypass | `GrokConfig.alwaysApprove` -> `--always-approve` | Grok's real flag (verified via `--help`): "Auto-approve all tool executions", i.e. its `bypassPermissions` mode. Config-level deny rules still apply on top. The Run button sends `true`, matching `runAntigravity()` and Claude's own `--dangerously-skip-permissions` default: Codeman sessions exist for autonomous work. |
| Multi-user clamp branch | only-if-sent (codex/antigravity branch) | A bare `grok` spawn is grok's own ask-mode default, which is already safe, so the clamp only needs to force a SENT `alwaysApprove` off. Contrast pi, whose absent default is an answerable prompt and therefore needs the materialize branch. Cron needs nothing for grok for the same reason (`clampCronExternalCliConfigs`). |
| Alt-screen strip | OUT of `isAltScreenStripMode()` | Grok is a fullscreen alternate-screen TUI with mouse support (its own scrollback pane, `pager.toml [terminal] alt_screen`), i.e. the opencode case, not the Ink repaint case. It falls through to the narrow tmux-attach strip like opencode/antigravity/pi. |
| Resolver | version probe, like pi | `grok` has npm squatters (the unrelated `@vibe-kit/grok-cli` installs a `grok` bin). Candidates must pass `grok --version`; `GROK_VERSION_REGEX` is exported and shared with the dependency registry so doctor and run mode cannot disagree. The probe cannot tell two version-printing `grok`s apart, so `GET /api/grok/status` surfaces path AND version. Search dirs: `~/.grok/bin` first (installer target), then `~/.local/bin`, `/usr/local/bin`, `~/bin`. |
| Env allowlist | `GROK_*` + `XAI_*` prefixes | `GROK_*` covers grok's documented inputs (`GROK_HOME`, `GROK_CONFIG`/`GROK_CONFIG_PATH`, `GROK_MEMORY`, `GROK_WORKFLOWS`, `GROK_SANDBOX`, `GROK_OIDC_*`, `GROK_AUTH_PROVIDER_COMMAND`). `XAI_*` is xAI's vendor namespace and carries `XAI_API_KEY`, grok's documented headless auth var: the same narrow-vendor-namespace reasoning that admitted `GOOGLE_*` for gemini. Foreign provider keys stay out, as always. |
| Resume | `--resume <id>` / `--continue`, id-regexed | Grok's `--resume` also matches session TITLES (arbitrary user strings, case-insensitive). The `^[a-zA-Z0-9._-]+$` regex doubles as the no-titles rule, so nothing free-form can reach the `bash -c` spawn line. A valid explicit id wins over `-c`, mirroring pi. |
| Local echo | `'buffer'` via the `_updateLocalEchoState` fallthrough | UNMEASURED against an authenticated session (see §4). If grok's composer turns out per-keystroke reactive like codex's, the fallback is one `'off'` branch; teaching `PredictiveEchoAddon` grok's composer row is the larger follow-up. |
| Truecolor | `COLORTERM=truecolor` + `unset NO_COLOR` | Rust TUI with themes; joins the codex/gemini/antigravity/pi list in `buildEnvExports()` and `buildMuxAttachEnv()`. |
| Docker credentials | per-file seed: `auth.json`, `config.toml`, `pager.toml` | `~/.grok` also holds `sessions/`, `memory/`, `completions/`, `docs/` and the ~160MB binary under `downloads/`; a whole-dir seed would copy all of it on every container start. Same trade-off as pi: in-container sessions are invisible host-side, so `grok -c` in a Docker case sees only that container's history. |
| Docker install | own Dockerfile step | Not an npm package. xAI's installer has no `--dir` override, so the step copies `/root/.grok/bin/grok` (through the symlink, `cp -L`) into `/usr/local/bin` and removes root's `~/.grok` in the same layer. |
| Remote SSH | `exec "$SHELL" -i -l -c 'grok'` | sshd's remote-command PATH does not include `~/.grok/bin`; same login-shell fix as every other agent CLI. |
| What is NOT wired | `--permission-mode`, `--allow`/`--deny`, `-p` headless, `--worktree`, `--sandbox`, `--reasoning-effort`, `-s/--session-id`, `--fork-session`, `--agent`, `--output-format` | Follow-ups. The flag surface is kept minimal on purpose; grok is pre-1.0-style fast-moving and every flag added is a flag validated forever. |
| npm package | `@earendil-works/pi-coding-agent`, latest **0.84.1** (2026-08-07; 0.84.0 was 2026-08-06); `legacy-node20` dist-tag at 0.74.2 |
| Install | `npm install -g --ignore-scripts @earendil-works/pi-coding-agent`, or `curl -fsSL https://pi.dev/install.sh \| sh` (the curl installer also goes through global npm, so both uninstall via npm) |
| Config dir | `~/.pi/agent` (override: `PI_CODING_AGENT_DIR`). Holds `auth.json`, `trust.json`, `settings.json`, `models.json` (user-defined providers), `models-store.json` (cached catalogs), `keybindings.json`, `extensions/`, `skills/`, `prompts/`, `themes/`, `AGENTS.md`, `SYSTEM.md`, and the package trees `npm/` + `git/` |
| Sessions | `~/.pi/agent/sessions/--<cwd with / replaced by ->--/<timestamp>_<uuid>.jsonl`, tree-structured (`id`/`parentId`), format v3. Overrides: `PI_CODING_AGENT_SESSION_DIR`, `--session-dir` |
| Credentials | `~/.pi/agent/auth.json` (OAuth subscriptions + API keys, auto-refresh), plus ~34 provider env vars with **no common prefix**. 0.84.1 adds `pi auth check` (auth preflight with optional credential output) |
| TUI | Default: **main screen with terminal-owned scrollback**. Since **0.84.0** an experimental fullscreen mode exists, selectable via `--tui-mode fullscreen`**or at runtime through `/settings`**; the default remains the main-screen mode |
| Providers | 15+ (Anthropic, OpenAI, Google, Azure, Bedrock, Mistral, Groq, xAI, OpenRouter, Copilot, Baseten since 0.84.0, ...). OAuth subscription login via `/login` for six: ChatGPT Plus/Pro, Claude Pro/Max, GitHub Copilot, xAI, OpenRouter, Radius |
| Permission model | **No permission prompts at all.** No built-in sandbox, no MCP (none planned), no sub-agents, no plan mode, no to-dos, no background bash. Tools run with the user's own permissions |
| Trust model | "Project trust" gates **loading** of project-local `.pi/` config/extensions/skills and **installing missing project packages**, not tool execution. Triggered only when the cwd (or an ancestor) contains `.pi/settings.json`, `.pi/extensions\|skills\|prompts\|themes`, `.pi/SYSTEM.md`/`.pi/APPEND_SYSTEM.md`, or `.agents/skills`; a bare `.pi/` directory does NOT prompt. Global `defaultProjectTrust`: `ask` (default) / `always` / `never` |
Three consequences shape the whole integration:
1. **There is no `--dangerously-skip-permissions` analog and none is needed.** Pi never prompts for
tool approval. The Claude/Codex/Gemini/Antigravity pattern of "send the bypass flag so the session
is not stuck on a modal" does not apply. Codeman must not invent a flag here.
2. **The one privileged knob is `--approve` / `-a`** (trust project-local files for this run), which
makes pi load and execute project `.pi/extensions` TypeScript **and run an npm install of missing
project packages**. That is the field the multi-user clamp has to cover. Its explicit inverse
`-na` / `--no-approve` exists, which lets the clamp force-deny rather than merely omit (§3, §5.2).
3. **Provider keys cannot ride the env allowlist.** Pi's provider key vars (`ANTHROPIC_API_KEY`,
`OPENAI_API_KEY`, `DEEPSEEK_API_KEY`, `HF_TOKEN`, `BASETEN_API_KEY`, ...) share no prefix, so
there is no way to admit them through `ALLOWED_ENV_PREFIXES` without widening the list for every
mode (§2.4).
---
## 2. Design decisions
### 2.1 Mode identity
`SessionMode` gains `'pi'`. Not a location overlay (unlike Docker/remote-SSH cases), not a web tab:
a real sixth CLI backend with its own PTY, tmux session and respawn behaviour, exactly like
`antigravity`. Append `pi` after `antigravity` in every enum/list to keep ordering consistent.
| Tab badge | `pi` (two-letter lowercase, like `sh`/`oc`/`cx`/`gm`/`ag`) |
| Run button label | `Run PI` (short-label ternary in `_applyRunMode`, pattern `Run AG`) |
| Kill-menu label | `Kill Tmux & Pi` |
| Identity color | **`#f472b6` (rose-400)**. Verified free: live computed values on the default skin are claude `#38b6f0`, opencode `#44b993`, codex `#2b8fd9`, gemini `#8ab4f8`, antigravity `#22d3ee`, shell `#98a2b1`, web `#38bdf8`; purple is codex's base hex and amber reads as the shell tab badge, so pink/rose (or orange `#fb923c`) are the only genuinely free hues. No `pi` CSS identifier collides anywhere (`mode-pi`, `.tab-mode.pi`, `.run-mode-dot.pi` all grep clean, re-checked at f39beb3) |
| Env prefix | `PI_` |
| Dependency id | `pi` |
| Status endpoint | `GET /api/pi/status` |
### 2.2 `isExternalCliMode()` yes, `isAltScreenStripMode()` no
Pi joins `isExternalCliMode()` (`session.ts:164-167`): its own TUI, its own output format, so the
Ralph tracker, `BashToolParser`, token/CLI-info scraping and the `❯` readiness probe all stay off
(gates at `session.ts:1100`, `:1701`, `:2000`, `:2103`), and readiness falls back to the output
stabilization used by the other external CLIs.
Pi stays **out** of `isAltScreenStripMode()` (`session.ts:197-199`, currently codex/claude/gemini;
antigravity and opencode are deliberately excluded). Pi's default TUI renders into the main screen
with terminal-owned scrollback, so there is nothing to strip. The fullscreen mode **shipped in
0.84.0 and is runtime-switchable via `/settings`**, so Codeman cannot assume a pi session stays
main-screen for its lifetime; staying out of the strip list is exactly what makes that safe (the alt
screen is load-bearing when the user flips to fullscreen, as it is for `opencode`). Putting pi IN
the strip list would corrupt fullscreen sessions. Three mirrors must stay consistent (all unchanged
for pi, i.e. pi appears in none of them): the replay-side strip in `session-routes.ts:2275`, the
live-stream twin in `session.ts`, and the frontend `_sessionUsesServerMouseStrip()` in
`terminal-ui.js` (usages `:3432`, `:3697`).
### 2.3 tmux required, no direct-PTY fallback, no per-mode configurator
Same rule as the other external CLIs: `pi` mode throws if tmux is unavailable. Add a fourth block to
the guard chain at `session.ts:1751-1768` (antigravity's is `:1765-1768`).
**No `_configurePi()` is needed.** Opencode/codex/gemini each have a tmux-`setenv` configurator
(`tmux-manager.ts:1709-1727`), but antigravity has none: it relies entirely on the generic
`applyEnvOverrides()` (`tmux-manager.ts:1643`, `VALID_KEY = /^[A-Z_][A-Z0-9_]*$/`), which runs for
every mode in both create (`:1880`) and respawn (`:2107`) and injects via socket-scoped
`tmux setenv`, never the spawn command line. Pi follows the antigravity precedent: `PI_*` overrides
flow through `applyEnvOverrides()` and nothing else.
Pi joins the truecolor branches: `buildEnvExports()` (`tmux-manager.ts:1604-1609`,
`export COLORTERM=truecolor` + `unset NO_COLOR` for codex/gemini/antigravity) and the attach-env
condition at `session.ts:1400-1402` (`buildMuxAttachEnv(...)`, whose comment says it must mirror
`buildEnvExports`). Add `|| mode === 'pi'` to both, or the tmux session and the attach client
disagree about color depth.
### 2.4 Env prefix: `PI_` only
Add `'PI_'` to `ALLOWED_ENV_PREFIXES` (`schemas.ts:125`) and to the prose error message at `:163`
(two edits: the message hardcodes the list, and since 1.12+ it also names the exact-key allowlist,
currently `...ANTIGRAVITY_* keys and CLAUDE_CONFIG_DIR are allowed.`; there is now a separate
`ALLOWED_ENV_KEYS` exact-key set alongside the prefix list, which pi does not need to touch). That
covers every documented variable pi reads: `PI_CODING_AGENT_DIR`, `PI_CODING_AGENT_SESSION_DIR`,
| `src/utils/index.ts` | Re-export the three (resolver block `:30-36`) |
| `src/types/session.ts` | `SessionMode` union `:46`; **both `Extract` lists**: `RemoteCommandMode``:48-51`, `DockerCommandMode``:157-161` (§2.8); new `PiConfig` after `AntigravityConfig` (`:325-333`); `SessionState.piConfig` after `:486`; `@fileoverview` mode list `:11` + config list `:17` |
| `src/mux-interface.ts` | `piConfig?: PiConfig` on `CreateSessionOptions` (config block ends `:78`) and `RespawnPaneOptions` (ends `:109`) |
| `src/session.ts` | `isExternalCliMode()``:164-167` (+pi); `getModeLabel()``:168-183` (+`'Pi'`); `_piConfig` field decl `:466-470`; ctor option `:556-563` + apply `:652-654`; `toState()``:1227-1230`; `_buildRespawnPaneOptions()``:1466-1469` (single source of truth shared by `startInteractive` and `reattachRemote`); `startInteractive()` createSessionOptions `:1680-1683`; COLORTERM attach-env condition `:1400-1402` (+pi); requires-tmux guard chain `:1751-1768` (new block: "Pi sessions require tmux for env override injection via setenv") |
| `src/tmux-manager.ts` | `buildPiCommand()` after `:736` per §3; `buildSpawnCommand()` signature `:770-779` + dispatch branch after `:822-825`; `appendResumeFlag()``:1030-1042` (`case 'pi': return \`${modeCommand} --session ${resumeId}\`;`); `buildEnvExports()` truecolor branches `:1604-1609` (+pi); `buildPathExport()` `:1680-1707` (+pi branch calling `resolvePiDir()`); missing-CLI error chain in `createSession` `:1788-1806` (+pi, install hint `npm install -g --ignore-scripts @earendil-works/pi-coding-agent`; note `respawnPane` deliberately has no such check); `piConfig` threading at the four sites `:1748`, `:1817`, `:2041`, `:2080`. **No `_configurePi`** (§2.3) |
| `src/docker-hosts.ts` | `defaultDockerCommandForMode``:138-149`: `pi: 'exec pi'`. `CRED_STORES``:597-605`: the `.pi/agent` seedFiles entry per §2.5 (nested `rel` already handled at `:613-645`). File unchanged since 2026-08-06 |
| `src/remote-hosts.ts` | `defaultRemoteCommandForMode``:92-118`: `pi: remoteLoginShellCommand('pi')` (`remoteLoginShellCommand` at `:88-90`). Login-shell routing is mandatory (the #209/e803186 lesson: ssh remote-command exec sees only sshd's minimal PATH, and npm's global bin is usually only on PATH via rc files) |
| `src/web/schemas.ts` | `'PI_'` in `ALLOWED_ENV_PREFIXES``:125`**and** the prose error message `:163` (which now also names `CLAUDE_CONFIG_DIR`; the `ALLOWED_ENV_KEYS` exact-key set needs no change); new `PiConfigSchema` after `AntigravityConfigSchema` (`:256-271`), mirroring §3's regexes, `.optional()`, not `.strict()`; `piConfig` on `CreateSessionSchema` (`:299` area) and `QuickStartSchema` (`:712` area); `'pi'` in all three mode enums (`:285`, `:708`, cron `agentType``:1214`; they are byte-identical and there is no fourth); `pi` key in `RemoteCommandOverridesSchema``:426-436` (it is `.strict()`, so an unknown key is a hard error today; one edit covers both remote `:501` and docker `:577` reuse) |
| `src/web/routes/session-routes.ts` | Thread `piConfig` through create (`POST /api/sessions`): disk-strip exclusion chain `:705-712`, availability gate `:782-790` (+`isPiAvailable` with install-hint error), model resolution `:825-838` (`mode === 'pi' ? body.piConfig?.model : ...`), clamp call `:845`, Session ctor `:860` (`piConfig: mode === 'pi' ? gatedPiConfig : undefined`). Quick-start (`POST /api/quick-start`, handler `:2559`): remote-case config rejection `:2614-2621` and docker-case `:2645-2652` (+`piConfig`: per-CLI config does not cross ssh or the bind mount), hooks-scaffold exclusions `:2801`/`:2809`, availability gate `:2744-2752` (local-case branch only), env-strip chains `:2833`/`:2863`, model resolution `:2885`, clamp `:2897`, ctor `:2913`. **Extend `clampExternalCliBypassForOwner()`** (`:305-336`, doc comment above): fifth param + return field; pi joins the **materialize** branch per §5.2. Alt-screen replay-strip at `:2275` unchanged (pi not in it, §2.2) |
| `src/web/routes/system-routes.ts` | `GET /api/pi/status` after the antigravity handler (`:418-426`; file unchanged since 2026-08-06), same shape plus `version` (§2.6); update the "CLI Integrations" prose comment `:377` |
| `src/web/server.ts` | Restore path: `piConfig: muxSession.mode === 'pi' ? savedState?.piConfig : undefined` after `:2636`. **`renderIndexHtml` CLI-availability injection `:1375-1407`**: add `isPiAvailable` to the dynamic-import tuple (`:1382`) and a `pi` key to the injected object (`:1399`). Per §2.8 a missing key reads as *available*, so this is a correctness edit, not polish |
### Phase 3: Frontend
The antigravity touchpoints are the template. Since the first draft, the settings-surface overhaul
moved most anchors and added one **new touchpoint** (the clone-repo Brain picker below).
`constants.js`, `api-client.js`, `ralph-wizard.js`, `cron-ui.js`, `webview-tabs.js` and `sw.js`
still need **no** changes (re-verified zero mode coupling at f39beb3; cron-ui reads the `<select>`
| `index.html` | Welcome button `welcomePiBtn` after Gemini's (antigravity's is `:347`; there is deliberately no codex welcome button), `display:none` default, `onclick="app.setRunMode('pi'); app.runPi()"`, text `Run Pi`; run-mode-option row with `.run-mode-dot.pi` after antigravity's (`:526-528`), before the `.run-mode-sep``:529`; cron `<option value="pi">Pi</option>` after `:803`; **NEW: the clone-repo "Brain" picker** (`cloneCaseBrain`, `:2476-2486`): add `<option value="pi" data-cli="pi">Pi</option>` after the antigravity option `:2483` (gating is automatic: session-ui.js `:2107-2115` hides options whose `data-cli` fails `isCliAvailable`, and `:2250` reads the value at clone time); docker image hint `:2624` (`claude/codex/gemini/opencode/agy` + pi). No per-CLI remote-command override field needed (only codex has one, `:2559`) |
| `session-ui.js` | `@fileoverview` mode list `:2`; `run()` dispatch branch after `:400-402`; `_refreshRunModeAvailability` list `:468` (+`'pi'` as a quoted literal, the static test in §6 demands it); short-label ternary `:565` (+`'Run PI'`); **the `runMode` setter whitelist `:2949-2960`** (§2.8, the deceptive one); new `runPi()` modeled on `runAntigravity()``:1170-1219`: same remote/docker skip, same `_beginSessionLaunchStatus` frame, probes `/api/pi/status` reading `(await res.json()).data.available` (envelope!), **sends no `piConfig` at all** (no bypass exists and trust defaults are pi's own; envOverrides still sent for local cases), install-hint error text matching Phase 1's; `isAltMode``:1233` and `isExternalCli``:1263` four-way comparisons (+pi) |
| `panels-ui.js` | Command-palette `labels` map `:430` (+`pi: 'Pi'`; the `\|\| mode` fallback means this is cosmetic, not load-bearing) |
| `mobile-overview.js`| `MOBILE_OVERVIEW_RUN_MODES``:55-62`: `{ mode: 'pi', label: 'Pi', short: 'Pi' }` after antigravity `:60`, before the shell entry. Nothing else: the Run-button badge (`:499`) and menu builder (`:554-556`) consume the list generically, and the buttons carry `btn-toolbar btn-run mode-pi`, which is exactly why they inherit the §2.9 cascade problem and its fix |
| `terminal-ui.js` | Badge-row comment `:1750` only (the badge itself is a raw `s.mode` passthrough, no list to extend). `_sessionUsesServerMouseStrip` unchanged (§2.2). `_updateLocalEchoState` unchanged for v1 (§2.10: pi lands on `'buffer'` via the fallthrough; only touch it if E2E forces the `'off'` fallback) |
| `i18n.js` | `'Run Pi': '运行 Pi'` in the zh-CN table (`:102-107`, matches the welcome-button text; short labels like `Run PI` are deliberately untranslated, as are the other modes') |
| `styles.css` | Tab badge `.session-tab .tab-mode.pi` after `:2157` (`background: rgba(244,114,182,0.2); color: #f472b6;`); add `.session-tab .tab-mode.pi` to the light-skin ink list `:325-336` (gemini + antigravity are its precedent, `:332`); welcome `.welcome-btn-pi` + `:hover` after antigravity's `:3366` block, rose family (e.g. base `linear-gradient(135deg, #33121f 0%, #9d174d 55%, #be185d 100%)`, border `rgba(244,114,182,0.4)`, text `#fce7f3`); toolbar gradient pair `.btn-toolbar.btn-run.mode-pi, .btn-toolbar.btn-run-gear.mode-pi` + `:hover` after `:4420`'s antigravity block; `.run-mode-dot.pi { background: #f472b6; }` in the dot list `:4506-4516`; **and the §2.9 rule inside the Daylight block** next to codex's `:13787` (e.g. `background: linear-gradient(135deg, #be185d, #f472b6); border-color: #be185d; color: #fff1f7;`). The dot needs no skin-block entry (the block overrides only claude/opencode/codex/shell dots; gemini/antigravity dots already fall through correctly) |
| `mobile.css` | Phone toolbar block after `:910` inside the `@media (max-width: 430px)` opened at `:338`: `mode-pi` base + `:active`, **with `!important` on background/border-color/color** (§2.9; antigravity's block `:895-910` omits it and is dead); light-skin override entry after `:2985` with the same four-skin `html:is(...)` prefix as its siblings |
### Phase 4: Docker image and installer
Both files are unchanged since the 2026-08-06 verification; all anchors stand.
- `docker/agent.Dockerfile`: a **separate**`RUN` step after the antigravity block (`:38-45`), not a
fifth line in the shared npm block (`:31-36`), because pi documents `--ignore-scripts` and that
flag must not silently change how the other four install:
```dockerfile
# Pi (pi.dev). Upstream documents --ignore-scripts (pi needs no lifecycle scripts);
# kept out of the shared npm block above so the flag cannot affect the other CLIs.
RUN npm install -g --ignore-scripts @earendil-works/pi-coding-agent \
&& npm cache clean --force \
&& pi --version
```
Implementation checklist item: the gid-0 pre-created dirs at `:64-68` include `.claude/projects`
and `.codex/sessions`; verify whether the cred-seed copy into `~/.pi/agent` creates its target
dir in a fresh container or whether `.pi/agent` must join that `mkdir` line. Rebuild with
`node scripts/build-agent-image.mjs --no-cache` (the script itself needs no change; nothing in it
is CLI-specific). The cached npm layer has silently frozen a CLI at a broken version before; see
`docs/docker-cases.md`.
- `install.sh` (six edit sites, all verified): `PI_SEARCH_PATHS` block after `:125` (mirror the
| `pi` resolves to an unrelated binary | `pi --version` + semver-shape check in the resolver (§2.6); path and version shown in `/api/pi/status` |
| Pi's TUI repaints in a way the browser terminal handles badly | Test scrollback and repaint early (step 3 of §7); pi's default is main-screen with terminal-owned scrollback, which is the friendly case |
| Fullscreen TUI mode (shipped 0.84.0, runtime-switchable) | Already designed for: pi stays OUT of the strip list, so a user flipping `/settings` to fullscreen gets opencode-like alt-screen behavior, not corruption. §7 step 7 tests the flip explicitly |
| The buffer local-echo overlay fights pi's live composer | §2.10: explicit E2E gate (§7 step 4) with the one-line `'off'` fallback; predictive echo for pi is a tracked follow-up, not a v1 blocker |
| Pi moves fast (pre-1.0; 9 releases in the 7 weeks before 0.84.1) | Keep the flag surface small; every flag validated and droppable; nothing pinned in the Dockerfile beyond the `--no-cache` rebuild cadence. Live example of the hazard: `--tui-mode` went from main-only docs to released between the two drafts of this plan |
| Docker image grows | Pi is an npm package; the layer is modest next to the ~190MB `agy` binary |
| Trust prompt blocks a session | Narrower than feared: only fires when `.pi/settings.json`, `.pi/extensions\|skills\|prompts\|themes`, `.pi/SYSTEM.md`/`APPEND_SYSTEM.md` or `.agents/skills` exists (bare `.pi/` does not). Documented; `approveProjectTrust` is the opt-in escape hatch; multi-user forces `--no-approve` (§5.2); the `project_trust` extension follow-up removes the prompt entirely |
| Provider auth is awkward without key prefixes in the allowlist | `/login` writes `~/.pi/agent/auth.json` once and Docker seeds it; the mode-aware allowlist follow-up removes the friction |
| Cron pi jobs mis-detect readiness | Known degradation, documented in §6; readiness falls through after the poll budget and the prompt still sends |
@@ -124,7 +124,7 @@ Agent use cases this unlocks: a lead session records intentions as the user stat
1. **Intent store + capture + intent endpoints + skill docs.** Immediately useful to agents even before any UI exists.
2. **Context assembler + predictor + predict endpoint + desktop button/modal.** The feature as pitched. The assembler ships with all collectors it can serve from day one (transcript, intent, git, run-summary, siblings); the approvals collector activates when PR #245 lands.
3. **Phone accessory key, rethink steering, alternates row.** Part 1 (shipped): the alternates row (tappable, swap into the field without losing edits; Rethink rejects the whole shown set), the phone 🧠 keyboard-accessory key (both bar templates, `rmm-enabled` marker class on the bar), and a phone-sized modal (small dialog, not full-screen). Part 2: rethink steering (the free-text steer note; the API already accepts `steer`).
3. **Phone accessory key, rethink steering, alternates row.** Part 1 (shipped): the alternates row (tappable, swap into the field without losing edits; Rethink rejects the whole shown set), the phone 🧠 keyboard-accessory key (both bar templates, `rmm-enabled` marker class on the bar), and a phone-sized modal (small dialog, not full-screen). Part 2 (shipped): rethink steering, the free-text steer note under the suggestions, sent as `steer`, visible whenever Rethink is live (ready and empty-result phases), cleared on each open; the empty-result copy points at the note, and the footer buttons moved to the styled `btn-toolbar` convention (the bare `btn btn-*` classes they shipped with match no CSS in this codebase and rendered as unstyled UA buttons).
4. Explicitly later: proactive predict-on-idle (ghost suggestion chip), auto-compaction of `recentPrompts` into `goals` via a cheap model, codex/gemini capture, cross-case "global" intent.
@@ -27,7 +27,7 @@ On a Claude session, press the brain button in the header (desktop) or the 🧠
- **Send** submits it to the session (with Enter).
- **Insert** drops it on the CLI composer *without* Enter, so you can edit it in the terminal before sending.
- **Rethink** re-runs with everything shown (the field and the alternates) recorded as rejected.
- **Rethink** re-runs with everything shown (the field and the alternates) recorded as rejected. An optional steer note below the suggestions ("no, I meant the mobile bug") rides along as your own words, the highest-authority signal the predictor gets; it stays in the field across re-runs until you clear it or reopen the modal.
- **Dismiss** closes; nothing happens.
A prediction takes 5-90 seconds and costs real tokens; one runs per session at a time. If the session is sitting on a permission/question dialog, the suggestion is usually an answer to that dialog: that is intentional.
@@ -38,7 +38,7 @@ A prediction takes 5-90 seconds and costs real tokens; one runs per session at a
Capture reads the Claude session transcript, not your keystrokes: when a user turn lands in the transcript, its text is folded into the case's profile. Filters applied on the way in:
- **Claude-mode sessions only.** Shell, OpenCode, Codex, Gemini, and Antigravity sessions are never captured (they have no transcript watcher).
- **Claude-mode sessions only.** Shell, OpenCode, Codex, Gemini, Antigravity, and Pi sessions are never captured (they have no transcript watcher).
- Tool results, local slash-command echo (`/model` and friends), system wrappers, and interrupt markers are skipped.
- Entries shorter than 3 characters are skipped (menu digits, Esc artifacts).
- Consecutive duplicates collapse (auto-resume's "continue" spam counts once per run).
@@ -87,7 +87,7 @@ The `codeman` agent skill documents the same verbs (SKILL.md §3 plus `reference
## What comes next
A steer-note input on Rethink ("no, I meant the mobile bug"; the API already accepts `steer`). Explicitly later: proactive predict-on-idle, auto-compaction of the prompt history into goals, non-Claude capture. See the phases section of [`readmymind-plan.md`](readmymind-plan.md).
Explicitly later: proactive predict-on-idle, auto-compaction of the prompt history into goals, non-Claude capture. See the phases section of [`readmymind-plan.md`](readmymind-plan.md).
| `RemoteCase` | A working directory on a host: `name`, `type: 'remote'`, `hostId`, `remotePath`. |
| `SessionRemote` (extends `RemoteSshOptions`) | The resolved bundle stamped onto a live session: host coordinates + `remotePath` + `commands`, plus **`owned?`** and **`remoteSessionName?`** (COD-105 — see [Ownership](#ownership-launched-vs-discovered-and-attached-cod-105)). Built by `toSessionRemote(host, case)` (sets `owned: true`) for the launch path, or `toAttachedSessionRemote(host, name, path)` (sets `owned: false`) for the attach path. Both copy the advanced SSH options through so every connection is identical. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity'>` — the modes that can run remotely. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity' \| 'pi' \| 'grok' \| 'deepseek' \| 'omp'>` — the modes that can run remotely. |
| `RemoteSessionInfo` (COD-105) | One discovered remote tmux session: `name` (always `codeman-*`), `attached` (a client is connected), `created` (epoch s), `windows`. Returned by `listRemoteCodemanSessions()`. |
Persistence is two flat JSON arrays in the instance data dir:
@@ -116,6 +116,11 @@ Key points:
the agent. The per-mode command comes from `remote.commands?.[mode]` or
⚠️ **claude and omp no longer take that path**: both have their own arm in
`buildRemoteLaunchCommand` so a respawn can continue the same conversation
(see [Respawn / reattach continuation](#respawn--reattach-continuation)), and
because the claude arm is an `a || b` pair under `-c`, its pane PID is the
**login shell**, not the agent.
- The **whole tmux invocation is a single shell-quoted ssh argument**, and the
pane command is independently quoted, so a `remotePath` with spaces is safe.
- Connection options come from the **same `buildSshConnectionArgs(remote)`** as
@@ -202,6 +207,151 @@ The early return is a structural guarantee that **no code path can ever issue a
remote `kill-session` for a session we don't own** — the only `kill-session` run is
on the local socket, which never reaches the remote socket.
## Respawn / reattach continuation
A dropped connection or a dead pane must reconnect to the **same conversation**,
not launch a fresh one — the whole point of a durable remote session.
- **Claude**: the launch command is idempotent — `claude --session-id <id> ||
claude --resume <id>` (see `buildRemoteLaunchCommand`'s claude branch). The
first run creates the conversation under the deterministic session id; every
later reattach/respawn re-runs the same line, `--session-id` fails
("already in use"), and the `||` fallback resumes it.
- **OMP**: `omp` has no equivalent idempotent single-line form, so
`Session._pinOmpRespawnId()` resolves and pins an explicit `--resume <id>`
before a respawn (mirroring the local/docker builders, rendered through the
same `buildSpawnCommandFromRegistry` engine — not a hand-rolled command and
not `appendResumeFlag()`, which is docker-only and cannot work here: appending
a flag after the quoted `-c 'omp'` hands the id to the login shell as `$0`
instead of to `omp`). ⚠️ **The resolver only ever reads THIS host's local
`~/.omp/agent/sessions/`**, which is meaningless for a remote session — the
conversation and its session file live on the remote host, under the remote
user's home. For a remote session, `_pinOmpRespawnId()` therefore skips local
resolution entirely and falls back to `omp`'s own ambiguous `--continue`
(`ompConfig.continueSession`), which the remote pane command already renders.
This is a known, accepted degradation versus the local/docker paths' exact
`--resume` pin — safe in practice because each remote respawn talks to
exactly one remote pane's own omp history, so "most recent" is normally
correct, but it can drift the same way `--continue` always could if two
remote sessions ever share one remote directory.
## Auto-reconnect vs. a clean agent exit
`remoteAutoReconnect` (default ON) watches for a dropped SSH connection and
reconnects with bounded backoff. It must **never** revive a session whose agent
exited cleanly (Ctrl-C, Ctrl-D, `exit`) — that tears down the durable remote
tmux session itself, and a transport-level `isPaneDead()` cannot tell that apart
from a plain network drop. `remoteTmuxSessionAlive()` (#355) resolves this by
probing the remote host directly: `tmux -L codeman-remote has-session -t
codeman-ssh-<id8>` over the same `buildSshConnectionArgs` as launch, classified
by **exit status alone** (`classifyRemoteAliveExit`: `0` = alive, ssh's `255` or
a timeout = unknown, anything else = gone) — `has-session` prints nothing on
success, so reading stdout would misclassify every live session as gone. An
unreachable host answers "unknown", which also means do not revive. The answer
is cached per session and cleared whenever the pane is next seen alive, so a
stale `true` from one transport drop can never revive the NEXT clean exit.
## File access over SSH
A remote case's `workingDir` is an absolute path on the **remote** host
(`Session.workingDir = RemoteCase.remotePath`), so the file routes cannot use local
`fs`: a local `realpathSync` on a remote-only path fails by construction, which is why
previewing a file used to answer `404 File not found` for a case that was working
perfectly (#415). `src/remote-files.ts` is the one module that reads remote bytes,
and it follows the same rule as the launch path: every ssh command line comes from
`buildSshConnectionArgs()` — **never** a hand-built ssh line.
| Request | What happens |
|---------|--------------|
| `GET /api/sessions/:id/file-raw` | Streamed over `ssh` (`cat`, or `tail -c +N \| head -c L` for a `Range`); the same 200/206/416 contract as a local file, so `<video>`/`<audio>` seeking works |
| `GET /api/sessions/:id/file-content` | `cat` into memory, capped by the existing text limit; `edit=1` answers `400` (see below) and `editable` is always `false` |
| `PUT /api/sessions/:id/file-content` | `400` before any path is looked at: the guard sits AHEAD of the local path validation, because with a same-named directory on the Codeman host (an `sshfs` mount) the write would otherwise land on the local twin |
| `GET /api/sessions/:id/file-preview` | Non-office files redirect to `file-raw` (which works remotely); docx/pptx answer `400` |
| `GET /api/sessions/:id/file-thumbnail` | `400` for remote files |
| `POST /api/sessions/:id/attachments` | Registers an absolute path that lives on the **remote** host (a clicked link pointing outside the case directory) by probing it there |
| `GET /api/sessions/:id/attachments/:attachmentId/raw` | Streams the registered remote file over ssh, same 200/206/416 contract; `preview` (office) and `thumbnail` answer `400` |
| `GET /api/sessions/:id/attachments/:attachmentId`, `GET …/attachments` (history) | Size/mtime/existence resolved over ssh, so a remote entry is not reported `missing`; the history list resolves EVERY entry in one batched probe, never one connection per entry |
⚠️ The attachment route is the one a clicked path takes when it is **outside** the case
directory (a remote `/tmp` scratchpad capture, a screenshot elsewhere in the home dir):
the frontend's `_isExternalPreviewPath()` sends every absolute path that is not under
`workingDir` there, so fixing only `file-raw` would leave exactly that half broken.
Guard order is deliberately **the same as locally**, and the checks are not weakened
by the transport:
1. Ownership (`findSessionOrFail` / the scope helper) — unchanged.
2. Lexical containment of `workingDir + path` — a `../` escape is refused before any
connection is opened.
3. ONE ssh round trip that returns `realpath`**and**`stat` for the path **and** the
workspace root (`remoteProbePaths`). Resolving the root remotely is what keeps the
boundary honest for a symlinked `remotePath`. The probe uses `readlink -f` when
available; on a host without it (macOS before 12.3) a POSIX fallback canonicalizes
the directory chain with `cd -P`/`pwd -P` and then follows the LAST component with
plain `readlink` for a bounded number of hops. ⚠️ **The fallback fails closed**: a
path it cannot fully resolve (a loop, a `readlink` failure, the hop cap) is reported
as unresolvable and answers 404, never as its own unresolved string. An earlier
version resolved only the directory chain, so `ws/notes.txt -> ~/.ssh/id_rsa` passed
containment under the link's own path while `cat` followed it to the key.
Records come back NUL-separated and index-keyed (`<index>|kind|size|mtime|realPath`,
after a leading NUL that fences off any login banner), so a filename containing a
newline cannot shift the alignment.
4. Containment of the remote realpath against the remote root. The sensitive-path
blocklist then applies on whichever routes already apply it locally (`/api/download`,
attachment registration, edit mode — where resolving symlinks first is what makes it
meaningful); the remote branch neither drops a guard the local path has nor invents a
stricter one. One entry of that blocklist is host-bound by construction: the three
home-anchored members (`~/.claude.json`, `~/.claude/settings.json`,
`~/.claude/settings.local.json`) are compared against the **Codeman host's** home
directory, so they do not match a remote home at a different path. Everything else in
the list is depth-anchored (`/.ssh/`, `/.aws/credentials`, `/.claude/.credentials.json`,
`/etc/shadow`, ...) and applies to a remote path unchanged.
5. Size cap (`CODEMAN_MAX_DOWNLOAD_BYTES`) applied to the **remote** size, before the
body is requested.
The path arrives from the browser (`?path=`) and is interpolated as a single
`shellescape`-quoted token, in a command that is itself shellescaped into the ssh
line; `BatchMode=yes` means a host needing a passphrase fails fast instead of hanging.
A failed connection is reported as **502** with the remote reason — never a 404, which
used to make an unreachable host look like a typo in the agent's output. The reason is
the first stderr line, the timeout, or the exit code; never Node's `Command failed: …`
message, which would carry the identity-file path and the probe script into the body.
**Connections are bounded.** Every probe and buffered read runs through a small global
semaphore (`src/remote-ssh-limiter.ts`, default 4, `CODEMAN_MAX_REMOTE_FILE_SSH`), the
attachment-history list resolves its whole history in one batched probe instead of one
handshake per entry, and probes are chunked at 40 paths per round trip. Terminal output
in a remote session is written on the remote host, so a prompt-injected agent printing
hundreds of `codeman://attach` links used to make the server fork one `ssh` per link,
each holding a 20 s probe timeout, and a 100-entry history re-listed on every
| `file-raw` | 2 GB (`CODEMAN_MAX_DOWNLOAD_BYTES`, `0` = unlimited) | inline MIME map; **`X-Content-Type-Options: nosniff` on all responses**; streamed, `Range`-aware (206 slices come from the same validated path, and the cap is checked before the range) |
| `GET /api/download` | same cap | forced `attachment`; sensitive‑path blocklist; streamed, `Range`-aware |
### SVG / content‑type XSS
@@ -340,7 +342,7 @@ the attachment guard below.
Live external attachments (`src/attachment-registry.ts`) mint an `att_<uuid>` id
for a host file so browser requests carry the id, never an absolute path. Serving
is by id (`GET /api/sessions/:id/attachments/:attachmentId/raw`, 50 MB cap,
is by id (`GET /api/sessions/:id/attachments/:attachmentId/raw`, same download cap,
`nosniff`) and re‑resolves the symlink + re‑checks the **attachment guard**
(`src/config/attachment-guard.ts`: the shared sensitive‑path blocklist **plus**
the `/root` and `/etc` trees, extendable via `attachmentBlockedPaths` /
@@ -489,7 +491,7 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never**`--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — and`~/.config/{gcloud,opencode}`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, five seeded files from `~/.pi/agent`, and three from `~/.grok`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
@@ -514,11 +516,12 @@ Full feature guide: [`docker-cases.md`](docker-cases.md).
## 10b. Web tabs (dashboard proxy)
A saved dashboard URL renders as a tab, served through Codeman's own origin at `/webview/<capability>/`. User guide: [`web-tabs.md`](web-tabs.md). Three properties carry the security weight:
A saved dashboard URL renders as a tab, served through Codeman's own origin at `/webview/<capability>/`. User guide: [`web-tabs.md`](web-tabs.md). Four properties carry the security weight:
- **The proxy is exempt from cookie auth and the Origin/CSRF guard, and that is deliberate.** The iframe is sandboxed without `allow-same-origin`, so it is opaque‑origin: its requests are cross‑site, meaning the `SameSite=lax` session cookie is never attached and its writes and WS upgrades arrive with `Origin: null`. The credential is instead a 192‑bit capability in the path, minted only by an authenticated `POST /api/webviews/:id/open`, held in memory (a restart invalidates every one), rolling TTL, bound to the minting user, and granting nothing but "relay bytes to this one saved URL". ⚠️ **The Host allowlist is NOT bypassed**, so DNS‑rebinding protection is unaffected. A second `Referer`‑keyed form exists for root‑absolute assets and is the only exemption decided by a request‑supplied header, so it is fenced to safe methods on non‑`/api`, non‑`/ws`, non‑`/q` paths. Edges pinned by `test/webview-auth-exemption.test.ts`.
- **The lost‑frame recovery page is the third unauthenticated 200, and the only one decided by request headers alone.** The proxy's runtime shim masks `/webview/<cap>/` off the page's own URL so a single‑page app routes on the path it expects; a navigation the page then starts itself (`location.reload()`, a root‑absolute `location.href`) lands on Codeman's root with no capability anywhere, no cookie (opaque origin) and a Referer naming the masked page. `serveLostWebviewFrame()` in `middleware/auth.ts` recognises it by shape (`GET`/`HEAD`, `Sec-Fetch-Dest: iframe` or `frame`, `Accept: text/html`, `Sec-Fetch-Mode: navigate` or absent) and answers, BEFORE the credential checks and without counting an auth failure, with a static page whose only content is a `postMessage` of the lost path to the parent tab (`default-src 'none'` plus the hash of that one script, `no-store`, `referrer: no-referrer`, no reflected input). It is fenced to paths that are NOT registered routes and never `/api/`, `/ws/` or `/q/`, with one carve‑out: `/` itself, because the landing page masks to exactly `/` and its reload otherwise rendered Codeman's app shell inside the web tab. `/` is admitted only when the request carries neither the `codeman_session` cookie nor an `Authorization` header: nothing in Codeman frames its own root and a sandboxed frame has neither, while a framed `/` that does carry credentials still gets the shell. On a passwordless install no auth hook runs, so the index route applies the same test itself (`isLostWebviewRootFrame`). ⚠️ Known property, accepted rather than mitigated: those headers are trivially set by a non‑browser client, so an unauthenticated caller can distinguish a registered route (401) from a non‑route (200) and enumerate the route table; the routes are public in `docs/api-reference.md`, so nothing is learned. Pinned by `test/webview-auth-exemption.test.ts` (password) and `test/webview-lost-root-frame.test.ts` (passwordless).
- **Sandboxed by default; `allow-same-origin` is an explicit per‑dashboard opt‑in.** A proxied page is same‑origin with Codeman, so without the sandbox its JavaScript could read the Codeman document and call the agent‑spawning API. ⚠️ In BOTH modes the `Authorization` header and the `codeman_session` cookie are stripped before the upstream request, because a trusted (same‑origin) frame makes the browser attach Codeman's own Basic‑auth credentials to every proxied request; forwarding them would hand `CODEMAN_PASSWORD` to the dashboard.
- **Not an open relay, and not a privilege boundary.**`resolveUpstreamUrl()` refuses anything leaving the saved origin, and cross‑origin redirects are handed back unchanged rather than followed. The proxy does reach whatever the SERVER can reach, which is not an escalation for someone who already commands `--dangerously-skip-permissions` agents, but in multi‑user mode it means a non‑admin's dashboard is fetched from the server's network position. Saved URLs are validated to plain http(s) with no embedded credentials, and there is deliberately **no magic‑link path**: terminal output can never create a webview (the mistake the attachment scanner had to be walled off from).
- **Not an open relay, and not a privilege boundary.**`resolveUpstreamUrl()` refuses anything leaving the saved origin, and cross‑origin redirects are handed back unchanged rather than followed. The proxy does reach whatever the SERVER can reach, which is not an escalation for someone who already commands `--dangerously-skip-permissions` agents, but in multi‑user mode it means a non‑admin's dashboard is fetched from the server's network position. Saved URLs are validated to plain http(s) with no embedded credentials, and there is deliberately **no magic‑link path**: terminal output can never create a webview (the mistake the attachment scanner had to be walled off from). The one refused destination class is link‑local and cloud‑metadata addresses (`169.254.0.0/16`, `fe80::/10`, `fd00:ec2::254`, `168.63.129.16`, `100.100.100.200`, `metadata.google.internal`): `webview-egress-policy.ts` refuses them at save time, and `webview-egress.ts` re‑judges the RESOLVED address at connect time through a `lookup` hook on the proxy's undici Agent and on its WebSocket client, so a DNS name pointing into those ranges is refused as well. Loopback and RFC1918 stay allowed on purpose. Capabilities are revoked on logout, admin logout and user deletion, and proxied responses carry `Referrer-Policy: same-origin` so a dashboard cannot hand the capability‑bearing URL to a third‑party host it links.
---
@@ -529,6 +532,7 @@ A saved dashboard URL renders as a tab, served through Codeman's own origin at `
| `CODEMAN_ALLOWED_HOSTS` | Extra `Host`/`Origin` allowlist entries for reverse proxies (comma‑separated; exact host, or leading‑dot `.suffix` for subdomains) — see §3 |
| `--base-url` / `CODEMAN_BASE_URL` | Sub‑path prefix Codeman is mounted under behind a reverse proxy, e.g. `/codeman` (default `/`); the proxy must forward the prefix unchanged. Independent of `CODEMAN_ALLOWED_HOSTS` |
| `--allow-unauthenticated-network` / `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK` | Acknowledge an unauthenticated non‑loopback bind (downgrades the warning) |
| `--https` | Enable TLS (adds HSTS) |
| `CODEMAN_INSTANCE` | Scope tmux socket + data dir for isolation |
# Session lineage lines (spawn lines between tabs)
**Goal:** when a session spawns another session (the `codeman` agent skill starting a
worker, or anything else that says who it is), draw the same kind of glowing connection
line the subagent windows already use, but **tab → tab**, so a glance at the strip shows
which tab spawned which.
Status: PLAN. Nothing implemented yet.
---
## 1. The blocking fact: no parent relationship exists today
There is no spawn-parent link between sessions anywhere in the codebase:
- `SessionState` (`src/types/session.ts:388`) has no `parentSessionId` / `spawnedBy` /
`createdBy`.
- `POST /api/quick-start` and `POST /api/sessions` record only `owner = ownerFor(req)`,
which is the multi-user **human**, not the calling session.
- The only parent links that do exist are `TeamConfig.leadSessionId` (agent teams) and
`subagent-parents.json` (a frontend **window-layout** store for subagent windows).
Neither says "session A spawned session B".
- Nothing in the HTTP request identifies the caller: an agent's spawn call is plain
`curl` from inside a tmux pane, so there is no socket-level identity to recover
(`SO_PEERCRED` needs a unix socket; the API is TCP).
So the caller has to **tell** us. It already knows its own id: every managed pane gets
`CODEMAN_SESSION_ID` exported by `session-cli-builder.ts` (and the skill's §0 preamble
already binds it to `$SELF`).
## 2. Wire format
Two ways in, because they serve different callers. Body wins when both are present.
| Where | Shape | Who uses it |
| --- | --- | --- |
| body field | `"parentSessionId": "<uuid>"` | anything hand-writing one create call |
| request header | `X-Codeman-Parent-Session: <uuid>` | the skill: added **once** to the `CURL` array in the §0 preamble, so every present and future create call carries it with no per-recipe edit |
Rules, all of them deliberate:
- **Advisory decoration only.** It never grants access, never scopes anything, never
affects lifecycle. A child is not killed when its parent dies; the line just stops
being drawn once the parent tab is gone.
- **Never fails a spawn.** An unknown / stale / foreign parent id is silently dropped
(field ends up `undefined`), not a `400`. A cosmetic field must not be able to break
worker creation.
- **Resolved, not trusted.** The id must match a live session the caller can already
see (`canAccessOwned`), and the resolved parent's `owner` must equal the new
session's `owner`. Otherwise a user could staple their session under another user's
tab in multi-user mode.
- Exact id match first; a `>= 8`-char **unique** prefix match as a fallback (ids appear
truncated in mux names and UI surfaces; ambiguous prefixes resolve to nothing).
## 3. Server changes
| File | Change |
| --- | --- |
| `src/types/session.ts` | `SessionState.parentSessionId?: string` with a doc comment saying it is UI decoration and never a permission signal |
| `src/session.ts` | constructor option `parentSessionId` → `_parentSessionId`, public getter, emitted from `toState()` (~line 1170) |
| `src/web/schemas.ts` | `parentSessionId: z.string().max(100).optional()` on `CreateSessionSchema` (272) and `QuickStartSchema` (680). Neither is `.strict()`, so this is additive |
| `src/web/route-helpers.ts` | new `resolveParentSessionId(ctx, req, bodyValue, owner)` implementing §2's rules; returns `string \| undefined`, never throws |
| `src/web/routes/session-routes.ts` | pass it into the three `new Session({...})` sites: `POST /api/sessions` (846), `POST /api/run` (2522), `POST /api/quick-start` (2896) |
| `src/web/server.ts` | recovery path (~2617): `parentSessionId: savedState?.parentSessionId` so the link survives a restart |
**No new SSE event.** `session_created` / `session_updated` broadcast
`getSessionStateWithRespawn(session)`, which is `toState()`-derived, so the field rides
along to the browser for free — and the frontend already does
`this.sessions.set(data.id, data)`, so `session.parentSessionId` is simply there.
Optional follow-up: surface it on `/api/sessions/unified` rows so the Session Manager
and the home rails can show "spawned by w3-claudeman".
## 4. Frontend rendering
### 4.1 Where the code goes
`_updateConnectionLinesImmediate()` (`subagent-windows.js:242`) is a strict
**batched read → batched write** pass, and it already has an extension point:
ultracode appends its own layer via `_appendUltracodeConnectionLines(svg, rects)` at
the end, sharing the `rects` cache so no layer forces a second reflow.
Lineage lines follow that exactly: a new module `src/web/public/session-lineage.js`
(load order 15.6, after `ultracode-windows.js`) exporting
`_appendLineageConnectionLines(svg, rects)` onto `CodemanApp.prototype`, called from the
same tail. **The core function keeps ownership of the read/write split**; the new layer
only reads through the shared `rects` map and only appends paths.
The path math itself lives in `constants.js` as a pure
`computeLineagePath(parentRect, childRect, stripRect, depth)` — same treatment as
`computeTabScrollLeft`, so the geometry is unit-testable without a browser.
### 4.2 Geometry
Both endpoints are tabs in one horizontal strip, so the subagent shape (tab-bottom →
window-top) does not apply. **One case**, a **U-bridge hanging below the strip** that
Status: **phases 0-2 implemented** on `feat/tui`; phases 3-4 remain follow-ups. The user guide is [`docs/tui.md`](tui.md); this document stays the design record.
- Phase 0: `src/cli-style.ts` (palette, glyphs, `heading`/`kv`/`table`/`spinner`/`confirm`) plus the mechanical fixes of §5, and `test/cli-commands.test.ts` now derives its inventory from the real commander `program` instead of parsing a fixture.
- Phases 1-2: `src/tui/`. `tui-app.ts` (main loop, attach handoff, verbs) and `tui-client.ts` (API, SSE, degraded enumeration) are the only IO; `tui-model`, `tui-layout`, `tui-render`, `tui-keys`, `tui-ansi`, `tui-composer`, `tui-approvals`, `tui-digest`, `tui-sse` and `tui-types` are pure and unit-tested, with an E2E suite driving the real binary under node-pty.
- Deferred with the rest of phase 3: `r` (resume a RECENT row) is not wired up, so the help overlay does not advertise it.
- Not started: phase 3 (mouse, `--pick` popup switcher, opt-in attach status line, OSC 9) and phase 4 (retiring the bash choosers).
The goal: replace Codeman's scattered terminal surfaces with one first-class TUI, `codeman tui`, that gives SSH/terminal users the same at-a-glance awareness the web UI gives browsers. The reference point is herdr (herdr.dev), the trending Rust "agent multiplexer" whose defining feature is a live agent-state sidebar. Codeman can match and beat that sidebar in the terminal because the states herdr infers from screen-scraping heuristics are states our server already computes from hooks, pane probing, and the approvals inbox.
---
## 1. What we have today (inventory)
Three disconnected surfaces, three visual idioms, two data sources:
| Surface | What it is | Data source | Idiom |
| --- | --- | --- | --- |
| `codeman` CLI (`src/cli.ts`, 1214 lines) | commander + chalk, ~20 commands | HTTP API + state files | `✓`/`✗` line-per-fact, no interactivity |
| `sc` (`scripts/tmux-chooser.sh`, 663 lines) | bash number-menu chooser, mobile-tuned (44 cols) | `tmux -L codeman` + `state.json` via jq | 256-color, numbered, full repaint per key |
Weaknesses found in the audit (file:line refs verified 2026-08-16):
1. **No interactive picker in the Node CLI at all.** Every `session stop`, `task status`, `session logs` requires a pasted UUID prefix. There is no `codeman attach <session>`; `codeman attach` is actually the attachment-card command (and `README.md:895` describes it wrongly).
2. **`sc` cannot reach sessions 10+ interactively**: entries are numbered globally (`tmux-chooser.sh:343`) but input accepts a single `[1-9]` keypress (`:487-493`). Page 2 shows items 8-14 that mostly cannot be selected.
3. **No cursor/selection concept in `sc`** (`BG_SEL` at `:90` is dead code); arrows only page.
4. The two bash tools can disagree about which sessions exist (different data files), and only `sc` is on PATH.
5. **Zero live feedback anywhere**: `codeman web -d` and `service install` block silently up to 30s (`daemon-control.ts:395-412`); no spinner exists in the codebase.
6. Styling drift: `doctor` is the only table and is deliberately monochrome with a colorize hook nobody wired up (`dependency-report.ts:5-7`); `codeman web` prints its "running at" line twice (colored `cli.ts:934`, plain `server.ts:2366`); the server's security warning is colorless `console.warn` while the CLI's version of the same warning is yellow; `tmux-manager.sh`'s header box is visibly misaligned; `padEnd(14)` overflows on "Antigravity CLI".
7. Bash TUIs emit raw escapes unconditionally (no TTY/NO_COLOR gate); `install.sh` and `postinstall.js` do it right.
8. Detach hint inconsistency: chooser says Ctrl+B D, `README.md:671` says Ctrl+A D.
9. Inside an attached session there is **no chrome at all**: Codeman turns the tmux status bar off (`tmux-manager.ts:1978`), so an SSH user in a pane has no session identity, no state, no way back to a picker except detach.
10. `test/cli-commands.test.ts` asserts against a hand-written fixture, not the real `program`, and that fixture already lists a `tui` command that does not exist (`:57-61`). The name is pre-approved by our own test file.
## 2. Research: how herdr does it
herdr (github.com/herdrdev/herdr, ~30k stars, single Rust binary, pre-1.0) is a background terminal multiplexer "your coding agents live on". What matters for us:
- **The agent-state sidebar is the product.** Every pane is classified live as `working` / `blocked` / `done` / `idle` and grouped in a sidebar, so you see who needs you without switching tabs. Reviews unanimously call this "the killer feature tmux can't match".
- **Detection is heuristic-first**: process-name matching + screen-manifest TOML rules parsing the visible frame; optional per-agent "integration install" adds lifecycle hooks over JSON-RPC on a unix socket for accurate states. Claude Code there is on the heuristic path and reviewers note blocked-state lag.
- **Model**: workspaces → tabs → panes, tmux-style prefix keys (Ctrl+B V split, arrows navigate, D detach), mouse-first (click select, drag resize, right-click menus, touch over SSH), adapts to narrow widths.
- **Agent-shaped API**: socket API with `pane read` (visible/recent/detection), `send-text`/`send-keys`/`run`, `agent start|prompt|wait|explain`, `pane wait-output` with regex, plugins placed as overlay/split/tab/popup.
- **Persistence**: sessions survive disconnects, reattach from any terminal / SSH.
- Weaknesses reviewers cite: pre-1.0 churn, bus factor 1, no session resurrection, rendering lag with many panes.
What is striking is how much of herdr Codeman already has, server-side: our hooks give exact `permission_prompt`/`stop`/`idle_prompt` events (herdr's "integration" path, but installed by default), `_confirmIdle()` does the screen-probe fallback, the approvals inbox parses the actual dialog options, and the agent skill + wait primitives are our socket API. What we lack is purely the presentation layer in the terminal.
Prior art for the architecture we want: **agent-deck** (Bubble Tea + tmux) proves the "TUI list + attach into tmux" model works great: session list with live glyphs (● ◐ ○ ✕), Enter attaches into a tmux pane, status polling, groups, fuzzy search. We take the shape, not the code.
Licensing note: herdr is reported variously as Apache-2.0/AGPL-3.0. Irrelevant either way: we copy concepts, never code.
### What we take / what we skip
Take: the four-state sidebar as the organizing principle; grouping by "needs you first"; narrow-width adaptation; mouse support; tmux-familiar keys; the "attention at a glance" framing.
Skip: being a multiplexer. tmux already backs every Codeman session and is a hard dependency; herdr had to build pane management because it owns terminals, we do not. Also skip (for now): plugin marketplace, split layouts, pane drag. Our TUI is a **dashboard + switchboard over tmux**, not a tmux replacement.
## 3. Design: `codeman tui`
One command, one full-screen client of the existing HTTP/SSE API.
**Positioning (owner decision, 2026-08-16): the web UI remains THE primary surface.** The TUI is strictly additive, for users who want a terminal workflow (SSH, Termius, tmux die-hards). Bare `codeman` keeps printing help; nothing existing changes behavior. The `sc` bash chooser also stays untouched for now; flipping its alias to `codeman tui` is deferred to a follow-up release once the TUI has mileage.
↑↓ select · ⏎ attach · 1-9 jump · y/n answer · p prompt · n new · x kill · / search · g digest
```
- **Header**: hostname/instance, server version, session count, plan-usage chip (same telemetry that feeds the web chip, when available). Degrades gracefully when the server is down (see §3.6).
- **Sidebar**: sessions grouped `NEEDS YOU` → `WORKING` → `IDLE` → `RECENT` (past sessions from the unified list, resumable). Within groups, reuse the activity ordering already built for the home screens in PR #303 (blocked first, running longest, quiet newest); that logic is pure and shared.
- **Preview pane**: live tail of the selected session, SGR colors preserved, cursor-movement stripped. When the selected session has a pending approval, the parsed dialog is rendered as a card above the tail with one-key answer bindings.
- **Footer**: contextual keymap (changes when a dialog/confirm is active).
### States and vocabulary
Exactly the web's language so the two surfaces read the same:
| NEEDS YOU (waiting for input) | `✋` | yellow | `idle_prompt` / waiting classification |
| WORKING | `✻` animating through `· ✢ ✳ ∗ ✻ ✽` at 2Hz | green | working classification (the same glyph family Claude itself draws, a deliberate nod) |
| IDLE | `○` | muted | idle |
| RECENT / done | `✔` | muted green | unified list history rows |
Nerd-font/glyph fallback exactly like `sc` does today (`[!] [w] [*] [-] [ok]` when the terminal is not known-capable), plus full NO_COLOR / `tput colors` degradation (8-color and mono renderings are designed, not accidental).
### Keymap
- `↑/↓` or `j/k` select · `Enter` attach · `1-9` jump-attach (parity with `sc`, but now the cursor covers 10+)
- `y`/`n` (or the digit keys) answer the selected session's pending approval right from the dashboard, via `POST /api/approvals/:id/answer`. The server already re-captures the pane and 409s if the dialog is gone, so this is safe by construction.
- `p` send a one-line prompt to the selected session without attaching (`POST /input` with `\r`, the composer opens in the footer)
- `n` new session (case picker → mode picker, drives `POST /api/quick-start`) · `x` kill with typed confirm (never bulk; refuses the session hosting the TUI itself, like tmux-manager.sh does)
- `/` fuzzy search across sessions/history/attachments (`GET /api/search`) · `g` away digest (`GET /api/away-digest`) rendered as a panel
- `r` resume selected RECENT row (unified list `resume-session` flow) · `?` help overlay · `q` quit
- Mouse (phase 3): SGR mouse reporting, click selects, wheel scrolls list/preview, click on footer keys triggers them. Works over SSH, same as herdr's touch story.
### Responsive behavior
The `sc` design constraint survives: below ~72 cols (Termius, iPhone portrait) the preview pane drops and the TUI is a single-column list with two-line rows, nearly identical to today's `sc` but with a cursor, live states, and the answer/prompt/new/kill verbs. The layout switch is width-driven at draw time, no mode flag.
### Attach model
Enter suspends the TUI (restore main screen + cooked mode), then hands the terminal to `tmux -L <socket> attach-session -t <name>` with `stdio: inherit`. On tmux exit/detach, the TUI resumes and refreshes. Full fidelity (mouse, paste, colors) is tmux's, we never proxy bytes.
- Inside tmux already: same socket → `switch-client -t`; different socket → warn about nesting and offer detach-first. `$TMUX` + `CODEMAN_MUX` detection.
- **Return path**: a tmux binding installed for codeman sessions (opt-in) runs `codeman tui --pick` inside `tmux display-popup -E`, a minimal picker-only mode (list + jump, no preview) so switching sessions from inside a pane is one keystroke, fzf-style.
- Optional per-attach chrome (opt-in setting, default off since `status off` at `tmux-manager.ts:1978` is deliberate): a minimal codeman-styled tmux status line showing `name · state · alert`, set on attach, restored on detach.
### Notifications
While the TUI is open and a session flips to NEEDS YOU: flash the row, ring BEL, and optionally emit OSC 9 (desktop notification in kitty/WezTerm/iTerm2, and it traverses SSH). This is the herdr sidebar promise delivered even when the terminal is backgrounded.
### Degraded mode (server down)
`sc` works without the server today and the TUI must too: when no server answers, enumerate `tmux -L codeman list-sessions` + read `state.json` (read-only), show a "server not running" header line, and offer attach only (no states, no approvals). This keeps the "web server crashed, get me to my sessions" path alive.
## 4. Architecture
### A client of the server, not a second brain
Everything live comes from the API the web UI already uses:
| Need | Endpoint |
| --- | --- |
| Session list + history | `GET /api/sessions/unified` |
| Live updates | SSE `GET /api/events` (heartbeat `sse:heartbeat` already exists; fall back to 2s polling) |
| Plan usage chip | latest status-telemetry snapshot (`plan-usage-latest`) |
Server discovery and auth reuse what exists: instance config from `src/config/instance.ts` (`CODEMAN_INSTANCE`, `CODEMAN_PORT`), the probe logic from `daemon-control.ts`, credentials from `~/.codeman/.env` (the established `codeman attach` pattern), self-signed HTTPS accepted for loopback probes (the hooks-on-HTTPS lesson). Multi-user scoping comes free: the API only returns what the authenticated user owns.
### Renderer: hand-rolled, zero new dependencies (decision)
Options considered:
- **Ink (React for CLIs)**: what Claude Code uses. Pros: layout engine, ecosystem. Cons: pulls React into a CLI that today ships only commander+chalk; rerender model fights the two things we care most about (a raw-ANSI preview region and 2Hz glyph animation without flicker); version-pins React for every `npm i -g aicodeman`.
- **blessed/neo-blessed**: unmaintained, skip.
- **Hand-rolled screen core** (recommended): this repo hand-rolls ANSI everywhere already and has the expertise (regex-patterns, stripAnsi, the xterm work). The core is small and boring: alt screen + raw mode + cursor-home full-frame repaint from an off-screen string buffer, throttled to state changes and the 2Hz animation tick, wrapped in DECSET 2026 (synchronized output) where supported so repaints are atomic in modern terminals (tmux, kitty, WezTerm, iTerm2). No diffing needed at these frame rates.
The one genuinely tricky pure function: SGR-aware line clipping for the preview (keep colors, strip cursor movement/OSC/DECSET, clip to width while carrying SGR state, reset at EOL). That is a pure module with exhaustive unit tests, and it is exactly the kind of function Ink would not have given us anyway.
### Module layout
```
src/tui/
tui-app.ts entry + main loop + attach handoff (IO)
tui-client.ts API + SSE client, degraded-mode enumeration (IO)
tui-model.ts pure: state store, grouping, ordering (reuses PR #303 helpers)
tui-layout.ts pure: responsive layout math, row building
tui-ansi.ts pure: SGR-aware clip/filter for the preview
```
Pure modules unit-test with no TTY. `cli.ts` gains one thin `tui` command registration (and `--list`/`<n>` fast paths for `sc -l` / `sc 2` parity, which must stay fast: they short-circuit before any screen setup).
## 5. CLI-wide polish (the rest of "make it much nicer")
A shared style kit, `src/cli-style.ts`: one palette (mirroring the web's status colors), one glyph set with fallback, `heading()`, `kv()`, `table()` (width-aware, fixes the Antigravity overflow), `spinner()` (finally: the 30s silent daemon/service waits get a live line), `confirm()` (used by `reset --force`'s missing prompt and `x` in the TUI). Then the mechanical fixes from §1: colorize `doctor` through the hook that already exists for it, dedupe the `codeman web` startup line, colorize the server's security warning, fix the README `codeman attach` description and the Ctrl+B/Ctrl+A detach drift, TTY/NO_COLOR gates everywhere.
## 6. Phasing
| Phase | Contents | Size |
| --- | --- | --- |
| 0 | `cli-style.ts` + mechanical fixes (§5), real CLI tests (retire the fixture parser in `test/cli-commands.test.ts`) | S |
| 1 | `codeman tui` core: list + states via SSE, cursor + 1-9, attach/return loop, kill w/ confirm, new session, narrow mode, degraded mode, `sc` alias flip + `--list`/`<n>` parity | M/L |
| 4 | Retire `tmux-chooser.sh`/fold `tmux-manager.sh` (keep as thin wrappers for one release), docs/README/wiki, screenshots for promo | S |
Phases 0-1 are the useful minimum; 2 is where it beats herdr's sidebar (answering approvals from the dashboard); 3 is delight.
## 7. Testing
- Pure modules (`tui-model/layout/render/keys/ansi`): plain vitest, frame snapshots as stripped strings plus targeted ANSI assertions.
- Interactive E2E: spawn the built TUI under `node-pty` (already a dependency), feed keys, assert on captured frames; the vitest tmux mock (`IS_TEST_MODE`) keeps attach paths inert. Port rules per CLAUDE.md (3150+, `app.inject()` where possible by testing `tui-client` against injected routes).
- tmux socket and data dir always via instance config (`dataPath()`, `-L codeman`); a beta instance TUI sees only its own world.
- Never bulk kill, always confirm, never touch another session implicitly, refuse killing the session the TUI runs in (w1/w2/w3 are sacred).
- Input is single-line with `\r`, via the server (never raw tmux send-keys from the TUI while the server owns the session).
- Approvals answering goes through the server's re-capture + 409 path, never blind keystrokes.
- `status off` on panes stays the default; any chrome is opt-in.
- No new runtime dependencies; the npm package stays light.
## 9. Decisions (resolved 2026-08-16)
1. **Bare `codeman` does NOT open the TUI** (owner decision): the web UI is the main thing, the TUI is additional. `codeman tui` only.
2. **`sc` stays the bash chooser for now**; the alias flip is a follow-up once the TUI has mileage. `codeman tui --list` / `codeman tui <n>` provide the same fast paths for people who want to switch.
3. Opt-in tmux status line: deferred to phase 3 along with the `--pick` popup switcher.
4. Preview tail goes over the API (auth/multi-user/remote-consistent); previews are simply unavailable in degraded server-down mode.
5. Name is `codeman tui` (the test fixture historically expected it).
Initial PR scope: phases 0-2. Phase 3 (mouse, popup switcher, status line, OSC 9) and phase 4 (bash chooser retirement) are follow-ups.
`codeman tui` is a full-screen dashboard for your Codeman sessions, in the terminal.
It shows every session grouped by whether it needs you, lets you answer a permission
dialog or send a prompt without switching anywhere, and puts you inside a session's
tmux pane with one keystroke.
It is **additional, not a replacement**: the web UI stays the primary surface and
gets every feature first. The TUI exists for the terminal workflow (SSH, Termius,
a tmux window you keep open all day), and it is a *client* of the running server,
so the two surfaces can never disagree about what a session is doing. It is also
not a multiplexer: tmux still owns every pane, and attaching hands the terminal to
tmux rather than proxying bytes.
## Starting it
```bash
codeman tui # the dashboard
codeman tui --list # print the numbered session list and exit
codeman tui 2 # attach straight to session 2 of that list
```
The two fast paths are the scriptable ones.
Neither sets up a screen, so both are as quick as the one API call they make, and
`--list` prints plain text when piped, so it composes with `grep`/`awk`.
What it needs:
| Needs | What you get |
| --- | --- |
| **Full features** | A running Codeman server (states, approvals, preview, prompts, search, digest). The TUI finds it the way `codeman attach` does: `CODEMAN_API_URL`, else loopback on `CODEMAN_PORT` for this `CODEMAN_INSTANCE`. The self-signed certificate an `--https` install generates is accepted, as it is everywhere else in the CLI. |
| **Server down** | It still starts, in **degraded mode**: sessions are enumerated straight from `tmux -L codeman` plus a read-only peek at `state.json`, and attach is the only verb. See [Troubleshooting](#troubleshooting). |
| **A terminal** | `codeman tui` refuses to run when stdin/stdout are not a TTY, and says to use `--list` instead. A cron job or a pipe therefore fails loudly rather than emitting escape codes into a log. |
## What it looks like
A real frame at 100x30 (`NO_COLOR`, trailing blank rows trimmed). The selected
session has a pending permission dialog, so the preview pane leads with the card:
│ 1 -import { SessionManager } from "../session-manager.js";
│ 2 +import type { SessionPort } from "../web/ports/session-
│
│ Bash(npm run typecheck)
│ └ tsc --noEmit: no errors
│
│ ✻ Actualizing… (2m 14s · ↓ 12.3k tokens)
↑↓ select · ⏎ attach · y approve · n deny · 1-9 option · p prompt · x kill · / search · g digest ·
```
- **Header**: the machine, the server version, how many sessions are live, and the
plan-usage chip (the same statusLine telemetry that feeds the web chip, when the
server has a snapshot). A `⚠ n` badge counts pending approvals.
- **Sidebar**: every session, grouped and numbered.
- **Preview**: a live tail of the selected session, its own colors preserved, with
the parsed dialog card on top when that session is blocked.
- **Footer**: only the keys that work right now. `n` reads `n new` normally and
`n deny` when the selected session has a dialog, because it cannot be both.
The same world through `--list`:
```
1 waiting w6-docs /home/you/dev/docs
2 blocked w4-api-refactor /home/you/dev/api
3 working w1-codeman /home/you/dev/codeman
4 working w2-gallery /home/you/dev/gallery
5 idle w3-promo /home/you/dev/promo
6 done api-hotfix /home/you/dev/api
```
The numbers are the same on both surfaces, so `codeman tui --list` then
`codeman tui 4` is one thought.
## The four groups
Groups are always in this order, and a session is in exactly one of them:
| Group | Glyph | Means | Comes from |
| --- | --- | --- | --- |
| **NEEDS YOU** | `⚠` | A permission or question dialog is blocking the agent | The approvals inbox (`permission_prompt` hooks, with the on-screen options parsed) |
| | `✋` | Waiting for your next instruction, or errored | `idle_prompt`, or an errored session (equally something only a human clears) |
| **WORKING** | `✻` animating | A turn is running | The same working classification the web dashboard uses |
| **IDLE** | `○` | Live, but sitting there | |
| **RECENT** | `✔` | A past session from the unified list | History rows, no live pane |
Ordering inside a group is "the one that has waited longest, first": blocked
sessions sort by how long the dialog has been up, working sessions by when their
turn started (the pane's last Enter, since a working pane repaints every second
and would otherwise always look freshly started), and quiet ones by last activity.
That is the ordering the web home screens already use.
The cursor sticks to a **session**, not a row number, so a session that jumps to
NEEDS YOU does not drag your selection with it. The number beside each row is what
`1-9` and `codeman tui <n>` mean, and it is renumbered on every re-sort.
When a new dialog appears, the terminal bell rings once, for that dialog only: the
same item announced twice does not ring twice.
## Keymap
| Key | Does |
| --- | --- |
| `↑``↓` or `j``k` | Move the cursor. PageUp/PageDown jump five rows. |
| `Enter` | Attach to the selected session (see [Attaching](#attaching)) |
| `1`-`9` | Jump to that row and attach. When a dialog is on screen, a digit answers it instead (see below). |
| `y` | Approve the selected session's dialog |
| `n` | Deny it, or **start a new session** when there is no dialog |
| `p` | Send one line to the selected session without attaching |
| `x` | Kill the selected session; `y` confirms, any other key cancels |
| `/` | Search sessions, events and files |
| `g` | Away digest: what happened while you were gone |
| `?` | Help overlay |
| `Esc` | Close whatever overlay is open |
| `q` or `Ctrl+C` | Quit, restoring the screen you started with |
Inside the `p` composer and the `/` query: `←``→``Home``End``Delete`
`Backspace` plus `Ctrl+A` / `Ctrl+E` / `Ctrl+U` / `Ctrl+W`, `Enter` to send or open,
`Esc` (or `Ctrl+C`) to cancel. In the kill confirmation you retype the session name;
anything else cancels. In the `n` pickers, type to filter, `Enter` chooses.
Verbs that need the server (`y`/`n`/`p`/`x`/`/`/`g`) say so in degraded mode
instead of failing silently; `Enter` and `1-9` keep working.
### `p` sends exactly one line
The composer is a single line by design, ending in a carriage return: that is the
input contract every Codeman path follows, because multi-line text breaks the
agent's own composer. Pasted newlines become spaces rather than being rejected, so
a paste cannot silently run a different command than the one you read.
## Answering approvals
This is the thing the terminal could not do before. Select a blocked session and:
- `y` approves.
- `n` picks the parsed "No" option, or sends Esc when the dialog did not parse one.
- A digit picks that numbered option, **but only a digit the dialog actually
offers**. A digit with no matching option falls through to the list's own
jump-and-attach binding, so it can never be typed at whatever has focus.
The answer goes through `POST /api/approvals/:id/answer`, which **re-captures the
pane before it types anything**. If the dialog is no longer on screen (you answered
it in tmux a moment ago, or the agent moved on), the server refuses with a 409 and
the TUI says `that dialog is no longer on screen` rather than pressing a key into a
live composer. The answer is scoped to the options the server parsed off the actual
frame, never to a guess.
An idle prompt (`✋`) is not a dialog: there is nothing to approve, so `p` is the
reply path and the footer says `p reply` instead of `p prompt`.
## Attaching
`Enter` suspends the dashboard (main screen back, cooked mode back) and hands the
terminal to tmux with `stdio: inherit`. Colors, mouse and paste are tmux's, at full
fidelity.
**Press `F1` to come back.** One key, no modifier to hold or release, nothing to
type in a particular order. tmux's own way out is a chord — press the prefix, let
go, then a letter — and beta testing showed that is genuinely hard to convey: the
bar first named the wrong letter (tmux binds lowercase `d` to `detach-client` and
capital `D` to `choose-client`), and once corrected it still failed for anyone who
kept Ctrl held, because that sends `Ctrl+D`, which tmux leaves unbound. So the TUI
claims `F1` in tmux's prefix-less key table for the length of the attach and gives
it back afterwards. The chord still works; it is simply not what you are told to
press.
You do not have to remember any of it. For as long as the attach lasts the pane
wears a bar across the top:
```
1 w3-codeman-… 2 w4-codeman-… 3 testcase … alt+1-9 switch · F1 back to the codeman dashboard
```
That is the **session strip**: the other sessions stay visible from inside a pane,
numbered exactly as the dashboard numbers them, with the one you are in inverted.
`Alt+1`..`Alt+9` switch between them without going back to the dashboard first. With
more sessions than fit, the strip shows a window around the current one and marks
each cut end with `…`; the way-out hint is measured first and always keeps its space.
Codeman keeps the status bar off on its panes (the web UI carries that information
around the terminal instead), so the TUI turns it on for the attach and puts it back
exactly as it was on detach, along with each window's size. Every session the strip
can switch to is dressed and sized the same way, so switching is instant and lands
in a pane that already fills your terminal.
Detaching leaves the agent running; typing `exit` or pressing `Ctrl+D` would end it,
which is the difference the bar exists to make obvious. If an agent does exit, its
pane stays as a corpse: the TUI refuses to attach to a dead pane and offers `r` to
resume the conversation in a fresh one instead.
Three cases:
| Where you are | What happens |
| --- | --- |
| Not in tmux | `tmux -L codeman attach-session` |
| Already in tmux on Codeman's socket | `switch-client`, so you do not nest |
| In tmux on a **different** socket | Refused, with an explanation: detach from that tmux first, then run `codeman tui` again |
A direct-PTY session has no pane to attach to, and says so.
**`Enter` on a RECENT row resumes that conversation** instead: there is no pane to
attach to, so the TUI creates a new claude session carrying the old transcript
(`resumeSessionId`, exactly what the web UI's "Resume Conversation" list does), in
the directory it originally ran in and under its old name, then attaches to it. It
is claude-only, and a row with no working directory or no conversation id says why
rather than resuming something else.
`x` never bulk-kills: it kills one session, only after you retype its name, never a
history row, and never the session the TUI itself is running in.
## Over SSH, and on a phone
The TUI is an ordinary terminal program with no local dependencies beyond tmux, so
`ssh box` then `codeman tui` works exactly like running it locally. There is no
separate remote mode.
Below 72 columns (Termius, an iPhone in portrait) the preview pane is dropped and
rows take two lines each, keeping the cursor, the live states and the
answer/prompt/kill verbs. The switch is
width-driven at draw time, so unfolding a foldable or resizing a window re-lays out
immediately; there is no mode flag to set.
## Troubleshooting
**"The Codeman server rejected these credentials."** The server has
`CODEMAN_PASSWORD` set. Export `CODEMAN_PASSWORD` (and `CODEMAN_USERNAME` if it is
not `admin`), or put them in the data dir's `.env` (`~/.codeman/.env`), which is
where `codeman attach` already reads them from.
**`server not running: attach only`** in a yellow banner. Nothing answered on the
expected port, so the TUI fell back to enumerating tmux. You get names and attach;
you do not get states, approvals or previews, because those only exist on the
server. Start the server (`codeman web -d`, or `systemctl --user start codeman-web`)
and the banner clears on its own: the TUI keeps re-probing.
**It found the wrong server, or none.** Discovery is instance-scoped. A beta
instance (`CODEMAN_INSTANCE=beta`) has its own data dir *and* its own tmux socket,
so its TUI sees only its own sessions. Set `CODEMAN_PORT` or `CODEMAN_API_URL`
explicitly when you run more than one.
**"this terminal is already inside tmux on socket ..."** You are in a tmux session
on a socket that is not Codeman's, so attaching would nest two multiplexers whose
prefix keys collide. Detach from that tmux and run `codeman tui` from outside.
**Boxes and glyphs render as garbage.** The TUI picks a glyph tier from the
environment: no `TERM` (or `dumb`), or a non-UTF-8 locale, gets the ASCII set
(`[!] [w] [*] [-]`, `+`/`-`/`|` frames). Force it either way with
`CODEMAN_TUI_GLYPHS=ascii|unicode|nerd`.
**Colors.** Standard `NO_COLOR` / `FORCE_COLOR` handling (chalk's, the same as the
rest of the CLI). Under `NO_COLOR` the frame is cursor addressing and text only,
and the preview's own colors are stripped too, so a session's output cannot repaint
the dashboard.
**It refuses to open at all**, saying it needs an interactive terminal. stdout or
stdin is not a TTY. That is the guard: use `codeman tui --list`.
## Related
- [`docs/tui-plan.md`](tui-plan.md): the design record. Why hand-rolled ANSI, why a
client and not a second brain, and what is deliberately deferred.
- [`docs/approvals-inbox-plan.md`](approvals-inbox-plan.md): where the parsed
dialogs and the answer endpoint come from.
- [`docs/remote-sessions.md`](remote-sessions.md): remote-SSH cases, which the TUI
> **Status: SHIPPED — deployed to prod + pushed to master, not yet released (2026-06-14).** App Settings → Display → **Plan Usage Limits** (`showPlanUsageLimits`). **Default changed in 1.9.3: desktop now defaults ON, handhelds stay OFF, resolved via `planUsageChipEnabled()`.** The per-device notes further down describing it as opt-in/synced record the original 2026-06-14 shape, not current behavior. Commits `c82f6c8` (feature) → `4d9d93d` (end-to-end fixes) → `eae225b` (per-user reconcile) → `95fb5fc` (init-snapshot replay). Full suite green (2869), CI green. No changeset/version bump yet.
>
> **2026-09-07 rework — the "Injection lifecycle" section below (disk-write reconcile via `applyStatusLineConfig`) is SUPERSEDED and describes the OLD mechanism, kept for history.** That disk write let a Codeman-marked `statusLine.command` in `.claude/settings.local.json` take precedence over the user's own global/project statusline for ANY `claude` run in that directory — including entirely outside Codeman — with no disclosure and no way to undo it (real bug, found 2026-08-31). The exporter is now injected as an EPHEMERAL `claude --settings` CLI flag at spawn (`resolveStatusLineCliCommand`/`ensureStatusLineExporterScript`, hooks-config.ts) — never written to disk — and it WRAPS the user's own real statusline (`findEffectiveUserStatusLineCommand`) rather than replacing it. `showPlanUsageLimits` now doubles as the telemetry COLLECTION switch too: `readPlanUsageTelemetryEnabled()` reads it fresh from `settings.json` at every claude session create/respawn (`TmuxManager.createSession`/`respawnPane`), so it applies uniformly across every claude-creation path — interactive Run, cron, the Ralph Loop API, quick-start — with no per-session state (a Codeman restart cannot silently kill it) and no per-request field on the wire at all. An absent key reads as ON (the reader resolves the default; `GET /api/settings` never writes), and a settings save carries the key only when it flips the chip on that device, so a handheld with the chip off cannot switch collection off for a desktop by saving something unrelated. The exporter prints nothing on failure rather than the bare word `codeman` (discussion #405).
@@ -110,33 +112,33 @@ Fixed path (sessionId in the **body**, not the URL) so the auth exemption is an
2. **Fresh load / reconnect:** server stores the latest in `plan-usage-latest.ts`; `getLightState()` includes it as `planUsage`; the per-connection **init snapshot** replays it; `handleInit` paints the chip immediately (authoritative over localStorage). Null until the first telemetry of the process.
### 5. Injection lifecycle — works for *any* user, never self-destructs
### 5. Injection lifecycle (SUPERSEDED 2026-09-07 — see header note; kept for history)
The setting `showPlanUsageLimits` is **synced** (in `settings.json`, not a per-device `displayKey`).
- **On toggle** (`PUT /api/settings`, `system-routes.ts`): reconcile the exporter across **all active Claude sessions' working dirs** — inject on enable, remove on disable. Server-side and authoritative, so existing sessions get the footer + feed the chip *immediately*, no new session needed, no dependency on a client's synced localStorage.
- **On session create** (`session-routes.ts`): **ADD-ONLY** — inject when `statusLineTelemetry` is true; **never remove**. Sessions in a repo share one `settings.local.json`, so a single create-with-false (e.g. a client whose synced setting hadn't loaded) must not yank the statusLine out from under other live sessions. Removal happens only via the explicit toggle.
- `applyStatusLineConfig()` is **`isOurs`-guarded** (matches`/api/status-telemetry`), so a user's own hand-authored statusLine is never touched, and it **updates an out-of-date ours-command** so fixes (e.g. `-k`) propagate. **No `CASES_DIR` gate** — runs for linked cases / real repos (where sessions actually run), mirroring `updateCaseModel`.
- ~~**On toggle** (`PUT /api/settings`, `system-routes.ts`): reconcile the exporter across **all active Claude sessions' working dirs** — inject on enable, remove on disable.~~ There is nothing to (re)inject into an already-running session under the new CLI-flag mechanism — the NEXT respawn (a Ralph cycle, `/clear`, a PTY-exit restart) already reads the setting fresh.
- ~~**On session create** (`session-routes.ts`): **ADD-ONLY** — inject when `statusLineTelemetry` is true; **never remove**.~~ There is no `statusLineTelemetry` request field anymore. `TmuxManager.createSession`/`respawnPane` read `readPlanUsageTelemetryEnabled()` fresh at spawn instead, uniformly across every claude-creation path.
- ~~`applyStatusLineConfig()` is **`isOurs`-guarded**~~ —`applyStatusLineConfig` still exists but only for the SELF-HEAL path now (`resolveStatusLineCliCommand` strips a legacy disk-written exporter the first time a session starts in a workspace an older Codeman build touched).
## Codeman-specific considerations
1. **Account-global limits.** The 5h/7d pools are shared across all sessions on the account → one shared header chip (freshest sample wins), not a per-tab bar.
2. **The footer is owned, by necessity.** A statusLine command always replaces Claude's default footer. Since `rate_limits`*only* arrives via statusLine, we reconstruct a useful **session-status** footer (model · tokens · ctx %) from the same payload rather than showing the limits there.
3. **`isOurs`-guarded.** Never removes/overwrites a user's own statusLine on disable; only manages the Codeman exporter.
3. **Never overwrites, now WRAPS.** The exporter composes with a user's own real statusline (`findEffectiveUserStatusLineCommand`) rather than replacing it; `applyStatusLineConfig`'s `isOurs`-guard now only backs the legacy self-heal removal path.
4. **Security envelope unchanged.** The exporter runs arbitrary shell every render — same trust model as the hook curls (localhost + `$CODEMAN_HOOK_SECRET_FILE`); reuses the hook-secret gate.
5. **Claude-only.** OpenCode/Codex emit no `rate_limits` JSON; injection is gated to`mode === 'claude'`.
5. **Claude-only, registry-gated.** Injection is gated on `getCli(mode)?.capabilities.statusLineTelemetry` (currently `true` only for claude) rather than a hardcoded`mode === 'claude'` string.
6. **Future — auto-resume synergy.** Live percentages would let `SessionAutoOps` pre-arm *before* the wall instead of reacting to the stall footer. Not built.
- `src/hooks-config.ts` — `resolveStatusLineCliCommand()`/`ensureStatusLineExporterScript()` (ephemeral CLI-flag injection, never disk), `findEffectiveUserStatusLineCommand()` (wrap the user's real statusline), `readPlanUsageTelemetryEnabled()` (fresh global-setting read), `applyStatusLineConfig()` (legacy self-heal removal only now).
- `src/session-cli-registry-bridge.ts` — merges the exporter path into the SAME `--settings` JSON object as effort/ultracode (Claude Code accepts only one `--settings` flag per invocation).
- [`docs/architecture-invariants.md`](https://github.com/Ark0N/Codeman/blob/master/docs/architecture-invariants.md) - the mechanisms behind all of this, for contributors.
| Claude Code plugin | `/plugin marketplace add Ark0N/Codeman`, then `/plugin install codeman@codeman` | Global, through Claude Code's plugin manager. `/plugin update codeman` follows releases. Pick this or `codeman skill install`, not both, or the skill is listed twice (`codeman` and `codeman:codeman`). |
| "What sessions are running right now?" | Lists them with name, mode, and status. Read-only. |
| "Start a shell worker on the `myapp` case, run the test suite, tell me if it passes." | Spawns, waits on a completion marker, reads the exit code, cleans up. |
| "Spin up 3 workers for lint, typecheck and tests, run them in parallel, report failures." | One session per task, all started first, then gathered as each finishes. |
| "Have a claude worker summarize `src/session.ts`, then close it." | Spawns, runs the readiness ladder, sends and waits, reads the answer, deletes the session. |
| "Watch session w4 and tell me if it gets stuck on a permission prompt." | Blocks on the `blocked` signal and surfaces the question to **you**. |
Sessions the agent creates get deleted when it is done. You can watch the tabs appear and
disappear in the dashboard while it works.
### What it will and will not do
- **It self-gates.** Outside a Codeman session it refuses to act and does not guess an API
URL, so a global install costs an unrelated Claude Code session nothing.
- **Unprompted, it may only** spawn sessions, prompt them, and delete ones **it created in
that conversation, by exact id**, behind a guard that refuses to delete the agent's own
session.
- **It will not** answer another session's permission prompt on your behalf. It surfaces the
question instead.
- **Deleting a case** (which erases a real directory of your code), bulk kills, respawn,
Ralph, cron, orchestrator, and settings writes all require you to ask, naming the target.
Turning the setting back off **does not remove already-injected copies**, because a
create-time sweep would yank the skill out from under other live sessions sharing that
directory. Remove them per case with `codeman skill uninstall --case <name>`.
The skill ships with the verb index always loaded, plus on-demand references for the verbs,
worked multi-worker recipes, endpoint tables, and cross-session messaging.
## The manual path
The same operations as raw HTTP, for a CI bot, a shell script, or an agent without skill
support.
### Detect that you are inside Codeman
These are set in every managed session. Read them rather than hardcoding anything:
| `CODEMAN_MUX=1` | You are in a managed tmux session. Never `tmux kill-session`, `pkill claude`, or `pkill tmux`: you will kill yourself or a sibling. |
| `CODEMAN_API_URL` | Base URL, with the correct scheme. |
| `CODEMAN_SESSION_ID` | Your own session id. Use it to avoid acting on yourself. |
| `CODEMAN_HOOK_SECRET_FILE` | Path to the hook secret. |
### Rules of the road
Read these before writing any code. Each one has cost somebody an afternoon.
1. **Input is single line and must end with `\r`.** Enter fires only when the payload
contains a carriage return. Without it the text sits unsubmitted on the prompt, the
request still succeeds, and a combined wait burns its full timeout on a turn that never
started. Embedded newlines are stripped rather than rejected, so `"echo A\necho B\r"` runs
the joined `echo Aecho B`. One line per call.
2. **Make input idempotent.** Send a stable `clientId` and a monotonic per-session `seq`. The
server deduplicates, so a retry after a dropped connection cannot double-deliver.
3. **Auth.** With `CODEMAN_PASSWORD` set, use HTTP Basic or the session cookie. A missing
`Origin` is allowed, so plain curl works. A `401` replies with the bare string
`Unauthorized`, **not** the JSON envelope, so piping it into `jq` throws a parse error
instead of showing the failure. Check the status before parsing.
4. **Envelope.** Most endpoints return `{ "success": true, "data": ... }`. A few legacy GETs
return bare bodies, so handle both: `body.data ?? body`.
5. **Wait instead of polling, and a timeout is not an error.** The wait endpoints answer
`200` with `wait.timedOut: true`. Loop over short waits rather than one long call, because
tunnels cut idle connections.
6. **Only `claude` sessions emit `stop` and `blocked`.** They come from Claude Code hooks.
Shell and the external CLIs accept only `idle`, `working`, and `exit`; asking for `stop`
explicitly there is a `400`, while omitting `until` is always safe. On a shell session
`idle` fires **once at startup and never again**, so synchronize hook-less sessions with an
output marker instead.
7. **Nothing reports "ready", so wait for it explicitly.** A new session answers
`{"signal":"exit","immediate":true}` until its PID exists, and that means *not started*,
not *crashed*. A Claude worker in a fresh case then sits on the CLI's trust dialog; prompt
it there and the wait resolves on idle in about two seconds looking exactly like a finished
turn, while your text sits stuck in the dialog.
### Recipes
```bash
API="${CODEMAN_API_URL:-http://localhost:3000}"
# Add -u admin:"$CODEMAN_PASSWORD" if a password is set, and -k on an HTTPS install.
| OS | macOS or Linux. Windows works through WSL2. |
| Node.js | 22 or newer. |
| tmux | Required. Sessions live in tmux, which is what makes them survive restarts. |
| An agent CLI | At least one of Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi. Plain shell sessions need none. |
| Network | Binds to `127.0.0.1` by default. Reaching it from another device is a deliberate step: see [Remote Access](Remote-Access). |
Codeman is MIT licensed, self-hosted, and sends no telemetry. Everything runs on your
machine.
## Getting help
- **Questions and setup help**: [Discussions](https://github.com/Ark0N/Codeman/discussions), especially [Q&A](https://github.com/Ark0N/Codeman/discussions/categories/q-a).
- **Bugs**: [Issues](https://github.com/Ark0N/Codeman/issues). Include your OS, install method, browser, and which CLI the session was running.
- **Ideas and roadmap**: [Ideas](https://github.com/Ark0N/Codeman/discussions/categories/ideas).
- **Security**: never a public issue. See [SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md).
| **macOS or Linux** | Windows works through WSL2. See [Windows](#windows-wsl) below. |
| **Node.js 22+** | The installer offers to install it if missing. |
| **tmux** | Not optional. Sessions live inside tmux, which is what makes them survive a server restart, a dropped connection, or a closed laptop. |
| **An agent CLI** | At least one of [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev). Plain shell sessions need none. See [Agent CLIs](Agent-CLIs). |
Codeman itself sends no telemetry and phones no home. The only network traffic is your
browser to your server, and whatever the agent CLI you chose does on its own.
## Route A: the installer (recommended)
```bash
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js and tmux if they are missing, clones Codeman into `~/.codeman/app`,
and builds it.
What it asks you:
1. **Permission for every system change.** Package installs and agent CLI downloads are
prompted individually. Nothing is installed silently.
2. **How the dashboard should be reachable.** Three choices:
- **Tailscale** (recommended for phone access): keeps the loopback bind and walks you
through `tailscale serve`, including the tailnet HTTPS toggle, then verifies the result
end to end.
- **Your local network** (`0.0.0.0`): prompts for a password. Skipping the password takes
an explicit confirmation and ends on a loud warning.
- **This machine only** (`127.0.0.1`): the safest option, and the default for a bare
`codeman web` regardless of what you pick here.
Which one is preselected depends on what the installer finds. A fresh install defaults to
the local network, unless Tailscale is already connected, in which case it defaults to
Tailscale. An existing loopback install defaults to keeping loopback, or to Tailscale when
a serve mapping for Codeman is already there. A bare Enter never pulls in new software,
and a non-interactive run always keeps the safe loopback default.
3. **What to do when it finishes.** Run in this terminal, install as a background service
that starts on boot, or do nothing yet.
Re-running the same one-liner **updates an existing install in place**. Local changes in
`~/.codeman/app` are stashed rather than discarded, a running service is restarted and
verified, and your existing network binding is preserved. An interrupted first install
resumes instead of restarting.
Two other entry points exist:
```bash
install.sh update # update only
install.sh uninstall # remove
install.sh tailscale # retrofit Tailscale access onto an existing install
```
**Automation and CI**: with no terminal attached, any step that would change the system
aborts with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to
approve those steps. `CODEMAN_TAILSCALE=1` preselects the Tailscale answer, and never
installs Tailscale itself non-interactively.
## Route B: npm
```bash
npm install -g aicodeman
codeman web
```
The npm package is named `aicodeman`; the product is Codeman. Both `codeman` and
`aicodeman` are installed as commands.
The trade-off against Route A: no guided network setup, and the in-app self-updater does
not apply. npm installs report as non-updatable in **App Settings → System → Updates**, and
you update with `npm update -g aicodeman`.
## Route C: git clone
For contributing, or for running unreleased code.
```bash
git clone https://github.com/Ark0N/Codeman.git
cd Codeman
npm install # postinstall builds the vendored xterm addon bundles
npm run dev # dev server on http://localhost:3000
```
For a production run from a clone:
```bash
npm run build
npm run start
```
`npm run dev` runs TypeScript directly through `tsx` with no build step. The frontend is
plain JavaScript served from `src/web/public/` with no bundler, so editing a `.js` or `.css`
file and reloading the page is enough. The one exception is `index.html`, which is read once
at server start, so markup changes need a restart.
See [Contributing](Contributing) for the rest of the development loop.
## Installing an agent CLI
Codeman drives CLIs, it does not bundle them. Install at least one:
| **Claude Code** | `npm i -g @anthropic-ai/claude-code` | The primary target. Some Codeman features are Claude-only: see [Agent CLIs](Agent-CLIs). |
| **OpenCode** | See [opencode.ai](https://opencode.ai) | |
| **Codex** | See [developers.openai.com/codex/cli](https://developers.openai.com/codex/cli) | |
| **Antigravity** | See [antigravity.google](https://antigravity.google) | Google's successor to the consumer Gemini CLI. |
| **Gemini CLI** | See [github.com/google-gemini/gemini-cli](https://github.com/google-gemini/gemini-cli) | Enterprise only since Google's June 2026 consumer cutover. |
| **Pi** | See [pi.dev](https://pi.dev) | No permission prompts and no sandbox by design. Read [Agent CLIs](Agent-CLIs) before using it on a repo you care about. |
Log each CLI in once, by hand, before pointing Codeman at it. Codeman never collects or
stores your CLI credentials.
## Verify the install
```bash
codeman doctor # checks Node, tmux, the agent CLIs, document converters
codeman --version
codeman web # then open http://localhost:3000
```
`codeman doctor --json` gives machine-readable output, and `--category core` narrows it to
the things a session cannot start without.
If the dashboard loads and **+ New Session** opens, you are done. Continue to
Codeman on a phone is not a shrunken desktop UI. It is the surface most of its design
attention has gone into, because checking on an agent from a bus is the thing this software
is for.
<p align="center">
<img src="https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/screenshots/mobile-session-keyboard-20260727.png" alt="Answering an agent prompt on a phone" width="300">
</p>
## Getting there
1. **Set up access.** Tailscale is the recommended route and gives you real HTTPS. See
[Remote Access](Remote-Access).
2. **Log in by QR.** Open the dashboard on your desktop and scan the code. No password
typing. Tokens are single use and rotate every 60 seconds.
3. **Install it to your home screen.** On iOS this is mandatory for push notifications;
Safari does not deliver push to tabs. On Android it makes the app full screen.
HTTPS matters for more than security here: microphone access and push notifications both
| **Create New** | Starting a fresh project. Creates `~/codeman-cases/<name>` and scaffolds a `CLAUDE.md` into it. |
| **Clone Repo** | Working on an existing public repo. Paste the URL; Codeman preflights it as you type, offers the repo's real branches and tags, and fills in the case name. |
| **Link Existing** | The code is already on disk. Point at the folder, with **Browse** if you would rather click than type. |
The gear next to the picker holds two per-case toggles: **Agent Teams** and
**1M Opus Context**. Both are off by default and both are safe to ignore for now.
**Create New** also has a checkbox for running the case inside a Docker container, and a
**Remote** panel for running it over SSH on another machine. Those are
[Docker Cases](Docker-Cases) and [Remote SSH Sessions](Remote-SSH-Sessions); skip them for
your first session.
## 4. Pick a run mode and hit Run
The **Run** button starts an agent in the selected case. The arrow next to it picks which
| `403 host not allowed` | Your domain is not in the allowlist. Set `CODEMAN_ALLOWED_HOSTS`. |
| Assets 404 / blank page under a sub-path | Start Codeman with `--base-url /<prefix>` and have the proxy forward the prefix unchanged (don't strip it). |
| Phone shows the login page but the terminal never connects | The proxy is not forwarding WebSocket upgrades. |
| Browser warns about the certificate | Expected with `--https` and its self-signed certificate. Tailscale gives you a real one instead. |
| LAN IP does not respond, but a tunnel to the same box works | The server is bound to loopback. That is the default. A tunnel reaches it; a LAN browser cannot. |
| Hooks stopped working after switching to HTTPS | Hook callbacks need `-k` for the self-signed certificate. Recent versions self-heal existing cases; if yours predates that, recreate the case's hooks. |
| Everything is slow over the tunnel | Quick tunnels route through Cloudflare's edge. Tailscale is usually a direct connection and much faster. |
## Read next
- [Security](Security) - the whole model, and the hardening checklist.
- [Mobile Guide](Mobile-Guide) - once you can reach it from the phone.
- [Running As A Service](Running-As-A-Service) - keeping server and tunnel up across reboots.
- [`docs/security-architecture.md`](https://github.com/Ark0N/Codeman/blob/master/docs/security-architecture.md) - the full model.
| **Host-header allowlist** | DNS rebinding. A domain rebound to `127.0.0.1` is rejected before any handler runs. Add your own domains with `CODEMAN_ALLOWED_HOSTS`. |
| **Cross-site Origin guard** | CSRF on state-changing requests. A *missing* Origin is allowed so curl, the CLI, and hooks keep working; a foreign or opaque one is rejected. |
| **Raw `text/plain` bodies** | The CORS simple-request CSRF vector, where a cross-site form could smuggle JSON into a write route with no preflight. |
| **WebSocket origin check** | Cross-site WebSocket hijacking. The terminal upgrade closes with code `4003` on failure. |
| **Security headers** | A strict content security policy, `nosniff`, frame options, and HSTS over HTTPS. CORS is reflected only for loopback origins. |
| **HTTP Basic** | `CODEMAN_USERNAME` (default `admin`) and `CODEMAN_PASSWORD`. |
| **Session cookie** | A 256-bit opaque token validated server side, so it cannot be forged offline. 24 hours, extended on activity, with a device-context audit trail. |
| **Rate limiting** | Ten failed attempts per IP produce a `429` with a 15 minute decay. A correct password or valid cookie recovers immediately even under attack, which matters because all tunnel traffic shares one loopback address. |
| **QR auth** | Single-use 60-second tokens with their own separate rate limiter, so a mistyped password cannot lock out QR login. |
| **Hook endpoints** | The hook and telemetry endpoints skip Basic auth because they are called from localhost by the CLI, but when auth is on, that bypass additionally requires a per-instance hook secret. |
## File access
Three separate file surfaces, each confined differently, because a single shared rule would
| **File Viewer** | Real path resolution before boundary checks, so symlinks cannot escape. Sensitive trees blocked. Edit mode adds an extension allowlist, a size cap, `.git` denial, and optimistic concurrency. It never creates files. |
| **Attachments** | An id-based registry, so browser requests never carry absolute paths. The magic-link scanner is prompt-injectable by nature and is therefore force-confined to the session's workspace. Extension allowlist, not a blocklist. |
| **Path picker** | Its own root allowlist rather than the workspace confinement. In multi-user mode a non-admin gets only their own user space, because per-user spaces live inside the home directory. |
| **Per-device, local** | Skin, WebGL renderer, local echo, CJK input, extended keyboard bar, File Viewer and Cron header buttons. Never sent to the server at all. |
| **Per-device policy** | Most `show*` toggles, plan usage chip, language. Stored server-side, but a device only takes the server value when it has no local one of its own. |
| **Synced** | Models, effort, CLI options, notification preferences, voice settings, display name, the agent skill and approvals toggles. |
The practical rule: **appearance and input are per device, behaviour is shared.** If a change
did not follow you, it is in one of the first two rows, and you change it again on that
device.
## App Settings
### Updates
Current version, a manual check, and the in-app updater. Covers git-clone installs
supervised by systemd or launchd; npm installs report as non-updatable. See
| Skin | Theme palettes, light ones included. Applied before first paint, so no flash of the wrong theme. |
| Entrance Animations | Per-surface animation styles for tabs, terminals, windows, and lineage lines. All default to the legacy no-animation behaviour. |
| Display Name | Your name in the UI. Cosmetic only; it never renames the package, CLI, API, or storage. |
| Interface Language | English or Simplified Chinese. Per device. |
| Session List Layout | Header tab strip (default) or a collapsible left sidebar. See [The Dashboard](The-Dashboard#session-list-layout). |
| Tall Tabs | Taller tab strip. |
| Pop-out Button on Tabs | Adds the detach control to tabs, with a per-tab override. |
| Spawn Lineage Lines | Arcs from a parent tab to sessions it spawned. Desktop only, on by default. |
| Overview Home Screen | The phone home screen. On by default. |
### Models
Claude model cards, the 1M context window switch, and the thinking effort segment. The cards
and the switch compose into one model choice, so there is no separate "which one wins"
question.
Model and effort are both **soft defaults**: the model is written into the case's
`.claude/settings.local.json` and effort is passed at start, so `/model` and `/effort`
| Startup Mode | Claude's permission mode for new sessions. Default skips prompts; `auto` uses Anthropic's classifier-guarded mode; `normal` prompts; or give an explicit allowed-tools list. |
| Allowed Tools | The list used by the explicit mode. |
| Ralph / Todo Tracker | Enables the Ralph loop surfaces. |
| Agent Teams | Experimental teams. Also needs the CLI's own environment flag. |
| Codeman Agent Skill | Injects the agent skill into new Claude sessions per case. Off by default. See [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent). |
| Remote auto-reconnect | Reattaches dropped remote SSH sessions. On by default. |
| Nice priority / value | Runs agent processes at a lower CPU priority. |
| Bypass approvals and sandbox | Pi's project trust. Read [Agent CLIs](Agent-CLIs) before enabling. |
| Animated status effects | Cosmetic. |
### Notifications
Master toggle, browser notifications, push subscription, audio alerts, and the idle
threshold that decides when a quiet session counts as needing you. See
[Notifications And Approvals](Notifications-And-Approvals).
### Voice
Active provider and the engine behind it, insert mode, language, domain keywords to bias
recognition, the Deepgram API key, and the opt-in switch for transcribing through this
server's Claude login, with its live credential status. See
[Input And Voice](Input-And-Voice).
### Shortcuts
Rebinding for the shortcut registry. See [Keyboard Shortcuts](Keyboard-Shortcuts).
### System
`CLAUDE.md` template for new cases, default working directory, the image watcher, and
Cloudflare tunnel controls including the tunnel and upload URLs. In multi-user mode, the
| **Header tab strip** | The default. Wraps to a second row on desktop, scrolls sideways on a phone. |
| **Left sidebar** | A vertical list with a filter box and a live session count. `Alt+B` collapses it to a narrow rail that keeps the status dots and task badges visible. On a phone it is an off-canvas drawer rather than a docked rail. |
It is the same list either way, just re-hosted: tab order, drag-to-reorder, the `Alt+1`
to `Alt+9` numbers and every status colour below behave identically in both. The setting is
per device, so a sidebar on your desktop does not force one onto your phone.
## Session tabs
One tab per session, in your order, and that order syncs across your devices.
**Status is carried by the dot and the tab's own styling:**
| Connection dot | Always on | SSE connection health. Green is connected. |
| Font size `-` / `+` | Always on | `Ctrl +` / `Ctrl -` do the same. |
| CPU / MEM bars | On | Server resource use. |
| File Viewer | On | Toggles the file browser panel. |
| Settings gear | Always on | App Settings. |
| Plan usage chip | On, desktop only | Live Claude subscription usage. Claude-only, and needs its telemetry exporter, which the same setting installs. |
| Session Manager | Off | The full session list, live and historical. |
| Approvals bell | Off | Cross-session queue of prompts waiting on a human. Appears only when the count is above zero. Never shown on phones. |
| Read My Mind 🧠 | Off | Predicts your next prompt for this case. Claude-only. |
| Attachments | Off | Registered external files. |
| Away Digest | Off | What happened while you were gone. |
| Last Response | Off | Readable view of the agent's last answer, useful on phones. |
| Ultracode / Workflow | Off | Live workflow-run agents. |
| Notifications | Off | Notification history and settings. |
| Lifecycle Log | Off | Session start, exit, and kill audit trail. |
| Cron ⏰ | Off | Scheduled jobs. |
| Multi-monitor | Off, macOS | Opens a window spanning every display. |
| Tunnel indicator | When a tunnel runs | Cloudflare tunnel status. |
| Admin panel | Multi-user only | User administration. |
New header controls never appear on phones. Phone layout is deliberately minimal and is
covered in [Mobile Guide](Mobile-Guide).
## Connection state
The dot in the header is the quick read. Two louder surfaces exist because a cached page
with no server behind it used to look identical to a page with no sessions:
- **A full-screen overlay** when the page has never loaded server state. There is nothing
behind it worth preserving.
- **A banner** when the connection drops after state had loaded, so your scrollback stays
readable.
Both wait about 2.5 seconds before appearing, so a deploy that restarts the server does not
flash a warning at you every time. If the browser reports itself offline, the grace period
is skipped.
There is also a watchdog for the case where the connection stops delivering without
erroring. If the server's heartbeat stops arriving, Codeman reconnects on its own rather
than sitting on a green dot showing frozen data.
## The terminal
A real terminal: xterm.js in the browser, a real PTY on the server, tmux in between. Full
TUIs render correctly.
Worth knowing:
- **Scrollback.** Agent/TUI sessions pull their entire tmux scrollback on first open.
Shell sessions open from a bounded recent tail so a large transcript cannot stall tab
switching; press **Load full history** to pull the rest explicitly. Ordinary Shell scrolling
and automatic output recovery stay within the bounded browser buffer.
- **Wheel and touch scrolling** are forwarded into Claude's own transcript on recent Claude
versions, so the wheel scrolls the conversation rather than the terminal. `Shift+Wheel` is
always local scrollback. Other CLIs scroll locally.
- **Selection copy.**`Ctrl+C` copies when text is selected and interrupts when it is not.
`Ctrl+Shift+C` always copies.
- **Zero-lag input.** On touch devices, keystrokes paint locally before the round trip. See
[Input And Voice](Input-And-Voice).
- **Renderer.** WebGL by default, with a watchdog that falls back to DOM rendering if the
GPU stalls. `?nowebgl` forces DOM rendering for one page load.
## The home screen
With no session selected you get the welcome screen: run buttons for the CLIs Codeman
found, a QR code when a password is set, cross-session search, and **Resume Conversation**,
which lists past sessions including Claude conversations started outside Codeman entirely.
Two extras depending on the device:
- **Desktop, wide windows**: your open tabs appear as a rail docked to the left edge, in tab
order, with created and last-active stamps. It needs at least 1180px of width; below that
it is hidden so it cannot overlap the search panel.
- **Phones**: tapping the "C" logo gives a session overview instead: NEEDS YOU first, then
# Warm worker pool: sub-second claude worker spawns
Design sketch. Status: **proposed**, not started. Opt-in (`workerPoolSize`, default 0 = off); a user who touches nothing sees no change at all.
---
## 1. Problem and numbers
Measured against prod 1.18.3 on 2026-08-15, AFTER the SKILL.md fast-path hardening
(no recon turns), on the identical "spawn two codeman workers" prompt:
- **Cold orchestrator** (fresh session, skill loaded from disk): **20.2 s** prompt to
final report. Breakdown: 3.9 s Skill-load turn, 6.4 s generating the one fused Bash
call, **4.4 s spawn call**, 5.5 s summary. Tabs appeared at 10.5 s.
- **Warm orchestrator** (skill already in context, no Skill turn): **12.8 s**, spawn
call 6.0 s.
- Inside the spawn call, session + tmux + case creation is cheap: the workers (and
their tabs) appeared 0.2-1.7 s in, both siblings within ~350 ms of each other. The
remaining **~4-5 s is claude CLI boot plus the composer-readiness wait**, paid again
on every cold spawn. That slice is the pool's entire target.
The honest framing after the hardening: model turns dominate the skill flow (~16 of
20 cold seconds) and no server feature can shrink those. The pool attacks the
tool-side floor, and it has two distinct beneficiaries:
- **Skill/API orchestration**: the spawn call drops from ~4.4-6 s to ~1 s. Cold runs
land ~16-17 s, warm ~8 s. Tab appearance barely moves for this consumer (it is
model-turn-bound at ~10 s cold / ~4 s warm).
- **The UI Run button and direct quick-start callers**: a click today waits the full
boot + readiness before the worker can take a prompt; a pooled claim makes the tab
appear and the worker READY sub-second. This is the most visible win, and it
involves no skill at all.
Target: hand out an already-ready worker in **under 1 s**.
## 2. Shape
A new `src/worker-pool.ts` singleton service, following the `CronService` pattern: it **reuses the existing session layer** (`SessionManager` create + the normal spawn path) and never rebuilds tmux logic.
A pool member is a real claude `Session`, pre-spawned in a reserved scratch case (`~/codeman-cases/.pool-<n>`, created with the standard scaffold + hooks), already past readiness: composer drawn, hooks installed, preamble file seeded. It sits idle at the composer costing no tokens.
The claim happens **transparently inside `POST /api/quick-start`**: when a request is pool-eligible (§3) and a healthy member is available, quick-start returns that member instead of cold-spawning. The agent skill, the UI Run button, and every existing caller change **nothing**. Ineligible or pool-empty requests cold-spawn exactly as today, so the pool is only ever a fast path, never a behavior change.
## 3. Eligibility gate
Claim only when ALL of these hold; otherwise fall through to a cold spawn:
- `mode === 'claude'` (external CLIs have different readiness semantics and inject secrets via `tmux setenv` at spawn; out of scope).
- No `envOverrides`, no `CLAUDE_CONFIG_DIR`, and `modelOverride`/`effort` unset or equal to what the pool member was spawned with. Env vars flow at spawn time and cannot be applied to a running CLI.
- The requested case is **fresh** (does not exist yet). A linked case, an existing directory, a remote-SSH case, or a Docker case means the caller wants a specific workspace; pool members cannot provide one.
- Single-user mode, or the requester owns the pool (v1 ships single-user only; §11).
## 4. What a claim does (~300 ms)
1. Pop a ready member (in-memory check-and-remove; Node's single thread makes this atomic, so two concurrent quick-starts cannot claim the same member).
2. Health-probe it: `isPaneDead` (the existing ~750 ms-cached mux probe) plus one `capturePaneText` asserting a clean composer. A dead, limit-paused, or dirty member is recycled, and the claim tries the next member or falls through to cold spawn.
3. Rename the session to the normal `w<n>-<case>` name, set `parentSessionId` via the existing `resolveParentSessionId()`, clear the pool flag, persist state.
4. Emit `session_created`**now** (it was suppressed at warm-spawn time, §5). The tab appears here, sub-second after the request.
5. Return the **pool case** as `casePath`/`workingDir` and do NOT create a directory under the requested name: an empty dir the worker's CLI does not run in is a trap (files written there are invisible to the worker at cwd), and the agent skill greps the RETURNED `casePath` for Codeman hooks before trusting the worker, so the response must point at the directory that really carries them.
6. Kick a background refill (§6).
**The identity wrinkle, stated honestly:** the session id, `CODEMAN_SESSION_ID` inside the pane, the seeded preamble file, and the CLI's cwd are all fixed at warm-spawn and survive the claim unchanged. So a claimed worker's `workingDir` is the pool dir, not `~/codeman-cases/<requested-name>`; the requested name is a **label**. The API must report the truthful `workingDir`. Transcript projHash, response viewer, subagent windows, and Read My Mind all key off the real path and keep working precisely because we do not lie about it. This is acceptable for the dominant use (ephemeral skill workers that are deleted after answering) and is documented in the skill; a caller that needs the real case as cwd is by definition not pool-eligible.
**Verified skill compatibility (zero preamble changes).** Checked against the shipped 1.18.3 preamble: `spawn_worker`'s readiness probe (`_composer_up`) is a `wait-output` call with `from=buffer`, which scans output that already scrolled past before blocking, so a pooled member's long-since-drawn composer matches instantly instead of stranding a fresh-stream wait. The trust-dialog fallback never fires (members passed the dialog at warm time), and the hooks grep passes because the pool case carries the standard scaffold. Pooled and cold spawns are indistinguishable to the skill except in speed and the additive `pooled: true`.
## 5. Hiding pre-claim members
Pool members must be invisible until claimed or they read as ghost tabs. `Session.isPoolWorker` gates, at minimum:
- `GET /api/sessions` and `GET /api/sessions/unified` (and therefore the Cmd+K palette and the session-history-index snapshot that feeds `/api/search`).
- `session_created` SSE at warm-spawn (deferred to claim time). All other per-session SSE for a hidden member is suppressed at the broadcast call sites it would reach.
- Push notifications and the Approvals Inbox (a warm member showing a trust dialog must recycle, not notify).
- The phone overview / home rail (both render from the session list, so the list filter covers them).
- The lifecycle log records `pool_warm` / `pool_claim` events rather than user-visible session history.
`maxSessions` (50) **counts** pool members, and the pool refuses to warm within `poolSize + 2` of the cap so it can never starve real session creation.
## 6. Refill, TTL, drain
- **Refill** after each claim, debounced, at most one warm spawn in flight (a claim burst falls back to cold spawns rather than forking N CLIs at once; same reasoning as the document-conversion limiter).
- **TTL ~30 min**: recycle members older than that so they cannot drift from settings, hooks config, or a self-updated CLI on disk.
- **Drain and respawn** on: `claudeModel` change, hooks-config regeneration, self-update, and `workerPoolSize` changes. On server shutdown, kill pool sessions (they are stateless and ours). On boot, kill any leftover `.pool-*` tmux sessions found via `mux-sessions.json` rather than adopting them; adoption buys nothing for stateless members.
## 7. Failure modes
| Failure | Handling |
| --- | --- |
| Member died idle (PTY exit, crash) | Health probe at claim catches it; recycle + try next; PTY-exit breaker applies unchanged |
| Member hit a usage limit while idle | `isLimitPaused` members are never handed out; recycle |
| Claim race | Impossible by construction (synchronous in-memory pop) |
| Warm spawn itself fails | Log, back off, retry on next refill tick; pool empty just means cold spawns |
## 8. Cost
Each warm member is one tmux session + one idle claude process (order 150-300 MB RSS; **measure before defaulting the size above 0**, including whether an idle CLI makes any background requests via its statusline refresh). Zero token cost while idle. Suggested starting size for users who opt in: 2.
## 9. Settings and API surface
- `workerPoolSize` (int, 0-4, default 0): **synced** setting in `SettingsUpdateSchema`. The watcher that resizes the pool on `PUT /api/settings` must resolve from `merged`, never the raw body (the partial-PUT gotcha in CLAUDE.md).
- One internal status endpoint, `GET /api/worker-pool` (size, members' ages, claims served, fall-through count), for debugging. No new SSE events: the claim emits the existing `session_created`.
- No new public API semantics: `/api/quick-start`'s contract is unchanged apart from a `pooled: true` field in the response data, which is additive.
## 10. Considered and rejected
- **Renaming the pool case dir to the requested name at claim.** Linux keeps the process cwd working across the rename (inode-based), but claude computed its transcript projHash from the old path string at boot, so transcripts, subagent windows, and the response viewer go blind, the exact failure mode the `CLAUDE_CONFIG_DIR` docs warn about. Truthful label semantics (§4) beat a clever rename.
- **A new explicit claim endpoint.** Transparency inside quick-start means the skill, the UI, and every existing script get the speedup with zero changes; a new endpoint means new docs, new drift, and callers that must know the pool exists.
- **Pooling external CLI modes.** Readiness there is output stabilization, secrets ride `tmux setenv` at spawn, and codex/pi composer semantics differ per CLI. Claude-only until someone measures a need.
- **Returning quick-start at creation instead of readiness (no pool).** Would move tabs earlier on cold spawns too, but `sendwait` immediately after would then race the composer; readiness is what makes immediate tasking safe, and the pool makes the whole question moot for eligible spawns.
## 11. Phasing
1. **v1**: single-user, claude-only, fixed-size pool, transparent claim, status endpoint. Everything above.
2. **v2**: per-owner pools for multi-user mode (pool members must carry an owner because ownership scoping is structural); possibly model-matched pools (one warm set per configured `claudeModel`).
3. **Explicitly out**: warming linked/repo cases (spawning where the work is has no hooks and is the skill's documented costliest mistake; a warm pool must not make it faster to reach).
## 12. Testing
- Unit: pool manager logic pure and mock-driven (eligibility gate, TTL, refill debounce, drain triggers), `MockSession` from `test/mocks/`.
- Route: `app.inject` on quick-start asserting claim vs cold-spawn per eligibility row in §3, plus the double-claim race (two concurrent injects, one pool member: exactly one `pooled: true`).
- Live: re-run the pinned baselines against a warmed beta instance. Before (2026-08-15, prod 1.18.3, post-hardening): cold orchestrator **20.2 s** / warm **12.8 s** end to end, spawn call 4.4-6.0 s. Acceptance: spawn call under 1 s, cold ~16-17 s, warm ~8-9 s, and a UI Run click to a READY worker in under 1 s.
warn "curl is not available, so $curl_only_skipped install command(s) that need it were left out of the menu below (still shown as hints if you skip)."
fi
for path in "${CODEX_SEARCH_PATHS[@]}";do
if[[ -x "$path"]];then
echo"$path"
return
fi
done
}
if[[${#offer_idx[@]} -eq 0]];then
warn "No AI CLI can be installed automatically here. Codeman will run, but sessions need a CLI to drive."
cli_catalog_print_install_hints
else
echo -e "${BOLD}Which AI CLI would you like to install?${NC}"
localn=0 idx
for idx in "${offer_idx[@]}";do
n=$((n +1))
echo -e "${CYAN}${n})${NC}${CLI_LABELS[$idx]}"
done
echo -e "${CYAN}s)${NC} Skip (I'll install one myself)"
echo""
check_gemini(){
ifcommand -v gemini &>/dev/null;then
return0
localcli_choice=""
if[["$NONINTERACTIVE"=="1"]]|| ! has_tty;then
# Explicit automation opt-in: default to the first offered entry,
# which is registry order, which is Claude Code (order 0) — the
# same default this prompt has always taken non-interactively.
cli_choice="1"
info "CODEMAN_NONINTERACTIVE=1: defaulting to ${CLI_LABELS[${offer_idx[0]}]}"
else
while true;do
echo -en "${CYAN}Choose [1-${n}, or s to skip]:${NC}" >&2
read_reply cli_choice ||{cli_choice="1"; break;}
case"$cli_choice" in
s|S)break;;
''|*[!0-9]*)echo"Please enter a number between 1 and ${n}, or s." >&2;;
- Mobile catches up: links open from a tap, terminal text can be selected and copied, long prompts stay visible while you type. Plus Files panel search, a bundled Nerd Font symbols fallback, and a per-device terminal font setting.
- **Terminal and chat links work on phones** (#321): tapping a URL or file path in terminal output now opens it (new tab, file preview, or log viewer), resolved through the same provider desktop hover uses, so tap and click can never disagree about what is a link. Dialog rows and the composer keep their existing meaning. Response-viewer links open in a new tab with `rel="noopener noreferrer"` instead of navigating the dashboard away. Wrapped links open whole: the logical-line reconstruction now stitches hard wraps through the indent their continuation carries, which also fixes desktop hover-click truncating wrapped URLs.
- **Terminal text can be copied on touch devices** (#321): long-press selects the token under the finger, drag or tap the other end to extend, and a small bar offers Copy, Line (the whole logical line, wraps included) and dismiss. Copy works on plain-HTTP installs too. Three guards keep the keyboard down and the selection alive through the browser's own long-press handling.
- **A long prompt stays visible on phones** (#321): the local-echo overlay grows upward once it would run past the last visible row (a prompt taller than the screen keeps its tail, where the cursor is), and the keyboard-driven padding shrink can no longer reclaim the space the fixed toolbar and accessory bar stand in.
- **Files panel search** (#324): `GET /api/sessions/:id/files?q=...` answers a flat match list (name or path substring, `*`/`?` globs), recursing past non-matching directories with its own match cap on top of the existing bounds; without `q` the response is byte-identical to before. Glob queries are matched without regex so a pathological pattern cannot stall the server.
- **Nerd Font prompt glyphs out of the box, custom terminal font** (#320): a bundled icons-only Symbols Nerd Font Mono fallback renders powerlevel10k/starship/oh-my-posh glyphs on every device with no font install, and App Settings gains a per-device terminal font family that is prepended to the built-in stack.
### Thanks
Three contributor PRs in one release: thanks to @rounakdatta (#321), @aakhter (#324) and @comzine (#320).
"description":"Instant keystroke feedback overlay for xterm.js: Mosh-inspired local echo that removes perceived input latency over SSH, tunnels and other high-RTT connections",
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.