Follow-ups to the delegating statusline shim from discussion #405, answering
the four design questions and the Docker one raised there.
Docker cases: the injected command is now a self-selecting shell guard,
`if [ -x <node> ] && [ -f <shim> ]; then exec <node> <shim>; fi;` followed by
the inline curl exporter. A Docker case bind-mounts the workspace, and with it
settings.local.json, at the same absolute path inside the container, but
neither the host's node nor ~/.codeman exists there, so a bare shim command
would have rendered a broken statusline in every container session. The same
string now runs the shim on the host and the curl inside the container. The
settings-save injection loop needs no docker guard for that reason; it does
skip remote attaches now, whose workingDir is a user@host pseudo-path.
Settings precedence: the shim reads exactly the three files Claude Code
documents, .claude/settings.local.json and .claude/settings.json under
workspace.project_dir (the launch directory), then ~/.claude/settings.json.
No ancestor walk and no user-level settings.local.json: delegating to a
command Claude Code would have ignored is the original failure in a new coat.
The bare word: all three paths that produced `codeman` are gone. The route
answers an unknown session with an empty body, formatSessionStatusText()
returns '' with nothing to show, and the inline fallback ends in `|| true`
(plus `curl -f`, so an HTTP error body never renders as the statusline). The
shim also treats a literal `codeman` from an older server as no telemetry and
prints nothing rather than a brand word when it has neither a delegate nor a
footer, which is what Claude Code shows a user with no statusline of their own.
Removal: turning the plan-usage chip OFF now takes the exporter out of the
workspaces of the caller's live Claude sessions. statusLineTelemetry:false is
sent only by the save that flips the chip off on that device
(statusLineTelemetryAction() in settings-ui.js), so a phone whose chip was
never on cannot strip the exporter a desktop's chip depends on; a second
device with the chip still on re-injects on its next save or session create.
Nothing in src/ called applyStatusLineConfig(dir, false) before.
Tests run the generated shim AND the injected command as real subprocesses
(the fallback half with the shim path pointed at nothing, the container's
view), plus both directions of the action field through PUT /api/settings.
Measured against the live server: a render costs ~85 ms through the shim
versus ~26 ms for the old inline curl (node start ~33 ms, the rest TLS to
the loopback HTTPS server plus the delegate spawn), off the input path.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude Code ranks a repo's .claude/settings.local.json above
~/.claude/settings.json, so the statusLine Codeman injects for the Plan
Usage chip shadows whatever statusline the user configured globally. The
inline exporter then printed Codeman's own footer in its place, and
running `claude` by hand in a managed repo rendered the bare word
`codeman` — the response the server returns for an unknown session id.
The exporter becomes a generated shim, following the deepseek-status-shim
pattern: a versioned .mjs in the data dir, refreshed on a marker change,
written through temp-and-rename so a live render cannot read a
half-written file. It forwards the same payload to /api/status-telemetry
and, concurrently, resolves the statusline it shadows and prints that.
Codeman's footer still appears when there is nothing to shadow, so the
exporter keeps its value on a machine with no statusline of its own.
The delegate resolves at render time, walking the settings files Claude
Code consults from the render directory upward and then the home ones,
skipping Codeman's own entry in either form. Late resolution means
editing a global statusline needs no reinjection.
Ownership now keys on the version-free `codeman-statusline-shim` token,
and applyStatusLineConfig still reads the old /api/status-telemetry
command as ours, so managed repos upgrade in place instead of being
mistaken for hand-authored. A hand-authored statusLine is left alone
exactly as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 641795821fdbb171f9f47c4909c6e8945465d05b)
The fixed 120px/96px caps on the two wrapped layouts were row counts in disguise: a
third row was clipped into a ~4px scroller, hiding tabs inside a container nothing
invites you to scroll, while the header had the page below it to grow into. Both
layouts now share one rule capped at var(--tab-strip-max-height, 40vh), a safety net
for an absurd session count rather than a row limit.
Verified before shipping: .header is min-height + flex-shrink: 0 so it can grow, and
terminal-ui's ResizeObserver refits the terminal when it does; updateTabOverflowMode()
returns early for any non-desktop viewport, and below 1024px mobile.css pins the header
to max-height: 48px, so this is desktop-only in effect; the selector is comma-grouped
rather than :is(), so each arm keeps (0,2,0) and mobile.css's overrides still win on
source order. PostCSS parses the file cleanly (prettier ignores styles.css).
Authored in a parallel session against this shared checkout.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
House style, and these land in the changelog. Only the sentences added in the
previous commit are touched; the em-dashes in contributor text and in the
pre-existing COD-54/COD-115 comments are left alone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#409 (Claude truecolor). The changeset becomes the changelog, and its premise
does not hold on tmux 3.2 or newer. Measured here on tmux 3.4: `default-terminal`
sits at its compiled default of `tmux-256color`, a live claude pane reports
`TERM=tmux-256color`, and supports-color reads that as 256 colors, where
rgb(55,55,55) lands on ESC[48;5;237m — visible, just not the color the theme
named. The invisible block the PR describes needs TERM to resolve to a 16-color
entry: tmux older than 3.2, or a ~/.tmux.conf setting `default-terminal screen`,
which Codeman's own tmux server does read (it passes no -f). Both the changeset
and the invariants paragraph now say that, so the next report here gets paired
with the reporter's tmux -V instead of being read as universal. The change itself
stands on the simpler argument: claude was one of two entries not asking for
truecolor while twelve do.
Also reorders buildClaudeEnv(). It applied the registry's unset/exports AFTER the
whole env was built, so a clis.json entry naming CODEMAN_HOOK_SECRET_FILE or PATH
would strip it on the direct-PTY path while the tmux pane kept it — buildEnvExports()
emits `...cliEnv` ahead of `export CODEMAN_MUX=1` and cannot. The block now runs
first and Codeman's own keys are assigned on top, matching the pane.
#404 (Ctrl+Z trap). Adds the missing changeset, and records what the trap does
not cover: an agent CLI already holds its tty with ISIG off (verified on three
live panes: `susp = ^Z -isig -icanon`), so this is defence for the startup window
rather than a fix for the steady state, and two input paths still reach the PTY
unfiltered — the mobile accessory bar's one-shot Ctrl and the CJK textarea.
#399 (path picker sort). The server sorts by name and cuts at 500, so the client
sorting those 500 by date gives "the newest of the first 500 by name", which is
wrong in exactly the >500-entry folder the date sort exists for. The status line
now says "(first 500 by name)" so the cut is legible, with the reasoning parked
on _sortEntries.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
buildClaudeEnv(), the direct-PTY fallback taken when mux creation fails, now
reads getCli('claude').env and applies its unset and exports lists. It used to
delete COLORTERM and CLAUDECODE from a hand-maintained list of its own, which
left it contradicting the registry entry that the tmux pane and the attach
client both read. An engine value needing a mux name has nothing to resolve
against on this path, so it is skipped rather than guessed.
Claude no longer unsets NO_COLOR. The invisible-background bug does not need
it, and unsetting it overrides a preference the user set deliberately, so a
user who exports NO_COLOR globally keeps monochrome panes. The other seven
truecolor CLIs still unset it; that inconsistency is intentional and the
comment on the entry says so.
The invariants doc gains a Terminal colour env paragraph under Session launch
modes, where a reader looking up Claude will find it — the previous sentence
sat under a heading that lists only the non-Claude CLIs. It now says the lists
are the stock catalog and a clis.json override replaces them wholesale, and
that the declarations reach the tmux pane, its attach client and the direct
PTY but not a remote pane, whose command carries no env exports at all. Docker
hands COLORTERM=truecolor to every mode, including the two the registry says
must unset it.
The changeset named six peer CLIs and there are seven: deepseek also exports
truecolor. A test beside the existing OpenCode assertion pins the new
behaviour, so a future registry edit cannot make the backgrounds vanish again
in silence.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude draws the user's own messages as a block of background color, and
inside a Codeman pane that block was invisible. tmux hands each pane
TERM=screen, which supports-color reads as 16 colors, and Claude's registry
entry deleted COLORTERM on top of that. Claude therefore quantized every RGB
color its theme asked for down to the basic palette, where rgb(55, 55, 55)
and every other dark background becomes ESC[40m, the terminal's own black.
Changing the color in a custom Claude theme moved nothing on screen.
Claude now exports COLORTERM=truecolor and unsets NO_COLOR, matching codex,
gemini, antigravity, pi, grok and omp. CLAUDECODE stays unset, because Claude
reads it as a signal that it is running nested inside itself. Both the tmux
session and the attach client read this one registry entry, so they cannot
disagree.
PR #3 introduced the unset in February, citing xterm.js#484 for the claim
that xterm.js mishandles truecolor. xterm.js closed that issue in April 2019,
Codeman now depends on @xterm/xterm 6, and TmuxManager already sets
terminal-overrides ",*:Tc" on its own tmux server, so 24-bit color reaches
the browser today for every CLI that asks for it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A Fable 5.1 reviewer read the whole release diff against 1.26.2 and returned
SHIP WITH FIXES. These are its findings, verified before acting on each.
**The changelog advertised a feature the code refuses (major).** The #401
changeset and docs/web-tabs.md both listed `*.localhost` in the loopback set.
The follow-up in 02b0e278 moved it out of the auto-route set on security
grounds and updated CLAUDE.md but neither of those, and that changeset becomes
the 1.27.0 CHANGELOG entry: a user would have read the release notes, tapped
`http://app.localhost:3000/` on a phone and got a connection error from a
documented feature. Both corrected, and the user guide now says why it is
excluded and that adding such a dashboard by hand still works.
**Dictation delivered its text twice (minor, #388).** `keydownSnapshot` started
`null`, so `keydownSnapshot ?? canonicalCount` at the input event read a counter
xterm had ALREADY bumped: on a fresh page load with no keydown yet, xterm's own
capture listener forwards the `insertText` itself (it is not gated behind a
keydown), then the snapshot equals the bumped count, `count > snapshot` is
false, and the controller emits the same text again. Reproduced directly
against the module: it emitted `hello` for input xterm had already delivered.
A `0` baseline restores that file's own invariant, that a missed recovery is
acceptable and a duplicated keystroke is not. Two regression tests, covering
both the xterm-already-delivered and genuinely-dropped halves.
**The sorted rail's arrow-key walk followed the DOM (minor).** `_tabKeydownHandler`
steps `querySelectorAll` order, which is `sessionOrder`, while a sorted rail
paints its rows with the flex `order` property, so ArrowDown from the top card
landed wherever that session happened to sit in the tab order. It now sorts its
node list by the COMPUTED order first: computed rather than inline, because web
tabs take their `order: 9999` from CSS and would otherwise read as 0 and lead
the walk. This is the one place that follows the paint; the Alt+N badge, the
drag model and the filter all still deliberately read the DOM.
**A trusted dashboard was auto-reused by a tapped link (minor, #401).** The
reuse loop skipped `managed` and direct-mode records but not `trusted`. A
trusted frame is mounted with `allow-same-origin`, i.e. on Codeman's origin
with the user's cookie, and these links come from agent output, which is the
threat model the loopback allowlist was just narrowed for. An agent that can
write into the dev server's tree could print a path that one tap opens inside
that privileged frame. Excluded from auto-reuse, with a test; opening it from
the Run dropdown is still an explicit action and unchanged.
**Two documentation claims that were no longer true.** CLAUDE.md said
test/location-overlay-commands.test.ts pins every remote pane command, but
remote claude and remote omp now have their own arm in `buildRemoteLaunchCommand`
and never reach `defaultRemoteCommandForMode`, which is what that test asserts,
so it pins nothing for them and changing either arm will not fail it. Named the
real pins instead. Also documented the arrow-key-walk exception in the rail
paragraph.
Left as follow-ups, deliberately: `POST /api/webviews` does not dedupe by URL
server-side, so two devices tapping one link concurrently can still save two
dashboards for one origin (pre-existing endpoint behaviour that #401 makes
reachable by a tap), and the location-overlay golden should assert the real
remote claude/omp commands rather than a branch neither reaches.
Full gate green: 359 files, 6869 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An agent prints `http://localhost:5173/` (a dev server, a preview it just
served) and the user taps it on a phone. That address only exists on the
Codeman box, so the link was a guaranteed connection error from any other
device — while the web-tab proxy fetches from the server, where it works.
A loopback link (`localhost`, `*.localhost`, 127/8, 0.0.0.0, ::1) activated
in the terminal or clicked in the Response Viewer now opens as a proxied
web tab whenever the Codeman page itself is not on that box. A saved
proxied dashboard on the same origin is reused, with the link's own path,
query and fragment opened inside it (a mounted frame is navigated, not torn
down, so its state survives); otherwise one is saved under its host:port,
sandboxed like any other web tab, so it is in the Run dropdown next time.
Only loopback is routed this way. A LAN or tailnet address may well be
reachable from the device (a VPN, the same Wi-Fi) and a direct open is the
cheaper, richer path, so those keep opening in a new browser tab; on the
box itself every link opens directly. The terminal link provider and the
viewer's click handler consult one hook and fall through to their existing
behaviour when it declines.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
The eye button rendered `data.text`, which is one row: the last assistant
row of the transcript. A Claude answer is a median of 3 model messages
(p90 11) split around tool calls, so the brief view usually showed the tail
of an answer ("Done.", "Let me look.") and the substance appeared only after
More. The full view was fine, which is why the brief one read as broken by
comparison.
The brief view now asks `?context=turn`. The reader answers with the
assistant messages of the last ANSWERED turn (`selectLastAnsweredTurn`: the
highest `turn` that has an assistant row, so a prompt queued after the
answer does not blank the view) and the frontend renders them exactly as
the full view renders that turn: one badge, then continuation segments,
gated on the numeric `turn` as before.
`data.text` is unchanged in every context — still the last assistant row,
never `messages.at(-1)` — because agent pollers hash it. Readers that emit
no turns (Codex, the pane parser, DeepSeek, an older server) return `text`
only for `context=turn`, and the brief view keeps its single card for them.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
The picker's current-folder line was a read-only breadcrumb, so reaching a
deep folder meant tapping through every level, and the listing was fixed to
name order, so the file an agent had just written was somewhere in a
500-entry list.
The current folder is now an editable field: Enter or Go jumps there, a full
file path lands in its folder with that file selected, and a path that does
not resolve keeps the listing you had and says so, instead of the reset to
the root that a stale initialPath gets. A Sort control orders the listing by
name or modified time in either direction, folders always first, and the
choice is remembered per device like the hidden toggle. Each entry shows a
compact modified time (time of day today, month-day this year, else the
date).
GET /api/filesystem/browse stamps every entry with mtimeMs to make that
possible; the stat that already fetched a file's size now serves both, so
it is still one stat per entry. Entries without an mtime (an older server,
the in-container listing) sort after dated ones and then by name.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
Which conversation a pane is on was re-derived by correlating
~/.claude/history.jsonl against Session.lastSubmitAt — and lastSubmitAt is
bumped only by input that flows through Codeman's own write path
(Session.write / writeViaMux). A user who attaches to the pane's tmux session
directly never set it, so resolveActiveClaudeSessionIdFromHistory() returned at
its first line for that pane's whole life and the response viewer stayed pinned
to the launch conversation, showing a pre-/clear transcript indefinitely.
A UserPromptSubmit hook reports the live conversation id from inside the CLI
process, delivered under the pane's own $CODEMAN_SESSION_ID. That binding is a
fact rather than a correlation: it never consults workingDir, so it cannot be
claimed by a sibling pane on the same folder, a closed tab, or a bare `claude`
in the user's terminal. A pane holding such an id skips the correlation
entirely, so the number of prompts eligible for cwd-based guessing goes DOWN,
never up — the naive alternative (relax the guard, or synthesize an anchor from
PTY activity) is the reverted bug the resolver's own comment describes.
The hook also stamps lastSubmitAt, so it finally means "a prompt was submitted"
rather than "typed into Codeman's web terminal". Conversations vouched for
first-hand — and only those — extend a persisted claudeSessionChain, whose tail
re-pins the conversation when a surviving tmux session is re-attached after a
restart. ⚠️ start() resets the id at THREE points and the last one runs
unconditionally after the mux branch, so the tail is applied there too; patching
only the mux branch looks right and silently does nothing.
⚠️ The hook's stdout is discarded with curl's own -o /dev/null. Claude Code
injects a UserPromptSubmit hook's stdout into the model's context ("Exit code 0
- stdout shown to Claude"), and a trailing >/dev/null does NOT work: curlCmd
already ends `... 2>/dev/null || true`, and in `pipeline || true >/dev/null` the
shell binds the redirection to `true`, which never runs on the success path. The
discard is opt-in so the five SSE-fed events keep byte-identical command text
and no workspace's settings file is rewritten for them. The staleness marker is
quote-free for the matching reason: hooksJson is JSON.stringify'd, so a quoted
needle never matches and the gate would rewrite every workspace on every spawn.
Existing workspaces heal on their next Claude spawn through the staleness sweep.
The Claude reader concatenated every assistant row between two human prompts
into one card, fusing up to 74 distinct model messages into a single card, and
it never read the attachment rows that hold a prompt typed while the agent was
working. Measured over 57 real transcripts on 2026-09-01, the viewer shows
1,806 messages instead of 356 and 353 user cards instead of 178, with the
assistant text sequence unchanged row for row and the response without
?context=full byte-identical on all 57 files.
One assistant row IS one whole model message: in that corpus no assistant row
carries more than one content block and no message id carries more than one
text block, so there was nothing to reassemble. Each row becomes its own
message carrying an additive {kind, label, turn}, and the frontend renders a
same-role run inside one turn as badge-less continuation segments — which is
what keeps a p90 of 11 messages per turn from reading as card spam. A numeric
turn gates that rendering, so Codex, the external-CLI pane parser and an older
server keep one badge per card.
A prompt typed while Claude is working is recorded ONLY as an
attachment/queued_command row. Taking it when origin.kind is 'human' and
commandMode is 'prompt' recovers 162 user cards from 163 such rows — one is a
verbatim repeat inside an unanswered user run and is collapsed by the existing
dedup guard — and restores the turn boundary whose absence let the assistant
runs fuse. The CLI's own queue entries are cleanly separable: of 322
queued_command rows, 159 are commandMode 'task-notification' and not one of
them carries an origin key.
This narrows #169 rather than reverting it: sidechain exclusion, the
restored-<uuid8> rebind, replayed-snapshot dedup and synthetic-row filtering
are all unchanged and still asserted.
mobile.css hides #createCaseModal's .set-foot below 860px, and that modal's
header — unlike Settings' — carries no set-head-save. So on a phone the
Create/Link button existed nowhere and the modal could not be submitted at all.
Adds the header button and drives both together through switchCaseModalTab()
and submitCaseModal(), so whichever one is pressed the other shows the same
pending state and is equally unclickable. Following the Settings pattern also
means Add Case picks up the existing .set-head-actions:has(.set-head-save) tray
and .set-head-save sizing with no new CSS; the mobile.css comment that still
listed Add Case as a lone-× sheet is corrected to match.
Small cleanup items from upstream review (Ark0N/Codeman#353):
- OMP_SEARCH_DIRS now leads with ~/.local/bin, matching omp.sh's real
installer target (~/.omp/bin was an earlier unverified guess, confirmed
wrong against a real --no-cache Docker build).
- docs/omp-integration.md: fixed the dead GitHub URL (can1357/omp ->
can1357/oh-my-pi), corrected the CLI count (ninth backend, tenth
SessionMode incl. shell -- not eighth), matched the install-path guidance
to the resolver fix, updated the version example to the actually-tested
18.0.8, and added a Docker-section caveat: --resume pinning does not
currently reach an in-container omp process, since Docker panes never see
ompConfig.
- docs/architecture-invariants.md: fixed a heading missing ", OMP" (CLAUDE.md
already linked to the -omp anchor, so the link was dead) and added an OMP
specifics paragraph -- the one external CLI missing an entry in this doc.
- .changeset/omp-backend.md: corrected the sibling-CLI list (was missing Pi,
Grok, and DeepSeek Harness) and the backend count.
- Removed a stray orphaned comment fragment in the quick-start docker branch
and split two CSS lines that had two declarations jammed onto one line.
install.sh installs a build toolchain on Linux (node-pty has no Linux
prebuild, so a stock Ubuntu 24 server died inside node-gyp with
"not found: make"), plus review hardening for #339: the write-queue
reset paths now release the one-chunk-in-flight gate.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolves the four advisories that reach the production dependency tree. The
other 16 npm audit reports are devDependencies-only (Remotion, Puppeteer,
postcss, the eslint/tsx toolchain) and never ship to users.
- @fastify/static 9.1.3 -> 10.1.3 GHSA-8pvw-jcv7-9cmj (authz bypass via
non-canonical URL paths). Covers <=10.1.1, so all of 9.x is affected and
the fix exists only on the 10.x line.
- find-my-way 9.6.0 -> 9.8.0 GHSA-c96f-x56v-gq3h (HTTP/2 DDoS)
- fast-uri 3.1.2 -> 3.1.5 GHSA-v2hh-gcrm-f6hx (host confusion)
- brace-expansion -> 5.0.9/1.1.18 GHSA-3jxr-9vmj-r5cp (expansion DoS)
The last three are transitive and needed only a lockfile re-resolve, so no
overrides were introduced.
The @fastify/static major changes setHeaders' first argument from a Node
ServerResponse to a FastifyReply. Two consequences:
1. res.setHeader() -> reply.header(). The v9 body throws TypeError from
inside the plugin on every static request.
2. Precedence flips, silently. The callback used to write to the raw
response and lose to the route's staged reply headers; it now writes to
the reply and wins. That gave /sw.js a year of immutable in place of the
no-cache, no-store its route sets, pinning a service worker on every
client with no server-side recovery. A route that already set
Cache-Control now keeps it.
Verified against v9 to confirm the sw.js behaviour is a regression and not
a pre-existing bug.
ws appears in npm audit but production is on 8.21.0, outside the vulnerable
range; the only affected copy is bundled under @remotion/renderer (dev-only,
and remotion is pinned at 4.0.473 because the compositor refuses to start on
a version mismatch).
Adds test/static-cache-headers.test.ts, which drives a real server and covers
a caching contract that had no test at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`npm test` ran config/vitest.config.ts, which includes the browser, visual and
perf suites. On any machine without chromium, a free port and per-machine PNG
baselines that fails ~87 tests on a clean master, so the repo's most obvious
command could not be used as a pass/fail signal. The workaround had spread into
four docs as "never run bare `npm test`" warnings.
`npm test` now runs config/vitest.ci.config.ts — byte-for-byte what CI runs — so
local green means CI green. Verified: 264 files, 5248 tests, exit 0.
The suites it leaves out are not abandoned; each has a command:
test:browser 5 Playwright files (chromium + a live server; codex-predictive-echo
also needs a real codex binary)
test:mobile unchanged — the above plus per-machine PNG baselines
test:perf 2 wall-clock benchmarks; need an otherwise idle machine
test:all the old everything-behaviour, kept reachable
test:ci is untouched (CI still calls it). test:watch and test:coverage follow
test onto the gate's config.
The more important half is the hole this closes. The exclusion list lived as
literals in one config and pointed one way only: a file excluded from CI and
added to no runner would be tested by NOTHING, silently, with every command
still green — vitest counts "no files matched a filter" as success. That is the
same shape as the #279/#280 blind spot already documented in CLAUDE.md.
So the globs moved to config/test-suites.ts, one array per REASON a suite cannot
run in CI, and all three configs derive from it. test/test-suite-partition.test.ts
then checks the arithmetic against the files on disk: it fails if any test file
is reachable by no runner, or by two. Confirmed it fires by orphaning a file and
watching it name it. The partition is exact today:
gate 264 + browser 5 + perf 2 + mobile 9 = 280 = every *.test.ts in the repo
⚠️ One sharp edge, deliberate and documented: a file filter must match its
runner. `npm test -- test/mobile/keyboard.test.ts` now matches nothing and exits
GREEN having run zero tests, because the gate's config excludes that path.
CLAUDE.md recommended exactly that command in the on-screen-keyboard note; that
line now says `npm run test:mobile -- <file>`, and the Testing section calls out
the trap, since a green run of zero tests is worse than a red one.
Docs synced: CLAUDE.md, AGENTS.md, .github/CONTRIBUTING.md, README.md,
README.zh-CN.md, and two ci.yml comments that claimed only test/mobile/** was
excluded — it is three suites, and 5 Playwright files rather than 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Session List Layout gains a third option. The old "Left sidebar" becomes
"Left sidebar simple" and is unchanged down to the byte; the new "Left sidebar"
puts on each row what the desktop home rail and the phone overview already show:
when the session was first created, how long it has been in the state it is in,
and a status pill naming that state.
A docked column is not a tab strip. It has width to spare and a row per session
either way, and "name + folder" is the whole story a TAB can tell, not the whole
story there is. This is the information that was missing, and it already existed
one surface over.
Both sidebar values are the same layout, and both set data-session-list="sidebar";
the row detail rides on a separate data-sidebar-detail attribute. That split is
the load-bearing decision here: every one of the ~25 isSessionSidebarActive()
call sites and every html[data-session-list="sidebar"] rule in styles.css and
mobile.css keeps matching both variants without being touched. A third
data-session-list value would have meant auditing and editing all of them.
- Stored values: 'header', 'sidebar' (simple), 'sidebar-rich'. Anyone already on
'sidebar' keeps exactly the layout they picked — the rename is label-only.
- State classification and the "how long has it been like this" anchor come from
_mobileOverviewState() / _mobileOverviewSince(), not re-derived, so the three
surfaces cannot disagree about what "working" means. A working pane repaints
~1/s, so its duration is measured from the turn's last Enter: a running turn
reads "working 12m", not "0m".
- Stamps refresh in place on a 20s clock rather than by re-rendering — a rebuild
would restart every load spinner and alert animation in the list, twice a
minute. The clock runs only while rich rows are on screen, and is stopped from
both render paths and from applySessionListLayout().
- The incremental render path updates the pill, the accent class and the since
anchor; a tick alone cannot see a state change, and a new turn re-stamps
lastSubmitAt without changing state.
- applySessionListLayout() now re-renders on a DETAIL change too. simple <-> rich
leaves data-session-list on 'sidebar' both times, and the meta line is emitted
by the row template rather than toggled by CSS, so the old layout-only test
would have flipped the setting and repainted nothing.
- Width: 300px for the extra line. The collapsed 44px rail and the handheld
drawer are both explicitly held back from it — the desktop rule is (0,3,1) and
would otherwise out-specify mobile.css's (0,2,1) drawer base and pin a 320px
phone's drawer to 300px.
- Missing/stale mobile-overview.js degrades to a row with no meta line rather
than throwing and taking the whole tab strip down.
15 new tests cover the attribute split, the solo-window override, the
detail-change re-render, the row model, both render paths, the clock lifecycle
and the mobile width guard.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Brings in https://github.com/christianhaberl/Codeman/pull/4 (three commits,
authorship preserved) and adapts it across the 211 commits master gained
since the branch was cut:
- App Settings control re-authored for the set-* surface (PR #278): a
set-row in Layout -> Tabs, replacing the old settings-item markup the
branch targeted. i18n description synced.
- Lineage arcs (PR #291, post-branch) are SKIPPED in sidebar layout:
computeLineagePath()'s U-bridge geometry hangs from the horizontal
strip's bottom edge and has no meaning against a vertical list. The
lineage strip-scroll listener now also redraws subagent/ultracode
connectors while the sidebar scrolls vertically.
- The desktop home tab rail (post-branch) defers to the sidebar: both dock
the session list flush left, and the rail would render z-ordered under it.
- Active-row reveal unified into _scrollActiveTabIntoView() (#257 landed on
master after the branch): sidebar mode branches to scrollIntoView
block:'nearest', and _fullRenderSessionTabs() restores scrollTop alongside
the #257 scrollLeft restore so ambient rebuilds cannot yank a mid-scroll
sidebar back to the top.
- Mobile active-tab hoisting the branch guarded against no longer exists on
master (removed by #257); kept master's order-stable render.
Verified: typecheck, lint, format:check, check:frontend-syntax,
check:public-assets, PostCSS parse of both merged stylesheets, the 26 new
jsdom tests, the structural guard suites, and the headless-Chromium harness
(scripts/verify-session-sidebar.mts) green across all seven layout states
at 1600/1000/393px against current master.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The changeset ships this as 1.18.2, so the stamp, the bootstrap's grep/write
condition, both re-source guards and the recipes guard all carry 1.18.2 now
instead of a version that would never exist.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fifteen review findings on the fast-path rewrite plus one caught live, all
verified against a real 1.18.1 server before landing:
- sendwait picks a fresh seq (the epoch second) instead of a fixed 2, so a
second prompt to the same worker is typed instead of silently swallowed as
an already-applied duplicate; explicit seq remains for deliberate resends
- sendwait self-heals stranded delivery: an Ink repaint occasionally eats the
Enter (observed live), so a timed-out short first wait sends one bare \r and
re-waits by resending the identical frame as a tagged duplicate
- spawn_worker verifies the resolved casePath carries Codeman hooks (the same
/api/hook-event marker the server checks), refusing names that resolve to
linked or pre-existing hook-less directories instead of running the job in
what may be the user's real repo
- spawn_worker probes the trust dialog after a short 5s composer wait, not the
full 45s, restoring the ladder staging verbs.md documents; on a readiness
miss it deletes the half-spawned session and returns 1 with empty stdout,
so a prompt can never be typed blind into a trust dialog
- spawn_workers refuses duplicate case names and empty argument lists, and
keys result files by index
- section 1 is bash 3.2 compatible (indexed arrays, no declare -A), prints the
full delivered/timedOut/signal tuple per worker with an explicit line for a
missing result, deletes only workers whose turn really ended (a timeout
means still working), cleans up spawned siblings when any spawn fails, and
guards its mktemp
- last_text takes the previous answer as an optional second argument for
consecutive-turn reads (the transcript briefly serves the prior answer
after a stop, observed live)
- the stale duplicate bullets in section 1's closing list are gone
- reference/verbs.md joins the mode-list drift guard's file list
- README's skill inventory covers verbs.md and the new SKILL.md shape
- the changeset is minor so the shipped release matches the 1.19.0 stamp
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured against a live 1.18.1 server, the API does the whole job in about ten
seconds: two cold claude workers spawned and ready in 6.3s, both tasked and both
answers read in 4.0s more. The slowness users reported was agent-side.
Three causes, all of them things the skill taught:
- It taught serial spawning. Nothing in the main document showed `&`/`wait`, so
"spawn two workers" read as "do the readiness ladder twice", which is one model
turn per worker.
- It had no spawn primitive. The happy path had to be reassembled on every run from
where-to-spawn, a four-stage readiness ladder, send-and-wait, the fan-out caveats
and a recipe with two variants. Each is a decision, and most carry a warning.
- It cost ~16k tokens before the first call, at 3.6:1 prose to code, with 25 warning
glyphs and 55 occurrences of "never". A document that is mostly failure modes
teaches caution, and caution bills as thinking tokens.
The preamble now defines the verbs rather than describing them: spawn_worker,
spawn_workers (concurrent), sendwait, last_text. Section 1 composes them into the
whole job in one Bash call and says to stop reading there.
Two ceremonies the measurements retired: the pid poll (one iteration, 33ms, and
wait-output already blocks on the composer) and reading settings.local.json to check
hooks for a case quick-start creates, which always has them. That check stays
required for linked cases and raw paths, where its absence silently breaks
send-and-wait.
The bootstrap's write condition now greps the version stamp, so a stale or truncated
preamble self-heals rather than failing and asking for a manual rm. The stamp line is
kept bare because the grep anchors on it with $; an inline comment there would rewrite
the file on every bootstrap.
Section 5 moved to reference/verbs.md behind an index, cutting the always-paid
SKILL.md from ~16.4k to ~7.6k tokens. Section numbers and anchor slugs are unchanged,
so existing references still resolve; all 201 anchors across the five files were
checked, with the checker positive-controlled against an injected bad link.
Verified by extracting the code blocks from the shipped file and running them against
the live server: bootstrap plus full fast path, two workers resolving on the
definitive stop signal, answers read and sessions deleted, in 6.8s.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`GET /api/pi/status` shipped undocumented in the agent skill, and only a human
reading the doc noticed. Turns out none of its five siblings were documented
either, so this adds the whole family in one place: spawning with a mode whose
CLI is absent fails with OPERATION_FAILED rather than falling back, which is
exactly what an agent picking a backend it did not choose needs to know. Pi's
extra `.data.version` is called out, since a false `available:false` there means
an unrelated `pi` is in front on PATH.
On whether the endpoint scanner should also check registered-to-documented:
measured, and NO for the general case. The skill documents 34 of 217 registered
endpoints deliberately (it is an agent guide, not an API reference), so a blanket
reverse check needs a 183-entry allowlist that would fail CI on unrelated route
work and get appended to mechanically, which is worse than the gap it closes.
Grouping by path shape does not save it either: the families that yields are
things like `DELETE /api/<any>/:id`, lumping cases, webviews and docker hosts
together, and it would not have caught this gap anyway (the family had zero
documented members).
What IS cheap is a family the schema can enumerate with no allowlist: the new
assertion derives the agent modes from the Zod enum and requires each one's
`/api/<mode>/status` to be documented, so a seventh backend fails here until it
is. The sibling scanner still proves the other direction, that nothing documented
is a 404. Both mutation-checked: dropping pi's probe fails the new guard, and
documenting a nonexistent probe fails the old one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Second review pass on #282, the three items left open after f4dcfbe.
1. `codeman doctor` and the run mode disagreed about pi. The registry entry
accepted a bare `which pi` hit while pi-cli-resolver demanded semver-shaped
`--version` output, so the Dependencies panel could report an installed Pi CLI
on a box where Run Pi stays hidden, which reads as a broken mode rather than a
missing install. Both sides now share one exported PI_VERSION_REGEX, and
PathResolver gains an opt-in `requireVersionMatch` so a binary that fails the
shape check is reported MISSING instead of installed-with-unknown-version.
Only pi sets it; every other tool keeps its current behaviour.
2. The isAltScreenStripMode comment justified excluding pi with "the alt screen
is load-bearing for its fullscreen TUI". That is not what exclusion does: pi
is tmux-backed, so it falls through to isMuxAltScreenOnlyStripMode, which
strips the alt-screen toggles anyway. What exclusion actually preserves is
`\x1b[3J` and the mouse DECSETs, which is the real reason (pi renders into the
main screen and is mouse-aware). Comment and changeset now say that, and state
the consequence: fullscreen pi paints into the main buffer, like vim in a tmux
shell session.
3. skills/codeman still enumerated the five pre-pi modes in nine places, telling
agents a backend does not exist and understating class-wide caveats by one
mode. All updated, plus stale session.ts line references refreshed.
Tests: a new static guard derives the mode set from the Zod schema (not a copy)
and fails when a skill enumeration lists a partial set of external CLIs, verified
by mutation. It also documents the one legitimate exception it found: the "writes
no transcript" lists drop codex, which does write a rollout Codeman reads back.
Plus doctor cases for an unrelated `pi` on PATH and registry/resolver regex parity.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review follow-ups on #282. All four are the same failure shape: a list that
enumerates run modes, missed by the sweep that added 'pi'.
1. Cron ignored pi's project-trust clamp. The PR widened CronJobBaseSchema's
agentType to accept 'pi' but not the matching clamp beside gemini's, so a
non-granted multi-user owner's cron pi job spawned bare `pi` (pi's own
defaultProjectTrust, an interactive prompt they can answer "yes" to, which
loads and EXECUTES repo-local .pi/extensions TypeScript) while the same
user's UI/API launch was forced to --no-approve. The clamp is now a pure
exported helper, clampCronExternalCliConfigs(), so both it and gemini's
previously untested materialization are pinned.
2. POST /api/sessions/:id/interactive auto-enabled the Ralph tracker for pi:
its denylist covered opencode/codex/gemini/antigravity only. The tracker is
never fed for an external CLI (_processExpensiveParsers returns early), so a
pi session reported ralphEnabled and Ralph UI state no sibling backend shows.
3. REMOTE_CLI_BIN had no pi entry, so buildRemoteCliVersionProbeCommand()
returned null and Session.cliVersion stayed blank for every remote-SSH pi
session, even though the PR wired the remote launch command and the
per-mode override schema field.
4. The desktop home rail's badge map had no pi entry, and its lookup falls back
to '', which is what claude renders. A pi session read as Claude there while
the tab strip and phone overview badged it correctly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>