Follow-up to #421 (remote-case file reads over ssh), addressing the review.
Symlink escape on a host without `readlink -f` (blocker). The probe's
portable fallback canonicalized only the directory chain and returned the
final component unresolved, so on macOS < 12.3 `ws/notes.txt -> ~/.ssh/id_rsa`
came back as `.../ws/notes.txt` (with the target's size), passed every
containment and blocklist check that runs on `realPath`, and `cat` followed
the link. The fallback now walks the directory chain with `cd -P`/`pwd -P`
and follows the LAST component with plain `readlink` for a bounded number of
hops, and anything it cannot fully resolve (a loop, a readlink failure, the
hop cap) is reported with an `x` marker that parses as null, i.e. 404. It
never returns the unresolved string. Measured on a real /bin/sh with
`readlink -f` shadowed: the pre-fix script reports `/ws/notes.txt`, the fixed
one `/secret/id_rsa`; both branches (native and fallback) now agree.
`PUT /api/sessions/:id/file-content` never had the remote guard the PR
described. It sits ahead of `validateSessionFilePath`, which resolves against
the LOCAL filesystem, because with a same-named directory on the Codeman host
(an sshfs mount of the remote tree, the documented stop-gap) the write landed
on the local twin while the viewer believed it edited the remote file.
ssh fan-out is bounded. `src/remote-ssh-limiter.ts` is a
document-conversion-limiter-shaped semaphore (default 4, env
`CODEMAN_MAX_REMOTE_FILE_SSH`) around every probe and buffered read; the
attachment-history list resolves its whole history in ONE batched probe
(`probeRemoteAttachmentHistory`, threaded into
`registerExternalAttachment({remoteProbes})` so the guards run unchanged)
instead of one handshake per entry; and probes chunk at 40 paths because the
whole script is one argv string. Terminal output in a remote session is
written on the remote host, so a prompt-injected agent printing hundreds of
`codeman://attach` links forked one ssh per link, each holding a 20 s
timeout, and a 100-entry history re-listed on every attachment:detected
tripped OpenSSH's default MaxStartups. Streams are deliberately not counted
(one per browser request, held for a whole playback, and gated behind a
counted probe anyway).
Smaller items from the same review: probe records are NUL-terminated and
index-keyed after a leading NUL (a newline in a filename can no longer shift
the alignment, and the banner is fenced off without last-N-lines guessing);
size comes from `stat -c %s || stat -f %z`; the three IO functions refuse
under VITEST instead of opening a connection; an unreachable host now reads
as unknown (missing: false) for detected AND external history entries, where
external used to fold its 502 into missing; a client that aborted during the
guard probe has its body's ssh child reaped (`reply.raw.destroyed` is checked
before the close listener is attached); `describeExecError` never returns
Node's `Command failed: <ssh line>` message, which carried the identity path
and the probe script into a 502 body; and the docs note that
`isSensitivePath`'s three home-anchored entries resolve against the Codeman
host's home, not the remote one.
Tests: the probe script runs on a real /bin/sh with a `readlink` shim that
rejects `-f` (the escape, a relative chain through a symlinked directory, a
loop, a newline filename, banner chatter that itself looks like a record),
the limiter's cap and FIFO order, and route tests for the PUT guard (local
twin untouched, no connection), the single batched history probe, the
unreachable-host alignment and the aborted-client reap. All four route tests
fail against the pre-fix file-routes.ts.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The runtime shim masks `/webview/<cap>/` off a proxied page's URL so its router
boots on the path it expects, and the landing page masks to exactly `/`. A
`location.reload()` there (a Vite dev server on a config change or a failed HMR
update, the likeliest case in the feature's own motivating scenario) therefore
asks for Codeman's root as an iframe navigation. `serveLostWebviewFrame()`
returned early for `/`, so on a passwordless install the frame received Codeman's
own app shell and rendered it inside the web tab, and with a password it got a
401 in the frame. Either way no `codeman:webview-lost` message was posted, and
because the document loaded fine the load handler cleared the failed-frame panel,
so the Reload / Open in new tab affordances never appeared. Before masking the
frame's URL was the prefixed one, so a reload worked; this was a regression.
`/` is the one lost-frame path a registered route also serves, so the route
table cannot tell that reload from a real navigation. Credentials can: nothing
in Codeman frames its own root, and a sandboxed frame is opaque-origin with no
cookie and no Authorization header. `carriesAuthCredentials()` (pure, in
webview-proxy.ts) makes that test, and `/` is now admitted by the auth hook only
when it fails; a framed `/` that does carry credentials still gets the shell.
Without a password no auth hook runs at all, so the index route applies the
same test itself (`isLostWebviewRootFrame`) before rendering the shell, and the
three places that emitted the recovery page share `sendLostWebviewFramePage()`.
Tests: the password form in webview-auth-exemption (recovery page for a
credential-free framed `/`, shell with valid Basic auth, 401 with a stale cookie
or a top-level navigation), the passwordless form against a real WebServer in
webview-lost-root-frame (port 3198), and the credential predicate in
webview-proxy. All three fail without the fix. Verified against a live isolated
instance as well: a framed `GET /` with no credentials answers the 470-byte
recovery page, a top-level `GET /` and a framed one carrying a cookie answer the
shell.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The lost-frame handler in webview-tabs.js remounts a web-tab frame at the path
the frame reports it lost. It promised "path only, never an origin" and collapsed
a leading run of slashes so `//host/x` could not jump the frame off the proxy,
but it left two spellings through that the WHATWG URL parser treats the same way:
a backslash, which is read as `/` for http(s) schemes, and an ASCII tab or
newline, which the parser deletes before it looks at anything, so `/\host/x` and
`/<tab>/host/x` both resolve to `https://host/x`. That mattered only in
direct-mode tabs, where `POST /api/webviews/:id/open` returns no embedUrl and the
recovered path is resolved with `new URL(path, src)` straight into the frame's
src; a page in such a tab could remount its own frame on a foreign origin.
Not an escalation (the page can already navigate itself anywhere, and the remount
carries no Codeman-origin access), but the comment did not hold and the existing
test only covered the form that already worked. The handler now strips tab, CR
and LF, collapses any leading run of `/` or `\` to one `/`, and refuses whatever
still opens a second separator. The new test drives the reachable direct-mode
branch with all four spellings and fails without the fix.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The capability list is DERIVED from what the scripts do (chown => CHOWN +
DAC_OVERRIDE, a setpriv uid/gid drop => SETUID + SETGID, `init: true` next to a
uid drop => KILL) and compared to docker-compose.yaml's cap_add, the
entrypoint's own required_caps diagnosis, and the lists quoted in docker/README.md
and CLAUDE.md, so the drift that shipped the missing CAP_KILL fails here rather
than on someone's server. Also pinned: the CLI prefix is appended to PATH in
server.Dockerfile and entrypoint.sh pins its PATH before its first command;
Start-Codeman.sh derives PUID/PGID before creating the cases dir, builds before
`down`, writes the source marker only after a refresh, and never aborts on a
failed volume removal.
git_head_commit is run as the script defines it, extracted by its own
delimiters into a real bash, against temp repos made with real git: a symbolic
ref with a loose ref file, a detached HEAD, packed refs after `git pack-refs`,
a linked worktree (which must resolve nothing rather than something wrong) and
a directory that is not a checkout.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every per-skin xterm palette declared its selection layer as `selection`,
the key xterm.js renamed to `selectionBackground` in v5. An ITheme is a
plain object handed straight to the terminal, so an unknown key is not an
error, it is dropped: all seven skins have been drawing xterm's built-in
default, rgba(255,255,255,0.3), rather than the colour sitting next to it
in the palette.
Nobody saw it on the dark skins, where white at 30% is close to what those
palettes asked for. On the four light skins it is white over a near-white
background: blended, Paper Gray's selection differs from its own background
by 3/255. That is not a subtle highlight, it is no highlight, and it looks
exactly like a selection gesture that failed, which is part of what #360
reports on Android Chrome.
test/skin-themes.test.ts pins both halves: the key name, and that the
blended selection stays at least 16/255 from the background on every skin,
plus the light-skin fallback landing under that floor, which is what makes
this a fix rather than a rename.
Refs #360
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two keys #408 adds to the mobile keyboard accessory bar send
Shift+Left and Shift+Right, which are Codex bindings (edit the last
queued message, step back through the prompt stack). They shipped on
both agent layouts, so a claude, pi, grok, omp, deepseek or gemini
session got two keys that do nothing. That was not only cosmetic: a tap
goes through sendNavKey(), which adds the session to
_echoPassthroughSessions and hands editing to plain PTY echo until Enter
or Ctrl+C, so on a phone a dead key also switched off the local echo
that makes typing feel instant there.
The reveal now follows the shape the 🧠 key already uses. The buttons
stay in both templates, carry an accessory-btn-codex marker class, and
are display:none in styles.css until the bar element carries
codex-enabled. The class has to live on the bar rather than on the keys
because setMode() rebuilds the buttons' innerHTML on every layout
switch. syncCodexKeys() toggles it from the active session's mode
(the same lookup _isShellSession() uses) and is called at init and from
refreshForActiveSession(), which selectSession() already invokes on
every switch. A session's mode is readonly on the server and fixed at
create, so no other event can change the answer; the welcome screen
(no active session) reads as not codex and hides the keys.
The frontend id-branching guard (test/cli-registry-no-id-branching.test.ts)
scans only src/**/*.ts, so the mode comparison in a public JS file is
in bounds, the same as the existing shell check beside it.
Tests: the new describe block in test/mobile-shell-keyboard.test.ts pins
the marker class in both templates, the CSS pair, the class for a codex
session in both layouts, its absence for claude/shell/pi/omp/deepseek,
the re-sync in both directions on a session switch, the no-session case,
and the init + refresh wiring. All six positive assertions fail without
the source change. README and the changeset now say the keys are
Codex-only.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A clicked path that points OUTSIDE the case directory goes through the attachment
routes (the frontend's `_isExternalPreviewPath` sends every absolute path not under
`workingDir` to `POST /attachments`), and those had the same local-`fs` assumption
as file-raw: `realpathSync`/`fs.stat` on a path that only exists on the remote host,
so the file never opened — the case the #415 report was actually about.
- `registerExternalAttachment()` accepts `remote` and resolves through
`remoteProbePaths` (canonical path, size/mtime, kind, plus the workspace root for
the confinement check). Everything around it — blocklist, extension allowlist,
workspace confinement, registry/dedupe — is now shared by both branches, so the
remote path cannot drift from the local one.
- The by-id routes (`raw`, `preview`, `thumbnail`), the metadata poll and the
attachment history list resolve over ssh too. `raw` streams with the same
Range contract as file-raw; `preview` (office) and `thumbnail` answer 400 for a
remote record; an unreachable host answers 502, a vanished file 404.
- Which host a record is read from follows the SESSION, never the path string: the
same absolute path is a different file on each host, and a remote session never
falls back to a local file with that name.
- Codex generated artifacts keep force-workspace confinement for a remote case: the
well-known artifact directories are anchored at THIS host's home, so only a file
inside the remote workspace is trusted.
Still local-only by design: writes, office conversion, thumbnails, the file
tree/picker and tail-file.
Rebased over #390, which moved the phone tier's cutoff from 430px to
600px. The palette's compound fold rule now lives in the 600-768px band
mobile.css pads, the cascade samples the palette inside that band, and
the closed iPhone Duo (466pt) is a phone rather than a small tablet while
the open one (626pt) stays a tablet. Comments in both stylesheets, the
device registry, CLAUDE.md and architecture-invariants say 600.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Ported from #416 (discussion #405): a statusline reading just `codeman`
is what a hand-run claude in a managed repo showed, and it reads as a
broken config rather than a footer. Three paths produced it and all
three now yield an empty footer: the exporter's `|| echo codeman`
fallback (now `curl -sfk ... || true`, with -f keeping an HTTP error
body off stdout), the unknown-session answer of POST /api/status-telemetry,
and formatSessionStatusText() with nothing to show. The exporter script
marker moves to V4 so live installs pick the new content up on the next
spawn.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Three small follow-ups from the #361 review.
A tmux setenv survives respawn-pane, so _configureStatusLineUserCommand
returning early when the user has no statusline left a previously
exported CODEMAN_USER_STATUSLINE_CMD in place: a user who deleted their
own statusline kept getting the stale one wrapped, and lost Codeman's
footer print-through, until the tmux session was recreated. It now
issues `setenv -u` in that case, the same shape as the effort-level
cleanup in applyEnvOverrides.
ensureStatusLineExporterScript truncated and rewrote a script that live
sessions execute on every statusline render, and chmod'd it after the
write. It now writes a temp file next to the target, chmods that, and
rename()s it into place.
The non-tmux direct-PTY fallback carries no exporter; that is now stated
at the spawn site and in the architecture-invariants paragraph rather
than left as a silent gap.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Two follow-ups to #361's sticky telemetry switch.
GET /api/settings reconciled an absent showPlanUsageLimits by persisting
true, but readJsonConfig() answers {} for ANY read failure (a parse
error, EACCES, EMFILE, a read landing inside PUT's non-atomic write), not
only ENOENT, and every page load calls this route, so one unlucky read
replaced the whole settings file with a one-key file. The route is a
plain read again and the default moved into the reader:
readPlanUsageTelemetryEnabled() treats an absent key as ON, the same way
readWorkspaceHooksEnabled() does, which is what the desktop chip already
shows for an install that never touched the setting.
saveAppSettings() sent showPlanUsageLimits on every save. The chip
defaults OFF on handhelds, so a phone saving its font size persisted
false and switched collection off for every desktop, whose chip then
went stale with no error anywhere. The key is now stripped like the
other per-device display keys and re-added only when the save FLIPS the
chip relative to what the device had (planUsageCollectionFlip), so an
explicit toggle on any device still writes it in either direction.
Tests pin both: the GET route with a mocked filesystem (absent, missing,
EACCES, garbage, explicit), the reader default, and the flip helper plus
its wiring in saveAppSettings.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
CLAUDE.md's folding-devices rule gains the two new invariants (a shape change
with the keyboard up baselines to window.innerHeight; a base gutter overridden
by a later @media block needs its own zero-base fold restatement, and a
compound rule written against a mobile.css shorthand is scoped to that band)
plus the architecture-invariants pointer it lacked; the new Folding devices
section there carries the mechanisms and the measurements. The device count
is 138 since the two Duo profiles landed (68 Playwright + 70 custom).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Three cascade problems in the fold reserved-region CSS, each measured by
computed style in headless Chromium (styles.css + mobile.css in index.html
link order):
- The unconditional .path-picker-overlay / .path-preview-overlay fold rules at
the end of the file beat the `padding: 0` both overlays set under 600px, so
every phone got a 16px and 18px gutter on dialogs built flush (393 and 500px:
edges floating off the screen). The fold strip is now restated on a ZERO
base inside the same media query: 0/0 without a fold, the strip alone with
one, 16/18 plus the strip from 626px up as before.
- .modal.command-palette-modal was unscoped, so outside the 430-768px band
(where mobile.css pads the palette with a shorthand) it ADDED 0.75rem with
no gutter to compose with and pushed the shell 6px off centre at 393, 900
and 1400px, while inside the band the shorthand beat the generic .modal rule
on the bottom side and the palette lost its block-end gutter. The compound
rule now lives inside that band and restates both sides.
- The tabletop cap on .response-viewer lost to mobile.css's `max-height:
92dvh` under 430px (same specificity, later file). mobile.css now carries an
identical twin at its end.
test/foldable-layout.test.ts simulates the padding cascade across both files
at every breakpoint, with and without the fold rules, and requires the two to
differ by exactly the fold strip; it also pins the palette rule to the band
mobile.css keys on and the response-viewer twin to the styles.css value. Its
model reproduces the Chromium numbers, and against the pre-fix stylesheets it
fails on all three problems.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A shape change with the keyboard up re-baselined initialViewportHeight to the
SHRUNK visual height, so heightDiff was 0 and the settle event the OS fires at
the new width (or any later address-bar drift) satisfied the hide branch and
ran onKeyboardHide() with the keyboard still on screen: accessory bar hidden,
toolbar lift dropped, main's padding cleared. It could not recover, since no
further 150px drop re-arms the show branch against a baseline already sitting
at the shrunk height.
Baseline to window.innerHeight instead when the keyboard is up: the page sets
no interactive-widget, so the keyboard shrinks only the visual viewport and
the layout viewport stays the display's full height on both engines, the same
fact updateLayoutForKeyboard() relies on.
The vm harness now models the two heights separately (resizeTo takes an
optional layout height) and pins the fold flavour (626x590, 466x378, 466x378),
the rotation flavour (393x359, 852x150, 852x160) and the eventual close. All
three fail against the old line.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Apple's "Designing for iPhone Duo" asks an app to adapt to both displays,
to stay continuous as the device opens and closes, and to treat the band a
partly-open display folds through as a reserved region. Three things here.
1. A visual-viewport resize that changes the WIDTH is the device changing
shape (a rotation, or a foldable opening or closing) and is never the
virtual keyboard, which only ever takes height. handleViewportResize()
read any height drop over 150px as the keyboard appearing, so closing a
Duo (890 to 678pt tall) latched keyboardVisible with no keyboard on
screen: the accessory bar appeared, main grew 84px of dead padding, and
updateAppHeight() stopped refreshing --app-height. The latch was sticky,
because clearing it needs the height back within 100px of a baseline
belonging to a display the user is no longer looking at. Rotating any
phone hit the same latch. The shape branch re-baselines instead, which
is also what lets a keyboard opened after the fold be detected.
2. The hinge is now a reserved region in CSS. --fold-inline-end and
--fold-block-end measure the strip to keep clear from the Viewport
Segments env() variables, and are 0px everywhere else, so the seven
centred overlays are inert by construction off a foldable. Each shrinks
its content box with padding rather than the box itself, so the backdrop
still covers the far side of the fold and still swallows taps there.
3. iPhone Duo (outer) and iPhone Duo (inner) join the mobile device
registry, derived from Apple's published pixel specs at 3x.
Verified in Chromium: flat, a dialog stays centred at 313 of a 626pt
viewport; in book pose it centres at 153 inside the 0-305 leading segment
with its right edge at 293, while the backdrop still spans all 626. The
3-term calc on the offline overlay resolves to 367px in tabletop pose and
20px flat.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three instructions a future contributor would follow literally were stale after
the last review round: the "Adding a CLI" checklist sent the agent-image reason to
AGENT_IMAGE_SPECIAL_CASES, a constant that no longer exists (it is
discovery.install.agentImageLayer on the entry in stock.ts), the trust-boundary
paragraph credited the embedded-commands pin to the invariants test when it is
test/cli-catalog-sync.test.ts, and install.sh claimed "the parity test" pinned the
DeepSeek Harness banner when no test did. That pin now exists: the invariants test
asserts the script's grep literal and the registry's discovery.identity.regex agree
on "DeepSeek Harness", and the comment names it.
docs/docker-cases.md separated the two reasons a CLI stays out of the shared npm
layer (no npmPackage at all versus an agentImageLayer entry), which it had folded
into one, and architecture-invariants no longer lists the agent image's CLI set by
hand.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Choosing "s" (Skip) in the new catalogue-driven install menu warned, printed the
install hints and then fell into the shared "The selected AI CLI failed to install"
gate one line below, because CLI_FOUND_COUNT is 0 by construction inside that block
and skipping does not change it. The AI CLI check runs before the clone and the
build, so a user who picked the documented skip option ended up with nothing
installed. The code this replaced guarded the gate with an elif on the skip choice.
The menu moves out of main() into offer_ai_cli_install() and the gate moves inside
the install branch: skipping continues to the clone, a chosen install that leaves
nothing behind is still fatal. Being a function, the interactive path can now be
driven with a stubbed read_reply, which is what nothing reached before: two
behavioural tests in test/install-sh-invariants.test.ts run the real function in a
real bash (skip continues with exit 0, a failed install dies with exit 1), and the
bash 3.2 CI step drives the skip path in the container as well.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Follow-up to #390. PHONE_MAX had become 599, an inclusive bound, while
three of its four consumers still read it as exclusive (width < PHONE_MAX
for phone); the one site that switched to <= disagreed with
getDeviceType(). It is 600 again with < at every site. The breakpoint
table in docs/mobile-testing-report.md says 600, and the three 430px
visual baselines are removed: they depict the tablet tier now, and the
visual suite recreates a missing baseline on its next run on the machine
that owns them. device-matrix.test.ts is also run through Prettier, which
the commit hook demanded and the format gate (src/ only) never did.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`scripts/pr-bot/` was maintainer tooling, not part of the server, the CLI or the
npm package: a Telegram bot that reviews open pull requests in Codeman sessions
and reports to the maintainer. It now lives in its own private repository and
keeps running unchanged, as a client of Codeman's HTTP API like any other.
It moved because it grew a second watcher, for GitHub Discussions, and shipping
that here would mean publishing the briefs it hands its review sessions, the
judgement calls in them and its safety model. None of that helps anyone
installing Codeman, and all of it is easier to change when it is not a public
interface. The move cost nothing structurally: the whole tree depended on one
external package plus Node builtins.
What this removes from the repo, and nothing else: the sources, their three test
files, `config/tsconfig.pr-bot.json`, `docs/pr-bot.md`, the `pr-bot` npm script,
the bot's globs in the typecheck/lint/format scripts, and its knip entry. CLAUDE.md
keeps a short pointer in place of the section, because the bot still constrains
work in here: it takes the `prbot-<n>` and `dscbot-<n>` session names on the local
Codeman, holds clones under `~/.codeman/pr-bot/`, and fetches pull-request heads
into `refs/pr-bot/*` of this checkout, which it must never check out or reset.
The CHANGELOG entries from 1.25.0 and earlier still describe it. That is history
rather than drift, and is left alone.
Verified after the removal: typecheck, lint and format:check clean, and the suite
passes 6843 tests across 357 files, which is the previous run minus exactly the
70 tests that moved out with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A remote case's workingDir is an absolute path on the remote host, but the
file read routes resolved it with local `fs`: `validateSessionFilePath`'s
realpathSync fails for a path that does not exist on the Codeman host, so
every preview of an agent-written file answered "File not found" (#415).
Add src/remote-files.ts as the single remote-read layer, built on the same
buildSshConnectionArgs() the launch uses:
- remoteProbePaths(): ONE round trip returning realpath + stat for the
requested path AND the workspace root, so containment is checked against a
remotely canonicalized root (a symlinked remotePath is ordinary).
- remoteCreateReadStream(): streams the body (cat, or tail -c +N | head -c L
for a Range) with nothing buffered in memory, and reaps the ssh child when
the response ends so an aborted download cannot orphan it.
- remoteReadFile(): bounded read for file-content.
file-raw, file-content, file-preview and file-thumbnail now share one local/
remote target resolution. Guards keep their local strength: lexical pre-check,
remote realpath, workspace containment, sensitive-path blocklist, and the size
cap applied to the remote size before any bytes are read. An unreachable host
answers 502 with the remote reason instead of a misleading 404. Nothing is ever
copied to the Codeman host and there is NO local fallback (an sshfs mount of
the same tree must not shadow the remote bytes).
Deliberately unchanged: writes (edit=1 / PUT now answer 400 explicitly while
the viewer hides its Edit affordance), office previews, thumbnails, file tree,
picker, external attachment registration and tail-file stay local-only.
With the repo root as the plugin root, `claude plugin install codeman@codeman`
copied the whole checkout into its cache and, because that root carries a
package.json, ran an npm install there: 832 MB, 511 packages and this repo's
postinstall build on every installer's machine (measured from a clean worktree
of the previous commit). A plugin root must be a directory without one.
The plugin is now `plugins/codeman/`: its manifest, a README, and a MIRROR of
`skills/codeman/`. A mirror rather than a symlink because the install copies
the plugin directory and a link pointing outside it would dangle; a mirror
rather than the source because every install path, injector and doc already
names `skills/codeman/`. `scripts/sync-plugin.mjs` (replacing
sync-plugin-version.mjs) mirrors the skill and syncs both manifest versions
inside `version-packages`; `test/plugin-manifest.test.ts` pins byte-identity,
the versions, the absence of a package.json in the plugin root and that the
repo root `.claude-plugin/` holds only the marketplace manifest.
`claude plugin validate --strict` now passes for both the plugin and the repo
root. Install commands are unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`.claude-plugin/marketplace.json` at the repo root makes
`/plugin marketplace add Ark0N/Codeman` work, and the one plugin it lists is
the repo itself (`source: "./"`), whose one component is `skills/codeman/`.
So `/plugin install codeman@codeman` is a third install route next to
`npx skills add` and `codeman skill install`, and the skill shows up in the
plugin directories that index Claude Code marketplaces.
Both manifests carry package.json's version: `scripts/sync-plugin-version.mjs`
rewrites them inside `version-packages`, right after `changeset version`, and
`test/plugin-manifest.test.ts` pins the equality, the skill's frontmatter name
(without it the installed skill would be named after a versioned cache dir),
and that no other plugin component (`commands/`, `agents/`, `hooks/`,
`.mcp.json`, `settings.json`) appears at the repo root, since an install would
silently ship it.
Verified with `claude plugin validate` (one expected warning: CLAUDE.md at a
plugin root is not plugin context) and a local marketplace add, install,
details, uninstall cycle against a clean checkout of this commit.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Bold text on the theme's default foreground carries exactly ONE cue, the
weight step. Claude Code marks its markdown bold with a bare ESC[1m and
changes no colour, and xterm substitutes a bright colour for bold only
when the foreground is a palette index 0-7, so the substitution never
fires for default-foreground text. A family shipping only a regular and
a bold face keeps that step small (measured on Consolas: glyph ink rises
from 14.25% to 16.57%), and picking a different family does not help,
because 400 stays 400 whatever the family. Lowering the NORMAL weight is
the only way to widen the gap.
Two per-device settings beside "Terminal font" in the Font group, each
defaulting to xterm's own value for its slot, so an untouched install
renders exactly as it did before. Both thread into the main terminal and
the Agent Teams panes, and apply on save without a reload.
The bundled face had to be unclamped in the same change or the settings
would look broken on a stock install. fonts/jetbrains-mono-variable.woff2
carries a wght axis of 100 to 800, but styles.css declared the face
`400 700`, and the descriptor is what the browser synthesizes from: at
that range 100, 200 and 300 rendered identically to 400 and 800
identically to 700 (measured in headless Chromium, both directions).
The two families ahead of it in the default stack, Fira Code and Cascadia
Code, exist only if the user installed them, so for most installs
"normal = 300" would have been a no-op. Declared `100 800`, every step is
distinct: 61%, 77% and 90% of the ink at 400, and 800 adds ~14% over 700.
Nothing in the stylesheets asks for a monospace weight outside 400-700,
so widening it changes nothing that rendered before.
Details that are easy to get wrong and are pinned by tests:
- Each slot falls back to its OWN xterm default, so an unset bold weight
can never inherit `normal` and become a visible change.
- A live save refreshes both echo overlays. They cache
terminal.options.fontWeight and paint it into their spans, so without
it the characters being typed keep the old weight while the rest of the
screen changes. Most visible on a phone, where local echo is on by
default.
- A live save reaches open Agent Teams panes, which read their options at
construction, exactly as applyTerminalSkin() propagates its own.
- A stored weight the picker does not list (a hand-set 350) is added to
the select rather than dropped, so merely opening App Settings cannot
reset it.
- _awaitTerminalFont() is untouched. CharSizeService measures through the
CSS `font` shorthand, which resets the weight, so the measured face is
always the 400 one and a weighted descriptor would request nothing new.
Verified end to end in a headless browser against a live server: the save
reaches the running terminal with no reload, the settings PUT stays 200
(both keys are display keys and are stripped before it, since
SettingsUpdateSchema is strict), the value survives a reload, and the
painted terminal really changes weight with the bundled font (lit-pixel
ink 0.83 / 0.95 / 1.00 / 1.13 / 1.21 at 100 / 300 / default / 700 / 800).
Proposed and analysed by @irisitymichaelgrundberg in discussion #403.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
buildClaudeEnv(), the direct-PTY fallback taken when mux creation fails, now
reads getCli('claude').env and applies its unset and exports lists. It used to
delete COLORTERM and CLAUDECODE from a hand-maintained list of its own, which
left it contradicting the registry entry that the tmux pane and the attach
client both read. An engine value needing a mux name has nothing to resolve
against on this path, so it is skipped rather than guessed.
Claude no longer unsets NO_COLOR. The invisible-background bug does not need
it, and unsetting it overrides a preference the user set deliberately, so a
user who exports NO_COLOR globally keeps monochrome panes. The other seven
truecolor CLIs still unset it; that inconsistency is intentional and the
comment on the entry says so.
The invariants doc gains a Terminal colour env paragraph under Session launch
modes, where a reader looking up Claude will find it — the previous sentence
sat under a heading that lists only the non-Claude CLIs. It now says the lists
are the stock catalog and a clis.json override replaces them wholesale, and
that the declarations reach the tmux pane, its attach client and the direct
PTY but not a remote pane, whose command carries no env exports at all. Docker
hands COLORTERM=truecolor to every mode, including the two the registry says
must unset it.
The changeset named six peer CLIs and there are seven: deepseek also exports
truecolor. A test beside the existing OpenCode assertion pins the new
behaviour, so a future registry edit cannot make the backgrounds vanish again
in silence.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Rebased onto current master (the one real conflict was the import line
in docker-hosts.ts Ark0N flagged; kept both), then addressed every
point from the review:
**1. Rebase.** Done — this branch now sits on current upstream/master.
**2. Agent-image special cases are data now, not an id-keyed table
outside stock.ts.** `AGENT_IMAGE_SPECIAL_CASE_IDS`/`AGENT_IMAGE_SPECIAL_CASES`
are gone. `CliDiscovery.install.agentImageLayer?: { kind: 'dedicated';
reason: string }` is a field on the registry entry itself (pi,
deepseek), `reason` is required by schema.ts, both producers
(docker-hosts.ts and cli-catalog.mjs) filter on its presence instead
of an id, and the coverage test reads it from the generated catalogue.
Also added the npm-package-name validation to the TS producer, which
only the .mjs one had — same SAFE_PACKAGE regex, duplicated
(necessarily, one side can't import the other) and now pinned
byte-identical by a new parity test.
**3. Changeset said five, it's eight.** (Not nine — see the DeepSeek
point below, which changes the true count.) Reworded to state it
structurally rather than pin a number that will go stale again.
Then the four behavior-changing findings:
- **DeepSeek was offered as a normal install option but can't actually
drive a pane.** `npm install -g @deepseek-ai/dsh` installs the
launcher only; DeepSeek ships no profile that can run standalone.
The generator now emits an empty install command for any
`launcherProfile` entry, so install.sh's menu (which requires a
non-empty command) skips it and falls through to its docs URL hint
instead — matching what the old hand-written code did before this
PR replaced it.
- **wget-only hosts lost every automatic install, including the npm
ones that never needed curl.** The menu-building loop now filters
PER ENTRY (only a command starting with `curl ` is held back) rather
than wiping the whole menu when DOWNLOADER != curl.
- **The DISPLAY/TRUSTED split and the catalogue refresh didn't hold up
under review** (refresh's only real write was the label; it ran
before the Node existence check; its own eval-detection test was
tripped by the word "eval'd" in a comment). Dropped entirely per
your own recommendation — embedded catalogue only, no network
fetch, no second array. install-sh-invariants.test.ts now asserts
the refresh/DISPLAY machinery does not exist rather than testing its
internals.
The three take-or-leave items, applied:
- `dsh_banner_probe`'s bash 3.2 empty-array bug: `${runner[@]}` →
`${runner[@]+"${runner[@]}"}`. Verified live in a real `bash:3.2.57`
container with `timeout` removed from PATH — crashed before, clean
now, full `detect_all_clis` path exercised end to end.
- `docker-agent-image-coverage.test.ts` now anchors on each layer's
`<binary> --version` proof line instead of `Dockerfile.includes(binary)`,
which stayed true if a layer were deleted but its comment survived.
- Doc drift: docs/docker-cases.md (four → five, and now describes the
data field), docker/agent.Dockerfile's "other four CLIs" comment (no
longer a magic number — CLI_NPM_PACKAGES is generated and can grow),
CLAUDE.md's install.sh size (104KB → ~112KB) and its stale mention of
the now-dropped refresh.
Verified: tsc clean, prettier clean, the full targeted suite (142
tests across the 8 affected files) green, and the full `npm test` gate
diffed BY TEST NAME against a clean upstream/master baseline run on
this same machine — identical 201-name failure set both sides (168
tests / 67 files, all pre-existing Windows-environment noise: symlinks,
PTY spawning, POSIX permission bits — none of it touching anything
this PR changes), zero new failures either side of the diff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
docker/agent.Dockerfile hardcoded the four npm-published CLIs it installs, one
of the several lists that had to be kept in step with the registry by hand.
It now takes them as `ARG CLI_NPM_PACKAGES`, supplied by
scripts/build-agent-image.mjs from config/clis.stock.json, with the default set
to today's list so a bare `docker build` still produces the same image. The arg
is expanded unquoted because word splitting is what turns the list into several
arguments, which is exactly why every token is validated against
^[@A-Za-z0-9][@A-Za-z0-9/._-]*$ on the producing side; a package name carrying a
space or a metacharacter is refused rather than reaching the RUN line. Verified
by building the layer: four packages in, four arguments out, and the default
still applies with no arg.
The list is filtered on each entry's `enabled` flag — the field whose absence
was the maintainer's §3 finding, where a CLI shipping disabled still got baked
into every image. No stock entry is disabled today, so that assertion would pass
vacuously; a unit test feeds the pure helper a fabricated disabled entry so the
fix is covered now rather than the first time someone ships one.
⚠️ It reads the STOCK catalogue, never the merged registry. A user's
~/.codeman/clis.json must not change what is inside an image tagged
codeman/agent:base, or two machines holding that tag hold different images.
Four CLIs keep hand-written layers because the registry cannot describe what
makes them special: pi's --ignore-scripts, deepseek's pnpm companion and dsh-tui
profile, and the three standalone installers. Rather than extend the schema for
a Docker-only benefit, the coverage test requires each to carry a written reason
AND still be present, so an exclusion cannot quietly become an omission.
There are two producers of this command line and there have to be — the .mjs
cannot import TypeScript, and src/docker-hosts.ts builds the same argv for the
in-app auto-build — so a parity test pins them together, package list, arg pairs
and rendered argv. Their order is pinned too: a different order is a different
RUN string and so a needless cache miss between the two build paths.
docker/server.Dockerfile is deliberately NOT edited (PRs #373 and #377 both
modify it); its narrower list is asserted as a declared omission list instead, so
the divergence is reviewable without touching the file.
Also fixes the in-app hint at index.html, which the new coverage test caught
still omitting omp.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12
install.sh carried nine search-path arrays, eighteen near-identical
check_<cli>/get_<cli>_path functions, and three separately hand-maintained
enumerations of all nine CLIs. They had to agree and did not: upstream b6d0f1fa
is "wire OMP into install.sh's CLI detection (it had none)", and the section
comment above the roll-call named six of the nine.
All of it now reads the generated catalogue. `detect_all_clis` resolves every
CLI in one memoized pass into CLI_FOUND_PATH/CLI_FOUND_COUNT; `check_cli` and
`get_cli_path` replace the eighteen pairs; the roll-call, the "no AI CLI found"
gate and the closing reminder become loops. Probe order per CLI is unchanged and
`test/install-sh-detection-parity.test.ts` proves it against the literals
transcribed from the arrays this deletes.
Behaviour changes worth naming:
- The install menu is built from the catalogue, so it offers every enabled CLI
that is not installed and ships a command — five instead of two. Gemini had a
command in the registry and appeared in NO list in this script.
- Its labels are now the registry's ("Claude" rather than "Claude Code"), the
same trade PR A made for `codeman doctor` rows. A suffix map would just be the
hand-maintained list again.
- On a wget-only host the menu prints commands instead of running them. The
registry's commands call curl, whereas the two literals this replaces went
through download_to_stdout; rewriting curl to wget inside a string we are
about to execute is the wrong instinct.
The trust boundary is mechanical, not a promise: CLI_INSTALL_CMD_TRUSTED is
written only from the generated per-platform arrays and is the only thing ever
executed; CLI_INSTALL_CMD_DISPLAY is what the optional, opt-in refresh may
rewrite. The refresh warns on all three failure shapes — empty body, unparseable
content, failed fetch — which is the silent-degradation bug from the review, and
it parses with node into tab-separated records read by `read`, never eval.
Bash 3.2 throughout (macOS ships it): parallel indexed arrays, offset/length
windows instead of delimiters, no associative arrays, namerefs, mapfile or
here-strings. Verified by executing the script under a real bash 3.2 container,
which is also now a CI step alongside `bash -n` and a catalogue `--check` — the
empty-window case (`shell` has no binaries) is a runtime `set -u` abort that
`bash -n` cannot see. Running it that way caught `detect_os` being called inside
the platform loop: ten forks, and ten copies of one error, since a `die` inside
`$( )` can only exit the subshell.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12
Two consumers of the registry cannot import TypeScript: `install.sh`, which runs
via `curl | bash` before any checkout exists, and `scripts/build-agent-image.mjs`.
Both currently hand-maintain their own CLI lists, and both have already drifted.
`scripts/generate-cli-catalog.mts` (`npm run generate:cli-catalog`, plus a
`--check` mode) emits from `STOCK_CLIS`:
- `config/clis.stock.json` for the `.mjs` and the tests. It carries `enabled` —
the field the earlier attempt omitted, which is how a disabled CLI's npm
package still got baked into every agent image.
- a marker-delimited block inside `install.sh`, embedded rather than fetched.
The embedded copy is the FULL catalogue on purpose: the earlier design fetched
it and fell back to a hardcoded two-CLI list, degrading silently on an empty
response. There is no degraded mode to fall into now.
The block is bash 3.2 safe: parallel indexed arrays, no associative arrays, no
namerefs, no mapfile. Variable-length lists use OFFSET/LENGTH windows into one
flat array rather than a delimiter, so a $HOME containing a space needs no IFS
handling and `shell` (no binaries) gets length 0 and is never iterated. Search
paths are emitted dir-major, matching the probe order the hand-written arrays
use and `test/install-sh-detection-parity.test.ts` pins.
Only fields the two consumers need are exported. `launch`/`env`/`capabilities`/
`overlays` are spawn-time concerns the server alone interprets, and a test
asserts they never leak into the artifact.
`main()` sits behind an `isMainModule()` guard so the sync test can import the
renderers. Without it, importing the module would rewrite the artifacts as a
side effect of checking them — passing always, guarding never.
This commit adds the block; it does not yet delete the hand-written arrays, so
the detection pin keeps measuring both against each other.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12
PR B replaces nine hand-written `*_SEARCH_PATHS` arrays in install.sh with one
block generated from `STOCK_CLIS`. This lands FIRST, against the hand-written
arrays, so the replacement has something to be measured against.
The arrays are not uniform, which is why "generate them from the registry" is a
claim rather than an obvious truth: claude alone has `~/.claude/local`, opencode
alone has `~/go/bin`, opencode/codex/gemini/pi/omp carry `~/.bun/bin` while
dsh/grok/agy do not, and omp's `~/.omp/bin` sits second rather than first. A
generated list that silently narrows leaves a user with that CLI installed being
told no AI CLI was found — upstream `b6d0f1fa` is that bug, fixed for omp by
hand after it shipped.
The test asserts a three-way identity: the pinned literals equal what install.sh
contains today, AND equal `searchDirs x binaries` from the registry, dir-major so
the probe ORDER is pinned too and not just the set. Both halves were verified to
fail independently — dropping one path from install.sh fails the first, changing
one `searchDirs` entry fails the second — because a pin that cannot fail is
worse than no pin. A fourth case asserts every stock CLI with a binary is
covered, which is the omp bug restated so it cannot recur silently.
No production code changes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12
CI caught a real regression: DEEPSEEK_API_KEY was added to deepseek's
privilegedEnvKeys alongside DEEPSEEK_BASE_URL on the theory that "the pair
travels together," but that contradicts the documented and tested design
(clampEnvOverridesForOwner()'s own docstring in session-routes.ts) — a
non-granted owner supplying their OWN DeepSeek key removes privilege
rather than granting it, since the exfiltration vector is the BASE URL
(which redirects the server's own forwarded key to a foreign host), not
the key itself. Removed it from the list; test/deepseek-mode.test.ts's
existing two clamp tests now pass again.
Also swapped that test's "unrelated override" example off CODEX_HOME,
which the earlier commit in this same PR legitimately made privileged
(closing a real pre-existing gap, documented in PR.md) — so it stopped
being a valid "unrelated" example the moment that fix landed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.
Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.
Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:
- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
undocumented GATEWAY AuthType gemini-cli selects once
GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
tried
- deepseek: reaches the server (env vars are read) but gets a consistent
HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)
Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).
deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its
native cloud backend, for a given session. Covers local hardware (llama.cpp,
Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter).
Off by default (customModelEndpointsEnabled, synced, default OFF).
- Registry: capabilities.customModelInjection per CLI entry (env /
configContentEnv / configDir / unsupported kinds)
- Pure injection builder (custom-model-injection.ts) turning an endpoint +
model id into the real env vars / config content per CLI
- Endpoint store + CRUD routes (custom-model-hosts.ts,
custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded
- Session integration: Session.setCustomModel()/restartCli()
(POST /api/sessions/:id/custom-model), reusing the existing
respawn-pane -k primitive to restart the CLI process with new env
- Multi-user hardening: every new redirect-capable env var added to its
CLI's privilegedEnvKeys, closing a pre-existing gap where several were
already reachable via the generic envOverrides field's prefix allowlist
- Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries
against a real endpoint outside the web UI, independent of tmux/sessions
- Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying
every CLI's injected values through a real HTTP shape
Real end-to-end validation against a live llama-swap server (inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries) found and
fixed three real bugs before they shipped:
- Codex's config.toml schema was wrong ([model].default table instead of
a top-level model string + [model_providers.custom]); fixing it then
surfaced a genuine, documented protocol incompatibility (Codex only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement)
- Claude Code's async session-title-generation call validates
ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and
hangs the whole -p invocation on an unrecognized name; documented for
chunk 6, worked around in the standalone script only (--bare is NOT
safe for a real interactive session, which needs hooks)
- The discovery route's authStyle: 'both' option (send both Authorization
and api-key headers) reliably hung a real server; removed the option
entirely rather than just changing the default
Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see
PR.md and deployment_plan.md for the full chunk breakdown and confidence
table.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
A Fable 5.1 reviewer read the whole release diff against 1.26.2 and returned
SHIP WITH FIXES. These are its findings, verified before acting on each.
**The changelog advertised a feature the code refuses (major).** The #401
changeset and docs/web-tabs.md both listed `*.localhost` in the loopback set.
The follow-up in 02b0e278 moved it out of the auto-route set on security
grounds and updated CLAUDE.md but neither of those, and that changeset becomes
the 1.27.0 CHANGELOG entry: a user would have read the release notes, tapped
`http://app.localhost:3000/` on a phone and got a connection error from a
documented feature. Both corrected, and the user guide now says why it is
excluded and that adding such a dashboard by hand still works.
**Dictation delivered its text twice (minor, #388).** `keydownSnapshot` started
`null`, so `keydownSnapshot ?? canonicalCount` at the input event read a counter
xterm had ALREADY bumped: on a fresh page load with no keydown yet, xterm's own
capture listener forwards the `insertText` itself (it is not gated behind a
keydown), then the snapshot equals the bumped count, `count > snapshot` is
false, and the controller emits the same text again. Reproduced directly
against the module: it emitted `hello` for input xterm had already delivered.
A `0` baseline restores that file's own invariant, that a missed recovery is
acceptable and a duplicated keystroke is not. Two regression tests, covering
both the xterm-already-delivered and genuinely-dropped halves.
**The sorted rail's arrow-key walk followed the DOM (minor).** `_tabKeydownHandler`
steps `querySelectorAll` order, which is `sessionOrder`, while a sorted rail
paints its rows with the flex `order` property, so ArrowDown from the top card
landed wherever that session happened to sit in the tab order. It now sorts its
node list by the COMPUTED order first: computed rather than inline, because web
tabs take their `order: 9999` from CSS and would otherwise read as 0 and lead
the walk. This is the one place that follows the paint; the Alt+N badge, the
drag model and the filter all still deliberately read the DOM.
**A trusted dashboard was auto-reused by a tapped link (minor, #401).** The
reuse loop skipped `managed` and direct-mode records but not `trusted`. A
trusted frame is mounted with `allow-same-origin`, i.e. on Codeman's origin
with the user's cookie, and these links come from agent output, which is the
threat model the loopback allowlist was just narrowed for. An agent that can
write into the dev server's tree could print a path that one tap opens inside
that privileged frame. Excluded from auto-reuse, with a test; opening it from
the Run dropdown is still an explicit action and unchanged.
**Two documentation claims that were no longer true.** CLAUDE.md said
test/location-overlay-commands.test.ts pins every remote pane command, but
remote claude and remote omp now have their own arm in `buildRemoteLaunchCommand`
and never reach `defaultRemoteCommandForMode`, which is what that test asserts,
so it pins nothing for them and changing either arm will not fail it. Named the
real pins instead. Also documented the arrow-key-walk exception in the rail
paragraph.
Left as follow-ups, deliberately: `POST /api/webviews` does not dedupe by URL
server-side, so two devices tapping one link concurrently can still save two
dashboards for one origin (pre-existing endpoint behaviour that #401 makes
reachable by a tap), and the location-overlay golden should assert the real
remote claude/omp commands rather than a branch neither reaches.
Full gate green: 359 files, 6869 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The vertical rail lists exactly the sessions both home screens list, so it now
answers their question the same way instead of showing the raw tab order.
Order: new per-device `tabRailSort` (App Settings -> Appearance -> Tabs ->
Vertical Rail Order, default "By activity"). It runs `CodemanSessionOrder` over
rows classified by `_mobileOverviewState`, i.e. literally the home screens'
comparator, `lastSubmitAt`-anchored running group included.
It is applied as the flex `order` property, never by reordering the DOM.
`#sessionTabs` stays in `sessionOrder`, which is what keeps the Alt+N badge
honest (it names a shortcut, not a row position, so it deliberately does NOT
run 1,2,3 down a sorted rail), and keeps drag-and-drop, the arrow-key walk, the
sidebar filter and `_scrollActiveTabIntoView()` all reading the list they
always read. A session changing state then moves one inline style instead of
forcing the full rebuild that would restart every card's animation on every SSE
tick. The incremental render path re-applies it, since a state flip adds no tab
and never reaches the full rebuild, and an empty string is what clears it when
sorting stops. Web tabs are pinned past the cards by a CSS `order: 9999`, since
`renderWebviewTabs()` emits the same markup for every layout and the flex
default of 0 would interleave them. Drag is switched off while sorting (the
drop rewrites `sessionOrder` correctly and the sort puts the card straight
back, so the affordance would be a lie); 'manual' is the way back.
Cards: detailed rail rows become bordered cards on `--bg-card`, with the stamps
line on its own full-width row and the pill at its right end. The state dot
goes 6px to 9px, keeps its orbiting ring while working and gains the green
halo; idle mutes toward `--text-muted` as the home rail does. Needs/error/
waiting reuse `home-sessions-blink-red`/`-yellow` rather than a second copy.
These card rules are RAIL-SCOPED and deliberately absent from the comma-grouped
selectors that carry both vertical surfaces: the rail is an occasional,
resizable list you scan, while the sidebar is a permanently-docked nav column
where 20 stacked cards read as a wall. Every state-dot rule also excludes
`.tab-alert-action`/`.tab-alert-idle` by hand, because those alert rules are
only (0,3,0) and these are (0,5,1)+.
Lines: the lineage bracket already drew in the rail, but its track sat 6px from
the left edge, so half of its 11px outer glow was clipped by the window frame
and it read as a thread pinned to the frame. It now runs at 10px, mid-channel
in the gutter the rail already reserves.
Tests: test/tab-rail-order.test.ts (17) drives the real `isTabRailSorted()` and
`_tabRailSortOrder()` out of app.js, covering the row model (a WORKING row
ranked by `lastSubmitAt`, which would otherwise rank every running turn as
freshly started and fail no rendering test), the Alt+N badge, and the opt-out.
Verified in Chromium across sorted/manual/simple/header-strip/sidebar with the
setting flipped at runtime: no page errors, and the header strip and sidebar
render byte-identically to before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Each item is from the pre-merge review of the PR it names, applied on master
rather than by pushing to a contributor branch.
#400 (response viewer, shenlvkang-collab)
- The brief view opened at `scrollTop = 0`, right when it was a single card
holding the last row. Now that it renders the whole turn, the top is the
turn's first narration line and the answer can be screens below it, while
loadFullContext already scrolls to the bottom of the same turn. A multi-row
turn now opens at its newest text; a single card still opens at the top.
#401 (loopback links as web tabs, shenlvkang-collab)
- Drop `*.localhost` from the auto-route set. Every other member is an address
literal that can only mean this box; a `*.localhost` DNS name is not one, and
a resolver with a search domain retries `evil.localhost` as
`evil.localhost.<search domain>`. The link source is agent-written terminal
output, so that set is the whole confinement on a tap that makes Codeman
fetch a URL server-side and persist it. The page-side test stays broader
(`isOnBoxHostname`), where a false positive only declines to proxy.
- A link to the origin root navigated nothing: the path was flattened to '',
which openWebview reads as "no deep link", leaving an open frame where it was.
- `this.webviews` being set does not mean it is loaded. initWebviews() assigns a
truthy empty map and only then awaits the list, so a tap during page load
found nothing to reuse and POSTed a duplicate record. Join the in-flight
refresh instead.
- One dashboard per dev server rather than per host spelling, which is what the
method's own comment already promised.
- Toast on the auto-create: it writes webviews.json, broadcasts over SSE and
adds a Run-dropdown row on every signed-in device, with a new tab as its only
previous signal.
#362 (remote omp continuation, timkjr)
- Accept the allowlisted `mode === 'omp'` arm as-is; a blanket registry render
would hand deepseek a locally-resolved --profile and bypass claude's own
overlay. A registry-declared switch is the follow-up if a third mode needs it.
- Revert the whole-file Prettier reformat of docs/remote-sessions.md (docs/ is
hand-formatted and outside `npm run format`), keeping only the two new
sections.
- Correct three stale passages: architecture-invariants' `exec claude
--dangerously-skip-permissions`, the `exec <cli>` paragraph (claude and omp
now have their own arms, and the claude pane's PID is the login shell), and
omp-integration's `-c 'omp'`. RemoteCommandMode gains deepseek and omp.
- Add the missing `_maybeCaptureOmpSessionId` remote-guard test; the sibling
guard in `_pinOmpRespawnId` had one and this path runs earlier, on the first
idle turn.
#388 (keyCode 229 recovery, aakhter)
- Gate notifyCanonicalData on shouldSuppressTerminalQueryResponse and
isTerminalFocusOrMouseReport. onData also carries the DA/DSR/CPR/OSC replies
xterm answers during Ink redraws and its SGR mouse and focus reports; any of
those landing between the keydown and the candidate's resolution was read as
"xterm spoke for this keystroke", standing the recovery down and leaving the
character dropped, worst on a busy agent pane. Reached through
window.CodemanTerminalInput: the predicates live in a module IIFE that closes
long before this call site, so bare references would throw into the
surrounding try/catch and stop the notify from ever running.
Every fix has a test that fails without it (verified by reverting each).
Full gate green on the combined tree: 358 files, 6849 tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>