* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (da07b38c, db4557d9) with no
corresponding update to the plan doc itself — every checklist still read
Status: TODO and every box unchecked. Brings the doc in line with the tree:
- A new "Status as of 2026-09-22" section up top: what's actually
implemented (verified by grepping the routes/schema/UI, not just trusting
the commit messages), the availability-flag staleness bug found and fixed
in this session (commit 0c77dd0a) with its devbox verification record, and
one real outstanding gap.
- The outstanding gap: a custom CLI created via Phase 5's write API has no
way to actually be launched. The Run menu is static per-mode markup with
no consumer of window.__codemanCliCatalog, so Phase 6's own "create a
custom entry, confirm it can be launched" verify step was never actually
exercised against this. Documented with two candidate fixes, neither
started.
- Each phase's checklist flipped to [x] where confirmed present in the tree,
Status lines updated from TODO to DONE, and the two originally-open
questions (Phase 2's installed source, Phase 5's PUT endpoint shape)
marked resolved against what actually shipped.
No code changes in this commit — documentation only, so a future session
(or the one already mid-flight on a separate checkout of this same branch)
picks up accurate status instead of a stale plan.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs: add the CLI-registry deployment plan and the parked Copilot plan
Both were sitting as untracked scratch files in the master checkout,
never committed to any branch. Moving them here rather than leaving them
loose:
- DEPLOYMENT_PLAN.md is the live tracker for the CLI-registry follow-up
series (PR A #347 merged, PR B #380 merged, PR B2 merged as #458) and
is where PR C (this branch's own CLI-management work) belongs.
- docs/copilot-integration-plan.md is explicitly PARKED, referenced by
name in docs/cli-enable-disable-plan.md's own header as a sibling plan
tracked separately — kept for continuity, not active on this branch.
The other scratch files found alongside these (PRA.md, PRB.md, PR-B2.md
and their review-response counterparts) described PR A/B/B2, all now
merged — deleted from the master checkout as stale rather than committed
anywhere, since their content is superseded by the real merged PRs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* fix(cli-registry): render enabled CLIs in launch surfaces
* test(cli-registry): update frontend branch guard
* fix(test): isolate suite from deployment environment
* fix(cli-registry): revise Decision 4 - claude is toggleable, shell stays permanent
shell/claude were both structurally un-disableable in the original plan
(Decision 4). Revised: shell keeps the hard backend guarantee (it is the
one non-agent mode several code paths assume always exists as a raw-
terminal fallback), but claude is now a normal toggleable entry like any
other CLI.
Safe to do because internal session creation (tmux-manager.ts, session.ts,
Ralph, plan-orchestrator) resolves a CLI via getCli(), which does not
check `enabled` at all - only the Run menu and the HTTP-facing
sessionModeSchema() (new session requests through the normal API) key off
it. Disabling claude therefore behaves identically in kind to disabling
any other CLI: no internal fallback path breaks, it just stops being
offered for new sessions until re-enabled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): hide shell's toggle entirely instead of greying it out
A permanently-disabled switch next to every other row's working toggle
read as broken rather than intentional. shell now renders no switch at
all - a plain "Always available" label - so there is nothing to click
that could look like it should work but doesn't. Backend guard is
unchanged (UNDISABLEABLE_IDS still refuses shell unconditionally); this
is UI-only.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): sort the Installed CLIs list, installed-first then alphabetical
renderCliList() previously rendered in registry order (each entry's fixed
order field). Now sorts installed CLIs first, then not-installed, each
group alphabetical by label - matches how a user actually scans the list
(what's ready to use, then what needs installing).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* style: prettier fixes from the master merge
* fix(cli-registry): install/edit take effect immediately, confirm before install, phone labels
Four gaps found verifying #476 against the #343 review trail:
- Installed or edited CLIs kept reading as missing/stale. Every binary lookup
(the nine per-CLI resolvers and the generic registry one) caches in its own
closure, with a negative-cache backoff of up to 5 minutes, and nothing
cleared them. invalidateCliExecutableResolvers(binaries) now drops those
caches per binary; install (success or failure), create, edit and delete
call it plus invalidateCliResolverCache(id). Before this, a CLI installed
from Settings could fail to launch for minutes, and an edited custom entry
kept launching its old binary until a restart.
- The Settings "installed" badge for a custom entry used a private `which`,
ignoring the entry's searchDirs and the login-shell lookup that spawn and
the Run menu use; it now asks the same generic resolver they do.
- Install ran on a single click. The #343 review asked for auto-install to
sit behind an explicit confirm; the confirm now names the exact command,
which GET /api/clis returns for stock entries only (installCommand).
- The phone Run button showed the two-letter tab badge ("CC", "CX") instead
of the word ("Claude", "Codex"). It uses the registry label again, which is
identical to the old static table for every stock CLI (now pinned).
14 new tests; 9 of them fail against the previous head and pass here.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): address #476 review — safe serialized writes, no id branches, docs
Must-fix:
- registry-writer: start fresh only on ENOENT; refuse (409) a clis.json that
does not parse or has group/world permission bits instead of overwriting it
(isUnsafePermissions now exported from registry.ts)
- mutateRegistryFile(): one promise chain for every mutation, with the
existence/duplicate checks inside the serialized step, plus a unique tmp
name per write
- docs: CLAUDE.md, architecture-invariants, cli-registry (new Settings
section) and api-reference (the six /api/clis routes)
- drop DEPLOYMENT_PLAN.md and docs/copilot-integration-plan.md
Smaller:
- PUT /api/clis/custom/:id keeps the entry's current enabled state when the
body omits it
- runMode setter falls back to the first enabled catalogue entry, not 'claude'
- shell guard keyed on kind === 'shell' (routes + Settings list); stock probe
map shared with server.ts via utils/cli-installed-probes.ts
- stock claude label is now 'Claude Code', so the Run menu / phone overview
label rewrites are gone (doctor row keeps "Claude CLI" via its override)
- welcome buttons are translatable again and read "Run Claude Code" /
"Run Shell"; zh-CN gains "Run Codex" / "Run OMP"
- install: per-id in-flight guard (409) and CODEMAN_* stripped from its env
- fileoverview / CliEnableSchema comments no longer say stock-only
- test-env isolation changes moved to their own PR
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* test(cli-registry): pin the #343/#347 findings #476 makes reachable
A CLI toggled or created through the routes is accepted or rejected by
CreateSessionSchema with no restart (#343 finding 2), and a custom CLI created
through the API renders a real local, remote and docker launch command
(#347 finding 5: no more `cd <path> && undefined`).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
25 KiB
The CLI registry
Every run mode Codeman can launch — Claude Code, Terminal/Shell, OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek Harness and OMP — is a CliEntry: a data record describing how to find the binary, how to build its command line, what environment it needs, and what it can do. Code that used to ask "which CLI is this?" asks the entry instead.
Where it lives
| File | What it holds |
|---|---|
types.ts |
The CliEntry interface and everything under it. Read this first. |
stock.ts |
The shipped catalog. The only file allowed to name a CLI id. |
schema.ts |
Zod validation, including the cross-field checks that reject an incoherent entry at LOAD time. |
argv.ts |
The argv engine: the only code that turns typed tokens into a command string. |
patterns.ts |
The NAMED value patterns (model, uuid, path-segment, …) and the regex-compilation guard. |
profiles.ts |
The names of behaviours that genuinely need code, kept import-free so schema.ts can validate one. |
registry.ts |
Loading, merging ~/.codeman/clis.json, and the accessors (getCli, enabledClis). |
src/session-cli-registry-bridge.ts maps the legacy per-mode option bag onto the engine, and src/utils/cli-resolver.ts / src/utils/cli-launcher.ts do registry-driven binary resolution and launcher-profile dispatch.
The override file
~/.codeman/clis.json (instance-scoped through dataPath()) holds overrides and custom entries only, never a copy of the stock catalog: { "clis": { "<id>": { ...partial entry... } } }. Objects merge key-wise onto the stock entry, arrays replace wholesale. The file must be mode 0600; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you chmod 600 it. Every reason a file was ignored or an entry dropped is logged once, prefixed [cli-registry], on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read after a change made through CLI management (below).
Managing CLIs from Settings
App Settings → Agents & CLIs → CLI management (cliManagementEnabled, default OFF; admin-only in multi-user mode) lists every entry with an installed/not-installed badge and:
- toggles any entry on or off. A
kind: 'shell'entry cannot be disabled, and the row shows no switch for it. A disabled CLI disappears from the Run menu, the welcome screen and the phone overview, and new session requests for it are rejected. - installs a missing stock CLI by running its shipped install command, after a confirm that names the exact command. Only one install per CLI runs at a time, and the command runs without any
CODEMAN_*variable in its environment. A custom entry's install command is never executed. - adds, edits and deletes custom entries (id, label, badge, binaries, launch argv). The server re-validates the whole assembled entry through
CliEntrySchema, so the form cannot bypass the load-time rules.
These are the only writes to clis.json. They are serialized, and a file that does not parse or has unsafe permissions is refused rather than overwritten; fix it (or chmod 600 it) and retry. The HTTP routes are listed in docs/api-reference.md under CLI management.
The shape of an entry
interface CliEntry {
id: CliId; // 'codex'
label: string; // 'Codex' — shown in menus
shortBadge: string; // tab badge, e.g. 'CX'
accent: string; // single hex colour
enabled: boolean;
stock: boolean; // set by the loader; a custom entry can never claim it
order: number;
kind: 'agent' | 'shell';
discovery: CliDiscovery; // how to find and prove the binary
launch: CliLaunch; // the structured argv template
env: CliEnv; // exports, tmux setenv keys, the env-override allowlist
capabilities: CliCapabilities; // what every call site reads instead of the id
// .workDetect?: { promptGlyph, workingLine, watchingLine?, watchingLines? } — how
// this CLI's pane shows work, and how it shows work it started in the background
overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store
}
capabilities is the important part. It is what isExternalCliMode(), isAltScreenStripMode(), hooksAvailableForMode() and every other former per-mode branch actually read.
Regexes that come from config
Three capability fields carry a regular expression an override file can set: discovery.version.regex, capabilities.workDetect.workingLine and capabilities.workDetect.watchingLine. All three go through compileVersionRegex(), which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns null rather than throwing so every caller degrades instead of crashing.
workingLine is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: schema.ts rejects the entry at LOAD time so a bad pattern never reaches a session, and _workingLinePattern() in session.ts compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused.
watchingLine reads a different row of the same screen. A CLI draws it while work the agent
itself started is still running — Claude prints ⏵⏵ bypass permissions on · 1 monitor · ← for agents while a monitor, a backgrounded shell or a cloud session is live. Codeman turns that
into Session.watching, and an idle prompt from such a session opens already acknowledged,
so a pane waiting for its own background work never raises an alert a human cannot answer.
Group 1 is the label, and a CLI that declares no pattern reports no background work.
Two CLIs declare such a row today, and they put it in different places. Claude writes its
chip on the last row of the screen, so it keeps the default one-row window and anchors on
the · its footer joins items with. Codex pins
1 background terminal running · /ps to view · /stop to close ABOVE its composer, which
puts the row third from the bottom once the status line and the composer are counted, so its
entry declares watchingLines: 3 and matches that row end to end. Both were measured
against live panes rather than read out of a binary, which is the standard for adding a
third.
That label is the one value in the registry that an AGENT can influence, because it comes off
the agent's own screen. Two things keep it honest, and both belong to whoever adds a pattern
for a new CLI. watchingLabel() in session-activity.ts searches only the last few
non-blank rows, which should be the part of the screen the CLI draws rather than the agent,
and the pattern should anchor on chrome only that CLI can produce. Keep the window as small
as the layout allows, since every row it adds is another row the agent may be able to write.
The label is also ANSI-stripped and length-capped at the source, and every interpolation of
it into markup goes through escapeHtml(), since it ends up on a badge and in an approval
card.
The two shipped entries do not sit equally well behind that rule, and the difference decides
what a pattern is allowed to do. Claude's chip is the last row, so its one-row window holds
nothing the agent can write — not even the status line above it, whose command a session
running with permissions bypassed can write into its own .claude/settings.json. Codex's row
shares its slot with the last row of the transcript whenever no terminal is running, so a
message ending in that exact line is matched. What keeps that harmless is hooks: 'none': no
hook event from a codex session reaches the approvals inbox, so a forged label costs a wrong
badge and cannot silence an alert. Before giving a CLI both hook signals and a pattern, make
sure its row is one the agent cannot write.
Three capabilities that must stay independent
external, hooks and altScreen describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. shell has no hooks but is not an external CLI, so a hooks predicate written as !isExternalCliMode() accepted until=stop on a shell session and then blocked the caller for their entire timeout. deepseek is the mirror image: it IS external and it DOES have hooks.
test/cli-capability-predicates.test.ts asserts that no two of the three are equivalent across the catalog, so collapsing them fails the build rather than a user's session.
Arg-template safety
The composed command line is interpolated into bash -c "…" inside tmux, which makes command construction a security boundary. Four independent layers keep config out of it:
- Config contains no shell text. There is no
command: "..."field anywhere in the schema. An entry declares a sequence of typed tokens;argv.tsis the only place that turns them into a string, and it owns every separator itself — one space between tokens,||between fallback variants. Neither can originate from config, because config has no field that could hold either. - Every literal is validated at LOAD time against a safe-word pattern (no space, quote, backtick,
$,;,&,|, redirection, parens, braces, newline or backslash). A bad literal rejects the whole entry rather than being dropped, because a silently dropped flag would change security-relevant behaviour — losing--no-approveis not a cosmetic difference. - Values resolve through NAMED patterns. A value placeholder selects a
TokenPattern(model,uuid,slug,path-segment,tool-list, …) frompatterns.ts; config can never supply its own regex for a value, so aclis.jsonstructurally cannot widen its own validation. A value that fails its pattern drops the whole argument, exactly as the hand-written builders did: an invalid--modelomits--model, it never substitutes something else. - Escaping is independent of validation.
renderToken()re-checks the resolved value before emitting it unquoted, and single-quotes anything else — so even a value that somehow bypassed validation is quoted, never concatenated raw.
The only config-supplied regexes are discovery.version.regex and discovery.identity.regex. Both run against command output rather than a shell token, both are compiled through compileVersionRegex() (length cap, nested-quantifier rejection, never the g flag), and the output they see is truncated first.
Named profiles: the escape hatch
Some differences genuinely need to run code rather than be described. Those are named profiles: a capability field holds a profile NAME, and the implementation lives in one place keyed by that name — never by CLI id.
discovery.launcherProfile— for a CLI whose binary is not the agent.dshboots$DSH_HOME/profiles/<name>, so "installed" and "runnable" have different answers; the profile answers both, plus why a specifically-named target will not work. Implemented inutils/cli-launcher.ts.env.setenvProfile— per-CLI environment setup that is more than a list of keys, such as DeepSeek's status bridge.capabilities.transcript— which on-disk history reader understands this CLI (claude-jsonl,codex-rollout,deepseek-zstd,omp-jsonl,none).capabilities.echo.predictProfile— the predictive-echo model a composer needs.
The names live in profiles.ts, which is kept free of imports so schema.ts can validate a name at load time. A profile this build does not implement is a load-time error naming the field, rather than a CLI that silently looks permanently uninstalled.
DeepSeek: the four assumptions it breaks
DeepSeek is worth reading before assuming an entry looks like its siblings — the schema carries four extensions because of it.
| What it breaks | How the registry expresses it |
|---|---|
dsh is a profile LAUNCHER, not the agent, so "installed" is not "runnable". |
discovery.launcherProfile + discovery.launcherTargetParam. |
Its permission switch is the DSH_PERMISSION_MODE env var, not a flag — the harness has none. |
env.configSetenv (so the ordinary privilegedParams clamp still reaches it) and capabilities.privilegedEnvKeys. |
| It is the only non-claude mode with real hook signals, and for it that is a per-SESSION question. | capabilities.hooks: 'supervised' — a third state, not a boolean. |
| Its transcript is zstd session files, one frame per write. | capabilities.transcript: 'deepseek-zstd'. |
Identity probes
discovery.identity asks the binary whether it is the program we meant, and it runs before the version probe, because a version probe cannot tell an impostor from the real thing. Debian ships an unrelated dsh (dancer's shell) that answers --version perfectly happily, and npm carries squatters for both pi and grok.
discovery.version.requireVersionMatch is the weaker companion: a binary whose version output has the wrong shape counts as ABSENT rather than present-with-unknown-version. That is what a short, generic binary name needs, and it is what keeps codeman doctor and the run mode from telling the user opposite things about the same binary — both read the same regex off the same entry.
The no-id-branching rule
test/cli-registry-no-id-branching.test.ts fails the build if a CLI id comparison appears outside the stock catalog. It builds its id list from the live catalog, blanks comment lines before scanning (comments legitimately quote the pattern to explain why a branch was removed, and blanking rather than dropping is what keeps reported line numbers pointing at the real file), and keeps an allowlist in which every entry carries its reason.
It matches four shapes, not one: mode === '<id>', mode !== '<id>', case '<id>':, and ['<id>', …].includes(mode). The first version matched === only, and that gap was not academic — the refactor it guards converted the === sites and left the negated ones, so 36 !== branches survived it, including a seven-mode chain auto-enabling Ralph under a comment asking the next person to keep it in step with a predicate by hand while the sibling code path already read the capability. A guard that sees half the shapes reports a count measured over the half it happens to catch.
The allowlist is not a formality. If a branch is about what a CLI can DO it belongs in CliCapabilities; the entries that remain are things that are not CLI-behaviour branches at all — chiefly the legacy per-mode <Mode>Config objects on POST /api/sessions, which are a fact about the public HTTP API rather than about any CLI, plus a few documented cases where mode === 'claude' is genuinely the right question (Read My Mind reads Claude's own transcript, so a capability there would be actively wrong).
test/frontend-cli-no-id-branching.test.ts is the same guard for the two frontend files the CLI registry's Run-menu consolidation touches, session-ui.js and mobile-overview.js — deliberately not the rest of src/web/public/, whose per-CLI rules stay out of scope for now (see "Fields declared for later" below). Its allowlist keys on <file>::<expression> with no line number, since a single unrelated edit to a contended file would otherwise shift every subsequent line and make every entry go stale at once, and each entry additionally carries the exact number of approved call sites — a bare key would let a brand-new branch reusing an already-approved expression land unreviewed. Its comparison shape differs from the backend guard's in one respect: the left-hand side may be any identifier, not only one named mode, id or agentType, because the review of #458 found const m = this._runMode; if (m === 'codex') slipping past the named form while the scanned file already filters with (m) => m !== 'shell'.
Two namespaces called param
launch.params keys, env.configSetenv[].fromParam and capabilities.privilegedParams[].param all name a launch param. The legacy wire field a param arrives as is a separate namespace, and launch.legacyConfigAliases is the only bridge between the two.
This matters because it is invisible when it is wrong. capabilities.privilegedParams[].param is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a name from the wrong namespace clamps nothing: no load error, no failing test, the clamp simply stops running. Codex is the entry where the two names differ (bypassApprovals as the param, dangerouslyBypassApprovals on the wire), so it is the one that catches a regression. schema.ts rejects any entry naming a param it never declared, on both configSetenv.fromParam and privilegedParams.param.
Fields declared for later
accent, capabilities.echo, capabilities.wheelForward, capabilities.keyboardAccessory and capabilities.maxFrameBytes are declared but not yet read. (shortBadge was on this list until the CLI management list in Settings started showing it.) They all describe frontend behaviour, and the frontend is deliberately untouched here: app.js, terminal-ui.js and styles.css keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
Treat those values as transcribed, not authoritative — nothing enforces that echo.policy matches _updateLocalEchoState's fallthrough, so re-measure before wiring one up. accent is the one exception: it was measured against styles.css on 2026-09-21 (method in the comment above CLAUDE in stock.ts), though nothing keeps it in step with the CSS either. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; test/cli-registry-no-id-branching.test.ts pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
overlays.credStore is in the same category, for a sharper reason: the Docker credential-seeding path still reads its own CRED_STORES table, because this shape allows ONE store per CLI and the live table needs two for gemini (.gemini for the CLI's own auth plus .config/gcloud for Vertex), while deepseek declares none here even though .dsh is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is simultaneously the worst thing here to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest.
Everything else in the interface is live, including overlays.remote / overlays.docker, which back defaultRemoteCommandForMode() and defaultDockerCommandForMode() directly. Those two used to be hardcoded Record<…CommandMode, string> tables duplicating the registry with nothing keeping the two in step; test/location-overlay-commands.test.ts pins every resulting command as a literal string.
Consumers outside the server
Two things need the catalogue but cannot import TypeScript, so npm run generate:cli-catalog
(scripts/generate-cli-catalog.mts) emits two artifacts from stock.ts. Both are committed,
and test/cli-catalog-sync.test.ts fails if either drifts from a fresh generation.
| Artifact | Consumer | Why it exists |
|---|---|---|
config/clis.stock.json |
scripts/lib/cli-catalog.mjs (Docker build args), tests |
A .mjs cannot import the registry. |
a marked block inside install.sh |
the installer itself | It runs via curl | bash before any checkout exists, so it can read neither. |
Only id, label, shortBadge, enabled, order, kind and discovery are exported.
launch, env, capabilities and overlays are spawn-time concerns the server alone
interprets, and a test asserts they never leak into the artifact — a second reading of the
launch model in a consumer that cannot be tested against a real spawn is exactly what this
registry exists to prevent.
The install.sh copy is embedded, not fetched, and is the FULL catalogue. An earlier design
fetched it and fell back to a hardcoded two-CLI list, which degraded silently on an empty
response; there is no degraded mode to fall into now, and no network fetch either — a curl | bash from master already carries a catalogue exactly as fresh as the script itself, so there is
nothing a refresh would buy that isn't already true. An earlier draft added an opt-in refresh
with a TRUSTED/DISPLAY array split to keep it from ever writing the executed command; it was
dropped before merge rather than shipped half-verified — the split's only actual write was the
label, DISPLAY never diverged from TRUSTED in practice, and the added surface (a second
array, a fetch path, three failure shapes to warn on) bought nothing the embedded copy didn't
already have.
The install-command trust boundary
Three rules, and the middle one is why the embed matters:
- The server never executes an entry's
install.command. Unchanged, and still enforced by nothing executing it: the field is display text (CliDiscovery.install.command). install.shexecutes only commands embedded in itself. Those arrive in the same file, over the same TLS fetch, in the same commit as thecurl \| bashline that fetched the script — identical trust to the hardcoded vendor one-liners it replaces.- Nothing fetched at install time is ever executed. There is no second code path that fetches anything after the script itself has been fetched.
That is mechanical rather than a promise. CLI_INSTALL_CMD_TRUSTED is written only from the
generated block and is the only array the installer ever runs or displays — there is no second
array a refresh could rewrite, because there is no refresh. test/cli-catalog-sync.test.ts
asserts that the embedded commands are exactly the registry's, and
test/install-sh-invariants.test.ts that nothing in install.sh evals.
bash 3.2
macOS ships bash 3.2 and the documented install is curl -fsSL <url> | bash under
set -euo pipefail, so a bash-4 construct is not a warning there — it kills the install. The
generated block therefore uses parallel indexed arrays with offset/length windows into one
flat array instead of delimiters (a $HOME containing a space needs no IFS handling, and an
entry with nothing to contribute gets length 0 and is never iterated). CI runs bash -n and
executes the script inside a real bash:3.2 container, because the empty-window case is a
runtime set -u abort that bash -n cannot see.
Resolve at call time, never at import
Anything reading the registry must resolve it when it is asked, not when its module is first imported. sessionModeSchema(), allowedEnvPrefixes(), dependencyRegistry() and each resolver's searchDirs thunk all re-read the catalog per call.
A module-level const freezes at first import, and the failure is asymmetric: a CLI enabled while the server is running moved the run menu but not the frozen surface, so validation rejected a mode the menu offered, or codeman doctor reported a catalog nobody had any more.
Adding a CLI
- Add a
CliEntrytostock.ts. - Run
npm run generate:cli-catalogand commit both artifacts (config/clis.stock.jsonandinstall.sh). The installer's detection, its install menu, its reminder text and the Docker agent image all follow from that one step — this is what makes upstreamb6d0f1fa("wire OMP into install.sh's CLI detection, it had none") impossible rather than merely fixed. - Add a golden spawn-command pin to
test/cli-registry-spawn-golden.test.ts, a row totest/cli-capability-predicates.test.ts, its remote/docker commands totest/location-overlay-commands.test.ts, and its search paths totest/install-sh-detection-parity.test.ts. - Only if it cannot install with a plain
npm install -g <pkg>: give it a layer indocker/agent.Dockerfileand setdiscovery.install.agentImageLayer: { kind: 'dedicated', reason }on its entry instock.ts.test/docker-agent-image-coverage.test.tsrequires both, so an exclusion cannot quietly become an omission. An entry with nonpmPackageneeds only the Dockerfile layer, since it never enters the shared npm layer in the first place. - That is usually all. If you find yourself wanting to add an
ifsomewhere, the guard test will tell you — and the answer is a capability field, or a named profile if it genuinely needs to run code.
See also
- Agent CLIs — the user-facing per-CLI guide.
docs/architecture-invariants.md— the mechanics and the history behind the rules above.docs/deepseek-integration.md— why DeepSeek is shaped the way it is.