mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 20:49:41 +02:00
1d85909a0667d6d1c0ac9102403deaeadc96c2e0
15
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0a52a99ca9 |
feat(cli-registry): CLI management write API + Settings UI (Phases 1-6) (#476)
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (
|
||
|
|
10f87428c3 |
fix(cli-registry): merge-time fixes for the run-button accents (#463)
- mobile.css: gemini and antigravity run/gear rules get `!important` like pi/omp/grok/deepseek, so the gear half no longer keeps the skin accent while the body takes the mode colour (two-tone button on the default skin). - test/skin-themes.test.ts: static guard that every run mode with a base `.btn-toolbar.btn-run.mode-<id>` rule also has a resting rule inside the `html:not([data-skin="og"])` block; ids are derived from the stylesheet. - stock.ts: grok's accent comment names zinc-300 (border/badge colour); gemini's accent is #8ab4f8 to match its tab badge and run-mode dot, noted as the one exception to the border-colour method. - types.ts: "(below)" -> "(above)". - docs/cli-registry.md, CLAUDE.md: `accent` is now measured, not transcribed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
05c788ce9d |
fix(watching): close the review findings on the label and its window
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief) found the trust boundary weaker than the comments around it claimed. Eleven findings, all applied. The two blockers were both about who can write the row the label is read from. Claude's window covered two rows, and the second one is the status line, whose command a session running with permissions bypassed can write into its own `.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of its own and silence its own idle alert. The default window is one row now, which is the footer and nothing else, and the constant says why. Separately, the label reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML, which is an injection sink for any config-supplied pattern whose capture group is permissive; it goes through escapeHtml() like every other untrusted string in that file. The Codex entry could not be fixed the same way, and now says so. Its row is third from the bottom only while a terminal runs; with none running that slot holds the last row of the transcript, so matching the complete row (with the `/stop to close` tail, window narrowed to three) raises the bar without closing it. What contains it is `hooks: 'none'`: no hook event from a codex session reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an alert. The registry comment, `docs/cli-registry.md` and the test all state that rather than claiming a guarantee the code does not have. Also from the review: the TUI header badge no longer counts an acknowledged item, which was the same gate the classifier fix already went through and was wrong for human acknowledgement too; the TUI approval card reads the quiet reason and drops to a new `info` tone instead of asking for a reply; the badge carries an aria-label, because the phone it was built for has no hover target; the schema refuses `watchingLines` without a `watchingLine`; and the pattern and its window are resolved together rather than one memoized and one not. Documentation moved with it. The mechanism now lives in `docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer, `docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped buzzing, and both that page and the changeset name the limitation neither did before: a question asked in plain prose is not a dialog, so it is silenced along with the false alarms while background work runs. Verified live again after the narrowing, on an isolated beta: a Claude session reported `1 monitor` and took its idle prompt acknowledged, and a Codex session reported `1 background terminal` against the full-row anchor. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
64c288a683 |
feat(codex): read Codex's own background-terminal row
Codex states background work too, and it says so in a different place. Claude writes `· 1 monitor ·` on the last row of the screen; Codex pins `1 background terminal running · /ps to view · /stop to close` ABOVE its composer, which puts that row third from the bottom once the status line and the composer are counted. So how far up the screen to look is now per-CLI data as well: `capabilities.workDetect.watchingLines`, bounded to 1..8 by the schema, and defaulting to Claude's two. That bound is the point. The window is half the injection guard, since every row it adds is another row the agent itself may be able to write, and the label is what silences an idle alert. The other half is the anchor, and Codex's is ` · /ps to view`: chrome naming a slash command only the CLI can offer, so a session that writes "I left 1 background terminal running for you" into its own output matches nothing. Measured against a live codex-cli 0.154.0 pane rather than read out of a binary. The row appears when the terminal starts, follows the composer down as the conversation grows, and is gone after `/stop`. Verified end to end on an isolated beta: the session payload carried `watching: "1 background terminal"` and the badge rendered with it, and both cleared when the terminal stopped. The fixtures in the tests are that capture verbatim. Codex has no hook signals, so no idle prompt and no false NEEDS YOU row: for a Codex session this is the badge alone, which is the case the maintainer said a registry field could cover and a hook never could. Cross-CLI tests pin that neither pattern fires on the other's screen, and that a CLI declaring nothing still reports nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
74884a20eb |
feat(approvals): let a watching session keep quiet, and fix the tui gate
The badge alone left the row in NEEDS YOU, which is the thing the issue was about. The fix is the alert that does not fire. An idle prompt from a session that is watching its own background work now opens ALREADY acknowledged. `hook-event-routes` passes `Session.watching` to `notePrompt()`, which sets `acknowledgedAt` and records why in a new `acknowledgedReason`. Nothing new suppresses anything: `acknowledge()` has always meant "the alert this prompt armed is spent", and the prompt itself stays pending, answerable and available as Read My Mind context. A wrong label therefore costs a card that does not blink, never an alert that was never created. Every surface follows from that. The broadcast carries the reason, so a live page declines to arm the tab alert and raises no desktop notification. The push is skipped, since a false alarm is hardest to ignore on a phone. A reloading page reads `acknowledgedAt` in `seedApprovals()`, which it already did. And `classifySession()` now reads it too, which is a pre-existing bug fixed here: acknowledging on one device cleared the alert everywhere except `codeman tui`. It re-arms for free, because the next idle prompt supersedes the item and is built fresh. Only `idle` is eligible, so a dialog that blocks the agent still goes red whatever else it started. The label is pane-derived and therefore prompt-injectable, so it is now read from the last two rows of the screen only, with Claude's pattern anchored on the `·` its footer joins items with, ANSI-stripped and length-capped at the source. An agent that prints `· 1 monitor ·` into its own output finds no match. Verified on an isolated beta: a session that armed a monitor took its idle prompt acknowledged with no alert on any surface, wore the badge, and showed "quiet, watching 1 monitor" on its still-answerable card; the same session with the monitor killed alerted normally on the next prompt. `test/watching-no-alert.test.ts` pins both directions across all four surfaces. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3f2cde2db7 |
feat(session): say when a session is watching its own background work
An agent that arms a monitor, backgrounds a shell or hands a task to a cloud session is told to end its turn. The pane then falls quiet, Claude Code's idle_prompt notification arrives a minute later, and every surface files the session under NEEDS YOU with nothing for a human to answer. Claude states what it is still running on the last row of its screen (`⏵⏵ bypass permissions on · 1 monitor · ← for agents`). That row is now `capabilities.workDetect.watchingLine` in the CLI registry, guarded by compileVersionRegex() like every other config regex, and the idle probe reads it off the capture it already takes: `watchingLabel()` in session-activity.ts searches the last five lines only, so a session that PRINTS "1 monitor" is not mistaken for one running it. The label lands on Session.watching and rides toLightDetailedState() out to every surface. The phone overview, the desktop home rail and the rich sidebar rows wear it as a `watching` badge in the accent colour, beside the state pill and never in place of it: an agent can arm a monitor and ask a question in the same breath, and only the pill says which. Verified end to end against a throwaway session on an isolated beta instance: the payload carried `watching: "1 monitor"` once the turn ended, the badge rendered next to a yellow `waiting` pill, and both cleared when the monitor died. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0f955327b2 |
fix(cli-registry): merge-time fixes for the run-menu consolidation (#458)
- test/opencode-resize.test.ts: retarget the launcher guard at the real code (this.selectSession(firstSessionId), any this.activeSessionId assignment) with an anti-vacuity check; the old strings existed nowhere, so it could never fail - session-ui.js: restore as comments the two invariants the merged bodies lost (deepseek leaves statusReporting unset, i.e. ON; no effort field for external CLIs, it is Claude-specific) - docs/cli-registry.md: move the frontend-guard paragraph below the two backend-guard paragraphs so they keep their antecedent, and note the widened comparison shape - test/frontend-cli-no-id-branching.test.ts: the comparison shape accepts any left-hand identifier (const m = this._runMode; m === 'codex' was invisible), normalized to `mode`; the two `m !== 'shell'` display filters are allowlisted and the remaining blind spots documented - test/run-mode-dispatch.test.ts: table-driven pin of run() dispatch (claude to runClaude, each RUN_MODE_LAUNCH id to _runCliMode(id), shell to runShell, unknown to runClaude, lock held and released) - CLAUDE.md: name the second CI-gated guard next to the backend one - server.ts: every </head> injection passes a replacer function; a clis.json label containing $' re-injected the rest of the document past escapeScriptJson (two render tests pin it, proven failing on the string form) - _isAltCliMode(): no reference anywhere in the tree, nothing to fix Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit 1ea363ff808a62861559bc141e724b163cc1c56e) |
||
|
|
d9e6ebb20a |
fix(cli-registry): address round-2 review on #458 — count-based allowlist, RUN_MODE_LAUNCH drift guard
Three of Ark0N's four "will take at merge" items, applied instead since they were straightforward to do properly: 1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on <file>::<expression> (fixed last round) closed the line-shift problem but opened a new one: every stock id was already allowlisted for session-ui.js in the `mode === '<id>'` form, so a BRAND NEW branch reusing that exact expression anywhere in the file passed unnoticed. Reproduced live (`if (this.mode === 'codex')` injected into runOpenCode()) — stayed green under the old version. Each allowlist entry now carries the exact count of approved call sites, and a new test asserts actual-vs-declared count for every key; a mismatch in either direction is real (higher = new unreviewed branch riding in on an existing approval, lower = a reviewed site was removed and the entry is now stale). Reproduced again against the fix: same injection now fails with an exact diagnostic (expected 2, found 3). 2. Added test/run-mode-launch-table-drift.test.ts. RUN_MODE_LAUNCH restates four things stock.ts already owns (label, install command, supportsCustomModel, the external-mode key set), and they agree today with nothing enforcing it. supportsCustomModel is the dangerous one: the Run-menu picker's rows come from the server-injected window.__codemanCustomModelClis (built from capabilities.customModelInjection.kind), so a CLI gaining a real injection recipe later would be OFFERED in the picker while _runCliMode silently drops the customModel field for it — the session launches on the vendor's cloud while the UI claims the local endpoint. Drives the real session-ui.js via JSDOM and compares RUN_MODE_LAUNCH against STOCK_CLIS on all four axes. 3. Inlined the "Open Question 7 in PR-B2.md" references in the allowlist reasons — PR-B2.md is a local planning doc, never part of the committed tree, so the reference was dead on arrival for anyone reading the repo. Points at the PR #458 review thread instead. 4. Added a sentence to docs/cli-registry.md naming the new frontend guard alongside the backend one it mirrors. Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format/ check:frontend-syntax all clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
a5cf1f6005 |
docs(cli-registry): name the real tests and fields the catalogue docs point at
Three instructions a future contributor would follow literally were stale after the last review round: the "Adding a CLI" checklist sent the agent-image reason to AGENT_IMAGE_SPECIAL_CASES, a constant that no longer exists (it is discovery.install.agentImageLayer on the entry in stock.ts), the trust-boundary paragraph credited the embedded-commands pin to the invariants test when it is test/cli-catalog-sync.test.ts, and install.sh claimed "the parity test" pinned the DeepSeek Harness banner when no test did. That pin now exists: the invariants test asserts the script's grep literal and the registry's discovery.identity.regex agree on "DeepSeek Harness", and the comment names it. docs/docker-cases.md separated the two reasons a CLI stays out of the shared npm layer (no npmPackage at all versus an agentImageLayer entry), which it had folded into one, and architecture-invariants no longer lists the agent image's CLI set by hand. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
a0628a40e8 |
fix(cli-registry): address maintainer review on #380
Rebased onto current master (the one real conflict was the import line
in docker-hosts.ts Ark0N flagged; kept both), then addressed every
point from the review:
**1. Rebase.** Done — this branch now sits on current upstream/master.
**2. Agent-image special cases are data now, not an id-keyed table
outside stock.ts.** `AGENT_IMAGE_SPECIAL_CASE_IDS`/`AGENT_IMAGE_SPECIAL_CASES`
are gone. `CliDiscovery.install.agentImageLayer?: { kind: 'dedicated';
reason: string }` is a field on the registry entry itself (pi,
deepseek), `reason` is required by schema.ts, both producers
(docker-hosts.ts and cli-catalog.mjs) filter on its presence instead
of an id, and the coverage test reads it from the generated catalogue.
Also added the npm-package-name validation to the TS producer, which
only the .mjs one had — same SAFE_PACKAGE regex, duplicated
(necessarily, one side can't import the other) and now pinned
byte-identical by a new parity test.
**3. Changeset said five, it's eight.** (Not nine — see the DeepSeek
point below, which changes the true count.) Reworded to state it
structurally rather than pin a number that will go stale again.
Then the four behavior-changing findings:
- **DeepSeek was offered as a normal install option but can't actually
drive a pane.** `npm install -g @deepseek-ai/dsh` installs the
launcher only; DeepSeek ships no profile that can run standalone.
The generator now emits an empty install command for any
`launcherProfile` entry, so install.sh's menu (which requires a
non-empty command) skips it and falls through to its docs URL hint
instead — matching what the old hand-written code did before this
PR replaced it.
- **wget-only hosts lost every automatic install, including the npm
ones that never needed curl.** The menu-building loop now filters
PER ENTRY (only a command starting with `curl ` is held back) rather
than wiping the whole menu when DOWNLOADER != curl.
- **The DISPLAY/TRUSTED split and the catalogue refresh didn't hold up
under review** (refresh's only real write was the label; it ran
before the Node existence check; its own eval-detection test was
tripped by the word "eval'd" in a comment). Dropped entirely per
your own recommendation — embedded catalogue only, no network
fetch, no second array. install-sh-invariants.test.ts now asserts
the refresh/DISPLAY machinery does not exist rather than testing its
internals.
The three take-or-leave items, applied:
- `dsh_banner_probe`'s bash 3.2 empty-array bug: `${runner[@]}` →
`${runner[@]+"${runner[@]}"}`. Verified live in a real `bash:3.2.57`
container with `timeout` removed from PATH — crashed before, clean
now, full `detect_all_clis` path exercised end to end.
- `docker-agent-image-coverage.test.ts` now anchors on each layer's
`<binary> --version` proof line instead of `Dockerfile.includes(binary)`,
which stayed true if a layer were deleted but its comment survived.
- Doc drift: docs/docker-cases.md (four → five, and now describes the
data field), docker/agent.Dockerfile's "other four CLIs" comment (no
longer a magic number — CLI_NPM_PACKAGES is generated and can grow),
CLAUDE.md's install.sh size (104KB → ~112KB) and its stale mention of
the now-dropped refresh.
Verified: tsc clean, prettier clean, the full targeted suite (142
tests across the 8 affected files) green, and the full `npm test` gate
diffed BY TEST NAME against a clean upstream/master baseline run on
this same machine — identical 201-name failure set both sides (168
tests / 67 files, all pre-existing Windows-environment noise: symlinks,
PTY spawning, POSIX permission bits — none of it touching anything
this PR changes), zero new failures either side of the diff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
|
||
|
|
c5c015d648 |
docs(cli-registry): document the catalogue's consumers and the trust boundary
Adds a "Consumers outside the server" section covering the two generated artifacts, why each exists (neither install.sh nor a .mjs can import TypeScript), what is deliberately NOT exported and why, the three-rule install command trust boundary, and the bash 3.2 constraint with the offset/length window shape it forces. The adding-a-CLI checklist gains the regenerate step, since forgetting it is how the installer would keep detecting the old set while the server offers the new one — the drift this change removes, one level out. docs/docker-cases.md gains how CLI_NPM_PACKAGES is derived, why it reads the stock catalogue and not the merged registry, and a table of the four documented Dockerfile special cases with their reasons. CLAUDE.md gains a command row and names the generated block, the bash 3.2 rule and the trust boundary in its install.sh paragraph. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12 |
||
|
|
f1b7283393 |
fix(cli-registry): guard workDetect.workingLine like every other config regex
#385 made the composer glyph and the working status line per-CLI registry data, which is right, but `workingLine` arrived as a config-supplied regex validated with a bare `new RegExp()`. That skips `compileVersionRegex()`, the helper the registry uses for exactly this: a `~/.codeman/clis.json` override can set the field, the compiled pattern is run against every accumulated PTY chunk and every pane capture, and a nested quantifier there backtracks on the event loop for the whole server rather than one session. Route it through the helper in both places, which are not redundant: the schema refine rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` compiles through the same helper so the runtime cannot hold a pattern the schema would have refused. The helper returns null instead of throwing, so the Claude-pattern fallback stops being a try/catch and becomes structural. Both shipped patterns compile unchanged, and Claude's is behaviourally identical to CLAUDE_WORKING_LINE_PATTERN. Also match the Codex footer case-insensitively on the E. It was characterised against codex-cli 0.152.1, which prints a lowercase `esc`; a version capitalising it would make the whole fix silently inert, since the pane would simply never look like it was working. Docs: CLAUDE.md, architecture-invariants and cli-registry.md all still stated the Claude-mode-only rule this PR retires, and none of them named the new capability or the regex guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a81e87f440 |
fix(cli-registry): log why clis.json was ignored, and say 0600 when that is the rule
The loader refuses a `clis.json` with any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored. That is a defensible posture for a file that chooses the binaries Codeman spawns, but two things around it made the override feature look dead: the warning said "group/world-writable", which a 0644 file is not, and `LoadResult.warnings` was returned to a caller nobody wired up, so nothing anywhere printed it. A user following the docs got silence. The message now names the rule and the command that satisfies it, the loader logs every warning once on first load (the result is memoized, so once per process), the module header stops claiming that nothing ever writes (the quarantine rename of a malformed file is a write, on first use) and the registry doc gains a short section on the override file with the 0600 requirement in it. Whether the check should relax to writable bits only is a separate decision; this keeps the shipped behaviour and makes it visible. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu |
||
|
|
c5b84fb5f4 |
docs(cli-registry): annotate overlays.credStore as declared-for-later
Review item 4 named THREE live tables duplicating registry data. Two are now read from the entry (`defaultRemoteCommandForMode`, `defaultDockerCommandForMode`); the third, `resolveDockerCredentialArtifacts`, is not — and it was left neither wired nor annotated, which is the state that item explicitly rules out. It is not wired because the shape cannot express the live table: `credStore` is ONE store per CLI, and `CRED_STORES` needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek's entry declares none at all even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is at once the worst thing in that file to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest. It belongs in its own change, measured against a real container. So it is annotated instead, at the field, in the type's declared-for-later header, in docs/cli-registry.md, and in the pinned DECLARED_FOR_LATER list — the last of which means wiring it later makes a test fail rather than leaving a stale comment behind. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
4830e662f9 |
refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. Code that used to ask "which CLI is this?" reads the entry instead. Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every spawn command as a literal string, captured from the hand-written builders before they were deleted, and `test/location-overlay-commands.test.ts` does the same for all 20 remote and in-container pane commands. Config can never contain shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only in this release. OMP is included as a registry entry rather than a tenth hand-written builder, so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of `buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen and doctor ladders all drop out. Guard rails: - `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id branching reappears outside `stock.ts`, in any of its four shapes (`===`, `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the negated forms, which is how 36 of them survived an earlier pass. Every allowlisted branch carries its reason. - `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities; deriving one from another shipped the `until=stop`-hangs-on-shell bug. - `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config` wire field is separate, bridged only by `legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. - Registry data resolves AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks). A module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. - Six fields are annotated DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/ `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured. A test pins the list so it cannot quietly grow. Three user-visible changes, all deliberate and named: - `probeDockerCliVersion()` derives the in-container binary from the registry rather than assuming it equals the mode name (`antigravity` runs `agy`). - The remote CLI version probe now covers grok and deepseek, which the hardcoded map it replaces omitted while its own comment said the rule was "every mode except shell". - `codeman doctor`'s CLI rows are generated from the entries, so Claude's install hint is the install command rather than a docs URL, five CLIs gain hints they never had, and the row order follows the catalog. Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars (matching the `cliId` pattern) before its failure message quotes the value back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading the hand-editable `clis.json`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |