App Settings now refuses a Default Codex model that the server would reject,
before anything is written to localStorage. SettingsUpdateSchema is .strict()
and checks codexModel with ^[a-zA-Z0-9._\-/]*$, so a value like gpt-oss:20b
400'd the whole settings PUT while the toast still said "Settings saved", and
because the bad value was already in the local blob every later save from that
device failed the same way. The client check uses the same pattern, shows an
error toast, focuses the field and keeps the modal open. The toast has a zh-CN
translation in i18n.js.
src/web/codex-launch-defaults.ts gets an @fileoverview (fill only unset fields,
re-validate persisted values, callers decide scope, never writes Codex config
files), as every module in src carries one.
Both new Codex rows in index.html carry has-field, like every other App
Settings field row, so on phones the input and the select stack under their
label instead of squeezing it into a narrow column.
The Agent CLIs wiki paragraph said the defaults apply to every local launch.
Scheduled (cron) codex jobs are built without a codexConfig and never get
them, while Resume goes through POST /api/sessions and does, so the sentence
now names the Run menu, Resume, POST /api/sessions and /api/quick-start, and
says cron jobs do not use them.
The Settings Reference lists the two new rows in the Agents & CLIs table. The
neighbouring "Bypass approvals and sandbox" row described Pi's project trust;
it is the Codex --dangerously-bypass-approvals-and-sandbox toggle, so its note
says that now.
The PR's own changeset is removed: the release writes one consolidated
changeset at COM, and the PR's text overstated the scope (it included cron).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Design B ("one primary + chips") won a two-judge panel over a launcher grid
and a command panel. Their defect list, fixed here:
- The tunnel spinner's track was the shared translucent white and vanished on
light skins; it now follows the link's own text colour.
- A running tunnel was only green text above its QR; it now shows as a quiet
green pill, and the resting link uses --text-dim (contrast).
- A long custom CLI label could push a horizontal scrollbar into the welcome
column: labels sit in a .welcome-label span that ellipsizes (still one text
node, so i18n.js matches it), and both buttons cap at the column width.
- Chips are 40px tall on coarse pointers (tablets reach this view).
- The OG primary's colours now name the toolbar Run rule they mirror.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The welcome overview drew every CLI as its own gradient pill (one colour,
gradient and glow per CLI, plus skin overrides that re-tinted three of them),
all at the same size with the same generic play icon, wrapping into an
unaligned cluster with Cloudflare Tunnel styled as one more launcher. Nothing
said which action mattered.
Now the screen has one obvious thing to press:
- A single primary button for the first AGENT in the registry catalog (Claude
Code on a stock install), in the accent fill the toolbar's Run button uses,
with the CLI's real logo on a small light disc (a brand mark straight on the
accent turns muddy) and the same translatable "Run <label>" text.
- Every other enabled CLI as a slim, uniform pill under it, in the Compact
header's language: control-bg surface, control-border, small type, the
logo from the shared run-mode-dot slot, the bare name, and "Run <label>" as
tooltip and accessible name. Catalog order, centred, wrapping when needed.
- Cloudflare Tunnel as a quiet text link under the launchers (same id,
handler and cloudflared gate; text colour carries the active and connecting
states), still directly above the QR it reveals.
The primary is chosen by catalog order and kind, never by an id, so with
Claude disabled the next agent takes the slot and a custom CLI listed first
gets it like any other. The buttons carry no per-CLI class any more: the id
travels only as data-mode and the logo slot, and all the per-CLI welcome
gradients and their skin overrides are gone. Everything is token-driven, so
it follows every skin; OG gets its own toolbar Run blue with light text
because its --accent-ink is an unused placeholder. Focus rings sit offset
from the fills, hover motion is off under prefers-reduced-motion, and the
phone rules (reached only with the phone overview off) span the primary and
keep the chips wrapping.
Tests: run-mode-ui pins catalog order across primary and chips, the
kind-based primary choice (Claude off, shell listed first, custom agent
first, shell only), the logo slot and absence of id classes, the token-only
CSS with no .welcome-btn rule left, and the tunnel link's place between the
launchers and the QR.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The owner's picks after testing the 1.36.0 beta:
- Header Stats Style defaults to Compact (two pills with rings) instead of
Tiles. The resolver, the pre-paint stamp and the App Settings option all
agree; an unknown value now reads as compact.
- The Tiles header button is ON by default everywhere but handhelds (their
defaults object keeps it off, and the button still needs a 1180px window).
The Ctrl+Shift+G gate in tileShortcutFor() now resolves an absent key
through the device defaults too, so the chord and the button cannot
disagree; before, it required a stored true.
- Tab Layout defaults to Classic (the single strip, as before). By state, By
case and Ledger stay available as opt-ins. The tile-grid beta the owner used
last had only this layout, and it is the one they wanted back.
zh-CN option labels follow the new "(default)" markers; docs and wiki updated.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Shortcuts tab stores overrides in the per-device localStorage blob, and
saveAppSettings() carries them over from the previous blob, but the strip
before the PUT never removed them. SettingsUpdateSchema is .strict() and does
not declare the key, so once a device had any override (even the empty {}
that Reset leaves behind) every App Settings save got a 400 and no synced
setting reached the server again, while the toast still read "Settings
saved" (_apiPut resolves on a 400). Pre-existing, but #560 now points users
at the Shortcuts tab to bind Close Session again, so it would be hit often.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
_syncCliLaunchCatalog() rebuilt window.__codemanCliCatalog from /api/clis rows, which carry no capabilities, so after App Settings loaded the CLI list isExternalCliSession('claude') fell back to the kind check and Respawn and Ralph disappeared again until a reload. The rebuild now carries external over from the served catalog; a newly created custom CLI has no previous entry and falls back to kind, which is right since custom entries are external. Tests pin the resync and the openSessionOptions() call site, and the /api/clis comment names the page catalog as the deliberate capabilities exception.
The tunnel Upload URL page (upload.html) has been broken since the response envelope landed in 458fb81c: it reads j.filename and j.files while the server answers { success, data: { filename } } and { success, data: { files } }, so every upload reported "Saved: undefined" and the recent list stayed empty. Nothing else reads ~/.codeman/screenshots/, and handing a file to an agent goes through POST /api/sessions/:id/paste-image into the session's own workspace, so the page, its Settings row, the suffix branch of the tunnel row helper (now folded into its one caller) and the Upload URL i18n key go.
The three /api/screenshots routes keep working unchanged and log one deprecation warning per process on first use, naming paste-image as the replacement. Per docs/versioning-policy.md they are removed in a later MAJOR, after at least one MINOR release that carries the warning; the docs, CLAUDE.md and the multi-user plan say so.
Static caching: with upload.html gone no HTML is served by @fastify/static any more (every page has its own no-cache route; on a built tree only the precompressed index.html.gz artifact is reachable, as application/gzip, and nothing requests it). The .html branch of setHeaders was therefore dead and goes with the test that fetched upload.html to reach it; a comment now says a new static HTML page needs its own route. The index.html no-cache assertion on the route stays.
showTileGridButton gets the full per-device treatment Split has: a header
chip in App Settings beside Split, its load and save lines, OFF by default
(and in the handheld defaults), a member of the displayKeys merge policy,
stripped from the settings PUT and never declared in the .strict()
SettingsUpdateSchema (sending it would 400 the whole save).
The setting also gates the Ctrl+Shift+G toggle (the applied default while the
owner's answer is pending; one line in tileShortcutFor to change): OFF, the
chord is inert and reaches the terminal like any unbound key; ON, it opens and
closes the grid where one can open. A grid that is open however it was opened
(Ctrl/Cmd+click, a dropped tab, "Open group as tiles") keeps all its chords,
the toggle that closes it included.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A Tiles button beside Split in the header, opt-in through the per-device
showTileGridButton setting (read in applyHeaderVisibilitySettings; the
settings checkbox, displayKeys membership and schema exclusion follow with
persistence) and hard-gated like Split: hidden by its --hidden marker, a JS
width check with a live media listener, a @media (max-width: 1179px) backstop
and never in a solo window. While the grid is open the button closes it and
reads as pressed.
Closed, it opens a picker: a checkbox per open session in tab order (never one
popped out to its own window; one with no PTY is offered, its tile shows the
Attach overlay), names as text, preselected with the grid this tab last left,
else the active session and an open split's two. Boxes past what the window
can fit are disabled with the count shown, and Open opens the grid on the
checked sessions, focusing the active one if checked. Escape and an outside
click close it; its close method is idempotent and the global Escape handler
calls it.
A tile's + lists the open sessions not yet tiled; picking one adds it and
focuses it (a human selection). A grid that already holds what the window fits
disables the entries. "New session in this case" waits for auto-join.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Settings (per device): Git status: max repositories (1-50, default 12) and git timeout (5-120 s, default 30, was a fixed 10). Both go to /git-status and /git-diff as maxRepos / timeout query parameters, clamped server-side (an empty value means the default). A repository whose git status fails stays in the list with the reason instead of being dropped silently, shows as '? N' in the indicator, and the truncation line now names the limit and the setting. The discovery cache is keyed by the limit.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JrzFKEdBLwVfu6ev2ZscJS
Tab Grouping becomes Tab Layout (tabArrangement: 'state' | 'case' |
'ledger' | 'classic', default 'state'), the first row of App Settings,
Appearance, Tabs, so the old and the new strip are one choice apart.
- By case (option A): each case's tabs sit in one .tab-cluster box in
first-appearance order, labelled with the case and its count, coloured
by a stable hash into the session palette. Membership is
_mobileOverviewCaseFor(), the home screens' own match. Inside a box with
company a generated w75-api-gateway reads w75; the -<case> stays in the
DOM in a .tab-name-case span only .tabs-clusters hides. The incremental
render path rebuilds only when the cluster structure key changes. The
rail and the sidebar get a labelled section per case; phones dissolve
the boxes into the chip row. Drag stays inside one box.
- Ledger (option B): CSS only on .tabs-ledger, desktop header strip: an
auto-fill column grid of equal cells in mono type with a 3px status
bar. Its markup is identical to classic's.
- State Order (tabStateOrder: 'urgent-first' | 'urgent-last'): flips the
by-state groups so needs you can be the bottom row.
- By state: the label column is measured to the widest label on screen
and the labels are right-aligned in it, instead of a fixed 92px gutter
that left short labels far from their tabs.
- Tiles: a three-row grid (label, value, bar) with pixel line-heights in
the bundled JetBrains Mono, 36px like the header. The bar used to lie
over a fixed 28px tile, and a taller system mono (SF Mono) pushed the
value into it. The WS tile's grid moved onto an inner .connection-tile
span because JS writes the indicator's display inline. Compact uses the
same font and a matched WS size.
Tests: test/tab-clusters.test.ts (new), plus the rename and the reversed
order in test/tab-triage.test.ts.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two directions from Discussion #426, each a per-device setting and the
new default.
Tab Grouping (tabGrouping: 'state' | 'none', default 'state', option C):
tabs are grouped needs you (red, plus failed sessions), waiting (yellow),
working and idle (also ended, exited panes and web tabs), most urgent on
top. The desktop header strip draws a row per group with its label and
count in a left gutter; the flat vertical rail and the sidebar draw a
section per group; tablets keep their scrolling row with inline dividers;
phones keep the chip row in group order without headings. Classification
is the home screens' own (_mobileOverviewState/_mobileOverviewExit), the
fold into four groups is pure in CodemanTabTriage (constants.js). It is
flex `order` plus aria-hidden heading/break elements reconciled in place
after both render paths, never a DOM reorder, so Alt+N, the keyboard walk
and drag keep reading tab order; a drop is refused across groups. Named
groups in the vertical rail take precedence.
Header Stats Style (headerStatsStyle: 'classic' | 'compact' | 'tiles',
default 'tiles', option G): tiles give WS, CPU, MEM and each plan window
a label-over-value tile with a bar underneath; compact is one WS/CPU/MEM
pill with sparklines plus a plan-ring pill; classic is the header as
before. Desktop only (classic below 768px and in solo windows). The
clustered styles move #connectionIndicator into #headerSystemStats and
WS stays out of a hidden System Stats pill. The extra parts are always
rendered and hidden by default in CSS, so classic is unchanged.
Tests: test/tab-triage.test.ts, test/header-stats-style.test.ts; three
source pins in test/tab-rail-order.test.ts follow renamed lines.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- The doctor now judges candidates like the run mode's resolver: the PATH
hit, then each search dir, each one version-checked on its own and
skipped on a mismatch (a wrong `pi`/`grok` on the PATH no longer hides
the real one in a search dir). A search-dir candidate must be an
absolute path to an executable regular file, so a relative dir or a
file without the x bit reads as missing, as it does in the Run menu.
`isExecutableRegularFile` is exported from cli-executable-resolver.ts
and reused rather than copied.
- Every doctor probe passes killSignal: 'SIGKILL'; a --version that
ignores SIGTERM held the probe for its full runtime (15 s vs 5 s
measured with a TERM-trapping script).
- README no longer claims parity with the Run menu or nvm prefixes.
- The Diagnostics panel marks a missing optional tool with ○, a missing
required one with ✗, as the terminal doctor does.
- expandSearchDir names its twin, expandHome() in cli-resolver.ts.
- test/doctor-cli-json.test.ts is hermetic: temp HOME, a PATH of only
`which` and `node`, and a clis.json that drops the registry's absolute
search dirs, so it never runs the machine's installed agent CLIs.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
feat(ui): git status indicator in the bottom bar, with a panel of uncommitted and unpushed work
# Conflicts:
# config/test-suites.ts
# docs/api-reference.md
The Git window shows each group's files under their folders, collapsed until clicked, with single-child folder chains merged and open folders surviving the refresh. App Settings → Bottom bar → 'Git status: group files by folder' (per device) switches back to the flat list.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JrzFKEdBLwVfu6ev2ZscJS
- doctor probes each CLI's discovery.searchDirs when which misses and runs --version on the resolved path, so a service with a minimal PATH no longer reports installed CLIs as missing
- GET /api/doctor shares one in-flight run per category
- Diagnostics group hidden from non-admins in multi-user mode (_applyDoctorAdminGate)
- 500 uses INTERNAL_ERROR; a killed child reports 'timed out after 30 s'
- browser test blocks service workers so page.route() is reliable
- wiki: Diagnostics sentence
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JrzFKEdBLwVfu6ev2ZscJS
Optional and per-device (showGitStatus, default off). GET /api/sessions/:id/git-status is
read-only and offline (no fetch, --no-optional-locks), skips remote and Docker sessions, caps its
lists, and single-flights concurrent polls. The toolbar indicator shows uncommitted files,
commits not pushed, or a check; clicking opens a draggable panel in the style of the Files window.
Which repositories: the enclosing one when there is one; otherwise every repository up to two
levels below the working directory (capped, skipping dot-folders and node_modules, never
following symlinks), each in a collapsible section, with the indicator summing them. A repository
that merely sits above the workspace and is the home folder or higher (a dotfiles repo) is
ignored. Git-supplied text is only ever written with textContent.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JrzFKEdBLwVfu6ev2ZscJS
Maintainer merge-time fixes for the MCP server sync (opt-in mcpSyncEnabled, synced, default OFF).
M1, parse errors echoed config text (secrets included) into the HTTP response and Settings:
smol-toml's TomlError carries a code frame of the offending lines and V8's JSON "Unexpected
token" errors quote source. Both catch sites now go through describeMcpSyncError(): a parse
failure is reported by line/column only ("not valid TOML (line 3, column 21)", "not valid
JSON"), an errno failure by Node's own message (code, syscall, path), the module's own
messages via a McpConfigError class, anything else as "unexpected error". Tests put a secret
on the broken line (TOML, both JSON message shapes, and a write refused at the re-parse that
would have quoted a copied server's env) and assert it is absent from the result and from the
route's response body; they fail against the old code.
M2, CODEX_HOME / CLAUDE_CONFIG_DIR / XDG_CONFIG_HOME were ignored, so a sync could create a
file the CLI never reads and report success: new optional registry field
capabilities.mcpConfig.relocation { envVar, path } (registry data, no id branch; schema
reuses the env-name and no-traversal path rules). Declared for claude (CLAUDE_CONFIG_DIR,
checked in the 2.1.289 binary), codex (CODEX_HOME), opencode (XDG_CONFIG_HOME) and gemini
(GEMINI_CLI_HOME, gemini-cli paths.ts); antigravity follows $HOME only (agy 1.1.12 has no
relocation var). Resolved from the server process env at call time: absolute moves the file,
empty means unset, anything else reports the target with the new status "skipped" plus the
reason and writes nothing. Dedupe is now by resolved file. When a caller overrides `home`
without passing `env`, process.env is not consulted, and the route tests clear those vars so
a CI runner's XDG_CONFIG_HOME can never aim a write outside the temp HOME.
M3, feature undocumented: CLAUDE.md Key Patterns paragraph (opt-in, admin-only, additive
only, backups, re-parse validation, 0600 for copied secrets, names-only responses with
position-only parse errors, capabilities.mcpConfig and relocation), a Settings-Reference row
in the wiki, and docs/cli-registry.md + docs/api-reference.md updated for relocation, the
"skipped" status and the error policy.
Nits:
- N1 Preview/Sync before Save: the UI remembers the saved value on open and says "Save
settings to turn MCP sync on first" instead of calling the routes; the 403 message also
says to turn it on and save.
- N2 non-admins in multi-user mode: _applyMcpSyncAdminGate() hides the whole MCP group, called
from applyMcpSyncVisibility() and the codeman:me event like the CLI-management gate.
- N3 scope chip says "synced".
- N4 "(1 servers)" pluralised; the unsupported list only names installed CLIs (route test
pins it with a per-test installed set).
- N5 McpSyncResult / McpSyncTargetResult moved to src/types/mcp-sync.ts (barrel export); only
the route imported them, so no churn.
Verified with an isolated instance (throwaway HOME, own instance and tmux socket) and
Playwright: chip, save-first message, preview rendering and the admin gate.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge-time fixes for the webhook notification channel (ntfy, Slack, Discord, generic JSON).
Minor 1, App Settings Save silently dropped webhook edits: the modal's main Save now
persists the webhook group beside the settings PUT, the same way it already saves the
model config (saveModelConfigFromSettings), but only when the group differs from what
loadWebhook() put on screen (_webhookPending), so an untouched group never re-PUTs. A
refusal (bad URL, enabled with no URL) shows a warning toast, keeps the modal open and
scrolls to the group with the pasted URL still in the box, instead of a success toast.
Send test now saves pending edits first, so it never tests the old URL while the box
shows a new one. The row says so in one line.
Minor 2, no test for the server.ts glue: new test/webhook-push-glue.test.ts drives the
private sendPushNotifications on a real (never started) WebServer with an EMPTY push
store and webhook.json in the instance data dir, delivering through the real
egress-guarded fetch to a local receiver: a permission prompt arrives with the
host-prefixed ntfy Title and body while Web Push is never called, an immediate repeat is
deduped, "response complete" is skipped under scope attention and sent under all, and a
disabled config or a non-push event sends nothing. Verified it fails when the webhook
call is moved below the "no subscriptions" return.
Minor 3, docs: webhook.json added to CLAUDE.md State Files; a Webhooks section in
docs/wiki/Notifications-And-Approvals.md (setup, what is sent, the secret URL, public
ntfy topics, local targets allowed, dedupe, instance-wide reach in multi-user mode) plus
a table row, and a line in Settings-Reference; new section 10c in
docs/security-architecture.md for the second outbound channel through the web-tab
egress guard.
Nits:
- Orphaned JSDoc: the webhook schema moved below the push schemas, so
PushSubscribeSchema has its comment back.
- Duplicated enums: WebhookUpdateSchema uses z.enum(WEBHOOK_KINDS/WEBHOOK_SCOPES), so
the schema cannot accept a kind the store would coerce away.
- describeError classifies egress refusals with isEgressBlockedError (the
CODEMAN_EGRESS_BLOCKED code anywhere in the cause chain) instead of a message regex;
tests pin a deep cause chain and that matching words alone are not a refusal.
- Markup: the URL input uses set-input, the whitespace-only line is gone, and the switch
row hints to pick a long random topic on public ntfy.sh.
- Remove a saved URL: a "Remove URL" button (shown only while a URL is saved, with a
confirm) sends { url: "", enabled: false }.
- Types placement: WEBHOOK_KINDS/SCOPES and WebhookKind/Scope/Urgency/Config/Result/Status
moved to src/types/push.ts (the IO-side WebhookMessage/Request/Fetch stay in the module).
Browser test extended: main Save persists a pending edit, a refused URL keeps the modal
open with the URL, Send test saves a newly pasted URL first, Remove URL clears it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude Code's advisor tool (code.claude.com/docs/en/advisor) lets the session's
main model consult a second, stronger model at decision points: before
committing to an approach, on a recurring error, and before declaring a task
done. Codeman can now start claude sessions with one.
- `advisorModel` field on POST /api/sessions, /api/quick-start and
/api/ralph-loop/start (fable, opus, sonnet or a full model id in those
families; haiku cannot advise and is refused). Stored on the session and
persisted, so respawn, boot restore and reboot restore keep it. Remote and
docker quick-starts refuse it, as they refuse effort.
- App Settings, Models, "Advisor" segment (Default / Sonnet / Opus / Fable),
synced as `claudeAdvisorModel`. Run, resume and the Ralph wizard send it.
Default sends nothing, leaving the CLI's own /advisor choice in charge.
- Carried as the `advisorModel` key in the launch's single --settings JSON,
merged with ultracode and the statusLine exporter, never the --advisor
flag: `claude --advisor haiku` exits 1 at launch, which would leave a dead
pane on every respawn, while the settings key degrades to no advisor. A
launch without an advisor is byte-identical to before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
capabilities.newline replaces choosing the Shift+Enter bytes in the send-key
route. Key tester shows the keydown/keypress/keyup a browser reports.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Opt-in (mcpSyncEnabled, default OFF; routes 403 until on). Review fixes:
- codex TOML read/validated with smol-toml: CRLF, inline tables and
command-less tables no longer yield a duplicate [mcp_servers.x]; the new
text is re-parsed before writing
- null-prototype tables and own-key checks; unsafe names ignored at every level
- servers switched off in their own CLI (codex/opencode/antigravity) are not copied
- only CLIs that are installed or already have a config file take part
- files receiving env/headers are left 0600; symlinked configs are written through
- one apply at a time (409), unique tmp files cleaned on failure, failed status
- routes set real HTTP status codes; api-reference section; format type single-sourced
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Adds capabilities.mcpConfig to the CLI registry (Claude, Gemini, Codex,
OpenCode), an additive src/mcp-sync.ts, GET/POST /api/mcp-sync and a
Settings > Agents & CLIs control. Never edits or removes an existing
server; backs up each file it changes; reports conflicts.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Adds claude-opus-5-5 to the App Settings model picker (base option with
data-ctx="1" plus its [1m] companion row, since Opus 5.5 has a 1M window)
and to the five task-routing selects, mirroring how Fable 5.1 was added.
Co-authored-by: Claude <noreply@anthropic.com>
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (da07b38c, db4557d9) with no
corresponding update to the plan doc itself — every checklist still read
Status: TODO and every box unchecked. Brings the doc in line with the tree:
- A new "Status as of 2026-09-22" section up top: what's actually
implemented (verified by grepping the routes/schema/UI, not just trusting
the commit messages), the availability-flag staleness bug found and fixed
in this session (commit 0c77dd0a) with its devbox verification record, and
one real outstanding gap.
- The outstanding gap: a custom CLI created via Phase 5's write API has no
way to actually be launched. The Run menu is static per-mode markup with
no consumer of window.__codemanCliCatalog, so Phase 6's own "create a
custom entry, confirm it can be launched" verify step was never actually
exercised against this. Documented with two candidate fixes, neither
started.
- Each phase's checklist flipped to [x] where confirmed present in the tree,
Status lines updated from TODO to DONE, and the two originally-open
questions (Phase 2's installed source, Phase 5's PUT endpoint shape)
marked resolved against what actually shipped.
No code changes in this commit — documentation only, so a future session
(or the one already mid-flight on a separate checkout of this same branch)
picks up accurate status instead of a stale plan.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs: add the CLI-registry deployment plan and the parked Copilot plan
Both were sitting as untracked scratch files in the master checkout,
never committed to any branch. Moving them here rather than leaving them
loose:
- DEPLOYMENT_PLAN.md is the live tracker for the CLI-registry follow-up
series (PR A #347 merged, PR B #380 merged, PR B2 merged as #458) and
is where PR C (this branch's own CLI-management work) belongs.
- docs/copilot-integration-plan.md is explicitly PARKED, referenced by
name in docs/cli-enable-disable-plan.md's own header as a sibling plan
tracked separately — kept for continuity, not active on this branch.
The other scratch files found alongside these (PRA.md, PRB.md, PR-B2.md
and their review-response counterparts) described PR A/B/B2, all now
merged — deleted from the master checkout as stale rather than committed
anywhere, since their content is superseded by the real merged PRs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* fix(cli-registry): render enabled CLIs in launch surfaces
* test(cli-registry): update frontend branch guard
* fix(test): isolate suite from deployment environment
* fix(cli-registry): revise Decision 4 - claude is toggleable, shell stays permanent
shell/claude were both structurally un-disableable in the original plan
(Decision 4). Revised: shell keeps the hard backend guarantee (it is the
one non-agent mode several code paths assume always exists as a raw-
terminal fallback), but claude is now a normal toggleable entry like any
other CLI.
Safe to do because internal session creation (tmux-manager.ts, session.ts,
Ralph, plan-orchestrator) resolves a CLI via getCli(), which does not
check `enabled` at all - only the Run menu and the HTTP-facing
sessionModeSchema() (new session requests through the normal API) key off
it. Disabling claude therefore behaves identically in kind to disabling
any other CLI: no internal fallback path breaks, it just stops being
offered for new sessions until re-enabled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): hide shell's toggle entirely instead of greying it out
A permanently-disabled switch next to every other row's working toggle
read as broken rather than intentional. shell now renders no switch at
all - a plain "Always available" label - so there is nothing to click
that could look like it should work but doesn't. Backend guard is
unchanged (UNDISABLEABLE_IDS still refuses shell unconditionally); this
is UI-only.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): sort the Installed CLIs list, installed-first then alphabetical
renderCliList() previously rendered in registry order (each entry's fixed
order field). Now sorts installed CLIs first, then not-installed, each
group alphabetical by label - matches how a user actually scans the list
(what's ready to use, then what needs installing).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* style: prettier fixes from the master merge
* fix(cli-registry): install/edit take effect immediately, confirm before install, phone labels
Four gaps found verifying #476 against the #343 review trail:
- Installed or edited CLIs kept reading as missing/stale. Every binary lookup
(the nine per-CLI resolvers and the generic registry one) caches in its own
closure, with a negative-cache backoff of up to 5 minutes, and nothing
cleared them. invalidateCliExecutableResolvers(binaries) now drops those
caches per binary; install (success or failure), create, edit and delete
call it plus invalidateCliResolverCache(id). Before this, a CLI installed
from Settings could fail to launch for minutes, and an edited custom entry
kept launching its old binary until a restart.
- The Settings "installed" badge for a custom entry used a private `which`,
ignoring the entry's searchDirs and the login-shell lookup that spawn and
the Run menu use; it now asks the same generic resolver they do.
- Install ran on a single click. The #343 review asked for auto-install to
sit behind an explicit confirm; the confirm now names the exact command,
which GET /api/clis returns for stock entries only (installCommand).
- The phone Run button showed the two-letter tab badge ("CC", "CX") instead
of the word ("Claude", "Codex"). It uses the registry label again, which is
identical to the old static table for every stock CLI (now pinned).
14 new tests; 9 of them fail against the previous head and pass here.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): address #476 review — safe serialized writes, no id branches, docs
Must-fix:
- registry-writer: start fresh only on ENOENT; refuse (409) a clis.json that
does not parse or has group/world permission bits instead of overwriting it
(isUnsafePermissions now exported from registry.ts)
- mutateRegistryFile(): one promise chain for every mutation, with the
existence/duplicate checks inside the serialized step, plus a unique tmp
name per write
- docs: CLAUDE.md, architecture-invariants, cli-registry (new Settings
section) and api-reference (the six /api/clis routes)
- drop DEPLOYMENT_PLAN.md and docs/copilot-integration-plan.md
Smaller:
- PUT /api/clis/custom/:id keeps the entry's current enabled state when the
body omits it
- runMode setter falls back to the first enabled catalogue entry, not 'claude'
- shell guard keyed on kind === 'shell' (routes + Settings list); stock probe
map shared with server.ts via utils/cli-installed-probes.ts
- stock claude label is now 'Claude Code', so the Run menu / phone overview
label rewrites are gone (doctor row keeps "Claude CLI" via its override)
- welcome buttons are translatable again and read "Run Claude Code" /
"Run Shell"; zh-CN gains "Run Codex" / "Run OMP"
- install: per-id in-flight guard (409) and CODEMAN_* stripped from its env
- fileoverview / CliEnableSchema comments no longer say stock-only
- test-env isolation changes moved to their own PR
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* test(cli-registry): pin the #343/#347 findings #476 makes reachable
A CLI toggled or created through the routes is accepted or rejected by
CreateSessionSchema with no restart (#343 finding 2), and a custom CLI created
through the API renders a real local, remote and docker launch command
(#347 finding 5: no more `cd <path> && undefined`).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Copying a paragraph out of a Claude Code or Codex pane puts that pane's own
two-column transcript gutter on the clipboard, so every pasted line arrives
indented. #451 shipped the trailing half of the copy clean and left the leading
half out, because deriving the width from the selection fires on 73% of ordinary
indented text and cannot tell a margin from content.
The width is DECLARED rather than derived. `capabilities.transcriptGutter` on
the CLI registry is a bounded integer; claude and codex each declare 2, measured
on live panes, and no other stock entry declares any, so a CLI whose transcript
layout nobody has measured is never touched. The server publishes the map as
`window.__codemanTranscriptGutter`, built by filtering `enabledClis()` on the
capability rather than by listing ids, and `_activeCliGutterColumns()` looks the
active session's mode up in it. The copy path reads no terminal buffer at all.
The declared width is a CEILING, not the answer: `clean()` strips the lesser of
it and the run every selected line shares. A block can therefore only shift as a
unit, the structure inside a selection survives by construction, and a selection
reaching column 0 loses nothing. That is what keeps a `git log` body at its own
four-space indent inside an agent's two-column gutter.
Codex was measured separately, because it renders nothing like Claude: it draws
boxes narrower than the pane and pushes its transcript into ordinary scrollback.
On a live 0.154.0 answer its `•`/`›`/`⚠` markers sit in the gutter, prose
continuations sit at 2, and a nested YAML block the model wrote rendered at
2/4/6/8 for its own 0/2/4/6. Replayed at 100, 120, 160, 198, 235 and 282 columns
its indents were 0, 2, 4, 6 and 8 at every one, never 1. Copying that YAML out
of a live Codex pane now yields 0/2/4/6: gutter gone, nesting intact.
Two derived versions were built and measured first, and both are recorded in the
code because both looked correct:
- Painted trailing padding — a full-screen TUI writes real spaces across the
unused part of a row, a shell leaves them never-written for xterm to trim —
has no false positives and never over-stripped. It is also a function of pane
WIDTH: the padding exists only while a rendered line stops short of the CLI's
own layout width, and Claude's prose wraps to fill it. Dragging the same two
prose rows of one live transcript at five window sizes, the share of padded
rows ran 44%, 6%, 6%, 7% and 87% at 123, 160, 198, 235 and 298 columns, so the
strip silently did nothing at every ordinary size while a corpus captured
entirely at 282 columns said it worked.
- Taking the narrowest indent on the rows around the selection fires at every
width and over-strips about 1% of selections, because a file listing inside
the transcript can be the narrowest thing on screen.
Measured over 1,392,281 selections — every 1, 2, 3, 5, 10 and 20-row window of
real Claude screens replayed from live PTY streams at 100, 120, 160, 198, 235
and 282 columns — the declared width over-strips none, breaks no relative indent
and alters no text, and serves 100% of the selections whose own indent covers
the gutter. Verified end to end in a browser with a real mouse drag and a real
Ctrl+C: Claude and Codex panes paste flush at 123, 198 and 298 columns, a shell
pane is untouched at every one.
The strip sits behind `copyStripMargin` (App Settings, Selection & clipboard),
per-device and default ON: a display key, absent from the .strict()
SettingsUpdateSchema, read as `!== false` because the desktop branch of
getDefaultSettings() returns {}. The toggle is checked before the map.
Two review findings from #451, handled:
- The mid-row flag governs ONE line now. `range.start.x > 0` excludes only the
first selected line, the one whose margin the mousedown genuinely cut off, so
the same three rows no longer produce three different clipboard results.
- The reversed-drag finding does not reproduce on the pinned xterm.
`getSelectionPosition()` reads `_selectionService.selectionStart`, whose
getter returns `SelectionModel.finalSelectionStart`, and that swaps the pair
when `areSelectionValuesReversed()` says so. A real upward mouse drag through
chromium against xterm 6.0 reports the same range as the downward drag.
`_normalisedSelectionRange()` keeps the ordering as a guard, because the model
one layer down exposes the unnormalised fields under the same two names.
Tests: test/terminal-copy-clean.test.ts (64, up from 31), plus the injected
script stripped in test/server-index-title.test.ts. Every guard is pinned:
removing any one of seven reds at least one test, including declaring the wrong
gutter width. Full suite green, 7,861 passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Issue #464, "text gets muffled sometimes, in both TUI default and fullscreen".
The screenshot is not a dropped frame or a frozen renderer — it is arithmetic.
Claude Code's TUI wraps its frame at the width the PTY reported and erases the
previous frame by walking the cursor up the rows it believes that frame took. A
browser terminal of a different width makes each logical line occupy more
physical rows than Ink counted, so `eraseLines(n)` clears too few and the new
frame paints over rows nothing erased: doubled lines, and short tool summaries
sitting inside longer prose rows with the prose's tail still visible.
Reproduced against this repo's own xterm before changing anything — a 120-column
PTY against a 62-column terminal renders every wrapped line twice. `test/
terminal-pty-geometry.test.ts` pins that, and pins the clean render at matching
widths beside it, so the assertion cannot be satisfied by code that fixes
nothing.
Four ways the two drifted apart, none of them observable from either end:
1. `fitAddon.fit()` resizes xterm to `proposeDimensions()` RAW while every
server-facing path reported those floored at 40x10. Measured in Chrome at
430px: font size 44 proposed 13 columns, the server was told 40, and xterm
stayed at 13. Three call sites each did their own fit-then-floor, and two
re-read the proposal after the fit — `_shrinkPaddingToFit()` runs exactly
there, so the container had moved.
2. `throttledResize` (keyboard up) and `sendResize` (session detached into its
own window) reflowed locally and withheld only the SIGWINCH. That is the one
combination that cannot be right: a reflow nothing is rendering for buys
nothing and costs correctness. Both now withhold everything, and the
keyboard's settle timer still sends the one resize that stops the PTY going
stale.
3. `setFontSize`/`setFontFamily`/`setFontWeight` move the cell size — a geometry
change — and told the server nothing at all, so raising the font on a phone
left the CLI wrapping at the old column count.
4. `Session.resize` DECLINES a small-viewport request while a desktop connection
holds an active sizing claim, and said nothing, because resize was write-only.
`syncTerminalGeometry()` is now the one function that may change the terminal's
size: it fits, floors and applies as a single step, so the numbers xterm holds
are the numbers the server is told. A test sweeps every module for a bare
`fit()` on the main terminal, and finds exactly one — the owner's own.
For (4) the client cannot win, so it is told the truth instead: both transports
answer a resize with `session.ptyCols`/`ptyRows` (`{"t":"zc"}` on the socket,
the body of the resize POST) and `_onPtyGeometryReport` adopts them. A terminal
that keeps a shape the PTY refused does not render "too narrow", it renders
garbled. Adopting can leave the pane wider than the screen and the container is
`overflow: hidden`, so `.pty-oversized` grants horizontal reach for exactly as
long as the mismatch lasts: correct-and-reachable beats correct-and-clipped
beats garbled. That rule sets both overflow axes and its own `touch-action`
because mobile.css loads later and sets `.terminal-container { overflow:
visible; touch-action: none }` — a bare `overflow-x` would leave overflow-y
computing to `auto` and hand the browser a vertical scroll container the
terminal's touch handler knows nothing about.
Verified in Chrome at 430px against a live server, with a desktop client holding
the claim: the phone adopts 198x43, gets `overflow-x: auto` / `overflow-y:
hidden` / `touch-action: pan-x`, 758px of reach to the right, and keeps its own
vertical scrolling. The pre-fix build was measured in the same harness for the
control.
Two things this deliberately does not do. It does not change who owns the pane
size — the desktop still wins, and `_startMobileResizeRetry` still takes it back
once that goes idle. And `throttledResize` still holds the PTY's shape for the
whole keyboard animation rather than sending a SIGWINCH per step; that decision
predates this and was not re-tested here.
Also in this commit, Ark0N's third-pass review items on #431:
- The response viewer's byte-buffer fallback and `_onSessionClearTerminal` both
used the no-param `/terminal` form, capped only by `terminalBufferMaxBytes`
(32MB) — the largest body the frontend asks for anywhere. One carried no
deadline at all and the other got the 15s tail budget. Both now take the
full-history budget.
- A `?full=1` capture that outruns its deadline falls back to the bounded tail.
The pane is blanked before that fetch, so an abort used to leave a black
rectangle, discard the queued live output and never reach `_connectWs`. A
failed load now still opens the socket, says one dim line where the content
would have been, and clears the tab's spinner — which nothing did, so a failed
select left `aria-busy="true"` set forever.
- `_wsOutputGapSession` is cleared at the repaint that settles it, not in a
`finally` that also ran on the catch. A reconcile that threw, or hit the new
deadline — the flaky link the marker exists for — dropped the gap with nothing
to retry it. `ws.onopen` no longer clears it up front either.
- The replay-clear invariant is pinned in the gate, which is the drift this PR
exists to fix: `_resetTerminalForReplay` must be a queued write and nothing
else, and no module may blank the terminal with a `clear()+reset()` pair.
- `DIAG_ENTRY_MAX_CHARS` replaces the hardcoded 300, bound through a local
first: `CodemanDiag?.x` still throws a ReferenceError when the identifier was
never declared, and that is the one function in the app that must not throw.
- panels-ui's two kill-all clears route through the same helper, and the
xterm-version guard's comment says "resolved lockfile version" rather than
"dependency RANGE", which is what it has pinned since the last round.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The badge alone left the row in NEEDS YOU, which is the thing the issue was
about. The fix is the alert that does not fire.
An idle prompt from a session that is watching its own background work now
opens ALREADY acknowledged. `hook-event-routes` passes `Session.watching` to
`notePrompt()`, which sets `acknowledgedAt` and records why in a new
`acknowledgedReason`. Nothing new suppresses anything: `acknowledge()` has
always meant "the alert this prompt armed is spent", and the prompt itself
stays pending, answerable and available as Read My Mind context. A wrong label
therefore costs a card that does not blink, never an alert that was never
created.
Every surface follows from that. The broadcast carries the reason, so a live
page declines to arm the tab alert and raises no desktop notification. The push
is skipped, since a false alarm is hardest to ignore on a phone. A reloading
page reads `acknowledgedAt` in `seedApprovals()`, which it already did. And
`classifySession()` now reads it too, which is a pre-existing bug fixed here:
acknowledging on one device cleared the alert everywhere except `codeman tui`.
It re-arms for free, because the next idle prompt supersedes the item and is
built fresh. Only `idle` is eligible, so a dialog that blocks the agent still
goes red whatever else it started.
The label is pane-derived and therefore prompt-injectable, so it is now read
from the last two rows of the screen only, with Claude's pattern anchored on
the `·` its footer joins items with, ANSI-stripped and length-capped at the
source. An agent that prints `· 1 monitor ·` into its own output finds no
match.
Verified on an isolated beta: a session that armed a monitor took its idle
prompt acknowledged with no alert on any surface, wore the badge, and showed
"quiet, watching 1 monitor" on its still-answerable card; the same session with
the monitor killed alerted normally on the next prompt. `test/watching-no-alert.test.ts`
pins both directions across all four surfaces.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Blocker from Ark0N's second PR #453 pass: moving showSplitButton into
settings-ui.js's per-device displayKeys set was only half of making it
per-device. saveAppSettings() still put it in the object PUT to
/api/settings, SettingsUpdateSchema (.strict()) does not declare it,
the server answered 400 INVALID_INPUT, and because the call site never
checked res.ok the UI still reported "Settings saved" while NOTHING
persisted — workspaceHooksEnabled, agentSkillEnabled, tunnelEnabled,
claudeModel, every toggle, on every save, on every device. Strip it
out via the same destructure every other per-device key goes through
(`showSplitButton: _ssp,`), drop the stray mention from a schemas.ts
comment (a mention there reads as "this is a real field" to the next
grep), and add a static guard test mirroring
test/terminal-auto-copy.test.ts's three-way rule.
Also finishes the desktop gate the first pass only did in CSS at
599px: SPLIT_PANE_MIN_WIDTH (1180, matching HOME_SESSIONS_MIN_WIDTH)
now backs an actual JS width check in _applySplitButtonVisibility,
with a matchMedia listener so a live window resize hides/shows the
button without a reload — the CSS backstop in styles.css is the
reverse-direction guarantee for when JS hasn't run.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ark0N's PR #453 review: nothing gated this feature to desktop even
though the design called for it (two 240px min-width panes plus the
divider need ~486px, and the divider has no touch handlers), and
showSplitButton was a SYNCED setting, so turning it on at a desk also
put the button in the phone header.
- Hard-hide .btn-split on phones in mobile.css regardless of the
setting, matching the other desktop-oriented header buttons in the
same @media (max-width: 599px) block.
- Move showSplitButton into settings-ui.js's per-device displayKeys
set and drop it from SettingsUpdateSchema entirely, matching the
showFileViewerButton/skin precedent (CLAUDE.md's "per-device keys
... must NOT be added to SettingsUpdateSchema" rule) — a desktop
opt-in must never sync onto a phone that never asked for it. Removes
the now-invalid server-round-trip test for the setting.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.
Blockers:
1. Every generated inline onclick was unparseable. JSON.stringify's own
double quotes terminated the double-quoted HTML attribute at the first
one, leaving btn.onclick null on every picker entry and every Discover/
Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
argument, the same idiom deleteCase's onclick already uses four lines
away in session-ui.js. This also closes the live-HTML-injection route
through modelId (server-controlled, from the endpoint's own /v1/models
reply): with quoting intact, a `>` inside it can no longer terminate the
<button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
like every other /api route (server.ts's preSerialization hook applies
to arrays too), so Array.isArray(hosts) was always false in production
and the picker/settings panel silently saw nothing. Both call sites now
go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
returns normally without ever changing activeSessionId, so the apply
step used to silently re-point and restart whatever session the user was
already looking at. runCustomModelEntry() now snapshots activeSessionId
before the launch and requires it to have actually changed.
Majors:
4. Routes the launch through run() itself via a temporary _runMode swap
(never persisted — setRunMode() would sync it to the server) instead of
a parallel hardcoded dispatch table, so a custom-model launch now holds
the same _runInFlight lock every other Run click gets. This also
resolves the "hardcoded runners map contradicts the PR's own design"
minor: dispatch is run()'s own, so a CLI whose customModelInjection
recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
parses markup this module generated itself) for exactly the DOM-level
facts the review said needed no Playwright and no tmux: a generated
button's onclick genuinely compiles and fires, a dangerous modelId never
produces a live element, the envelope unwrap works, the session-changed
guard holds, run() actually gets called (proving the in-flight lock
engages), and _runMode is restored afterward. Confirmed against the
pre-fix code first (reproduces btn.onclick === null exactly) so this
isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
render-index-html.test.ts for the other fixes below.
Minors:
- Generated entries now filter through isCliAvailable(), matching
_refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
(applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
and to settings-modal open) instead of always rendering; the endpoint GET
no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
redactApiKey() replaces the field with a computed apiKeySet: boolean, and
a PUT with no apiKey now keeps the stored one server-side
(applyStoredApiKey()) instead of the client resending a value it was
never given. New tests cover both directions (kept vs. replaced) by
observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
(_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
event, since the real role can resolve after settings were first opened)
— endpoint writes were already admin-only server-side, but the button
used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
(CLAUDE.md already records that exact literal turning the settings
preview into a grey slab on light skins), .run-mode-custom-models gets
the same gap: 2px .run-mode-menu's own flex gap only applies one level
up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
</script> (CliEntry.label is user-clis.json-settable, unlike
__codemanCliAvailable's booleans-only payload) via a new exported
escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
docs/api-reference.md; left the "no zh-CN for the new Models-section
group" minor unaddressed only insofar as the wider Models section (task
routing, thinking effort, etc.) has never had zh-CN coverage either —
everything this PR itself introduces (labels, hints, button text, the
Run-menu's "Custom Endpoints" header) IS translated in i18n.js.
Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Follow-up to #393, picking up the work Ark0N invited in his merge comment:
"generate those entries from the saved profiles rather than a fixed
duplicate per harness, and put it in a follow-up PR so this one stays the
backend... The Run-menu picker is yours if you want it."
Adds the frontend surface the backend has been waiting on:
- Run menu: a "Custom Endpoints" section lists one entry per (harness that
supports customModelInjection, saved endpoint) pair, e.g.
"Claude Code (llama.cpp)". The harness list comes from
window.__codemanCustomModelClis, injected at page render straight off the
CLI registry's own capabilities (never a hardcoded id list in the
frontend), so a CLI whose injection recipe lands later appears with no
frontend change. Picking an entry runs that harness's own existing run*()
function unmodified (case creation, env overrides, everything, forced to
a single instance) and then applies the endpoint's default model to the
session it creates via the existing POST /api/sessions/:id/custom-model
route. Entries are hidden for a remote/docker active case, since that
route already refuses both.
- Settings: App Settings -> Models gets a "Custom model endpoints" group
wiring up the customModelEndpointsEnabled toggle (declared since #393,
read by nothing until now) plus CRUD against the existing
/api/model-endpoints routes: list, add/edit (inline form), delete,
discover models.
- Backend: CustomModelHost gains an optional defaultModelId, the model the
picker applies with no further choice per endpoint (one generated menu
entry per CLI+endpoint pair, not per CLI+endpoint+model). The route
refuses a value that isn't one of the endpoint's own discovered models,
and a fresh discovery drops a default that no longer appears rather than
carrying an invalid one forward.
Docs: docs/custom-model-endpoints.md describes the new picker and settings
panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the
"backend-only" status note and documents the picker's generation mechanism.
Tests: four new route tests cover defaultModelId validation, acceptance,
and the drop/keep behaviour across a re-discovery; a new render-index-html
test pins the __codemanCustomModelClis injection (present, agent CLIs
supporting the capability, antigravity and shell excluded) and its
solo-window skip. No browser test was added for the Run-menu picker itself
or the settings CRUD panel (this box has no tmux, so the live server used
by test:browser/test:mobile could not be exercised here) -- worth a
Playwright pass before merge, same as any other frontend PR.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Finishes #376. The contributed keystroke tracker sat on the raw byte stream
and named tabs wrong five ways (every prompt, every write path, a bare Esc
eating the next prompt's first character, pasted newlines as Enter, any CSI
clearing the draft) and replaced the whole name, which dropped the case from
the tab and reset the w<n> counter. This lands the feature with each of those
closed:
- First prompt means the first: applyAutoName() flips a placeholder to
`auto` whether or not the string changed. nameSource is now the tri-state
placeholder | auto | manual; the name setter is the only manual path.
- Only user-originated input counts: write()/writeViaMux() take
SessionWriteOptions.fromUser, set by the browser WS path and POST /input
only, so Ralph, respawn, cron, approvals and the trust-dialog keys can
never name a tab. A startMode 'shell' CLI never feeds the tracker (a
capability, not an id check); the send-key route feeds trackUserInput()
because its line feed bypasses the session.
- Prefix form `w3-case: title`: parseSessionPrefix() already renders it as
the title with the prefix in the tooltip and the next-session counter
still matches it. Composed within MAX_SESSION_NAME_LENGTH.
- Tracker rules per key: bare Esc resolves at chunk end; mouse/focus
reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R
taint the draft so Enter submits nothing rather than a fragment;
bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space;
the draft keeps its head past 8192 code points; an escape past 64 bytes
is abandoned.
- Title: slash commands by shape (a path is a prompt), `!` escapes
refused, first sentence only past 8 code points ("e.g." is not a title),
72 code points on a word boundary.
- Synced `autoNameSessions` setting, default OFF (the prompt reaches
mux-sessions.json, session:updated and /api/search), App Settings ->
Appearance -> Tabs, read fresh per prompt after the eligibility check.
Tests: test/session-auto-name.test.ts (tracker, title, composition,
ownership, emit gating), the wiring test (once, prefix, setting off,
manual protected), test/routes/session-name-routes.test.ts (PUT /name
flips to manual and persists). Verified live on an isolated instance: API
and browser-typed prompts name the tab, a second prompt does not, shells
and renamed tabs are untouched, nameSource survives a restart.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Two follow-ups to #361's sticky telemetry switch.
GET /api/settings reconciled an absent showPlanUsageLimits by persisting
true, but readJsonConfig() answers {} for ANY read failure (a parse
error, EACCES, EMFILE, a read landing inside PUT's non-atomic write), not
only ENOENT, and every page load calls this route, so one unlucky read
replaced the whole settings file with a one-key file. The route is a
plain read again and the default moved into the reader:
readPlanUsageTelemetryEnabled() treats an absent key as ON, the same way
readWorkspaceHooksEnabled() does, which is what the desktop chip already
shows for an install that never touched the setting.
saveAppSettings() sent showPlanUsageLimits on every save. The chip
defaults OFF on handhelds, so a phone saving its font size persisted
false and switched collection off for every desktop, whose chip then
went stale with no error anywhere. The key is now stripped like the
other per-device display keys and re-added only when the save FLIPS the
chip relative to what the device had (planUsageCollectionFlip), so an
explicit toggle on any device still writes it in either direction.
Tests pin both: the GET route with a mocked filesystem (absent, missing,
EACCES, garbage, explicit), the reader default, and the flip helper plus
its wiring in saveAppSettings.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>