Two directions from Discussion #426, each a per-device setting and the
new default.
Tab Grouping (tabGrouping: 'state' | 'none', default 'state', option C):
tabs are grouped needs you (red, plus failed sessions), waiting (yellow),
working and idle (also ended, exited panes and web tabs), most urgent on
top. The desktop header strip draws a row per group with its label and
count in a left gutter; the flat vertical rail and the sidebar draw a
section per group; tablets keep their scrolling row with inline dividers;
phones keep the chip row in group order without headings. Classification
is the home screens' own (_mobileOverviewState/_mobileOverviewExit), the
fold into four groups is pure in CodemanTabTriage (constants.js). It is
flex `order` plus aria-hidden heading/break elements reconciled in place
after both render paths, never a DOM reorder, so Alt+N, the keyboard walk
and drag keep reading tab order; a drop is refused across groups. Named
groups in the vertical rail take precedence.
Header Stats Style (headerStatsStyle: 'classic' | 'compact' | 'tiles',
default 'tiles', option G): tiles give WS, CPU, MEM and each plan window
a label-over-value tile with a bar underneath; compact is one WS/CPU/MEM
pill with sparklines plus a plan-ring pill; classic is the header as
before. Desktop only (classic below 768px and in solo windows). The
clustered styles move #connectionIndicator into #headerSystemStats and
WS stays out of a hidden System Stats pill. The extra parts are always
rendered and hidden by default in CSS, so classic is unchanged.
Tests: test/tab-triage.test.ts, test/header-stats-style.test.ts; three
source pins in test/tab-rail-order.test.ts follow renamed lines.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Route test hygiene: each test works in its own mkdtemp folder, every
deletion goes through safeRmHomeTree, and the suite refuses to start
outside test/setup.ts's temp HOME, so a raw `npx vitest` can no longer
delete a real ~/projects or the live linked-cases registry.
- Path policy: the symlink-resolved target is also judged against the
resolved home, data dir and system roots (home reached through a link,
macOS /etc -> /private/etc); test expectations are realpath-safe.
- Refuse a target equal to or inside the caller's or the shared cases
directory, pointing at plain Create New (it would list twice, and
deleting the local copy removes files).
- The registry re-read comment no longer claims to prevent the
lost-update race; documented as narrowing it, like /api/cases/link.
- UI: the success toast names the folder the server created, the
"under ~/codeman-cases" blurb and name hint change while a custom
folder is ticked, a "/" parent previews and sends /<name> instead of
an empty path, and the new labels have zh-CN entries.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
feat(ui): git status indicator in the bottom bar, with a panel of uncommitted and unpushed work
# Conflicts:
# config/test-suites.ts
# docs/api-reference.md
The Git window shows each group's files under their folders, collapsed until clicked, with single-child folder chains merged and open folders surviving the refresh. App Settings → Bottom bar → 'Git status: group files by folder' (per device) switches back to the flat list.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JrzFKEdBLwVfu6ev2ZscJS
Optional and per-device (showGitStatus, default off). GET /api/sessions/:id/git-status is
read-only and offline (no fetch, --no-optional-locks), skips remote and Docker sessions, caps its
lists, and single-flights concurrent polls. The toolbar indicator shows uncommitted files,
commits not pushed, or a check; clicking opens a draggable panel in the style of the Files window.
Which repositories: the enclosing one when there is one; otherwise every repository up to two
levels below the working directory (capped, skipping dot-folders and node_modules, never
following symlinks), each in a collapsible section, with the indicator summing them. A repository
that merely sits above the workspace and is the home folder or higher (a dotfiles repo) is
ignored. Git-supplied text is only ever written with textContent.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JrzFKEdBLwVfu6ev2ZscJS
POST /api/cases takes an optional path; Add Case > Create New gets a 'Create in a
custom folder' option with Browse. The folder is created (or an empty one filled),
scaffolded like a normal case and registered as a linked case. System, home,
credential and Codeman folders are refused; a folder with files is Link Existing's
job; a failure after the first write undoes what this call created. Admin only in
multi-user mode, like Link Existing.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Maintainer merge-time fixes for the MCP server sync (opt-in mcpSyncEnabled, synced, default OFF).
M1, parse errors echoed config text (secrets included) into the HTTP response and Settings:
smol-toml's TomlError carries a code frame of the offending lines and V8's JSON "Unexpected
token" errors quote source. Both catch sites now go through describeMcpSyncError(): a parse
failure is reported by line/column only ("not valid TOML (line 3, column 21)", "not valid
JSON"), an errno failure by Node's own message (code, syscall, path), the module's own
messages via a McpConfigError class, anything else as "unexpected error". Tests put a secret
on the broken line (TOML, both JSON message shapes, and a write refused at the re-parse that
would have quoted a copied server's env) and assert it is absent from the result and from the
route's response body; they fail against the old code.
M2, CODEX_HOME / CLAUDE_CONFIG_DIR / XDG_CONFIG_HOME were ignored, so a sync could create a
file the CLI never reads and report success: new optional registry field
capabilities.mcpConfig.relocation { envVar, path } (registry data, no id branch; schema
reuses the env-name and no-traversal path rules). Declared for claude (CLAUDE_CONFIG_DIR,
checked in the 2.1.289 binary), codex (CODEX_HOME), opencode (XDG_CONFIG_HOME) and gemini
(GEMINI_CLI_HOME, gemini-cli paths.ts); antigravity follows $HOME only (agy 1.1.12 has no
relocation var). Resolved from the server process env at call time: absolute moves the file,
empty means unset, anything else reports the target with the new status "skipped" plus the
reason and writes nothing. Dedupe is now by resolved file. When a caller overrides `home`
without passing `env`, process.env is not consulted, and the route tests clear those vars so
a CI runner's XDG_CONFIG_HOME can never aim a write outside the temp HOME.
M3, feature undocumented: CLAUDE.md Key Patterns paragraph (opt-in, admin-only, additive
only, backups, re-parse validation, 0600 for copied secrets, names-only responses with
position-only parse errors, capabilities.mcpConfig and relocation), a Settings-Reference row
in the wiki, and docs/cli-registry.md + docs/api-reference.md updated for relocation, the
"skipped" status and the error policy.
Nits:
- N1 Preview/Sync before Save: the UI remembers the saved value on open and says "Save
settings to turn MCP sync on first" instead of calling the routes; the 403 message also
says to turn it on and save.
- N2 non-admins in multi-user mode: _applyMcpSyncAdminGate() hides the whole MCP group, called
from applyMcpSyncVisibility() and the codeman:me event like the CLI-management gate.
- N3 scope chip says "synced".
- N4 "(1 servers)" pluralised; the unsupported list only names installed CLIs (route test
pins it with a per-test installed set).
- N5 McpSyncResult / McpSyncTargetResult moved to src/types/mcp-sync.ts (barrel export); only
the route imported them, so no churn.
Verified with an isolated instance (throwaway HOME, own instance and tmux socket) and
Playwright: chip, save-first message, preview rendering and the admin gate.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Merge-time fixes for the webhook notification channel (ntfy, Slack, Discord, generic JSON).
Minor 1, App Settings Save silently dropped webhook edits: the modal's main Save now
persists the webhook group beside the settings PUT, the same way it already saves the
model config (saveModelConfigFromSettings), but only when the group differs from what
loadWebhook() put on screen (_webhookPending), so an untouched group never re-PUTs. A
refusal (bad URL, enabled with no URL) shows a warning toast, keeps the modal open and
scrolls to the group with the pasted URL still in the box, instead of a success toast.
Send test now saves pending edits first, so it never tests the old URL while the box
shows a new one. The row says so in one line.
Minor 2, no test for the server.ts glue: new test/webhook-push-glue.test.ts drives the
private sendPushNotifications on a real (never started) WebServer with an EMPTY push
store and webhook.json in the instance data dir, delivering through the real
egress-guarded fetch to a local receiver: a permission prompt arrives with the
host-prefixed ntfy Title and body while Web Push is never called, an immediate repeat is
deduped, "response complete" is skipped under scope attention and sent under all, and a
disabled config or a non-push event sends nothing. Verified it fails when the webhook
call is moved below the "no subscriptions" return.
Minor 3, docs: webhook.json added to CLAUDE.md State Files; a Webhooks section in
docs/wiki/Notifications-And-Approvals.md (setup, what is sent, the secret URL, public
ntfy topics, local targets allowed, dedupe, instance-wide reach in multi-user mode) plus
a table row, and a line in Settings-Reference; new section 10c in
docs/security-architecture.md for the second outbound channel through the web-tab
egress guard.
Nits:
- Orphaned JSDoc: the webhook schema moved below the push schemas, so
PushSubscribeSchema has its comment back.
- Duplicated enums: WebhookUpdateSchema uses z.enum(WEBHOOK_KINDS/WEBHOOK_SCOPES), so
the schema cannot accept a kind the store would coerce away.
- describeError classifies egress refusals with isEgressBlockedError (the
CODEMAN_EGRESS_BLOCKED code anywhere in the cause chain) instead of a message regex;
tests pin a deep cause chain and that matching words alone are not a refusal.
- Markup: the URL input uses set-input, the whitespace-only line is gone, and the switch
row hints to pick a long random topic on public ntfy.sh.
- Remove a saved URL: a "Remove URL" button (shown only while a URL is saved, with a
confirm) sends { url: "", enabled: false }.
- Types placement: WEBHOOK_KINDS/SCOPES and WebhookKind/Scope/Urgency/Config/Result/Status
moved to src/types/push.ts (the IO-side WebhookMessage/Request/Fetch stay in the module).
Browser test extended: main Save persists a pending edit, a refused URL keeps the modal
open with the URL, Send test saves a newly pasted URL first, Remove URL clears it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude Code's advisor tool (code.claude.com/docs/en/advisor) lets the session's
main model consult a second, stronger model at decision points: before
committing to an approach, on a recurring error, and before declaring a task
done. Codeman can now start claude sessions with one.
- `advisorModel` field on POST /api/sessions, /api/quick-start and
/api/ralph-loop/start (fable, opus, sonnet or a full model id in those
families; haiku cannot advise and is refused). Stored on the session and
persisted, so respawn, boot restore and reboot restore keep it. Remote and
docker quick-starts refuse it, as they refuse effort.
- App Settings, Models, "Advisor" segment (Default / Sonnet / Opus / Fable),
synced as `claudeAdvisorModel`. Run, resume and the Ralph wizard send it.
Default sends nothing, leaving the CLI's own /advisor choice in charge.
- Carried as the `advisorModel` key in the launch's single --settings JSON,
merged with ultracode and the statusLine exporter, never the --advisor
flag: `claude --advisor haiku` exits 1 at launch, which would leave a dead
pane on every respawn, while the settings key degrades to no advisor. A
launch without an advisor is byte-identical to before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- app.js: the shortcut dispatcher returns early for events aimed at a data-raw-keys
field, so Ctrl+W / Ctrl+L / Escape / Alt+1 / Ctrl+K pressed in the Key tester no
longer kill the session, clear the terminal or close Settings
- stock.ts: drop Codex's esc-enter (a line feed works); no stock CLI declares a chord.
The esc-enter path is tested through a clis.json override
- tests: unused port (3194), Ctrl+Enter asserts no keypress, shortcut-isolation test
(verified to fail without the guard)
- docs/comments point at capabilities.newline; set-input class, trailing whitespace
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
capabilities.newline replaces choosing the Shift+Enter bytes in the send-key
route. Key tester shows the keydown/keypress/keyup a browser reports.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Opt-in (mcpSyncEnabled, default OFF; routes 403 until on). Review fixes:
- codex TOML read/validated with smol-toml: CRLF, inline tables and
command-less tables no longer yield a duplicate [mcp_servers.x]; the new
text is re-parsed before writing
- null-prototype tables and own-key checks; unsafe names ignored at every level
- servers switched off in their own CLI (codex/opencode/antigravity) are not copied
- only CLIs that are installed or already have a config file take part
- files receiving env/headers are left 0600; symlinked configs are written through
- one apply at a time (409), unique tmp files cleaned on failure, failed status
- routes set real HTTP status codes; api-reference section; format type single-sourced
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Adds capabilities.mcpConfig to the CLI registry (Claude, Gemini, Codex,
OpenCode), an additive src/mcp-sync.ts, GET/POST /api/mcp-sync and a
Settings > Agents & CLIs control. Never edits or removes an existing
server; backs up each file it changes; reports conflicts.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The vertical tab rail now reads the owner's tab layout (GET /api/tab-layout)
and draws its groups as collapsible sections. This is the first frontend
consumer of the tab-layout backend and it is read-only: nothing in the
browser writes the layout yet.
- tab-layout-browser.js (new, pure, loaded before app.js): projects the
layout onto the live sessions and open web tabs, renders the grouped
markup, stores collapse per device, and sequences loads newest-wins with
a bounded retry on failure.
- app.js: loads the layout on init and on tab:layoutChanged, renders the
grouped rail from the same per-row markup the flat rail uses, falls
through to a full render whenever the grouping structure changes, and
withholds drag-reorder in the grouped rail.
- Grouping is opt-in by construction. With no layout, a failed read, a
layout without groups, or a horizontal strip, the rail renders exactly
as before (byte-identical markup).
- Grouping is a render layer only: sessionOrder, Alt+N, Ctrl+Tab and the
palette keep reading the server-projected order, and row badges keep
their Alt+N slot.
- A collapsed group still shows the active row; lineage arcs to a hidden
session anchor to its group header.
- webview-tabs.js: renderWebviewTab() extracted so a single web tab can be
placed into its group with unchanged markup.
Clicking a .md in the Files panel showed wrapped source with an Edit
pencil and no way to see it rendered, although marked + DOMPurify were
already on the page for the Response Viewer. The viewer now renders
.md/.markdown through that same pipeline (one parser, one click
delegate) with an MD pill back to source, and the plain-text view gains
Lines (CSS-counter gutter) and Wrap toggles. All three persist per device
in their own localStorage keys.
- Relative images are rebased onto the workspace-confined file-raw route
under the document's directory, built inside a <template> so no fetch
fires before the rewrite; a failed load degrades to alt text. Relative
links become a.rv-path so the existing delegate opens them in the
viewer; fragment and http(s) links are untouched.
- The rendered container carries data-i18n-skip so the translator does
not rewrite the document's prose.
- Markdown fetches the route's 10000-line ceiling; other text keeps 500.
- avif renders inline (file-content image set, file-raw MIME map), and
avif/ico printed paths open the viewer instead of tailing bytes. .md
deliberately stays with the tail viewer for printed paths.
WebKit on iOS does not show text being composed by an IME inside the
terminal, so users type blind until it commits. Add mobile-ime-preview.js,
a visual-only controller that renders the composition in the xterm helper
layer and holds a committed chunk until local echo, parsed terminal output
or a 2s fallback shows it. Wire it into terminal-ui.js, the script order,
the build minify/hash lists and styles, with unit and wiring tests.
The phone case picker had no way to narrow a long case list, so finding
one meant scrolling a sheet that showed about six rows at a time.
- A search field filters rows by name (every typed word must match, any
order, case-insensitive), with a "No matching cases" state. Enter picks
the case when exactly one row is left; Escape clears, then closes.
- The field is not auto-focused, so opening the picker does not raise the
keyboard. The list holds its unfiltered height while searching so the
sheet does not jump, and the input is 16px so iOS Safari does not zoom.
- Layout: the sheet padded the home-indicator inset on top of the footer
already doing so, leaving a dead band under Create New Case; the sheet
now grows to 80dvh and the list fills it instead of a separate 50vh cap.
- Opening scrolls the list (its own box, not scrollIntoView) to the
currently selected case.
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
The Run bar's case picker only loaded /api/cases at page load, so folders
deleted or created on disk stayed listed until a reload. It now refetches
on open and every 5 seconds while open, repainting only when the list
changed and falling back to another case if the selected one was removed.
The Manage tab of Add Case gains a search box filtering by name or path.
Reorder arrows are disabled while a filter is active so a swap cannot
involve a hidden case.
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Adds claude-opus-5-5 to the App Settings model picker (base option with
data-ctx="1" plus its [1m] companion row, since Opus 5.5 has a 1M window)
and to the five task-routing selects, mirroring how Fable 5.1 was added.
Co-authored-by: Claude <noreply@anthropic.com>
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (da07b38c, db4557d9) with no
corresponding update to the plan doc itself — every checklist still read
Status: TODO and every box unchecked. Brings the doc in line with the tree:
- A new "Status as of 2026-09-22" section up top: what's actually
implemented (verified by grepping the routes/schema/UI, not just trusting
the commit messages), the availability-flag staleness bug found and fixed
in this session (commit 0c77dd0a) with its devbox verification record, and
one real outstanding gap.
- The outstanding gap: a custom CLI created via Phase 5's write API has no
way to actually be launched. The Run menu is static per-mode markup with
no consumer of window.__codemanCliCatalog, so Phase 6's own "create a
custom entry, confirm it can be launched" verify step was never actually
exercised against this. Documented with two candidate fixes, neither
started.
- Each phase's checklist flipped to [x] where confirmed present in the tree,
Status lines updated from TODO to DONE, and the two originally-open
questions (Phase 2's installed source, Phase 5's PUT endpoint shape)
marked resolved against what actually shipped.
No code changes in this commit — documentation only, so a future session
(or the one already mid-flight on a separate checkout of this same branch)
picks up accurate status instead of a stale plan.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs: add the CLI-registry deployment plan and the parked Copilot plan
Both were sitting as untracked scratch files in the master checkout,
never committed to any branch. Moving them here rather than leaving them
loose:
- DEPLOYMENT_PLAN.md is the live tracker for the CLI-registry follow-up
series (PR A #347 merged, PR B #380 merged, PR B2 merged as #458) and
is where PR C (this branch's own CLI-management work) belongs.
- docs/copilot-integration-plan.md is explicitly PARKED, referenced by
name in docs/cli-enable-disable-plan.md's own header as a sibling plan
tracked separately — kept for continuity, not active on this branch.
The other scratch files found alongside these (PRA.md, PRB.md, PR-B2.md
and their review-response counterparts) described PR A/B/B2, all now
merged — deleted from the master checkout as stale rather than committed
anywhere, since their content is superseded by the real merged PRs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* fix(cli-registry): render enabled CLIs in launch surfaces
* test(cli-registry): update frontend branch guard
* fix(test): isolate suite from deployment environment
* fix(cli-registry): revise Decision 4 - claude is toggleable, shell stays permanent
shell/claude were both structurally un-disableable in the original plan
(Decision 4). Revised: shell keeps the hard backend guarantee (it is the
one non-agent mode several code paths assume always exists as a raw-
terminal fallback), but claude is now a normal toggleable entry like any
other CLI.
Safe to do because internal session creation (tmux-manager.ts, session.ts,
Ralph, plan-orchestrator) resolves a CLI via getCli(), which does not
check `enabled` at all - only the Run menu and the HTTP-facing
sessionModeSchema() (new session requests through the normal API) key off
it. Disabling claude therefore behaves identically in kind to disabling
any other CLI: no internal fallback path breaks, it just stops being
offered for new sessions until re-enabled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): hide shell's toggle entirely instead of greying it out
A permanently-disabled switch next to every other row's working toggle
read as broken rather than intentional. shell now renders no switch at
all - a plain "Always available" label - so there is nothing to click
that could look like it should work but doesn't. Backend guard is
unchanged (UNDISABLEABLE_IDS still refuses shell unconditionally); this
is UI-only.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): sort the Installed CLIs list, installed-first then alphabetical
renderCliList() previously rendered in registry order (each entry's fixed
order field). Now sorts installed CLIs first, then not-installed, each
group alphabetical by label - matches how a user actually scans the list
(what's ready to use, then what needs installing).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* style: prettier fixes from the master merge
* fix(cli-registry): install/edit take effect immediately, confirm before install, phone labels
Four gaps found verifying #476 against the #343 review trail:
- Installed or edited CLIs kept reading as missing/stale. Every binary lookup
(the nine per-CLI resolvers and the generic registry one) caches in its own
closure, with a negative-cache backoff of up to 5 minutes, and nothing
cleared them. invalidateCliExecutableResolvers(binaries) now drops those
caches per binary; install (success or failure), create, edit and delete
call it plus invalidateCliResolverCache(id). Before this, a CLI installed
from Settings could fail to launch for minutes, and an edited custom entry
kept launching its old binary until a restart.
- The Settings "installed" badge for a custom entry used a private `which`,
ignoring the entry's searchDirs and the login-shell lookup that spawn and
the Run menu use; it now asks the same generic resolver they do.
- Install ran on a single click. The #343 review asked for auto-install to
sit behind an explicit confirm; the confirm now names the exact command,
which GET /api/clis returns for stock entries only (installCommand).
- The phone Run button showed the two-letter tab badge ("CC", "CX") instead
of the word ("Claude", "Codex"). It uses the registry label again, which is
identical to the old static table for every stock CLI (now pinned).
14 new tests; 9 of them fail against the previous head and pass here.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): address #476 review — safe serialized writes, no id branches, docs
Must-fix:
- registry-writer: start fresh only on ENOENT; refuse (409) a clis.json that
does not parse or has group/world permission bits instead of overwriting it
(isUnsafePermissions now exported from registry.ts)
- mutateRegistryFile(): one promise chain for every mutation, with the
existence/duplicate checks inside the serialized step, plus a unique tmp
name per write
- docs: CLAUDE.md, architecture-invariants, cli-registry (new Settings
section) and api-reference (the six /api/clis routes)
- drop DEPLOYMENT_PLAN.md and docs/copilot-integration-plan.md
Smaller:
- PUT /api/clis/custom/:id keeps the entry's current enabled state when the
body omits it
- runMode setter falls back to the first enabled catalogue entry, not 'claude'
- shell guard keyed on kind === 'shell' (routes + Settings list); stock probe
map shared with server.ts via utils/cli-installed-probes.ts
- stock claude label is now 'Claude Code', so the Run menu / phone overview
label rewrites are gone (doctor row keeps "Claude CLI" via its override)
- welcome buttons are translatable again and read "Run Claude Code" /
"Run Shell"; zh-CN gains "Run Codex" / "Run OMP"
- install: per-id in-flight guard (409) and CODEMAN_* stripped from its env
- fileoverview / CliEnableSchema comments no longer say stock-only
- test-env isolation changes moved to their own PR
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* test(cli-registry): pin the #343/#347 findings #476 makes reachable
A CLI toggled or created through the routes is accepted or rejected by
CreateSessionSchema with no restart (#343 finding 2), and a custom CLI created
through the API renders a real local, remote and docker launch command
(#347 finding 5: no more `cd <path> && undefined`).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Review fixes for #469.
The Ctrl+C branch cleaned the selection to decide whether to copy and then
passed that cleaned string to copyTerminalSelection(), which cleans again. The
trailing trim is a fixed point, so that was safe until this PR; the margin
strip is not, because it takes the lesser of the declared width and the run
every line shares, so a second pass takes up to `margin` columns more. The
branch now gates on the cleaned string and hands the raw one on. Verified in
chromium with a real drag, a real Ctrl+C and a real clipboard read on a live
claude pane: an on-screen ` fix(terminal): trim it` reaches the clipboard
as ` fix(terminal): trim it`, and reverting the branch reproduces the
reported ` fix(terminal): trim it`.
Pane B of a split resolves its own width. `_cliGutterColumns()` and
`_normalisedSelectionRange()` take the session and the terminal to read,
defaulting to the primary pane's, so Pane B looks its own run mode up instead
of keeping a margin Pane A drops on the same keystroke. Verified live with two
claude panes open side by side.
A detached session window (`/session/:id`) receives the gutter map. The
injection sat inside the block that skips the run menu's payloads for a solo
window, so the toggle worked in the main window and did nothing in the popup on
the same device. It needs no availability probe, so it moved below that block
and the solo window still carries none of the payloads it skipped before.
The settings description said the width is measured and named Codex as exempt.
Nothing is measured, and Codex is one of the two panes that are stripped.
docs/wiki/Settings-Reference.md gains the row every Terminal and Input toggle
carries. CLAUDE.md no longer says the clean touches trailing runs "and nothing
else" one sentence before the leading-margin rule, and both it and
docs/architecture-invariants.md record that the strip is not idempotent.
Two round-trip tests run on a mode that declares a gutter, which the existing
copyTerminalSelection cases could not, since they all use the harness default
mode that declares none. The Ctrl+C branch itself is pinned at the source,
because it lives inside initTerminal's attachCustomKeyEventHandler closure over
a real xterm the vm harness cannot build. Both pins fail on the reintroduced
bug. Gate: 7865 passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copying a paragraph out of a Claude Code or Codex pane puts that pane's own
two-column transcript gutter on the clipboard, so every pasted line arrives
indented. #451 shipped the trailing half of the copy clean and left the leading
half out, because deriving the width from the selection fires on 73% of ordinary
indented text and cannot tell a margin from content.
The width is DECLARED rather than derived. `capabilities.transcriptGutter` on
the CLI registry is a bounded integer; claude and codex each declare 2, measured
on live panes, and no other stock entry declares any, so a CLI whose transcript
layout nobody has measured is never touched. The server publishes the map as
`window.__codemanTranscriptGutter`, built by filtering `enabledClis()` on the
capability rather than by listing ids, and `_activeCliGutterColumns()` looks the
active session's mode up in it. The copy path reads no terminal buffer at all.
The declared width is a CEILING, not the answer: `clean()` strips the lesser of
it and the run every selected line shares. A block can therefore only shift as a
unit, the structure inside a selection survives by construction, and a selection
reaching column 0 loses nothing. That is what keeps a `git log` body at its own
four-space indent inside an agent's two-column gutter.
Codex was measured separately, because it renders nothing like Claude: it draws
boxes narrower than the pane and pushes its transcript into ordinary scrollback.
On a live 0.154.0 answer its `•`/`›`/`⚠` markers sit in the gutter, prose
continuations sit at 2, and a nested YAML block the model wrote rendered at
2/4/6/8 for its own 0/2/4/6. Replayed at 100, 120, 160, 198, 235 and 282 columns
its indents were 0, 2, 4, 6 and 8 at every one, never 1. Copying that YAML out
of a live Codex pane now yields 0/2/4/6: gutter gone, nesting intact.
Two derived versions were built and measured first, and both are recorded in the
code because both looked correct:
- Painted trailing padding — a full-screen TUI writes real spaces across the
unused part of a row, a shell leaves them never-written for xterm to trim —
has no false positives and never over-stripped. It is also a function of pane
WIDTH: the padding exists only while a rendered line stops short of the CLI's
own layout width, and Claude's prose wraps to fill it. Dragging the same two
prose rows of one live transcript at five window sizes, the share of padded
rows ran 44%, 6%, 6%, 7% and 87% at 123, 160, 198, 235 and 298 columns, so the
strip silently did nothing at every ordinary size while a corpus captured
entirely at 282 columns said it worked.
- Taking the narrowest indent on the rows around the selection fires at every
width and over-strips about 1% of selections, because a file listing inside
the transcript can be the narrowest thing on screen.
Measured over 1,392,281 selections — every 1, 2, 3, 5, 10 and 20-row window of
real Claude screens replayed from live PTY streams at 100, 120, 160, 198, 235
and 282 columns — the declared width over-strips none, breaks no relative indent
and alters no text, and serves 100% of the selections whose own indent covers
the gutter. Verified end to end in a browser with a real mouse drag and a real
Ctrl+C: Claude and Codex panes paste flush at 123, 198 and 298 columns, a shell
pane is untouched at every one.
The strip sits behind `copyStripMargin` (App Settings, Selection & clipboard),
per-device and default ON: a display key, absent from the .strict()
SettingsUpdateSchema, read as `!== false` because the desktop branch of
getDefaultSettings() returns {}. The toggle is checked before the map.
Two review findings from #451, handled:
- The mid-row flag governs ONE line now. `range.start.x > 0` excludes only the
first selected line, the one whose margin the mousedown genuinely cut off, so
the same three rows no longer produce three different clipboard results.
- The reversed-drag finding does not reproduce on the pinned xterm.
`getSelectionPosition()` reads `_selectionService.selectionStart`, whose
getter returns `SelectionModel.finalSelectionStart`, and that swaps the pair
when `areSelectionValuesReversed()` says so. A real upward mouse drag through
chromium against xterm 6.0 reports the same range as the downward drag.
`_normalisedSelectionRange()` keeps the ordering as a guard, because the model
one layer down exposes the unnormalised fields under the same two names.
Tests: test/terminal-copy-clean.test.ts (64, up from 31), plus the injected
script stripped in test/server-index-title.test.ts. Every guard is pinned:
removing any one of seven reds at least one test, including declaring the wrong
gutter width. Full suite green, 7,861 passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Pane B had no attachCustomKeyEventHandler of its own, so the document
capture-phase shortcut handler's preventDefault() (which does not stop
xterm) left Ctrl+K/Alt+1/Alt+B ALSO writing their raw byte/escape
sequence into Pane B's live PTY on top of whatever the app action did
to Pane A. Pane B now installs the same registry-aware gates the
primary pane's own attachCustomKeyEventHandler uses. Ctrl+V is left on
xterm's default paste — no image-paste trap to route it to.
Plus the rest of the review's smaller items:
- Narrowing the window past the desktop gate now closes an open split
instead of leaving it stranded on screen.
- Split is refused while a web tab is active (activeWebviewId), which
used to open Pane B's socket behind a hidden container.
- Pane B now handles the server's `{t:'r'}` refresh frame via a shared
_loadBuffer() helper (also used by connect()), instead of ignoring it.
- The divider drag now uses pointer events + setPointerCapture (mirrors
tab-rail-resize.js), a button!==0 guard, preventDefault, and a
body.split-pane-resizing cursor/selection lock — a plain mousedown
drag selected the text under the cursor as it crossed both terminals.
- Pane B's close control and the picker rows are real <button>s now
(keyboard-reachable), with matching CSS chrome resets.
- Dropped the redundant CodemanBase.base prefix on the buffer fetch
(the global fetch wrapper already applies it).
- data-preview-order for the Split settings chip moved from a collision
with Ultracode Agents (both 15/12) to 11.5, matching its real
position between Multi-monitor and Ultracode Agents in the header;
widened test/app-settings-structure.test.ts's regex to allow the
decimal (Number() already parses it fine for the preview sort).
- Added zh-CN i18n entries for the Split button and empty-picker text.
- Dropped the stray unused `vi` import Ark0N flagged as unrelated to
this feature (vitest's `globals: true` makes it ambient anyway).
- Documented the fix and the deliberate no-cid/seq choice in the
split-pane-sessions architecture-invariants entry.
Added a real-Chromium regression test asserting Ctrl+K/Alt+1/Alt+B
dispatched at Pane B's own textarea send no `{t:'i'}` frame over its
WebSocket. Full CI gate green (409 files, 7736 tests) plus all 8
split-pane browser tests.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two majors:
- The divider drag was unthrottled: every mousemove did a full xterm
reflow on BOTH panes and sent Pane B a {t:'z'} resize frame with no
unchanged-dimensions skip, fanning out into a `tmux resize-window`
child plus a SIGWINCH per event — ~50 of each dragging across half a
wide viewport. SplitTerminalPane.fit() is now split into localFit()
(reflow only) and fit() (reflow + send); the drag coalesces moves
into one localFit() per animation frame via requestAnimationFrame,
and sends the real resize for both panes exactly once, at drag end,
matching the primary pane's own throttledResize convention.
- Pane B pulled the FULL scrollback unchunked for every session mode,
writing it in one terminal.write() call. Mirrors the primary pane's
own mode check (app.js's selectSession): shell sessions get a
bounded 1MiB ?tail= fetch instead of ?full=1, and the fetched buffer
is written through a minimal chunked writer (32KB slices, yielding a
frame between each) instead of one primary-pane chunkedTerminalWrite
this simpler, independently created/destroyed pane has no equivalent
of (no session-switch generation counters or live-output gate).
Smaller items from the same review:
- Pane B now follows live appearance changes (applyTerminalSkin,
applyTerminalFontFamily, applyTerminalFontWeights, setFontSize all
propagate to it, matching the teammateTerminals pattern) and reads
the real codeman-font-size/terminalFontFamily/weights/DEFAULT_SCROLLBACK
settings at construction instead of hardcoding fontSize 14 / scrollback 5000.
- The Pane-B-promotion path now skips selectSession() when
_closingSessions already owns this delete (the user closing Pane A's
own tab), matching _onSessionDeleted's own active-session-handoff guard.
- Detaching a session AFTER a split is already open now yields the PTY
size in _sendResize() too (not just at picker-open time), mirroring
sendResize's own detachedElsewhere guard.
- .btn-split joins the body.solo-mode hide list, next to .btn-multimonitor.
- The split row was 6px wider than its container (two flex-shrink:0
50% panes plus a 6px divider): both panes are now flex-shrink 1.
- Pane B's header and the split-picker rows are marked so i18n.js's
exact-string lookup skips them, matching .session-name elsewhere —
a session literally named e.g. "Sessions" was translatable on zh-CN.
- The Split button now reflects open/closed state via a `.split-open`
accent style, aria-pressed, and a title/aria-label that says which
behaviour the next click gets.
- _splitPane.connect() is no longer an unawaited call with no .catch().
- terminal-split.js's fileoverview pointed at a doc path that was
renamed away in the previous push; @dependency now credits
constants.js for CodemanTerminalFont, not terminal-ui.js.
- index.html's Split settings chip no longer reuses data-preview-order
"12" (already the Ultracode Agents chip's slot in the same "header"
preview group).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Resolves CLAUDE.md count tables (route counts recounted on the merged
tree: 235 handlers, sessions 37) and keeps both the host-wake and the
reboot-restore banner in index.html.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.
- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
registry entry declares contextLengthVar (currently only claude) and
the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
(40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
quick-start customModel path) check this before the swap-conflict
check and before launching/restarting anything, returning
{requiresContextWarning, modelId, contextLength, minSafeContextTokens}
— skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
_resolveContextWarningConfirm (session-ui.js), wired into both
_quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
(the path Claude actually uses) ahead of the swap-confirmation check.
Explains the fix in-modal: give the model an explicit larger -c/
--ctx-size in llama-swap instead of relying on --fit-ctx, which
optimizes for the biggest model that fits rather than the biggest
context.
Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
The llama-swap "this will unload it for session X" warning used a native
browser confirm() popup, which looks out of place next to the rest of the
app's own modals.
Adds #customModelSwapConfirmModal (index.html) with Cancel/Switch-anyway
buttons, styled to match the app. _confirmModelSwap(message) shows it and
returns a promise that resolves true/false the same way confirm() would;
_resolveModelSwapConfirm(proceed) (wired to both buttons and the backdrop
click) settles it. Both llama-swap conflict call sites
(_quickStartWithCustomModelConfirm for the one-shot launch path,
_runCustomModelEntryViaRestart for Claude's restart path) now await this
instead of calling confirm() directly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
A host reboot takes the tmux server down with it, so every pane dies,
reconciliation finds nothing to attach to, and the board comes up empty.
Picking yesterday's work back up meant finding each conversation in history
and resuming it by hand, one at a time.
The boot pass now works out what the reboot killed and leaves it on offer.
It runs inside restoreMuxSessions(), in the window where reconciliation has
reported the dead sessions and cleanupStaleSessions() has not pruned their
records yet, which is the only place the records can still be read. The
board shows a banner, and nothing is created until the user clicks it.
A click rather than an automatic restore is what makes the reboot heuristic
acceptable. The heuristic cannot tell a reboot from a crash that took tmux
down inside the same window, so it decides whether to ASK, never whether to
act: a wrong yes costs a line of text the user dismisses instead of N CLI
processes nobody asked for.
Four things are re-checked when the click arrives rather than trusted from
boot, because hours can pass and the board moves on. The owner's privilege
grant re-resolves through the env clamp. The workspace must still be on
disk. A conversation the user already resumed by hand from the Resume list
is skipped, since two panes running --resume on one conversation would
fight over the same transcript. Entries leave the plan synchronously before
the first await, and the route is single-flighted, so a double-click or two
devices cannot both reach the same entry.
A restored session comes back attached, idle and disarmed. Respawn
controllers and Ralph loops are deliberately not re-armed: a machine that
just came up is the worst moment to turn an autonomous run loose. Its
workspace hooks are installed by the restore route itself, because the
boot-time sweep sits behind a gate that is false after a reboot and has
finished long before the click; without them a session goes silently blind,
with no stop or idle events for respawn, no Approvals Inbox item and no red
tab on a blocking dialog. Stats collection starts the same way.
The pane is new, so the conversation continues and the terminal scrollback
does not. The banner says so rather than letting an empty pane read as a
broken restore.
The plan lives in memory only. A server restart drops it, which costs the
convenience this adds and never the conversation: the conversation is the
transcript under ~/.claude/projects, which the Welcome screen's Resume list
and the Session Manager already read, so a dropped plan returns the user to
resuming by hand.
clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the
question it answers is about session privilege rather than about HTTP and
it now has a caller outside the route layer. Its test hook stays re-exported
from session-routes.ts.
Claude sessions only for this pass. The other CLIs name their thread in
their own config object, which this does not thread through yet. Remote and
docker sessions are skipped on purpose, because both need another host or a
container to be up and a freshly booted machine cannot promise either.
Refs #411
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The dialog had no max-height at all, so an endpoint with many discovered
models grew it past the viewport with nothing to scroll — reported live as
both "takes up the full page" and "the list is truncated", which turn out
to be the same bug. Gives #customModelPickModal .modal-content the same
bounded-height + scrollable-body shape cronModal's .modal-lg already uses
(max-height + flex column on the content, overflow-y:auto + flex:1 on the
body), scoped by id rather than folded into the shared .modal-sm class
three other modals already use for short, fixed content.
max-height: min(70vh, 520px) scales with the viewport (a phone gets 70% of
its height; a 4K display never gets a needlessly tall dialog) rather than
committing to one fixed pixel value that would be wrong at either end.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Two enhancements requested after live-validating PR #430 against a real
llama.cpp server:
1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to
apply the endpoint's defaultModelId (or the first discovered model)
silently. Now, via the new selectCustomModelEntry() (session-ui.js):
- exactly one discovered model launches straight away, same as before
- two or more open a new #customModelPickModal listing every discovered
model; defaultModelId (if set) is marked but never auto-chosen, since
the point of asking is letting ONE launch deliberately differ from
the saved default, not just confirming it
The endpoint is re-fetched at click time rather than trusting anything
cached from the dropdown's own render, since the model list can have
changed (the sweep below, or a settings-panel edit) since it opened.
runCustomModelEntry() itself — the actual launch, routed through run()
for the in-flight lock, snapshot-guarded against applying to the wrong
session — is unchanged; it now just always receives an explicit model
id from one of these two paths instead of computing one itself.
2. Periodic re-discovery. Every saved endpoint's models now refresh
automatically every 5 minutes in the background
(CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way
as the Codex plan-usage poll it sits beside — this.cleanup.setInterval,
off under testMode), so a model the server starts or stops serving shows
up without another manual "Discover" click. The manual POST
.../discover-models route and the new refreshAllCustomModelHosts()
sweep (custom-model-routes.ts) now share one pure merge step
(applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId
that no longer appears) rather than two copies that could drift. The
sweep is best-effort per host — one endpoint being unreachable on a
cycle never blocks the others — and re-reads the store before each
host's write, keyed by id, so a concurrent edit or delete from the
settings panel always wins over a sweep that started before it.
Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated
file for the sweep (kept separate from custom-model-routes.test.ts because
that file's data dir is shared across every test in it — one temp HOME per
FILE, not per test — which would make a sweep-touches-every-host assertion
meaningless there). test/custom-model-run-menu-ui.test.ts gained a new
describe block driving the real picker modal through JSDOM: single-model
bypass, multi-model dialog with the default marked-not-chosen, picking a
row closes the modal and launches with that exact model, the endpoint
re-fetch, and the two "vanished by click time" toast paths.
Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md,
docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated
— the last of these also caught up two sentences that had gone stale after
the draft-review fixes landed (the picker routes through run() now, not a
raw run*() call).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.
Blockers:
1. Every generated inline onclick was unparseable. JSON.stringify's own
double quotes terminated the double-quoted HTML attribute at the first
one, leaving btn.onclick null on every picker entry and every Discover/
Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
argument, the same idiom deleteCase's onclick already uses four lines
away in session-ui.js. This also closes the live-HTML-injection route
through modelId (server-controlled, from the endpoint's own /v1/models
reply): with quoting intact, a `>` inside it can no longer terminate the
<button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
like every other /api route (server.ts's preSerialization hook applies
to arrays too), so Array.isArray(hosts) was always false in production
and the picker/settings panel silently saw nothing. Both call sites now
go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
returns normally without ever changing activeSessionId, so the apply
step used to silently re-point and restart whatever session the user was
already looking at. runCustomModelEntry() now snapshots activeSessionId
before the launch and requires it to have actually changed.
Majors:
4. Routes the launch through run() itself via a temporary _runMode swap
(never persisted — setRunMode() would sync it to the server) instead of
a parallel hardcoded dispatch table, so a custom-model launch now holds
the same _runInFlight lock every other Run click gets. This also
resolves the "hardcoded runners map contradicts the PR's own design"
minor: dispatch is run()'s own, so a CLI whose customModelInjection
recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
parses markup this module generated itself) for exactly the DOM-level
facts the review said needed no Playwright and no tmux: a generated
button's onclick genuinely compiles and fires, a dangerous modelId never
produces a live element, the envelope unwrap works, the session-changed
guard holds, run() actually gets called (proving the in-flight lock
engages), and _runMode is restored afterward. Confirmed against the
pre-fix code first (reproduces btn.onclick === null exactly) so this
isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
render-index-html.test.ts for the other fixes below.
Minors:
- Generated entries now filter through isCliAvailable(), matching
_refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
(applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
and to settings-modal open) instead of always rendering; the endpoint GET
no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
redactApiKey() replaces the field with a computed apiKeySet: boolean, and
a PUT with no apiKey now keeps the stored one server-side
(applyStoredApiKey()) instead of the client resending a value it was
never given. New tests cover both directions (kept vs. replaced) by
observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
(_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
event, since the real role can resolve after settings were first opened)
— endpoint writes were already admin-only server-side, but the button
used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
(CLAUDE.md already records that exact literal turning the settings
preview into a grey slab on light skins), .run-mode-custom-models gets
the same gap: 2px .run-mode-menu's own flex gap only applies one level
up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
</script> (CliEntry.label is user-clis.json-settable, unlike
__codemanCliAvailable's booleans-only payload) via a new exported
escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
docs/api-reference.md; left the "no zh-CN for the new Models-section
group" minor unaddressed only insofar as the wider Models section (task
routing, thinking effort, etc.) has never had zh-CN coverage either —
everything this PR itself introduces (labels, hints, button text, the
Run-menu's "Custom Endpoints" header) IS translated in i18n.js.
Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Follow-up to #393, picking up the work Ark0N invited in his merge comment:
"generate those entries from the saved profiles rather than a fixed
duplicate per harness, and put it in a follow-up PR so this one stays the
backend... The Run-menu picker is yours if you want it."
Adds the frontend surface the backend has been waiting on:
- Run menu: a "Custom Endpoints" section lists one entry per (harness that
supports customModelInjection, saved endpoint) pair, e.g.
"Claude Code (llama.cpp)". The harness list comes from
window.__codemanCustomModelClis, injected at page render straight off the
CLI registry's own capabilities (never a hardcoded id list in the
frontend), so a CLI whose injection recipe lands later appears with no
frontend change. Picking an entry runs that harness's own existing run*()
function unmodified (case creation, env overrides, everything, forced to
a single instance) and then applies the endpoint's default model to the
session it creates via the existing POST /api/sessions/:id/custom-model
route. Entries are hidden for a remote/docker active case, since that
route already refuses both.
- Settings: App Settings -> Models gets a "Custom model endpoints" group
wiring up the customModelEndpointsEnabled toggle (declared since #393,
read by nothing until now) plus CRUD against the existing
/api/model-endpoints routes: list, add/edit (inline form), delete,
discover models.
- Backend: CustomModelHost gains an optional defaultModelId, the model the
picker applies with no further choice per endpoint (one generated menu
entry per CLI+endpoint pair, not per CLI+endpoint+model). The route
refuses a value that isn't one of the endpoint's own discovered models,
and a fresh discovery drops a default that no longer appears rather than
carrying an invalid one forward.
Docs: docs/custom-model-endpoints.md describes the new picker and settings
panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the
"backend-only" status note and documents the picker's generation mechanism.
Tests: four new route tests cover defaultModelId validation, acceptance,
and the drop/keep behaviour across a re-discovery; a new render-index-html
test pins the __codemanCustomModelClis injection (present, agent CLIs
supporting the capability, antigravity and shell excluded) and its
solo-window skip. No browser test was added for the Run-menu picker itself
or the settings CRUD panel (this box has no tmux, so the live server used
by test:browser/test:mobile could not be exercised here) -- worth a
Playwright pass before merge, same as any other frontend PR.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Finishes #376. The contributed keystroke tracker sat on the raw byte stream
and named tabs wrong five ways (every prompt, every write path, a bare Esc
eating the next prompt's first character, pasted newlines as Enter, any CSI
clearing the draft) and replaced the whole name, which dropped the case from
the tab and reset the w<n> counter. This lands the feature with each of those
closed:
- First prompt means the first: applyAutoName() flips a placeholder to
`auto` whether or not the string changed. nameSource is now the tri-state
placeholder | auto | manual; the name setter is the only manual path.
- Only user-originated input counts: write()/writeViaMux() take
SessionWriteOptions.fromUser, set by the browser WS path and POST /input
only, so Ralph, respawn, cron, approvals and the trust-dialog keys can
never name a tab. A startMode 'shell' CLI never feeds the tracker (a
capability, not an id check); the send-key route feeds trackUserInput()
because its line feed bypasses the session.
- Prefix form `w3-case: title`: parseSessionPrefix() already renders it as
the title with the prefix in the tooltip and the next-session counter
still matches it. Composed within MAX_SESSION_NAME_LENGTH.
- Tracker rules per key: bare Esc resolves at chunk end; mouse/focus
reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R
taint the draft so Enter submits nothing rather than a fragment;
bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space;
the draft keeps its head past 8192 code points; an escape past 64 bytes
is abandoned.
- Title: slash commands by shape (a path is a prompt), `!` escapes
refused, first sentence only past 8 code points ("e.g." is not a title),
72 code points on a word boundary.
- Synced `autoNameSessions` setting, default OFF (the prompt reaches
mux-sessions.json, session:updated and /api/search), App Settings ->
Appearance -> Tabs, read fresh per prompt after the eligibility check.
Tests: test/session-auto-name.test.ts (tracker, title, composition,
ownership, emit gating), the wiring test (once, prefix, setting off,
manual protected), test/routes/session-name-routes.test.ts (PUT /name
flips to manual and persists). Verified live on an isolated instance: API
and browser-typed prompts name the tab, a second prompt does not, shells
and renamed tabs are untouched, nameSource survives a restart.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The reactive wake (typing into a session whose host slept) left the state invisible:
nothing told the user the machine was asleep, and with no wake target configured
there was nothing to do about it. Adds:
- RemoteHost.wakeMac (comma-separated) - Codeman builds and broadcasts the magic
packet itself (UDP port 9), so the common case needs no external script. The
existing wakeCommand stays as the explicit override.
- GET /api/sessions/:id/reachability - probes (throttled, cached, and it never
wakes) and reports HOW the host can be woken, or that nothing is configured.
- POST /api/sessions/:id/wake - wakes, waits, reattaches the pane and flushes
buffered input; 400 with a routable message when no target is configured.
- The amber host-unreachable banner + its 'Wake' / 'Configure WoL' action, and a
small config dialog that saves via PUT /api/remote-hosts/:id.
- RemoteWakeDeps.resolveRemote: host config is re-resolved for LIVE sessions
(throttled + cached), so saving the dialog takes effect without a restart.
The backend already lets one adopted container back several cases pointing at
different in-container directories, but using it meant retyping the container
name, host and workspace one by one — exactly the friction that leaves a
capability unused. Picking an existing case from a dropdown now carries those
three over, leaving only the two fields that must differ: the case name and the
in-container directory.
Clearing those two is the point of the feature, not a convenience: keeping the
old name is refused by the server as "case already exists", and keeping the old
directory is refused as "a twin case on the same container and directory". Both
errors are clear, but a form pre-filled with values that are guaranteed to be
rejected is a trap. Focus lands on the in-container directory — the thing the
user came here to change.
⚠️ Only adopted containers are listed (docker.owned === false). A Codeman-built
container's lifecycle belongs to its one case — a second case would be torn out
by that case's recreate or delete — so the server refuses it anyway, and listing
it here would only manufacture a baffling error. `owned` may be absent and absent
means owned, so the test is `!== false`, not truthiness.
CaseInfo.docker gains containerWorkdir and owned for this: the former is the
"which directory does this case use" half of the picker, without which the user
cannot tell what to change it to; the latter backs the filter above. ⚠️ Both
places that build a docker CaseInfo (the list endpoint and the single-case query)
must set them — filling in only one makes the picker work or not depending on
which read path was taken, and a test pins "exactly two".
(cherry picked from commit f1ed3a58e1)