Compare commits

...
Author SHA1 Message Date
Codeman maintainer 61037082d1 fix(mobile): make the phone header tab strip read as live tabs
On a phone every inactive tab rendered transparent: grey 11px text
floating in unmarked gaps, a boxed Alt+N digit in each tab (a phone has
no Alt key), names capped at 50px so a shared `w1-` prefix was most of
what showed, and the tab that did not fit was chopped mid-word against
the connection dot. The strip looked like a row of disabled labels.

Phone block of mobile.css only:
- Every header tab is a chip, filled and bordered from the skin's
  --control-* tokens, name in --text at weight 500. Written
  `:where(.header) .session-tab` so it stays at (0,1,0): the per-colour
  left border still wins, and sidebar layout (where the list leaves the
  header) is untouched.
- The Alt+N digit is hidden in the header; inactive tabs drop their
  empty .tab-actions container, which padded the chip's right side.
- Name cap 50px -> 80px, status dot 4px -> 6px, strip gap 2px -> 6px.
- Scroll-driven edge fade: a mask on the strip whose widths follow its
  own inline scroll timeline (registered @property lengths), so the
  clipped tab dissolves into the edge. No JS; a strip that does not
  overflow gets no mask, and browsers without scroll timelines keep the
  old hard edge.

The tap-zone arithmetic comment is updated for the numberless phone
tabs and the bigger dot (the required reserve drops from 38px to 36px;
the 44px min-width stays). test/mobile-tab-strip-chips.test.ts pins the
(0,1,0) selector, the top-level @property registration and the
timeline-after-shorthand order, each of which fails silently otherwise.
test/mobile/tabs.test.ts follows the new name cap.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 15:23:22 +02:00
Codeman maintainer 45ea2e1d32 docs(readme): ask readers to star the project
Adds a centered star call-to-action under the badge row in both the
English and Simplified Chinese READMEs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-26 04:39:13 +02:00
Codeman maintainer 5ae574374f docs(changelog): add the Thanks section to 1.33.1
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 23:47:08 +02:00
Codeman maintainer 47f209bf0a chore: version packages (1.33.1)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 23:20:42 +02:00
Ark0NandCodeman maintainer d81a4a76de feat(mobile): search box in the Select Case picker (#488)
The phone case picker had no way to narrow a long case list, so finding
one meant scrolling a sheet that showed about six rows at a time.

- A search field filters rows by name (every typed word must match, any
  order, case-insensitive), with a "No matching cases" state. Enter picks
  the case when exactly one row is left; Escape clears, then closes.
- The field is not auto-focused, so opening the picker does not raise the
  keyboard. The list holds its unfiltered height while searching so the
  sheet does not jump, and the input is 16px so iOS Safari does not zoom.
- Layout: the sheet padded the home-indicator inset on top of the footer
  already doing so, leaving a dead band under Create New Case; the sheet
  now grows to 80dvh and the list fills it instead of a separate 50vh cap.
- Opening scrolls the list (its own box, not scrollIntoView) to the
  currently selected case.

Co-authored-by: Codeman maintainer <noreply@anthropic.com>
2026-09-24 23:08:53 +02:00
DevvynandClaude Sonnet 5 8841bcc93f feat(cases): refresh the case picker and add search to Manage (#483)
The Run bar's case picker only loaded /api/cases at page load, so folders
deleted or created on disk stayed listed until a reload. It now refetches
on open and every 5 seconds while open, repainting only when the list
changed and falling back to another case if the selected one was removed.

The Manage tab of Add Case gains a search box filtering by name or path.
Reorder arrows are disabled while a filter is active so a swap cannot
involve a hidden case.


Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 23:08:47 +02:00
DevvynandClaude Sonnet 5 77ba41f8da feat(docker): install uv/uvx, libsecret-1-0 and pnpm (#487)
* fix(docker): install pnpm in the Compose server image

`dsh plugin` spawns a literal `pnpm` with no npm fallback, so the Run
menu's "DeepSeek - add a terminal profile" button failed with
`dsh: pnpm not found on PATH` (exit 127) on the server image. The agent
image already installs pnpm for the same reason (#352). Pin pnpm@12.6.0
in the runtime-writable CLI prefix and note it in the DeepSeek doc.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* feat(docker): install uv and uvx in server and agent images

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* feat(docker): install libsecret-1-0 for the Azure DevOps MCP

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* feat(docker): add sudo to the agent image

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* feat(docker): install sudo with passwordless access for the agent user

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* Revert "feat(docker): install sudo with passwordless access for the agent user"

This reverts commit b070c9ee65.

* Revert "feat(docker): add sudo to the agent image"

This reverts commit e98127a804.

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 23:08:41 +02:00
Michael GrundbergandClaude Opus 5.5 b80d47aff8 feat(session): close sessions whose agent exited cleanly (#486)
* fix(cleanup): keep .claude-images while a sibling session uses the same dir

cleanupSession() recursively removes {workingDir}/.claude-images. That
directory belongs to the working directory rather than to the session, and
several sessions routinely share one case directory, so closing one session
deleted the pasted images a live sibling still referred to.

The removal now runs only when no other live session has the same working
directory. A session that is itself being cleaned up does not count as live,
so two sessions of one case closed together still remove the dir.

Split out ahead of the exited-agent sweep for Ark0N/Codeman#446, which closes
sessions unattended and would otherwise make the loss routine.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* feat(session): close sessions whose agent exited cleanly (#446)

Part 2 of Ark0N/Codeman#446. Part 1 records an exited agent as
SessionState.paneExit. A session whose agent the user ended with /exit is
now closed through cleanupSession(), the same path the X button takes, so
finished sessions stop piling up on the board. The lifecycle log records
the reason as "agent exited cleanly (status 0)", and the conversation stays
resumable from the Resume list.

shouldCloseCleanlyExitedSession() in the new pure module pane-exit-sweep.ts
holds the rule. It closes a session only when all of these hold:

- The exit status is an explicit numeric 0 with no signal. An absent status
  is how a SIGKILL presents on tmux 3.2a, so it counts as unknown and the
  row stays. A non-zero status or any signal also keeps the row, with the
  exit code on the tab.
- Two authoritative pane reads agreed on that exit.
  TmuxManager.getPaneExitReadCount() counts them, and a failed, empty or
  skipped read neither confirms nor resets the count.
- No start, attach or relaunch is running for the pane.
  Session.paneLifecycleInFlight covers _setupOrAttachMuxSession(), whose
  dead-pane branch revives an exited pane on purpose, and restartCli().

setPaneExit() already scopes paneExit to local mux-backed sessions, so
remote, docker and direct-PTY sessions are never closed.

planRebootRestore() now refuses a record whose persisted paneExit is a
clean exit. That covers an agent that exited just before a reboot, before
the sweep reached it. A crashed agent's record stays eligible, like its row.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(web): show "exited" on the phone overview and desktop home rail (#446)

Part 1 of Ark0N/Codeman#446 taught the tab strip and the rich rail rows
to say that a session's agent has exited. The phone overview and the
desktop home rail still said "idle", beside a green or pulsing dot.

_mobileOverviewExit() in mobile-overview.js is now the one rule for all
three surfaces, and _sidebarRichRow() uses it as well. It changes what a
row shows and leaves the row's state alone, because the state still picks
the section and the sort order. An exited row gets an "exited" pill, a
neutral dot and row accent, and a duration measured from when the server
first saw the pane dead. A pending permission prompt or question still
wins, as it does on the tab.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cleanup): close the gaps review found in the #446 sweep and image guard

Four fixes from a dual review of Ark0N/Codeman#446 part 2.

- The .claude-images guard compares canonical paths, so a sibling that
  reaches the same directory through a symlink keeps it. Its comment used to
  say that case only missed a deletion; it caused one.
- A detached session counts as a live sibling. DELETE ?killMux=false removes
  it from the server's map while its pane keeps running, so the guard now
  reads persisted records too, and exempts only sessions being killed rather
  than every session in cleaningUp.
- A session being closed refuses startInteractive() and startShell(). The
  /interactive route awaits listener setup before the start, and a start
  that raced the close could launch a CLI in a tmux session whose record was
  then deleted. A failed close clears the mark again.
- The clean-exit sweep tries each exit once, keyed by session id and the
  exit's at stamp, so a close that fails is not retried and logged every
  two seconds.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(session): keep a clean exit that lands within 10 s of a pane start (#446)

A CLI that prints a startup error ("not logged in", a bad profile, a config
error) and exits 0 used to lose its tab, and the error with it, about 4 s
after launch. The sweep now keeps any clean exit that lands within
CLEAN_EXIT_MIN_PANE_LIFETIME_MS (10 s) of the last start, attach or relaunch
finishing (Session.paneStartedAt, stamped when _withPaneLifecycle ends). The
row stays as "exited (0)" for the user to read and close.

Verified on an isolated instance: a shell that ran `exit 0` 2 s after start
kept its row, one that exited after 13 s was closed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 22:18:07 +02:00
Codeman maintainer d67da5c9d0 fix(input): an oversized paste no longer poisons the durable input queue (#484)
A single input over MAX_INPUT_LENGTH (64 KiB) was queued for reliable
delivery, refused by both transports (the WebSocket silently, POST with a
400), and never dropped: the client treated the 400 as transient, so the
frame was re-sent every 2 s forever, blocked every later input for that
session, and came back from localStorage on every reload.

- Client: a paste over the frame limit is split into in-limit frames
  (never cutting a surrogate pair) delivered in seq order; over 1 MiB, or
  an oversized mux write, it is refused with a toast and never queued.
- Client: the POST drain drops a frame answered 400/413; a WS error ACK
  drops it too; frames over the limit persisted by an older build are
  pruned on load.
- Server: the WebSocket answers an oversized sequenced frame with
  {t:'ia',seq,err:'too_large',max} instead of silence (an older client
  reads that as a plain ACK and drops it); the POST schema uses
  MAX_INPUT_LENGTH instead of a second 100000 limit.

Verified end to end on an isolated instance: a 110 KB paste reached the
PTY byte-identical over both the WebSocket and the POST path, a poisoned
120 KB persisted frame was pruned on load, and a 2 MB paste showed the
refusal toast with nothing queued.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 18:17:06 +02:00
Codeman maintainer e6ddb0485a chore: version packages (1.33.0)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 01:57:55 +02:00
Ark0NandClaude 334884e96a feat(models): offer Opus 5.5 in the model picker and task routing (#480)
Adds claude-opus-5-5 to the App Settings model picker (base option with
data-ctx="1" plus its [1m] companion row, since Opus 5.5 has a 1M window)
and to the five task-routing selects, mirroring how Fable 5.1 was added.

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-24 01:57:35 +02:00
69a71287e6 fix(sessions): stop pinning the w1-myapp placeholder as Claude's /resume title (#457)
* fix(sessions): stop pinning the w1-myapp placeholder as Claude's /resume title

Local claude spawns passed the tab name as `--name`. That flag is not only the
cross-session peer name: it is also the prompt-box label, the `/resume` picker
entry and the terminal title, and a pinned title stops Claude generating its own
(`customTitle ?? aiTitle`). So every conversation of a case was listed in
`/resume` as the same `w1-myapp`, and none of them got a generated title. On one
workspace, 34 of 34 conversations spawned with `--name` had no ai-title, while
every conversation spawned without it had one.

Only a name the user chose is pinned now: `Session.cliPinnedName` is the name
when `nameSource === 'manual'`, carried to the builders as a separate `cliName`
so the tab/mux name is untouched. Placeholder and auto names let Claude title
the conversation again.

A rename in Codeman also reaches `/resume`: the new name is appended to the
conversation's transcript as the `custom-title` row `/rename` writes (never
creating the file, never writing an empty title). For a pane spawned without
`--name` this holds immediately; a pane spawned with one re-appends its own
title each turn, so there the new name holds from the next spawn.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(sessions): skip no-op renames and docker sessions when syncing the /resume title

A same-name PUT (the Session Options field saves on blur and recomposes the
unchanged placeholder) no longer flips nameSource to manual or appends a
custom-title row, and docker sessions skip the host transcript scan since their
transcript lives in the container. The skill pages no longer use a w<N>- name
as the peer-name example, and the changeset notes the re-append caveat.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: record that nameSource decides --name and renames reach /resume

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: codeman-local <codeman@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-24 01:48:33 +02:00
Julian MartinezandClaude Opus 5.5 7485afecaf fix(session-manager): carry mode, claudeSessionId and resumeId into rows (#477)
_loadSessionManagerList() re-projects each unified item into the
history-record shape _buildHistoryItem renders, and dropped these three
fields. The row's own onActivate still read them from the unified item, but
everything built from the record did not: the ⋯ menu's "Resume session"
relaunched a codex row as claude (no mode, no resumeId), a resumed session
lost its conversation id, and Cmd+K rows showed no mode badge. Same class
of bug as the worktree fields the re-projection already carries (#266).

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 01:48:29 +02:00
DevvynandClaude Opus 5.5 0a52a99ca9 feat(cli-registry): CLI management write API + Settings UI (Phases 1-6) (#476)
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis

Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI

Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).

Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.

Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).

Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.

Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.

Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.

27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else

window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.

Fixed in two places:

- server.ts: after building `available`, intersect the nine real
  SessionMode ids against `enabledClis()`. git/cloudflared (utility
  binaries, not CLI registry entries) and deepseekBinary (a secondary
  installed-only flag for the "add a profile" affordance) are deliberately
  left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
  `window.__codemanCliAvailable` in place and refreshes the welcome screen,
  the mobile overview and an already-open Run menu, mirroring the existing
  `installDeepSeekProfile()` pattern for the same "injected once, needs an
  explicit patch" reason — without this half, the server-side fix alone
  still left every surface stale until the next reload.

New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.

Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap

Phases 1-6 were implemented across two commits (da07b38c, db4557d9) with no
corresponding update to the plan doc itself — every checklist still read
Status: TODO and every box unchecked. Brings the doc in line with the tree:

- A new "Status as of 2026-09-22" section up top: what's actually
  implemented (verified by grepping the routes/schema/UI, not just trusting
  the commit messages), the availability-flag staleness bug found and fixed
  in this session (commit 0c77dd0a) with its devbox verification record, and
  one real outstanding gap.

- The outstanding gap: a custom CLI created via Phase 5's write API has no
  way to actually be launched. The Run menu is static per-mode markup with
  no consumer of window.__codemanCliCatalog, so Phase 6's own "create a
  custom entry, confirm it can be launched" verify step was never actually
  exercised against this. Documented with two candidate fixes, neither
  started.

- Each phase's checklist flipped to [x] where confirmed present in the tree,
  Status lines updated from TODO to DONE, and the two originally-open
  questions (Phase 2's installed source, Phase 5's PUT endpoint shape)
  marked resolved against what actually shipped.

No code changes in this commit — documentation only, so a future session
(or the one already mid-flight on a separate checkout of this same branch)
picks up accurate status instead of a stale plan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* docs: add the CLI-registry deployment plan and the parked Copilot plan

Both were sitting as untracked scratch files in the master checkout,
never committed to any branch. Moving them here rather than leaving them
loose:

- DEPLOYMENT_PLAN.md is the live tracker for the CLI-registry follow-up
  series (PR A #347 merged, PR B #380 merged, PR B2 merged as #458) and
  is where PR C (this branch's own CLI-management work) belongs.
- docs/copilot-integration-plan.md is explicitly PARKED, referenced by
  name in docs/cli-enable-disable-plan.md's own header as a sibling plan
  tracked separately — kept for continuity, not active on this branch.

The other scratch files found alongside these (PRA.md, PRB.md, PR-B2.md
and their review-response counterparts) described PR A/B/B2, all now
merged — deleted from the master checkout as stale rather than committed
anywhere, since their content is superseded by the real merged PRs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD

* fix(cli-registry): render enabled CLIs in launch surfaces

* test(cli-registry): update frontend branch guard

* fix(test): isolate suite from deployment environment

* fix(cli-registry): revise Decision 4 - claude is toggleable, shell stays permanent

shell/claude were both structurally un-disableable in the original plan
(Decision 4). Revised: shell keeps the hard backend guarantee (it is the
one non-agent mode several code paths assume always exists as a raw-
terminal fallback), but claude is now a normal toggleable entry like any
other CLI.

Safe to do because internal session creation (tmux-manager.ts, session.ts,
Ralph, plan-orchestrator) resolves a CLI via getCli(), which does not
check `enabled` at all - only the Run menu and the HTTP-facing
sessionModeSchema() (new session requests through the normal API) key off
it. Disabling claude therefore behaves identically in kind to disabling
any other CLI: no internal fallback path breaks, it just stops being
offered for new sessions until re-enabled.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): hide shell's toggle entirely instead of greying it out

A permanently-disabled switch next to every other row's working toggle
read as broken rather than intentional. shell now renders no switch at
all - a plain "Always available" label - so there is nothing to click
that could look like it should work but doesn't. Backend guard is
unchanged (UNDISABLEABLE_IDS still refuses shell unconditionally); this
is UI-only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): sort the Installed CLIs list, installed-first then alphabetical

renderCliList() previously rendered in registry order (each entry's fixed
order field). Now sorts installed CLIs first, then not-installed, each
group alphabetical by label - matches how a user actually scans the list
(what's ready to use, then what needs installing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* style: prettier fixes from the master merge

* fix(cli-registry): install/edit take effect immediately, confirm before install, phone labels

Four gaps found verifying #476 against the #343 review trail:

- Installed or edited CLIs kept reading as missing/stale. Every binary lookup
  (the nine per-CLI resolvers and the generic registry one) caches in its own
  closure, with a negative-cache backoff of up to 5 minutes, and nothing
  cleared them. invalidateCliExecutableResolvers(binaries) now drops those
  caches per binary; install (success or failure), create, edit and delete
  call it plus invalidateCliResolverCache(id). Before this, a CLI installed
  from Settings could fail to launch for minutes, and an edited custom entry
  kept launching its old binary until a restart.
- The Settings "installed" badge for a custom entry used a private `which`,
  ignoring the entry's searchDirs and the login-shell lookup that spawn and
  the Run menu use; it now asks the same generic resolver they do.
- Install ran on a single click. The #343 review asked for auto-install to
  sit behind an explicit confirm; the confirm now names the exact command,
  which GET /api/clis returns for stock entries only (installCommand).
- The phone Run button showed the two-letter tab badge ("CC", "CX") instead
  of the word ("Claude", "Codex"). It uses the registry label again, which is
  identical to the old static table for every stock CLI (now pinned).

14 new tests; 9 of them fail against the previous head and pass here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* fix(cli-registry): address #476 review — safe serialized writes, no id branches, docs

Must-fix:
- registry-writer: start fresh only on ENOENT; refuse (409) a clis.json that
  does not parse or has group/world permission bits instead of overwriting it
  (isUnsafePermissions now exported from registry.ts)
- mutateRegistryFile(): one promise chain for every mutation, with the
  existence/duplicate checks inside the serialized step, plus a unique tmp
  name per write
- docs: CLAUDE.md, architecture-invariants, cli-registry (new Settings
  section) and api-reference (the six /api/clis routes)
- drop DEPLOYMENT_PLAN.md and docs/copilot-integration-plan.md

Smaller:
- PUT /api/clis/custom/:id keeps the entry's current enabled state when the
  body omits it
- runMode setter falls back to the first enabled catalogue entry, not 'claude'
- shell guard keyed on kind === 'shell' (routes + Settings list); stock probe
  map shared with server.ts via utils/cli-installed-probes.ts
- stock claude label is now 'Claude Code', so the Run menu / phone overview
  label rewrites are gone (doctor row keeps "Claude CLI" via its override)
- welcome buttons are translatable again and read "Run Claude Code" /
  "Run Shell"; zh-CN gains "Run Codex" / "Run OMP"
- install: per-id in-flight guard (409) and CODEMAN_* stripped from its env
- fileoverview / CliEnableSchema comments no longer say stock-only
- test-env isolation changes moved to their own PR

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

* test(cli-registry): pin the #343/#347 findings #476 makes reachable

A CLI toggled or created through the routes is accepted or rejected by
CreateSessionSchema with no restart (#343 finding 2), and a custom CLI created
through the API renders a real local, remote and docker launch command
(#347 finding 5: no more `cd <path> && undefined`).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 01:48:26 +02:00
Ark0NandClaude Opus 5.5 c46e87fd7a fix(self-update): stalled status and hung shutdown on launchd-daemon installs (#478)
* fix(self-update): stop a stalled status from blocking every later update

A Homebrew node upgrade under a long-running server deletes the versioned
Cellar path the server passes as --node, so every status write from the
updater failed. The update itself still built and restarted (npm and the
build use node from PATH), but update-status.json stayed "queued" forever.
The boot reconcile ran one minute after the restart, inside its 15 min
window, and isInFlight() had no age limit, so "An update is already in
progress." blocked every later update until the next server restart.

- self-update.sh falls back to node on PATH when --node is not executable.
- expireStalledStatus() (pure) fails an in-flight status whose last write
  is older than the stale window; applied on every read (start + status
  poll) and persisted. The live updater heartbeats every few seconds, so a
  running update never trips it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(self-update): a hung graceful shutdown no longer leaves a LaunchDaemon install down

On a KeepAlive LaunchDaemon (headless macOS) the updater restarts by sending
the server SIGTERM and letting launchd respawn it. launchd only respawns once
the process EXITS, and nothing escalates a stuck stop (systemd would SIGKILL
after TimeoutStopSec). Observed after an update to 1.32.1: the server closed
port 3000, server.stop() never resolved, the process stayed alive and the
service stayed down until it was killed by hand.

- cli.ts: the signal handler arms an unref'd 10s timer that force-exits if
  server.stop() hangs.
- self-update.sh (launchd-daemon): wait up to 30s for the server pid to exit,
  then SIGKILL it. tmux sessions live outside the server and survive.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Codeman maintainer <noreply@anthropic.com>
2026-09-24 01:35:27 +02:00
DevvynandClaude Opus 5.5 dd230b0b6e fix(test): strip every inherited CODEMAN_* var and move quick-start off 3099 (#479)
Split out of #476. A Docker Compose deployment exports CODEMAN_CASES_PATH,
which bypasses the temp HOME, so route tests wrote into the real case root.


Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 01:35:23 +02:00
github-actions[bot]Claude Opus 5.5github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
120d780267 chore: version packages (#474)
* chore: version packages

* docs: sync CLAUDE.md version to 1.32.1

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
2026-09-23 12:34:33 +02:00
Codeman maintainer 0af925fe82 Merge remote-tracking branch 'origin/master' into land/1.32.1
# Conflicts:
#	CLAUDE.md
#	docs/architecture-invariants.md
2026-09-23 12:25:58 +02:00
Codeman maintainer b404dacfde chore: changeset for the merge-time fixes and thanks
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:41:11 +02:00
Codeman maintainer cbd1fa639d fix(tmux): merge-time fixes for the exited-agent report (#466)
- docs/wiki/The-Dashboard.md: the tab-appearance table gains the exited
  state (muted dot plus an `exited (137)` badge) and explains the bare
  `exited` variant.
- The detailed sidebar and rail no longer pair the muted dot with an "idle"
  pill: an exited session's pill reads "exited" (neutral styling) and its
  since stamp measures from the observed exit. This is a label override on
  the row model, not a new state, so SESSION_ACTIVITY_RANK and the home
  screen order are untouched, and a pending alert still keeps its own pill.
  The row signature includes the flag so the incremental path repaints it.
- The exited badge is aria-hidden like its sibling badges, and the exit is
  appended to the tab's aria-label in both render paths through one helper.
- test/tmux-manager.test.ts re-adds the junk-trailing-field parser case
  against parsePaneRows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:27 +02:00
Codeman maintainer 43d4be8eeb fix(session): merge-time fixes for the dead-pane resume pin (#467)
- test/setup.ts strips CLAUDE_CONFIG_DIR (pinned in test-env-isolation), so
  transcript-fixture tests such as session-custom-model-restart no longer go
  red on a machine that exports it for a separate Claude account (#255).
- The vanished-tmux-session branch of _setupOrAttachMuxSession() relaunches
  the CLI through createSession() just like a failed respawn, so it now takes
  the same resume pin. A genuinely new session is unaffected.
- After a dead-pane respawn of a fallback-chain CLI, _claudeSessionId names
  the conversation the walk actually pinned instead of the chain tail, which
  the walk may have passed over for lack of a transcript.
- _claudeConfigDir() trims the override like claudeProjectsDir() does.
- The remote-reattach test is labelled as documentation, since the pin
  builder's own remote guard would make it pass either way.
- CLAUDE.md: the create-path pin persists through toState() as
  resumeSessionId, and the end of the walk adds no pin rather than clearing
  the launch seed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:27 +02:00
Codeman maintainer f4d1ee8027 fix(terminal): merge-time fixes for dropped-output recovery (#470)
- The TERMINAL DROP crash-trail line moves behind the scheduler's debounce
  guard, so it is written once per window rather than once per dropped
  frame. At the server's 8ms batching, one second of drops evicted the whole
  50-entry trail, including the recovery lines that explain it.
- A refresh that failed at the capture fetch deadline now returns
  'deadline', and the scheduler does not retry it: that is a stalled link,
  not contention, and each retry was another ?full=1 capture waiting out a
  deadline of up to two minutes. The early-return retries are unchanged.
  CLAUDE.md and the code comments no longer claim every skip reason is
  transient contention.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer 7a30a31430 fix(terminal): merge-time fixes for the silent-failure paths (#431)
- While another device holds the pane width (_paneWidthRefused), a resize
  now asks for the container's width without applying it locally
  (_geometryForResizeRequest: rows follow the container, columns stay at
  the PTY's). Fitting first re-wrapped the whole buffer to the container
  and back on every 30s mobile retry, and throttledResize ran the
  scrollback clear for a resize that brings no redraw. selectSession
  clears the flag, since it belongs to the previous pane. New unit tests
  run the real mixin against a fake terminal and fail without the fix.
- Session seeds _ptyCols/_ptyRows at spawn (_notePtySpawnGeometry), so a
  reattached pane reports its tmux window's real size through ptyGeometry.
- Session.resize's declined-branch comment names ptyGeometry, not the
  deleted ptyCols/ptyRows getters.
- Delete the dead terminalGeometryAgrees() and its window export.
- test/xterm-private-api.test.ts header: it pins the exact locked version,
  so any bump fails, not only a major.
- The main-terminal fit sweep also matches fitAddon?.fit?.(), and
  CLAUDE.md names the modules it actually covers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer fdfcc15c10 docs(docker): merge-time note for the gh/az sign-in in multi-user mode (#472)
Clone Repo clears the credential helpers for a non-admin, but a non-admin's
Docker case with credential seeding on still receives a copy of the server
account's gh/az sign-in when the agent-image switches are on, the same as
the Claude and Codex credentials. Say so in the multi-user notes so the docs
do not read as a stronger guarantee than they are.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer 697b05b118 fix(docker): merge-time fixes for Update-Codeman.sh (#465)
- Remove exactly the codeman-node-modules/codeman-dist volumes by Compose
  label after a plain `down`, instead of `down --volumes` (which also takes
  any volume an override file declares while the message named two).
  `down --volumes` remains only as a warned fallback when the project name
  cannot be resolved.
- Report a failing first `docker compose config --format json` call with a
  clear error instead of exiting silently under `set -e`.
- Filter empty label lines in the collision guard so an unlabelled container
  cannot hide a real collision; name the moved-checkout exit in its error.
- Comments no longer cite a guard or incident in Start-Codeman.sh that does
  not exist; the README states the real gap (a Node base-image bump leaves
  codeman-node-modules stale because the lockfile did not move).
- docs: Update-Codeman.sh in the docker-self-update.md short-version table
  and a mention in docker-compose.md; "Major updates" moved under "Updating"
  in docker/README.md.
- test: smoke test covers the new sequence, the config failure and the
  empty-line case; quiet stdio; @fileoverview names the fourth concern.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer 0462a5d5a0 fix(approvals): merge-time fixes for the watching badge (#473)
- session.ts: a pane capture that fails now CLEARS the watching label
  (and emits watchingChanged so pages drop the badge) instead of keeping
  the last one, so a failed capture degrades toward an alert rather than
  pre-acknowledging the next real idle prompt. Test updated; invariant
  noted in architecture-invariants.
- approvals-ui.js: the header bell counts only unacknowledged items
  (pendingApprovalsCount), matching codeman tui's pendingApprovalCount();
  pinned in watching-no-alert.test.ts.
- mobile-overview.js: move the orphaned "Pill copy per state" JSDoc back
  onto MOBILE_OVERVIEW_PILL_LABEL.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer da6fa663e7 fix(terminal): merge-time fixes for the copy gutter strip (#469)
- stock.ts: claude is no longer the only entry declaring transcriptGutter;
  codex declares it too.
- architecture-invariants: the strip applies when the session's CLI declares
  a margin (not detection), and a note that it keys on the session's launch
  mode, not on what is running in the pane (a claude pane dropped to a shell
  still loses up to two columns; copyStripMargin is the escape hatch).
- render-index-html test: the gutter map is injected for a solo
  /session/:id render as well.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer 10f87428c3 fix(cli-registry): merge-time fixes for the run-button accents (#463)
- mobile.css: gemini and antigravity run/gear rules get `!important` like
  pi/omp/grok/deepseek, so the gear half no longer keeps the skin accent
  while the body takes the mode colour (two-tone button on the default skin).
- test/skin-themes.test.ts: static guard that every run mode with a base
  `.btn-toolbar.btn-run.mode-<id>` rule also has a resting rule inside the
  `html:not([data-skin="og"])` block; ids are derived from the stylesheet.
- stock.ts: grok's accent comment names zinc-300 (border/badge colour);
  gemini's accent is #8ab4f8 to match its tab badge and run-mode dot, noted
  as the one exception to the border-colour method.
- types.ts: "(below)" -> "(above)".
- docs/cli-registry.md, CLAUDE.md: `accent` is now measured, not transcribed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer 13e652e43f chore: changesets for the 1.32.1 batch
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:40:03 +02:00
Codeman maintainer 2afb1c2c2e docs: trim CLAUDE.md from 265 KB to 142 KB, detail moved to architecture-invariants
CLAUDE.md loads into every session, and its Architecture section had grown
feature write-ups (history, measurements, rationale) that belong in
docs/architecture-invariants.md per the file's own header. Each long block
now keeps what the feature is, where it lives, its setting/default and the
rules that prevent real bugs, and links to its invariants section. Everything
removed was moved there: 29 new sections, extra facts appended to the
existing ones.

Also: hard-coded counts (SSE events, route handlers, module/file counts,
device profiles) replaced by pointers to the source of truth, and the
Debugging commands fixed to use the codeman tmux socket and HTTPS for prod.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 11:32:34 +02:00
Codeman maintainer 2fb744f865 Merge pull request #470 from rounakdatta/fix/dropped-output-recovery
fix(terminal): recover a dropped output frame, do not merely schedule it
2026-09-23 11:32:14 +02:00
Codeman maintainer de4b1db490 Merge pull request #431 from rounakdatta/feat/mobile-terminal-resilience
fix(terminal): four silent-failure paths — renderer freeze, replay race, reconnect gap, unbounded fetches
2026-09-23 11:32:14 +02:00
Codeman maintainer 8536aaef7b Merge pull request #473 from irisitymichaelgrundberg/feat/session-watching-badge
feat(approvals): let a session watching its own background work keep quiet (#468)

# Conflicts:
#	src/config/cli-registry/stock.ts
2026-09-23 11:32:11 +02:00
Codeman maintainer 94b093b617 Merge pull request #469 from irisitymichaelgrundberg/feat/copy-dedent-pane-margin
feat(terminal): take the transcript gutter off a copy, at the width the CLI declares
2026-09-23 11:32:02 +02:00
Codeman maintainer 6a01412af9 Merge pull request #466 from irisitymichaelgrundberg/feat/pane-exit-reporting
feat(tmux): report that a pane's agent has exited (#446, part 1)
2026-09-23 11:32:02 +02:00
Codeman maintainer bc04b6457e Merge pull request #467 from irisitymichaelgrundberg/fix/respawn-session-id-collision
fix(session): resume the conversation when respawning a dead pane
2026-09-23 11:32:01 +02:00
Codeman maintainer 5b5e932ec4 Merge pull request #465 from opticon454/chore/docker-major-update-script
chore(docker): add Update-Codeman.sh for scripted major-update rebuilds

# Conflicts:
#	docker/README.md
2026-09-23 11:32:00 +02:00
Codeman maintainer f1dfbcdd65 Merge pull request #472 from opticon454/feature/git-host-auth-clis
feat(docker): opt-in gh + az CLIs with git credential helpers so Clone Repo and Docker cases can reach private repos
2026-09-23 11:31:53 +02:00
Codeman maintainer 233af33dac Merge pull request #463 from opticon454/fix/cli-accent-colours
fix(cli-registry): correct accent colours, and a real gemini/antigravity/omp rendering bug
2026-09-23 11:31:52 +02:00
Codeman maintainer 993e5e021c Merge pull request #471 from DodgyBadger/fix/mobile-blank-long-press
fix(mobile): swallow blank-space terminal long presses
2026-09-23 11:31:52 +02:00
DevvynandClaude Opus 5.5 02e40f506b fix(docker): gate gh/az seeding on its switch; no shared git sign-in for non-admin clones
Addresses the review on #472.

- CRED_STORES: `.config/gh` and `.azure` now carry `enabledByEnv`
  (CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ), and resolveDockerCredentialArtifacts
  skips a store unless that variable is exactly `1`, read at container
  create. A host that merely has ~/.config/gh/hosts.yml or a plaintext MSAL
  cache no longer copies them into every case container. Tests: the default
  environment seeds neither even with the files present, and each store
  follows only its own switch.
- Multi-user mode: a non-admin's Clone Repo clone and preflight run with
  `git -c credential.helper=` (GIT_NO_CREDENTIAL_HELPERS, placed before the
  subcommand), so the server account's helpers are never lent to them.
  Verified against a real private repo that it also clears the URL-scoped
  credential.<url>.helper entries, and that public clones still work.
  Tests: the argv in test/git-clone.test.ts, and the route decision
  (non-admin cleared; admin and single-user kept) in
  test/routes/case-clone-credential-helpers.test.ts.
- Docs: recreate the case container to pick up seeds (docker/README.md,
  Docker-Cases wiki, docker-cases.md); the multi-user behaviour in
  docker/README.md and security-architecture.md; "functionally unchanged"
  instead of "unchanged" for an image built with both switches off
  (server.Dockerfile comment, README, changeset).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw
2026-09-23 14:44:26 +08:00
Michael GrundbergandClaude Opus 5 e558264977 chore: leave the changeset to the maintainer
CONTRIBUTING says releases are handled by the maintainer via changesets after
merge, and every `.changeset/*.md` on master was written by him or by the
release bot — including the ones covering other people's pull requests. The
summary this file carried moves to the pull-request description, where it is
the maintainer's to reuse or rewrite.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 08:28:25 +02:00
Michael GrundbergandClaude Opus 5 9a48c43aa1 docs(watching): a restart is not a gap, and here is the measurement
Claimed after a manual test that a session comes back from a server restart
without its badge until it next produces output. Measured instead of assumed,
and it is wrong: a codex session with a background terminal still running had
its label back within about 20 seconds of the restart, with no input from
anyone. Reconciliation re-attaches the pane, the attach repaint carries the
composer glyph, the idle confirmation arms on it, and the probe re-reads the
label — the ordinary path, doing the ordinary thing.

What produced the false claim was a session whose monitor had simply expired
while it sat there. Its footer carries no chip, so `watching: null` was the
right answer and there was nothing missing to restore.

Recorded at the field and in the invariants, because the shape of this invites
exactly one wrong fix: a polling timer to keep a value fresh that the pane
already refreshes by itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 08:00:16 +02:00
DevvynandClaude Opus 5.5 5cf5a45438 feat(docker): opt-in gh + az CLIs with git credential helpers for private repos
Add Case -> Clone Repo could only reach public repositories in the Docker
deployment. This lets a deployment opt in to the GitHub CLI and the Azure
CLI (+ azure-devops extension) as git credential helpers. Codeman itself
still collects no credentials.

- server.Dockerfile / agent.Dockerfile: CODEMAN_INSTALL_GH /
  CODEMAN_INSTALL_AZ build args (0 or 1, default 0; anything else stops the
  build). Off leaves no apt repository, package, extension, helper script
  or credential entry, so a default build is unchanged. On installs from
  the vendors' apt repositories and configures system gitconfig helpers:
  github.com / gist.github.com -> `gh auth git-credential`, dev.azure.com /
  *.visualstudio.com -> new docker/git-credential-azure-cli (an Entra ID
  token from `az account get-access-token`, or AZURE_DEVOPS_EXT_PAT).
  A helper whose CLI is not signed in prints nothing, so a private clone
  still fails fast.
- The extension lives in AZURE_EXTENSION_DIR outside HOME
  (/opt/codeman-az-extensions, runtime-owned; /opt/az-extensions, gid-0
  group-writable in the agent image).
- Hosts turn them on in docker-compose.override.yml: `build: args:` for the
  server image, `environment:` CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ for the
  agent image. build-agent-image.mjs and the in-app auto-build share one
  env -> ARG table (pinned by the parity test) and pass nothing when unset.
  docker-compose.yaml is untouched; .env.example only gains a comment, so
  the self-updater's environment gate sees no new keys.
- Docker cases seed the gh sign-in (~/.config/gh/hosts.yml, config.yml) and
  the az sign-in files from ~/.azure per file, read-only, like pi/grok.
- The Clone Repo AUTH_REQUIRED message says how to sign the server's git
  in instead of claiming private repositories cannot be cloned.
- Docs: docker/README.md "Private repositories", docker-compose.md,
  docker-cases.md, the Quick-Start / Core-Concepts / Docker-Cases wiki
  pages, security-architecture.md, architecture-invariants.md, changeset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw
2026-09-23 08:57:09 +08:00
DodgyBadger 110c4696ad fix(mobile): swallow blank-space long presses 2026-09-22 17:09:01 +00:00
Michael GrundbergandClaude Opus 5 ac6236b268 fix(terminal): clean a copy once, and reach every pane that copies
Review fixes for #469.

The Ctrl+C branch cleaned the selection to decide whether to copy and then
passed that cleaned string to copyTerminalSelection(), which cleans again. The
trailing trim is a fixed point, so that was safe until this PR; the margin
strip is not, because it takes the lesser of the declared width and the run
every line shares, so a second pass takes up to `margin` columns more. The
branch now gates on the cleaned string and hands the raw one on. Verified in
chromium with a real drag, a real Ctrl+C and a real clipboard read on a live
claude pane: an on-screen `      fix(terminal): trim it` reaches the clipboard
as `    fix(terminal): trim it`, and reverting the branch reproduces the
reported `  fix(terminal): trim it`.

Pane B of a split resolves its own width. `_cliGutterColumns()` and
`_normalisedSelectionRange()` take the session and the terminal to read,
defaulting to the primary pane's, so Pane B looks its own run mode up instead
of keeping a margin Pane A drops on the same keystroke. Verified live with two
claude panes open side by side.

A detached session window (`/session/:id`) receives the gutter map. The
injection sat inside the block that skips the run menu's payloads for a solo
window, so the toggle worked in the main window and did nothing in the popup on
the same device. It needs no availability probe, so it moved below that block
and the solo window still carries none of the payloads it skipped before.

The settings description said the width is measured and named Codex as exempt.
Nothing is measured, and Codex is one of the two panes that are stripped.
docs/wiki/Settings-Reference.md gains the row every Terminal and Input toggle
carries. CLAUDE.md no longer says the clean touches trailing runs "and nothing
else" one sentence before the leading-margin rule, and both it and
docs/architecture-invariants.md record that the strip is not idempotent.

Two round-trip tests run on a mode that declares a gutter, which the existing
copyTerminalSelection cases could not, since they all use the harness default
mode that declares none. The Ctrl+C branch itself is pinned at the source,
because it lives inside initTerminal's attachCustomKeyEventHandler closure over
a real xterm the vm harness cannot build. Both pins fail on the reintroduced
bug. Gate: 7865 passed, 0 failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 19:02:55 +02:00
Rounak DattaandClaude Opus 5 00f022ccf8 fix(terminal): recover a dropped output frame, do not merely schedule it
`_onSessionTerminal` drops an incoming frame once the app-owned render queues
already hold 128KB. That is the right call — the alternative is an unbounded
backlog — but a hole in a TUI byte stream is a desynced cursor, and a desynced
cursor is muffled text (#464). The drop was only half of it.

The recovery was a fire-and-forget timer: it nulled its own handle and then
called `_onSessionNeedsRefresh()`, which opens with four early returns. Two of
them — a buffer load already in flight, a refresh already owning this session —
are MOST likely to be true during exactly the output burst that caused the
drop. So the recovery was skipped precisely when it was needed, with nothing
left to retry it, and the dropped bytes were never replayed.

`_onSessionNeedsRefresh` reports whether it actually repainted now, and
`_scheduleDroppedOutputRecovery` re-arms while it has not. Bounded by
`DROP_RECOVERY_MAX_ATTEMPTS`, because every reason the refresh can be skipped is
transient contention that clears in seconds and a permanently failing refresh
must not become a loop against the API; giving up at the cap leaves exactly what
the old code left, so the floor is no worse than before. The same 2s debounce
still collapses a burst of drops into one attempt.

This is the principle Ark0N established reviewing #431 for the WebSocket
output-gap marker — only a repaint that actually happened settles the recovery —
applied to the one recovery path that still trusted a timer having fired.

The retry decision is a pure function in constants.js so the gate can reach it,
and the scheduler itself is driven from app.js under a fake clock. The retry
case and the no-retry case only pin the fix AS A PAIR: either alone passes
against something wrong, one against the old fire-and-forget timer and the other
against retrying forever. Checked by reverting app.js to the old shape, where
three of the twelve fail.

Two harness details that would otherwise have made the tests lie. The vm context
baked in the real `setTimeout`, so `vi.useFakeTimers()` could not reach the
scheduler and every case reported zero calls; it delegates lazily now. And
app.js reached `CodemanDroppedOutput` as a bare global, which resolves in a
browser but not in the vm — worth fixing beyond the test, because that call sits
inside a timer where a ReferenceError is swallowed and would take the recovery
with it. It reads through `window.` like terminal-ui.js does with its own
constants.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 21:34:03 +05:30
Michael GrundbergandClaude Opus 5 b13672596f fix(watching): tell the page when the badge goes away
Reported from a manual test: a Codex session went on showing the watching badge
after its background terminal had finished. The server was right and the page
was stale — `Session.watching` changes while the session's status does not, and
nothing broadcast it.

The label is usually SET on the idle transition, which broadcasts anyway, so the
badge always appeared correctly. It CLEARS when the work ends, and a CLI can end
background work without taking a turn: codex repaints its background-terminal
row away and stays idle, so `_confirmIdle()` concludes without emitting `idle`
(that emit is guarded by `wasWorking || isInitialReady`) and no other event
fires. Every open page kept drawing a badge the server had already dropped.

`_readWatching()` now emits `watchingChanged` when, and only when, the label
really changes, and the wiring pushes the session state on it. No new SSE event:
the badge reads off the session payload every surface already has.

A/B measured on an isolated beta with the page loaded and then left untouched.
Without this commit the server dropped the label at t+50s and the page still
showed the badge at t+100s; with it, page and server cleared in the same
ten-second window.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 17:59:45 +02:00
Michael GrundbergandClaude Opus 5 8d45b92eba docs(codex): record that a sub-agent leaves no row to read
Measured on the same beta, codex-cli 0.154.0: a sub-agent started without
waiting outlives the turn exactly as a background terminal does — the sandboxed
process was still running — and codex shows nothing for it. The last rows of the
pane are the composer and the status line, and `Sub-agents running` belongs to
the on-demand `/subagents` panel rather than to the row above the composer.

So there is no second codex label to add. A codex session waiting on a sub-agent
reads as plainly idle, which misfiles nothing (codex raises no idle prompts) and
simply leaves that one kind of quiet unexplained until codex pins a row for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 17:35:56 +02:00
Michael GrundbergandClaude Opus 5 05c788ce9d fix(watching): close the review findings on the label and its window
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief)
found the trust boundary weaker than the comments around it claimed. Eleven
findings, all applied.

The two blockers were both about who can write the row the label is read from.
Claude's window covered two rows, and the second one is the status line, whose
command a session running with permissions bypassed can write into its own
`.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of
its own and silence its own idle alert. The default window is one row now, which
is the footer and nothing else, and the constant says why. Separately, the label
reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML,
which is an injection sink for any config-supplied pattern whose capture group is
permissive; it goes through escapeHtml() like every other untrusted string in
that file.

The Codex entry could not be fixed the same way, and now says so. Its row is
third from the bottom only while a terminal runs; with none running that slot
holds the last row of the transcript, so matching the complete row (with the
`/stop to close` tail, window narrowed to three) raises the bar without closing
it. What contains it is `hooks: 'none'`: no hook event from a codex session
reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an
alert. The registry comment, `docs/cli-registry.md` and the test all state that
rather than claiming a guarantee the code does not have.

Also from the review: the TUI header badge no longer counts an acknowledged
item, which was the same gate the classifier fix already went through and was
wrong for human acknowledgement too; the TUI approval card reads the quiet
reason and drops to a new `info` tone instead of asking for a reply; the badge
carries an aria-label, because the phone it was built for has no hover target;
the schema refuses `watchingLines` without a `watchingLine`; and the pattern and
its window are resolved together rather than one memoized and one not.

Documentation moved with it. The mechanism now lives in
`docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer,
`docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped
buzzing, and both that page and the changeset name the limitation neither did
before: a question asked in plain prose is not a dialog, so it is silenced along
with the false alarms while background work runs.

Verified live again after the narrowing, on an isolated beta: a Claude session
reported `1 monitor` and took its idle prompt acknowledged, and a Codex session
reported `1 background terminal` against the full-row anchor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 17:11:33 +02:00
Michael GrundbergandClaude Opus 5 9c286eeddf fix(session): persist an exit retraction, and let tests reach the watcher
Ten findings from a two-model review of this branch. Both reviewers cleared the
detection logic itself; everything here is a gap around it.

A route that starts a command in a pane now PERSISTS as well as broadcasts.
`/interactive` and `/shell` did neither before, and the pane-exit watcher cannot
cover for them: its next tick finds `paneExit` already cleared in memory,
reports no change and writes nothing, so `state.json` kept saying the agent had
exited for as long as the session stayed quiet. Nothing reads that record for a
decision yet, which is exactly why it had to be fixed now — part 2 is designed
to read it. The `clearPaneExitForNewPane()` docstring claimed its callers
already persisted; that claim was false for these two, and now says what the
caller owes instead.

The watcher's four guards were unreachable by any test. `refreshPaneExits()`
opened with `if (IS_TEST_MODE) return;`, so the read gate, the in-flight
suppression, the generation counter and the empty-read rule could each be
deleted with the whole suite green. The tmux call moves into `readPaneRows()`,
which a test subclass overrides — the shape `runRemoteReconnectTick` already
uses in this file for the same reason — and the test-mode gate moves with it, so
what a test cannot do is spawn a process rather than exercise the bookkeeping.
Each of the four guards now has a test that fails when it is deleted.

The muted status dot turned out to be a specificity fight on three surfaces, not
two. `.tab-status.error` was not excluded, so a session whose agent exited and
whose PTY-exit breaker then tripped lost its red dot to the mute — the state the
browser answers with a "restart it?" confirm, and a needs-you colour by the same
argument that protects the two alert classes. And mobile.css gives a `busy` dot
a 9px size and a green glow with `!important`, while `status` stays `busy` for a
pane whose agent died mid-turn, so a phone rendered a grey dot still wearing the
green halo beside a badge reading "exited". Both measured against the real
stylesheets, both now excluded, and the CSS test reads mobile.css too instead of
being structurally blind to half the problem.

Six comments said things that were not true. Two named the stats collector as
what replaces a restored reading, which is the opposite of the design. The
interval constant argued that 2000 ms keeps a read inside a tick, when the
5000 ms exec timeout means it cannot — which is why the in-flight guard exists.
`MuxSession.discovered` did not say the flag is permanent, though `saveSessions()`
serializes it. The empty-read docstring claimed a distinction that `|| true`
makes impossible. The invariants doc promised more than its drift test delivers.
And CLAUDE.md had no pointer at all, leaving its two hardest prohibitions
("never set `status: 'error'`", "never null the pid") only in the file it is
meant to route people to.

Refs Ark0N/Codeman#446.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 15:24:09 +02:00
Rounak DattaandClaude Opus 5 e1e7dc5bd8 fix(terminal): Ark0N's read of the #464 geometry work
Five items, two of which he could only see by running it, plus six smaller
ones. Taking the two blockers first, because both were wrong in ways the
existing tests could not catch.

**Adopting the PTY's rows put the CLI's input line off-screen.** A phone that
took a desktop's 43 rows into a viewport with room for 18 painted an
`.xterm-screen` far taller than its container; xterm's own viewport then had
nothing to scroll, so the bottom of the frame sat below the container with no
gesture able to reach it. Output visible, typing invisible, for as long as the
desktop kept the claim hot. `reconcilePtyGeometry` adopts COLUMNS ONLY now:
width is the axis Ink's wrap and `eraseLines` arithmetic depend on, and keeping
the local row count keeps the composer at the bottom of a viewport that
scrolls. Measured at his geometry — a 360x300 container against a 198x43 pane
now keeps 13 rows, takes 198 columns, paints 202px into a 210px container, and
the input line is inside the box.

**`capture-geometry-retry.browser.test.ts` failed, and CI could not see it**
because the file is in `BROWSER_TEST_GLOBS`. Its premise WAS the clamp —
`getTerminalDimensions()` floored while `fitAddon.fit()` did not — which this
work removes at the source, so it can never hold again at any viewport. The
case survives on its own terms: a pane already drawing at the requested size
must not be replayed. Its premise is now the #464 invariant itself, that the
floored report and the terminal agree, which is a stronger guard because the
clamp coming back fails it here rather than silently restoring the replay loop.
The helper docblock that repeated the old premise is corrected too.

**A session with no pane reported 120x40 and the client adopted it.**
`resize()` writes `_ptyCols`/`_ptyRows` only when `ptyProcess` is set and
nothing seeds them from the spawn geometry, so a dead-pane session still held
the constructor defaults — clicking that tab resized the browser terminal to
120x40 and, on anything narrower, claimed another device owned the pane when
none existed. `Session.ptyGeometry` returns null without a pane, the HTTP route
answers `{}` and the socket sends no frame at all. The raw `ptyCols`/`ptyRows`
getters are deleted rather than left available to be misused again.

**The 40-column floor clipped the pane with nothing able to reach it.** The
affordance keyed on a PTY mismatch, and the floor produces no mismatch — xterm
and the PTY agree throughout, the terminal is simply wider than the box. It
keys on what does not FIT now, MEASURED (`.xterm-screen` against the container,
on the next frame, because the screen takes its width with the render) rather
than derived from cell arithmetic. Measured at 360px: font 24 applies 40
columns and paints 560px, and all 200px of the overhang is reachable.
`.pty-oversized` is renamed `.term-overflows-x`, because after this the old
name describes only one of the two causes.

**"Scroll sideways" did not work on touch for the sessions it targets.**
`touch-action: pan-x` is cancelled before it starts by the `preventDefault()`
`touchstart` calls on every 'content' tap. The terminal's own touchmove handler
pans the container now, with the axis locked once per gesture so a diagonal
cannot pan and scroll at once, and the CSS grants no `touch-action` at all —
handing the browser a pan AS WELL would move the pane twice for one finger on
the taps where that preventDefault does not run. Measured under real touch
dispatch: a 140px swipe reaches `scrollLeft` 140 where it reached 0 before, the
buffer does not move with it, and a vertical swipe still scrolls the scrollback.

Three defects in the above, found while checking it rather than by being told:

- `canPanHorizontally` first tested `scrollWidth > clientWidth` alone, which is
  true of a container that is not a scroller — a sideways swipe would have
  locked the axis, done nothing, AND suppressed the vertical scroll it should
  have been. Gated on the class as well.
- The notice advised scrolling sideways whenever the PTY was wider, including
  when it still fitted and nothing scrolled. It is gated on measured overflow,
  and on a comparison against the width this container WOULD request rather
  than the one it currently holds — once adopted those are equal, so the second
  question answers itself false while the condition is still true.
- `_syncTerminalOverflowAffordance` could throw out of `document.getElementById`
  before reaching its try block. It runs off every geometry change, so a
  cosmetic affordance could have taken the resize down with it.

The smaller items:

- `docs/architecture-invariants.md` no longer explains the equality guard as a
  clamp signature; it records what the clamp used to do and why it cannot any
  more. Edited by hand — that file is outside the Prettier glob, and letting
  Prettier near it rewrote eleven unrelated emphasis markers.
- `throttledResize`'s HTTP fallback reads the reply. It is the path where a
  declined resize is least likely to be noticed, because no socket means no
  `{"t":"zc"}` frame either.
- The changeset covers the whole release: the geometry work, the queued replay
  clear, the renderer watchdog, the body-covering fetch deadline, the WebSocket
  output-gap reconcile, the build-generated service-worker precache and
  per-build cache key, and the crash-trail hygiene.
- `@xterm/headless` is declared in the root devDependencies instead of being
  reached through workspace hoisting.
- The output-gap marker is cleared after any response arrives, not only when
  the capture was non-empty: a server that answers with an empty capture HAS
  reconciled us, and leaving the marker set refetched on every reconnect.
- `e587d845`'s message claimed a test asserted the failed-load copy against the
  built asset. It did not — that assertion lived in a probe deleted with the
  other scratch scripts, so the claim was false when it was written. There is a
  real test now, and it reads the source rather than `dist/`, because `dist/` is
  not committed and a test that skips when it is absent would pass for the wrong
  reason in CI.

`Session.ptyGeometry` gets behavioural coverage against the real class in
`session-resize-arbitration.test.ts` rather than a source guard, including the
contrast — a pane that does exist still reports, and still follows a resize —
so "always null" would fail it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 18:53:46 +05:30
Michael GrundbergandClaude Opus 5 ce80b7a212 feat(terminal): take the transcript gutter off a copy, at the width the CLI declares
Copying a paragraph out of a Claude Code or Codex pane puts that pane's own
two-column transcript gutter on the clipboard, so every pasted line arrives
indented. #451 shipped the trailing half of the copy clean and left the leading
half out, because deriving the width from the selection fires on 73% of ordinary
indented text and cannot tell a margin from content.

The width is DECLARED rather than derived. `capabilities.transcriptGutter` on
the CLI registry is a bounded integer; claude and codex each declare 2, measured
on live panes, and no other stock entry declares any, so a CLI whose transcript
layout nobody has measured is never touched. The server publishes the map as
`window.__codemanTranscriptGutter`, built by filtering `enabledClis()` on the
capability rather than by listing ids, and `_activeCliGutterColumns()` looks the
active session's mode up in it. The copy path reads no terminal buffer at all.

The declared width is a CEILING, not the answer: `clean()` strips the lesser of
it and the run every selected line shares. A block can therefore only shift as a
unit, the structure inside a selection survives by construction, and a selection
reaching column 0 loses nothing. That is what keeps a `git log` body at its own
four-space indent inside an agent's two-column gutter.

Codex was measured separately, because it renders nothing like Claude: it draws
boxes narrower than the pane and pushes its transcript into ordinary scrollback.
On a live 0.154.0 answer its `•`/`›`/`⚠` markers sit in the gutter, prose
continuations sit at 2, and a nested YAML block the model wrote rendered at
2/4/6/8 for its own 0/2/4/6. Replayed at 100, 120, 160, 198, 235 and 282 columns
its indents were 0, 2, 4, 6 and 8 at every one, never 1. Copying that YAML out
of a live Codex pane now yields 0/2/4/6: gutter gone, nesting intact.

Two derived versions were built and measured first, and both are recorded in the
code because both looked correct:

- Painted trailing padding — a full-screen TUI writes real spaces across the
  unused part of a row, a shell leaves them never-written for xterm to trim —
  has no false positives and never over-stripped. It is also a function of pane
  WIDTH: the padding exists only while a rendered line stops short of the CLI's
  own layout width, and Claude's prose wraps to fill it. Dragging the same two
  prose rows of one live transcript at five window sizes, the share of padded
  rows ran 44%, 6%, 6%, 7% and 87% at 123, 160, 198, 235 and 298 columns, so the
  strip silently did nothing at every ordinary size while a corpus captured
  entirely at 282 columns said it worked.
- Taking the narrowest indent on the rows around the selection fires at every
  width and over-strips about 1% of selections, because a file listing inside
  the transcript can be the narrowest thing on screen.

Measured over 1,392,281 selections — every 1, 2, 3, 5, 10 and 20-row window of
real Claude screens replayed from live PTY streams at 100, 120, 160, 198, 235
and 282 columns — the declared width over-strips none, breaks no relative indent
and alters no text, and serves 100% of the selections whose own indent covers
the gutter. Verified end to end in a browser with a real mouse drag and a real
Ctrl+C: Claude and Codex panes paste flush at 123, 198 and 298 columns, a shell
pane is untouched at every one.

The strip sits behind `copyStripMargin` (App Settings, Selection & clipboard),
per-device and default ON: a display key, absent from the .strict()
SettingsUpdateSchema, read as `!== false` because the desktop branch of
getDefaultSettings() returns {}. The toggle is checked before the map.

Two review findings from #451, handled:

- The mid-row flag governs ONE line now. `range.start.x > 0` excludes only the
  first selected line, the one whose margin the mousedown genuinely cut off, so
  the same three rows no longer produce three different clipboard results.
- The reversed-drag finding does not reproduce on the pinned xterm.
  `getSelectionPosition()` reads `_selectionService.selectionStart`, whose
  getter returns `SelectionModel.finalSelectionStart`, and that swaps the pair
  when `areSelectionValuesReversed()` says so. A real upward mouse drag through
  chromium against xterm 6.0 reports the same range as the downward drag.
  `_normalisedSelectionRange()` keeps the ordering as a guard, because the model
  one layer down exposes the unnormalised fields under the same two names.

Tests: test/terminal-copy-clean.test.ts (64, up from 31), plus the injected
script stripped in test/server-index-title.test.ts. Every guard is pinned:
removing any one of seven reds at least one test, including declaring the wrong
gutter width. Full suite green, 7,861 passed, 0 failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 13:33:50 +02:00
Rounak DattaandClaude Opus 5 e587d84590 fix(terminal): make the failed-load notice fit the narrowest terminal
A third pass in a real browser, at the widths this app actually renders at.

The notice a failed history load writes into the blanked pane was one
70-character sentence. At 430px that exactly filled the line; at 320px it
wrapped and left a lone '.' on a line of its own. The floor this app will
render at is 40 columns — reachable today by raising the font on a phone — so
the notice is three lines now, none over 25 columns, one fact each: what
failed, that the session is still alive, and what to do.

It says RELOAD rather than "reopen the tab" because `selectSession`
early-returns when the session is already active, so clicking the tab you are
already on retries nothing. The earlier wording named no next step at all,
which left a mostly-empty terminal and no way out of it.

CLAUDE.md no longer cites "758px reachable to the right" as evidence: that
figure is a property of the test content, not of the fix, and the file's value
is that a reader can trust a claim without re-deriving it. What is pinned
instead is the invariant that survives any content — the full pane width is
reachable, and removing the class returns scrollLeft to 0, so a resolved
mismatch cannot leave the pane parked off-screen.

Verified at 430, 360 and 320px against the shipped bundle, with the test
asserting the built asset carries the copy so an edit that never reached the
build fails rather than passing on the source's wording.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 13:57:40 +05:30
Rounak DattaandClaude Opus 5 d9fa9ba1eb test(terminal): follow the existing suites to the one geometry owner
The gate caught fourteen failures the focused tests could not: every harness
that builds a partial app out of cherry-picked mixin methods, and every source
guard that named `fitAddon.fit()` by hand.

Most are wiring — `syncTerminalGeometry`, `_refitAfterCellSizeChange` and
`_resizeTerminalTo` added to the fakes so the real chain runs rather than a
stub of it. `file-browser-search` is the one that shows why it matters: without
the method on the fake, selectSession's unconditional call threw into its own
catch and every later assertion in the file measured a load that never
happened.

Two are not wiring.

`detached-session-pane-sizing` pinned the behaviour this change deliberately
reverses. It asserted the LOCAL fit still runs for a session owned by its own
window — "withhold the send, never the reflow" — so the assertion is restated
rather than patched, with the reason beside it and in the file's docblock: a
reflow the PTY is never told about leaves this xterm rendering a CLI's frames
against a shape that does not exist, and the popup that owns the pane is
drawing for its own width regardless. The old rule bought a garbled frame, not
a correct one.

`mobile-prompt-composer` sliced `_cleanupSessionData` as a fixed 1200-character
window, so the assertion depended on how much unrelated code sat above the line
it cared about. It reads the whole method now.

`terminal-scroll-intent` records `syncTerminalGeometry` rather than `fit`,
under its own name: recording a bare fit there would name the very thing the
subject was changed to stop doing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 13:22:28 +05:30
Rounak DattaandClaude Opus 5 abf1d1f1ca fix(terminal): the PTY and the browser terminal must never disagree about size
Issue #464, "text gets muffled sometimes, in both TUI default and fullscreen".
The screenshot is not a dropped frame or a frozen renderer — it is arithmetic.
Claude Code's TUI wraps its frame at the width the PTY reported and erases the
previous frame by walking the cursor up the rows it believes that frame took. A
browser terminal of a different width makes each logical line occupy more
physical rows than Ink counted, so `eraseLines(n)` clears too few and the new
frame paints over rows nothing erased: doubled lines, and short tool summaries
sitting inside longer prose rows with the prose's tail still visible.

Reproduced against this repo's own xterm before changing anything — a 120-column
PTY against a 62-column terminal renders every wrapped line twice. `test/
terminal-pty-geometry.test.ts` pins that, and pins the clean render at matching
widths beside it, so the assertion cannot be satisfied by code that fixes
nothing.

Four ways the two drifted apart, none of them observable from either end:

1. `fitAddon.fit()` resizes xterm to `proposeDimensions()` RAW while every
   server-facing path reported those floored at 40x10. Measured in Chrome at
   430px: font size 44 proposed 13 columns, the server was told 40, and xterm
   stayed at 13. Three call sites each did their own fit-then-floor, and two
   re-read the proposal after the fit — `_shrinkPaddingToFit()` runs exactly
   there, so the container had moved.
2. `throttledResize` (keyboard up) and `sendResize` (session detached into its
   own window) reflowed locally and withheld only the SIGWINCH. That is the one
   combination that cannot be right: a reflow nothing is rendering for buys
   nothing and costs correctness. Both now withhold everything, and the
   keyboard's settle timer still sends the one resize that stops the PTY going
   stale.
3. `setFontSize`/`setFontFamily`/`setFontWeight` move the cell size — a geometry
   change — and told the server nothing at all, so raising the font on a phone
   left the CLI wrapping at the old column count.
4. `Session.resize` DECLINES a small-viewport request while a desktop connection
   holds an active sizing claim, and said nothing, because resize was write-only.

`syncTerminalGeometry()` is now the one function that may change the terminal's
size: it fits, floors and applies as a single step, so the numbers xterm holds
are the numbers the server is told. A test sweeps every module for a bare
`fit()` on the main terminal, and finds exactly one — the owner's own.

For (4) the client cannot win, so it is told the truth instead: both transports
answer a resize with `session.ptyCols`/`ptyRows` (`{"t":"zc"}` on the socket,
the body of the resize POST) and `_onPtyGeometryReport` adopts them. A terminal
that keeps a shape the PTY refused does not render "too narrow", it renders
garbled. Adopting can leave the pane wider than the screen and the container is
`overflow: hidden`, so `.pty-oversized` grants horizontal reach for exactly as
long as the mismatch lasts: correct-and-reachable beats correct-and-clipped
beats garbled. That rule sets both overflow axes and its own `touch-action`
because mobile.css loads later and sets `.terminal-container { overflow:
visible; touch-action: none }` — a bare `overflow-x` would leave overflow-y
computing to `auto` and hand the browser a vertical scroll container the
terminal's touch handler knows nothing about.

Verified in Chrome at 430px against a live server, with a desktop client holding
the claim: the phone adopts 198x43, gets `overflow-x: auto` / `overflow-y:
hidden` / `touch-action: pan-x`, 758px of reach to the right, and keeps its own
vertical scrolling. The pre-fix build was measured in the same harness for the
control.

Two things this deliberately does not do. It does not change who owns the pane
size — the desktop still wins, and `_startMobileResizeRetry` still takes it back
once that goes idle. And `throttledResize` still holds the PTY's shape for the
whole keyboard animation rather than sending a SIGWINCH per step; that decision
predates this and was not re-tested here.

Also in this commit, Ark0N's third-pass review items on #431:

- The response viewer's byte-buffer fallback and `_onSessionClearTerminal` both
  used the no-param `/terminal` form, capped only by `terminalBufferMaxBytes`
  (32MB) — the largest body the frontend asks for anywhere. One carried no
  deadline at all and the other got the 15s tail budget. Both now take the
  full-history budget.
- A `?full=1` capture that outruns its deadline falls back to the bounded tail.
  The pane is blanked before that fetch, so an abort used to leave a black
  rectangle, discard the queued live output and never reach `_connectWs`. A
  failed load now still opens the socket, says one dim line where the content
  would have been, and clears the tab's spinner — which nothing did, so a failed
  select left `aria-busy="true"` set forever.
- `_wsOutputGapSession` is cleared at the repaint that settles it, not in a
  `finally` that also ran on the catch. A reconcile that threw, or hit the new
  deadline — the flaky link the marker exists for — dropped the gap with nothing
  to retry it. `ws.onopen` no longer clears it up front either.
- The replay-clear invariant is pinned in the gate, which is the drift this PR
  exists to fix: `_resetTerminalForReplay` must be a queued write and nothing
  else, and no module may blank the terminal with a `clear()+reset()` pair.
- `DIAG_ENTRY_MAX_CHARS` replaces the hardcoded 300, bound through a local
  first: `CodemanDiag?.x` still throws a ReferenceError when the identifier was
  never declared, and that is the one function in the app that must not throw.
- panels-ui's two kill-all clears route through the same helper, and the
  xterm-version guard's comment says "resolved lockfile version" rather than
  "dependency RANGE", which is what it has pinned since the last round.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 13:07:33 +05:30
Michael GrundbergandClaude Opus 5 90fd0a5a15 fix(tmux): gate the pane-exit read, and mute the dot on the rich rail too
Four changes the maintainer asked for on Ark0N/Codeman#446 before merging.

The pane-exit watcher stays always-on, but a tick now costs nothing when there
is nothing to observe. `hasObservablePaneSession()` skips the tmux exec while
every session on the manager is one of the shapes `Session.paneExitApplies`
already forces to UNKNOWN: a remote SSH session (its local pane holds the ssh
client), a docker case (a `docker exec` into the container's own tmux), and a
record rebuilt from the socket (no provenance at all). The timer is untouched.
Skipping retracts nothing, for the same reason a failed read does not: the map
still holds the last real reading, and every path that puts a new command in a
pane calls `clearPaneExit()` itself. The two copies of that rule are pinned
against each other in `test/session-pane-exit.test.ts`, because drift between
them is silent in both directions.

`DEFAULT_PANE_EXIT_INTERVAL_MS` was already a constant beside the stats and
remote-reconnect intervals; its comment now says why the watcher owns its own
cadence and why the number is what it is.

The never-default-an-absent-status rule is written where `PaneExit` is declared.
It names `status ?? 0` as the thing never to write, and says that an agent the
OOM killer took would otherwise read as a user typing `/exit` — which is what
absent-stays-absent keeps a later clean-exit sweep away from. Nothing fails when
somebody adds that `??`, which is why the sentence is there rather than a test.

Checking the dot's specificity found a second fight, and it was losing. On the
tab strip the alert rules win as intended: a session that exits with a
permission dialog pending still renders red, and yellow for an idle alert. On
the rich vertical tab rail they did not — that rail's own `tab-state-*` dot
rules are (0,9,1) against the strip's mute at (0,5,0), so an exited session
there kept a full green dot AND the working halo beside a badge reading
"exited". The rail twin matches that specificity exactly and therefore must stay
below those rules in source order; it clears the halo as well, which the strip's
rule never had to think about.

`test/session-pane-exit-ui.test.ts` now resolves the real stylesheet in jsdom
rather than matching selector text: postcss collects every rule that paints
`.tab-status`, a real engine decides, and the tests read back the answer. Two
mutations were run against it to prove it has teeth — dropping the hand-written
alert exclusions fails three cases, and moving the rail twin above the state
rules fails one.

Refs Ark0N/Codeman#446.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 09:19:55 +02:00
Michael GrundbergandClaude Opus 5 64c288a683 feat(codex): read Codex's own background-terminal row
Codex states background work too, and it says so in a different place. Claude
writes `· 1 monitor ·` on the last row of the screen; Codex pins
`1 background terminal running · /ps to view · /stop to close` ABOVE its
composer, which puts that row third from the bottom once the status line and
the composer are counted.

So how far up the screen to look is now per-CLI data as well:
`capabilities.workDetect.watchingLines`, bounded to 1..8 by the schema, and
defaulting to Claude's two. That bound is the point. The window is half the
injection guard, since every row it adds is another row the agent itself may be
able to write, and the label is what silences an idle alert. The other half is
the anchor, and Codex's is ` · /ps to view`: chrome naming a slash command only
the CLI can offer, so a session that writes "I left 1 background terminal
running for you" into its own output matches nothing.

Measured against a live codex-cli 0.154.0 pane rather than read out of a
binary. The row appears when the terminal starts, follows the composer down as
the conversation grows, and is gone after `/stop`. Verified end to end on an
isolated beta: the session payload carried `watching: "1 background terminal"`
and the badge rendered with it, and both cleared when the terminal stopped. The
fixtures in the tests are that capture verbatim.

Codex has no hook signals, so no idle prompt and no false NEEDS YOU row: for a
Codex session this is the badge alone, which is the case the maintainer said a
registry field could cover and a hook never could. Cross-CLI tests pin that
neither pattern fires on the other's screen, and that a CLI declaring nothing
still reports nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 08:58:52 +02:00
Rounak DattaandClaude Opus 5 abd39318e6 fix(terminal): deadline must cover the body, precache must ignore the cache-bust query
Review fixes. Two of these are defects in the previous commit.

1. The fetch deadline only covered time-to-headers. `await fetch()` settles on
   response headers, so clearing the abort timer in a finally around it left the
   body — the multi-megabyte `?full=1` capture the deadline exists for —
   completely unbounded; it only ever bounded a server that accepts a connection
   and never replies. Measured against a server that sends headers immediately
   and stalls the body 4s under a 1s deadline: fetch resolved at 30ms, timer
   cleared there, body completed at 4026ms unaborted. Now the body is read
   inside `_fetchTerminalCapture`, which returns {json, headers, headersAt} —
   headers because two callers read server-timing, headersAt because those same
   callers measure header-vs-body time and can no longer observe that moment.
   `_terminalCaptureInflight` is scoped the same way, so a body still streaming
   counts toward a capture starting beside it. Same test now aborts at 1005ms.

2. The precache could never be hit, and the previous commit made that expensive
   rather than free. `renderIndexHtml` runs `cacheBustAssets`, which appends
   `?v=<mtime>` to every same-origin .js/.css reference INCLUDING content-hashed
   names — confirmed against a running instance:
   `vendor/xterm-zerolag-input.6fee72f2.js?v=1789402869101`. `caches.match` is
   query-sensitive, so entries keyed on the bare hashed path were unreachable;
   deriving the list from the manifest turned cheap 404s into ~1.3MB downloaded
   at every install that nothing could read back, once per deploy now that
   CACHE_NAME rotates. The fallback match takes `{ ignoreSearch: true }`, which
   also lets runtime-cached entries survive an mtime change.

3. `_wsOutputGapSession` was only cleared in ws.onopen, so paths that already
   repaint the buffer left it set and the socket replayed everything a second
   time. `selectSession` loads the buffer and only THEN calls `_connectWs`, so
   neither the _isLoadingBuffer nor the _terminalRefreshOwner guard applied.
   `_markTerminalBufferReconciled()` is now called from _onSessionNeedsRefresh's
   finally, from selectSession after its load, and from _cleanupSessionData.

   The scope claim was also wrong and is corrected in the comment: when the
   network drops, SSE drops with it and handleInit's keepTerminal branch already
   reconciles. The genuinely uncovered case is the WS dying while SSE stays up,
   where _onSSETerminal discards SSE terminal frames until _wsReady flips in
   onclose — up to the ping+pong window of output nothing writes.

4. CLAUDE.md said "all of them measured rather than reasoned", which the PR's
   own "not verified" section contradicted. Split explicitly: the replay race is
   measured, the watchdog mechanism is verified against xterm 6.0.0 under jsdom
   (field path resolves, a forced stale handle makes refreshRows a no-op, the
   kick schedules a fresh frame), and the iOS rAF-discard premise is reasoned
   and still wants a device. Adds the two missing entries — the WebSocket
   reconcile and the sw.js/build.mjs "keep these in sync or the build throws"
   contract.

Also: test/xterm-private-api.test.ts pins the RESOLVED lockfile version instead
of the declared `^6.0.0` range, which was the wrong assertion in both directions
— a real upgrade to 6.4.0 can rename a private field while resolving inside the
range, and an innocuous range edit failed while changing nothing installed. And
test/sw-precache-manifest.test.ts now parses HASHABLE out of scripts/build.mjs
rather than hand-copying it, which was the same drift this PR exists to fix; the
parse is guarded against silently matching nothing.

The deadline fix has a behavioural test against a real socket plus a source
guard asserting `await res.json()` precedes the finally — verified to fail when
the helper is reverted to the old shape, so it is not vacuous.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 12:25:36 +05:30
Rounak DattaandClaude Opus 5 c0422c4e21 feat(terminal): renderer watchdog, atomic replay clear, fetch deadlines, reconnect recovery
Four ways the terminal can silently stop being correct — in each case the
buffer keeps updating, nothing throws, and the only recourse is a reload.

1. Renderer freeze after backgrounding. iOS DISCARDS scheduled rAF callbacks
   when a PWA backgrounds, and xterm's RenderDebouncer only clears its
   `_animationFrame` handle from inside that callback — so one drop leaves it
   permanently set and every later refresh() early-returns. Parsing is
   decoupled from rendering, so bytes keep filling the buffer correctly while
   nothing paints. Codeman has exactly ONE xterm for the whole page load, so a
   single backgrounding wedges it until a reload. Adds a 2s liveness poll and
   `_kickRenderer()`, which does what the dropped `_innerRefresh` would have.

2. Replay clears raced live output. xterm's write() is async-queued while
   reset() is synchronous and, per upstream, "does not clear input buffers and
   does not reset the parser" — so bytes queued before a reset are parsed after
   it and fuse into the snapshot. Verified against the real xterm 6 here:
   write('p8'); reset(); write('rmissions') renders "p8rmissions". The main
   path was already safe via a queued erase; the needsRefresh and clearTerminal
   paths were not. All three now share one queued `\x1bc` (RIS), which unlike
   3J/H/2J also resets modes, charsets, scroll regions and SGR state.

3. Output lost on WebSocket reconnect. Input frames carry seq+cid and are
   delivered exactly once; output frames carry nothing. ws.onopen re-sends dims
   and flushes queued input, and needsRefresh only fires on external-CLI
   startup and SSE backpressure drain — never on reconnect. Output produced
   while offline was simply absent afterwards. Interim fix: reaching onclose
   means the drop was unintentional, so the session is marked and the next open
   reconciles from the server buffer. Sequencing output is the follow-up.

4. Terminal captures had no deadline. No AbortController anywhere in the
   frontend, including `?full=1`, which the code itself calls "unbounded-ish
   work: at the default history limit it can be megabytes". Adds a budget that
   scales with full-vs-tail and with captures in flight, degrading to a plain
   fetch where AbortController is missing.

Also: the service-worker precache was dead — the build content-hashes assets
but sw.js listed pre-hash names, so 15 of 23 entries 404'd (verified against a
running instance) and cache.add().catch() hid it. Offline still worked via
runtime caching, but CACHE_NAME was a constant so activate's cleanup never
deleted anything and every past release's assets accumulated. Both are now
derived from the build manifest. Crash-trail entries are flattened and capped,
since they are joined with \n into one value and one call site interpolates a
server-controlled WS close reason.

The watchdog reads xterm privates — there is no public API. Every access is
optional-chained so a shape change degrades to a no-op. `_renderService` only
exists after open(), which needs a real DOM, so the gate cannot assert the
field path; test/xterm-private-api.test.ts pins the dependency range instead.

Tests: 23 new (terminal-resilience, sw-precache-manifest, xterm-private-api),
all pure/static so they run in the gate, which excludes the mobile suite. One
static source guard in history-truncation-notice updated for the renamed call;
the behaviour it pins is unchanged.

Not verified: no browser available, so no runtime reproduction of the freeze
and no real-device test of the reconnect path. Both warrant a device pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 12:24:52 +05:30
Michael GrundbergandClaude Opus 5 74884a20eb feat(approvals): let a watching session keep quiet, and fix the tui gate
The badge alone left the row in NEEDS YOU, which is the thing the issue was
about. The fix is the alert that does not fire.

An idle prompt from a session that is watching its own background work now
opens ALREADY acknowledged. `hook-event-routes` passes `Session.watching` to
`notePrompt()`, which sets `acknowledgedAt` and records why in a new
`acknowledgedReason`. Nothing new suppresses anything: `acknowledge()` has
always meant "the alert this prompt armed is spent", and the prompt itself
stays pending, answerable and available as Read My Mind context. A wrong label
therefore costs a card that does not blink, never an alert that was never
created.

Every surface follows from that. The broadcast carries the reason, so a live
page declines to arm the tab alert and raises no desktop notification. The push
is skipped, since a false alarm is hardest to ignore on a phone. A reloading
page reads `acknowledgedAt` in `seedApprovals()`, which it already did. And
`classifySession()` now reads it too, which is a pre-existing bug fixed here:
acknowledging on one device cleared the alert everywhere except `codeman tui`.
It re-arms for free, because the next idle prompt supersedes the item and is
built fresh. Only `idle` is eligible, so a dialog that blocks the agent still
goes red whatever else it started.

The label is pane-derived and therefore prompt-injectable, so it is now read
from the last two rows of the screen only, with Claude's pattern anchored on
the `·` its footer joins items with, ANSI-stripped and length-capped at the
source. An agent that prints `· 1 monitor ·` into its own output finds no
match.

Verified on an isolated beta: a session that armed a monitor took its idle
prompt acknowledged with no alert on any surface, wore the badge, and showed
"quiet, watching 1 monitor" on its still-answerable card; the same session with
the monitor killed alerted normally on the next prompt. `test/watching-no-alert.test.ts`
pins both directions across all four surfaces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 08:10:05 +02:00
Michael GrundbergandClaude Opus 5 1cb0441bd8 fix(session): degrade the resume pin to the session id, not to nothing
A single pin that failed its transcript gate returned the options untouched,
so `resumeSessionId` fell back to `_resumeSessionId` — undefined for an
ordinary session — and the renderer emitted the bare
`claude --dangerously-skip-permissions --session-id "<this.id>"`. Every
session prompted before its first `/clear` owns a transcript under that id,
so the dropped pin handed back exactly the refusal this branch removes, with
no `||` branch to catch it. It was also a regression against master on the
`restartCli()` path, which pinned `_claudeSessionId ?? this.id` and, since the
constructor seeds that field, could never land unpinned.

The pin now walks three candidates in priority order — the conversation
chain's tail, the launch seed, then the session's own id — and takes the first
one a transcript backs. A candidate that misses is passed over rather than
ending the walk.

Falling off the end pins nothing, which also settles the second half of the
problem: the old code skipped the transcript check whenever the pin was the
session's own id, so a genuinely new pane rendered the two-branch form after
all. That costs a brand-new session claude's "No conversation found" line in
its scrollback, and `wrapWithNice()` prefixes only the first branch of the
rendered `a || b`, so the branch that actually runs loses its priority for the
life of the session. With no transcript anywhere the bare `--session-id` is
the correct command, so the comment claiming an unchanged shape is now true.

The transcript lookup reads the server process's own `CLAUDE_CONFIG_DIR` when
a session declares none. A pane inherits the server environment through tmux,
so on an install that exports it the CLI writes its transcripts there and
every lookup under `~/.claude` was a false negative — which under the old code
meant the colliding command. `claudeCredentialsPath()` and
`realClaudeConfigDir()` resolve the same directory the same way. The header
sentence calling a skipped resume "the safe direction" described the opposite
of what happens at this call site, and says so now.

The create-path fallback writes `_resumeSessionId` alongside the create
options. That branch leaves `isRestored` false, so `_claudeSessionId` is
recomputed from the launch fields and settled on `this.id` while the CLI
resumed the chain tail; the response viewer, Read My Mind and the unified-list
alias map read that field until the next first-hand hook.

Four new tests: a chain tail with no transcript while the session id has one,
no transcript anywhere, the create path's alias, and the process-env lookup.
All four fail against the previous commit. Two existing tests move with the
gate — the guess-refusal test now backs the session's own id, and the
custom-model restart test gives its working pane the transcript that makes
`--session-id` collide in the first place, alongside a new one pinning the
no-transcript case.

CLAUDE.md described the pin as a `restartCli()`-only thing sourced from the
live conversation id. All three halves of that moved here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 18:21:21 +02:00
Michael GrundbergandClaude Opus 5 3f2cde2db7 feat(session): say when a session is watching its own background work
An agent that arms a monitor, backgrounds a shell or hands a task to a
cloud session is told to end its turn. The pane then falls quiet, Claude
Code's idle_prompt notification arrives a minute later, and every surface
files the session under NEEDS YOU with nothing for a human to answer.

Claude states what it is still running on the last row of its screen
(`⏵⏵ bypass permissions on · 1 monitor · ← for agents`). That row is now
`capabilities.workDetect.watchingLine` in the CLI registry, guarded by
compileVersionRegex() like every other config regex, and the idle probe
reads it off the capture it already takes: `watchingLabel()` in
session-activity.ts searches the last five lines only, so a session that
PRINTS "1 monitor" is not mistaken for one running it.

The label lands on Session.watching and rides toLightDetailedState() out
to every surface. The phone overview, the desktop home rail and the rich
sidebar rows wear it as a `watching` badge in the accent colour, beside
the state pill and never in place of it: an agent can arm a monitor and
ask a question in the same breath, and only the pill says which.

Verified end to end against a throwaway session on an isolated beta
instance: the payload carried `watching: "1 monitor"` once the turn
ended, the badge rendered next to a yellow `waiting` pill, and both
cleared when the monitor died.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 18:00:30 +02:00
Michael GrundbergandClaude Opus 5 47e7935274 fix(session): resume the conversation when respawning a dead pane
A CLI that launches with `--session-id <id>` refuses an id that is already in
use (claude: `Error: Session ID ... is already in use.`), and every session
whose agent has been prompted owns a transcript under that id. The dead-pane
respawn in `_setupOrAttachMuxSession()` passed the bare launch line, so
recovering such a session relaunched a CLI that died on startup, the pane went
dead again at once, and the conversation was stranded behind a tab that looked
merely idle.

`restartCli()` has pinned a resume id against this since the custom-model work,
and its comment states the assumption that made the other path look safe:
"Unlike the dead-pane respawn, this one kills a WORKING pane whose conversation
already has a transcript". A pane whose agent exited has a transcript too.

Both relaunch paths now build options through
`_buildRespawnPaneOptionsWithResumePin()`, and so does the create-path fallback
after a failed respawn, which otherwise met the same refusal that made it the
fallback. Four gates guard the pin, each standing for a way of resuming the
WRONG conversation or of making a working relaunch fail.

A remote or docker session is never pinned. Unlike `restartCli()`, whose route
refuses both, the dead-pane respawn is reached by every session shape. Their
pane commands already render a self-healing `--session-id || --resume`, and
both flip to resume-first once the resume id differs; the conversation lives on
the far side, so a local id resolves to nothing there and the `--session-id`
fallback then collides with the transcript the far side does hold.

The id comes from the conversation CHAIN rather than `_claudeSessionId`, which
also holds history-correlated guesses keyed on the working directory.
`_recordClaudeSessionInChain()` refuses those so they cannot "write a foreign
conversation into this pane's permanent record", and launching from one is
worse than the display bug that rule prevents. The chain tail also outranks the
launch seed, which is written once at construction and never moves off a
`/clear`.

A pin no transcript backs is dropped, because the fallback branch keeps
`--session-id <this.id>` and would collide. A synthetic `restored-<fragment>`
id from socket discovery is dropped too, and logged: it fails claude's `uuid`
token pattern, so the renderer would emit the unpinned command while the caller
believed otherwise.

Tests cover each gate and the rendered command. Four of them fail against the
unfixed source; the remote and docker ones were separately checked against a
build with only that guard removed, since they pass on master for the wrong
reason.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 15:36:17 +02:00
DevvynandClaude Sonnet 5 a2dcc91ddf fix(docker): add and correct the cross-checkout collision guard for Update-Codeman.sh
Same guard as Start-Codeman.sh's own (docs/docker-self-update.md-adjacent
incident, 2026-09-21): docker-compose.yaml hard-codes `name: codeman`, so a
second checkout run without COMPOSE_PROJECT_NAME resolves to the SAME
Compose project as any other checkout on the host. It has to live here too,
not just in Start-Codeman.sh: this script's own --no-cache build and its
`down`/`down --volumes` both run BEFORE the handoff at the bottom of the
file, so Start-Codeman.sh's copy of the guard would only fire after this
script's own destructive calls already ran — and its default `down
--volumes` is more destructive than Start-Codeman.sh's own targeted
refresh, clearing every named volume the resolved project has.

Also fixes a real bug the same guard shipped with: under `set -o pipefail`,
`grep -v` legitimately exits 1 when nothing survives the filter (the
ordinary, no-collision case), and without `|| true` on the pipeline that
non-zero status propagates through the command substitution and `set -e`
aborts the WHOLE script at the guard — every time, collision or not. Caught
only by actually executing the guard end-to-end against a stub `docker`
(the existing smoke-test harness), never by a static text/regex check on
the source; the stub's `config --format json` response was also fixed to
pretty-print like real Compose does, since a compact one-liner silently
resolved project_name to empty and exercised neither script's guard the
way production output does.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-21 21:23:39 +08:00
Michael GrundbergandClaude Opus 5 c67c130caa feat(web): mark a session tab whose agent has exited
The tab now reads "exited (137)" beside the session name, drawn from the
`paneExit` field the server publishes. `applyPaneExitBadge()` owns the DOM
work, called from the incremental render path — the only path a live session
ever takes, since going from live to exited adds and removes no tab and so
never reaches the full rebuild.

An unknown answer draws nothing. A death tmux could not explain reads "exited"
with no number rather than "exited (0)", so an unexplained death and a clean
exit do not look alike. A signal death reads "exited (signal 9)".

The badge carries `data-i18n-skip`, like the status pills: it is generated
text, `i18n.js` walks inserted content, and a dictionary entry added later
would fight the renderer, whose in-place comparison is against English.

The tab also carries a `tab-agent-exited` class that mutes the status dot. That
dot is drawn from `status`, which stays `idle` or `busy` for an exited pane as
the issue requires, so without this a green or pulsing dot sits beside a badge
saying the agent is gone — the first thing a tester asked about. `status`
itself is untouched, so this is a rendering rule only. The CSS excludes the two
alert classes by hand, following the convention the rich-rail dot rules
document: a dot turning red or yellow because a session is blocked on a human
outranks "the agent exited".

The tab keeps its click behavior. X still closes it, and nothing here closes,
sweeps or restarts anything.

`docs/architecture-invariants.md` gains the mechanism under "Session data and
lifecycle", where every comparable one already lives: what the tri-state means,
the four shapes it is absent for, why the watcher cannot ride the stats
collector, why an absent `#{pane_dead_status}` is not 0, and the three things
that must never happen to an exited pane.

Refs Ark0N/Codeman#446.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 10:52:15 +02:00
Michael GrundbergandClaude Opus 5 90a95f562b feat(session): publish and persist a local pane's agent exit
The mux layer now knows a pane's agent has exited. This puts it on the session
record, where the board and, later, the reboot restore can see it.

`SessionState.paneExit` carries `{ status?, signal?, at }` and rides the
existing `session:updated` broadcast through `toState()`. No new SSE event. The
server pulls each answer from `mux.getPaneExit()` rather than off a broadcast
payload, so the raw reading never reaches a browser: for a remote or docker
session that reading is the death of an ssh client or a `docker exec`, not of
the agent.

The field is tri-state, and the third state is its absence: `undefined` means
Codeman does not know, and it never reads as alive. `Session.setPaneExit()`
forces that unknown for every shape a dead local pane does not describe. A
direct-PTY session owns no pane. A remote SSH session's local pane holds the
ssh client, whose death means a transport drop OR an exit, which is the
ambiguity PR #355 settled by not guessing. A docker case's local pane holds a
`docker exec` into the container's own tmux. And a session rebuilt from the
socket has no provenance at all: `reconcileSessions()` gives it a synthetic
`restored-<fragment>` id that matches no `state.json` entry, so a remote
session rediscovered after `mux-sessions.json` was lost arrives with no
`remote` field and looks local — `MuxSession.discovered` marks it, and absent
metadata there counts as unproven rather than as proof. The scoping lives on
`Session` rather than in `TmuxManager` so there is one copy of the rule.

`status` and `pid` are untouched. `status: 'error'` belongs to the PTY-exit
circuit breaker and makes the browser offer a restart, and a null `pid` is what
makes the browser re-attach and launch a fresh CLI. A reading that repeats the
previous answer writes nothing and broadcasts nothing.

An unknown answer never reads as alive, but a stale KNOWN one would keep
reading as exited, so `clearPaneExitForNewPane()` retracts it on every path
that puts a new command in the pane: the start/attach path, the `restartCli()`
relaunch behind a custom-model switch, and the remote reattach. Without the
second of those, switching an endpoint on an exited session launched a new
command and then persisted and broadcast the old exit straight back onto it.

`toState()` is also what `state.json` persists, so the record survives a
reboot, which is the only thing that does: a reboot takes the tmux server, and
with it every live signal and every `mux-sessions.json` entry. Nothing reads it
there yet — making the restore refuse such a session is a behavior change that
belongs with the part that closes them. Recovery threads the saved value back
through the constructor so the first persist after boot cannot blank it.

Refs Ark0N/Codeman#446.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 10:01:09 +02:00
Michael GrundbergandClaude Opus 5 02dc46dcd7 feat(tmux): report a dead pane's exit from the batched pane list
Codeman creates every tmux pane with `remain-on-exit on`. When the agent exits,
tmux keeps the pane, the tmux session, and the `tmux attach-session` process
Codeman records as the session's pid, so no PTY exit handler fires and nothing
writes the exit down. tmux itself knows: it marks the pane dead and reports the
exit status. This reads that.

`PANE_LIST_FORMAT` gains `#{pane_dead}`, `#{pane_dead_status}` and
`#{pane_dead_signal}`, and `startPaneExitWatcher()` refreshes a
muxName-to-observation map from ONE batched `tmux list-panes -a` per tick. Boot
reconciliation already ran that same call, so it now fills the map too and
recovery starts with a reading.

The watcher owns its own interval rather than riding `startStatsCollection()`,
which the issue suggested. That collector is armed when a browser opens the
Monitor panel and DISARMED when it closes it, and boot skips it entirely unless
recovery found a live session, so a session created on a freshly booted server
would publish nothing and one browser could turn detection off for every other.
Measured on an isolated instance: a dead pane with status 0 reported nothing
until `POST /api/mux-sessions/stats/start` was called by hand. It is still one
batched read per tick; only the timer changed.

Three rules keep a positive answer trustworthy. A session answers only when
tmux listed exactly one pane for it, because Codeman never splits a pane and a
session the user split by hand has none that speaks for the agent. A pane
answers only when `#{pane_dead}` said 1 or 0, because an empty field is a tmux
that did not answer. An absent status stays absent rather than becoming 0:
measured on tmux 3.2a, a SIGKILLed pane reports neither a status nor a signal,
and calling that a clean exit would be wrong in the direction that matters.

Two guards stop a slow read undoing a fast one. `EXEC_TIMEOUT_MS` is 5000 ms
against a 2000 ms interval, so a read can outlive two ticks: one already in
flight suppresses the next, and a generation counter that every
`clearPaneExit()` bumps discards a read that started before a respawn or a
kill. An observation also carries its pane pid, so a second command in the same
pane that exits the same way starts a new timestamp rather than inheriting the
first death's.

A non-empty read of `list-panes -a` is authoritative for the whole socket, so
sessions missing from it are pruned, which also bounds the map as tmux sessions
come and go outside `killSession()`. A failed or empty read retracts nothing.

The manager reports the raw pane reading and applies no session-shape scoping,
because the remote-reconnect watcher beside it needs exactly that raw reading.

`parsePaneList` becomes `parsePaneRows`, returning one row per pane instead of
a name-to-pid map; reconciliation builds its map from the rows. The parser's
existing cases carry over unchanged, including the launchd/systemd literal-tab
regression from PR #71.

Refs Ark0N/Codeman#446.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 10:00:54 +02:00
DevvynandClaude Sonnet 5 5b4878df3b fix(docker): address Ark0N's PR review on Update-Codeman.sh — fix the handoff, the build/down ordering, and default-clear the build volumes
Three blockers, all fixed and verified by actually running the script (not
just string-matching it):

1. `exec "$script_dir/Start-Codeman.sh"` failed EACCES/exit 126 on every
   checkout, since Start-Codeman.sh is committed non-executable (100644) —
   the same fact my own second commit on this branch established. Fixed to
   `exec bash "$script_dir/Start-Codeman.sh"`.

2. `down` ran before `build --no-cache`, so Codeman and every session it was
   running were offline for the entire rebuild, and a build failure left the
   stack down with nothing to bring it back — the exact ordering mistake
   Start-Codeman.sh's own "Build BEFORE taking the stack down" comment exists
   to prevent. Reordered to build, then down, then hand off.

3. The default path could throw the rebuild away: codeman-node-modules/
   codeman-dist only re-seed from the image while EMPTY, Start-Codeman.sh
   only clears them when it detects the checkout's HEAD or package-lock.json
   moved, and neither condition is true for the Dockerfile-only change this
   script exists for — so a plain `bash docker/Update-Codeman.sh` rebuilt an
   image whose fresh node_modules/dist then sat unused behind the old
   volumes. Made clearing them the default; `--keep-volumes` opts out
   (replaces the old `--volumes`/`-v` flag, which is no longer needed since
   clearing is now the default).

Smaller items from the same review, also fixed:

- The --no-cache build now derives PUID/PGID from CODEMAN_APPDATA_PATH's
  owner first, via the identical owner_of() helper Start-Codeman.sh uses
  (parity-tested) — without it, the build used Compose's default 1000:1000
  regardless of the real appdata owner (99:100 on the Unraid layout
  docker/README.md documents), and Start-Codeman.sh's own correctly-PUID'd
  build during the handoff would then rebuild those layers anyway, so the
  --no-cache image never actually shipped.
- docker/README.md's "rebuilds ... only when it detects ... moved" wrongly
  described BOTH the rebuild and the volume-clearing as conditional;
  Start-Codeman.sh rebuilds on every start, only the volume-clearing is
  conditional. Corrected, and reworded around the new default.
- --help/-h now prints usage and exits 0 instead of falling into the
  unrecognised-argument branch.
- "the ONLY named volumes this stack declares" now says docker-compose.yaml
  specifically, since a docker-compose.override.yml could add more.

New tests: PUID/PGID derivation parity with Start-Codeman.sh's owner_of(),
--help handling, and — the one that actually catches blocker #1, which five
source-string-matching tests did not — a real end-to-end smoke test: a
synthetic deployment, a stub `docker` on PATH logging every invocation, the
real script executed via a real subprocess. Confirms the real command
sequence (build --no-cache, then down --volumes or plain down, then evidence
the handoff genuinely ran Start-Codeman.sh) and that a working handoff fails
honestly at Start-Codeman.sh's own later check rather than with EACCES.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
2026-09-21 15:19:00 +08:00
DevvynandClaude Sonnet 5 3b714446b4 fix(docker): drop the wrong executable-bit assertion for Update-Codeman.sh
Start-Codeman.sh, its sibling and the script it hands off to, is itself
committed non-executable (100644) upstream — it's documented and invoked
as `bash docker/Start-Codeman.sh`, never `./docker/Start-Codeman.sh`. The
"is executable" test I'd added for Update-Codeman.sh asserted the opposite
convention, which the file correctly does not follow.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
2026-09-21 13:58:45 +08:00
DevvynandClaude Sonnet 5 9ba90a674a chore(docker): add Update-Codeman.sh for scripted major-update rebuilds
docker/README.md and docs/docker-self-update.md both already point operators
at "stop the stack, rebuild, restart" for anything the in-app updater refuses
to apply (a changed server.Dockerfile, a changed docker-compose.yaml, or a
new required .env key) — but that was a manual, hand-typed procedure with no
script of its own, unlike every other start/update path this deployment has.

docker/Update-Codeman.sh scripts it: `docker compose down`, then an
unconditional `docker compose build --no-cache` (a major update should be
certain of what actually ships, not reuse whatever layers happened to be
cached), then hands off to the existing Start-Codeman.sh for the same
careful PUID/PGID, override-file and fingerprint handling every other start
already goes through — rather than reimplementing any of that by hand and
risking it drifting out of step.

An optional --volumes/-v flag also removes the codeman-node-modules/
codeman-dist named volumes, the scripted form of the "Resetting the build
artefacts" procedure docs/docker-self-update.md already documents by hand.
Safe: those two are the only named volumes this stack declares; application
data and case workspaces are host bind mounts, never touched by
`docker compose down` either way.

Docs updated: a "Major updates" section in docker/README.md, and a pointer
from docs/docker-self-update.md's existing "Resetting the build artefacts"
troubleshooting entry.

Tests: extended test/docker-entrypoint.test.ts (the existing home for
Start-Codeman.sh's own static checks) with a bash -n parse check, the
down-before-build-before-handoff ordering, the --volumes flag's effect,
unrecognised-argument handling, and byte-for-byte agreement with
Start-Codeman.sh's own override-file resolution logic (so `down` here and
`up` there can never target different Compose files).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
2026-09-21 13:57:54 +08:00
DevvynandClaude Sonnet 5 d3ee9f23c2 fix(cli-registry): correct accent colours, and a real gemini/antigravity/omp rendering bug
Two related fixes, found while re-measuring stock.ts's `accent` field
against the actual rendered UI (docs/cli-registry.md flags this field as
"transcribed, not authoritative — re-measure before wiring one up"):

1. A real, user-visible bug: `.btn-toolbar.btn-run.mode-gemini`,
   `.mode-antigravity` and `.mode-omp` had no override rule inside the
   `html:not([data-skin="og"])` block, unlike codex/pi/grok/deepseek, which
   do. The generic `.btn-toolbar.btn-run` rule in that block resolves at
   higher specificity than the base sheet's per-mode pair, so all three
   rendered as plain claude-blue on every skin except `og` — including
   `daylight-blue`, which is the actual DEFAULT skin for a fresh install
   (index.html's pre-paint script), not an edge case. Added the three
   missing rules, sourced from each CLI's own already-designed og-skin
   colours (no new colours invented), mirroring the exact pattern
   pi/grok/deepseek already use. Also corrected the stale comment on the
   pi rule, which claimed this was still broken for gemini/antigravity.

2. `stock.ts`'s `accent` field was simply wrong for most CLIs — e.g. claude
   was registered as Anthropic's brand orange (#d97757) while its button
   renders blue, antigravity was registered purple while it renders cyan,
   pi was registered green while it renders pink. Measured each CLI's real
   `border-color` from its own `.mode-<id>` rule on the og skin (the
   cleanest single representative hex each entry's gradient resolves
   around) and corrected all 9 non-shell entries to match. `accent` has no
   reader yet (confirmed via the DECLARED_FOR_LATER guard test), so this
   changes no rendered output — it's a data-accuracy fix, matching the
   registry's own "transcribed, not authoritative" warning taken literally.
   Also fixed a false claim in types.ts's doc comment for the field
   ("CSS derives every per-CLI gradient from it via --cli-accent") — no
   such CSS variable exists anywhere in the codebase.

Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format:check/
check:public-assets/check:frontend-syntax all clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-21 12:23:19 +08:00
github-actions[bot]Claude Fable 5.1github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
9466acfc1a chore: version packages (#461)
* chore: version packages

* chore: sync CLAUDE.md version to 1.32.0

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
2026-09-21 06:00:49 +02:00
Codeman maintainer e899af4305 chore: changeset for the merge-time fixes
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 04:53:19 +02:00
Codeman maintainer 299a21d5f5 fix(split-pane): merge-time fixes for split-pane sessions (#453)
The maintainer's promised merge-time fixes from the final review of #453:

1. closeSplitPane() tears down a divider drag still in progress, so a split
   that collapses mid-drag no longer leaves body.split-pane-resizing (the
   page-wide col-resize cursor and user-select lock) set until a reload.
2. openSplitPane() re-applies the picker's own exclusions (detached session,
   pid === null, no session record) for a row that went stale while the
   menu sat open, refusing silently like its neighbouring gates.
3. architecture-invariants: the hard-hide of .btn-split is the
   @media (max-width: 1179px) rule in styles.css, not mobile.css.
4. SplitTerminalPane.destroy() nulls onclose (and onerror) beside onopen
   and onmessage.
5. Picker rows drop the data-session-id attribute nothing read.
6. The Pane-A-ends branch collapses with skipPrimaryResize, so the closing
   resize is no longer aimed at the session the server just removed.
7. The {t:'r'} refresh path is single-flight across the fetch and the
   chunked write, coalescing a mid-replay refresh into one trailing re-run.

Tests: split-pane-auto-collapse-unit gains the drag-teardown, exclusion and
skip-resize cases; the new split-pane-terminal-unit covers destroy() and the
refresh single-flight. All were run against the pre-fix module to confirm
they fail there.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit dbd39aed015ae5ae5870aba398bf4b4ab5118e47)
2026-09-21 04:53:19 +02:00
Codeman maintainer d8e85285c9 fix(mobile): merge-time fixes for the prompt composer (#444)
- styles.css: restate the composer overlay's own bottom gutter after the fold rules
  (the generic .paste-overlay longhand erased it: 0px flat, hinge strip replacing it
  folded) and subtract the fold strip from the dialog's max-height
- test/foldable-layout.test.ts: simulate the cascade for
  .paste-overlay.prompt-composer-overlay (fails without the CSS fix); pin the palette
  anchor by name instead of ELEMENTS.at(-1)
- keyboard-accessory.js: guard the app global in refreshForActiveSession() like the
  rest of the file
- keyboard-accessory.js: a whitespace-only draft is empty (Send no longer submits
  blank lines); the text still goes out untrimmed
- keyboard-accessory.js: derive _composerMaxLength and the frame refusal from one
  64 KiB frame limit minus both bracketed-paste markers so they cannot drift
- keyboard-accessory.js: translate the textarea placeholder and label at build time,
  since the DOM translator skips <textarea> subtrees
- i18n.js: zh-CN entries for the composer dialog copy
- docs/wiki/Mobile-Guide.md: describe the Compose key instead of a clipboard key
- CLAUDE.md: a "Mobile prompt composer" paragraph after the accessory bar one
- test/mobile-prompt-composer.test.ts: pin the whitespace rule and the derived budget

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit f6725ba52da17b0bdbee8be3b5011e7cae514f69)
2026-09-21 04:39:28 +02:00
Codeman maintainer 0f955327b2 fix(cli-registry): merge-time fixes for the run-menu consolidation (#458)
- test/opencode-resize.test.ts: retarget the launcher guard at the real code (this.selectSession(firstSessionId), any this.activeSessionId assignment) with an anti-vacuity check; the old strings existed nowhere, so it could never fail
- session-ui.js: restore as comments the two invariants the merged bodies lost (deepseek leaves statusReporting unset, i.e. ON; no effort field for external CLIs, it is Claude-specific)
- docs/cli-registry.md: move the frontend-guard paragraph below the two backend-guard paragraphs so they keep their antecedent, and note the widened comparison shape
- test/frontend-cli-no-id-branching.test.ts: the comparison shape accepts any left-hand identifier (const m = this._runMode; m === 'codex' was invisible), normalized to `mode`; the two `m !== 'shell'` display filters are allowlisted and the remaining blind spots documented
- test/run-mode-dispatch.test.ts: table-driven pin of run() dispatch (claude to runClaude, each RUN_MODE_LAUNCH id to _runCliMode(id), shell to runShell, unknown to runClaude, lock held and released)
- CLAUDE.md: name the second CI-gated guard next to the backend one
- server.ts: every </head> injection passes a replacer function; a clis.json label containing $' re-injected the rest of the document past escapeScriptJson (two render tests pin it, proven failing on the string form)
- _isAltCliMode(): no reference anywhere in the tree, nothing to fix

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 1ea363ff808a62861559bc141e724b163cc1c56e)
2026-09-21 04:37:45 +02:00
Codeman maintainer d3f2ec0220 fix(custom-model): merge-time fixes for the promoted-model picker (#459)
The maintainer's promised follow-ups to opticon454's picker promotion,
applied on the landing branch after the merge (ecb95b5d):

- session-ui.js: the promotion tag ("Currently loaded" / "Last used") and
  the "Default" pill are two separate spans, so a promoted row that is
  also the endpoint's defaultModelId shows both instead of silently
  losing its Default marking; two tests pin it (both fail on the old
  exclusive-slot rendering).
- styles.css: a dedicated #customModelPickModal .set-scope rule, since
  the pill was only styled inside the three settings modals and rendered
  as plain body text here; same skin tokens, modal layout untouched.
- docs/wiki/Custom-Model-Endpoints.md: describe the promotion (currently
  loaded, else last used per device), the separate Default pill, and
  that nothing is ever auto-chosen.
- CLAUDE.md + docs/custom-model-endpoints.md: credit the real "Last used"
  writers (_runCustomModelEntryViaRestart and
  _quickStartWithCustomModelConfirm; runCustomModelEntry only dispatches
  since 88e5b7b2) and drop the now-wrong "both defer to Default" sentence.
- Not done: moving the one-shot "last used" write into
  _runCustomModelEntryOneShot, because the existing one-shot tests assert
  that _quickStartWithCustomModelConfirm writes the key itself.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 0cb0f911adc14a852ba5c2951768a4aa87c25657)
2026-09-21 04:30:01 +02:00
Codeman maintainer 6ef71ec3b9 chore: thanks for 1.32.0
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 04:28:22 +02:00
Codeman maintainer d47f93abdb chore: changesets for #453, #444, #459 and #458
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 04:27:17 +02:00
Codeman maintainer dcf9437308 Merge pull request #460 from Ark0N/feat/installer-v2
feat(install): three questions up front, an unattended build, and a URL you can scan
2026-09-21 04:23:15 +02:00
Codeman maintainer aa13af1f7f Merge pull request #444 from DodgyBadger/feat/mobile-prompt-composer
feat(mobile): add manual prompt composer
2026-09-21 04:23:15 +02:00
Codeman maintainer 9a2e14a93a Merge pull request #453 from timkjr/feat/split-pane-sessions
feat: split-pane sessions — view two live terminals side by side
2026-09-21 04:23:14 +02:00
Codeman maintainer ecb95b5d67 Merge pull request #459 from opticon454/feature/run-menu-picker-currently-loaded-model
feat(custom-model): promote the currently-loaded/last-used model in the Run-menu picker
2026-09-21 04:23:14 +02:00
Codeman maintainer a7452dc046 Merge pull request #458 from opticon454/followups
feat(cli-registry): drive the run-menu frontend from the CLI catalogue (PR B2)
2026-09-21 04:23:13 +02:00
Codeman maintainer 72d437ab63 fix(install): fold in both reviews of #460
The two reviews on the PR (DeepSeek Harness, then Claude) found one class of
bug twice and a list of smaller ones; all of them land here, each pinned in
test/install-sh-invariants.test.ts and, where it is bash logic, driven in the
bash:3.2 CI step as well.

The Start line the done screen prints is now composed in one place
(start_command_hint) from every non-default value, the same five the exec
branch exports through export_bind_env, so "do not start" under a sub-path or
a custom port no longer prints a bare `codeman web`. The --lan / --tailscale /
env preset paths read ${CODEMAN_PASSWORD:-$EXISTING_PASSWORD}: a flag re-run on
a unit that carried a password used to rewrite it without the password and
with the unauthenticated ack. --password and --port flip RECONFIGURE so they
reach the unit instead of taking the quiet update path, and `install.sh name`
re-syncs the unit's base URL after the mapping is re-added.

Also: the sudo keepalive is ended before the exec into the foreground server
(exec skips the EXIT trap, and the loop keys on $$); Ctrl+C in the HTTPS-toggle
poll is trapped for the poll only and skips Tailscale for the run instead of
killing the installer; uninstall asks before removing a LaunchDaemon this
installer never wrote; a foreign daemon gets a launchctl kickstart hint and the
done screen stops claiming the new build is running; the preflight summary
reads the Tailscale state with a line grep when node is not installed yet; the
LAN security notice uses the configured port; a bare re-run ends on the done
screen; a build failure after a rename names the install.sh tailscale
recovery; TS_JOINED_HERE (written, never read) is gone; the plan doc and
architecture-invariants say what the code does. A minor changeset is included.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 03:03:31 +02:00
Codeman maintainer 1ba0684438 docs(install): describe installer v2 and the Tailscale naming options
README, the Installation / Remote-Access / Running-As-A-Service wiki pages,
docs/security-architecture.md and CLAUDE.md describe the three-question flow,
the flags, the subcommands, the sub-path answer for an occupied :443 and why
the rename is opt-in. docs/installer-v2-plan.md is the design and the
verification record (what was measured, what still needs a fresh machine);
docs/tailscale-installer-plan.md points at it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 21:38:12 +02:00
Codeman maintainer af744bdb54 feat(install): ask three questions up front, then install unattended and end on the URL with a QR code
The installer used to ask about ten things, half of them after a multi-minute
build, and the question that matters most (how do I reach the dashboard) came
last. It now looks at what is on the machine, asks at most three questions
(access, an optional tailnet name, service), and does the rest unattended.

- Every step that needs a human runs before the build: one consent for all
  missing packages, one sudo prompt kept warm for the run, the AI CLI menu,
  and the Tailscale install/login/operator/HTTPS-toggle preflight (the toggle
  is polled and the admin page opened in a browser, instead of "re-check
  now?").
- The build, the service and `tailscale serve` run behind spinners with their
  output in ~/.codeman/install.log; the tail is shown on failure and a failed
  dependency install names its step.
- The done screen leads with the URL (tailnet, network, this machine) and a
  terminal QR code from the qrcode package Codeman already ships.
  `install.sh status` prints it again.
- Tailscale is two halves: tailscale_prepare (question phase) decides the
  serve SHAPE, tailscale_apply (after the build) issues the one serve command.
  When :443 already belongs to another app, Codeman goes under a sub-path
  (serve --set-path /codeman + CODEMAN_BASE_URL in the unit; serve strips the
  prefix, Codeman's ingress tolerates that, --base-url covers the URLs it
  emits) or a second port, instead of replace-or-nothing.
- Renaming the node to codeman-<hostname> is opt-in and defaults to no
  everywhere (the tailnet name is the machine's ssh identity); --name and
  `install.sh name` do it, uninstall offers the old name back. Serve config is
  keyed by the DNS name, so a rename takes our mapping down first and re-adds
  it under the new name.
- Flags pipe through `bash -s --`: --tailscale|--lan|--local, --name|--no-rename,
  --service|--run|--no-start, --yes, --password, --port. --port is now also
  written into the service file.
- npm install runs with CODEMAN_NO_AUTOSTART=1: postinstall otherwise builds
  and starts a detached `codeman web` on 127.0.0.1:3000, which made the
  service crash-loop on EADDRINUSE while the done screen reported "running"
  off the orphan (fresh Ubuntu 24 sandbox).
- The LAN address comes from the default route, not the first interface.
- A foreign /Library/LaunchDaemons/com.codeman.web.plist is left alone
  instead of being replaced by a LaunchAgent.
- The cloudflared question leaves the main flow (`install.sh cloudflared`).
- "Continue WITHOUT a password?" defaults to yes (owner decision).

Tests: the invariants test pins no `serve reset`, no funnel, no Tailscale
Service, every serve mutation through ts_cmd_serve, rename before shape,
flag/header parity, the rename default and the NO_AUTOSTART opt-out; the CI
bash 3.2 step drives the question phase with stubbed tailscale state.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-20 21:38:11 +02:00
timkjrandClaude Sonnet 5 46d8b92049 fix(split-pane): port Ctrl+Shift+C's never-falls-through guarantee to Pane B
The smart-copy gate only entered its selection-check block behind
hasSelection(), so a selection-less Ctrl+Shift+C skipped straight to
`return true` and ceded the keystroke to the browser's own handling
(e.g. Chrome's Inspect-Element binding) instead of matching Pane A's
"never falls through" contract for that chord.

Verified live in a real browser that this is a UX-parity fix, not an
interrupt-safety one: xterm's evaluateKeyboardEvent never emits PTY
data for a shifted ctrl-letter regardless of any gate (only "_" and
"@" get special-cased), so no accidental 0x03 was ever at risk. The
regression test added here asserts on the dispatched event's
defaultPrevented rather than the absence of a WS frame, since the
frame-count check passes vacuously for this exact key combo whether
or not the gate fires.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:34 -05:00
timkjrandClaude Sonnet 5 0b3e086334 fix(split-pane): address Ark0N's fourth pass — PTY-less picker exclusion, hollow chord test, remaining key gates
- buildSplitPickerSessions() now excludes any session with pid === null
  (exited CLI, tripped PTY-exit breaker, a restore that never re-attached).
  Pane B has no equivalent of selectSession()'s auto re-attach POST, so a
  split opened onto one had nothing reading its tmux pane: no terminal
  events ever arrived and Session.write() silently dropped every keystroke
  with no ack either way, while the socket itself reported healthy.
- Fixed the hollow chord regression test: the synthetic keydowns carried no
  keyCode, which is what xterm's evaluateKeyboardEvent switches on to
  produce a data frame at all, so the assertion held regardless of whether
  the gate fired. Adding real keyCodes surfaced a second, real bug in the
  Alt+B case: the event bubbles to app.js's own document-level shortcut
  dispatcher, which really toggles the sidebar and resets the layout
  attribute the gate reads before Pane B's own (later, non-capture) handler
  ever sees it — fixed by driving the app's real settings cache instead of
  only the DOM attribute.
- Ported the two remaining primary-pane gates with real consequences:
  Ctrl+Z (SIGTSTP) is swallowed for every non-shell session, matching
  terminal-ui.js's reasoning (an Ink/TUI agent loop stops dead with no
  visible output otherwise), and Shift/Ctrl+Enter now POSTs to
  /api/sessions/:id/send-key for THIS pane's own session instead of
  letting xterm send a bare \r, which used to submit an incomplete prompt
  instead of inserting a newline. Smart-copy Ctrl+C is re-implemented
  against Pane B's own terminal (copying app.copyTerminalSelection() would
  have copied Pane A's selection instead).
- Updated docs/architecture-invariants.md and docs/split-pane-sessions-plan.md
  to match, and added CLAUDE.md's missing .split-picker-menu z-index entry.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:34 -05:00
timkjrandClaude Sonnet 5 fafef0aa00 fix(split-pane): gate app-level chords out of Pane B, address Ark0N's third pass
Pane B had no attachCustomKeyEventHandler of its own, so the document
capture-phase shortcut handler's preventDefault() (which does not stop
xterm) left Ctrl+K/Alt+1/Alt+B ALSO writing their raw byte/escape
sequence into Pane B's live PTY on top of whatever the app action did
to Pane A. Pane B now installs the same registry-aware gates the
primary pane's own attachCustomKeyEventHandler uses. Ctrl+V is left on
xterm's default paste — no image-paste trap to route it to.

Plus the rest of the review's smaller items:
- Narrowing the window past the desktop gate now closes an open split
  instead of leaving it stranded on screen.
- Split is refused while a web tab is active (activeWebviewId), which
  used to open Pane B's socket behind a hidden container.
- Pane B now handles the server's `{t:'r'}` refresh frame via a shared
  _loadBuffer() helper (also used by connect()), instead of ignoring it.
- The divider drag now uses pointer events + setPointerCapture (mirrors
  tab-rail-resize.js), a button!==0 guard, preventDefault, and a
  body.split-pane-resizing cursor/selection lock — a plain mousedown
  drag selected the text under the cursor as it crossed both terminals.
- Pane B's close control and the picker rows are real <button>s now
  (keyboard-reachable), with matching CSS chrome resets.
- Dropped the redundant CodemanBase.base prefix on the buffer fetch
  (the global fetch wrapper already applies it).
- data-preview-order for the Split settings chip moved from a collision
  with Ultracode Agents (both 15/12) to 11.5, matching its real
  position between Multi-monitor and Ultracode Agents in the header;
  widened test/app-settings-structure.test.ts's regex to allow the
  decimal (Number() already parses it fine for the preview sort).
- Added zh-CN i18n entries for the Split button and empty-picker text.
- Dropped the stray unused `vi` import Ark0N flagged as unrelated to
  this feature (vitest's `globals: true` makes it ambient anyway).
- Documented the fix and the deliberate no-cid/seq choice in the
  split-pane-sessions architecture-invariants entry.

Added a real-Chromium regression test asserting Ctrl+K/Alt+1/Alt+B
dispatched at Pane B's own textarea send no `{t:'i'}` frame over its
WebSocket. Full CI gate green (409 files, 7736 tests) plus all 8
split-pane browser tests.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:33 -05:00
timkjrandClaude Sonnet 5 3152ec801d docs(split-pane): short CLAUDE.md rule, stale module count, wiki entries, shortcut-handler caveat
CLAUDE.md previously only mentioned split-pane in the load-order list,
with nothing in the Architecture/frontend prose the way every other
feature gets, and its own module count was one stale (34, should have
been bumped to 35 when terminal-split.js was added). Add a short
pointer-style paragraph next to the other terminal features, fix the
count.

docs/wiki/The-Dashboard.md's header button table and
docs/wiki/Settings-Reference.md's header chips list are the two
user-facing surfaces that never mention Split at all; added both, plus
a note that the feature is desktop-only regardless of the setting.

docs/split-pane-sessions-plan.md: recorded the one design note that
isn't a code change — the global capture-phase shortcut handler always
resolves against Pane A, so Ctrl+L/Ctrl+W typed into Pane B affects the
other session. Not fixed for v1, same reasoning as the rest of the
"deliberately plainer" section.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:33 -05:00
timkjrandClaude Sonnet 5 2e3e245cc6 fix(split-pane): throttle the drag, chunk the scrollback, and the rest of Ark0N's second pass
Two majors:
- The divider drag was unthrottled: every mousemove did a full xterm
  reflow on BOTH panes and sent Pane B a {t:'z'} resize frame with no
  unchanged-dimensions skip, fanning out into a `tmux resize-window`
  child plus a SIGWINCH per event — ~50 of each dragging across half a
  wide viewport. SplitTerminalPane.fit() is now split into localFit()
  (reflow only) and fit() (reflow + send); the drag coalesces moves
  into one localFit() per animation frame via requestAnimationFrame,
  and sends the real resize for both panes exactly once, at drag end,
  matching the primary pane's own throttledResize convention.
- Pane B pulled the FULL scrollback unchunked for every session mode,
  writing it in one terminal.write() call. Mirrors the primary pane's
  own mode check (app.js's selectSession): shell sessions get a
  bounded 1MiB ?tail= fetch instead of ?full=1, and the fetched buffer
  is written through a minimal chunked writer (32KB slices, yielding a
  frame between each) instead of one primary-pane chunkedTerminalWrite
  this simpler, independently created/destroyed pane has no equivalent
  of (no session-switch generation counters or live-output gate).

Smaller items from the same review:
- Pane B now follows live appearance changes (applyTerminalSkin,
  applyTerminalFontFamily, applyTerminalFontWeights, setFontSize all
  propagate to it, matching the teammateTerminals pattern) and reads
  the real codeman-font-size/terminalFontFamily/weights/DEFAULT_SCROLLBACK
  settings at construction instead of hardcoding fontSize 14 / scrollback 5000.
- The Pane-B-promotion path now skips selectSession() when
  _closingSessions already owns this delete (the user closing Pane A's
  own tab), matching _onSessionDeleted's own active-session-handoff guard.
- Detaching a session AFTER a split is already open now yields the PTY
  size in _sendResize() too (not just at picker-open time), mirroring
  sendResize's own detachedElsewhere guard.
- .btn-split joins the body.solo-mode hide list, next to .btn-multimonitor.
- The split row was 6px wider than its container (two flex-shrink:0
  50% panes plus a 6px divider): both panes are now flex-shrink 1.
- Pane B's header and the split-picker rows are marked so i18n.js's
  exact-string lookup skips them, matching .session-name elsewhere —
  a session literally named e.g. "Sessions" was translatable on zh-CN.
- The Split button now reflects open/closed state via a `.split-open`
  accent style, aria-pressed, and a title/aria-label that says which
  behaviour the next click gets.
- _splitPane.connect() is no longer an unawaited call with no .catch().
- terminal-split.js's fileoverview pointed at a doc path that was
  renamed away in the previous push; @dependency now credits
  constants.js for CodemanTerminalFont, not terminal-ui.js.
- index.html's Split settings chip no longer reuses data-preview-order
  "12" (already the Ultracode Agents chip's slot in the same "header"
  preview group).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:32 -05:00
timkjrandClaude Sonnet 5 165cfb52d6 fix(split-pane): stop breaking every settings save, finish the desktop gate
Blocker from Ark0N's second PR #453 pass: moving showSplitButton into
settings-ui.js's per-device displayKeys set was only half of making it
per-device. saveAppSettings() still put it in the object PUT to
/api/settings, SettingsUpdateSchema (.strict()) does not declare it,
the server answered 400 INVALID_INPUT, and because the call site never
checked res.ok the UI still reported "Settings saved" while NOTHING
persisted — workspaceHooksEnabled, agentSkillEnabled, tunnelEnabled,
claudeModel, every toggle, on every save, on every device. Strip it
out via the same destructure every other per-device key goes through
(`showSplitButton: _ssp,`), drop the stray mention from a schemas.ts
comment (a mention there reads as "this is a real field" to the next
grep), and add a static guard test mirroring
test/terminal-auto-copy.test.ts's three-way rule.

Also finishes the desktop gate the first pass only did in CSS at
599px: SPLIT_PANE_MIN_WIDTH (1180, matching HOME_SESSIONS_MIN_WIDTH)
now backs an actual JS width check in _applySplitButtonVisibility,
with a matchMedia listener so a live window resize hides/shows the
button without a reload — the CSS backstop in styles.css is the
reverse-direction guarantee for when JS hasn't run.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:32 -05:00
timkjrandClaude Sonnet 5 6c8bd6c606 fix(split-pane): gate the Split button to desktop, make it per-device
Ark0N's PR #453 review: nothing gated this feature to desktop even
though the design called for it (two 240px min-width panes plus the
divider need ~486px, and the divider has no touch handlers), and
showSplitButton was a SYNCED setting, so turning it on at a desk also
put the button in the phone header.

- Hard-hide .btn-split on phones in mobile.css regardless of the
  setting, matching the other desktop-oriented header buttons in the
  same @media (max-width: 599px) block.
- Move showSplitButton into settings-ui.js's per-device displayKeys
  set and drop it from SettingsUpdateSchema entirely, matching the
  showFileViewerButton/skin precedent (CLAUDE.md's "per-device keys
  ... must NOT be added to SettingsUpdateSchema" rule) — a desktop
  opt-in must never sync onto a phone that never asked for it. Removes
  the now-invalid server-round-trip test for the setting.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:32 -05:00
timkjrandClaude Sonnet 5 8d3bde5469 fix(split-pane): address the rest of Ark0N's PR #453 review
- Exclude popped-out (detached) sessions from the split picker:
  SplitTerminalPane._sendResize() has no yield-to-detached-window check
  the way the primary pane's sendResize() does, so splitting against a
  detached session put its own window and Pane B in a fight over the
  same PTY's dimensions. Simplest fix per the review: keep them out of
  buildSplitPickerSessions() entirely.
- Show a visible dead state when Pane B's WebSocket drops. onData
  already silently discards keystrokes while the socket isn't OPEN
  (there is no reconnect for v1), so a dropped socket left the pane
  looking normal while it quietly ate everything typed into it.
- openSplitPane() returns early with no active session, so a split
  triggered from the home screen no longer creates and connects Pane B
  behind the opaque welcome overlay with nothing to show for it.
- onMove() during a divider drag now bails when the split has
  auto-collapsed mid-drag (the other pane's session ending) instead of
  throwing on `divider.parentElement` being null.
- Promote Pane B via `selectSession(id, { auto: true })` when Pane A's
  session ends — this is an app-driven selection, not the user clicking
  a tab, so it must not spend the promoted session's idle alert.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:31 -05:00
timkjrandClaude Sonnet 5 78dcb0aa24 fix(split-pane): refit Pane B when the window/sidebar/tab-rail resizes
Ark0N's PR #453 review: fit() was only ever called from the divider
drag, and the trailing-edge ResizeObserver callback in terminal-ui.js
(throttledResize) only ever measured Pane A's own container. Split at
a wide viewport, shrink the window (or toggle the Alt+B sidebar, or
drag the tab rail), and Pane A's cols changed while Pane B silently
kept its stale PTY size in both xterm and the real pane.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:31 -05:00
timkjrandClaude Sonnet 5 7fc66e8161 docs(split-pane): keep the design spec, drop the task-plan scaffolding
Per Ark0N's review on PR #453: rename the design spec to
docs/split-pane-sessions-plan.md, matching every other feature's
*-plan.md convention, and drop the 957-line implementation task plan
(docs/superpowers/plans/2026-09-15-split-pane-sessions.md) — workflow
scaffolding for the subagent-driven-development run, not repo
documentation. Fixes the now-dangling link in architecture-invariants.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:30 -05:00
timkjrandClaude Sonnet 5 1678386f50 test(split-pane): cover the blank-Pane-B and stale-width-Pane-A fixes
Real-browser regression coverage for the previous commit:

- SplitTerminalPane connects onto an already-quiet session and shows its
  existing scrollback with no new output, proving the ?full=1 fetch (not
  a live echo) populated the pane.
- openSplitPane() force-resizes Pane A synchronously as part of opening
  a split.
- Dragging the divider force-resizes Pane A once, at drag end.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:30 -05:00
timkjrandClaude Sonnet 5 3859506f9b fix(split-pane): populate Pane B history and force-resize Pane A on split changes
Pane B's SplitTerminalPane.connect() only opened a WebSocket and waited for
live output — ws-routes.ts's terminal socket sends nothing on connect, only
future 'terminal' events — so it stayed blank until the target session
happened to produce new output. It looked intermittent rather than
always-broken because a resize sent by _sendResize() often nudges the
session's real tmux window to a new size, and tmux repaints its current
screen on resize; that incidental repaint was what usually populated the
pane. When Pane B's computed dimensions already matched the session's
last-known size, Session.resize() skipped the resize as a no-op and the
pane stayed empty. Fetch the existing scrollback (?full=1) before opening
the socket, same as the primary pane does.

Pane A never told its own session's PTY/tmux about a size change at all,
relying purely on the passive 300ms-debounced ResizeObserver in
terminal-ui.js. openSplitPane() now force-resizes Pane A immediately on
entering split (mirroring closeSplitPane()'s existing symmetric call), and
the divider-drag handler force-resizes it once at drag end (matching the
codebase's established trailing-edge debounce convention rather than
flooding a resize per mousemove).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:29 -05:00
timkjrandClaude Sonnet 5 33b2605815 fix(split-pane): stop leaking document listeners on repeated split-picker toggles
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:29 -05:00
timkjrandClaude Sonnet 5 d0a887d98a test(split-pane): add fast unit coverage for the _onSessionDeleted auto-collapse ordering
test/split-pane-auto-collapse.browser.test.ts covers "Pane B's session ends"
in a real Chromium, but that suite is excluded from the npm test CI gate.
The "Pane A's session ends, Pane B gets promoted" branch had no coverage
anywhere, and it is the one branch whose correctness depends on exact
ordering: _splitSessionId must be captured BEFORE closeSplitPane() runs
(which nulls it) or the promoted session id is lost. Loads terminal-split.js
via `vm` against a minimal fake CodemanApp (same technique as
test/session-close-fallback.test.ts), and pins all three branches (Pane A
ends, Pane B ends, unrelated session ends) plus that the original
_onSessionDeleted always still fires.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:29 -05:00
timkjrandClaude Sonnet 5 24a92c8f3e fix(split-pane): refuse to split a session against itself
Nothing stopped a stale picker click (opened before switching tabs) or
clicking Pane B's own session tab while split from landing on
openSplitPane(sessionId) with sessionId === activeSessionId, or from
selectSession() rebinding the primary pane onto the session Pane B was
already showing — either way, two live WebSockets to one session, each
independently claiming PTY dimensions via its own {t:'z',...} resize frame.
openSplitPane() now refuses early when the target is already the active
session, and a new selectSession() prototype patch (same top-level pattern
as the existing _onSessionDeleted patch) closes an active split BEFORE the
primary pane rebinds to the session Pane B holds.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:28 -05:00
timkjrandClaude Sonnet 5 f7852081b7 docs(split-pane): fix orphaned Session list layout section
The new "Split-pane sessions" section was inserted between the "Session
list layout (header strip vs. left sidebar)" heading and that section's own
body paragraphs, orphaning the heading from its content. Move "Split-pane
sessions" to after the Session list layout section's full body, before
"Gesture control: the setting" — no change to the Session list layout prose
itself, only where the new section sits relative to it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:28 -05:00
timkjrandClaude Sonnet 5 f0e24d8ce2 fix(split-pane): hide the split container together with the rest of the terminal on a web tab
.main.webview-active hid .terminal-wrap when a web tab became active, but
.terminal-wrap is reparented INSIDE .terminal-split-container while a split
is open, so Pane B and the divider stayed stranded on screen over the
dashboard iframe. Hide the whole split container as one unit, mirroring the
existing .terminal-wrap rule; no state is destroyed, so returning to the
session tab shows the split intact.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:28 -05:00
timkjrandClaude Sonnet 5 2724c922ce fix(split-pane): style and dismiss the split-picker menu
.split-picker-menu/-item/-empty (created in openSplitPicker()) had zero CSS
and could not be dismissed except by picking an item — a default-path defect
since the Split button ships enabled to anyone who flips showSplitButton on.
Add CSS matching the sibling .run-mode-menu popover's look (floating-bg
backdrop blur, border, shadow, z-index 1000 above the header's 100), and
dismiss on outside click or Escape via the same one-shot listener pattern
session-ui.js already uses for its other transient popovers
(toggleCaseSettings(), toggleRunModeMenu()). Picking an item now routes
through the same _dismissSplitPicker() method as the outside-click/Escape
handlers, so the listeners never outlive the menu.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:27 -05:00
timkjrandClaude Sonnet 5 57406f6c14 fix(split-pane): stop clamping Pane B's resize dimensions to a 40x10 floor
_sendResize() clamped Pane B's proposed cols/rows to a 40/10 floor before
sending the {t:'z',...} resize frame, so the PTY was misinformed of Pane B's
real width at the divider's own reachable 20% position, causing real
output-wrapping bugs. The primary pane (terminal-ui.js's
getTerminalDimensions()) sends fitAddon.proposeDimensions() unclamped and
lets the server enforce its own valid range ([1,500]/[1,200] in
ws-routes.ts); Pane B now matches that convention.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:27 -05:00
timkjrandClaude Sonnet 5 8903a72662 fix(split-pane): hide the Split header button when showSplitButton is off
.btn-split--hidden had no matching CSS rule anywhere, so the opt-in Split
header button shipped visible to every user on every viewport regardless of
the setting. Add the `display: none !important` rule alongside its sibling
marker classes (.btn-multimonitor--hidden etc.), plus a static regression
guard (test/split-pane-hidden-button-css.test.ts) that fails if any future
"*--hidden" marker class in index.html is missing a matching CSS rule.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:27 -05:00
timkjr ba7b8b7bef docs: add split-pane sessions architecture-invariants entry 2026-09-20 13:10:26 -05:00
timkjr 8c73128cd6 feat(split-pane): auto-collapse split when either session ends 2026-09-20 13:10:26 -05:00
timkjrandClaude Sonnet 5 2abf328db8 feat(split-pane): add open/close orchestration, picker, and divider drag
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:25 -05:00
timkjrandClaude Sonnet 5 fa8bb13a27 fix(split-pane): reset _wsReady on WS close/error in SplitTerminalPane
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:25 -05:00
timkjrandClaude Sonnet 5 6f64e557e5 docs(plan): fix session-creation test bug found by Task 4's implementer
Task 4's implementer found two real bugs in this plan's browser-test
helpers: POST /api/sessions nests the id at data.session.id (not
data.id), and mode:'shell' needs a follow-up POST .../shell to actually
spawn a PTY. Fixed in Task 4's own snippet (documentation accuracy —
already fixed in the real committed code) and pre-emptively in Tasks
5/6's createShellSession() helper before either was dispatched, so
neither implementer has to rediscover it independently. Also corrected
the <script> tag snippet to defer, matching the real file's convention.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:25 -05:00
timkjrandClaude Sonnet 5 97a1238c85 feat(split-pane): add SplitTerminalPane class for Pane B
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:24 -05:00
timkjr aa0521602d feat(split-pane): add split container/divider/pane-b CSS 2026-09-20 13:10:24 -05:00
timkjr 2bc16d5fd9 feat(split-pane): add showSplitButton setting and header button 2026-09-20 13:10:23 -05:00
timkjr d60a164025 feat(split-pane): add pure divider-clamp and picker-list helpers 2026-09-20 13:10:23 -05:00
timkjrandClaude Sonnet 5 727817410c docs(plan): fix Task 6's SSE handler patch to target the prototype
Monkey-patching the instance's _onSessionDeleted inside a
DOMContentLoaded listener races connectSSE()'s handler-wrapper cache,
which captures the function reference by value on first connect and
never re-reads it. Patching CodemanApp.prototype at module-evaluation
time (synchronous script-tag order) is unraceable: it completes before
any instance exists or connectSSE() ever runs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:22 -05:00
timkjrandClaude Sonnet 5 28d3bd7da8 docs(plan): fix Task 2/4/5/6 tests against real test infrastructure
Preflight scan for SDD execution caught two classes of defect before
dispatch: Task 2's test invented a buildTestApp() helper and response
envelope that don't exist for /api/settings; Tasks 4-6 used
@playwright/test's runner against a test/browser/ directory that
doesn't exist in this codebase. Both corrected against real patterns
found in existing tests (system-routes-settings-partial-put.test.ts,
terminal-copy-shortcut.test.ts, tab-rail-resize.browser.test.ts).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:22 -05:00
timkjrandClaude Sonnet 5 d0f9bdd251 docs: fix plan wording and add execution-environment note
Global Constraints previously read as if local-echo/CJK/accessory-bar
were desktop features; they are mobile-only, and split-pane is the
desktop-only side of that equation. Also names the exact spec section
instead of a loose paraphrase, and adds a worktree/branch note so an
executing subagent knows where this plan runs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:22 -05:00
timkjrandClaude Sonnet 5 ff98006471 docs: add split-pane sessions implementation plan
7 tasks: pure divider/picker helpers, showSplitButton header wiring,
split-container CSS, SplitTerminalPane (Pane B's independent xterm+WS),
open/close orchestration with picker and divider drag, auto-collapse on
either session ending, and an architecture-invariants entry.

Also folds in the "detach session" prior art discovered mid-brainstorm
into the spec's architecture section.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:21 -05:00
timkjrandClaude Sonnet 5 5f1be90ae9 docs: fix tab/pane terminology in split-pane spec
The Problem paragraph and the architecture section used "tab" where
"pane" was meant, colliding with the browser's own tab concept.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:21 -05:00
timkjrandClaude Sonnet 5 3da8bb7046 docs: add split-pane sessions design spec
Scopes v1 of an in-app split view (two live session panes side-by-side,
draggable divider) after multi-monitor spanning turned out to solve a
different problem than showing multiple panes at once.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 13:10:20 -05:00
DodgyBadger ac574d6c64 fix(mobile): compact the compose key 2026-09-20 17:31:10 +00:00
DevvynandClaude Sonnet 5 d9e6ebb20a fix(cli-registry): address round-2 review on #458 — count-based allowlist, RUN_MODE_LAUNCH drift guard
Three of Ark0N's four "will take at merge" items, applied instead since
they were straightforward to do properly:

1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on
   <file>::<expression> (fixed last round) closed the line-shift problem
   but opened a new one: every stock id was already allowlisted for
   session-ui.js in the `mode === '<id>'` form, so a BRAND NEW branch
   reusing that exact expression anywhere in the file passed unnoticed.
   Reproduced live (`if (this.mode === 'codex')` injected into
   runOpenCode()) — stayed green under the old version. Each allowlist
   entry now carries the exact count of approved call sites, and a new
   test asserts actual-vs-declared count for every key; a mismatch in
   either direction is real (higher = new unreviewed branch riding in on
   an existing approval, lower = a reviewed site was removed and the
   entry is now stale). Reproduced again against the fix: same injection
   now fails with an exact diagnostic (expected 2, found 3).

2. Added test/run-mode-launch-table-drift.test.ts. RUN_MODE_LAUNCH
   restates four things stock.ts already owns (label, install command,
   supportsCustomModel, the external-mode key set), and they agree today
   with nothing enforcing it. supportsCustomModel is the dangerous one:
   the Run-menu picker's rows come from the server-injected
   window.__codemanCustomModelClis (built from
   capabilities.customModelInjection.kind), so a CLI gaining a real
   injection recipe later would be OFFERED in the picker while
   _runCliMode silently drops the customModel field for it — the session
   launches on the vendor's cloud while the UI claims the local endpoint.
   Drives the real session-ui.js via JSDOM and compares RUN_MODE_LAUNCH
   against STOCK_CLIS on all four axes.

3. Inlined the "Open Question 7 in PR-B2.md" references in the allowlist
   reasons — PR-B2.md is a local planning doc, never part of the
   committed tree, so the reference was dead on arrival for anyone
   reading the repo. Points at the PR #458 review thread instead.

4. Added a sentence to docs/cli-registry.md naming the new frontend guard
   alongside the backend one it mirrors.

Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format/
check:frontend-syntax all clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-20 21:37:05 +08:00
DevvynandClaude Sonnet 5 88e5b7b200 fix(custom-model): address Ark0N's PR review — client-side probe timeout, defer "last used" past confirmation, docs, zh-CN
Four things from the maintainer's review on PR #459, all fixed:

1. Bound _getCustomModelCurrentlyLoaded's probe client-side (~800ms via
   Promise.race, on top of — never instead of — the route's own 5s
   server-side timeout). Without it, an asleep/firewalled endpoint behind
   a saved model list left the picker completely invisible for up to 5s
   after the Run menu had already closed, with no spinner or toast.
   `timeoutMs` is an optional param (default 800, real callers never pass
   it) so a test can drive it in milliseconds, same pattern as
   `_watchLlamaSwapLoading`'s own `pollIntervalMs` — this code runs in a
   JSDOM window's own realm, whose setTimeout vi.useFakeTimers() cannot
   patch.

2. "Last used" is now written only once a launch actually applies, never
   on the mere click. It moved out of runCustomModelEntry (unconditional)
   and into each path's own success point: _quickStartWithCustomModelConfirm
   after the final post succeeds, and _runCustomModelEntryViaRestart right
   after the apply's success check. A context-window-warning decline means
   this exact model cannot work with this CLI at all, so the old
   unconditional write would promote, next time the picker opened, the one
   model guaranteed to fail again.

3. Documented the promotion/tag precedence and the new
   codeman:customModelLastUsed:<mode>:<endpointId> localStorage key in both
   CLAUDE.md's Custom Model Endpoint Profiles section and
   docs/custom-model-endpoints.md's Run-menu picker section.

4. Added zh-CN entries for "Currently loaded" and "Last used" in i18n.js,
   next to this modal's existing "Choose a model"/"Custom Endpoints" pair.

New tests: the client-side timeout (endpoint that never answers, one that
answers within the bound, and a rejected-after-timeout probe settling
quietly), and "last used" recording on success vs. NOT recording on either
confirmation's decline, for both the restart and one-shot paths.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
2026-09-20 21:19:42 +08:00
DevvynandClaude Sonnet 5 73607663fd fix(custom-model): guard the picker's async open against a slower, superseded probe
Code review (high effort) on the previous commit found a real race: making
_openCustomModelPickModal async (it now awaits the currently-loaded-model
probe before rendering) meant a second, faster call for a different
endpoint could render first, only for the first call's slower probe to
resolve afterwards and overwrite the modal with the wrong endpoint's model
list — while _pendingCustomModelPick (set synchronously, before either
await) still named the second, correct endpoint. Picking a model in that
state would launch/apply the wrong model on the wrong endpoint.

Fixed with the same mutable-generation-counter guard
_watchLlamaSwapLoading already uses for an identical async-superseded-by-
newer-call shape: every DOM write, including _pendingCustomModelPick
itself, is deferred until after the awaited probe, and a call that finds
its generation already superseded bails out untouched instead of clobbering
whatever a newer call already rendered.

Added a regression test driving two overlapping opens with a controlled
promise so the earlier, slower probe resolves after the later, faster one
renders, asserting the late response is a no-op.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
2026-09-20 19:27:07 +08:00
DevvynandClaude Sonnet 5 458ca578e7 feat(custom-model): promote the currently-loaded/last-used model in the Run-menu picker
Custom Model Endpoint Profiles' "which model" picker (session-ui.js's
_openCustomModelPickModal) always listed models in their raw discovery
order, so on a host with several downloaded GGUFs the user had to
remember (or eyeball the "Default" tag) which one llama-swap actually
had hot before picking — the whole point of the picker being fast is
undone if it makes you think first.

The picker now promotes exactly one model to the top of the list:

- If llama-swap reports a model from this host's own list `ready`
  right now (via the existing GET /api/model-endpoints/:id/running-status
  route), that model is promoted and tagged "Currently loaded" — it's
  what a launch attaches to with zero wait.
- Otherwise, the last model actually launched on this exact
  (harness, endpoint) pair is promoted and tagged "Last used", read
  from a new per-device localStorage key
  (codeman:customModelLastUsed:<mode>:<endpointId>), written by
  runCustomModelEntry on every launch attempt regardless of outcome.
- A plain (non-llama-swap) OpenAI-compatible server, an unreachable
  endpoint, or a loaded-but-not-yet-ready model never promotes
  anything — the rest of the list keeps its discovery order.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
2026-09-20 19:20:45 +08:00
DodgyBadger 884713cca5 fix(mobile): retain oversized composer drafts 2026-09-20 07:21:20 +00:00
DodgyBadger a220c28a14 fix(input): count code points when clearing prompts 2026-09-20 07:18:49 +00:00
DodgyBadger 0761de3dae fix(mobile): use raw fallback for composed prompts 2026-09-20 07:18:23 +00:00
DodgyBadger c2eaba990b fix(mobile): release echo passthrough after compose 2026-09-20 07:17:41 +00:00
DodgyBadger e214691429 fix(mobile): preserve composer delivery after replay 2026-09-20 07:16:59 +00:00
DodgyBadger 773b405429 feat(mobile): add manual prompt composer 2026-09-20 07:16:59 +00:00
DevvynandClaude Sonnet 5 2df9355367 fix(cli-registry): address PR B2 review — fix two test guards, drop unused catalogue
Two required fixes from Ark0N's review of #458:

1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on
   <file>::<line>::<expression>. A single inserted line anywhere above an
   entry shifted every subsequent line number, so all 21 entries went stale
   simultaneously and the same 21 branches were reported as "new" — on a
   file six other open PRs also touch. Dropped the line number from the key
   (<file>::<expression>, matching the backend guard's own design), which
   collapses 21 line-keyed entries to 11 or-collapse where the same
   expression recurs at multiple call sites in the same file.

2. test/run-mode-ui.test.ts's terminal-ownership guard scanned method
   bodies via `^ {2}async (run[A-Za-z]*)\(\) \{$`, which matched the 8
   one-line run<Mode>() wrappers PR B2 introduced but not _runCliMode(mode),
   where the real logic (and the actual risk the guard exists to catch) now
   lives. Fixed the regex to `^ {2}async (_?run[A-Za-z]*)\(\w*\) \{$` and
   added _runCliMode to the sanity list. Same-class fix in
   test/opencode-resize.test.ts, which had the identical blind spot via
   runOpenCode.toString().

Both reproduced live before fixing (inserted the same comment line; added
this.terminal.clear() to _runCliMode) to confirm the bug, then confirmed
the fix catches it and the suite stays green otherwise.

Also resolves Open Question 2 by dropping window.__codemanCliCatalog
entirely: nothing consumed it, and a registry DECLARED_FOR_LATER field
costs nothing until read while an unconsumed script tag on every page
render is a different trade. Reverts Phase 1 cleanly — server.ts's
injection, shortBadge back in types.ts's DECLARED_FOR_LATER list and the
pinned guard test, and the three associated render-index-html.test.ts /
server-index-title.test.ts assertions.

Full gate: 405 files / 7717 tests / 0 failures (net unchanged), typecheck/
lint/format clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-20 02:04:37 +08:00
DevvynandClaude Sonnet 5 cd64b0a3f7 feat(cli-registry): drive the run-menu frontend from the CLI catalogue (PR B2)
PR #380 (PR B) held back the frontend half of the CLI registry refactor,
explicitly deferring window.__codemanCliCatalog and making session-ui.js /
mobile-overview.js catalogue-driven as "PR B2".

- Inject window.__codemanCliCatalog in renderIndexHtml(), following the
  existing __codemanCustomModelClis pattern (escapeScriptJson-guarded,
  resolved per-request). Reading CliEntry.shortBadge here is what makes it
  genuinely read, so it drops out of types.ts's DECLARED_FOR_LATER list.
- Consolidate session-ui.js's 8 near-duplicate run<Mode>() launch functions
  (opencode/codex/gemini/antigravity/pi/omp/grok/deepseek) into one shared
  _runCliMode() plus a local RUN_MODE_LAUNCH config table. The 8 method
  names stay as thin wrappers (index.html calls them by name; tests assert
  on the name). Also collapses a duplicated 8-way isAltMode/isExternalCli
  OR-chain (same expression, copy-pasted twice in openSessionOptions) into
  one EXTERNAL_CLI_MODES check.
- Add test/frontend-cli-no-id-branching.test.ts, a guard scoped to
  session-ui.js/mobile-overview.js only (not the rest of src/web/public/,
  which stays explicitly out of scope per CLAUDE.md), mirroring the
  backend's own no-id-branching guard.

mobile-overview.js and the wiring of accent/echo/wheelForward/
keyboardAccessory were investigated and deliberately left alone: the first
is already a single, tested, gated table (not duplicated logic); the second
set belongs to terminal-ui.js/keyboard-accessory.js/styles.css, files
outside this PR's mandate.

Verified on a tmux-capable devbox (this sandbox has no tmux): full CI gate
at 405 files / 7717 tests / 0 failures, typecheck clean, 94 targeted tests
covering exact per-CLI wire-body shapes unmodified and passing, and a live
anti-vacuity check on the new guard (injected a real branch, confirmed it
fails, reverted, confirmed green).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-19 21:16:31 +08:00
Devvyn 2d573d8a34 Merge branch 'master' of https://github.com/Ark0N/Codeman into followups 2026-09-19 19:43:45 +08:00
github-actions[bot]Claude Opus 5github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
51b4a1b758 chore: version packages (1.31.0)
* chore: version packages

* chore: sync CLAUDE.md version to 1.31.0

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
2026-09-19 13:17:25 +02:00
Codeman maintainer 4205f6930f fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine.

**The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext`
and `confirmedSwap` changed the wire field without moving three assertions that check
it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts`
(the swap modal and the context modal, each of which already receives exactly the right
per-question flag). Moved, with the titles.

**Worse, my own tests for the split never ran.** The four cases in
`session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`,
which was declared inside a sibling `describe`, so they threw a ReferenceError during
setup. The split would have shipped with no passing server-side coverage while the gate
reported the failure as four broken tests rather than as four tests that were never
written. `mockRunning` is hoisted to the outer describe.

**The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved
its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so
the other eight modes fell back to claude's `❯`. That is also starship's default shell
prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line
`❯ npm run build` sits on screen for as long as the command runs, the verifier reads it
as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to
nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N
prompt, `read -p`, an installer or a pager, where it takes the default. The module's own
fileoverview already stated the rule this broke. Now `?? ''`, which
`promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI
that actually declares a composer.

**My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing
three numbered rules.

The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared
in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told
an integrator to retry with `confirmed: true` for both questions, which is precisely the
thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md`
and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR`
multi-user consequence, both user-visible and both previously absent, and #454's gained
the one exception to its own claim: a Custom Endpoints launch ignores the Instance count
stepper and always starts one session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:58:33 +02:00
Codeman maintainer 12de3c5164 docs(custom-model): record the CLAUDE_CONFIG_DIR clamp in architecture-invariants
CLAUDE.md gained the admin-only note when the key joined claude's privilegedEnvKeys;
architecture-invariants, which is where the exact-key allowlist rule is documented in
depth, still described the pre-change world. The reboot-restore half is the one worth
writing down: a non-granted owner's already-persisted CLAUDE_CONFIG_DIR is stripped on
restore, which moves that session back to the default Claude account with no error.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:35:32 +02:00
Codeman maintainer 9af12afb57 docs(custom-model): make the docs match the code, and trim the changeset
More from the review of 5fc391a4, all documentation rather than behaviour.

The changeset was 1602 words of development log, written as the PR grew, with bullets
and loose paragraphs interleaved. That text becomes CHANGELOG.md and the GitHub release
body verbatim, so it is now one user-facing account of what the feature does and what
the real-server work bought, at roughly a fifth the length.

docs/api-reference.md promised a `cmd` field on running-status that the route
deliberately strips (it carries model paths and can carry --api-key).

Two places claimed the apply routes validate `modelId` against the endpoint's
discovered models. Neither does. Dropped the claim rather than adding the check:
discovery can be up to five minutes stale, so a 400 there would refuse a launch that
actually works, and a typo'd id already fails on the CLI's own first request. CLAUDE.md
now says so explicitly, since the absence is the surprising part.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:34:24 +02:00
Codeman maintainer fe3bd0074c fix(custom-model): split the two confirmation questions, and seed the API key the way claude reads it
Two findings from the review of 5fc391a4, both fixed here rather than sent back.

**The API-key trust seed never matched a real key.** `seedApiKeyTrustFile()` wrote the
key verbatim into `customApiKeyResponses.approved`, but Claude Code stores and compares
only the last 20 characters (`key.trim().slice(-20)`, applied on both the write and the
lookup). For any real key the seed missed, so claude stopped at the interactive
"Detected a custom API key in your environment" prompt, whose default is
"No (recommended)": the launch hangs, or silently refuses the key this feature just
injected and falls through to an OAuth login the isolated config dir does not have. It
survived review because a keyless llama.cpp/llama-swap endpoint uses DEFAULT_API_KEY
('local-dummy-key', 15 chars), where slice(-20) returns the whole string and the seed
matches by accident, and every test used a key shorter than that. Now truncated through
`truncateApiKeyForTrustFile()`, with a test using a 57-character key that also asserts
the full credential never reaches that second file.

**One `confirmed` flag answered two different questions.** The context-floor warning
("this model's window is below what this CLI needs") and the swap-conflict warning
("loading this unloads the model another session is using") shared it, and the context
check runs first, so a user clicking "launch anyway" past the context warning silently
consented to evicting someone else's model. They are about different people, so an
answer to one is not consent to the other. Both routes now read `confirmedContext` and
`confirmedSwap` independently; the legacy `confirmed` still means both, because it
shipped in this feature's HTTP-API-only cut and an existing caller must keep working.
The frontend answers each question with its own flag and accumulates them, on the
one-shot path, the restart path and the batch carry-forward alike.

Also from the same review: the swap-confirm dialog no longer renders " are currently
using ..." when multi-user scoping leaves the affected-session list empty (the swap is
blocked regardless of ownership; only the NAMES are scoped), and the per-endpoint
llama-swap log tails are closed in `WebServer.stop()` instead of only by the idle sweep
whose interval that same teardown disposes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:32:50 +02:00
Codeman maintainer 3b55957d79 fix(custom-model): merge-time fixes for the Run-menu picker
Conflict resolution against the five PRs that landed while this was in review, plus
the items left for merge on the thread.

The real one was `session-ui.js`. #454 refactored all eight non-Claude `run*()`
functions to funnel through one `_launchQuickStartInstances()` helper that does the
POST itself, while this PR replaced that same POST in each of them with
`_quickStartWithCustomModelConfirm()`. Resolved in the helper rather than seven times
over: the helper now goes through the confirm path, and each body builder carries the
`customModel` spread. `runAntigravity` deliberately does NOT, since antigravity's
`customModelInjection` is `unsupported`; parity with this PR's own per-mode choices is
asserted rather than assumed.

That merge creates a question neither feature had alone: the confirm dialog now runs
inside a loop that can launch up to 20 instances. Both questions it can ask (context
window too small, and loading this will unload the model another session is using) are
decisions about the ENDPOINT, and every instance in a batch targets the same one, so
the answer is taken once and carried to the rest. Without that a 20-instance launch
asks the same question 20 times.

Also: `sse-events.ts` is 161 constants (master added two for remote wake, this adds
one, verified by counting rather than by arithmetic), `server.ts` keeps both new SSE
prefixes, the two comments pointing at code that no longer exists are corrected, and
CLAUDE.md's SSE and route counts move to 161 / ~236 / custom-model (6).

`pumpLlamaSwapLogTail`'s unparsed remainder is now capped at 64 KiB. It only shrank at
a `\n\n` frame boundary, so a backend that streams without one would grow it for the
life of a deliberately indefinite connection.

NOT changed, deliberately: the context warning and the swap-conflict warning still
share one `confirmed` flag with the context check first, so confirming "launch anyway"
on a too-small context also skips the "this unloads it for another session" ask. That
is the author's documented choice and the reviewer's own note calls it minor. Both
fixes are worse to make here than to defer: separate flags are new wire surface landed
unreviewed during a release, and reordering the checks adds a network round trip to a
path that currently short-circuits. Raised as a follow-up instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:27:01 +02:00
Codeman maintainer 1a99b5836c Merge pull request #430 from opticon454/custom-model-run-menu 2026-09-19 12:25:11 +02:00
Codeman maintainer 035bfbc2fe fix(remote): merge-time fixes for Wake-on-LAN
The MAC-count limit lived in two places that disagreed. RemoteHostSchema.wakeMac's
128-character cap admits seven comma-separated MACs while parseMacList takes at most
four, all-or-nothing, so a five-MAC value validated, was written to remote-hosts.json,
and then resolved to NO wake target: POST /api/sessions/:id/wake answered
"No wake-on-LAN target configured for this host" and the banner offered "Configure WoL"
for a host the user had just configured. MAX_WAKE_MACS now lives in
src/config/remote-wake-limits.ts and both sides refine against it. Its own module
because src/remote-wake.ts is import-fenced to session-routes.ts and server.ts (the
wiring guard that stops a watcher waking a host), and because schemas.ts must not drag
dgram/net/child_process into every request-validating module.

The documented 40 s request budget also omitted the wake's own cost. A `command` target
is bounded by REMOTE_WAKE_COMMAND_TIMEOUT_MS and runs BEFORE the readiness poll, so a
slow one pushed a wakeCommand host's worst case to ~68 s, past the 60 s
proxy_read_timeout the budget exists to stay under. _wakeAndWait now subtracts the
wake's measured elapsed time from the readiness budget, floored at one poll interval so
a wake that ate the whole budget still gets one probe. A magic packet is effectively
instant and is unaffected, which is why live testing never saw it.

Also: the two new endpoints are documented in docs/api-reference.md with the import
fence stated as the rule it is, CLAUDE.md's frontend module count moves to 34, and the
release changesets carry the Thanks section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:40 +02:00
Codeman maintainer 4c705094f7 fix(terminal): ship the copy clean as a trailing trim, without the shared dedent
#451 cleaned two things on copy. The trailing trim is right and every native
terminal does it. The shared leading-indent strip is this project's own rule,
and it is dropped here rather than shipped.

Measured against the shipped transform over 401,445 three-row windows across
1,010 tracked files in this repo, it fired on 73% of them: 92% inside a YAML
workflow, 76% over `git log` output, 48% in a TypeScript source. No width
threshold separates a margin from content because they are the same widths, a
live Claude Code pane's own margins measuring 2 and 5 columns while the most
common non-TUI shared run is 4. The failure modes are not symmetric either: a
wrong trailing trim costs nothing, while a wrong dedent silently deletes
information that was on the screen, with nothing in the clipboard to hint at
it, on git log bodies, on indented code read out of cat (semantic in Python),
on git diff context rows where the leading space is the marker, and on stack
traces.

It also could not be made self-consistent cheaply. Whether the first row joined
the measurement depended on the mousedown COLUMN, which the user never sees, so
one block of three rows produced three different clipboard results; and the
flag read getSelectionPosition().start, which is xterm's mousedown anchor and
is never normalised, so dragging UP through a block read it off the bottom row.
The PR's test stub hardcoded a downward drag, so its suite could not express
that case.

The transform, the wiring, the tests, the invariants, CLAUDE.md, the wiki page
and the changeset all move together. The test block now pins the ABSENCE as a
contract, with the git log, Python and git diff cases as its examples, so this
is not re-derived later. If it is ever revisited, the one qualification that
measured clean is painted trailing padding: zero false positives over all
401,445 windows.

Also from the review: the comments and invariant rule justifying the
padding-only clear described the pre-change code (the Ctrl+C gate reads the
CLEANED selection now, so such a selection falls through to the PTY on its own
and the clear is feedback rather than protection), the new 'Nothing to copy'
toast gained its zh-CN entry, and the invariants paragraph no longer repeats
its own opening sentence.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:39 +02:00
Codeman maintainer c376534a50 fix(run,terminal): merge-time fixes for the Instance count stepper and capture geometry
#454: the behaviour the PR adds had no test, so a regression test drives
runGrok() at tabCount 3 and asserts three quick-start POSTs with sequential
w<n>-<case> names (verified to fail against master's session-ui.js). Each
caller now reads the count BEFORE its opening banner and announces it there,
the way runClaude() already did, so a launch no longer prints two headers and
a launch with another session already active still says how many are starting.
runClaude() calls the shared _readTabCount() instead of its own copy of the
1..20 clamp, and that helper optional-chains the element read, since hoisting
it above each caller's try block would otherwise let a missing #tabCount throw
where the launch-error path cannot report it.

#435: sizeMovedUnderLoad derived from data.source alone. `mux-visible` is not
sufficient: a failed display-message cursor query makes capturePaneBuffer skip
the snapshot repaint and return the raw capture, which the route still labels
mux-visible, so a size that moved during such a load bought a full forced
reload to repair a frame that was never positioned. It now tests
Number.isFinite(data.captureRows) like its two siblings.

Plus the invariants and CLAUDE.md lines promised on #435: a visible capture
reports its geometry and omits it when nothing was positioned, the comparison
runs on mux-visible only, and the replay is capped at one attempt and latches
per session when it cannot converge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 12:18:39 +02:00
Ark0N 2c3ccdf030 Merge pull request #439
feat(remote): wake a sleeping host (Wake-on-LAN) from input, banner and native magic packet
2026-09-19 12:18:18 +02:00
Ark0N 475436242c Merge pull request #455
fix(input): make sure a prompt sent through the API actually leaves the composer
2026-09-19 12:18:13 +02:00
Ark0N 613b774bf1 Merge pull request #451
fix(terminal): trim the padding and shared indent out of a copied selection
2026-09-19 12:18:08 +02:00
Ark0N 60c9af0599 Merge pull request #435
fix(terminal): replay a pane capture at the geometry it was taken at
2026-09-19 12:18:03 +02:00
Ark0N 2d842ded35 Merge pull request #454
fix(run): make the Instance count stepper work for every non-Claude mode
2026-09-19 12:17:58 +02:00
DevvynandClaude Sonnet 5 5fc391a47c fix(custom-model): address fourth pre-merge review + merge upstream master (Ark0N)
Merged upstream/master (22 commits: reboot-restore recovery feature,
terminal keycode229 recovery work, install.sh/CLI-catalog generator
changes, CHANGELOG/version bump to 1.30.0) into this branch. No
conflicts; git auto-merged every overlapping file (CLAUDE.md,
docs/api-reference.md, app.js, index.html, styles.css, routes/index.ts,
session-routes.ts, schemas.ts, server.ts).

Two required fixes from the latest review:

1. privilegedEnvKeys widening (stock.ts) changes behaviour outside this
   feature. The reviewer decided to keep both CLAUDE_CODE_MAX_CONTEXT_TOKENS
   and CLAUDE_CONFIG_DIR listed (types.ts's rule that every traffic-
   redirecting var this feature introduces must appear there stays
   literally true), and asked for the real consequences documented
   instead of hidden:
   - Corrected session-env-clamp.ts's fileoverview, which stated the
     opposite of what the code now does (reboot-restore's clamp call
     used to be able to strip nothing for claude; it now strips a
     persisted CLAUDE_CONFIG_DIR for a non-granted owner).
   - Corrected the rationale comments in stock.ts: privilegedEnvKeys
     has exactly one consumer (ownerClampedEnvKeys, feeding the
     generic envOverrides clamp on create/quick-start/reboot-restore),
     not the custom-model routes.
   - Added a CLAUDE.md line to the CLAUDE_CONFIG_DIR gotcha covering
     the admin-only-in-multi-user-mode and reboot-restore-strips-it
     consequences.
   - Added a "Claude multi-user clamp" test next to the existing
     DeepSeek/OMP ones, pinning the new stripping behaviour.

2. GET .../running-status (custom-model-routes.ts) no longer passes
   the raw llama-swap `cmd` field (the literal launch line, which can
   carry model paths and --api-key) to the browser -- the frontend
   only ever reads model/state, cmd exists solely for server-side
   parseCtxFromCmd() during discovery. Added a test asserting the
   response never contains cmd or a planted secret.

Also regenerated config/clis.stock.json and install.sh's catalogue
block (npm run generate:cli-catalog) to clear drift introduced by the
upstream merge, since it was failing the sync check.

Left to the reviewer, as they said they'd take at merge: the two
"comments pointing at removed code" cleanups, the two stale CLAUDE.md
counts, and the small items list (mode==='claude' frontend branch,
isCliAvailable() unknown-id gap, shared confirmed flag ordering,
one-shot cancel toast severity, pumpLlamaSwapLogTail buffer cap).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-19 18:05:51 +08:00
RandalixandClaude Opus 5 5bb489addb fix(remote): authorize the attach wake first; tell the caller what happened to its bytes
Review round 3 on #439.

- The attachRemoteSession branch of POST /api/sessions ran `ensureHostAwake`
  before the multi-user gates, so a non-admin could have any configured
  host's `wakeCommand` spawned (or a packet broadcast) and the request held
  for the wake budget, then be refused for the workingDir. The admin gate
  now comes first, before the host is even looked up; remote hosts are
  admin-only infrastructure everywhere else. Route test: wake spy empty,
  403.
- The non-wait input route answers `{buffered:true}` when the registry took
  the chunk and `{buffered:true, dropped:true}` when it was over the cap
  and is gone (`RemoteInputOutcome` gains 'dropped'); additive to the bare
  `{}`.
- The send-and-wait path answers OPERATION_FAILED when the host never comes
  back, like create and attach, instead of writing into the stalled pane
  and reporting delivered:true plus a timeout.
- The flush writes with `fromUser: true`, so a first prompt buffered
  through a wake can still name the tab.

Docs: api-reference (input route), remote-sessions.md (two invariants),
CLAUDE.md key pattern.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-19 11:39:50 +02:00
Michael GrundbergandClaude Opus 5 95dc6fe944 fix(terminal): remember a geometry replay that did not converge
`resizeRetry` caps the recursion inside one select and says nothing about the
next one, so a pane this browser cannot size reported the same mismatch on
every select and bought the same failed repair each time: two fetches per tab
switch for the life of the page, measured as a running count of 2, 4, 6 across
three selects. That is the case this branch describes as happening every time
rather than occasionally, a phone whose resize `Session.resize` declines while a
desktop claim is live, and it is not the only one — any pane Codeman cannot size
lands there, including one a second tmux client is also holding. Each wasted
pass costs another `capture-pane`, which is `execSync` and blocks the server's
event loop, plus a reset and chunked rewrite, a discarded snapshot and cache
entry, and a dropped and reopened WebSocket.

`_geometryRetryUseless` mirrors the existing `_fullHistoryRepullUseless`: a
retry pass whose frame still does not fit adds the session, geometry that fits
removes it, and the replay gate consults it. The proof has to come from a retry
pass rather than a first one, because the retry ran at the size that stuck and
the pane ignored it. Clearing on a fitting frame is what stops a pane that
becomes sizeable again, once the desktop tab closes or its claim goes idle, from
staying permanently unrepaired. The race case never reaches the latch, since it
converges on its first attempt.

The new browser case walks all of that: three selects reading 2, 3, 4 instead of
2, 4, 6, then a fitting frame, then a mismatch diagnosed afresh. Without the
gate it fails on the second switch with `expected 4 to be 3`.

Rebased onto master, which has moved to 1.30.0 and taken #436. The one conflict
was `config/test-suites.ts`, where both branches appended a glob to
`BROWSER_TEST_GLOBS`; both are kept. Everything else merged clean, #436's own
changes to the same buffer-load path included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:56:58 +02:00
Michael GrundbergandClaude Opus 5 383f834704 fix(terminal): flush unsent local echo before the geometry replay
On a touch device the characters the user has typed live only in the local-echo
overlay until Enter; they have never reached the PTY. The replay re-enters
`selectSession` with `forceReload` on the session that is still active, and that
branch nulled `activeSessionId` before `_cleanupPreviousSession` ran. The flush
there is guarded on a session it can still see, so it was skipped, and the
unconditional `_localEchoOverlay.clear()` that follows took the characters with
it. Measured in chromium against the previous head: typing into the overlay and
then making the call the replay makes left `pendingText` empty with nothing
crossing into the delivery layer on either transport.

The flush moves into `_flushLocalEchoTo(sessionId)`, called from both
`_cleanupPreviousSession` and the `forceReload` branch before it nulls the id.
The session is a parameter because the two callers mean different ones: cleanup
flushes to the tab being left, the branch to the tab being reloaded.

This was reachable before this branch, through the one gesture that already
takes the `forceReload` path on an active session. What is new is that nothing
the user does triggers it. The replay fires on its own the moment a tab switch
finishes, which is exactly when someone typing into a still-loading terminal has
text in the overlay, and on a phone beside an active desktop tab that is every
tab switch.

A seventh browser case pins it: it forces the overlay on, since headless
chromium reports no touch support and the case would otherwise pass vacuously,
asserts the typed characters really are sitting unsent, then triggers the replay
and asserts they reached the session. Without the fix it fails with nothing
delivered at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:44:35 +02:00
Michael GrundbergandClaude Opus 5 e0d4477edc fix(terminal): keep the geometry replay to the pass that can converge
Three follow-ups to the source gate, each one measured rather than reasoned.

A pane already drawing at the size the client just requested is left alone. The
replay runs at `dimsAfterLoad`, so it can only change what is on screen if the
pane was drawing at some other size; when the reported geometry already IS that
size, the second pass captures the identical frame and pays a full reload to do
it, including a visible re-flash, a dropped and reopened WebSocket and a deleted
xterm snapshot. That equality is the signature of a clamp rather than a race:
`getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` does not, so a
terminal narrower than 40 columns or shorter than 10 rows reports a pane
permanently bigger than itself and replayed on every tab switch without ever
converging. A race never produces the equality, since its premise is that the
pane was still at the size it was asked to leave. The declined-resize case does
not produce it either, so that one still costs the single capped attempt and
needs the pane-ownership question this does not touch.

The full-history re-arm is unreachable and now says so. A pass that consumed the
flag sent `full=1`, and the route answers `full=1` with `mux-full-history` or
`history`, never `mux-visible`, so the source gate already rules out every such
pass. The line stays for the invariant, but its comment no longer reads as if a
page load retries, and the suite pins that it does not.

The response no longer reports geometry for a body that carries no capture. The
full-history path writes `capturedGeometry` from the cursor query and then
returns '' for a pane holding nothing visible, which drops the source to
`history` with the geometry already recorded: a `full=1` request whose capture
reported 100x50 and returned nothing answered `source: "history"` with both
fields set. Nothing acted on it, because the client ignores geometry on any
other source, but the field said a frame had been drawn at a size when none had.

The browser stub now derives `source` from the request the way the route does,
rather than answering `full=1` with `mux-visible`, which the route cannot
produce. Each case reaches a visible-frame response the way production does, by
not being the first select of the page. Three cases pin the new behaviour and
each fails without its guard: the clamp case sees two fetches instead of one,
the scope case and the full-history case both see a replay the gate forbids, and
the width case sees one fetch instead of two.

The changeset now describes the change from 1.29.x rather than the difference
between the two commits on this branch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:44:35 +02:00
Michael GrundbergandClaude Opus 5 5cfb98fb8b fix(terminal): compare capture geometry only on a visible-frame response
Only a visible-frame capture positions its rows absolutely, so only that frame
can be damaged by a terminal of the wrong size. A `full=1` body is linear
scrollback closed by a relative cursor move, which is relative precisely so the
browser's row count need not match the pane's, and a `history` body is the byte
stream, which carries no row alignment to protect. The geometry comparison ran
on all three, so it fired most often on the one response it cannot help:
`_fullHistoryLoaded` is empty on the first select of every non-shell session per
page, and a session whose pane a desktop tab holds too tall to ever fit then
paid a second whole-scrollback capture, reset and replay on every page load and
every first tab switch.

`framePositionsRowsAbsolutely` gates both the captured-geometry comparison and
`sizeMovedUnderLoad`. A size that moved under a byte-stream or scrollback replay
is healed by xterm's own reflow plus the SIGWINCH the trailing `sendResize`
already sends.

A pane WIDER than the terminal damages the same frame a second way, so
`captureCols` is now compared rather than only logged. `formatPaneSnapshot`
paints each row out to the pane's own width, so a narrower browser wraps every
painted row, and the wrap on the last one scrolls the whole frame up by a row.

The terminal response no longer falls back to `session.ptyCols`/`ptyRows` when
the capture reported no geometry. The cursor query is what produces the absolute
addressing in the first place, so a capture that lost it returned a raw frame
that was never positioned, and a byte-history response was never positioned
either. Naming the session's own PTY size there described a frame that does not
exist and invited a repair for damage that is not present. `_ptyCols` is also
written only by `resize()` while the PTY is spawned at the size queried from
tmux, so it can be wrong on its own terms. Both fields are now absent instead,
and the `Session` getters added for that fallback go with it.

Two browser cases cover the new behaviour and each fails without its fix: a
`mux-full-history` response with both dimensions mismatched asserts one fetch
(two without the gate), and a `mux-visible` response wider than the terminal
but short enough to fit asserts two (one without the width comparison).

Corrects a claim in the comment above `capturedGeometry` in tmux-manager.ts.
Both replay paths do not address rows absolutely; the full-history one ends in a
relative move, which is the whole reason the gate is right.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:44:35 +02:00
Michael GrundbergandClaude Opus 5 3edf9aae2f fix(terminal): replay a pane capture at the geometry it was taken at
A visible-frame capture repaints each row at an absolute position, counting up
to the pane's height. A terminal shorter than that clamps every address past its
own height onto its last line. The overflow rows then overwrite one another, and
the rows underneath are lost. Replaying a real 50-row capture into a 30-row
terminal rendered 28 lines of a 45-line command and drew the frame twice.

Nothing in the response said what height the frame was built for, so the client
could not detect this. A capture now reports the geometry it was really taken at
through `capturedGeometry` on `PaneCaptureOptions`, and the terminal response
carries it as `captureCols` and `captureRows`. When the captured pane is taller
than the terminal, or the size that produced the capture did not survive the
load, `selectSession` replays once at the size that stuck. `resizeRetry` caps
that at one attempt, so two competing fits cannot trade replays forever.

The retry re-arms the full-history flag only when the pass that ran had consumed
it. A tab switch takes the bounded tail, so its retry takes the tail too:
clearing the flag unconditionally would upgrade that switch into a fresh
scrollback capture the user never asked for, which the route's own comments put
at tens of megabytes.

What this repairs is a capture that won a race against the resize meant to
precede it. It does not repair a capture whose pane was too tall because
`Session.resize` declined the resize outright, which it does for a small
viewport while a desktop viewport's size claim is live. The retry re-sends the
same declined resize and captures the same pane, and `resizeRetry` then stops
it. Repairing that means changing who owns the pane size, which is a policy
question this does not touch. The reported geometry still helps there, because
the client can see the mismatch at all rather than being blind to it.

Follows #395, #396 and #397, which fixed the other ways the replayed frame and
the terminal could disagree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 10:44:35 +02:00
Devvyn 55a80eab86 Merge branch 'master' of https://github.com/Ark0N/Codeman into followups 2026-09-19 08:12:17 +08:00
timkjrandClaude Sonnet 5 358aef16e3 fix(run): make the Instance count stepper work for every non-Claude mode
runOpenCode(), runCodex(), runGemini(), runAntigravity(), runPi(), runOmp(),
runGrok(), and runDeepSeek() all ignored the "Instance count" stepper next
to the Run button and hardcoded a single quick-start call — bumping the
counter to 2 or 3 while on any of these modes silently launched exactly one
session, with no error. Only runClaude() ever read it.

Extract the shared launch-N-sessions-and-select-the-first loop into
_launchQuickStartInstances(), reused by all eight modes, and _readTabCount()
for the shared clamp-and-parse. Each mode still builds its own quick-start
body (config differs per CLI), just via a closure passed to the shared
loop instead of a single inline fetch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-18 18:16:46 -05:00
RandalixandClaude Opus 5 1040f6c489 fix(remote): a proxied host is reachability-unknown; scope remote: SSE per session
Review round 2 on #439.

1. The bare TCP probe connects to host:port, which a host behind a jump host
   or SOCKS proxy does not answer even while ssh works. Acting on that
   verdict drew a permanent banner over a healthy session, replaced a real
   "needs tmux" error with "not reachable" in quick-start, and - with a wake
   target - buffered every HTTP input for the life of the session, since the
   readiness poll could never succeed. `WakeableRemote` now carries
   `jumpHost`/`socksProxy`/`extraSshOptions`, and `isProbeable()` turns such
   a host into reachability-UNKNOWN: input is delivered, `checkReachable` /
   `checkHostReachable` answer `null` (never `false`), `ensureHostAwake`
   returns `'unprobeable'` (handled like `'no-target'`), the quick-start gate
   fires on `=== false` only, and `GET …/reachability` reports
   `reachable: null, probeable: false` so the banner has nothing to key on.
   A wake target can still be fired for it, blind: no readiness poll, no
   reattach, no toast - the response says only whether the packet went out.

2. `'remote:'` joins the session-scoped SSE prefixes. The create/attach wake
   has no session yet, so the registry names the requesting user
   (`ensureHostAwake({ requestedBy })` -> `username` in the payload) and
   `deriveSseHint` routes on it; with neither it fails closed to admins.
   Single-user mode is unaffected.

Smaller, from the same review:

- A flush write that fails now drops the remaining buffer (logged) instead
  of retaining it: the wake still resolved and marked the host reachable, so
  the retained chunk waited for the NEXT wake and was replayed hours later,
  after everything typed since. Same policy as the oversized paste.
- The banner polls on tab activation (a user action) and on its 30 s timer
  only for a host with a wake target; a timer connecting to a host Codeman
  cannot wake is the traffic invariant #2 rejects keepalives for. A proxied
  host is never polled.
- `probeRemoteHostReachable`, `runRemoteWakeCommand` and the default UDP
  socket refuse under VITEST, as remote-files.ts does. The guard caught a
  leak on the spot: `createDefaultRemoteWakeDeps({ probe })` overrode the
  probe but still polled readiness with the real one, so the shutdown test
  had been connecting to a production address. The poll now uses the
  injected probe.
- docs/remote-sessions.md is additions only again (the reformatting is
  gone); the architecture-invariants overlap resolved itself in the merge.

Live, against a throwaway instance with a non-routable ghost host: proxied
-> no probe, no wake, the genuine ssh error after 10 s; direct (control) ->
probe, magic packet, "did not come back" after the 40 s budget.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-18 22:46:11 +02:00
RandalixandClaude Opus 5 e271a65e79 Merge origin/master into feat/remote-host-wake
Resolves CLAUDE.md count tables (route counts recounted on the merged
tree: 235 handlers, sessions 37) and keeps both the host-wake and the
reboot-restore banner in index.html.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
2026-09-18 22:20:41 +02:00
Devvyn 56209e7829 Merge remote-tracking branch 'upstream/master' into feature/run-menu-custom-model-picker 2026-09-19 03:25:26 +08:00
DevvynandClaude Sonnet 5 afb6754453 fix(custom-model): address third pre-merge review (Ark0N)
Blocker 1: the loading banner hides itself ~200ms after it reopens.

- _showCenterStatus reuses one shared DOM node; dismiss() scheduled
  el.hidden = true 200ms later with nothing to cancel it. On the
  Claude path, switchingToast.dismiss() is followed by one same-
  origin request (5-30ms locally) before _watchLlamaSwapLoading opens
  the new banner -- well inside that window -- so the stale timer
  fired against the shared node and hid the fresh banner, leaving the
  whole model-load wait with no progress text, no log line and no
  reachable Cancel button.
- Fixed by parking the pending timeout on the element and clearing it
  at the top of _showCenterStatus. Added a regression test that
  reproduces the exact repro (open, dismiss, reopen 20ms later,
  advance past 200ms) alongside the existing Cancel-button DOM tests;
  confirmed it fails without the fix and passes with it.

Blocker 2: the swap-conflict warning named other users' sessions.

- Both affectedSessions scans (POST .../custom-model and quick-start)
  walked the whole session map with no ownership filter, so in multi-
  user mode a non-admin pointing their own session at a shared
  endpoint learned another user's session name and id -- which with
  autoNameSessions on is that user's own prompt.
- The swap is still blocked pending confirmation regardless of
  ownership (a foreign session is just as real a disruption); only
  which ones get NAMED back to the caller is scoped, via the
  already-imported canAccessOwned. Added a two-owner test to
  test/routes/session-custom-model.test.ts covering both the
  foreign-owner (blocked, not named) and same-owner (named) cases.

Smaller ride-along fixes:

- server.ts boot recovery now passes contextLength into
  applyCustomModelInjection, so CLAUDE_CODE_MAX_CONTEXT_TOKENS is
  correctly rebuilt into _envOverrides after a restart instead of
  surviving only because tmux retains the old setenv.
- pumpLlamaSwapLogTail's finally now deletes by IDENTITY, not just by
  key, so an aborted pump finishing after a newer entry was created
  for the same endpoint can no longer delete that newer entry and
  orphan its connection.
- docs/custom-model-endpoints.md now notes that clearing a custom
  model removes injected keys by name, including CLAUDE_CONFIG_DIR --
  so a session that also had CLAUDE_CONFIG_DIR set via envOverrides
  (the per-client-account case) silently falls back to the default
  account on clear.

Left for later, as flagged in the review itself: the quick-start
case-scaffolding/cancel ordering (real behavioural reordering across
a large handler, too risky to make without a live re-test), and
retiring runCustomModelEntry's mode === 'claude' branch behind a
launchStrategy registry field (explicitly deferred by the reviewer to
"the next one").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-19 03:09:36 +08:00
Michael GrundbergandClaude Opus 5 f9edb33d15 fix(terminal): trim the padding and shared indent out of a copied selection
xterm hands back whole screen rows and trims only the cells that were
never written to, so the real spaces a full-screen TUI paints across the
unused part of a row count as content and reach the clipboard. Measured
against Claude Code in a 282-column pane, single lines arrived carrying
138 trailing spaces, and every line carried the two-space transcript
indent as well. Windows Terminal, iTerm2 and GNOME Terminal all trim that
for you, decideAutoCopy already calls a wall of spaces "never what the
gesture meant", and _selectTouchSelectionLine already treats those cells
as padding — the mouse and keyboard paths never had the same rule.

CodemanCopySelection.clean lives in constants.js beside decideAutoCopy,
its pure sibling. It drops the trailing run from each line, and removes
the leading run only where every selected row shares one. A selection of
a single row keeps its run, because one row shares nothing with anything
and stripping it would silently reindent one line of `git log` body text
or one line out of `less`. A drag that began inside a row keeps its
partial first line untouched and out of the measurement, which otherwise
pins the shared run to zero and leaves every following row indented.

Every pass over a line is a scan rather than a regex. `/[ \t]+(\r?)$/` is
quadratic on a line whose spaces are followed by a non-space character,
which is what right-aligned or centred TUI content looks like: measured
over 50 000 rows with a 280-column run it took 2.9s, against 1.3ms for
the scan, and a 2 000-column run took 16s. The scan is also the faster of
the two on an ordinary padded row.

cleanedTerminalSelection in terminal-ui.js is the half that needs the
live terminal. It returns a COLUMN selection untouched: Alt+drag makes
one, and a rectangle's rows lining up is the point of the gesture, so
both halves of the clean would destroy it. xterm exposes the mode nowhere
public, so the check reads terminal._core._selectionService, the way this
file already reads terminal._core for cell dimensions, and cleans
normally if a future xterm renames the field. A test pins that assumption
against the library rather than against a stub repeating the literal.

The Ctrl+C chord decides on the cleaned selection, not the raw one. A
drag across the blank part of a row selects real padding spaces, so the
raw text is truthy, and testing it would spend that press on a copy of
nothing and make the user press again to interrupt. A padding-only
selection is now dropped and the press falls through to the PTY, while
Ctrl+Shift+C still never falls through. copyTerminalSelection gates on
trim() for the same reason, since a multi-row drag across padding cleans
to line breaks alone and a bare newline pasted into a chat composer
submits it.

All four of the main terminal's copy paths go through it: the Ctrl+C
chord, right-click, the phone selection button and Auto Copy. The
browser's own Edit menu copy, a disabled copy shortcut and the subagent
windows still copy raw rows, as they did before, and the invariants doc
now says so rather than claiming every copy is cleaned. Auto Copy
resolves its own toggle before it reads the selection, since it is off by
default and a selection can run to the 50 000-row scrollback ceiling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 20:18:25 +02:00
DevvynandClaude Sonnet 5 9982a1325f fix(custom-model): address second pre-merge review (Ark0N)
Blocker: .center-status-banner never actually disappears.

- Add `.center-status-banner[hidden] { display: none; }`, same trap as
  `.home-sessions[hidden]`: the author-level `display: flex` beat the
  UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
  card stayed laid out at `opacity: 0` with its text/cancel/close
  children still `pointer-events: auto` -- an invisible 442x67 click
  blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
  banner (10001) and the swap-confirm/context-warning modals (10010)
  in CLAUDE.md's Z-index layers list.

Stale wording pointed at the reverted sticky-toast default:

- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
  `.toast-message` comment in styles.css all still said "toasts
  default to sticky" after 1f32128c put the flat 3s default back.
  Reworded all three to describe the actual behaviour: one call site
  passes an explicit `duration: 0`.

Smaller items from the same review:

- docs/api-reference.md said discovery failures answer
  `502 OPERATION_FAILED`; OPERATION_FAILED is 422 per src/types/api.ts
  and the error-code table earlier in the same file.
- The periodic re-discovery sweep (server.ts) never read
  customModelEndpointsEnabled, so turning the feature off left
  Codeman polling every saved endpoint forever. Added
  readCustomModelEndpointsEnabled() (custom-model-routes.ts, same
  shape as readPlanUsageTelemetryEnabled) and gated the interval
  callback on it.
- Reverted the formatting-only Prettier pass docs/api-reference.md
  picked up (table padding, *x* to _x_, JSON re-indent) by re-merging
  the new Custom Model Endpoints section onto the pre-PR file, so the
  diff is reviewable. No prose content was lost -- verified by diffing
  the result against the pre-revert file (formatting-only) and against
  the merge-base file (only the new section added).
- docs/custom-model-endpoints.md now states that a custom-model Claude
  session's isolated CLAUDE_CONFIG_DIR loses the user's global
  settings.json, user-level skills/agents/commands, and MCP servers
  from ~/.claude.json -- only `projects` is symlinked back.

Design question left open in the review (does `confirmed: true` need
to be two flags so "launch anyway" on the context warning doesn't also
skip the llama-swap displacement warning): keeping the single flag, as
offered. The 20s displacement sweep still catches a resulting swap
after the fact, so it's a surprise rather than a silent failure, and
splitting it is real behavioural surface I have no way to verify live
in this environment.

`npm run test:browser` could not be run in this environment (no tmux,
no downloaded Playwright browser binary) -- none of its suite's files
touch code this fix changes, but it still needs a real pass before
merge, same as any frontend change.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-18 21:45:50 +08:00
DevvynandClaude Sonnet 5 1f32128ca9 fix(custom-model): address PR #430 pre-merge review (Ark0N)
Four blockers from the 2026-09-18 review:

- PUT /api/model-endpoints/:id now merges modelContextLengths/
  modelSizesGB back in from the stored record instead of trusting the
  editor's body, so renaming an endpoint or changing its default model
  no longer silently drops the context-window floor check and
  CLAUDE_CODE_MAX_CONTEXT_TOKENS injection.
- custom-model:swapped-out is now session-scoped (added to
  SESSION_PREFIXES) instead of broadcasting to every connected client.
- The quick-start custom-model path now hands setCustomModel() only
  the endpoint's own injected env vars, not the full merged set,
  matching the restart-in-place path — the full set put
  CLAUDE_CODE_EFFORT_LEVEL back after the Session constructor had
  already stripped it.
- The quick-start launchModel override for pi/grok/omp is now applied
  generically via the registry's legacyConfigField, mirroring
  Session._withCustomModelLaunchModel, instead of three hardcoded
  mode === '<id>' branches a future CLI's injection recipe would miss.

Also scopes the sticky-toast default (item 5): reverted the blanket
"all error toasts are sticky" default, which had no container cap or
eviction, back to a flat 3s; the one message that needs a moment to
read (a failed custom-model apply) now passes an explicit
duration: 0 at its own call site.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
2026-09-18 20:16:07 +08:00
DevvynandClaude Sonnet 5 e2034177c5 fix(custom-model): root-cause and fix DeepSeek's HTTP_404 (missing /v1)
DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.

Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
  published dependencies) into a scratch dir purely to read
  dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
  completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
  the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
  completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
  same endpoint. dsh's own error template ("DeepSeek API error (HTTP
  ${status})") reproduces the originally-reported
  "dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.

- New registry field `appendV1Suffix` (env kind only, deepseek's entry
  alone — claude/gemini must NOT get it, since claude was already
  confirmed working against the unmodified baseUrl). When set,
  buildCustomModelInjection runs endpoint.baseUrl through the same
  withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
  use, instead of writing it verbatim.

Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.

2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:41:19 +08:00
DevvynandClaude Sonnet 5 8520925e76 docs(custom-model): bring CLAUDE.md and api-reference.md up to date
Full documentation review pass across the branch's 30 commits.
CLAUDE.md's Custom Model Endpoint Profiles entry hadn't been touched
since the initial backend+picker cut (3 early commits) despite 27
follow-up commits adding real behavior — it described restart-in-place
as universal (now claude-only; 7 other CLIs launch one-shot) and
claimed codex's Responses-API gap as a flat protocol break (now
re-verified as a more precise tool-calling gap). Corrected both and
added a new paragraph covering everything landed since: the llama-swap
conflict check, the after-the-fact swap-displacement sweep, the
/running-cmd-based context-length fix, the context-window floor
warning, skipFirstRunPrompts, the real-time /api/events-based log
status, and the countdown-to-Cancel-button change.

docs/api-reference.md's custom-model-endpoints section was missing the
running-status route, the requiresConfirmation/requiresContextWarning
response shapes, and POST /api/quick-start's customModel field
entirely (the primary launch path for 7 of 8 supported CLIs) — added
all three. Also fixed a real markdown bug in custom-model-endpoints.md:
an inline code span (`POST <baseUrl>/v1/chat/completions`) split across
a line break, which CommonMark renders with the line ending collapsed
to a space, so it displayed as ".../v1/chat/ completions" with a
spurious space inside the path.

Verified: origin/master and upstream/master are both already an
ancestor of this branch (identical at bd286bf5, no new commits since
this branch was cut) — nothing to merge, no conflicts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:18:04 +08:00
DevvynandClaude Sonnet 5 db9729e1fc feat(custom-model): remove loading-banner countdown, add manual Cancel
Replaces the size-scaled expected-time estimate + matching auto-timeout
with a generic hardware/model-size disclaimer and a user-driven Cancel
button, per explicit request. Real load time depends on hardware this
feature has no way to know (VRAM, storage speed, GPU contention), so
the old estimate/timeout was a guess dressed up as a fact — worse, one
that could kill a genuinely slow load partway through on slower
hardware.

- _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline
  entirely — polls indefinitely until ready or cancelled, no automatic
  give-up. Message is now "Loading <model> (<size>) on <endpoint> —
  this can take a while depending on your hardware and the model
  size.", with the real llama.cpp log line still on its own second
  line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/
  _formatRemaining (dead code once the countdown is gone) —
  _lookupModelSizeGB is kept, the GB figure still shows.
- _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real
  "Cancel" button (distinct from the error-type "×" close button,
  since Cancel has a real consequence) that calls it on click. Caller
  owns what cancelling actually means, same split as the swap-confirm
  modal's promise-resolving buttons.
- Cancelling dismisses the banner, shows an info toast (not an error —
  this was deliberate), and closes the session, mirroring what the old
  timeout used to do automatically but now on the user's own call.
- New .center-status-cancel CSS (bordered pill button, distinct from
  the plain "×" close glyph).

Test changes: removed the now-invalid timeout-auto-close/estimate
tests, added cancel-flow tests (dismiss/toast-type/session-close,
never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests
for the new Cancel button (bootAppWithRealCenterStatus, evaluating
panels-ui.js instead of stubbing _showCenterStatus, since this button
is worth verifying for real rather than just through the stub every
other test in the file uses). Typecheck/lint/frontend-syntax clean;
full suite shows no new regressions.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 13:00:50 +08:00
DevvynandClaude Sonnet 5 2d3fc65758 feat(custom-model): show real-time llama.cpp backend status in the loading banner
Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.

- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
  (custom-model-routes.ts): one persistent GET /api/events (SSE)
  connection held open per endpoint, parsing logData frames and
  keeping the latest source:"upstream" (backend llama-server) line —
  filtering out llama-swap's own source:"proxy" request-access lines.
  Idle-closed after 30s of no polling, same 20s sweep as the existing
  swap-displacement check.
- running-status route now returns logLine alongside the existing
  isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
  ("llama.cpp: <line>", bootlog timestamp/level/component prefix
  stripped for display) that stays on the last real thing llama.cpp
  said rather than clearing to blank between polls.

⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).

12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 12:22:45 +08:00
DevvynandClaude Sonnet 5 5ddc028a2f feat(custom-model): detect and notify when a session's model gets swapped out later
The llama-swap conflict check on the apply/create routes only ever runs
at THAT session's own launch/apply moment, and cannot see a swap caused
by a DIFFERENT session's later, ordinary use. Confirmed live: a second
Codex session picking a different model launched with no warning at
all — nothing conflicted at that exact instant — yet it silently
evicted the first session's model regardless (llama.cpp runs one model
at a time). Reproduced and root-caused via direct API calls against a
live test-picker instance rather than guessing.

- detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups
  live sessions with a customModel by endpointId, checks each group's
  endpoint via GET /running once, and flags a session whose own modelId
  is no longer in the running list. Read-only, best-effort per endpoint
  like refreshAllCustomModelHosts's sibling sweep.
- Notifies once per displacement via a caller-owned de-dupe Set: a
  session id is added when displaced, removed once its own model is
  loaded/ready again, so a later genuinely-new displacement can notify
  again.
- New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS,
  20s — much shorter than the 5-minute model-list refresh, since this
  is time-sensitive) broadcasts a new custom-model:swapped-out SSE
  event per displacement. De-dupe Set cleared per-session on session
  cleanup to avoid an unbounded leak.
- Frontend: global toast (not tied to the displaced session's tab,
  since the point is warning before the user types into it) naming the
  session, its previous model, and what's currently loaded.

Chose the "detect after the fact" scope (vs. checking before every
message send, which would add a round-trip to every turn on every
custom-model session) per explicit user decision after being presented
the trade-off.

9 new tests for the detection logic (flag/clear/re-flag cycle,
unreachable/deleted endpoints, non-llama-swap servers, multiple
sessions on one endpoint). SSE registry bumped 158->159, parity test
passing. Typecheck/lint/frontend-syntax clean; full suite shows no new
regressions (9 more passing than baseline, matching the new tests;
same pre-existing Windows-environment failures).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 11:09:24 +08:00
DevvynandClaude Sonnet 5 470f75b08c docs(custom-model): record live findings on codex's model-metadata warning
Investigated the user's report of "Model metadata for <id> not found.
Defaulting to fallback metadata..." on every custom-endpoint codex
launch, live against the test-picker's llama-swap deployment (codex
0.152.1):

- The warning is cosmetic. `codex exec 'reply with just OK'` against the
  isolated CODEX_HOME still printed the warning and still returned a
  real reply.
- The isolated CODEX_HOME never gets a models_cache.json written into
  it at all, even after extended real use (inspected a live, actively-
  used directory) — codex can't reach OpenAI's own hosted model catalog
  for this session and silently falls back every time, with no local
  file to create or clean up. There is also no config.toml override for
  a model's metadata.
- Fabricating a fake catalog entry to suppress it would mean copying the
  SHAPE of OpenAI's own proprietary models_cache.json schema, including
  real per-model system-prompt content visible in a genuine entry — not
  something to build for a warning confirmed to have no effect.
- More importantly: a real tool-call attempt against the same setup came
  back as agent_message TEXT (the tool-call JSON printed as the answer)
  rather than an executable function_call item, confirmed via
  `codex exec --json`'s raw event stream. Tool execution is what makes
  codex a coding agent, so it remains not usable for real work regardless
  of the metadata warning — a more precise, re-verified update to the
  existing "Responses API protocol gap" finding (which reported a harder
  Reconnecting/high-demand failure on a different llama-swap deployment;
  this one answers /v1/responses for plain chat but still can't execute
  tools).

No code changes — recipe/comment/confidence-table documentation only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 10:33:48 +08:00
DevvynandClaude Sonnet 5 211b872335 feat(custom-model): skip Claude Code's first-run wizard on custom-model launches
A fresh, isolated CLAUDE_CONFIG_DIR (used to keep an injected API key
away from a stored claude.ai OAuth login) looks like a brand-new Claude
Code profile to the CLI, so it replays its ENTIRE first-run sequence on
every single launch: the theme picker, the security-notes screen, the
per-project "trust this folder?" dialog, and (running with
--dangerously-skip-permissions) a one-time bypass-permissions warning —
confirmed live, none of which a real, already-onboarded profile shows
again.

- New registry-declared env-kind field `skipFirstRunPrompts` (alongside
  apiKeyTrustFile, which it reuses) — claude's entry only, carried
  through buildCustomModelInjection (pure) into
  applyCustomModelInjection (IO).
- seedFirstRunOnboardingState(): merges hasCompletedOnboarding: true and
  this session's own projects[workingDir].hasTrustDialogAccepted: true
  into the same <configDir>/.claude.json the API-key trust file already
  writes to — other projects and other fields on this session's own
  entry are left untouched.
- seedSkipBypassPermissionsPrompt(): merges
  skipDangerousModePermissionPrompt: true into <configDir>/settings.json,
  a separate file, same corrupt-tolerant merge behavior.
- applyCustomModelInjection() gains an optional workingDir parameter,
  threaded from session.workingDir (dedicated apply route) /
  resolvedCasePath (quick-start route) — boot recovery omits it
  (a dialog already answered once needs no re-seed on the same,
  persisted isolated directory).

Tests added at the pure-builder, IO-wrapper (including merge-preserves-
other-fields and corrupt-file-tolerance cases), and existing directory-
listing assertions updated for the new settings.json file. Typecheck/
lint/format clean; full suite shows no new regressions (baseline
pre-existing Windows-environment failures unchanged, 8 more passing
tests than before — the ones added here).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 09:21:36 +08:00
DevvynandClaude Sonnet 5 2c89359d42 fix(custom-model): Cancel/Launch-anyway buttons stacked instead of side by side
Neither dialog's footer had a row layout of its own to override, and
.btn-toolbar is display:flex (a block-level flex container with no
explicit inline-flex), so with no flex row context each button took its
own full-width line and the two stacked vertically. The swap-confirm
modal already had a .modal-footer rule (flex-end); the context-warning
modal had none at all. Both now share one row-layout rule, centred
rather than flex-end per feedback.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 09:01:28 +08:00
DevvynandClaude Sonnet 5 962029bb3d fix(custom-model): context-warning/swap-confirm modals hidden behind status banner
Both dialogs can appear while the centred llama-swap status banner is
still on screen (right after "Claude started — switching to
llama-swap…") — the banner's z-index is 10001, .modal's base z-index is
only 1000, so the dialog rendered fully behind it. Reported live against
the context-window-too-small modal; the swap-confirm modal has the same
structural bug for the same reason, so both get the fix.

Also: both messages ARE the modal's whole explanatory content, not a
one-line caption under a form field, so .form-hint's 0.65rem caption
size read as illegibly small — worst on the multi-sentence
context-window explanation. Bumped to 0.85rem/1.5 line-height/--text.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 08:57:12 +08:00
DevvynandClaude Sonnet 5 b45a96358e feat(custom-model): warn before launching Claude on a model too small for its own overhead
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.

- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
  registry entry declares contextLengthVar (currently only claude) and
  the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
  (40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
  quick-start customModel path) check this before the swap-conflict
  check and before launching/restarting anything, returning
  {requiresContextWarning, modelId, contextLength, minSafeContextTokens}
  — skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
  _resolveContextWarningConfirm (session-ui.js), wired into both
  _quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
  (the path Claude actually uses) ahead of the swap-confirmation check.
  Explains the fix in-modal: give the model an explicit larger -c/
  --ctx-size in llama-swap instead of relying on --fit-ctx, which
  optimizes for the biggest model that fits rather than the biggest
  context.

Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-17 07:36:57 +08:00
Randalix 29984c639d fix(remote): stop the flush losing a chunk, and reset the host form's wake fields
Own review pass over the PR:

- `_flush` took the chunk out of the buffer only AFTER awaiting the write. Input
  arriving during that await is enqueued (`waking` is still set, so it takes the
  buffer path), and the 4 KB cap then drops the OLDEST chunk — which is the one
  already on its way to the pane. The `shift()` that followed removed the NEXT
  chunk instead, so the drop-oldest bookkeeping silently lost a chunk that was
  never written, while the log line blamed the one that was. The chunk is now
  removed before the await and re-inserted at the FRONT on a failed write, so the
  order of the queue behind it is preserved. Regression test: a chunk enqueued
  during the first write of a full buffer must still reach the pane (red against
  the old order).
- `showCreateCaseModal()` reset the remote-host form fields but not the two new
  wake inputs, so one host's MAC/command carried over into the next host that
  form saved.
- The banner's pre-poll `wakeConfigured` labelled a command-only host as 'mac'.
  Nothing reads the distinction, but the field is documented as which path is
  configured, so it says the truth until the first poll corrects it.
- Stale `resolveRemote` comment ("only for sessions that have no usable target of
  their own"): after the host config became authoritative in both directions it is
  consulted on the TTL regardless.
2026-09-16 21:06:25 +02:00
Randalix acb8d4b0aa docs(remote): correct what the wake PR moved
- `host-wake-ui.js` joins the documented load order (12.2) and gets its
  `@dependency`/`@loadorder` tags; the frontend module count is 33, not 32.
- `remote-wake` is not "(pure)" — the module uses `dgram`/`net`/`child_process`.
- SSE counts: 160 constants, and the category is "Remote auto-reconnect / wake
  (5)"; the route table's per-file counts are refreshed (sessions 37, cases 34).
- The CLAUDE.md wake rule now names the create/attach wake, the 40 s request
  budget, the whole-chunk paste drop, the registry's lifetime (drop on cleanup,
  stop on shutdown) and the deliberately non-wake-aware WebSocket keystroke
  path — that paragraph is what the next person reads.
- Reverted the eight lines of unrelated Prettier markdown churn in
  `docs/architecture-invariants.md` (docs/ is not in the format glob, so it was
  an editor): only the new wake paragraph remains in the diff.
2026-09-16 20:44:48 +02:00
Randalix 7b947fa3f1 fix(remote): close the wake-state leaks and the dishonest wake budget
Review follow-up on the wake-on-LAN PR (five findings, all of them about the
state the feature keeps and the budgets it inherits):

- Wake state is dropped by `WebServer.cleanupSession` instead of the two delete
  routes, so it now goes with the session on EVERY cleanup path (cron, admin,
  scheduled-run teardown, error paths) instead of surviving with up to 4 KB of
  the user's buffered keystrokes. `registerSessionRoutes` returns the registry
  so the server can own its lifetime without the wake-capable code living in
  `server.ts`; the wiring guard is updated to allow that and gains a second
  assertion that `server.ts` calls nothing but `drop`/`stop` on it.
- `_effectiveRemote` returns before `_state`, so a LOCAL session no longer gets
  a wake-state entry — the input gate runs on every keystroke, so that entry
  used to be allocated for every session the user types in.
- An input chunk larger than the 4 KB cap is dropped OUTRIGHT instead of being
  head-trimmed and then written as a fragment: one paste is one `input` value
  and was never typed character by character, so its tail is a partial command
  the user never sent. The drop is logged.
- The manual wake button passes `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS` (40 s)
  like the create/attach paths, instead of inheriting the 90 s session default
  that the dashboard's reverse proxy cuts off at 60 s.
- `RemoteWakeRegistry.stop()` aborts in-flight readiness polls (abortable
  sleep) and refuses new wakes, and `WebServer.stop()` calls it, so a restart
  during a wake no longer waits the poll out.
- The banner/toast wording keys off a new `queuedInput` flag on the two SSE
  events, which is true only when the server actually holds bytes: browser
  keystrokes travel over the WebSocket, which never passes through the
  registry, so the wake BUTTON must not promise queued input. The failed-wake
  path also stops pattern-matching the error message (it re-asks the
  reachability route) and the WoL dialog says "admin-only" instead of "host not
  found" for a non-admin in multi-user mode.
2026-09-16 20:44:39 +02:00
DevvynandClaude Sonnet 5 993710263d fix(custom-model): stop trusting /props's n_ctx, parse the real context size from /running's cmd
Root cause of the context-overflow regression reported live: "API Error: 400
request (36437 tokens) exceeds the available context size (16384 tokens)".
Discovery had stored modelContextLengths.qwen3.8-27b-ud-q4_k_xl = 154112,
so CLAUDE_CODE_MAX_CONTEXT_TOKENS told Claude Code it had a huge window and
it never compacted - but the real llama-swap server was launched with
--fit-ctx 16384 (confirmed against /running's own cmd field) and refused
the request right at that real limit.

/props?model=<id>'s n_ctx (the field discovery read) is confirmed live to
be unreliable for a --fit-ctx-launched backend: it reported 154112 for the
same model /running says was launched with --fit-ctx 16384 - appears to
report the model's theoretical/trained maximum context, not the runtime-
configured one.

discoverModels() now parses the REAL configured size straight out of
llama-swap's own launch command instead (parseCtxFromCmd(), reading
/running's cmd field - --fit-ctx first, then the plain llama.cpp -c/
--ctx-size a hand-written command might use), and only falls back to the
old /props probe when cmd states no recognizable flag at all. One /running
call now covers every loaded model's context length in a single request,
same as it already did for the swap-conflict check and the load trigger.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 20:47:12 +08:00
DevvynandClaude Sonnet 5 7bbe408e44 feat(custom-model): live countdown on the loading banner; timeout is now an error
The loading banner now shows a live countdown against its own timeout
(updated every poll, so every second by default) instead of a static
"this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically
~1-3 min) on llama-swap - 47s remaining".

If the countdown reaches zero and the model still isn't ready, this is now
treated as a real failure rather than a "keep waiting" shrug:
- The banner turns into a sticky error (_showCenterStatus gains a `type`
  option - 'error' drops the spinner and adds a close button, since nothing
  is "in progress" anymore and a sticky message needs a way to dismiss it),
  naming the llama-swap server's own logs as where to look for detail.
- The session that load was for is closed automatically (closeSession) -
  requested explicitly: a console left open and pointed at a model that
  never finished loading is worse than no console at all. Both apply paths
  now thread the new session's id through to _watchLlamaSwapLoading for
  this (new required 3rd parameter, after endpointId/modelId).

_watchLlamaSwapGeneration's existing stale-call guard extends naturally to
this: a superseded call's own eventual timeout recognises it no longer owns
the banner and neither shows the error nor closes a session that may by
then belong to a different, newer launch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 20:22:02 +08:00
DevvynandClaude Sonnet 5 55dae31530 feat(custom-model): estimate model load time from its discovered size
Discovery now also parses a GB figure out of an auto-discovered model's own
description (llama-swap writes "Auto-discovered 16.35 GB - parameters
auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike
context length this needs no /props probe (the figure is right there in
/v1/models) so it is populated for every model regardless of loaded state.
A hand-configured profile's own description has no such figure and
correctly gets no entry.

The loading banner (_watchLlamaSwapLoading) now looks this up and, when
known, shows it plus a rough estimate from a small size->time matrix
(_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) -
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on
llama-swap... this can take a while" - and uses that same estimate's own
bracket to scale the banner's default give-up timeout for a very large
model, instead of a flat 5 minutes for everything. Explicitly labelled as
an UNMEASURED, typical-hardware estimate in every relevant comment - this
is not benchmarked against any real endpoint's actual storage/GPU, just a
reasonable expectation-setter. A model with no discoverable size (a
hand-configured profile) gets no size/estimate shown at all, matching the
"never a guess" convention modelContextLengths already established.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 19:58:46 +08:00
DevvynandClaude Sonnet 5 0af233c96c fix(custom-model): poll llama-swap readiness every 1s, check immediately, extend the cap
Reported: the "Loading..." banner stayed up past 2 minutes even though
llama-swap itself had already finished loading the model. Three fixes:

1. pollIntervalMs default 3000ms -> 1000ms (as asked).
2. The loop now checks readiness IMMEDIATELY on entry rather than sleeping
   a full interval first - a model that's already ready (a fast load, or a
   re-apply onto one already loaded) shouldn't sit on "Loading..." at all.
3. maxWaitMs default 120000ms (2 min) -> 300000ms (5 min): a large (20GB+)
   model reading from disk can genuinely take longer than 2 minutes, which
   would have looked identical to the reported symptom - "still stuck past
   the point it should have resolved" - except it would have actually
   flipped to a "still waiting" warning toast at the 2-minute mark rather
   than staying on "Loading" indefinitely, so this alone doesn't explain
   what was reported, but is a real, separate improvement worth making.

Also fixes a real, separate bug this surfaced while reasoning through the
report: _showCenterStatus's banner is ONE shared, reused DOM node. A second
call to _watchLlamaSwapLoading (e.g. switching models again before the
first switch's loop had finished) would take over that shared banner, but
the FIRST loop was still running and would eventually dismiss or overwrite
it once ITS OWN deadline or readiness check resolved - clobbering whatever
the second, current loop had put there. A generation counter
(_watchLlamaSwapGeneration) now lets each call recognise when it no longer
owns the banner and stop touching it silently, rather than only the last
call to actually start ever safely reading or writing it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 19:44:37 +08:00
Randalix a7f74f374f fix(remote): keep the wake banner hidden after switching to a local session
refreshHostWakeBanner clears _hostWake before calling _hostWakeTick, so the
clear branch's `if (this._hostWake)` guard skipped the repaint: once the
banner had appeared for an unreachable remote session it stayed up on every
chat (local ones included) until a reload, and the 30s ticker never cleared
it either. Render unconditionally in that branch — _renderHostWakeBanner is
idempotent with a null state.

Reproduced in a real browser (Puppeteer, mobile viewport): state went null
but banner.hidden stayed false. Regression test added in
test/host-wake-banner.test.ts (red before, green after).
2026-09-16 10:39:48 +02:00
DevvynandClaude Sonnet 5 0929694012 fix(custom-model): actually trigger the llama-swap load, not just watch for it
Root cause of "it doesn't look like llama-swap is actually switching the
model" (confirmed live: no load_model line in llama-swap's own logs after
applying a selection). llama-swap has no "switch model" admin endpoint - the
ONLY thing that starts a swap is a real inference request naming the model.
Every previous fix (the conflict check, the loading banner) assumed a swap
would start on its own; nothing ever actually asked llama-swap to load
anything until the launched CLI's first real prompt did, which could be
much later than "applying the selection" implied.

Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest
real request that will start a load - POST <baseUrl>/v1/chat/completions,
max_tokens: 1, one throwaway message - fire-and-forget (never awaited by
the caller; the frontend's own running-status polling is what actually
confirms readiness). Wired into both apply paths (the dedicated restart
route and the one-shot quick-start route), fired whenever the target model
isn't already the one loaded and ready - a broader condition than the
existing swapNeeded (which only gates the "this will evict another
session's model" confirmation ask and deliberately stays narrow to that).
modelSwapInProgress in both routes' responses now reflects this same
broader condition too, so the frontend's loading banner actually correlates
with a real in-flight load rather than only firing when something else
happened to be loaded already.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:53:34 +08:00
DevvynandClaude Sonnet 5 01b32ee6cd fix(custom-model): move the switching/loading status to a centred banner
The "Claude started - switching to <endpoint>..." and "Loading <model> on
<endpoint>... this can take a while" messages lived in the top-right toast
corner along with everything else, easy to miss given they can each sit on
screen for well over a minute (a real llama-swap model load).

Adds _showCenterStatus() (panels-ui.js): a single, reused, screen-centred
banner with a spinner, non-blocking (no backdrop, pointer-events: none on
the wrapper) so it never gets in the way of using the app while it's up.
Both call sites (_runCustomModelEntryViaRestart's switching message,
_watchLlamaSwapLoading's loading message) now use it instead of showToast.
Every OTHER status in these two flows - the llama-swap conflict warning
already moved to its own modal, apply failures, cancellation, and
_watchLlamaSwapLoading's own final "ready"/"still waiting" outcome - stays
exactly where it was, in the corner.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:40:43 +08:00
DevvynandClaude Sonnet 5 2936ba6e3d fix(custom-model): replace the native confirm() popup with an in-app modal
The llama-swap "this will unload it for session X" warning used a native
browser confirm() popup, which looks out of place next to the rest of the
app's own modals.

Adds #customModelSwapConfirmModal (index.html) with Cancel/Switch-anyway
buttons, styled to match the app. _confirmModelSwap(message) shows it and
returns a promise that resolves true/false the same way confirm() would;
_resolveModelSwapConfirm(proceed) (wired to both buttons and the backdrop
click) settles it. Both llama-swap conflict call sites
(_quickStartWithCustomModelConfirm for the one-shot launch path,
_runCustomModelEntryViaRestart for Claude's restart path) now await this
instead of calling confirm() directly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:29:38 +08:00
Devvyn b1db5515d7 Merge branch 'master' of https://github.com/Ark0N/Codeman into followups 2026-09-16 15:23:52 +08:00
DevvynandClaude Sonnet 5 83033b4299 fix(custom-model): show a status toast during Claude's native-boot-then-restart window
Claude stays on the launch-then-restart path (see runCustomModelEntry's own
comment for why), but with nothing on screen during that window, a native
boot that briefly talks to the cloud model read as "the endpoint didn't
apply" rather than "the switch hasn't happened yet".

A sticky "Claude started - switching to <endpoint>..." toast now covers the
whole window from the native launch through the apply call, updated in
place (never stacked) as the outcome resolves: dismissed on cancel or
failure (replaced by the existing cancellation/error toast), handed off to
_watchLlamaSwapLoading's own sticky toast when a model swap is in progress,
or updated to the existing "Pointed at ... - restarting" message and
auto-dismissed after 3s on a plain success.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:21:09 +08:00
DevvynandClaude Sonnet 5 f865f74a0f feat(custom-model): launch directly on the endpoint, no restart, for 7 of 8 CLIs
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.

POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.

Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.

Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 15:04:10 +08:00
DevvynandClaude Sonnet 5 fbee1b2d82 docs(changeset): add changeset for the Run-menu custom-model picker PR
Covers #430's full scope so far: the picker itself, the model-selection
dialog, periodic re-discovery, and the session-busy/toast/CLAUDE_CONFIG_DIR/
context-length/llama-swap-conflict fixes found through live validation
against a real llama-swap server.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 14:39:56 +08:00
DevvynandClaude Sonnet 5 bcebc81fcd feat(custom-model): detect llama-swap model conflicts before switching
Root-caused the user's earlier confusion ('the terminal says opus even though
something is waiting for llama to load'): llama.cpp runs exactly one model at
a time, and llama-swap unloads/reloads it on demand - a swap can take
anywhere from a few seconds to well over a minute, during which a session
looks indistinguishable from one still on the native backend.

1. Feature-detects llama-swap (vs. plain llama.cpp/any OpenAI-compatible
   server) via its own GET /running, which plain llama.cpp has no concept of
   at all. New GET /api/model-endpoints/:id/running-status route exposes this
   read-only, for the frontend's polling loop below.

2. Before applying a selection, POST /api/sessions/:id/custom-model now checks
   what llama-swap currently has loaded. If it differs from the requested
   model AND another live session's own customModel selection is actively
   using that loaded model, the apply is refused with a
   {requiresConfirmation, currentlyLoadedModel, affectedSessions} payload
   instead of silently switching. A "confirmed: true" field on the retry
   skips the check. Switching with nothing else affected proceeds
   immediately, no confirmation asked, only ever when there is something to
   warn about.

3. The frontend (runCustomModelEntry) shows a native confirm() naming the
   affected session(s) and the model they'd lose, matching this codebase's
   existing convention for this class of decision (delete case, kill
   session, etc.) rather than a new modal. On a successful apply the response
   also carries modelSwapInProgress; when true, a new _watchLlamaSwapLoading
   poll shows a sticky "Loading <model>..." toast via the new running-status
   route until llama-swap reports the target model ready (bounded at 2
   minutes), so a prompt sent mid-swap reads as "loading", never as silence
   or an answer from whatever was loaded a moment before.

Checks are read-only against llama-swap's own /running - never /props, which
takes a ?model= and can itself trigger a load as a side effect of asking.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 14:35:07 +08:00
DevvynandClaude Sonnet 5 25f22b9839 test(custom-model): update session-custom-model route test for CLAUDE_CONFIG_DIR isolation
Fixes the CI failure on the last two commits: this route test asserted an
exact envKeys list for a claude-mode apply that predates the
CLAUDE_CONFIG_DIR isolation fix, so it failed on the new CLAUDE_CONFIG_DIR
entry it correctly started appending. Updates the expected list and adds
assertions for the isolated config dir path and the pre-seeded
.claude.json trust-approval file, matching the behavior added in the two
prior commits rather than just tolerating it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 13:23:54 +08:00
DevvynandClaude Sonnet 5 97464bfa27 fix(custom-model): pre-approve the injected API key in the isolated Claude config dir
The CLAUDE_CONFIG_DIR isolation from the previous commit fixed the cosmetic
auth warning but introduced a real regression: an otherwise-empty config
directory has none of a real profile's prior custom-API-key approvals, so
Claude Code stops at an interactive 'Detected a custom API key - use it?'
prompt on every single launch. Confirmed live. With nobody at a TTY to
answer, the prompt's own default ('No') silently refuses the very key this
feature just injected, which looks like the endpoint being ignored.

Adds apiKeyTrustFile to the env-kind customModelInjection capability shape
({relPath, shape: 'claude-api-key-responses'}), set on claude's entry to
{relPath: '.claude.json', shape: 'claude-api-key-responses'}. The apply step
merges customApiKeyResponses.approved: [apiKey] into
<isolatedConfigDir>/.claude.json - the exact field a real answered prompt
itself writes to (confirmed against a real ~/.claude.json after answering by
hand once), so this answers the prompt in advance rather than bypassing it.
Merges onto whatever the CLI already wrote into that file on an earlier
launch in the same isolated directory rather than overwriting it; a missing
or corrupt file is treated as empty rather than failing the apply.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 13:08:07 +08:00
DevvynandClaude Sonnet 5 0e8b1981af fix(custom-model): isolate Claude config dir and inject real context length
Addresses two live-validation findings on the Run-menu custom-model picker:

1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still
   coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same
   config directory and warns about it (confirmed cosmetic - the API key
   wins for actual requests, verified via a real session's own API Usage
   Billing line). A custom-model claude session now gets an isolated
   CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty,
   no files written into it) so there is nothing to conflict with. projects
   is symlinked (junction on Windows) back into the real config dir so the
   response viewer, subagent windows and Read My Mind keep working for that
   session, best-effort.

2. Context-window overflow. Claude Code assumes a large default context
   window for a model id it doesn't recognise and never compacts, so a
   custom endpoint's real, much smaller context (verified live: a 400
   exceeding a 16384-token llama-swap model with a stock ~33.7K-token system
   prompt) silently overflows. Discovery now also learns each model's real
   context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx),
   but ONLY for a model llama-swap's own /v1/models response already marks
   status.value === 'loaded' - never an unloaded one, since llama-swap
   treats ?model= as a routing hint and probing an unloaded model risks
   triggering an actual, slow, GPU-swapping load as a side effect of
   read-only discovery. A server with no status field at all gets no
   enrichment rather than a guess; a model not probed this round keeps its
   previously-learned value until it disappears from the list entirely.
   Stored per model (CustomModelHost.modelContextLengths) and applied via a
   new contextLengthVar registry field, set to
   CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude.

Both new fields live on the existing env-kind customModelInjection
capability shape, declared only on claude's registry entry - every other
CLI's injection is unaffected (pinned by test).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 12:08:26 +08:00
DevvynandClaude Sonnet 5 5c25a52f95 fix(custom-model): wait for a freshly launched session to go idle before applying
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.

Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.

Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 10:38:37 +08:00
DevvynandClaude Sonnet 5 409a6e65f9 fix(custom-model,toast): surface the real apply error, and make error toasts sticky with a close button
Two related fixes, both needed to actually diagnose 'Session started on
the native backend — could not apply the custom endpoint' reports from
live testing:

1. runCustomModelEntry()'s apply call went through _apiJson(), which
   unwraps a success body but SWALLOWS a failure response entirely and
   returns null — discarding the one thing (error, errorCode) that would
   tell 'endpoint unreachable' apart from 'not a discovered model',
   'remote/Docker session', or a dozen other real causes the apply route
   already reports distinctly. Switched to _api() so the actual response
   body is read on failure too, and the toast now includes the real
   message.
2. showToast() defaulted every toast, error or not, to a 3s auto-dismiss
   with no way to read it again — exactly what made the above generic
   message impossible to act on even before the fix above. Error toasts
   now default to sticky (duration: 0, no auto-dismiss) unless a caller
   opts into a duration, and every toast — sticky or not — gets an
   explicit close (x) button, since a sticky toast with no way to
   dismiss it would just accumulate across repeated failures.

Tests: custom-model-run-menu-ui.test.ts's two apply tests updated for the
_api() switch (their mocks previously stubbed _apiJson, which the apply
call no longer goes through), plus a new test pinning that the real
server error string reaches the toast on a failure.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 09:55:03 +08:00
DevvynandClaude Sonnet 5 9a9e542a7d fix(custom-model): bound the model-picker dialog's height and make its list scroll
The dialog had no max-height at all, so an endpoint with many discovered
models grew it past the viewport with nothing to scroll — reported live as
both "takes up the full page" and "the list is truncated", which turn out
to be the same bug. Gives #customModelPickModal .modal-content the same
bounded-height + scrollable-body shape cronModal's .modal-lg already uses
(max-height + flex column on the content, overflow-y:auto + flex:1 on the
body), scoped by id rather than folded into the shared .modal-sm class
three other modals already use for short, fixed content.

max-height: min(70vh, 520px) scales with the viewport (a phone gets 70% of
its height; a 4K display never gets a needlessly tall dialog) rather than
committing to one fixed pixel value that would be wrong at either end.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 09:23:14 +08:00
DevvynandClaude Sonnet 5 5a9ff07f57 feat(custom-model): ask which model on launch when an endpoint has more than one, and re-discover models every 5 minutes
Two enhancements requested after live-validating PR #430 against a real
llama.cpp server:

1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to
   apply the endpoint's defaultModelId (or the first discovered model)
   silently. Now, via the new selectCustomModelEntry() (session-ui.js):
   - exactly one discovered model launches straight away, same as before
   - two or more open a new #customModelPickModal listing every discovered
     model; defaultModelId (if set) is marked but never auto-chosen, since
     the point of asking is letting ONE launch deliberately differ from
     the saved default, not just confirming it
   The endpoint is re-fetched at click time rather than trusting anything
   cached from the dropdown's own render, since the model list can have
   changed (the sweep below, or a settings-panel edit) since it opened.
   runCustomModelEntry() itself — the actual launch, routed through run()
   for the in-flight lock, snapshot-guarded against applying to the wrong
   session — is unchanged; it now just always receives an explicit model
   id from one of these two paths instead of computing one itself.

2. Periodic re-discovery. Every saved endpoint's models now refresh
   automatically every 5 minutes in the background
   (CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way
   as the Codex plan-usage poll it sits beside — this.cleanup.setInterval,
   off under testMode), so a model the server starts or stops serving shows
   up without another manual "Discover" click. The manual POST
   .../discover-models route and the new refreshAllCustomModelHosts()
   sweep (custom-model-routes.ts) now share one pure merge step
   (applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId
   that no longer appears) rather than two copies that could drift. The
   sweep is best-effort per host — one endpoint being unreachable on a
   cycle never blocks the others — and re-reads the store before each
   host's write, keyed by id, so a concurrent edit or delete from the
   settings panel always wins over a sweep that started before it.

Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated
file for the sweep (kept separate from custom-model-routes.test.ts because
that file's data dir is shared across every test in it — one temp HOME per
FILE, not per test — which would make a sweep-touches-every-host assertion
meaningless there). test/custom-model-run-menu-ui.test.ts gained a new
describe block driving the real picker modal through JSDOM: single-model
bypass, multi-model dialog with the default marked-not-chosen, picking a
row closes the modal and launches with that exact model, the endpoint
re-fetch, and the two "vanished by click time" toast paths.

Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md,
docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated
— the last of these also caught up two sentences that had gone stale after
the draft-review fixes landed (the picker routes through run() now, not a
raw run*() call).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 08:50:18 +08:00
DevvynandClaude Sonnet 5 60e1bd52f7 fix(custom-model): act on the draft review — unparseable onclick, unwrapped envelope, wrong-session apply, missing lock, no tests
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.

Blockers:

1. Every generated inline onclick was unparseable. JSON.stringify's own
   double quotes terminated the double-quoted HTML attribute at the first
   one, leaving btn.onclick null on every picker entry and every Discover/
   Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
   argument, the same idiom deleteCase's onclick already uses four lines
   away in session-ui.js. This also closes the live-HTML-injection route
   through modelId (server-controlled, from the endpoint's own /v1/models
   reply): with quoting intact, a `>` inside it can no longer terminate the
   <button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
   like every other /api route (server.ts's preSerialization hook applies
   to arrays too), so Array.isArray(hosts) was always false in production
   and the picker/settings panel silently saw nothing. Both call sites now
   go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
   returns normally without ever changing activeSessionId, so the apply
   step used to silently re-point and restart whatever session the user was
   already looking at. runCustomModelEntry() now snapshots activeSessionId
   before the launch and requires it to have actually changed.

Majors:

4. Routes the launch through run() itself via a temporary _runMode swap
   (never persisted — setRunMode() would sync it to the server) instead of
   a parallel hardcoded dispatch table, so a custom-model launch now holds
   the same _runInFlight lock every other Run click gets. This also
   resolves the "hardcoded runners map contradicts the PR's own design"
   minor: dispatch is run()'s own, so a CLI whose customModelInjection
   recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
   against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
   parses markup this module generated itself) for exactly the DOM-level
   facts the review said needed no Playwright and no tmux: a generated
   button's onclick genuinely compiles and fires, a dangerous modelId never
   produces a live element, the envelope unwrap works, the session-changed
   guard holds, run() actually gets called (proving the in-flight lock
   engages), and _runMode is restored afterward. Confirmed against the
   pre-fix code first (reproduces btn.onclick === null exactly) so this
   isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
   render-index-html.test.ts for the other fixes below.

Minors:

- Generated entries now filter through isCliAvailable(), matching
  _refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
  (applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
  and to settings-modal open) instead of always rendering; the endpoint GET
  no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
  redactApiKey() replaces the field with a computed apiKeySet: boolean, and
  a PUT with no apiKey now keeps the stored one server-side
  (applyStoredApiKey()) instead of the client resending a value it was
  never given. New tests cover both directions (kept vs. replaced) by
  observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
  (_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
  event, since the real role can resolve after settings were first opened)
  — endpoint writes were already admin-only server-side, but the button
  used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
  toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
  route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
  that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
  (CLAUDE.md already records that exact literal turning the settings
  preview into a grey slab on light skins), .run-mode-custom-models gets
  the same gap: 2px .run-mode-menu's own flex gap only applies one level
  up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
  </script> (CliEntry.label is user-clis.json-settable, unlike
  __codemanCliAvailable's booleans-only payload) via a new exported
  escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
  docs/api-reference.md; left the "no zh-CN for the new Models-section
  group" minor unaddressed only insofar as the wider Models section (task
  routing, thinking effort, etc.) has never had zh-CN coverage either —
  everything this PR itself introduces (labels, hints, button text, the
  Run-menu's "Custom Endpoints" header) IS translated in i18n.js.

Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
DevvynandClaude Sonnet 5 fed6582d3e fix(test): strip the custom-model Run-menu picker's injected script too
CI on PR #430 failed test/server-index-title.test.ts's byte-identity
check: renderIndexHtml now injects a second unconditional <script> before
</head> (window.__codemanCustomModelClis, added alongside the existing
__codemanCliAvailable one), and the test only knew to strip the older one
before comparing the rendered HTML against the raw template.

Strip both. Unlike __codemanCliAvailable (an object, historically injected
only where something resolved), the new one is a plain array injected
unconditionally, possibly empty, so it needs stripping on every machine,
not just one with CLIs installed.

Verified the two replace() calls compose correctly against the exact
strings server.ts actually produces (simulated in isolation; this box has
no tmux, so the real WebServer-backed test file cannot run here at all --
same environment gap noted throughout this PR's review).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
DevvynandClaude Sonnet 5 98d26e14d9 docs(wiki): document Custom Model Endpoints and the Run-menu picker
New docs/wiki/Custom-Model-Endpoints.md (auto-synced to the live GitHub
wiki on push to master, per docs/wiki/Contributing.md) covers turning the
feature on, adding an endpoint, the Run-menu picker's one-off-run
behaviour, the per-harness confidence table, and what it deliberately does
not do yet (remote/Docker sessions, live hot-swap). Linked from the
sidebar, from Agent-CLIs.md's "Read next" list plus a short pointer
section, and from Settings-Reference.md's Models section.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
DevvynandClaude Sonnet 5 25fae9ad10 feat(custom-model): generate Run-menu entries from saved endpoint profiles
Follow-up to #393, picking up the work Ark0N invited in his merge comment:
"generate those entries from the saved profiles rather than a fixed
duplicate per harness, and put it in a follow-up PR so this one stays the
backend... The Run-menu picker is yours if you want it."

Adds the frontend surface the backend has been waiting on:

- Run menu: a "Custom Endpoints" section lists one entry per (harness that
  supports customModelInjection, saved endpoint) pair, e.g.
  "Claude Code (llama.cpp)". The harness list comes from
  window.__codemanCustomModelClis, injected at page render straight off the
  CLI registry's own capabilities (never a hardcoded id list in the
  frontend), so a CLI whose injection recipe lands later appears with no
  frontend change. Picking an entry runs that harness's own existing run*()
  function unmodified (case creation, env overrides, everything, forced to
  a single instance) and then applies the endpoint's default model to the
  session it creates via the existing POST /api/sessions/:id/custom-model
  route. Entries are hidden for a remote/docker active case, since that
  route already refuses both.
- Settings: App Settings -> Models gets a "Custom model endpoints" group
  wiring up the customModelEndpointsEnabled toggle (declared since #393,
  read by nothing until now) plus CRUD against the existing
  /api/model-endpoints routes: list, add/edit (inline form), delete,
  discover models.
- Backend: CustomModelHost gains an optional defaultModelId, the model the
  picker applies with no further choice per endpoint (one generated menu
  entry per CLI+endpoint pair, not per CLI+endpoint+model). The route
  refuses a value that isn't one of the endpoint's own discovered models,
  and a fresh discovery drops a default that no longer appears rather than
  carrying an invalid one forward.

Docs: docs/custom-model-endpoints.md describes the new picker and settings
panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the
"backend-only" status note and documents the picker's generation mechanism.

Tests: four new route tests cover defaultModelId validation, acceptance,
and the drop/keep behaviour across a re-discovery; a new render-index-html
test pins the __codemanCustomModelClis injection (present, agent CLIs
supporting the capability, antigravity and shell excluded) and its
solo-window skip. No browser test was added for the Run-menu picker itself
or the settings CRUD panel (this box has no tmux, so the live server used
by test:browser/test:mobile could not be exercised here) -- worth a
Playwright pass before merge, same as any other frontend PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
2026-09-16 07:01:23 +08:00
Randalix 4a30f510e6 fix(remote): let the host config turn wake-on-LAN OFF for a live session too
Found by driving the real UI: with a MAC configured in remote-hosts.json, removing it
(here: to reach the "Configure WoL" dialog) changed nothing for a running session —
_effectiveRemote short-circuited on the session's own snapshot whenever that snapshot
HAD a target, so the resolver was only ever consulted in the one direction where the
feature was missing. The documented "host config is authoritative" promise therefore
failed in the direction a user can actually observe, and a wake target could live on
invisibly after being deleted from the config.

The resolver is now consulted on the TTL regardless, and wins for the wake fields in
both directions. Also adds a route test for the browser's real input shape: one POST
per keystroke, all buffered during a wake, replayed IN ORDER.
2026-09-15 23:21:50 +02:00
Randalix d0a5a583cd feat(remote): wake a sleeping host when a session is created or attached
Pressing Run on a remote case whose host was asleep failed with
`could not verify tmux on remote host 192.168.50.137: …` — an ssh error that
blames tmux for a machine that is merely suspended. The only wake paths were
typed input on an established session and the banner's Wake button, so OPENING a
session (the moment the user actually decides to use that host) had none.

`RemoteWakeRegistry.ensureHostAwake()` reuses the existing probe/wake/readiness
machinery for a host that has no session yet, and is wired into the two
user-initiated create paths: `POST /api/quick-start` for a remote case (before
the tmux prereq probe, which is what surfaced the misleading error) and
`POST /api/sessions` with `attachRemoteSession`. A host without a wake target is
not even probed, so its behavior and latency are byte-identical. The wake is
blocking — the caller gets the session or an error — but bounded by
REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS (40 s) instead of the 90 s session default,
because the dashboard sits behind a reverse proxy whose default
`proxy_read_timeout` is 60 s: a longer wait would be cut off at the proxy while
the session was still being created. The budget has to cover the whole request
(40 s wake + 1.5 s probe + the tmux probe's own 15 s = 56.5 s worst case), which
is why it is 40 s and not 45. A timeout now says the host did not come back, and
an unreachable host without a wake target says so instead of pointing at tmux.

The wiring is deliberately in the HTTP ROUTE, never in the shared session
service: `cron-service.ts` builds sessions there with nobody waiting on the
answer, and a wake on that path would power the host on for every schedule —
the timer-driven re-wake invariant #1 exists to prevent. Both halves are asserted
(importers of `remote-wake`, and `ensureHostAwake` having exactly one caller
file), so a future caller has to come through the guard test. A rejection from
the wake IO is caught too: a broken target must fail the wake, not the route.

`remote:hostWaking`/`remote:hostWakeFailed` now carry `forNewSession` for the
session-less case, where "input is queued" would be untrue; the toast then reads
"the session starts when it is back".

Live wake numbers are unchanged (this reuses the measured ~12 s S3 path); the
route behavior is covered by new tests in session-routes.test.ts with an injected
registry, so no test opens a real socket or ssh.
2026-09-15 22:37:37 +02:00
Randalix 8dfc965d13 fix(remote): stop the wake handlers shadowing each other; enforce the input cap
Two findings from a final review pass over the wake-on-LAN feature.

`_onRemoteHostWaking` / `_onRemoteHostWakeFailed` were defined in BOTH
`panels-ui.js` (toasts) and `host-wake-ui.js` (banner). Both files mix into
`CodemanApp.prototype` and `host-wake-ui.js` loads later, so the panels-ui copies
were silently shadowed: the toast never fired, and a wake started for a BACKGROUND
session (input on a non-active tab) produced no notification at all, since the
banner handler only acts on the active session. The handlers now live only in
`host-wake-ui.js`, show the toast unconditionally, and update the banner when the
woken session is the active one.

`appendBoundedPending` dropped only WHOLE chunks, so a single input value over the
cap (one large paste is one `input` value, up to the 100 KB input schema) was kept
in full: "bounded at 4 KB" held per chunk, not per session, and nothing was logged.
The surviving chunk's head is now trimmed too, code-point aware so a multi-byte
character is never split into a replacement char.

Adds the guard that would have caught the first one: every SSE dispatch handler must
be defined in exactly ONE frontend module. The existing test only asserts a handler
EXISTS somewhere, which two modules both satisfy while one is shadowed.
2026-09-15 21:01:53 +02:00
Randalix 1380b023e2 fix(remote): make the wake banner's poller page-wide and independent of tab switches
Reported as 'the tab shows no banner' while the host was verifiably unreachable: the
banner only started polling from selectSession, which RETURNS EARLY for the tab you
are already on (so a page loaded with the remote tab active never polled), and a
long-lived tab keeps running the JS it loaded — the feature was invisible to anyone
who did not switch tabs after the deploy.

The poller is now page-wide: one interval (created on init and on the first session
switch), re-targeted whenever the active session changes, plus a visibilitychange
wake-up. It no longer depends on any single selection path running.

Also adds test/sse-dispatch-table.test.ts: a static guard that every
[SSE_EVENTS.X, '_onFoo'] entry names an event constants.js defines AND a handler some
module defines. Both halves fail silently (a typo'd constant is an undefined table
key; a renamed handler just never runs), which is exactly how a new banner can never
appear with no error anywhere.
2026-09-15 15:11:36 +02:00
Randalix e8f7772320 fix(remote): offer the WoL config dialog after a failed wake too
A configured-but-broken target (host replaced NIC, command removed) had no way
out: the dialog hung off the 'no target configured' branch only, so the banner
would keep offering a Wake button that keeps failing.
2026-09-15 14:38:24 +02:00
Randalix 2f61be6e74 fix(remote): bind the wake socket before enabling broadcast
setBroadcast() on an unbound dgram socket throws EBADF on Linux and the following
send fails with EACCES, so the magic packet silently never left the machine — the
feature reported a wake that never happened. Caught by waking a real sleeping host
(a unit test with a real UDP broadcast would not be welcome in CI, so the socket is
injectable and the bind-before-setBroadcast ORDER is asserted).
2026-09-15 14:24:13 +02:00
Randalix 8b5a13435a feat(remote): host-unreachable banner, manual wake, and native MAC wake-on-LAN
The reactive wake (typing into a session whose host slept) left the state invisible:
nothing told the user the machine was asleep, and with no wake target configured
there was nothing to do about it. Adds:

- RemoteHost.wakeMac (comma-separated) - Codeman builds and broadcasts the magic
  packet itself (UDP port 9), so the common case needs no external script. The
  existing wakeCommand stays as the explicit override.
- GET /api/sessions/:id/reachability - probes (throttled, cached, and it never
  wakes) and reports HOW the host can be woken, or that nothing is configured.
- POST /api/sessions/:id/wake - wakes, waits, reattaches the pane and flushes
  buffered input; 400 with a routable message when no target is configured.
- The amber host-unreachable banner + its 'Wake' / 'Configure WoL' action, and a
  small config dialog that saves via PUT /api/remote-hosts/:id.
- RemoteWakeDeps.resolveRemote: host config is re-resolved for LIVE sessions
  (throttled + cached), so saving the dialog takes effect without a restart.
2026-09-15 14:20:31 +02:00
Randalix 3f0bfde54a docs(remote): document the wake-on-LAN invariants; drop wake state on bulk delete
Self-review pass: the input-ladder's two 'buffer' branches were the same three
lines, and bulk delete left a session's (bounded, per-random-uuid) wake state
behind. Documents the design where the code refers to it - remote-sessions.md
section, the architecture invariant, and the CLAUDE.md key pattern.
2026-09-15 10:45:01 +02:00
Randalix 0f3eea2fb5 fix(remote): refresh wake command from host config when restoring sessions
A session's remote block is persisted at launch time and recovery uses that
snapshot, so a wakeCommand added to remote-hosts.json afterwards never reached
an already-running session - not even across a Codeman restart (observed: the
live Hufflepuff session came back with no wakeCommand). Merge the host-level
field in on restore, with the host config authoritative.
2026-09-15 10:35:58 +02:00
Randalix a81f430e41 feat(remote): wake a sleeping host from user input (Wake-on-LAN)
A durable remote session survives SSH drops (COD-104/108), but nothing brought
the HOST back: after the remote machine suspended, the local tmux pane's ssh
child stalled silently and `send-keys` SUCCEEDS against it, so typed input
vanished with no error anywhere.

Add an optional per-host `wakeCommand` (Wake-on-LAN wrapper, e.g. whuff) that
the input route runs when a wake-enabled host is unreachable: input is buffered,
the host is woken, the pane is reattached, and the buffer is flushed in order.
Detection is a throttled bare TCP probe on wake-enabled hosts only, and only
REAL user input may wake a host - the auto-reconnect watcher and boot recovery
deliberately cannot, or the host would be re-woken seconds after every suspend
and could never stay asleep.
2026-09-15 10:25:52 +02:00
228 changed files with 35945 additions and 2005 deletions
@@ -1,5 +0,0 @@
---
"aicodeman": patch
---
fix(terminal): keep the output a pane capture could not contain. Opening a session, a backpressure refresh, a clear-terminal reload and a full-history re-pull all load the screen from a tmux pane capture, and anything the CLI printed between that capture and the end of the load used to be dropped, so its next partial redraw landed on a frame the terminal had never seen: missing or garbled output right after a tab switch or a refresh, plainest in a shell session. Each load now replays exactly the output that arrived after the capture, through one shared rule for all four paths, and a refresh that restores your scroll position no longer snaps back to the bottom afterwards.
-5
View File
@@ -1,5 +0,0 @@
---
"aicodeman": patch
---
fix(input): make sure a prompt sent through the API actually leaves the composer. Claude Code 2.1.277 started ignoring Enter for the first 30 to 50 seconds after the composer paints while still accepting the typed text, so a prompt sent right after a session came up sat unsent in the pane and every waiter (send-and-wait, the agent skill, cron, the maintainer bot) burned its whole timeout on a turn that never started. The server now reads the pane after every programmatic write that carried Enter and presses Enter again, on a 2 to 60 second schedule, only while the composer verifiably still holds the text it sent; an empty composer, other text, or a pane with no composer at all ends it. The agent skill's `sendwait` gets the same loop for servers that predate this, and its preamble version moves to 1.30.1 so an already-seeded agent picks up the fresh copy.
+1 -1
View File
@@ -10,7 +10,7 @@
"name": "codeman",
"source": "./plugins/codeman",
"description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.30.0",
"version": "1.33.1",
"author": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
+44
View File
@@ -100,6 +100,50 @@ jobs:
fi
echo "bash $BASH_VERSION: dsh identity probe survives a missing timeout"
'
# Installer v2: the question phase runs before the build, and every decision it
# takes is bash logic over stubbed tailscale state. Drive the flags, the launch
# default, the occupied-:443 menu and the rename question with canned answers,
# so a bash-4 construct or a flipped default in any of them fails here, not on a
# Mac. The JSON parsers need node (absent in this image) and are stubbed; their
# own coverage is test/install-sh-invariants.test.ts plus the vitest gate.
docker run --rm -v "$PWD":/w -w /w -e CODEMAN_INSTALL_SH_LIB=1 -e HOME=/tmp/h bash:3.2 bash -c '
set -euo pipefail
mkdir -p /tmp/h
. /w/install.sh
parse_flags --tailscale --service --name Build-Box --port 4000
[[ "$CODEMAN_TAILSCALE" == "1" && "$LAUNCH_PRESET" == "2" && "$TS_NAME" == "Build-Box" && "$CODEMAN_PORT" == "4000" ]]
[[ "$(ts_sanitize_name "$TS_NAME")" == "build-box" ]]
has_tty() { return 0; }
ANSWER=""; read_reply() { eval "$1=\"\$ANSWER\""; }
systemctl() { return 0; }
LAUNCH_PRESET=""; NONINTERACTIVE=0
choose_launch_mode linux >/dev/null 2>&1
[[ "$LAUNCH_CHOICE" == "2" ]]
check_tailscale() { return 0; }
ts_status_field() { case "$1" in "s.BackendState") printf Running ;; "s.Self && s.Self.DNSName") printf "box.tail.ts.net." ;; esac; }
ts_backend_state() { printf Running; }
ts_dns_name() { printf box.tail.ts.net; }
ts_serve_443_target_port() { printf 8080; }
ts_serve_find_port_mapping() { :; }
ts_serve_port_used() { return 1; }
detect_tailscale_serve_url() { :; }
tailscale_choose_mapping >/dev/null 2>&1
[[ "$TS_SERVE_MODE" == "path" && "$BIND_BASE_URL" == "/codeman" ]]
RENAMED=""; tailscale_rename_node() { RENAMED="$1"; }
TS_NAME=""; tailscale_choose_name >/dev/null 2>&1
[[ -z "$RENAMED" ]]
# A flag re-run keeps the password the unit already carries (and so
# never writes the unauthenticated ack), and the hand-start line the
# done screen prints carries every non-default value.
read_existing_binding() { EXISTING_FOUND=1; EXISTING_HOST=0.0.0.0; EXISTING_PASSWORD=s3cret; EXISTING_ACK=0; EXISTING_BASE_URL=""; }
CODEMAN_HOST=0.0.0.0; CODEMAN_TAILSCALE=0; unset CODEMAN_PASSWORD; BIND_ACK=0
choose_network_binding >/dev/null 2>&1
[[ "$BIND_PASSWORD" == "s3cret" && "$BIND_ACK" == "0" ]]
BIND_HOST=0.0.0.0; BIND_PASSWORD=x; BIND_ACK=0; BIND_BASE_URL=/codeman; CODEMAN_PORT=4000
[[ "$(start_command_hint)" == "CODEMAN_HOST=0.0.0.0 CODEMAN_PASSWORD="*" CODEMAN_BASE_URL=/codeman CODEMAN_PORT=4000 codeman web" ]]
RECONFIGURE=0; parse_flags --port 4001; [[ "$RECONFIGURE" == "1" ]]
echo "bash $BASH_VERSION: question phase (flags, launch default, occupied :443, rename opt-in, kept password, start line) ok"
'
- name: CLI catalogue artifacts are in sync with stock.ts
run: npm run generate:cli-catalog -- --check
+216
View File
@@ -1,5 +1,221 @@
# aicodeman
## 1.33.1
### Patch Changes
- ### Thanks
- @irisitymichaelgrundberg for closing sessions whose agent exited cleanly (#486), built carefully around every way a pane exit can lie (a SIGKILL with no status, a single misread), with the `.claude-images` guard split into its own commit as asked.
- @opticon454 for the live-refreshing case picker and Manage search (#483), and for the uv/uvx, libsecret and pnpm additions to the Docker images (#487, #485).
**Finished sessions close themselves (#486).** A session whose agent you ended with `/exit` is now closed the same way the X button closes it, so finished sessions stop piling up on the board; the conversation stays resumable from the Resume list and the lifecycle log records "agent exited cleanly (status 0)". Only an explicit exit status 0 with no signal, confirmed by two pane reads, qualifies: a crashed or OOM-killed agent keeps its row with the exit code on the tab. The phone overview and desktop home rail now say `exited` instead of `idle`, reboot restore no longer offers to rebuild a session whose agent had exited, and closing one session no longer deletes the `.claude-images` directory that a sibling session in the same case still uses. Thanks @irisitymichaelgrundberg.
**Search in the phone Select Case sheet (#488).** The bottom sheet gains a "Search cases" field that filters by name (every word must match, any order, ignoring case), Enter picks the case when exactly one row is left, and Escape clears then closes. Also fixes a dead band under Create New Case and a list shorter than the sheet could show.
**An oversized paste no longer jams a session's input (#484).** A single input over the 64 KiB frame limit used to be refused by both transports, retried every 2 s forever, block every later input for that session and come back from localStorage on each reload. Pastes over the limit are now split into in-limit frames delivered in order (up to 1 MiB; larger ones are refused with a toast and never queued), a refused frame is dropped instead of retried, frames persisted by an older build are pruned on load, and the WebSocket answers an oversized frame with an explicit `too_large` error instead of silence.
- 8841bcc: Add a search box to the Manage tab of the Add Case dialog. It filters the case list by name or path, and the reorder arrows are disabled while a filter is active so a swap cannot involve a hidden case.
- 8841bcc: The case picker now refreshes its list from `/api/cases` when it opens and every 5 seconds while it stays open, so folders deleted or created on disk appear without a page reload. If the selected case has been removed, the picker falls back to another case without saving it as the last-used one.
Thanks @opticon454.
- 77ba41f: Install `uv` and `uvx` in the Compose server image and the agent image, so MCP servers launched with `uvx` (such as the Nginx Proxy Manager MCP) can be enabled by Codex instead of failing with `uvx` not found. Both images also install `libsecret-1-0`, the native library the `keytar` dependency of the Azure DevOps MCP (`@azure-devops/mcp`) needs; without it the server crashes before answering the MCP initialize handshake.
The Compose server image now also carries `pnpm`: `dsh plugin` spawns a literal `pnpm` with no npm fallback, so the Run menu's "DeepSeek - add a terminal profile" button failed with `dsh: pnpm not found on PATH` there. Because this release changes `server.Dockerfile`, the in-app updater asks Compose deployments to rebuild the image (`Update-Codeman.sh`) rather than applying it in place.
Thanks @opticon454 (#487, #485).
## 1.33.0
### Minor Changes
- CLI management from Settings (#476, finishing the CLI registry work from #343). `~/.codeman/clis.json` used to be hand-edit only; with the new opt-in `cliManagementEnabled` switch (synced, default OFF) App Settings → Agents & CLIs can enable or disable any CLI, install a missing stock CLI with its vetted install command, and add, edit or remove custom CLIs. Six new endpoints back it (`GET`/`POST /api/clis`, `PUT /api/clis/:id`, `POST /api/clis/:id/install`, `PUT /api/clis/custom/:id`, `DELETE /api/clis/:id`), documented in `docs/api-reference.md`. Every write is refused while the switch is off, is admin-only in multi-user mode, is serialized on one queue, and refuses to overwrite a `clis.json` that does not parse or has group/world permission bits. A custom entry is re-validated through the same schema as the stock ones and its install text is never executed. `shell` cannot be disabled. The Run menu and the welcome screen are now built from the enabled catalogue, so the welcome screen also offers Codex, Shell and any custom CLI, and the stock Claude entry is labelled "Claude Code".
Models: Opus 5.5 (`claude-opus-5-5`, 1M context capable) is offered in App Settings → Models and in task routing (#480).
Self-update: on a macOS `launchd-daemon` install, a Homebrew node upgrade could leave `update-status.json` stuck at `queued`, which made every later update fail with "An update is already in progress." The updater now falls back to `node` on PATH when the server's own node binary is gone, and an in-flight status that has not been written for 15 minutes is failed on the next read. A graceful shutdown that hangs is now force-exited after 10 s (and the launchd updater SIGKILLs a server that has not exited after 30 s), so launchd can start the new build instead of leaving the service down (#478). Both fixes protect updates that start FROM this release.
Session Manager (Cmd+K): rows keep their `mode`, `claudeSessionId` and `resumeId`, so the ⋯ menu's Resume session relaunches a Codex row as Codex on its own conversation, and the mode badge shows as it does on the home list (#477).
Maintainer fixes applied while landing #457: renaming a tab to the name it already has (the Session Options field saves on blur) is now a no-op, so it no longer pins the placeholder as the `/resume` title again; Docker sessions skip the transcript title sync, since their transcript lives in the container; and the agent skill's messaging examples no longer use a `w<N>-` name as the peer name.
Tests: the suite strips every inherited `CODEMAN_*` variable, so running it inside a Docker Compose deployment no longer writes into the deployment's real case root (#479).
### Thanks
- @opticon454 for CLI management (#476), the last piece of the CLI registry, with every review item answered in one round, and for splitting the test isolation fix out into #479.
- @shenlvkang-collab for the `/resume` title fix (#457) and the careful diagnosis behind it.
- @julian3xl for the Session Manager row fix (#477), their first contribution.
### Patch Changes
- 69a7128: fix(sessions): stop pinning the `w1-myapp` placeholder as Claude's session title. Local Claude spawns passed the tab name as `--name`, which is also the `/resume` picker entry and the terminal title, and a pinned title stops Claude generating its own, so every conversation of a case showed up in `/resume` as the same `w1-myapp` and none got a generated title. Only a name the user chose is pinned now; placeholder and auto-named tabs let Claude title the conversation again. Renaming a Claude tab also reaches `/resume`: the new name is appended to the conversation's transcript as the `custom-title` row `/rename` writes (a tab that was spawned with `--name` keeps re-appending its own title until its next respawn, so the rename wins from then on). Orchestrators that rely on a fixed peer name should give workers a descriptive `sessionName` rather than a `w<N>-` one.
## 1.32.1
### Patch Changes
- 13e652e: Terminal copy: copying text out of a Claude Code or Codex pane no longer puts the pane's two-column transcript gutter on the clipboard, so pasted lines arrive flush instead of indented (#469). The width comes from the CLI registry (`capabilities.transcriptGutter`, 2 for claude and codex, measured on live panes) and is only a ceiling: a selection only ever shifts as a block, so its own indentation survives. Other CLIs and shells are untouched. It works in split panes and detached session windows too, and can be turned off per device in App Settings under Selection & clipboard.
- 13e652e: Sessions: recovering a Claude session whose tmux pane had died relaunched `claude --session-id <id>`, which Claude refuses once that id has a transcript, so the pane died again straight away and the conversation was stranded. The relaunch now resumes the conversation (`--resume <id> || --session-id <id>`), including when tmux lost the whole session (#467).
- 00f022c: Terminal: when a burst of output overflows the render queue and a frame has to be dropped, the repaint that repairs it is now retried until it actually happens, instead of being scheduled once and silently skipped when another load was in flight (#470).
- 13e652e: Mobile: a long press on blank terminal space on Android Chrome no longer opens the keyboard and blanks the terminal (#471, fixes #360). The long-press guards are now armed before the press is checked for selectable text, so a press on empty space is swallowed the same way a press on a word already was.
- 13e652e: Sessions: a tab whose agent has exited (the CLI quit, but tmux kept the pane) now says so with a muted dot and an `exited (137)` badge, instead of looking like an idle session (#466, part 1 of #446). The state is published as `paneExit` on the session and survives a restart. Nothing closes such sessions yet; that is part 2.
- 13e652e: Docker: optional GitHub CLI and Azure CLI for private repositories (#472). Both are off by default. With `CODEMAN_INSTALL_GH=1` / `CODEMAN_INSTALL_AZ=1` as build args in `docker-compose.override.yml`, the server image gets `gh` and/or `az` (with the `azure-devops` extension) wired in as git credential helpers, so after one `gh auth login` or `az login` from a shell session, Add Case → Clone Repo can clone private GitHub and Azure DevOps repositories. `CODEMAN_AGENT_IMAGE_INSTALL_GH` / `_AZ` do the same for the Docker-case agent image, and only then are the sign-ins copied into new case containers. In multi-user mode a non-admin's clone runs with the credential helpers cleared. This changes `server.Dockerfile`, so Compose deployments need a `Start-Codeman.sh` rebuild rather than an in-app update.
- 13e652e: Run menu: the Gemini, Antigravity and OMP run buttons now show their own colours on every skin; they rendered in Claude blue on all skins except OG (#463). The CLI registry's `accent` values were also corrected to the colours the UI really paints, and a test now guards the stylesheet trap that caused it.
- 13e652e: Terminal: five ways the browser terminal could silently stop being correct are fixed (#431, #464). The browser terminal and the PTY can no longer disagree about their width, which is what produced doubled lines and half-overwritten text ("text gets muffled sometimes"): there is now one function that sizes the terminal, and every resize is answered with the geometry the PTY really holds. A replay clear goes through the terminal's own queue, so bytes written just before it no longer fuse into the next snapshot. A renderer that stops painting after an iOS PWA is backgrounded heals itself instead of needing a reload. Every terminal capture has a deadline that also covers the response body, and a capture that runs out of time during a tab switch falls back to the bounded tail instead of leaving a blank pane. Output lost to a half-open WebSocket is repainted on the next successful open. The service worker's precache list is now generated by the build and its cache is rotated per build, so old releases' assets no longer pile up.
- 13e652e: Docker: new `docker/Update-Codeman.sh` for the major-update path the docs used to describe by hand (#465). It rebuilds the image with `--no-cache` before taking the stack down, clears the build-artefact volumes, refuses to run when another checkout's Compose project already owns the same name, and then hands over to `Start-Codeman.sh`.
- 13e652e: Approvals: a session that is idle only because it is waiting on its own background work (Claude Code's `1 monitor` footer chip, or a Codex background terminal) no longer raises the yellow NEEDS YOU alert or a push (#473, fixes #468). Its idle item is opened already acknowledged, and the tab, the home screens and the rail show a small `watching` badge next to the state instead. The item still exists in the Approvals Inbox, and the TUI's pending count now leaves acknowledged items out.
- b404dac: Maintainer fixes applied while landing this batch:
- Terminal (#431): while another device holds the pane's width, a resize retry no longer re-fits xterm to the container and re-wraps the whole buffer every 30 s, and no longer clears scrollback for a redraw that never comes. The PTY's spawn geometry is now recorded at attach, so `ptyGeometry` never reports a size the PTY never held.
- Terminal (#470): the `TERMINAL DROP` crash-trail line is logged once per recovery window instead of once per dropped frame (which wiped the rest of the trail within a second), and a refresh that died at its fetch deadline is no longer retried.
- Sessions (#467): the resume pin also covers the branch where tmux lost the whole session, the conversation id Codeman reports follows what the relaunch actually resumed, and the test setup strips `CLAUDE_CONFIG_DIR` so the suite stays green for anyone running a separate Claude config dir.
- Sessions (#466): detailed sidebar and rail rows show an `exited` pill instead of `idle`, the exit is announced to screen readers, and the user manual's tab-appearance table lists the new state.
- Approvals (#473): a failed pane capture clears the `watching` badge rather than keeping a stale one (a failure now falls toward an alert, not toward silence), and the header bell's count leaves acknowledged items out, matching the TUI.
- Run menu (#463): the Gemini and Antigravity run buttons no longer render two-tone on phones, Gemini's registry accent matches its tab badge, and a test now guards the stylesheet trap for every run mode.
- Docker (#465): `Update-Codeman.sh` removes exactly the two build-artefact volumes it names instead of every named volume in the project, reports a failing `docker compose` instead of exiting silently, and its docs and comments were corrected. (#472): the multi-user notes say that a non-admin's seeded Docker case also receives the gh/az sign-in when those switches are on.
### Thanks
- @irisitymichaelgrundberg for four PRs in this release: the `watching` badge that stops background work from raising false alerts (#473, from their own report #468), the exited-agent badge (#466) and the dead-pane resume fix (#467), both from their report #446, and the transcript-gutter strip for copied text (#469), a follow-up to their #451.
- @rounakdatta for the terminal resilience work (#431) and the dropped-frame recovery (#470), both from their report #464, and for answering four rounds of review in full.
- @opticon454 for private-repository support in the Docker images (#472), the `Update-Codeman.sh` script (#465) and the run-button colour fix (#463).
- @DodgyBadger for the Android long-press fix (#471), from their own report #360.
## 1.32.0
### Minor Changes
- d47f93a: feat(custom-model): the model picker puts the ready model first
When a custom endpoint has more than one model, the Run menu's picker now promotes one row to the top instead of showing raw discovery order: the model llama-swap reports loaded and ready right now (tagged "Currently loaded", the one a launch attaches to with zero wait), else the model you last launched on that harness and endpoint (tagged "Last used", remembered per device). The endpoint's default keeps its own pill, nothing is ever auto-chosen, and a plain OpenAI-compatible server or an endpoint that does not answer within a second simply keeps the old order. The probe is bounded on the client too, so a GPU box that is off no longer holds the picker closed for five seconds.
- d47f93a: feat(split-pane): view two live sessions side by side
A new Split button in the header (opt-in in App Settings, off by default, desktop only at 1180px and wider) opens a picker and shows a second live session beside the active one: its own terminal, its own WebSocket, and a divider you can drag. When either session ends the view collapses back to one pane, with Pane B promoted to the primary when it is Pane A that ended. Nothing is persisted on purpose in this first cut, so a page reload always returns to a single pane. Pane B is deliberately plainer than the primary pane (no local-echo overlay, CJK input, touch handling or keyboard accessory bar); the design and the v2 boundaries are in discussion #452.
- 72d437a: Installer v2. `curl -fsSL https://getcodeman.com/install | bash` now looks at the machine first, asks at most three questions up front (how the dashboard is reached, optionally what to call the machine on your tailnet, whether to run Codeman as a background service), does the install unattended behind progress spinners with the output in `~/.codeman/install.log`, and ends on the URL with a QR code to scan. One consent covers every missing package and sudo asks for your password once. Flags pipe through `bash -s --` (`--tailscale | --lan | --local`, `--name <n> | --no-rename`, `--service | --run | --no-start`, `--yes`, `--password`, `--port`), `install.sh status` prints the URL and the QR code again, and the cloudflared question moved out of the main flow into `install.sh cloudflared`. On the Tailscale route, a `:443` that already belongs to another app gets Codeman under `https://<node>/codeman` (or on a second port) instead of a dead end, the node can be renamed opt-in (`--name`, `install.sh name`, undone by uninstall), and the HTTPS-certificates toggle is polled with the admin page opened for you. Also fixed on the way: the installer's own `npm install` no longer lets the postinstall start a stray server on port 3000 (the service crash-looped on EADDRINUSE while the done screen said "running"), the LAN address comes from the default route rather than the first interface, a hand-written LaunchDaemon on a headless Mac is left alone, a flag re-run keeps an existing dashboard password, and the done screen's start command carries the sub-path and port it was installed with.
- d47f93a: feat(mobile): a Compose key for writing prompts on a phone
The agent keyboard bars on phones replace their Paste key with Compose: a real multiline editor with autocorrect and spellcheck, per-session drafts kept in memory only, image attach that never writes into the terminal early, and a Send that delivers the text as one paste followed by Enter, so a long prompt no longer has to be typed blind into the terminal composer. Anything you had already typed into the terminal is picked up into the editor. Shell sessions keep the direct Paste key. This is the manual first slice from #359; the auto-open setting and terminal tap routing are a separate follow-up.
### Patch Changes
- d47f93a: refactor(run-menu): one table-driven launcher for every external CLI
The eight near-identical per-CLI launch functions in the Run menu collapsed into one launcher driven by a table that a CI test keeps in step with the CLI registry, and a second no-id-branching guard now covers the frontend the way the backend guard covers the server. No behaviour change: the refactor was verified byte-identical across 288 launch permutations against the previous code.
- e899af4: Maintainer fixes applied while landing the above. The model picker's promoted row keeps its Default pill (the promotion tag and the default marker are two pills now, and they render as pills in the picker rather than as plain text). The phone composer keeps its bottom gutter on folding devices (the generic fold rule used to erase it), a whitespace-only draft is no longer sent, and its dialog is translated on a zh-CN UI. A split that collapses mid-drag no longer leaves the page stuck in resize-cursor mode, Pane B refuses a session that has no live process, and a burst of refresh frames replays once instead of twice. The `</head>` script injections on the page render use replacer functions, so a CLI label containing `$'` can no longer splice the document into the inline script, and the frontend no-id-branching guard now catches comparisons on any variable name.
- 6ef71ec: ### Thanks
- @timkjr for split-pane sessions (#453): five review rounds turned around in two days, and the pointer-capture edge case measured in a real browser rather than reasoned about.
- @DodgyBadger for the mobile prompt composer (#444), a first contribution that took the scope back down to one slice when asked, and that verified the delivery path against a live tmux pane and a live Claude Code composer instead of trusting the diff.
- @opticon454 for putting the ready model first in the picker (#459) and for collapsing the eight Run-menu launch functions into one (#458), proven byte-identical across 288 launch permutations instead of argued.
## 1.31.0
### Minor Changes
- 035bfbc: feat(remote): wake a sleeping remote host from Codeman
A remote SSH case pointing at a machine that suspends used to fail the same way every
time: the session was there, the host was not, and typing into it went nowhere. A host
can now carry a wake target, either a MAC address for Wake-on-LAN (Codeman builds the
magic packet itself, so nothing reaches a shell) or a wake command of your own, and
Codeman uses it when you ask for the host: when you type into a sleeping session, when
you press the wake button on the banner, or when you start or attach a session on that
host. Input you type while it wakes is buffered and flushed once it is back, up to 4 KB,
and a chunk over that is refused outright rather than delivered as a fragment.
Waking only ever happens because you asked. No watcher, dropped-session handler or
boot-recovery path can reach it, since a machine woken by a reconnect watcher would come
back seconds after every suspend.
- fbee1b2: feat(custom-model): pick a custom endpoint straight from the Run menu
#393 landed the backend for custom model endpoints and left it reachable only over the
HTTP API. This is the rest of it. Turn on Custom model endpoints in App Settings, save
an endpoint, and the Run dropdown grows a Custom Endpoints section built live off the
CLI registry, one entry per harness that can actually redirect plus each endpoint you
saved. Pick one and it launches that harness pointed at your server, asking which model
first when the endpoint has more than one. Endpoints re-discover themselves every five
minutes, and one unreachable endpoint never blocks the others. App Settings gains full
add, edit and delete for endpoints.
Seven of the harnesses (opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP) now launch
directly onto the endpoint with no restart at all, where before you watched a native
boot followed immediately by a second one. Claude still launches and then restarts in
place, which its own resume makes far less jarring.
Most of this release's work went into things that only show up against a real server,
and each was found that way rather than in tests: a freshly launched CLI reporting
itself busy for its own startup and getting refused; Claude Code assuming a large
context window for a model it does not recognise and silently overflowing a small one;
a model whose real context is below what Claude Code's own system prompt costs, which
no setting can fix and which now warns before launching into a certain failure; and the
big one, llama.cpp running exactly one model at a time, so applying a selection can
unload the model another session is using. That last case now asks first, tells you
which session it affects, and keeps a "loading model" notice on screen for the whole
swap window, so a prompt sent mid-swap reads as loading rather than as an answer from
whatever was loaded a moment ago. A background sweep also catches the reverse: your
session's model being evicted later by somebody else's ordinary use.
Two things worth knowing if you drive this over the HTTP API or run multi-user. The two
questions an apply can ask (the model's context window is too small, and loading it will
unload the model another session is using) are now answered by separate
`confirmedContext` and `confirmedSwap` fields rather than one `confirmed`. They shared a
flag until now, and since the context check runs first, confirming that one silently
agreed to evict another session's model as well. The old `confirmed` still means both.
And `CLAUDE_CONFIG_DIR` is now admin-only in multi-user mode: it joined claude's
privileged env keys, so a non-granted owner can no longer set it through `envOverrides`,
and an already-persisted one is dropped on reboot-restore, which returns that session to
the default Claude account rather than the per-client one it was pointed at. Single-user
installs are unaffected.
Remote SSH and Docker sessions are refused for now, since their restart reattaches a
durable tmux rather than relaunching the agent.
### Patch Changes
- c9515b1: fix(terminal): keep the output a pane capture could not contain. Opening a session, a backpressure refresh, a clear-terminal reload and a full-history re-pull all load the screen from a tmux pane capture, and anything the CLI printed between that capture and the end of the load used to be dropped, so its next partial redraw landed on a frame the terminal had never seen: missing or garbled output right after a tab switch or a refresh, plainest in a shell session. Each load now replays exactly the output that arrived after the capture, through one shared rule for all four paths, and a refresh that restores your scroll position no longer snaps back to the bottom afterwards.
- 3edf9aa: fix(terminal): replay a pane capture at the geometry it was taken at
Opening a session could draw a frame built for a pane bigger than your terminal. A
taller pane wrote its overflow rows onto the last line and lost the rows underneath
(against a 50-row pane, a 30-row terminal rendered 28 of a 45-line command and drew
the survivors twice), and a wider one wrapped every row and scrolled the whole frame
up by one. The terminal response now reports the geometry the capture was really
taken at, so the browser can see the mismatch and replay once at the size that stuck.
A pane that cannot be sized to fit is diagnosed once per session instead of on every
tab switch.
- 035bfbc: ### Thanks
- @irisitymichaelgrundberg for three terminal fixes in one release: keeping the output a pane capture could not contain (#436), replaying a capture at the geometry it was taken at (#435, five rounds and a Playwright suite that fails against the merge base), and trimming the padding out of a copied selection (#451), where the scan-instead-of-regex call avoided a 2.9s freeze nobody would have traced back to a copy.
- @timkjr for a first contribution that found a real silent failure: the Instance count stepper next to the Run button had only ever applied to Claude, so on the other eight run modes it launched one session and said nothing (#454).
- @Randalix for Wake-on-LAN on remote hosts (#439), built and live-tested against a real sleeping machine, and for reading the whole diff again between rounds rather than only the parts that were asked about.
- @opticon454 for turning #393's backend-only custom model endpoints into the whole feature (#430), and for validating it against a real llama-swap box rather than against the tests: the `/props` versus `/running` context discrepancy and the DeepSeek `/v1` root cause were both tracked down to the SDK source instead of guessed at.
- c376534: fix(run): make the Instance count stepper work for every non-Claude mode
The Instance count stepper next to the Run button only ever applied to Claude.
Setting it to 3 and launching OpenCode, Codex, Gemini, Antigravity, Pi, OMP, Grok or
DeepSeek started exactly one session, with no error and no hint that the control had
done nothing. All eight now launch the count you asked for, and the opening banner
says how many are starting. The one exception is a launch started from the Custom
Endpoints section of the Run menu, which always starts a single session.
- 19ffe9b: fix(input): make sure a prompt sent through the API actually leaves the composer. Claude Code 2.1.277 started ignoring Enter for the first 30 to 50 seconds after the composer paints while still accepting the typed text, so a prompt sent right after a session came up sat unsent in the pane and every waiter (send-and-wait, the agent skill, cron, the maintainer bot) burned its whole timeout on a turn that never started. The server now reads the pane after every programmatic write that carried Enter and presses Enter again, on a 2 to 60 second schedule, only while the composer verifiably still holds the text it sent; an empty composer, other text, or a pane with no composer at all ends it. The agent skill's `sendwait` gets the same loop for servers that predate this, and its preamble version moves to 1.30.1 so an already-seeded agent picks up the fresh copy.
- f9edb33: fix(terminal): trim the padding out of a copied selection
Copying out of a pane put a wall of spaces on the clipboard. xterm hands back
whole screen rows and trims only the cells that were never written to, so the
real spaces a full-screen program paints across the unused part of a row count
as content: measured against Claude Code in a 282-column pane, single lines
arrived carrying 138 trailing spaces. Pasting that into a chat client or an
editor meant deleting the whitespace by hand, while Windows Terminal, iTerm2 and
GNOME Terminal all trim it for you. A copy now drops the trailing run from every
line, on all four paths (the Ctrl+C chord, right-click, the phone selection
button and Auto Copy), while leading indentation is left exactly as it is. An
Alt+drag rectangular selection is copied verbatim, because its columns lining up
is the point of that gesture. A selection holding nothing but padding is refused
rather than copied as bare line breaks.
## 1.30.0
### Minor Changes
+90 -68
View File
File diff suppressed because one or more lines are too long
+11 -6
View File
@@ -19,6 +19,10 @@
<a href="https://github.com/Ark0N/Codeman/commits/master"><img src="https://img.shields.io/github/commit-activity/t/Ark0N/Codeman?style=flat-square&color=1e3a5f" alt="Total commits"></a>
</p>
<p align="center">
⭐ <strong>Like Codeman? <a href="https://github.com/Ark0N/Codeman">Give it a star on GitHub!</a></strong> It takes one click and helps more people find the project. ⭐
</p>
<p align="center">
<strong>English</strong> &bull; <a href="README.zh-CN.md">简体中文</a>
</p>
@@ -61,12 +65,13 @@ The installer asks before every system change, and re-running the same line upda
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js, tmux and a build toolchain if missing (node-pty ships no Linux prebuilds, so it compiles from source), clones Codeman to `~/.codeman/app`, and builds it. A few things worth knowing:
This installs Node.js, tmux and a build toolchain if missing (node-pty ships no Linux prebuilds, so it compiles from source), clones Codeman to `~/.codeman/app`, and builds it. It looks at what is already on the machine, asks at most three questions, then does all the work unattended and ends on the URL with a QR code for your phone. A few things worth knowing:
- **It asks first.** Every system change (package installs, AI CLI download) is prompted, and a menu at the end lets you choose: run Codeman in this terminal, install it as a background service (systemd/launchd, auto-start on boot), or don't start yet. Nothing runs in the background unless you pick it.
- **How it's reachable, your choice.** The installer offers three ways to reach the dashboard: **Tailscale** (loopback bind fronted by `tailscale serve`, so you get `https://<machine>.<tailnet>.ts.net` with a real certificate and your tailnet as the login, no password needed), **any device on your network** (`0.0.0.0`, with a strongly recommended password prompt), or **this machine only** (`127.0.0.1`, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. The highlighted default reflects what is already on the machine (Tailscale when it is already in use, your existing binding on a re-run), and a bare Enter never pulls in new software. A bare `codeman web` started by hand still defaults to loopback.
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh update` and `install.sh uninstall` also exist.
- **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
- **Three questions, all up front.** How the dashboard is reached, optionally what to call this machine on your tailnet, and whether to run Codeman as a background service (systemd/launchd, auto-start on boot; Enter says yes). Everything that needs you, including one consent for all missing packages, one sudo password, and the Tailscale login, happens before the build, so you can walk away while it compiles.
- **How it's reachable, your choice.** **Tailscale** (loopback bind fronted by `tailscale serve`, so you get `https://<machine>.<tailnet>.ts.net` with a real certificate and your tailnet as the login, no password needed), **any device on your network** (`0.0.0.0`, with a strongly recommended password prompt), or **this machine only** (`127.0.0.1`, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. The highlighted default reflects what is already on the machine (Tailscale when it is already connected, your existing binding on a re-run), and a bare Enter never pulls in new software. If another app already owns `:443` on your node, Codeman goes under `https://<machine>.<tailnet>.ts.net/codeman` or on a second port instead of replacing it. A bare `codeman web` started by hand still defaults to loopback.
- **The name is yours to choose.** By default the URL uses the machine's existing tailnet name. Answering yes to the second question renames the machine to `codeman-<hostname>` (which also renames it for SSH, so the default is no); `install.sh name` does it later.
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh status` prints the URLs and the QR code again; `install.sh update`, `install.sh tailscale` and `install.sh uninstall` also exist.
- **Flags for the impatient.** `curl -fsSL https://getcodeman.com/install | bash -s -- --tailscale --service` answers the questions from the command line (`--lan`, `--local`, `--run`, `--no-start`, `--name <n>`, `--port <n>`, `--yes` too). **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently; set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), or [OMP](https://github.com/can1357/oh-my-pi) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the nine is present; if none is found, it offers to install any of them from a menu (DeepSeek excepted, since its npm package installs only a launcher with no runnable profile), or you can skip and install one yourself later. After install:
@@ -219,7 +224,7 @@ codeman web --https
# Open on your phone: https://<your-ip>:3000
```
> `localhost` works over plain HTTP. Use `--https` when accessing from another device, or use [Tailscale](https://tailscale.com/) (recommended): the installer can set it up for you (choose **Tailscale** at the network-access prompt, or run `bash ~/.codeman/app/install.sh tailscale` on an existing install). That gives you `https://<your-machine>.<tailnet>.ts.net` with a real certificate: private to your tailnet, no password required, and PWA install + push notifications work on your phone.
> `localhost` works over plain HTTP. Use `--https` when accessing from another device, or use [Tailscale](https://tailscale.com/) (recommended): the installer can set it up for you (choose **Tailscale** at the network-access prompt, or run `bash ~/.codeman/app/install.sh tailscale` on an existing install). That gives you `https://<your-machine>.<tailnet>.ts.net` with a real certificate: private to your tailnet, no password required, and PWA install + push notifications work on your phone. The installer ends on that URL with a QR code to scan, and `bash ~/.codeman/app/install.sh status` prints it again any time.
### Secure QR Code Authentication
+4
View File
@@ -23,6 +23,10 @@
<a href="https://github.com/Ark0N/Codeman/commits/master"><img src="https://img.shields.io/github/commit-activity/t/Ark0N/Codeman?style=flat-square&color=1e3a5f" alt="Total commits"></a>
</p>
<p align="center">
⭐ <strong>喜欢 Codeman?<a href="https://github.com/Ark0N/Codeman">在 GitHub 上给它点个 Star 吧!</a></strong>只需轻点一下,就能帮助更多人发现这个项目。⭐
</p>
<p align="center">
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — 并行子智能体可视化" width="900">
</p>
+1 -1
View File
@@ -1,7 +1,7 @@
[
{
"id": "claude",
"label": "Claude",
"label": "Claude Code",
"shortBadge": "CC",
"enabled": true,
"order": 0,
+4
View File
@@ -28,7 +28,11 @@ export const BROWSER_TEST_GLOBS = [
'test/terminal-copy-shortcut.test.ts',
'test/terminal-keycode229-recovery.browser.test.ts',
'test/capture-load-window.browser.test.ts',
'test/capture-geometry-retry.browser.test.ts',
'test/codex-predictive-echo.test.ts', // also needs a real codex binary
'test/split-pane-terminal.browser.test.ts',
'test/split-pane-orchestration.browser.test.ts',
'test/split-pane-auto-collapse.browser.test.ts',
];
/**
+8
View File
@@ -50,6 +50,14 @@ CODEMAN_USERNAME=admin
# README.md, "Reverse-proxy host allowlist".
# CODEMAN_ALLOWED_HOSTS=codeman.example.com,.internal.example.com
# The GitHub CLI (gh) and the Azure CLI (az, with the azure-devops extension)
# can be built into the images as git credential helpers, so Codeman can clone
# private GitHub and Azure DevOps repositories. Both are OFF by default and are
# NOT set here: turn them on in docker-compose.override.yml with the build args
# CODEMAN_INSTALL_GH / CODEMAN_INSTALL_AZ and, for the Docker-case agent image,
# the environment variables CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ. See
# README.md, "Private repositories".
# Optional: authenticate Gemini CLI without an interactive login.
GEMINI_API_KEY=
+88
View File
@@ -41,6 +41,94 @@ Releases that change `server.Dockerfile`, `docker-compose.yaml`, or add a key to
changed, and asks you to run `Start-Codeman.sh` here on the host instead. Details:
[`../docs/docker-self-update.md`](../docs/docker-self-update.md).
### Major updates
`Start-Codeman.sh` rebuilds the image on every start, but with the layer cache,
and it refreshes the build-artefact volumes selectively: `codeman-dist` when
the checkout's HEAD moved, `codeman-node-modules` only when `package-lock.json`
changed. That is right for an ordinary `git pull`. It is not enough when a
`server.Dockerfile` change bumps the Node base image without touching the
lockfile: `node-pty` is compiled from source (there is no Linux prebuild), so
the old `codeman-node-modules` volume would keep a build made for the previous
Node version. For that case, or whenever you want to be certain of what ships,
`docker/Update-Codeman.sh` force-rebuilds the image with no layer cache, stops
the stack, removes the `codeman-node-modules` and `codeman-dist` volumes, then
hands off to `Start-Codeman.sh` for the usual start:
```sh
bash docker/Update-Codeman.sh
```
Pass `--keep-volumes` to skip clearing them (safe only if you know the
rebuilt image's `node_modules`/`dist` did not change). The scripted default
is the "Resetting the build artefacts" procedure in
[`../docs/docker-self-update.md`](../docs/docker-self-update.md). Only those
two volumes are removed, by name within this Compose project; any volume a
`docker-compose.override.yml` adds is left alone, and application data and
case workspaces are host bind mounts, never touched either way.
## Private repositories (GitHub and Azure DevOps)
The images can include the GitHub CLI (`gh`) and the Azure CLI (`az`, with the `azure-devops` extension), wired into the system Git configuration as credential helpers, so Codeman can clone private repositories. Both are **opt-in and off by default**, and are turned on per host in `docker-compose.override.yml`.
### Turning them on
Add the build arguments to `docker-compose.override.yml` (see [Local customisation](#local-customisation)), then rebuild with `Start-Codeman.sh`. Set only the one you need:
```yaml
services:
codeman:
build:
args:
CODEMAN_INSTALL_GH: '1'
CODEMAN_INSTALL_AZ: '1'
environment:
# The same two switches for the Docker-case agent image Codeman builds.
CODEMAN_AGENT_IMAGE_INSTALL_GH: '1'
CODEMAN_AGENT_IMAGE_INSTALL_AZ: '1'
```
The `build: args:` pair controls the Codeman server image. The `environment:` pair controls the agent image for [Docker cases](../docs/docker-cases.md), which Codeman builds on the first Docker case; an agent image that already exists is not rebuilt by this, so run `node scripts/build-agent-image.mjs --no-cache` inside the container afterwards. The same variables work in front of that command when building it by hand. Values must be `0` or `1`; anything else stops the build with an error naming the argument.
They are not `.env` settings: turning a CLI on is a per-host choice, which is what the override file is for, and a new `.env.example` key makes the in-app updater refuse to update every existing installation until its `.env` gains the key.
The Azure CLI is the large one, about 600 MB of the roughly 670 MB the pair adds. A CLI left off leaves nothing functional behind: no apt repository, no package, no `azure-devops` extension and no credential-helper entry, so git for that host behaves exactly as it does without this feature. With both off the image is functionally unchanged; it still carries the `AZURE_EXTENSION_DIR` variable, an empty extensions directory and one small layer that copies and then removes the helper script.
### Signing in
With a CLI on, the system Git configuration routes credentials through it:
| Host | Credential helper | Sign in with |
| ----------------------------------------------------- | ----------------------------------------- | ---------------------------- |
| `https://github.com`, `https://gist.github.com` | `gh auth git-credential` | `gh auth login` |
| `https://dev.azure.com`, `https://*.visualstudio.com` | `/usr/local/bin/git-credential-azure-cli` | `az login --use-device-code` |
Codeman itself still collects no Git credentials. Sign the container in once from a **Terminal / Shell** session (Run menu). The session runs as the runtime account, so the sign-in is stored under `CODEMAN_APPDATA_PATH` (`~/.config/gh`, `~/.azure`) and survives rebuilds and container recreation:
```sh
gh auth login # GitHub.com -> HTTPS -> "Login with a web browser" (device code)
az login --use-device-code # then: az devops configure --defaults organization=https://dev.azure.com/<org>
```
After that, **Add Case → Clone Repo** accepts private `https://` URLs on those hosts, and `git clone` works from any session. Until a CLI is signed in its helper prints nothing, so a private clone fails immediately with the usual authentication error rather than waiting on a prompt.
**Multi-user mode:** every Codeman user's git runs as the same server account, so these sign-ins would otherwise be shared. Clone Repo therefore runs a **non-admin**'s clone and preflight with every git credential helper cleared (`git -c credential.helper=`): a non-admin can clone public repositories and anything their own SSH setup allows, but not a private https repository through the admin's `gh`/`az` sign-in. Admins, and single-user mode, keep the helpers. A non-admin's own agent sessions still run as that same account, and with the agent-image `gh`/`az` switches on, a non-admin's Docker case with credential seeding on also receives the server account's `gh`/`az` sign-in, the same as the Claude and Codex credentials; see `docs/security-architecture.md`, multi-user mode.
Azure DevOps is authenticated with an Entra ID access token that the helper requests from `az` for each Git operation, so nothing is written to disk beyond `az`'s own sign-in. An account that has to use a personal access token can set `AZURE_DEVOPS_EXT_PAT` for the container instead (for example under `environment:` in `docker-compose.override.yml`); the helper prefers it when present. SSH remotes are unaffected by any of this and keep using the account's own keys.
Docker cases copy these sign-ins into a case container only when the matching agent-image switch is on (`CODEMAN_AGENT_IMAGE_INSTALL_GH=1` for `~/.config/gh/hosts.yml` and `config.yml`, `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` for the sign-in files from `~/.azure`) and the case has credential seeding on. With a switch off they are never copied, even when the files exist, because a GitHub token or an Azure refresh token is usable by anything in the container. The copies are made when the container is **created**, so an existing case container never picks them up: after turning a switch on, signing in, or rebuilding the agent image, **recreate the case container** (remove it; the next session in that case creates a fresh one).
The GitHub agent skill for `gh` installs into the runtime account's home in the same session:
```sh
gh skill install cli/cli gh --scope user
gh skill update gh # after a later gh release
```
### Versions
Both CLIs, and the extension, are installed from their vendors' repositories with no version pinned, so they arrive at whatever is current when that build step runs. Docker caches the step, though: `Start-Codeman.sh` rebuilds with the cache, which keeps the versions from the first build until the Dockerfile changes at or above that step or the image is rebuilt with `--no-cache`. They are apt packages owned by root, so they cannot be upgraded from a session; `az extension update --name azure-devops` is the exception and works without a rebuild.
## Local customisation
Compose merges `docker-compose.override.yml` on top of `docker-compose.yaml`. Keep host-specific changes there rather than editing `docker-compose.yaml`, so this repository can be updated without losing them. Both `docker-compose.override.yml` and `docker-compose.override.yaml` are ignored by Git.
+255
View File
@@ -0,0 +1,255 @@
#!/usr/bin/env bash
#
# The scripted major-update path for the Docker Compose deployment.
#
# docker/README.md and docs/docker-self-update.md both point operators here for
# anything the in-app updater itself refuses to apply: a changed
# `server.Dockerfile`, a changed `docker-compose.yaml`, or a new required
# `.env.example` key. None of those can be applied by a container restarting
# itself — a restart reuses the existing image and configuration (see "The
# environment gate" in docs/docker-self-update.md) — so this script does the
# three things an in-place update cannot: force a real image rebuild with no
# layer cache, stop the stack, then hand off to Start-Codeman.sh for the same
# careful PUID/PGID, override-file and fingerprint handling every other start
# goes through.
#
# ⚠️ Build BEFORE stopping the stack, deliberately, same reasoning as
# Start-Codeman.sh's own build-then-down ordering: the build needs nothing
# stopped, so a slow --no-cache rebuild costs no downtime, and a build failure
# (a bad Dockerfile edit, a network blip pulling a base image) leaves the
# ALREADY-RUNNING stack untouched instead of stopped with nothing to bring it
# back.
#
# ⚠️ Clears the codeman-node-modules/codeman-dist named volumes by DEFAULT.
# Docker seeds a named volume from the image only while that volume is EMPTY,
# so a rebuilt image's fresh node_modules/dist otherwise sit unused behind a
# volume's old content and the container comes back up looking unchanged —
# exactly wrong for a script whose whole point is "be certain of what ships".
# Start-Codeman.sh clears codeman-dist when the checkout's HEAD moved and
# codeman-node-modules only when `package-lock.json` changed. A released
# server.Dockerfile change arrives through `git pull`, so HEAD moves and dist
# is refreshed, but a Dockerfile change that bumps the Node base image leaves
# the lockfile untouched while every native module (node-pty is compiled from
# source, there is no Linux prebuild) has to be rebuilt against the new Node
# ABI. Start-Codeman.sh would keep the old codeman-node-modules volume, and it
# never builds with --no-cache. This script clears BOTH volumes, and ONLY
# those two (targeted `docker volume rm` by Compose label, never
# `down --volumes`, which would also take any volume an override file adds).
# Pass --keep-volumes to opt out and reuse whatever is already in them.
#
# Usage: docker/Update-Codeman.sh [--keep-volumes]
# --keep-volumes Do not clear codeman-node-modules/codeman-dist. Safe to
# combine with a source change Start-Codeman.sh's own
# detection would have cleared anyway; unsafe if the reason
# you are here is a change to server.Dockerfile alone.
set -euo pipefail
script_dir=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
env_file="$script_dir/.env"
compose_file="$script_dir/docker-compose.yaml"
keep_volumes=0
for arg in "$@"; do
case "$arg" in
--keep-volumes)
keep_volumes=1
;;
--help | -h)
printf 'Usage: bash %s [--keep-volumes]\n' "$0"
exit 0
;;
*)
printf 'Error: unrecognised argument: %s\n' "$arg" >&2
printf 'Usage: bash %s [--keep-volumes]\n' "$0" >&2
exit 1
;;
esac
done
if [[ ! -f "$env_file" ]]; then
printf 'Error: Docker environment file is missing: %s\n' "$env_file" >&2
printf 'Create it from %s/.env.example before running this script.\n' "$script_dir" >&2
exit 1
fi
# Same override-file discovery as Start-Codeman.sh, and deliberately kept in
# step with it: a stack built here and started there must resolve to the exact
# same Compose files, or this script's build could target a configuration the
# handoff's own `up` never actually uses. Compose's own precedence (measured on
# v5.5.0 with both present: it uses .yml and ignores .yaml).
override_yml="$script_dir/docker-compose.override.yml"
override_yaml="$script_dir/docker-compose.override.yaml"
if [[ -f "$override_yml" && -f "$override_yaml" ]]; then
printf 'Warning: both %s and %s exist; Compose uses .yml and ignores .yaml.\n' \
"$override_yml" "$override_yaml" >&2
fi
compose_files=(-f "$compose_file")
for override_file in "$override_yml" "$override_yaml"; do
if [[ -f "$override_file" ]]; then
compose_files+=(-f "$override_file")
printf 'Using Compose override file: %s\n' "$override_file"
break
fi
done
compose_command=(docker compose --env-file "$env_file" "${compose_files[@]}")
# Collision guard. Start-Codeman.sh has no equivalent; this is the only one,
# and it has to run before this script's own --no-cache build, `down` and
# volume removal below. docker-compose.yaml hard-codes `name: codeman`, so a
# second checkout run without COMPOSE_PROJECT_NAME resolves to the SAME Compose
# project as any other checkout on the host and would operate on ITS
# containers and volumes.
#
# The project name is read from the resolved config's top-level `name` key
# (the first `name` in the output; nested ones come later), the same parse
# Start-Codeman.sh uses. `--format json` needs Compose v2.3+. This is the first
# `docker` call the script makes, so its failure is reported here rather than
# left to `set -e`, which would exit with no output at all.
if ! project_config=$("${compose_command[@]}" config --format json); then
printf 'Error: `docker compose config --format json` failed (see the message above, if any).\n' >&2
printf 'Check that Docker and Compose v2.3+ are installed and on PATH, and that\n' >&2
printf '%s and the Compose files in %s are valid.\n' "$env_file" "$script_dir" >&2
exit 1
fi
project_name=$(
printf '%s\n' "$project_config" |
sed -n 's/^[[:space:]]*"name":[[:space:]]*"\([^"]*\)".*$/\1/p' | head -n1
)
if [[ -n "$project_name" ]]; then
# `|| true` on the pipeline's LAST command: under `set -o pipefail`, `grep -v`
# exits 1 when nothing survives the filter — the ordinary, no-collision case,
# since `docker ps` finds nothing at all on a first-ever deployment or a
# single matching (own) working_dir gets filtered out. Without it, that exit
# status propagates through the command substitution and `set -e` aborts the
# WHOLE script right here, every time, regardless of whether a collision
# actually exists — caught only by actually running this end-to-end (a
# static text/regex check on the source cannot see it). The empty-line
# filter keeps a container with no working_dir label from winning head -n1
# and hiding a real collision behind it.
other_working_dir=$(
docker ps -a --filter "label=com.docker.compose.project=$project_name" \
--format '{{.Label "com.docker.compose.project.working_dir"}}' 2>/dev/null |
grep -v -F -x -- "$script_dir" | grep -v '^$' | head -n1 || true
)
if [[ -n "$other_working_dir" ]]; then
printf 'Error: Compose project "%s" is already in use by a DIFFERENT checkout:\n' "$project_name" >&2
printf ' %s\n' "$other_working_dir" >&2
printf 'This checkout is:\n' >&2
printf ' %s\n' "$script_dir" >&2
printf '\n' >&2
printf 'docker-compose.yaml hard-codes `name: %s`, so two checkouts on the same host\n' "$project_name" >&2
printf 'collide unless each one sets a distinct COMPOSE_PROJECT_NAME. Continuing would\n' >&2
printf 'rebuild and stop the OTHER checkout'"'"'s running container and, by default,\n' >&2
printf 'delete its codeman-node-modules/codeman-dist volumes.\n' >&2
printf '\n' >&2
printf 'Fix: export COMPOSE_PROJECT_NAME=<something-unique-to-this-checkout> before\n' >&2
printf 'running this script, then retry.\n' >&2
printf '\n' >&2
printf 'If instead THIS checkout was moved or renamed after its container was created,\n' >&2
printf 'the path above is its own old location: remove the old container (for example\n' >&2
printf '`docker rm -f <container>` for the codeman container) and retry, rather than\n' >&2
printf 'setting COMPOSE_PROJECT_NAME, which would start a second project beside it.\n' >&2
exit 1
fi
fi
# Same owner-detection Start-Codeman.sh uses to derive PUID/PGID for its own
# build — without it, the --no-cache build below gets Compose's untouched
# default of 1000:1000, and on any host whose appdata owner differs (99:100 on
# Unraid, per docker/README.md's chown example), Start-Codeman.sh's own
# correctly-PUID'd build during the handoff then rebuilds those layers with the
# right values anyway — so the "no cache, certain of what ships" image this
# script produces is not the one that actually ends up running.
#
# Deliberately NOT the same as Start-Codeman.sh's own handling of a MISSING
# appdata directory (which creates it): this script updates an EXISTING
# deployment, so a missing appdata path means there is nothing here yet to
# update, and creating one would just be this script quietly doing
# Start-Codeman.sh's first-run job worse.
appdata_path=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "CODEMAN_APPDATA_PATH" { sub(/^[^=]*=/, ""); print; exit }'
)
if [[ -z "$appdata_path" || ! -d "$appdata_path" ]]; then
printf 'Error: CODEMAN_APPDATA_PATH is not set or does not exist: %s\n' "${appdata_path:-<unset>}" >&2
printf 'Run docker/Start-Codeman.sh first to set up a new deployment.\n' >&2
exit 1
fi
# `stat -c` is GNU, `stat -f` is BSD/macOS; the bind source lives on the Docker
# host, so both need to work. Identical to Start-Codeman.sh's own helper.
owner_of() {
stat -c '%u:%g' -- "$1" 2>/dev/null || stat -f '%u:%g' "$1" 2>/dev/null
}
if ! owner_ids=$(owner_of "$appdata_path"); then
printf 'Error: Cannot determine the owner of CODEMAN_APPDATA_PATH: %s\n' "$appdata_path" >&2
exit 1
fi
export PUID=${owner_ids%%:*}
export PGID=${owner_ids##*:}
if [[ "$PUID" == '0' ]]; then
printf 'Error: CODEMAN_APPDATA_PATH is owned by root: %s\n' "$appdata_path" >&2
printf 'Change the directory ownership to the unprivileged account that should run Codeman.\n' >&2
exit 1
fi
# --no-cache, always: a plain `build` reuses cached layers (npm install, apt
# packages, the CLI installs baked into the image) and can silently keep them
# frozen at whatever they were the day the cache was populated — exactly wrong
# for a major update, whose whole point is being certain of what actually
# ships. `scripts/build-agent-image.mjs` makes the same call for the same
# reason (see its entry in CLAUDE.md's Additional Commands table). Runs BEFORE
# the stack is stopped — see the header comment for why.
printf 'Building a fresh image (--no-cache)...\n'
"${compose_command[@]}" build --no-cache
printf 'Stopping the stack...\n'
if [[ "$keep_volumes" == '1' || -n "$project_name" ]]; then
"${compose_command[@]}" down
else
# No resolvable project name means the label filter below could match
# nothing, so fall back to Compose's own removal, and say what it really does.
printf 'Warning: could not resolve the Compose project name; clearing EVERY named volume\n' >&2
printf 'in this Compose project (override file included) with `down --volumes` instead.\n' >&2
"${compose_command[@]}" down --volumes
fi
# Targeted removal of exactly the two build-artefact volumes, scoped by label to
# THIS project (the volume key alone is shared by any other stack declaring the
# same key). Same lookup as Start-Codeman.sh's refresh. A failure is reported,
# not fatal: the stack is already down, and the handoff below is what brings
# it back up.
if [[ "$keep_volumes" != '1' && -n "$project_name" ]]; then
printf 'Clearing the codeman-node-modules/codeman-dist volumes (pass --keep-volumes to skip).\n'
for key in codeman-node-modules codeman-dist; do
volume_name=$(
docker volume ls -q \
--filter "label=com.docker.compose.volume=$key" \
--filter "label=com.docker.compose.project=$project_name" |
head -n1
) || volume_name=''
if [[ -n "$volume_name" ]] && ! docker volume rm -- "$volume_name"; then
printf 'Warning: could not remove volume %s; the container may keep serving the\n' "$volume_name" >&2
printf 'previous build from it. Remove it by hand and rerun this script.\n' >&2
fi
done
fi
# Start-Codeman.sh does everything a plain `up -d` does not: re-derives
# PUID/PGID, pre-creates CODEMAN_CASES_PATH with the right ownership, resolves
# DOCKER_SOCKET_GID, records the server.Dockerfile/docker-compose.yaml
# fingerprint the in-app updater's gate reads on every future update, and
# starts the (already freshly built) image. Reimplementing any of that here
# would only risk drifting out of step with it — hand off instead, exactly as
# docs/docker-self-update.md's own reset procedure does.
#
# ⚠️ `bash`, not a bare exec of the path: Start-Codeman.sh is committed
# non-executable (100644), the same as this script, and is documented
# everywhere as `bash docker/Start-Codeman.sh` rather than
# `./docker/Start-Codeman.sh` — execing the bare path fails with EACCES.
printf 'Handing off to Start-Codeman.sh...\n'
exec bash "$script_dir/Start-Codeman.sh"
+87
View File
@@ -17,6 +17,7 @@ FROM node:22-bookworm-slim
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
git \
libsecret-1-0 \
tmux \
ripgrep \
curl \
@@ -26,6 +27,88 @@ RUN apt-get update \
openssh-client \
&& rm -rf /var/lib/apt/lists/*
# GitHub CLI and Azure CLI (+ the azure-devops extension) with the same system
# git credential helpers as docker/server.Dockerfile, so an agent in a Docker
# case can clone and push to private GitHub / Azure DevOps repositories. The
# sign-ins themselves are NOT baked in: `~/.config/gh` and `~/.azure` are seeded
# per container at launch like every other CLI's credentials (CRED_STORES in
# src/docker-hosts.ts), and a helper whose CLI is not signed in prints nothing,
# so git fails fast instead of prompting. See server.Dockerfile for why the
# vendor apt repositories are configured here rather than via deb_install.sh.
#
# Each is OPT-IN and OFF by default, like the server image: CODEMAN_INSTALL_GH=1
# / CODEMAN_INSTALL_AZ=1 turn one on; off leaves no repository, package,
# extension or helper entry. scripts/build-agent-image.mjs and the in-app
# auto-build pass them from CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ in their own
# environment (for the Compose deployment: `environment:` in
# docker-compose.override.yml), and pass nothing when those are unset, so
# these defaults (off) apply.
ARG CODEMAN_INSTALL_GH=0
ARG CODEMAN_INSTALL_AZ=0
RUN set -eux; \
for flag in "CODEMAN_INSTALL_GH=${CODEMAN_INSTALL_GH}" "CODEMAN_INSTALL_AZ=${CODEMAN_INSTALL_AZ}"; do \
case "${flag#*=}" in 0|1) ;; *) echo "${flag%%=*} must be 0 or 1, got '${flag#*=}'" >&2; exit 1;; esac; \
done; \
codename="$(. /etc/os-release && echo "${VERSION_CODENAME}")"; \
arch="$(dpkg --print-architecture)"; \
pkgs=""; \
install -d -m 0755 /etc/apt/keyrings; \
if [ "${CODEMAN_INSTALL_GH}" = 1 ]; then \
curl -fsSL -o /etc/apt/keyrings/githubcli-archive-keyring.gpg \
https://cli.github.com/packages/githubcli-archive-keyring.gpg; \
chmod go+r /etc/apt/keyrings/githubcli-archive-keyring.gpg; \
echo "deb [arch=${arch} signed-by=/etc/apt/keyrings/githubcli-archive-keyring.gpg] https://cli.github.com/packages stable main" \
> /etc/apt/sources.list.d/github-cli.list; \
pkgs="${pkgs} gh"; \
fi; \
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
curl -fsSL -o /etc/apt/keyrings/microsoft.asc \
https://packages.microsoft.com/keys/microsoft.asc; \
chmod go+r /etc/apt/keyrings/microsoft.asc; \
echo "deb [arch=${arch} signed-by=/etc/apt/keyrings/microsoft.asc] https://packages.microsoft.com/repos/azure-cli/ ${codename} main" \
> /etc/apt/sources.list.d/azure-cli.list; \
pkgs="${pkgs} azure-cli"; \
fi; \
if [ -n "${pkgs}" ]; then \
apt-get update; \
apt-get install -y --no-install-recommends ${pkgs}; \
rm -rf /var/lib/apt/lists/*; \
fi; \
if [ "${CODEMAN_INSTALL_GH}" = 1 ]; then gh --version; fi; \
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then az version --output none; fi
# Outside HOME so the seeded `~/.azure` (auth files only) never has to carry
# extensions. gid 0 + group-writable, the same arbitrary-uid convention as HOME
# below, so `az extension update` works as whatever uid the container runs as.
# Created even without az; an empty directory costs nothing.
ENV AZURE_EXTENSION_DIR=/opt/az-extensions
RUN set -eux; \
install -d -m 0755 "${AZURE_EXTENSION_DIR}"; \
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
az extension add --name azure-devops --only-show-errors; \
rm -rf /root/.azure; \
fi; \
chgrp -R 0 "${AZURE_EXTENSION_DIR}"; \
chmod -R g=u "${AZURE_EXTENSION_DIR}"
# Only an installed CLI gets a helper entry (see server.Dockerfile).
COPY docker/git-credential-azure-cli /usr/local/bin/git-credential-azure-cli
RUN set -eux; \
if [ "${CODEMAN_INSTALL_GH}" = 1 ]; then \
for host in https://github.com https://gist.github.com; do \
git config --system "credential.${host}.helper" '!/usr/bin/gh auth git-credential'; \
done; \
fi; \
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
chmod 0755 /usr/local/bin/git-credential-azure-cli; \
for host in https://dev.azure.com 'https://*.visualstudio.com'; do \
git config --system "credential.${host}.helper" /usr/local/bin/git-credential-azure-cli; \
git config --system "credential.${host}.useHttpPath" true; \
done; \
else \
rm -f /usr/local/bin/git-credential-azure-cli; \
fi
# The npm-published agent CLIs, supplied by scripts/build-agent-image.mjs from
# config/clis.stock.json so a new stock CLI needs no edit here. The default is
# today's literal list, so a bare `docker build` still produces the same image.
@@ -44,6 +127,10 @@ RUN apt-get update \
# A different order is a different RUN string, which is a different layer hash and
# so a needless cache miss between a bare `docker build` and a scripted one.
ARG CLI_NPM_PACKAGES="@anthropic-ai/claude-code opencode-ai @openai/codex @google/gemini-cli"
# uv/uvx: MCP servers are commonly launched with `uvx <package>` (e.g. the Nginx
# Proxy Manager MCP), and Codex failed to enable them with "uvx not found". Copied
# from the pinned upstream image into root-owned /usr/local/bin, never pip-installed.
COPY --from=ghcr.io/astral-sh/uv:0.9 /uv /uvx /usr/local/bin/
RUN npm install -g ${CLI_NPM_PACKAGES} \
&& npm cache clean --force
+34
View File
@@ -0,0 +1,34 @@
#!/bin/sh
# Git credential helper for Azure DevOps, backed by the signed-in Azure CLI.
#
# Configured in the image's system gitconfig for https://dev.azure.com and
# https://*.visualstudio.com (see server.Dockerfile). On `get` it answers with
# an Entra ID access token for the Azure DevOps resource as the password, the
# same token type Git Credential Manager uses for Azure Repos. It never prompts:
# when `az` is not signed in it prints nothing, so git fails fast with its own
# authentication error instead of hanging a request that has no terminal.
#
# AZURE_DEVOPS_EXT_PAT, the azure-devops extension's own PAT variable, is used
# instead when it is set, for accounts that authenticate with a PAT.
# `store` and `erase` are no-ops: the token belongs to az, which refreshes it.
[ "$1" = "get" ] || exit 0
# Drain the request git writes on stdin; the host scoping is in gitconfig.
cat >/dev/null
if [ -n "${AZURE_DEVOPS_EXT_PAT:-}" ]; then
printf 'username=pat\npassword=%s\n' "$AZURE_DEVOPS_EXT_PAT"
exit 0
fi
command -v az >/dev/null 2>&1 || exit 0
# 499b84ac-1321-427f-aa17-267ca6975798 is the fixed application ID of Azure
# DevOps: https://learn.microsoft.com/azure/devops/integrate/get-started/authentication/service-principal-managed-identity
token="$(az account get-access-token \
--resource 499b84ac-1321-427f-aa17-267ca6975798 \
--query accessToken --output tsv 2>/dev/null)" || exit 0
[ -n "$token" ] || exit 0
printf 'username=azure-cli\npassword=%s\n' "$token"
+119 -1
View File
@@ -39,6 +39,7 @@ RUN apt-get update \
curl \
g++ \
git \
libsecret-1-0 \
make \
openssh-client \
procps \
@@ -68,6 +69,111 @@ COPY --from=docker:29-cli \
/usr/local/libexec/docker/cli-plugins/docker-buildx \
/usr/local/libexec/docker/cli-plugins/docker-buildx
# GitHub CLI and Azure CLI (with the azure-devops extension), so a user can sign
# this container in to GitHub and Azure DevOps from a Codeman shell session and
# then clone PRIVATE repositories, both from that session and through Add Case
# -> Clone Repo. Codeman still collects no Git credentials itself: the clone
# path (src/git-clone.ts) only inherits HOME and git's config, so whatever the
# user signs in to here is what authenticates, and nothing when they have not
# (the clone then fails fast with AUTH_REQUIRED, exactly as before).
#
# Each is OPT-IN and OFF by default: the image is functionally unchanged
# unless the build gets CODEMAN_INSTALL_GH=1 and/or CODEMAN_INSTALL_AZ=1, which
# a deployment sets under `build: args:` in docker-compose.override.yml
# (docker/README.md, "Private repositories"). Off installs no apt repository,
# package, extension or credential-helper entry; all that remains is the
# AZURE_EXTENSION_DIR variable, its empty directory and one layer that copies
# and then removes the helper script. The Azure CLI is the heavy one (~600 MB,
# mostly its bundled Python). The base docker-compose.yaml
# and .env deliberately do not carry them: turning a CLI on is a per-host
# choice, which is what the override file is for, and a new .env.example key
# would make the self-updater refuse existing installs until their .env gained
# it (docs/docker-self-update.md).
#
# Both come from their vendors' own apt repositories, the same ones the
# documented one-liners configure (https://github.com/cli/cli/blob/trunk/docs/install_linux.md
# and https://learn.microsoft.com/cli/azure/install-azure-cli-linux?pivots=apt).
# Microsoft's `deb_install.sh` is deliberately not piped into the build: it does
# exactly this plus a `gnupg` install, and a remote script run at build time is
# the one step a reviewer cannot read in this file. apt reads an ASCII-armoured
# `.asc` key directly, which is what keeps `gnupg` out of the image.
#
# Not pinned, unlike the agent CLIs below: nothing in Codeman depends on a
# particular gh or az behaviour, so the pinning argument there does not apply.
# The layer cache still keeps whatever version the first build fetched until a
# --no-cache rebuild.
ARG CODEMAN_INSTALL_GH=0
ARG CODEMAN_INSTALL_AZ=0
RUN set -eux; \
for flag in "CODEMAN_INSTALL_GH=${CODEMAN_INSTALL_GH}" "CODEMAN_INSTALL_AZ=${CODEMAN_INSTALL_AZ}"; do \
case "${flag#*=}" in 0|1) ;; *) echo "${flag%%=*} must be 0 or 1, got '${flag#*=}'" >&2; exit 1;; esac; \
done; \
codename="$(. /etc/os-release && echo "${VERSION_CODENAME}")"; \
arch="$(dpkg --print-architecture)"; \
pkgs=""; \
install -d -m 0755 /etc/apt/keyrings; \
if [ "${CODEMAN_INSTALL_GH}" = 1 ]; then \
curl -fsSL -o /etc/apt/keyrings/githubcli-archive-keyring.gpg \
https://cli.github.com/packages/githubcli-archive-keyring.gpg; \
chmod go+r /etc/apt/keyrings/githubcli-archive-keyring.gpg; \
echo "deb [arch=${arch} signed-by=/etc/apt/keyrings/githubcli-archive-keyring.gpg] https://cli.github.com/packages stable main" \
> /etc/apt/sources.list.d/github-cli.list; \
pkgs="${pkgs} gh"; \
fi; \
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
curl -fsSL -o /etc/apt/keyrings/microsoft.asc \
https://packages.microsoft.com/keys/microsoft.asc; \
chmod go+r /etc/apt/keyrings/microsoft.asc; \
echo "deb [arch=${arch} signed-by=/etc/apt/keyrings/microsoft.asc] https://packages.microsoft.com/repos/azure-cli/ ${codename} main" \
> /etc/apt/sources.list.d/azure-cli.list; \
pkgs="${pkgs} azure-cli"; \
fi; \
if [ -n "${pkgs}" ]; then \
apt-get update; \
apt-get install -y --no-install-recommends ${pkgs}; \
rm -rf /var/lib/apt/lists/*; \
fi
# The azure-devops extension goes into a SYSTEM directory rather than the
# default ~/.azure/cliextensions: HOME is the application-data bind mount, which
# hides anything installed there at build time. The directory is handed to the
# runtime account below (next to /opt/codeman-cli) so `az extension update`
# works from a session. Nothing that runs as root executes from it. It is
# created even without az, so the chown below does not have to know.
ENV AZURE_EXTENSION_DIR=/opt/codeman-az-extensions
RUN set -eux; \
install -d -m 0755 "${AZURE_EXTENSION_DIR}"; \
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
az extension add --name azure-devops --only-show-errors; \
rm -rf /root/.azure; \
fi
# Git credential helpers, in the SYSTEM gitconfig so they apply to every
# account and survive a fresh application-data directory. Each one answers only
# for its own host and prints nothing when its CLI is not signed in, so git
# falls through to its normal non-interactive failure. Only an installed CLI
# gets an entry: a helper naming a missing binary would print an error on every
# clone from that host.
# github.com `gh auth git-credential`, what `gh auth setup-git` configures.
# Azure DevOps an Entra ID token from `az login` (git-credential-azure-cli),
# for both dev.azure.com and the legacy *.visualstudio.com hosts.
COPY docker/git-credential-azure-cli /usr/local/bin/git-credential-azure-cli
RUN set -eux; \
if [ "${CODEMAN_INSTALL_GH}" = 1 ]; then \
for host in https://github.com https://gist.github.com; do \
git config --system "credential.${host}.helper" '!/usr/bin/gh auth git-credential'; \
done; \
fi; \
if [ "${CODEMAN_INSTALL_AZ}" = 1 ]; then \
chmod 0755 /usr/local/bin/git-credential-azure-cli; \
for host in https://dev.azure.com 'https://*.visualstudio.com'; do \
git config --system "credential.${host}.helper" /usr/local/bin/git-credential-azure-cli; \
git config --system "credential.${host}.useHttpPath" true; \
done; \
else \
rm -f /usr/local/bin/git-credential-azure-cli; \
fi
# Keep credentials out of the image. Users authenticate these CLIs at runtime
# through Codeman sessions, and the configured host bind mount retains state.
#
@@ -107,13 +213,25 @@ COPY --from=docker:29-cli \
# minimal image of this exact shape). The four CLIs live only in this prefix,
# so they still resolve; entrypoint.sh additionally pins its own PATH to the
# system directories for the root part of the start.
# uv/uvx: MCP servers are commonly launched with `uvx <package>` (e.g. the Nginx
# Proxy Manager MCP), and Codex failed to enable them with "uvx not found". Copied
# from the pinned upstream image into root-owned /usr/local/bin, never pip-installed.
COPY --from=ghcr.io/astral-sh/uv:0.9 /uv /uvx /usr/local/bin/
ENV NPM_CONFIG_PREFIX=/opt/codeman-cli
ENV PATH=$PATH:/opt/codeman-cli/bin
# pnpm is not an agent CLI: it is here because `dsh plugin` (DeepSeek Harness, which
# this image leaves to be installed at runtime, see SERVER_INTENTIONAL_OMISSIONS in
# test/docker-agent-image-coverage.test.ts) spawns a literal `pnpm` with no npm
# fallback, so the Run menu's "DeepSeek - add a terminal profile" button failed
# with `dsh: pnpm not found on PATH` (exit 127) on this image. The agent image
# already carries it for the same reason (#352). It lives in the same
# runtime-writable prefix as the CLIs, so a session can update it in place.
RUN npm install --global \
@anthropic-ai/claude-code@2.1.258 \
@google/gemini-cli@0.58.0 \
@openai/codex@0.152.1 \
opencode-ai@1.18.26 \
pnpm@12.6.0 \
&& npm cache clean --force
# Keep the web server and every local Codeman session unprivileged. PUID and
@@ -153,7 +271,7 @@ RUN set -eux; \
--shell /bin/bash \
"${CODEMAN_RUNTIME_USER}"; \
fi; \
chown -R "${PUID}:${PGID}" /opt/codeman-cli
chown -R "${PUID}:${PGID}" /opt/codeman-cli /opt/codeman-az-extensions
WORKDIR /opt/codeman
+10
View File
@@ -757,3 +757,13 @@ works, and its replies arrive tagged `from-name="w9-msgtest"` (a derived-name
worker's replies carry no `from-name`). A quick-start without `sessionName` has an
empty Codeman name, so the peer name stays derived: agents should name their
workers. Tests: `test/name-flag-injection.test.ts`.
Later narrowing: `--name` is not only the peer name but also the `/resume` picker
entry and the terminal title, and a pinned title stops Claude generating its own, so
pinning the `w1-myapp` placeholder listed every conversation of a case under the same
name in `/resume`. Only a manual name is pinned now (`Session.cliPinnedName`,
`nameSource === 'manual'`, carried to the builders as `cliName`); placeholder and auto
names leave Claude to title the conversation. A rename in Codeman appends a
`custom-title` row to the conversation's transcript (`claude-session-title.ts`), the
row `/rename` writes. Tests: `test/claude-resume-title.test.ts`,
`test/routes/session-name-routes.test.ts`.
+146
View File
@@ -324,6 +324,30 @@ from the session's current state rather than requiring a new transition: the
original turn may be long over. It comes back as
`"delivered": false, "duplicate": true`.
**Wake-on-LAN hosts** (`docs/remote-sessions.md` §Wake-on-LAN): when the session's
remote host has a wake target and is asleep, the non-wait form answers `200` with
`{"buffered": true}` — the bytes are held and flushed after the host is back — or
`{"buffered": true, "dropped": true}` for a chunk over the 4 KB wake buffer, which
is gone (never delivered as a fragment). Both fields are additive to the historical
bare `{}`. With `wait`, the route blocks on the wake instead and answers
`422 OPERATION_FAILED` ("did not come back after a wake-on-LAN request — nothing was
sent") when the host never returns, rather than writing into the stalled pane and
reporting `delivered:true` plus a timeout.
Two endpoints back that flow directly, both scoped to one session's remote host and
both refusing a session that is not remote (`400 INVALID_INPUT`):
| Method | Path | Purpose |
| --- | --- | --- |
| `GET` | `/api/sessions/:id/reachability` | Whether the session's remote host answers SSH right now, plus whether a wake target is configured. Read-only: it never wakes. `{"reachable": true\|false\|null, "wakeConfigured": "mac"\|"command"\|"none"}`, where `null` means the answer is unknown (a proxied host, where a TCP probe proves nothing). |
| `POST` | `/api/sessions/:id/wake` | Wake the host and wait for it to accept SSH again, bounded by the request budget. `422 OPERATION_FAILED` when it does not come back; `400 INVALID_INPUT` with "No wake-on-LAN target configured for this host" when nothing is set. |
⚠️ Waking is deliberately reachable only from an explicit user action (this route, a
session create/attach, or typing into a sleeping session). No watcher, dropped-session
handler or boot-recovery path may wake a host, or a suspended machine would be woken
again seconds after every suspend; `test/remote-wake.test.ts` pins that as an import
fence around `src/remote-wake.ts`.
### Response
All three nest the wait result under `data.wait`, so one client helper works against
@@ -558,6 +582,128 @@ All four enforce session ownership in multi-user mode; a foreign session id
answers `404 NOT_FOUND` (no existence leak), and profiles of two owners of the
same directory are distinct by construction.
## Custom Model Endpoints
Points a session's harness at a user-configured OpenAI-compatible endpoint —
local (llama.cpp, vLLM, DGX Spark) or cloud (Azure AI Foundry, OpenRouter) —
instead of its native cloud backend, gated by the opt-in
`customModelEndpointsEnabled` setting (default OFF). Endpoints are
machine-level infra, like remote/docker hosts: writes are admin-only in
multi-user mode. Design: [`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md);
user guide: [`custom-model-endpoints.md`](custom-model-endpoints.md).
- `GET /api/v1/model-endpoints` -> `CustomModelHost[]`, an unwrapped bare
array like every other list route (still riding the standard `{success,
data}` envelope on the wire — unwrap it the same way). Answers `[]` for a
non-admin in multi-user mode. `apiKey` is never returned; `apiKeySet:
boolean` reports whether one is stored, so a client can render "unchanged
if left blank" without ever holding the real value.
- `POST /api/v1/model-endpoints` with `{ id, label, baseUrl, apiKey?,
authStyle?, defaultModelId? }` creates one. `id` must match
`^[a-zA-Z0-9_-]+$`; `authStyle` is `bearer` (default) or `api-key`, never
both (a real server hung indefinitely when sent both headers on one
request); `baseUrl` must be `http(s)`, carry no embedded credentials, and
is refused if it points at (or resolves to) a link-local or
cloud-metadata address. `409 ALREADY_EXISTS` on a duplicate id.
- `PUT /api/v1/model-endpoints/:id` updates one. An **absent** `apiKey`
keeps the stored one rather than clearing it — the client never receives
the real value to resend deliberately unchanged, so omission is the only
way to say "leave it alone"; there is no way to clear a key back to unset
this way. `defaultModelId`, when set, must be one of that endpoint's own
`models` (`400 INVALID_INPUT` otherwise).
- `DELETE /api/v1/model-endpoints/:id` removes one.
- `POST /api/v1/model-endpoints/:id/discover-models` fetches the endpoint's
own `GET /v1/models` and stores the result as `models`, updating
`lastDiscoveredAt`, plus (best-effort, only for a model llama-swap's own
response already reports loaded) `modelContextLengths` and `modelSizesGB`.
A `defaultModelId` that no longer appears in the fresh list is dropped
rather than carried forward invalid. Failures answer `422 OPERATION_FAILED`
with the underlying connection error, or a named egress refusal if the
resolved address turned out to be blocked. The same refresh also runs
automatically for every saved endpoint every 5 minutes in the background
(`refreshAllCustomModelHosts()`, `custom-model-routes.ts`, started from
`server.ts`), so there is no route for triggering "refresh all" — one
endpoint being unreachable on a cycle never blocks the others.
- `GET /api/v1/model-endpoints/:id/running-status` -> `{ isLlamaSwap,
running: [{model, state}], logLine? }`, read-only, no admin gate
(any session owner who could already point a session at this endpoint can
equally ask what it currently has loaded). `isLlamaSwap` is
feature-detected via the endpoint's own `GET /running` — a plain
llama.cpp/OpenAI-compatible server has none and always answers `false`.
`logLine`, present only when `isLlamaSwap` is true, is the most recent
REAL backend `llama-server` process log line (`load_model: ...`,
`llama_server: model loaded`, etc.), sourced from the endpoint's own
`GET /api/events` SSE stream and filtered to `source: "upstream"` frames
only (never llama-swap's own `source: "proxy"` request-access log) — one
connection is held open per endpoint and reused across every poller,
idle-closed after 30s of nobody asking. This is what the Run-menu
picker's loading banner polls once a second while a model is loading.
- `POST /api/v1/sessions/:id/custom-model` with `{ endpointId, modelId,
confirmed? } | { clear: true }` applies (or clears) the session's
selection and **restarts the session's CLI process in place** — every
supported harness reads its endpoint config at process start, never per
turn, so there is no live hot-swap. (`POST /api/v1/quick-start`'s own
`customModel: { endpointId, modelId, confirmed? }` field is the
no-restart equivalent for a session that doesn't exist yet — see below.)
A Claude session resumes its existing conversation across the restart;
pi/omp/grok additionally get a forced `--model`/`-m` value, since for
those three the config file alone does not select it. `400 INVALID_INPUT`
for a remote (SSH) or Docker session — both restart their agent
differently under the hood, and applying to one would report success
while changing nothing. Two more responses replace the normal
`{customModel, restarted}` shape, neither an error, and neither restarts
or creates anything on the first ask. ⚠️ **Each is answered by its OWN
flag on the retry, and answering one is not consent to the other**: they
are questions about different people, and while they shared a single flag
a caller who confirmed the context warning silently agreed to evict
another session's model as well. Send `confirmedContext: true` to proceed
past the context warning, `confirmedSwap: true` past the swap conflict,
and both when both were asked (they accumulate, so the second retry still
carries the first answer). The original `confirmed: true` still means
BOTH and is still accepted, because it shipped in this feature's
HTTP-API-only cut; new callers should send the specific one:
- `{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}` —
llama.cpp/llama-swap only runs one model at a time, and switching would
unload a model another **live session's own selection** is actively
using. Never returned for a plain (non-llama-swap) server, and never
just because a swap is needed at all — only when it would disrupt
someone else.
- `{requiresContextWarning: true, modelId, contextLength,
minSafeContextTokens}` — Claude Code's own fixed per-turn overhead
(system prompt + tool schemas) can exceed a small model's entire
discovered context on its own, before any conversation history exists
to compact, guaranteeing the very first message fails regardless of
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`. Gated on the CLI registry declaring a
`contextLengthVar` (claude only today), so it never fires for another
harness.
- `POST /api/v1/quick-start`'s `customModel: { endpointId, modelId,
confirmed?, confirmedContext?, confirmedSwap? }` field (alongside its
normal `caseName`/`mode`/etc. body)
computes the same injection **before** the session exists and launches
directly on the endpoint — no restart, because there was never a
native-backend boot to restart away from. Runs the identical checks as
the dedicated route above (`requiresConfirmation`/`requiresContextWarning`,
same shapes, same per-question `confirmedContext`/`confirmedSwap` retry),
and is refused the same way
for a remote or Docker case. This is what the Run-menu picker uses for
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP; Claude still uses the
dedicated restart route above (its `--resume`-based restart is far less
jarring than a full relaunch, and folding it into the one-shot path is
separate work — see `docs/custom-model-endpoints-plan.md`).
## CLI management
Read and write the CLI registry (`docs/cli-registry.md`). Every **write** route answers `403 FORBIDDEN` while `cliManagementEnabled` is off (the default), and for a non-admin in multi-user mode. A write that would overwrite a `clis.json` which does not parse, or which has group/world permission bits, is refused with `409 CONFLICT` and a message naming the fix; the file is left untouched.
| Method | Path | Body | Notes |
| -------- | ----------------------------- | ------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
| `GET` | `/api/clis` | none | Every entry, disabled ones included: `id`, `label`, `shortBadge`, `order`, `kind`, `enabled`, `stock`, `installed`, and `installCommand` for a stock entry. Not gated; a non-admin in multi-user mode gets `[]`. |
| `PUT` | `/api/clis/:id` | `{ enabled }` | Toggle an existing entry, stock or custom. `404` for an unknown id; `400 INVALID_INPUT` when disabling a `kind: 'shell'` entry. |
| `POST` | `/api/clis/:id/install` | none | Run a **stock** entry's install command (never a custom one: `400`). `409 CONFLICT` while an install for the same id is running; `422 OPERATION_FAILED` with the output tail when it fails. Never enables the entry. |
| `POST` | `/api/clis` | `{ id, label, shortBadge, binaries, argv, enabled? }` | Create a custom entry. `409 ALREADY_EXISTS` for a stock id or an existing custom id. `enabled` defaults to `true`. |
| `PUT` | `/api/clis/custom/:id` | `{ label, shortBadge, binaries, argv, enabled? }` | Replace an existing custom entry. An absent `enabled` keeps the entry's current state. `400` for a stock id, `404` for an unknown one. |
| `DELETE` | `/api/clis/:id` | none | Delete a custom entry. `400` for a stock id, `404` for an unknown one. |
## Voice dictation
Browser dictation transcribed through this server's Claude Code login, i.e. the
File diff suppressed because one or more lines are too long
+314
View File
@@ -0,0 +1,314 @@
# CLI management Settings UI + write API — plan
> Tracked separately from `DEPLOYMENT_PLAN.md` (PR B2, merged) and `docs/copilot-integration-plan.md`
> (parked). This is "PR C" from the original #343 review: *"settings UI + write endpoints +
> auto-install, once we've settled the trust model... I want to make that call on its own, not
> inside a 100-file diff."*
>
> **Phase 0 is CLOSED as of 2026-09-21** — all three original pieces are IN SCOPE (expanded from
> this plan's first draft, which recommended #2/#3 as separate/out-of-scope; the user chose full
> scope instead, with the risk called out explicitly for #3 before confirming). See "Decisions"
> below for the full record.
## Status as of 2026-09-22
**Phases 1–6 are ALL IMPLEMENTED** (commits `da07b38c` "add cliManagementEnabled flag and GET
/api/clis" and `db4557d9` "Phases 3-6 - write API + custom entries + Settings UI", both on this
branch, `feat/cli-management`). Confirmed present in the tree: `cliManagementEnabled` in
`SettingsUpdateSchema`; `GET /api/clis`, `PUT /api/clis/:id`, `POST /api/clis/:id/install`,
`POST /api/clis`, `PUT /api/clis/custom/:id`, `DELETE /api/clis/:id` in
`src/web/routes/cli-registry-routes.ts`; the `shell`/`claude` `UNDISABLEABLE_IDS` backend guard;
`isAdmin(req)` gating on both the list and write routes; `appendAdminAudit` wired into the install
route; tmp+rename+`0o600` writes in `registry-writer.ts`; the full Settings UI (row list, toggle,
Install button, custom-entry create/edit/delete form) in `settings-ui.js` + `index.html`.
`test/routes/cli-registry-routes.test.ts` (425 lines) and `test/cli-registry-no-id-branching.test.ts`
cover it. This status section, plus the fix and gap below, is the one piece of that work done in
a *different* session from the one that wrote Phases 1–6 — reviewed by reading the diff and
verifying each claim against the actual routes/tests, not by re-implementing anything.
### Gotcha found and fixed (commit `0c77dd0a`)
**Toggling a CLI off in Settings had no effect anywhere except the Settings row itself.**
`window.__codemanCliAvailable` — the flag `isCliAvailable()` reads client-side to gate the
welcome-screen buttons, the Run-menu dropdown and the mobile overview — is injected **once**, at
initial page render (`server.ts`), built purely from each CLI's own installed-on-PATH resolver
(`isClaudeAvailable()` etc.), with **no reference to the registry's `enabled` flag at all**. So
disabling a CLI here updated its own row and nothing else — every launch surface kept offering it,
both live and after a full page reload, since even a *fresh* render never consulted the registry.
Root-caused and reported by the user testing the live feature ("toggle those off, they still
appear in that menu and on the front main screen").
Fixed two places:
- `server.ts`: after building `available`, intersect the nine real `SessionMode` ids against
`enabledClis()`. `git`/`cloudflared` (utility binaries, not CLI registry entries) and
`deepseekBinary` (a secondary installed-only flag for the "add a profile" affordance) are
deliberately left alone — they were never registry-gated to begin with.
- `settings-ui.js`: `toggleCliEnabled()` now patches `window.__codemanCliAvailable` in place and
refreshes the welcome screen, the mobile overview and an already-open Run menu, mirroring the
existing `installDeepSeekProfile()` pattern for the same "injected once, needs an explicit
patch" reason — the server-side fix alone still left every surface stale until the next reload.
New test in `test/render-index-html.test.ts`: an installed-but-disabled CLI (codex, forced via
`clis.json` + `reloadCliRegistry()`) reads as unavailable, while an installed-and-enabled one
(claude) is unaffected by the override.
**Verified on the Debian devbox** (`codeman-devbox`, real tmux — this sandbox has none and
`WebServer`'s constructor hard-requires it): typecheck clean, the new test passes (17/17 in
`render-index-html.test.ts`), the CLI-registry suites pass (86/86), and the **full CI gate is
green — 415 test files, 7855 tests, 0 failures**.
### Launch-surface registry integration — completed
The welcome screen, desktop Run menu and mobile Run picker now use the same injected CLI catalog.
Every enabled registry entry is rendered; unavailable binaries remain hidden as before. Settings
updates the catalog and availability flags in place after enable/disable, create, edit or delete,
so the launch surfaces update without a page reload. A custom entry uses the generic quick-start
path, while stock entries retain their existing per-CLI launch settings.
Not otherwise re-verified line-by-line against every Phase 1–6 checklist item below (e.g. the
exact wording of toasts, the "same PR" sequencing notes) — the checklists are left as originally
written; treat the **Status** section above as authoritative for what exists.
---
## Background
`src/config/cli-registry/registry.ts` is READ-ONLY today, and says so in its own header comment:
> "⚠️ READ-ONLY. Nothing in this module writes, creates or migrates the file... there is no
> settings UI and no write API yet... A `seededStockIds` ratchet belongs with the write API that
> needs it."
Confirmed on `master` (2026-09-21): no `/api/clis` route exists at all (read or write);
`~/.codeman/clis.json` is hand-edit-only; `resolveInstallCommandForPlatform()` is documented
"Display text only — never executed" — nothing runs an install command server-side today. The
original #343 review flagged the opposite (`spawn(command, {shell: true})`, `env.allowedPrefixes`
contributed from a write) as needing its own trust-model decision; that decision was never made
after the split, just dropped. This plan makes it.
**Closest existing precedent, and the template this plan follows for the read/write API**:
`src/web/routes/custom-model-routes.ts` + `src/custom-model-hosts.ts` (#393/#430/#459) — a small
per-item JSON store, Settings-UI-driven, admin-gated in multi-user mode, tmp+rename+0600 writes.
**Precedent for the new master feature flag (Phase 1)**: `customModelEndpointsEnabled` —
`z.boolean().optional()` in `SettingsUpdateSchema` (`schemas.ts:1319`), a checkbox read/written by
id in `openAppSettings()`/`saveAppSettings()` (`settings-ui.js:401`/`:2120`). SYNCED, not
per-device (present in the schema, absent from `displayKeys`), default OFF.
**Spec refs for the whole plan:**
- `src/config/cli-registry/registry.ts` — the read path; `resolveRegistry()`'s merge semantics
(`deepMerge`, `UNMERGEABLE_KEYS`) apply unchanged to whatever this plan writes
- `docs/cli-registry.md` — registry shape, "The override file", "Arg-template safety" (the four
layers Phase 5's custom-entry validation must not weaken), "Adding a CLI" (the 5-step recipe a
custom entry does NOT get to skip just because it arrives via UI instead of a stock.ts edit)
- `src/web/routes/custom-model-routes.ts` + `src/custom-model-hosts.ts` — read/write API template
- `docs/multi-user-plan.md`, `docs/security-architecture.md` — admin-gating conventions
- `CLAUDE.md` §Multi-user mode, §"Settings surface", §"Per-device vs synced settings"
---
## Decisions (Phase 0, closed 2026-09-21)
1. **Enable/disable a stock CLI's `enabled` flag** — IN SCOPE. Plus a **master feature flag**
(`cliManagementEnabled`, synced, default OFF) gating the whole Settings UI section's visibility,
matching this codebase's standing convention for new admin-facing surfaces.
2. **Auto-install** (stock CLIs' already-shipped, already-vetted install commands) — IN SCOPE,
same PR.
3. **Custom CLI entries via the UI** — IN SCOPE, **typed-argv only**: a custom entry goes through
the exact same schema/argv-safety path stock entries do (named token patterns, no raw shell-text
field). Its install command stays **display-only text**, same as every stock entry today — Phase
4's auto-install NEVER executes a custom entry's install command, only a stock one's. This is
the one place scope was deliberately narrowed relative to what was agreed in principle, because
`docs/cli-registry.md`'s arg-template-safety section exists specifically to keep config free of
shell text, and a free-text install command for a user-defined entry would reopen exactly that.
4. **`shell`/`claude` un-disableable** — enforced at the **backend**, not just the UI (a
frontend-only guard is bypassable with curl).
5. **Non-admin visibility in multi-user mode** — the CLI-management Settings section is **hidden
entirely** for a non-admin, not shown-empty.
6. **`seededStockIds` ratchet** — not needed. `deepMerge()` only overrides a key the file actually
sets, so a CLI absent from `clis.json.clis` always falls through to its stock `enabled` value
with no special-casing. (Carried over from the first draft, not re-litigated.)
---
## Phase 1 — Master feature flag: `cliManagementEnabled`
**Status:** DONE (commit `da07b38c`) — verified present in `SettingsUpdateSchema`, `index.html`,
`openAppSettings()`/`saveAppSettings()`.
**Spec refs:**
- `schemas.ts:1319` (`customModelEndpointsEnabled`) — the exact pattern to mirror: `z.boolean().optional()`
in `SettingsUpdateSchema`
- `settings-ui.js:401`/`:2120` — checkbox read/write by id in `openAppSettings()`/`saveAppSettings()`
- `CLAUDE.md` §"Adding Features" → "App setting" — decide per-device vs synced FIRST (this one is
synced: a feature toggle, not a display preference) and add to `displayKeys` NEVER for a synced
setting
**Checklist:**
- [x] Add `cliManagementEnabled: z.boolean().optional()` to `SettingsUpdateSchema`
- [x] Add the checkbox to `index.html`'s `#settings-clis` section, above where Phase 6's per-CLI
list will render — reads/writes via `openAppSettings()`/`saveAppSettings()` by id, same as
`customModelEndpointsEnabled`
- [x] `readCliManagementEnabled()` helper (mirrors `readCustomModelEndpointsEnabled()` in
`custom-model-routes.ts:609`) for the route file(s) in Phases 2-5 to gate on
- [x] When OFF: `GET /api/clis` still exists but the Settings UI section stays hidden
(`applyCliManagementVisibility()`); the write endpoints reject (see Phase 3)
**Verify:** `npm run typecheck` passes; a unit test confirms `SettingsUpdateSchema` accepts/rejects
the field correctly; toggling it in a fresh browser profile shows/hides the Settings section with
no server restart.
---
## Phase 2 — Read endpoint: `GET /api/clis`
**Status:** DONE (commit `da07b38c`) — verified present in `src/web/routes/cli-registry-routes.ts`.
**Spec refs:**
- `src/web/routes/custom-model-routes.ts:730` (`GET /api/model-endpoints`) — multi-user read
gating: empty list for a non-admin, never a 403
- `src/config/cli-registry/registry.ts` — `listClis()` (every entry, including disabled stock
ones — this is an admin/settings surface, unlike `enabledClis()`)
- `window.__codemanCliAvailable`'s resolvers (`isClaudeAvailable()` etc.) — candidate `installed`
source; confirm whether to reuse directly or the response needs its own probe (Open Question 4,
carried from the first draft — still genuinely open, decide during this phase not before)
**Checklist:**
- [x] New route file `cli-registry-routes.ts`
- [x] Response excludes `launch`/`env`/`capabilities`/`overlays`/`discovery`
- [x] `isMultiUserMode() && !isAdmin(req)` → `[]`
- [x] Unit tests in `test/routes/cli-registry-routes.test.ts` (admin/non-admin/single-user,
disabled stock CLI still present)
**Verify:** `npm test -- test/routes/cli-registry-routes.test.ts` passes; `curl localhost:3000/api/clis | jq`
shows every stock CLI including disabled ones.
---
## Phase 3 — Write endpoint: `PUT /api/clis/:id` (stock enable/disable)
**Status:** DONE (commit `db4557d9`) — `UNDISABLEABLE_IDS`, admin gate, tmp+rename+0600 all
confirmed present.
**Spec refs:**
- `src/web/routes/custom-model-routes.ts:753` + `src/custom-model-hosts.ts:91` — write-path
template: `adminOnly` gate, read-modify-write the WHOLE file, tmp+rename+0600
- `registry.ts:47` (`filePath()` = `dataPath(...)`) and `reloadCliRegistry()` — write to the same
resolved path, invalidate the cache on every successful write or the change is invisible until
restart
**Checklist:**
- [x] Body: `{ enabled: boolean }`. Zod schema in `schemas.ts`
- [x] Gate order: `cliManagementEnabled` → `adminOnly` → shell/claude guard → stock-only guard
- [x] Rejects disabling `shell` or `claude` (`UNDISABLEABLE_IDS`)
- [x] Rejects a write for an id that isn't a stock CLI
- [x] Deep-merges `{ clis: { [id]: { enabled } } }`, preserving other override keys
- [x] tmp+rename+0600 write, `reloadCliRegistry()` on success
- [x] Unit tests (`test/routes/cli-registry-routes.test.ts`)
**Verify:** `npm test` full gate green; `curl -X PUT localhost:3000/api/clis/grok -d '{"enabled":false}'`
then `GET /api/clis` shows the change with no restart; same against `shell`/`claude` returns an
error and changes nothing; `ls -la ~/.codeman/clis.json` shows mode 0600.
---
## Phase 4 — Auto-install: `POST /api/clis/:id/install` (stock CLIs only)
**Status:** DONE (commit `db4557d9`) — route present, `appendAdminAudit` wired in.
**Spec refs:**
- `registry.ts:231` (`resolveInstallCommandForPlatform`) — currently "Display text only — never
executed"; this phase is what changes that, for stock entries only, with Decision 2's sign-off
- Original #343 review's exact concern re: `env.allowedPrefixes` contributed from a write — stays
out of scope; this phase only ever runs a command, never touches the env allowlist
**Checklist:**
- [x] Separate endpoint from Phase 3's toggle
- [x] Gate order: `cliManagementEnabled` → `adminOnly` → stock-entry-only guard
- [x] `resolveInstallCommandForPlatform(entry)` for the target
- [x] Bounded execution (timeout, captured stdout/stderr)
- [x] Does NOT auto-enable on successful install
- [x] Audit-logged via `appendAdminAudit`
- [x] Unit tests
**Verify:** a real install triggered via the endpoint against a CLI not currently installed,
`GET /api/clis`'s `installed` field flips true with no restart; audit log entry present; attempting
install against a custom entry's id fails with a clear error; full CI gate green.
---
## Phase 5 — Custom CLI entries: create / update / delete via API
**Status:** DONE (commit `db4557d9`) — `POST /api/clis`, `PUT /api/clis/custom/:id`,
`DELETE /api/clis/:id` all present. Open Question 2 resolved: a **separate** endpoint
(`PUT /api/clis/custom/:id`), not Phase 3's `PUT /api/clis/:id` widened.
**Spec refs:**
- `docs/cli-registry.md` §"Arg-template safety" (all four layers), §"Adding a CLI" (the 5-step
recipe) — a custom entry created via this API must satisfy the SAME schema (`CliEntrySchema`)
every stock entry does; there is no relaxed path for UI-originated entries
- `registry.ts`'s `resolveRegistry()` — the custom-entry branch (`stock: false`, dropped with a
warning on validation failure, never falls back silently) already exists and is unchanged by
this phase; this phase only adds a way to WRITE what that branch reads
**Checklist:**
- [x] `POST /api/clis` (create), full `CliEntrySchema` validation
- [x] `PUT /api/clis/custom/:id` (update) — separate endpoint from Phase 3's stock toggle
- [x] `DELETE /api/clis/:id` refuses for any stock id
- [x] `id` collision check against existing stock ids
- [x] `discovery.install.command` on a custom entry stays DISPLAY-ONLY
- [x] Same tmp+rename+0600 write pattern, `reloadCliRegistry()` on every successful mutation
- [x] Unit tests
**Verify:** `npm test` full gate green; create a custom entry via curl, confirm it appears in
`GET /api/clis` — **confirm it appears in the Run menu is UNVERIFIED and currently FALSE, see
"Outstanding" above**; delete it, confirm it's gone and `clis.json` no longer references it.
---
## Phase 6 — Settings UI
**Status:** DONE (commit `db4557d9`) — `#cliListGroup`, row rendering, toggle, Install button,
custom-entry create/edit/delete form all present in `settings-ui.js`/`index.html`. Manual browser
verification per the phase's own "Verify" step (flag on/off, non-admin hidden, toggle stops the
Run menu offering a CLI, create/enable/launch a custom entry, delete it, shell/claude undisableable)
has **not** been re-run in this session — the toggle→Run-menu leg specifically was BROKEN until the
gotcha fix above, and the create→launch leg for a custom entry is the confirmed gap in
"Outstanding".
**Spec refs:**
- `index.html:2357` (`#settings-clis`) — the existing home; Phase 1's master toggle at the top,
then the per-CLI list, then (if `cliManagementEnabled`) a "custom CLI" creation form, all above
the existing Codex-only groups
- `CLAUDE.md` §"Settings surface" — App Settings scrolls, it does not tab-switch
- `admin-ui.js` — pattern for an admin-only-VISIBLE section (not just admin-only-writable),
needed here per Decision 5
**Checklist:**
- [x] Whole section hidden when `cliManagementEnabled` is OFF, and separately hidden for a
non-admin in multi-user mode (`_applyCliManagementAdminGate`)
- [x] Fetches `GET /api/clis` when the section becomes visible; renders one row per CLI
- [x] Stock rows: enabled toggle only; `shell`/`claude` rows show the toggle disabled/greyed
- [x] Custom rows: enabled toggle plus edit/delete affordances
- [x] "Add custom CLI" form (id/label/badge/binary/argv)
- [x] Toggle/edit/delete update the row in place
**Verify:** manual browser test per `CLAUDE.md`'s "Always Test Before Deploying" rule — **not yet
re-run end-to-end in this session**; do this before considering the feature ready to ship, and
expect the custom-entry-launch step to fail until the Outstanding gap above is closed.
---
## Remaining Open Questions
1. **Phase 2's `installed` source** — resolved: reuses `window.__codemanCliAvailable`'s existing
resolvers via `GET /api/clis`'s own probe (confirmed by reading the route).
2. **Phase 5's `PUT` endpoint shape** — resolved: a **separate** endpoint
(`PUT /api/clis/custom/:id`), not Phase 3's toggle route widened.
3. **Sequencing against the parked Copilot plan** — unchanged, still not blocking.
4. **NEW: custom-CLI Run-menu integration** — see "Outstanding" above. Not decided or started.
---
Implementation is underway (see Status above); this line is left for history rather than removed —
the plan was originally approved before Phases 1–6 landed.
+54 -5
View File
@@ -18,7 +18,17 @@ Every run mode Codeman can launch — Claude Code, Terminal/Shell, OpenCode, Cod
## The override file
`~/.codeman/clis.json` (instance-scoped through `dataPath()`) holds overrides and custom entries only, never a copy of the stock catalog: `{ "clis": { "<id>": { ...partial entry... } } }`. Objects merge key-wise onto the stock entry, arrays replace wholesale. **The file must be mode 0600**; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you `chmod 600` it. Every reason a file was ignored or an entry dropped is logged once, prefixed `[cli-registry]`, on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read only on restart.
`~/.codeman/clis.json` (instance-scoped through `dataPath()`) holds overrides and custom entries only, never a copy of the stock catalog: `{ "clis": { "<id>": { ...partial entry... } } }`. Objects merge key-wise onto the stock entry, arrays replace wholesale. **The file must be mode 0600**; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you `chmod 600` it. Every reason a file was ignored or an entry dropped is logged once, prefixed `[cli-registry]`, on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read after a change made through CLI management (below).
## Managing CLIs from Settings
App Settings → Agents & CLIs → **CLI management** (`cliManagementEnabled`, default OFF; admin-only in multi-user mode) lists every entry with an installed/not-installed badge and:
- toggles any entry on or off. A `kind: 'shell'` entry cannot be disabled, and the row shows no switch for it. A disabled CLI disappears from the Run menu, the welcome screen and the phone overview, and new session requests for it are rejected.
- installs a missing **stock** CLI by running its shipped install command, after a confirm that names the exact command. Only one install per CLI runs at a time, and the command runs without any `CODEMAN_*` variable in its environment. A custom entry's install command is never executed.
- adds, edits and deletes **custom** entries (id, label, badge, binaries, launch argv). The server re-validates the whole assembled entry through `CliEntrySchema`, so the form cannot bypass the load-time rules.
These are the only writes to `clis.json`. They are serialized, and a file that does not parse or has unsafe permissions is refused rather than overwritten; fix it (or `chmod 600` it) and retry. The HTTP routes are listed in `docs/api-reference.md` under *CLI management*.
## The shape of an entry
@@ -36,7 +46,8 @@ interface CliEntry {
launch: CliLaunch; // the structured argv template
env: CliEnv; // exports, tmux setenv keys, the env-override allowlist
capabilities: CliCapabilities; // what every call site reads instead of the id
// .workDetect?: { promptGlyph, workingLine } — how this CLI's pane shows work
// .workDetect?: { promptGlyph, workingLine, watchingLine?, watchingLines? } — how
// this CLI's pane shows work, and how it shows work it started in the background
overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store
}
```
@@ -45,10 +56,46 @@ interface CliEntry {
### Regexes that come from config
Two capability fields carry a regular expression an override file can set: `discovery.version.regex` and `capabilities.workDetect.workingLine`. Both go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
Three capability fields carry a regular expression an override file can set: `discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`. All three go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
`workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused.
`watchingLine` reads a different row of the same screen. A CLI draws it while work the agent
itself started is still running — Claude prints `⏵⏵ bypass permissions on · 1 monitor · ← for
agents` while a monitor, a backgrounded shell or a cloud session is live. Codeman turns that
into `Session.watching`, and an idle prompt from such a session opens already acknowledged,
so a pane waiting for its own background work never raises an alert a human cannot answer.
Group 1 is the label, and a CLI that declares no pattern reports no background work.
Two CLIs declare such a row today, and they put it in different places. Claude writes its
chip on the last row of the screen, so it keeps the default one-row window and anchors on
the `·` its footer joins items with. Codex pins
`1 background terminal running · /ps to view · /stop to close` ABOVE its composer, which
puts the row third from the bottom once the status line and the composer are counted, so its
entry declares `watchingLines: 3` and matches that row end to end. Both were measured
against live panes rather than read out of a binary, which is the standard for adding a
third.
That label is the one value in the registry that an AGENT can influence, because it comes off
the agent's own screen. Two things keep it honest, and both belong to whoever adds a pattern
for a new CLI. `watchingLabel()` in `session-activity.ts` searches only the last few
non-blank rows, which should be the part of the screen the CLI draws rather than the agent,
and the pattern should anchor on chrome only that CLI can produce. Keep the window as small
as the layout allows, since every row it adds is another row the agent may be able to write.
The label is also ANSI-stripped and length-capped at the source, and every interpolation of
it into markup goes through `escapeHtml()`, since it ends up on a badge and in an approval
card.
The two shipped entries do not sit equally well behind that rule, and the difference decides
what a pattern is allowed to do. Claude's chip is the last row, so its one-row window holds
nothing the agent can write — not even the status line above it, whose command a session
running with permissions bypassed can write into its own `.claude/settings.json`. Codex's row
shares its slot with the last row of the transcript whenever no terminal is running, so a
message ending in that exact line is matched. What keeps that harmless is `hooks: 'none'`: no
hook event from a codex session reaches the approvals inbox, so a forged label costs a wrong
badge and cannot silence an alert. Before giving a CLI both hook signals and a pattern, make
sure its row is one the agent cannot write.
### Three capabilities that must stay independent
`external`, `hooks` and `altScreen` describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. `shell` has no hooks but is **not** an external CLI, so a hooks predicate written as `!isExternalCliMode()` accepted `until=stop` on a shell session and then blocked the caller for their entire timeout. `deepseek` is the mirror image: it IS external and it DOES have hooks.
@@ -102,6 +149,8 @@ It matches four shapes, not one: `mode === '<id>'`, `mode !== '<id>'`, `case '<i
The allowlist is not a formality. If a branch is about what a CLI can DO it belongs in `CliCapabilities`; the entries that remain are things that are not CLI-behaviour branches at all — chiefly the legacy per-mode `<Mode>Config` objects on `POST /api/sessions`, which are a fact about the public HTTP API rather than about any CLI, plus a few documented cases where `mode === 'claude'` is genuinely the right question (Read My Mind reads Claude's _own_ transcript, so a capability there would be actively wrong).
`test/frontend-cli-no-id-branching.test.ts` is the same guard for the two frontend files the CLI registry's Run-menu consolidation touches, `session-ui.js` and `mobile-overview.js` — deliberately not the rest of `src/web/public/`, whose per-CLI rules stay out of scope for now (see "Fields declared for later" below). Its allowlist keys on `<file>::<expression>` with no line number, since a single unrelated edit to a contended file would otherwise shift every subsequent line and make every entry go stale at once, and each entry additionally carries the exact number of approved call sites — a bare key would let a brand-new branch reusing an already-approved expression land unreviewed. Its comparison shape differs from the backend guard's in one respect: the left-hand side may be any identifier, not only one named `mode`, `id` or `agentType`, because the review of #458 found `const m = this._runMode; if (m === 'codex')` slipping past the named form while the scanned file already filters with `(m) => m !== 'shell'`.
## Two namespaces called `param`
`launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a **launch param**. The **legacy wire field** a param arrives as is a separate namespace, and `launch.legacyConfigAliases` is the only bridge between the two.
@@ -110,9 +159,9 @@ This matters because it is invisible when it is wrong. `capabilities.privilegedP
## Fields declared for later
`shortBadge`, `accent`, `capabilities.echo`, `capabilities.wheelForward`, `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` are **declared but not yet read**. They all describe frontend behaviour, and the frontend is deliberately untouched here: `app.js`, `terminal-ui.js` and `styles.css` keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
`accent`, `capabilities.echo`, `capabilities.wheelForward`, `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` are **declared but not yet read**. (`shortBadge` was on this list until the CLI management list in Settings started showing it.) They all describe frontend behaviour, and the frontend is deliberately untouched here: `app.js`, `terminal-ui.js` and `styles.css` keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
Treat those values as **transcribed, not authoritative** — nothing enforces that `echo.policy` matches `_updateLocalEchoState`'s fallthrough, or that `accent` matches the gradient CSS paints, so re-measure before wiring one up. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; `test/cli-registry-no-id-branching.test.ts` pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
Treat those values as **transcribed, not authoritative** — nothing enforces that `echo.policy` matches `_updateLocalEchoState`'s fallthrough, so re-measure before wiring one up. `accent` is the one exception: it was measured against styles.css on 2026-09-21 (method in the comment above `CLAUDE` in `stock.ts`), though nothing keeps it in step with the CSS either. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; `test/cli-registry-no-id-branching.test.ts` pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
`overlays.credStore` is in the same category, for a sharper reason: the Docker credential-seeding path still reads its own `CRED_STORES` table, because this shape allows ONE store per CLI and the live table needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek declares none here even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is simultaneously the worst thing here to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest.
+25 -13
View File
@@ -104,17 +104,17 @@ declared capability, never an `if (mode === 'claude')` branch.
## Per-CLI injection recipes (confidence-ranked)
| CLI | Mechanism | Confidence |
| ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol CONFIRMED BROKEN against llama.cpp/llama-swap**: codex only speaks the Responses API (`wire_api = "responses"`, the only value it accepts since it dropped `"chat"` support in Feb 2026), and a real llama-swap server does not implement `/v1/responses` — a live run against it failed with repeated `Reconnecting...` then `high demand` errors. Codex support therefore needs a Responses-API-compatible endpoint (most local llama.cpp/Ollama/vLLM setups do not qualify); do not present this as working against a generic OpenAI-Chat-Completions box |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Confirmed reaching the server, but failing — unresolved.** A real run against llama-swap returns `dsh: HTTP_404: DeepSeek API error (HTTP 404)` consistently (confirmed the env vars are read: the request reaches the network rather than failing locally). Root cause not identified — plausible explanation by analogy with codex's Responses-API gap is that `dsh --profile headless` expects DeepSeek's official API response shape/path structure rather than a generic OpenAI-compatible `/v1/chat/completions` endpoint, but this was not confirmed by reading dsh's own bundled source (unlike pi/grok, where that grep resolved the question directly). Documented as best-effort/unknown, not shipped as verified working |
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
| CLI | Mechanism | Confidence |
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol picture more nuanced than a flat break, re-verified live twice on 2026-09-17 against a llama-swap deployment that DOES answer `/v1/responses`** (an earlier test's `Reconnecting...`/`high demand` failure does not reproduce against every llama-swap setup): a plain, no-tool-call chat turn (`codex exec 'reply with just OK'`) returned a real reply. But a real tool-call attempt (`run the shell command: echo hello`) came back as an `agent_message` TEXT item — the tool-call JSON printed as the model's answer, not a `function_call` item codex would actually execute (confirmed via `codex exec --json`'s raw event stream: `item.completed`/`agent_message`, never `function_call`). Since tool execution is what makes codex a coding agent at all, this remains **not usable for real work**, just with a different, more specific failure mode than previously documented — still do not present this as working. Separately, EVERY custom-endpoint codex session also prints `warning: Model metadata for '<id>' not found. Defaulting to fallback metadata...` on launch (confirmed harmless — the successful plain-text reply above still had it): codex's per-model metadata (reasoning tiers, system-prompt templates, context-window figures) comes from `models_cache.json`, a LOCAL CACHE of OpenAI's own hosted model catalog that a custom model can never appear in by construction. No config.toml override exists for it, and the isolated `CODEX_HOME` never gets a `models_cache.json` written into it at all (confirmed: inspected a live, actively-used isolated dir — codex evidently can't reach OpenAI's catalog endpoint for this session and just falls back silently every time, with no file left behind to fix or clean up). Fabricating a fake catalog entry to suppress the warning would mean copying the _shape_ of OpenAI's own proprietary schema — including their real per-model system-prompt content, visible in a genuine `models_cache.json` — for a warning confirmed to have no effect on the actual (broken) tool-calling outcome; not worth building |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not** `PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
| `deepseek` | Reuse the **existing** `DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`), now with `appendV1Suffix: true` (see confidence). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Root cause of the original `HTTP_404` found and fixed, by reading dsh's own bundled source — the same bar pi/grok's fixes were held to.** Installed `@deepseek-ai/dsh` (all its real published dependencies) into a scratch directory purely to read `@deepseek-ai/dsh-llm-deepseek/lib/index.js`: it builds its request as `fetch(\`${connection.baseURL}/chat/completions\`, ...)`with`baseURL`read straight from`DEEPSEEK_BASE_URL`(or defaulting to DeepSeek's real public API root,`https://api.deepseek.com`, which also carries no `/v1`) — no `/v1` insertion of dsh's own, unlike the OpenAI-SDK convention this recipe originally assumed. llama-swap/llama.cpp only ever serves the OpenAI-conventional `/v1/chat/completions`. Confirmed live: `POST <baseUrl>/chat/completions` → `404`, `POST <baseUrl>/v1/chat/completions` → `200`, on the exact same endpoint — and dsh's own error-message template, `DeepSeek API error (HTTP ${status})`, reproduces the originally reported `dsh: HTTP_404: DeepSeek API error (HTTP 404)` precisely. Fixed by adding `appendV1Suffix` (env kind only, deepseek's entry alone — claude/gemini must NOT get it, since claude was already confirmed working against the unmodified `baseUrl`), which runs `endpoint.baseUrl` through the same `withV1Suffix()` helper `configDir`-kind CLIs already use. ⚠️ Not yet re-run end-to-end with a real `dsh` binary — no install available in this environment (no npm-installed CLI binary in `PATH`, and the `codeman-test-picker` container doesn't bundle it either); the fix is source-confirmed and live-verified at the HTTP level, but a genuine "hello world" reply through `dsh` itself is the remaining step before promoting this to **verified** alongside claude/opencode/pi/grok/omp |
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
Everything web-researched-but-unverified gets implemented but must be
smoke-tested against real installs of those CLIs before being called done —
@@ -208,6 +208,14 @@ extra per-model configuration on Codeman's side at all.
### 4. Toolbar UI
> **Superseded.** This section describes the toolbar-button design as originally
> planned. What actually shipped is a Run-menu picker instead: one generated entry
> per (capable harness, saved endpoint) pair directly in the existing `#runModeMenu`
> dropdown, rather than a separate `#customModelBtn`/`#customModelMenu` surface. See
> [`docs/custom-model-endpoints.md`](custom-model-endpoints.md#the-run-menu-picker)
> for the current design; the sections below (session-restart mechanics, security)
> remain accurate regardless of which UI calls the underlying route.
- New header/toolbar button (e.g. `#customModelBtn`, `btn-toolbar
btn-custom-model`), marker-hidden by default (`btn-custom-model--hidden`)
and revealed by `applyHeaderVisibilitySettings()` only when
@@ -344,8 +352,12 @@ pure unit tests and the live manual checks in Verification:
up automatically with zero edits to the script). Already run to
completion against the author's llama-swap server (a LAN address,
inside a `codeman/agent:llm-test` Docker image with all 9 CLI binaries):
claude/opencode/pi/grok/omp **PASS**, codex **FAILs as expected**
(Responses-API protocol gap, not a bug), gemini/deepseek **UNCONFIRMED**
claude/opencode/pi/grok/omp **PASS**, codex **partially works and still
isn't usable** (plain chat succeeds against a llama-swap deployment that
answers `/v1/responses`, but a real tool-call attempt comes back as
inert text rather than an executable `function_call` — see the
confidence table row for the full, re-verified picture), gemini/deepseek
**UNCONFIRMED**
(reach the server, fail for undiagnosed reasons — see their table rows),
antigravity **SKIP** (no mechanism). Re-run this against a real cloud
endpoint (e.g. an Azure AI Foundry deployment) once one is available, to
+434 -18
View File
@@ -11,20 +11,22 @@ company gateway) — anything answering `GET /v1/models` and
recipe confidence table, and security reasoning:
[`custom-model-endpoints-plan.md`](custom-model-endpoints-plan.md).
> **Status**: backend is implemented and tested (registry capability, the
> injection engine, the endpoint store + discovery route, the session
> restart route). The toolbar picker / settings UI described below as the
> intended surface is **not yet built** — until it lands, use the HTTP API
> directly (examples below). Antigravity has no known custom-endpoint
> mechanism and is not supported.
> **Status**: fully wired end to end — registry capability, the injection
> engine, the endpoint store + discovery route, both the restart-in-place
> apply route (Claude) and the one-shot quick-start launch path (every
> other supported harness), a settings-panel CRUD surface, and the Run-menu
> picker described below. Antigravity has no known custom-endpoint
> mechanism and is not supported. The HTTP API (examples below) still works
> directly and is what the picker itself calls under the hood.
## Turning it on
App Settings → Agents & CLIs → **Custom Model Endpoints** (synced setting
`customModelEndpointsEnabled`, default **OFF**). Until the toolbar picker
lands, nothing reads this setting: the HTTP routes below work whether it is
on or off, and it exists now only so the picker has a switch to hang off
when it ships. The API equivalent:
App Settings → Models → **Custom model endpoints** (synced setting
`customModelEndpointsEnabled`, default **OFF**). Turning it on does two
things: it reveals the endpoint list/add/edit/discover panel in that same
settings section, and it makes the Run menu offer a generated entry per
(harness, endpoint) pair — see "The Run-menu picker" below. The API
equivalent:
```bash
curl -sk -X PUT https://localhost:3000/api/settings \
@@ -34,6 +36,9 @@ curl -sk -X PUT https://localhost:3000/api/settings \
## Adding an endpoint
Via App Settings → Models → Custom model endpoints → **+ Add endpoint**, or
directly:
```bash
curl -sk -X POST https://localhost:3000/api/model-endpoints \
-H 'Content-Type: application/json' \
@@ -62,7 +67,201 @@ configured, `PUT`/`DELETE /api/model-endpoints/:id` update or remove one.
Endpoint management is admin-only in multi-user mode, same as remote/docker
hosts — these are machine-level infra, not per-user settings.
## Applying a model to a session
**Context length is discovered too, opportunistically and safely.** The plain
`GET /v1/models` response has no context-window field. Discovery only ever
looks for one for a model llama-swap's own response already reports
`status.value === "loaded"` for — never for an unloaded one, because
llama-swap treats `?model=` as a routing hint and asking about a model that
isn't loaded risks triggering an actual (slow, GPU-swapping) load as a side
effect of what should be read-only discovery. A server with no `status` field
on any entry at all (not llama-swap) gets no context-length enrichment,
rather than guessing. A model's previously-learned context length survives a
later cycle where it wasn't the loaded one; it's dropped only once the model
disappears from the endpoint's list entirely. Stored per model in
`modelContextLengths` and applied automatically (see "Applying a model to a
session" below) so a CLI that would otherwise assume a large default context
window for an unrecognized model id stops silently overflowing a much
smaller real one.
**Where that number actually comes from matters, and got this wrong once
already.** The first cut read it from llama.cpp's own
`GET /props?model=<id>` (`n_ctx`) — plausible, and it worked in testing, but
confirmed live to be actively WRONG for a `--fit-ctx`-launched llama-swap
backend: `/props` reported `n_ctx: 154112` for a model llama-swap itself had
launched with `--fit-ctx 16384`, and the real server then refused a request
right at that real 16384-token limit — `/props`'s `n_ctx` appears to report
the model's theoretical/trained maximum there, not the runtime-configured
one. Discovery now parses the REAL configured size straight out of
llama-swap's own launch command instead (`GET /running`'s `cmd` field —
`--fit-ctx <N>` first, then the plain llama.cpp `-c`/`--ctx-size` a
hand-written command might use), and only falls back to the `/props` probe
when `cmd` states no recognizable flag at all.
**File size is discovered too, when the server states one.** llama-swap
writes a GB figure into an auto-discovered model's own `description`
(`"Auto-discovered 16.35 GB - parameters auto-fitted by llama.cpp"`), parsed
into `modelSizesGB` — unlike context length, this needs no `/props` probe
(the figure is right there in the `/v1/models` response) and so is populated
for every model regardless of loaded state. A hand-configured profile's own
description has no such figure and correctly gets no entry, never a guess.
Used only to label the Run-menu picker's "loading model" banner (e.g.
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB) on llama-swap..."); never anything
a server-side check relies on.
**The loading banner is unbounded by design, and says so — no countdown, no
automatic give-up.** An earlier version scaled an expected-time estimate and
a timeout off the model's file size and auto-closed the session once that
elapsed, but a real load's actual duration depends on hardware this feature
has no way to know (VRAM, storage speed, whatever else is contending for the
GPU) — any fixed number was a guess dressed up as a fact, and a model that
genuinely takes 10+ minutes on slower hardware would just get killed
mid-load by its own display. The banner now says outright that it can take a
while depending on hardware and model size, polls
`GET /api/model-endpoints/:id/running-status` every second for as long as it
takes, and carries a **Cancel** button (rendered on the banner itself) that
ends the wait and closes the session the load was for — the user's own call
on when it's taking too long, not a fixed number baked into the client.
**The banner's second line is the real backend log line, not a guess.**
llama-swap's `GET /api/events` SSE stream carries the actual `llama-server`
process's own stdout — `load_model: loading model '<path>'`,
`llama_server: model loaded`, tokenizer warnings, all of it — tagged
`source: "upstream"`, distinct from llama-swap's own `source: "proxy"`
request-access lines. `running-status`'s response now includes `logLine`
(via `getLatestLlamaSwapLogLine`), and the banner shows it on its own line
under the disclaimer, e.g. "llama.cpp: load_model: loading model '...'" —
confirmed live end-to-end through a real forced swap, sequentially showing
the model path, a tokenizer warning, then staying on whatever llama.cpp last
printed once the load goes quiet (never cleared back to blank). ⚠️
**`GET /logs` — the endpoint this feature's own first cut was built
against — turns out to carry ONLY llama-swap's own proxy request-access
log.** Confirmed live it never showed a single backend line, even seconds
after a real, verified model swap; `/api/events`'s `logData` frames are the
only source that actually has it, and its own `source` field (`upstream` vs
`proxy`) is what `getLatestLlamaSwapLogLine` filters on. One `/api/events`
connection is held open per endpoint and reused across every session
watching a load on it (confirmed live to stay open indefinitely, unlike
`/logs`, which closes after a fixed ~100KB), idle-closed after 30s of nobody
polling it (`pruneIdleLlamaSwapLogTails`, same 20s sweep as the
swap-displacement check below).
`defaultModelId` names which discovered model the picker pre-marks for that
endpoint — the settings panel's Edit form exposes it as a select populated
from the endpoint's own discovered `models`, and the route refuses a value
that isn't one of them. It is applied automatically only when the endpoint
has exactly one discovered model (nothing to choose); with two or more it
is a pre-selection in the model-picker dialog below, never a silent default.
Re-discovering drops a default that no longer appears in the fresh list
rather than carrying an invalid one forward.
**Model lists refresh themselves.** A background sweep (`server.ts`,
`CUSTOM_MODEL_REDISCOVER_INTERVAL_MS`, every 5 minutes) re-discovers every
saved endpoint the same way the manual `POST .../discover-models` route
does, best-effort per endpoint — one being unreachable on a given cycle
never blocks the others. Off under `npm test`, same reasoning as the Codex
plan-usage poll it sits beside: no real network to hit, no server instance
to keep the timer alive for.
## The Run-menu picker
With the setting on and at least one endpoint carrying a discovered model,
the toolbar's Run dropdown grows a **Custom Endpoints** section: one entry
per (harness that can redirect to a custom endpoint, saved endpoint) pair,
e.g. "Claude Code (llama.cpp)". The harness list is read off the CLI
registry's own `capabilities.customModelInjection` at page render
(`window.__codemanCustomModelClis`, `server.ts`) — never a hardcoded id list
in the frontend — so a CLI whose injection recipe lands later shows up with
no frontend change, and Antigravity (`unsupported`) never does.
Picking an entry re-fetches the endpoint (`selectCustomModelEntry()`,
`session-ui.js`) rather than trusting anything cached from the dropdown's
own render — the model list can have changed via the 5-minute sweep above
or a settings-panel edit since the menu opened. With exactly one discovered
model it runs straight away; with two or more, a small modal
(`#customModelPickModal`) lists them and asks which one to use for this
launch, with the endpoint's `defaultModelId` marked but not auto-chosen —
the point of asking is letting one launch deliberately differ from the
saved default, not just confirming it.
The modal promotes exactly one row to the top of the list rather than
always showing raw discovery order, so the zero-wait choice is the one
under your thumb:
- **"Currently loaded"** — a model from this host's own list that
llama-swap reports `ready` right now, queried via
`GET /api/model-endpoints/:id/running-status`. Bounded client-side to
~800ms (`Promise.race`), on top of the route's own 5s server-side
timeout, so an endpoint that is asleep or firewalled cannot leave the
modal invisible for the full 5s after the Run menu has already closed.
- **"Last used"** — shown only when nothing is currently loaded: the model
actually launched last for this exact (harness, endpoint) pair, read
from the per-device `codeman:customModelLastUsed:<mode>:<endpointId>`
localStorage key. Written by `_runCustomModelEntryViaRestart` (claude)
and `_quickStartWithCustomModelConfirm` (every one-shot launch; the
`runCustomModelEntry` entry point itself only dispatches between the
two) only once the model is actually applied, never on the mere click —
declining the context-window warning means this exact model cannot work
with this CLI at all, so promoting it next time would be actively wrong,
not just premature.
Neither tag reorders anything past that one promoted row. The "Default"
pill is a separate span, not a third value of the same slot: a promoted
row that is also the endpoint's `defaultModelId` shows both tags (on a
single-purpose GPU box that is the common case, and an exclusive slot
silently dropped the Default marking for exactly that row), and a row
with neither promotion nor default shows no tag at all.
**How the launch itself applies the endpoint depends on the harness.** For
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP (`runCustomModelEntry` →
`_runCustomModelEntryOneShot`), the endpoint/model is folded into the SAME
`POST /api/quick-start` call that creates the session (`customModel` field),
so the session launches directly on the endpoint — no restart, no visible
relaunch. Claude (`_runCustomModelEntryViaRestart`) still uses the original
two-step design: the launch runs a single native session exactly the way its
own Run-menu entry would, then **waits for the new session to go idle**
(`GET .../wait?until=idle`, bounded at 20s — a normal 200 either way, never
an error, per the wait endpoint's own contract) before applying the endpoint
via the restart route below. That wait exists because a freshly launched CLI
reports itself as `busy` for its own startup (a boot spinner, a
workspace-trust check) well before the apply call would otherwise reach it,
and the apply route correctly refuses to restart a session mid-turn — a
fresh boot looks exactly like one from the outside. A session still busy
after the wait reaches the apply call anyway and gets that route's own
honest `SESSION_BUSY` error, now visible as a sticky toast with a close
button rather than a generic message that vanished in three seconds. Claude
stays on this path because its own restart (`--resume`-based, keeping the
conversation) is far less jarring than the other seven's, and `runClaude()`'s
multi-tab launch and docker-config-drift confirm/retry loop make folding it
into the one-shot path separate work. It is a
one-off "try this endpoint" action, not a sticky mode: the plain Run button
still means "this harness, native cloud" afterward. Entries are hidden
entirely for a remote or Docker active case, since the apply route refuses
both (see the next section).
## Launching directly on an endpoint (no restart)
```bash
curl -sk -X POST https://localhost:3000/api/quick-start \
-H 'Content-Type: application/json' \
-d '{"caseName": "myapp", "mode": "codex", "customModel": {"endpointId": "llama-box", "modelId": "qwen3"}}'
```
`POST /api/quick-start`'s `customModel` field (`{endpointId, modelId,
confirmed?}`) computes the same injection the restart route below does, but
BEFORE the session exists — the session is minted its own id up front
(`crypto.randomUUID()`), the injection (env vars, and for a `configDir`-kind
CLI, the written config file) targets that real id, and the session launches
already pointed at the endpoint. No restart, because there was never a
native-backend launch to restart away from. Runs the same llama-swap
conflict check as the restart route (below) — a `409`-shaped
`{requiresConfirmation, currentlyLoadedModel, affectedSessions}` response
with no session created, resolved by retrying with `confirmedSwap: true` — and
is refused the same way for a remote or Docker case. This is what the
Run-menu picker uses for opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP;
Claude still uses the restart route below (see "The Run-menu picker" above
for why).
## Applying a model to an ALREADY-RUNNING session
```bash
curl -sk -X POST https://localhost:3000/api/sessions/<sessionId>/custom-model \
@@ -84,6 +283,179 @@ since for those three the config file alone does not switch the model.
reattaches the durable remote/in-container tmux rather than relaunching the
agent, so the selection would report success and change nothing.
**Claude gets two more env vars when known/applicable, both declared on its
registry entry (`contextLengthVar`/`configDirVar`), not hardcoded here:**
- `CLAUDE_CODE_MAX_CONTEXT_TOKENS` is set to `modelId`'s discovered context
length (see the discovery section above) whenever one is known. Without
it, Claude Code assumes a large (200k) window for any unrecognized custom
model id and never compacts, which reliably overflows a much smaller real
local context — confirmed live: a stock ~33.7K-token system prompt against
a 16384-token llama-swap model failed with `exceeds the available context
size`. No entry for the model in `modelContextLengths` means the var is
simply omitted, never a guess. ⚠️ **This var only affects when Claude
Code compacts conversation _history_ — it cannot fix a model whose real
context is smaller than Claude Code's own fixed per-turn overhead**
(system prompt + tool schemas, empirically ~36.4K tokens, confirmed live
via an `in:0 out:0` failure on the very first message, before any
history exists to compact). No context-length declaration changes that
fixed overhead, so a model below the safe floor fails outright on
message one regardless of what this var says. See "Context-window floor
warning" below for how Codeman catches this case before launching
instead of after.
- `CLAUDE_CONFIG_DIR` is pointed at the same isolated per-session directory
the `configDir`-kind CLIs use (empty, no files written into it), so the
injected `ANTHROPIC_API_KEY` never shares a directory with a stored
claude.ai OAuth login. Claude Code still prints "Both claude.ai and
ANTHROPIC_API_KEY set" when the two coexist in the same config directory —
cosmetic (confirmed live: the API key wins for actual requests either way,
visible in the terminal's own `API Usage Billing` line) but worth
eliminating rather than living with. The directory's `projects`
subdirectory is symlinked (a junction on Windows) back to the real
`~/.claude/projects` so the response viewer, subagent windows and Read My
Mind keep working for that session — the same trade-off and fix documented
for a manually-set `CLAUDE_CONFIG_DIR` in
[`docs/wiki/Agent-CLIs.md`](wiki/Agent-CLIs.md), just applied
automatically here. Best-effort: a platform that refuses the symlink keeps
the pre-existing blind-response-viewer side effect rather than failing the
whole custom-model apply over it. ⚠️ **This relocates the whole `.claude`
tree, not just transcripts**: a custom-model Claude session also loses the
user's global `settings.json`, user-level skills (the codeman agent skill
included), user-level agents and commands, and the MCP servers configured
in `~/.claude.json` — none of those are symlinked back, only `projects` is.
A fine trade for "point this session at my local llama.cpp," but worth
knowing before it surprises you mid-session.
**That isolated directory needed one more fix to actually be usable
non-interactively.** An otherwise-empty `CLAUDE_CONFIG_DIR` has none of a
real profile's prior "Detected a custom API key — use it?" approvals, so
without more, Claude Code stops and asks that on _every single launch_ —
confirmed live, and with nobody at a TTY to answer, its own default answer
("No") silently refuses the very key this feature just injected, which
looks like the endpoint being ignored entirely. `customModelInjection`'s
`apiKeyTrustFile` (`{ relPath: '.claude.json', shape:
'claude-api-key-responses' }` on claude's entry) pre-seeds that exact
approval: the apply step merges `customApiKeyResponses.approved: [apiKey]`
into `<configDir>/.claude.json`, the same field a real answered prompt
itself writes to (confirmed against a real file after answering by hand
once) — this answers the prompt in advance rather than bypassing it. The
merge preserves whatever else the CLI already wrote into that file on an
earlier launch in the same isolated directory (`userID`, `numStartups`,
earlier approved keys), and a missing or corrupt file is treated as empty
rather than failing the apply.
**A fresh `CLAUDE_CONFIG_DIR` isn't just missing that one approval — Claude
Code treats it as a brand-new profile and replays its ENTIRE first-run
sequence on every launch: the theme picker, the security-notes screen, the
per-project "trust this folder?" dialog, and (running with
`--dangerously-skip-permissions`) a one-time warning about bypassing
permissions.** Confirmed live: none of these show up again for a real,
already-onboarded profile, but every custom-model session gets a fresh,
otherwise-empty isolated directory, so it saw all four every single time.
`customModelInjection`'s `skipFirstRunPrompts` (`true` on claude's entry,
requires `apiKeyTrustFile` since it reuses the same file) pre-seeds the
state a real profile accumulates from answering all of that once:
`hasCompletedOnboarding: true` and the launching session's own
`projects[workingDir].hasTrustDialogAccepted: true` go into the same
`<configDir>/.claude.json` the API-key approval above already merges into
(other projects, and other fields on this session's own project entry, are
left untouched), and `skipDangerousModePermissionPrompt: true` goes into
`<configDir>/settings.json` — a different file, merged the same
corrupt-tolerant way. `workingDir` is used exactly as the session was
launched with as its cwd, never realpath'd or slash-normalized, since
that's the literal string Claude Code itself uses as the project key.
**llama-swap gets two more fixes on top of the context-length/config-dir
ones above, both from watching a real switch live.** llama.cpp only ever
runs one model at a time; llama-swap swaps the backing process on demand,
which can take anywhere from a few seconds to well over a minute:
- **The conflict check.** Both apply routes (the restart one here and the
one-shot `POST /api/quick-start` above) call llama-swap's own
`GET /running` first — feature-detected, so a plain llama.cpp/OpenAI-
compatible server (no such endpoint) is simply never checked. If a
_different_ model is currently loaded and ready, and another **live
session's own selection** is using it, the apply returns
`{requiresConfirmation: true, currentlyLoadedModel, affectedSessions}`
instead of silently switching — nothing is applied or created yet.
Retrying with `confirmedSwap: true` skips the check (the legacy `confirmed: true`
still means both questions). Switching with nothing
else affected proceeds immediately; this is a warning about disrupting
another session, never a gate on the switch itself.
- **Actually starting the load.** llama-swap has no "switch model" admin
call — the only thing that starts a swap is a real inference request
naming the model, and confirmed live: applying a selection alone never
reached llama-swap at all (nothing in its own server logs), since nothing
had actually asked it to load anything yet. Both apply routes now also
send the smallest real request that will —
`POST <baseUrl>/v1/chat/completions` with `max_tokens: 1` and one
throwaway message — whenever the
target model isn't already the one loaded and ready, fire-and-forget (its
response is never read; `GET /api/model-endpoints/:id/running-status`,
polled client-side, is what actually confirms readiness). The response
also carries `modelSwapInProgress: true` in that case, which is what
drives the Run-menu picker's own "loading model" status banner.
## Catching a swap after the fact
The conflict check above only runs at the moment a session is created or a
model is applied — it has no way to catch a swap that happens **later**.
Confirmed live: a session created while nothing else conflicted at that
exact instant can still get silently displaced afterward, once a
_different_ session's own normal use (or its own create-time load trigger)
asks llama-swap to load something else. llama-swap has no push
notification of its own for this, so a background sweep
(`detectCustomModelSwapDisplacements`, `CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS`
= 20s in `server.ts`) polls `GET /running` once per distinct endpoint that
has at least one live custom-model session, and compares each such
session's own `modelId` against what is actually loaded. A session whose
model is no longer in that list gets a `custom-model:swapped-out` SSE event
(`{sessionId, sessionName, endpointId, previousModel, currentlyLoadedModel}`),
shown as a global toast — global rather than tied to that session's tab,
since the whole point is telling the user before they type into it
expecting the model they picked. Notifies **once per displacement**: the
same de-dupe `Set` clears a session's flag once its own model is loaded and
ready again, so a later, genuinely new displacement notifies again rather
than the session staying silently un-notified forever after the first one.
## Context-window floor warning
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
empirically ~36.4K tokens) can exceed a small local model's _entire_ real
context on its own, before any conversation history exists to fill it —
confirmed live twice, both as an `in:0 out:0` failure on the very first
message sent. `CLAUDE_CODE_MAX_CONTEXT_TOKENS` (above) cannot fix this: it
only governs when Claude Code compacts conversation history, and there is
no history yet on message one. Applying such a model would look like the
endpoint being ignored, or the wrong model being used, when in fact the
endpoint applied correctly and the model is simply too small for this CLI.
Both apply routes (the restart route and the one-shot `POST
/api/quick-start`) now check for this **before** launching or restarting
anything, gated on the CLI's registry entry declaring a `contextLengthVar`
(currently only claude — the check is a no-op for every other CLI by
construction, never a hardcoded mode check). If the model's discovered
context (`modelContextLengths`, from discovery above) is below
`CLAUDE_MIN_SAFE_CONTEXT_TOKENS` (40000, comfortably above the measured
~36.4K overhead), the response is `{requiresContextWarning: true, modelId,
contextLength, minSafeContextTokens}` instead of applying — nothing is
restarted or created yet. A context length that was never discovered at
all skips the check entirely (nothing to compare, so it fails open rather
than warning on every model an endpoint hasn't reported a size for).
Retrying with `confirmedContext: true` launches anyway (the legacy `confirmed: true` still means both questions).
The Run-menu picker shows this as an in-app modal
(`#customModelContextWarningModal`, matching the llama-swap conflict
modal's look) naming the model, its discovered context, and the safe
floor, and explaining the fix: reconfigure llama-swap to give that model
(or a smaller one) an explicit larger context instead of relying on
auto-fit (`--fit-ctx`), which optimizes for the biggest _model_ that fits
rather than the biggest _context_ — e.g. adding `-c 65536` (or as large a
`--ctx-size` as the hardware holds) to that model's llama-swap config
entry. A smaller model at a much larger explicit context often fits in
the same VRAM a bigger model's auto-fit context gets shrunk to make room
for.
Clear back to the harness's native cloud default with:
```bash
@@ -101,6 +473,14 @@ id, model and injected key NAMES are persisted, the values are re-derived
from the endpoint store on recovery, and the pane keeps running against the
endpoint in between because tmux retains its environment.
⚠️ Clearing removes injected keys **by name**, and `CLAUDE_CONFIG_DIR` is one
of the names claude's selection injects — so a session that ALSO had
`CLAUDE_CONFIG_DIR` set through the generic `envOverrides` field (the
per-client-account case) loses that override on clear too, and silently
falls back to the server's default Claude account. If you route a session
to a specific account this way, re-apply the override after clearing a
custom-model selection from it.
**New sessions always default back to the harness's native backend.** A
custom-endpoint selection is a per-session choice, never a sticky global
default — starting a fresh session doesn't inherit whatever the last one was
@@ -115,17 +495,53 @@ automatically). Results:
- **Claude, opencode, Pi, Grok, OMP** — verified: a real "hello world" reply
came back through the endpoint.
- **Codex** — the config is structurally correct, but Codex only speaks the
Responses API since Feb 2026, which llama.cpp/llama-swap don't implement.
This is a real protocol incompatibility, not a bug here; Codex support
needs a Responses-API-compatible endpoint.
- **Codex** — the config is structurally correct, and against a llama-swap
server that DOES answer `/v1/responses` (confirmed live: a plain,
no-tool-call chat turn returned a real reply), the picture is more
nuanced than a flat failure. A real tool-call attempt (`run the shell
command: echo hello`) came back as `agent_message` TEXT — literally the
tool-call JSON printed as the model's answer — instead of a
`function_call` item Codex would actually execute (confirmed via `codex
exec --json`'s raw event stream). So plain chat can work while the thing
that makes Codex a coding agent — actually running commands and editing
files — does not; treat Codex as still unreliable for real work against a
llama.cpp/llama-swap endpoint, tool-calling gap included, not just the
earlier-documented `wire_api` mismatch (which not every deployment hits
the same way — some legitimately have no `/v1/responses` route at all).
Separately, EVERY custom-endpoint Codex session prints `Model metadata
for '<id>' not found. Defaulting to fallback metadata...` on launch —
confirmed harmless (the reply above still came back correctly): Codex's
model metadata (reasoning-tier options, per-model system-prompt
templates, context-window figures) comes from `models_cache.json`, a
local cache of OpenAI's own hosted model catalog that a custom local
model can never appear in by construction, since it isn't one of
OpenAI's models. There's no config.toml override for a model's metadata,
and fabricating a fake catalog entry would mean copying the _shape_ of
OpenAI's own proprietary schema (their per-model system-prompt content
included) for a warning that doesn't otherwise affect behavior — not
something to build into discovery.
- **Gemini** — fails with `Invalid auth method selected`, traced to an
undocumented `GATEWAY` auth path gemini-cli selects once
`GOOGLE_GEMINI_BASE_URL` is set. Unresolved after real investigation
(several auth workarounds were tried and ruled out); do not rely on
Gemini support yet.
- **DeepSeek** — the request reaches the server (env vars are read) but
gets a consistent `HTTP_404`. Root cause not identified; best-effort only.
- **DeepSeek** — root cause of the `HTTP_404` found and fixed. DeepSeek
Harness's own bundled provider module (`@deepseek-ai/dsh-llm-deepseek`)
builds its request URL as `${DEEPSEEK_BASE_URL}/chat/completions` with no
`/v1` insertion of its own (its real public API, `https://api.deepseek.com`,
expects the caller's base URL to already carry any needed prefix) —
confirmed by reading its own source and, live, that
`POST <baseUrl>/chat/completions` 404s against llama-swap while
`POST <baseUrl>/v1/chat/completions` succeeds; the harness's own error
template (`DeepSeek API error (HTTP ${status})`) matches the originally
reported symptom exactly. `customModelInjection`'s new `appendV1Suffix`
(deepseek's entry only — claude/gemini must NOT get it, since claude was
already confirmed working against the raw `baseUrl`) fixes it by writing
`DEEPSEEK_BASE_URL` with `/v1` appended. Not yet re-run end-to-end with a
real `dsh` binary (no install available in this environment) — the fix
is source-confirmed and live-verified at the HTTP level, but a real
"hello world" reply through `dsh` itself is still outstanding before
calling this fully verified like the harnesses above.
- **Antigravity** — no known custom-endpoint mechanism at all; unsupported.
See the confidence table in `custom-model-endpoints-plan.md` for the full detail behind
+3 -1
View File
@@ -53,7 +53,9 @@ that spawns a literal `pnpm` with no npm fallback, so without one it exits 127 w
surfaces that same line as the install error. `npm install -g pnpm` (or
`corepack enable pnpm`) is the fix. This is what broke the Docker agent image in
[#352](https://github.com/Ark0N/Codeman/issues/352); the image now installs pnpm
alongside `dsh`.
alongside `dsh`. The Compose server image (`docker/server.Dockerfile`) does not
ship `dsh`, since it is installed at runtime, but it does ship pnpm so the UI
button works there too.
Codeman's default is `@deepseek-harness-tui/dsh-tui` because it is by a wide
margin the most used community TUI, it is MIT, and it implements the status
+2
View File
@@ -81,6 +81,8 @@ Antigravity (`agy`) and Grok (`grok`) are the two CLIs not installed from npm (G
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md). Grok is seeded per-file for the same reason (`auth.json`, `config.toml`, `pager.toml` out of `~/.grok`, which also holds `sessions/`, `memory/` and the ~160MB binary under `downloads/`), with the same consequence for `grok -c`. See [`grok-integration.md`](./grok-integration.md). OMP is the one CLI in this family where `sessions/` is the EXCEPTION rather than the rule: `~/.omp/agent/{config.yml,mcp.json,models.yml,settings.yml}` are seeded per-file (the dir also holds SQLite caches and `terminal-sessions/`), but `~/.omp/agent/sessions/` is shared RW like codex's, not seeded, because Codeman reads it host-side for history recovery and `--resume` pinning. See [`omp-integration.md`](./omp-integration.md).
The image can also carry the GitHub CLI (`gh`) and the Azure CLI (`az` + the `azure-devops` extension, in `AZURE_EXTENSION_DIR=/opt/az-extensions` so it stays out of the seeded HOME), wired into the system git config as credential helpers for github.com and dev.azure.com / *.visualstudio.com, exactly as in `docker/server.Dockerfile`. Their sign-ins are seeded per-FILE like pi's: `~/.config/gh/{hosts.yml,config.yml}` and `~/.azure/{azureProfile.json,msal_token_cache.json,service_principal_entries.json,clouds.config,config}`, never `~/.azure`'s logs, command index or extensions. A token kept in a desktop keyring, or in the encrypted MSAL cache az uses on Windows/macOS, is not in those files and does not carry. None of the three is version-pinned; the `--no-cache` rebuild recommended above is also what refreshes them. Both CLIs are opt-in and OFF by default: `CODEMAN_AGENT_IMAGE_INSTALL_GH=1` / `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` in the environment of `scripts/build-agent-image.mjs`, or of the Codeman server for its own auto-build (in the Compose deployment, `environment:` in `docker-compose.override.yml`), become the `CODEMAN_INSTALL_GH` / `CODEMAN_INSTALL_AZ` build args and put that CLI, its extension and its helper entry into the image. Unset passes nothing, so a default build's argv is unchanged and the image has neither. The sign-in seeds follow the same switches, read when a case container is created: `.config/gh` only with `CODEMAN_AGENT_IMAGE_INSTALL_GH=1`, `.azure` only with `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` (`enabledByEnv` in `CRED_STORES`), never merely because the files exist. Seeds are create-time mounts and deliberately not part of the config hash (hashing them would trip the drift gate for every case), so an existing case container picks them up only when it is recreated.
## Quickest path: one-click "Run in Docker"
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
+3 -1
View File
@@ -6,6 +6,8 @@ For the Compose configuration, environment settings, storage migration, and macv
The image includes Claude Code, Codex, Gemini CLI, and OpenCode. Authenticate a CLI from its Codeman session; credentials are never baked into the image.
It can also include the GitHub CLI (`gh`) and the Azure CLI (`az`) with the `azure-devops` extension, wired in as Git credential helpers, so Clone Repo and `git clone` reach private GitHub and Azure DevOps repositories once they are signed in. Both are off by default; [Turning them on](../docker/README.md#turning-them-on) shows the `docker-compose.override.yml` settings.
## Prerequisites
- Docker Engine or Docker Desktop with Docker Compose v2
@@ -67,7 +69,7 @@ If that directory was created by an earlier root-running image, change its owner
Codeman updates itself from **App Settings → Updates**, as it does on a bare host. The checkout mounted at `/opt/codeman` is the same directory Compose builds from, so the update's `git checkout` and rebuild land on the host and survive container recreation; the restart is the server exiting, which `restart: unless-stopped` turns into a relaunch on the new build.
That applies application code only. A release that changes `docker/server.Dockerfile`, `docker/docker-compose.yaml`, or adds a key to `docker/.env.example` needs the image rebuilt or the container recreated, which a container cannot do to itself. The updater detects each case and refuses with a message naming what changed; run `docker/Start-Codeman.sh` on the host to apply those.
That applies application code only. A release that changes `docker/server.Dockerfile`, `docker/docker-compose.yaml`, or adds a key to `docker/.env.example` needs the image rebuilt or the container recreated, which a container cannot do to itself. The updater detects each case and refuses with a message naming what changed; run `docker/Start-Codeman.sh` on the host to apply those. For a major update, or a base-image change `Start-Codeman.sh` does not fully pick up, `docker/Update-Codeman.sh` rebuilds with no layer cache and clears the two build-artefact volumes before handing off to it (see "Major updates" in `docker/README.md`).
`CODEMAN_REPO_PATH` overrides which checkout is mounted. It defaults to the compose project's parent directory, so it normally needs no setting. Point it at a directory that is not a git checkout and in-app updates are reported as unavailable.
+14 -8
View File
@@ -10,15 +10,18 @@ this file covers only what the container changes.
## The short version
| Change in the release | Applied by |
| -------------------------------- | ------------------------------------------------ |
| Application code | The in-app updater |
| `docker/server.Dockerfile` | `docker/Start-Codeman.sh` on the host |
| `docker/docker-compose.yaml` | `docker/Start-Codeman.sh` on the host |
| New key in `docker/.env.example` | Add it to `docker/.env`, then `Start-Codeman.sh` |
| Change in the release | Applied by |
| ----------------------------------------- | ------------------------------------------------------------------------------- |
| Application code | The in-app updater |
| `docker/server.Dockerfile` | `docker/Start-Codeman.sh` on the host |
| `docker/docker-compose.yaml` | `docker/Start-Codeman.sh` on the host |
| New key in `docker/.env.example` | Add it to `docker/.env`, then `Start-Codeman.sh` |
| A major update, or a Node base-image bump | `docker/Update-Codeman.sh` on the host (no-cache rebuild + fresh build volumes) |
The in-app updater detects all three of the bottom rows itself and refuses with a
The in-app updater detects the three middle rows itself and refuses with a
message naming what changed, so you never have to work out which case you are in.
`Update-Codeman.sh` is the heavier option for when `Start-Codeman.sh` is not
enough: see "Major updates" in `docker/README.md`.
## Why the container needs its own path
@@ -224,7 +227,10 @@ the host and the in-app path works from then on.
**Resetting the build artefacts** — `docker compose down -v`, then
`Start-Codeman.sh`. This discards the named volumes and re-seeds them from a fresh
image.
image. `docker/Update-Codeman.sh` scripts the same reset by default for the two
build-artefact volumes (`codeman-node-modules`, `codeman-dist`) only, plus an
unconditional `--no-cache` rebuild, which a plain `Start-Codeman.sh` run does not
force on its own. See "Major updates" in `docker/README.md`.
## Disabling it
+380
View File
@@ -0,0 +1,380 @@
# Installer v2: three questions, then a URL you can open on your phone (Plan)
Status: **Phase 1 IMPLEMENTED (2026-09-20)**, phases 2 and 3 open. It builds on
`docs/tailscale-installer-plan.md` (implemented 2026-08-04), which made Tailscale a
guided option; this round makes it the thing the install ENDS on, and makes the whole
installer shorter to sit through. Owner decisions taken before implementation: rename
is opt-in and **defaults to no everywhere** (the machine name is used for other things);
the URL keeps the node name unless asked; `codeman-<hostname>` is the suggested name;
sub-path is the default for an occupied `:443`.
Verification record for phase 1 (all on the maintainer's box, 2026-09-20):
- `test/install-sh-invariants.test.ts` (28 tests, incl. the new Tailscale safety pins)
and the detection-parity test pass; `bash -n` passes.
- Every new decision function driven with stubbed tailscale state under **bash 5.2 and
bash 3.2** (the `bash:3.2` container CI uses): flags, the launch default, the serve
shape for free / ours / occupied `:443` (all four answers plus the non-interactive
default), the three serve commands, the rename question (Enter keeps the name; `--yes`
and non-interactive never rename; `codeman-*` nodes are skipped; `--name` is
sanitized), `run_step` success/failure/stdin, the unit round-trip of
`CODEMAN_BASE_URL`/`CODEMAN_PORT`/an escaped password, and the done screen.
- A full non-interactive install into a sandboxed `HOME` with `CODEMAN_TAILSCALE=1`:
preflight summary, kept the existing prod mapping (no serve mutation), clone 2 s,
`npm install` 18 s, build 23 s, symlink, done screen; `install.sh status` on a pty
renders the QR code. Nothing on the real system changed.
- **Sub-path mode end to end over the real tailnet**: an isolated Codeman
(`CODEMAN_INSTANCE`, port 3999, `--base-url /codeman`) behind
`tailscale serve --https=8445 --set-path /codeman 3999` answered `/codeman/api/status`,
`/codeman/` (with `<base href="/codeman/">` and `__CODEMAN_BASE__="/codeman"`), the
hashed CSS/JS, `/codeman` without a slash, and the SSE stream; mapping and server
removed afterwards. **Correction to section 2**: serve STRIPS the mount prefix
before proxying (a direct `/codeman/api/status` on the server is 404 while the same
path through serve is 200). That is fine because Codeman's ingress tolerates
unprefixed requests; `--base-url` is needed for the URLs Codeman EMITS, not for
what it receives.
- Not yet exercised on a fresh machine (unchanged from the previous plan): Tailscale
absent / logged out / HTTPS toggle off, the rename against a real node (the
off-rename-re-add order is implemented but only unit-driven), macOS, uninstall. The
Mac mini and a throwaway VM are the venues; see section 8.
- **Review fixes (2026-09-21)**, from the two reviews on PR #460 (DeepSeek Harness, then
Claude): the done screen's Start line is composed from every non-default value
(`start_command_hint`, shared with the exec branch as `export_bind_env`), so "do not
start" under a sub-path or a custom port no longer prints a bare `codeman web`; the
`--lan`/`--tailscale`/env preset paths keep an existing password instead of rewriting
the unit open; `--password`/`--port` flip `RECONFIGURE` so they reach the unit;
`install.sh name` re-syncs the unit's base URL after a rename; the sudo keepalive is
ended before the `exec` into the foreground server; Ctrl+C in the HTTPS-toggle poll
skips Tailscale instead of killing the run; `uninstall` asks before removing a
LaunchDaemon it never wrote; a foreign LaunchDaemon gets a restart hint and the done
screen stops claiming the new build is running; the preflight summary reads the
Tailscale state without node; the LAN security notice uses the configured port; a
bare re-run ends on the done screen; a build failure after a rename names the
`install.sh tailscale` recovery; `TS_JOINED_HERE` is gone.
Goal, in one sentence: a user runs the one-liner, answers at most three questions, walks
away during the build, and comes back to `https://<name>.<tailnet>.ts.net` printed with a
QR code, already answering, on every device in their tailnet. That is exactly the
maintainer's own production setup (`tnode.tailf80371.ts.net` fronting `127.0.0.1:3000`),
and the installer should produce it without the user knowing what `tailscale serve` is.
## 1. Where the installer is today
Facts from reading `install.sh` (2886 lines, 19 `prompt_yes_no` sites) and the live
Tailscale state on the maintainer's box (tailscale 1.102.2, user-owned node, MagicDNS +
HTTPS certs on, serve mapping `443 -> https+insecure://localhost:3000`).
**The order is backwards for a human.** The flow is: detect -> ask about git -> ask about
node -> ask about tmux -> ask about build tools -> AI CLI menu -> ask about cloudflared ->
clone -> `npm install` -> build (minutes) -> **then** the network-access question -> the
Tailscale sub-steps (install? login URL, sudo for operator, admin-console toggle loop) ->
the launch menu (no default; a bare Enter re-prompts) -> tunnel-service question. A fresh
Ubuntu server taking the Tailscale route answers roughly ten prompts plus two to four sudo
password prompts, split around a multi-minute build. The user cannot walk away at any
point, and the question that matters most (how do I reach it) comes last.
**The Tailscale flow works but was never exercised on a fresh machine.** The previous
plan's manual matrix still lists items 1-4, 7 and 10-12 (Tailscale absent, logged out,
HTTPS toggle off, port 443 occupied, macOS, uninstall, phone PWA) as untested. The
maintainer's own verification was the idempotent "kept as-is" path.
**The URL is the machine's name, full stop.** `setup_tailscale_serve` derives it from
`.Self.DNSName`, and nothing lets the user influence it. A second Codeman on the same
tailnet is `macminis-mac-mini.tailf80371.ts.net`, which tells you nothing about Codeman.
**Port 443 taken means give up or clobber.** If another app already owns the root of
`:443`, the only offer is "replace it?" (default no), and declining falls back to
local-only. Codeman already supports running under a sub-path (`--base-url`), and
Tailscale serve supports mounting a path (`--set-path`), so there is a third answer nobody
is offered.
**The result is invisible afterwards.** Once the terminal scrolls away, nothing in the app
or the CLI tells the user their Tailscale URL again. `codeman doctor` does not probe
Tailscale; App Settings -> Remote access shows only the Cloudflare tunnel.
**Two service writers exist.** `install.sh` carries its own plist/unit generator (~180
lines) next to `codeman service install` (`src/service-installer.ts`). They agree on the
job name by design, but the bash copy is the one that writes `CODEMAN_PASSWORD` into the
unit, so they cannot simply be merged. Left as-is in this plan (see section 9).
## 2. What Tailscale makes possible for the name (researched 2026-09-20)
| Option | Resulting URL | What it needs | Side effects | Verdict |
| ------ | ------------- | ------------- | ------------ | ------- |
| **A. Node name** (today) | `https://tnode.tailf80371.ts.net` | `tailscale serve --bg 3000` | none | **Default.** Zero admin-console work, matches the maintainer's prod. |
| **B. Rename the node** | `https://codeman-tnode.tailf80371.ts.net` | `tailscale set --hostname codeman-<host>` (operator or root) | Renames the machine tailnet-wide: ssh targets, other serve URLs, the admin console entry. Tailscale de-dups a clash as `-1`. The cert follows the new name. | **Opt-in, default NO everywhere** (owner decision 2026-09-20: the machine is used for other things, so a bare Enter never renames it). The proposal was YES when the installer itself had just joined the tailnet; rejected. |
| **C. Tailscale Service** | `https://codeman.tailf80371.ts.net` | tailscale >= 1.86 on the host; the host must have a **tag-based identity** ("You cannot use a device authenticated with a user account as a Service host"); the service is defined in the admin console first; the host is then approved there (or via `autoApprovers.services`). Public beta since 2025-10-28, all plans. | Re-authenticating a personal machine as a tagged node changes its identity (SSH ACLs, user attribution). Known daemon quirk: approval is not picked up until `serve clear` + re-advertise (tailscale/tailscale#18821). | **Detect and hint only** in this round. The maintainer's own node has `Self.Tags: null`, so it could not host one without re-tagging. Worth a real flow once someone with a tagged fleet asks. |
| **D. Sub-path** | `https://tnode.tailf80371.ts.net/codeman` | `tailscale serve --bg --set-path /codeman 3000` plus `--base-url /codeman` on the server | Codeman runs under a prefix. Hooks are unaffected (they hit the raw port with no prefix, which `rewriteUrl` already tolerates). Serve forwards the prefix unchanged, which is exactly the shape `--base-url` was built for. | **The answer when `:443` root is already taken.** Replaces today's replace-or-nothing prompt. |
| **E. Second port** | `https://tnode.tailf80371.ts.net:8443` | `tailscale serve --bg --https=8443 3000` | Port in the URL; the beta-preview recipe already uses this. | Fallback when the user rejects D. |
| Funnel (public internet) | `https://tnode.tailf80371.ts.net` from anywhere | `tailscale funnel` | Public exposure; different risk class. | **Out of scope**, as before. Docs only, with the password warning. |
Sources: Tailscale Services docs (`tailscale.com/docs/features/tailscale-services`), the
Services beta announcement (`tailscale.com/blog/services-beta`), machine names
(`tailscale.com/kb/1098/machine-names`), the serve CLI reference
(`tailscale.com/docs/reference/tailscale-cli/serve`), the macOS variants page
(`tailscale.com/docs/concepts/macos-variants`), and `tailscale serve --help` on 1.102.2
(which lists `--service`, `--set-path`, `--yes`, `advertise`, `get-config`/`set-config`).
**Trap for option B (verify on the Mac mini before shipping):** the serve config is keyed
by `host:port` using the DNS name at configuration time (`"Web": {"tnode.tailf80371.ts.net:443": ...}`
in `serve status --json`). Renaming a node after serve is configured most likely orphans that
entry: the handler lookup uses the current name and never matches the old key, and the only
tool that removes a stale key is `serve reset`, which this installer must never run. So the
order is **rename first, then configure serve** on a fresh install, and on a retrofit
(`install.sh name`) **turn our mapping off, rename, wait for `.Self.DNSName` to change,
re-add**.
## 3. Target UX
### 3.1 Three questions, then walk away
```
Codeman installer
Found: git, Node 22.14, tmux 3.4, build tools Missing: nothing
AI CLIs: Claude Code (~/.local/bin/claude)
Tailscale: connected as tnode (tailf80371.ts.net)
Existing: none
1/3 How should the dashboard be reachable?
1) Tailscale https://tnode.tailf80371.ts.net (recommended, already connected)
2) Any device on your network (0.0.0.0, password required)
3) This machine only (127.0.0.1)
Choose [1/2/3] (default 1):
2/3 Name this machine "codeman-tnode" on your tailnet? [y/N]
(only shown for option 1; default no, always)
3/3 Run Codeman as a background service that starts on boot? [Y/n]
Installing… this takes a few minutes. You can leave this running.
✓ dependencies ✓ clone ✓ build (2m 41s) ✓ service ✓ tailscale serve
```
Rules that make this work:
- **Every step that needs a human runs BEFORE the build.** The dependency consent, the
AI CLI menu, the Tailscale install consent, the `tailscale up` login URL, the operator
grant, and the tailnet HTTPS toggle all move into the question phase. The build, the
service, `tailscale serve` and the verification are unattended.
- **One consent for all missing system packages.** "Install git, Node 22 and build tools
now? [Y/n]" replaces four separate prompts. Each package still runs its own
distro-specific installer.
- **One sudo prompt.** When anything needs root (packages, the Tailscale installer,
`tailscale up`, the operator grant), the installer says so once, runs `sudo -v`, and keeps
the timestamp alive in a background loop until it exits. macOS needs no sudo for the
Tailscale GUI-app CLI and the pattern still holds for Homebrew packages.
- **Service is the default.** Enter on the last question installs the service; "run in
this terminal" and "don't start" stay reachable by answering, and by flag.
- **The cloudflared question is gone from the main flow.** It is optional, defaults to
no, and has an in-app toggle (App Settings -> Remote access). The done screen mentions it
only when `cloudflared` is already installed. The Linux tunnel-service prompt goes with it.
- **The HTTPS-certificates toggle no longer asks "re-check now?"** The installer prints the
admin URL, opens it in a browser when one is available (`xdg-open` / `open`, never on a
headless box), and polls `tailscale status --json` every 5 s for up to 5 minutes. Ctrl+C or
the timeout falls back exactly as today.
- **Progress, not silence.** `npm install` and `npm run build` run behind one line each
with elapsed time; their output goes to `~/.codeman/install.log` and is printed only on
failure, with the exact retry command.
### 3.2 The done screen
One block, the URL first, a QR code the phone can scan, and nothing the user does not need
right now.
```
✓ Codeman 1.31.0 is running
Your tailnet: https://codeman-tnode.tailf80371.ts.net (HTTPS, any of your devices)
This machine: http://localhost:3000
▄▄▄▄▄▄▄ ▄ ▄▄ ▄▄▄▄▄▄▄
█ ▄▄▄ █ ▄▄▀ ▄ █ ▄▄▄ █ scan with your phone
█ ███ █ ███▀▀ █ ███ █
█▄▄▄▄▄█ █ ▄ █ █▄▄▄▄▄█
Manage systemctl --user restart codeman-web · journalctl --user -u codeman-web -f
Update re-run the install line, or App Settings → System → Updates
Docs https://github.com/Ark0N/Codeman/wiki
Security: Codeman binds 127.0.0.1. Tailscale authenticates every device before a
packet reaches it. Details: docs/security-architecture.md
```
The QR comes from the `qrcode` package Codeman already depends on
(`node -e "require('qrcode').toString(url, {type:'terminal', small:true}, …)"` from
`$INSTALL_DIR`, verified locally: 17 rows by 45 columns). Skipped when the terminal has no
color support or fewer than 50 columns. The QR encodes the plain URL, not an auth token:
the tailnet is the login.
### 3.3 Express mode and flags
Env vars stay (`CODEMAN_TAILSCALE=1`, `CODEMAN_HOST`, `CODEMAN_PASSWORD`,
`CODEMAN_NONINTERACTIVE=1`, `CODEMAN_PORT`). Flags are added because they are
discoverable from the one-liner and pipe through `bash -s --`:
```bash
curl -fsSL https://getcodeman.com/install | bash -s -- --tailscale --service
curl -fsSL https://getcodeman.com/install | bash -s -- --lan --password 'x' --service
curl -fsSL https://getcodeman.com/install | bash -s -- --local --run
curl -fsSL https://getcodeman.com/install | bash -s -- --tailscale --name codeman-build --yes
```
| Flag | Meaning |
| ---- | ------- |
| `--tailscale` / `--lan` / `--local` | Answer 1/3 (same semantics as `CODEMAN_TAILSCALE=1`, `CODEMAN_HOST=0.0.0.0`, `CODEMAN_HOST=127.0.0.1`) |
| `--name <n>` / `--no-rename` | Answer 2/3: rename the node to `<n>`, or never ask |
| `--service` / `--run` / `--no-start` | Answer 3/3 |
| `--yes` | Accept every default, still prompt for a login URL (a human must open it) |
| `--password <p>` | Same as `CODEMAN_PASSWORD` |
| `--port <n>` | Same as `CODEMAN_PORT`; the serve target follows it |
`--yes` differs from `CODEMAN_NONINTERACTIVE=1`: it is the interactive user saying "I trust
the defaults", so it may install software and may wait on a login URL. Non-interactive stays
the CI contract and never installs Tailscale.
## 4. The Tailscale flow, v2
The state machine from the previous plan stays; these are the changes.
1. **Preflight, before the build** (`tailscale_preflight`): installed? -> install
(Linux: official script; macOS: brew cask, else download link and wait). Logged in? ->
`tailscale up` with the URL printed prominently and a 5-minute poll. Operator (Linux):
grant once under the single sudo session. HTTPS certs: poll instead of ask (Ctrl+C
during the poll skips Tailscale for this run rather than ending the installer). The
rename default does not depend on whether this run performed the login (decided NO
everywhere), so nothing records it.
2. **Name** (`tailscale_choose_name`, question 2/3): shown only on the Tailscale route.
Default `codeman-<oshostname>` sanitized to `[a-z0-9-]`, max 63. Applied with
`ts_cmd_serve set --hostname`, then poll `.Self.DNSName` until it carries the new name
(up to 60 s). Order matters: this runs before any serve mutation (section 2 trap).
Declining keeps the node name. On a re-run against a node already named `codeman-*`,
the question is skipped.
3. **Serve, after the service is up** (`setup_tailscale_serve`): unchanged idempotent
"kept as-is" path first. When `:443` root belongs to another target, the new prompt is:
```
tailscale serve already sends https://tnode.tailf80371.ts.net to port 8080.
1) Add Codeman under a path: https://tnode.tailf80371.ts.net/codeman (default)
2) Use another port: https://tnode.tailf80371.ts.net:8443
3) Replace the existing mapping with Codeman
4) Skip Tailscale for now
```
Option 1 writes `--base-url /codeman` into the service unit (it is a `WebLaunchOptions`
field already, and `buildWebArgs` carries it) and runs
`tailscale serve --bg --set-path /codeman <port>`. Option 2 runs `--https=8443`.
`detect_tailscale_serve_url` learns to recognize all three shapes (root, path, port) so
uninstall, the security notice and the re-run default keep working.
4. **Warm the certificate.** Right after serve is configured, fire one background
`curl -sk https://<url>/api/status` so Let's Encrypt issuance overlaps the rest of the
install instead of adding 30 s to the verify step.
5. **Verify** as today (200 or 401 on `/api/status`), with the path-aware URL.
6. **Services hint** (option C): when `.Self.Tags` is non-empty and `serve --help`
lists `--service`, the done screen adds one line: "This is a tagged node, so it can also
host `https://codeman.<tailnet>.ts.net` as a Tailscale Service: see Remote Access in the
wiki." No flow, no prompt.
7. **macOS**: the App Store and Standalone variants cannot run before login, so a
LaunchAgent plus serve only comes back after someone logs in. The done screen says so on
macOS. The Mac mini (`arbbot`, headless, system LaunchDaemon) is the reference for the
"headless Mac" caveat, and `install.sh` must keep refusing to replace a LaunchDaemon it
did not write (today it removes one; that is a bug for the Mac mini and is fixed here:
detect `UserName` in the daemon plist and leave it alone with a message).
8. **Uninstall** additionally offers to restore the original node name when this installer
renamed it (the original is recorded in `~/.codeman/install.json`, the one marker file
this feature adds, because tailscaled does not remember previous names).
9. **Subcommands**: `install.sh tailscale` (unchanged purpose, now runs the v2 flow),
`install.sh name [<n>]` (rename with the off/rename/re-add dance), `install.sh status`
(prints the done screen again, URL and QR included, for the "what was my URL" moment).
## 5. In-app: the URL stays discoverable
Small, read-only, and the first server-side code this feature has ever needed.
- **`GET /api/system/remote-access`** returns
`{ tailscale: { installed, connected, dnsName, url, mode: 'root'|'path'|'port'|null } }`
by running `tailscale status --json` and `tailscale serve status --json` through
`execFile` with the existing exec timeout, cached 30 s, resolved through the same
`get_tailscale_path` search as the installer (PATH, then the macOS app bundle), and a
no-op under `VITEST` like every other IO probe. Never mutates serve config.
- **App Settings -> Remote access** gains a **Tailscale** row above the Cloudflare toggle:
the URL as a copy chip, a QR button reusing `showTunnelQR`'s modal, and when nothing is
configured a one-line hint with `bash ~/.codeman/app/install.sh tailscale`. The welcome
screen's "open on your phone" affordance shows the same QR.
- **`codeman doctor`** grows a `tailscale` entry under `other` in
`config/dependency-registry.ts`: installed, connected, serving Codeman (URL). Pure
engine, injectable probe host, like the existing rows.
- No new SSE event, no settings key, no state.json change.
## 6. Security posture
Nothing widens. The bind stays loopback; the tailnet is the authentication boundary;
`.ts.net` is already in `DEFAULT_TRUSTED_HOST_SUFFIXES`. New surfaces are read-only
probes. `install.sh` still never runs `tailscale serve reset`, still touches only the
mapping it created, and gains one more never: it never advertises a Tailscale Service or
runs `tailscale funnel`. The sudo keep-alive loop is killed by the existing `cleanup` trap.
The rename records the previous name locally and offers the reversal at uninstall.
## 7. Implementation inventory
| File | Change |
| ---- | ------ |
| `install.sh` | New `parse_flags`, `preflight_summary`, `ask_everything` (the three questions), `sudo_session`, `run_step` (spinner + log), `tailscale_preflight`, `tailscale_choose_name`, `tailscale_rename_node`, `print_done_screen`, `print_qr`, `status` subcommand, `name` subcommand. Modified: `main` (reordered into ask -> work -> done), `choose_network_binding` (question 1/3, same defaults), `setup_tailscale_serve` (path/port options), `detect_tailscale_serve_url` (three shapes), `setup_systemd_service`/`setup_launchd_service` (`--base-url`, LaunchDaemon guard), `uninstall` (rename reversal), header docs (flags). Removed from the main flow: the cloudflared prompt, the tunnel-service prompt. bash 3.2 rules unchanged. |
| `src/web/routes/system-routes.ts` | `GET /api/system/remote-access` |
| `src/tailscale-status.ts` (new) | Pure parser for the two JSON shapes + the IO wrapper; unit-tested against captured `serve status --json` fixtures (root, path, port, foreign target, none) |
| `src/config/dependency-registry.ts`, `src/utils/dependency-checker.ts` | `tailscale` doctor row |
| `src/web/public/index.html`, `settings-ui.js`, `panels-ui.js` | Tailscale row + QR, welcome-screen QR |
| `test/install-sh-invariants.test.ts` | Extend: flags documented in the header, no `serve reset`, no `funnel`, no `--service` advertise, every serve mutation goes through `ts_cmd_serve`, rename happens before serve in `main` (static order check) |
| `.github/workflows/ci.yml` | The bash 3.2 step additionally sources the script with stubbed `ts_cmd`/`ts_cmd_serve`/`read_reply` and drives `ask_everything` through all three answers and the 443-occupied menu |
| `test/tailscale-status.test.ts`, `test/routes/system-routes-remote-access.test.ts` | Parser + route |
| Docs | README install + remote-access sections, `docs/wiki/Installation.md`, `Remote-Access.md` (naming options table, Services caveat, path/port variants), `Mobile-Guide.md`, `Running-As-A-Service.md` (macOS login caveat), `FAQ.md`, `docs/security-architecture.md` §A, CLAUDE.md Scripts & Tunnel paragraph, `docs/tailscale-installer-plan.md` gets a pointer here. getcodeman.com copy lives outside the repo (maintainer handbook). |
Changeset: `minor` (new flags, new subcommands, new API route).
## 8. Test plan
Automated (the gate): the static invariants above, the bash 3.2 container drive of the
question phase, the JSON parser fixtures, the route test.
Manual matrix, on a fresh Ubuntu 24 VM and on the Mac mini, since the previous plan's
items never ran on a fresh machine:
1. Tailscale absent, declined -> local-only, done screen shows the retrofit command.
2. Tailscale absent, accepted -> install, login URL, operator, certs toggle polled, rename
question shown (default no), service, serve, URL verified, QR scans on a phone, PWA installs.
3. Tailscale present and logged in on a pre-existing node -> rename default NO, URL is the
node name, `serve status` gains exactly one entry.
4. `:443` root occupied -> path option -> `https://<node>/codeman` answers, hooks still
fire (raw port), `install.sh status` prints the path URL.
5. Rename on a node that already has our serve mapping (`install.sh name`) -> off, rename,
re-add, `serve status` has no stale key.
6. Re-run the one-liner -> quiet update, binding and name preserved, no prompts.
7. `--yes` end to end; `CODEMAN_NONINTERACTIVE=1` end to end (no software installed).
8. Uninstall -> mapping removed, other mappings intact, rename reversal offered.
9. Mac mini: LaunchDaemon left alone with the message; done screen carries the login caveat.
## 9. Phasing and open decisions
**Phase 1 (this round):** the reorder, the three questions, one consent + one sudo, flags,
the done screen with QR, Tailscale preflight-before-build, the path/port answer for an
occupied 443, the rename step, `status` and `name` subcommands, docs.
**Phase 2:** the in-app Tailscale row + QR, `codeman doctor` row, the `remote-access`
route. Independent of phase 1 and useful on its own for existing installs.
**Phase 3 (optional):** replace the bash service writers with `codeman service install`
once that command can carry `CODEMAN_PASSWORD` behind an explicit flag; and a Tailscale
Services flow if a tagged-fleet user asks for `codeman.<tailnet>.ts.net`.
Decisions for the maintainer:
1. **Rename default.** Decided 2026-09-20: always NO; the yes answer, `--name` and
`install.sh name` are the ways in. (The proposal was YES only when this run had joined
the tailnet, NO otherwise; rejected because the host is used for other things.)
2. **Name pattern.** `codeman-<hostname>` (proposed; unique per machine, and two Codemans
on one tailnet stay distinguishable) versus plain `codeman` (nicer once, collides on the
second install, Tailscale silently appends `-1`).
3. **Path versus port** as the default answer for an occupied 443. Proposed: path, because
the URL has no port and `--base-url` already exists for exactly this proxy shape.
4. **Whether Phase 2 ships in the same release.** It is the part that helps people who
installed months ago.
+27
View File
@@ -66,6 +66,31 @@ each `(clientId, seq)` at most once, so a resend can't type the prompt twice.
(the 200 is the client's ACK). `curl`/legacy callers omit the fields and always
apply.
## Oversized input (issue #484)
Delivery has a third outcome besides "applied" and "retry": **refused for good**.
Both transports refuse a frame longer than `MAX_INPUT_LENGTH` (64 KiB,
`src/config/terminal-limits.ts`; the POST schema uses the same constant). Before
#484 the client treated that like a transient failure, so an oversized paste sat
at the head of the queue, was re-sent every 2 s forever, blocked every later
input for the session, and came back from localStorage on each reload.
- `_sendInputAsync()` splits a paste over the frame limit into in-limit frames
(`CodemanInputLimit.split`, constants.js, never cutting a surrogate pair). They
go out in seq order, so the PTY sees one contiguous stream. A paste over
`PASTE_MAX_CHARS` (1 MiB), or an oversized `useMux` write (line-oriented, never
split), is refused with a toast and never queued.
- The WebSocket answers an oversized sequenced frame with
`{t:'ia', seq, err:'too_large', max}`; the client drops it with a toast. A
client that predates `err` reads it as a plain ACK and drops it too.
- The POST drain drops a frame answered `400`/`413` (`401`/`403` stay transient:
an expired login delivers once the user signs in again).
- `_loadReliableState()` prunes persisted frames over the limit, so a queue
poisoned by an older build heals on the first load after upgrading.
- ⚠️ The frontend limit (`INPUT_FRAME_MAX_CHARS`) and the composer's
`COMPOSER_INPUT_FRAME_LIMIT` must equal `MAX_INPUT_LENGTH`; pinned by
`test/input-size-limit.test.ts`.
## Known limitation
Dedup state is in-memory on the server. A **server restart** between a write and
@@ -79,3 +104,5 @@ across the narrow restart window.
semantics (monotonic, per-client, gap-tolerant, eviction-safe).
- `test/routes/session-routes.test.ts` — POST `/input` applies a tagged
`(clientId, seq)` once on redelivery; untagged input always applies.
- `test/input-size-limit.test.ts`: one input limit on both sides, frame
splitting, and dropping (never retrying) a frame refused for good (#484).
+175
View File
@@ -352,6 +352,177 @@ path but the SESSION (`session.remote`): a remote session never falls back to lo
`fs`, and a local session never opens an ssh connection — including for attachment
records, which are keyed to the session that registered them.
## Wake-on-LAN from user input
A durable remote session survives an SSH drop (COD-104/108), but nothing brought the
HOST back. When the remote machine suspended, the local pane's `ssh` child **stalled**
rather than exited: `tmux send-keys` SUCCEEDS against a stalled pane, so typed input
vanished with no error anywhere, and without a keepalive the pane could look alive for
the OS TCP timeout. The only recovery was waiting for the reconnect watcher, which
gave up after ~13 minutes and, once exhausted, never retried.
An **optional** `wakeMac` (one or more MAC addresses, comma-separated) or `wakeCommand` on a
remote host closes that: on user input, `POST /api/sessions/:id/input` probes the host, and if
it is unreachable it wakes it, polls until the host answers, reattaches the pane
(`Session.reattachRemote()`, which idempotently attaches the still-running remote tmux — the
agent conversation is not restarted), and flushes the input that arrived meanwhile.
Implementation: `src/remote-wake.ts`.
The same wake path also serves **opening** a session, which is where a sleeping host used to
be a dead end: pressing Run on a remote case (`POST /api/quick-start`) or Attach on a
discovered remote tmux session (`POST /api/sessions` + `attachRemoteSession`) probes the host
first, and on a sleeping one wakes it, waits for SSH and only then runs the tmux prereq probe.
Without that the run failed with `could not verify tmux on remote host …` — an ssh error that
blames tmux for a machine that is merely suspended. The wait is **blocking** (the caller gets
the session or the error) but bounded by `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS` (40 s) rather
than the 90 s session default, because the dashboard sits behind a reverse proxy whose default
`proxy_read_timeout` is 60 s: a longer wait would be cut off at the proxy while the session was
still being created. The budget covers the whole request, not just the wait (40 s wake + 1.5 s
probe + the tmux prereq probe's own 15 s timeout = 56.5 s worst case). A host with no wake target is not even probed on this path, so nothing
changes for it, and `remote:hostWaking` is broadcast without a `sessionId` (the toast then reads
"the session starts when it is back" — there is no session yet, and no input queued behind it).
Two wake paths, `wakeCommand` first because it is the explicit override:
- **`wakeMac`** — Codeman builds the magic packet itself (`buildMagicPacket`, six `0xFF`
bytes then the MAC repeated 16×; the shape is asserted byte-for-byte) and broadcasts it
over UDP port 9 (`sendWakePackets`). This is the normal case: no external script, and one
MAC list per host instead of one per consumer.
- **`wakeCommand`** — a single executable path, run WITHOUT a shell. For hosts that need a
router/another machine to send the packet.
**UI**: a banner (`#hostWakeBanner`, `host-wake-ui.js`) appears while the ACTIVE remote
session's host is unreachable — amber, since the Codeman session is healthy and only the
machine is asleep. With a wake target the action is **Wake** (`POST /api/sessions/:id/wake`);
with none it is **Configure WoL** and opens `#wakeConfigModal`, a small form for that host's
`wakeMac`/`wakeCommand` that saves with `PUT /api/remote-hosts/:id` (in multi-user mode that
GET is admin-only, so a non-admin is told the setting is admin-only instead of "host not
found"). Reachability for the banner comes from `GET /api/sessions/:id/reachability`: once
when the remote tab is activated (a user action), and every 30 s while the tab is visible
**only for a host with a wake target** — each poll is a TCP connect to the host, and a timer
that connects to a host Codeman could not wake anyway is exactly the timer-driven traffic
the keepalive rule below rejects (it cannot wake a host, but it can keep an activity-based
suspend timer from firing). A host the probe cannot reach (see the next section) is never
polled. ⚠️ The button is pressed from the SAME
dashboard as Run/Attach, so it holds its request open under the same proxy and uses the same
40 s budget — and it **queues nothing**: browser keystrokes travel over the WebSocket, which
deliberately does not pass through the registry (that is the hot path this feature keeps its
hands off), so the banner says "waiting for the host to come back" for the button and only
claims "input is queued" when the HTTP input path actually buffered bytes
(`queuedInput` on the two SSE events).
**Hosts behind a jump host or SOCKS proxy are reachability-UNKNOWN.** The probe is a bare
TCP connect to `host:port`, and a host reached through `jumpHost`, `socksProxy` or a
`ProxyCommand`/`ProxyJump` in `extraSshOptions` does not answer that even while ssh works —
the direct address may not route at all (the cloudflared case). Acting on the resulting
"unreachable" verdict was wrong three times over: a permanent banner over a healthy session,
a create-path error that replaced a genuine "needs tmux" with "not reachable", and — with a
wake target configured — every HTTP input buffered for the life of the session, because the
readiness poll could never succeed. `isProbeable()` (`remote-wake.ts`) decides from the
proxy fields, which travel on `WakeableRemote`; for such a host the registry delivers input
unchanged, `GET …/reachability` answers `reachable: null, probeable: false` (unknown is not
`false`, and only a proven `false` raises the banner), the create/attach path is not gated
(`ensureHostAwake` → `'unprobeable'`, handled like `'no-target'`), and the quick-start
"not reachable" message is reserved for a **proven** unreachable host (`=== false`). A wake
target can still be fired for it through `POST /api/sessions/:id/wake`, blind: the packet or
command goes out and the response says only whether it did — no readiness poll, no reattach
(the COD-108 watcher owns the pane once ssh works again), no "waking" toast.
The invariants worth keeping:
- **Authorization comes before the wake.** In multi-user mode the attach path
(`POST /api/sessions` + `attachRemoteSession`) answers `403` to a non-admin BEFORE the
host is looked up or probed: remote hosts are admin-only infrastructure everywhere else
(the list is `[]` for a non-admin, write and discovery routes are `adminOnly`), and the
wake spawns the host's `wakeCommand` or broadcasts a packet — a gate that came after the
wake handed an unprivileged account a way to run that executable for any configured
`hostId`, hold the request for the wake budget, and only then be refused for the
workingDir. The quick-start path resolves its remote case through `canAccessOwned`
first. Pinned in `test/routes/session-remote-wake.test.ts` (wake spy stays empty).
- **The caller is told what happened to its bytes.** The non-wait input route answers
`{buffered:true}` when the registry took the chunk and `{buffered:true, dropped:true}`
when it was over the cap and is gone; the send-and-wait route answers `OPERATION_FAILED`
when the host never comes back, like the create and attach paths, instead of writing
into the stalled pane and reporting `delivered:true` plus a timeout. Flushed chunks are
written with `fromUser`, so a first prompt that was buffered through a wake can still
name the tab.
- **Only an EXPLICIT request may wake a host:** user input on an established session, the wake
button, or the user's own session create/attach request (`ensureHostAwake`). Everything that
runs on a TIMER must never wake one — the COD-108 watcher, the server's dropped-session
handler, boot recovery and session discovery have no access to the wake registry, and neither
has the shared session service, because `cron-service.ts` builds sessions there with nobody
waiting on the answer; a wake on such a path would re-wake the host seconds after every
suspend, so it could never stay asleep (the same failure `hufflepuff-mcp-lazy` exists to
prevent for MCP keepalives). A reachability check, a discovery listing and the tmux prereq
probe never wake: they are questions, not actions. All of it is enforced by tests in
`test/remote-wake.test.ts` (two wiring guards: one pins the importers — the route module and
`server.ts`, which holds the registry for its LIFETIME only, `drop()` on session cleanup and
`stop()` on shutdown — and one asserts `server.ts` calls nothing but those two, while
`ensureHostAwake` has exactly one caller file) and `test/routes/session-remote-wake.test.ts`,
not by comments.
- **Detection is a bare TCP connect** to the SSH port (then the configured `port`, else 22),
throttled per session, and only for wake-enabled hosts. No `ServerAliveInterval` is added to
the launch command: keepalives push bytes into an otherwise idle connection every interval,
which is exactly what a byte-threshold idle detector must not count as activity. A probe is
~200 bytes per 30 s, orders of magnitude below any such threshold, and the SYN alone cannot
wake a host.
- **Input is buffered while a wake is in flight** (`REMOTE_WAKE_PENDING_MAX_BYTES`,
oldest whole chunks dropped, bounded so user input cannot grow memory) and flushed in
order after the reattach, with a settle delay so bytes cannot land in a still-connecting
pane. ⚠️ A chunk LARGER than the cap (one big paste is one `input` value) is dropped
**outright**, never trimmed: it was never typed character by character, so its tail is not
"what the user just typed" but a fragment of a command they never sent — the drop is logged
instead. ⚠️ Only the HTTP input route reaches the registry; the **WebSocket keystroke path
is deliberately NOT wake-aware**, so typing into a sleeping host sends nothing and queues
nothing (the banner's Wake button is the recovery for that case, which is why it must not
promise queued input). The **send-and-wait** path blocks on the wake instead — its response
is open anyway, and buffering would break the wait contract. ⚠️ A flush write that FAILS
drops the whole remaining buffer (logged) rather than retaining it: the wake still resolves
and marks the host reachable, so the next input takes the deliver path while a retained
chunk would wait for the NEXT wake — replayed hours later, after everything typed since,
possibly ending in a carriage return. Same policy as the oversized paste.
- **The command runs without a shell** (`spawn(path, [], { stdio: 'ignore' })` — `shell`
defaults to `false`), the schema
requires a single executable path (no arguments, no `$`/backtick), and `wakeMac` is a
structural hex-pair allowlist. A broken or missing wake target fails the wake, never the
input route.
- **`wakeMac`/`wakeCommand` are host-level config, refreshed on recovery AND live**
(`rehydrateRemoteHostFields` in `src/remote-hosts.ts` plus `RemoteWakeDeps.resolveRemote`).
A session's `remote` block is persisted at launch time, so a field added to
`remote-hosts.json` later would otherwise never reach an already-running session — not even
across a Codeman restart, and certainly not right after saving the banner's config dialog.
Recovery rehydration covers restarts, the (throttled, cache-backed) resolver covers the live
session; the host config is authoritative for both (removing the field disables the feature
again). Other host-level fields deliberately stay as persisted, so neither path can
silently re-point an existing pane's SSH options.
- **UI/SSE**: `remote:hostWaking` and `remote:hostWakeFailed` (plus the reused
`remote:sessionReconnected`) drive the banner and toasts, all from `host-wake-ui.js` —
its handlers are the ONLY definitions, since a second one in another mixin would be
silently shadowed by script order. Both carry `queuedInput`, which is true only when the
server actually holds bytes for that session — the wording keys off that, not off "a wake
is running", so the button path never claims input is queued. In multi-user mode the
whole `remote:` family is **session-scoped** (`deriveSseHint`, `server.ts`): an event with
a `sessionId` reaches that session's owner, and the create/attach wake — which has no
session yet — carries the requesting `username` instead (`ensureHostAwake({ requestedBy })`),
since its payload names a `hostId`/`label` that `GET /api/remote-hosts` withholds from
non-admins. With neither, it reaches admins only.
- **No real IO under vitest.** `probeRemoteHostReachable`, `runRemoteWakeCommand` and the
default UDP socket of `sendWakePackets` throw under `VITEST` (as `remote-files.ts` does),
so a test that reaches the defaults fails loudly instead of connecting, spawning or
broadcasting from CI. Every consumer injects its IO (`RemoteWakeDeps`, the socket
factory); `createDefaultRemoteWakeDeps({ probe })` also polls readiness with THAT probe,
which is the leak the guard found.
Tests: `test/remote-wake.test.ts` (decision/throttle table, single-flight registry,
buffering + flush order, MAC parsing/magic packet, live host-config resolution, the proxied
host, SSE payload routing, the vitest IO guard, and the wiring guard),
`test/routes/session-remote-wake.test.ts` (the input route buffers instead of writing into a
sleeping host — and writes straight into a proxied one —, the reachability route never wakes
and reports a proxied host as unknown, and the wake route reports the no-target case the UI
turns into "configure WoL"), `test/sse-routing-remote.test.ts` (multi-user routing of the
`remote:` family) and `test/host-wake-banner.test.ts` (banner visibility and when the poller
may connect).
## API
Routes are registered in `src/web/routes/case-routes.ts`:
@@ -365,6 +536,10 @@ Routes are registered in `src/web/routes/case-routes.ts`:
| `GET` | `/api/remote-hosts/:hostId/sessions` | Discover `codeman-*` sessions on the host (COD-105; `listRemoteCodemanSessions`, never errors) |
| `POST` | `/api/cases/remote-link` | Link a case to a remote host (creates the `RemoteCase`) |
`RemoteHost` accepts the optional `wakeMac` (magic packet, sent by Codeman) and `wakeCommand`
(single executable path, run without a shell, takes precedence) — see **Wake-on-LAN from user
input** above.
Attaching to a discovered session is a **session-create** path, not a host route:
`POST /api/sessions` accepts `attachRemoteSession: { hostId, remoteSessionName }`
(schema in `schemas.ts`; `remoteSessionName` must match `^codeman-[a-zA-Z0-9._-]+$`),
+11 -4
View File
@@ -270,9 +270,15 @@ tailscale serve --bg 3000 # HTTPS at https://<node>.<tailnet>.ts.net
Only devices on your tailnet can reach it; Tailscale handles identity and
terminates TLS with a real Let's Encrypt certificate (so PWA install and web
push work). No app password and no `0.0.0.0` bind required. (This is the
maintainer's production setup.) `CODEMAN_TAILSCALE=1` presets the choice for
automation; the installer never runs `tailscale serve reset` and never touches
serve mappings other than `443 -> Codeman's port`.
maintainer's production setup.) `CODEMAN_TAILSCALE=1` or `--tailscale` presets
the choice for automation. When `:443` on the node already belongs to another
app, the installer mounts Codeman under `/codeman` (`tailscale serve --set-path`
plus `--base-url`, which keeps the loopback bind and the same host guard) or on a
second port rather than replacing it. The installer never runs `tailscale serve
reset`, never touches serve mappings other than the one it created, never opens a
`tailscale funnel` (public internet, a different risk class) and never advertises
a Tailscale Service. Renaming the node (`--name`, `install.sh name`) is opt-in
and defaults to no, because the tailnet name is also the machine's SSH identity.
### B. Authenticated cloudflared tunnel + password
@@ -491,7 +497,7 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never** `--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, five seeded files from `~/.pi/agent`, and three from `~/.grok`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, five seeded files from `~/.pi/agent`, three from `~/.grok`, and, only when their opt-in switches `CODEMAN_AGENT_IMAGE_INSTALL_GH` / `_AZ` are `1`, `~/.config/gh/{hosts.yml,config.yml}` and the sign-in files from `~/.azure`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
@@ -510,6 +516,7 @@ Full feature guide: [`docker-cases.md`](docker-cases.md).
- **Auth is a parallel branch** (`middleware/auth.ts`) that leaves the single‑user path untouched: per‑user scrypt verify (`timingSafeEqual`, timing‑equalized against user enumeration), identity‑carrying cookies, a per‑username failure bucket (a botnet can't brute one account across IPs; one NATed user can't lock out the rest), and a `mustChangePassword` lockbox. The hook‑secret loopback bypass, host guard, and Origin/CSRF guard are unchanged (hooks authenticate the INSTANCE, not a user).
- **Ownership is enforced server‑side only** and fails closed: `req.authUser` (a synthetic admin in single‑user), `findSessionOrFail` returns NOT_FOUND (never 403) for a foreign session, list/SSE/WS/file‑preview/search all filter by `session.owner`, and SSE routing defaults session‑scoped events to their owner (unresolved owner → withheld). The load‑bearing rule is **non‑admin `workingDir` confinement**: a non‑admin's session/one‑shot working dir must realpath‑resolve inside `~/codeman-users/<name>/cases`, checked BEFORE any disk write.
- **Privileged actions are a one‑bit grant** (`canBypassPermissions`, default off): only granted users (and admins) get `--dangerously-skip-permissions` (others are silently downgraded to `--permission-mode auto`), shell‑mode sessions, cron `launchCommand`, and other CLIs' bypass flags. Machine‑level resources (remote/Docker host definitions, tunnel, self‑update, settings writes) are admin‑only.
- **Clone Repo does not lend the server's git sign-in to non-admins.** A clone writes only inside the caller's own case space, so it is not admin-gated, but the server account's git credential helpers (the Docker image's opt-in `gh`/`az` helpers, or any `gh auth setup-git`) are shared by every user. A non-admin's clone and preflight therefore run with `git -c credential.helper=`, which empties the helper list including the URL-scoped entries (`cloneWithoutCredentialHelpers` in `case-routes.ts`, argv pinned in `test/git-clone.test.ts`). This closes the Clone Repo path only: the account's SSH keys still apply to an `ssh://` URL, and a non-admin's agent sessions run as the same account, consistent with the first bullet above. Docker cases are a second route to the same sign-in: with `CODEMAN_AGENT_IMAGE_INSTALL_GH`/`_AZ` on, a non-admin's Docker case with credential seeding on (the default) receives a copy of the server account's `gh`/`az` sign-in, exactly as it receives the Claude and Codex credentials.
- **Admin actions are audited** append‑only to `~/.codeman/admin-audit.jsonl` (acting admin, action, target, IP). Passwords set by an admin create/reset are one‑time (returned once, force change). Under Basic auth, `logout` only truly ends QR‑issued sessions — to lock someone out, disable the account or reset the password (a proper login form is a deferred Phase 6).
---
+198
View File
@@ -0,0 +1,198 @@
# Split-Pane Sessions — Design Spec
**Status**: Implemented (v1)
**Author**: Claude (session with Tim), 2026-09-15
**Scope**: v1 only. v2 items are named and explicitly deferred, not designed.
## Problem
Codeman's terminal area shows exactly one active session (pane) at a time —
switching panes re-binds the single xterm instance and the single WebSocket
to a different session. Multi-monitor spanning (`scripts/span-codeman.sh` /
`span-codeman.ps1`) turned out to solve a different problem: it makes one
browser window bigger, but that window still shows one session; floating
subagent windows are draggable overlays on top of it, not tiled panes. There
is no way today to see two live sessions (e.g. `w1-codeman` and
`w1-mcp-memory`) side-by-side in one window, even on a monitor wide enough to
fit both.
## Goal (v1)
From the active session, open a **second, independent, fully live session**
in a pane beside it — draggable divider, side-by-side only. Closing the
second pane collapses back to today's normal single-pane view. No
persistence: a page reload always returns to single-pane. Floating
subagent/Ultracode windows keep their current behavior unchanged (global,
unconstrained across the whole viewport, split or not).
Explicitly out of scope for v1 (v2 candidates, not designed here):
- More than 2 panes / grid layouts
- Vertical (stacked) splits
- Drag-a-tab-to-split as a trigger (v1 trigger is an explicit button + picker)
- Persisting the split layout across reload or across devices
- Mobile/tablet layouts (viewport is too narrow for this to make sense; gated
to desktop widths the same way `home-sessions.js`'s rail is)
- Feature parity between the two panes (see "Pane B is deliberately plainer"
below)
## Current architecture (why this isn't a CSS change)
`terminal-ui.js` is built entirely around **singleton** state: `this.terminal`
(one xterm instance), `this._ws`/`this._wsSessionId` (one WebSocket, rebound
on every pane switch via `_disconnectWs()` + `_connectWs(newId)`), a
`this._xtermSnapshots` map used only to restore scrollback into that one
terminal when switching back to a session. Roughly 280 references to this
singleton state exist across the file (input handling, resize/fit, sizing-
token claims, mobile touch gestures, CJK IME, local-echo overlay wiring,
keyboard accessory bar, link providers, etc.).
Showing two sessions at once therefore requires a second, independently
alive xterm + WebSocket pair running concurrently — not a layout change to
one shared instance.
**Related prior art**: `detachSession(id)` (app.js) already opens one session
in a genuinely separate browser window (`isSoloWindow` mode) with its own
independent WebSocket, and two of those can already be snapped side-by-side
today with zero new code. That covers "two sessions visible at once" but not
what this spec is for: one Codeman window with two panes and a divider you
can drag without leaving your seat, each still a full participant in that
window's floating subagent windows, header, and settings. This spec builds
past detach, not a duplicate of it.
**Server-side check (done, not just assumed)**: `MAX_WS_PER_SESSION = 5`
(`src/web/routes/ws-routes.ts`), scoped by `clientId:tabNonce`
(`ws-connection-registry.ts`). Splitting always opens a *different* session
in the second pane (self-splitting is disallowed, see below), so this is two
sessions each getting their normal one connection — the existing cap is
irrelevant here and needs no server change.
## Key design decision: Pane B is deliberately plainer than Pane A
Porting all ~280 singleton behaviors to a second, symmetric pane is not
worth it for v1 — most of that code is input-quality-of-life for **mobile/
touch** (local-echo overlay, CJK IME textarea, touch gesture handling,
keyboard accessory bar), and this feature is desktop-only by nature (a split
view needs a wide viewport). So:
- **Pane A** (the session that was already active when you opened the split)
stays exactly what it is today — `this.terminal`, `this._ws`, unchanged
code path, zero regression risk.
- **Pane B** is a new, smaller `SplitTerminalPane` object: its own xterm
instance + fit addon, its own WebSocket to `/ws/sessions/:id/terminal`,
resize-on-divider-drag, and plain keyboard input. It does **not** get the
local-echo overlay, CJK IME composition, touch/mobile handlers, or the
keyboard accessory bar. On a desktop, typing directly into an xterm
instance with no overlay is exactly how Codeman behaved before the local-
echo overlay existed for touch devices — normal, not degraded, for a
keyboard-and-mouse user.
If this asymmetry actually bothers you in daily use, promoting Pane B to full
parity is a scoped v2 (extract the shared logic already once you have two
call sites to compare, rather than guessing the right abstraction now).
One more asymmetry worth naming here rather than discovering by surprise:
while both panes accept keyboard input, the global capture-phase shortcut
handler (`app.js`) always resolves against Pane A — it has no notion of
which pane currently has focus. So Ctrl+L or Ctrl+W typed while Pane B has
focus clears or closes Pane A, not the session you were actually typing
into. Not fixed for v1, same reasoning as the rest of this section.
## Components
### 1. `SplitTerminalPane` (new, `terminal-split.js`)
A small class, one instance per secondary pane:
- `constructor(sessionId, mountEl)`
- `connect()` — creates the xterm instance (same theme/font config as the
primary, read from the same settings so it doesn't visually clash), opens
`/ws/sessions/:id/terminal`, wires input → WS, WS → terminal write
- `fit()` — calls the fit addon; called on divider drag (rAF-throttled) and
on window resize
- `destroy()` — disposes the xterm instance, closes the WS cleanly
No snapshot/scrollback-restore map is needed the way `_xtermSnapshots` exists
for Pane A — Pane B is destroyed on close, not hidden-and-restored, since
there's no persistence requirement.
### 2. Split container (layout)
```
.terminal-split-container (flex row, only rendered when split is active)
├── .terminal-wrap (existing element, Pane A — untouched)
├── .split-divider (new, draggable seam)
└── .terminal-pane-b (new, hosts SplitTerminalPane's xterm + a
small header: session name + × close button)
```
When not split, `.terminal-wrap` renders exactly as it does today (no
wrapping container at all, to keep the no-split path byte-identical to
current behavior). Splitting inserts the container and reparents
`.terminal-wrap` into it as the first child — same reparenting pattern
already used by `applySessionListLayout()` for `#sessionTabs`, so this isn't
a new pattern for the codebase.
Default split is 50/50 (`flex-basis: 50%` each). Divider drag updates both
panes' `flex-basis` live (rAF-throttled) and calls `fit()` on **both**
terminals per tick, clamped to 20%/80% so neither pane can be dragged into an
unusably thin sliver.
### 3. Trigger UI
A **"Split"** button (header, opt-in like the other header buttons —
`showSplitButton`, default off, same pattern as `showMultiMonitorButton`)
opens a small picker listing your other open sessions (reuses
`this.sessions`/`sessionOrder`, filtered to exclude the currently active
session — you cannot split a session against itself). Picking one:
1. Creates the split container, reparents `.terminal-wrap`
2. Instantiates `SplitTerminalPane` for the chosen session in `.terminal-pane-b`
3. Button state flips to "close split" (or Pane B's own header × does it)
Closing (via Pane B's × or the header button toggling off):
1. `SplitTerminalPane.destroy()`
2. Removes `.terminal-split-container`, reparents `.terminal-wrap` back to
its original location at 100% width
3. Fires a resize/fit on Pane A (same `ResizeObserver`-driven fit already in
place today — no new code needed here, it fires naturally once the
container's size changes)
v2 note (not designed): dragging a session tab onto the active pane as an
alternate trigger. You confirmed right-click doesn't work today (Codeman
doesn't intercept it) and declined a keybind, so v1 is button+picker only.
### 4. Failure / edge cases
- **The Pane B session ends or is deleted while split is active** → treat
identically to the user closing Pane B manually: destroy the pane, collapse
to Pane A at full width.
- **The Pane A session ends while split is active** → Pane B is promoted:
it becomes the new single full-width pane (reusing today's normal
single-pane code path means Pane B's `SplitTerminalPane` must hand off to
a real `this.terminal`/`this._ws` binding — simplest correct approach is
to just collapse the split and let normal session-select logic reopen
Pane B's session as the new primary, rather than trying to promote the
lightweight pane object in place).
- **Both end** → falls through to today's normal "no active session" /
welcome-screen state.
- **Subagent/Ultracode floating windows** → no design work needed; they're
already positioned independent of `.terminal-wrap`'s layout, so they
continue to float over whichever pane(s) are on screen, unconstrained,
exactly as today.
## Testing
- Unit: `SplitTerminalPane` connect/fit/destroy lifecycle (mock WS, like
existing terminal tests use `TEST_PTY_SCRIPT`).
- Route/integration: opening two WS connections to two different sessions
from one simulated client concurrently — confirms the existing per-session
cap and connection registry need no changes.
- Browser (Playwright, `test/browser` since this is desktop-viewport-gated
UI): open split via button+picker, verify both panes render live output
independently, drag divider and confirm both refit, close Pane B and
confirm Pane A returns to full width, kill the Pane B session externally
and confirm auto-collapse.
## Open questions for review
None blocking — the scope-narrowing decisions above (Pane B feature parity,
no persistence, side-by-side only, button+picker trigger) came directly from
your answers during brainstorming. Flag anything here you want reconsidered.
+5
View File
@@ -1,5 +1,10 @@
# Tailscale Setup in the Installer (Plan)
> Superseded in part by [`installer-v2-plan.md`](installer-v2-plan.md) (2026-09-20), which
> moved every human step before the build, added the sub-path / second-port answer for an
> occupied `:443`, the opt-in rename, flags, and the done screen with a QR code. The
> state machine and safety rules below still hold.
Goal: make "Codeman over Tailscale, with real HTTPS" a first-class, guided path in
`install.sh`, instead of a one-line hint pointing at the docs. Today the safest
recommended deployment (loopback bind + `tailscale serve`) is exactly what the
+7
View File
@@ -273,9 +273,16 @@ into the case's `.claude/settings.local.json` so that `/model` keeps working.
- **Shell** for the times you want a terminal on your phone with no agent at all. It is a
genuinely useful mode, not a fallback.
## Pointing one at your own server
Most of these harnesses can also run against a custom OpenAI-compatible endpoint instead of
their native cloud backend, for one session at a time, an opt-in feature covered in full on
[Custom Model Endpoints](Custom-Model-Endpoints).
## Read next
- [Core Concepts](Core-Concepts) - run modes versus location overlays.
- [Custom Model Endpoints](Custom-Model-Endpoints) - run a harness against your own server.
- [Settings Reference](Settings-Reference) - model, effort, and permission-mode settings.
- [Keeping Agents Running](Keeping-Agents-Running) - what idle detection does per mode.
- [Security](Security) - what skipping permission prompts actually means.
+8 -1
View File
@@ -20,12 +20,19 @@ Three ways to get one, all under **+** next to the case picker:
| How | Result |
| ----------------- | ------------------------------------------------------------------------------------------------------ |
| **Create New** | A fresh `~/codeman-cases/<name>` with a scaffolded `CLAUDE.md`. |
| **Clone Repo** | A public repo cloned into `~/codeman-cases/<name>` and registered as a case. |
| **Clone Repo** | A repo cloned into `~/codeman-cases/<name>` and registered as a case. Private repos need this machine's own git credentials (see below). |
| **Link Existing** | An existing folder anywhere on disk, registered in place. Nothing is copied or moved. |
Linked cases keep living where they are. Deleting a case in Codeman removes the
registration, and for a linked case that is all it removes.
**Clone Repo never asks for credentials.** It uses whatever the server's own git already has:
an ssh key, or a credential helper such as `gh auth setup-git`. The Docker image can include
helpers for GitHub (`gh`) and Azure DevOps (`az`), turned on in `docker-compose.override.yml`;
then signing those CLIs in once from a shell session is enough. See the private repositories
section of `docker/README.md`. Without credentials a private repo fails straight away with an
authentication error.
**Cases created from scratch are the only copy of that code.** Uninstalling Codeman does not
delete `~/codeman-cases/`, but treat that directory as real work, not scratch space.
+189
View File
@@ -0,0 +1,189 @@
# Custom Model Endpoints
Point a harness at your own OpenAI-compatible server instead of its native cloud backend, for
one session at a time. "Custom endpoint" covers **local** hardware (llama.cpp, Ollama, vLLM,
a home GPU rig, DGX Spark, Strix Halo) and **cloud** services (Azure AI Foundry's
OpenAI-compatible endpoint, OpenRouter, a company gateway) alike, anything answering
`GET /v1/models` and `POST /v1/chat/completions` in the standard shape.
**Off by default.** Turn it on in App Settings → Models → **Custom model endpoints**.
## Adding an endpoint
Still in App Settings → Models → Custom model endpoints:
1. **+ Add endpoint** — give it an id, a label, and the base URL (`http://192.168.1.50:8080`,
say). An API key is optional; most local servers don't check one.
2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it.
3. Pick a **default model** from what was discovered. This is the model the Run-menu entry
applies directly when only one model is discovered; with two or more, it's just the one
pre-marked in the picker dialog described below, not a silent default.
Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker
hosts — these are machine-level infra, not a per-user setting.
**Model lists refresh themselves.** Every saved endpoint is re-discovered automatically every
5 minutes in the background, so a model the server starts serving later — or stops serving —
shows up without another manual click of **Discover**. One endpoint being unreachable on a
given cycle (powered off, wrong network) never blocks the others from refreshing.
**Context length is picked up automatically where it can be, safely.** Against a
llama.cpp/llama-swap server, discovery also learns each _currently loaded_ model's real
context window and applies it to the launched session (Claude Code today — see below), so
the harness stops assuming a large default window for a model name it doesn't recognise and
overflowing a much smaller real one. It's deliberately never probed for a model that isn't
already loaded, since asking a llama-swap server about an unloaded model can trigger an
actual, slow model swap as a side effect — a model just not currently loaded keeps whatever
context length an earlier cycle already learned for it instead.
## Running a session against one
With the setting on and at least one endpoint carrying a discovered model, the **Run**
dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a
custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a
session on that harness exactly the way its own entry would. It is a one-off "try this
endpoint" action, not a sticky mode — the plain **Run** button still means "this harness,
native cloud" afterward, and a fresh session never inherits whatever the last one was
pointed at.
**Which model it uses depends on how many the endpoint has discovered.** With exactly one,
the session launches straight away on that model — nothing to choose. With two or more, a
small dialog asks which one to use for this launch before starting the session; the
endpoint's default model, if set, is marked but not auto-picked, so a launch can deliberately
use a different one without changing the saved default. The list is not raw discovery order
either: the model llama-swap reports loaded and ready is moved to the top and tagged
**Currently loaded**, and when nothing is loaded, the model you last launched on this harness
and endpoint pair is moved up instead and tagged **Last used** (a per-device browser value, so
another device starts from its own history). The default model keeps its own **Default** pill
in both cases, and nothing is ever auto-chosen: the promoted row is simply the one under your
thumb.
**For opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP, picking an entry launches
straight onto the endpoint** — no restart, because the endpoint is applied before the
session's process ever starts. **Claude still restarts the harness's process in place** —
same tab, same conversation (`--resume`) — after a normal native launch, since that restart
is far less jarring for Claude than for the other seven, whose own TUI can fully
reinitialize on a restart. Either way, every supported harness reads its endpoint config at
process start, never per turn, so there is no live hot-swap while a turn is running.
Picking an entry that launches a **brand-new** Claude session waits (up to 20 seconds) for it to
finish its own startup before applying — a freshly started CLI reports itself as busy for its
boot sequence, and applying to a genuinely busy session is refused so a real, in-progress
turn is never interrupted out from under you. A session that is still busy after that wait
(a very slow-starting CLI, or one you started typing into right away) surfaces that refusal
as an ordinary error, which now stays on screen with a close button instead of vanishing
after a few seconds — read it, it names the actual reason rather than a generic failure.
Entries are hidden entirely for a session in a **remote (SSH) or Docker case** — support for
redirecting those hasn't landed yet, see below. The picker also only appears in the desktop
**Run** dropdown; the phone home screen builds its own run picker separately and does not
currently offer these entries.
**Against llama-swap, applying a selection also starts the actual model load, rather than
waiting on your first prompt to do it.** llama-swap has no "switch model" button of its own
— the only thing that starts a swap is a real request naming the model, and confirmed live:
just applying a selection never reached llama-swap's own logs at all until something asked
it to load. Picking an entry now also sends the smallest real request that will trigger
that load, in the background, the moment the target model isn't already loaded and ready.
**The centred loading banner has no countdown and no automatic timeout — it waits as long as
it takes, and tells you so.** When it knows the model's discovered file size (its GB figure,
when llama-swap states one) it's shown too, e.g. "Loading qwen3.8-27b (16.4 GB) on
llama-swap — this can take a while depending on your hardware and the model size." An
earlier version tried to estimate and enforce a time limit, but real load time depends on
hardware this feature has no way to know, so a fixed number was always a guess — worse, one
that could kill a genuinely slow load partway through. If it really is taking too long, a
**Cancel** button right on the banner ends the wait and **closes the session that load was
for**, on your own call rather than a guessed deadline.
**The banner also shows a real, live second line of what llama.cpp itself is doing** — not
a made-up progress phase, the actual next line the `llama-server` process printed, e.g.
"llama.cpp: load_model: loading model '/models/.../Qwen3.8-27B.gguf'" then later
"llama.cpp: llama_server: model loaded". It comes straight from llama-swap's own event
feed, filtered down to just the backend process's own output (not llama-swap's own request
logging), and stays on whatever it last said once the load goes quiet, rather than
clearing back to nothing.
**You'll also be told if a session's model gets swapped out from under it later, not just
at launch.** The conflict warning above only fires at the moment you launch or apply a
model — llama.cpp only runs one model at a time, so if a DIFFERENT session using the same
endpoint later triggers its own load, whatever was loaded before (including a session you
already had running) gets silently evicted, with no warning at that instant since nothing
conflicted when it was first set up. A background check (every 20 seconds) catches this
after the fact and shows a toast naming which session lost its model and what's loaded now
— so you know before typing into that session that it's about to reload (and, in turn,
evict whatever displaced it).
**Claude Code specifically gets three extra fixes applied automatically:**
- Its discovered context length (see above) is passed through as
`CLAUDE_CODE_MAX_CONTEXT_TOKENS`, so it doesn't send a full-size prompt against a much
smaller real local context and overflow it.
- Its session runs with an isolated `CLAUDE_CONFIG_DIR`, so the injected API key never sits
in the same directory as a stored claude.ai login — that combination is harmless for actual
requests (the API key wins) but the CLI still prints a "both claude.ai and
ANTHROPIC_API_KEY set" warning about it, which this avoids entirely. The isolated directory
keeps a link back to your real session history so the response viewer and similar features
still work for that session. That isolated directory starts with no prior approvals of its
own, so Codeman also pre-approves the injected key the same way answering Claude Code's own
"Detected a custom API key" prompt once would — without it, that prompt would otherwise
reappear on every single launch with nobody there to answer it.
- **That same fresh isolated directory also looks like a brand-new Claude Code profile**, so
without this fix it replayed the WHOLE first-run sequence every single launch: the theme
picker, the security-notes screen, the "trust this folder?" dialog, and a one-time warning
about running with permissions bypassed — none of which a real, already-used profile shows
again. Codeman now pre-seeds that same "already been through this once" state (onboarding
completed, this session's own project marked trusted, the bypass-permissions warning
acknowledged) so a custom-model launch reaches the actual conversation exactly as fast as a
native cloud one does, instead of stopping at a wizard with nobody there to click through it.
**If a model's real context is too small for Claude Code to even get started, you get a
warning instead of a confusing failure.** Claude Code's own system prompt and tools take up
roughly 40K tokens on their own, before you've typed anything — a small local model with a
smaller real context than that fails outright on the very first message, no matter what
context size Codeman tells it to expect (raising the declared context only changes when
Claude Code trims _conversation history_, and there is none yet on message one). Picking
such a model now shows an in-app dialog naming the model, its discovered context and what's
needed, before anything launches or restarts, with the fix spelled out: reconfigure
llama-swap to give that model (or a smaller one) an explicit larger context instead of
relying on auto-fit (`--fit-ctx`), which sizes the context around fitting the biggest model
rather than the biggest context — for example adding `-c 65536` to that model's llama-swap
entry. "Launch anyway" is still there if you want to try regardless.
## Which harnesses actually work
| Harness | Status |
| ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
| **Codex** | Config is correct, and plain chat can work against a server that speaks the Responses API — but a real tool-call attempt comes back as inert text instead of running, so it's still not usable for real coding work. |
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
| **DeepSeek** | The original 404 is root-caused and fixed (DeepSeek Harness's own code was missing a `/v1` most local servers require) — not yet re-run against a real `dsh` install to confirm end-to-end. |
| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. |
Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry,
not a fixed list here, so this table can go stale before this page does — a greyed-out or
missing entry is the more current answer.
## What it does not do
- **No remote or Docker sessions yet.** Both restart their agent differently under the hood
(reattaching a durable tmux session rather than relaunching the process), so redirecting
them needs its own plumbing that hasn't been built.
- **No live hot-swap mid-conversation.** Applying a selection always restarts the process.
- **No button to un-point a session from the UI yet.** Clearing back to native cloud is an
HTTP call (`POST .../custom-model {"clear": true}`) or deleting the session; the settings
panel manages saved endpoints, not what a running session is currently pointed at.
- **Nothing is shared with your real cloud credentials.** The endpoint's own key, if any,
never touches your Anthropic/OpenAI/Google login — a custom endpoint is a separate,
explicit choice per session.
## Security
An endpoint's base URL can't point at a link-local or cloud-metadata address (both at save
time and against the address it actually resolves to), the same guard Web Tabs uses for
saved dashboards. Endpoint records and any per-session config files a harness needs are
written with owner-only permissions. See
[custom-model-endpoints-plan.md](https://github.com/Ark0N/Codeman/blob/master/docs/custom-model-endpoints-plan.md)
in the repository for the full design reasoning, including why this feature closed a
pre-existing gap in how session environment overrides were guarded rather than opening a new
one.
+29
View File
@@ -118,6 +118,35 @@ invisible from the host (`pi -c` and `grok -c` inside a docker case see only tha
container's history). OMP's `sessions/` is the exception and is shared read-write, because
Codeman reads it host-side for history and resume.
**Git hosts.** The agent image can also include the GitHub CLI (`gh`) and the Azure CLI (`az`,
with the `azure-devops` extension), off by default, and its git then uses them as credential
helpers for github.com and Azure DevOps. When the matching switch is on, their sign-ins are
seeded like everything else, file by file: `~/.config/gh/hosts.yml` and `config.yml`, and the
sign-in files from `~/.azure` (not its logs or extensions). With a switch off they are never
copied in, even if the files exist. So once a switch is on and `gh auth login` / `az login`
have been run where Codeman runs, agents in a Docker case can clone and push private repos on
those hosts. Two limits:
- A token held in a desktop keyring or an encrypted token cache (Windows, macOS) is not
inside those files and does not carry in. Sign in inside the container instead. The Docker
server image and a headless Linux host keep it in the files, so they carry.
- The sign-ins are mounted when a case container is **created**, so an existing container
never picks them up. After turning a switch on, signing in, or rebuilding the agent image,
**recreate the case container**: remove it, and the next session in that case creates a
fresh one. (Or sign in inside the existing container instead.)
This hands a GitHub token and an Azure sign-in to every agent in a seeded Docker case, the
same trust you already give it with Claude, Codex or gcloud. Turn seeding off for a case that
should not have them.
Both CLIs are opt-in. To build the agent image with them, set
`CODEMAN_AGENT_IMAGE_INSTALL_GH=1` and/or `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` where the image
is built: in front of `node scripts/build-agent-image.mjs`, or in the Codeman server's
environment for the image it builds automatically (in the Docker deployment, `environment:`
in `docker-compose.override.yml`), then rebuild the image with `--no-cache`.
`docker/README.md` ("Private repositories") has the details and the matching switches for
the server image.
## Isolation
Every container runs hardened by default:
+47 -13
View File
@@ -24,16 +24,22 @@ This installs Node.js, tmux and a build toolchain if they are missing (node-pty
Linux prebuild, so it compiles from source), clones Codeman into `~/.codeman/app`, and
builds it.
What it asks you:
It starts by printing what it found (git, Node, tmux, build tools, agent CLIs, Tailscale,
an existing install), then asks everything it needs up front, then does the work
unattended. You can leave while it builds. What it asks you:
1. **Permission for every system change.** Package installs and agent CLI downloads are
prompted individually. Nothing is installed silently. If no agent CLI is found, a menu
offers to install any of them (DeepSeek excepted: its npm package installs only a
launcher with no runnable profile), or you skip and install one yourself later.
1. **One consent for the missing packages.** Git, Node.js, tmux and (on Linux) the build
toolchain are installed after a single yes, and sudo asks for your password once for
the whole run. Nothing is installed silently. If no agent CLI is found, a menu offers
to install any of them (DeepSeek excepted: its npm package installs only a launcher
with no runnable profile), or you skip and install one yourself later.
2. **How the dashboard should be reachable.** Three choices:
- **Tailscale** (recommended for phone access): keeps the loopback bind and walks you
through `tailscale serve`, including the tailnet HTTPS toggle, then verifies the result
end to end.
- **Tailscale** (recommended for phone access): keeps the loopback bind, installs
Tailscale if needed, logs in, enables the tailnet HTTPS toggle (it opens the admin
page for you and waits; Ctrl+C there skips Tailscale for this run), then configures `tailscale serve` after the build and
verifies the result end to end. If another app already owns `:443` on your node,
you choose between a sub-path (`https://<machine>.<tailnet>.ts.net/codeman`, the
default), a second port, replacing the other mapping, or skipping.
- **Your local network** (`0.0.0.0`): prompts for a password. Skipping the password takes
an explicit confirmation and ends on a loud warning.
- **This machine only** (`127.0.0.1`): the safest option, and the default for a bare
@@ -44,26 +50,54 @@ What it asks you:
Tailscale. An existing loopback install defaults to keeping loopback, or to Tailscale when
a serve mapping for Codeman is already there. A bare Enter never pulls in new software,
and a non-interactive run always keeps the safe loopback default.
3. **What to do when it finishes.** Run in this terminal, install as a background service
that starts on boot, or do nothing yet.
3. **What to call this machine on your tailnet** (Tailscale route only). By default the URL
uses the machine's existing name. Answer yes to rename it `codeman-<hostname>`; the
default is no, because the tailnet name is also what SSH and everything else on that
machine are reached by.
4. **Whether to run Codeman in the background.** Enter installs a systemd user service or a
macOS LaunchAgent that starts on boot; answering no offers to start it in this terminal
instead, or not at all.
It ends on a screen with the URL (your tailnet, your network, or this machine), a QR code to
scan with your phone, and the two commands you need to manage the service.
Re-running the same one-liner **updates an existing install in place**. Local changes in
`~/.codeman/app` are stashed rather than discarded, a running service is restarted and
verified, and your existing network binding is preserved. An interrupted first install
resumes instead of restarting.
Two other entry points exist:
Other entry points:
```bash
install.sh status # print the URLs, the QR code and the manage commands again
install.sh update # update only
install.sh uninstall # remove
install.sh uninstall # remove (offers to undo a rename it performed)
install.sh tailscale # retrofit Tailscale access onto an existing install
install.sh name [<n>] # rename this machine on your tailnet (default codeman-<hostname>)
install.sh cloudflared # install cloudflared for the in-app Cloudflare tunnel
```
**Flags** answer the questions from the command line and pipe through `bash -s --`:
```bash
curl -fsSL https://getcodeman.com/install | bash -s -- --tailscale --service
curl -fsSL https://getcodeman.com/install | bash -s -- --lan --password 'x' --service
curl -fsSL https://getcodeman.com/install | bash -s -- --local --run
```
`--tailscale` / `--lan` / `--local` answer the access question, `--name <n>` / `--no-rename`
the name, `--service` / `--run` / `--no-start` the last one. `--yes` takes every default
(it still waits on a Tailscale login URL, and a network bind still asks for a password).
`--port <n>` moves Codeman off 3000; the service file and the serve mapping follow it. On an
existing install, `--port` and `--password` re-run the setup so the service file picks them up,
and a re-run with `--lan` or `--tailscale` keeps the password the service already has.
**Automation and CI**: with no terminal attached, any step that would change the system
aborts with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to
approve those steps. `CODEMAN_TAILSCALE=1` preselects the Tailscale answer, and never
installs Tailscale itself non-interactively.
installs Tailscale itself non-interactively; a non-interactive run never renames the
machine and never starts a service. Everything the unattended steps print goes to
`~/.codeman/install.log`, and the last lines of it are shown when a step fails.
## Route B: npm
+2
View File
@@ -34,6 +34,8 @@ Press `Ctrl+?` in the app for the same list in a floating overlay.
| Right-click | Copy the selection. With nothing selected the native menu is left alone. |
| `Ctrl+Z` | Swallowed in agent sessions so a running CLI cannot be suspended. Normal job control in a shell. |
Anything you copy is cleaned on the way to the clipboard: each line loses the padding spaces a full-screen program paints across the rest of the row. Leading indentation is left exactly as it is, so indented code, a `git log` message body and `git diff` context lines paste back the way they looked on screen. An `Alt+drag` rectangular selection is copied exactly as it looks, so its columns stay lined up.
## Everything else
| Shortcut | Action |
+12 -6
View File
@@ -60,14 +60,20 @@ On by default; it can be turned off in settings.
A row of keys above the virtual keyboard, and what it contains depends on the session.
**Agent sessions** get quick actions: `/init`, `/clear`, `/compact`, a clipboard key, `Esc`,
a path picker, an image key, and 🧠 when Read My Mind is on. Destructive commands need a
double press, so you cannot fire `/clear` with a stray thumb. On Codex sessions the bar also
shows `⇧←` and `⇧→`, the Shift-modified arrows Codex binds to editing the last queued
message and walking the prompt stack.
**Agent sessions** get quick actions: `/init`, `/clear`, `/compact`, a Compose key, `Esc`,
a path picker, and 🧠 when Read My Mind is on. Compose opens a multiline editor with
autocorrect: Enter adds a new line, and only Send delivers the text, as one paste followed
by Enter, so your line breaks reach the agent intact. Anything already typed on the terminal
prompt moves into the editor when it opens. Drafts are kept per session and in memory only,
so switching tabs keeps them and a page reload forgets them; a dot on the key shows a draft
is parked. The editor's Image button attaches photos and puts their paths into the draft.
Destructive commands need a double press, so you cannot fire `/clear` with a stray thumb. On
Codex sessions the bar also shows `⇧←` and `⇧→`, the Shift-modified arrows Codex binds to
editing the last queued message and walking the prompt stack.
**Shell sessions** automatically swap it for terminal controls: `Ctrl`, `Esc`, `Tab`, four
arrows, paste, and dismiss. Your normal preference is remembered and restored when you
arrows, a direct Paste key (shell input is not an agent prompt, so there is no Compose
there), and dismiss. Your normal preference is remembered and restored when you
switch back to an agent session, so a settings change during a shell session cannot strip
the bar away permanently.
+24
View File
@@ -103,6 +103,30 @@ locked phone and the agent continues.
With the inbox off, the buttons are stripped from the notification payload entirely rather
than being shown and failing.
## When a session is watching its own work
An agent that starts a monitor, puts a shell in the background or hands a task to a cloud
session is told by its CLI to end the turn and wait to be notified. The pane then goes
quiet, and the CLI's idle notification arrives about a minute later — for a session that
wants nothing from you.
Codeman reads what the CLI prints about its own background work and treats that prompt
differently. It raises no tab alert, no desktop notification and no push, the session stays
out of NEEDS YOU on every surface, and the row wears a blue **watching** badge instead. Hover
it, or read it on a phone through your screen reader, and it says what is running: "1
monitor", "2 shells", "1 background terminal".
The prompt itself is not thrown away. It sits in the Approvals drawer as an ordinary card,
still answerable, with a line reading "quiet, watching 1 monitor" where a card you had
already looked at would say nothing. The next time that session goes quiet for an ordinary
reason, it alerts you exactly as before.
Two limits are worth knowing. A permission prompt or a question dialog still goes red
whatever else the agent started, because that one blocks it outright. A question asked in
plain prose is not a dialog, so an agent that starts a monitor and then writes "which branch
should I target?" is quiet along with the rest — check a watching session yourself if it has
been quiet longer than the work it is waiting for should take.
## The phone overview
On phones, tapping the "C" logo gives a session overview with **NEEDS YOU** first, then
+1 -1
View File
@@ -43,7 +43,7 @@ To make a new one, click **+** next to the picker. The Add Case dialog has three
| Tab | Use it when |
| ----------------- | ------------------------------------------------------------------------------------------------------------------ |
| **Create New** | Starting a fresh project. Creates `~/codeman-cases/<name>` and scaffolds a `CLAUDE.md` into it. |
| **Clone Repo** | Working on an existing public repo. Paste the URL; Codeman preflights it as you type, offers the repo's real branches and tags, and fills in the case name. |
| **Clone Repo** | Working on an existing repo: public, or private once this machine's git can authenticate (the Docker image can include `gh`/`az` helpers for this). Paste the URL; Codeman preflights it as you type, offers the repo's real branches and tags, and fills in the case name. |
| **Link Existing** | The code is already on disk. Point at the folder, with **Browse** if you would rather click than type. |
The gear next to the picker holds two per-case toggles: **Agent Teams** and
+24 -3
View File
@@ -33,11 +33,12 @@ Your devices join a private network, and Codeman stays bound to loopback. Nothin
published to the internet, and you get real HTTPS with a real certificate.
The installer sets this up for you, including installing Tailscale, logging in, enabling
tailnet HTTPS, and verifying the result end to end. To retrofit it onto an existing
install:
tailnet HTTPS, and verifying the result end to end. It ends on the URL with a QR code to
scan. To retrofit it onto an existing install, or to see the URL and QR code again:
```bash
install.sh tailscale
install.sh status
```
By hand:
@@ -49,6 +50,22 @@ tailscale serve status
Then open `https://<machine>.<tailnet>.ts.net` from any device on your tailnet.
### The name in the URL
The URL is the machine's MagicDNS name, so on a machine called `tnode` it is
`https://tnode.<tailnet>.ts.net`. Three ways to influence that, from least to most work:
| You want | How |
| ------------------------------------------ | ----------------------------------------------------------------------------------------------------- |
| The machine's existing name (default) | Nothing. This is what the installer does unless you say otherwise. |
| `https://codeman-<hostname>.<tailnet>.ts.net` | Answer yes to the installer's name question, pass `--name codeman-<hostname>`, or run `install.sh name`. This renames the machine tailnet-wide (SSH included), which is why the installer defaults to no. `install.sh uninstall` offers to rename it back. |
| `https://codeman.<tailnet>.ts.net` | A [Tailscale Service](https://tailscale.com/docs/features/tailscale-services). Only a **tagged** node can host one (a device signed in with a user account cannot), the service is defined and approved in the admin console, and the feature is in beta. The installer does not set this up; it is a `tailscale serve --service=svc:codeman --https=443 127.0.0.1:3000` on a tagged host once the service exists. |
If `:443` on your node already belongs to another app, the installer offers Codeman under
`https://<machine>.<tailnet>.ts.net/codeman` (the default, via `tailscale serve --set-path`
plus Codeman's `--base-url`), on a second port (`https://<machine>.<tailnet>.ts.net:8443`),
or replacing the other mapping. It never replaces anything without asking.
Notes:
- Keep the loopback bind. `tailscale serve` connects to `127.0.0.1:3000` locally, so
@@ -58,7 +75,11 @@ Notes:
- Codeman's Host-header allowlist already accepts `.ts.net`, so no extra configuration is
needed.
- The installer never resets or rewrites `serve` mappings other than the one pointing at
Codeman's port, so unrelated serve configuration is left alone.
Codeman's port, so unrelated serve configuration is left alone. It also never opens a
`tailscale funnel` (that is the public internet) and never advertises a Tailscale Service.
- On macOS, the App Store and standalone Tailscale apps only run once someone is logged in,
so a headless Mac needs the open-source `tailscaled` for the URL to come back after a
reboot on its own.
## Cloudflare tunnel
+8
View File
@@ -67,6 +67,14 @@ On Linux, if you want the service running while you are not logged in:
loginctl enable-linger $USER
```
On macOS, a LaunchAgent starts when you log in, not at boot. A headless Mac (no GUI login)
needs a system LaunchDaemon instead, written by hand as root. The installer recognises an
existing `/Library/LaunchDaemons/com.codeman.web.plist` and leaves it alone rather than
installing a LaunchAgent next to it, since the two would fight over the port; remove the
daemon first if you want to switch. The same login caveat applies to the App Store and
standalone Tailscale apps, so on a headless Mac the Tailscale URL only comes back after a
reboot if the open-source `tailscaled` is used.
### Writing the unit by hand
**Linux (systemd user unit):**
+9 -2
View File
@@ -46,6 +46,7 @@ supervised by systemd or launchd; npm installs report as non-updatable. See
| Extended Keyboard Bar | Per device | Which accessory bar phones get. Shell sessions override it while they are active. |
| Wheel Scrolls Local History | Off | Keeps the wheel on the local buffer instead of forwarding it to the CLI. |
| Auto Copy Selection | Off | Copies highlighted terminal text to the clipboard the moment you finish selecting it. Ctrl+C still copies on demand. |
| Trim The Pane Margin On Copy | On | Takes the left margin a full-screen agent CLI paints down its own edge off a copy, so the text pastes flush. Each CLI declares its own width, and the strip never exceeds the indent every selected line shares, so nesting is kept. Claude Code and Codex declare a margin; a shell does not. |
| Normal / Bold font weight | xterm defaults | Per device, each slot from 100 to 900. The bundled JetBrains Mono renders every step, so a lighter normal weight makes Claude's bold headings stand out. Applies live to the terminal, both echo overlays and open team panes. |
| WebGL Renderer | On | With a GPU-stall watchdog that falls back to DOM rendering. |
| Gesture Control | Off | Camera hand tracking. Also needs `CODEMAN_GESTURE=1` on the server. |
@@ -55,12 +56,14 @@ supervised by systemd or launchd; npm installs report as non-updatable. See
Chips for every optional header control, with a live preview of the resulting header:
Run, Font Size, System Stats, Redraw Terminal, Response Viewer, Away Digest, Session
Manager, Attachments, File Viewer, Multi-monitor, Plan Usage, Lifecycle Log, Monitor,
Manager, Attachments, File Viewer, Multi-monitor, Split, Plan Usage, Lifecycle Log, Monitor,
Project Insights, File Browser, Subagents, Approvals Inbox, Read My Mind, Ultracode Agents,
Ultracode Windows, Cron.
Most default to off. The stock desktop header is system stats, File Viewer, and the gear.
New header controls never appear on phones.
New header controls never appear on phones. Split is desktop-only regardless of this
setting — the button and the feature both stay off below a ~1180px viewport, where two
resizable panes plus their divider have nowhere to go.
This section also holds background-agent tracking, including whether to track agents for
every session or only the active tab.
@@ -92,6 +95,10 @@ Model and effort are both **soft defaults**: the model is written into the case'
`.claude/settings.local.json` and effort is passed at start, so `/model` and `/effort`
inside a session override them at any time.
**Custom model endpoints** (off by default) adds a saved-endpoint list plus a matching
section to the Run dropdown, for pointing a harness at your own OpenAI-compatible server
instead of its native cloud backend. See [Custom Model Endpoints](Custom-Model-Endpoints).
### Agents & CLIs
| Setting | Notes |
+2
View File
@@ -48,6 +48,7 @@ One tab per session, in your order, and that order syncs across your devices.
| Yellow tab, blinking | The agent is waiting for input from you. |
| Red tab, blinking | A question or permission prompt is blocking the session. |
| No dot | The session is not running. |
| Muted grey dot plus an `exited (137)` badge | The agent inside the pane has exited, with that exit code (or `exited (signal 9)`). A bare `exited` means tmux saw the pane die but did not report how, which is not the same as a clean `exited (0)`. Detailed sidebar and rail rows read `exited` in their pill. |
![Tab alerts](https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/tab-alerts-20260815.png)
@@ -116,6 +117,7 @@ The right side of the header. Almost all of these are off until you enable them
| Lifecycle Log | Off | Session start, exit, and kill audit trail. |
| Cron ⏰ | Off | Scheduled jobs. |
| Multi-monitor | Off, macOS | Opens a window spanning every display. |
| Split | Off, desktop only | View a second session beside the active one, with a draggable divider. |
| Tunnel indicator | When a tunnel runs | Cloudflare tunnel status. |
| Admin panel | Multi-user only | User administration. |
+1
View File
@@ -12,6 +12,7 @@
- [The Dashboard](The-Dashboard)
- [Agent CLIs](Agent-CLIs)
- [Custom Model Endpoints](Custom-Model-Endpoints)
- [Working With Files](Working-With-Files)
- [Input And Voice](Input-And-Voice)
- [Mobile Guide](Mobile-Guide)
+1479 -554
View File
File diff suppressed because it is too large Load Diff
+3 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.30.0",
"version": "1.33.1",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.30.0",
"version": "1.33.1",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
@@ -55,6 +55,7 @@
"@types/web-push": "^3.6.4",
"@types/ws": "^8.18.1",
"@vitest/coverage-v8": "^4.1.8",
"@xterm/headless": "^6.0.0",
"agent-browser": "^0.6.0",
"esbuild": "^0.27.3",
"eslint": "^9.0.0",
+2 -1
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.30.0",
"version": "1.33.1",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -123,6 +123,7 @@
"@types/web-push": "^3.6.4",
"@types/ws": "^8.18.1",
"@vitest/coverage-v8": "^4.1.8",
"@xterm/headless": "^6.0.0",
"agent-browser": "^0.6.0",
"esbuild": "^0.27.3",
"eslint": "^9.0.0",
+1 -1
View File
@@ -1,7 +1,7 @@
{
"name": "codeman",
"description": "Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.30.0",
"version": "1.33.1",
"author": {
"name": "Ark0N",
"url": "https://github.com/Ark0N"
@@ -340,7 +340,7 @@ ESC=$(printf '\033')
### Starting a worker
`POST /api/v1/quick-start` body (all optional):
`{"caseName":"worker-1","mode":"claude","sessionName":"w9-worker","effort":"high"}`
`{"caseName":"worker-1","mode":"claude","sessionName":"auth-worker","effort":"high"}`
, `mode` ∈ `claude|shell|opencode|codex|gemini|antigravity|pi|grok|deepseek|omp`; response is
`.data.{sessionId, caseName, casePath}`. Creates the case directory (a real directory
on the user's disk) if missing, do not retry it in a loop, and remember the name.
@@ -101,11 +101,16 @@ the case name, read it from the listing.
From Codeman 1.16 a LOCAL claude spawn passes `--name <session name>` when the local
CLI is 2.1.224+ (`buildNameCliArgs`, `session-cli-builder.ts:97-101`, wired in at
`tmux-manager.ts:797`), so a worker's peer name usually IS its Codeman session name
(verified live: quick-start with `sessionName: "w9-msgtest"` listed as `w9-msgtest`,
and its messages arrive tagged `from-name="w9-msgtest"`; a derived-name worker's
(verified live: a quick-start `sessionName` is listed as that exact peer name, and
the worker's messages arrive tagged `from-name="<that name>"`; a derived-name worker's
messages carry no `from-name`). Name your workers: a quick-start WITHOUT
`sessionName` leaves the Codeman name empty, so there is nothing to pass and the
peer name stays derived. The flag is fail-closed (older/unknown CLI omits it, because an
peer name stays derived. ⚠️ Give them a DESCRIPTIVE name: only a name the user chose
is pinned (`Session.cliPinnedName`), because `--name` is also the conversation's
`/resume` title and terminal title and suppresses Claude's own generated title. A
placeholder-shaped name (`w9-msgtest`, anything matching `isGeneratedSessionName`)
and an auto name are NOT passed, so such a worker's peer name is derived; use
`msgtest-worker` rather than `w9-msgtest`. The flag is fail-closed (older/unknown CLI omits it, because an
unknown flag aborts startup and would kill every spawn) and allowlist-sanitized (a name of
only unsafe characters is dropped), and the docker/remote builders never see it at all
(`tmux-manager.ts:782-789`), which is why the `tmux` column stays the canonical join key
@@ -196,8 +201,8 @@ idle:
The contract an orchestrator follows for any fleet of two or more messaging workers.
Every topology in the next section is this protocol plus a wiring diagram.
1. **Spawn with a name, and confirm hooks.** Use `quick-start` with `sessionName` (the
`--name` gate above). Session create installs the hooks block into the workspace
1. **Spawn with a name, and confirm hooks.** Use `quick-start` with a descriptive,
non-`w<N>-` `sessionName` (the `--name` gate above). Session create installs the hooks block into the workspace
whatever kind it is, so a linked case and a raw `POST /api/sessions` path both get
`stop`/`blocked` by default. ⚠️ Not unconditionally: the operator can turn
`workspaceHooksEnabled` off, remote SSH sessions never get hooks, and a session from
@@ -692,8 +692,9 @@ The shape, each step verified live (probes, failure modes and safety detail in
[§5.2](#52-readiness)).
2. `ListAgents`: find the worker's row by its `tmux codeman-<first 8 of session id>`
column; the row's `name [ref]` is the address. On Codeman 1.16+ with claude
2.1.224+ a worker's peer name is its Codeman session name, so pass `sessionName`
in quick-start to pick it; older setups list a name derived from the case folder.
2.1.224+ a worker's peer name is its Codeman session name, so pass a DESCRIPTIVE
`sessionName` in quick-start to pick it (a `w<N>-` placeholder-shaped name is not
pinned, so it lists derived); older setups list a name derived from the case folder.
No row = messaging is off for that worker (it is feature-flagged even on matching
CLI versions, observed live): fall back to the HTTP recipes without complaint.
3. `SendMessage` the task; first contact must use the `name [ref]` form copied from
+34
View File
@@ -146,10 +146,44 @@ console.log('\n[build] content-hash cache busting');
html = html.replaceAll(`"${original}"`, `"${hashed}"`);
}
writeFileSync(join(distPublic, 'index.html'), html);
// Rewrite sw.js from the SAME manifest that just renamed the files.
//
// The service worker's precache list used to be maintained by hand with the
// pre-hash names, so after this step every entry in it pointed at a file that
// no longer existed and `cache.add(...).catch(() => {})` hid it. Deriving it
// here is the only way the two cannot drift.
//
// The cache key gets the build hash for the same reason: `activate` deletes
// every cache that is not the current one, so a constant key meant that
// cleanup never ran and hashed assets from every past release piled up.
const swPath = join(distPublic, 'sw.js');
let sw = readFileSync(swPath, 'utf8');
const hashedAssets = Object.values(manifest);
const buildId = createHash('md5').update(hashedAssets.join('|')).digest('hex').slice(0, 12);
// Rewrite the two declarations. Anchored on the full `const … = …;` text so
// each pattern occurs exactly once and cannot collide with prose in sw.js's
// own comments — an earlier cut used bare `__BUILD_ID__` sentinels and the
// first match landed in the comment that documented them, leaving the real
// constant untouched and still producing a plausible-looking cache key.
const swEdits = [
["const BUILD_ID = 'dev';", `const BUILD_ID = '${buildId}';`],
['const HASHED_ASSETS = [];', `const HASHED_ASSETS = [${hashedAssets.map((p) => JSON.stringify(p)).join(', ')}];`],
];
for (const [from, to] of swEdits) {
const hits = sw.split(from).length - 1;
if (hits !== 1) {
throw new Error(`sw.js: expected exactly one \`${from}\`, found ${hits} — precache would ship stale`);
}
sw = sw.replace(from, to);
}
writeFileSync(swPath, sw);
console.log(' Hashed files:');
for (const [orig, hashed] of Object.entries(manifest)) {
console.log(` ${orig} -> ${hashed}`);
}
console.log(` sw.js: cache bucket codeman-${buildId}, ${hashedAssets.length} precached assets`);
}
// 6. Compress with gzip + brotli
+30 -3
View File
@@ -55,9 +55,36 @@ export function agentImageNpmPackages(catalog) {
return packages;
}
/** The `--build-arg` pairs the agent image takes. PURE. */
export function agentImageBuildArgPairs(catalog) {
return [['CLI_NPM_PACKAGES', agentImageNpmPackages(catalog).join(' ')]];
/**
* Environment variable → agent.Dockerfile ARG for the optional git-host CLIs (gh, az).
* ⚠️ Mirrored by `GIT_HOST_CLI_BUILD_ARGS` in `src/docker-hosts.ts`; the parity test pins them.
*/
export const GIT_HOST_CLI_BUILD_ARGS = [
['CODEMAN_AGENT_IMAGE_INSTALL_GH', 'CODEMAN_INSTALL_GH'],
['CODEMAN_AGENT_IMAGE_INSTALL_AZ', 'CODEMAN_INSTALL_AZ'],
];
/**
* The `--build-arg` pairs for the optional git-host CLIs. PURE. An unset or empty variable
* contributes NOTHING, so the Dockerfile's own default (off) applies and the argv is the same
* as before these existed; anything other than 0/1 is refused rather than guessed at.
*/
export function gitHostCliBuildArgPairs(env) {
const pairs = [];
for (const [envName, argName] of GIT_HOST_CLI_BUILD_ARGS) {
const value = env[envName];
if (value === undefined || value === '') continue;
if (value !== '0' && value !== '1') {
throw new Error(`${envName} must be 0 or 1, got ${JSON.stringify(value)}`);
}
pairs.push([argName, value]);
}
return pairs;
}
/** The `--build-arg` pairs the agent image takes. PURE given `env`. */
export function agentImageBuildArgPairs(catalog, env = process.env) {
return [['CLI_NPM_PACKAGES', agentImageNpmPackages(catalog).join(' ')], ...gitHostCliBuildArgPairs(env)];
}
/** Read the committed catalogue. IO. */
+20 -1
View File
@@ -74,6 +74,15 @@ echo "[self-update] $(date) start tag=$TAG supervisor=$SUPERVISOR repo=$REPO"
export PATH="$(dirname "$NODE"):$HOME/.local/bin:$HOME/.npm-global/bin:/usr/local/bin:/opt/homebrew/bin:$PATH"
export GIT_TERMINAL_PROMPT=0
# --node is the server's process.execPath, a VERSIONED path (Homebrew resolves it
# into Cellar/node/<ver>/). A `brew upgrade node` under a long-running server
# deletes it, and every status write then failed, so the status stayed "queued"
# forever. Fall back to whatever node is on PATH.
if [ ! -x "$NODE" ]; then
echo "[self-update] WARN: $NODE is not executable, falling back to node on PATH"
NODE="$(command -v node || echo node)"
fi
TO_VERSION="${TAG##*@}" # codeman@0.9.4 → 0.9.4 (tag is validated upstream)
STASH_REF=""
MANUAL_CMD=""
@@ -276,7 +285,17 @@ case "$SUPERVISOR" in
# domain needs root, but we don't need it — kill the server and launchd
# respawns it on the new dist/ within ThrottleInterval seconds.
if [[ -n "$SERVER_PID" ]] && kill "$SERVER_PID" 2>/dev/null; then
: # respawn is launchd's job from here
# Respawn is launchd's job, but only once the old process EXITS. A graceful
# shutdown that hangs leaves the port closed and the service down, so
# escalate to SIGKILL (tmux sessions live outside the server and survive).
for _ in $(seq 1 30); do
kill -0 "$SERVER_PID" 2>/dev/null || break
sleep 1
done
if kill -0 "$SERVER_PID" 2>/dev/null; then
echo "[self-update] server pid $SERVER_PID still alive 30s after SIGTERM, sending SIGKILL"
kill -9 "$SERVER_PID" 2>/dev/null || true
fi
else
MANUAL_CMD="sudo launchctl kickstart -k system/com.codeman.web"
write_status "completed-needs-manual-restart" "Update staged — restart Codeman to apply v$TO_VERSION."
+1 -1
View File
@@ -340,7 +340,7 @@ ESC=$(printf '\033')
### Starting a worker
`POST /api/v1/quick-start` body (all optional):
`{"caseName":"worker-1","mode":"claude","sessionName":"w9-worker","effort":"high"}`
`{"caseName":"worker-1","mode":"claude","sessionName":"auth-worker","effort":"high"}`
, `mode` ∈ `claude|shell|opencode|codex|gemini|antigravity|pi|grok|deepseek|omp`; response is
`.data.{sessionId, caseName, casePath}`. Creates the case directory (a real directory
on the user's disk) if missing, do not retry it in a loop, and remember the name.
+10 -5
View File
@@ -101,11 +101,16 @@ the case name, read it from the listing.
From Codeman 1.16 a LOCAL claude spawn passes `--name <session name>` when the local
CLI is 2.1.224+ (`buildNameCliArgs`, `session-cli-builder.ts:97-101`, wired in at
`tmux-manager.ts:797`), so a worker's peer name usually IS its Codeman session name
(verified live: quick-start with `sessionName: "w9-msgtest"` listed as `w9-msgtest`,
and its messages arrive tagged `from-name="w9-msgtest"`; a derived-name worker's
(verified live: a quick-start `sessionName` is listed as that exact peer name, and
the worker's messages arrive tagged `from-name="<that name>"`; a derived-name worker's
messages carry no `from-name`). Name your workers: a quick-start WITHOUT
`sessionName` leaves the Codeman name empty, so there is nothing to pass and the
peer name stays derived. The flag is fail-closed (older/unknown CLI omits it, because an
peer name stays derived. ⚠️ Give them a DESCRIPTIVE name: only a name the user chose
is pinned (`Session.cliPinnedName`), because `--name` is also the conversation's
`/resume` title and terminal title and suppresses Claude's own generated title. A
placeholder-shaped name (`w9-msgtest`, anything matching `isGeneratedSessionName`)
and an auto name are NOT passed, so such a worker's peer name is derived; use
`msgtest-worker` rather than `w9-msgtest`. The flag is fail-closed (older/unknown CLI omits it, because an
unknown flag aborts startup and would kill every spawn) and allowlist-sanitized (a name of
only unsafe characters is dropped), and the docker/remote builders never see it at all
(`tmux-manager.ts:782-789`), which is why the `tmux` column stays the canonical join key
@@ -196,8 +201,8 @@ idle:
The contract an orchestrator follows for any fleet of two or more messaging workers.
Every topology in the next section is this protocol plus a wiring diagram.
1. **Spawn with a name, and confirm hooks.** Use `quick-start` with `sessionName` (the
`--name` gate above). Session create installs the hooks block into the workspace
1. **Spawn with a name, and confirm hooks.** Use `quick-start` with a descriptive,
non-`w<N>-` `sessionName` (the `--name` gate above). Session create installs the hooks block into the workspace
whatever kind it is, so a linked case and a raw `POST /api/sessions` path both get
`stop`/`blocked` by default. ⚠️ Not unconditionally: the operator can turn
`workspaceHooksEnabled` off, remote SSH sessions never get hooks, and a session from
+3 -2
View File
@@ -692,8 +692,9 @@ The shape, each step verified live (probes, failure modes and safety detail in
[§5.2](#52-readiness)).
2. `ListAgents`: find the worker's row by its `tmux codeman-<first 8 of session id>`
column; the row's `name [ref]` is the address. On Codeman 1.16+ with claude
2.1.224+ a worker's peer name is its Codeman session name, so pass `sessionName`
in quick-start to pick it; older setups list a name derived from the case folder.
2.1.224+ a worker's peer name is its Codeman session name, so pass a DESCRIPTIVE
`sessionName` in quick-start to pick it (a `w<N>-` placeholder-shaped name is not
pinned, so it lists derived); older setups list a name derived from the case folder.
No row = messaging is off for that worker (it is feature-flagged even on matching
CLI versions, observed live): fall back to the HTTP recipes without complaint.
3. `SendMessage` the task; first contact must use the `name [ref]` form copied from
+45
View File
@@ -0,0 +1,45 @@
/**
* @fileoverview Carry a Codeman rename into Claude Code's own session title.
*
* Claude Code keeps a conversation's title in its transcript as a
* `{"type":"custom-title"}` row (what `/rename` writes), last row wins, and the
* `/resume` picker shows `customTitle ?? aiTitle`. Renaming a tab in Codeman
* used to change only the tab, so `/resume` kept listing the old name.
*
* Appending the row is enough for a pane that was spawned WITHOUT `--name`
* (every placeholder- or auto-named tab, see `Session.cliPinnedName`): that
* process holds no title of its own and never writes one back. A process that
* WAS spawned with `--name` re-appends its in-memory title after each turn, so
* there the new title holds from the next spawn, which pins the new name.
*
* @module claude-session-title
*/
import fs from 'node:fs/promises';
/**
* Append a `custom-title` row for `conversationId` to an existing transcript.
* Never creates the file: a missing transcript means the conversation has not
* been written yet, and a file of only a title row would show up in `/resume`
* as an empty conversation. Returns whether a row was written.
*/
export async function appendClaudeCustomTitle(
transcriptPath: string,
conversationId: string,
title: string
): Promise<boolean> {
const customTitle = title.trim();
// Claude reads the row through `customTitle ?? aiTitle`, so an empty string
// would blank the picker entry rather than fall back to the generated title.
if (!customTitle) return false;
try {
if (!(await fs.stat(transcriptPath)).isFile()) return false;
} catch {
return false;
}
// One O_APPEND write of one line, the same way Claude appends its own rows,
// so it cannot interleave with a row the live process is writing.
const row = JSON.stringify({ type: 'custom-title', customTitle, sessionId: conversationId });
await fs.appendFile(transcriptPath, `${row}\n`);
return true;
}
+10
View File
@@ -129,6 +129,8 @@ program
/** Same registry the server resolves case names through (mirrors `case-routes.ts`). */
const LINKED_CASES_FILE = dataPath('linked-cases.json');
/** Graceful shutdown budget before the process force-exits (see the SIGTERM handler). */
const SHUTDOWN_FORCE_EXIT_MS = 10_000;
/**
* Case name to directory, checking `linked-cases.json` FIRST and falling back to the
@@ -1002,6 +1004,14 @@ webCmd.action(async (options) => {
if (shuttingDown) return;
shuttingDown = true;
console.log(palette.warn(`\n${signal} received, shutting down gracefully...`));
// A hung stop() must not keep the process alive: the listener is already
// closed by then, and a KeepAlive LaunchDaemon only respawns the server once
// it EXITS (systemd would SIGKILL after TimeoutStopSec; launchd does not).
// Seen after a self-update on macOS: port closed, process alive, service down.
setTimeout(() => {
console.error(palette.err(`Shutdown did not finish in ${SHUTDOWN_FORCE_EXIT_MS / 1000}s, forcing exit`));
process.exit(1);
}, SHUTDOWN_FORCE_EXIT_MS).unref();
try {
await server.stop();
} catch (err) {
+115
View File
@@ -0,0 +1,115 @@
/**
* @fileoverview Write side of the CLI registry (docs/cli-enable-disable-plan.md, Phases 3/5).
*
* Kept deliberately SEPARATE from `registry.ts`, whose reading path does no writes on import
* (`schemas.ts` imports it, transitively). Only `cli-registry-routes.ts` imports this module,
* so that property still holds for every OTHER importer of the registry.
*
* Every mutation goes through `mutateRegistryFile()`, which does three things the #476 review
* found missing:
*
* - **Serialized.** Mutations run one at a time on a single promise chain, and each one
* reads, changes, writes and reloads before the next starts. Unserialized read-modify-write
* lost toggles when three `PUT /api/clis/:id` calls ran in parallel.
* - **Refuses a file it must not trust.** The reader ignores a `clis.json` with any
* group/world permission bit and quarantines one that does not parse. The writer used to
* treat both as "start fresh", so one Settings click replaced a hand-edited file with a
* one-key file, or rewrote a refused file as 0600 and so trusted it. It now starts fresh
* ONLY on ENOENT and otherwise throws `RegistryWriteRefusedError`, leaving the file alone.
* - **Unique temp file.** Every write gets its own tmp name before the rename, so two writes
* can never rename each other's temp file away (the ENOENT-on-rename 500s).
*
* Same tmp+rename+0600 shape as `custom-model-hosts.ts`. The file is hand-editable, so a
* write must never leave it half-written, and 0600 is the mode `isUnsafePermissions()`
* requires on the next read.
*/
import { randomUUID } from 'node:crypto';
import { existsSync, mkdirSync } from 'node:fs';
import fs from 'node:fs/promises';
import { dirname } from 'node:path';
import { isUnsafePermissions, registryFilePath, reloadCliRegistry } from './registry.js';
import type { CliRegistryFile } from './types.js';
/** A write refused because the existing `clis.json` must not be overwritten. The message is user-facing. */
export class RegistryWriteRefusedError extends Error {
constructor(message: string) {
super(message);
this.name = 'RegistryWriteRefusedError';
}
}
/**
* Read the raw override file for mutation. Only a MISSING file starts fresh. A file with
* unsafe permissions, one that cannot be read, or one that does not parse is refused rather
* than overwritten, because the user's hand-edit is worth more than one toggle.
*/
export async function readRegistryFileForWrite(): Promise<CliRegistryFile> {
const path = registryFilePath();
let raw: string;
try {
raw = await fs.readFile(path, 'utf-8');
} catch (err) {
if ((err as NodeJS.ErrnoException).code === 'ENOENT') return { schemaVersion: 1, clis: {} };
throw new RegistryWriteRefusedError(`Cannot read ${path} (${(err as Error).message}); not changing it.`);
}
if (isUnsafePermissions(path)) {
throw new RegistryWriteRefusedError(
`${path} has group/world permission bits, so Codeman ignores it. Run \`chmod 600 ${path}\` and check its contents before changing CLIs here.`
);
}
let parsed: unknown;
try {
parsed = JSON.parse(raw);
} catch (err) {
throw new RegistryWriteRefusedError(
`${path} is not valid JSON (${(err as Error).message}). Fix or remove it before changing CLIs here.`
);
}
const clis = (parsed as { clis?: unknown } | null)?.clis;
if (typeof parsed !== 'object' || parsed === null || typeof clis !== 'object' || clis === null) {
throw new RegistryWriteRefusedError(`${path} has no "clis" object. Fix or remove it before changing CLIs here.`);
}
return parsed as CliRegistryFile;
}
export async function writeRegistryFile(file: CliRegistryFile): Promise<void> {
const target = registryFilePath();
const dir = dirname(target);
if (!existsSync(dir)) mkdirSync(dir, { recursive: true });
const tmp = `${target}.${process.pid}.${randomUUID()}.tmp`;
try {
await fs.writeFile(tmp, JSON.stringify(file, null, 2), { mode: 0o600 });
await fs.rename(tmp, target);
} catch (err) {
await fs.rm(tmp, { force: true }).catch(() => {});
throw err;
}
}
let mutationChain: Promise<unknown> = Promise.resolve();
/**
* Run one registry mutation. The chain holds exactly one at a time: `fn` receives the
* current file and returns `{ file, result }`. If `file` is set it is written and the
* registry reloaded before the next mutation starts; if not, nothing is written, which is
* how a validation failure returns early. Checks made inside `fn` (does this id exist,
* is it a duplicate) therefore see every earlier mutation's result.
*
* A failed mutation rejects its own caller only. The chain keeps going.
*/
export function mutateRegistryFile<T>(
fn: (file: CliRegistryFile) => Promise<{ file?: CliRegistryFile; result: T }> | { file?: CliRegistryFile; result: T }
): Promise<T> {
const run = mutationChain.then(async () => {
const current = await readRegistryFileForWrite();
const { file, result } = await fn(current);
if (file) {
await writeRegistryFile(file);
reloadCliRegistry();
}
return result;
});
mutationChain = run.catch(() => {});
return run;
}
+14 -1
View File
@@ -48,6 +48,16 @@ function filePath(): string {
return dataPath('clis.json');
}
/**
* The resolved path of `~/.codeman/clis.json`, exported for the write API
* (`cli-registry-writer.ts`, docs/cli-enable-disable-plan.md Phases 3/5) so both the read and
* write sides resolve the SAME path through the SAME instance-scoped helper — never a second
* `dataPath('clis.json')` call that could drift from this one under a future `dataPath()` change.
*/
export function registryFilePath(): string {
return filePath();
}
/**
* Keys that must never be merged out of a hand-editable JSON file.
*
@@ -90,8 +100,11 @@ export interface LoadResult {
* as mode 0o666 there regardless of its actual ACL), so this check would flag every file on
* Windows and silently ignore all user config. `win32` relies on NTFS ACLs instead, which
* this check cannot see and does not attempt to.
*
* Exported for `registry-writer.ts`, which must refuse the same files: rewriting a refused
* file as 0600 would silently turn it into trusted config.
*/
function isUnsafePermissions(path: string): boolean {
export function isUnsafePermissions(path: string): boolean {
if (process.platform === 'win32') return false;
try {
const mode = statSync(path).mode & 0o777;
+67
View File
@@ -294,6 +294,23 @@ const capabilitiesSchema = z
effort: z.boolean(),
agentSkillInjection: z.boolean(),
statusLineTelemetry: z.boolean(),
// How many columns this CLI indents its transcript body by, so a copy can take
// that much off the clipboard. Bounded, because it is the whole strip: a copy
// never removes more than this, nor more than every selected line shares.
//
// ⚠ DECLARED, not measured off the pane, and two measured attempts are why.
// Asking whether the pane painted spaces across the unused part of each row
// separates a TUI from a shell perfectly where it fires and never
// over-stripped, but it is a function of pane WIDTH: that padding exists
// only while a rendered line stops short of the CLI's own layout width, and
// Claude Code's prose wraps to fill it — the share of padded rows on one
// live transcript ran 44%, 6%, 6%, 7% and 87% at 123, 160, 198, 235 and 298
// columns, so the strip did nothing at any ordinary size. Taking the
// narrowest indent on screen instead fires everywhere and over-strips, since
// a file listing inside the transcript can be the narrowest thing on it.
// A declared width cannot do either. Absent means no strip, so a CLI whose
// transcript layout nobody has measured is never touched.
transcriptGutter: z.number().int().min(1).max(8).optional(),
workDetect: z
.object({
promptGlyph: z.string().min(1).max(8),
@@ -308,8 +325,29 @@ const capabilitiesSchema = z
(src) => compileVersionRegex(src) !== null,
'workingLine must be a regex compileVersionRegex() accepts: at most 200 characters, no nested quantifiers'
),
// Same guard, same reasons: this one runs over the foot of a pane capture every
// time a session settles, and ~/.codeman/clis.json can set it.
watchingLine: z
.string()
.min(1)
.refine(
(src) => compileVersionRegex(src) !== null,
'watchingLine must be a regex compileVersionRegex() accepts: at most 200 characters, no nested quantifiers'
)
.optional(),
// Bounded hard: this is how far up the screen a config file may push the search,
// and every row it adds is one more row the agent itself may be able to write.
watchingLines: z.number().int().min(1).max(8).optional(),
})
.strict()
// A window with nothing to search is a typo, not a configuration. Refused at LOAD
// time for the same reason `privilegedParams[].param` is checked against the params
// the entry declares: the failure is otherwise silent and looks like a feature that
// simply never fires.
.refine(
(v) => v.watchingLines === undefined || v.watchingLine !== undefined,
'watchingLines has nothing to bound without a watchingLine'
)
.optional(),
model: z
.object({ source: z.enum(['flag', 'claude-settings-file', 'none']), param: z.string().optional() })
@@ -340,6 +378,35 @@ const capabilitiesSchema = z
// an env var, so it declares baseUrl/apiKey injection with no model var at all.
modelVars: z.array(envName).max(8),
launchModel: launchModelTemplate,
// Optional: the env var to carry a discovered per-model context-window size
// (claude's CLAUDE_CODE_MAX_CONTEXT_TOKENS), and/or the env var that isolates
// this session's config/credential directory from the user's real one (claude's
// CLAUDE_CONFIG_DIR) so an injected API key never collides with a stored OAuth
// session. See the customModelInjection doc comment in cli-registry/types.ts.
contextLengthVar: envName.optional(),
configDirVar: envName.optional(),
// Relative path, WITHIN the isolated configDirVar directory, of a trust-dialog
// seed file the CLI itself owns the shape of — claude's `.claude.json`
// `customApiKeyResponses.approved` list, the same field an interactive "Detected
// a custom API key — use it?" prompt writes to on a real terminal. Only makes
// sense alongside configDirVar (an isolated, otherwise-empty directory has none
// of a real profile's prior approvals), and only implemented for the
// 'claude-api-key-responses' shape today — see custom-model-injection-apply.ts.
apiKeyTrustFile: z
.object({ relPath: z.string().min(1).max(80), shape: z.literal('claude-api-key-responses') })
.strict()
.optional(),
// An isolated config directory replays the CLI's whole first-run sequence (theme
// picker, security notes, per-project trust dialog, bypass-permissions warning)
// on every launch, same root cause as apiKeyTrustFile above — this reuses that
// same file to pre-seed the state a real, already-onboarded profile carries. See
// the customModelInjection doc comment in cli-registry/types.ts.
skipFirstRunPrompts: z.boolean().optional(),
// DeepSeek-only, confirmed by reading its own bundled SDK source: it concatenates
// "/chat/completions" onto baseUrlVar's value with no "/v1" of its own, while
// llama-swap/llama.cpp only serves the "/v1/..." path — claude/gemini must NOT
// get this. See the customModelInjection doc comment in cli-registry/types.ts.
appendV1Suffix: z.boolean().optional(),
})
.strict(),
z
+145 -21
View File
@@ -75,11 +75,24 @@ function agentDefaults(): Pick<
};
}
// `accent` on every entry below (except SHELL, which the frontend renders no
// distinct color for) is measured from the actual `.btn-toolbar.btn-run.mode-<id>`
// CSS rule's `border-color` on the OG skin (styles.css) — the single cleanest
// representative hex each entry's own multi-stop gradient resolves around.
// Corrected 2026-09-21 after PR #458's review found several were simply wrong
// (e.g. claude was registered as Anthropic's brand orange, `#d97757`, but the
// button renders blue): `docs/cli-registry.md`'s own "transcribed, not
// authoritative, re-measure before wiring one up" warning for this
// DECLARED-FOR-LATER field, taken literally. The one exception is GEMINI, whose
// run-button border (#60a5fa) is the only one that disagrees with its own tab badge
// and run-mode dot (#8ab4f8); it takes the badge colour, so every accent names the
// same hex the frontend uses as that CLI's flat identity. This is a data-accuracy fix only —
// `accent` still has no reader, so nothing rendered changes because of it.
const CLAUDE: CliEntry = {
id: 'claude' as CliEntry['id'],
label: 'Claude',
label: 'Claude Code',
shortBadge: 'CC',
accent: '#d97757',
accent: '#3b82f6',
enabled: true,
stock: true,
order: 0,
@@ -200,12 +213,31 @@ const CLAUDE: CliEntry = {
},
capabilities: {
external: false,
// Claude indents its transcript body two columns and puts its own ●/✻/❯ markers
// in them, so a copy can drop two and paste flush. Claude and codex are the only
// entries that declare this, because theirs are the only gutters that have been measured.
transcriptGutter: 2,
// The historical hard-coded pair, now stated as data. `workingLine` matches both the
// `✻ Actualizing… (39s · ↓ 2.0k tokens)` status line and the bare `esc to interrupt`
// footer, because tmux repaints partially and only one of the two may land in a chunk.
workDetect: {
promptGlyph: '❯',
workingLine: String.raw`…\s*\((?:\d+h\s+)?(?:\d+m\s+)?\d+s\b|esc to interrupt`,
// Claude prints what it started in the background on the footer row beneath its
// composer, as `⏵⏵ bypass permissions on · 1 monitor · ← for agents`. The labels are
// the CLI's own words for each kind of background task, and group 1 is the one
// Codeman badges the session with. Verified against a live 2.1.278 pane on
// 2026-09-21.
// ⚠️ Two things keep an agent from writing its own label here, and both matter.
// The footer is the LAST row, so the default one-row window (`WATCHING_TAIL_LINES`)
// holds nothing but Ink's own chrome — in particular it leaves out the status line
// directly above, whose content comes from a `statusLine` command a bypassed
// session can write into its own `.claude/settings.json`. And the leading `·` keeps
// the match on the footer's own item list rather than on any text that happens to
// carry a count. A footer that ever drew the chip as its only item would report no
// watching rather than open that door. See `watchingLabel()` in
// `session-activity.ts`.
watchingLine: String.raw`·\s*(\d+ (?:monitors?|shells?|teams?|local agents?|cloud sessions?|MCP tasks?|background tasks?|(?:background|remote) dynamic workflows?|Artifact comment monitors?))`,
},
requiresMux: false,
// Claude installs Codeman's own hooks block into every workspace it runs in, so its
@@ -228,15 +260,31 @@ const CLAUDE: CliEntry = {
privilegedParams: [],
// ANTHROPIC_* is NOT in allowedPrefixes/allowedKeys above (deliberately — see the
// allowedPrefixes comment nearby), so these are unreachable via plain envOverrides
// today; listed here only so the dedicated custom-model route (docs/custom-model-endpoints-plan.md
// chunk 5) clamps them for a non-granted multi-user owner the same way every other
// CLI's injection vars are clamped, the day that route widens who can set them.
// today. privilegedEnvKeys has exactly one consumer, ownerClampedEnvKeys() in
// session-env-clamp.ts, which feeds the generic envOverrides clamp on
// POST /api/sessions, POST /api/quick-start and reboot-restore — no custom-model
// route reads this field at all, and the values it injects are merged in AFTER
// that clamp runs regardless of what's listed here.
privilegedEnvKeys: [
'ANTHROPIC_BASE_URL',
'ANTHROPIC_API_KEY',
'ANTHROPIC_DEFAULT_SONNET_MODEL',
'ANTHROPIC_DEFAULT_HAIKU_MODEL',
'ANTHROPIC_DEFAULT_OPUS_MODEL',
// CLAUDE_CODE_MAX_CONTEXT_TOKENS already matches the CLAUDE_CODE_* allowedPrefix, and
// CLAUDE_CONFIG_DIR is already an allowed exact key (docs/wiki/Agent-CLIs.md), so both
// were already reachable via plain envOverrides before this pair existed and this
// feature does not strictly need either listed. They stay listed anyway, because
// types.ts's rule ("every traffic-redirecting var this feature introduces MUST also
// appear in privilegedEnvKeys") is meant to hold literally, not with an exception
// carved out for the two vars that happen not to need it today. The real
// consequence lands on the GENERIC envOverrides clamp above, not on this feature:
// a non-granted multi-user owner can no longer set CLAUDE_CONFIG_DIR through
// envOverrides at all (the per-client-account override, #255), and a PERSISTED one
// is now stripped on reboot-restore for such an owner too — see
// session-env-clamp.ts's own fileoverview.
'CLAUDE_CODE_MAX_CONTEXT_TOKENS',
'CLAUDE_CONFIG_DIR',
],
gates: { nameFlag: { minVersion: '2.1.224', failClosed: true } },
// Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md) — verified by hand against a real
@@ -247,6 +295,32 @@ const CLAUDE: CliEntry = {
baseUrlVar: 'ANTHROPIC_BASE_URL',
apiKeyVar: 'ANTHROPIC_API_KEY',
modelVars: ['ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL'],
// Verified via Claude Code's own docs: CLAUDE_CODE_MAX_CONTEXT_TOKENS overrides the
// assumed context window and applies directly for a model name Claude Code doesn't
// recognize as one of its own — exactly the custom-model case. Without it, Claude Code
// assumes a large (200k) window for any unrecognized model id and never compacts,
// eventually overflowing a much smaller real local context (see plan doc reasoning
// above the interface for the confirmed failure).
contextLengthVar: 'CLAUDE_CODE_MAX_CONTEXT_TOKENS',
// Isolates this session's config/credential directory so an injected ANTHROPIC_API_KEY
// never shares a directory with a stored claude.ai OAuth login — see the doc comment on
// customModelInjection in cli-registry/types.ts for the traded-off side effect.
configDirVar: 'CLAUDE_CONFIG_DIR',
// ⚠️ Required alongside configDirVar, not optional in practice: verified live that an
// isolated, otherwise-empty config directory makes claude stop at an interactive
// "Detected a custom API key — use it?" prompt on EVERY launch, defaulting to "No" with
// no one at the TTY to answer — silently refusing the very key this feature injected.
// Pre-seeding this file's customApiKeyResponses.approved list (verified against a real
// ~/.claude.json after answering the prompt once by hand) answers it in advance instead.
apiKeyTrustFile: { relPath: '.claude.json', shape: 'claude-api-key-responses' },
// ⚠️ Same isolated-directory root cause, one step further: verified live that on top
// of the API-key prompt above, a fresh CLAUDE_CONFIG_DIR also replays claude's ENTIRE
// first-run sequence on every launch — the theme picker, the security-notes screen,
// the per-project "trust this folder?" dialog, and (running with
// --dangerously-skip-permissions) a one-time bypass-permissions warning — none of
// which a real, already-onboarded profile shows again. Pre-seeds that same
// already-onboarded state instead of leaving a human to click through it.
skipFirstRunPrompts: true,
},
},
overlays: {
@@ -326,7 +400,7 @@ const OPENCODE: CliEntry = {
id: 'opencode' as CliEntry['id'],
label: 'OpenCode',
shortBadge: 'OC',
accent: '#f59e0b',
accent: '#10b981',
enabled: true,
stock: true,
order: 10,
@@ -412,7 +486,7 @@ const CODEX: CliEntry = {
id: 'codex' as CliEntry['id'],
label: 'Codex',
shortBadge: 'CX',
accent: '#6b7fd7',
accent: '#a855f7',
enabled: true,
stock: true,
order: 20,
@@ -476,7 +550,43 @@ const CODEX: CliEntry = {
// `Working (2m 49s • esc to interrupt)` above it while a turn runs. It animates no
// braille spinner, and it never prints `esc to interrupt` at rest, so that phrase
// alone separates a running turn from an idle one.
workDetect: { promptGlyph: '›', workingLine: '[Ee]sc to interrupt' },
// Codex pins a row of its own while a background terminal it started is still
// running: ` 1 background terminal running · /ps to view · /stop to close`. Unlike
// Claude's footer chip that row sits ABOVE the composer, which puts it third from the
// bottom once the status line and the composer are counted, hence `watchingLines`.
// Measured against a live codex-cli 0.154.0 pane on 2026-09-22: the row appears when
// the terminal starts, follows the composer down as the conversation grows, and is
// gone after `/stop`.
// ⚠️ This entry CANNOT promise what Claude's does, and the difference is Codex's
// layout rather than its pattern. The third row from the bottom is the chip only
// while a terminal runs; with none running it is the last row of the transcript,
// which the agent writes. Matching the complete row raises the bar — an assistant
// message has to end with this exact line, to the character — but nothing here makes
// forging it impossible, so do not read the Claude comment above as applying here.
// What contains it is that codex declares `hooks: 'none'`: no hook event from a codex
// session ever reaches `notePrompt()`, so there is no idle item to pre-acknowledge
// and a forged label costs a wrong badge and nothing else. A CLI that gains hook
// signals must not keep a pattern this soft.
// ⚠️ Background TERMINALS are the only background work codex advertises on screen.
// A sub-agent started without waiting outlives the turn just as a terminal does —
// measured 2026-09-22, the sandboxed process was still running — and the pane shows
// nothing at all for it: the last rows are the composer and the status line, and
// `Sub-agents running` lives in the on-demand `/subagents` panel, not above the
// composer. So a codex session waiting on a sub-agent reads as plainly idle here.
// Nothing is misfiled by that (codex raises no idle prompts), and there is no row to
// match until codex pins one.
workDetect: {
promptGlyph: '›',
workingLine: '[Ee]sc to interrupt',
watchingLine: String.raw`^\s{0,4}(\d+ background terminals?) running · /ps to view · /stop to close$`,
watchingLines: 3,
},
// Two columns, like claude's, measured on a live 0.154.0 answer: the `•`/`›`/`⚠`
// markers sit in the gutter, prose continuations sit at 2, and a nested YAML block
// the model wrote rendered at 2/4/6/8 for its own 0/2/4/6. Replayed at 100, 120,
// 160, 198, 235 and 282 columns the indents were 0, 2, 4, 6 and 8 at every one,
// never 1, so the width is not a function of the pane.
transcriptGutter: 2,
transcript: 'codex-rollout',
altScreen: 'strip-full',
echo: { policy: 'predict', anchor: { kind: 'cursor' }, predictProfile: 'codex' },
@@ -523,7 +633,8 @@ const GEMINI: CliEntry = {
id: 'gemini' as CliEntry['id'],
label: 'Gemini',
shortBadge: 'GM',
accent: '#4285f4',
// The tab badge / run-mode-dot colour, not the run-button border (see the note above CLAUDE).
accent: '#8ab4f8',
enabled: true,
stock: true,
order: 30,
@@ -615,7 +726,7 @@ const ANTIGRAVITY: CliEntry = {
id: 'antigravity' as CliEntry['id'],
label: 'Antigravity',
shortBadge: 'AG',
accent: '#8b5cf6',
accent: '#22d3ee',
enabled: true,
stock: true,
order: 40,
@@ -685,7 +796,7 @@ const PI: CliEntry = {
id: 'pi' as CliEntry['id'],
label: 'Pi',
shortBadge: 'PI',
accent: '#10b981',
accent: '#f472b6',
enabled: true,
stock: true,
order: 50,
@@ -809,10 +920,10 @@ const GROK: CliEntry = {
shortBadge: 'GK',
// Upstream hand-authored a charcoal GRADIENT across 4+ CSS spots (welcome button, tab
// badge, run-mode dot, mobile skin overrides) rather than one flat colour; our registry's
// `accent` is a single hex, so this is the closest single value (the run-mode-dot colour,
// zinc-400). Nothing reads `accent` yet — the frontend is untouched in this change and
// keeps its own hand-authored CSS; the field is here so the entry is complete.
accent: '#a1a1aa',
// `accent` is a single hex, so this is the closest single value (zinc-300, the run-button
// border and tab-badge colour). Nothing reads `accent` yet: the frontend keeps its own
// hand-authored CSS; the field is here so the entry is complete.
accent: '#d4d4d8',
enabled: true,
stock: true,
order: 70,
@@ -944,7 +1055,7 @@ const DEEPSEEK: CliEntry = {
id: 'deepseek' as CliEntry['id'],
label: 'DeepSeek',
shortBadge: 'DS',
accent: '#4d6bfe',
accent: '#7c93ff',
enabled: true,
stock: true,
order: 80,
@@ -1071,15 +1182,28 @@ const DEEPSEEK: CliEntry = {
// privilege rather than granting it, and clamping it here was a real regression
// (test/deepseek-mode.test.ts) fixed before this shipped.
privilegedEnvKeys: ['DSH_PERMISSION_MODE', 'DSH_HOME', 'DEEPSEEK_BASE_URL'],
// Web-researched, unverified, partial: reuses the already-existing DEEPSEEK_BASE_URL/
// DEEPSEEK_API_KEY keys above. No modelVars — dsh's model is a profile-composition
// entry (see `model: { source: 'none' }` above), not an env var, so forcing a specific
// model name may not fully work; verify against a real profile before shipping.
// Reuses the already-existing DEEPSEEK_BASE_URL/DEEPSEEK_API_KEY keys above. No
// modelVars — dsh's model is a profile-composition entry (see `model: { source: 'none'
// }` above), not an env var, so forcing a specific model name may not fully work;
// verify against a real profile before shipping.
//
// ⚠️ appendV1Suffix is REQUIRED, not optional-nice-to-have: without it every request
// 404s. Confirmed live and by reading dsh's own bundled source
// (@deepseek-ai/dsh-llm-deepseek): it builds the request URL as
// `${DEEPSEEK_BASE_URL}/chat/completions` with no "/v1" of its own (its real public
// API, https://api.deepseek.com, expects the caller's base URL to already carry any
// needed prefix), while llama-swap/llama.cpp only serves the OpenAI-conventional
// "/v1/chat/completions" — a bare POST to ".../chat/completions" 404s live, and the
// 404 reported here originally ("dsh: HTTP_404: DeepSeek API error (HTTP 404)")
// matches dsh's own error-message template for exactly this failure. See the
// customModelInjection doc comment in cli-registry/types.ts for the full reasoning,
// including why claude/gemini must NOT get this.
customModelInjection: {
kind: 'env',
baseUrlVar: 'DEEPSEEK_BASE_URL',
apiKeyVar: 'DEEPSEEK_API_KEY',
modelVars: [],
appendV1Suffix: true,
},
},
overlays: {
@@ -1096,7 +1220,7 @@ const OMP: CliEntry = {
id: 'omp' as CliEntry['id'],
label: 'OMP',
shortBadge: 'OM',
accent: '#7c9cf5',
accent: '#818cf8',
enabled: true,
stock: true,
order: 90,
+120 -6
View File
@@ -344,7 +344,48 @@ export interface CliCapabilities {
promptGlyph: string;
/** Source of a regex matching the status line this CLI draws while a turn runs. */
workingLine: string;
/**
* Source of a regex matching the row this CLI draws while work it started in the
* background is still running, e.g. Claude's `· 1 monitor ·` footer chip or Codex's
* `1 background terminal running · /ps to view`. Capture group 1 is the label Codeman
* shows, and the whole match stands in when the pattern declares no group. A CLI that
* omits this reports no background work, which is what every CLI did before the field
* existed.
*/
watchingLine?: string;
/**
* How many rows at the FOOT of the screen that row can appear in, counting non-blank
* rows only. Claude writes its chip on the last row and keeps the default; Codex pins
* its own above the composer, which puts it third from the bottom, so it declares
* more. Keep each number as small as that CLI's layout allows: every extra row is
* another row an agent might be able to write, and the label is what silences an
* alert. See `watchingLabel()` in `session-activity.ts`.
*/
watchingLines?: number;
};
/**
* How many columns this CLI indents its transcript body by, so a copy taken from its
* pane can drop that much and paste flush. Claude Code indents two and puts its own
* markers in those columns.
*
* ⚠ DECLARED rather than measured off the pane, and two measured attempts are why.
* Asking whether the pane painted real spaces across the unused part of each row
* separates a TUI from a shell perfectly where it fires and never over-stripped; it
* is also a function of pane WIDTH, because that padding exists only while a
* rendered line stops short of the CLI's own layout width and Claude Code's prose
* wraps to fill it. On one live transcript the share of padded rows ran 44%, 6%, 6%,
* 7% and 87% at 123, 160, 198, 235 and 298 columns, so at any ordinary window size
* the strip silently did nothing. Taking the narrowest indent on the surrounding
* rows instead fires at every width and over-strips on roughly 1% of selections,
* because a file listing inside the transcript can be the narrowest thing on screen.
*
* A declared width can do neither. The strip is the lesser of this and what every
* selected line shares, so a block can only ever shift as a unit, and it can never
* shift further than the CLI itself says its gutter is.
*
* Absent means no strip at all, the same fail-safe direction `workDetect` takes.
*/
transcriptGutter?: number;
/** No direct-PTY fallback: the CLI must run inside tmux (secrets ride tmux setenv). */
requiresMux: boolean;
/**
@@ -496,9 +537,74 @@ export interface CliCapabilities {
* declares). Absent = the config alone selects the model (claude's env vars,
* opencode's blob, codex's top-level `model` key). Applied by the session's
* respawn options through the entry's `legacyConfigField`, never by id.
*
* `contextLengthVar` (env kind only): the env var a discovered per-model context-window
* size is written to when known (claude's `CLAUDE_CODE_MAX_CONTEXT_TOKENS`) — without it,
* a CLI that assumes a large default window for an unrecognized model name keeps sending
* full-size prompts against a much smaller local server and eventually overflows its real
* context (verified: a 33.7K-token system prompt against a 16384-token llama-swap model).
* Absent when the CLI has no such override, or the value is unknown for this model.
*
* `configDirVar` (env kind only): the env var that redirects this session's config/
* credential directory to an isolated, per-session one (claude's `CLAUDE_CONFIG_DIR`), so
* an injected API key never coexists with a stored claude.ai OAuth session in the same
* directory — the CLI still warns "both claude.ai and ANTHROPIC_API_KEY set" when they
* share a directory even though the API key wins for actual requests. Isolating it trades
* that cosmetic warning for a documented side effect: a relocated config directory writes
* transcripts outside `~/.claude/projects`, blinding the response viewer, subagent
* windows, and Read My Mind for that session (see docs/wiki/Agent-CLIs.md).
*
* `apiKeyTrustFile` (env kind only, alongside configDirVar): an isolated config directory
* has none of a real profile's prior "detected a custom API key, use it?" approvals, so
* without this the CLI stops and asks interactively on every single launch — with no one
* at a TTY to answer, that's a hang, not a warning (confirmed live: claude's own default
* answer, "No", would silently refuse to use the very key this feature just injected).
* `relPath`/`shape` name the file (claude's `.claude.json`) and its
* `customApiKeyResponses.approved` field this pre-seeds — the exact field a real answered
* prompt itself writes to, so this isn't bypassing the check, just answering it the same
* way a one-off prior approval on a shared profile already would.
*
* `skipFirstRunPrompts` (env kind only, alongside apiKeyTrustFile): an isolated config
* directory is not just missing API-key approvals — it is a brand-new profile as far as
* the CLI is concerned, so it also replays its ENTIRE first-run sequence on every launch:
* the theme picker, the security-notes screen, the per-project "trust this folder?"
* dialog, and (running with a bypass-permissions flag) a one-time warning about it —
* confirmed live, none of which a real, long-used profile ever shows again. `true`
* pre-seeds the same state a real profile accumulates from having answered all of that
* once: `hasCompletedOnboarding` and the launching session's own project entry in the
* `apiKeyTrustFile` (claude's `.claude.json`), plus `skipDangerousModePermissionPrompt`
* in claude's `settings.json` — see `seedFirstRunState`/`seedSkipBypassPermissionsPrompt`
* in custom-model-injection-apply.ts. Requires `apiKeyTrustFile` to be set too, since it
* reuses that file.
*
* `appendV1Suffix` (env kind only): the raw `endpoint.baseUrl` gets `withV1Suffix()`
* applied before being written to `baseUrlVar`, instead of being used verbatim.
* DeepSeek needs this and claude/gemini must NOT get it — a per-CLI asymmetry confirmed
* by reading each SDK's own request-building source, not assumed: DeepSeek Harness's
* bundled `@deepseek-ai/dsh-llm-deepseek` concatenates `${connection.baseURL}/chat/
* completions` with no `/v1` insertion of its own (its real public API base,
* `https://api.deepseek.com`, expects the caller's base URL to already carry any
* needed prefix), while llama-swap/llama.cpp only ever serves the OpenAI-conventional
* `/v1/chat/completions` — confirmed live: a bare `POST <baseUrl>/chat/completions`
* 404s, `POST <baseUrl>/v1/chat/completions` succeeds, and the harness's own error
* message template (`DeepSeek API error (HTTP ${status})`) reproduces the exact
* `HTTP_404` this feature originally shipped with unexplained. Claude Code's own SDK,
* by contrast, was already confirmed working end-to-end against the RAW `baseUrl` with
* no suffix — appending one there would be wrong, not just redundant.
*/
customModelInjection:
| { kind: 'env'; baseUrlVar: string; apiKeyVar: string; modelVars: string[]; launchModel?: string }
| {
kind: 'env';
baseUrlVar: string;
apiKeyVar: string;
modelVars: string[];
launchModel?: string;
contextLengthVar?: string;
apiKeyTrustFile?: { relPath: string; shape: 'claude-api-key-responses' };
configDirVar?: string;
skipFirstRunPrompts?: boolean;
appendV1Suffix?: boolean;
}
| { kind: 'configContentEnv'; envVar: string; template: 'opencode-json'; launchModel?: string }
| {
kind: 'configDir';
@@ -564,12 +670,14 @@ export interface CliOverlays {
/**
* ⚠️ DECLARED-FOR-LATER: fields no code reads yet.
*
* `shortBadge`, `accent`, `overlays.credStore`, `capabilities.echo`, `capabilities.wheelForward`,
* `accent`, `overlays.credStore`, `capabilities.echo`, `capabilities.wheelForward`,
* `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` all describe FRONTEND
* behaviour, and the frontend is deliberately untouched by the change that introduced this
* registry — `app.js`, `terminal-ui.js`, `styles.css` and friends keep their own
* behaviour, and most of the frontend is deliberately untouched by the change that introduced
* this registry — `app.js`, `terminal-ui.js`, `styles.css` and friends keep their own
* hand-authored per-CLI rules, and moving them is its own piece of work with its own way of
* being verified (a mobile/browser suite the CI gate cannot see).
* being verified (a mobile/browser suite the CI gate cannot see). `shortBadge` graduated out of
* this list (docs/cli-enable-disable-plan.md, Phase 2): `GET /api/clis` reads it for the
* CLI-management Settings list.
*
* They are declared now because each entry should describe its CLI completely, and because
* transcribing them while the hand-written source is still on screen is when the values are
@@ -586,7 +694,13 @@ export interface CliEntry {
label: string;
/** Two-ish character tab badge, e.g. 'OC'. */
shortBadge: string;
/** Single hex colour. CSS derives every per-CLI gradient from it via --cli-accent. */
/**
* Single hex colour, measured from the CLI's actual `.btn-toolbar.btn-run.mode-<id>`
* gradient in styles.css (see stock.ts's comment above `CLAUDE` for the exact
* methodology). DECLARED-FOR-LATER (above) — no code reads this yet; styles.css's
* gradients are still hand-authored per id, not derived from this field via any
* CSS custom property. There is no `--cli-accent` variable in the codebase.
*/
accent: string;
enabled: boolean;
/** Set by the loader from the shipped catalog; a user entry can never claim it. */
+3 -1
View File
@@ -69,7 +69,9 @@ const ALL: ProbeEnvironment[] = ['linux', 'darwin', 'wsl', 'win32'];
* shown to the user and claude's does not follow the pattern.
*/
const DOCTOR_ROW_OVERRIDES: Record<string, { id?: string; label?: string; usedBy: string[] }> = {
claude: { usedBy: ['Claude Code sessions (default backend)'] },
// The label override keeps the doctor row's historical "Claude CLI" spelling now that
// the registry label is the product name, "Claude Code".
claude: { label: 'Claude CLI', usedBy: ['Claude Code sessions (default backend)'] },
opencode: { usedBy: ['OpenCode sessions'] },
codex: { usedBy: ['Codex sessions'] },
gemini: { usedBy: ['Gemini sessions'] },
+23
View File
@@ -0,0 +1,23 @@
/**
* @fileoverview Limits shared between Wake-on-LAN parsing and its request schema.
*
* Its own module because `src/remote-wake.ts` is import-fenced: only
* `web/routes/session-routes.ts` and `web/server.ts` may import it, so that no
* watcher or boot-recovery path can WAKE a host (pinned by the wiring guard in
* `test/remote-wake.test.ts`). `web/schemas.ts` needs the same MAC-count limit and
* must not become a third importer, and it would drag `dgram`/`net`/`child_process`
* into every module that validates a request body. A plain constant satisfies both.
*/
/**
* How many comma-separated MACs one `wakeMac` may carry.
*
* ⚠ Single source for `parseMacList()` and `RemoteHostSchema.wakeMac`. The two used
* to disagree: the schema's 128-character cap admits seven MACs while the parser
* rejected more than four all-or-nothing, so a five-MAC value validated, persisted to
* `remote-hosts.json`, and then resolved to NO wake target. The host read as
* unconfigured and the banner offered "Configure WoL" for a host the user had just
* configured, which is the worst shape a validation gap can take: accepted, stored,
* silently inert.
*/
export const MAX_WAKE_MACS = 4;
+32
View File
@@ -40,6 +40,38 @@ export interface CustomModelHost {
authStyle?: CustomModelAuthStyle;
models?: string[];
lastDiscoveredAt?: string;
/**
* The model the Run-menu picker (docs/custom-model-endpoints-plan.md) applies when
* this endpoint is picked with no further choice — one generated menu entry per
* (CLI, endpoint) pair, not per (CLI, endpoint, model), so it needs a single answer.
* Must be a member of `models` when set; the picker falls back to `models[0]` when
* this is unset, and disables the entry entirely when `models` is empty (nothing to
* default to). Never auto-set on discovery — the previous default staying valid
* after a re-discover is a property worth keeping even if the model list changes.
*/
defaultModelId?: string;
/**
* Discovered context-window size (tokens) per model id, keyed by the same strings as
* `models`. Populated opportunistically during discovery (`custom-model-routes.ts`) from
* llama.cpp/llama-swap's `GET /props?model=<id>` — the plain OpenAI-shaped `/v1/models`
* response has no such field. Only ever probed for a model the server already reports as
* loaded (llama-swap's `status.value === 'loaded'`); an unloaded one is deliberately never
* probed, since llama-swap treats `/props?model=` as a routing hint that can trigger an
* actual (slow, GPU-swapping) model load as a side effect of merely asking. A model this
* has no entry for simply gets no context-length env override applied — never a guess.
*/
modelContextLengths?: Record<string, number>;
/**
* Discovered file size (GB) per model id, keyed by the same strings as `models`.
* Populated during discovery by parsing llama-swap's own `description` field for an
* auto-discovered model ("Auto-discovered 16.35 GB - parameters auto-fitted by
* llama.cpp") — a hand-configured profile's own description has no such figure and
* correctly gets no entry, never a guess. Used only to label the Run-menu picker's
* "loading model" banner with a rough, unmeasured expected-time estimate
* (the Run-menu picker's loading banner in session-ui.js) — never a guarantee, and never anything a
* server-side check relies on.
*/
modelSizesGB?: Record<string, number>;
}
export function customModelHostsPath(configDir: string): string {
+197 -5
View File
@@ -10,7 +10,8 @@
* cli-registry changes" requirement it was written against.
*/
import { chmodSync, mkdirSync, writeFileSync, rmSync } from 'node:fs';
import { chmodSync, existsSync, mkdirSync, readFileSync, writeFileSync, rmSync, symlinkSync } from 'node:fs';
import { homedir, platform } from 'node:os';
import { join, dirname } from 'node:path';
import { dataPath } from './config/instance.js';
import type { CliEntry } from './config/cli-registry/types.js';
@@ -48,6 +49,168 @@ export function applyConfigDirInjection(baseDir: string, injection: ConfigDirInj
return { [injection.dirEnvVar]: baseDir, ...injection.extraEnv };
}
/**
* Real, shared Claude config directory Codeman's own host process runs under — honors
* `CLAUDE_CONFIG_DIR` the same way `claude-credentials.ts`'s `claudeCredentialsPath()`
* does, so the symlink below points at wherever `~/.claude/projects` actually lives
* rather than assuming the plain default.
*/
function realClaudeConfigDir(): string {
const configured = typeof process.env.CLAUDE_CONFIG_DIR === 'string' && process.env.CLAUDE_CONFIG_DIR.trim();
return configured || join(homedir(), '.claude');
}
/**
* Symlinks `<isolatedDir>/projects` back to the real, shared `~/.claude/projects`, so an
* isolated `CLAUDE_CONFIG_DIR` (used to keep an injected API key away from a stored OAuth
* session — see `configDirVar` on customModelInjection) doesn't also blind the response
* viewer, subagent windows, and Read My Mind for that session (docs/wiki/Agent-CLIs.md).
* Best-effort: a platform that refuses symlinks (unprivileged Windows without a junction
* fallback working, e.g.) just keeps the pre-existing documented side effect instead of
* failing the whole custom-model apply over a nice-to-have.
*/
function linkSharedProjectsDir(isolatedDir: string): void {
const link = join(isolatedDir, 'projects');
if (existsSync(link)) return; // already linked (idempotent re-apply) or real dir wrote one
try {
symlinkSync(join(realClaudeConfigDir(), 'projects'), link, platform() === 'win32' ? 'junction' : 'dir');
} catch {
// best-effort only — response viewer/subagent windows go blind for this session instead
}
}
/**
* Pre-approves the injected API key in an isolated config directory's trust-dialog state
* (`customModelInjection.apiKeyTrustFile`), so an otherwise-empty directory doesn't make the
* CLI stop at an interactive "Detected a custom API key — use it?" prompt on every single
* launch. Confirmed live: with nobody at the TTY to answer, that prompt's own default
* ("No") silently refuses the very key this feature just injected — this isn't bypassing
* the check, it's answering it the same field a real answered prompt itself writes to
* (verified against a real `~/.claude.json` after answering by hand once).
*
* Merges rather than overwrites: the file may already carry fields the CLI itself wrote on
* an earlier launch in this same isolated directory (machineID, userID, other approved
* keys), and a corrupt or partially-written file (a crash mid-write) is treated as absent
* rather than failing the whole apply over a nice-to-have.
*/
/**
* The form Claude Code actually stores an approved key in: the trimmed last 20
* characters. Mirrors the CLI's own `e.trim().slice(-20)`, which is applied on BOTH
* the write and the lookup, so anything else never matches.
*/
export function truncateApiKeyForTrustFile(apiKey: string): string {
return apiKey.trim().slice(-20);
}
function seedApiKeyTrustFile(
configDir: string,
trustFile: { relPath: string; shape: 'claude-api-key-responses' },
apiKey: string
): void {
const filePath = join(configDir, trustFile.relPath);
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(readFileSync(filePath, 'utf8')) as Record<string, unknown>;
} catch {
existing = {};
}
const responses = (existing.customApiKeyResponses ?? {}) as { approved?: unknown; rejected?: unknown };
const approved = new Set(Array.isArray(responses.approved) ? (responses.approved as string[]) : []);
// ⚠ Claude Code stores and compares only the LAST 20 CHARACTERS of a key, never the
// whole thing: its lookup is `approved.includes(key.trim().slice(-20))` (decompiled
// from the 2.1.278 bundle, and corroborated by real `~/.claude.json` files, whose
// customApiKeyResponses entries are all exactly 20 characters). Seeding the full key
// therefore never matches for a REAL key, and claude stops at the interactive
// "Detected a custom API key in your environment" prompt, whose default is
// "No (recommended)" — so the launch hangs or silently refuses the key this feature
// just injected. It went unnoticed because a keyless llama.cpp/llama-swap endpoint
// uses DEFAULT_API_KEY ('local-dummy-key', 15 chars), where slice(-20) is the whole
// string and the seed matches by accident. Truncating here also keeps a full
// third-party credential from being written into a second file on disk.
approved.add(truncateApiKeyForTrustFile(apiKey));
const rejected = Array.isArray(responses.rejected) ? responses.rejected : [];
existing.customApiKeyResponses = { approved: [...approved], rejected };
try {
writeFileSync(filePath, JSON.stringify(existing, null, 2), { encoding: 'utf8', mode: 0o600 });
chmodSync(filePath, 0o600);
} catch {
// best-effort only — the interactive prompt returns instead of a hard failure here
}
}
/**
* Pre-seeds the two remaining pieces of "already been onboarded" state a fresh
* `CLAUDE_CONFIG_DIR` has none of (`customModelInjection.skipFirstRunPrompts`, alongside
* apiKeyTrustFile): claude replays its whole first-run sequence — the theme picker, the
* security-notes screen, and (per-project) the "trust this folder?" dialog — against ANY
* config directory that has never completed it, confirmed live against a genuinely fresh
* isolated directory. `hasCompletedOnboarding` skips the theme/security-notes screens
* outright; `projects[workingDir].hasTrustDialogAccepted` answers the trust dialog for
* THIS session's own working directory the same way a real profile's own prior approval
* would — other projects in the file are left alone, and `workingDir` is used verbatim
* (never realpath'd or slash-normalized) since that's the literal string claude itself
* uses as the project key, being whatever string the session was actually launched with
* as its cwd.
*
* Same merge-not-overwrite and corrupt-file-tolerant behavior as `seedApiKeyTrustFile`
* (same file, so a second sequential read-modify-write here is deliberate rather than
* folding both into one pass — keeps each seed independently testable and optional).
*/
function seedFirstRunOnboardingState(
configDir: string,
trustFile: { relPath: string; shape: 'claude-api-key-responses' },
workingDir: string
): void {
const filePath = join(configDir, trustFile.relPath);
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(readFileSync(filePath, 'utf8')) as Record<string, unknown>;
} catch {
existing = {};
}
existing.hasCompletedOnboarding = true;
const projects =
existing.projects && typeof existing.projects === 'object' && !Array.isArray(existing.projects)
? (existing.projects as Record<string, Record<string, unknown>>)
: {};
const existingProject = projects[workingDir] && typeof projects[workingDir] === 'object' ? projects[workingDir] : {};
projects[workingDir] = { ...existingProject, hasTrustDialogAccepted: true };
existing.projects = projects;
try {
writeFileSync(filePath, JSON.stringify(existing, null, 2), { encoding: 'utf8', mode: 0o600 });
chmodSync(filePath, 0o600);
} catch {
// best-effort only — the interactive dialogs return instead of a hard failure here
}
}
/**
* Pre-seeds the "skip the bypass-permissions warning" setting (`customModelInjection.
* skipFirstRunPrompts`, alongside apiKeyTrustFile) into an isolated config directory's
* `settings.json` — a real, already-onboarded profile answers claude's one-time warning
* about running with a bypass-permissions flag once and never sees it again, but every
* custom-model session launches with a fresh, otherwise-empty CLAUDE_CONFIG_DIR that
* carries none of that (confirmed live). A different file from apiKeyTrustFile's
* `.claude.json` — this is claude's own global `settings.json`, not project-keyed —
* so it gets its own merge-not-overwrite read-modify-write.
*/
function seedSkipBypassPermissionsPrompt(configDir: string): void {
const filePath = join(configDir, 'settings.json');
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(readFileSync(filePath, 'utf8')) as Record<string, unknown>;
} catch {
existing = {};
}
existing.skipDangerousModePermissionPrompt = true;
try {
writeFileSync(filePath, JSON.stringify(existing, null, 2), { encoding: 'utf8', mode: 0o600 });
chmodSync(filePath, 0o600);
} catch {
// best-effort only — the interactive warning returns instead of a hard failure here
}
}
/** Best-effort recursive removal of a previously-written configDir. Never throws. */
export function removeConfigDir(dir: string | undefined): void {
if (!dir) return;
@@ -79,14 +242,43 @@ export function applyCustomModelInjection(
entry: Pick<CliEntry, 'capabilities'>,
endpoint: CustomModelEndpoint,
modelId: string,
sessionId: string
sessionId: string,
/** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */
contextLength?: number,
/**
* The session's own working directory — only used for `skipFirstRunPrompts`'s per-project
* trust-dialog seed, and only when provided (boot recovery, which has no reason to
* re-answer a dialog that already fired once, omits it rather than re-deriving it).
*/
workingDir?: string
): AppliedCustomModel | undefined {
const injection = buildCustomModelInjection(entry, endpoint, modelId);
const injection = buildCustomModelInjection(entry, endpoint, modelId, contextLength);
if (injection.kind === 'unsupported') return undefined;
if (injection.kind === 'env') {
// `configDirVar` (claude's CLAUDE_CONFIG_DIR): point it at the same isolated,
// per-session directory the `configDir` kind uses, but write no files into it — an
// empty directory has no stored OAuth credential to conflict with the injected API
// key, which is the whole point. Reusing the same path keyed by sessionId keeps this
// idempotent across a boot-recovery re-apply, same as the configDir kind below.
let envOverrides = injection.envOverrides;
let configDir: string | undefined;
if (injection.configDirVar) {
configDir = customModelConfigDir(sessionId);
mkdirSync(configDir, { recursive: true, mode: 0o700 });
linkSharedProjectsDir(configDir);
if (injection.apiKeyTrustFile && injection.apiKey) {
seedApiKeyTrustFile(configDir, injection.apiKeyTrustFile, injection.apiKey);
}
if (injection.skipFirstRunPrompts && injection.apiKeyTrustFile) {
if (workingDir) seedFirstRunOnboardingState(configDir, injection.apiKeyTrustFile, workingDir);
seedSkipBypassPermissionsPrompt(configDir);
}
envOverrides = { ...envOverrides, [injection.configDirVar]: configDir };
}
return {
envOverrides: injection.envOverrides,
envKeys: Object.keys(injection.envOverrides),
envOverrides,
envKeys: Object.keys(envOverrides),
configDir,
launchModel: injection.launchModel,
};
}
+48 -16
View File
@@ -16,14 +16,21 @@
* shape was rejected by a real codex binary with "invalid type: map,
* expected a string" — caught by `scripts/test-local-llm-harnesses.ts`),
* but `wire_api = "responses"` is the only value codex still accepts
* (support for `"chat"` was dropped in Feb 2026), and a plain OpenAI
* Chat-Completions server (llama.cpp, llama-swap, most local setups) does
* NOT implement the Responses API — so codex may still fail at the
* PROTOCOL level even with a correctly-shaped config file. That gap is
* real and current, not a stale warning; see docs/custom-model-endpoints-plan.md. The rest
* (gemini/pi/grok/deepseek/omp) have their ONE-SHOT INVOCATION flags
* confirmed against real installed binaries' own `--help` output, but
* their custom-endpoint env/config conventions remain web-researched,
* (support for `"chat"` was dropped in Feb 2026). ⚠️ Re-verified live
* against a llama-swap deployment that DOES answer `/v1/responses`: a
* plain, no-tool-call turn gets a real reply, but a real tool-call attempt
* comes back as `agent_message` TEXT (the tool-call JSON printed as the
* answer) rather than a `function_call` item codex would execute —
* confirmed via `codex exec --json`'s raw event stream. Tool execution is
* what makes codex a coding agent, so this remains not usable for real
* work even where plain chat succeeds; see docs/custom-model-endpoints-plan.md
* for the full picture (including the harmless `Model metadata ... not
* found` warning every custom-endpoint codex session prints — sourced from
* a local cache of OpenAI's OWN hosted model catalog that a custom model
* can never appear in, confirmed to have no effect on the outcome above).
* The rest (gemini/pi/grok/deepseek/omp) have their ONE-SHOT INVOCATION
* flags confirmed against real installed binaries' own `--help` output,
* but their custom-endpoint env/config conventions remain web-researched,
* unverified.
*/
@@ -44,6 +51,20 @@ export interface EnvInjection {
envOverrides: Record<string, string>;
/** See {@link ConfigDirInjection.launchModel}. */
launchModel?: string;
/**
* Name of the env var the caller should point at an isolated, credential-free config
* directory for this session (claude's `CLAUDE_CONFIG_DIR`), from the registry entry's
* `customModelInjection.configDirVar`. The actual directory value isn't computed here —
* this module is pure and has no sessionId to derive one from — the IO wrapper
* (`custom-model-injection-apply.ts`) creates it and adds it to `envOverrides`.
*/
configDirVar?: string;
/** See `customModelInjection.apiKeyTrustFile` — carried through so the IO wrapper can seed it. */
apiKeyTrustFile?: { relPath: string; shape: 'claude-api-key-responses' };
/** The literal API key value this injection used, for `apiKeyTrustFile` to pre-approve. */
apiKey?: string;
/** See `customModelInjection.skipFirstRunPrompts` — carried through so the IO wrapper can seed it. */
skipFirstRunPrompts?: boolean;
}
export interface ConfigDirInjection {
@@ -90,7 +111,9 @@ function quoted(value: string): string {
export function buildCustomModelInjection(
entry: Pick<CliEntry, 'capabilities'>,
endpoint: CustomModelEndpoint,
modelId: string
modelId: string,
/** Discovered context-window size for `modelId`, if known — see `contextLengthVar`. */
contextLength?: number
): CustomModelInjectionResult {
const cap = entry.capabilities.customModelInjection;
const apiKey = endpoint.apiKey?.trim() || DEFAULT_API_KEY;
@@ -98,11 +121,18 @@ export function buildCustomModelInjection(
switch (cap.kind) {
case 'env': {
const envOverrides: Record<string, string> = {
[cap.baseUrlVar]: endpoint.baseUrl,
[cap.baseUrlVar]: cap.appendV1Suffix ? withV1Suffix(endpoint.baseUrl) : endpoint.baseUrl,
[cap.apiKeyVar]: apiKey,
};
for (const modelVar of cap.modelVars) envOverrides[modelVar] = modelId;
return withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId);
if (cap.contextLengthVar && contextLength !== undefined && Number.isFinite(contextLength)) {
envOverrides[cap.contextLengthVar] = String(Math.trunc(contextLength));
}
let result: EnvInjection = withLaunchModel({ kind: 'env', envOverrides }, cap.launchModel, modelId);
if (cap.configDirVar) result = { ...result, configDirVar: cap.configDirVar };
if (cap.apiKeyTrustFile) result = { ...result, apiKeyTrustFile: cap.apiKeyTrustFile, apiKey };
if (cap.skipFirstRunPrompts) result = { ...result, skipFirstRunPrompts: true };
return result;
}
case 'configContentEnv': {
@@ -177,11 +207,13 @@ function renderConfigFile(
// `env_key`, the NAME of an env var it reads the credential from at runtime, so the
// actual value must ride along as an extra env var, never embedded in the file.
// ⚠️ `wire_api = "responses"` is the only value codex still accepts (it dropped
// `"chat"` support in Feb 2026) — a plain OpenAI Chat-Completions server (llama.cpp,
// llama-swap, most local setups) does NOT implement the Responses API, so this
// recipe may still fail at the PROTOCOL level even though the file now parses
// correctly. That is a real, currently-unresolved compatibility gap, not a syntax
// bug — track it before calling codex support done.
// `"chat"` support in Feb 2026). Even against a llama-swap deployment that DOES
// answer `/v1/responses`, a real tool-call attempt came back as plain TEXT (the
// tool-call JSON printed as the model's answer) rather than an executable
// `function_call` item — confirmed live via `codex exec --json`. Tool execution is
// what makes codex a coding agent, so this remains not usable for real work even
// where plain chat succeeds — see the confidence table in
// docs/custom-model-endpoints-plan.md, not a syntax bug in this file.
const content = [
`model = ${quoted(modelId)}`,
`model_provider = "custom"`,
+76 -5
View File
@@ -606,9 +606,36 @@ export function agentImageNpmPackages(): string[] {
return packages;
}
/** The `--build-arg` pairs the agent image takes. */
export function agentImageBuildArgPairs(): Array<[string, string]> {
return [['CLI_NPM_PACKAGES', agentImageNpmPackages().join(' ')]];
/**
* Environment variable → agent.Dockerfile ARG for the optional git-host CLIs (gh, az).
* ⚠️ Mirrors `GIT_HOST_CLI_BUILD_ARGS` in `scripts/lib/cli-catalog.mjs`; the parity test pins them.
*/
export const GIT_HOST_CLI_BUILD_ARGS: ReadonlyArray<readonly [string, string]> = [
['CODEMAN_AGENT_IMAGE_INSTALL_GH', 'CODEMAN_INSTALL_GH'],
['CODEMAN_AGENT_IMAGE_INSTALL_AZ', 'CODEMAN_INSTALL_AZ'],
];
/**
* The `--build-arg` pairs for the optional git-host CLIs. PURE. An unset or empty variable
* contributes NOTHING, so the Dockerfile's own default (off) applies and the argv is the same
* as before these existed; anything other than 0/1 is refused rather than guessed at.
*/
export function gitHostCliBuildArgPairs(env: NodeJS.ProcessEnv): Array<[string, string]> {
const pairs: Array<[string, string]> = [];
for (const [envName, argName] of GIT_HOST_CLI_BUILD_ARGS) {
const value = env[envName];
if (value === undefined || value === '') continue;
if (value !== '0' && value !== '1') {
throw new Error(`${envName} must be 0 or 1, got ${JSON.stringify(value)}`);
}
pairs.push([argName, value]);
}
return pairs;
}
/** The `--build-arg` pairs the agent image takes. PURE given `env`. */
export function agentImageBuildArgPairs(env: NodeJS.ProcessEnv = process.env): Array<[string, string]> {
return [['CLI_NPM_PACKAGES', agentImageNpmPackages().join(' ')], ...gitHostCliBuildArgPairs(env)];
}
// ========== Credential mount resolution (IO) ==========
@@ -785,6 +812,14 @@ interface CredStorePolicy {
seedFiles?: string[];
/** Seed the WHOLE dir (RO mount → cp -a) — for stores with no shared/host-read state. */
seedWhole?: boolean;
/**
* Seed this store ONLY when this environment variable is exactly `1`, read when
* the container is created. For credentials that belong to an opt-in tool rather
* than to an agent CLI every case already trusts: they are not inert just because
* the image lacks the tool (a gh `hosts.yml` token or an Azure refresh token is
* usable by anything in the container, and the agent in it is prompt-injectable).
*/
enabledByEnv?: string;
}
const CRED_STORES: CredStorePolicy[] = [
@@ -829,6 +864,31 @@ const CRED_STORES: CredStorePolicy[] = [
},
{ rel: '.config/gcloud', seedWhole: true },
{ rel: '.config/opencode', seedWhole: true },
// GitHub CLI: `hosts.yml` holds the token wherever no system keyring exists (the
// Docker server image, a headless Linux host), `config.yml` the preferences. An
// agent image built with CODEMAN_INSTALL_GH=1 routes github.com git credentials
// through `gh`, so this seed is what lets an agent clone/push a private repo. A
// token that lives in a desktop keyring is not in `hosts.yml` and does not carry
// in; sign `gh` in inside the container. OPT-IN: seeded only when the same switch
// that builds gh into the agent image is on, never merely because the file exists.
{ rel: '.config/gh', seedFiles: ['hosts.yml', 'config.yml'], enabledByEnv: 'CODEMAN_AGENT_IMAGE_INSTALL_GH' },
// Azure CLI: only the sign-in state. `~/.azure` also accumulates `logs/`,
// `commands/`, telemetry and (on a bare host) `cliextensions/`, none of which is
// needed to authenticate; the agent image carries its own extensions outside HOME.
// `msal_token_cache.json` is plaintext only on Linux (Windows/macOS encrypt it), so
// this carries a sign-in from the Docker server image or a Linux host.
// OPT-IN like gh: the MSAL cache holds refresh tokens for the whole Azure account.
{
rel: '.azure',
enabledByEnv: 'CODEMAN_AGENT_IMAGE_INSTALL_AZ',
seedFiles: [
'azureProfile.json',
'msal_token_cache.json',
'service_principal_entries.json',
'clouds.config',
'config',
],
},
// OMP keeps its config in `~/.omp/agent` (config.yml/mcp.json/models.yml/
// settings.yml — small, no bigger than grok's config.toml/pager.toml), but
// that dir ALSO holds agent.db/history.db/models.db (SQLite caches) and
@@ -853,10 +913,14 @@ const CRED_STORES: CredStorePolicy[] = [
* session state back into the host). Every path is existsSync-gated (on most hosts
* only a subset exists). Pure-ish IO (no writes; just existence checks + mount specs).
*/
export function resolveDockerCredentialArtifacts(home: string = homedir()): DockerClaudeArtifacts {
export function resolveDockerCredentialArtifacts(
home: string = homedir(),
env: NodeJS.ProcessEnv = process.env
): DockerClaudeArtifacts {
const mounts: DockerMount[] = [];
const seedCopies: DockerSeedCopy[] = [];
for (const store of CRED_STORES) {
if (store.enabledByEnv && env[store.enabledByEnv] !== '1') continue;
const hostBase = join(home, store.rel);
if (!existsSync(hostBase)) continue;
const containerBase = `${CONTAINER_HOME}/${store.rel}`;
@@ -1136,10 +1200,17 @@ function buildAgentImage(
error: `docker/agent.Dockerfile not found in this install; clone the repo or build ${image} manually`,
});
}
let buildArgPairs: Array<[string, string]>;
try {
buildArgPairs = agentImageBuildArgPairs();
} catch (err) {
// A malformed CODEMAN_AGENT_IMAGE_INSTALL_* value: report it like any other build failure.
return Promise.resolve({ ok: false, built: false, alreadyPresent: false, error: String((err as Error).message) });
}
const argv = dockerEngineArgv(docker);
const args = [
...argv.slice(1),
...agentImageBuildArgs(resolved.dockerfile, image, resolved.contextDir, opts.noCache, agentImageBuildArgPairs()),
...agentImageBuildArgs(resolved.dockerfile, image, resolved.contextDir, opts.noCache, buildArgPairs),
];
return new Promise<EnsureImageResult>((resolve) => {
// async spawn (NEVER spawnSync) so a multi-minute build never wedges the event loop.
+30 -6
View File
@@ -195,6 +195,8 @@ export interface CloneOptions {
/** `--depth 1`: history-less but much faster on large repos. */
shallow?: boolean;
timeoutMs?: number;
/** Clear every git credential helper for this run (see `GIT_NO_CREDENTIAL_HELPERS`). */
withoutCredentialHelpers?: boolean;
}
export type CloneResult = { ok: true; stderr: string } | { ok: false; failure: GitFailure };
@@ -436,12 +438,27 @@ export function isSafeGitRef(ref: string): boolean {
// ─── Pure: argv + env ────────────────────────────────────────────────────────
/**
* Global git options that empty the credential-helper list for one run.
*
* Every Codeman user in multi-user mode runs git as the SAME OS account, so a
* helper that account has (the Docker image's opt-in `gh`/`az` helpers, or a
* user's own `gh auth setup-git`) would read private repositories on the
* signed-in admin's behalf for anyone who can reach Clone Repo. An empty
* `credential.helper` resets the helper list, and a command-line `-c` is read
* last, so it also drops the URL-scoped `credential.<url>.helper` entries the
* image configures (verified against a real private repo: refs with the helper,
* `could not read Username` with it cleared). Public repositories are
* unaffected. It must precede the subcommand.
*/
export const GIT_NO_CREDENTIAL_HELPERS: readonly string[] = ['-c', 'credential.helper='];
/**
* argv for the clone. `--` separates flags from operands so neither the
* repository nor the destination can ever be read as an option.
*/
export function buildCloneArgs(opts: CloneOptions): string[] {
const args = ['clone'];
const args = [...(opts.withoutCredentialHelpers ? GIT_NO_CREDENTIAL_HELPERS : []), 'clone'];
// `--single-branch` is what makes "just this tag/branch" cheap on a big repo.
if (opts.ref) args.push('--single-branch', '--branch', opts.ref);
if (opts.shallow) args.push('--depth', '1');
@@ -450,8 +467,14 @@ export function buildCloneArgs(opts: CloneOptions): string[] {
}
/** argv for the preflight. `--symref` is what reveals the remote's default branch. */
export function buildLsRemoteArgs(repository: string): string[] {
return ['ls-remote', '--symref', '--', repository];
export function buildLsRemoteArgs(repository: string, opts: { withoutCredentialHelpers?: boolean } = {}): string[] {
return [
...(opts.withoutCredentialHelpers ? GIT_NO_CREDENTIAL_HELPERS : []),
'ls-remote',
'--symref',
'--',
repository,
];
}
/**
@@ -579,7 +602,7 @@ export function classifyGitFailure(stderr: string, timedOut: boolean, spawnError
return {
code: 'AUTH_REQUIRED',
message:
'That repository needs authentication. Codeman clones without credentials, so private repositories have to be cloned outside Codeman and added with Link Existing.',
"That repository needs authentication. Codeman never asks for credentials, so sign this server's git in first (for example `gh auth login` or `az login` from a shell session; the Docker image can include both, see docker/README.md), or clone it outside Codeman and add it with Link Existing.",
stderr: clean,
};
}
@@ -790,7 +813,8 @@ export function isGitAvailable(): boolean {
*/
export async function probeGitRemote(
repository: string,
timeoutMs = GIT_LS_REMOTE_TIMEOUT_MS
timeoutMs = GIT_LS_REMOTE_TIMEOUT_MS,
opts: { withoutCredentialHelpers?: boolean } = {}
): Promise<GitRemoteProbe> {
if (!isGitAvailable()) {
return {
@@ -800,7 +824,7 @@ export async function probeGitRemote(
failure: classifyGitFailure('', false, 'ENOENT: git not found'),
};
}
const run = await runGit(buildLsRemoteArgs(repository), timeoutMs, MAX_LS_REMOTE_BYTES);
const run = await runGit(buildLsRemoteArgs(repository, opts), timeoutMs, MAX_LS_REMOTE_BYTES);
if (run.code !== 0 || run.spawnError) {
return {
reachable: false,
+69 -1
View File
@@ -24,6 +24,7 @@ import type {
OmpConfig,
SessionRemote,
SessionDocker,
PaneExit,
} from './types.js';
/**
@@ -56,6 +57,23 @@ export interface MuxSession {
respawnConfig?: PersistedRespawnConfig;
/** Whether Ralph / Todo tracking is enabled */
ralphEnabled?: boolean;
/**
* This record was rebuilt from the tmux socket rather than from Codeman's own
* bookkeeping, so everything on it but the name and the pid is a guess. Its
* synthetic `restored-<fragment>` id cannot find the session's `state.json`
* entry either, which means a remote or docker session rediscovered this way
* arrives with no `remote`/`docker` metadata and looks local. Anything that
* would be WRONG about such a session rather than merely vague must fail
* closed on this flag.
*
* ⚠ It is PERMANENT, not merely true for the boot that rediscovered the
* session: `saveSessions()` serializes the whole record to
* `mux-sessions.json` and `loadSessions()` restores it, so a genuinely local
* session rediscovered once stays opted out of everything keyed on this for
* the life of that record. That is the safe direction to fail, and it costs
* only the guess Codeman is declining to make.
*/
discovered?: boolean;
}
/**
@@ -72,6 +90,13 @@ export interface CreateSessionOptions {
workingDir: string;
mode: SessionMode;
name?: string;
/**
* Name pinned on a claude spawn as `--name` (version-gated, sanitized, local only).
* Deliberately NOT `name`: `--name` owns the prompt-box label, the `/resume` picker
* entry and the terminal title, and a pinned title stops Claude generating its own,
* so only a user-chosen name belongs here (see `Session.cliPinnedName`).
*/
cliName?: string;
niceConfig?: NiceConfig;
model?: string;
claudeMode?: ClaudeMode;
@@ -105,8 +130,15 @@ export interface RespawnPaneOptions {
sessionId: string;
workingDir: string;
mode: SessionMode;
/** Session display name; a respawned claude keeps its `--name` peer name (version-gated, local only). */
/** Session display name (tab name). */
name?: string;
/**
* Name pinned on a respawned claude as `--name` (version-gated, sanitized, local only).
* Deliberately NOT `name`: `--name` owns the prompt-box label, the `/resume` picker
* entry and the terminal title, and a pinned title stops Claude generating its own,
* so only a user-chosen name belongs here (see `Session.cliPinnedName`).
*/
cliName?: string;
niceConfig?: NiceConfig;
model?: string;
claudeMode?: ClaudeMode;
@@ -159,6 +191,15 @@ export interface PaneCaptureOptions {
* the 1MB execSync default (ENOBUFS).
*/
maxCaptureBytes?: number;
/**
* Filled in by the implementation with the pane geometry the capture was
* really taken at, which is not always the geometry the caller last asked
* for: a resize and a capture can race, and a pane whose size a desktop
* viewport has claimed ignores a smaller client's resize outright. A
* visible-frame capture addresses every row absolutely, so a consumer
* rendering it needs the real height to know the frame fits.
*/
capturedGeometry?: { cols: number; rows: number };
}
/**
@@ -171,6 +212,7 @@ export interface PaneCaptureOptions {
* - `sessionKilled` (data: { sessionId: string }) - Session terminated
* - `sessionDied` (data: { sessionId: string }) - Session died unexpectedly
* - `statsUpdated` (sessions: MuxSessionWithStats[]) - Stats refreshed
* - `paneExitsUpdated` () - A pane read finished; ask `getPaneExit()` per session
*/
export interface TerminalMultiplexer extends EventEmitter {
/** Which backend this instance uses */
@@ -285,6 +327,32 @@ export interface TerminalMultiplexer extends EventEmitter {
/** Check if the pane in a session is dead (command exited but remain-on-exit keeps it alive) */
isPaneDead(muxName: string): boolean;
/**
* What the last pane read saw of this session's agent, or `undefined` for
* UNKNOWN (Ark0N/Codeman#446). Unlike `isPaneDead()` this costs nothing: it
* reads a map the batched watcher fills, so it answers no fresher than that
* watcher's interval and the three synchronous `isPaneDead()` callers still
* need their own probe. See {@link PaneExit}.
*/
getPaneExit?(muxName: string): PaneExit | undefined;
/**
* How many authoritative pane reads have agreed on the exit `getPaneExit()`
* reports, or 0 when it reports none. The exited-agent sweep closes a session
* only once this reaches `CLEAN_EXIT_CONFIRMING_READS` (`pane-exit-sweep.ts`),
* and a multiplexer without this method never has a session closed by it.
*/
getPaneExitReadCount?(muxName: string): number;
/** Forget a session's exit observation, e.g. once its pane has been respawned. */
clearPaneExit?(muxName: string): void;
/** Start polling every pane on the socket for an exited agent. */
startPaneExitWatcher?(intervalMs?: number): void;
/** Stop the pane-exit watcher. */
stopPaneExitWatcher?(): void;
/** Respawn a dead pane with a fresh command. Returns the new PID or null on failure. */
respawnPane(options: RespawnPaneOptions): Promise<number | null>;
+99
View File
@@ -0,0 +1,99 @@
/**
* @fileoverview The exited-agent sweep's decision rule (Ark0N/Codeman#446).
*
* Codeman creates every tmux pane with `remain-on-exit on`, so `/exit` ends the
* CLI while the pane, the tmux session and the `tmux attach-session` process
* all live on. Part 1 of #446 records that as `SessionState.paneExit`. This
* module decides when such a session is closed, the way the X button closes
* it, so finished sessions stop piling up on the board.
*
* The rule closes a session only on a POSITIVE observation of a clean exit:
*
* - The exit status must be an explicit numeric 0 with no signal. An absent
* status is UNKNOWN, never 0: on tmux 3.2a a SIGKILLed pane reports neither a
* status nor a signal, so reading absence as clean would sweep an agent the
* OOM killer took. A non-zero status or any signal keeps the row, marked with
* the exit, as the crash evidence #210 was filed to keep.
* - At least {@link CLEAN_EXIT_CONFIRMING_READS} authoritative pane reads must
* have agreed on that exit. A failed, empty or skipped read counts for
* nothing, because unknown never closes anything.
* - No start, attach or relaunch may be in flight for the session. The
* dead-pane branch of `Session._setupOrAttachMuxSession()` respawns an exited
* pane on purpose, and for a few seconds that pane still reads as dead.
* - The exit must land at least {@link CLEAN_EXIT_MIN_PANE_LIFETIME_MS} after
* the last start, attach or relaunch finished. A CLI that prints a startup
* error ("not logged in", a bad profile, a config error) and exits 0 would
* otherwise lose its tab, and the error with it, seconds after launch. Its
* row stays, marked `exited (0)`, for the user to read and close.
*
* Scoping to local mux-backed sessions happens before this rule runs:
* `Session.setPaneExit()` forces the field to UNKNOWN for direct-PTY, remote,
* docker and discovered sessions, so their `paneExit` never reaches here.
*
* Pure, so the rule is unit-tested without a server (test/pane-exit-sweep.test.ts).
*/
import type { PaneExit } from './types/index.js';
/**
* How many authoritative pane reads must agree on a clean exit before the
* session is closed. At the watcher's 2 s cadence two reads mean a finished
* session disappears within about four seconds of its agent exiting.
*/
export const CLEAN_EXIT_CONFIRMING_READS = 2;
/**
* How long a pane must have been up before a clean exit closes its session.
* An exit sooner than this after the last pane start is read as a startup
* failure rather than a user ending the agent, and the row is kept.
*/
export const CLEAN_EXIT_MIN_PANE_LIFETIME_MS = 10_000;
/** The lifecycle-log reason recorded when the sweep closes a session. */
export const CLEAN_EXIT_CLOSE_REASON = 'agent exited cleanly (status 0)';
/**
* Is this exit a clean one? True only for an explicit numeric status of 0 with
* no signal reported.
*
* ⚠ Never widen this to `(exit.status ?? 0) === 0` or to "no signal, so it was
* clean". An absent status is how a signal death presents on tmux 3.2a, and
* that shortcut would close crashed agents with nothing failing to warn you.
*/
export function isCleanPaneExit(exit: PaneExit | undefined): boolean {
if (!exit) return false;
if (exit.signal !== undefined) return false;
return exit.status === 0;
}
/** Everything the sweep needs to know about one session. */
export interface CleanExitSweepCandidate {
/** The session's published exit, already scoped by `Session.setPaneExit()`. */
paneExit: PaneExit | undefined;
/** Authoritative pane reads that agreed on that exit (`getPaneExitReadCount()`). */
confirmingReads: number;
/** A start, attach or relaunch is running for this session's pane. */
paneLifecycleInFlight: boolean;
/** The session is already being closed or detached. */
closing: boolean;
/**
* When the last start, attach or relaunch of this pane finished
* (`Session.paneStartedAt`), or 0 when none has run in this process.
*/
paneStartedAt: number;
}
/** Should the sweep close this session now? See the file overview for the rule. */
export function shouldCloseCleanlyExitedSession(candidate: CleanExitSweepCandidate): boolean {
if (candidate.closing) return false;
if (candidate.paneLifecycleInFlight) return false;
if (!isCleanPaneExit(candidate.paneExit)) return false;
// `at` is when this server first read the pane dead, so an exit during the
// start itself lands BEFORE `paneStartedAt` and is kept too.
if (
candidate.paneStartedAt > 0 &&
candidate.paneExit!.at - candidate.paneStartedAt < CLEAN_EXIT_MIN_PANE_LIFETIME_MS
) {
return false;
}
return candidate.confirmingReads >= CLEAN_EXIT_CONFIRMING_READS;
}
+27 -14
View File
@@ -19,19 +19,22 @@
* touching its status, so a pinned session a reboot killed still reads `idle` or
* `busy` and stays eligible.
*
* ⚠️ Ending the AGENT rather than the session is a shape this module CANNOT
* recognise today, and a reboot restores it. `/exit` ends the CLI inside the
* pane, `remain-on-exit` keeps the pane, and the PTY Codeman owns is the
* `tmux attach-session` process, which stays alive throughout — so no exit
* handler runs, no lifecycle `exit` is logged, and the record keeps both its pid
* and `status: 'idle'`. Nothing durable distinguishes it from a session that was
* simply idle when the power went. Ark0N/Codeman#446 covers making Codeman
* notice the dead pane; until a record can say the agent is gone, this pass will
* offer those sessions back, and the user dismisses or closes them.
* ⚠️ Ending the AGENT rather than the session leaves no trace in `status` or
* `pid`. `/exit` ends the CLI inside the pane, `remain-on-exit` keeps the pane,
* and the PTY Codeman owns is the `tmux attach-session` process, which stays
* alive throughout — so no exit handler runs, no lifecycle `exit` is logged,
* and the record keeps both its pid and `status: 'idle'`. Ark0N/Codeman#446
* handles it in two steps. The pane-exit watcher persists `paneExit`, and the
* clean-exit sweep (`pane-exit-sweep.ts`) closes a session whose agent exited
* with status 0 through `cleanupSession()`, which leaves the durable record
* described above. This module also refuses a record whose persisted
* `paneExit` is a clean exit, which covers a session that exited moments
* before the power went, before the sweep reached it. A crashed agent's record
* stays eligible, like the row the sweep leaves on the board for it.
*
* The `pid` check below is therefore NOT that rule. It refuses a record whose
* attach process was already gone, which is a session that never started or
* whose pane died outright.
* The `pid` check below is NOT that rule. It refuses a record whose attach
* process was already gone, which is a session that never started or whose
* pane died outright.
*
* @dependencies types (SessionState), config/cli-registry
* @consumedby web/server (plan build at boot), web/routes/reboot-restore-routes
@@ -41,6 +44,7 @@
import type { SessionState } from './types.js';
import { getCli } from './config/cli-registry/registry.js';
import { isCleanPaneExit } from './pane-exit-sweep.js';
/** Session statuses a reboot restore may rebuild. `stopped` is the kill marker. */
const RESTORABLE_STATUSES: ReadonlySet<string> = new Set(['idle', 'busy', 'error']);
@@ -109,7 +113,7 @@ export function resolveResumeConversationId(state: SessionState): string {
/**
* Why one session was passed over. Reported for logging and shown to the user.
*
* The first seven are decided before anything is built. `capacity-reached` and
* All but the last two are decided before anything is built. `capacity-reached` and
* `rebuild-failed` can only happen once a click is spending the plan, and they
* are the two the banner must not confuse with a missing workspace: one means
* "try again after closing something", the other means the CLI would not start.
@@ -120,6 +124,7 @@ export interface RebootRestoreRejection {
| 'no-persisted-record'
| 'intentionally-ended'
| 'not-running'
| 'agent-exited'
| 'respawn-blocked'
| 'remote-or-docker'
| 'unsupported-mode'
@@ -191,7 +196,8 @@ export function planRebootRestore(
//
// ⚠️ This does NOT catch a session the user ended with `/exit`. See the
// module header: that leaves the pid in place, because the pid is the tmux
// attach process and `remain-on-exit` keeps it alive.
// attach process and `remain-on-exit` keeps it alive. The `paneExit` check
// below catches it instead.
//
// Conservative on purpose. A session that somehow persisted no pid while
// genuinely running is not offered, and its conversation stays reachable
@@ -200,6 +206,13 @@ export function planRebootRestore(
skipped.push({ sessionId, reason: 'not-running' });
continue;
}
if (isCleanPaneExit(state.paneExit)) {
// The user ended the agent, and the clean-exit sweep would have closed the
// session had the power not gone first (Ark0N/Codeman#446). The same
// explicit-0 rule applies: an absent status is unknown, not clean.
skipped.push({ sessionId, reason: 'agent-exited' });
continue;
}
if (state.respawnBlocked === true) {
// The crash-loop breaker tripped on this pane. Re-creating it restarts the loop.
skipped.push({ sessionId, reason: 'respawn-blocked' });
+34
View File
@@ -540,6 +540,32 @@ export function remoteDisplayPath(
return `${remote.username}@${remote.host}:${path}`;
}
/**
* Refresh HOST-level config on a RESTORED `SessionRemote`.
*
* A session's `remote` block is persisted at launch time (mux-sessions.json /
* state.json) and recovery uses that snapshot, so a field ADDED to the host config
* later never reaches an already-running session — not even across a Codeman
* restart. That is exactly how a `wakeCommand` added to `remote-hosts.json` would
* silently do nothing until the session is relaunched (which for an owned remote
* session means killing the remote tmux).
*
* Deliberately narrow: ONLY `wakeCommand`/`wakeMac` are taken from the host config,
* and the host is authoritative for them (removing one in the config turns that
* wake path off again). The other host-level fields (`commands`, ssh options) stay as
* persisted so this cannot silently change how an existing pane connects.
*/
export function rehydrateRemoteHostFields<T extends { hostId: string; wakeCommand?: string; wakeMac?: string }>(
remote: T | undefined,
hostsById: ReadonlyMap<string, RemoteHost>
): T | undefined {
if (!remote) return remote;
const host = hostsById.get(remote.hostId);
if (!host) return remote;
if (remote.wakeCommand === host.wakeCommand && remote.wakeMac === host.wakeMac) return remote;
return { ...remote, wakeCommand: host.wakeCommand, wakeMac: host.wakeMac };
}
export function toSessionRemote(host: RemoteHost, remoteCase: RemoteCase): SessionRemote {
return {
hostId: host.id,
@@ -549,6 +575,10 @@ export function toSessionRemote(host: RemoteHost, remoteCase: RemoteCase): Sessi
port: host.port,
remotePath: remoteCase.remotePath,
commands: host.commands,
// Wake-on-LAN command/MAC travel with the session so the input route can wake a
// sleeping host without a second config read (see remote-wake.ts).
wakeCommand: host.wakeCommand,
wakeMac: host.wakeMac,
// COD-105 — the COD-104 launch path creates the remote session, so we own it
// (an explicit kill may propagate a remote kill-session). Discovered+attached
// sessions go through `toAttachedSessionRemote` with `owned: false`.
@@ -587,6 +617,10 @@ export function toAttachedSessionRemote(
port: host.port,
remotePath,
commands: host.commands,
// An attached session can be woken exactly the same way — the identity of the
// creator does not change whether the host is asleep.
wakeCommand: host.wakeCommand,
wakeMac: host.wakeMac,
// Discovered + attached — another Codeman created it. Detach-not-kill.
owned: false,
remoteSessionName,
+1035
View File
File diff suppressed because it is too large Load Diff
+67
View File
@@ -21,6 +21,8 @@
* in 12/12 windows and the four idle ones in 0/12.
*/
import { stripAnsi } from './utils/regex-patterns.js';
/**
* A gap longer than this ends a run of continuous output. Claude repaints at
* least once a second while working, so this leaves generous headroom.
@@ -91,3 +93,68 @@ export function isSustainedActivity(streak: ActivityStreak | null, streakMs: num
export function isPaneQuiet(lastActivityAt: number, now: number, silenceMs: number = IDLE_SILENCE_MS): boolean {
return now - lastActivityAt >= silenceMs;
}
/**
* How many rows at the foot of a pane capture may hold the background-work row, for a
* CLI that declares no number of its own (`capabilities.workDetect.watchingLines`).
*
* One, because the tightest window is the right default and Claude Code needs no more:
* it draws its chip on the LAST row of the screen. Blank rows are dropped before the
* window is taken, so a trailing blank costs nothing, and a CLI that ever prints a row
* BELOW its chip loses the badge rather than gaining a hole.
*
* ⚠️ The size of this window is a trust boundary, not a tidiness measure, and the row
* it excludes first is the one that taught us so: Claude's status line sits directly
* above the footer, its content comes from a `statusLine` command, and a session running
* with permissions bypassed can write that command into `.claude/settings.json` in its
* own workspace. A window of two therefore let an agent print `· 1 monitor ·` onto a row
* of its own and silence its own idle alert. Every row added here is another row
* somebody may be able to write, so widen this only for a CLI whose layout forces it,
* and never to a whole-pane search.
*/
export const WATCHING_TAIL_LINES = 1;
/** Longest label a badge will carry. A footer chip is a handful of words. */
export const MAX_WATCHING_LABEL_CHARS = 40;
/**
* What a pane says is still running in the background, e.g. `1 monitor` or `2 shells`.
*
* The CLI writes that chip while a monitor, a backgrounded shell or a cloud session it
* started is still going, which is exactly the case where the agent has ended its turn
* without wanting anything from the user. `pattern` comes from the CLI's own registry
* entry (`capabilities.workDetect.watchingLine`); group 1 is the label when the pattern
* declares one, and the whole match stands in when it does not.
*
* Each candidate row is tested on its own, bottom row first, so a pattern can anchor
* itself with `^` or `$` against a single row rather than against a joined block. Blank
* rows are dropped before the window is taken, because a CLI that leaves a blank line
* between its chrome rows would otherwise spend the window on nothing. The answer is
* stripped of ANSI and capped, because it ends up on a badge and in an approval card.
*
* @param tailLines how many non-blank rows from the bottom to look at, defaulting to
* `WATCHING_TAIL_LINES`; a CLI declares its own when its row is not the last one
* @returns the label, or null when the pane shows no background work
*/
export function watchingLabel(
paneText: string | null | undefined,
pattern: RegExp,
tailLines: number = WATCHING_TAIL_LINES
): string | null {
if (!paneText) return null;
const lines = stripAnsi(paneText)
.split('\n')
.map((line) => line.trimEnd())
.filter((line) => line !== '');
for (const line of lines.slice(-Math.max(1, tailLines)).reverse()) {
// A pattern compiled by compileVersionRegex() never carries the `g` flag, but a
// caller reaching in from a test or a config reload might, and a stale lastIndex
// would make the same screen match every other call.
pattern.lastIndex = 0;
const match = pattern.exec(line);
if (!match) continue;
const label = (match[1] ?? match[0]).trim().slice(0, MAX_WATCHING_LABEL_CHARS);
if (label) return label;
}
return null;
}
+12 -7
View File
@@ -6,13 +6,18 @@
* stripped before the session is built. The create and resume routes are what
* this bites on: they clamp what a request asked for.
*
* The reboot-restore route calls it as defence in depth, and today it can strip
* nothing. `Session.getEnvOverridesForPersist()` keeps only `CLAUDE_CODE_*` and
* `CLAUDE_CONFIG_DIR` out of a session's overrides, claude's `privilegedEnvKeys`
* are the five `ANTHROPIC_*` names, and that pass admits claude alone — so a
* persisted record cannot carry a clamped key. The call is there for the day the
* persisted set widens. The grant re-resolution that does bite on that path is
* `resolveClaudeModeForUsername`, which recomputes the permission mode.
* The reboot-restore route calls it as defence in depth, and it CAN strip
* something today: `Session.getEnvOverridesForPersist()` keeps only
* `CLAUDE_CODE_*` and `CLAUDE_CONFIG_DIR` out of a session's overrides, and
* claude's `privilegedEnvKeys` now includes both `CLAUDE_CODE_MAX_CONTEXT_TOKENS`
* and `CLAUDE_CONFIG_DIR` (Custom Model Endpoint Profiles, since both can
* redirect a claude session's traffic — see stock.ts's own comment on why they
* are listed despite not needing the clamp for that feature). So a non-granted
* owner's persisted `CLAUDE_CONFIG_DIR` (the per-client-account override, #255)
* is now stripped on reboot-restore, silently returning that session to the
* default Claude account rather than the account it was pointed at. The grant
* re-resolution that ALSO bites on that path is `resolveClaudeModeForUsername`,
* which recomputes the permission mode.
*
* This lives outside `web/routes` on purpose. The question it answers is about
* session privilege rather than about HTTP, and `cron/cron-service.ts` sets the
+520 -24
View File
@@ -60,8 +60,11 @@ import {
type SessionDocker,
type SessionNameSource,
type SessionWriteOptions,
type PaneExit,
} from './types.js';
import { resolveAndClaimOmpSessionId } from './utils/omp-session-resolver.js';
import { claudeTranscriptExists } from './utils/claude-transcript.js';
import { matchesPattern } from './config/cli-registry/patterns.js';
import { probeDockerCliVersion } from './docker-hosts.js';
import { probeRemoteCliVersion } from './remote-hosts.js';
import type { TerminalMultiplexer, MuxSession } from './mux-interface.js';
@@ -81,6 +84,8 @@ import {
trackActivityStreak,
isSustainedActivity,
isPaneQuiet,
watchingLabel,
WATCHING_TAIL_LINES,
IDLE_RECHECK_MS,
PANE_PROBE_MIN_INTERVAL_MS,
PANE_PROBE_RECHECK_MS,
@@ -497,8 +502,35 @@ export class Session extends EventEmitter {
private _activityStreak: ActivityStreak | null = null; // Unbroken run of PTY repaints (working detection)
private _lastPaneProbeAt = 0; // Throttle for the tmux screen probe
private _lastPaneProbeWorking: boolean | null = null; // Its last verdict (null = could not read)
/**
* Background work the pane's own footer reports, e.g. `1 monitor`; null for none.
*
* Cached BESIDE `_lastPaneProbeWorking` and refreshed only by a capture that really
* happened, so it goes stale exactly as that verdict does. The probe returns its
* cached boolean without re-capturing inside `PANE_PROBE_MIN_INTERVAL_MS`, and a
* label derived from a capture nobody took would be a guess wearing a fact's clothes.
*
* ⚠️ It then FREEZES once `_confirmIdle()` concludes: `activityTimeout` is null from
* there, and nothing looks at the pane again until it produces output. That is
* correct rather than merely tolerable, because work ending repaints the pane either
* way — a monitor firing wakes the agent, and codex drops its background-terminal row
* on its own. Do not add a timer to keep this fresh; it would spend a `capture-pane`
* per idle session per tick to learn nothing.
*
* A server restart is not a hole in that either, though it looks like one: this field
* is live state and starts empty. Reconciliation re-attaches the pane, the attach
* repaint carries the composer glyph, and the idle confirmation that arms on it probes
* and re-reads the label with no input from anyone — measured 2026-09-23 on a restarted
* instance, back within ~20 s for a session whose background terminal was still
* running. A session that comes back with no label has no chip on its screen.
*/
private _watching: string | null = null;
/** Lazily compiled `capabilities.workDetect.workingLine`. See _workingLinePattern(). */
private _workingLineRe: RegExp | undefined = undefined;
/** Lazily compiled `capabilities.workDetect.watchingLine`. See _watchingLinePattern(). */
private _watchingLineRe: RegExp | null | undefined = undefined;
/** Resolved with the pattern above: how many rows at the foot of the screen to search. */
private _watchingWindow = WATCHING_TAIL_LINES;
private _trustDialogAccepted: boolean = false; // Stops the trust-dialog scan (answered, or given up)
private _trustDialogAttempts = 0; // Keystrokes sent at the trust dialog
private _lastTrustDialogScanAt = 0; // Throttle for the trust-dialog screen read
@@ -542,6 +574,35 @@ export class Session extends EventEmitter {
private _mux: TerminalMultiplexer | null = null;
private _muxSession: MuxSession | null = null;
private _useMux: boolean = false;
/**
* The agent in this session's local tmux pane has exited (Ark0N/Codeman#446).
* `null` is the UNKNOWN arm of the tri-state and is what {@link setPaneExit}
* stores for every session shape the field does not apply to. See
* {@link PaneExit} for the shapes and for why an unknown answer must never be
* rendered as "alive".
*/
private _paneExit: PaneExit | null = null;
/**
* How many starts, attaches or relaunches are running for this session's
* pane. While one is, a dead-pane reading may describe a pane that is being
* revived on purpose, so the exited-agent sweep leaves the session alone
* (Ark0N/Codeman#446). A counter rather than a flag, so two overlapping
* operations cannot clear each other's mark.
*/
private _paneLifecycleOps = 0;
/** When the last pane start, attach or relaunch finished (ms), 0 when none has run. */
private _paneStartedAt = 0;
/**
* The server has started closing this session, so no start or attach may
* begin (see {@link markClosing}).
*/
private _closing = false;
/**
* This session was rebuilt from the tmux socket rather than from Codeman's
* own records, so its `remote`/`docker` metadata is missing rather than known
* to be absent. See {@link MuxSession.discovered}.
*/
private _discoveredMuxSession = false;
// Flag to prevent new timers after session is stopped
private _isStopped: boolean = false;
@@ -718,6 +779,10 @@ export class Session extends EventEmitter {
lastSubmitAt?: number;
/** Restored conversation chain, oldest first (see `claudeSessionChain`). */
claudeSessionChain?: string[];
/** Restored agent-exit observation for this session's pane (see `paneExit`). */
paneExit?: PaneExit;
/** This session was rebuilt from the tmux socket, so its metadata is a guess. */
discoveredMuxSession?: boolean;
/** Restored wall-clock ms of the pane's last output (recovery only; see `_wireActivityAt`). */
lastActivityAt?: number;
/** Remote execution metadata for sessions launched through SSH inside local tmux. */
@@ -870,6 +935,16 @@ export class Session extends EventEmitter {
this._remote = config.remote;
this._docker = config.docker;
this._owner = config.owner;
this._discoveredMuxSession = config.discoveredMuxSession === true;
// Restored so a record that says the agent exited survives a server restart
// rather than being blanked by the first persist after boot. It runs here
// because the scoping reads `_remote`, `_docker` and the mux fields, all of
// which are set by now. It is a claim about a pane this process has not
// looked at yet, so every path that starts or re-attaches a pane drops it
// (see `_setupOrAttachMuxSession`) and the pane-exit watcher's own tick
// replaces it with a first-hand reading. NOT the stats collector, which a
// browser panel arms and disarms — see `startPaneExitWatcher`.
this.setPaneExit(config.paneExit);
// Never self-parent: a session pointing at itself would draw a zero-length
// lineage arc under its own tab. Only reachable via the recovery path, where
// both the id and the saved parent come from disk.
@@ -1097,6 +1172,118 @@ export class Session extends EventEmitter {
return this._muxSession?.muxName ?? null;
}
/**
* True when a tmux pane's death would mean THIS session's agent has exited.
*
* Four shapes fail the test, and each would otherwise publish a death that is
* not the agent's. A direct-PTY session owns no pane at all. A remote SSH
* session's local pane holds the ssh client, whose death means a transport
* drop OR an exit, which is the ambiguity PR #355 was about. A docker case's
* local pane holds a `docker exec` into the container's own tmux.
*
* The fourth is a session rebuilt from the socket. Absent `remote`/`docker`
* normally means "this is local", but on a discovered record it only means
* "Codeman never found the metadata": the synthetic `restored-<fragment>` id
* matches no `state.json` entry, so a remote session rediscovered after
* `mux-sessions.json` was lost arrives looking local, and its next transport
* drop would be published as an agent exit. Unproven locality fails closed.
*/
private get paneExitApplies(): boolean {
if (this._discoveredMuxSession) return false;
return this._useMux && this._muxSession !== null && !this._remote && !this._docker;
}
/** What Codeman last observed of this pane's agent, or undefined for UNKNOWN. */
get paneExit(): PaneExit | undefined {
return this._paneExit ?? undefined;
}
/**
* True while a start, attach or relaunch is running for this session's pane.
* The exited-agent sweep reads it (see `pane-exit-sweep.ts`): the dead-pane
* branch of {@link _setupOrAttachMuxSession} respawns an exited pane, and
* until it finishes and clears the exit, the pane still reads as dead.
*/
get paneLifecycleInFlight(): boolean {
return this._paneLifecycleOps > 0;
}
/**
* When the last start, attach or relaunch of this pane finished, or 0 when
* none has run in this process. The exited-agent sweep keeps an exit that
* lands within `CLEAN_EXIT_MIN_PANE_LIFETIME_MS` of it, since that reads as a
* CLI failing at startup rather than a user ending it. An attach to a pane
* that was already running stamps it too, which only costs a user who
* `/exit`s within seconds of a server restart a row to close by hand.
*/
get paneStartedAt(): number {
return this._paneStartedAt;
}
/**
* Mark this session as being closed, or clear the mark after a close that
* failed. While it is set, {@link startInteractive} and {@link startShell}
* refuse to run. A start that raced a close would otherwise launch a CLI in a
* tmux session whose record is about to be deleted (Ark0N/Codeman#446).
*/
markClosing(closing: boolean): void {
this._closing = closing;
}
/** Run one pane start, attach or relaunch with {@link paneLifecycleInFlight} raised. */
private async _withPaneLifecycle<T>(op: () => Promise<T>): Promise<T> {
this._paneLifecycleOps++;
try {
return await op();
} finally {
this._paneLifecycleOps--;
this._paneStartedAt = Date.now();
}
}
/**
* Forget this pane's exit, on both this record and the mux layer's cache.
* Every path that starts or relaunches a command in the pane calls it, and
* the mux half also invalidates a pane read already in flight.
*
* It does not persist or broadcast by itself; the caller owns both. ⚠ That
* caller MUST persist, and the pane-exit watcher is not a fallback for it:
* the watcher's next tick reads UNKNOWN, finds this field already cleared,
* reports no change and therefore writes nothing, so a caller that only
* broadcasts leaves `state.json` saying the agent exited for as long as the
* session stays quiet. `/interactive` and `/shell` did exactly that until
* Ark0N/Codeman#446 review; both now persist on their success path.
*/
private clearPaneExitForNewPane(): void {
this.setPaneExit(undefined);
if (this._muxSession) this._mux?.clearPaneExit?.(this._muxSession.muxName);
}
/**
* Record what the mux layer observed of this pane's agent, and say whether
* that changed the answer. The caller persists and broadcasts on a true.
*
* A session the field does not apply to is forced to UNKNOWN here rather than
* at the reporting end, so the rule lives in one place and the mux layer stays
* free to report the raw pane reading its own remote-reconnect watcher needs.
*/
setPaneExit(next: PaneExit | undefined): boolean {
const resolved = this.paneExitApplies ? (next ?? null) : null;
const prev = this._paneExit;
if (prev === resolved) return false;
if (
prev !== null &&
resolved !== null &&
prev.status === resolved.status &&
prev.signal === resolved.signal &&
prev.at === resolved.at
) {
return false;
}
this._paneExit = resolved;
return true;
}
/**
* True when this session's PTY is a tmux client rather than the program itself.
* Read by the replay-side alt-screen strip, which must apply the same
@@ -1118,6 +1305,16 @@ export class Session extends EventEmitter {
return this._isWorking;
}
/**
* What the pane says is still running in the background, e.g. `1 monitor`, or null when
* nothing is. A session with a label here has ended its turn without wanting anything
* from the user, so a surface that would otherwise file it under "needs you" can say
* what it is waiting for instead.
*/
get watching(): string | null {
return this._watching;
}
/**
* Check if the session's process tree has active child processes beyond Claude itself.
* Detects running bash tools, test suites, builds, servers, etc. that Claude spawned.
@@ -1421,6 +1618,19 @@ export class Session extends EventEmitter {
return this._nameSource;
}
/**
* The name to pin on the Claude CLI as `--name`, or undefined to let Claude
* title the conversation itself. `--name` is the prompt-box label, the
* `/resume` picker entry and the terminal title all at once, and a pinned
* title stops Claude generating its own, so only a name the user chose is
* worth pinning. Pinning the `w1-myapp` placeholder gave every conversation
* in a case the same `/resume` entry; an auto name is a cut of the first
* prompt, which Claude's own generated title already beats.
*/
get cliPinnedName(): string | undefined {
return this._nameSource === 'manual' ? this._name : undefined;
}
setAutoClear(enabled: boolean, threshold?: number): void {
this._autoOps.setAutoClear(enabled, threshold);
}
@@ -1620,6 +1830,10 @@ export class Session extends EventEmitter {
// by the constructor: a Codeman restart starts with a fresh breaker so boot
// recovery can re-attach.
respawnBlocked: this._respawnBlocked || undefined,
// Ark0N/Codeman#446 — the agent in this pane has exited, published here so
// it rides the existing `session:updated` broadcast and lands in state.json
// through the same persist. `status` and `pid` above stay untouched by it.
paneExit: this._paneExit ?? undefined,
attachmentHistory: this.attachmentHistory.length > 0 ? this.attachmentHistory : undefined,
lastSubmitAt: this._lastSubmitAt || undefined,
// Only a chain the CLI's own hooks vouched for is persisted, and only when
@@ -1675,6 +1889,7 @@ export class Session extends EventEmitter {
totalCost: this._totalCost,
messageCount: this._messages.length,
isWorking: this._isWorking,
watching: this._watching,
lastPromptTime: this._lastPromptTime,
// Buffer statistics for monitoring long-running sessions
bufferStats: {
@@ -1731,41 +1946,90 @@ export class Session extends EventEmitter {
respawnPaneOptions: import('./mux-interface.js').RespawnPaneOptions;
createSessionOptions: import('./mux-interface.js').CreateSessionOptions;
spawnErrLabel: string;
}): Promise<{ isRestored: boolean }> {
}): Promise<{ isRestored: boolean; respawnedResumeId?: string; respawnedDeadPane: boolean }> {
return this._withPaneLifecycle(() => this._doSetupOrAttachMuxSession(options));
}
private async _doSetupOrAttachMuxSession(options: {
respawnPaneOptions: import('./mux-interface.js').RespawnPaneOptions;
createSessionOptions: import('./mux-interface.js').CreateSessionOptions;
spawnErrLabel: string;
}): Promise<{ isRestored: boolean; respawnedResumeId?: string; respawnedDeadPane: boolean }> {
const mux = this._mux!;
// Verify stale mux session — tmux may have been destroyed (e.g., killed externally)
// Verify stale mux session — tmux may have been destroyed (e.g., killed externally).
// A session that HAD a mux session relaunches its CLI below just like a failed
// respawn does (tmux kill-server, a tmux crash, an external kill-session), so
// its transcript collides with the bare `--session-id` the same way. A
// genuinely new session starts with `_muxSession` null and never sets this.
let muxSessionVanished = false;
if (this._muxSession && !mux.muxSessionExists(this._muxSession.muxName)) {
console.log('[Session] Stale mux session detected (tmux gone):', this._muxSession.muxName);
this._muxSession = null;
muxSessionVanished = true;
}
// Check if session exists but pane is dead (remain-on-exit keeps it alive)
// Respawn the pane instead of creating a whole new session — preserves tmux scrollback
let needsNewSession = false;
let respawnedDeadPane = false;
let respawnedResumeId: string | undefined;
if (this._muxSession && mux.isPaneDead(this._muxSession.muxName)) {
console.log('[Session] Dead pane detected, respawning:', this._muxSession.muxName);
// Confirmed dead — safe to resolve/pin now (see `_pinOmpRespawnId()`).
// `options.respawnPaneOptions` was built eagerly before this dead-pane
// check ran, so it still carries the pre-pin ompConfig; rebuild it.
this._pinOmpRespawnId();
const newPid = await mux.respawnPane(this._buildRespawnPaneOptions());
const respawnOptions = await this._buildRespawnPaneOptionsWithResumePin();
const newPid = await mux.respawnPane(respawnOptions);
if (!newPid) {
console.error('[Session] Failed to respawn pane, will create new session');
needsNewSession = true;
} else {
respawnedDeadPane = true;
respawnedResumeId = respawnOptions.resumeSessionId;
this._pendingEnvUnsets.clear();
// Wait a moment for the respawned process to fully start
await new Promise((resolve) => setTimeout(resolve, MUX_STARTUP_DELAY_MS));
}
}
// Whatever the last reading said about the OLD command in this pane is now
// history: the branch above either respawned the pane or found it alive, and
// the branch below creates a new one. The paths that reach here are boot
// recovery and an explicit start, NOT a click on an exited tab — the browser
// re-attaches only on a null pid, and the premise of Ark0N/Codeman#446 is
// that an exited pane keeps its pid. `restartCli()` clears separately.
this.clearPaneExitForNewPane();
// Check if we already have a mux session (restored session)
const isRestored = this._muxSession !== null && !needsNewSession;
if (isRestored) {
console.log('[Session] Attaching to existing mux session:', this._muxSession!.muxName);
} else {
// Create a new mux session
// Create a new mux session. When this is the FALLBACK after a failed
// respawn, the eagerly-built create options still carry the unpinned
// launch seed, so a session whose transcript exists would meet the same
// `--session-id ... already in use` refusal the respawn just lost to —
// the recovery of last resort failing for the very reason it was needed.
// A genuinely new session has no transcript under any of its candidate
// ids, so nothing is pinned and its command shape is unchanged.
//
// `_resumeSessionId` is written alongside, not just the create options:
// this branch leaves `isRestored` false, so the block that sets
// `_claudeSessionId` below reads that field and would otherwise settle on
// `this.id` while the CLI resumes the chain tail. The response viewer,
// Read My Mind and the unified-list alias map all read `_claudeSessionId`
// until the next first-hand hook, so the two have to name the same
// conversation. The vanished-tmux-session branch above relaunches for the
// same reason and takes the same pin.
if (needsNewSession || muxSessionVanished) {
const pinned = (await this._buildRespawnPaneOptionsWithResumePin()).resumeSessionId;
if (pinned) {
options.createSessionOptions.resumeSessionId = pinned;
this._resumeSessionId = pinned;
}
}
this._muxSession = await mux.createSession(options.createSessionOptions);
console.log('[Session] Created mux session:', this._muxSession.muxName);
// No extra sleep — createSession() already waits for tmux readiness
@@ -1800,13 +2064,14 @@ export class Session extends EventEmitter {
env: buildMuxAttachEnv(cliExportsTruecolor(this.mode)),
})
);
this._notePtySpawnGeometry(ptyCols, ptyRows);
} catch (spawnErr) {
console.error(`[Session] Failed to spawn PTY for ${options.spawnErrLabel}:`, spawnErr);
this.emit('error', `Failed to attach to mux session: ${spawnErr}`);
throw spawnErr;
}
return { isRestored };
return { isRestored, respawnedResumeId, respawnedDeadPane };
}
/**
@@ -1844,6 +2109,9 @@ export class Session extends EventEmitter {
console.error('[Session] reattachRemote: respawnPane failed for', this._muxSession.muxName);
return false;
}
// No-op for the record (a remote session's field is always UNKNOWN), but the
// mux layer's cache is keyed by muxName and this pane now runs a new client.
this.clearPaneExitForNewPane();
console.log('[Session] reattachRemote: reattached remote session', this._muxSession.muxName, 'pid', newPid);
return true;
}
@@ -1869,6 +2137,10 @@ export class Session extends EventEmitter {
* the mux session is gone — see {@link reattachRemote} for that reasoning).
*/
async restartCli(): Promise<boolean> {
return this._withPaneLifecycle(() => this._doRestartCli());
}
private async _doRestartCli(): Promise<boolean> {
if (!this._useMux || !this._mux || !this._muxSession) return false;
const mux = this._mux;
@@ -1878,26 +2150,16 @@ export class Session extends EventEmitter {
}
this._pinOmpRespawnId();
const options = this._buildRespawnPaneOptions();
// Unlike the dead-pane respawn, this one kills a WORKING pane whose conversation
// already has a transcript, and a CLI that launches with `--session-id <id>` refuses
// an id that is already in use (claude: `Error: Session ID ... is already in use.`),
// which turned an endpoint switch into a dead pane and a lost session. A launch that
// declares a `fallback` chain renders `resume || new` once a resume id is set, the
// same `--resume <id> || --session-id <id>` shape the docker and remote pane commands
// already use, so pin the live conversation id for THIS respawn only. The registry
// shape is the gate, not the CLI's name: an entry whose resume id is minted by the
// CLI itself (codex/pi/omp/grok) never declares that chain, and its resume field is
// read from its own `<Mode>Config` rather than this top-level one anyway.
if (!options.resumeSessionId && getCli(this.mode)?.launch.chain === 'fallback') {
options.resumeSessionId = this._claudeSessionId ?? this.id;
}
const newPid = await mux.respawnPane(options);
const newPid = await mux.respawnPane(await this._buildRespawnPaneOptionsWithResumePin());
if (!newPid) {
console.error('[Session] restartCli: respawnPane failed for', this._muxSession.muxName);
return false;
}
this._pendingEnvUnsets.clear();
// A relaunch in the same pane, so any exit observed of the previous command
// is history. Without this the caller's persist-and-broadcast writes the old
// exit straight back onto a session that is running again.
this.clearPaneExitForNewPane();
console.log('[Session] restartCli: restarted CLI for', this._muxSession.muxName, 'pid', newPid);
return true;
}
@@ -1914,6 +2176,7 @@ export class Session extends EventEmitter {
workingDir: this.workingDir,
mode: this.mode,
name: this._name,
cliName: this.cliPinnedName,
niceConfig: this._niceConfig,
model: this._model,
claudeMode: this._claudeMode,
@@ -1947,6 +2210,112 @@ export class Session extends EventEmitter {
return this._withCustomModelLaunchModel(options);
}
/**
* Respawn options for a pane whose command is being REPLACED, with the
* conversation pinned so the relaunch resumes rather than collides.
*
* A CLI that launches with `--session-id <id>` refuses an id that is already
* in use (claude: `Error: Session ID ... is already in use.`), and every
* session whose agent has been prompted owns a transcript under that id. So
* relaunching such a pane with the bare launch line fails, the pane dies
* again immediately, and the user's conversation is stranded. A launch that
* declares a `fallback` chain renders `resume || new` once a resume id is
* set, which is the shape that survives both cases.
*
* Three candidates are tried in priority order — the conversation chain's
* tail, the launch seed, then the session's own id — and the first one a
* transcript backs is pinned. Four conditions gate that walk, each protecting
* against a way of resuming the WRONG conversation or of making a working
* relaunch fail.
*
* ⚠️ **A remote or docker session is never pinned.** Unlike `restartCli()`,
* whose route refuses both, the dead-pane respawn is reached by every session
* shape. Their pane commands (`buildRemoteLaunchCommand`,
* `claudeDockerPaneCommand`) already render a SELF-HEALING
* `--session-id <sid> || --resume <sid>`, and both flip to resume-first the
* moment the resume id differs from the session id. The conversation lives on
* the far side, so a local id pinned onto it resolves to nothing there, the
* resume fails, and the `--session-id` fallback then collides with the
* transcript the far side really does hold — both branches fail and the pane
* dies. `_pinOmpRespawnId()` refuses remote for the same reason.
*
* ⚠️ **The candidates come from the conversation CHAIN, never from
* `_claudeSessionId`.** That field holds either a first-hand id from the
* CLI's own hook payload or a history correlation, which is a guess keyed on
* the working directory. `_recordClaudeSessionInChain()` refuses a guess
* precisely so it cannot "write a foreign conversation into this pane's
* permanent record", and launching from one would do worse than the display
* bug that rule exists to prevent: the relaunched CLI would open and WRITE to
* a conversation that was never this pane's. The chain's tail is the live
* conversation and is hook-vouched, so it leads the walk, ahead of the launch
* seed, which is written once at construction and never moves off a `/clear`.
*
* ⚠️ **Every candidate must be backed by a transcript, the session's own id
* included, and a candidate that has none is passed over rather than ending
* the walk.** A pin that differs from the session id leaves
* `--session-id <this.id>` in the fallback branch, so a resume that finds
* nothing collides there and the pane dies exactly as it did before this
* pinning existed. Pinning `this.id` renders the self-healing
* `--resume <id> || --session-id <id>`, which is correct whether or not a
* transcript exists, but a pane that has none pays for the shape twice:
* claude prints "No conversation found" into the scrollback of a session that
* is brand new, and `wrapWithNice()` prefixes only the FIRST branch of the
* rendered `a || b`, so the branch that actually runs loses its priority for
* the life of the session. Falling off the end of the walk therefore adds
* no pin (the options keep any launch seed they already carried), which is
* the right answer: with no transcript anywhere there is nothing for the
* bare `--session-id <this.id>` to collide with.
*
* The create route pre-validates a resume id for the same reason, though it
* additionally requires the transcript be substantial — here mere existence
* is the question, because a one-line transcript still makes `--session-id`
* collide.
*
* The registry shape is the last gate, not the CLI's name: an entry whose
* resume id is minted by the CLI itself (codex/pi/omp/grok) declares no
* `fallback` chain and reads its resume field from its own `<Mode>Config`.
*
* `reattachRemote()` deliberately does NOT call this. It re-runs the remote
* session command, which attaches to the durable remote tmux with the agent
* still inside it and renders no local `--session-id` to collide.
*/
private async _buildRespawnPaneOptionsWithResumePin(): Promise<import('./mux-interface.js').RespawnPaneOptions> {
const options = this._buildRespawnPaneOptions();
if (this._remote || this._docker) return options;
const entry = getCli(this.mode);
if (entry?.launch.chain !== 'fallback') return options;
const resumeIdPattern = entry.launch.params?.resumeId;
const configDir = this._claudeConfigDir();
const chainTail = this._claudeSessionChain[this._claudeSessionChain.length - 1];
const candidates = [chainTail, options.resumeSessionId, this.id].filter((v): v is string => !!v);
for (const candidate of candidates) {
// A session Codeman DISCOVERED on the socket rather than created carries a
// synthetic `restored-<fragment>` id, which fails claude's `uuid` token
// pattern. The renderer would silently drop the resume flag and emit the
// unpinned command, so say so here rather than letting the caller believe
// the pane was pinned.
if (resumeIdPattern?.type === 'token' && !matchesPattern(resumeIdPattern.pattern, candidate)) {
console.log(`[Session] Not pinning resume id ${candidate} for relaunch: the CLI cannot accept that id shape`);
continue;
}
if (!(await claudeTranscriptExists(candidate, configDir))) continue;
options.resumeSessionId = candidate;
return options;
}
// Nothing on disk to collide with, so the bare `--session-id <this.id>` the
// unpinned options already carry is the correct command.
return options;
}
/** The session's Claude config dir when it has been relocated (#255), else undefined. */
private _claudeConfigDir(): string | undefined {
// Trimmed like `claudeProjectsDir()` trims the process-wide override: the
// envOverrides schema validates keys only, and a whitespace-only value would
// otherwise resolve to a relative path and read "no transcript" for everything.
return this._envOverrides?.CLAUDE_CONFIG_DIR?.trim() || undefined;
}
/**
* Force the custom-model selection's `launchModel` (pi/omp `custom/<id>`, grok's
* `[model.<name>]` block name) onto the CLI's `model` launch param. Where that param
@@ -2160,6 +2529,9 @@ export class Session extends EventEmitter {
if (this.ptyProcess) {
throw new Error('Session already has a running process');
}
if (this._closing) {
throw new Error('Session is being closed');
}
// Bounds the workspace-trust scan (see _maybeAcceptTrustDialog). Stamped here
// rather than at PTY spawn so a slow mux attach still counts as startup.
@@ -2268,7 +2640,7 @@ export class Session extends EventEmitter {
// If mux wrapping is enabled, create or attach to a mux session
if (this._useMux && this._mux) {
try {
const { isRestored } = await this._setupOrAttachMuxSession({
const { isRestored, respawnedResumeId, respawnedDeadPane } = await this._setupOrAttachMuxSession({
// Single source of truth shared with reattachRemote() (COD-108).
respawnPaneOptions: this._buildRespawnPaneOptions(),
createSessionOptions: {
@@ -2276,6 +2648,7 @@ export class Session extends EventEmitter {
workingDir: this.workingDir,
mode: this.mode,
name: this._name,
cliName: this.cliPinnedName,
niceConfig: this._niceConfig,
model: this._model,
claudeMode: this._claudeMode,
@@ -2315,7 +2688,18 @@ export class Session extends EventEmitter {
// persisted chain's tail is that conversation, reported first-hand by
// the CLI's own hook, so it outranks every fallback here. A NEW pane has
// an empty chain and falls through to the resume/alias fallbacks.
restoredConversation = isRestored ? this._claudeSessionChain[this._claudeSessionChain.length - 1] : undefined;
//
// A dead-pane respawn is NOT that case for a CLI whose relaunch the resume
// pin walk governs (`launch.chain === 'fallback'`): the CLI did stop, and
// the walk may have passed over a chain tail with no transcript behind it,
// so the conversation is whatever the respawn actually resumed. Undefined
// there means the pane launched unpinned, which the fallbacks below name.
const pinGovernsRespawn = respawnedDeadPane && getCli(this.mode)?.launch.chain === 'fallback';
restoredConversation = pinGovernsRespawn
? respawnedResumeId
: isRestored
? this._claudeSessionChain[this._claudeSessionChain.length - 1]
: undefined;
this._claudeSessionId =
restoredConversation ||
this._resumeSessionId ||
@@ -2403,7 +2787,7 @@ export class Session extends EventEmitter {
this._model,
this._allowedTools,
this._effort,
this._name,
this.cliPinnedName,
getClaudeCliVersion()
);
this.ptyProcess = spawnPtyWithHelperRepair(() =>
@@ -2416,6 +2800,7 @@ export class Session extends EventEmitter {
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
})
);
this._notePtySpawnGeometry(120, 40);
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY:', spawnErr);
this._status = 'stopped';
@@ -2709,9 +3094,63 @@ export class Session extends EventEmitter {
this._lastPaneProbeAt = now;
const text = this._mux.capturePaneText?.(this._muxSession.muxName) ?? null;
this._lastPaneProbeWorking = text === null ? null : this._workingLinePattern().test(text);
this._readWatching(text);
return this._lastPaneProbeWorking;
}
/**
* Read the background-work chip off the same capture the working probe just took.
*
* The two questions are different. A turn that is running is work the user is waiting
* for; a monitor, a backgrounded shell or a cloud session the agent started is work
* the AGENT is waiting for, and it is the reason a pane can sit at its composer with
* nothing to say and still not want anything from the user. `_confirmIdle` takes this
* capture at exactly the moment the turn ends, which is the moment the answer starts
* mattering.
*
* A capture that could not be read CLEARS the label rather than keeping the last one.
* The two wrong answers are not symmetric: a stale label opens the next idle prompt
* already acknowledged, so a failed capture would silence a real alert, while a dropped
* label only costs a card and an alert that the next readable capture takes back.
* Degrading toward the alert is the rule the whole signal is built on.
*/
private _readWatching(paneText: string | null): void {
const pattern = this._watchingLinePattern();
if (!pattern) return;
// `null` is "the screen could not be read". That is no evidence either way, so the
// label falls to null (and the change is announced below like any other).
const label = paneText === null ? null : watchingLabel(paneText, pattern, this._watchingWindow);
if (label === this._watching) return;
this._watching = label;
// ⚠️ This CHANGES while the session's status does not, so it needs an event of its
// own. The label is usually set on the idle transition, which broadcasts anyway, but
// it CLEARS when the work ends — and for a CLI whose background work ends without
// taking a turn (measured on codex: a background terminal finishing repaints the row
// away and nothing else happens) the session is idle before and after. Without this,
// the server knew the badge was gone and every open page went on drawing it until
// some unrelated event arrived.
this.emit('watchingChanged');
}
/**
* The regex matching this CLI's background-work chip, or null for a CLI whose registry
* entry declares none. Compiled once per session, like the working-line pattern, and
* null rather than a fallback: no other CLI has been measured drawing such a chip, and
* guessing one would badge sessions on the strength of an unread screen.
*/
private _watchingLinePattern(): RegExp | null {
if (this._watchingLineRe === undefined) {
// The pattern and the window it runs over are one decision, so they are resolved
// together: how far up the screen a CLI's row can sit is as much a property of its
// layout as the row itself. Claude writes on the last row and keeps the default,
// Codex pins one above its composer and declares more.
const detect = getCli(this.mode)?.capabilities.workDetect;
this._watchingLineRe = detect?.watchingLine ? compileVersionRegex(detect.watchingLine) : null;
this._watchingWindow = detect?.watchingLines ?? WATCHING_TAIL_LINES;
}
return this._watchingLineRe;
}
/**
* The regex matching this CLI's "a turn is running" status line.
*
@@ -2914,6 +3353,9 @@ export class Session extends EventEmitter {
if (this.ptyProcess) {
throw new Error('Session already has a running process');
}
if (this._closing) {
throw new Error('Session is being closed');
}
this._resetBuffers();
@@ -2984,6 +3426,7 @@ export class Session extends EventEmitter {
env: buildShellEnv(this.id),
})
);
this._notePtySpawnGeometry(120, 40);
} catch (spawnErr) {
console.error('[Session] Failed to spawn shell PTY:', spawnErr);
this._status = 'stopped';
@@ -3093,6 +3536,7 @@ export class Session extends EventEmitter {
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
})
);
this._notePtySpawnGeometry(120, 40);
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY for runPrompt:', spawnErr);
this.emit(
@@ -3758,7 +4202,18 @@ export class Session extends EventEmitter {
? null
: this._mux.capturePaneText?.(this._muxSession.muxName),
sendEnter: () => this._mux?.sendInput(this.id, '\r'),
glyph: () => getCli(this.mode)?.capabilities.workDetect?.promptGlyph ?? '❯',
// ⚠ NO fallback glyph here, unlike the screen-reading probe elsewhere in this file.
// Only claude and codex declare a promptGlyph; the other eight modes would fall back
// to claude's `❯`, which is ALSO starship's default shell prompt (and pure's, and
// spaceship's, and p10k lean's). On a shell session the line `❯ npm run build` sits
// on screen for as long as the command runs, promptStillInComposer() reads that as
// "still unsubmitted", and the verifier presses Enter into the running program's
// stdin on its 2s..60s schedule. Mostly a stray blank line; not harmless against a
// y/N prompt, `read -p`, an installer or a pager, where it takes the default.
// promptStillInComposer() returns undefined for an empty glyph, so this makes the
// verifier inert for every CLI that does not declare one, which is what the Claude
// Code 2.1.277 defect it exists for actually calls for.
glyph: () => getCli(this.mode)?.capabilities.workDetect?.promptGlyph ?? '',
log: (m) => console.log(`[Session ${this.id.slice(0, 8)}] ${m}`),
});
this._submitVerifier.arm(text);
@@ -3768,6 +4223,40 @@ export class Session extends EventEmitter {
private _ptyCols = 120;
private _ptyRows = 40;
/**
* Record the geometry a PTY was just spawned at. A reattached pane keeps the
* tmux window's size, not the constructor's 120x40, and without this
* `ptyGeometry` reported the old numbers for a live pane and the dedupe in
* `resize()` skipped a real resize that happened to match them.
*/
private _notePtySpawnGeometry(cols: number, rows: number): void {
this._ptyCols = cols;
this._ptyRows = rows;
}
/**
* The geometry the CLI is actually drawing for, or null when nothing is
* drawing.
*
* Exposed because `resize()` can decline a request outright (arbitration
* below) and the asking client has no other way to find out: a browser
* terminal that keeps a shape the PTY refused renders garbled output rather
* than wrong-sized output, because Claude Code's repaints are computed from
* the width it was told (issue #464). Both transports report this back.
*
* ⚠️ NULL WITHOUT A PANE, never the field values. The fields are seeded at
* spawn (`_notePtySpawnGeometry`) and moved by `resize()`, but a session with
* a dead pane (or one created through the API and never started) still
* holds the constructor defaults of 120x40, or the size of a pane that is
* gone. Reporting those made a client adopt a size no process had ever
* been told, and on anything narrower than 120 columns it claimed another
* device owned the pane when none existed. `reconcilePtyGeometry` treats a
* report with no finite numbers as no evidence, which is the truth here.
*/
get ptyGeometry(): { cols: number; rows: number } | null {
return this.ptyProcess ? { cols: this._ptyCols, rows: this._ptyRows } : null;
}
/**
* Live WebSocket connections that have announced a desktop viewport for this
* session. While at least one is registered, small-viewport (mobile/tablet)
@@ -3853,6 +4342,10 @@ export class Session extends EventEmitter {
}
if (isSmallViewport && this._desktopSizeClaims.size > 0) {
if (Date.now() - this._lastDesktopActivityAt < Session.DESKTOP_CLAIM_IDLE_MS) {
// Declined. The caller is told nothing here on purpose — the decision
// belongs to the session, not the socket — but the caller MUST report
// `ptyGeometry` back afterwards so the asking client can adopt
// the shape it did not get. Both transports do; see issue #464.
return;
}
this._mobileSizeOverride = true;
@@ -3952,6 +4445,9 @@ export class Session extends EventEmitter {
async stop(killMux: boolean = true): Promise<void> {
// Set stopped flag first to prevent new timers from being created
this._isStopped = true;
// A pane that is gone is watching nothing. Nothing probes a stopped session, so
// without this the last chip it drew would ride along on its row forever.
this._watching = null;
this._clearAllTimers();
+404 -19
View File
@@ -58,6 +58,7 @@ import {
type SessionRemote,
type SessionDocker,
type DockerCommandMode,
type PaneExit,
} from './types.js';
import { getCli } from './config/cli-registry/registry.js';
import { missingCliMessage, resolveCliBinDir } from './utils/cli-resolver.js';
@@ -152,6 +153,16 @@ const GRACEFUL_SHUTDOWN_WAIT_MS = 100;
/** Default stats collection interval (2 seconds) */
const DEFAULT_STATS_INTERVAL_MS = 2000;
/**
* How often the pane-exit watcher re-reads every pane on the socket. The
* watcher owns this cadence: it does NOT ride `startStatsCollection()`, whose
* lifetime a browser panel controls (see {@link TmuxManager.startPaneExitWatcher}).
* Matched to the stats cadence above because both cost one batched tmux read.
* ⚠ It does NOT bound a read: EXEC_TIMEOUT_MS is 5000 ms, so a slow read can
* outlive two ticks, which is exactly why `paneExitReadInFlight` exists.
*/
const DEFAULT_PANE_EXIT_INTERVAL_MS = 2000;
/** Default remote-reconnect watcher poll interval (5 seconds) — COD-108 */
const DEFAULT_REMOTE_RECONNECT_INTERVAL_MS = 5000;
@@ -219,8 +230,18 @@ const DEFAULT_CODEMAN_TMUX_SOCKET = DEFAULT_TMUX_SOCKET;
*/
const PANE_LIST_SEP = '|';
/** Format string for `tmux list-panes -F`. Keep in sync with {@link parsePaneList}. */
const PANE_LIST_FORMAT = `#{session_name}${PANE_LIST_SEP}#{pane_pid}`;
/**
* Format string for `tmux list-panes -F`. Keep in sync with {@link parsePaneRows}.
*
* The three `pane_dead*` fields carry the agent-exit signal of Ark0N/Codeman#446.
* Appending them is backward compatible in both directions. A tmux that does not
* know a variable substitutes the empty string rather than failing, which is how
* tmux 3.2a answers `#{pane_dead_signal}` (added in 3.4), and the parser reads a
* short row as "pid known, deadness unknown" rather than discarding it.
*/
const PANE_LIST_FORMAT =
`#{session_name}${PANE_LIST_SEP}#{pane_pid}` +
`${PANE_LIST_SEP}#{pane_dead}${PANE_LIST_SEP}#{pane_dead_status}${PANE_LIST_SEP}#{pane_dead_signal}`;
/**
* 构建 pane 启动前的 nofile 修复命令。
@@ -235,26 +256,154 @@ export function buildNofileLimitCommand(targetLimit = CLAUDE_CODE_NOFILE_LIMIT):
return `ulimit -Sn ${safeLimit} 2>/dev/null || ulimit -n ${safeLimit} 2>/dev/null || true`;
}
/** One pane of one tmux session, as {@link parsePaneRows} reads it off the wire. */
export interface PaneRow {
/** tmux session this pane belongs to. Repeats once per pane of a split session. */
sessionName: string;
/** `#{pane_pid}` — the process tmux started in the pane. */
pid: number;
/** `#{pane_dead}` — true for 1, false for 0, undefined when tmux said nothing. */
dead?: boolean;
/** `#{pane_dead_status}` — the exit code, absent when tmux reported none. */
exitStatus?: number;
/** `#{pane_dead_signal}` — the killing signal, absent before tmux 3.4 and when unsignalled. */
exitSignal?: number;
}
/**
* Parse the output of `tmux list-panes -a -F '#{session_name}|#{pane_pid}'`
* into a Map of session-name → pane pid. Exported for unit testing.
* One pane-exit reading, with the pane pid that produced it.
*
* The pid never leaves this module. It is what distinguishes "the same dead
* pane, seen again" from "a second command in the same pane that also exited
* with the same status", so the `at` stamp can hold across the first and must
* not across the second. {@link PaneExit} itself stays free of it: the pid on
* the session record is the attach client's, and a second pid there would
* invite exactly the confusion Ark0N/Codeman#446 is about.
*/
export interface PaneExitObservation {
/** `#{pane_pid}` of the pane this reading came from. */
panePid: number;
/** What to publish on the session record. */
exit: PaneExit;
}
/**
* A {@link PaneExitObservation} as the manager stores it, with a count of the
* authoritative reads that have seen this same exit. The count is what lets
* the exited-agent sweep act only on a death that more than one read agreed on
* (`CLEAN_EXIT_CONFIRMING_READS` in `pane-exit-sweep.ts`). A failed or skipped
* read never reaches {@link TmuxManager.applyPaneExits}, so it neither raises
* the count nor resets it.
*/
interface TrackedPaneExit extends PaneExitObservation {
/** Authoritative reads that saw this exit, counting the first. */
reads: number;
}
/** Read one optional numeric field; a blank or non-numeric value is "not reported". */
function paneField(fields: string[], index: number): number | undefined {
const raw = fields[index];
if (raw === undefined || raw === '') return undefined;
const value = parseInt(raw, 10);
return Number.isNaN(value) ? undefined : value;
}
/**
* Parse the output of `tmux list-panes -a -F` under {@link PANE_LIST_FORMAT}
* into one row per pane, in tmux's own order. Exported for unit testing.
*
* - Skips empty lines and lines without the separator.
* - Skips entries with a non-numeric pid or empty name.
* - Leaves every field after the pid undefined when it is blank or absent, so a
* row from an older tmux still yields its pid.
*/
export function parsePaneList(output: string): Map<string, number> {
const result = new Map<string, number>();
export function parsePaneRows(output: string): PaneRow[] {
const rows: PaneRow[] = [];
for (const line of output.split('\n')) {
if (!line) continue;
const sep = line.indexOf(PANE_LIST_SEP);
if (sep === -1) continue;
const name = line.slice(0, sep);
const pid = parseInt(line.slice(sep + 1), 10);
if (name && !Number.isNaN(pid)) {
result.set(name, pid);
}
if (!line.includes(PANE_LIST_SEP)) continue;
const fields = line.split(PANE_LIST_SEP);
const sessionName = fields[0];
const pid = parseInt(fields[1] ?? '', 10);
if (!sessionName || Number.isNaN(pid)) continue;
const deadFlag = fields[2];
rows.push({
sessionName,
pid,
dead: deadFlag === '1' ? true : deadFlag === '0' ? false : undefined,
exitStatus: paneField(fields, 3),
exitSignal: paneField(fields, 4),
});
}
return result;
return rows;
}
/**
* Decide, from every pane tmux listed, which tmux sessions have an exited agent.
* Exported for unit testing. Returns one entry per session with a known answer;
* a session absent from the map is UNKNOWN, which must never render as alive.
*
* Two rules make a positive answer trustworthy:
*
* A session answers only when tmux listed EXACTLY ONE pane for it. Codeman
* creates one pane per session and `isPaneDead()` reads one pane, so a session
* the user has split by hand has no single "the agent" to report on, and
* guessing which of its panes speaks for the session could report a live
* session as exited.
*
* A pane answers only when `#{pane_dead}` said 1 or 0. An empty field is a tmux
* that did not answer, not a live pane.
*
* `status` and `signal` stay absent when tmux did not report them. Measured on
* tmux 3.2a, a SIGKILLed pane reports neither, so folding an absent status into
* 0 would turn an unexplained death into a clean exit.
*/
export function derivePaneExits(rows: PaneRow[], now: number): Map<string, PaneExitObservation> {
const panesPerSession = new Map<string, number>();
for (const row of rows) {
panesPerSession.set(row.sessionName, (panesPerSession.get(row.sessionName) ?? 0) + 1);
}
const exits = new Map<string, PaneExitObservation>();
for (const row of rows) {
if (panesPerSession.get(row.sessionName) !== 1) continue;
if (row.dead !== true) continue;
exits.set(row.sessionName, {
panePid: row.pid,
exit: {
...(row.exitStatus !== undefined ? { status: row.exitStatus } : {}),
...(row.exitSignal !== undefined ? { signal: row.exitSignal } : {}),
at: now,
},
});
}
return exits;
}
/**
* Could any of these tmux sessions ever produce a pane-exit answer? Exported
* for unit testing.
*
* Mirrors `Session.paneExitApplies`, which is where the rule is enforced. A
* remote session's local pane holds the ssh client, a docker case's holds a
* `docker exec` into the container's own tmux, and a record rebuilt from the
* socket carries no provenance at all, so the session end forces all three to
* UNKNOWN whatever tmux reports. A tick that sees only those has nothing to
* learn, and `refreshPaneExits()` skips its tmux read rather than paying for
* the answer.
*
* ⚠ This gates the READ, never the watcher. The watcher is always-on by
* design (see {@link TmuxManager.startPaneExitWatcher}), so it keeps ticking
* with nothing to observe and picks the read straight back up as soon as one
* local session exists.
*/
export function hasObservablePaneSession(sessions: Iterable<MuxSession>): boolean {
for (const session of sessions) {
if (session.remote) continue;
if (session.docker) continue;
if (session.discovered === true) continue;
return true;
}
return false;
}
/**
@@ -735,7 +884,7 @@ export function buildSpawnCommand(options: {
effort?: EffortLevel;
/** Resolved by resolveStatusLineCliCommand (hooks-config.ts) — undefined skips the exporter. Claude only. */
statusLineCommand?: string;
/** Codeman session name, passed to claude as `--name` (version-gated, sanitized; local spawns only). */
/** Name pinned on claude as `--name` (version-gated, sanitized; local spawns only). Only a user-chosen name: see `Session.cliPinnedName`. */
sessionName?: string;
/**
* Claude CLI version for the `--name` gate. Omitted = probe the local CLI
@@ -1527,6 +1676,30 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
private mouseSyncInterval: NodeJS.Timeout | null = null;
/** Track last-known pane count per session to avoid unnecessary tmux set-option calls */
private lastPaneCount: Map<string, number> = new Map();
/**
* muxName → the exited agent the pane-exit watcher last observed
* (Ark0N/Codeman#446). Absence is the UNKNOWN arm of the tri-state, so an
* entry goes the moment tmux stops reporting the pane dead, and the map is
* empty until the first read runs. The manager reports what tmux says and
* nothing more: the scoping that hides this for remote and docker sessions
* lives on `Session`, because the remote-reconnect watcher above needs the
* raw pane reading.
*/
private paneExits: Map<string, TrackedPaneExit> = new Map();
/** The pane-exit watcher's own interval. Runs whether or not stats are on. */
private paneExitInterval: NodeJS.Timeout | null = null;
/**
* True while a pane read is in flight. `EXEC_TIMEOUT_MS` is 5000 ms against a
* poll interval of 2000 ms, so without this a slow read overlaps the next two
* and the older one can resolve last and win.
*/
private paneExitReadInFlight = false;
/**
* Bumped by every deliberate {@link clearPaneExit}. A read that started before
* a clear carries the older generation and is discarded rather than writing
* the death back over the pane that has just replaced it.
*/
private paneExitGeneration = 0;
// ── COD-108 remote-reconnect watcher state ────────────────────────────────
/** Periodic watcher that re-establishes dropped remote sessions. */
@@ -1894,6 +2067,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
workingDir,
mode,
name,
cliName,
niceConfig,
model,
claudeMode,
@@ -1997,7 +2171,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
resumeSessionId,
effort,
statusLineCommand,
sessionName: name,
sessionName: cliName,
});
const config = niceConfig || DEFAULT_NICE_CONFIG;
@@ -2225,7 +2399,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
effort,
remote,
docker,
name,
cliName,
} = options;
const session = this.sessions.get(sessionId);
if (!session) return null;
@@ -2262,7 +2436,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
resumeSessionId,
effort,
statusLineCommand,
sessionName: name,
sessionName: cliName,
});
const config = niceConfig || DEFAULT_NICE_CONFIG;
const cmd = wrapWithNice(baseCmd, config);
@@ -2297,6 +2471,11 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
);
// Wait for the respawned process to start
await new Promise((resolve) => setTimeout(resolve, TMUX_CREATION_WAIT_MS));
// The pane now runs a fresh command, so whatever the last read observed of
// the old one is history. Clearing it here rather than waiting for the next
// poll also invalidates any read already in flight, which would otherwise
// write the old death back over the pane that just replaced it.
this.clearPaneExit(muxName);
const pid = this.getPanePid(muxName);
if (pid) session.pid = pid;
return pid;
@@ -2508,6 +2687,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
}
this.lastPaneCount.delete(session.muxName);
this.clearPaneExit(session.muxName);
this.sessions.delete(sessionId);
this.clearRemoteReconnectState(sessionId);
this.saveSessions();
@@ -2618,6 +2798,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
this.lastPaneCount.delete(session.muxName);
this.clearPaneExit(session.muxName);
this.sessions.delete(sessionId);
this.clearRemoteReconnectState(sessionId);
this.saveSessions();
@@ -2671,7 +2852,14 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
active = parsePaneList(output);
const rows = parsePaneRows(output);
active = new Map(rows.map((row) => [row.sessionName, row.pid]));
// The same read answers both questions, so recovery starts with a pane-exit
// reading rather than waiting for the first stats tick — which may never
// come, since the collector only starts when boot found a live session.
if (rows.length > 0) {
this.applyPaneExits(derivePaneExits(rows, Date.now()));
}
} catch (err) {
console.error('[TmuxManager] Failed to list tmux panes:', err);
active = new Map();
@@ -2686,6 +2874,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
} else {
dead.push(sessionId);
this.sessions.delete(sessionId);
this.clearPaneExit(session.muxName);
this.clearRemoteReconnectState(sessionId);
this.emit('sessionDied', { sessionId });
}
@@ -2722,6 +2911,13 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
mode: 'claude',
attached: false,
name: `Restored: ${sessionName}`,
// Every field above except the name and the pid is a guess: this record
// was rebuilt from the socket because Codeman's own bookkeeping did not
// have it. The synthetic id also cannot find the session's state.json
// entry, so a remote or docker session rediscovered this way arrives
// looking local. Consumers that would be wrong about such a session
// read this flag and fail closed — see `Session.paneExitApplies`.
discovered: true,
};
this.sessions.set(sessionId, session);
knownMuxNames.add(sessionName);
@@ -2882,6 +3078,185 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}));
}
/**
* What the last pane read saw of this tmux session's agent. `undefined` is
* the UNKNOWN answer and must never be rendered as "alive": it covers a pane
* that is running, a tmux session that no longer exists, a probe that failed,
* and every poll that has not run yet. See {@link PaneExit}.
*/
getPaneExit(muxName: string): PaneExit | undefined {
return this.paneExits.get(muxName)?.exit;
}
/**
* How many authoritative pane reads have agreed on the exit that
* {@link getPaneExit} reports, or 0 when it reports none. A new observation
* starts at 1, and every later read that sees the same pane with the same
* status and signal adds one.
*/
getPaneExitReadCount(muxName: string): number {
return this.paneExits.get(muxName)?.reads ?? 0;
}
/**
* Re-read every pane on the socket and refresh {@link paneExits}. ONE batched
* `tmux list-panes -a` answers for every session at once, which is why this
* polls rather than probing per session.
*
* A failed or empty probe leaves the previous answers ALONE rather than
* clearing them, because the two cannot be told apart: the command ends in
* `|| true`, so a tmux that errored and a socket with genuinely no panes both
* arrive as empty output. Treating that as "tmux did not answer" is the
* conservative reading — clearing on it would turn a transient failure into a
* silent retraction of a death Codeman had already observed, and the cost of
* being wrong the other way is one stale entry for a socket that no longer
* has the pane. A NON-empty read is different: `list-panes -a` lists
* every pane on the socket, so it is authoritative and {@link applyPaneExits}
* prunes against it.
*
* Two guards keep a slow read from undoing a fast one. A read already in
* flight suppresses the next poll, and a read that started before a
* {@link clearPaneExit} is discarded when it lands.
*
* A third guard skips the read entirely while no session on this manager
* could produce an answer ({@link hasObservablePaneSession}). Skipping
* retracts nothing, for the same reason a failed read does not: the map
* still holds what the last real read saw, and every path that puts a new
* command in a pane calls {@link clearPaneExit} itself.
*/
async refreshPaneExits(now: number = Date.now()): Promise<void> {
// Nothing on this socket could answer, so do not read tmux to find that
// out. See `hasObservablePaneSession`: the watcher above still ticks.
if (!hasObservablePaneSession(this.sessions.values())) return;
if (this.paneExitReadInFlight) return;
const generation = this.paneExitGeneration;
this.paneExitReadInFlight = true;
let rows: PaneRow[];
try {
rows = await this.readPaneRows();
} finally {
this.paneExitReadInFlight = false;
}
if (rows.length === 0) return;
// A pane was respawned or killed while this read was out, so what it saw is
// already history. Dropping it is what stops a freshly respawned pane from
// being republished as exited.
if (generation !== this.paneExitGeneration) return;
this.applyPaneExits(derivePaneExits(rows, now));
}
/**
* Read every pane on the socket. The ONLY part of the pane-exit watcher that
* touches tmux, which is what lets a test subclass drive the guards in
* {@link refreshPaneExits} — the in-flight suppression, the generation
* check, the empty-read retraction rule and the read gate — against rows it
* chooses. Split out for the reason `runRemoteReconnectTick` is: a guard no
* test can reach is a guard that can be deleted without anything failing.
*
* A failed read answers with NO rows, which the caller treats as "tmux did
* not answer" and which therefore retracts nothing.
*/
protected async readPaneRows(): Promise<PaneRow[]> {
// The test-mode gate lives HERE rather than at the top of the tick, so that
// what tests cannot do is spawn a process, not exercise the bookkeeping.
if (IS_TEST_MODE) return [];
try {
// execAsync, not execSync: this runs on a 2000 ms timer, and a synchronous
// exec freezes the port while the process stays alive (see the
// event-loop-monitor note in CLAUDE.md). The three `isPaneDead()` callers
// stay synchronous because each is answering one request right then.
const { stdout } = await execAsync(`${this.tmux()} list-panes -a -F '${PANE_LIST_FORMAT}' 2>/dev/null || true`, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
});
return parsePaneRows(stdout.trim());
} catch (err) {
console.error('[TmuxManager] Failed to read pane exit state:', err);
return [];
}
}
/**
* Fold one authoritative observation into {@link paneExits}. Split out from
* the tmux call so the merge rules are unit-testable.
*
* `observed` comes from a read of EVERY pane on the socket, so a session
* missing from it has no exit to report and its entry goes. That is what
* keeps the map from growing without bound as tmux sessions come and go
* outside `killSession()`. Only the caller may decide a read is authoritative:
* a failed or empty one never reaches here.
*
* An entry keeps the `at` of the FIRST read that saw that exit, so the stamp
* says when the agent was found gone rather than when the last poll ran. A
* changed status, a changed signal, or a different pane pid all start a new
* observation — the pid is what catches a second command in the same pane
* that happened to exit the same way.
*
* The same rule decides the read count: a repeat of the stored exit adds one,
* and anything that starts a new observation starts the count again at 1.
*/
applyPaneExits(observed: Map<string, PaneExitObservation>): void {
for (const muxName of [...this.paneExits.keys()]) {
if (!observed.has(muxName)) this.paneExits.delete(muxName);
}
for (const [muxName, next] of observed) {
const prev = this.paneExits.get(muxName);
const sameExit =
prev !== undefined &&
prev.panePid === next.panePid &&
prev.exit.status === next.exit.status &&
prev.exit.signal === next.exit.signal;
this.paneExits.set(muxName, sameExit ? { ...prev, reads: prev.reads + 1 } : { ...next, reads: 1 });
}
}
/**
* Forget a session's exit observation, e.g. once its pane has been respawned.
* Also invalidates any read already in flight, so the answer this retracts
* cannot be written back a moment later.
*/
clearPaneExit(muxName: string): void {
this.paneExits.delete(muxName);
this.paneExitGeneration++;
}
/**
* Poll for exited agents, on the manager's own interval.
*
* Deliberately NOT part of `startStatsCollection()`. That collector is armed
* when the browser opens the Monitor panel and DISARMED when it closes it
* (`panels-ui.js`), and it is skipped at boot entirely when no session was
* recovered — so riding it would leave a session created on a freshly booted
* server reporting nothing at all, and would let one browser turn exit
* detection off for every other. Started unconditionally, like the mouse-mode
* sync and the remote-reconnect watcher below.
*
* The `paneExitsUpdated` event is internal to the server; nothing here adds an
* SSE event, and the field reaches the browser on `session:updated`.
*/
startPaneExitWatcher(intervalMs: number = DEFAULT_PANE_EXIT_INTERVAL_MS): void {
if (this.paneExitInterval) {
clearInterval(this.paneExitInterval);
}
this.paneExitInterval = setInterval(() => {
// No IS_TEST_MODE guard: `readPaneRows()` is the only thing that would
// spawn a process and it refuses under test, so a test can drive this
// whole loop with fake timers instead of being locked out of it.
void this.refreshPaneExits()
.then(() => this.emit('paneExitsUpdated'))
.catch((err) => console.error('[TmuxManager] Pane exit watcher error:', err));
}, intervalMs);
}
stopPaneExitWatcher(): void {
if (this.paneExitInterval) {
clearInterval(this.paneExitInterval);
this.paneExitInterval = null;
}
}
startStatsCollection(intervalMs: number = DEFAULT_STATS_INTERVAL_MS): void {
if (this.statsInterval) {
clearInterval(this.statsInterval);
@@ -3101,6 +3476,8 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
destroy(): void {
this.stopStatsCollection();
this.stopPaneExitWatcher();
this.paneExits.clear();
this.stopMouseModeSync();
this.stopRemoteReconnectWatcher();
this.reconnectState.clear();
@@ -3482,6 +3859,14 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
{ encoding: 'utf-8', timeout: EXEC_TIMEOUT_MS }
)
);
// Report the size the pane was really drawing at. The visible-frame path
// below addresses every row absolutely, so a consumer whose terminal is
// shorter than this piles the overflow rows onto its last line and loses
// the rows it overwrote. The full-history path instead ends in a RELATIVE
// cursor move, which costs it nothing when the two sizes disagree, so the
// geometry is reported there for diagnosis rather than for repair. Only
// the caller can see both sizes, so hand it this one.
if (opts && geometry) opts.capturedGeometry = { cols: geometry.cols, rows: geometry.rows };
if (fullHistory) {
// Without geometry there is no cursor move, so fall back to the old trim.
+18 -5
View File
@@ -28,8 +28,11 @@
import type { ApprovalItem, ApprovalOption } from '../web/approval-inbox.js';
import type { TuiApprovalAnswer } from './tui-client.js';
/** Card severity, in the same red/yellow vocabulary the web inbox uses. */
export type TuiApprovalTone = 'err' | 'warn';
/**
* Card severity, in the same red/yellow vocabulary the web inbox uses, plus the quiet
* third case: an item that opened acknowledged asks for nothing and reads grey.
*/
export type TuiApprovalTone = 'err' | 'warn' | 'info';
export interface TuiApprovalCard {
tone: TuiApprovalTone;
@@ -50,8 +53,14 @@ function clean(text: string | undefined): string {
return (text ?? '').replace(/\s+/g, ' ').trim().slice(0, MAX_CARD_TEXT);
}
/**
* How loud the card is. An idle prompt the inbox opened ALREADY acknowledged is not
* asking for anything — the session is watching work it started itself — so it drops to
* `info` and out of the warning vocabulary the other two share with the web inbox.
*/
export function approvalTone(item: ApprovalItem): TuiApprovalTone {
return item.kind === 'idle' ? 'warn' : 'err';
if (item.kind !== 'idle') return 'err';
return item.acknowledgedReason ? 'info' : 'warn';
}
/**
@@ -65,9 +74,13 @@ export function approvalCard(item: ApprovalItem): TuiApprovalCard {
const summary = clean(item.toolSummary) || clean(item.toolName);
if (item.kind === 'idle') {
// Say what it is waiting for rather than asking for a reply, in the same words the
// web drawer uses for the same item. The prompt is still answerable, so the hint
// stays either way.
const quiet = clean(item.acknowledgedReason);
return {
tone: 'warn',
title: message || 'waiting for your reply',
tone: approvalTone(item),
title: quiet ? `quiet, ${quiet}` : message || 'waiting for your reply',
detail: [],
options: [],
hint: 'p to reply',
+18 -2
View File
@@ -74,13 +74,25 @@ export function isLiveRow(session: TuiSessionRow): boolean {
* outranks a stale `busy` status because the hook is the newer signal. An
* errored session has no state of its own here and joins the waiting tier,
* since it is equally something only a human can clear.
*
* ⚠️ An ACKNOWLEDGED item no longer decides the row. `acknowledgedAt` means the
* alert this prompt armed has been spent, either because somebody opened the
* session on another device or because the inbox opened the item that way for a
* session watching its own background work. The item itself stays pending and
* answerable, so the row keeps carrying it and the approval card still renders;
* it simply stops dragging the session into NEEDS YOU. The web has honoured
* that since acknowledgement existed (`approvals-ui.js` clears the pending hook
* that `_mobileOverviewState` reads), and this gate is where the TUI had been
* reading past it: acknowledging on a phone cleared the alert everywhere except
* here. Only `idle` can be acknowledged, so a permission or question dialog is
* unaffected by construction, and both are checked ahead of the flag anyway.
*/
export function classifySession(session: TuiSessionRow, approval?: ApprovalItem): TuiSessionState {
if (!isLiveRow(session)) return 'recent';
if (approval) {
if (approval.kind === 'permission') return 'blocked-permission';
if (approval.kind === 'question') return 'blocked-question';
return 'waiting';
if (!approval.acknowledgedAt) return 'waiting';
}
if (session.status === 'error') return 'waiting';
if (session.isWorking === true || session.status === 'busy') return 'working';
@@ -96,7 +108,11 @@ export function classifySession(session: TuiSessionRow, approval?: ApprovalItem)
* turn's own start is the pane's last Enter.
*/
export function stateSince(state: TuiSessionState, session: TuiSessionRow, approval?: ApprovalItem): number {
if (approval) return approval.createdAt;
// The prompt's own age measures the state only while the prompt is what put the
// row in that state. An acknowledged item still rides along on a row that is
// plainly idle or working, and dating such a row from it would report how long
// ago the prompt arrived as though it were how long the session has been quiet.
if (approval && STATE_GROUP[state] === 'needs-you') return approval.createdAt;
if (state === 'working') return session.lastSubmitAt ?? session.createdAt ?? 0;
return session.lastActivityAt ?? session.createdAt ?? 0;
}
+15 -5
View File
@@ -462,7 +462,8 @@ export function computeListWindow(
* The pending dialog, drawn above the tail: the question, the parsed options
* with their digits, and the keys that answer them. Red for a permission or
* question prompt, yellow for an idle one, the same severity vocabulary the web
* inbox uses.
* inbox uses. An idle prompt that opened acknowledged carries neither: it reads
* grey with the idle glyph, because nothing about it wants the reader.
*/
export function renderApprovalCard(
item: ApprovalItem,
@@ -472,8 +473,8 @@ export function renderApprovalCard(
): string[] {
const paint = painterFor(opts.color);
const card = approvalCard(item);
const color = card.tone === 'err' ? SGR.red : SGR.yellow;
const glyph = card.tone === 'err' ? glyphs.blockedPermission : glyphs.waiting;
const color = card.tone === 'err' ? SGR.red : card.tone === 'warn' ? SGR.yellow : SGR.gray;
const glyph = card.tone === 'err' ? glyphs.blockedPermission : card.tone === 'warn' ? glyphs.waiting : glyphs.idle;
const lines: string[] = [];
const push = (text: string, style: string): void => {
lines.push(padDisplay(paint(clipStyledLine(text, width), style), width));
@@ -571,10 +572,19 @@ function previewBody(
// Chrome
// ─────────────────────────────────────────────────────────────────────────────
/** Sessions with a prompt waiting on a human, which is what the badge counts. */
/**
* Sessions with a prompt waiting on a human, which is what the badge counts.
*
* An ACKNOWLEDGED item is not one of them. Its alert has been spent, either by somebody
* opening the session elsewhere or because the inbox opened it that way for a session
* watching its own background work, and the row has already left NEEDS YOU by the same
* flag (`classifySession`). Counting it here would put a number in the header for a
* group the reader can see is empty.
*/
export function pendingApprovalCount(model: TuiRenderModel): number {
let count = 0;
for (const group of model.groups()) for (const row of group.rows) if (row.approval) count++;
for (const group of model.groups())
for (const row of group.rows) if (row.approval && !row.approval.acknowledgedAt) count++;
return count;
}
+90
View File
@@ -115,6 +115,25 @@ export interface RemoteHost extends RemoteSshOptions {
username: string;
port?: number;
commands?: Partial<Record<RemoteCommandMode, string>>;
/**
* Optional Wake-on-LAN MAC address(es), comma-separated (e.g.
* `04:d9:f5:80:c6:58`). Codeman sends the magic packet itself (UDP port 9
* broadcast), so the common case needs no external script. A SLEEPING host's
* port-22 probe still fails, which is what triggers the wake — this only
* controls HOW the host is woken.
*/
wakeMac?: string;
/**
* Optional Wake-on-LAN command that powers this host on from SLEEP (e.g. a
* wrapper script like `/home/joe/bin/whuff`). TAKES PRECEDENCE over `wakeMac`
* (an explicit override for hosts that need a router/other-host wake). Absent
* = no wake support and today's behavior exactly. Executed WITHOUT a shell (a
* single executable path, never a command line), only from user input or an
* explicit wake request on a session whose host is unreachable — never from
* the auto-reconnect/boot-recovery path, which would re-wake a host seconds
* after each suspend.
*/
wakeCommand?: string;
}
export interface RemoteCase {
@@ -155,6 +174,13 @@ export interface SessionRemote extends RemoteSshOptions {
* session was created elsewhere. Only meaningful when `owned === false`.
*/
remoteSessionName?: string;
/**
* Wake-on-LAN command carried over from the host config (see `RemoteHost.wakeCommand`)
* so the input route can wake a sleeping host without re-reading the host list.
*/
wakeCommand?: string;
/** Wake-on-LAN MAC address(es) from the host config (see `RemoteHost.wakeMac`). */
wakeMac?: string;
}
/**
@@ -599,6 +625,55 @@ export interface CustomModelBookkeeping extends CustomModelSelection {
launchModel?: string;
}
/**
* The agent inside a LOCAL tmux pane has exited, and the pane survived it.
*
* Codeman creates every pane with `remain-on-exit on`, so `/exit` ends the CLI
* while tmux keeps the pane, the tmux session and the `tmux attach-session`
* process Codeman records as the session's pid. No PTY exit handler runs, so
* without this record the session reads as a live idle one (Ark0N/Codeman#446).
*
* The field is TRI-STATE, and the third state is the absence of the field:
* `undefined` means Codeman does not know, and it must never be rendered as
* "alive". It is absent for a direct-PTY session (no pane exists), for a remote
* SSH session (the local pane holds the ssh client, whose death means transport
* drop OR exit) and for a docker case (the local pane holds a `docker exec`
* into the container's own tmux).
*
* `status` and `signal` are independently optional because tmux may know that
* the pane died without reporting how. Measured on tmux 3.2a: a SIGKILLed pane
* reports `pane_dead=1` with BOTH `#{pane_dead_status}` and `#{pane_dead_signal}`
* empty, and `#{pane_dead_signal}` does not exist at all before tmux 3.4. So an
* absent `status` means "the exit code is unknown", never "the exit code is 0".
*
* ⚠ AN ABSENT `status` STAYS ABSENT. Never write `status ?? 0`, and never read
* "no signal was reported" as "the exit must have been clean". On tmux 3.2a
* the absent status IS how a signal death presents, so absent-stays-absent is
* the only thing keeping the clean-exit sweep (`pane-exit-sweep.ts`) away
* from crashed agents: an agent SIGKILLed by the OOM killer would otherwise
* read as a user typing `/exit` and be closed. Nothing here fails when
* somebody adds that `??` — the types allow it and the label still renders.
* The rule is enforced in `derivePaneExits()` (`tmux-manager.ts`), which omits
* the key rather than defaulting it, and again in `isCleanPaneExit()`, which
* accepts only an explicit 0.
*/
export interface PaneExit {
/**
* tmux `#{pane_dead_status}` — the command's exit code. Absent when tmux
* reported none, which means UNKNOWN and never 0. See the ⚠ above before
* giving this a default anywhere.
*/
status?: number;
/** tmux `#{pane_dead_signal}` — the signal that killed the command. Absent when unsignalled or unsupported. */
signal?: number;
/**
* Wall-clock ms when THIS server process first observed the pane dead. It is
* not when the agent exited, which nothing records, and a restart that finds
* the pane still dead respawns it rather than re-timing the old exit.
*/
at: number;
}
export interface SessionState {
/** Unique session identifier */
id: string;
@@ -754,6 +829,21 @@ export interface SessionState {
* (COD-118). Runtime-only: never restored on boot (fresh server = fresh breaker).
*/
respawnBlocked?: boolean;
/**
* The agent in this session's LOCAL tmux pane has exited (Ark0N/Codeman#446).
* See {@link PaneExit} for the tri-state rule and for which session shapes
* leave it absent. `status` and `pid` are deliberately untouched by it: the
* PTY-exit breaker owns `status: 'error'`, and a null `pid` is what makes the
* browser re-attach and launch a fresh CLI.
*
* Persisted so a reboot restore can tell a session whose agent exited from one
* that was merely idle when the power went. `reboot-restore.ts` reads the
* persisted record and never builds a `Session`, so the record is the only
* place that survives the reboot to carry it. Nothing reads it there YET:
* making the restore refuse such a session is a behavior change, and it
* belongs with the part of Ark0N/Codeman#446 that closes exited sessions.
*/
paneExit?: PaneExit;
}
/**
+75
View File
@@ -0,0 +1,75 @@
/**
* @fileoverview Does a Claude conversation transcript exist on this host?
*
* Claude writes one `<conversation-id>.jsonl` per conversation under
* `<config dir>/projects/<mangled cwd>/`. Two launch decisions turn on whether
* such a file exists: `--resume <id>` needs one, and `--session-id <id>` is
* REFUSED when one exists (`Error: Session ID ... is already in use.`).
*
* The project directory name is derived from the working directory, and a case
* that has been moved or renamed leaves its transcript under the OLD name, so
* the search is across every project directory rather than the one that matches
* the pane's cwd today.
*
* ⚠️ Existence is the whole question here, with no size floor. The create route
* additionally requires ~4 KB before it will resume, which is a "is this
* conversation worth resuming" judgement; for a relaunch the question is the
* opposite one — a one-line transcript still makes `--session-id` collide.
*
* ⚠️ A false answer is not the conservative one. Skipping a resume leaves the
* relaunch on `--session-id <id>`, which is safe only when no transcript backs
* that id either, so a lookup that misses the real config dir turns a
* recoverable pane into the collision this module exists to prevent.
*
* @dependencies none
* @consumedby session (relaunch resume pinning)
*
* @module utils/claude-transcript
*/
import { readdir, stat } from 'node:fs/promises';
import { homedir } from 'node:os';
import { join } from 'node:path';
/**
* `<config dir>/projects`, honouring a session's relocated `CLAUDE_CONFIG_DIR`
* (#255) and, failing that, the server process's own.
*
* ⚠️ The process env is not optional here. A pane inherits the server's
* environment through tmux, so on an install that exports `CLAUDE_CONFIG_DIR`
* the CLI writes its transcripts there and a lookup under `~/.claude` answers
* "no transcript" for every conversation on the host. `claudeCredentialsPath()`
* (claude-credentials.ts) and `realClaudeConfigDir()`
* (custom-model-injection-apply.ts) resolve the same directory the same way.
*/
export function claudeProjectsDir(configDir?: string): string {
const fromEnv = typeof process.env.CLAUDE_CONFIG_DIR === 'string' && process.env.CLAUDE_CONFIG_DIR.trim();
return join(configDir || fromEnv || join(homedir(), '.claude'), 'projects');
}
/**
* True when a transcript for `conversationId` exists under any project
* directory. Returns false for a missing projects dir or an unreadable one,
* which leaves the caller unpinned: safe where nothing else can collide with
* the bare `--session-id`, and the reason the caller walks its candidates down
* to the session's own id rather than treating one false answer as final.
*/
export async function claudeTranscriptExists(conversationId: string, configDir?: string): Promise<boolean> {
if (!conversationId) return false;
const projectsDir = claudeProjectsDir(configDir);
let projectDirs: string[];
try {
projectDirs = await readdir(projectsDir);
} catch {
return false;
}
for (const projectDir of projectDirs) {
try {
await stat(join(projectsDir, projectDir, `${conversationId}.jsonl`));
return true;
} catch {
// Not in this project directory; keep looking.
}
}
return false;
}
+31
View File
@@ -204,6 +204,28 @@ export function createProductionCliResolverHost(options: ProductionCliResolverHo
};
}
/**
* Per-binary invalidation generation, bumped by `invalidateCliExecutableResolvers()`.
*
* Every resolver instance (each per-CLI module's private one AND the generic registry
* resolver in cli-resolver.ts) is built by the factory below and caches in its own
* closure, so there is no instance to reach from outside. Keying on the BINARY name is
* what lets one call reach all of them: the CLI-management install/update routes know
* which binaries just changed, and every resolver knows its own.
*/
const binaryGenerations = new Map<string, number>();
/**
* Forget every cached result — success and negative-cache backoff alike — for these
* binaries, so the next `resolve()` re-runs the chain immediately. For an action that
* just changed what is on disk (an install) or what a CLI's binary IS (editing a custom
* entry): without it a CLI installed from Settings kept reading as missing for up to the
* 5-minute backoff, and an edited entry kept launching its old binary until a restart.
*/
export function invalidateCliExecutableResolvers(binaries: readonly string[]): void {
for (const binary of binaries) binaryGenerations.set(binary, (binaryGenerations.get(binary) ?? 0) + 1);
}
export function createCliExecutableResolver<T = undefined>(
options: {
binary: string;
@@ -235,6 +257,8 @@ export function createCliExecutableResolver<T = undefined>(
let failures = 0;
/** Timestamp of the most recent miss. */
let lastFailureAt = 0;
/** The invalidation generation the cached state above belongs to. */
let generation = binaryGenerations.get(options.binary) ?? 0;
const accept = (path: string | null, source: CliResolutionSource): CliResolution<T> | null => {
if (!path || !isAbsolute(path) || !host.exists(path)) return null;
const validation = options.validateCandidate?.(path) ?? ({ accepted: true } as CandidateValidation<T>);
@@ -249,6 +273,13 @@ export function createCliExecutableResolver<T = undefined>(
return {
resolve() {
const current = binaryGenerations.get(options.binary) ?? 0;
if (current !== generation) {
generation = current;
cached = null;
failures = 0;
lastFailureAt = 0;
}
if (cached) return cached;
// Negative cache: a miss is remembered and the chain — whose login-shell
// tail is a synchronous 5s-bounded spawn — is not re-run until the
+69
View File
@@ -0,0 +1,69 @@
/**
* @fileoverview One answer to "is this CLI installed here?", shared by the page render
* (`renderIndexHtml` in server.ts, which injects `window.__codemanCliAvailable` and
* `window.__codemanCliCatalog`) and `GET /api/clis` (the Settings list's badge).
*
* The two used to keep their own copies of the per-CLI probe map, so they could drift apart
* and the badge could disagree with the Run menu.
*
* Every probe is a memoized resolver, so this is cheap to call per request. Dynamic imports
* keep the nine resolvers out of any module that never asks.
*/
import type { CliEntry } from '../config/cli-registry/types.js';
import { isCliAvailable as isRegistryCliAvailable } from './cli-resolver.js';
/**
* The stock CLIs whose own resolver answers availability. It keeps the resolver's specific
* semantics (pi/grok/deepseek identity probes). DeepSeek reports RUNNABLE here, not merely
* installed: `dsh` is a profile launcher, and a dsh with no pane-capable profile would
* offer a Run button that spawns a pane which dies on arrival.
*/
export async function probeStockCliAvailability(): Promise<Record<string, boolean>> {
const [
{ isClaudeAvailable },
{ isOpenCodeAvailable },
{ isCodexAvailable },
{ isGeminiAvailable },
{ isAntigravityAvailable },
{ isPiAvailable },
{ isGrokAvailable },
{ isDeepSeekRunnable },
{ isOmpAvailable },
] = await Promise.all([
import('./claude-cli-resolver.js'),
import('./opencode-cli-resolver.js'),
import('./codex-cli-resolver.js'),
import('./gemini-cli-resolver.js'),
import('./antigravity-cli-resolver.js'),
import('./pi-cli-resolver.js'),
import('./grok-cli-resolver.js'),
import('./deepseek-cli-resolver.js'),
import('./omp-cli-resolver.js'),
]);
return {
claude: isClaudeAvailable(),
opencode: isOpenCodeAvailable(),
codex: isCodexAvailable(),
gemini: isGeminiAvailable(),
antigravity: isAntigravityAvailable(),
pi: isPiAvailable(),
grok: isGrokAvailable(),
deepseek: isDeepSeekRunnable(),
omp: isOmpAvailable(),
};
}
/**
* Is `entry` installed? A shell entry has no binary to probe, since it is the server's own
* login shell. A stock entry with a dedicated resolver uses `stockAvailability`. Anything
* else, custom entries included, uses the registry's GENERIC resolver. That is the one a
* session spawn uses, and it understands the entry's declared binaries and search dirs.
*/
export function isCliEntryInstalled(entry: CliEntry, stockAvailability: Record<string, boolean>): boolean {
if (entry.kind === 'shell') return true;
const id = entry.id as string;
return Object.prototype.hasOwnProperty.call(stockAvailability, id)
? stockAvailability[id]
: isRegistryCliAvailable(id);
}
+35
View File
@@ -71,6 +71,15 @@ export interface ApprovalItem {
* and reach the user's other devices. See `acknowledge()`.
*/
acknowledgedAt?: number;
/**
* Why the item arrived already acknowledged, for display only: the inbox
* writes `watching 1 monitor` for a session that went quiet because work it
* started itself is still running. A human acknowledgement leaves this unset,
* so a card can say "quiet, watching 1 monitor" rather than implying somebody
* looked. ⚠️ Pane-derived text, so it is bounded at the source and must not
* reach the DOM as markup — see `watchingLabel()` in `session-activity.ts`.
*/
acknowledgedReason?: string;
/**
* Present only when the frame parsed confidently. Gates which digits the
* answer endpoint accepts; absent → only approve('1')/deny(Esc) are allowed.
@@ -93,6 +102,12 @@ interface NotePromptArgs {
toolSummary?: string;
message?: string;
cwd?: string;
/**
* What the session's pane says is still running in the background
* (`Session.watching`, e.g. `1 monitor`). An idle prompt from such a session
* opens ALREADY acknowledged: see `notePrompt()`.
*/
watching?: string | null;
/** Returns the raw (ANSI-bearing) pane frame, or null when unavailable. */
capture?: () => string | null;
}
@@ -217,6 +232,22 @@ export class ApprovalInbox {
* Record a prompt for a session, superseding any previous item, and return
* the new item. Captures context immediately and once more after a short
* delay (see RECAPTURE_DELAY_MS).
*
* ⚠️ An idle prompt from a session that is WATCHING its own background work
* opens already acknowledged (`args.watching`). Claude Code ends the turn
* after arming a monitor or backgrounding a shell and then reports the pane
* idle a minute later, so the alert that follows asks a human to look at a
* session that wants nothing from them. Acknowledging is deliberately what
* happens here rather than skipping the item: the prompt is real and stays
* pending, answerable and available as Read My Mind context, and only the
* alert it would have armed is spent. A wrong label therefore costs a card
* that does not blink, never an alert that was never created.
*
* It re-arms by itself. The next idle prompt supersedes this item and builds
* a fresh one, so once the background work ends and the session goes quiet
* for an ordinary reason, that item carries no acknowledgement and alerts
* normally. Only `idle` is eligible: a permission or question dialog blocks
* the agent whatever else it started, so its alert must survive.
*/
notePrompt(args: NotePromptArgs): ApprovalItem {
this.resolveForSession(args.sessionId, 'superseded');
@@ -231,6 +262,10 @@ export class ApprovalInbox {
message: args.message,
cwd: args.cwd,
};
if (args.kind === 'idle' && args.watching) {
item.acknowledgedAt = item.createdAt;
item.acknowledgedReason = `watching ${args.watching}`;
}
this.applyCapture(item, args.capture);
this.items.set(args.sessionId, item);
if (args.capture) this.captures.set(args.sessionId, args.capture);
+69 -1
View File
@@ -12,7 +12,8 @@
* image dir.
*/
import fs from 'node:fs/promises';
import { join } from 'node:path';
import { realpathSync } from 'node:fs';
import { join, resolve } from 'node:path';
import type { SessionPort } from './ports/index.js';
const MAX_AGE_MS = 7 * 24 * 60 * 60 * 1000; // 7 days
@@ -53,6 +54,73 @@ export async function sweepPasteImagesOnce(
return { scanned, deleted };
}
/**
* The path two sessions must share to share a paste-image dir: the canonical
* path when it can be resolved, so a sibling that reaches the same directory
* through a symlink matches, and the normalised path otherwise (a directory
* that no longer exists has nothing left to protect).
*/
function canonicalDir(dir: string): string {
try {
return realpathSync(dir);
} catch {
return resolve(dir);
}
}
/** One session the paste-image guard weighs: its id, directory and, for a persisted record, its status. */
export interface PasteImageDirUser {
id: string;
workingDir: string;
status?: string;
}
/**
* Does another live session still use this working directory's paste-image
* dir? Deleting a session removes `{workingDir}/.claude-images` recursively,
* and several sessions routinely share one case directory, so without this
* check closing one session deletes the pasted images a sibling in the same
* case still refers to.
*
* Two kinds of sibling count as live:
*
* - a session in the server's map, unless it is itself being killed;
* - a persisted record whose status is not `stopped`. That covers a session
* detached with `killMux=false`, which leaves the server's map while its
* tmux pane keeps running, and a session whose detach is still in progress.
*
* A session being KILLED does not count. Without that exemption, killing two
* sessions of one case concurrently (a bulk delete, or the exited-agent sweep
* closing two panes on one tick) would have each defer to the other, and
* neither would remove the dir.
*
* Erring toward "in use" only costs a missed deletion, which the periodic
* sweep above ages out. A pinned record whose tmux session is gone keeps its
* status through boot pruning, so it holds the dir this way until unpinned.
*/
export function pasteImageDirInUseByOtherSession(input: {
live: Iterable<PasteImageDirUser>;
persisted: Iterable<PasteImageDirUser>;
closingId: string;
workingDir: string;
killing: ReadonlySet<string>;
}): boolean {
const target = canonicalDir(input.workingDir);
const matches = (user: PasteImageDirUser): boolean =>
user.id !== input.closingId &&
!input.killing.has(user.id) &&
!!user.workingDir &&
canonicalDir(user.workingDir) === target;
for (const user of input.live) {
if (matches(user)) return true;
}
for (const user of input.persisted) {
if (user.status === 'stopped') continue;
if (matches(user)) return true;
}
return false;
}
export function startPasteImageGc(ctx: Pick<SessionPort, 'sessions'>): () => void {
const initial = setTimeout(() => {
void sweepPasteImagesOnce(ctx);
+767 -91
View File
File diff suppressed because it is too large Load Diff
+26 -1
View File
@@ -194,8 +194,21 @@ Object.assign(CodemanApp.prototype, {
document.querySelector('.btn-approvals')?.setAttribute('aria-expanded', 'false');
},
/**
* Items still waiting on a human. An acknowledged item (a human already looked, or the
* session opened it acknowledged because it is watching its own background work) keeps
* its card but arms no alert, so it must not light the bell either; this is the same
* count `pendingApprovalCount()` gives `codeman tui`.
*/
pendingApprovalsCount() {
if (!this.approvals) return 0;
let count = 0;
for (const item of this.approvals.values()) if (!item.acknowledgedAt) count++;
return count;
},
renderApprovals() {
const count = this.approvals ? this.approvals.size : 0;
const count = this.pendingApprovalsCount();
const btn = document.querySelector('.btn-approvals');
if (btn) {
// Marker-class visibility (base header rules are display !important):
@@ -222,6 +235,15 @@ Object.assign(CodemanApp.prototype, {
return;
}
list.innerHTML = items.map((item) => this._approvalCardHtml(item)).join('');
// The quiet reason is the only pane-derived string on a card, and it is the one
// an agent could write itself (it prints its own footer row), so it reaches the
// DOM as text and never as markup. The card leaves an empty span for it.
for (const item of items) {
if (!item.acknowledgedReason) continue;
const card = list.querySelector(`[data-approval-id="${CSS.escape(item.id)}"]`);
const slot = card && card.querySelector('.approval-quiet');
if (slot) slot.textContent = 'quiet, ' + item.acknowledgedReason;
}
},
_approvalCardHtml(item) {
@@ -260,6 +282,9 @@ Object.assign(CodemanApp.prototype, {
`<span class="approval-session" data-i18n-skip>${escapeHtml(item.sessionName || item.sessionId.slice(0, 8))}</span>` +
`<span class="approval-age" data-i18n-skip>${age}</span>` +
`</div>` +
// Filled by renderApprovalsDrawer through textContent, never here: see the note
// there. An item a human acknowledged carries no reason and gets no line.
(item.acknowledgedReason ? `<div class="approval-quiet" data-i18n-skip></div>` : '') +
(summary ? `<div class="approval-summary" data-i18n-skip>${escapeHtml(summary)}</div>` : '') +
(item.context ? `<pre class="approval-context">${escapeHtml(item.context)}</pre>` : '') +
`<div class="approval-actions">${actions}</div>` +
+435
View File
@@ -762,6 +762,49 @@ function resolveTerminalFontWeights(settings) {
// without a terminal, a clipboard, or a browser.
// ---------------------------------------------------------------------------
/**
* Largest single input frame the server accepts, in UTF-16 code units.
* ⚠️ Must equal MAX_INPUT_LENGTH in src/config/terminal-limits.ts (pinned by
* test/input-size-limit.test.ts). Both transports reject a longer frame, and
* before issue #484 the durable input queue retried such a frame forever.
*/
const INPUT_FRAME_MAX_CHARS = 64 * 1024;
/**
* Largest paste the client will deliver at all. Anything up to this is split
* into INPUT_FRAME_MAX_CHARS frames that go out in seq order, so the PTY sees
* one contiguous byte stream (bracketed-paste markers included). Past it the
* input is refused with a toast rather than queued: every frame is persisted
* and retried until ACKed, so a multi-megabyte paste would pin the queue.
*/
const INPUT_PASTE_MAX_CHARS = 1024 * 1024;
/**
* Split input into frames no longer than `max` code units, never cutting a
* surrogate pair in half (a lone surrogate reaches the PTY as U+FFFD).
*
* @param {string} data
* @param {number} [max]
* @returns {string[]}
*/
function splitInputFrames(data, max = INPUT_FRAME_MAX_CHARS) {
if (typeof data !== 'string' || data.length === 0) return [];
if (!(max >= 2)) max = 2;
if (data.length <= max) return [data];
const frames = [];
let start = 0;
while (start < data.length) {
let end = Math.min(start + max, data.length);
if (end < data.length) {
const code = data.charCodeAt(end - 1);
if (code >= 0xd800 && code <= 0xdbff) end--; // keep the pair together
}
frames.push(data.slice(start, end));
start = end;
}
return frames;
}
/**
* Upper bound on an AUTO-copied selection.
*
@@ -806,6 +849,96 @@ function decideAutoCopy({ enabled, text, lastCopied, pending } = {}) {
return 'copy';
}
// The text a copy should put on the clipboard, given xterm's raw selection.
// Pure: the caller reads the selection and decides the mode, this transforms.
//
// xterm hands back whole screen ROWS, and its own trim only drops cells that
// were never written to. A full-screen TUI writes real spaces across the part
// of a row it is not using, so that padding counts as content and rides along
// to the clipboard: measured against Claude Code in a 282-column pane, single
// lines arrived carrying 138 trailing spaces. Native terminals trim it on copy
// (Windows Terminal, iTerm2 and GNOME Terminal all do), decideAutoCopy above
// already calls a wall of spaces "never what the gesture meant", and
// _selectTouchSelectionLine already treats those cells as padding. This is that
// same rule for the mouse and keyboard paths, which never had it.
//
// A LEADING margin is stripped too, but only the one the CLI in the pane
// DECLARES as its transcript gutter, passed in as `options.margin`. Called with
// no options this trims trailing padding and nothing else, which is what keeps
// every caller that has no declared gutter on the old behaviour.
//
// ⚠ The failure modes are not symmetrical, and that asymmetry sets how much
// evidence a leading strip has to show before it fires. A wrong trailing trim
// costs nothing. A wrong dedent silently deletes information that was on the
// screen, with no signal to the user and nothing in the clipboard to hint at
// it, and it is wrong on `git log` bodies, on indented code read out of `cat`
// (semantic in Python), on `git diff` context rows where the leading space is
// the marker, and on stack traces.
//
// ⚠ The declared gutter is a CEILING, not the answer. The strip is the lesser
// of it and the run every selected line shares, so a block can only ever shift
// as a unit: the relative structure inside a selection survives by
// construction, and a selection reaching column 0 loses nothing at all.
//
// ⚠ Deriving the width from the text instead is what fails, twice over. The
// selection's own shared indent cannot tell a margin from content, because a
// three-row window of nested YAML shares an indent for the same reason a margin
// does — it fired on 73% of ordinary indented text. Taking the narrowest indent
// on the surrounding rows fails more quietly: a file listing inside the
// transcript can be the narrowest thing on screen, which over-stripped about 1%
// of selections across six pane widths.
function cleanCopiedSelection(text, options) {
if (typeof text !== 'string' || !text) return '';
// Split on \n and leave any \r in place: xterm joins rows with \r\n on
// Windows, and the clipboard should keep the endings xterm chose.
// Scanned rather than matched. A selection can run to the 50 000-row
// scrollback ceiling, and `/[ \t]+(\r?)$/` is QUADRATIC on a line whose spaces
// are followed by any non-space character, which is what right-aligned or
// centred TUI content looks like: the engine retries the run from every
// whitespace position and backtracks over it. Measured over 50 000 rows with a
// 280-column run, that regex took 2.9s against 1.3ms for the scan below, and a
// 2 000-column run took 16s. It is also the faster of the two on an ordinary
// padded row. A length is returned rather than a trimmed string so a
// \r-terminated line costs no substring either.
const trimEnd = (line) => {
let end = line.length;
if (end > 0 && line[end - 1] === '\r') end--;
let cut = end;
while (cut > 0 && (line[cut - 1] === ' ' || line[cut - 1] === '\t')) cut--;
return cut === end ? line : line.slice(0, cut) + line.slice(end);
};
const lines = text.split('\n');
for (let i = 0; i < lines.length; i++) lines[i] = trimEnd(lines[i]);
const margin = Math.max(0, Math.trunc(Number(options?.margin) || 0));
if (!margin) return lines.join('\n');
// The first line of a selection that began mid-row carries no margin — the
// mousedown cut it off — so it neither votes on the shared indent nor gets
// stripped. This is the ONE thing the mousedown column still decides, and it
// decides it for that line alone. Whether the rest of the block is dedented
// no longer depends on where the click landed, which is what made the same
// three rows produce three different clipboard results before.
const from = options?.firstLinePartial === true ? 1 : 0;
// The pane's margin is a ceiling, not the answer. Strip the narrower of it
// and what every selected line shares, so the block shifts as a unit and no
// line can lose indentation another line keeps.
let shared = margin;
for (let i = from; i < lines.length && shared > 0; i++) {
const line = lines[i];
if (!line || line === '\r') continue; // a padding-only row, already trimmed away
let run = 0;
while (run < line.length && line[run] === ' ') run++;
if (run < shared) shared = run;
}
if (!shared) return lines.join('\n');
for (let i = from; i < lines.length; i++) {
if (lines[i] && lines[i] !== '\r') lines[i] = lines[i].slice(shared);
}
return lines.join('\n');
}
if (typeof window !== 'undefined') {
window.WEBGL_FALLBACK = WEBGL_FALLBACK;
window.evaluateWebGLLongTaskTrip = evaluateWebGLLongTaskTrip;
@@ -850,10 +983,18 @@ if (typeof window !== 'undefined') {
compare: compareSessionActivity,
sort: sortSessionsByActivity,
};
window.CodemanInputLimit = {
FRAME_MAX_CHARS: INPUT_FRAME_MAX_CHARS,
PASTE_MAX_CHARS: INPUT_PASTE_MAX_CHARS,
split: splitInputFrames,
};
window.CodemanAutoCopy = {
decide: decideAutoCopy,
MAX_CHARS: AUTO_COPY_MAX_CHARS,
};
window.CodemanCopySelection = {
clean: cleanCopiedSelection,
};
window.CodemanTerminalFont = {
DEFAULT_STACK: TERMINAL_FONT_DEFAULT_STACK,
resolve: resolveTerminalFontFamily,
@@ -1057,6 +1198,9 @@ const SSE_EVENTS = {
REMOTE_SESSION_DROPPED: 'remote:sessionDropped',
REMOTE_SESSION_RECONNECTED: 'remote:sessionReconnected',
REMOTE_RECONNECT_EXHAUSTED: 'remote:reconnectExhausted',
// Wake-on-LAN from user input on a sleeping remote host
REMOTE_HOST_WAKING: 'remote:hostWaking',
REMOTE_HOST_WAKE_FAILED: 'remote:hostWakeFailed',
// Ralph
SESSION_RALPH_LOOP_UPDATE: 'session:ralphLoopUpdate',
@@ -1094,6 +1238,9 @@ const SSE_EVENTS = {
APPROVAL_UPDATED: 'approval:updated',
APPROVAL_RESOLVED: 'approval:resolved',
// Custom Model Endpoint Profiles
CUSTOM_MODEL_SWAPPED_OUT: 'custom-model:swapped-out',
// Subagents (Claude Code background agents)
SUBAGENT_DISCOVERED: 'subagent:discovered',
SUBAGENT_UPDATED: 'subagent:updated',
@@ -1442,8 +1589,296 @@ function terminalLogicalLine(buffer, row, cols, maxRows) {
return { startRow, endRow, text, offsetToCell, cellToOffset };
}
// ═══════════════════════════════════════════════════════════════
// Split-Pane Sessions — pure helpers (divider math, picker list)
// ═══════════════════════════════════════════════════════════════
// Desktop-only, same reasoning and same threshold as HOME_SESSIONS_MIN_WIDTH
// (home-sessions.js): two 240px min-width panes plus the divider need ~486px,
// which a phone or narrow tablet cannot give them, and the divider has no
// touch handlers. A dedicated constant rather than reusing
// HOME_SESSIONS_MIN_WIDTH directly — that name lives in home-sessions.js,
// which loads AFTER this file (load order 12.56 vs 7.5), so referencing it
// from module-evaluation-time code here would be a ReferenceError.
const SPLIT_PANE_MIN_WIDTH = 1180;
function clampDividerPercent(rawPercent, min = 20, max = 80) {
if (rawPercent < min) return min;
if (rawPercent > max) return max;
return rawPercent;
}
function buildSplitPickerSessions(sessions, sessionOrder, excludeId, detachedIds) {
const result = [];
for (const id of sessionOrder) {
if (id === excludeId) continue;
// A detached (popped-out) session's own window already yields its PTY
// size (see sendResize's detachedElsewhere guard in terminal-ui.js) —
// Pane B's SplitTerminalPane._sendResize() has no such check, so letting
// one into the picker put its detached window and Pane B in a fight over
// the same PTY's dimensions.
if (detachedIds?.has?.(id)) continue;
const session = sessions.get(id);
if (!session) continue;
// A session with no PTY attached (exited CLI, a crash-looped session
// whose breaker tripped, a restore that failed to re-attach) has nothing
// reading its tmux pane. SplitTerminalPane never does selectSession()'s
// re-attach POST, so its socket would open onto a pane nothing feeds:
// no terminal events, and Session.write() silently drops every keystroke
// with no ack either way (Pane B sends no `seq`), so the loss is
// invisible — the healthy socket never trips the disconnect banner.
if (session.pid === null) continue;
result.push({ id, label: session.name || 'Session' });
}
return result;
}
// ── Renderer liveness ──────────────────────────────────────────────────────
//
// iOS DISCARDS scheduled requestAnimationFrame callbacks when a PWA goes to
// the background — not deferred, never delivered. xterm's RenderDebouncer only
// clears its `_animationFrame` handle from INSIDE that callback:
//
// refresh() {
// if (this._animationFrame !== undefined) return; // <- stale forever
// this._animationFrame = requestAnimationFrame(() => this._innerRefresh());
// }
// _innerRefresh() { this._animationFrame = undefined; ... } // never runs
//
// So after one backgrounding the handle is permanently non-undefined and EVERY
// later render request returns on line one. Parsing is decoupled from
// rendering, so bytes keep filling the buffer correctly and nothing throws —
// the terminal is simply frozen. Closing and reopening fixes it because that
// constructs a new Terminal, and therefore a new debouncer.
//
// Codeman is MORE exposed than a per-session-terminal app: there is exactly one
// xterm instance for the whole page load, so a single backgrounding can wedge
// it until a full reload.
//
// This is the pure decision half. The signature that distinguishes this from
// every other way a terminal can look stuck is that bytes were WRITTEN and the
// element is VISIBLE, yet onRender has not fired since:
//
// frozen = wroteAt > renderedAt && now - wroteAt >= threshold && visible
//
// Deliberately NOT a "no output at all" check: a quiet terminal is the normal
// state and must never be kicked. And `visible` is required because a hidden
// terminal legitimately stops rendering (xterm pauses it), so kicking there
// would fire constantly on every backgrounded tab.
const RENDER_STALL_MS = 4000;
// How often the watchdog checks. Deliberately coarse: the failure it catches is
// permanent until healed, so detecting it a second late costs nothing, while a
// tight interval would burn a wakeup per second on every idle phone.
const RENDER_LIVENESS_POLL_MS = 2000;
/**
* Should the renderer be kicked? Pure so the CI gate can cover it — the DOM
* half (cancelling the stale handle) lives in terminal-ui.js.
*
* @param {{wroteAt:number, renderedAt:number, now:number, visible:boolean,
* thresholdMs?:number}} s
* @returns {boolean}
*/
function shouldKickRenderer(s) {
if (!s || !s.visible) return false;
const wroteAt = Number(s.wroteAt) || 0;
const renderedAt = Number(s.renderedAt) || 0;
const now = Number(s.now) || 0;
// Nothing written yet — a fresh terminal has no render to be missing.
if (wroteAt <= 0) return false;
// A render landed at or after the last write: the pipeline is alive.
if (renderedAt >= wroteAt) return false;
const threshold = Number.isFinite(s.thresholdMs) && s.thresholdMs > 0 ? s.thresholdMs : RENDER_STALL_MS;
return now - wroteAt >= threshold;
}
// ── Fetch deadlines ────────────────────────────────────────────────────────
//
// No terminal fetch carried any deadline, including `?full=1`, which the code
// itself describes as "unbounded-ish work: at the default history limit it can
// be megabytes". On a stalled mobile link that request hangs on the browser
// default with no retry and no path back to a usable terminal short of a
// reload.
//
// A single fixed timeout is wrong in both directions — too short for a full
// scrollback capture on a slow uplink, too long for a small tail on a dead
// connection. So the deadline is scaled by what is actually being asked for,
// and by how many captures are already in flight: on a slow link those bytes
// must drain before this request's own bytes start moving, and its timer is
// already running the whole time.
const FETCH_DEADLINE_TAIL_MS = 15000;
const FETCH_DEADLINE_FULL_MS = 45000;
const FETCH_DEADLINE_MAX_MS = 120000;
/**
* Deadline in ms for a terminal capture.
*
* @param {{full?:boolean, inflight?:number}} s - `full` = the ?full=1 capture;
* `inflight` = captures already running (this one included or not, it only
* scales the budget).
* @returns {number}
*/
function terminalFetchDeadlineMs(s) {
const full = !!(s && s.full);
const base = full ? FETCH_DEADLINE_FULL_MS : FETCH_DEADLINE_TAIL_MS;
const inflight = Math.max(0, Number(s && s.inflight) || 0);
// Each already-queued capture gets the newcomer one more base budget to wait
// through. Linear rather than clever: the point is only that eight tabs
// resuming do not all time out together because each assumed it was alone.
return Math.min(FETCH_DEADLINE_MAX_MS, base * (1 + inflight));
}
// ── Diagnostics hygiene ────────────────────────────────────────────────────
//
// The crash trail is joined with '\n' into ONE localStorage value and beaconed
// to the server, and at least one call site interpolates server-controlled text
// (a WebSocket close `reason`). An embedded newline there forges extra entries
// in the trail; an unbounded string can fill the storage quota. Both are cheap
// to close, and the trail is something a user may be asked to paste into an
// issue.
const DIAG_ENTRY_MAX_CHARS = 300;
/** Flatten a diagnostic message to one bounded, newline-free line. */
function sanitizeDiagEntry(msg) {
return String(msg == null ? '' : msg)
.replace(/[\r\n\u2028\u2029]+/g, ' ')
.slice(0, DIAG_ENTRY_MAX_CHARS);
}
// ── Recovering a dropped output frame ──────────────────────────────────────
//
// `_onSessionTerminal` drops an incoming frame when the app-owned render queues
// already hold 128KB, which is the right call — the alternative is an unbounded
// backlog — but a hole in a TUI byte stream is a desynced cursor, and a desynced
// cursor is muffled text (issue #464). So the drop is only half of it: the
// recovery has to actually happen.
//
// ⚠️ It used to be a fire-and-forget timer. `_onSessionNeedsRefresh` opens with
// four early returns, and two of them — a buffer load in flight, a refresh
// already owning this session — are MOST likely to be true during exactly the
// output burst that caused the drop. The timer nulled itself before the call,
// so a skipped refresh lost the recovery silently and the dropped bytes were
// never replayed.
//
// Bounded, because the early returns it retries past are transient contention
// that clears in seconds, and a permanently failing refresh must not become a
// forever-loop against the API. A refresh that hit the capture fetch DEADLINE
// is not contention but a stalled link, and is not retried at all: each retry
// would be another `?full=1` capture waiting out a deadline of up to two
// minutes, where the old code cost exactly one. Giving up after the cap leaves
// exactly the garbled frames the old code left, so the floor is no worse.
const DROP_RECOVERY_DELAY_MS = 2000;
const DROP_RECOVERY_MAX_ATTEMPTS = 5;
/**
* Should a dropped-output recovery run again?
*
* @param {{repainted: boolean, timedOut?: boolean, attempt: number, stillActive: boolean}} state
* `repainted` — whether `_onSessionNeedsRefresh` actually rewrote the buffer.
* `timedOut` - whether it failed at the capture fetch deadline.
* `attempt` — how many have already run, zero-based.
* `stillActive` — whether the dropped session is still the one on screen.
* @returns {boolean}
*/
function shouldRetryDroppedOutputRecovery({ repainted, timedOut = false, attempt, stillActive }) {
// Switched away: `selectSession` repaints from the server on its own, so a
// retry here would be a second replay of a buffer that is about to be written.
if (!stillActive) return false;
if (repainted) return false;
if (timedOut) return false;
return attempt + 1 < DROP_RECOVERY_MAX_ATTEMPTS;
}
// ── Terminal geometry: xterm and the PTY must never disagree ───────────────
//
// Issue #464 ("text gets muffled"). Claude Code's TUI repaints by wrapping its
// frame at the width the PTY reported and walking the cursor up that many
// ROWS. So a browser terminal whose width differs from the PTY's makes every
// repaint arithmetic wrong: a logical line occupies more physical rows than
// Ink counted, `eraseLines(n)` clears too few of them, and the new frame paints
// over rows that were never erased. Measured against a real xterm — a PTY
// believing 120 columns against a 62-column terminal renders each wrapped line
// twice, and a shorter replacement line leaves the tail of the old one behind.
// That is exactly the doubled rows and half-overwritten prose in the report.
//
// The floor exists because a PTY a handful of columns wide makes any CLI wrap
// every word; it is NOT a display preference, so the browser terminal has to
// honour it too. Three separate call sites used to fit xterm to the RAW
// proposal and report the CLAMPED one, which is how the two drifted apart with
// nothing to notice: resize is write-only, so nobody could see the disagreement.
const TERMINAL_MIN_COLS = 40;
const TERMINAL_MIN_ROWS = 10;
/**
* The geometry to apply AND report — there is only ever one answer to both.
* @param {{cols: number, rows: number}|null|undefined} proposed
* @returns {{cols: number, rows: number}|null}
*/
function clampTerminalDimensions(proposed) {
if (!proposed || !Number.isFinite(proposed.cols) || !Number.isFinite(proposed.rows)) return null;
return {
cols: Math.max(Math.trunc(proposed.cols), TERMINAL_MIN_COLS),
rows: Math.max(Math.trunc(proposed.rows), TERMINAL_MIN_ROWS),
};
}
/**
* What to do when the server reports the PTY's real geometry.
*
* The server is the authority: it owns the PTY the CLI is drawing for, and it
* can refuse a resize outright (`Session.resize` ignores small-viewport
* requests while a desktop connection holds an active sizing claim) without
* the asking client ever being told. A terminal that keeps its own WIDTH after
* such a refusal renders garbage, because Ink wraps its frame and counts its
* erase rows at the width it was told.
*
* ⚠️ COLUMNS ONLY. Rows are deliberately left alone, and adopting them was a
* real regression: a phone that took a desktop's 43 rows into a viewport with
* room for 18 painted an `.xterm-screen` far taller than its container, and
* because xterm's own viewport then had nothing to scroll, the bottom of the
* frame — the CLI's input line — sat below the container with no gesture that
* could reach it. Output visible, typing invisible, for as long as the claim
* stayed hot. Width is the axis the wrap arithmetic depends on; rows only
* decide how much is on screen at once, and keeping the local row count keeps
* the composer at the bottom of a viewport that scrolls.
*
* @param {{cols: number, rows: number}|null} local - what xterm currently holds
* @param {{cols: number, rows: number}|null} pty - what the server just reported
* @returns {{adopt: boolean, cols: number|null}}
*/
function reconcilePtyGeometry(local, pty) {
if (!pty || !Number.isFinite(pty.cols)) return { adopt: false, cols: null };
if (!local || !Number.isFinite(local.cols) || local.cols === pty.cols) return { adopt: false, cols: null };
return { adopt: true, cols: pty.cols };
}
if (typeof window !== 'undefined') {
window.CodemanHistoryFormat = { formatHistoryBytes, computeHistoryTruncationNotice, computeRewriteScrollLine };
window.CodemanFilePaths = { absoluteFilePathPattern, previewsInFileViewer, FILE_PREVIEW_EXTENSIONS };
window.CodemanTerminalLines = { terminalLogicalLine };
window.CodemanSplitPane = {
clampDividerPercent,
buildSplitPickerSessions,
SPLIT_PANE_MIN_WIDTH,
};
window.CodemanRenderLiveness = { shouldKickRenderer, RENDER_STALL_MS, RENDER_LIVENESS_POLL_MS };
window.CodemanFetchDeadline = {
terminalFetchDeadlineMs,
FETCH_DEADLINE_TAIL_MS,
FETCH_DEADLINE_FULL_MS,
FETCH_DEADLINE_MAX_MS,
};
window.CodemanDiag = { sanitizeDiagEntry, DIAG_ENTRY_MAX_CHARS };
window.CodemanDroppedOutput = {
shouldRetryDroppedOutputRecovery,
DROP_RECOVERY_DELAY_MS,
DROP_RECOVERY_MAX_ATTEMPTS,
};
window.CodemanTerminalGeometry = {
clampTerminalDimensions,
reconcilePtyGeometry,
TERMINAL_MIN_COLS,
TERMINAL_MIN_ROWS,
};
}
+20 -5
View File
@@ -201,6 +201,8 @@ Object.assign(CodemanApp.prototype, {
const session = this.sessions.get(id);
const matched = this._mobileOverviewCaseFor(session.workingDir, cases);
const state = this._mobileOverviewState(session, this.pendingHooks?.get(id));
// Guarded: a stale cached mobile-overview.js may predate the helper.
const exit = this._mobileOverviewExit ? this._mobileOverviewExit(state, session) : null;
const mode = session.mode || 'claude';
return {
id,
@@ -211,7 +213,13 @@ Object.assign(CodemanApp.prototype, {
caseName: matched ? matched.name : '',
dir: this._shortenHomePath ? this._shortenHomePath(session.workingDir) : session.workingDir || '',
state,
pill: HOME_SESSIONS_PILL_LABEL[state] || state,
// What the row's dot, accent and pill show. It differs from `state` only
// for an exited agent (Ark0N/Codeman#446), whose state still sorts it.
display: exit ? 'exited' : state,
pill: exit ? 'exited' : HOME_SESSIONS_PILL_LABEL[state] || state,
// What the pane's footer says is still running in the background, straight off
// the session payload. Same field, same meaning as on the phone overview.
watching: typeof session.watching === 'string' ? session.watching : '',
// Epoch ms, straight off the session payload; formatting happens at
// render time so the clock below can redo it without a re-render.
createdAt: Number(session.createdAt) || 0,
@@ -221,7 +229,7 @@ Object.assign(CodemanApp.prototype, {
lastSubmitAt: Number(session.lastSubmitAt) || 0,
// "how long has it been like this", resolved by the phone overview's
// helper so both home screens label the same stamp with the same word.
since: this._mobileOverviewSince(state, session),
since: exit ? exit.since : this._mobileOverviewSince(state, session),
};
});
@@ -389,7 +397,8 @@ Object.assign(CodemanApp.prototype, {
_buildHomeSessionRow(row) {
const item = document.createElement('button');
item.type = 'button';
item.className = 'home-sessions-row home-sessions-row--' + row.state;
const display = row.display || row.state;
item.className = 'home-sessions-row home-sessions-row--' + display;
item.dataset.hsAction = 'session';
item.dataset.hsSession = row.id;
item.title = row.dir ? `${row.name} (${row.dir})` : row.name;
@@ -406,7 +415,7 @@ Object.assign(CodemanApp.prototype, {
}
const dot = document.createElement('span');
dot.className = 'home-sessions-dot home-sessions-dot--' + row.state;
dot.className = 'home-sessions-dot home-sessions-dot--' + display;
dot.setAttribute('aria-hidden', 'true');
item.appendChild(dot);
@@ -438,7 +447,7 @@ Object.assign(CodemanApp.prototype, {
item.appendChild(body);
const pill = document.createElement('span');
pill.className = 'home-sessions-pill home-sessions-pill--' + row.state;
pill.className = 'home-sessions-pill home-sessions-pill--' + display;
// Skipped by i18n on purpose: generic single words ("idle", "done", "error")
// that collide with state strings on other surfaces.
pill.setAttribute('data-i18n-skip', '');
@@ -450,6 +459,12 @@ Object.assign(CodemanApp.prototype, {
// what stops it ellipsizing.
const meta = this._buildHomeSessionsMeta(row);
meta.appendChild(pill);
// Built by the phone overview so both home screens word the badge identically.
// Guarded like every other cross-file call here: a stale cached mobile-overview.js
// must cost the badge, not the rail.
if (row.watching && typeof this._buildWatchingBadge === 'function') {
meta.appendChild(this._buildWatchingBadge(row.watching, 'home-sessions-pill'));
}
item.appendChild(meta);
return item;
+440
View File
@@ -0,0 +1,440 @@
/**
* @fileoverview Remote-host wake-on-LAN: the "host unreachable" banner + its config dialog.
*
* A sleeping remote host does not fail loudly. The local tmux pane runs `ssh`, and when
* the machine suspends, that ssh child stalls: `tmux send-keys` still SUCCEEDS, so typed
* input disappears with no error and the pane looks alive. The server side
* (`src/remote-wake.ts`) buffers input and wakes the host when the user types; this
* module makes the state VISIBLE and gives it a button, which is what turns "why is
* nothing happening" into one click.
*
* Behavior:
* - Asks `GET /api/sessions/:id/reachability` for the ACTIVE remote session only:
* once when the tab is activated (a user action), and every `POLL_MS` while the tab
* is visible ONLY for a host with a wake target. The timer is the one thing here that
* is not user-driven, and each poll is a TCP connect to the host — the same
* timer-driven traffic invariant #2 rejects keepalives for: it cannot wake a host,
* but it can keep an activity-based suspend timer from firing. So a host Codeman
* could not wake anyway is never polled on a timer. A host behind a jump host or
* SOCKS proxy (`probeable: false`) is never polled at all: the probe cannot reach
* it, so its answer would only ever be a false "asleep". The endpoint shares the
* server's probe cache with the input path, so opening the tab also primes the
* wake path.
* - Unreachable + a configured wake target → "Wake" button → `POST /api/sessions/:id/wake`
* (which wakes, waits, reattaches the pane and flushes buffered input).
* - Unreachable + NO wake target → "Configure WoL" → `#wakeConfigModal`, a small form
* for this host's MAC/command that saves via `PUT /api/remote-hosts/:id`. The server
* re-resolves host config while the session is live, so saving takes effect without
* restarting the session.
* - SSE (`remote:hostWaking`, `remote:hostWakeFailed`, `remote:sessionReconnected`)
* keeps the banner in sync while a wake is running.
*
* @mixin Extends CodemanApp.prototype via Object.assign
* @dependency app.js (CodemanApp class, this.sessions, this.activeSessionId, showToast)
* @dependency constants.js (SSE_EVENTS — the remote:hostWaking / remote:hostWakeFailed names)
* @loadorder 12.2 — loaded after session-ui.js, before webview-tabs.js
*/
const HOST_WAKE_POLL_MS = 30_000;
Object.assign(CodemanApp.prototype, {
/** Per-tab banner state (single active session at a time). */
_hostWake: null,
/** The page-wide poller interval (created once, see `_ensureHostWakePoller`). */
_hostWakeTimer: null,
/** Fresh state for a session we just switched to. */
_hostWakeState() {
return {
sessionId: null,
/** Last reachability answer, or null before the first poll. */
reachable: null,
/** 'command' | 'mac' | 'none' — what the banner action should do. */
wakeConfigured: 'none',
host: '',
label: '',
/**
* False for a host the server's probe cannot reach (behind a jump host or SOCKS
* proxy): its reachability is unknown, so there is no banner and no polling.
*/
probeable: true,
/** True between clicking Wake and the answer coming back. */
waking: false,
/**
* True only when the server is actually holding bytes for this session (the typing
* path buffers them). Browser keystrokes go over the WebSocket, which never passes
* through the wake registry — so the Wake BUTTON must not claim input is queued.
*/
queuedInput: false,
/** Set when the last wake attempt or poll failed. */
error: '',
};
},
/**
* Entry point from the session switcher — called for every active session, remote or
* not, so it must be cheap and must clear the banner for local sessions.
*
* ⚠️ The POLLER is page-wide and independent of this call on purpose: a session
* switch is not the only way the active tab changes (boot restore, a page loaded with
* the tab already active, and `selectSession`'s own early return for the tab you are
* already on), and the banner must not depend on any single one of those paths
* running — that is exactly how it could silently never appear.
*/
refreshHostWakeBanner(sessionId) {
this._ensureHostWakePoller();
const state = this._hostWake;
if (state && state.sessionId && state.sessionId !== sessionId) this._hostWake = null;
this._hostWakeTick();
},
/** Create the page-wide poller once (interval + a visibility wake-up). */
_ensureHostWakePoller() {
if (this._hostWakeTimer) return;
this._hostWakeTimer = setInterval(() => this._hostWakeTick({ periodic: true }), HOST_WAKE_POLL_MS);
document.addEventListener('visibilitychange', () => {
if (document.visibilityState === 'visible') this._hostWakeTick({ periodic: true });
});
},
/**
* One poller tick: resolve the ACTIVE session, reset the banner when it changed, and
* ask the server. No-op while the page is hidden (a background tab must not poll).
*
* `periodic` marks the timer (and the visibility wake-up) as opposed to a tab
* activation: a periodic tick polls only a host with a wake target, see the module
* comment. The activation poll is what still offers "Configure WoL" for a sleeping
* host that has none — one connect, on a user action.
*/
_hostWakeTick({ periodic = false } = {}) {
if (typeof document !== 'undefined' && document.visibilityState === 'hidden') return;
const sessionId = this.activeSessionId;
const session = sessionId && this.sessions ? this.sessions.get(sessionId) : null;
if (!sessionId || !session || !session.remote) {
// Render unconditionally: `refreshHostWakeBanner` clears `_hostWake` BEFORE
// calling this tick, so a guard here would skip the repaint and leave the
// banner up on every chat (the clear and the repaint must not be coupled to
// whoever cleared the state). Idempotent — with a null state it just hides.
this._hostWake = null;
this._renderHostWakeBanner();
return;
}
let state = this._hostWake;
let fresh = false;
if (!state || state.sessionId !== sessionId) {
fresh = true;
state = this._hostWake = this._hostWakeState();
state.sessionId = sessionId;
state.host = session.remote.host || '';
state.label = session.remote.label || 'Remote host';
// Text from the session payload first (instant, no round trip), corrected by the
// poll — a session whose wake config was added after launch only knows it after
// the server resolves host config. The kind matters: the payload can say WHICH
// path is configured, so a command-only host is not mislabelled 'mac' until the
// first poll lands.
state.wakeConfigured = session.remote.wakeMac ? 'mac' : session.remote.wakeCommand ? 'command' : 'none';
// Known from the payload already: a proxied host is not probeable (the server
// says so too, on every answer), so not even the activation poll is worth a
// round trip whose verdict could only be a wrong "asleep".
state.probeable = !(session.remote.jumpHost || session.remote.socksProxy);
this._renderHostWakeBanner();
}
if (!state.probeable) return;
if (periodic && !fresh && state.wakeConfigured === 'none') return;
this._pollHostReachability();
},
/** One reachability check for the active remote session. */
async _pollHostReachability(force = false) {
const state = this._hostWake;
if (!state || !state.sessionId) return;
const sessionId = state.sessionId;
try {
const res = await fetch(`/api/sessions/${encodeURIComponent(sessionId)}/reachability${force ? '?force=1' : ''}`);
const data = await res.json();
if (!data.success) return;
// The tab may have changed while this was in flight.
if (this._hostWake !== state || state.sessionId !== sessionId) return;
// `reachable` is `null` (unknown, not unreachable) for a host the probe cannot
// reach — only a PROVEN `false` may raise the banner.
state.reachable = data.data.reachable !== false;
if (data.data.probeable === false) state.probeable = false;
state.wakeConfigured = data.data.wakeConfigured || 'none';
if (data.data.host) state.host = data.data.host;
if (data.data.label) state.label = data.data.label;
if (state.reachable) {
state.waking = false;
state.error = '';
}
this._renderHostWakeBanner();
} catch {
/* A failed poll is not a state change: leave the banner as it was. */
}
},
/** Draw the banner from `_hostWake`. */
_renderHostWakeBanner() {
const state = this._hostWake;
const banner = this.$('hostWakeBanner');
const text = this.$('hostWakeBannerText');
const detail = this.$('hostWakeBannerDetail');
const action = this.$('hostWakeBannerAction');
if (!banner || !text || !action) return;
const visible = Boolean(state && state.sessionId && state.reachable === false);
banner.hidden = !visible;
if (!visible) return;
const hasTarget = state.wakeConfigured !== 'none';
const target = state.label || state.host || 'Remote host';
if (state.waking) {
text.textContent = `Waking ${target} …`;
} else if (state.error) {
text.textContent = `${target} did not wake up`;
} else {
text.textContent = `${target} is not reachable`;
}
if (detail) {
detail.textContent = state.waking
? state.queuedInput
? 'input is queued until it is back'
: 'waiting for the host to come back'
: hasTarget
? `ssh ${state.host}`
: 'no wake-on-LAN configured';
}
// After a FAILED wake the only useful next step is fixing the target (wrong MAC,
// host moved NIC, command gone) — otherwise a configured-but-broken host would be
// stuck behind a button that keeps failing with no way to edit it.
const offerConfig = !hasTarget || Boolean(state.error);
action.textContent = state.waking ? 'Waking …' : offerConfig ? 'Configure WoL' : 'Wake';
action.disabled = state.waking;
},
/** Banner button: wake the host, or open the setup dialog when nothing is configured. */
hostWakeAction() {
const state = this._hostWake;
if (!state || !state.sessionId || state.waking) return;
if (state.wakeConfigured === 'none' || state.error) {
this.openWakeConfigDialog();
return;
}
this.wakeRemoteHost();
},
/** POST the manual wake for the active session and follow the result. */
async wakeRemoteHost() {
const state = this._hostWake;
if (!state || !state.sessionId) return;
const sessionId = state.sessionId;
state.waking = true;
// The button path holds nothing: whatever the user typed went into the stalled pane
// over the WebSocket and is gone. Saying otherwise is a promise the next keystroke
// disproves.
state.queuedInput = false;
state.error = '';
this._renderHostWakeBanner();
try {
const res = await fetch(`/api/sessions/${encodeURIComponent(sessionId)}/wake`, { method: 'POST' });
const data = await res.json();
if (this._hostWake !== state || state.sessionId !== sessionId) return;
state.waking = false;
if (!data.success) {
// The ROUTE is the authority on whether a target is configured, so ask it again
// (`/reachability` reports `wakeConfigured`) rather than pattern-matching the
// error message: the message is prose, and the code is generic (`INVALID_INPUT`
// covers "Not a remote session" too).
state.error = data.error || 'Wake failed';
this._renderHostWakeBanner();
await this._pollHostReachability(true);
return;
}
state.reachable = data.data.reachable !== false;
state.wakeConfigured = data.data.wakeConfigured || state.wakeConfigured;
if (state.reachable) {
this.showToast(`${state.label || 'Remote host'} is awake`, 'success');
} else {
state.error = 'timeout';
}
this._renderHostWakeBanner();
} catch (err) {
if (this._hostWake !== state) return;
state.waking = false;
state.error = err && err.message ? err.message : 'Wake failed';
this._renderHostWakeBanner();
}
},
/**
* Why the host could not be read. In multi-user mode `GET /api/remote-hosts` returns
* `[]` to a non-admin, so "Remote host not found" would blame a config the user simply
* is not allowed to see — the save is admin-only, and that is what it should say.
*/
_wakeConfigUnavailableMessage() {
const me = window.__codemanUser || {};
return me.multiUser && me.role !== 'admin' ? 'Wake-on-LAN configuration is admin-only' : 'Remote host not found';
},
/** Open the small WoL dialog for the banner's host, pre-filled from the host config. */
async openWakeConfigDialog() {
const state = this._hostWake;
const session = state && state.sessionId && this.sessions ? this.sessions.get(state.sessionId) : null;
if (!session || !session.remote) return;
const hostId = session.remote.hostId;
const label = this.$('wakeConfigHostLabel');
const mac = this.$('wakeConfigMac');
const command = this.$('wakeConfigCommand');
const status = this.$('wakeConfigStatus');
if (!mac || !command) return;
mac.value = session.remote.wakeMac || '';
command.value = session.remote.wakeCommand || '';
if (label) label.textContent = session.remote.label || hostId;
if (status) status.textContent = '';
this._wakeConfigHostId = hostId;
const modal = this.$('wakeConfigModal');
if (modal) modal.classList.add('active');
// Read the saved host so the dialog shows what is actually persisted (the session
// payload may predate a change made in another tab).
try {
const res = await fetch('/api/remote-hosts');
const data = await res.json();
const hosts = data.success ? data.data : [];
const host = Array.isArray(hosts) ? hosts.find((item) => item.id === hostId) : null;
if (host && this._wakeConfigHostId === hostId) {
mac.value = host.wakeMac || '';
command.value = host.wakeCommand || '';
} else if (!host && this._wakeConfigHostId === hostId && status) {
// Say it up front rather than only when Save fails.
status.textContent = this._wakeConfigUnavailableMessage();
}
} catch {
/* The form is already usable from the session payload. */
}
},
closeWakeConfigDialog() {
const modal = this.$('wakeConfigModal');
if (modal) modal.classList.remove('active');
this._wakeConfigHostId = null;
},
/** Save MAC/command for the host, then re-check whether the session can wake now. */
async saveWakeConfig() {
const hostId = this._wakeConfigHostId;
const mac = this.$('wakeConfigMac');
const command = this.$('wakeConfigCommand');
const status = this.$('wakeConfigStatus');
const save = this.$('wakeConfigSave');
if (!hostId || !mac || !command) return;
const macValue = mac.value.trim();
const commandValue = command.value.trim();
if (
macValue &&
!/^[0-9a-fA-F]{2}([:-][0-9a-fA-F]{2}){5}(\s*,\s*[0-9a-fA-F]{2}([:-][0-9a-fA-F]{2}){5})*$/.test(macValue)
) {
if (status) status.textContent = 'MAC must look like 04:d9:f5:80:c6:58 (comma-separated for several).';
return;
}
if (commandValue && /\s/.test(commandValue)) {
if (status) status.textContent = 'The wake command must be a single executable path (no arguments).';
return;
}
if (save) save.disabled = true;
if (status) status.textContent = 'Saving …';
try {
const listRes = await fetch('/api/remote-hosts');
const listData = await listRes.json();
const hosts = listData.success ? listData.data : [];
const host = Array.isArray(hosts) ? hosts.find((item) => item.id === hostId) : null;
if (!host) throw new Error(this._wakeConfigUnavailableMessage());
// PUT takes the whole host (schema-validated), so send back everything we know and
// only replace the wake fields. `undefined` drops the key entirely.
const payload = {
...host,
wakeMac: macValue || undefined,
wakeCommand: commandValue || undefined,
};
const res = await fetch(`/api/remote-hosts/${encodeURIComponent(hostId)}`, {
method: 'PUT',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(payload),
});
const data = await res.json();
if (!data.success) throw new Error(data.error || 'Save failed');
this.showToast('Wake settings saved', 'success');
this.closeWakeConfigDialog();
// The server re-resolves host config for live sessions, so the banner can offer
// the wake right away — probe fresh instead of waiting out the poll interval.
await this._pollHostReachability(true);
} catch (err) {
if (status) status.textContent = err && err.message ? err.message : 'Save failed';
} finally {
if (save) save.disabled = false;
}
},
/**
* SSE `remote:hostWaking` — a wake is running (ours or one started by typing).
*
* ⚠️ The ONLY definition of this handler: `panels-ui.js` must not define it too.
* Both mix into `Codeman.prototype` and this file loads later, so a second copy
* would be silently shadowed (the guard in `sse-dispatch-table.test.ts` sees that a
* handler exists, not that two modules claim the same name). The toast is
* deliberately UNCONDITIONAL — a wake can start for a background session (input on
* a non-active tab) where there is no banner to update.
*/
_onRemoteHostWaking(data) {
const label = data && data.label ? data.label : 'Remote host';
// A create-path wake (the user pressed Run / Attach) has no session yet, so
// nothing is queued behind it — the wording has to say what actually happens.
const forNewSession = Boolean(data && data.forNewSession);
// Only the typing path buffers bytes; the wake button and the send-and-wait path
// hold none, and a browser keystroke never reaches the registry at all.
const queuedInput = Boolean(data && data.queuedInput);
// Long enough to cover the wake + attach (~10s measured on a warm S3), and it
// is replaced by `remote:sessionReconnected` the moment the pane is back.
this.showToast(
forNewSession
? `Waking ${label} … the session starts when it is back`
: queuedInput
? `Waking ${label} … input is queued`
: `Waking ${label} … waiting for it to come back`,
'info',
{ duration: 12000 }
);
const state = this._hostWake;
if (!state || !data || state.sessionId !== data.sessionId) return;
state.waking = true;
state.queuedInput = queuedInput;
state.error = '';
if (data.label) state.label = data.label;
this._renderHostWakeBanner();
},
/** SSE `remote:hostWakeFailed` — the host did not come back in time. */
_onRemoteHostWakeFailed(data) {
const label = data && data.label ? data.label : 'Remote host';
const forNewSession = Boolean(data && data.forNewSession);
const queuedInput = Boolean(data && data.queuedInput);
this.showToast(
forNewSession
? `${label} did not wake up — no session was started`
: queuedInput
? `${label} did not wake up — queued input is still held`
: `${label} did not wake up`,
'error',
{ duration: 15000 }
);
const state = this._hostWake;
if (!state || !data || state.sessionId !== data.sessionId) return;
state.waking = false;
state.queuedInput = queuedInput;
state.error = 'timeout';
state.reachable = false;
this._renderHostWakeBanner();
},
});
+50
View File
@@ -73,6 +73,10 @@
'File Viewer': '文件查看器',
'Open file viewer': '打开文件查看器',
'Open Codeman across all displays': '在所有显示器上打开 {name}',
'Split: open a second session beside this one': '分屏:在旁边打开第二个会话',
'Split: close the second session': '分屏:关闭第二个会话',
'Close split': '关闭分屏',
'No other sessions to split with': '没有其他可用于分屏的会话',
'Ultracode / Workflow agents': 'Ultracode / Workflow 智能体',
'Open ultracode workflow agents': '打开 Ultracode 工作流智能体',
Notifications: '通知',
@@ -102,16 +106,20 @@
'Manage AI Coding tools in persistent tmux sessions.': '在持久化 tmux 会话中管理 AI 编程工具。',
'Select case': '选择案例',
'Select Case': '选择案例',
'Search cases': '搜索案例',
'No matching cases': '没有匹配的案例',
'All cases': '全部案例',
'No directory': '未选择目录',
Run: '运行',
'Run Claude Code': '运行 Claude Code',
'Run OpenCode': '运行 OpenCode',
'Run Codex': '运行 Codex',
'Run Gemini': '运行 Gemini',
'Run Antigravity': '运行 Antigravity',
'Run Pi': '运行 Pi',
'Run Grok': '运行 Grok',
'Run DeepSeek': '运行 DeepSeek',
'Run OMP': '运行 OMP',
'Run Shell': '运行 Shell',
'Select AI backend': '选择 AI 后端',
'Create New Case': '新建案例',
@@ -286,6 +294,36 @@
'Prompt sent': '提示已发送',
'Inserted, press Enter in the terminal to send': '已插入,在终端中按 Enter 发送',
'Could not reach the session': '无法连接到会话',
'Custom model endpoints': '自定义模型端点',
'Point a harness at your own OpenAI-compatible server (llama.cpp, vLLM, DGX Spark, Azure AI Foundry, OpenRouter) instead of its native cloud backend. When on, the Run menu offers an extra entry per harness that supports it, per saved endpoint.':
'让工具指向您自己的兼容 OpenAI 服务器(llama.cpp、vLLM、DGX Spark、Azure AI Foundry、OpenRouter),而非其原生云端后端。开启后,"运行"菜单会为每个支持此功能的工具、每个已保存的端点新增一个条目。',
'Enable custom model endpoints': '启用自定义模型端点',
'Adds a per-endpoint entry to the Run menu for every harness that can redirect to one.':
'为每个可重定向到端点的工具,在"运行"菜单中添加对应条目。',
'No endpoints yet. Add one below to point a harness at a local or cloud OpenAI-compatible server.':
'暂无端点。请在下方添加一个,以便将工具指向本地或云端的兼容 OpenAI 服务器。',
Discover: '发现模型',
'+ Add endpoint': '+ 添加端点',
'Add endpoint': '添加端点',
Id: 'ID',
'Short, stable — used in URLs, never shown to the CLI.': '简短且固定 — 用于 URL,不会展示给 CLI。',
Label: '标签',
'Base URL': '基础 URL',
'API key': 'API 密钥',
'Optional. Left blank on edit keeps the existing key.': '可选。编辑时留空将保留现有密钥。',
'Auth header': '认证请求头',
'Never send both — some servers hang indefinitely.': '切勿同时发送两者 — 部分服务器会因此无限期挂起。',
'Authorization: Bearer (default)': 'Authorization: Bearer(默认)',
'api-key header (Azure)': 'api-key 请求头(Azure)',
'Default model': '默认模型',
'What the Run-menu picker applies for this endpoint. Discover models first.':
'运行菜单选择器会为此端点应用该模型。请先发现可用模型。',
'Custom Endpoints': '自定义端点',
'Choose a model': '选择模型',
'Currently loaded': '当前已加载',
'Last used': '上次使用',
'That endpoint no longer exists': '该端点已不存在',
'No models discovered for this endpoint yet': '此端点尚未发现任何模型',
'Subagent Options': '子智能体选项',
'Enable Tracking': '启用跟踪',
'Active Tab Only': '仅活动标签页',
@@ -417,6 +455,17 @@
'Show Shortcuts': '显示快捷键',
'Full shortcut reference': '完整快捷键参考',
// Mobile prompt composer (keyboard-accessory.js). The textarea's own
// placeholder and label are looked up by the module at build time, since
// the DOM translator skips <textarea> subtrees.
'Compose prompt': '撰写提示词',
'Compose prompt, draft saved': '撰写提示词,草稿已保存',
'Resume saved prompt draft': '继续编辑已保存的提示词草稿',
'Enter adds a new line': '按 Enter 换行',
'Write your prompt…': '请输入提示词…',
'Use terminal keyboard': '使用终端键盘',
'Uploading…': '上传中…',
// Mobile overview (phone home screen)
'Needs you': '需要你',
'Current sessions': '当前会话',
@@ -521,6 +570,7 @@
'Respawn Blocked': '重生已阻止',
'Task Complete': '任务完成',
'Copied to clipboard': '已复制到剪贴板',
'Nothing to copy': '没有可复制的内容',
// Terminal touch-selection bar (long-press to select). The bar is a sibling of
// `.xterm`, not a descendant, so SKIP_SELECTOR does not cover it and these apply.
Copy: '复制',
+8 -4
View File
@@ -124,12 +124,15 @@ Object.assign(CodemanApp.prototype, {
// 20 photos don't crawl through serially.
_uploadConcurrency: 3,
async _uploadAndInsertImages(fileList) {
/** Upload a batch and normally insert its paths into the active terminal.
* The prompt composer passes `{ insert: false }` so it can put those paths
* into its textarea instead. Returns successful paths in selection order. */
async _uploadAndInsertImages(fileList, options = {}) {
const sessionId = this.activeSessionId;
if (!sessionId) return;
if (!sessionId) return [];
let files = Array.from(fileList || []);
if (files.length === 0) return;
if (files.length === 0) return [];
// Cap the batch and tell the user what got dropped (no silent truncation).
let capped = false;
@@ -175,7 +178,7 @@ Object.assign(CodemanApp.prototype, {
await Promise.all(Array.from({ length: Math.min(this._uploadConcurrency, total) }, () => worker()));
const paths = results.filter(Boolean);
if (paths.length > 0) {
if (paths.length > 0 && options.insert !== false) {
// Insert all paths in one shot, space-separated, in selection order.
await this.sendInput(paths.join(' '));
}
@@ -187,6 +190,7 @@ Object.assign(CodemanApp.prototype, {
if (capped) parts.push(`max ${this._maxBatchImages} per batch`);
const tone = paths.length > 0 ? (failed > 0 || capped ? 'info' : 'success') : 'error';
this.showToast(parts.join(' · ') || 'No images uploaded', tone);
return paths;
},
async _uploadPasteImage(sessionId, file) {
+260 -60
View File
@@ -191,6 +191,7 @@
</button>
<button class="btn-icon-header btn-file-viewer" onclick="app.toggleFileBrowserButton()" title="File Viewer" aria-label="Open file viewer" aria-expanded="false"><svg width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M3 7a2 2 0 0 1 2-2h4l2 2h8a2 2 0 0 1 2 2v8a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2z"/></svg></button>
<button class="btn-icon-header btn-multimonitor btn-multimonitor--hidden" onclick="app.launchMultiMonitor()" title="Open Codeman across all displays" aria-label="Open Codeman across all displays"><svg width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><rect x="2" y="4" width="13" height="9" rx="1.5"/><rect x="11" y="9" width="11" height="8" rx="1.5"/></svg></button>
<button class="btn-icon-header btn-split btn-split--hidden" onclick="app.openSplitPicker(event)" title="Split: open a second session beside this one" aria-label="Split: open a second session beside this one" aria-pressed="false"><svg width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><rect x="2" y="3" width="20" height="18" rx="2"/><line x1="12" y1="3" x2="12" y2="21"/></svg></button>
<button class="btn-icon-header btn-ultracode-agents btn-ultracode-agents--hidden" onclick="app.toggleUltracodeAgentsPanel()" title="Ultracode / Workflow agents" aria-label="Open ultracode workflow agents"><svg width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><circle cx="6" cy="6" r="2.5"/><circle cx="6" cy="18" r="2.5"/><circle cx="18" cy="12" r="2.5"/><path d="M8.2 7.2 15.6 11M8.2 16.8 15.6 13"/></svg></button>
<div class="header-plan-usage header-plan-usage--hidden" id="planUsageChip" title="Claude and Codex plan usage limits">—</div>
<button class="btn-icon-header btn-notifications" onclick="app.toggleNotifications()" title="Notifications" aria-label="Toggle notifications" style="display:none;">
@@ -213,6 +214,18 @@
<button class="offline-banner-retry" id="offlineBannerRetry" onclick="app.retryConnection()">Retry now</button>
</div>
<!-- Remote-host unreachable: the machine SLEEPS, the local ssh pane stalls
silently (send-keys succeeds against it, so typed input would vanish) and
Codeman can wake it. Amber, not red: the session is fine, the host is
asleep. Without a configured wake target the action becomes "Configure
WoL" and opens the small config dialog. -->
<div class="offline-banner host-wake-banner" id="hostWakeBanner" role="status" hidden>
<span class="offline-banner-dot" aria-hidden="true"></span>
<span class="offline-banner-text" id="hostWakeBannerText">Remote host is unreachable</span>
<span class="offline-banner-detail" id="hostWakeBannerDetail"></span>
<button class="offline-banner-retry" id="hostWakeBannerAction" onclick="app.hostWakeAction()">Wake</button>
</div>
<!-- Reboot-restore offer: shown when the server found sessions a host reboot
killed and is asking whether to rebuild them. Populated by
reboot-restore-ui.js; nothing is created until the user clicks. -->
@@ -438,42 +451,11 @@
<h1 class="welcome-title">Codeman</h1>
<p class="welcome-desc">Manage AI Coding tools in persistent tmux sessions.</p>
<div class="welcome-actions">
<button class="welcome-btn welcome-btn-claude" id="welcomeClaudeBtn" style="display: none;" onclick="app.setRunMode('claude'); app.runClaude()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run Claude Code
</button>
<div class="welcome-cli-actions" id="welcomeCliActions"></div>
<button class="welcome-btn welcome-btn-tunnel" id="welcomeTunnelBtn" style="display: none;" onclick="app.toggleTunnelFromWelcome()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path d="M12 2L2 7l10 5 10-5-10-5z"/><path d="M2 17l10 5 10-5"/><path d="M2 12l10 5 10-5"/></svg>
Cloudflare Tunnel
</button>
<button class="welcome-btn welcome-btn-opencode" id="welcomeOpencodeBtn" style="display: none;" onclick="app.setRunMode('opencode'); app.runOpenCode()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run OpenCode
</button>
<button class="welcome-btn welcome-btn-antigravity" id="welcomeAntigravityBtn" style="display: none;" onclick="app.setRunMode('antigravity'); app.runAntigravity()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run Antigravity
</button>
<button class="welcome-btn welcome-btn-gemini" id="welcomeGeminiBtn" style="display: none;" onclick="app.setRunMode('gemini'); app.runGemini()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run Gemini
</button>
<button class="welcome-btn welcome-btn-pi" id="welcomePiBtn" style="display: none;" onclick="app.setRunMode('pi'); app.runPi()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run Pi
</button>
<button class="welcome-btn welcome-btn-grok" id="welcomeGrokBtn" style="display: none;" onclick="app.setRunMode('grok'); app.runGrok()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run Grok
</button>
<button class="welcome-btn welcome-btn-deepseek" id="welcomeDeepSeekBtn" style="display: none;" onclick="app.setRunMode('deepseek'); app.runDeepSeek()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run DeepSeek
</button>
<button class="welcome-btn welcome-btn-omp" id="welcomeOmpBtn" style="display: none;" onclick="app.setRunMode('omp'); app.runOmp()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run OMP
</button>
</div>
<div class="welcome-qr" id="welcomeQr" onclick="app.toggleWelcomeQrSize()">
<div class="welcome-qr-inner" id="welcomeQrInner"></div>
@@ -635,39 +617,21 @@
<svg width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path d="M6 9l6 6 6-6"/></svg>
</button>
<div class="run-mode-menu" id="runModeMenu">
<button class="run-mode-option" data-mode="claude" onclick="app.setRunMode('claude')">
<span class="run-mode-dot claude"></span>Claude Code
</button>
<button class="run-mode-option" data-mode="opencode" onclick="app.setRunMode('opencode')">
<span class="run-mode-dot opencode"></span>OpenCode
</button>
<button class="run-mode-option" data-mode="codex" onclick="app.setRunMode('codex')">
<span class="run-mode-dot codex"></span>Codex
</button>
<button class="run-mode-option" data-mode="gemini" onclick="app.setRunMode('gemini')">
<span class="run-mode-dot gemini"></span>Gemini
</button>
<button class="run-mode-option" data-mode="antigravity" onclick="app.setRunMode('antigravity')">
<span class="run-mode-dot antigravity"></span>Antigravity
</button>
<button class="run-mode-option" data-mode="pi" onclick="app.setRunMode('pi')">
<span class="run-mode-dot pi"></span>Pi
</button>
<button class="run-mode-option" data-mode="grok" onclick="app.setRunMode('grok')">
<span class="run-mode-dot grok"></span>Grok
</button>
<button class="run-mode-option" data-mode="deepseek" onclick="app.setRunMode('deepseek')">
<span class="run-mode-dot deepseek"></span>DeepSeek
</button>
<div class="run-mode-cli-options" id="runModeCliOptions"></div>
<!-- Shown only when `dsh` is installed but no pane-capable profile is:
DeepSeek ships no terminal front door, so the fix is an install,
not a greyed-out entry the user cannot act on. -->
<button class="run-mode-option run-mode-option-install" data-action="deepseek-install" id="runModeDeepSeekInstall" style="display: none;" onclick="app.installDeepSeekProfile()">
<span class="run-mode-dot deepseek"></span>DeepSeek — add a terminal profile…
</button>
<button class="run-mode-option" data-mode="omp" onclick="app.setRunMode('omp')">
<span class="run-mode-dot omp"></span>OMP
</button>
<!-- Custom Model Endpoint Profiles (docs/custom-model-endpoints-plan.md): one
generated entry per (harness, saved endpoint) pair, e.g. "Claude Code
(llama.cpp)". Built entirely by _refreshCustomModelRunOptions() — hidden
when the feature is off or no endpoint has a usable default model, never
a fixed per-harness duplicate in this markup. -->
<div class="run-mode-sep" id="runModeCustomModelSep" style="display: none;"></div>
<div class="run-mode-header" id="runModeCustomModelHeader" style="display: none;">Custom Endpoints</div>
<div class="run-mode-custom-models" id="runModeCustomModels"></div>
<div class="run-mode-sep"></div>
<button class="run-mode-option" data-mode="shell" onclick="app.setRunMode('shell')">
<span class="run-mode-dot shell"></span>Terminal / Shell
@@ -920,6 +884,67 @@
</div>
</div>
<!-- Custom Model Endpoint Profiles: "which model" picker (docs/custom-model-endpoints-plan.md).
Shown only when the chosen endpoint has more than one discovered model — see
selectCustomModelEntry() in session-ui.js, which skips straight to launch otherwise. -->
<div class="modal" id="customModelPickModal">
<div class="modal-backdrop" onclick="app.closeCustomModelPickModal()"></div>
<div class="modal-content modal-sm">
<div class="modal-header">
<h3 id="customModelPickTitle">Choose a model</h3>
<button class="modal-close" onclick="app.closeCustomModelPickModal()" aria-label="Close model picker">&times;</button>
</div>
<div class="modal-body">
<p class="form-hint" id="customModelPickHint"></p>
<div id="customModelPickList" class="run-mode-custom-models"></div>
</div>
</div>
</div>
<!-- Custom Model Endpoint Profiles: llama-swap model-swap confirmation
(docs/custom-model-endpoints-plan.md) — replaces a native confirm()
popup, shown when switching would unload a model another live
session is actively using. See _confirmModelSwap() in session-ui.js. -->
<div class="modal" id="customModelSwapConfirmModal">
<div class="modal-backdrop" onclick="app._resolveModelSwapConfirm(false)"></div>
<div class="modal-content modal-sm">
<div class="modal-header">
<h3>Switch models?</h3>
<button class="modal-close" onclick="app._resolveModelSwapConfirm(false)" aria-label="Cancel">&times;</button>
</div>
<div class="modal-body">
<p class="form-hint" id="customModelSwapConfirmMessage"></p>
</div>
<div class="modal-footer">
<button class="btn-toolbar" onclick="app._resolveModelSwapConfirm(false)">Cancel</button>
<button class="btn-toolbar btn-primary" onclick="app._resolveModelSwapConfirm(true)">Switch anyway</button>
</div>
</div>
</div>
<!-- Custom Model Endpoint Profiles: context-window-too-small warning
(docs/custom-model-endpoints-plan.md) — shown before launching a CLI
whose own fixed system-prompt/tool-schema overhead exceeds the
model's real discovered context, which guarantees a first-message
failure regardless of CLAUDE_CODE_MAX_CONTEXT_TOKENS. See
_confirmContextWarning() in session-ui.js. -->
<div class="modal" id="customModelContextWarningModal">
<div class="modal-backdrop" onclick="app._resolveContextWarningConfirm(false)"></div>
<div class="modal-content modal-sm">
<div class="modal-header">
<h3>Context window too small</h3>
<button class="modal-close" onclick="app._resolveContextWarningConfirm(false)" aria-label="Cancel">&times;</button>
</div>
<div class="modal-body">
<p class="form-hint" id="customModelContextWarningMessage" style="white-space: pre-wrap;"></p>
</div>
<div class="modal-footer">
<button class="btn-toolbar" onclick="app._resolveContextWarningConfirm(false)">Cancel</button>
<button class="btn-toolbar btn-primary" onclick="app._resolveContextWarningConfirm(true)">Launch anyway</button>
</div>
</div>
</div>
<!-- Cron Jobs Modal -->
<div class="modal" id="cronModal">
<div class="modal-backdrop" onclick="app.closeCron()"></div>
@@ -1709,6 +1734,13 @@
</div>
<label class="switch switch-sm"><input type="checkbox" id="appSettingsAutoCopySelection"><span class="slider"></span></label>
</div>
<div class="set-row" data-search="copy indent margin gutter dedent trim leading whitespace paste">
<div class="set-row-text">
<span class="set-row-label">Trim the pane margin on copy</span>
<span class="set-row-desc">Drop the left margin a full-screen agent CLI paints down its own edge, so copied text pastes flush instead of indented. Each CLI declares its own width, and the strip is never wider than the indent every selected line shares, so nesting inside the selection is kept. Claude Code and Codex declare a margin today; a shell, and any CLI that declares none, is left alone.</span>
</div>
<label class="switch switch-sm"><input type="checkbox" id="appSettingsCopyStripMargin"><span class="slider"></span></label>
</div>
</div>
</div>
@@ -1843,6 +1875,7 @@
<label class="set-chip" data-preview="header" data-preview-order="9"><input type="checkbox" id="appSettingsShowAttachmentsButton"><svg class="set-chip-ico" width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.9" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="m21.44 11.05-9.19 9.19a6 6 0 0 1-8.49-8.49l9.19-9.19a4 4 0 0 1 5.66 5.66l-9.2 9.19a2 2 0 0 1-2.83-2.83l8.49-8.48"/></svg><span>Attachments</span></label>
<label class="set-chip" data-preview="header" data-preview-order="10"><input type="checkbox" id="appSettingsShowFileViewerButton"><svg class="set-chip-ico" width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.9" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M3 7a2 2 0 0 1 2-2h4l2 2h8a2 2 0 0 1 2 2v8a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2z"/></svg><span>File Viewer</span></label>
<label class="set-chip" data-preview="header" data-preview-order="11"><input type="checkbox" id="appSettingsShowMultiMonitorButton"><svg class="set-chip-ico" width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.9" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><rect x="2" y="4" width="13" height="9" rx="1.5"/><rect x="11" y="9" width="11" height="8" rx="1.5"/></svg><span>Multi-monitor</span></label>
<label class="set-chip" data-preview="header" data-preview-order="11.5"><input type="checkbox" id="appSettingsShowSplitButton"><svg class="set-chip-ico" width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.9" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><rect x="2" y="3" width="20" height="18" rx="2"/><line x1="12" y1="3" x2="12" y2="21"/></svg><span>Split</span></label>
<label class="set-chip" data-preview="header" data-preview-order="13" data-preview-text="42%"><input type="checkbox" id="appSettingsShowPlanUsageLimits"><svg class="set-chip-ico" width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.9" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M4 18a8 8 0 1 1 16 0"/><path d="M12 18l4.5-5"/></svg><span>Plan Usage</span></label>
<label class="set-chip" data-preview="header" data-preview-order="14"><input type="checkbox" id="appSettingsShowLifecycleLog"><svg class="set-chip-ico" width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.9" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z"/><polyline points="14 2 14 8 20 8"/><line x1="16" y1="13" x2="8" y2="13"/><line x1="16" y1="17" x2="8" y2="17"/></svg><span>Lifecycle Log</span></label>
</div>
@@ -2104,6 +2137,8 @@
<option value="claude-fable-5-1[1m]" data-variant="1m" data-base="claude-fable-5-1">Fable 5.1 (1M context)</option>
<option value="claude-fable-5" data-meta="Most powerful" data-base="claude-fable-5" data-ctx="1">Fable 5</option>
<option value="claude-fable-5[1m]" data-variant="1m" data-base="claude-fable-5">Fable 5 (1M context)</option>
<option value="claude-opus-5-5" data-meta="Latest Opus" data-base="claude-opus-5-5" data-ctx="1">Opus 5.5</option>
<option value="claude-opus-5-5[1m]" data-variant="1m" data-base="claude-opus-5-5">Opus 5.5 (1M context)</option>
<option value="opus" data-meta="Most capable" data-base="opus" data-ctx="1">Opus</option>
<option value="opus[1m]" data-variant="1m" data-base="opus">Opus (1M context)</option>
<option value="claude-opus-4-6" data-meta="Previous generation" data-base="claude-opus-4-6" data-ctx="1">Opus 4.6</option>
@@ -2114,7 +2149,7 @@
<div class="set-row" id="appSettingsContextRow" data-search="1m context window opus long">
<div class="set-row-text">
<span class="set-row-label">1M context window</span>
<span class="set-row-desc" id="appSettingsContextDesc">Available for Fable 5.1, Fable 5, Opus and Opus 4.6.</span>
<span class="set-row-desc" id="appSettingsContextDesc">Available for Fable 5.1, Fable 5, Opus 5.5, Opus and Opus 4.6.</span>
</div>
<label class="switch switch-sm"><input type="checkbox" id="appSettingsOpusContext1m"><span class="slider"></span></label>
</div>
@@ -2156,6 +2191,7 @@
<option value="">Default (CLI default)</option>
<option value="claude-fable-5-1">Fable 5.1 (Latest)</option>
<option value="claude-fable-5">Fable 5 (Most powerful)</option>
<option value="claude-opus-5-5">Opus 5.5 (Latest Opus)</option>
<option value="opus">Opus (Most capable)</option>
<option value="sonnet">Sonnet (Balanced)</option>
<option value="haiku">Haiku (Fast &amp; cheap)</option>
@@ -2171,6 +2207,7 @@
<option value="opus">Opus</option>
<option value="claude-fable-5-1">Fable 5.1</option>
<option value="claude-fable-5">Fable 5</option>
<option value="claude-opus-5-5">Opus 5.5</option>
</select>
</div>
<div class="set-mini">
@@ -2182,6 +2219,7 @@
<option value="opus">Opus</option>
<option value="claude-fable-5-1">Fable 5.1</option>
<option value="claude-fable-5">Fable 5</option>
<option value="claude-opus-5-5">Opus 5.5</option>
</select>
</div>
<div class="set-mini">
@@ -2193,6 +2231,7 @@
<option value="opus">Opus</option>
<option value="claude-fable-5-1">Fable 5.1</option>
<option value="claude-fable-5">Fable 5</option>
<option value="claude-opus-5-5">Opus 5.5</option>
</select>
</div>
<div class="set-mini">
@@ -2204,6 +2243,7 @@
<option value="opus">Opus</option>
<option value="claude-fable-5-1">Fable 5.1</option>
<option value="claude-fable-5">Fable 5</option>
<option value="claude-opus-5-5">Opus 5.5</option>
</select>
</div>
</div>
@@ -2216,6 +2256,60 @@
</div>
</div>
</div>
<div class="set-group" id="customModelEndpointsGroup">
<div class="set-group-head"><h4>Custom model endpoints</h4><span class="set-scope">synced</span></div>
<p class="set-group-hint">Point a harness at your own OpenAI-compatible server (llama.cpp, vLLM, DGX Spark, Azure AI Foundry, OpenRouter) instead of its native cloud backend. When on, the Run menu offers an extra entry per harness that supports it, per saved endpoint.</p>
<div class="set-group-body">
<div class="set-row" data-search="custom model endpoint llama.cpp local llm run menu picker">
<div class="set-row-text">
<span class="set-row-label">Enable custom model endpoints</span>
<span class="set-row-desc">Adds a per-endpoint entry to the Run menu for every harness that can redirect to one.</span>
</div>
<label class="switch switch-sm"><input type="checkbox" id="appSettingsCustomModelEndpoints" onchange="app.applyCustomModelEndpointsVisibility()"><span class="slider"></span></label>
</div>
<!-- Gated on the toggle above (applyCustomModelEndpointsVisibility): with the
feature off, a list of endpoints that do nothing is worse than nothing. -->
<div id="customModelEndpointsBody" style="display:none">
<div id="customModelHostsList" class="set-group-body" data-search="endpoints"></div>
<button type="button" class="btn-toolbar btn-sm" id="customModelHostAddBtn" onclick="app.openCustomModelHostEditor()">+ Add endpoint</button>
<div id="customModelHostEditor" class="set-inline-form" style="display:none">
<h5 id="customModelHostEditorTitle">Add endpoint</h5>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Id</span><span class="set-row-desc">Short, stable — used in URLs, never shown to the CLI.</span></div>
<input type="text" id="customModelHostId" class="set-input" placeholder="llama-cpp-local">
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Label</span></div>
<input type="text" id="customModelHostLabel" class="set-input" placeholder="llama.cpp (local)">
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Base URL</span></div>
<input type="text" id="customModelHostBaseUrl" class="set-input" placeholder="http://192.168.1.50:8080">
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">API key</span><span class="set-row-desc">Optional. Left blank on edit keeps the existing key.</span></div>
<input type="password" id="customModelHostApiKey" class="set-input" autocomplete="new-password">
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Auth header</span><span class="set-row-desc">Never send both — some servers hang indefinitely.</span></div>
<select id="customModelHostAuthStyle" class="set-select">
<option value="bearer">Authorization: Bearer (default)</option>
<option value="api-key">api-key header (Azure)</option>
</select>
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Default model</span><span class="set-row-desc">What the Run-menu picker applies for this endpoint. Discover models first.</span></div>
<select id="customModelHostDefaultModel" class="set-select" disabled></select>
</div>
<div class="set-row-actions">
<button type="button" class="btn-toolbar btn-sm" onclick="app.saveCustomModelHostFromEditor()">Save</button>
<button type="button" class="btn-toolbar btn-sm" onclick="app.closeCustomModelHostEditor()">Cancel</button>
</div>
</div>
</div>
</div>
</div>
</section>
<!-- ══ Agents &amp; CLIs ═════════════════════════════════════════ -->
@@ -2226,6 +2320,64 @@
</div>
<p class="set-section-blurb">Launch flags for the CLIs Codeman spawns.</p>
<div class="set-group" id="cliManagementGroup">
<div class="set-group-head"><h4>CLI management</h4><span class="set-scope">synced</span></div>
<p class="set-group-hint">Enable/disable a CLI, install one that's missing, or add your own — without hand-editing ~/.codeman/clis.json.</p>
<div class="set-group-body">
<div class="set-row" data-search="cli management enable disable install custom">
<div class="set-row-text">
<span class="set-row-label">Enable CLI management</span>
<span class="set-row-desc">Adds the list below and its write endpoints. Off by default: this changes machine configuration, not just what you see.</span>
</div>
<label class="switch switch-sm"><input type="checkbox" id="appSettingsCliManagement" onchange="app.applyCliManagementVisibility()"><span class="slider"></span></label>
</div>
</div>
</div>
<div class="set-group" id="cliListGroup" style="display: none;">
<div class="set-group-head"><h4>Installed CLIs</h4></div>
<div class="set-group-body">
<div id="cliListRows"></div>
<div class="set-row" data-search="add custom cli">
<div class="set-row-text">
<span class="set-row-label">Add a custom CLI</span>
<span class="set-row-desc">A launch command Codeman doesn't ship — id, label, badge, binary and its bare argv.</span>
</div>
<button type="button" class="btn btn-xs" id="cliCustomAddToggle" onclick="app.openCliCustomForm()">Add</button>
</div>
<form id="cliCustomForm" style="display: none;" onsubmit="app.submitCliCustomForm(event)">
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Id</span></div>
<input type="text" id="cliCustomId" class="set-input" placeholder="my-cli" maxlength="24">
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Label</span></div>
<input type="text" id="cliCustomLabel" class="set-input" placeholder="My CLI" maxlength="60">
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Badge</span></div>
<input type="text" id="cliCustomBadge" class="set-input" placeholder="MC" maxlength="6">
</div>
<div class="set-row has-field">
<div class="set-row-text"><span class="set-row-label">Binary</span></div>
<input type="text" id="cliCustomBinary" class="set-input" placeholder="my-cli">
</div>
<div class="set-row has-field">
<div class="set-row-text">
<span class="set-row-label">Launch argv</span>
<span class="set-row-desc">Space-separated bare words, e.g. "my-cli --flag". No quoting or shell syntax.</span>
</div>
<input type="text" id="cliCustomArgv" class="set-input" placeholder="my-cli --flag">
</div>
<div class="set-row">
<button type="submit" class="btn btn-xs" id="cliCustomSubmit">Create</button>
<button type="button" class="btn btn-xs" id="cliCustomCancel" onclick="app.closeCliCustomForm()">Cancel</button>
</div>
<div id="cliCustomFormError" class="set-row-desc" style="color: var(--error, #e5484d); display: none;"></div>
</form>
</div>
</div>
<div class="set-group">
<div class="set-group-head"><h4>Claude</h4><span class="set-scope">synced</span></div>
<div class="set-group-body">
@@ -2878,6 +3030,11 @@
<input type="number" id="remoteHostPort" placeholder="22" min="1" max="65535" autocomplete="off">
<span class="form-hint">Optional. Leave blank for the default port 22.</span>
</div>
<div class="form-row">
<label>Wake-on-LAN MAC</label>
<input type="text" id="remoteHostWakeMac" placeholder="04:d9:f5:80:c6:58" autocomplete="off" autocapitalize="off" spellcheck="false">
<span class="form-hint">Optional. Comma-separated for several NICs. Codeman sends the magic packet itself so a sleeping host can be woken from the session banner.</span>
</div>
<div class="form-row">
<label>Codex Command Override</label>
<input type="text" id="remoteHostCodexCommand" placeholder="exec codx personal" autocomplete="off" autocapitalize="off" spellcheck="false">
@@ -2886,6 +3043,11 @@
<details class="advanced-options">
<summary><svg class="set-adv-chev" width="12" height="12" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.4" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M6 9l6 6 6-6"/></svg><span>Advanced SSH</span></summary>
<div class="advanced-options-content">
<div class="form-row">
<label>Wake Command</label>
<input type="text" id="remoteHostWakeCommand" placeholder="/home/user/bin/wake-this-host" autocomplete="off" autocapitalize="off" autocorrect="off" spellcheck="false">
<span class="form-hint">Optional override for the MAC above (takes precedence). A single executable path, run without a shell — use it when the host needs a router/other machine to send the packet.</span>
</div>
<div class="form-row">
<label>Identity File</label>
<input type="text" id="remoteHostIdentityFile" placeholder="~/.ssh/remote_ed25519" autocomplete="off" autocapitalize="off" autocorrect="off" spellcheck="false">
@@ -3013,6 +3175,7 @@
<h2>Manage</h2>
</div>
<p class="set-section-blurb">Reorder or remove cases, and pick up anything exported from a docker case.</p>
<input type="search" id="caseManageSearch" class="set-input" placeholder="Search cases by name or path" autocomplete="off" spellcheck="false" aria-label="Search cases" oninput="app.setCaseManageFilter(this.value)" style="margin-bottom: 8px;">
<div class="case-manage-list" id="caseManageList">
<!-- Populated by JS -->
</div>
@@ -3042,7 +3205,12 @@
<h3>Select Case</h3>
<button class="modal-close" onclick="app.closeMobileCasePicker()" aria-label="Close case picker">&times;</button>
</div>
<div class="mobile-case-picker-search">
<svg width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" aria-hidden="true"><circle cx="11" cy="11" r="7"/><line x1="21" y1="21" x2="16.65" y2="16.65"/></svg>
<input type="search" id="mobileCaseSearch" placeholder="Search cases" aria-label="Search cases" autocomplete="off" autocapitalize="off" autocorrect="off" spellcheck="false" enterkeyhint="go" oninput="app.filterMobileCases()" onkeydown="app.onMobileCaseSearchKey(event)">
</div>
<div class="mobile-case-picker-body">
<div class="mobile-case-empty" id="mobileCaseEmpty" hidden>No matching cases</div>
<div class="mobile-case-list" id="mobileCaseList">
<!-- Cases populated by JS -->
</div>
@@ -3481,6 +3649,36 @@
text is set via value/textContent only: predictor output derives from
observable (injectable) content, and the explicit click here is the
security boundary (nothing is ever auto-sent). -->
<!-- Wake-on-LAN setup for a remote host whose session cannot be woken yet. Kept
deliberately small (host is fixed, only the wake fields are editable) so it can
be opened from the banner with one click. Persists via PUT /api/remote-hosts/:id. -->
<div class="modal" id="wakeConfigModal">
<div class="modal-backdrop" onclick="app.closeWakeConfigDialog()"></div>
<div class="modal-content">
<div class="modal-header">
<h3>Wake-on-LAN &middot; <span id="wakeConfigHostLabel"></span></h3>
<button class="modal-close" onclick="app.closeWakeConfigDialog()" aria-label="Close">&times;</button>
</div>
<div class="modal-body">
<div class="form-row">
<label>MAC address(es)</label>
<input type="text" id="wakeConfigMac" placeholder="04:d9:f5:80:c6:58" autocomplete="off" autocapitalize="off" autocorrect="off" spellcheck="false">
<span class="form-hint">Comma-separated for several NICs. Codeman sends the magic packet itself (UDP port 9, broadcast).</span>
</div>
<div class="form-row">
<label>Wake command (optional)</label>
<input type="text" id="wakeConfigCommand" placeholder="/home/user/bin/wake-this-host" autocomplete="off" autocapitalize="off" autocorrect="off" spellcheck="false">
<span class="form-hint">Takes precedence over the MAC. A single executable path, run without a shell.</span>
</div>
<div class="form-hint" id="wakeConfigStatus"></div>
</div>
<div class="modal-footer">
<button class="btn-toolbar" onclick="app.closeWakeConfigDialog()">Cancel</button>
<button class="btn-toolbar btn-primary" id="wakeConfigSave" onclick="app.saveWakeConfig()">Save</button>
</div>
</div>
</div>
<div class="modal" id="readMyMindModal">
<div class="modal-backdrop" onclick="app.closeReadMyMind()"></div>
<div class="modal-content readmymind-modal">
@@ -3544,6 +3742,7 @@
<script defer src="app.js"></script>
<script defer src="tab-rail-resize.js"></script>
<script defer src="terminal-ui.js"></script>
<script defer src="terminal-split.js"></script>
<script defer src="respawn-ui.js"></script>
<script defer src="ralph-panel.js"></script>
<script defer src="orchestrator-panel.js"></script>
@@ -3556,6 +3755,7 @@
<script defer src="reboot-restore-ui.js"></script>
<script defer src="admin-ui.js"></script>
<script defer src="session-ui.js"></script>
<script defer src="host-wake-ui.js"></script>
<script defer src="webview-tabs.js"></script>
<script defer src="mobile-overview.js"></script>
<script defer src="home-sessions.js"></script>

Some files were not shown because too many files have changed in this diff Show More