Adds claude-opus-5-5 to the App Settings model picker (base option with
data-ctx="1" plus its [1m] companion row, since Opus 5.5 has a 1M window)
and to the five task-routing selects, mirroring how Fable 5.1 was added.
* fix(sessions): stop pinning the w1-myapp placeholder as Claude's /resume title
Local claude spawns passed the tab name as `--name`. That flag is not only the
cross-session peer name: it is also the prompt-box label, the `/resume` picker
entry and the terminal title, and a pinned title stops Claude generating its own
(`customTitle ?? aiTitle`). So every conversation of a case was listed in
`/resume` as the same `w1-myapp`, and none of them got a generated title. On one
workspace, 34 of 34 conversations spawned with `--name` had no ai-title, while
every conversation spawned without it had one.
Only a name the user chose is pinned now: `Session.cliPinnedName` is the name
when `nameSource === 'manual'`, carried to the builders as a separate `cliName`
so the tab/mux name is untouched. Placeholder and auto names let Claude title
the conversation again.
A rename in Codeman also reaches `/resume`: the new name is appended to the
conversation's transcript as the `custom-title` row `/rename` writes (never
creating the file, never writing an empty title). For a pane spawned without
`--name` this holds immediately; a pane spawned with one re-appends its own
title each turn, so there the new name holds from the next spawn.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(sessions): skip no-op renames and docker sessions when syncing the /resume title
A same-name PUT (the Session Options field saves on blur and recomposes the
unchanged placeholder) no longer flips nameSource to manual or appends a
custom-title row, and docker sessions skip the host transcript scan since their
transcript lives in the container. The skill pages no longer use a w<N>- name
as the peer-name example, and the changeset notes the re-append caveat.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs: record that nameSource decides --name and renames reach /resume
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: codeman-local <codeman@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
_loadSessionManagerList() re-projects each unified item into the
history-record shape _buildHistoryItem renders, and dropped these three
fields. The row's own onActivate still read them from the unified item, but
everything built from the record did not: the ⋯ menu's "Resume session"
relaunched a codex row as claude (no mode, no resumeId), a resumed session
lost its conversation id, and Cmd+K rows showed no mode badge. Same class
of bug as the worktree fields the re-projection already carries (#266).
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (da07b38c, db4557d9) with no
corresponding update to the plan doc itself — every checklist still read
Status: TODO and every box unchecked. Brings the doc in line with the tree:
- A new "Status as of 2026-09-22" section up top: what's actually
implemented (verified by grepping the routes/schema/UI, not just trusting
the commit messages), the availability-flag staleness bug found and fixed
in this session (commit 0c77dd0a) with its devbox verification record, and
one real outstanding gap.
- The outstanding gap: a custom CLI created via Phase 5's write API has no
way to actually be launched. The Run menu is static per-mode markup with
no consumer of window.__codemanCliCatalog, so Phase 6's own "create a
custom entry, confirm it can be launched" verify step was never actually
exercised against this. Documented with two candidate fixes, neither
started.
- Each phase's checklist flipped to [x] where confirmed present in the tree,
Status lines updated from TODO to DONE, and the two originally-open
questions (Phase 2's installed source, Phase 5's PUT endpoint shape)
marked resolved against what actually shipped.
No code changes in this commit — documentation only, so a future session
(or the one already mid-flight on a separate checkout of this same branch)
picks up accurate status instead of a stale plan.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs: add the CLI-registry deployment plan and the parked Copilot plan
Both were sitting as untracked scratch files in the master checkout,
never committed to any branch. Moving them here rather than leaving them
loose:
- DEPLOYMENT_PLAN.md is the live tracker for the CLI-registry follow-up
series (PR A #347 merged, PR B #380 merged, PR B2 merged as #458) and
is where PR C (this branch's own CLI-management work) belongs.
- docs/copilot-integration-plan.md is explicitly PARKED, referenced by
name in docs/cli-enable-disable-plan.md's own header as a sibling plan
tracked separately — kept for continuity, not active on this branch.
The other scratch files found alongside these (PRA.md, PRB.md, PR-B2.md
and their review-response counterparts) described PR A/B/B2, all now
merged — deleted from the master checkout as stale rather than committed
anywhere, since their content is superseded by the real merged PRs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* fix(cli-registry): render enabled CLIs in launch surfaces
* test(cli-registry): update frontend branch guard
* fix(test): isolate suite from deployment environment
* fix(cli-registry): revise Decision 4 - claude is toggleable, shell stays permanent
shell/claude were both structurally un-disableable in the original plan
(Decision 4). Revised: shell keeps the hard backend guarantee (it is the
one non-agent mode several code paths assume always exists as a raw-
terminal fallback), but claude is now a normal toggleable entry like any
other CLI.
Safe to do because internal session creation (tmux-manager.ts, session.ts,
Ralph, plan-orchestrator) resolves a CLI via getCli(), which does not
check `enabled` at all - only the Run menu and the HTTP-facing
sessionModeSchema() (new session requests through the normal API) key off
it. Disabling claude therefore behaves identically in kind to disabling
any other CLI: no internal fallback path breaks, it just stops being
offered for new sessions until re-enabled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): hide shell's toggle entirely instead of greying it out
A permanently-disabled switch next to every other row's working toggle
read as broken rather than intentional. shell now renders no switch at
all - a plain "Always available" label - so there is nothing to click
that could look like it should work but doesn't. Backend guard is
unchanged (UNDISABLEABLE_IDS still refuses shell unconditionally); this
is UI-only.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): sort the Installed CLIs list, installed-first then alphabetical
renderCliList() previously rendered in registry order (each entry's fixed
order field). Now sorts installed CLIs first, then not-installed, each
group alphabetical by label - matches how a user actually scans the list
(what's ready to use, then what needs installing).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* style: prettier fixes from the master merge
* fix(cli-registry): install/edit take effect immediately, confirm before install, phone labels
Four gaps found verifying #476 against the #343 review trail:
- Installed or edited CLIs kept reading as missing/stale. Every binary lookup
(the nine per-CLI resolvers and the generic registry one) caches in its own
closure, with a negative-cache backoff of up to 5 minutes, and nothing
cleared them. invalidateCliExecutableResolvers(binaries) now drops those
caches per binary; install (success or failure), create, edit and delete
call it plus invalidateCliResolverCache(id). Before this, a CLI installed
from Settings could fail to launch for minutes, and an edited custom entry
kept launching its old binary until a restart.
- The Settings "installed" badge for a custom entry used a private `which`,
ignoring the entry's searchDirs and the login-shell lookup that spawn and
the Run menu use; it now asks the same generic resolver they do.
- Install ran on a single click. The #343 review asked for auto-install to
sit behind an explicit confirm; the confirm now names the exact command,
which GET /api/clis returns for stock entries only (installCommand).
- The phone Run button showed the two-letter tab badge ("CC", "CX") instead
of the word ("Claude", "Codex"). It uses the registry label again, which is
identical to the old static table for every stock CLI (now pinned).
14 new tests; 9 of them fail against the previous head and pass here.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): address #476 review — safe serialized writes, no id branches, docs
Must-fix:
- registry-writer: start fresh only on ENOENT; refuse (409) a clis.json that
does not parse or has group/world permission bits instead of overwriting it
(isUnsafePermissions now exported from registry.ts)
- mutateRegistryFile(): one promise chain for every mutation, with the
existence/duplicate checks inside the serialized step, plus a unique tmp
name per write
- docs: CLAUDE.md, architecture-invariants, cli-registry (new Settings
section) and api-reference (the six /api/clis routes)
- drop DEPLOYMENT_PLAN.md and docs/copilot-integration-plan.md
Smaller:
- PUT /api/clis/custom/:id keeps the entry's current enabled state when the
body omits it
- runMode setter falls back to the first enabled catalogue entry, not 'claude'
- shell guard keyed on kind === 'shell' (routes + Settings list); stock probe
map shared with server.ts via utils/cli-installed-probes.ts
- stock claude label is now 'Claude Code', so the Run menu / phone overview
label rewrites are gone (doctor row keeps "Claude CLI" via its override)
- welcome buttons are translatable again and read "Run Claude Code" /
"Run Shell"; zh-CN gains "Run Codex" / "Run OMP"
- install: per-id in-flight guard (409) and CODEMAN_* stripped from its env
- fileoverview / CliEnableSchema comments no longer say stock-only
- test-env isolation changes moved to their own PR
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* test(cli-registry): pin the #343/#347 findings #476 makes reachable
A CLI toggled or created through the routes is accepted or rejected by
CreateSessionSchema with no restart (#343 finding 2), and a custom CLI created
through the API renders a real local, remote and docker launch command
(#347 finding 5: no more `cd <path> && undefined`).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(self-update): stop a stalled status from blocking every later update
A Homebrew node upgrade under a long-running server deletes the versioned
Cellar path the server passes as --node, so every status write from the
updater failed. The update itself still built and restarted (npm and the
build use node from PATH), but update-status.json stayed "queued" forever.
The boot reconcile ran one minute after the restart, inside its 15 min
window, and isInFlight() had no age limit, so "An update is already in
progress." blocked every later update until the next server restart.
- self-update.sh falls back to node on PATH when --node is not executable.
- expireStalledStatus() (pure) fails an in-flight status whose last write
is older than the stale window; applied on every read (start + status
poll) and persisted. The live updater heartbeats every few seconds, so a
running update never trips it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(self-update): a hung graceful shutdown no longer leaves a LaunchDaemon install down
On a KeepAlive LaunchDaemon (headless macOS) the updater restarts by sending
the server SIGTERM and letting launchd respawn it. launchd only respawns once
the process EXITS, and nothing escalates a stuck stop (systemd would SIGKILL
after TimeoutStopSec). Observed after an update to 1.32.1: the server closed
port 3000, server.stop() never resolved, the process stayed alive and the
service stayed down until it was killed by hand.
- cli.ts: the signal handler arms an unref'd 10s timer that force-exits if
server.stop() hangs.
- self-update.sh (launchd-daemon): wait up to 30s for the server pid to exit,
then SIGKILL it. tmux sessions live outside the server and survive.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
- docs/wiki/The-Dashboard.md: the tab-appearance table gains the exited
state (muted dot plus an `exited (137)` badge) and explains the bare
`exited` variant.
- The detailed sidebar and rail no longer pair the muted dot with an "idle"
pill: an exited session's pill reads "exited" (neutral styling) and its
since stamp measures from the observed exit. This is a label override on
the row model, not a new state, so SESSION_ACTIVITY_RANK and the home
screen order are untouched, and a pending alert still keeps its own pill.
The row signature includes the flag so the incremental path repaints it.
- The exited badge is aria-hidden like its sibling badges, and the exit is
appended to the tab's aria-label in both render paths through one helper.
- test/tmux-manager.test.ts re-adds the junk-trailing-field parser case
against parsePaneRows.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- test/setup.ts strips CLAUDE_CONFIG_DIR (pinned in test-env-isolation), so
transcript-fixture tests such as session-custom-model-restart no longer go
red on a machine that exports it for a separate Claude account (#255).
- The vanished-tmux-session branch of _setupOrAttachMuxSession() relaunches
the CLI through createSession() just like a failed respawn, so it now takes
the same resume pin. A genuinely new session is unaffected.
- After a dead-pane respawn of a fallback-chain CLI, _claudeSessionId names
the conversation the walk actually pinned instead of the chain tail, which
the walk may have passed over for lack of a transcript.
- _claudeConfigDir() trims the override like claudeProjectsDir() does.
- The remote-reattach test is labelled as documentation, since the pin
builder's own remote guard would make it pass either way.
- CLAUDE.md: the create-path pin persists through toState() as
resumeSessionId, and the end of the walk adds no pin rather than clearing
the launch seed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- The TERMINAL DROP crash-trail line moves behind the scheduler's debounce
guard, so it is written once per window rather than once per dropped
frame. At the server's 8ms batching, one second of drops evicted the whole
50-entry trail, including the recovery lines that explain it.
- A refresh that failed at the capture fetch deadline now returns
'deadline', and the scheduler does not retry it: that is a stalled link,
not contention, and each retry was another ?full=1 capture waiting out a
deadline of up to two minutes. The early-return retries are unchanged.
CLAUDE.md and the code comments no longer claim every skip reason is
transient contention.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- While another device holds the pane width (_paneWidthRefused), a resize
now asks for the container's width without applying it locally
(_geometryForResizeRequest: rows follow the container, columns stay at
the PTY's). Fitting first re-wrapped the whole buffer to the container
and back on every 30s mobile retry, and throttledResize ran the
scrollback clear for a resize that brings no redraw. selectSession
clears the flag, since it belongs to the previous pane. New unit tests
run the real mixin against a fake terminal and fail without the fix.
- Session seeds _ptyCols/_ptyRows at spawn (_notePtySpawnGeometry), so a
reattached pane reports its tmux window's real size through ptyGeometry.
- Session.resize's declined-branch comment names ptyGeometry, not the
deleted ptyCols/ptyRows getters.
- Delete the dead terminalGeometryAgrees() and its window export.
- test/xterm-private-api.test.ts header: it pins the exact locked version,
so any bump fails, not only a major.
- The main-terminal fit sweep also matches fitAddon?.fit?.(), and
CLAUDE.md names the modules it actually covers.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Clone Repo clears the credential helpers for a non-admin, but a non-admin's
Docker case with credential seeding on still receives a copy of the server
account's gh/az sign-in when the agent-image switches are on, the same as
the Claude and Codex credentials. Say so in the multi-user notes so the docs
do not read as a stronger guarantee than they are.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Remove exactly the codeman-node-modules/codeman-dist volumes by Compose
label after a plain `down`, instead of `down --volumes` (which also takes
any volume an override file declares while the message named two).
`down --volumes` remains only as a warned fallback when the project name
cannot be resolved.
- Report a failing first `docker compose config --format json` call with a
clear error instead of exiting silently under `set -e`.
- Filter empty label lines in the collision guard so an unlabelled container
cannot hide a real collision; name the moved-checkout exit in its error.
- Comments no longer cite a guard or incident in Start-Codeman.sh that does
not exist; the README states the real gap (a Node base-image bump leaves
codeman-node-modules stale because the lockfile did not move).
- docs: Update-Codeman.sh in the docker-self-update.md short-version table
and a mention in docker-compose.md; "Major updates" moved under "Updating"
in docker/README.md.
- test: smoke test covers the new sequence, the config failure and the
empty-line case; quiet stdio; @fileoverview names the fourth concern.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- session.ts: a pane capture that fails now CLEARS the watching label
(and emits watchingChanged so pages drop the badge) instead of keeping
the last one, so a failed capture degrades toward an alert rather than
pre-acknowledging the next real idle prompt. Test updated; invariant
noted in architecture-invariants.
- approvals-ui.js: the header bell counts only unacknowledged items
(pendingApprovalsCount), matching codeman tui's pendingApprovalCount();
pinned in watching-no-alert.test.ts.
- mobile-overview.js: move the orphaned "Pill copy per state" JSDoc back
onto MOBILE_OVERVIEW_PILL_LABEL.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- stock.ts: claude is no longer the only entry declaring transcriptGutter;
codex declares it too.
- architecture-invariants: the strip applies when the session's CLI declares
a margin (not detection), and a note that it keys on the session's launch
mode, not on what is running in the pane (a claude pane dropped to a shell
still loses up to two columns; copyStripMargin is the escape hatch).
- render-index-html test: the gutter map is injected for a solo
/session/:id render as well.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- mobile.css: gemini and antigravity run/gear rules get `!important` like
pi/omp/grok/deepseek, so the gear half no longer keeps the skin accent
while the body takes the mode colour (two-tone button on the default skin).
- test/skin-themes.test.ts: static guard that every run mode with a base
`.btn-toolbar.btn-run.mode-<id>` rule also has a resting rule inside the
`html:not([data-skin="og"])` block; ids are derived from the stylesheet.
- stock.ts: grok's accent comment names zinc-300 (border/badge colour);
gemini's accent is #8ab4f8 to match its tab badge and run-mode dot, noted
as the one exception to the border-colour method.
- types.ts: "(below)" -> "(above)".
- docs/cli-registry.md, CLAUDE.md: `accent` is now measured, not transcribed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CLAUDE.md loads into every session, and its Architecture section had grown
feature write-ups (history, measurements, rationale) that belong in
docs/architecture-invariants.md per the file's own header. Each long block
now keeps what the feature is, where it lives, its setting/default and the
rules that prevent real bugs, and links to its invariants section. Everything
removed was moved there: 29 new sections, extra facts appended to the
existing ones.
Also: hard-coded counts (SSE events, route handlers, module/file counts,
device profiles) replaced by pointers to the source of truth, and the
Debugging commands fixed to use the codeman tmux socket and HTTPS for prod.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Addresses the review on #472.
- CRED_STORES: `.config/gh` and `.azure` now carry `enabledByEnv`
(CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ), and resolveDockerCredentialArtifacts
skips a store unless that variable is exactly `1`, read at container
create. A host that merely has ~/.config/gh/hosts.yml or a plaintext MSAL
cache no longer copies them into every case container. Tests: the default
environment seeds neither even with the files present, and each store
follows only its own switch.
- Multi-user mode: a non-admin's Clone Repo clone and preflight run with
`git -c credential.helper=` (GIT_NO_CREDENTIAL_HELPERS, placed before the
subcommand), so the server account's helpers are never lent to them.
Verified against a real private repo that it also clears the URL-scoped
credential.<url>.helper entries, and that public clones still work.
Tests: the argv in test/git-clone.test.ts, and the route decision
(non-admin cleared; admin and single-user kept) in
test/routes/case-clone-credential-helpers.test.ts.
- Docs: recreate the case container to pick up seeds (docker/README.md,
Docker-Cases wiki, docker-cases.md); the multi-user behaviour in
docker/README.md and security-architecture.md; "functionally unchanged"
instead of "unchanged" for an image built with both switches off
(server.Dockerfile comment, README, changeset).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw
CONTRIBUTING says releases are handled by the maintainer via changesets after
merge, and every `.changeset/*.md` on master was written by him or by the
release bot — including the ones covering other people's pull requests. The
summary this file carried moves to the pull-request description, where it is
the maintainer's to reuse or rewrite.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claimed after a manual test that a session comes back from a server restart
without its badge until it next produces output. Measured instead of assumed,
and it is wrong: a codex session with a background terminal still running had
its label back within about 20 seconds of the restart, with no input from
anyone. Reconciliation re-attaches the pane, the attach repaint carries the
composer glyph, the idle confirmation arms on it, and the probe re-reads the
label — the ordinary path, doing the ordinary thing.
What produced the false claim was a session whose monitor had simply expired
while it sat there. Its footer carries no chip, so `watching: null` was the
right answer and there was nothing missing to restore.
Recorded at the field and in the invariants, because the shape of this invites
exactly one wrong fix: a polling timer to keep a value fresh that the pane
already refreshes by itself.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add Case -> Clone Repo could only reach public repositories in the Docker
deployment. This lets a deployment opt in to the GitHub CLI and the Azure
CLI (+ azure-devops extension) as git credential helpers. Codeman itself
still collects no credentials.
- server.Dockerfile / agent.Dockerfile: CODEMAN_INSTALL_GH /
CODEMAN_INSTALL_AZ build args (0 or 1, default 0; anything else stops the
build). Off leaves no apt repository, package, extension, helper script
or credential entry, so a default build is unchanged. On installs from
the vendors' apt repositories and configures system gitconfig helpers:
github.com / gist.github.com -> `gh auth git-credential`, dev.azure.com /
*.visualstudio.com -> new docker/git-credential-azure-cli (an Entra ID
token from `az account get-access-token`, or AZURE_DEVOPS_EXT_PAT).
A helper whose CLI is not signed in prints nothing, so a private clone
still fails fast.
- The extension lives in AZURE_EXTENSION_DIR outside HOME
(/opt/codeman-az-extensions, runtime-owned; /opt/az-extensions, gid-0
group-writable in the agent image).
- Hosts turn them on in docker-compose.override.yml: `build: args:` for the
server image, `environment:` CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ for the
agent image. build-agent-image.mjs and the in-app auto-build share one
env -> ARG table (pinned by the parity test) and pass nothing when unset.
docker-compose.yaml is untouched; .env.example only gains a comment, so
the self-updater's environment gate sees no new keys.
- Docker cases seed the gh sign-in (~/.config/gh/hosts.yml, config.yml) and
the az sign-in files from ~/.azure per file, read-only, like pi/grok.
- The Clone Repo AUTH_REQUIRED message says how to sign the server's git
in instead of claiming private repositories cannot be cloned.
- Docs: docker/README.md "Private repositories", docker-compose.md,
docker-cases.md, the Quick-Start / Core-Concepts / Docker-Cases wiki
pages, security-architecture.md, architecture-invariants.md, changeset.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw
Review fixes for #469.
The Ctrl+C branch cleaned the selection to decide whether to copy and then
passed that cleaned string to copyTerminalSelection(), which cleans again. The
trailing trim is a fixed point, so that was safe until this PR; the margin
strip is not, because it takes the lesser of the declared width and the run
every line shares, so a second pass takes up to `margin` columns more. The
branch now gates on the cleaned string and hands the raw one on. Verified in
chromium with a real drag, a real Ctrl+C and a real clipboard read on a live
claude pane: an on-screen ` fix(terminal): trim it` reaches the clipboard
as ` fix(terminal): trim it`, and reverting the branch reproduces the
reported ` fix(terminal): trim it`.
Pane B of a split resolves its own width. `_cliGutterColumns()` and
`_normalisedSelectionRange()` take the session and the terminal to read,
defaulting to the primary pane's, so Pane B looks its own run mode up instead
of keeping a margin Pane A drops on the same keystroke. Verified live with two
claude panes open side by side.
A detached session window (`/session/:id`) receives the gutter map. The
injection sat inside the block that skips the run menu's payloads for a solo
window, so the toggle worked in the main window and did nothing in the popup on
the same device. It needs no availability probe, so it moved below that block
and the solo window still carries none of the payloads it skipped before.
The settings description said the width is measured and named Codex as exempt.
Nothing is measured, and Codex is one of the two panes that are stripped.
docs/wiki/Settings-Reference.md gains the row every Terminal and Input toggle
carries. CLAUDE.md no longer says the clean touches trailing runs "and nothing
else" one sentence before the leading-margin rule, and both it and
docs/architecture-invariants.md record that the strip is not idempotent.
Two round-trip tests run on a mode that declares a gutter, which the existing
copyTerminalSelection cases could not, since they all use the harness default
mode that declares none. The Ctrl+C branch itself is pinned at the source,
because it lives inside initTerminal's attachCustomKeyEventHandler closure over
a real xterm the vm harness cannot build. Both pins fail on the reintroduced
bug. Gate: 7865 passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`_onSessionTerminal` drops an incoming frame once the app-owned render queues
already hold 128KB. That is the right call — the alternative is an unbounded
backlog — but a hole in a TUI byte stream is a desynced cursor, and a desynced
cursor is muffled text (#464). The drop was only half of it.
The recovery was a fire-and-forget timer: it nulled its own handle and then
called `_onSessionNeedsRefresh()`, which opens with four early returns. Two of
them — a buffer load already in flight, a refresh already owning this session —
are MOST likely to be true during exactly the output burst that caused the
drop. So the recovery was skipped precisely when it was needed, with nothing
left to retry it, and the dropped bytes were never replayed.
`_onSessionNeedsRefresh` reports whether it actually repainted now, and
`_scheduleDroppedOutputRecovery` re-arms while it has not. Bounded by
`DROP_RECOVERY_MAX_ATTEMPTS`, because every reason the refresh can be skipped is
transient contention that clears in seconds and a permanently failing refresh
must not become a loop against the API; giving up at the cap leaves exactly what
the old code left, so the floor is no worse than before. The same 2s debounce
still collapses a burst of drops into one attempt.
This is the principle Ark0N established reviewing #431 for the WebSocket
output-gap marker — only a repaint that actually happened settles the recovery —
applied to the one recovery path that still trusted a timer having fired.
The retry decision is a pure function in constants.js so the gate can reach it,
and the scheduler itself is driven from app.js under a fake clock. The retry
case and the no-retry case only pin the fix AS A PAIR: either alone passes
against something wrong, one against the old fire-and-forget timer and the other
against retrying forever. Checked by reverting app.js to the old shape, where
three of the twelve fail.
Two harness details that would otherwise have made the tests lie. The vm context
baked in the real `setTimeout`, so `vi.useFakeTimers()` could not reach the
scheduler and every case reported zero calls; it delegates lazily now. And
app.js reached `CodemanDroppedOutput` as a bare global, which resolves in a
browser but not in the vm — worth fixing beyond the test, because that call sits
inside a timer where a ReferenceError is swallowed and would take the recovery
with it. It reads through `window.` like terminal-ui.js does with its own
constants.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Reported from a manual test: a Codex session went on showing the watching badge
after its background terminal had finished. The server was right and the page
was stale — `Session.watching` changes while the session's status does not, and
nothing broadcast it.
The label is usually SET on the idle transition, which broadcasts anyway, so the
badge always appeared correctly. It CLEARS when the work ends, and a CLI can end
background work without taking a turn: codex repaints its background-terminal
row away and stays idle, so `_confirmIdle()` concludes without emitting `idle`
(that emit is guarded by `wasWorking || isInitialReady`) and no other event
fires. Every open page kept drawing a badge the server had already dropped.
`_readWatching()` now emits `watchingChanged` when, and only when, the label
really changes, and the wiring pushes the session state on it. No new SSE event:
the badge reads off the session payload every surface already has.
A/B measured on an isolated beta with the page loaded and then left untouched.
Without this commit the server dropped the label at t+50s and the page still
showed the badge at t+100s; with it, page and server cleared in the same
ten-second window.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Measured on the same beta, codex-cli 0.154.0: a sub-agent started without
waiting outlives the turn exactly as a background terminal does — the sandboxed
process was still running — and codex shows nothing for it. The last rows of the
pane are the composer and the status line, and `Sub-agents running` belongs to
the on-demand `/subagents` panel rather than to the row above the composer.
So there is no second codex label to add. A codex session waiting on a sub-agent
reads as plainly idle, which misfiles nothing (codex raises no idle prompts) and
simply leaves that one kind of quiet unexplained until codex pins a row for it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief)
found the trust boundary weaker than the comments around it claimed. Eleven
findings, all applied.
The two blockers were both about who can write the row the label is read from.
Claude's window covered two rows, and the second one is the status line, whose
command a session running with permissions bypassed can write into its own
`.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of
its own and silence its own idle alert. The default window is one row now, which
is the footer and nothing else, and the constant says why. Separately, the label
reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML,
which is an injection sink for any config-supplied pattern whose capture group is
permissive; it goes through escapeHtml() like every other untrusted string in
that file.
The Codex entry could not be fixed the same way, and now says so. Its row is
third from the bottom only while a terminal runs; with none running that slot
holds the last row of the transcript, so matching the complete row (with the
`/stop to close` tail, window narrowed to three) raises the bar without closing
it. What contains it is `hooks: 'none'`: no hook event from a codex session
reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an
alert. The registry comment, `docs/cli-registry.md` and the test all state that
rather than claiming a guarantee the code does not have.
Also from the review: the TUI header badge no longer counts an acknowledged
item, which was the same gate the classifier fix already went through and was
wrong for human acknowledgement too; the TUI approval card reads the quiet
reason and drops to a new `info` tone instead of asking for a reply; the badge
carries an aria-label, because the phone it was built for has no hover target;
the schema refuses `watchingLines` without a `watchingLine`; and the pattern and
its window are resolved together rather than one memoized and one not.
Documentation moved with it. The mechanism now lives in
`docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer,
`docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped
buzzing, and both that page and the changeset name the limitation neither did
before: a question asked in plain prose is not a dialog, so it is silenced along
with the false alarms while background work runs.
Verified live again after the narrowing, on an isolated beta: a Claude session
reported `1 monitor` and took its idle prompt acknowledged, and a Codex session
reported `1 background terminal` against the full-row anchor.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ten findings from a two-model review of this branch. Both reviewers cleared the
detection logic itself; everything here is a gap around it.
A route that starts a command in a pane now PERSISTS as well as broadcasts.
`/interactive` and `/shell` did neither before, and the pane-exit watcher cannot
cover for them: its next tick finds `paneExit` already cleared in memory,
reports no change and writes nothing, so `state.json` kept saying the agent had
exited for as long as the session stayed quiet. Nothing reads that record for a
decision yet, which is exactly why it had to be fixed now — part 2 is designed
to read it. The `clearPaneExitForNewPane()` docstring claimed its callers
already persisted; that claim was false for these two, and now says what the
caller owes instead.
The watcher's four guards were unreachable by any test. `refreshPaneExits()`
opened with `if (IS_TEST_MODE) return;`, so the read gate, the in-flight
suppression, the generation counter and the empty-read rule could each be
deleted with the whole suite green. The tmux call moves into `readPaneRows()`,
which a test subclass overrides — the shape `runRemoteReconnectTick` already
uses in this file for the same reason — and the test-mode gate moves with it, so
what a test cannot do is spawn a process rather than exercise the bookkeeping.
Each of the four guards now has a test that fails when it is deleted.
The muted status dot turned out to be a specificity fight on three surfaces, not
two. `.tab-status.error` was not excluded, so a session whose agent exited and
whose PTY-exit breaker then tripped lost its red dot to the mute — the state the
browser answers with a "restart it?" confirm, and a needs-you colour by the same
argument that protects the two alert classes. And mobile.css gives a `busy` dot
a 9px size and a green glow with `!important`, while `status` stays `busy` for a
pane whose agent died mid-turn, so a phone rendered a grey dot still wearing the
green halo beside a badge reading "exited". Both measured against the real
stylesheets, both now excluded, and the CSS test reads mobile.css too instead of
being structurally blind to half the problem.
Six comments said things that were not true. Two named the stats collector as
what replaces a restored reading, which is the opposite of the design. The
interval constant argued that 2000 ms keeps a read inside a tick, when the
5000 ms exec timeout means it cannot — which is why the in-flight guard exists.
`MuxSession.discovered` did not say the flag is permanent, though `saveSessions()`
serializes it. The empty-read docstring claimed a distinction that `|| true`
makes impossible. The invariants doc promised more than its drift test delivers.
And CLAUDE.md had no pointer at all, leaving its two hardest prohibitions
("never set `status: 'error'`", "never null the pid") only in the file it is
meant to route people to.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Five items, two of which he could only see by running it, plus six smaller
ones. Taking the two blockers first, because both were wrong in ways the
existing tests could not catch.
**Adopting the PTY's rows put the CLI's input line off-screen.** A phone that
took a desktop's 43 rows into a viewport with room for 18 painted an
`.xterm-screen` far taller than its container; xterm's own viewport then had
nothing to scroll, so the bottom of the frame sat below the container with no
gesture able to reach it. Output visible, typing invisible, for as long as the
desktop kept the claim hot. `reconcilePtyGeometry` adopts COLUMNS ONLY now:
width is the axis Ink's wrap and `eraseLines` arithmetic depend on, and keeping
the local row count keeps the composer at the bottom of a viewport that
scrolls. Measured at his geometry — a 360x300 container against a 198x43 pane
now keeps 13 rows, takes 198 columns, paints 202px into a 210px container, and
the input line is inside the box.
**`capture-geometry-retry.browser.test.ts` failed, and CI could not see it**
because the file is in `BROWSER_TEST_GLOBS`. Its premise WAS the clamp —
`getTerminalDimensions()` floored while `fitAddon.fit()` did not — which this
work removes at the source, so it can never hold again at any viewport. The
case survives on its own terms: a pane already drawing at the requested size
must not be replayed. Its premise is now the #464 invariant itself, that the
floored report and the terminal agree, which is a stronger guard because the
clamp coming back fails it here rather than silently restoring the replay loop.
The helper docblock that repeated the old premise is corrected too.
**A session with no pane reported 120x40 and the client adopted it.**
`resize()` writes `_ptyCols`/`_ptyRows` only when `ptyProcess` is set and
nothing seeds them from the spawn geometry, so a dead-pane session still held
the constructor defaults — clicking that tab resized the browser terminal to
120x40 and, on anything narrower, claimed another device owned the pane when
none existed. `Session.ptyGeometry` returns null without a pane, the HTTP route
answers `{}` and the socket sends no frame at all. The raw `ptyCols`/`ptyRows`
getters are deleted rather than left available to be misused again.
**The 40-column floor clipped the pane with nothing able to reach it.** The
affordance keyed on a PTY mismatch, and the floor produces no mismatch — xterm
and the PTY agree throughout, the terminal is simply wider than the box. It
keys on what does not FIT now, MEASURED (`.xterm-screen` against the container,
on the next frame, because the screen takes its width with the render) rather
than derived from cell arithmetic. Measured at 360px: font 24 applies 40
columns and paints 560px, and all 200px of the overhang is reachable.
`.pty-oversized` is renamed `.term-overflows-x`, because after this the old
name describes only one of the two causes.
**"Scroll sideways" did not work on touch for the sessions it targets.**
`touch-action: pan-x` is cancelled before it starts by the `preventDefault()`
`touchstart` calls on every 'content' tap. The terminal's own touchmove handler
pans the container now, with the axis locked once per gesture so a diagonal
cannot pan and scroll at once, and the CSS grants no `touch-action` at all —
handing the browser a pan AS WELL would move the pane twice for one finger on
the taps where that preventDefault does not run. Measured under real touch
dispatch: a 140px swipe reaches `scrollLeft` 140 where it reached 0 before, the
buffer does not move with it, and a vertical swipe still scrolls the scrollback.
Three defects in the above, found while checking it rather than by being told:
- `canPanHorizontally` first tested `scrollWidth > clientWidth` alone, which is
true of a container that is not a scroller — a sideways swipe would have
locked the axis, done nothing, AND suppressed the vertical scroll it should
have been. Gated on the class as well.
- The notice advised scrolling sideways whenever the PTY was wider, including
when it still fitted and nothing scrolled. It is gated on measured overflow,
and on a comparison against the width this container WOULD request rather
than the one it currently holds — once adopted those are equal, so the second
question answers itself false while the condition is still true.
- `_syncTerminalOverflowAffordance` could throw out of `document.getElementById`
before reaching its try block. It runs off every geometry change, so a
cosmetic affordance could have taken the resize down with it.
The smaller items:
- `docs/architecture-invariants.md` no longer explains the equality guard as a
clamp signature; it records what the clamp used to do and why it cannot any
more. Edited by hand — that file is outside the Prettier glob, and letting
Prettier near it rewrote eleven unrelated emphasis markers.
- `throttledResize`'s HTTP fallback reads the reply. It is the path where a
declined resize is least likely to be noticed, because no socket means no
`{"t":"zc"}` frame either.
- The changeset covers the whole release: the geometry work, the queued replay
clear, the renderer watchdog, the body-covering fetch deadline, the WebSocket
output-gap reconcile, the build-generated service-worker precache and
per-build cache key, and the crash-trail hygiene.
- `@xterm/headless` is declared in the root devDependencies instead of being
reached through workspace hoisting.
- The output-gap marker is cleared after any response arrives, not only when
the capture was non-empty: a server that answers with an empty capture HAS
reconciled us, and leaving the marker set refetched on every reconnect.
- `e587d845`'s message claimed a test asserted the failed-load copy against the
built asset. It did not — that assertion lived in a probe deleted with the
other scratch scripts, so the claim was false when it was written. There is a
real test now, and it reads the source rather than `dist/`, because `dist/` is
not committed and a test that skips when it is absent would pass for the wrong
reason in CI.
`Session.ptyGeometry` gets behavioural coverage against the real class in
`session-resize-arbitration.test.ts` rather than a source guard, including the
contrast — a pane that does exist still reports, and still follows a resize —
so "always null" would fail it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copying a paragraph out of a Claude Code or Codex pane puts that pane's own
two-column transcript gutter on the clipboard, so every pasted line arrives
indented. #451 shipped the trailing half of the copy clean and left the leading
half out, because deriving the width from the selection fires on 73% of ordinary
indented text and cannot tell a margin from content.
The width is DECLARED rather than derived. `capabilities.transcriptGutter` on
the CLI registry is a bounded integer; claude and codex each declare 2, measured
on live panes, and no other stock entry declares any, so a CLI whose transcript
layout nobody has measured is never touched. The server publishes the map as
`window.__codemanTranscriptGutter`, built by filtering `enabledClis()` on the
capability rather than by listing ids, and `_activeCliGutterColumns()` looks the
active session's mode up in it. The copy path reads no terminal buffer at all.
The declared width is a CEILING, not the answer: `clean()` strips the lesser of
it and the run every selected line shares. A block can therefore only shift as a
unit, the structure inside a selection survives by construction, and a selection
reaching column 0 loses nothing. That is what keeps a `git log` body at its own
four-space indent inside an agent's two-column gutter.
Codex was measured separately, because it renders nothing like Claude: it draws
boxes narrower than the pane and pushes its transcript into ordinary scrollback.
On a live 0.154.0 answer its `•`/`›`/`⚠` markers sit in the gutter, prose
continuations sit at 2, and a nested YAML block the model wrote rendered at
2/4/6/8 for its own 0/2/4/6. Replayed at 100, 120, 160, 198, 235 and 282 columns
its indents were 0, 2, 4, 6 and 8 at every one, never 1. Copying that YAML out
of a live Codex pane now yields 0/2/4/6: gutter gone, nesting intact.
Two derived versions were built and measured first, and both are recorded in the
code because both looked correct:
- Painted trailing padding — a full-screen TUI writes real spaces across the
unused part of a row, a shell leaves them never-written for xterm to trim —
has no false positives and never over-stripped. It is also a function of pane
WIDTH: the padding exists only while a rendered line stops short of the CLI's
own layout width, and Claude's prose wraps to fill it. Dragging the same two
prose rows of one live transcript at five window sizes, the share of padded
rows ran 44%, 6%, 6%, 7% and 87% at 123, 160, 198, 235 and 298 columns, so the
strip silently did nothing at every ordinary size while a corpus captured
entirely at 282 columns said it worked.
- Taking the narrowest indent on the rows around the selection fires at every
width and over-strips about 1% of selections, because a file listing inside
the transcript can be the narrowest thing on screen.
Measured over 1,392,281 selections — every 1, 2, 3, 5, 10 and 20-row window of
real Claude screens replayed from live PTY streams at 100, 120, 160, 198, 235
and 282 columns — the declared width over-strips none, breaks no relative indent
and alters no text, and serves 100% of the selections whose own indent covers
the gutter. Verified end to end in a browser with a real mouse drag and a real
Ctrl+C: Claude and Codex panes paste flush at 123, 198 and 298 columns, a shell
pane is untouched at every one.
The strip sits behind `copyStripMargin` (App Settings, Selection & clipboard),
per-device and default ON: a display key, absent from the .strict()
SettingsUpdateSchema, read as `!== false` because the desktop branch of
getDefaultSettings() returns {}. The toggle is checked before the map.
Two review findings from #451, handled:
- The mid-row flag governs ONE line now. `range.start.x > 0` excludes only the
first selected line, the one whose margin the mousedown genuinely cut off, so
the same three rows no longer produce three different clipboard results.
- The reversed-drag finding does not reproduce on the pinned xterm.
`getSelectionPosition()` reads `_selectionService.selectionStart`, whose
getter returns `SelectionModel.finalSelectionStart`, and that swaps the pair
when `areSelectionValuesReversed()` says so. A real upward mouse drag through
chromium against xterm 6.0 reports the same range as the downward drag.
`_normalisedSelectionRange()` keeps the ordering as a guard, because the model
one layer down exposes the unnormalised fields under the same two names.
Tests: test/terminal-copy-clean.test.ts (64, up from 31), plus the injected
script stripped in test/server-index-title.test.ts. Every guard is pinned:
removing any one of seven reds at least one test, including declaring the wrong
gutter width. Full suite green, 7,861 passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A third pass in a real browser, at the widths this app actually renders at.
The notice a failed history load writes into the blanked pane was one
70-character sentence. At 430px that exactly filled the line; at 320px it
wrapped and left a lone '.' on a line of its own. The floor this app will
render at is 40 columns — reachable today by raising the font on a phone — so
the notice is three lines now, none over 25 columns, one fact each: what
failed, that the session is still alive, and what to do.
It says RELOAD rather than "reopen the tab" because `selectSession`
early-returns when the session is already active, so clicking the tab you are
already on retries nothing. The earlier wording named no next step at all,
which left a mostly-empty terminal and no way out of it.
CLAUDE.md no longer cites "758px reachable to the right" as evidence: that
figure is a property of the test content, not of the fix, and the file's value
is that a reader can trust a claim without re-deriving it. What is pinned
instead is the invariant that survives any content — the full pane width is
reachable, and removing the class returns scrollLeft to 0, so a resolved
mismatch cannot leave the pane parked off-screen.
Verified at 430, 360 and 320px against the shipped bundle, with the test
asserting the built asset carries the copy so an edit that never reached the
build fails rather than passing on the source's wording.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The gate caught fourteen failures the focused tests could not: every harness
that builds a partial app out of cherry-picked mixin methods, and every source
guard that named `fitAddon.fit()` by hand.
Most are wiring — `syncTerminalGeometry`, `_refitAfterCellSizeChange` and
`_resizeTerminalTo` added to the fakes so the real chain runs rather than a
stub of it. `file-browser-search` is the one that shows why it matters: without
the method on the fake, selectSession's unconditional call threw into its own
catch and every later assertion in the file measured a load that never
happened.
Two are not wiring.
`detached-session-pane-sizing` pinned the behaviour this change deliberately
reverses. It asserted the LOCAL fit still runs for a session owned by its own
window — "withhold the send, never the reflow" — so the assertion is restated
rather than patched, with the reason beside it and in the file's docblock: a
reflow the PTY is never told about leaves this xterm rendering a CLI's frames
against a shape that does not exist, and the popup that owns the pane is
drawing for its own width regardless. The old rule bought a garbled frame, not
a correct one.
`mobile-prompt-composer` sliced `_cleanupSessionData` as a fixed 1200-character
window, so the assertion depended on how much unrelated code sat above the line
it cared about. It reads the whole method now.
`terminal-scroll-intent` records `syncTerminalGeometry` rather than `fit`,
under its own name: recording a bare fit there would name the very thing the
subject was changed to stop doing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Issue #464, "text gets muffled sometimes, in both TUI default and fullscreen".
The screenshot is not a dropped frame or a frozen renderer — it is arithmetic.
Claude Code's TUI wraps its frame at the width the PTY reported and erases the
previous frame by walking the cursor up the rows it believes that frame took. A
browser terminal of a different width makes each logical line occupy more
physical rows than Ink counted, so `eraseLines(n)` clears too few and the new
frame paints over rows nothing erased: doubled lines, and short tool summaries
sitting inside longer prose rows with the prose's tail still visible.
Reproduced against this repo's own xterm before changing anything — a 120-column
PTY against a 62-column terminal renders every wrapped line twice. `test/
terminal-pty-geometry.test.ts` pins that, and pins the clean render at matching
widths beside it, so the assertion cannot be satisfied by code that fixes
nothing.
Four ways the two drifted apart, none of them observable from either end:
1. `fitAddon.fit()` resizes xterm to `proposeDimensions()` RAW while every
server-facing path reported those floored at 40x10. Measured in Chrome at
430px: font size 44 proposed 13 columns, the server was told 40, and xterm
stayed at 13. Three call sites each did their own fit-then-floor, and two
re-read the proposal after the fit — `_shrinkPaddingToFit()` runs exactly
there, so the container had moved.
2. `throttledResize` (keyboard up) and `sendResize` (session detached into its
own window) reflowed locally and withheld only the SIGWINCH. That is the one
combination that cannot be right: a reflow nothing is rendering for buys
nothing and costs correctness. Both now withhold everything, and the
keyboard's settle timer still sends the one resize that stops the PTY going
stale.
3. `setFontSize`/`setFontFamily`/`setFontWeight` move the cell size — a geometry
change — and told the server nothing at all, so raising the font on a phone
left the CLI wrapping at the old column count.
4. `Session.resize` DECLINES a small-viewport request while a desktop connection
holds an active sizing claim, and said nothing, because resize was write-only.
`syncTerminalGeometry()` is now the one function that may change the terminal's
size: it fits, floors and applies as a single step, so the numbers xterm holds
are the numbers the server is told. A test sweeps every module for a bare
`fit()` on the main terminal, and finds exactly one — the owner's own.
For (4) the client cannot win, so it is told the truth instead: both transports
answer a resize with `session.ptyCols`/`ptyRows` (`{"t":"zc"}` on the socket,
the body of the resize POST) and `_onPtyGeometryReport` adopts them. A terminal
that keeps a shape the PTY refused does not render "too narrow", it renders
garbled. Adopting can leave the pane wider than the screen and the container is
`overflow: hidden`, so `.pty-oversized` grants horizontal reach for exactly as
long as the mismatch lasts: correct-and-reachable beats correct-and-clipped
beats garbled. That rule sets both overflow axes and its own `touch-action`
because mobile.css loads later and sets `.terminal-container { overflow:
visible; touch-action: none }` — a bare `overflow-x` would leave overflow-y
computing to `auto` and hand the browser a vertical scroll container the
terminal's touch handler knows nothing about.
Verified in Chrome at 430px against a live server, with a desktop client holding
the claim: the phone adopts 198x43, gets `overflow-x: auto` / `overflow-y:
hidden` / `touch-action: pan-x`, 758px of reach to the right, and keeps its own
vertical scrolling. The pre-fix build was measured in the same harness for the
control.
Two things this deliberately does not do. It does not change who owns the pane
size — the desktop still wins, and `_startMobileResizeRetry` still takes it back
once that goes idle. And `throttledResize` still holds the PTY's shape for the
whole keyboard animation rather than sending a SIGWINCH per step; that decision
predates this and was not re-tested here.
Also in this commit, Ark0N's third-pass review items on #431:
- The response viewer's byte-buffer fallback and `_onSessionClearTerminal` both
used the no-param `/terminal` form, capped only by `terminalBufferMaxBytes`
(32MB) — the largest body the frontend asks for anywhere. One carried no
deadline at all and the other got the 15s tail budget. Both now take the
full-history budget.
- A `?full=1` capture that outruns its deadline falls back to the bounded tail.
The pane is blanked before that fetch, so an abort used to leave a black
rectangle, discard the queued live output and never reach `_connectWs`. A
failed load now still opens the socket, says one dim line where the content
would have been, and clears the tab's spinner — which nothing did, so a failed
select left `aria-busy="true"` set forever.
- `_wsOutputGapSession` is cleared at the repaint that settles it, not in a
`finally` that also ran on the catch. A reconcile that threw, or hit the new
deadline — the flaky link the marker exists for — dropped the gap with nothing
to retry it. `ws.onopen` no longer clears it up front either.
- The replay-clear invariant is pinned in the gate, which is the drift this PR
exists to fix: `_resetTerminalForReplay` must be a queued write and nothing
else, and no module may blank the terminal with a `clear()+reset()` pair.
- `DIAG_ENTRY_MAX_CHARS` replaces the hardcoded 300, bound through a local
first: `CodemanDiag?.x` still throws a ReferenceError when the identifier was
never declared, and that is the one function in the app that must not throw.
- panels-ui's two kill-all clears route through the same helper, and the
xterm-version guard's comment says "resolved lockfile version" rather than
"dependency RANGE", which is what it has pinned since the last round.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four changes the maintainer asked for on Ark0N/Codeman#446 before merging.
The pane-exit watcher stays always-on, but a tick now costs nothing when there
is nothing to observe. `hasObservablePaneSession()` skips the tmux exec while
every session on the manager is one of the shapes `Session.paneExitApplies`
already forces to UNKNOWN: a remote SSH session (its local pane holds the ssh
client), a docker case (a `docker exec` into the container's own tmux), and a
record rebuilt from the socket (no provenance at all). The timer is untouched.
Skipping retracts nothing, for the same reason a failed read does not: the map
still holds the last real reading, and every path that puts a new command in a
pane calls `clearPaneExit()` itself. The two copies of that rule are pinned
against each other in `test/session-pane-exit.test.ts`, because drift between
them is silent in both directions.
`DEFAULT_PANE_EXIT_INTERVAL_MS` was already a constant beside the stats and
remote-reconnect intervals; its comment now says why the watcher owns its own
cadence and why the number is what it is.
The never-default-an-absent-status rule is written where `PaneExit` is declared.
It names `status ?? 0` as the thing never to write, and says that an agent the
OOM killer took would otherwise read as a user typing `/exit` — which is what
absent-stays-absent keeps a later clean-exit sweep away from. Nothing fails when
somebody adds that `??`, which is why the sentence is there rather than a test.
Checking the dot's specificity found a second fight, and it was losing. On the
tab strip the alert rules win as intended: a session that exits with a
permission dialog pending still renders red, and yellow for an idle alert. On
the rich vertical tab rail they did not — that rail's own `tab-state-*` dot
rules are (0,9,1) against the strip's mute at (0,5,0), so an exited session
there kept a full green dot AND the working halo beside a badge reading
"exited". The rail twin matches that specificity exactly and therefore must stay
below those rules in source order; it clears the halo as well, which the strip's
rule never had to think about.
`test/session-pane-exit-ui.test.ts` now resolves the real stylesheet in jsdom
rather than matching selector text: postcss collects every rule that paints
`.tab-status`, a real engine decides, and the tests read back the answer. Two
mutations were run against it to prove it has teeth — dropping the hand-written
alert exclusions fails three cases, and moving the rail twin above the state
rules fails one.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Codex states background work too, and it says so in a different place. Claude
writes `· 1 monitor ·` on the last row of the screen; Codex pins
`1 background terminal running · /ps to view · /stop to close` ABOVE its
composer, which puts that row third from the bottom once the status line and
the composer are counted.
So how far up the screen to look is now per-CLI data as well:
`capabilities.workDetect.watchingLines`, bounded to 1..8 by the schema, and
defaulting to Claude's two. That bound is the point. The window is half the
injection guard, since every row it adds is another row the agent itself may be
able to write, and the label is what silences an idle alert. The other half is
the anchor, and Codex's is ` · /ps to view`: chrome naming a slash command only
the CLI can offer, so a session that writes "I left 1 background terminal
running for you" into its own output matches nothing.
Measured against a live codex-cli 0.154.0 pane rather than read out of a
binary. The row appears when the terminal starts, follows the composer down as
the conversation grows, and is gone after `/stop`. Verified end to end on an
isolated beta: the session payload carried `watching: "1 background terminal"`
and the badge rendered with it, and both cleared when the terminal stopped. The
fixtures in the tests are that capture verbatim.
Codex has no hook signals, so no idle prompt and no false NEEDS YOU row: for a
Codex session this is the badge alone, which is the case the maintainer said a
registry field could cover and a hook never could. Cross-CLI tests pin that
neither pattern fires on the other's screen, and that a CLI declaring nothing
still reports nothing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review fixes. Two of these are defects in the previous commit.
1. The fetch deadline only covered time-to-headers. `await fetch()` settles on
response headers, so clearing the abort timer in a finally around it left the
body — the multi-megabyte `?full=1` capture the deadline exists for —
completely unbounded; it only ever bounded a server that accepts a connection
and never replies. Measured against a server that sends headers immediately
and stalls the body 4s under a 1s deadline: fetch resolved at 30ms, timer
cleared there, body completed at 4026ms unaborted. Now the body is read
inside `_fetchTerminalCapture`, which returns {json, headers, headersAt} —
headers because two callers read server-timing, headersAt because those same
callers measure header-vs-body time and can no longer observe that moment.
`_terminalCaptureInflight` is scoped the same way, so a body still streaming
counts toward a capture starting beside it. Same test now aborts at 1005ms.
2. The precache could never be hit, and the previous commit made that expensive
rather than free. `renderIndexHtml` runs `cacheBustAssets`, which appends
`?v=<mtime>` to every same-origin .js/.css reference INCLUDING content-hashed
names — confirmed against a running instance:
`vendor/xterm-zerolag-input.6fee72f2.js?v=1789402869101`. `caches.match` is
query-sensitive, so entries keyed on the bare hashed path were unreachable;
deriving the list from the manifest turned cheap 404s into ~1.3MB downloaded
at every install that nothing could read back, once per deploy now that
CACHE_NAME rotates. The fallback match takes `{ ignoreSearch: true }`, which
also lets runtime-cached entries survive an mtime change.
3. `_wsOutputGapSession` was only cleared in ws.onopen, so paths that already
repaint the buffer left it set and the socket replayed everything a second
time. `selectSession` loads the buffer and only THEN calls `_connectWs`, so
neither the _isLoadingBuffer nor the _terminalRefreshOwner guard applied.
`_markTerminalBufferReconciled()` is now called from _onSessionNeedsRefresh's
finally, from selectSession after its load, and from _cleanupSessionData.
The scope claim was also wrong and is corrected in the comment: when the
network drops, SSE drops with it and handleInit's keepTerminal branch already
reconciles. The genuinely uncovered case is the WS dying while SSE stays up,
where _onSSETerminal discards SSE terminal frames until _wsReady flips in
onclose — up to the ping+pong window of output nothing writes.
4. CLAUDE.md said "all of them measured rather than reasoned", which the PR's
own "not verified" section contradicted. Split explicitly: the replay race is
measured, the watchdog mechanism is verified against xterm 6.0.0 under jsdom
(field path resolves, a forced stale handle makes refreshRows a no-op, the
kick schedules a fresh frame), and the iOS rAF-discard premise is reasoned
and still wants a device. Adds the two missing entries — the WebSocket
reconcile and the sw.js/build.mjs "keep these in sync or the build throws"
contract.
Also: test/xterm-private-api.test.ts pins the RESOLVED lockfile version instead
of the declared `^6.0.0` range, which was the wrong assertion in both directions
— a real upgrade to 6.4.0 can rename a private field while resolving inside the
range, and an innocuous range edit failed while changing nothing installed. And
test/sw-precache-manifest.test.ts now parses HASHABLE out of scripts/build.mjs
rather than hand-copying it, which was the same drift this PR exists to fix; the
parse is guarded against silently matching nothing.
The deadline fix has a behavioural test against a real socket plus a source
guard asserting `await res.json()` precedes the finally — verified to fail when
the helper is reverted to the old shape, so it is not vacuous.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four ways the terminal can silently stop being correct — in each case the
buffer keeps updating, nothing throws, and the only recourse is a reload.
1. Renderer freeze after backgrounding. iOS DISCARDS scheduled rAF callbacks
when a PWA backgrounds, and xterm's RenderDebouncer only clears its
`_animationFrame` handle from inside that callback — so one drop leaves it
permanently set and every later refresh() early-returns. Parsing is
decoupled from rendering, so bytes keep filling the buffer correctly while
nothing paints. Codeman has exactly ONE xterm for the whole page load, so a
single backgrounding wedges it until a reload. Adds a 2s liveness poll and
`_kickRenderer()`, which does what the dropped `_innerRefresh` would have.
2. Replay clears raced live output. xterm's write() is async-queued while
reset() is synchronous and, per upstream, "does not clear input buffers and
does not reset the parser" — so bytes queued before a reset are parsed after
it and fuse into the snapshot. Verified against the real xterm 6 here:
write('p8'); reset(); write('rmissions') renders "p8rmissions". The main
path was already safe via a queued erase; the needsRefresh and clearTerminal
paths were not. All three now share one queued `\x1bc` (RIS), which unlike
3J/H/2J also resets modes, charsets, scroll regions and SGR state.
3. Output lost on WebSocket reconnect. Input frames carry seq+cid and are
delivered exactly once; output frames carry nothing. ws.onopen re-sends dims
and flushes queued input, and needsRefresh only fires on external-CLI
startup and SSE backpressure drain — never on reconnect. Output produced
while offline was simply absent afterwards. Interim fix: reaching onclose
means the drop was unintentional, so the session is marked and the next open
reconciles from the server buffer. Sequencing output is the follow-up.
4. Terminal captures had no deadline. No AbortController anywhere in the
frontend, including `?full=1`, which the code itself calls "unbounded-ish
work: at the default history limit it can be megabytes". Adds a budget that
scales with full-vs-tail and with captures in flight, degrading to a plain
fetch where AbortController is missing.
Also: the service-worker precache was dead — the build content-hashes assets
but sw.js listed pre-hash names, so 15 of 23 entries 404'd (verified against a
running instance) and cache.add().catch() hid it. Offline still worked via
runtime caching, but CACHE_NAME was a constant so activate's cleanup never
deleted anything and every past release's assets accumulated. Both are now
derived from the build manifest. Crash-trail entries are flattened and capped,
since they are joined with \n into one value and one call site interpolates a
server-controlled WS close reason.
The watchdog reads xterm privates — there is no public API. Every access is
optional-chained so a shape change degrades to a no-op. `_renderService` only
exists after open(), which needs a real DOM, so the gate cannot assert the
field path; test/xterm-private-api.test.ts pins the dependency range instead.
Tests: 23 new (terminal-resilience, sw-precache-manifest, xterm-private-api),
all pure/static so they run in the gate, which excludes the mobile suite. One
static source guard in history-truncation-notice updated for the renamed call;
the behaviour it pins is unchanged.
Not verified: no browser available, so no runtime reproduction of the freeze
and no real-device test of the reconnect path. Both warrant a device pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The badge alone left the row in NEEDS YOU, which is the thing the issue was
about. The fix is the alert that does not fire.
An idle prompt from a session that is watching its own background work now
opens ALREADY acknowledged. `hook-event-routes` passes `Session.watching` to
`notePrompt()`, which sets `acknowledgedAt` and records why in a new
`acknowledgedReason`. Nothing new suppresses anything: `acknowledge()` has
always meant "the alert this prompt armed is spent", and the prompt itself
stays pending, answerable and available as Read My Mind context. A wrong label
therefore costs a card that does not blink, never an alert that was never
created.
Every surface follows from that. The broadcast carries the reason, so a live
page declines to arm the tab alert and raises no desktop notification. The push
is skipped, since a false alarm is hardest to ignore on a phone. A reloading
page reads `acknowledgedAt` in `seedApprovals()`, which it already did. And
`classifySession()` now reads it too, which is a pre-existing bug fixed here:
acknowledging on one device cleared the alert everywhere except `codeman tui`.
It re-arms for free, because the next idle prompt supersedes the item and is
built fresh. Only `idle` is eligible, so a dialog that blocks the agent still
goes red whatever else it started.
The label is pane-derived and therefore prompt-injectable, so it is now read
from the last two rows of the screen only, with Claude's pattern anchored on
the `·` its footer joins items with, ANSI-stripped and length-capped at the
source. An agent that prints `· 1 monitor ·` into its own output finds no
match.
Verified on an isolated beta: a session that armed a monitor took its idle
prompt acknowledged with no alert on any surface, wore the badge, and showed
"quiet, watching 1 monitor" on its still-answerable card; the same session with
the monitor killed alerted normally on the next prompt. `test/watching-no-alert.test.ts`
pins both directions across all four surfaces.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A single pin that failed its transcript gate returned the options untouched,
so `resumeSessionId` fell back to `_resumeSessionId` — undefined for an
ordinary session — and the renderer emitted the bare
`claude --dangerously-skip-permissions --session-id "<this.id>"`. Every
session prompted before its first `/clear` owns a transcript under that id,
so the dropped pin handed back exactly the refusal this branch removes, with
no `||` branch to catch it. It was also a regression against master on the
`restartCli()` path, which pinned `_claudeSessionId ?? this.id` and, since the
constructor seeds that field, could never land unpinned.
The pin now walks three candidates in priority order — the conversation
chain's tail, the launch seed, then the session's own id — and takes the first
one a transcript backs. A candidate that misses is passed over rather than
ending the walk.
Falling off the end pins nothing, which also settles the second half of the
problem: the old code skipped the transcript check whenever the pin was the
session's own id, so a genuinely new pane rendered the two-branch form after
all. That costs a brand-new session claude's "No conversation found" line in
its scrollback, and `wrapWithNice()` prefixes only the first branch of the
rendered `a || b`, so the branch that actually runs loses its priority for the
life of the session. With no transcript anywhere the bare `--session-id` is
the correct command, so the comment claiming an unchanged shape is now true.
The transcript lookup reads the server process's own `CLAUDE_CONFIG_DIR` when
a session declares none. A pane inherits the server environment through tmux,
so on an install that exports it the CLI writes its transcripts there and
every lookup under `~/.claude` was a false negative — which under the old code
meant the colliding command. `claudeCredentialsPath()` and
`realClaudeConfigDir()` resolve the same directory the same way. The header
sentence calling a skipped resume "the safe direction" described the opposite
of what happens at this call site, and says so now.
The create-path fallback writes `_resumeSessionId` alongside the create
options. That branch leaves `isRestored` false, so `_claudeSessionId` is
recomputed from the launch fields and settled on `this.id` while the CLI
resumed the chain tail; the response viewer, Read My Mind and the unified-list
alias map read that field until the next first-hand hook.
Four new tests: a chain tail with no transcript while the session id has one,
no transcript anywhere, the create path's alias, and the process-env lookup.
All four fail against the previous commit. Two existing tests move with the
gate — the guess-refusal test now backs the session's own id, and the
custom-model restart test gives its working pane the transcript that makes
`--session-id` collide in the first place, alongside a new one pinning the
no-transcript case.
CLAUDE.md described the pin as a `restartCli()`-only thing sourced from the
live conversation id. All three halves of that moved here.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An agent that arms a monitor, backgrounds a shell or hands a task to a
cloud session is told to end its turn. The pane then falls quiet, Claude
Code's idle_prompt notification arrives a minute later, and every surface
files the session under NEEDS YOU with nothing for a human to answer.
Claude states what it is still running on the last row of its screen
(`⏵⏵ bypass permissions on · 1 monitor · ← for agents`). That row is now
`capabilities.workDetect.watchingLine` in the CLI registry, guarded by
compileVersionRegex() like every other config regex, and the idle probe
reads it off the capture it already takes: `watchingLabel()` in
session-activity.ts searches the last five lines only, so a session that
PRINTS "1 monitor" is not mistaken for one running it.
The label lands on Session.watching and rides toLightDetailedState() out
to every surface. The phone overview, the desktop home rail and the rich
sidebar rows wear it as a `watching` badge in the accent colour, beside
the state pill and never in place of it: an agent can arm a monitor and
ask a question in the same breath, and only the pill says which.
Verified end to end against a throwaway session on an isolated beta
instance: the payload carried `watching: "1 monitor"` once the turn
ended, the badge rendered next to a yellow `waiting` pill, and both
cleared when the monitor died.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A CLI that launches with `--session-id <id>` refuses an id that is already in
use (claude: `Error: Session ID ... is already in use.`), and every session
whose agent has been prompted owns a transcript under that id. The dead-pane
respawn in `_setupOrAttachMuxSession()` passed the bare launch line, so
recovering such a session relaunched a CLI that died on startup, the pane went
dead again at once, and the conversation was stranded behind a tab that looked
merely idle.
`restartCli()` has pinned a resume id against this since the custom-model work,
and its comment states the assumption that made the other path look safe:
"Unlike the dead-pane respawn, this one kills a WORKING pane whose conversation
already has a transcript". A pane whose agent exited has a transcript too.
Both relaunch paths now build options through
`_buildRespawnPaneOptionsWithResumePin()`, and so does the create-path fallback
after a failed respawn, which otherwise met the same refusal that made it the
fallback. Four gates guard the pin, each standing for a way of resuming the
WRONG conversation or of making a working relaunch fail.
A remote or docker session is never pinned. Unlike `restartCli()`, whose route
refuses both, the dead-pane respawn is reached by every session shape. Their
pane commands already render a self-healing `--session-id || --resume`, and
both flip to resume-first once the resume id differs; the conversation lives on
the far side, so a local id resolves to nothing there and the `--session-id`
fallback then collides with the transcript the far side does hold.
The id comes from the conversation CHAIN rather than `_claudeSessionId`, which
also holds history-correlated guesses keyed on the working directory.
`_recordClaudeSessionInChain()` refuses those so they cannot "write a foreign
conversation into this pane's permanent record", and launching from one is
worse than the display bug that rule prevents. The chain tail also outranks the
launch seed, which is written once at construction and never moves off a
`/clear`.
A pin no transcript backs is dropped, because the fallback branch keeps
`--session-id <this.id>` and would collide. A synthetic `restored-<fragment>`
id from socket discovery is dropped too, and logged: it fails claude's `uuid`
token pattern, so the renderer would emit the unpinned command while the caller
believed otherwise.
Tests cover each gate and the rendered command. Four of them fail against the
unfixed source; the remote and docker ones were separately checked against a
build with only that guard removed, since they pass on master for the wrong
reason.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Same guard as Start-Codeman.sh's own (docs/docker-self-update.md-adjacent
incident, 2026-09-21): docker-compose.yaml hard-codes `name: codeman`, so a
second checkout run without COMPOSE_PROJECT_NAME resolves to the SAME
Compose project as any other checkout on the host. It has to live here too,
not just in Start-Codeman.sh: this script's own --no-cache build and its
`down`/`down --volumes` both run BEFORE the handoff at the bottom of the
file, so Start-Codeman.sh's copy of the guard would only fire after this
script's own destructive calls already ran — and its default `down
--volumes` is more destructive than Start-Codeman.sh's own targeted
refresh, clearing every named volume the resolved project has.
Also fixes a real bug the same guard shipped with: under `set -o pipefail`,
`grep -v` legitimately exits 1 when nothing survives the filter (the
ordinary, no-collision case), and without `|| true` on the pipeline that
non-zero status propagates through the command substitution and `set -e`
aborts the WHOLE script at the guard — every time, collision or not. Caught
only by actually executing the guard end-to-end against a stub `docker`
(the existing smoke-test harness), never by a static text/regex check on
the source; the stub's `config --format json` response was also fixed to
pretty-print like real Compose does, since a compact one-liner silently
resolved project_name to empty and exercised neither script's guard the
way production output does.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
The tab now reads "exited (137)" beside the session name, drawn from the
`paneExit` field the server publishes. `applyPaneExitBadge()` owns the DOM
work, called from the incremental render path — the only path a live session
ever takes, since going from live to exited adds and removes no tab and so
never reaches the full rebuild.
An unknown answer draws nothing. A death tmux could not explain reads "exited"
with no number rather than "exited (0)", so an unexplained death and a clean
exit do not look alike. A signal death reads "exited (signal 9)".
The badge carries `data-i18n-skip`, like the status pills: it is generated
text, `i18n.js` walks inserted content, and a dictionary entry added later
would fight the renderer, whose in-place comparison is against English.
The tab also carries a `tab-agent-exited` class that mutes the status dot. That
dot is drawn from `status`, which stays `idle` or `busy` for an exited pane as
the issue requires, so without this a green or pulsing dot sits beside a badge
saying the agent is gone — the first thing a tester asked about. `status`
itself is untouched, so this is a rendering rule only. The CSS excludes the two
alert classes by hand, following the convention the rich-rail dot rules
document: a dot turning red or yellow because a session is blocked on a human
outranks "the agent exited".
The tab keeps its click behavior. X still closes it, and nothing here closes,
sweeps or restarts anything.
`docs/architecture-invariants.md` gains the mechanism under "Session data and
lifecycle", where every comparable one already lives: what the tri-state means,
the four shapes it is absent for, why the watcher cannot ride the stats
collector, why an absent `#{pane_dead_status}` is not 0, and the three things
that must never happen to an exited pane.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mux layer now knows a pane's agent has exited. This puts it on the session
record, where the board and, later, the reboot restore can see it.
`SessionState.paneExit` carries `{ status?, signal?, at }` and rides the
existing `session:updated` broadcast through `toState()`. No new SSE event. The
server pulls each answer from `mux.getPaneExit()` rather than off a broadcast
payload, so the raw reading never reaches a browser: for a remote or docker
session that reading is the death of an ssh client or a `docker exec`, not of
the agent.
The field is tri-state, and the third state is its absence: `undefined` means
Codeman does not know, and it never reads as alive. `Session.setPaneExit()`
forces that unknown for every shape a dead local pane does not describe. A
direct-PTY session owns no pane. A remote SSH session's local pane holds the
ssh client, whose death means a transport drop OR an exit, which is the
ambiguity PR #355 settled by not guessing. A docker case's local pane holds a
`docker exec` into the container's own tmux. And a session rebuilt from the
socket has no provenance at all: `reconcileSessions()` gives it a synthetic
`restored-<fragment>` id that matches no `state.json` entry, so a remote
session rediscovered after `mux-sessions.json` was lost arrives with no
`remote` field and looks local — `MuxSession.discovered` marks it, and absent
metadata there counts as unproven rather than as proof. The scoping lives on
`Session` rather than in `TmuxManager` so there is one copy of the rule.
`status` and `pid` are untouched. `status: 'error'` belongs to the PTY-exit
circuit breaker and makes the browser offer a restart, and a null `pid` is what
makes the browser re-attach and launch a fresh CLI. A reading that repeats the
previous answer writes nothing and broadcasts nothing.
An unknown answer never reads as alive, but a stale KNOWN one would keep
reading as exited, so `clearPaneExitForNewPane()` retracts it on every path
that puts a new command in the pane: the start/attach path, the `restartCli()`
relaunch behind a custom-model switch, and the remote reattach. Without the
second of those, switching an endpoint on an exited session launched a new
command and then persisted and broadcast the old exit straight back onto it.
`toState()` is also what `state.json` persists, so the record survives a
reboot, which is the only thing that does: a reboot takes the tmux server, and
with it every live signal and every `mux-sessions.json` entry. Nothing reads it
there yet — making the restore refuse such a session is a behavior change that
belongs with the part that closes them. Recovery threads the saved value back
through the constructor so the first persist after boot cannot blank it.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Codeman creates every tmux pane with `remain-on-exit on`. When the agent exits,
tmux keeps the pane, the tmux session, and the `tmux attach-session` process
Codeman records as the session's pid, so no PTY exit handler fires and nothing
writes the exit down. tmux itself knows: it marks the pane dead and reports the
exit status. This reads that.
`PANE_LIST_FORMAT` gains `#{pane_dead}`, `#{pane_dead_status}` and
`#{pane_dead_signal}`, and `startPaneExitWatcher()` refreshes a
muxName-to-observation map from ONE batched `tmux list-panes -a` per tick. Boot
reconciliation already ran that same call, so it now fills the map too and
recovery starts with a reading.
The watcher owns its own interval rather than riding `startStatsCollection()`,
which the issue suggested. That collector is armed when a browser opens the
Monitor panel and DISARMED when it closes it, and boot skips it entirely unless
recovery found a live session, so a session created on a freshly booted server
would publish nothing and one browser could turn detection off for every other.
Measured on an isolated instance: a dead pane with status 0 reported nothing
until `POST /api/mux-sessions/stats/start` was called by hand. It is still one
batched read per tick; only the timer changed.
Three rules keep a positive answer trustworthy. A session answers only when
tmux listed exactly one pane for it, because Codeman never splits a pane and a
session the user split by hand has none that speaks for the agent. A pane
answers only when `#{pane_dead}` said 1 or 0, because an empty field is a tmux
that did not answer. An absent status stays absent rather than becoming 0:
measured on tmux 3.2a, a SIGKILLed pane reports neither a status nor a signal,
and calling that a clean exit would be wrong in the direction that matters.
Two guards stop a slow read undoing a fast one. `EXEC_TIMEOUT_MS` is 5000 ms
against a 2000 ms interval, so a read can outlive two ticks: one already in
flight suppresses the next, and a generation counter that every
`clearPaneExit()` bumps discards a read that started before a respawn or a
kill. An observation also carries its pane pid, so a second command in the same
pane that exits the same way starts a new timestamp rather than inheriting the
first death's.
A non-empty read of `list-panes -a` is authoritative for the whole socket, so
sessions missing from it are pruned, which also bounds the map as tmux sessions
come and go outside `killSession()`. A failed or empty read retracts nothing.
The manager reports the raw pane reading and applies no session-shape scoping,
because the remote-reconnect watcher beside it needs exactly that raw reading.
`parsePaneList` becomes `parsePaneRows`, returning one row per pane instead of
a name-to-pid map; reconciliation builds its map from the rows. The parser's
existing cases carry over unchanged, including the launchd/systemd literal-tab
regression from PR #71.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three blockers, all fixed and verified by actually running the script (not
just string-matching it):
1. `exec "$script_dir/Start-Codeman.sh"` failed EACCES/exit 126 on every
checkout, since Start-Codeman.sh is committed non-executable (100644) —
the same fact my own second commit on this branch established. Fixed to
`exec bash "$script_dir/Start-Codeman.sh"`.
2. `down` ran before `build --no-cache`, so Codeman and every session it was
running were offline for the entire rebuild, and a build failure left the
stack down with nothing to bring it back — the exact ordering mistake
Start-Codeman.sh's own "Build BEFORE taking the stack down" comment exists
to prevent. Reordered to build, then down, then hand off.
3. The default path could throw the rebuild away: codeman-node-modules/
codeman-dist only re-seed from the image while EMPTY, Start-Codeman.sh
only clears them when it detects the checkout's HEAD or package-lock.json
moved, and neither condition is true for the Dockerfile-only change this
script exists for — so a plain `bash docker/Update-Codeman.sh` rebuilt an
image whose fresh node_modules/dist then sat unused behind the old
volumes. Made clearing them the default; `--keep-volumes` opts out
(replaces the old `--volumes`/`-v` flag, which is no longer needed since
clearing is now the default).
Smaller items from the same review, also fixed:
- The --no-cache build now derives PUID/PGID from CODEMAN_APPDATA_PATH's
owner first, via the identical owner_of() helper Start-Codeman.sh uses
(parity-tested) — without it, the build used Compose's default 1000:1000
regardless of the real appdata owner (99:100 on the Unraid layout
docker/README.md documents), and Start-Codeman.sh's own correctly-PUID'd
build during the handoff would then rebuild those layers anyway, so the
--no-cache image never actually shipped.
- docker/README.md's "rebuilds ... only when it detects ... moved" wrongly
described BOTH the rebuild and the volume-clearing as conditional;
Start-Codeman.sh rebuilds on every start, only the volume-clearing is
conditional. Corrected, and reworded around the new default.
- --help/-h now prints usage and exits 0 instead of falling into the
unrecognised-argument branch.
- "the ONLY named volumes this stack declares" now says docker-compose.yaml
specifically, since a docker-compose.override.yml could add more.
New tests: PUID/PGID derivation parity with Start-Codeman.sh's owner_of(),
--help handling, and — the one that actually catches blocker #1, which five
source-string-matching tests did not — a real end-to-end smoke test: a
synthetic deployment, a stub `docker` on PATH logging every invocation, the
real script executed via a real subprocess. Confirms the real command
sequence (build --no-cache, then down --volumes or plain down, then evidence
the handoff genuinely ran Start-Codeman.sh) and that a working handoff fails
honestly at Start-Codeman.sh's own later check rather than with EACCES.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Start-Codeman.sh, its sibling and the script it hands off to, is itself
committed non-executable (100644) upstream — it's documented and invoked
as `bash docker/Start-Codeman.sh`, never `./docker/Start-Codeman.sh`. The
"is executable" test I'd added for Update-Codeman.sh asserted the opposite
convention, which the file correctly does not follow.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
docker/README.md and docs/docker-self-update.md both already point operators
at "stop the stack, rebuild, restart" for anything the in-app updater refuses
to apply (a changed server.Dockerfile, a changed docker-compose.yaml, or a
new required .env key) — but that was a manual, hand-typed procedure with no
script of its own, unlike every other start/update path this deployment has.
docker/Update-Codeman.sh scripts it: `docker compose down`, then an
unconditional `docker compose build --no-cache` (a major update should be
certain of what actually ships, not reuse whatever layers happened to be
cached), then hands off to the existing Start-Codeman.sh for the same
careful PUID/PGID, override-file and fingerprint handling every other start
already goes through — rather than reimplementing any of that by hand and
risking it drifting out of step.
An optional --volumes/-v flag also removes the codeman-node-modules/
codeman-dist named volumes, the scripted form of the "Resetting the build
artefacts" procedure docs/docker-self-update.md already documents by hand.
Safe: those two are the only named volumes this stack declares; application
data and case workspaces are host bind mounts, never touched by
`docker compose down` either way.
Docs updated: a "Major updates" section in docker/README.md, and a pointer
from docs/docker-self-update.md's existing "Resetting the build artefacts"
troubleshooting entry.
Tests: extended test/docker-entrypoint.test.ts (the existing home for
Start-Codeman.sh's own static checks) with a bash -n parse check, the
down-before-build-before-handoff ordering, the --volumes flag's effect,
unrecognised-argument handling, and byte-for-byte agreement with
Start-Codeman.sh's own override-file resolution logic (so `down` here and
`up` there can never target different Compose files).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Two related fixes, found while re-measuring stock.ts's `accent` field
against the actual rendered UI (docs/cli-registry.md flags this field as
"transcribed, not authoritative — re-measure before wiring one up"):
1. A real, user-visible bug: `.btn-toolbar.btn-run.mode-gemini`,
`.mode-antigravity` and `.mode-omp` had no override rule inside the
`html:not([data-skin="og"])` block, unlike codex/pi/grok/deepseek, which
do. The generic `.btn-toolbar.btn-run` rule in that block resolves at
higher specificity than the base sheet's per-mode pair, so all three
rendered as plain claude-blue on every skin except `og` — including
`daylight-blue`, which is the actual DEFAULT skin for a fresh install
(index.html's pre-paint script), not an edge case. Added the three
missing rules, sourced from each CLI's own already-designed og-skin
colours (no new colours invented), mirroring the exact pattern
pi/grok/deepseek already use. Also corrected the stale comment on the
pi rule, which claimed this was still broken for gemini/antigravity.
2. `stock.ts`'s `accent` field was simply wrong for most CLIs — e.g. claude
was registered as Anthropic's brand orange (#d97757) while its button
renders blue, antigravity was registered purple while it renders cyan,
pi was registered green while it renders pink. Measured each CLI's real
`border-color` from its own `.mode-<id>` rule on the og skin (the
cleanest single representative hex each entry's gradient resolves
around) and corrected all 9 non-shell entries to match. `accent` has no
reader yet (confirmed via the DECLARED_FOR_LATER guard test), so this
changes no rendered output — it's a data-accuracy fix, matching the
registry's own "transcribed, not authoritative" warning taken literally.
Also fixed a false claim in types.ts's doc comment for the field
("CSS derives every per-CLI gradient from it via --cli-accent") — no
such CSS variable exists anywhere in the codebase.
Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format:check/
check:public-assets/check:frontend-syntax all clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
The maintainer's promised merge-time fixes from the final review of #453:
1. closeSplitPane() tears down a divider drag still in progress, so a split
that collapses mid-drag no longer leaves body.split-pane-resizing (the
page-wide col-resize cursor and user-select lock) set until a reload.
2. openSplitPane() re-applies the picker's own exclusions (detached session,
pid === null, no session record) for a row that went stale while the
menu sat open, refusing silently like its neighbouring gates.
3. architecture-invariants: the hard-hide of .btn-split is the
@media (max-width: 1179px) rule in styles.css, not mobile.css.
4. SplitTerminalPane.destroy() nulls onclose (and onerror) beside onopen
and onmessage.
5. Picker rows drop the data-session-id attribute nothing read.
6. The Pane-A-ends branch collapses with skipPrimaryResize, so the closing
resize is no longer aimed at the session the server just removed.
7. The {t:'r'} refresh path is single-flight across the fetch and the
chunked write, coalescing a mid-replay refresh into one trailing re-run.
Tests: split-pane-auto-collapse-unit gains the drag-teardown, exclusion and
skip-resize cases; the new split-pane-terminal-unit covers destroy() and the
refresh single-flight. All were run against the pre-fix module to confirm
they fail there.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit dbd39aed015ae5ae5870aba398bf4b4ab5118e47)
- styles.css: restate the composer overlay's own bottom gutter after the fold rules
(the generic .paste-overlay longhand erased it: 0px flat, hinge strip replacing it
folded) and subtract the fold strip from the dialog's max-height
- test/foldable-layout.test.ts: simulate the cascade for
.paste-overlay.prompt-composer-overlay (fails without the CSS fix); pin the palette
anchor by name instead of ELEMENTS.at(-1)
- keyboard-accessory.js: guard the app global in refreshForActiveSession() like the
rest of the file
- keyboard-accessory.js: a whitespace-only draft is empty (Send no longer submits
blank lines); the text still goes out untrimmed
- keyboard-accessory.js: derive _composerMaxLength and the frame refusal from one
64 KiB frame limit minus both bracketed-paste markers so they cannot drift
- keyboard-accessory.js: translate the textarea placeholder and label at build time,
since the DOM translator skips <textarea> subtrees
- i18n.js: zh-CN entries for the composer dialog copy
- docs/wiki/Mobile-Guide.md: describe the Compose key instead of a clipboard key
- CLAUDE.md: a "Mobile prompt composer" paragraph after the accessory bar one
- test/mobile-prompt-composer.test.ts: pin the whitespace rule and the derived budget
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit f6725ba52da17b0bdbee8be3b5011e7cae514f69)
- test/opencode-resize.test.ts: retarget the launcher guard at the real code (this.selectSession(firstSessionId), any this.activeSessionId assignment) with an anti-vacuity check; the old strings existed nowhere, so it could never fail
- session-ui.js: restore as comments the two invariants the merged bodies lost (deepseek leaves statusReporting unset, i.e. ON; no effort field for external CLIs, it is Claude-specific)
- docs/cli-registry.md: move the frontend-guard paragraph below the two backend-guard paragraphs so they keep their antecedent, and note the widened comparison shape
- test/frontend-cli-no-id-branching.test.ts: the comparison shape accepts any left-hand identifier (const m = this._runMode; m === 'codex' was invisible), normalized to `mode`; the two `m !== 'shell'` display filters are allowlisted and the remaining blind spots documented
- test/run-mode-dispatch.test.ts: table-driven pin of run() dispatch (claude to runClaude, each RUN_MODE_LAUNCH id to _runCliMode(id), shell to runShell, unknown to runClaude, lock held and released)
- CLAUDE.md: name the second CI-gated guard next to the backend one
- server.ts: every </head> injection passes a replacer function; a clis.json label containing $' re-injected the rest of the document past escapeScriptJson (two render tests pin it, proven failing on the string form)
- _isAltCliMode(): no reference anywhere in the tree, nothing to fix
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 1ea363ff808a62861559bc141e724b163cc1c56e)
The maintainer's promised follow-ups to opticon454's picker promotion,
applied on the landing branch after the merge (ecb95b5d):
- session-ui.js: the promotion tag ("Currently loaded" / "Last used") and
the "Default" pill are two separate spans, so a promoted row that is
also the endpoint's defaultModelId shows both instead of silently
losing its Default marking; two tests pin it (both fail on the old
exclusive-slot rendering).
- styles.css: a dedicated #customModelPickModal .set-scope rule, since
the pill was only styled inside the three settings modals and rendered
as plain body text here; same skin tokens, modal layout untouched.
- docs/wiki/Custom-Model-Endpoints.md: describe the promotion (currently
loaded, else last used per device), the separate Default pill, and
that nothing is ever auto-chosen.
- CLAUDE.md + docs/custom-model-endpoints.md: credit the real "Last used"
writers (_runCustomModelEntryViaRestart and
_quickStartWithCustomModelConfirm; runCustomModelEntry only dispatches
since 88e5b7b2) and drop the now-wrong "both defer to Default" sentence.
- Not done: moving the one-shot "last used" write into
_runCustomModelEntryOneShot, because the existing one-shot tests assert
that _quickStartWithCustomModelConfirm writes the key itself.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 0cb0f911adc14a852ba5c2951768a4aa87c25657)
The two reviews on the PR (DeepSeek Harness, then Claude) found one class of
bug twice and a list of smaller ones; all of them land here, each pinned in
test/install-sh-invariants.test.ts and, where it is bash logic, driven in the
bash:3.2 CI step as well.
The Start line the done screen prints is now composed in one place
(start_command_hint) from every non-default value, the same five the exec
branch exports through export_bind_env, so "do not start" under a sub-path or
a custom port no longer prints a bare `codeman web`. The --lan / --tailscale /
env preset paths read ${CODEMAN_PASSWORD:-$EXISTING_PASSWORD}: a flag re-run on
a unit that carried a password used to rewrite it without the password and
with the unauthenticated ack. --password and --port flip RECONFIGURE so they
reach the unit instead of taking the quiet update path, and `install.sh name`
re-syncs the unit's base URL after the mapping is re-added.
Also: the sudo keepalive is ended before the exec into the foreground server
(exec skips the EXIT trap, and the loop keys on $$); Ctrl+C in the HTTPS-toggle
poll is trapped for the poll only and skips Tailscale for the run instead of
killing the installer; uninstall asks before removing a LaunchDaemon this
installer never wrote; a foreign daemon gets a launchctl kickstart hint and the
done screen stops claiming the new build is running; the preflight summary
reads the Tailscale state with a line grep when node is not installed yet; the
LAN security notice uses the configured port; a bare re-run ends on the done
screen; a build failure after a rename names the install.sh tailscale
recovery; TS_JOINED_HERE (written, never read) is gone; the plan doc and
architecture-invariants say what the code does. A minor changeset is included.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
README, the Installation / Remote-Access / Running-As-A-Service wiki pages,
docs/security-architecture.md and CLAUDE.md describe the three-question flow,
the flags, the subcommands, the sub-path answer for an occupied :443 and why
the rename is opt-in. docs/installer-v2-plan.md is the design and the
verification record (what was measured, what still needs a fresh machine);
docs/tailscale-installer-plan.md points at it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The installer used to ask about ten things, half of them after a multi-minute
build, and the question that matters most (how do I reach the dashboard) came
last. It now looks at what is on the machine, asks at most three questions
(access, an optional tailnet name, service), and does the rest unattended.
- Every step that needs a human runs before the build: one consent for all
missing packages, one sudo prompt kept warm for the run, the AI CLI menu,
and the Tailscale install/login/operator/HTTPS-toggle preflight (the toggle
is polled and the admin page opened in a browser, instead of "re-check
now?").
- The build, the service and `tailscale serve` run behind spinners with their
output in ~/.codeman/install.log; the tail is shown on failure and a failed
dependency install names its step.
- The done screen leads with the URL (tailnet, network, this machine) and a
terminal QR code from the qrcode package Codeman already ships.
`install.sh status` prints it again.
- Tailscale is two halves: tailscale_prepare (question phase) decides the
serve SHAPE, tailscale_apply (after the build) issues the one serve command.
When :443 already belongs to another app, Codeman goes under a sub-path
(serve --set-path /codeman + CODEMAN_BASE_URL in the unit; serve strips the
prefix, Codeman's ingress tolerates that, --base-url covers the URLs it
emits) or a second port, instead of replace-or-nothing.
- Renaming the node to codeman-<hostname> is opt-in and defaults to no
everywhere (the tailnet name is the machine's ssh identity); --name and
`install.sh name` do it, uninstall offers the old name back. Serve config is
keyed by the DNS name, so a rename takes our mapping down first and re-adds
it under the new name.
- Flags pipe through `bash -s --`: --tailscale|--lan|--local, --name|--no-rename,
--service|--run|--no-start, --yes, --password, --port. --port is now also
written into the service file.
- npm install runs with CODEMAN_NO_AUTOSTART=1: postinstall otherwise builds
and starts a detached `codeman web` on 127.0.0.1:3000, which made the
service crash-loop on EADDRINUSE while the done screen reported "running"
off the orphan (fresh Ubuntu 24 sandbox).
- The LAN address comes from the default route, not the first interface.
- A foreign /Library/LaunchDaemons/com.codeman.web.plist is left alone
instead of being replaced by a LaunchAgent.
- The cloudflared question leaves the main flow (`install.sh cloudflared`).
- "Continue WITHOUT a password?" defaults to yes (owner decision).
Tests: the invariants test pins no `serve reset`, no funnel, no Tailscale
Service, every serve mutation through ts_cmd_serve, rename before shape,
flag/header parity, the rename default and the NO_AUTOSTART opt-out; the CI
bash 3.2 step drives the question phase with stubbed tailscale state.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The smart-copy gate only entered its selection-check block behind
hasSelection(), so a selection-less Ctrl+Shift+C skipped straight to
`return true` and ceded the keystroke to the browser's own handling
(e.g. Chrome's Inspect-Element binding) instead of matching Pane A's
"never falls through" contract for that chord.
Verified live in a real browser that this is a UX-parity fix, not an
interrupt-safety one: xterm's evaluateKeyboardEvent never emits PTY
data for a shifted ctrl-letter regardless of any gate (only "_" and
"@" get special-cased), so no accidental 0x03 was ever at risk. The
regression test added here asserts on the dispatched event's
defaultPrevented rather than the absence of a WS frame, since the
frame-count check passes vacuously for this exact key combo whether
or not the gate fires.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- buildSplitPickerSessions() now excludes any session with pid === null
(exited CLI, tripped PTY-exit breaker, a restore that never re-attached).
Pane B has no equivalent of selectSession()'s auto re-attach POST, so a
split opened onto one had nothing reading its tmux pane: no terminal
events ever arrived and Session.write() silently dropped every keystroke
with no ack either way, while the socket itself reported healthy.
- Fixed the hollow chord regression test: the synthetic keydowns carried no
keyCode, which is what xterm's evaluateKeyboardEvent switches on to
produce a data frame at all, so the assertion held regardless of whether
the gate fired. Adding real keyCodes surfaced a second, real bug in the
Alt+B case: the event bubbles to app.js's own document-level shortcut
dispatcher, which really toggles the sidebar and resets the layout
attribute the gate reads before Pane B's own (later, non-capture) handler
ever sees it — fixed by driving the app's real settings cache instead of
only the DOM attribute.
- Ported the two remaining primary-pane gates with real consequences:
Ctrl+Z (SIGTSTP) is swallowed for every non-shell session, matching
terminal-ui.js's reasoning (an Ink/TUI agent loop stops dead with no
visible output otherwise), and Shift/Ctrl+Enter now POSTs to
/api/sessions/:id/send-key for THIS pane's own session instead of
letting xterm send a bare \r, which used to submit an incomplete prompt
instead of inserting a newline. Smart-copy Ctrl+C is re-implemented
against Pane B's own terminal (copying app.copyTerminalSelection() would
have copied Pane A's selection instead).
- Updated docs/architecture-invariants.md and docs/split-pane-sessions-plan.md
to match, and added CLAUDE.md's missing .split-picker-menu z-index entry.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Pane B had no attachCustomKeyEventHandler of its own, so the document
capture-phase shortcut handler's preventDefault() (which does not stop
xterm) left Ctrl+K/Alt+1/Alt+B ALSO writing their raw byte/escape
sequence into Pane B's live PTY on top of whatever the app action did
to Pane A. Pane B now installs the same registry-aware gates the
primary pane's own attachCustomKeyEventHandler uses. Ctrl+V is left on
xterm's default paste — no image-paste trap to route it to.
Plus the rest of the review's smaller items:
- Narrowing the window past the desktop gate now closes an open split
instead of leaving it stranded on screen.
- Split is refused while a web tab is active (activeWebviewId), which
used to open Pane B's socket behind a hidden container.
- Pane B now handles the server's `{t:'r'}` refresh frame via a shared
_loadBuffer() helper (also used by connect()), instead of ignoring it.
- The divider drag now uses pointer events + setPointerCapture (mirrors
tab-rail-resize.js), a button!==0 guard, preventDefault, and a
body.split-pane-resizing cursor/selection lock — a plain mousedown
drag selected the text under the cursor as it crossed both terminals.
- Pane B's close control and the picker rows are real <button>s now
(keyboard-reachable), with matching CSS chrome resets.
- Dropped the redundant CodemanBase.base prefix on the buffer fetch
(the global fetch wrapper already applies it).
- data-preview-order for the Split settings chip moved from a collision
with Ultracode Agents (both 15/12) to 11.5, matching its real
position between Multi-monitor and Ultracode Agents in the header;
widened test/app-settings-structure.test.ts's regex to allow the
decimal (Number() already parses it fine for the preview sort).
- Added zh-CN i18n entries for the Split button and empty-picker text.
- Dropped the stray unused `vi` import Ark0N flagged as unrelated to
this feature (vitest's `globals: true` makes it ambient anyway).
- Documented the fix and the deliberate no-cid/seq choice in the
split-pane-sessions architecture-invariants entry.
Added a real-Chromium regression test asserting Ctrl+K/Alt+1/Alt+B
dispatched at Pane B's own textarea send no `{t:'i'}` frame over its
WebSocket. Full CI gate green (409 files, 7736 tests) plus all 8
split-pane browser tests.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
CLAUDE.md previously only mentioned split-pane in the load-order list,
with nothing in the Architecture/frontend prose the way every other
feature gets, and its own module count was one stale (34, should have
been bumped to 35 when terminal-split.js was added). Add a short
pointer-style paragraph next to the other terminal features, fix the
count.
docs/wiki/The-Dashboard.md's header button table and
docs/wiki/Settings-Reference.md's header chips list are the two
user-facing surfaces that never mention Split at all; added both, plus
a note that the feature is desktop-only regardless of the setting.
docs/split-pane-sessions-plan.md: recorded the one design note that
isn't a code change — the global capture-phase shortcut handler always
resolves against Pane A, so Ctrl+L/Ctrl+W typed into Pane B affects the
other session. Not fixed for v1, same reasoning as the rest of the
"deliberately plainer" section.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two majors:
- The divider drag was unthrottled: every mousemove did a full xterm
reflow on BOTH panes and sent Pane B a {t:'z'} resize frame with no
unchanged-dimensions skip, fanning out into a `tmux resize-window`
child plus a SIGWINCH per event — ~50 of each dragging across half a
wide viewport. SplitTerminalPane.fit() is now split into localFit()
(reflow only) and fit() (reflow + send); the drag coalesces moves
into one localFit() per animation frame via requestAnimationFrame,
and sends the real resize for both panes exactly once, at drag end,
matching the primary pane's own throttledResize convention.
- Pane B pulled the FULL scrollback unchunked for every session mode,
writing it in one terminal.write() call. Mirrors the primary pane's
own mode check (app.js's selectSession): shell sessions get a
bounded 1MiB ?tail= fetch instead of ?full=1, and the fetched buffer
is written through a minimal chunked writer (32KB slices, yielding a
frame between each) instead of one primary-pane chunkedTerminalWrite
this simpler, independently created/destroyed pane has no equivalent
of (no session-switch generation counters or live-output gate).
Smaller items from the same review:
- Pane B now follows live appearance changes (applyTerminalSkin,
applyTerminalFontFamily, applyTerminalFontWeights, setFontSize all
propagate to it, matching the teammateTerminals pattern) and reads
the real codeman-font-size/terminalFontFamily/weights/DEFAULT_SCROLLBACK
settings at construction instead of hardcoding fontSize 14 / scrollback 5000.
- The Pane-B-promotion path now skips selectSession() when
_closingSessions already owns this delete (the user closing Pane A's
own tab), matching _onSessionDeleted's own active-session-handoff guard.
- Detaching a session AFTER a split is already open now yields the PTY
size in _sendResize() too (not just at picker-open time), mirroring
sendResize's own detachedElsewhere guard.
- .btn-split joins the body.solo-mode hide list, next to .btn-multimonitor.
- The split row was 6px wider than its container (two flex-shrink:0
50% panes plus a 6px divider): both panes are now flex-shrink 1.
- Pane B's header and the split-picker rows are marked so i18n.js's
exact-string lookup skips them, matching .session-name elsewhere —
a session literally named e.g. "Sessions" was translatable on zh-CN.
- The Split button now reflects open/closed state via a `.split-open`
accent style, aria-pressed, and a title/aria-label that says which
behaviour the next click gets.
- _splitPane.connect() is no longer an unawaited call with no .catch().
- terminal-split.js's fileoverview pointed at a doc path that was
renamed away in the previous push; @dependency now credits
constants.js for CodemanTerminalFont, not terminal-ui.js.
- index.html's Split settings chip no longer reuses data-preview-order
"12" (already the Ultracode Agents chip's slot in the same "header"
preview group).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Blocker from Ark0N's second PR #453 pass: moving showSplitButton into
settings-ui.js's per-device displayKeys set was only half of making it
per-device. saveAppSettings() still put it in the object PUT to
/api/settings, SettingsUpdateSchema (.strict()) does not declare it,
the server answered 400 INVALID_INPUT, and because the call site never
checked res.ok the UI still reported "Settings saved" while NOTHING
persisted — workspaceHooksEnabled, agentSkillEnabled, tunnelEnabled,
claudeModel, every toggle, on every save, on every device. Strip it
out via the same destructure every other per-device key goes through
(`showSplitButton: _ssp,`), drop the stray mention from a schemas.ts
comment (a mention there reads as "this is a real field" to the next
grep), and add a static guard test mirroring
test/terminal-auto-copy.test.ts's three-way rule.
Also finishes the desktop gate the first pass only did in CSS at
599px: SPLIT_PANE_MIN_WIDTH (1180, matching HOME_SESSIONS_MIN_WIDTH)
now backs an actual JS width check in _applySplitButtonVisibility,
with a matchMedia listener so a live window resize hides/shows the
button without a reload — the CSS backstop in styles.css is the
reverse-direction guarantee for when JS hasn't run.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ark0N's PR #453 review: nothing gated this feature to desktop even
though the design called for it (two 240px min-width panes plus the
divider need ~486px, and the divider has no touch handlers), and
showSplitButton was a SYNCED setting, so turning it on at a desk also
put the button in the phone header.
- Hard-hide .btn-split on phones in mobile.css regardless of the
setting, matching the other desktop-oriented header buttons in the
same @media (max-width: 599px) block.
- Move showSplitButton into settings-ui.js's per-device displayKeys
set and drop it from SettingsUpdateSchema entirely, matching the
showFileViewerButton/skin precedent (CLAUDE.md's "per-device keys
... must NOT be added to SettingsUpdateSchema" rule) — a desktop
opt-in must never sync onto a phone that never asked for it. Removes
the now-invalid server-round-trip test for the setting.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Exclude popped-out (detached) sessions from the split picker:
SplitTerminalPane._sendResize() has no yield-to-detached-window check
the way the primary pane's sendResize() does, so splitting against a
detached session put its own window and Pane B in a fight over the
same PTY's dimensions. Simplest fix per the review: keep them out of
buildSplitPickerSessions() entirely.
- Show a visible dead state when Pane B's WebSocket drops. onData
already silently discards keystrokes while the socket isn't OPEN
(there is no reconnect for v1), so a dropped socket left the pane
looking normal while it quietly ate everything typed into it.
- openSplitPane() returns early with no active session, so a split
triggered from the home screen no longer creates and connects Pane B
behind the opaque welcome overlay with nothing to show for it.
- onMove() during a divider drag now bails when the split has
auto-collapsed mid-drag (the other pane's session ending) instead of
throwing on `divider.parentElement` being null.
- Promote Pane B via `selectSession(id, { auto: true })` when Pane A's
session ends — this is an app-driven selection, not the user clicking
a tab, so it must not spend the promoted session's idle alert.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ark0N's PR #453 review: fit() was only ever called from the divider
drag, and the trailing-edge ResizeObserver callback in terminal-ui.js
(throttledResize) only ever measured Pane A's own container. Split at
a wide viewport, shrink the window (or toggle the Alt+B sidebar, or
drag the tab rail), and Pane A's cols changed while Pane B silently
kept its stale PTY size in both xterm and the real pane.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Per Ark0N's review on PR #453: rename the design spec to
docs/split-pane-sessions-plan.md, matching every other feature's
*-plan.md convention, and drop the 957-line implementation task plan
(docs/superpowers/plans/2026-09-15-split-pane-sessions.md) — workflow
scaffolding for the subagent-driven-development run, not repo
documentation. Fixes the now-dangling link in architecture-invariants.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Real-browser regression coverage for the previous commit:
- SplitTerminalPane connects onto an already-quiet session and shows its
existing scrollback with no new output, proving the ?full=1 fetch (not
a live echo) populated the pane.
- openSplitPane() force-resizes Pane A synchronously as part of opening
a split.
- Dragging the divider force-resizes Pane A once, at drag end.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Pane B's SplitTerminalPane.connect() only opened a WebSocket and waited for
live output — ws-routes.ts's terminal socket sends nothing on connect, only
future 'terminal' events — so it stayed blank until the target session
happened to produce new output. It looked intermittent rather than
always-broken because a resize sent by _sendResize() often nudges the
session's real tmux window to a new size, and tmux repaints its current
screen on resize; that incidental repaint was what usually populated the
pane. When Pane B's computed dimensions already matched the session's
last-known size, Session.resize() skipped the resize as a no-op and the
pane stayed empty. Fetch the existing scrollback (?full=1) before opening
the socket, same as the primary pane does.
Pane A never told its own session's PTY/tmux about a size change at all,
relying purely on the passive 300ms-debounced ResizeObserver in
terminal-ui.js. openSplitPane() now force-resizes Pane A immediately on
entering split (mirroring closeSplitPane()'s existing symmetric call), and
the divider-drag handler force-resizes it once at drag end (matching the
codebase's established trailing-edge debounce convention rather than
flooding a resize per mousemove).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
test/split-pane-auto-collapse.browser.test.ts covers "Pane B's session ends"
in a real Chromium, but that suite is excluded from the npm test CI gate.
The "Pane A's session ends, Pane B gets promoted" branch had no coverage
anywhere, and it is the one branch whose correctness depends on exact
ordering: _splitSessionId must be captured BEFORE closeSplitPane() runs
(which nulls it) or the promoted session id is lost. Loads terminal-split.js
via `vm` against a minimal fake CodemanApp (same technique as
test/session-close-fallback.test.ts), and pins all three branches (Pane A
ends, Pane B ends, unrelated session ends) plus that the original
_onSessionDeleted always still fires.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Nothing stopped a stale picker click (opened before switching tabs) or
clicking Pane B's own session tab while split from landing on
openSplitPane(sessionId) with sessionId === activeSessionId, or from
selectSession() rebinding the primary pane onto the session Pane B was
already showing — either way, two live WebSockets to one session, each
independently claiming PTY dimensions via its own {t:'z',...} resize frame.
openSplitPane() now refuses early when the target is already the active
session, and a new selectSession() prototype patch (same top-level pattern
as the existing _onSessionDeleted patch) closes an active split BEFORE the
primary pane rebinds to the session Pane B holds.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The new "Split-pane sessions" section was inserted between the "Session
list layout (header strip vs. left sidebar)" heading and that section's own
body paragraphs, orphaning the heading from its content. Move "Split-pane
sessions" to after the Session list layout section's full body, before
"Gesture control: the setting" — no change to the Session list layout prose
itself, only where the new section sits relative to it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
.main.webview-active hid .terminal-wrap when a web tab became active, but
.terminal-wrap is reparented INSIDE .terminal-split-container while a split
is open, so Pane B and the divider stayed stranded on screen over the
dashboard iframe. Hide the whole split container as one unit, mirroring the
existing .terminal-wrap rule; no state is destroyed, so returning to the
session tab shows the split intact.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
.split-picker-menu/-item/-empty (created in openSplitPicker()) had zero CSS
and could not be dismissed except by picking an item — a default-path defect
since the Split button ships enabled to anyone who flips showSplitButton on.
Add CSS matching the sibling .run-mode-menu popover's look (floating-bg
backdrop blur, border, shadow, z-index 1000 above the header's 100), and
dismiss on outside click or Escape via the same one-shot listener pattern
session-ui.js already uses for its other transient popovers
(toggleCaseSettings(), toggleRunModeMenu()). Picking an item now routes
through the same _dismissSplitPicker() method as the outside-click/Escape
handlers, so the listeners never outlive the menu.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
_sendResize() clamped Pane B's proposed cols/rows to a 40/10 floor before
sending the {t:'z',...} resize frame, so the PTY was misinformed of Pane B's
real width at the divider's own reachable 20% position, causing real
output-wrapping bugs. The primary pane (terminal-ui.js's
getTerminalDimensions()) sends fitAddon.proposeDimensions() unclamped and
lets the server enforce its own valid range ([1,500]/[1,200] in
ws-routes.ts); Pane B now matches that convention.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
.btn-split--hidden had no matching CSS rule anywhere, so the opt-in Split
header button shipped visible to every user on every viewport regardless of
the setting. Add the `display: none !important` rule alongside its sibling
marker classes (.btn-multimonitor--hidden etc.), plus a static regression
guard (test/split-pane-hidden-button-css.test.ts) that fails if any future
"*--hidden" marker class in index.html is missing a matching CSS rule.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Task 4's implementer found two real bugs in this plan's browser-test
helpers: POST /api/sessions nests the id at data.session.id (not
data.id), and mode:'shell' needs a follow-up POST .../shell to actually
spawn a PTY. Fixed in Task 4's own snippet (documentation accuracy —
already fixed in the real committed code) and pre-emptively in Tasks
5/6's createShellSession() helper before either was dispatched, so
neither implementer has to rediscover it independently. Also corrected
the <script> tag snippet to defer, matching the real file's convention.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Monkey-patching the instance's _onSessionDeleted inside a
DOMContentLoaded listener races connectSSE()'s handler-wrapper cache,
which captures the function reference by value on first connect and
never re-reads it. Patching CodemanApp.prototype at module-evaluation
time (synchronous script-tag order) is unraceable: it completes before
any instance exists or connectSSE() ever runs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Preflight scan for SDD execution caught two classes of defect before
dispatch: Task 2's test invented a buildTestApp() helper and response
envelope that don't exist for /api/settings; Tasks 4-6 used
@playwright/test's runner against a test/browser/ directory that
doesn't exist in this codebase. Both corrected against real patterns
found in existing tests (system-routes-settings-partial-put.test.ts,
terminal-copy-shortcut.test.ts, tab-rail-resize.browser.test.ts).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Global Constraints previously read as if local-echo/CJK/accessory-bar
were desktop features; they are mobile-only, and split-pane is the
desktop-only side of that equation. Also names the exact spec section
instead of a loose paraphrase, and adds a worktree/branch note so an
executing subagent knows where this plan runs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
7 tasks: pure divider/picker helpers, showSplitButton header wiring,
split-container CSS, SplitTerminalPane (Pane B's independent xterm+WS),
open/close orchestration with picker and divider drag, auto-collapse on
either session ending, and an architecture-invariants entry.
Also folds in the "detach session" prior art discovered mid-brainstorm
into the spec's architecture section.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The Problem paragraph and the architecture section used "tab" where
"pane" was meant, colliding with the browser's own tab concept.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Scopes v1 of an in-app split view (two live session panes side-by-side,
draggable divider) after multi-monitor spanning turned out to solve a
different problem than showing multiple panes at once.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Three of Ark0N's four "will take at merge" items, applied instead since
they were straightforward to do properly:
1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on
<file>::<expression> (fixed last round) closed the line-shift problem
but opened a new one: every stock id was already allowlisted for
session-ui.js in the `mode === '<id>'` form, so a BRAND NEW branch
reusing that exact expression anywhere in the file passed unnoticed.
Reproduced live (`if (this.mode === 'codex')` injected into
runOpenCode()) — stayed green under the old version. Each allowlist
entry now carries the exact count of approved call sites, and a new
test asserts actual-vs-declared count for every key; a mismatch in
either direction is real (higher = new unreviewed branch riding in on
an existing approval, lower = a reviewed site was removed and the
entry is now stale). Reproduced again against the fix: same injection
now fails with an exact diagnostic (expected 2, found 3).
2. Added test/run-mode-launch-table-drift.test.ts. RUN_MODE_LAUNCH
restates four things stock.ts already owns (label, install command,
supportsCustomModel, the external-mode key set), and they agree today
with nothing enforcing it. supportsCustomModel is the dangerous one:
the Run-menu picker's rows come from the server-injected
window.__codemanCustomModelClis (built from
capabilities.customModelInjection.kind), so a CLI gaining a real
injection recipe later would be OFFERED in the picker while
_runCliMode silently drops the customModel field for it — the session
launches on the vendor's cloud while the UI claims the local endpoint.
Drives the real session-ui.js via JSDOM and compares RUN_MODE_LAUNCH
against STOCK_CLIS on all four axes.
3. Inlined the "Open Question 7 in PR-B2.md" references in the allowlist
reasons — PR-B2.md is a local planning doc, never part of the
committed tree, so the reference was dead on arrival for anyone
reading the repo. Points at the PR #458 review thread instead.
4. Added a sentence to docs/cli-registry.md naming the new frontend guard
alongside the backend one it mirrors.
Full gate: 406 files / 7721 tests / 0 failures, typecheck/lint/format/
check:frontend-syntax all clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
Four things from the maintainer's review on PR #459, all fixed:
1. Bound _getCustomModelCurrentlyLoaded's probe client-side (~800ms via
Promise.race, on top of — never instead of — the route's own 5s
server-side timeout). Without it, an asleep/firewalled endpoint behind
a saved model list left the picker completely invisible for up to 5s
after the Run menu had already closed, with no spinner or toast.
`timeoutMs` is an optional param (default 800, real callers never pass
it) so a test can drive it in milliseconds, same pattern as
`_watchLlamaSwapLoading`'s own `pollIntervalMs` — this code runs in a
JSDOM window's own realm, whose setTimeout vi.useFakeTimers() cannot
patch.
2. "Last used" is now written only once a launch actually applies, never
on the mere click. It moved out of runCustomModelEntry (unconditional)
and into each path's own success point: _quickStartWithCustomModelConfirm
after the final post succeeds, and _runCustomModelEntryViaRestart right
after the apply's success check. A context-window-warning decline means
this exact model cannot work with this CLI at all, so the old
unconditional write would promote, next time the picker opened, the one
model guaranteed to fail again.
3. Documented the promotion/tag precedence and the new
codeman:customModelLastUsed:<mode>:<endpointId> localStorage key in both
CLAUDE.md's Custom Model Endpoint Profiles section and
docs/custom-model-endpoints.md's Run-menu picker section.
4. Added zh-CN entries for "Currently loaded" and "Last used" in i18n.js,
next to this modal's existing "Choose a model"/"Custom Endpoints" pair.
New tests: the client-side timeout (endpoint that never answers, one that
answers within the bound, and a rejected-after-timeout probe settling
quietly), and "last used" recording on success vs. NOT recording on either
confirmation's decline, for both the restart and one-shot paths.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Code review (high effort) on the previous commit found a real race: making
_openCustomModelPickModal async (it now awaits the currently-loaded-model
probe before rendering) meant a second, faster call for a different
endpoint could render first, only for the first call's slower probe to
resolve afterwards and overwrite the modal with the wrong endpoint's model
list — while _pendingCustomModelPick (set synchronously, before either
await) still named the second, correct endpoint. Picking a model in that
state would launch/apply the wrong model on the wrong endpoint.
Fixed with the same mutable-generation-counter guard
_watchLlamaSwapLoading already uses for an identical async-superseded-by-
newer-call shape: every DOM write, including _pendingCustomModelPick
itself, is deferred until after the awaited probe, and a call that finds
its generation already superseded bails out untouched instead of clobbering
whatever a newer call already rendered.
Added a regression test driving two overlapping opens with a controlled
promise so the earlier, slower probe resolves after the later, faster one
renders, asserting the late response is a no-op.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Custom Model Endpoint Profiles' "which model" picker (session-ui.js's
_openCustomModelPickModal) always listed models in their raw discovery
order, so on a host with several downloaded GGUFs the user had to
remember (or eyeball the "Default" tag) which one llama-swap actually
had hot before picking — the whole point of the picker being fast is
undone if it makes you think first.
The picker now promotes exactly one model to the top of the list:
- If llama-swap reports a model from this host's own list `ready`
right now (via the existing GET /api/model-endpoints/:id/running-status
route), that model is promoted and tagged "Currently loaded" — it's
what a launch attaches to with zero wait.
- Otherwise, the last model actually launched on this exact
(harness, endpoint) pair is promoted and tagged "Last used", read
from a new per-device localStorage key
(codeman:customModelLastUsed:<mode>:<endpointId>), written by
runCustomModelEntry on every launch attempt regardless of outcome.
- A plain (non-llama-swap) OpenAI-compatible server, an unreachable
endpoint, or a loaded-but-not-yet-ready model never promotes
anything — the rest of the list keeps its discovery order.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
Two required fixes from Ark0N's review of #458:
1. test/frontend-cli-no-id-branching.test.ts's ALLOWED_BRANCHES keyed on
<file>::<line>::<expression>. A single inserted line anywhere above an
entry shifted every subsequent line number, so all 21 entries went stale
simultaneously and the same 21 branches were reported as "new" — on a
file six other open PRs also touch. Dropped the line number from the key
(<file>::<expression>, matching the backend guard's own design), which
collapses 21 line-keyed entries to 11 or-collapse where the same
expression recurs at multiple call sites in the same file.
2. test/run-mode-ui.test.ts's terminal-ownership guard scanned method
bodies via `^ {2}async (run[A-Za-z]*)\(\) \{$`, which matched the 8
one-line run<Mode>() wrappers PR B2 introduced but not _runCliMode(mode),
where the real logic (and the actual risk the guard exists to catch) now
lives. Fixed the regex to `^ {2}async (_?run[A-Za-z]*)\(\w*\) \{$` and
added _runCliMode to the sanity list. Same-class fix in
test/opencode-resize.test.ts, which had the identical blind spot via
runOpenCode.toString().
Both reproduced live before fixing (inserted the same comment line; added
this.terminal.clear() to _runCliMode) to confirm the bug, then confirmed
the fix catches it and the suite stays green otherwise.
Also resolves Open Question 2 by dropping window.__codemanCliCatalog
entirely: nothing consumed it, and a registry DECLARED_FOR_LATER field
costs nothing until read while an unconsumed script tag on every page
render is a different trade. Reverts Phase 1 cleanly — server.ts's
injection, shortBadge back in types.ts's DECLARED_FOR_LATER list and the
pinned guard test, and the three associated render-index-html.test.ts /
server-index-title.test.ts assertions.
Full gate: 405 files / 7717 tests / 0 failures (net unchanged), typecheck/
lint/format clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
PR #380 (PR B) held back the frontend half of the CLI registry refactor,
explicitly deferring window.__codemanCliCatalog and making session-ui.js /
mobile-overview.js catalogue-driven as "PR B2".
- Inject window.__codemanCliCatalog in renderIndexHtml(), following the
existing __codemanCustomModelClis pattern (escapeScriptJson-guarded,
resolved per-request). Reading CliEntry.shortBadge here is what makes it
genuinely read, so it drops out of types.ts's DECLARED_FOR_LATER list.
- Consolidate session-ui.js's 8 near-duplicate run<Mode>() launch functions
(opencode/codex/gemini/antigravity/pi/omp/grok/deepseek) into one shared
_runCliMode() plus a local RUN_MODE_LAUNCH config table. The 8 method
names stay as thin wrappers (index.html calls them by name; tests assert
on the name). Also collapses a duplicated 8-way isAltMode/isExternalCli
OR-chain (same expression, copy-pasted twice in openSessionOptions) into
one EXTERNAL_CLI_MODES check.
- Add test/frontend-cli-no-id-branching.test.ts, a guard scoped to
session-ui.js/mobile-overview.js only (not the rest of src/web/public/,
which stays explicitly out of scope per CLAUDE.md), mirroring the
backend's own no-id-branching guard.
mobile-overview.js and the wiring of accent/echo/wheelForward/
keyboardAccessory were investigated and deliberately left alone: the first
is already a single, tested, gated table (not duplicated logic); the second
set belongs to terminal-ui.js/keyboard-accessory.js/styles.css, files
outside this PR's mandate.
Verified on a tmux-capable devbox (this sandbox has no tmux): full CI gate
at 405 files / 7717 tests / 0 failures, typecheck clean, 94 targeted tests
covering exact per-CLI wire-body shapes unmodified and passing, and a live
anti-vacuity check on the new guard (injected a real branch, confirmed it
fails, reverted, confirmed green).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
A full review of the release tree found seven things, and four of them were mine.
**The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext`
and `confirmedSwap` changed the wire field without moving three assertions that check
it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts`
(the swap modal and the context modal, each of which already receives exactly the right
per-question flag). Moved, with the titles.
**Worse, my own tests for the split never ran.** The four cases in
`session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`,
which was declared inside a sibling `describe`, so they threw a ReferenceError during
setup. The split would have shipped with no passing server-side coverage while the gate
reported the failure as four broken tests rather than as four tests that were never
written. `mockRunning` is hoisted to the outer describe.
**The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved
its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so
the other eight modes fell back to claude's `❯`. That is also starship's default shell
prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line
`❯ npm run build` sits on screen for as long as the command runs, the verifier reads it
as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to
nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N
prompt, `read -p`, an installer or a pager, where it takes the default. The module's own
fileoverview already stated the rule this broke. Now `?? ''`, which
`promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI
that actually declares a composer.
**My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing
three numbered rules.
The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared
in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told
an integrator to retry with `confirmed: true` for both questions, which is precisely the
thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md`
and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR`
multi-user consequence, both user-visible and both previously absent, and #454's gained
the one exception to its own claim: a Custom Endpoints launch ignores the Instance count
stepper and always starts one session.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CLAUDE.md gained the admin-only note when the key joined claude's privilegedEnvKeys;
architecture-invariants, which is where the exact-key allowlist rule is documented in
depth, still described the pre-change world. The reboot-restore half is the one worth
writing down: a non-granted owner's already-persisted CLAUDE_CONFIG_DIR is stripped on
restore, which moves that session back to the default Claude account with no error.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
More from the review of 5fc391a4, all documentation rather than behaviour.
The changeset was 1602 words of development log, written as the PR grew, with bullets
and loose paragraphs interleaved. That text becomes CHANGELOG.md and the GitHub release
body verbatim, so it is now one user-facing account of what the feature does and what
the real-server work bought, at roughly a fifth the length.
docs/api-reference.md promised a `cmd` field on running-status that the route
deliberately strips (it carries model paths and can carry --api-key).
Two places claimed the apply routes validate `modelId` against the endpoint's
discovered models. Neither does. Dropped the claim rather than adding the check:
discovery can be up to five minutes stale, so a 400 there would refuse a launch that
actually works, and a typo'd id already fails on the CLI's own first request. CLAUDE.md
now says so explicitly, since the absence is the surprising part.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two findings from the review of 5fc391a4, both fixed here rather than sent back.
**The API-key trust seed never matched a real key.** `seedApiKeyTrustFile()` wrote the
key verbatim into `customApiKeyResponses.approved`, but Claude Code stores and compares
only the last 20 characters (`key.trim().slice(-20)`, applied on both the write and the
lookup). For any real key the seed missed, so claude stopped at the interactive
"Detected a custom API key in your environment" prompt, whose default is
"No (recommended)": the launch hangs, or silently refuses the key this feature just
injected and falls through to an OAuth login the isolated config dir does not have. It
survived review because a keyless llama.cpp/llama-swap endpoint uses DEFAULT_API_KEY
('local-dummy-key', 15 chars), where slice(-20) returns the whole string and the seed
matches by accident, and every test used a key shorter than that. Now truncated through
`truncateApiKeyForTrustFile()`, with a test using a 57-character key that also asserts
the full credential never reaches that second file.
**One `confirmed` flag answered two different questions.** The context-floor warning
("this model's window is below what this CLI needs") and the swap-conflict warning
("loading this unloads the model another session is using") shared it, and the context
check runs first, so a user clicking "launch anyway" past the context warning silently
consented to evicting someone else's model. They are about different people, so an
answer to one is not consent to the other. Both routes now read `confirmedContext` and
`confirmedSwap` independently; the legacy `confirmed` still means both, because it
shipped in this feature's HTTP-API-only cut and an existing caller must keep working.
The frontend answers each question with its own flag and accumulates them, on the
one-shot path, the restart path and the batch carry-forward alike.
Also from the same review: the swap-confirm dialog no longer renders " are currently
using ..." when multi-user scoping leaves the affected-session list empty (the swap is
blocked regardless of ownership; only the NAMES are scoped), and the per-endpoint
llama-swap log tails are closed in `WebServer.stop()` instead of only by the idle sweep
whose interval that same teardown disposes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Conflict resolution against the five PRs that landed while this was in review, plus
the items left for merge on the thread.
The real one was `session-ui.js`. #454 refactored all eight non-Claude `run*()`
functions to funnel through one `_launchQuickStartInstances()` helper that does the
POST itself, while this PR replaced that same POST in each of them with
`_quickStartWithCustomModelConfirm()`. Resolved in the helper rather than seven times
over: the helper now goes through the confirm path, and each body builder carries the
`customModel` spread. `runAntigravity` deliberately does NOT, since antigravity's
`customModelInjection` is `unsupported`; parity with this PR's own per-mode choices is
asserted rather than assumed.
That merge creates a question neither feature had alone: the confirm dialog now runs
inside a loop that can launch up to 20 instances. Both questions it can ask (context
window too small, and loading this will unload the model another session is using) are
decisions about the ENDPOINT, and every instance in a batch targets the same one, so
the answer is taken once and carried to the rest. Without that a 20-instance launch
asks the same question 20 times.
Also: `sse-events.ts` is 161 constants (master added two for remote wake, this adds
one, verified by counting rather than by arithmetic), `server.ts` keeps both new SSE
prefixes, the two comments pointing at code that no longer exists are corrected, and
CLAUDE.md's SSE and route counts move to 161 / ~236 / custom-model (6).
`pumpLlamaSwapLogTail`'s unparsed remainder is now capped at 64 KiB. It only shrank at
a `\n\n` frame boundary, so a backend that streams without one would grow it for the
life of a deliberately indefinite connection.
NOT changed, deliberately: the context warning and the swap-conflict warning still
share one `confirmed` flag with the context check first, so confirming "launch anyway"
on a too-small context also skips the "this unloads it for another session" ask. That
is the author's documented choice and the reviewer's own note calls it minor. Both
fixes are worse to make here than to defer: separate flags are new wire surface landed
unreviewed during a release, and reordering the checks adds a network round trip to a
path that currently short-circuits. Raised as a follow-up instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The MAC-count limit lived in two places that disagreed. RemoteHostSchema.wakeMac's
128-character cap admits seven comma-separated MACs while parseMacList takes at most
four, all-or-nothing, so a five-MAC value validated, was written to remote-hosts.json,
and then resolved to NO wake target: POST /api/sessions/:id/wake answered
"No wake-on-LAN target configured for this host" and the banner offered "Configure WoL"
for a host the user had just configured. MAX_WAKE_MACS now lives in
src/config/remote-wake-limits.ts and both sides refine against it. Its own module
because src/remote-wake.ts is import-fenced to session-routes.ts and server.ts (the
wiring guard that stops a watcher waking a host), and because schemas.ts must not drag
dgram/net/child_process into every request-validating module.
The documented 40 s request budget also omitted the wake's own cost. A `command` target
is bounded by REMOTE_WAKE_COMMAND_TIMEOUT_MS and runs BEFORE the readiness poll, so a
slow one pushed a wakeCommand host's worst case to ~68 s, past the 60 s
proxy_read_timeout the budget exists to stay under. _wakeAndWait now subtracts the
wake's measured elapsed time from the readiness budget, floored at one poll interval so
a wake that ate the whole budget still gets one probe. A magic packet is effectively
instant and is unaffected, which is why live testing never saw it.
Also: the two new endpoints are documented in docs/api-reference.md with the import
fence stated as the rule it is, CLAUDE.md's frontend module count moves to 34, and the
release changesets carry the Thanks section.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#451 cleaned two things on copy. The trailing trim is right and every native
terminal does it. The shared leading-indent strip is this project's own rule,
and it is dropped here rather than shipped.
Measured against the shipped transform over 401,445 three-row windows across
1,010 tracked files in this repo, it fired on 73% of them: 92% inside a YAML
workflow, 76% over `git log` output, 48% in a TypeScript source. No width
threshold separates a margin from content because they are the same widths, a
live Claude Code pane's own margins measuring 2 and 5 columns while the most
common non-TUI shared run is 4. The failure modes are not symmetric either: a
wrong trailing trim costs nothing, while a wrong dedent silently deletes
information that was on the screen, with nothing in the clipboard to hint at
it, on git log bodies, on indented code read out of cat (semantic in Python),
on git diff context rows where the leading space is the marker, and on stack
traces.
It also could not be made self-consistent cheaply. Whether the first row joined
the measurement depended on the mousedown COLUMN, which the user never sees, so
one block of three rows produced three different clipboard results; and the
flag read getSelectionPosition().start, which is xterm's mousedown anchor and
is never normalised, so dragging UP through a block read it off the bottom row.
The PR's test stub hardcoded a downward drag, so its suite could not express
that case.
The transform, the wiring, the tests, the invariants, CLAUDE.md, the wiki page
and the changeset all move together. The test block now pins the ABSENCE as a
contract, with the git log, Python and git diff cases as its examples, so this
is not re-derived later. If it is ever revisited, the one qualification that
measured clean is painted trailing padding: zero false positives over all
401,445 windows.
Also from the review: the comments and invariant rule justifying the
padding-only clear described the pre-change code (the Ctrl+C gate reads the
CLEANED selection now, so such a selection falls through to the PTY on its own
and the clear is feedback rather than protection), the new 'Nothing to copy'
toast gained its zh-CN entry, and the invariants paragraph no longer repeats
its own opening sentence.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#454: the behaviour the PR adds had no test, so a regression test drives
runGrok() at tabCount 3 and asserts three quick-start POSTs with sequential
w<n>-<case> names (verified to fail against master's session-ui.js). Each
caller now reads the count BEFORE its opening banner and announces it there,
the way runClaude() already did, so a launch no longer prints two headers and
a launch with another session already active still says how many are starting.
runClaude() calls the shared _readTabCount() instead of its own copy of the
1..20 clamp, and that helper optional-chains the element read, since hoisting
it above each caller's try block would otherwise let a missing #tabCount throw
where the launch-error path cannot report it.
#435: sizeMovedUnderLoad derived from data.source alone. `mux-visible` is not
sufficient: a failed display-message cursor query makes capturePaneBuffer skip
the snapshot repaint and return the raw capture, which the route still labels
mux-visible, so a size that moved during such a load bought a full forced
reload to repair a frame that was never positioned. It now tests
Number.isFinite(data.captureRows) like its two siblings.
Plus the invariants and CLAUDE.md lines promised on #435: a visible capture
reports its geometry and omits it when nothing was positioned, the comparison
runs on mux-visible only, and the replay is capped at one attempt and latches
per session when it cannot converge.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Merged upstream/master (22 commits: reboot-restore recovery feature,
terminal keycode229 recovery work, install.sh/CLI-catalog generator
changes, CHANGELOG/version bump to 1.30.0) into this branch. No
conflicts; git auto-merged every overlapping file (CLAUDE.md,
docs/api-reference.md, app.js, index.html, styles.css, routes/index.ts,
session-routes.ts, schemas.ts, server.ts).
Two required fixes from the latest review:
1. privilegedEnvKeys widening (stock.ts) changes behaviour outside this
feature. The reviewer decided to keep both CLAUDE_CODE_MAX_CONTEXT_TOKENS
and CLAUDE_CONFIG_DIR listed (types.ts's rule that every traffic-
redirecting var this feature introduces must appear there stays
literally true), and asked for the real consequences documented
instead of hidden:
- Corrected session-env-clamp.ts's fileoverview, which stated the
opposite of what the code now does (reboot-restore's clamp call
used to be able to strip nothing for claude; it now strips a
persisted CLAUDE_CONFIG_DIR for a non-granted owner).
- Corrected the rationale comments in stock.ts: privilegedEnvKeys
has exactly one consumer (ownerClampedEnvKeys, feeding the
generic envOverrides clamp on create/quick-start/reboot-restore),
not the custom-model routes.
- Added a CLAUDE.md line to the CLAUDE_CONFIG_DIR gotcha covering
the admin-only-in-multi-user-mode and reboot-restore-strips-it
consequences.
- Added a "Claude multi-user clamp" test next to the existing
DeepSeek/OMP ones, pinning the new stripping behaviour.
2. GET .../running-status (custom-model-routes.ts) no longer passes
the raw llama-swap `cmd` field (the literal launch line, which can
carry model paths and --api-key) to the browser -- the frontend
only ever reads model/state, cmd exists solely for server-side
parseCtxFromCmd() during discovery. Added a test asserting the
response never contains cmd or a planted secret.
Also regenerated config/clis.stock.json and install.sh's catalogue
block (npm run generate:cli-catalog) to clear drift introduced by the
upstream merge, since it was failing the sync check.
Left to the reviewer, as they said they'd take at merge: the two
"comments pointing at removed code" cleanups, the two stale CLAUDE.md
counts, and the small items list (mode==='claude' frontend branch,
isCliAvailable() unknown-id gap, shared confirmed flag ordering,
one-shot cancel toast severity, pumpLlamaSwapLogTail buffer cap).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
The first shape of the Enter loop read the composer BETWEEN two short waits,
and tested `wait.ended` (the session exiting) where it meant `timedOut`. A
`stop` that fired while no wait was open was lost, since signals have no
history, and a re-wait that had already resolved on `stop` fell through into
another wait that could never see the edge again: measured twice, the answer
was on screen and sendwait ran its whole 580 s slice anyway.
The long re-wait (a tagged duplicate of the original frame) is now registered
first and kept open in the background for the rest of the call; the loop reads
the composer and re-sends Enter beside it, stops when the prompt has left or
the wait's response has landed, then returns that response. Measured: the
stranded prompt got one extra Enter and sendwait returned on `stop` at 36 s.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Review round 3 on #439.
- The attachRemoteSession branch of POST /api/sessions ran `ensureHostAwake`
before the multi-user gates, so a non-admin could have any configured
host's `wakeCommand` spawned (or a packet broadcast) and the request held
for the wake budget, then be refused for the workingDir. The admin gate
now comes first, before the host is even looked up; remote hosts are
admin-only infrastructure everywhere else. Route test: wake spy empty,
403.
- The non-wait input route answers `{buffered:true}` when the registry took
the chunk and `{buffered:true, dropped:true}` when it was over the cap
and is gone (`RemoteInputOutcome` gains 'dropped'); additive to the bare
`{}`.
- The send-and-wait path answers OPERATION_FAILED when the host never comes
back, like create and attach, instead of writing into the stalled pane
and reporting delivered:true plus a timeout.
- The flush writes with `fromUser: true`, so a first prompt buffered
through a wake can still name the tab.
Docs: api-reference (input route), remote-sessions.md (two invariants),
CLAUDE.md key pattern.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
Claude Code 2.1.277 takes typed text the moment its composer paints but
ignores Enter for the first 30 to 50 seconds after it (measured 2026-09-19
through the input route: an Enter at 28 s stranded the prompt, one at 51 s
submitted it). The text+Enter pair `sendInput` sends 50 ms apart therefore
left every programmatic prompt sitting unsent, and every waiter burned its
timeout on a turn that never started.
Server: `SubmitVerifier` (session-submit-verifier.ts), armed from
`writeViaMux` for every mux write that carried a carriage return, reads the
pane on a 2 s to 60 s schedule and re-sends Enter only while the last
composer line (the CLI's own prompt glyph) still holds the head of what was
sent. An empty composer, other text, or no composer line at all ends it; a
newer write replaces the schedule.
Skill: `sendwait` gets the same loop (`_composer_text`, no-break space
stripped by its bytes for BSD sed) for servers that predate this, and the
preamble version moves to 1.30.1 so seeded agents pick up the fresh copy.
SKILL.md's heredoc and the plugin mirror are regenerated.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`resizeRetry` caps the recursion inside one select and says nothing about the
next one, so a pane this browser cannot size reported the same mismatch on
every select and bought the same failed repair each time: two fetches per tab
switch for the life of the page, measured as a running count of 2, 4, 6 across
three selects. That is the case this branch describes as happening every time
rather than occasionally, a phone whose resize `Session.resize` declines while a
desktop claim is live, and it is not the only one — any pane Codeman cannot size
lands there, including one a second tmux client is also holding. Each wasted
pass costs another `capture-pane`, which is `execSync` and blocks the server's
event loop, plus a reset and chunked rewrite, a discarded snapshot and cache
entry, and a dropped and reopened WebSocket.
`_geometryRetryUseless` mirrors the existing `_fullHistoryRepullUseless`: a
retry pass whose frame still does not fit adds the session, geometry that fits
removes it, and the replay gate consults it. The proof has to come from a retry
pass rather than a first one, because the retry ran at the size that stuck and
the pane ignored it. Clearing on a fitting frame is what stops a pane that
becomes sizeable again, once the desktop tab closes or its claim goes idle, from
staying permanently unrepaired. The race case never reaches the latch, since it
converges on its first attempt.
The new browser case walks all of that: three selects reading 2, 3, 4 instead of
2, 4, 6, then a fitting frame, then a mismatch diagnosed afresh. Without the
gate it fails on the second switch with `expected 4 to be 3`.
Rebased onto master, which has moved to 1.30.0 and taken #436. The one conflict
was `config/test-suites.ts`, where both branches appended a glob to
`BROWSER_TEST_GLOBS`; both are kept. Everything else merged clean, #436's own
changes to the same buffer-load path included.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
On a touch device the characters the user has typed live only in the local-echo
overlay until Enter; they have never reached the PTY. The replay re-enters
`selectSession` with `forceReload` on the session that is still active, and that
branch nulled `activeSessionId` before `_cleanupPreviousSession` ran. The flush
there is guarded on a session it can still see, so it was skipped, and the
unconditional `_localEchoOverlay.clear()` that follows took the characters with
it. Measured in chromium against the previous head: typing into the overlay and
then making the call the replay makes left `pendingText` empty with nothing
crossing into the delivery layer on either transport.
The flush moves into `_flushLocalEchoTo(sessionId)`, called from both
`_cleanupPreviousSession` and the `forceReload` branch before it nulls the id.
The session is a parameter because the two callers mean different ones: cleanup
flushes to the tab being left, the branch to the tab being reloaded.
This was reachable before this branch, through the one gesture that already
takes the `forceReload` path on an active session. What is new is that nothing
the user does triggers it. The replay fires on its own the moment a tab switch
finishes, which is exactly when someone typing into a still-loading terminal has
text in the overlay, and on a phone beside an active desktop tab that is every
tab switch.
A seventh browser case pins it: it forces the overlay on, since headless
chromium reports no touch support and the case would otherwise pass vacuously,
asserts the typed characters really are sitting unsent, then triggers the replay
and asserts they reached the session. Without the fix it fails with nothing
delivered at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three follow-ups to the source gate, each one measured rather than reasoned.
A pane already drawing at the size the client just requested is left alone. The
replay runs at `dimsAfterLoad`, so it can only change what is on screen if the
pane was drawing at some other size; when the reported geometry already IS that
size, the second pass captures the identical frame and pays a full reload to do
it, including a visible re-flash, a dropped and reopened WebSocket and a deleted
xterm snapshot. That equality is the signature of a clamp rather than a race:
`getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` does not, so a
terminal narrower than 40 columns or shorter than 10 rows reports a pane
permanently bigger than itself and replayed on every tab switch without ever
converging. A race never produces the equality, since its premise is that the
pane was still at the size it was asked to leave. The declined-resize case does
not produce it either, so that one still costs the single capped attempt and
needs the pane-ownership question this does not touch.
The full-history re-arm is unreachable and now says so. A pass that consumed the
flag sent `full=1`, and the route answers `full=1` with `mux-full-history` or
`history`, never `mux-visible`, so the source gate already rules out every such
pass. The line stays for the invariant, but its comment no longer reads as if a
page load retries, and the suite pins that it does not.
The response no longer reports geometry for a body that carries no capture. The
full-history path writes `capturedGeometry` from the cursor query and then
returns '' for a pane holding nothing visible, which drops the source to
`history` with the geometry already recorded: a `full=1` request whose capture
reported 100x50 and returned nothing answered `source: "history"` with both
fields set. Nothing acted on it, because the client ignores geometry on any
other source, but the field said a frame had been drawn at a size when none had.
The browser stub now derives `source` from the request the way the route does,
rather than answering `full=1` with `mux-visible`, which the route cannot
produce. Each case reaches a visible-frame response the way production does, by
not being the first select of the page. Three cases pin the new behaviour and
each fails without its guard: the clamp case sees two fetches instead of one,
the scope case and the full-history case both see a replay the gate forbids, and
the width case sees one fetch instead of two.
The changeset now describes the change from 1.29.x rather than the difference
between the two commits on this branch.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Only a visible-frame capture positions its rows absolutely, so only that frame
can be damaged by a terminal of the wrong size. A `full=1` body is linear
scrollback closed by a relative cursor move, which is relative precisely so the
browser's row count need not match the pane's, and a `history` body is the byte
stream, which carries no row alignment to protect. The geometry comparison ran
on all three, so it fired most often on the one response it cannot help:
`_fullHistoryLoaded` is empty on the first select of every non-shell session per
page, and a session whose pane a desktop tab holds too tall to ever fit then
paid a second whole-scrollback capture, reset and replay on every page load and
every first tab switch.
`framePositionsRowsAbsolutely` gates both the captured-geometry comparison and
`sizeMovedUnderLoad`. A size that moved under a byte-stream or scrollback replay
is healed by xterm's own reflow plus the SIGWINCH the trailing `sendResize`
already sends.
A pane WIDER than the terminal damages the same frame a second way, so
`captureCols` is now compared rather than only logged. `formatPaneSnapshot`
paints each row out to the pane's own width, so a narrower browser wraps every
painted row, and the wrap on the last one scrolls the whole frame up by a row.
The terminal response no longer falls back to `session.ptyCols`/`ptyRows` when
the capture reported no geometry. The cursor query is what produces the absolute
addressing in the first place, so a capture that lost it returned a raw frame
that was never positioned, and a byte-history response was never positioned
either. Naming the session's own PTY size there described a frame that does not
exist and invited a repair for damage that is not present. `_ptyCols` is also
written only by `resize()` while the PTY is spawned at the size queried from
tmux, so it can be wrong on its own terms. Both fields are now absent instead,
and the `Session` getters added for that fallback go with it.
Two browser cases cover the new behaviour and each fails without its fix: a
`mux-full-history` response with both dimensions mismatched asserts one fetch
(two without the gate), and a `mux-visible` response wider than the terminal
but short enough to fit asserts two (one without the width comparison).
Corrects a claim in the comment above `capturedGeometry` in tmux-manager.ts.
Both replay paths do not address rows absolutely; the full-history one ends in a
relative move, which is the whole reason the gate is right.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A visible-frame capture repaints each row at an absolute position, counting up
to the pane's height. A terminal shorter than that clamps every address past its
own height onto its last line. The overflow rows then overwrite one another, and
the rows underneath are lost. Replaying a real 50-row capture into a 30-row
terminal rendered 28 lines of a 45-line command and drew the frame twice.
Nothing in the response said what height the frame was built for, so the client
could not detect this. A capture now reports the geometry it was really taken at
through `capturedGeometry` on `PaneCaptureOptions`, and the terminal response
carries it as `captureCols` and `captureRows`. When the captured pane is taller
than the terminal, or the size that produced the capture did not survive the
load, `selectSession` replays once at the size that stuck. `resizeRetry` caps
that at one attempt, so two competing fits cannot trade replays forever.
The retry re-arms the full-history flag only when the pass that ran had consumed
it. A tab switch takes the bounded tail, so its retry takes the tail too:
clearing the flag unconditionally would upgrade that switch into a fresh
scrollback capture the user never asked for, which the route's own comments put
at tens of megabytes.
What this repairs is a capture that won a race against the resize meant to
precede it. It does not repair a capture whose pane was too tall because
`Session.resize` declined the resize outright, which it does for a small
viewport while a desktop viewport's size claim is live. The retry re-sends the
same declined resize and captures the same pane, and `resizeRetry` then stops
it. Repairing that means changing who owns the pane size, which is a policy
question this does not touch. The reported geometry still helps there, because
the client can see the mismatch at all rather than being blind to it.
Follows #395, #396 and #397, which fixed the other ways the replayed frame and
the terminal could disagree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
runOpenCode(), runCodex(), runGemini(), runAntigravity(), runPi(), runOmp(),
runGrok(), and runDeepSeek() all ignored the "Instance count" stepper next
to the Run button and hardcoded a single quick-start call — bumping the
counter to 2 or 3 while on any of these modes silently launched exactly one
session, with no error. Only runClaude() ever read it.
Extract the shared launch-N-sessions-and-select-the-first loop into
_launchQuickStartInstances(), reused by all eight modes, and _readTabCount()
for the shared clamp-and-parse. Each mode still builds its own quick-start
body (config differs per CLI), just via a closure passed to the shared
loop instead of a single inline fetch.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Review round 2 on #439.
1. The bare TCP probe connects to host:port, which a host behind a jump host
or SOCKS proxy does not answer even while ssh works. Acting on that
verdict drew a permanent banner over a healthy session, replaced a real
"needs tmux" error with "not reachable" in quick-start, and - with a wake
target - buffered every HTTP input for the life of the session, since the
readiness poll could never succeed. `WakeableRemote` now carries
`jumpHost`/`socksProxy`/`extraSshOptions`, and `isProbeable()` turns such
a host into reachability-UNKNOWN: input is delivered, `checkReachable` /
`checkHostReachable` answer `null` (never `false`), `ensureHostAwake`
returns `'unprobeable'` (handled like `'no-target'`), the quick-start gate
fires on `=== false` only, and `GET …/reachability` reports
`reachable: null, probeable: false` so the banner has nothing to key on.
A wake target can still be fired for it, blind: no readiness poll, no
reattach, no toast - the response says only whether the packet went out.
2. `'remote:'` joins the session-scoped SSE prefixes. The create/attach wake
has no session yet, so the registry names the requesting user
(`ensureHostAwake({ requestedBy })` -> `username` in the payload) and
`deriveSseHint` routes on it; with neither it fails closed to admins.
Single-user mode is unaffected.
Smaller, from the same review:
- A flush write that fails now drops the remaining buffer (logged) instead
of retaining it: the wake still resolved and marked the host reachable, so
the retained chunk waited for the NEXT wake and was replayed hours later,
after everything typed since. Same policy as the oversized paste.
- The banner polls on tab activation (a user action) and on its 30 s timer
only for a host with a wake target; a timer connecting to a host Codeman
cannot wake is the traffic invariant #2 rejects keepalives for. A proxied
host is never polled.
- `probeRemoteHostReachable`, `runRemoteWakeCommand` and the default UDP
socket refuse under VITEST, as remote-files.ts does. The guard caught a
leak on the spot: `createDefaultRemoteWakeDeps({ probe })` overrode the
probe but still polled readiness with the real one, so the shutdown test
had been connecting to a production address. The poll now uses the
injected probe.
- docs/remote-sessions.md is additions only again (the reformatting is
gone); the architecture-invariants overlap resolved itself in the merge.
Live, against a throwaway instance with a non-routable ghost host: proxied
-> no probe, no wake, the genuine ssh error after 10 s; direct (control) ->
probe, magic packet, "did not come back" after the 40 s budget.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
Resolves CLAUDE.md count tables (route counts recounted on the merged
tree: 235 handlers, sessions 37) and keeps both the host-wake and the
reboot-restore banner in index.html.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG
The four edits the review said would be folded in at merge, none of them
code: the changeset becomes one user-facing paragraph, since it is what
CHANGELOG.md and the release notes print; the `_bufferLoadFinishOpts` comment
now names the second contributor to the duplicate window (`captureActivePaneBuffer`
is `execSync`, so anything painted into the pane before the server read it is
in the capture and is broadcast after the reply) and says why a `history`
payload keeps the pre-existing discard when its exposure is the same; the
`_finishBufferLoad` doc block moves from above `_beginBufferLoad` onto the
function it documents; and the test file's header describes both rules the
file now pins instead of only COD-144.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Blocker 1: the loading banner hides itself ~200ms after it reopens.
- _showCenterStatus reuses one shared DOM node; dismiss() scheduled
el.hidden = true 200ms later with nothing to cancel it. On the
Claude path, switchingToast.dismiss() is followed by one same-
origin request (5-30ms locally) before _watchLlamaSwapLoading opens
the new banner -- well inside that window -- so the stale timer
fired against the shared node and hid the fresh banner, leaving the
whole model-load wait with no progress text, no log line and no
reachable Cancel button.
- Fixed by parking the pending timeout on the element and clearing it
at the top of _showCenterStatus. Added a regression test that
reproduces the exact repro (open, dismiss, reopen 20ms later,
advance past 200ms) alongside the existing Cancel-button DOM tests;
confirmed it fails without the fix and passes with it.
Blocker 2: the swap-conflict warning named other users' sessions.
- Both affectedSessions scans (POST .../custom-model and quick-start)
walked the whole session map with no ownership filter, so in multi-
user mode a non-admin pointing their own session at a shared
endpoint learned another user's session name and id -- which with
autoNameSessions on is that user's own prompt.
- The swap is still blocked pending confirmation regardless of
ownership (a foreign session is just as real a disruption); only
which ones get NAMED back to the caller is scoped, via the
already-imported canAccessOwned. Added a two-owner test to
test/routes/session-custom-model.test.ts covering both the
foreign-owner (blocked, not named) and same-owner (named) cases.
Smaller ride-along fixes:
- server.ts boot recovery now passes contextLength into
applyCustomModelInjection, so CLAUDE_CODE_MAX_CONTEXT_TOKENS is
correctly rebuilt into _envOverrides after a restart instead of
surviving only because tmux retains the old setenv.
- pumpLlamaSwapLogTail's finally now deletes by IDENTITY, not just by
key, so an aborted pump finishing after a newer entry was created
for the same endpoint can no longer delete that newer entry and
orphan its connection.
- docs/custom-model-endpoints.md now notes that clearing a custom
model removes injected keys by name, including CLAUDE_CONFIG_DIR --
so a session that also had CLAUDE_CONFIG_DIR set via envOverrides
(the per-client-account case) silently falls back to the default
account on clear.
Left for later, as flagged in the review itself: the quick-start
case-scaffolding/cancel ordering (real behavioural reordering across
a large handler, too risky to make without a live re-test), and
retiring runCustomModelEntry's mode === 'claude' branch behind a
launchStrategy registry field (explicitly deferred by the reviewer to
"the next one").
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
xterm hands back whole screen rows and trims only the cells that were
never written to, so the real spaces a full-screen TUI paints across the
unused part of a row count as content and reach the clipboard. Measured
against Claude Code in a 282-column pane, single lines arrived carrying
138 trailing spaces, and every line carried the two-space transcript
indent as well. Windows Terminal, iTerm2 and GNOME Terminal all trim that
for you, decideAutoCopy already calls a wall of spaces "never what the
gesture meant", and _selectTouchSelectionLine already treats those cells
as padding — the mouse and keyboard paths never had the same rule.
CodemanCopySelection.clean lives in constants.js beside decideAutoCopy,
its pure sibling. It drops the trailing run from each line, and removes
the leading run only where every selected row shares one. A selection of
a single row keeps its run, because one row shares nothing with anything
and stripping it would silently reindent one line of `git log` body text
or one line out of `less`. A drag that began inside a row keeps its
partial first line untouched and out of the measurement, which otherwise
pins the shared run to zero and leaves every following row indented.
Every pass over a line is a scan rather than a regex. `/[ \t]+(\r?)$/` is
quadratic on a line whose spaces are followed by a non-space character,
which is what right-aligned or centred TUI content looks like: measured
over 50 000 rows with a 280-column run it took 2.9s, against 1.3ms for
the scan, and a 2 000-column run took 16s. The scan is also the faster of
the two on an ordinary padded row.
cleanedTerminalSelection in terminal-ui.js is the half that needs the
live terminal. It returns a COLUMN selection untouched: Alt+drag makes
one, and a rectangle's rows lining up is the point of the gesture, so
both halves of the clean would destroy it. xterm exposes the mode nowhere
public, so the check reads terminal._core._selectionService, the way this
file already reads terminal._core for cell dimensions, and cleans
normally if a future xterm renames the field. A test pins that assumption
against the library rather than against a stub repeating the literal.
The Ctrl+C chord decides on the cleaned selection, not the raw one. A
drag across the blank part of a row selects real padding spaces, so the
raw text is truthy, and testing it would spend that press on a copy of
nothing and make the user press again to interrupt. A padding-only
selection is now dropped and the press falls through to the PTY, while
Ctrl+Shift+C still never falls through. copyTerminalSelection gates on
trim() for the same reason, since a multi-row drag across padding cleans
to line breaks alone and a bare newline pasted into a chat composer
submits it.
All four of the main terminal's copy paths go through it: the Ctrl+C
chord, right-click, the phone selection button and Auto Copy. The
browser's own Edit menu copy, a disabled copy shortcut and the subagent
windows still copy raw rows, as they did before, and the invariants doc
now says so rather than claiming every copy is cleaned. Auto Copy
resolves its own toggle before it reads the selection, since it is off by
default and a selection can run to the 50 000-row scrollback ceiling.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The JSDoc on `_syncStickyScrollBaseline` said `selectSession` deliberately ends
at the bottom, so the baseline the replay samples is already true there. It
does not. `selectSession` calls `scrollToBottom()` after the write and then
ends at `scrollToLastNonEmptyLine()` (app.js:6512), which targets
`lastNonEmptyLine - rows + 2` and therefore parks ABOVE `baseY` whenever the
replayed frame keeps trailing blank rows — which a full capture does on
purpose, since no transform that can delete a line may run over one.
Its baseline really is a stale true. What covers it is the sticky snap itself:
since de864e7d that snap fires only when the flush found the viewport already
at the bottom (`preserveViewportY === null`), which a parked selectSession
viewport is not. That commit landed on master after this branch was cut, so
the guard arrives with the merge rather than being present here.
`_onSessionClearTerminal` is unchanged in the comment and was correct: it
resets and rewrites with no scroll afterwards, so it does end at the bottom.
Comment only; no behaviour change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The first version of this fix decided the flush policy in `selectSession`
alone, and a later pass found it still covering one path of four. Nothing in
the CI gate stops a fifth path, or an inlined `{ flushQueued: true }`, from
splitting that policy up again — the browser suite that would notice is
excluded from `npm test`.
A static scan over `selectSession`, `_onSessionNeedsRefresh`,
`_onSessionClearTerminal` and `_maybeRefetchFullHistory` asserts each one asks
`_bufferLoadFinishOpts`, reusing the `methodBody` slice the sticky-scroll guard
already needed. Verified by inlining the policy back into
`_onSessionClearTerminal`, which fails it by name.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A capture load now replays its queued tail, and that replay runs through
`batchTerminalWrite`, which samples `_wasAtBottomBeforeWrite` before it queues.
It runs inside `chunkedTerminalWrite`, before that promise resolves, with the
terminal freshly reset and rewritten — so the sample is always true. The caller
then restored the reader's position and the next `flushPendingWrites` scrolled
straight back to the bottom off the latched flag, undoing it. The only thing in
the way was `_hasRecentUserScrollUp()`, a 1500ms window a server-triggered
refresh is usually past.
`_syncStickyScrollBaseline()` re-takes the flag from wherever the viewport now
sits, and the two paths that restore a position call it right after doing so:
`_onSessionNeedsRefresh` and `_maybeRefetchFullHistory`. Those are the paths
#259 and #205 exist for, and they are also where a non-empty queue is most
likely, since a needsRefresh fires when output is flooding. Re-taking rather
than suppressing the sampling: suppressing leaves whatever stale value the flag
held from before the load, which on the full-history re-pull has no reason to
be false. `selectSession` and `_onSessionClearTerminal` deliberately end at the
bottom, so the sampled true is already the truth there and they do not call it.
`_bufferLoadFinishOpts` gains the coverage the CI gate can see: both mux
sources flush, `history` does not, and a payload naming no source does not.
Its only coverage was the browser suite, which CI does not run.
The JSDoc and the changeset now record the one duplicate window this cutoff
cannot close. The server appends output to the byte buffer in the same tick it
emits, but broadcasts on a batch timer — 8ms over WebSocket, 16 to 50ms over
SSE — so a batch pending when `capture-pane` ran leaves the server after the
reply and is replayed although the capture holds it. It is one batch interval
wide against a recovery window spanning the whole chunked write, and closing it
means flushing that batch server side before the capture.
The second browser test asserts its session was created, so a failed create
fails it instead of passing with zero hits.
docs/architecture-invariants.md no longer claims the replay leaves the
queued-event discard window alone. That clause now describes what decides how a
load ends, the baseline rule, the batch window, and the three covering tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Blocker: .center-status-banner never actually disappears.
- Add `.center-status-banner[hidden] { display: none; }`, same trap as
`.home-sessions[hidden]`: the author-level `display: flex` beat the
UA `[hidden]` rule, so `dismiss()` set `el.hidden = true` and the
card stayed laid out at `opacity: 0` with its text/cancel/close
children still `pointer-events: auto` -- an invisible 442x67 click
blocker dead centre over the terminal until the page reloaded.
- Added a regression test pinning the CSS rule, and documented the
banner (10001) and the swap-confirm/context-warning modals (10010)
in CLAUDE.md's Z-index layers list.
Stale wording pointed at the reverted sticky-toast default:
- .changeset/run-menu-custom-model-picker.md, CLAUDE.md, and the
`.toast-message` comment in styles.css all still said "toasts
default to sticky" after 1f32128c put the flat 3s default back.
Reworded all three to describe the actual behaviour: one call site
passes an explicit `duration: 0`.
Smaller items from the same review:
- docs/api-reference.md said discovery failures answer
`502 OPERATION_FAILED`; OPERATION_FAILED is 422 per src/types/api.ts
and the error-code table earlier in the same file.
- The periodic re-discovery sweep (server.ts) never read
customModelEndpointsEnabled, so turning the feature off left
Codeman polling every saved endpoint forever. Added
readCustomModelEndpointsEnabled() (custom-model-routes.ts, same
shape as readPlanUsageTelemetryEnabled) and gated the interval
callback on it.
- Reverted the formatting-only Prettier pass docs/api-reference.md
picked up (table padding, *x* to _x_, JSON re-indent) by re-merging
the new Custom Model Endpoints section onto the pre-PR file, so the
diff is reviewable. No prose content was lost -- verified by diffing
the result against the pre-revert file (formatting-only) and against
the merge-base file (only the new section added).
- docs/custom-model-endpoints.md now states that a custom-model Claude
session's isolated CLAUDE_CONFIG_DIR loses the user's global
settings.json, user-level skills/agents/commands, and MCP servers
from ~/.claude.json -- only `projects` is symlinked back.
Design question left open in the review (does `confirmed: true` need
to be two flags so "launch anyway" on the context warning doesn't also
skip the llama-swap displacement warning): keeping the single flag, as
offered. The 20s displacement sweep still catches a resulting swap
after the fact, so it's a surprise rather than a silent failure, and
splitting it is real behavioural surface I have no way to verify live
in this environment.
`npm run test:browser` could not be run in this environment (no tmux,
no downloaded Playwright browser binary) -- none of its suite's files
touch code this fix changes, but it still needs a real pass before
merge, same as any frontend change.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
Four blockers from the 2026-09-18 review:
- PUT /api/model-endpoints/:id now merges modelContextLengths/
modelSizesGB back in from the stored record instead of trusting the
editor's body, so renaming an endpoint or changing its default model
no longer silently drops the context-window floor check and
CLAUDE_CODE_MAX_CONTEXT_TOKENS injection.
- custom-model:swapped-out is now session-scoped (added to
SESSION_PREFIXES) instead of broadcasting to every connected client.
- The quick-start custom-model path now hands setCustomModel() only
the endpoint's own injected env vars, not the full merged set,
matching the restart-in-place path — the full set put
CLAUDE_CODE_EFFORT_LEVEL back after the Session constructor had
already stripped it.
- The quick-start launchModel override for pi/grok/omp is now applied
generically via the registry's legacyConfigField, mirroring
Session._withCustomModelLaunchModel, instead of three hardcoded
mode === '<id>' branches a future CLI's injection recipe would miss.
Also scopes the sticky-toast default (item 5): reverted the blanket
"all error toasts are sticky" default, which had no container cap or
eviction, back to a flat 3s; the one message that needs a moment to
read (a failed custom-model apply) now passes an explicit
duration: 0 at its own call site.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ea59JhUmHBm1gRCsiYF33R
* chore: version packages
* chore: sync the CLAUDE.md version line to 1.30.0
The changesets bot does not touch this line, and pushing it to master
after merging the version PR starts a second Release run that has raced
the first before. Riding the bot's own branch keeps it to one push.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
Changeset text becomes user-facing CHANGELOG, so the #429 entry is cut
from five bullets of internal bash-array detail down to what the change
does for someone running the installer, as promised on the PR. The #441
entry loses its em-dashes, which are not house style. Adds an entry for
the maintainer fixes applied while landing #442, and the Thanks section.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Seven things, none of which changes what the feature does.
1. The rebuilt Session dropped `nameSource`, so the constructor re-inferred
it from the name: a session the user renamed by hand to something shaped
like `w<n>-<case>` came back as `placeholder`, and with auto-naming on the
next prompt overwrote their name. The route persists right after, so the
loss went to disk. `restoreMuxSessions()` already passes it.
2. The already-live sets were snapshotted once before a loop that awaits a
real `startInteractive()` per entry, so by the tenth entry the snapshot
was tens of seconds old and a conversation resumed by hand from the
Resume list in that window was invisible to it: two panes on one
transcript, the exact thing the check exists to prevent. Both sets are
now read per iteration, and the late case is spent rather than re-offered
for the same reason the batch case is.
3. Auto-resume no longer re-arms the pre-reboot `autoResumeAt` on this path.
The stamp predates the reboot and the pane is new, so honouring it meant
one click had every restored session type `continue` into itself about a
minute later, unattended, against the route header's own promise that a
restored session comes back idle and disarmed. The setting stays ENABLED,
so it re-arms on the next real limit message. A Codeman restart still
re-arms from the stamp, because the limit footer will not reprint on its
own; the new option exists only to tell the two paths apart.
4. `discardPartiallyBuiltSession()` now also calls `recordSessionStopped()`
and `ralphTracker.fullReset()`, the two teardown steps `_doCleanupSession`
performs that it was missing. Cosmetic, but a run left open reads as
still going in the away digest.
5. A restored claude session gets `seedAgentSessionPreamble()` like both
create paths, so the agent skill's bootstrap stays a two-line loader.
6. The heuristic's container comment was wrong in one direction and quiet
about the real gap: after a genuine host reboot a containerized Codeman
sees the host's short uptime and the banner does appear. What it cannot
see is a container-only restart, which is where this would help most.
7. The banner is hidden in a solo window, which shows one session and has
no tab strip to put restored ones in.
Also reverts 17 of the 18 hunks in docs/api-reference.md, which were
Prettier reformatting of prose the PR does not otherwise touch (docs/ is
outside the format glob), keeping only the Reboot restore section and
repairing the two continuation lines that reformat de-indented; renumbers
reboot-restore-ui.js to @loadorder 11.65, since 11.7 is admin-ui.js, which
loads after it; and gives the feature its CLAUDE.md entry plus a route
test for the multi-user workspace-forbidden branch, the only new rule that
had nothing behind it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The unit harness proves WHICH candidate gets forwarded; the ordering is
the half that shipped the bug, and only a real xterm shows it. The new
browser case dispatches the character's keydown, its composed insertText
and Enter's keydown in ONE page task, the shape an Android soft keyboard
delivers through a single InputConnection transaction, and asserts what
reaches the send path.
Verified in both directions on this machine: with the drain in place the
wire is `o\r`; with the drain removed (master's behaviour) it is `\r` and
the character is gone entirely, because by the time the zero-delay timer
runs xterm has emitted the `\r` and bumped the canonical counter past the
candidate's snapshot, so the candidate stands down. The other four cases
pass in both states.
CLAUDE.md now names the decision point, what it costs (a keydown decides
with less evidence than the timer did) and why that is safe for Enter,
and says that the pin lives in a suite the CI gate does not run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The caveat #429 added ends with "see the docs above", and "the docs
above" is CLI_DOCS[$i], which for DeepSeek is the upstream harness repo.
Per docs/deepseek-integration.md the harness ships only the web,
headless and base profiles, so following that link and running
`npm install -g @deepseek-ai/dsh` leaves the reader exactly where the
caveat is warning them about: a dsh that cannot drive a pane. What
actually resolves it is Codeman's own Run dropdown, which offers
"DeepSeek: add a terminal profile..." and installs one in a click.
The new wording stays generic for any future launcherProfile entry,
since Codeman is the thing being installed at all three call sites.
Also flips one word in the generator: the comment said "see
installCommandFor below" and that function is defined above it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Each of these was measurable and wrong: the CI note listed 5 excluded
Playwright tests where config/test-suites.ts has 9, never mentioned the
packages/xterm-zerolag-input run that follows the gate, and never
mentioned wiki-sync.yml at all; the format glob note omitted that lint
covers only src/**/*.ts; app.js is ~6.9K lines, not ~6.7K, and
voice-pcm-worklet.js is fetched from JS rather than sitting in the load
order; src/config/ holds 23 files plus the cli-registry/ subdir, not 21,
and nothing said that the repo-root config/ is a different directory;
the route count is ~232 with cases at 34, not ~228 with cases at 30.
Also adds the pointer to docs/wiki/ as the user-facing manual, which the
header describes every other doc surface but not that one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A second real reboot disproved the mechanism the previous commit was built
on. Typing `/exit` does not persist `pid: null`, and the session was
restored anyway.
The pid a session record carries is its `tmux attach-session` process, not
the agent. `/exit` ends the CLI inside the pane, `remain-on-exit` keeps the
pane, and the attach process stays alive throughout — so Codeman's PTY never
exits, no exit handler runs, and the record keeps both its pid and
`status: 'idle'`. The lifecycle log for the session that came back shows
created, started, stale_cleaned and recovered, with no exit event at all,
which is the proof: Codeman never learned the agent was gone.
So nothing durable distinguishes an exited agent from a session that was
idle when the power went, and this pass restores both. Ark0N/Codeman#446 is
about making Codeman notice the dead pane; contrary to what the previous
commit's message claimed, this genuinely does wait on that. Until a record
can say the agent is gone, the user dismisses or closes those sessions.
The rule itself is kept, because a record with no attach process does
describe a session that never started or whose pane died outright, and
refusing it is right. Only its documentation was wrong. The module header,
the branch comment and the test names now say what it recognises instead of
claiming the case it cannot see.
Refs #411
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Found by a real reboot, which is the first thing to catch it. Typing `/exit`
ends the CLI process and leaves the session record behind, and the
process-exit handler persists `pid: null` with `status: 'idle'` before
anything else runs. By status alone that is indistinguishable from a session
sitting idle when the power went, so the boot pass offered those sessions
back and a click spawned the agents the user had deliberately closed — the
exact case the eligibility rule exists to exclude.
The absent pid is what tells the two apart, and the plan step now refuses a
record without one, under its own `not-running` reason so the boot log says
why. On a healthy board every running session carries a pid; a record with
none describes an agent that is already gone.
Deliberately the conservative direction. A session that somehow persisted no
pid while genuinely running is not offered, and its conversation stays
reachable from the Resume list, which is where every session would be
without this feature. The opposite error spawns processes nobody asked for.
Ark0N/Codeman#446 covers the dead panes those exits leave behind, but this
does not wait on it: the rule belongs here whether or not the record's shape
changes later.
Refs #411
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The plan build runs inside the try that decides whether restoreMuxSessions()
succeeded, so a throw would be caught there, report restoration as failed,
and block the stale cleanup and layout reconciliation that follow. An
optional convenience would then break the recovery it exists to help. It is
guarded on its own now: the correct way for this to fail is an offer nobody
gets.
Refs #411
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ran the feature against a real server for the first time, on an isolated
instance, and two claims in the code turned out to be wrong.
A rebuild that fails after the session is registered was documented as
commonly caused by a CLI binary missing from a freshly booted machine's
PATH. It is not: the resolver finds its binary by absolute path, so PATH
never enters into it, and a server started without claude on PATH restored
every session normally. Nor does an un-enterable workspace fail — tmux falls
back to another directory and the pane comes up there. Neither obvious cause
throws, so the discard path is defended rather than expected, and the
comments now say that instead of naming a cause that cannot happen.
The four review rounds that shaped this path all reasoned about a trigger
none of them could test. The path itself is still worth having, since a mux
failure would reach it, but its comments should not claim a likelihood the
machine disagrees with.
Refs #411
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DeepSeek Harness's own bundled provider module
(@deepseek-ai/dsh-llm-deepseek) builds its request URL as
`${DEEPSEEK_BASE_URL}/chat/completions` with no `/v1` insertion of its
own (its real public API, https://api.deepseek.com, expects the
caller's base URL to already carry any needed prefix), while
llama-swap/llama.cpp only ever serves the OpenAI-conventional
`/v1/chat/completions`.
Confirmed two ways:
- Installed the real @deepseek-ai/dsh package (all its actual
published dependencies) into a scratch dir purely to read
dsh-llm-deepseek's source: `fetch(`${connection.baseURL}/chat/
completions`, ...)`, baseURL read straight from DEEPSEEK_BASE_URL —
the same grep-the-real-source bar pi/grok's fixes were held to.
- Live against the test-picker's llama-swap: `POST <baseUrl>/chat/
completions` -> 404, `POST <baseUrl>/v1/chat/completions` -> 200,
same endpoint. dsh's own error template ("DeepSeek API error (HTTP
${status})") reproduces the originally-reported
"dsh: HTTP_404: DeepSeek API error (HTTP 404)" exactly.
- New registry field `appendV1Suffix` (env kind only, deepseek's entry
alone — claude/gemini must NOT get it, since claude was already
confirmed working against the unmodified baseUrl). When set,
buildCustomModelInjection runs endpoint.baseUrl through the same
withV1Suffix() helper configDir-kind CLIs (pi/grok/codex) already
use, instead of writing it verbatim.
Not yet re-run end-to-end through a real dsh binary — no install
available in this environment (not in PATH, and the test-picker
container doesn't bundle it) — so this is source-confirmed and
live-verified at the HTTP level, not yet promoted to "verified"
alongside claude/opencode/pi/grok/omp. Docs (custom-model-endpoints.md,
the plan doc's confidence table, the wiki page, CLAUDE.md) all updated
to reflect this precisely rather than leaving the old "root cause not
identified" claim in place.
2 new/updated tests for the /v1 suffix (including idempotency against
a baseUrl that already ends in /v1) plus a corrected mock-server
contract test. Typecheck/lint clean; full suite shows no new
regressions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Full documentation review pass across the branch's 30 commits.
CLAUDE.md's Custom Model Endpoint Profiles entry hadn't been touched
since the initial backend+picker cut (3 early commits) despite 27
follow-up commits adding real behavior — it described restart-in-place
as universal (now claude-only; 7 other CLIs launch one-shot) and
claimed codex's Responses-API gap as a flat protocol break (now
re-verified as a more precise tool-calling gap). Corrected both and
added a new paragraph covering everything landed since: the llama-swap
conflict check, the after-the-fact swap-displacement sweep, the
/running-cmd-based context-length fix, the context-window floor
warning, skipFirstRunPrompts, the real-time /api/events-based log
status, and the countdown-to-Cancel-button change.
docs/api-reference.md's custom-model-endpoints section was missing the
running-status route, the requiresConfirmation/requiresContextWarning
response shapes, and POST /api/quick-start's customModel field
entirely (the primary launch path for 7 of 8 supported CLIs) — added
all three. Also fixed a real markdown bug in custom-model-endpoints.md:
an inline code span (`POST <baseUrl>/v1/chat/completions`) split across
a line break, which CommonMark renders with the line ending collapsed
to a space, so it displayed as ".../v1/chat/ completions" with a
spurious space inside the path.
Verified: origin/master and upstream/master are both already an
ancestor of this branch (identical at bd286bf5, no new commits since
this branch was cut) — nothing to merge, no conflicts.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Replaces the size-scaled expected-time estimate + matching auto-timeout
with a generic hardware/model-size disclaimer and a user-driven Cancel
button, per explicit request. Real load time depends on hardware this
feature has no way to know (VRAM, storage speed, GPU contention), so
the old estimate/timeout was a guess dressed up as a fact — worse, one
that could kill a genuinely slow load partway through on slower
hardware.
- _watchLlamaSwapLoading (session-ui.js): dropped maxWaitMs/deadline
entirely — polls indefinitely until ready or cancelled, no automatic
give-up. Message is now "Loading <model> (<size>) on <endpoint> —
this can take a while depending on your hardware and the model
size.", with the real llama.cpp log line still on its own second
line. Removed _MODEL_LOAD_TIME_MATRIX/_estimateModelLoad/
_formatRemaining (dead code once the countdown is gone) —
_lookupModelSizeGB is kept, the GB figure still shows.
- _showCenterStatus (panels-ui.js) gains opts.onCancel: renders a real
"Cancel" button (distinct from the error-type "×" close button,
since Cancel has a real consequence) that calls it on click. Caller
owns what cancelling actually means, same split as the swap-confirm
modal's promise-resolving buttons.
- Cancelling dismisses the banner, shows an info toast (not an error —
this was deliberate), and closes the session, mirroring what the old
timeout used to do automatically but now on the user's own call.
- New .center-status-cancel CSS (bordered pill button, distinct from
the plain "×" close glyph).
Test changes: removed the now-invalid timeout-auto-close/estimate
tests, added cancel-flow tests (dismiss/toast-type/session-close,
never-closes-with-no-sessionId, unbounded-polling), and real-DOM tests
for the new Cancel button (bootAppWithRealCenterStatus, evaluating
panels-ui.js instead of stubbing _showCenterStatus, since this button
is worth verifying for real rather than just through the stub every
other test in the file uses). Typecheck/lint/frontend-syntax clean;
full suite shows no new regressions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Answers the underlying request behind investigating llama.cpp log
access: surface what the backend is actually doing, live, on top of
the existing countdown timer during a model load.
- getLatestLlamaSwapLogLine()/pruneIdleLlamaSwapLogTails()
(custom-model-routes.ts): one persistent GET /api/events (SSE)
connection held open per endpoint, parsing logData frames and
keeping the latest source:"upstream" (backend llama-server) line —
filtering out llama-swap's own source:"proxy" request-access lines.
Idle-closed after 30s of no polling, same 20s sweep as the existing
swap-displacement check.
- running-status route now returns logLine alongside the existing
isLlamaSwap/running fields.
- Frontend: _watchLlamaSwapLoading's banner gains a second line
("llama.cpp: <line>", bootlog timestamp/level/component prefix
stripped for display) that stays on the last real thing llama.cpp
said rather than clearing to blank between polls.
⚠️ Caught and fixed before merge, not after: the first cut targeted
GET /logs (the endpoint the name suggests), shipped a working-looking
implementation with passing tests, and only failed a live check against
the real Nemesis llama-swap deployment — /logs turns out to carry ONLY
llama-swap's own proxy request-access log and never once showed a
single backend line, even seconds after a real, confirmed model swap
triggered via a direct API call. GET /api/events's logData frames
(with an explicit source field distinguishing upstream from proxy) are
the only source that actually has backend output; corrected and
re-verified live end-to-end through an actual forced swap before
writing this commit, confirmed live to hold its connection open
indefinitely (unlike /logs, which closes after a fixed ~100KB).
12 tests for the corrected /api/events parsing (SSE frame buffering
across chunk boundaries, source filtering, malformed/wrong-type frames,
connection reuse, idle pruning) plus 2 for the frontend banner
rendering. Typecheck/lint/frontend-syntax clean; full suite shows no
new regressions (14 more passing than baseline, matching the new
tests; same pre-existing Windows-environment failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
The llama-swap conflict check on the apply/create routes only ever runs
at THAT session's own launch/apply moment, and cannot see a swap caused
by a DIFFERENT session's later, ordinary use. Confirmed live: a second
Codex session picking a different model launched with no warning at
all — nothing conflicted at that exact instant — yet it silently
evicted the first session's model regardless (llama.cpp runs one model
at a time). Reproduced and root-caused via direct API calls against a
live test-picker instance rather than guessing.
- detectCustomModelSwapDisplacements() (custom-model-routes.ts): groups
live sessions with a customModel by endpointId, checks each group's
endpoint via GET /running once, and flags a session whose own modelId
is no longer in the running list. Read-only, best-effort per endpoint
like refreshAllCustomModelHosts's sibling sweep.
- Notifies once per displacement via a caller-owned de-dupe Set: a
session id is added when displaced, removed once its own model is
loaded/ready again, so a later genuinely-new displacement can notify
again.
- New periodic sweep in server.ts (CUSTOM_MODEL_SWAP_CHECK_INTERVAL_MS,
20s — much shorter than the 5-minute model-list refresh, since this
is time-sensitive) broadcasts a new custom-model:swapped-out SSE
event per displacement. De-dupe Set cleared per-session on session
cleanup to avoid an unbounded leak.
- Frontend: global toast (not tied to the displaced session's tab,
since the point is warning before the user types into it) naming the
session, its previous model, and what's currently loaded.
Chose the "detect after the fact" scope (vs. checking before every
message send, which would add a round-trip to every turn on every
custom-model session) per explicit user decision after being presented
the trade-off.
9 new tests for the detection logic (flag/clear/re-flag cycle,
unreachable/deleted endpoints, non-llama-swap servers, multiple
sessions on one endpoint). SSE registry bumped 158->159, parity test
passing. Typecheck/lint/frontend-syntax clean; full suite shows no new
regressions (9 more passing than baseline, matching the new tests;
same pre-existing Windows-environment failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Investigated the user's report of "Model metadata for <id> not found.
Defaulting to fallback metadata..." on every custom-endpoint codex
launch, live against the test-picker's llama-swap deployment (codex
0.152.1):
- The warning is cosmetic. `codex exec 'reply with just OK'` against the
isolated CODEX_HOME still printed the warning and still returned a
real reply.
- The isolated CODEX_HOME never gets a models_cache.json written into
it at all, even after extended real use (inspected a live, actively-
used directory) — codex can't reach OpenAI's own hosted model catalog
for this session and silently falls back every time, with no local
file to create or clean up. There is also no config.toml override for
a model's metadata.
- Fabricating a fake catalog entry to suppress it would mean copying the
SHAPE of OpenAI's own proprietary models_cache.json schema, including
real per-model system-prompt content visible in a genuine entry — not
something to build for a warning confirmed to have no effect.
- More importantly: a real tool-call attempt against the same setup came
back as agent_message TEXT (the tool-call JSON printed as the answer)
rather than an executable function_call item, confirmed via
`codex exec --json`'s raw event stream. Tool execution is what makes
codex a coding agent, so it remains not usable for real work regardless
of the metadata warning — a more precise, re-verified update to the
existing "Responses API protocol gap" finding (which reported a harder
Reconnecting/high-demand failure on a different llama-swap deployment;
this one answers /v1/responses for plain chat but still can't execute
tools).
No code changes — recipe/comment/confidence-table documentation only.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
A fresh, isolated CLAUDE_CONFIG_DIR (used to keep an injected API key
away from a stored claude.ai OAuth login) looks like a brand-new Claude
Code profile to the CLI, so it replays its ENTIRE first-run sequence on
every single launch: the theme picker, the security-notes screen, the
per-project "trust this folder?" dialog, and (running with
--dangerously-skip-permissions) a one-time bypass-permissions warning —
confirmed live, none of which a real, already-onboarded profile shows
again.
- New registry-declared env-kind field `skipFirstRunPrompts` (alongside
apiKeyTrustFile, which it reuses) — claude's entry only, carried
through buildCustomModelInjection (pure) into
applyCustomModelInjection (IO).
- seedFirstRunOnboardingState(): merges hasCompletedOnboarding: true and
this session's own projects[workingDir].hasTrustDialogAccepted: true
into the same <configDir>/.claude.json the API-key trust file already
writes to — other projects and other fields on this session's own
entry are left untouched.
- seedSkipBypassPermissionsPrompt(): merges
skipDangerousModePermissionPrompt: true into <configDir>/settings.json,
a separate file, same corrupt-tolerant merge behavior.
- applyCustomModelInjection() gains an optional workingDir parameter,
threaded from session.workingDir (dedicated apply route) /
resolvedCasePath (quick-start route) — boot recovery omits it
(a dialog already answered once needs no re-seed on the same,
persisted isolated directory).
Tests added at the pure-builder, IO-wrapper (including merge-preserves-
other-fields and corrupt-file-tolerance cases), and existing directory-
listing assertions updated for the new settings.json file. Typecheck/
lint/format clean; full suite shows no new regressions (baseline
pre-existing Windows-environment failures unchanged, 8 more passing
tests than before — the ones added here).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Neither dialog's footer had a row layout of its own to override, and
.btn-toolbar is display:flex (a block-level flex container with no
explicit inline-flex), so with no flex row context each button took its
own full-width line and the two stacked vertically. The swap-confirm
modal already had a .modal-footer rule (flex-end); the context-warning
modal had none at all. Both now share one row-layout rule, centred
rather than flex-end per feedback.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Both dialogs can appear while the centred llama-swap status banner is
still on screen (right after "Claude started — switching to
llama-swap…") — the banner's z-index is 10001, .modal's base z-index is
only 1000, so the dialog rendered fully behind it. Reported live against
the context-window-too-small modal; the swap-confirm modal has the same
structural bug for the same reason, so both get the fix.
Also: both messages ARE the modal's whole explanatory content, not a
one-line caption under a form field, so .form-hint's 0.65rem caption
size read as illegibly small — worst on the multi-sentence
context-window explanation. Bumped to 0.85rem/1.5 line-height/--text.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Claude Code's own fixed per-turn overhead (system prompt + tool schemas,
~36.4K tokens measured live) can exceed a small local model's entire real
context before any conversation history exists to compact — confirmed
live twice as an in:0 out:0 failure on the very first message sent.
CLAUDE_CODE_MAX_CONTEXT_TOKENS cannot fix this: it only governs when
history gets compacted, and there is none on message one.
- exceedsSafeContextFloor() (custom-model-routes.ts): true when a CLI's
registry entry declares contextLengthVar (currently only claude) and
the model's discovered context is below CLAUDE_MIN_SAFE_CONTEXT_TOKENS
(40000). A no-op for every other CLI by construction.
- Both apply routes (POST /api/sessions/:id/custom-model and the
quick-start customModel path) check this before the swap-conflict
check and before launching/restarting anything, returning
{requiresContextWarning, modelId, contextLength, minSafeContextTokens}
— skipped when confirmed:true.
- Frontend: #customModelContextWarningModal + _confirmContextWarning/
_resolveContextWarningConfirm (session-ui.js), wired into both
_quickStartWithCustomModelConfirm and _runCustomModelEntryViaRestart
(the path Claude actually uses) ahead of the swap-confirmation check.
Explains the fix in-modal: give the model an explicit larger -c/
--ctx-size in llama-swap instead of relying on --fit-ctx, which
optimizes for the biggest model that fits rather than the biggest
context.
Tests added for the route-level warning/confirm/skip cases and the
frontend modal + launch-flow wiring. Docs updated (custom-model-
endpoints.md, wiki/Custom-Model-Endpoints.md) and the PR's running
changeset extended.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Own review pass over the PR:
- `_flush` took the chunk out of the buffer only AFTER awaiting the write. Input
arriving during that await is enqueued (`waking` is still set, so it takes the
buffer path), and the 4 KB cap then drops the OLDEST chunk — which is the one
already on its way to the pane. The `shift()` that followed removed the NEXT
chunk instead, so the drop-oldest bookkeeping silently lost a chunk that was
never written, while the log line blamed the one that was. The chunk is now
removed before the await and re-inserted at the FRONT on a failed write, so the
order of the queue behind it is preserved. Regression test: a chunk enqueued
during the first write of a full buffer must still reach the pane (red against
the old order).
- `showCreateCaseModal()` reset the remote-host form fields but not the two new
wake inputs, so one host's MAC/command carried over into the next host that
form saved.
- The banner's pre-poll `wakeConfigured` labelled a command-only host as 'mac'.
Nothing reads the distinction, but the field is documented as which path is
configured, so it says the truth until the first poll corrects it.
- Stale `resolveRemote` comment ("only for sessions that have no usable target of
their own"): after the host config became authoritative in both directions it is
consulted on the TTL regardless.
- `host-wake-ui.js` joins the documented load order (12.2) and gets its
`@dependency`/`@loadorder` tags; the frontend module count is 33, not 32.
- `remote-wake` is not "(pure)" — the module uses `dgram`/`net`/`child_process`.
- SSE counts: 160 constants, and the category is "Remote auto-reconnect / wake
(5)"; the route table's per-file counts are refreshed (sessions 37, cases 34).
- The CLAUDE.md wake rule now names the create/attach wake, the 40 s request
budget, the whole-chunk paste drop, the registry's lifetime (drop on cleanup,
stop on shutdown) and the deliberately non-wake-aware WebSocket keystroke
path — that paragraph is what the next person reads.
- Reverted the eight lines of unrelated Prettier markdown churn in
`docs/architecture-invariants.md` (docs/ is not in the format glob, so it was
an editor): only the new wake paragraph remains in the diff.
Review follow-up on the wake-on-LAN PR (five findings, all of them about the
state the feature keeps and the budgets it inherits):
- Wake state is dropped by `WebServer.cleanupSession` instead of the two delete
routes, so it now goes with the session on EVERY cleanup path (cron, admin,
scheduled-run teardown, error paths) instead of surviving with up to 4 KB of
the user's buffered keystrokes. `registerSessionRoutes` returns the registry
so the server can own its lifetime without the wake-capable code living in
`server.ts`; the wiring guard is updated to allow that and gains a second
assertion that `server.ts` calls nothing but `drop`/`stop` on it.
- `_effectiveRemote` returns before `_state`, so a LOCAL session no longer gets
a wake-state entry — the input gate runs on every keystroke, so that entry
used to be allocated for every session the user types in.
- An input chunk larger than the 4 KB cap is dropped OUTRIGHT instead of being
head-trimmed and then written as a fragment: one paste is one `input` value
and was never typed character by character, so its tail is a partial command
the user never sent. The drop is logged.
- The manual wake button passes `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS` (40 s)
like the create/attach paths, instead of inheriting the 90 s session default
that the dashboard's reverse proxy cuts off at 60 s.
- `RemoteWakeRegistry.stop()` aborts in-flight readiness polls (abortable
sleep) and refuses new wakes, and `WebServer.stop()` calls it, so a restart
during a wake no longer waits the poll out.
- The banner/toast wording keys off a new `queuedInput` flag on the two SSE
events, which is true only when the server actually holds bytes: browser
keystrokes travel over the WebSocket, which never passes through the
registry, so the wake BUTTON must not promise queued input. The failed-wake
path also stops pattern-matching the error message (it re-asks the
reachability route) and the WoL dialog says "admin-only" instead of "host not
found" for a non-admin in multi-user mode.
Fourth review of the reboot-restore branch, and the third to find a defect
in the previous round's fix. This one is the same shape as its predecessor:
a counter keyed on one thing, compared against a set keyed on another.
The generation counter was indexed by the entry's owner, while the in-flight
set holds the caller doing the restoring. Those are the same person exactly
when a user restores their own sessions, which is every case the tests
covered. The route deliberately supports the other case: an admin may spend
another user's entries. So when an admin restored Bob's sessions and Bob
dismissed the banner, nothing matched, the entries came back, and a plan Bob
had explicitly dismissed was re-armed for another twenty-four hours.
Rather than reconcile the two key spaces, the counter is gone. `take()` now
parks the entries it hands out, remembering which caller is spending them,
and they stay parked until that restore ends. A dismiss filters the parked
entries by `canAccess(entry.owner)` — the same predicate it already applies
to the plan — so it reaches them wherever they are. `releaseFlight()` puts
back only what is still parked. Expiry and a fresh boot plan unpark
everything, for the same reason. There is one key space now, the entry's
owner, and the spender is only ever used to tell two concurrent flights
apart. That removes `generations`, `snapshotGenerations()`, `bump()`,
`bumpAll()` and the argument threaded through the route.
The discard grew the teardown it still lacked. A rebuild can fail after
startInteractive() resolved, and a restored workspace still carries
Codeman's hooks, so the CLI can post a hook event within milliseconds; the
transcript watcher that starts from it, the attachment registry, the wait
registry and the approvals inbox all outlive the listeners and would meet
the retry, which reuses the session id by design. Its steps also run in
reverse order now, so no live listener can reach a tracker that has already
stopped, and the mux kill has its own guard, because stop() kills the pane
in its last block after destroying four trackers.
Tests. The run-summary test named an interval and asserted a map entry, so
dropping stop() left it green; it now spies on stop(). Nothing pinned that
before-spawn must precede setupSessionListeners, which reads the flag that
phase restores, so swapping the two lines was silent; the ordering test now
includes the listener setup. The retry assertion was a tautology and now
asserts a different refs object. Both strengthened tests were verified by
reverting their fix. Two new tests cover the admin-restores-another-owner
cases this round was about. The server in the discard test is built once and
stopped, since its constructor registers handlers on module-level watchers,
and the workspace is removed through safeRmHomeTree.
Refs #411
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Third review of the reboot-restore branch. The narrow discard the previous
commit introduced avoided everything cleanupSession() did wrongly, and in
dropping so much of it also dropped four things it had to keep.
The worst broke the retry the whole design rests on. setupSessionListeners()
returns early while sessionListenerRefs still holds the session id, and the
discard never cleared that entry. So the advertised flow — a rebuild fails
because the agent binary is missing, the user fixes their PATH and clicks
again — reused the same id, wired no listeners at all, and produced a tab
that never showed output, never updated its status and never persisted. That
is worse than the leak the discard was added to prevent. Three more
registrations leaked with it: a RunSummaryTracker and its interval, an image
watcher on the workspace, and the Ralph fix-plan watcher. The discard now
undoes each registration setupSessionListeners() makes, in its order, and
the per-session custom-model config directory, which holds the endpoint's
API key literally and which nothing else would ever remove.
The image-watcher flag was restored after the code that reads it, so a
session came back reporting the feature as on with nothing watching. It
moves to the before-spawn phase, and that phase now runs before the
listeners rather than after them.
The generation counter that lets a mid-restore dismiss win was global while
clear() is ownership-scoped, so one user's dismiss discarded another user's
unspent entries, permanently, because nothing rebuilds an in-memory plan. It
is now per owner. Bumping only the owners of entries the dismiss removed was
not enough either: take() has already emptied the plan by then, so a dismiss
landing mid-restore saw nothing of that owner's to remove and invalidated
nothing. The owners that matter are those with a restore in flight, filtered
by what the dismissing user may access, and that is what clear() now bumps.
Plan expiry bumps too, so a restore straddling the 24-hour boundary cannot
hand entries back and give an expired plan another full day.
Tests. discardPartiallyBuiltSession had no test at all: the only
implementation any test ran was the mock's one-line stub, which is why every
defect above was invisible. test/discard-partially-built-session.ts drives
the real WebServer, and the retry assertion fails if the listener refs are
left behind — verified by reverting the fix. The dismiss-race test drove the
registry by hand, so deleting the route's generation argument left it green;
it now goes through the route, and two further tests cover the multi-user
cases.
The mock context has now gone stale twice, because route tests pass it as
`ctx as never` and tsconfig.json includes only src, so nothing ever compares
it to the ports. A type-level guard is therefore inert — I wrote one and
confirmed it never fires. test/mocks/mock-route-context-completeness.ts
compares the mock's keys against WebServer.createRouteContext() at runtime
instead, and names what is missing.
Also: the API reference now says workspace-forbidden is judged against the
owner's grant, the banner's module header no longer claims Restore always
dismisses it, and the detail span gets the same min-width: 0 the phone rule
already needed.
Refs #411
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A second review of the previous commit found that its own repair for the
session leak introduced three defects, all from reaching for
cleanupSession() to undo a half-built session. That function is the
user-initiated delete, not an undo.
It banked the session's historical token and cost totals into the lifetime
figures, and a reboot never runs cleanup, so those totals had never been
counted before; every failed rebuild added them again. It saw the pin that
had just been restored and demoted the record to `stopped`, which this pass
reads as the durable marker of a deliberate kill, so a pinned session whose
rebuild failed became permanently unrestorable. And it recursively removed
`.claude-images` from the working directory, which belongs to the workspace
rather than to the session, so a failed rebuild destroyed the pasted images
of any other live session in that repo.
discardPartiallyBuiltSession() now undoes only what the construction did:
the map entry, the tab-layout slot, the listeners and any pane the launch
created before throwing. The persisted record, the lifetime totals, the
Ralph state and the workspace's files are left alone.
Re-applying the persisted state also splits in two, which removes the first
two defects at the root rather than only at the call site. The half that
shapes the pane, the custom-model environment and the nice priority, still
runs before the spawn. The half that is the session's own history now runs
after it, so a session whose pane never started carries no totals and no pin
for anything downstream to misread.
The rest of that review. The multi-user workspace confinement re-check read
the requesting user's grant, and returns true for an admin, so the case its
own comment described was the one it missed; it now resolves the entry
owner's grant through isWorkingDirAllowedForUsername, the way cron does. A
forbidden workspace goes back on offer, matching both the registry's stated
contract and the API reference. The client re-reads the plan after a restore
instead of blanking the banner, so entries the server put back stay
reachable, and a 409 now says a restore is already running rather than
reporting a failure. A dismiss arriving mid-restore wins, through a
generation counter the route carries across its take. The re-application
also restores the tab colour, the image-watcher flag and the original
pinnedAt, via a new Session.restorePin that does not re-stamp the pin time.
The phone breakpoint gains min-width: 0, without which a nowrap flex item
never shrinks and the buttons still overflow, and it folds into the existing
phone block.
Ralph's loop configuration still does not survive a restore, because
toState() reads it off a live tracker and there is no way to keep it without
arming the loop. The method now says so rather than leaving it implied.
Tests. The capacity test could not fail on the property it existed for: it
filled the board past the cap before the loop, so a single pre-loop check
would have passed it. It now leaves one seat, so only a per-iteration check
restores exactly one entry. New tests cover the ordering around the spawn,
a throw before the loop returning the whole plan and releasing the flight,
the dismiss-during-restore race, and that the failure path calls the narrow
discard rather than the delete. The shared mock context gains the port
method it was missing, which is what made the first run of these tests fail
for the wrong reason.
Refs #411
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Root cause of the context-overflow regression reported live: "API Error: 400
request (36437 tokens) exceeds the available context size (16384 tokens)".
Discovery had stored modelContextLengths.qwen3.8-27b-ud-q4_k_xl = 154112,
so CLAUDE_CODE_MAX_CONTEXT_TOKENS told Claude Code it had a huge window and
it never compacted - but the real llama-swap server was launched with
--fit-ctx 16384 (confirmed against /running's own cmd field) and refused
the request right at that real limit.
/props?model=<id>'s n_ctx (the field discovery read) is confirmed live to
be unreliable for a --fit-ctx-launched backend: it reported 154112 for the
same model /running says was launched with --fit-ctx 16384 - appears to
report the model's theoretical/trained maximum context, not the runtime-
configured one.
discoverModels() now parses the REAL configured size straight out of
llama-swap's own launch command instead (parseCtxFromCmd(), reading
/running's cmd field - --fit-ctx first, then the plain llama.cpp -c/
--ctx-size a hand-written command might use), and only falls back to the
old /props probe when cmd states no recognizable flag at all. One /running
call now covers every loaded model's context length in a single request,
same as it already did for the swap-conflict check and the load trigger.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
The loading banner now shows a live countdown against its own timeout
(updated every poll, so every second by default) instead of a static
"this can take a while" — e.g. "Loading qwen3.8-27b (16.4 GB, typically
~1-3 min) on llama-swap - 47s remaining".
If the countdown reaches zero and the model still isn't ready, this is now
treated as a real failure rather than a "keep waiting" shrug:
- The banner turns into a sticky error (_showCenterStatus gains a `type`
option - 'error' drops the spinner and adds a close button, since nothing
is "in progress" anymore and a sticky message needs a way to dismiss it),
naming the llama-swap server's own logs as where to look for detail.
- The session that load was for is closed automatically (closeSession) -
requested explicitly: a console left open and pointed at a model that
never finished loading is worse than no console at all. Both apply paths
now thread the new session's id through to _watchLlamaSwapLoading for
this (new required 3rd parameter, after endpointId/modelId).
_watchLlamaSwapGeneration's existing stale-call guard extends naturally to
this: a superseded call's own eventual timeout recognises it no longer owns
the banner and neither shows the error nor closes a session that may by
then belong to a different, newer launch.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Discovery now also parses a GB figure out of an auto-discovered model's own
description (llama-swap writes "Auto-discovered 16.35 GB - parameters
auto-fitted by llama.cpp"), stored per model as modelSizesGB - unlike
context length this needs no /props probe (the figure is right there in
/v1/models) so it is populated for every model regardless of loaded state.
A hand-configured profile's own description has no such figure and
correctly gets no entry.
The loading banner (_watchLlamaSwapLoading) now looks this up and, when
known, shows it plus a rough estimate from a small size->time matrix
(_estimateModelLoad/_MODEL_LOAD_TIME_MATRIX, session-ui.js) -
"Loading qwen3.8-27b-ud-q4_k_xl (16.4 GB, typically ~1-3 min) on
llama-swap... this can take a while" - and uses that same estimate's own
bracket to scale the banner's default give-up timeout for a very large
model, instead of a flat 5 minutes for everything. Explicitly labelled as
an UNMEASURED, typical-hardware estimate in every relevant comment - this
is not benchmarked against any real endpoint's actual storage/GPU, just a
reasonable expectation-setter. A model with no discoverable size (a
hand-configured profile) gets no size/estimate shown at all, matching the
"never a guess" convention modelContextLengths already established.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Reported: the "Loading..." banner stayed up past 2 minutes even though
llama-swap itself had already finished loading the model. Three fixes:
1. pollIntervalMs default 3000ms -> 1000ms (as asked).
2. The loop now checks readiness IMMEDIATELY on entry rather than sleeping
a full interval first - a model that's already ready (a fast load, or a
re-apply onto one already loaded) shouldn't sit on "Loading..." at all.
3. maxWaitMs default 120000ms (2 min) -> 300000ms (5 min): a large (20GB+)
model reading from disk can genuinely take longer than 2 minutes, which
would have looked identical to the reported symptom - "still stuck past
the point it should have resolved" - except it would have actually
flipped to a "still waiting" warning toast at the 2-minute mark rather
than staying on "Loading" indefinitely, so this alone doesn't explain
what was reported, but is a real, separate improvement worth making.
Also fixes a real, separate bug this surfaced while reasoning through the
report: _showCenterStatus's banner is ONE shared, reused DOM node. A second
call to _watchLlamaSwapLoading (e.g. switching models again before the
first switch's loop had finished) would take over that shared banner, but
the FIRST loop was still running and would eventually dismiss or overwrite
it once ITS OWN deadline or readiness check resolved - clobbering whatever
the second, current loop had put there. A generation counter
(_watchLlamaSwapGeneration) now lets each call recognise when it no longer
owns the banner and stop touching it silently, rather than only the last
call to actually start ever safely reading or writing it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Fifteen findings from two independent reviews of #442, three of them
blocking. Every one is addressed here.
The three blockers all sat in the restore route. A rebuild that threw after
addSession left a registered session with no pane behind it, visible on the
board, holding a layout slot and written to state.json, with its plan entry
already spent; the catch now cleans the session up and puts the entry back.
The loop checked neither the global nor the per-user session cap, so one
click could take a board past a documented limit; capacity is now re-checked
per iteration, because the loop is itself creating the sessions it counts.
Worst of the three, a rebuilt session carried none of the state its
constructor has no parameter for and then persisted itself over the record
that held it, zeroing token and cost totals and dropping the pin. The pin
matters most: pruning keeps a record only while it is pinned, so discarding
it handed the record to the next stale sweep. A new
reapplyPersistedSessionState() on the session port restores the pin, the
token totals, auto-compact, auto-clear, auto-resume, nice priority, the
flicker filter and the custom-model selection, and it runs before both
startInteractive and the first persist.
The rest, in the order they bite a user. Every rebuild failure was reported
as workspace-missing, so the banner told users their repo was gone when the
agent had simply failed to start; there are now distinct reasons, and the
toast names each one. The client read restored and skipped off the outer
response object rather than through the uniform envelope, so every count
came back zero and neither toast ever fired. A board left open across the
reboot never learned an offer existed, because the banner was seeded only on
the page-load path; it now re-reads on every SSE init. The workspace check
was existence-only, skipping the multi-user confinement that the create
route applies, so a withdrawn grant would not be noticed. The banner had no
phone breakpoint while its text was nowrap and its buttons could not shrink.
Smaller: a missing workspace is now re-offered rather than dropped, while an
already-open conversation is dropped rather than re-offered forever; a throw
anywhere in the route returns the unspent entries instead of discarding the
plan; the single flight is keyed by owner, since take() already stops two
callers receiving one entry; the env clamp's header no longer claims a
protection it cannot provide on this path today, and names the check that
does bite; the three endpoints are documented in docs/api-reference.md; and
the module header now says that os.uptime() reads the host's clock, so the
feature is effectively off inside a container.
The review also explained why the tests missed all of this: they proved the
construction claim through their own copy of the construction rather than
through the route, and the route tests used workspaces that did not exist,
so no Session was ever built. test/routes/reboot-restore-rebuild-failure.ts
mocks the Session module to drive the route's real path, and covers the
cleanup, the reason reported, the re-application ordering, the broadcast and
the caps. The mock route context gains the port method and the mux call the
route needs.
Refs #411
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
refreshHostWakeBanner clears _hostWake before calling _hostWakeTick, so the
clear branch's `if (this._hostWake)` guard skipped the repaint: once the
banner had appeared for an unreachable remote session it stayed up on every
chat (local ones included) until a reload, and the 30s ticker never cleared
it either. Render unconditionally in that branch — _renderHostWakeBanner is
idempotent with a null state.
Reproduced in a real browser (Puppeteer, mobile viewport): state went null
but banner.hidden stayed false. Regression test added in
test/host-wake-banner.test.ts (red before, green after).
Root cause of "it doesn't look like llama-swap is actually switching the
model" (confirmed live: no load_model line in llama-swap's own logs after
applying a selection). llama-swap has no "switch model" admin endpoint - the
ONLY thing that starts a swap is a real inference request naming the model.
Every previous fix (the conflict check, the loading banner) assumed a swap
would start on its own; nothing ever actually asked llama-swap to load
anything until the launched CLI's first real prompt did, which could be
much later than "applying the selection" implied.
Adds triggerLlamaSwapLoad() (custom-model-routes.ts): sends the smallest
real request that will start a load - POST <baseUrl>/v1/chat/completions,
max_tokens: 1, one throwaway message - fire-and-forget (never awaited by
the caller; the frontend's own running-status polling is what actually
confirms readiness). Wired into both apply paths (the dedicated restart
route and the one-shot quick-start route), fired whenever the target model
isn't already the one loaded and ready - a broader condition than the
existing swapNeeded (which only gates the "this will evict another
session's model" confirmation ask and deliberately stays narrow to that).
modelSwapInProgress in both routes' responses now reflects this same
broader condition too, so the frontend's loading banner actually correlates
with a real in-flight load rather than only firing when something else
happened to be loaded already.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
The "Claude started - switching to <endpoint>..." and "Loading <model> on
<endpoint>... this can take a while" messages lived in the top-right toast
corner along with everything else, easy to miss given they can each sit on
screen for well over a minute (a real llama-swap model load).
Adds _showCenterStatus() (panels-ui.js): a single, reused, screen-centred
banner with a spinner, non-blocking (no backdrop, pointer-events: none on
the wrapper) so it never gets in the way of using the app while it's up.
Both call sites (_runCustomModelEntryViaRestart's switching message,
_watchLlamaSwapLoading's loading message) now use it instead of showToast.
Every OTHER status in these two flows - the llama-swap conflict warning
already moved to its own modal, apply failures, cancellation, and
_watchLlamaSwapLoading's own final "ready"/"still waiting" outcome - stays
exactly where it was, in the corner.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
The llama-swap "this will unload it for session X" warning used a native
browser confirm() popup, which looks out of place next to the rest of the
app's own modals.
Adds #customModelSwapConfirmModal (index.html) with Cancel/Switch-anyway
buttons, styled to match the app. _confirmModelSwap(message) shows it and
returns a promise that resolves true/false the same way confirm() would;
_resolveModelSwapConfirm(proceed) (wired to both buttons and the backdrop
click) settles it. Both llama-swap conflict call sites
(_quickStartWithCustomModelConfirm for the one-shot launch path,
_runCustomModelEntryViaRestart for Claude's restart path) now await this
instead of calling confirm() directly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Claude stays on the launch-then-restart path (see runCustomModelEntry's own
comment for why), but with nothing on screen during that window, a native
boot that briefly talks to the cloud model read as "the endpoint didn't
apply" rather than "the switch hasn't happened yet".
A sticky "Claude started - switching to <endpoint>..." toast now covers the
whole window from the native launch through the apply call, updated in
place (never stacked) as the outcome resolves: dismissed on cancel or
failure (replaced by the existing cancellation/error toast), handed off to
_watchLlamaSwapLoading's own sticky toast when a model swap is in progress,
or updated to the existing "Pointed at ... - restarting" message and
auto-dismissed after 3s on a plain success.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Fixes the visible double-launch reported on Codex: picking a custom-model
Run-menu entry launched natively first, waited for it to settle, then
restarted it in place with the endpoint applied. Necessary for the design at
the time, but visibly a native boot immediately followed by a second one -
worst on a CLI whose TUI fully reinitializes on a restart, confirmed live on
Codex.
POST /api/quick-start gains an optional customModel field
({endpointId, modelId, confirmed?}). When present, the route mints the
session's id itself (crypto.randomUUID()) before constructing it, computes
the same injection the existing POST /api/sessions/:id/custom-model route
computes (including the llama-swap conflict check from the last commit -
same {requiresConfirmation, currentlyLoadedModel, affectedSessions} shape,
no session created until confirmed), and launches the session already
pointed at the endpoint: env vars via the constructor, and the launchModel
override merged onto piConfig/grokConfig/ompConfig using the registry's own
launch.legacyConfigField the same way session.ts's restart path already
does. No restart at all - setCustomModel() afterward is bookkeeping only.
Wired into 7 of 8 launch functions (session-ui.js): openCode, codex, gemini,
pi, grok, deepseek, omp. Claude stays on the original launch-then-restart
path for now: its own --resume-based restart is far less jarring than the
other seven's, and runClaude()'s multi-tab launch plus docker-config-drift
confirm/retry loop make folding it into the one-shot path separate,
higher-risk work than the other seven's each-a-single-simple-launch shape.
Also fixes a pre-existing 'mode === omp' branch flagged by the CLI-id
static guard (test/cli-registry-no-id-branching.test.ts) - the ompConfig
launchModel merge is the same 'legacy <Mode>Config plumbing' category as
the six sibling branches already allowlisted there, just newly literal
where it was previously only inside resolveOmpConfigForCreate's own check.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Covers #430's full scope so far: the picker itself, the model-selection
dialog, periodic re-discovery, and the session-busy/toast/CLAUDE_CONFIG_DIR/
context-length/llama-swap-conflict fixes found through live validation
against a real llama-swap server.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Root-caused the user's earlier confusion ('the terminal says opus even though
something is waiting for llama to load'): llama.cpp runs exactly one model at
a time, and llama-swap unloads/reloads it on demand - a swap can take
anywhere from a few seconds to well over a minute, during which a session
looks indistinguishable from one still on the native backend.
1. Feature-detects llama-swap (vs. plain llama.cpp/any OpenAI-compatible
server) via its own GET /running, which plain llama.cpp has no concept of
at all. New GET /api/model-endpoints/:id/running-status route exposes this
read-only, for the frontend's polling loop below.
2. Before applying a selection, POST /api/sessions/:id/custom-model now checks
what llama-swap currently has loaded. If it differs from the requested
model AND another live session's own customModel selection is actively
using that loaded model, the apply is refused with a
{requiresConfirmation, currentlyLoadedModel, affectedSessions} payload
instead of silently switching. A "confirmed: true" field on the retry
skips the check. Switching with nothing else affected proceeds
immediately, no confirmation asked, only ever when there is something to
warn about.
3. The frontend (runCustomModelEntry) shows a native confirm() naming the
affected session(s) and the model they'd lose, matching this codebase's
existing convention for this class of decision (delete case, kill
session, etc.) rather than a new modal. On a successful apply the response
also carries modelSwapInProgress; when true, a new _watchLlamaSwapLoading
poll shows a sticky "Loading <model>..." toast via the new running-status
route until llama-swap reports the target model ready (bounded at 2
minutes), so a prompt sent mid-swap reads as "loading", never as silence
or an answer from whatever was loaded a moment before.
Checks are read-only against llama-swap's own /running - never /props, which
takes a ?model= and can itself trigger a load as a side effect of asking.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
A host reboot takes the tmux server down with it, so every pane dies,
reconciliation finds nothing to attach to, and the board comes up empty.
Picking yesterday's work back up meant finding each conversation in history
and resuming it by hand, one at a time.
The boot pass now works out what the reboot killed and leaves it on offer.
It runs inside restoreMuxSessions(), in the window where reconciliation has
reported the dead sessions and cleanupStaleSessions() has not pruned their
records yet, which is the only place the records can still be read. The
board shows a banner, and nothing is created until the user clicks it.
A click rather than an automatic restore is what makes the reboot heuristic
acceptable. The heuristic cannot tell a reboot from a crash that took tmux
down inside the same window, so it decides whether to ASK, never whether to
act: a wrong yes costs a line of text the user dismisses instead of N CLI
processes nobody asked for.
Four things are re-checked when the click arrives rather than trusted from
boot, because hours can pass and the board moves on. The owner's privilege
grant re-resolves through the env clamp. The workspace must still be on
disk. A conversation the user already resumed by hand from the Resume list
is skipped, since two panes running --resume on one conversation would
fight over the same transcript. Entries leave the plan synchronously before
the first await, and the route is single-flighted, so a double-click or two
devices cannot both reach the same entry.
A restored session comes back attached, idle and disarmed. Respawn
controllers and Ralph loops are deliberately not re-armed: a machine that
just came up is the worst moment to turn an autonomous run loose. Its
workspace hooks are installed by the restore route itself, because the
boot-time sweep sits behind a gate that is false after a reboot and has
finished long before the click; without them a session goes silently blind,
with no stop or idle events for respawn, no Approvals Inbox item and no red
tab on a blocking dialog. Stats collection starts the same way.
The pane is new, so the conversation continues and the terminal scrollback
does not. The banner says so rather than letting an empty pane read as a
broken restore.
The plan lives in memory only. A server restart drops it, which costs the
convenience this adds and never the conversation: the conversation is the
transcript under ~/.claude/projects, which the Welcome screen's Resume list
and the Session Manager already read, so a dropped plan returns the user to
resuming by hand.
clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the
question it answers is about session privilege rather than about HTTP and
it now has a caller outside the route layer. Its test hook stays re-exported
from session-routes.ts.
Claude sessions only for this pass. The other CLIs name their thread in
their own config object, which this does not thread through yet. Remote and
docker sessions are skipped on purpose, because both need another host or a
container to be up and a freshly booted machine cannot promise either.
Refs #411
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fixes the CI failure on the last two commits: this route test asserted an
exact envKeys list for a claude-mode apply that predates the
CLAUDE_CONFIG_DIR isolation fix, so it failed on the new CLAUDE_CONFIG_DIR
entry it correctly started appending. Updates the expected list and adds
assertions for the isolated config dir path and the pre-seeded
.claude.json trust-approval file, matching the behavior added in the two
prior commits rather than just tolerating it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
The CLAUDE_CONFIG_DIR isolation from the previous commit fixed the cosmetic
auth warning but introduced a real regression: an otherwise-empty config
directory has none of a real profile's prior custom-API-key approvals, so
Claude Code stops at an interactive 'Detected a custom API key - use it?'
prompt on every single launch. Confirmed live. With nobody at a TTY to
answer, the prompt's own default ('No') silently refuses the very key this
feature just injected, which looks like the endpoint being ignored.
Adds apiKeyTrustFile to the env-kind customModelInjection capability shape
({relPath, shape: 'claude-api-key-responses'}), set on claude's entry to
{relPath: '.claude.json', shape: 'claude-api-key-responses'}. The apply step
merges customApiKeyResponses.approved: [apiKey] into
<isolatedConfigDir>/.claude.json - the exact field a real answered prompt
itself writes to (confirmed against a real ~/.claude.json after answering by
hand once), so this answers the prompt in advance rather than bypassing it.
Merges onto whatever the CLI already wrote into that file on an earlier
launch in the same isolated directory rather than overwriting it; a missing
or corrupt file is treated as empty rather than failing the apply.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Addresses two live-validation findings on the Run-menu custom-model picker:
1. Both claude.ai and ANTHROPIC_API_KEY set warning. Claude Code still
coexists an OAuth login with an injected ANTHROPIC_API_KEY in the same
config directory and warns about it (confirmed cosmetic - the API key
wins for actual requests, verified via a real session's own API Usage
Billing line). A custom-model claude session now gets an isolated
CLAUDE_CONFIG_DIR (registry-declared via a new configDirVar field, empty,
no files written into it) so there is nothing to conflict with. projects
is symlinked (junction on Windows) back into the real config dir so the
response viewer, subagent windows and Read My Mind keep working for that
session, best-effort.
2. Context-window overflow. Claude Code assumes a large default context
window for a model id it doesn't recognise and never compacts, so a
custom endpoint's real, much smaller context (verified live: a 400
exceeding a 16384-token llama-swap model with a stock ~33.7K-token system
prompt) silently overflows. Discovery now also learns each model's real
context length from llama.cpp/llama-swap's GET /props?model=<id> (n_ctx),
but ONLY for a model llama-swap's own /v1/models response already marks
status.value === 'loaded' - never an unloaded one, since llama-swap
treats ?model= as a routing hint and probing an unloaded model risks
triggering an actual, slow, GPU-swapping load as a side effect of
read-only discovery. A server with no status field at all gets no
enrichment rather than a guess; a model not probed this round keeps its
previously-learned value until it disappears from the list entirely.
Stored per model (CustomModelHost.modelContextLengths) and applied via a
new contextLengthVar registry field, set to
CLAUDE_CODE_MAX_CONTEXT_TOKENS for claude.
Both new fields live on the existing env-kind customModelInjection
capability shape, declared only on claude's registry entry - every other
CLI's injection is unaffected (pinned by test).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Root cause of every 'Session is busy' apply failure reported from live
testing: a just-launched CLI reports itself 'busy' for its own startup
(boot spinner, workspace-trust check) well before runCustomModelEntry's
apply call could reach it, and the apply route's isBusy() guard correctly
cannot tell that apart from a real turn in progress — it exists precisely
to refuse restarting a session mid-turn, and a fresh boot looks exactly
like one from the outside. Confirmed live: replaying the identical apply
call by hand against the same session, once it had settled, succeeded
immediately.
Fixed by waiting on the session's own readiness signal before applying:
GET /api/sessions/:id/wait?until=idle&timeout=20000, one GET already built
for exactly this ('Agent wait primitives', CLAUDE.md) rather than inventing
a client-side poll loop. A timeout there is a normal 200 per that
endpoint's own contract, never an error, so a session still busy after 20s
just reaches the apply call anyway and gets the route's own honest error —
now visible, since the previous commit made error toasts sticky and
stopped discarding the real error text.
Tests: new case in custom-model-run-menu-ui.test.ts pins the ordering (the
wait call happens, and strictly before the apply call) and its exact query
string.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Two related fixes, both needed to actually diagnose 'Session started on
the native backend — could not apply the custom endpoint' reports from
live testing:
1. runCustomModelEntry()'s apply call went through _apiJson(), which
unwraps a success body but SWALLOWS a failure response entirely and
returns null — discarding the one thing (error, errorCode) that would
tell 'endpoint unreachable' apart from 'not a discovered model',
'remote/Docker session', or a dozen other real causes the apply route
already reports distinctly. Switched to _api() so the actual response
body is read on failure too, and the toast now includes the real
message.
2. showToast() defaulted every toast, error or not, to a 3s auto-dismiss
with no way to read it again — exactly what made the above generic
message impossible to act on even before the fix above. Error toasts
now default to sticky (duration: 0, no auto-dismiss) unless a caller
opts into a duration, and every toast — sticky or not — gets an
explicit close (x) button, since a sticky toast with no way to
dismiss it would just accumulate across repeated failures.
Tests: custom-model-run-menu-ui.test.ts's two apply tests updated for the
_api() switch (their mocks previously stubbed _apiJson, which the apply
call no longer goes through), plus a new test pinning that the real
server error string reaches the toast on a failure.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Every message typed on an Android phone lost its last character.
An Android soft keyboard commits the last typed character and sends the
Enter key in ONE InputConnection transaction, so the committed-text
`input` event and the Enter keydown are both processed before any
zero-delay timer runs. The orphaned-input recovery from #388 resolved
its candidate only on such a timer, and that lost the character twice
over:
* ORDER — xterm emits `\r` synchronously from the Enter keydown, and
the local-echo composer submits `pendingText` right there. The
recovered character arrived one macrotask too late to be part of the
prompt.
* LOSS — that same `\r` bumps the canonical counter, so by the time
the candidate resolved, `canonicalCount > snapshot` read as "xterm
spoke for this keystroke" and stood the recovery down. The character
was not merely late, it was dropped.
Drain pending candidates synchronously at the next keydown instead, from
xterm's custom key handler, which runs before xterm processes that key.
The counter then still holds the value it had while the candidate's own
keystroke was current, so the stand-down decision is made against the
right keystroke, and the recovered byte reaches the composer ahead of
whatever the new key emits. The timer stays as the fallback for a
keystroke with no key after it.
Physical keyboards are unaffected: there the timer has already resolved
the candidate long before the next key arrives.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The dialog had no max-height at all, so an endpoint with many discovered
models grew it past the viewport with nothing to scroll — reported live as
both "takes up the full page" and "the list is truncated", which turn out
to be the same bug. Gives #customModelPickModal .modal-content the same
bounded-height + scrollable-body shape cronModal's .modal-lg already uses
(max-height + flex column on the content, overflow-y:auto + flex:1 on the
body), scoped by id rather than folded into the shared .modal-sm class
three other modals already use for short, fixed content.
max-height: min(70vh, 520px) scales with the viewport (a phone gets 70% of
its height; a 4K display never gets a needlessly tall dialog) rather than
committing to one fixed pixel value that would be wrong at either end.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Two enhancements requested after live-validating PR #430 against a real
llama.cpp server:
1. Model picker dialog. Picking a Run-menu Custom Endpoints entry used to
apply the endpoint's defaultModelId (or the first discovered model)
silently. Now, via the new selectCustomModelEntry() (session-ui.js):
- exactly one discovered model launches straight away, same as before
- two or more open a new #customModelPickModal listing every discovered
model; defaultModelId (if set) is marked but never auto-chosen, since
the point of asking is letting ONE launch deliberately differ from
the saved default, not just confirming it
The endpoint is re-fetched at click time rather than trusting anything
cached from the dropdown's own render, since the model list can have
changed (the sweep below, or a settings-panel edit) since it opened.
runCustomModelEntry() itself — the actual launch, routed through run()
for the in-flight lock, snapshot-guarded against applying to the wrong
session — is unchanged; it now just always receives an explicit model
id from one of these two paths instead of computing one itself.
2. Periodic re-discovery. Every saved endpoint's models now refresh
automatically every 5 minutes in the background
(CUSTOM_MODEL_REDISCOVER_INTERVAL_MS, server.ts, registered the same way
as the Codex plan-usage poll it sits beside — this.cleanup.setInterval,
off under testMode), so a model the server starts or stops serving shows
up without another manual "Discover" click. The manual POST
.../discover-models route and the new refreshAllCustomModelHosts()
sweep (custom-model-routes.ts) now share one pure merge step
(applyDiscoveredModels: stamps lastDiscoveredAt, drops a defaultModelId
that no longer appears) rather than two copies that could drift. The
sweep is best-effort per host — one endpoint being unreachable on a
cycle never blocks the others — and re-reads the store before each
host's write, keyed by id, so a concurrent edit or delete from the
settings panel always wins over a sweep that started before it.
Tests: test/custom-model-endpoint-rediscovery.test.ts is a new, dedicated
file for the sweep (kept separate from custom-model-routes.test.ts because
that file's data dir is shared across every test in it — one temp HOME per
FILE, not per test — which would make a sweep-touches-every-host assertion
meaningless there). test/custom-model-run-menu-ui.test.ts gained a new
describe block driving the real picker modal through JSDOM: single-model
bypass, multi-model dialog with the default marked-not-chosen, picking a
row closes the modal and launches with that exact model, the endpoint
re-fetch, and the two "vanished by click time" toast paths.
Docs: docs/custom-model-endpoints.md, docs/wiki/Custom-Model-Endpoints.md,
docs/api-reference.md and CLAUDE.md's dense feature paragraph all updated
— the last of these also caught up two sentences that had gone stale after
the draft-review fixes landed (the picker routes through run() now, not a
raw run*() call).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Addresses every blocker, both majors, and all but one minor from the
maintainer's review of the draft PR.
Blockers:
1. Every generated inline onclick was unparseable. JSON.stringify's own
double quotes terminated the double-quoted HTML attribute at the first
one, leaving btn.onclick null on every picker entry and every Discover/
Edit/Delete button. Fixed with escapeHtml(JSON.stringify(...)) per
argument, the same idiom deleteCase's onclick already uses four lines
away in session-ui.js. This also closes the live-HTML-injection route
through modelId (server-controlled, from the endpoint's own /v1/models
reply): with quoting intact, a `>` inside it can no longer terminate the
<button> tag early.
2. GET /api/model-endpoints wraps its body in the {success,data} envelope
like every other /api route (server.ts's preSerialization hook applies
to arrays too), so Array.isArray(hosts) was always false in production
and the picker/settings panel silently saw nothing. Both call sites now
go through _apiJson(), which already exists for exactly this.
3. A failed or declined run*() (missing CLI, isBusy, a caught exception)
returns normally without ever changing activeSessionId, so the apply
step used to silently re-point and restart whatever session the user was
already looking at. runCustomModelEntry() now snapshots activeSessionId
before the launch and requires it to have actually changed.
Majors:
4. Routes the launch through run() itself via a temporary _runMode swap
(never persisted — setRunMode() would sync it to the server) instead of
a parallel hardcoded dispatch table, so a custom-model launch now holds
the same _runInFlight lock every other Run click gets. This also
resolves the "hardcoded runners map contradicts the PR's own design"
minor: dispatch is run()'s own, so a CLI whose customModelInjection
recipe lands later needs no update here.
5. New test/custom-model-run-menu-ui.test.ts drives the real session-ui.js
against a JSDOM window (runScripts:"dangerously" — this JSDOM only ever
parses markup this module generated itself) for exactly the DOM-level
facts the review said needed no Playwright and no tmux: a generated
button's onclick genuinely compiles and fires, a dangerous modelId never
produces a live element, the envelope unwrap works, the session-changed
guard holds, run() actually gets called (proving the in-flight lock
engages), and _runMode is restored afterward. Confirmed against the
pre-fix code first (reproduces btn.onclick === null exactly) so this
isn't a vacuous pass. Plus new tests in custom-model-routes.test.ts and
render-index-html.test.ts for the other fixes below.
Minors:
- Generated entries now filter through isCliAvailable(), matching
_refreshRunModeAvailability's own gating of the stock entries.
- The CRUD panel is now gated on customModelEndpointsEnabled
(applyCustomModelEndpointsVisibility(), wired to the toggle's onchange
and to settings-modal open) instead of always rendering; the endpoint GET
no longer fires unconditionally either.
- API keys are never handed back to the browser on GET, POST or PUT —
redactApiKey() replaces the field with a computed apiKeySet: boolean, and
a PUT with no apiKey now keeps the stored one server-side
(applyStoredApiKey()) instead of the client resending a value it was
never given. New tests cover both directions (kept vs. replaced) by
observing the actual auth header a subsequent discovery request sends.
- "+ Add endpoint" hides for a non-admin in multi-user mode
(_applyCustomModelAdminGate(), also wired to admin-ui.js's codeman:me
event, since the real role can resolve after settings were first opened)
— endpoint writes were already admin-only server-side, but the button
used to render for everyone and eat a 403.
- design doc (custom-model-endpoints-plan.md §4) now says up front that its
toolbar-button design was superseded by the Run-menu picker.
- docs/api-reference.md gained a Custom Model Endpoints section (every
route, the apiKeySet/defaultModelId contract, the restart mechanics).
- Wiki page now covers un-pointing a session (curl/delete, no UI yet) and
that the picker is desktop-only for now.
- .set-inline-form uses --control-bg instead of a hardcoded black alpha
(CLAUDE.md already records that exact literal turning the settings
preview into a grey slab on light skins), .run-mode-custom-models gets
the same gap: 2px .run-mode-menu's own flex gap only applies one level
up, and the index.html comment naming the wrong function is fixed.
- __codemanCustomModelClis's JSON is now escaped against a literal
</script> (CliEntry.label is user-clis.json-settable, unlike
__codemanCliAvailable's booleans-only payload) via a new exported
escapeScriptJson(), pure and unit-tested without needing a WebServer.
- Added defaultModelId + the new /v1/model-endpoints routes to
docs/api-reference.md; left the "no zh-CN for the new Models-section
group" minor unaddressed only insofar as the wider Models section (task
routing, thinking effort, etc.) has never had zh-CN coverage either —
everything this PR itself introduces (labels, hints, button text, the
Run-menu's "Custom Endpoints" header) IS translated in i18n.js.
Regression caught while fixing #4: the admin-gate's codeman:me listener is
a module-level document.addEventListener() call, which threw in
run-mode-ui.test.ts's minimal vm-context fake document and failed all 10
of that file's tests. Fixed with optional chaining before it ever reached
the branch this commit lands on; full targeted suite (route tests,
structural guards, every settings-ui.js-loading frontend test) reverified
green afterward.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
CI on PR #430 failed test/server-index-title.test.ts's byte-identity
check: renderIndexHtml now injects a second unconditional <script> before
</head> (window.__codemanCustomModelClis, added alongside the existing
__codemanCliAvailable one), and the test only knew to strip the older one
before comparing the rendered HTML against the raw template.
Strip both. Unlike __codemanCliAvailable (an object, historically injected
only where something resolved), the new one is a plain array injected
unconditionally, possibly empty, so it needs stripping on every machine,
not just one with CLIs installed.
Verified the two replace() calls compose correctly against the exact
strings server.ts actually produces (simulated in isolation; this box has
no tmux, so the real WebServer-backed test file cannot run here at all --
same environment gap noted throughout this PR's review).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
New docs/wiki/Custom-Model-Endpoints.md (auto-synced to the live GitHub
wiki on push to master, per docs/wiki/Contributing.md) covers turning the
feature on, adding an endpoint, the Run-menu picker's one-off-run
behaviour, the per-harness confidence table, and what it deliberately does
not do yet (remote/Docker sessions, live hot-swap). Linked from the
sidebar, from Agent-CLIs.md's "Read next" list plus a short pointer
section, and from Settings-Reference.md's Models section.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Follow-up to #393, picking up the work Ark0N invited in his merge comment:
"generate those entries from the saved profiles rather than a fixed
duplicate per harness, and put it in a follow-up PR so this one stays the
backend... The Run-menu picker is yours if you want it."
Adds the frontend surface the backend has been waiting on:
- Run menu: a "Custom Endpoints" section lists one entry per (harness that
supports customModelInjection, saved endpoint) pair, e.g.
"Claude Code (llama.cpp)". The harness list comes from
window.__codemanCustomModelClis, injected at page render straight off the
CLI registry's own capabilities (never a hardcoded id list in the
frontend), so a CLI whose injection recipe lands later appears with no
frontend change. Picking an entry runs that harness's own existing run*()
function unmodified (case creation, env overrides, everything, forced to
a single instance) and then applies the endpoint's default model to the
session it creates via the existing POST /api/sessions/:id/custom-model
route. Entries are hidden for a remote/docker active case, since that
route already refuses both.
- Settings: App Settings -> Models gets a "Custom model endpoints" group
wiring up the customModelEndpointsEnabled toggle (declared since #393,
read by nothing until now) plus CRUD against the existing
/api/model-endpoints routes: list, add/edit (inline form), delete,
discover models.
- Backend: CustomModelHost gains an optional defaultModelId, the model the
picker applies with no further choice per endpoint (one generated menu
entry per CLI+endpoint pair, not per CLI+endpoint+model). The route
refuses a value that isn't one of the endpoint's own discovered models,
and a fresh discovery drops a default that no longer appears rather than
carrying an invalid one forward.
Docs: docs/custom-model-endpoints.md describes the new picker and settings
panel; CLAUDE.md's Custom Model Endpoint Profiles entry drops the
"backend-only" status note and documents the picker's generation mechanism.
Tests: four new route tests cover defaultModelId validation, acceptance,
and the drop/keep behaviour across a re-discovery; a new render-index-html
test pins the __codemanCustomModelClis injection (present, agent CLIs
supporting the capability, antigravity and shell excluded) and its
solo-window skip. No browser test was added for the Run-menu picker itself
or the settings CRUD panel (this box has no tmux, so the live server used
by test:browser/test:mobile could not be exercised here) -- worth a
Playwright pass before merge, same as any other frontend PR.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
Found by driving the real UI: with a MAC configured in remote-hosts.json, removing it
(here: to reach the "Configure WoL" dialog) changed nothing for a running session —
_effectiveRemote short-circuited on the session's own snapshot whenever that snapshot
HAD a target, so the resolver was only ever consulted in the one direction where the
feature was missing. The documented "host config is authoritative" promise therefore
failed in the direction a user can actually observe, and a wake target could live on
invisibly after being deleted from the config.
The resolver is now consulted on the TTL regardless, and wins for the wake fields in
both directions. Also adds a route test for the browser's real input shape: one POST
per keystroke, all buffered during a wake, replayed IN ORDER.
Pressing Run on a remote case whose host was asleep failed with
`could not verify tmux on remote host 192.168.50.137: …` — an ssh error that
blames tmux for a machine that is merely suspended. The only wake paths were
typed input on an established session and the banner's Wake button, so OPENING a
session (the moment the user actually decides to use that host) had none.
`RemoteWakeRegistry.ensureHostAwake()` reuses the existing probe/wake/readiness
machinery for a host that has no session yet, and is wired into the two
user-initiated create paths: `POST /api/quick-start` for a remote case (before
the tmux prereq probe, which is what surfaced the misleading error) and
`POST /api/sessions` with `attachRemoteSession`. A host without a wake target is
not even probed, so its behavior and latency are byte-identical. The wake is
blocking — the caller gets the session or an error — but bounded by
REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS (40 s) instead of the 90 s session default,
because the dashboard sits behind a reverse proxy whose default
`proxy_read_timeout` is 60 s: a longer wait would be cut off at the proxy while
the session was still being created. The budget has to cover the whole request
(40 s wake + 1.5 s probe + the tmux probe's own 15 s = 56.5 s worst case), which
is why it is 40 s and not 45. A timeout now says the host did not come back, and
an unreachable host without a wake target says so instead of pointing at tmux.
The wiring is deliberately in the HTTP ROUTE, never in the shared session
service: `cron-service.ts` builds sessions there with nobody waiting on the
answer, and a wake on that path would power the host on for every schedule —
the timer-driven re-wake invariant #1 exists to prevent. Both halves are asserted
(importers of `remote-wake`, and `ensureHostAwake` having exactly one caller
file), so a future caller has to come through the guard test. A rejection from
the wake IO is caught too: a broken target must fail the wake, not the route.
`remote:hostWaking`/`remote:hostWakeFailed` now carry `forNewSession` for the
session-less case, where "input is queued" would be untrue; the toast then reads
"the session starts when it is back".
Live wake numbers are unchanged (this reuses the measured ~12 s S3 path); the
route behavior is covered by new tests in session-routes.test.ts with an injected
registry, so no test opens a real socket or ssh.
Two findings from a final review pass over the wake-on-LAN feature.
`_onRemoteHostWaking` / `_onRemoteHostWakeFailed` were defined in BOTH
`panels-ui.js` (toasts) and `host-wake-ui.js` (banner). Both files mix into
`CodemanApp.prototype` and `host-wake-ui.js` loads later, so the panels-ui copies
were silently shadowed: the toast never fired, and a wake started for a BACKGROUND
session (input on a non-active tab) produced no notification at all, since the
banner handler only acts on the active session. The handlers now live only in
`host-wake-ui.js`, show the toast unconditionally, and update the banner when the
woken session is the active one.
`appendBoundedPending` dropped only WHOLE chunks, so a single input value over the
cap (one large paste is one `input` value, up to the 100 KB input schema) was kept
in full: "bounded at 4 KB" held per chunk, not per session, and nothing was logged.
The surviving chunk's head is now trimmed too, code-point aware so a multi-byte
character is never split into a replacement char.
Adds the guard that would have caught the first one: every SSE dispatch handler must
be defined in exactly ONE frontend module. The existing test only asserts a handler
EXISTS somewhere, which two modules both satisfy while one is shadowed.
The wiki was written for seven run modes and never received Grok Build, DeepSeek
Harness or OMP. They now appear everywhere the others do: the modes table and
per-CLI notes, install commands, environment prefixes, the Quick Start table, the
requirements rows, the vocabulary, and every "seven modes" count.
The 1.27 to 1.29.0 changes land on the pages that own them: attaching a case to an
existing container, multi-case adoption and the copy-a-case picker (Docker Cases);
file reads over ssh in remote cases and what stays unavailable (Remote SSH Sessions,
Working With Files, Security); single-page app routing, frame recovery, localhost
links as tabs and the egress guard (Web Tabs); DeepSeek as the one non-Claude mode
with real stop/blocked signals and Approvals items, Codex's own work detection,
last-response, the model-endpoint routes and refreshed counts (HTTP API, Driving
From An Agent, Hooks, Notifications, Keeping Agents Running, Core Concepts);
Shift+drag, right-click copy, Auto Copy, the Ctrl+Z guard, font weight, the vertical
rail and its activity sort (Keyboard Shortcuts, Input And Voice, The Dashboard,
Settings Reference); the 600px phone cutoff, Codex shift arrows and iPhone Duo
(Mobile Guide); the Docker Compose route and its update rule (Installation, Running
As A Service); four new symptom entries and a "which CLIs" question (Troubleshooting,
FAQ).
Custom model endpoints are deliberately left to #430, which adds that page and edits
Agent CLIs, Settings Reference and the sidebar; these edits stay out of the regions
#430, #428 and #376 touch, and all three still merge cleanly on top.
Both READMEs: the web-tab menu entry is labelled "Add URL" in the UI, not
"Add dashboard".
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* chore: version packages
* chore: sync the CLAUDE.md version line to 1.29.1
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
Live terminal events are queued while a buffer load runs, and the load discards
that queue when it ends. That is right when the loaded buffer is the server's
accumulated byte history. The route appends to that history right up to the
moment it serializes the response, so a queued event already appears in it and
replaying it would duplicate output, most visibly Ink's cursor-up redraws.
A tmux pane capture is a photograph, current only as of the instant
`capture-pane` ran. Output printed afterwards was queued and then dropped, and
nothing scheduled a re-fetch to recover it: `_onSessionNeedsRefresh` is wired
only to the 128KB overflow path. The CLI's next partial redraw then landed on a
frame the terminal never received.
How much went missing depended on which capture the route served. A `?full=1`
load returns the capture alone, with no history in front of it, so it lost
everything from the capture to the end of the chunked write. A `?tail=` load
returns history, a clear, and then the capture, and the route reads that history
after the capture, so it lost everything from the response to the end of that
write. The chunked write dominates either way. An agent CLI hides the loss on
its next full redraw; a shell session does not, because its output is linear and
nothing repaints it.
Queue entries now carry their arrival time, and `_finishBufferLoad` takes a
`since` cutoff, so a capture load replays exactly the tail that arrived after
the response headers. The earlier events stay dropped, because a payload that
carries history does hold those.
All four paths that fetch a terminal buffer and write it now decide this the
same way, through one `_bufferLoadFinishOpts` helper, so they cannot drift
apart: `selectSession`, `_onSessionNeedsRefresh`, `_onSessionClearTerminal` and
`_maybeRefetchFullHistory`. The second of those is the one that stings. It
exists to restore output the client already dropped once under backpressure, and
it was dropping more output while performing that recovery. The cache-hit write
inside `selectSession` stays on discard deliberately: it runs before the fetch,
so its queue holds only events the capture that follows already contains.
Two further things had to change for that tail to still exist when the load
ends, and a browser test is what found both. `chunkedTerminalWrite` is what ends
the load for every non-empty buffer, so the flush policy travels to its own
finish calls; the call in `selectSession` runs only when the write was skipped.
`_beginBufferLoad` no longer empties the queue when one load re-enters it, which
it does on every write, because that reset discarded the whole fetch window
before anything could replay it.
The response already distinguishes the sources. `source` reads `mux-visible` or
`mux-full-history` for a capture and `history` for the byte stream.
Follows #395, #396 and #397, which fixed the ways the replayed frame itself
could disagree with the terminal.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Finishes #376. The contributed keystroke tracker sat on the raw byte stream
and named tabs wrong five ways (every prompt, every write path, a bare Esc
eating the next prompt's first character, pasted newlines as Enter, any CSI
clearing the draft) and replaced the whole name, which dropped the case from
the tab and reset the w<n> counter. This lands the feature with each of those
closed:
- First prompt means the first: applyAutoName() flips a placeholder to
`auto` whether or not the string changed. nameSource is now the tri-state
placeholder | auto | manual; the name setter is the only manual path.
- Only user-originated input counts: write()/writeViaMux() take
SessionWriteOptions.fromUser, set by the browser WS path and POST /input
only, so Ralph, respawn, cron, approvals and the trust-dialog keys can
never name a tab. A startMode 'shell' CLI never feeds the tracker (a
capability, not an id check); the send-key route feeds trackUserInput()
because its line feed bypasses the session.
- Prefix form `w3-case: title`: parseSessionPrefix() already renders it as
the title with the prefix in the tooltip and the next-session counter
still matches it. Composed within MAX_SESSION_NAME_LENGTH.
- Tracker rules per key: bare Esc resolves at chunk end; mouse/focus
reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R
taint the draft so Enter submits nothing rather than a fragment;
bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space;
the draft keeps its head past 8192 code points; an escape past 64 bytes
is abandoned.
- Title: slash commands by shape (a path is a prompt), `!` escapes
refused, first sentence only past 8 code points ("e.g." is not a title),
72 code points on a word boundary.
- Synced `autoNameSessions` setting, default OFF (the prompt reaches
mux-sessions.json, session:updated and /api/search), App Settings ->
Appearance -> Tabs, read fresh per prompt after the eligibility check.
Tests: test/session-auto-name.test.ts (tracker, title, composition,
ownership, emit gating), the wiring test (once, prefix, setting off,
manual protected), test/routes/session-name-routes.test.ts (PUT /name
flips to manual and persists). Verified live on an isolated instance: API
and browser-typed prompts name the tab, a second prompt does not, shells
and renamed tabs are untouched, nameSource survives a restart.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Reported as 'the tab shows no banner' while the host was verifiably unreachable: the
banner only started polling from selectSession, which RETURNS EARLY for the tab you
are already on (so a page loaded with the remote tab active never polled), and a
long-lived tab keeps running the JS it loaded — the feature was invisible to anyone
who did not switch tabs after the deploy.
The poller is now page-wide: one interval (created on init and on the first session
switch), re-targeted whenever the active session changes, plus a visibilitychange
wake-up. It no longer depends on any single selection path running.
Also adds test/sse-dispatch-table.test.ts: a static guard that every
[SSE_EVENTS.X, '_onFoo'] entry names an event constants.js defines AND a handler some
module defines. Both halves fail silently (a typo'd constant is an undefined table
key; a renamed handler just never runs), which is exactly how a new banner can never
appear with no error anywhere.
A configured-but-broken target (host replaced NIC, command removed) had no way
out: the dialog hung off the 'no target configured' branch only, so the banner
would keep offering a Wake button that keeps failing.
setBroadcast() on an unbound dgram socket throws EBADF on Linux and the following
send fails with EACCES, so the magic packet silently never left the machine — the
feature reported a wake that never happened. Caught by waking a real sleeping host
(a unit test with a real UDP broadcast would not be welcome in CI, so the socket is
injectable and the bind-before-setBroadcast ORDER is asserted).
The reactive wake (typing into a session whose host slept) left the state invisible:
nothing told the user the machine was asleep, and with no wake target configured
there was nothing to do about it. Adds:
- RemoteHost.wakeMac (comma-separated) - Codeman builds and broadcasts the magic
packet itself (UDP port 9), so the common case needs no external script. The
existing wakeCommand stays as the explicit override.
- GET /api/sessions/:id/reachability - probes (throttled, cached, and it never
wakes) and reports HOW the host can be woken, or that nothing is configured.
- POST /api/sessions/:id/wake - wakes, waits, reattaches the pane and flushes
buffered input; 400 with a routable message when no target is configured.
- The amber host-unreachable banner + its 'Wake' / 'Configure WoL' action, and a
small config dialog that saves via PUT /api/remote-hosts/:id.
- RemoteWakeDeps.resolveRemote: host config is re-resolved for LIVE sessions
(throttled + cached), so saving the dialog takes effect without a restart.
Self-review pass: the input-ladder's two 'buffer' branches were the same three
lines, and bulk delete left a session's (bounded, per-random-uuid) wake state
behind. Documents the design where the code refers to it - remote-sessions.md
section, the architecture invariant, and the CLAUDE.md key pattern.
A session's remote block is persisted at launch time and recovery uses that
snapshot, so a wakeCommand added to remote-hosts.json afterwards never reached
an already-running session - not even across a Codeman restart (observed: the
live Hufflepuff session came back with no wakeCommand). Merge the host-level
field in on restore, with the host config authoritative.
A durable remote session survives SSH drops (COD-104/108), but nothing brought
the HOST back: after the remote machine suspended, the local tmux pane's ssh
child stalled silently and `send-keys` SUCCEEDS against it, so typed input
vanished with no error anywhere.
Add an optional per-host `wakeCommand` (Wake-on-LAN wrapper, e.g. whuff) that
the input route runs when a wake-enabled host is unreachable: input is buffered,
the host is woken, the pane is reattached, and the buffer is flushed in order.
Detection is a throttled bare TCP probe on wake-enabled hosts only, and only
REAL user input may wake a host - the auto-reconnect watcher and boot recovery
deliberately cannot, or the host would be re-woken seconds after every suspend
and could never stay asleep.
Addresses the "left as they are"/"worth knowing" items Ark0N named when
merging #380 (the CLI-catalogue-driven install.sh + Docker agent image
PR), none of which were correctness-blocking but all of which were real:
- Removed install.sh's dead _cli_index/check_cli/get_cli_path helpers:
the catalogue-driven menu and hints stopped calling them and nothing
else ever did.
- The generator no longer emits CLI_KIND/CLI_NPM, two bash arrays
install.sh never read (the .mjs/docker-hosts.ts producers already
read the JSON catalogue's kind/npmPackage fields directly, so only
the bash copies were dead).
- detect_all_clis now skips a disabled entry's probe entirely instead
of running it and filtering the result downstream. No stock entry
ships disabled today, so this closes a latent inefficiency before it
is a latent bug rather than fixing an observed one.
- The install hint for a launcherProfile entry (DeepSeek today) now
explains in one line why it's a docs link and not a command: its own
docs page documents `npm install -g @deepseek-ai/dsh`, which installs
the launcher only and can't drive a pane, the exact trap the menu
already avoids by withholding the command. Driven by a new generated
CLI_LAUNCHER_ONLY array (from discovery.launcherProfile), not an id
check, so any future launcherProfile entry gets the same caveat free.
- Corrected the non-interactive-default comment: on a wget-only host,
Claude's curl one-liner is filtered out of the offered list first, so
the default becomes whichever npm-based entry sorts earliest instead
(Codex today), not always Claude. Behaviour is unchanged — it was
already printed, never silent — only the comment overclaimed.
Tests: extended test/install-sh-invariants.test.ts with a positive
guard for the new array and the trimmed array list, a negative guard
that CLI_KIND/CLI_NPM/the three dead helpers cannot come back, and two
real-bash tests (driven the same way the existing skip-menu tests are)
proving a disabled entry is genuinely never probed rather than merely
filtered after the fact.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011WzDjJnbK7zug8iQWnCc9z
DeepSeek Harness joins every CLI list it was missing from (tagline, intro,
run-mode table, Multi-CLI bullet with its env prefixes, security allowlist,
architecture diagram), and the 1.27 to 1.29.0 features get their bullets:
custom model endpoints (HTTP API only, with the verified and gapped CLIs
named), web tabs, attaching a case to an existing container, remote SSH file
access, the plan-usage chip, the sidebar and activity-sorted rail, font
weight, skins and entrance animations, Approvals Inbox, Read My Mind,
Claude-login voice dictation, Shift+drag select and right-click copy. The
agent guide's rule 7 now counts deepseek among the hook-signalling modes and
the recipes read answers through last-response first; the API section carries
the new routes and current counts; the download cap reads 2 GB instead of the
retired 50 MB; the zerolag package test count is the measured 238.
The English file had three spots where the OMP merge of 2026-08-18 left two
copies of a line joined without a newline (the Docker credentials bullet, rule
7 of the agent guide, the CLI node of the mermaid diagram). All three are
single lines again.
The Chinese file was further behind: besides the above it had never received
the daemon and service block, the Tailscale install option, the Compose
paragraph, the Tab Alerts section, the codeman tui section, the agent-skill
walkthrough, the Community section or the closing star paragraph. Those are
translated in, so both files now share one section structure.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
flushPendingWrites() captured the viewport of a user who was reading
scrollback, called terminal.write(), and restored the anchor on the next
line. xterm parses on its own schedule, so at that point the buffer has not
moved: the guard `viewportY !== preserveViewportY` was false, scrollToLine
was never called at all, and the Codex redraw landed a tick later and took
the viewport to the live bottom with nothing left to pull it back. Scrolling
up during a stream still got dragged down, which is what #358 reports, and a
refresh was the only way back to a coherent view.
The restore moves inside xterm's write callback, the first moment the
redraw's effect exists, and runs before _scheduleTerminalWriteFlush() so a
deferred remainder re-captures the restored anchor rather than the bottom.
Two things follow from it running later:
- A live anchor now wins over the sticky scroll-to-bottom. The two are
captured at different moments (_wasAtBottomBeforeWrite at the frame's
first batchTerminalWrite, the anchor at flush time), so a scroll-up in
between leaves both set, and running both would jump to the bottom and
come back a frame later instead of staying put.
- The anchor is dropped if the active session changed or a buffer load
started while the write was in flight. It indexes the buffer it was
captured from, and selectSession() resets the terminal and chunk-loads a
different scrollback.
The existing regression passed throughout, because its write mock moved the
viewport synchronously, which real xterm never does. The harness now models
an asynchronous parse (redraw lands, then the callback fires), and all five
of the anchor tests fail against the old code.
Fixes#358
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three behaviours landed from #375 without their doc entries: the
duplicate input ACK now carries `dup:true` and the server's watermark
(`docs/reliable-input-delivery.md` still described a bare ACK), Shift+drag
and right-click copy in the terminal (the shortcut list did not know
them), and one adopted container backing several cases at different
in-container directories (the Docker cases paragraph still implied one
case per container for adopted containers too).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The cherry-picked "copy an existing case" commit declared a second
`CaseInfo.docker.owned` and emitted `owned: true|false` on every docker
case, while master had meanwhile shipped the same field from the
adopted-container work with a narrower wire shape: `owned` is present
only when false, absent means owned. Two declarations failed typecheck,
and two emit styles on one response would have made the picker's answer
depend on which read path filled it.
Keep master's shape at both response sites (the case list and the
single-case lookup, which lacked the field entirely), fold the picker's
reason for the field into the existing doc comment, and repoint the test
that pinned "set on exactly two sites" at the surviving form, adding a
negative pin so the duplicate style cannot come back.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The previous version cleared the case name and the in-container directory on the
grounds that they must differ. That left a form with three fields mysteriously
filled and two empty, and turned the most common operation — changing
/srv/app/api to /srv/app/web — into retyping a long path.
Both are now pre-filled, with focus on the in-container directory and the caret
at the end, since the tail is what changes. What stops an unmodified submit is no
longer an empty field but a guard: the values applied are recorded, compared at
submit time, and if nothing changed the reason is stated next to the field and
focus moves to it, without sending a request that is certain to be refused.
The server refuses these anyway (a duplicate case name, a twin case on the same
container and directory) and its errors are clear; but making a round trip to be
told "you forgot to edit the field you are looking at" is worse than saying so on
the spot. The guard only applies when a source case was actually selected, so
filling the adopt form from scratch is unaffected.
⚠️ The status text is written into dockerLinkStatus. My first version referenced
an id that does not exist (dockerAdoptStatus), which made the explanation vanish
silently and left only a toast. The test now extracts that id from the code and
looks it up in index.html, pinning that it must really exist.
(cherry picked from commit ba21ae11f4)
The backend already lets one adopted container back several cases pointing at
different in-container directories, but using it meant retyping the container
name, host and workspace one by one — exactly the friction that leaves a
capability unused. Picking an existing case from a dropdown now carries those
three over, leaving only the two fields that must differ: the case name and the
in-container directory.
Clearing those two is the point of the feature, not a convenience: keeping the
old name is refused by the server as "case already exists", and keeping the old
directory is refused as "a twin case on the same container and directory". Both
errors are clear, but a form pre-filled with values that are guaranteed to be
rejected is a trap. Focus lands on the in-container directory — the thing the
user came here to change.
⚠️ Only adopted containers are listed (docker.owned === false). A Codeman-built
container's lifecycle belongs to its one case — a second case would be torn out
by that case's recreate or delete — so the server refuses it anyway, and listing
it here would only manufacture a baffling error. `owned` may be absent and absent
means owned, so the test is `!== false`, not truthiness.
CaseInfo.docker gains containerWorkdir and owned for this: the former is the
"which directory does this case use" half of the picker, without which the user
cannot tell what to change it to; the latter backs the filter above. ⚠️ Both
places that build a docker CaseInfo (the list endpoint and the single-case query)
must set them — filling in only one makes the picker work or not depending on
which read path was taken, and a test pins "exactly two".
(cherry picked from commit f1ed3a58e1)
Once a container is adopted, it could not be adopted a second time. But a
container usually holds more than one project directory, and opening a case for
another one had no path forward except starting a second container — precisely
what adoption exists to avoid.
The original reason was in a comment: two cases sharing an adopted container
would make one case's teardown race the other's launch on the same tmux server.
That reason does not hold. The in-container tmux session name is
dockerTmuxSessionName(sessionId), i.e. codeman-dkr-<id8>, keyed by SESSION and
not by case, and buildDockerKillCommand tears down exactly that name, so killing
A never touches B — hosting multiple sessions is what a tmux server is for.
The other three routes into an adopted container's lifecycle do not pass through
here either, confirmed one by one: the stop and remove builders throw outright;
recreate refuses `owned === false` before it even resolves the container name;
and orphan reaping filters on `label=codeman.managed=1`, which a user-built
container does not carry — a structural exclusion.
That leaves exactly three cases worth refusing, none of them tmux-related, split
into the pure, unit-tested classifyAdoptContainerConflict:
- owned-case the container belongs to a Codeman-created case, whose lifecycle
Codeman manages: one recreate or delete there would pull the
container out from under the adopting case.
⚠️ `owned` may be absent and absent means owned (cases predate
the field), so the test is `!== false`, not truthiness.
- other-owner already adopted by a different user. Adoption hands out a shell
inside someone else's container.
- duplicate same container, same directory. The second case would behave
identically to the first, so name the existing one rather than
silently minting a twin. A different in-container directory is
the case this change exists to support and passes.
(cherry picked from commit 1cb6bde891)
`backdrop-filter` promotes an element to its own compositing layer. A
position:fixed full-screen layer that is created and then hidden was measured to
leave a stale hit-test region behind in Chrome: the page renders perfectly, but
pointer events across the viewport go nowhere.
The report came from a long-lived tab connected to a remote server, where a
connection blip shows and then hides #offlineOverlay. The symptoms were a
terminal that would not scroll and, at the same time, an unrelated
click-to-expand that also stopped responding, while a freshly opened tab was
fine; a read-only console command (getComputedStyle + elementFromPoint, both of
which force a hit-test recomputation) then cured it. Two unrelated features
dying together and one read-only command fixing both points at hit-testing
itself rather than at either feature.
So the `backdrop-filter` moves onto the actually-visible selector and the layer
is never created while hidden. Only the two persistent overlays change:
offline-overlay (toggled with [hidden]) and file-preview-overlay (toggled with
.visible). path-picker and path-preview are created and removed by JS, leave
nothing behind, and are untouched.
⚠️ This is an evidence-based inference, not a fix verified by reproduction:
reproducing it needs a long-lived page that has been through a connection blip,
which I could not manufacture in a controlled environment. The guard test pins
both halves — no such property while hidden, and a real blur while shown — so a
later cleanup cannot quietly delete the effect.
(cherry picked from commit 08442dfee1)
handleInit() did not distinguish a first load from an SSE reconnect: it always
cleared the terminal caches in _resetAllAppState() and re-ran selectSession() for
the session that was already on screen. Every reconnect therefore refetched up to
1 MiB of buffer and reset+rewrote xterm. On a link that drops a connection about
once a minute (measured at ~57s intervals against a healthy server) that reads as
the page refreshing itself and throwing away your reading position.
A reconnect that lands back on the still-open session now keeps the terminal
caches and activeSessionId and resyncs through _onSessionNeedsRefresh(). That
path still reloads the buffer, so output produced during the outage is not lost,
but it preserves distance-from-bottom — the same rule #259 established for a
refresh the server triggered rather than the user. The WS is reconnected
explicitly when it is not already on that session, since skipping selectSession()
skips its _connectWs() call.
First load (gen === 1) takes exactly the path it took before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rv24Pk4qzrsDYdVyDyJQmT
(cherry picked from commit 435569c76e)
Browser input is delivered exactly once by (clientId, seq). The server records a
watermark per clientId and discards anything not above it as a duplicate — but
acknowledged it with an ACK indistinguishable from "applied". The client then
dropped the record from its queue, the UI looked perfectly normal, and the
terminal received nothing at all.
The counter is persisted to localStorage through a debounced write. Kill the page
between "sent" and "persisted" and the restored counter is below the server's
watermark, after which every keystroke lands under it, is discarded, and is
ACKed. Reloading does not help: the clientId is restored from localStorage
alongside that stale counter. Measured on a real session — typing into the same
session from a fresh browser (new clientId, no watermark on the server) worked
perfectly, which is what localised the fault to client state.
Three changes:
- on rejection the server replies {"t":"ia",seq,"dup":true,"last":<watermark>}.
It still ACKs, so the client can drop the record from its queue, but it now
says the input was not applied and supplies the number needed to climb out.
- on `dup` the client lifts its counter above the watermark and re-queues.
⚠️ Only records whose FIRST delivery is being retried are re-sent: a retry
judged duplicate means the mechanism is working (the original did arrive), and
re-sending would type the same text twice.
- the counter is now persisted synchronously. The queue payload can stay
debounced, but the counter is the thing that has to survive a crash, and
leaving it on the lossiest path cancels the only guarantee there is.
⚠️ Reading the watermark is defensive: the session arrives through a structured
port, and a port missing that method must not take the whole input path down —
a throw inside the handler means the ACK is never sent and the record is stuck in
the client queue forever, which is worse than the ambiguity being fixed. A mock
port's test timeout is what exposed this.
(cherry picked from commit 05bb7081cc)
claude advertises "Jump to bottom (ctrl+End)", so that chord has to actually
reach it. But PASSTHROUGH_KEYS carried only the bare forms (End -> \x1b[F) and
CTRL_KEYS held just six letters (c/d/l/z/a/e), which cannot express End. Ctrl+End
therefore failed in both directions:
- with an empty composer it went out as a bare \x1b[F, the modifier silently
dropped, so the CLI received a plain End;
- with text in the composer the forwarding branch requires empty, so nothing was
forwarded and the browser default applied — the caret jumped to the end of the
draft, which is the "the shortcut now edits my input box" the user saw.
Encode them as CSI 1;<mod><final> instead, and forward Ctrl/Alt-modified
navigation keys whether or not the composer is empty: they are commands for the
CLI, and the composer has no editing semantics for them worth preserving (bare
Home/End still use the old table and edit locally).
⚠️ Bare Shift is deliberately excluded: Shift+arrow selects text in the composer,
a real editing gesture that must stay local. Shift held together with Ctrl/Alt is
still encoded into the modifier mask.
(cherry picked from commit 3fbaadadfb)
In a native terminal running a TUI with mouse tracking on (claude, codex), Shift
is the "let me select text" modifier: it bypasses the application's mouse
reporting so the emulator selects locally. Users bring that habit here, where it
did nothing — measured, `hasSelection` was already false during a Shift+drag and
no clearSelection call ran at all, because there was never a selection to clear.
The mismatch is that the two Shifts mean different things. xterm reads Shift as
"force selection", but that path is only taken when the application really has
mouse tracking on. The server strips the mouse DECSETs for claude/codex/gemini
(isAltScreenStripMode), so xterm's mouseTrackingMode is permanently `none`, that
branch is unreachable, and Shift instead lands in _onIncrementalClick — which
EXTENDS an existing selection. Extension is a no-op while selectionStart is
empty, so the drag had no anchor.
So plant the anchor xterm is missing. The listener sits on the capture phase of
the `.xterm` root, an ancestor of the `.xterm-screen` that SelectionService binds
to, and therefore runs before xterm's own mousedown; xterm then extends from our
anchor and the drag behaves like any other. Length is 0 so a Shift+click without
a drag does not select a stray character. An existing selection is left alone —
that is a genuine extend gesture, and xterm handles it correctly.
Right-click copies the selection (the mintty/PuTTY convention), completing the
gesture: until now there was nowhere for a finished selection to go. With no
selection the native menu is not hijacked — taking it away while offering
nothing in return is a pure loss.
(cherry picked from commit 7ab5015737)
scripts/test-local-llm-harnesses.ts (#393) sits outside tsconfig.json's
include, so nothing type-checked it. config/tsconfig.scripts.json pulls
it in; npm run typecheck now runs both projects, the way the pr-bot
config used to be chained before the bot moved out of the repo.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Custom Model Endpoint Profiles (#393) let a session point its CLI at a
custom OpenAI-compatible endpoint by injecting env vars or a config file
and restarting the CLI in place. Review of the apply path found four
things, two of them destructive. This lands all four plus the smaller
items from the same review.
1. Clearing a selection did not clear it. The injected vars reach the CLI
via `tmux setenv`, which persists at the tmux-session level and is
inherited by `respawn-pane` (measured: `setenv FOO bar` survived two
successive `respawn-pane -k`), so deleting the keys from the session's
envOverrides relaunched the CLI still pointed at the old endpoint, and
for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just
been deleted. `Session.setCustomModel()` now reports the removed keys,
queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys`
carries them into `applyEnvOverrides()`, which `setenv -u`s them before
re-applying the live overrides, on the same path that already unsets
the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket
that `setenv -u HOME` hands the next respawn the global HOME back.
2. Applying a model to a local claude session killed the pane. The
relaunch was `claude --session-id <id>` and Claude refuses an id that
already has a transcript, and unlike the dead-pane respawn this one
kills a working pane first. `restartCli()` now pins the live
conversation id as the resume id for that respawn when the CLI's launch
declares a `fallback` chain, which renders the same
`--resume <id> || --session-id <id>` shape the docker and remote pane
commands use. Gated on the registry shape, not the CLI id: an entry
whose resume id is minted by the CLI itself never declares that chain.
3. pi, omp and grok wrote their config file and then launched without the
`--model` that selects it, so the file was ignored. The registry entry
now declares `customModelInjection.launchModel` (`custom/{modelId}` for
pi and omp, grok's `[model.codeman-custom]` block name), the builder
renders it, and `_withCustomModelLaunchModel()` applies it onto the
respawn options through `legacyConfigField`, leaving the stored
<Mode>Config untouched so a clear falls back to the user's own model.
A model id the CLI's `model` token pattern cannot carry is refused
with a 400 rather than silently dropped by the argv engine.
4. Remote (SSH) and Docker sessions reported `restarted: true` and changed
nothing: their `restartCli()` reattaches the durable tmux rather than
relaunching the agent, and the env lands on the local pane. Both are
refused with a 400 until those paths are plumbed.
Smaller items from the same review:
- The selection survives a Codeman restart as the disk-only `__customModel`
bookkeeping (endpoint, model, injected key NAMES, config dir, launch
model; never the values, which carry the API key). Recovery re-derives
the values from the endpoint store through the same apply path the route
uses and keeps the bookkeeping even when the endpoint is gone, so a
later clear still has keys to unset.
- Discovery goes through `webviewFetch()`, so the RESOLVED address is
judged by the same egress guard the web-tab proxy uses, and `baseUrl`
reuses `webviewUrlSchema` (http(s) only, no embedded credentials,
link-local and cloud-metadata addresses refused). undici's `fetch failed`
wrapper is unwrapped so the user sees the ECONNREFUSED underneath.
- `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session
config dir 0700/0600 (pi and omp embed the key literally), and that dir
is removed with the session.
- `PR.md` is gone from the repo root and the design doc moved to
`docs/custom-model-endpoints-plan.md` with the LAN address and the
personal name scrubbed; every reference follows. The guide's `authStyle`
text matches the shipped schema (`bearer | api-key`, default `bearer`)
and says that `customModelEndpointsEnabled` is read by nothing until
the picker lands.
- `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts`
(four real type errors fixed). It is not yet wired into `npm run typecheck`
because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json`
there is the one-line follow-up.
Tests: `test/session-custom-model-restart.test.ts` drives a real Session and
fails on the unfixed code for items 1 to 3; the route suite covers item 4
and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets
run before the overrides and that a shell-metachar key never reaches tmux.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Follow-up to #421 (remote-case file reads over ssh), addressing the review.
Symlink escape on a host without `readlink -f` (blocker). The probe's
portable fallback canonicalized only the directory chain and returned the
final component unresolved, so on macOS < 12.3 `ws/notes.txt -> ~/.ssh/id_rsa`
came back as `.../ws/notes.txt` (with the target's size), passed every
containment and blocklist check that runs on `realPath`, and `cat` followed
the link. The fallback now walks the directory chain with `cd -P`/`pwd -P`
and follows the LAST component with plain `readlink` for a bounded number of
hops, and anything it cannot fully resolve (a loop, a readlink failure, the
hop cap) is reported with an `x` marker that parses as null, i.e. 404. It
never returns the unresolved string. Measured on a real /bin/sh with
`readlink -f` shadowed: the pre-fix script reports `/ws/notes.txt`, the fixed
one `/secret/id_rsa`; both branches (native and fallback) now agree.
`PUT /api/sessions/:id/file-content` never had the remote guard the PR
described. It sits ahead of `validateSessionFilePath`, which resolves against
the LOCAL filesystem, because with a same-named directory on the Codeman host
(an sshfs mount of the remote tree, the documented stop-gap) the write landed
on the local twin while the viewer believed it edited the remote file.
ssh fan-out is bounded. `src/remote-ssh-limiter.ts` is a
document-conversion-limiter-shaped semaphore (default 4, env
`CODEMAN_MAX_REMOTE_FILE_SSH`) around every probe and buffered read; the
attachment-history list resolves its whole history in ONE batched probe
(`probeRemoteAttachmentHistory`, threaded into
`registerExternalAttachment({remoteProbes})` so the guards run unchanged)
instead of one handshake per entry; and probes chunk at 40 paths because the
whole script is one argv string. Terminal output in a remote session is
written on the remote host, so a prompt-injected agent printing hundreds of
`codeman://attach` links forked one ssh per link, each holding a 20 s
timeout, and a 100-entry history re-listed on every attachment:detected
tripped OpenSSH's default MaxStartups. Streams are deliberately not counted
(one per browser request, held for a whole playback, and gated behind a
counted probe anyway).
Smaller items from the same review: probe records are NUL-terminated and
index-keyed after a leading NUL (a newline in a filename can no longer shift
the alignment, and the banner is fenced off without last-N-lines guessing);
size comes from `stat -c %s || stat -f %z`; the three IO functions refuse
under VITEST instead of opening a connection; an unreachable host now reads
as unknown (missing: false) for detected AND external history entries, where
external used to fold its 502 into missing; a client that aborted during the
guard probe has its body's ssh child reaped (`reply.raw.destroyed` is checked
before the close listener is attached); `describeExecError` never returns
Node's `Command failed: <ssh line>` message, which carried the identity path
and the probe script into a 502 body; and the docs note that
`isSensitivePath`'s three home-anchored entries resolve against the Codeman
host's home, not the remote one.
Tests: the probe script runs on a real /bin/sh with a `readlink` shim that
rejects `-f` (the escape, a relative chain through a symlinked directory, a
loop, a newline filename, banner chatter that itself looks like a record),
the limiter's cap and FIFO order, and route tests for the PUT guard (local
twin untouched, no connection), the single batched history probe, the
unreachable-host alignment and the aborted-client reap. All four route tests
fail against the pre-fix file-routes.ts.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The lost-frame recovery page is answered ahead of the credential checks in both
auth hooks, which makes it the third unauthenticated 200 beside the two hook
routes, and the only one decided by request headers alone. CLAUDE.md's security
table listed exactly two, and docs/web-tabs.md is not where anyone auditing that
looks, so it now has a row in the table and a fourth property in
docs/security-architecture.md section 10b, including the `/` carve-out and its
credential-free condition. Both state the property that comes with it: a
non-browser client can set those headers, so an unauthenticated caller can tell a
registered route (401) from a non-route (200) and enumerate the route table,
accepted because the routes are public in docs/api-reference.md.
docs/web-tabs.md gains the landing-page case in layer 6 and a Known limits entry:
masking trades away the Referer safety net, only HTML is rewritten server-side,
and a root-absolute url() inside an inline <style> block has the masked document
as its Referer, so it 404s where the Referer fallback used to rescue it. External
stylesheets are unaffected.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The runtime shim masks `/webview/<cap>/` off a proxied page's URL so its router
boots on the path it expects, and the landing page masks to exactly `/`. A
`location.reload()` there (a Vite dev server on a config change or a failed HMR
update, the likeliest case in the feature's own motivating scenario) therefore
asks for Codeman's root as an iframe navigation. `serveLostWebviewFrame()`
returned early for `/`, so on a passwordless install the frame received Codeman's
own app shell and rendered it inside the web tab, and with a password it got a
401 in the frame. Either way no `codeman:webview-lost` message was posted, and
because the document loaded fine the load handler cleared the failed-frame panel,
so the Reload / Open in new tab affordances never appeared. Before masking the
frame's URL was the prefixed one, so a reload worked; this was a regression.
`/` is the one lost-frame path a registered route also serves, so the route
table cannot tell that reload from a real navigation. Credentials can: nothing
in Codeman frames its own root, and a sandboxed frame is opaque-origin with no
cookie and no Authorization header. `carriesAuthCredentials()` (pure, in
webview-proxy.ts) makes that test, and `/` is now admitted by the auth hook only
when it fails; a framed `/` that does carry credentials still gets the shell.
Without a password no auth hook runs at all, so the index route applies the
same test itself (`isLostWebviewRootFrame`) before rendering the shell, and the
three places that emitted the recovery page share `sendLostWebviewFramePage()`.
Tests: the password form in webview-auth-exemption (recovery page for a
credential-free framed `/`, shell with valid Basic auth, 401 with a stale cookie
or a top-level navigation), the passwordless form against a real WebServer in
webview-lost-root-frame (port 3198), and the credential predicate in
webview-proxy. All three fail without the fix. Verified against a live isolated
instance as well: a framed `GET /` with no credentials answers the 470-byte
recovery page, a top-level `GET /` and a framed one carrying a cookie answer the
shell.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The lost-frame handler in webview-tabs.js remounts a web-tab frame at the path
the frame reports it lost. It promised "path only, never an origin" and collapsed
a leading run of slashes so `//host/x` could not jump the frame off the proxy,
but it left two spellings through that the WHATWG URL parser treats the same way:
a backslash, which is read as `/` for http(s) schemes, and an ASCII tab or
newline, which the parser deletes before it looks at anything, so `/\host/x` and
`/<tab>/host/x` both resolve to `https://host/x`. That mattered only in
direct-mode tabs, where `POST /api/webviews/:id/open` returns no embedUrl and the
recovered path is resolved with `new URL(path, src)` straight into the frame's
src; a page in such a tab could remount its own frame on a foreign origin.
Not an escalation (the page can already navigate itself anywhere, and the remount
carries no Codeman-origin access), but the comment did not hold and the existing
test only covered the form that already worked. The handler now strips tab, CR
and LF, collapses any leading run of `/` or `\` to one `/`, and refuses whatever
still opens a second separator. The new test drives the reachable direct-mode
branch with all four spellings and fails without the fix.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The capability list is DERIVED from what the scripts do (chown => CHOWN +
DAC_OVERRIDE, a setpriv uid/gid drop => SETUID + SETGID, `init: true` next to a
uid drop => KILL) and compared to docker-compose.yaml's cap_add, the
entrypoint's own required_caps diagnosis, and the lists quoted in docker/README.md
and CLAUDE.md, so the drift that shipped the missing CAP_KILL fails here rather
than on someone's server. Also pinned: the CLI prefix is appended to PATH in
server.Dockerfile and entrypoint.sh pins its PATH before its first command;
Start-Codeman.sh derives PUID/PGID before creating the cases dir, builds before
`down`, writes the source marker only after a refresh, and never aborts on a
failed volume removal.
git_head_commit is run as the script defines it, extracted by its own
delimiters into a real bash, against temp repos made with real git: a symbolic
ref with a loose ref file, a detached HEAD, packed refs after `git pack-refs`,
a linked worktree (which must resolve nothing rather than something wrong) and
a directory that is not a checkout.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
docker/README.md and docs/docker-compose.md now say that the container starts
as root, corrects a daemon-created bind source and drops to PUID:PGID with
setpriv, which capabilities that needs, and that a compose file written
elsewhere must carry them. The README's PowerShell example runs Compose from
inside docker/ so the override file is discovered, instead of the `-f
docker/docker-compose.yaml` form its own Local customisation section warns
silently drops it, and the reverse-proxy section no longer asks for an override
file now that docker-compose.yaml forwards CODEMAN_ALLOWED_HOSTS itself.
.dockerignore excludes docker-compose.override.* everywhere: it is the
documented home for host-specific settings and rode `COPY . .` into the image,
the same shape as the docker/.env exclusion above it (verified with a scratch
build context: the override files and docker/.env are absent, .env.example and
the compose file present).
CLAUDE.md's Compose paragraph carries the corrected cap list, the writability
probe, and the two traps behind it (KILL is for tini, the CLI prefix is
appended to PATH).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Start-Codeman.sh created CODEMAN_CASES_PATH with a plain `mkdir -p` BEFORE it
derived PUID/PGID from the appdata directory, so the new directory landed as
the invoking user's uid and primary gid. On a host set up the way the README
suggests (`chown -R 99:100 <appdata>`) that gid is not PGID, and the container
refused to start on a directory the script had just made. PUID/PGID are now
derived first and the directory is chowned to them right after creation, with
a clear host-side error when that is not possible. As root this always works,
which also retires the old "refusing to create as root" branch for this path.
The build-artefact volume refresh had three holes. The docker-build-source.json
marker was written whether or not a volume had actually been removed, and the
project name came from a sed over `docker compose config --format json` keyed
on two-space indentation: an empty name made the label filter match nothing,
nothing was removed, and the marker recorded the new HEAD, so the check never
fired again while the stale volume kept serving old code. The name is now
parsed indentation-agnostically, an empty result falls back to `down --volumes`
(the documented reset; both volumes re-seed from the image by a plain copy),
the marker is written only after a successful refresh, and a failed `docker
volume rm` warns and leaves the marker alone instead of aborting under set -e
with the stack down. The image is also built BEFORE `down`, so the deployment
is offline only for the recreate rather than for the whole rebuild.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Three changes to how the Compose container starts as root and drops to
PUID:PGID, each reproduced on Docker 29.1.3 / Compose v5.5.0 with a minimal
image of the same shape as server.Dockerfile.
- cap_add gains KILL. `init: true` makes tini PID 1, and tini stays root while
the entrypoint drops the server to PUID. Signalling a process of a different
uid needs CAP_KILL, and `cap_drop: ALL` had removed it, so every `docker
compose down`/`restart` ended in `[FATAL tini (1)] Unexpected error when
forwarding signal: 'Operation not permitted'` and the server being SIGKILLed
instead of running `server.stop()`. Measured: without KILL the trap never
fires, with it the child logs `GOT SIGTERM`.
- /opt/codeman-cli/bin is appended to PATH, never prepended, and entrypoint.sh
pins its own PATH to the system directories before its first command. The
prefix is chowned to the runtime account so sessions can update the agent
CLIs in place, and the root entrypoint resolved stat/chown/setpriv by bare
name through it: a `setpriv` planted there by the unprivileged uid ran as
uid 0 at the next start. The image's full PATH is handed back to the server
at the exec (`env PATH=...`), since Codeman resolves the CLIs through it.
- The ownership gate becomes a writability probe. A directory owned by neither
root nor PUID:PGID is no longer refused on ownership alone; it is tested with
`setpriv --reuid PUID --regid PGID --groups <same groups> test -w`, the exact
identity the server gets, so a group-writable tree, an ACL or a CIFS/NFS
mount reporting some unrelated uid all pass, and the refusal names path,
owner and PUID:PGID. Root-owned directories are still chowned first.
Also: a pre-flight runs the drop before touching anything and, when it fails,
prints the cap_add list the compose file needs, so an out-of-tree compose file
(Unraid's Compose Manager) gets a one-line diagnosis instead of a restart loop;
`--bounding-set -all` is gone, since it is a silent no-op without CAP_SETPCAP;
a root:root Docker socket now produces a warning that Docker cases will not
work rather than silently losing group 0 at the drop; and CODEMAN_ALLOWED_HOSTS
is forwarded from .env with an empty default (documented as a commented entry
in .env.example so the parity test and the updater's env gate both stay quiet).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The bash 3.2 job added with #380 cannot reach dsh_banner_probe, which is
the function #382 was filed against: this image ships `timeout`, so the
optional-prefix array is never empty, and with no `dsh` binary anywhere on
PATH the probe is not called at all. The fix landed in 1.28.2 with nothing
guarding it, and the failure mode is a runtime abort under `set -u` that
`bash -n` cannot see, which is precisely why the reporter had to find it by
reading the source rather than by running anything.
So call the probe directly, with `timeout` hidden behind a narrowed PATH,
and refuse to pass if `timeout` is still reachable (a guard that silently
stops exercising its branch is worse than no guard). Both directions are
asserted: a real DeepSeek Harness banner is accepted, and Debian's unrelated
`dsh` is refused, so the check covers the identity half too.
Verified by reverting install.sh to the pre-fix expansion, where the step
fails with the exact error from the issue, `runner[@]: unbound variable`.
Refs #382
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The two keys #408 adds to the mobile keyboard accessory bar send
Shift+Left and Shift+Right, which are Codex bindings (edit the last
queued message, step back through the prompt stack). They shipped on
both agent layouts, so a claude, pi, grok, omp, deepseek or gemini
session got two keys that do nothing. That was not only cosmetic: a tap
goes through sendNavKey(), which adds the session to
_echoPassthroughSessions and hands editing to plain PTY echo until Enter
or Ctrl+C, so on a phone a dead key also switched off the local echo
that makes typing feel instant there.
The reveal now follows the shape the 🧠 key already uses. The buttons
stay in both templates, carry an accessory-btn-codex marker class, and
are display:none in styles.css until the bar element carries
codex-enabled. The class has to live on the bar rather than on the keys
because setMode() rebuilds the buttons' innerHTML on every layout
switch. syncCodexKeys() toggles it from the active session's mode
(the same lookup _isShellSession() uses) and is called at init and from
refreshForActiveSession(), which selectSession() already invokes on
every switch. A session's mode is readonly on the server and fixed at
create, so no other event can change the answer; the welcome screen
(no active session) reads as not codex and hides the keys.
The frontend id-branching guard (test/cli-registry-no-id-branching.test.ts)
scans only src/**/*.ts, so the mode comparison in a public JS file is
in bounds, the same as the existing shell check beside it.
Tests: the new describe block in test/mobile-shell-keyboard.test.ts pins
the marker class in both templates, the CSS pair, the class for a codex
session in both layouts, its absence for claude/shell/pi/omp/deepseek,
the re-sync in both directions on a session switch, the no-session case,
and the init + refresh wiring. All six positive assertions fail without
the source change. README and the changeset now say the keys are
Codex-only.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A clicked path that points OUTSIDE the case directory goes through the attachment
routes (the frontend's `_isExternalPreviewPath` sends every absolute path not under
`workingDir` to `POST /attachments`), and those had the same local-`fs` assumption
as file-raw: `realpathSync`/`fs.stat` on a path that only exists on the remote host,
so the file never opened — the case the #415 report was actually about.
- `registerExternalAttachment()` accepts `remote` and resolves through
`remoteProbePaths` (canonical path, size/mtime, kind, plus the workspace root for
the confinement check). Everything around it — blocklist, extension allowlist,
workspace confinement, registry/dedupe — is now shared by both branches, so the
remote path cannot drift from the local one.
- The by-id routes (`raw`, `preview`, `thumbnail`), the metadata poll and the
attachment history list resolve over ssh too. `raw` streams with the same
Range contract as file-raw; `preview` (office) and `thumbnail` answer 400 for a
remote record; an unreachable host answers 502, a vanished file 404.
- Which host a record is read from follows the SESSION, never the path string: the
same absolute path is a different file on each host, and a remote session never
falls back to a local file with that name.
- Codex generated artifacts keep force-workspace confinement for a remote case: the
well-known artifact directories are anchored at THIS host's home, so only a file
inside the remote workspace is trusted.
Still local-only by design: writes, office conversion, thumbnails, the file
tree/picker and tail-file.
A remote case's workingDir is an absolute path on the remote host, but the
file read routes resolved it with local `fs`: `validateSessionFilePath`'s
realpathSync fails for a path that does not exist on the Codeman host, so
every preview of an agent-written file answered "File not found" (#415).
Add src/remote-files.ts as the single remote-read layer, built on the same
buildSshConnectionArgs() the launch uses:
- remoteProbePaths(): ONE round trip returning realpath + stat for the
requested path AND the workspace root, so containment is checked against a
remotely canonicalized root (a symlinked remotePath is ordinary).
- remoteCreateReadStream(): streams the body (cat, or tail -c +N | head -c L
for a Range) with nothing buffered in memory, and reaps the ssh child when
the response ends so an aborted download cannot orphan it.
- remoteReadFile(): bounded read for file-content.
file-raw, file-content, file-preview and file-thumbnail now share one local/
remote target resolution. Guards keep their local strength: lexical pre-check,
remote realpath, workspace containment, sensitive-path blocklist, and the size
cap applied to the remote size before any bytes are read. An unreachable host
answers 502 with the remote reason instead of a misleading 404. Nothing is ever
copied to the Codeman host and there is NO local fallback (an sshfs mount of
the same tree must not shadow the remote bytes).
Deliberately unchanged: writes (edit=1 / PUT now answer 400 explicitly while
the viewer hides its Edit affordance), office previews, thumbnails, file tree,
picker, external attachment registration and tail-file stay local-only.
CI caught a real regression: DEEPSEEK_API_KEY was added to deepseek's
privilegedEnvKeys alongside DEEPSEEK_BASE_URL on the theory that "the pair
travels together," but that contradicts the documented and tested design
(clampEnvOverridesForOwner()'s own docstring in session-routes.ts) — a
non-granted owner supplying their OWN DeepSeek key removes privilege
rather than granting it, since the exfiltration vector is the BASE URL
(which redirects the server's own forwarded key to a foreign host), not
the key itself. Removed it from the list; test/deepseek-mode.test.ts's
existing two clamp tests now pass again.
Also swapped that test's "unrelated override" example off CODEX_HOME,
which the earlier commit in this same PR legitimately made privileged
(closing a real pre-existing gap, documented in PR.md) — so it stopped
being a valid "unrelated" example the moment that fix landed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.
Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.
Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:
- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
undocumented GATEWAY AuthType gemini-cli selects once
GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
tried
- deepseek: reaches the server (env vars are read) but gets a consistent
HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)
Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).
deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its
native cloud backend, for a given session. Covers local hardware (llama.cpp,
Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter).
Off by default (customModelEndpointsEnabled, synced, default OFF).
- Registry: capabilities.customModelInjection per CLI entry (env /
configContentEnv / configDir / unsupported kinds)
- Pure injection builder (custom-model-injection.ts) turning an endpoint +
model id into the real env vars / config content per CLI
- Endpoint store + CRUD routes (custom-model-hosts.ts,
custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded
- Session integration: Session.setCustomModel()/restartCli()
(POST /api/sessions/:id/custom-model), reusing the existing
respawn-pane -k primitive to restart the CLI process with new env
- Multi-user hardening: every new redirect-capable env var added to its
CLI's privilegedEnvKeys, closing a pre-existing gap where several were
already reachable via the generic envOverrides field's prefix allowlist
- Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries
against a real endpoint outside the web UI, independent of tmux/sessions
- Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying
every CLI's injected values through a real HTTP shape
Real end-to-end validation against a live llama-swap server (inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries) found and
fixed three real bugs before they shipped:
- Codex's config.toml schema was wrong ([model].default table instead of
a top-level model string + [model_providers.custom]); fixing it then
surfaced a genuine, documented protocol incompatibility (Codex only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement)
- Claude Code's async session-title-generation call validates
ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and
hangs the whole -p invocation on an unrecognized name; documented for
chunk 6, worked around in the standalone script only (--bare is NOT
safe for a real interactive session, which needs hooks)
- The discovery route's authStyle: 'both' option (send both Authorization
and api-key headers) reliably hung a real server; removed the option
entirely rather than just changing the default
Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see
PR.md and deployment_plan.md for the full chunk breakdown and confidence
table.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
/opt/codeman-cli is chowned to PUID:PGID once, at image build time, from
the PUID/PGID build args. That bake only happens when the image is
actually rebuilt (`docker compose up --build`, which Start-Codeman.sh
always does) — a deployment that runs the compose file directly instead
(Unraid's Compose Manager, a native systemd unit, any plain
`docker compose up`/`restart`) can change PUID/PGID in .env and restart
without ever rebuilding. The container then runs as the NEW uid via
entrypoint's setpriv (Linux needs no /etc/passwd entry to setuid to an
arbitrary number) while the CLI directory is still owned by the OLD one
baked into the image layer — silently breaking the self-update-a-CLI-
in-place fix that directory exists for.
Unlike HOME/CODEMAN_CASES_PATH, this one is pure image content Codeman
itself populated, never host data that might legitimately belong to
someone else, so there is no ownership to be careful about — it is
always correct for it to be owned by whoever the container is about to
run as. Re-assert it unconditionally on every start.
Verified live: built an image with PUID=99/PGID=100, ran it with
PUID=1234/PGID=4321 (no rebuild, simulating a changed .env restarted
directly), confirmed /opt/codeman-cli ends up 1234:4321-owned and is
genuinely writable by the running process.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
Two real bugs the review caught, both verified live against a real
build on the Unraid host:
1. entrypoint.sh's chown fired on ANY ownership mismatch, not just a
directory the daemon itself created root-owned. A host tree
legitimately owned by some other account - an existing
CODEMAN_CASES_PATH the README already allows pointing at a normal
projects directory, or appdata under a different PUID/PGID
convention than the one in use - got silently recursively re-owned
with one log line to explain it. Now gated on the target actually
being root-owned; anything else is a clean refusal naming the
directory, its owner, and PUID/PGID. Start-Codeman.sh also now
pre-creates CODEMAN_CASES_PATH the same way it already did
CODEMAN_APPDATA_PATH, so Compose never has to materialise a missing
bind source as root in the first place - the in-container chown
becomes a safety net, not the primary mechanism.
2. The CLI-update chown (chown -R .../node_modules /usr/local/bin)
handed the runtime account write access to entrypoint.sh itself
(root-owned, executed as root on every container start with
CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node binary - owning the
DIRECTORY is enough to rename it aside and drop a replacement, which
would let a compromised session arrange for its own script to run
as root at the next restart. The four CLIs now install into a
dedicated /opt/codeman-cli prefix (NPM_CONFIG_PREFIX); only that
directory is chowned, /usr/local stays root-owned throughout.
Smaller fixes from the same review:
- Start-Codeman.sh's volume-refresh label filter wasn't project-scoped:
a second Compose stack on the same host sharing the `codeman-dist`
volume KEY could have had ITS volume deleted. Added a
com.docker.compose.project filter, resolved from this stack's own
`compose config --format json`.
- Override-file precedence was backwards (checked .yaml before .yml;
Compose actually prefers .yml) - swapped, plus a warning when both
exist.
- entrypoint.sh's setpriv now also passes --bounding-set -all, so
CapBnd actually clears post-drop rather than just CapPrm/CapEff.
- A comment on git_head_commit() noting it returns nothing for a
worktree checkout (.git as a file), consistent with the script's
existing -d .git convention elsewhere.
- Doc drift: CLAUDE.md's Docker Compose section still described the
old pre-created-and-chowned-by-hand model and didn't mention the
root-then-drop entrypoint; the state-files list was missing
docker-build-source.json; docs/docker-compose.md and
docker/.env.example still had the pre-rename `Coding/codeman` path
in one place each.
Verified end to end against a real build on the Unraid host: a
root-owned bind source is corrected as before; a directory owned by
neither root nor PUID:PGID is refused rather than silently rewritten;
a correctly-owned directory is left alone entirely; the four CLIs
resolve via PATH from /opt/codeman-cli while /usr/local/bin,
/usr/local/lib/node_modules and entrypoint.sh itself stay root-owned;
CapBnd is fully cleared post-drop.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded
from the image only while empty, so a rebuilt image's fresh dist/
node_modules sat unused behind old volume content until something
cleared it. The in-app self-updater never hit this (it rebuilds INSIDE
the running container, into the very volume already in use), but a
`docker compose build` triggered from outside it — Start-Codeman.sh,
after a manual `git pull` — did: the container came back up looking
unchanged, serving stale compiled routes against current source.
Start-Codeman.sh now compares the checkout's HEAD commit and
package-lock.json hash against a recorded marker
(docker-build-source.json) and clears just the affected volume(s)
before its own --build when either moved.
The in-place self-update path writes that same marker after a
successful build, so the two mechanisms agree on what the volumes
currently reflect — without it, the next plain Start-Codeman.sh run
would see the HEAD self-update just checked out, not recognise it as
already accounted for, and wipe the volumes self-update just correctly
rebuilt right back to the older baked image.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
The four CLIs (claude, gemini, codex, opencode) are npm-installed
globally as root during the image build, before the unprivileged
runtime account exists. A session running as that account (e.g. a
codex-mode terminal) then hits EACCES the moment it tries to update
one in place, because npm renames the old package directory aside
before installing the new one, which needs write access to the
parent (/usr/local/lib/node_modules), not just the target package.
Chown that tree plus /usr/local/bin's CLI symlinks to PUID:PGID in
the same step that creates/renames the runtime account.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
CODEMAN_ALLOWED_HOSTS is a real, documented application setting (the Host-
header allowlist in network-auth-policy.ts), but docker-compose.yaml does not
forward it from .env into the container - Compose only passes through
variables explicitly listed under environment:, and this is not one of them.
Set without that passthrough, any request through a reverse proxy is rejected
with 403 Forbidden: host not allowed before it reaches any handler, and
nothing in the Docker deployment docs said why.
Document the variable and the override needed to forward it, using the
Local customisation mechanism already described above it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
CODEMAN_RUNTIME_USER defaulted to `opencode`, which no longer matches the
project and is confusing in a deployment whose every other identifier is
codeman. Rename the default in .env.example and in the Dockerfile ARG that
mirrors it, and correct the example comment that referred to
/home/opencode/codeman-cases.
Also drop the `Coding/` component from the example application-data path.
CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH now suggest /mnt/user/appdata/codeman
and its codeman-cases child, matching the account name and removing a directory
level that meant nothing outside the original author's host. README.md is
updated to match, including the chown example.
The npm package `opencode-ai` and the references to the OpenCode CLI are
deliberately left alone: those name a different tool, not this account.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Naming a Compose file with -f disables Compose's automatic discovery of the
override file, so Start-Codeman.sh silently ignored docker-compose.override.yml.
Any local customisation placed in the conventional override file was dropped
without warning, and the only way to notice was to inspect the running
container.
Collect the -f arguments into an array, append the override file when one is
present, and reuse that array for the final launch so the two cannot drift
apart again. Both .yml and .yaml are checked, in Compose's own precedence
order, and the chosen file is reported on startup.
Document the override file in docker/README.md, including the two things that
are easy to get wrong: it is ignored when -f is passed without naming it, and
it cannot remove a key such as ports, which Compose concatenates. Add the
override file to .gitignore.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Compose binds CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH from the host. When
either path does not exist yet - a first run, a cleared application-data
directory, a restored backup - the Docker daemon creates it owned by root. The
server runs unprivileged as CODEMAN_RUNTIME_USER, so it cannot create its own
state directory, and the container restarts forever on:
Failed to start web server: EACCES: permission denied, mkdir '/home/<user>/.codeman'
Start-Codeman.sh already worked around this by preparing the directory on the
host, so the failure only appears when Compose is run directly, which the README
documents as a supported path.
Add docker/entrypoint.sh, which starts as root, corrects the ownership of both
bind mounts, then drops to PUID:PGID with setpriv. The Dockerfile's USER
instruction is replaced by that entrypoint and CMD is unchanged.
docker-compose.yaml adds back only the four capabilities the chown and the
privilege drop require, so cap_drop: ALL continues to remove everything else.
Two guards keep existing deployments working:
- A container started with an explicit `user:` is left alone. The entrypoint
execs straight through, with no elevation and no chown.
- A chown that fails is a warning, not an error. Bind mounts backed by NFS,
CIFS or a rootless daemon can refuse chown while remaining perfectly
writable, and those deployments must keep starting.
PUID and PGID are also exported as runtime environment defaults so the image
behaves correctly when run without Compose, rather than depending on build args
alone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A dashboard served through a web tab saw `/webview/<cap>/` as its
`location.pathname`, and no app has a route for that: a React Router, Vue
Router or Vite dev-server page painted its HTML and CSS and then replaced
them with its own "page not found" the moment its script ran (reproduced
with a minimal history-routed page).
The proxy's runtime shim now rewrites the history entry to the path the
page would see on its own origin, before any page script runs. The base
element still resolves relative URLs inside the prefix and every root-
absolute sink is rewritten back into it, so only what the page READS
changes. With the document URL masked the Referer-keyed 404 rescue can no
longer help a request the shim misses, so the remaining URL-taking entry
points (`Worker`, `SharedWorker`, `navigator.sendBeacon`, `window.open`)
are covered by the shim as well.
A navigation the page starts itself afterwards — `location.reload()`
(a dev server's full-reload HMR), a root-absolute `location.href` — lands
on Codeman's root with no capability anywhere: no prefix in the path, no
cookie in an opaque-origin frame, a Referer naming the masked page. It is
recognised by shape (a top-level iframe navigation asking for HTML, for a
path Codeman does not serve) and answered with a static page whose only
script posts `{type:'codeman:webview-lost', path}` to the parent; the tab
that owns the frame (matched by `event.source`, never by the payload)
remounts it inside the prefix at that path, bounded per frame. The
unauthenticated form is answered in the auth middleware before the
credential checks, so a dev server that reloads on every save cannot
rate-limit its own user out of Codeman; the authenticated form (Basic
auth, trusted mode) is answered by the 404 handler.
Verified end to end against a history-routed page: boots on `/`, its
API call succeeds, a reload inside the frame comes back routed on the
path it had pushed, `location.href = '/about'` comes back on `/about`,
and a deep link opens on its path.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
fix(sessions): stop pinning the `w1-myapp` placeholder as Claude's session title. Local Claude spawns passed the tab name as `--name`, which is also the `/resume` picker entry and the terminal title, and a pinned title stops Claude generating its own, so every conversation of a case showed up in `/resume` as the same `w1-myapp` and none got a generated title. Only a name the user chose is pinned now; placeholder and auto-named tabs let Claude title the conversation again. Renaming a Claude tab also reaches `/resume`: the new name is appended to the conversation's transcript as the `custom-title` row `/rename` writes (a tab that was spawned with `--name` keeps re-appending its own title until its next respawn, so the rename wins from then on). Orchestrators that rely on a fixed peer name should give workers a descriptive `sessionName` rather than a `w<N>-` one.
Fix an invisible terminal text selection on the light skins (#360). Every xterm palette declared its selection colour under the key `selection`, which xterm.js renamed to `selectionBackground` in v5. An `ITheme` is a plain object, so the unknown key was dropped without an error and every skin fell back to xterm's own default of `rgba(255,255,255,0.3)`: unnoticeable on the dark skins, which wanted roughly that anyway, and effectively invisible on Paper Gray, Solarized Light, Catppuccin Latte and Rosé Pine Dawn, where white at 30% over a near-white background moves a channel by about 3/255. Selecting text on those skins now highlights it, with desktop drag-select and the mobile long-press both fixed by the same rename.
"description":"Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
- 13e652e: Terminal copy: copying text out of a Claude Code or Codex pane no longer puts the pane's two-column transcript gutter on the clipboard, so pasted lines arrive flush instead of indented (#469). The width comes from the CLI registry (`capabilities.transcriptGutter`, 2 for claude and codex, measured on live panes) and is only a ceiling: a selection only ever shifts as a block, so its own indentation survives. Other CLIs and shells are untouched. It works in split panes and detached session windows too, and can be turned off per device in App Settings under Selection & clipboard.
- 13e652e: Sessions: recovering a Claude session whose tmux pane had died relaunched `claude --session-id <id>`, which Claude refuses once that id has a transcript, so the pane died again straight away and the conversation was stranded. The relaunch now resumes the conversation (`--resume <id> || --session-id <id>`), including when tmux lost the whole session (#467).
- 00f022c: Terminal: when a burst of output overflows the render queue and a frame has to be dropped, the repaint that repairs it is now retried until it actually happens, instead of being scheduled once and silently skipped when another load was in flight (#470).
- 13e652e: Mobile: a long press on blank terminal space on Android Chrome no longer opens the keyboard and blanks the terminal (#471, fixes #360). The long-press guards are now armed before the press is checked for selectable text, so a press on empty space is swallowed the same way a press on a word already was.
- 13e652e: Sessions: a tab whose agent has exited (the CLI quit, but tmux kept the pane) now says so with a muted dot and an `exited (137)` badge, instead of looking like an idle session (#466, part 1 of #446). The state is published as `paneExit` on the session and survives a restart. Nothing closes such sessions yet; that is part 2.
- 13e652e: Docker: optional GitHub CLI and Azure CLI for private repositories (#472). Both are off by default. With `CODEMAN_INSTALL_GH=1` / `CODEMAN_INSTALL_AZ=1` as build args in `docker-compose.override.yml`, the server image gets `gh` and/or `az` (with the `azure-devops` extension) wired in as git credential helpers, so after one `gh auth login` or `az login` from a shell session, Add Case → Clone Repo can clone private GitHub and Azure DevOps repositories. `CODEMAN_AGENT_IMAGE_INSTALL_GH` / `_AZ` do the same for the Docker-case agent image, and only then are the sign-ins copied into new case containers. In multi-user mode a non-admin's clone runs with the credential helpers cleared. This changes `server.Dockerfile`, so Compose deployments need a `Start-Codeman.sh` rebuild rather than an in-app update.
- 13e652e: Run menu: the Gemini, Antigravity and OMP run buttons now show their own colours on every skin; they rendered in Claude blue on all skins except OG (#463). The CLI registry's `accent` values were also corrected to the colours the UI really paints, and a test now guards the stylesheet trap that caused it.
- 13e652e: Terminal: five ways the browser terminal could silently stop being correct are fixed (#431, #464). The browser terminal and the PTY can no longer disagree about their width, which is what produced doubled lines and half-overwritten text ("text gets muffled sometimes"): there is now one function that sizes the terminal, and every resize is answered with the geometry the PTY really holds. A replay clear goes through the terminal's own queue, so bytes written just before it no longer fuse into the next snapshot. A renderer that stops painting after an iOS PWA is backgrounded heals itself instead of needing a reload. Every terminal capture has a deadline that also covers the response body, and a capture that runs out of time during a tab switch falls back to the bounded tail instead of leaving a blank pane. Output lost to a half-open WebSocket is repainted on the next successful open. The service worker's precache list is now generated by the build and its cache is rotated per build, so old releases' assets no longer pile up.
- 13e652e: Docker: new `docker/Update-Codeman.sh` for the major-update path the docs used to describe by hand (#465). It rebuilds the image with `--no-cache` before taking the stack down, clears the build-artefact volumes, refuses to run when another checkout's Compose project already owns the same name, and then hands over to `Start-Codeman.sh`.
- 13e652e: Approvals: a session that is idle only because it is waiting on its own background work (Claude Code's `1 monitor` footer chip, or a Codex background terminal) no longer raises the yellow NEEDS YOU alert or a push (#473, fixes #468). Its idle item is opened already acknowledged, and the tab, the home screens and the rail show a small `watching` badge next to the state instead. The item still exists in the Approvals Inbox, and the TUI's pending count now leaves acknowledged items out.
- b404dac: Maintainer fixes applied while landing this batch:
- Terminal (#431): while another device holds the pane's width, a resize retry no longer re-fits xterm to the container and re-wraps the whole buffer every 30 s, and no longer clears scrollback for a redraw that never comes. The PTY's spawn geometry is now recorded at attach, so `ptyGeometry` never reports a size the PTY never held.
- Terminal (#470): the `TERMINAL DROP` crash-trail line is logged once per recovery window instead of once per dropped frame (which wiped the rest of the trail within a second), and a refresh that died at its fetch deadline is no longer retried.
- Sessions (#467): the resume pin also covers the branch where tmux lost the whole session, the conversation id Codeman reports follows what the relaunch actually resumed, and the test setup strips `CLAUDE_CONFIG_DIR` so the suite stays green for anyone running a separate Claude config dir.
- Sessions (#466): detailed sidebar and rail rows show an `exited` pill instead of `idle`, the exit is announced to screen readers, and the user manual's tab-appearance table lists the new state.
- Approvals (#473): a failed pane capture clears the `watching` badge rather than keeping a stale one (a failure now falls toward an alert, not toward silence), and the header bell's count leaves acknowledged items out, matching the TUI.
- Run menu (#463): the Gemini and Antigravity run buttons no longer render two-tone on phones, Gemini's registry accent matches its tab badge, and a test now guards the stylesheet trap for every run mode.
- Docker (#465): `Update-Codeman.sh` removes exactly the two build-artefact volumes it names instead of every named volume in the project, reports a failing `docker compose` instead of exiting silently, and its docs and comments were corrected. (#472): the multi-user notes say that a non-admin's seeded Docker case also receives the gh/az sign-in when those switches are on.
### Thanks
- @irisitymichaelgrundberg for four PRs in this release: the `watching` badge that stops background work from raising false alerts (#473, from their own report #468), the exited-agent badge (#466) and the dead-pane resume fix (#467), both from their report #446, and the transcript-gutter strip for copied text (#469), a follow-up to their #451.
- @rounakdatta for the terminal resilience work (#431) and the dropped-frame recovery (#470), both from their report #464, and for answering four rounds of review in full.
- @opticon454 for private-repository support in the Docker images (#472), the `Update-Codeman.sh` script (#465) and the run-button colour fix (#463).
- @DodgyBadger for the Android long-press fix (#471), from their own report #360.
## 1.32.0
### Minor Changes
- d47f93a: feat(custom-model): the model picker puts the ready model first
When a custom endpoint has more than one model, the Run menu's picker now promotes one row to the top instead of showing raw discovery order: the model llama-swap reports loaded and ready right now (tagged "Currently loaded", the one a launch attaches to with zero wait), else the model you last launched on that harness and endpoint (tagged "Last used", remembered per device). The endpoint's default keeps its own pill, nothing is ever auto-chosen, and a plain OpenAI-compatible server or an endpoint that does not answer within a second simply keeps the old order. The probe is bounded on the client too, so a GPU box that is off no longer holds the picker closed for five seconds.
- d47f93a: feat(split-pane): view two live sessions side by side
A new Split button in the header (opt-in in App Settings, off by default, desktop only at 1180px and wider) opens a picker and shows a second live session beside the active one: its own terminal, its own WebSocket, and a divider you can drag. When either session ends the view collapses back to one pane, with Pane B promoted to the primary when it is Pane A that ended. Nothing is persisted on purpose in this first cut, so a page reload always returns to a single pane. Pane B is deliberately plainer than the primary pane (no local-echo overlay, CJK input, touch handling or keyboard accessory bar); the design and the v2 boundaries are in discussion #452.
- 72d437a: Installer v2. `curl -fsSL https://getcodeman.com/install | bash` now looks at the machine first, asks at most three questions up front (how the dashboard is reached, optionally what to call the machine on your tailnet, whether to run Codeman as a background service), does the install unattended behind progress spinners with the output in `~/.codeman/install.log`, and ends on the URL with a QR code to scan. One consent covers every missing package and sudo asks for your password once. Flags pipe through `bash -s --` (`--tailscale | --lan | --local`, `--name <n> | --no-rename`, `--service | --run | --no-start`, `--yes`, `--password`, `--port`), `install.sh status` prints the URL and the QR code again, and the cloudflared question moved out of the main flow into `install.sh cloudflared`. On the Tailscale route, a `:443` that already belongs to another app gets Codeman under `https://<node>/codeman` (or on a second port) instead of a dead end, the node can be renamed opt-in (`--name`, `install.sh name`, undone by uninstall), and the HTTPS-certificates toggle is polled with the admin page opened for you. Also fixed on the way: the installer's own `npm install` no longer lets the postinstall start a stray server on port 3000 (the service crash-looped on EADDRINUSE while the done screen said "running"), the LAN address comes from the default route rather than the first interface, a hand-written LaunchDaemon on a headless Mac is left alone, a flag re-run keeps an existing dashboard password, and the done screen's start command carries the sub-path and port it was installed with.
- d47f93a: feat(mobile): a Compose key for writing prompts on a phone
The agent keyboard bars on phones replace their Paste key with Compose: a real multiline editor with autocorrect and spellcheck, per-session drafts kept in memory only, image attach that never writes into the terminal early, and a Send that delivers the text as one paste followed by Enter, so a long prompt no longer has to be typed blind into the terminal composer. Anything you had already typed into the terminal is picked up into the editor. Shell sessions keep the direct Paste key. This is the manual first slice from #359; the auto-open setting and terminal tap routing are a separate follow-up.
### Patch Changes
- d47f93a: refactor(run-menu): one table-driven launcher for every external CLI
The eight near-identical per-CLI launch functions in the Run menu collapsed into one launcher driven by a table that a CI test keeps in step with the CLI registry, and a second no-id-branching guard now covers the frontend the way the backend guard covers the server. No behaviour change: the refactor was verified byte-identical across 288 launch permutations against the previous code.
- e899af4: Maintainer fixes applied while landing the above. The model picker's promoted row keeps its Default pill (the promotion tag and the default marker are two pills now, and they render as pills in the picker rather than as plain text). The phone composer keeps its bottom gutter on folding devices (the generic fold rule used to erase it), a whitespace-only draft is no longer sent, and its dialog is translated on a zh-CN UI. A split that collapses mid-drag no longer leaves the page stuck in resize-cursor mode, Pane B refuses a session that has no live process, and a burst of refresh frames replays once instead of twice. The `</head>` script injections on the page render use replacer functions, so a CLI label containing `$'` can no longer splice the document into the inline script, and the frontend no-id-branching guard now catches comparisons on any variable name.
- 6ef71ec: ### Thanks
- @timkjr for split-pane sessions (#453): five review rounds turned around in two days, and the pointer-capture edge case measured in a real browser rather than reasoned about.
- @DodgyBadger for the mobile prompt composer (#444), a first contribution that took the scope back down to one slice when asked, and that verified the delivery path against a live tmux pane and a live Claude Code composer instead of trusting the diff.
- @opticon454 for putting the ready model first in the picker (#459) and for collapsing the eight Run-menu launch functions into one (#458), proven byte-identical across 288 launch permutations instead of argued.
## 1.31.0
### Minor Changes
- 035bfbc: feat(remote): wake a sleeping remote host from Codeman
A remote SSH case pointing at a machine that suspends used to fail the same way every
time: the session was there, the host was not, and typing into it went nowhere. A host
can now carry a wake target, either a MAC address for Wake-on-LAN (Codeman builds the
magic packet itself, so nothing reaches a shell) or a wake command of your own, and
Codeman uses it when you ask for the host: when you type into a sleeping session, when
you press the wake button on the banner, or when you start or attach a session on that
host. Input you type while it wakes is buffered and flushed once it is back, up to 4 KB,
and a chunk over that is refused outright rather than delivered as a fragment.
Waking only ever happens because you asked. No watcher, dropped-session handler or
boot-recovery path can reach it, since a machine woken by a reconnect watcher would come
back seconds after every suspend.
- fbee1b2: feat(custom-model): pick a custom endpoint straight from the Run menu
#393 landed the backend for custom model endpoints and left it reachable only over the
HTTP API. This is the rest of it. Turn on Custom model endpoints in App Settings, save
an endpoint, and the Run dropdown grows a Custom Endpoints section built live off the
CLI registry, one entry per harness that can actually redirect plus each endpoint you
saved. Pick one and it launches that harness pointed at your server, asking which model
first when the endpoint has more than one. Endpoints re-discover themselves every five
minutes, and one unreachable endpoint never blocks the others. App Settings gains full
add, edit and delete for endpoints.
Seven of the harnesses (opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP) now launch
directly onto the endpoint with no restart at all, where before you watched a native
boot followed immediately by a second one. Claude still launches and then restarts in
place, which its own resume makes far less jarring.
Most of this release's work went into things that only show up against a real server,
and each was found that way rather than in tests: a freshly launched CLI reporting
itself busy for its own startup and getting refused; Claude Code assuming a large
context window for a model it does not recognise and silently overflowing a small one;
a model whose real context is below what Claude Code's own system prompt costs, which
no setting can fix and which now warns before launching into a certain failure; and the
big one, llama.cpp running exactly one model at a time, so applying a selection can
unload the model another session is using. That last case now asks first, tells you
which session it affects, and keeps a "loading model" notice on screen for the whole
swap window, so a prompt sent mid-swap reads as loading rather than as an answer from
whatever was loaded a moment ago. A background sweep also catches the reverse: your
session's model being evicted later by somebody else's ordinary use.
Two things worth knowing if you drive this over the HTTP API or run multi-user. The two
questions an apply can ask (the model's context window is too small, and loading it will
unload the model another session is using) are now answered by separate
`confirmedContext` and `confirmedSwap` fields rather than one `confirmed`. They shared a
flag until now, and since the context check runs first, confirming that one silently
agreed to evict another session's model as well. The old `confirmed` still means both.
And `CLAUDE_CONFIG_DIR` is now admin-only in multi-user mode: it joined claude's
privileged env keys, so a non-granted owner can no longer set it through `envOverrides`,
and an already-persisted one is dropped on reboot-restore, which returns that session to
the default Claude account rather than the per-client one it was pointed at. Single-user
installs are unaffected.
Remote SSH and Docker sessions are refused for now, since their restart reattaches a
durable tmux rather than relaunching the agent.
### Patch Changes
- c9515b1: fix(terminal): keep the output a pane capture could not contain. Opening a session, a backpressure refresh, a clear-terminal reload and a full-history re-pull all load the screen from a tmux pane capture, and anything the CLI printed between that capture and the end of the load used to be dropped, so its next partial redraw landed on a frame the terminal had never seen: missing or garbled output right after a tab switch or a refresh, plainest in a shell session. Each load now replays exactly the output that arrived after the capture, through one shared rule for all four paths, and a refresh that restores your scroll position no longer snaps back to the bottom afterwards.
- 3edf9aa: fix(terminal): replay a pane capture at the geometry it was taken at
Opening a session could draw a frame built for a pane bigger than your terminal. A
taller pane wrote its overflow rows onto the last line and lost the rows underneath
(against a 50-row pane, a 30-row terminal rendered 28 of a 45-line command and drew
the survivors twice), and a wider one wrapped every row and scrolled the whole frame
up by one. The terminal response now reports the geometry the capture was really
taken at, so the browser can see the mismatch and replay once at the size that stuck.
A pane that cannot be sized to fit is diagnosed once per session instead of on every
tab switch.
- 035bfbc: ### Thanks
- @irisitymichaelgrundberg for three terminal fixes in one release: keeping the output a pane capture could not contain (#436), replaying a capture at the geometry it was taken at (#435, five rounds and a Playwright suite that fails against the merge base), and trimming the padding out of a copied selection (#451), where the scan-instead-of-regex call avoided a 2.9s freeze nobody would have traced back to a copy.
- @timkjr for a first contribution that found a real silent failure: the Instance count stepper next to the Run button had only ever applied to Claude, so on the other eight run modes it launched one session and said nothing (#454).
- @Randalix for Wake-on-LAN on remote hosts (#439), built and live-tested against a real sleeping machine, and for reading the whole diff again between rounds rather than only the parts that were asked about.
- @opticon454 for turning #393's backend-only custom model endpoints into the whole feature (#430), and for validating it against a real llama-swap box rather than against the tests: the `/props` versus `/running` context discrepancy and the DeepSeek `/v1` root cause were both tracked down to the SDK source instead of guessed at.
- c376534: fix(run): make the Instance count stepper work for every non-Claude mode
The Instance count stepper next to the Run button only ever applied to Claude.
Setting it to 3 and launching OpenCode, Codex, Gemini, Antigravity, Pi, OMP, Grok or
DeepSeek started exactly one session, with no error and no hint that the control had
done nothing. All eight now launch the count you asked for, and the opening banner
says how many are starting. The one exception is a launch started from the Custom
Endpoints section of the Run menu, which always starts a single session.
- 19ffe9b: fix(input): make sure a prompt sent through the API actually leaves the composer. Claude Code 2.1.277 started ignoring Enter for the first 30 to 50 seconds after the composer paints while still accepting the typed text, so a prompt sent right after a session came up sat unsent in the pane and every waiter (send-and-wait, the agent skill, cron, the maintainer bot) burned its whole timeout on a turn that never started. The server now reads the pane after every programmatic write that carried Enter and presses Enter again, on a 2 to 60 second schedule, only while the composer verifiably still holds the text it sent; an empty composer, other text, or a pane with no composer at all ends it. The agent skill's `sendwait` gets the same loop for servers that predate this, and its preamble version moves to 1.30.1 so an already-seeded agent picks up the fresh copy.
- f9edb33: fix(terminal): trim the padding out of a copied selection
Copying out of a pane put a wall of spaces on the clipboard. xterm hands back
whole screen rows and trims only the cells that were never written to, so the
real spaces a full-screen program paints across the unused part of a row count
as content: measured against Claude Code in a 282-column pane, single lines
arrived carrying 138 trailing spaces. Pasting that into a chat client or an
editor meant deleting the whitespace by hand, while Windows Terminal, iTerm2 and
GNOME Terminal all trim it for you. A copy now drops the trailing run from every
line, on all four paths (the Ctrl+C chord, right-click, the phone selection
button and Auto Copy), while leading indentation is left exactly as it is. An
Alt+drag rectangular selection is copied verbatim, because its columns lining up
is the point of that gesture. A selection holding nothing but padding is refused
rather than copied as bare line breaks.
## 1.30.0
### Minor Changes
- da933d7: Offer to rebuild the sessions a host reboot destroyed. A reboot takes the tmux server down with it, so every pane dies and the board comes up empty. Codeman now works out what was running, and the board offers to restore it behind a click. The conversations come back; the terminal scrollback does not, and the banner says so.
### Patch Changes
- a1c35da: Stop a phone keyboard losing the last character of every message it sends. Android soft keyboards commit the last typed character and send the Enter key in one InputConnection transaction, so the `input` event and the Enter keydown are both processed before any zero-delay timer runs. The orphaned-input recovery from #388 only resolved its candidate on such a timer, and lost it both ways: xterm emits `\r` synchronously from the Enter keydown, so the local-echo composer submitted the prompt before the recovered character existed, and that `\r` bumped the "did xterm speak for this keystroke" counter, so the candidate then stood itself down and dropped the character outright. Pending candidates are now drained synchronously at the next keydown, from xterm's custom key handler, which runs before xterm processes that key, so the counter still holds the value it had while the candidate's own keystroke was current, and the recovered byte reaches the composer ahead of the Enter. Typing on a physical keyboard is unaffected: there, the timer has already resolved the candidate before the next key arrives.
- 3f2928a: The installer's hint for a launcher-only CLI (DeepSeek today) now says why it is a docs link rather than a command you can run, and points at the thing that resolves it: the package installs a launcher that still needs a terminal profile, and Codeman's Run menu can add one in a click. Driven by a generated `CLI_LAUNCHER_ONLY` flag rather than an id check, so it covers any future entry of that shape. Also removes three dead lookup helpers and two never-read generated arrays from `install.sh`, skips a disabled entry's probe instead of filtering it afterwards, and corrects a comment that claimed the non-interactive default is always Claude Code (on a wget-only host its curl one-liner is filtered out first).
- 0e1191b: Maintainer fixes applied while landing the above. A session restored after a reboot keeps the name you gave it (the rebuild dropped the field that records who named a session, so a hand-renamed session came back looking auto-named and the next prompt overwrote it), and no longer types `continue` into itself on its own: a pending auto-resume stamp from before the reboot is dropped rather than re-armed, since the pane is new and one click could otherwise arm several unattended prompts at once. Auto-resume itself stays on and re-arms on the next real usage-limit message. The restore offer is also hidden in a detached single-session window, which has no tab strip to put restored sessions in, and a conversation that goes live while an earlier session in the same batch is starting is no longer restored a second time.
- 0e1191b: ### Thanks
- @irisitymichaelgrundberg for the reboot-restore banner (#442), and for the three real reboots behind it rather than a mocked one.
- @shenlvkang-collab for tracking down why Android keyboards lost the last character of every message (#441), including the half where the character was not late but gone.
- @opticon454 for going back and closing out the loose ends left as "worth knowing rather than fixing" after #380 (#429).
- de864e7: Keep the terminal anchored where you are reading while an agent streams (#358). Scrolling up during a Codex response could still be dragged back to the live bottom by the next redraw: the flush captured the viewport before writing and restored it immediately after, but xterm parses asynchronously, so at that moment the buffer had not moved yet, the restore compared the anchor against itself and did nothing, and the redraw landed a tick later with nothing left to pull the view back. The restore now runs inside xterm's own write callback, which is the first point at which the redraw's effect exists, and it holds across consecutive and chunked redraws. It is dropped if you switch sessions or a history replay starts before the write parses, since the anchor indexes the buffer it was captured from.
## 1.29.1
### Patch Changes
- 5b920cb: Auto-name sessions from the first prompt (#376, opt-in). With the new synced **Auto-name Sessions** setting on (App Settings → Appearance → Tabs, default off), a tab that still carries its generated name takes a title from the first real prompt you submit, keeping the case prefix: `w3-myapp` becomes `w3-myapp: fix the login redirect`. The strip shows the title with the prefix in the tooltip, and the next session in that case still counts up. It happens once per session, only for prompts you type or send through the input API (never a Ralph, respawn, cron or approval answer), never for shells, and a name you set yourself is never touched. Slash commands such as `/clear` do not become titles. The title is derived locally from the prompt's first sentence; no text leaves the machine. `nameSource` (`placeholder` / `auto` / `manual`) is a new additive field on session state.
Landed with the fixes the review of #376 asked for: first prompt only (not every prompt), a user-input gate so Ralph, respawn, cron and approval writes cannot name a tab, the prefix form so the case identity and `w<n>` counter survive, and a keystroke tracker that handles a bare Esc, bracketed pastes, wheel reports, Tab and history recall instead of mis-titling the tab.
### Thanks
- @shenlvkang-collab for #376, the auto-naming idea and the ownership plumbing (`nameSource`, the listener wiring, the restore path) it shipped with.
## 1.29.0
### Minor Changes
- **Custom model endpoints, HTTP API first** (#393). Any run mode that has a mechanism for it can be pointed at a custom OpenAI-compatible endpoint (a local llama.cpp, llama-swap, Ollama or vLLM, or a cloud gateway) instead of its native backend, per session. Endpoints are stored in `~/.codeman/custom-model-hosts.json` (`GET/POST/PUT/DELETE /api/model-endpoints`, admin-only in multi-user mode), their model lists are discovered from the endpoint's own `/v1/models`, and `POST /api/sessions/:id/custom-model` applies one to a session by restarting its CLI in place. The mechanism is per-CLI registry data (`capabilities.customModelInjection`): env vars for Claude, Gemini, Grok and DeepSeek, `OPENCODE_CONFIG_CONTENT` for opencode, an isolated config dir for Codex, Pi and OMP, unsupported for Antigravity. Verified live against a llama-swap server for claude, opencode, pi, grok and omp; gemini and deepseek reach the server and fail for reasons not yet understood, and codex only speaks the Responses API, so a plain chat-completions server cannot serve it. Those three are documented as gaps rather than shipped as working. The toolbar picker is a follow-up; until it lands the feature is HTTP-API only (`docs/custom-model-endpoints.md`), and the `customModelEndpointsEnabled` setting is declared but read by nothing yet. Merged with maintainer follow-ups: clearing a selection now actually clears it (the injected vars are delivered by `tmux setenv`, which `respawn-pane` inherits, so the relaunched CLI came back still pointed at the endpoint; retired keys are now `setenv -u`'d before the respawn), applying a model to a local claude session no longer kills the pane (the relaunch pins `--resume <id>` with the `--session-id` fallback, since Claude Code refuses a session id that already has a transcript), pi, omp and grok now select the generated model through a registry-declared `launchModel` (`custom/<id>`, `-m codeman-custom`) instead of writing a config the CLI then ignored, remote and Docker sessions are refused with a clear 400 until those paths are plumbed, the selection survives a Codeman restart, discovery goes through the egress-guarded `webviewFetch()`, key-bearing files are written 0600 and the per-session config dir is removed with the session, and the design plan moved from the repo root to `docs/custom-model-endpoints-plan.md`. Along the way the multi-user clamp learned about `GOOGLE_GEMINI_BASE_URL`, `GROK_BASE_URL`, `CODEX_HOME`, `PI_CONFIG_DIR` and `OPENCODE_CONFIG_CONTENT`, which were already reachable through `envOverrides` and now count as privileged keys.
**Single-page apps work as web tabs, and a frame that reloads comes back** (#402). A history-routed dashboard (React Router, Vue Router, a Vite dev server) read `/webview/<cap>/` as its `location.pathname` and rendered its own "page not found" the moment its script ran. The proxy's runtime shim now masks the prefix off the document URL before any page script runs, while every URL the page emits still goes through the rewrite layers (now including `Worker`, `SharedWorker`, `sendBeacon` and `window.open`). A navigation the page starts itself afterwards (a dev server's full reload, a root-absolute `location.href`) used to land on Codeman's root with no capability; it is now recognised by shape, answered with a static recovery page that posts the lost path to the owning tab, and the frame is remounted inside the prefix at that path, bounded to five recoveries a minute per frame. Merged with maintainer follow-ups: the recovery path is sanitised properly (a leading backslash, or a tab/newline the URL parser deletes before parsing, resolved `/\evil.com` to a foreign origin in a direct-mode tab); a reload on the dashboard's landing page is recovered too, on password-protected and passwordless installs alike (it used to render Codeman's own shell inside the web tab); and the recovery page is written down as the third unauthenticated 200 in the security table and `docs/security-architecture.md`, with the route-enumeration property it implies stated rather than left to be discovered.
**Shift arrows for Codex on the phone keyboard bar** (#408). Two keys, `⇧←` and `⇧→`, send the Shift-modified arrows Codex binds to editing the last queued message and walking the prompt stack (verified against Codex 0.154.0's `/keymap`). Merged with a maintainer follow-up: the keys are shown only on Codex sessions (a `codex-enabled` class on the bar, the same shape as the Read My Mind key), because tapping one in any other session did nothing except hand that session to plain PTY echo for the rest of the prompt.
**Remote (SSH) cases can finally show you their files** (#421, fixes #415). File previews, downloads, text reads and the out-of-workspace attachment path resolved every path against the Codeman host's own filesystem, so in a remote case every click ended in "File not found" while the file plainly existed on the other machine. A single new ssh read layer (`src/remote-files.ts`, built on the same `buildSshConnectionArgs()` the launch uses) probes realpath and stat for the file and the workspace root in one round trip, then streams the body with `cat` (or a `tail`/`head` slice for a `Range`), so the 200/206/416 contract holds and nothing is buffered on the server. Symlinks are resolved on the host that can resolve them, containment is checked against the resolved remote root, the size cap applies to the remote size before a byte is requested, an unreachable host is a 502 rather than a 404, and there is deliberately no local fallback: a same-named file on the Codeman host is never served under a remote name. Writes, Office previews and generated thumbnails answer 400 for a remote case instead of a misleading 404. Merged with maintainer follow-ups: the `readlink -f` fallback resolved only the directory chain, so on a host without it a symlink's final component was returned unresolved and `ws/notes.txt -> ~/.ssh/id_rsa` passed containment while `cat` served the key; it now follows the last component with plain `readlink` for a bounded number of hops and fails closed (404) on a loop or the cap; `PUT /api/sessions/:id/file-content` answers 400 for a remote case as the PR already claimed (it still validated against the local filesystem, so a same-named local directory took the write); ssh children are bounded by a small semaphore (`CODEMAN_MAX_REMOTE_FILE_SSH`, default 4) covering the attachment-history fan-out, which now probes the whole history in one batched call, and the fire-and-forget magic-link registrations an injected agent could use to fork hundreds of `ssh` processes; probe records are NUL-delimited and index-keyed so a newline in a filename cannot shift one path's result onto the next; and a 502 body never carries the ssh command line.
**Docker Compose: bind-mount ownership, override files, a `codeman` runtime account, and no more stale volumes** (#377). A missing bind source (first run, cleared appdata, restored backup) is created root-owned by the daemon, and the unprivileged server crash-looped on `EACCES` when Compose was run directly; the image now starts through an entrypoint that corrects a root-owned bind mount and drops to `PUID:PGID` with `setpriv`, and the compose file adds back only the capabilities that needs. `Start-Codeman.sh` honours `docker-compose.override.yml` (naming a Compose file with `-f` silently disables Compose's own discovery of it), pre-creates the cases directory like it already did for appdata, and detects when the checkout's HEAD or lockfile moved under the `codeman-node-modules`/`codeman-dist` volumes and refreshes them, which used to leave a `docker compose build` serving stale compiled routes. The default runtime account is named `codeman` (it was `opencode`), the four global agent CLIs live in their own `/opt/codeman-cli` prefix so the runtime account can update them in place without owning `/usr/local/bin`, and `CODEMAN_ALLOWED_HOSTS` is documented and forwarded. Merged with maintainer follow-ups: `cap_add` gains `KILL` (with `init: true` tini runs as root while the server runs as `PUID`, and without CAP_KILL its SIGTERM forward failed and the server was SIGKILLed on every `compose down`/`restart`); the CLI prefix is appended to `PATH` rather than prepended and the root entrypoint pins its own `PATH`, since a `PUID`-writable directory ahead of `/usr/bin` let the runtime account plant a `setpriv` that ran as root on the next start; the entrypoint decides with a real writability probe as the runtime identity instead of an owner comparison, so ACLs, group-writable trees and NFS/CIFS mounts work and only a genuinely unwritable directory is refused, by name; the cases directory is created with the runtime owner after `PUID`/`PGID` are known; the build-source marker is written only when a refresh actually happened, an empty Compose project name falls back to `down --volumes`, the build runs before the `down` so the stack is offline only for the recreate, `docker-compose.override.*` stays out of the image, and `test/docker-entrypoint.test.ts` pins `cap_add` against what the entrypoint needs. ⚠️ Compose users: run `Start-Codeman.sh` once for this release rather than a plain `docker compose up`, so the rebuilt image, the refreshed volumes and the new entrypoint arrive together.
**Selected text is visible again on the light skins** (#423, part of #360). Every skin palette named its selection layer `selection`, the key xterm renamed to `selectionBackground` in v5, so all seven skins had been painting xterm's default white at 30% instead of the colour next to it in the palette. Dark skins hid it; on the four light skins a selection was white on near-white. The key is renamed and `test/skin-themes.test.ts` pins it. CI additionally exercises `install.sh`'s dsh identity probe with `timeout` missing under bash 3.2 (#422), the guard #382's fix shipped without.
**Eight fixes salvaged from #375** (dignfei; landed with the author's commits preserved, the rest of that PR is covered below). Shift+drag starts a text selection in a pane whose mouse reports go to the CLI, and right-click copies the selection. Ctrl- and Alt-modified navigation keys typed through the CJK composer reach the CLI as the modified sequences instead of plain arrows. A browser whose reliable-input sequence counter fell behind the server's watermark (a restored tab, a cleared localStorage) now recovers: the duplicate ACK carries `dup: true` plus the watermark, the client lifts its counter and re-sends, so a session that had silently stopped accepting typed prompts accepts them again. An SSE reconnect that lands on the session you are already looking at keeps its terminal buffer and resyncs instead of resetting the whole terminal. The hidden offline overlay and the file-preview overlay only apply `backdrop-filter` while shown, which removes a stale compositing layer that swallowed clicks. One adopted Docker container can back several cases at different in-container directories, and the adopt panel gains a "copy an existing case" picker. Of the PR's 27 commits, 14 had already shipped through #357, the selection theme key rename shipped as #423, and foreign tmux adoption plus SSH password auth stay with the author.
### Thanks
- **@opticon454** for custom model endpoints (#393), including the part nobody enjoys: working out each CLI's real endpoint mechanism against real binaries and writing down which ones do not work yet instead of claiming they do; and for the Docker Compose deployment fixes (#377), rebased and reworked through three review rounds.
- **@shenlvkang-collab** for making single-page apps route inside web tabs and recovering a frame that reloads (#402), the best-engineered PR of this batch, and for the Codex Shift arrows on the phone keyboard bar (#408), verified against Codex's own keymap.
- **@dignfei** for the eight fixes salvaged from #375 (terminal selection and copy, CJK navigation keys, input recovery, SSE reconnect, overlay compositing, multi-case adopted containers), landed under their own name.
- **@Randalix** for reporting #415 and then fixing it themselves with the whole missing ssh read side for remote cases (#421), with a real-shell test for the probe script and a full route suite.
### Patch Changes
- 349a89e: fix(webview): let a proxied single-page app route on its own path, and recover a frame that reloads
A dashboard served through a web tab saw `/webview/<cap>/` as its `location.pathname`, and
no app has a route for that: a React Router, Vue Router or Vite dev-server page painted its
HTML and CSS and then replaced them with its own "page not found" the moment its script ran.
The proxy's runtime shim now rewrites the history entry to the path the page would see on its
own origin before any page script runs, while every URL the page emits still goes through
the existing rewrite layers (plus `Worker`, `sendBeacon` and `window.open`, which the masked
Referer can no longer rescue). A navigation the page starts itself afterwards — a dev
server's full-reload HMR, a root-absolute `location.href` — lands on Codeman's root with no
capability; it is recognised by shape (an iframe navigation asking for HTML for a path Codeman
does not serve), answered with a static page that tells the owning tab which path was lost,
and the tab remounts the frame inside the prefix at that path. That answer is served before
the credential checks, so it never counts as a failed login.
- 013a5d9: File previews, downloads and text reads now work in a **remote (SSH) case**.
A remote case's working directory is an absolute path on the _remote_ host, but the
file routes resolved it with local `fs` — so a clicked path (or the File Viewer) always
failed as "File not found" even though the file existed and the session was clearly
working in that directory. `GET /api/sessions/:id/file-raw`, `file-content`,
`file-preview` and `file-thumbnail` now resolve and read through the same
`buildSshConnectionArgs()` connection the launch uses (`src/remote-files.ts`, one
`realpath`+`stat` probe per request returning both the file and the workspace root).
Clicked paths that point OUTSIDE the case directory (a remote `/tmp` scratchpad capture,
a screenshot elsewhere in the remote home) go through the attachment routes, which had
the same local-`fs` assumption: registration, the by-id `raw` stream, the metadata poll
and the attachment history list now resolve over ssh as well, so the click-path works
whether the file sits inside or outside the case. Which host a record is read from
follows the SESSION, never the path string — the same absolute path means a different
file on each host, and a remote session never falls back to a local file.
The guards are unchanged in strength: the workspace boundary is still enforced (now
resolved on the host that can actually resolve it), the sensitive-path blocklist and
the size cap (`CODEMAN_MAX_DOWNLOAD_BYTES`) still apply before any bytes are read, and
`Range` requests keep working, so remote `<video>`/`<audio>` seeking behaves like a
local file. An unreachable host is reported as `502` with the remote reason instead of
a misleading 404. Nothing is ever copied to the Codeman host.
Still not available for remote cases, and now said explicitly instead of 404-ing:
editing a file (`edit=1` / `PUT` answer 400, the viewer hides its Edit affordance),
office-document previews and generated thumbnails (both need the bytes on the server's
disk), the file tree / path picker, and `tail-file`. Docker cases are unaffected (their
workspace is bind-mounted at the same absolute path).
- b357fe8: Add Shift+Left and Shift+Right buttons to the default and extended mobile agent keyboard bars, shown only on Codex sessions, enabling Codex queued-message editing and prompt-stack navigation. Flush locally buffered drafts before navigation and keep terminal focus after taps.
- 9acc5aa: Fix an invisible terminal text selection on the light skins (#360). Every xterm palette declared its selection colour under the key `selection`, which xterm.js renamed to `selectionBackground` in v5. An `ITheme` is a plain object, so the unknown key was dropped without an error and every skin fell back to xterm's own default of `rgba(255,255,255,0.3)`: unnoticeable on the dark skins, which wanted roughly that anyway, and effectively invisible on Paper Gray, Solarized Light, Catppuccin Latte and Rosé Pine Dawn, where white at 30% over a near-white background moves a channel by about 3/255. Selecting text on those skins now highlights it, with desktop drag-select and the mobile long-press both fixed by the same rename.
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, DeepSeek Harness, or OMP inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
Get started in one line (macOS & Linux, Windows via WSL):
@@ -42,7 +42,7 @@ codeman web
The installer asks before every system change, and re-running the same line updates in place. Full details: [Quick Start - Installation](#quick-start---installation).
- **One dashboard, eight CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **One dashboard, nine CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, DeepSeek, or OMP](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions), with your own dashboards open as [web tabs](#more-features) beside them
- **Truly phone-friendly** - a [touch-optimized terminal](#mobile-optimized-web-ui) with instant local echo, QR login, swipe navigation, and push notifications
- **Runs while you sleep** - [idle detection + respawn cycling](#respawn-controller) and auto-resume when a subscription limit resets, for 24+ hour unattended runs
- **See your agents think** - [live floating windows](#live-agent-visualization) for every subagent and teammate, with real-time transcripts
@@ -61,14 +61,15 @@ The installer asks before every system change, and re-running the same line upda
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js, tmux and a build toolchain if missing (node-pty ships no Linux prebuilds, so it compiles from source), clones Codeman to `~/.codeman/app`, and builds it. A few things worth knowing:
This installs Node.js, tmux and a build toolchain if missing (node-pty ships no Linux prebuilds, so it compiles from source), clones Codeman to `~/.codeman/app`, and builds it. It looks at what is already on the machine, asks at most three questions, then does all the work unattended and ends on the URL with a QR code for your phone. A few things worth knowing:
- **It asks first.** Every system change (package installs, AI CLI download) is prompted, and a menu at the end lets you choose: run Codeman in this terminal, install it as a background service (systemd/launchd, auto-start on boot), or don't start yet. Nothing runs in the background unless you pick it.
- **How it's reachable, your choice.** The installer offers three ways to reach the dashboard:**Tailscale** (loopback bind fronted by `tailscale serve`, so you get `https://<machine>.<tailnet>.ts.net` with a real certificate and your tailnet as the login, no password needed), **any device on your network** (`0.0.0.0`, with a strongly recommended password prompt), or **this machine only** (`127.0.0.1`, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. The highlighted default reflects what is already on the machine (Tailscale when it is already in use, your existing binding on a re-run), and a bare Enter never pulls in new software. A bare `codeman web` started by hand still defaults to loopback.
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh update` and `install.sh uninstall` also exist.
- **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
- **Three questions, all up front.** How the dashboard is reached, optionally what to call this machine on your tailnet, and whether to run Codeman as a background service (systemd/launchd, auto-start on boot; Enter says yes). Everything that needs you, including one consent for all missing packages, one sudo password, and the Tailscale login, happens before the build, so you can walk away while it compiles.
- **How it's reachable, your choice.** **Tailscale** (loopback bind fronted by `tailscale serve`, so you get `https://<machine>.<tailnet>.ts.net` with a real certificate and your tailnet as the login, no password needed), **any device on your network** (`0.0.0.0`, with a strongly recommended password prompt), or **this machine only** (`127.0.0.1`, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. The highlighted default reflects what is already on the machine (Tailscale when it is already connected, your existing binding on a re-run), and a bare Enter never pulls in new software. If another app already owns `:443` on your node, Codeman goes under `https://<machine>.<tailnet>.ts.net/codeman` or on a second port instead of replacing it. A bare `codeman web` started by hand still defaults to loopback.
- **The name is yours to choose.** By default the URL uses the machine's existing tailnet name. Answering yes to the second question renames the machine to `codeman-<hostname>` (which also renames it for SSH, so the default is no); `install.sh name` does it later.
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh status` prints the URLs and the QR code again; `install.sh update`, `install.sh tailscale` and `install.sh uninstall` also exist.
- **Flags for the impatient.** `curl -fsSL https://getcodeman.com/install | bash -s -- --tailscale --service` answers the questions from the command line (`--lan`, `--local`, `--run`, `--no-start`, `--name <n>`, `--port <n>`, `--yes` too). **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently; set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), or [OMP](https://github.com/can1357/oh-my-pi) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the nine is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), or [OMP](https://github.com/can1357/oh-my-pi) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the nine is present; if none is found, it offers to install any of them from a menu (DeepSeek excepted, since its npm package installs only a launcher with no runnable profile), or you can skip and install one yourself later. After install:
```bash
codeman web
@@ -82,7 +83,7 @@ codeman users add alice --admin # create the first admin account
codeman web --multiuser # named logins + per-user case spaces
```
**Prefer Docker Compose?** A local-image Compose deployment ships in `docker/`: copy `docker/.env.example` to `docker/.env`, set `CODEMAN_PASSWORD`, then run `bash docker/Start-Codeman.sh` on Linux. Codeman runs in a container and spawns Docker cases as sibling containers through the host socket. See the [Docker deployment guide](docker/README.md) for direct Compose commands, storage and networking options.
**Prefer Docker Compose?** A local-image Compose deployment ships in `docker/`: copy `docker/.env.example` to `docker/.env`, set `CODEMAN_PASSWORD`, then run `bash docker/Start-Codeman.sh` on Linux. Codeman runs in a container and spawns Docker cases as sibling containers through the host socket. After updating, run the script again rather than a plain `docker compose up`, so the rebuilt image, refreshed volumes and entrypoint arrive together. See the [Docker deployment guide](docker/README.md) for direct Compose commands, storage and networking options.
Details in [Multi-User Mode](#multi-user-mode-opt-in) below.
@@ -209,17 +210,17 @@ The most responsive AI coding agent experience on any phone. Full xterm.js termi
<tr><td>Password typing on phone</td><td><b>QR code scan — instant auth</b></td></tr>
</table>
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard; destructive commands require a double-press to confirm, so you never fire one by accident
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard; destructive commands require a double-press to confirm, so you never fire one by accident; on Codex sessions the bar also shows `⇧←` / `⇧→` (Shift+Left / Shift+Right: edit the last queued message / return through the prompt stack)
- **Dedicated Enter button** — replays the keypress through the terminal, so text buffered by local echo is flushed first rather than stranded
- **Swipe navigation & smart keyboard handling** — swipe left/right to switch sessions; toolbar and terminal shift up when the keyboard opens (`visualViewport` API)
- **Built for phones** — safe-area insets for notch and home indicator, 44px touch targets, bottom-sheet case picker, native momentum scrolling
- **Built for phones** — safe-area insets for notch and home indicator, 44px touch targets, bottom-sheet case picker, native momentum scrolling; on a folding phone (iPhone Duo) dialogs stay clear of the hinge, and opening or closing the device is never mistaken for the keyboard
```bash
codeman web --https
# Open on your phone: https://<your-ip>:3000
```
> `localhost` works over plain HTTP. Use `--https` when accessing from another device, or use [Tailscale](https://tailscale.com/) (recommended): the installer can set it up for you (choose **Tailscale** at the network-access prompt, or run `bash ~/.codeman/app/install.sh tailscale` on an existing install). That gives you `https://<your-machine>.<tailnet>.ts.net` with a real certificate: private to your tailnet, no password required, and PWA install + push notifications work on your phone.
> `localhost` works over plain HTTP. Use `--https` when accessing from another device, or use [Tailscale](https://tailscale.com/) (recommended): the installer can set it up for you (choose **Tailscale** at the network-access prompt, or run `bash ~/.codeman/app/install.sh tailscale` on an existing install). That gives you `https://<your-machine>.<tailnet>.ts.net` with a real certificate: private to your tailnet, no password required, and PWA install + push notifications work on your phone. The installer ends on that URL with a QR code to scan, and `bash ~/.codeman/app/install.sh status` prints it again any time.
### Secure QR Code Authentication
@@ -255,7 +256,7 @@ Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. **Add Case** creates one from scratch, links an existing folder, or clones a GitHub repo straight into one (**Clone Repo**). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, `Pi`, `Grok`, `OMP`, or `Terminal` (plain shell). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, `Pi`, `Grok`,`DeepSeek`,`OMP`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Models → New Claude sessions). A soft default — `/model` still works in-session. |
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
@@ -263,7 +264,7 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
### 3. Read the dashboard
- **Tabs (top)** — one per session. `Alt+1`-`9` to jump, `Ctrl+Tab` for next, drag to reorder (tab order syncs across your devices).
- **Tabs (top)** — one per session. `Alt+1`-`9` to jump, `Ctrl+Tab` for next, drag to reorder (tab order syncs across your devices). Prefer a list? **App Settings → Appearance → Tabs** moves it into a left sidebar with a filter box (`Alt+B` collapses it) or a vertical rail whose rows sort by activity: blocked on you first, then longest running, then most recently quiet.
- **Terminal (center)** — a real `xterm.js` terminal; full TUIs render correctly. Type directly and press **Enter** to send. `Shift+Enter` inserts a newline.
- **Side panels** — Respawn, Orchestrator, Cron, Subagents, Settings (toggled from the toolbar).
@@ -271,8 +272,10 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
- **Type prompts** straight into the terminal — input is delivered exactly-once even across reconnects (a dropped link never loses or double-sends a prompt).
- **Paste or drag-and-drop images** directly into the session.
- **Voice input** — `Ctrl+Shift+V` (Deepgram Nova-3, with auto-silence stop).
- **Attachments** — register external files/docs and preview Office/PDF inline.
- **Voice input** — `Ctrl+Shift+V` (Deepgram Nova-3, or this machine's Claude Code login with no API key; auto-silence stop).
- **Attachments** — register external files/docs and preview Office/PDF inline; any file path an agent prints is clickable, in the terminal and in the chat view.
- **When it needs you** — the tab turns yellow (waiting for input) or red (a question is blocking). The **Approvals Inbox**_(opt-in)_ queues every pending prompt across sessions, answerable from the header bell or the phone home screen, and 🧠 **Read My Mind**_(opt-in)_ drafts your next prompt from the case's goals and recent work.
- **Copy what you see** — `Shift+drag` selects text even while the CLI owns the mouse, right-click copies it, and Auto Copy _(opt-in)_ copies a selection the moment you release it.
### 5. Make it autonomous
@@ -291,7 +294,7 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
### 7. Operate & maintain
- **App Settings** — model, effort, permission startup mode, theme/skin, notifications, display toggles, per-CLI options, a synced custom display name, and per-device English/Simplified Chinese UI language.
- **App Settings** — model, effort, permission startup mode, theme/skin, terminal font family and weight, entrance animations, notifications, display toggles, per-CLI options, a synced custom display name, and per-device English/Simplified Chinese UI language.
- **Run it in the background** — `codeman web -d` detaches from your shell (`--status`, `--stop`); `codeman service install` makes it a systemd user unit / macOS LaunchAgent that survives reboots. Both verify the server actually answers before reporting success, and both refuse to start a second server on one data dir. See [Keep it running in the background](#quick-start---installation).
- **Self-update** — git-clone installs update in place from **App Settings → System → Updates**.
- **Deploy your own changes** — see [Development](#development).
- **Background daemon & service install** — `codeman web -d` runs the server detached with a pidfile, `~/.codeman/web.log`, and verified startup (it polls the server until it answers, so a port clash never reads as success); `codeman service install` writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew `node`, `tmux` and `claude` are actually found. Secrets are never written into unit files
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → System → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Clone a GitHub repo as a case** — paste a repository URL into **Add Case → Clone Repo** and Codeman clones it into `~/codeman-cases/<name>` and registers it as a normal case, ready to run an agent in. It preflights the URL while you type (tells you whether it can be cloned anonymously and offers the repo's real branches and tags for the optional branch/tag field), fills the case name in from the URL, and lets you pick which CLI the Run button should use. Public repositories over `https://`; Codeman never collects or stores credentials
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, **Pi**, **Grok**, or **OMP** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*` vs `GROK_*`/`XAI_*` vs `OMP_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md), [`docs/pi-integration.md`](docs/pi-integration.md), [`docs/grok-integration.md`](docs/grok-integration.md) and [`docs/omp-integration.md`](docs/omp-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, **Pi**, **Grok**,**DeepSeek Harness**, or **OMP** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*` vs `GROK_*`/`XAI_*` vs `DSH_*`/`DEEPSEEK_*` vs `OMP_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md), [`docs/pi-integration.md`](docs/pi-integration.md), [`docs/grok-integration.md`](docs/grok-integration.md), [`docs/deepseek-integration.md`](docs/deepseek-integration.md) and [`docs/omp-integration.md`](docs/omp-integration.md)
- **Custom model endpoints** _(new in 1.29.0, HTTP API for now)_ — point a session's CLI at any OpenAI-compatible endpointinstead of its native backend: a local llama.cpp, llama-swap, Ollama or vLLM box, or a cloud gateway such as Azure AI Foundry or OpenRouter. Save an endpoint once (`POST /api/model-endpoints`; its models are discovered from `/v1/models`), apply it to a session (`POST /api/sessions/:id/custom-model`), and the CLI restarts in place on that endpoint. Verified live for Claude, OpenCode, Pi, Grok and OMP; Codex, Gemini and DeepSeek have documented gaps, Antigravity has no mechanism. A toolbar picker is the follow-up. See [`docs/custom-model-endpoints.md`](docs/custom-model-endpoints.md)
- **Web tabs** — open Grafana, Uptime Kuma, a Vite dev server or any dashboard URL as a tab beside your sessions (Run dropdown → **Web / URL** → **Add URL**). Dashboards are proxied through Codeman's own origin, so an `http://` target works from a phone over HTTPS and through the tunnel, single-page apps route on their own paths, and a frame that reloads recovers itself. A `localhost` link an agent prints opens as a web tab automatically. See [`docs/web-tabs.md`](docs/web-tabs.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container, or attach a case to a container you already run; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host; file previews and downloads come over the same ssh connection. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
- **Voice input** — dictate prompts with Deepgram Nova-3 (Web Speech API fallback): toggle recording, auto-silence stop, live level meter (`Ctrl+Shift+V`)
- **Voice input** — dictate prompts with Deepgram Nova-3, or through this machine's Claude Code login with no API key at all (App Settings → Voice; Web Speech API fallback): toggle recording, auto-silence stop, live level meter (`Ctrl+Shift+V`)
- **Image input** — paste or drag-and-drop images straight into a session
- **Gesture control** _(opt-in)_ — a MediaPipe hand-tracking overlay to grab/drag session windows and pinch buttons, hands-free. Enable with `CODEMAN_GESTURE=1` + App Settings → Terminal & Input
- **Multi-monitor span** _(macOS)_ — one click opens a browser window maximized across all displays, so floating agent/gesture panels can cross the physical seam
- **File Viewer button** _(opt-in)_ — a header button that toggles the built-in file browser panel with one tap; enable under App Settings → Header & Panels → Header buttons
- **CJK / IME input** — full composition support for Chinese / Japanese / Korean
- **CJK / IME input** — full composition support for Chinese / Japanese / Korean, with Ctrl- and Alt-modified navigation keys passed through to the CLI
- **Plan usage in the header** — live Claude subscription usage (the 5-hour and weekly windows) from a statusline exporter Codeman hands to `claude` at spawn and never writes into your settings files, plus Codex limits from its own app-server; per device, on for desktops and off for phones
- **Session list, your way** — the header strip, a left sidebar with a filter box, or a vertical rail whose detailed rows carry created and state stamps and sort by activity; the phone home screen and the desktop home rail use the same order
- **Terminal looks** — seven skins, four of them light, per-device font family and weight (the bundled JetBrains Mono covers weights 100 to 800), and opt-in entrance animations for tabs, agent windows, the terminal pane and connection lines
- **OS notifications & hostname-aware titles** — desktop alerts and tab titles are prefixed `codeman:<host>` so multi-host setups stay unambiguous
---
@@ -461,8 +469,9 @@ Run a case inside its own hardened Docker container instead of directly on your
- **Resource templates** — expand the checkbox for a **Small / Medium / Large / GPU** preset (memory, CPUs, GPU), or set your own. **Disk is elastic** — storage grows as data flows in, no fixed cap.
- **Shared per-case container** — many sessions can `docker exec` into the same container; killing one session never tears the container out from under the others.
- **Hardened by default** — non-root, `--cap-drop ALL`, `no-new-privileges`, PID/memory caps, never `--privileged` or the docker socket; a **sealed** profile (no host credentials, network off) is one toggle away.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / Pi logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / OMP logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / Pi / Grok / OMP logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Attach to a container you already run** — tick **Attach to an existing container** on the Docker panel to link a case to it instead of creating one. Codeman only `exec`s into it and never starts, stops, restarts or removes it; one adopted container can back several cases at different directories, and **copy an existing case** pre-fills the form from a sibling. Admin-only in multi-user mode, since the container's mounts belong to whoever started it.
- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Durable** — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.
Prerequisite: just Docker (or Podman). The agent base image builds itself automatically on first use, with progress streamed to the UI (or pre-build it with `node scripts/build-agent-image.mjs`). Full guide: [`docs/docker-cases.md`](docs/docker-cases.md).
@@ -478,6 +487,7 @@ Point a case at another machine and run the agent **there**, over SSH, with the
- **Discover & attach**: list the `codeman-*` sessions already running on a host (started by that machine's own Codeman, or by another operator) and attach to one. Attached sessions you don't own **detach on tab close, never kill**.
- **Shared sessions**: several clients can attach the same remote session at different window sizes without clamping each other; discovery shows a "shared" badge with the client count.
- **Injection-safe**: every ssh command line flows through a single shell-escaping builder, and host/path/identity fields are schema-guarded.
- **Files too**: previews, downloads and text reads in a remote case go over the same ssh connection (one `realpath` + `stat` probe, then a streamed `cat`, `Range` seeking included), so a clicked path opens the file on the machine the agent is on. Nothing is copied to the Codeman host; editing and Office previews answer a clear 400 instead of a misleading 404.
Set it up under **New Case → Remote** (host, user, identity file, optional jump host). Full design: [`docs/remote-sessions.md`](docs/remote-sessions.md).
@@ -647,8 +657,8 @@ These run for **every** request — before auth, even on the default no-password
### Input, files & headers
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` / `PI_*`env-prefix allowlist gates which settings each CLI can receive
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` / `PI_*`/ `GROK_*` / `XAI_*` / `DSH_*` / `DEEPSEEK_*` / `OMP_*` env-prefix allowlist gates which settings each CLI can receive, and the keys that could redirect a CLI's traffic (base URLs, config homes) are clamped for non-admin users
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 2 GB raw & download (`CODEMAN_MAX_DOWNLOAD_BYTES`; bodies stream and answer `Range` requests, so the cap is a sanity bound rather than memory protection); `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
- **Security headers** — `Content-Security-Policy` (`default-src 'self'`, every exception enumerated), `X-Content-Type-Options: nosniff`, `X-Frame-Options: SAMEORIGIN`, HSTS over HTTPS, and CORS reflected **only** for `localhost` / `127.0.0.1` / `::1`
### Supply chain & isolation
@@ -698,6 +708,10 @@ The web UI remains the primary surface; see **[docs/tui.md](docs/tui.md)** for t
| `Ctrl/Cmd +` / `-` | Font size |
| `Ctrl/Cmd+?` | Keyboard help |
| `Shift+Enter` | Insert newline (sent to terminal) |
| `Shift+drag` | Select text in a pane whose mouse events go to the CLI |
| Right-click | Copy the selection (the native menu stays when nothing is selected) |
| `Shift+Wheel` | Scroll the local scrollback while the wheel is forwarded to the CLI |
| `Ctrl+Z` | Swallowed in agent sessions so a running CLI cannot be suspended; normal job control in a shell |
| `Escape` | Close panels & modals |
---
@@ -762,7 +776,7 @@ Those `DONE_<task>_<random>` strings are the skill's **split marker** trick, and
| [`SKILL.md`](skills/codeman/SKILL.md) | Safety rules, the ready-made fast path (spawn N workers, task them, collect), and the verb index. Always loaded. |
| [`reference/verbs.md`](skills/codeman/reference/verbs.md) | The 14 verbs in detail: readiness, send-and-wait, markers, interrupts, cleanup. On demand. |
| [`reference/recipes.md`](skills/codeman/reference/recipes.md) | 6 worked multi-worker flows (fan-out, blocked-worker watch, messaging fan-out). On demand. |
| [`reference/recipes.md`](skills/codeman/reference/recipes.md) | 8 worked flows: claude, DeepSeek Harness and shell workers, fan-out, blocked-worker watch, messaging fan-out. On demand. |
| [`reference/endpoints.md`](skills/codeman/reference/endpoints.md) | Full endpoint tables, error codes, per-mode signal table, capacity limits. On demand. |
| [`reference/messaging.md`](skills/codeman/reference/messaging.md) | Talking to claude workers directly via Claude Code cross-session messaging. On demand. |
@@ -798,8 +812,8 @@ When a CLI runs in a Codeman-managed session, these environment variables are se
4.**Response envelope.** Most endpoints return `{ "success": true, "data": … }` (errors: `{ "success": false, "error", "errorCode" }`). A few legacy GETs return bare bodies — **handle both** (`body.data ?? body`).
5.**`/api/v1/*`** is a stable alias of `/api/*`.
6.**Wait instead of polling, and don't treat a timeout as an error.** The wait endpoints answer with HTTP `200` and `wait.timedOut: true` when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. `wait.timeoutMs` tells you the timeout the server actually applied after clamping (600s ceiling).
7.**Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity/pi) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
7.**Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity/omp) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.8.**Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
7.**Only `claude` and `deepseek` sessions emit `stop` and `blocked`.** Those two come from hooks (Claude Code's own, and the DeepSeek Harness status bridge); `shell` and the other external CLIs (opencode/codex/gemini/antigravity/pi/grok/omp) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
8.**Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
# 6. Stream live events (session output, agent activity, status)
@@ -914,7 +939,7 @@ Codeman registers Claude Code hooks that `POST /api/hook-event` (`permission_pro
## API
REST over Fastify — **~200 handlers across 21 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
REST over Fastify — **~230 handlers across 25 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
### Sessions
@@ -925,11 +950,13 @@ REST over Fastify — **~200 handlers across 21 route modules**, plus an SSE str
| `GET` | `/api/sessions/:id/terminal` | Read terminal output (`?tail=<bytes>`, `?full=1`); the read path for interactive sessions |
| `GET` | `/api/sessions/:id/output` | Parsed one-shot output (`textOutput` is empty for tmux-backed sessions) |
| `GET` | `/api/sessions/:id/last-response` | The last answer as clean text, read from the transcript (claude, codex, deepseek) |
| `GET` | `/api/sessions/:id/wait` | Block until a signal fires (`?until=stop,idle,exit&timeout=&fresh=`); a timeout is a `200` |
| `GET` | `/api/sessions/:id/wait-output` | Block until a literal string appears (`?match=&nocase=&from=now\|buffer&timeout=`) |
| `GET` | `/api/sessions/unified` | Unified live + history list (Session Manager) — `?q=&limit=` |
| `POST` | `/api/sessions/:id/pin` | Pin/unpin in the Session Manager (`{pinned}`) |
| `PUT` | `/api/session-order` | Sync tab order across devices (`{order: [ids]}`) |
| `POST` | `/api/sessions/:id/custom-model` | Restart the session's CLI on a saved custom endpoint (`{endpointId, modelId}`; `{clear: true}` returns to the native backend) |
> **Building something on top of Codeman?** [`docs/extending-codeman.md`](docs/extending-codeman.md) is the integration guide: render your own UI as a tab, subscribe to the SSE event stream to react when an agent needs you, drive Codeman from a script, and the traps worth knowing before you start. Codeman has no plugin runtime on purpose, so an integration is just your own process talking HTTP.
Instant keystroke feedback overlay for xterm.js. Eliminates perceived input latency over high-RTT connections by rendering typed characters immediately as a pixel-perfect DOM overlay. Zero dependencies, 6.1 kB gzipped, configurable prompt detection, CJK/emoji wide-character support, full state machine with 175 tests.
Instant keystroke feedback overlay for xterm.js. Eliminates perceived input latency over high-RTT connections by rendering typed characters immediately as a pixel-perfect DOM overlay. Zero dependencies, 6.1 kB gzipped, configurable prompt detection, CJK/emoji wide-character support, full state machine with 238 tests.
- **需要你的时候** —— 标签会变黄(等待输入)或变红(有个问题挡住了它)。**审批收件箱(Approvals Inbox)**(可选启用)把所有会话里等着你的提示排成一个队列,可以从页头的铃铛或手机首页直接作答;🧠 **Read My Mind**(可选启用)会根据这个 case 的目标和最近的工作替你起草下一条提示。
- **语音输入** —— 用 Deepgram Nova-3 口述提示(带 Web Speech API 回退):切换录音、自动静音停止、实时音量表(`Ctrl+Shift+V`)
- **语音输入** —— 用 Deepgram Nova-3 口述提示,或者干脆用这台机器的 Claude Code 登录、不需要任何 API key(App Settings → Voice;带 Web Speech API 回退):切换录音、自动静音停止、实时音量表(`Ctrl+Shift+V`)
On PowerShell, use the following commands instead. Running Compose from inside `docker/` with no `-f` lets it discover `docker-compose.override.yml` on its own (see [Local customisation](#local-customisation)); naming the file with `-f docker/docker-compose.yaml` from the repository root silently drops the override unless it is named too.
Every required value is defined and explained in `.env.example`. `GEMINI_API_KEY` is intentionally optional and may remain blank.
The container starts as root so `entrypoint.sh` can correct the ownership of a bind source the Docker daemon created (it creates a missing one as `root:root`), then drops to `PUID:PGID` with `setpriv` before the server starts, so Codeman itself never runs privileged. That drop needs `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against the file's `cap_drop: ALL`; a compose file written elsewhere (Unraid's Compose Manager, a hand-written unit) must carry the same additions, and the entrypoint names them when they are missing. A directory owned by neither root nor `PUID:PGID` is never re-owned: it is probed for writability as the runtime account and refused with a clear message if that fails. Setting `user:` in Compose skips the whole step.
On Linux, `Start-Codeman.sh` stops with an error when required paths are missing. It creates the application-data directory when safe, detects its numeric owner as `PUID:PGID`, and detects `DOCKER_SOCKET_GID` from the configured Docker socket. It rejects a root-owned application-data directory because Codeman and its local CLI sessions must remain unprivileged.
Codeman, Claude, OpenCode, and other local sessions run as the unprivileged account named by `CODEMAN_RUNTIME_USER`, which defaults to `opencode`. When Compose is run directly, `PUID` and `PGID` default to `1000:1000`; set them in `.env` when the application-data directory has a different owner. The Bash start script determines them automatically instead.
Codeman, Claude, OpenCode, and other local sessions run as the unprivileged account named by `CODEMAN_RUNTIME_USER`, which defaults to `codeman`. When Compose is run directly, `PUID` and `PGID` default to `1000:1000`; set them in `.env` when the application-data directory has a different owner. The Bash start script determines them automatically instead.
To retain Docker-case support without root when running Compose directly, set `DOCKER_SOCKET_GID` to the numeric group ID of the host socket. On a standard Linux Docker host, obtain it with `stat -c '%g' /var/run/docker.sock`. The Bash start script detects it automatically.
@@ -38,6 +41,131 @@ Releases that change `server.Dockerfile`, `docker-compose.yaml`, or add a key to
changed, and asks you to run `Start-Codeman.sh` here on the host instead. Details:
`Start-Codeman.sh` rebuilds the image on every start, but with the layer cache,
and it refreshes the build-artefact volumes selectively: `codeman-dist` when
the checkout's HEAD moved, `codeman-node-modules` only when `package-lock.json`
changed. That is right for an ordinary `git pull`. It is not enough when a
`server.Dockerfile` change bumps the Node base image without touching the
lockfile: `node-pty` is compiled from source (there is no Linux prebuild), so
the old `codeman-node-modules` volume would keep a build made for the previous
Node version. For that case, or whenever you want to be certain of what ships,
`docker/Update-Codeman.sh` force-rebuilds the image with no layer cache, stops
the stack, removes the `codeman-node-modules` and `codeman-dist` volumes, then
hands off to `Start-Codeman.sh` for the usual start:
```sh
bash docker/Update-Codeman.sh
```
Pass `--keep-volumes` to skip clearing them (safe only if you know the
rebuilt image's `node_modules`/`dist` did not change). The scripted default
is the "Resetting the build artefacts" procedure in
[`../docs/docker-self-update.md`](../docs/docker-self-update.md). Only those
two volumes are removed, by name within this Compose project; any volume a
`docker-compose.override.yml` adds is left alone, and application data and
case workspaces are host bind mounts, never touched either way.
## Private repositories (GitHub and Azure DevOps)
The images can include the GitHub CLI (`gh`) and the Azure CLI (`az`, with the `azure-devops` extension), wired into the system Git configuration as credential helpers, so Codeman can clone private repositories. Both are **opt-in and off by default**, and are turned on per host in `docker-compose.override.yml`.
### Turning them on
Add the build arguments to `docker-compose.override.yml` (see [Local customisation](#local-customisation)), then rebuild with `Start-Codeman.sh`. Set only the one you need:
```yaml
services:
codeman:
build:
args:
CODEMAN_INSTALL_GH:'1'
CODEMAN_INSTALL_AZ:'1'
environment:
# The same two switches for the Docker-case agent image Codeman builds.
CODEMAN_AGENT_IMAGE_INSTALL_GH:'1'
CODEMAN_AGENT_IMAGE_INSTALL_AZ:'1'
```
The `build: args:` pair controls the Codeman server image. The `environment:` pair controls the agent image for [Docker cases](../docs/docker-cases.md), which Codeman builds on the first Docker case; an agent image that already exists is not rebuilt by this, so run `node scripts/build-agent-image.mjs --no-cache` inside the container afterwards. The same variables work in front of that command when building it by hand. Values must be `0` or `1`; anything else stops the build with an error naming the argument.
They are not `.env` settings: turning a CLI on is a per-host choice, which is what the override file is for, and a new `.env.example` key makes the in-app updater refuse to update every existing installation until its `.env` gains the key.
The Azure CLI is the large one, about 600 MB of the roughly 670 MB the pair adds. A CLI left off leaves nothing functional behind: no apt repository, no package, no `azure-devops` extension and no credential-helper entry, so git for that host behaves exactly as it does without this feature. With both off the image is functionally unchanged; it still carries the `AZURE_EXTENSION_DIR` variable, an empty extensions directory and one small layer that copies and then removes the helper script.
### Signing in
With a CLI on, the system Git configuration routes credentials through it:
Codeman itself still collects no Git credentials. Sign the container in once from a **Terminal / Shell** session (Run menu). The session runs as the runtime account, so the sign-in is stored under `CODEMAN_APPDATA_PATH` (`~/.config/gh`, `~/.azure`) and survives rebuilds and container recreation:
```sh
gh auth login # GitHub.com -> HTTPS -> "Login with a web browser" (device code)
az login --use-device-code # then: az devops configure --defaults organization=https://dev.azure.com/<org>
```
After that, **Add Case → Clone Repo** accepts private `https://` URLs on those hosts, and `git clone` works from any session. Until a CLI is signed in its helper prints nothing, so a private clone fails immediately with the usual authentication error rather than waiting on a prompt.
**Multi-user mode:** every Codeman user's git runs as the same server account, so these sign-ins would otherwise be shared. Clone Repo therefore runs a **non-admin**'s clone and preflight with every git credential helper cleared (`git -c credential.helper=`): a non-admin can clone public repositories and anything their own SSH setup allows, but not a private https repository through the admin's `gh`/`az` sign-in. Admins, and single-user mode, keep the helpers. A non-admin's own agent sessions still run as that same account, and with the agent-image `gh`/`az` switches on, a non-admin's Docker case with credential seeding on also receives the server account's `gh`/`az` sign-in, the same as the Claude and Codex credentials; see `docs/security-architecture.md`, multi-user mode.
Azure DevOps is authenticated with an Entra ID access token that the helper requests from `az` for each Git operation, so nothing is written to disk beyond `az`'s own sign-in. An account that has to use a personal access token can set `AZURE_DEVOPS_EXT_PAT` for the container instead (for example under `environment:` in `docker-compose.override.yml`); the helper prefers it when present. SSH remotes are unaffected by any of this and keep using the account's own keys.
Docker cases copy these sign-ins into a case container only when the matching agent-image switch is on (`CODEMAN_AGENT_IMAGE_INSTALL_GH=1` for `~/.config/gh/hosts.yml` and `config.yml`, `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` for the sign-in files from `~/.azure`) and the case has credential seeding on. With a switch off they are never copied, even when the files exist, because a GitHub token or an Azure refresh token is usable by anything in the container. The copies are made when the container is **created**, so an existing case container never picks them up: after turning a switch on, signing in, or rebuilding the agent image, **recreate the case container** (remove it; the next session in that case creates a fresh one).
The GitHub agent skill for `gh` installs into the runtime account's home in the same session:
```sh
gh skill install cli/cli gh --scope user
gh skill update gh # after a later gh release
```
### Versions
Both CLIs, and the extension, are installed from their vendors' repositories with no version pinned, so they arrive at whatever is current when that build step runs. Docker caches the step, though: `Start-Codeman.sh` rebuilds with the cache, which keeps the versions from the first build until the Dockerfile changes at or above that step or the image is rebuilt with `--no-cache`. They are apt packages owned by root, so they cannot be upgraded from a session; `az extension update --name azure-devops` is the exception and works without a rebuild.
## Local customisation
Compose merges `docker-compose.override.yml` on top of `docker-compose.yaml`. Keep host-specific changes there rather than editing `docker-compose.yaml`, so this repository can be updated without losing them. Both `docker-compose.override.yml` and `docker-compose.override.yaml` are ignored by Git.
`Start-Codeman.sh` names the Compose file explicitly, which disables Compose's automatic discovery of the override file, so the script adds it back when one is present and prints the file it used. Running `docker compose` from this folder without any `-f` option finds it automatically. When passing `-f docker/docker-compose.yaml` from the repository root, add `-f docker/docker-compose.override.yml` as well, or the override is silently ignored.
An override file adds to and replaces individual settings. It cannot delete a key from `docker-compose.yaml`, and Compose concatenates rather than replaces `ports`, so removing a published port still requires editing `docker-compose.yaml`. The example below replaces the restart policy and adds a mount, leaving every other setting in place:
```yaml
services:
codeman:
restart:always
volumes:
- /srv/projects:/srv/projects
```
### Reverse-proxy host allowlist
Codeman rejects any request whose `Host` header is not on its own allowlist - a
DNS-rebinding guard, not a Compose or Docker concern. Loopback, any IP literal,
the configured `--host`, and a few tunnel-provider suffixes are allowed by
default; a reverse-proxied domain is not, and is rejected with
`403 Forbidden: host not allowed` before the request reaches any handler.
Add the domain with `CODEMAN_ALLOWED_HOSTS` in `.env`:
`docker-compose.yaml` forwards it into the container (Compose only passes
through the environment keys it explicitly lists, and this is one of them, with
an empty default so the line is optional in `.env`).
See the application's own `docs/wiki/Remote-Access.md` for the full allowlist
format and the tunnel providers it accepts by default.
## Application data storage
The default configuration uses a host-folder bind mount:
@@ -49,7 +177,7 @@ volumes:
target:/home/${CODEMAN_RUNTIME_USER}
```
Set `CODEMAN_APPDATA_PATH` in `.env` to a directory that the Docker daemon can access. The example value is `/mnt/user/appdata/Coding/codeman`.
Set `CODEMAN_APPDATA_PATH` in `.env` to a directory that the Docker daemon can access. The example value is `/mnt/user/appdata/codeman`.
`CODEMAN_CASES_PATH` is the separate host directory for managed case workspaces. It is mounted into Codeman at the same absolute path, allowing the host Docker daemon to bind it into an isolated case container. Set it to a child directory of `CODEMAN_APPDATA_PATH` unless you deliberately store workspaces elsewhere.
@@ -60,7 +188,7 @@ Set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` when `docker info` reports `SwapLimit=
For an existing installation created by a root-running image, change ownership of the application-data directory before upgrading so the configured `PUID` and `PGID` can read the saved credentials and state:
```sh
chown -R 99:100 /mnt/user/appdata/Coding/codeman
chown -R 99:100 /mnt/user/appdata/codeman
```
Replace `99:100` and the path with the values from your `.env` file.
@@ -69,7 +197,7 @@ Do not replace this bind mount with a Docker-managed named volume when Docker ca
## Static macvlan networking
The default configuration publishes a host port. It does not use `network_mode: host`. To attach Codeman directly to an existing external macvlan network with a static IP address and MAC address, remove the `ports:` section and add the following to the `codeman` service:
The default configuration publishes a host port. It does not use `network_mode: host`. To attach Codeman directly to an existing external macvlan network with a static IP address and MAC address, remove the `ports:` section from `docker-compose.yaml` and add the following to the `codeman` service. The service and network additions can instead be placed in `docker-compose.override.yml`, but the `ports:` removal cannot, as described under [Local customisation](#local-customisation):
@@ -324,6 +324,30 @@ from the session's current state rather than requiring a new transition: the
original turn may be long over. It comes back as
`"delivered": false, "duplicate": true`.
**Wake-on-LAN hosts** (`docs/remote-sessions.md` §Wake-on-LAN): when the session's
remote host has a wake target and is asleep, the non-wait form answers `200` with
`{"buffered": true}` — the bytes are held and flushed after the host is back — or
`{"buffered": true, "dropped": true}` for a chunk over the 4 KB wake buffer, which
is gone (never delivered as a fragment). Both fields are additive to the historical
bare `{}`. With `wait`, the route blocks on the wake instead and answers
`422 OPERATION_FAILED` ("did not come back after a wake-on-LAN request — nothing was
sent") when the host never returns, rather than writing into the stalled pane and
reporting `delivered:true` plus a timeout.
Two endpoints back that flow directly, both scoped to one session's remote host and
both refusing a session that is not remote (`400 INVALID_INPUT`):
| Method | Path | Purpose |
| --- | --- | --- |
| `GET` | `/api/sessions/:id/reachability` | Whether the session's remote host answers SSH right now, plus whether a wake target is configured. Read-only: it never wakes. `{"reachable": true\|false\|null, "wakeConfigured": "mac"\|"command"\|"none"}`, where `null` means the answer is unknown (a proxied host, where a TCP probe proves nothing). |
| `POST` | `/api/sessions/:id/wake` | Wake the host and wait for it to accept SSH again, bounded by the request budget. `422 OPERATION_FAILED` when it does not come back; `400 INVALID_INPUT` with "No wake-on-LAN target configured for this host" when nothing is set. |
⚠️ Waking is deliberately reachable only from an explicit user action (this route, a
session create/attach, or typing into a sleeping session). No watcher, dropped-session
handler or boot-recovery path may wake a host, or a suspended machine would be woken
again seconds after every suspend; `test/remote-wake.test.ts` pins that as an import
fence around `src/remote-wake.ts`.
### Response
All three nest the wait result under `data.wait`, so one client helper works against
@@ -479,6 +503,48 @@ re-captured, or the item acknowledged), `approval:resolved` (`{ id, sessionId, k
`resolution` one of `answered | resolved_in_terminal | superseded |
session_ended | dismissed | expired`).
## Reboot restore
A host reboot takes the tmux server down with it, so every pane dies and the
board comes up empty. At boot Codeman works out which sessions the reboot
destroyed and holds that plan in memory, and these endpoints let a client offer
it to the user. Nothing creates a pane until the user asks: the boot-time reboot
heuristic decides whether to ASK, never whether to act.
Claude-mode sessions only (others carry their conversation id in their own
config object); remote and docker sessions are never offered, because both need
another host or container to be up. The plan is in-memory, so a server restart
drops it and the offer is gone; the conversations themselves are unaffected,
since they live in the CLI's own transcript store and stay reachable from the
Resume list. A plan nobody spends expires after 24 hours.
confirmed?, confirmedContext?, confirmedSwap? }` field (alongside its
normal `caseName`/`mode`/etc. body)
computes the same injection **before** the session exists and launches
directly on the endpoint — no restart, because there was never a
native-backend boot to restart away from. Runs the identical checks as
the dedicated route above (`requiresConfirmation`/`requiresContextWarning`,
same shapes, same per-question `confirmedContext`/`confirmedSwap` retry),
and is refused the same way
for a remote or Docker case. This is what the Run-menu picker uses for
opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP; Claude still uses the
dedicated restart route above (its `--resume`-based restart is far less
jarring than a full relaunch, and folding it into the one-shot path is
separate work — see `docs/custom-model-endpoints-plan.md`).
## CLI management
Read and write the CLI registry (`docs/cli-registry.md`). Every **write** route answers `403 FORBIDDEN` while `cliManagementEnabled` is off (the default), and for a non-admin in multi-user mode. A write that would overwrite a `clis.json` which does not parse, or which has group/world permission bits, is refused with `409 CONFLICT` and a message naming the fix; the file is left untouched.
| `GET` | `/api/clis` | none | Every entry, disabled ones included: `id`, `label`, `shortBadge`, `order`, `kind`, `enabled`, `stock`, `installed`, and `installCommand` for a stock entry. Not gated; a non-admin in multi-user mode gets `[]`. |
| `PUT` | `/api/clis/:id` | `{ enabled }` | Toggle an existing entry, stock or custom. `404` for an unknown id; `400 INVALID_INPUT` when disabling a `kind: 'shell'` entry. |
| `POST` | `/api/clis/:id/install` | none | Run a **stock** entry's install command (never a custom one: `400`). `409 CONFLICT` while an install for the same id is running; `422 OPERATION_FAILED` with the output tail when it fails. Never enables the entry. |
| `POST` | `/api/clis` | `{ id, label, shortBadge, binaries, argv, enabled? }` | Create a custom entry. `409 ALREADY_EXISTS` for a stock id or an existing custom id. `enabled` defaults to `true`. |
| `PUT` | `/api/clis/custom/:id` | `{ label, shortBadge, binaries, argv, enabled? }` | Replace an existing custom entry. An absent `enabled` keeps the entry's current state. `400` for a stock id, `404` for an unknown one. |
| `DELETE` | `/api/clis/:id` | none | Delete a custom entry. `400` for a stock id, `404` for an unknown one. |
## Voice dictation
Browser dictation transcribed through this server's Claude Code login, i.e. the
@@ -18,7 +18,17 @@ Every run mode Codeman can launch — Claude Code, Terminal/Shell, OpenCode, Cod
## The override file
`~/.codeman/clis.json` (instance-scoped through `dataPath()`) holds overrides and custom entries only, never a copy of the stock catalog: `{ "clis": { "<id>": { ...partial entry... } } }`. Objects merge key-wise onto the stock entry, arrays replace wholesale. **The file must be mode 0600**; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you `chmod 600` it. Every reason a file was ignored or an entry dropped is logged once, prefixed `[cli-registry]`, on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read only on restart.
`~/.codeman/clis.json` (instance-scoped through `dataPath()`) holds overrides and custom entries only, never a copy of the stock catalog: `{ "clis": { "<id>": { ...partial entry... } } }`. Objects merge key-wise onto the stock entry, arrays replace wholesale. **The file must be mode 0600**; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you `chmod 600` it. Every reason a file was ignored or an entry dropped is logged once, prefixed `[cli-registry]`, on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read after a change made through CLI management (below).
## Managing CLIs from Settings
App Settings → Agents & CLIs → **CLI management** (`cliManagementEnabled`, default OFF; admin-only in multi-user mode) lists every entry with an installed/not-installed badge and:
- toggles any entry on or off. A `kind: 'shell'` entry cannot be disabled, and the row shows no switch for it. A disabled CLI disappears from the Run menu, the welcome screen and the phone overview, and new session requests for it are rejected.
- installs a missing **stock** CLI by running its shipped install command, after a confirm that names the exact command. Only one install per CLI runs at a time, and the command runs without any `CODEMAN_*` variable in its environment. A custom entry's install command is never executed.
- adds, edits and deletes **custom** entries (id, label, badge, binaries, launch argv). The server re-validates the whole assembled entry through `CliEntrySchema`, so the form cannot bypass the load-time rules.
These are the only writes to `clis.json`. They are serialized, and a file that does not parse or has unsafe permissions is refused rather than overwritten; fix it (or `chmod 600` it) and retry. The HTTP routes are listed in `docs/api-reference.md` under *CLI management*.
## The shape of an entry
@@ -36,7 +46,8 @@ interface CliEntry {
launch: CliLaunch; // the structured argv template
env: CliEnv; // exports, tmux setenv keys, the env-override allowlist
capabilities: CliCapabilities; // what every call site reads instead of the id
// .workDetect?: { promptGlyph, workingLine } — how this CLI's pane shows work
// .workDetect?: { promptGlyph, workingLine, watchingLine?, watchingLines? } — how
// this CLI's pane shows work, and how it shows work it started in the background
overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store
}
```
@@ -45,10 +56,46 @@ interface CliEntry {
### Regexes that come from config
Two capability fields carry a regular expression an override file can set: `discovery.version.regex` and `capabilities.workDetect.workingLine`. Both go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
Three capability fields carry a regular expression an override file can set: `discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`. All three go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
`workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused.
`watchingLine` reads a different row of the same screen. A CLI draws it while work the agent
itself started is still running — Claude prints `⏵⏵ bypass permissions on · 1 monitor · ← for
agents` while a monitor, a backgrounded shell or a cloud session is live. Codeman turns that
into `Session.watching`, and an idle prompt from such a session opens already acknowledged,
so a pane waiting for its own background work never raises an alert a human cannot answer.
Group 1 is the label, and a CLI that declares no pattern reports no background work.
Two CLIs declare such a row today, and they put it in different places. Claude writes its
chip on the last row of the screen, so it keeps the default one-row window and anchors on
the `·` its footer joins items with. Codex pins
`1 background terminal running · /ps to view · /stop to close` ABOVE its composer, which
puts the row third from the bottom once the status line and the composer are counted, so its
entry declares `watchingLines: 3` and matches that row end to end. Both were measured
against live panes rather than read out of a binary, which is the standard for adding a
third.
That label is the one value in the registry that an AGENT can influence, because it comes off
the agent's own screen. Two things keep it honest, and both belong to whoever adds a pattern
for a new CLI. `watchingLabel()` in `session-activity.ts` searches only the last few
non-blank rows, which should be the part of the screen the CLI draws rather than the agent,
and the pattern should anchor on chrome only that CLI can produce. Keep the window as small
as the layout allows, since every row it adds is another row the agent may be able to write.
The label is also ANSI-stripped and length-capped at the source, and every interpolation of
it into markup goes through `escapeHtml()`, since it ends up on a badge and in an approval
card.
The two shipped entries do not sit equally well behind that rule, and the difference decides
what a pattern is allowed to do. Claude's chip is the last row, so its one-row window holds
nothing the agent can write — not even the status line above it, whose command a session
running with permissions bypassed can write into its own `.claude/settings.json`. Codex's row
shares its slot with the last row of the transcript whenever no terminal is running, so a
message ending in that exact line is matched. What keeps that harmless is `hooks: 'none'`: no
hook event from a codex session reaches the approvals inbox, so a forged label costs a wrong
badge and cannot silence an alert. Before giving a CLI both hook signals and a pattern, make
sure its row is one the agent cannot write.
### Three capabilities that must stay independent
`external`, `hooks` and `altScreen` describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. `shell` has no hooks but is **not** an external CLI, so a hooks predicate written as `!isExternalCliMode()` accepted `until=stop` on a shell session and then blocked the caller for their entire timeout. `deepseek` is the mirror image: it IS external and it DOES have hooks.
@@ -102,6 +149,8 @@ It matches four shapes, not one: `mode === '<id>'`, `mode !== '<id>'`, `case '<i
The allowlist is not a formality. If a branch is about what a CLI can DO it belongs in `CliCapabilities`; the entries that remain are things that are not CLI-behaviour branches at all — chiefly the legacy per-mode `<Mode>Config` objects on `POST /api/sessions`, which are a fact about the public HTTP API rather than about any CLI, plus a few documented cases where `mode === 'claude'` is genuinely the right question (Read My Mind reads Claude's _own_ transcript, so a capability there would be actively wrong).
`test/frontend-cli-no-id-branching.test.ts` is the same guard for the two frontend files the CLI registry's Run-menu consolidation touches, `session-ui.js` and `mobile-overview.js` — deliberately not the rest of `src/web/public/`, whose per-CLI rules stay out of scope for now (see "Fields declared for later" below). Its allowlist keys on `<file>::<expression>` with no line number, since a single unrelated edit to a contended file would otherwise shift every subsequent line and make every entry go stale at once, and each entry additionally carries the exact number of approved call sites — a bare key would let a brand-new branch reusing an already-approved expression land unreviewed. Its comparison shape differs from the backend guard's in one respect: the left-hand side may be any identifier, not only one named `mode`, `id` or `agentType`, because the review of #458 found `const m = this._runMode; if (m === 'codex')` slipping past the named form while the scanned file already filters with `(m) => m !== 'shell'`.
## Two namespaces called `param`
`launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a **launch param**. The **legacy wire field** a param arrives as is a separate namespace, and `launch.legacyConfigAliases` is the only bridge between the two.
@@ -110,9 +159,9 @@ This matters because it is invisible when it is wrong. `capabilities.privilegedP
## Fields declared for later
`shortBadge`, `accent`, `capabilities.echo`, `capabilities.wheelForward`, `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` are **declared but not yet read**. They all describe frontend behaviour, and the frontend is deliberately untouched here: `app.js`, `terminal-ui.js` and `styles.css` keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
`accent`, `capabilities.echo`, `capabilities.wheelForward`, `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` are **declared but not yet read**. (`shortBadge` was on this list until the CLI management list in Settings started showing it.) They all describe frontend behaviour, and the frontend is deliberately untouched here: `app.js`, `terminal-ui.js` and `styles.css` keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
Treat those values as **transcribed, not authoritative** — nothing enforces that `echo.policy` matches `_updateLocalEchoState`'s fallthrough, or that `accent` matches the gradient CSS paints, so re-measure before wiring one up. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; `test/cli-registry-no-id-branching.test.ts` pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
Treat those values as **transcribed, not authoritative** — nothing enforces that `echo.policy` matches `_updateLocalEchoState`'s fallthrough, so re-measure before wiring one up. `accent` is the one exception: it was measured against styles.css on 2026-09-21 (method in the comment above `CLAUDE` in `stock.ts`), though nothing keeps it in step with the CSS either. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; `test/cli-registry-no-id-branching.test.ts` pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
`overlays.credStore` is in the same category, for a sharper reason: the Docker credential-seeding path still reads its own `CRED_STORES` table, because this shape allows ONE store per CLI and the live table needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek declares none here even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is simultaneously the worst thing here to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest.
| `claude` | Env vars: `ANTHROPIC_BASE_URL`, `ANTHROPIC_API_KEY`, `ANTHROPIC_DEFAULT_SONNET_MODEL`/`_HAIKU_MODEL`/`_OPUS_MODEL` (all set to the chosen model/deployment name) | **Verified end-to-end** against a real llama-swap server — a real "hello world" reply came back. ⚠️ Non-interactive (`-p`) invocations also fire an async session-title-generation call that reuses `ANTHROPIC_DEFAULT_HAIKU_MODEL` and validates it against Claude Code's OWN internal recognized-model list, printing `[claude-code:unrecognized_model]` and, in `-p` mode, hanging the whole invocation rather than just warning. `--settings '{"autoTitle":false}'` does NOT stop this (confirmed); `--bare` does (the warning still prints, but the real prompt runs) — but `--bare` ALSO disables hooks, LSP, plugin sync, and CLAUDE.md auto-discovery, so it is only safe for the standalone one-shot test script, NEVER for a real interactive Codeman session (which depends on hooks for idle detection, trust-dialog auto-accept, etc. — see the External CLI modes section of CLAUDE.md). Whether an INTERACTIVE claude session with a custom model hits the same hang (vs. just a background warning) is untested and should be checked before calling chunk 5/6 done for claude |
| `opencode` | `OPENCODE_CONFIG_CONTENT` env var (already a registry mechanism, `stock.ts:342`) holding a JSON blob: `{"provider":{"custom":{"options":{"baseURL":...,"apiKey":...},"models":{"<name>":{}}}},"model":"custom/<name>"}` | **Verified by user** |
| `codex` | TOML `config.toml`: top-level `model = "<id>"` + `[model_providers.custom]` (`base_url`, `env_key` naming an env var the real API key rides in — never a literal TOML field, since codex's schema has no such field). Written to an isolated dir via `CODEX_HOME` (`stock.ts:405-415`) so the user's own `~/.codex/config.toml` is never touched | **Config STRUCTURE verified** against a real codex binary (an earlier `[model].default` table shape was rejected: "invalid type: map, expected a string" — caught live). **Protocol picture more nuanced than a flat break, re-verified live twice on 2026-09-17 against a llama-swap deployment that DOES answer `/v1/responses`** (an earlier test's `Reconnecting...`/`high demand` failure does not reproduce against every llama-swap setup): a plain, no-tool-call chat turn (`codex exec 'reply with just OK'`) returned a real reply. But a real tool-call attempt (`run the shell command: echo hello`) came back as an `agent_message` TEXT item — the tool-call JSON printed as the model's answer, not a `function_call` item codex would actually execute (confirmed via `codex exec --json`'s raw event stream: `item.completed`/`agent_message`, never `function_call`). Since tool execution is what makes codex a coding agent at all, this remains **not usable for real work**, just with a different, more specific failure mode than previously documented — still do not present this as working. Separately, EVERY custom-endpoint codex session also prints `warning: Model metadata for '<id>' not found. Defaulting to fallback metadata...` on launch (confirmed harmless — the successful plain-text reply above still had it): codex's per-model metadata (reasoning tiers, system-prompt templates, context-window figures) comes from `models_cache.json`, a LOCAL CACHE of OpenAI's own hosted model catalog that a custom model can never appear in by construction. No config.toml override exists for it, and the isolated `CODEX_HOME` never gets a `models_cache.json` written into it at all (confirmed: inspected a live, actively-used isolated dir — codex evidently can't reach OpenAI's catalog endpoint for this session and just falls back silently every time, with no file left behind to fix or clean up). Fabricating a fake catalog entry to suppress the warning would mean copying the _shape_ of OpenAI's own proprietary schema — including their real per-model system-prompt content, visible in a genuine `models_cache.json` — for a warning confirmed to have no effect on the actual (broken) tool-calling outcome; not worth building |
| `gemini` | Env vars `GOOGLE_GEMINI_BASE_URL` + `GEMINI_API_KEY` + `GEMINI_MODEL`; CLI needs a restart to pick them up | **Confirmed BROKEN against llama.cpp/llama-swap, unresolved after real investigation.** Setting `GOOGLE_GEMINI_BASE_URL` makes gemini-cli internally select an `AuthType.GATEWAY` auth path (undocumented — inferred from behaviour) with validation requirements distinct from every normal auth mode; a real run against llama-swap fails with `Invalid auth method selected` regardless of what key/format is supplied. Tried and all failed: a Google-format dummy API key, `GOOGLE_GENAI_USE_VERTEXAI=false`, a `GEMINI_DEFAULT_AUTH_TYPE` override, and hand-writing `settings.json` directly. `--skip-trust` was a real, separate fix (without it a trust-folder check silently overrides `--approval-mode yolo` back to `default`) but does not touch this auth failure. Documented as an open gap, not shipped as working — the registry entry and injection code exist and are exercised by the test script, but end-to-end gemini support needs upstream investigation of `GATEWAY` AuthType before it can be called done |
| `pi` | Config file `~/.pi/agent/models.json` with a custom provider whose `models` is an **array** of `{id}` objects (not an object keyed by id) plus `authHeader: true`. Redirected via the child process's own `HOME` env var, isolated per test/session — **not**`PI_CONFIG_DIR`, which does nothing for pi (grepped pi's entire bundled JS source: the string appears nowhere) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. Two real bugs found and fixed before this worked: (1) `PI_CONFIG_DIR` is not read by pi at all — pi hardcodes `~/.pi/agent/models.json` with no dedicated override, so the actual redirect has to be the child process's `HOME`; (2) `models` must be an array of `{id}` objects per pi's own bundled `docs/models.md`, not an object keyed by model id (silently loaded zero models). Also requires an explicit `--model custom/<id>` on invocation — without it pi falls back to its own default provider and fails with "No API key found for the selected model" |
| `grok` | TOML `config.toml`: a fixed `[model.codeman-custom]` block (`base_url`, `env_key` naming an env var the key rides in, never a literal TOML field) written to an isolated dir via `GROK_HOME`. Invoked with `-m codeman-custom` | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back. The ORIGINAL recipe in this table (env vars `GROK_BASE_URL`/`XAI_API_KEY`/`GROK_MODEL`) was flat-out **wrong**, not just unverified: it produced "Not signed in" against a real binary. Grok's real mechanism, confirmed against xAI's own docs and a live binary, is a `config.toml` with a `[model.<name>]` block, redirected via `GROK_HOME`; the key still rides as an env var (`XAI_API_KEY` via `env_key`), just referenced from the TOML rather than read directly |
| `deepseek` | Reuse the **existing**`DEEPSEEK_BASE_URL` + `DEEPSEEK_API_KEY` keys (already declared in `stock.ts`), now with `appendV1Suffix: true` (see confidence). Only `DEEPSEEK_BASE_URL` is in `privilegedEnvKeys` — `DEEPSEEK_API_KEY` deliberately stays clamp-exempt, since a non-granted owner supplying their OWN key removes privilege rather than granting it (adding it to the clamp list was a real regression, caught by `test/deepseek-mode.test.ts` and fixed before merge). No model-selection var — dsh model is a profile composition entry, not a flag/env var | **Root cause of the original `HTTP_404` found and fixed, by reading dsh's own bundled source — the same bar pi/grok's fixes were held to.** Installed `@deepseek-ai/dsh` (all its real published dependencies) into a scratch directory purely to read `@deepseek-ai/dsh-llm-deepseek/lib/index.js`: it builds its request as `fetch(\`${connection.baseURL}/chat/completions\`, ...)`with`baseURL`read straight from`DEEPSEEK_BASE_URL`(or defaulting to DeepSeek's real public API root,`https://api.deepseek.com`, which also carries no `/v1`) — no `/v1` insertion of dsh's own, unlike the OpenAI-SDK convention this recipe originally assumed. llama-swap/llama.cpp only ever serves the OpenAI-conventional `/v1/chat/completions`. Confirmed live: `POST <baseUrl>/chat/completions` → `404`, `POST <baseUrl>/v1/chat/completions` → `200`, on the exact same endpoint — and dsh's own error-message template, `DeepSeek API error (HTTP ${status})`, reproduces the originally reported `dsh: HTTP_404: DeepSeek API error (HTTP 404)` precisely. Fixed by adding `appendV1Suffix` (env kind only, deepseek's entry alone — claude/gemini must NOT get it, since claude was already confirmed working against the unmodified `baseUrl`), which runs `endpoint.baseUrl` through the same `withV1Suffix()` helper `configDir`-kind CLIs already use. ⚠️ Not yet re-run end-to-end with a real `dsh` binary — no install available in this environment (no npm-installed CLI binary in `PATH`, and the `codeman-test-picker` container doesn't bundle it either); the fix is source-confirmed and live-verified at the HTTP level, but a genuine "hello world" reply through `dsh` itself is the remaining step before promoting this to **verified** alongside claude/opencode/pi/grok/omp |
| `omp` | Config file `~/.omp/agent/models.yml` with the same array-shaped `models` + `authHeader: true` fix as pi. Redirected via `HOME`, same reasoning as pi (`PI_CONFIG_DIR` does not relocate omp's config either, despite an earlier CLAUDE.md note claiming it does) | **Verified end-to-end** against a real llama-swap server — real "hello world" reply came back, after applying the same two fixes as pi (array-shaped `models`, `HOME`-redirect instead of `PI_CONFIG_DIR`) plus an explicit `--model custom/<id>` on invocation. Unverified against omp's own official docs (none are bundled in the install), but empirically confirmed working live |
| `antigravity` | No CLI/env/config mechanism found — Antigravity's docs describe only a GUI settings panel, and explicitly say a custom endpoint "cannot currently" become the core reasoning model. **Not implemented**; toolbar entry stays disabled for this mode with an explanatory tooltip | No known mechanism |
Everything web-researched-but-unverified gets implemented but must be
smoke-tested against real installs of those CLIs before being called done —
call this out explicitly when implementing, don't just ship on faith.
**Cloud-endpoint specifics** to keep in mind per recipe above: an Azure AI
Foundry-style endpoint typically wants the API key in an `api-key` header
rather than (or in addition to) `Authorization: Bearer`, and its "model" is
often a deployment name rather than the underlying model family name — the
discovery step (`GET /v1/models`) still works the same way against Azure AI
Foundry's OpenAI-compatible endpoint shape, but a user may need to type the
deployment name manually if it isn't returned as expected.
## Architecture
### 1. Registry: new `capabilities.customModelInjection` field
Extend `src/config/cli-registry/types.ts` / `schema.ts` with a discriminated
- `scripts/test-local-llm-harnesses.ts` (already added, this branch; run via `npx tsx`) — the standalone real-CLI-and-real-endpoint smoke test, supporting any `--base-url` (local or cloud). Dynamic: derives its harness list and every env var/config it injects from the live CLI registry + `buildCustomModelInjection()` rather than a second hand-maintained copy — only the one-shot invocation flags (`ONE_SHOT` table) are CLI-specific info the registry doesn't model and stay hand-maintained
- `docs/custom-model-endpoints.md` (new) + a CLAUDE.md pointer bullet under External CLI modes / envOverrides
## Mock-server validation strategy (CI-runnable, no real CLI binaries needed)
Spawning nine real CLI binaries in CI isn't realistic, and neither the author's
llama.cpp box nor a real cloud subscription can be a CI dependency. So the
injection _logic_ gets a tier of automated coverage that sits between the
pure unit tests and the live manual checks in Verification:
1. **`test/fixtures/mock-openai-server.ts`** — a small in-process HTTP
server (plain `http.createServer`, no external deps, port picked per the
existing `const PORT = 3150+` convention) that:
- Serves `GET /v1/models` → a fixed fake model list (`{data:[{id:'qwen3'},...]}`),
@@ -81,6 +81,8 @@ Antigravity (`agy`) and Grok (`grok`) are the two CLIs not installed from npm (G
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md). Grok is seeded per-file for the same reason (`auth.json`, `config.toml`, `pager.toml` out of `~/.grok`, which also holds `sessions/`, `memory/` and the ~160MB binary under `downloads/`), with the same consequence for `grok -c`. See [`grok-integration.md`](./grok-integration.md). OMP is the one CLI in this family where `sessions/` is the EXCEPTION rather than the rule: `~/.omp/agent/{config.yml,mcp.json,models.yml,settings.yml}` are seeded per-file (the dir also holds SQLite caches and `terminal-sessions/`), but `~/.omp/agent/sessions/` is shared RW like codex's, not seeded, because Codeman reads it host-side for history recovery and `--resume` pinning. See [`omp-integration.md`](./omp-integration.md).
The image can also carry the GitHub CLI (`gh`) and the Azure CLI (`az` + the `azure-devops` extension, in `AZURE_EXTENSION_DIR=/opt/az-extensions` so it stays out of the seeded HOME), wired into the system git config as credential helpers for github.com and dev.azure.com / *.visualstudio.com, exactly as in `docker/server.Dockerfile`. Their sign-ins are seeded per-FILE like pi's: `~/.config/gh/{hosts.yml,config.yml}` and `~/.azure/{azureProfile.json,msal_token_cache.json,service_principal_entries.json,clouds.config,config}`, never `~/.azure`'s logs, command index or extensions. A token kept in a desktop keyring, or in the encrypted MSAL cache az uses on Windows/macOS, is not in those files and does not carry. None of the three is version-pinned; the `--no-cache` rebuild recommended above is also what refreshes them. Both CLIs are opt-in and OFF by default: `CODEMAN_AGENT_IMAGE_INSTALL_GH=1` / `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` in the environment of `scripts/build-agent-image.mjs`, or of the Codeman server for its own auto-build (in the Compose deployment, `environment:` in `docker-compose.override.yml`), become the `CODEMAN_INSTALL_GH` / `CODEMAN_INSTALL_AZ` build args and put that CLI, its extension and its helper entry into the image. Unset passes nothing, so a default build's argv is unchanged and the image has neither. The sign-in seeds follow the same switches, read when a case container is created: `.config/gh` only with `CODEMAN_AGENT_IMAGE_INSTALL_GH=1`, `.azure` only with `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` (`enabledByEnv` in `CRED_STORES`), never merely because the files exist. Seeds are create-time mounts and deliberately not part of the config hash (hashing them would trip the drift gate for every case), so an existing case container picks them up only when it is recreated.
## Quickest path: one-click "Run in Docker"
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
@@ -6,6 +6,8 @@ For the Compose configuration, environment settings, storage migration, and macv
The image includes Claude Code, Codex, Gemini CLI, and OpenCode. Authenticate a CLI from its Codeman session; credentials are never baked into the image.
It can also include the GitHub CLI (`gh`) and the Azure CLI (`az`) with the `azure-devops` extension, wired in as Git credential helpers, so Clone Repo and `git clone` reach private GitHub and Azure DevOps repositories once they are signed in. Both are off by default; [Turning them on](../docker/README.md#turning-them-on) shows the `docker-compose.override.yml` settings.
## Prerequisites
- Docker Engine or Docker Desktop with Docker Compose v2
@@ -15,7 +17,7 @@ The application container mounts the Docker daemon socket so Codeman can create
## Start
Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/Coding/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes.
Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes.
```sh
cp docker/.env.example docker/.env
@@ -33,12 +35,14 @@ On Linux, run the stack with the start script. It determines `PUID` and `PGID` f
bash docker/Start-Codeman.sh
```
On other platforms, run Compose directly. `PUID` and `PGID` default to `1000:1000`; set them in `docker/.env` when the application-data directory has a different owner.
On other platforms, run Compose directly. `PUID` and `PGID` default to `1000:1000`; set them in `docker/.env` when the application-data directory has a different owner. Naming the file with `-f` disables Compose's own discovery of `docker/docker-compose.override.yml`, so add a second `-f` for it when you keep one (see `docker/README.md`, Local customisation).
```sh
docker compose --env-file docker/.env -f docker/docker-compose.yaml up --build -d
```
The container starts as root, corrects the ownership of a bind source the daemon had to create, and drops to `PUID:PGID` with `setpriv` before Codeman starts; the capabilities that needs are declared in `docker/docker-compose.yaml` and named by the entrypoint when a compose file written elsewhere lacks them.
Open `http://localhost:3000` and sign in with the username and password from `docker/.env`.
## Operations
@@ -65,7 +69,7 @@ If that directory was created by an earlier root-running image, change its owner
Codeman updates itself from **App Settings → Updates**, as it does on a bare host. The checkout mounted at `/opt/codeman` is the same directory Compose builds from, so the update's `git checkout` and rebuild land on the host and survive container recreation; the restart is the server exiting, which `restart: unless-stopped` turns into a relaunch on the new build.
That applies application code only. A release that changes `docker/server.Dockerfile`, `docker/docker-compose.yaml`, or adds a key to `docker/.env.example` needs the image rebuilt or the container recreated, which a container cannot do to itself. The updater detects each case and refuses with a message naming what changed; run `docker/Start-Codeman.sh` on the host to apply those.
That applies application code only. A release that changes `docker/server.Dockerfile`, `docker/docker-compose.yaml`, or adds a key to `docker/.env.example` needs the image rebuilt or the container recreated, which a container cannot do to itself. The updater detects each case and refuses with a message naming what changed; run `docker/Start-Codeman.sh` on the host to apply those. For a major update, or a base-image change `Start-Codeman.sh` does not fully pick up, `docker/Update-Codeman.sh` rebuilds with no layer cache and clears the two build-artefact volumes before handing off to it (see "Major updates" in `docker/README.md`).
`CODEMAN_REPO_PATH` overrides which checkout is mounted. It defaults to the compose project's parent directory, so it normally needs no setting. Point it at a directory that is not a git checkout and in-app updates are reported as unavailable.
| `docker/server.Dockerfile` | `docker/Start-Codeman.sh` on the host |
| `docker/docker-compose.yaml` | `docker/Start-Codeman.sh` on the host |
| New key in `docker/.env.example` | Add it to `docker/.env`, then `Start-Codeman.sh` |
| A major update, or a Node base-image bump | `docker/Update-Codeman.sh` on the host (no-cache rebuild + fresh build volumes) |
The in-app updater detects all three of the bottom rows itself and refuses with a
The in-app updater detects the three middle rows itself and refuses with a
message naming what changed, so you never have to work out which case you are in.
`Update-Codeman.sh` is the heavier option for when `Start-Codeman.sh` is not
enough: see "Major updates" in `docker/README.md`.
## Why the container needs its own path
@@ -59,6 +62,7 @@ unchanged. The container path is a new `SupervisorKind`, not a new updater.
| `CODEMAN_RESTART_BY_EXIT=1` | The Compose file's declaration of that policy, so the updater may exit even with no Docker socket. |
| Toolchain + devDependencies in the image | Lets `npm install` and `npm run build` run inside the container. |
| `docker-env-applied.json` | Fingerprint baseline, written by `Start-Codeman.sh` on every start. |
| `docker-build-source.json` | What HEAD/`package-lock.json` the build artefact volumes currently reflect. Written by both `Start-Codeman.sh` and this in-place update, so the two agree on whether those volumes are stale. |
### Why build artefacts are in named volumes
@@ -72,6 +76,20 @@ Docker seeds an empty named volume from the image, so the first start inherits t
image's already-built `node_modules` and `dist` and pays no bootstrap cost.
`docker compose down -v` is the supported reset: the next start re-seeds them.
That seeding-only-while-empty behaviour has a second, less obvious edge: it also
means a plain `docker compose build` triggered from OUTSIDE the container (for
example `Start-Codeman.sh`, after a `git pull` done by hand rather than through
this in-app updater) produces a fresh image whose freshly-built `dist`/
`node_modules` then sit unused behind the volumes' OLD content — the container
comes back up looking unchanged. `Start-Codeman.sh` detects this by comparing the
checkout's current HEAD and `package-lock.json` hash against `docker-build-source.json`,
and clears just the affected volume(s) before its own `--build` if they moved.
This in-place update writes that same file after a successful build precisely so
that comparison does not fire on stale information: without it, the next plain
`Start-Codeman.sh` run would see the HEAD this update just checked out, not
recognise it as already accounted for, and wipe the volumes this update just
correctly rebuilt right back to the OLDER image.
### Why the runtime image carries a build toolchain
`npm run build` is `tsc` plus `esbuild`, both devDependencies, so the image no
@@ -209,7 +227,10 @@ the host and the in-app path works from then on.
**Resetting the build artefacts** — `docker compose down -v`, then
`Start-Codeman.sh`. This discards the named volumes and re-seeds them from a fresh
image.
image.`docker/Update-Codeman.sh` scripts the same reset by default for the two
build-artefact volumes (`codeman-node-modules`, `codeman-dist`) only, plus an
unconditional `--no-cache` rebuild, which a plain `Start-Codeman.sh` run does not
force on its own. See "Major updates" in `docker/README.md`.
| **B. Rename the node** | `https://codeman-tnode.tailf80371.ts.net` | `tailscale set --hostname codeman-<host>` (operator or root) | Renames the machine tailnet-wide: ssh targets, other serve URLs, the admin console entry. Tailscale de-dups a clash as `-1`. The cert follows the new name. | **Opt-in, default NO everywhere** (owner decision 2026-09-20: the machine is used for other things, so a bare Enter never renames it). The proposal was YES when the installer itself had just joined the tailnet; rejected. |
| **C. Tailscale Service** | `https://codeman.tailf80371.ts.net` | tailscale >= 1.86 on the host; the host must have a **tag-based identity** ("You cannot use a device authenticated with a user account as a Service host"); the service is defined in the admin console first; the host is then approved there (or via `autoApprovers.services`). Public beta since 2025-10-28, all plans. | Re-authenticating a personal machine as a tagged node changes its identity (SSH ACLs, user attribution). Known daemon quirk: approval is not picked up until `serve clear` + re-advertise (tailscale/tailscale#18821). | **Detect and hint only** in this round. The maintainer's own node has `Self.Tags: null`, so it could not host one without re-tagging. Worth a real flow once someone with a tagged fleet asks. |
| **D. Sub-path** | `https://tnode.tailf80371.ts.net/codeman` | `tailscale serve --bg --set-path /codeman 3000` plus `--base-url /codeman` on the server | Codeman runs under a prefix. Hooks are unaffected (they hit the raw port with no prefix, which `rewriteUrl` already tolerates). Serve forwards the prefix unchanged, which is exactly the shape `--base-url` was built for. | **The answer when `:443` root is already taken.** Replaces today's replace-or-nothing prompt. |
| **E. Second port** | `https://tnode.tailf80371.ts.net:8443` | `tailscale serve --bg --https=8443 3000` | Port in the URL; the beta-preview recipe already uses this. | Fallback when the user rejects D. |
| Funnel (public internet) | `https://tnode.tailf80371.ts.net` from anywhere | `tailscale funnel` | Public exposure; different risk class. | **Out of scope**, as before. Docs only, with the password warning. |
Sources: Tailscale Services docs (`tailscale.com/docs/features/tailscale-services`), the
| `src/tailscale-status.ts` (new) | Pure parser for the two JSON shapes + the IO wrapper; unit-tested against captured `serve status --json` fixtures (root, path, port, foreign target, none) |
| `src/config/dependency-registry.ts`, `src/utils/dependency-checker.ts` | `tailscale` doctor row |
| `test/install-sh-invariants.test.ts` | Extend: flags documented in the header, no `serve reset`, no `funnel`, no `--service` advertise, every serve mutation goes through `ts_cmd_serve`, rename happens before serve in `main` (static order check) |
| `.github/workflows/ci.yml` | The bash 3.2 step additionally sources the script with stubbed `ts_cmd`/`ts_cmd_serve`/`read_reply` and drives `ask_everything` through all three answers and the 443-occupied menu |
@@ -251,6 +251,278 @@ unreachable host answers "unknown", which also means do not revive. The answer
is cached per session and cleared whenever the pane is next seen alive, so a
stale `true` from one transport drop can never revive the NEXT clean exit.
## File access over SSH
A remote case's `workingDir` is an absolute path on the **remote** host
(`Session.workingDir = RemoteCase.remotePath`), so the file routes cannot use local
`fs`: a local `realpathSync` on a remote-only path fails by construction, which is why
previewing a file used to answer `404 File not found` for a case that was working
perfectly (#415). `src/remote-files.ts` is the one module that reads remote bytes,
and it follows the same rule as the launch path: every ssh command line comes from
`buildSshConnectionArgs()` — **never** a hand-built ssh line.
| Request | What happens |
|---------|--------------|
| `GET /api/sessions/:id/file-raw` | Streamed over `ssh` (`cat`, or `tail -c +N \| head -c L` for a `Range`); the same 200/206/416 contract as a local file, so `<video>`/`<audio>` seeking works |
| `GET /api/sessions/:id/file-content` | `cat` into memory, capped by the existing text limit; `edit=1` answers `400` (see below) and `editable` is always `false` |
| `PUT /api/sessions/:id/file-content` | `400` before any path is looked at: the guard sits AHEAD of the local path validation, because with a same-named directory on the Codeman host (an `sshfs` mount) the write would otherwise land on the local twin |
| `GET /api/sessions/:id/file-preview` | Non-office files redirect to `file-raw` (which works remotely); docx/pptx answer `400` |
| `GET /api/sessions/:id/file-thumbnail` | `400` for remote files |
| `POST /api/sessions/:id/attachments` | Registers an absolute path that lives on the **remote** host (a clicked link pointing outside the case directory) by probing it there |
| `GET /api/sessions/:id/attachments/:attachmentId/raw` | Streams the registered remote file over ssh, same 200/206/416 contract; `preview` (office) and `thumbnail` answer `400` |
| `GET /api/sessions/:id/attachments/:attachmentId`, `GET …/attachments` (history) | Size/mtime/existence resolved over ssh, so a remote entry is not reported `missing`; the history list resolves EVERY entry in one batched probe, never one connection per entry |
⚠️ The attachment route is the one a clicked path takes when it is **outside** the case
directory (a remote `/tmp` scratchpad capture, a screenshot elsewhere in the home dir):
the frontend's `_isExternalPreviewPath()` sends every absolute path that is not under
`workingDir` there, so fixing only `file-raw` would leave exactly that half broken.
Guard order is deliberately **the same as locally**, and the checks are not weakened
by the transport:
1. Ownership (`findSessionOrFail` / the scope helper) — unchanged.
2. Lexical containment of `workingDir + path` — a `../` escape is refused before any
connection is opened.
3. ONE ssh round trip that returns `realpath`**and**`stat` for the path **and** the
workspace root (`remoteProbePaths`). Resolving the root remotely is what keeps the
boundary honest for a symlinked `remotePath`. The probe uses `readlink -f` when
available; on a host without it (macOS before 12.3) a POSIX fallback canonicalizes
the directory chain with `cd -P`/`pwd -P` and then follows the LAST component with
plain `readlink` for a bounded number of hops. ⚠️ **The fallback fails closed**: a
path it cannot fully resolve (a loop, a `readlink` failure, the hop cap) is reported
as unresolvable and answers 404, never as its own unresolved string. An earlier
version resolved only the directory chain, so `ws/notes.txt -> ~/.ssh/id_rsa` passed
containment under the link's own path while `cat` followed it to the key.
Records come back NUL-separated and index-keyed (`<index>|kind|size|mtime|realPath`,
after a leading NUL that fences off any login banner), so a filename containing a
newline cannot shift the alignment.
4. Containment of the remote realpath against the remote root. The sensitive-path
blocklist then applies on whichever routes already apply it locally (`/api/download`,
attachment registration, edit mode — where resolving symlinks first is what makes it
meaningful); the remote branch neither drops a guard the local path has nor invents a
stricter one. One entry of that blocklist is host-bound by construction: the three
home-anchored members (`~/.claude.json`, `~/.claude/settings.json`,
`~/.claude/settings.local.json`) are compared against the **Codeman host's** home
directory, so they do not match a remote home at a different path. Everything else in
the list is depth-anchored (`/.ssh/`, `/.aws/credentials`, `/.claude/.credentials.json`,
`/etc/shadow`, ...) and applies to a remote path unchanged.
5. Size cap (`CODEMAN_MAX_DOWNLOAD_BYTES`) applied to the **remote** size, before the
body is requested.
The path arrives from the browser (`?path=`) and is interpolated as a single
`shellescape`-quoted token, in a command that is itself shellescaped into the ssh
line; `BatchMode=yes` means a host needing a passphrase fails fast instead of hanging.
A failed connection is reported as **502** with the remote reason — never a 404, which
used to make an unreachable host look like a typo in the agent's output. The reason is
the first stderr line, the timeout, or the exit code; never Node's `Command failed: …`
message, which would carry the identity-file path and the probe script into the body.
**Connections are bounded.** Every probe and buffered read runs through a small global
semaphore (`src/remote-ssh-limiter.ts`, default 4, `CODEMAN_MAX_REMOTE_FILE_SSH`), the
attachment-history list resolves its whole history in one batched probe instead of one
handshake per entry, and probes are chunked at 40 paths per round trip. Terminal output
in a remote session is written on the remote host, so a prompt-injected agent printing
hundreds of `codeman://attach` links used to make the server fork one `ssh` per link,
each holding a 20 s probe timeout, and a 100-entry history re-listed on every
Only devices on your tailnet can reach it; Tailscale handles identity and
terminates TLS with a real Let's Encrypt certificate (so PWA install and web
push work). No app password and no `0.0.0.0` bind required. (This is the
maintainer's production setup.) `CODEMAN_TAILSCALE=1`presets the choice for
automation; the installer never runs `tailscale serve reset` and never touches
serve mappings other than `443 -> Codeman's port`.
maintainer's production setup.) `CODEMAN_TAILSCALE=1`or `--tailscale` presets
the choice for automation. When `:443` on the node already belongs to another
app, the installer mounts Codeman under `/codeman` (`tailscale serve --set-path`
plus `--base-url`, which keeps the loopback bind and the same host guard) or on a
second port rather than replacing it. The installer never runs `tailscale serve
reset`, never touches serve mappings other than the one it created, never opens a
`tailscale funnel` (public internet, a different risk class) and never advertises
a Tailscale Service. Renaming the node (`--name`, `install.sh name`) is opt-in
and defaults to no, because the tailnet name is also the machine's SSH identity.
### B. Authenticated cloudflared tunnel + password
@@ -489,7 +497,7 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never**`--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, five seeded files from `~/.pi/agent`, and three from `~/.grok`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, five seeded files from `~/.pi/agent`, three from `~/.grok`, and, only when their opt-in switches `CODEMAN_AGENT_IMAGE_INSTALL_GH` / `_AZ` are `1`, `~/.config/gh/{hosts.yml,config.yml}` and the sign-in files from `~/.azure`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
@@ -508,15 +516,17 @@ Full feature guide: [`docker-cases.md`](docker-cases.md).
- **Auth is a parallel branch** (`middleware/auth.ts`) that leaves the single‑user path untouched: per‑user scrypt verify (`timingSafeEqual`, timing‑equalized against user enumeration), identity‑carrying cookies, a per‑username failure bucket (a botnet can't brute one account across IPs; one NATed user can't lock out the rest), and a `mustChangePassword` lockbox. The hook‑secret loopback bypass, host guard, and Origin/CSRF guard are unchanged (hooks authenticate the INSTANCE, not a user).
- **Ownership is enforced server‑side only** and fails closed: `req.authUser` (a synthetic admin in single‑user), `findSessionOrFail` returns NOT_FOUND (never 403) for a foreign session, list/SSE/WS/file‑preview/search all filter by `session.owner`, and SSE routing defaults session‑scoped events to their owner (unresolved owner → withheld). The load‑bearing rule is **non‑admin `workingDir` confinement**: a non‑admin's session/one‑shot working dir must realpath‑resolve inside `~/codeman-users/<name>/cases`, checked BEFORE any disk write.
- **Privileged actions are a one‑bit grant** (`canBypassPermissions`, default off): only granted users (and admins) get `--dangerously-skip-permissions` (others are silently downgraded to `--permission-mode auto`), shell‑mode sessions, cron `launchCommand`, and other CLIs' bypass flags. Machine‑level resources (remote/Docker host definitions, tunnel, self‑update, settings writes) are admin‑only.
- **Clone Repo does not lend the server's git sign-in to non-admins.** A clone writes only inside the caller's own case space, so it is not admin-gated, but the server account's git credential helpers (the Docker image's opt-in `gh`/`az` helpers, or any `gh auth setup-git`) are shared by every user. A non-admin's clone and preflight therefore run with `git -c credential.helper=`, which empties the helper list including the URL-scoped entries (`cloneWithoutCredentialHelpers` in `case-routes.ts`, argv pinned in `test/git-clone.test.ts`). This closes the Clone Repo path only: the account's SSH keys still apply to an `ssh://` URL, and a non-admin's agent sessions run as the same account, consistent with the first bullet above. Docker cases are a second route to the same sign-in: with `CODEMAN_AGENT_IMAGE_INSTALL_GH`/`_AZ` on, a non-admin's Docker case with credential seeding on (the default) receives a copy of the server account's `gh`/`az` sign-in, exactly as it receives the Claude and Codex credentials.
- **Admin actions are audited** append‑only to `~/.codeman/admin-audit.jsonl` (acting admin, action, target, IP). Passwords set by an admin create/reset are one‑time (returned once, force change). Under Basic auth, `logout` only truly ends QR‑issued sessions — to lock someone out, disable the account or reset the password (a proper login form is a deferred Phase 6).
---
## 10b. Web tabs (dashboard proxy)
A saved dashboard URL renders as a tab, served through Codeman's own origin at `/webview/<capability>/`. User guide: [`web-tabs.md`](web-tabs.md). Three properties carry the security weight:
A saved dashboard URL renders as a tab, served through Codeman's own origin at `/webview/<capability>/`. User guide: [`web-tabs.md`](web-tabs.md). Four properties carry the security weight:
- **The proxy is exempt from cookie auth and the Origin/CSRF guard, and that is deliberate.** The iframe is sandboxed without `allow-same-origin`, so it is opaque‑origin: its requests are cross‑site, meaning the `SameSite=lax` session cookie is never attached and its writes and WS upgrades arrive with `Origin: null`. The credential is instead a 192‑bit capability in the path, minted only by an authenticated `POST /api/webviews/:id/open`, held in memory (a restart invalidates every one), rolling TTL, bound to the minting user, and granting nothing but "relay bytes to this one saved URL". ⚠️ **The Host allowlist is NOT bypassed**, so DNS‑rebinding protection is unaffected. A second `Referer`‑keyed form exists for root‑absolute assets and is the only exemption decided by a request‑supplied header, so it is fenced to safe methods on non‑`/api`, non‑`/ws`, non‑`/q` paths. Edges pinned by `test/webview-auth-exemption.test.ts`.
- **The lost‑frame recovery page is the third unauthenticated 200, and the only one decided by request headers alone.** The proxy's runtime shim masks `/webview/<cap>/` off the page's own URL so a single‑page app routes on the path it expects; a navigation the page then starts itself (`location.reload()`, a root‑absolute `location.href`) lands on Codeman's root with no capability anywhere, no cookie (opaque origin) and a Referer naming the masked page. `serveLostWebviewFrame()` in `middleware/auth.ts` recognises it by shape (`GET`/`HEAD`, `Sec-Fetch-Dest: iframe` or `frame`, `Accept: text/html`, `Sec-Fetch-Mode: navigate` or absent) and answers, BEFORE the credential checks and without counting an auth failure, with a static page whose only content is a `postMessage` of the lost path to the parent tab (`default-src 'none'` plus the hash of that one script, `no-store`, `referrer: no-referrer`, no reflected input). It is fenced to paths that are NOT registered routes and never `/api/`, `/ws/` or `/q/`, with one carve‑out: `/` itself, because the landing page masks to exactly `/` and its reload otherwise rendered Codeman's app shell inside the web tab. `/` is admitted only when the request carries neither the `codeman_session` cookie nor an `Authorization` header: nothing in Codeman frames its own root and a sandboxed frame has neither, while a framed `/` that does carry credentials still gets the shell. On a passwordless install no auth hook runs, so the index route applies the same test itself (`isLostWebviewRootFrame`). ⚠️ Known property, accepted rather than mitigated: those headers are trivially set by a non‑browser client, so an unauthenticated caller can distinguish a registered route (401) from a non‑route (200) and enumerate the route table; the routes are public in `docs/api-reference.md`, so nothing is learned. Pinned by `test/webview-auth-exemption.test.ts` (password) and `test/webview-lost-root-frame.test.ts` (passwordless).
- **Sandboxed by default; `allow-same-origin` is an explicit per‑dashboard opt‑in.** A proxied page is same‑origin with Codeman, so without the sandbox its JavaScript could read the Codeman document and call the agent‑spawning API. ⚠️ In BOTH modes the `Authorization` header and the `codeman_session` cookie are stripped before the upstream request, because a trusted (same‑origin) frame makes the browser attach Codeman's own Basic‑auth credentials to every proxied request; forwarding them would hand `CODEMAN_PASSWORD` to the dashboard.
- **Not an open relay, and not a privilege boundary.**`resolveUpstreamUrl()` refuses anything leaving the saved origin, and cross‑origin redirects are handed back unchanged rather than followed. The proxy does reach whatever the SERVER can reach, which is not an escalation for someone who already commands `--dangerously-skip-permissions` agents, but in multi‑user mode it means a non‑admin's dashboard is fetched from the server's network position. Saved URLs are validated to plain http(s) with no embedded credentials, and there is deliberately **no magic‑link path**: terminal output can never create a webview (the mistake the attachment scanner had to be walled off from). The one refused destination class is link‑local and cloud‑metadata addresses (`169.254.0.0/16`, `fe80::/10`, `fd00:ec2::254`, `168.63.129.16`, `100.100.100.200`, `metadata.google.internal`): `webview-egress-policy.ts` refuses them at save time, and `webview-egress.ts` re‑judges the RESOLVED address at connect time through a `lookup` hook on the proxy's undici Agent and on its WebSocket client, so a DNS name pointing into those ranges is refused as well. Loopback and RFC1918 stay allowed on purpose. Capabilities are revoked on logout, admin logout and user deletion, and proxied responses carry `Referrer-Policy: same-origin` so a dashboard cannot hand the capability‑bearing URL to a third‑party host it links.
| Precise idle detection | Yes | Codex: same screen check, via its own prompt and working line. DeepSeek: reports its state itself. Others: outputstabilization, coarser |
| Auto-resume when a usage limit resets | Yes | No |
| **Create New** | A fresh `~/codeman-cases/<name>` with a scaffolded `CLAUDE.md`. |
| **Clone Repo** | A public repo cloned into `~/codeman-cases/<name>` and registered as a case. |
| **Clone Repo** | A repo cloned into `~/codeman-cases/<name>` and registered as a case. Private repos need this machine's own git credentials (see below). |
| **Link Existing** | An existing folder anywhere on disk, registered in place. Nothing is copied or moved. |
Linked cases keep living where they are. Deleting a case in Codeman removes the
registration, and for a linked case that is all it removes.
**Clone Repo never asks for credentials.** It uses whatever the server's own git already has:
an ssh key, or a credential helper such as `gh auth setup-git`. The Docker image can include
helpers for GitHub (`gh`) and Azure DevOps (`az`), turned on in `docker-compose.override.yml`;
then signing those CLIs in once from a shell session is enough. See the private repositories
section of `docker/README.md`. Without credentials a private repo fails straight away with an
authentication error.
**Cases created from scratch are the only copy of that code.** Uninstalling Codeman does not
delete `~/codeman-cases/`, but treat that directory as real work, not scratch space.
@@ -50,10 +57,10 @@ A session carries state the case does not:
## Run mode
The **run mode** is which CLI the session runs: `claude`, `opencode`, `codex`, `gemini`,
`antigravity`, `pi`, or `shell`. It is chosen at start and does not change afterwards; to
`antigravity`, `pi`,`grok`, `deepseek`, `omp`, or `shell`. It is chosen at start and does not change afterwards; to
switch, start another session.
Claude is the reference mode. Six of the seven are not Claude, and a number of Codeman
Claude is the reference mode. Nine of the ten are not Claude, and a number of Codeman
features are Claude-only for structural reasons rather than missing effort: they depend on
Claude Code's hook system or on parsing its terminal output. Every such feature is labelled
Claude-only where it appears, and [Agent CLIs](Agent-CLIs) lists them in one place.
@@ -68,8 +75,8 @@ Where a case runs is **separate from** which CLI it runs. There are three locati
| **Docker** | One long-lived container per case; sessions `docker exec` into it. See [Docker Cases](Docker-Cases). |
| **Remote SSH** | A durable tmux server on the remote host, fronted by a local pane running `ssh`. See [Remote SSH Sessions](Remote-SSH-Sessions). |
This matters because it is a common source of confusion: Docker is **not** an eighth run
mode. All seven run modes work in all three locations. A case is docker-backed or
This matters because it is a common source of confusion: Docker is **not** an eleventh run
mode. All ten run modes work in all three locations. A case is docker-backed or
ssh-backed; a session is claude or codex or shell.
**Web tabs** are the other thing that is not a session. A saved dashboard URL renders as a
@@ -155,9 +162,11 @@ report events back: a permission prompt appeared, the turn finished, the agent w
task completed. Those events drive tab alerts, the Approvals Inbox, notifications, and the
wait primitives.
This is why some features are Claude-only. The other CLIs have no equivalent hook system,
so for them Codeman falls back to watching terminal output, which is coarser: it can see
that something happened, not what it was.
This is why some features are Claude-only. The one partial exception is DeepSeek Harness,
whose terminal front door reports idle, working and blocked to Codeman over the harness's
own supervisor contract, so it gets the hook-driven signals without a hook file. The other
CLIs have no equivalent, so for them Codeman falls back to watching terminal output, which
is coarser: it can see that something happened, not what it was.
See [Hooks And Integrations](Hooks-And-Integrations).
@@ -167,7 +176,7 @@ See [Hooks And Integrations](Hooks-And-Integrations).
- [`docs/architecture-invariants.md`](https://github.com/Ark0N/Codeman/blob/master/docs/architecture-invariants.md) - the mechanisms behind all of this, for contributors.
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
| **Codex** | Config is correct, and plain chat can work against a server that speaks the Responses API — but a real tool-call attempt comes back as inert text instead of running, so it's still not usable for real coding work. |
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
| **DeepSeek** | The original 404 is root-caused and fixed (DeepSeek Harness's own code was missing a `/v1` most local servers require) — not yet re-run against a real `dsh` install to confirm end-to-end. |
| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. |
Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry,
not a fixed list here, so this table can go stale before this page does — a greyed-out or
missing entry is the more current answer.
## What it does not do
- **No remote or Docker sessions yet.** Both restart their agent differently under the hood
(reattaching a durable tmux session rather than relaunching the process), so redirecting
them needs its own plumbing that hasn't been built.
- **No live hot-swap mid-conversation.** Applying a selection always restarts the process.
- **No button to un-point a session from the UI yet.** Clearing back to native cloud is an
HTTP call (`POST .../custom-model {"clear": true}`) or deleting the session; the settings
panel manages saved endpoints, not what a running session is currently pointed at.
- **Nothing is shared with your real cloud credentials.** The endpoint's own key, if any,
never touches your Anthropic/OpenAI/Google login — a custom endpoint is a separate,
explicit choice per session.
## Security
An endpoint's base URL can't point at a link-local or cloud-metadata address (both at save
time and against the address it actually resolves to), the same guard Web Tabs uses for
saved dashboards. Endpoint records and any per-session config files a harness needs are
@@ -122,7 +122,7 @@ codeman web # then open http://localhost:3000
| OS | macOS or Linux. Windows works through WSL2. |
| Node.js | 22 or newer. |
| tmux | Required. Sessions live in tmux, which is what makes them survive restarts. |
| An agent CLI | At least one of Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi. Plain shell sessions need none. |
| An agent CLI | At least one of Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi, Grok Build, DeepSeek Harness, OMP. Plain shell sessions need none. |
| Network | Binds to `127.0.0.1` by default. Reaching it from another device is a deliberate step: see [Remote Access](Remote-Access). |
Codeman is MIT licensed, self-hosted, and sends no telemetry. Everything runs on your
@@ -9,7 +9,7 @@ Getting Codeman onto a machine, verifying it works, updating it, and removing it
| **macOS or Linux** | Windows works through WSL2. See [Windows](#windows-wsl) below. |
| **Node.js 22+** | The installer offers to install it if missing. |
| **tmux** | Not optional. Sessions live inside tmux, which is what makes them survive a server restart, a dropped connection, or a closed laptop. |
| **An agent CLI** | At least one of [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev). Plain shell sessions need none. See [Agent CLIs](Agent-CLIs). |
| **An agent CLI** | At least one of [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), [OMP](https://github.com/can1357/oh-my-pi). Plain shell sessions need none. See [Agent CLIs](Agent-CLIs). |
Codeman itself sends no telemetry and phones no home. The only network traffic is your
browser to your server, and whatever the agent CLI you chose does on its own.
@@ -20,17 +20,26 @@ browser to your server, and whatever the agent CLI you chose does on its own.
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js and tmux if they are missing, clones Codeman into `~/.codeman/app`,
and builds it.
This installs Node.js, tmux and a build toolchain if they are missing (node-pty ships no
Linux prebuild, so it compiles from source), clones Codeman into `~/.codeman/app`, and
builds it.
What it asks you:
It starts by printing what it found (git, Node, tmux, build tools, agent CLIs, Tailscale,
an existing install), then asks everything it needs up front, then does the work
unattended. You can leave while it builds. What it asks you:
1. **Permission for every system change.**Package installs and agent CLI downloads are
prompted individually. Nothing is installed silently.
1. **One consent for the missing packages.**Git, Node.js, tmux and (on Linux) the build
toolchain are installed after a single yes, and sudo asks for your password once for
the whole run. Nothing is installed silently. If no agent CLI is found, a menu offers
to install any of them (DeepSeek excepted: its npm package installs only a launcher
with no runnable profile), or you skip and install one yourself later.
2. **How the dashboard should be reachable.** Three choices:
- **Tailscale** (recommended for phone access): keeps the loopback bind and walks you
through `tailscale serve`, including the tailnet HTTPS toggle, then verifies the result
end to end.
- **Tailscale** (recommended for phone access): keeps the loopback bind, installs
Tailscale if needed, logs in, enables the tailnet HTTPS toggle (it opens the admin
page for you and waits; Ctrl+C there skips Tailscale for this run), then configures `tailscale serve` after the build and
verifies the result end to end. If another app already owns `:443` on your node,
you choose between a sub-path (`https://<machine>.<tailnet>.ts.net/codeman`, the
default), a second port, replacing the other mapping, or skipping.
- **Your local network** (`0.0.0.0`): prompts for a password. Skipping the password takes
an explicit confirmation and ends on a loud warning.
- **This machine only** (`127.0.0.1`): the safest option, and the default for a bare
@@ -41,26 +50,54 @@ What it asks you:
Tailscale. An existing loopback install defaults to keeping loopback, or to Tailscale when
a serve mapping for Codeman is already there. A bare Enter never pulls in new software,
and a non-interactive run always keeps the safe loopback default.
3. **What to do when it finishes.** Run in this terminal, install as a background service
that starts on boot, or do nothing yet.
3. **What to call this machine on your tailnet** (Tailscale route only). By default the URL
uses the machine's existing name. Answer yes to rename it `codeman-<hostname>`; the
default is no, because the tailnet name is also what SSH and everything else on that
machine are reached by.
4. **Whether to run Codeman in the background.** Enter installs a systemd user service or a
macOS LaunchAgent that starts on boot; answering no offers to start it in this terminal
instead, or not at all.
It ends on a screen with the URL (your tailnet, your network, or this machine), a QR code to
scan with your phone, and the two commands you need to manage the service.
Re-running the same one-liner **updates an existing install in place**. Local changes in
`~/.codeman/app` are stashed rather than discarded, a running service is restarted and
verified, and your existing network binding is preserved. An interrupted first install
resumes instead of restarting.
Two other entry points exist:
Other entry points:
```bash
install.sh status # print the URLs, the QR code and the manage commands again
install.sh update # update only
install.sh uninstall # remove
install.sh uninstall # remove (offers to undo a rename it performed)
install.sh tailscale # retrofit Tailscale access onto an existing install
install.sh name [<n>] # rename this machine on your tailnet (default codeman-<hostname>)
install.sh cloudflared # install cloudflared for the in-app Cloudflare tunnel
```
**Flags** answer the questions from the command line and pipe through `bash -s --`:
Codeman drives CLIs, it does not bundle them. Install at least one:
@@ -113,6 +165,9 @@ Codeman drives CLIs, it does not bundle them. Install at least one:
| **Antigravity** | See [antigravity.google](https://antigravity.google) | Google's successor to the consumer Gemini CLI. |
| **Gemini CLI** | See [github.com/google-gemini/gemini-cli](https://github.com/google-gemini/gemini-cli) | Enterprise only since Google's June 2026 consumer cutover. |
| **Pi** | See [pi.dev](https://pi.dev) | No permission prompts and no sandbox by design. Read [Agent CLIs](Agent-CLIs) before using it on a repo you care about. |
| **DeepSeek Harness** | `npm i -g @deepseek-ai/dsh pnpm`, then a terminal profile | The npm package is only a launcher. Codeman's Run menu installs the community terminal profile for you. See [Agent CLIs](Agent-CLIs). |
| **OMP** | `curl -fsSL https://omp.sh/install \| sh` | Oh My Pi. Run it once by hand to finish its own onboarding. |
Log each CLI in once, by hand, before pointing Codeman at it. Codeman never collects or
stores your CLI credentials.
@@ -161,6 +216,7 @@ Full detail, including logs and the self-updater, is in
| Installer | Re-run the one-liner, or **App Settings → System → Updates** in the UI. |
| npm | `npm update -g aicodeman` |
| git clone | `git pull && npm install && npm run build`, then restart. |
| Docker Compose | Re-run `Start-Codeman.sh`. The in-app updater works too, and refuses a release that changes the container definition until you re-run the script. |
The in-app updater covers git-clone installs supervised by systemd or launchd. It restarts
the process that is running it, so the actual work happens in a detached script and the
@@ -30,6 +30,11 @@ Press `Ctrl+?` in the app for the same list in a floating overlay.
| `Ctrl+Shift+R` | Restore terminal size. |
| `Ctrl``+` / `Ctrl``-` | Font size. |
| `Shift+Wheel` | Scroll the local buffer, even where the wheel is forwarded to the CLI. |
| `Shift+drag` | Start a selection in a pane whose mouse events go to the CLI. |
| Right-click | Copy the selection. With nothing selected the native menu is left alone. |
| `Ctrl+Z` | Swallowed in agent sessions so a running CLI cannot be suspended. Normal job control in a shell. |
Anything you copy is cleaned on the way to the clipboard: each line loses the padding spaces a full-screen program paints across the rest of the row. Leading indentation is left exactly as it is, so indented code, a `git log` message body and `git diff` context lines paste back the way they looked on screen. An `Alt+drag` rectangular selection is copied exactly as it looks, so its columns stay lined up.
| **Create New** | Starting a fresh project. Creates `~/codeman-cases/<name>` and scaffolds a `CLAUDE.md` into it. |
| **Clone Repo** | Working on an existing public repo. Paste the URL; Codeman preflights it as you type, offers the repo's real branches and tags, and fills in the case name. |
| **Clone Repo** | Working on an existing repo: public, or private once this machine's git can authenticate (the Docker image can include `gh`/`az` helpers for this). Paste the URL; Codeman preflights it as you type, offers the repo's real branches and tags, and fills in the case name. |
| **Link Existing** | The code is already on disk. Point at the folder, with **Browse** if you would rather click than type. |
The gear next to the picker holds two per-case toggles: **Agent Teams** and
@@ -67,6 +67,9 @@ one:
| **Gemini** | Enterprise only since Google's consumer cutover. |
| **Antigravity** | Google's successor to the consumer Gemini CLI. |
| **Pi** | No permission prompts and no sandbox by design. |
| **Grok Build** | xAI's CLI. |
| **DeepSeek Harness** | Needs a terminal profile; the menu offers to install one. |
| **OMP** | Oh My Pi, configured entirely through its own `~/.omp`. |
| **Terminal / Shell** | A plain shell, no agent. Also the **Run Shell** button. |
The dropdown also lists any saved dashboard URLs ([Web Tabs](Web-Tabs)) and your recent
| The machine's existing name (default) | Nothing. This is what the installer does unless you say otherwise. |
| `https://codeman-<hostname>.<tailnet>.ts.net` | Answer yes to the installer's name question, pass `--name codeman-<hostname>`, or run `install.sh name`. This renames the machine tailnet-wide (SSH included), which is why the installer defaults to no. `install.sh uninstall` offers to rename it back. |
| `https://codeman.<tailnet>.ts.net` | A [Tailscale Service](https://tailscale.com/docs/features/tailscale-services). Only a **tagged** node can host one (a device signed in with a user account cannot), the service is defined and approved in the admin console, and the feature is in beta. The installer does not set this up; it is a `tailscale serve --service=svc:codeman --https=443 127.0.0.1:3000` on a tagged host once the service exists. |
If `:443` on your node already belongs to another app, the installer offers Codeman under
`https://<machine>.<tailnet>.ts.net/codeman` (the default, via `tailscale serve --set-path`
plus Codeman's `--base-url`), on a second port (`https://<machine>.<tailnet>.ts.net:8443`),
or replacing the other mapping. It never replaces anything without asking.
Notes:
- Keep the loopback bind. `tailscale serve` connects to `127.0.0.1:3000` locally, so
@@ -58,7 +75,11 @@ Notes:
- Codeman's Host-header allowlist already accepts `.ts.net`, so no extra configuration is
needed.
- The installer never resets or rewrites `serve` mappings other than the one pointing at
Codeman's port, so unrelated serve configuration is left alone.
Codeman's port, so unrelated serve configuration is left alone. It also never opens a
`tailscale funnel` (that is the public internet) and never advertises a Tailscale Service.
- On macOS, the App Store and standalone Tailscale apps only run once someone is logged in,
so a headless Mac needs the open-source `tailscaled` for the URL to come back after a
@@ -67,6 +67,7 @@ be wrong for at least one of them:
| **File Viewer** | Real path resolution before boundary checks, so symlinks cannot escape. Sensitive trees blocked. Edit mode adds an extension allowlist, a size cap, `.git` denial, and optimistic concurrency. It never creates files. |
| **Attachments** | An id-based registry, so browser requests never carry absolute paths. The magic-link scanner is prompt-injectable by nature and is therefore force-confined to the session's workspace. Extension allowlist, not a blocklist. |
| **Path picker** | Its own root allowlist rather than the workspace confinement. In multi-user mode a non-admin gets only their own user space, because per-user spaces live inside the home directory. |
| **Remote cases** | Reads go over the session's own ssh connection and are resolved and contained on the remote host, with a bounded number of ssh children. Nothing is copied to the Codeman host; writes, Office previews and thumbnails are refused. |
@@ -46,6 +46,8 @@ supervised by systemd or launchd; npm installs report as non-updatable. See
| Extended Keyboard Bar | Per device | Which accessory bar phones get. Shell sessions override it while they are active. |
| Wheel Scrolls Local History | Off | Keeps the wheel on the local buffer instead of forwarding it to the CLI. |
| Auto Copy Selection | Off | Copies highlighted terminal text to the clipboard the moment you finish selecting it. Ctrl+C still copies on demand. |
| Trim The Pane Margin On Copy | On | Takes the left margin a full-screen agent CLI paints down its own edge off a copy, so the text pastes flush. Each CLI declares its own width, and the strip never exceeds the indent every selected line shares, so nesting is kept. Claude Code and Codex declare a margin; a shell does not. |
| Normal / Bold font weight | xterm defaults | Per device, each slot from 100 to 900. The bundled JetBrains Mono renders every step, so a lighter normal weight makes Claude's bold headings stand out. Applies live to the terminal, both echo overlays and open team panes. |
| WebGL Renderer | On | With a GPU-stall watchdog that falls back to DOM rendering. |
| Gesture Control | Off | Camera hand tracking. Also needs `CODEMAN_GESTURE=1` on the server. |
@@ -54,12 +56,14 @@ supervised by systemd or launchd; npm installs report as non-updatable. See
Chips for every optional header control, with a live preview of the resulting header:
Run, Font Size, System Stats, Redraw Terminal, Response Viewer, Away Digest, Session
Manager, Attachments, File Viewer, Multi-monitor, Plan Usage, Lifecycle Log, Monitor,
Most default to off. The stock desktop header is system stats, File Viewer, and the gear.
New header controls never appear on phones.
New header controls never appear on phones. Split is desktop-only regardless of this
setting — the button and the feature both stay off below a ~1180px viewport, where two
resizable panes plus their divider have nowhere to go.
This section also holds background-agent tracking, including whether to track agents for
every session or only the active tab.
@@ -72,10 +76,13 @@ every session or only the active tab.
| Entrance Animations | Per-surface animation styles for tabs, terminals, windows, and lineage lines. All default to the legacy no-animation behaviour. |
| Display Name | Your name in the UI. Cosmetic only; it never renames the package, CLI, API, or storage. |
| Interface Language | English or Simplified Chinese. Per device. |
| Session List Layout | Header tab strip (default) or a collapsible left sidebar. See [The Dashboard](The-Dashboard#session-list-layout). |
| Session List Layout | Header tab strip (default), a collapsible left sidebar, or the sidebar with detailed rows. See [The Dashboard](The-Dashboard#session-list-layout). |
| Tab Orientation | Keeps the header list but turns the strip vertical beside the terminal, resizable, with detailed rows by default. Desktop and tablet only. |
| Vertical Rail Order | *By activity* (default) sorts the rail the way the home screens are sorted; *Manual* keeps your tab order and drag-reordering. |
| Tall Tabs | Taller tab strip. |
| Pop-out Button on Tabs | Adds the detach control to tabs, with a per-tab override. |
| Spawn Lineage Lines | Arcs from a parent tab to sessions it spawned. Desktop only, on by default. |
| Auto-name Sessions | Titles a new tab after its first prompt, keeping the case prefix (`w3-myapp: fix the login redirect`). Synced, off by default. See [The Dashboard](The-Dashboard#automatic-session-names). |
| Overview Home Screen | The phone home screen. On by default. |
### Models
@@ -88,6 +95,10 @@ Model and effort are both **soft defaults**: the model is written into the case'
`.claude/settings.local.json` and effort is passed at start, so `/model` and `/effort`
inside a session override them at any time.
**Custom model endpoints** (off by default) adds a saved-endpoint list plus a matching
section to the Run dropdown, for pointing a harness at your own OpenAI-compatible server
instead of its native cloud backend. See [Custom Model Endpoints](Custom-Model-Endpoints).
### Agents & CLIs
| Setting | Notes |
@@ -155,6 +166,9 @@ Some things are configured before the server starts, not in the UI:
| `CODEMAN_DOCKER_BRIDGE_HOOKS` | Lets in-container hooks reach the host on a loopback bind. |
| `CODEMAN_FILE_PICKER_ROOTS` | Extra roots for the path picker. |
| `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK` | Acknowledges exposing the server with no password. |
| `CODEMAN_BASE_URL` | Mounts Codeman under a sub-path behind a reverse proxy that forwards the prefix unchanged. See [Remote Access](Remote-Access). |
| `CODEMAN_MAX_DOWNLOAD_BYTES` | Cap on raw file bodies and downloads. 2 GB by default, `0` for none. |
| `CODEMAN_MAX_REMOTE_FILE_SSH` | Concurrent ssh reads for files in remote cases. 4 by default. |
| **Header tab strip** | The default. Wraps to a second row on desktop, scrolls sideways on a phone. |
| **Left sidebar** | A vertical list with a filter box and a live session count. `Alt+B` collapses it to a narrow rail that keeps the status dots and task badges visible. On a phone it is an off-canvas drawer rather than a docked rail. |
| **Left sidebar** | A vertical list with a filter box and a live session count. `Alt+B` collapses it to a narrow rail that keeps the status dots and task badges visible. On a phone it is an off-canvas drawer rather than a docked rail. A detailed variant adds the home screen's per-session line (`created 3d ago · working 12m`) and a status pill. |
| **Vertical rail** | The strip turned vertical beside the terminal, resizable, with detailed rows by default. **Vertical Rail Order** sorts it by activity (blocked on you first, then longest running, then most recently quiet), the same order as the home screens; pick *Manual* to get your own order and drag-reordering back. Desktop and tablet only. |
It is the same list either way, just re-hosted: tab order, drag-to-reorder, the `Alt+1`
to `Alt+9` numbers and every status colour below behave identically in both. The setting is
@@ -46,6 +48,7 @@ One tab per session, in your order, and that order syncs across your devices.
| Yellow tab, blinking | The agent is waiting for input from you. |
| Red tab, blinking | A question or permission prompt is blocking the session. |
| No dot | The session is not running. |
| Muted grey dot plus an `exited (137)` badge | The agent inside the pane has exited, with that exit code (or `exited (signal 9)`). A bare `exited` means tmux saw the pane die but did not report how, which is not the same as a clean `exited (0)`. Detailed sidebar and rail rows read `exited` in their pill. |
"description":"Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"_comment":"Copy this file to local-llm-test.config.json (gitignored) and fill in your own values. CLI flags on scripts/test-local-llm-harnesses.mjs always override these. Any field can be omitted. apiKey is OPTIONAL — omit it entirely (or delete this line) for an endpoint like llama.cpp that doesn't check one; it defaults to a harmless placeholder either way.",
note:'config STRUCTURE verified; codex only speaks the Responses API (dropped Chat-Completions Feb 2026) — expect FAIL against a plain OpenAI-compatible server, that is a protocol gap, not a bug here.',
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.